From d3380539ee54dc6099c5ebb3e56df527d684ef1f Mon Sep 17 00:00:00 2001 From: noctarius aka Christoph Engelbert Date: Tue, 15 Sep 2026 08:46:15 +0200 Subject: [PATCH 001/206] feat(operator): StorageNode moves to v1alpha2, and its operations are rebuilt (#542) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * feat(operator): StorageNode moves to v1alpha2, and its operations are rebuilt Both kinds shipped in a shape that predates the group's conventions, so this adds a v1alpha2 hub, a v1alpha1 spoke, and the conversion between them, and rewrites the operations controller against design-storagenode.md in internal/controllers/node/. What moves on the entity is §15.1. spec.storageNodeSetRef becomes spec.clusterRef plus spec.nodeSet, spec.overrides becomes spec.config and stops being an override of anything, spec.socketIndex becomes spec.slot, the failure domain becomes a label rather than an index on both the spec and the status, and the four per-node fields that reached nothing move to StorageCluster.spec.storageNodes. Most of spec.config is immutable by marker, and the three fields with exactly one legitimate writer are guarded by a webhook instead. Status gains a typed phase, a provisioning step, observedGeneration, and a device summary that is two counts rather than a string — which also corrects a rendering that reported total/online against a documented online/total. Appendix C lands with it. StorageCluster.spec.storageNodes is the workload every storage node runs as, which storagecluster_types.go recorded as deliberately absent while this kind was still v1alpha1. The operation is the larger half. It gains a status.step holding a declared statemachine graph per action, an Aborted phase and the spec.abort that reaches it, and the seventh HostMaintenance action that retires the node-drain coordinator. status.triggered goes with nothing replacing it: every step completes on a predicate over current state, and every call is skipped when its target is already at or past what the call would produce, so a step recorded without its side effect having fired is safe to re-enter. The two drain counters regroup under status.drain, where volumesTotal is written once and replaces a pending count that had to be kept in step with it. Two of the conversion's rows needed more than an assignment. spec.clusterRef is not on the stored object and a conversion has no client, so it is read from the controller owner reference the upgrade's reparent step puts there before the storage version moves. And status.subPhase is not a function of status.step: Migrating means the volume drain under Remove and the relocation restart under Migrate, and Restarting means the wait for the node under Migrate, so the action is an input to both directions — which is the defect §6.3 splits the value in two to remove. StorageNodeMetrics joins the four readings in metrics.simplyblock.io/v1alpha2, served rather than stored for the reason they are. It sits between two of them: a device's reading is one drive and a cluster's is the whole fleet, and a node's is the machine, which is the unit placement is decided against and the unit a drain moves volumes off. StorageNode.status.resources.capacity carries the same pair with hysteresis; this is the same measurement without the damping. One thing is not what the design specifies, and §15.4 records it. §8.4 fans the drain out as one PersistentVolumeOps per volume, and that kind has not been written — the StoragePool rework hit the same wall. The fan-out is the VolumeMigration that exists and works, tracked by the same label and cleaned up by the same cascade, and it becomes a PersistentVolumeOps when that kind lands. Blocking a drain that works today would have been the worse trade. One key deliberately does not move to the storage.simplyblock.io prefix. The per-slot storage-node-uuid label on a worker Node is what external-provisioner caches in CSINode, and it hard-errors CreateVolume when a live Node's topology keys do not match the cached set, so §5.2's own rule that the key must never change for a worker's lifetime outranks the prefix migration. The upgrade tool's key rewrite moves it in step with the CSI driver that reads it. Co-Authored-By: Claude Opus 5 (1M context) * feat(operator): the node domain moves to controllers/node, and StorageNodeSet retires design-crd-model.md §7.10 assigns StorageNode, StorageNodeOps, StorageDevice, StorageDeviceOps, and the storage-node workload to one package, and only the operation had moved. The rest follows here, and with it the retirement §15.3 was blocked on. The entity controller arrives against the provisioning machine of §4.2. Adding a backend node is not idempotent and the call adds every socket of a worker at once, so two objects for one worker that both observe an empty status.uuid would add it twice; the claim is made in Kubernetes first, as an optimistic-lock patch on the transition into Posting, and status.postedAt and the List over siblings that stood in for it are gone. Adoption is a branch of the same machine rather than a second path, reached from the host check by an upgrade Secret and from either gate by a backend node already at the worker's address. HostMaintenance stops being unreachable. §10 has the entity controller raise it when it sees a worker cordoned, which is what the Node watch is for, and the eight-phase drain coordinator in a fleet object's status goes with it. The workload becomes the cluster's. The DaemonSet, the two Services, the EndpointSlice, the serving certificates, the ServiceAccount and its role, and the per-node ConfigMap are children of the StorageCluster now, driven by a second controller on that kind which writes none of its status. Several sets per cluster collapse to one workload, because growth is nodes rather than sets and what differs between hardware generations is per node already. The per-node ConfigMap loses its merge with it: a node carries its whole configuration, so an entry is a rendering of one object rather than a resolution of two. skipKubeletConfiguration is written out rather than substituted. It is the one rename in the migration that also inverts, so a mechanical one would have turned kubelet configuration on for every cluster that never mentioned it. StorageDevice moves in the same change, because §7.10 puts it here and a device cannot discover itself. Its cluster label stops going through the set: a node names its own cluster, so a label on a device no longer depends on a third object being readable. The latency controller's reading moves onto the nodes. A fleet-wide list made every node's measurement a write to one object shared by all of them, and a stale snapshot of that list silently dropped entries a concurrent reconcile had written. Its own package move belongs to controllers/volume, which design-persistentvolumeops.md creates. internal/controllers/testsupport is the shared test package §7.10 names. cluster and pool each rolled a local copy of the same four helpers; the moved device suites use this one rather than adding a third. The storage-version guard is what found the rest. Marking v1alpha2 as stored put StorageNode into TestTheOperatorReadsEveryKindAtItsStoredVersion's guarded set, and every remaining v1alpha1 reader — the device mirror, the node subscription, the validating webhook, and the deployment band's worker check — now reads at the version a fresh install answers. Co-Authored-By: Claude Opus 5 (1M context) * fix(operator): regenerate the installer and clear what the linter found dist/install.yaml is generated from the same kustomize build the CRDs and the chart are, and the node domain's move changed all three. Only the first two were regenerated, so the installer still granted the retired StorageNodeSet and still withheld storagenodemetrics from the aggregated view role. The linter's five findings are the move's own loose ends. Two functions stopped being able to fail once the StorageNodeSet lookup behind them went, so both lose an error nothing could return. The metric labels a terminal phase reports under get names rather than sharing a spelling with the control plane's device status, which is a different vocabulary that happens to use one of the same words. The conversion fixtures' cluster name becomes a constant now that a third file names it. And the status-subresource constant goes with the device suites that moved out of the package. `make build`, `make lint`, and `make test` are the repository's own entry points and are what should have been run: `go build` alone does not regenerate, and it is the regeneration that was missing. * feat(operator): the StorageNode admission guard, and the sizing stamp its conversion needs §3.2 and §3.4 are the guard. It answers two different kinds of question, and the split is which of them a request can decide on its own. On update it holds the three fields with exactly one legitimate writer. spec.workerNode, spec.config.pcieAllowList, and spec.config.sizing are the operator's, so a +k8s:immutable marker would lock the operator out along with everyone else and no marker at all would let a user invalidate a layout claim by editing a string. The guard already existed for the worker; the other two are new, and the allow list is guarded rather than marked because a migration merges the drives bound on the target host into it. On create it resolves spec.clusterRef and checks the node against the cluster it names. The reference is immutable from creation, so a node naming a cluster that is not there can never be corrected and the rejection asks for the delete-and- rewrite that is the only remedy. Then the device class: a cluster is built out of one class of backend storage, because an erasure-coding stripe placed across both is written and rebuilt at the slower one's rate — so a list mixing PCI addresses with paths is refused, a list of the class the cluster is not is refused, and the PCI filters are refused on a LogicalBlock cluster rather than silently selecting nothing. And the sizing: a user's node is held to the fleet's, while the operator may write a node that differs, because a rolling hardware upgrade is exactly the case where it should. The step is what the conversion's own doc comment promised. spec.config.sizing is required on the hub and has no v1alpha1 spelling at all, so a node stored as v1alpha1 converts up without one and the next write of it is refused. A conversion cannot fill it in — it has no client and runs inside the API server's request path — so the value travels in the stash annotation ConvertTo already reads, and stamp-storage-node-sizing is what writes it. It runs after reparent-storage-nodes, because the cluster it reads is the owner that step establishes, and it refuses rather than writes when the cluster states no core count: a stamp built from nothing would be refused by the field's own minimum one step later and be harder to attribute there. One bug the tests found before the code shipped. The guard compared sizing blocks with ==, and the core count is a pointer, so two separately decoded objects held two pointers to the same number and every update read as a re-size. It refused every edit a user is entitled to make. --------- Co-authored-by: Claude Opus 5 (1M context) --- ...mplyblock.io_clusterdeploymentconfigs.yaml | 7 +- ...torage.simplyblock.io_storageclusters.yaml | 275 ++ ...storage.simplyblock.io_storagenodeops.yaml | 140 +- .../storage.simplyblock.io_storagenodes.yaml | 1145 +++++-- .../templates/roles/manager_role.yaml | 38 +- ...metrics_apiserver_aggregate_view_role.yaml | 1 + .../simplyblock-operator-webhook.yaml | 6 +- .../v1alpha2/storagenodemetrics_types.go | 129 + .../metrics/v1alpha2/zz_generated.deepcopy.go | 78 + .../metrics/v1alpha2/zz_generated.openapi.go | 174 + .../v1alpha1/controlplane_conversion_test.go | 4 + operator/api/v1alpha1/hub_roundtrip_test.go | 16 +- .../v1alpha1/storagebackup_conversion_test.go | 18 +- .../api/v1alpha1/storagenode_conversion.go | 391 +++ .../v1alpha1/storagenode_conversion_test.go | 254 ++ .../api/v1alpha1/storagenodeops_conversion.go | 260 +- .../v1alpha2/clusterdeploymentconfig_types.go | 23 +- operator/api/v1alpha2/storagecluster_types.go | 129 +- operator/api/v1alpha2/storagenode_types.go | 535 ++++ operator/api/v1alpha2/storagenodeops_types.go | 253 +- .../api/v1alpha2/zz_generated.deepcopy.go | 405 +++ operator/cmd/main.go | 71 +- operator/config/conversion-webhook/rbac.yaml | 1 + ...mplyblock.io_clusterdeploymentconfigs.yaml | 7 +- ...torage.simplyblock.io_storageclusters.yaml | 275 ++ ...storage.simplyblock.io_storagenodeops.yaml | 140 +- .../storage.simplyblock.io_storagenodes.yaml | 1145 +++++-- operator/config/crd/converted-kinds.txt | 1 + ...metrics_apiserver_aggregate_view_role.yaml | 1 + operator/config/rbac/role.yaml | 38 +- operator/config/webhook/manifests.yaml | 6 +- operator/dist/install.yaml | 1210 +++++-- .../controller/nodedrain_controller.go | 1539 --------- .../nodedrain_controller_unit_test.go | 1316 -------- .../simplyblockstoragenodeset_controller.go | 1708 ---------- ...mplyblockstoragenodeset_controller_test.go | 31 - ...lockstoragenodeset_controller_unit_test.go | 2844 ----------------- .../simplyblockstoragenodeset_drain.go | 305 -- ...mplyblockstoragenodeset_drain_unit_test.go | 311 -- ...simplyblockstoragenodeset_pernodeconfig.go | 260 -- .../simplyblockstoragenodeset_storagenode.go | 582 ---- ...ockstoragenodeset_storagenode_unit_test.go | 492 --- .../controller/storagenode_controller.go | 1206 ------- .../storagenode_controller_unit_test.go | 904 ------ .../storagenode_latency_controller.go | 217 +- .../controller/storagenodeops_controller.go | 1786 ----------- .../storagenodeops_controller_unit_test.go | 596 ---- ...storagenodeops_migrate_config_unit_test.go | 182 -- .../internal/controller/test_helpers_test.go | 4 - .../deployment/operatorops_controller.go | 3 +- .../deployment/operatorops_unit_test.go | 4 +- operator/internal/controllers/node/actions.go | 155 + .../internal/controllers/node/classify.go | 309 ++ .../internal/controllers/node/controlplane.go | 309 ++ operator/internal/controllers/node/events.go | 86 + operator/internal/controllers/node/graphs.go | 358 +++ .../internal/controllers/node/graphs_test.go | 305 ++ .../controllers/node/hostmaintenance.go | 261 ++ operator/internal/controllers/node/metrics.go | 153 + operator/internal/controllers/node/migrate.go | 298 ++ .../controllers/node/pernodeconfig.go | 279 ++ operator/internal/controllers/node/remove.go | 474 +++ .../node}/storagedevice_collector.go | 2 +- .../node}/storagedevice_collector_test.go | 5 +- .../node}/storagedevice_controller.go | 54 +- ...storagedevice_controller_reporting_test.go | 9 +- .../storagedevice_controller_unit_test.go | 23 +- .../node}/storagedevice_metrics.go | 2 +- .../node/storagenode_controller.go | 1288 ++++++++ .../node/storagenodeops_controller.go | 1025 ++++++ .../internal/controllers/node/workload.go | 519 +++ .../controllers/node/workload_controller.go | 403 +++ .../controllers/testsupport/testsupport.go | 99 + .../internal/cpinformer/subscriptions/node.go | 4 +- operator/internal/metricsapi/install.go | 14 +- operator/internal/metricsapi/nodestorage.go | 374 +++ operator/internal/metricsapi/scheme.go | 2 + operator/internal/metricsapi/server.go | 2 + operator/internal/upgrade/catalog/catalog.go | 1 + ...mplyblock.io_clusterdeploymentconfigs.yaml | 7 +- ...torage.simplyblock.io_storageclusters.yaml | 275 ++ ...storage.simplyblock.io_storagenodeops.yaml | 140 +- .../storage.simplyblock.io_storagenodes.yaml | 1145 +++++-- operator/internal/upgrade/steps/sizing.go | 246 ++ .../internal/upgrade/steps/sizing_test.go | 237 ++ operator/internal/utils/storage_node_api.go | 71 + ...nodeset_ds.go => storage_node_workload.go} | 88 +- ..._test.go => storage_node_workload_test.go} | 34 +- operator/internal/webhook/conversion.go | 1 + operator/internal/webhook/conversion_trust.go | 2 +- .../internal/webhook/storagenode_validator.go | 326 +- .../webhook/storagenode_validator_test.go | 344 +- 92 files changed, 15252 insertions(+), 15943 deletions(-) create mode 100644 operator/api/metrics/v1alpha2/storagenodemetrics_types.go create mode 100644 operator/api/v1alpha1/storagenode_conversion.go create mode 100644 operator/api/v1alpha1/storagenode_conversion_test.go create mode 100644 operator/api/v1alpha2/storagenode_types.go delete mode 100644 operator/internal/controller/nodedrain_controller.go delete mode 100644 operator/internal/controller/nodedrain_controller_unit_test.go delete mode 100644 operator/internal/controller/simplyblockstoragenodeset_controller.go delete mode 100644 operator/internal/controller/simplyblockstoragenodeset_controller_test.go delete mode 100644 operator/internal/controller/simplyblockstoragenodeset_controller_unit_test.go delete mode 100644 operator/internal/controller/simplyblockstoragenodeset_drain.go delete mode 100644 operator/internal/controller/simplyblockstoragenodeset_drain_unit_test.go delete mode 100644 operator/internal/controller/simplyblockstoragenodeset_pernodeconfig.go delete mode 100644 operator/internal/controller/simplyblockstoragenodeset_storagenode.go delete mode 100644 operator/internal/controller/simplyblockstoragenodeset_storagenode_unit_test.go delete mode 100644 operator/internal/controller/storagenode_controller.go delete mode 100644 operator/internal/controller/storagenode_controller_unit_test.go delete mode 100644 operator/internal/controller/storagenodeops_controller.go delete mode 100644 operator/internal/controller/storagenodeops_controller_unit_test.go delete mode 100644 operator/internal/controller/storagenodeops_migrate_config_unit_test.go create mode 100644 operator/internal/controllers/node/actions.go create mode 100644 operator/internal/controllers/node/classify.go create mode 100644 operator/internal/controllers/node/controlplane.go create mode 100644 operator/internal/controllers/node/events.go create mode 100644 operator/internal/controllers/node/graphs.go create mode 100644 operator/internal/controllers/node/graphs_test.go create mode 100644 operator/internal/controllers/node/hostmaintenance.go create mode 100644 operator/internal/controllers/node/metrics.go create mode 100644 operator/internal/controllers/node/migrate.go create mode 100644 operator/internal/controllers/node/pernodeconfig.go create mode 100644 operator/internal/controllers/node/remove.go rename operator/internal/{controller => controllers/node}/storagedevice_collector.go (99%) rename operator/internal/{controller => controllers/node}/storagedevice_collector_test.go (98%) rename operator/internal/{controller => controllers/node}/storagedevice_controller.go (95%) rename operator/internal/{controller => controllers/node}/storagedevice_controller_reporting_test.go (98%) rename operator/internal/{controller => controllers/node}/storagedevice_controller_unit_test.go (95%) rename operator/internal/{controller => controllers/node}/storagedevice_metrics.go (99%) create mode 100644 operator/internal/controllers/node/storagenode_controller.go create mode 100644 operator/internal/controllers/node/storagenodeops_controller.go create mode 100644 operator/internal/controllers/node/workload.go create mode 100644 operator/internal/controllers/node/workload_controller.go create mode 100644 operator/internal/controllers/testsupport/testsupport.go create mode 100644 operator/internal/metricsapi/nodestorage.go create mode 100644 operator/internal/upgrade/steps/sizing.go create mode 100644 operator/internal/upgrade/steps/sizing_test.go create mode 100644 operator/internal/utils/storage_node_api.go rename operator/internal/utils/{storage_nodeset_ds.go => storage_node_workload.go} (85%) rename operator/internal/utils/{storage_nodeset_ds_test.go => storage_node_workload_test.go} (76%) diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml index 30144c3c1..695d12254 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml @@ -259,11 +259,14 @@ spec: description: Count is the number of journal managers to configure. format: int32 + minimum: 1 type: integer percentPerDevice: - description: PercentPerDevice is the journal manager - capacity percentage per device. + description: PercentPerDevice is the share of each + device given to the journal. format: int32 + maximum: 100 + minimum: 1 type: integer type: object mgmtInterface: diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusters.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusters.yaml index 278c299a9..99cc036e1 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusters.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusters.yaml @@ -856,6 +856,281 @@ spec: x-kubernetes-validations: - message: field is immutable rule: self == oldSelf + storageNodes: + description: |- + StorageNodes is the Kubernetes workload the cluster's storage nodes run + as, and the cluster owns every object in it by controller reference: a + cluster deleted takes its DaemonSet, Services, certificate, and per-node + ConfigMap with it. One workload serves the whole cluster, because growth is + nodes rather than sets and what differs between hardware generations is per + node already. + properties: + containerResources: + description: |- + ContainerResources sets requests and limits for the storage-node container. + Unset enforces no limits. + properties: + claims: + description: |- + Claims lists the names of resources, defined in spec.resourceClaims, + that are used by this container. + + This field depends on the + DynamicResourceAllocation feature gate. + + This field is immutable. It can only be set for containers. + items: + description: ResourceClaim references one entry in PodSpec.ResourceClaims. + properties: + name: + description: |- + Name must match the name of one entry in pod.spec.resourceClaims of + the Pod where this field is used. It makes that resource available + inside a container. + type: string + request: + description: |- + Request is the name chosen for a request in the referenced claim. + If empty, everything from the claim is made available, otherwise + only the result of this request. + type: string + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + limits: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Limits describes the maximum amount of compute resources allowed. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + requests: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Requests describes the minimum amount of compute resources required. + If Requests is omitted for a container, it defaults to Limits if that is explicitly specified, + otherwise to an implementation-defined value. Requests cannot exceed Limits. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + type: object + dataInterfaces: + description: DataInterfaces are the data-plane network interfaces. + items: + type: string + type: array + enableCpuTopology: + description: EnableCpuTopology turns on topology-aware CPU assignment. + type: boolean + enableFormat4K: + description: |- + EnableFormat4K formats NVMe devices to a 4K block size where the device + supports it. + type: boolean + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + enableJournalDevice: + description: |- + EnableJournalDevice dedicates the smallest NVMe device on each node to the + journal manager, instead of carving a journal partition out of every + device. + type: boolean + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + enableKubeletConfiguration: + description: |- + EnableKubeletConfiguration lets the storage node apply the kubelet + configuration changes it needs. Off by default, which is the behavior the + retired skipKubeletConfiguration expressed by being set. + type: boolean + image: + description: |- + Image is the storage-node container image. Defaults to the ControlPlane + singleton's spec.image when unset, so a deployment states the version once. + pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ + type: string + imagePullPolicy: + default: IfNotPresent + description: ImagePullPolicy controls when that image is pulled. + enum: + - Always + - Never + - IfNotPresent + type: string + initContainerResources: + description: InitContainerResources does the same for the init container. + properties: + claims: + description: |- + Claims lists the names of resources, defined in spec.resourceClaims, + that are used by this container. + + This field depends on the + DynamicResourceAllocation feature gate. + + This field is immutable. It can only be set for containers. + items: + description: ResourceClaim references one entry in PodSpec.ResourceClaims. + properties: + name: + description: |- + Name must match the name of one entry in pod.spec.resourceClaims of + the Pod where this field is used. It makes that resource available + inside a container. + type: string + request: + description: |- + Request is the name chosen for a request in the referenced claim. + If empty, everything from the claim is made available, otherwise + only the result of this request. + type: string + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + limits: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Limits describes the maximum amount of compute resources allowed. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + requests: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Requests describes the minimum amount of compute resources required. + If Requests is omitted for a container, it defaults to Limits if that is explicitly specified, + otherwise to an implementation-defined value. Requests cannot exceed Limits. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + type: object + maxParallelNodeAdds: + default: 1 + description: |- + MaxParallelNodeAdds limits how many workers may be in the node-add process + at once, counted by distinct worker rather than by object so that a + two-socket host consumes one slot. Workers hosting a FoundationDB pod are + always sequential regardless of this value, because a node add reboots the + host and two simultaneous FoundationDB reboots reduce the control plane's + own fault tolerance. + format: int32 + minimum: 1 + type: integer + mgmtInterface: + description: MgmtInterface is the management network interface storage nodes bind. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + nodesPerSocket: + default: 1 + description: NodesPerSocket is how many storage nodes run per NUMA socket. + format: int32 + minimum: 1 + type: integer + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + openShiftCluster: + description: OpenShiftCluster states that the Kubernetes distribution is OpenShift. + type: boolean + openShiftMachineConfigPool: + default: worker + description: |- + OpenShiftMachineConfigPool names the pool generated MachineConfig objects + are labeled into. + type: string + reservedSystemCPU: + description: ReservedSystemCPU is the CPU set held back from SPDK for system workloads. + type: string + socketsToUse: + description: |- + SocketsToUse restricts deployment to selected NUMA sockets. Empty means + socket 0 alone. + items: + type: string + type: array + tolerations: + description: Tolerations are applied to the storage-node pods. + items: + description: |- + The pod this Toleration is attached to tolerates any taint that matches + the triple using the matching operator . + properties: + effect: + description: |- + Effect indicates the taint effect to match. Empty means match all taint effects. + When specified, allowed values are NoSchedule, PreferNoSchedule and NoExecute. + type: string + key: + description: |- + Key is the taint key that the toleration applies to. Empty means match all taint keys. + If the key is empty, operator must be Exists; this combination means to match all values and all keys. + type: string + operator: + description: |- + Operator represents a key's relationship to the value. + Valid operators are Exists, Equal, Lt, and Gt. Defaults to Equal. + Exists is equivalent to wildcard for value, so that a pod can + tolerate all taints of a particular category. + Lt and Gt perform numeric comparisons (requires feature gate TaintTolerationComparisonOperators). + type: string + tolerationSeconds: + description: |- + TolerationSeconds represents the period of time the toleration (which must be + of effect NoExecute, otherwise this field is ignored) tolerates the taint. By default, + it is not set, which means tolerate the taint forever (do not evict). Zero and + negative values will be treated as 0 (evict immediately) by the system. + format: int64 + type: integer + value: + description: |- + Value is the taint value the toleration matches to. + If the operator is Exists, the value should be empty, otherwise just a regular string. + type: string + type: object + type: array + ubuntuHost: + description: |- + UbuntuHost states that the worker's host OS is Ubuntu, which changes how + the node configures huge pages and the kernel modules it loads. + type: boolean + type: object + x-kubernetes-validations: + - message: field mgmtInterface is immutable once set + rule: '!has(oldSelf.mgmtInterface) || has(self.mgmtInterface)' + - message: field nodesPerSocket is immutable once set + rule: '!has(oldSelf.nodesPerSocket) || has(self.nodesPerSocket)' + - message: field enableJournalDevice is immutable once set + rule: '!has(oldSelf.enableJournalDevice) || has(self.enableJournalDevice)' + - message: field enableFormat4K is immutable once set + rule: '!has(oldSelf.enableFormat4K) || has(self.enableFormat4K)' stripe: description: |- Stripe is the erasure-coding layout every volume in the cluster is diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagenodeops.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagenodeops.yaml index a74479f86..9d63d8ae7 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagenodeops.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagenodeops.yaml @@ -195,11 +195,12 @@ spec: - jsonPath: .status.phase name: Phase type: string - - jsonPath: .status.subPhase - name: SubPhase + - jsonPath: .status.step.state + name: Step type: string - jsonPath: .status.message name: Message + priority: 1 type: string - jsonPath: .metadata.creationTimestamp name: Age @@ -208,10 +209,9 @@ spec: schema: openAPIV3Schema: description: |- - StorageNodeOps is a one-shot operational CR targeting a single StorageNode. - Analogous to a Kubernetes Job — it drives an action (Shutdown, Restart, Suspend, - Resume, Remove, Migrate) to completion and records the result. Only one - StorageNodeOps can be active per StorageNode at a time. + StorageNodeOps is a single operation performed against one StorageNode. It runs + to a terminal phase and stays afterward as the audit record of what was done, to + which node, with which parameters, and how it ended. properties: apiVersion: description: |- @@ -231,10 +231,18 @@ spec: metadata: type: object spec: - description: StorageNodeOpsSpec defines the desired state of a StorageNodeOps. + description: StorageNodeOpsSpec is one operation to perform against one StorageNode. properties: + abort: + description: |- + Abort asks a running operation to stop at its next step and unwind. It is + the only mutable field on this spec, because it is the only thing about an + operation that can legitimately be decided after it started. Whether an + abort is expressible from the current step is declared by that action's + graph rather than checked here. + type: boolean action: - description: Action is the operation to perform. Immutable. + description: Action is the operation to perform. enum: - Shutdown - Restart @@ -242,28 +250,30 @@ spec: - Resume - Remove - Migrate + - HostMaintenance type: string x-kubernetes-validations: - message: field is immutable rule: self == oldSelf force: - description: Force enables forced execution where the backend supports it. + description: |- + Force passes the control plane's force flag where the action supports it. + Migrate defaults it to true, because the control plane rejects a non-forced + restart of a node that is not already offline. type: boolean migrate: description: Migrate parameterizes action Migrate and is ignored by the others. properties: newSsdPcie: description: |- - NewSsdPcie lists additional NVMe PCIe addresses to bind on the target host - during a migration. Passed through to the control-plane restart as - new_ssd_pcie. + NewSsdPcie lists additional NVMe PCI addresses to bind on the target host, + passed through to the control-plane restart as new_ssd_pcie and merged into + the node's effective allow list so they survive a later rebuild. items: type: string type: array targetWorkerNode: - description: |- - TargetWorkerNode is the Kubernetes worker hostname the storage node is - relocated onto. + description: TargetWorkerNode is the Kubernetes worker the node is relocated onto. type: string x-kubernetes-validations: - message: field is immutable @@ -272,25 +282,28 @@ spec: - targetWorkerNode type: object nodeRef: - description: NodeRef is the name of the target StorageNode. Immutable. + description: |- + NodeRef names the StorageNode this operation acts on. The operation never + owns its target, because deleting the record of an operation must not delete + the node it operated on. type: string x-kubernetes-validations: - message: field is immutable rule: self == oldSelf reattachVolume: description: |- - ReattachVolume reattaches volumes during the node restart. - Applicable when action=Restart or action=Migrate. + ReattachVolume asks the control plane to reattach this node's volumes as + part of a restart. Applies to Restart, Migrate, and HostMaintenance. type: boolean remove: description: Remove parameterizes action Remove and is ignored by the others. properties: systemVolumeFilterRegex: + default: ^sb-fio-baseline-.* description: |- - SystemVolumeFilterRegex is a Go regular expression matched against backend - volume names. Matching volumes are treated as system volumes: excluded from - drain migration and deleted inline during the Verifying phase. - Defaults to `^sb-fio-baseline-.*`. + SystemVolumeFilterRegex matches backend volume names that are system + volumes: excluded from the drain's migration and deleted during + verification rather than blocking it. type: string type: object required: @@ -298,50 +311,79 @@ spec: - nodeRef type: object status: - description: StorageNodeOpsStatus holds the observed state of a StorageNodeOps. + description: StorageNodeOpsStatus is the observed state of one node operation. properties: completedAt: - description: CompletedAt is when the operation finished (successfully or not). + description: CompletedAt is when it reached a terminal phase. format: date-time type: string + drain: + description: |- + Drain is the drain's progress over the node's volumes, set only for action + Remove. + properties: + volumesMigrated: + description: VolumesMigrated is how many of them have completed. + format: int32 + minimum: 0 + type: integer + volumesTotal: + description: |- + VolumesTotal is the number of PV-managed volumes the drain has to move, + written once at the end of Validating and not modified afterward. + format: int32 + minimum: 0 + type: integer + required: + - volumesMigrated + - volumesTotal + type: object message: - description: Message is a human-readable description of the current state or failure reason. + description: |- + Message is the reason the phase is what it is: one sentence, replaced as the + operation moves, and never a log. type: string + observedGeneration: + description: |- + ObservedGeneration is the generation the rest of this status was computed + from, so a stale status can be told from a current one. + format: int64 + type: integer phase: - description: Phase is the high-level lifecycle phase. + description: Phase is the operation's own progress. enum: - Pending - Running - Succeeded - Failed + - Aborted type: string startedAt: - description: StartedAt is when the operation began. + description: StartedAt is when the operation acquired its target's lock. format: date-time type: string - subPhase: - description: SubPhase tracks the active drain step when action=Remove and phase=Running. - enum: - - Validating - - Suspending - - Migrating - - Verifying - - Removing - - Preparing - - Restarting - - Promoting - type: string - triggered: + step: description: |- - Triggered indicates the backend action POST has been sent (used during - Suspending to avoid duplicate POSTs across reconcile iterations). - type: boolean - volumesMigrated: - description: VolumesMigrated is the count of volumes successfully migrated (drain only). - type: integer - volumesPending: - description: VolumesPending is the count of volumes awaiting migration (drain only). - type: integer + Step is the position of the running action's state machine, as the shared + statemachine.KubeSnapshot. The rule is what an Enum marker would do if a + marker could reach a field of a shared type. + properties: + deadline: + description: |- + Deadline is when that state expires, absent when it has none. It is an + absolute instant, so a state whose deadline passed while the controller + was down restores as already expired. + format: date-time + type: string + state: + description: |- + State is the state the machine was in. Empty means the resource has not + been reconciled yet, and restores to the graph's initial state. + type: string + type: object + x-kubernetes-validations: + - message: unknown step + rule: '!has(self.state) || self.state in [''Requesting'',''Awaiting'',''Validating'',''Suspending'',''MigratingVolumes'',''Verifying'',''Removing'',''Preparing'',''Relocating'',''AwaitingNode'',''Promoting'',''Holding'',''ShuttingDown'',''Releasing'',''AwaitingHost'',''Restarting'',''Cleanup'']' type: object type: object served: true diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagenodes.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagenodes.yaml index b6a7f5f88..afcbe3faf 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagenodes.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagenodes.yaml @@ -12,339 +12,826 @@ spec: listKind: StorageNodeList plural: storagenodes shortNames: - - sn + - sn singular: storagenode scope: Namespaced versions: - - additionalPrinterColumns: - - jsonPath: .spec.workerNode - name: Worker - type: string - - jsonPath: .spec.socketId - name: Socket - type: string - - jsonPath: .spec.nodeIndex - name: NodeIdx - type: integer - - jsonPath: .status.failureDomain - name: FD - priority: 1 - type: integer - - jsonPath: .status.uuid - name: UUID - type: string - - jsonPath: .status.status - name: Status - type: string - - jsonPath: .status.health - name: Health - type: boolean - - jsonPath: .metadata.creationTimestamp - name: Age - type: date - name: v1alpha1 - schema: - openAPIV3Schema: - description: |- - StorageNode is the Schema for a single backend storage node instance. - One StorageNode CR exists per (workerNode, socketIndex) pair and is owned - by the parent StorageNodeSet. - properties: - apiVersion: - description: |- - APIVersion defines the versioned schema of this representation of an object. - Servers should convert recognized schemas to the latest internal value, and - may reject unrecognized values. - More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources - type: string - kind: - description: |- - Kind is a string value representing the REST resource this object represents. - Servers may infer this from the endpoint the client submits requests to. - Cannot be updated. - In CamelCase. - More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds - type: string - metadata: - type: object - spec: - description: StorageNodeSpec defines the desired state of a StorageNode. - properties: - nodeIndex: - description: NodeIndex is the per-socket node index (0..nodesPerSocket-1). - Immutable. - format: int32 - type: integer - x-kubernetes-validations: - - message: field is immutable - rule: self == oldSelf - overrides: - description: |- - Overrides holds per-node configuration propagated from - StorageNodeSet.spec.nodeConfigs[workerNode] on every reconcile. - properties: - deviceNames: - description: |- - DeviceNames explicitly defines the NVMe namespace names to use on this node - (e.g. ["nvme0n1","nvme1n1"]). - items: + - additionalPrinterColumns: + - jsonPath: .spec.workerNode + name: Worker + type: string + - jsonPath: .spec.socketId + name: Socket + type: string + - jsonPath: .spec.nodeIndex + name: NodeIdx + type: integer + - jsonPath: .status.failureDomain + name: FD + priority: 1 + type: integer + - jsonPath: .status.uuid + name: UUID + type: string + - jsonPath: .status.status + name: Status + type: string + - jsonPath: .status.health + name: Health + type: boolean + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha1 + schema: + openAPIV3Schema: + description: |- + StorageNode is the Schema for a single backend storage node instance. + One StorageNode CR exists per (workerNode, socketIndex) pair and is owned + by the parent StorageNodeSet. + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: StorageNodeSpec defines the desired state of a StorageNode. + properties: + nodeIndex: + description: NodeIndex is the per-socket node index (0..nodesPerSocket-1). Immutable. + format: int32 + type: integer + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + overrides: + description: |- + Overrides holds per-node configuration propagated from + StorageNodeSet.spec.nodeConfigs[workerNode] on every reconcile. + properties: + deviceNames: + description: |- + DeviceNames explicitly defines the NVMe namespace names to use on this node + (e.g. ["nvme0n1","nvme1n1"]). + items: + type: string + type: array + driveSizeRange: + description: DriveSizeRange overrides the drive size range filter for this node. + type: string + enableCpuTopology: + description: EnableCpuTopology overrides topology-aware CPU handling for this node. + type: boolean + expand: + description: |- + Expand marks this node as a cluster-expansion add. When true the backend + node-add endpoint receives expand=true, triggering rebalancing behaviour + appropriate for in-place cluster growth. Overrides StorageNodeSet.spec.expand. + type: boolean + failureDomain: + description: |- + FailureDomain is the failure-domain group index (≥ 0) for this node. + Required when the parent StorageCluster has enableFailureDomains=true. + Overrides StorageNodeSet.spec.nodeFailureDomains[workerNode] when both are set. + format: int32 + minimum: 0 + type: integer + journalManager: + description: JournalManagerSpec overrides journal manager tuning for this node. + properties: + count: + description: Count is the number of journal managers to configure. + format: int32 + type: integer + percentPerDevice: + description: PercentPerDevice is the journal manager capacity percentage per device. + format: int32 + type: integer + type: object + pcieAllowList: + description: PcieAllowList overrides the list of PCI addresses allowed for use on this node. + items: + type: string + type: array + pcieDenyList: + description: PcieDenyList overrides the list of PCI addresses excluded from use on this node. + items: + type: string + type: array + pcieModel: + description: PcieModel overrides the PCI model filter for this node. + type: string + reservedSystemCPU: + description: ReservedSystemCPU overrides the CPUs reserved for system workloads on this node. + type: string + skipKubeletConfiguration: + description: |- + SkipKubeletConfiguration overrides whether kubelet configuration changes are + skipped for this node. + type: boolean + spdkImage: + description: SpdkImage overrides the SPDK image for this node (e.g. for phased rollouts). + type: string + spdkProxyImage: + description: SpdkProxyImage overrides the SPDK proxy image for this node. type: string - type: array - driveSizeRange: - description: DriveSizeRange overrides the drive size range filter - for this node. - type: string - enableCpuTopology: - description: EnableCpuTopology overrides topology-aware CPU handling - for this node. - type: boolean - expand: - description: |- - Expand marks this node as a cluster-expansion add. When true the backend - node-add endpoint receives expand=true, triggering rebalancing behaviour - appropriate for in-place cluster growth. Overrides StorageNodeSet.spec.expand. - type: boolean - failureDomain: - description: |- - FailureDomain is the failure-domain group index (≥ 0) for this node. - Required when the parent StorageCluster has enableFailureDomains=true. - Overrides StorageNodeSet.spec.nodeFailureDomains[workerNode] when both are set. - format: int32 - minimum: 0 - type: integer - journalManager: - description: JournalManagerSpec overrides journal manager tuning - for this node. - properties: - count: - description: Count is the number of journal managers to configure. - format: int32 - type: integer - percentPerDevice: - description: PercentPerDevice is the journal manager capacity - percentage per device. - format: int32 - type: integer - type: object - pcieAllowList: - description: PcieAllowList overrides the list of PCI addresses - allowed for use on this node. - items: + spdkSystemMemory: + description: |- + SpdkSystemMemory overrides the SPDK huge-page memory allocation for this node + (e.g. "4G", "512M"). + pattern: ^[0-9]+(G|GI|GB|GiB|M|MI|MB|MiB|g|gi|gb|gib|m|mi|mb|mib)?$ type: string - type: array - pcieDenyList: - description: PcieDenyList overrides the list of PCI addresses - excluded from use on this node. - items: + ubuntuHost: + description: UbuntuHost overrides the Ubuntu host OS flag for this node. + type: boolean + type: object + socketId: + description: SocketID is the NUMA socket identifier from spec.socketsToUse (e.g. "0", "1"). Immutable. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + socketIndex: + description: |- + SocketIndex is the global ordinal (socketPosition × nodesPerSocket + nodeIndex). + Used internally by the operator to select the correct backend node from the + RPC-port-sorted list in pollUUIDFromBackend. Immutable. + format: int32 + type: integer + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + storageNodeSetRef: + description: StorageNodeSetRef is the name of the owning StorageNodeSet. Immutable. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + workerNode: + description: |- + WorkerNode is the Kubernetes node hostname this StorageNode runs on. + Users may not change it directly — it is re-pointed only by the operator + during a node migration (StorageNodeOps action=migrate). The + StorageNode validating webhook rejects user-driven changes to this field. + type: string + required: + - storageNodeSetRef + - workerNode + type: object + x-kubernetes-validations: + - message: field socketId is immutable once set + rule: '!has(oldSelf.socketId) || has(self.socketId)' + - message: field nodeIndex is immutable once set + rule: '!has(oldSelf.nodeIndex) || has(self.nodeIndex)' + - message: field socketIndex is immutable once set + rule: '!has(oldSelf.socketIndex) || has(self.socketIndex)' + status: + description: StorageNodeStatus holds the observed state of a StorageNode. + properties: + activeOpsRef: + description: |- + ActiveOpsRef is the name of the currently active StorageNodeOps CR targeting + this node. Empty when no operation is in progress. Used for mutual exclusion. + type: string + failureDomain: + description: |- + FailureDomain is the effective failure-domain group index for this node + as reported by the backend (≥ 0). Nil when the backend has not assigned one. + format: int32 + type: integer + health: + description: Health is the backend-reported node health flag. + type: boolean + hostname: + description: Hostname is the node hostname as reported by the backend. + type: string + latencyMetrics: + description: |- + LatencyMetrics holds the fio-measured baseline NVMe-oF latency for this node, + used by the volume rebalancer to make data-placement decisions. + properties: + baselineMeasuredAt: + description: BaselineMeasuredAt is when the baseline was established. + format: date-time type: string - type: array - pcieModel: - description: PcieModel overrides the PCI model filter for this - node. - type: string - reservedSystemCPU: - description: ReservedSystemCPU overrides the CPUs reserved for - system workloads on this node. - type: string - skipKubeletConfiguration: - description: |- - SkipKubeletConfiguration overrides whether kubelet configuration changes are - skipped for this node. - type: boolean - spdkImage: - description: SpdkImage overrides the SPDK image for this node - (e.g. for phased rollouts). - type: string - spdkProxyImage: - description: SpdkProxyImage overrides the SPDK proxy image for - this node. - type: string - spdkSystemMemory: - description: |- - SpdkSystemMemory overrides the SPDK huge-page memory allocation for this node - (e.g. "4G", "512M"). - pattern: ^[0-9]+(G|GI|GB|GiB|M|MI|MB|MiB|g|gi|gb|gib|m|mi|mb|mib)?$ - type: string - ubuntuHost: - description: UbuntuHost overrides the Ubuntu host OS flag for - this node. - type: boolean - type: object - socketId: - description: SocketID is the NUMA socket identifier from spec.socketsToUse - (e.g. "0", "1"). Immutable. - type: string - x-kubernetes-validations: - - message: field is immutable - rule: self == oldSelf - socketIndex: - description: |- - SocketIndex is the global ordinal (socketPosition × nodesPerSocket + nodeIndex). - Used internally by the operator to select the correct backend node from the - RPC-port-sorted list in pollUUIDFromBackend. Immutable. - format: int32 - type: integer - x-kubernetes-validations: - - message: field is immutable - rule: self == oldSelf - storageNodeSetRef: - description: StorageNodeSetRef is the name of the owning StorageNodeSet. - Immutable. - type: string - x-kubernetes-validations: - - message: field is immutable - rule: self == oldSelf - workerNode: - description: |- - WorkerNode is the Kubernetes node hostname this StorageNode runs on. - Users may not change it directly — it is re-pointed only by the operator - during a node migration (StorageNodeOps action=migrate). The - StorageNode validating webhook rejects user-driven changes to this field. - type: string - required: - - storageNodeSetRef - - workerNode - type: object - x-kubernetes-validations: - - message: field socketId is immutable once set - rule: '!has(oldSelf.socketId) || has(self.socketId)' - - message: field nodeIndex is immutable once set - rule: '!has(oldSelf.nodeIndex) || has(self.nodeIndex)' - - message: field socketIndex is immutable once set - rule: '!has(oldSelf.socketIndex) || has(self.socketIndex)' - status: - description: StorageNodeStatus holds the observed state of a StorageNode. - properties: - activeOpsRef: - description: |- - ActiveOpsRef is the name of the currently active StorageNodeOps CR targeting - this node. Empty when no operation is in progress. Used for mutual exclusion. - type: string - failureDomain: - description: |- - FailureDomain is the effective failure-domain group index for this node - as reported by the backend (≥ 0). Nil when the backend has not assigned one. - format: int32 - type: integer - health: - description: Health is the backend-reported node health flag. - type: boolean - hostname: - description: Hostname is the node hostname as reported by the backend. - type: string - latencyMetrics: - description: |- - LatencyMetrics holds the fio-measured baseline NVMe-oF latency for this node, - used by the volume rebalancer to make data-placement decisions. - properties: - baselineMeasuredAt: - description: BaselineMeasuredAt is when the baseline was established. - format: date-time - type: string - baselineP50NS: - description: BaselineP50NS is the p50 write latency (nanoseconds) - from the initial empty-cluster benchmark. - format: int64 - type: integer - baselineP99NS: - description: BaselineP99NS is the p99 write latency (nanoseconds) - from the initial empty-cluster benchmark. - format: int64 - type: integer - nodeUUID: - description: NodeUUID is the backend storage node UUID. - type: string - required: - - nodeUUID - type: object - ports: - description: Ports groups network connectivity fields (addresses and - ports). - properties: - lvol: - description: Lvol is the logical-volume subsystem port. - format: int32 - type: integer - management: - description: Management is the management IP address of the node. - type: string - nvmeof: - description: NvmeOf is the NVMe-oF fabric port. - format: int32 - type: integer - rpc: - description: Rpc is the RPC/management API port. - format: int32 - type: integer - type: object - postedAt: - description: |- - PostedAt is the timestamp when the node-add POST was sent. - Used as a provisioning guard against duplicate POSTs. - format: date-time - type: string - resources: - description: Resources groups compute and storage resource metrics. - properties: - capacity: - description: |- - Capacity is how much of the node's storage is in use, summed over its - devices. It is a measurement rather than a declaration, so it is absent - until something has measured it, and it lags reality by the interval at - which the control plane's metrics are scraped. - properties: - sampledAt: - description: |- - SampledAt is when the control plane took the reading. It is not when the - object was written, and it may be considerably older if metrics - collection has stopped. - format: date-time + baselineP50NS: + description: BaselineP50NS is the p50 write latency (nanoseconds) from the initial empty-cluster benchmark. + format: int64 + type: integer + baselineP99NS: + description: BaselineP99NS is the p99 write latency (nanoseconds) from the initial empty-cluster benchmark. + format: int64 + type: integer + nodeUUID: + description: NodeUUID is the backend storage node UUID. + type: string + required: + - nodeUUID + type: object + ports: + description: Ports groups network connectivity fields (addresses and ports). + properties: + lvol: + description: Lvol is the logical-volume subsystem port. + format: int32 + type: integer + management: + description: Management is the management IP address of the node. + type: string + nvmeof: + description: NvmeOf is the NVMe-oF fabric port. + format: int32 + type: integer + rpc: + description: Rpc is the RPC/management API port. + format: int32 + type: integer + type: object + postedAt: + description: |- + PostedAt is the timestamp when the node-add POST was sent. + Used as a provisioning guard against duplicate POSTs. + format: date-time + type: string + resources: + description: Resources groups compute and storage resource metrics. + properties: + capacity: + description: |- + Capacity is how much of the node's storage is in use, summed over its + devices. It is a measurement rather than a declaration, so it is absent + until something has measured it, and it lags reality by the interval at + which the control plane's metrics are scraped. + properties: + sampledAt: + description: |- + SampledAt is when the control plane took the reading. It is not when the + object was written, and it may be considerably older if metrics + collection has stopped. + format: date-time + type: string + totalBytes: + description: TotalBytes is the storage the node's devices provide. + format: int64 + minimum: 0 + type: integer + usedBytes: + description: UsedBytes is what they currently hold. + format: int64 + minimum: 0 + type: integer + type: object + cpu: + description: CPU is the number of SPDK CPU cores allocated to this node. + format: int32 + type: integer + devices: + description: Devices is the device summary (online/total) reported by the backend. + type: string + memory: + description: Memory is the SPDK memory allocation reported by the backend. + type: string + volumes: + description: Volumes is the current number of logical volumes on this node. + format: int32 + type: integer + type: object + status: + description: Status is the backend-reported node status (e.g. online, suspended, offline). + type: string + uptime: + description: Uptime is the node uptime as reported by the backend. + type: string + uuid: + description: UUID is the backend storage node UUID. Set once after node-add completes. + type: string + type: object + type: object + served: true + storage: false + subresources: + status: {} + - additionalPrinterColumns: + - jsonPath: .spec.clusterRef + name: Cluster + type: string + - jsonPath: .spec.workerNode + name: Worker + type: string + - jsonPath: .spec.socketId + name: Socket + type: string + - jsonPath: .spec.slot + name: Slot + type: integer + - jsonPath: .status.phase + name: Phase + type: string + - jsonPath: .status.step.state + name: Step + type: string + - jsonPath: .status.status + name: Status + type: string + - jsonPath: .status.health + name: Health + type: boolean + - jsonPath: .status.uuid + name: UUID + priority: 1 + type: string + - jsonPath: .status.failureDomain + name: FD + priority: 1 + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha2 + schema: + openAPIV3Schema: + description: |- + StorageNode is one backend storage node: one SPDK process bound to one NUMA + socket of one Kubernetes worker. One object exists per (workerNode, slot) pair, + owned by the StorageCluster it belongs to. + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: |- + StorageNodeSpec is the desired state of one backend storage node, meaning one + SPDK process bound to one NUMA socket of one Kubernetes worker. + properties: + clusterRef: + description: |- + ClusterRef names the StorageCluster this node belongs to. The cluster also + owns this object by controller reference, so deleting the cluster deletes + its nodes. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + config: + description: |- + Config is this node's complete configuration, copied from the + ClusterDeploymentConfig entry that produced it. It is a copy rather than a + projection, because that document is ephemeral: nothing reads it once the + node exists, deleting it changes nothing, and editing it reaches only nodes + created afterward. + properties: + deviceNames: + description: |- + DeviceNames names the devices to use. An entry is a PCI address + ("0000:5e:00.0") or a device path ("/dev/sdb,") which are the two classes + simplyblock accepts as backend storage, and a bare name ("nvme0n1") is read + as a path under /dev. One list carries both spellings, and every entry is of + the class its cluster declares in StorageCluster.spec.deviceClass: a list + mixing the two, or naming the class the cluster is not, is rejected by the + StorageNode validating webhook. Set explicitly, it overrides every filter + below. Immutable: it selects which physical devices the node owns. + items: + pattern: ^([0-9a-fA-F]{4}:[0-9a-fA-F]{2}:[0-9a-fA-F]{2}\.[0-9a-fA-F]|/dev/[a-zA-Z0-9._/-]+|[a-zA-Z0-9._-]+)$ type: string - totalBytes: - description: TotalBytes is the storage the node's devices - provide. - format: int64 - minimum: 0 - type: integer - usedBytes: - description: UsedBytes is what they currently hold. - format: int64 - minimum: 0 - type: integer - type: object - cpu: - description: CPU is the number of SPDK CPU cores allocated to - this node. - format: int32 - type: integer - devices: - description: Devices is the device summary (online/total) reported - by the backend. - type: string - memory: - description: Memory is the SPDK memory allocation reported by - the backend. - type: string - volumes: - description: Volumes is the current number of logical volumes - on this node. - format: int32 - type: integer - type: object - status: - description: Status is the backend-reported node status (e.g. online, - suspended, offline). - type: string - uptime: - description: Uptime is the node uptime as reported by the backend. - type: string - uuid: - description: UUID is the backend storage node UUID. Set once after - node-add completes. - type: string - type: object - type: object - served: true - storage: true - subresources: - status: {} + type: array + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + driveSizeRange: + description: DriveSizeRange filters devices by size. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + expand: + description: |- + Expand marks this node as an addition to an already-active cluster, which + the control plane reads as a request to rebalance onto it rather than to + treat it as part of an initial layout. Immutable once set: it describes how + the node joined rather than what it is. + type: boolean + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + failureDomain: + description: |- + FailureDomain is the label of the fault group this node belongs to + such as rack-b, naming the physical grouping it shares with its peers rather + than indexing it. Required when the cluster has enableFailureDomains set, + and provisioning is held with a FailureDomainMissing event until it is + present. Immutable once set, which is what makes it fillable later and then + frozen: chunk placement was computed from it. + + The value takes the shape of a Kubernetes label value, because that is what + it is seeded from where a cluster carries topology labels at all. + maxLength: 63 + pattern: ^[a-zA-Z0-9]([-_.a-zA-Z0-9]*[a-zA-Z0-9])?$ + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + journalManager: + description: |- + JournalManager tunes the journal manager count and per-device capacity + share for this node. Immutable: both are on-disk layout, fixed when the + devices were partitioned. + properties: + count: + description: Count is the number of journal managers to configure. + format: int32 + minimum: 1 + type: integer + percentPerDevice: + description: PercentPerDevice is the share of each device given to the journal. + format: int32 + maximum: 100 + minimum: 1 + type: integer + type: object + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + pcieAllowList: + description: |- + PcieAllowList selects devices by PCI address. It is the one device field a + migration writes, merging spec.migrate.newSsdPcie into it so devices added + on the target host survive a later rebuild, so it is guarded by the + StorageNode validating webhook rather than by a marker. This and the two + PCI filters below belong to an NVMe cluster: the webhook rejects them on a + cluster whose deviceClass is LogicalBlock, because a logical block device + has no PCI address to match. + items: + type: string + type: array + pcieDenyList: + description: PcieDenyList excludes devices by PCI address. + items: + type: string + type: array + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + pcieModel: + description: PcieModel filters devices by PCI model string. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + sizing: + description: |- + Sizing is what this node's huge pages and core layout were sized from. + Writable by the operator alone. + properties: + minHugePagesSize: + description: |- + MinHugePagesSize is the smallest huge-page allocation this node makes, as a + size string such as 100G or 1T, where a bare number is gigabytes. It is a + floor rather than a limit: the effective allocation is the larger of this + value and the minimum the node's device and subsystem count requires. + type: string + vcpuCount: + description: |- + VCPUCount is the number of vCPUs allocated to SPDK on this node, as an + explicit core count rather than a percentage. + format: int32 + minimum: 4 + type: integer + required: + - vcpuCount + type: object + spdkImage: + description: |- + SpdkImage overrides the SPDK image the control plane starts for this node, + which is what makes a phased image rollout expressible per node. + type: string + spdkProxyImage: + description: SpdkProxyImage overrides the SPDK proxy image for this node. + type: string + spdkSystemMemory: + description: |- + SpdkSystemMemory is the memory the control plane starts this node's SPDK + with, as a size string such as 4G or 512M. Mutable: a node whose device + count grew legitimately needs to raise it. + pattern: ^[0-9]+(G|GI|GB|GiB|M|MI|MB|MiB|g|gi|gb|gib|m|mi|mb|mib)?$ + type: string + required: + - sizing + type: object + x-kubernetes-validations: + - message: field journalManager is immutable once set + rule: '!has(oldSelf.journalManager) || has(self.journalManager)' + - message: field deviceNames is immutable once set + rule: '!has(oldSelf.deviceNames) || has(self.deviceNames)' + - message: field pcieDenyList is immutable once set + rule: '!has(oldSelf.pcieDenyList) || has(self.pcieDenyList)' + - message: field pcieModel is immutable once set + rule: '!has(oldSelf.pcieModel) || has(self.pcieModel)' + - message: field driveSizeRange is immutable once set + rule: '!has(oldSelf.driveSizeRange) || has(self.driveSizeRange)' + - message: field failureDomain is immutable once set + rule: '!has(oldSelf.failureDomain) || has(self.failureDomain)' + - message: field expand is immutable once set + rule: '!has(oldSelf.expand) || has(self.expand)' + nodeIndex: + description: |- + NodeIndex is the position among the nodes sharing this socket, in + 0..nodesPerSocket-1. See SocketID. + format: int32 + minimum: 0 + type: integer + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + nodeSet: + description: |- + NodeSet is the name of the group in ClusterDeploymentConfig.nodeSets[] this + node was declared under. It is a label rather than a reference: nothing is + fetched by it, and it exists so that a node can be traced back to the + document that produced it. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + slot: + description: |- + Slot is which storage-node slot on this worker the object occupies, counted + from zero. A worker runs one node per socket per nodesPerSocket, and the + slot is the position among them. It is the identity the operator keys on: + the topology label the CSI driver reads is + storage.simplyblock.io/storage-node-uuid.., and the slot + outlives the node filling it, because only the UUID behind it changes when a + node is replaced or relocated. + format: int32 + minimum: 0 + type: integer + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + socketId: + description: |- + SocketID is the NUMA socket this node is bound to, as declared in the node + set's socket list, so 0 or 1. With NodeIndex it decomposes Slot into the + pair a person reads; nothing but a print column consumes either. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + workerNode: + description: |- + WorkerNode is the Kubernetes worker hostname this node runs on. It is not + marked immutable, because a migration re-points it, but the StorageNode + validating webhook rejects any change made by an identity outside the + operator's namespace. + type: string + required: + - clusterRef + - config + - workerNode + type: object + x-kubernetes-validations: + - message: field nodeSet is immutable once set + rule: '!has(oldSelf.nodeSet) || has(self.nodeSet)' + - message: field socketId is immutable once set + rule: '!has(oldSelf.socketId) || has(self.socketId)' + - message: field nodeIndex is immutable once set + rule: '!has(oldSelf.nodeIndex) || has(self.nodeIndex)' + - message: field slot is immutable once set + rule: '!has(oldSelf.slot) || has(self.slot)' + status: + description: StorageNodeStatus is the observed state of one storage node. + properties: + activeOpsRef: + description: |- + ActiveOpsRef names the StorageNodeOps currently allowed to touch this node. + Empty when none is running. + type: string + failureDomain: + description: |- + FailureDomain is the failure-domain label the control plane actually + assigned, which is not necessarily the one spec.config.failureDomain + requested. + type: string + health: + description: Health is the health flag the control plane reports. + type: boolean + hostname: + description: Hostname is the node hostname as the control plane reports it. + type: string + latencyMetrics: + description: |- + LatencyMetrics holds the fio-measured NVMe-oF baseline the volume + rebalancer reads. + properties: + baselineMeasuredAt: + description: BaselineMeasuredAt is when the baseline was established. + format: date-time + type: string + baselineP50NS: + description: |- + BaselineP50NS is the p50 write latency, in nanoseconds, of the initial + empty-cluster benchmark. + format: int64 + minimum: 0 + type: integer + baselineP99NS: + description: |- + BaselineP99NS is the p99 write latency, in nanoseconds, of the same + benchmark. + format: int64 + minimum: 0 + type: integer + nodeUUID: + description: |- + NodeUUID is the backend storage node the reading was taken against. It is + carried beside the reading rather than inferred from status.uuid, because a + baseline measured against one backend node stops describing the slot once a + replacement fills it. + type: string + required: + - nodeUUID + type: object + message: + description: |- + Message is the reason the phase is what it is: one sentence, replaced as the + node moves, and never a log. + type: string + observedGeneration: + description: |- + ObservedGeneration is the generation the rest of this status was computed + from, so a stale status can be told from a current one. + format: int64 + type: integer + phase: + description: |- + Phase is the operator's own view of this node, and the field its + provisioning branches on. + enum: + - Pending + - Provisioning + - Online + - Removing + - Offline + - Degraded + - Failed + type: string + ports: + description: Ports groups the reported addresses and ports. + properties: + lvol: + description: Lvol is the logical-volume subsystem port. + format: int32 + type: integer + management: + description: Management is the management IP address of the node. + type: string + nvmeof: + description: The NVMe-oF fabric port. + format: int32 + type: integer + rpc: + description: Rpc is the RPC and management API port. + format: int32 + type: integer + type: object + resources: + description: Resources groups the reported compute and storage figures. + properties: + capacity: + description: |- + Capacity is how much of the node's storage is in use, summed over its + devices. It is a measurement rather than a declaration, so it is absent + until something has measured it, and it lags reality by the interval at + which the control plane's metrics are scraped. + properties: + sampledAt: + description: |- + SampledAt is when the control plane took the reading. It is not when the + object was written, and it may be considerably older if metrics collection + has stopped. + format: date-time + type: string + totalBytes: + description: TotalBytes is the storage the node's devices provide. + format: int64 + minimum: 0 + type: integer + usedBytes: + description: UsedBytes is what they currently hold. + format: int64 + minimum: 0 + type: integer + type: object + cpu: + description: CPU is the number of SPDK cores allocated to this node. + format: int32 + type: integer + devices: + description: |- + Devices summarizes the node's NVMe devices. Absent until the control plane + has reported, which is what tells a node that has not reported from one that + genuinely has no devices. + properties: + online: + description: |- + Online is how many of the node's devices the control plane reports as + usable. + format: int32 + minimum: 0 + type: integer + total: + description: Total is how many devices the node has. + format: int32 + minimum: 0 + type: integer + required: + - online + - total + type: object + memory: + description: Memory is the SPDK memory allocation the control plane reports. + type: string + volumes: + description: Volumes is the current number of logical volumes on this node. + format: int32 + type: integer + type: object + status: + description: |- + Status is the lifecycle the control plane reports: online, suspended, + offline, in_creation, in_restart, in_shutdown, unreachable, or timeout. The + values are the control plane's, which is why they are neither PascalCase nor + constrained by an Enum here. + type: string + step: + description: |- + Step is the position of the provisioning machine, as the shared + statemachine.KubeSnapshot. The rule is what an Enum marker would do if a + marker could reach a field of a shared type. + properties: + deadline: + description: |- + Deadline is when that state expires, absent when it has none. It is an + absolute instant, so a state whose deadline passed while the controller + was down restores as already expired. + format: date-time + type: string + state: + description: |- + State is the state the machine was in. Empty means the resource has not + been reconciled yet, and restores to the graph's initial state. + type: string + type: object + x-kubernetes-validations: + - message: unknown step + rule: '!has(self.state) || self.state in [''CheckingHost'',''CheckingConfig'',''AwaitingSlot'',''Posting'',''Resolving'',''Adopting'']' + uptime: + description: Uptime is the node uptime as the control plane reports it. + type: string + uuid: + description: |- + UUID is the backend node UUID. Empty means the node has neither been + provisioned nor adopted, and non-empty means steady state. + type: string + type: object + type: object + served: true + storage: true + subresources: + status: {} + conversion: + strategy: Webhook + webhook: + conversionReviewVersions: + - v1 + clientConfig: + service: + namespace: simplyblock-operator-system + name: simplyblock-operator-conversion-webhook-service + path: /convert diff --git a/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml b/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml index 6757f6fb8..e335a7885 100644 --- a/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml @@ -19,16 +19,6 @@ rules: - patch - update - watch -- apiGroups: - - "" - resources: - - events - verbs: - - create - - get - - list - - patch - - watch - apiGroups: - "" resources: @@ -74,9 +64,15 @@ rules: - delete - get - list - - patch - - update - watch +- apiGroups: + - "" + - events.k8s.io + resources: + - events + verbs: + - create + - patch - apiGroups: - admissionregistration.k8s.io resources: @@ -103,6 +99,7 @@ rules: - storageclusterops.storage.simplyblock.io - storageclusters.storage.simplyblock.io - storagenodeops.storage.simplyblock.io + - storagenodes.storage.simplyblock.io - storagepools.storage.simplyblock.io resources: - customresourcedefinitions @@ -173,13 +170,6 @@ rules: - patch - update - watch -- apiGroups: - - events.k8s.io - resources: - - events - verbs: - - create - - patch - apiGroups: - policy resources: @@ -262,7 +252,6 @@ rules: - storagedevices - storagenodeops - storagenodes - - storagenodesets - storagepoolops - storagepools - tasks @@ -293,7 +282,6 @@ rules: - storageclusters/finalizers - storagenodeops/finalizers - storagenodes/finalizers - - storagenodesets/finalizers - storagepoolops/finalizers - storagepools/finalizers - tasks/finalizers @@ -342,3 +330,11 @@ rules: - patch - update - watch +- apiGroups: + - storage.simplyblock.io + resources: + - storagenodesets + verbs: + - get + - list + - watch diff --git a/helm-charts/charts/simplyblock-operator/templates/roles/metrics_apiserver_aggregate_view_role.yaml b/helm-charts/charts/simplyblock-operator/templates/roles/metrics_apiserver_aggregate_view_role.yaml index af7e52983..71e7ffce1 100644 --- a/helm-charts/charts/simplyblock-operator/templates/roles/metrics_apiserver_aggregate_view_role.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/roles/metrics_apiserver_aggregate_view_role.yaml @@ -35,6 +35,7 @@ rules: - storagedevicemetrics - storagepoolmetrics - storageclustermetrics + - storagenodemetrics verbs: - get - list diff --git a/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml b/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml index 659888c7a..ed983e159 100644 --- a/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml @@ -209,15 +209,17 @@ webhooks: service: name: simplyblock-operator-webhook-service namespace: {{ .Release.Namespace }} - path: /validate-storage-simplyblock-io-v1alpha1-storagenode + path: /validate-storage-simplyblock-io-v1alpha2-storagenode failurePolicy: Fail + matchPolicy: Equivalent name: vstoragenode.simplyblock.io rules: - apiGroups: - storage.simplyblock.io apiVersions: - - v1alpha1 + - v1alpha2 operations: + - CREATE - UPDATE resources: - storagenodes diff --git a/operator/api/metrics/v1alpha2/storagenodemetrics_types.go b/operator/api/metrics/v1alpha2/storagenodemetrics_types.go new file mode 100644 index 000000000..d4dc9619f --- /dev/null +++ b/operator/api/metrics/v1alpha2/storagenodemetrics_types.go @@ -0,0 +1,129 @@ +// StorageNodeMetrics: how much of a storage node is in use, as the control +// plane's exporter last measured it. +// +// The kind exists for the reason the four readings beside it do, and it sits +// between two of them. A device's reading is one drive and a cluster's is the +// whole fleet; a node's is the machine, which is the unit placement is decided +// against and the unit a drain moves volumes off. StorageNode.status.resources +// .capacity carries the same pair, and it carries it with hysteresis: it is +// written only when the figure has moved by a percent of the node's own total, +// because a status rewritten on every sample wakes every watcher of the kind for +// a number nothing reconciles toward (design-crd-model.md §7.13). This is the +// same measurement without that damping, computed when a client asks and never +// persisted. +// +// The consequences are the same as the other four kinds' and are load-bearing +// rather than incidental: +// +// - There is no spec and no status. The fields sit at the top level of the +// object, because neither half of the split means anything for a reading +// nobody wrote. +// - An object exists only while a StorageNode does. There is no deletion and +// no tombstone: a node that goes away stops being listed. +// - A deployment with no reachable Prometheus serves no readings at all, +// rather than readings of zero. A zero is the reading of an empty node and +// not the absence of a reading. +// +// Which device or which volume is filling a node up is deliberately not here. +// A node reports its own totals, and the breakdown behind them is +// StorageDeviceMetrics and LogicalVolumeMetrics, which is what keeps the +// cardinality of the workload out of this list. +// +// design-storagenode.md §3.3 and §12 are the specification. + +package v1alpha2 + +import ( + "k8s.io/apimachinery/pkg/api/resource" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" +) + +// StorageNodeCapacity is what one storage node holds. Every size is in bytes and +// is quoted as a resource.Quantity so that kubectl prints it the way it prints a +// PersistentVolumeClaim's capacity. +// +// It carries the same fields a device's reading does, because a node's total is +// the sum of its devices' and a reader comparing the two should not have to +// reconcile different shapes. +// +// +k8s:openapi-gen=true +type StorageNodeCapacity struct { + // Total is the space the node's devices provide, as the exporter measured + // it. + Total resource.Quantity `json:"total"` + // Used is the space they currently hold. + Used resource.Quantity `json:"used"` + // Free is the node's unallocated remainder as the control plane accounts for + // it. It is reported rather than derived, so it need not equal Total minus + // Used. + Free resource.Quantity `json:"free"` + // Provisioned is the space promised out of the node, which on a + // thin-provisioned pool may exceed Total. + Provisioned resource.Quantity `json:"provisioned"` + // UtilizationPercent is the control plane's own utilization figure, from 0 to + // 100. It is taken verbatim rather than recomputed from Used and Total, so + // that it agrees with what the control plane's own interfaces report. + UtilizationPercent int32 `json:"utilizationPercent"` +} + +// +kubebuilder:object:root=true + +// StorageNodeMetrics is one storage node's capacity reading. +// +// The object is named after the StorageNode it measures and lives in that +// object's namespace, so somebody who has the node's name needs to learn nothing +// else to ask for it, and ordinary namespaced RBAC confines a reader to the +// namespaces they already have. A backend node with no StorageNode object is +// therefore not listed: it has no name in this API and no namespace to be +// authorized against. +// +// +k8s:openapi-gen=true +type StorageNodeMetrics struct { + metav1.TypeMeta `json:",inline"` + + // metadata is standard object metadata. Name and namespace are the + // StorageNode's. The creationTimestamp is the node object's rather than the + // reading's. + // +optional + metav1.ObjectMeta `json:"metadata,omitzero"` + + // Timestamp is when the control plane sampled these values, which is older + // than the moment the request was served and may be considerably older if its + // exporter has stopped being scraped. It is the zero time when the node has + // never been sampled, so that "never measured" does not read as "measured in + // 1970." + Timestamp metav1.Time `json:"timestamp"` + + // NodeID is the control plane's identifier for the backend node. It is the + // join key back to the control plane's own exporter and to its API, and it is + // the field that changes when a slot is refilled by a replacement node while + // the object's name stays where it was. + NodeID string `json:"nodeID"` + + // StorageCluster is the name of the StorageCluster object the node belongs + // to, so a reading says which cluster it is about without a second lookup. + StorageCluster string `json:"storageCluster"` + + // WorkerNode is the Kubernetes worker the node runs on, which is what a + // reader correlating a full node with a machine is actually looking for. + WorkerNode string `json:"workerNode,omitempty"` + + // Capacity is the reading itself. + Capacity StorageNodeCapacity `json:"capacity"` +} + +// +kubebuilder:object:root=true + +// StorageNodeMetricsList is a list of readings. It carries no continue token: the +// whole set is served from memory in one pass, so there is nothing to page +// through. +// +// +k8s:openapi-gen=true +type StorageNodeMetricsList struct { + metav1.TypeMeta `json:",inline"` + // The tag is omitempty rather than the omitzero the CRD kinds in this + // repository use, because openapi-gen enforces the streaming-list convention + // on a type it generates definitions for and that convention names omitempty. + metav1.ListMeta `json:"metadata,omitempty"` + Items []StorageNodeMetrics `json:"items"` +} diff --git a/operator/api/metrics/v1alpha2/zz_generated.deepcopy.go b/operator/api/metrics/v1alpha2/zz_generated.deepcopy.go index ebdad1606..66cf962cc 100644 --- a/operator/api/metrics/v1alpha2/zz_generated.deepcopy.go +++ b/operator/api/metrics/v1alpha2/zz_generated.deepcopy.go @@ -258,6 +258,84 @@ func (in *StorageDeviceMetricsList) DeepCopyObject() runtime.Object { return nil } +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *StorageNodeCapacity) DeepCopyInto(out *StorageNodeCapacity) { + *out = *in + out.Total = in.Total.DeepCopy() + out.Used = in.Used.DeepCopy() + out.Free = in.Free.DeepCopy() + out.Provisioned = in.Provisioned.DeepCopy() +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new StorageNodeCapacity. +func (in *StorageNodeCapacity) DeepCopy() *StorageNodeCapacity { + if in == nil { + return nil + } + out := new(StorageNodeCapacity) + in.DeepCopyInto(out) + return out +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *StorageNodeMetrics) DeepCopyInto(out *StorageNodeMetrics) { + *out = *in + out.TypeMeta = in.TypeMeta + in.ObjectMeta.DeepCopyInto(&out.ObjectMeta) + in.Timestamp.DeepCopyInto(&out.Timestamp) + in.Capacity.DeepCopyInto(&out.Capacity) +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new StorageNodeMetrics. +func (in *StorageNodeMetrics) DeepCopy() *StorageNodeMetrics { + if in == nil { + return nil + } + out := new(StorageNodeMetrics) + in.DeepCopyInto(out) + return out +} + +// DeepCopyObject is an autogenerated deepcopy function, copying the receiver, creating a new runtime.Object. +func (in *StorageNodeMetrics) DeepCopyObject() runtime.Object { + if c := in.DeepCopy(); c != nil { + return c + } + return nil +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *StorageNodeMetricsList) DeepCopyInto(out *StorageNodeMetricsList) { + *out = *in + out.TypeMeta = in.TypeMeta + in.ListMeta.DeepCopyInto(&out.ListMeta) + if in.Items != nil { + in, out := &in.Items, &out.Items + *out = make([]StorageNodeMetrics, len(*in)) + for i := range *in { + (*in)[i].DeepCopyInto(&(*out)[i]) + } + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new StorageNodeMetricsList. +func (in *StorageNodeMetricsList) DeepCopy() *StorageNodeMetricsList { + if in == nil { + return nil + } + out := new(StorageNodeMetricsList) + in.DeepCopyInto(out) + return out +} + +// DeepCopyObject is an autogenerated deepcopy function, copying the receiver, creating a new runtime.Object. +func (in *StorageNodeMetricsList) DeepCopyObject() runtime.Object { + if c := in.DeepCopy(); c != nil { + return c + } + return nil +} + // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *StoragePoolCapacity) DeepCopyInto(out *StoragePoolCapacity) { *out = *in diff --git a/operator/api/metrics/v1alpha2/zz_generated.openapi.go b/operator/api/metrics/v1alpha2/zz_generated.openapi.go index 875640f02..1bfeace48 100644 --- a/operator/api/metrics/v1alpha2/zz_generated.openapi.go +++ b/operator/api/metrics/v1alpha2/zz_generated.openapi.go @@ -40,6 +40,9 @@ func GetOpenAPIDefinitions(ref common.ReferenceCallback) map[string]common.OpenA "github.com/simplyblock/simplyblock-operator/api/metrics/v1alpha2.StorageDeviceCapacity": schema_simplyblock_operator_api_metrics_v1alpha2_StorageDeviceCapacity(ref), "github.com/simplyblock/simplyblock-operator/api/metrics/v1alpha2.StorageDeviceMetrics": schema_simplyblock_operator_api_metrics_v1alpha2_StorageDeviceMetrics(ref), "github.com/simplyblock/simplyblock-operator/api/metrics/v1alpha2.StorageDeviceMetricsList": schema_simplyblock_operator_api_metrics_v1alpha2_StorageDeviceMetricsList(ref), + "github.com/simplyblock/simplyblock-operator/api/metrics/v1alpha2.StorageNodeCapacity": schema_simplyblock_operator_api_metrics_v1alpha2_StorageNodeCapacity(ref), + "github.com/simplyblock/simplyblock-operator/api/metrics/v1alpha2.StorageNodeMetrics": schema_simplyblock_operator_api_metrics_v1alpha2_StorageNodeMetrics(ref), + "github.com/simplyblock/simplyblock-operator/api/metrics/v1alpha2.StorageNodeMetricsList": schema_simplyblock_operator_api_metrics_v1alpha2_StorageNodeMetricsList(ref), "github.com/simplyblock/simplyblock-operator/api/metrics/v1alpha2.StoragePoolCapacity": schema_simplyblock_operator_api_metrics_v1alpha2_StoragePoolCapacity(ref), "github.com/simplyblock/simplyblock-operator/api/metrics/v1alpha2.StoragePoolMetrics": schema_simplyblock_operator_api_metrics_v1alpha2_StoragePoolMetrics(ref), "github.com/simplyblock/simplyblock-operator/api/metrics/v1alpha2.StoragePoolMetricsList": schema_simplyblock_operator_api_metrics_v1alpha2_StoragePoolMetricsList(ref), @@ -600,6 +603,177 @@ func schema_simplyblock_operator_api_metrics_v1alpha2_StorageDeviceMetricsList(r } } +func schema_simplyblock_operator_api_metrics_v1alpha2_StorageNodeCapacity(ref common.ReferenceCallback) common.OpenAPIDefinition { + return common.OpenAPIDefinition{ + Schema: spec.Schema{ + SchemaProps: spec.SchemaProps{ + Description: "StorageNodeCapacity is what one storage node holds. Every size is in bytes and is quoted as a resource.Quantity so that kubectl prints it the way it prints a PersistentVolumeClaim's capacity.\n\nIt carries the same fields a device's reading does, because a node's total is the sum of its devices' and a reader comparing the two should not have to reconcile different shapes.", + Type: []string{"object"}, + Properties: map[string]spec.Schema{ + "total": { + SchemaProps: spec.SchemaProps{ + Description: "Total is the space the node's devices provide, as the exporter measured it.", + Ref: ref(resource.Quantity{}.OpenAPIModelName()), + }, + }, + "used": { + SchemaProps: spec.SchemaProps{ + Description: "Used is the space they currently hold.", + Ref: ref(resource.Quantity{}.OpenAPIModelName()), + }, + }, + "free": { + SchemaProps: spec.SchemaProps{ + Description: "Free is the node's unallocated remainder as the control plane accounts for it. It is reported rather than derived, so it need not equal Total minus Used.", + Ref: ref(resource.Quantity{}.OpenAPIModelName()), + }, + }, + "provisioned": { + SchemaProps: spec.SchemaProps{ + Description: "Provisioned is the space promised out of the node, which on a thin-provisioned pool may exceed Total.", + Ref: ref(resource.Quantity{}.OpenAPIModelName()), + }, + }, + "utilizationPercent": { + SchemaProps: spec.SchemaProps{ + Description: "UtilizationPercent is the control plane's own utilization figure, from 0 to 100. It is taken verbatim rather than recomputed from Used and Total, so that it agrees with what the control plane's own interfaces report.", + Default: 0, + Type: []string{"integer"}, + Format: "int32", + }, + }, + }, + Required: []string{"total", "used", "free", "provisioned", "utilizationPercent"}, + }, + }, + Dependencies: []string{ + resource.Quantity{}.OpenAPIModelName()}, + } +} + +func schema_simplyblock_operator_api_metrics_v1alpha2_StorageNodeMetrics(ref common.ReferenceCallback) common.OpenAPIDefinition { + return common.OpenAPIDefinition{ + Schema: spec.Schema{ + SchemaProps: spec.SchemaProps{ + Description: "StorageNodeMetrics is one storage node's capacity reading.\n\nThe object is named after the StorageNode it measures and lives in that object's namespace, so somebody who has the node's name needs to learn nothing else to ask for it, and ordinary namespaced RBAC confines a reader to the namespaces they already have. A backend node with no StorageNode object is therefore not listed: it has no name in this API and no namespace to be authorized against.", + Type: []string{"object"}, + Properties: map[string]spec.Schema{ + "kind": { + SchemaProps: spec.SchemaProps{ + Description: "Kind is a string value representing the REST resource this object represents. Servers may infer this from the endpoint the client submits requests to. Cannot be updated. In CamelCase. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds", + Type: []string{"string"}, + Format: "", + }, + }, + "apiVersion": { + SchemaProps: spec.SchemaProps{ + Description: "APIVersion defines the versioned schema of this representation of an object. Servers should convert recognized schemas to the latest internal value, and may reject unrecognized values. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources", + Type: []string{"string"}, + Format: "", + }, + }, + "metadata": { + SchemaProps: spec.SchemaProps{ + Description: "metadata is standard object metadata. Name and namespace are the StorageNode's. The creationTimestamp is the node object's rather than the reading's.", + Default: map[string]interface{}{}, + Ref: ref(v1.ObjectMeta{}.OpenAPIModelName()), + }, + }, + "timestamp": { + SchemaProps: spec.SchemaProps{ + Description: "Timestamp is when the control plane sampled these values, which is older than the moment the request was served and may be considerably older if its exporter has stopped being scraped. It is the zero time when the node has never been sampled, so that \"never measured\" does not read as \"measured in 1970.\"", + Ref: ref(v1.Time{}.OpenAPIModelName()), + }, + }, + "nodeID": { + SchemaProps: spec.SchemaProps{ + Description: "NodeID is the control plane's identifier for the backend node. It is the join key back to the control plane's own exporter and to its API, and it is the field that changes when a slot is refilled by a replacement node while the object's name stays where it was.", + Default: "", + Type: []string{"string"}, + Format: "", + }, + }, + "storageCluster": { + SchemaProps: spec.SchemaProps{ + Description: "StorageCluster is the name of the StorageCluster object the node belongs to, so a reading says which cluster it is about without a second lookup.", + Default: "", + Type: []string{"string"}, + Format: "", + }, + }, + "workerNode": { + SchemaProps: spec.SchemaProps{ + Description: "WorkerNode is the Kubernetes worker the node runs on, which is what a reader correlating a full node with a machine is actually looking for.", + Type: []string{"string"}, + Format: "", + }, + }, + "capacity": { + SchemaProps: spec.SchemaProps{ + Description: "Capacity is the reading itself.", + Default: map[string]interface{}{}, + Ref: ref("github.com/simplyblock/simplyblock-operator/api/metrics/v1alpha2.StorageNodeCapacity"), + }, + }, + }, + Required: []string{"timestamp", "nodeID", "storageCluster", "capacity"}, + }, + }, + Dependencies: []string{ + "github.com/simplyblock/simplyblock-operator/api/metrics/v1alpha2.StorageNodeCapacity", v1.ObjectMeta{}.OpenAPIModelName(), v1.Time{}.OpenAPIModelName()}, + } +} + +func schema_simplyblock_operator_api_metrics_v1alpha2_StorageNodeMetricsList(ref common.ReferenceCallback) common.OpenAPIDefinition { + return common.OpenAPIDefinition{ + Schema: spec.Schema{ + SchemaProps: spec.SchemaProps{ + Description: "StorageNodeMetricsList is a list of readings. It carries no continue token: the whole set is served from memory in one pass, so there is nothing to page through.", + Type: []string{"object"}, + Properties: map[string]spec.Schema{ + "kind": { + SchemaProps: spec.SchemaProps{ + Description: "Kind is a string value representing the REST resource this object represents. Servers may infer this from the endpoint the client submits requests to. Cannot be updated. In CamelCase. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds", + Type: []string{"string"}, + Format: "", + }, + }, + "apiVersion": { + SchemaProps: spec.SchemaProps{ + Description: "APIVersion defines the versioned schema of this representation of an object. Servers should convert recognized schemas to the latest internal value, and may reject unrecognized values. More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources", + Type: []string{"string"}, + Format: "", + }, + }, + "metadata": { + SchemaProps: spec.SchemaProps{ + Description: "The tag is omitempty rather than the omitzero the CRD kinds in this repository use, because openapi-gen enforces the streaming-list convention on a type it generates definitions for and that convention names omitempty.", + Default: map[string]interface{}{}, + Ref: ref(v1.ListMeta{}.OpenAPIModelName()), + }, + }, + "items": { + SchemaProps: spec.SchemaProps{ + Type: []string{"array"}, + Items: &spec.SchemaOrArray{ + Schema: &spec.Schema{ + SchemaProps: spec.SchemaProps{ + Default: map[string]interface{}{}, + Ref: ref("github.com/simplyblock/simplyblock-operator/api/metrics/v1alpha2.StorageNodeMetrics"), + }, + }, + }, + }, + }, + }, + Required: []string{"items"}, + }, + }, + Dependencies: []string{ + "github.com/simplyblock/simplyblock-operator/api/metrics/v1alpha2.StorageNodeMetrics", v1.ListMeta{}.OpenAPIModelName()}, + } +} + func schema_simplyblock_operator_api_metrics_v1alpha2_StoragePoolCapacity(ref common.ReferenceCallback) common.OpenAPIDefinition { return common.OpenAPIDefinition{ Schema: spec.Schema{ diff --git a/operator/api/v1alpha1/controlplane_conversion_test.go b/operator/api/v1alpha1/controlplane_conversion_test.go index 639dc53c4..b3e7cb01d 100644 --- a/operator/api/v1alpha1/controlplane_conversion_test.go +++ b/operator/api/v1alpha1/controlplane_conversion_test.go @@ -22,6 +22,10 @@ const testSystemVolumeFilter = "^system-.*" const testImage = "quay.io/simplyblock-io/simplyblock:26.2.2" +// testCluster is the cluster name every conversion fixture names, so the +// literal appears once rather than in each of them. +const testCluster = "production" + func TestControlPlaneConvertToRegroupsImage(t *testing.T) { src := &ControlPlane{ ObjectMeta: metav1.ObjectMeta{Name: "simplyblock", Namespace: "sb"}, diff --git a/operator/api/v1alpha1/hub_roundtrip_test.go b/operator/api/v1alpha1/hub_roundtrip_test.go index 5a4717aeb..3481cf559 100644 --- a/operator/api/v1alpha1/hub_roundtrip_test.go +++ b/operator/api/v1alpha1/hub_roundtrip_test.go @@ -206,13 +206,15 @@ func TestStorageNodeOpsRoundTripsFromTheHub(t *testing.T) { Remove: &v1alpha2.RemoveSpec{SystemVolumeFilterRegex: &filter}, }, Status: v1alpha2.StorageNodeOpsStatus{ - Phase: v1alpha2.StorageNodeOpsPhaseRunning, - SubPhase: v1alpha2.StorageNodeOpsSubPhaseRestarting, - Message: "waiting for node-1", - VolumesMigrated: 7, - VolumesPending: 3, - Triggered: true, - StartedAt: &started, + Phase: v1alpha2.StorageNodeOpsPhaseRunning, + // AwaitingNode is the sharper half of the pair this version spells as + // one Restarting, so it is the value that proves the stash carries + // what the projection cannot. + Step: statemachine.KubeSnapshot{State: string(v1alpha2.StorageNodeOpsStepAwaitingNode)}, + Message: "waiting for node-1", + Drain: &v1alpha2.DrainStatus{VolumesTotal: 10, VolumesMigrated: 7}, + ObservedGeneration: 3, + StartedAt: &started, }, } diff --git a/operator/api/v1alpha1/storagebackup_conversion_test.go b/operator/api/v1alpha1/storagebackup_conversion_test.go index 6341455ad..f694ca7d1 100644 --- a/operator/api/v1alpha1/storagebackup_conversion_test.go +++ b/operator/api/v1alpha1/storagebackup_conversion_test.go @@ -32,7 +32,7 @@ func TestStorageBackupConvertToRegroupsTheStatus(t *testing.T) { src := &StorageBackup{ ObjectMeta: metav1.ObjectMeta{Name: "backup-1", Namespace: "sb"}, Spec: StorageBackupSpec{ - ClusterName: "production", + ClusterName: testCluster, PVCRef: &PersistentVolumeClaimRef{Name: "claim-1", Namespace: "apps"}, }, Status: StorageBackupStatus{ @@ -64,8 +64,8 @@ func TestStorageBackupConvertToRegroupsTheStatus(t *testing.T) { t.Fatalf("ConvertTo: %v", err) } - if got := dst.Spec.ClusterRef; got != "production" { - t.Errorf("spec.clusterRef = %q, want %q", got, "production") + if got := dst.Spec.ClusterRef; got != testCluster { + t.Errorf("spec.clusterRef = %q, want %q", got, testCluster) } // The store's identifier is the object's identity in the hub, and v1alpha1 // only ever held it in status. @@ -111,7 +111,7 @@ func TestStorageBackupConvertToRegroupsTheStatus(t *testing.T) { // each would be two objects of nothing that every reader then has to check // past. func TestStorageBackupConvertToLeavesEmptyGroupsAbsent(t *testing.T) { - src := &StorageBackup{Spec: StorageBackupSpec{ClusterName: "production"}} + src := &StorageBackup{Spec: StorageBackupSpec{ClusterName: testCluster}} var dst v1alpha2.StorageBackup if err := src.ConvertTo(&dst); err != nil { @@ -127,7 +127,7 @@ func TestStorageBackupConvertToLeavesEmptyGroupsAbsent(t *testing.T) { func TestStorageBackupConvertFromFlattensTheStatus(t *testing.T) { src := &v1alpha2.StorageBackup{ - Spec: v1alpha2.StorageBackupSpec{ClusterRef: "production", BackupID: "backup-uuid"}, + Spec: v1alpha2.StorageBackupSpec{ClusterRef: testCluster, BackupID: "backup-uuid"}, Status: v1alpha2.StorageBackupStatus{ Phase: v1alpha2.StorageBackupPhaseCreating, ClusterID: "cluster-uuid", @@ -141,8 +141,8 @@ func TestStorageBackupConvertFromFlattensTheStatus(t *testing.T) { t.Fatalf("ConvertFrom: %v", err) } - if got := dst.Spec.ClusterName; got != "production" { - t.Errorf("spec.clusterName = %q, want %q", got, "production") + if got := dst.Spec.ClusterName; got != testCluster { + t.Errorf("spec.clusterName = %q, want %q", got, testCluster) } if dst.Spec.PVCRef == nil || dst.Spec.PVCRef.Name != "claim-1" { t.Errorf("spec.pvcRef = %+v, want claim-1 in apps", dst.Spec.PVCRef) @@ -317,7 +317,7 @@ func TestStorageBackupKeepsTheRequestApartFromWhatHappened(t *testing.T) { // for values that exist, and writing an empty one would put a key on every // object of the kind for a fact none of them has. func TestStorageBackupStashesNothingForAnEmptyObject(t *testing.T) { - obj := &StorageBackup{Spec: StorageBackupSpec{ClusterName: "production"}} + obj := &StorageBackup{Spec: StorageBackupSpec{ClusterName: testCluster}} var hub v1alpha2.StorageBackup if err := obj.ConvertTo(&hub); err != nil { @@ -367,7 +367,7 @@ func storedBackup() *StorageBackup { return &StorageBackup{ ObjectMeta: metav1.ObjectMeta{Name: "backup-1", Namespace: "sb"}, Spec: StorageBackupSpec{ - ClusterName: "production", + ClusterName: testCluster, PVCRef: &PersistentVolumeClaimRef{Name: "claim-1", Namespace: "apps"}, SnapshotName: "snap-1", SourceClusterUUID: "source-cluster-uuid", diff --git a/operator/api/v1alpha1/storagenode_conversion.go b/operator/api/v1alpha1/storagenode_conversion.go new file mode 100644 index 000000000..a7f6d685c --- /dev/null +++ b/operator/api/v1alpha1/storagenode_conversion.go @@ -0,0 +1,391 @@ +// Conversion of StorageNode between this version and the v1alpha2 hub. +// +// design-storagenode.md §15.1 is the delta, and four of its rows need more than +// an assignment. +// +// - The parent. spec.storageNodeSetRef names a StorageNodeSet the redesign +// retires; the hub names its StorageCluster and keeps the set's name as +// spec.nodeSet, a label nothing is fetched by. The cluster is not on the +// stored object, and a conversion has no client to look one up with, so it is +// read from the controller owner reference. That is a pure function of the +// subject, and the upgrade's reparent-storage-nodes step puts the reference +// there before the storage version moves (design-api-upgrade.md §20), which +// is the ordering that makes the read answer. +// - The sizing. spec.config.sizing is required on the hub and has no v1alpha1 +// spelling at all, because the two values lived on the StorageCluster. A node +// the hub never wrote converts up with none, and the upgrade's +// stamp-storage-node-sizing step is what fills it from the cluster before +// anything writes the object again. +// - The failure domain. It was an index on both the spec and the status and is +// a label on both here, so the digits are what a mechanical conversion can +// carry and the meaning that lived outside the API is not. A label that is not +// a number has nowhere to go on the way down and is stashed whole. +// - The device summary. status.resources.devices was one string and is two +// counts. The string is read as total/online rather than as the online/total +// its own doc comment claimed, because both call sites passed the device +// count before the online count: a node with three of four devices online +// stored 4/3. The conversion reads what was written rather than what was +// documented, since only one of those is in etcd. +// +// Four spec fields travel the other way. ubuntuHost, skipKubeletConfiguration, +// enableCpuTopology, and reservedSystemCPU are declared per node here and reach +// nothing, because their only consumers are environment variables in a DaemonSet +// pod template and a DaemonSet is one object for every node it schedules. They +// move to StorageCluster.spec.storageNodes (§5.1) and stash on the way up under +// storage.simplyblock.io/v1alpha1-, so a node converted back down carries +// what it carried. status.postedAt stashes for the same reason: the persisted +// step replaced it (§3.3). + +package v1alpha1 + +import ( + "strconv" + "strings" + + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "sigs.k8s.io/controller-runtime/pkg/conversion" + + "github.com/simplyblock/atlas/statemachine" + + "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// clusterKind is the kind an owner reference carries when it is the parent the +// hub names, and it is spelled here rather than derived so that the conversion +// package needs no scheme. +const clusterKind = "StorageCluster" + +// The annotations holding the five fields this version has and the hub removed. +const ( + annoV1Alpha1NodeUbuntuHost = "storage.simplyblock.io/v1alpha1-spec.overrides.ubuntuHost" + annoV1Alpha1NodeSkipKubelet = "storage.simplyblock.io/v1alpha1-spec.overrides.skipKubeletConfiguration" + annoV1Alpha1NodeCPUTopology = "storage.simplyblock.io/v1alpha1-spec.overrides.enableCpuTopology" + annoV1Alpha1NodeReservedCPUs = "storage.simplyblock.io/v1alpha1-spec.overrides.reservedSystemCPU" + annoV1Alpha1NodePostedAt = "storage.simplyblock.io/v1alpha1-status.postedAt" +) + +// The annotations holding the hub fields this version cannot express. +const ( + annoNodeCluster = "storage.simplyblock.io/conversion-spec.clusterRef" + annoNodeSizing = "storage.simplyblock.io/conversion-spec.config.sizing" + annoNodeSpecDomain = "storage.simplyblock.io/conversion-spec.config.failureDomain" + annoNodeStatDomain = "storage.simplyblock.io/conversion-status.failureDomain" + annoNodeStep = "storage.simplyblock.io/conversion-status.step" + annoNodePhase = "storage.simplyblock.io/conversion-status.phase" + annoNodeMessage = "storage.simplyblock.io/conversion-status.message" + annoNodeObserved = "storage.simplyblock.io/conversion-status.observedGeneration" +) + +// ConvertTo converts this StorageNode to the v1alpha2 hub. +func (src *StorageNode) ConvertTo(dstRaw conversion.Hub) error { + dst := dstRaw.(*v1alpha2.StorageNode) + + dst.ObjectMeta = *src.ObjectMeta.DeepCopy() + + overrides := src.Spec.Overrides + if overrides == nil { + overrides = &StorageNodeOverrides{} + } + stashRemovedBool(&dst.ObjectMeta, annoV1Alpha1NodeUbuntuHost, overrides.UbuntuHost) + stashRemovedBool(&dst.ObjectMeta, annoV1Alpha1NodeSkipKubelet, overrides.SkipKubeletConfiguration) + stashRemovedBool(&dst.ObjectMeta, annoV1Alpha1NodeCPUTopology, overrides.EnableCpuTopology) + stashRemoved(&dst.ObjectMeta, annoV1Alpha1NodeReservedCPUs, overrides.ReservedSystemCPU) + if err := stash(&dst.ObjectMeta, annoV1Alpha1NodePostedAt, src.Status.PostedAt); err != nil { + return err + } + + dst.Spec = v1alpha2.StorageNodeSpec{ + ClusterRef: controllingClusterName(&src.ObjectMeta), + NodeSet: src.Spec.StorageNodeSetRef, + WorkerNode: src.Spec.WorkerNode, + SocketID: src.Spec.SocketID, + NodeIndex: src.Spec.NodeIndex, + Slot: src.Spec.SocketIndex, + Config: v1alpha2.StorageNodeConfig{ + SpdkImage: overrides.SpdkImage, + SpdkProxyImage: overrides.SpdkProxyImage, + SpdkSystemMemory: overrides.SpdkSystemMemory, + PcieAllowList: overrides.PcieAllowList, + PcieDenyList: overrides.PcieDenyList, + PcieModel: overrides.PcieModel, + DriveSizeRange: overrides.DriveSizeRange, + DeviceNames: overrides.DeviceNames, + FailureDomain: formatInt32(overrides.FailureDomain), + Expand: overrides.Expand, + }, + } + if jm := overrides.JournalManagerSpec; jm != nil { + dst.Spec.Config.JournalManager = &v1alpha2.JournalManagerSpec{ + Count: jm.Count, + PercentPerDevice: jm.PercentPerDevice, + } + } + + dst.Status = v1alpha2.StorageNodeStatus{ + UUID: src.Status.UUID, + Status: src.Status.Status, + Health: src.Status.Health, + Hostname: src.Status.Hostname, + Uptime: src.Status.Uptime, + FailureDomain: formatInt32(src.Status.FailureDomain), + ActiveOpsRef: src.Status.ActiveOpsRef, + } + if lm := src.Status.LatencyMetrics; lm != nil { + dst.Status.LatencyMetrics = &v1alpha2.NodeLatencyMetrics{ + NodeUUID: lm.NodeUUID, + BaselineP50NS: lm.BaselineP50NS, + BaselineP99NS: lm.BaselineP99NS, + BaselineMeasuredAt: lm.BaselineMeasuredAt, + } + } + if res := src.Status.Resources; res != nil { + dst.Status.Resources = &v1alpha2.StorageNodeResources{ + CPU: res.CPU, + Memory: res.Memory, + Volumes: res.Volumes, + Devices: parseDeviceSummary(res.Devices), + } + if cap := res.Capacity; cap != nil { + dst.Status.Resources.Capacity = &v1alpha2.StorageNodeCapacity{ + TotalBytes: cap.TotalBytes, + UsedBytes: cap.UsedBytes, + SampledAt: cap.SampledAt, + } + } + } + if ports := src.Status.Ports; ports != nil { + dst.Status.Ports = &v1alpha2.StorageNodePorts{ + Management: ports.Management, + NvmeOf: ports.NvmeOf, + Lvol: ports.Lvol, + Rpc: ports.Rpc, + } + } + + return restoreNodeHubOnly(&dst.ObjectMeta, dst) +} + +// ConvertFrom converts the v1alpha2 hub into this StorageNode. +func (dst *StorageNode) ConvertFrom(srcRaw conversion.Hub) error { + src := srcRaw.(*v1alpha2.StorageNode) + + dst.ObjectMeta = *src.ObjectMeta.DeepCopy() + + config := src.Spec.Config + dst.Spec = StorageNodeSpec{ + StorageNodeSetRef: src.Spec.NodeSet, + WorkerNode: src.Spec.WorkerNode, + SocketID: src.Spec.SocketID, + NodeIndex: src.Spec.NodeIndex, + SocketIndex: src.Spec.Slot, + Overrides: &StorageNodeOverrides{ + SpdkImage: config.SpdkImage, + SpdkProxyImage: config.SpdkProxyImage, + SpdkSystemMemory: config.SpdkSystemMemory, + PcieAllowList: config.PcieAllowList, + PcieDenyList: config.PcieDenyList, + PcieModel: config.PcieModel, + DriveSizeRange: config.DriveSizeRange, + DeviceNames: config.DeviceNames, + FailureDomain: parseInt32(config.FailureDomain), + Expand: config.Expand, + UbuntuHost: unstashRemovedBool(&dst.ObjectMeta, annoV1Alpha1NodeUbuntuHost), + SkipKubeletConfiguration: unstashRemovedBool(&dst.ObjectMeta, annoV1Alpha1NodeSkipKubelet), + EnableCpuTopology: unstashRemovedBool(&dst.ObjectMeta, annoV1Alpha1NodeCPUTopology), + ReservedSystemCPU: unstashRemoved(&dst.ObjectMeta, annoV1Alpha1NodeReservedCPUs), + }, + } + if jm := config.JournalManager; jm != nil { + dst.Spec.Overrides.JournalManagerSpec = &JournalManagerSpec{ + Count: jm.Count, + PercentPerDevice: jm.PercentPerDevice, + } + } + + dst.Status = StorageNodeStatus{ + UUID: src.Status.UUID, + Status: src.Status.Status, + Health: src.Status.Health, + Hostname: src.Status.Hostname, + Uptime: src.Status.Uptime, + FailureDomain: parseInt32(src.Status.FailureDomain), + ActiveOpsRef: src.Status.ActiveOpsRef, + } + if err := unstash(&dst.ObjectMeta, annoV1Alpha1NodePostedAt, &dst.Status.PostedAt); err != nil { + return err + } + if lm := src.Status.LatencyMetrics; lm != nil { + dst.Status.LatencyMetrics = &NodeLatencyMetrics{ + NodeUUID: lm.NodeUUID, + BaselineP50NS: lm.BaselineP50NS, + BaselineP99NS: lm.BaselineP99NS, + BaselineMeasuredAt: lm.BaselineMeasuredAt, + } + } + if res := src.Status.Resources; res != nil { + dst.Status.Resources = &StorageNodeResources{ + CPU: res.CPU, + Memory: res.Memory, + Volumes: res.Volumes, + Devices: formatDeviceSummary(res.Devices), + } + if cap := res.Capacity; cap != nil { + dst.Status.Resources.Capacity = &StorageNodeCapacity{ + TotalBytes: cap.TotalBytes, + UsedBytes: cap.UsedBytes, + SampledAt: cap.SampledAt, + } + } + } + if ports := src.Status.Ports; ports != nil { + dst.Status.Ports = &StorageNodePorts{ + Management: ports.Management, + NvmeOf: ports.NvmeOf, + Lvol: ports.Lvol, + Rpc: ports.Rpc, + } + } + + return stashNodeHubOnly(&dst.ObjectMeta, src) +} + +// controllingClusterName reads the parent the hub names off the controller owner +// reference, which is where the upgrade's ownership phase puts it. +// +// A node that has not been reparented yet resolves to the empty string rather +// than to an error. The object is still readable that way, which is what a read +// during an upgrade needs, and the field is Required so the next write of it is +// refused until the reparent has run — which is the ordering the upgrade already +// enforces and a louder failure than a node silently joining no cluster. +func controllingClusterName(meta *metav1.ObjectMeta) string { + for _, ref := range meta.OwnerReferences { + if ref.Controller != nil && *ref.Controller && ref.Kind == clusterKind { + return ref.Name + } + } + return "" +} + +// parseDeviceSummary reads this version's device string as the hub's two counts. +// +// The stored order is total/online. Both call sites formatted the device count +// before the online count, against a doc comment claiming online/total, so a node +// with three of four devices online stored 4/3 (§15.1). Reading it the documented +// way would report every degraded node as having more devices online than it has. +// +// A string that is not two numbers is an absent summary rather than zeroes. A +// node the control plane has never reported on and one that genuinely has no +// devices are different answers, and the absent parent is how the hub tells them +// apart. +func parseDeviceSummary(summary string) *v1alpha2.StorageNodeDevices { + total, online, found := strings.Cut(summary, "/") + if !found { + return nil + } + totalN, err := strconv.ParseInt(strings.TrimSpace(total), 10, 32) + if err != nil { + return nil + } + onlineN, err := strconv.ParseInt(strings.TrimSpace(online), 10, 32) + if err != nil { + return nil + } + return &v1alpha2.StorageNodeDevices{Online: int32(onlineN), Total: int32(totalN)} +} + +// formatDeviceSummary writes the two counts back in the order this version's +// readers expect, which is the order parseDeviceSummary reads. +func formatDeviceSummary(devices *v1alpha2.StorageNodeDevices) string { + if devices == nil { + return "" + } + return strconv.FormatInt(int64(devices.Total), 10) + "/" + + strconv.FormatInt(int64(devices.Online), 10) +} + +// stashNodeHubOnly writes every hub field this version has nowhere to put. A +// field at its zero value writes no annotation, so a node that stated none of +// them is not given metadata it never had. +func stashNodeHubOnly(meta *metav1.ObjectMeta, src *v1alpha2.StorageNode) error { + for _, field := range []struct { + key string + value any + }{ + {annoNodeCluster, src.Spec.ClusterRef}, + {annoNodeSizing, src.Spec.Config.Sizing}, + {annoNodeStep, src.Status.Step}, + {annoNodePhase, string(src.Status.Phase)}, + {annoNodeMessage, src.Status.Message}, + {annoNodeObserved, src.Status.ObservedGeneration}, + } { + if err := stash(meta, field.key, field.value); err != nil { + return err + } + } + + // A failure domain that is a number survives the narrowing to an index and + // needs no note. One that is a label — which is what every domain written + // against the hub is — has nowhere to go, and only that case is recorded. + stashNonNumericDomain(meta, annoNodeSpecDomain, src.Spec.Config.FailureDomain) + stashNonNumericDomain(meta, annoNodeStatDomain, src.Status.FailureDomain) + return nil +} + +// stashNonNumericDomain records a failure-domain label this version's int32 field +// cannot hold. +func stashNonNumericDomain(meta *metav1.ObjectMeta, key, domain string) { + if domain == "" || parseInt32(domain) != nil { + clear(meta, key) + return + } + stashRemoved(meta, key, domain) +} + +// restoreNodeHubOnly reads them back and removes the annotations, so an object +// converted up carries the fields rather than both the fields and the notes about +// them. +func restoreNodeHubOnly(meta *metav1.ObjectMeta, dst *v1alpha2.StorageNode) error { + // The stash wins over the owner reference, because a cluster the hub stated + // is what the hub stated, and the reference is the derivation that answers + // for a node no hub has ever written. + var cluster string + if err := unstash(meta, annoNodeCluster, &cluster); err != nil { + return err + } + if cluster != "" { + dst.Spec.ClusterRef = cluster + } + + var step statemachine.KubeSnapshot + if err := unstash(meta, annoNodeStep, &step); err != nil { + return err + } + dst.Status.Step = step + + var phase string + if err := unstash(meta, annoNodePhase, &phase); err != nil { + return err + } + dst.Status.Phase = v1alpha2.StorageNodePhase(phase) + + for _, field := range []struct { + key string + target any + }{ + {annoNodeSizing, &dst.Spec.Config.Sizing}, + {annoNodeMessage, &dst.Status.Message}, + {annoNodeObserved, &dst.Status.ObservedGeneration}, + } { + if err := unstash(meta, field.key, field.target); err != nil { + return err + } + } + + if domain := unstashRemoved(meta, annoNodeSpecDomain); domain != "" { + dst.Spec.Config.FailureDomain = domain + } + if domain := unstashRemoved(meta, annoNodeStatDomain); domain != "" { + dst.Status.FailureDomain = domain + } + return nil +} diff --git a/operator/api/v1alpha1/storagenode_conversion_test.go b/operator/api/v1alpha1/storagenode_conversion_test.go new file mode 100644 index 000000000..7323b1ea2 --- /dev/null +++ b/operator/api/v1alpha1/storagenode_conversion_test.go @@ -0,0 +1,254 @@ +// What the StorageNode conversion has to get right, beyond the mechanical +// assignment the round trip in hub_roundtrip_test.go covers. +// +// Four rows of design-storagenode.md §15.1 need more than a copy, and each has a +// test here because each is a place the two shapes genuinely disagree: the parent +// that is not on the stored object, the sizing the stored shape has no field for, +// the failure domain that changes type, and the device summary whose stored order +// is not the order its own documentation claimed. + +package v1alpha1 + +import ( + "testing" + + "github.com/google/go-cmp/cmp" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + + "github.com/simplyblock/atlas/ptr" + + "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// A node stored as v1alpha1 names a StorageNodeSet and nothing else, so the +// cluster the hub requires comes from the controller owner reference the +// upgrade's reparent step puts there. +func TestStorageNodeReadsItsClusterFromTheControllerOwner(t *testing.T) { + stored := &StorageNode{ + ObjectMeta: metav1.ObjectMeta{ + Name: "production-7f3a9c", + Namespace: "simplyblock", + OwnerReferences: []metav1.OwnerReference{ + {Kind: "StorageNodeSet", Name: "rack-a", Controller: ptr.To(false)}, + {Kind: "StorageCluster", Name: testCluster, Controller: ptr.To(true)}, + }, + }, + Spec: StorageNodeSpec{StorageNodeSetRef: "rack-a", WorkerNode: "worker-3"}, + } + + var hub v1alpha2.StorageNode + if err := stored.ConvertTo(&hub); err != nil { + t.Fatalf("ConvertTo: %v", err) + } + + if hub.Spec.ClusterRef != testCluster { + t.Errorf("spec.clusterRef = %q, want the controlling StorageCluster %q", + hub.Spec.ClusterRef, testCluster) + } + // The set's name survives as a label rather than as a reference, so a node can + // still be traced back to the document that produced it. + if hub.Spec.NodeSet != "rack-a" { + t.Errorf("spec.nodeSet = %q, want the set the node was declared under", hub.Spec.NodeSet) + } +} + +// A node that has not been reparented yet converts with an empty cluster rather +// than failing. The object stays readable, which is what a read during an upgrade +// needs, and the field is Required so the next write of it is refused until the +// reparent has run. +func TestStorageNodeWithNoControllerConvertsWithNoCluster(t *testing.T) { + stored := &StorageNode{ + ObjectMeta: metav1.ObjectMeta{Name: "production-7f3a9c", Namespace: "simplyblock"}, + Spec: StorageNodeSpec{StorageNodeSetRef: "rack-a", WorkerNode: "worker-3"}, + } + + var hub v1alpha2.StorageNode + if err := stored.ConvertTo(&hub); err != nil { + t.Fatalf("ConvertTo: %v", err) + } + if hub.Spec.ClusterRef != "" { + t.Errorf("spec.clusterRef = %q, want it empty for a node nothing has reparented", + hub.Spec.ClusterRef) + } +} + +// The stored device summary is total/online, which is the order both call sites +// rendered it in against a doc comment claiming the reverse. Reading it the +// documented way would report every degraded node as having more devices online +// than it has. +func TestStorageNodeDeviceSummaryIsReadInTheOrderItWasWritten(t *testing.T) { + stored := &StorageNode{ + ObjectMeta: metav1.ObjectMeta{Name: "n", Namespace: "sb"}, + Status: StorageNodeStatus{ + Resources: &StorageNodeResources{Devices: "4/3"}, + }, + } + + var hub v1alpha2.StorageNode + if err := stored.ConvertTo(&hub); err != nil { + t.Fatalf("ConvertTo: %v", err) + } + + want := &v1alpha2.StorageNodeDevices{Online: 3, Total: 4} + if diff := cmp.Diff(want, hub.Status.Resources.Devices); diff != "" { + t.Errorf("the device summary was read wrongly (-want +got):\n%s", diff) + } +} + +// A summary that is not two numbers is an absent block rather than two zeroes. A +// node the control plane has never reported on and one that genuinely has no +// devices are different answers, and the absent parent is how the hub tells them +// apart. +func TestStorageNodeUnparsableDeviceSummaryHasNoBlock(t *testing.T) { + for _, summary := range []string{"", "unknown", "4", "4/x"} { + stored := &StorageNode{ + ObjectMeta: metav1.ObjectMeta{Name: "n", Namespace: "sb"}, + Status: StorageNodeStatus{ + Resources: &StorageNodeResources{Devices: summary}, + }, + } + var hub v1alpha2.StorageNode + if err := stored.ConvertTo(&hub); err != nil { + t.Fatalf("ConvertTo(%q): %v", summary, err) + } + if hub.Status.Resources.Devices != nil { + t.Errorf("summary %q produced %+v, want no device block", + summary, hub.Status.Resources.Devices) + } + } +} + +// The failure domain was an index on both the spec and the status and is a label +// on both here. The digits are what a mechanical conversion can carry, and a label +// that is not a number has nowhere to go on the way down, so it is stashed. +func TestStorageNodeFailureDomainLabelSurvivesTheTripDown(t *testing.T) { + hub := &v1alpha2.StorageNode{ + ObjectMeta: metav1.ObjectMeta{Name: "n", Namespace: "sb"}, + Spec: v1alpha2.StorageNodeSpec{ + ClusterRef: testCluster, + WorkerNode: "worker-3", + Config: v1alpha2.StorageNodeConfig{ + Sizing: v1alpha2.StorageNodeSizing{VCPUCount: ptr.To(int32(8))}, + FailureDomain: "rack-b", + }, + }, + Status: v1alpha2.StorageNodeStatus{FailureDomain: "rack-b"}, + } + + var stored StorageNode + if err := stored.ConvertFrom(hub); err != nil { + t.Fatalf("ConvertFrom: %v", err) + } + // The int32 field cannot hold a label, so it stays absent rather than holding + // a number nobody wrote. + if stored.Spec.Overrides.FailureDomain != nil { + t.Errorf("spec.overrides.failureDomain = %v, want it absent for a label", + *stored.Spec.Overrides.FailureDomain) + } + + var back v1alpha2.StorageNode + if err := stored.ConvertTo(&back); err != nil { + t.Fatalf("ConvertTo: %v", err) + } + if back.Spec.Config.FailureDomain != "rack-b" { + t.Errorf("spec.config.failureDomain = %q, want the label back", + back.Spec.Config.FailureDomain) + } + if back.Status.FailureDomain != "rack-b" { + t.Errorf("status.failureDomain = %q, want the label back", back.Status.FailureDomain) + } +} + +// An index a real v1alpha1 object holds converts up to its digits, which is a +// valid label value and the only thing a mechanical conversion can say about it. +func TestStorageNodeFailureDomainIndexBecomesItsDigits(t *testing.T) { + stored := &StorageNode{ + ObjectMeta: metav1.ObjectMeta{Name: "n", Namespace: "sb"}, + Spec: StorageNodeSpec{ + WorkerNode: "worker-3", + Overrides: &StorageNodeOverrides{FailureDomain: ptr.To(int32(1))}, + }, + Status: StorageNodeStatus{FailureDomain: ptr.To(int32(1))}, + } + + var hub v1alpha2.StorageNode + if err := stored.ConvertTo(&hub); err != nil { + t.Fatalf("ConvertTo: %v", err) + } + if hub.Spec.Config.FailureDomain != "1" { + t.Errorf("spec.config.failureDomain = %q, want the index's digits", + hub.Spec.Config.FailureDomain) + } + if hub.Status.FailureDomain != "1" { + t.Errorf("status.failureDomain = %q, want the index's digits", hub.Status.FailureDomain) + } +} + +// The four per-node fields that reached nothing move to the cluster, and they +// stash on the way up so a node converted back down carries what it carried. +func TestStorageNodeDeadPerNodeFieldsSurviveTheRoundTrip(t *testing.T) { + stored := &StorageNode{ + ObjectMeta: metav1.ObjectMeta{Name: "n", Namespace: "sb"}, + Spec: StorageNodeSpec{ + WorkerNode: "worker-3", + Overrides: &StorageNodeOverrides{ + UbuntuHost: ptr.To(true), + SkipKubeletConfiguration: ptr.To(true), + EnableCpuTopology: ptr.To(true), + ReservedSystemCPU: "0-1", + }, + }, + } + + var hub v1alpha2.StorageNode + if err := stored.ConvertTo(&hub); err != nil { + t.Fatalf("ConvertTo: %v", err) + } + var back StorageNode + if err := back.ConvertFrom(&hub); err != nil { + t.Fatalf("ConvertFrom: %v", err) + } + + got := back.Spec.Overrides + if got == nil { + t.Fatal("spec.overrides is absent, want the four stashed fields back") + } + if !ptr.BoolFromOrFalse(got.UbuntuHost) || + !ptr.BoolFromOrFalse(got.SkipKubeletConfiguration) || + !ptr.BoolFromOrFalse(got.EnableCpuTopology) || + got.ReservedSystemCPU != "0-1" { + t.Errorf("the four fields did not survive: %+v", got) + } +} + +// The sizing has no v1alpha1 spelling at all, so it stashes on the way down and +// comes back on the way up. A node the hub never wrote has none, which is what the +// upgrade's sizing stamp exists to fill. +func TestStorageNodeSizingSurvivesTheTripDown(t *testing.T) { + hub := &v1alpha2.StorageNode{ + ObjectMeta: metav1.ObjectMeta{Name: "n", Namespace: "sb"}, + Spec: v1alpha2.StorageNodeSpec{ + ClusterRef: testCluster, + WorkerNode: "worker-3", + Config: v1alpha2.StorageNodeConfig{ + Sizing: v1alpha2.StorageNodeSizing{ + VCPUCount: ptr.To(int32(8)), + MinHugePagesSize: "100G", + }, + }, + }, + } + + var stored StorageNode + if err := stored.ConvertFrom(hub); err != nil { + t.Fatalf("ConvertFrom: %v", err) + } + var back v1alpha2.StorageNode + if err := stored.ConvertTo(&back); err != nil { + t.Fatalf("ConvertTo: %v", err) + } + + if diff := cmp.Diff(hub.Spec.Config.Sizing, back.Spec.Config.Sizing); diff != "" { + t.Errorf("the sizing did not survive (-want +got):\n%s", diff) + } +} diff --git a/operator/api/v1alpha1/storagenodeops_conversion.go b/operator/api/v1alpha1/storagenodeops_conversion.go index ab9bbaea6..cdf91c81f 100644 --- a/operator/api/v1alpha1/storagenodeops_conversion.go +++ b/operator/api/v1alpha1/storagenodeops_conversion.go @@ -1,23 +1,71 @@ // Conversion of StorageNodeOps between this version and the v1alpha2 hub. // -// Four properties move (design-property-renames.md §2.1, §2.4, and §2.5): +// Four properties are renames (design-property-renames.md §2.1, §2.4, and §2.5): // storageNodeRef becomes nodeRef, drain becomes remove, targetWorkerNode and // newSsdPcie regroup under migrate, and the action enum is recased. See // controlplane_conversion.go for why an unmapped enum value is passed through // rather than rejected. +// +// The rest is the step machine of design-storagenode.md §6.3 and §7, which this +// version cannot express, and it divides in three: +// +// - status.subPhase against status.step. The two carry the same position and +// convert by table rather than by a stash, except that this version's table +// is not a function: Migrating means "move the volumes off" under Remove and +// "issue the relocation restart" under Migrate, and Restarting means "wait +// for the node to come back" under Migrate and "restart it" under +// HostMaintenance. The action is therefore an input to both directions, which +// is the whole reason §6.3 splits the value in two. +// - The drain counters. status.volumesMigrated and status.volumesPending +// become status.drain, and the arithmetic is exact in both directions: +// volumesTotal is the sum of the two, and volumesPending is what is left of +// it. A pending count that has to be kept in step with a total is what the +// regrouping removes. +// - Everything else. status.step's deadline, spec.abort, +// status.observedGeneration, and the Aborted phase have nowhere to go here, +// so they are stashed on the way down and restored on the way up. The Aborted +// phase is the sharp one: this version's Enum marker does not accept the +// value, so writing it would make the object rejected at admission rather +// than merely odd, and it is narrowed to Failed with the true phase in its +// annotation. +// +// One field travels the other way. status.triggered is this version's and the hub +// removed it (§7.2), so it stashes on the way up under +// storage.simplyblock.io/v1alpha1-status.triggered. What is preserved is the text +// that was stored rather than any behavior: the hub's controller reads the +// persisted step instead, and every call it makes is skipped when its target is +// already at or past what that call would produce. package v1alpha1 import ( + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" "sigs.k8s.io/controller-runtime/pkg/conversion" + "github.com/simplyblock/atlas/statemachine" + "github.com/simplyblock/simplyblock-operator/api/v1alpha2" ) +// The annotation holding the one status field the hub removed. +const annoV1Alpha1NodeOpsTriggered = "storage.simplyblock.io/v1alpha1-status.triggered" + +// The annotations holding the hub fields this version cannot express. +const ( + annoNodeOpsStep = "storage.simplyblock.io/conversion-status.step" + annoNodeOpsPhase = "storage.simplyblock.io/conversion-status.phase" + annoNodeOpsObserved = "storage.simplyblock.io/conversion-status.observedGeneration" + annoNodeOpsAbort = "storage.simplyblock.io/conversion-spec.abort" +) + // storageNodeOpsActionToHub maps this version's lowercase actions onto the hub's -// PascalCase ones. HostMaintenance has no v1alpha1 spelling: it is an action the -// redesign adds rather than renames, so it converts down to itself through the -// pass-through rule. +// PascalCase ones. Every value is listed, because unlike a phase this enum is +// user-authored and appears in scripts and runbooks. +// +// HostMaintenance has no row: the action is the hub's addition and this version +// never accepted it, so it passes through and is rejected by this version's own +// Enum marker, which is the correct outcome for an operation this version cannot +// perform. var storageNodeOpsActionToHub = map[string]string{ "shutdown": string(v1alpha2.StorageNodeOpsActionShutdown), "restart": string(v1alpha2.StorageNodeOpsActionRestart), @@ -27,20 +75,72 @@ var storageNodeOpsActionToHub = map[string]string{ "migrate": string(v1alpha2.StorageNodeOpsActionMigrate), } -// storageNodeOpsActionFromHub is the inverse, derived so the two cannot disagree. +// storageNodeOpsActionFromHub is the inverse, derived so the two cannot disagree +// about a value. var storageNodeOpsActionFromHub = invertStringMap(storageNodeOpsActionToHub) +// subPhaseToStep reads this version's sub-phase as the hub's step, keyed by the +// hub action the operation is performing. +// +// It is keyed rather than flat because two of this version's eight values are +// ambiguous by construction, which is the defect §6.3 names: Migrating is the +// volume drain under Remove and the relocation restart under Migrate, and +// Restarting is the wait for the node to come back under Migrate. Reading either +// without the action produces a step belonging to the other workflow, which the +// hub's per-action graph then refuses. +var subPhaseToStep = map[v1alpha2.StorageNodeOpsAction]map[string]string{ + v1alpha2.StorageNodeOpsActionRemove: { + string(StorageNodeOpsSubPhaseValidating): string(v1alpha2.StorageNodeOpsStepValidating), + string(StorageNodeOpsSubPhaseSuspending): string(v1alpha2.StorageNodeOpsStepSuspending), + string(StorageNodeOpsSubPhaseMigrating): string(v1alpha2.StorageNodeOpsStepMigratingVolumes), + string(StorageNodeOpsSubPhaseVerifying): string(v1alpha2.StorageNodeOpsStepVerifying), + string(StorageNodeOpsSubPhaseRemoving): string(v1alpha2.StorageNodeOpsStepRemoving), + }, + v1alpha2.StorageNodeOpsActionMigrate: { + string(StorageNodeOpsSubPhasePreparing): string(v1alpha2.StorageNodeOpsStepPreparing), + // This version issues the restart in Restarting and waits for the node + // in the same value. The hub splits the two, and the half a stored + // object is in is the wait: the call was made on entry. + string(StorageNodeOpsSubPhaseRestarting): string(v1alpha2.StorageNodeOpsStepAwaitingNode), + string(StorageNodeOpsSubPhasePromoting): string(v1alpha2.StorageNodeOpsStepPromoting), + }, +} + +// stepToSubPhase projects a hub step back onto a sub-phase this version's Enum +// accepts, so that a v1alpha1 reader sees where the operation is rather than an +// empty field. +// +// It is not the inverse of subPhaseToStep and cannot be. Four of the hub's steps +// have no v1alpha1 spelling at all: Requesting and Awaiting, because the four +// single-step actions never had a sub-phase here, and Relocating and AwaitingNode, +// which are the two halves this version wrote as one Restarting. Every step of +// HostMaintenance is likewise absent, since the action is. Those read as an empty +// sub-phase, and the stash carries the real step. +var stepToSubPhase = map[string]StorageNodeOpsSubPhase{ + string(v1alpha2.StorageNodeOpsStepValidating): StorageNodeOpsSubPhaseValidating, + string(v1alpha2.StorageNodeOpsStepSuspending): StorageNodeOpsSubPhaseSuspending, + string(v1alpha2.StorageNodeOpsStepMigratingVolumes): StorageNodeOpsSubPhaseMigrating, + string(v1alpha2.StorageNodeOpsStepVerifying): StorageNodeOpsSubPhaseVerifying, + string(v1alpha2.StorageNodeOpsStepRemoving): StorageNodeOpsSubPhaseRemoving, + string(v1alpha2.StorageNodeOpsStepPreparing): StorageNodeOpsSubPhasePreparing, + string(v1alpha2.StorageNodeOpsStepRelocating): StorageNodeOpsSubPhaseRestarting, + string(v1alpha2.StorageNodeOpsStepAwaitingNode): StorageNodeOpsSubPhaseRestarting, + string(v1alpha2.StorageNodeOpsStepPromoting): StorageNodeOpsSubPhasePromoting, +} + // ConvertTo converts this StorageNodeOps to the v1alpha2 hub. func (src *StorageNodeOps) ConvertTo(dstRaw conversion.Hub) error { dst := dstRaw.(*v1alpha2.StorageNodeOps) - dst.ObjectMeta = src.ObjectMeta + dst.ObjectMeta = *src.ObjectMeta.DeepCopy() + stashOpsFlag(&dst.ObjectMeta, annoV1Alpha1NodeOpsTriggered, src.Status.Triggered) + action := v1alpha2.StorageNodeOpsAction( + mapOrPassThrough(storageNodeOpsActionToHub, src.Spec.Action), + ) dst.Spec = v1alpha2.StorageNodeOpsSpec{ - NodeRef: src.Spec.StorageNodeRef, - Action: v1alpha2.StorageNodeOpsAction( - mapOrPassThrough(storageNodeOpsActionToHub, src.Spec.Action), - ), + NodeRef: src.Spec.StorageNodeRef, + Action: action, Force: src.Spec.Force, ReattachVolume: src.Spec.ReattachVolume, } @@ -62,24 +162,38 @@ func (src *StorageNodeOps) ConvertTo(dstRaw conversion.Hub) error { } dst.Status = v1alpha2.StorageNodeOpsStatus{ - Phase: v1alpha2.StorageNodeOpsPhase(src.Status.Phase), - SubPhase: v1alpha2.StorageNodeOpsSubPhase(src.Status.SubPhase), - Message: src.Status.Message, - VolumesMigrated: src.Status.VolumesMigrated, - VolumesPending: src.Status.VolumesPending, - Triggered: src.Status.Triggered, - StartedAt: src.Status.StartedAt, - CompletedAt: src.Status.CompletedAt, + Phase: v1alpha2.StorageNodeOpsPhase(src.Status.Phase), + Message: src.Status.Message, + StartedAt: src.Status.StartedAt, + CompletedAt: src.Status.CompletedAt, + } + // A sub-phase this version's tables do not place under this action is + // dropped rather than passed through. The target is a field of a shared + // type, so an unrecognized step fails the machine's restore rather than the + // object's admission (§6.3), and an empty step restores to the graph's + // initial state, which is where an operation with nothing recorded belongs. + if step, ok := subPhaseToStep[action][string(src.Status.SubPhase)]; ok { + dst.Status.Step = statemachine.KubeSnapshot{State: step} + } + // The drain block exists only for an operation that had one. The counters + // are declared "drain only" here and default to zero on every other action, + // so allocating a block for them would give a Shutdown a drain that says it + // moved none of no volumes. + if src.Status.VolumesMigrated != 0 || src.Status.VolumesPending != 0 { + dst.Status.Drain = &v1alpha2.DrainStatus{ + VolumesTotal: int32(src.Status.VolumesMigrated + src.Status.VolumesPending), + VolumesMigrated: int32(src.Status.VolumesMigrated), + } } - return nil + return restoreNodeOpsHubOnly(&dst.ObjectMeta, dst) } // ConvertFrom converts the v1alpha2 hub into this StorageNodeOps. func (dst *StorageNodeOps) ConvertFrom(srcRaw conversion.Hub) error { src := srcRaw.(*v1alpha2.StorageNodeOps) - dst.ObjectMeta = src.ObjectMeta + dst.ObjectMeta = *src.ObjectMeta.DeepCopy() dst.Spec = StorageNodeOpsSpec{ StorageNodeRef: src.Spec.NodeRef, @@ -96,15 +210,105 @@ func (dst *StorageNodeOps) ConvertFrom(srcRaw conversion.Hub) error { } dst.Status = StorageNodeOpsStatus{ - Phase: StorageNodeOpsPhase(src.Status.Phase), - SubPhase: StorageNodeOpsSubPhase(src.Status.SubPhase), - Message: src.Status.Message, - VolumesMigrated: src.Status.VolumesMigrated, - VolumesPending: src.Status.VolumesPending, - Triggered: src.Status.Triggered, - StartedAt: src.Status.StartedAt, - CompletedAt: src.Status.CompletedAt, + Phase: narrowNodeOpsPhase(src.Status.Phase), + SubPhase: stepToSubPhase[src.Status.Step.State], + Message: src.Status.Message, + Triggered: unstashRemoved(&dst.ObjectMeta, annoV1Alpha1NodeOpsTriggered) == stashedTrue, + StartedAt: src.Status.StartedAt, + CompletedAt: src.Status.CompletedAt, + } + if d := src.Status.Drain; d != nil { + dst.Status.VolumesMigrated = int(d.VolumesMigrated) + dst.Status.VolumesPending = int(d.VolumesTotal - d.VolumesMigrated) } + return stashNodeOpsHubOnly(&dst.ObjectMeta, src) +} + +// stashNodeOpsHubOnly writes every hub field this version has nowhere to put. A +// field at its zero value writes no annotation, so an operation that reached none +// of them is not given metadata it never had. +func stashNodeOpsHubOnly(meta *metav1.ObjectMeta, src *v1alpha2.StorageNodeOps) error { + if err := stashNodeOpsStep(meta, src.Spec.Action, src.Status.Step); err != nil { + return err + } + if err := stash(meta, annoNodeOpsObserved, src.Status.ObservedGeneration); err != nil { + return err + } + + // Only the phase this version cannot spell is recorded. Every other value + // survives the narrowing, and an annotation for each would be noise on every + // operation that ever ran. + if src.Status.Phase == v1alpha2.StorageNodeOpsPhaseAborted { + stashRemoved(meta, annoNodeOpsPhase, string(src.Status.Phase)) + } else { + clear(meta, annoNodeOpsPhase) + } + + // spec.abort is a bool rather than a pointer, so false and unset are the same + // value and only true is worth a note. + if src.Spec.Abort { + stashRemoved(meta, annoNodeOpsAbort, stashedTrue) + } else { + clear(meta, annoNodeOpsAbort) + } return nil } + +// stashNodeOpsStep records the step only when subPhase cannot carry it. +// +// The projection is one-way, so the five Remove steps and two of the four Migrate +// steps survive being written to subPhase and read back, and annotating those +// would put a note on every drain that ever ran. What does not survive is a step +// with a deadline, either half of the Relocating and AwaitingNode pair this +// version spelled as one Restarting, and every step of an action subPhase never +// covered. +func stashNodeOpsStep( + meta *metav1.ObjectMeta, + action v1alpha2.StorageNodeOpsAction, + step statemachine.KubeSnapshot, +) error { + roundTrips := subPhaseToStep[action][string(stepToSubPhase[step.State])] == step.State + if step.Deadline == nil && roundTrips { + clear(meta, annoNodeOpsStep) + return nil + } + return stash(meta, annoNodeOpsStep, step) +} + +// restoreNodeOpsHubOnly reads them back and removes the annotations, so an object +// converted up carries the fields rather than both the fields and the notes about +// them. +func restoreNodeOpsHubOnly(meta *metav1.ObjectMeta, dst *v1alpha2.StorageNodeOps) error { + // The stash wins where there is one, because only it carries a deadline and + // the steps subPhase cannot spell. ConvertTo has already read + // status.subPhase for an object a real v1alpha1 client wrote. + var step statemachine.KubeSnapshot + if err := unstash(meta, annoNodeOpsStep, &step); err != nil { + return err + } + if step.State != "" || step.Deadline != nil { + dst.Status.Step = step + } + + if err := unstash(meta, annoNodeOpsObserved, &dst.Status.ObservedGeneration); err != nil { + return err + } + + if phase := unstashRemoved(meta, annoNodeOpsPhase); phase != "" { + dst.Status.Phase = v1alpha2.StorageNodeOpsPhase(phase) + } + dst.Spec.Abort = unstashRemoved(meta, annoNodeOpsAbort) == stashedTrue + return nil +} + +// narrowNodeOpsPhase maps a hub phase onto one this version's Enum accepts. Only +// Aborted needs it, and it reads as Failed here: the operation did stop before +// finishing, which is the closest true statement this version can make, and the +// annotation carries the distinction. +func narrowNodeOpsPhase(phase v1alpha2.StorageNodeOpsPhase) StorageNodeOpsPhase { + if phase == v1alpha2.StorageNodeOpsPhaseAborted { + return StorageNodeOpsPhaseFailed + } + return StorageNodeOpsPhase(phase) +} diff --git a/operator/api/v1alpha2/clusterdeploymentconfig_types.go b/operator/api/v1alpha2/clusterdeploymentconfig_types.go index aa27ce189..edaea6046 100644 --- a/operator/api/v1alpha2/clusterdeploymentconfig_types.go +++ b/operator/api/v1alpha2/clusterdeploymentconfig_types.go @@ -23,23 +23,12 @@ import ( "github.com/simplyblock/atlas/statemachine" ) -// JournalManagerSpec configures the journal managers on a set of nodes. -// -// It is declared here rather than borrowed from v1alpha1, as -// design-clusterdeploymentconfig.md §Appendix A spells it. A v1alpha2 type whose -// fields are v1alpha1 types cannot be reshaped by the redesign without changing -// this version's wire format, and it makes the import cycle that the conversion -// webhook needs impossible: the spoke's ConvertTo and ConvertFrom have to be -// methods on the v1alpha1 type, so v1alpha1 imports v1alpha2 and v1alpha2 cannot -// import back. -type JournalManagerSpec struct { - // Count is the number of journal managers to configure. - // +optional - Count *int32 `json:"count,omitempty"` - // PercentPerDevice is the journal manager capacity percentage per device. - // +optional - PercentPerDevice *int32 `json:"percentPerDevice,omitempty"` -} +// JournalManagerSpec, the journal tuning this document's node-set template +// states, is StorageNode's own type in storagenode_types.go. It was declared +// here while that kind was still v1alpha1, for the reason StripeSpec was, and +// moved to the kind that owns the concept once it arrived: a journal count and a +// per-device share are one node's on-disk layout, fixed when its devices were +// partitioned. // StripeSpec, the erasure-coding layout this document's template states, is // StorageCluster's own type in storagecluster_types.go. It was declared here diff --git a/operator/api/v1alpha2/storagecluster_types.go b/operator/api/v1alpha2/storagecluster_types.go index 57d4e5bff..a71d443df 100644 --- a/operator/api/v1alpha2/storagecluster_types.go +++ b/operator/api/v1alpha2/storagecluster_types.go @@ -11,11 +11,10 @@ // `enable` fields and one removal, and status gaining a typed phase, a creation // step, the control plane's task window, and observedGeneration. // -// One field of Appendix A is deliberately absent. spec.storageNodes is the -// Kubernetes workload the cluster's nodes run as, and its type belongs to -// design-storagenode.md Appendix C, which has not been written yet: StorageNode -// is still v1alpha1 and StorageNodeSet still owns the workload. It lands with -// that kind's move rather than here, where it could only be an empty block. +// spec.storageNodes is the Kubernetes workload the cluster's nodes run as. Its +// type is design-storagenode.md Appendix C and it arrived with that kind's move +// to v1alpha2, which is what retired StorageNodeSet and left the DaemonSet, the +// Services, the certificates, and the per-node ConfigMap without an owner. package v1alpha2 @@ -439,6 +438,117 @@ type ClusterTask struct { Retry int32 `json:"retry,omitempty"` } +// StorageNodesSpec is the Kubernetes workload every storage node in the cluster +// runs as: a DaemonSet, a headless Service and its EndpointSlices, a serving +// certificate, a ServiceAccount, and the ConfigMap the init container reads its +// per-node configuration out of. +// +// Every field here is cluster-uniform by construction, because a DaemonSet is one +// object for every node it schedules and its pod template cannot differ per node. +// What can differ is in StorageNode.spec.config: the two images, the SPDK system +// memory, and the sizing block, which are per node precisely so that an image +// rollout and a hardware re-size can walk the fleet one machine at a time. +type StorageNodesSpec struct { + // Image is the storage-node container image. Defaults to the ControlPlane + // singleton's spec.image when unset, so a deployment states the version once. + // +kubebuilder:validation:Pattern=`^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$` + // +optional + Image string `json:"image,omitempty"` + + // ImagePullPolicy controls when that image is pulled. + // +kubebuilder:validation:Enum=Always;Never;IfNotPresent + // +kubebuilder:default=IfNotPresent + // +optional + ImagePullPolicy corev1.PullPolicy `json:"imagePullPolicy,omitempty"` + + // MgmtInterface is the management network interface storage nodes bind. + // +optional + // +k8s:immutable + MgmtInterface string `json:"mgmtInterface,omitempty"` + + // DataInterfaces are the data-plane network interfaces. + // +optional + DataInterfaces []string `json:"dataInterfaces,omitempty"` + + // SocketsToUse restricts deployment to selected NUMA sockets. Empty means + // socket 0 alone. + // +optional + SocketsToUse []string `json:"socketsToUse,omitempty"` + + // NodesPerSocket is how many storage nodes run per NUMA socket. + // +kubebuilder:validation:Minimum=1 + // +kubebuilder:default=1 + // +optional + // +k8s:immutable + NodesPerSocket *int32 `json:"nodesPerSocket,omitempty"` + + // MaxParallelNodeAdds limits how many workers may be in the node-add process + // at once, counted by distinct worker rather than by object so that a + // two-socket host consumes one slot. Workers hosting a FoundationDB pod are + // always sequential regardless of this value, because a node add reboots the + // host and two simultaneous FoundationDB reboots reduce the control plane's + // own fault tolerance. + // +kubebuilder:validation:Minimum=1 + // +kubebuilder:default=1 + // +optional + MaxParallelNodeAdds *int32 `json:"maxParallelNodeAdds,omitempty"` + + // EnableJournalDevice dedicates the smallest NVMe device on each node to the + // journal manager, instead of carving a journal partition out of every + // device. + // +optional + // +k8s:immutable + EnableJournalDevice *bool `json:"enableJournalDevice,omitempty"` + + // EnableFormat4K formats NVMe devices to a 4K block size where the device + // supports it. + // +optional + // +k8s:immutable + EnableFormat4K *bool `json:"enableFormat4K,omitempty"` + + // EnableCpuTopology turns on topology-aware CPU assignment. + // +optional + EnableCpuTopology *bool `json:"enableCpuTopology,omitempty"` + + // ReservedSystemCPU is the CPU set held back from SPDK for system workloads. + // +optional + ReservedSystemCPU string `json:"reservedSystemCPU,omitempty"` + + // EnableKubeletConfiguration lets the storage node apply the kubelet + // configuration changes it needs. Off by default, which is the behavior the + // retired skipKubeletConfiguration expressed by being set. + // +optional + EnableKubeletConfiguration *bool `json:"enableKubeletConfiguration,omitempty"` + + // UbuntuHost states that the worker's host OS is Ubuntu, which changes how + // the node configures huge pages and the kernel modules it loads. + // +optional + UbuntuHost *bool `json:"ubuntuHost,omitempty"` + + // OpenShiftCluster states that the Kubernetes distribution is OpenShift. + // +optional + OpenShiftCluster *bool `json:"openShiftCluster,omitempty"` + + // OpenShiftMachineConfigPool names the pool generated MachineConfig objects + // are labeled into. + // +kubebuilder:default=worker + // +optional + OpenShiftMachineConfigPool string `json:"openShiftMachineConfigPool,omitempty"` + + // Tolerations are applied to the storage-node pods. + // +optional + Tolerations []corev1.Toleration `json:"tolerations,omitempty"` + + // ContainerResources sets requests and limits for the storage-node container. + // Unset enforces no limits. + // +optional + ContainerResources corev1.ResourceRequirements `json:"containerResources,omitempty"` + + // InitContainerResources does the same for the init container. + // +optional + InitContainerResources corev1.ResourceRequirements `json:"initContainerResources,omitempty"` +} + // StorageClusterSpec is the desired state of one simplyblock backend cluster. // +kubebuilder:validation:XValidation:rule="!has(oldSelf.kms) || self.kms == oldSelf.kms",message="kms is immutable once set" // +kubebuilder:validation:XValidation:rule="!(has(self.enableAtomic4kWrites) && self.enableAtomic4kWrites) || (has(self.enableChecksumValidation) && self.enableChecksumValidation)",message="enableAtomic4kWrites requires enableChecksumValidation to be true" @@ -586,6 +696,15 @@ type StorageClusterSpec struct { // +optional MaxConcurrentWorkerRestarts *int32 `json:"maxConcurrentWorkerRestarts,omitempty"` + // StorageNodes is the Kubernetes workload the cluster's storage nodes run + // as, and the cluster owns every object in it by controller reference: a + // cluster deleted takes its DaemonSet, Services, certificate, and per-node + // ConfigMap with it. One workload serves the whole cluster, because growth is + // nodes rather than sets and what differs between hardware generations is per + // node already. + // +optional + StorageNodes *StorageNodesSpec `json:"storageNodes,omitempty"` + // Backup is the S3 location this cluster's backups live in, and it is both // the target copies are written to and the inventory the operator walks to // produce StorageBackup objects. Mutable: a cluster that has never had a diff --git a/operator/api/v1alpha2/storagenode_types.go b/operator/api/v1alpha2/storagenode_types.go new file mode 100644 index 000000000..63d412930 --- /dev/null +++ b/operator/api/v1alpha2/storagenode_types.go @@ -0,0 +1,535 @@ +// StorageNode in the shape design-storagenode.md Appendix A specifies: one +// backend storage node, meaning one SPDK process bound to one NUMA socket of one +// Kubernetes worker. +// +// What moves against the registered v1alpha1 type is §15.1 of that document, and +// the conversion in api/v1alpha1/storagenode_conversion.go is the whole of the +// translation. The headline changes are spec.storageNodeSetRef becoming +// spec.clusterRef plus spec.nodeSet, spec.overrides becoming spec.config and +// stopping being an override of anything, spec.socketIndex becoming spec.slot, +// the failure domain becoming a label rather than an index, the four per-node +// fields that reached nothing moving to StorageCluster.spec.storageNodes, and +// status gaining a typed phase, a provisioning step, a device summary that is two +// counts rather than a string, and observedGeneration. +// +// NodeLatencyMetrics is declared here rather than on the retired StorageNodeSet, +// because the reading is one node's and the fleet object that used to collect +// them is gone (§15.3). + +package v1alpha2 + +import ( + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + + "github.com/simplyblock/atlas/statemachine" +) + +// StorageNodePhase is where the operator has got to with this node. The first two +// values are the operator's own provisioning path; the rest are its reading of the +// lifecycle status.status carries in the control plane's own spelling. +// +kubebuilder:validation:Enum=Pending;Provisioning;Online;Removing;Offline;Degraded;Failed +type StorageNodePhase string + +const ( + // StorageNodePhasePending: the object exists and no slot has been claimed + // for it yet. + StorageNodePhasePending StorageNodePhase = "Pending" + + // StorageNodePhaseProvisioning: the provisioning machine is running. + StorageNodePhaseProvisioning StorageNodePhase = "Provisioning" + + // StorageNodePhaseOnline: the control plane reports the node online and + // carrying its share. + StorageNodePhaseOnline StorageNodePhase = "Online" + + // StorageNodePhaseRemoving: a StorageNodeOps with action Remove is draining + // it. + StorageNodePhaseRemoving StorageNodePhase = "Removing" + + // StorageNodePhaseOffline: out of service and reachable, which is where + // Shutdown, Suspend, and a host maintenance window leave it. + StorageNodePhaseOffline StorageNodePhase = "Offline" + + // StorageNodePhaseDegraded: serving with less than its devices, which is the + // node-level half of what StorageDevice reports per device. + StorageNodePhaseDegraded StorageNodePhase = "Degraded" + + // StorageNodePhaseFailed: unreachable, timed out, or provisioning that will + // not complete. + StorageNodePhaseFailed StorageNodePhase = "Failed" +) + +// StorageNodeStep is one step of the provisioning path. There is one graph rather +// than a MultiConfig, because an entity has no spec.action to key one on. +// +kubebuilder:validation:Enum=CheckingHost;CheckingConfig;AwaitingSlot;Posting;Resolving;Adopting +type StorageNodeStep string + +const ( + // StorageNodeStepCheckingHost waits for the worker's storage-node API to + // answer, which is the precondition for adding the node at all. + StorageNodeStepCheckingHost StorageNodeStep = "CheckingHost" + + // StorageNodeStepCheckingConfig holds until the node declares a fault group, + // where its cluster requires one. It is a gate rather than a validation: the + // value can arrive later, and holding is what makes filling it in sufficient. + StorageNodeStepCheckingConfig StorageNodeStep = "CheckingConfig" + + // StorageNodeStepAwaitingSlot holds until the cluster is under its + // parallel-add limit and no FoundationDB worker is in flight. + StorageNodeStepAwaitingSlot StorageNodeStep = "AwaitingSlot" + + // StorageNodeStepPosting is the claim: the transition into it is the + // optimistic-lock patch that makes the node-add single-shot, and the POST + // follows it. + StorageNodeStepPosting StorageNodeStep = "Posting" + + // StorageNodeStepResolving matches this node's slot against the cluster's + // node list until the backend UUID appears. + StorageNodeStepResolving StorageNodeStep = "Resolving" + + // StorageNodeStepAdopting takes over a backend node the operator did not add. + StorageNodeStepAdopting StorageNodeStep = "Adopting" +) + +// JournalManagerSpec tunes the journal managers on one storage node. +type JournalManagerSpec struct { + // Count is the number of journal managers to configure. + // +kubebuilder:validation:Minimum=1 + // +optional + Count *int32 `json:"count,omitempty"` + + // PercentPerDevice is the share of each device given to the journal. + // +kubebuilder:validation:Minimum=1 + // +kubebuilder:validation:Maximum=100 + // +optional + PercentPerDevice *int32 `json:"percentPerDevice,omitempty"` +} + +// StorageNodeSizing is what this node's SPDK core layout and huge-page floor were +// sized from. It is stamped from StorageCluster.spec when the node is created and +// is equal across the fleet in steady state; a rolling hardware upgrade is what +// makes two nodes differ, and only for as long as the roll takes. The StorageNode +// validating webhook admits a change from the operator and rejects it from +// everyone else, because unmanaged divergence is what stops the control plane +// placing erasure-coding chunks evenly. +// +// The cluster's maxSubsystemCount is not copied in here. It bounds how many +// volumes a node can serve rather than describing the host the node runs on, so +// it is the same for every node of a cluster and is read from the cluster when the +// node's configuration is generated. +type StorageNodeSizing struct { + // VCPUCount is the number of vCPUs allocated to SPDK on this node, as an + // explicit core count rather than a percentage. + // +kubebuilder:validation:Minimum=4 + // +kubebuilder:validation:Required + VCPUCount *int32 `json:"vcpuCount"` + + // MinHugePagesSize is the smallest huge-page allocation this node makes, as a + // size string such as 100G or 1T, where a bare number is gigabytes. It is a + // floor rather than a limit: the effective allocation is the larger of this + // value and the minimum the node's device and subsystem count requires. + // +optional + MinHugePagesSize string `json:"minHugePagesSize,omitempty"` +} + +// StorageNodeConfig is a storage node's complete configuration, copied from the +// ClusterDeploymentConfig entry that produced the node. It is a copy rather than a +// projection, because that document is ephemeral and nothing reads it once the +// node exists. +// +// Most of it is immutable: by marker where a field has no legitimate writer, and +// by the validating webhook where it has exactly one. +type StorageNodeConfig struct { + // Sizing is what this node's huge pages and core layout were sized from. + // Writable by the operator alone. + // +kubebuilder:validation:Required + Sizing StorageNodeSizing `json:"sizing"` + + // SpdkImage overrides the SPDK image the control plane starts for this node, + // which is what makes a phased image rollout expressible per node. + // +optional + SpdkImage string `json:"spdkImage,omitempty"` + + // SpdkProxyImage overrides the SPDK proxy image for this node. + // +optional + SpdkProxyImage string `json:"spdkProxyImage,omitempty"` + + // SpdkSystemMemory is the memory the control plane starts this node's SPDK + // with, as a size string such as 4G or 512M. Mutable: a node whose device + // count grew legitimately needs to raise it. + // +kubebuilder:validation:Pattern=`^[0-9]+(G|GI|GB|GiB|M|MI|MB|MiB|g|gi|gb|gib|m|mi|mb|mib)?$` + // +optional + SpdkSystemMemory string `json:"spdkSystemMemory,omitempty"` + + // JournalManager tunes the journal manager count and per-device capacity + // share for this node. Immutable: both are on-disk layout, fixed when the + // devices were partitioned. + // +optional + // +k8s:immutable + JournalManager *JournalManagerSpec `json:"journalManager,omitempty"` + + // DeviceNames names the devices to use. An entry is a PCI address + // ("0000:5e:00.0") or a device path ("/dev/sdb,") which are the two classes + // simplyblock accepts as backend storage, and a bare name ("nvme0n1") is read + // as a path under /dev. One list carries both spellings, and every entry is of + // the class its cluster declares in StorageCluster.spec.deviceClass: a list + // mixing the two, or naming the class the cluster is not, is rejected by the + // StorageNode validating webhook. Set explicitly, it overrides every filter + // below. Immutable: it selects which physical devices the node owns. + // +kubebuilder:validation:items:Pattern=`^([0-9a-fA-F]{4}:[0-9a-fA-F]{2}:[0-9a-fA-F]{2}\.[0-9a-fA-F]|/dev/[a-zA-Z0-9._/-]+|[a-zA-Z0-9._-]+)$` + // +optional + // +k8s:immutable + DeviceNames []string `json:"deviceNames,omitempty"` + + // PcieAllowList selects devices by PCI address. It is the one device field a + // migration writes, merging spec.migrate.newSsdPcie into it so devices added + // on the target host survive a later rebuild, so it is guarded by the + // StorageNode validating webhook rather than by a marker. This and the two + // PCI filters below belong to an NVMe cluster: the webhook rejects them on a + // cluster whose deviceClass is LogicalBlock, because a logical block device + // has no PCI address to match. + // +optional + PcieAllowList []string `json:"pcieAllowList,omitempty"` + + // PcieDenyList excludes devices by PCI address. + // +optional + // +k8s:immutable + PcieDenyList []string `json:"pcieDenyList,omitempty"` + + // PcieModel filters devices by PCI model string. + // +optional + // +k8s:immutable + PcieModel string `json:"pcieModel,omitempty"` + + // DriveSizeRange filters devices by size. + // +optional + // +k8s:immutable + DriveSizeRange string `json:"driveSizeRange,omitempty"` + + // FailureDomain is the label of the fault group this node belongs to + // such as rack-b, naming the physical grouping it shares with its peers rather + // than indexing it. Required when the cluster has enableFailureDomains set, + // and provisioning is held with a FailureDomainMissing event until it is + // present. Immutable once set, which is what makes it fillable later and then + // frozen: chunk placement was computed from it. + // + // The value takes the shape of a Kubernetes label value, because that is what + // it is seeded from where a cluster carries topology labels at all. + // +kubebuilder:validation:MaxLength=63 + // +kubebuilder:validation:Pattern=`^[a-zA-Z0-9]([-_.a-zA-Z0-9]*[a-zA-Z0-9])?$` + // +optional + // +k8s:immutable + FailureDomain string `json:"failureDomain,omitempty"` + + // Expand marks this node as an addition to an already-active cluster, which + // the control plane reads as a request to rebalance onto it rather than to + // treat it as part of an initial layout. Immutable once set: it describes how + // the node joined rather than what it is. + // +optional + // +k8s:immutable + Expand *bool `json:"expand,omitempty"` +} + +// StorageNodeSpec is the desired state of one backend storage node, meaning one +// SPDK process bound to one NUMA socket of one Kubernetes worker. +type StorageNodeSpec struct { + // ClusterRef names the StorageCluster this node belongs to. The cluster also + // owns this object by controller reference, so deleting the cluster deletes + // its nodes. + // +kubebuilder:validation:Required + // +k8s:immutable + ClusterRef string `json:"clusterRef"` + + // NodeSet is the name of the group in ClusterDeploymentConfig.nodeSets[] this + // node was declared under. It is a label rather than a reference: nothing is + // fetched by it, and it exists so that a node can be traced back to the + // document that produced it. + // +optional + // +k8s:immutable + NodeSet string `json:"nodeSet,omitempty"` + + // WorkerNode is the Kubernetes worker hostname this node runs on. It is not + // marked immutable, because a migration re-points it, but the StorageNode + // validating webhook rejects any change made by an identity outside the + // operator's namespace. + // +kubebuilder:validation:Required + WorkerNode string `json:"workerNode"` + + // SocketID is the NUMA socket this node is bound to, as declared in the node + // set's socket list, so 0 or 1. With NodeIndex it decomposes Slot into the + // pair a person reads; nothing but a print column consumes either. + // +optional + // +k8s:immutable + SocketID string `json:"socketId,omitempty"` + + // NodeIndex is the position among the nodes sharing this socket, in + // 0..nodesPerSocket-1. See SocketID. + // +kubebuilder:validation:Minimum=0 + // +optional + // +k8s:immutable + NodeIndex *int32 `json:"nodeIndex,omitempty"` + + // Slot is which storage-node slot on this worker the object occupies, counted + // from zero. A worker runs one node per socket per nodesPerSocket, and the + // slot is the position among them. It is the identity the operator keys on: + // the topology label the CSI driver reads is + // storage.simplyblock.io/storage-node-uuid.., and the slot + // outlives the node filling it, because only the UUID behind it changes when a + // node is replaced or relocated. + // +kubebuilder:validation:Minimum=0 + // +optional + // +k8s:immutable + Slot *int32 `json:"slot,omitempty"` + + // Config is this node's complete configuration, copied from the + // ClusterDeploymentConfig entry that produced it. It is a copy rather than a + // projection, because that document is ephemeral: nothing reads it once the + // node exists, deleting it changes nothing, and editing it reaches only nodes + // created afterward. + // +kubebuilder:validation:Required + Config StorageNodeConfig `json:"config"` +} + +// NodeLatencyMetrics is the fio-measured 4K NVMe-oF write latency of one backend +// storage node, which is the denominator of the rebalancer's deviation signal. +// +// It is a node's reading and it lives on the node. The retired StorageNodeSet +// collected one entry per node in a fleet-wide list, which made every node's +// measurement a write to one object shared by all of them. +type NodeLatencyMetrics struct { + // NodeUUID is the backend storage node the reading was taken against. It is + // carried beside the reading rather than inferred from status.uuid, because a + // baseline measured against one backend node stops describing the slot once a + // replacement fills it. + // +kubebuilder:validation:Required + NodeUUID string `json:"nodeUUID"` + + // BaselineP50NS is the p50 write latency, in nanoseconds, of the initial + // empty-cluster benchmark. + // +kubebuilder:validation:Minimum=0 + // +optional + BaselineP50NS int64 `json:"baselineP50NS,omitempty"` + + // BaselineP99NS is the p99 write latency, in nanoseconds, of the same + // benchmark. + // +kubebuilder:validation:Minimum=0 + // +optional + BaselineP99NS int64 `json:"baselineP99NS,omitempty"` + + // BaselineMeasuredAt is when the baseline was established. + // +optional + BaselineMeasuredAt *metav1.Time `json:"baselineMeasuredAt,omitempty"` +} + +// StorageNodeDevices counts the NVMe devices on a node and how many of them are +// online. It is a summary rather than an inventory: per-device capacity, health, +// and conditions belong to StorageDevice. +// +// Neither field takes omitempty. Zero online devices is the condition worth +// seeing, and a field that disappears at zero would report it as nothing at all. A +// node the control plane has not reported on is the absent parent instead. +type StorageNodeDevices struct { + // Online is how many of the node's devices the control plane reports as + // usable. + // +kubebuilder:validation:Minimum=0 + Online int32 `json:"online"` + + // Total is how many devices the node has. + // +kubebuilder:validation:Minimum=0 + Total int32 `json:"total"` +} + +// StorageNodeCapacity is a node's storage occupancy, as the control plane last +// measured it. +// +// It carries the same two numbers as a device's capacity, because a node's is the +// sum of its devices' and a reader comparing the two should not have to reconcile +// different shapes. It is written only when the reading has moved materially: a +// sample that changed by a few blocks is not worth an etcd write, and writing +// every sample would make the reconciler retrigger itself on its own status +// update. +type StorageNodeCapacity struct { + // TotalBytes is the storage the node's devices provide. + // +kubebuilder:validation:Minimum=0 + // +optional + TotalBytes *int64 `json:"totalBytes,omitempty"` + + // UsedBytes is what they currently hold. + // +kubebuilder:validation:Minimum=0 + // +optional + UsedBytes *int64 `json:"usedBytes,omitempty"` + + // SampledAt is when the control plane took the reading. It is not when the + // object was written, and it may be considerably older if metrics collection + // has stopped. + // +optional + SampledAt *metav1.Time `json:"sampledAt,omitempty"` +} + +// StorageNodeResources groups the compute and storage figures the control plane +// reports for a node. +type StorageNodeResources struct { + // CPU is the number of SPDK cores allocated to this node. + // +optional + CPU *int32 `json:"cpu,omitempty"` + + // Memory is the SPDK memory allocation the control plane reports. + // +optional + Memory string `json:"memory,omitempty"` + + // Volumes is the current number of logical volumes on this node. + // +optional + Volumes *int32 `json:"volumes,omitempty"` + + // Devices summarizes the node's NVMe devices. Absent until the control plane + // has reported, which is what tells a node that has not reported from one that + // genuinely has no devices. + // +optional + Devices *StorageNodeDevices `json:"devices,omitempty"` + + // Capacity is how much of the node's storage is in use, summed over its + // devices. It is a measurement rather than a declaration, so it is absent + // until something has measured it, and it lags reality by the interval at + // which the control plane's metrics are scraped. + // +optional + Capacity *StorageNodeCapacity `json:"capacity,omitempty"` +} + +// StorageNodePorts groups the addresses and ports a node listens on. +type StorageNodePorts struct { + // Management is the management IP address of the node. + // +optional + Management string `json:"management,omitempty"` + + // The NVMe-oF fabric port. + // +optional + NvmeOf *int32 `json:"nvmeof,omitempty"` + + // Lvol is the logical-volume subsystem port. + // +optional + Lvol *int32 `json:"lvol,omitempty"` + + // Rpc is the RPC and management API port. + // +optional + Rpc *int32 `json:"rpc,omitempty"` +} + +// StorageNodeStatus is the observed state of one storage node. +type StorageNodeStatus struct { + // Phase is the operator's own view of this node, and the field its + // provisioning branches on. + // +optional + Phase StorageNodePhase `json:"phase,omitempty"` + + // Step is the position of the provisioning machine, as the shared + // statemachine.KubeSnapshot. The rule is what an Enum marker would do if a + // marker could reach a field of a shared type. + // +kubebuilder:validation:XValidation:rule="!has(self.state) || self.state in ['CheckingHost','CheckingConfig','AwaitingSlot','Posting','Resolving','Adopting']",message="unknown step" + // +optional + Step statemachine.KubeSnapshot `json:"step,omitempty"` + + // UUID is the backend node UUID. Empty means the node has neither been + // provisioned nor adopted, and non-empty means steady state. + // +optional + UUID string `json:"uuid,omitempty"` + + // Status is the lifecycle the control plane reports: online, suspended, + // offline, in_creation, in_restart, in_shutdown, unreachable, or timeout. The + // values are the control plane's, which is why they are neither PascalCase nor + // constrained by an Enum here. + // +optional + Status string `json:"status,omitempty"` + + // Health is the health flag the control plane reports. + // +optional + Health bool `json:"health,omitempty"` + + // Hostname is the node hostname as the control plane reports it. + // +optional + Hostname string `json:"hostname,omitempty"` + + // Uptime is the node uptime as the control plane reports it. + // +optional + Uptime string `json:"uptime,omitempty"` + + // Resources groups the reported compute and storage figures. + // +optional + Resources *StorageNodeResources `json:"resources,omitempty"` + + // Ports groups the reported addresses and ports. + // +optional + Ports *StorageNodePorts `json:"ports,omitempty"` + + // FailureDomain is the failure-domain label the control plane actually + // assigned, which is not necessarily the one spec.config.failureDomain + // requested. + // +optional + FailureDomain string `json:"failureDomain,omitempty"` + + // ActiveOpsRef names the StorageNodeOps currently allowed to touch this node. + // Empty when none is running. + // +optional + ActiveOpsRef string `json:"activeOpsRef,omitempty"` + + // LatencyMetrics holds the fio-measured NVMe-oF baseline the volume + // rebalancer reads. + // +optional + LatencyMetrics *NodeLatencyMetrics `json:"latencyMetrics,omitempty"` + + // Message is the reason the phase is what it is: one sentence, replaced as the + // node moves, and never a log. + // +optional + Message string `json:"message,omitempty"` + + // ObservedGeneration is the generation the rest of this status was computed + // from, so a stale status can be told from a current one. + // +optional + ObservedGeneration int64 `json:"observedGeneration,omitempty"` +} + +// v1alpha2 is the storage version in the manifests this repository ships, which +// are the ones a fresh install applies. An upgrade of an existing cluster reaches +// it through the storage rewrite rather than through this marker, for the reason +// storagenodeops_types.go states. +// +kubebuilder:storageversion +// +kubebuilder:object:root=true +// +kubebuilder:subresource:status +// +kubebuilder:resource:scope=Namespaced,shortName=sn +// +kubebuilder:printcolumn:name="Cluster",type=string,JSONPath=".spec.clusterRef" +// +kubebuilder:printcolumn:name="Worker",type=string,JSONPath=".spec.workerNode" +// +kubebuilder:printcolumn:name="Socket",type=string,JSONPath=".spec.socketId" +// +kubebuilder:printcolumn:name="Slot",type=integer,JSONPath=".spec.slot" +// +kubebuilder:printcolumn:name="Phase",type=string,JSONPath=".status.phase" +// +kubebuilder:printcolumn:name="Step",type=string,JSONPath=".status.step.state" +// +kubebuilder:printcolumn:name="Status",type=string,JSONPath=".status.status" +// +kubebuilder:printcolumn:name="Health",type=boolean,JSONPath=".status.health" +// +kubebuilder:printcolumn:name="UUID",type=string,JSONPath=".status.uuid",priority=1 +// +kubebuilder:printcolumn:name="FD",type=string,JSONPath=".status.failureDomain",priority=1 +// +kubebuilder:printcolumn:name="Age",type=date,JSONPath=".metadata.creationTimestamp" + +// StorageNode is one backend storage node: one SPDK process bound to one NUMA +// socket of one Kubernetes worker. One object exists per (workerNode, slot) pair, +// owned by the StorageCluster it belongs to. +type StorageNode struct { + metav1.TypeMeta `json:",inline"` + metav1.ObjectMeta `json:"metadata,omitempty"` + + Spec StorageNodeSpec `json:"spec,omitempty"` + Status StorageNodeStatus `json:"status,omitempty"` +} + +// Hub marks this version as the conversion hub for StorageNode. +func (*StorageNode) Hub() {} + +// +kubebuilder:object:root=true + +// StorageNodeList contains a list of StorageNode. +type StorageNodeList struct { + metav1.TypeMeta `json:",inline"` + metav1.ListMeta `json:"metadata,omitempty"` + Items []StorageNode `json:"items"` +} + +func init() { + SchemeBuilder.Register(&StorageNode{}, &StorageNodeList{}) +} diff --git a/operator/api/v1alpha2/storagenodeops_types.go b/operator/api/v1alpha2/storagenodeops_types.go index d1e0ca8e2..5cdf2d209 100644 --- a/operator/api/v1alpha2/storagenodeops_types.go +++ b/operator/api/v1alpha2/storagenodeops_types.go @@ -1,40 +1,29 @@ -// StorageNodeOps with the four properties the CRD redesign renames or regroups -// on it (design-storagenode.md §6.1 and §6.3): +// StorageNodeOps in the shape design-storagenode.md Appendix B specifies: a +// single operation performed against one StorageNode, which runs to a terminal +// phase and stays afterward as the audit record of what was done. // -// - spec.storageNodeRef becomes spec.nodeRef. The kind is already named for the -// node, so the prefix repeated it. -// - spec.drain becomes spec.remove, matching the action it parameterizes, and -// DrainOpsSpec becomes RemoveSpec with it. -// - spec.targetWorkerNode and spec.newSsdPcie regroup under spec.migrate, so -// that a parameter block belongs to the action that reads it. -// - spec.action becomes a named enum whose values are PascalCase. -// -// The step machine is not here. status.subPhase becoming a status.step object, -// the Aborted phase, spec.abort, the HostMaintenance action, and the drain status -// regrouping are all the Ops shape of design-crd-model.md §9.5 rather than -// renames (design-property-renames.md §2.7). +// What moves against the registered v1alpha1 type is §15.2 of that document. The +// property renames landed first (spec.storageNodeRef becoming spec.nodeRef, +// spec.drain becoming spec.remove, the migrate parameters regrouping under +// spec.migrate, and the action enum becoming PascalCase), and the step machine is +// what arrives here: status.subPhase becomes a status.step holding a declared +// statemachine graph per action, status.triggered goes with nothing replacing it, +// the Aborted phase and the spec.abort that reaches it arrive, the seventh +// HostMaintenance action retires a controller of its own, and the two drain +// counters regroup under status.drain. package v1alpha2 import ( metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" -) - -// StorageNodeOpsPhase is the lifecycle phase of a StorageNodeOps. -// +kubebuilder:validation:Enum=Pending;Running;Succeeded;Failed -type StorageNodeOpsPhase string -const ( - StorageNodeOpsPhasePending StorageNodeOpsPhase = "Pending" - StorageNodeOpsPhaseRunning StorageNodeOpsPhase = "Running" - StorageNodeOpsPhaseSucceeded StorageNodeOpsPhase = "Succeeded" - StorageNodeOpsPhaseFailed StorageNodeOpsPhase = "Failed" + "github.com/simplyblock/atlas/statemachine" ) -// StorageNodeOpsAction is the operation a StorageNodeOps performs. The values are -// PascalCase, as every enum in the group is; v1alpha1 spelled them lowercase, and -// the conversion maps between the two. -// +kubebuilder:validation:Enum=Shutdown;Restart;Suspend;Resume;Remove;Migrate +// StorageNodeOpsAction is the operation a StorageNodeOps performs. Values are +// PascalCase, which is the casing every enum this API group defines carries; +// v1alpha1 spelled them lowercase, and the conversion maps between the two. +// +kubebuilder:validation:Enum=Shutdown;Restart;Suspend;Resume;Remove;Migrate;HostMaintenance type StorageNodeOpsAction string const ( @@ -44,90 +33,128 @@ const ( StorageNodeOpsActionResume StorageNodeOpsAction = "Resume" StorageNodeOpsActionRemove StorageNodeOpsAction = "Remove" StorageNodeOpsActionMigrate StorageNodeOpsAction = "Migrate" + + // StorageNodeOpsActionHostMaintenance takes a node down deliberately so that + // its Kubernetes worker can be drained and rebooted, then brings it back. The + // operator raises it when it sees the worker cordoned, and a user creating + // one by hand behaves identically. + StorageNodeOpsActionHostMaintenance StorageNodeOpsAction = "HostMaintenance" ) -// StorageNodeOpsSubPhase is the active sub-phase during a running op: the drain -// steps when action=Remove, and the Preparing → Migrating → Promoting steps when -// action=Migrate. -// +kubebuilder:validation:Enum=Validating;Suspending;Migrating;Verifying;Removing;Preparing;Restarting;Promoting -type StorageNodeOpsSubPhase string +// StorageNodeOpsPhase is the operation's own progress. Aborted is terminal and +// distinct from Failed, because a canceled operation did not go wrong. +// +kubebuilder:validation:Enum=Pending;Running;Succeeded;Failed;Aborted +type StorageNodeOpsPhase string const ( - StorageNodeOpsSubPhaseValidating StorageNodeOpsSubPhase = "Validating" - StorageNodeOpsSubPhaseSuspending StorageNodeOpsSubPhase = "Suspending" - StorageNodeOpsSubPhaseMigrating StorageNodeOpsSubPhase = "Migrating" - StorageNodeOpsSubPhaseVerifying StorageNodeOpsSubPhase = "Verifying" - StorageNodeOpsSubPhaseRemoving StorageNodeOpsSubPhase = "Removing" - // StorageNodeOpsSubPhasePreparing marks that a migrate op is preparing the - // target worker: cloning per-node config, labeling it into the storage - // plane, and waiting until its storage-node-api pod is Ready and its per-pod - // DNS name is published in the EndpointSlice — the precondition for the - // control-plane restart to resolve node_address. - StorageNodeOpsSubPhasePreparing StorageNodeOpsSubPhase = "Preparing" - // StorageNodeOpsSubPhaseRestarting marks that a migrate op has issued the - // control-plane restart and confirmed the node entered in_restart; it is - // now waiting for the node to come back online on the target host. The - // restart is asynchronous, so the op only advances to Promoting after the - // node has left online (restart started) and returned to online (restart - // finished) — issuing /promote earlier races the in-flight restart's node - // writes and leaves the relocated devices stuck in `new`. - StorageNodeOpsSubPhaseRestarting StorageNodeOpsSubPhase = "Restarting" - // StorageNodeOpsSubPhasePromoting marks that a migrate op has issued the - // control-plane /promote for the relocated node (guards against re-promoting). - StorageNodeOpsSubPhasePromoting StorageNodeOpsSubPhase = "Promoting" + // StorageNodeOpsPhasePending: the operation holds no lock and has issued + // nothing. + StorageNodeOpsPhasePending StorageNodeOpsPhase = "Pending" + + // StorageNodeOpsPhaseRunning: it holds its node's lock and its first side + // effect may have been issued. + StorageNodeOpsPhaseRunning StorageNodeOpsPhase = "Running" + + StorageNodeOpsPhaseSucceeded StorageNodeOpsPhase = "Succeeded" + StorageNodeOpsPhaseFailed StorageNodeOpsPhase = "Failed" + + // StorageNodeOpsPhaseAborted: called off rather than gone wrong. A drain an + // administrator stops after an hour reported as Failed would sit in the same + // bucket as one the control plane rejected. + StorageNodeOpsPhaseAborted StorageNodeOpsPhase = "Aborted" +) + +// StorageNodeOpsStep is one step of a running node operation. The enum is the +// union of every action's steps; which steps belong to which action is declared by +// the graph rather than by this type. +// +kubebuilder:validation:Enum=Requesting;Awaiting;Validating;Suspending;MigratingVolumes;Verifying;Removing;Preparing;Relocating;AwaitingNode;Promoting;Holding;ShuttingDown;Releasing;AwaitingHost;Restarting;Cleanup +type StorageNodeOpsStep string + +const ( + // Shutdown, Restart, Suspend, and Resume. + StorageNodeOpsStepRequesting StorageNodeOpsStep = "Requesting" + StorageNodeOpsStepAwaiting StorageNodeOpsStep = "Awaiting" + + // Remove. + StorageNodeOpsStepValidating StorageNodeOpsStep = "Validating" + StorageNodeOpsStepSuspending StorageNodeOpsStep = "Suspending" + StorageNodeOpsStepMigratingVolumes StorageNodeOpsStep = "MigratingVolumes" + StorageNodeOpsStepVerifying StorageNodeOpsStep = "Verifying" + StorageNodeOpsStepRemoving StorageNodeOpsStep = "Removing" + + // Migrate. + StorageNodeOpsStepPreparing StorageNodeOpsStep = "Preparing" + StorageNodeOpsStepRelocating StorageNodeOpsStep = "Relocating" + StorageNodeOpsStepAwaitingNode StorageNodeOpsStep = "AwaitingNode" + StorageNodeOpsStepPromoting StorageNodeOpsStep = "Promoting" + + // HostMaintenance. + StorageNodeOpsStepHolding StorageNodeOpsStep = "Holding" + StorageNodeOpsStepShuttingDown StorageNodeOpsStep = "ShuttingDown" + StorageNodeOpsStepReleasing StorageNodeOpsStep = "Releasing" + StorageNodeOpsStepAwaitingHost StorageNodeOpsStep = "AwaitingHost" + StorageNodeOpsStepRestarting StorageNodeOpsStep = "Restarting" + StorageNodeOpsStepCleanup StorageNodeOpsStep = "Cleanup" ) // MigrateSpec parameterizes the Migrate action and is ignored by the others. // -// A migration is NOT a drain or a remove: the storage node keeps its backend UUID -// and its partition and logical-volume assignments follow it. The operator issues -// a control-plane restart pointed at the target host's storage-node-api -// (node_address), waits for the node to come back online there, then promotes it -// (starting a rebalance) and re-points the StorageNode's spec.workerNode and the -// owning StorageNodeSet.workerNodes from the source worker to this one. No fresh -// storage node is provisioned and no VolumeMigration CRs are created. +// A migration is not a removal followed by an add: the node keeps its backend +// UUID, its partitions, and its logical-volume assignments, and what changes is +// the machine the SPDK process runs on. No PersistentVolumeOps is created. type MigrateSpec struct { - // TargetWorkerNode is the Kubernetes worker hostname the storage node is - // relocated onto. + // TargetWorkerNode is the Kubernetes worker the node is relocated onto. // +kubebuilder:validation:Required // +k8s:immutable TargetWorkerNode string `json:"targetWorkerNode"` - // NewSsdPcie lists additional NVMe PCIe addresses to bind on the target host - // during a migration. Passed through to the control-plane restart as - // new_ssd_pcie. + // NewSsdPcie lists additional NVMe PCI addresses to bind on the target host, + // passed through to the control-plane restart as new_ssd_pcie and merged into + // the node's effective allow list so they survive a later rebuild. // +optional NewSsdPcie []string `json:"newSsdPcie,omitempty"` } // RemoveSpec parameterizes the Remove action and is ignored by the others. type RemoveSpec struct { - // SystemVolumeFilterRegex is a Go regular expression matched against backend - // volume names. Matching volumes are treated as system volumes: excluded from - // drain migration and deleted inline during the Verifying phase. - // Defaults to `^sb-fio-baseline-.*`. + // SystemVolumeFilterRegex matches backend volume names that are system + // volumes: excluded from the drain's migration and deleted during + // verification rather than blocking it. + // +kubebuilder:default=`^sb-fio-baseline-.*` // +optional SystemVolumeFilterRegex *string `json:"systemVolumeFilterRegex,omitempty"` } -// StorageNodeOpsSpec defines the desired state of a StorageNodeOps. +// StorageNodeOpsSpec is one operation to perform against one StorageNode. type StorageNodeOpsSpec struct { - // NodeRef is the name of the target StorageNode. Immutable. + // NodeRef names the StorageNode this operation acts on. The operation never + // owns its target, because deleting the record of an operation must not delete + // the node it operated on. // +kubebuilder:validation:Required // +k8s:immutable NodeRef string `json:"nodeRef"` - // Action is the operation to perform. Immutable. + // Action is the operation to perform. // +kubebuilder:validation:Required // +k8s:immutable Action StorageNodeOpsAction `json:"action"` - // Force enables forced execution where the backend supports it. + // Abort asks a running operation to stop at its next step and unwind. It is + // the only mutable field on this spec, because it is the only thing about an + // operation that can legitimately be decided after it started. Whether an + // abort is expressible from the current step is declared by that action's + // graph rather than checked here. + // +optional + Abort bool `json:"abort,omitempty"` + + // Force passes the control plane's force flag where the action supports it. + // Migrate defaults it to true, because the control plane rejects a non-forced + // restart of a node that is not already offline. // +optional Force *bool `json:"force,omitempty"` - // ReattachVolume reattaches volumes during the node restart. - // Applicable when action=Restart or action=Migrate. + // ReattachVolume asks the control plane to reattach this node's volumes as + // part of a restart. Applies to Restart, Migrate, and HostMaintenance. // +optional ReattachVolume *bool `json:"reattachVolume,omitempty"` @@ -140,8 +167,8 @@ type StorageNodeOpsSpec struct { Remove *RemoveSpec `json:"remove,omitempty"` } -// MigrateParams returns the Migrate block, or a zero-valued one when the -// operation did not set it. +// MigrateParams returns the Migrate block, or a zero-valued one when the operation +// did not set it. // // The block is optional on every action, so four of the migrate workflow's steps // would otherwise each need their own nil check. A Migrate operation that arrives @@ -155,38 +182,63 @@ func (s *StorageNodeOpsSpec) MigrateParams() MigrateSpec { return *s.Migrate } -// StorageNodeOpsStatus holds the observed state of a StorageNodeOps. +// RemoveParams returns the Remove block, or a zero-valued one when the operation +// did not set it, for the reason MigrateParams does. +func (s *StorageNodeOpsSpec) RemoveParams() RemoveSpec { + if s.Remove == nil { + return RemoveSpec{} + } + return *s.Remove +} + +// DrainStatus is the drain's progress over the volumes on the node being removed. +// Neither field takes omitempty: zero is meaningful for both, and a field that +// disappears at zero makes "nothing to move" and "not yet counted" the same wire +// value. +type DrainStatus struct { + // VolumesTotal is the number of PV-managed volumes the drain has to move, + // written once at the end of Validating and not modified afterward. + // +kubebuilder:validation:Minimum=0 + VolumesTotal int32 `json:"volumesTotal"` + + // VolumesMigrated is how many of them have completed. + // +kubebuilder:validation:Minimum=0 + VolumesMigrated int32 `json:"volumesMigrated"` +} + +// StorageNodeOpsStatus is the observed state of one node operation. type StorageNodeOpsStatus struct { - // Phase is the high-level lifecycle phase. + // Phase is the operation's own progress. // +optional Phase StorageNodeOpsPhase `json:"phase,omitempty"` - // SubPhase tracks the active drain step when action=Remove and phase=Running. + // Step is the position of the running action's state machine, as the shared + // statemachine.KubeSnapshot. The rule is what an Enum marker would do if a + // marker could reach a field of a shared type. + // +kubebuilder:validation:XValidation:rule="!has(self.state) || self.state in ['Requesting','Awaiting','Validating','Suspending','MigratingVolumes','Verifying','Removing','Preparing','Relocating','AwaitingNode','Promoting','Holding','ShuttingDown','Releasing','AwaitingHost','Restarting','Cleanup']",message="unknown step" // +optional - SubPhase StorageNodeOpsSubPhase `json:"subPhase,omitempty"` + Step statemachine.KubeSnapshot `json:"step,omitempty"` - // Message is a human-readable description of the current state or failure reason. + // Message is the reason the phase is what it is: one sentence, replaced as the + // operation moves, and never a log. // +optional Message string `json:"message,omitempty"` - // VolumesMigrated is the count of volumes successfully migrated (drain only). - // +optional - VolumesMigrated int `json:"volumesMigrated,omitempty"` - - // VolumesPending is the count of volumes awaiting migration (drain only). + // Drain is the drain's progress over the node's volumes, set only for action + // Remove. // +optional - VolumesPending int `json:"volumesPending,omitempty"` + Drain *DrainStatus `json:"drain,omitempty"` - // Triggered indicates the backend action POST has been sent (used during - // Suspending to avoid duplicate POSTs across reconcile iterations). + // ObservedGeneration is the generation the rest of this status was computed + // from, so a stale status can be told from a current one. // +optional - Triggered bool `json:"triggered,omitempty"` + ObservedGeneration int64 `json:"observedGeneration,omitempty"` - // StartedAt is when the operation began. + // StartedAt is when the operation acquired its target's lock. // +optional StartedAt *metav1.Time `json:"startedAt,omitempty"` - // CompletedAt is when the operation finished (successfully or not). + // CompletedAt is when it reached a terminal phase. // +optional CompletedAt *metav1.Time `json:"completedAt,omitempty"` } @@ -209,14 +261,13 @@ type StorageNodeOpsStatus struct { // +kubebuilder:printcolumn:name="Node",type=string,JSONPath=".spec.nodeRef" // +kubebuilder:printcolumn:name="Action",type=string,JSONPath=".spec.action" // +kubebuilder:printcolumn:name="Phase",type=string,JSONPath=".status.phase" -// +kubebuilder:printcolumn:name="SubPhase",type=string,JSONPath=".status.subPhase" -// +kubebuilder:printcolumn:name="Message",type=string,JSONPath=".status.message" +// +kubebuilder:printcolumn:name="Step",type=string,JSONPath=".status.step.state" +// +kubebuilder:printcolumn:name="Message",type=string,JSONPath=".status.message",priority=1 // +kubebuilder:printcolumn:name="Age",type=date,JSONPath=".metadata.creationTimestamp" -// StorageNodeOps is a one-shot operational CR targeting a single StorageNode. -// Analogous to a Kubernetes Job — it drives an action (Shutdown, Restart, Suspend, -// Resume, Remove, Migrate) to completion and records the result. Only one -// StorageNodeOps can be active per StorageNode at a time. +// StorageNodeOps is a single operation performed against one StorageNode. It runs +// to a terminal phase and stays afterward as the audit record of what was done, to +// which node, with which parameters, and how it ended. type StorageNodeOps struct { metav1.TypeMeta `json:",inline"` metav1.ObjectMeta `json:"metadata,omitempty"` diff --git a/operator/api/v1alpha2/zz_generated.deepcopy.go b/operator/api/v1alpha2/zz_generated.deepcopy.go index efdc3bc81..6ce441753 100644 --- a/operator/api/v1alpha2/zz_generated.deepcopy.go +++ b/operator/api/v1alpha2/zz_generated.deepcopy.go @@ -581,6 +581,21 @@ func (in *DiscoverSpec) DeepCopy() *DiscoverSpec { return out } +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *DrainStatus) DeepCopyInto(out *DrainStatus) { + *out = *in +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new DrainStatus. +func (in *DrainStatus) DeepCopy() *DrainStatus { + if in == nil { + return nil + } + out := new(DrainStatus) + in.DeepCopyInto(out) + return out +} + // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *DriverTLS) DeepCopyInto(out *DriverTLS) { *out = *in @@ -721,6 +736,25 @@ func (in *NodeGroup) DeepCopy() *NodeGroup { return out } +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *NodeLatencyMetrics) DeepCopyInto(out *NodeLatencyMetrics) { + *out = *in + if in.BaselineMeasuredAt != nil { + in, out := &in.BaselineMeasuredAt, &out.BaselineMeasuredAt + *out = (*in).DeepCopy() + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new NodeLatencyMetrics. +func (in *NodeLatencyMetrics) DeepCopy() *NodeLatencyMetrics { + if in == nil { + return nil + } + out := new(NodeLatencyMetrics) + in.DeepCopyInto(out) + return out +} + // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *NodeLoadMetrics) DeepCopyInto(out *NodeLoadMetrics) { *out = *in @@ -1739,6 +1773,11 @@ func (in *StorageClusterSpec) DeepCopyInto(out *StorageClusterSpec) { *out = new(int32) **out = **in } + if in.StorageNodes != nil { + in, out := &in.StorageNodes, &out.StorageNodes + *out = new(StorageNodesSpec) + (*in).DeepCopyInto(*out) + } if in.Backup != nil { in, out := &in.Backup, &out.Backup *out = new(BackupStoreSpec) @@ -1930,6 +1969,150 @@ func (in *StorageDeviceStatus) DeepCopy() *StorageDeviceStatus { return out } +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *StorageNode) DeepCopyInto(out *StorageNode) { + *out = *in + out.TypeMeta = in.TypeMeta + in.ObjectMeta.DeepCopyInto(&out.ObjectMeta) + in.Spec.DeepCopyInto(&out.Spec) + in.Status.DeepCopyInto(&out.Status) +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new StorageNode. +func (in *StorageNode) DeepCopy() *StorageNode { + if in == nil { + return nil + } + out := new(StorageNode) + in.DeepCopyInto(out) + return out +} + +// DeepCopyObject is an autogenerated deepcopy function, copying the receiver, creating a new runtime.Object. +func (in *StorageNode) DeepCopyObject() runtime.Object { + if c := in.DeepCopy(); c != nil { + return c + } + return nil +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *StorageNodeCapacity) DeepCopyInto(out *StorageNodeCapacity) { + *out = *in + if in.TotalBytes != nil { + in, out := &in.TotalBytes, &out.TotalBytes + *out = new(int64) + **out = **in + } + if in.UsedBytes != nil { + in, out := &in.UsedBytes, &out.UsedBytes + *out = new(int64) + **out = **in + } + if in.SampledAt != nil { + in, out := &in.SampledAt, &out.SampledAt + *out = (*in).DeepCopy() + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new StorageNodeCapacity. +func (in *StorageNodeCapacity) DeepCopy() *StorageNodeCapacity { + if in == nil { + return nil + } + out := new(StorageNodeCapacity) + in.DeepCopyInto(out) + return out +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *StorageNodeConfig) DeepCopyInto(out *StorageNodeConfig) { + *out = *in + in.Sizing.DeepCopyInto(&out.Sizing) + if in.JournalManager != nil { + in, out := &in.JournalManager, &out.JournalManager + *out = new(JournalManagerSpec) + (*in).DeepCopyInto(*out) + } + if in.DeviceNames != nil { + in, out := &in.DeviceNames, &out.DeviceNames + *out = make([]string, len(*in)) + copy(*out, *in) + } + if in.PcieAllowList != nil { + in, out := &in.PcieAllowList, &out.PcieAllowList + *out = make([]string, len(*in)) + copy(*out, *in) + } + if in.PcieDenyList != nil { + in, out := &in.PcieDenyList, &out.PcieDenyList + *out = make([]string, len(*in)) + copy(*out, *in) + } + if in.Expand != nil { + in, out := &in.Expand, &out.Expand + *out = new(bool) + **out = **in + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new StorageNodeConfig. +func (in *StorageNodeConfig) DeepCopy() *StorageNodeConfig { + if in == nil { + return nil + } + out := new(StorageNodeConfig) + in.DeepCopyInto(out) + return out +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *StorageNodeDevices) DeepCopyInto(out *StorageNodeDevices) { + *out = *in +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new StorageNodeDevices. +func (in *StorageNodeDevices) DeepCopy() *StorageNodeDevices { + if in == nil { + return nil + } + out := new(StorageNodeDevices) + in.DeepCopyInto(out) + return out +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *StorageNodeList) DeepCopyInto(out *StorageNodeList) { + *out = *in + out.TypeMeta = in.TypeMeta + in.ListMeta.DeepCopyInto(&out.ListMeta) + if in.Items != nil { + in, out := &in.Items, &out.Items + *out = make([]StorageNode, len(*in)) + for i := range *in { + (*in)[i].DeepCopyInto(&(*out)[i]) + } + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new StorageNodeList. +func (in *StorageNodeList) DeepCopy() *StorageNodeList { + if in == nil { + return nil + } + out := new(StorageNodeList) + in.DeepCopyInto(out) + return out +} + +// DeepCopyObject is an autogenerated deepcopy function, copying the receiver, creating a new runtime.Object. +func (in *StorageNodeList) DeepCopyObject() runtime.Object { + if c := in.DeepCopy(); c != nil { + return c + } + return nil +} + // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *StorageNodeOps) DeepCopyInto(out *StorageNodeOps) { *out = *in @@ -2027,6 +2210,12 @@ func (in *StorageNodeOpsSpec) DeepCopy() *StorageNodeOpsSpec { // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *StorageNodeOpsStatus) DeepCopyInto(out *StorageNodeOpsStatus) { *out = *in + in.Step.DeepCopyInto(&out.Step) + if in.Drain != nil { + in, out := &in.Drain, &out.Drain + *out = new(DrainStatus) + **out = **in + } if in.StartedAt != nil { in, out := &in.StartedAt, &out.StartedAt *out = (*in).DeepCopy() @@ -2047,6 +2236,222 @@ func (in *StorageNodeOpsStatus) DeepCopy() *StorageNodeOpsStatus { return out } +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *StorageNodePorts) DeepCopyInto(out *StorageNodePorts) { + *out = *in + if in.NvmeOf != nil { + in, out := &in.NvmeOf, &out.NvmeOf + *out = new(int32) + **out = **in + } + if in.Lvol != nil { + in, out := &in.Lvol, &out.Lvol + *out = new(int32) + **out = **in + } + if in.Rpc != nil { + in, out := &in.Rpc, &out.Rpc + *out = new(int32) + **out = **in + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new StorageNodePorts. +func (in *StorageNodePorts) DeepCopy() *StorageNodePorts { + if in == nil { + return nil + } + out := new(StorageNodePorts) + in.DeepCopyInto(out) + return out +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *StorageNodeResources) DeepCopyInto(out *StorageNodeResources) { + *out = *in + if in.CPU != nil { + in, out := &in.CPU, &out.CPU + *out = new(int32) + **out = **in + } + if in.Volumes != nil { + in, out := &in.Volumes, &out.Volumes + *out = new(int32) + **out = **in + } + if in.Devices != nil { + in, out := &in.Devices, &out.Devices + *out = new(StorageNodeDevices) + **out = **in + } + if in.Capacity != nil { + in, out := &in.Capacity, &out.Capacity + *out = new(StorageNodeCapacity) + (*in).DeepCopyInto(*out) + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new StorageNodeResources. +func (in *StorageNodeResources) DeepCopy() *StorageNodeResources { + if in == nil { + return nil + } + out := new(StorageNodeResources) + in.DeepCopyInto(out) + return out +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *StorageNodeSizing) DeepCopyInto(out *StorageNodeSizing) { + *out = *in + if in.VCPUCount != nil { + in, out := &in.VCPUCount, &out.VCPUCount + *out = new(int32) + **out = **in + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new StorageNodeSizing. +func (in *StorageNodeSizing) DeepCopy() *StorageNodeSizing { + if in == nil { + return nil + } + out := new(StorageNodeSizing) + in.DeepCopyInto(out) + return out +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *StorageNodeSpec) DeepCopyInto(out *StorageNodeSpec) { + *out = *in + if in.NodeIndex != nil { + in, out := &in.NodeIndex, &out.NodeIndex + *out = new(int32) + **out = **in + } + if in.Slot != nil { + in, out := &in.Slot, &out.Slot + *out = new(int32) + **out = **in + } + in.Config.DeepCopyInto(&out.Config) +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new StorageNodeSpec. +func (in *StorageNodeSpec) DeepCopy() *StorageNodeSpec { + if in == nil { + return nil + } + out := new(StorageNodeSpec) + in.DeepCopyInto(out) + return out +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *StorageNodeStatus) DeepCopyInto(out *StorageNodeStatus) { + *out = *in + in.Step.DeepCopyInto(&out.Step) + if in.Resources != nil { + in, out := &in.Resources, &out.Resources + *out = new(StorageNodeResources) + (*in).DeepCopyInto(*out) + } + if in.Ports != nil { + in, out := &in.Ports, &out.Ports + *out = new(StorageNodePorts) + (*in).DeepCopyInto(*out) + } + if in.LatencyMetrics != nil { + in, out := &in.LatencyMetrics, &out.LatencyMetrics + *out = new(NodeLatencyMetrics) + (*in).DeepCopyInto(*out) + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new StorageNodeStatus. +func (in *StorageNodeStatus) DeepCopy() *StorageNodeStatus { + if in == nil { + return nil + } + out := new(StorageNodeStatus) + in.DeepCopyInto(out) + return out +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *StorageNodesSpec) DeepCopyInto(out *StorageNodesSpec) { + *out = *in + if in.DataInterfaces != nil { + in, out := &in.DataInterfaces, &out.DataInterfaces + *out = make([]string, len(*in)) + copy(*out, *in) + } + if in.SocketsToUse != nil { + in, out := &in.SocketsToUse, &out.SocketsToUse + *out = make([]string, len(*in)) + copy(*out, *in) + } + if in.NodesPerSocket != nil { + in, out := &in.NodesPerSocket, &out.NodesPerSocket + *out = new(int32) + **out = **in + } + if in.MaxParallelNodeAdds != nil { + in, out := &in.MaxParallelNodeAdds, &out.MaxParallelNodeAdds + *out = new(int32) + **out = **in + } + if in.EnableJournalDevice != nil { + in, out := &in.EnableJournalDevice, &out.EnableJournalDevice + *out = new(bool) + **out = **in + } + if in.EnableFormat4K != nil { + in, out := &in.EnableFormat4K, &out.EnableFormat4K + *out = new(bool) + **out = **in + } + if in.EnableCpuTopology != nil { + in, out := &in.EnableCpuTopology, &out.EnableCpuTopology + *out = new(bool) + **out = **in + } + if in.EnableKubeletConfiguration != nil { + in, out := &in.EnableKubeletConfiguration, &out.EnableKubeletConfiguration + *out = new(bool) + **out = **in + } + if in.UbuntuHost != nil { + in, out := &in.UbuntuHost, &out.UbuntuHost + *out = new(bool) + **out = **in + } + if in.OpenShiftCluster != nil { + in, out := &in.OpenShiftCluster, &out.OpenShiftCluster + *out = new(bool) + **out = **in + } + if in.Tolerations != nil { + in, out := &in.Tolerations, &out.Tolerations + *out = make([]v1.Toleration, len(*in)) + for i := range *in { + (*in)[i].DeepCopyInto(&(*out)[i]) + } + } + in.ContainerResources.DeepCopyInto(&out.ContainerResources) + in.InitContainerResources.DeepCopyInto(&out.InitContainerResources) +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new StorageNodesSpec. +func (in *StorageNodesSpec) DeepCopy() *StorageNodesSpec { + if in == nil { + return nil + } + out := new(StorageNodesSpec) + in.DeepCopyInto(out) + return out +} + // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *StoragePool) DeepCopyInto(out *StoragePool) { *out = *in diff --git a/operator/cmd/main.go b/operator/cmd/main.go index b18ed7985..469ff09f1 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -59,6 +59,7 @@ import ( clustercontroller "github.com/simplyblock/simplyblock-operator/internal/controllers/cluster" "github.com/simplyblock/simplyblock-operator/internal/controllers/deployment" "github.com/simplyblock/simplyblock-operator/internal/controllers/driver" + nodecontroller "github.com/simplyblock/simplyblock-operator/internal/controllers/node" "github.com/simplyblock/simplyblock-operator/internal/controllers/pool" "github.com/simplyblock/simplyblock-operator/internal/csilink" "github.com/simplyblock/simplyblock-operator/internal/utils" @@ -345,6 +346,18 @@ func main() { // LeaderOnly: these subscriptions feed reconcilers that write StorageDevice // and StorageNode objects, and two replicas writing the same object would // fight over it. + // The storage-plane side of a node, shared by the three reconcilers that touch + // it. The EndpointSlice a migration blocks on is read straight from the API + // server: a stale informer cache can miss a freshly published endpoint, and a + // migration would then wait on DNS forever while the name has in fact resolved + // for minutes (design-storagenode.md §5.4). + storageNodeWorkload := &nodecontroller.Workload{ + Client: mgr.GetClient(), + Uncached: mgr.GetAPIReader(), + TLSEnabled: tlsEnabled, + TLSMutualEnabled: tlsMutualEnabled, + } + cpSubscriptions := cpinformer.NewSubscriptionManager(streamCfg, ctrl.Log.WithName("cpinformer"), cpinformer.LeaderOnly) deviceSubscription := subscriptions.NewDeviceSubscription() deviceScopes := cpSubscriptions.AddSubscription(deviceSubscription) @@ -375,13 +388,13 @@ func main() { // so it comes from Prometheus rather than from the API or the stream. An // endpoint that cannot be reached leaves the capacity absent from the // status and everything else in it correct. - var nodeCapacity controller.NodeCapacitySource + var nodeCapacity nodecontroller.NodeCapacitySource if provider, err := atlasprom.New(prometheusURL); err != nil { setupLog.Error(err, "storage-node capacity will be absent", "prometheusURL", prometheusURL) } else { nodeCapacity = provider } - if err := (&controller.StorageDeviceReconciler{ + if err := (&nodecontroller.StorageDeviceReconciler{ Client: mgr.GetClient(), Scheme: mgr.GetScheme(), Devices: deviceSubscription, @@ -394,13 +407,13 @@ func main() { // continuously and comes from the same Prometheus the node capacity does. The // collector publishes it as a gauge on a timer and warns about a device over // its cluster's threshold, without writing any of it to an object. - var deviceCapacity controller.DeviceCapacitySource + var deviceCapacity nodecontroller.DeviceCapacitySource if provider, err := atlasprom.New(prometheusURL); err != nil { setupLog.Error(err, "storage-device capacity will be absent", "prometheusURL", prometheusURL) } else { deviceCapacity = provider } - if err := mgr.Add(&controller.StorageDeviceCollector{ + if err := mgr.Add(&nodecontroller.StorageDeviceCollector{ Client: mgr.GetClient(), Recorder: mgr.GetEventRecorder("storagedevice-collector"), Capacity: deviceCapacity, @@ -492,16 +505,17 @@ func main() { setupLog.Error(err, "unable to create controller", "controller", "StorageCluster") os.Exit(1) } - if err := (&controller.StorageNodeSetReconciler{ + if err := (&nodecontroller.StorageNodeWorkloadReconciler{ Client: mgr.GetClient(), Scheme: mgr.GetScheme(), + Recorder: mgr.GetEventRecorder("storagenode-workload-controller"), Namespace: operatorNamespace, TLSEnabled: tlsEnabled, TLSProvider: tlsProvider, TLSMutualEnabled: tlsMutualEnabled, - Recorder: mgr.GetEventRecorder("storagenodeset-controller"), + Workload: storageNodeWorkload, }).SetupWithManager(mgr); err != nil { - setupLog.Error(err, "unable to create controller", "controller", "StorageNodeSet") + setupLog.Error(err, "unable to create controller", "controller", "StorageNodeWorkload") os.Exit(1) } if err := (&pool.StoragePoolReconciler{ @@ -528,16 +542,6 @@ func main() { setupLog.Error(err, "unable to create controller", "controller", "Task") os.Exit(1) } - if err := (&controller.NodeDrainCoordinatorReconciler{ - Client: mgr.GetClient(), - Scheme: mgr.GetScheme(), - ManagerNodeName: os.Getenv("NODE_NAME"), - TLSEnabled: tlsEnabled, - TLSMutualEnabled: tlsMutualEnabled, - }).SetupWithManager(mgr); err != nil { - setupLog.Error(err, "unable to create controller", "controller", "NodeDrainCoordinator") - os.Exit(1) - } // The data-protection band. The mirror is what creates every StorageBackup // object, so nothing here takes a backup: a policy tells the control plane // to, and the operations kind reads one back into a claim. @@ -618,26 +622,30 @@ func main() { setupLog.Error(err, "unable to create controller", "controller", "PersistentVolumeClaim") os.Exit(1) } - if err := (&controller.StorageNodeReconciler{ - Client: mgr.GetClient(), - Scheme: mgr.GetScheme(), - Recorder: mgr.GetEventRecorder("storagenode-controller"), - TLSEnabled: tlsEnabled, - TLSMutualEnabled: tlsMutualEnabled, - DeviceScopes: deviceScopes, - NodeRegistries: []controller.NodeObjectRegistry{ + if err := (&nodecontroller.StorageNodeReconciler{ + Client: mgr.GetClient(), + Scheme: mgr.GetScheme(), + Recorder: mgr.GetEventRecorder("storagenode-controller"), + API: nodecontroller.NewControlPlane(), + Nodes: nodeSubscription, + Registries: []nodecontroller.NodeObjectRegistry{ deviceSubscription, nodeSubscription, }, - Nodes: nodeSubscription, - Capacity: nodeCapacity, + DeviceScopes: deviceScopes, + Capacity: nodeCapacity, + Workload: storageNodeWorkload, }).SetupWithManager(mgr); err != nil { setupLog.Error(err, "unable to create controller", "controller", "StorageNode") os.Exit(1) } - if err := (&controller.StorageNodeOpsReconciler{ + if err := (&nodecontroller.StorageNodeOpsReconciler{ Client: mgr.GetClient(), Scheme: mgr.GetScheme(), Recorder: mgr.GetEventRecorder("storagenodeops-controller"), + API: nodecontroller.NewControlPlane(), + Nodes: nodeSubscription, + Clusters: clusterSubscription, + Workload: storageNodeWorkload, }).SetupWithManager(mgr); err != nil { setupLog.Error(err, "unable to create controller", "controller", "StorageNodeOps") os.Exit(1) @@ -771,8 +779,11 @@ func main() { &webhook.Admission{Handler: &internalwebhook.SimplyblockRebalancerInjector{Client: mgr.GetClient()}}) setupLog.Info("registered simplyblock-rebalancer mutating webhook") - mgr.GetWebhookServer().Register("/validate-storage-simplyblock-io-v1alpha1-storagenode", - &webhook.Admission{Handler: &internalwebhook.StorageNodeValidator{OperatorNamespace: operatorNamespace}}) + mgr.GetWebhookServer().Register("/validate-storage-simplyblock-io-v1alpha2-storagenode", + &webhook.Admission{Handler: &internalwebhook.StorageNodeValidator{ + Client: mgr.GetClient(), + OperatorNamespace: operatorNamespace, + }}) setupLog.Info("registered storagenode validating webhook") mgr.GetWebhookServer().Register("/validate-storage-simplyblock-io-v1alpha1-replicationops", diff --git a/operator/config/conversion-webhook/rbac.yaml b/operator/config/conversion-webhook/rbac.yaml index 6f3f16a49..6b3096bef 100644 --- a/operator/config/conversion-webhook/rbac.yaml +++ b/operator/config/conversion-webhook/rbac.yaml @@ -36,6 +36,7 @@ rules: - storageclusterops.storage.simplyblock.io - storageclusters.storage.simplyblock.io - storagenodeops.storage.simplyblock.io + - storagenodes.storage.simplyblock.io - storagepools.storage.simplyblock.io verbs: ["get", "update", "patch"] --- diff --git a/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml b/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml index 30144c3c1..695d12254 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml @@ -259,11 +259,14 @@ spec: description: Count is the number of journal managers to configure. format: int32 + minimum: 1 type: integer percentPerDevice: - description: PercentPerDevice is the journal manager - capacity percentage per device. + description: PercentPerDevice is the share of each + device given to the journal. format: int32 + maximum: 100 + minimum: 1 type: integer type: object mgmtInterface: diff --git a/operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml b/operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml index 278c299a9..99cc036e1 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml @@ -856,6 +856,281 @@ spec: x-kubernetes-validations: - message: field is immutable rule: self == oldSelf + storageNodes: + description: |- + StorageNodes is the Kubernetes workload the cluster's storage nodes run + as, and the cluster owns every object in it by controller reference: a + cluster deleted takes its DaemonSet, Services, certificate, and per-node + ConfigMap with it. One workload serves the whole cluster, because growth is + nodes rather than sets and what differs between hardware generations is per + node already. + properties: + containerResources: + description: |- + ContainerResources sets requests and limits for the storage-node container. + Unset enforces no limits. + properties: + claims: + description: |- + Claims lists the names of resources, defined in spec.resourceClaims, + that are used by this container. + + This field depends on the + DynamicResourceAllocation feature gate. + + This field is immutable. It can only be set for containers. + items: + description: ResourceClaim references one entry in PodSpec.ResourceClaims. + properties: + name: + description: |- + Name must match the name of one entry in pod.spec.resourceClaims of + the Pod where this field is used. It makes that resource available + inside a container. + type: string + request: + description: |- + Request is the name chosen for a request in the referenced claim. + If empty, everything from the claim is made available, otherwise + only the result of this request. + type: string + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + limits: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Limits describes the maximum amount of compute resources allowed. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + requests: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Requests describes the minimum amount of compute resources required. + If Requests is omitted for a container, it defaults to Limits if that is explicitly specified, + otherwise to an implementation-defined value. Requests cannot exceed Limits. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + type: object + dataInterfaces: + description: DataInterfaces are the data-plane network interfaces. + items: + type: string + type: array + enableCpuTopology: + description: EnableCpuTopology turns on topology-aware CPU assignment. + type: boolean + enableFormat4K: + description: |- + EnableFormat4K formats NVMe devices to a 4K block size where the device + supports it. + type: boolean + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + enableJournalDevice: + description: |- + EnableJournalDevice dedicates the smallest NVMe device on each node to the + journal manager, instead of carving a journal partition out of every + device. + type: boolean + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + enableKubeletConfiguration: + description: |- + EnableKubeletConfiguration lets the storage node apply the kubelet + configuration changes it needs. Off by default, which is the behavior the + retired skipKubeletConfiguration expressed by being set. + type: boolean + image: + description: |- + Image is the storage-node container image. Defaults to the ControlPlane + singleton's spec.image when unset, so a deployment states the version once. + pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ + type: string + imagePullPolicy: + default: IfNotPresent + description: ImagePullPolicy controls when that image is pulled. + enum: + - Always + - Never + - IfNotPresent + type: string + initContainerResources: + description: InitContainerResources does the same for the init container. + properties: + claims: + description: |- + Claims lists the names of resources, defined in spec.resourceClaims, + that are used by this container. + + This field depends on the + DynamicResourceAllocation feature gate. + + This field is immutable. It can only be set for containers. + items: + description: ResourceClaim references one entry in PodSpec.ResourceClaims. + properties: + name: + description: |- + Name must match the name of one entry in pod.spec.resourceClaims of + the Pod where this field is used. It makes that resource available + inside a container. + type: string + request: + description: |- + Request is the name chosen for a request in the referenced claim. + If empty, everything from the claim is made available, otherwise + only the result of this request. + type: string + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + limits: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Limits describes the maximum amount of compute resources allowed. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + requests: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Requests describes the minimum amount of compute resources required. + If Requests is omitted for a container, it defaults to Limits if that is explicitly specified, + otherwise to an implementation-defined value. Requests cannot exceed Limits. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + type: object + maxParallelNodeAdds: + default: 1 + description: |- + MaxParallelNodeAdds limits how many workers may be in the node-add process + at once, counted by distinct worker rather than by object so that a + two-socket host consumes one slot. Workers hosting a FoundationDB pod are + always sequential regardless of this value, because a node add reboots the + host and two simultaneous FoundationDB reboots reduce the control plane's + own fault tolerance. + format: int32 + minimum: 1 + type: integer + mgmtInterface: + description: MgmtInterface is the management network interface storage nodes bind. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + nodesPerSocket: + default: 1 + description: NodesPerSocket is how many storage nodes run per NUMA socket. + format: int32 + minimum: 1 + type: integer + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + openShiftCluster: + description: OpenShiftCluster states that the Kubernetes distribution is OpenShift. + type: boolean + openShiftMachineConfigPool: + default: worker + description: |- + OpenShiftMachineConfigPool names the pool generated MachineConfig objects + are labeled into. + type: string + reservedSystemCPU: + description: ReservedSystemCPU is the CPU set held back from SPDK for system workloads. + type: string + socketsToUse: + description: |- + SocketsToUse restricts deployment to selected NUMA sockets. Empty means + socket 0 alone. + items: + type: string + type: array + tolerations: + description: Tolerations are applied to the storage-node pods. + items: + description: |- + The pod this Toleration is attached to tolerates any taint that matches + the triple using the matching operator . + properties: + effect: + description: |- + Effect indicates the taint effect to match. Empty means match all taint effects. + When specified, allowed values are NoSchedule, PreferNoSchedule and NoExecute. + type: string + key: + description: |- + Key is the taint key that the toleration applies to. Empty means match all taint keys. + If the key is empty, operator must be Exists; this combination means to match all values and all keys. + type: string + operator: + description: |- + Operator represents a key's relationship to the value. + Valid operators are Exists, Equal, Lt, and Gt. Defaults to Equal. + Exists is equivalent to wildcard for value, so that a pod can + tolerate all taints of a particular category. + Lt and Gt perform numeric comparisons (requires feature gate TaintTolerationComparisonOperators). + type: string + tolerationSeconds: + description: |- + TolerationSeconds represents the period of time the toleration (which must be + of effect NoExecute, otherwise this field is ignored) tolerates the taint. By default, + it is not set, which means tolerate the taint forever (do not evict). Zero and + negative values will be treated as 0 (evict immediately) by the system. + format: int64 + type: integer + value: + description: |- + Value is the taint value the toleration matches to. + If the operator is Exists, the value should be empty, otherwise just a regular string. + type: string + type: object + type: array + ubuntuHost: + description: |- + UbuntuHost states that the worker's host OS is Ubuntu, which changes how + the node configures huge pages and the kernel modules it loads. + type: boolean + type: object + x-kubernetes-validations: + - message: field mgmtInterface is immutable once set + rule: '!has(oldSelf.mgmtInterface) || has(self.mgmtInterface)' + - message: field nodesPerSocket is immutable once set + rule: '!has(oldSelf.nodesPerSocket) || has(self.nodesPerSocket)' + - message: field enableJournalDevice is immutable once set + rule: '!has(oldSelf.enableJournalDevice) || has(self.enableJournalDevice)' + - message: field enableFormat4K is immutable once set + rule: '!has(oldSelf.enableFormat4K) || has(self.enableFormat4K)' stripe: description: |- Stripe is the erasure-coding layout every volume in the cluster is diff --git a/operator/config/crd/bases/storage.simplyblock.io_storagenodeops.yaml b/operator/config/crd/bases/storage.simplyblock.io_storagenodeops.yaml index a74479f86..9d63d8ae7 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_storagenodeops.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_storagenodeops.yaml @@ -195,11 +195,12 @@ spec: - jsonPath: .status.phase name: Phase type: string - - jsonPath: .status.subPhase - name: SubPhase + - jsonPath: .status.step.state + name: Step type: string - jsonPath: .status.message name: Message + priority: 1 type: string - jsonPath: .metadata.creationTimestamp name: Age @@ -208,10 +209,9 @@ spec: schema: openAPIV3Schema: description: |- - StorageNodeOps is a one-shot operational CR targeting a single StorageNode. - Analogous to a Kubernetes Job — it drives an action (Shutdown, Restart, Suspend, - Resume, Remove, Migrate) to completion and records the result. Only one - StorageNodeOps can be active per StorageNode at a time. + StorageNodeOps is a single operation performed against one StorageNode. It runs + to a terminal phase and stays afterward as the audit record of what was done, to + which node, with which parameters, and how it ended. properties: apiVersion: description: |- @@ -231,10 +231,18 @@ spec: metadata: type: object spec: - description: StorageNodeOpsSpec defines the desired state of a StorageNodeOps. + description: StorageNodeOpsSpec is one operation to perform against one StorageNode. properties: + abort: + description: |- + Abort asks a running operation to stop at its next step and unwind. It is + the only mutable field on this spec, because it is the only thing about an + operation that can legitimately be decided after it started. Whether an + abort is expressible from the current step is declared by that action's + graph rather than checked here. + type: boolean action: - description: Action is the operation to perform. Immutable. + description: Action is the operation to perform. enum: - Shutdown - Restart @@ -242,28 +250,30 @@ spec: - Resume - Remove - Migrate + - HostMaintenance type: string x-kubernetes-validations: - message: field is immutable rule: self == oldSelf force: - description: Force enables forced execution where the backend supports it. + description: |- + Force passes the control plane's force flag where the action supports it. + Migrate defaults it to true, because the control plane rejects a non-forced + restart of a node that is not already offline. type: boolean migrate: description: Migrate parameterizes action Migrate and is ignored by the others. properties: newSsdPcie: description: |- - NewSsdPcie lists additional NVMe PCIe addresses to bind on the target host - during a migration. Passed through to the control-plane restart as - new_ssd_pcie. + NewSsdPcie lists additional NVMe PCI addresses to bind on the target host, + passed through to the control-plane restart as new_ssd_pcie and merged into + the node's effective allow list so they survive a later rebuild. items: type: string type: array targetWorkerNode: - description: |- - TargetWorkerNode is the Kubernetes worker hostname the storage node is - relocated onto. + description: TargetWorkerNode is the Kubernetes worker the node is relocated onto. type: string x-kubernetes-validations: - message: field is immutable @@ -272,25 +282,28 @@ spec: - targetWorkerNode type: object nodeRef: - description: NodeRef is the name of the target StorageNode. Immutable. + description: |- + NodeRef names the StorageNode this operation acts on. The operation never + owns its target, because deleting the record of an operation must not delete + the node it operated on. type: string x-kubernetes-validations: - message: field is immutable rule: self == oldSelf reattachVolume: description: |- - ReattachVolume reattaches volumes during the node restart. - Applicable when action=Restart or action=Migrate. + ReattachVolume asks the control plane to reattach this node's volumes as + part of a restart. Applies to Restart, Migrate, and HostMaintenance. type: boolean remove: description: Remove parameterizes action Remove and is ignored by the others. properties: systemVolumeFilterRegex: + default: ^sb-fio-baseline-.* description: |- - SystemVolumeFilterRegex is a Go regular expression matched against backend - volume names. Matching volumes are treated as system volumes: excluded from - drain migration and deleted inline during the Verifying phase. - Defaults to `^sb-fio-baseline-.*`. + SystemVolumeFilterRegex matches backend volume names that are system + volumes: excluded from the drain's migration and deleted during + verification rather than blocking it. type: string type: object required: @@ -298,50 +311,79 @@ spec: - nodeRef type: object status: - description: StorageNodeOpsStatus holds the observed state of a StorageNodeOps. + description: StorageNodeOpsStatus is the observed state of one node operation. properties: completedAt: - description: CompletedAt is when the operation finished (successfully or not). + description: CompletedAt is when it reached a terminal phase. format: date-time type: string + drain: + description: |- + Drain is the drain's progress over the node's volumes, set only for action + Remove. + properties: + volumesMigrated: + description: VolumesMigrated is how many of them have completed. + format: int32 + minimum: 0 + type: integer + volumesTotal: + description: |- + VolumesTotal is the number of PV-managed volumes the drain has to move, + written once at the end of Validating and not modified afterward. + format: int32 + minimum: 0 + type: integer + required: + - volumesMigrated + - volumesTotal + type: object message: - description: Message is a human-readable description of the current state or failure reason. + description: |- + Message is the reason the phase is what it is: one sentence, replaced as the + operation moves, and never a log. type: string + observedGeneration: + description: |- + ObservedGeneration is the generation the rest of this status was computed + from, so a stale status can be told from a current one. + format: int64 + type: integer phase: - description: Phase is the high-level lifecycle phase. + description: Phase is the operation's own progress. enum: - Pending - Running - Succeeded - Failed + - Aborted type: string startedAt: - description: StartedAt is when the operation began. + description: StartedAt is when the operation acquired its target's lock. format: date-time type: string - subPhase: - description: SubPhase tracks the active drain step when action=Remove and phase=Running. - enum: - - Validating - - Suspending - - Migrating - - Verifying - - Removing - - Preparing - - Restarting - - Promoting - type: string - triggered: + step: description: |- - Triggered indicates the backend action POST has been sent (used during - Suspending to avoid duplicate POSTs across reconcile iterations). - type: boolean - volumesMigrated: - description: VolumesMigrated is the count of volumes successfully migrated (drain only). - type: integer - volumesPending: - description: VolumesPending is the count of volumes awaiting migration (drain only). - type: integer + Step is the position of the running action's state machine, as the shared + statemachine.KubeSnapshot. The rule is what an Enum marker would do if a + marker could reach a field of a shared type. + properties: + deadline: + description: |- + Deadline is when that state expires, absent when it has none. It is an + absolute instant, so a state whose deadline passed while the controller + was down restores as already expired. + format: date-time + type: string + state: + description: |- + State is the state the machine was in. Empty means the resource has not + been reconciled yet, and restores to the graph's initial state. + type: string + type: object + x-kubernetes-validations: + - message: unknown step + rule: '!has(self.state) || self.state in [''Requesting'',''Awaiting'',''Validating'',''Suspending'',''MigratingVolumes'',''Verifying'',''Removing'',''Preparing'',''Relocating'',''AwaitingNode'',''Promoting'',''Holding'',''ShuttingDown'',''Releasing'',''AwaitingHost'',''Restarting'',''Cleanup'']' type: object type: object served: true diff --git a/operator/config/crd/bases/storage.simplyblock.io_storagenodes.yaml b/operator/config/crd/bases/storage.simplyblock.io_storagenodes.yaml index b6a7f5f88..afcbe3faf 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_storagenodes.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_storagenodes.yaml @@ -12,339 +12,826 @@ spec: listKind: StorageNodeList plural: storagenodes shortNames: - - sn + - sn singular: storagenode scope: Namespaced versions: - - additionalPrinterColumns: - - jsonPath: .spec.workerNode - name: Worker - type: string - - jsonPath: .spec.socketId - name: Socket - type: string - - jsonPath: .spec.nodeIndex - name: NodeIdx - type: integer - - jsonPath: .status.failureDomain - name: FD - priority: 1 - type: integer - - jsonPath: .status.uuid - name: UUID - type: string - - jsonPath: .status.status - name: Status - type: string - - jsonPath: .status.health - name: Health - type: boolean - - jsonPath: .metadata.creationTimestamp - name: Age - type: date - name: v1alpha1 - schema: - openAPIV3Schema: - description: |- - StorageNode is the Schema for a single backend storage node instance. - One StorageNode CR exists per (workerNode, socketIndex) pair and is owned - by the parent StorageNodeSet. - properties: - apiVersion: - description: |- - APIVersion defines the versioned schema of this representation of an object. - Servers should convert recognized schemas to the latest internal value, and - may reject unrecognized values. - More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources - type: string - kind: - description: |- - Kind is a string value representing the REST resource this object represents. - Servers may infer this from the endpoint the client submits requests to. - Cannot be updated. - In CamelCase. - More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds - type: string - metadata: - type: object - spec: - description: StorageNodeSpec defines the desired state of a StorageNode. - properties: - nodeIndex: - description: NodeIndex is the per-socket node index (0..nodesPerSocket-1). - Immutable. - format: int32 - type: integer - x-kubernetes-validations: - - message: field is immutable - rule: self == oldSelf - overrides: - description: |- - Overrides holds per-node configuration propagated from - StorageNodeSet.spec.nodeConfigs[workerNode] on every reconcile. - properties: - deviceNames: - description: |- - DeviceNames explicitly defines the NVMe namespace names to use on this node - (e.g. ["nvme0n1","nvme1n1"]). - items: + - additionalPrinterColumns: + - jsonPath: .spec.workerNode + name: Worker + type: string + - jsonPath: .spec.socketId + name: Socket + type: string + - jsonPath: .spec.nodeIndex + name: NodeIdx + type: integer + - jsonPath: .status.failureDomain + name: FD + priority: 1 + type: integer + - jsonPath: .status.uuid + name: UUID + type: string + - jsonPath: .status.status + name: Status + type: string + - jsonPath: .status.health + name: Health + type: boolean + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha1 + schema: + openAPIV3Schema: + description: |- + StorageNode is the Schema for a single backend storage node instance. + One StorageNode CR exists per (workerNode, socketIndex) pair and is owned + by the parent StorageNodeSet. + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: StorageNodeSpec defines the desired state of a StorageNode. + properties: + nodeIndex: + description: NodeIndex is the per-socket node index (0..nodesPerSocket-1). Immutable. + format: int32 + type: integer + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + overrides: + description: |- + Overrides holds per-node configuration propagated from + StorageNodeSet.spec.nodeConfigs[workerNode] on every reconcile. + properties: + deviceNames: + description: |- + DeviceNames explicitly defines the NVMe namespace names to use on this node + (e.g. ["nvme0n1","nvme1n1"]). + items: + type: string + type: array + driveSizeRange: + description: DriveSizeRange overrides the drive size range filter for this node. + type: string + enableCpuTopology: + description: EnableCpuTopology overrides topology-aware CPU handling for this node. + type: boolean + expand: + description: |- + Expand marks this node as a cluster-expansion add. When true the backend + node-add endpoint receives expand=true, triggering rebalancing behaviour + appropriate for in-place cluster growth. Overrides StorageNodeSet.spec.expand. + type: boolean + failureDomain: + description: |- + FailureDomain is the failure-domain group index (≥ 0) for this node. + Required when the parent StorageCluster has enableFailureDomains=true. + Overrides StorageNodeSet.spec.nodeFailureDomains[workerNode] when both are set. + format: int32 + minimum: 0 + type: integer + journalManager: + description: JournalManagerSpec overrides journal manager tuning for this node. + properties: + count: + description: Count is the number of journal managers to configure. + format: int32 + type: integer + percentPerDevice: + description: PercentPerDevice is the journal manager capacity percentage per device. + format: int32 + type: integer + type: object + pcieAllowList: + description: PcieAllowList overrides the list of PCI addresses allowed for use on this node. + items: + type: string + type: array + pcieDenyList: + description: PcieDenyList overrides the list of PCI addresses excluded from use on this node. + items: + type: string + type: array + pcieModel: + description: PcieModel overrides the PCI model filter for this node. + type: string + reservedSystemCPU: + description: ReservedSystemCPU overrides the CPUs reserved for system workloads on this node. + type: string + skipKubeletConfiguration: + description: |- + SkipKubeletConfiguration overrides whether kubelet configuration changes are + skipped for this node. + type: boolean + spdkImage: + description: SpdkImage overrides the SPDK image for this node (e.g. for phased rollouts). + type: string + spdkProxyImage: + description: SpdkProxyImage overrides the SPDK proxy image for this node. type: string - type: array - driveSizeRange: - description: DriveSizeRange overrides the drive size range filter - for this node. - type: string - enableCpuTopology: - description: EnableCpuTopology overrides topology-aware CPU handling - for this node. - type: boolean - expand: - description: |- - Expand marks this node as a cluster-expansion add. When true the backend - node-add endpoint receives expand=true, triggering rebalancing behaviour - appropriate for in-place cluster growth. Overrides StorageNodeSet.spec.expand. - type: boolean - failureDomain: - description: |- - FailureDomain is the failure-domain group index (≥ 0) for this node. - Required when the parent StorageCluster has enableFailureDomains=true. - Overrides StorageNodeSet.spec.nodeFailureDomains[workerNode] when both are set. - format: int32 - minimum: 0 - type: integer - journalManager: - description: JournalManagerSpec overrides journal manager tuning - for this node. - properties: - count: - description: Count is the number of journal managers to configure. - format: int32 - type: integer - percentPerDevice: - description: PercentPerDevice is the journal manager capacity - percentage per device. - format: int32 - type: integer - type: object - pcieAllowList: - description: PcieAllowList overrides the list of PCI addresses - allowed for use on this node. - items: + spdkSystemMemory: + description: |- + SpdkSystemMemory overrides the SPDK huge-page memory allocation for this node + (e.g. "4G", "512M"). + pattern: ^[0-9]+(G|GI|GB|GiB|M|MI|MB|MiB|g|gi|gb|gib|m|mi|mb|mib)?$ type: string - type: array - pcieDenyList: - description: PcieDenyList overrides the list of PCI addresses - excluded from use on this node. - items: + ubuntuHost: + description: UbuntuHost overrides the Ubuntu host OS flag for this node. + type: boolean + type: object + socketId: + description: SocketID is the NUMA socket identifier from spec.socketsToUse (e.g. "0", "1"). Immutable. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + socketIndex: + description: |- + SocketIndex is the global ordinal (socketPosition × nodesPerSocket + nodeIndex). + Used internally by the operator to select the correct backend node from the + RPC-port-sorted list in pollUUIDFromBackend. Immutable. + format: int32 + type: integer + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + storageNodeSetRef: + description: StorageNodeSetRef is the name of the owning StorageNodeSet. Immutable. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + workerNode: + description: |- + WorkerNode is the Kubernetes node hostname this StorageNode runs on. + Users may not change it directly — it is re-pointed only by the operator + during a node migration (StorageNodeOps action=migrate). The + StorageNode validating webhook rejects user-driven changes to this field. + type: string + required: + - storageNodeSetRef + - workerNode + type: object + x-kubernetes-validations: + - message: field socketId is immutable once set + rule: '!has(oldSelf.socketId) || has(self.socketId)' + - message: field nodeIndex is immutable once set + rule: '!has(oldSelf.nodeIndex) || has(self.nodeIndex)' + - message: field socketIndex is immutable once set + rule: '!has(oldSelf.socketIndex) || has(self.socketIndex)' + status: + description: StorageNodeStatus holds the observed state of a StorageNode. + properties: + activeOpsRef: + description: |- + ActiveOpsRef is the name of the currently active StorageNodeOps CR targeting + this node. Empty when no operation is in progress. Used for mutual exclusion. + type: string + failureDomain: + description: |- + FailureDomain is the effective failure-domain group index for this node + as reported by the backend (≥ 0). Nil when the backend has not assigned one. + format: int32 + type: integer + health: + description: Health is the backend-reported node health flag. + type: boolean + hostname: + description: Hostname is the node hostname as reported by the backend. + type: string + latencyMetrics: + description: |- + LatencyMetrics holds the fio-measured baseline NVMe-oF latency for this node, + used by the volume rebalancer to make data-placement decisions. + properties: + baselineMeasuredAt: + description: BaselineMeasuredAt is when the baseline was established. + format: date-time type: string - type: array - pcieModel: - description: PcieModel overrides the PCI model filter for this - node. - type: string - reservedSystemCPU: - description: ReservedSystemCPU overrides the CPUs reserved for - system workloads on this node. - type: string - skipKubeletConfiguration: - description: |- - SkipKubeletConfiguration overrides whether kubelet configuration changes are - skipped for this node. - type: boolean - spdkImage: - description: SpdkImage overrides the SPDK image for this node - (e.g. for phased rollouts). - type: string - spdkProxyImage: - description: SpdkProxyImage overrides the SPDK proxy image for - this node. - type: string - spdkSystemMemory: - description: |- - SpdkSystemMemory overrides the SPDK huge-page memory allocation for this node - (e.g. "4G", "512M"). - pattern: ^[0-9]+(G|GI|GB|GiB|M|MI|MB|MiB|g|gi|gb|gib|m|mi|mb|mib)?$ - type: string - ubuntuHost: - description: UbuntuHost overrides the Ubuntu host OS flag for - this node. - type: boolean - type: object - socketId: - description: SocketID is the NUMA socket identifier from spec.socketsToUse - (e.g. "0", "1"). Immutable. - type: string - x-kubernetes-validations: - - message: field is immutable - rule: self == oldSelf - socketIndex: - description: |- - SocketIndex is the global ordinal (socketPosition × nodesPerSocket + nodeIndex). - Used internally by the operator to select the correct backend node from the - RPC-port-sorted list in pollUUIDFromBackend. Immutable. - format: int32 - type: integer - x-kubernetes-validations: - - message: field is immutable - rule: self == oldSelf - storageNodeSetRef: - description: StorageNodeSetRef is the name of the owning StorageNodeSet. - Immutable. - type: string - x-kubernetes-validations: - - message: field is immutable - rule: self == oldSelf - workerNode: - description: |- - WorkerNode is the Kubernetes node hostname this StorageNode runs on. - Users may not change it directly — it is re-pointed only by the operator - during a node migration (StorageNodeOps action=migrate). The - StorageNode validating webhook rejects user-driven changes to this field. - type: string - required: - - storageNodeSetRef - - workerNode - type: object - x-kubernetes-validations: - - message: field socketId is immutable once set - rule: '!has(oldSelf.socketId) || has(self.socketId)' - - message: field nodeIndex is immutable once set - rule: '!has(oldSelf.nodeIndex) || has(self.nodeIndex)' - - message: field socketIndex is immutable once set - rule: '!has(oldSelf.socketIndex) || has(self.socketIndex)' - status: - description: StorageNodeStatus holds the observed state of a StorageNode. - properties: - activeOpsRef: - description: |- - ActiveOpsRef is the name of the currently active StorageNodeOps CR targeting - this node. Empty when no operation is in progress. Used for mutual exclusion. - type: string - failureDomain: - description: |- - FailureDomain is the effective failure-domain group index for this node - as reported by the backend (≥ 0). Nil when the backend has not assigned one. - format: int32 - type: integer - health: - description: Health is the backend-reported node health flag. - type: boolean - hostname: - description: Hostname is the node hostname as reported by the backend. - type: string - latencyMetrics: - description: |- - LatencyMetrics holds the fio-measured baseline NVMe-oF latency for this node, - used by the volume rebalancer to make data-placement decisions. - properties: - baselineMeasuredAt: - description: BaselineMeasuredAt is when the baseline was established. - format: date-time - type: string - baselineP50NS: - description: BaselineP50NS is the p50 write latency (nanoseconds) - from the initial empty-cluster benchmark. - format: int64 - type: integer - baselineP99NS: - description: BaselineP99NS is the p99 write latency (nanoseconds) - from the initial empty-cluster benchmark. - format: int64 - type: integer - nodeUUID: - description: NodeUUID is the backend storage node UUID. - type: string - required: - - nodeUUID - type: object - ports: - description: Ports groups network connectivity fields (addresses and - ports). - properties: - lvol: - description: Lvol is the logical-volume subsystem port. - format: int32 - type: integer - management: - description: Management is the management IP address of the node. - type: string - nvmeof: - description: NvmeOf is the NVMe-oF fabric port. - format: int32 - type: integer - rpc: - description: Rpc is the RPC/management API port. - format: int32 - type: integer - type: object - postedAt: - description: |- - PostedAt is the timestamp when the node-add POST was sent. - Used as a provisioning guard against duplicate POSTs. - format: date-time - type: string - resources: - description: Resources groups compute and storage resource metrics. - properties: - capacity: - description: |- - Capacity is how much of the node's storage is in use, summed over its - devices. It is a measurement rather than a declaration, so it is absent - until something has measured it, and it lags reality by the interval at - which the control plane's metrics are scraped. - properties: - sampledAt: - description: |- - SampledAt is when the control plane took the reading. It is not when the - object was written, and it may be considerably older if metrics - collection has stopped. - format: date-time + baselineP50NS: + description: BaselineP50NS is the p50 write latency (nanoseconds) from the initial empty-cluster benchmark. + format: int64 + type: integer + baselineP99NS: + description: BaselineP99NS is the p99 write latency (nanoseconds) from the initial empty-cluster benchmark. + format: int64 + type: integer + nodeUUID: + description: NodeUUID is the backend storage node UUID. + type: string + required: + - nodeUUID + type: object + ports: + description: Ports groups network connectivity fields (addresses and ports). + properties: + lvol: + description: Lvol is the logical-volume subsystem port. + format: int32 + type: integer + management: + description: Management is the management IP address of the node. + type: string + nvmeof: + description: NvmeOf is the NVMe-oF fabric port. + format: int32 + type: integer + rpc: + description: Rpc is the RPC/management API port. + format: int32 + type: integer + type: object + postedAt: + description: |- + PostedAt is the timestamp when the node-add POST was sent. + Used as a provisioning guard against duplicate POSTs. + format: date-time + type: string + resources: + description: Resources groups compute and storage resource metrics. + properties: + capacity: + description: |- + Capacity is how much of the node's storage is in use, summed over its + devices. It is a measurement rather than a declaration, so it is absent + until something has measured it, and it lags reality by the interval at + which the control plane's metrics are scraped. + properties: + sampledAt: + description: |- + SampledAt is when the control plane took the reading. It is not when the + object was written, and it may be considerably older if metrics + collection has stopped. + format: date-time + type: string + totalBytes: + description: TotalBytes is the storage the node's devices provide. + format: int64 + minimum: 0 + type: integer + usedBytes: + description: UsedBytes is what they currently hold. + format: int64 + minimum: 0 + type: integer + type: object + cpu: + description: CPU is the number of SPDK CPU cores allocated to this node. + format: int32 + type: integer + devices: + description: Devices is the device summary (online/total) reported by the backend. + type: string + memory: + description: Memory is the SPDK memory allocation reported by the backend. + type: string + volumes: + description: Volumes is the current number of logical volumes on this node. + format: int32 + type: integer + type: object + status: + description: Status is the backend-reported node status (e.g. online, suspended, offline). + type: string + uptime: + description: Uptime is the node uptime as reported by the backend. + type: string + uuid: + description: UUID is the backend storage node UUID. Set once after node-add completes. + type: string + type: object + type: object + served: true + storage: false + subresources: + status: {} + - additionalPrinterColumns: + - jsonPath: .spec.clusterRef + name: Cluster + type: string + - jsonPath: .spec.workerNode + name: Worker + type: string + - jsonPath: .spec.socketId + name: Socket + type: string + - jsonPath: .spec.slot + name: Slot + type: integer + - jsonPath: .status.phase + name: Phase + type: string + - jsonPath: .status.step.state + name: Step + type: string + - jsonPath: .status.status + name: Status + type: string + - jsonPath: .status.health + name: Health + type: boolean + - jsonPath: .status.uuid + name: UUID + priority: 1 + type: string + - jsonPath: .status.failureDomain + name: FD + priority: 1 + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha2 + schema: + openAPIV3Schema: + description: |- + StorageNode is one backend storage node: one SPDK process bound to one NUMA + socket of one Kubernetes worker. One object exists per (workerNode, slot) pair, + owned by the StorageCluster it belongs to. + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: |- + StorageNodeSpec is the desired state of one backend storage node, meaning one + SPDK process bound to one NUMA socket of one Kubernetes worker. + properties: + clusterRef: + description: |- + ClusterRef names the StorageCluster this node belongs to. The cluster also + owns this object by controller reference, so deleting the cluster deletes + its nodes. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + config: + description: |- + Config is this node's complete configuration, copied from the + ClusterDeploymentConfig entry that produced it. It is a copy rather than a + projection, because that document is ephemeral: nothing reads it once the + node exists, deleting it changes nothing, and editing it reaches only nodes + created afterward. + properties: + deviceNames: + description: |- + DeviceNames names the devices to use. An entry is a PCI address + ("0000:5e:00.0") or a device path ("/dev/sdb,") which are the two classes + simplyblock accepts as backend storage, and a bare name ("nvme0n1") is read + as a path under /dev. One list carries both spellings, and every entry is of + the class its cluster declares in StorageCluster.spec.deviceClass: a list + mixing the two, or naming the class the cluster is not, is rejected by the + StorageNode validating webhook. Set explicitly, it overrides every filter + below. Immutable: it selects which physical devices the node owns. + items: + pattern: ^([0-9a-fA-F]{4}:[0-9a-fA-F]{2}:[0-9a-fA-F]{2}\.[0-9a-fA-F]|/dev/[a-zA-Z0-9._/-]+|[a-zA-Z0-9._-]+)$ type: string - totalBytes: - description: TotalBytes is the storage the node's devices - provide. - format: int64 - minimum: 0 - type: integer - usedBytes: - description: UsedBytes is what they currently hold. - format: int64 - minimum: 0 - type: integer - type: object - cpu: - description: CPU is the number of SPDK CPU cores allocated to - this node. - format: int32 - type: integer - devices: - description: Devices is the device summary (online/total) reported - by the backend. - type: string - memory: - description: Memory is the SPDK memory allocation reported by - the backend. - type: string - volumes: - description: Volumes is the current number of logical volumes - on this node. - format: int32 - type: integer - type: object - status: - description: Status is the backend-reported node status (e.g. online, - suspended, offline). - type: string - uptime: - description: Uptime is the node uptime as reported by the backend. - type: string - uuid: - description: UUID is the backend storage node UUID. Set once after - node-add completes. - type: string - type: object - type: object - served: true - storage: true - subresources: - status: {} + type: array + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + driveSizeRange: + description: DriveSizeRange filters devices by size. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + expand: + description: |- + Expand marks this node as an addition to an already-active cluster, which + the control plane reads as a request to rebalance onto it rather than to + treat it as part of an initial layout. Immutable once set: it describes how + the node joined rather than what it is. + type: boolean + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + failureDomain: + description: |- + FailureDomain is the label of the fault group this node belongs to + such as rack-b, naming the physical grouping it shares with its peers rather + than indexing it. Required when the cluster has enableFailureDomains set, + and provisioning is held with a FailureDomainMissing event until it is + present. Immutable once set, which is what makes it fillable later and then + frozen: chunk placement was computed from it. + + The value takes the shape of a Kubernetes label value, because that is what + it is seeded from where a cluster carries topology labels at all. + maxLength: 63 + pattern: ^[a-zA-Z0-9]([-_.a-zA-Z0-9]*[a-zA-Z0-9])?$ + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + journalManager: + description: |- + JournalManager tunes the journal manager count and per-device capacity + share for this node. Immutable: both are on-disk layout, fixed when the + devices were partitioned. + properties: + count: + description: Count is the number of journal managers to configure. + format: int32 + minimum: 1 + type: integer + percentPerDevice: + description: PercentPerDevice is the share of each device given to the journal. + format: int32 + maximum: 100 + minimum: 1 + type: integer + type: object + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + pcieAllowList: + description: |- + PcieAllowList selects devices by PCI address. It is the one device field a + migration writes, merging spec.migrate.newSsdPcie into it so devices added + on the target host survive a later rebuild, so it is guarded by the + StorageNode validating webhook rather than by a marker. This and the two + PCI filters below belong to an NVMe cluster: the webhook rejects them on a + cluster whose deviceClass is LogicalBlock, because a logical block device + has no PCI address to match. + items: + type: string + type: array + pcieDenyList: + description: PcieDenyList excludes devices by PCI address. + items: + type: string + type: array + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + pcieModel: + description: PcieModel filters devices by PCI model string. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + sizing: + description: |- + Sizing is what this node's huge pages and core layout were sized from. + Writable by the operator alone. + properties: + minHugePagesSize: + description: |- + MinHugePagesSize is the smallest huge-page allocation this node makes, as a + size string such as 100G or 1T, where a bare number is gigabytes. It is a + floor rather than a limit: the effective allocation is the larger of this + value and the minimum the node's device and subsystem count requires. + type: string + vcpuCount: + description: |- + VCPUCount is the number of vCPUs allocated to SPDK on this node, as an + explicit core count rather than a percentage. + format: int32 + minimum: 4 + type: integer + required: + - vcpuCount + type: object + spdkImage: + description: |- + SpdkImage overrides the SPDK image the control plane starts for this node, + which is what makes a phased image rollout expressible per node. + type: string + spdkProxyImage: + description: SpdkProxyImage overrides the SPDK proxy image for this node. + type: string + spdkSystemMemory: + description: |- + SpdkSystemMemory is the memory the control plane starts this node's SPDK + with, as a size string such as 4G or 512M. Mutable: a node whose device + count grew legitimately needs to raise it. + pattern: ^[0-9]+(G|GI|GB|GiB|M|MI|MB|MiB|g|gi|gb|gib|m|mi|mb|mib)?$ + type: string + required: + - sizing + type: object + x-kubernetes-validations: + - message: field journalManager is immutable once set + rule: '!has(oldSelf.journalManager) || has(self.journalManager)' + - message: field deviceNames is immutable once set + rule: '!has(oldSelf.deviceNames) || has(self.deviceNames)' + - message: field pcieDenyList is immutable once set + rule: '!has(oldSelf.pcieDenyList) || has(self.pcieDenyList)' + - message: field pcieModel is immutable once set + rule: '!has(oldSelf.pcieModel) || has(self.pcieModel)' + - message: field driveSizeRange is immutable once set + rule: '!has(oldSelf.driveSizeRange) || has(self.driveSizeRange)' + - message: field failureDomain is immutable once set + rule: '!has(oldSelf.failureDomain) || has(self.failureDomain)' + - message: field expand is immutable once set + rule: '!has(oldSelf.expand) || has(self.expand)' + nodeIndex: + description: |- + NodeIndex is the position among the nodes sharing this socket, in + 0..nodesPerSocket-1. See SocketID. + format: int32 + minimum: 0 + type: integer + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + nodeSet: + description: |- + NodeSet is the name of the group in ClusterDeploymentConfig.nodeSets[] this + node was declared under. It is a label rather than a reference: nothing is + fetched by it, and it exists so that a node can be traced back to the + document that produced it. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + slot: + description: |- + Slot is which storage-node slot on this worker the object occupies, counted + from zero. A worker runs one node per socket per nodesPerSocket, and the + slot is the position among them. It is the identity the operator keys on: + the topology label the CSI driver reads is + storage.simplyblock.io/storage-node-uuid.., and the slot + outlives the node filling it, because only the UUID behind it changes when a + node is replaced or relocated. + format: int32 + minimum: 0 + type: integer + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + socketId: + description: |- + SocketID is the NUMA socket this node is bound to, as declared in the node + set's socket list, so 0 or 1. With NodeIndex it decomposes Slot into the + pair a person reads; nothing but a print column consumes either. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + workerNode: + description: |- + WorkerNode is the Kubernetes worker hostname this node runs on. It is not + marked immutable, because a migration re-points it, but the StorageNode + validating webhook rejects any change made by an identity outside the + operator's namespace. + type: string + required: + - clusterRef + - config + - workerNode + type: object + x-kubernetes-validations: + - message: field nodeSet is immutable once set + rule: '!has(oldSelf.nodeSet) || has(self.nodeSet)' + - message: field socketId is immutable once set + rule: '!has(oldSelf.socketId) || has(self.socketId)' + - message: field nodeIndex is immutable once set + rule: '!has(oldSelf.nodeIndex) || has(self.nodeIndex)' + - message: field slot is immutable once set + rule: '!has(oldSelf.slot) || has(self.slot)' + status: + description: StorageNodeStatus is the observed state of one storage node. + properties: + activeOpsRef: + description: |- + ActiveOpsRef names the StorageNodeOps currently allowed to touch this node. + Empty when none is running. + type: string + failureDomain: + description: |- + FailureDomain is the failure-domain label the control plane actually + assigned, which is not necessarily the one spec.config.failureDomain + requested. + type: string + health: + description: Health is the health flag the control plane reports. + type: boolean + hostname: + description: Hostname is the node hostname as the control plane reports it. + type: string + latencyMetrics: + description: |- + LatencyMetrics holds the fio-measured NVMe-oF baseline the volume + rebalancer reads. + properties: + baselineMeasuredAt: + description: BaselineMeasuredAt is when the baseline was established. + format: date-time + type: string + baselineP50NS: + description: |- + BaselineP50NS is the p50 write latency, in nanoseconds, of the initial + empty-cluster benchmark. + format: int64 + minimum: 0 + type: integer + baselineP99NS: + description: |- + BaselineP99NS is the p99 write latency, in nanoseconds, of the same + benchmark. + format: int64 + minimum: 0 + type: integer + nodeUUID: + description: |- + NodeUUID is the backend storage node the reading was taken against. It is + carried beside the reading rather than inferred from status.uuid, because a + baseline measured against one backend node stops describing the slot once a + replacement fills it. + type: string + required: + - nodeUUID + type: object + message: + description: |- + Message is the reason the phase is what it is: one sentence, replaced as the + node moves, and never a log. + type: string + observedGeneration: + description: |- + ObservedGeneration is the generation the rest of this status was computed + from, so a stale status can be told from a current one. + format: int64 + type: integer + phase: + description: |- + Phase is the operator's own view of this node, and the field its + provisioning branches on. + enum: + - Pending + - Provisioning + - Online + - Removing + - Offline + - Degraded + - Failed + type: string + ports: + description: Ports groups the reported addresses and ports. + properties: + lvol: + description: Lvol is the logical-volume subsystem port. + format: int32 + type: integer + management: + description: Management is the management IP address of the node. + type: string + nvmeof: + description: The NVMe-oF fabric port. + format: int32 + type: integer + rpc: + description: Rpc is the RPC and management API port. + format: int32 + type: integer + type: object + resources: + description: Resources groups the reported compute and storage figures. + properties: + capacity: + description: |- + Capacity is how much of the node's storage is in use, summed over its + devices. It is a measurement rather than a declaration, so it is absent + until something has measured it, and it lags reality by the interval at + which the control plane's metrics are scraped. + properties: + sampledAt: + description: |- + SampledAt is when the control plane took the reading. It is not when the + object was written, and it may be considerably older if metrics collection + has stopped. + format: date-time + type: string + totalBytes: + description: TotalBytes is the storage the node's devices provide. + format: int64 + minimum: 0 + type: integer + usedBytes: + description: UsedBytes is what they currently hold. + format: int64 + minimum: 0 + type: integer + type: object + cpu: + description: CPU is the number of SPDK cores allocated to this node. + format: int32 + type: integer + devices: + description: |- + Devices summarizes the node's NVMe devices. Absent until the control plane + has reported, which is what tells a node that has not reported from one that + genuinely has no devices. + properties: + online: + description: |- + Online is how many of the node's devices the control plane reports as + usable. + format: int32 + minimum: 0 + type: integer + total: + description: Total is how many devices the node has. + format: int32 + minimum: 0 + type: integer + required: + - online + - total + type: object + memory: + description: Memory is the SPDK memory allocation the control plane reports. + type: string + volumes: + description: Volumes is the current number of logical volumes on this node. + format: int32 + type: integer + type: object + status: + description: |- + Status is the lifecycle the control plane reports: online, suspended, + offline, in_creation, in_restart, in_shutdown, unreachable, or timeout. The + values are the control plane's, which is why they are neither PascalCase nor + constrained by an Enum here. + type: string + step: + description: |- + Step is the position of the provisioning machine, as the shared + statemachine.KubeSnapshot. The rule is what an Enum marker would do if a + marker could reach a field of a shared type. + properties: + deadline: + description: |- + Deadline is when that state expires, absent when it has none. It is an + absolute instant, so a state whose deadline passed while the controller + was down restores as already expired. + format: date-time + type: string + state: + description: |- + State is the state the machine was in. Empty means the resource has not + been reconciled yet, and restores to the graph's initial state. + type: string + type: object + x-kubernetes-validations: + - message: unknown step + rule: '!has(self.state) || self.state in [''CheckingHost'',''CheckingConfig'',''AwaitingSlot'',''Posting'',''Resolving'',''Adopting'']' + uptime: + description: Uptime is the node uptime as the control plane reports it. + type: string + uuid: + description: |- + UUID is the backend node UUID. Empty means the node has neither been + provisioned nor adopted, and non-empty means steady state. + type: string + type: object + type: object + served: true + storage: true + subresources: + status: {} + conversion: + strategy: Webhook + webhook: + conversionReviewVersions: + - v1 + clientConfig: + service: + namespace: simplyblock-operator-system + name: simplyblock-operator-conversion-webhook-service + path: /convert diff --git a/operator/config/crd/converted-kinds.txt b/operator/config/crd/converted-kinds.txt index caf416cd6..72ca64ca8 100644 --- a/operator/config/crd/converted-kinds.txt +++ b/operator/config/crd/converted-kinds.txt @@ -19,4 +19,5 @@ storagebackups.storage.simplyblock.io storageclusterops.storage.simplyblock.io storageclusters.storage.simplyblock.io storagenodeops.storage.simplyblock.io +storagenodes.storage.simplyblock.io storagepools.storage.simplyblock.io diff --git a/operator/config/rbac/metrics_apiserver_aggregate_view_role.yaml b/operator/config/rbac/metrics_apiserver_aggregate_view_role.yaml index 862282d33..da5b250bc 100644 --- a/operator/config/rbac/metrics_apiserver_aggregate_view_role.yaml +++ b/operator/config/rbac/metrics_apiserver_aggregate_view_role.yaml @@ -35,6 +35,7 @@ rules: - storagedevicemetrics - storagepoolmetrics - storageclustermetrics + - storagenodemetrics verbs: - get - list diff --git a/operator/config/rbac/role.yaml b/operator/config/rbac/role.yaml index d03e263e7..07f03f8ac 100644 --- a/operator/config/rbac/role.yaml +++ b/operator/config/rbac/role.yaml @@ -19,16 +19,6 @@ rules: - patch - update - watch -- apiGroups: - - "" - resources: - - events - verbs: - - create - - get - - list - - patch - - watch - apiGroups: - "" resources: @@ -74,9 +64,15 @@ rules: - delete - get - list - - patch - - update - watch +- apiGroups: + - "" + - events.k8s.io + resources: + - events + verbs: + - create + - patch - apiGroups: - admissionregistration.k8s.io resources: @@ -103,6 +99,7 @@ rules: - storageclusterops.storage.simplyblock.io - storageclusters.storage.simplyblock.io - storagenodeops.storage.simplyblock.io + - storagenodes.storage.simplyblock.io - storagepools.storage.simplyblock.io resources: - customresourcedefinitions @@ -173,13 +170,6 @@ rules: - patch - update - watch -- apiGroups: - - events.k8s.io - resources: - - events - verbs: - - create - - patch - apiGroups: - policy resources: @@ -262,7 +252,6 @@ rules: - storagedevices - storagenodeops - storagenodes - - storagenodesets - storagepoolops - storagepools - tasks @@ -293,7 +282,6 @@ rules: - storageclusters/finalizers - storagenodeops/finalizers - storagenodes/finalizers - - storagenodesets/finalizers - storagepoolops/finalizers - storagepools/finalizers - tasks/finalizers @@ -342,3 +330,11 @@ rules: - patch - update - watch +- apiGroups: + - storage.simplyblock.io + resources: + - storagenodesets + verbs: + - get + - list + - watch diff --git a/operator/config/webhook/manifests.yaml b/operator/config/webhook/manifests.yaml index 5c27a6dda..115ca5a8e 100644 --- a/operator/config/webhook/manifests.yaml +++ b/operator/config/webhook/manifests.yaml @@ -190,15 +190,17 @@ webhooks: service: name: webhook-service namespace: system - path: /validate-storage-simplyblock-io-v1alpha1-storagenode + path: /validate-storage-simplyblock-io-v1alpha2-storagenode failurePolicy: Fail + matchPolicy: Equivalent name: vstoragenode.simplyblock.io rules: - apiGroups: - storage.simplyblock.io apiVersions: - - v1alpha1 + - v1alpha2 operations: + - CREATE - UPDATE resources: - storagenodes diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index 5f2031f15..0f8a6c159 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -932,11 +932,14 @@ spec: description: Count is the number of journal managers to configure. format: int32 + minimum: 1 type: integer percentPerDevice: - description: PercentPerDevice is the journal manager - capacity percentage per device. + description: PercentPerDevice is the share of each + device given to the journal. format: int32 + maximum: 100 + minimum: 1 type: integer type: object mgmtInterface: @@ -4949,6 +4952,286 @@ spec: x-kubernetes-validations: - message: field is immutable rule: self == oldSelf + storageNodes: + description: |- + StorageNodes is the Kubernetes workload the cluster's storage nodes run + as, and the cluster owns every object in it by controller reference: a + cluster deleted takes its DaemonSet, Services, certificate, and per-node + ConfigMap with it. One workload serves the whole cluster, because growth is + nodes rather than sets and what differs between hardware generations is per + node already. + properties: + containerResources: + description: |- + ContainerResources sets requests and limits for the storage-node container. + Unset enforces no limits. + properties: + claims: + description: |- + Claims lists the names of resources, defined in spec.resourceClaims, + that are used by this container. + + This field depends on the + DynamicResourceAllocation feature gate. + + This field is immutable. It can only be set for containers. + items: + description: ResourceClaim references one entry in PodSpec.ResourceClaims. + properties: + name: + description: |- + Name must match the name of one entry in pod.spec.resourceClaims of + the Pod where this field is used. It makes that resource available + inside a container. + type: string + request: + description: |- + Request is the name chosen for a request in the referenced claim. + If empty, everything from the claim is made available, otherwise + only the result of this request. + type: string + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + limits: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Limits describes the maximum amount of compute resources allowed. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + requests: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Requests describes the minimum amount of compute resources required. + If Requests is omitted for a container, it defaults to Limits if that is explicitly specified, + otherwise to an implementation-defined value. Requests cannot exceed Limits. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + type: object + dataInterfaces: + description: DataInterfaces are the data-plane network interfaces. + items: + type: string + type: array + enableCpuTopology: + description: EnableCpuTopology turns on topology-aware CPU assignment. + type: boolean + enableFormat4K: + description: |- + EnableFormat4K formats NVMe devices to a 4K block size where the device + supports it. + type: boolean + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + enableJournalDevice: + description: |- + EnableJournalDevice dedicates the smallest NVMe device on each node to the + journal manager, instead of carving a journal partition out of every + device. + type: boolean + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + enableKubeletConfiguration: + description: |- + EnableKubeletConfiguration lets the storage node apply the kubelet + configuration changes it needs. Off by default, which is the behavior the + retired skipKubeletConfiguration expressed by being set. + type: boolean + image: + description: |- + Image is the storage-node container image. Defaults to the ControlPlane + singleton's spec.image when unset, so a deployment states the version once. + pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ + type: string + imagePullPolicy: + default: IfNotPresent + description: ImagePullPolicy controls when that image is pulled. + enum: + - Always + - Never + - IfNotPresent + type: string + initContainerResources: + description: InitContainerResources does the same for the init + container. + properties: + claims: + description: |- + Claims lists the names of resources, defined in spec.resourceClaims, + that are used by this container. + + This field depends on the + DynamicResourceAllocation feature gate. + + This field is immutable. It can only be set for containers. + items: + description: ResourceClaim references one entry in PodSpec.ResourceClaims. + properties: + name: + description: |- + Name must match the name of one entry in pod.spec.resourceClaims of + the Pod where this field is used. It makes that resource available + inside a container. + type: string + request: + description: |- + Request is the name chosen for a request in the referenced claim. + If empty, everything from the claim is made available, otherwise + only the result of this request. + type: string + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + limits: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Limits describes the maximum amount of compute resources allowed. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + requests: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Requests describes the minimum amount of compute resources required. + If Requests is omitted for a container, it defaults to Limits if that is explicitly specified, + otherwise to an implementation-defined value. Requests cannot exceed Limits. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + type: object + maxParallelNodeAdds: + default: 1 + description: |- + MaxParallelNodeAdds limits how many workers may be in the node-add process + at once, counted by distinct worker rather than by object so that a + two-socket host consumes one slot. Workers hosting a FoundationDB pod are + always sequential regardless of this value, because a node add reboots the + host and two simultaneous FoundationDB reboots reduce the control plane's + own fault tolerance. + format: int32 + minimum: 1 + type: integer + mgmtInterface: + description: MgmtInterface is the management network interface + storage nodes bind. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + nodesPerSocket: + default: 1 + description: NodesPerSocket is how many storage nodes run per + NUMA socket. + format: int32 + minimum: 1 + type: integer + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + openShiftCluster: + description: OpenShiftCluster states that the Kubernetes distribution + is OpenShift. + type: boolean + openShiftMachineConfigPool: + default: worker + description: |- + OpenShiftMachineConfigPool names the pool generated MachineConfig objects + are labeled into. + type: string + reservedSystemCPU: + description: ReservedSystemCPU is the CPU set held back from SPDK + for system workloads. + type: string + socketsToUse: + description: |- + SocketsToUse restricts deployment to selected NUMA sockets. Empty means + socket 0 alone. + items: + type: string + type: array + tolerations: + description: Tolerations are applied to the storage-node pods. + items: + description: |- + The pod this Toleration is attached to tolerates any taint that matches + the triple using the matching operator . + properties: + effect: + description: |- + Effect indicates the taint effect to match. Empty means match all taint effects. + When specified, allowed values are NoSchedule, PreferNoSchedule and NoExecute. + type: string + key: + description: |- + Key is the taint key that the toleration applies to. Empty means match all taint keys. + If the key is empty, operator must be Exists; this combination means to match all values and all keys. + type: string + operator: + description: |- + Operator represents a key's relationship to the value. + Valid operators are Exists, Equal, Lt, and Gt. Defaults to Equal. + Exists is equivalent to wildcard for value, so that a pod can + tolerate all taints of a particular category. + Lt and Gt perform numeric comparisons (requires feature gate TaintTolerationComparisonOperators). + type: string + tolerationSeconds: + description: |- + TolerationSeconds represents the period of time the toleration (which must be + of effect NoExecute, otherwise this field is ignored) tolerates the taint. By default, + it is not set, which means tolerate the taint forever (do not evict). Zero and + negative values will be treated as 0 (evict immediately) by the system. + format: int64 + type: integer + value: + description: |- + Value is the taint value the toleration matches to. + If the operator is Exists, the value should be empty, otherwise just a regular string. + type: string + type: object + type: array + ubuntuHost: + description: |- + UbuntuHost states that the worker's host OS is Ubuntu, which changes how + the node configures huge pages and the kernel modules it loads. + type: boolean + type: object + x-kubernetes-validations: + - message: field mgmtInterface is immutable once set + rule: '!has(oldSelf.mgmtInterface) || has(self.mgmtInterface)' + - message: field nodesPerSocket is immutable once set + rule: '!has(oldSelf.nodesPerSocket) || has(self.nodesPerSocket)' + - message: field enableJournalDevice is immutable once set + rule: '!has(oldSelf.enableJournalDevice) || has(self.enableJournalDevice)' + - message: field enableFormat4K is immutable once set + rule: '!has(oldSelf.enableFormat4K) || has(self.enableFormat4K)' stripe: description: |- Stripe is the erasure-coding layout every volume in the cluster is @@ -5824,11 +6107,12 @@ spec: - jsonPath: .status.phase name: Phase type: string - - jsonPath: .status.subPhase - name: SubPhase + - jsonPath: .status.step.state + name: Step type: string - jsonPath: .status.message name: Message + priority: 1 type: string - jsonPath: .metadata.creationTimestamp name: Age @@ -5837,10 +6121,9 @@ spec: schema: openAPIV3Schema: description: |- - StorageNodeOps is a one-shot operational CR targeting a single StorageNode. - Analogous to a Kubernetes Job — it drives an action (Shutdown, Restart, Suspend, - Resume, Remove, Migrate) to completion and records the result. Only one - StorageNodeOps can be active per StorageNode at a time. + StorageNodeOps is a single operation performed against one StorageNode. It runs + to a terminal phase and stays afterward as the audit record of what was done, to + which node, with which parameters, and how it ended. properties: apiVersion: description: |- @@ -5860,10 +6143,19 @@ spec: metadata: type: object spec: - description: StorageNodeOpsSpec defines the desired state of a StorageNodeOps. + description: StorageNodeOpsSpec is one operation to perform against one + StorageNode. properties: + abort: + description: |- + Abort asks a running operation to stop at its next step and unwind. It is + the only mutable field on this spec, because it is the only thing about an + operation that can legitimately be decided after it started. Whether an + abort is expressible from the current step is declared by that action's + graph rather than checked here. + type: boolean action: - description: Action is the operation to perform. Immutable. + description: Action is the operation to perform. enum: - Shutdown - Restart @@ -5871,13 +6163,16 @@ spec: - Resume - Remove - Migrate + - HostMaintenance type: string x-kubernetes-validations: - message: field is immutable rule: self == oldSelf force: - description: Force enables forced execution where the backend supports - it. + description: |- + Force passes the control plane's force flag where the action supports it. + Migrate defaults it to true, because the control plane rejects a non-forced + restart of a node that is not already offline. type: boolean migrate: description: Migrate parameterizes action Migrate and is ignored by @@ -5885,16 +6180,15 @@ spec: properties: newSsdPcie: description: |- - NewSsdPcie lists additional NVMe PCIe addresses to bind on the target host - during a migration. Passed through to the control-plane restart as - new_ssd_pcie. + NewSsdPcie lists additional NVMe PCI addresses to bind on the target host, + passed through to the control-plane restart as new_ssd_pcie and merged into + the node's effective allow list so they survive a later rebuild. items: type: string type: array targetWorkerNode: - description: |- - TargetWorkerNode is the Kubernetes worker hostname the storage node is - relocated onto. + description: TargetWorkerNode is the Kubernetes worker the node + is relocated onto. type: string x-kubernetes-validations: - message: field is immutable @@ -5903,26 +6197,29 @@ spec: - targetWorkerNode type: object nodeRef: - description: NodeRef is the name of the target StorageNode. Immutable. + description: |- + NodeRef names the StorageNode this operation acts on. The operation never + owns its target, because deleting the record of an operation must not delete + the node it operated on. type: string x-kubernetes-validations: - message: field is immutable rule: self == oldSelf reattachVolume: description: |- - ReattachVolume reattaches volumes during the node restart. - Applicable when action=Restart or action=Migrate. + ReattachVolume asks the control plane to reattach this node's volumes as + part of a restart. Applies to Restart, Migrate, and HostMaintenance. type: boolean remove: description: Remove parameterizes action Remove and is ignored by the others. properties: systemVolumeFilterRegex: + default: ^sb-fio-baseline-.* description: |- - SystemVolumeFilterRegex is a Go regular expression matched against backend - volume names. Matching volumes are treated as system volumes: excluded from - drain migration and deleted inline during the Verifying phase. - Defaults to `^sb-fio-baseline-.*`. + SystemVolumeFilterRegex matches backend volume names that are system + volumes: excluded from the drain's migration and deleted during + verification rather than blocking it. type: string type: object required: @@ -5930,55 +6227,80 @@ spec: - nodeRef type: object status: - description: StorageNodeOpsStatus holds the observed state of a StorageNodeOps. + description: StorageNodeOpsStatus is the observed state of one node operation. properties: completedAt: - description: CompletedAt is when the operation finished (successfully - or not). + description: CompletedAt is when it reached a terminal phase. format: date-time type: string + drain: + description: |- + Drain is the drain's progress over the node's volumes, set only for action + Remove. + properties: + volumesMigrated: + description: VolumesMigrated is how many of them have completed. + format: int32 + minimum: 0 + type: integer + volumesTotal: + description: |- + VolumesTotal is the number of PV-managed volumes the drain has to move, + written once at the end of Validating and not modified afterward. + format: int32 + minimum: 0 + type: integer + required: + - volumesMigrated + - volumesTotal + type: object message: - description: Message is a human-readable description of the current - state or failure reason. + description: |- + Message is the reason the phase is what it is: one sentence, replaced as the + operation moves, and never a log. type: string + observedGeneration: + description: |- + ObservedGeneration is the generation the rest of this status was computed + from, so a stale status can be told from a current one. + format: int64 + type: integer phase: - description: Phase is the high-level lifecycle phase. + description: Phase is the operation's own progress. enum: - Pending - Running - Succeeded - Failed + - Aborted type: string startedAt: - description: StartedAt is when the operation began. + description: StartedAt is when the operation acquired its target's + lock. format: date-time type: string - subPhase: - description: SubPhase tracks the active drain step when action=Remove - and phase=Running. - enum: - - Validating - - Suspending - - Migrating - - Verifying - - Removing - - Preparing - - Restarting - - Promoting - type: string - triggered: + step: description: |- - Triggered indicates the backend action POST has been sent (used during - Suspending to avoid duplicate POSTs across reconcile iterations). - type: boolean - volumesMigrated: - description: VolumesMigrated is the count of volumes successfully - migrated (drain only). - type: integer - volumesPending: - description: VolumesPending is the count of volumes awaiting migration - (drain only). - type: integer + Step is the position of the running action's state machine, as the shared + statemachine.KubeSnapshot. The rule is what an Enum marker would do if a + marker could reach a field of a shared type. + properties: + deadline: + description: |- + Deadline is when that state expires, absent when it has none. It is an + absolute instant, so a state whose deadline passed while the controller + was down restores as already expired. + format: date-time + type: string + state: + description: |- + State is the state the machine was in. Empty means the resource has not + been reconciled yet, and restores to the graph's initial state. + type: string + type: object + x-kubernetes-validations: + - message: unknown step + rule: '!has(self.state) || self.state in [''Requesting'',''Awaiting'',''Validating'',''Suspending'',''MigratingVolumes'',''Verifying'',''Removing'',''Preparing'',''Relocating'',''AwaitingNode'',''Promoting'',''Holding'',''ShuttingDown'',''Releasing'',''AwaitingHost'',''Restarting'',''Cleanup'']' type: object type: object served: true @@ -5993,6 +6315,16 @@ metadata: controller-gen.kubebuilder.io/version: v0.21.0 name: storagenodes.storage.simplyblock.io spec: + conversion: + strategy: Webhook + webhook: + clientConfig: + service: + name: simplyblock-operator-conversion-webhook-service + namespace: simplyblock-operator-system + path: /convert + conversionReviewVersions: + - v1 group: storage.simplyblock.io names: kind: StorageNode @@ -6065,82 +6397,483 @@ spec: x-kubernetes-validations: - message: field is immutable rule: self == oldSelf - overrides: + overrides: + description: |- + Overrides holds per-node configuration propagated from + StorageNodeSet.spec.nodeConfigs[workerNode] on every reconcile. + properties: + deviceNames: + description: |- + DeviceNames explicitly defines the NVMe namespace names to use on this node + (e.g. ["nvme0n1","nvme1n1"]). + items: + type: string + type: array + driveSizeRange: + description: DriveSizeRange overrides the drive size range filter + for this node. + type: string + enableCpuTopology: + description: EnableCpuTopology overrides topology-aware CPU handling + for this node. + type: boolean + expand: + description: |- + Expand marks this node as a cluster-expansion add. When true the backend + node-add endpoint receives expand=true, triggering rebalancing behaviour + appropriate for in-place cluster growth. Overrides StorageNodeSet.spec.expand. + type: boolean + failureDomain: + description: |- + FailureDomain is the failure-domain group index (≥ 0) for this node. + Required when the parent StorageCluster has enableFailureDomains=true. + Overrides StorageNodeSet.spec.nodeFailureDomains[workerNode] when both are set. + format: int32 + minimum: 0 + type: integer + journalManager: + description: JournalManagerSpec overrides journal manager tuning + for this node. + properties: + count: + description: Count is the number of journal managers to configure. + format: int32 + type: integer + percentPerDevice: + description: PercentPerDevice is the journal manager capacity + percentage per device. + format: int32 + type: integer + type: object + pcieAllowList: + description: PcieAllowList overrides the list of PCI addresses + allowed for use on this node. + items: + type: string + type: array + pcieDenyList: + description: PcieDenyList overrides the list of PCI addresses + excluded from use on this node. + items: + type: string + type: array + pcieModel: + description: PcieModel overrides the PCI model filter for this + node. + type: string + reservedSystemCPU: + description: ReservedSystemCPU overrides the CPUs reserved for + system workloads on this node. + type: string + skipKubeletConfiguration: + description: |- + SkipKubeletConfiguration overrides whether kubelet configuration changes are + skipped for this node. + type: boolean + spdkImage: + description: SpdkImage overrides the SPDK image for this node + (e.g. for phased rollouts). + type: string + spdkProxyImage: + description: SpdkProxyImage overrides the SPDK proxy image for + this node. + type: string + spdkSystemMemory: + description: |- + SpdkSystemMemory overrides the SPDK huge-page memory allocation for this node + (e.g. "4G", "512M"). + pattern: ^[0-9]+(G|GI|GB|GiB|M|MI|MB|MiB|g|gi|gb|gib|m|mi|mb|mib)?$ + type: string + ubuntuHost: + description: UbuntuHost overrides the Ubuntu host OS flag for + this node. + type: boolean + type: object + socketId: + description: SocketID is the NUMA socket identifier from spec.socketsToUse + (e.g. "0", "1"). Immutable. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + socketIndex: + description: |- + SocketIndex is the global ordinal (socketPosition × nodesPerSocket + nodeIndex). + Used internally by the operator to select the correct backend node from the + RPC-port-sorted list in pollUUIDFromBackend. Immutable. + format: int32 + type: integer + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + storageNodeSetRef: + description: StorageNodeSetRef is the name of the owning StorageNodeSet. + Immutable. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + workerNode: + description: |- + WorkerNode is the Kubernetes node hostname this StorageNode runs on. + Users may not change it directly — it is re-pointed only by the operator + during a node migration (StorageNodeOps action=migrate). The + StorageNode validating webhook rejects user-driven changes to this field. + type: string + required: + - storageNodeSetRef + - workerNode + type: object + x-kubernetes-validations: + - message: field socketId is immutable once set + rule: '!has(oldSelf.socketId) || has(self.socketId)' + - message: field nodeIndex is immutable once set + rule: '!has(oldSelf.nodeIndex) || has(self.nodeIndex)' + - message: field socketIndex is immutable once set + rule: '!has(oldSelf.socketIndex) || has(self.socketIndex)' + status: + description: StorageNodeStatus holds the observed state of a StorageNode. + properties: + activeOpsRef: + description: |- + ActiveOpsRef is the name of the currently active StorageNodeOps CR targeting + this node. Empty when no operation is in progress. Used for mutual exclusion. + type: string + failureDomain: + description: |- + FailureDomain is the effective failure-domain group index for this node + as reported by the backend (≥ 0). Nil when the backend has not assigned one. + format: int32 + type: integer + health: + description: Health is the backend-reported node health flag. + type: boolean + hostname: + description: Hostname is the node hostname as reported by the backend. + type: string + latencyMetrics: + description: |- + LatencyMetrics holds the fio-measured baseline NVMe-oF latency for this node, + used by the volume rebalancer to make data-placement decisions. + properties: + baselineMeasuredAt: + description: BaselineMeasuredAt is when the baseline was established. + format: date-time + type: string + baselineP50NS: + description: BaselineP50NS is the p50 write latency (nanoseconds) + from the initial empty-cluster benchmark. + format: int64 + type: integer + baselineP99NS: + description: BaselineP99NS is the p99 write latency (nanoseconds) + from the initial empty-cluster benchmark. + format: int64 + type: integer + nodeUUID: + description: NodeUUID is the backend storage node UUID. + type: string + required: + - nodeUUID + type: object + ports: + description: Ports groups network connectivity fields (addresses and + ports). + properties: + lvol: + description: Lvol is the logical-volume subsystem port. + format: int32 + type: integer + management: + description: Management is the management IP address of the node. + type: string + nvmeof: + description: NvmeOf is the NVMe-oF fabric port. + format: int32 + type: integer + rpc: + description: Rpc is the RPC/management API port. + format: int32 + type: integer + type: object + postedAt: + description: |- + PostedAt is the timestamp when the node-add POST was sent. + Used as a provisioning guard against duplicate POSTs. + format: date-time + type: string + resources: + description: Resources groups compute and storage resource metrics. + properties: + capacity: + description: |- + Capacity is how much of the node's storage is in use, summed over its + devices. It is a measurement rather than a declaration, so it is absent + until something has measured it, and it lags reality by the interval at + which the control plane's metrics are scraped. + properties: + sampledAt: + description: |- + SampledAt is when the control plane took the reading. It is not when the + object was written, and it may be considerably older if metrics + collection has stopped. + format: date-time + type: string + totalBytes: + description: TotalBytes is the storage the node's devices + provide. + format: int64 + minimum: 0 + type: integer + usedBytes: + description: UsedBytes is what they currently hold. + format: int64 + minimum: 0 + type: integer + type: object + cpu: + description: CPU is the number of SPDK CPU cores allocated to + this node. + format: int32 + type: integer + devices: + description: Devices is the device summary (online/total) reported + by the backend. + type: string + memory: + description: Memory is the SPDK memory allocation reported by + the backend. + type: string + volumes: + description: Volumes is the current number of logical volumes + on this node. + format: int32 + type: integer + type: object + status: + description: Status is the backend-reported node status (e.g. online, + suspended, offline). + type: string + uptime: + description: Uptime is the node uptime as reported by the backend. + type: string + uuid: + description: UUID is the backend storage node UUID. Set once after + node-add completes. + type: string + type: object + type: object + served: true + storage: false + subresources: + status: {} + - additionalPrinterColumns: + - jsonPath: .spec.clusterRef + name: Cluster + type: string + - jsonPath: .spec.workerNode + name: Worker + type: string + - jsonPath: .spec.socketId + name: Socket + type: string + - jsonPath: .spec.slot + name: Slot + type: integer + - jsonPath: .status.phase + name: Phase + type: string + - jsonPath: .status.step.state + name: Step + type: string + - jsonPath: .status.status + name: Status + type: string + - jsonPath: .status.health + name: Health + type: boolean + - jsonPath: .status.uuid + name: UUID + priority: 1 + type: string + - jsonPath: .status.failureDomain + name: FD + priority: 1 + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha2 + schema: + openAPIV3Schema: + description: |- + StorageNode is one backend storage node: one SPDK process bound to one NUMA + socket of one Kubernetes worker. One object exists per (workerNode, slot) pair, + owned by the StorageCluster it belongs to. + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: |- + StorageNodeSpec is the desired state of one backend storage node, meaning one + SPDK process bound to one NUMA socket of one Kubernetes worker. + properties: + clusterRef: + description: |- + ClusterRef names the StorageCluster this node belongs to. The cluster also + owns this object by controller reference, so deleting the cluster deletes + its nodes. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + config: description: |- - Overrides holds per-node configuration propagated from - StorageNodeSet.spec.nodeConfigs[workerNode] on every reconcile. + Config is this node's complete configuration, copied from the + ClusterDeploymentConfig entry that produced it. It is a copy rather than a + projection, because that document is ephemeral: nothing reads it once the + node exists, deleting it changes nothing, and editing it reaches only nodes + created afterward. properties: deviceNames: description: |- - DeviceNames explicitly defines the NVMe namespace names to use on this node - (e.g. ["nvme0n1","nvme1n1"]). + DeviceNames names the devices to use. An entry is a PCI address + ("0000:5e:00.0") or a device path ("/dev/sdb,") which are the two classes + simplyblock accepts as backend storage, and a bare name ("nvme0n1") is read + as a path under /dev. One list carries both spellings, and every entry is of + the class its cluster declares in StorageCluster.spec.deviceClass: a list + mixing the two, or naming the class the cluster is not, is rejected by the + StorageNode validating webhook. Set explicitly, it overrides every filter + below. Immutable: it selects which physical devices the node owns. items: + pattern: ^([0-9a-fA-F]{4}:[0-9a-fA-F]{2}:[0-9a-fA-F]{2}\.[0-9a-fA-F]|/dev/[a-zA-Z0-9._/-]+|[a-zA-Z0-9._-]+)$ type: string type: array + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf driveSizeRange: - description: DriveSizeRange overrides the drive size range filter - for this node. + description: DriveSizeRange filters devices by size. type: string - enableCpuTopology: - description: EnableCpuTopology overrides topology-aware CPU handling - for this node. - type: boolean + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf expand: description: |- - Expand marks this node as a cluster-expansion add. When true the backend - node-add endpoint receives expand=true, triggering rebalancing behaviour - appropriate for in-place cluster growth. Overrides StorageNodeSet.spec.expand. + Expand marks this node as an addition to an already-active cluster, which + the control plane reads as a request to rebalance onto it rather than to + treat it as part of an initial layout. Immutable once set: it describes how + the node joined rather than what it is. type: boolean + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf failureDomain: description: |- - FailureDomain is the failure-domain group index (≥ 0) for this node. - Required when the parent StorageCluster has enableFailureDomains=true. - Overrides StorageNodeSet.spec.nodeFailureDomains[workerNode] when both are set. - format: int32 - minimum: 0 - type: integer + FailureDomain is the label of the fault group this node belongs to + such as rack-b, naming the physical grouping it shares with its peers rather + than indexing it. Required when the cluster has enableFailureDomains set, + and provisioning is held with a FailureDomainMissing event until it is + present. Immutable once set, which is what makes it fillable later and then + frozen: chunk placement was computed from it. + + The value takes the shape of a Kubernetes label value, because that is what + it is seeded from where a cluster carries topology labels at all. + maxLength: 63 + pattern: ^[a-zA-Z0-9]([-_.a-zA-Z0-9]*[a-zA-Z0-9])?$ + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf journalManager: - description: JournalManagerSpec overrides journal manager tuning - for this node. + description: |- + JournalManager tunes the journal manager count and per-device capacity + share for this node. Immutable: both are on-disk layout, fixed when the + devices were partitioned. properties: count: description: Count is the number of journal managers to configure. format: int32 + minimum: 1 type: integer percentPerDevice: - description: PercentPerDevice is the journal manager capacity - percentage per device. + description: PercentPerDevice is the share of each device + given to the journal. format: int32 + maximum: 100 + minimum: 1 type: integer type: object + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf pcieAllowList: - description: PcieAllowList overrides the list of PCI addresses - allowed for use on this node. + description: |- + PcieAllowList selects devices by PCI address. It is the one device field a + migration writes, merging spec.migrate.newSsdPcie into it so devices added + on the target host survive a later rebuild, so it is guarded by the + StorageNode validating webhook rather than by a marker. This and the two + PCI filters below belong to an NVMe cluster: the webhook rejects them on a + cluster whose deviceClass is LogicalBlock, because a logical block device + has no PCI address to match. items: type: string type: array pcieDenyList: - description: PcieDenyList overrides the list of PCI addresses - excluded from use on this node. + description: PcieDenyList excludes devices by PCI address. items: type: string type: array + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf pcieModel: - description: PcieModel overrides the PCI model filter for this - node. - type: string - reservedSystemCPU: - description: ReservedSystemCPU overrides the CPUs reserved for - system workloads on this node. + description: PcieModel filters devices by PCI model string. type: string - skipKubeletConfiguration: + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + sizing: description: |- - SkipKubeletConfiguration overrides whether kubelet configuration changes are - skipped for this node. - type: boolean + Sizing is what this node's huge pages and core layout were sized from. + Writable by the operator alone. + properties: + minHugePagesSize: + description: |- + MinHugePagesSize is the smallest huge-page allocation this node makes, as a + size string such as 100G or 1T, where a bare number is gigabytes. It is a + floor rather than a limit: the effective allocation is the larger of this + value and the minimum the node's device and subsystem count requires. + type: string + vcpuCount: + description: |- + VCPUCount is the number of vCPUs allocated to SPDK on this node, as an + explicit core count rather than a percentage. + format: int32 + minimum: 4 + type: integer + required: + - vcpuCount + type: object spdkImage: - description: SpdkImage overrides the SPDK image for this node - (e.g. for phased rollouts). + description: |- + SpdkImage overrides the SPDK image the control plane starts for this node, + which is what makes a phased image rollout expressible per node. type: string spdkProxyImage: description: SpdkProxyImage overrides the SPDK proxy image for @@ -6148,105 +6881,174 @@ spec: type: string spdkSystemMemory: description: |- - SpdkSystemMemory overrides the SPDK huge-page memory allocation for this node - (e.g. "4G", "512M"). + SpdkSystemMemory is the memory the control plane starts this node's SPDK + with, as a size string such as 4G or 512M. Mutable: a node whose device + count grew legitimately needs to raise it. pattern: ^[0-9]+(G|GI|GB|GiB|M|MI|MB|MiB|g|gi|gb|gib|m|mi|mb|mib)?$ type: string - ubuntuHost: - description: UbuntuHost overrides the Ubuntu host OS flag for - this node. - type: boolean + required: + - sizing type: object - socketId: - description: SocketID is the NUMA socket identifier from spec.socketsToUse - (e.g. "0", "1"). Immutable. - type: string + x-kubernetes-validations: + - message: field journalManager is immutable once set + rule: '!has(oldSelf.journalManager) || has(self.journalManager)' + - message: field deviceNames is immutable once set + rule: '!has(oldSelf.deviceNames) || has(self.deviceNames)' + - message: field pcieDenyList is immutable once set + rule: '!has(oldSelf.pcieDenyList) || has(self.pcieDenyList)' + - message: field pcieModel is immutable once set + rule: '!has(oldSelf.pcieModel) || has(self.pcieModel)' + - message: field driveSizeRange is immutable once set + rule: '!has(oldSelf.driveSizeRange) || has(self.driveSizeRange)' + - message: field failureDomain is immutable once set + rule: '!has(oldSelf.failureDomain) || has(self.failureDomain)' + - message: field expand is immutable once set + rule: '!has(oldSelf.expand) || has(self.expand)' + nodeIndex: + description: |- + NodeIndex is the position among the nodes sharing this socket, in + 0..nodesPerSocket-1. See SocketID. + format: int32 + minimum: 0 + type: integer x-kubernetes-validations: - message: field is immutable rule: self == oldSelf - socketIndex: + nodeSet: description: |- - SocketIndex is the global ordinal (socketPosition × nodesPerSocket + nodeIndex). - Used internally by the operator to select the correct backend node from the - RPC-port-sorted list in pollUUIDFromBackend. Immutable. + NodeSet is the name of the group in ClusterDeploymentConfig.nodeSets[] this + node was declared under. It is a label rather than a reference: nothing is + fetched by it, and it exists so that a node can be traced back to the + document that produced it. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + slot: + description: |- + Slot is which storage-node slot on this worker the object occupies, counted + from zero. A worker runs one node per socket per nodesPerSocket, and the + slot is the position among them. It is the identity the operator keys on: + the topology label the CSI driver reads is + storage.simplyblock.io/storage-node-uuid.., and the slot + outlives the node filling it, because only the UUID behind it changes when a + node is replaced or relocated. format: int32 + minimum: 0 type: integer x-kubernetes-validations: - message: field is immutable rule: self == oldSelf - storageNodeSetRef: - description: StorageNodeSetRef is the name of the owning StorageNodeSet. - Immutable. + socketId: + description: |- + SocketID is the NUMA socket this node is bound to, as declared in the node + set's socket list, so 0 or 1. With NodeIndex it decomposes Slot into the + pair a person reads; nothing but a print column consumes either. type: string x-kubernetes-validations: - message: field is immutable rule: self == oldSelf workerNode: description: |- - WorkerNode is the Kubernetes node hostname this StorageNode runs on. - Users may not change it directly — it is re-pointed only by the operator - during a node migration (StorageNodeOps action=migrate). The - StorageNode validating webhook rejects user-driven changes to this field. + WorkerNode is the Kubernetes worker hostname this node runs on. It is not + marked immutable, because a migration re-points it, but the StorageNode + validating webhook rejects any change made by an identity outside the + operator's namespace. type: string required: - - storageNodeSetRef + - clusterRef + - config - workerNode type: object x-kubernetes-validations: + - message: field nodeSet is immutable once set + rule: '!has(oldSelf.nodeSet) || has(self.nodeSet)' - message: field socketId is immutable once set rule: '!has(oldSelf.socketId) || has(self.socketId)' - message: field nodeIndex is immutable once set rule: '!has(oldSelf.nodeIndex) || has(self.nodeIndex)' - - message: field socketIndex is immutable once set - rule: '!has(oldSelf.socketIndex) || has(self.socketIndex)' + - message: field slot is immutable once set + rule: '!has(oldSelf.slot) || has(self.slot)' status: - description: StorageNodeStatus holds the observed state of a StorageNode. + description: StorageNodeStatus is the observed state of one storage node. properties: activeOpsRef: description: |- - ActiveOpsRef is the name of the currently active StorageNodeOps CR targeting - this node. Empty when no operation is in progress. Used for mutual exclusion. + ActiveOpsRef names the StorageNodeOps currently allowed to touch this node. + Empty when none is running. type: string failureDomain: description: |- - FailureDomain is the effective failure-domain group index for this node - as reported by the backend (≥ 0). Nil when the backend has not assigned one. - format: int32 - type: integer + FailureDomain is the failure-domain label the control plane actually + assigned, which is not necessarily the one spec.config.failureDomain + requested. + type: string health: - description: Health is the backend-reported node health flag. + description: Health is the health flag the control plane reports. type: boolean hostname: - description: Hostname is the node hostname as reported by the backend. + description: Hostname is the node hostname as the control plane reports + it. type: string latencyMetrics: description: |- - LatencyMetrics holds the fio-measured baseline NVMe-oF latency for this node, - used by the volume rebalancer to make data-placement decisions. + LatencyMetrics holds the fio-measured NVMe-oF baseline the volume + rebalancer reads. properties: baselineMeasuredAt: description: BaselineMeasuredAt is when the baseline was established. format: date-time type: string baselineP50NS: - description: BaselineP50NS is the p50 write latency (nanoseconds) - from the initial empty-cluster benchmark. + description: |- + BaselineP50NS is the p50 write latency, in nanoseconds, of the initial + empty-cluster benchmark. format: int64 + minimum: 0 type: integer baselineP99NS: - description: BaselineP99NS is the p99 write latency (nanoseconds) - from the initial empty-cluster benchmark. + description: |- + BaselineP99NS is the p99 write latency, in nanoseconds, of the same + benchmark. format: int64 + minimum: 0 type: integer nodeUUID: - description: NodeUUID is the backend storage node UUID. + description: |- + NodeUUID is the backend storage node the reading was taken against. It is + carried beside the reading rather than inferred from status.uuid, because a + baseline measured against one backend node stops describing the slot once a + replacement fills it. type: string required: - nodeUUID type: object + message: + description: |- + Message is the reason the phase is what it is: one sentence, replaced as the + node moves, and never a log. + type: string + observedGeneration: + description: |- + ObservedGeneration is the generation the rest of this status was computed + from, so a stale status can be told from a current one. + format: int64 + type: integer + phase: + description: |- + Phase is the operator's own view of this node, and the field its + provisioning branches on. + enum: + - Pending + - Provisioning + - Online + - Removing + - Offline + - Degraded + - Failed + type: string ports: - description: Ports groups network connectivity fields (addresses and - ports). + description: Ports groups the reported addresses and ports. properties: lvol: description: Lvol is the logical-volume subsystem port. @@ -6256,22 +7058,16 @@ spec: description: Management is the management IP address of the node. type: string nvmeof: - description: NvmeOf is the NVMe-oF fabric port. + description: The NVMe-oF fabric port. format: int32 type: integer rpc: - description: Rpc is the RPC/management API port. + description: Rpc is the RPC and management API port. format: int32 type: integer type: object - postedAt: - description: |- - PostedAt is the timestamp when the node-add POST was sent. - Used as a provisioning guard against duplicate POSTs. - format: date-time - type: string resources: - description: Resources groups compute and storage resource metrics. + description: Resources groups the reported compute and storage figures. properties: capacity: description: |- @@ -6283,8 +7079,8 @@ spec: sampledAt: description: |- SampledAt is when the control plane took the reading. It is not when the - object was written, and it may be considerably older if metrics - collection has stopped. + object was written, and it may be considerably older if metrics collection + has stopped. format: date-time type: string totalBytes: @@ -6300,17 +7096,35 @@ spec: type: integer type: object cpu: - description: CPU is the number of SPDK CPU cores allocated to - this node. + description: CPU is the number of SPDK cores allocated to this + node. format: int32 type: integer devices: - description: Devices is the device summary (online/total) reported - by the backend. - type: string + description: |- + Devices summarizes the node's NVMe devices. Absent until the control plane + has reported, which is what tells a node that has not reported from one that + genuinely has no devices. + properties: + online: + description: |- + Online is how many of the node's devices the control plane reports as + usable. + format: int32 + minimum: 0 + type: integer + total: + description: Total is how many devices the node has. + format: int32 + minimum: 0 + type: integer + required: + - online + - total + type: object memory: - description: Memory is the SPDK memory allocation reported by - the backend. + description: Memory is the SPDK memory allocation the control + plane reports. type: string volumes: description: Volumes is the current number of logical volumes @@ -6319,15 +7133,42 @@ spec: type: integer type: object status: - description: Status is the backend-reported node status (e.g. online, - suspended, offline). + description: |- + Status is the lifecycle the control plane reports: online, suspended, + offline, in_creation, in_restart, in_shutdown, unreachable, or timeout. The + values are the control plane's, which is why they are neither PascalCase nor + constrained by an Enum here. type: string + step: + description: |- + Step is the position of the provisioning machine, as the shared + statemachine.KubeSnapshot. The rule is what an Enum marker would do if a + marker could reach a field of a shared type. + properties: + deadline: + description: |- + Deadline is when that state expires, absent when it has none. It is an + absolute instant, so a state whose deadline passed while the controller + was down restores as already expired. + format: date-time + type: string + state: + description: |- + State is the state the machine was in. Empty means the resource has not + been reconciled yet, and restores to the graph's initial state. + type: string + type: object + x-kubernetes-validations: + - message: unknown step + rule: '!has(self.state) || self.state in [''CheckingHost'',''CheckingConfig'',''AwaitingSlot'',''Posting'',''Resolving'',''Adopting'']' uptime: - description: Uptime is the node uptime as reported by the backend. + description: Uptime is the node uptime as the control plane reports + it. type: string uuid: - description: UUID is the backend storage node UUID. Set once after - node-add completes. + description: |- + UUID is the backend node UUID. Empty means the node has neither been + provisioned nor adopted, and non-empty means steady state. type: string type: object type: object @@ -8232,16 +9073,6 @@ rules: - patch - update - watch -- apiGroups: - - "" - resources: - - events - verbs: - - create - - get - - list - - patch - - watch - apiGroups: - "" resources: @@ -8287,9 +9118,15 @@ rules: - delete - get - list - - patch - - update - watch +- apiGroups: + - "" + - events.k8s.io + resources: + - events + verbs: + - create + - patch - apiGroups: - admissionregistration.k8s.io resources: @@ -8316,6 +9153,7 @@ rules: - storageclusterops.storage.simplyblock.io - storageclusters.storage.simplyblock.io - storagenodeops.storage.simplyblock.io + - storagenodes.storage.simplyblock.io - storagepools.storage.simplyblock.io resources: - customresourcedefinitions @@ -8386,13 +9224,6 @@ rules: - patch - update - watch -- apiGroups: - - events.k8s.io - resources: - - events - verbs: - - create - - patch - apiGroups: - policy resources: @@ -8475,7 +9306,6 @@ rules: - storagedevices - storagenodeops - storagenodes - - storagenodesets - storagepoolops - storagepools - tasks @@ -8506,7 +9336,6 @@ rules: - storageclusters/finalizers - storagenodeops/finalizers - storagenodes/finalizers - - storagenodesets/finalizers - storagepoolops/finalizers - storagepools/finalizers - tasks/finalizers @@ -8555,6 +9384,14 @@ rules: - patch - update - watch +- apiGroups: + - storage.simplyblock.io + resources: + - storagenodesets + verbs: + - get + - list + - watch --- apiVersion: rbac.authorization.k8s.io/v1 kind: ClusterRole @@ -8652,6 +9489,7 @@ rules: - storagedevicemetrics - storagepoolmetrics - storageclustermetrics + - storagenodemetrics verbs: - get - list @@ -9279,15 +10117,17 @@ webhooks: service: name: simplyblock-operator-webhook-service namespace: simplyblock-operator-system - path: /validate-storage-simplyblock-io-v1alpha1-storagenode + path: /validate-storage-simplyblock-io-v1alpha2-storagenode failurePolicy: Fail + matchPolicy: Equivalent name: vstoragenode.simplyblock.io rules: - apiGroups: - storage.simplyblock.io apiVersions: - - v1alpha1 + - v1alpha2 operations: + - CREATE - UPDATE resources: - storagenodes diff --git a/operator/internal/controller/nodedrain_controller.go b/operator/internal/controller/nodedrain_controller.go deleted file mode 100644 index aef5ca062..000000000 --- a/operator/internal/controller/nodedrain_controller.go +++ /dev/null @@ -1,1539 +0,0 @@ -/* -Copyright 2025. - -Licensed under the Apache License, Version 2.0 (the "License"); -you may not use this file except in compliance with the License. -You may obtain a copy of the License at - - http://www.apache.org/licenses/LICENSE-2.0 - -Unless required by applicable law or agreed to in writing, software -distributed under the License is distributed on an "AS IS" BASIS, -WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. -See the License for the specific language governing permissions and -limitations under the License. -*/ - -package controller - -import ( - "context" - "encoding/json" - "fmt" - "net/http" - "slices" - "time" - - corev1 "k8s.io/api/core/v1" - policyv1 "k8s.io/api/policy/v1" - apierrors "k8s.io/apimachinery/pkg/api/errors" - metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" - "k8s.io/apimachinery/pkg/runtime" - "k8s.io/apimachinery/pkg/types" - "k8s.io/apimachinery/pkg/util/intstr" - "k8s.io/client-go/util/retry" - ctrl "sigs.k8s.io/controller-runtime" - "sigs.k8s.io/controller-runtime/pkg/client" - "sigs.k8s.io/controller-runtime/pkg/handler" - logf "sigs.k8s.io/controller-runtime/pkg/log" - "sigs.k8s.io/controller-runtime/pkg/reconcile" - - simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" - "github.com/simplyblock/simplyblock-operator/internal/utils" - "github.com/simplyblock/simplyblock-operator/internal/webapi" -) - -const ( - // drainNodeLabelKey is patched onto the storage pod on a draining node so the - // per-node PodDisruptionBudget can select it precisely. - drainNodeLabelKey = "simplyblock.io/drain-node" - - // drainPDBPrefix is the prefix used for per-node PodDisruptionBudget names. - drainPDBPrefix = "simplyblock-drain-" - - // managerPDBName is the name of the temporary PDB that protects the manager - // pod from eviction while it sets up storage PDB protection on its own node. - managerPDBName = "simplyblock-operator-self" - - nodeStatusOnline = "online" - nodeStatusOffline = "offline" - nodeStatusInRestart = "in_restart" - nodeStatusInShutdown = "in_shutdown" - nodeStatusInCreation = "in_creation" -) - -// NodeDrainCoordinatorReconciler coordinates graceful simplyblock node shutdown -// and restart during Kubernetes node drain events such as rolling OS upgrades. -// -// The controller implements a requeue-based state machine tracked in -// StorageNodeSet.status.drainCoordination. The full per-node flow is: -// -// 1. Detect – k8s node cordoned (spec.unschedulable=true); wait for drain slot -// 2. Shutdown – label storage pod, create blocking PDB (maxUnavailable=0), -// call simplyblock shutdown API -// 3. Confirm – poll until backend node status == nodeStatusOffline -// 4. Release – relax PDB to maxUnavailable=1; drain proceeds and pod is evicted -// 5. Reboot – node reboots (OS upgrade applied); wait for SPDK to restart -// 6. Restart – call simplyblock restart API once snode/info is reachable -// 7. Confirm – poll until backend node status == "online" -// 8. Cleanup – delete PDB, remove drain label, mark phase "complete" -// -// Concurrency is controlled by StorageCluster.spec.maxFaultTolerance: at most -// that many nodes per StorageNodeSet CR may be in the active drain window -// (phases shutdown_called, draining, restart_called) simultaneously. When -// failure domains are enabled (StorageCluster.spec.enableFailureDomains), -// MaxFaultTolerance counts distinct active failure domains instead of raw -// node count — see fdDrainGate for the full gating rule, including the -// under-provisioned-domains fallback. -// -// OpenShift MachineConfigPool (MCP) pausing is not implemented here; instead, -// set MachineConfigPool.spec.maxUnavailable to a high value and rely on the -// PDB as the actual throttle. MCP integration can be added via a dynamic client -// watching machineconfiguration.openshift.io/v1.MachineConfigPool objects. -type NodeDrainCoordinatorReconciler struct { - client.Client - Scheme *runtime.Scheme - // ManagerNodeName is the Kubernetes node this manager pod is running on, - // injected via the downward API (spec.nodeName). When set, the controller - // will create a temporary self-PDB to prevent premature eviction while - // setting up storage PDB protection on the same node. - ManagerNodeName string - TLSEnabled bool - TLSMutualEnabled bool -} - -// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodesets,verbs=get;list;watch;update;patch -// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodesets/status,verbs=get;update;patch -// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storageclusters,verbs=get;list;watch -// +kubebuilder:rbac:groups="",resources=nodes,verbs=get;list;watch -// +kubebuilder:rbac:groups="",resources=pods,verbs=get;list;watch;update;patch -// +kubebuilder:rbac:groups="",resources=secrets,verbs=get;list;watch -// +kubebuilder:rbac:groups=policy,resources=poddisruptionbudgets,verbs=get;list;watch;create;update;patch;delete - -func (r *NodeDrainCoordinatorReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) { - log := logf.FromContext(ctx) - - snCR := &simplyblockv1alpha1.StorageNodeSet{} - if err := r.Get(ctx, req.NamespacedName, snCR); err != nil { - return ctrl.Result{}, client.IgnoreNotFound(err) - } - - // Cluster UUID must be available before we can call the simplyblock API. - clusterUUID, err := utils.ResolveClusterUUID(ctx, r.Client, snCR.Namespace, snCR.Spec.ClusterName) - if err != nil { - // Not yet provisioned; nothing to gate. - return ctrl.Result{RequeueAfter: 30 * time.Second}, nil - } - - // Clean up any stale manager self-PDB left over from a previous crash - // (e.g., manager was killed after creating its PDB but before deleting it). - if r.ManagerNodeName != "" { - r.cleanupManagerPDBIfStale(ctx, snCR) - } - - // MaxConcurrentWorkerRestarts caps how many nodes can be simultaneously in the drain window. - clusterCR, err := utils.ResolveClusterCR(ctx, r.Client, snCR.Namespace, snCR.Spec.ClusterName) - if err != nil { - log.Info("Cluster CR not ready, requeuing", "cluster", snCR.Spec.ClusterName) - return ctrl.Result{RequeueAfter: 15 * time.Second}, nil - } - - // Do not start drain coordination while the cluster is still unready. - // During initial provisioning the cluster status is "unready" until enough - // nodes join for the first activation; running drain/restart logic at that - // point would interfere with the node-add flow and produce spurious errors. - if clusterCR.Status.Status == utils.ClusterStatusUnready { - log.Info("Drain coordinator skipping — cluster is unready", - "cluster", snCR.Spec.ClusterName) - return ctrl.Result{RequeueAfter: 120 * time.Second}, nil - } - - maxFaultTolerance := 1 - if clusterCR.Status.MaxConcurrentWorkerRestarts != nil && *clusterCR.Status.MaxConcurrentWorkerRestarts > 0 { - maxFaultTolerance = int(*clusterCR.Status.MaxConcurrentWorkerRestarts) - } - - // Failure-domain mode changes what MaxFaultTolerance counts against: with FD - // enabled, placement guarantees at most one erasure-coding chunk per domain, - // so the tolerable unit of loss is a whole domain, not an individual node - // (see handleDetected, which delegates the accounting to fdDrainGate). - fdEnabled := clusterCR.Spec.EnableFailureDomains != nil && *clusterCR.Spec.EnableFailureDomains - - // domainsNeededForFullDisjoint (ndcs+npcs) is the number of distinct failure - // domains required to place every stripe's chunks one-per-domain. Below that, - // at least one domain necessarily carries more than one chunk, so losing - // more than one domain at once can no longer be assumed safe — see - // fdDrainGate's chunksPerDomain calculation. 0 means "unknown" (scheme not - // yet reported); fdDrainGate currently clamps that to chunksPerDomain=1 (the - // same as a fully-disjoint layout) rather than treating it conservatively. - domainsNeededForFullDisjoint := 0 - // npcs is the failure-domain risk budget fdDrainGate spends against — - // see fdDrainGate for the full accounting rule. - npcs := 0 - if fdEnabled { - if n, err := utils.RequiredNodesFromErasureCodingScheme(clusterCR.Status.ErasureCodingScheme); err == nil { - domainsNeededForFullDisjoint = n - } else { - log.Info("Could not parse erasure coding scheme for failure-domain gate; treating as under-provisioned", - "scheme", clusterCR.Status.ErasureCodingScheme, "err", err) - } - if n, err := utils.ParityChunksFromErasureCodingScheme(clusterCR.Status.ErasureCodingScheme); err == nil { - npcs = n - } else { - log.Info("Could not parse npcs from erasure coding scheme for failure-domain gate", - "scheme", clusterCR.Status.ErasureCodingScheme, "err", err) - } - } - - apiClient := webapi.NewClient() - nextRequeue := time.Duration(0) - - // Pre-create blocking PDBs for worker nodes that are already online in the - // backend. This protects SPDK/FDB/webappapi pods from being evicted by MCP - // before the drain state machine fires, while avoiding blocking the kubelet - // reboot that is required when a new node applies its KubeletConfig/MachineConfig - // for the first time (i.e. node add flow). - for _, workerName := range snCR.Spec.WorkerNodes { - // Skip nodes that are already in an active drain — their PDB lifecycle - // is managed by the drain state machine (deleted after offline confirmed). - state := getDrainState(snCR, workerName) - if state != nil { - switch state.Phase { - case simplyblockv1alpha1.DrainPhaseShutdownCalled, - simplyblockv1alpha1.DrainPhaseDraining, - simplyblockv1alpha1.DrainPhaseRestartCalled: - continue - } - } - // Only pre-create the PDB once the storage node is online — during the - // initial node-add flow the worker must be allowed to reboot freely to - // apply KubeletConfig/MachineConfig changes. - if !isWorkerOnline(snCR, workerName) { - continue - } - if err := r.labelStoragePod(ctx, snCR.Namespace, workerName); err != nil { - log.Error(err, "Failed to pre-label storage pods", "node", workerName) - } - if err := r.ensurePDB(ctx, snCR.Namespace, workerName, 0); err != nil { - log.Error(err, "Failed to pre-create blocking PDB", "node", workerName) - } - } - - for _, workerName := range snCR.Spec.WorkerNodes { - requeue, shouldBreak := r.processWorker( - ctx, snCR, workerName, apiClient, clusterUUID, maxFaultTolerance, fdEnabled, domainsNeededForFullDisjoint, npcs, - ) - if requeue > 0 && (nextRequeue == 0 || requeue < nextRequeue) { - nextRequeue = requeue - } - if shouldBreak { - break - } - } - - // Persist the computed drain state. Use RetryOnConflict so that a 409 - // (another actor updated the CR between our snapshot and now) does not - // silently discard the phase transitions computed above — which would - // cause processWorker to re-run from the previous phase and call backend - // shutdown/restart APIs a second time. - // - // On each conflict attempt: re-read the latest CR, overlay only the drain - // coordination changes (preserving any concurrent updates to other fields), - // and retry the patch. - computedCoordination := snCR.Status.DrainCoordination - if retryErr := retry.RetryOnConflict(retry.DefaultRetry, func() error { - var latest simplyblockv1alpha1.StorageNodeSet - if err := r.Get(ctx, req.NamespacedName, &latest); err != nil { - return err - } - latestPatch := client.MergeFrom(latest.DeepCopy()) - latest.Status.DrainCoordination = computedCoordination - return r.Status().Patch(ctx, &latest, latestPatch) - }); retryErr != nil { - log.Error(retryErr, "Failed to patch drain coordination status after retries") - return ctrl.Result{RequeueAfter: 5 * time.Second}, nil - } - - if nextRequeue > 0 { - return ctrl.Result{RequeueAfter: nextRequeue}, nil - } - return ctrl.Result{}, nil -} - -// processWorker handles drain coordination for a single worker node in one -// reconcile pass. It returns the desired requeue duration and whether the -// outer worker loop should stop processing further nodes. -func (r *NodeDrainCoordinatorReconciler) processWorker( - ctx context.Context, - snCR *simplyblockv1alpha1.StorageNodeSet, - workerName string, - apiClient *webapi.Client, - clusterUUID string, - maxFaultTolerance int, - fdEnabled bool, - domainsNeededForFullDisjoint int, - npcs int, -) (requeue time.Duration, shouldBreak bool) { - log := logf.FromContext(ctx) - - node := &corev1.Node{} - if err := r.Get(ctx, types.NamespacedName{Name: workerName}, node); err != nil { - if !apierrors.IsNotFound(err) { - log.Error(err, "Failed to get worker node", "node", workerName) - } - return 0, false - } - - state := getDrainState(snCR, workerName) - cordoned := node.Spec.Unschedulable - - // Normal operation: no cordon, no state. - if !cordoned && state == nil { - return 0, false - } - - // Node was uncordoned (possibly after reboot + MCO uncordon). - if !cordoned && state != nil { - return r.processUncordoned(ctx, snCR, workerName, state, apiClient, clusterUUID) - } - - // Node is cordoned: initialise state if first observation. - if state == nil { - // Do not start drain coordination for a node that has never been online. - // MCP cordons new nodes for the initial KubeletConfig/MachineConfig reboot - // before the storage node is added — triggering drain here would create a - // blocking PDB and prevent the reboot from completing. - if !isWorkerOnline(snCR, workerName) { - log.Info("Node cordoned but not yet online — skipping drain coordination (node add reboot)", "node", workerName) - return 0, false - } - log.Info("Node cordoned — starting drain coordination", "node", workerName) - newState := simplyblockv1alpha1.NodeDrainState{ - Hostname: workerName, - Phase: simplyblockv1alpha1.DrainPhaseDetected, - StartedAt: metav1.Now(), - } - upsertDrainState(snCR, newState) - state = getDrainState(snCR, workerName) - } - - prevPhase := state.Phase - advRequeue, advErr := r.advanceStateMachine( - ctx, snCR, state, apiClient, clusterUUID, maxFaultTolerance, fdEnabled, domainsNeededForFullDisjoint, npcs, - ) - if advErr != nil { - log.Error(advErr, "Drain state machine error", "node", workerName, "phase", state.Phase) - state.Phase = simplyblockv1alpha1.DrainPhaseFailed - state.Message = advErr.Error() - } - upsertDrainState(snCR, *state) - - // If this node just entered the active drain window, stop processing - // further workers in this reconcile pass so the slot gate sees the - // correct active count on the next requeue. - if prevPhase == simplyblockv1alpha1.DrainPhaseDetected && - state.Phase == simplyblockv1alpha1.DrainPhaseShutdownCalled { - log.Info("Node entered active drain window; deferring remaining workers to next reconcile", "node", workerName) - if advRequeue == 0 { - advRequeue = 10 * time.Second - } - return advRequeue, true - } - - // If this node just completed, stop processing further workers so the - // next node's shutdown only begins once completion is persisted. - if prevPhase != simplyblockv1alpha1.DrainPhaseComplete && - prevPhase != simplyblockv1alpha1.DrainPhaseFailed && - (state.Phase == simplyblockv1alpha1.DrainPhaseComplete || state.Phase == simplyblockv1alpha1.DrainPhaseFailed) { - log.Info("Node drain complete; deferring next node shutdown to next reconcile", "node", workerName, "phase", state.Phase) - if advRequeue == 0 { - advRequeue = 5 * time.Second - } - return advRequeue, true - } - - return advRequeue, false -} - -// processUncordoned handles a worker node that has been uncordoned while drain -// state is still present (e.g., after reboot or admin intervention). -func (r *NodeDrainCoordinatorReconciler) processUncordoned( - ctx context.Context, - snCR *simplyblockv1alpha1.StorageNodeSet, - workerName string, - state *simplyblockv1alpha1.NodeDrainState, - apiClient *webapi.Client, - clusterUUID string, -) (requeue time.Duration, shouldBreak bool) { - log := logf.FromContext(ctx) - - switch state.Phase { - case simplyblockv1alpha1.DrainPhaseComplete, simplyblockv1alpha1.DrainPhaseFailed: - removeDrainState(snCR, workerName) - log.Info("Cleared terminal drain state after uncordon", "node", workerName, "phase", state.Phase) - case simplyblockv1alpha1.DrainPhaseDraining: - // Node uncordoned after reboot — call restart. - log.Info("Node uncordoned after drain; calling restart", "node", workerName) - requeue = r.handleDraining(ctx, snCR, state, apiClient, clusterUUID) - upsertDrainState(snCR, *state) - case simplyblockv1alpha1.DrainPhaseRestartCalled: - // Node uncordoned while restart polling is in progress — this is expected - // (MCP uncordons after reboot). Continue polling for online + health. - log.Info("Node uncordoned during restart polling; continuing", "node", workerName) - prevPhase := state.Phase - var err error - requeue, err = r.handleRestartCalled(ctx, snCR, state, apiClient, clusterUUID) - if err != nil { - log.Error(err, "handleRestartCalled failed after uncordon", "node", workerName) - } - upsertDrainState(snCR, *state) - // If this node just completed, stop the loop so the next node's shutdown - // only begins on the next reconcile cycle after completion is persisted. - if prevPhase != simplyblockv1alpha1.DrainPhaseComplete && - (state.Phase == simplyblockv1alpha1.DrainPhaseComplete || state.Phase == simplyblockv1alpha1.DrainPhaseFailed) { - log.Info("Node drain complete via uncordon path; deferring next node to next reconcile", "node", workerName) - if requeue == 0 { - requeue = 5 * time.Second - } - return requeue, true - } - case simplyblockv1alpha1.DrainPhaseShutdownCalled: - // Node uncordoned while waiting for backend to go offline — continue polling. - log.Info("Node uncordoned during shutdown polling; continuing", "node", workerName) - var err error - requeue, err = r.handleShutdownCalled(ctx, snCR, state, apiClient, clusterUUID) - if err != nil { - log.Error(err, "handleShutdownCalled failed after uncordon", "node", workerName) - } - upsertDrainState(snCR, *state) - default: - // Unexpected uncordon mid-sequence (e.g., admin intervention at detected phase). - log.Info("Node uncordoned mid-drain; aborting coordination", "node", workerName, "phase", state.Phase) - if cleanupErr := r.cleanupDrainResources(ctx, snCR.Namespace, workerName); cleanupErr != nil { - log.Error(cleanupErr, "Failed to clean up drain resources on abort", "node", workerName) - } - removeDrainState(snCR, workerName) - } - return requeue, false -} - -// advanceStateMachine dispatches to the handler for the current drain phase. -func (r *NodeDrainCoordinatorReconciler) advanceStateMachine( - ctx context.Context, - snCR *simplyblockv1alpha1.StorageNodeSet, - state *simplyblockv1alpha1.NodeDrainState, - apiClient *webapi.Client, - clusterUUID string, - maxFaultTolerance int, - fdEnabled bool, - domainsNeededForFullDisjoint int, - npcs int, -) (time.Duration, error) { - switch state.Phase { - case simplyblockv1alpha1.DrainPhaseDetected: - return r.handleDetected(ctx, snCR, state, apiClient, clusterUUID, maxFaultTolerance, fdEnabled, domainsNeededForFullDisjoint, npcs) - case simplyblockv1alpha1.DrainPhaseShutdownCalled: - return r.handleShutdownCalled(ctx, snCR, state, apiClient, clusterUUID) - case simplyblockv1alpha1.DrainPhaseDraining: - // Restart is only triggered from the uncordon path — not while the node - // is still cordoned. Just wait here. - state.Message = "waiting for node to be uncordoned after reboot" - return 15 * time.Second, nil - case simplyblockv1alpha1.DrainPhaseRestartCalled: - return r.handleRestartCalled(ctx, snCR, state, apiClient, clusterUUID) - case simplyblockv1alpha1.DrainPhaseComplete, simplyblockv1alpha1.DrainPhaseFailed: - return 0, nil - } - return 0, fmt.Errorf("unknown drain phase %q", state.Phase) -} - -// handleDetected waits for a drain slot then initiates the simplyblock shutdown. -// -// Gate without failure domains: number of nodes currently in -// {shutdown_called, draining, restart_called} must be less than MaxFaultTolerance. -// -// Gate with failure domains enabled: see fdDrainGate for the full accounting -// rule (a per-domain risk budget of npcs, confirmed against the backend -// team's stated requirements for 2/3/4-domain 2+2 layouts). -// -// See activeDrainWorkers/activeDrainDomainCounts. -// -// Importantly, the storage pod is labelled and a blocking PDB (maxUnavailable=0) -// is created BEFORE the slot check, so that MCP/kubectl-drain cannot evict the -// pod while this node is queued behind another drain in progress. -func (r *NodeDrainCoordinatorReconciler) handleDetected( - ctx context.Context, - snCR *simplyblockv1alpha1.StorageNodeSet, - state *simplyblockv1alpha1.NodeDrainState, - apiClient *webapi.Client, - clusterUUID string, - maxFaultTolerance int, - fdEnabled bool, - domainsNeededForFullDisjoint int, - npcs int, -) (time.Duration, error) { - log := logf.FromContext(ctx) - - // If the manager is running on the node being drained, protect it first with - // a self-PDB so MCP cannot evict it while we set up storage protection below. - managerOnThisNode := r.ManagerNodeName != "" && r.ManagerNodeName == state.Hostname - if managerOnThisNode { - if err := r.ensureManagerPDB(ctx, snCR.Namespace); err != nil { - return 10 * time.Second, fmt.Errorf("create manager self-PDB: %w", err) - } - log.Info("Manager is on draining node; created self-PDB to prevent premature eviction", "node", state.Hostname) - } - - // Label the storage pod and install a blocking PDB immediately — before the - // slot check — so MCP cannot evict this node's storage pod while it is - // waiting for a drain slot behind another in-progress drain. - if err := r.labelStoragePod(ctx, snCR.Namespace, state.Hostname); err != nil { - return 10 * time.Second, fmt.Errorf("label storage pod: %w", err) - } - if err := r.ensurePDB(ctx, snCR.Namespace, state.Hostname, 0); err != nil { - return 10 * time.Second, fmt.Errorf("create blocking PDB: %w", err) - } - - // Storage is now protected. Release the manager self-PDB so MCP can evict - // and reschedule the manager onto another node, from which it will continue - // drain coordination with the storage PDB already in place. - if managerOnThisNode { - if err := r.deleteManagerPDB(ctx, snCR.Namespace); err != nil { - log.Error(err, "Failed to delete manager self-PDB; will retry") - return 10 * time.Second, nil - } - log.Info("Storage PDB in place; released manager self-PDB — manager will migrate to another node", "node", state.Hostname) - } - - if fdEnabled { - myDomain, hasDomain := workerFailureDomain(snCR, state.Hostname) - if !hasDomain { - // FD is enabled cluster-wide but this node has no domain assignment - // yet (should not normally happen — add_node enforces the pairing). - // Fall back to the plain node-count gate rather than let an - // unassigned node bypass the budget entirely. - activeDrains := countActiveDrains(ctx, snCR, apiClient, clusterUUID) - if activeDrains >= maxFaultTolerance { - state.Message = fmt.Sprintf("waiting for drain slot (%d/%d active, no failure-domain assignment)", activeDrains, maxFaultTolerance) - log.Info("No drain slot available (node has no FD assignment; using node count)", "node", state.Hostname, "active", activeDrains, "max", maxFaultTolerance) - return 10 * time.Second, nil - } - } else { - activeWorkers := activeDrainWorkers(ctx, snCR, apiClient, clusterUUID) - activeDomainCounts := activeDrainDomainCounts(snCR, activeWorkers) - domainsAvailable := distinctDomainCount(snCR) - if blocked, reason := fdDrainGate(activeDomainCounts, myDomain, domainsAvailable, domainsNeededForFullDisjoint, npcs); blocked { - state.Message = fmt.Sprintf("waiting for drain slot (%s)", reason) - log.Info("No drain slot available for failure domain", "node", state.Hostname, "domain", myDomain, "reason", reason) - return 10 * time.Second, nil - } - } - } else { - activeDrains := countActiveDrains(ctx, snCR, apiClient, clusterUUID) - if activeDrains >= maxFaultTolerance { - state.Message = fmt.Sprintf("waiting for drain slot (%d/%d active)", activeDrains, maxFaultTolerance) - log.Info("No drain slot available, blocking PDB in place", "node", state.Hostname, "active", activeDrains, "max", maxFaultTolerance) - return 10 * time.Second, nil - } - } - - nodeUUIDs := findAllNodeUUIDs(snCR, state.Hostname) - if len(nodeUUIDs) == 0 { - return 15 * time.Second, fmt.Errorf("node %s not yet registered with backend (UUID missing)", state.Hostname) - } - - // Check backend status for all nodes before calling shutdown — they may - // already be in_shutdown or offline from a previous (possibly crashed) - // reconcile run. Handles multi-socket workers where each NUMA socket has - // its own backend node entry. - type nodeStatusEntry struct{ uuid, status string } - statuses := make([]nodeStatusEntry, 0, len(nodeUUIDs)) - for _, uuid := range nodeUUIDs { - nodeInfo, err := getBackendNodeInfo(ctx, apiClient, clusterUUID, uuid) - if err != nil { - log.Info("Failed to query backend node status before shutdown, retrying", "node", state.Hostname, "err", err) - return 15 * time.Second, nil - } - statuses = append(statuses, nodeStatusEntry{uuid: uuid, status: nodeInfo.Status}) - } - - // Fast-path: all nodes already offline — skip shutdown, allow drain. - allOffline := true - for _, s := range statuses { - if s.status != nodeStatusOffline { - allOffline = false - break - } - } - if allOffline { - log.Info("All backend nodes already offline; skipping shutdown API call", "node", state.Hostname) - if err := r.cleanupPDB(ctx, snCR.Namespace, state.Hostname); err != nil { - return 10 * time.Second, fmt.Errorf("delete PDB: %w", err) - } - state.Phase = simplyblockv1alpha1.DrainPhaseDraining - state.ActiveNodeUUID = "" - state.Message = "all backend nodes already offline; drain allowed" - return 15 * time.Second, nil - } - - // Find the first node that is not yet offline and initiate shutdown on it. - // Shutdown proceeds one node at a time: handleShutdownCalled waits for each - // node to go offline before calling shutdown on the next socket node. - for _, s := range statuses { - if s.status == nodeStatusOffline { - continue - } - state.ActiveNodeUUID = s.uuid - if s.status == nodeStatusInShutdown { - log.Info("First pending node already in_shutdown; advancing to shutdown_called", "node", state.Hostname, "nodeUUID", s.uuid) - state.Message = fmt.Sprintf("node %s already in_shutdown; waiting for offline", s.uuid) - } else { - endpoint := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/shutdown", clusterUUID, s.uuid) - body, httpStatus, err := apiClient.Do(ctx, http.MethodPost, endpoint, nil) - if err != nil || httpStatus >= 300 { - if err == nil { - err = fmt.Errorf("status %d: %s", httpStatus, string(body)) - } - return 15 * time.Second, fmt.Errorf("shutdown API for node %s: %w", s.uuid, err) - } - log.Info("Shutdown API called for first node", "node", state.Hostname, "nodeUUID", s.uuid) - state.Message = fmt.Sprintf("shutdown called for node %s; waiting for offline", s.uuid) - } - break - } - - state.Phase = simplyblockv1alpha1.DrainPhaseShutdownCalled - return 10 * time.Second, nil -} - -// handleShutdownCalled polls the active node until it is nodeStatusOffline, then calls -// shutdown on the next socket node. Only after every node on the worker is -// offline is the PDB removed and drain allowed to proceed. -func (r *NodeDrainCoordinatorReconciler) handleShutdownCalled( - ctx context.Context, - snCR *simplyblockv1alpha1.StorageNodeSet, - state *simplyblockv1alpha1.NodeDrainState, - apiClient *webapi.Client, - clusterUUID string, -) (time.Duration, error) { - log := logf.FromContext(ctx) - - nodeUUIDs := findAllNodeUUIDs(snCR, state.Hostname) - if len(nodeUUIDs) == 0 { - return 10 * time.Second, nil - } - - // Fall back to the first UUID if ActiveNodeUUID is somehow unset. - activeUUID := state.ActiveNodeUUID - if activeUUID == "" { - activeUUID = nodeUUIDs[0] - state.ActiveNodeUUID = activeUUID - } - - nodeInfo, err := getBackendNodeInfo(ctx, apiClient, clusterUUID, activeUUID) - if err != nil { - log.Info("Failed to poll backend node status, retrying", "node", state.Hostname, "err", err) - return 10 * time.Second, nil - } - if nodeInfo.Status != nodeStatusOffline { - state.Message = fmt.Sprintf("waiting for node %s to go offline, current: %s", activeUUID, nodeInfo.Status) - log.Info("Node not offline yet", "node", state.Hostname, "nodeUUID", activeUUID, "status", nodeInfo.Status) - return 10 * time.Second, nil - } - - log.Info("Node offline", "node", state.Hostname, "nodeUUID", activeUUID) - - // Advance to the next socket node in sequence. - if nextUUID := nextUUIDInList(nodeUUIDs, activeUUID); nextUUID != "" { - endpoint := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/shutdown", clusterUUID, nextUUID) - body, httpStatus, err := apiClient.Do(ctx, http.MethodPost, endpoint, nil) - if err != nil || httpStatus >= 300 { - if err == nil { - err = fmt.Errorf("status %d: %s", httpStatus, string(body)) - } - return 10 * time.Second, fmt.Errorf("shutdown API for node %s: %w", nextUUID, err) - } - log.Info("Shutdown API called for next node", "node", state.Hostname, "nodeUUID", nextUUID) - state.ActiveNodeUUID = nextUUID - state.Message = fmt.Sprintf("shutdown called for node %s; waiting for offline", nextUUID) - return 10 * time.Second, nil - } - - // All nodes offline — delete the PDB entirely to allow MCP to evict all pods. - if err := r.cleanupPDB(ctx, snCR.Namespace, state.Hostname); err != nil { - return 10 * time.Second, fmt.Errorf("delete PDB: %w", err) - } - - log.Info("All nodes offline; PDB removed — drain can proceed", "node", state.Hostname) - state.Phase = simplyblockv1alpha1.DrainPhaseDraining - state.ActiveNodeUUID = "" - state.Message = "shutdown confirmed; drain allowed" - return 15 * time.Second, nil -} - -// handleDraining is called once the node has been uncordoned and is Ready after -// reboot. It verifies SPDK is reachable before calling the simplyblock restart API. -func (r *NodeDrainCoordinatorReconciler) handleDraining( - ctx context.Context, - snCR *simplyblockv1alpha1.StorageNodeSet, - state *simplyblockv1alpha1.NodeDrainState, - apiClient *webapi.Client, - clusterUUID string, -) time.Duration { - log := logf.FromContext(ctx) - - // Verify SPDK is reachable before calling restart. - if err := checkNodeInfoReachable(ctx, state.Hostname, snCR.Namespace, r.TLSEnabled, r.TLSMutualEnabled); err != nil { - state.Message = "waiting for SPDK to become reachable after reboot" - log.Info("SPDK not yet reachable, will retry", "node", state.Hostname) - return 15 * time.Second - } - - nodeUUIDs := findAllNodeUUIDs(snCR, state.Hostname) - if len(nodeUUIDs) == 0 { - return 15 * time.Second - } - - // Call restart for the first node only. handleRestartCalled will wait for it - // to be online+healthy, then call restart on the next socket node in sequence. - firstUUID := nodeUUIDs[0] - - // Hold off if the secondary is currently restarting or shutting down. - busy, err := isPeerBusy(ctx, apiClient, clusterUUID, firstUUID) - if err != nil { - log.Info("Failed to check peer node status, retrying", "node", state.Hostname, "nodeUUID", firstUUID, "err", err) - return 15 * time.Second - } - if busy { - state.Message = fmt.Sprintf("waiting for peer node %s to finish restart/shutdown", firstUUID) - log.Info("Peer node is busy; deferring restart", "node", state.Hostname, "nodeUUID", firstUUID) - return 15 * time.Second - } - - firstNodeInfo, err := getBackendNodeInfo(ctx, apiClient, clusterUUID, firstUUID) - if err != nil { - log.Info("Failed to check node status before restart, retrying", "node", state.Hostname, "nodeUUID", firstUUID, "err", err) - return 15 * time.Second - } - if firstNodeInfo.Status == nodeStatusInRestart { - log.Info("Node already in_restart; advancing to restart_called without re-calling API", "node", state.Hostname, "nodeUUID", firstUUID) - state.ActiveNodeUUID = firstUUID - state.Phase = simplyblockv1alpha1.DrainPhaseRestartCalled - state.Message = fmt.Sprintf("node %s already in_restart; waiting for online", firstUUID) - return 10 * time.Second - } - - restartPayload := map[string]any{ - "force": true, - "node_address": utils.StorageNodeSetAPIAddress(state.Hostname, snCR.Namespace), - } - endpoint := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/restart", clusterUUID, firstUUID) - body, status, err := apiClient.Do(ctx, http.MethodPost, endpoint, restartPayload) - if err != nil || status >= 300 { - if err == nil { - err = fmt.Errorf("status %d: %s", status, string(body)) - } - log.Error(err, "Restart API failed, will retry", "node", state.Hostname, "nodeUUID", firstUUID) - return 15 * time.Second - } - log.Info("Restart API called for first node", "node", state.Hostname, "nodeUUID", firstUUID) - - state.ActiveNodeUUID = firstUUID - state.Phase = simplyblockv1alpha1.DrainPhaseRestartCalled - state.Message = fmt.Sprintf("restart called for node %s; waiting for online", firstUUID) - return 10 * time.Second -} - -// handleRestartCalled polls the backend until ALL nodes on the worker are -// "online" and healthy (covers multi-socket workers with one backend node per -// NUMA socket). Only then is the worker considered ready and the next worker -// allowed to drain, subject to maxFaultTolerance. -func (r *NodeDrainCoordinatorReconciler) handleRestartCalled( - ctx context.Context, - snCR *simplyblockv1alpha1.StorageNodeSet, - state *simplyblockv1alpha1.NodeDrainState, - apiClient *webapi.Client, - clusterUUID string, -) (time.Duration, error) { - log := logf.FromContext(ctx) - - nodeUUIDs := findAllNodeUUIDs(snCR, state.Hostname) - if len(nodeUUIDs) == 0 { - return 10 * time.Second, nil - } - - // First verify the Kubernetes node itself is Ready. - node := &corev1.Node{} - if err := r.Get(ctx, types.NamespacedName{Name: state.Hostname}, node); err != nil { - log.Info("Failed to get node for readiness check, retrying", "node", state.Hostname, "err", err) - return 10 * time.Second, nil - } - if !isNodeReady(node) { - state.Message = "waiting for Kubernetes node to become Ready" - log.Info("Node not Ready yet", "node", state.Hostname) - return 10 * time.Second, nil - } - - // Poll the active node until it is online and healthy, then call restart on - // the next socket node. The next worker is only allowed to drain once every - // socket node on this worker is back online and healthy. - activeUUID := state.ActiveNodeUUID - if activeUUID == "" { - activeUUID = nodeUUIDs[0] - state.ActiveNodeUUID = activeUUID - } - - nodeInfo, err := getBackendNodeInfo(ctx, apiClient, clusterUUID, activeUUID) - if err != nil { - log.Info("Failed to poll backend node status after restart, retrying", "node", state.Hostname, "err", err) - return 10 * time.Second, nil - } - if nodeInfo.Status != nodeStatusOnline { - state.Message = fmt.Sprintf("Kubernetes node Ready; waiting for node %s backend online, current: %s", activeUUID, nodeInfo.Status) - log.Info("Node not online yet", "node", state.Hostname, "nodeUUID", activeUUID, "status", nodeInfo.Status) - return 10 * time.Second, nil - } - if !nodeInfo.Healthy { - state.Message = fmt.Sprintf("node %s online; waiting for health check to pass", activeUUID) - log.Info("Storage node health check not passing yet", "node", state.Hostname, "nodeUUID", activeUUID) - return 10 * time.Second, nil - } - - log.Info("Node online and healthy", "node", state.Hostname, "nodeUUID", activeUUID) - - // Advance to the next socket node in sequence. - if nextUUID := nextUUIDInList(nodeUUIDs, activeUUID); nextUUID != "" { - // Hold off if the next node's secondary is currently restarting or shutting down. - busy, err := isPeerBusy(ctx, apiClient, clusterUUID, nextUUID) - if err != nil { - log.Info("Failed to check peer node status, retrying", "node", state.Hostname, "nodeUUID", nextUUID, "err", err) - return 15 * time.Second, nil - } - if busy { - state.Message = fmt.Sprintf("waiting for peer node %s to finish restart/shutdown", nextUUID) - log.Info("Peer node is busy; deferring restart", "node", state.Hostname, "nodeUUID", nextUUID) - return 15 * time.Second, nil - } - - nextNodeInfo, err := getBackendNodeInfo(ctx, apiClient, clusterUUID, nextUUID) - if err != nil { - log.Info("Failed to check node status before restart, retrying", "node", state.Hostname, "nodeUUID", nextUUID, "err", err) - return 15 * time.Second, nil - } - if nextNodeInfo.Status == nodeStatusInRestart { - log.Info("Next node already in_restart; skipping restart API call", "node", state.Hostname, "nodeUUID", nextUUID) - state.ActiveNodeUUID = nextUUID - state.Message = fmt.Sprintf("node %s already in_restart; waiting for online", nextUUID) - return 10 * time.Second, nil - } - - restartPayload := map[string]any{ - "force": true, - "node_address": utils.StorageNodeSetAPIAddress(state.Hostname, snCR.Namespace), - } - endpoint := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/restart", clusterUUID, nextUUID) - body, httpStatus, err := apiClient.Do(ctx, http.MethodPost, endpoint, restartPayload) - if err != nil || httpStatus >= 300 { - if err == nil { - err = fmt.Errorf("status %d: %s", httpStatus, string(body)) - } - log.Error(err, "Restart API failed for next node, will retry", "node", state.Hostname, "nodeUUID", nextUUID) - return 15 * time.Second, nil - } - log.Info("Restart API called for next node", "node", state.Hostname, "nodeUUID", nextUUID) - state.ActiveNodeUUID = nextUUID - state.Message = fmt.Sprintf("restart called for node %s; waiting for online", nextUUID) - return 10 * time.Second, nil - } - - // All socket nodes are back online — wait for the cluster to finish - // rebalancing before releasing this drain slot to the next worker. - rebalancing, err := isClusterRebalancing(ctx, apiClient, clusterUUID) - if err != nil { - log.Info("Failed to check cluster rebalancing status, retrying", "node", state.Hostname, "err", err) - return 30 * time.Second, nil - } - if rebalancing { - state.Message = "all nodes online; waiting for cluster rebalancing to complete" - log.Info("Cluster is rebalancing; holding drain slot before next worker", "node", state.Hostname) - return 30 * time.Second, nil - } - - // Clean up PDB and pod label. - if err := r.cleanupDrainResources(ctx, snCR.Namespace, state.Hostname); err != nil { - log.Error(err, "Cleanup failed, will retry", "node", state.Hostname) - return 10 * time.Second, nil - } - - log.Info("All storage nodes back online and cluster rebalancing complete; drain coordination complete", "node", state.Hostname, "nodeCount", len(nodeUUIDs)) - state.Phase = simplyblockv1alpha1.DrainPhaseComplete - state.ActiveNodeUUID = "" - state.Message = "all storage nodes back online; cluster rebalancing complete" - return 0, nil -} - -// ensurePDB creates or updates the per-node PodDisruptionBudget. -// maxUnavailable=0 blocks eviction; maxUnavailable=1 allows it. -func (r *NodeDrainCoordinatorReconciler) ensurePDB( - ctx context.Context, - namespace, nodeName string, - maxUnavailable int, -) error { - pdbName := drainPDBPrefix + sanitizeLabelValue(nodeName) - maxUnavailableVal := intstr.FromInt32(int32(maxUnavailable)) - - desired := &policyv1.PodDisruptionBudget{ - ObjectMeta: metav1.ObjectMeta{ - Name: pdbName, - Namespace: namespace, - Labels: map[string]string{ - "simplyblock.io/managed-by": "drain-coordinator", - }, - }, - Spec: policyv1.PodDisruptionBudgetSpec{ - MaxUnavailable: &maxUnavailableVal, - Selector: &metav1.LabelSelector{ - MatchLabels: map[string]string{ - drainNodeLabelKey: nodeName, - }, - }, - }, - } - - existing := &policyv1.PodDisruptionBudget{} - err := r.Get(ctx, types.NamespacedName{Name: pdbName, Namespace: namespace}, existing) - if apierrors.IsNotFound(err) { - return r.Create(ctx, desired) - } - if err != nil { - return err - } - - patch := client.MergeFrom(existing.DeepCopy()) - existing.Spec.MaxUnavailable = &maxUnavailableVal - return r.Patch(ctx, existing, patch) -} - -// labelStoragePod patches the SPDK pod (role=simplyblock-storage-node) and -// simplyblock-webappapi pods running on nodeName with drainNodeLabelKey so the -// per-node PDB can select and protect them during drain coordination. -// These are regular pods (not DaemonSet-owned), so PDB eviction blocking works. -func (r *NodeDrainCoordinatorReconciler) labelStoragePod( - ctx context.Context, - namespace, nodeName string, -) error { - // Label selectors for pods that must be protected during drain. - // FDB pods are included so that coordinators and log processes are not - // evicted before the simplyblock shutdown is confirmed — losing an FDB - // process mid-shutdown causes transaction timeouts in the webAPI. - targetSelectors := []map[string]string{ - {"role": "simplyblock-storage-node"}, - {"app": "simplyblock-webappapi"}, - {"foundationdb.org/fdb-cluster-name": "simplyblock-fdb-cluster"}, - } - - for _, selector := range targetSelectors { - podList := &corev1.PodList{} - if err := r.List(ctx, podList, - client.InNamespace(namespace), - client.MatchingLabels(selector), - ); err != nil { - return fmt.Errorf("list pods %v: %w", selector, err) - } - - for i := range podList.Items { - pod := &podList.Items[i] - if pod.Spec.NodeName != nodeName { - continue - } - if pod.Labels[drainNodeLabelKey] == sanitizeLabelValue(nodeName) { - continue // already labelled - } - patch := client.MergeFrom(pod.DeepCopy()) - if pod.Labels == nil { - pod.Labels = map[string]string{} - } - pod.Labels[drainNodeLabelKey] = sanitizeLabelValue(nodeName) - if err := r.Patch(ctx, pod, patch); err != nil { - return fmt.Errorf("patch pod %s: %w", pod.Name, err) - } - } - } - return nil -} - -// cleanupDrainResources deletes the per-node PDB and removes the drain label -// cleanupPDB deletes the per-node PodDisruptionBudget. -func (r *NodeDrainCoordinatorReconciler) cleanupPDB( - ctx context.Context, - namespace, nodeName string, -) error { - pdbName := drainPDBPrefix + sanitizeLabelValue(nodeName) - pdb := &policyv1.PodDisruptionBudget{} - if err := r.Get(ctx, types.NamespacedName{Name: pdbName, Namespace: namespace}, pdb); err == nil { - if err := r.Delete(ctx, pdb); err != nil && !apierrors.IsNotFound(err) { - return fmt.Errorf("delete PDB: %w", err) - } - } else if !apierrors.IsNotFound(err) { - return fmt.Errorf("get PDB: %w", err) - } - return nil -} - -// cleanupDrainResources deletes the per-node PDB and removes the drain label -// from any pods that still carry it. -func (r *NodeDrainCoordinatorReconciler) cleanupDrainResources( - ctx context.Context, - namespace, nodeName string, -) error { - if err := r.cleanupPDB(ctx, namespace, nodeName); err != nil { - return err - } - - // Remove drain label from any surviving pods (e.g., if eviction didn't happen). - podList := &corev1.PodList{} - if err := r.List(ctx, podList, - client.InNamespace(namespace), - client.MatchingLabels{drainNodeLabelKey: sanitizeLabelValue(nodeName)}, - ); err != nil { - return fmt.Errorf("list labelled pods: %w", err) - } - for i := range podList.Items { - pod := &podList.Items[i] - patch := client.MergeFrom(pod.DeepCopy()) - delete(pod.Labels, drainNodeLabelKey) - if err := r.Patch(ctx, pod, patch); err != nil && !apierrors.IsNotFound(err) { - return fmt.Errorf("remove drain label from pod %s: %w", pod.Name, err) - } - } - return nil -} - -// ensureManagerPDB creates a blocking PDB (maxUnavailable=0) for the manager -// pod itself, preventing MCP from evicting the manager while storage PDB -// protection is being set up on the same node. -func (r *NodeDrainCoordinatorReconciler) ensureManagerPDB(ctx context.Context, namespace string) error { - maxUnavailable := intstr.FromInt32(0) - desired := &policyv1.PodDisruptionBudget{ - ObjectMeta: metav1.ObjectMeta{ - Name: managerPDBName, - Namespace: namespace, - Labels: map[string]string{ - "simplyblock.io/managed-by": "drain-coordinator", - }, - }, - Spec: policyv1.PodDisruptionBudgetSpec{ - MaxUnavailable: &maxUnavailable, - Selector: &metav1.LabelSelector{ - MatchLabels: map[string]string{ - "app": "simplyblock-operator", - }, - }, - }, - } - existing := &policyv1.PodDisruptionBudget{} - err := r.Get(ctx, types.NamespacedName{Name: managerPDBName, Namespace: namespace}, existing) - if apierrors.IsNotFound(err) { - return r.Create(ctx, desired) - } - return err -} - -// deleteManagerPDB removes the manager self-PDB, allowing MCP to evict and -// reschedule the manager onto a non-draining node. -func (r *NodeDrainCoordinatorReconciler) deleteManagerPDB(ctx context.Context, namespace string) error { - pdb := &policyv1.PodDisruptionBudget{} - err := r.Get(ctx, types.NamespacedName{Name: managerPDBName, Namespace: namespace}, pdb) - if apierrors.IsNotFound(err) { - return nil - } - if err != nil { - return err - } - return r.Delete(ctx, pdb) -} - -// cleanupManagerPDBIfStale deletes the manager self-PDB when the manager's own -// node is no longer in the detected drain phase. This handles crash-recovery -// where the PDB was created but the manager was killed before deleting it. -func (r *NodeDrainCoordinatorReconciler) cleanupManagerPDBIfStale(ctx context.Context, snCR *simplyblockv1alpha1.StorageNodeSet) { - log := logf.FromContext(ctx) - state := getDrainState(snCR, r.ManagerNodeName) - // PDB is only needed transiently during DrainPhaseDetected on the manager's node. - if state != nil && state.Phase == simplyblockv1alpha1.DrainPhaseDetected { - return - } - if err := r.deleteManagerPDB(ctx, snCR.Namespace); err != nil { - log.Error(err, "Failed to clean up stale manager self-PDB") - } -} - -// SetupWithManager wires the controller to watch StorageNodeSet CRs, k8s Nodes, -// and the pods that labelStoragePod tracks. Watching pods ensures that when a -// tracked pod is recreated (e.g. after a crash) the reconcile fires immediately -// to re-apply the drain label, keeping PDB protection continuous. -func (r *NodeDrainCoordinatorReconciler) SetupWithManager(mgr ctrl.Manager) error { - return ctrl.NewControllerManagedBy(mgr). - For(&simplyblockv1alpha1.StorageNodeSet{}). - Named("nodedrain"). - Watches( - &corev1.Node{}, - handler.EnqueueRequestsFromMapFunc(r.nodeToStorageNodeSetRequests), - ). - Watches( - &corev1.Pod{}, - handler.EnqueueRequestsFromMapFunc(r.trackedPodToStorageNodeSetRequests), - ). - Complete(r) -} - -// trackedPodToStorageNodeSetRequests maps a Pod event to the StorageNodeSet CR(s) -// whose workerNodes contains the pod's node. Only fires for the pod selectors -// that labelStoragePod tracks, so unrelated pods cause no reconcile churn. -func (r *NodeDrainCoordinatorReconciler) trackedPodToStorageNodeSetRequests( - ctx context.Context, - obj client.Object, -) []reconcile.Request { - pod := obj.(*corev1.Pod) - - if pod.Spec.NodeName == "" { - return nil - } - - labels := pod.GetLabels() - tracked := labels["role"] == "simplyblock-storage-node" || - labels["app"] == "simplyblock-webappapi" || - labels["foundationdb.org/fdb-cluster-name"] == "simplyblock-fdb-cluster" - if !tracked { - return nil - } - - var snList simplyblockv1alpha1.StorageNodeSetList - if err := r.List(ctx, &snList); err != nil { - return nil - } - - var requests []reconcile.Request - for _, sn := range snList.Items { - if slices.Contains(sn.Spec.WorkerNodes, pod.Spec.NodeName) { - requests = append(requests, reconcile.Request{ - NamespacedName: types.NamespacedName{ - Namespace: sn.Namespace, - Name: sn.Name, - }, - }) - } - } - return requests -} - -// nodeToStorageNodeSetRequests maps a k8s Node event to the StorageNodeSet CR(s) -// that list the node in spec.workerNodes. -func (r *NodeDrainCoordinatorReconciler) nodeToStorageNodeSetRequests( - ctx context.Context, - obj client.Object, -) []reconcile.Request { - node := obj.(*corev1.Node) - - var snList simplyblockv1alpha1.StorageNodeSetList - if err := r.List(ctx, &snList); err != nil { - return nil - } - - var requests []reconcile.Request - for _, sn := range snList.Items { - if slices.Contains(sn.Spec.WorkerNodes, node.Name) { - requests = append(requests, reconcile.Request{ - NamespacedName: types.NamespacedName{ - Namespace: sn.Namespace, - Name: sn.Name, - }, - }) - } - } - return requests -} - -/* -------------------- helpers -------------------- */ - -// backendNodeInfo holds the relevant fields from the storage-nodes API response. -type backendNodeInfo struct { - UUID string `json:"id"` - Status string `json:"status"` - Healthy bool `json:"health_check"` - SecondaryNodeID string `json:"secondary_node_id"` -} - -// getBackendNodeInfo fetches status and health_check for a node UUID. -func getBackendNodeInfo( - ctx context.Context, - apiClient *webapi.Client, - clusterUUID, nodeUUID string, -) (backendNodeInfo, error) { - endpoint := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s", clusterUUID, nodeUUID) - body, statusCode, err := apiClient.Do(ctx, http.MethodGet, endpoint, nil) - if err != nil || statusCode >= 300 { - if err == nil { - err = fmt.Errorf("status %d: %s", statusCode, string(body)) - } - return backendNodeInfo{}, err - } - - var info backendNodeInfo - if err := json.Unmarshal(body, &info); err != nil { - return backendNodeInfo{}, fmt.Errorf("unmarshal node info: %w", err) - } - return info, nil -} - -// getDrainState returns the drain state entry for the given hostname, or nil. -func getDrainState(snCR *simplyblockv1alpha1.StorageNodeSet, hostname string) *simplyblockv1alpha1.NodeDrainState { - for i := range snCR.Status.DrainCoordination { - if snCR.Status.DrainCoordination[i].Hostname == hostname { - return &snCR.Status.DrainCoordination[i] - } - } - return nil -} - -// upsertDrainState inserts or updates the drain state entry for state.Hostname. -func upsertDrainState(snCR *simplyblockv1alpha1.StorageNodeSet, state simplyblockv1alpha1.NodeDrainState) { - for i := range snCR.Status.DrainCoordination { - if snCR.Status.DrainCoordination[i].Hostname == state.Hostname { - snCR.Status.DrainCoordination[i] = state - return - } - } - snCR.Status.DrainCoordination = append(snCR.Status.DrainCoordination, state) -} - -// removeDrainState removes the drain state entry for the given hostname. -func removeDrainState(snCR *simplyblockv1alpha1.StorageNodeSet, hostname string) { - filtered := snCR.Status.DrainCoordination[:0] - for _, s := range snCR.Status.DrainCoordination { - if s.Hostname != hostname { - filtered = append(filtered, s) - } - } - snCR.Status.DrainCoordination = filtered -} - -// countActiveDrains returns the number of *workers* currently in the active -// drain window. maxFaultTolerance is a worker-level limit: with tolerance=2, -// two workers may drain simultaneously regardless of how many socket nodes each -// worker carries. It takes the maximum of: -// -// 1. Controller drain state (phases: shutdown_called, draining, restart_called) -// — one DrainCoordination entry per worker, so this is already worker-level. -// 2. Backend node status (in_shutdown, in_restart) grouped by worker hostname -// — prevents a controller crash/restart from resetting the count to zero -// while nodes are still transitioning in the backend. -func countActiveDrains( - ctx context.Context, - snCR *simplyblockv1alpha1.StorageNodeSet, - apiClient *webapi.Client, - clusterUUID string, -) int { - // Count workers in the active drain window from controller state. - // DrainCoordination has one entry per worker, so no grouping needed. - controllerCount := 0 - for _, s := range snCR.Status.DrainCoordination { - switch s.Phase { - case simplyblockv1alpha1.DrainPhaseShutdownCalled, - simplyblockv1alpha1.DrainPhaseDraining, - simplyblockv1alpha1.DrainPhaseRestartCalled: - controllerCount++ - } - } - - // Count workers with at least one backend node in transition, grouped by - // hostname so a 2-socket worker counts as 1, not 2. - // On API error, conservatively mark that worker active to prevent a - // transient failure from opening a drain slot. - activeWorkers := map[string]bool{} - for _, n := range snCR.Status.Nodes { - if n.UUID == "" || n.Hostname == "" { - continue - } - if activeWorkers[n.Hostname] { - continue // already counted this worker - } - info, err := getBackendNodeInfo(ctx, apiClient, clusterUUID, n.UUID) - if err != nil { - activeWorkers[n.Hostname] = true - continue - } - switch info.Status { - case nodeStatusInShutdown, nodeStatusInRestart: - activeWorkers[n.Hostname] = true - } - } - backendCount := len(activeWorkers) - - if backendCount > controllerCount { - return backendCount - } - return controllerCount -} - -// activeDrainWorkers returns the set of worker hostnames currently in the -// active drain window, merging the same two sources as countActiveDrains -// (controller-tracked phases and backend node status) into an actual set -// rather than a count — needed so the failure-domain gate can test which -// domains those workers belong to, not just how many there are. -func activeDrainWorkers( - ctx context.Context, - snCR *simplyblockv1alpha1.StorageNodeSet, - apiClient *webapi.Client, - clusterUUID string, -) map[string]bool { - active := map[string]bool{} - for _, s := range snCR.Status.DrainCoordination { - switch s.Phase { - case simplyblockv1alpha1.DrainPhaseShutdownCalled, - simplyblockv1alpha1.DrainPhaseDraining, - simplyblockv1alpha1.DrainPhaseRestartCalled: - active[s.Hostname] = true - } - } - - seen := map[string]bool{} - for _, n := range snCR.Status.Nodes { - if n.UUID == "" || n.Hostname == "" || seen[n.Hostname] { - continue - } - seen[n.Hostname] = true - info, err := getBackendNodeInfo(ctx, apiClient, clusterUUID, n.UUID) - if err != nil { - // Conservatively mark active on a query error, same as countActiveDrains. - active[n.Hostname] = true - continue - } - switch info.Status { - case nodeStatusInShutdown, nodeStatusInRestart: - active[n.Hostname] = true - } - } - return active -} - -// workerFailureDomain returns the effective failure-domain group for a worker -// hostname from status.nodes[].failureDomain, which is populated by the backend -// API response and is therefore authoritative even for nodes added outside the -// operator. Returns (0, false) when the backend has not assigned a domain. -func workerFailureDomain(snCR *simplyblockv1alpha1.StorageNodeSet, hostname string) (int32, bool) { - for _, ns := range snCR.Status.Nodes { - if ns.Hostname == hostname && ns.FailureDomain != nil { - return *ns.FailureDomain, true - } - } - return 0, false -} - -// activeDrainDomainCounts maps a set of active worker hostnames to how many -// of them fall in each failure domain — the per-domain node count fdDrainGate -// needs to compute risk. Workers with no domain assignment are excluded — the -// caller falls back to the plain node-count gate for those. -func activeDrainDomainCounts(snCR *simplyblockv1alpha1.StorageNodeSet, activeWorkers map[string]bool) map[int32]int { - counts := map[int32]int{} - for w := range activeWorkers { - if d, ok := workerFailureDomain(snCR, w); ok { - counts[d]++ - } - } - return counts -} - -// distinctDomainCount returns the number of distinct effective failure-domain -// groups currently assigned across a StorageNodeSet's workers, from -// status.nodes[].failureDomain — the same authoritative source as -// workerFailureDomain, not the deploy-time-requested spec.nodeFailureDomains. -func distinctDomainCount(snCR *simplyblockv1alpha1.StorageNodeSet) int { - domains := map[int32]bool{} - for _, ns := range snCR.Status.Nodes { - if ns.FailureDomain != nil { - domains[*ns.FailureDomain] = true - } - } - return len(domains) -} - -// fdDrainGate decides whether a candidate node in myDomain may proceed past -// the Detected phase when failure domains are enabled. Kept as a pure -// decision function (all inputs precomputed by the caller) so the policy is -// directly unit-testable without the k8s/backend side effects in -// handleDetected. -// -// The risk budget is npcs, spent per domain at a rate of -// min(nodesDownInThatDomain, chunksPerDomain), where chunksPerDomain = -// ceil(domainsNeededForFullDisjoint / domainsAvailable) — placement spreads a -// stripe's domainsNeededForFullDisjoint (ndcs+npcs) chunks as evenly as -// possible across the domains that actually exist, so with fewer domains -// than chunks, at least one domain holds more than one chunk. A domain that -// already has chunksPerDomain nodes down has maxed its contribution to the -// budget — further nodes in THAT SAME domain are free — but a node in a -// domain that hasn't maxed out yet is only safe if the combined risk across -// every affected domain still leaves room. This reduces to the familiar "up -// to npcs whole domains are free" rule once there are >= ndcs+npcs domains -// (chunksPerDomain == 1). -// -// Confirmed against the backend team's stated requirements for a 2+2 layout: -// 2 domains -> "1 whole domain" or "1 node in each of the 2 domains" are the -// only safe combinations; 3 domains -> only 1 domain may be fully down, not -// 2; 4 domains -> 2 domains may be fully down (the well-provisioned case). -func fdDrainGate( - activeDomainCounts map[int32]int, - myDomain int32, - domainsAvailable int, - domainsNeededForFullDisjoint int, - npcs int, -) (blocked bool, reason string) { - chunksPerDomain := domainsNeededForFullDisjoint - if domainsAvailable > 0 { - chunksPerDomain = (domainsNeededForFullDisjoint + domainsAvailable - 1) / domainsAvailable // ceil - } - if chunksPerDomain < 1 { - chunksPerDomain = 1 - } - - if activeDomainCounts[myDomain] >= chunksPerDomain { - return false, "" - } - - currentRisk := 0 - for _, count := range activeDomainCounts { - risk := count - if risk > chunksPerDomain { - risk = chunksPerDomain - } - currentRisk += risk - } - if currentRisk+1 > npcs { - return true, fmt.Sprintf( - "failure domain %d: %d/%d failure-domain risk budget already committed "+ - "(%d domain(s) available, %d chunk(s)/domain worst case)", - myDomain, currentRisk, npcs, domainsAvailable, chunksPerDomain) - } - return false, "" -} - -// findNodeUUID returns the backend UUID for the given hostname from StorageNodeSet status. -func findNodeUUID(snCR *simplyblockv1alpha1.StorageNodeSet, hostname string) string { - for _, n := range snCR.Status.Nodes { - if n.Hostname == hostname { - return n.UUID - } - } - return "" -} - -// findAllNodeUUIDs returns all backend UUIDs registered for the given hostname. -// For multi-socket workers (socketsToUse configured) each NUMA socket produces -// a separate backend node entry sharing the same hostname, so this may return -// more than one UUID. Returns a single-element slice for standard workers. -func findAllNodeUUIDs(snCR *simplyblockv1alpha1.StorageNodeSet, hostname string) []string { - var uuids []string - for _, n := range snCR.Status.Nodes { - if n.Hostname == hostname && n.UUID != "" { - uuids = append(uuids, n.UUID) - } - } - return uuids -} - -// nextUUIDInList returns the element immediately after current in uuids, -// or an empty string if current is the last element or not found. -func nextUUIDInList(uuids []string, current string) string { - for i, u := range uuids { - if u == current && i+1 < len(uuids) { - return uuids[i+1] - } - } - return "" -} - -// isNodeReady returns true when the node has a Ready condition with status True. -// isWorkerOnline returns true if the worker node has status "online" in the -// StorageNodeSet status, meaning it has been fully added and is serving I/O. -func isWorkerOnline(snCR *simplyblockv1alpha1.StorageNodeSet, workerName string) bool { - for _, n := range snCR.Status.Nodes { - if n.Hostname == workerName { - return n.Status == nodeStatusOnline - } - } - return false -} - -func isNodeReady(node *corev1.Node) bool { - for _, cond := range node.Status.Conditions { - if cond.Type == corev1.NodeReady { - return cond.Status == corev1.ConditionTrue - } - } - return false -} - -// isPeerBusy lists all storage nodes in the cluster and finds the one that -// has nodeUUID as its secondary_node_id. That node is the HA peer (primary) -// of the node we want to restart. Returns true if the peer is in_restart or -// in_shutdown, which means we must defer the restart. -func isPeerBusy( - ctx context.Context, - apiClient *webapi.Client, - clusterUUID, nodeUUID string, -) (bool, error) { - endpoint := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes", clusterUUID) - body, statusCode, err := apiClient.Do(ctx, http.MethodGet, endpoint, nil) - if err != nil || statusCode >= 300 { - if err == nil { - err = fmt.Errorf("status %d: %s", statusCode, string(body)) - } - return false, fmt.Errorf("list storage nodes: %w", err) - } - - var nodes []backendNodeInfo - if err := json.Unmarshal(body, &nodes); err != nil { - return false, fmt.Errorf("unmarshal storage nodes: %w", err) - } - - for _, n := range nodes { - if n.SecondaryNodeID == nodeUUID { - return n.Status == nodeStatusInRestart || n.Status == nodeStatusInShutdown, nil - } - } - return false, nil -} - -// isClusterRebalancing returns true when the simplyblock cluster is actively -// rebalancing data. The drain slot is held until rebalancing completes so that -// the next worker drain does not start while the cluster is already under load. -func isClusterRebalancing( - ctx context.Context, - apiClient *webapi.Client, - clusterUUID string, -) (bool, error) { - endpoint := fmt.Sprintf("/api/v2/clusters/%s", clusterUUID) - body, statusCode, err := apiClient.Do(ctx, http.MethodGet, endpoint, nil) - if err != nil || statusCode >= 300 { - if err == nil { - err = fmt.Errorf("status %d: %s", statusCode, string(body)) - } - return false, fmt.Errorf("get cluster info: %w", err) - } - - var info struct { - Rebalancing bool `json:"is_re_balancing"` - } - if err := json.Unmarshal(body, &info); err != nil { - return false, fmt.Errorf("unmarshal cluster info: %w", err) - } - return info.Rebalancing, nil -} - -// sanitizeLabelValue truncates to 63 chars (Kubernetes label value limit). -func sanitizeLabelValue(name string) string { - if len(name) > 63 { - return name[:63] - } - return name -} diff --git a/operator/internal/controller/nodedrain_controller_unit_test.go b/operator/internal/controller/nodedrain_controller_unit_test.go deleted file mode 100644 index eca2c02f4..000000000 --- a/operator/internal/controller/nodedrain_controller_unit_test.go +++ /dev/null @@ -1,1316 +0,0 @@ -package controller - -import ( - "context" - "net/http" - "net/http/httptest" - "testing" - "time" - - simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" - simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" - "github.com/simplyblock/simplyblock-operator/internal/utils" - "github.com/simplyblock/simplyblock-operator/internal/webapi" - corev1 "k8s.io/api/core/v1" - policyv1 "k8s.io/api/policy/v1" - apierrors "k8s.io/apimachinery/pkg/api/errors" - metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" - "k8s.io/apimachinery/pkg/util/intstr" - ctrl "sigs.k8s.io/controller-runtime" - "sigs.k8s.io/controller-runtime/pkg/client" - "sigs.k8s.io/controller-runtime/pkg/client/fake" - "sigs.k8s.io/controller-runtime/pkg/client/interceptor" -) - -// ---- pure helpers ---- - -func TestGetDrainState(t *testing.T) { - snCR := &simplyblockv1alpha1.StorageNodeSet{ - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - DrainCoordination: []simplyblockv1alpha1.NodeDrainState{ - {Hostname: "node-a", Phase: simplyblockv1alpha1.DrainPhaseDetected}, - {Hostname: "node-b", Phase: simplyblockv1alpha1.DrainPhaseDraining}, - }, - }, - } - - s := getDrainState(snCR, "node-a") - if s == nil || s.Phase != simplyblockv1alpha1.DrainPhaseDetected { - t.Fatalf("expected DrainPhaseDetected for node-a, got %v", s) - } - - if getDrainState(snCR, "node-c") != nil { - t.Fatalf("expected nil for unknown hostname") - } -} - -func TestUpsertDrainState(t *testing.T) { - snCR := &simplyblockv1alpha1.StorageNodeSet{} - - upsertDrainState(snCR, simplyblockv1alpha1.NodeDrainState{Hostname: "node-a", Phase: simplyblockv1alpha1.DrainPhaseDetected}) - if len(snCR.Status.DrainCoordination) != 1 { - t.Fatalf("expected 1 entry after insert, got %d", len(snCR.Status.DrainCoordination)) - } - - upsertDrainState(snCR, simplyblockv1alpha1.NodeDrainState{Hostname: "node-a", Phase: simplyblockv1alpha1.DrainPhaseShutdownCalled}) - if len(snCR.Status.DrainCoordination) != 1 { - t.Fatalf("expected 1 entry after update (no duplicate), got %d", len(snCR.Status.DrainCoordination)) - } - if snCR.Status.DrainCoordination[0].Phase != simplyblockv1alpha1.DrainPhaseShutdownCalled { - t.Fatalf("expected updated phase, got %q", snCR.Status.DrainCoordination[0].Phase) - } - - upsertDrainState(snCR, simplyblockv1alpha1.NodeDrainState{Hostname: "node-b", Phase: simplyblockv1alpha1.DrainPhaseDraining}) - if len(snCR.Status.DrainCoordination) != 2 { - t.Fatalf("expected 2 entries after inserting second node, got %d", len(snCR.Status.DrainCoordination)) - } -} - -func TestRemoveDrainState(t *testing.T) { - snCR := &simplyblockv1alpha1.StorageNodeSet{ - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - DrainCoordination: []simplyblockv1alpha1.NodeDrainState{ - {Hostname: "node-a", Phase: simplyblockv1alpha1.DrainPhaseComplete}, - {Hostname: "node-b", Phase: simplyblockv1alpha1.DrainPhaseDraining}, - }, - }, - } - - removeDrainState(snCR, "node-a") - if len(snCR.Status.DrainCoordination) != 1 { - t.Fatalf("expected 1 entry after remove, got %d", len(snCR.Status.DrainCoordination)) - } - if snCR.Status.DrainCoordination[0].Hostname != "node-b" { - t.Fatalf("expected node-b to remain, got %q", snCR.Status.DrainCoordination[0].Hostname) - } - - // removing an absent entry is a no-op - removeDrainState(snCR, "node-missing") - if len(snCR.Status.DrainCoordination) != 1 { - t.Fatalf("expected no change when removing absent hostname") - } -} - -func TestFindNodeUUID(t *testing.T) { - snCR := &simplyblockv1alpha1.StorageNodeSet{ - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - Nodes: []simplyblockv1alpha1.NodeStatus{ - {Hostname: "node-a", UUID: "uuid-a"}, - {Hostname: "node-b", UUID: "uuid-b"}, - }, - }, - } - - if got := findNodeUUID(snCR, "node-a"); got != "uuid-a" { - t.Fatalf("expected uuid-a, got %q", got) - } - if got := findNodeUUID(snCR, "node-missing"); got != "" { - t.Fatalf("expected empty string for unknown hostname, got %q", got) - } -} - -func TestIsWorkerOnline(t *testing.T) { - snCR := &simplyblockv1alpha1.StorageNodeSet{ - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - Nodes: []simplyblockv1alpha1.NodeStatus{ - {Hostname: "node-online", Status: "online"}, - {Hostname: "node-offline", Status: "offline"}, - }, - }, - } - - if !isWorkerOnline(snCR, "node-online") { - t.Fatalf("expected node-online to be online") - } - if isWorkerOnline(snCR, "node-offline") { - t.Fatalf("expected node-offline to not be online") - } - if isWorkerOnline(snCR, "node-missing") { - t.Fatalf("expected missing node to not be online") - } -} - -func TestIsNodeReady(t *testing.T) { - ready := &corev1.Node{ - Status: corev1.NodeStatus{ - Conditions: []corev1.NodeCondition{ - {Type: corev1.NodeReady, Status: corev1.ConditionTrue}, - }, - }, - } - notReady := &corev1.Node{ - Status: corev1.NodeStatus{ - Conditions: []corev1.NodeCondition{ - {Type: corev1.NodeReady, Status: corev1.ConditionFalse}, - }, - }, - } - noConditions := &corev1.Node{} - - if !isNodeReady(ready) { - t.Fatalf("expected ready node to return true") - } - if isNodeReady(notReady) { - t.Fatalf("expected not-ready node to return false") - } - if isNodeReady(noConditions) { - t.Fatalf("expected node with no conditions to return false") - } -} - -func TestSanitizeLabelValue(t *testing.T) { - short := "node-abc" - if got := sanitizeLabelValue(short); got != short { - t.Fatalf("short value should be unchanged, got %q", got) - } - - exactly63 := "aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa"[:63] - if got := sanitizeLabelValue(exactly63); got != exactly63 { - t.Fatalf("63-char value should be unchanged") - } - - long := "a" + exactly63 // 64 chars - got := sanitizeLabelValue(long) - if len(got) != 63 { - t.Fatalf("expected truncation to 63 chars, got len=%d", len(got)) - } -} - -func TestCountActiveDrainsControllerState(t *testing.T) { - snCR := &simplyblockv1alpha1.StorageNodeSet{ - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - DrainCoordination: []simplyblockv1alpha1.NodeDrainState{ - {Hostname: "n1", Phase: simplyblockv1alpha1.DrainPhaseShutdownCalled}, - {Hostname: "n2", Phase: simplyblockv1alpha1.DrainPhaseDraining}, - {Hostname: "n3", Phase: simplyblockv1alpha1.DrainPhaseRestartCalled}, - {Hostname: "n4", Phase: simplyblockv1alpha1.DrainPhaseComplete}, - {Hostname: "n5", Phase: simplyblockv1alpha1.DrainPhaseDetected}, - {Hostname: "n6", Phase: simplyblockv1alpha1.DrainPhaseFailed}, - }, - // No Nodes with UUIDs → no backend calls. - }, - } - - got := countActiveDrains(context.Background(), snCR, webapi.NewClient("http://127.0.0.1:1"), "cluster") - if got != 3 { - t.Fatalf("expected 3 active drains (shutdown_called, draining, restart_called), got %d", got) - } -} - -func TestCountActiveDrainsBackendConservative(t *testing.T) { - // Backend API unreachable → node is counted as active (conservative). - snCR := &simplyblockv1alpha1.StorageNodeSet{ - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - Nodes: []simplyblockv1alpha1.NodeStatus{ - {Hostname: "node-a", UUID: "uuid-a"}, - }, - }, - } - - // Use an unreachable address to force backend error. - got := countActiveDrains(context.Background(), snCR, webapi.NewClient("http://127.0.0.1:1"), "cluster") - if got < 1 { - t.Fatalf("expected at least 1 (conservative count on API error), got %d", got) - } -} - -func TestCountActiveDrainsBackendTakesPrecedence(t *testing.T) { - // Backend reports 2 in_shutdown; controller state has 0 active. - // countActiveDrains should return the backend count when it's higher. - srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) { - w.WriteHeader(http.StatusOK) - _, _ = w.Write([]byte(`{"status":"in_shutdown","health_check":false}`)) - })) - defer srv.Close() - - snCR := &simplyblockv1alpha1.StorageNodeSet{ - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - Nodes: []simplyblockv1alpha1.NodeStatus{ - {Hostname: "node-a", UUID: "uuid-a"}, - {Hostname: "node-b", UUID: "uuid-b"}, - }, - }, - } - - got := countActiveDrains(context.Background(), snCR, webapi.NewClient(srv.URL), "cluster") - if got != 2 { - t.Fatalf("expected backend count of 2, got %d", got) - } -} - -func TestActiveDrainWorkersUnionsControllerAndBackend(t *testing.T) { - // Controller state flags node-a; backend independently flags node-b (e.g. - // a real failure with no coordinator-driven drain in progress). Both must - // appear in the set so the failure-domain gate sees each affected domain. - srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) { - w.WriteHeader(http.StatusOK) - _, _ = w.Write([]byte(`{"status":"in_shutdown","health_check":false}`)) - })) - defer srv.Close() - - snCR := &simplyblockv1alpha1.StorageNodeSet{ - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - DrainCoordination: []simplyblockv1alpha1.NodeDrainState{ - {Hostname: "node-a", Phase: simplyblockv1alpha1.DrainPhaseShutdownCalled}, - }, - Nodes: []simplyblockv1alpha1.NodeStatus{ - {Hostname: "node-b", UUID: "uuid-b"}, - }, - }, - } - - got := activeDrainWorkers(context.Background(), snCR, webapi.NewClient(srv.URL), "cluster") - if !got["node-a"] || !got["node-b"] || len(got) != 2 { - t.Fatalf("expected {node-a, node-b}, got %v", got) - } -} - -func TestActiveDrainWorkersConservativeOnBackendError(t *testing.T) { - snCR := &simplyblockv1alpha1.StorageNodeSet{ - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - Nodes: []simplyblockv1alpha1.NodeStatus{ - {Hostname: "node-a", UUID: "uuid-a"}, - }, - }, - } - - got := activeDrainWorkers(context.Background(), snCR, webapi.NewClient("http://127.0.0.1:1"), "cluster") - if !got["node-a"] { - t.Fatalf("expected node-a marked active conservatively on API error, got %v", got) - } -} - -func TestDistinctDomainCount(t *testing.T) { - fd1, fd2, fd3 := int32(1), int32(2), int32(3) - snCR := &simplyblockv1alpha1.StorageNodeSet{ - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - Nodes: []simplyblockv1alpha1.NodeStatus{ - {Hostname: "a", FailureDomain: &fd1}, - {Hostname: "b", FailureDomain: &fd1}, - {Hostname: "c", FailureDomain: &fd2}, - {Hostname: "d", FailureDomain: &fd3}, - {Hostname: "e"}, // unassigned -> excluded - }, - }, - } - - if got := distinctDomainCount(snCR); got != 3 { - t.Fatalf("expected 3 distinct domains, got %d", got) - } - if got := distinctDomainCount(&simplyblockv1alpha1.StorageNodeSet{}); got != 0 { - t.Fatalf("expected 0 for empty status, got %d", got) - } -} - -func TestActiveDrainDomainCountsTallies(t *testing.T) { - fd1, fd2 := int32(1), int32(2) - snCR := &simplyblockv1alpha1.StorageNodeSet{ - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - Nodes: []simplyblockv1alpha1.NodeStatus{ - {Hostname: "node-a", FailureDomain: &fd1}, - {Hostname: "node-b", FailureDomain: &fd1}, - {Hostname: "node-c", FailureDomain: &fd2}, - // node-d intentionally has no status entry -> excluded - }, - }, - } - activeWorkers := map[string]bool{"node-a": true, "node-b": true, "node-c": true, "node-d": true} - counts := activeDrainDomainCounts(snCR, activeWorkers) - if counts[1] != 2 || counts[2] != 1 || len(counts) != 2 { - t.Fatalf("expected {1:2, 2:1}, got %v", counts) - } -} - -// fdDrainGate is confirmed against the backend team's stated requirements -// (2026-08, Dmitrii Iakovlev) for a 2+2 layout: -// - 2 domains: "1 whole domain" or "1 node in each of the 2 domains" are -// the only safe combinations. -// - 3 domains: only 1 domain may be fully down, not 2. -// - 4 domains: 2 domains may be fully down (well-provisioned case). - -func TestFdDrainGate2DomainsOneWholeDomainSafe(t *testing.T) { - // chunksPerDomain = ceil(4/2) = 2. Domain 1 already has 3 nodes down - // (maxed at chunksPerDomain regardless of how many more) -- a 4th is free. - counts := map[int32]int{1: 3} - if blocked, _ := fdDrainGate(counts, 1, 2, 4, 2); blocked { - t.Fatalf("expected piling further within an already-maxed domain to proceed") - } -} - -func TestFdDrainGate2DomainsOneNodePerDomainSafe(t *testing.T) { - counts := map[int32]int{1: 1} - if blocked, _ := fdDrainGate(counts, 2, 2, 4, 2); blocked { - t.Fatalf("expected 1 node in each of the 2 domains to proceed") - } -} - -func TestFdDrainGate2DomainsWholeDomainPlusOtherIsUnsafe(t *testing.T) { - // Domain 1 fully down (4 nodes, maxed at chunksPerDomain=2) -- a node in - // domain 2 must now be blocked (not one of the two safe combinations). - counts := map[int32]int{1: 4} - blocked, reason := fdDrainGate(counts, 2, 2, 4, 2) - if !blocked || reason == "" { - t.Fatalf("expected domain 2 to be blocked once domain 1 is fully down, got blocked=%v reason=%q", blocked, reason) - } -} - -func TestFdDrainGate2DomainsOnePerDomainPlusExtraInEitherIsUnsafe(t *testing.T) { - // 1 node down in each of domains 1 and 2 already (the safe combo) -- - // piling a SECOND node onto EITHER domain must now be blocked, even - // though that domain is already "active". This is the exact gap the old - // unconditional-piling logic missed. - counts := map[int32]int{1: 1, 2: 1} - if blocked, _ := fdDrainGate(counts, 2, 2, 4, 2); !blocked { - t.Fatalf("expected a 2nd node in domain 2 to be blocked once 1+1 is already committed") - } - if blocked, _ := fdDrainGate(counts, 1, 2, 4, 2); !blocked { - t.Fatalf("expected a 2nd node in domain 1 to be blocked once 1+1 is already committed") - } -} - -func TestFdDrainGate3DomainsOneWholeDomainSafeTwoUnsafe(t *testing.T) { - // chunksPerDomain = ceil(4/3) = 2, same per-domain cap as 2 domains, but - // spread across 3. 1 domain fully down is safe; opening a 2nd is not. - counts := map[int32]int{1: 2} // domain 1 already maxed - if blocked, _ := fdDrainGate(counts, 1, 3, 4, 2); blocked { - t.Fatalf("expected piling within the already-maxed domain 1 to proceed") - } - if blocked, _ := fdDrainGate(counts, 2, 3, 4, 2); !blocked { - t.Fatalf("expected opening domain 2 to be blocked once domain 1 has maxed the risk budget") - } -} - -func TestFdDrainGate4DomainsTwoWholeDomainsSafeThreeUnsafe(t *testing.T) { - // chunksPerDomain = ceil(4/4) = 1 (well-provisioned): up to npcs=2 whole - // domains may be fully down. - counts := map[int32]int{1: 1} // domain 1 already maxed (chunksPerDomain=1) - if blocked, _ := fdDrainGate(counts, 2, 4, 4, 2); blocked { - t.Fatalf("expected opening a 2nd domain to proceed while under the npcs=2 domain budget") - } - countsTwoActive := map[int32]int{1: 1, 2: 1} - if blocked, _ := fdDrainGate(countsTwoActive, 3, 4, 4, 2); !blocked { - t.Fatalf("expected opening a 3rd domain to be blocked once 2 domains already max the npcs=2 budget") - } -} - -func TestFdDrainGateUnknownSchemeTreatedAsSingleDomainChunk(t *testing.T) { - // domainsNeededForFullDisjoint=0 signals the scheme could not be parsed; - // chunksPerDomain must fall back to >= 1, not 0 (which would divide by - // zero / always-block via a degenerate cap). - counts := map[int32]int{1: 1} - blocked, _ := fdDrainGate(counts, 2, 5, 0, 2) - if blocked { - t.Fatalf("expected an unparsed scheme to still allow a 2nd domain within the npcs budget, got blocked") - } -} - -// ---- reconciler tests ---- - -func TestNodeDrainReconcileNotFound(t *testing.T) { - r := newNodeDrainTestReconciler(t) - - res, err := r.Reconcile(context.Background(), ctrl.Request{ - NamespacedName: client.ObjectKey{Name: "missing", Namespace: "default"}, - }) - if err != nil { - t.Fatalf("expected no error for missing CR, got %v", err) - } - if res.RequeueAfter != 0 { - t.Fatalf("expected no requeue for missing CR, got %+v", res) - } -} - -func TestNodeDrainReconcileNoClusterAuthRequeues(t *testing.T) { - snCR := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-no-auth", Namespace: "default"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - ClusterName: "cluster-missing", - }, - } - r := newNodeDrainTestReconciler(t, snCR) - - res, err := r.Reconcile(context.Background(), ctrl.Request{NamespacedName: client.ObjectKeyFromObject(snCR)}) - if err != nil { - t.Fatalf("expected no error when cluster auth unavailable, got %v", err) - } - if res.RequeueAfter == 0 { - t.Fatalf("expected delayed requeue when cluster auth is unavailable") - } -} - -func TestNodeDrainReconcileNoClusterCRRequeues(t *testing.T) { - const clusterName = "cluster-no-cr" - - snCR := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-no-cluster-cr", Namespace: "default"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - ClusterName: clusterName, - }, - } - r := newNodeDrainTestReconciler(t, snCR) - - res, err := r.Reconcile(context.Background(), ctrl.Request{NamespacedName: client.ObjectKeyFromObject(snCR)}) - if err != nil { - t.Fatalf("expected no error when cluster CR unavailable, got %v", err) - } - if res.RequeueAfter == 0 { - t.Fatalf("expected delayed requeue when cluster CR is missing") - } -} - -func TestNodeDrainReconcileSkipsWhenClusterUnready(t *testing.T) { - const clusterName = "cluster-unready" - - clusterCR := &simplyblockv1alpha2.StorageCluster{ - ObjectMeta: metav1.ObjectMeta{Name: clusterName, Namespace: "default"}, - Status: simplyblockv1alpha2.StorageClusterStatus{ - Status: utils.ClusterStatusUnready, - }, - } - snCR := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-unready", Namespace: "default"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - ClusterName: clusterName, - }, - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - Nodes: []simplyblockv1alpha1.NodeStatus{ - {UUID: "node-1", Status: utils.NodeStatusOnline, Hostname: "worker-1"}, - }, - }, - } - r := newNodeDrainTestReconciler(t, snCR, clusterCR) - - res, err := r.Reconcile(context.Background(), ctrl.Request{NamespacedName: client.ObjectKeyFromObject(snCR)}) - if err != nil { - t.Fatalf("expected no error when cluster is unready, got %v", err) - } - if res.RequeueAfter == 0 { - t.Fatal("expected a delayed requeue when cluster is unready") - } - - // No PDBs should have been created — drain coordinator must not run while cluster is unready. - var pdbList policyv1.PodDisruptionBudgetList - if err := r.List(context.Background(), &pdbList); err != nil { - t.Fatalf("list PDBs: %v", err) - } - if len(pdbList.Items) != 0 { - t.Errorf("expected no PDBs when cluster is unready, got %d", len(pdbList.Items)) - } -} - -func TestEnsurePDBCreatesWhenMissing(t *testing.T) { - r := newNodeDrainTestReconciler(t) - - if err := r.ensurePDB(context.Background(), "default", "node-a", 0); err != nil { - t.Fatalf("ensurePDB returned error: %v", err) - } - - pdb := &policyv1.PodDisruptionBudget{} - if err := r.Get(context.Background(), client.ObjectKey{ - Name: drainPDBPrefix + "node-a", - Namespace: "default", - }, pdb); err != nil { - t.Fatalf("PDB should have been created: %v", err) - } - if pdb.Spec.MaxUnavailable == nil || pdb.Spec.MaxUnavailable.IntValue() != 0 { - t.Fatalf("expected maxUnavailable=0, got %v", pdb.Spec.MaxUnavailable) - } -} - -func TestEnsurePDBUpdatesExisting(t *testing.T) { - maxUnavailable := intstr.FromInt32(0) - existing := &policyv1.PodDisruptionBudget{ - ObjectMeta: metav1.ObjectMeta{ - Name: drainPDBPrefix + "node-b", - Namespace: "default", - }, - Spec: policyv1.PodDisruptionBudgetSpec{ - MaxUnavailable: &maxUnavailable, - Selector: &metav1.LabelSelector{ - MatchLabels: map[string]string{drainNodeLabelKey: "node-b"}, - }, - }, - } - r := newNodeDrainTestReconciler(t, existing) - - if err := r.ensurePDB(context.Background(), "default", "node-b", 1); err != nil { - t.Fatalf("ensurePDB update returned error: %v", err) - } - - pdb := &policyv1.PodDisruptionBudget{} - if err := r.Get(context.Background(), client.ObjectKey{ - Name: drainPDBPrefix + "node-b", - Namespace: "default", - }, pdb); err != nil { - t.Fatalf("failed to fetch PDB: %v", err) - } - if pdb.Spec.MaxUnavailable == nil || pdb.Spec.MaxUnavailable.IntValue() != 1 { - t.Fatalf("expected maxUnavailable=1 after update, got %v", pdb.Spec.MaxUnavailable) - } -} - -func TestCleanupPDBDeletesWhenPresent(t *testing.T) { - maxUnavailable := intstr.FromInt32(0) - pdb := &policyv1.PodDisruptionBudget{ - ObjectMeta: metav1.ObjectMeta{ - Name: drainPDBPrefix + "node-c", - Namespace: "default", - }, - Spec: policyv1.PodDisruptionBudgetSpec{ - MaxUnavailable: &maxUnavailable, - Selector: &metav1.LabelSelector{MatchLabels: map[string]string{drainNodeLabelKey: "node-c"}}, - }, - } - r := newNodeDrainTestReconciler(t, pdb) - - if err := r.cleanupPDB(context.Background(), "default", "node-c"); err != nil { - t.Fatalf("cleanupPDB returned error: %v", err) - } - - out := &policyv1.PodDisruptionBudget{} - err := r.Get(context.Background(), client.ObjectKey{Name: drainPDBPrefix + "node-c", Namespace: "default"}, out) - if err == nil { - t.Fatalf("expected PDB to be deleted") - } -} - -func TestCleanupPDBNoopWhenMissing(t *testing.T) { - r := newNodeDrainTestReconciler(t) - if err := r.cleanupPDB(context.Background(), "default", "node-missing"); err != nil { - t.Fatalf("cleanupPDB should be no-op for missing PDB, got error: %v", err) - } -} - -func TestLabelStoragePodLabelsMatchingPod(t *testing.T) { - pod := &corev1.Pod{ - ObjectMeta: metav1.ObjectMeta{ - Name: "spdk-pod", - Namespace: "default", - Labels: map[string]string{"role": "simplyblock-storage-node"}, - }, - Spec: corev1.PodSpec{NodeName: "node-d"}, - } - r := newNodeDrainTestReconciler(t, pod) - - if err := r.labelStoragePod(context.Background(), "default", "node-d"); err != nil { - t.Fatalf("labelStoragePod returned error: %v", err) - } - - out := &corev1.Pod{} - if err := r.Get(context.Background(), client.ObjectKeyFromObject(pod), out); err != nil { - t.Fatalf("failed to fetch pod: %v", err) - } - if out.Labels[drainNodeLabelKey] != sanitizeLabelValue("node-d") { - t.Fatalf("expected drain label to be set, got %q", out.Labels[drainNodeLabelKey]) - } -} - -func TestLabelStoragePodSkipsDifferentNode(t *testing.T) { - pod := &corev1.Pod{ - ObjectMeta: metav1.ObjectMeta{ - Name: "spdk-pod-other", - Namespace: "default", - Labels: map[string]string{"role": "simplyblock-storage-node"}, - }, - Spec: corev1.PodSpec{NodeName: "node-other"}, - } - r := newNodeDrainTestReconciler(t, pod) - - if err := r.labelStoragePod(context.Background(), "default", "node-target"); err != nil { - t.Fatalf("labelStoragePod returned error: %v", err) - } - - out := &corev1.Pod{} - if err := r.Get(context.Background(), client.ObjectKeyFromObject(pod), out); err != nil { - t.Fatalf("failed to fetch pod: %v", err) - } - if _, ok := out.Labels[drainNodeLabelKey]; ok { - t.Fatalf("expected drain label NOT to be set on pod from different node") - } -} - -func TestLabelStoragePodIdempotent(t *testing.T) { - nodeName := "node-e" - pod := &corev1.Pod{ - ObjectMeta: metav1.ObjectMeta{ - Name: "spdk-pod-idempotent", - Namespace: "default", - Labels: map[string]string{ - "role": "simplyblock-storage-node", - drainNodeLabelKey: sanitizeLabelValue(nodeName), - }, - }, - Spec: corev1.PodSpec{NodeName: nodeName}, - } - r := newNodeDrainTestReconciler(t, pod) - - // Should succeed without error (patch is skipped for already-labeled pods). - if err := r.labelStoragePod(context.Background(), "default", nodeName); err != nil { - t.Fatalf("labelStoragePod returned error on idempotent call: %v", err) - } -} - -func TestCleanupDrainResources(t *testing.T) { - nodeName := "node-f" - maxUnavailable := intstr.FromInt32(0) - pdb := &policyv1.PodDisruptionBudget{ - ObjectMeta: metav1.ObjectMeta{ - Name: drainPDBPrefix + nodeName, - Namespace: "default", - }, - Spec: policyv1.PodDisruptionBudgetSpec{ - MaxUnavailable: &maxUnavailable, - Selector: &metav1.LabelSelector{MatchLabels: map[string]string{drainNodeLabelKey: nodeName}}, - }, - } - pod := &corev1.Pod{ - ObjectMeta: metav1.ObjectMeta{ - Name: "spdk-pod-cleanup", - Namespace: "default", - Labels: map[string]string{ - drainNodeLabelKey: sanitizeLabelValue(nodeName), - }, - }, - } - r := newNodeDrainTestReconciler(t, pdb, pod) - - if err := r.cleanupDrainResources(context.Background(), "default", nodeName); err != nil { - t.Fatalf("cleanupDrainResources returned error: %v", err) - } - - // PDB should be gone. - out := &policyv1.PodDisruptionBudget{} - if err := r.Get(context.Background(), client.ObjectKey{Name: drainPDBPrefix + nodeName, Namespace: "default"}, out); err == nil { - t.Fatalf("expected PDB to be deleted") - } - - // Drain label should be removed from the pod. - outPod := &corev1.Pod{} - if err := r.Get(context.Background(), client.ObjectKeyFromObject(pod), outPod); err != nil { - t.Fatalf("failed to fetch pod: %v", err) - } - if _, ok := outPod.Labels[drainNodeLabelKey]; ok { - t.Fatalf("expected drain label to be removed from pod") - } -} - -func TestEnsureManagerPDBCreates(t *testing.T) { - r := newNodeDrainTestReconciler(t) - - if err := r.ensureManagerPDB(context.Background(), "default"); err != nil { - t.Fatalf("ensureManagerPDB returned error: %v", err) - } - - pdb := &policyv1.PodDisruptionBudget{} - if err := r.Get(context.Background(), client.ObjectKey{Name: managerPDBName, Namespace: "default"}, pdb); err != nil { - t.Fatalf("manager PDB should have been created: %v", err) - } - if pdb.Spec.MaxUnavailable == nil || pdb.Spec.MaxUnavailable.IntValue() != 0 { - t.Fatalf("expected manager PDB maxUnavailable=0") - } -} - -func TestDeleteManagerPDBDeletesWhenPresent(t *testing.T) { - maxUnavailable := intstr.FromInt32(0) - pdb := &policyv1.PodDisruptionBudget{ - ObjectMeta: metav1.ObjectMeta{Name: managerPDBName, Namespace: "default"}, - Spec: policyv1.PodDisruptionBudgetSpec{ - MaxUnavailable: &maxUnavailable, - Selector: &metav1.LabelSelector{MatchLabels: map[string]string{"app": "simplyblock-operator"}}, - }, - } - r := newNodeDrainTestReconciler(t, pdb) - - if err := r.deleteManagerPDB(context.Background(), "default"); err != nil { - t.Fatalf("deleteManagerPDB returned error: %v", err) - } - - out := &policyv1.PodDisruptionBudget{} - if err := r.Get(context.Background(), client.ObjectKey{Name: managerPDBName, Namespace: "default"}, out); err == nil { - t.Fatalf("expected manager PDB to be deleted") - } -} - -func TestDeleteManagerPDBNoopWhenMissing(t *testing.T) { - r := newNodeDrainTestReconciler(t) - if err := r.deleteManagerPDB(context.Background(), "default"); err != nil { - t.Fatalf("deleteManagerPDB should be no-op when missing, got error: %v", err) - } -} - -func TestCleanupManagerPDBIfStaleRemovesWhenNotInDetectedPhase(t *testing.T) { - maxUnavailable := intstr.FromInt32(0) - pdb := &policyv1.PodDisruptionBudget{ - ObjectMeta: metav1.ObjectMeta{Name: managerPDBName, Namespace: "default"}, - Spec: policyv1.PodDisruptionBudgetSpec{ - MaxUnavailable: &maxUnavailable, - Selector: &metav1.LabelSelector{MatchLabels: map[string]string{"app": "simplyblock-operator"}}, - }, - } - // Manager node is NOT in detected phase (no drain state at all). - snCR := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Namespace: "default"}, - } - r := newNodeDrainTestReconciler(t, pdb) - r.ManagerNodeName = "manager-node" - - r.cleanupManagerPDBIfStale(context.Background(), snCR) - - out := &policyv1.PodDisruptionBudget{} - if err := r.Get(context.Background(), client.ObjectKey{Name: managerPDBName, Namespace: "default"}, out); err == nil { - t.Fatalf("expected stale manager PDB to be deleted") - } -} - -func TestCleanupManagerPDBIfStaleKeepsWhenDetected(t *testing.T) { - maxUnavailable := intstr.FromInt32(0) - pdb := &policyv1.PodDisruptionBudget{ - ObjectMeta: metav1.ObjectMeta{Name: managerPDBName, Namespace: "default"}, - Spec: policyv1.PodDisruptionBudgetSpec{ - MaxUnavailable: &maxUnavailable, - Selector: &metav1.LabelSelector{MatchLabels: map[string]string{"app": "simplyblock-operator"}}, - }, - } - // Manager node IS in the detected phase — PDB should be kept. - snCR := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Namespace: "default"}, - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - DrainCoordination: []simplyblockv1alpha1.NodeDrainState{ - {Hostname: "manager-node", Phase: simplyblockv1alpha1.DrainPhaseDetected}, - }, - }, - } - r := newNodeDrainTestReconciler(t, pdb) - r.ManagerNodeName = "manager-node" - - r.cleanupManagerPDBIfStale(context.Background(), snCR) - - out := &policyv1.PodDisruptionBudget{} - if err := r.Get(context.Background(), client.ObjectKey{Name: managerPDBName, Namespace: "default"}, out); err != nil { - t.Fatalf("expected manager PDB to be kept during detected phase: %v", err) - } -} - -func TestProcessWorkerUncordonedNoState(t *testing.T) { - node := &corev1.Node{ - ObjectMeta: metav1.ObjectMeta{Name: "node-g"}, - Spec: corev1.NodeSpec{Unschedulable: false}, - } - snCR := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn", Namespace: "default"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{WorkerNodes: []string{"node-g"}}, - } - r := newNodeDrainTestReconciler(t, snCR, node) - - requeue, shouldBreak := r.processWorker( - context.Background(), snCR, "node-g", - webapi.NewClient("http://127.0.0.1:1"), "cluster", 1, false, 0, 0, - ) - if requeue != 0 || shouldBreak { - t.Fatalf("expected (0, false) for uncordoned node with no state, got (%v, %v)", requeue, shouldBreak) - } - if getDrainState(snCR, "node-g") != nil { - t.Fatalf("expected no drain state to be created") - } -} - -func TestProcessWorkerSkipsCordonedNotYetOnline(t *testing.T) { - node := &corev1.Node{ - ObjectMeta: metav1.ObjectMeta{Name: "node-h"}, - Spec: corev1.NodeSpec{Unschedulable: true}, - } - // No Nodes in status → isWorkerOnline returns false. - snCR := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-h", Namespace: "default"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{WorkerNodes: []string{"node-h"}}, - } - r := newNodeDrainTestReconciler(t, snCR, node) - - requeue, shouldBreak := r.processWorker( - context.Background(), snCR, "node-h", - webapi.NewClient("http://127.0.0.1:1"), "cluster", 1, false, 0, 0, - ) - if requeue != 0 || shouldBreak { - t.Fatalf("expected (0, false) for cordoned node not yet online, got (%v, %v)", requeue, shouldBreak) - } - if getDrainState(snCR, "node-h") != nil { - t.Fatalf("expected no drain state created for node that was never online") - } -} - -func TestProcessWorkerCordonedOnlineInitializesState(t *testing.T) { - node := &corev1.Node{ - ObjectMeta: metav1.ObjectMeta{Name: "node-i"}, - Spec: corev1.NodeSpec{Unschedulable: true}, - } - // Node is online in backend status. - snCR := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-i", Namespace: "default"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{WorkerNodes: []string{"node-i"}}, - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - Nodes: []simplyblockv1alpha1.NodeStatus{ - {Hostname: "node-i", Status: "online", UUID: ""}, - }, - }, - } - r := newNodeDrainTestReconciler(t, snCR, node) - - r.processWorker( - context.Background(), snCR, "node-i", - webapi.NewClient("http://127.0.0.1:1"), "cluster", 1, false, 0, 0, - ) - - // Drain state must have been initialized (phase may be detected or failed - // depending on backend reachability, but the entry must exist). - if getDrainState(snCR, "node-i") == nil { - t.Fatalf("expected drain state to be created for cordoned online node") - } -} - -func TestIsClusterRebalancingTrue(t *testing.T) { - srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) { - w.WriteHeader(http.StatusOK) - _, _ = w.Write([]byte(`{"is_re_balancing":true}`)) - })) - defer srv.Close() - - rebalancing, err := isClusterRebalancing(context.Background(), webapi.NewClient(srv.URL), "cluster-uuid") - if err != nil { - t.Fatalf("unexpected error: %v", err) - } - if !rebalancing { - t.Fatalf("expected rebalancing=true, got false") - } -} - -func TestIsClusterRebalancingFalse(t *testing.T) { - srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) { - w.WriteHeader(http.StatusOK) - _, _ = w.Write([]byte(`{"is_re_balancing":false}`)) - })) - defer srv.Close() - - rebalancing, err := isClusterRebalancing(context.Background(), webapi.NewClient(srv.URL), "cluster-uuid") - if err != nil { - t.Fatalf("unexpected error: %v", err) - } - if rebalancing { - t.Fatalf("expected rebalancing=false, got true") - } -} - -func TestIsClusterRebalancingAPIError(t *testing.T) { - // Unreachable address → error expected. - _, err := isClusterRebalancing(context.Background(), webapi.NewClient("http://127.0.0.1:1"), "cluster-uuid") - if err == nil { - t.Fatalf("expected error when API is unreachable") - } -} - -func TestIsClusterRebalancingNonOKStatus(t *testing.T) { - srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) { - w.WriteHeader(http.StatusInternalServerError) - _, _ = w.Write([]byte(`internal error`)) - })) - defer srv.Close() - - _, err := isClusterRebalancing(context.Background(), webapi.NewClient(srv.URL), "cluster-uuid") - if err == nil { - t.Fatalf("expected error on non-2xx response") - } -} - -func TestHandleRestartCalledHoldsDrainSlotWhileRebalancing(t *testing.T) { - // Scenario: all socket nodes are online+healthy but cluster is still - // rebalancing. handleRestartCalled must NOT mark phase complete — it should - // requeue and keep the message about rebalancing. - const nodeName = "node-rebal" - const nodeUUID = "uuid-rebal" - - // Backend: node is online and healthy; cluster is rebalancing. - callCount := 0 - srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { - w.WriteHeader(http.StatusOK) - callCount++ - // Node info endpoint returns online+healthy. - // Cluster info endpoint returns rebalancing=true. - if r.URL.Path == "/api/v2/clusters/cluster-uuid" && r.Method == http.MethodGet { - _, _ = w.Write([]byte(`{"is_re_balancing":true}`)) - } else { - _, _ = w.Write([]byte(`{"status":"online","health_check":true}`)) - } - })) - defer srv.Close() - - k8sNode := &corev1.Node{ - ObjectMeta: metav1.ObjectMeta{Name: nodeName}, - Status: corev1.NodeStatus{ - Conditions: []corev1.NodeCondition{ - {Type: corev1.NodeReady, Status: corev1.ConditionTrue}, - }, - }, - } - snCR := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-rebal", Namespace: "default"}, - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - Nodes: []simplyblockv1alpha1.NodeStatus{ - {Hostname: nodeName, UUID: nodeUUID, Status: "online"}, - }, - }, - } - state := &simplyblockv1alpha1.NodeDrainState{ - Hostname: nodeName, - Phase: simplyblockv1alpha1.DrainPhaseRestartCalled, - ActiveNodeUUID: nodeUUID, - } - - r := newNodeDrainTestReconciler(t, snCR, k8sNode) - requeue, err := r.handleRestartCalled(context.Background(), snCR, state, webapi.NewClient(srv.URL), "cluster-uuid") - - if err != nil { - t.Fatalf("unexpected error: %v", err) - } - if requeue == 0 { - t.Fatalf("expected non-zero requeue while cluster is rebalancing") - } - if state.Phase == simplyblockv1alpha1.DrainPhaseComplete { - t.Fatalf("drain phase must NOT be complete while cluster is rebalancing") - } - if state.Message == "" { - t.Fatalf("expected a status message about rebalancing") - } -} - -func TestHandleRestartCalledCompletesWhenNotRebalancing(t *testing.T) { - // Scenario: all socket nodes are online+healthy and cluster is NOT - // rebalancing. handleRestartCalled should mark phase complete. - const nodeName = "node-done" - const nodeUUID = "uuid-done" - - srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { - w.WriteHeader(http.StatusOK) - if r.URL.Path == "/api/v2/clusters/cluster-uuid" && r.Method == http.MethodGet { - _, _ = w.Write([]byte(`{"is_re_balancing":false}`)) - } else { - _, _ = w.Write([]byte(`{"status":"online","health_check":true}`)) - } - })) - defer srv.Close() - - k8sNode := &corev1.Node{ - ObjectMeta: metav1.ObjectMeta{Name: nodeName}, - Status: corev1.NodeStatus{ - Conditions: []corev1.NodeCondition{ - {Type: corev1.NodeReady, Status: corev1.ConditionTrue}, - }, - }, - } - snCR := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-done", Namespace: "default"}, - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - Nodes: []simplyblockv1alpha1.NodeStatus{ - {Hostname: nodeName, UUID: nodeUUID, Status: "online"}, - }, - }, - } - state := &simplyblockv1alpha1.NodeDrainState{ - Hostname: nodeName, - Phase: simplyblockv1alpha1.DrainPhaseRestartCalled, - ActiveNodeUUID: nodeUUID, - } - - r := newNodeDrainTestReconciler(t, snCR, k8sNode) - requeue, err := r.handleRestartCalled(context.Background(), snCR, state, webapi.NewClient(srv.URL), "cluster-uuid") - - if err != nil { - t.Fatalf("unexpected error: %v", err) - } - if state.Phase != simplyblockv1alpha1.DrainPhaseComplete { - t.Fatalf("expected DrainPhaseComplete when node is online+healthy and cluster not rebalancing, got %q", state.Phase) - } - if requeue != 0 { - t.Fatalf("expected zero requeue on completion, got %v", requeue) - } -} - -// ---- 409 conflict retry test ---- - -// TestNodeDrainStatusPatch409RetryPreservesDrainState verifies that a 409 -// Conflict on the final Status().Patch() does NOT discard the drain phase -// transitions computed during the reconcile. Without RetryOnConflict the -// controller would silently revert to the pre-reconcile state. -// -// Setup: one worker already in DrainPhaseComplete (no backend HTTP calls -// needed) so processWorker is a pure no-op. The interesting behaviour is in -// the final patch: the interceptor returns 409 on the first attempt and -// succeeds on the second, verifying that RetryOnConflict re-reads and retries -// rather than logging and returning the 5-second requeue. -func TestNodeDrainStatusPatch409RetryPreservesDrainState(t *testing.T) { - const ( - ns = "default" - clusterName = "cluster-drain-409" - clusterUUID = "uuid-drain-409" - workerName = "worker-409.example.com" - snsName = "sn-drain-409" - ) - - clusterCR := &simplyblockv1alpha2.StorageCluster{ - ObjectMeta: metav1.ObjectMeta{Name: clusterName, Namespace: ns}, - Status: simplyblockv1alpha2.StorageClusterStatus{ - Status: utils.ClusterStatusActive, - UUID: clusterUUID, - }, - } - snCR := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: snsName, Namespace: ns}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - ClusterName: clusterName, - WorkerNodes: []string{workerName}, - }, - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - Nodes: []simplyblockv1alpha1.NodeStatus{ - {Hostname: workerName, Status: utils.NodeStatusOnline, UUID: "backend-uuid-409"}, - }, - // DrainPhaseComplete → processWorker is a pure no-op (no backend calls). - // The patch carries this state; the interceptor will conflict on the - // first attempt and succeed on the second. - DrainCoordination: []simplyblockv1alpha1.NodeDrainState{ - { - Hostname: workerName, - Phase: simplyblockv1alpha1.DrainPhaseComplete, - }, - }, - }, - } - - scheme := newTestScheme( - t, - simplyblockv1alpha1.AddToScheme, - corev1.AddToScheme, - policyv1.AddToScheme, - ) - - patchCalls := 0 - conflictErr := apierrors.NewConflict( - simplyblockv1alpha1.GroupVersion.WithResource("storagenodesets").GroupResource(), - snsName, nil, - ) - - cl := fake.NewClientBuilder(). - WithScheme(scheme). - WithStatusSubresource(&simplyblockv1alpha1.StorageNodeSet{}, &simplyblockv1alpha2.StorageCluster{}). - WithObjects(clusterCR, snCR). - WithInterceptorFuncs(interceptor.Funcs{ - SubResourcePatch: func( - ctx context.Context, - c client.Client, - subResourceName string, - obj client.Object, - patch client.Patch, - opts ...client.SubResourcePatchOption, - ) error { - if subResourceName == "status" { - patchCalls++ - if patchCalls == 1 { - return conflictErr - } - } - return c.Status().Patch(ctx, obj, patch, opts...) - }, - }). - Build() - - r := &NodeDrainCoordinatorReconciler{Client: cl, Scheme: scheme} - res, err := r.Reconcile(context.Background(), ctrl.Request{ - NamespacedName: client.ObjectKey{Name: snsName, Namespace: ns}, - }) - - if err != nil { - t.Fatalf("expected no error despite initial 409, got: %v", err) - } - // RetryOnConflict should succeed — must NOT return the 5-second requeue - // that the old bare-patch code returned on any error. - if res.RequeueAfter == 5*time.Second { - t.Fatal("reconcile returned the 5s conflict requeue — RetryOnConflict did not succeed") - } - if patchCalls < 2 { - t.Fatalf("expected ≥2 Status.Patch calls (conflict + retry), got %d", patchCalls) - } - - // Drain state must be persisted after the conflict retry. - var updated simplyblockv1alpha1.StorageNodeSet - if err := cl.Get(context.Background(), client.ObjectKey{Name: snsName, Namespace: ns}, &updated); err != nil { - t.Fatalf("failed to get updated CR: %v", err) - } - state := getDrainState(&updated, workerName) - if state == nil { - t.Fatal("drain state was lost after 409 — RetryOnConflict did not preserve the phase") - return - } - if state.Phase != simplyblockv1alpha1.DrainPhaseComplete { - t.Fatalf("expected DrainPhaseComplete after retry, got %q", state.Phase) - } -} - -// ---- failure-domain gate tests ---- - -// snCRWithFD builds a minimal StorageNodeSet with two nodes whose failure domains -// are set in status.nodes (populated from the backend API), and node-a already -// in an active drain phase. -func snCRWithFD(nodeADomain, nodeBDomain int32) *simplyblockv1alpha1.StorageNodeSet { - return &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-fd", Namespace: "default"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - WorkerNodes: []string{"node-a", "node-b"}, - }, - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - // FailureDomain sourced from backend API response, stored in status. - // No UUIDs so activeDrainWorkers makes no backend API calls. - Nodes: []simplyblockv1alpha1.NodeStatus{ - {Hostname: "node-a", FailureDomain: &nodeADomain}, - {Hostname: "node-b", FailureDomain: &nodeBDomain}, - }, - DrainCoordination: []simplyblockv1alpha1.NodeDrainState{ - {Hostname: "node-a", Phase: simplyblockv1alpha1.DrainPhaseShutdownCalled}, - }, - }, - } -} - -// TestHandleDetectedFDDisabledUsesGlobalGate verifies that when fdEnabled=false -// the existing node-count gate blocks a second drain when activeDrains >= maxFaultTolerance. -func TestHandleDetectedFDDisabledUsesGlobalGate(t *testing.T) { - snCR := snCRWithFD(1, 2) - state := &simplyblockv1alpha1.NodeDrainState{Hostname: "node-b", Phase: simplyblockv1alpha1.DrainPhaseDetected} - r := newNodeDrainTestReconciler(t, snCR) - - requeue, err := r.handleDetected( - context.Background(), snCR, state, - webapi.NewClient("http://127.0.0.1:1"), "cluster", - 1, // maxFaultTolerance=1 → node-a already consumes the slot - false, // fdEnabled=false → global gate - 0, // domainsNeededForFullDisjoint unused when fdEnabled=false - 0, // npcs unused when fdEnabled=false - ) - - if err != nil { - t.Fatalf("expected nil error when blocked, got %v", err) - } - if requeue != 10*time.Second { - t.Fatalf("expected 10s requeue when blocked by global gate, got %v", requeue) - } - if state.Phase == simplyblockv1alpha1.DrainPhaseShutdownCalled { - t.Fatalf("phase must not advance while drain slot is unavailable") - } -} - -// TestHandleDetectedSameDomainParallelAllowed verifies that when fdEnabled=true, -// a worker in the same failure domain as an already-draining worker is allowed -// past the gate without waiting. -func TestHandleDetectedSameDomainParallelAllowed(t *testing.T) { - // Both node-a and node-b are in failure domain 1. - snCR := snCRWithFD(1, 1) - state := &simplyblockv1alpha1.NodeDrainState{Hostname: "node-b", Phase: simplyblockv1alpha1.DrainPhaseDetected} - r := newNodeDrainTestReconciler(t, snCR) - - requeue, err := r.handleDetected( - context.Background(), snCR, state, - webapi.NewClient("http://127.0.0.1:1"), "cluster", - 1, // maxFaultTolerance=1 — unused inside the fdEnabled branch - true, // fdEnabled=true - 1, // domainsNeededForFullDisjoint=1, domainsAvailable=1 -> chunksPerDomain=1, - // so node-a (already active in domain 1) has already maxed the domain's - // contribution -- node-b proceeds unconditionally regardless of npcs. - 1, // npcs -- irrelevant here since the maxed-domain shortcut fires first - ) - - // The gate passes; the function proceeds until it finds no UUID for node-b - // and returns a 15s requeue with a non-nil error (UUID missing). That proves - // it was NOT held back at the drain-slot check. - blocked := err == nil && requeue == 10*time.Second - if blocked { - t.Fatalf("node in the same failure domain must not be blocked; state.Message=%q", state.Message) - } -} - -// TestHandleDetectedCrossDomainGated verifies that when fdEnabled=true, a worker -// in a different failure domain from the already-draining worker is blocked when -// the active domain count meets maxFaultTolerance. -func TestHandleDetectedCrossDomainGated(t *testing.T) { - // node-a is in domain 1 (already draining); node-b is in domain 2. - snCR := snCRWithFD(1, 2) - state := &simplyblockv1alpha1.NodeDrainState{Hostname: "node-b", Phase: simplyblockv1alpha1.DrainPhaseDetected} - r := newNodeDrainTestReconciler(t, snCR) - - requeue, err := r.handleDetected( - context.Background(), snCR, state, - webapi.NewClient("http://127.0.0.1:1"), "cluster", - 1, // maxFaultTolerance=1 — unused inside the fdEnabled branch - true, // fdEnabled=true - 2, // domainsNeededForFullDisjoint=2, domainsAvailable=2 -> chunksPerDomain=1 (well-provisioned) - 1, // npcs=1: currentRisk(1)+1=2 > npcs(1) -> correctly blocked - ) - - if err != nil { - t.Fatalf("expected nil error when blocked by domain gate, got %v", err) - } - if requeue != 10*time.Second { - t.Fatalf("expected 10s requeue when cross-domain is gated, got %v", requeue) - } - if state.Phase == simplyblockv1alpha1.DrainPhaseShutdownCalled { - t.Fatalf("phase must not advance when cross-domain drain slot is unavailable") - } -} - -// TestWorkerFailureDomainFromStatus verifies that the failure domain is read -// from status.nodes[].failureDomain (populated from the backend API response). -func TestWorkerFailureDomainFromStatus(t *testing.T) { - fd := int32(5) - snCR := &simplyblockv1alpha1.StorageNodeSet{ - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - Nodes: []simplyblockv1alpha1.NodeStatus{ - {Hostname: "worker-a", FailureDomain: &fd}, - }, - }, - } - - got, ok := workerFailureDomain(snCR, "worker-a") - if !ok { - t.Fatal("expected domain to be found in status") - } - if got != 5 { - t.Fatalf("expected 5 from status.nodes, got %d", got) - } -} - -// TestWorkerFailureDomainUnassigned verifies that a worker with no domain -// assignment returns (0, false). -func TestWorkerFailureDomainUnassigned(t *testing.T) { - snCR := &simplyblockv1alpha1.StorageNodeSet{} - _, ok := workerFailureDomain(snCR, "worker-missing") - if ok { - t.Fatal("expected (0, false) for unassigned worker") - } -} - -// ---- helper ---- - -func newNodeDrainTestReconciler(t *testing.T, objects ...client.Object) *NodeDrainCoordinatorReconciler { - t.Helper() - - scheme := newTestScheme( - t, - simplyblockv1alpha1.AddToScheme, - corev1.AddToScheme, - policyv1.AddToScheme, - ) - cl := newTestClient(t, scheme, []client.Object{ - &simplyblockv1alpha1.StorageNodeSet{}, - &simplyblockv1alpha2.StorageCluster{}, - }, objects...) - - return &NodeDrainCoordinatorReconciler{ - Client: cl, - Scheme: scheme, - } -} diff --git a/operator/internal/controller/simplyblockstoragenodeset_controller.go b/operator/internal/controller/simplyblockstoragenodeset_controller.go deleted file mode 100644 index 9104b7514..000000000 --- a/operator/internal/controller/simplyblockstoragenodeset_controller.go +++ /dev/null @@ -1,1708 +0,0 @@ -/* -Copyright 2025. - -Licensed under the Apache License, Version 2.0 (the "License"); -you may not use this file except in compliance with the License. -You may obtain a copy of the License at - - http://www.apache.org/licenses/LICENSE-2.0 - -Unless required by applicable law or agreed to in writing, software -distributed under the License is distributed on an "AS IS" BASIS, -WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. -See the License for the specific language governing permissions and -limitations under the License. -*/ - -package controller - -import ( - "context" - "encoding/json" - "fmt" - "net/http" - "reflect" - - "strconv" - "strings" - - "time" - - appsv1 "k8s.io/api/apps/v1" - corev1 "k8s.io/api/core/v1" - discoveryv1 "k8s.io/api/discovery/v1" - apierrors "k8s.io/apimachinery/pkg/api/errors" - metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" - "k8s.io/apimachinery/pkg/runtime" - "k8s.io/apimachinery/pkg/types" - "k8s.io/client-go/tools/events" - ctrl "sigs.k8s.io/controller-runtime" - "sigs.k8s.io/controller-runtime/pkg/builder" - "sigs.k8s.io/controller-runtime/pkg/client" - "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil" - "sigs.k8s.io/controller-runtime/pkg/event" - "sigs.k8s.io/controller-runtime/pkg/handler" - logf "sigs.k8s.io/controller-runtime/pkg/log" - "sigs.k8s.io/controller-runtime/pkg/predicate" - "sigs.k8s.io/controller-runtime/pkg/reconcile" - - "github.com/simplyblock/atlas/kube" - "github.com/simplyblock/atlas/ptr" - - simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" - simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" - clustercontroller "github.com/simplyblock/simplyblock-operator/internal/controllers/cluster" - "github.com/simplyblock/simplyblock-operator/internal/tlsutil" - "github.com/simplyblock/simplyblock-operator/internal/utils" - "github.com/simplyblock/simplyblock-operator/internal/webapi" -) - -// StorageNodeSetReconciler reconciles a StorageNodeSet object -type StorageNodeSetReconciler struct { - client.Client - Scheme *runtime.Scheme - Namespace string // operator namespace, used to look up the singleton ControlPlane CR - TLSEnabled bool - TLSProvider string - TLSMutualEnabled bool - Recorder events.EventRecorder -} - -type SNODEAPIResponse struct { - UUID string `json:"id"` - Status string `json:"status"` - IP string `json:"mgmt_ip"` - Health bool `json:"health_check"` - Hostname string `json:"hostname"` - DevicesCount int `json:"device_count"` - OnlineDevicesCount int `json:"online_device_count"` - CPU int `json:"cpu_spdk_count"` - Memory int64 `json:"spdk_mem"` - Volumes int `json:"lvols"` - RPC_PORT int `json:"rpc_port"` - LVOL_PORT int `json:"lvol_subsys_port"` - NVMF_PORT int `json:"nvmf_port"` - FailureDomain int `json:"failure_domain"` -} - -var ( - waitForNodeInfoReachableCheckFn = checkNodeInfoReachable - waitForNodeInfoReachableMaxRetries = 12 - waitForNodeInfoReachableRetryDelay = 10 * time.Second - - waitForNodeOnlineRetries = 60 - waitForNodeOnlineWaitInterval = 10 * time.Second - waitForNodeOnlineActivationDelay = 120 * time.Second - waitForNodeOnlineSleepFn = func(ctx context.Context, d time.Duration) error { - select { - case <-time.After(d): - return nil - case <-ctx.Done(): - return ctx.Err() - } - } - - syncNodeStatusInterval = 30 * time.Second - - spdkPodEventDelay = 20 * time.Second -) - -// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodesets,verbs=get;list;watch;create;update;patch;delete -// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodesets/status,verbs=get;update;patch -// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodesets/finalizers,verbs=update -// +kubebuilder:rbac:groups="",resources=services,verbs=get;list;watch;create;update;patch;delete -// +kubebuilder:rbac:groups="",resources=pods,verbs=get;list;watch -// +kubebuilder:rbac:groups=discovery.k8s.io,resources=endpointslices,verbs=get;list;watch;create;update;patch;delete -// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storageclusters,verbs=get;list;watch -// +kubebuilder:rbac:groups="",resources=nodes,verbs=get;list;watch;update;patch -// +kubebuilder:rbac:groups="",resources=secrets,verbs=get;list;watch -// +kubebuilder:rbac:groups=apps,resources=daemonsets,verbs=get;list;watch;create;update;patch;delete -// +kubebuilder:rbac:groups="",resources=serviceaccounts,verbs=get;list;watch;create;update;patch;delete -// +kubebuilder:rbac:groups=rbac.authorization.k8s.io,resources=clusterroles,verbs=get;list;watch;create;update;patch;delete -// +kubebuilder:rbac:groups=rbac.authorization.k8s.io,resources=clusterrolebindings,verbs=get;list;watch;create;update;patch;delete -// +kubebuilder:rbac:groups=cert-manager.io,resources=certificates,verbs=get;list;watch;create;update;patch;delete -// +kubebuilder:rbac:groups="",resources=events,verbs=get;list;watch -// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=volumemigrations,verbs=get;list;watch;create;delete -// +kubebuilder:rbac:groups="",resources=persistentvolumes,verbs=get;list;watch -// +kubebuilder:rbac:groups="",resources=persistentvolumeclaims,verbs=get;list;watch;update;patch - -// Reconcile is part of the main kubernetes reconciliation loop which aims to -// move the current state of the cluster closer to the desired state. -// TODO(user): Modify the Reconcile function to compare the state specified by -// the StorageNodeSet object against the actual cluster state, and then -// perform operations to make the cluster state reflect the state specified by -// the user. -// -// For more details, check Reconcile and its Result here: -// - https://pkg.go.dev/sigs.k8s.io/controller-runtime@v0.22.4/pkg/reconcile -func (r *StorageNodeSetReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) { - log := logf.FromContext(ctx) - - snCR := &simplyblockv1alpha1.StorageNodeSet{} - if err := r.Get(ctx, req.NamespacedName, snCR); err != nil { - return ctrl.Result{}, client.IgnoreNotFound(err) - } - - clusterUUID, err := utils.ResolveClusterUUID( - ctx, - r.Client, - snCR.Namespace, - snCR.Spec.ClusterName, - ) - - if err != nil { - log.Info("Cluster UUID not ready yet, requeuing", - "cluster", snCR.Spec.ClusterName, - ) - return ctrl.Result{RequeueAfter: 10 * time.Second}, nil - } - - /* -------------------- Deletion -------------------- */ - if updated, err := r.handleDeletion(ctx, snCR); updated || err != nil { - return ctrl.Result{}, err - } - - /* -------------------- Finalizer -------------------- */ - if updated, err := r.ensureFinalizer(ctx, snCR); updated || err != nil { - return ctrl.Result{}, err - } - - apiClient := webapi.NewClient() - - if err := labelWorkerNodes(ctx, r.Client, r.Recorder, snCR, clusterUUID); err != nil { - return ctrl.Result{}, err - } - - if err := r.reconcileRBAC(ctx, snCR); err != nil { - return ctrl.Result{}, err - } - - if err := r.reconcileService(ctx, snCR); err != nil { - return ctrl.Result{}, err - } - - if err := r.reconcileSpdkProxyService(ctx, snCR); err != nil { - return ctrl.Result{}, err - } - - // Reconcile certificates before the DaemonSet so the TLS Secret is more - // likely to exist when reconcileDaemonSet reads its resourceVersion to - // stamp it as a pod-template annotation. - if err := r.reconcileServingCertificates(ctx, snCR); err != nil { - return ctrl.Result{}, err - } - - // Reconcile per-node ConfigMap BEFORE the DaemonSet so that pods never - // start without the ConfigMap already present. - if err := r.reconcilePerNodeConfigMap(ctx, snCR); err != nil { - log.Error(err, "failed to reconcile per-node ConfigMap") - } - - if err := r.reconcileDaemonSet(ctx, snCR); err != nil { - return ctrl.Result{}, err - } - - if err := r.reconcileEndpointSlice(ctx, snCR); err != nil { - return ctrl.Result{}, err - } - - if err := r.reconcileSpdkProxyEndpointSlices(ctx, snCR); err != nil { - return ctrl.Result{}, err - } - - expectedPerHost := utils.ExpectedNodesPerHost(snCR) - - // Phase-1 bridge: create/sync/delete owned StorageNode CRs to match - // spec.workerNodes × spec.socketsToUse. The StorageNodeReconciler owns - // the per-node provisioning; this only manages CR lifecycle. - if err := r.reconcileStorageNodeCRs(ctx, snCR); err != nil { - log.Error(err, "failed to reconcile StorageNode CRs") - } - - if res, err := r.reconcileWorkerNodes(ctx, snCR, clusterUUID, apiClient, expectedPerHost); err != nil || res.RequeueAfter > 0 { - return res, err - } - - if err := r.syncTrackedNodesStatus(ctx, apiClient, clusterUUID, snCR); err != nil { - log.Error(err, "Failed to sync storage node status") - } - - // Sync manually created StorageNode CRs (not in spec.workerNodes) into - // StorageNodeSet.status.nodes[] so their status is visible in the fleet view. - if err := r.syncManualStorageNodeStatus(ctx, snCR); err != nil { - log.Error(err, "Failed to sync manual StorageNode status") - } - - // On every reconcile, check whether the cluster is still unready and if - // the activation conditions are now met. This catches cases where the - // operator restarted after nodes came online but before activation fired, - // or where the activation trigger was missed during concurrent reconciles. - if clusterCR, err := utils.ResolveClusterCR(ctx, r.Client, snCR.Namespace, snCR.Spec.ClusterName); err == nil { - if clusterCR.Status.Status == utils.ClusterStatusUnready { - if activateErr := maybeActivateCluster(ctx, apiClient, clusterUUID, snCR, r); activateErr != nil { - log.Info("Activation conditions not yet met", "reason", activateErr.Error()) - } - } - } - - hasTracked := false - for _, n := range snCR.Status.Nodes { - if n.UUID != "" { - hasTracked = true - break - } - } - if !hasTracked { - return ctrl.Result{}, nil - } - return ctrl.Result{RequeueAfter: syncNodeStatusInterval}, nil -} - -// reconcileWorkerNode handles provisioning and online-wait for a single worker node. -func (r *StorageNodeSetReconciler) reconcileWorkerNode( - ctx context.Context, - snCR *simplyblockv1alpha1.StorageNodeSet, - nodeName, clusterUUID string, - apiClient *webapi.Client, - expectedPerHost int, -) (ctrl.Result, error) { - log := logf.FromContext(ctx) - - // Count status entries that have a UUID — these represent backend nodes - // confirmed online at least once. Skip when all socket nodes are tracked. - trackedCount := 0 - for _, n := range snCR.Status.Nodes { - if n.Hostname == nodeName && n.UUID != "" { - trackedCount++ - } - } - if trackedCount >= expectedPerHost { - return ctrl.Result{}, nil - } - - ip, err := getNodeInternalIP(ctx, r.Client, nodeName) - if err != nil { - log.Error(err, "failed to get internal IP", "node", nodeName) - return ctrl.Result{RequeueAfter: time.Second * 10}, nil - } - - // StorageNodeReconciler is the sole owner of provisioning. Only call - // pollNodeOnline once ALL StorageNode CRs for this worker have their UUID set - // (all nodes confirmed online). Until then requeue — calling pollNodeOnline - // before all nodes are posted would time out waiting for expectedPerHost nodes. - if r.storageNodeAlreadyPosted(ctx, snCR.Namespace, nodeName) { - if r.allStorageNodesOnline(ctx, snCR.Namespace, nodeName, expectedPerHost) { - return r.pollNodeOnline(ctx, apiClient, clusterUUID, ip, nodeName, expectedPerHost, snCR) - } - return ctrl.Result{RequeueAfter: waitForNodeOnlineWaitInterval}, nil - } - - // StorageNodeReconciler is the sole owner of provisioning. If it hasn't - // POSTed yet, requeue and wait — never POST from here. - return ctrl.Result{RequeueAfter: 10 * time.Second}, nil -} - -// SetupWithManager sets up the controller with the Manager. -func (r *StorageNodeSetReconciler) SetupWithManager(mgr ctrl.Manager) error { - return ctrl.NewControllerManagedBy(mgr). - For(&simplyblockv1alpha1.StorageNodeSet{}). - Named("storagenodeset"). - Watches( - &corev1.Pod{}, - handler.EnqueueRequestsFromMapFunc(r.spdkProxyPodToStorageNodeSetRequests), - builder.WithPredicates( - predicate.NewPredicateFuncs(isSpdkProxyPod), - predicate.Funcs{ - UpdateFunc: func(e event.UpdateEvent) bool { - oldPod, ok := e.ObjectOld.(*corev1.Pod) - if !ok { - return true - } - newPod, ok := e.ObjectNew.(*corev1.Pod) - if !ok { - return true - } - return oldPod.Status.Phase != newPod.Status.Phase - }, - }, - ), - ). - Watches( - &corev1.Secret{}, - handler.EnqueueRequestsFromMapFunc(r.tlsSecretToStorageNodeSetRequests), - builder.WithPredicates(predicate.NewPredicateFuncs(isStorageNodeSetTLSSecret)), - ). - Watches( - &simplyblockv1alpha2.ControlPlane{}, - handler.EnqueueRequestsFromMapFunc(r.controlPlaneToStorageNodeSetRequests), - builder.WithPredicates(predicate.NewPredicateFuncs(isSimplyblockControlPlane)), - ). - Complete(r) -} - -func isSpdkProxyPod(obj client.Object) bool { - return obj.GetLabels()["role"] == utils.LabelSpdkProxyRole -} - -func isStorageNodeSetTLSSecret(obj client.Object) bool { - return obj.GetName() == utils.SecretNameStorageNodeSetAPITLS -} - -func isSimplyblockControlPlane(obj client.Object) bool { - return obj.GetName() == SingletonControlPlaneName -} - -func (r *StorageNodeSetReconciler) controlPlaneToStorageNodeSetRequests( - ctx context.Context, - obj client.Object, -) []reconcile.Request { - var snList simplyblockv1alpha1.StorageNodeSetList - if err := r.List(ctx, &snList, client.InNamespace(obj.GetNamespace())); err != nil { - return nil - } - reqs := make([]reconcile.Request, 0, len(snList.Items)) - for _, sn := range snList.Items { - reqs = append(reqs, reconcile.Request{ - NamespacedName: types.NamespacedName{Namespace: sn.Namespace, Name: sn.Name}, - }) - } - return reqs -} - -// tlsSecretToStorageNodeSetRequests enqueues every StorageNodeSet CR in the -// Secret's namespace when the storage-node-api TLS Secret changes. Coupled -// with the resourceVersion annotation stamped on the DaemonSet pod template, -// this drives a rolling restart whenever cert-manager (or OpenShift's -// service-ca) rotates the Secret. -func (r *StorageNodeSetReconciler) tlsSecretToStorageNodeSetRequests( - ctx context.Context, - obj client.Object, -) []reconcile.Request { - var snList simplyblockv1alpha1.StorageNodeSetList - if err := r.List(ctx, &snList, client.InNamespace(obj.GetNamespace())); err != nil { - return nil - } - reqs := make([]reconcile.Request, 0, len(snList.Items)) - for _, sn := range snList.Items { - reqs = append(reqs, reconcile.Request{ - NamespacedName: types.NamespacedName{Namespace: sn.Namespace, Name: sn.Name}, - }) - } - return reqs -} - -// spdkProxyPodToStorageNodeSetRequests enqueues every StorageNodeSet CR in the Pod's -// namespace when a spdk-proxy pod changes. Pods are created by the backend, not -// by the operator, so there is no forward owner reference — fanning out within -// the namespace is the simplest correct mapping and cheap in practice (one CR -// per namespace is typical). -func (r *StorageNodeSetReconciler) spdkProxyPodToStorageNodeSetRequests( - ctx context.Context, - obj client.Object, -) []reconcile.Request { - var snList simplyblockv1alpha1.StorageNodeSetList - if err := r.List(ctx, &snList, client.InNamespace(obj.GetNamespace())); err != nil { - return nil - } - reqs := make([]reconcile.Request, 0, len(snList.Items)) - for _, sn := range snList.Items { - reqs = append(reqs, reconcile.Request{ - NamespacedName: types.NamespacedName{Namespace: sn.Namespace, Name: sn.Name}, - }) - } - return reqs -} - -func (r *StorageNodeSetReconciler) handleDeletion( - ctx context.Context, - snCR *simplyblockv1alpha1.StorageNodeSet, -) (bool, error) { - - if snCR.DeletionTimestamp.IsZero() { - return false, nil - } - - if !controllerutil.ContainsFinalizer(snCR, utils.FinalizerStorageNodeSet) { - return true, nil - } - - controllerutil.RemoveFinalizer(snCR, utils.FinalizerStorageNodeSet) - return true, r.Update(ctx, snCR) -} - -func (r *StorageNodeSetReconciler) ensureFinalizer( - ctx context.Context, - snCR *simplyblockv1alpha1.StorageNodeSet, -) (bool, error) { - - if controllerutil.ContainsFinalizer(snCR, utils.FinalizerStorageNodeSet) { - return false, nil - } - - controllerutil.AddFinalizer(snCR, utils.FinalizerStorageNodeSet) - return true, r.Update(ctx, snCR) -} - -// storageNodeUUIDLabelPrefix marks a worker Node with the SimplyBlock storage-node -// instance(s) co-located on it. The label KEY is "." -// and must stay stable for the Node's lifetime — Kubernetes' external-provisioner -// caches the *set* of topology keys in the CSINode object at node-plugin -// registration time and only refreshes it when the node-driver pod restarts, then -// hard-errors CreateVolume if a live Node's topology-label keys don't match that -// cached set. The label VALUE is the storage-node UUID, which changes freely (e.g. -// when a node is replaced) since values are always read fresh — only the key must -// never depend on anything that can change post-registration. Cluster-scoping the -// key (not just the value) also stops a worker hosting instances from more than one -// SimplyBlock cluster from having one cluster's slot collide with another's. -// Consumed by the CSI node plugin (csi-driver/internal/csi/node) to advertise -// CSI topology, and by the CSI controller (createVolume) to co-locate a new volume's -// primary with whichever worker the consuming Pod is scheduled to. Keep this literal -// in sync with topologyKeyStorageNodeUUIDPrefix in csi-driver. -const storageNodeUUIDLabelPrefix = "simplyblock.io/storage-node-uuid." - -// labelWorkerNodes applies the storage-plane node labels — the DaemonSet node -// selector (io.simplyblock.storagenodeset) plus the per-slot storage-node-uuid -// labels — to every worker owned by the StorageNodeSet. -// It is a free function (not a method) so both the StorageNodeSet reconciler and -// the StorageNodeOps migration flow drive the identical labeling; a migration -// target that is not yet in spec.workerNodes is passed via extraWorkers so its -// DaemonSet pod schedules before the topology swap. -func labelWorkerNodes( - ctx context.Context, - c client.Client, - recorder events.EventRecorder, - sn *simplyblockv1alpha1.StorageNodeSet, - clusterUUID string, - extraWorkers ...string, -) error { - // Collect all workers: spec.workerNodes, any explicitly requested extras - // (e.g. a migration target not yet in the spec), plus any manually created - // StorageNode CRs that reference this StorageNodeSet but are not in spec.workerNodes. - workers := make(map[string]struct{}, len(sn.Spec.WorkerNodes)+len(extraWorkers)) - for _, w := range sn.Spec.WorkerNodes { - workers[w] = struct{}{} - } - for _, w := range extraWorkers { - workers[w] = struct{}{} - } - - // slotsByWorker maps worker -> "." -> storage-node UUID, - // built from every owned StorageNode CR that has come online at least once - // (Status.UUID set). The slot key is stable; only the UUID value churns. - slotsByWorker := make(map[string]map[string]string) - - var snList simplyblockv1alpha1.StorageNodeList - // A failed List must abort: slotsByWorker is the source of truth for which - // per-slot storage-node-uuid labels are desired, so an empty map from a - // transient API error or a missing field index would make the cleanup loop - // below delete every simplyblock.io/storage-node-uuid..* label - // from the workers, breaking CSI topology. Return the error and requeue. - if err := c.List(ctx, &snList, - client.InNamespace(sn.Namespace), - client.MatchingFields{"spec.storageNodeSetRef": sn.Name}, - ); err != nil { - return fmt.Errorf("listing StorageNodes for StorageNodeSet %s: %w", sn.Name, err) - } - for _, snCR := range snList.Items { - workers[snCR.Spec.WorkerNode] = struct{}{} - - if snCR.Status.UUID == "" { - continue - } - ordinal := int32(0) - if snCR.Spec.SocketIndex != nil { - ordinal = *snCR.Spec.SocketIndex - } - if slotsByWorker[snCR.Spec.WorkerNode] == nil { - slotsByWorker[snCR.Spec.WorkerNode] = map[string]string{} - } - slotKey := fmt.Sprintf("%s.%d", clusterUUID, ordinal) - slotsByWorker[snCR.Spec.WorkerNode][slotKey] = snCR.Status.UUID - } - - // Per-StorageNodeSet label: used as the DaemonSet node selector so that - // each StorageNodeSet owns its own DaemonSet and per-node ConfigMap, - // enabling multiple StorageNodeSets per cluster for node grouping. - snsLabelKey := kube.LabelStorageNodeSet - snsLabelVal := sn.Name - - for nodeName := range workers { - var node corev1.Node - if err := c.Get(ctx, client.ObjectKey{Name: nodeName}, &node); err != nil { - recorder.Eventf(sn, nil, corev1.EventTypeWarning, "WorkerNodeNotFound", "WorkerNodeNotFound", - "worker node %q: %v", nodeName, err) - return err - } - - if node.Labels == nil { - node.Labels = map[string]string{} - } - - changed := false - if node.Labels[snsLabelKey] != snsLabelVal { - node.Labels[snsLabelKey] = snsLabelVal - changed = true - } - - desired := slotsByWorker[nodeName] - for k, v := range node.Labels { - if !strings.HasPrefix(k, storageNodeUUIDLabelPrefix) { - continue - } - slot := strings.TrimPrefix(k, storageNodeUUIDLabelPrefix) - sep := strings.LastIndex(slot, ".") - if sep < 0 || slot[:sep] != clusterUUID { - // Slot belongs to a different SimplyBlock cluster (a worker can host - // storage-node instances from more than one) or is malformed — leave - // it untouched; this reconcile only owns clusterUUID's slots. - continue - } - if desired[slot] != v { - delete(node.Labels, k) - changed = true - } - } - for slot, uuid := range desired { - k := storageNodeUUIDLabelPrefix + slot - if node.Labels[k] != uuid { - node.Labels[k] = uuid - changed = true - } - } - - if !changed { - continue - } - - if err := c.Update(ctx, &node); err != nil { - return err - } - } - - return nil -} - -func (r *StorageNodeSetReconciler) reconcileDaemonSet( - ctx context.Context, - snCR *simplyblockv1alpha1.StorageNodeSet, -) error { - - if snCR.Spec.ClusterImage == "" { - // Read at v1alpha2, the stored version, rather than at this type's own - // v1alpha1. A read of the retired version is answered only by the - // conversion webhook, which a fresh install does not deploy, and the - // cache it would be served from lists empty instead of failing: the - // fallback would report the singleton missing on a cluster that has it. - cp := &simplyblockv1alpha2.ControlPlane{} - if err := r.Get(ctx, types.NamespacedName{Namespace: r.Namespace, Name: SingletonControlPlaneName}, cp); err != nil { - return fmt.Errorf("clusterImage not set and ControlPlane %q not found: %w", SingletonControlPlaneName, err) - } - image := "" - if cp.Spec.Source != nil && cp.Spec.Source.Managed != nil { - image = cp.Spec.Source.Managed.Image - } - if image == "" { - return fmt.Errorf( - "clusterImage not set and ControlPlane %q has no spec.source.managed.image", - SingletonControlPlaneName) - } - snCR = snCR.DeepCopy() - snCR.Spec.ClusterImage = image - } - - tlsSecretRV, err := r.getTLSSecretResourceVersion(ctx, snCR.Namespace) - if err != nil { - return err - } - - ds := utils.BuildStorageNodeSetDaemonSet(snCR, r.TLSEnabled, r.TLSMutualEnabled, r.TLSProvider, tlsSecretRV) - - if err := controllerutil.SetControllerReference(snCR, ds, r.Scheme); err != nil { - return err - } - - var existing appsv1.DaemonSet - err = r.Get(ctx, client.ObjectKeyFromObject(ds), &existing) - if apierrors.IsNotFound(err) { - return r.Create(ctx, ds) - } - if err != nil { - return err - } - - ds.ResourceVersion = existing.ResourceVersion - return r.Update(ctx, ds) -} - -// getTLSSecretResourceVersion returns the storage-node-api TLS Secret's -// metadata.resourceVersion, or "" if TLS is disabled or the Secret has not -// been provisioned yet. The value is stamped onto the DaemonSet's pod -// template so that cert rotations (where the Secret object changes but its -// name does not) trigger a rolling restart. -func (r *StorageNodeSetReconciler) getTLSSecretResourceVersion( - ctx context.Context, - namespace string, -) (string, error) { - if !r.TLSEnabled { - return "", nil - } - var sec corev1.Secret - err := r.Get(ctx, types.NamespacedName{ - Namespace: namespace, - Name: utils.SecretNameStorageNodeSetAPITLS, - }, &sec) - if apierrors.IsNotFound(err) { - return "", nil - } - if err != nil { - return "", err - } - return sec.ResourceVersion, nil -} - -func (r *StorageNodeSetReconciler) reconcileService( - ctx context.Context, - snCR *simplyblockv1alpha1.StorageNodeSet, -) error { - svc := utils.BuildStorageNodeSetService(snCR, r.TLSEnabled, r.TLSProvider) - if err := controllerutil.SetControllerReference(snCR, svc, r.Scheme); err != nil { - return fmt.Errorf("failed to set Service owner reference: %w", err) - } - - var existing corev1.Service - err := r.Get(ctx, client.ObjectKeyFromObject(svc), &existing) - if apierrors.IsNotFound(err) { - return r.Create(ctx, svc) - } - if err != nil { - return err - } - - svc.ResourceVersion = existing.ResourceVersion - svc.Spec.ClusterIP = existing.Spec.ClusterIP - return r.Update(ctx, svc) -} - -func (r *StorageNodeSetReconciler) reconcileServingCertificates( - ctx context.Context, - snCR *simplyblockv1alpha1.StorageNodeSet, -) error { - if !r.TLSEnabled || !utils.IsCertManagerTLSProvider(r.TLSProvider) { - return nil - } - - certificates := []struct { - serviceName string - secretName string - }{ - { - serviceName: "simplyblock-storage-node-api", - secretName: utils.SecretNameStorageNodeSetAPITLS, - }, - { - serviceName: "simplyblock-spdk-proxy", - secretName: utils.SecretNameSpdkProxyTLS, - }, - } - - for _, cert := range certificates { - if err := r.reconcileServingCertificate(ctx, snCR, cert.serviceName, cert.secretName); err != nil { - return err - } - } - - return nil -} - -func (r *StorageNodeSetReconciler) reconcileServingCertificate( - ctx context.Context, - snCR *simplyblockv1alpha1.StorageNodeSet, - serviceName, secretName string, -) error { - cert := utils.BuildServiceServingCertificate(snCR.Namespace, serviceName, secretName) - if _, err := controllerutil.CreateOrUpdate(ctx, r.Client, cert, func() error { - desired := utils.BuildServiceServingCertificate(snCR.Namespace, serviceName, secretName) - cert.Object["spec"] = desired.Object["spec"] - return controllerutil.SetControllerReference(snCR, cert, r.Scheme) - }); err != nil { - return fmt.Errorf("failed to apply serving Certificate for %s: %w", serviceName, err) - } - - return nil -} - -func (r *StorageNodeSetReconciler) reconcileEndpointSlice( - ctx context.Context, - snCR *simplyblockv1alpha1.StorageNodeSet, -) error { - log := logf.FromContext(ctx) - - // Start with workers from spec.workerNodes. - nodeIPs := make(map[string]string) - for _, nodeName := range snCR.Spec.WorkerNodes { - ip, err := getNodeInternalIP(ctx, r.Client, nodeName) - if err != nil { - log.Error(err, "failed to get internal IP for EndpointSlice, skipping node", "node", nodeName) - continue - } - nodeIPs[nodeName] = ip - } - - // Also include workers from manually created StorageNode CRs so their - // per-node DNS hostname resolves and checkNodeInfoReachable succeeds. - var snList simplyblockv1alpha1.StorageNodeList - if err := r.List(ctx, &snList, - client.InNamespace(snCR.Namespace), - client.MatchingFields{"spec.storageNodeSetRef": snCR.Name}, - ); err == nil { - for _, sn := range snList.Items { - if _, ok := nodeIPs[sn.Spec.WorkerNode]; ok { - continue // already covered - } - ip, err := getNodeInternalIP(ctx, r.Client, sn.Spec.WorkerNode) - if err != nil { - log.Error(err, "failed to get IP for manual StorageNode worker, skipping", "worker", sn.Spec.WorkerNode) - continue - } - nodeIPs[sn.Spec.WorkerNode] = ip - } - } - - // Also include every node currently labeled into THIS StorageNodeSet. A - // StorageNodeOps migrate labels the target worker (so the storage-node - // DaemonSet schedules a pod there) before the backend restart, but the - // target is not yet in spec.workerNodes or any StorageNode CR — those are - // only updated by reconcileMigratedTopology AFTER the node comes online on - // the target. Without publishing the target's per-pod DNS name here, the - // control plane cannot resolve node_address, the restart fails, the node - // never comes online, and the topology swap never runs: a deadlock. Keying - // off the label breaks it — the target's DNS entry appears as soon as it is - // labeled. (This mirrored behavior was lost in the StorageNodeSet/StorageNode - // split.) The label is scoped per set, so a second StorageNodeSet's workers - // are not pulled into this set's slice. - var nodeList corev1.NodeList - if err := r.List(ctx, &nodeList, client.MatchingLabels{ - kube.LabelStorageNodeSet: snCR.Name, - }); err == nil { - for i := range nodeList.Items { - nodeName := nodeList.Items[i].Name - if _, ok := nodeIPs[nodeName]; ok { - continue // already covered - } - ip, err := getNodeInternalIP(ctx, r.Client, nodeName) - if err != nil { - log.Error(err, "failed to get IP for storage-plane-labeled node, skipping", "worker", nodeName) - continue - } - nodeIPs[nodeName] = ip - } - } else { - log.Error(err, "failed to list storage-plane-labeled nodes for EndpointSlice") - } - - return r.applyStorageNodeSetEndpointSlice(ctx, snCR, nodeIPs) -} - -// applyStorageNodeSetEndpointSlice creates or updates the storage-node-api -// EndpointSlice with the supplied nodeIPs map. -func (r *StorageNodeSetReconciler) applyStorageNodeSetEndpointSlice( - ctx context.Context, - snCR *simplyblockv1alpha1.StorageNodeSet, - nodeIPs map[string]string, -) error { - eps := utils.BuildStorageNodeSetEndpointSlice(snCR, nodeIPs) - if err := controllerutil.SetControllerReference(snCR, eps, r.Scheme); err != nil { - return fmt.Errorf("failed to set EndpointSlice owner reference: %w", err) - } - - var existing discoveryv1.EndpointSlice - err := r.Get(ctx, client.ObjectKeyFromObject(eps), &existing) - if apierrors.IsNotFound(err) { - return r.Create(ctx, eps) - } - if err != nil { - return err - } - - eps.ResourceVersion = existing.ResourceVersion - return r.Update(ctx, eps) -} - -func (r *StorageNodeSetReconciler) reconcileSpdkProxyService( - ctx context.Context, - snCR *simplyblockv1alpha1.StorageNodeSet, -) error { - svc := utils.BuildSpdkProxyService(snCR, r.TLSEnabled, r.TLSProvider) - if err := controllerutil.SetControllerReference(snCR, svc, r.Scheme); err != nil { - return fmt.Errorf("failed to set spdk-proxy Service owner reference: %w", err) - } - - var existing corev1.Service - err := r.Get(ctx, client.ObjectKeyFromObject(svc), &existing) - if apierrors.IsNotFound(err) { - return r.Create(ctx, svc) - } - if err != nil { - return err - } - - svc.ResourceVersion = existing.ResourceVersion - svc.Spec.ClusterIP = existing.Spec.ClusterIP - return r.Update(ctx, svc) -} - -func (r *StorageNodeSetReconciler) reconcileSpdkProxyEndpointSlices( - ctx context.Context, - snCR *simplyblockv1alpha1.StorageNodeSet, -) error { - log := logf.FromContext(ctx) - - var pods corev1.PodList - if err := r.List(ctx, &pods, - client.InNamespace(snCR.Namespace), - client.MatchingLabels{"role": utils.LabelSpdkProxyRole}, - ); err != nil { - return fmt.Errorf("failed to list spdk-proxy pods: %w", err) - } - - // portsWithAnyPod tracks every RPC port that has a matching pod object AT - // ALL, ready or not -- computed separately from byPort (ready pods only) - // so the delete pass below can tell "pod is genuinely gone" apart from - // "pod exists but isn't ready this instant". RPC_PORT is a static env var - // on the pod spec, readable the moment the pod is scheduled, well before - // it ever becomes ready, so this is safe to compute from the full list. - byPort := map[int32][]utils.SpdkProxyEndpoint{} - portsWithAnyPod := map[int32]bool{} - for i := range pods.Items { - pod := &pods.Items[i] - rpcPort, ok := extractSpdkProxyRpcPort(pod) - if !ok { - log.Info("skipping spdk-proxy pod: unable to determine RPC_PORT", "pod", pod.Name) - continue - } - portsWithAnyPod[rpcPort] = true - if !isSpdkProxyPodReady(pod) { - continue - } - byPort[rpcPort] = append(byPort[rpcPort], utils.SpdkProxyEndpoint{ - NodeName: pod.Spec.NodeName, - PodIP: pod.Status.PodIP, - RpcPort: rpcPort, - }) - } - - for rpcPort, endpoints := range byPort { - eps, err := utils.BuildSpdkProxyEndpointSlice(snCR, rpcPort, endpoints) - if err != nil { - return err - } - if err := controllerutil.SetControllerReference(snCR, eps, r.Scheme); err != nil { - return fmt.Errorf("failed to set spdk-proxy EndpointSlice owner reference: %w", err) - } - - var existing discoveryv1.EndpointSlice - err = r.Get(ctx, client.ObjectKeyFromObject(eps), &existing) - if apierrors.IsNotFound(err) { - if err := r.Create(ctx, eps); err != nil { - return err - } - continue - } - if err != nil { - return err - } - eps.ResourceVersion = existing.ResourceVersion - if err := r.Update(ctx, eps); err != nil { - return err - } - } - - // Delete orphaned slices whose RPC_PORT has no matching pod AT ALL. - // - // Deliberately checked against portsWithAnyPod, NOT byPort: a pod that's - // merely not-ready this instant (isSpdkProxyPodReady can flip false for - // a single missed probe tick on either container, well short of what - // would restart the container or emit an Unhealthy event) must not have - // its DNS entry deleted -- the previous incident's root cause. Only - // delete when the pod for that port is genuinely gone (scaled down, - // node removed, rescheduled to a different port); a transiently - // not-ready pod simply keeps its last-known-good EndpointSlice in place - // until the next reconcile finds it ready and refreshes it via the - // create/update pass above. - var existingSlices discoveryv1.EndpointSliceList - if err := r.List(ctx, &existingSlices, - client.InNamespace(snCR.Namespace), - client.MatchingLabels{"kubernetes.io/service-name": "simplyblock-spdk-proxy"}, - ); err != nil { - return fmt.Errorf("failed to list existing spdk-proxy EndpointSlices: %w", err) - } - for i := range existingSlices.Items { - slice := &existingSlices.Items[i] - if !metav1.IsControlledBy(slice, snCR) { - continue - } - keep := false - for _, p := range slice.Ports { - if p.Port != nil && portsWithAnyPod[*p.Port] { - keep = true - break - } - } - if keep { - continue - } - if err := r.Delete(ctx, slice); err != nil && !apierrors.IsNotFound(err) { - return fmt.Errorf("failed to delete stale spdk-proxy EndpointSlice %s: %w", slice.Name, err) - } - } - - return nil -} - -// workerIsInFlight returns true if a node-add POST has already been sent for -// nodeName and is still being tracked — either via PendingNodeAdds (primary) -// or the legacy UUID=="" placeholder (backward compatibility). -func workerIsInFlight(snCR *simplyblockv1alpha1.StorageNodeSet, nodeName string) bool { - if _, ok := snCR.Status.PendingNodeAdds[nodeName]; ok { - return true - } - for _, n := range snCR.Status.Nodes { - if n.Hostname == nodeName && n.UUID == "" { - return true - } - } - return false -} - -// recordSpdkPodEvents finds the worker's pending SPDK pod, fetches its most -// recent Kubernetes event, and surfaces it on the StorageNodeSet CR status so -// operators can see why a pod is stuck without running kubectl describe. -func (r *StorageNodeSetReconciler) recordSpdkPodEvents( - ctx context.Context, - snCR *simplyblockv1alpha1.StorageNodeSet, - nodeName string, -) { - log := logf.FromContext(ctx) - - var podList corev1.PodList - if err := r.List(ctx, &podList, - client.InNamespace(snCR.Namespace), - client.MatchingLabels{"role": utils.LabelSpdkProxyRole}, - ); err != nil { - log.Error(err, "recordSpdkPodEvents: failed to list SPDK pods", "node", nodeName) - return - } - - var targetPod *corev1.Pod - for i := range podList.Items { - pod := &podList.Items[i] - if pod.Status.Phase != corev1.PodPending { - continue - } - if pod.Spec.NodeName == nodeName || - pod.Spec.NodeSelector["kubernetes.io/hostname"] == nodeName { - targetPod = pod - break - } - } - if targetPod == nil { - return - } - - var eventList corev1.EventList - if err := r.List(ctx, &eventList, client.InNamespace(snCR.Namespace)); err != nil { - log.Error(err, "recordSpdkPodEvents: failed to list events", "node", nodeName) - return - } - - var latest *corev1.Event - for i := range eventList.Items { - ev := &eventList.Items[i] - if ev.InvolvedObject.Name != targetPod.Name { - continue - } - if latest == nil || ev.LastTimestamp.After(latest.LastTimestamp.Time) { - latest = ev - } - } - if latest == nil { - return - } - - r.Recorder.Eventf(snCR, nil, corev1.EventTypeWarning, latest.Reason, latest.Reason, - "worker %s: %s", nodeName, latest.Message) - r.emitOnStorageNodeForWorker(ctx, snCR, nodeName, corev1.EventTypeWarning, latest.Reason, latest.Message) - - // Persist the flag so the recovery event is emitted correctly even if the - // operator restarts before the node comes online. - patch := client.MergeFrom(snCR.DeepCopy()) - if snCR.Status.SchedulingFailedWorkers == nil { - snCR.Status.SchedulingFailedWorkers = make(map[string]bool) - } - snCR.Status.SchedulingFailedWorkers[nodeName] = true - if err := r.Status().Patch(ctx, snCR, patch); err != nil { - log.Error(err, "recordSpdkPodEvents: failed to persist scheduling failure flag", "node", nodeName) - } -} - -// reconcileWorkerNodes fans out the node-add loop across parallel (non-FDB) and -// sequential (FDB) workers, respecting MaxParallelNodeAdds. -// MaxParallelNodeAdds carries a +kubebuilder:default=1 marker so the API server -// always populates it before the CR is stored — it is safe to dereference directly. -func (r *StorageNodeSetReconciler) reconcileWorkerNodes( - ctx context.Context, - snCR *simplyblockv1alpha1.StorageNodeSet, - clusterUUID string, - apiClient *webapi.Client, - expectedPerHost int, -) (ctrl.Result, error) { - fdbWorkers := r.fdbWorkerSet(ctx, snCR) - - var parallelWorkers, sequentialWorkers []string - for _, nodeName := range snCR.Spec.WorkerNodes { - if fdbWorkers[nodeName] { - sequentialWorkers = append(sequentialWorkers, nodeName) - } else { - parallelWorkers = append(parallelWorkers, nodeName) - } - } - - maxParallel := int(*snCR.Spec.MaxParallelNodeAdds) - - inFlight := 0 - for _, nodeName := range parallelWorkers { - if workerIsInFlight(snCR, nodeName) { - inFlight++ - } - } - availableSlots := maxParallel - inFlight - - var parallelRequeueAfter time.Duration - for _, nodeName := range parallelWorkers { - alreadyInFlight := workerIsInFlight(snCR, nodeName) - if !alreadyInFlight { - if availableSlots <= 0 { - if waitForNodeOnlineWaitInterval > parallelRequeueAfter { - parallelRequeueAfter = waitForNodeOnlineWaitInterval - } - continue - } - } - res, err := r.reconcileWorkerNode(ctx, snCR, nodeName, clusterUUID, apiClient, expectedPerHost) - if err != nil { - return ctrl.Result{}, err - } - // Only count the slot if the POST was genuinely sent (PendingNodeAdds - // was set). A transient failure (e.g. checkNodeInfoReachable) clears - // PendingNodeAdds immediately, so the slot should not be consumed. - if !alreadyInFlight && workerIsInFlight(snCR, nodeName) { - availableSlots-- - } - if res.RequeueAfter > parallelRequeueAfter { - parallelRequeueAfter = res.RequeueAfter - } - } - - for _, nodeName := range sequentialWorkers { - res, err := r.reconcileWorkerNode(ctx, snCR, nodeName, clusterUUID, apiClient, expectedPerHost) - if err != nil { - return ctrl.Result{}, err - } - // Always return early for sequential (FDB) workers on any requeue. - // If a concurrent reconcile's PendingNodeAdds persist failed (conflict), - // workerIsInFlight would be false on the local snCR even though the - // worker is effectively claimed — continuing the loop would process the - // next FDB worker in parallel, violating the one-at-a-time guarantee. - if res.RequeueAfter > 0 { - return res, nil - } - } - - return ctrl.Result{RequeueAfter: parallelRequeueAfter}, nil -} - -// reconcileRBAC ensures the ServiceAccount, ClusterRole, and ClusterRoleBinding -// required by the storage-node DaemonSet are present and up to date. -func (r *StorageNodeSetReconciler) reconcileRBAC(ctx context.Context, snCR *simplyblockv1alpha1.StorageNodeSet) error { - sa := utils.BuildStorageNodeSetServiceAccount(snCR.Namespace) - if err := controllerutil.SetControllerReference(snCR, sa, r.Scheme); err != nil { - return fmt.Errorf("failed to set ServiceAccount owner reference: %w", err) - } - desiredSAOwnerRefs := sa.OwnerReferences - if _, err := controllerutil.CreateOrUpdate(ctx, r.Client, sa, func() error { - sa.OwnerReferences = desiredSAOwnerRefs - return nil - }); err != nil { - return fmt.Errorf("failed to apply ServiceAccount: %w", err) - } - - cr := utils.BuildStorageNodeSetClusterRole(ptr.BoolFromOrFalse(snCR.Spec.OpenShiftCluster)) - desiredCRRules := cr.Rules - if _, err := controllerutil.CreateOrUpdate(ctx, r.Client, cr, func() error { - cr.Rules = desiredCRRules - return nil - }); err != nil { - return fmt.Errorf("failed to apply ClusterRole: %w", err) - } - - crb := utils.BuildStorageNodeSetClusterRoleBinding(snCR.Namespace) - desiredCRBSubjects := crb.Subjects - desiredCRBRoleRef := crb.RoleRef - if _, err := controllerutil.CreateOrUpdate(ctx, r.Client, crb, func() error { - crb.Subjects = desiredCRBSubjects - crb.RoleRef = desiredCRBRoleRef - return nil - }); err != nil { - return fmt.Errorf("failed to apply ClusterRoleBinding: %w", err) - } - return nil -} - -// fdbWorkerSet returns the set of worker node names (from snCR.Spec.WorkerNodes) -// that currently host at least one FDB pod. These workers must be added -// sequentially to avoid simultaneous reboots that reduce FDB fault tolerance. -func (r *StorageNodeSetReconciler) fdbWorkerSet(ctx context.Context, snCR *simplyblockv1alpha1.StorageNodeSet) map[string]bool { - workerSet := make(map[string]bool, len(snCR.Spec.WorkerNodes)) - for _, w := range snCR.Spec.WorkerNodes { - workerSet[w] = false - } - - var podList corev1.PodList - if err := r.List(ctx, &podList, - client.InNamespace(snCR.Namespace), - client.HasLabels{utils.LabelFDBClusterName}, - ); err != nil { - return workerSet - } - - fdbWorkers := make(map[string]bool) - for _, pod := range podList.Items { - if pod.Spec.NodeName != "" { - if _, isWorker := workerSet[pod.Spec.NodeName]; isWorker { - fdbWorkers[pod.Spec.NodeName] = true - } - } - } - return fdbWorkers -} - -func isSpdkProxyPodReady(pod *corev1.Pod) bool { - if pod.Status.Phase != corev1.PodRunning { - return false - } - if pod.Spec.NodeName == "" || pod.Status.PodIP == "" { - return false - } - for _, cs := range pod.Status.ContainerStatuses { - if !cs.Ready { - return false - } - } - return len(pod.Status.ContainerStatuses) > 0 -} - -// extractSpdkProxyRpcPort reads RPC_PORT from the spdk-proxy-container env; as -// a defensive fallback it parses the pod name pattern -// snode-spdk-pod--. -func extractSpdkProxyRpcPort(pod *corev1.Pod) (int32, bool) { - for _, c := range pod.Spec.Containers { - if c.Name != "spdk-proxy-container" { - continue - } - for _, e := range c.Env { - if e.Name != "RPC_PORT" || e.Value == "" { - continue - } - n, err := strconv.ParseInt(e.Value, 10, 32) - if err != nil { - return 0, false - } - return int32(n), true - } - } - - const prefix = "snode-spdk-pod-" - if rest, ok := strings.CutPrefix(pod.Name, prefix); ok { - if dash := strings.Index(rest, "-"); dash > 0 { - if n, err := strconv.ParseInt(rest[:dash], 10, 32); err == nil { - return int32(n), true - } - } - } - return 0, false -} - -func getNodeInternalIP(ctx context.Context, c client.Client, nodeName string) (string, error) { - var node corev1.Node - if err := c.Get(ctx, client.ObjectKey{Name: nodeName}, &node); err != nil { - return "", fmt.Errorf("failed to get node %s: %w", nodeName, err) - } - - for _, addr := range node.Status.Addresses { - if addr.Type == corev1.NodeInternalIP { - return addr.Address, nil - } - } - - return "", fmt.Errorf("node %s has no InternalIP", nodeName) -} - -func checkNodeInfoReachable(ctx context.Context, nodeName, namespace string, tlsEnabled, tlsMutualEnabled bool) error { - scheme := "http" - httpClient := &http.Client{Timeout: 3 * time.Second} - if tlsEnabled { - scheme = "https" - certPath, keyPath := "", "" - if tlsMutualEnabled { - certPath = tlsutil.ServiceClientCertificatePath - keyPath = tlsutil.ServiceClientKeyPath - } - c, err := tlsutil.BuildStorageNodeSetAPIClient(namespace, tlsutil.ServiceCABundlePath, certPath, keyPath) - if err != nil { - return fmt.Errorf("build storage-node TLS client: %w", err) - } - httpClient = c - } - - url := fmt.Sprintf("%s://%s/snode/info", scheme, utils.StorageNodeSetAPIAddress(nodeName, namespace)) - - req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil) - if err != nil { - return err - } - - resp, err := httpClient.Do(req) - if err != nil { - return fmt.Errorf("node info endpoint not reachable: %w", err) - } - defer func() { - if cerr := resp.Body.Close(); cerr != nil { - fmt.Printf("warning: failed to close response body: %v\n", cerr) - } - }() - - if resp.StatusCode != http.StatusOK { - return fmt.Errorf("node info endpoint returned %d", resp.StatusCode) - } - - return nil -} - -func waitForNodeInfoReachable( - ctx context.Context, - nodeName string, - namespace string, //nolint:unparam - tlsEnabled, tlsMutualEnabled bool, //nolint:unparam -) error { - log := logf.FromContext(ctx) - - var lastErr error - - for i := 1; i <= waitForNodeInfoReachableMaxRetries; i++ { - - if err := waitForNodeInfoReachableCheckFn(ctx, nodeName, namespace, tlsEnabled, tlsMutualEnabled); err == nil { - log.Info("Storage node API is reachable", - "node", nodeName, - "attempt", i, - ) - return nil - } else { - lastErr = err - log.V(1).Info("Storage node API not reachable yet, retrying", - "node", nodeName, - "attempt", i, - "error", err.Error(), - ) - } - - select { - case <-time.After(waitForNodeInfoReachableRetryDelay): - case <-ctx.Done(): - return ctx.Err() - } - } - - return fmt.Errorf( - "storage node API not reachable after %d retries: %w", - waitForNodeInfoReachableMaxRetries, - lastErr, - ) -} - -// pollNodeOnline performs a single non-blocking check of whether the node is -// online, returning RequeueAfter if it isn't yet. This replaces the old -// blocking waitForNodeOnline loop so the reconcile worker goroutine stays free. -func (r *StorageNodeSetReconciler) pollNodeOnline( - ctx context.Context, - apiClient *webapi.Client, - clusterUUID, ip, nodeName string, - expectedPerHost int, - snCR *simplyblockv1alpha1.StorageNodeSet, -) (ctrl.Result, error) { - log := logf.FromContext(ctx) - endpoint := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/", clusterUUID) - - body, status, err := apiClient.Do(ctx, http.MethodGet, endpoint, nil) - log.Info("SNODE LIST raw API response", "endpoint", endpoint, "status", status, "body", string(body)) - - if err != nil || status >= 300 { - if err == nil { - err = fmt.Errorf("unexpected status %d", status) - } - log.Error(err, "Failed to get storage node statuses", "node", nodeName, "status", status, "response", string(body)) - return ctrl.Result{RequeueAfter: waitForNodeOnlineWaitInterval}, nil - } - - if strings.TrimSpace(string(body)) == "[]" { - log.Info("Storage node list is empty", "node", nodeName) - return r.nodeOnlineRequeueOrTimeout(ctx, nodeName, ip, snCR) - } - - var apiResp []SNODEAPIResponse - if err := json.Unmarshal(body, &apiResp); err != nil { - return ctrl.Result{}, fmt.Errorf("failed to unmarshal storage node response for %s: %v", nodeName, err) - } - - // Collect all backend nodes for this host IP that are online+healthy. - // When socketsToUse is set the backend creates one node per socket, all - // sharing the same mgmt IP and Hostname — we need all of them online. - onlineForHost := make([]SNODEAPIResponse, 0, expectedPerHost) - for _, res := range apiResp { - if res.IP == ip && res.Status == utils.NodeStatusOnline && res.Health { - onlineForHost = append(onlineForHost, res) - } - } - - if len(onlineForHost) < expectedPerHost { - log.Info("Not all socket nodes online yet", - "node", nodeName, - "online", len(onlineForHost), - "expected", expectedPerHost, - ) - return r.nodeOnlineRequeueOrTimeout(ctx, nodeName, ip, snCR) - } - - // All socket nodes are online — sync status and check cluster activation. - if err := onAllSocketNodesOnline(ctx, apiClient, clusterUUID, snCR, nodeName, onlineForHost, r); err != nil { - return ctrl.Result{}, err - } - log.Info("Storage node created successfully", "node", nodeName) - return ctrl.Result{}, nil -} - -// nodeOnlineRequeueOrTimeout returns RequeueAfter when the node is still -// within the allowed wait window, or marks it as timed-out and returns done. -func (r *StorageNodeSetReconciler) nodeOnlineRequeueOrTimeout( - ctx context.Context, - nodeName, ip string, - snCR *simplyblockv1alpha1.StorageNodeSet, -) (ctrl.Result, error) { - log := logf.FromContext(ctx) - timeout := time.Duration(waitForNodeOnlineRetries) * waitForNodeOnlineWaitInterval - - // Read the post timestamp from the StorageNode CR (set by StorageNodeReconciler). - // Fall back to PendingNodeAdds (legacy) and status.nodes[].PostedAt for - // deployments that pre-date the StorageNodeReconciler. - var postedAt *metav1.Time - if t := r.storageNodePostedAt(ctx, snCR.Namespace, nodeName); t != nil { - postedAt = t - } else if t2, ok := snCR.Status.PendingNodeAdds[nodeName]; ok { - postedAt = &t2 - } else { - for i := range snCR.Status.Nodes { - n := &snCR.Status.Nodes[i] - if n.Hostname == nodeName && n.UUID == "" && n.PostedAt != nil { - postedAt = n.PostedAt - break - } - } - } - - if postedAt != nil { - if time.Since(postedAt.Time) <= timeout { - if time.Since(postedAt.Time) >= spdkPodEventDelay { - r.recordSpdkPodEvents(ctx, snCR, nodeName) - } - return ctrl.Result{RequeueAfter: waitForNodeOnlineWaitInterval}, nil - } - } - - // Timed out (or no post timestamp found — treat as timed-out). - log.Error(nil, "Timeout waiting for node to become online", "node", nodeName) - updated := false - for i := range snCR.Status.Nodes { - if snCR.Status.Nodes[i].Hostname == nodeName { - snCR.Status.Nodes[i].Status = "timeout" - snCR.Status.Nodes[i].MgmtIp = ip - updated = true - } - } - if !updated { - snCR.Status.Nodes = append(snCR.Status.Nodes, simplyblockv1alpha1.NodeStatus{ - Hostname: nodeName, - MgmtIp: ip, - Status: "timeout", - }) - } - if err := r.Status().Update(ctx, snCR); err != nil { - log.Error(err, "Failed to update node status after timeout", "node", nodeName) - } - return ctrl.Result{}, nil -} - -// onAllSocketNodesOnline syncs the StorageNodeSet status entries for all online -// socket nodes and triggers cluster activation when conditions are met. -func onAllSocketNodesOnline( - ctx context.Context, - apiClient *webapi.Client, - clusterUUID string, - snCR *simplyblockv1alpha1.StorageNodeSet, - nodeName string, - onlineForHost []SNODEAPIResponse, - r *StorageNodeSetReconciler, -) error { - log := logf.FromContext(ctx) - - patch := client.MergeFrom(snCR.DeepCopy()) - changed := false - - for _, res := range onlineForHost { - updated := simplyblockv1alpha1.NodeStatus{ - Hostname: nodeName, - UUID: res.UUID, - Health: res.Health, - Status: res.Status, - MgmtIp: res.IP, - Devices: fmt.Sprintf("%d/%d", res.DevicesCount, res.OnlineDevicesCount), - CPU: ptr.To(int32(res.CPU)), - Memory: utils.HumanBytes(res.Memory, "iec"), - Volumes: ptr.To(int32(res.Volumes)), - RpcPort: ptr.To(int32(res.RPC_PORT)), - LvolPort: ptr.To(int32(res.LVOL_PORT)), - NvmfPort: ptr.To(int32(res.NVMF_PORT)), - FailureDomain: fdPtr(res.FailureDomain), - } - - // Try to find existing entry by UUID first, then fall back to the - // placeholder entry (UUID=="") created after the POST. - matched := false - for i := range snCR.Status.Nodes { - n := &snCR.Status.Nodes[i] - if n.Hostname == nodeName && (n.UUID == res.UUID || n.UUID == "") { - if !reflect.DeepEqual(*n, updated) { - *n = updated - changed = true - } - matched = true - break - } - } - if !matched { - snCR.Status.Nodes = append(snCR.Status.Nodes, updated) - changed = true - } - } - - // All socket nodes confirmed online — remove the pending marker so the - // worker is no longer considered in-flight. - if _, ok := snCR.Status.PendingNodeAdds[nodeName]; ok { - delete(snCR.Status.PendingNodeAdds, nodeName) - changed = true - } - // Emit a recovery event only if the worker previously had a scheduling - // failure, then clear the flag. - if snCR.Status.SchedulingFailedWorkers[nodeName] { - r.Recorder.Eventf(snCR, nil, corev1.EventTypeNormal, "NodeOnline", "NodeOnline", - "worker %s: SPDK pod is now online after previous scheduling failure", nodeName) - r.emitOnStorageNodeForWorker(ctx, snCR, nodeName, corev1.EventTypeNormal, "NodeOnline", - fmt.Sprintf("SPDK pod is now online after previous scheduling failure on %s", nodeName)) - delete(snCR.Status.SchedulingFailedWorkers, nodeName) - changed = true - } - if changed { - if err := r.Status().Patch(ctx, snCR, patch); err != nil { - log.Error(err, "Failed to patch node status to online", "node", nodeName) - } - } - - log.Info("All socket nodes online", "node", nodeName, "count", len(onlineForHost)) - - return maybeActivateCluster(ctx, apiClient, clusterUUID, snCR, r) -} - -// syncTrackedNodesStatus refreshes all tracked (UUID != "") NodeStatus entries -// from the backend API. It is called on every completed reconcile pass to keep -// Health, Status, LvolPort and the other fields up-to-date after initial -// provisioning. PostedAt is preserved because it is a creation timestamp. -func (r *StorageNodeSetReconciler) syncTrackedNodesStatus( - ctx context.Context, - apiClient *webapi.Client, - clusterUUID string, - snCR *simplyblockv1alpha1.StorageNodeSet, -) error { - log := logf.FromContext(ctx) - - hasTracked := false - for _, n := range snCR.Status.Nodes { - if n.UUID != "" { - hasTracked = true - break - } - } - if !hasTracked { - return nil - } - - endpoint := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/", clusterUUID) - body, status, err := apiClient.Do(ctx, http.MethodGet, endpoint, nil) - if err != nil || status >= 300 { - if err == nil { - err = fmt.Errorf("unexpected status %d", status) - } - return fmt.Errorf("sync: failed to list storage nodes: %w", err) - } - - var apiResp []SNODEAPIResponse - if err := json.Unmarshal(body, &apiResp); err != nil { - return fmt.Errorf("sync: failed to unmarshal storage node response: %w", err) - } - - byUUID := make(map[string]SNODEAPIResponse, len(apiResp)) - for _, res := range apiResp { - byUUID[res.UUID] = res - } - - patch := client.MergeFrom(snCR.DeepCopy()) - changed := false - - for i := range snCR.Status.Nodes { - n := &snCR.Status.Nodes[i] - if n.UUID == "" { - continue - } - res, ok := byUUID[n.UUID] - if !ok { - continue - } - updated := simplyblockv1alpha1.NodeStatus{ - Hostname: n.Hostname, - UUID: res.UUID, - Health: res.Health, - Status: res.Status, - MgmtIp: res.IP, - Devices: fmt.Sprintf("%d/%d", res.DevicesCount, res.OnlineDevicesCount), - CPU: ptr.To(int32(res.CPU)), - Memory: utils.HumanBytes(res.Memory, "iec"), - Volumes: ptr.To(int32(res.Volumes)), - RpcPort: ptr.To(int32(res.RPC_PORT)), - LvolPort: ptr.To(int32(res.LVOL_PORT)), - NvmfPort: ptr.To(int32(res.NVMF_PORT)), - PostedAt: n.PostedAt, - Uptime: n.Uptime, - FailureDomain: fdPtr(res.FailureDomain), - } - if !reflect.DeepEqual(*n, updated) { - *n = updated - changed = true - } - } - - if changed { - if err := r.Status().Patch(ctx, snCR, patch); err != nil { - log.Error(err, "Failed to patch storage node status during sync") - return err - } - log.Info("Storage node status synced") - } - return nil -} - -// fdPtr converts a failure_domain integer from the backend API into a *int32 -// suitable for status fields. Returns nil only for negative values (unset sentinel). -func fdPtr(fd int) *int32 { - if fd < 0 { - return nil - } - v := int32(fd) - return &v -} - -// maybeActivateCluster activates the cluster when online-node conditions are met. -func maybeActivateCluster( - ctx context.Context, - apiClient *webapi.Client, - clusterUUID string, - snCR *simplyblockv1alpha1.StorageNodeSet, - r *StorageNodeSetReconciler, -) error { - log := logf.FromContext(ctx) - - clusterCR, err := utils.ResolveClusterCR(ctx, r.Client, snCR.Namespace, snCR.Spec.ClusterName) - if err != nil { - log.Info("Cluster not found yet for activation check") - return fmt.Errorf("cluster not found yet") - } - - if utils.ClusterAlreadyActive(clusterCR) { - log.Info("Cluster already active, skipping activation") - return nil - } - - if utils.ClusterInExpansion(clusterCR) { - log.Info("Cluster In expansion, skipping activation") - return nil - } - - onlineHealthy := utils.CountOnlineHealthyNodes(snCR.Status.Nodes) - log.Info("Evaluating cluster activation conditions", - "erasureCodingScheme", clusterCR.Status.ErasureCodingScheme, - "onlineHealthy", onlineHealthy, - ) - - requiredEc, err := utils.RequiredNodesFromErasureCodingScheme(clusterCR.Status.ErasureCodingScheme) - if err != nil { - log.Error(err, "Invalid erasure coding scheme") - return err - } - - if utils.ShouldActivateCluster(requiredEc, onlineHealthy, snCR) { - // Failure-domain readiness gate: mirrors the one in - // StorageClusterOpsReconciler.reconcileActivate - // (controllers/cluster/actions.go). ShouldActivateCluster only - // counts online-healthy nodes against the erasure-coding scheme -- it - // has no notion of failure domains, so without this check a cluster - // with enough nodes but too few/unbalanced FDs gets POSTed to - // /activate every time this reconciler runs, which the backend - // synchronously rejects and reverts (unready -> in_activation -> - // unready), producing a permanent activation-retry loop instead of - // quietly waiting. No-op for FD-disabled clusters. - if ptr.BoolFromOrFalse(clusterCR.Spec.EnableFailureDomains) { - hostDomains, err := clustercontroller.FailureDomainHosts(ctx, r.Client, clusterCR.Namespace, clusterCR.Name) - if err != nil { - log.Error(err, "Failed to list StorageNodeSets for failure-domain readiness check", - "cluster", clusterCR.Name) - return err - } - npcs := clustercontroller.StripeParityChunks(clusterCR.Spec.Stripe) - if reason := clustercontroller.ActivationDomainCountViolation(npcs, hostDomains); reason != "" { - log.Info("Not activating yet, waiting on failure-domain readiness", - "cluster", clusterCR.Name, "reason", reason) - r.Recorder.Eventf(snCR, nil, corev1.EventTypeWarning, "FailureDomainNotReady", "FailureDomainNotReady", - "activation waiting on failure-domain readiness: %s", reason) - return nil - } - } - - if err := waitForNodeOnlineSleepFn(ctx, waitForNodeOnlineActivationDelay); err != nil { - return err - } - log.Info("Activation conditions met — activating cluster") - if err := utils.ActivateClusterAndWait(ctx, apiClient, clusterUUID); err != nil { - log.Error(err, "Cluster activation did not complete") - return err - } - log.Info("Cluster successfully activated") - } - - return nil -} diff --git a/operator/internal/controller/simplyblockstoragenodeset_controller_test.go b/operator/internal/controller/simplyblockstoragenodeset_controller_test.go deleted file mode 100644 index df45beecd..000000000 --- a/operator/internal/controller/simplyblockstoragenodeset_controller_test.go +++ /dev/null @@ -1,31 +0,0 @@ -/* -Copyright 2025. - -Licensed under the Apache License, Version 2.0 (the "License"); -you may not use this file except in compliance with the License. -You may obtain a copy of the License at - - http://www.apache.org/licenses/LICENSE-2.0 - -Unless required by applicable law or agreed to in writing, software -distributed under the License is distributed on an "AS IS" BASIS, -WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. -See the License for the specific language governing permissions and -limitations under the License. -*/ - -package controller - -import ( - . "github.com/onsi/ginkgo/v2" -) - -var _ = Describe("StorageNodeSet Controller", func() { - It("should ignore not-found resources and return no requeue", func() { - controllerReconciler := &StorageNodeSetReconciler{ - Client: k8sClient, - Scheme: k8sClient.Scheme(), - } - expectIgnoreNotFoundNoRequeue(controllerReconciler, "missing-node") - }) -}) diff --git a/operator/internal/controller/simplyblockstoragenodeset_controller_unit_test.go b/operator/internal/controller/simplyblockstoragenodeset_controller_unit_test.go deleted file mode 100644 index c8886f548..000000000 --- a/operator/internal/controller/simplyblockstoragenodeset_controller_unit_test.go +++ /dev/null @@ -1,2844 +0,0 @@ -package controller - -import ( - "context" - "errors" - "fmt" - "net/http" - "net/http/httptest" - "strings" - "testing" - "time" - - "github.com/simplyblock/atlas/kube" - simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" - simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" - "github.com/simplyblock/simplyblock-operator/internal/utils" - "github.com/simplyblock/simplyblock-operator/internal/webapi" - webapimock "github.com/simplyblock/simplyblock-operator/internal/webapi/mock" - appsv1 "k8s.io/api/apps/v1" - corev1 "k8s.io/api/core/v1" - discoveryv1 "k8s.io/api/discovery/v1" - rbacv1 "k8s.io/api/rbac/v1" - apierrors "k8s.io/apimachinery/pkg/api/errors" - metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" - "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" - "k8s.io/client-go/tools/events" - ctrl "sigs.k8s.io/controller-runtime" - "sigs.k8s.io/controller-runtime/pkg/client" - "sigs.k8s.io/controller-runtime/pkg/client/fake" -) - -const ( - statusOnline = "online" - mgmtIP = "10.0.0.1" - tlsVolumeName = "tls" - caVolumeName = "certificate-authority" - nodeUUID1 = "node-uuid-1" -) - -func TestStorageNodeSetFinalizerLifecycleHelpers(t *testing.T) { - now := metav1.NewTime(time.Now()) - - t.Run("ensureFinalizer adds finalizer when missing", func(t *testing.T) { - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{ - Name: "sn-finalizer-add", - Namespace: "default", - }, - } - r := newStorageNodeSetStateTestReconciler(t, sn) - - updated, err := r.ensureFinalizer(context.Background(), sn) - if err != nil { - t.Fatalf("ensureFinalizer returned error: %v", err) - } - if !updated { - t.Fatalf("expected ensureFinalizer to report update") - } - if !contains(sn.Finalizers, utils.FinalizerStorageNodeSet) { - t.Fatalf("expected storagenodeset finalizer to be set") - } - }) - - t.Run("handleDeletion removes finalizer when deletion timestamp is set", func(t *testing.T) { - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{ - Name: "sn-finalizer-del", - Namespace: "default", - Finalizers: []string{utils.FinalizerStorageNodeSet}, - DeletionTimestamp: &now, - }, - } - r := newStorageNodeSetStateTestReconciler(t, sn) - - updated, err := r.handleDeletion(context.Background(), sn) - if err != nil { - t.Fatalf("handleDeletion returned error: %v", err) - } - if !updated { - t.Fatalf("expected handleDeletion to report update") - } - if contains(sn.Finalizers, utils.FinalizerStorageNodeSet) { - t.Fatalf("expected storagenodeset finalizer to be removed") - } - }) -} - -func TestStorageNodeSetLabelingHelpers(t *testing.T) { - t.Run("labelWorkerNodes labels all configured workers", func(t *testing.T) { - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{ - Name: "sn-label-all", - Namespace: "default", - }, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - ClusterName: "cluster-a", - WorkerNodes: []string{"node-a", "node-b"}, - }, - } - nodeA := &corev1.Node{ObjectMeta: metav1.ObjectMeta{Name: "node-a"}} - nodeB := &corev1.Node{ObjectMeta: metav1.ObjectMeta{Name: "node-b"}} - r := newStorageNodeUUIDLabelTestReconciler(t, sn, nodeA, nodeB) - - if err := labelWorkerNodes(context.Background(), r.Client, r.Recorder, sn, testClusterUUID); err != nil { - t.Fatalf("labelWorkerNodes returned error: %v", err) - } - - for _, nodeName := range []string{"node-a", "node-b"} { - var n corev1.Node - if err := r.Get(context.Background(), client.ObjectKey{Name: nodeName}, &n); err != nil { - t.Fatalf("failed to fetch node %s: %v", nodeName, err) - } - got := n.Labels[kube.LabelStorageNodeSet] - want := sn.Name - if got != want { - t.Fatalf("node %s label mismatch: got %q want %q", nodeName, got, want) - } - // io.simplyblock.node-type marked exactly the nodes the per-set label - // marks, so it was retired. Without this nothing fails if it returns. - if value, ok := n.Labels["io.simplyblock.node-type"]; ok { - t.Fatalf("node %s carries the retired node-type label with value %q", nodeName, value) - } - } - }) - - t.Run("labelWorkerNodes records an event when a worker node is missing", func(t *testing.T) { - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{ - Name: "sn-label-missing", - Namespace: "default", - }, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - ClusterName: "cluster-a", - WorkerNodes: []string{"vm12.simplyblock4.localdomain"}, - }, - } - r := newStorageNodeUUIDLabelTestReconciler(t, sn) - - if err := labelWorkerNodes(context.Background(), r.Client, r.Recorder, sn, testClusterUUID); err == nil { - t.Fatalf("expected labelWorkerNodes to return an error for a missing node") - } - - recorder := r.Recorder.(*events.FakeRecorder) - select { - case event := <-recorder.Events: - if !strings.Contains(event, "WorkerNodeNotFound") || !strings.Contains(event, "vm12.simplyblock4.localdomain") { - t.Fatalf("unexpected event content: %q", event) - } - default: - t.Fatalf("expected a WorkerNodeNotFound event to be recorded") - } - }) - - t.Run("labelWorkerNodes labels a single co-located storage-node UUID", func(t *testing.T) { - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-label-uuid", Namespace: "default"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - ClusterName: "cluster-a", - WorkerNodes: []string{"node-a"}, - }, - } - socketIndex := int32(0) - storageNode := &simplyblockv1alpha1.StorageNode{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-label-uuid-0", Namespace: "default"}, - Spec: simplyblockv1alpha1.StorageNodeSpec{ - StorageNodeSetRef: "sn-label-uuid", - WorkerNode: "node-a", - SocketIndex: &socketIndex, - }, - Status: simplyblockv1alpha1.StorageNodeStatus{UUID: "sock0-storage-node-uuid"}, - } - nodeA := &corev1.Node{ObjectMeta: metav1.ObjectMeta{Name: "node-a"}} - r := newStorageNodeUUIDLabelTestReconciler(t, sn, storageNode, nodeA) - - if err := labelWorkerNodes(context.Background(), r.Client, r.Recorder, sn, testClusterUUID); err != nil { - t.Fatalf("labelWorkerNodes returned error: %v", err) - } - - var n corev1.Node - if err := r.Get(context.Background(), client.ObjectKey{Name: "node-a"}, &n); err != nil { - t.Fatalf("failed to fetch node: %v", err) - } - got := n.Labels[storageNodeUUIDLabelPrefix+testClusterUUID+".0"] - want := "sock0-storage-node-uuid" - if got != want { - t.Fatalf("storage-node-uuid label mismatch: got %q want %q", got, want) - } - }) - - t.Run("labelWorkerNodes labels every socket instance on a multi-socket worker", func(t *testing.T) { - const snsName = "sn-label-multi" - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: snsName, Namespace: "default"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - ClusterName: "cluster-a", - WorkerNodes: []string{"node-a"}, - }, - } - socket0, socket1 := int32(0), int32(1) - snSocket0 := &simplyblockv1alpha1.StorageNode{ - ObjectMeta: metav1.ObjectMeta{Name: snsName + "-0", Namespace: "default"}, - Spec: simplyblockv1alpha1.StorageNodeSpec{ - StorageNodeSetRef: snsName, - WorkerNode: "node-a", - SocketIndex: &socket0, - }, - Status: simplyblockv1alpha1.StorageNodeStatus{UUID: "uuid-socket-0"}, - } - snSocket1 := &simplyblockv1alpha1.StorageNode{ - ObjectMeta: metav1.ObjectMeta{Name: snsName + "-1", Namespace: "default"}, - Spec: simplyblockv1alpha1.StorageNodeSpec{ - StorageNodeSetRef: snsName, - WorkerNode: "node-a", - SocketIndex: &socket1, - }, - Status: simplyblockv1alpha1.StorageNodeStatus{UUID: "uuid-socket-1"}, - } - nodeA := &corev1.Node{ObjectMeta: metav1.ObjectMeta{Name: "node-a"}} - r := newStorageNodeUUIDLabelTestReconciler(t, sn, snSocket0, snSocket1, nodeA) - - if err := labelWorkerNodes(context.Background(), r.Client, r.Recorder, sn, testClusterUUID); err != nil { - t.Fatalf("labelWorkerNodes returned error: %v", err) - } - - var n corev1.Node - if err := r.Get(context.Background(), client.ObjectKey{Name: "node-a"}, &n); err != nil { - t.Fatalf("failed to fetch node: %v", err) - } - if got, want := n.Labels[storageNodeUUIDLabelPrefix+testClusterUUID+".0"], "uuid-socket-0"; got != want { - t.Fatalf("socket-0 label mismatch: got %q want %q", got, want) - } - if got, want := n.Labels[storageNodeUUIDLabelPrefix+testClusterUUID+".1"], "uuid-socket-1"; got != want { - t.Fatalf("socket-1 label mismatch: got %q want %q", got, want) - } - }) - - t.Run("labelWorkerNodes updates the UUID value in place when a slot's storage node is replaced", func(t *testing.T) { - // The label KEY (".") must stay stable - // across a UUID replacement — Kubernetes' external-provisioner caches the - // set of topology KEYS in CSINode and hard-errors CreateVolume if a live - // Node's label keys ever diverge from that cached set. Only the UUID - // value is expected to churn. - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-label-replace", Namespace: "default"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - ClusterName: "cluster-a", - WorkerNodes: []string{"node-a"}, - }, - } - socketIndex := int32(0) - storageNode := &simplyblockv1alpha1.StorageNode{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-label-replace-0", Namespace: "default"}, - Spec: simplyblockv1alpha1.StorageNodeSpec{ - StorageNodeSetRef: "sn-label-replace", - WorkerNode: "node-a", - SocketIndex: &socketIndex, - }, - Status: simplyblockv1alpha1.StorageNodeStatus{UUID: "uuid-new"}, - } - slotKey := storageNodeUUIDLabelPrefix + testClusterUUID + ".0" - nodeA := &corev1.Node{ - ObjectMeta: metav1.ObjectMeta{ - Name: "node-a", - Labels: map[string]string{slotKey: "uuid-old"}, - }, - } - r := newStorageNodeUUIDLabelTestReconciler(t, sn, storageNode, nodeA) - - if err := labelWorkerNodes(context.Background(), r.Client, r.Recorder, sn, testClusterUUID); err != nil { - t.Fatalf("labelWorkerNodes returned error: %v", err) - } - - var n corev1.Node - if err := r.Get(context.Background(), client.ObjectKey{Name: "node-a"}, &n); err != nil { - t.Fatalf("failed to fetch node: %v", err) - } - if got, want := n.Labels[slotKey], "uuid-new"; got != want { - t.Fatalf("slot label value mismatch: got %q want %q", got, want) - } - }) - - t.Run("labelWorkerNodes removes a slot label when its StorageNode CR no longer exists", func(t *testing.T) { - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-label-stale", Namespace: "default"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - ClusterName: "cluster-a", - WorkerNodes: []string{"node-a"}, - }, - } - // node-a still carries a slot label for a storage node that no longer - // exists (removed, not replaced) — it must be cleaned up, not just left - // to accumulate. - nodeA := &corev1.Node{ - ObjectMeta: metav1.ObjectMeta{ - Name: "node-a", - Labels: map[string]string{ - storageNodeUUIDLabelPrefix + testClusterUUID + ".0": "uuid-old", - }, - }, - } - r := newStorageNodeUUIDLabelTestReconciler(t, sn, nodeA) - - if err := labelWorkerNodes(context.Background(), r.Client, r.Recorder, sn, testClusterUUID); err != nil { - t.Fatalf("labelWorkerNodes returned error: %v", err) - } - - var n corev1.Node - if err := r.Get(context.Background(), client.ObjectKey{Name: "node-a"}, &n); err != nil { - t.Fatalf("failed to fetch node: %v", err) - } - if _, ok := n.Labels[storageNodeUUIDLabelPrefix+testClusterUUID+".0"]; ok { - t.Fatalf("expected stale storage-node-uuid label to be removed, got labels: %v", n.Labels) - } - }) -} - -// newStorageNodeUUIDLabelTestReconciler builds a fake-client-backed reconciler -// with the "spec.storageNodeSetRef" index registered, which labelWorkerNodes' -// StorageNode list requires. newStorageNodeSetStateTestReconciler does not -// register it, and a List against an unindexed field errors; because -// labelWorkerNodes now returns that error rather than swallowing it, any test -// that drives labelWorkerNodes must use this builder (both to see StorageNode -// objects for the storage-node-uuid path and to avoid a spurious List error). -func newStorageNodeUUIDLabelTestReconciler(t *testing.T, objects ...client.Object) *StorageNodeSetReconciler { - t.Helper() - scheme := newTestScheme(t, simplyblockv1alpha1.AddToScheme, corev1.AddToScheme) - cl := fake.NewClientBuilder(). - WithScheme(scheme). - WithStatusSubresource(&simplyblockv1alpha1.StorageNode{}, &simplyblockv1alpha1.StorageNodeSet{}). - WithObjects(objects...). - WithIndex(&simplyblockv1alpha1.StorageNode{}, "spec.storageNodeSetRef", func(obj client.Object) []string { - sn := obj.(*simplyblockv1alpha1.StorageNode) - return []string{sn.Spec.StorageNodeSetRef} - }). - Build() - return &StorageNodeSetReconciler{ - Client: cl, - Scheme: scheme, - Recorder: events.NewFakeRecorder(16), - } -} - -func TestStorageNodeSetDaemonSetReconcileCreatesWhenMissing(t *testing.T) { - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{ - Name: "sn-ds-create", - Namespace: "default", - UID: "uid-create", - }, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ClusterName: "cluster-a"}, - } - r := newStorageNodeSetStateTestReconciler(t, sn) - - if err := r.reconcileDaemonSet(context.Background(), sn); err != nil { - t.Fatalf("reconcileDaemonSet returned error: %v", err) - } - - var ds appsv1.DaemonSet - if err := r.Get(context.Background(), client.ObjectKey{Name: "simplyblock-storage-node-ds-sn-ds-create", Namespace: "default"}, &ds); err != nil { - t.Fatalf("daemonset should be created: %v", err) - } - if len(ds.OwnerReferences) == 0 || ds.OwnerReferences[0].Name != sn.Name { - t.Fatalf("expected daemonset to be owned by storagenodeset") - } -} - -// Regression: 2026-09-11-nodeset-image-fallback-reads-retired-version. A -// StorageNodeSet that names no clusterImage takes the image from the singleton -// ControlPlane, and that read stayed on v1alpha1 after v1alpha2 became the -// stored version. The API server answers a v1alpha1 read only through the -// conversion webhook, which a fresh install does not deploy, so the fallback -// returned NotFound and no DaemonSet was ever created. -func TestStorageNodeSetDaemonSetImageFallsBackToControlPlane(t *testing.T) { - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{ - Name: "sn-ds-fallback", - Namespace: "default", - UID: "uid-fallback", - }, - // No ClusterImage: the ControlPlane is what supplies it. - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ClusterName: "cluster-a"}, - } - r := newStorageNodeSetStateTestReconciler(t, sn) - - if err := r.reconcileDaemonSet(context.Background(), sn); err != nil { - t.Fatalf("reconcileDaemonSet returned error: %v", err) - } - - var ds appsv1.DaemonSet - if err := r.Get(context.Background(), client.ObjectKey{ - Name: "simplyblock-storage-node-ds-sn-ds-fallback", Namespace: "default", - }, &ds); err != nil { - t.Fatalf("daemonset should be created: %v", err) - } - for _, c := range ds.Spec.Template.Spec.Containers { - if c.Image != "test-image:latest" { - t.Fatalf("container %q: expected the ControlPlane image, got %q", c.Name, c.Image) - } - } -} - -func TestStorageNodeSetDaemonSetReconcileUpdatesExisting(t *testing.T) { - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{ - Name: "sn-ds-update", - Namespace: "default", - UID: "uid-update", - }, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ClusterName: "cluster-a"}, - } - existing := &appsv1.DaemonSet{ - ObjectMeta: metav1.ObjectMeta{ - Name: "simplyblock-storage-node-ds-sn-ds-update", - Namespace: "default", - }, - } - r := newStorageNodeSetStateTestReconciler(t, sn, existing) - - if err := r.reconcileDaemonSet(context.Background(), sn); err != nil { - t.Fatalf("reconcileDaemonSet returned error: %v", err) - } - - var ds appsv1.DaemonSet - if err := r.Get(context.Background(), client.ObjectKey{Name: "simplyblock-storage-node-ds-sn-ds-update", Namespace: "default"}, &ds); err != nil { - t.Fatalf("failed to fetch daemonset: %v", err) - } - if len(ds.OwnerReferences) == 0 || ds.OwnerReferences[0].Name != sn.Name { - t.Fatalf("expected updated daemonset to carry owner reference") - } -} - -func TestStorageNodeSetDaemonSetReconcileTLSDisabled(t *testing.T) { - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-ds-tls-off", Namespace: "default", UID: "uid-tls-off"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ClusterName: "cluster-a"}, - } - r := newStorageNodeSetStateTestReconciler(t, sn) - r.TLSEnabled = false - - if err := r.reconcileDaemonSet(context.Background(), sn); err != nil { - t.Fatalf("reconcileDaemonSet returned error: %v", err) - } - - var ds appsv1.DaemonSet - if err := r.Get(context.Background(), client.ObjectKey{Name: "simplyblock-storage-node-ds-sn-ds-tls-off", Namespace: "default"}, &ds); err != nil { - t.Fatalf("failed to fetch daemonset: %v", err) - } - for _, v := range ds.Spec.Template.Spec.Volumes { - if v.Name == tlsVolumeName || v.Name == caVolumeName { - t.Fatalf("unexpected TLS volume present: %s", v.Name) - } - } - for _, c := range ds.Spec.Template.Spec.InitContainers { - for _, m := range c.VolumeMounts { - if m.Name == tlsVolumeName || m.Name == caVolumeName { - t.Fatalf("unexpected TLS mount on init container: %s", m.Name) - } - } - } - for _, c := range ds.Spec.Template.Spec.Containers { - for _, m := range c.VolumeMounts { - if m.Name == tlsVolumeName || m.Name == caVolumeName { - t.Fatalf("unexpected TLS mount on main container: %s", m.Name) - } - } - if c.ReadinessProbe == nil || c.ReadinessProbe.HTTPGet == nil { - t.Fatalf("expected HTTPGet readiness probe") - } - if c.ReadinessProbe.HTTPGet.Scheme != "" && c.ReadinessProbe.HTTPGet.Scheme != corev1.URISchemeHTTP { - t.Fatalf("expected default/HTTP probe scheme when TLS disabled, got %q", c.ReadinessProbe.HTTPGet.Scheme) - } - if _, ok := envValue(c.Env, "SB_TLS_CONNECT"); ok { - t.Fatalf("unexpected SB_TLS_CONNECT env when TLS disabled") - } - } -} - -func checkTLSMounts(t *testing.T, label string, mounts []corev1.VolumeMount) { - t.Helper() - var gotTLS bool - for _, m := range mounts { - switch m.Name { - case tlsVolumeName: - gotTLS = true - if m.MountPath != "/etc/simplyblock/tls" || m.SubPath != "" || !m.ReadOnly { - t.Fatalf("%s: tls mount shape wrong: %#v", label, m) - } - case caVolumeName: - t.Fatalf("%s: unexpected separate certificate-authority mount: %#v", label, m) - } - } - if !gotTLS { - t.Fatalf("%s: expected tls mount", label) - } -} - -func envValue(env []corev1.EnvVar, name string) (string, bool) { - for _, item := range env { - if item.Name == name { - return item.Value, true - } - } - return "", false -} - -func TestStorageNodeSetDaemonSetReconcileTLSEnabled(t *testing.T) { - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-ds-tls-on", Namespace: "default", UID: "uid-tls-on"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ClusterName: "cluster-a"}, - } - r := newStorageNodeSetStateTestReconciler(t, sn) - r.TLSEnabled = true - r.TLSProvider = utils.TLSProviderOpenShift - r.TLSMutualEnabled = true - - if err := r.reconcileDaemonSet(context.Background(), sn); err != nil { - t.Fatalf("reconcileDaemonSet returned error: %v", err) - } - - var ds appsv1.DaemonSet - if err := r.Get(context.Background(), client.ObjectKey{Name: "simplyblock-storage-node-ds-sn-ds-tls-on", Namespace: "default"}, &ds); err != nil { - t.Fatalf("failed to fetch daemonset: %v", err) - } - - var tlsVol *corev1.Volume - for i := range ds.Spec.Template.Spec.Volumes { - v := &ds.Spec.Template.Spec.Volumes[i] - switch v.Name { - case tlsVolumeName: - tlsVol = v - case caVolumeName: - t.Fatalf("unexpected separate certificate-authority volume: %#v", v) - } - } - if tlsVol == nil || tlsVol.Projected == nil { - t.Fatalf("expected projected tls volume, got %#v", tlsVol) - } - var gotSecret, gotCA bool - for _, src := range tlsVol.Projected.Sources { - switch { - case src.Secret != nil && src.Secret.Name == "simplyblock-storage-node-api-tls": - gotSecret = true - case src.ConfigMap != nil && src.ConfigMap.Name == "simplyblock-certificate-authority": - gotCA = true - if len(src.ConfigMap.Items) != 1 || src.ConfigMap.Items[0].Key != "service-ca.crt" || src.ConfigMap.Items[0].Path != "ca.crt" { - t.Fatalf("ca configmap projection wrong: %#v", src.ConfigMap.Items) - } - } - } - if !gotSecret || !gotCA { - t.Fatalf("expected projected sources for secret and ca configmap, got secret=%v ca=%v", gotSecret, gotCA) - } - - if len(ds.Spec.Template.Spec.InitContainers) != 2 { - t.Fatalf("expected 2 init containers (node-env-writer + s-node-api-config-generator)") - } - checkTLSMounts(t, "init container", ds.Spec.Template.Spec.InitContainers[1].VolumeMounts) - if len(ds.Spec.Template.Spec.Containers) != 1 { - t.Fatalf("expected single main container") - } - checkTLSMounts(t, "main container", ds.Spec.Template.Spec.Containers[0].VolumeMounts) - - probe := ds.Spec.Template.Spec.Containers[0].ReadinessProbe - if probe == nil || probe.TCPSocket == nil { - t.Fatalf("expected TCPSocket readiness probe under mutual TLS, got %#v", probe) - } - if probe.HTTPGet != nil { - t.Fatalf("did not expect HTTPGet readiness probe under mutual TLS, got %#v", probe.HTTPGet) - } - if got, ok := envValue(ds.Spec.Template.Spec.Containers[0].Env, "SB_TLS_CONNECT"); !ok || got != "authenticated" { - t.Fatalf("expected SB_TLS_CONNECT=authenticated on main container, got value=%q present=%v", got, ok) - } -} - -func TestStorageNodeSetDaemonSetReconcileTLSCertManagerProvider(t *testing.T) { - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-ds-tls-cert-manager", Namespace: "default", UID: "uid-tls-cert-manager"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ClusterName: "cluster-a"}, - } - r := newStorageNodeSetStateTestReconciler(t, sn) - r.TLSEnabled = true - r.TLSProvider = utils.TLSProviderCertManager - r.TLSMutualEnabled = false - - if err := r.reconcileDaemonSet(context.Background(), sn); err != nil { - t.Fatalf("reconcileDaemonSet returned error: %v", err) - } - - var ds appsv1.DaemonSet - if err := r.Get(context.Background(), client.ObjectKey{Name: "simplyblock-storage-node-ds-sn-ds-tls-cert-manager", Namespace: "default"}, &ds); err != nil { - t.Fatalf("failed to fetch daemonset: %v", err) - } - - var tlsVol *corev1.Volume - for i := range ds.Spec.Template.Spec.Volumes { - v := &ds.Spec.Template.Spec.Volumes[i] - switch v.Name { - case tlsVolumeName: - tlsVol = v - case caVolumeName: - t.Fatalf("unexpected separate certificate-authority volume: %#v", v) - } - } - if tlsVol == nil { - t.Fatalf("expected tls volume, got none") - return - } - if tlsVol.Projected != nil { - t.Fatalf("expected plain Secret volume for cert-manager provider, got projected: %#v", tlsVol.Projected) - } - if tlsVol.Secret == nil || tlsVol.Secret.SecretName != "simplyblock-storage-node-api-tls" { - t.Fatalf("expected Secret volume referencing simplyblock-storage-node-api-tls, got %#v", tlsVol.Secret) - } - - if len(ds.Spec.Template.Spec.InitContainers) != 2 { - t.Fatalf("expected 2 init containers (node-env-writer + s-node-api-config-generator)") - } - checkTLSMounts(t, "init container", ds.Spec.Template.Spec.InitContainers[1].VolumeMounts) - if len(ds.Spec.Template.Spec.Containers) != 1 { - t.Fatalf("expected single main container") - } - checkTLSMounts(t, "main container", ds.Spec.Template.Spec.Containers[0].VolumeMounts) - - probe := ds.Spec.Template.Spec.Containers[0].ReadinessProbe - if probe == nil || probe.HTTPGet == nil { - t.Fatalf("expected HTTPGet readiness probe under server-only TLS, got %#v", probe) - } - if probe.HTTPGet.Scheme != corev1.URISchemeHTTPS { - t.Fatalf("expected readiness probe scheme HTTPS, got %q", probe.HTTPGet.Scheme) - } -} - -func TestGetNodeInternalIP(t *testing.T) { - node := &corev1.Node{ - ObjectMeta: metav1.ObjectMeta{Name: "node-ip"}, - Status: corev1.NodeStatus{ - Addresses: []corev1.NodeAddress{ - {Type: corev1.NodeHostName, Address: "node-ip"}, - {Type: corev1.NodeInternalIP, Address: "10.1.2.3"}, - }, - }, - } - r := newStorageNodeSetStateTestReconciler(t, node) - - got, err := getNodeInternalIP(context.Background(), r.Client, "node-ip") - if err != nil { - t.Fatalf("getNodeInternalIP returned error: %v", err) - } - if got != "10.1.2.3" { - t.Fatalf("expected internal IP 10.1.2.3, got %q", got) - } -} - -func TestGetNodeInternalIPNoAddress(t *testing.T) { - node := &corev1.Node{ - ObjectMeta: metav1.ObjectMeta{Name: "node-no-ip"}, - } - r := newStorageNodeSetStateTestReconciler(t, node) - - _, err := getNodeInternalIP(context.Background(), r.Client, "node-no-ip") - if err == nil { - t.Fatalf("expected error when node has no internal IP") - } -} - -func TestStorageNodeSetHandleDeletionNoopWithoutDeletionTimestamp(t *testing.T) { - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{ - Name: "sn-no-delete", - Namespace: "default", - }, - } - r := newStorageNodeSetStateTestReconciler(t, sn) - - updated, err := r.handleDeletion(context.Background(), sn) - if err != nil { - t.Fatalf("handleDeletion returned error: %v", err) - } - if updated { - t.Fatalf("expected no update when deletion timestamp is zero") - } -} - -func TestStorageNodeSetHandleDeletionDoneWithoutFinalizer(t *testing.T) { - now := metav1.NewTime(time.Now()) - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{ - Name: "sn-delete-done", - Namespace: "default", - DeletionTimestamp: &now, - }, - } - r := newStorageNodeSetStateTestReconciler(t) - - updated, err := r.handleDeletion(context.Background(), sn) - if err != nil { - t.Fatalf("handleDeletion returned error: %v", err) - } - if !updated { - t.Fatalf("expected deletion flow to be treated as handled without finalizer") - } -} - -func TestStorageNodeSetReconcileClusterUnavailableRequeues(t *testing.T) { - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{ - Name: "sn-reconcile-no-cluster", - Namespace: "default", - }, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ClusterName: "cluster-missing"}, - } - r := newStorageNodeSetStateTestReconciler(t, sn) - - res, err := r.Reconcile(context.Background(), ctrl.Request{NamespacedName: client.ObjectKeyFromObject(sn)}) - if err != nil { - t.Fatalf("reconcile returned error: %v", err) - } - if res.RequeueAfter == 0 { - t.Fatalf("expected delayed requeue when cluster UUID is unavailable") - } -} - -func TestStorageNodeSetReconcileWithClusterUUIDProceeds(t *testing.T) { - // With SA-token auth, the cluster secret is no longer required. - // Reconcile should proceed (not requeue waiting for a secret) when - // the cluster UUID is available. - cluster := &simplyblockv1alpha2.StorageCluster{ - ObjectMeta: metav1.ObjectMeta{Name: "cluster-a", Namespace: "default"}, - Spec: simplyblockv1alpha2.StorageClusterSpec{}, - Status: simplyblockv1alpha2.StorageClusterStatus{UUID: "cluster-uuid-no-secret"}, - } - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{ - Name: "sn-reconcile-no-secret", - Namespace: "default", - }, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ClusterName: "cluster-a"}, - } - r := newStorageNodeSetStateTestReconciler(t, sn, cluster) - - _, err := r.Reconcile(context.Background(), ctrl.Request{NamespacedName: client.ObjectKeyFromObject(sn)}) - if err != nil { - t.Fatalf("reconcile returned unexpected error: %v", err) - } - // No assertion on RequeueAfter — the reconciler may requeue for other - // reasons (e.g., waiting for nodes to join), but it must not error. -} - -func TestStorageNodeSetReconcileNotFoundReturnsNil(t *testing.T) { - r := newStorageNodeSetStateTestReconciler(t) - - res, err := r.Reconcile(context.Background(), ctrl.Request{ - NamespacedName: client.ObjectKey{Name: "missing", Namespace: "default"}, - }) - if err != nil { - t.Fatalf("reconcile returned unexpected error: %v", err) - } - if res.RequeueAfter != 0 { - t.Fatalf("expected no requeue for missing object, got %+v", res) - } -} - -func TestStorageNodeSetReconcileDeletionFlow(t *testing.T) { - const namespace = "default" - const clusterName = "cluster-del" - const clusterUUID = "cluster-uuid-del" - now := metav1.NewTime(time.Now()) - - cluster := &simplyblockv1alpha2.StorageCluster{ - ObjectMeta: metav1.ObjectMeta{Name: clusterName, Namespace: namespace}, - Spec: simplyblockv1alpha2.StorageClusterSpec{}, - Status: simplyblockv1alpha2.StorageClusterStatus{UUID: clusterUUID}, - } - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{ - Name: "sn-delete-flow", - Namespace: namespace, - Finalizers: []string{utils.FinalizerStorageNodeSet}, - DeletionTimestamp: &now, - }, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ClusterName: clusterName}, - } - - r := newStorageNodeSetStateTestReconciler(t, sn, cluster) - res, err := r.Reconcile(context.Background(), ctrl.Request{NamespacedName: client.ObjectKeyFromObject(sn)}) - if err != nil { - t.Fatalf("reconcile returned error: %v", err) - } - if res.RequeueAfter != 0 { - t.Fatalf("expected deletion flow to complete without requeue, got %+v", res) - } - - current := &simplyblockv1alpha1.StorageNodeSet{} - if err := r.Get(context.Background(), client.ObjectKeyFromObject(sn), current); err != nil { - if !apierrors.IsNotFound(err) { - t.Fatalf("failed to fetch storagenodeset: %v", err) - } - return - } - if contains(current.Finalizers, utils.FinalizerStorageNodeSet) { - t.Fatalf("expected finalizer to be removed during deletion flow") - } -} - -func TestStorageNodeSetReconcileAddsFinalizer(t *testing.T) { - const namespace = "default" - const clusterName = "cluster-finalizer" - const clusterUUID = "cluster-uuid-finalizer" - - cluster := &simplyblockv1alpha2.StorageCluster{ - ObjectMeta: metav1.ObjectMeta{Name: clusterName, Namespace: namespace}, - Spec: simplyblockv1alpha2.StorageClusterSpec{}, - Status: simplyblockv1alpha2.StorageClusterStatus{UUID: clusterUUID}, - } - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{ - Name: "sn-finalizer-flow", - Namespace: namespace, - }, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ClusterName: clusterName}, - } - - r := newStorageNodeSetStateTestReconciler(t, sn, cluster) - res, err := r.Reconcile(context.Background(), ctrl.Request{NamespacedName: client.ObjectKeyFromObject(sn)}) - if err != nil { - t.Fatalf("reconcile returned error: %v", err) - } - if res.RequeueAfter != 0 { - t.Fatalf("expected finalizer add path to return without requeue, got %+v", res) - } - - current := &simplyblockv1alpha1.StorageNodeSet{} - if err := r.Get(context.Background(), client.ObjectKeyFromObject(sn), current); err != nil { - t.Fatalf("failed to fetch storagenodeset: %v", err) - } - if !contains(current.Finalizers, utils.FinalizerStorageNodeSet) { - t.Fatalf("expected finalizer to be added by reconcile") - } -} - -func TestStorageNodeSetReconcileLabelWorkerNodesFailure(t *testing.T) { - const namespace = "default" - const clusterName = "cluster-label-fail" - const clusterUUID = "cluster-uuid-label-fail" - - cluster := &simplyblockv1alpha2.StorageCluster{ - ObjectMeta: metav1.ObjectMeta{Name: clusterName, Namespace: namespace}, - Spec: simplyblockv1alpha2.StorageClusterSpec{}, - Status: simplyblockv1alpha2.StorageClusterStatus{UUID: clusterUUID}, - } - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{ - Name: "sn-label-fail", - Namespace: namespace, - Finalizers: []string{utils.FinalizerStorageNodeSet}, - }, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - ClusterName: clusterName, - WorkerNodes: []string{"missing-worker"}, - }, - } - - r := newStorageNodeSetStateTestReconciler(t, sn, cluster) - _, err := r.Reconcile(context.Background(), ctrl.Request{NamespacedName: client.ObjectKeyFromObject(sn)}) - if err == nil { - t.Fatalf("expected reconcile to fail when worker node lookup fails") - } -} - -func TestStorageNodeSetReconcileKnownWorkerSkipsProvisioning(t *testing.T) { - const namespace = "default" - const clusterName = "cluster-known-worker" - const clusterUUID = "cluster-uuid-known-worker" - const workerName = "node-known" - - cluster := &simplyblockv1alpha2.StorageCluster{ - ObjectMeta: metav1.ObjectMeta{Name: clusterName, Namespace: namespace}, - Spec: simplyblockv1alpha2.StorageClusterSpec{}, - Status: simplyblockv1alpha2.StorageClusterStatus{UUID: clusterUUID}, - } - node := &corev1.Node{ - ObjectMeta: metav1.ObjectMeta{Name: workerName}, - } - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{ - Name: "sn-known-worker", - Namespace: namespace, - Finalizers: []string{utils.FinalizerStorageNodeSet}, - }, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - ClusterName: clusterName, - WorkerNodes: []string{workerName}, - }, - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - Nodes: []simplyblockv1alpha1.NodeStatus{ - { - Hostname: workerName, - MgmtIp: "10.0.0.10", - Status: statusOnline, - UUID: "node-uuid-known", - }, - }, - }, - } - - r := newStorageNodeSetStateTestReconciler(t, sn, cluster, node) - res, err := r.Reconcile(context.Background(), ctrl.Request{NamespacedName: client.ObjectKeyFromObject(sn)}) - if err != nil { - t.Fatalf("reconcile returned error: %v", err) - } - if res.RequeueAfter != syncNodeStatusInterval { - t.Fatalf("expected requeue after %v for status sync, got %+v", syncNodeStatusInterval, res) - } -} - -func TestStorageNodeSetReconcileServiceAccountHasOwnerReference(t *testing.T) { - const namespace = "default" - const clusterName = "cluster-ownerref-sa" - const clusterUUID = "cluster-uuid-ownerref-sa" - - cluster := &simplyblockv1alpha2.StorageCluster{ - ObjectMeta: metav1.ObjectMeta{Name: clusterName, Namespace: namespace}, - Spec: simplyblockv1alpha2.StorageClusterSpec{}, - Status: simplyblockv1alpha2.StorageClusterStatus{UUID: clusterUUID}, - } - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{ - Name: "sn-ownerref-sa", - Namespace: namespace, - Finalizers: []string{utils.FinalizerStorageNodeSet}, - }, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - ClusterName: clusterName, - WorkerNodes: []string{}, - }, - } - - r := newStorageNodeSetStateTestReconciler(t, sn, cluster) - _, err := r.Reconcile(context.Background(), ctrl.Request{NamespacedName: client.ObjectKeyFromObject(sn)}) - if err != nil { - t.Fatalf("reconcile returned error: %v", err) - } - - sa := &corev1.ServiceAccount{} - if err := r.Get(context.Background(), client.ObjectKey{ - Name: "simplyblock-storage-node-sa", - Namespace: namespace, - }, sa); err != nil { - t.Fatalf("failed to fetch serviceaccount: %v", err) - } - - if len(sa.OwnerReferences) == 0 { - t.Fatalf("expected ServiceAccount to carry ownerReference to storagenodeset CR") - } -} - -func TestStorageNodeSetReconcileCreatesNamespaceSpecificClusterRoleBindings(t *testing.T) { - const clusterUUID1 = "cluster-uuid-one" - const clusterUUID2 = "cluster-uuid-two" - - cluster1 := &simplyblockv1alpha2.StorageCluster{ - ObjectMeta: metav1.ObjectMeta{Name: "cluster1", Namespace: "cluster1"}, - Spec: simplyblockv1alpha2.StorageClusterSpec{}, - Status: simplyblockv1alpha2.StorageClusterStatus{UUID: clusterUUID1}, - } - cluster2 := &simplyblockv1alpha2.StorageCluster{ - ObjectMeta: metav1.ObjectMeta{Name: "cluster2", Namespace: "cluster2"}, - Spec: simplyblockv1alpha2.StorageClusterSpec{}, - Status: simplyblockv1alpha2.StorageClusterStatus{UUID: clusterUUID2}, - } - sn1 := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{ - Name: "sn-cluster1", - Namespace: "cluster1", - Finalizers: []string{utils.FinalizerStorageNodeSet}, - }, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ClusterName: "cluster1"}, - } - sn2 := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{ - Name: "sn-cluster2", - Namespace: "cluster2", - Finalizers: []string{utils.FinalizerStorageNodeSet}, - }, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ClusterName: "cluster2"}, - } - - r := newStorageNodeSetStateTestReconciler(t, sn1, sn2, cluster1, cluster2) - for _, sn := range []*simplyblockv1alpha1.StorageNodeSet{sn2, sn1} { - if _, err := r.Reconcile(context.Background(), ctrl.Request{NamespacedName: client.ObjectKeyFromObject(sn)}); err != nil { - t.Fatalf("reconcile %s/%s returned error: %v", sn.Namespace, sn.Name, err) - } - } - - for _, namespace := range []string{"cluster1", "cluster2"} { - binding := &rbacv1.ClusterRoleBinding{} - key := client.ObjectKey{Name: "simplyblock-storage-node-binding-" + namespace} - if err := r.Get(context.Background(), key, binding); err != nil { - t.Fatalf("failed to fetch ClusterRoleBinding %s: %v", key.Name, err) - } - if len(binding.Subjects) != 1 || binding.Subjects[0].Namespace != namespace { - t.Fatalf("expected binding %s to target namespace %s, got %#v", key.Name, namespace, binding.Subjects) - } - } -} - -func TestStorageNodeSetReconcileMissingInternalIPRequeues(t *testing.T) { - const namespace = "default" - const clusterName = "cluster-missing-ip" - const clusterUUID = "cluster-uuid-missing-ip" - const workerName = "node-no-ip" - - cluster := &simplyblockv1alpha2.StorageCluster{ - ObjectMeta: metav1.ObjectMeta{Name: clusterName, Namespace: namespace}, - Spec: simplyblockv1alpha2.StorageClusterSpec{}, - Status: simplyblockv1alpha2.StorageClusterStatus{UUID: clusterUUID}, - } - node := &corev1.Node{ - ObjectMeta: metav1.ObjectMeta{Name: workerName}, - } - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{ - Name: "sn-missing-ip", - Namespace: namespace, - Finalizers: []string{utils.FinalizerStorageNodeSet}, - }, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - ClusterName: clusterName, - WorkerNodes: []string{workerName}, - }, - } - - r := newStorageNodeSetStateTestReconciler(t, sn, cluster, node) - res, err := r.Reconcile(context.Background(), ctrl.Request{NamespacedName: client.ObjectKeyFromObject(sn)}) - if err != nil { - t.Fatalf("reconcile returned error: %v", err) - } - if res.RequeueAfter == 0 { - t.Fatalf("expected delayed requeue when worker has no internal IP") - } -} - -func TestStorageNodeSetReconcileUnreachableNodeInfoRequeues(t *testing.T) { - const namespace = "default" - const clusterName = "cluster-unreachable-info" - const clusterUUID = "cluster-uuid-unreachable-info" - const workerName = "node-bad-ip" - - cluster := &simplyblockv1alpha2.StorageCluster{ - ObjectMeta: metav1.ObjectMeta{Name: clusterName, Namespace: namespace}, - Spec: simplyblockv1alpha2.StorageClusterSpec{}, - Status: simplyblockv1alpha2.StorageClusterStatus{UUID: clusterUUID}, - } - node := &corev1.Node{ - ObjectMeta: metav1.ObjectMeta{Name: workerName}, - Status: corev1.NodeStatus{ - Addresses: []corev1.NodeAddress{ - { - Type: corev1.NodeInternalIP, - Address: "bad ip", - }, - }, - }, - } - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{ - Name: "sn-unreachable-info", - Namespace: namespace, - Finalizers: []string{utils.FinalizerStorageNodeSet}, - }, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - ClusterName: clusterName, - WorkerNodes: []string{workerName}, - }, - } - - r := newStorageNodeSetStateTestReconciler(t, sn, cluster, node) - res, err := r.Reconcile(context.Background(), ctrl.Request{NamespacedName: client.ObjectKeyFromObject(sn)}) - if err != nil { - t.Fatalf("reconcile returned error: %v", err) - } - if res.RequeueAfter == 0 { - t.Fatalf("expected delayed requeue when node info endpoint is unreachable") - } -} - -func TestCheckNodeInfoReachable(t *testing.T) { - // Use an unroutable test-net address to deterministically exercise error path. - err := checkNodeInfoReachable(context.Background(), "192.0.2.1", "default", false, false) - if err == nil { - t.Fatalf("expected error when node info endpoint is unreachable") - } -} - -func TestCheckNodeInfoReachableTLSMissingCA(t *testing.T) { - // With TLS enabled and the default CA path (which won't exist in unit tests), - // the function must surface a build-client error before attempting any I/O. - err := checkNodeInfoReachable(context.Background(), "192.0.2.1", "default", true, false) - if err == nil { - t.Fatalf("expected error when CA bundle is missing") - } - if !strings.Contains(err.Error(), "build storage-node TLS client") { - t.Fatalf("expected TLS client build error, got: %v", err) - } -} - -func TestWaitForNodeInfoReachable(t *testing.T) { - origCheckFn := waitForNodeInfoReachableCheckFn - origRetries := waitForNodeInfoReachableMaxRetries - origDelay := waitForNodeInfoReachableRetryDelay - t.Cleanup(func() { - waitForNodeInfoReachableCheckFn = origCheckFn - waitForNodeInfoReachableMaxRetries = origRetries - waitForNodeInfoReachableRetryDelay = origDelay - }) - - t.Run("returns nil on first successful check", func(t *testing.T) { - attempts := 0 - waitForNodeInfoReachableMaxRetries = 3 - waitForNodeInfoReachableRetryDelay = time.Millisecond - waitForNodeInfoReachableCheckFn = func(context.Context, string, string, bool, bool) error { - attempts++ - return nil - } - - if err := waitForNodeInfoReachable(context.Background(), "node-a", "default", false, false); err != nil { - t.Fatalf("waitForNodeInfoReachable returned error: %v", err) - } - if attempts != 1 { - t.Fatalf("expected one attempt, got %d", attempts) - } - }) - - t.Run("retries and then succeeds", func(t *testing.T) { - attempts := 0 - waitForNodeInfoReachableMaxRetries = 4 - waitForNodeInfoReachableRetryDelay = time.Millisecond - waitForNodeInfoReachableCheckFn = func(context.Context, string, string, bool, bool) error { - attempts++ - if attempts < 3 { - return errors.New("temporary failure") - } - return nil - } - - if err := waitForNodeInfoReachable(context.Background(), "node-b", "default", false, false); err != nil { - t.Fatalf("waitForNodeInfoReachable returned error: %v", err) - } - if attempts != 3 { - t.Fatalf("expected three attempts, got %d", attempts) - } - }) - - t.Run("returns context cancellation", func(t *testing.T) { - waitForNodeInfoReachableMaxRetries = 5 - waitForNodeInfoReachableRetryDelay = time.Second - waitForNodeInfoReachableCheckFn = func(context.Context, string, string, bool, bool) error { - return errors.New("still down") - } - ctx, cancel := context.WithCancel(context.Background()) - cancel() - - err := waitForNodeInfoReachable(ctx, "node-c", "default", false, false) - if !errors.Is(err, context.Canceled) { - t.Fatalf("expected context canceled error, got %v", err) - } - }) - - t.Run("returns wrapped error after max retries", func(t *testing.T) { - waitForNodeInfoReachableMaxRetries = 3 - waitForNodeInfoReachableRetryDelay = time.Millisecond - waitForNodeInfoReachableCheckFn = func(context.Context, string, string, bool, bool) error { - return errors.New("permanent failure") - } - - err := waitForNodeInfoReachable(context.Background(), "node-d", "default", false, false) - if err == nil { - t.Fatalf("expected timeout error after retries") - } - if !strings.Contains(err.Error(), fmt.Sprintf("after %d retries", waitForNodeInfoReachableMaxRetries)) { - t.Fatalf("unexpected retry error message: %v", err) - } - if !strings.Contains(err.Error(), "permanent failure") { - t.Fatalf("expected wrapped failure message, got: %v", err) - } - }) -} - -func TestPollNodeOnlinePaths(t *testing.T) { - origActivationDelay := waitForNodeOnlineActivationDelay - origSleepFn := waitForNodeOnlineSleepFn - t.Cleanup(func() { - waitForNodeOnlineActivationDelay = origActivationDelay - waitForNodeOnlineSleepFn = origSleepFn - }) - waitForNodeOnlineActivationDelay = 0 - waitForNodeOnlineSleepFn = func(context.Context, time.Duration) error { return nil } - - t.Run("updates node status and returns done when cluster already active", func(t *testing.T) { - const clusterName = "cluster-a" - const clusterUUID = "cluster-uuid-online" - - mock := webapimock.NewSpecServerFromFile(t, "../../../shared/openapi.json", true) - defer mock.Close() - mock.Register( - http.MethodGet, - "/api/v2/clusters/"+clusterUUID+"/storage-nodes/", - webapimock.RouteResponse{ - Status: http.StatusOK, - Body: `[ - { - "id":"node-uuid-1", - "status":"online", - "mgmt_ip":"10.0.0.1", - "health_check":true, - "hostname":"node-a", - "online_devices":"nvme0n1", - "cpu":4, - "spdk_mem":2147483648, - "lvols":3, - "rpc_port":9000, - "lvol_subsys_port":9001, - "nvmf_port":9002 - } - ]`, - Headers: map[string]string{"Content-Type": "application/json"}, - }, - ) - apiClient := webapi.NewClient(mock.URL()) - - cluster := &simplyblockv1alpha2.StorageCluster{ - ObjectMeta: metav1.ObjectMeta{Name: clusterName, Namespace: "default"}, - Spec: simplyblockv1alpha2.StorageClusterSpec{}, - Status: simplyblockv1alpha2.StorageClusterStatus{ - Status: "active", - ErasureCodingScheme: "1x0", - }, - } - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-online", Namespace: "default"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - ClusterName: clusterName, - WorkerNodes: []string{"node-a"}, - }, - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - Nodes: []simplyblockv1alpha1.NodeStatus{ - {Hostname: "node-a", MgmtIp: mgmtIP, Status: "in_creation"}, - }, - }, - } - r := newStorageNodeSetStateTestReconciler(t, cluster, sn) - - res, err := r.pollNodeOnline(context.Background(), apiClient, clusterUUID, mgmtIP, "node-a", 1, sn) - if err != nil { - t.Fatalf("pollNodeOnline returned error: %v", err) - } - if res.RequeueAfter != 0 { - t.Fatalf("expected done result, got requeue: %v", res) - } - if len(sn.Status.Nodes) != 1 { - t.Fatalf("unexpected node status length: %d", len(sn.Status.Nodes)) - } - got := sn.Status.Nodes[0] - if got.Status != utils.NodeStatusOnline || got.UUID != nodeUUID1 { - t.Fatalf("node status not updated as expected: %#v", got) - } - }) - - t.Run("appends node status entry when node missing in status list", func(t *testing.T) { - const clusterName = "cluster-b" - const clusterUUID = "cluster-uuid-missing-status" - - mock := webapimock.NewSpecServerFromFile(t, "../../../shared/openapi.json", true) - defer mock.Close() - mock.Register( - http.MethodGet, - "/api/v2/clusters/"+clusterUUID+"/storage-nodes/", - webapimock.RouteResponse{ - Status: http.StatusOK, - Body: `[ - { - "id":"node-uuid-2", - "status":"online", - "mgmt_ip":"10.0.0.2", - "health_check":true, - "hostname":"node-b", - "online_devices":"nvme0n2", - "cpu":8, - "spdk_mem":4294967296, - "lvols":1, - "rpc_port":9100, - "lvol_subsys_port":9101, - "nvmf_port":9102 - } - ]`, - Headers: map[string]string{"Content-Type": "application/json"}, - }, - ) - apiClient := webapi.NewClient(mock.URL()) - - cluster := &simplyblockv1alpha2.StorageCluster{ - ObjectMeta: metav1.ObjectMeta{Name: clusterName, Namespace: "default"}, - Spec: simplyblockv1alpha2.StorageClusterSpec{}, - Status: simplyblockv1alpha2.StorageClusterStatus{ - Status: "active", - ErasureCodingScheme: "1x0", - }, - } - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-missing-status", Namespace: "default"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - ClusterName: clusterName, - WorkerNodes: []string{"node-b"}, - }, - } - r := newStorageNodeSetStateTestReconciler(t, cluster, sn) - - res, err := r.pollNodeOnline(context.Background(), apiClient, clusterUUID, "10.0.0.2", "node-b", 1, sn) - if err != nil { - t.Fatalf("pollNodeOnline returned unexpected error: %v", err) - } - if res.RequeueAfter != 0 { - t.Fatalf("expected done result, got requeue: %v", res) - } - if len(sn.Status.Nodes) != 1 { - t.Fatalf("expected 1 status entry, got %d", len(sn.Status.Nodes)) - } - got := sn.Status.Nodes[0] - if got.Status != statusOnline || got.UUID != "node-uuid-2" || got.Hostname != "node-b" { - t.Fatalf("unexpected appended node status: %#v", got) - } - }) - - t.Run("returns RequeueAfter when node not yet online and within timeout window", func(t *testing.T) { - const clusterUUID = "cluster-uuid-not-yet-online" - mock := webapimock.NewSpecServerFromFile(t, "../../../shared/openapi.json", true) - defer mock.Close() - mock.Register( - http.MethodGet, - "/api/v2/clusters/"+clusterUUID+"/storage-nodes/", - webapimock.RouteResponse{ - Status: http.StatusOK, - Body: `[]`, - Headers: map[string]string{"Content-Type": "application/json"}, - }, - ) - - postedAt := metav1.Now() - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-not-yet-online", Namespace: "default"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ClusterName: "cluster-a"}, - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - Nodes: []simplyblockv1alpha1.NodeStatus{ - {Hostname: "node-a", MgmtIp: mgmtIP, Status: "in_creation", PostedAt: &postedAt}, - }, - }, - } - r := newStorageNodeSetStateTestReconciler(t, sn) - - res, err := r.pollNodeOnline(context.Background(), webapi.NewClient(mock.URL()), clusterUUID, mgmtIP, "node-a", 1, sn) - if err != nil { - t.Fatalf("pollNodeOnline returned unexpected error: %v", err) - } - if res.RequeueAfter == 0 { - t.Fatalf("expected RequeueAfter, got done result") - } - }) -} - -func TestPollNodeOnlineErrorAndTimeoutPaths(t *testing.T) { - origActivationDelay := waitForNodeOnlineActivationDelay - origSleepFn := waitForNodeOnlineSleepFn - t.Cleanup(func() { - waitForNodeOnlineActivationDelay = origActivationDelay - waitForNodeOnlineSleepFn = origSleepFn - }) - waitForNodeOnlineActivationDelay = 0 - waitForNodeOnlineSleepFn = func(context.Context, time.Duration) error { return nil } - - t.Run("returns error on invalid storage-node payload", func(t *testing.T) { - const clusterUUID = "cluster-uuid-wfno-invalid-json" - mock := webapimock.NewSpecServerFromFile(t, "../../../shared/openapi.json", true) - defer mock.Close() - mock.Register( - http.MethodGet, - "/api/v2/clusters/"+clusterUUID+"/storage-nodes/", - webapimock.RouteResponse{ - Status: http.StatusOK, - Body: `{`, - Headers: map[string]string{ - "Content-Type": "application/json", - }, - }, - ) - - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-wfno-invalid-json", Namespace: "default"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - ClusterName: "cluster-a", - }, - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - Nodes: []simplyblockv1alpha1.NodeStatus{ - {Hostname: "node-a", MgmtIp: mgmtIP, Status: "in_creation"}, - }, - }, - } - r := newStorageNodeSetStateTestReconciler(t, sn) - - _, err := r.pollNodeOnline(context.Background(), webapi.NewClient(mock.URL()), clusterUUID, mgmtIP, "node-a", 1, sn) - if err == nil { - t.Fatalf("expected unmarshal error for invalid payload") - } - if !strings.Contains(err.Error(), "failed to unmarshal") { - t.Fatalf("unexpected error: %v", err) - } - }) - - t.Run("returns cluster-not-found when activation precheck cannot resolve cluster CR", func(t *testing.T) { - const clusterUUID = "cluster-uuid-wfno-cluster-missing" - mock := webapimock.NewSpecServerFromFile(t, "../../../shared/openapi.json", true) - defer mock.Close() - mock.Register( - http.MethodGet, - "/api/v2/clusters/"+clusterUUID+"/storage-nodes/", - webapimock.RouteResponse{ - Status: http.StatusOK, - Body: `[ - { - "id":"node-uuid-3", - "status":"online", - "mgmt_ip":"10.0.0.3", - "health_check":true, - "hostname":"node-c", - "online_devices":"nvme1n1", - "cpu":4, - "spdk_mem":2147483648, - "lvols":1, - "rpc_port":9200, - "lvol_subsys_port":9201, - "nvmf_port":9202 - } - ]`, - Headers: map[string]string{"Content-Type": "application/json"}, - }, - ) - - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-wfno-cluster-missing", Namespace: "default"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - ClusterName: "cluster-missing", - }, - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - Nodes: []simplyblockv1alpha1.NodeStatus{ - {Hostname: "node-c", MgmtIp: "10.0.0.3", Status: "in_creation"}, - }, - }, - } - r := newStorageNodeSetStateTestReconciler(t, sn) - - _, err := r.pollNodeOnline(context.Background(), webapi.NewClient(mock.URL()), clusterUUID, "10.0.0.3", "node-c", 1, sn) - if err == nil { - t.Fatalf("expected cluster resolution error") - } - if !strings.Contains(err.Error(), "cluster not found yet") { - t.Fatalf("unexpected error: %v", err) - } - }) - - t.Run("writes timeout node status when PostedAt is expired", func(t *testing.T) { - const clusterUUID = "cluster-uuid-wfno-timeout" - mock := webapimock.NewSpecServerFromFile(t, "../../../shared/openapi.json", true) - defer mock.Close() - mock.Register( - http.MethodGet, - "/api/v2/clusters/"+clusterUUID+"/storage-nodes/", - webapimock.RouteResponse{ - Status: http.StatusOK, - Body: `[]`, - Headers: map[string]string{ - "Content-Type": "application/json", - }, - }, - ) - - expiredAt := metav1.NewTime(time.Now().Add(-2 * time.Hour)) - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-wfno-timeout", Namespace: "default"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - ClusterName: "cluster-a", - }, - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - Nodes: []simplyblockv1alpha1.NodeStatus{ - {Hostname: "node-timeout", MgmtIp: "10.0.0.4", Status: "in_creation", PostedAt: &expiredAt}, - }, - }, - } - r := newStorageNodeSetStateTestReconciler(t, sn) - - res, err := r.pollNodeOnline(context.Background(), webapi.NewClient(mock.URL()), clusterUUID, "10.0.0.4", "node-timeout", 1, sn) - if err != nil { - t.Fatalf("expected no error on timeout, got: %v", err) - } - if res.RequeueAfter != 0 { - t.Fatalf("expected done result after timeout, got requeue: %v", res) - } - if len(sn.Status.Nodes) != 1 { - t.Fatalf("expected timeout status node entry, got %d", len(sn.Status.Nodes)) - } - if sn.Status.Nodes[0].Hostname != "node-timeout" || sn.Status.Nodes[0].Status != "timeout" { - t.Fatalf("unexpected timeout node status: %#v", sn.Status.Nodes[0]) - } - }) -} - -// testOperatorNamespace is the namespace the test reconciler pretends to run in. -// It must match the namespace of the seeded singleton ControlPlane CR below. -const testOperatorNamespace = "default" - -func newStorageNodeSetStateTestReconciler( - t *testing.T, - objects ...client.Object, -) *StorageNodeSetReconciler { - t.Helper() - - scheme := newTestScheme( - t, - simplyblockv1alpha1.AddToScheme, - corev1.AddToScheme, - appsv1.AddToScheme, - rbacv1.AddToScheme, - discoveryv1.AddToScheme, - ) - - // Mirror real-cluster state: the Helm chart always creates the singleton - // ControlPlane CR before any StorageNodeSet CR is reconciled. - // - // It is seeded at v1alpha2 because that is the version an API server stores - // and serves. Seeding v1alpha1 would assert a read that no cluster answers: - // the retired version is served only through the conversion webhook, which a - // fresh install does not deploy. - singleton := &simplyblockv1alpha2.ControlPlane{ - ObjectMeta: metav1.ObjectMeta{ - Name: SingletonControlPlaneName, - Namespace: testOperatorNamespace, - }, - Spec: simplyblockv1alpha2.ControlPlaneSpec{ - Source: &simplyblockv1alpha2.ControlPlaneSource{ - Managed: &simplyblockv1alpha2.ManagedControlPlane{ - Image: "test-image:latest", - }, - }, - }, - } - // Simulate kubebuilder defaults that the API server would apply. - for _, obj := range objects { - if sn, ok := obj.(*simplyblockv1alpha1.StorageNodeSet); ok && sn.Spec.MaxParallelNodeAdds == nil { - v := int32(1) - sn.Spec.MaxParallelNodeAdds = &v - } - } - allObjects := append([]client.Object{singleton}, objects...) - - // Register the "spec.storageNodeSetRef" index that SetupWithManager wires up - // in production; labelWorkerNodes (reached via Reconcile) lists StorageNodes - // on that field and now returns an error when the index is absent instead of - // silently proceeding, so the fake client must mirror the real one. - cl := fake.NewClientBuilder(). - WithScheme(scheme). - WithStatusSubresource( - &simplyblockv1alpha1.StorageNodeSet{}, - &simplyblockv1alpha2.StorageCluster{}, - &simplyblockv1alpha2.ControlPlane{}, - &appsv1.DaemonSet{}, - ). - WithObjects(allObjects...). - WithIndex(&simplyblockv1alpha1.StorageNode{}, "spec.storageNodeSetRef", func(obj client.Object) []string { - return []string{obj.(*simplyblockv1alpha1.StorageNode).Spec.StorageNodeSetRef} - }). - Build() - - return &StorageNodeSetReconciler{ - Client: cl, - Scheme: scheme, - Namespace: testOperatorNamespace, - Recorder: events.NewFakeRecorder(32), - } -} - -func TestReconcileSpdkProxyService(t *testing.T) { - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn", Namespace: "ns", UID: "sn-uid"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ClusterName: "cluster-a"}, - } - - r := newStorageNodeSetStateTestReconciler(t, sn) - r.TLSEnabled = true - r.TLSProvider = utils.TLSProviderOpenShift - - if err := r.reconcileSpdkProxyService(context.Background(), sn); err != nil { - t.Fatalf("reconcileSpdkProxyService: %v", err) - } - - var svc corev1.Service - if err := r.Get(context.Background(), client.ObjectKey{Namespace: "ns", Name: "simplyblock-spdk-proxy"}, &svc); err != nil { - t.Fatalf("expected simplyblock-spdk-proxy Service to be created: %v", err) - } - if svc.Spec.ClusterIP != "None" { - t.Fatalf("expected headless Service, got ClusterIP=%q", svc.Spec.ClusterIP) - } - if len(svc.Spec.Ports) != 0 { - t.Fatalf("expected no ports on Service, got %#v", svc.Spec.Ports) - } - if got := svc.Annotations["service.beta.openshift.io/serving-cert-secret-name"]; got != "simplyblock-spdk-proxy-tls" { - t.Fatalf("missing/incorrect serving-cert annotation: %q", got) - } - if len(svc.OwnerReferences) != 1 || svc.OwnerReferences[0].UID != "sn-uid" { - t.Fatalf("expected owner reference to StorageNodeSet, got %#v", svc.OwnerReferences) - } - - // Second pass with a simulated ClusterIP already assigned must preserve it. - svc.Spec.ClusterIP = "None" - if err := r.reconcileSpdkProxyService(context.Background(), sn); err != nil { - t.Fatalf("second reconcileSpdkProxyService: %v", err) - } -} - -func TestSyncTrackedNodesStatus(t *testing.T) { - const clusterUUID = "cluster-sync-uuid" - - apiBody := func(uuid, status, ip string, health bool) string { - return fmt.Sprintf(`[{ - "id":%q, - "status":%q, - "mgmt_ip":%q, - "health_check":%v, - "hostname":"node-a", - "device_count":2, - "online_device_count":2, - "cpu_spdk_count":4, - "spdk_mem":2147483648, - "lvols":3, - "rpc_port":9000, - "lvol_subsys_port":9001, - "nvmf_port":9002 - }]`, uuid, status, ip, health) - } - - t.Run("no-op when no tracked nodes", func(t *testing.T) { - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-sync-noop", Namespace: "default"}, - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - Nodes: []simplyblockv1alpha1.NodeStatus{ - {Hostname: "node-a", UUID: ""}, - }, - }, - } - r := newStorageNodeSetStateTestReconciler(t, sn) - // Unreachable server — if the function makes an HTTP call it will fail. - c := webapi.NewClient("http://127.0.0.1:1") - if err := r.syncTrackedNodesStatus(context.Background(), c, clusterUUID, sn); err != nil { - t.Fatalf("expected no error when no tracked nodes, got: %v", err) - } - }) - - t.Run("updates tracked node fields by UUID", func(t *testing.T) { - postedAt := metav1.Now() - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-sync-update", Namespace: "default"}, - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - Nodes: []simplyblockv1alpha1.NodeStatus{ - { - Hostname: "node-a", - UUID: nodeUUID1, - Status: "in_creation", - Health: false, - MgmtIp: "10.0.0.1", - PostedAt: &postedAt, - Uptime: "1d2h", - }, - }, - }, - } - - srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) { - w.Header().Set("Content-Type", "application/json") - w.WriteHeader(http.StatusOK) - _, _ = w.Write([]byte(apiBody(nodeUUID1, statusOnline, "10.0.0.99", true))) - })) - defer srv.Close() - - r := newStorageNodeSetStateTestReconciler(t, sn) - if err := r.syncTrackedNodesStatus(context.Background(), webapi.NewClient(srv.URL), clusterUUID, sn); err != nil { - t.Fatalf("syncTrackedNodesStatus returned error: %v", err) - } - - n := sn.Status.Nodes[0] - if n.Status != statusOnline { - t.Errorf("expected Status %q, got %q", statusOnline, n.Status) - } - if !n.Health { - t.Errorf("expected Health=true") - } - if n.MgmtIp != "10.0.0.99" { - t.Errorf("expected MgmtIp 10.0.0.99, got %q", n.MgmtIp) - } - if n.UUID != nodeUUID1 { - t.Errorf("expected UUID preserved, got %q", n.UUID) - } - }) - - t.Run("preserves PostedAt and Uptime across sync", func(t *testing.T) { - postedAt := metav1.NewTime(time.Now().Add(-1 * time.Hour).Truncate(time.Second)) - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-sync-preserve", Namespace: "default"}, - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - Nodes: []simplyblockv1alpha1.NodeStatus{ - { - Hostname: "node-a", - UUID: "node-uuid-2", - Status: statusOnline, - PostedAt: &postedAt, - Uptime: "3d4h", - }, - }, - }, - } - - srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) { - w.Header().Set("Content-Type", "application/json") - w.WriteHeader(http.StatusOK) - _, _ = w.Write([]byte(apiBody("node-uuid-2", "online", "10.0.0.2", true))) - })) - defer srv.Close() - - r := newStorageNodeSetStateTestReconciler(t, sn) - if err := r.syncTrackedNodesStatus(context.Background(), webapi.NewClient(srv.URL), clusterUUID, sn); err != nil { - t.Fatalf("syncTrackedNodesStatus returned error: %v", err) - } - - n := sn.Status.Nodes[0] - if n.PostedAt == nil || !n.PostedAt.Equal(&postedAt) { - t.Errorf("expected PostedAt to be preserved, got %v", n.PostedAt) - } - if n.Uptime != "3d4h" { - t.Errorf("expected Uptime to be preserved as %q, got %q", "3d4h", n.Uptime) - } - }) - - t.Run("skips nodes whose UUID is absent from API response", func(t *testing.T) { - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-sync-missing", Namespace: "default"}, - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - Nodes: []simplyblockv1alpha1.NodeStatus{ - {Hostname: "node-a", UUID: "node-uuid-known", Status: "in_creation"}, - {Hostname: "node-b", UUID: "node-uuid-gone", Status: "in_creation"}, - }, - }, - } - - srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) { - w.Header().Set("Content-Type", "application/json") - w.WriteHeader(http.StatusOK) - // Only return node-uuid-known; node-uuid-gone is absent. - _, _ = w.Write([]byte(apiBody("node-uuid-known", statusOnline, "10.0.0.3", true))) - })) - defer srv.Close() - - r := newStorageNodeSetStateTestReconciler(t, sn) - if err := r.syncTrackedNodesStatus(context.Background(), webapi.NewClient(srv.URL), clusterUUID, sn); err != nil { - t.Fatalf("syncTrackedNodesStatus returned error: %v", err) - } - - if sn.Status.Nodes[0].Status != statusOnline { - t.Errorf("expected known node to be updated, got status %q", sn.Status.Nodes[0].Status) - } - if sn.Status.Nodes[1].Status != "in_creation" { - t.Errorf("expected absent node to be left unchanged, got status %q", sn.Status.Nodes[1].Status) - } - }) - - t.Run("returns error when API call fails", func(t *testing.T) { - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-sync-apierr", Namespace: "default"}, - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - Nodes: []simplyblockv1alpha1.NodeStatus{ - {Hostname: "node-a", UUID: "node-uuid-err"}, - }, - }, - } - - srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) { - w.WriteHeader(http.StatusInternalServerError) - })) - defer srv.Close() - - r := newStorageNodeSetStateTestReconciler(t, sn) - err := r.syncTrackedNodesStatus(context.Background(), webapi.NewClient(srv.URL), clusterUUID, sn) - if err == nil { - t.Fatalf("expected error on API failure") - } - if !strings.Contains(err.Error(), "sync: failed to list storage nodes") { - t.Errorf("unexpected error message: %v", err) - } - }) - - t.Run("returns error on invalid JSON response", func(t *testing.T) { - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-sync-badjson", Namespace: "default"}, - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - Nodes: []simplyblockv1alpha1.NodeStatus{ - {Hostname: "node-a", UUID: "node-uuid-json"}, - }, - }, - } - - srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) { - w.WriteHeader(http.StatusOK) - _, _ = w.Write([]byte(`{`)) - })) - defer srv.Close() - - r := newStorageNodeSetStateTestReconciler(t, sn) - err := r.syncTrackedNodesStatus(context.Background(), webapi.NewClient(srv.URL), clusterUUID, sn) - if err == nil { - t.Fatalf("expected error on invalid JSON") - } - if !strings.Contains(err.Error(), "sync: failed to unmarshal") { - t.Errorf("unexpected error message: %v", err) - } - }) -} - -func TestReconcileServicesAndServingCertificatesForCertManagerProvider(t *testing.T) { - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn", Namespace: "ns", UID: "sn-uid"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ClusterName: "cluster-a"}, - } - - r := newStorageNodeSetStateTestReconciler(t, sn) - r.TLSEnabled = true - r.TLSProvider = utils.TLSProviderCertManager - - if err := r.reconcileService(context.Background(), sn); err != nil { - t.Fatalf("reconcileService: %v", err) - } - if err := r.reconcileSpdkProxyService(context.Background(), sn); err != nil { - t.Fatalf("reconcileSpdkProxyService: %v", err) - } - if err := r.reconcileServingCertificates(context.Background(), sn); err != nil { - t.Fatalf("reconcileServingCertificates: %v", err) - } - - for _, serviceName := range []string{"simplyblock-storage-node-api", "simplyblock-spdk-proxy"} { - var svc corev1.Service - if err := r.Get(context.Background(), client.ObjectKey{Namespace: "ns", Name: serviceName}, &svc); err != nil { - t.Fatalf("expected Service %s to be created: %v", serviceName, err) - } - if got := svc.Annotations[utils.OpenShiftServingCertAnnotation]; got != "" { - t.Fatalf("unexpected OpenShift serving-cert annotation on %s: %q", serviceName, got) - } - } - - for serviceName, secretName := range map[string]string{ - "simplyblock-storage-node-api": "simplyblock-storage-node-api-tls", - "simplyblock-spdk-proxy": "simplyblock-spdk-proxy-tls", - } { - cert := &unstructured.Unstructured{} - cert.SetAPIVersion("cert-manager.io/v1") - cert.SetKind("Certificate") - if err := r.Get(context.Background(), client.ObjectKey{Namespace: "ns", Name: serviceName}, cert); err != nil { - t.Fatalf("expected Certificate %s to be created: %v", serviceName, err) - } - - gotSecret, found, err := unstructured.NestedString(cert.Object, "spec", "secretName") - if err != nil || !found { - t.Fatalf("expected secretName on Certificate %s, err=%v found=%v", serviceName, err, found) - } - if gotSecret != secretName { - t.Fatalf("Certificate %s secretName = %q, want %q", serviceName, gotSecret, secretName) - } - - gotIssuer, found, err := unstructured.NestedString(cert.Object, "spec", "issuerRef", "name") - if err != nil || !found { - t.Fatalf("expected issuerRef.name on Certificate %s, err=%v found=%v", serviceName, err, found) - } - if gotIssuer != utils.CertManagerClusterIssuerName { - t.Fatalf("Certificate %s issuerRef.name = %q, want %q", serviceName, gotIssuer, utils.CertManagerClusterIssuerName) - } - - dnsNames, found, err := unstructured.NestedStringSlice(cert.Object, "spec", "dnsNames") - if err != nil || !found { - t.Fatalf("expected dnsNames on Certificate %s, err=%v found=%v", serviceName, err, found) - } - if !contains(dnsNames, serviceName) || !contains(dnsNames, serviceName+".ns.svc.cluster.local") { - t.Fatalf("Certificate %s dnsNames = %#v", serviceName, dnsNames) - } - } -} - -func TestReconcileSpdkProxyEndpointSlices(t *testing.T) { - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn", Namespace: "ns", UID: "sn-uid"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ClusterName: "cluster-a"}, - } - - podReady := func(name, node, ip, rpcPort string) *corev1.Pod { - return &corev1.Pod{ - ObjectMeta: metav1.ObjectMeta{ - Name: name, - Namespace: "ns", - Labels: map[string]string{"role": "simplyblock-storage-node"}, - }, - Spec: corev1.PodSpec{ - NodeName: node, - Containers: []corev1.Container{ - { - Name: "spdk-proxy-container", - Env: []corev1.EnvVar{{Name: "RPC_PORT", Value: rpcPort}}, - }, - }, - }, - Status: corev1.PodStatus{ - Phase: corev1.PodRunning, - PodIP: ip, - ContainerStatuses: []corev1.ContainerStatus{ - {Name: "spdk-proxy-container", Ready: true}, - }, - }, - } - } - - pod1 := podReady("snode-spdk-pod-9001-cid", "node-a", mgmtIP, "9001") - pod2 := podReady("snode-spdk-pod-9002-cid", "node-a", mgmtIP, "9002") - pod3 := podReady("snode-spdk-pod-9001-cid-b", "node-b", "10.0.0.2", "9001") - - // wrong label — must be ignored - ignored := podReady("other", "node-c", "10.0.0.3", "9001") - ignored.Labels = map[string]string{"role": "other"} - - // not ready — must be ignored - notReady := podReady("not-ready", "node-a", mgmtIP, "9003") - notReady.Status.ContainerStatuses[0].Ready = false - - r := newStorageNodeSetStateTestReconciler(t, sn, pod1, pod2, pod3, ignored, notReady) - - ctx := context.Background() - if err := r.reconcileSpdkProxyEndpointSlices(ctx, sn); err != nil { - t.Fatalf("reconcileSpdkProxyEndpointSlices: %v", err) - } - - var slices discoveryv1.EndpointSliceList - if err := r.List(ctx, &slices, - client.InNamespace("ns"), - client.MatchingLabels{"kubernetes.io/service-name": "simplyblock-spdk-proxy"}, - ); err != nil { - t.Fatalf("list slices: %v", err) - } - if len(slices.Items) != 2 { - t.Fatalf("expected 2 EndpointSlices, got %d", len(slices.Items)) - } - - byName := map[string]discoveryv1.EndpointSlice{} - for _, s := range slices.Items { - byName[s.Name] = s - } - - slice9001, ok := byName["spdk-proxy-endpoints-9001"] - if !ok { - t.Fatalf("missing slice spdk-proxy-endpoints-9001; got %v", sliceNames(slices.Items)) - } - if len(slice9001.Endpoints) != 2 { - t.Fatalf("slice 9001: expected 2 endpoints, got %d", len(slice9001.Endpoints)) - } - gotHostnames := map[string]string{} - for _, ep := range slice9001.Endpoints { - if ep.Hostname == nil || len(ep.Addresses) != 1 { - t.Fatalf("slice 9001: malformed endpoint %#v", ep) - } - gotHostnames[*ep.Hostname] = ep.Addresses[0] - } - if gotHostnames["node-a"] != mgmtIP || - gotHostnames["node-b"] != "10.0.0.2" { - t.Fatalf("slice 9001: unexpected hostname/address map %#v", gotHostnames) - } - if len(slice9001.Ports) != 1 || slice9001.Ports[0].Port == nil || *slice9001.Ports[0].Port != 9001 { - t.Fatalf("slice 9001: expected port 9001, got %#v", slice9001.Ports) - } - if !metav1.IsControlledBy(&slice9001, sn) { - t.Fatalf("slice 9001: expected owner reference to StorageNodeSet") - } - - slice9002 := byName["spdk-proxy-endpoints-9002"] - if len(slice9002.Endpoints) != 1 || *slice9002.Endpoints[0].Hostname != "node-a" { - t.Fatalf("slice 9002: unexpected endpoints %#v", slice9002.Endpoints) - } - - // Delete pod2 (the only pod on port 9002) and reconcile again — the stale - // slice should be removed. - if err := r.Delete(ctx, pod2); err != nil { - t.Fatalf("delete pod2: %v", err) - } - if err := r.reconcileSpdkProxyEndpointSlices(ctx, sn); err != nil { - t.Fatalf("second reconcileSpdkProxyEndpointSlices: %v", err) - } - - slices = discoveryv1.EndpointSliceList{} - if err := r.List(ctx, &slices, - client.InNamespace("ns"), - client.MatchingLabels{"kubernetes.io/service-name": "simplyblock-spdk-proxy"}, - ); err != nil { - t.Fatalf("list slices after delete: %v", err) - } - if len(slices.Items) != 1 || slices.Items[0].Name != "spdk-proxy-endpoints-9001" { - t.Fatalf("expected only spdk-proxy-endpoints-9001 after pod2 deletion, got %v", sliceNames(slices.Items)) - } -} - -// TestReconcileSpdkProxyEndpointSlices_TransientNotReadyKeepsSlice guards the -// fix for the incident where a live, healthy pod's readiness condition -// flickered false for a single reconcile pass (well short of what would -// restart the container or emit an Unhealthy event) and the reconciler -// deleted its EndpointSlice outright, breaking DNS resolution for that node's -// SPDK-proxy hostname (used for TLS SNI/cert-hostname matching in -// StorageNode.rpc_client) for however long it took the next reconcile to see -// the pod ready again. -func TestReconcileSpdkProxyEndpointSlices_TransientNotReadyKeepsSlice(t *testing.T) { - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn", Namespace: "ns", UID: "sn-uid"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ClusterName: "cluster-a"}, - } - - pod := &corev1.Pod{ - ObjectMeta: metav1.ObjectMeta{ - Name: "snode-spdk-pod-9001-cid", - Namespace: "ns", - Labels: map[string]string{"role": "simplyblock-storage-node"}, - }, - Spec: corev1.PodSpec{ - NodeName: "node-a", - Containers: []corev1.Container{ - { - Name: "spdk-proxy-container", - Env: []corev1.EnvVar{{Name: "RPC_PORT", Value: "9001"}}, - }, - }, - }, - Status: corev1.PodStatus{ - Phase: corev1.PodRunning, - PodIP: mgmtIP, - ContainerStatuses: []corev1.ContainerStatus{ - {Name: "spdk-proxy-container", Ready: true}, - }, - }, - } - - r := newStorageNodeSetStateTestReconciler(t, sn, pod) - ctx := context.Background() - - if err := r.reconcileSpdkProxyEndpointSlices(ctx, sn); err != nil { - t.Fatalf("reconcileSpdkProxyEndpointSlices: %v", err) - } - - listSlices := func() discoveryv1.EndpointSliceList { - var slices discoveryv1.EndpointSliceList - if err := r.List(ctx, &slices, - client.InNamespace("ns"), - client.MatchingLabels{"kubernetes.io/service-name": "simplyblock-spdk-proxy"}, - ); err != nil { - t.Fatalf("list slices: %v", err) - } - return slices - } - - if slices := listSlices(); len(slices.Items) != 1 { - t.Fatalf("expected 1 EndpointSlice after first reconcile, got %d", len(slices.Items)) - } - - // Pod flips transiently not-ready (a single missed probe tick) but is - // still Running with the same IP -- must NOT be treated the same as a - // genuinely deleted pod. - pod.Status.ContainerStatuses[0].Ready = false - if err := r.Update(ctx, pod); err != nil { - t.Fatalf("update pod to not-ready: %v", err) - } - if err := r.reconcileSpdkProxyEndpointSlices(ctx, sn); err != nil { - t.Fatalf("second reconcileSpdkProxyEndpointSlices: %v", err) - } - - slices := listSlices() - if len(slices.Items) != 1 || slices.Items[0].Name != "spdk-proxy-endpoints-9001" { - t.Fatalf("expected the EndpointSlice to survive a transient not-ready pod, got %v", sliceNames(slices.Items)) - } - // The stale (last-known-good) endpoint must still be present -- DNS for - // node-a keeps resolving through the blip. - if len(slices.Items[0].Endpoints) != 1 || *slices.Items[0].Endpoints[0].Hostname != "node-a" { - t.Fatalf("expected node-a's endpoint to remain, got %#v", slices.Items[0].Endpoints) - } - - // Now the pod is genuinely gone -- the slice must be deleted as before. - if err := r.Delete(ctx, pod); err != nil { - t.Fatalf("delete pod: %v", err) - } - if err := r.reconcileSpdkProxyEndpointSlices(ctx, sn); err != nil { - t.Fatalf("third reconcileSpdkProxyEndpointSlices: %v", err) - } - if slices := listSlices(); len(slices.Items) != 0 { - t.Fatalf("expected the slice to be deleted once the pod is genuinely gone, got %v", sliceNames(slices.Items)) - } -} - -func TestReconcileSpdkProxyEndpointSlices_DuplicateFirstSegment(t *testing.T) { - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn", Namespace: "ns", UID: "sn-uid"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ClusterName: "cluster-a"}, - } - - podReady := func(name, node, ip, rpcPort string) *corev1.Pod { - return &corev1.Pod{ - ObjectMeta: metav1.ObjectMeta{ - Name: name, - Namespace: "ns", - Labels: map[string]string{"role": "simplyblock-storage-node"}, - }, - Spec: corev1.PodSpec{ - NodeName: node, - Containers: []corev1.Container{ - { - Name: "spdk-proxy-container", - Env: []corev1.EnvVar{{Name: "RPC_PORT", Value: rpcPort}}, - }, - }, - }, - Status: corev1.PodStatus{ - Phase: corev1.PodRunning, - PodIP: ip, - ContainerStatuses: []corev1.ContainerStatus{ - {Name: "spdk-proxy-container", Ready: true}, - }, - }, - } - } - - // Two distinct nodes whose first DNS label collides. - pod1 := podReady("snode-spdk-pod-9001-a", "worker.us-east-1.local", "10.0.0.1", "9001") - pod2 := podReady("snode-spdk-pod-9001-b", "worker.eu-west-1.local", "10.0.0.2", "9001") - - r := newStorageNodeSetStateTestReconciler(t, sn, pod1, pod2) - - ctx := context.Background() - err := r.reconcileSpdkProxyEndpointSlices(ctx, sn) - if err == nil { - t.Fatalf("expected collision error, got nil") - } - msg := err.Error() - if !strings.Contains(msg, "worker.us-east-1.local") || !strings.Contains(msg, "worker.eu-west-1.local") { - t.Fatalf("expected error to name both colliding nodes, got %q", msg) - } - - var slices discoveryv1.EndpointSliceList - if err := r.List(ctx, &slices, - client.InNamespace("ns"), - client.MatchingLabels{"kubernetes.io/service-name": "simplyblock-spdk-proxy"}, - ); err != nil { - t.Fatalf("list slices: %v", err) - } - if len(slices.Items) != 0 { - t.Fatalf("expected no slices to be created on collision, got %v", sliceNames(slices.Items)) - } -} - -func TestExtractSpdkProxyRpcPort_FallbackToPodName(t *testing.T) { - pod := &corev1.Pod{ - ObjectMeta: metav1.ObjectMeta{Name: "snode-spdk-pod-9004-mycluster"}, - Spec: corev1.PodSpec{ - Containers: []corev1.Container{ - {Name: "spdk-proxy-container"}, // no RPC_PORT env - }, - }, - } - got, ok := extractSpdkProxyRpcPort(pod) - if !ok || got != 9004 { - t.Fatalf("expected (9004,true) from pod-name fallback, got (%d,%v)", got, ok) - } -} - -func sliceNames(items []discoveryv1.EndpointSlice) []string { - out := make([]string, 0, len(items)) - for _, s := range items { - out = append(out, s.Name) - } - return out -} - -func TestStorageNodeSetDaemonSetTLSSecretRevisionAnnotation(t *testing.T) { - const ( - ns = "default" - clusterName = "cluster-a" - dsName = "simplyblock-storage-node-ds-sn-ds-rv" - ) - - cases := []struct { - name string - tlsEnabled bool - seedSecret bool - secretRV string - wantValue string - wantSet bool - }{ - { - name: "tls enabled with secret stamps annotation", - tlsEnabled: true, - seedSecret: true, - secretRV: "12345", - wantValue: "12345", - wantSet: true, - }, - { - name: "tls enabled but secret missing leaves annotation unset", - tlsEnabled: true, - seedSecret: false, - wantSet: false, - }, - { - name: "tls disabled leaves annotation unset even if secret exists", - tlsEnabled: false, - seedSecret: true, - secretRV: "67890", - wantSet: false, - }, - } - - for _, tc := range cases { - t.Run(tc.name, func(t *testing.T) { - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-ds-rv", Namespace: ns, UID: "uid-rv"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ClusterName: clusterName}, - } - objs := []client.Object{sn} - if tc.seedSecret { - objs = append(objs, &corev1.Secret{ - ObjectMeta: metav1.ObjectMeta{ - Name: utils.SecretNameStorageNodeSetAPITLS, - Namespace: ns, - ResourceVersion: tc.secretRV, - }, - }) - } - r := newStorageNodeSetStateTestReconciler(t, objs...) - r.TLSEnabled = tc.tlsEnabled - r.TLSProvider = utils.TLSProviderCertManager - - if err := r.reconcileDaemonSet(context.Background(), sn); err != nil { - t.Fatalf("reconcileDaemonSet returned error: %v", err) - } - - var ds appsv1.DaemonSet - if err := r.Get(context.Background(), client.ObjectKey{Name: dsName, Namespace: ns}, &ds); err != nil { - t.Fatalf("failed to fetch daemonset: %v", err) - } - - got, ok := ds.Spec.Template.Annotations[utils.AnnotationTLSSecretRevision] - switch { - case tc.wantSet && !ok: - t.Fatalf("expected pod-template annotation %q to be set", utils.AnnotationTLSSecretRevision) - case tc.wantSet && got != tc.wantValue: - t.Fatalf("annotation value: want %q, got %q", tc.wantValue, got) - case !tc.wantSet && ok: - t.Fatalf("expected pod-template annotation %q to be unset, got %q", utils.AnnotationTLSSecretRevision, got) - } - }) - } -} - -func TestStorageNodeSetDaemonSetReconcileRollsOnTLSSecretRevisionChange(t *testing.T) { - const ns = "default" - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-ds-roll", Namespace: ns, UID: "uid-roll"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ClusterName: "cluster-a"}, - } - secret := &corev1.Secret{ - ObjectMeta: metav1.ObjectMeta{ - Name: utils.SecretNameStorageNodeSetAPITLS, - Namespace: ns, - ResourceVersion: "1", - }, - } - r := newStorageNodeSetStateTestReconciler(t, sn, secret) - r.TLSEnabled = true - r.TLSProvider = utils.TLSProviderCertManager - - if err := r.reconcileDaemonSet(context.Background(), sn); err != nil { - t.Fatalf("first reconcileDaemonSet: %v", err) - } - - dsKey := client.ObjectKey{Name: "simplyblock-storage-node-ds-sn-ds-roll", Namespace: ns} - var first appsv1.DaemonSet - if err := r.Get(context.Background(), dsKey, &first); err != nil { - t.Fatalf("fetch first daemonset: %v", err) - } - - // Simulate cert-manager rotating the Secret: any Update bumps - // metadata.resourceVersion via the fake client's bookkeeping. - if err := r.Get(context.Background(), client.ObjectKey{Namespace: ns, Name: utils.SecretNameStorageNodeSetAPITLS}, secret); err != nil { - t.Fatalf("refetch secret: %v", err) - } - secret.Data = map[string][]byte{"tls.crt": []byte("rotated")} - if err := r.Update(context.Background(), secret); err != nil { - t.Fatalf("rotate secret: %v", err) - } - - if err := r.reconcileDaemonSet(context.Background(), sn); err != nil { - t.Fatalf("second reconcileDaemonSet: %v", err) - } - - var second appsv1.DaemonSet - if err := r.Get(context.Background(), dsKey, &second); err != nil { - t.Fatalf("fetch second daemonset: %v", err) - } - - firstRV := first.Spec.Template.Annotations[utils.AnnotationTLSSecretRevision] - secondRV := second.Spec.Template.Annotations[utils.AnnotationTLSSecretRevision] - if firstRV == "" || secondRV == "" { - t.Fatalf("expected pod-template annotation set in both passes, got first=%q second=%q", firstRV, secondRV) - } - if firstRV == secondRV { - t.Fatalf("expected pod-template annotation to change after Secret rotation, both still %q", firstRV) - } -} - -func TestStorageNodeSetDaemonSetSBTLSServeEnv(t *testing.T) { - cases := []struct { - name string - tlsEnabled bool - tlsProvider string - wantServe string - wantServeSet bool - wantProvider string - wantProviderSet bool - }{ - { - name: "tls enabled with cert-manager", - tlsEnabled: true, - tlsProvider: utils.TLSProviderCertManager, - wantServe: "true", - wantServeSet: true, - wantProvider: utils.TLSProviderCertManager, - wantProviderSet: true, - }, - { - name: "tls enabled with OpenShift", - tlsEnabled: true, - tlsProvider: utils.TLSProviderOpenShift, - wantServe: "true", - wantServeSet: true, - wantProvider: utils.TLSProviderOpenShift, - wantProviderSet: true, - }, - { - name: "tls disabled omits TLS env vars", - tlsEnabled: false, - tlsProvider: utils.TLSProviderCertManager, - wantServeSet: false, - }, - } - for _, tc := range cases { - t.Run(tc.name, func(t *testing.T) { - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-env", Namespace: "default", UID: "uid-env"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ClusterName: "cluster-a"}, - } - r := newStorageNodeSetStateTestReconciler(t, sn) - r.TLSEnabled = tc.tlsEnabled - r.TLSProvider = tc.tlsProvider - - if err := r.reconcileDaemonSet(context.Background(), sn); err != nil { - t.Fatalf("reconcileDaemonSet returned error: %v", err) - } - - var ds appsv1.DaemonSet - if err := r.Get(context.Background(), client.ObjectKey{Name: "simplyblock-storage-node-ds-sn-env", Namespace: "default"}, &ds); err != nil { - t.Fatalf("failed to fetch daemonset: %v", err) - } - if len(ds.Spec.Template.Spec.Containers) != 1 { - t.Fatalf("expected single main container, got %d", len(ds.Spec.Template.Spec.Containers)) - } - - envByName := map[string]string{} - envSeen := map[string]bool{} - for _, e := range ds.Spec.Template.Spec.Containers[0].Env { - envByName[e.Name] = e.Value - envSeen[e.Name] = true - } - - switch { - case tc.wantServeSet && !envSeen["SB_TLS_SERVE"]: - t.Fatalf("expected SB_TLS_SERVE env var to be set on main container") - case tc.wantServeSet && envByName["SB_TLS_SERVE"] != tc.wantServe: - t.Fatalf("SB_TLS_SERVE: want %q, got %q", tc.wantServe, envByName["SB_TLS_SERVE"]) - case !tc.wantServeSet && envSeen["SB_TLS_SERVE"]: - t.Fatalf("expected SB_TLS_SERVE env var to be absent, got %q", envByName["SB_TLS_SERVE"]) - } - - switch { - case tc.wantProviderSet && !envSeen["SB_TLS_PROVIDER"]: - t.Fatalf("expected SB_TLS_PROVIDER env var to be set on main container") - case tc.wantProviderSet && envByName["SB_TLS_PROVIDER"] != tc.wantProvider: - t.Fatalf("SB_TLS_PROVIDER: want %q, got %q", tc.wantProvider, envByName["SB_TLS_PROVIDER"]) - case !tc.wantProviderSet && envSeen["SB_TLS_PROVIDER"]: - t.Fatalf("expected SB_TLS_PROVIDER env var to be absent, got %q", envByName["SB_TLS_PROVIDER"]) - } - }) - } -} - -func TestIsStorageNodeSetTLSSecretPredicate(t *testing.T) { - cases := []struct { - name string - obj client.Object - want bool - }{ - { - name: "matches storage-node-api TLS secret", - obj: &corev1.Secret{ObjectMeta: metav1.ObjectMeta{Name: utils.SecretNameStorageNodeSetAPITLS}}, - want: true, - }, - { - name: "ignores spdk-proxy TLS secret", - obj: &corev1.Secret{ObjectMeta: metav1.ObjectMeta{Name: utils.SecretNameSpdkProxyTLS}}, - want: false, - }, - { - name: "ignores unrelated secret", - obj: &corev1.Secret{ObjectMeta: metav1.ObjectMeta{Name: "some-other-secret"}}, - want: false, - }, - } - for _, tc := range cases { - t.Run(tc.name, func(t *testing.T) { - if got := isStorageNodeSetTLSSecret(tc.obj); got != tc.want { - t.Fatalf("isStorageNodeSetTLSSecret(%q) = %v, want %v", tc.obj.GetName(), got, tc.want) - } - }) - } -} - -func TestTLSSecretToStorageNodeSetRequestsEnqueuesAllInNamespace(t *testing.T) { - const ns = "ns" - snA := &simplyblockv1alpha1.StorageNodeSet{ObjectMeta: metav1.ObjectMeta{Name: "sn-a", Namespace: ns}} - snB := &simplyblockv1alpha1.StorageNodeSet{ObjectMeta: metav1.ObjectMeta{Name: "sn-b", Namespace: ns}} - otherNS := &simplyblockv1alpha1.StorageNodeSet{ObjectMeta: metav1.ObjectMeta{Name: "sn-c", Namespace: "other"}} - - r := newStorageNodeSetStateTestReconciler(t, snA, snB, otherNS) - - secret := &corev1.Secret{ObjectMeta: metav1.ObjectMeta{ - Name: utils.SecretNameStorageNodeSetAPITLS, - Namespace: ns, - }} - reqs := r.tlsSecretToStorageNodeSetRequests(context.Background(), secret) - - got := make(map[string]bool, len(reqs)) - for _, req := range reqs { - got[req.Namespace+"/"+req.Name] = true - } - - want := map[string]bool{ns + "/sn-a": true, ns + "/sn-b": true} - if len(got) != len(want) { - t.Fatalf("expected %d requests, got %d (%v)", len(want), len(got), got) - } - for k := range want { - if !got[k] { - t.Fatalf("missing reconcile request for %q", k) - } - } - if got[ns+"/sn-c"] || got["other/sn-c"] { - t.Fatalf("did not expect cross-namespace StorageNodeSet to be enqueued: %v", got) - } -} - -// --------------------------------------------------------------------------- -// FDB worker detection -// --------------------------------------------------------------------------- - -func TestFDBWorkerSet(t *testing.T) { - const namespace = "default" - - makeFDBPod := func(name, nodeName string) *corev1.Pod { - return &corev1.Pod{ - ObjectMeta: metav1.ObjectMeta{ - Name: name, - Namespace: namespace, - Labels: map[string]string{utils.LabelFDBClusterName: "simplyblock-fdb-cluster"}, - }, - Spec: corev1.PodSpec{NodeName: nodeName}, - } - } - - makeSN := func(name string, workers ...string) *simplyblockv1alpha1.StorageNodeSet { - return &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: namespace}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{WorkerNodes: workers}, - } - } - - t.Run("worker with FDB pod is detected", func(t *testing.T) { - sn := makeSN("sn", "worker-1", "worker-2") - r := newStorageNodeSetStateTestReconciler(t, sn, makeFDBPod("fdb-log-1", "worker-1")) - - got := r.fdbWorkerSet(context.Background(), sn) - - if !got["worker-1"] { - t.Error("expected worker-1 to be in FDB set") - } - if got["worker-2"] { - t.Error("expected worker-2 to NOT be in FDB set") - } - }) - - t.Run("FDB pod on non-worker node is ignored", func(t *testing.T) { - sn := makeSN("sn", "worker-1") - r := newStorageNodeSetStateTestReconciler(t, sn, makeFDBPod("fdb-log-1", "infra-node")) - - got := r.fdbWorkerSet(context.Background(), sn) - - if got["worker-1"] { - t.Error("expected worker-1 to NOT be in FDB set") - } - }) - - t.Run("no FDB pods returns empty set", func(t *testing.T) { - sn := makeSN("sn", "worker-1", "worker-2") - r := newStorageNodeSetStateTestReconciler(t, sn) - - got := r.fdbWorkerSet(context.Background(), sn) - - for _, w := range sn.Spec.WorkerNodes { - if got[w] { - t.Errorf("expected %q to NOT be in FDB set", w) - } - } - }) - - t.Run("pod without FDB label is not counted", func(t *testing.T) { - sn := makeSN("sn", "worker-1") - otherPod := &corev1.Pod{ - ObjectMeta: metav1.ObjectMeta{ - Name: "other-pod", - Namespace: namespace, - Labels: map[string]string{"app": "something-else"}, - }, - Spec: corev1.PodSpec{NodeName: "worker-1"}, - } - r := newStorageNodeSetStateTestReconciler(t, sn, otherPod) - - got := r.fdbWorkerSet(context.Background(), sn) - - if got["worker-1"] { - t.Error("expected worker-1 to NOT be in FDB set") - } - }) - - t.Run("multiple FDB pods on same worker counted once", func(t *testing.T) { - sn := makeSN("sn", "worker-1") - r := newStorageNodeSetStateTestReconciler(t, sn, - makeFDBPod("fdb-log-1", "worker-1"), - makeFDBPod("fdb-storage-1", "worker-1"), - ) - - got := r.fdbWorkerSet(context.Background(), sn) - - if !got["worker-1"] { - t.Error("expected worker-1 to be in FDB set") - } - }) -} - -// --------------------------------------------------------------------------- -// PendingNodeAdds guard -// --------------------------------------------------------------------------- - -func TestPendingNodeAddsBlocksDuplicatePost(t *testing.T) { - const namespace = "default" - const clusterUUID = "cluster-uuid-pending" - const workerName = "worker-pending" - - postCalled := false - srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, req *http.Request) { - if req.Method == http.MethodPost { - postCalled = true - } - w.Header().Set("Content-Type", "application/json") - w.WriteHeader(http.StatusOK) - _, _ = w.Write([]byte("[]")) - })) - defer srv.Close() - - now := metav1.Now() - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-pending", Namespace: namespace, Finalizers: []string{utils.FinalizerStorageNodeSet}}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{WorkerNodes: []string{workerName}}, - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - PendingNodeAdds: map[string]metav1.Time{workerName: now}, - }, - } - node := &corev1.Node{ - ObjectMeta: metav1.ObjectMeta{Name: workerName}, - Status: corev1.NodeStatus{Addresses: []corev1.NodeAddress{{Type: corev1.NodeInternalIP, Address: "10.0.0.1"}}}, - } - - r := newStorageNodeSetStateTestReconciler(t, sn, node) - res, err := r.reconcileWorkerNode( - context.Background(), - sn, workerName, clusterUUID, webapi.NewClient(srv.URL), 1, - ) - if err != nil { - t.Fatalf("unexpected error: %v", err) - } - if postCalled { - t.Error("POST should not be called when PendingNodeAdds entry exists") - } - if res.RequeueAfter == 0 { - t.Error("expected RequeueAfter while node is not yet online") - } -} - -func TestPendingNodeAddsLegacyPlaceholderBlocksPost(t *testing.T) { - const namespace = "default" - const clusterUUID = "cluster-uuid-legacy" - const workerName = "worker-legacy" - - postCalled := false - srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, req *http.Request) { - if req.Method == http.MethodPost { - postCalled = true - } - w.Header().Set("Content-Type", "application/json") - w.WriteHeader(http.StatusOK) - _, _ = w.Write([]byte("[]")) - })) - defer srv.Close() - - postedAt := metav1.Now() - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-legacy", Namespace: namespace, Finalizers: []string{utils.FinalizerStorageNodeSet}}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{WorkerNodes: []string{workerName}}, - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - // No PendingNodeAdds — only the legacy UUID=="" placeholder. - Nodes: []simplyblockv1alpha1.NodeStatus{ - {Hostname: workerName, UUID: "", Status: "in_creation", PostedAt: &postedAt}, - }, - }, - } - node := &corev1.Node{ - ObjectMeta: metav1.ObjectMeta{Name: workerName}, - Status: corev1.NodeStatus{Addresses: []corev1.NodeAddress{{Type: corev1.NodeInternalIP, Address: "10.0.0.1"}}}, - } - - r := newStorageNodeSetStateTestReconciler(t, sn, node) - _, err := r.reconcileWorkerNode( - context.Background(), - sn, workerName, clusterUUID, webapi.NewClient(srv.URL), 1, - ) - if err != nil { - t.Fatalf("unexpected error: %v", err) - } - if postCalled { - t.Error("POST should not be called when legacy UUID=empty placeholder exists") - } -} - -// --------------------------------------------------------------------------- -// Parallel vs sequential node add split -// --------------------------------------------------------------------------- - -func TestParallelNodeAddContinuesPastPendingWorker(t *testing.T) { - // Two non-FDB workers. worker-1 is pending (PendingNodeAdds set, not yet - // online). worker-2 has no placeholder yet. The reconcile loop must - // continue past worker-1 and reach worker-2 in the same pass. - const namespace = "default" - const clusterUUID = "cluster-uuid-parallel" - - srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, req *http.Request) { - w.Header().Set("Content-Type", "application/json") - w.WriteHeader(http.StatusOK) - _, _ = w.Write([]byte("[]")) - })) - defer srv.Close() - - now := metav1.Now() - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-parallel", Namespace: namespace, Finalizers: []string{utils.FinalizerStorageNodeSet}}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{WorkerNodes: []string{"worker-1", "worker-2"}}, - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - // worker-1 is already in-flight. - PendingNodeAdds: map[string]metav1.Time{"worker-1": now}, - }, - } - node1 := &corev1.Node{ - ObjectMeta: metav1.ObjectMeta{Name: "worker-1"}, - Status: corev1.NodeStatus{Addresses: []corev1.NodeAddress{{Type: corev1.NodeInternalIP, Address: "10.0.0.1"}}}, - } - node2 := &corev1.Node{ - ObjectMeta: metav1.ObjectMeta{Name: "worker-2"}, - Status: corev1.NodeStatus{Addresses: []corev1.NodeAddress{{Type: corev1.NodeInternalIP, Address: "10.0.0.2"}}}, - } - - r := newStorageNodeSetStateTestReconciler(t, sn, node1, node2) - apiClient := webapi.NewClient(srv.URL) - - // worker-1: pending — must return RequeueAfter without touching worker-2. - res1, err := r.reconcileWorkerNode(context.Background(), sn, "worker-1", clusterUUID, apiClient, 1) - if err != nil { - t.Fatalf("worker-1: unexpected error: %v", err) - } - if res1.RequeueAfter == 0 { - t.Error("worker-1: expected RequeueAfter while in-flight") - } - // worker-1's marker must still be set (we didn't clear it). - if _, ok := sn.Status.PendingNodeAdds["worker-1"]; !ok { - t.Error("worker-1 PendingNodeAdds entry should not have been cleared") - } - - // In the parallel loop we continue — process worker-2 in the same pass. - // worker-2 has no marker so it enters the !isPending branch and writes - // PendingNodeAdds["worker-2"] before attempting the POST. checkNodeInfoReachable - // will fail (no real snode API in tests), so the marker is cleared and - // RequeueAfter is returned — but worker-2 WAS reached and processed. - res2, err := r.reconcileWorkerNode(context.Background(), sn, "worker-2", clusterUUID, apiClient, 1) - if err != nil { - t.Fatalf("worker-2: unexpected error: %v", err) - } - // worker-2 must have been processed (checkNodeInfoReachable fails → RequeueAfter). - if res2.RequeueAfter == 0 { - t.Error("worker-2: expected RequeueAfter after processing") - } -} - -// maybeActivateCluster is the OTHER path that can fire an activate call -// (alongside the Activate action, gated in controllers/cluster/actions.go). -// ShouldActivateCluster only counts online-healthy nodes against the erasure -// coding scheme -- it has no notion of failure domains, so this reconciler -// needs the same npcs+2 readiness gate or it POSTs /activate every time it -// runs whenever enough nodes are online but the FDs aren't ready yet (the -// 2026-08-06 repeated unready->in_activation->unready incident). -func newActivationTestClusterAndNodeSet(clusterName string, nodeFDs []int32) ( - *simplyblockv1alpha2.StorageCluster, *simplyblockv1alpha1.StorageNodeSet, -) { - parity := int32(2) - cluster := &simplyblockv1alpha2.StorageCluster{ - ObjectMeta: metav1.ObjectMeta{Name: clusterName, Namespace: "default"}, - Spec: simplyblockv1alpha2.StorageClusterSpec{ - EnableFailureDomains: &[]bool{true}[0], - Stripe: &simplyblockv1alpha2.StripeSpec{ParityChunks: &parity}, - }, - Status: simplyblockv1alpha2.StorageClusterStatus{ - ErasureCodingScheme: "1x1", // requiredEc=2, required=3 in ShouldActivateCluster - }, - } - - workerNodes := make([]string, 0, len(nodeFDs)) - nodes := make([]simplyblockv1alpha1.NodeStatus, 0, len(nodeFDs)) - for i, fdv := range nodeFDs { - host := fmt.Sprintf("w%d", i) - workerNodes = append(workerNodes, host) - fd := fdv - nodes = append(nodes, simplyblockv1alpha1.NodeStatus{ - Hostname: host, - MgmtIp: fmt.Sprintf("10.0.0.%d", i+1), - Status: utils.NodeStatusOnline, - Health: true, - FailureDomain: &fd, - }) - } - nodeSet := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "set-" + clusterName, Namespace: "default"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - ClusterName: clusterName, - WorkerNodes: workerNodes, - }, - Status: simplyblockv1alpha1.StorageNodeSetStatus{Nodes: nodes}, - } - return cluster, nodeSet -} - -func TestMaybeActivateClusterWaitsForFailureDomainReadiness(t *testing.T) { - // 3 online/healthy nodes (enough to satisfy ShouldActivateCluster), but - // only 2 distinct failure domains for npcs=2 -- must NOT be allowed - // through (requires npcs+2 = 4). No POST should ever reach the backend. - cluster, nodeSet := newActivationTestClusterAndNodeSet( - "cluster-maybe-activate-wait", []int32{0, 0, 1}) - - r := newStorageNodeSetStateTestReconciler(t, cluster, nodeSet) - - called := false - srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, req *http.Request) { - called = true - w.WriteHeader(http.StatusOK) - })) - defer srv.Close() - apiClient := webapi.NewClient(srv.URL) - - err := maybeActivateCluster(context.Background(), apiClient, "cluster-uuid-wait", nodeSet, r) - if err != nil { - t.Fatalf("maybeActivateCluster returned unexpected error: %v", err) - } - if called { - t.Fatalf("expected no activate call to reach the backend while failure domains aren't ready") - } -} - -func TestMaybeActivateClusterProceedsOnceFailureDomainsAreReady(t *testing.T) { - // 4 online/healthy nodes across 4 distinct, equally-sized failure domains - // for npcs=2 -- satisfies npcs+2 = 4, so the gate must let this through - // to the real activation call. - cluster, nodeSet := newActivationTestClusterAndNodeSet( - "cluster-maybe-activate-ready", []int32{0, 1, 2, 3}) - - r := newStorageNodeSetStateTestReconciler(t, cluster, nodeSet) - - origSleepFn := waitForNodeOnlineSleepFn - t.Cleanup(func() { waitForNodeOnlineSleepFn = origSleepFn }) - waitForNodeOnlineSleepFn = func(context.Context, time.Duration) error { return nil } - - activateCalled := false - srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, req *http.Request) { - if req.Method == http.MethodPost { - activateCalled = true - w.WriteHeader(http.StatusOK) - return - } - w.Header().Set("Content-Type", "application/json") - w.WriteHeader(http.StatusOK) - _, _ = w.Write([]byte(`{"id":"cluster-uuid-ready","status":"active"}`)) - })) - defer srv.Close() - apiClient := webapi.NewClient(srv.URL) - - err := maybeActivateCluster(context.Background(), apiClient, "cluster-uuid-ready", nodeSet, r) - if err != nil { - t.Fatalf("maybeActivateCluster returned unexpected error: %v", err) - } - if !activateCalled { - t.Fatalf("expected the gate to let activation proceed once failure domains are ready") - } -} diff --git a/operator/internal/controller/simplyblockstoragenodeset_drain.go b/operator/internal/controller/simplyblockstoragenodeset_drain.go deleted file mode 100644 index b65bb2240..000000000 --- a/operator/internal/controller/simplyblockstoragenodeset_drain.go +++ /dev/null @@ -1,305 +0,0 @@ -/* -Copyright 2025. - -Licensed under the Apache License, Version 2.0 (the "License"); -you may not use this file except in compliance with the License. -You may obtain a copy of the License at - - http://www.apache.org/licenses/LICENSE-2.0 - -Unless required by applicable law or agreed to in writing, software -distributed under the License is distributed on an "AS IS" BASIS, -WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. -See the License for the specific language governing permissions and -limitations under the License. -*/ - -package controller - -import ( - "context" - "encoding/json" - "fmt" - "net/http" - "regexp" - "strings" - "time" - - corev1 "k8s.io/api/core/v1" - "k8s.io/apimachinery/pkg/types" - "sigs.k8s.io/controller-runtime/pkg/client" - logf "sigs.k8s.io/controller-runtime/pkg/log" - - "github.com/simplyblock/atlas/kube" - - simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" - "github.com/simplyblock/simplyblock-operator/internal/utils" - "github.com/simplyblock/simplyblock-operator/internal/webapi" -) - -// defaultSystemVolumeFilter is compiled once at package init from the -// well-known default pattern. A MustCompile panics at startup if the constant -// is malformed — intentional fast-fail for a hardcoded value. -var defaultSystemVolumeFilter = regexp.MustCompile(simplyblockv1alpha1.DefaultSystemVolumeFilterRegex) - -// Requeue intervals used by the drain state machine. -const ( - drainRequeueImmediate = 1 * time.Second - drainRequeueSuspend = 10 * time.Second - drainRequeueMigrate = 15 * time.Second - drainRequeueMigrateNew = 10 * time.Second - drainRequeueVerify = 30 * time.Second - drainRequeueBlocking = 60 * time.Second - drainRequeueValidate = 30 * time.Second -) - -// fetchPoolVolumes fetches all pools and returns (pools, nodeVolumes, err). -// Callers that need both the pool list (e.g. for cleanup) and the node volumes -// should call this once and reuse the returned pools, avoiding a second -// GetStoragePools round-trip within the same reconcile. -func fetchPoolVolumes( - ctx context.Context, - apiClient *webapi.Client, - clusterUUID string, - nodeUUID string, -) (pools []webapi.StoragePoolInfo, nodeVols []webapi.VolumeInfo, err error) { - pools, err = apiClient.GetStoragePools(ctx, clusterUUID) - if err != nil { - return nil, nil, fmt.Errorf("listNodeVolumes: %w", err) - } - for _, pool := range pools { - vols, err := apiClient.GetPoolVolumes(ctx, clusterUUID, pool.UUID) - if err != nil { - return nil, nil, fmt.Errorf("listNodeVolumes: pool %s: %w", pool.UUID, err) - } - for _, v := range vols { - if v.PrimaryNodeUUID != nodeUUID { - continue - } - // Skip volumes already being deleted — backend deletion is async so - // the volume may still appear in the list briefly after DELETE 204. - if v.Status == "in_deletion" { - continue - } - nodeVols = append(nodeVols, v) - } - } - return pools, nodeVols, nil -} - -// listNodeVolumes returns volumes on nodeUUID. Use fetchPoolVolumes when the -// pool list is also needed (e.g. drainVerify cleanup) to avoid a double fetch. -func listNodeVolumes( - ctx context.Context, - apiClient *webapi.Client, - clusterUUID string, - nodeUUID string, -) ([]webapi.VolumeInfo, error) { - _, vols, err := fetchPoolVolumes(ctx, apiClient, clusterUUID, nodeUUID) - return vols, err -} - -// matchVolumesToPVs classifies each backend volume into one of three buckets: -// - pvManaged: volume UUID has a corresponding simplyblock PV, not pinned -// - pinned: volume UUID has a PV but the PVC has the pinned-volume annotation -// - unmanaged: volume UUID has no corresponding PV (and is not a system volume) -// -// System volumes (those matching filterRegex by name) are skipped entirely. -// pvNameByVolumeUUID maps volume UUID → PV name for the pvManaged and pinned buckets. -// matchVolumesToPVs classifies each backend volume into pvManaged, pinned, or -// unmanaged buckets. System volumes matching filterRegex are skipped entirely. -// -// Note: if the PVC fetch for a PV-backed volume fails (e.g. API server -// temporarily unavailable), that volume is conservatively placed in the -// unmanaged bucket. This will block drain with an UnmanagedVolumeBlocking -// event until the next reconcile succeeds. It is a transient false-positive, -// not a permanent classification. -// matchVolumesToPVs classifies backend volumes and additionally returns -// pvcFetchFailed=true when at least one PV-backed volume could not be classified -// because its PVC GET failed transiently. Callers in Migrating must requeue on -// pvcFetchFailed to avoid silently skipping volumes that would stall Verifying. -func matchVolumesToPVs( - ctx context.Context, - c client.Client, - volumes []webapi.VolumeInfo, - sysFilter *regexp.Regexp, -) (pvManaged, pinned, unmanaged []string, pvNameByVolumeUUID map[string]string, pvcFetchFailed bool, err error) { - log := logf.FromContext(ctx) - - pvNameByVolumeUUID = make(map[string]string) - - // Build a map: volumeUUID → pvName for all simplyblock PVs. - var pvList corev1.PersistentVolumeList - if err = c.List(ctx, &pvList); err != nil { - return nil, nil, nil, nil, false, fmt.Errorf("matchVolumesToPVs: list PVs: %w", err) - } - - // pvByVolumeUUID maps volume UUID → PV object for simplyblock PVs. - pvByVolumeUUID := make(map[string]*corev1.PersistentVolume, len(pvList.Items)) - for i := range pvList.Items { - pv := &pvList.Items[i] - if pv.Spec.CSI == nil || pv.Spec.CSI.Driver != utils.CSIProvisioner { - continue - } - // VolumeHandle format: clusterUUID:poolUUID:volumeUUID — extract last segment. - volHandle := pv.Spec.CSI.VolumeHandle - if volHandle != "" { - parts := strings.SplitN(volHandle, ":", 3) - volumeUUID := parts[len(parts)-1] - if volumeUUID != "" { - pvByVolumeUUID[volumeUUID] = pv - } - } - } - - for _, vol := range volumes { - // System volume: skip entirely. - if sysFilter.MatchString(vol.Name) { - continue - } - - pv, isManagedByCSI := pvByVolumeUUID[vol.UUID] - if !isManagedByCSI { - unmanaged = append(unmanaged, vol.UUID) - continue - } - - // PV exists — check if the PVC has the pinned annotation. - if pv.Spec.ClaimRef == nil { - // PV has no claim; treat as PV-managed (not pinned). - pvManaged = append(pvManaged, vol.UUID) - pvNameByVolumeUUID[vol.UUID] = pv.Name - continue - } - - var pvc corev1.PersistentVolumeClaim - if err := c.Get(ctx, types.NamespacedName{ - Namespace: pv.Spec.ClaimRef.Namespace, - Name: pv.Spec.ClaimRef.Name, - }, &pvc); err != nil { - log.Error(err, "matchVolumesToPVs: failed to get PVC, treating volume as unmanaged", - "pvc", pv.Spec.ClaimRef.Name, "namespace", pv.Spec.ClaimRef.Namespace) - unmanaged = append(unmanaged, vol.UUID) - pvcFetchFailed = true - continue - } - - if kube.IsPinnedVolume(pvc.Annotations) { - pinned = append(pinned, vol.UUID) - pvNameByVolumeUUID[vol.UUID] = pv.Name - } else { - pvManaged = append(pvManaged, vol.UUID) - pvNameByVolumeUUID[vol.UUID] = pv.Name - } - } - - return pvManaged, pinned, unmanaged, pvNameByVolumeUUID, pvcFetchFailed, nil -} - -// getNodeBackendStatus fetches the current status string of a single storage -// node directly from the backend API. -func getNodeBackendStatus( - ctx context.Context, - apiClient *webapi.Client, - clusterUUID, nodeUUID string, -) (string, error) { - endpoint := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s", clusterUUID, nodeUUID) - body, status, err := apiClient.Do(ctx, http.MethodGet, endpoint, nil) - if err != nil { - return "", fmt.Errorf("getNodeBackendStatus: %w", err) - } - if status >= 300 { - return "", fmt.Errorf("getNodeBackendStatus: status %d", status) - } - var resp utils.NodeStatusResponse - if err := json.Unmarshal(body, &resp); err != nil { - return "", fmt.Errorf("getNodeBackendStatus: unmarshal: %w", err) - } - return resp.Status, nil -} - -// roundRobinTargetNodes lists all online nodes (excluding the drained node) and -// assigns each PV name a target node UUID using round-robin order. The i-th PV -// in pvNames is assigned to onlineNodes[i % len(onlineNodes)], distributing -// migrations evenly across the cluster without requiring persistent state. -// Returns an error if no online peer node is available. -func roundRobinTargetNodes( - ctx context.Context, - apiClient *webapi.Client, - clusterUUID string, - excludeNodeUUID string, - pvNames []string, -) (map[string]string, error) { - nodes, err := apiClient.GetStorageNodes(ctx, clusterUUID) - if err != nil { - return nil, fmt.Errorf("roundRobinTargetNodes: %w", err) - } - - var online []string - for _, n := range nodes { - if n.UUID != excludeNodeUUID && n.Status == utils.NodeStatusOnline { - online = append(online, n.UUID) - } - } - if len(online) == 0 { - return nil, fmt.Errorf("roundRobinTargetNodes: no online node available other than %s", excludeNodeUUID) - } - - assignment := make(map[string]string, len(pvNames)) - for i, pv := range pvNames { - assignment[pv] = online[i%len(online)] - } - return assignment, nil -} - -// drainMigrationName builds a DNS-label-safe name for a VolumeMigration CR. -func drainMigrationName(nodeUUID, pvName string) string { - prefix := "drain-" - if len(nodeUUID) >= 8 { - prefix += nodeUUID[:8] + "-" - } - name := prefix + pvName - name = strings.ToLower(name) - // Replace invalid chars with '-'. - var result []byte - for i := 0; i < len(name); i++ { - c := name[i] - if (c >= 'a' && c <= 'z') || (c >= '0' && c <= '9') || c == '-' { - result = append(result, c) - } else { - result = append(result, '-') - } - } - s := strings.Trim(string(result), "-") - - // Guard against name collisions when two PV names share a long common prefix - // that gets truncated to the same 63-char string. Append a 6-char FNV-32 - // hash of the original pvName before truncating so each PV always maps to a - // unique CR name regardless of length. - const maxLen = 63 - if len(s) > maxLen { - h := fnv32Hash(pvName) - suffix := fmt.Sprintf("-%06x", h) // 7 chars: '-' + 6 hex digits - keep := maxLen - len(suffix) - if keep < 0 { - keep = 0 - } - s = s[:keep] + suffix - } - return s -} - -// fnv32Hash returns a non-cryptographic 32-bit FNV-1a hash of s, used as a -// short disambiguation suffix in drainMigrationName. -func fnv32Hash(s string) uint32 { - const ( - offset32 uint32 = 2166136261 - prime32 uint32 = 16777619 - ) - h := offset32 - for i := 0; i < len(s); i++ { - h ^= uint32(s[i]) - h *= prime32 - } - return h -} diff --git a/operator/internal/controller/simplyblockstoragenodeset_drain_unit_test.go b/operator/internal/controller/simplyblockstoragenodeset_drain_unit_test.go deleted file mode 100644 index 3ef60016f..000000000 --- a/operator/internal/controller/simplyblockstoragenodeset_drain_unit_test.go +++ /dev/null @@ -1,311 +0,0 @@ -package controller - -import ( - "context" - "net/http" - "regexp" - "strings" - "testing" - - corev1 "k8s.io/api/core/v1" - metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" - "k8s.io/client-go/tools/events" - "sigs.k8s.io/controller-runtime/pkg/client" - - "github.com/simplyblock/atlas/kube" - - simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" - simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" - "github.com/simplyblock/simplyblock-operator/internal/utils" - "github.com/simplyblock/simplyblock-operator/internal/webapi" - webapimock "github.com/simplyblock/simplyblock-operator/internal/webapi/mock" -) - -// ── helpers ────────────────────────────────────────────────────────────────── - -const ( - drainTestNS = "test" - drainTestCluster = "test-cluster" - drainTestClusterUUID = "cccc0000-0000-0000-0000-000000000001" - drainTestNodeUUID = "aaaa0000-0000-0000-0000-000000000001" - drainTestNodeUUID2 = "aaaa0000-0000-0000-0000-000000000002" -) - -func newDrainReconciler(t *testing.T, objects ...client.Object) *StorageNodeSetReconciler { - t.Helper() - scheme := newTestScheme(t, - simplyblockv1alpha1.AddToScheme, - corev1.AddToScheme, - ) - cluster := testCluster(drainTestNS, drainTestCluster, drainTestClusterUUID) - all := append([]client.Object{cluster}, objects...) - cl := newTestClient(t, scheme, []client.Object{ - &simplyblockv1alpha1.StorageNodeSet{}, - &simplyblockv1alpha2.StorageCluster{}, - &simplyblockv1alpha1.VolumeMigration{}, - }, all...) - return &StorageNodeSetReconciler{ - Client: cl, - Scheme: scheme, - Namespace: drainTestNS, - Recorder: events.NewFakeRecorder(32), - } -} - -func TestRoundRobinDistributesEvenly(t *testing.T) { - mock := webapimock.NewSpecServerFromFile(t, "../../../shared/openapi.json", true) - defer mock.Close() - mock.Register(http.MethodGet, - "/api/v2/clusters/"+drainTestClusterUUID+"/storage-nodes/", - webapimock.RouteResponse{Status: http.StatusOK, Body: `[ - {"id":"node-1","status":"online"}, - {"id":"node-2","status":"online"}, - {"id":"node-3","status":"online"} - ]`}, - ) - - pvNames := []string{"pv-a", "pv-b", "pv-c", "pv-d", "pv-e", "pv-f"} - excluded := "node-1" - assignment, err := roundRobinTargetNodes(context.Background(), webapi.NewClient(mock.URL()), drainTestClusterUUID, excluded, pvNames) - if err != nil { - t.Fatalf("unexpected error: %v", err) - } - // excluded node should not appear as a target - for pv, target := range assignment { - if target == excluded { - t.Errorf("pv %s assigned to excluded node %s", pv, excluded) - } - } - // all pvNames must be assigned - if len(assignment) != len(pvNames) { - t.Errorf("expected %d assignments, got %d", len(pvNames), len(assignment)) - } - // each of node-2 and node-3 should appear 3 times (6 pvs / 2 nodes) - counts := map[string]int{} - for _, target := range assignment { - counts[target]++ - } - for _, node := range []string{"node-2", "node-3"} { - if counts[node] != 3 { - t.Errorf("node %s expected 3 assignments, got %d", node, counts[node]) - } - } -} - -func TestRoundRobinErrorsWhenNoTargetAvailable(t *testing.T) { - mock := webapimock.NewSpecServerFromFile(t, "../../../shared/openapi.json", true) - defer mock.Close() - // Only one node, and it is the excluded one. - mock.Register(http.MethodGet, - "/api/v2/clusters/"+drainTestClusterUUID+"/storage-nodes/", - webapimock.RouteResponse{Status: http.StatusOK, Body: `[ - {"id":"node-1","status":"online"} - ]`}, - ) - - _, err := roundRobinTargetNodes(context.Background(), webapi.NewClient(mock.URL()), drainTestClusterUUID, "node-1", []string{"pv-a"}) - if err == nil { - t.Fatal("expected error when no online peer node is available") - } -} - -func TestRoundRobinSkipsOfflineNodes(t *testing.T) { - mock := webapimock.NewSpecServerFromFile(t, "../../../shared/openapi.json", true) - defer mock.Close() - mock.Register(http.MethodGet, - "/api/v2/clusters/"+drainTestClusterUUID+"/storage-nodes/", - webapimock.RouteResponse{Status: http.StatusOK, Body: `[ - {"id":"node-1","status":"online"}, - {"id":"node-2","status":"offline"}, - {"id":"node-3","status":"online"} - ]`}, - ) - - assignment, err := roundRobinTargetNodes(context.Background(), webapi.NewClient(mock.URL()), drainTestClusterUUID, "node-1", []string{"pv-a", "pv-b"}) - if err != nil { - t.Fatalf("unexpected error: %v", err) - } - for pv, target := range assignment { - if target == "node-2" { - t.Errorf("pv %s assigned to offline node-2", pv) - } - if target == "node-1" { - t.Errorf("pv %s assigned to excluded node-1", pv) - } - } - _ = assignment -} - -// ── matchVolumesToPVs ───────────────────────────────────────────────────────── - -func newPV(name, volumeUUID string) *corev1.PersistentVolume { - pv := &corev1.PersistentVolume{ - ObjectMeta: metav1.ObjectMeta{Name: name}, - } - pv.Spec.CSI = &corev1.CSIPersistentVolumeSource{ - Driver: utils.CSIProvisioner, - VolumeHandle: drainTestClusterUUID + ":pool-1:" + volumeUUID, - } - pv.Spec.ClaimRef = &corev1.ObjectReference{ - Namespace: drainTestNS, - Name: name + "-pvc", - } - return pv -} - -func newPVC(name string, pinned bool) *corev1.PersistentVolumeClaim { - pvc := &corev1.PersistentVolumeClaim{ - ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: drainTestNS}, - } - if pinned { - pvc.Annotations = map[string]string{ - kube.AnnoSelectedStorageNode: "true", - } - } - return pvc -} - -func TestMatchVolumesToPVs_PVManaged(t *testing.T) { - pv := newPV("pv-a", "vol-1111") - pvc := newPVC("pv-a-pvc", false) - r := newDrainReconciler(t, pv, pvc) - - vols := []webapi.VolumeInfo{{UUID: "vol-1111", Name: "pvc-something"}} - pvManaged, pinned, unmanaged, byUUID, _, err := matchVolumesToPVs(context.Background(), r.Client, vols, regexp.MustCompile("^never-matches$")) - if err != nil { - t.Fatalf("unexpected error: %v", err) - } - if len(pvManaged) != 1 || pvManaged[0] != "vol-1111" { - t.Errorf("expected vol-1111 in pvManaged, got %v", pvManaged) - } - if len(pinned) != 0 || len(unmanaged) != 0 { - t.Errorf("expected no pinned/unmanaged, got pinned=%v unmanaged=%v", pinned, unmanaged) - } - if byUUID["vol-1111"] != "pv-a" { - t.Errorf("expected pvName=pv-a, got %q", byUUID["vol-1111"]) - } -} - -func TestMatchVolumesToPVs_Pinned(t *testing.T) { - pv := newPV("pv-b", "vol-2222") - pvc := newPVC("pv-b-pvc", true) // pinned - r := newDrainReconciler(t, pv, pvc) - - vols := []webapi.VolumeInfo{{UUID: "vol-2222", Name: "pvc-something"}} - pvManaged, pinned, unmanaged, _, _, err := matchVolumesToPVs(context.Background(), r.Client, vols, regexp.MustCompile("^never-matches$")) - if err != nil { - t.Fatalf("unexpected error: %v", err) - } - if len(pinned) != 1 || pinned[0] != "vol-2222" { - t.Errorf("expected vol-2222 in pinned, got %v", pinned) - } - if len(pvManaged) != 0 || len(unmanaged) != 0 { - t.Errorf("expected no pvManaged/unmanaged, got pvManaged=%v unmanaged=%v", pvManaged, unmanaged) - } -} - -func TestMatchVolumesToPVs_Unmanaged(t *testing.T) { - r := newDrainReconciler(t) // no PVs in cluster - - vols := []webapi.VolumeInfo{{UUID: "vol-orphan", Name: "manually-created"}} - pvManaged, pinned, unmanaged, _, _, err := matchVolumesToPVs(context.Background(), r.Client, vols, regexp.MustCompile("^never-matches$")) - if err != nil { - t.Fatalf("unexpected error: %v", err) - } - if len(unmanaged) != 1 || unmanaged[0] != "vol-orphan" { - t.Errorf("expected vol-orphan in unmanaged, got %v", unmanaged) - } - if len(pvManaged) != 0 || len(pinned) != 0 { - t.Errorf("unexpected pvManaged/pinned: %v / %v", pvManaged, pinned) - } -} - -func TestMatchVolumesToPVs_SystemVolumeSkipped(t *testing.T) { - r := newDrainReconciler(t) // no PVs — if not filtered, would be unmanaged - - vols := []webapi.VolumeInfo{{UUID: "vol-bench", Name: "sb-fio-baseline-xyz"}} - pvManaged, pinned, unmanaged, _, _, err := matchVolumesToPVs(context.Background(), r.Client, vols, defaultSystemVolumeFilter) - if err != nil { - t.Fatalf("unexpected error: %v", err) - } - if len(pvManaged)+len(pinned)+len(unmanaged) != 0 { - t.Errorf("system volume should be skipped entirely, got pvManaged=%v pinned=%v unmanaged=%v", pvManaged, pinned, unmanaged) - } -} - -func TestMatchVolumesToPVs_EmptyNodeSkipsMigration(t *testing.T) { - r := newDrainReconciler(t) - pvManaged, pinned, unmanaged, _, _, err := matchVolumesToPVs(context.Background(), r.Client, nil, defaultSystemVolumeFilter) - if err != nil { - t.Fatalf("unexpected error: %v", err) - } - if len(pvManaged)+len(pinned)+len(unmanaged) != 0 { - t.Errorf("empty node should produce no buckets") - } -} - -func TestMatchVolumesToPVs_OnlySystemVolumes(t *testing.T) { - r := newDrainReconciler(t) - vols := []webapi.VolumeInfo{ - {UUID: "v1", Name: "sb-fio-baseline-read"}, - {UUID: "v2", Name: "sb-fio-baseline-write"}, - } - pvManaged, pinned, unmanaged, _, _, err := matchVolumesToPVs(context.Background(), r.Client, vols, defaultSystemVolumeFilter) - if err != nil { - t.Fatalf("unexpected error: %v", err) - } - if len(pvManaged)+len(pinned)+len(unmanaged) != 0 { - t.Errorf("system-only node should produce no drain work") - } -} - -func TestDrainMigrationNameNoCollisionOnLongPVNames(t *testing.T) { - // Two PV names that share a 60+ char common prefix must produce distinct CR - // names after sanitisation and truncation (collision guard via FNV suffix). - longBase := "pvc-" + strings.Repeat("a", 55) // 59 chars — produces a 63-char name when prefixed - pv1 := longBase + "1" - pv2 := longBase + "2" - nodeUUID := "aaaabbbb-cccc-dddd-eeee-ffffffffffff" - - name1 := drainMigrationName(nodeUUID, pv1) - name2 := drainMigrationName(nodeUUID, pv2) - - if name1 == name2 { - t.Errorf("collision: both PVs produced the same CR name %q", name1) - } - if len(name1) > 63 { - t.Errorf("name1 too long: %d chars", len(name1)) - } - if len(name2) > 63 { - t.Errorf("name2 too long: %d chars", len(name2)) - } -} - -func TestDrainMigrationNameIsDNSValid(t *testing.T) { - cases := []struct { - nodeUUID string - pvName string - }{ - {"afc7286e-ca84-42f1-bc8f-c582ad2a9a9e", "pvc-a62c57bc-f64c-4385-ace4-f84b729fc8ee"}, - {"short", "pvc-simple"}, - {"", "pvc-no-node"}, - {"uuid", "PVC-Upper-Case"}, - } - for _, tc := range cases { - name := drainMigrationName(tc.nodeUUID, tc.pvName) - if len(name) > 63 { - t.Errorf("name too long (%d): %q", len(name), name) - } - if len(name) == 0 { - t.Errorf("empty name for nodeUUID=%q pvName=%q", tc.nodeUUID, tc.pvName) - } - for _, c := range name { - if (c < 'a' || c > 'z') && (c < '0' || c > '9') && c != '-' { - t.Errorf("invalid char %q in name %q", c, name) - } - } - if name[0] == '-' || name[len(name)-1] == '-' { - t.Errorf("name starts or ends with '-': %q", name) - } - } -} diff --git a/operator/internal/controller/simplyblockstoragenodeset_pernodeconfig.go b/operator/internal/controller/simplyblockstoragenodeset_pernodeconfig.go deleted file mode 100644 index 99a5fd8ae..000000000 --- a/operator/internal/controller/simplyblockstoragenodeset_pernodeconfig.go +++ /dev/null @@ -1,260 +0,0 @@ -/* -Copyright 2025. - -Licensed under the Apache License, Version 2.0 (the "License"); -you may not use this file except in compliance with the License. -You may obtain a copy of the License at - - http://www.apache.org/licenses/LICENSE-2.0 - -Unless required by applicable law or agreed to in writing, software -distributed under the License is distributed on an "AS IS" BASIS, -WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. -See the License for the specific language governing permissions and -limitations under the License. -*/ - -package controller - -// reconcilePerNodeConfigMap creates or updates a single ConfigMap that holds -// per-worker-node effective configuration values. The DaemonSet init container -// mounts this ConfigMap and sources the file matching its hostname so that -// fields like deviceNames, pcieAllowList, etc. differ per node without -// requiring a separate DaemonSet per node. -// -// MAX_SUBSYS_COUNT, MAX_HUGE_PAGES_SIZE and VCPU_COUNT are cluster-scoped -// (StorageCluster.spec) rather than per-node: they size huge pages and the SPDK -// core layout, which the control plane assumes uniform across the cluster. They -// are still written into every per-node entry because the init container reads -// its whole configuration from this one file. -// -// ConfigMap structure: -// -// data: -// vm02.example.com: | -// MAX_SUBSYS_COUNT=20 -// MAX_HUGE_PAGES_SIZE= -// VCPU_COUNT=8 -// ... -// vm03.example.com: | -// MAX_SUBSYS_COUNT=20 -// ... - -import ( - "context" - "fmt" - "strings" - - "github.com/simplyblock/simplyblock-operator/internal/utils" - corev1 "k8s.io/api/core/v1" - apierrors "k8s.io/apimachinery/pkg/api/errors" - metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" - "sigs.k8s.io/controller-runtime/pkg/client" - "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil" - logf "sigs.k8s.io/controller-runtime/pkg/log" - - "github.com/simplyblock/atlas/kube" - "github.com/simplyblock/atlas/ptr" - - simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" - simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" -) - -// PerNodeConfigMapName returns the name of the per-node ConfigMap for a StorageNodeSet. -func PerNodeConfigMapName(snsName string) string { - return snsName + "-per-node-config" -} - -// reconcilePerNodeConfigMap creates or updates the per-node ConfigMap with the -// effective (fleet defaults merged with nodeConfigs overrides) values for every -// worker in the StorageNodeSet. -func (r *StorageNodeSetReconciler) reconcilePerNodeConfigMap( - ctx context.Context, - sns *simplyblockv1alpha1.StorageNodeSet, -) error { - log := logf.FromContext(ctx) - name := PerNodeConfigMapName(sns.Name) - - // MAX_SUBSYS_COUNT, MAX_HUGE_PAGES_SIZE and VCPU_COUNT come from the StorageCluster, - // so every entry this ConfigMap holds shares them. - var cluster simplyblockv1alpha2.StorageCluster - if err := r.Get(ctx, client.ObjectKey{ - Name: sns.Spec.ClusterName, - Namespace: sns.Namespace, - }, &cluster); err != nil { - return fmt.Errorf("getting StorageCluster %q for per-node ConfigMap: %w", sns.Spec.ClusterName, err) - } - - // maxSubsystemCount and vcpuCount are required by the CRD schema, so they can - // only be nil on a StorageCluster admitted before that requirement existed. - // Refuse to write a config the node cannot boot from: an empty MAX_SUBSYS_COUNT - // reaches node_configure.py as --max-subsys-count=0 and fails config - // generation there, far from the cause. - if cluster.Spec.MaxSubsystemCount == nil || cluster.Spec.VCPUCount == nil { - return fmt.Errorf( - "StorageCluster %q is missing required node sizing: set spec.maxSubsystemCount and spec.vcpuCount", - sns.Spec.ClusterName) - } - - data := make(map[string]string, len(sns.Spec.WorkerNodes)) - for _, worker := range sns.Spec.WorkerNodes { - data[worker] = buildPerNodeEnvFile(&cluster, sns, worker) - } - - // Also include manually created StorageNode CRs that reference this StorageNodeSet - // but whose worker is not in spec.workerNodes. Their overrides are merged with - // fleet defaults so the DaemonSet init container gets the right per-node config. - var snList simplyblockv1alpha1.StorageNodeList - if err := r.List(ctx, &snList, - client.InNamespace(sns.Namespace), - client.MatchingFields{"spec.storageNodeSetRef": sns.Name}, - ); err == nil { - for _, sn := range snList.Items { - if _, ok := data[sn.Spec.WorkerNode]; ok { - continue // already covered by spec.workerNodes - } - // Manually created: use its overrides on top of fleet defaults. - snsCopy := sns.DeepCopy() - if sn.Spec.Overrides != nil { - if snsCopy.Spec.NodeConfigs == nil { - snsCopy.Spec.NodeConfigs = make(map[string]simplyblockv1alpha1.StorageNodeOverrides) - } - snsCopy.Spec.NodeConfigs[sn.Spec.WorkerNode] = *sn.Spec.Overrides - } - data[sn.Spec.WorkerNode] = buildPerNodeEnvFile(&cluster, snsCopy, sn.Spec.WorkerNode) - } - } - - var existing corev1.ConfigMap - err := r.Get(ctx, client.ObjectKey{Name: name, Namespace: sns.Namespace}, &existing) - if err != nil && !apierrors.IsNotFound(err) { - return fmt.Errorf("getting per-node ConfigMap: %w", err) - } - - // Include workers labeled into this StorageNodeSet but not yet in - // spec.workerNodes or backed by a StorageNode CR — an in-flight StorageNodeOps - // migration labels its target and ensureMigratedWorkerConfig authors the - // target's per-node entry (source clone + newSsdPcie merged into PCI_ALLOWED) - // during Preparing, before reconcileMigratedTopology folds the target into - // spec.workerNodes once it is online. Without this, the full rebuild above - // drops that entry, and when the target's DaemonSet pod (re)starts its init - // container sources an empty env and config generation fails on max-lvol=0. - // Prefer the existing entry so the clone/newSsdPcie survive the rebuild; fall - // back to fleet defaults so a target whose entry was already lost still boots - // with a valid config. Mirrors reconcileEndpointSlice's label-based inclusion. - var labeledNodes corev1.NodeList - if listErr := r.List(ctx, &labeledNodes, client.MatchingLabels{kube.LabelStorageNodeSet: sns.Name}); listErr != nil { - return fmt.Errorf("listing storage-plane nodes for per-node ConfigMap: %w", listErr) - } - for i := range labeledNodes.Items { - worker := labeledNodes.Items[i].Name - if _, ok := data[worker]; ok { - continue // already covered by spec.workerNodes or a StorageNode CR - } - if entry, ok := existing.Data[worker]; ok { - data[worker] = entry - } else { - data[worker] = buildPerNodeEnvFile(&cluster, sns, worker) - } - } - - if apierrors.IsNotFound(err) { - cm := &corev1.ConfigMap{ - ObjectMeta: metav1.ObjectMeta{ - Name: name, - Namespace: sns.Namespace, - }, - Data: data, - } - if setErr := controllerutil.SetControllerReference(sns, cm, r.Scheme); setErr != nil { - return fmt.Errorf("setting owner reference on per-node ConfigMap: %w", setErr) - } - if createErr := r.Create(ctx, cm); createErr != nil { - return fmt.Errorf("creating per-node ConfigMap: %w", createErr) - } - log.Info("created per-node ConfigMap", "name", name) - return nil - } - - // Update if data changed. - patch := client.MergeFrom(existing.DeepCopy()) - existing.Data = data - if patchErr := r.Patch(ctx, &existing, patch); patchErr != nil { - return fmt.Errorf("patching per-node ConfigMap: %w", patchErr) - } - return nil -} - -// buildPerNodeEnvFile returns a shell-sourceable env file string with the -// effective per-node values for the given worker, merging fleet defaults from -// the StorageNodeSet spec with any nodeConfigs overrides. The huge-page and -// core-sizing values (MAX_SUBSYS_COUNT, MAX_HUGE_PAGES_SIZE, VCPU_COUNT) come from the -// StorageCluster and are therefore identical in every entry. -func buildPerNodeEnvFile( - cluster *simplyblockv1alpha2.StorageCluster, - sns *simplyblockv1alpha1.StorageNodeSet, - worker string, -) string { - // Start with fleet defaults. - eff := simplyblockv1alpha1.StorageNodeOverrides{ - SpdkSystemMemory: sns.Spec.SpdkSystemMemory, - JournalManagerSpec: sns.Spec.JournalManagerSpec, - PcieAllowList: sns.Spec.PcieAllowList, - PcieDenyList: sns.Spec.PcieDenyList, - PcieModel: sns.Spec.PcieModel, - DriveSizeRange: sns.Spec.DriveSizeRange, - DeviceNames: sns.Spec.DeviceNames, - EnableCpuTopology: sns.Spec.EnableCpuTopology, - ReservedSystemCPU: sns.Spec.ReservedSystemCPU, - } - - // Apply per-node overrides if present. - if o, ok := sns.Spec.NodeConfigs[worker]; ok { - if o.SpdkSystemMemory != "" { - eff.SpdkSystemMemory = o.SpdkSystemMemory - } - if o.JournalManagerSpec != nil { - eff.JournalManagerSpec = o.JournalManagerSpec - } - if len(o.PcieAllowList) > 0 { - eff.PcieAllowList = o.PcieAllowList - } - if len(o.PcieDenyList) > 0 { - eff.PcieDenyList = o.PcieDenyList - } - if o.PcieModel != "" { - eff.PcieModel = o.PcieModel - } - if o.DriveSizeRange != "" { - eff.DriveSizeRange = o.DriveSizeRange - } - if len(o.DeviceNames) > 0 { - eff.DeviceNames = o.DeviceNames - } - if o.EnableCpuTopology != nil { - eff.EnableCpuTopology = o.EnableCpuTopology - } - if o.ReservedSystemCPU != "" { - eff.ReservedSystemCPU = o.ReservedSystemCPU - } - } - - var b strings.Builder - // Cluster-scoped: identical for every worker in every set of this cluster. - fmt.Fprintf(&b, "MAX_SUBSYS_COUNT=%s\n", ptr.StringOrDefault(cluster.Spec.MaxSubsystemCount, "")) - fmt.Fprintf(&b, "MAX_HUGE_PAGES_SIZE=%s\n", utils.ShellQuote(cluster.Spec.MinHugePagesSize)) - fmt.Fprintf(&b, "VCPU_COUNT=%s\n", ptr.StringOrDefault(cluster.Spec.VCPUCount, "")) - fmt.Fprintf(&b, "PCI_ALLOWED=%s\n", utils.ShellQuote(strings.Join(eff.PcieAllowList, ","))) - fmt.Fprintf(&b, "PCI_BLOCKED=%s\n", utils.ShellQuote(strings.Join(eff.PcieDenyList, ","))) - fmt.Fprintf(&b, "NVME_DEVICES=%s\n", utils.ShellQuote(strings.Join(eff.DeviceNames, ","))) - fmt.Fprintf(&b, "DEVICE_MODEL=%s\n", utils.ShellQuote(eff.PcieModel)) - fmt.Fprintf(&b, "SIZE_RANGE=%s\n", utils.ShellQuote(eff.DriveSizeRange)) - if eff.JournalManagerSpec != nil { - fmt.Fprintf(&b, "JM_PERCENT=%s\n", ptr.StringOrDefault(eff.JournalManagerSpec.PercentPerDevice, "")) - fmt.Fprintf(&b, "HA_JM_COUNT=%s\n", ptr.StringOrDefault(eff.JournalManagerSpec.Count, "")) - } else { - b.WriteString("JM_PERCENT=\n") - b.WriteString("HA_JM_COUNT=\n") - } - return b.String() -} diff --git a/operator/internal/controller/simplyblockstoragenodeset_storagenode.go b/operator/internal/controller/simplyblockstoragenodeset_storagenode.go deleted file mode 100644 index 0d96bbe8f..000000000 --- a/operator/internal/controller/simplyblockstoragenodeset_storagenode.go +++ /dev/null @@ -1,582 +0,0 @@ -/* -Copyright 2025. - -Licensed under the Apache License, Version 2.0 (the "License"); -you may not use this file except in compliance with the License. -You may obtain a copy of the License at - - http://www.apache.org/licenses/LICENSE-2.0 - -Unless required by applicable law or agreed to in writing, software -distributed under the License is distributed on an "AS IS" BASIS, -WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. -See the License for the specific language governing permissions and -limitations under the License. -*/ - -package controller - -// reconcileStorageNodeCRs is the Phase-1 bridge that creates, updates, and -// deletes StorageNode CRs to match StorageNodeSet.spec.workerNodes × -// spec.socketsToUse. The StorageNodeSet is the single source of truth — this -// function only creates and garbage-collects; the StorageNodeReconciler owns -// the per-node provisioning and status-sync loops. -// -// Called from StorageNodeSetReconciler.Reconcile after fleet infrastructure -// (DaemonSet, RBAC, Services) has been reconciled. - -import ( - "context" - "fmt" - "strings" - - apierrors "k8s.io/apimachinery/pkg/api/errors" - metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" - "sigs.k8s.io/controller-runtime/pkg/client" - "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil" - logf "sigs.k8s.io/controller-runtime/pkg/log" - - atlaskube "github.com/simplyblock/atlas/kube" - simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" - "github.com/simplyblock/simplyblock-operator/internal/utils" -) - -// workerOrdinal uniquely identifies a backend storage node per worker host. -// ordinal = socketPosition * nodesPerSocket + nodeIndex. -type workerOrdinal struct { - worker string - ordinal int -} - -// reconcileStorageNodeCRs creates a StorageNode CR for every (worker, socket) -// pair in the StorageNodeSet spec and deletes any owned StorageNode CRs that no -// longer correspond to a configured worker/socket. Overrides are synced from -// spec.nodeConfigs on every call. -func (r *StorageNodeSetReconciler) reconcileStorageNodeCRs( - ctx context.Context, - sns *simplyblockv1alpha1.StorageNodeSet, -) error { - log := logf.FromContext(ctx) - - // Compute nodesPerSocket (default 1). - nodesPerSocket := 1 - if sns.Spec.NodesPerSocket != nil && *sns.Spec.NodesPerSocket > 1 { - nodesPerSocket = int(*sns.Spec.NodesPerSocket) - } - sockets := effectiveSockets(sns) - - // Build the expected set of (worker, globalOrdinal) pairs. - // globalOrdinal = socketPosition * nodesPerSocket + nodeIndex - // This uniquely identifies each backend storage node per worker host. - expected := make(map[workerOrdinal]struct{}, len(sns.Spec.WorkerNodes)*len(sockets)*nodesPerSocket) - for _, worker := range sns.Spec.WorkerNodes { - for si, socket := range sockets { - _ = socket // socket string is embedded in the ordinal calculation - for ni := 0; ni < nodesPerSocket; ni++ { - ordinal := si*nodesPerSocket + ni - expected[workerOrdinal{worker, ordinal}] = struct{}{} - } - } - } - - // List all StorageNode CRs owned by this StorageNodeSet. - var owned simplyblockv1alpha1.StorageNodeList - if err := r.List(ctx, &owned, - client.InNamespace(sns.Namespace), - client.MatchingFields{"spec.storageNodeSetRef": sns.Name}, - ); err != nil { - return fmt.Errorf("listing owned StorageNode CRs: %w", err) - } - - // Delete stale CRs (worker removed from spec.workerNodes, socket removed, - // or nodesPerSocket reduced). Manually created CRs (no OwnerReference) are kept. - r.deleteStaleStorageNodeCRs(ctx, sns, owned.Items, expected) - - // Index owned CRs by (worker, ordinal). Since the CR name is now a random id - // (simplyblock-node-) rather than derived from the worker/socket, a slot - // is identified by its spec fields, not by a computable name. - existingBySlot := make(map[workerOrdinal]*simplyblockv1alpha1.StorageNode, len(owned.Items)) - for i := range owned.Items { - sn := &owned.Items[i] - ord := 0 - if sn.Spec.SocketIndex != nil { - ord = int(*sn.Spec.SocketIndex) - } - existingBySlot[workerOrdinal{sn.Spec.WorkerNode, ord}] = sn - } - - // Determine which workers host FDB pods — added sequentially. - fdbWorkers := r.fdbWorkerSet(ctx, sns) - - fdbInFlight := false - for _, sn := range owned.Items { - if !fdbWorkers[sn.Spec.WorkerNode] { - continue - } - // A node is in-flight if it is actively being provisioned (in_creation) - // or has just been created (no UUID and no status yet). A timeout is a - // terminal failure — treat it like done so the next node can proceed. - if sn.Status.Status == utils.NodeStatusTimeout { - continue - } - if sn.Status.UUID == "" || sn.Status.Status == utils.NodeStatusInCreation { - fdbInFlight = true - break - } - } - - // Create or sync one StorageNode CR per (worker, socket, nodeIdx). - // globalOrdinal = socketPosition × nodesPerSocket + nodeIdx is stored as - // SocketIndex and used by pollUUIDFromBackend to select the correct backend node. - // - // FDB workers are created one at a time: after creating a new FDB CR in this - // reconcile pass we flip fdbInFlight=true so all subsequent FDB workers are - // deferred. This prevents the cache-lag race where multiple new CRs are - // created simultaneously and the StorageNode controller reconciles them all - // before any PostedAt is set, bypassing the sequential add gate. - for _, worker := range sns.Spec.WorkerNodes { - for si, socket := range sockets { - for ni := 0; ni < nodesPerSocket; ni++ { - globalOrdinal := si*nodesPerSocket + ni - existing := existingBySlot[workerOrdinal{worker, globalOrdinal}] - advanced := r.reconcileOneStorageNodeCR(ctx, sns, existing, worker, socket, ni, globalOrdinal, fdbWorkers, fdbInFlight) - if advanced { - fdbInFlight = true - } - } - } - } - - // Aggregate status: TotalNodes / OnlineNodes / OfflineNodes. - if err := r.aggregateStorageNodeStatus(ctx, sns); err != nil { - log.Error(err, "failed to aggregate StorageNode status into StorageNodeSet") - } - - return nil -} - -// deleteStaleStorageNodeCRs deletes owned StorageNode CRs whose (worker, ordinal) -// pair is not present in expected. Manually created CRs (no OwnerReference) are kept. -func (r *StorageNodeSetReconciler) deleteStaleStorageNodeCRs( - ctx context.Context, - sns *simplyblockv1alpha1.StorageNodeSet, - items []simplyblockv1alpha1.StorageNode, - expected map[workerOrdinal]struct{}, -) { - log := logf.FromContext(ctx) - for i := range items { - sn := &items[i] - ordinal := 0 - if sn.Spec.SocketIndex != nil { - ordinal = int(*sn.Spec.SocketIndex) - } - if _, ok := expected[workerOrdinal{sn.Spec.WorkerNode, ordinal}]; ok { - continue - } - owner := metav1.GetControllerOf(sn) - if owner == nil || owner.Name != sns.Name { - continue // manually created — preserve - } - if err := r.Delete(ctx, sn); err != nil && !apierrors.IsNotFound(err) { - log.Error(err, "failed to delete stale StorageNode CR", "name", sn.Name) - } else { - log.Info("deleted stale StorageNode CR", "name", sn.Name, - "worker", sn.Spec.WorkerNode, "ordinal", ordinal) - } - } -} - -// reconcileOneStorageNodeCR handles the FDB gate check, existedBefore check, and -// ensureStorageNodeCR call for a single (worker, socket, nodeIdx, globalOrdinal) tuple. -// existing is the CR already occupying this (worker, globalOrdinal) slot, or nil. -// It returns true if a brand-new FDB CR was just created (so the caller can set -// fdbInFlight=true to defer subsequent FDB workers). -func (r *StorageNodeSetReconciler) reconcileOneStorageNodeCR( - ctx context.Context, - sns *simplyblockv1alpha1.StorageNodeSet, - existing *simplyblockv1alpha1.StorageNode, - worker, socket string, - nodeIdx, globalOrdinal int, - fdbWorkers map[string]bool, - fdbInFlight bool, -) (newFDBInFlight bool) { - log := logf.FromContext(ctx) - - // The CR name is a random id (simplyblock-node-), so a slot is identified - // by its spec fields (worker, globalOrdinal) via existing, not by name. - if fdbWorkers[worker] && fdbInFlight && existing == nil { - log.Info("FDB worker: deferring StorageNode CR creation until previous FDB node is online", - "worker", worker, "socket", socket, "nodeIdx", nodeIdx) - return false - } - - existedBefore := existing != nil - - if err := r.ensureStorageNodeCR(ctx, sns, existing, worker, socket, nodeIdx, globalOrdinal); err != nil { - log.Error(err, "failed to ensure StorageNode CR", - "worker", worker, "socket", socket, "nodeIdx", nodeIdx) - return false - } - - // If this was a brand-new FDB CR (didn't exist before this reconcile), - // mark fdbInFlight so the remaining FDB workers are deferred until the - // next reconcile — preventing multiple new CRs from being created and - // reconciled simultaneously by the StorageNode controller. - if fdbWorkers[worker] && !existedBefore { - return true - } - return false -} - -// ensureStorageNodeCR creates or patches a StorageNode CR. -// socket is the NUMA socket identifier (from socketsToUse), nodeIdx is the -// per-socket node index (0..nodesPerSocket-1), and globalOrdinal is used as -// SocketIndex for backend node lookup in pollUUIDFromBackend. -func (r *StorageNodeSetReconciler) ensureStorageNodeCR( - ctx context.Context, - sns *simplyblockv1alpha1.StorageNodeSet, - existing *simplyblockv1alpha1.StorageNode, - worker, socket string, - nodeIdx, globalOrdinal int, -) error { - if existing == nil { - return r.createStorageNodeCR(ctx, sns, worker, socket, nodeIdx, globalOrdinal) - } - - // Sync overrides from nodeConfigs — the StorageNodeSet is the source of truth. - overrides, hasConfig := sns.Spec.NodeConfigs[worker] - desired := existing.Spec.Overrides - if hasConfig { - desired = &overrides - } - - patch := client.MergeFrom(existing.DeepCopy()) - existing.Spec.Overrides = desired - if err := r.Patch(ctx, existing, patch); err != nil && !apierrors.IsNotFound(err) { - return fmt.Errorf("patching StorageNode overrides %s: %w", existing.Name, err) - } - return nil -} - -// createStorageNodeCR creates a new StorageNode CR with a random, -// DNS-label-safe name (simplyblock-node-). Because the name is random, a -// collision with an existing object is possible; on AlreadyExists we regenerate -// the id and retry. -func (r *StorageNodeSetReconciler) createStorageNodeCR( - ctx context.Context, - sns *simplyblockv1alpha1.StorageNodeSet, - worker, socket string, - nodeIdx, globalOrdinal int, -) error { - log := logf.FromContext(ctx) - - const maxAttempts = 5 - var sn *simplyblockv1alpha1.StorageNode - var createErr error - for attempt := 0; attempt < maxAttempts; attempt++ { - name := storageNodeCRName(sns.Name) - sn = buildStorageNodeCR(sns, name, worker, socket, nodeIdx, globalOrdinal) - if err := controllerutil.SetControllerReference(sns, sn, r.Scheme); err != nil { - return fmt.Errorf("setting owner reference on StorageNode %s: %w", name, err) - } - createErr = r.Create(ctx, sn) - if createErr == nil { - log.Info("created StorageNode CR", "name", name, "worker", worker, "socket", socket, "nodeIdx", nodeIdx) - break - } - if !apierrors.IsAlreadyExists(createErr) { - return fmt.Errorf("creating StorageNode %s: %w", name, createErr) - } - log.Info("StorageNode CR name collision, regenerating id", "name", name) - } - if createErr != nil { - return fmt.Errorf("creating StorageNode for worker %s after %d attempts: %w", worker, maxAttempts, createErr) - } - - return nil -} - -// aggregateStorageNodeStatus rolls up online/offline counts from owned -// StorageNode CRs into StorageNodeSet.status. -func (r *StorageNodeSetReconciler) aggregateStorageNodeStatus( - ctx context.Context, - sns *simplyblockv1alpha1.StorageNodeSet, -) error { - var owned simplyblockv1alpha1.StorageNodeList - if err := r.List(ctx, &owned, - client.InNamespace(sns.Namespace), - client.MatchingFields{"spec.storageNodeSetRef": sns.Name}, - ); err != nil { - return err - } - - var online, offline, suspended, creating, removed int - for _, sn := range owned.Items { - switch sn.Status.Status { - case utils.NodeStatusOnline: - online++ - case "offline": - offline++ - case "suspended": - suspended++ - case nodeStatusInCreation: - creating++ - case "removed": - removed++ - } - } - - patch := client.MergeFrom(sns.DeepCopy()) - sns.Status.TotalNodes = len(owned.Items) - sns.Status.OnlineNodes = online - sns.Status.OfflineNodes = offline - sns.Status.SuspendedNodes = suspended - sns.Status.CreatingNodes = creating - sns.Status.RemovedNodes = removed - return r.Status().Patch(ctx, sns, patch) -} - -// storageNodeCRName builds a DNS-label-safe name for a StorageNode CR: -// "{sns}-{id}" where id is a random short id (atlas kube.NameWithID). The NUMA -// socket and per-socket node index are NOT encoded in the name — they live in -// the CR spec (SocketID/NodeIndex/SocketIndex) — so the name is stable when a -// storage node migrates between workers. Callers create with retry-on-collision. -func storageNodeCRName(snsName string) string { - return atlaskube.NameWithID(snsName) -} - -// sanitiseDNSLabel replaces characters not valid in a DNS label with '-' and -// strips leading/trailing hyphens. -func sanitiseDNSLabel(s string) string { - var b strings.Builder - for _, c := range strings.ToLower(s) { - if (c >= 'a' && c <= 'z') || (c >= '0' && c <= '9') || c == '-' || c == '.' { - b.WriteRune(c) - } else { - b.WriteByte('-') - } - } - return strings.Trim(b.String(), "-.") -} - -// buildStorageNodeCR constructs a new StorageNode CR for the given worker, -// NUMA socket identifier, per-socket node index, and global ordinal. -func buildStorageNodeCR( - sns *simplyblockv1alpha1.StorageNodeSet, - name, worker, socketID string, - nodeIdx, ordinal int, -) *simplyblockv1alpha1.StorageNode { - globalOrdinal := int32(ordinal) - ni := int32(nodeIdx) - - sn := &simplyblockv1alpha1.StorageNode{ - ObjectMeta: metav1.ObjectMeta{ - Name: name, - Namespace: sns.Namespace, - Labels: map[string]string{ - "storage.simplyblock.io/storagenodeset": sns.Name, - "storage.simplyblock.io/worker": sanitiseDNSLabel(worker), - }, - }, - Spec: simplyblockv1alpha1.StorageNodeSpec{ - StorageNodeSetRef: sns.Name, - WorkerNode: worker, - SocketID: socketID, - NodeIndex: &ni, - SocketIndex: &globalOrdinal, - }, - } - - if overrides, ok := sns.Spec.NodeConfigs[worker]; ok { - sn.Spec.Overrides = &overrides - } - - return sn -} - -// emitOnStorageNodeForWorker emits an event on the StorageNode CR for the given -// worker, mirroring events that are emitted on the StorageNodeSet. -func (r *StorageNodeSetReconciler) emitOnStorageNodeForWorker( - ctx context.Context, - sns *simplyblockv1alpha1.StorageNodeSet, - workerNode string, - eventType, reason, message string, -) { - var snList simplyblockv1alpha1.StorageNodeList - if err := r.List(ctx, &snList, - client.InNamespace(sns.Namespace), - client.MatchingFields{"spec.workerNode": workerNode}, - ); err != nil { - return - } - for i := range snList.Items { - if snList.Items[i].Spec.StorageNodeSetRef == sns.Name { - r.Recorder.Eventf(&snList.Items[i], nil, eventType, reason, reason, "%s", message) - return - } - } -} - -// syncManualStorageNodeStatus merges manually created StorageNode CRs (those -// without a controller OwnerReference pointing to this StorageNodeSet, i.e. -// not in spec.workerNodes) into StorageNodeSet.status.nodes[] so their status -// is visible in the fleet view alongside operator-managed nodes. -func (r *StorageNodeSetReconciler) syncManualStorageNodeStatus( - ctx context.Context, - sns *simplyblockv1alpha1.StorageNodeSet, -) error { - var snList simplyblockv1alpha1.StorageNodeList - if err := r.List(ctx, &snList, - client.InNamespace(sns.Namespace), - client.MatchingFields{"spec.storageNodeSetRef": sns.Name}, - ); err != nil { - return err - } - - // Build a set of workers already covered by spec.workerNodes. - managed := make(map[string]struct{}, len(sns.Spec.WorkerNodes)) - for _, w := range sns.Spec.WorkerNodes { - managed[w] = struct{}{} - } - - changed := false - patch := client.MergeFrom(sns.DeepCopy()) - - for _, sn := range snList.Items { - if _, ok := managed[sn.Spec.WorkerNode]; ok { - continue // already tracked by reconcileWorkerNodes - } - if sn.Status.UUID == "" { - continue // not yet provisioned - } - - // Check if this node is already in status.nodes[]. - found := false - for i := range sns.Status.Nodes { - if sns.Status.Nodes[i].UUID == sn.Status.UUID { - // Update in-place if fields changed. - n := &sns.Status.Nodes[i] - if n.Status != sn.Status.Status || n.Health != sn.Status.Health { - n.Status = sn.Status.Status - n.Health = sn.Status.Health - n.Hostname = sn.Status.Hostname - if p := sn.Status.Ports; p != nil { - n.MgmtIp = p.Management - n.RpcPort = p.Rpc - n.LvolPort = p.Lvol - n.NvmfPort = p.NvmeOf - } - if r := sn.Status.Resources; r != nil { - n.CPU = r.CPU - n.Volumes = r.Volumes - } - changed = true - } - found = true - break - } - } - if !found { - entry := simplyblockv1alpha1.NodeStatus{ - Hostname: sn.Spec.WorkerNode, - UUID: sn.Status.UUID, - Status: sn.Status.Status, - Health: sn.Status.Health, - } - if p := sn.Status.Ports; p != nil { - entry.MgmtIp = p.Management - entry.RpcPort = p.Rpc - entry.LvolPort = p.Lvol - entry.NvmfPort = p.NvmeOf - } - if r := sn.Status.Resources; r != nil { - entry.CPU = r.CPU - entry.Volumes = r.Volumes - } - sns.Status.Nodes = append(sns.Status.Nodes, entry) - changed = true - } - } - - if !changed { - return nil - } - return r.Status().Patch(ctx, sns, patch) -} - -// storageNodePostedAt returns the PostedAt timestamp from the StorageNode CR -// for the given worker, or nil if not found / not yet set. -func (r *StorageNodeSetReconciler) storageNodePostedAt( - ctx context.Context, - namespace, workerNode string, -) *metav1.Time { - var snList simplyblockv1alpha1.StorageNodeList - if err := r.List(ctx, &snList, - client.InNamespace(namespace), - client.MatchingFields{"spec.workerNode": workerNode}, - ); err != nil { - return nil - } - for _, sn := range snList.Items { - if sn.Status.PostedAt != nil { - return sn.Status.PostedAt - } - } - return nil -} - -// allStorageNodesOnline returns true if the number of StorageNode CRs for the -// given worker that have a non-empty UUID equals expectedPerHost. Used to gate -// pollNodeOnline so it is only called once every node has been posted AND -// received its UUID from the backend, avoiding premature timeout. -func (r *StorageNodeSetReconciler) allStorageNodesOnline( - ctx context.Context, - namespace, workerNode string, - expectedPerHost int, -) bool { - var snList simplyblockv1alpha1.StorageNodeList - if err := r.List(ctx, &snList, - client.InNamespace(namespace), - client.MatchingFields{"spec.workerNode": workerNode}, - ); err != nil { - return false - } - online := 0 - for _, sn := range snList.Items { - if sn.Status.UUID != "" { - online++ - } - } - return online >= expectedPerHost -} - -// storageNodeAlreadyPosted returns true if the StorageNode CR for the given -// worker node already has PostedAt set, meaning StorageNodeReconciler has taken -// over provisioning and this reconciler must not duplicate the POST. -func (r *StorageNodeSetReconciler) storageNodeAlreadyPosted( - ctx context.Context, - namespace, workerNode string, -) bool { - var snList simplyblockv1alpha1.StorageNodeList - if err := r.List(ctx, &snList, - client.InNamespace(namespace), - client.MatchingFields{"spec.workerNode": workerNode}, - ); err != nil { - return false - } - for _, sn := range snList.Items { - if sn.Status.PostedAt != nil { - return true - } - } - return false -} - -// effectiveSockets returns the list of socket identifiers to use. When -// SocketsToUse is empty, a single socket "0" is assumed. -func effectiveSockets(sns *simplyblockv1alpha1.StorageNodeSet) []string { - if len(sns.Spec.SocketsToUse) == 0 { - return []string{"0"} - } - return sns.Spec.SocketsToUse -} diff --git a/operator/internal/controller/simplyblockstoragenodeset_storagenode_unit_test.go b/operator/internal/controller/simplyblockstoragenodeset_storagenode_unit_test.go deleted file mode 100644 index 2cce5bd9d..000000000 --- a/operator/internal/controller/simplyblockstoragenodeset_storagenode_unit_test.go +++ /dev/null @@ -1,492 +0,0 @@ -package controller - -import ( - "context" - "strings" - "testing" - - corev1 "k8s.io/api/core/v1" - metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" - "k8s.io/apimachinery/pkg/types" - "k8s.io/client-go/tools/events" - "sigs.k8s.io/controller-runtime/pkg/client" - "sigs.k8s.io/controller-runtime/pkg/client/fake" - - "github.com/simplyblock/atlas/ptr" - simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" - simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" - "github.com/simplyblock/simplyblock-operator/internal/utils" -) - -const ( - snsTestNS = "test" - snsTestCluster = "cluster-a" -) - -func newSNSReconciler(t *testing.T, objects ...client.Object) *StorageNodeSetReconciler { - t.Helper() - // corev1: the reconciler reads/writes the per-node ConfigMap and lists Nodes. - scheme := newTestScheme(t, simplyblockv1alpha1.AddToScheme, corev1.AddToScheme) - cl := fake.NewClientBuilder(). - WithScheme(scheme). - WithStatusSubresource( - &simplyblockv1alpha1.StorageNode{}, - &simplyblockv1alpha1.StorageNodeSet{}, - ). - WithObjects(objects...). - WithIndex(&simplyblockv1alpha1.StorageNode{}, "spec.storageNodeSetRef", func(obj client.Object) []string { - sn := obj.(*simplyblockv1alpha1.StorageNode) - return []string{sn.Spec.StorageNodeSetRef} - }). - Build() - return &StorageNodeSetReconciler{ - Client: cl, - Scheme: scheme, - Recorder: events.NewFakeRecorder(16), - } -} - -// ── TestStorageNodeCRName ────────────────────────────────────────────────────── - -func TestStorageNodeCRName_SimpleCase(t *testing.T) { - name := storageNodeCRName("my-sns") - if name == "" { - t.Fatal("expected non-empty name") - } - if len(name) > 63 { - t.Errorf("name exceeds 63 chars: %q (%d)", name, len(name)) - } - if name != strings.ToLower(name) { - t.Errorf("name is not lowercase: %q", name) - } - if !strings.HasPrefix(name, "my-sns-") { - t.Errorf("name %q does not carry the sns prefix", name) - } -} - -func TestStorageNodeCRName_TruncatesLongNames(t *testing.T) { - longSNS := "simplyblock-node-" + strings.Repeat("a", 80) - name := storageNodeCRName(longSNS) - if len(name) > 63 { - t.Errorf("name exceeds 63 chars: len=%d", len(name)) - } -} - -func TestStorageNodeCRName_IsRandomPerCall(t *testing.T) { - // The id suffix is random, so repeated calls must (with overwhelming - // probability) produce distinct names — this is what the create-retry loop - // relies on to resolve collisions. - seen := map[string]struct{}{} - for i := 0; i < 50; i++ { - seen[storageNodeCRName("sns")] = struct{}{} - } - if len(seen) < 50 { - t.Errorf("expected 50 distinct random names, got %d", len(seen)) - } -} - -func TestStorageNodeCRName_IsDNSLabelSafe(t *testing.T) { - name := storageNodeCRName("my-sns") - for _, c := range name { - if (c < 'a' || c > 'z') && (c < '0' || c > '9') && c != '-' && c != '.' { - t.Errorf("invalid character %q in name %q", c, name) - } - } -} - -// ── TestSanitiseDNSLabel ─────────────────────────────────────────────────────── - -func TestSanitiseDNSLabel_ReplacesInvalidChars(t *testing.T) { - got := sanitiseDNSLabel("vm_01.EXAMPLE.com") - if strings.ContainsAny(got, "_ABCDEFGHIJKLMNOPQRSTUVWXYZ") { - t.Errorf("unsanitised result: %q", got) - } -} - -func TestSanitiseDNSLabel_StripsLeadingTrailingHyphens(t *testing.T) { - got := sanitiseDNSLabel("-bad-label-") - if strings.HasPrefix(got, "-") || strings.HasSuffix(got, "-") { - t.Errorf("result has leading/trailing hyphen: %q", got) - } -} - -// ── TestBuildPerNodeEnvFile ─────────────────────────────────────────────────── - -// newSizingStorageCluster returns a StorageCluster carrying the cluster-scoped -// node sizing values that buildPerNodeEnvFile reads. -func newSizingStorageCluster(maxSubsys, vcpuCount *int32, maxHugePages string) *simplyblockv1alpha2.StorageCluster { - return &simplyblockv1alpha2.StorageCluster{ - ObjectMeta: metav1.ObjectMeta{Name: snsTestCluster, Namespace: snsTestNS}, - Spec: simplyblockv1alpha2.StorageClusterSpec{ - MaxSubsystemCount: maxSubsys, - VCPUCount: vcpuCount, - MinHugePagesSize: maxHugePages, - }, - } -} - -func TestBuildPerNodeEnvFile_UsesClusterSizingValues(t *testing.T) { - maxSubsys := int32(20) - vcpuCount := int32(8) - cluster := newSizingStorageCluster(&maxSubsys, &vcpuCount, "100G") - sns := &simplyblockv1alpha1.StorageNodeSet{ - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - ClusterName: snsTestCluster, - SpdkSystemMemory: "4G", - }, - } - env := buildPerNodeEnvFile(cluster, sns, "worker-a.example.com") - for _, want := range []string{"MAX_SUBSYS_COUNT=20", "VCPU_COUNT=8", "MAX_HUGE_PAGES_SIZE='100G'"} { - if !strings.Contains(env, want) { - t.Errorf("missing %q in env:\n%s", want, env) - } - } -} - -// Cluster sizing values are not overridable per node: nodeConfigs may narrow -// device selection but never the huge-page or core layout. -func TestBuildPerNodeEnvFile_ClusterSizingIdenticalAcrossWorkers(t *testing.T) { - maxSubsys := int32(20) - vcpuCount := int32(8) - cluster := newSizingStorageCluster(&maxSubsys, &vcpuCount, "") - sns := &simplyblockv1alpha1.StorageNodeSet{ - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - ClusterName: snsTestCluster, - NodeConfigs: map[string]simplyblockv1alpha1.StorageNodeOverrides{ - "worker-b": {DriveSizeRange: "50G-1T"}, - }, - }, - } - plain := buildPerNodeEnvFile(cluster, sns, "worker-a") - overridden := buildPerNodeEnvFile(cluster, sns, "worker-b") - - for _, want := range []string{"MAX_SUBSYS_COUNT=20", "VCPU_COUNT=8"} { - if !strings.Contains(plain, want) || !strings.Contains(overridden, want) { - t.Errorf("expected %q in both entries:\n%s\n---\n%s", want, plain, overridden) - } - } - if !strings.Contains(overridden, "SIZE_RANGE='50G-1T'") { - t.Errorf("expected per-node driveSizeRange override to apply:\n%s", overridden) - } -} - -// maxSubsystemCount and vcpuCount are required by the CRD schema; a cluster -// missing them (admitted before that requirement) must not yield a ConfigMap the -// node cannot boot from, since an empty MAX_SUBSYS_COUNT only fails later inside -// node_configure.py. -func TestReconcilePerNodeConfigMap_RejectsClusterMissingRequiredSizing(t *testing.T) { - vcpuCount := int32(8) - cases := map[string]*simplyblockv1alpha2.StorageCluster{ - "both unset": newSizingStorageCluster(nil, nil, ""), - "maxSubsystemCount": newSizingStorageCluster(nil, &vcpuCount, ""), - "vcpuCount": newSizingStorageCluster(ptr.To(int32(20)), nil, ""), - } - for name, cluster := range cases { - t.Run(name, func(t *testing.T) { - sns := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sns", Namespace: snsTestNS}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - ClusterName: snsTestCluster, - WorkerNodes: []string{"worker-a"}, - }, - } - r := newSNSReconciler(t, cluster, sns) - err := r.reconcilePerNodeConfigMap(context.Background(), sns) - if err == nil { - t.Fatal("expected an error for a cluster missing required node sizing") - } - if !strings.Contains(err.Error(), "maxSubsystemCount") { - t.Errorf("error should name the fields to set, got: %v", err) - } - }) - } -} - -func TestReconcilePerNodeConfigMap_WritesClusterSizingForEveryWorker(t *testing.T) { - cluster := newSizingStorageCluster(ptr.To(int32(20)), ptr.To(int32(8)), "") - sns := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sns", Namespace: snsTestNS}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - ClusterName: snsTestCluster, - WorkerNodes: []string{"worker-a", "worker-b"}, - }, - } - r := newSNSReconciler(t, cluster, sns) - if err := r.reconcilePerNodeConfigMap(context.Background(), sns); err != nil { - t.Fatalf("reconcilePerNodeConfigMap: %v", err) - } - - var cm corev1.ConfigMap - if err := r.Get(context.Background(), types.NamespacedName{ - Name: PerNodeConfigMapName(sns.Name), - Namespace: snsTestNS, - }, &cm); err != nil { - t.Fatalf("get ConfigMap: %v", err) - } - for _, worker := range []string{"worker-a", "worker-b"} { - entry, ok := cm.Data[worker] - if !ok { - t.Fatalf("no entry for %s", worker) - } - for _, want := range []string{"MAX_SUBSYS_COUNT=20", "VCPU_COUNT=8"} { - if !strings.Contains(entry, want) { - t.Errorf("%s: missing %q in:\n%s", worker, want, entry) - } - } - } -} - -func TestBuildPerNodeEnvFile_ContainsAllRequiredKeys(t *testing.T) { - cluster := newSizingStorageCluster(ptr.To(int32(20)), ptr.To(int32(8)), "") - sns := &simplyblockv1alpha1.StorageNodeSet{ - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ClusterName: snsTestCluster}, - } - env := buildPerNodeEnvFile(cluster, sns, "any-worker") - required := []string{"MAX_SUBSYS_COUNT=", "MAX_HUGE_PAGES_SIZE=", "VCPU_COUNT=", - "PCI_ALLOWED=", "PCI_BLOCKED=", "NVME_DEVICES=", - "DEVICE_MODEL=", "SIZE_RANGE=", "JM_PERCENT=", "HA_JM_COUNT="} - for _, key := range required { - if !strings.Contains(env, key) { - t.Errorf("missing key %q in env:\n%s", key, env) - } - } -} - -// ── TestCountInFlightNodes ──────────────────────────────────────────────────── - -func TestCountInFlightNodes_ZeroWhenNonePosted(t *testing.T) { - sn1 := newStorageNode("sn-1", snsTestNS, "sns", "worker-1.example.com") - sn2 := newStorageNode("sn-2", snsTestNS, "sns", "worker-2.example.com") - r := newSNReconciler(t, sn1, sn2) - - count, err := r.countInFlightNodes(context.Background(), snsTestNS, "sns", "worker-1.example.com") - if err != nil { - t.Fatalf("unexpected error: %v", err) - } - if count != 0 { - t.Errorf("expected 0 in-flight, got %d", count) - } -} - -func TestCountInFlightNodes_CountsSiblingsWithPostedAtAndNoUUID(t *testing.T) { - now := metav1.Now() - sn1 := newStorageNode("sn-1", snsTestNS, "sns", "worker-1.example.com") - sn2 := newStorageNode("sn-2", snsTestNS, "sns", "worker-2.example.com") - sn2.Status.PostedAt = &now // sn-2 is in-flight - sn3 := newStorageNode("sn-3", snsTestNS, "sns", "worker-3.example.com") - sn3.Status.PostedAt = &now - sn3.Status.UUID = "already-online-uuid" // sn-3 is done - r := newSNReconciler(t, sn1, sn2, sn3) - - count, err := r.countInFlightNodes(context.Background(), snsTestNS, "sns", "worker-1.example.com") - if err != nil { - t.Fatalf("unexpected error: %v", err) - } - if count != 1 { - t.Errorf("expected 1 in-flight (sn-2), got %d", count) - } -} - -func TestCountInFlightNodes_ExcludesSelf(t *testing.T) { - now := metav1.Now() - sn1 := newStorageNode("sn-1", snsTestNS, "sns", "worker-1.example.com") - sn1.Status.PostedAt = &now // self is in-flight - r := newSNReconciler(t, sn1) - - count, err := r.countInFlightNodes(context.Background(), snsTestNS, "sns", "worker-1.example.com") - if err != nil { - t.Fatalf("unexpected error: %v", err) - } - if count != 0 { - t.Errorf("self should not be counted, got %d", count) - } -} - -// ── TestSyncUUIDFromNodeSet ─────────────────────────────────────────────────── - -func TestSyncUUIDFromNodeSet_CopiesUUIDWhenFound(t *testing.T) { - sn := newStorageNode("sn-1", snsTestNS, "sns", snTestWorker) - sns := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sns", Namespace: snsTestNS}, - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - Nodes: []simplyblockv1alpha1.NodeStatus{ - {Hostname: snTestWorker, UUID: "backend-uuid-123", Status: utils.NodeStatusOnline, Health: true}, - }, - }, - } - r := newSNReconciler(t, sn, sns) - - if err := r.syncUUIDFromNodeSet(context.Background(), sn, sns); err != nil { - t.Fatalf("unexpected error: %v", err) - } - - var updated simplyblockv1alpha1.StorageNode - _ = r.Get(context.Background(), types.NamespacedName{Name: "sn-1", Namespace: snsTestNS}, &updated) - if updated.Status.UUID != "backend-uuid-123" { - t.Errorf("UUID not synced: got %q", updated.Status.UUID) - } - if updated.Status.Status != utils.NodeStatusOnline { - t.Errorf("status not synced: got %q", updated.Status.Status) - } -} - -func TestSyncUUIDFromNodeSet_NoopWhenWorkerNotInNodes(t *testing.T) { - sn := newStorageNode("sn-1", snsTestNS, "sns", snTestWorker) - sns := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sns", Namespace: snsTestNS}, - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - Nodes: []simplyblockv1alpha1.NodeStatus{ - {Hostname: "other-worker.example.com", UUID: "other-uuid"}, - }, - }, - } - r := newSNReconciler(t, sn, sns) - - if err := r.syncUUIDFromNodeSet(context.Background(), sn, sns); err != nil { - t.Fatalf("unexpected error: %v", err) - } - - var updated simplyblockv1alpha1.StorageNode - _ = r.Get(context.Background(), types.NamespacedName{Name: "sn-1", Namespace: snsTestNS}, &updated) - if updated.Status.UUID != "" { - t.Errorf("UUID should remain empty, got %q", updated.Status.UUID) - } -} - -func TestSyncUUIDFromNodeSet_SkipsEmptyUUID(t *testing.T) { - sn := newStorageNode("sn-1", snsTestNS, "sns", snTestWorker) - sns := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sns", Namespace: snsTestNS}, - Status: simplyblockv1alpha1.StorageNodeSetStatus{ - Nodes: []simplyblockv1alpha1.NodeStatus{ - {Hostname: snTestWorker, UUID: ""}, // placeholder, not yet online - }, - }, - } - r := newSNReconciler(t, sn, sns) - - if err := r.syncUUIDFromNodeSet(context.Background(), sn, sns); err != nil { - t.Fatalf("unexpected error: %v", err) - } - - var updated simplyblockv1alpha1.StorageNode - _ = r.Get(context.Background(), types.NamespacedName{Name: "sn-1", Namespace: snsTestNS}, &updated) - if updated.Status.UUID != "" { - t.Errorf("UUID should remain empty when node entry has empty UUID, got %q", updated.Status.UUID) - } -} - -// ── TestSyncManualStorageNodeStatus ─────────────────────────────────────────── - -func TestSyncManualStorageNodeStatus_AddsManualNodeToSNSStatus(t *testing.T) { - // A StorageNode without OwnerReference (manual) that has a UUID - sn := newStorageNode("manual-sn", snsTestNS, "sns", "manual-worker.example.com") - sn.Status.UUID = "manual-uuid-456" - sn.Status.Status = utils.NodeStatusOnline - sn.Status.Health = true - - sns := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sns", Namespace: snsTestNS}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ClusterName: snsTestCluster}, - // WorkerNodes does NOT contain manual-worker - } - r := newSNSReconciler(t, sn, sns) - - if err := r.syncManualStorageNodeStatus(context.Background(), sns); err != nil { - t.Fatalf("unexpected error: %v", err) - } - - var updated simplyblockv1alpha1.StorageNodeSet - _ = r.Get(context.Background(), types.NamespacedName{Name: "sns", Namespace: snsTestNS}, &updated) - - found := false - for _, n := range updated.Status.Nodes { - if n.UUID == "manual-uuid-456" { - found = true - if n.Status != utils.NodeStatusOnline { - t.Errorf("status not synced: got %q", n.Status) - } - } - } - if !found { - t.Error("manual StorageNode UUID not added to StorageNodeSet.status.nodes[]") - } -} - -func TestSyncManualStorageNodeStatus_SkipsUnprovisionedNodes(t *testing.T) { - sn := newStorageNode("manual-sn", snsTestNS, "sns", "manual-worker.example.com") - // UUID is empty — not yet provisioned - - sns := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sns", Namespace: snsTestNS}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ClusterName: snsTestCluster}, - } - r := newSNSReconciler(t, sn, sns) - - if err := r.syncManualStorageNodeStatus(context.Background(), sns); err != nil { - t.Fatalf("unexpected error: %v", err) - } - - var updated simplyblockv1alpha1.StorageNodeSet - _ = r.Get(context.Background(), types.NamespacedName{Name: "sns", Namespace: snsTestNS}, &updated) - if len(updated.Status.Nodes) != 0 { - t.Errorf("expected empty status.nodes[], got %d entries", len(updated.Status.Nodes)) - } -} - -func TestSyncManualStorageNodeStatus_SkipsWorkerInSpecWorkerNodes(t *testing.T) { - // Worker is in spec.workerNodes — it's operator-managed, not manual - sn := newStorageNode("managed-sn", snsTestNS, "sns", snTestWorker) - sn.Status.UUID = "managed-uuid" - - sns := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sns", Namespace: snsTestNS}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - ClusterName: snsTestCluster, - WorkerNodes: []string{snTestWorker}, // operator-managed - }, - } - r := newSNSReconciler(t, sn, sns) - - if err := r.syncManualStorageNodeStatus(context.Background(), sns); err != nil { - t.Fatalf("unexpected error: %v", err) - } - - var updated simplyblockv1alpha1.StorageNodeSet - _ = r.Get(context.Background(), types.NamespacedName{Name: "sns", Namespace: snsTestNS}, &updated) - if len(updated.Status.Nodes) != 0 { - t.Errorf("operator-managed node should not be added by syncManualStorageNodeStatus") - } -} - -func TestSyncManualStorageNodeStatus_IdempotentOnSecondCall(t *testing.T) { - sn := newStorageNode("manual-sn", snsTestNS, "sns", "manual-worker.example.com") - sn.Status.UUID = "manual-uuid-789" - sn.Status.Status = utils.NodeStatusOnline - - sns := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sns", Namespace: snsTestNS}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ClusterName: snsTestCluster}, - } - r := newSNSReconciler(t, sn, sns) - - // First call - _ = r.syncManualStorageNodeStatus(context.Background(), sns) - - // Re-fetch and call again - var sns2 simplyblockv1alpha1.StorageNodeSet - _ = r.Get(context.Background(), types.NamespacedName{Name: "sns", Namespace: snsTestNS}, &sns2) - _ = r.syncManualStorageNodeStatus(context.Background(), &sns2) - - var updated simplyblockv1alpha1.StorageNodeSet - _ = r.Get(context.Background(), types.NamespacedName{Name: "sns", Namespace: snsTestNS}, &updated) - count := 0 - for _, n := range updated.Status.Nodes { - if n.UUID == "manual-uuid-789" { - count++ - } - } - if count != 1 { - t.Errorf("expected exactly 1 entry for manual node, got %d (not idempotent)", count) - } -} diff --git a/operator/internal/controller/storagenode_controller.go b/operator/internal/controller/storagenode_controller.go deleted file mode 100644 index 37453b992..000000000 --- a/operator/internal/controller/storagenode_controller.go +++ /dev/null @@ -1,1206 +0,0 @@ -/* -Copyright 2025. - -Licensed under the Apache License, Version 2.0 (the "License"); -you may not use this file except in compliance with the License. -You may obtain a copy of the License at - - http://www.apache.org/licenses/LICENSE-2.0 - -Unless required by applicable law or agreed to in writing, software -distributed under the License is distributed on an "AS IS" BASIS, -WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. -See the License for the specific language governing permissions and -limitations under the License. -*/ - -package controller - -import ( - "context" - "encoding/json" - "fmt" - "net/http" - "slices" - "time" - - corev1 "k8s.io/api/core/v1" - apierrors "k8s.io/apimachinery/pkg/api/errors" - metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" - "k8s.io/apimachinery/pkg/runtime" - "k8s.io/apimachinery/pkg/types" - "k8s.io/client-go/tools/events" - ctrl "sigs.k8s.io/controller-runtime" - "sigs.k8s.io/controller-runtime/pkg/client" - "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil" - "sigs.k8s.io/controller-runtime/pkg/event" - "sigs.k8s.io/controller-runtime/pkg/handler" - logf "sigs.k8s.io/controller-runtime/pkg/log" - "sigs.k8s.io/controller-runtime/pkg/reconcile" - "sigs.k8s.io/controller-runtime/pkg/source" - - "github.com/simplyblock/atlas/prometheus" - "github.com/simplyblock/atlas/ptr" - - simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" - simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" - "github.com/simplyblock/simplyblock-operator/internal/cpinformer" - "github.com/simplyblock/simplyblock-operator/internal/cpinformer/subscriptions" - "github.com/simplyblock/simplyblock-operator/internal/utils" - "github.com/simplyblock/simplyblock-operator/internal/webapi" -) - -const ( - storageNodeFinalizer = "storage.simplyblock.io/storagenode-finalizer" - // storageNodeSyncInterval is how soon a node is looked at again after a - // failed read. It stays short because it is a retry rather than a poll. - storageNodeSyncInterval = 30 * time.Second - // storageNodeBackstopInterval is how soon a healthy node is re-read from the - // control plane when nothing has pushed. The stream carries a status change - // within a second, so this is the correctness floor rather than the - // mechanism: it covers the operator's own stream goroutine wedging, and the - // control plane's own note that a change written by a pre-upgrade component - // can take 30 seconds to reach a stream at all. - storageNodeBackstopInterval = 3 * time.Minute -) - -// StorageNodeReconciler reconciles StorageNode objects. -// It owns the per-node provisioning loop: node-add POST, online polling, status -// sync, and triggering a StorageNodeOps(action=remove) on deletion. -type StorageNodeReconciler struct { - client.Client - Scheme *runtime.Scheme - Recorder events.EventRecorder - TLSEnabled bool - TLSMutualEnabled bool - - // DeviceScopes, if set, receives this node's (cluster, node) scope so the - // control-plane SSE manager streams the node's devices. The control plane - // offers no cluster-wide device stream, so the node rather than the cluster - // is what drives a device subscription. Optional (nil in tests). - DeviceScopes *cpinformer.ScopeSet - // NodeRegistries learn which StorageNode object a backend node id belongs - // to. Both the device and the storage-node subscriptions need that mapping - // to name what they stream, and this reconciler is where the backend id and - // the object are known at once. Optional (empty in tests). - NodeRegistries []NodeObjectRegistry - // Capacity, if set, supplies the node's storage occupancy. It is separate - // from the control-plane API and from the stream because neither carries - // the number: a node's capacity exists only in the metrics the control - // plane exports. Optional (nil leaves status.resources.capacity absent). - Capacity NodeCapacitySource - // Nodes, if set, is the storage-node subscription's cache. Status is read - // from it in preference to the control-plane API: the stream has already - // delivered the same DTO, so a request would ask for what is in memory. - // Optional (nil in tests, and nil leaves the reconciler polling). - Nodes NodeCache -} - -// NodeObjectRegistry is how the StorageNode reconciler tells a subscription -// which object a backend node id names. It is an interface so the reconciler -// does not depend on any subscription's concrete type. -type NodeObjectRegistry interface { - RegisterNode(nodeID string, node types.NamespacedName) - UnregisterNode(nodeID string) -} - -// NodeCapacitySource supplies how full a node's storage is. It is satisfied by -// atlas-lib's prometheus.Provider, and it is an interface here so that a test -// needs no Prometheus. -type NodeCapacitySource interface { - // NodeCapacity returns the sample for every node of a cluster, keyed by - // backend node UUID. A node with no sample is absent. - NodeCapacity(ctx context.Context, clusterUUID string) (map[string]prometheus.Capacity, error) -} - -// NodeCache is the read surface the reconciler needs from the storage-node -// subscription: the node the control plane last reported, and whether the -// cluster's snapshot has arrived at all. -type NodeCache interface { - // Lookup returns the cached node with the given backend id, or ok=false - // when the control plane no longer reports it. - Lookup(nodeID string) (cpinformer.Scope, subscriptions.NodeDTO, bool) - // List returns every node of a cluster, which is what a StorageNode with - // no backend id yet has to search to find its own. - List(scope cpinformer.Scope) []subscriptions.NodeDTO - // Synced reports whether the cluster's initial snapshot has been applied. - Synced(scope cpinformer.Scope) bool - // Triggers is the reconcile-trigger stream; each event names a StorageNode. - Triggers() <-chan event.GenericEvent -} - -// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodes,verbs=get;list;watch;create;update;patch;delete -// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodes/status,verbs=get;update;patch -// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodes/finalizers,verbs=update - -func (r *StorageNodeReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) { - log := logf.FromContext(ctx) - - var sn simplyblockv1alpha1.StorageNode - if err := r.Get(ctx, req.NamespacedName, &sn); err != nil { - if apierrors.IsNotFound(err) { - return ctrl.Result{}, nil - } - return ctrl.Result{}, err - } - - // Fetch the parent StorageNodeSet for fleet config. - var sns simplyblockv1alpha1.StorageNodeSet - if err := r.Get(ctx, types.NamespacedName{ - Name: sn.Spec.StorageNodeSetRef, - Namespace: sn.Namespace, - }, &sns); err != nil { - if apierrors.IsNotFound(err) { - log.Info("parent StorageNodeSet not found, requeuing", "ref", sn.Spec.StorageNodeSetRef) - return ctrl.Result{RequeueAfter: 10 * time.Second}, nil - } - return ctrl.Result{}, err - } - - // Resolve cluster UUID early — needed for both provisioning and status sync. - clusterUUID, err := utils.ResolveClusterUUID(ctx, r.Client, sn.Namespace, sns.Spec.ClusterName) - if err != nil { - log.Info("cluster UUID not ready yet, requeuing", "cluster", sns.Spec.ClusterName) - return ctrl.Result{RequeueAfter: 10 * time.Second}, nil - } - - // Handle deletion. - if !sn.DeletionTimestamp.IsZero() { - return r.handleDeletion(ctx, &sn, clusterUUID) - } - - // Ensure finalizer. - if !controllerutil.ContainsFinalizer(&sn, storageNodeFinalizer) { - controllerutil.AddFinalizer(&sn, storageNodeFinalizer) - if err := r.Update(ctx, &sn); err != nil { - return ctrl.Result{}, err - } - return ctrl.Result{Requeue: true}, nil - } - - // Sync overrides from the parent StorageNodeSet. - if err := r.syncOverrides(ctx, &sn, &sns); err != nil { - return ctrl.Result{}, err - } - - apiClient := webapi.NewClient() - - if sn.Status.UUID == "" { - // Fast path: old StorageNodeSetReconciler recorded UUID in status.nodes[]. - if err := r.syncUUIDFromNodeSet(ctx, &sn, &sns); err != nil { - return ctrl.Result{}, err - } - - if sn.Status.UUID == "" { - alreadyPosted := sn.Status.PostedAt != nil - // An adopted node already has a backend record that - // provisionNode's POST can't match, so check by IP first. - // Only during an upgrade (upgrade secret present) — normal - // nodes skip this extra check. - adopting := !alreadyPosted && r.isUpgradeAdoption(ctx, sn.Namespace, sns.Spec.ClusterName) - if alreadyPosted || adopting { - if err := r.pollUUIDFromBackend(ctx, &sn, clusterUUID, apiClient); err != nil { - return ctrl.Result{}, err - } - } - } - - if sn.Status.UUID == "" && sn.Status.PostedAt == nil { - return r.provisionNode(ctx, &sn, &sns, clusterUUID, apiClient) - } - if sn.Status.UUID == "" { - // POST already sent (by us or a sibling socket); keep polling. - return ctrl.Result{RequeueAfter: 10 * time.Second}, nil - } - } - - // Node provisioned → its devices can be streamed. The name is registered - // before the scope, so the subscription can name what the first snapshot - // delivers. - r.registerStreams(clusterUUID, &sn) - - // Node provisioned → sync status periodically. - return r.syncStatus(ctx, &sn, clusterUUID, apiClient) -} - -// registerStreams makes the node's control-plane events nameable and opens its -// device stream. The name is registered before the scope, so a subscription can -// name whatever its first snapshot delivers. -// -// The storage-node stream needs no scope registered here: it is per cluster, so -// the cluster's own controller opens it, and this node's events arrive on it -// whether or not this reconciler has run. -func (r *StorageNodeReconciler) registerStreams(clusterUUID string, sn *simplyblockv1alpha1.StorageNode) { - if sn.Status.UUID == "" { - return - } - for _, registry := range r.NodeRegistries { - registry.RegisterNode(sn.Status.UUID, client.ObjectKeyFromObject(sn)) - } - if r.DeviceScopes != nil { - r.DeviceScopes.Add(cpinformer.Scope{clusterUUID, sn.Status.UUID}) - } -} - -// unregisterStreams closes the node's device stream and stops naming events -// after it. The scope goes first: no further device events can arrive once the -// stream is closed, so the name mappings are dropped second and nothing is left -// naming objects after a node on its way out. -func (r *StorageNodeReconciler) unregisterStreams(clusterUUID string, sn *simplyblockv1alpha1.StorageNode) { - if sn.Status.UUID == "" { - return - } - if r.DeviceScopes != nil { - r.DeviceScopes.Remove(cpinformer.Scope{clusterUUID, sn.Status.UUID}) - } - for _, registry := range r.NodeRegistries { - registry.UnregisterNode(sn.Status.UUID) - } -} - -// pollUUIDFromBackend lists all backend nodes for the cluster, finds the ones -// matching the worker's internal IP, and assigns the UUID to this StorageNode -// based on its socketIndex. For multi-socket workers (multiple nodes per IP), -// backend nodes are sorted by RPC port (ascending) and matched by position -// to the socketIndex — socket 0 → lowest RPC port, socket 1 → next, etc. -// Called every 10s while PostedAt is set but UUID is still empty; stops as -// soon as the UUID is assigned. -func (r *StorageNodeReconciler) pollUUIDFromBackend( - ctx context.Context, - sn *simplyblockv1alpha1.StorageNode, - clusterUUID string, - apiClient *webapi.Client, -) error { - log := logf.FromContext(ctx) - - ip, err := getNodeInternalIP(ctx, r.Client, sn.Spec.WorkerNode) - if err != nil { - log.V(1).Info("pollUUIDFromBackend: could not get worker IP, retrying", - "worker", sn.Spec.WorkerNode, "error", err.Error()) - return nil - } - - allNodes, ok := r.clusterNodes(ctx, clusterUUID, apiClient) - if !ok { - return nil // transient — requeue silently - } - - // Collect all backend nodes for this worker's IP. - // Multi-socket: one backend node per socket, each with a different RPC port. - var matching []SNODEAPIResponse - for _, n := range allNodes { - if n.IP == ip && n.UUID != "" { - matching = append(matching, n) - } - } - if len(matching) == 0 { - return nil // node not yet visible on backend — requeue - } - - // Sort by RPC port ascending: socket 0 → lowest port, socket 1 → next, etc. - slices.SortFunc(matching, func(a, b SNODEAPIResponse) int { - return a.RPC_PORT - b.RPC_PORT - }) - - socketIdx := 0 - if sn.Spec.SocketIndex != nil { - socketIdx = int(*sn.Spec.SocketIndex) - } - if socketIdx >= len(matching) { - log.V(1).Info("pollUUIDFromBackend: socket not yet online", - "worker", sn.Spec.WorkerNode, "socketIndex", socketIdx, "found", len(matching)) - return nil - } - n := matching[socketIdx] - - cpu := int32(n.CPU) - volumes := int32(n.Volumes) - rpcPort := int32(n.RPC_PORT) - lvolPort := int32(n.LVOL_PORT) - nvmfPort := int32(n.NVMF_PORT) - - patch := client.MergeFrom(sn.DeepCopy()) - sn.Status.UUID = n.UUID - sn.Status.Status = n.Status - sn.Status.Health = n.Health - sn.Status.Hostname = n.Hostname - sn.Status.FailureDomain = fdPtr(n.FailureDomain) - sn.Status.Resources = &simplyblockv1alpha1.StorageNodeResources{ - CPU: &cpu, - Volumes: &volumes, - } - sn.Status.Ports = &simplyblockv1alpha1.StorageNodePorts{ - Management: n.IP, - NvmeOf: &nvmfPort, - Lvol: &lvolPort, - Rpc: &rpcPort, - } - if err := r.Status().Patch(ctx, sn, patch); err != nil && !apierrors.IsNotFound(err) { - return fmt.Errorf("pollUUIDFromBackend: %w", err) - } - log.Info("pollUUIDFromBackend: UUID assigned", - "worker", sn.Spec.WorkerNode, "socketIndex", socketIdx, - "uuid", n.UUID, "status", n.Status) - return nil -} - -// isUpgradeAdoption reports whether the "simplyblock--upgrade" -// secret exists — the same signal controllers/cluster/storagecluster_controller.go -// uses to adopt the StorageCluster instead of creating a new one. -func (r *StorageNodeReconciler) isUpgradeAdoption(ctx context.Context, namespace, clusterName string) bool { - secretName := fmt.Sprintf("simplyblock-%s-upgrade", clusterName) - var secret corev1.Secret - err := r.Get(ctx, types.NamespacedName{Name: secretName, Namespace: namespace}, &secret) - return err == nil -} - -// syncUUIDFromNodeSet copies the backend UUID from StorageNodeSet.status.nodes[] -// into StorageNode.status.uuid. This is the Phase 1 bridge: the old -// StorageNodeSetReconciler owns provisioning and tracks UUIDs in its own status; -// the StorageNodeReconciler reads that status so it doesn't re-POST. -func (r *StorageNodeReconciler) syncUUIDFromNodeSet( - ctx context.Context, - sn *simplyblockv1alpha1.StorageNode, - sns *simplyblockv1alpha1.StorageNodeSet, -) error { - for _, ns := range sns.Status.Nodes { - if ns.Hostname != sn.Spec.WorkerNode || ns.UUID == "" { - continue - } - patch := client.MergeFrom(sn.DeepCopy()) - sn.Status.UUID = ns.UUID - sn.Status.Status = ns.Status - sn.Status.Health = ns.Health - sn.Status.FailureDomain = ns.FailureDomain - if err := r.Status().Patch(ctx, sn, patch); err != nil && !apierrors.IsNotFound(err) { - return fmt.Errorf("syncing UUID for StorageNode %s: %w", sn.Name, err) - } - return nil - } - return nil -} - -// syncOverrides propagates StorageNodeSet.spec.nodeConfigs[worker] into -// StorageNode.spec.overrides. The StorageNodeSet is the single source of truth. -func (r *StorageNodeReconciler) syncOverrides( - ctx context.Context, - sn *simplyblockv1alpha1.StorageNode, - sns *simplyblockv1alpha1.StorageNodeSet, -) error { - overrides, ok := sns.Spec.NodeConfigs[sn.Spec.WorkerNode] - if !ok { - return nil - } - patch := client.MergeFrom(sn.DeepCopy()) - sn.Spec.Overrides = &overrides - if err := r.Patch(ctx, sn, patch); err != nil && !apierrors.IsNotFound(err) { - return fmt.Errorf("syncing overrides for %s: %w", sn.Name, err) - } - return nil -} - -// provisionNode posts the node to the backend API. StorageNodeReconciler is the -// sole owner of provisioning — the old StorageNodeSetReconciler skips nodes -// whose StorageNode CR already has PostedAt set. -func (r *StorageNodeReconciler) provisionNode( - ctx context.Context, - sn *simplyblockv1alpha1.StorageNode, - sns *simplyblockv1alpha1.StorageNodeSet, - clusterUUID string, - apiClient *webapi.Client, -) (ctrl.Result, error) { - log := logf.FromContext(ctx) - - // POST already sent — poll until the UUID appears via syncUUIDFromNodeSet. - if sn.Status.PostedAt != nil { - return ctrl.Result{RequeueAfter: 10 * time.Second}, nil - } - - // One POST per worker: the backend add_node handles all sockets internally. - // If any sibling CR for the same worker already has PostedAt, skip the POST - // and mark ourselves so pollUUIDFromBackend can pick up our UUID. - if r.workerAlreadyPosted(ctx, sn) { - log.Info("worker already posted by sibling socket — skipping POST, polling for UUID", - "worker", sn.Spec.WorkerNode) - now := metav1.Now() - patch := client.MergeFrom(sn.DeepCopy()) - sn.Status.PostedAt = &now - if err := r.Status().Patch(ctx, sn, patch); err != nil { - log.Error(err, "failed to patch PostedAt for non-primary socket") - } - return ctrl.Result{RequeueAfter: 10 * time.Second}, nil - } - - // Respect MaxParallelNodeAdds: count sibling StorageNode CRs in this set - // that are in-flight (PostedAt set, UUID not yet assigned) and block if the - // limit is reached. Defaults to 1 (sequential) when not set, which is safe - // for FDB clusters where simultaneous node reboots reduce fault tolerance. - maxParallel := 1 - if sns.Spec.MaxParallelNodeAdds != nil { - maxParallel = int(*sns.Spec.MaxParallelNodeAdds) - } - inFlight, err := r.countInFlightNodes(ctx, sn.Namespace, sn.Spec.StorageNodeSetRef, sn.Spec.WorkerNode) - if err == nil && inFlight >= maxParallel { - log.Info("parallel node add limit reached, requeuing", - "inFlight", inFlight, "max", maxParallel) - return ctrl.Result{RequeueAfter: waitForNodeOnlineWaitInterval}, nil - } - - // FDB workers must be added sequentially to avoid simultaneous reboots that - // would reduce FDB fault tolerance. If this worker hosts an FDB pod and any - // other FDB worker in the same StorageNodeSet is currently in-flight, block. - if r.isWorkerFDB(ctx, sn.Namespace, sn.Spec.WorkerNode) { - if blocked, err := r.isFDBWorkerBlocked(ctx, sn); err == nil && blocked { - log.Info("FDB worker: another FDB node is in-flight, requeuing sequentially", - "worker", sn.Spec.WorkerNode) - return ctrl.Result{RequeueAfter: waitForNodeOnlineWaitInterval}, nil - } - } - - // Guard: failure domain must be set if the feature is enabled. - if err := r.checkFailureDomain(ctx, sn, sns); err != nil { - r.Recorder.Eventf(sn, nil, "Warning", "FailureDomainMissing", "FailureDomainMissing", "%s", err.Error()) - log.Info("blocking node-add: "+err.Error(), "node", sn.Name) - return ctrl.Result{RequeueAfter: 60 * time.Second}, nil - } - - // Wait until the node's SPDK API endpoint is reachable. - if err := checkNodeInfoReachable(ctx, sn.Spec.WorkerNode, sn.Namespace, r.TLSEnabled, r.TLSMutualEnabled); err != nil { - log.V(1).Info("storage node API not reachable yet, requeuing", - "worker", sn.Spec.WorkerNode, "error", err.Error()) - return ctrl.Result{RequeueAfter: 10 * time.Second}, nil - } - - // Merge fleet defaults with per-node overrides — overrides always win. - eff := effectiveNodeConfig(sn, sns) - - nodeAddress := utils.StorageNodeSetAPIAddress(sn.Spec.WorkerNode, sn.Namespace) - params := utils.StorageNodeSetAddParams{ - NodeAddress: nodeAddress, - InterfaceName: sns.Spec.MgmtIfname, - SPDKImage: eff.SpdkImage, - SPDKProxyImage: eff.SpdkProxyImage, - DataNics: sns.Spec.DataIfname, - Namespace: sn.Namespace, - JMPercent: journalManagerPercentPerDeviceFromSpec(eff.JournalManagerSpec), - Partitions: partitionsPerDevice(sns), - HaJMCount: journalManagerCountFromSpec(eff.JournalManagerSpec), - CRName: sns.Name, - CRNameSpace: sns.Namespace, - CRPlural: "storagenodesets", - Format4K: ptr.BoolFromOrFalse(sns.Spec.ForceFormat4K), - SpdkSystemMemory: eff.SpdkSystemMemory, - FailureDomain: effectiveFailureDomainPtr(sn, sns), - Expand: ptr.BoolFromOrFalse(eff.Expand), - } - - // Re-read the in-flight count immediately before the POST to narrow the - // check-then-act race window. The first check (above) filters the common - // case; this final re-check reduces the window to the round-trip of a - // single List call, making concurrent overshoot extremely unlikely. - if recheck, recheckErr := r.countInFlightNodes(ctx, sn.Namespace, sn.Spec.StorageNodeSetRef, sn.Spec.WorkerNode); recheckErr == nil && recheck >= maxParallel { - log.Info("parallel node add limit reached on re-check, requeuing", - "inFlight", recheck, "max", maxParallel) - return ctrl.Result{RequeueAfter: waitForNodeOnlineWaitInterval}, nil - } - - endpoint := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes", clusterUUID) - body, status, err := apiClient.Do(ctx, http.MethodPost, endpoint, params) - if err != nil || status >= 300 { - if err == nil { - err = fmt.Errorf("unexpected status %d", status) - } - log.Error(err, "storage node add failed", "status", status, "response", string(body)) - return ctrl.Result{RequeueAfter: 20 * time.Second}, nil - } - - log.Info("storage node add POST sent", "endpoint", endpoint, "status", status) - - now := metav1.Now() - patch := client.MergeFrom(sn.DeepCopy()) - sn.Status.PostedAt = &now - if err := r.Status().Patch(ctx, sn, patch); err != nil { - log.Error(err, "failed to patch PostedAt") - } - return ctrl.Result{RequeueAfter: 10 * time.Second}, nil -} - -// isWorkerFDB returns true if the given worker node currently hosts at least -// one FDB pod. -func (r *StorageNodeReconciler) isWorkerFDB(ctx context.Context, namespace, workerNode string) bool { - var podList corev1.PodList - if err := r.List(ctx, &podList, - client.InNamespace(namespace), - client.HasLabels{utils.LabelFDBClusterName}, - client.MatchingFields{"spec.nodeName": workerNode}, - ); err != nil { - return false - } - return len(podList.Items) > 0 -} - -// isFDBWorkerBlocked returns true if any sibling StorageNode in the same -// StorageNodeSet is an FDB worker currently in-flight (PostedAt set, UUID -// empty). Used to enforce sequential adds for FDB nodes. -func (r *StorageNodeReconciler) isFDBWorkerBlocked( - ctx context.Context, - sn *simplyblockv1alpha1.StorageNode, -) (bool, error) { - var snList simplyblockv1alpha1.StorageNodeList - if err := r.List(ctx, &snList, - client.InNamespace(sn.Namespace), - client.MatchingFields{"spec.storageNodeSetRef": sn.Spec.StorageNodeSetRef}, - ); err != nil { - return false, err - } - for _, sibling := range snList.Items { - if sibling.Name == sn.Name { - continue - } - if sibling.Status.PostedAt == nil || sibling.Status.UUID != "" { - continue - } - // Sibling is in-flight — check if it's also an FDB worker. - if r.isWorkerFDB(ctx, sn.Namespace, sibling.Spec.WorkerNode) { - return true, nil - } - } - return false, nil -} - -// countInFlightNodes returns how many distinct workers (physical hosts) in the -// same StorageNodeSet are still being provisioned, excluding the calling node's -// own worker. A worker is considered in-flight if any of its StorageNode CRs -// has PostedAt set and either has no UUID yet or is still "in_creation." -// Counting distinct workers (not individual CRs) ensures maxParallelNodeAdds -// matches its documented meaning regardless of nodesPerSocket — without this, -// each in-flight host would consume nodesPerSocket slots instead of one. -func (r *StorageNodeReconciler) countInFlightNodes( - ctx context.Context, - namespace, snsRef, excludeWorker string, -) (int, error) { - var snList simplyblockv1alpha1.StorageNodeList - if err := r.List(ctx, &snList, - client.InNamespace(namespace), - client.MatchingFields{"spec.storageNodeSetRef": snsRef}, - ); err != nil { - return 0, err - } - inFlightWorkers := make(map[string]struct{}) - for _, sn := range snList.Items { - if sn.Spec.WorkerNode == excludeWorker { - continue - } - if sn.Status.PostedAt != nil && - sn.Status.Status != utils.NodeStatusTimeout && - (sn.Status.UUID == "" || sn.Status.Status == utils.NodeStatusInCreation) { - inFlightWorkers[sn.Spec.WorkerNode] = struct{}{} - } - } - return len(inFlightWorkers), nil -} - -// workerAlreadyPosted returns true if a sibling StorageNode for the same -// worker already has PostedAt or UUID set (UUID too: an adopted sibling -// skips PostedAt entirely). Enforces one POST per worker — the backend -// add_node handles all sockets from a single call. -func (r *StorageNodeReconciler) workerAlreadyPosted(ctx context.Context, sn *simplyblockv1alpha1.StorageNode) bool { - var snList simplyblockv1alpha1.StorageNodeList - if err := r.List(ctx, &snList, - client.InNamespace(sn.Namespace), - client.MatchingFields{"spec.storageNodeSetRef": sn.Spec.StorageNodeSetRef}, - ); err != nil { - return false - } - for _, sibling := range snList.Items { - if sibling.Name == sn.Name || sibling.Spec.WorkerNode != sn.Spec.WorkerNode { - continue - } - if sibling.Status.PostedAt != nil || sibling.Status.UUID != "" { - return true - } - } - return false -} - -// journalManagerPercentPerDeviceFromSpec returns JM percent from the effective -// JournalManagerSpec, defaulting to 3 when nil. -func journalManagerPercentPerDeviceFromSpec(spec *simplyblockv1alpha1.JournalManagerSpec) int { - if spec == nil { - return 3 - } - return ptr.IntFrom(spec.PercentPerDevice, 3) -} - -// journalManagerCountFromSpec returns JM count from the effective -// JournalManagerSpec, defaulting to 3 when nil. -func journalManagerCountFromSpec(spec *simplyblockv1alpha1.JournalManagerSpec) int { - if spec == nil { - return 3 - } - return ptr.IntFrom(spec.Count, 3) -} - -// syncStatus fetches the current node status from the backend and updates StorageNode.status. -func (r *StorageNodeReconciler) syncStatus( - ctx context.Context, - sn *simplyblockv1alpha1.StorageNode, - clusterUUID string, - apiClient *webapi.Client, -) (ctrl.Result, error) { - log := logf.FromContext(ctx) - - // The stream has already delivered this node's DTO, so asking the control - // plane for it would fetch what is in memory. The cache is only trusted for - // a node it actually holds: an absent one may be gone, or may be a scope - // whose snapshot has not arrived, and the request below tells those apart - // by returning either the node or a 404. - if dto, ok := r.cachedNode(sn.Status.UUID); ok { - return r.applyNodeStatus(ctx, sn, clusterUUID, nodeResponseFrom(dto)) - } - - endpoint := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s", clusterUUID, sn.Status.UUID) - body, status, err := apiClient.Do(ctx, http.MethodGet, endpoint, nil) - if err != nil || status >= 300 { - if status == http.StatusNotFound { - // The stored UUID no longer exists on the backend — the cluster may - // have been reset and nodes re-created with new UUIDs. - // Clear only UUID (PostedAt is intentionally preserved): provisionNode - // checks PostedAt and returns early without POSTing, while - // syncUUIDFromNodeSet and pollUUIDFromBackend find the new UUID from - // StorageNodeSet.status.nodes[] or by querying the backend by IP. - log.Info("backend node not found (404) — clearing stale UUID for re-adoption", - "staleUUID", sn.Status.UUID) - patch := client.MergeFrom(sn.DeepCopy()) - sn.Status.UUID = "" - sn.Status.Status = "" - sn.Status.Health = false - if patchErr := r.Status().Patch(ctx, sn, patch); patchErr != nil { - log.Error(patchErr, "failed to clear stale UUID from StorageNode") - } - return ctrl.Result{Requeue: true}, nil - } - if err == nil { - err = fmt.Errorf("status %d", status) - } - log.Error(err, "failed to GET node status", "uuid", sn.Status.UUID) - return ctrl.Result{RequeueAfter: storageNodeSyncInterval}, nil - } - - var resp SNODEAPIResponse - if err := json.Unmarshal(body, &resp); err != nil { - log.Error(err, "failed to unmarshal node status response") - return ctrl.Result{RequeueAfter: storageNodeSyncInterval}, nil - } - - return r.applyNodeStatus(ctx, sn, clusterUUID, resp) -} - -// capacityWriteThreshold is how much a node's used size has to move before the -// new reading is worth recording: one percent of the node's own total, so a -// larger node tolerates a larger absolute drift. -// -// Some threshold is required rather than merely economical. The reconciler -// watches its own objects, so every status write schedules another reconcile; -// writing a freshly sampled number every time would make the node reconcile -// itself in a loop for as long as any I/O was happening, bounded only by the -// workqueue's rate limiter. -const capacityWriteThresholdPercent = 1 - -// existingCapacity returns what the object already records, so an unchanged -// reading can be carried forward rather than rewritten. -func existingCapacity(sn *simplyblockv1alpha1.StorageNode) *simplyblockv1alpha1.StorageNodeCapacity { - if sn.Status.Resources == nil { - return nil - } - return sn.Status.Resources.Capacity -} - -// worthWriting reports whether a sample says something the object does not -// already say. A first reading always does; after that the used size has to -// have moved by at least capacityWriteThresholdPercent of the total, or the -// total itself has to have changed, which happens when a device joins or -// leaves the node. -func worthWriting( - current *simplyblockv1alpha1.StorageNodeCapacity, - sample prometheus.Capacity, -) bool { - if !sample.Sampled() { - return false // nothing has measured this node; say nothing about it - } - if current == nil || current.UsedBytes == nil || current.TotalBytes == nil { - return true - } - if *current.TotalBytes != sample.Total { - return true - } - drift := *current.UsedBytes - sample.Used - if drift < 0 { - drift = -drift - } - return drift*100 >= sample.Total*capacityWriteThresholdPercent -} - -// nodeCapacity reads one node's occupancy, or reports that there is none to -// read. A failure is not an error the caller has to handle: the rest of the -// status is correct without it, and a node whose capacity is momentarily -// unknown is better published than not published at all. -func (r *StorageNodeReconciler) nodeCapacity( - ctx context.Context, - clusterUUID, nodeUUID string, -) (prometheus.Capacity, bool) { - if r.Capacity == nil || clusterUUID == "" || nodeUUID == "" { - return prometheus.Capacity{}, false - } - samples, err := r.Capacity.NodeCapacity(ctx, clusterUUID) - if err != nil { - logf.FromContext(ctx).V(1).Info("no capacity sample for this node", - "cluster", clusterUUID, "node", nodeUUID, "err", err.Error()) - return prometheus.Capacity{}, false - } - sample, ok := samples[nodeUUID] - return sample, ok -} - -// clusterNodes returns every backend node of the cluster, from the stream's -// cache once it has delivered the cluster's snapshot and from the control plane -// until then. -// -// The cache is preferred because the stream has already delivered exactly this -// list, so requesting it again asks for what is in memory. It is only trusted -// once the scope is synced: an empty unsynced cache and a cluster with no nodes -// look identical, and adoption would read the first as the second and keep -// waiting for a node that is already there. -func (r *StorageNodeReconciler) clusterNodes( - ctx context.Context, - clusterUUID string, - apiClient *webapi.Client, -) ([]SNODEAPIResponse, bool) { - if r.Nodes != nil && r.Nodes.Synced(cpinformer.Scope{clusterUUID}) { - cached := r.Nodes.List(cpinformer.Scope{clusterUUID}) - out := make([]SNODEAPIResponse, 0, len(cached)) - for _, dto := range cached { - out = append(out, nodeResponseFrom(dto)) - } - return out, true - } - - endpoint := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/", clusterUUID) - body, httpStatus, err := apiClient.Do(ctx, http.MethodGet, endpoint, nil) - if err != nil || httpStatus >= 300 { - return nil, false - } - var allNodes []SNODEAPIResponse - if err := json.Unmarshal(body, &allNodes); err != nil { - return nil, false - } - return allNodes, true -} - -// cachedNode returns the node the storage-node stream last reported, when there -// is a subscription and it holds one. -func (r *StorageNodeReconciler) cachedNode(nodeID string) (subscriptions.NodeDTO, bool) { - if r.Nodes == nil || nodeID == "" { - return subscriptions.NodeDTO{}, false - } - _, dto, ok := r.Nodes.Lookup(nodeID) - return dto, ok -} - -// nodeResponseFrom converts a streamed node into the shape the status -// projection takes. The two are the same wire schema read through different -// transports, so converting is what keeps one projection serving both and makes -// a pushed status and a polled one indistinguishable in the object. -func nodeResponseFrom(dto subscriptions.NodeDTO) SNODEAPIResponse { - return SNODEAPIResponse{ - UUID: dto.ID, - Status: dto.Status, - IP: dto.ManagementIP, - Health: dto.HealthCheck, - Hostname: dto.Hostname, - CPU: int(dto.CPUCount), - Volumes: int(dto.Volumes), - RPC_PORT: int(dto.RPCPort), - LVOL_PORT: int(dto.LvolPort), - NVMF_PORT: int(dto.NVMeOFPort), - FailureDomain: dto.FailureDomain, - } -} - -// applyNodeStatus writes what the control plane reports into the CR's status, -// whichever transport it arrived by. -// -// The requeue is the backstop rather than the mechanism: a status change reaches -// the stream in about a second, and this interval only bounds how long a wedged -// stream can go unnoticed. -func (r *StorageNodeReconciler) applyNodeStatus( - ctx context.Context, - sn *simplyblockv1alpha1.StorageNode, - clusterUUID string, - resp SNODEAPIResponse, -) (ctrl.Result, error) { - log := logf.FromContext(ctx) - - cpu := int32(resp.CPU) - volumes := int32(resp.Volumes) - rpcPort := int32(resp.RPC_PORT) - lvolPort := int32(resp.LVOL_PORT) - nvmfPort := int32(resp.NVMF_PORT) - - patch := client.MergeFrom(sn.DeepCopy()) - // Carry the previous capacity forward by default. It is replaced below only - // when the reading has moved materially, so an unchanged one produces an - // empty patch and therefore no write and no retrigger. - capacity := existingCapacity(sn) - if sample, ok := r.nodeCapacity(ctx, clusterUUID, sn.Status.UUID); ok { - if worthWriting(capacity, sample) { - capacity = &simplyblockv1alpha1.StorageNodeCapacity{ - TotalBytes: ptr.To(sample.Total), - UsedBytes: ptr.To(sample.Used), - SampledAt: ptr.To(metav1.NewTime(sample.SampledAt)), - } - } - } - - sn.Status.Status = resp.Status - sn.Status.Health = resp.Health - sn.Status.Hostname = resp.Hostname - sn.Status.FailureDomain = fdPtr(resp.FailureDomain) - sn.Status.Resources = &simplyblockv1alpha1.StorageNodeResources{ - CPU: &cpu, - Volumes: &volumes, - Capacity: capacity, - } - sn.Status.Ports = &simplyblockv1alpha1.StorageNodePorts{ - Management: resp.IP, - NvmeOf: &nvmfPort, - Lvol: &lvolPort, - Rpc: &rpcPort, - } - - if err := r.Status().Patch(ctx, sn, patch); err != nil { - log.Error(err, "failed to patch StorageNode status") - } - return ctrl.Result{RequeueAfter: storageNodeBackstopInterval}, nil -} - -// checkFailureDomain returns an error if the parent cluster has -// enableFailureDomains=true but this node has no failureDomain set. -func (r *StorageNodeReconciler) checkFailureDomain( - ctx context.Context, - sn *simplyblockv1alpha1.StorageNode, - sns *simplyblockv1alpha1.StorageNodeSet, -) error { - var cluster simplyblockv1alpha2.StorageCluster - if err := r.Get(ctx, types.NamespacedName{ - Name: sns.Spec.ClusterName, - Namespace: sn.Namespace, - }, &cluster); err != nil { - return nil // can't determine; don't block - } - if cluster.Spec.EnableFailureDomains == nil || !*cluster.Spec.EnableFailureDomains { - return nil - } - if effectiveFailureDomainSet(sn, sns) { - return nil - } - return fmt.Errorf( - "failureDomain not set for worker %q; add nodeConfigs[%s].failureDomain to StorageNodeSet %q", - sn.Spec.WorkerNode, sn.Spec.WorkerNode, sns.Name, - ) -} - -// partitionsPerDevice translates spec.enableJournalDevice into the backend's -// partitions-per-device count: 0 dedicates a whole NVMe device to the journal -// manager, 1 carves a journal partition out of each storage device. Unset -// defaults to 1, preserving the behavior of the spec.partitions field this -// replaced. -func partitionsPerDevice(sns *simplyblockv1alpha1.StorageNodeSet) int { - if ptr.BoolFromOrFalse(sns.Spec.EnableJournalDevice) { - return 0 - } - return 1 -} - -// effectiveNodeConfig returns the merged config for a node: fleet defaults -// overridden by any per-node values from StorageNode.spec.overrides. -func effectiveNodeConfig(sn *simplyblockv1alpha1.StorageNode, sns *simplyblockv1alpha1.StorageNodeSet) simplyblockv1alpha1.StorageNodeOverrides { - eff := simplyblockv1alpha1.StorageNodeOverrides{ - SpdkImage: sns.Spec.SpdkImage, - SpdkProxyImage: sns.Spec.SpdkProxyImage, - SpdkSystemMemory: sns.Spec.SpdkSystemMemory, - JournalManagerSpec: sns.Spec.JournalManagerSpec, - PcieAllowList: sns.Spec.PcieAllowList, - PcieDenyList: sns.Spec.PcieDenyList, - PcieModel: sns.Spec.PcieModel, - DriveSizeRange: sns.Spec.DriveSizeRange, - DeviceNames: sns.Spec.DeviceNames, - EnableCpuTopology: sns.Spec.EnableCpuTopology, - ReservedSystemCPU: sns.Spec.ReservedSystemCPU, - UbuntuHost: sns.Spec.UbuntuHost, - Expand: sns.Spec.Expand, - } - if sn.Spec.Overrides == nil { - return eff - } - o := sn.Spec.Overrides - if o.SpdkImage != "" { - eff.SpdkImage = o.SpdkImage - } - if o.SpdkProxyImage != "" { - eff.SpdkProxyImage = o.SpdkProxyImage - } - if o.SpdkSystemMemory != "" { - eff.SpdkSystemMemory = o.SpdkSystemMemory - } - if o.JournalManagerSpec != nil { - eff.JournalManagerSpec = o.JournalManagerSpec - } - if len(o.PcieAllowList) > 0 { - eff.PcieAllowList = o.PcieAllowList - } - if len(o.PcieDenyList) > 0 { - eff.PcieDenyList = o.PcieDenyList - } - if o.PcieModel != "" { - eff.PcieModel = o.PcieModel - } - if o.DriveSizeRange != "" { - eff.DriveSizeRange = o.DriveSizeRange - } - if len(o.DeviceNames) > 0 { - eff.DeviceNames = o.DeviceNames - } - if o.EnableCpuTopology != nil { - eff.EnableCpuTopology = o.EnableCpuTopology - } - if o.ReservedSystemCPU != "" { - eff.ReservedSystemCPU = o.ReservedSystemCPU - } - if o.UbuntuHost != nil { - eff.UbuntuHost = o.UbuntuHost - } - if o.FailureDomain != nil { - eff.FailureDomain = o.FailureDomain - } - if o.Expand != nil { - eff.Expand = o.Expand - } - return eff -} - -// effectiveFailureDomainSet reports whether a failure domain has been explicitly -// assigned to the node via spec.overrides.failureDomain or spec.nodeFailureDomains. -func effectiveFailureDomainSet(sn *simplyblockv1alpha1.StorageNode, sns *simplyblockv1alpha1.StorageNodeSet) bool { - if sn.Spec.Overrides != nil && sn.Spec.Overrides.FailureDomain != nil { - return true - } - _, ok := sns.Spec.NodeFailureDomains[sn.Spec.WorkerNode] - return ok -} - -// effectiveFailureDomain returns the failure domain for the node: -// StorageNode.spec.overrides.failureDomain takes precedence over -// StorageNodeSet.spec.nodeFailureDomains[worker]. Only meaningful when -// effectiveFailureDomainSet reports true -- the zero return here also covers -// "unset", so callers that must distinguish the two (e.g. anything crossing -// a JSON boundary, where 0 and absent are different wire values) should use -// effectiveFailureDomainPtr instead. -func effectiveFailureDomain(sn *simplyblockv1alpha1.StorageNode, sns *simplyblockv1alpha1.StorageNodeSet) int { - if sn.Spec.Overrides != nil && sn.Spec.Overrides.FailureDomain != nil { - return int(*sn.Spec.Overrides.FailureDomain) - } - if v, ok := sns.Spec.NodeFailureDomains[sn.Spec.WorkerNode]; ok { - return int(v) - } - return 0 -} - -// effectiveFailureDomainPtr returns the same value as effectiveFailureDomain, -// but as *int so "domain 0" and "not configured" stay distinguishable across -// a JSON boundary (nil is omitted by `omitempty`; Ptr(0) serializes as 0). -// Use this instead of effectiveFailureDomain wherever the result crosses -// such a boundary, e.g. StorageNodeSetAddParams.FailureDomain. -func effectiveFailureDomainPtr(sn *simplyblockv1alpha1.StorageNode, sns *simplyblockv1alpha1.StorageNodeSet) *int { - if !effectiveFailureDomainSet(sn, sns) { - return nil - } - v := effectiveFailureDomain(sn, sns) - return &v -} - -// handleDeletion ensures a StorageNodeOps(action=remove) exists for this node -// if it is online, then removes the finalizer once the ops CR completes. -func (r *StorageNodeReconciler) handleDeletion( - ctx context.Context, - sn *simplyblockv1alpha1.StorageNode, - clusterUUID string, -) (ctrl.Result, error) { - log := logf.FromContext(ctx) - - // The node is going away, so stop streaming its devices. Its StorageDevice - // objects are garbage-collected by the owner reference rather than deleted - // here. - r.unregisterStreams(clusterUUID, sn) - - // If the node was never provisioned, skip ops and remove finalizer immediately. - if sn.Status.UUID == "" { - controllerutil.RemoveFinalizer(sn, storageNodeFinalizer) - return ctrl.Result{}, r.Update(ctx, sn) - } - - if sn.Status.Status == utils.NodeStatusSuspended || - sn.Status.Status == utils.ClusterStatusActive || - sn.Status.Status == utils.NodeStatusOnline { - if err := r.ensureRemoveOps(ctx, sn); err != nil { - return ctrl.Result{}, err - } - } - - // If an ops is still active, requeue and wait. - if sn.Status.ActiveOpsRef != "" { - log.Info("waiting for StorageNodeOps to complete before finalizer removal", - "ops", sn.Status.ActiveOpsRef) - return ctrl.Result{RequeueAfter: 15 * time.Second}, nil - } - - // A Failed remove ops clears ActiveOpsRef (releaseLock, in - // storagenodeops_controller.go) exactly like a Succeeded one -- but - // Failed means the backend node was never actually removed (blocked by - // a precondition, or a rejected DELETE that resumed the node instead of - // removing it). Removing the finalizer here would let Kubernetes delete - // this StorageNode CR anyway, orphaning a live, healthy backend node - // the operator no longer tracks at all. Block deletion and surface it - // instead: a human must either fix whatever blocked the removal and - // delete the failed ops to retry, or restore the worker to - // spec.workerNodes to keep it. - opsName := sn.Name + "-remove" - var removeOps simplyblockv1alpha2.StorageNodeOps - if err := r.Get(ctx, types.NamespacedName{Name: opsName, Namespace: sn.Namespace}, &removeOps); err == nil { - if removeOps.Status.Phase == simplyblockv1alpha2.StorageNodeOpsPhaseFailed { - r.Recorder.Eventf(sn, nil, "Warning", "RemoveOpsFailed", "RemoveOpsFailed", - "node removal failed (%s); the node was NOT removed and this StorageNode "+ - "will not be deleted -- delete StorageNodeOps/%s to retry, or restore "+ - "%s to spec.workerNodes to keep it", - removeOps.Status.Message, opsName, sn.Spec.WorkerNode) - return ctrl.Result{RequeueAfter: 60 * time.Second}, nil - } - } else if !apierrors.IsNotFound(err) { - return ctrl.Result{}, err - } - - controllerutil.RemoveFinalizer(sn, storageNodeFinalizer) - return ctrl.Result{}, r.Update(ctx, sn) -} - -// ensureRemoveOps creates a StorageNodeOps(action=remove) for this StorageNode -// if one does not already exist. -func (r *StorageNodeReconciler) ensureRemoveOps( - ctx context.Context, - sn *simplyblockv1alpha1.StorageNode, -) error { - opsName := sn.Name + "-remove" - var existing simplyblockv1alpha2.StorageNodeOps - err := r.Get(ctx, types.NamespacedName{Name: opsName, Namespace: sn.Namespace}, &existing) - if err == nil { - return nil // already exists - } - if !apierrors.IsNotFound(err) { - return err - } - - ops := simplyblockv1alpha2.StorageNodeOps{} - ops.Name = opsName - ops.Namespace = sn.Namespace - ops.Spec.NodeRef = sn.Name - ops.Spec.Action = simplyblockv1alpha2.StorageNodeOpsActionRemove - if err := controllerutil.SetControllerReference(sn, &ops, r.Scheme); err != nil { - return err - } - return r.Create(ctx, &ops) -} - -// storageNodeSetToStorageNodeRequests maps a StorageNodeSet change to all -// owned StorageNode reconcile requests. -func (r *StorageNodeReconciler) storageNodeSetToStorageNodeRequests( - ctx context.Context, - obj client.Object, -) []reconcile.Request { - var snList simplyblockv1alpha1.StorageNodeList - if err := r.List(ctx, &snList, - client.InNamespace(obj.GetNamespace()), - client.MatchingFields{"spec.storageNodeSetRef": obj.GetName()}, - ); err != nil { - return nil - } - reqs := make([]reconcile.Request, len(snList.Items)) - for i, sn := range snList.Items { - reqs[i] = reconcile.Request{NamespacedName: types.NamespacedName{ - Name: sn.Name, - Namespace: sn.Namespace, - }} - } - return reqs -} - -// SetupWithManager registers the StorageNodeReconciler with the controller manager. -func (r *StorageNodeReconciler) SetupWithManager(mgr ctrl.Manager) error { - if err := mgr.GetFieldIndexer().IndexField( - context.Background(), - &simplyblockv1alpha1.StorageNode{}, - "spec.storageNodeSetRef", - func(obj client.Object) []string { - sn := obj.(*simplyblockv1alpha1.StorageNode) - return []string{sn.Spec.StorageNodeSetRef} - }, - ); err != nil { - return err - } - - if err := mgr.GetFieldIndexer().IndexField( - context.Background(), - &simplyblockv1alpha1.StorageNode{}, - "spec.workerNode", - func(obj client.Object) []string { - sn := obj.(*simplyblockv1alpha1.StorageNode) - return []string{sn.Spec.WorkerNode} - }, - ); err != nil { - return err - } - - // Index Pods by spec.nodeName for efficient FDB worker detection. - if err := mgr.GetFieldIndexer().IndexField( - context.Background(), - &corev1.Pod{}, - "spec.nodeName", - func(obj client.Object) []string { - pod := obj.(*corev1.Pod) - if pod.Spec.NodeName == "" { - return nil - } - return []string{pod.Spec.NodeName} - }, - ); err != nil { - return err - } - - builder := ctrl.NewControllerManagedBy(mgr). - For(&simplyblockv1alpha1.StorageNode{}). - Named("storagenode"). - Watches( - &simplyblockv1alpha1.StorageNodeSet{}, - handler.EnqueueRequestsFromMapFunc(r.storageNodeSetToStorageNodeRequests), - ). - Owns(&simplyblockv1alpha2.StorageNodeOps{}) - - // A pushed control-plane change reconciles the node it named. The events - // already carry the object's own name, so they need no map function. - if r.Nodes != nil { - builder = builder.WatchesRawSource( - source.Channel(r.Nodes.Triggers(), &handler.EnqueueRequestForObject{}), - ) - } - - return builder.Complete(r) -} diff --git a/operator/internal/controller/storagenode_controller_unit_test.go b/operator/internal/controller/storagenode_controller_unit_test.go deleted file mode 100644 index 28c8ecc7f..000000000 --- a/operator/internal/controller/storagenode_controller_unit_test.go +++ /dev/null @@ -1,904 +0,0 @@ -package controller - -import ( - "context" - "errors" - "net/http" - "net/http/httptest" - "sync/atomic" - "testing" - "time" - - corev1 "k8s.io/api/core/v1" - metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" - "k8s.io/apimachinery/pkg/types" - "k8s.io/client-go/tools/events" - "sigs.k8s.io/controller-runtime/pkg/client" - "sigs.k8s.io/controller-runtime/pkg/client/fake" - - "github.com/simplyblock/atlas/prometheus" - "github.com/simplyblock/atlas/ptr" - - simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" - simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" - "github.com/simplyblock/simplyblock-operator/internal/cpinformer" - "github.com/simplyblock/simplyblock-operator/internal/cpinformer/subscriptions" - "github.com/simplyblock/simplyblock-operator/internal/utils" - "github.com/simplyblock/simplyblock-operator/internal/webapi" -) - -// ── helpers ────────────────────────────────────────────────────────────────── - -const ( - snTestNS = "test" - snTestCluster = "cluster-a" - snTestWorker = "worker-1.example.com" -) - -func newSNReconciler(t *testing.T, objects ...client.Object) *StorageNodeReconciler { - t.Helper() - scheme := newTestScheme(t, - simplyblockv1alpha1.AddToScheme, - corev1.AddToScheme, - ) - cl := fake.NewClientBuilder(). - WithScheme(scheme). - WithStatusSubresource( - &simplyblockv1alpha1.StorageNode{}, - &simplyblockv1alpha2.StorageNodeOps{}, - &simplyblockv1alpha2.StorageCluster{}, - &simplyblockv1alpha1.StorageNodeSet{}, - ). - WithObjects(objects...). - WithIndex(&simplyblockv1alpha1.StorageNode{}, "spec.storageNodeSetRef", func(obj client.Object) []string { - sn := obj.(*simplyblockv1alpha1.StorageNode) - return []string{sn.Spec.StorageNodeSetRef} - }). - WithIndex(&simplyblockv1alpha1.StorageNode{}, "spec.workerNode", func(obj client.Object) []string { - sn := obj.(*simplyblockv1alpha1.StorageNode) - return []string{sn.Spec.WorkerNode} - }). - Build() - return &StorageNodeReconciler{ - Client: cl, - Scheme: scheme, - Recorder: events.NewFakeRecorder(16), - } -} - -//nolint:unparam -func newStorageNodeSet(name, ns, cluster string, nodeConfigs map[string]simplyblockv1alpha1.StorageNodeOverrides) *simplyblockv1alpha1.StorageNodeSet { - return &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: ns}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - ClusterName: cluster, - WorkerNodes: []string{snTestWorker}, - NodeConfigs: nodeConfigs, - }, - } -} - -//nolint:unparam -func newStorageNode(name, ns, snsRef, worker string) *simplyblockv1alpha1.StorageNode { - return &simplyblockv1alpha1.StorageNode{ - ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: ns}, - Spec: simplyblockv1alpha1.StorageNodeSpec{ - StorageNodeSetRef: snsRef, - WorkerNode: worker, - }, - } -} - -// ── TestSyncOverrides ───────────────────────────────────────────────────────── - -func TestSyncOverrides_PropagatesNodeConfigs(t *testing.T) { - sns := newStorageNodeSet("sns", snTestNS, snTestCluster, map[string]simplyblockv1alpha1.StorageNodeOverrides{ - snTestWorker: {DriveSizeRange: "50G-1T", SpdkSystemMemory: "8G"}, - }) - sn := newStorageNode("sn-1", snTestNS, "sns", snTestWorker) - r := newSNReconciler(t, sns, sn) - - if err := r.syncOverrides(context.Background(), sn, sns); err != nil { - t.Fatalf("syncOverrides returned error: %v", err) - } - - var updated simplyblockv1alpha1.StorageNode - if err := r.Get(context.Background(), types.NamespacedName{Name: "sn-1", Namespace: snTestNS}, &updated); err != nil { - t.Fatalf("failed to fetch updated StorageNode: %v", err) - } - if updated.Spec.Overrides == nil { - t.Fatal("expected Overrides to be set") - } - if updated.Spec.Overrides.SpdkSystemMemory != "8G" { - t.Errorf("SpdkSystemMemory: got %q want %q", updated.Spec.Overrides.SpdkSystemMemory, "8G") - } - if updated.Spec.Overrides.DriveSizeRange != "50G-1T" { - t.Errorf("DriveSizeRange: got %q want %q", updated.Spec.Overrides.DriveSizeRange, "50G-1T") - } -} - -func TestSyncOverrides_NoopWhenWorkerNotInNodeConfigs(t *testing.T) { - sns := newStorageNodeSet("sns", snTestNS, snTestCluster, nil) - sn := newStorageNode("sn-1", snTestNS, "sns", snTestWorker) - r := newSNReconciler(t, sns, sn) - - if err := r.syncOverrides(context.Background(), sn, sns); err != nil { - t.Fatalf("unexpected error: %v", err) - } - - var updated simplyblockv1alpha1.StorageNode - _ = r.Get(context.Background(), types.NamespacedName{Name: "sn-1", Namespace: snTestNS}, &updated) - if updated.Spec.Overrides != nil { - t.Error("expected Overrides to remain nil when worker not in nodeConfigs") - } -} - -// ── TestEffectiveNodeConfig ─────────────────────────────────────────────────── - -func TestEffectiveNodeConfig_OverridesTakePrecedence(t *testing.T) { - fleetMem := "4G" - overrideMem := "16G" - sns := &simplyblockv1alpha1.StorageNodeSet{ - Spec: simplyblockv1alpha1.StorageNodeSetSpec{SpdkSystemMemory: fleetMem}, - } - sn := &simplyblockv1alpha1.StorageNode{ - Spec: simplyblockv1alpha1.StorageNodeSpec{ - Overrides: &simplyblockv1alpha1.StorageNodeOverrides{SpdkSystemMemory: overrideMem}, - }, - } - eff := effectiveNodeConfig(sn, sns) - if eff.SpdkSystemMemory != overrideMem { - t.Errorf("expected override %q, got %q", overrideMem, eff.SpdkSystemMemory) - } -} - -func TestEffectiveNodeConfig_FallsBackToFleetDefault(t *testing.T) { - fleetMem := "4G" - sns := &simplyblockv1alpha1.StorageNodeSet{ - Spec: simplyblockv1alpha1.StorageNodeSetSpec{SpdkSystemMemory: fleetMem}, - } - sn := &simplyblockv1alpha1.StorageNode{} // no overrides - eff := effectiveNodeConfig(sn, sns) - if eff.SpdkSystemMemory != fleetMem { - t.Errorf("expected fleet default %q, got %q", fleetMem, eff.SpdkSystemMemory) - } -} - -// ── TestEffectiveFailureDomain ──────────────────────────────────────────────── - -func TestEffectiveFailureDomain_OverrideTakesPrecedenceOverMap(t *testing.T) { - fd := int32(3) - sns := &simplyblockv1alpha1.StorageNodeSet{ - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - NodeFailureDomains: map[string]int32{snTestWorker: 1}, - }, - } - sn := &simplyblockv1alpha1.StorageNode{ - Spec: simplyblockv1alpha1.StorageNodeSpec{ - WorkerNode: snTestWorker, - Overrides: &simplyblockv1alpha1.StorageNodeOverrides{FailureDomain: &fd}, - }, - } - if got := effectiveFailureDomain(sn, sns); got != 3 { - t.Errorf("expected 3 from override, got %d", got) - } -} - -func TestEffectiveFailureDomain_FallsBackToMap(t *testing.T) { - sns := &simplyblockv1alpha1.StorageNodeSet{ - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - NodeFailureDomains: map[string]int32{snTestWorker: 2}, - }, - } - sn := &simplyblockv1alpha1.StorageNode{ - Spec: simplyblockv1alpha1.StorageNodeSpec{WorkerNode: snTestWorker}, - } - if got := effectiveFailureDomain(sn, sns); got != 2 { - t.Errorf("expected 2 from map, got %d", got) - } -} - -func TestEffectiveFailureDomain_ZeroWhenNotSet(t *testing.T) { - sns := &simplyblockv1alpha1.StorageNodeSet{} - sn := &simplyblockv1alpha1.StorageNode{ - Spec: simplyblockv1alpha1.StorageNodeSpec{WorkerNode: snTestWorker}, - } - if got := effectiveFailureDomain(sn, sns); got != 0 { - t.Errorf("expected 0, got %d", got) - } -} - -// ── TestEffectiveFailureDomainPtr ───────────────────────────────────────────── -// -// Regression for the 2026-08-27 incident: StorageNodeSetAddParams.FailureDomain -// used to be a plain int with `omitempty`, so a node explicitly assigned -// domain 0 had the field silently dropped from the JSON POSTed to the -// backend -- the backend then saw no failure_domain at all, and that node's -// add_node never completed (it sat with no status forever, blocking every -// FDB worker queued behind it). effectiveFailureDomainPtr exists so the -// caller can tell "domain 0" (a real, valid *int) apart from "not -// configured" (nil, correctly omitted). - -func TestEffectiveFailureDomainPtr_NilWhenNotSet(t *testing.T) { - sns := &simplyblockv1alpha1.StorageNodeSet{} - sn := &simplyblockv1alpha1.StorageNode{ - Spec: simplyblockv1alpha1.StorageNodeSpec{WorkerNode: snTestWorker}, - } - if got := effectiveFailureDomainPtr(sn, sns); got != nil { - t.Errorf("expected nil (not configured), got %v", *got) - } -} - -func TestEffectiveFailureDomainPtr_NonNilForDomainZero(t *testing.T) { - sns := &simplyblockv1alpha1.StorageNodeSet{ - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - NodeFailureDomains: map[string]int32{snTestWorker: 0}, - }, - } - sn := &simplyblockv1alpha1.StorageNode{ - Spec: simplyblockv1alpha1.StorageNodeSpec{WorkerNode: snTestWorker}, - } - got := effectiveFailureDomainPtr(sn, sns) - if got == nil { - t.Fatal("expected a non-nil pointer to 0 (explicitly configured), got nil") - } - if *got != 0 { - t.Errorf("expected *got == 0, got %d", *got) - } -} - -func TestEffectiveFailureDomainPtr_NonNilForNonZeroDomain(t *testing.T) { - sns := &simplyblockv1alpha1.StorageNodeSet{ - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - NodeFailureDomains: map[string]int32{snTestWorker: 2}, - }, - } - sn := &simplyblockv1alpha1.StorageNode{ - Spec: simplyblockv1alpha1.StorageNodeSpec{WorkerNode: snTestWorker}, - } - got := effectiveFailureDomainPtr(sn, sns) - if got == nil || *got != 2 { - t.Errorf("expected pointer to 2, got %v", got) - } -} - -// ── TestCheckFailureDomain ──────────────────────────────────────────────────── - -func TestCheckFailureDomain_BlocksWhenEnabledAndNotSet(t *testing.T) { - enabled := true - cluster := &simplyblockv1alpha2.StorageCluster{ - ObjectMeta: metav1.ObjectMeta{Name: snTestCluster, Namespace: snTestNS}, - Spec: simplyblockv1alpha2.StorageClusterSpec{EnableFailureDomains: &enabled}, - } - sns := newStorageNodeSet("sns", snTestNS, snTestCluster, nil) - sn := newStorageNode("sn-1", snTestNS, "sns", snTestWorker) - r := newSNReconciler(t, cluster, sns, sn) - - err := r.checkFailureDomain(context.Background(), sn, sns) - if err == nil { - t.Fatal("expected error when failureDomain not set and enableFailureDomains=true") - } -} - -func TestCheckFailureDomain_AllowsWhenFailureDomainSet(t *testing.T) { - enabled := true - fd := int32(1) - cluster := &simplyblockv1alpha2.StorageCluster{ - ObjectMeta: metav1.ObjectMeta{Name: snTestCluster, Namespace: snTestNS}, - Spec: simplyblockv1alpha2.StorageClusterSpec{EnableFailureDomains: &enabled}, - } - sns := newStorageNodeSet("sns", snTestNS, snTestCluster, nil) - sn := newStorageNode("sn-1", snTestNS, "sns", snTestWorker) - sn.Spec.Overrides = &simplyblockv1alpha1.StorageNodeOverrides{FailureDomain: &fd} - r := newSNReconciler(t, cluster, sns, sn) - - if err := r.checkFailureDomain(context.Background(), sn, sns); err != nil { - t.Fatalf("unexpected error: %v", err) - } -} - -func TestCheckFailureDomain_SkipsWhenFeatureDisabled(t *testing.T) { - disabled := false - cluster := &simplyblockv1alpha2.StorageCluster{ - ObjectMeta: metav1.ObjectMeta{Name: snTestCluster, Namespace: snTestNS}, - Spec: simplyblockv1alpha2.StorageClusterSpec{EnableFailureDomains: &disabled}, - } - sns := newStorageNodeSet("sns", snTestNS, snTestCluster, nil) - sn := newStorageNode("sn-1", snTestNS, "sns", snTestWorker) // no failureDomain - r := newSNReconciler(t, cluster, sns, sn) - - if err := r.checkFailureDomain(context.Background(), sn, sns); err != nil { - t.Fatalf("expected no error when feature disabled, got: %v", err) - } -} - -// ── TestEnsureRemoveOps ─────────────────────────────────────────────────────── - -func TestEnsureRemoveOps_CreatesOpsWhenMissing(t *testing.T) { - sn := newStorageNode("sn-1", snTestNS, "sns", snTestWorker) - sn.Status.UUID = "uuid-1" - sns := newStorageNodeSet("sns", snTestNS, snTestCluster, nil) - r := newSNReconciler(t, sn, sns) - - if err := r.ensureRemoveOps(context.Background(), sn); err != nil { - t.Fatalf("ensureRemoveOps returned error: %v", err) - } - - var ops simplyblockv1alpha2.StorageNodeOps - if err := r.Get(context.Background(), types.NamespacedName{ - Name: "sn-1-remove", Namespace: snTestNS, - }, &ops); err != nil { - t.Fatalf("expected StorageNodeOps to be created: %v", err) - } - if ops.Spec.Action != simplyblockv1alpha2.StorageNodeOpsActionRemove { - t.Errorf("expected action=Remove, got %q", ops.Spec.Action) - } - if ops.Spec.NodeRef != "sn-1" { - t.Errorf("expected nodeRef=sn-1, got %q", ops.Spec.NodeRef) - } -} - -func TestEnsureRemoveOps_IdempotentWhenAlreadyExists(t *testing.T) { - sn := newStorageNode("sn-1", snTestNS, "sns", snTestWorker) - sn.Status.UUID = "uuid-1" - existingOps := &simplyblockv1alpha2.StorageNodeOps{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-1-remove", Namespace: snTestNS}, - Spec: simplyblockv1alpha2.StorageNodeOpsSpec{NodeRef: "sn-1", Action: simplyblockv1alpha2.StorageNodeOpsActionRemove}, - } - sns := newStorageNodeSet("sns", snTestNS, snTestCluster, nil) - r := newSNReconciler(t, sn, sns, existingOps) - - // Should not return an error on second call. - if err := r.ensureRemoveOps(context.Background(), sn); err != nil { - t.Fatalf("ensureRemoveOps should be idempotent, got: %v", err) - } -} - -// ── TestHandleDeletion ──────────────────────────────────────────────────────── - -func TestHandleDeletion_RemovesFinalizerWhenNeverProvisioned(t *testing.T) { - sn := newStorageNode("sn-1", snTestNS, "sns", snTestWorker) - sn.Finalizers = []string{storageNodeFinalizer} - // status.UUID is empty — node was never provisioned - sns := newStorageNodeSet("sns", snTestNS, snTestCluster, nil) - r := newSNReconciler(t, sn, sns) - - _, err := r.handleDeletion(context.Background(), sn, snTestCluster) - if err != nil { - t.Fatalf("handleDeletion returned error: %v", err) - } - - var updated simplyblockv1alpha1.StorageNode - _ = r.Get(context.Background(), types.NamespacedName{Name: "sn-1", Namespace: snTestNS}, &updated) - for _, f := range updated.Finalizers { - if f == storageNodeFinalizer { - t.Error("finalizer should have been removed for unprovisioned node") - } - } -} - -// TestHandleDeletion_FailedRemoveOpsBlocksFinalizerRemoval guards the gap -// found live 2026-08-13: a Failed remove ops clears ActiveOpsRef exactly -// like a Succeeded one (releaseLock, storagenodeops_controller.go), but -// Failed means the backend node was never actually removed -- blocked by a -// precondition, or resumed instead of removed after a rejected DELETE. -// Letting the finalizer come off here would delete this StorageNode CR -// from Kubernetes while the backend node is still alive and online, with -// nothing left to track it. -func TestHandleDeletion_FailedRemoveOpsBlocksFinalizerRemoval(t *testing.T) { - sn := newStorageNode("sn-1", snTestNS, "sns", snTestWorker) - sn.Finalizers = []string{storageNodeFinalizer} - sn.Status.UUID = "node-uuid-1" - sn.Status.Status = utils.NodeStatusOnline - // ActiveOpsRef already cleared by releaseLock when the ops failed. - sn.Status.ActiveOpsRef = "" - sns := newStorageNodeSet("sns", snTestNS, snTestCluster, nil) - ops := &simplyblockv1alpha2.StorageNodeOps{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-1-remove", Namespace: snTestNS}, - Spec: simplyblockv1alpha2.StorageNodeOpsSpec{NodeRef: "sn-1", Action: simplyblockv1alpha2.StorageNodeOpsActionRemove}, - Status: simplyblockv1alpha2.StorageNodeOpsStatus{ - Phase: simplyblockv1alpha2.StorageNodeOpsPhaseFailed, - Message: "blocked: failure-domain balance: ...", - }, - } - r := newSNReconciler(t, sn, sns, ops) - - result, err := r.handleDeletion(context.Background(), sn, snTestCluster) - if err != nil { - t.Fatalf("handleDeletion returned error: %v", err) - } - if result.RequeueAfter == 0 { - t.Error("expected a requeue, not a one-shot pass-through") - } - - var updated simplyblockv1alpha1.StorageNode - _ = r.Get(context.Background(), types.NamespacedName{Name: "sn-1", Namespace: snTestNS}, &updated) - found := false - for _, f := range updated.Finalizers { - if f == storageNodeFinalizer { - found = true - } - } - if !found { - t.Error("finalizer must NOT be removed while the remove ops is Failed -- " + - "the backend node was never actually removed") - } -} - -// TestHandleDeletion_SucceededRemoveOpsAllowsFinalizerRemoval is the -// contrasting happy path: once the remove ops actually Succeeded, deletion -// must proceed exactly as before this fix. -func TestHandleDeletion_SucceededRemoveOpsAllowsFinalizerRemoval(t *testing.T) { - sn := newStorageNode("sn-1", snTestNS, "sns", snTestWorker) - sn.Finalizers = []string{storageNodeFinalizer} - sn.Status.UUID = "node-uuid-1" - sn.Status.Status = utils.NodeStatusOnline - sn.Status.ActiveOpsRef = "" - sns := newStorageNodeSet("sns", snTestNS, snTestCluster, nil) - ops := &simplyblockv1alpha2.StorageNodeOps{ - ObjectMeta: metav1.ObjectMeta{Name: "sn-1-remove", Namespace: snTestNS}, - Spec: simplyblockv1alpha2.StorageNodeOpsSpec{NodeRef: "sn-1", Action: simplyblockv1alpha2.StorageNodeOpsActionRemove}, - Status: simplyblockv1alpha2.StorageNodeOpsStatus{Phase: simplyblockv1alpha2.StorageNodeOpsPhaseSucceeded}, - } - r := newSNReconciler(t, sn, sns, ops) - - _, err := r.handleDeletion(context.Background(), sn, snTestCluster) - if err != nil { - t.Fatalf("handleDeletion returned error: %v", err) - } - - var updated simplyblockv1alpha1.StorageNode - _ = r.Get(context.Background(), types.NamespacedName{Name: "sn-1", Namespace: snTestNS}, &updated) - for _, f := range updated.Finalizers { - if f == storageNodeFinalizer { - t.Error("finalizer should have been removed once the remove ops succeeded") - } - } -} - -// TestCountInFlightNodes_DeduplicatesByWorker verifies that countInFlightNodes -// counts distinct physical hosts (WorkerNode), not individual StorageNode CRs. -// With nodesPerSocket=2 each host has two CRs; both get PostedAt stamped -// (primary posts, secondary copies via workerAlreadyPosted fast-path). Without -// deduplication each host would consume 2 slots instead of 1, causing -// maxParallelNodeAdds=6 to allow only ~3 hosts concurrently. -func TestCountInFlightNodes_DeduplicatesByWorker(t *testing.T) { - const ( - ns = snTestNS - snsRef = "sns-dedup" - ) - - now := metav1.Now() - postedAt := &now - - makeNode := func(name, worker, uuid, status string, posted *metav1.Time) *simplyblockv1alpha1.StorageNode { - sn := newStorageNode(name, ns, snsRef, worker) - sn.Status.PostedAt = posted - sn.Status.UUID = uuid - sn.Status.Status = status - return sn - } - - // Two sockets on worker-A: both in-flight (PostedAt set, no UUID yet). - snA1 := makeNode("sn-a1", "worker-a", "", "", postedAt) - snA2 := makeNode("sn-a2", "worker-a", "", "", postedAt) - - // Two sockets on worker-B: both in-flight. - snB1 := makeNode("sn-b1", "worker-b", "", "", postedAt) - snB2 := makeNode("sn-b2", "worker-b", "", "", postedAt) - - // worker-C: one socket done (UUID assigned), one still in_creation. - snC1 := makeNode("sn-c1", "worker-c", "uuid-c", utils.NodeStatusInCreation, postedAt) - snC2 := makeNode("sn-c2", "worker-c", "uuid-c", utils.NodeStatusInCreation, postedAt) - - // worker-D: timed out — should NOT count. - snD1 := makeNode("sn-d1", "worker-d", "", utils.NodeStatusTimeout, postedAt) - - // worker-E: no PostedAt — not yet started, should NOT count. - snE1 := makeNode("sn-e1", "worker-e", "", "", nil) - - r := newSNReconciler(t, snA1, snA2, snB1, snB2, snC1, snC2, snD1, snE1) - - t.Run("counts distinct in-flight workers, not CRs", func(t *testing.T) { - // Exclude worker-a (the calling node's worker). Expect worker-b and worker-c = 2. - count, err := r.countInFlightNodes(context.Background(), ns, snsRef, "worker-a") - if err != nil { - t.Fatalf("countInFlightNodes returned error: %v", err) - } - if count != 2 { - t.Fatalf("expected 2 distinct in-flight workers (b, c), got %d", count) - } - }) - - t.Run("excludes the calling node's own worker", func(t *testing.T) { - // Exclude worker-b. Expect worker-a and worker-c = 2. - count, err := r.countInFlightNodes(context.Background(), ns, snsRef, "worker-b") - if err != nil { - t.Fatalf("countInFlightNodes returned error: %v", err) - } - if count != 2 { - t.Fatalf("expected 2 distinct in-flight workers (a, c), got %d", count) - } - }) -} - -// The reconciler serves a node's status from the stream's cache, so a change -// the control plane pushed costs no request to read back. The server here -// fails every request and counts them: a reconciler that polled would both -// call it and fail to write a status. -func TestSyncStatusReadsThePushedNodeInsteadOfAskingForIt(t *testing.T) { - var calls int32 - srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) { - atomic.AddInt32(&calls, 1) - w.WriteHeader(http.StatusInternalServerError) - })) - defer srv.Close() - - const nodeUUID = "fd687dfd-9b5d-4eca-8cb1-23bcf550ad21" - sn := &simplyblockv1alpha1.StorageNode{ - ObjectMeta: metav1.ObjectMeta{Namespace: "default", Name: "simplyblock-node-asxeub"}, - Status: simplyblockv1alpha1.StorageNodeStatus{UUID: nodeUUID}, - } - - // The stream reported the node online, as its snapshot would. - sub := subscriptions.NewNodeSubscription() - sub.RegisterNode(nodeUUID, client.ObjectKeyFromObject(sn)) - err := sub.Ingest(context.Background(), cpinformer.Event{ - Kind: cpinformer.EventSnapshot, - Scope: cpinformer.Scope{"22222222-2222-2222-2222-222222222222"}, - Data: []byte(`[{"id":"` + nodeUUID + `","status":"online","mgmt_ip":"192.168.10.112", - "health_check":true,"hostname":"vm02_4420","cpu_spdk_count":6,"lvols":3, - "rpc_port":4420,"lvol_subsys_port":4426,"nvmf_port":4421,"failure_domain":-1}]`), - }) - if err != nil { - t.Fatalf("ingest snapshot: %v", err) - } - - r := newSNReconciler(t, sn) - r.Nodes = sub - - res, err := r.syncStatus(context.Background(), sn, "22222222-2222-2222-2222-222222222222", - webapi.NewClient(srv.URL)) - if err != nil { - t.Fatalf("syncStatus: %v", err) - } - - if n := atomic.LoadInt32(&calls); n != 0 { - t.Errorf("made %d control-plane request(s), want none for a node the stream already delivered", n) - } - - var got simplyblockv1alpha1.StorageNode - key := client.ObjectKeyFromObject(sn) - if err := r.Get(context.Background(), key, &got); err != nil { - t.Fatalf("get node: %v", err) - } - if got.Status.Status != "online" || !got.Status.Health { - t.Errorf("status/health = %q/%v, want online/true", got.Status.Status, got.Status.Health) - } - if got.Status.Hostname != snBackendHostname { - t.Errorf("hostname = %q", got.Status.Hostname) - } - if got.Status.Ports == nil || got.Status.Ports.Management != "192.168.10.112" { - t.Errorf("ports = %+v", got.Status.Ports) - } - if got.Status.Resources == nil || *got.Status.Resources.Volumes != 3 { - t.Errorf("resources = %+v", got.Status.Resources) - } - - // The poll survives as the correctness floor, relaxed because it is no - // longer how a change is noticed. - if res.RequeueAfter < time.Minute { - t.Errorf("RequeueAfter = %s, want the relaxed backstop rather than a poll", res.RequeueAfter) - } -} - -// A node the stream has not delivered still has to be readable, or a cold cache -// would leave the CR's status frozen until the first snapshot lands. -func TestSyncStatusFallsBackToTheControlPlaneForAnUncachedNode(t *testing.T) { - var calls int32 - srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) { - atomic.AddInt32(&calls, 1) - w.Header().Set("Content-Type", "application/json") - _, _ = w.Write([]byte(`{"id":"aaaa","status":"in_creation","health_check":false, - "hostname":"vm09_4420","cpu_spdk_count":4,"lvols":0,"mgmt_ip":"10.0.0.9", - "rpc_port":4420,"lvol_subsys_port":4426,"nvmf_port":4421,"failure_domain":-1}`)) - })) - defer srv.Close() - - sn := &simplyblockv1alpha1.StorageNode{ - ObjectMeta: metav1.ObjectMeta{Namespace: "default", Name: "simplyblock-node-new"}, - Status: simplyblockv1alpha1.StorageNodeStatus{UUID: "aaaa"}, - } - - r := newSNReconciler(t, sn) - r.Nodes = subscriptions.NewNodeSubscription() // connected, nothing delivered yet - - if _, err := r.syncStatus(context.Background(), sn, "cluster-1", webapi.NewClient(srv.URL)); err != nil { - t.Fatalf("syncStatus: %v", err) - } - if n := atomic.LoadInt32(&calls); n != 1 { - t.Errorf("made %d request(s), want exactly 1 for a node the stream has not delivered", n) - } - - var got simplyblockv1alpha1.StorageNode - if err := r.Get(context.Background(), client.ObjectKeyFromObject(sn), &got); err != nil { - t.Fatalf("get node: %v", err) - } - if got.Status.Status != nodeStatusInCreation || got.Status.Hostname != "vm09_4420" { - t.Errorf("status/hostname = %q/%q", got.Status.Status, got.Status.Hostname) - } -} - -// fakeNodeCapacity stands in for the metrics endpoint. -type fakeNodeCapacity struct { - samples map[string]prometheus.Capacity - err error -} - -func (f *fakeNodeCapacity) NodeCapacity( - _ context.Context, _ string, -) (map[string]prometheus.Capacity, error) { - if f.err != nil { - return nil, f.err - } - return f.samples, nil -} - -const ( - // snBackendHostname is how the control plane names a node, which is not the - // worker's Kubernetes name: the two are separate fields for that reason. - snBackendHostname = "vm02_4420" - snCapNode = "fd687dfd-9b5d-4eca-8cb1-23bcf550ad21" - snCapTotal = int64(112303538176) - snCapUsed = int64(422576128) -) - -func nodeWithUUID() *simplyblockv1alpha1.StorageNode { - return &simplyblockv1alpha1.StorageNode{ - ObjectMeta: metav1.ObjectMeta{Namespace: "default", Name: "simplyblock-node-cap"}, - Status: simplyblockv1alpha1.StorageNodeStatus{UUID: snCapNode}, - } -} - -func nodeRecording(used, total int64) *simplyblockv1alpha1.StorageNode { - sn := nodeWithUUID() - sn.Status.Resources = &simplyblockv1alpha1.StorageNodeResources{ - Capacity: &simplyblockv1alpha1.StorageNodeCapacity{ - TotalBytes: ptr.To(total), - UsedBytes: ptr.To(used), - }, - } - return sn -} - -func applyWithCapacity( - t *testing.T, sn *simplyblockv1alpha1.StorageNode, src NodeCapacitySource, -) simplyblockv1alpha1.StorageNode { - t.Helper() - r := newSNReconciler(t, sn) - r.Capacity = src - if _, err := r.applyNodeStatus(context.Background(), sn, "cluster-1", - SNODEAPIResponse{Status: nodeStatusOnline, Hostname: snBackendHostname}); err != nil { - t.Fatalf("applyNodeStatus: %v", err) - } - var got simplyblockv1alpha1.StorageNode - if err := r.Get(context.Background(), client.ObjectKeyFromObject(sn), &got); err != nil { - t.Fatalf("get node: %v", err) - } - return got -} - -// A node's occupancy is in neither the control plane's API nor its stream, only -// in the metrics it exports, so a first reading is published from there. -func TestNodeStatusPublishesTheCapacitySample(t *testing.T) { - sampledAt := time.Unix(1788423117, 0).UTC() - got := applyWithCapacity(t, nodeWithUUID(), &fakeNodeCapacity{ - samples: map[string]prometheus.Capacity{ - snCapNode: {Total: snCapTotal, Used: snCapUsed, SampledAt: sampledAt}, - }, - }) - - c := got.Status.Resources.Capacity - if c == nil { - t.Fatal("capacity was not published") - } - if *c.TotalBytes != snCapTotal || *c.UsedBytes != snCapUsed { - t.Errorf("total/used = %d/%d", *c.TotalBytes, *c.UsedBytes) - } - if c.SampledAt == nil || !c.SampledAt.Time.Equal(sampledAt) { - t.Errorf("sampledAt = %v, want %s", c.SampledAt, sampledAt) - } -} - -// The reconciler watches its own objects, so a status write schedules another -// reconcile. Recording every sample would make a node reconcile itself for as -// long as any I/O was happening, so a reading that has barely moved is left -// alone and the patch stays empty. -func TestASampleThatBarelyMovedIsNotWritten(t *testing.T) { - // One mebibyte more on a 104 GiB node: far under one percent. - got := applyWithCapacity(t, nodeRecording(snCapUsed, snCapTotal), &fakeNodeCapacity{ - samples: map[string]prometheus.Capacity{ - snCapNode: { - Total: snCapTotal, Used: snCapUsed + 1048576, - SampledAt: time.Unix(1788423200, 0).UTC(), - }, - }, - }) - - if *got.Status.Resources.Capacity.UsedBytes != snCapUsed { - t.Errorf("used was rewritten to %d for a sub-threshold move", - *got.Status.Resources.Capacity.UsedBytes) - } -} - -// A move worth knowing about is recorded. Without this the suppression above -// would be indistinguishable from never updating at all. -func TestASampleThatMovedMateriallyIsWritten(t *testing.T) { - want := snCapUsed + 10*1024*1024*1024 // ten gibibytes: over one percent - got := applyWithCapacity(t, nodeRecording(snCapUsed, snCapTotal), &fakeNodeCapacity{ - samples: map[string]prometheus.Capacity{ - snCapNode: {Total: snCapTotal, Used: want, SampledAt: time.Unix(1788423200, 0).UTC()}, - }, - }) - - if *got.Status.Resources.Capacity.UsedBytes != want { - t.Errorf("used = %d, want %d", *got.Status.Resources.Capacity.UsedBytes, want) - } -} - -// A device joining or leaving changes the total, which is worth recording -// however little the used size moved. -func TestAChangedTotalIsAlwaysWritten(t *testing.T) { - const grown = int64(168455307264) - got := applyWithCapacity(t, nodeRecording(snCapUsed, snCapTotal), &fakeNodeCapacity{ - samples: map[string]prometheus.Capacity{ - snCapNode: {Total: grown, Used: snCapUsed, SampledAt: time.Unix(1788423200, 0).UTC()}, - }, - }) - - if *got.Status.Resources.Capacity.TotalBytes != grown { - t.Errorf("total = %d, want the new one", *got.Status.Resources.Capacity.TotalBytes) - } -} - -// Prometheus being unreachable leaves the rest of the status correct. A node -// whose occupancy is momentarily unknown is better published than not. -func TestAFailingNodeCapacitySourceStillPublishesTheNode(t *testing.T) { - got := applyWithCapacity(t, nodeWithUUID(), - &fakeNodeCapacity{err: errors.New("connection refused")}) - - if got.Status.Status != nodeStatusOnline || got.Status.Hostname != snBackendHostname { - t.Errorf("status/hostname = %q/%q", got.Status.Status, got.Status.Hostname) - } - if got.Status.Resources.Capacity != nil { - t.Errorf("capacity = %+v, want absent rather than zero", got.Status.Resources.Capacity) - } -} - -// A node the exporter has never measured has no capacity to publish, and zeros -// would read as an empty node rather than an unmeasured one. -func TestAnUnsampledNodePublishesNoCapacity(t *testing.T) { - got := applyWithCapacity(t, nodeWithUUID(), &fakeNodeCapacity{ - samples: map[string]prometheus.Capacity{ - snCapNode: {Total: 0, Used: 0}, // no SampledAt: never measured - }, - }) - - if got.Status.Resources.Capacity != nil { - t.Errorf("capacity = %+v, want absent", got.Status.Resources.Capacity) - } -} - -// Adoption matches a StorageNode to its backend node by the worker's address, -// and the node stream already holds every node of the cluster. Reading the list -// over HTTP would ask the control plane for what is in memory, so the server -// here fails every request and counts them. -func TestAdoptionMatchesTheBackendNodeFromTheStreamCache(t *testing.T) { - const ( - cluster = "22222222-2222-2222-2222-222222222222" - nodeUUID = "fd687dfd-9b5d-4eca-8cb1-23bcf550ad21" - workerIP = "192.168.10.112" - ) - var calls int32 - srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) { - atomic.AddInt32(&calls, 1) - w.WriteHeader(http.StatusInternalServerError) - })) - defer srv.Close() - - worker := &corev1.Node{ - ObjectMeta: metav1.ObjectMeta{Name: "vm02.example.com"}, - Status: corev1.NodeStatus{Addresses: []corev1.NodeAddress{ - {Type: corev1.NodeInternalIP, Address: workerIP}, - }}, - } - sn := &simplyblockv1alpha1.StorageNode{ - ObjectMeta: metav1.ObjectMeta{Namespace: "default", Name: "simplyblock-node-adopt"}, - Spec: simplyblockv1alpha1.StorageNodeSpec{WorkerNode: worker.Name}, - } - - // The stream has delivered the cluster's nodes, which is what makes the - // scope authoritative enough to match against. - sub := subscriptions.NewNodeSubscription() - err := sub.Ingest(context.Background(), cpinformer.Event{ - Kind: cpinformer.EventSnapshot, - Scope: cpinformer.Scope{cluster}, - Data: []byte(`[{"id":"` + nodeUUID + `","status":"online","mgmt_ip":"` + workerIP + `", - "health_check":true,"hostname":"vm02_4420","cpu_spdk_count":6,"lvols":0, - "rpc_port":4420,"lvol_subsys_port":4426,"nvmf_port":4421,"failure_domain":-1}]`), - }) - if err != nil { - t.Fatalf("ingest snapshot: %v", err) - } - - r := newSNReconciler(t, sn, worker) - r.Nodes = sub - - if err := r.pollUUIDFromBackend(context.Background(), sn, cluster, - webapi.NewClient(srv.URL)); err != nil { - t.Fatalf("pollUUIDFromBackend: %v", err) - } - - if n := atomic.LoadInt32(&calls); n != 0 { - t.Errorf("listed the cluster's nodes over HTTP %d time(s), want none", n) - } - - var got simplyblockv1alpha1.StorageNode - if err := r.Get(context.Background(), client.ObjectKeyFromObject(sn), &got); err != nil { - t.Fatalf("get node: %v", err) - } - if got.Status.UUID != nodeUUID { - t.Errorf("adopted UUID = %q, want %q", got.Status.UUID, nodeUUID) - } - if got.Status.Status != nodeStatusOnline || got.Status.Hostname != snBackendHostname { - t.Errorf("status/hostname = %q/%q", got.Status.Status, got.Status.Hostname) - } -} - -// An unsynced cache is not evidence that the cluster has no nodes, so adoption -// asks the control plane rather than concluding there is nothing to adopt. -func TestAdoptionFallsBackToTheControlPlaneBeforeTheSnapshotArrives(t *testing.T) { - const ( - cluster = "22222222-2222-2222-2222-222222222222" - nodeUUID = "fd687dfd-9b5d-4eca-8cb1-23bcf550ad21" - workerIP = "192.168.10.112" - ) - var calls int32 - srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) { - atomic.AddInt32(&calls, 1) - w.Header().Set("Content-Type", "application/json") - _, _ = w.Write([]byte(`[{"id":"` + nodeUUID + `","status":"online","mgmt_ip":"` + workerIP + `", - "health_check":true,"hostname":"vm02_4420","rpc_port":4420,"failure_domain":-1}]`)) - })) - defer srv.Close() - - worker := &corev1.Node{ - ObjectMeta: metav1.ObjectMeta{Name: "vm02.example.com"}, - Status: corev1.NodeStatus{Addresses: []corev1.NodeAddress{ - {Type: corev1.NodeInternalIP, Address: workerIP}, - }}, - } - sn := &simplyblockv1alpha1.StorageNode{ - ObjectMeta: metav1.ObjectMeta{Namespace: "default", Name: "simplyblock-node-adopt"}, - Spec: simplyblockv1alpha1.StorageNodeSpec{WorkerNode: worker.Name}, - } - - r := newSNReconciler(t, sn, worker) - r.Nodes = subscriptions.NewNodeSubscription() // connected, no snapshot yet - - if err := r.pollUUIDFromBackend(context.Background(), sn, cluster, - webapi.NewClient(srv.URL)); err != nil { - t.Fatalf("pollUUIDFromBackend: %v", err) - } - - if n := atomic.LoadInt32(&calls); n != 1 { - t.Errorf("made %d request(s), want exactly 1 while the cache is cold", n) - } - var got simplyblockv1alpha1.StorageNode - if err := r.Get(context.Background(), client.ObjectKeyFromObject(sn), &got); err != nil { - t.Fatalf("get node: %v", err) - } - if got.Status.UUID != nodeUUID { - t.Errorf("adopted UUID = %q, want %q", got.Status.UUID, nodeUUID) - } -} diff --git a/operator/internal/controller/storagenode_latency_controller.go b/operator/internal/controller/storagenode_latency_controller.go index cdc425d27..07b00ec67 100644 --- a/operator/internal/controller/storagenode_latency_controller.go +++ b/operator/internal/controller/storagenode_latency_controller.go @@ -36,7 +36,6 @@ import ( "sigs.k8s.io/controller-runtime/pkg/controller" logf "sigs.k8s.io/controller-runtime/pkg/log" - simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" "github.com/simplyblock/simplyblock-operator/internal/autoplacement" "github.com/simplyblock/simplyblock-operator/internal/utils" @@ -69,8 +68,8 @@ type StorageNodeLatencyReconciler struct { // Set to WebAPIBenchmarkProvisioner for test environments that require explicit provisioning. Provisioner BenchmarkProvisioner - // APIClient queries the SimplyBlock REST API to resolve a storage node's - // data-network IP (the /nics endpoint). Independent of the provisioner. + // APIClient queries the simplyblock REST API to resolve a storage node's + // data-network IP from its NIC listing. Independent of the provisioner. APIClient *webapi.Client } @@ -85,7 +84,7 @@ type StorageNodeLatencyReconciler struct { func (r *StorageNodeLatencyReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) { log := logf.FromContext(ctx) - snode := &simplyblockv1alpha1.StorageNodeSet{} + snode := &simplyblockv1alpha2.StorageCluster{} if err := r.Get(ctx, req.NamespacedName, snode); err != nil { return ctrl.Result{}, client.IgnoreNotFound(err) } @@ -93,7 +92,7 @@ func (r *StorageNodeLatencyReconciler) Reconcile(ctx context.Context, req ctrl.R clusterCR := &simplyblockv1alpha2.StorageCluster{} if err := r.Get(ctx, types.NamespacedName{ Namespace: req.Namespace, - Name: snode.Spec.ClusterName, + Name: snode.Name, }, clusterCR); err != nil { if apierrors.IsNotFound(err) { return ctrl.Result{RequeueAfter: 30 * time.Second}, nil @@ -132,36 +131,34 @@ func (r *StorageNodeLatencyReconciler) Reconcile(ctx context.Context, req ctrl.R return ctrl.Result{RequeueAfter: 30 * time.Second}, nil } - poolUUID, err := r.Provisioner.EnsurePool(ctx, snode.Namespace, snode.Spec.ClusterName) + poolUUID, err := r.Provisioner.EnsurePool(ctx, snode.Namespace, snode.Name) if err != nil { log.Error(err, "Cannot ensure benchmark pool") return ctrl.Result{RequeueAfter: 30 * time.Second}, nil } - // One baseline Job per node UUID. On NUMA hosts multiple backend nodes share the - // same k8s hostname but have independent NVMe devices and independent latency - // characteristics, so every node UUID is measured separately. - nodesByUUID := map[string]simplyblockv1alpha1.NodeStatus{} - for _, n := range snode.Status.Nodes { - if n.UUID == "" || n.Status != nodeStatusOnline || !n.Health || n.Hostname == "" { - continue - } - if _, seen := nodesByUUID[n.UUID]; !seen { - nodesByUUID[n.UUID] = n - } + // One baseline Job per backend node. On NUMA hosts several of them share a + // Kubernetes hostname while having independent NVMe devices and independent + // latency characteristics, so each is measured separately — which is what makes + // the StorageNode object, rather than the host, the thing a reading belongs to. + var nodes simplyblockv1alpha2.StorageNodeList + if err := r.List(ctx, &nodes, client.InNamespace(snode.Namespace)); err != nil { + return ctrl.Result{RequeueAfter: 30 * time.Second}, nil } - latencyMetrics := r.copyLatencyMetrics(snode.Status.LatencyMetrics) - // hostConfigs accumulates per-node configs keyed by k8s hostname so the sidecar - // (one pod per host) receives a JSON array covering all NUMA nodes on its host. + // hostConfigs accumulates per-node configs keyed by Kubernetes hostname so the + // sidecar, which is one pod per host, receives a JSON array covering every NUMA + // node on its host. hostConfigs := map[string][]autoplacement.NodeConfig{} - changed := false - for _, node := range nodesByUUID { - nodeChanged := r.processNodeBaseline(ctx, snode, clusterCR, poolUUID, node, rebalancerImage, &latencyMetrics, hostConfigs) - if nodeChanged { - changed = true + for i := range nodes.Items { + node := &nodes.Items[i] + if node.Spec.ClusterRef != snode.Name || node.Status.UUID == "" || + node.Status.Status != utils.NodeStatusOnline || !node.Status.Health || + node.Status.Hostname == "" { + continue } + r.processNodeBaseline(ctx, snode, clusterCR, poolUUID, node, rebalancerImage, hostConfigs) } configData := make(map[string]string, len(hostConfigs)) @@ -169,23 +166,10 @@ func (r *StorageNodeLatencyReconciler) Reconcile(ctx context.Context, req ctrl.R raw, _ := json.Marshal(cfgs) configData[hostname] = string(raw) } - if err := r.reconcileConfigMap(ctx, snode.Namespace, snode.Spec.ClusterName, configData); err != nil { + if err := r.reconcileConfigMap(ctx, snode.Namespace, snode.Name, configData); err != nil { log.Error(err, "Cannot reconcile simplyblock-rebalancer ConfigMap") } - if changed { - if err := r.patchLatencyStatus(ctx, snode, latencyMetrics); err != nil { - if apierrors.IsConflict(err) { - // Stale snapshot — the StorageNode status was updated concurrently. The - // optimistic lock prevented clobbering existing baselines; requeue to - // recompute from fresh state. - return ctrl.Result{Requeue: true}, nil - } - log.Error(err, "Failed to patch StorageNode latency status") - return ctrl.Result{RequeueAfter: 10 * time.Second}, nil - } - } - return ctrl.Result{RequeueAfter: benchInterval}, nil } @@ -194,59 +178,62 @@ func (r *StorageNodeLatencyReconciler) Reconcile(ctx context.Context, req ctrl.R // Returns true when latencyMetrics was changed and needs to be patched. func (r *StorageNodeLatencyReconciler) processNodeBaseline( ctx context.Context, - snode *simplyblockv1alpha1.StorageNodeSet, + snode *simplyblockv1alpha2.StorageCluster, clusterCR *simplyblockv1alpha2.StorageCluster, poolUUID string, - node simplyblockv1alpha1.NodeStatus, + node *simplyblockv1alpha2.StorageNode, image string, - latencyMetrics *[]simplyblockv1alpha1.NodeLatencyMetrics, hostConfigs map[string][]autoplacement.NodeConfig, -) bool { +) { log := logf.FromContext(ctx) - m := r.findOrCreateEntry(*latencyMetrics, node.UUID) + nodeUUID := node.Status.UUID + measured := node.Status.LatencyMetrics volumeUUID, err := r.Provisioner.EnsureVolume( - ctx, snode.Namespace, snode.Spec.ClusterName, poolUUID, - "simplyblock-rebalancer-"+node.UUID, node.UUID, + ctx, snode.Namespace, snode.Name, poolUUID, + "simplyblock-rebalancer-"+nodeUUID, nodeUUID, ) if err != nil { - log.Error(err, "Cannot ensure benchmark volume", "node", node.UUID) - return false + log.Error(err, "Cannot ensure benchmark volume", "node", nodeUUID) + return } conn := benchmarkConnInfo{ NQN: r.Provisioner.BenchmarkNQN(clusterCR.Status.NQN, volumeUUID), - Addr: node.MgmtIp, + Addr: managementAddress(node), Port: logicalVolumeConnectionPort(node), } // The lvol's NVMe-oF subsystem listens on the node's data NIC, not its management - // IP, so targeting node.MgmtIp fails with "connection refused". Resolve the node's - // data-network address from the /nics endpoint; fall back to the management address + // IP, so targeting the management address fails with a refused connection. + // Resolve the node's data-network address from its NIC listing, and fall back to + // the management address // only when it cannot be resolved. - if dataAddr, err := r.nodeDataAddr(ctx, clusterCR.Status.UUID, node.UUID); err != nil { + if dataAddr, err := r.nodeDataAddr(ctx, clusterCR.Status.UUID, nodeUUID); err != nil { log.Info("Could not resolve data-network address; falling back to management IP", - "node", node.UUID, "addr", conn.Addr, "error", err.Error()) + "node", nodeUUID, "addr", conn.Addr, "error", err.Error()) } else { conn.Addr = dataAddr } - changed := false - if m.BaselineP99NS == 0 { - baseline, jobChanged, err := r.reconcileBaselineJob(ctx, snode, node, conn, image) + if measured == nil || measured.BaselineP99NS == 0 { + baseline, _, err := r.reconcileBaselineJob(ctx, snode, node, conn, image) if err != nil { - log.Error(err, "Baseline job error", "node", node.UUID) + log.Error(err, "Baseline job error", "node", nodeUUID) } if baseline != nil { now := metav1.NewTime(time.Now()) - m.BaselineP50NS = baseline.P50NS - m.BaselineP99NS = baseline.P99NS - m.BaselineMeasuredAt = &now - log.Info("Baseline measured", "node", node.UUID, "p50ns", baseline.P50NS, "p99ns", baseline.P99NS) - jobChanged = true - } - if jobChanged { - changed = true + measured = &simplyblockv1alpha2.NodeLatencyMetrics{ + NodeUUID: nodeUUID, + BaselineP50NS: baseline.P50NS, + BaselineP99NS: baseline.P99NS, + BaselineMeasuredAt: &now, + } + log.Info("Baseline measured", "node", nodeUUID, + "p50ns", baseline.P50NS, "p99ns", baseline.P99NS) + if err := r.patchLatencyStatus(ctx, node, measured); err != nil { + log.Error(err, "Failed to record the node's baseline", "node", nodeUUID) + } } } @@ -254,18 +241,16 @@ func (r *StorageNodeLatencyReconciler) processNodeBaseline( // This prevents the sidecar's continuous fio loop from running concurrently // with the one-shot baseline Job — both would write to the same NVMe device // and corrupt each other's measurements. - if m.BaselineP99NS > 0 { - hostConfigs[node.Hostname] = append(hostConfigs[node.Hostname], autoplacement.NodeConfig{ + if measured != nil && measured.BaselineP99NS > 0 { + host := node.Status.Hostname + hostConfigs[host] = append(hostConfigs[host], autoplacement.NodeConfig{ NQN: conn.NQN, Addr: conn.Addr, Port: conn.Port, - NodeUUID: node.UUID, + NodeUUID: nodeUUID, ClusterUUID: clusterCR.Status.UUID, }) } - - *latencyMetrics = r.setEntry(*latencyMetrics, m) - return changed } // benchmarkConnInfo holds the NVMe-oF connection parameters for the benchmark volume. @@ -281,12 +266,12 @@ type benchmarkConnInfo struct { // for a single backend node. Returns the parsed result once the Job succeeds. func (r *StorageNodeLatencyReconciler) reconcileBaselineJob( ctx context.Context, - snode *simplyblockv1alpha1.StorageNodeSet, - node simplyblockv1alpha1.NodeStatus, + snode *simplyblockv1alpha2.StorageCluster, + node *simplyblockv1alpha2.StorageNode, conn benchmarkConnInfo, image string, ) (*autoplacement.LatencyResult, bool, error) { - jobName := baselineJobNamePrefix + safeNodeID(node.UUID) + jobName := baselineJobNamePrefix + safeNodeID(node.Status.UUID) job := &batchv1.Job{} err := r.Get(ctx, types.NamespacedName{Namespace: snode.Namespace, Name: jobName}, job) @@ -317,7 +302,7 @@ func (r *StorageNodeLatencyReconciler) reconcileBaselineJob( } // No job yet — create one if all connection info is available. - if node.Hostname == "" || conn.Addr == "" || conn.NQN == "" { + if node.Status.Hostname == "" || conn.Addr == "" || conn.NQN == "" { return nil, false, nil } if createErr := r.createBaselineJob(ctx, snode, node, conn, image); createErr != nil { @@ -328,8 +313,8 @@ func (r *StorageNodeLatencyReconciler) reconcileBaselineJob( func (r *StorageNodeLatencyReconciler) createBaselineJob( ctx context.Context, - snode *simplyblockv1alpha1.StorageNodeSet, - node simplyblockv1alpha1.NodeStatus, + snode *simplyblockv1alpha2.StorageCluster, + node *simplyblockv1alpha2.StorageNode, conn benchmarkConnInfo, image string, ) error { @@ -340,14 +325,14 @@ func (r *StorageNodeLatencyReconciler) createBaselineJob( return r.Create(ctx, &batchv1.Job{ ObjectMeta: metav1.ObjectMeta{ - Name: baselineJobNamePrefix + safeNodeID(node.UUID), + Name: baselineJobNamePrefix + safeNodeID(node.Status.UUID), Namespace: snode.Namespace, Labels: map[string]string{ baselineJobLabelKey: "true", - baselineJobNodeLabelKey: node.UUID, + baselineJobNodeLabelKey: node.Status.UUID, }, OwnerReferences: []metav1.OwnerReference{ - *metav1.NewControllerRef(snode, simplyblockv1alpha1.GroupVersion.WithKind("StorageNodeSet")), + *metav1.NewControllerRef(snode, simplyblockv1alpha2.GroupVersion.WithKind("StorageCluster")), }, }, Spec: batchv1.JobSpec{ @@ -356,7 +341,7 @@ func (r *StorageNodeLatencyReconciler) createBaselineJob( Template: corev1.PodTemplateSpec{ Spec: corev1.PodSpec{ RestartPolicy: corev1.RestartPolicyNever, - NodeSelector: map[string]string{"kubernetes.io/hostname": node.Hostname}, + NodeSelector: map[string]string{"kubernetes.io/hostname": node.Status.Hostname}, HostNetwork: true, Volumes: []corev1.Volume{ { @@ -397,7 +382,7 @@ func (r *StorageNodeLatencyReconciler) createBaselineJob( }) } -// nodeDataAddr resolves a storage node's data-network IP via the /nics endpoint, +// nodeDataAddr resolves a storage node's data-network IP from its NIC listing, // returning the first interface that is UP with a non-empty address. The lvol // subsystem listens on the data NIC, so the fio baseline must target this address // rather than the node's management IP. Returns an error when no API client is @@ -419,10 +404,9 @@ func (r *StorageNodeLatencyReconciler) nodeDataAddr(ctx context.Context, cluster } // logicalVolumeConnectionPort returns the NVMe/TCP connection port for a node, falling back to 4430 if not reported. -func logicalVolumeConnectionPort(node simplyblockv1alpha1.NodeStatus) int32 { - port := ptr.IntFromOrZero(node.LvolPort) - if port > 0 { - return int32(port) +func logicalVolumeConnectionPort(node *simplyblockv1alpha2.StorageNode) int32 { + if node.Status.Ports != nil && node.Status.Ports.Lvol != nil && *node.Status.Ports.Lvol > 0 { + return *node.Status.Ports.Lvol } return 4430 } @@ -453,7 +437,7 @@ func (r *StorageNodeLatencyReconciler) readJobResult(ctx context.Context, job *b return nil, fmt.Errorf("no termination message for job %s", job.Name) } -// reconcileConfigMap creates or updates the per-cluster ConfigMap that maps k8s +// reconcileConfigMap creates or updates the per-cluster ConfigMap that maps Kubernetes // node hostname → benchmark volume config JSON consumed by the simplyblock-rebalancer sidecar. func (r *StorageNodeLatencyReconciler) reconcileConfigMap( ctx context.Context, @@ -494,51 +478,30 @@ func (r *StorageNodeLatencyReconciler) jobFailed(job *batchv1.Job) bool { return false } +// patchLatencyStatus records one node's baseline on the node itself. +// +// The reading moved off the retired fleet object for the reason the reading is one +// node's: a fleet-wide list made every node's measurement a write to one object +// shared by all of them, and a stale snapshot of that list silently dropped +// entries a concurrent reconcile had written. One node, one write, and no list to +// lose an entry from. func (r *StorageNodeLatencyReconciler) patchLatencyStatus( ctx context.Context, - snode *simplyblockv1alpha1.StorageNodeSet, - latencyMetrics []simplyblockv1alpha1.NodeLatencyMetrics, + node *simplyblockv1alpha2.StorageNode, + measured *simplyblockv1alpha2.NodeLatencyMetrics, ) error { - orig := snode.DeepCopy() - snode.Status.LatencyMetrics = latencyMetrics - // Optimistic lock: the LatencyMetrics array is replaced wholesale by this merge patch, - // so a stale snapshot would silently drop entries written by a concurrent reconcile. - // Pinning the resourceVersion turns that lost update into a Conflict the caller requeues on. - return r.Status().Patch(ctx, snode, client.MergeFromWithOptions(orig, client.MergeFromWithOptimisticLock{})) -} - -func (r *StorageNodeLatencyReconciler) copyLatencyMetrics( - src []simplyblockv1alpha1.NodeLatencyMetrics, -) []simplyblockv1alpha1.NodeLatencyMetrics { - out := make([]simplyblockv1alpha1.NodeLatencyMetrics, len(src)) - copy(out, src) - return out + patch := client.MergeFrom(node.DeepCopy()) + node.Status.LatencyMetrics = measured + return r.Status().Patch(ctx, node, patch) } -func (r *StorageNodeLatencyReconciler) findOrCreateEntry( - metrics []simplyblockv1alpha1.NodeLatencyMetrics, - nodeUUID string, -) *simplyblockv1alpha1.NodeLatencyMetrics { - for i := range metrics { - if metrics[i].NodeUUID == nodeUUID { - cp := metrics[i] - return &cp - } - } - return &simplyblockv1alpha1.NodeLatencyMetrics{NodeUUID: nodeUUID} -} - -func (r *StorageNodeLatencyReconciler) setEntry( - metrics []simplyblockv1alpha1.NodeLatencyMetrics, - m *simplyblockv1alpha1.NodeLatencyMetrics, -) []simplyblockv1alpha1.NodeLatencyMetrics { - for i := range metrics { - if metrics[i].NodeUUID == m.NodeUUID { - metrics[i] = *m - return metrics - } +// managementAddress is the node's reported management IP, which is the fallback a +// benchmark connects on when the data NIC cannot be resolved. +func managementAddress(node *simplyblockv1alpha2.StorageNode) string { + if node.Status.Ports == nil { + return "" } - return append(metrics, *m) + return node.Status.Ports.Management } func (r *StorageNodeLatencyReconciler) SetupWithManager(mgr ctrl.Manager) error { @@ -546,7 +509,7 @@ func (r *StorageNodeLatencyReconciler) SetupWithManager(mgr ctrl.Manager) error r.Provisioner = &AutomaticBenchmarkProvisioner{} } return ctrl.NewControllerManagedBy(mgr). - For(&simplyblockv1alpha1.StorageNodeSet{}). + For(&simplyblockv1alpha2.StorageCluster{}). Owns(&batchv1.Job{}). Named("storagenodelatency"). WithOptions(controller.Options{MaxConcurrentReconciles: 1}). diff --git a/operator/internal/controller/storagenodeops_controller.go b/operator/internal/controller/storagenodeops_controller.go deleted file mode 100644 index b8f59a918..000000000 --- a/operator/internal/controller/storagenodeops_controller.go +++ /dev/null @@ -1,1786 +0,0 @@ -/* -Copyright 2025. - -Licensed under the Apache License, Version 2.0 (the "License"); -you may not use this file except in compliance with the License. -You may obtain a copy of the License at - - http://www.apache.org/licenses/LICENSE-2.0 - -Unless required by applicable law or agreed to in writing, software -distributed under the License is distributed on an "AS IS" BASIS, -WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. -See the License for the specific language governing permissions and -limitations under the License. -*/ - -package controller - -import ( - "context" - "encoding/json" - "fmt" - "net/http" - "regexp" - "strings" - "time" - - corev1 "k8s.io/api/core/v1" - discoveryv1 "k8s.io/api/discovery/v1" - apierrors "k8s.io/apimachinery/pkg/api/errors" - metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" - "k8s.io/apimachinery/pkg/runtime" - "k8s.io/apimachinery/pkg/types" - "k8s.io/client-go/tools/events" - ctrl "sigs.k8s.io/controller-runtime" - "sigs.k8s.io/controller-runtime/pkg/client" - "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil" - "sigs.k8s.io/controller-runtime/pkg/handler" - logf "sigs.k8s.io/controller-runtime/pkg/log" - "sigs.k8s.io/controller-runtime/pkg/reconcile" - - "github.com/simplyblock/atlas/kube" - "github.com/simplyblock/atlas/ptr" - - simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" - simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" - clustercontroller "github.com/simplyblock/simplyblock-operator/internal/controllers/cluster" - "github.com/simplyblock/simplyblock-operator/internal/utils" - "github.com/simplyblock/simplyblock-operator/internal/webapi" -) - -// StorageNodeOpsReconciler drives all imperative StorageNode operations. -// It replaces the existing action-handling in StorageNodeSetReconciler and -// owns VolumeMigration CRs during drain (action=remove). -type StorageNodeOpsReconciler struct { - client.Client - Scheme *runtime.Scheme - Recorder events.EventRecorder - // apiReader is an uncached reader (mgr.GetAPIReader) used for the migrate - // DNS gate (endpointSliceHasWorker). A stale informer cache could otherwise - // miss the target worker's freshly-published storage-node-api endpoint and - // wedge the migration in Preparing on "waiting for DNS" indefinitely. - apiReader client.Reader -} - -// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodeops,verbs=get;list;watch;create;update;patch;delete -// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodeops/status,verbs=get;update;patch -// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodeops/finalizers,verbs=update -// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodes,verbs=get;list;watch;update;patch -// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodes/status,verbs=get;update;patch -// +kubebuilder:rbac:groups="",resources=nodes,verbs=get;list;watch;update;patch -// +kubebuilder:rbac:groups="",resources=pods,verbs=get;list;watch -// +kubebuilder:rbac:groups=events.k8s.io,resources=events,verbs=create;patch - -func (r *StorageNodeOpsReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) { - log := logf.FromContext(ctx) - - var ops simplyblockv1alpha2.StorageNodeOps - if err := r.Get(ctx, req.NamespacedName, &ops); err != nil { - if apierrors.IsNotFound(err) { - return ctrl.Result{}, nil - } - return ctrl.Result{}, err - } - - // Terminal — nothing left to do. - if ops.Status.Phase == simplyblockv1alpha2.StorageNodeOpsPhaseSucceeded || - ops.Status.Phase == simplyblockv1alpha2.StorageNodeOpsPhaseFailed { - return ctrl.Result{}, nil - } - - // Fetch the target StorageNode. - var sn simplyblockv1alpha1.StorageNode - if err := r.Get(ctx, types.NamespacedName{ - Name: ops.Spec.NodeRef, - Namespace: ops.Namespace, - }, &sn); err != nil { - if apierrors.IsNotFound(err) { - return r.failOps(ctx, &ops, "target StorageNode not found") - } - return ctrl.Result{}, err - } - - // Fetch the parent StorageNodeSet for cluster config. - var sns simplyblockv1alpha1.StorageNodeSet - if err := r.Get(ctx, types.NamespacedName{ - Name: sn.Spec.StorageNodeSetRef, - Namespace: sn.Namespace, - }, &sns); err != nil { - return ctrl.Result{}, client.IgnoreNotFound(err) - } - - // Resolve cluster UUID. - clusterUUID, err := utils.ResolveClusterUUID(ctx, r.Client, sn.Namespace, sns.Spec.ClusterName) - if err != nil { - log.Info("cluster UUID not ready, requeuing", "cluster", sns.Spec.ClusterName) - return ctrl.Result{RequeueAfter: 10 * time.Second}, nil - } - - apiClient := webapi.NewClient() - - // Mutual exclusion: only one ops may run per StorageNode at a time. - if ops.Status.Phase == "" || ops.Status.Phase == simplyblockv1alpha2.StorageNodeOpsPhasePending { - return r.acquireLock(ctx, &ops, &sn) - } - - // Cluster pause check for drain operations. - if ops.Spec.Action == simplyblockv1alpha2.StorageNodeOpsActionRemove { - if res, paused := r.clusterPauseCheck(ctx, &ops, apiClient); paused { - return res, nil - } - } - - log.Info("dispatching ops", "action", ops.Spec.Action, "subPhase", ops.Status.SubPhase) - return r.dispatch(ctx, &ops, &sn, &sns, clusterUUID, apiClient) -} - -// acquireLock attempts to set StorageNode.status.activeOpsRef to this ops. -// Requeues if another ops holds the lock. -func (r *StorageNodeOpsReconciler) acquireLock( - ctx context.Context, - ops *simplyblockv1alpha2.StorageNodeOps, - sn *simplyblockv1alpha1.StorageNode, -) (ctrl.Result, error) { - log := logf.FromContext(ctx) - - if sn.Status.ActiveOpsRef != "" && sn.Status.ActiveOpsRef != ops.Name { - log.Info("another ops is active, requeuing", "activeOps", sn.Status.ActiveOpsRef) - return ctrl.Result{RequeueAfter: 15 * time.Second}, nil - } - - snPatch := client.MergeFrom(sn.DeepCopy()) - sn.Status.ActiveOpsRef = ops.Name - if err := r.Status().Patch(ctx, sn, snPatch); err != nil { - return ctrl.Result{}, fmt.Errorf("setting activeOpsRef: %w", err) - } - - now := metav1.Now() - opsPatch := client.MergeFrom(ops.DeepCopy()) - ops.Status.Phase = simplyblockv1alpha2.StorageNodeOpsPhaseRunning - ops.Status.StartedAt = &now - if ops.Spec.Action == simplyblockv1alpha2.StorageNodeOpsActionRemove { - ops.Status.SubPhase = simplyblockv1alpha2.StorageNodeOpsSubPhaseValidating - } - if err := r.Status().Patch(ctx, ops, opsPatch); err != nil { - return ctrl.Result{}, err - } - return ctrl.Result{Requeue: true}, nil -} - -// dispatch routes the ops to the correct handler. -func (r *StorageNodeOpsReconciler) dispatch( - ctx context.Context, - ops *simplyblockv1alpha2.StorageNodeOps, - sn *simplyblockv1alpha1.StorageNode, - sns *simplyblockv1alpha1.StorageNodeSet, - clusterUUID string, - apiClient *webapi.Client, -) (ctrl.Result, error) { - switch ops.Spec.Action { - case simplyblockv1alpha2.StorageNodeOpsActionRemove: - return r.runDrain(ctx, ops, sn, clusterUUID, apiClient) - case simplyblockv1alpha2.StorageNodeOpsActionMigrate: - return r.runMigrate(ctx, ops, sn, sns, clusterUUID, apiClient) - case "shutdown", "restart", "suspend", "resume": - return r.runSimpleAction(ctx, ops, sn, sns, clusterUUID, apiClient) - default: - return r.failOps(ctx, ops, fmt.Sprintf("unknown action %q", ops.Spec.Action)) - } -} - -// runSimpleAction handles shutdown / restart / suspend / resume by posting to -// the backend and polling until the node reaches its terminal status. -func (r *StorageNodeOpsReconciler) runSimpleAction( - ctx context.Context, - ops *simplyblockv1alpha2.StorageNodeOps, - sn *simplyblockv1alpha1.StorageNode, - _ *simplyblockv1alpha1.StorageNodeSet, - clusterUUID string, - apiClient *webapi.Client, -) (ctrl.Result, error) { - log := logf.FromContext(ctx) - nodeUUID := sn.Status.UUID - action := ops.Spec.Action - - // POST the action if not yet triggered. - if !ops.Status.Triggered { - endpoint := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/%s", - clusterUUID, nodeUUID, action) - body := map[string]interface{}{} - if ops.Spec.Force != nil && *ops.Spec.Force { - body["force"] = true - } - if action == "restart" && ops.Spec.ReattachVolume != nil { - body["reattach_volume"] = *ops.Spec.ReattachVolume - } - _, status, err := apiClient.Do(ctx, http.MethodPost, endpoint, body) - if err != nil || status >= 300 { - if err == nil { - err = fmt.Errorf("status %d", status) - } - log.Error(err, "action POST failed", "action", action, "nodeUUID", nodeUUID) - return ctrl.Result{RequeueAfter: 10 * time.Second}, nil - } - patch := client.MergeFrom(ops.DeepCopy()) - ops.Status.Triggered = true - ops.Status.Message = fmt.Sprintf("%s request sent, waiting for node", action) - if err := r.Status().Patch(ctx, ops, patch); err != nil { - return ctrl.Result{}, err - } - return ctrl.Result{RequeueAfter: 5 * time.Second}, nil - } - - // Poll node status until terminal. - terminalStatus := map[simplyblockv1alpha2.StorageNodeOpsAction]string{ - simplyblockv1alpha2.StorageNodeOpsActionSuspend: utils.NodeStatusSuspended, - simplyblockv1alpha2.StorageNodeOpsActionResume: utils.NodeStatusOnline, - simplyblockv1alpha2.StorageNodeOpsActionRestart: utils.NodeStatusOnline, - simplyblockv1alpha2.StorageNodeOpsActionShutdown: "offline", - } - want := terminalStatus[action] - - currentStatus, err := getNodeBackendStatus(ctx, apiClient, clusterUUID, nodeUUID) - if err != nil { - log.Error(err, "failed to get node status during action poll") - return ctrl.Result{RequeueAfter: 10 * time.Second}, nil - } - - if currentStatus == want { - return r.succeedOps(ctx, ops, sn) - } - log.Info("waiting for node to reach terminal status", - "want", want, "current", currentStatus, "action", action) - return ctrl.Result{RequeueAfter: 5 * time.Second}, nil -} - -// runMigrate relocates a storage node onto a different worker host. A migration -// is not a drain: the node keeps its UUID and its partitions/logical-volume -// assignments follow it. It is a restart with a target node_address — the same -// primitive the node-drain coordinator uses to bring a node back after a reboot -// (see nodedrain_controller.go), but pointed at a different worker's -// storage-node-api pod instead of the current one. -// -// The target worker must already be part of the StorageNodeSet topology (labeled -// and running a storage-node-api pod), otherwise node_address is unreachable. -func (r *StorageNodeOpsReconciler) runMigrate( - ctx context.Context, - ops *simplyblockv1alpha2.StorageNodeOps, - sn *simplyblockv1alpha1.StorageNode, - sns *simplyblockv1alpha1.StorageNodeSet, - clusterUUID string, - apiClient *webapi.Client, -) (ctrl.Result, error) { - nodeUUID := sn.Status.UUID - target := ops.Spec.MigrateParams().TargetWorkerNode - - // Validate the request. - if target == "" { - return r.failOps(ctx, ops, "targetWorkerNode is required for action=migrate") - } - if target == sn.Spec.WorkerNode { - return r.failOps(ctx, ops, fmt.Sprintf("targetWorkerNode %q is the node's current worker", target)) - } - - // runMigrate is a four-phase state machine tracked via ops.Status.SubPhase: - // - // Preparing → Migrating → Restarting → Promoting - // - // Preparing readies the target (config, label, pod, DNS); Migrating issues - // the restart and waits for the node to enter in_restart; Restarting waits - // for it to return online; Promoting issues /promote and re-points topology. - // Each phase issues its one-shot control-plane call gated by - // ops.Status.Triggered, which advanceSubPhase resets to false on every - // transition, so a requeue within a phase never repeats the POST. - switch ops.Status.SubPhase { - case "": - // Enter the state machine. Persist Preparing first so the phase is - // observable before any preparation work begins. - return r.advanceSubPhase(ctx, ops, simplyblockv1alpha2.StorageNodeOpsSubPhasePreparing) - case simplyblockv1alpha2.StorageNodeOpsSubPhasePreparing: - return r.migratePrepare(ctx, ops, sn, sns, target, clusterUUID) - case simplyblockv1alpha2.StorageNodeOpsSubPhaseMigrating: - return r.migrateRestart(ctx, ops, sn, target, clusterUUID, nodeUUID, apiClient) - case simplyblockv1alpha2.StorageNodeOpsSubPhaseRestarting: - return r.migrateAwaitOnline(ctx, ops, target, clusterUUID, nodeUUID, apiClient) - case simplyblockv1alpha2.StorageNodeOpsSubPhasePromoting: - return r.migratePromote(ctx, ops, sn, sns, target, clusterUUID, nodeUUID, apiClient) - default: - return r.failOps(ctx, ops, fmt.Sprintf("migrate: unexpected sub-phase %q", ops.Status.SubPhase)) - } -} - -// migratePrepare runs the Preparing sub-phase: it clones the source worker's -// per-node config onto the target, labels the target into the storage plane so -// the DaemonSet schedules a storage-node-api pod there, then blocks until that -// pod is Ready AND the target's per-pod DNS name is published in the -// storage-node-api EndpointSlice. Only then can the control-plane restart -// resolve node_address, so the phase does not advance to Migrating until the -// DNS precondition holds — otherwise the restart fails name resolution and the -// control plane resets the node to OFFLINE. -func (r *StorageNodeOpsReconciler) migratePrepare( - ctx context.Context, - ops *simplyblockv1alpha2.StorageNodeOps, - sn *simplyblockv1alpha1.StorageNode, - sns *simplyblockv1alpha1.StorageNodeSet, - target string, - clusterUUID string, -) (ctrl.Result, error) { - log := logf.FromContext(ctx) - - var node corev1.Node - if err := r.Get(ctx, types.NamespacedName{Name: target}, &node); err != nil { - if apierrors.IsNotFound(err) { - return r.failOps(ctx, ops, fmt.Sprintf("target worker node %q not found in the cluster", target)) - } - return ctrl.Result{RequeueAfter: 10 * time.Second}, nil - } - if !isNodeReady(&node) { - return r.failOps(ctx, ops, fmt.Sprintf("target worker node %q is not Ready", target)) - } - - // Clone the source worker's per-node config onto the target before the - // storage-node pod is scheduled there, so it boots with the same effective - // configuration as the node being migrated. Any additional NVMe devices - // requested via spec.newSsdPcie are merged into the cloned PCI_ALLOWED so the - // target host binds them on start. Done before labeling so the entry exists - // by the time the pod's init container sources it. - if err := r.ensureMigratedWorkerConfig(ctx, sns, sn.Spec.WorkerNode, target, ops.Spec.MigrateParams().NewSsdPcie); err != nil { - log.Error(err, "migrate: failed to clone per-node config to target worker", - "source", sn.Spec.WorkerNode, "target", target) - return ctrl.Result{RequeueAfter: 10 * time.Second}, nil - } - - // Label the target so the storage-node DaemonSet schedules a storage-node-api - // pod there and the StorageNodeSet reconcile publishes its per-pod DNS name - // in the EndpointSlice. Reuse the canonical StorageNodeSet labeler so the - // migration target receives the exact DaemonSet node-selector label - // (io.simplyblock.storagenodeset) — the target is not yet in - // sns.Spec.WorkerNodes (that swap happens in the Promoting phase), so it is - // injected via extraWorkers. - if node.Labels[kube.LabelStorageNodeSet] != sns.Name { - if err := labelWorkerNodes(ctx, r.Client, r.Recorder, sns, clusterUUID, target); err != nil { - log.Error(err, "migrate: failed to label target worker", "worker", target) - return ctrl.Result{RequeueAfter: 10 * time.Second}, nil - } - r.Recorder.Eventf(ops, nil, corev1.EventTypeNormal, "TargetWorkerLabeled", "TargetWorkerLabeled", - "labeled worker %s for storage plane of cluster %s", target, sns.Spec.ClusterName) - r.emitOnStorageNode(ctx, ops, corev1.EventTypeNormal, "TargetWorkerLabeled", - fmt.Sprintf("labeled worker %s for storage plane of cluster %s", target, sns.Spec.ClusterName)) - } - - // Wait for the storage-node-api pod to be Running+Ready on the target. - ready, err := r.storageNodePodReady(ctx, sns.Namespace, sns.Name, target) - if err != nil { - return ctrl.Result{RequeueAfter: 10 * time.Second}, nil - } - if !ready { - return r.migrateWaiting(ctx, ops, fmt.Sprintf("waiting for storage-node pod on worker %s", target)) - } - - // Gate: the control plane resolves node_address via the target's per-pod - // headless DNS name, which only exists once the target is published in the - // storage-node-api EndpointSlice (built from labeled storage-plane nodes). - // Hold in Preparing until the entry appears so the restart can resolve it. - inSlice, err := r.endpointSliceHasWorker(ctx, sns.Namespace, sns.Name, target) - if err != nil { - log.Error(err, "migrate: failed to read storage-node-api EndpointSlice", "target", target) - return ctrl.Result{RequeueAfter: 10 * time.Second}, nil - } - if !inSlice { - return r.migrateWaiting(ctx, ops, - fmt.Sprintf("waiting for worker %s DNS to be published before restart", target)) - } - - return r.advanceSubPhase(ctx, ops, simplyblockv1alpha2.StorageNodeOpsSubPhaseMigrating) -} - -// migrateRestart runs the Migrating sub-phase: it issues the control-plane -// restart pointed at the target worker's node_address (once, gated by -// Triggered), then waits until the node is observed to LEAVE online (enter -// in_restart). Because /restart is asynchronous — the POST returns immediately -// while the old primary keeps reporting online — a plain "status == online" -// check would pass on the pre-restart status and let /promote fire into the -// in-flight restart. Confirming the node entered in_restart here means the -// Restarting phase's later "back to online" is the genuine post-restart state. -func (r *StorageNodeOpsReconciler) migrateRestart( - ctx context.Context, - ops *simplyblockv1alpha2.StorageNodeOps, - sn *simplyblockv1alpha1.StorageNode, - target, clusterUUID, nodeUUID string, - apiClient *webapi.Client, -) (ctrl.Result, error) { - log := logf.FromContext(ctx) - - if !ops.Status.Triggered { - payload := map[string]any{ - // Migration relocates a still-online node, so the control-plane - // restart must run with force=true — a non-forced restart is rejected - // unless the node is already OFFLINE ("Node must be offline"). Default - // to true and honor an explicit spec.force override only when set. - "force": ops.Spec.Force == nil || *ops.Spec.Force, - "node_address": utils.StorageNodeSetAPIAddress(target, sn.Namespace), - } - if ops.Spec.ReattachVolume != nil { - payload["reattach_volume"] = *ops.Spec.ReattachVolume - } - if len(ops.Spec.MigrateParams().NewSsdPcie) > 0 { - payload["new_ssd_pcie"] = ops.Spec.MigrateParams().NewSsdPcie - } - endpoint := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/restart", clusterUUID, nodeUUID) - respBody, status, err := apiClient.Do(ctx, http.MethodPost, endpoint, payload) - if err != nil || status >= 300 { - if err == nil { - err = fmt.Errorf("status %d: %s", status, string(respBody)) - } - log.Error(err, "migrate: restart POST failed", "nodeUUID", nodeUUID, "target", target) - return ctrl.Result{RequeueAfter: 10 * time.Second}, nil - } - patch := client.MergeFrom(ops.DeepCopy()) - ops.Status.Triggered = true - ops.Status.Message = fmt.Sprintf("migrating node %s to worker %s, waiting for restart to begin", nodeUUID, target) - if err := r.Status().Patch(ctx, ops, patch); err != nil { - return ctrl.Result{}, err - } - r.Recorder.Eventf(ops, nil, corev1.EventTypeNormal, "MigrateStarted", "MigrateStarted", - "restart with target worker %s issued for node %s", target, nodeUUID) - r.emitOnStorageNode(ctx, ops, corev1.EventTypeNormal, "MigrateStarted", - fmt.Sprintf("restart with target worker %s issued for node %s", target, nodeUUID)) - // Poll quickly so the (typically tens-of-seconds) in_restart window is - // reliably observed before the node returns online. - return ctrl.Result{RequeueAfter: migrateRestartPoll}, nil - } - - currentStatus, err := getNodeBackendStatus(ctx, apiClient, clusterUUID, nodeUUID) - if err != nil { - log.Error(err, "migrate: failed to get node status during restart-start poll") - return ctrl.Result{RequeueAfter: migrateRestartPoll}, nil - } - if currentStatus == utils.NodeStatusOnline { - // Still the pre-restart online — the async restart has not taken the - // node down yet. Keep waiting; do NOT advance, or /promote would race - // the in-flight restart. - return r.migrateWaitingAfter(ctx, ops, migrateRestartPoll, - fmt.Sprintf("waiting for restart of node %s to begin on worker %s", nodeUUID, target)) - } - - // Node has left online (in_restart / offline) — the restart is underway. - return r.advanceSubPhase(ctx, ops, simplyblockv1alpha2.StorageNodeOpsSubPhaseRestarting) -} - -// migrateAwaitOnline runs the Restarting sub-phase: the restart is confirmed -// in-flight (the node left online in Migrating), so it waits for the node to -// return to online on the target host — the genuine post-restart state — and -// only then advances to Promoting. -func (r *StorageNodeOpsReconciler) migrateAwaitOnline( - ctx context.Context, - ops *simplyblockv1alpha2.StorageNodeOps, - target, clusterUUID, nodeUUID string, - apiClient *webapi.Client, -) (ctrl.Result, error) { - log := logf.FromContext(ctx) - - currentStatus, err := getNodeBackendStatus(ctx, apiClient, clusterUUID, nodeUUID) - if err != nil { - log.Error(err, "migrate: failed to get node status during online poll") - return ctrl.Result{RequeueAfter: 10 * time.Second}, nil - } - if currentStatus != utils.NodeStatusOnline { - return r.migrateWaiting(ctx, ops, - fmt.Sprintf("waiting for node %s to come online on worker %s (status %s)", nodeUUID, target, currentStatus)) - } - - return r.advanceSubPhase(ctx, ops, simplyblockv1alpha2.StorageNodeOpsSubPhasePromoting) -} - -// migratePromote runs the Promoting sub-phase: it issues the control-plane -// /promote for the relocated node (once, gated by Triggered), then re-points the -// Kubernetes topology onto the target worker and completes the op. /promote -// activates the new host's devices, fails and migrates the origin host's devices -// (starting a rebalance), sets the primary, and re-homes the logical volumes -// onto the relocated node. -func (r *StorageNodeOpsReconciler) migratePromote( - ctx context.Context, - ops *simplyblockv1alpha2.StorageNodeOps, - sn *simplyblockv1alpha1.StorageNode, - sns *simplyblockv1alpha1.StorageNodeSet, - target, clusterUUID, nodeUUID string, - apiClient *webapi.Client, -) (ctrl.Result, error) { - log := logf.FromContext(ctx) - - if !ops.Status.Triggered { - endpoint := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/promote", clusterUUID, nodeUUID) - respBody, status, err := apiClient.Do(ctx, http.MethodPost, endpoint, nil) - if err != nil || status >= 300 { - if err == nil { - err = fmt.Errorf("status %d: %s", status, string(respBody)) - } - log.Error(err, "migrate: promote POST failed", "nodeUUID", nodeUUID, "target", target) - return ctrl.Result{RequeueAfter: 10 * time.Second}, nil - } - patch := client.MergeFrom(ops.DeepCopy()) - ops.Status.Triggered = true - ops.Status.Message = fmt.Sprintf("promoted node %s on worker %s; rebalance started", nodeUUID, target) - if err := r.Status().Patch(ctx, ops, patch); err != nil { - return ctrl.Result{}, err - } - log.Info("migrate: promoted relocated node; rebalance started", "nodeUUID", nodeUUID, "target", target) - r.Recorder.Eventf(ops, nil, corev1.EventTypeNormal, "MigratePromoted", "MigratePromoted", - "promote issued for node %s on worker %s; rebalance started", nodeUUID, target) - r.emitOnStorageNode(ctx, ops, corev1.EventTypeNormal, "MigratePromoted", - fmt.Sprintf("promote issued for node %s on worker %s; rebalance started", nodeUUID, target)) - return ctrl.Result{RequeueAfter: 5 * time.Second}, nil - } - - // Promote issued — re-point the Kubernetes topology: update this StorageNode's - // spec.workerNode and swap the owning StorageNodeSet's worker list (and status) - // from the source to the target worker. - if err := r.reconcileMigratedTopology(ctx, sn, sns, target, ops.Spec.MigrateParams().NewSsdPcie); err != nil { - log.Error(err, "migrate: failed to reconcile topology after migration", "target", target) - return ctrl.Result{RequeueAfter: 10 * time.Second}, nil - } - - r.Recorder.Eventf(ops, nil, corev1.EventTypeNormal, "MigrateCompleted", "MigrateCompleted", - "node %s is online on worker %s", nodeUUID, target) - r.emitOnStorageNode(ctx, ops, corev1.EventTypeNormal, "MigrateCompleted", - fmt.Sprintf("node %s is online on worker %s", nodeUUID, target)) - return r.succeedOps(ctx, ops, sn) -} - -// migrateRestartPoll is the requeue interval while waiting for the asynchronous -// restart to be observed entering in_restart. It is short so the (typically -// tens-of-seconds) in_restart window is reliably caught before the node returns -// online — otherwise the op could miss it and stall in Migrating. -const migrateRestartPoll = 3 * time.Second - -// migrateWaiting patches the op's status message and requeues (after 10s) -// without changing the sub-phase — used while a migrate precondition is pending. -func (r *StorageNodeOpsReconciler) migrateWaiting( - ctx context.Context, - ops *simplyblockv1alpha2.StorageNodeOps, - msg string, -) (ctrl.Result, error) { - return r.migrateWaitingAfter(ctx, ops, 10*time.Second, msg) -} - -// migrateWaitingAfter is migrateWaiting with a caller-chosen requeue interval. -func (r *StorageNodeOpsReconciler) migrateWaitingAfter( - ctx context.Context, - ops *simplyblockv1alpha2.StorageNodeOps, - after time.Duration, - msg string, -) (ctrl.Result, error) { - logf.FromContext(ctx).Info("migrate: " + msg) - patch := client.MergeFrom(ops.DeepCopy()) - ops.Status.Message = msg - if err := r.Status().Patch(ctx, ops, patch); err != nil { - // Surface the failure so the reconciler backs off and retries rather - // than silently spinning with a stale status message. - return ctrl.Result{}, fmt.Errorf("migrate: updating status message %q: %w", msg, err) - } - return ctrl.Result{RequeueAfter: after}, nil -} - -// endpointSliceHasWorker reports whether the storage-node-api EndpointSlice -// publishes the given worker's per-pod DNS hostname with at least one address — -// i.e. whether .simplyblock-storage-node-api..svc resolves. -func (r *StorageNodeOpsReconciler) endpointSliceHasWorker( - ctx context.Context, - namespace, storageNodeSetName, worker string, -) (bool, error) { - log := logf.FromContext(ctx) - var eps discoveryv1.EndpointSlice - // Uncached read via the APIReader: this is a liveness gate — if it reads a - // stale slice and misses the target's freshly-published endpoint, the - // migration wedges in Preparing on "waiting for DNS" with no error. Reading - // straight from the API server removes any dependence on informer freshness. - if err := r.apiReader.Get(ctx, types.NamespacedName{ - Name: kube.StorageNodeSetAPIEndpointSliceName(storageNodeSetName), - Namespace: namespace, - }, &eps); err != nil { - if apierrors.IsNotFound(err) { - return false, nil - } - return false, err - } - want := utils.NodeHostnameLabel(worker) - for i := range eps.Endpoints { - e := eps.Endpoints[i] - if e.Hostname != nil && *e.Hostname == want && len(e.Addresses) > 0 { - return true, nil - } - } - // Diagnostic: dump exactly what the read returned so a recurrence tells us - // whether the worker was absent (slice not yet published) or present but - // unmatched (hostname/address mismatch) — the two have different root causes. - seen := make([]string, 0, len(eps.Endpoints)) - for i := range eps.Endpoints { - h := "" - if eps.Endpoints[i].Hostname != nil { - h = *eps.Endpoints[i].Hostname - } - seen = append(seen, fmt.Sprintf("%s=%v", h, eps.Endpoints[i].Addresses)) - } - log.Info("migrate: target worker not present in storage-node-api EndpointSlice", - "want", want, "sliceResourceVersion", eps.ResourceVersion, "endpoints", seen) - return false, nil -} - -// reconcileMigratedTopology re-points the operator's Kubernetes topology at the -// target worker after the backend node has moved. It updates this StorageNode's -// spec.workerNode in place (permitted for the operator service account; blocked -// for users by the StorageNode validating webhook) and swaps the owning -// StorageNodeSet.spec.workerNodes from the source worker to the target. It also -// migrates the per-node config source of truth: spec.nodeConfigs[source] is -// cloned onto spec.nodeConfigs[target] with newSsdPcie merged into its -// PcieAllowList, and the source entry is dropped. Persisting this keeps the -// StorageNodeSet reconciler from rebuilding the target's ConfigMap entry back to -// fleet defaults (which would drop the newly bound devices) and removes the -// stale source-host config. -func (r *StorageNodeOpsReconciler) reconcileMigratedTopology( - ctx context.Context, - sn *simplyblockv1alpha1.StorageNode, - sns *simplyblockv1alpha1.StorageNodeSet, - target string, - newSsdPcie []string, -) error { - source := sn.Spec.WorkerNode - - // 1. Re-point the StorageNode CR at the target worker. The UUID and status - // are preserved — the same backend node simply runs on a different host. - // The name no longer encodes the worker, so the worker label is the - // human-facing indicator and is refreshed alongside the spec. - if sn.Spec.WorkerNode != target { - patch := client.MergeFrom(sn.DeepCopy()) - sn.Spec.WorkerNode = target - if sn.Labels == nil { - sn.Labels = map[string]string{} - } - sn.Labels["storage.simplyblock.io/worker"] = sanitiseDNSLabel(target) - if err := r.Patch(ctx, sn, patch); err != nil { - return fmt.Errorf("re-pointing StorageNode %s to worker %s: %w", sn.Name, target, err) - } - } - - // 2. Reconcile the StorageNodeSet worker list and per-node config: drop the - // source worker, add the target, and move nodeConfigs[source] to - // nodeConfigs[target] (with newSsdPcie merged in). Re-fetch to patch - // against the latest version. - var fresh simplyblockv1alpha1.StorageNodeSet - if err := r.Get(ctx, types.NamespacedName{Name: sns.Name, Namespace: sns.Namespace}, &fresh); err != nil { - return err - } - - // New worker list: source removed, target present. - workers := make([]string, 0, len(fresh.Spec.WorkerNodes)+1) - hasTarget := false - workersChanged := false - for _, w := range fresh.Spec.WorkerNodes { - if w == source { - workersChanged = true - continue - } - if w == target { - hasTarget = true - } - workers = append(workers, w) - } - if !hasTarget { - workers = append(workers, target) - workersChanged = true - } - - // Migrated per-node config for the target. Prefer an existing target entry - // (idempotent re-runs), otherwise clone the source's overrides. The effective - // PcieAllowList is the entry's own list, or the fleet default when unset; - // newSsdPcie is merged into it so the added devices persist across rebuilds. - targetCfg, hadTargetCfg := fresh.Spec.NodeConfigs[target] - _, hadSourceCfg := fresh.Spec.NodeConfigs[source] - if !hadTargetCfg { - targetCfg = fresh.Spec.NodeConfigs[source] // zero value if source has none - } - if len(newSsdPcie) > 0 { - effPcie := targetCfg.PcieAllowList - if len(effPcie) == 0 { - effPcie = fresh.Spec.PcieAllowList - } - targetCfg.PcieAllowList = mergePcieList(effPcie, newSsdPcie) - } - setTarget := hadTargetCfg || hadSourceCfg || len(newSsdPcie) > 0 - - if !workersChanged && !hadSourceCfg && !setTarget { - // Spec already reconciled (idempotent re-run); still prune the stale - // source-host status entry left behind by the move. - if err := r.pruneMigratedSourceStatus(ctx, fresh.Name, fresh.Namespace, source, sn.Status.UUID); err != nil { - return err - } - r.removeSourceWorkerLabels(ctx, source, sn.Status.UUID, fresh.Namespace) - return nil - } - - patch := client.MergeFrom(fresh.DeepCopy()) - fresh.Spec.WorkerNodes = workers - if hadSourceCfg { - delete(fresh.Spec.NodeConfigs, source) - } - if setTarget { - if fresh.Spec.NodeConfigs == nil { - fresh.Spec.NodeConfigs = map[string]simplyblockv1alpha1.StorageNodeOverrides{} - } - fresh.Spec.NodeConfigs[target] = targetCfg - } - if err := r.Patch(ctx, &fresh, patch); err != nil { - return fmt.Errorf("reconciling StorageNodeSet %s topology: %w", fresh.Name, err) - } - - // 3. Prune the stale source-host entry from status. The status sync keys - // Status.Nodes by hostname, so once the migrated node reports from the - // target it is appended as a new entry while the source entry (same - // backend UUID) lingers. Drop it so the set reflects only live hosts. - if err := r.pruneMigratedSourceStatus(ctx, fresh.Name, fresh.Namespace, source, sn.Status.UUID); err != nil { - return err - } - - // 4. Remove storage-plane labels from the source K8s Node now that no - // storage node runs there. Best-effort: the migration itself has already - // succeeded at this point. - r.removeSourceWorkerLabels(ctx, source, sn.Status.UUID, fresh.Namespace) - return nil -} - -// pruneMigratedSourceStatus removes the StorageNodeSet.Status.Nodes entry left -// on the source host after a node migrated away — matched by the source -// hostname and the migrated node's backend UUID so sibling nodes still on that -// host (multi-socket) and unrelated in-flight (UUID=="") entries are untouched. -func (r *StorageNodeOpsReconciler) pruneMigratedSourceStatus( - ctx context.Context, - name, namespace, source, uuid string, -) error { - if uuid == "" { - return nil // cannot match safely without the migrated node's UUID - } - var sns simplyblockv1alpha1.StorageNodeSet - if err := r.Get(ctx, types.NamespacedName{Name: name, Namespace: namespace}, &sns); err != nil { - return err - } - kept := make([]simplyblockv1alpha1.NodeStatus, 0, len(sns.Status.Nodes)) - removed := false - for _, n := range sns.Status.Nodes { - if n.Hostname == source && n.UUID == uuid { - removed = true - continue - } - kept = append(kept, n) - } - if !removed { - return nil - } - patch := client.MergeFrom(sns.DeepCopy()) - sns.Status.Nodes = kept - if err := r.Status().Patch(ctx, &sns, patch); err != nil { - return fmt.Errorf("pruning migrated source status entry for %s on %s: %w", uuid, source, err) - } - return nil -} - -// removeSourceWorkerLabels removes storage-plane K8s Node labels from the source -// worker after a successful migration. It removes: -// - the UUID slot label whose value matches migratedUUID (identifies this node's slot) -// - io.simplyblock.storagenodeset if no other StorageNode CRs in the namespace -// still target this worker (multi-socket guard) -// -// Best-effort: errors are logged but do not block the migration result. -func (r *StorageNodeOpsReconciler) removeSourceWorkerLabels( - ctx context.Context, - source, migratedUUID, namespace string, -) { - log := logf.FromContext(ctx) - - var snList simplyblockv1alpha1.StorageNodeList - if err := r.List(ctx, &snList, client.InNamespace(namespace)); err != nil { - log.Error(err, "removeSourceWorkerLabels: failed to list StorageNodes", "source", source) - return - } - hasOtherSNs := false - for _, sn := range snList.Items { - if sn.Spec.WorkerNode == source { - hasOtherSNs = true - break - } - } - - var node corev1.Node - if err := r.Get(ctx, client.ObjectKey{Name: source}, &node); err != nil { - if !apierrors.IsNotFound(err) { - log.Error(err, "removeSourceWorkerLabels: failed to get k8s Node", "source", source) - } - return - } - - patch := client.MergeFrom(node.DeepCopy()) - changed := false - - // Remove the UUID slot label that identifies this storage node's slot. - for k, v := range node.Labels { - if strings.HasPrefix(k, storageNodeUUIDLabelPrefix) && v == migratedUUID { - delete(node.Labels, k) - changed = true - } - } - - // Remove cluster-level labels only when no sibling storage nodes remain on this host. - if !hasOtherSNs { - if _, ok := node.Labels[kube.LabelStorageNodeSet]; ok { - delete(node.Labels, kube.LabelStorageNodeSet) - changed = true - } - } - - if !changed { - return - } - if err := r.Patch(ctx, &node, patch); err != nil { - log.Error(err, "removeSourceWorkerLabels: failed to patch k8s Node", "source", source) - return - } - log.Info("Removed storage-plane labels from migrated source node", "node", source) -} - -// ensureMigratedWorkerConfig clones the source worker's entry in the per-node -// ConfigMap onto the target worker so the storage-node pod scheduled on the -// target boots with the same effective configuration. When newSsdPcie is -// non-empty those PCIe addresses are merged into the cloned entry's PCI_ALLOWED. -// -// It writes the ConfigMap directly because the target is not yet in -// spec.workerNodes, so the StorageNodeSet reconciler would not build an entry -// for it. reconcileMigratedTopology later persists the durable source of truth -// into spec.nodeConfigs[target]. Idempotent: an existing target entry is left -// untouched. -func (r *StorageNodeOpsReconciler) ensureMigratedWorkerConfig( - ctx context.Context, - sns *simplyblockv1alpha1.StorageNodeSet, - source, target string, - newSsdPcie []string, -) error { - name := PerNodeConfigMapName(sns.Name) - var cm corev1.ConfigMap - if err := r.Get(ctx, types.NamespacedName{Name: name, Namespace: sns.Namespace}, &cm); err != nil { - return fmt.Errorf("getting per-node ConfigMap %s: %w", name, err) - } - if _, ok := cm.Data[target]; ok { - return nil // already cloned - } - srcEntry, ok := cm.Data[source] - if !ok { - return fmt.Errorf("per-node ConfigMap %s has no entry for source worker %q", name, source) - } - - patch := client.MergeFrom(cm.DeepCopy()) - if cm.Data == nil { - cm.Data = map[string]string{} - } - cm.Data[target] = mergePcieAllowedIntoEnvFile(srcEntry, newSsdPcie) - if err := r.Patch(ctx, &cm, patch); err != nil { - return fmt.Errorf("cloning per-node config %q -> %q: %w", source, target, err) - } - return nil -} - -// mergePcieAllowedIntoEnvFile returns the per-node env-file text with extra PCIe -// addresses merged into its PCI_ALLOWED= line (deduplicated, order preserved, -// shell-quoted to match buildPerNodeEnvFile). The input is returned unchanged -// when extra is empty. A PCI_ALLOWED line is appended if none is present. -func mergePcieAllowedIntoEnvFile(envFile string, extra []string) string { - if len(extra) == 0 { - return envFile - } - const prefix = "PCI_ALLOWED=" - lines := strings.Split(envFile, "\n") - for i, line := range lines { - if !strings.HasPrefix(line, prefix) { - continue - } - current := parseShellCSV(strings.TrimPrefix(line, prefix)) - lines[i] = prefix + utils.ShellQuote(strings.Join(mergePcieList(current, extra), ",")) - return strings.Join(lines, "\n") - } - // No PCI_ALLOWED line: append one, preserving a single trailing newline. - trimmed := strings.TrimRight(envFile, "\n") - return trimmed + "\n" + prefix + utils.ShellQuote(strings.Join(mergePcieList(nil, extra), ",")) + "\n" -} - -// parseShellCSV parses a shell-quoted, comma-separated value (as produced by -// utils.ShellQuote) into its elements, dropping empties. -func parseShellCSV(v string) []string { - v = strings.TrimSpace(v) - if len(v) >= 2 && strings.HasPrefix(v, "'") && strings.HasSuffix(v, "'") { - v = v[1 : len(v)-1] - v = strings.ReplaceAll(v, `'\''`, "'") // undo ShellQuote escaping - } - var out []string - for _, p := range strings.Split(v, ",") { - if p = strings.TrimSpace(p); p != "" { - out = append(out, p) - } - } - return out -} - -// mergePcieList concatenates base and extra, dropping empties and duplicates -// while preserving first-seen order. -func mergePcieList(base, extra []string) []string { - seen := make(map[string]struct{}, len(base)+len(extra)) - out := make([]string, 0, len(base)+len(extra)) - for _, group := range [][]string{base, extra} { - for _, s := range group { - if s == "" { - continue - } - if _, ok := seen[s]; ok { - continue - } - seen[s] = struct{}{} - out = append(out, s) - } - } - return out -} - -// storageNodePodReady reports whether the storage-node DaemonSet pod on the given -// worker is Running and Ready, which is the precondition for the control plane to -// reach that host's storage-node-api at node_address. It selects on the per-set -// label so it inspects only the target StorageNodeSet's pods, not every set in -// the cluster. -func (r *StorageNodeOpsReconciler) storageNodePodReady( - ctx context.Context, - namespace, storageNodeSetName, workerName string, -) (bool, error) { - var pods corev1.PodList - if err := r.List(ctx, &pods, - client.InNamespace(namespace), - client.MatchingLabels{kube.LabelApp: kube.AppStorageNode, kube.LabelStorageNodeSet: storageNodeSetName}, - ); err != nil { - return false, err - } - for i := range pods.Items { - p := &pods.Items[i] - if p.Spec.NodeName != workerName || p.Status.Phase != corev1.PodRunning { - continue - } - for _, c := range p.Status.Conditions { - if c.Type == corev1.PodReady && c.Status == corev1.ConditionTrue { - return true, nil - } - } - } - return false, nil -} - -// ───────────────────────────────────────────────────────────────────────────── -// Drain state machine (action=remove) -// Phases: Validating → Suspending → Migrating → Verifying → Removing -// ───────────────────────────────────────────────────────────────────────────── - -func (r *StorageNodeOpsReconciler) runDrain( - ctx context.Context, - ops *simplyblockv1alpha2.StorageNodeOps, - sn *simplyblockv1alpha1.StorageNode, - clusterUUID string, - apiClient *webapi.Client, -) (ctrl.Result, error) { - switch ops.Status.SubPhase { - case simplyblockv1alpha2.StorageNodeOpsSubPhaseValidating: - return r.drainValidate(ctx, ops, sn, clusterUUID, apiClient) - case simplyblockv1alpha2.StorageNodeOpsSubPhaseSuspending: - return r.drainSuspend(ctx, ops, sn, clusterUUID, apiClient) - case simplyblockv1alpha2.StorageNodeOpsSubPhaseMigrating: - return r.drainMigrate(ctx, ops, sn, clusterUUID, apiClient) - case simplyblockv1alpha2.StorageNodeOpsSubPhaseVerifying: - return r.drainVerify(ctx, ops, sn, clusterUUID, apiClient) - case simplyblockv1alpha2.StorageNodeOpsSubPhaseRemoving: - return r.drainRemove(ctx, ops, sn, clusterUUID, apiClient) - default: - return r.failOps(ctx, ops, fmt.Sprintf("unknown drain sub-phase %q", ops.Status.SubPhase)) - } -} - -func (r *StorageNodeOpsReconciler) drainValidate( - ctx context.Context, - ops *simplyblockv1alpha2.StorageNodeOps, - sn *simplyblockv1alpha1.StorageNode, - clusterUUID string, - apiClient *webapi.Client, -) (ctrl.Result, error) { - log := logf.FromContext(ctx) - nodeUUID := sn.Status.UUID - - volumes, err := listNodeVolumes(ctx, apiClient, clusterUUID, nodeUUID) - if err != nil { - log.Error(err, "drain: failed to list volumes during validation") - return ctrl.Result{RequeueAfter: drainRequeueImmediate}, nil - } - - sysFilter, err := r.resolveOpsSystemVolumeFilter(ops) - if err != nil { - return r.failOps(ctx, ops, "invalid systemVolumeFilterRegex: "+err.Error()) - } - - _, pinned, unmanaged, _, _, err := matchVolumesToPVs(ctx, r.Client, volumes, sysFilter) - if err != nil { - log.Error(err, "drain: matchVolumesToPVs failed during validation") - return ctrl.Result{RequeueAfter: drainRequeueImmediate}, nil - } - - if len(pinned) > 0 { - r.Recorder.Eventf(ops, nil, corev1.EventTypeWarning, "PinnedVolumeBlocking", "PinnedVolumeBlocking", - "drain blocked: %d pinned volume(s) on node %s — remove the %s annotation to proceed", - len(pinned), nodeUUID, kube.AnnoSelectedStorageNode) - r.emitOnStorageNode(ctx, ops, corev1.EventTypeWarning, "PinnedVolumeBlocking", fmt.Sprintf("drain blocked: %d pinned volume(s) on node %s — remove the %s annotation to proceed", len(pinned), nodeUUID, kube.AnnoSelectedStorageNode)) - patch := client.MergeFrom(ops.DeepCopy()) - ops.Status.Message = fmt.Sprintf("blocked: %d pinned volume(s) — remove %s annotation", len(pinned), kube.AnnoSelectedStorageNode) - _ = r.Status().Patch(ctx, ops, patch) - return ctrl.Result{RequeueAfter: 60 * time.Second}, nil - } - - if len(unmanaged) > 0 { - r.Recorder.Eventf(ops, nil, corev1.EventTypeWarning, "UnmanagedVolumeBlocking", "UnmanagedVolumeBlocking", - "drain blocked: %d unmanaged volume(s) on node %s — remove them manually", - len(unmanaged), nodeUUID) - r.emitOnStorageNode(ctx, ops, corev1.EventTypeWarning, "UnmanagedVolumeBlocking", fmt.Sprintf("drain blocked: %d unmanaged volume(s) on node %s — remove them manually", len(unmanaged), nodeUUID)) - patch := client.MergeFrom(ops.DeepCopy()) - ops.Status.Message = fmt.Sprintf("blocked: %d unmanaged volume(s) — remove manually", len(unmanaged)) - _ = r.Status().Patch(ctx, ops, patch) - return ctrl.Result{RequeueAfter: 60 * time.Second}, nil - } - - // Failure-domain balance gate: the backend's own admission check - // (check_fd_admission_for_remove) will refuse the DELETE call later in - // this drain if removing this node would violate the +/-1 balance rule - // -- but by then the node has already been suspended (Suspending runs - // before Removing), and that suspension has no path back on its own. - // Checking it here, before Suspending, means an infeasible removal - // never suspends the node in the first place. - // - // Fails outright rather than blocking-and-requeuing like the - // pinned/unmanaged-volume checks above: those resolve by acting ON THIS - // NODE (drop the annotation, delete the volume), so the same ops can - // just notice and proceed. Restoring failure-domain balance never does - // -- it needs a deliberate cluster-wide change (add a host, or remove a - // different node instead) that this ops has no way to detect on its - // own, so silently polling every 60s would leave a permanently-stuck - // Running ops easy to miss in `kubectl get storagenodeops`. Failing is - // safe here specifically because handleDeletion (storagenode_controller.go) - // now refuses to remove the StorageNode's finalizer while its remove - // ops is Failed -- the CR stays, nothing gets orphaned, and a human - // must either fix the imbalance and delete this ops to retry, or - // restore the worker to spec.workerNodes to keep it. - reason, err := r.fdRemovalBalanceCheck(ctx, sn) - if err != nil { - log.Error(err, "drain: failed to check failure-domain balance for removal") - return ctrl.Result{RequeueAfter: drainRequeueImmediate}, nil - } - if reason != "" { - return r.failOps(ctx, ops, fmt.Sprintf( - "removing node %s would violate failure-domain balance: %s", nodeUUID, reason)) - } - - return r.advanceSubPhase(ctx, ops, simplyblockv1alpha2.StorageNodeOpsSubPhaseSuspending) -} - -// fdRemovalBalanceCheck reports whether removing sn would violate the -// cluster's failure-domain balance rule, mirroring the backend's -// check_fd_admission_for_remove (simplyblock_core), including its very -// first early-out: a no-op when the cluster doesn't have failure domains -// enabled at all. Re-fetches the parent StorageNodeSet (and StorageCluster) -// rather than threading them through runDrain's whole dispatch chain -- -// Validating is the only sub-phase that needs them. Returns ("", nil) when -// removal is fine (including when FD data isn't populated yet, same as the -// backend's own early-outs); a non-empty reason means drainValidate must -// fail rather than advance to Suspending. -func (r *StorageNodeOpsReconciler) fdRemovalBalanceCheck( - ctx context.Context, sn *simplyblockv1alpha1.StorageNode, -) (string, error) { - var sns simplyblockv1alpha1.StorageNodeSet - if err := r.Get(ctx, types.NamespacedName{ - Name: sn.Spec.StorageNodeSetRef, - Namespace: sn.Namespace, - }, &sns); err != nil { - return "", fmt.Errorf("fetching StorageNodeSet %s: %w", sn.Spec.StorageNodeSetRef, err) - } - var cluster simplyblockv1alpha2.StorageCluster - if err := r.Get(ctx, types.NamespacedName{ - Name: sns.Spec.ClusterName, - Namespace: sn.Namespace, - }, &cluster); err != nil { - return "", fmt.Errorf("fetching StorageCluster %s: %w", sns.Spec.ClusterName, err) - } - if !ptr.BoolFromOrFalse(cluster.Spec.EnableFailureDomains) { - return "", nil - } - hostDomains, err := clustercontroller.FailureDomainHosts(ctx, r.Client, sn.Namespace, sns.Spec.ClusterName) - if err != nil { - return "", fmt.Errorf("computing failure-domain host map: %w", err) - } - // Counts are built from the PRE-removal host map so every domain that - // currently has a host stays represented -- including at zero, once - // decremented below -- rather than being derived from a post-removal - // hostDomains that has the affected host's key deleted outright (which - // would drop the domain from the map entirely once it had no - // surviving host, and hide that from fdRemovalBalanceViolation). - counts := map[int32]int{} - for _, fd := range hostDomains { - counts[fd]++ - } - if sn.Status.Ports != nil { - if removedDomain, ok := hostDomains[sn.Status.Ports.Management]; ok && - !clustercontroller.HostHasSurvivingSibling(ctx, r.Client, sn.Namespace, sns.Spec.ClusterName, sn.Status.Ports.Management, sn.Status.UUID) { - // Multi-node hosts (spec.socketsToUse / spec.nodesPerSocket > 1) - // run more than one StorageNode per physical host, all sharing - // this management IP -- only decrement the domain's count when - // no sibling StorageNode at this IP survives the removal, so - // removing one node on a shared host doesn't make the whole - // host (and a sibling still running on it) vanish from the count. - counts[removedDomain]-- - } - } - return clustercontroller.RemovalBalanceViolation(counts), nil -} - -func (r *StorageNodeOpsReconciler) drainSuspend( - ctx context.Context, - ops *simplyblockv1alpha2.StorageNodeOps, - sn *simplyblockv1alpha1.StorageNode, - clusterUUID string, - apiClient *webapi.Client, -) (ctrl.Result, error) { - log := logf.FromContext(ctx) - nodeUUID := sn.Status.UUID - - if !ops.Status.Triggered { - currentStatus, err := getNodeBackendStatus(ctx, apiClient, clusterUUID, nodeUUID) - if err != nil { - log.Error(err, "drain: could not read node status before suspend, retrying") - return ctrl.Result{RequeueAfter: drainRequeueSuspend}, nil - } - if currentStatus == utils.NodeStatusSuspended { - log.Info("drain: node already suspended, advancing without POST") - patch := client.MergeFrom(ops.DeepCopy()) - ops.Status.Triggered = true - ops.Status.Message = "node already suspended" - _ = r.Status().Patch(ctx, ops, patch) - return ctrl.Result{RequeueAfter: drainRequeueImmediate}, nil - } - - endpoint := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/suspend", clusterUUID, nodeUUID) - _, status, err := apiClient.Do(ctx, http.MethodPost, endpoint, nil) - if err != nil || status >= 300 { - if err == nil { - err = fmt.Errorf("suspend API returned status %d", status) - } - log.Error(err, "drain: suspend POST failed") - return ctrl.Result{RequeueAfter: drainRequeueSuspend}, nil - } - patch := client.MergeFrom(ops.DeepCopy()) - ops.Status.Triggered = true - ops.Status.Message = "suspend request sent, waiting for node to suspend" - _ = r.Status().Patch(ctx, ops, patch) - return ctrl.Result{RequeueAfter: drainRequeueSuspend}, nil - } - - // Poll node status. - endpoint := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s", clusterUUID, nodeUUID) - body, status, err := apiClient.Do(ctx, http.MethodGet, endpoint, nil) - if err != nil || status >= 300 { - if err == nil { - err = fmt.Errorf("status %d", status) - } - log.Error(err, "drain: failed to GET node status during suspend poll") - return ctrl.Result{RequeueAfter: drainRequeueSuspend}, nil - } - var nodeResp utils.NodeStatusResponse - if err := json.Unmarshal(body, &nodeResp); err != nil { - log.Error(err, "drain: failed to unmarshal node status") - return ctrl.Result{RequeueAfter: drainRequeueSuspend}, nil - } - if nodeResp.Status != utils.NodeStatusSuspended { - r.Recorder.Eventf(ops, nil, corev1.EventTypeWarning, "DrainSuspendPending", "DrainSuspendPending", - "waiting for node %s to suspend (current status: %s)", nodeUUID, nodeResp.Status) - r.emitOnStorageNode(ctx, ops, corev1.EventTypeWarning, "DrainSuspendPending", fmt.Sprintf("waiting for node %s to suspend (current status: %s)", nodeUUID, nodeResp.Status)) - return ctrl.Result{RequeueAfter: drainRequeueSuspend}, nil - } - return r.advanceSubPhase(ctx, ops, simplyblockv1alpha2.StorageNodeOpsSubPhaseMigrating) -} - -func (r *StorageNodeOpsReconciler) drainMigrate( - ctx context.Context, - ops *simplyblockv1alpha2.StorageNodeOps, - sn *simplyblockv1alpha1.StorageNode, - clusterUUID string, - apiClient *webapi.Client, -) (ctrl.Result, error) { - log := logf.FromContext(ctx) - nodeUUID := sn.Status.UUID - - var vmigList simplyblockv1alpha1.VolumeMigrationList - if err := r.List(ctx, &vmigList, - client.InNamespace(ops.Namespace), - client.MatchingLabels{"storage.simplyblock.io/drain-node": nodeUUID}, - ); err != nil { - log.Error(err, "drain: failed to list VolumeMigration CRs") - return ctrl.Result{RequeueAfter: drainRequeueMigrate}, nil - } - - // Handle failed migrations. - if res, handled := r.handleFailedVolumeMigrations(ctx, ops, apiClient, vmigList.Items); handled { - return res, nil - } - - completed, inProgress := 0, 0 - for i := range vmigList.Items { - if vmigList.Items[i].Status.Phase == simplyblockv1alpha1.VolumeMigrationPhaseCompleted { - completed++ - } else { - inProgress++ - } - } - - existingVMNames := make(map[string]struct{}, len(vmigList.Items)) - for i := range vmigList.Items { - existingVMNames[vmigList.Items[i].Name] = struct{}{} - } - - if len(vmigList.Items) == 0 || r.hasMissingVolumeMigrationsOps(ctx, apiClient, clusterUUID, nodeUUID, ops, existingVMNames) { - return r.createMissingVolumeMigrationsOps(ctx, apiClient, clusterUUID, ops, sn, vmigList.Items, existingVMNames) - } - - if inProgress == 0 && completed == len(vmigList.Items) { - patch := client.MergeFrom(ops.DeepCopy()) - ops.Status.VolumesMigrated = completed - ops.Status.VolumesPending = 0 - _ = r.Status().Patch(ctx, ops, patch) - - for i := range vmigList.Items { - vm := &vmigList.Items[i] - if err := r.Delete(ctx, vm); err != nil { - log.Error(err, "drain: failed to delete completed VolumeMigration", "name", vm.Name) - } - } - r.Recorder.Eventf(ops, nil, corev1.EventTypeNormal, "MigrationCompleted", "MigrationCompleted", - "all %d volume migrations completed", completed) - r.emitOnStorageNode(ctx, ops, corev1.EventTypeNormal, "MigrationCompleted", fmt.Sprintf("all %d volume migrations completed", completed)) - return r.advanceSubPhase(ctx, ops, simplyblockv1alpha2.StorageNodeOpsSubPhaseVerifying) - } - - patch := client.MergeFrom(ops.DeepCopy()) - ops.Status.VolumesMigrated = completed - ops.Status.VolumesPending = inProgress - ops.Status.Message = fmt.Sprintf("Migrating: %d of %d volumes migrated", completed, len(vmigList.Items)) - _ = r.Status().Patch(ctx, ops, patch) - return ctrl.Result{RequeueAfter: drainRequeueMigrate}, nil -} - -func (r *StorageNodeOpsReconciler) handleFailedVolumeMigrations( - ctx context.Context, - ops *simplyblockv1alpha2.StorageNodeOps, - apiClient *webapi.Client, - items []simplyblockv1alpha1.VolumeMigration, -) (ctrl.Result, bool) { - log := logf.FromContext(ctx) - var failed []simplyblockv1alpha1.VolumeMigration - for i := range items { - if items[i].Status.Phase == simplyblockv1alpha1.VolumeMigrationPhaseFailed || - items[i].Status.Phase == simplyblockv1alpha1.VolumeMigrationPhaseAborted { - failed = append(failed, items[i]) - } - } - if len(failed) == 0 { - return ctrl.Result{}, false - } - - // Check if the cluster is paused — if so, delete and wait. - if res, paused := r.clusterPauseCheck(ctx, ops, apiClient); paused { - for i := range failed { - _ = r.Delete(ctx, &failed[i]) - } - log.Info("drain: cluster not ready, deleted failed VMs and pausing", "count", len(failed)) - return res, true - } - - // Cluster ready: delete failed CRs and let createMissingVolumeMigrationsOps recreate them. - for i := range failed { - vm := &failed[i] - if err := r.Delete(ctx, vm); err != nil { - log.Error(err, "drain: failed to delete failed VolumeMigration", "name", vm.Name) - continue - } - r.Recorder.Eventf(ops, nil, corev1.EventTypeWarning, "MigrationRetry", "MigrationRetry", - "VolumeMigration %s failed, deleted and will retry with new target", vm.Name) - r.emitOnStorageNode(ctx, ops, corev1.EventTypeWarning, "MigrationRetry", fmt.Sprintf("VolumeMigration %s failed, deleted and will retry with new target", vm.Name)) - } - return ctrl.Result{RequeueAfter: drainRequeueImmediate}, true -} - -func (r *StorageNodeOpsReconciler) hasMissingVolumeMigrationsOps( - ctx context.Context, - apiClient *webapi.Client, - clusterUUID, nodeUUID string, - ops *simplyblockv1alpha2.StorageNodeOps, - existingVMNames map[string]struct{}, -) bool { - vols, err := listNodeVolumes(ctx, apiClient, clusterUUID, nodeUUID) - if err != nil { - return false - } - sf, err := r.resolveOpsSystemVolumeFilter(ops) - if err != nil { - return false - } - pvm, _, _, pvByVol, _, err := matchVolumesToPVs(ctx, r.Client, vols, sf) - if err != nil { - return false - } - for _, volUUID := range pvm { - if pvName, ok := pvByVol[volUUID]; ok { - if _, exists := existingVMNames[drainMigrationName(nodeUUID, pvName)]; !exists { - return true - } - } - } - return false -} - -func (r *StorageNodeOpsReconciler) createMissingVolumeMigrationsOps( - ctx context.Context, - apiClient *webapi.Client, - clusterUUID string, - ops *simplyblockv1alpha2.StorageNodeOps, - sn *simplyblockv1alpha1.StorageNode, - existingItems []simplyblockv1alpha1.VolumeMigration, - existingVMNames map[string]struct{}, -) (ctrl.Result, error) { - log := logf.FromContext(ctx) - nodeUUID := sn.Status.UUID - - volumes, err := listNodeVolumes(ctx, apiClient, clusterUUID, nodeUUID) - if err != nil { - log.Error(err, "drain: failed to list volumes for migration creation") - return ctrl.Result{RequeueAfter: drainRequeueMigrateNew}, nil - } - - sysFilter, err := r.resolveOpsSystemVolumeFilter(ops) - if err != nil { - return r.failOps(ctx, ops, "invalid systemVolumeFilterRegex: "+err.Error()) - } - - pvManaged, _, _, pvNameByVolumeUUID, pvcFetchFailed, err := matchVolumesToPVs(ctx, r.Client, volumes, sysFilter) - if err != nil { - log.Error(err, "drain: matchVolumesToPVs failed") - return ctrl.Result{RequeueAfter: drainRequeueMigrateNew}, nil - } - if pvcFetchFailed { - log.Info("drain: PVC fetch failed — retrying to avoid skipping volumes") - return ctrl.Result{RequeueAfter: drainRequeueMigrateNew}, nil - } - - if len(pvManaged) == 0 && len(existingItems) == 0 { - return r.advanceSubPhase(ctx, ops, simplyblockv1alpha2.StorageNodeOpsSubPhaseVerifying) - } - - pvNames := make([]string, 0, len(pvManaged)) - for _, volUUID := range pvManaged { - pvName, ok := pvNameByVolumeUUID[volUUID] - if !ok { - continue - } - if _, exists := existingVMNames[drainMigrationName(nodeUUID, pvName)]; !exists { - pvNames = append(pvNames, pvName) - } - } - if len(pvNames) == 0 { - return ctrl.Result{RequeueAfter: drainRequeueMigrate}, nil - } - - targetByPV, err := roundRobinTargetNodes(ctx, apiClient, clusterUUID, nodeUUID, pvNames) - if err != nil { - log.Error(err, "drain: no available target nodes for migration") - r.Recorder.Eventf(ops, nil, corev1.EventTypeWarning, "DrainNoMigrationTarget", "DrainNoMigrationTarget", - "drain stalled: no online storage node available as migration target for node %s", nodeUUID) - r.emitOnStorageNode(ctx, ops, corev1.EventTypeWarning, "DrainNoMigrationTarget", fmt.Sprintf("drain stalled: no online storage node available as migration target for node %s", nodeUUID)) - return ctrl.Result{RequeueAfter: drainRequeueMigrateNew}, nil - } - - createdCount := 0 - for _, volUUID := range pvManaged { - pvName, ok := pvNameByVolumeUUID[volUUID] - if !ok { - continue - } - migName := drainMigrationName(nodeUUID, pvName) - if _, exists := existingVMNames[migName]; exists { - continue - } - vmig := &simplyblockv1alpha1.VolumeMigration{ - ObjectMeta: metav1.ObjectMeta{ - Name: migName, - Namespace: ops.Namespace, - Labels: map[string]string{"storage.simplyblock.io/drain-node": nodeUUID}, - }, - Spec: simplyblockv1alpha1.VolumeMigrationSpec{ - PVName: pvName, - TargetNodeUUID: targetByPV[pvName], - }, - } - if err := controllerutil.SetControllerReference(ops, vmig, r.Scheme); err != nil { - log.Error(err, "drain: failed to set controller reference", "name", migName) - continue - } - if err := r.Create(ctx, vmig); err != nil { - log.Error(err, "drain: failed to create VolumeMigration", "name", migName) - continue - } - createdCount++ - } - - patch := client.MergeFrom(ops.DeepCopy()) - ops.Status.VolumesPending = createdCount - ops.Status.VolumesMigrated = 0 - ops.Status.Message = fmt.Sprintf("Migrating: 0 of %d volumes migrated", createdCount) - _ = r.Status().Patch(ctx, ops, patch) - return ctrl.Result{RequeueAfter: drainRequeueMigrateNew}, nil -} - -func (r *StorageNodeOpsReconciler) drainVerify( - ctx context.Context, - ops *simplyblockv1alpha2.StorageNodeOps, - sn *simplyblockv1alpha1.StorageNode, - clusterUUID string, - apiClient *webapi.Client, -) (ctrl.Result, error) { - log := logf.FromContext(ctx) - nodeUUID := sn.Status.UUID - - pools, volumes, err := fetchPoolVolumes(ctx, apiClient, clusterUUID, nodeUUID) - if err != nil { - log.Error(err, "drain: failed to list volumes during verification") - return ctrl.Result{RequeueAfter: drainRequeueVerify}, nil - } - - sysFilter, err := r.resolveOpsSystemVolumeFilter(ops) - if err != nil { - return r.failOps(ctx, ops, "invalid systemVolumeFilterRegex: "+err.Error()) - } - - var nonSystem, systemVols []string - for _, vol := range volumes { - if sysFilter.MatchString(vol.Name) { - systemVols = append(systemVols, vol.UUID) - } else { - nonSystem = append(nonSystem, vol.UUID) - } - } - - if len(nonSystem) > 0 { - r.Recorder.Eventf(ops, nil, corev1.EventTypeWarning, "DrainVerifyPending", "DrainVerifyPending", - "node %s still has %d non-system volume(s) after migration; waiting for backend to confirm empty", - nodeUUID, len(nonSystem)) - r.emitOnStorageNode(ctx, ops, corev1.EventTypeWarning, "DrainVerifyPending", fmt.Sprintf("node %s still has %d non-system volume(s) after migration; waiting for backend to confirm empty", nodeUUID, len(nonSystem))) - return ctrl.Result{RequeueAfter: drainRequeueVerify}, nil - } - - if len(systemVols) > 0 { - poolByVol := make(map[string]string) - for _, pool := range pools { - vols, err := apiClient.GetPoolVolumes(ctx, clusterUUID, pool.UUID) - if err != nil { - continue - } - for _, v := range vols { - poolByVol[v.UUID] = pool.UUID - } - } - for _, volUUID := range systemVols { - poolUUID, ok := poolByVol[volUUID] - if !ok { - continue - } - endpoint := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/", - clusterUUID, poolUUID, volUUID) - _, delStatus, delErr := apiClient.Do(ctx, http.MethodDelete, endpoint, nil) - delClass := webapi.ClassifyError(delErr, delStatus) - switch { - case delErr == nil && (delStatus == http.StatusOK || delStatus == http.StatusNoContent || delStatus == http.StatusNotFound): - log.Info("drain: deleted system volume", "volUUID", volUUID) - case delClass.Retryable: - log.Error(delErr, "drain: transient error deleting system volume, retrying", "volUUID", volUUID) - default: - return r.resumeAndFail(ctx, ops, sn, apiClient, clusterUUID, - fmt.Sprintf("system volume %s delete rejected by backend (status %d)", volUUID, delStatus)) - } - } - return ctrl.Result{RequeueAfter: drainRequeueVerify}, nil - } - - return r.advanceSubPhase(ctx, ops, simplyblockv1alpha2.StorageNodeOpsSubPhaseRemoving) -} - -func (r *StorageNodeOpsReconciler) drainRemove( - ctx context.Context, - ops *simplyblockv1alpha2.StorageNodeOps, - sn *simplyblockv1alpha1.StorageNode, - clusterUUID string, - apiClient *webapi.Client, -) (ctrl.Result, error) { - log := logf.FromContext(ctx) - nodeUUID := sn.Status.UUID - - endpoint := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s?force_remove=false", - clusterUUID, nodeUUID) - _, status, err := apiClient.Do(ctx, http.MethodDelete, endpoint, nil) - - if err == nil && (status == http.StatusOK || status == http.StatusNoContent || status == http.StatusNotFound) { - r.Recorder.Eventf(ops, nil, corev1.EventTypeNormal, "NodeRemoved", "NodeRemoved", - "storage node %s removed successfully", nodeUUID) - r.emitOnStorageNode(ctx, ops, corev1.EventTypeNormal, "NodeRemoved", fmt.Sprintf("storage node %s removed successfully", nodeUUID)) - return r.succeedOps(ctx, ops, sn) - } - - class := webapi.ClassifyError(err, status) - if class.Retryable { - log.Error(err, "drain: transient error on node DELETE, retrying", "status", status) - return ctrl.Result{RequeueAfter: drainRequeueSuspend}, nil - } - return r.resumeAndFail(ctx, ops, sn, apiClient, clusterUUID, - fmt.Sprintf("DELETE node returned status %d", status)) -} - -func (r *StorageNodeOpsReconciler) resumeAndFail( - ctx context.Context, - ops *simplyblockv1alpha2.StorageNodeOps, - sn *simplyblockv1alpha1.StorageNode, - apiClient *webapi.Client, - clusterUUID, reason string, -) (ctrl.Result, error) { - log := logf.FromContext(ctx) - nodeUUID := sn.Status.UUID - - resumeEndpoint := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/resume", clusterUUID, nodeUUID) - _, resumeStatus, resumeErr := apiClient.Do(ctx, http.MethodPost, resumeEndpoint, nil) - resumeClass := webapi.ClassifyError(resumeErr, resumeStatus) - if resumeClass.Retryable { - log.Error(resumeErr, "drain: transient error resuming node, will retry", "status", resumeStatus) - patch := client.MergeFrom(ops.DeepCopy()) - ops.Status.Message = fmt.Sprintf("resume pending after failure: %s", reason) - _ = r.Status().Patch(ctx, ops, patch) - return ctrl.Result{RequeueAfter: drainRequeueSuspend}, nil - } - r.Recorder.Eventf(ops, nil, corev1.EventTypeWarning, "NodeResumed", "NodeResumed", - "drain failed, attempted resume of node %s: %s", nodeUUID, reason) - r.emitOnStorageNode(ctx, ops, corev1.EventTypeWarning, "NodeResumed", fmt.Sprintf("drain failed, attempted resume of node %s: %s", nodeUUID, reason)) - return r.failOps(ctx, ops, reason) -} - -// clusterPauseCheck returns (requeue, true) if the cluster is not ready for drain operations. -func (r *StorageNodeOpsReconciler) clusterPauseCheck( - ctx context.Context, - ops *simplyblockv1alpha2.StorageNodeOps, - _ *webapi.Client, -) (ctrl.Result, bool) { - log := logf.FromContext(ctx) - - // Resolve the StorageNode to get the namespace and cluster name. - var sn simplyblockv1alpha1.StorageNode - if err := r.Get(ctx, types.NamespacedName{Name: ops.Spec.NodeRef, Namespace: ops.Namespace}, &sn); err != nil { - return ctrl.Result{RequeueAfter: drainRequeueSuspend}, false - } - var sns simplyblockv1alpha1.StorageNodeSet - if err := r.Get(ctx, types.NamespacedName{Name: sn.Spec.StorageNodeSetRef, Namespace: sn.Namespace}, &sns); err != nil { - return ctrl.Result{RequeueAfter: drainRequeueSuspend}, false - } - - clusterCR, err := utils.ResolveClusterCR(ctx, r.Client, ops.Namespace, sns.Spec.ClusterName) - if err != nil { - log.Error(err, "drain: could not resolve cluster CR") - return ctrl.Result{RequeueAfter: drainRequeueSuspend}, false - } - - var reason string - if clusterCR.Status.Status != "" && clusterCR.Status.Status != utils.ClusterStatusActive { - reason = fmt.Sprintf("cluster status is %q (not active)", clusterCR.Status.Status) - } else if clusterCR.Status.Rebalancing != nil && *clusterCR.Status.Rebalancing { - reason = "cluster is rebalancing" - } - - if reason == "" { - return ctrl.Result{}, false - } - - patch := client.MergeFrom(ops.DeepCopy()) - ops.Status.Message = "drain paused: " + reason - _ = r.Status().Patch(ctx, ops, patch) - r.Recorder.Eventf(ops, nil, corev1.EventTypeWarning, "DrainPaused", "DrainPaused", - "drain paused: %s — will resume when cluster is active", reason) - r.emitOnStorageNode(ctx, ops, corev1.EventTypeWarning, "DrainPaused", fmt.Sprintf("drain paused: %s — will resume when cluster is active", reason)) - log.Info("drain: pausing — cluster not ready", "reason", reason) - return ctrl.Result{RequeueAfter: 60 * time.Second}, true -} - -// advanceSubPhase patches ops.status.subPhase and requeues immediately. -func (r *StorageNodeOpsReconciler) advanceSubPhase( - ctx context.Context, - ops *simplyblockv1alpha2.StorageNodeOps, - next simplyblockv1alpha2.StorageNodeOpsSubPhase, -) (ctrl.Result, error) { - patch := client.MergeFrom(ops.DeepCopy()) - ops.Status.SubPhase = next - ops.Status.Triggered = false - ops.Status.Message = fmt.Sprintf("entering phase %s", next) - if err := r.Status().Patch(ctx, ops, patch); err != nil { - return ctrl.Result{}, err - } - return ctrl.Result{RequeueAfter: drainRequeueImmediate}, nil -} - -// succeedOps marks the ops as Succeeded and releases the lock on the StorageNode. -func (r *StorageNodeOpsReconciler) succeedOps( - ctx context.Context, - ops *simplyblockv1alpha2.StorageNodeOps, - sn *simplyblockv1alpha1.StorageNode, -) (ctrl.Result, error) { - now := metav1.Now() - patch := client.MergeFrom(ops.DeepCopy()) - ops.Status.Phase = simplyblockv1alpha2.StorageNodeOpsPhaseSucceeded - ops.Status.SubPhase = "" - ops.Status.CompletedAt = &now - if err := r.Status().Patch(ctx, ops, patch); err != nil { - return ctrl.Result{}, err - } - return ctrl.Result{}, r.releaseLock(ctx, sn, ops.Name) -} - -// failOps marks the ops as Failed with the given reason and releases the lock. -func (r *StorageNodeOpsReconciler) failOps( - ctx context.Context, - ops *simplyblockv1alpha2.StorageNodeOps, - reason string, -) (ctrl.Result, error) { - log := logf.FromContext(ctx) - log.Error(nil, "ops failed", "ops", ops.Name, "reason", reason) - r.Recorder.Eventf(ops, nil, "Warning", "OpsFailed", "OpsFailed", "%s", reason) - r.emitOnStorageNode(ctx, ops, "Warning", "OpsFailed", reason) - - now := metav1.Now() - patch := client.MergeFrom(ops.DeepCopy()) - ops.Status.Phase = simplyblockv1alpha2.StorageNodeOpsPhaseFailed - ops.Status.SubPhase = "" - ops.Status.Message = reason - ops.Status.CompletedAt = &now - if err := r.Status().Patch(ctx, ops, patch); err != nil { - return ctrl.Result{}, err - } - - var sn simplyblockv1alpha1.StorageNode - if err := r.Get(ctx, types.NamespacedName{ - Name: ops.Spec.NodeRef, - Namespace: ops.Namespace, - }, &sn); err == nil { - _ = r.releaseLock(ctx, &sn, ops.Name) - } - return ctrl.Result{}, nil -} - -// emitOnStorageNode emits an event on the StorageNode that this ops targets, -// mirroring events that are also emitted on the StorageNodeOps CR itself. -func (r *StorageNodeOpsReconciler) emitOnStorageNode( - ctx context.Context, - ops *simplyblockv1alpha2.StorageNodeOps, - eventType, reason, message string, -) { - var sn simplyblockv1alpha1.StorageNode - if err := r.Get(ctx, types.NamespacedName{Name: ops.Spec.NodeRef, Namespace: ops.Namespace}, &sn); err != nil { - return - } - r.Recorder.Eventf(&sn, nil, eventType, reason, reason, "%s", message) -} - -// releaseLock clears StorageNode.status.activeOpsRef if it still points to opsName. -func (r *StorageNodeOpsReconciler) releaseLock( - ctx context.Context, - sn *simplyblockv1alpha1.StorageNode, - opsName string, -) error { - if sn.Status.ActiveOpsRef != opsName { - return nil - } - patch := client.MergeFrom(sn.DeepCopy()) - sn.Status.ActiveOpsRef = "" - return r.Status().Patch(ctx, sn, patch) -} - -// resolveOpsSystemVolumeFilter compiles the system volume filter regex from the ops, -// falling back to the default pattern. -func (r *StorageNodeOpsReconciler) resolveOpsSystemVolumeFilter( - ops *simplyblockv1alpha2.StorageNodeOps, -) (*regexp.Regexp, error) { - pattern := simplyblockv1alpha1.DefaultSystemVolumeFilterRegex - if ops.Spec.Remove != nil && ops.Spec.Remove.SystemVolumeFilterRegex != nil { - pattern = *ops.Spec.Remove.SystemVolumeFilterRegex - } - return regexp.Compile(pattern) -} - -// storageNodeToOpsRequests maps a StorageNode change to any pending -// StorageNodeOps that targets it, so ops waiting on lock acquisition requeue -// immediately when activeOpsRef is cleared rather than waiting for the poll timer. -func (r *StorageNodeOpsReconciler) storageNodeToOpsRequests( - ctx context.Context, - obj client.Object, -) []reconcile.Request { - var opsList simplyblockv1alpha2.StorageNodeOpsList - if err := r.List(ctx, &opsList, - client.InNamespace(obj.GetNamespace()), - client.MatchingFields{"spec.storageNodeRef": obj.GetName()}, - ); err != nil { - return nil - } - reqs := make([]reconcile.Request, 0, len(opsList.Items)) - for _, ops := range opsList.Items { - if ops.Status.Phase == simplyblockv1alpha2.StorageNodeOpsPhasePending || - ops.Status.Phase == "" { - reqs = append(reqs, reconcile.Request{NamespacedName: types.NamespacedName{ - Name: ops.Name, - Namespace: ops.Namespace, - }}) - } - } - return reqs -} - -// SetupWithManager registers the StorageNodeOpsReconciler with the controller manager. -func (r *StorageNodeOpsReconciler) SetupWithManager(mgr ctrl.Manager) error { - // Index StorageNodeOps by their target StorageNode for efficient watch lookups. - if err := mgr.GetFieldIndexer().IndexField( - context.Background(), - &simplyblockv1alpha2.StorageNodeOps{}, - "spec.storageNodeRef", - func(obj client.Object) []string { - ops := obj.(*simplyblockv1alpha2.StorageNodeOps) - return []string{ops.Spec.NodeRef} - }, - ); err != nil { - return err - } - - // Uncached reader for the migrate DNS gate (endpointSliceHasWorker). - r.apiReader = mgr.GetAPIReader() - - return ctrl.NewControllerManagedBy(mgr). - For(&simplyblockv1alpha2.StorageNodeOps{}). - Named("storagenodeops"). - Watches( - &simplyblockv1alpha1.StorageNode{}, - handler.EnqueueRequestsFromMapFunc(r.storageNodeToOpsRequests), - ). - Owns(&simplyblockv1alpha1.VolumeMigration{}). - Complete(r) -} diff --git a/operator/internal/controller/storagenodeops_controller_unit_test.go b/operator/internal/controller/storagenodeops_controller_unit_test.go deleted file mode 100644 index 861cb17fa..000000000 --- a/operator/internal/controller/storagenodeops_controller_unit_test.go +++ /dev/null @@ -1,596 +0,0 @@ -package controller - -import ( - "context" - "net/http" - "net/http/httptest" - "testing" - - corev1 "k8s.io/api/core/v1" - discoveryv1 "k8s.io/api/discovery/v1" - metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" - "k8s.io/apimachinery/pkg/types" - "k8s.io/client-go/tools/events" - "sigs.k8s.io/controller-runtime/pkg/client" - - simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" - simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" - "github.com/simplyblock/simplyblock-operator/internal/utils" - "github.com/simplyblock/simplyblock-operator/internal/webapi" -) - -// ── helpers ─────────────────────────────────────────────────────────────────── - -const ( - opsTestNS = "test" - opsTestCluster = "cluster-a" - opsTestWorker = "worker-1.example.com" - opsTestNodeUUID = "aaaa0000-0000-0000-0000-000000000001" - opsTestOpsName = "ops-1" - opsTestOtherOps = "ops-other" -) - -func newOpsReconciler(t *testing.T, objects ...client.Object) *StorageNodeOpsReconciler { - t.Helper() - scheme := newTestScheme(t, - simplyblockv1alpha1.AddToScheme, - corev1.AddToScheme, - ) - cl := newTestClient(t, scheme, - []client.Object{ - &simplyblockv1alpha1.StorageNode{}, - &simplyblockv1alpha2.StorageNodeOps{}, - &simplyblockv1alpha1.StorageNodeSet{}, - &simplyblockv1alpha2.StorageCluster{}, - &simplyblockv1alpha1.VolumeMigration{}, - }, - objects..., - ) - return &StorageNodeOpsReconciler{ - Client: cl, - Scheme: scheme, - Recorder: events.NewFakeRecorder(16), - apiReader: cl, - } -} - -//nolint:unparam -func newTestStorageNode(name, ns, snsRef, worker, uuid string) *simplyblockv1alpha1.StorageNode { - sn := &simplyblockv1alpha1.StorageNode{ - ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: ns}, - Spec: simplyblockv1alpha1.StorageNodeSpec{ - StorageNodeSetRef: snsRef, - WorkerNode: worker, - }, - } - sn.Status.UUID = uuid - return sn -} - -//nolint:unparam -func newTestStorageNodeOps(name, ns, snRef string, action simplyblockv1alpha2.StorageNodeOpsAction) *simplyblockv1alpha2.StorageNodeOps { - return &simplyblockv1alpha2.StorageNodeOps{ - ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: ns}, - Spec: simplyblockv1alpha2.StorageNodeOpsSpec{ - NodeRef: snRef, - Action: action, - }, - } -} - -// ── TestAcquireLock ─────────────────────────────────────────────────────────── - -func TestAcquireLock_SetsActiveOpsRefAndTransitionsToRunning(t *testing.T) { - sn := newTestStorageNode("sn-1", opsTestNS, "sns", opsTestWorker, opsTestNodeUUID) - ops := newTestStorageNodeOps(opsTestOpsName, opsTestNS, "sn-1", simplyblockv1alpha2.StorageNodeOpsActionSuspend) - r := newOpsReconciler(t, sn, ops) - - _, err := r.acquireLock(context.Background(), ops, sn) - if err != nil { - t.Fatalf("acquireLock returned error: %v", err) - } - - // Check StorageNode.status.activeOpsRef was set. - var updatedSN simplyblockv1alpha1.StorageNode - _ = r.Get(context.Background(), types.NamespacedName{Name: "sn-1", Namespace: opsTestNS}, &updatedSN) - if updatedSN.Status.ActiveOpsRef != opsTestOpsName { - t.Errorf("activeOpsRef: got %q want ops-1", updatedSN.Status.ActiveOpsRef) - } - - // Check ops phase was set to Running. - var updatedOps simplyblockv1alpha2.StorageNodeOps - _ = r.Get(context.Background(), types.NamespacedName{Name: opsTestOpsName, Namespace: opsTestNS}, &updatedOps) - if updatedOps.Status.Phase != simplyblockv1alpha2.StorageNodeOpsPhaseRunning { - t.Errorf("phase: got %q want Running", updatedOps.Status.Phase) - } -} - -func TestAcquireLock_RequeuesWhenAnotherOpsActive(t *testing.T) { - sn := newTestStorageNode("sn-1", opsTestNS, "sns", opsTestWorker, opsTestNodeUUID) - sn.Status.ActiveOpsRef = opsTestOtherOps - ops := newTestStorageNodeOps(opsTestOpsName, opsTestNS, "sn-1", simplyblockv1alpha2.StorageNodeOpsActionSuspend) - r := newOpsReconciler(t, sn, ops) - - result, err := r.acquireLock(context.Background(), ops, sn) - if err != nil { - t.Fatalf("acquireLock returned error: %v", err) - } - if result.RequeueAfter == 0 { - t.Error("expected requeue when another ops is active") - } - - // StorageNode.activeOpsRef must NOT be changed. - var updatedSN simplyblockv1alpha1.StorageNode - _ = r.Get(context.Background(), types.NamespacedName{Name: "sn-1", Namespace: opsTestNS}, &updatedSN) - if updatedSN.Status.ActiveOpsRef != opsTestOtherOps { - t.Errorf("activeOpsRef should not change: got %q", updatedSN.Status.ActiveOpsRef) - } -} - -// ── TestFdRemovalBalanceCheck ──────────────────────────────────────────────── -// -// Mirrors the live 2026-08-13 incident's exact topology: 7 nodes split -// FD1=2/FD2=2/FD3=3. Removing a node from the already-smallest domain (FD1) -// drops it to 1 while FD3 stays at 3 -- a spread the backend's own -// check_fd_admission_for_remove correctly refuses (populations {1,2,3}). -// Removing instead from the domain with slack (FD3) leaves 2/2/2, which is -// fine. This is the gate that must fire in drainValidate BEFORE Suspending, -// so an infeasible removal never suspends the node in the first place. - -//nolint:unparam -func newTestStorageNodeSet(name, ns, clusterName string, nodes ...simplyblockv1alpha1.NodeStatus) *simplyblockv1alpha1.StorageNodeSet { - return &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: ns}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ClusterName: clusterName}, - Status: simplyblockv1alpha1.StorageNodeSetStatus{Nodes: nodes}, - } -} - -//nolint:unparam -func newTestStorageClusterWithFD(name, ns string, enableFD bool) *simplyblockv1alpha2.StorageCluster { - return &simplyblockv1alpha2.StorageCluster{ - ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: ns}, - Spec: simplyblockv1alpha2.StorageClusterSpec{EnableFailureDomains: &enableFD}, - } -} - -// sevenNodeTopology returns the live 2026-08-13 incident's exact topology: 7 -// nodes split FD1=2/FD2=2/FD3=3, sn-1 in FD1 at mgmtIP. The entry at mgmtIP -// gets UUID=opsTestNodeUUID -- matching sn.Status.UUID in every caller -- -// so hostHasSurvivingSibling correctly recognizes it as sn-1 itself rather -// than an unrelated sibling on the same host (the other six entries are -// distinct hosts at distinct IPs, so their empty UUID never collides with -// a real lookup at a different IP). -func sevenNodeTopology(mgmtIP string, snFD int32) []simplyblockv1alpha1.NodeStatus { - fd := func(v int32) *int32 { return &v } - all := []struct { - ip string - domain int32 - }{ - {"10.0.0.1", 1}, {"10.0.0.2", 1}, - {"10.0.0.3", 2}, {"10.0.0.4", 2}, - {"10.0.0.5", 3}, {"10.0.0.6", 3}, {"10.0.0.7", 3}, - } - nodes := make([]simplyblockv1alpha1.NodeStatus, 0, len(all)) - for _, n := range all { - domain := n.domain - uuid := "" - if n.ip == mgmtIP { - domain = snFD // let the caller override sn-1's own domain - uuid = opsTestNodeUUID - } - nodes = append(nodes, simplyblockv1alpha1.NodeStatus{MgmtIp: n.ip, FailureDomain: fd(domain), UUID: uuid}) - } - return nodes -} - -func TestFdRemovalBalanceCheck_RemovingFromSlackDomainAllowed(t *testing.T) { - // Removing from FD3 (2/2/3 -> 2/2/2) is fine. - sn := newTestStorageNode("sn-1", opsTestNS, "sns", opsTestWorker, opsTestNodeUUID) - sn.Status.Ports = &simplyblockv1alpha1.StorageNodePorts{Management: "10.0.0.7"} - sns := newTestStorageNodeSet("sns", opsTestNS, opsTestCluster, sevenNodeTopology("10.0.0.7", 3)...) - cluster := newTestStorageClusterWithFD(opsTestCluster, opsTestNS, true) - r := newOpsReconciler(t, sn, sns, cluster) - - reason, err := r.fdRemovalBalanceCheck(context.Background(), sn) - if err != nil { - t.Fatalf("fdRemovalBalanceCheck returned error: %v", err) - } - if reason != "" { - t.Errorf("expected removal from FD3 (2/2/3 -> 2/2/2) to be allowed, got blocked: %q", reason) - } -} - -func TestFdRemovalBalanceCheck_RemovingFromThinDomainBlocked(t *testing.T) { - // Removing from FD1 (2/2/3 -> 1/2/3) violates the +/-1 rule. - sn := newTestStorageNode("sn-1", opsTestNS, "sns", opsTestWorker, opsTestNodeUUID) - sn.Status.Ports = &simplyblockv1alpha1.StorageNodePorts{Management: "10.0.0.2"} - sns := newTestStorageNodeSet("sns", opsTestNS, opsTestCluster, sevenNodeTopology("10.0.0.2", 1)...) - cluster := newTestStorageClusterWithFD(opsTestCluster, opsTestNS, true) - r := newOpsReconciler(t, sn, sns, cluster) - - reason, err := r.fdRemovalBalanceCheck(context.Background(), sn) - if err != nil { - t.Fatalf("fdRemovalBalanceCheck returned error: %v", err) - } - if reason == "" { - t.Error("expected removal from FD1 (2/2/3 -> 1/2/3) to be blocked, got none") - } -} - -// TestFdRemovalBalanceCheck_NoOpWhenFailureDomainsDisabled locks in the -// gate's very first early-out, mirroring check_fd_admission_for_remove's -// own first line (simplyblock_core): with EnableFailureDomains unset/false, -// this must never block a removal, regardless of topology -- the exact -// same 1/2/3 split that TestFdRemovalBalanceCheck_RemovingFromThinDomainBlocked -// correctly blocks above must be a no-op when the cluster hasn't opted in. -func TestFdRemovalBalanceCheck_NoOpWhenFailureDomainsDisabled(t *testing.T) { - sn := newTestStorageNode("sn-1", opsTestNS, "sns", opsTestWorker, opsTestNodeUUID) - sn.Status.Ports = &simplyblockv1alpha1.StorageNodePorts{Management: "10.0.0.2"} - sns := newTestStorageNodeSet("sns", opsTestNS, opsTestCluster, sevenNodeTopology("10.0.0.2", 1)...) - cluster := newTestStorageClusterWithFD(opsTestCluster, opsTestNS, false) - r := newOpsReconciler(t, sn, sns, cluster) - - reason, err := r.fdRemovalBalanceCheck(context.Background(), sn) - if err != nil { - t.Fatalf("fdRemovalBalanceCheck returned error: %v", err) - } - if reason != "" { - t.Errorf("expected no-op with failure domains disabled, got blocked: %q", reason) - } -} - -// TestFdRemovalBalanceCheck_MultiNodeHostSiblingSurvivesAllowsRemoval is a -// regression for the 2026-08-26 finding: a host running more than one -// StorageNode (spec.socketsToUse / spec.nodesPerSocket > 1) must not vanish -// from the failure-domain host count when only one of its nodes is removed -// -- the sibling node still lives there. 3 domains, 2 hosts each (balanced -// 2/2/2); FD3's second host (10.0.0.6) runs two StorageNodes sharing that -// management IP. Removing one of them must leave FD3 at 2 hosts, not drop -// it to 1 -- the old delete(hostDomains, ip) removed the whole host and -// produced a spurious "failure domain 3 would drop to 1 host(s)" block. -func TestFdRemovalBalanceCheck_MultiNodeHostSiblingSurvivesAllowsRemoval(t *testing.T) { - fd := func(v int32) *int32 { return &v } - const siblingUUID = "bbbb0000-0000-0000-0000-000000000002" - nodes := []simplyblockv1alpha1.NodeStatus{ - {MgmtIp: "10.0.0.1", FailureDomain: fd(1), UUID: "n1"}, - {MgmtIp: "10.0.0.2", FailureDomain: fd(1), UUID: "n2"}, - {MgmtIp: "10.0.0.3", FailureDomain: fd(2), UUID: "n3"}, - {MgmtIp: "10.0.0.4", FailureDomain: fd(2), UUID: "n4"}, - {MgmtIp: "10.0.0.5", FailureDomain: fd(3), UUID: "n5"}, - {MgmtIp: "10.0.0.6", FailureDomain: fd(3), UUID: opsTestNodeUUID}, // node being removed - {MgmtIp: "10.0.0.6", FailureDomain: fd(3), UUID: siblingUUID}, // surviving sibling, same host - } - sn := newTestStorageNode("sn-1", opsTestNS, "sns", opsTestWorker, opsTestNodeUUID) - sn.Status.Ports = &simplyblockv1alpha1.StorageNodePorts{Management: "10.0.0.6"} - sns := newTestStorageNodeSet("sns", opsTestNS, opsTestCluster, nodes...) - cluster := newTestStorageClusterWithFD(opsTestCluster, opsTestNS, true) - r := newOpsReconciler(t, sn, sns, cluster) - - reason, err := r.fdRemovalBalanceCheck(context.Background(), sn) - if err != nil { - t.Fatalf("fdRemovalBalanceCheck returned error: %v", err) - } - if reason != "" { - t.Errorf("expected removal to be allowed (sibling survives on the host), got blocked: %q", reason) - } -} - -// TestFdRemovalBalanceCheck_DomainDroppingToZeroBlocked is a regression for -// the 2026-08-26 finding: A=2,B=2,C=1 -- removing C's only host must be -// blocked. It would drop the cluster from 3 domains to 2, exactly the -// topology fdActivationDomainCountViolation itself calls unsupported; the -// removal gate must not permit what the activation gate would refuse. The -// old code derived counts from a hostDomains map with C's key deleted, so C -// vanished from the map instead of appearing at count zero, and neither the -// +/-1 spread nor the 2-per-domain floor check ever saw it. -func TestFdRemovalBalanceCheck_DomainDroppingToZeroBlocked(t *testing.T) { - fd := func(v int32) *int32 { return &v } - nodes := []simplyblockv1alpha1.NodeStatus{ - {MgmtIp: "10.0.0.1", FailureDomain: fd(1), UUID: "n1"}, - {MgmtIp: "10.0.0.2", FailureDomain: fd(1), UUID: "n2"}, - {MgmtIp: "10.0.0.3", FailureDomain: fd(2), UUID: "n3"}, - {MgmtIp: "10.0.0.4", FailureDomain: fd(2), UUID: "n4"}, - {MgmtIp: "10.0.0.5", FailureDomain: fd(3), UUID: opsTestNodeUUID}, // domain 3's only host - } - sn := newTestStorageNode("sn-1", opsTestNS, "sns", opsTestWorker, opsTestNodeUUID) - sn.Status.Ports = &simplyblockv1alpha1.StorageNodePorts{Management: "10.0.0.5"} - sns := newTestStorageNodeSet("sns", opsTestNS, opsTestCluster, nodes...) - cluster := newTestStorageClusterWithFD(opsTestCluster, opsTestNS, true) - r := newOpsReconciler(t, sn, sns, cluster) - - reason, err := r.fdRemovalBalanceCheck(context.Background(), sn) - if err != nil { - t.Fatalf("fdRemovalBalanceCheck returned error: %v", err) - } - if reason == "" { - t.Error("expected removal to be blocked (would drop domain 3 to zero hosts), got allowed") - } -} - -// TestDrainValidate_FailureDomainBalanceViolationFailsOps exercises the full -// drainValidate wiring, not just fdRemovalBalanceCheck: a blocked removal -// must land in Phase=Failed (not a silent 60s-requeue Running loop) -- -// deliberately different from the pinned/unmanaged-volume checks above it, -// since restoring failure-domain balance needs a cluster-wide human action -// this ops has no way to detect on its own. Safe to fail here specifically -// because handleDeletion (storagenode_controller.go) now refuses to remove -// the StorageNode's finalizer while its remove ops is Failed. -func TestDrainValidate_FailureDomainBalanceViolationFailsOps(t *testing.T) { - // Empty storage-pools list -> fetchPoolVolumes returns zero volumes -> - // pinned/unmanaged checks pass through, reaching the FD-balance gate. - srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { - w.Header().Set("Content-Type", "application/json") - _, _ = w.Write([]byte("[]")) - })) - defer srv.Close() - - sn := newTestStorageNode("sn-1", opsTestNS, "sns", opsTestWorker, opsTestNodeUUID) - sn.Status.Ports = &simplyblockv1alpha1.StorageNodePorts{Management: "10.0.0.2"} - ops := newTestStorageNodeOps(opsTestOpsName, opsTestNS, "sn-1", simplyblockv1alpha2.StorageNodeOpsActionRemove) - sns := newTestStorageNodeSet("sns", opsTestNS, opsTestCluster, sevenNodeTopology("10.0.0.2", 1)...) - cluster := newTestStorageClusterWithFD(opsTestCluster, opsTestNS, true) - r := newOpsReconciler(t, sn, ops, sns, cluster) - - _, err := r.drainValidate(context.Background(), ops, sn, "cluster-uuid", webapi.NewClient(srv.URL)) - if err != nil { - t.Fatalf("drainValidate returned error: %v", err) - } - - var updated simplyblockv1alpha2.StorageNodeOps - if err := r.Get(context.Background(), types.NamespacedName{Name: opsTestOpsName, Namespace: opsTestNS}, &updated); err != nil { - t.Fatalf("failed to fetch updated ops: %v", err) - } - if updated.Status.Phase != simplyblockv1alpha2.StorageNodeOpsPhaseFailed { - t.Errorf("Phase: got %q, want Failed", updated.Status.Phase) - } - if updated.Status.Message == "" { - t.Error("expected a failure message explaining the balance violation") - } -} - -func TestAcquireLock_RemoveDrainSetsValidatingSubPhase(t *testing.T) { - sn := newTestStorageNode("sn-1", opsTestNS, "sns", opsTestWorker, opsTestNodeUUID) - ops := newTestStorageNodeOps("ops-drain", opsTestNS, "sn-1", simplyblockv1alpha2.StorageNodeOpsActionRemove) - r := newOpsReconciler(t, sn, ops) - - _, err := r.acquireLock(context.Background(), ops, sn) - if err != nil { - t.Fatalf("acquireLock returned error: %v", err) - } - - var updated simplyblockv1alpha2.StorageNodeOps - _ = r.Get(context.Background(), types.NamespacedName{Name: "ops-drain", Namespace: opsTestNS}, &updated) - if updated.Status.SubPhase != simplyblockv1alpha2.StorageNodeOpsSubPhaseValidating { - t.Errorf("subPhase: got %q want Validating", updated.Status.SubPhase) - } -} - -// ── TestSucceedOps ──────────────────────────────────────────────────────────── - -func TestSucceedOps_SetsPhaseAndClearsLock(t *testing.T) { - sn := newTestStorageNode("sn-1", opsTestNS, "sns", opsTestWorker, opsTestNodeUUID) - sn.Status.ActiveOpsRef = opsTestOpsName - ops := newTestStorageNodeOps(opsTestOpsName, opsTestNS, "sn-1", simplyblockv1alpha2.StorageNodeOpsActionSuspend) - ops.Status.Phase = simplyblockv1alpha2.StorageNodeOpsPhaseRunning - r := newOpsReconciler(t, sn, ops) - - _, err := r.succeedOps(context.Background(), ops, sn) - if err != nil { - t.Fatalf("succeedOps returned error: %v", err) - } - - var updatedOps simplyblockv1alpha2.StorageNodeOps - _ = r.Get(context.Background(), types.NamespacedName{Name: opsTestOpsName, Namespace: opsTestNS}, &updatedOps) - if updatedOps.Status.Phase != simplyblockv1alpha2.StorageNodeOpsPhaseSucceeded { - t.Errorf("phase: got %q want Succeeded", updatedOps.Status.Phase) - } - if updatedOps.Status.CompletedAt == nil { - t.Error("expected CompletedAt to be set") - } - - var updatedSN simplyblockv1alpha1.StorageNode - _ = r.Get(context.Background(), types.NamespacedName{Name: "sn-1", Namespace: opsTestNS}, &updatedSN) - if updatedSN.Status.ActiveOpsRef != "" { - t.Errorf("activeOpsRef should be cleared, got %q", updatedSN.Status.ActiveOpsRef) - } -} - -// ── TestFailOps ─────────────────────────────────────────────────────────────── - -func TestFailOps_SetsPhaseAndClearsLock(t *testing.T) { - sn := newTestStorageNode("sn-1", opsTestNS, "sns", opsTestWorker, opsTestNodeUUID) - sn.Status.ActiveOpsRef = opsTestOpsName - ops := newTestStorageNodeOps(opsTestOpsName, opsTestNS, "sn-1", simplyblockv1alpha2.StorageNodeOpsActionSuspend) - ops.Status.Phase = simplyblockv1alpha2.StorageNodeOpsPhaseRunning - r := newOpsReconciler(t, sn, ops) - - _, err := r.failOps(context.Background(), ops, "something went wrong") - if err != nil { - t.Fatalf("failOps returned error: %v", err) - } - - var updatedOps simplyblockv1alpha2.StorageNodeOps - _ = r.Get(context.Background(), types.NamespacedName{Name: opsTestOpsName, Namespace: opsTestNS}, &updatedOps) - if updatedOps.Status.Phase != simplyblockv1alpha2.StorageNodeOpsPhaseFailed { - t.Errorf("phase: got %q want Failed", updatedOps.Status.Phase) - } - if updatedOps.Status.Message != "something went wrong" { - t.Errorf("message: got %q", updatedOps.Status.Message) - } - - var updatedSN simplyblockv1alpha1.StorageNode - _ = r.Get(context.Background(), types.NamespacedName{Name: "sn-1", Namespace: opsTestNS}, &updatedSN) - if updatedSN.Status.ActiveOpsRef != "" { - t.Errorf("activeOpsRef should be cleared after failure, got %q", updatedSN.Status.ActiveOpsRef) - } -} - -// ── TestReleaseLock ─────────────────────────────────────────────────────────── - -func TestReleaseLock_OnlyClearsIfOwner(t *testing.T) { - sn := newTestStorageNode("sn-1", opsTestNS, "sns", opsTestWorker, opsTestNodeUUID) - sn.Status.ActiveOpsRef = opsTestOtherOps - r := newOpsReconciler(t, sn) - - // Releasing with a different name should be a no-op. - if err := r.releaseLock(context.Background(), sn, opsTestOpsName); err != nil { - t.Fatalf("releaseLock returned error: %v", err) - } - - var updated simplyblockv1alpha1.StorageNode - _ = r.Get(context.Background(), types.NamespacedName{Name: "sn-1", Namespace: opsTestNS}, &updated) - if updated.Status.ActiveOpsRef != opsTestOtherOps { - t.Error("releaseLock should not clear a lock it does not own") - } -} - -// ── TestAdvanceSubPhase ─────────────────────────────────────────────────────── - -func TestAdvanceSubPhase_UpdatesSubPhaseAndResetsTrigger(t *testing.T) { - ops := newTestStorageNodeOps("ops-drain", opsTestNS, "sn-1", simplyblockv1alpha2.StorageNodeOpsActionRemove) - ops.Status.Phase = simplyblockv1alpha2.StorageNodeOpsPhaseRunning - ops.Status.SubPhase = simplyblockv1alpha2.StorageNodeOpsSubPhaseValidating - ops.Status.Triggered = true - r := newOpsReconciler(t, ops) - - _, err := r.advanceSubPhase(context.Background(), ops, simplyblockv1alpha2.StorageNodeOpsSubPhaseSuspending) - if err != nil { - t.Fatalf("advanceSubPhase returned error: %v", err) - } - - var updated simplyblockv1alpha2.StorageNodeOps - _ = r.Get(context.Background(), types.NamespacedName{Name: "ops-drain", Namespace: opsTestNS}, &updated) - if updated.Status.SubPhase != simplyblockv1alpha2.StorageNodeOpsSubPhaseSuspending { - t.Errorf("subPhase: got %q want Suspending", updated.Status.SubPhase) - } - if updated.Status.Triggered { - t.Error("Triggered should be reset to false on phase advance") - } -} - -// ── TestDispatch ────────────────────────────────────────────────────────────── - -func TestDispatch_UnknownActionFails(t *testing.T) { - sn := newTestStorageNode("sn-1", opsTestNS, "sns", opsTestWorker, opsTestNodeUUID) - sns := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sns", Namespace: opsTestNS}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ClusterName: opsTestCluster}, - } - ops := newTestStorageNodeOps(opsTestOpsName, opsTestNS, "sn-1", "bogus-action") - ops.Status.Phase = simplyblockv1alpha2.StorageNodeOpsPhaseRunning - r := newOpsReconciler(t, sn, sns, ops) - - _, err := r.dispatch(context.Background(), ops, sn, sns, "cluster-uuid", nil) - if err != nil { - t.Fatalf("dispatch returned unexpected error: %v", err) - } - - var updated simplyblockv1alpha2.StorageNodeOps - _ = r.Get(context.Background(), types.NamespacedName{Name: opsTestOpsName, Namespace: opsTestNS}, &updated) - if updated.Status.Phase != simplyblockv1alpha2.StorageNodeOpsPhaseFailed { - t.Errorf("expected Failed for unknown action, got %q", updated.Status.Phase) - } -} - -// ── TestResolveOpsSystemVolumeFilter ───────────────────────────────────────── - -func TestResolveOpsSystemVolumeFilter_UsesDefaultWhenNoDrain(t *testing.T) { - ops := newTestStorageNodeOps(opsTestOpsName, opsTestNS, "sn-1", simplyblockv1alpha2.StorageNodeOpsActionRemove) - r := newOpsReconciler(t, ops) - - re, err := r.resolveOpsSystemVolumeFilter(ops) - if err != nil { - t.Fatalf("unexpected error: %v", err) - } - // Default pattern matches sb-fio-baseline-* names. - if !re.MatchString("sb-fio-baseline-read") { - t.Error("default filter should match sb-fio-baseline-read") - } - if re.MatchString("user-volume") { - t.Error("default filter should not match user volumes") - } -} - -func TestResolveOpsSystemVolumeFilter_UsesCustomPattern(t *testing.T) { - custom := "^bench-.*" - ops := newTestStorageNodeOps(opsTestOpsName, opsTestNS, "sn-1", simplyblockv1alpha2.StorageNodeOpsActionRemove) - ops.Spec.Remove = &simplyblockv1alpha2.RemoveSpec{SystemVolumeFilterRegex: &custom} - r := newOpsReconciler(t, ops) - - re, err := r.resolveOpsSystemVolumeFilter(ops) - if err != nil { - t.Fatalf("unexpected error: %v", err) - } - if !re.MatchString("bench-read") { - t.Error("custom filter should match bench-read") - } - if re.MatchString("sb-fio-baseline-read") { - t.Error("custom filter should not match sb-fio-baseline-read") - } -} - -func TestResolveOpsSystemVolumeFilter_InvalidPatternReturnsError(t *testing.T) { - bad := "[" - ops := newTestStorageNodeOps(opsTestOpsName, opsTestNS, "sn-1", simplyblockv1alpha2.StorageNodeOpsActionRemove) - ops.Spec.Remove = &simplyblockv1alpha2.RemoveSpec{SystemVolumeFilterRegex: &bad} - r := newOpsReconciler(t, ops) - - _, err := r.resolveOpsSystemVolumeFilter(ops) - if err == nil { - t.Fatal("expected error for invalid regex pattern") - } -} - -// TestEndpointSliceHasWorker_MatchesBuilderOutput guards the coupling between the -// EndpointSlice builder and the migrate flow's DNS gate: a slice built by -// BuildStorageNodeSetEndpointSlice must be found by endpointSliceHasWorker. The -// two independently encoded the slice name and hostname, and a rename that -// touched only the builder silently wedged migrations at "waiting for DNS". -func TestEndpointSliceHasWorker_MatchesBuilderOutput(t *testing.T) { - const ns = "test" - const worker = "worker-5.ocp.simplyblock.ai" - - sns := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "simplyblock-node", Namespace: ns}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ClusterName: "cluster-a"}, - } - // Build the slice exactly as reconcileEndpointSlice does. - slice := utils.BuildStorageNodeSetEndpointSlice(sns, map[string]string{worker: "10.0.0.15"}) - - scheme := newTestScheme(t, - simplyblockv1alpha1.AddToScheme, - corev1.AddToScheme, - discoveryv1.AddToScheme, - ) - cl := newTestClient(t, scheme, nil, slice) - r := &StorageNodeOpsReconciler{Client: cl, Scheme: scheme, Recorder: events.NewFakeRecorder(16), apiReader: cl} - - // The enrolled worker is found — this is what the drifted name broke. - ok, err := r.endpointSliceHasWorker(context.Background(), ns, sns.Name, worker) - if err != nil { - t.Fatalf("endpointSliceHasWorker returned error: %v", err) - } - if !ok { - t.Fatalf("expected worker %q to be found in slice %q built for StorageNodeSet %q", worker, slice.Name, sns.Name) - } - - // A worker not published is not found. - ok, err = r.endpointSliceHasWorker(context.Background(), ns, sns.Name, "worker-0.ocp.simplyblock.ai") - if err != nil { - t.Fatalf("endpointSliceHasWorker returned error: %v", err) - } - if ok { - t.Fatal("did not expect an unpublished worker to be found") - } - - // A different StorageNodeSet name resolves to a different (absent) slice, so - // the worker is not found — guards the per-set name derivation. - ok, err = r.endpointSliceHasWorker(context.Background(), ns, "other-set", worker) - if err != nil { - t.Fatalf("endpointSliceHasWorker returned error: %v", err) - } - if ok { - t.Fatal("expected miss when querying under the wrong StorageNodeSet name") - } -} diff --git a/operator/internal/controller/storagenodeops_migrate_config_unit_test.go b/operator/internal/controller/storagenodeops_migrate_config_unit_test.go deleted file mode 100644 index b6decb1c5..000000000 --- a/operator/internal/controller/storagenodeops_migrate_config_unit_test.go +++ /dev/null @@ -1,182 +0,0 @@ -package controller - -import ( - "context" - "reflect" - "strings" - "testing" - - corev1 "k8s.io/api/core/v1" - metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" - "k8s.io/apimachinery/pkg/types" - - simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" -) - -const migrateSrcEnvFile = `MAX_SUBSYS_COUNT=10 -MAX_HUGE_PAGES_SIZE='' -VCPU_COUNT=8 -RESERVED_SYSTEM_CPUS='' -CPU_TOPOLOGY_ENABLED=true -PCI_ALLOWED='0000:02:00.0,0000:03:00.0' -PCI_BLOCKED='' -NVME_DEVICES='' -DEVICE_MODEL='' -SIZE_RANGE='' -JM_PERCENT= -HA_JM_COUNT= -` - -func TestMergePcieList(t *testing.T) { - got := mergePcieList([]string{"a", "b", ""}, []string{"b", "c", "c"}) - want := []string{"a", "b", "c"} - if !reflect.DeepEqual(got, want) { - t.Fatalf("mergePcieList = %v, want %v", got, want) - } -} - -func TestParseShellCSV(t *testing.T) { - cases := map[string][]string{ - `'0000:02:00.0,0000:03:00.0'`: {"0000:02:00.0", "0000:03:00.0"}, - `''`: nil, - ``: nil, - `0000:02:00.0`: {"0000:02:00.0"}, - `'a, b ,c'`: {"a", "b", "c"}, - } - for in, want := range cases { - if got := parseShellCSV(in); !reflect.DeepEqual(got, want) { - t.Errorf("parseShellCSV(%q) = %v, want %v", in, got, want) - } - } -} - -func TestMergePcieAllowedIntoEnvFile(t *testing.T) { - // Empty extra: unchanged. - if got := mergePcieAllowedIntoEnvFile(migrateSrcEnvFile, nil); got != migrateSrcEnvFile { - t.Fatalf("empty extra changed the env file:\n%s", got) - } - - // Merge new address, dedupe an already-present one, leave other lines intact. - got := mergePcieAllowedIntoEnvFile(migrateSrcEnvFile, - []string{"0000:03:00.0", "0000:04:00.0"}) - if !strings.Contains(got, `PCI_ALLOWED='0000:02:00.0,0000:03:00.0,0000:04:00.0'`) { - t.Fatalf("PCI_ALLOWED not merged as expected:\n%s", got) - } - // Only the PCI_ALLOWED line should differ. - for _, line := range strings.Split(migrateSrcEnvFile, "\n") { - if strings.HasPrefix(line, "PCI_ALLOWED=") || line == "" { - continue - } - if !strings.Contains(got, line) { - t.Errorf("line unexpectedly changed/removed: %q", line) - } - } - - // No PCI_ALLOWED line present: one is appended. - appended := mergePcieAllowedIntoEnvFile("MAX_SUBSYS_COUNT=10\n", []string{"0000:05:00.0"}) - if !strings.Contains(appended, "MAX_SUBSYS_COUNT=10") || - !strings.Contains(appended, `PCI_ALLOWED='0000:05:00.0'`) { - t.Fatalf("PCI_ALLOWED not appended:\n%s", appended) - } -} - -func TestEnsureMigratedWorkerConfig(t *testing.T) { - sns := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "simplyblock-node", Namespace: opsTestNS}, - } - cm := &corev1.ConfigMap{ - ObjectMeta: metav1.ObjectMeta{ - Name: PerNodeConfigMapName(sns.Name), - Namespace: opsTestNS, - }, - Data: map[string]string{"worker-1": migrateSrcEnvFile}, - } - r := newOpsReconciler(t, sns, cm) - ctx := context.Background() - - if err := r.ensureMigratedWorkerConfig(ctx, sns, "worker-1", "worker-4", - []string{"0000:04:00.0"}); err != nil { - t.Fatalf("ensureMigratedWorkerConfig: %v", err) - } - - var got corev1.ConfigMap - if err := r.Get(ctx, types.NamespacedName{Name: cm.Name, Namespace: opsTestNS}, &got); err != nil { - t.Fatalf("get cm: %v", err) - } - entry, ok := got.Data["worker-4"] - if !ok { - t.Fatal("worker-4 entry not created") - } - if !strings.Contains(entry, `PCI_ALLOWED='0000:02:00.0,0000:03:00.0,0000:04:00.0'`) { - t.Fatalf("worker-4 PCI_ALLOWED not merged:\n%s", entry) - } - - // Idempotent: a pre-existing target entry is left untouched. - got.Data["worker-4"] = "SENTINEL=1\n" - if err := r.Update(ctx, &got); err != nil { - t.Fatalf("seed sentinel: %v", err) - } - if err := r.ensureMigratedWorkerConfig(ctx, sns, "worker-1", "worker-4", nil); err != nil { - t.Fatalf("ensureMigratedWorkerConfig (idempotent): %v", err) - } - var again corev1.ConfigMap - _ = r.Get(ctx, types.NamespacedName{Name: cm.Name, Namespace: opsTestNS}, &again) - if again.Data["worker-4"] != "SENTINEL=1\n" { - t.Fatalf("existing target entry was overwritten: %q", again.Data["worker-4"]) - } - - // Missing source entry is an error. - if err := r.ensureMigratedWorkerConfig(ctx, sns, "worker-nope", "worker-5", nil); err == nil { - t.Fatal("expected error for missing source entry") - } -} - -func TestReconcileMigratedTopologyMigratesNodeConfig(t *testing.T) { - source, target := "worker-1", "worker-4" - sns := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "simplyblock-node", Namespace: opsTestNS}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - WorkerNodes: []string{source, "worker-2"}, - PcieAllowList: []string{"0000:02:00.0"}, - NodeConfigs: map[string]simplyblockv1alpha1.StorageNodeOverrides{ - source: {PcieAllowList: []string{"0000:02:00.0", "0000:03:00.0"}}, - }, - }, - } - sn := newTestStorageNode("simplyblock-node-x", opsTestNS, sns.Name, source, opsTestNodeUUID) - r := newOpsReconciler(t, sns, sn) - ctx := context.Background() - - if err := r.reconcileMigratedTopology(ctx, sn, sns, target, []string{"0000:04:00.0"}); err != nil { - t.Fatalf("reconcileMigratedTopology: %v", err) - } - - var fresh simplyblockv1alpha1.StorageNodeSet - if err := r.Get(ctx, types.NamespacedName{Name: sns.Name, Namespace: opsTestNS}, &fresh); err != nil { - t.Fatalf("get sns: %v", err) - } - - // Worker list: source dropped, target added. - if contains(fresh.Spec.WorkerNodes, source) || !contains(fresh.Spec.WorkerNodes, target) { - t.Fatalf("worker list not swapped: %v", fresh.Spec.WorkerNodes) - } - // nodeConfigs: source removed, target holds source's list + newSsdPcie. - if _, ok := fresh.Spec.NodeConfigs[source]; ok { - t.Error("source nodeConfig not removed") - } - tc, ok := fresh.Spec.NodeConfigs[target] - if !ok { - t.Fatal("target nodeConfig not set") - } - want := []string{"0000:02:00.0", "0000:03:00.0", "0000:04:00.0"} - if !reflect.DeepEqual(tc.PcieAllowList, want) { - t.Fatalf("target PcieAllowList = %v, want %v", tc.PcieAllowList, want) - } - - // StorageNode re-pointed at the target worker. - var freshSN simplyblockv1alpha1.StorageNode - _ = r.Get(ctx, types.NamespacedName{Name: sn.Name, Namespace: opsTestNS}, &freshSN) - if freshSN.Spec.WorkerNode != target { - t.Fatalf("StorageNode workerNode = %q, want %q", freshSN.Spec.WorkerNode, target) - } -} diff --git a/operator/internal/controller/test_helpers_test.go b/operator/internal/controller/test_helpers_test.go index 43d961734..1025dd935 100644 --- a/operator/internal/controller/test_helpers_test.go +++ b/operator/internal/controller/test_helpers_test.go @@ -12,10 +12,6 @@ import ( simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" ) -// statusSubresource is the name a client passes to a SubResourceUpdate -// interceptor for a status write, which is how a test makes one fail. -const statusSubresource = "status" - // newTestScheme builds a scheme carrying both simplyblock API versions, plus // whatever else the caller adds. // diff --git a/operator/internal/controllers/deployment/operatorops_controller.go b/operator/internal/controllers/deployment/operatorops_controller.go index e8c2ce6d0..461b39d10 100644 --- a/operator/internal/controllers/deployment/operatorops_controller.go +++ b/operator/internal/controllers/deployment/operatorops_controller.go @@ -42,7 +42,6 @@ import ( logf "sigs.k8s.io/controller-runtime/pkg/log" "github.com/simplyblock/atlas/inventory" - simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" discoverypkg "github.com/simplyblock/simplyblock-operator/internal/discovery" "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" @@ -256,7 +255,7 @@ func (r *OperatorOpsReconciler) workersAlreadyTaken( ctx context.Context, namespace string, ) (map[string]struct{}, error) { - var nodes simplyblockv1alpha1.StorageNodeList + var nodes simplyblockv1alpha2.StorageNodeList if err := r.List(ctx, &nodes, client.InNamespace(namespace)); err != nil { return nil, err } diff --git a/operator/internal/controllers/deployment/operatorops_unit_test.go b/operator/internal/controllers/deployment/operatorops_unit_test.go index fde17f24f..a3a9ef84d 100644 --- a/operator/internal/controllers/deployment/operatorops_unit_test.go +++ b/operator/internal/controllers/deployment/operatorops_unit_test.go @@ -275,9 +275,9 @@ func TestDiscoverSkipsWorkersAStorageNodeAlreadyRunsOn(t *testing.T) { // A run reports only what is unclaimed, which is what makes re-running it // useful: a run against a deployed fleet finds the machines nobody has // taken yet. - taken := &simplyblockv1alpha1.StorageNode{ + taken := &simplyblockv1alpha2.StorageNode{ ObjectMeta: metav1.ObjectMeta{Name: "sn-1", Namespace: opsNamespace}, - Spec: simplyblockv1alpha1.StorageNodeSpec{WorkerNode: "worker-1"}, + Spec: simplyblockv1alpha2.StorageNodeSpec{WorkerNode: "worker-1"}, } r := newRunner(t, discoverRun(nil), worker("worker-1"), worker("worker-2"), taken) diff --git a/operator/internal/controllers/node/actions.go b/operator/internal/controllers/node/actions.go new file mode 100644 index 000000000..5c3bff952 --- /dev/null +++ b/operator/internal/controllers/node/actions.go @@ -0,0 +1,155 @@ +// What each step of each action does, and how it knows it is finished. +// +// Every entry here obeys the same two rules, and both come from §7.2: +// +// - A step completes on a state, not on a transition. The "finished" half of +// each step is a predicate over what the control plane reports now, never an +// observation of a change, because a coalescing stream delivers current truth +// rather than an edit log and a node that moved through three states between +// two readings arrives as the last one once. +// - A call is skipped when its target is already at or past the state that call +// would produce. That is what makes a step recorded without its side effect +// having fired safe to re-enter: the re-entry reads the state and either acts +// or advances, so the question a `triggered` flag answers is one this +// controller never asks. +// +// The four single-step actions live here in full. The three multi-step ones are +// each a file of their own, because a drain, a relocation, and a maintenance +// window are each a workflow rather than a call. +// +// design-storagenode.md §7.3 is the specification for what is here. + +package node + +import ( + "context" + "fmt" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// perform runs one step of one action and reports whether it has finished. +// +// A step that has not finished is waiting on the control plane, and the caller +// requeues. An ordinary error is retried; a terminalStepError is not, and a +// blockedStepError holds with an event. +func (r *StorageNodeOpsReconciler) perform( + ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, current step, +) (bool, error) { + switch current { + case stepRequesting: + return r.request(ctx, ops) + case stepAwaiting: + return r.await(ctx, ops) + + case stepValidating, stepSuspending, stepMigratingVolumes, stepVerifying, stepRemoving: + return r.performRemoveStep(ctx, ops, current) + + case stepPreparing, stepRelocating, stepAwaitingNode, stepPromoting: + return r.performMigrateStep(ctx, ops, current) + + case stepHolding, stepShuttingDown, stepReleasing, stepAwaitingHost, + stepRestarting, stepCleanup: + return r.performMaintenanceStep(ctx, ops, current) + + default: + return false, fatalf("step %s belongs to no action this operator runs", current) + } +} + +// request issues the one call the four single-step actions make. Each is skipped +// when the node is already where the call would put it, which is what makes +// re-entering the step after a crash harmless. +func (r *StorageNodeOpsReconciler) request( + ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, +) (bool, error) { + clusterID, nodeID, err := r.target(ctx, ops) + if err != nil { + return false, err + } + reading, err := r.nodeReading(ctx, clusterID, nodeID) + if err != nil { + return false, err + } + + switch ops.Spec.Action { + case simplyblockv1alpha2.StorageNodeOpsActionShutdown: + if reading.Status == nodeStatusOffline { + return true, nil + } + if err := r.API.ShutdownNode(ctx, clusterID, nodeID); err != nil { + return false, fmt.Errorf("shut down node %s: %w", ops.Spec.NodeRef, err) + } + + case simplyblockv1alpha2.StorageNodeOpsActionRestart: + // A restart has no state of its own to skip on: a node is online before + // it and online after it. What guards the second call is the step record + // and the completion condition below, which does not report finished + // until the node is back. + params := RestartParams{ + Force: boolValue(ops.Spec.Force), + ReattachVolume: boolValue(ops.Spec.ReattachVolume), + } + if err := r.API.RestartNode(ctx, clusterID, nodeID, params); err != nil { + return false, fmt.Errorf("restart node %s: %w", ops.Spec.NodeRef, err) + } + + case simplyblockv1alpha2.StorageNodeOpsActionSuspend: + if reading.Status == nodeStatusSuspended { + return true, nil + } + if err := r.API.Suspend(ctx, clusterID, nodeID); err != nil { + return false, fmt.Errorf("suspend node %s: %w", ops.Spec.NodeRef, err) + } + + case simplyblockv1alpha2.StorageNodeOpsActionResume: + if reading.Status == nodeStatusOnline { + return true, nil + } + if err := r.API.Resume(ctx, clusterID, nodeID); err != nil { + return false, fmt.Errorf("resume node %s: %w", ops.Spec.NodeRef, err) + } + + default: + return false, fatalf("action %s does not issue a single request", ops.Spec.Action) + } + return true, nil +} + +// await is the completion condition of the four single-step actions, read against +// the control plane's current answer rather than against a transition it might +// have missed. +func (r *StorageNodeOpsReconciler) await( + ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, +) (bool, error) { + clusterID, nodeID, err := r.target(ctx, ops) + if err != nil { + return false, err + } + reading, err := r.nodeReading(ctx, clusterID, nodeID) + if err != nil { + return false, err + } + + wanted, ok := completionStatus[ops.Spec.Action] + if !ok { + return false, fatalf("action %s declares no completion condition", ops.Spec.Action) + } + return reading.Status == wanted, nil +} + +// completionStatus is what each single-step action waits for the node to report. +// The values are the control plane's own status strings, because a backend status +// is its vocabulary rather than this group's (§7.3). +var completionStatus = map[simplyblockv1alpha2.StorageNodeOpsAction]string{ + simplyblockv1alpha2.StorageNodeOpsActionShutdown: nodeStatusOffline, + simplyblockv1alpha2.StorageNodeOpsActionRestart: nodeStatusOnline, + simplyblockv1alpha2.StorageNodeOpsActionSuspend: nodeStatusSuspended, + simplyblockv1alpha2.StorageNodeOpsActionResume: nodeStatusOnline, +} + +// boolValue reads an optional flag, absent meaning false. Both flags this is used +// for are modifiers the control plane defaults itself, so not sending one is not +// the same as sending false — which is why the spec fields are pointers and only +// the value that was stated travels. +func boolValue(v *bool) bool { return v != nil && *v } diff --git a/operator/internal/controllers/node/classify.go b/operator/internal/controllers/node/classify.go new file mode 100644 index 000000000..8c786007b --- /dev/null +++ b/operator/internal/controllers/node/classify.go @@ -0,0 +1,309 @@ +// Sorting a node's backend volumes into the four buckets a drain acts on, and +// choosing where the movable ones go. +// +// It is separate from the drain's steps because it is a pure question about state +// — which volumes are on this node and what Kubernetes knows about each — and the +// steps are what act on the answer. Every one of the drain's five steps asks it, +// and two of them only ask it. +// +// design-storagenode.md §8.1 is the specification. + +package node + +import ( + "context" + "fmt" + "regexp" + "sort" + "strings" + + corev1 "k8s.io/api/core/v1" + "k8s.io/apimachinery/pkg/types" + logf "sigs.k8s.io/controller-runtime/pkg/log" + + "github.com/simplyblock/atlas/kube" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/cpinformer" + "github.com/simplyblock/simplyblock-operator/internal/utils" + "github.com/simplyblock/simplyblock-operator/internal/webapi" +) + +// The reason labels the blocked-volume gauge carries. They are lowercase because +// a metric label value is not an API enum. +const ( + blockedPinned = "pinned" + blockedUnmanaged = "unmanaged" +) + +// volumeCensus is what one classification pass found on a node. +// +// The four buckets are disjoint and every volume the control plane reports on the +// node is in exactly one of them, which is what makes "the drain is done" a +// statement about the census rather than about a counter. +type volumeCensus struct { + // Managed are the volumes a PersistentVolume accounts for and nothing pins, + // paired with the name of that PersistentVolume, because the migration is + // addressed by the Kubernetes object rather than by the backend volume. + Managed []managedVolume + + // Pinned carries a claim with the selected-storage-node annotation. It blocks: + // moving it would violate the pin, and the operator does not remove the + // annotation on the user's behalf, because a pin is a placement decision + // somebody made deliberately (§8.1). + Pinned []string + + // System matched spec.remove.systemVolumeFilterRegex. It is skipped by the + // migration and deleted during verification, because these are the + // rebalancer's own per-node benchmark volumes and moving one to a peer would + // produce a benchmark measuring the wrong node. + System []systemVolume + + // Unmanaged has no PersistentVolume behind it. It blocks, and blocking is the + // only safe answer: migrating it moves data nothing in Kubernetes is tracking, + // and deleting it destroys data nothing in Kubernetes is tracking. + Unmanaged []string + + // Incomplete says at least one volume could not be classified because its + // claim could not be read. Those volumes are counted as unmanaged for safety, + // and this flag is what stops the drain acting on a census it knows is a + // transient false positive. + Incomplete bool +} + +// managedVolume is one movable volume and the PersistentVolume that accounts for +// it. +type managedVolume struct { + VolumeUUID string + PVName string +} + +// systemVolume is one benchmark volume, carried with its pool because deleting a +// volume is addressed by both. +type systemVolume struct { + VolumeUUID string + PoolUUID string + Name string +} + +// classify lists every volume on one node and sorts it. +// +// It walks the cluster's pools because the control plane offers no per-node +// volume list, and it drops volumes already being deleted: the backend's deletion +// is asynchronous, so one stays in the list briefly after its DELETE returned. +func (r *StorageNodeOpsReconciler) classify( + ctx context.Context, + ops *simplyblockv1alpha2.StorageNodeOps, + clusterID, nodeID string, +) (volumeCensus, error) { + filter, err := systemVolumeFilter(ops) + if err != nil { + return volumeCensus{}, err + } + + pools, err := r.API.StoragePools(ctx, clusterID) + if err != nil { + return volumeCensus{}, fmt.Errorf("list the cluster's pools: %w", err) + } + + byVolumeUUID, err := r.persistentVolumesByVolumeUUID(ctx) + if err != nil { + return volumeCensus{}, err + } + + var census volumeCensus + for _, pool := range pools { + volumes, err := r.API.PoolVolumes(ctx, clusterID, pool.UUID) + if err != nil { + return volumeCensus{}, fmt.Errorf("list the volumes of pool %s: %w", pool.UUID, err) + } + for _, volume := range volumes { + if volume.PrimaryNodeUUID != nodeID || volume.Status == volumeStatusInDeletion { + continue + } + r.sortVolume(ctx, volume, pool.UUID, filter, byVolumeUUID, &census) + } + } + return census, nil +} + +// volumeStatusInDeletion is the control plane's own spelling for a volume whose +// delete has been accepted and not yet completed. +const volumeStatusInDeletion = "in_deletion" + +// sortVolume places one volume in the census. +func (r *StorageNodeOpsReconciler) sortVolume( + ctx context.Context, + volume webapi.VolumeInfo, + poolUUID string, + filter *regexp.Regexp, + byVolumeUUID map[string]*corev1.PersistentVolume, + census *volumeCensus, +) { + if filter.MatchString(volume.Name) { + census.System = append(census.System, systemVolume{ + VolumeUUID: volume.UUID, PoolUUID: poolUUID, Name: volume.Name, + }) + return + } + + pv, accounted := byVolumeUUID[volume.UUID] + if !accounted { + census.Unmanaged = append(census.Unmanaged, volume.UUID) + return + } + + // A PersistentVolume with no claim cannot be pinned, because the annotation + // lives on the claim. It is movable. + if pv.Spec.ClaimRef == nil { + census.Managed = append(census.Managed, + managedVolume{VolumeUUID: volume.UUID, PVName: pv.Name}) + return + } + + var claim corev1.PersistentVolumeClaim + key := types.NamespacedName{ + Namespace: pv.Spec.ClaimRef.Namespace, + Name: pv.Spec.ClaimRef.Name, + } + if err := r.Get(ctx, key, &claim); err != nil { + // Counted as unmanaged, which blocks, and recorded as incomplete, which + // is what tells the caller this is a transient false positive rather than + // a volume nothing accounts for. An API server that is briefly away must + // not be read as permission to move a volume nobody could classify. + logf.FromContext(ctx).V(1).Info("a volume's claim could not be read; the census is incomplete", + "volume", volume.UUID, "claim", key.String(), "err", err.Error()) + census.Unmanaged = append(census.Unmanaged, volume.UUID) + census.Incomplete = true + return + } + + if kube.IsPinnedVolume(claim.Annotations) { + census.Pinned = append(census.Pinned, volume.UUID) + return + } + census.Managed = append(census.Managed, + managedVolume{VolumeUUID: volume.UUID, PVName: pv.Name}) +} + +// persistentVolumesByVolumeUUID indexes every simplyblock PersistentVolume in the +// cluster by the backend volume it names. +// +// The handle is clusterUUID:poolUUID:volumeUUID and the last segment is the +// volume, which is what makes the map key the same identity the control plane +// reports. +func (r *StorageNodeOpsReconciler) persistentVolumesByVolumeUUID( + ctx context.Context, +) (map[string]*corev1.PersistentVolume, error) { + var volumes corev1.PersistentVolumeList + if err := r.List(ctx, &volumes); err != nil { + return nil, fmt.Errorf("list persistent volumes: %w", err) + } + out := make(map[string]*corev1.PersistentVolume, len(volumes.Items)) + for i := range volumes.Items { + pv := &volumes.Items[i] + if pv.Spec.CSI == nil || pv.Spec.CSI.Driver != utils.CSIProvisioner { + continue + } + handle := pv.Spec.CSI.VolumeHandle + if handle == "" { + continue + } + parts := strings.SplitN(handle, ":", 3) + if volumeUUID := parts[len(parts)-1]; volumeUUID != "" { + out[volumeUUID] = pv + } + } + return out, nil +} + +// systemVolumeFilter compiles the operation's system-volume pattern. +// +// A pattern that does not compile is fatal rather than retried: the expression is +// in the spec and no number of passes will make it parse. The default is applied +// by the CRD, so an operation that states nothing arrives with the benchmark +// pattern already in place, and the fallback here covers an object written before +// the default existed. +func systemVolumeFilter(ops *simplyblockv1alpha2.StorageNodeOps) (*regexp.Regexp, error) { + pattern := defaultSystemVolumePattern + if stated := ops.Spec.RemoveParams().SystemVolumeFilterRegex; stated != nil && *stated != "" { + pattern = *stated + } + filter, err := regexp.Compile(pattern) + if err != nil { + return nil, fatalf("spec.remove.systemVolumeFilterRegex does not compile: %v", err) + } + return filter, nil +} + +// defaultSystemVolumePattern matches the rebalancer's benchmark volumes by name, +// which is a convention rather than a guarantee: a volume a user happens to name +// this way is deleted during a drain's verification, and §16 Q1 is whether a label +// applied at creation should replace it. +const defaultSystemVolumePattern = `^sb-fio-baseline-.*` + +// peerTargets assigns each movable volume an online peer to move to, round-robin. +// +// Round-robin spreads the drained node's volumes rather than concentrating them on +// whichever peer sorts first. The order is the peers' own UUIDs sorted, so the +// assignment is stable across passes: a volume that was assigned to one peer and +// whose migration then failed is reassigned by the caller deliberately rather than +// by the list having reshuffled. +// +// A drain with no online peer is a stall rather than a failure, which is why this +// reports a blockedStepError: the condition is resolved by another node coming +// back, and failing the operation would only mean starting it again afterward +// (§8.2). +func (r *StorageNodeOpsReconciler) peerTargets( + ctx context.Context, clusterID, nodeID string, volumes []managedVolume, +) (map[string]string, error) { + readings, err := r.clusterNodes(ctx, clusterID) + if err != nil { + return nil, err + } + + peers := make([]string, 0, len(readings)) + for _, reading := range readings { + if reading.UUID == nodeID || reading.Status != nodeStatusOnline { + continue + } + peers = append(peers, reading.UUID) + } + if len(peers) == 0 { + return nil, blockedf(NoMigrationTarget, + "no online peer to move this node's volumes to; the drain resumes when one returns") + } + sort.Strings(peers) + + targets := make(map[string]string, len(volumes)) + for i, volume := range volumes { + targets[volume.PVName] = peers[i%len(peers)] + } + return targets, nil +} + +// clusterNodes returns every backend node of the cluster, from the stream's cache +// once it has delivered the cluster's snapshot and from the control plane until +// then. +// +// The gate is the snapshot rather than a preference: an empty unsynced cache and a +// cluster with no nodes look identical, and reading the first as the second would +// report a drain with no peers when every peer is there. +func (r *StorageNodeOpsReconciler) clusterNodes( + ctx context.Context, clusterID string, +) ([]NodeReading, error) { + if r.Nodes != nil && r.Nodes.Synced(scopeOf(clusterID)) { + cached := r.Nodes.List(scopeOf(clusterID)) + out := make([]NodeReading, 0, len(cached)) + for _, dto := range cached { + out = append(out, readingFromDTO(dto)) + } + return out, nil + } + return r.API.StorageNodes(ctx, clusterID) +} + +// scopeOf is the storage-node stream's scope for one cluster. Every subscription +// in this package is per cluster, so the scope is the cluster's UUID and nothing +// else. +func scopeOf(clusterID string) cpinformer.Scope { return cpinformer.Scope{clusterID} } diff --git a/operator/internal/controllers/node/controlplane.go b/operator/internal/controllers/node/controlplane.go new file mode 100644 index 000000000..f7c2615d3 --- /dev/null +++ b/operator/internal/controllers/node/controlplane.go @@ -0,0 +1,309 @@ +// The control-plane surface this package needs, and the HTTP client that +// satisfies it. +// +// It is an interface so that a test can drive both controllers, including a +// whole drain and a whole relocation, without an HTTP server, and so that the +// one place a URL is spelled is the implementation below rather than scattered +// through the reconcilers. What it declares is the operator's question rather +// than the client's vocabulary: "suspend this node" rather than "POST this path." +// +// The endpoints are design-storagenode.md §12. Two of them are specified there as +// a `?watch=true` subscription, and the storage-node stream serves both: what +// remains here is the fallback each reader takes until its scope reports synced, +// plus the calls that have no streamed counterpart at all — the add, the delete, +// the five actions, and the volume list a drain classifies. +// +// A predicate over current state reads the same whether the state arrived by +// stream or by poll (design-crd-model.md §7.7), which is what lets a step's +// completion condition be written once. + +package node + +import ( + "context" + "encoding/json" + "errors" + "fmt" + "net/http" + + "github.com/simplyblock/simplyblock-operator/internal/utils" + "github.com/simplyblock/simplyblock-operator/internal/webapi" +) + +// NodeReading is one backend storage node as the control plane reports it. It is +// this package's own type rather than the client's, because every completion +// condition in the package is a predicate over it and the fields it needs are the +// ones named here. +type NodeReading struct { + UUID string `json:"id"` + Status string `json:"status"` + ManagementIP string `json:"mgmt_ip"` + Health bool `json:"health_check"` + Hostname string `json:"hostname"` + Uptime string `json:"uptime"` + DevicesCount int32 `json:"device_count"` + OnlineDevicesCount int32 `json:"online_device_count"` + CPUCount int32 `json:"cpu_spdk_count"` + Memory int64 `json:"spdk_mem"` + Volumes int32 `json:"lvols"` + RPCPort int32 `json:"rpc_port"` + LvolPort int32 `json:"lvol_subsys_port"` + NVMeOFPort int32 `json:"nvmf_port"` + + // FailureDomain is the group the control plane actually assigned. It is an + // integer on the wire and a label on the object, so the status write renders + // the digits: the control plane's own vocabulary is what it is, and the + // operator does not invent a name the control plane never said (§3.3). + FailureDomain int `json:"failure_domain"` +} + +// The lifecycle values the control plane reports, in its own spelling. They are +// neither PascalCase nor an API enum for that reason: a backend status is the +// control plane's vocabulary rather than this group's (§7.3). +const ( + nodeStatusOnline = "online" + nodeStatusSuspended = "suspended" + nodeStatusOffline = "offline" + nodeStatusInCreation = "in_creation" + nodeStatusInRestart = "in_restart" + nodeStatusActive = "active" +) + +// RestartParams are what the control plane's restart endpoint takes. Three +// actions use it and each fills a different subset: a plain Restart passes only +// the two flags, a Migrate passes the target's address and any drives being +// bound on it, and a HostMaintenance passes neither address nor drives because +// the node is coming back on the host it left. +type RestartParams struct { + // NodeAddress is the per-pod DNS name the control plane resolves itself. A + // name that does not resolve makes the restart fail inside the control plane, + // whose response is to reset the node to offline, which is why the migration + // blocks on the EndpointSlice before this is sent (§5.4). + NodeAddress string `json:"node_address,omitempty"` + Force bool `json:"force,omitempty"` + ReattachVolume bool `json:"reattach_volume,omitempty"` + NewSsdPcie []string `json:"new_ssd_pcie,omitempty"` +} + +// ControlPlane is everything the two reconcilers in this package ask of the +// simplyblock control plane. +type ControlPlane interface { + // AddNode adds every storage node of one worker at once and is not + // idempotent, which is why the provisioning machine claims its slot in + // Kubernetes before calling it (§4.2). + AddNode(ctx context.Context, clusterID string, params utils.StorageNodeSetAddParams) error + + // StorageNodes are the cluster's nodes as the control plane reports them, + // which is what adoption matches against and what the fallback of every + // completion condition reads. It is the poll behind the storage-node stream. + StorageNodes(ctx context.Context, clusterID string) ([]NodeReading, error) + + // StorageNode reads one node by its UUID. A node the control plane no longer + // knows about reports false rather than an error, because a stored UUID that + // has gone is the ordinary answer after a cluster was reset. + StorageNode(ctx context.Context, clusterID, nodeID string) (NodeReading, bool, error) + + // The five node actions. Each must tolerate a repeat, because a step + // recorded without its call having fired re-issues it (§7.2). + Suspend(ctx context.Context, clusterID, nodeID string) error + Resume(ctx context.Context, clusterID, nodeID string) error + ShutdownNode(ctx context.Context, clusterID, nodeID string) error + RestartNode(ctx context.Context, clusterID, nodeID string, params RestartParams) error + + // Promote is the migration's last control-plane call and the one that cannot + // be undone: it activates the target host's devices, fails and migrates the + // origin host's, starts a rebalance, and re-homes the logical volumes (§9). + Promote(ctx context.Context, clusterID, nodeID string) error + + // RemoveNode is the drain's last step. A 404 is success: a node the control + // plane no longer knows about is a node that has been removed, and a retry + // after a lost response is the common way to arrive there (§8.2). + RemoveNode(ctx context.Context, clusterID, nodeID string) error + + // StoragePools are the cluster's pools, which is the list a drain walks to + // find the volumes on one node. + StoragePools(ctx context.Context, clusterID string) ([]webapi.StoragePoolInfo, error) + + // PoolVolumes are the volumes of one pool. The control plane offers no + // per-node volume list, so a drain reads every pool and keeps the volumes + // whose node is the one being removed. + PoolVolumes(ctx context.Context, clusterID, poolID string) ([]webapi.VolumeInfo, error) + + // DeleteVolume removes one volume, which verification does to the system + // volumes a drain skipped. A 404 is success, for the reason RemoveNode's is. + DeleteVolume(ctx context.Context, clusterID, poolID, volumeID string) error +} + +// httpControlPlane is the ControlPlane the operator runs with: the shared webapi +// client, with one method per endpoint of §12. +type httpControlPlane struct{ client *webapi.Client } + +// NewControlPlane returns the HTTP-backed control-plane surface. +func NewControlPlane() ControlPlane { return &httpControlPlane{client: webapi.NewClient()} } + +func (c *httpControlPlane) AddNode( + ctx context.Context, clusterID string, params utils.StorageNodeSetAddParams, +) error { + return c.post(ctx, fmt.Sprintf("/api/v2/clusters/%s/storage-nodes", clusterID), params) +} + +func (c *httpControlPlane) StorageNodes( + ctx context.Context, clusterID string, +) ([]NodeReading, error) { + body, err := c.call(ctx, http.MethodGet, + fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/", clusterID), nil) + if err != nil { + return nil, err + } + var nodes []NodeReading + if err := json.Unmarshal(body, &nodes); err != nil { + return nil, fmt.Errorf("read the storage node list: %w", err) + } + return nodes, nil +} + +func (c *httpControlPlane) StorageNode( + ctx context.Context, clusterID, nodeID string, +) (NodeReading, bool, error) { + body, err := c.call(ctx, http.MethodGet, + fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s", clusterID, nodeID), nil) + var refusal *ControlPlaneError + if errors.As(err, &refusal) && refusal.Status == http.StatusNotFound { + return NodeReading{}, false, nil + } + if err != nil { + return NodeReading{}, false, err + } + var node NodeReading + if err := json.Unmarshal(body, &node); err != nil { + return NodeReading{}, false, fmt.Errorf("read storage node %s: %w", nodeID, err) + } + return node, true, nil +} + +func (c *httpControlPlane) Suspend(ctx context.Context, clusterID, nodeID string) error { + return c.post(ctx, c.nodePath(clusterID, nodeID, "suspend"), nil) +} + +func (c *httpControlPlane) Resume(ctx context.Context, clusterID, nodeID string) error { + return c.post(ctx, c.nodePath(clusterID, nodeID, "resume"), nil) +} + +func (c *httpControlPlane) ShutdownNode(ctx context.Context, clusterID, nodeID string) error { + return c.post(ctx, c.nodePath(clusterID, nodeID, "shutdown"), nil) +} + +func (c *httpControlPlane) RestartNode( + ctx context.Context, clusterID, nodeID string, params RestartParams, +) error { + return c.post(ctx, c.nodePath(clusterID, nodeID, "restart"), params) +} + +func (c *httpControlPlane) Promote(ctx context.Context, clusterID, nodeID string) error { + return c.post(ctx, c.nodePath(clusterID, nodeID, "promote"), nil) +} + +// RemoveNode does not force the removal. The drain has already moved every +// volume off the node and verified there are none left, so a removal the control +// plane refuses is one of its own admission checks saying the cluster cannot +// afford to lose this node — which is an answer to report rather than to +// override (§8.2). +func (c *httpControlPlane) RemoveNode(ctx context.Context, clusterID, nodeID string) error { + err := c.delete(ctx, fmt.Sprintf( + "/api/v2/clusters/%s/storage-nodes/%s?force_remove=false", clusterID, nodeID)) + return ignoreGone(err) +} + +func (c *httpControlPlane) StoragePools( + ctx context.Context, clusterID string, +) ([]webapi.StoragePoolInfo, error) { + body, err := c.call(ctx, http.MethodGet, + fmt.Sprintf("/api/v2/clusters/%s/storage-pools/", clusterID), nil) + if err != nil { + return nil, err + } + var pools []webapi.StoragePoolInfo + if err := json.Unmarshal(body, &pools); err != nil { + return nil, fmt.Errorf("read the storage pool list: %w", err) + } + return pools, nil +} + +func (c *httpControlPlane) PoolVolumes( + ctx context.Context, clusterID, poolID string, +) ([]webapi.VolumeInfo, error) { + body, err := c.call(ctx, http.MethodGet, fmt.Sprintf( + "/api/v2/clusters/%s/storage-pools/%s/volumes/", clusterID, poolID), nil) + if err != nil { + return nil, err + } + var volumes []webapi.VolumeInfo + if err := json.Unmarshal(body, &volumes); err != nil { + return nil, fmt.Errorf("read the volumes of pool %s: %w", poolID, err) + } + return volumes, nil +} + +func (c *httpControlPlane) DeleteVolume( + ctx context.Context, clusterID, poolID, volumeID string, +) error { + err := c.delete(ctx, fmt.Sprintf( + "/api/v2/clusters/%s/storage-pools/%s/volumes/%s/", clusterID, poolID, volumeID)) + return ignoreGone(err) +} + +// nodePath is the action endpoint of one node. The segments keep the control +// plane's own lowercase spelling, because a URL is its vocabulary rather than +// this group's. +func (c *httpControlPlane) nodePath(clusterID, nodeID, action string) string { + return fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/%s", clusterID, nodeID, action) +} + +// post issues a call whose body is not read, which every action endpoint is. +func (c *httpControlPlane) post(ctx context.Context, path string, body any) error { + _, err := c.call(ctx, http.MethodPost, path, body) + return err +} + +func (c *httpControlPlane) delete(ctx context.Context, path string) error { + _, err := c.call(ctx, http.MethodDelete, path, nil) + return err +} + +// call performs one request and turns a non-2xx into a ControlPlaneError carrying +// the status and the body, which is what makes a refusal visible in +// `kubectl describe` without reading the operator's log. +func (c *httpControlPlane) call( + ctx context.Context, method, path string, body any, +) ([]byte, error) { + response, status, err := c.client.Do(ctx, method, path, body) + if err != nil { + return nil, fmt.Errorf("%s %s: %w", method, path, err) + } + if status >= 300 { + return nil, &ControlPlaneError{Status: status, Body: string(response)} + } + return response, nil +} + +// ignoreGone reads a 404 as success. A thing the control plane no longer knows +// about is the outcome the delete asked for, reached either by this call's own +// lost response or by somebody else (§8.2). +func ignoreGone(err error) error { + var refusal *ControlPlaneError + if errors.As(err, &refusal) && refusal.Status == http.StatusNotFound { + return nil + } + return err +} + +// ControlPlaneError is a request the control plane refused. It carries the status +// and the whole body, because the body is where the control plane says why. +type ControlPlaneError struct { + Status int + Body string +} + +func (e *ControlPlaneError) Error() string { + return fmt.Sprintf("the control plane answered %d: %s", e.Status, e.Body) +} diff --git a/operator/internal/controllers/node/events.go b/operator/internal/controllers/node/events.go new file mode 100644 index 000000000..716585662 --- /dev/null +++ b/operator/internal/controllers/node/events.go @@ -0,0 +1,86 @@ +// The reasons this package emits events under. +// +// They are collected here rather than declared beside the code that raises them +// because a reason is a contract with whoever is reading `kubectl describe`: it +// is what an administrator greps for and what an alert matches on, so it outlives +// the function that happens to raise it today. +// +// Events need a target object and this pairing has two candidates that are both +// right for different things. An event about the node's own lifecycle goes on the +// StorageNode, which is what an administrator looking at a worker has open. An +// event about an operation goes on the StorageNodeOps, which outlives the +// operation as its audit record. An operation's events are mirrored onto its +// target node as well, because the node is where someone investigating a stuck +// cluster starts and the operation's name is not something they know yet. +// +// design-storagenode.md §13.1 is the specification. + +package node + +const ( + // The node's own lifecycle, raised on the StorageNode. + // + // The three holding reasons are the load-bearing ones. A provisioning node + // waiting for a slot, one waiting for a fault group, and one whose worker + // does not answer are all correct behavior that looks exactly like a stalled + // controller, and the event is the only thing that distinguishes them. + ClusterNotReady = "ClusterNotReady" + FailureDomainMissing = "FailureDomainMissing" + HostUnreachable = "HostUnreachable" + AwaitingSlot = "AwaitingSlot" + + // NodeAdopted says an existing backend node was taken over rather than + // added, which is the difference between a migration and a mistake. + NodeAdopted = "NodeAdopted" + + // NodeOnline says the node came up and is carrying its share. + NodeOnline = "NodeOnline" + + // PodSchedulingFailed says the node's storage pod cannot be placed, which is + // the one Kubernetes-side failure a node's own status cannot show. + PodSchedulingFailed = "PodSchedulingFailed" + + // The operation reasons, raised on the StorageNodeOps rather than the node. + // + // OperationSucceeded is one reason for all seven actions rather than one + // each: the action is already in spec.action and on a print column, so + // encoding it in the reason name tells a reader nothing and gives anyone + // alerting on completion seven reasons to match instead of one. + OperationQueued = "OperationQueued" + OperationStarted = "OperationStarted" + OperationSucceeded = "OperationSucceeded" + OperationFailed = "OperationFailed" + OperationAborted = "OperationAborted" + + // StepDeadlineExceeded distinguishes an operation still working from one + // that stopped, which is the distinction status.message cannot express. + StepDeadlineExceeded = "StepDeadlineExceeded" + + // DrainBlocked is one reason for two conditions, with the volume names and + // the resolution in the message. A pinned volume and an unmanaged one are the + // same situation from an alerting perspective — a drain that will not proceed + // until somebody acts — and the difference is what the message says to do. + DrainBlocked = "DrainBlocked" + + // NoMigrationTarget says a drain has no online peer to move volumes to. It is + // a stall rather than a failure: the condition is resolved by another node + // coming back, and failing the operation would only mean starting it again + // afterward. + NoMigrationTarget = "NoMigrationTarget" + + // MigrationRetried says one volume's move failed and is being retried + // against a fresh target. + MigrationRetried = "MigrationRetried" + + // DrainCompleted says every volume has been migrated off the node. + DrainCompleted = "DrainCompleted" + + // NodeResumeFailed is the one that cannot be retried away. The unwind of §8.3 + // is best-effort, so a resume that fails leaves a node suspended and out of + // service, and this event is the only place that is visible. + NodeResumeFailed = "NodeResumeFailed" + + // MaintenanceQueued says a maintenance window is holding for another worker, + // which is correct behavior and looks like a stalled controller without it. + MaintenanceQueued = "MaintenanceQueued" +) diff --git a/operator/internal/controllers/node/graphs.go b/operator/internal/controllers/node/graphs.go new file mode 100644 index 000000000..225695c9a --- /dev/null +++ b/operator/internal/controllers/node/graphs.go @@ -0,0 +1,358 @@ +// The state graph of every StorageNodeOps action, and of the StorageNode's own +// provisioning path, both declared as data. +// +// One status.step field serves all seven operations, so nothing in the API type +// prevents a Remove from reporting Promoting. The per-action graph makes that an +// IllegalTransitionError at the point of the write rather than an accepted status +// (design-crd-model.md §3.1), which is the whole reason the steps are declared +// here instead of switched on in the reconciler. +// +// MultiConfig validates every declared graph whenever a machine is built for any +// of them, so a bad edge in HostMaintenance — the action that only runs during an +// OS upgrade and is the most expensive to exercise — is caught by any test that +// builds a machine at all. +// +// The entity's graph is a Config rather than a MultiConfig, because a StorageNode +// has no spec.action to key one on: there is one provisioning path and adoption +// is a branch within it. +// +// design-storagenode.md §4.2, §6.3, §8.2, §9, and §10 are the specification. + +package node + +import ( + "context" + "time" + + "github.com/simplyblock/atlas/statemachine" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// step is the operation's step type, aliased so the graph literals below read as +// the graphs rather than as a wall of package qualifiers. +type step = simplyblockv1alpha2.StorageNodeOpsStep + +const ( + stepRequesting = simplyblockv1alpha2.StorageNodeOpsStepRequesting + stepAwaiting = simplyblockv1alpha2.StorageNodeOpsStepAwaiting + stepValidating = simplyblockv1alpha2.StorageNodeOpsStepValidating + stepSuspending = simplyblockv1alpha2.StorageNodeOpsStepSuspending + stepMigratingVolumes = simplyblockv1alpha2.StorageNodeOpsStepMigratingVolumes + stepVerifying = simplyblockv1alpha2.StorageNodeOpsStepVerifying + stepRemoving = simplyblockv1alpha2.StorageNodeOpsStepRemoving + stepPreparing = simplyblockv1alpha2.StorageNodeOpsStepPreparing + stepRelocating = simplyblockv1alpha2.StorageNodeOpsStepRelocating + stepAwaitingNode = simplyblockv1alpha2.StorageNodeOpsStepAwaitingNode + stepPromoting = simplyblockv1alpha2.StorageNodeOpsStepPromoting + stepHolding = simplyblockv1alpha2.StorageNodeOpsStepHolding + stepShuttingDown = simplyblockv1alpha2.StorageNodeOpsStepShuttingDown + stepReleasing = simplyblockv1alpha2.StorageNodeOpsStepReleasing + stepAwaitingHost = simplyblockv1alpha2.StorageNodeOpsStepAwaitingHost + stepRestarting = simplyblockv1alpha2.StorageNodeOpsStepRestarting + stepCleanup = simplyblockv1alpha2.StorageNodeOpsStepCleanup +) + +// nodeStep is the entity's step type, aliased for the same reason. +type nodeStep = simplyblockv1alpha2.StorageNodeStep + +const ( + stepCheckingHost = simplyblockv1alpha2.StorageNodeStepCheckingHost + stepCheckingConfig = simplyblockv1alpha2.StorageNodeStepCheckingConfig + stepAwaitingSlot = simplyblockv1alpha2.StorageNodeStepAwaitingSlot + stepPosting = simplyblockv1alpha2.StorageNodeStepPosting + stepResolving = simplyblockv1alpha2.StorageNodeStepResolving + stepAdopting = simplyblockv1alpha2.StorageNodeStepAdopting +) + +// How long each operation step may take before it is reported as stuck. +// +// They differ by what the step is waiting on rather than by preference. A request +// is one HTTP call. A node coming back from a restart is a data-plane operation +// across every device it owns. Two are deliberately generous and finite, and §16 +// records that neither number is known to be right: Validating holds on a human +// removing a pin, and MigratingVolumes holds on a node's worth of volumes moving, +// which on a hundred large volumes is hours. What a deadline separates there is a +// drain waiting by design from one waiting because of a bug, and without one the +// two look identical. +const ( + requestingDeadline = 2 * time.Minute + awaitingDeadline = 30 * time.Minute + validatingDeadline = 24 * time.Hour + suspendingDeadline = 15 * time.Minute + migratingDeadline = 12 * time.Hour + verifyingDeadline = 30 * time.Minute + removingDeadline = 30 * time.Minute + preparingDeadline = 15 * time.Minute + relocatingDeadline = 15 * time.Minute + nodeRestartDeadline = 45 * time.Minute + promotingDeadline = 30 * time.Minute + holdingDeadline = 6 * time.Hour + releasingDeadline = 15 * time.Minute + + // awaitingHostDeadline is the step nobody controls the length of: an OS + // upgrade and a reboot take as long as they take, and a firmware update is + // the case that sets the number. An expiry fails the operation and leaves the + // node offline, needing a Restart to recover, so this is a detection + // mechanism rather than a recovery one (§10). + awaitingHostDeadline = 4 * time.Hour + + cleanupDeadline = 5 * time.Minute +) + +// How long each step of the entity's provisioning path may take. +// +// Resolving is the one that matters most: a node add the control plane accepted +// and then failed to complete leaves the object polling for a UUID that never +// arrives, which is the failure mode status.status: timeout names today with no +// bound behind it (§4.2). +const ( + checkingHostDeadline = 30 * time.Minute + checkingConfigDeadline = 24 * time.Hour + awaitingSlotDeadline = 4 * time.Hour + postingDeadline = 10 * time.Minute + resolvingDeadline = 45 * time.Minute + adoptingDeadline = 10 * time.Minute +) + +// deadline is the entry hook every state here carries: it sets the step's budget +// and performs nothing. The side effect of a step is performed on the pass that +// follows, against the step the entry's patch persisted, which is where the +// write-ahead record is needed and what it records. +func deadline[S comparable](d time.Duration) statemachine.TransitionFunc[S] { + return func(context.Context, S, S) (time.Duration, error) { return d, nil } +} + +// graphs declares one state graph per operation action over one step type. +func graphs() statemachine.MultiConfig[step] { + // requestAndWait is the two-step line the four single-step actions share: + // post the action, then wait for the completion condition. It is a function + // rather than a shared value because MultiConfig copies the graph when a + // machine is built and a shared map would be one graph under four keys. + requestAndWait := func() statemachine.Config[step] { + return statemachine.Config[step]{ + Initial: stepRequesting, + States: map[step]statemachine.StateDef[step]{ + stepRequesting: {To: []step{stepAwaiting}, OnEnter: deadline[step](requestingDeadline)}, + stepAwaiting: {OnEnter: deadline[step](awaitingDeadline)}, + }, + } + } + + return statemachine.MultiConfig[step]{ + action(simplyblockv1alpha2.StorageNodeOpsActionShutdown): requestAndWait(), + action(simplyblockv1alpha2.StorageNodeOpsActionRestart): requestAndWait(), + action(simplyblockv1alpha2.StorageNodeOpsActionSuspend): requestAndWait(), + action(simplyblockv1alpha2.StorageNodeOpsActionResume): requestAndWait(), + + // Validation runs before the suspend, and that ordering is the design: a + // suspended node accepts no new volume placement, so suspending one whose + // drain cannot complete takes capacity out of the cluster and leaves it + // out for as long as the blocker goes unnoticed (§8.2). + action(simplyblockv1alpha2.StorageNodeOpsActionRemove): { + Initial: stepValidating, + States: map[step]statemachine.StateDef[step]{ + stepValidating: { + To: []step{stepSuspending}, + OnEnter: deadline[step](validatingDeadline), + }, + stepSuspending: { + To: []step{stepMigratingVolumes}, + OnEnter: deadline[step](suspendingDeadline), + }, + stepMigratingVolumes: { + To: []step{stepVerifying}, + OnEnter: deadline[step](migratingDeadline), + }, + stepVerifying: { + To: []step{stepRemoving}, + OnEnter: deadline[step](verifyingDeadline), + }, + stepRemoving: {OnEnter: deadline[step](removingDeadline)}, + }, + }, + + // Relocating and AwaitingNode are two steps because one would race. The + // restart is asynchronous, so a node still reporting online immediately + // after the call may be reporting the state from before it: Relocating + // completes when the node has left online, and AwaitingNode when it is + // back. Collapsing them means /promote can be issued while the restart's + // own node writes are in flight, which leaves the relocated devices stuck + // in `new` (§9). + action(simplyblockv1alpha2.StorageNodeOpsActionMigrate): { + Initial: stepPreparing, + States: map[step]statemachine.StateDef[step]{ + stepPreparing: { + To: []step{stepRelocating}, + OnEnter: deadline[step](preparingDeadline), + }, + stepRelocating: { + To: []step{stepAwaitingNode}, + OnEnter: deadline[step](relocatingDeadline), + }, + stepAwaitingNode: { + To: []step{stepPromoting}, + OnEnter: deadline[step](nodeRestartDeadline), + }, + stepPromoting: {OnEnter: deadline[step](promotingDeadline)}, + }, + }, + + action(simplyblockv1alpha2.StorageNodeOpsActionHostMaintenance): { + Initial: stepHolding, + States: map[step]statemachine.StateDef[step]{ + stepHolding: { + To: []step{stepShuttingDown}, + OnEnter: deadline[step](holdingDeadline), + }, + stepShuttingDown: { + To: []step{stepReleasing}, + OnEnter: deadline[step](suspendingDeadline), + }, + stepReleasing: { + To: []step{stepAwaitingHost}, + OnEnter: deadline[step](releasingDeadline), + }, + stepAwaitingHost: { + To: []step{stepRestarting}, + OnEnter: deadline[step](awaitingHostDeadline), + }, + stepRestarting: { + To: []step{stepCleanup}, + OnEnter: deadline[step](nodeRestartDeadline), + }, + stepCleanup: {OnEnter: deadline[step](cleanupDeadline)}, + }, + }, + } +} + +// provisioningGraph is the entity's own machine (§4.2). +// +// CheckingHost declares three successors because adoption diverts from it: an +// upgrade Secret or a backend node already at the worker's address sends the node +// to Adopting, and everything else continues to the configuration gate. Adopting +// and Resolving are both terminal, because both end with a UUID on the object and +// the node in steady state. +func provisioningGraph() statemachine.Config[nodeStep] { + return statemachine.Config[nodeStep]{ + Initial: stepCheckingHost, + States: map[nodeStep]statemachine.StateDef[nodeStep]{ + stepCheckingHost: { + To: []nodeStep{stepCheckingConfig, stepAdopting}, + OnEnter: deadline[nodeStep](checkingHostDeadline), + }, + stepCheckingConfig: { + To: []nodeStep{stepAwaitingSlot, stepAdopting}, + OnEnter: deadline[nodeStep](checkingConfigDeadline), + }, + stepAwaitingSlot: { + To: []nodeStep{stepPosting, stepResolving}, + OnEnter: deadline[nodeStep](awaitingSlotDeadline), + }, + stepPosting: { + To: []nodeStep{stepResolving}, + OnEnter: deadline[nodeStep](postingDeadline), + }, + stepResolving: {OnEnter: deadline[nodeStep](resolvingDeadline)}, + stepAdopting: {OnEnter: deadline[nodeStep](adoptingDeadline)}, + }, + } +} + +// initialDeadlines are the budgets of the step each action's machine is born in. +// A machine is already in its initial state when it is built, so that state's +// OnEnter never runs and the graph's deadline for it is never set. Setting it +// explicitly is what stops the first step of every operation from being the one +// step that cannot time out. +var initialDeadlines = map[statemachine.Action]time.Duration{ + action(simplyblockv1alpha2.StorageNodeOpsActionShutdown): requestingDeadline, + action(simplyblockv1alpha2.StorageNodeOpsActionRestart): requestingDeadline, + action(simplyblockv1alpha2.StorageNodeOpsActionSuspend): requestingDeadline, + action(simplyblockv1alpha2.StorageNodeOpsActionResume): requestingDeadline, + action(simplyblockv1alpha2.StorageNodeOpsActionRemove): validatingDeadline, + action(simplyblockv1alpha2.StorageNodeOpsActionMigrate): preparingDeadline, + action(simplyblockv1alpha2.StorageNodeOpsActionHostMaintenance): holdingDeadline, +} + +// stepBudgets is what each step's deadline was set from, which is the other half +// of the arithmetic that measures how long a step took. It is derived from the +// graph's deadlines rather than restated, so a budget changed in one place moves +// both. +var stepBudgets = map[step]time.Duration{ + stepRequesting: requestingDeadline, + stepAwaiting: awaitingDeadline, + stepValidating: validatingDeadline, + stepSuspending: suspendingDeadline, + stepMigratingVolumes: migratingDeadline, + stepVerifying: verifyingDeadline, + stepRemoving: removingDeadline, + stepPreparing: preparingDeadline, + stepRelocating: relocatingDeadline, + stepAwaitingNode: nodeRestartDeadline, + stepPromoting: promotingDeadline, + stepHolding: holdingDeadline, + stepShuttingDown: suspendingDeadline, + stepReleasing: releasingDeadline, + stepAwaitingHost: awaitingHostDeadline, + stepRestarting: nodeRestartDeadline, + stepCleanup: cleanupDeadline, +} + +// abortableSteps are the steps from which an abort stops the operation cleanly. +// +// The line is whether anything is currently down or half-done. Requesting has +// issued nothing. Validating performs no side effect at all and is the step +// before a drain touches the node, which is why an abort there is an Aborted +// directly rather than an unwind (§8.3). Preparing has labeled a target host and +// nothing more. The three drain steps past the suspend are abortable because +// their unwind exists: the resume the graph already performs on every other +// terminal outcome from Suspending onward. +// +// Promoting is the clearest refusal. The promote has activated the target host's +// devices, failed and migrated the origin's, started a rebalance, and re-homed +// the logical volumes, so there is nothing to unwind and the operation is what +// finishes the relocation (§9). Relocating and AwaitingNode are the same rule: +// the node is mid-restart and this operation is the only thing watching it back. +// HostMaintenance past Holding is the third: the node is being taken down for a +// reboot nothing else will bring it back from. +// +// It is a table beside the graph rather than an edge in it, because a terminal +// Aborted step would be an eighteenth value in the API and the phase already +// carries that meaning. A test asserts every step here is one some graph +// declares, and another asserts that no step between a node's shutdown and its +// restart appears, so the two cannot drift. +var abortableSteps = map[step]bool{ + stepRequesting: true, + stepValidating: true, + stepSuspending: true, + stepMigratingVolumes: true, + stepVerifying: true, + stepPreparing: true, + stepHolding: true, +} + +// abortable reports whether an abort asked for while the operation sits on this +// step can be honored. +func abortable(current step) bool { return abortableSteps[current] } + +// unwinds reports whether an abort or a failure from this step owes the node a +// resume before the operation ends. Everything from Suspending onward in a drain +// does: the node is not serving, and an operation that stopped there and left it +// that way would take capacity out of the cluster indefinitely (§8.3). +func unwinds(current step) bool { + switch current { + case stepSuspending, stepMigratingVolumes, stepVerifying, stepRemoving: + return true + default: + return false + } +} + +// action converts the API's action enum into the MultiConfig's key. The +// conversion exists because statemachine.Action is a concrete string type rather +// than a second type parameter, and doing it in one place keeps the graph literal +// readable. +func action(a simplyblockv1alpha2.StorageNodeOpsAction) statemachine.Action { + return statemachine.Action(a) +} diff --git a/operator/internal/controllers/node/graphs_test.go b/operator/internal/controllers/node/graphs_test.go new file mode 100644 index 000000000..640879085 --- /dev/null +++ b/operator/internal/controllers/node/graphs_test.go @@ -0,0 +1,305 @@ +// The graphs as data: what they declare, and the three places that have to agree +// about it. +// +// A shared statemachine.KubeSnapshot cannot carry an Enum marker, so the step +// values live in three places — the graph, the kind's Enum marker, and the CEL +// rule on status.step — and nothing but a test makes them agree +// (design-storagenode.md §6.3). These are that test. +// +// The rest is what declaring a graph as data buys: a transition table is cheap to +// exercise exhaustively, so every illegal edge is a unit test rather than a review +// comment, and the refusal to abort a Promoting migration is one assertion rather +// than a code path. + +package node + +import ( + "context" + "slices" + "strings" + "testing" + + "github.com/google/go-cmp/cmp" + + "github.com/simplyblock/atlas/statemachine" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// everyStep is the Enum marker's list, transcribed. It is written out rather than +// derived so that the assertion below compares two independent statements of the +// same set: deriving it from the graph would make the test agree with itself. +var everyStep = []string{ + "Awaiting", "AwaitingHost", "AwaitingNode", "Cleanup", "Holding", + "MigratingVolumes", "Preparing", "Promoting", "Relocating", "Releasing", + "Removing", "Requesting", "Restarting", "ShuttingDown", "Suspending", + "Validating", "Verifying", +} + +func TestTheStepEnumCoversEveryDeclaredState(t *testing.T) { + declared := statemachine.DeclaredMultiStates(graphs()) + want := slices.Clone(everyStep) + slices.Sort(want) + if diff := cmp.Diff(want, declared); diff != "" { + t.Errorf("the graphs and the Enum marker disagree (-marker +graphs):\n%s", diff) + } +} + +// The CEL rule on status.step is what an Enum marker would do if a marker could +// reach a field of a type another module declares. It is a literal list in a +// struct tag, so nothing but this compares it against the graph — and it is +// compared both ways, because a rule that names a step no graph declares admits a +// status no controller can resume from. +func TestTheCELRuleCoversEveryDeclaredState(t *testing.T) { + declared := statemachine.DeclaredMultiStates(graphs()) + for _, state := range declared { + if !strings.Contains(opsStepCELRule, "'"+state+"'") { + t.Errorf("status.step's CEL rule does not accept the declared step %q", state) + } + } + for _, named := range celRuleValues(opsStepCELRule) { + if !slices.Contains(declared, named) { + t.Errorf("status.step's CEL rule accepts %q, which no graph declares", named) + } + } +} + +// The entity's own machine carries the same three-way agreement, over a Config +// rather than a MultiConfig because a StorageNode has no spec.action to key one +// on. +func TestTheNodeStepEnumAndRuleCoverTheProvisioningGraph(t *testing.T) { + declared := statemachine.DeclaredStates(provisioningGraph()) + want := []string{ + "Adopting", "AwaitingSlot", "CheckingConfig", "CheckingHost", "Posting", "Resolving", + } + if diff := cmp.Diff(want, declared); diff != "" { + t.Errorf("the provisioning graph and the Enum marker disagree (-marker +graph):\n%s", diff) + } + for _, state := range declared { + if !strings.Contains(nodeStepCELRule, "'"+state+"'") { + t.Errorf("status.step's CEL rule does not accept the declared step %q", state) + } + } + for _, named := range celRuleValues(nodeStepCELRule) { + if !slices.Contains(declared, named) { + t.Errorf("status.step's CEL rule accepts %q, which the graph does not declare", named) + } + } +} + +// celRuleValues reads the quoted values out of the rule's `in` list. +func celRuleValues(rule string) []string { + var values []string + for _, part := range strings.Split(rule, "'") { + if part != "" && !strings.ContainsAny(part, "[],| ") { + values = append(values, part) + } + } + return values +} + +// The two rules as the types declare them. Keeping a copy here is the cost of a +// rule living in a struct tag; the tests above are what make the copies worth +// having. +const opsStepCELRule = "!has(self.state) || self.state in " + + "['Requesting','Awaiting','Validating','Suspending','MigratingVolumes','Verifying'," + + "'Removing','Preparing','Relocating','AwaitingNode','Promoting','Holding'," + + "'ShuttingDown','Releasing','AwaitingHost','Restarting','Cleanup']" + +const nodeStepCELRule = "!has(self.state) || self.state in " + + "['CheckingHost','CheckingConfig','AwaitingSlot','Posting','Resolving','Adopting']" + +// Every action the API accepts needs a graph, or an operation of that action fails +// at its first pass with ErrUnknownAction rather than doing anything. +func TestEveryActionDeclaresAGraph(t *testing.T) { + declared := graphs() + for _, a := range []simplyblockv1alpha2.StorageNodeOpsAction{ + simplyblockv1alpha2.StorageNodeOpsActionShutdown, + simplyblockv1alpha2.StorageNodeOpsActionRestart, + simplyblockv1alpha2.StorageNodeOpsActionSuspend, + simplyblockv1alpha2.StorageNodeOpsActionResume, + simplyblockv1alpha2.StorageNodeOpsActionRemove, + simplyblockv1alpha2.StorageNodeOpsActionMigrate, + simplyblockv1alpha2.StorageNodeOpsActionHostMaintenance, + } { + if _, ok := declared[action(a)]; !ok { + t.Errorf("action %s declares no graph", a) + } + if _, ok := initialDeadlines[action(a)]; !ok { + t.Errorf("action %s has no budget for the step its machine is born in", a) + } + } +} + +// An abortable step that no graph declares is a table that has drifted from the +// graphs beside it, and the consequence is an abort that is silently never +// honored. +func TestEveryAbortableStepIsDeclared(t *testing.T) { + declared := statemachine.DeclaredMultiStates(graphs()) + for state := range abortableSteps { + if !slices.Contains(declared, string(state)) { + t.Errorf("step %q is marked abortable and no graph declares it", state) + } + } +} + +// The line the abort table draws is whether anything is currently down or +// half-done. These four are the sharpest cases and each would leave the node in a +// state nothing else drives it out of. +func TestNoStepPastThePointOfNoReturnIsAbortable(t *testing.T) { + for _, state := range []step{ + // The promote has re-homed the logical volumes: there is nothing to + // unwind, and the operation is what finishes the relocation. + stepPromoting, + // The node is mid-restart on a host it is being moved to, and this + // operation is the only thing watching it back. + stepRelocating, + stepAwaitingNode, + // The node is down for a reboot nothing else will bring it back from. + stepShuttingDown, + stepReleasing, + stepAwaitingHost, + stepRestarting, + } { + if abortable(state) { + t.Errorf("step %q is abortable and the node is not in a state an abort can leave it in", state) + } + } +} + +// Every terminal outcome from Suspending onward owes the node a resume, because a +// node past the suspend is not serving and an operation that stopped there would +// take capacity out of the cluster for as long as nobody noticed (§8.3). +func TestTheDrainStepsPastTheSuspendUnwind(t *testing.T) { + for _, state := range []step{ + stepSuspending, stepMigratingVolumes, stepVerifying, stepRemoving, + } { + if !unwinds(state) { + t.Errorf("step %q leaves the node suspended and owes it a resume", state) + } + } + // Validating performs no side effect at all, which is what makes an abort + // there an Aborted directly rather than an unwind. + if unwinds(stepValidating) { + t.Error("Validating touches nothing and must not issue a resume") + } +} + +func TestEveryStepHasABudget(t *testing.T) { + for _, state := range statemachine.DeclaredMultiStates(graphs()) { + if _, ok := stepBudgets[step(state)]; !ok { + t.Errorf("step %q has no budget, so no duration can be measured for it", state) + } + } +} + +// One status.step field serves all seven actions, so nothing in the API type +// prevents a Remove from reporting Promoting. The per-action graph is what makes +// that a refusal at the point of the write. +func TestAStepOfAnotherActionIsRejected(t *testing.T) { + _, err := graphs().FromSnapshot(context.Background(), + action(simplyblockv1alpha2.StorageNodeOpsActionRemove), + statemachine.Snapshot[step]{State: stepPromoting}) + if err == nil { + t.Fatal("a Remove restored into Promoting, which belongs to Migrate") + } +} + +// The drain's graph is the line §8.2 states, and its ordering is the design: +// validation before the suspend, so a drain that cannot complete never takes +// capacity out of the cluster. +func TestTheRemoveGraphValidatesBeforeItSuspends(t *testing.T) { + assertLine(t, simplyblockv1alpha2.StorageNodeOpsActionRemove, []step{ + stepValidating, stepSuspending, stepMigratingVolumes, stepVerifying, stepRemoving, + }) +} + +// Relocating and AwaitingNode are two steps because one would race, which is the +// whole reason the migration's graph is four steps rather than three. +func TestTheMigrateGraphSplitsTheRestartFromTheWait(t *testing.T) { + assertLine(t, simplyblockv1alpha2.StorageNodeOpsActionMigrate, []step{ + stepPreparing, stepRelocating, stepAwaitingNode, stepPromoting, + }) +} + +func TestTheHostMaintenanceGraphIsTheSixStepWindow(t *testing.T) { + assertLine(t, simplyblockv1alpha2.StorageNodeOpsActionHostMaintenance, []step{ + stepHolding, stepShuttingDown, stepReleasing, stepAwaitingHost, + stepRestarting, stepCleanup, + }) +} + +// The four single-step actions share one two-step line, which is what keeps a +// change to it from landing in one of four copies. +func TestTheSingleStepActionsShareOneLine(t *testing.T) { + for _, a := range []simplyblockv1alpha2.StorageNodeOpsAction{ + simplyblockv1alpha2.StorageNodeOpsActionShutdown, + simplyblockv1alpha2.StorageNodeOpsActionRestart, + simplyblockv1alpha2.StorageNodeOpsActionSuspend, + simplyblockv1alpha2.StorageNodeOpsActionResume, + } { + assertLine(t, a, []step{stepRequesting, stepAwaiting}) + } +} + +// assertLine walks an action's graph from its initial state and checks it is the +// straight line the design states, ending terminal. +func assertLine(t *testing.T, a simplyblockv1alpha2.StorageNodeOpsAction, want []step) { + t.Helper() + ctx := context.Background() + + machine, err := graphs().FromSnapshot(ctx, action(a), statemachine.Snapshot[step]{}) + if err != nil { + t.Fatalf("build the %s machine: %v", a, err) + } + defer machine.Close() + + if got := machine.CurrentState(); got != want[0] { + t.Fatalf("%s starts at %q, want %q", a, got, want[0]) + } + for i := 1; i < len(want); i++ { + if err := machine.TransitionTo(ctx, want[i]); err != nil { + t.Fatalf("%s cannot move from %q to %q: %v", a, want[i-1], want[i], err) + } + } + if !machine.IsTerminal() { + t.Errorf("%s does not end at %q", a, want[len(want)-1]) + } +} + +// The provisioning machine branches rather than running in a line: adoption +// diverts from the host check and from the configuration gate alike, which is what +// lets an upgrade Secret and a backend node found at the worker's address reach the +// same step. +func TestAdoptionIsReachableFromBothGates(t *testing.T) { + ctx := context.Background() + for _, from := range []nodeStep{stepCheckingHost, stepCheckingConfig} { + config := provisioningGraph() + machine, err := statemachine.NewFromSnapshot(ctx, config, + statemachine.Snapshot[nodeStep]{State: from}) + if err != nil { + t.Fatalf("build the provisioning machine at %q: %v", from, err) + } + if err := machine.TransitionTo(ctx, stepAdopting); err != nil { + t.Errorf("adoption is not reachable from %q: %v", from, err) + } + machine.Close() + } +} + +// The second socket of a worker must not post again: its sibling's claim is what +// it reads, and it enters Resolving directly. That edge is what makes the skip +// expressible at all. +func TestAwaitingSlotMayGoStraightToResolving(t *testing.T) { + ctx := context.Background() + machine, err := statemachine.NewFromSnapshot(ctx, provisioningGraph(), + statemachine.Snapshot[nodeStep]{State: stepAwaitingSlot}) + if err != nil { + t.Fatalf("build the provisioning machine: %v", err) + } + defer machine.Close() + + if err := machine.TransitionTo(ctx, stepResolving); err != nil { + t.Errorf("a sibling's claim cannot be read as a slot already posted: %v", err) + } +} diff --git a/operator/internal/controllers/node/hostmaintenance.go b/operator/internal/controllers/node/hostmaintenance.go new file mode 100644 index 000000000..0d9e821d3 --- /dev/null +++ b/operator/internal/controllers/node/hostmaintenance.go @@ -0,0 +1,261 @@ +// The HostMaintenance action: surviving a Kubernetes node drain. +// +// A Kubernetes worker being drained for an OS upgrade takes its storage-node pod +// with it. Left alone, the SPDK process is killed underneath a running backend +// node, which the control plane sees as a node that vanished. This action is how +// the node is taken down deliberately, allowed out of the way, and brought back. +// +// Holding ──► ShuttingDown ──► Releasing ──► AwaitingHost ──► Restarting ──► Cleanup +// +// The operator raises it, and a user does not. The trigger is the worker being +// cordoned, which the StorageNode controller sees by watching Kubernetes Node +// objects. A user creating one by hand is accepted and behaves identically, which +// is what makes the flow testable without cordoning anything. +// +// Modeling it as a StorageNodeOps rather than as its own controller is what gives +// it the discipline the other actions have: it takes the node's lock, so nothing +// else touches a node whose host is rebooting, its position is a persisted step +// rather than a phase string in a fleet object's status, and it is an audit record +// of a maintenance window afterward. That retires the eight-phase drain +// coordinator the retired StorageNodeSet carried (§15.3). +// +// The PodDisruptionBudget runs backward from the usual one. A per-node budget with +// no disruption allowed is created *before* the shutdown, so `kubectl drain` blocks +// on it while the backend node is being taken down gracefully. Relaxing it in +// Releasing is what lets the drain proceed. The budget's job is therefore to hold +// the eviction until the storage node is safely offline, rather than to keep a +// replica count up. +// +// design-storagenode.md §10 is the specification. + +package node + +import ( + "context" + "fmt" + + "k8s.io/apimachinery/pkg/types" + "sigs.k8s.io/controller-runtime/pkg/client" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// performMaintenanceStep runs one step of a maintenance window. +func (r *StorageNodeOpsReconciler) performMaintenanceStep( + ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, current step, +) (bool, error) { + node, err := r.node(ctx, ops) + if err != nil { + return false, err + } + + // Holding performs nothing and asks only about its peers, so it is the one + // step that does not need the node to exist in the control plane yet. + if current == stepHolding { + return r.maintenanceHold(ctx, ops, node) + } + + clusterID, nodeID, err := r.target(ctx, ops) + if err != nil { + return false, err + } + + switch current { + case stepShuttingDown: + return r.maintenanceShutDown(ctx, ops, node, clusterID, nodeID) + case stepReleasing: + return r.maintenanceRelease(ctx, node) + case stepAwaitingHost: + return r.maintenanceAwaitHost(ctx, node) + case stepRestarting: + return r.maintenanceRestart(ctx, ops, clusterID, nodeID) + case stepCleanup: + return r.maintenanceCleanup(ctx, node) + default: + return false, fatalf("step %s does not belong to the HostMaintenance action", current) + } +} + +// maintenanceHold is the concurrency gate, and it is a cluster-wide count. +// +// How many workers may be in maintenance at once is +// StorageCluster.status.maxConcurrentWorkerRestarts, which is the smaller of what +// the cluster asks for and the fault tolerance the control plane reports. Counting +// operations rather than nodes is what makes the gate correct across a +// multi-socket worker: two nodes on one host go into maintenance together, and the +// pair is one worker's worth of unavailability — so the count is by distinct +// worker. +func (r *StorageNodeOpsReconciler) maintenanceHold( + ctx context.Context, + ops *simplyblockv1alpha2.StorageNodeOps, + node *simplyblockv1alpha2.StorageNode, +) (bool, error) { + var cluster simplyblockv1alpha2.StorageCluster + key := types.NamespacedName{Name: node.Spec.ClusterRef, Namespace: node.Namespace} + if err := r.Get(ctx, key, &cluster); err != nil { + return false, fmt.Errorf("read cluster %s: %w", node.Spec.ClusterRef, err) + } + + limit := int32(1) + if effective := cluster.Status.MaxConcurrentWorkerRestarts; effective != nil && *effective > 0 { + limit = *effective + } + + inFlight, err := r.workersInMaintenance(ctx, ops, node) + if err != nil { + return false, err + } + if int32(len(inFlight)) >= limit { + return false, blockedf(MaintenanceQueued, + "%d of %d maintenance slots are taken by %v; this window waits its turn", + len(inFlight), limit, inFlight) + } + return true, nil +} + +// workersInMaintenance names the distinct workers whose maintenance window has +// passed Holding, excluding this operation's own. +// +// A window still in Holding holds no slot: it is queued behind the same gate, and +// counting it would let a deployment deadlock with every window waiting for every +// other. +func (r *StorageNodeOpsReconciler) workersInMaintenance( + ctx context.Context, + ops *simplyblockv1alpha2.StorageNodeOps, + node *simplyblockv1alpha2.StorageNode, +) ([]string, error) { + var operations simplyblockv1alpha2.StorageNodeOpsList + if err := r.List(ctx, &operations, client.InNamespace(ops.Namespace)); err != nil { + return nil, fmt.Errorf("list the namespace's node operations: %w", err) + } + + workers := map[string]struct{}{} + for i := range operations.Items { + other := &operations.Items[i] + if other.Name == ops.Name || + other.Spec.Action != simplyblockv1alpha2.StorageNodeOpsActionHostMaintenance || + terminalOps(other.Status.Phase) || + other.Status.Step.State == "" || + other.Status.Step.State == string(stepHolding) { + continue + } + var target simplyblockv1alpha2.StorageNode + key := types.NamespacedName{Name: other.Spec.NodeRef, Namespace: other.Namespace} + if err := r.Get(ctx, key, &target); err != nil { + continue + } + if target.Spec.ClusterRef != node.Spec.ClusterRef { + continue + } + // The worker this operation is about to take down is already counted when + // its sibling socket's window is running, which is the multi-socket case + // the gate exists to get right. + if target.Spec.WorkerNode == node.Spec.WorkerNode { + continue + } + workers[target.Spec.WorkerNode] = struct{}{} + } + + names := make([]string, 0, len(workers)) + for worker := range workers { + names = append(names, worker) + } + return names, nil +} + +// maintenanceShutDown blocks the eviction and takes the backend node down. +// +// The budget is created before the shutdown is issued, and that ordering is the +// whole mechanism: a drain that reaches the pod before the budget exists evicts it +// under a running SPDK process. +func (r *StorageNodeOpsReconciler) maintenanceShutDown( + ctx context.Context, + ops *simplyblockv1alpha2.StorageNodeOps, + node *simplyblockv1alpha2.StorageNode, + clusterID, nodeID string, +) (bool, error) { + err := r.Workload.BlockEviction(ctx, node.Namespace, node.Spec.ClusterRef, node.Spec.WorkerNode) + if err != nil { + return false, fmt.Errorf("hold the eviction of worker %s: %w", node.Spec.WorkerNode, err) + } + + reading, err := r.nodeReading(ctx, clusterID, nodeID) + if err != nil { + return false, err + } + if reading.Status == nodeStatusOffline { + return true, nil + } + if reading.Status == nodeStatusInRestart { + // Mid-restart is not a state to shut down from: the call would be refused + // and the node is on its way somewhere anyway. Waiting is what lets the + // next pass see where it landed. + return false, nil + } + if err := r.API.ShutdownNode(ctx, clusterID, nodeID); err != nil { + return false, fmt.Errorf("shut down node %s for maintenance: %w", ops.Spec.NodeRef, err) + } + return false, nil +} + +// maintenanceRelease relaxes the budget so the eviction the drain is waiting on +// can proceed, and completes when the pod has actually gone. +func (r *StorageNodeOpsReconciler) maintenanceRelease( + ctx context.Context, node *simplyblockv1alpha2.StorageNode, +) (bool, error) { + err := r.Workload.AllowEviction(ctx, node.Namespace, node.Spec.ClusterRef, node.Spec.WorkerNode) + if err != nil { + return false, fmt.Errorf("release the eviction of worker %s: %w", node.Spec.WorkerNode, err) + } + return r.Workload.PodGone(ctx, node.Namespace, node.Spec.ClusterRef, node.Spec.WorkerNode) +} + +// maintenanceAwaitHost is the step whose length nobody controls. An OS upgrade and +// a reboot take as long as they take, and the node's lock is held throughout. That +// is correct, and it is why this step's deadline is generous enough for a firmware +// update: an expiry fails the operation and leaves the node offline, needing a +// Restart to recover, so the deadline is a detection mechanism rather than a +// recovery one. +func (r *StorageNodeOpsReconciler) maintenanceAwaitHost( + ctx context.Context, node *simplyblockv1alpha2.StorageNode, +) (bool, error) { + return r.Workload.HostAnswers(ctx, node.Namespace, node.Spec.WorkerNode) +} + +// maintenanceRestart brings the backend node back on the host it left. The call is +// skipped when the node is already restarting or online, which is what makes +// re-entering the step harmless. +func (r *StorageNodeOpsReconciler) maintenanceRestart( + ctx context.Context, + ops *simplyblockv1alpha2.StorageNodeOps, + clusterID, nodeID string, +) (bool, error) { + reading, err := r.nodeReading(ctx, clusterID, nodeID) + if err != nil { + return false, err + } + if reading.Status == nodeStatusOnline { + return true, nil + } + if reading.Status == nodeStatusInRestart { + return false, nil + } + params := RestartParams{ + Force: boolValue(ops.Spec.Force), + ReattachVolume: boolValue(ops.Spec.ReattachVolume), + } + if err := r.API.RestartNode(ctx, clusterID, nodeID, params); err != nil { + return false, fmt.Errorf("restart node %s after maintenance: %w", ops.Spec.NodeRef, err) + } + return false, nil +} + +// maintenanceCleanup removes what the window put in place, so the worker is +// drainable by the ordinary rules again. +func (r *StorageNodeOpsReconciler) maintenanceCleanup( + ctx context.Context, node *simplyblockv1alpha2.StorageNode, +) (bool, error) { + err := r.Workload.ClearEvictionBudget(ctx, node.Namespace, + node.Spec.ClusterRef, node.Spec.WorkerNode) + return err == nil, err +} diff --git a/operator/internal/controllers/node/metrics.go b/operator/internal/controllers/node/metrics.go new file mode 100644 index 000000000..704bb33b7 --- /dev/null +++ b/operator/internal/controllers/node/metrics.go @@ -0,0 +1,153 @@ +// The metrics the operator publishes about storage nodes and the operations +// performed against them. +// +// They exist because nothing in the operator measured how long a drain takes, how +// long a node is locked, or how often a maintenance window holds for a peer. +// Three of them answer questions nothing else in the design can: +// +// - operation_active_state is the alert for a leaked lock. The release paths +// are idempotent and run from three places precisely because a lock held by a +// terminal operation would block the node forever, and a gauge is how that is +// noticed rather than reported. +// - drain_blocked_volumes_count turns the most common support question about a +// stalled drain into a dashboard panel. +// - operation_step_deadline_exceeded_total distinguishes an operation still +// working from one that stopped. +// +// provisioning_duration_seconds is the one to watch when a cluster is being +// expanded, because maxParallelNodeAdds and the FoundationDB serialization of +// §4.2 mean the time to add ten workers is not ten times the time to add one, and +// nothing today says what it actually is. +// +// An operation's metrics are the entity's: a StorageNodeOps records against +// simplyblock_storagenode_operations_total rather than a subsystem of its own, +// because the question an operator asks is what has happened to a node and the +// action label already says which operation it was (design-crd-model.md §7.12). +// Every series carries `cluster`, so one dashboard covers a multi-cluster +// deployment. +// +// design-storagenode.md §13.2 is the specification. + +package node + +import ( + "github.com/prometheus/client_golang/prometheus" + ctrlmetrics "sigs.k8s.io/controller-runtime/pkg/metrics" +) + +// operationBuckets span the range a node operation actually takes: a suspend that +// returns in seconds, a restart in minutes, and a drain of a node holding a +// hundred large volumes in hours. A linear set would put every interesting +// operation in one bucket. +var operationBuckets = prometheus.ExponentialBuckets(1, 3, 10) + +var ( + operationDurationSeconds = prometheus.NewHistogramVec( + prometheus.HistogramOpts{ + Name: "simplyblock_storagenode_operation_duration_seconds", + Help: "How long an operation took, from acquiring the node's lock to a terminal phase.", + Buckets: operationBuckets, + }, + []string{"cluster", "action", "result"}, + ) + + operationsTotal = prometheus.NewCounterVec( + prometheus.CounterOpts{ + Name: "simplyblock_storagenode_operations_total", + Help: "Operations that reached a terminal phase, by succeeded, failed, and aborted.", + }, + []string{"cluster", "action", "result"}, + ) + + operationStepDurationSeconds = prometheus.NewHistogramVec( + prometheus.HistogramOpts{ + Name: "simplyblock_storagenode_operation_step_duration_seconds", + Help: "How long one step of an operation took, which is where a slow operation is actually slow.", + Buckets: operationBuckets, + }, + []string{"cluster", "action", "step"}, + ) + + operationStepDeadlineExceededTotal = prometheus.NewCounterVec( + prometheus.CounterOpts{ + Name: "simplyblock_storagenode_operation_step_deadline_exceeded_total", + Help: "Steps that ran out of time, including those that expired while the operator was down.", + }, + []string{"cluster", "action", "step"}, + ) + + operationLockWaitSeconds = prometheus.NewHistogramVec( + prometheus.HistogramOpts{ + Name: "simplyblock_storagenode_operation_lock_wait_seconds", + Help: "How long an operation spent Pending behind another operation's lock.", + Buckets: operationBuckets, + }, + []string{"cluster", "action"}, + ) + + operationActiveState = prometheus.NewGaugeVec( + prometheus.GaugeOpts{ + Name: "simplyblock_storagenode_operation_active_state", + Help: "1 while a node's status.activeOpsRef is set, so a lock held by a finished operation is visible.", + }, + []string{"cluster", "node"}, + ) + + drainBlockedVolumesCount = prometheus.NewGaugeVec( + prometheus.GaugeOpts{ + Name: "simplyblock_storagenode_drain_blocked_volumes_count", + Help: "Volumes blocking a drain, by pinned and unmanaged.", + }, + []string{"cluster", "reason"}, + ) + + drainVolumesMigratedTotal = prometheus.NewCounterVec( + prometheus.CounterOpts{ + Name: "simplyblock_storagenode_drain_volumes_migrated_total", + Help: "Volumes moved off a node by a drain, so a drain's rate is graphable against its total.", + }, + []string{"cluster"}, + ) + + maintenanceHoldSeconds = prometheus.NewHistogramVec( + prometheus.HistogramOpts{ + Name: "simplyblock_storagenode_maintenance_hold_seconds", + Help: "How long a maintenance window held before it got its concurrency slot.", + Buckets: operationBuckets, + }, + []string{"cluster"}, + ) + + nodePhaseState = prometheus.NewGaugeVec( + prometheus.GaugeOpts{ + Name: "simplyblock_storagenode_phase_state", + Help: "1 for the node's current phase and 0 for the rest, so a node stuck in Provisioning is alertable.", + }, + []string{"cluster", "node", "phase"}, + ) + + provisioningDurationSeconds = prometheus.NewHistogramVec( + prometheus.HistogramOpts{ + Name: "simplyblock_storagenode_provisioning_duration_seconds", + Help: "How long a node took from object creation to online, which is what a node add costs.", + Buckets: operationBuckets, + }, + []string{"cluster"}, + ) +) + +func init() { + ctrlmetrics.Registry.MustRegister( + operationDurationSeconds, + operationsTotal, + operationStepDurationSeconds, + operationStepDeadlineExceededTotal, + operationLockWaitSeconds, + operationActiveState, + drainBlockedVolumesCount, + drainVolumesMigratedTotal, + maintenanceHoldSeconds, + nodePhaseState, + provisioningDurationSeconds, + ) +} diff --git a/operator/internal/controllers/node/migrate.go b/operator/internal/controllers/node/migrate.go new file mode 100644 index 000000000..c1cb8c835 --- /dev/null +++ b/operator/internal/controllers/node/migrate.go @@ -0,0 +1,298 @@ +// The Migrate action: relocating a node onto another worker. +// +// A migration moves a storage node to a different worker host without draining +// it. The node keeps its backend UUID, its partitions, and its logical-volume +// assignments, and what changes is the machine the SPDK process runs on. It is +// therefore not a removal followed by an add, and no volume migration is created. +// +// Preparing ──► Relocating ──► AwaitingNode ──► Promoting +// +// The mechanism is a control-plane restart pointed at a different node_address, +// which is the same primitive a host maintenance uses aimed somewhere else. +// +// Relocating and AwaitingNode are two steps because one would race. The restart is +// asynchronous, so a node still reporting online immediately after the call may be +// reporting the state from before it. This is the one place in the design where a +// step's completion condition is a negative predicate, and it is worth naming as a +// weakness: a stream that coalesces can deliver online before and online after +// without ever delivering what is in between, so the departure can be missed. What +// makes it tolerable is that Promoting re-reads the node before issuing, so a +// promote that would race is refused there. §16 Q4 is the control-plane request +// that would make the observation positive instead. +// +// design-storagenode.md §9 is the specification. + +package node + +import ( + "context" + "fmt" + + corev1 "k8s.io/api/core/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" + "k8s.io/apimachinery/pkg/types" + "sigs.k8s.io/controller-runtime/pkg/client" + + "github.com/simplyblock/atlas/ptr" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// performMigrateStep runs one step of the relocation. +func (r *StorageNodeOpsReconciler) performMigrateStep( + ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, current step, +) (bool, error) { + target := ops.Spec.MigrateParams().TargetWorkerNode + if target == "" { + return false, fatalf("spec.migrate.targetWorkerNode is empty; there is nowhere to relocate to") + } + + node, err := r.node(ctx, ops) + if err != nil { + return false, err + } + if node.Spec.WorkerNode == target && current == stepPreparing { + // The node is already where the operation would move it, which is what a + // re-run of a finished migration looks like. There is nothing to prepare + // and nothing to relocate, and saying so beats issuing a restart that + // moves a node onto the host it is on. + return true, nil + } + + clusterID, nodeID, err := r.target(ctx, ops) + if err != nil { + return false, err + } + + switch current { + case stepPreparing: + return r.migratePrepare(ctx, ops, node, target) + case stepRelocating: + return r.migrateRelocate(ctx, ops, node, target, clusterID, nodeID) + case stepAwaitingNode: + return r.migrateAwaitNode(ctx, clusterID, nodeID) + case stepPromoting: + return r.migratePromote(ctx, ops, node, target, clusterID, nodeID) + default: + return false, fatalf("step %s does not belong to the Migrate action", current) + } +} + +// migratePrepare puts the target host into the storage plane and holds until the +// control plane will be able to resolve it. +// +// It blocks on DNS, not on readiness. The control plane resolves node_address +// itself, and a name that does not resolve makes the restart fail inside the +// control plane, whose response is to reset the node to offline. Pod readiness +// happens before the EndpointSlice is published, so waiting on readiness alone +// leaves a window in which the restart is issued against a name that does not yet +// exist (§9). +func (r *StorageNodeOpsReconciler) migratePrepare( + ctx context.Context, + ops *simplyblockv1alpha2.StorageNodeOps, + node *simplyblockv1alpha2.StorageNode, + target string, +) (bool, error) { + var worker corev1.Node + if err := r.Get(ctx, types.NamespacedName{Name: target}, &worker); err != nil { + if apierrors.IsNotFound(err) { + return false, fatalf("target worker %s is not a node of this Kubernetes cluster", target) + } + return false, fmt.Errorf("read target worker %s: %w", target, err) + } + if !workerReady(&worker) { + return false, fatalf("target worker %s is not Ready", target) + } + + // The target's per-node configuration is written before it is labeled, so the + // entry exists by the time the pod's init container sources it. Any drives + // spec.migrate.newSsdPcie names are merged into the cloned allow list, so the + // target host binds them on start and they survive a later rebuild (§3.2). + err := r.Workload.CloneWorkerConfig(ctx, node.Namespace, node.Spec.ClusterRef, + node.Spec.WorkerNode, target, ops.Spec.MigrateParams().NewSsdPcie) + if err != nil { + return false, fmt.Errorf("clone the node's configuration onto worker %s: %w", target, err) + } + + if err := r.Workload.LabelWorker(ctx, node.Namespace, node.Spec.ClusterRef, target); err != nil { + return false, fmt.Errorf("label worker %s into the storage plane: %w", target, err) + } + + ready, err := r.Workload.PodReady(ctx, node.Namespace, node.Spec.ClusterRef, target) + if err != nil { + return false, err + } + if !ready { + return false, nil + } + + // The EndpointSlice is read through an uncached reader. A stale informer cache + // can miss a freshly published endpoint, and the consequence is a migration + // that waits on DNS forever while the name has in fact resolved for minutes + // (§5.4). + return r.Workload.PublishedInDNS(ctx, node.Namespace, node.Spec.ClusterRef, target) +} + +// migrateRelocate issues the restart pointed at the target and completes when the +// node has left online, which is the observation that the restart has actually +// started. +// +// The restart is forced by default. A migration relocates a node that is still +// online, and the control plane rejects a non-forced restart of a node that is not +// already offline. spec.force is honored when it is set explicitly, which is the +// one place a default of true is the right one (§9). +func (r *StorageNodeOpsReconciler) migrateRelocate( + ctx context.Context, + ops *simplyblockv1alpha2.StorageNodeOps, + node *simplyblockv1alpha2.StorageNode, + target, clusterID, nodeID string, +) (bool, error) { + reading, err := r.nodeReading(ctx, clusterID, nodeID) + if err != nil { + return false, err + } + if reading.Status != nodeStatusOnline { + return true, nil + } + + force := true + if ops.Spec.Force != nil { + force = *ops.Spec.Force + } + params := RestartParams{ + NodeAddress: r.Workload.NodeAddress(target, node.Namespace), + Force: force, + ReattachVolume: boolValue(ops.Spec.ReattachVolume), + NewSsdPcie: ops.Spec.MigrateParams().NewSsdPcie, + } + if err := r.API.RestartNode(ctx, clusterID, nodeID, params); err != nil { + return false, fmt.Errorf("relocate node %s onto worker %s: %w", + ops.Spec.NodeRef, target, err) + } + return false, nil +} + +// migrateAwaitNode waits for the node to be online again, on the host it was +// relocated onto. It performs nothing: the restart is running and this step is the +// half of the pair that observes it finish. +func (r *StorageNodeOpsReconciler) migrateAwaitNode( + ctx context.Context, clusterID, nodeID string, +) (bool, error) { + reading, err := r.nodeReading(ctx, clusterID, nodeID) + if err != nil { + return false, err + } + return reading.Status == nodeStatusOnline, nil +} + +// migratePromote activates the relocated node and then re-points the Kubernetes +// view of where it runs. +// +// The promote is the step with no way back: it activates the target host's +// devices, fails and migrates the origin host's, starts a rebalance, and re-homes +// the logical volumes. The graph declares no edge from Promoting to an aborted +// state for that reason. +// +// The topology re-point happens after the promote and never before. Doing it first +// would leave the Kubernetes view describing a relocation the control plane had +// not performed. +func (r *StorageNodeOpsReconciler) migratePromote( + ctx context.Context, + ops *simplyblockv1alpha2.StorageNodeOps, + node *simplyblockv1alpha2.StorageNode, + target, clusterID, nodeID string, +) (bool, error) { + // The promote is separately guarded, which is what makes the negative + // predicate in Relocating tolerable: a node that is not online has not + // finished its restart, and promoting into an in-flight one leaves the + // relocated devices stuck in `new` (§9). + reading, err := r.nodeReading(ctx, clusterID, nodeID) + if err != nil { + return false, err + } + if reading.Status != nodeStatusOnline { + return false, nil + } + + // A node whose object already names the target has been promoted by an + // earlier pass: the re-point below is the last thing this step does, so its + // presence is the record that the promote landed. + if node.Spec.WorkerNode != target { + if err := r.API.Promote(ctx, clusterID, nodeID); err != nil { + return false, fmt.Errorf("promote node %s on worker %s: %w", + ops.Spec.NodeRef, target, err) + } + if err := r.repointTopology(ctx, node, target, ops.Spec.MigrateParams().NewSsdPcie); err != nil { + return false, err + } + } + + // The source worker keeps its storage-plane labels while any other node still + // runs there, and loses them when none does. Removing them from a worker that + // still hosts a node would unschedule it. + err = r.Workload.ReleaseWorker(ctx, node.Namespace, node.Spec.ClusterRef, node.Name) + return err == nil, err +} + +// repointTopology moves the Kubernetes half of the relocation: the node's worker, +// the drives the migration bound on the target, and the storage-plane labels that +// follow the worker. +// +// The allow-list merge is why spec.config.pcieAllowList is guarded by the webhook +// rather than by an immutability marker: this is the one legitimate writer of it, +// and the addresses have to survive a later rebuild of the node (§3.2). +func (r *StorageNodeOpsReconciler) repointTopology( + ctx context.Context, + node *simplyblockv1alpha2.StorageNode, + target string, + newSsdPcie []string, +) error { + patch := client.MergeFrom(node.DeepCopy()) + node.Spec.WorkerNode = target + node.Spec.Config.PcieAllowList = mergePCIAddresses(node.Spec.Config.PcieAllowList, newSsdPcie) + if err := r.Patch(ctx, node, patch); err != nil { + return fmt.Errorf("re-point node %s onto worker %s: %w", node.Name, target, err) + } + return r.Workload.LabelWorker(ctx, node.Namespace, node.Spec.ClusterRef, target) +} + +// mergePCIAddresses adds the addresses a migration bound to the list the node +// already had, keeping the existing order and appending what is new. +// +// Order is preserved rather than sorted because the list is a user's, and a field +// the operator rewrites should come back recognizable to whoever wrote it. +func mergePCIAddresses(existing, added []string) []string { + if len(added) == 0 { + return existing + } + seen := make(map[string]struct{}, len(existing)) + for _, address := range existing { + seen[address] = struct{}{} + } + merged := existing + for _, address := range added { + if address == "" { + continue + } + if _, ok := seen[address]; ok { + continue + } + seen[address] = struct{}{} + merged = append(merged, address) + } + return merged +} + +// workerReady reports the Kubernetes Ready condition of a worker. +func workerReady(worker *corev1.Node) bool { + for _, condition := range worker.Status.Conditions { + if condition.Type == corev1.NodeReady { + return condition.Status == corev1.ConditionTrue + } + } + return false +} + +// ensurePtr keeps the atlas-lib pointer helper imported where the package uses it +// for the optional flags a restart carries. +var _ = ptr.To[bool] diff --git a/operator/internal/controllers/node/pernodeconfig.go b/operator/internal/controllers/node/pernodeconfig.go new file mode 100644 index 000000000..f86062491 --- /dev/null +++ b/operator/internal/controllers/node/pernodeconfig.go @@ -0,0 +1,279 @@ +// The per-node configuration the storage-node DaemonSet's init container reads. +// +// One ConfigMap holds one entry per worker, keyed by hostname, each a +// shell-sourceable env file. One ConfigMap rather than one per node is what lets a +// single DaemonSet serve nodes that differ. +// +// It is derived rather than authoritative. Every value in it comes from a +// StorageNode's own spec.config or from its cluster, so it can be rebuilt from +// them at any time, and nothing reads it back. +// +// The merge that used to happen here is gone. The retired StorageNodeSet held +// fleet defaults that each node's overrides were layered onto, which made the set +// the source of truth and the node a cache of it. A node carries its whole +// configuration now (§3.1), so an entry is a rendering of one object rather than a +// resolution of two. +// +// design-storagenode.md §5.3 is the specification. + +package node + +import ( + "context" + "fmt" + "sort" + "strings" + + corev1 "k8s.io/api/core/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil" + + atlaskube "github.com/simplyblock/atlas/kube" + "github.com/simplyblock/atlas/ptr" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/utils" +) + +// PerNodeConfigMapName is the ConfigMap of one cluster's storage nodes. It is +// named for the cluster now rather than for a node set, which is the retirement's +// one visible trace in an object name (§15.3). +func PerNodeConfigMapName(cluster string) string { + return cluster + "-per-node-config" +} + +// ReconcileConfig writes the ConfigMap from the cluster's nodes. +// +// It is written before the DaemonSet on every pass. A pod that starts against a +// missing or empty entry reaches the node configuration script with +// --max-subsys-count=0 and fails there, which is a long way from the cause. For +// the same reason, a cluster missing its required sizing is refused with an error +// naming the fields rather than written out as blanks. +func (w *Workload) ReconcileConfig( + ctx context.Context, cluster *simplyblockv1alpha2.StorageCluster, +) error { + if cluster.Spec.MaxSubsystemCount == nil || cluster.Spec.VCPUCount == nil { + return fmt.Errorf( + "cluster %s is missing the node sizing its workers boot from: "+ + "set spec.maxSubsystemCount and spec.vcpuCount", cluster.Name) + } + + var nodes simplyblockv1alpha2.StorageNodeList + if err := w.List(ctx, &nodes, client.InNamespace(cluster.Namespace)); err != nil { + return fmt.Errorf("list cluster %s's nodes: %w", cluster.Name, err) + } + + data := map[string]string{} + for i := range nodes.Items { + node := &nodes.Items[i] + if node.Spec.ClusterRef != cluster.Name { + continue + } + data[node.Spec.WorkerNode] = renderNodeConfig(cluster, node) + } + + return w.applyConfigMap(ctx, cluster, data) +} + +// CloneWorkerConfig copies one worker's entry onto another, merging any drives a +// migration is binding on the target into the cloned allow list. +// +// It runs before the target is labeled, so the entry exists by the time the pod's +// init container sources it. Persisting it here rather than letting the next +// ordinary pass rebuild it is what stops that pass writing the target's entry back +// from a node whose spec.workerNode still names the source (§9). +func (w *Workload) CloneWorkerConfig( + ctx context.Context, namespace, cluster, source, target string, newSsdPcie []string, +) error { + var clusterObject simplyblockv1alpha2.StorageCluster + key := client.ObjectKey{Namespace: namespace, Name: cluster} + if err := w.Get(ctx, key, &clusterObject); err != nil { + return fmt.Errorf("read cluster %s: %w", cluster, err) + } + + var configMap corev1.ConfigMap + configKey := client.ObjectKey{Namespace: namespace, Name: PerNodeConfigMapName(cluster)} + if err := w.Get(ctx, configKey, &configMap); err != nil { + if apierrors.IsNotFound(err) { + // Nothing to clone from. The ordinary pass writes the map, and the + // migration's Preparing step holds until the target's pod is ready, + // which cannot happen before the entry exists. + return nil + } + return fmt.Errorf("read cluster %s's per-node configuration: %w", cluster, err) + } + + entry, ok := configMap.Data[source] + if !ok { + return fmt.Errorf("worker %s has no per-node configuration to clone", source) + } + cloned := mergeAllowedIntoEntry(entry, newSsdPcie) + if configMap.Data[target] == cloned { + return nil + } + + patch := client.MergeFrom(configMap.DeepCopy()) + if configMap.Data == nil { + configMap.Data = map[string]string{} + } + configMap.Data[target] = cloned + if err := w.Patch(ctx, &configMap, patch); err != nil { + return fmt.Errorf("write worker %s's cloned configuration: %w", target, err) + } + return nil +} + +// applyConfigMap creates or updates the ConfigMap, owned by the cluster. +func (w *Workload) applyConfigMap( + ctx context.Context, cluster *simplyblockv1alpha2.StorageCluster, data map[string]string, +) error { + name := PerNodeConfigMapName(cluster.Name) + + var existing corev1.ConfigMap + key := client.ObjectKey{Namespace: cluster.Namespace, Name: name} + err := w.Get(ctx, key, &existing) + if apierrors.IsNotFound(err) { + created := &corev1.ConfigMap{ + ObjectMeta: metav1.ObjectMeta{ + Name: name, + Namespace: cluster.Namespace, + Labels: map[string]string{ + atlaskube.LabelApp: atlaskube.AppStorageNode, + atlaskube.LabelStorageNodeSet: cluster.Name, + }, + }, + Data: data, + } + if err := controllerutil.SetControllerReference(cluster, created, w.Scheme()); err != nil { + return fmt.Errorf("own the per-node configuration: %w", err) + } + if err := w.Create(ctx, created); err != nil && !apierrors.IsAlreadyExists(err) { + return fmt.Errorf("create the per-node configuration: %w", err) + } + return nil + } + if err != nil { + return fmt.Errorf("read the per-node configuration: %w", err) + } + + // A worker a migration cloned an entry onto is not in the list until the + // node's spec.workerNode has moved, and the re-point happens after the + // promote. Dropping the entry in between would rebuild the target's + // configuration from nothing while its pod is running against it, so entries + // this pass did not produce are carried forward rather than removed. + merged := make(map[string]string, len(existing.Data)+len(data)) + for worker, entry := range existing.Data { + merged[worker] = entry + } + for worker, entry := range data { + merged[worker] = entry + } + if equalConfigData(existing.Data, merged) { + return nil + } + + patch := client.MergeFrom(existing.DeepCopy()) + existing.Data = merged + if err := w.Patch(ctx, &existing, patch); err != nil { + return fmt.Errorf("write the per-node configuration: %w", err) + } + return nil +} + +// renderNodeConfig is one worker's entry: a shell-sourceable env file. +// +// MAX_SUBSYS_COUNT comes from the cluster and is identical in every entry, because +// the cap bounds how many volumes any node can serve and is therefore the +// cluster's rather than a stamp on the node (§3.1). That is also what makes a +// change to it reach the nodes that already exist rather than only the next one +// created. MAX_HUGE_PAGES_SIZE and VCPU_COUNT come from the node's own sizing, +// which is equal across the fleet in steady state and deliberately unequal for the +// duration of a rolling hardware upgrade. +func renderNodeConfig( + cluster *simplyblockv1alpha2.StorageCluster, node *simplyblockv1alpha2.StorageNode, +) string { + config := node.Spec.Config + + var entry strings.Builder + fmt.Fprintf(&entry, "MAX_SUBSYS_COUNT=%s\n", ptr.StringOrDefault(cluster.Spec.MaxSubsystemCount, "")) + fmt.Fprintf(&entry, "MAX_HUGE_PAGES_SIZE=%s\n", utils.ShellQuote(config.Sizing.MinHugePagesSize)) + fmt.Fprintf(&entry, "VCPU_COUNT=%s\n", ptr.StringOrDefault(config.Sizing.VCPUCount, "")) + fmt.Fprintf(&entry, "PCI_ALLOWED=%s\n", utils.ShellQuote(strings.Join(config.PcieAllowList, ","))) + fmt.Fprintf(&entry, "PCI_BLOCKED=%s\n", utils.ShellQuote(strings.Join(config.PcieDenyList, ","))) + fmt.Fprintf(&entry, "NVME_DEVICES=%s\n", utils.ShellQuote(strings.Join(config.DeviceNames, ","))) + fmt.Fprintf(&entry, "DEVICE_MODEL=%s\n", utils.ShellQuote(config.PcieModel)) + fmt.Fprintf(&entry, "SIZE_RANGE=%s\n", utils.ShellQuote(config.DriveSizeRange)) + if jm := config.JournalManager; jm != nil { + fmt.Fprintf(&entry, "JM_PERCENT=%s\n", ptr.StringOrDefault(jm.PercentPerDevice, "")) + fmt.Fprintf(&entry, "HA_JM_COUNT=%s\n", ptr.StringOrDefault(jm.Count, "")) + } else { + entry.WriteString("JM_PERCENT=\n") + entry.WriteString("HA_JM_COUNT=\n") + } + return entry.String() +} + +// mergeAllowedIntoEntry rewrites one entry's PCI_ALLOWED to include the addresses +// a migration is binding on the target host. +// +// It edits the rendered text rather than re-rendering from the node, because the +// node's spec.config.pcieAllowList has not been rewritten yet: that happens after +// the promote, and this runs before the relocation. +func mergeAllowedIntoEntry(entry string, added []string) string { + if len(added) == 0 { + return entry + } + lines := strings.Split(entry, "\n") + for i, line := range lines { + if !strings.HasPrefix(line, "PCI_ALLOWED=") { + continue + } + existing := parseShellList(strings.TrimPrefix(line, "PCI_ALLOWED=")) + merged := mergePCIAddresses(existing, added) + lines[i] = "PCI_ALLOWED=" + utils.ShellQuote(strings.Join(merged, ",")) + return strings.Join(lines, "\n") + } + // The entry predates the field. Appending it is what a node whose allow list + // was empty needs, and the init container reads the last assignment. + return entry + "PCI_ALLOWED=" + utils.ShellQuote(strings.Join(added, ",")) + "\n" +} + +// parseShellList reads back a comma-separated value this file wrote, stripping the +// quoting ShellQuote applied. +func parseShellList(value string) []string { + value = strings.TrimSpace(value) + value = strings.Trim(value, `'"`) + if value == "" { + return nil + } + parts := strings.Split(value, ",") + out := make([]string, 0, len(parts)) + for _, part := range parts { + if trimmed := strings.TrimSpace(part); trimmed != "" { + out = append(out, trimmed) + } + } + return out +} + +// equalConfigData compares two entry sets, so a pass that changed nothing writes +// nothing. A map comparison is enough because both halves are rendered by the same +// function from the same objects. +func equalConfigData(a, b map[string]string) bool { + if len(a) != len(b) { + return false + } + keys := make([]string, 0, len(a)) + for key := range a { + keys = append(keys, key) + } + sort.Strings(keys) + for _, key := range keys { + if a[key] != b[key] { + return false + } + } + return true +} diff --git a/operator/internal/controllers/node/remove.go b/operator/internal/controllers/node/remove.go new file mode 100644 index 000000000..e1ec515b4 --- /dev/null +++ b/operator/internal/controllers/node/remove.go @@ -0,0 +1,474 @@ +// The Remove action: draining a node before it leaves. +// +// Removing a storage node destroys it, and every logical volume whose data lives +// on it has to be somewhere else first. The drain is the part of the operation +// that makes that true, and the removal is the last step rather than the +// operation. +// +// Validating ──► Suspending ──► MigratingVolumes ──► Verifying ──► Removing +// +// Validation runs before the suspend, and that ordering is the design. A suspended +// node accepts no new volume placement, so suspending one whose drain cannot +// complete takes capacity out of the cluster and leaves it out for as long as the +// blocker goes unnoticed. Blocking first leaves the node fully operational while +// somebody decides what to do about the pinned claim. +// +// Every terminal outcome from Suspending onward resumes the node first, which is +// the unwind the graph's abort edges are declared against. It lives in the +// reconciler rather than here because a failure of any kind owes it, not only a +// failure of a step in this file. +// +// One thing here is not what the design specifies. §8.4 fans the migration out as +// one PersistentVolumeOps per volume, and that kind has not been written yet — the +// StoragePool rework recorded the same gap. The fan-out is the VolumeMigration +// that exists and works, tracked by the same label and cleaned up by the same +// cascade, and it becomes a PersistentVolumeOps when that kind lands. §15.4 +// records it. +// +// design-storagenode.md §8 is the specification. + +package node + +import ( + "context" + "fmt" + "strings" + + corev1 "k8s.io/api/core/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil" + logf "sigs.k8s.io/controller-runtime/pkg/log" + + "github.com/simplyblock/atlas/kube" + + simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// drainNodeLabel is what a migration this drain created carries, so that a List +// selects the fan-out of one node's drain and a watch maps a completion back to +// the operation that asked for it (§8.4). +const drainNodeLabel = "storage.simplyblock.io/drain-node" + +// performRemoveStep runs one step of the drain. +func (r *StorageNodeOpsReconciler) performRemoveStep( + ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, current step, +) (bool, error) { + clusterID, nodeID, err := r.target(ctx, ops) + if err != nil { + return false, err + } + + switch current { + case stepValidating: + return r.drainValidate(ctx, ops, clusterID, nodeID) + case stepSuspending: + return r.drainSuspend(ctx, ops, clusterID, nodeID) + case stepMigratingVolumes: + return r.drainMigrate(ctx, ops, clusterID, nodeID) + case stepVerifying: + return r.drainVerify(ctx, ops, clusterID, nodeID) + case stepRemoving: + return r.drainRemove(ctx, ops, clusterID, nodeID) + default: + return false, fatalf("step %s does not belong to the Remove action", current) + } +} + +// drainValidate classifies the node's volumes and refuses to go on while any of +// them is pinned or unmanaged. It performs no side effect at all, which is what +// makes an abort here an Aborted directly rather than an unwind. +// +// It also writes status.drain.volumesTotal, once, at the end. That is the number +// every later step's progress is reported against, and fixing it here is what +// stops a pending count having to be kept in step with a total. +func (r *StorageNodeOpsReconciler) drainValidate( + ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, clusterID, nodeID string, +) (bool, error) { + census, err := r.classify(ctx, ops, clusterID, nodeID) + if err != nil { + return false, err + } + if census.Incomplete { + // A claim that could not be read put a volume in the unmanaged bucket for + // safety. Blocking on that would report a drain blocked by a volume that + // is in fact accounted for, so the pass is retried instead. + return false, fmt.Errorf( + "a volume's claim could not be read; the classification is retried") + } + + cluster := r.clusterLabel(ctx, ops) + drainBlockedVolumesCount.WithLabelValues(cluster, blockedPinned). + Set(float64(len(census.Pinned))) + drainBlockedVolumesCount.WithLabelValues(cluster, blockedUnmanaged). + Set(float64(len(census.Unmanaged))) + + if len(census.Pinned) > 0 { + return false, blockedf(DrainBlocked, + "blocked: %d pinned %s, remove the %s annotation from %s", + len(census.Pinned), plural(len(census.Pinned), "volume", "volumes"), + kube.AnnoSelectedStorageNode, strings.Join(census.Pinned, ", ")) + } + if len(census.Unmanaged) > 0 { + return false, blockedf(DrainBlocked, + "blocked: %d unmanaged %s no PersistentVolume accounts for, remove %s by hand", + len(census.Unmanaged), plural(len(census.Unmanaged), "volume", "volumes"), + strings.Join(census.Unmanaged, ", ")) + } + + total := int32(len(census.Managed)) + err = r.writeStatus(ctx, ops, func(status *simplyblockv1alpha2.StorageNodeOpsStatus) { + status.Drain = &simplyblockv1alpha2.DrainStatus{VolumesTotal: total, VolumesMigrated: 0} + }) + return err == nil, err +} + +// drainSuspend takes the node out of service so that no new volume is placed on +// it while its own are being moved. The call is skipped when the node is already +// suspended or past it, which is what makes re-entering the step harmless. +func (r *StorageNodeOpsReconciler) drainSuspend( + ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, clusterID, nodeID string, +) (bool, error) { + reading, err := r.nodeReading(ctx, clusterID, nodeID) + if err != nil { + return false, err + } + if reading.Status == nodeStatusSuspended { + return true, nil + } + // A node already offline is past the state a suspend would produce: it is + // serving nothing, which is what the suspend exists to achieve. + if reading.Status == nodeStatusOffline { + return true, nil + } + if err := r.API.Suspend(ctx, clusterID, nodeID); err != nil { + return false, fmt.Errorf("suspend node %s: %w", ops.Spec.NodeRef, err) + } + return false, nil +} + +// drainMigrate moves every PV-managed volume to a peer, one migration object per +// volume, and completes when all of them have. +// +// Completed objects are deleted immediately, which is what keeps a hundred-volume +// drain from leaving a hundred objects behind. status.drain is the progress record +// rather than the objects' presence, which is why the counter is written before +// the delete rather than derived from a List (§8.4). +func (r *StorageNodeOpsReconciler) drainMigrate( + ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, clusterID, nodeID string, +) (bool, error) { + log := logf.FromContext(ctx) + + migrations, err := r.migrationsOf(ctx, ops, nodeID) + if err != nil { + return false, err + } + + // A failed migration is deleted and replaced against a fresh target, rather + // than failing the drain: the volume is still on the node, and another peer + // may take it. + if retried, err := r.retryFailedMigrations(ctx, ops, migrations); err != nil { + return false, err + } else if retried > 0 { + return false, nil + } + + census, err := r.classify(ctx, ops, clusterID, nodeID) + if err != nil { + return false, err + } + if census.Incomplete { + return false, fmt.Errorf( + "a volume's claim could not be read; the migration fan-out is retried") + } + + // Every movable volume that has no migration gets one. That covers the first + // pass, an object deleted out of band, and a volume that arrived on the node + // after the count was taken. + existing := make(map[string]struct{}, len(migrations)) + for i := range migrations { + existing[migrations[i].Name] = struct{}{} + } + var missing []managedVolume + for _, volume := range census.Managed { + if _, ok := existing[migrationName(nodeID, volume.PVName)]; !ok { + missing = append(missing, volume) + } + } + + if len(missing) > 0 { + targets, err := r.peerTargets(ctx, clusterID, nodeID, missing) + if err != nil { + return false, err + } + for _, volume := range missing { + if err := r.createMigration(ctx, ops, nodeID, volume, targets[volume.PVName]); err != nil { + log.Error(err, "a volume's migration could not be created", + "volume", volume.VolumeUUID, "persistentVolume", volume.PVName) + } + } + return false, nil + } + + // No movable volume is left and no migration is outstanding: everything that + // was going to move has moved. The census is the authority rather than the + // counter, because the counter is a record of what this operation did and the + // census is what is actually on the node. + if len(census.Managed) == 0 && len(migrations) == 0 { + r.emit(ctx, ops, corev1.EventTypeNormal, DrainCompleted, + "Every volume has been migrated off the node") + return true, nil + } + + completed, running := 0, 0 + for i := range migrations { + if migrations[i].Status.Phase == simplyblockv1alpha1.VolumeMigrationPhaseCompleted { + completed++ + } else { + running++ + } + } + + if running > 0 { + return false, r.recordDrainProgress(ctx, ops, completed) + } + + // Every migration finished. The counter is written before the objects go, so + // a crash between the two leaves the progress recorded rather than lost. + if err := r.recordDrainProgress(ctx, ops, completed); err != nil { + return false, err + } + drainVolumesMigratedTotal.WithLabelValues(r.clusterLabel(ctx, ops)).Add(float64(completed)) + for i := range migrations { + if err := r.Delete(ctx, &migrations[i]); err != nil && !apierrors.IsNotFound(err) { + log.Error(err, "a completed migration could not be deleted", + "migration", migrations[i].Name) + } + } + return false, nil +} + +// recordDrainProgress writes how many volumes have moved. The total stays as +// Validating fixed it: a drain that finds one more volume than it counted moves it +// too, and reporting eleven of ten is more honest than silently raising the total. +func (r *StorageNodeOpsReconciler) recordDrainProgress( + ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, completed int, +) error { + return r.writeStatus(ctx, ops, func(status *simplyblockv1alpha2.StorageNodeOpsStatus) { + if status.Drain == nil { + status.Drain = &simplyblockv1alpha2.DrainStatus{} + } + status.Drain.VolumesMigrated = int32(completed) + }) +} + +// drainVerify deletes the system volumes the migration skipped and completes when +// the node reports no volumes at all. +// +// They are deleted rather than migrated because they are per-node benchmark +// artifacts: moving one to a peer would produce a benchmark volume measuring the +// wrong node. A delete the control plane refuses for a reason other than "already +// gone" fails the operation, because a volume that cannot be deleted and cannot be +// migrated is a volume the removal would destroy (§8.2). +func (r *StorageNodeOpsReconciler) drainVerify( + ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, clusterID, nodeID string, +) (bool, error) { + census, err := r.classify(ctx, ops, clusterID, nodeID) + if err != nil { + return false, err + } + if census.Incomplete { + return false, fmt.Errorf( + "a volume's claim could not be read; the verification is retried") + } + + for _, volume := range census.System { + if err := r.API.DeleteVolume(ctx, clusterID, volume.PoolUUID, volume.VolumeUUID); err != nil { + return false, fatalf("system volume %s could not be deleted and the node still holds it: %v", + volume.Name, err) + } + } + + remaining := len(census.Managed) + len(census.Pinned) + len(census.Unmanaged) + if remaining > 0 { + return false, blockedf(DrainBlocked, + "%d %s still on the node after the migration; the removal is held", + remaining, plural(remaining, "volume", "volumes")) + } + // The system volumes were deleted on this pass and the control plane's + // deletion is asynchronous, so the node is not empty until a later pass says + // so. Reporting unfinished is what makes the next pass re-read rather than + // trust this one's arithmetic. + return len(census.System) == 0, nil +} + +// drainRemove deletes the backend node. A 404 is success, since a node the control +// plane no longer knows about is a node that has been removed. +func (r *StorageNodeOpsReconciler) drainRemove( + ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, clusterID, nodeID string, +) (bool, error) { + if err := r.API.RemoveNode(ctx, clusterID, nodeID); err != nil { + // The control plane's own admission refused the removal, which is its + // answer about what the cluster can afford to lose. Retrying cannot change + // it, so the operation fails and the resume of §8.3 puts the node back + // into service. + return false, fatalf("the control plane refused to remove node %s: %v", + ops.Spec.NodeRef, err) + } + return true, nil +} + +// migrationsOf lists the fan-out of this drain. +func (r *StorageNodeOpsReconciler) migrationsOf( + ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, nodeID string, +) ([]simplyblockv1alpha1.VolumeMigration, error) { + var migrations simplyblockv1alpha1.VolumeMigrationList + err := r.List(ctx, &migrations, + client.InNamespace(ops.Namespace), + client.MatchingLabels{drainNodeLabel: nodeID}) + if err != nil { + return nil, fmt.Errorf("list this drain's volume migrations: %w", err) + } + return migrations.Items, nil +} + +// createMigration raises one volume's move, owned by the operation that asked for +// it so that deleting the drain cascades to its fan-out. +func (r *StorageNodeOpsReconciler) createMigration( + ctx context.Context, + ops *simplyblockv1alpha2.StorageNodeOps, + nodeID string, + volume managedVolume, + target string, +) error { + migration := &simplyblockv1alpha1.VolumeMigration{ + ObjectMeta: metav1.ObjectMeta{ + Name: migrationName(nodeID, volume.PVName), + Namespace: ops.Namespace, + Labels: map[string]string{drainNodeLabel: nodeID}, + }, + Spec: simplyblockv1alpha1.VolumeMigrationSpec{ + PVName: volume.PVName, + TargetNodeUUID: target, + }, + } + if err := controllerutil.SetControllerReference(ops, migration, r.Scheme); err != nil { + return fmt.Errorf("own the migration of %s: %w", volume.PVName, err) + } + if err := r.Create(ctx, migration); err != nil && !apierrors.IsAlreadyExists(err) { + return fmt.Errorf("create the migration of %s: %w", volume.PVName, err) + } + return nil +} + +// retryFailedMigrations deletes every migration that failed and reports how many, +// so the caller's next pass recreates them against a fresh round-robin target. +// +// Retrying rather than failing the drain is the design: the volume is still on the +// node, and the peer it could not reach is not the only peer. +func (r *StorageNodeOpsReconciler) retryFailedMigrations( + ctx context.Context, + ops *simplyblockv1alpha2.StorageNodeOps, + migrations []simplyblockv1alpha1.VolumeMigration, +) (int, error) { + retried := 0 + for i := range migrations { + migration := &migrations[i] + if migration.Status.Phase != simplyblockv1alpha1.VolumeMigrationPhaseFailed { + continue + } + r.emit(ctx, ops, corev1.EventTypeWarning, MigrationRetried, fmt.Sprintf( + "The migration of %s failed and is being retried against another peer: %s", + migration.Spec.PVName, migration.Status.ErrorMessage)) + if err := r.Delete(ctx, migration); err != nil && !apierrors.IsNotFound(err) { + return retried, fmt.Errorf("delete the failed migration of %s: %w", + migration.Spec.PVName, err) + } + retried++ + } + return retried, nil +} + +// cascadeMigrations aborts the fan-out of a drain that is being deleted and +// reports whether any of it is still running. +// +// Aborting before deleting is what makes the cascade's own deletes admissible and +// what stops a deleted operation leaving migrations running behind it (§8.4). +func (r *StorageNodeOpsReconciler) cascadeMigrations( + ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, +) (bool, error) { + node, err := r.node(ctx, ops) + if apierrors.IsNotFound(err) { + return false, nil + } + if err != nil { + return false, err + } + migrations, err := r.migrationsOf(ctx, ops, node.Status.UUID) + if err != nil { + return false, err + } + pending := false + for i := range migrations { + migration := &migrations[i] + if !terminalMigration(migration.Status.Phase) { + pending = true + continue + } + if err := r.Delete(ctx, migration); err != nil && !apierrors.IsNotFound(err) { + return true, fmt.Errorf("delete the migration of %s: %w", migration.Spec.PVName, err) + } + } + return pending, nil +} + +// abortMigrations is the abort path's half of the cascade: a drain called off +// mid-flight leaves nothing moving behind it. +func (r *StorageNodeOpsReconciler) abortMigrations( + ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, current step, +) { + if current != stepMigratingVolumes && current != stepVerifying { + return + } + if _, err := r.cascadeMigrations(ctx, ops); err != nil { + logf.FromContext(ctx).Error(err, "the drain's migrations could not all be stopped", + "operation", ops.Name) + } +} + +// terminalMigration reports a phase a volume migration can never leave. +func terminalMigration(phase simplyblockv1alpha1.VolumeMigrationPhase) bool { + switch phase { + case simplyblockv1alpha1.VolumeMigrationPhaseCompleted, + simplyblockv1alpha1.VolumeMigrationPhaseFailed, + simplyblockv1alpha1.VolumeMigrationPhaseAborted: + return true + default: + return false + } +} + +// migrationFormula names one volume's move. It is derived rather than generated +// so that the fan-out is idempotent: a pass that runs again finds the object it +// made rather than making a second. +// +// The formula is atlas-lib's rather than a local truncation, which is what keeps +// two long PersistentVolume names that share a prefix from colliding on one +// object name: the digest is part of the formula rather than something a caller +// remembers to append. +var migrationFormula = kube.Formula{Prefix: "drain-"} + +func migrationName(nodeID, pvName string) string { + return migrationFormula.Derive(nodeID, pvName).Value +} + +// plural picks the noun for a count, so a message reads 1 pinned volume rather +// than 1 pinned volume(s). +func plural(n int, one, many string) string { + if n == 1 { + return one + } + return many +} diff --git a/operator/internal/controller/storagedevice_collector.go b/operator/internal/controllers/node/storagedevice_collector.go similarity index 99% rename from operator/internal/controller/storagedevice_collector.go rename to operator/internal/controllers/node/storagedevice_collector.go index e33956fa7..6b91c4ec8 100644 --- a/operator/internal/controller/storagedevice_collector.go +++ b/operator/internal/controllers/node/storagedevice_collector.go @@ -15,7 +15,7 @@ // object went away, and a series nobody deletes reports a failed drive as // healthy forever. -package controller +package node import ( "context" diff --git a/operator/internal/controller/storagedevice_collector_test.go b/operator/internal/controllers/node/storagedevice_collector_test.go similarity index 98% rename from operator/internal/controller/storagedevice_collector_test.go rename to operator/internal/controllers/node/storagedevice_collector_test.go index 3b5b3d075..99a04368d 100644 --- a/operator/internal/controller/storagedevice_collector_test.go +++ b/operator/internal/controllers/node/storagedevice_collector_test.go @@ -6,7 +6,7 @@ // scraping, because what matters is the value under a label set and not the // exposition format. -package controller +package node import ( "context" @@ -24,6 +24,7 @@ import ( "github.com/simplyblock/atlas/ptr" simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/testsupport" ) // fakeCapacitySource is a static device-capacity provider, or a broken one when @@ -86,7 +87,7 @@ func newCollector( ) *StorageDeviceCollector { t.Helper() resetDeviceMetrics() - scheme := newTestScheme(t, simplyblockv1alpha1.AddToScheme, simplyblockv1alpha2.AddToScheme) + scheme := testsupport.NewScheme(t, simplyblockv1alpha1.AddToScheme, simplyblockv1alpha2.AddToScheme) c := fake.NewClientBuilder().WithScheme(scheme).WithObjects(objs...).Build() return &StorageDeviceCollector{ Client: c, diff --git a/operator/internal/controller/storagedevice_controller.go b/operator/internal/controllers/node/storagedevice_controller.go similarity index 95% rename from operator/internal/controller/storagedevice_controller.go rename to operator/internal/controllers/node/storagedevice_controller.go index f3b3fdee0..76e830549 100644 --- a/operator/internal/controller/storagedevice_controller.go +++ b/operator/internal/controllers/node/storagedevice_controller.go @@ -4,7 +4,7 @@ // update and delete ones, which is what makes it different from the reconcilers // that converge a user's spec. -package controller +package node import ( "context" @@ -28,7 +28,6 @@ import ( "sigs.k8s.io/controller-runtime/pkg/source" "github.com/simplyblock/atlas/ptr" - simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" "github.com/simplyblock/simplyblock-operator/internal/cpinformer" "github.com/simplyblock/simplyblock-operator/internal/cpinformer/subscriptions" @@ -46,7 +45,7 @@ const StorageNodeUUIDIndex = "status.uuid" // does. A node with no id yet is not indexed: it has nothing a device could // match against. func IndexStorageNodeUUID(o client.Object) []string { - node, ok := o.(*simplyblockv1alpha1.StorageNode) + node, ok := o.(*simplyblockv1alpha2.StorageNode) if !ok || node.Status.UUID == "" { return nil } @@ -107,7 +106,7 @@ type StorageDeviceReconciler struct { // control-plane changes). Both enqueue a StorageDevice to reconcile. func (r *StorageDeviceReconciler) SetupWithManager(mgr ctrl.Manager) error { if err := mgr.GetFieldIndexer().IndexField( - context.Background(), &simplyblockv1alpha1.StorageNode{}, StorageNodeUUIDIndex, IndexStorageNodeUUID, + context.Background(), &simplyblockv1alpha2.StorageNode{}, StorageNodeUUIDIndex, IndexStorageNodeUUID, ); err != nil { return err } @@ -187,7 +186,7 @@ func (r *StorageDeviceReconciler) unreported( // and Removed records a departure that has already happened. Unknown replaces // Online and Degraded, which are observations of a device that was serving. func (r *StorageDeviceReconciler) markUnobservable( - ctx context.Context, sd *simplyblockv1alpha2.StorageDevice, node *simplyblockv1alpha1.StorageNode, + ctx context.Context, sd *simplyblockv1alpha2.StorageDevice, node *simplyblockv1alpha2.StorageNode, ) error { previous := sd.Status.Phase message := fmt.Sprintf( @@ -236,7 +235,7 @@ func (r *StorageDeviceReconciler) markUnobservable( // removed it, so the two events distinguish an orderly departure from an abrupt // one and no more than that. func (r *StorageDeviceReconciler) announceDeparture( - sd *simplyblockv1alpha2.StorageDevice, node *simplyblockv1alpha1.StorageNode, + sd *simplyblockv1alpha2.StorageDevice, node *simplyblockv1alpha2.StorageNode, ) { if node == nil { return // the node is gone too, and garbage collection is the whole story @@ -285,7 +284,7 @@ func (r *StorageDeviceReconciler) announcePhase( // node qualify, and an unrecognized state does not: a status the operator does // not know is an absence of information, the same way an unrecognized device // status is (see [devicePhase]). -func nodeSeesItsDevices(node *simplyblockv1alpha1.StorageNode) bool { +func nodeSeesItsDevices(node *simplyblockv1alpha2.StorageNode) bool { switch nodeState(node) { case utils.NodeStatusOnline, utils.NodeStatusSuspended, utils.NodeStatusRemoved: return true @@ -296,7 +295,7 @@ func nodeSeesItsDevices(node *simplyblockv1alpha1.StorageNode) bool { // nodeState is the node's control-plane status folded to lower case, or // "unknown" when it has none yet. -func nodeState(node *simplyblockv1alpha1.StorageNode) string { +func nodeState(node *simplyblockv1alpha2.StorageNode) string { if node.Status.Status == "" { return "unknown" } @@ -321,10 +320,7 @@ func (r *StorageDeviceReconciler) upsert( return ctrl.Result{RequeueAfter: deviceRetry}, nil } - labels, err := r.deviceLabels(ctx, node) - if err != nil { - return ctrl.Result{}, err - } + labels := r.deviceLabels(node) spec := simplyblockv1alpha2.StorageDeviceSpec{NodeRef: node.Name, DeviceID: dto.ID} status := simplyblockv1alpha2.StorageDeviceStatus{ @@ -419,28 +415,20 @@ func (r *StorageDeviceReconciler) upsert( // absent label is a selector that matches nothing, and a wrong one is a selector // that matches the wrong devices. func (r *StorageDeviceReconciler) deviceLabels( - ctx context.Context, node *simplyblockv1alpha1.StorageNode, -) (map[string]string, error) { + node *simplyblockv1alpha2.StorageNode, +) map[string]string { labels := map[string]string{simplyblockv1alpha2.DeviceLabelNode: node.Name} if worker := node.Labels[simplyblockv1alpha2.DeviceLabelWorker]; worker != "" { labels[simplyblockv1alpha2.DeviceLabelWorker] = worker } - if node.Spec.StorageNodeSetRef == "" { - return labels, nil - } - - var set simplyblockv1alpha1.StorageNodeSet - key := client.ObjectKey{Namespace: node.Namespace, Name: node.Spec.StorageNodeSetRef} - switch err := r.Get(ctx, key, &set); { - case apierrors.IsNotFound(err): - return labels, nil - case err != nil: - return nil, err - } - if set.Spec.ClusterName != "" { - labels[simplyblockv1alpha2.DeviceLabelCluster] = set.Spec.ClusterName + // The cluster is on the node itself now. It used to be reached through the + // StorageNodeSet the node belonged to, which made a label on a device depend on + // a third object being readable; a node names its own cluster, so there is + // nothing left to look up (design-storagenode.md §3.1). + if node.Spec.ClusterRef != "" { + labels[simplyblockv1alpha2.DeviceLabelCluster] = node.Spec.ClusterRef } - return labels, nil + return labels } // deviceOwnedLabels are the keys the mirror writes and is therefore responsible @@ -566,8 +554,8 @@ func deviceTrouble(dto subscriptions.DeviceDTO) string { // resolve a backend id from. func (r *StorageDeviceReconciler) nodeNamed( ctx context.Context, namespace, name string, -) (*simplyblockv1alpha1.StorageNode, error) { - var node simplyblockv1alpha1.StorageNode +) (*simplyblockv1alpha2.StorageNode, error) { + var node simplyblockv1alpha2.StorageNode switch err := r.Get(ctx, client.ObjectKey{Namespace: namespace, Name: name}, &node); { case apierrors.IsNotFound(err): return nil, nil @@ -580,8 +568,8 @@ func (r *StorageDeviceReconciler) nodeNamed( // nodeFor returns the StorageNode carrying the given backend node id, or nil // when none does. A nil node is not an error: the node's own object may not have // been created yet, or may already be on its way out. -func (r *StorageDeviceReconciler) nodeFor(ctx context.Context, namespace, nodeID string) (*simplyblockv1alpha1.StorageNode, error) { - var nodes simplyblockv1alpha1.StorageNodeList +func (r *StorageDeviceReconciler) nodeFor(ctx context.Context, namespace, nodeID string) (*simplyblockv1alpha2.StorageNode, error) { + var nodes simplyblockv1alpha2.StorageNodeList if err := r.List(ctx, &nodes, client.InNamespace(namespace), client.MatchingFields{StorageNodeUUIDIndex: nodeID}, diff --git a/operator/internal/controller/storagedevice_controller_reporting_test.go b/operator/internal/controllers/node/storagedevice_controller_reporting_test.go similarity index 98% rename from operator/internal/controller/storagedevice_controller_reporting_test.go rename to operator/internal/controllers/node/storagedevice_controller_reporting_test.go index b5b6de930..83a279436 100644 --- a/operator/internal/controller/storagedevice_controller_reporting_test.go +++ b/operator/internal/controllers/node/storagedevice_controller_reporting_test.go @@ -7,7 +7,7 @@ // because that file is about the mirror's create, update, and delete paths, // which are the same three whatever the device turns out to be. -package controller +package node import ( "context" @@ -22,6 +22,7 @@ import ( simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/testsupport" "github.com/simplyblock/simplyblock-operator/internal/cpinformer/subscriptions" "github.com/simplyblock/simplyblock-operator/internal/utils" ) @@ -38,10 +39,10 @@ func sdNodeSet() *simplyblockv1alpha1.StorageNodeSet { // sdNodeWithStatus is the owning node in a given control-plane state, wired to // its set and carrying the worker label the device copies. -func sdNodeWithStatus(status string) *simplyblockv1alpha1.StorageNode { +func sdNodeWithStatus(status string) *simplyblockv1alpha2.StorageNode { node := sdNodeObject() node.Labels = map[string]string{simplyblockv1alpha2.DeviceLabelWorker: "worker-3"} - node.Spec.StorageNodeSetRef = "production-set" + node.Spec.ClusterRef = "production" node.Status.Status = status return node } @@ -444,7 +445,7 @@ func TestTheUnknownTransitionRetriesOnConflict(t *testing.T) { ctx context.Context, c client.Client, subResourceName string, obj client.Object, opts ...client.SubResourceUpdateOption, ) error { - if subResourceName == statusSubresource { + if subResourceName == testsupport.StatusSubresource { statusUpdates++ if statusUpdates == 1 { return sdConflictErr() diff --git a/operator/internal/controller/storagedevice_controller_unit_test.go b/operator/internal/controllers/node/storagedevice_controller_unit_test.go similarity index 95% rename from operator/internal/controller/storagedevice_controller_unit_test.go rename to operator/internal/controllers/node/storagedevice_controller_unit_test.go index 85d440921..040831761 100644 --- a/operator/internal/controller/storagedevice_controller_unit_test.go +++ b/operator/internal/controllers/node/storagedevice_controller_unit_test.go @@ -2,7 +2,7 @@ // device status to a typed phase, and the create, update, and delete paths of // the reconciler that publishes it. -package controller +package node import ( "context" @@ -21,6 +21,7 @@ import ( "github.com/simplyblock/atlas/ptr" simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/testsupport" "github.com/simplyblock/simplyblock-operator/internal/cpinformer" "github.com/simplyblock/simplyblock-operator/internal/cpinformer/subscriptions" ) @@ -52,10 +53,10 @@ func (f *fakeDeviceCache) Lookup(key types.NamespacedName) (cpinformer.Scope, su return sdScope(), dto, true } -func sdNodeObject() *simplyblockv1alpha1.StorageNode { - return &simplyblockv1alpha1.StorageNode{ +func sdNodeObject() *simplyblockv1alpha2.StorageNode { + return &simplyblockv1alpha2.StorageNode{ ObjectMeta: metav1.ObjectMeta{Namespace: "sb", Name: sdNodeCR}, - Status: simplyblockv1alpha1.StorageNodeStatus{UUID: sdNodeID}, + Status: simplyblockv1alpha2.StorageNodeStatus{UUID: sdNodeID}, } } @@ -64,11 +65,11 @@ func sdNodeObject() *simplyblockv1alpha1.StorageNode { // fake client only answers a MatchingFields query for an index it was given. func sdReconciler(t *testing.T, cache DeviceCache, objs ...client.Object) *StorageDeviceReconciler { t.Helper() - scheme := newTestScheme(t, simplyblockv1alpha1.AddToScheme, simplyblockv1alpha2.AddToScheme) + scheme := testsupport.NewScheme(t, simplyblockv1alpha1.AddToScheme, simplyblockv1alpha2.AddToScheme) c := fake.NewClientBuilder(). WithScheme(scheme). WithStatusSubresource(&simplyblockv1alpha2.StorageDevice{}). - WithIndex(&simplyblockv1alpha1.StorageNode{}, StorageNodeUUIDIndex, IndexStorageNodeUUID). + WithIndex(&simplyblockv1alpha2.StorageNode{}, StorageNodeUUIDIndex, IndexStorageNodeUUID). WithObjects(objs...). Build() return &StorageDeviceReconciler{ @@ -266,9 +267,9 @@ func TestTheMirrorCreatesTheDeviceInTheOwningNodesNamespace(t *testing.T) { // all: it learns the node's, which is the whole of the fix. const nodeNamespace = "default" - node := &simplyblockv1alpha1.StorageNode{ + node := &simplyblockv1alpha2.StorageNode{ ObjectMeta: metav1.ObjectMeta{Namespace: nodeNamespace, Name: sdNodeCR}, - Status: simplyblockv1alpha1.StorageNodeStatus{UUID: sdNodeID}, + Status: simplyblockv1alpha2.StorageNodeStatus{UUID: sdNodeID}, } // The real subscription, wired the way cmd/main.go wires it, so the object @@ -316,11 +317,11 @@ func sdReconcilerWithInterceptor( t *testing.T, cache DeviceCache, funcs interceptor.Funcs, objs ...client.Object, ) *StorageDeviceReconciler { t.Helper() - scheme := newTestScheme(t, simplyblockv1alpha1.AddToScheme, simplyblockv1alpha2.AddToScheme) + scheme := testsupport.NewScheme(t, simplyblockv1alpha1.AddToScheme, simplyblockv1alpha2.AddToScheme) c := fake.NewClientBuilder(). WithScheme(scheme). WithStatusSubresource(&simplyblockv1alpha2.StorageDevice{}). - WithIndex(&simplyblockv1alpha1.StorageNode{}, StorageNodeUUIDIndex, IndexStorageNodeUUID). + WithIndex(&simplyblockv1alpha2.StorageNode{}, StorageNodeUUIDIndex, IndexStorageNodeUUID). WithObjects(objs...). WithInterceptorFuncs(funcs). Build() @@ -368,7 +369,7 @@ func TestStorageDeviceStatusUpdateRetriesOnConflict(t *testing.T) { ctx context.Context, c client.Client, subResourceName string, obj client.Object, opts ...client.SubResourceUpdateOption, ) error { - if subResourceName == statusSubresource { + if subResourceName == testsupport.StatusSubresource { statusUpdates++ if statusUpdates == 1 { return sdConflictErr() diff --git a/operator/internal/controller/storagedevice_metrics.go b/operator/internal/controllers/node/storagedevice_metrics.go similarity index 99% rename from operator/internal/controller/storagedevice_metrics.go rename to operator/internal/controllers/node/storagedevice_metrics.go index fffbf3338..f2ff7082f 100644 --- a/operator/internal/controller/storagedevice_metrics.go +++ b/operator/internal/controllers/node/storagedevice_metrics.go @@ -12,7 +12,7 @@ // than joined from the control plane's own exporter, so nothing needs the // backend id to line the two up. -package controller +package node import ( "github.com/prometheus/client_golang/prometheus" diff --git a/operator/internal/controllers/node/storagenode_controller.go b/operator/internal/controllers/node/storagenode_controller.go new file mode 100644 index 000000000..571f8eae4 --- /dev/null +++ b/operator/internal/controllers/node/storagenode_controller.go @@ -0,0 +1,1288 @@ +// The StorageNode reconciler: it owns one backend node's own lifecycle, and +// nothing else. Every imperative operation belongs to StorageNodeOps, and the +// Kubernetes workload the node runs in belongs to the cluster. +// +// Four paths leave this reconcile, and status.uuid is what chooses between them: +// empty means the provisioning machine of §4.2 is running, and non-empty means +// steady-state synchronization against the storage-node stream. Deletion and +// adoption are the other two. +// +// Provisioning is a persisted machine rather than a sequence of nullable fields +// because adding a backend node is not idempotent, and the call adds every socket +// of the worker at once rather than one node. Two StorageNode objects for the same +// worker that both observe an empty status.uuid would add the worker twice. The +// claim is therefore made in Kubernetes before the control plane is touched: the +// transition into Posting is an optimistic-lock patch, which succeeds for exactly +// one reconciler at a given resourceVersion and returns 409 to the rest. That +// replaces both status.postedAt and the List over sibling objects that stood in +// for it. +// +// design-storagenode.md §4 is the specification. + +package node + +import ( + "context" + "errors" + "fmt" + "slices" + "strconv" + "time" + + corev1 "k8s.io/api/core/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/runtime" + "k8s.io/apimachinery/pkg/types" + "k8s.io/client-go/tools/events" + "k8s.io/client-go/util/retry" + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil" + "sigs.k8s.io/controller-runtime/pkg/handler" + logf "sigs.k8s.io/controller-runtime/pkg/log" + "sigs.k8s.io/controller-runtime/pkg/reconcile" + + "github.com/simplyblock/atlas/prometheus" + "github.com/simplyblock/atlas/ptr" + "github.com/simplyblock/atlas/statemachine" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/cpinformer" + "github.com/simplyblock/simplyblock-operator/internal/utils" +) + +const ( + // NodeFinalizer holds the object while its backend node is drained. A node + // that is online has data on it, and deleting the object must not delete the + // node underneath without moving that data first (§4.5). + NodeFinalizer = "storage.simplyblock.io/storagenode-finalizer" + + // nodeRetry is the slow backstop. Steady state arrives on the storage-node + // stream, so this is what covers a stream nothing has yet noticed is dead, + // and the interval at which a held provisioning step looks again. + nodeRetry = 30 * time.Second + + // nodeAdvance is how long a pass that moved the machine forward waits before + // the next one. The status write this pass made is itself a change the + // controller watches, so this is the backstop for the event rather than the + // path the next step normally arrives on. + nodeAdvance = time.Second + + // clusterRefField indexes nodes by the cluster they belong to, so a cluster + // event wakes its nodes rather than every node in the deployment. + clusterRefField = "spec.clusterRef" +) + +// StorageNodeReconciler reconciles a StorageNode. +type StorageNodeReconciler struct { + client.Client + Scheme *runtime.Scheme + Recorder events.EventRecorder + API ControlPlane + + // Nodes is the storage-node stream's cache, which steady state is read from + // and which adoption matches against. It is optional: a deployment without the + // control-plane informer falls back to reading the control plane directly. + Nodes NodeCache + + // Registries are told which Kubernetes object a backend node id belongs to, + // once this reconcile has resolved one. The control plane knows nothing of + // object names, so a stream cannot enqueue a reconcile for a node it has never + // been told about. + Registries []NodeObjectRegistry + + // DeviceScopes is the device stream's scope set. A device stream is per node + // rather than per cluster, so a scope is opened when a node resolves its UUID + // and closed when the node goes: nothing else knows a node exists to stream + // the devices of. + DeviceScopes *cpinformer.ScopeSet + + // Capacity is where a node's occupancy is read from. Neither the node list nor + // the node stream carries how full a node is; the numbers exist only in the + // metrics the control plane exports (§12). + Capacity NodeCapacitySource + + // Workload is the storage-plane side: the worker labels, the per-node + // configuration, and the probe that asks whether a worker's storage-node API + // answers. + Workload *Workload +} + +// NodeObjectRegistry is told the Kubernetes object behind a backend node id. +type NodeObjectRegistry interface { + RegisterNode(nodeID string, object types.NamespacedName) + UnregisterNode(nodeID string) +} + +// NodeCapacitySource supplies a node's occupancy, satisfied by atlas-lib's +// prometheus.Provider. +type NodeCapacitySource interface { + NodeCapacity(ctx context.Context, clusterUUID string) (map[string]prometheus.Capacity, error) +} + +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodes,verbs=get;list;watch;create;update;patch;delete +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodes/status,verbs=get;update;patch +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodes/finalizers,verbs=update +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodeops,verbs=get;list;watch;create;delete +// +kubebuilder:rbac:groups="",resources=secrets,verbs=get;list;watch + +// SetupWithManager registers the controller. +// +// It watches three things beside its own kind. A StorageCluster event wakes its +// nodes, because the cluster's UUID is what provisioning waits for. A Kubernetes +// Node event is how a cordon is seen, which is what raises a maintenance window +// (§10). And the storage-node stream's triggers are how steady state arrives at +// all. +func (r *StorageNodeReconciler) SetupWithManager(mgr ctrl.Manager) error { + err := mgr.GetFieldIndexer().IndexField(context.Background(), + &simplyblockv1alpha2.StorageNode{}, clusterRefField, + func(object client.Object) []string { + node, ok := object.(*simplyblockv1alpha2.StorageNode) + if !ok { + return nil + } + return []string{node.Spec.ClusterRef} + }) + if err != nil { + return fmt.Errorf("index nodes by their cluster: %w", err) + } + + return ctrl.NewControllerManagedBy(mgr). + For(&simplyblockv1alpha2.StorageNode{}). + Named("storagenode"). + Watches(&simplyblockv1alpha2.StorageCluster{}, + handler.EnqueueRequestsFromMapFunc(r.nodesOf)). + Watches(&corev1.Node{}, + handler.EnqueueRequestsFromMapFunc(r.nodesOn)). + Complete(r) +} + +// nodesOf enqueues every node of a cluster. +func (r *StorageNodeReconciler) nodesOf( + ctx context.Context, cluster client.Object, +) []reconcile.Request { + var nodes simplyblockv1alpha2.StorageNodeList + err := r.List(ctx, &nodes, + client.InNamespace(cluster.GetNamespace()), + client.MatchingFields{clusterRefField: cluster.GetName()}) + if err != nil { + return nil + } + requests := make([]reconcile.Request, 0, len(nodes.Items)) + for i := range nodes.Items { + requests = append(requests, reconcile.Request{ + NamespacedName: client.ObjectKeyFromObject(&nodes.Items[i]), + }) + } + return requests +} + +// nodesOn enqueues every storage node running on one Kubernetes worker, which is +// how a cordon reaches the nodes it is about. +func (r *StorageNodeReconciler) nodesOn( + ctx context.Context, worker client.Object, +) []reconcile.Request { + var nodes simplyblockv1alpha2.StorageNodeList + if err := r.List(ctx, &nodes); err != nil { + return nil + } + var requests []reconcile.Request + for i := range nodes.Items { + if nodes.Items[i].Spec.WorkerNode != worker.GetName() { + continue + } + requests = append(requests, reconcile.Request{ + NamespacedName: client.ObjectKeyFromObject(&nodes.Items[i]), + }) + } + return requests +} + +func (r *StorageNodeReconciler) Reconcile( + ctx context.Context, req ctrl.Request, +) (ctrl.Result, error) { + var node simplyblockv1alpha2.StorageNode + if err := r.Get(ctx, req.NamespacedName, &node); err != nil { + return ctrl.Result{}, client.IgnoreNotFound(err) + } + + cluster, err := r.cluster(ctx, &node) + if err != nil { + return ctrl.Result{}, err + } + + if !node.DeletionTimestamp.IsZero() { + return r.teardown(ctx, &node, cluster) + } + + if !controllerutil.ContainsFinalizer(&node, NodeFinalizer) { + controllerutil.AddFinalizer(&node, NodeFinalizer) + return ctrl.Result{}, r.Update(ctx, &node) + } + + // The cluster owns the node, which is what makes a cluster delete cascade to + // its nodes (§3.1). It is established here rather than by whatever created the + // object, so a node written by hand joins the spine too. + if err := r.adopt(ctx, &node, cluster); err != nil { + return ctrl.Result{}, err + } + + if node.Status.UUID == "" { + return r.provision(ctx, &node, cluster) + } + + // A cordoned worker takes its storage-node pod with it, so the node is taken + // down deliberately rather than killed underneath a running SPDK process. The + // window is raised before the status sync, because a node whose host is going + // away should not first be reported healthy. + if raised, err := r.raiseMaintenance(ctx, &node); err != nil || raised { + return ctrl.Result{RequeueAfter: nodeAdvance}, err + } + + return r.syncStatus(ctx, &node, cluster) +} + +// cluster resolves the node's parent. A node whose cluster does not exist is held +// rather than failed: admission refuses such a node on create (§3.4), so one seen +// here is a cluster deleted out from under a node that outlived it. +func (r *StorageNodeReconciler) cluster( + ctx context.Context, node *simplyblockv1alpha2.StorageNode, +) (*simplyblockv1alpha2.StorageCluster, error) { + var cluster simplyblockv1alpha2.StorageCluster + key := types.NamespacedName{Name: node.Spec.ClusterRef, Namespace: node.Namespace} + if err := r.Get(ctx, key, &cluster); err != nil { + if apierrors.IsNotFound(err) { + return nil, nil + } + return nil, err + } + return &cluster, nil +} + +// adopt establishes the cluster as the node's controller owner, which is both the +// ownership spine and what the conversion reads spec.clusterRef back from. +func (r *StorageNodeReconciler) adopt( + ctx context.Context, + node *simplyblockv1alpha2.StorageNode, + cluster *simplyblockv1alpha2.StorageCluster, +) error { + if cluster == nil || metav1.IsControlledBy(node, cluster) { + return nil + } + patch := client.MergeFrom(node.DeepCopy()) + if err := controllerutil.SetControllerReference(cluster, node, r.Scheme); err != nil { + return fmt.Errorf("own node %s by cluster %s: %w", node.Name, cluster.Name, err) + } + return r.Patch(ctx, node, patch) +} + +// provision runs the machine of §4.2 forward by at most one step. +func (r *StorageNodeReconciler) provision( + ctx context.Context, + node *simplyblockv1alpha2.StorageNode, + cluster *simplyblockv1alpha2.StorageCluster, +) (ctrl.Result, error) { + if cluster == nil || cluster.Status.UUID == "" { + r.emit(node, corev1.EventTypeNormal, ClusterNotReady, + "The cluster has no UUID yet, so there is nothing to add this node to") + return ctrl.Result{RequeueAfter: nodeRetry}, r.hold(ctx, node, + "waiting for the cluster to be created in the control plane") + } + + machine, err := statemachine.NewFromSnapshot(ctx, provisioningGraph(), + statemachine.FromKube[nodeStep](node.Status.Step)) + if err != nil { + // An unrecognized step is a downgrade, a hand-edited object, or a rename + // that shipped without a conversion, and none of them resolve by + // reconciling again. + return ctrl.Result{}, r.fail(ctx, node, + fmt.Sprintf("provisioning cannot be resumed: %v", err)) + } + defer machine.Close() + + // A machine is born already in its initial state, so that state's entry hook + // never runs and no deadline is set for it. Setting one on the first pass is + // what stops the first step being the one step that cannot time out. + if node.Status.Step.State == "" { + deadline := metav1.NewTime(time.Now().Add(checkingHostDeadline)) + return ctrl.Result{RequeueAfter: nodeAdvance}, + r.recordStep(ctx, node, machine.CurrentState(), &deadline) + } + + current := machine.CurrentState() + if machine.TimeoutReached() { + return ctrl.Result{}, r.fail(ctx, node, + fmt.Sprintf("step %s outlived its deadline", current)) + } + + next, done, err := r.performNodeStep(ctx, node, cluster, current) + if err != nil { + var blocked *blockedStepError + if errors.As(err, &blocked) { + r.emit(node, corev1.EventTypeWarning, blocked.reason, blocked.message) + return ctrl.Result{RequeueAfter: nodeRetry}, r.hold(ctx, node, blocked.message) + } + logf.FromContext(ctx).Error(err, "the provisioning step could not be advanced", + "node", node.Name, "step", current) + return ctrl.Result{RequeueAfter: nodeRetry}, r.hold(ctx, node, err.Error()) + } + if !done { + return ctrl.Result{RequeueAfter: nodeRetry}, r.hold(ctx, node, + fmt.Sprintf("waiting on %s", current)) + } + + if machine.IsTerminal() { + // Resolving and Adopting both end with a UUID on the object, which the + // step that reached them has already written. The next pass is steady + // state. + return ctrl.Result{RequeueAfter: nodeAdvance}, nil + } + + if err := machine.TransitionTo(ctx, next); err != nil { + return ctrl.Result{}, fmt.Errorf("enter step %s: %w", next, err) + } + snapshot := statemachine.ToKube(machine.Snapshot()) + return ctrl.Result{RequeueAfter: nodeAdvance}, + r.recordStep(ctx, node, next, snapshot.Deadline) +} + +// performNodeStep runs one step and reports the step that follows it and whether +// this one has finished. +// +// The next step is returned rather than read off the graph's single edge, because +// two of the six branch: CheckingHost and CheckingConfig both divert to Adopting, +// and AwaitingSlot skips Posting when a sibling socket has already claimed the +// worker. +func (r *StorageNodeReconciler) performNodeStep( + ctx context.Context, + node *simplyblockv1alpha2.StorageNode, + cluster *simplyblockv1alpha2.StorageCluster, + current nodeStep, +) (nodeStep, bool, error) { + switch current { + case stepCheckingHost: + return r.checkHost(ctx, node, cluster) + case stepCheckingConfig: + return r.checkConfig(ctx, node, cluster) + case stepAwaitingSlot: + return r.awaitSlot(ctx, node, cluster) + case stepPosting: + return stepResolving, true, r.postNode(ctx, node, cluster) + case stepResolving: + done, err := r.resolveUUID(ctx, node, cluster) + return stepResolving, done, err + case stepAdopting: + done, err := r.resolveUUID(ctx, node, cluster) + return stepAdopting, done, err + default: + return current, false, fmt.Errorf("step %s belongs to no provisioning path", current) + } +} + +// checkHost holds until the worker's storage-node API answers, and diverts to +// adoption when this deployment is being taken over wholesale or when a backend +// node is already at the worker's address. +func (r *StorageNodeReconciler) checkHost( + ctx context.Context, + node *simplyblockv1alpha2.StorageNode, + cluster *simplyblockv1alpha2.StorageCluster, +) (nodeStep, bool, error) { + // An upgrade Secret declares that this deployment is being adopted wholesale, + // which is the migration route off a Helm deployment. It diverts before the + // host check, because an adopted node is already running and its API answering + // is not this operator's precondition to establish. + if r.upgradeAdoption(ctx, node.Namespace, cluster.Name) { + return stepAdopting, true, nil + } + + // A backend node already at the worker's address covers a POST whose response + // was lost after the control plane committed, as well as a node this operator + // never added. + if _, found, err := r.matchBackendNode(ctx, node, cluster); err != nil { + return stepCheckingHost, false, err + } else if found { + return stepAdopting, true, nil + } + + answers, err := r.Workload.HostAnswers(ctx, node.Namespace, node.Spec.WorkerNode) + if err != nil { + return stepCheckingHost, false, err + } + if !answers { + return stepCheckingHost, false, blockedf(HostUnreachable, + "the storage-node API on worker %s does not answer yet", node.Spec.WorkerNode) + } + return stepCheckingConfig, true, nil +} + +// checkConfig is a gate rather than a validation. A cluster with +// enableFailureDomains set requires every node to declare a fault group, and a +// node that does not is held rather than rejected: the value can arrive later, and +// holding is what makes filling it in sufficient (§4.2). +func (r *StorageNodeReconciler) checkConfig( + ctx context.Context, + node *simplyblockv1alpha2.StorageNode, + cluster *simplyblockv1alpha2.StorageCluster, +) (nodeStep, bool, error) { + if _, found, err := r.matchBackendNode(ctx, node, cluster); err != nil { + return stepCheckingConfig, false, err + } else if found { + return stepAdopting, true, nil + } + + if ptr.BoolFromOrFalse(cluster.Spec.EnableFailureDomains) && + node.Spec.Config.FailureDomain == "" { + return stepCheckingConfig, false, blockedf(FailureDomainMissing, + "cluster %s requires a fault group and this node declares none; "+ + "set spec.config.failureDomain", cluster.Name) + } + return stepAwaitingSlot, true, nil +} + +// awaitSlot is where two independent serialization rules live. +// +// maxParallelNodeAdds caps how many workers may be in flight at once, counted by +// distinct worker rather than by object so that a two-socket host consumes one +// slot. Workers hosting a FoundationDB pod are always sequential regardless of +// that cap, because a node add reboots the host and two simultaneous FoundationDB +// reboots reduce the control plane's own fault tolerance. Both are predicates over +// the current state of the cluster's other nodes, so the step re-evaluates them on +// every pass and holds rather than failing. +// +// The sibling check is the third thing here and it is not a serialization rule: a +// second socket of a worker some other object has already claimed must not post +// again, and enters Resolving directly. +func (r *StorageNodeReconciler) awaitSlot( + ctx context.Context, + node *simplyblockv1alpha2.StorageNode, + cluster *simplyblockv1alpha2.StorageCluster, +) (nodeStep, bool, error) { + siblings, err := r.clusterNodeObjects(ctx, node) + if err != nil { + return stepAwaitingSlot, false, err + } + + // One POST adds every socket of the worker, so a sibling at Posting or beyond + // means the worker has been claimed. + for i := range siblings { + sibling := &siblings[i] + if sibling.Name == node.Name || sibling.Spec.WorkerNode != node.Spec.WorkerNode { + continue + } + if claimedWorker(sibling) { + return stepResolving, true, nil + } + } + + limit := int32(1) + if spec := cluster.Spec.StorageNodes; spec != nil && spec.MaxParallelNodeAdds != nil { + limit = *spec.MaxParallelNodeAdds + } + + inFlight := map[string]struct{}{} + for i := range siblings { + sibling := &siblings[i] + if sibling.Name == node.Name || sibling.Spec.WorkerNode == node.Spec.WorkerNode { + continue + } + if claimedWorker(sibling) && sibling.Status.UUID == "" { + inFlight[sibling.Spec.WorkerNode] = struct{}{} + } + } + if int32(len(inFlight)) >= limit { + return stepAwaitingSlot, false, blockedf(AwaitingSlot, + "waiting for a node-add slot, %d of %d in flight", len(inFlight), limit) + } + + // A FoundationDB worker waits for every other FoundationDB worker, whatever + // the cap says. + if r.hostsFoundationDB(ctx, node.Namespace, node.Spec.WorkerNode) { + for worker := range inFlight { + if r.hostsFoundationDB(ctx, node.Namespace, worker) { + return stepAwaitingSlot, false, blockedf(AwaitingSlot, + "worker %s hosts FoundationDB and worker %s is already being added", + node.Spec.WorkerNode, worker) + } + } + } + return stepPosting, true, nil +} + +// claimedWorker reports whether an object has claimed its worker, which is the +// transition into Posting or anything past it. +func claimedWorker(node *simplyblockv1alpha2.StorageNode) bool { + switch nodeStep(node.Status.Step.State) { + case stepPosting, stepResolving, stepAdopting: + return true + default: + return node.Status.UUID != "" + } +} + +// postNode adds the worker's nodes. The claim was made by the transition into +// this step, so the call is made once and the step that follows polls for the UUID +// it produces. +func (r *StorageNodeReconciler) postNode( + ctx context.Context, + node *simplyblockv1alpha2.StorageNode, + cluster *simplyblockv1alpha2.StorageCluster, +) error { + params := r.addParams(node, cluster) + if err := r.API.AddNode(ctx, cluster.Status.UUID, params); err != nil { + return fmt.Errorf("add node %s on worker %s: %w", + node.Name, node.Spec.WorkerNode, err) + } + return nil +} + +// resolveUUID matches this node's slot against the cluster's node list and writes +// the UUID when it appears. +// +// The match is positional rather than by identity: the control plane's nodes for +// one worker are sorted by RPC port ascending, and position in that list is the +// socket ordinal, because the ports are assigned in socket order at node-add time +// (§4.3). +func (r *StorageNodeReconciler) resolveUUID( + ctx context.Context, + node *simplyblockv1alpha2.StorageNode, + cluster *simplyblockv1alpha2.StorageCluster, +) (bool, error) { + reading, found, err := r.matchBackendNode(ctx, node, cluster) + if err != nil || !found { + return false, err + } + + err = r.writeStatus(ctx, node, func(status *simplyblockv1alpha2.StorageNodeStatus) { + status.UUID = reading.UUID + applyReading(status, reading) + status.Phase = phaseOf(reading) + status.Step = statemachine.KubeSnapshot{} + status.Message = "" + }) + if err != nil { + return false, err + } + + r.register(node, cluster.Status.UUID, reading.UUID) + if nodeStep(node.Status.Step.State) == stepAdopting { + r.emit(node, corev1.EventTypeNormal, NodeAdopted, + fmt.Sprintf("Backend node %s was adopted rather than added", reading.UUID)) + } + provisioningDurationSeconds.WithLabelValues(node.Spec.ClusterRef). + Observe(time.Since(node.CreationTimestamp.Time).Seconds()) + return true, nil +} + +// matchBackendNode finds the backend node filling this object's slot, by the +// worker's internal IP and the slot's position in the port-sorted list. +func (r *StorageNodeReconciler) matchBackendNode( + ctx context.Context, + node *simplyblockv1alpha2.StorageNode, + cluster *simplyblockv1alpha2.StorageCluster, +) (NodeReading, bool, error) { + address, err := r.workerAddress(ctx, node.Spec.WorkerNode) + if err != nil || address == "" { + return NodeReading{}, false, err + } + + readings, err := r.clusterNodes(ctx, cluster.Status.UUID) + if err != nil { + return NodeReading{}, false, err + } + + var onWorker []NodeReading + for _, reading := range readings { + if reading.ManagementIP == address && reading.UUID != "" { + onWorker = append(onWorker, reading) + } + } + if len(onWorker) == 0 { + return NodeReading{}, false, nil + } + slices.SortFunc(onWorker, func(a, b NodeReading) int { + return int(a.RPCPort) - int(b.RPCPort) + }) + + slot := 0 + if node.Spec.Slot != nil { + slot = int(*node.Spec.Slot) + } + if slot >= len(onWorker) { + // The socket is not online yet. The other sockets of the worker may be, + // which is why this is "not found" rather than an error. + return NodeReading{}, false, nil + } + return onWorker[slot], true, nil +} + +// syncStatus writes what the control plane reports, and returns without patching +// when nothing has moved. +func (r *StorageNodeReconciler) syncStatus( + ctx context.Context, + node *simplyblockv1alpha2.StorageNode, + cluster *simplyblockv1alpha2.StorageCluster, +) (ctrl.Result, error) { + if cluster == nil { + return ctrl.Result{RequeueAfter: nodeRetry}, nil + } + r.register(node, cluster.Status.UUID, node.Status.UUID) + + reading, found, err := r.nodeReading(ctx, cluster.Status.UUID, node.Status.UUID) + if err != nil { + return ctrl.Result{RequeueAfter: nodeRetry}, nil + } + if !found { + // The stored UUID no longer exists. The cluster may have been reset and + // its nodes recreated, so the object goes back to provisioning rather than + // reporting a node that is not there. + return ctrl.Result{RequeueAfter: nodeAdvance}, + r.writeStatus(ctx, node, func(status *simplyblockv1alpha2.StorageNodeStatus) { + status.UUID = "" + status.Status = "" + status.Health = false + status.Phase = simplyblockv1alpha2.StorageNodePhasePending + status.Message = "the control plane no longer reports this node" + }) + } + + wasOnline := node.Status.Phase == simplyblockv1alpha2.StorageNodePhaseOnline + sample, sampled := r.capacitySample(ctx, cluster.Status.UUID, node.Status.UUID) + + err = r.writeStatus(ctx, node, func(status *simplyblockv1alpha2.StorageNodeStatus) { + applyReading(status, reading) + status.Phase = phaseOf(reading) + if sampled && worthWriting(status.Resources.Capacity, sample) { + status.Resources.Capacity = &simplyblockv1alpha2.StorageNodeCapacity{ + TotalBytes: ptr.To(sample.Total), + UsedBytes: ptr.To(sample.Used), + SampledAt: ptr.To(metav1.NewTime(sample.SampledAt)), + } + } + }) + if err != nil { + return ctrl.Result{}, err + } + + if !wasOnline && node.Status.Phase == simplyblockv1alpha2.StorageNodePhaseOnline { + r.emit(node, corev1.EventTypeNormal, NodeOnline, + fmt.Sprintf("Node %s is online and carrying its share", node.Status.UUID)) + } + r.observePhase(node) + return ctrl.Result{RequeueAfter: nodeRetry}, nil +} + +// applyReading copies what the control plane says into the status, carrying the +// previous capacity forward: it is measured elsewhere and on its own schedule. +func applyReading(status *simplyblockv1alpha2.StorageNodeStatus, reading NodeReading) { + previous := (*simplyblockv1alpha2.StorageNodeCapacity)(nil) + if status.Resources != nil { + previous = status.Resources.Capacity + } + + status.Status = reading.Status + status.Health = reading.Health + status.Hostname = reading.Hostname + status.Uptime = reading.Uptime + status.FailureDomain = fmt.Sprintf("%d", reading.FailureDomain) + status.Resources = &simplyblockv1alpha2.StorageNodeResources{ + CPU: ptr.To(reading.CPUCount), + Volumes: ptr.To(reading.Volumes), + Capacity: previous, + } + if reading.Memory > 0 { + status.Resources.Memory = fmt.Sprintf("%d", reading.Memory) + } + // The device summary is absent until the control plane has reported, which is + // what tells a node that has not reported from one that genuinely has no + // devices. The stream carries neither count, so a streamed reading leaves + // whatever the last listed one said. + if reading.DevicesCount > 0 || reading.OnlineDevicesCount > 0 { + status.Resources.Devices = &simplyblockv1alpha2.StorageNodeDevices{ + Online: reading.OnlineDevicesCount, + Total: reading.DevicesCount, + } + } + status.Ports = &simplyblockv1alpha2.StorageNodePorts{ + Management: reading.ManagementIP, + NvmeOf: ptr.To(reading.NVMeOFPort), + Lvol: ptr.To(reading.LvolPort), + Rpc: ptr.To(reading.RPCPort), + } +} + +// phaseOf is the operator's reading of the lifecycle the control plane reports. +// +// The two are deliberately separate: one says how far the operator has got, and +// the other says what the control plane reports, in its own spelling (§3.3). +func phaseOf(reading NodeReading) simplyblockv1alpha2.StorageNodePhase { + switch reading.Status { + case nodeStatusOnline, nodeStatusActive: + if reading.Resources().degraded() { + return simplyblockv1alpha2.StorageNodePhaseDegraded + } + return simplyblockv1alpha2.StorageNodePhaseOnline + case nodeStatusSuspended, nodeStatusOffline: + return simplyblockv1alpha2.StorageNodePhaseOffline + case nodeStatusInCreation, nodeStatusInRestart: + return simplyblockv1alpha2.StorageNodePhaseProvisioning + default: + // unreachable and timeout, plus anything the control plane adds later. A + // value this operator does not know is a node it cannot vouch for. + return simplyblockv1alpha2.StorageNodePhaseFailed + } +} + +// deviceHealth is the node-level half of what StorageDevice reports per device. +type deviceHealth struct{ online, total int32 } + +func (d deviceHealth) degraded() bool { return d.total > 0 && d.online < d.total } + +// Resources is the device pair a phase is decided from. +func (n NodeReading) Resources() deviceHealth { + return deviceHealth{online: n.OnlineDevicesCount, total: n.DevicesCount} +} + +// capacityWriteThresholdPercent is how much a node's used size has to move before +// the new reading is worth recording: one percent of the node's own total, so a +// larger node tolerates a larger absolute drift. +// +// Some threshold is required rather than merely economical. The reconciler watches +// its own objects, so every status write schedules another reconcile; writing a +// freshly sampled number every time would make the node reconcile itself in a loop +// for as long as any I/O was happening. +const capacityWriteThresholdPercent = 1 + +// worthWriting reports whether a sample says something the object does not already +// say. A first reading always does; after that the used size has to have moved by +// at least one percent of the total, or the total itself has to have changed, +// which is what a device joining or leaving looks like. +func worthWriting( + current *simplyblockv1alpha2.StorageNodeCapacity, sample prometheus.Capacity, +) bool { + if !sample.Sampled() { + return false // nothing has measured this node; say nothing about it + } + if current == nil || current.UsedBytes == nil || current.TotalBytes == nil { + return true + } + if *current.TotalBytes != sample.Total { + return true + } + drift := *current.UsedBytes - sample.Used + if drift < 0 { + drift = -drift + } + return drift*100 >= sample.Total*capacityWriteThresholdPercent +} + +// capacitySample reads one node's occupancy, or reports that there is none. +// +// A failure is not an error the caller has to handle: the rest of the status is +// correct without it, and a node whose capacity is momentarily unknown is worth +// publishing with the figure absent rather than not published at all (§12). +func (r *StorageNodeReconciler) capacitySample( + ctx context.Context, clusterID, nodeID string, +) (prometheus.Capacity, bool) { + if r.Capacity == nil || clusterID == "" || nodeID == "" { + return prometheus.Capacity{}, false + } + samples, err := r.Capacity.NodeCapacity(ctx, clusterID) + if err != nil { + logf.FromContext(ctx).V(1).Info("no capacity sample for this node", + "cluster", clusterID, "node", nodeID, "err", err.Error()) + return prometheus.Capacity{}, false + } + sample, ok := samples[nodeID] + return sample, ok +} + +// raiseMaintenance raises a HostMaintenance operation when the node's worker has +// been cordoned, and reports whether it did. +// +// The operator raises this and a user does not, which is what §10 means by the +// trigger being the cordon. A user creating one by hand is accepted and behaves +// identically, which is what makes the flow testable without cordoning anything. +func (r *StorageNodeReconciler) raiseMaintenance( + ctx context.Context, node *simplyblockv1alpha2.StorageNode, +) (bool, error) { + var worker corev1.Node + if err := r.Get(ctx, types.NamespacedName{Name: node.Spec.WorkerNode}, &worker); err != nil { + return false, client.IgnoreNotFound(err) + } + if !worker.Spec.Unschedulable { + return false, nil + } + return r.ensureOps(ctx, node, simplyblockv1alpha2.StorageNodeOpsActionHostMaintenance, + node.Name+"-maintenance") +} + +// teardown drains the node before the object goes. +// +// A node with no status.uuid has no backend node behind it, so its finalizer is +// removed immediately. One that has data on it gets a Remove operation, owned by +// the node through a controller reference, and the finalizer is held until the +// node's lock is clear (§4.5). +func (r *StorageNodeReconciler) teardown( + ctx context.Context, + node *simplyblockv1alpha2.StorageNode, + cluster *simplyblockv1alpha2.StorageCluster, +) (ctrl.Result, error) { + clusterID := "" + if cluster != nil { + clusterID = cluster.Status.UUID + } + if !controllerutil.ContainsFinalizer(node, NodeFinalizer) { + return ctrl.Result{}, nil + } + + if node.Status.UUID == "" { + r.unregister(node, clusterID) + controllerutil.RemoveFinalizer(node, NodeFinalizer) + return ctrl.Result{}, r.Update(ctx, node) + } + + if _, err := r.ensureOps(ctx, node, + simplyblockv1alpha2.StorageNodeOpsActionRemove, node.Name+"-remove"); err != nil { + return ctrl.Result{}, err + } + + // The lock being clear is what says the drain has finished, whatever its + // outcome. A failed removal leaves the lock released and the operation as the + // record of why, so the object is not held forever by a drain nobody is going + // to retry. + if node.Status.ActiveOpsRef != "" { + return ctrl.Result{RequeueAfter: nodeRetry}, nil + } + + r.unregister(node, clusterID) + controllerutil.RemoveFinalizer(node, NodeFinalizer) + return ctrl.Result{}, r.Update(ctx, node) +} + +// ensureOps raises one operation the entity created for itself, idempotently by +// name, and reports whether it created one. +// +// The entity owns the operation it raised for itself. That is the one direction +// ownership runs between the two categories: an operation never owns its target, +// and one an entity created for itself is a subordinate of it (§4.5). +func (r *StorageNodeReconciler) ensureOps( + ctx context.Context, + node *simplyblockv1alpha2.StorageNode, + act simplyblockv1alpha2.StorageNodeOpsAction, + name string, +) (bool, error) { + var existing simplyblockv1alpha2.StorageNodeOps + key := types.NamespacedName{Name: name, Namespace: node.Namespace} + err := r.Get(ctx, key, &existing) + if err == nil { + return false, nil + } + if !apierrors.IsNotFound(err) { + return false, err + } + + ops := &simplyblockv1alpha2.StorageNodeOps{ + ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: node.Namespace}, + Spec: simplyblockv1alpha2.StorageNodeOpsSpec{ + NodeRef: node.Name, + Action: act, + }, + } + if err := controllerutil.SetControllerReference(node, ops, r.Scheme); err != nil { + return false, fmt.Errorf("own the %s operation on node %s: %w", act, node.Name, err) + } + if err := r.Create(ctx, ops); err != nil { + if apierrors.IsAlreadyExists(err) { + return false, nil + } + return false, fmt.Errorf("raise the %s operation on node %s: %w", act, node.Name, err) + } + return true, nil +} + +// register tells every stream which object a backend node id belongs to, and opens +// the node's device stream. +func (r *StorageNodeReconciler) register( + node *simplyblockv1alpha2.StorageNode, clusterID, nodeID string, +) { + if nodeID == "" { + return + } + for _, registry := range r.Registries { + registry.RegisterNode(nodeID, client.ObjectKeyFromObject(node)) + } + if r.DeviceScopes != nil && clusterID != "" { + r.DeviceScopes.Add(cpinformer.Scope{clusterID, nodeID}) + } +} + +// unregister closes the node's device stream and stops naming events after it. +// +// The scope goes first: no further device event can arrive once the stream is +// closed, so the name mappings are dropped second and nothing is left naming +// objects after a node on its way out. +func (r *StorageNodeReconciler) unregister( + node *simplyblockv1alpha2.StorageNode, clusterID string, +) { + if node.Status.UUID == "" { + return + } + if r.DeviceScopes != nil && clusterID != "" { + r.DeviceScopes.Remove(cpinformer.Scope{clusterID, node.Status.UUID}) + } + for _, registry := range r.Registries { + registry.UnregisterNode(node.Status.UUID) + } +} + +// nodeReading is what the control plane currently says about the node, from the +// stream's cache once it has delivered its snapshot and from the control plane +// until then. +func (r *StorageNodeReconciler) nodeReading( + ctx context.Context, clusterID, nodeID string, +) (NodeReading, bool, error) { + if r.Nodes != nil && r.Nodes.Synced(scopeOf(clusterID)) { + if _, dto, ok := r.Nodes.Lookup(nodeID); ok { + return readingFromDTO(dto), true, nil + } + return NodeReading{}, false, nil + } + return r.API.StorageNode(ctx, clusterID, nodeID) +} + +// clusterNodes returns every backend node of the cluster, preferring the stream's +// cache once its snapshot has arrived. The gate is the snapshot rather than a +// preference: an empty unsynced cache and a cluster with no nodes look identical, +// and adoption would read the first as the second and keep waiting for a node that +// is already there. +func (r *StorageNodeReconciler) clusterNodes( + ctx context.Context, clusterID string, +) ([]NodeReading, error) { + if r.Nodes != nil && r.Nodes.Synced(scopeOf(clusterID)) { + cached := r.Nodes.List(scopeOf(clusterID)) + out := make([]NodeReading, 0, len(cached)) + for _, dto := range cached { + out = append(out, readingFromDTO(dto)) + } + return out, nil + } + return r.API.StorageNodes(ctx, clusterID) +} + +// clusterNodeObjects are this node's siblings: every StorageNode of the same +// cluster. +func (r *StorageNodeReconciler) clusterNodeObjects( + ctx context.Context, node *simplyblockv1alpha2.StorageNode, +) ([]simplyblockv1alpha2.StorageNode, error) { + var nodes simplyblockv1alpha2.StorageNodeList + err := r.List(ctx, &nodes, + client.InNamespace(node.Namespace), + client.MatchingFields{clusterRefField: node.Spec.ClusterRef}) + if err != nil { + return nil, fmt.Errorf("list the cluster's nodes: %w", err) + } + return nodes.Items, nil +} + +// workerAddress is the worker's internal IP, which is what the control plane +// reports as a backend node's management address. +func (r *StorageNodeReconciler) workerAddress( + ctx context.Context, worker string, +) (string, error) { + var object corev1.Node + if err := r.Get(ctx, types.NamespacedName{Name: worker}, &object); err != nil { + return "", client.IgnoreNotFound(err) + } + for _, address := range object.Status.Addresses { + if address.Type == corev1.NodeInternalIP { + return address.Address, nil + } + } + return "", nil +} + +// hostsFoundationDB reports whether a worker runs a FoundationDB pod, which is +// what makes its node add sequential regardless of the parallel-add cap. +func (r *StorageNodeReconciler) hostsFoundationDB( + ctx context.Context, namespace, worker string, +) bool { + var pods corev1.PodList + err := r.List(ctx, &pods, + client.InNamespace(namespace), client.HasLabels{utils.LabelFDBClusterName}) + if err != nil { + // An unreadable list is treated as a yes, which serializes rather than + // parallelizes. The + // cost of being wrong that way is a slower expansion; the other way it is + // two simultaneous reboots of the control plane's own store. + return true + } + for i := range pods.Items { + if pods.Items[i].Spec.NodeName == worker { + return true + } + } + return false +} + +// upgradeAdoption reports whether this deployment is being adopted wholesale, +// which is the same signal the cluster's own creation path reads. +func (r *StorageNodeReconciler) upgradeAdoption( + ctx context.Context, namespace, cluster string, +) bool { + var secret corev1.Secret + key := types.NamespacedName{ + Name: fmt.Sprintf("simplyblock-%s-upgrade", cluster), + Namespace: namespace, + } + return r.Get(ctx, key, &secret) == nil +} + +// addParams is what the node-add call carries. The node describes itself, so every +// value but the subsystem cap comes from its own spec.config (§3.1). +func (r *StorageNodeReconciler) addParams( + node *simplyblockv1alpha2.StorageNode, + cluster *simplyblockv1alpha2.StorageCluster, +) utils.StorageNodeSetAddParams { + config := node.Spec.Config + workload := cluster.Spec.StorageNodes + if workload == nil { + workload = &simplyblockv1alpha2.StorageNodesSpec{} + } + + params := utils.StorageNodeSetAddParams{ + NodeAddress: r.Workload.NodeAddress(node.Spec.WorkerNode, node.Namespace), + InterfaceName: workload.MgmtInterface, + SPDKImage: config.SpdkImage, + SPDKProxyImage: config.SpdkProxyImage, + DataNics: workload.DataInterfaces, + Namespace: node.Namespace, + JMPercent: journalPercent(config.JournalManager), + Partitions: partitionsPerDevice(workload), + HaJMCount: journalCount(config.JournalManager), + CRName: cluster.Name, + CRNameSpace: cluster.Namespace, + CRPlural: "storageclusters", + Format4K: ptr.BoolFromOrFalse(workload.EnableFormat4K), + SpdkSystemMemory: config.SpdkSystemMemory, + Expand: ptr.BoolFromOrFalse(config.Expand), + } + + // The control plane's failure domain is an integer, and this API's is a label. + // Only a label that is a number can be sent, which is what a domain seeded + // from an index looks like; anything else is a name the control plane has no + // field for and is left to it to assign. + if index := domainIndex(config.FailureDomain); index != nil { + params.FailureDomain = index + } + return params +} + +// domainIndex reads a failure-domain label as the integer the control plane's own +// field takes, and reports nil for a label that is not one. +// +// The two vocabularies genuinely differ: this API names a fault group after the +// rack, the zone, or the power feed somebody would say out loud, and the control +// plane indexes one. A label seeded from an index sends its number, and a name +// the control plane has no field for is left to it to assign — which is why +// status.failureDomain reports what was assigned rather than what was asked for +// (§3.3). +func domainIndex(domain string) *int { + if domain == "" { + return nil + } + index, err := strconv.Atoi(domain) + if err != nil { + return nil + } + return &index +} + +// partitionsPerDevice is how many partitions each device is carved into, which is +// one when the journal has a device of its own and two when it shares. +func partitionsPerDevice(workload *simplyblockv1alpha2.StorageNodesSpec) int { + if ptr.BoolFromOrFalse(workload.EnableJournalDevice) { + return 1 + } + return 2 +} + +func journalPercent(spec *simplyblockv1alpha2.JournalManagerSpec) int { + if spec == nil { + return 3 + } + return ptr.IntFrom(spec.PercentPerDevice, 3) +} + +func journalCount(spec *simplyblockv1alpha2.JournalManagerSpec) int { + if spec == nil { + return 3 + } + return ptr.IntFrom(spec.Count, 3) +} + +// recordStep persists the step the machine is about to be in, with the instant it +// expires. Both travel together, because a step persisted without its deadline +// restores as a step that can never time out. +func (r *StorageNodeReconciler) recordStep( + ctx context.Context, + node *simplyblockv1alpha2.StorageNode, + next nodeStep, + deadline *metav1.Time, +) error { + return r.writeStatus(ctx, node, func(status *simplyblockv1alpha2.StorageNodeStatus) { + status.Phase = simplyblockv1alpha2.StorageNodePhaseProvisioning + status.Step = statemachine.KubeSnapshot{State: string(next), Deadline: deadline} + }) +} + +// hold reports a provisioning step that is waiting on something outside this +// process. It is not a failure, so the deadline keeps running. +func (r *StorageNodeReconciler) hold( + ctx context.Context, node *simplyblockv1alpha2.StorageNode, message string, +) error { + return r.writeStatus(ctx, node, func(status *simplyblockv1alpha2.StorageNodeStatus) { + if status.Phase == "" { + status.Phase = simplyblockv1alpha2.StorageNodePhasePending + } + status.Message = message + }) +} + +// fail ends provisioning. A node whose machine cannot be resumed, or whose step +// outlived its budget, is Failed with the reason in its message rather than +// retrying forever. +func (r *StorageNodeReconciler) fail( + ctx context.Context, node *simplyblockv1alpha2.StorageNode, message string, +) error { + r.emit(node, corev1.EventTypeWarning, HostUnreachable, message) + return r.writeStatus(ctx, node, func(status *simplyblockv1alpha2.StorageNodeStatus) { + status.Phase = simplyblockv1alpha2.StorageNodePhaseFailed + status.Message = message + }) +} + +// writeStatus applies the mutation and patches only when something changed, which +// is what keeps a node serving I/O from reconciling itself in a loop. +func (r *StorageNodeReconciler) writeStatus( + ctx context.Context, + node *simplyblockv1alpha2.StorageNode, + mutate func(*simplyblockv1alpha2.StorageNodeStatus), +) error { + return retry.RetryOnConflict(retry.DefaultRetry, func() error { + var fresh simplyblockv1alpha2.StorageNode + if err := r.Get(ctx, client.ObjectKeyFromObject(node), &fresh); err != nil { + return err + } + + desired := *fresh.Status.DeepCopy() + mutate(&desired) + desired.ObservedGeneration = fresh.Generation + + if equalNodeStatus(fresh.Status, desired) { + node.Status = desired + node.ResourceVersion = fresh.ResourceVersion + return nil + } + + patch := client.MergeFromWithOptions(fresh.DeepCopy(), + client.MergeFromWithOptimisticLock{}) + fresh.Status = desired + if err := r.Status().Patch(ctx, &fresh, patch); err != nil { + return err + } + node.Status = fresh.Status + node.ResourceVersion = fresh.ResourceVersion + return nil + }) +} + +// emit raises an event about the node's own lifecycle, which is what an +// administrator looking at a worker has open. +func (r *StorageNodeReconciler) emit( + node *simplyblockv1alpha2.StorageNode, eventType, reason, message string, +) { + r.Recorder.Eventf(node, nil, eventType, reason, reason, "%s", message) +} + +// observePhase publishes the node's phase as a gauge, so a node stuck in +// Provisioning is alertable rather than merely visible. +func (r *StorageNodeReconciler) observePhase(node *simplyblockv1alpha2.StorageNode) { + for _, phase := range []simplyblockv1alpha2.StorageNodePhase{ + simplyblockv1alpha2.StorageNodePhasePending, + simplyblockv1alpha2.StorageNodePhaseProvisioning, + simplyblockv1alpha2.StorageNodePhaseOnline, + simplyblockv1alpha2.StorageNodePhaseRemoving, + simplyblockv1alpha2.StorageNodePhaseOffline, + simplyblockv1alpha2.StorageNodePhaseDegraded, + simplyblockv1alpha2.StorageNodePhaseFailed, + } { + value := 0.0 + if node.Status.Phase == phase { + value = 1 + } + nodePhaseState.WithLabelValues(node.Spec.ClusterRef, node.Name, string(phase)).Set(value) + } +} + +// equalNodeStatus compares two statuses for the purpose of deciding whether to +// write. It is spelled out rather than reflect.DeepEqual because the status +// carries pointers, and two equal values behind two pointers are not deeply equal. +func equalNodeStatus(a, b simplyblockv1alpha2.StorageNodeStatus) bool { + if a.Phase != b.Phase || a.UUID != b.UUID || a.Status != b.Status || + a.Health != b.Health || a.Hostname != b.Hostname || a.Uptime != b.Uptime || + a.FailureDomain != b.FailureDomain || a.ActiveOpsRef != b.ActiveOpsRef || + a.Message != b.Message || a.ObservedGeneration != b.ObservedGeneration || + a.Step.State != b.Step.State || !equalTime(a.Step.Deadline, b.Step.Deadline) { + return false + } + return equalResources(a.Resources, b.Resources) && equalPorts(a.Ports, b.Ports) +} + +func equalResources(a, b *simplyblockv1alpha2.StorageNodeResources) bool { + if a == nil || b == nil { + return a == b + } + if !equalInt32(a.CPU, b.CPU) || a.Memory != b.Memory || + !equalInt32(a.Volumes, b.Volumes) { + return false + } + if (a.Devices == nil) != (b.Devices == nil) { + return false + } + if a.Devices != nil && *a.Devices != *b.Devices { + return false + } + return equalCapacity(a.Capacity, b.Capacity) +} + +func equalCapacity(a, b *simplyblockv1alpha2.StorageNodeCapacity) bool { + if a == nil || b == nil { + return a == b + } + return equalInt64(a.TotalBytes, b.TotalBytes) && + equalInt64(a.UsedBytes, b.UsedBytes) && + equalTime(a.SampledAt, b.SampledAt) +} + +func equalPorts(a, b *simplyblockv1alpha2.StorageNodePorts) bool { + if a == nil || b == nil { + return a == b + } + return a.Management == b.Management && equalInt32(a.NvmeOf, b.NvmeOf) && + equalInt32(a.Lvol, b.Lvol) && equalInt32(a.Rpc, b.Rpc) +} + +func equalInt32(a, b *int32) bool { + if a == nil || b == nil { + return a == b + } + return *a == *b +} + +func equalInt64(a, b *int64) bool { + if a == nil || b == nil { + return a == b + } + return *a == *b +} diff --git a/operator/internal/controllers/node/storagenodeops_controller.go b/operator/internal/controllers/node/storagenodeops_controller.go new file mode 100644 index 000000000..4e150c1fc --- /dev/null +++ b/operator/internal/controllers/node/storagenodeops_controller.go @@ -0,0 +1,1025 @@ +// The StorageNodeOps reconciler: it drives one operation against one StorageNode +// to a terminal phase and leaves the object behind as the record of what was +// done, to which node, with which parameters, and how it ended. +// +// Two machines run, not one. The outer phase — Pending, Running, and the three +// terminal values — is identical for every action, so folding it into each +// action's graph would copy that spine seven times and a later fix would land in +// one copy (design-crd-model.md §3.1). The inner one is the action's steps, +// declared in graphs.go. +// +// Nothing here blocks. One reconcile advances at most one step: it asks whether +// the current step has finished, and either requeues or writes the next step down +// and enters it. A step that has not finished is waiting on something outside this +// process, and waiting for it inline would hold a worker for as long as a drain +// takes. +// +// The persisted position is the write-ahead record, and no flag sits beside it +// (§7.2). A step is written before the side effect that step performs, so a +// process dying between the two restarts into a state saying the call may already +// have landed. That is safe rather than merely tolerated: every step's completion +// condition is a predicate over current state, and every call is skipped when its +// target is already at or past the state that call would produce, so a node +// already suspended receives no second suspend. +// +// design-storagenode.md §7 is the specification. + +package node + +import ( + "context" + "errors" + "fmt" + "time" + + corev1 "k8s.io/api/core/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/runtime" + "k8s.io/apimachinery/pkg/types" + "k8s.io/client-go/tools/events" + "k8s.io/client-go/util/retry" + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil" + "sigs.k8s.io/controller-runtime/pkg/handler" + logf "sigs.k8s.io/controller-runtime/pkg/log" + "sigs.k8s.io/controller-runtime/pkg/reconcile" + + "github.com/simplyblock/atlas/statemachine" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/cpinformer" + "github.com/simplyblock/simplyblock-operator/internal/cpinformer/subscriptions" + "github.com/simplyblock/simplyblock-operator/internal/utils" +) + +const ( + // OpsFinalizer is what guarantees the node's lock is released even when the + // object is deleted mid-flight. Without it a `kubectl delete` on a running + // drain would leave the node locked by an object that no longer exists, and + // nothing would ever unlock it (§11). + OpsFinalizer = "storage.simplyblock.io/storagenodeops-finalizer" + + // opsRetry is how long an operation waits before looking again at something + // it cannot hurry: a lock another operation holds, a cluster that is not + // ready, or a step waiting on the control plane. A queued operation is + // normally woken by its node rather than by this, and this is the backstop + // for when that event is missed. + opsRetry = 15 * time.Second + + // opsAdvance is how long a pass that moved the operation forward waits + // before the next one. It is short because there is nothing to wait for: the + // status write this pass made is itself a change the controller watches, so + // this is the backstop for the event rather than the path the next step + // normally arrives on. + opsAdvance = time.Second + + // nodeRefField is the index a node event is mapped back through. It is what + // makes a released lock wake the queue immediately rather than after a + // requeue interval. + nodeRefField = "spec.nodeRef" +) + +// StorageNodeOpsReconciler reconciles a StorageNodeOps. +type StorageNodeOpsReconciler struct { + client.Client + Scheme *runtime.Scheme + Recorder events.EventRecorder + API ControlPlane + + // The two stream caches every completion condition in this package is + // evaluated against (§4.4). Each is optional: a deployment without the + // control-plane informer, and every unit test that does not script one, falls + // back to reading the control plane directly. + // + // They are the same caches the StorageNode reconciler reads, and that is the + // point of a cache rather than a second stream: the node's status and the + // operation's completion condition are one reading, so two controllers asking + // the same question get the same answer. + Nodes NodeCache + Clusters ClusterCache + + // Workload is the storage-plane side of a node: the worker labels, the + // storage-node pod, its published DNS name, and the eviction budget a + // maintenance window holds. Three of the seven actions touch it, and the + // objects behind it belong to the StorageCluster (§5.1). + Workload *Workload +} + +// NodeCache is the part of the storage-node subscription this package reads. +type NodeCache interface { + // Lookup returns one node by its backend id, which is what every completion + // condition in this package is a predicate over. The scope it comes back with + // is the cluster the node was streamed under, which this package already knows + // and does not read. + Lookup(nodeID string) (cpinformer.Scope, subscriptions.NodeDTO, bool) + + // List returns every node of a cluster, which adoption matches against and + // which the parallel-add and peer-selection predicates walk. + List(scope cpinformer.Scope) []subscriptions.NodeDTO + + // Synced reports whether the cluster's initial snapshot has been applied. An + // empty unsynced cache and a cluster with no nodes look identical, and + // reading the first as the second would report a drain complete before it + // started. + Synced(scope cpinformer.Scope) bool +} + +// ClusterCache is the part of the cluster subscription this package reads. The +// gate of §7.1 is the only thing that asks: whether the node's cluster is active +// and not mid-rebalance. +type ClusterCache interface { + Lookup(clusterID string) (subscriptions.ClusterDTO, bool) + SyncedRoot() bool +} + +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodeops,verbs=get;list;watch;create;update;patch;delete +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodeops/status,verbs=get;update;patch +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodeops/finalizers,verbs=update +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodes,verbs=get;list;watch;update;patch +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodes/status,verbs=get;update;patch +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storageclusters,verbs=get;list;watch +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=volumemigrations,verbs=get;list;watch;create;update;patch;delete +// +kubebuilder:rbac:groups="",resources=persistentvolumes,verbs=get;list;watch +// +kubebuilder:rbac:groups="",resources=persistentvolumeclaims,verbs=get;list;watch +// +kubebuilder:rbac:groups="",resources=nodes,verbs=get;list;watch;update;patch +// +kubebuilder:rbac:groups="",resources=pods,verbs=get;list;watch;delete +// +kubebuilder:rbac:groups="",resources=configmaps,verbs=get;list;watch;update;patch +// +kubebuilder:rbac:groups=discovery.k8s.io,resources=endpointslices,verbs=get;list;watch +// +kubebuilder:rbac:groups=policy,resources=poddisruptionbudgets,verbs=get;list;watch;create;update;patch;delete +// +kubebuilder:rbac:groups=events.k8s.io,resources=events,verbs=create;patch + +// SetupWithManager registers the controller and maps an event on a StorageNode +// back to every operation targeting it. +// +// That mapping is what makes the queue move. An operation waiting on a lock has +// nothing of its own to react to, so a controller watching only its own kind +// would leave every queued operation waiting out a requeue interval after the +// lock frees (design-crd-model.md §3.2). +func (r *StorageNodeOpsReconciler) SetupWithManager(mgr ctrl.Manager) error { + err := mgr.GetFieldIndexer().IndexField(context.Background(), + &simplyblockv1alpha2.StorageNodeOps{}, nodeRefField, + func(object client.Object) []string { + ops, ok := object.(*simplyblockv1alpha2.StorageNodeOps) + if !ok { + return nil + } + return []string{ops.Spec.NodeRef} + }) + if err != nil { + return fmt.Errorf("index node operations by their target: %w", err) + } + + return ctrl.NewControllerManagedBy(mgr). + For(&simplyblockv1alpha2.StorageNodeOps{}). + Named("storagenodeops"). + Watches(&simplyblockv1alpha2.StorageNode{}, + handler.EnqueueRequestsFromMapFunc(r.operationsOn)). + Complete(r) +} + +// operationsOn enqueues every operation naming this node that has not finished. +// A terminal one has nothing to react to. +func (r *StorageNodeOpsReconciler) operationsOn( + ctx context.Context, node client.Object, +) []reconcile.Request { + var operations simplyblockv1alpha2.StorageNodeOpsList + err := r.List(ctx, &operations, + client.InNamespace(node.GetNamespace()), + client.MatchingFields{nodeRefField: node.GetName()}) + if err != nil { + return nil + } + var requests []reconcile.Request + for i := range operations.Items { + if terminalOps(operations.Items[i].Status.Phase) { + continue + } + requests = append(requests, reconcile.Request{ + NamespacedName: client.ObjectKeyFromObject(&operations.Items[i]), + }) + } + return requests +} + +func (r *StorageNodeOpsReconciler) Reconcile( + ctx context.Context, req ctrl.Request, +) (ctrl.Result, error) { + var ops simplyblockv1alpha2.StorageNodeOps + if err := r.Get(ctx, req.NamespacedName, &ops); err != nil { + return ctrl.Result{}, client.IgnoreNotFound(err) + } + + if !ops.DeletionTimestamp.IsZero() { + return r.teardown(ctx, &ops) + } + + if !controllerutil.ContainsFinalizer(&ops, OpsFinalizer) { + controllerutil.AddFinalizer(&ops, OpsFinalizer) + return ctrl.Result{}, r.Update(ctx, &ops) + } + + // A terminal operation is a record, and a record does nothing. Releasing the + // lock here as well as on the transition is what covers the pass that crashed + // between persisting the phase and clearing activeOpsRef, which would + // otherwise leave the node locked by a finished operation forever. + if terminalOps(ops.Status.Phase) { + return ctrl.Result{}, r.releaseLock(ctx, &ops) + } + + acquired, err := r.acquireLock(ctx, &ops) + if err != nil { + return ctrl.Result{}, err + } + if !acquired { + return ctrl.Result{RequeueAfter: opsRetry}, nil + } + + return r.advance(ctx, &ops) +} + +// advance runs the action's machine forward by at most one step. +func (r *StorageNodeOpsReconciler) advance( + ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, +) (ctrl.Result, error) { + machine, err := graphs().FromSnapshot(ctx, action(ops.Spec.Action), + statemachine.FromKube[step](ops.Status.Step)) + if err != nil { + // An unrecognized step or action is a downgrade, a hand-edited object, or + // a rename that shipped without a conversion, and none of them resolve by + // reconciling again. The operation is terminal with the reason in + // status.message, which leaves a record saying so (§6.3). + return r.finish(ctx, ops, simplyblockv1alpha2.StorageNodeOpsPhaseFailed, + fmt.Sprintf("the operation cannot be resumed: %v", err)) + } + defer machine.Close() + + // A machine is born already in its initial state, so that state's entry hook + // never runs and no deadline is set for it. Setting one on the first pass is + // what stops the first step being the one step that cannot time out. + if ops.Status.Step.State == "" { + return r.enterInitialStep(ctx, ops, machine) + } + + current := machine.CurrentState() + + if ops.Spec.Abort { + return r.unwind(ctx, ops, current) + } + + // The cluster gate is not the same as the lock. A node operation runs inside + // a cluster, and one whose cluster is mid-rebalance or not active will either + // be rejected by the control plane or succeed into an inconsistent layout. It + // holds rather than fails, and resumes when the cluster does (§7.1). + if ready, reason, err := r.clusterReady(ctx, ops); err != nil { + return ctrl.Result{RequeueAfter: opsRetry}, r.note(ctx, ops, err.Error()) + } else if !ready { + r.emit(ctx, ops, corev1.EventTypeWarning, ClusterNotReady, reason) + return ctrl.Result{RequeueAfter: opsRetry}, r.note(ctx, ops, reason) + } + + if machine.TimeoutReached() { + operationStepDeadlineExceededTotal. + WithLabelValues(r.clusterLabel(ctx, ops), string(ops.Spec.Action), string(current)).Inc() + r.emit(ctx, ops, corev1.EventTypeWarning, StepDeadlineExceeded, + fmt.Sprintf("Step %s outlived its deadline", current)) + return r.fail(ctx, ops, current, + fmt.Sprintf("step %s outlived its deadline", current)) + } + + done, err := r.perform(ctx, ops, current) + if err != nil { + var fatal *terminalStepError + if errors.As(err, &fatal) { + return r.fail(ctx, ops, current, fatal.Error()) + } + var blocked *blockedStepError + if errors.As(err, &blocked) { + // A blocked step is correct behavior waiting on a human or on + // another node coming back, and the event is what distinguishes it + // from a stalled controller (§13.1). It is not a failure, so the + // deadline keeps running and the operation keeps looking. + r.emit(ctx, ops, corev1.EventTypeWarning, blocked.reason, blocked.message) + return r.waitOn(machine), r.note(ctx, ops, blocked.message) + } + logf.FromContext(ctx).Error(err, "the step could not be advanced", + "operation", ops.Name, "step", current) + return ctrl.Result{RequeueAfter: opsRetry}, r.note(ctx, ops, err.Error()) + } + if !done { + return r.waitOn(machine), r.note(ctx, ops, r.waitingMessage(ops, current)) + } + + r.observeStep(ctx, ops, current) + + if machine.IsTerminal() { + return r.finish(ctx, ops, simplyblockv1alpha2.StorageNodeOpsPhaseSucceeded, + r.successMessage(ops)) + } + + next, err := r.nextStep(machine) + if err != nil { + return r.fail(ctx, ops, current, err.Error()) + } + return r.enterStep(ctx, ops, machine, next) +} + +// enterStep moves the machine into the next step and writes the step and the +// deadline its entry hook armed in one patch. +// +// The two go together because a step with no deadline is a step nothing can ever +// time out: TimeoutReached reads the stored deadline, so a crash between a patch +// carrying the state and a later one carrying the deadline would restore an +// operation that retries on the fallback interval for good and never reports the +// failure its budget exists to produce. +// +// Writing after the transition rather than before it costs nothing, because every +// OnEnter in graphs.go returns a duration and performs nothing. The side effect of +// a step is performed on the pass that follows, against the step this patch +// persisted, which is where the write-ahead record is needed and what it records. +func (r *StorageNodeOpsReconciler) enterStep( + ctx context.Context, + ops *simplyblockv1alpha2.StorageNodeOps, + machine *statemachine.Machine[step], + next step, +) (ctrl.Result, error) { + if err := machine.TransitionTo(ctx, next); err != nil { + return ctrl.Result{}, fmt.Errorf("enter step %s: %w", next, err) + } + snapshot := statemachine.ToKube(machine.Snapshot()) + return ctrl.Result{RequeueAfter: opsAdvance}, r.recordStep(ctx, ops, next, snapshot.Deadline) +} + +// enterInitialStep sets the first step's deadline and moves the operation to +// Running. +func (r *StorageNodeOpsReconciler) enterInitialStep( + ctx context.Context, + ops *simplyblockv1alpha2.StorageNodeOps, + machine *statemachine.Machine[step], +) (ctrl.Result, error) { + budget, ok := initialDeadlines[action(ops.Spec.Action)] + if !ok { + budget = requestingDeadline + } + deadline := metav1.NewTime(time.Now().Add(budget)) + if err := r.recordStep(ctx, ops, machine.CurrentState(), &deadline); err != nil { + return ctrl.Result{}, err + } + r.emit(ctx, ops, corev1.EventTypeNormal, OperationStarted, + fmt.Sprintf("The operation acquired the lock on node %s and started", ops.Spec.NodeRef)) + return ctrl.Result{RequeueAfter: opsAdvance}, nil +} + +// nextStep is the step that follows the current one. Every graph in this package +// is a line, so the first edge is the only edge. +func (r *StorageNodeOpsReconciler) nextStep( + machine *statemachine.Machine[step], +) (step, error) { + current := machine.CurrentState() + for next := range machine.AllowedTransitions() { + return next, nil + } + return current, fmt.Errorf("step %s declares no successor and is not terminal", current) +} + +// unwind honors spec.abort where the graph allows it, and reports an abort that +// arrived too late rather than half-undoing the work. +// +// The refusal is the point. A step with no abort edge has already asked the +// control plane for something it is part-way through, and stopping there would +// leave nothing driving the node back to a state somebody can reason about. +// Promoting is the clearest case: the promote has re-homed the logical volumes, +// so there is nothing to unwind and the operation is what finishes the relocation +// (§9). +func (r *StorageNodeOpsReconciler) unwind( + ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, current step, +) (ctrl.Result, error) { + if !abortable(current) { + // Not a failure of the operation: it carries on. What the user asked for + // cannot be done, and saying so is the whole of the response. + return ctrl.Result{RequeueAfter: opsRetry}, r.note(ctx, ops, fmt.Sprintf( + "the abort arrived at step %s, which the control plane is part-way through "+ + "and cannot be stopped; the operation is running on", current)) + } + + // Every terminal outcome from Suspending onward resumes the node first. A + // node past the suspend is not serving, and an operation that stopped there + // and left it that way would take capacity out of the cluster for as long as + // nobody noticed (§8.3). + r.resumeNode(ctx, ops, current) + r.abortMigrations(ctx, ops, current) + + r.emit(ctx, ops, corev1.EventTypeNormal, OperationAborted, + fmt.Sprintf("The operation was aborted at step %s", current)) + return r.finish(ctx, ops, simplyblockv1alpha2.StorageNodeOpsPhaseAborted, + fmt.Sprintf("aborted at step %s", current)) +} + +// fail ends the operation, resuming the node first where the step it failed on +// left it suspended. +func (r *StorageNodeOpsReconciler) fail( + ctx context.Context, + ops *simplyblockv1alpha2.StorageNodeOps, + current step, + message string, +) (ctrl.Result, error) { + r.resumeNode(ctx, ops, current) + return r.finish(ctx, ops, simplyblockv1alpha2.StorageNodeOpsPhaseFailed, message) +} + +// resumeNode is the unwind of §8.3, and it is best-effort on purpose. +// +// A resume that itself fails leaves the node suspended, which is visible in +// status.status and in the NodeResumeFailed event. Retrying it forever would mean +// an operation that can never reach a terminal phase and a lock that is never +// released, and §16 Q3 is whether that is the right trade. +func (r *StorageNodeOpsReconciler) resumeNode( + ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, current step, +) { + if !unwinds(current) { + return + } + clusterID, nodeID, err := r.target(ctx, ops) + if err != nil || nodeID == "" { + return + } + if err := r.API.Resume(ctx, clusterID, nodeID); err != nil { + r.emit(ctx, ops, corev1.EventTypeWarning, NodeResumeFailed, fmt.Sprintf( + "Node %s could not be resumed and is left suspended: %v", ops.Spec.NodeRef, err)) + } +} + +// waitOn requeues for whatever is left of the current step's deadline, so that a +// step with a long budget is looked at when it expires rather than on a fixed +// interval, and a step with a short one is not left waiting past it. +func (r *StorageNodeOpsReconciler) waitOn(machine *statemachine.Machine[step]) ctrl.Result { + if remaining, bounded := machine.RequeueAfter(); bounded && remaining < opsRetry { + return ctrl.Result{RequeueAfter: remaining} + } + return ctrl.Result{RequeueAfter: opsRetry} +} + +// finish writes a terminal phase, releases the node's lock, and records what the +// operation cost. +func (r *StorageNodeOpsReconciler) finish( + ctx context.Context, + ops *simplyblockv1alpha2.StorageNodeOps, + phase simplyblockv1alpha2.StorageNodeOpsPhase, + message string, +) (ctrl.Result, error) { + now := metav1.Now() + err := r.writeStatus(ctx, ops, func(status *simplyblockv1alpha2.StorageNodeOpsStatus) { + status.Phase = phase + status.Message = message + status.CompletedAt = &now + }) + if err != nil { + return ctrl.Result{}, err + } + + switch phase { + case simplyblockv1alpha2.StorageNodeOpsPhaseSucceeded: + r.emit(ctx, ops, corev1.EventTypeNormal, OperationSucceeded, message) + case simplyblockv1alpha2.StorageNodeOpsPhaseFailed: + r.emit(ctx, ops, corev1.EventTypeWarning, OperationFailed, message) + } + + r.observeOperation(ctx, ops, phase) + return ctrl.Result{}, r.releaseLock(ctx, ops) +} + +// observeOperation records the operation's outcome and how long it ran. +func (r *StorageNodeOpsReconciler) observeOperation( + ctx context.Context, + ops *simplyblockv1alpha2.StorageNodeOps, + phase simplyblockv1alpha2.StorageNodeOpsPhase, +) { + cluster, act, result := r.clusterLabel(ctx, ops), string(ops.Spec.Action), resultOf(phase) + operationsTotal.WithLabelValues(cluster, act, result).Inc() + if started := ops.Status.StartedAt; started != nil { + operationDurationSeconds.WithLabelValues(cluster, act, result). + Observe(time.Since(started.Time).Seconds()) + } + drainBlockedVolumesCount.DeleteLabelValues(cluster, blockedPinned) + drainBlockedVolumesCount.DeleteLabelValues(cluster, blockedUnmanaged) +} + +// observeStep records how long one step took. The start is the step's entry, +// which is its deadline minus the budget the graph gives it, so no second +// timestamp has to be persisted for a measurement. +func (r *StorageNodeOpsReconciler) observeStep( + ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, current step, +) { + deadline, bounded := ops.Status.Step.KubeDeadline() + if !bounded { + return + } + budget, ok := stepBudgets[current] + if !ok { + return + } + elapsed := time.Since(deadline.Add(-budget)).Seconds() + if elapsed < 0 { + return + } + cluster := r.clusterLabel(ctx, ops) + operationStepDurationSeconds. + WithLabelValues(cluster, string(ops.Spec.Action), string(current)).Observe(elapsed) + if current == stepHolding { + maintenanceHoldSeconds.WithLabelValues(cluster).Observe(elapsed) + } +} + +// The metric labels a terminal phase is reported under, lowercased because a +// label value is not an API enum. They are this package's own vocabulary rather +// than the control plane's, which is why they are not the device status constants +// beside them that happen to spell one of the words the same way. +const ( + resultSucceeded = "succeeded" + resultAborted = "aborted" + resultFailed = "failed" +) + +// resultOf is the metric label for a terminal phase. +func resultOf(phase simplyblockv1alpha2.StorageNodeOpsPhase) string { + switch phase { + case simplyblockv1alpha2.StorageNodeOpsPhaseSucceeded: + return resultSucceeded + case simplyblockv1alpha2.StorageNodeOpsPhaseAborted: + return resultAborted + default: + return resultFailed + } +} + +// teardown releases the lock, aborts whatever the operation fanned out, and lets +// the object go. +// +// It is the third release path and the one that matters most, because +// `kubectl delete` on a running operation would otherwise leave the node locked +// by an object that no longer exists. Aborting the fan-out first is what stops a +// deleted drain leaving migrations running behind it (§8.4). +func (r *StorageNodeOpsReconciler) teardown( + ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, +) (ctrl.Result, error) { + if !controllerutil.ContainsFinalizer(ops, OpsFinalizer) { + return ctrl.Result{}, nil + } + if pending, err := r.cascadeMigrations(ctx, ops); err != nil { + return ctrl.Result{}, err + } else if pending { + return ctrl.Result{RequeueAfter: opsRetry}, nil + } + if err := r.releaseLock(ctx, ops); err != nil { + return ctrl.Result{}, err + } + controllerutil.RemoveFinalizer(ops, OpsFinalizer) + return ctrl.Result{}, r.Update(ctx, ops) +} + +// acquireLock takes the node's status.activeOpsRef, and reports whether this +// operation now holds it. +// +// Acquisition is an optimistic-lock patch rather than a plain one, which is what +// makes the read-then-write safe: two operations can both read an empty field and +// both conclude the lock is free, and the patch succeeds for exactly one of them +// at a given resourceVersion and returns 409 to the rest (§11). +func (r *StorageNodeOpsReconciler) acquireLock( + ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, +) (bool, error) { + node, err := r.node(ctx, ops) + if apierrors.IsNotFound(err) { + _, err := r.finish(ctx, ops, simplyblockv1alpha2.StorageNodeOpsPhaseFailed, + fmt.Sprintf("StorageNode %s does not exist", ops.Spec.NodeRef)) + return false, err + } + if err != nil { + return false, err + } + + if held := node.Status.ActiveOpsRef; held != "" && held != ops.Name { + r.emit(ctx, ops, corev1.EventTypeNormal, OperationQueued, fmt.Sprintf( + "Node %s is held by operation %s; this one is waiting", node.Name, held)) + return false, r.hold(ctx, ops, fmt.Sprintf( + "waiting for operation %s to release node %s", held, node.Name)) + } + + if node.Status.ActiveOpsRef != ops.Name { + patch := client.MergeFromWithOptions(node.DeepCopy(), + client.MergeFromWithOptimisticLock{}) + node.Status.ActiveOpsRef = ops.Name + if err := r.Status().Patch(ctx, node, patch); err != nil { + if apierrors.IsConflict(err) { + // Somebody else moved the object between the read and the write. + // Whether that was another operation taking the lock is decided + // by reading it again rather than guessed at here. + return false, nil + } + return false, fmt.Errorf("acquire the lock on node %s: %w", node.Name, err) + } + operationActiveState.WithLabelValues(node.Spec.ClusterRef, node.Name).Set(1) + } + + if ops.Status.Phase == "" || ops.Status.Phase == simplyblockv1alpha2.StorageNodeOpsPhasePending { + now := metav1.Now() + // How long the operation waited behind another one's lock. The object's + // creation is the start, because Pending is where an operation both + // begins and waits. + operationLockWaitSeconds.WithLabelValues(node.Spec.ClusterRef, string(ops.Spec.Action)). + Observe(now.Sub(ops.CreationTimestamp.Time).Seconds()) + err := r.writeStatus(ctx, ops, func(status *simplyblockv1alpha2.StorageNodeOpsStatus) { + status.Phase = simplyblockv1alpha2.StorageNodeOpsPhaseRunning + status.StartedAt = &now + status.Message = "The operation holds the node and is running" + }) + if err != nil { + return false, err + } + } + return true, nil +} + +// releaseLock clears the node's status.activeOpsRef, but only while it still +// names this operation. +// +// The ownership check is what makes a late release safe. A pass that started +// before the lock changed hands would otherwise clear a lock somebody else now +// holds, which is worse than not releasing at all: two operations would then be +// running against one node with neither of them knowing. +func (r *StorageNodeOpsReconciler) releaseLock( + ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, +) error { + node, err := r.node(ctx, ops) + if apierrors.IsNotFound(err) { + return nil + } + if err != nil { + return err + } + if node.Status.ActiveOpsRef != ops.Name { + return nil + } + + patch := client.MergeFromWithOptions(node.DeepCopy(), client.MergeFromWithOptimisticLock{}) + node.Status.ActiveOpsRef = "" + if err := r.Status().Patch(ctx, node, patch); err != nil { + // A conflict is reported rather than swallowed, and that is the whole + // point of returning an error here. Somebody else wrote the node's status + // between the read and the write, so the lock this operation still holds + // was not cleared; treating that as a release lets the caller reach a + // terminal phase and the finalizer go, and the node stays locked by an + // object that no longer exists. Reporting it retries on the next pass. + return fmt.Errorf("release the lock on node %s: %w", node.Name, err) + } + operationActiveState.WithLabelValues(node.Spec.ClusterRef, node.Name).Set(0) + return nil +} + +// node reads the operation's target. +func (r *StorageNodeOpsReconciler) node( + ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, +) (*simplyblockv1alpha2.StorageNode, error) { + var node simplyblockv1alpha2.StorageNode + key := types.NamespacedName{Name: ops.Spec.NodeRef, Namespace: ops.Namespace} + if err := r.Get(ctx, key, &node); err != nil { + return nil, err + } + return &node, nil +} + +// target resolves the operation to the pair every control-plane call takes: the +// cluster's UUID and the backend node's. +func (r *StorageNodeOpsReconciler) target( + ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, +) (clusterID, nodeID string, err error) { + node, err := r.node(ctx, ops) + if err != nil { + return "", "", err + } + if node.Status.UUID == "" { + return "", "", fatalf("node %s has not been provisioned in the control plane yet", + ops.Spec.NodeRef) + } + var cluster simplyblockv1alpha2.StorageCluster + key := types.NamespacedName{Name: node.Spec.ClusterRef, Namespace: node.Namespace} + if err := r.Get(ctx, key, &cluster); err != nil { + return "", "", err + } + if cluster.Status.UUID == "" { + return "", "", fatalf("cluster %s has not been created in the control plane yet", + node.Spec.ClusterRef) + } + return cluster.Status.UUID, node.Status.UUID, nil +} + +// clusterReady is the gate of §7.1: a node operation runs inside a cluster, and +// one whose cluster is not active or is mid-rebalance will either be rejected by +// the control plane or succeed into an inconsistent layout. +// +// It reports a reason rather than an error, because holding is the response and a +// reason is what an event carries. A cluster whose reading cannot be taken at all +// is an error, which the caller retries. +func (r *StorageNodeOpsReconciler) clusterReady( + ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, +) (bool, string, error) { + node, err := r.node(ctx, ops) + if err != nil { + return false, "", err + } + var cluster simplyblockv1alpha2.StorageCluster + key := types.NamespacedName{Name: node.Spec.ClusterRef, Namespace: node.Namespace} + if err := r.Get(ctx, key, &cluster); err != nil { + return false, "", err + } + if cluster.Status.UUID == "" { + return false, fmt.Sprintf( + "cluster %s has not been created in the control plane yet", cluster.Name), nil + } + + // The cache is trusted only once the one cluster stream has delivered its + // snapshot. A cluster missing from an unsynced cache and one the control + // plane has forgotten look identical, and reading the first as the second + // would hold every node operation in the deployment. + if r.Clusters == nil || !r.Clusters.SyncedRoot() { + return true, "", nil + } + reading, ok := r.Clusters.Lookup(cluster.Status.UUID) + if !ok { + return false, fmt.Sprintf( + "the control plane no longer reports cluster %s", cluster.Name), nil + } + if reading.Status != utils.ClusterStatusActive { + return false, fmt.Sprintf("cluster %s is %s rather than active", + cluster.Name, reading.Status), nil + } + if reading.Rebalancing { + return false, fmt.Sprintf( + "cluster %s is rebalancing; the operation resumes when it settles", cluster.Name), nil + } + return true, "", nil +} + +// clusterLabel is the metric label for the operation's cluster. It is the object +// name rather than the UUID, matching every other series in this package, and it +// is empty for an operation whose node has gone rather than an error: a metric is +// not worth failing a reconcile over. +func (r *StorageNodeOpsReconciler) clusterLabel( + ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, +) string { + node, err := r.node(ctx, ops) + if err != nil { + return "" + } + return node.Spec.ClusterRef +} + +// nodeReading is what the control plane currently says about the node, from the +// stream's cache once it has delivered its snapshot and from the control plane +// until then. +// +// An operation asks this on every pass of every step, so a drain that takes an +// hour is hundreds of reads the informer is already holding. The gate is the +// snapshot rather than a preference, because a node missing from an unsynced +// cache and one the control plane has forgotten look identical, and reading the +// first as the second would report a shutdown complete while the node is still up. +func (r *StorageNodeOpsReconciler) nodeReading( + ctx context.Context, clusterID, nodeID string, +) (NodeReading, error) { + if r.Nodes != nil && r.Nodes.Synced(cpinformer.Scope{clusterID}) { + if _, dto, ok := r.Nodes.Lookup(nodeID); ok { + return readingFromDTO(dto), nil + } + return NodeReading{}, fmt.Errorf("the control plane no longer reports node %s", nodeID) + } + + reading, ok, err := r.API.StorageNode(ctx, clusterID, nodeID) + if err != nil { + return NodeReading{}, fmt.Errorf("read node %s: %w", nodeID, err) + } + if !ok { + return NodeReading{}, fmt.Errorf("the control plane no longer reports node %s", nodeID) + } + return reading, nil +} + +// readingFromDTO projects a streamed node onto the reading every predicate in +// this package is written against. The stream carries less than the list does — +// no device counts, no memory, and no uptime — so those stay at their zero values +// and no completion condition reads them. +func readingFromDTO(dto subscriptions.NodeDTO) NodeReading { + return NodeReading{ + UUID: dto.ID, + Status: dto.Status, + ManagementIP: dto.ManagementIP, + Health: dto.HealthCheck, + Hostname: dto.Hostname, + CPUCount: dto.CPUCount, + Volumes: dto.Volumes, + RPCPort: dto.RPCPort, + LvolPort: dto.LvolPort, + NVMeOFPort: dto.NVMeOFPort, + FailureDomain: dto.FailureDomain, + } +} + +// hold reports an operation that is admitted, holds nothing, and is waiting. +// Pending is both where an operation starts and where it waits, and the +// OperationQueued event is the only thing that separates the two. +func (r *StorageNodeOpsReconciler) hold( + ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, message string, +) error { + return r.writeStatus(ctx, ops, func(status *simplyblockv1alpha2.StorageNodeOpsStatus) { + if status.Phase == "" { + status.Phase = simplyblockv1alpha2.StorageNodeOpsPhasePending + } + status.Message = message + }) +} + +// note replaces status.message without moving anything else. It is one sentence +// about where the operation is, replaced as it moves, and never a log. +func (r *StorageNodeOpsReconciler) note( + ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, message string, +) error { + return r.writeStatus(ctx, ops, func(status *simplyblockv1alpha2.StorageNodeOpsStatus) { + status.Message = message + }) +} + +// recordStep persists the step the operation is about to be in, with the instant +// it expires. Both travel together, because a step persisted without its deadline +// restores as a step that can never time out. +func (r *StorageNodeOpsReconciler) recordStep( + ctx context.Context, + ops *simplyblockv1alpha2.StorageNodeOps, + next step, + deadline *metav1.Time, +) error { + return r.writeStatus(ctx, ops, func(status *simplyblockv1alpha2.StorageNodeOpsStatus) { + status.Phase = simplyblockv1alpha2.StorageNodeOpsPhaseRunning + status.Step = statemachine.KubeSnapshot{State: string(next), Deadline: deadline} + }) +} + +// writeStatus applies the mutation and patches only when something changed. +// +// observedGeneration is written here rather than by each caller, and on this kind +// its second advance is precisely the signal that spec.abort has been observed: a +// user who sets it and sees an unchanged status cannot otherwise tell a controller +// that has not looked from one that looked and declined. +func (r *StorageNodeOpsReconciler) writeStatus( + ctx context.Context, + ops *simplyblockv1alpha2.StorageNodeOps, + mutate func(*simplyblockv1alpha2.StorageNodeOpsStatus), +) error { + // Retried rather than swallowed on a conflict. A caller that read nil would + // take the write for done, and finish does: it releases the node's lock + // straight afterward, so a dropped terminal status would free the node for + // the next operation while this one still reported Running. + return retry.RetryOnConflict(retry.DefaultRetry, func() error { + var fresh simplyblockv1alpha2.StorageNodeOps + if err := r.Get(ctx, client.ObjectKeyFromObject(ops), &fresh); err != nil { + return err + } + + desired := *fresh.Status.DeepCopy() + mutate(&desired) + desired.ObservedGeneration = fresh.Generation + + if equalOpsStatus(fresh.Status, desired) { + // Still published to the caller, which reads the object it passed in + // on the next line of its own logic. + ops.Status = desired + ops.ResourceVersion = fresh.ResourceVersion + return nil + } + + patch := client.MergeFromWithOptions(fresh.DeepCopy(), + client.MergeFromWithOptimisticLock{}) + fresh.Status = desired + if err := r.Status().Patch(ctx, &fresh, patch); err != nil { + return err + } + ops.Status = fresh.Status + ops.ResourceVersion = fresh.ResourceVersion + return nil + }) +} + +// emit raises an event on the operation and mirrors it onto the target node. +// +// The mirror is not duplication. An operation's events belong on the object that +// outlives it as its audit record, and the node is where somebody investigating a +// stuck cluster starts: they have the worker's name and not the operation's +// (§13.1). +func (r *StorageNodeOpsReconciler) emit( + ctx context.Context, + ops *simplyblockv1alpha2.StorageNodeOps, + eventType, reason, message string, +) { + r.Recorder.Eventf(ops, nil, eventType, reason, reason, "%s", message) + if node, err := r.node(ctx, ops); err == nil { + r.Recorder.Eventf(node, nil, eventType, reason, reason, "%s", message) + } +} + +// terminalOps reports a phase the operation can never leave. +func terminalOps(phase simplyblockv1alpha2.StorageNodeOpsPhase) bool { + switch phase { + case simplyblockv1alpha2.StorageNodeOpsPhaseSucceeded, + simplyblockv1alpha2.StorageNodeOpsPhaseFailed, + simplyblockv1alpha2.StorageNodeOpsPhaseAborted: + return true + default: + return false + } +} + +// terminalStepError is a step failure that retrying cannot fix: an action the +// control plane refused outright, a node with no UUID, a regular expression that +// does not compile. It is a distinct type so that the reconcile loop can tell it +// from a control plane that is briefly unreachable, which is the same shape of +// error and the opposite response. +type terminalStepError struct{ reason string } + +func (e *terminalStepError) Error() string { return e.reason } + +func fatalf(format string, args ...any) error { + return &terminalStepError{reason: fmt.Sprintf(format, args...)} +} + +// blockedStepError is a step that cannot proceed and has not gone wrong: a drain +// held by a pinned claim, one with no online peer to move to, a maintenance +// window waiting for its concurrency slot. +// +// It is a distinct type because the response is the opposite of a failure's. The +// operation holds, keeps its deadline running, and emits the reason it carries, +// because correct behavior that looks exactly like a stalled controller is what +// the event surface exists for (§13.1). +type blockedStepError struct { + reason string + message string +} + +func (e *blockedStepError) Error() string { return e.message } + +func blockedf(reason, format string, args ...any) error { + return &blockedStepError{reason: reason, message: fmt.Sprintf(format, args...)} +} + +// equalOpsStatus compares two statuses for the purpose of deciding whether to +// write. It is spelled out rather than reflect.DeepEqual because the status +// carries pointers to timestamps, and two equal instants behind two pointers are +// not deeply equal. +func equalOpsStatus(a, b simplyblockv1alpha2.StorageNodeOpsStatus) bool { + if a.Phase != b.Phase || a.Message != b.Message || + a.ObservedGeneration != b.ObservedGeneration || + a.Step.State != b.Step.State || + !equalTime(a.Step.Deadline, b.Step.Deadline) || + !equalTime(a.StartedAt, b.StartedAt) || + !equalTime(a.CompletedAt, b.CompletedAt) { + return false + } + if (a.Drain == nil) != (b.Drain == nil) { + return false + } + if a.Drain == nil { + return true + } + return *a.Drain == *b.Drain +} + +func equalTime(a, b *metav1.Time) bool { + if a == nil || b == nil { + return a == b + } + return a.Equal(b) +} + +// successMessage is what a finished operation says it did. A drain has its own, +// because how many volumes moved is the only number a reader wants from it. +func (r *StorageNodeOpsReconciler) successMessage( + ops *simplyblockv1alpha2.StorageNodeOps, +) string { + if ops.Spec.Action == simplyblockv1alpha2.StorageNodeOpsActionRemove { + moved := int32(0) + if d := ops.Status.Drain; d != nil { + moved = d.VolumesMigrated + } + return fmt.Sprintf("node %s removed after moving %d volumes", ops.Spec.NodeRef, moved) + } + return fmt.Sprintf("the %s completed on node %s", ops.Spec.Action, ops.Spec.NodeRef) +} + +// waitingMessage is what an unfinished step says. A drain's says how far through +// its volumes it is as well as which step, because the step alone does not locate +// a drain that runs for hours. +func (r *StorageNodeOpsReconciler) waitingMessage( + ops *simplyblockv1alpha2.StorageNodeOps, current step, +) string { + if d := ops.Status.Drain; d != nil && current == stepMigratingVolumes { + return fmt.Sprintf("%d of %d volumes migrated", d.VolumesMigrated, d.VolumesTotal) + } + return fmt.Sprintf("waiting on %s", current) +} diff --git a/operator/internal/controllers/node/workload.go b/operator/internal/controllers/node/workload.go new file mode 100644 index 000000000..885f21346 --- /dev/null +++ b/operator/internal/controllers/node/workload.go @@ -0,0 +1,519 @@ +// The storage-node workload, as the two reconcilers in this package touch it. +// +// A backend storage node is an SPDK process on a worker, and something has to put +// it there: a DaemonSet, a headless Service, an EndpointSlice per pod, a serving +// certificate, a ServiceAccount with its role, and a ConfigMap the init container +// reads its per-node configuration out of. The retired StorageNodeSet owned all of +// them, which is why deleting one tore the storage plane down and why the kind +// could not be retired until the ownership moved. They are children of the +// StorageCluster now, established by controller reference at the point each is +// created (§5.1). +// +// What is here is the surface the node's own reconcile and its operations ask of +// that workload: label a worker into the storage plane, take one out, ask whether +// a pod is ready, whether its per-pod DNS name is published, whether the worker's +// storage-node API answers, and hold or release an eviction. The objects +// themselves are reconciled in workload_objects.go, by the cluster that owns them. +// +// One key deliberately does not move. Every other annotation and label key in this +// product is migrating to the storage.simplyblock.io prefix (design-crd-model.md +// §7.3), and the per-slot storage-node-uuid label on a worker Node is the +// exception: Kubernetes' external-provisioner caches the set of topology keys in +// the CSINode object when the node plugin registers and hard-errors CreateVolume +// when a live Node's keys do not match the cached set, so the key must not change +// for a worker's lifetime (§5.2). The upgrade tool's key rewrite is what moves it, +// in step with the CSI driver that reads it, and until then this writes the +// spelling the driver knows. +// +// design-storagenode.md §5 is the specification. + +package node + +import ( + "context" + "fmt" + "strings" + + corev1 "k8s.io/api/core/v1" + discoveryv1 "k8s.io/api/discovery/v1" + policyv1 "k8s.io/api/policy/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/util/intstr" + "sigs.k8s.io/controller-runtime/pkg/client" + logf "sigs.k8s.io/controller-runtime/pkg/log" + + atlaskube "github.com/simplyblock/atlas/kube" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/utils" +) + +// storageNodeUUIDLabelPrefix is the per-slot topology label a worker Node carries, +// one key per slot, whose value is the backend node currently filling that slot. +// +// The key is scoped by cluster UUID and slot ordinal, both of which are fixed for +// the slot, while the UUID that says which backend node occupies it is the value +// and is always read fresh. That is what lets one worker host nodes of two +// simplyblock clusters without one cluster's slot colliding with the other's, and +// what makes a node replaced or relocated a change of value alone (§5.2). +const storageNodeUUIDLabelPrefix = "simplyblock.io/storage-node-uuid." + +// Workload is the storage-node workload of one cluster, reachable from a +// reconciler that has a client. +// +// It carries an uncached reader beside the ordinary client for one read. A stale +// informer cache can miss a freshly published EndpointSlice, and the consequence +// is a migration that waits on DNS forever while the name has in fact resolved for +// minutes (§5.4). Every other read here is served from the cache. +type Workload struct { + client.Client + + // Uncached reads straight from the API server. It may be nil, in which case + // the cached client answers: a unit test with a fake client has one reader and + // no cache to be stale. + Uncached client.Reader + + // TLSEnabled and TLSMutualEnabled decide how the worker's storage-node API is + // probed, because a deployment serving TLS refuses a plaintext request and a + // refused request is not the same answer as an unreachable host. + TLSEnabled bool + TLSMutualEnabled bool +} + +// NodeAddress is the per-pod DNS name the control plane is given as node_address +// when a node is added or restarted. +// +// It is the precondition for both: a restart issued against a name that does not +// yet resolve fails name resolution inside the control plane, and the control +// plane's response to that is to reset the node to offline (§5.4). +func (w *Workload) NodeAddress(worker, namespace string) string { + return utils.StorageNodeSetAPIAddress(worker, namespace) +} + +// LabelWorker puts one worker into a cluster's storage plane and rewrites the +// per-slot labels of every node on it. +// +// Rebuilding the label set from a failed List is the failure mode that matters: an +// empty result read as "no slots are desired" would delete every storage-node-uuid +// label from every worker and break CSI topology across the cluster. The reconcile +// aborts on a List error rather than proceeding with a partial view (§5.2). +func (w *Workload) LabelWorker(ctx context.Context, namespace, cluster, worker string) error { + desired, err := w.desiredSlotLabels(ctx, namespace, cluster, worker) + if err != nil { + return err + } + return w.applyWorkerLabels(ctx, namespace, cluster, worker, desired) +} + +// ReleaseWorker removes a cluster's storage-plane labels from the worker a node +// has just left, and only when no other node of that cluster remains on it. +// +// Removing them from a worker that still hosts a node would unschedule it, which +// is why the emptiness check is here rather than at the call site: a relocation +// off a two-socket host leaves the second socket behind. +func (w *Workload) ReleaseWorker( + ctx context.Context, namespace, cluster, movedNode string, +) error { + var nodes simplyblockv1alpha2.StorageNodeList + if err := w.List(ctx, &nodes, client.InNamespace(namespace)); err != nil { + return fmt.Errorf("list the cluster's nodes: %w", err) + } + + // The moved node's object already names the worker it went to, so the worker + // it left is simply one that carries the cluster's label and that no node of + // the cluster is on any more. Counting the occupied workers and stripping the + // rest finds it without having to remember where the node used to be. + occupied := map[string]struct{}{} + for i := range nodes.Items { + node := &nodes.Items[i] + if node.Spec.ClusterRef != cluster { + continue + } + occupied[node.Spec.WorkerNode] = struct{}{} + } + + var workers corev1.NodeList + if err := w.List(ctx, &workers); err != nil { + return fmt.Errorf("list the Kubernetes workers: %w", err) + } + for i := range workers.Items { + worker := &workers.Items[i] + if _, still := occupied[worker.Name]; still { + continue + } + if !w.carriesStoragePlane(worker, cluster) { + continue + } + if err := w.stripStoragePlane(ctx, worker, cluster); err != nil { + return err + } + } + return nil +} + +// desiredSlotLabels is the per-slot label set one worker should carry for one +// cluster: a key per slot that has come online at least once, valued with the +// backend node filling it. +func (w *Workload) desiredSlotLabels( + ctx context.Context, namespace, cluster, worker string, +) (map[string]string, error) { + var clusterObject simplyblockv1alpha2.StorageCluster + key := client.ObjectKey{Namespace: namespace, Name: cluster} + if err := w.Get(ctx, key, &clusterObject); err != nil { + return nil, fmt.Errorf("read cluster %s: %w", cluster, err) + } + if clusterObject.Status.UUID == "" { + // A cluster the control plane has not created has no UUID, so the keys its + // nodes will carry are not derivable yet. The DaemonSet selector is still + // applied, which is what gets a pod onto the worker in the first place. + return map[string]string{}, nil + } + + var nodes simplyblockv1alpha2.StorageNodeList + if err := w.List(ctx, &nodes, client.InNamespace(namespace)); err != nil { + return nil, fmt.Errorf("list the cluster's nodes: %w", err) + } + + desired := map[string]string{} + for i := range nodes.Items { + node := &nodes.Items[i] + if node.Spec.ClusterRef != cluster || node.Spec.WorkerNode != worker { + continue + } + if node.Status.UUID == "" { + continue + } + slot := int32(0) + if node.Spec.Slot != nil { + slot = *node.Spec.Slot + } + desired[fmt.Sprintf("%s.%d", clusterObject.Status.UUID, slot)] = node.Status.UUID + } + return desired, nil +} + +// applyWorkerLabels writes the DaemonSet selector and the slot labels onto one +// worker, and removes the slot keys of this cluster that are no longer wanted. +// +// A slot key belonging to another simplyblock cluster is left alone: a worker may +// host nodes of more than one, and this pass owns only the cluster it was called +// for. +func (w *Workload) applyWorkerLabels( + ctx context.Context, namespace, cluster, worker string, desired map[string]string, +) error { + var node corev1.Node + if err := w.Get(ctx, client.ObjectKey{Name: worker}, &node); err != nil { + return fmt.Errorf("read worker %s: %w", worker, err) + } + if node.Labels == nil { + node.Labels = map[string]string{} + } + + var clusterObject simplyblockv1alpha2.StorageCluster + if err := w.Get(ctx, client.ObjectKey{Namespace: namespace, Name: cluster}, &clusterObject); err != nil { + return fmt.Errorf("read cluster %s: %w", cluster, err) + } + + changed := false + if node.Labels[atlaskube.LabelStorageNodeSet] != cluster { + node.Labels[atlaskube.LabelStorageNodeSet] = cluster + changed = true + } + + if uuid := clusterObject.Status.UUID; uuid != "" { + for key, value := range node.Labels { + if !strings.HasPrefix(key, storageNodeUUIDLabelPrefix) { + continue + } + slot := strings.TrimPrefix(key, storageNodeUUIDLabelPrefix) + separator := strings.LastIndex(slot, ".") + if separator < 0 || slot[:separator] != uuid { + continue + } + if desired[slot] != value { + delete(node.Labels, key) + changed = true + } + } + for slot, uuid := range desired { + key := storageNodeUUIDLabelPrefix + slot + if node.Labels[key] != uuid { + node.Labels[key] = uuid + changed = true + } + } + } + + if !changed { + return nil + } + return w.Update(ctx, &node) +} + +// carriesStoragePlane reports whether a worker is labeled into this cluster's +// storage plane. +func (w *Workload) carriesStoragePlane(worker *corev1.Node, cluster string) bool { + return worker.Labels[atlaskube.LabelStorageNodeSet] == cluster +} + +// stripStoragePlane takes a worker out of the storage plane, selector and slot +// keys together. +func (w *Workload) stripStoragePlane( + ctx context.Context, worker *corev1.Node, cluster string, +) error { + delete(worker.Labels, atlaskube.LabelStorageNodeSet) + for key := range worker.Labels { + if strings.HasPrefix(key, storageNodeUUIDLabelPrefix) { + delete(worker.Labels, key) + } + } + if err := w.Update(ctx, worker); err != nil { + return fmt.Errorf("remove worker %s from cluster %s's storage plane: %w", + worker.Name, cluster, err) + } + return nil +} + +// PodReady reports whether the storage-node pod on one worker is running and +// ready. +func (w *Workload) PodReady( + ctx context.Context, namespace, cluster, worker string, +) (bool, error) { + pod, found, err := w.podOn(ctx, namespace, cluster, worker) + if err != nil || !found { + return false, err + } + if pod.Status.Phase != corev1.PodRunning { + return false, nil + } + for _, condition := range pod.Status.Conditions { + if condition.Type == corev1.PodReady { + return condition.Status == corev1.ConditionTrue, nil + } + } + return false, nil +} + +// PodGone reports whether the storage-node pod has left the worker, which is what +// a maintenance window waits for once it has relaxed the budget. +func (w *Workload) PodGone( + ctx context.Context, namespace, cluster, worker string, +) (bool, error) { + _, found, err := w.podOn(ctx, namespace, cluster, worker) + return !found, err +} + +// podOn finds the storage-node pod scheduled onto one worker. +func (w *Workload) podOn( + ctx context.Context, namespace, cluster, worker string, +) (*corev1.Pod, bool, error) { + var pods corev1.PodList + err := w.List(ctx, &pods, client.InNamespace(namespace), client.MatchingLabels{ + atlaskube.LabelApp: atlaskube.AppStorageNode, + atlaskube.LabelStorageNodeSet: cluster, + }) + if err != nil { + return nil, false, fmt.Errorf("list the cluster's storage-node pods: %w", err) + } + for i := range pods.Items { + if pods.Items[i].Spec.NodeName == worker && pods.Items[i].DeletionTimestamp.IsZero() { + return &pods.Items[i], true, nil + } + } + return nil, false, nil +} + +// PublishedInDNS reports whether the worker's per-pod DNS name is in the headless +// Service's EndpointSlice, which is what the control plane resolves node_address +// through. +// +// The read is uncached for the reason the type's doc comment states. +func (w *Workload) PublishedInDNS( + ctx context.Context, namespace, cluster, worker string, +) (bool, error) { + reader := client.Reader(w.Client) + if w.Uncached != nil { + reader = w.Uncached + } + + var slice discoveryv1.EndpointSlice + key := client.ObjectKey{ + Namespace: namespace, + Name: atlaskube.StorageNodeSetAPIEndpointSliceName(cluster), + } + if err := reader.Get(ctx, key, &slice); err != nil { + if apierrors.IsNotFound(err) { + // The slice is written by the same reconcile that labels the worker, + // so its absence is "not yet" rather than an error. + return false, nil + } + return false, fmt.Errorf("read the storage-node API EndpointSlice: %w", err) + } + + wanted := utils.NodeHostnameLabel(worker) + for i := range slice.Endpoints { + endpoint := &slice.Endpoints[i] + if endpoint.Hostname != nil && *endpoint.Hostname == wanted && + len(endpoint.Addresses) > 0 { + return true, nil + } + } + + // Which of the two it is matters, and the two have different causes: a worker + // absent from the slice means it has not been published, and one present + // without an address means the pod has no IP yet. + published := make([]string, 0, len(slice.Endpoints)) + for i := range slice.Endpoints { + name := "" + if slice.Endpoints[i].Hostname != nil { + name = *slice.Endpoints[i].Hostname + } + published = append(published, fmt.Sprintf("%s=%v", name, slice.Endpoints[i].Addresses)) + } + logf.FromContext(ctx).V(1).Info("the target worker is not published in the storage-node API EndpointSlice", + "want", wanted, "resourceVersion", slice.ResourceVersion, "published", published) + return false, nil +} + +// HostAnswers reports whether the worker's storage-node API is reachable, which is +// the one read in this package with no streamed counterpart: it is a Kubernetes- +// side check against a pod rather than a control-plane object (§4.4). +func (w *Workload) HostAnswers(ctx context.Context, namespace, worker string) (bool, error) { + if err := utils.StorageNodeAPIReachable( + ctx, worker, namespace, w.TLSEnabled, w.TLSMutualEnabled, + ); err != nil { + logf.FromContext(ctx).V(1).Info("the worker's storage-node API does not answer yet", + "worker", worker, "err", err.Error()) + return false, nil + } + return true, nil +} + +// BlockEviction labels the worker's storage pod and creates a budget that allows +// no disruption, so `kubectl drain` blocks on it while the backend node is being +// taken down gracefully (§10). +func (w *Workload) BlockEviction( + ctx context.Context, namespace, cluster, worker string, +) error { + pod, found, err := w.podOn(ctx, namespace, cluster, worker) + if err != nil { + return err + } + if !found { + // Nothing to hold. The pod has already gone, which is the state Releasing + // waits for, so the window is further along than it thought. + return nil + } + if pod.Labels[maintenanceLabel] != worker { + patch := client.MergeFrom(pod.DeepCopy()) + if pod.Labels == nil { + pod.Labels = map[string]string{} + } + pod.Labels[maintenanceLabel] = worker + if err := w.Patch(ctx, pod, patch); err != nil { + return fmt.Errorf("label the storage pod on worker %s: %w", worker, err) + } + } + return w.setBudget(ctx, namespace, cluster, worker, 0) +} + +// AllowEviction relaxes the budget to permit the one eviction the drain is waiting +// on. +func (w *Workload) AllowEviction( + ctx context.Context, namespace, cluster, worker string, +) error { + return w.setBudget(ctx, namespace, cluster, worker, 1) +} + +// ClearEvictionBudget removes the budget and the label the window put in place, so +// the worker is drainable by the ordinary rules again. +// +// A budget left behind by a crashed operator is what would make a worker +// undrainable forever, which is why this runs from the window's terminal step +// rather than only from its success. +func (w *Workload) ClearEvictionBudget( + ctx context.Context, namespace, cluster, worker string, +) error { + budget := &policyv1.PodDisruptionBudget{ObjectMeta: metav1.ObjectMeta{ + Name: maintenanceBudgetName(cluster, worker), + Namespace: namespace, + }} + if err := w.Delete(ctx, budget); err != nil && !apierrors.IsNotFound(err) { + return fmt.Errorf("delete the maintenance budget for worker %s: %w", worker, err) + } + + pod, found, err := w.podOn(ctx, namespace, cluster, worker) + if err != nil || !found { + return err + } + if _, labeled := pod.Labels[maintenanceLabel]; !labeled { + return nil + } + patch := client.MergeFrom(pod.DeepCopy()) + delete(pod.Labels, maintenanceLabel) + if err := w.Patch(ctx, pod, patch); err != nil { + return fmt.Errorf("unlabel the storage pod on worker %s: %w", worker, err) + } + return nil +} + +// setBudget creates or updates the per-worker budget with the given allowance. +func (w *Workload) setBudget( + ctx context.Context, namespace, cluster, worker string, allowed int32, +) error { + desired := &policyv1.PodDisruptionBudget{ + ObjectMeta: metav1.ObjectMeta{ + Name: maintenanceBudgetName(cluster, worker), + Namespace: namespace, + Labels: map[string]string{ + atlaskube.LabelApp: atlaskube.AppStorageNode, + atlaskube.LabelStorageNodeSet: cluster, + }, + }, + Spec: policyv1.PodDisruptionBudgetSpec{ + MaxUnavailable: &intstr.IntOrString{Type: intstr.Int, IntVal: allowed}, + Selector: &metav1.LabelSelector{ + MatchLabels: map[string]string{maintenanceLabel: worker}, + }, + }, + } + + var existing policyv1.PodDisruptionBudget + key := client.ObjectKeyFromObject(desired) + err := w.Get(ctx, key, &existing) + if apierrors.IsNotFound(err) { + if err := w.Create(ctx, desired); err != nil && !apierrors.IsAlreadyExists(err) { + return fmt.Errorf("create the maintenance budget for worker %s: %w", worker, err) + } + return nil + } + if err != nil { + return fmt.Errorf("read the maintenance budget for worker %s: %w", worker, err) + } + + if existing.Spec.MaxUnavailable != nil && existing.Spec.MaxUnavailable.IntVal == allowed { + return nil + } + patch := client.MergeFrom(existing.DeepCopy()) + existing.Spec.MaxUnavailable = desired.Spec.MaxUnavailable + existing.Spec.Selector = desired.Spec.Selector + if err := w.Patch(ctx, &existing, patch); err != nil { + return fmt.Errorf("relax the maintenance budget for worker %s: %w", worker, err) + } + return nil +} + +// maintenanceLabel marks the one pod a maintenance window's budget selects. It is +// the worker's own name rather than a fixed value, so two windows on two workers +// each hold their own pod and neither budget selects the other's. +const maintenanceLabel = "storage.simplyblock.io/maintenance-worker" + +// maintenanceBudgetName is the budget of one window. It is derived through +// atlas-lib's formula so that two long worker names cannot collide on one object +// name, which is what would make a window relax somebody else's budget. +var maintenanceBudgetFormula = atlaskube.Formula{Prefix: "sb-maintenance-"} + +func maintenanceBudgetName(cluster, worker string) string { + return maintenanceBudgetFormula.Derive(cluster, worker).Value +} diff --git a/operator/internal/controllers/node/workload_controller.go b/operator/internal/controllers/node/workload_controller.go new file mode 100644 index 000000000..ec69e2086 --- /dev/null +++ b/operator/internal/controllers/node/workload_controller.go @@ -0,0 +1,403 @@ +// The reconciler that puts a cluster's storage-node workload on the cluster. +// +// It watches StorageCluster and owns everything a storage node needs to run: the +// DaemonSet, the headless Service and its EndpointSlice, the spdk-proxy Service, +// the serving certificates, the ServiceAccount and its role, and the per-node +// ConfigMap. Every object is established as a child of the cluster by controller +// reference, which is what keeps the property that made them owned in the first +// place: deleting the thing they exist for tears them down (§5.1). +// +// It is a second controller on StorageCluster rather than a branch of the +// cluster's own reconciler, and the two write disjoint things. The cluster's +// reconciler owns StorageCluster.status and never touches a DaemonSet; this one +// owns the workload and never writes the cluster. Splitting them is what keeps the +// workload in the package design-crd-model.md §7.10 assigns it to, beside the +// nodes it exists to run. +// +// The retired StorageNodeSet owned all of this, and allowed several sets per +// cluster, each with its own DaemonSet selected by a per-set node label. That +// collapses to one workload per cluster, because growth is nodes rather than sets +// and what differs between hardware generations — the two images, the SPDK memory, +// and the sizing — is per node already (§5.1). + +package node + +import ( + "context" + "fmt" + + appsv1 "k8s.io/api/apps/v1" + corev1 "k8s.io/api/core/v1" + discoveryv1 "k8s.io/api/discovery/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" + "k8s.io/apimachinery/pkg/runtime" + "k8s.io/apimachinery/pkg/types" + "k8s.io/client-go/tools/events" + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil" + "sigs.k8s.io/controller-runtime/pkg/handler" + logf "sigs.k8s.io/controller-runtime/pkg/log" + "sigs.k8s.io/controller-runtime/pkg/reconcile" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/utils" +) + +// SingletonControlPlaneName is the one ControlPlane a deployment has, which is +// where a cluster that states no storage-node image takes one from. +const SingletonControlPlaneName = "simplyblock" + +// StorageNodeWorkloadReconciler reconciles the workload of one StorageCluster. +type StorageNodeWorkloadReconciler struct { + client.Client + Scheme *runtime.Scheme + Recorder events.EventRecorder + + // Namespace is where the operator runs, which is where the ControlPlane + // singleton the default image comes from lives. + Namespace string + + TLSEnabled bool + TLSMutualEnabled bool + TLSProvider string + + // Workload writes the per-node ConfigMap, which is the one object here whose + // contents come from the nodes rather than from the cluster. + Workload *Workload +} + +// +kubebuilder:rbac:groups=apps,resources=daemonsets,verbs=get;list;watch;create;update;patch;delete +// +kubebuilder:rbac:groups="",resources=services;serviceaccounts;configmaps,verbs=get;list;watch;create;update;patch;delete +// +kubebuilder:rbac:groups=discovery.k8s.io,resources=endpointslices,verbs=get;list;watch;create;update;patch;delete +// +kubebuilder:rbac:groups=rbac.authorization.k8s.io,resources=clusterroles;clusterrolebindings,verbs=get;list;watch;create;update;patch +// +kubebuilder:rbac:groups=cert-manager.io,resources=certificates,verbs=get;list;watch;create;update;patch;delete + +// SetupWithManager registers the controller. +// +// It watches the nodes as well as the cluster, because two of the objects here are +// built from them: the per-node ConfigMap holds one entry per worker, and the +// EndpointSlice publishes one DNS name per worker. A node that arrives has to +// reach both before its pod can start. +func (r *StorageNodeWorkloadReconciler) SetupWithManager(mgr ctrl.Manager) error { + return ctrl.NewControllerManagedBy(mgr). + For(&simplyblockv1alpha2.StorageCluster{}). + Named("storagenode-workload"). + Owns(&appsv1.DaemonSet{}). + Owns(&corev1.Service{}). + Owns(&corev1.ConfigMap{}). + Owns(&discoveryv1.EndpointSlice{}). + Watches(&simplyblockv1alpha2.StorageNode{}, + handler.EnqueueRequestsFromMapFunc(r.clusterOf)). + Complete(r) +} + +// clusterOf maps a node back to the cluster whose workload runs it. +func (r *StorageNodeWorkloadReconciler) clusterOf( + _ context.Context, object client.Object, +) []reconcile.Request { + node, ok := object.(*simplyblockv1alpha2.StorageNode) + if !ok || node.Spec.ClusterRef == "" { + return nil + } + return []reconcile.Request{{NamespacedName: types.NamespacedName{ + Name: node.Spec.ClusterRef, + Namespace: node.Namespace, + }}} +} + +func (r *StorageNodeWorkloadReconciler) Reconcile( + ctx context.Context, req ctrl.Request, +) (ctrl.Result, error) { + var cluster simplyblockv1alpha2.StorageCluster + if err := r.Get(ctx, req.NamespacedName, &cluster); err != nil { + return ctrl.Result{}, client.IgnoreNotFound(err) + } + + // Kubernetes garbage collection tears the workload down with the cluster, so + // a cluster on its way out needs nothing done to it here. + if !cluster.DeletionTimestamp.IsZero() { + return ctrl.Result{}, nil + } + + // The ConfigMap is written before the DaemonSet on every pass. A pod that + // starts against a missing or empty entry reaches the node configuration + // script with --max-subsys-count=0 and fails there, which is a long way from + // the cause (§5.3). + if err := r.Workload.ReconcileConfig(ctx, &cluster); err != nil { + return ctrl.Result{RequeueAfter: nodeRetry}, err + } + + for _, step := range []struct { + what string + run func(context.Context, *simplyblockv1alpha2.StorageCluster) error + }{ + {"the service account and its role", r.reconcileRBAC}, + {"the serving certificates", r.reconcileCertificates}, + {"the headless service", r.reconcileService}, + {"the endpoint slice", r.reconcileEndpointSlice}, + {"the daemon set", r.reconcileDaemonSet}, + } { + if err := step.run(ctx, &cluster); err != nil { + return ctrl.Result{RequeueAfter: nodeRetry}, + fmt.Errorf("reconcile %s: %w", step.what, err) + } + } + return ctrl.Result{}, nil +} + +// reconcileDaemonSet applies the pod template every storage node runs under. +// +// The TLS Secret's resourceVersion is stamped onto the template so that a +// certificate rotation rolls the pods. The Secret's name does not change when it +// rotates, so nothing else would notice (§5.4). +func (r *StorageNodeWorkloadReconciler) reconcileDaemonSet( + ctx context.Context, cluster *simplyblockv1alpha2.StorageCluster, +) error { + image, err := r.image(ctx, cluster) + if err != nil { + return err + } + secretVersion, err := r.tlsSecretVersion(ctx, cluster.Namespace) + if err != nil { + return err + } + + desired := utils.BuildStorageNodeDaemonSet(cluster, + r.TLSEnabled, r.TLSMutualEnabled, r.TLSProvider, secretVersion, image) + if err := controllerutil.SetControllerReference(cluster, desired, r.Scheme); err != nil { + return err + } + + var existing appsv1.DaemonSet + err = r.Get(ctx, client.ObjectKeyFromObject(desired), &existing) + if apierrors.IsNotFound(err) { + return r.Create(ctx, desired) + } + if err != nil { + return err + } + desired.ResourceVersion = existing.ResourceVersion + return r.Update(ctx, desired) +} + +// image is the storage-node container image, defaulting to the ControlPlane +// singleton's so that a deployment states the version once (§5.1). +func (r *StorageNodeWorkloadReconciler) image( + ctx context.Context, cluster *simplyblockv1alpha2.StorageCluster, +) (string, error) { + if wl := cluster.Spec.StorageNodes; wl != nil && wl.Image != "" { + return wl.Image, nil + } + + // Read at v1alpha2, the stored version, rather than at the retired v1alpha1. A + // read of the retired version is answered only by the conversion webhook, + // which a fresh install does not deploy, and the cache it would be served from + // lists empty instead of failing: the fallback would report the singleton + // missing on a cluster that has it. + var controlPlane simplyblockv1alpha2.ControlPlane + key := types.NamespacedName{Namespace: r.Namespace, Name: SingletonControlPlaneName} + if err := r.Get(ctx, key, &controlPlane); err != nil { + return "", fmt.Errorf( + "spec.storageNodes.image is unset and ControlPlane %s cannot be read: %w", + SingletonControlPlaneName, err) + } + if source := controlPlane.Spec.Source; source != nil && source.Managed != nil && + source.Managed.Image != "" { + return source.Managed.Image, nil + } + return "", fmt.Errorf( + "spec.storageNodes.image is unset and ControlPlane %s states no managed image", + SingletonControlPlaneName) +} + +// tlsSecretVersion is what a certificate rotation is noticed by. +func (r *StorageNodeWorkloadReconciler) tlsSecretVersion( + ctx context.Context, namespace string, +) (string, error) { + if !r.TLSEnabled { + return "", nil + } + var secret corev1.Secret + key := types.NamespacedName{Namespace: namespace, Name: utils.SecretNameStorageNodeSetAPITLS} + err := r.Get(ctx, key, &secret) + if apierrors.IsNotFound(err) { + return "", nil + } + if err != nil { + return "", err + } + return secret.ResourceVersion, nil +} + +// reconcileService applies the headless Service the per-pod DNS names hang off. +func (r *StorageNodeWorkloadReconciler) reconcileService( + ctx context.Context, cluster *simplyblockv1alpha2.StorageCluster, +) error { + for _, desired := range []*corev1.Service{ + utils.BuildStorageNodeService(cluster, r.TLSEnabled, r.TLSProvider), + utils.BuildSpdkProxyService(cluster, r.TLSEnabled, r.TLSProvider), + } { + if err := controllerutil.SetControllerReference(cluster, desired, r.Scheme); err != nil { + return err + } + var existing corev1.Service + err := r.Get(ctx, client.ObjectKeyFromObject(desired), &existing) + if apierrors.IsNotFound(err) { + if err := r.Create(ctx, desired); err != nil { + return err + } + continue + } + if err != nil { + return err + } + // The cluster IP is assigned by Kubernetes and may not be rewritten, so + // it is carried forward rather than re-stated. + desired.ResourceVersion = existing.ResourceVersion + desired.Spec.ClusterIP = existing.Spec.ClusterIP + if err := r.Update(ctx, desired); err != nil { + return err + } + } + return nil +} + +// reconcileEndpointSlice publishes one per-pod DNS name per worker the cluster's +// nodes run on. +// +// The slice is built from the StorageNode objects rather than from a list on the +// cluster, because the nodes are what say which workers are in the storage plane +// now: a relocation adds the target before the node's own spec.workerNode moves. +func (r *StorageNodeWorkloadReconciler) reconcileEndpointSlice( + ctx context.Context, cluster *simplyblockv1alpha2.StorageCluster, +) error { + log := logf.FromContext(ctx) + + var nodes simplyblockv1alpha2.StorageNodeList + if err := r.List(ctx, &nodes, client.InNamespace(cluster.Namespace)); err != nil { + return fmt.Errorf("list the cluster's nodes: %w", err) + } + + addresses := map[string]string{} + for i := range nodes.Items { + node := &nodes.Items[i] + if node.Spec.ClusterRef != cluster.Name { + continue + } + if _, known := addresses[node.Spec.WorkerNode]; known { + continue + } + address, err := r.workerAddress(ctx, node.Spec.WorkerNode) + if err != nil || address == "" { + log.V(1).Info("a worker has no internal address yet and is not published", + "worker", node.Spec.WorkerNode) + continue + } + addresses[node.Spec.WorkerNode] = address + } + + desired := utils.BuildStorageNodeEndpointSlice(cluster, addresses) + if err := controllerutil.SetControllerReference(cluster, desired, r.Scheme); err != nil { + return err + } + + var existing discoveryv1.EndpointSlice + err := r.Get(ctx, client.ObjectKeyFromObject(desired), &existing) + if apierrors.IsNotFound(err) { + return r.Create(ctx, desired) + } + if err != nil { + return err + } + desired.ResourceVersion = existing.ResourceVersion + return r.Update(ctx, desired) +} + +// workerAddress is the worker's internal IP, which is what an endpoint carries. +func (r *StorageNodeWorkloadReconciler) workerAddress( + ctx context.Context, worker string, +) (string, error) { + var object corev1.Node + if err := r.Get(ctx, types.NamespacedName{Name: worker}, &object); err != nil { + return "", client.IgnoreNotFound(err) + } + for _, address := range object.Status.Addresses { + if address.Type == corev1.NodeInternalIP { + return address.Address, nil + } + } + return "", nil +} + +// reconcileRBAC applies the ServiceAccount the storage-node pods run as and the +// cluster role they need. +// +// The role and its binding are cluster-scoped, so they carry a managed-by label +// rather than an owner reference: a cluster-scoped object cannot be owned by a +// namespaced one, and Kubernetes garbage-collects one that tries. +func (r *StorageNodeWorkloadReconciler) reconcileRBAC( + ctx context.Context, cluster *simplyblockv1alpha2.StorageCluster, +) error { + account := utils.BuildStorageNodeSetServiceAccount(cluster.Namespace) + if err := controllerutil.SetControllerReference(cluster, account, r.Scheme); err != nil { + return err + } + if err := r.apply(ctx, account); err != nil { + return err + } + + isOpenShift := false + if wl := cluster.Spec.StorageNodes; wl != nil && wl.OpenShiftCluster != nil { + isOpenShift = *wl.OpenShiftCluster + } + if err := r.apply(ctx, utils.BuildStorageNodeSetClusterRole(isOpenShift)); err != nil { + return err + } + return r.apply(ctx, utils.BuildStorageNodeSetClusterRoleBinding(cluster.Namespace)) +} + +// reconcileCertificates applies the serving certificates cert-manager issues for +// the two Services, where the deployment uses cert-manager at all. OpenShift's +// service-ca issues its own from an annotation on the Service, so there is nothing +// to apply there. +func (r *StorageNodeWorkloadReconciler) reconcileCertificates( + ctx context.Context, cluster *simplyblockv1alpha2.StorageCluster, +) error { + if !r.TLSEnabled || !utils.IsCertManagerTLSProvider(r.TLSProvider) { + return nil + } + + for _, certificate := range []struct{ service, secret string }{ + {"simplyblock-storage-node-api", utils.SecretNameStorageNodeSetAPITLS}, + {"simplyblock-spdk-proxy", utils.SecretNameSpdkProxyTLS}, + } { + object := utils.BuildServiceServingCertificate( + cluster.Namespace, certificate.service, certificate.secret) + _, err := controllerutil.CreateOrUpdate(ctx, r.Client, object, func() error { + desired := utils.BuildServiceServingCertificate( + cluster.Namespace, certificate.service, certificate.secret) + object.Object["spec"] = desired.Object["spec"] + return controllerutil.SetControllerReference(cluster, object, r.Scheme) + }) + if err != nil { + return fmt.Errorf("apply the serving certificate for %s: %w", certificate.service, err) + } + } + return nil +} + +// apply creates an object or updates it in place, which is what every object here +// but the two with server-assigned fields needs. +func (r *StorageNodeWorkloadReconciler) apply(ctx context.Context, desired client.Object) error { + existing := desired.DeepCopyObject().(client.Object) + err := r.Get(ctx, client.ObjectKeyFromObject(desired), existing) + if apierrors.IsNotFound(err) { + return r.Create(ctx, desired) + } + if err != nil { + return err + } + desired.SetResourceVersion(existing.GetResourceVersion()) + return r.Update(ctx, desired) +} diff --git a/operator/internal/controllers/testsupport/testsupport.go b/operator/internal/controllers/testsupport/testsupport.go new file mode 100644 index 000000000..3bea0c98e --- /dev/null +++ b/operator/internal/controllers/testsupport/testsupport.go @@ -0,0 +1,99 @@ +// The helpers every controller suite in this repository uses, as a real package +// rather than a _test.go file. +// +// The distinction is not stylistic. A _test.go file is compiled only into its own +// package's test binary, so a helper declared in one cannot be imported by +// another, and the controllers are one package per domain +// (design-crd-model.md §7.10). Keeping them here is what stops a third copy of +// the same scheme builder appearing with each domain that moves. +// +// There is no shared *controller* package beside it, and that is deliberate: what +// once looked like one held only the auto-rebalancer's Job scaffolding, which +// belongs to the volume domain. + +package testsupport + +import ( + "testing" + + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// StatusSubresource is the name a client passes to a SubResourceUpdate +// interceptor for a status write, which is how a test makes one fail. +const StatusSubresource = "status" + +// NewScheme builds a scheme carrying both simplyblock API versions, plus whatever +// else the caller adds. +// +// Both versions are unconditional because the group has two of them: a kind with a +// renamed property is read as v1alpha2 by its controller and may still be written +// as v1alpha1 by a controller that has not moved yet, and a fake client that knows +// only one of them panics on the other. +func NewScheme(t *testing.T, addToScheme ...func(*runtime.Scheme) error) *runtime.Scheme { + t.Helper() + + scheme := runtime.NewScheme() + for _, add := range append( + []func(*runtime.Scheme) error{ + simplyblockv1alpha1.AddToScheme, + simplyblockv1alpha2.AddToScheme, + }, + addToScheme..., + ) { + if err := add(scheme); err != nil { + t.Fatalf("failed to add scheme: %v", err) + } + } + return scheme +} + +// NewClient builds a fake client over the scheme, with the given kinds served +// through a status subresource. +// +// The subresource list is not optional decoration: without it a fake client +// applies a status write to the whole object, so a test asserting that a spec is +// left alone passes against a controller that overwrites it. +func NewClient( + t *testing.T, + scheme *runtime.Scheme, + statusSubresources []client.Object, + objects ...client.Object, +) client.Client { + t.Helper() + + builder := fake.NewClientBuilder().WithScheme(scheme) + if len(statusSubresources) > 0 { + builder = builder.WithStatusSubresource(statusSubresources...) + } + if len(objects) > 0 { + builder = builder.WithObjects(objects...) + } + return builder.Build() +} + +// Contains reports whether a slice holds a value, which is what an assertion over +// a set of names or reasons asks. +func Contains(items []string, want string) bool { + for _, item := range items { + if item == want { + return true + } + } + return false +} + +// Cluster is a StorageCluster the control plane has already created, which is the +// precondition almost every controller in this repository holds on. +func Cluster(namespace, name, uuid string) *simplyblockv1alpha2.StorageCluster { + return &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: namespace}, + Status: simplyblockv1alpha2.StorageClusterStatus{UUID: uuid}, + } +} diff --git a/operator/internal/cpinformer/subscriptions/node.go b/operator/internal/cpinformer/subscriptions/node.go index a25de0d35..1899b9ef9 100644 --- a/operator/internal/cpinformer/subscriptions/node.go +++ b/operator/internal/cpinformer/subscriptions/node.go @@ -16,7 +16,7 @@ import ( "sigs.k8s.io/controller-runtime/pkg/event" - simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" "github.com/simplyblock/simplyblock-operator/internal/cpinformer" ) @@ -106,7 +106,7 @@ func (s *NodeSubscription) enqueue(ctx context.Context, nodeID string) { if !ok { return } - sn := &simplyblockv1alpha1.StorageNode{} + sn := &simplyblockv1alpha2.StorageNode{} sn.SetNamespace(key.Namespace) sn.SetName(key.Name) select { diff --git a/operator/internal/metricsapi/install.go b/operator/internal/metricsapi/install.go index 64039f22d..1ed309b3d 100644 --- a/operator/internal/metricsapi/install.go +++ b/operator/internal/metricsapi/install.go @@ -68,34 +68,36 @@ func Install( // exports. // // An unconfigured or unbuildable endpoint is not fatal, and it costs the - // four kinds differently. A volume reading keeps its provisioned size and + // five kinds differently. A volume reading keeps its provisioned size and // loses what it occupies, because the first is known without measuring. A - // device, a pool, and a cluster reading are measurement throughout, so none - // of them is served at all. Serving what can be answered beats serving - // nothing over a dependency this API can answer partially without. + // device, a node, a pool, and a cluster reading are measurement throughout, + // so none of them is served at all. Serving what can be answered beats + // serving nothing over a dependency this API can answer partially without. var capacity CapacitySource var deviceCapacity DeviceCapacitySource var poolCapacity PoolCapacitySource var clusterCapacity ClusterCapacitySource + var nodeCapacity NodeCapacitySource if prometheusURL == "" { log.Info("no Prometheus endpoint configured; capacity samples will be absent") } else if provider, err := prometheus.New(prometheusURL); err != nil { log.Error(err, "capacity samples will be absent", "prometheusURL", prometheusURL) } else { - // One provider satisfies all four: the volume, device, pool, and + // One provider satisfies all five: the volume, device, node, pool, and // cluster readings are the same exporter's gauges under different // prefixes. capacity = provider deviceCapacity = provider poolCapacity = provider clusterCapacity = provider + nodeCapacity = provider } go func() { <-ready server, err := NewServer( Options{BindPort: port, CertDir: CertDir}, volumes, mgr.GetCache(), capacity, deviceCapacity, poolCapacity, - clusterCapacity, log, + clusterCapacity, nodeCapacity, log, ) if err != nil { log.Error(err, "the aggregated metrics API will not be served") diff --git a/operator/internal/metricsapi/nodestorage.go b/operator/internal/metricsapi/nodestorage.go new file mode 100644 index 000000000..37760c49c --- /dev/null +++ b/operator/internal/metricsapi/nodestorage.go @@ -0,0 +1,374 @@ +// The REST storage behind metrics.simplyblock.io/v1alpha2 storagenodemetrics, +// and the `kubectl get snm` columns that go with it. +// +// It is the same shape as the storagedevicemetrics storage next door and joins +// one level up. A device's identity comes from its own StorageDevice object and +// carries the cluster id with it; a node's comes from its StorageNode object, +// which names its cluster rather than the cluster's UUID, so the join reads the +// StorageCluster once per cluster per request to learn the id Prometheus keys on. +// +// A node with no object is not served. The object is the identity, and a reading +// under a name nothing else in the cluster knows is a reading nobody can +// correlate. A node with no sample is not served either: zeros are the reading of +// an empty node rather than the absence of a reading. + +package metricsapi + +import ( + "context" + "fmt" + + apierrors "k8s.io/apimachinery/pkg/api/errors" + "k8s.io/apimachinery/pkg/api/meta" + "k8s.io/apimachinery/pkg/api/resource" + metainternalversion "k8s.io/apimachinery/pkg/apis/meta/internalversion" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/labels" + "k8s.io/apimachinery/pkg/runtime" + "k8s.io/apiserver/pkg/endpoints/request" + "k8s.io/apiserver/pkg/registry/rest" + "sigs.k8s.io/controller-runtime/pkg/client" + logf "sigs.k8s.io/controller-runtime/pkg/log" + + "github.com/go-logr/logr" + + "github.com/simplyblock/atlas/prometheus" + + metricsv1alpha2 "github.com/simplyblock/simplyblock-operator/api/metrics/v1alpha2" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// NodeResourceName is the plural resource these readings are served under. It is +// its own singular as well, the way `endpoints` is: "metrics" is already the +// noun, and "storagenodemetric" is not a word. +const NodeResourceName = "storagenodemetrics" + +// NodeShortName is the abbreviation `kubectl get snm` resolves. +const NodeShortName = "snm" + +// NodeCapacitySource supplies a node's occupancy. It is satisfied by atlas-lib's +// prometheus.Provider, and it is an interface here so that a test needs no +// Prometheus and a deployment without one can pass nil. +// +// The control plane is not the source. Neither its node list nor its node stream +// carries how full a node is; the numbers exist only in the metrics the same +// service exports (design-storagenode.md §12). +type NodeCapacitySource interface { + // NodeCapacity returns the sample for every node of a cluster, keyed by + // control-plane node UUID. A node with no sample is absent. + NodeCapacity(ctx context.Context, clusterUUID string) (map[string]prometheus.Capacity, error) +} + +// NodeStorage serves storagenodemetrics. It holds no state of its own: the +// identities come from the manager's Kubernetes cache and the readings from +// Prometheus, and it is the join of the two. +type NodeStorage struct { + reader client.Reader + capacity NodeCapacitySource +} + +// NewNodeStorage returns the REST storage over the given Kubernetes reader. +// +// capacity may be nil, and a deployment with no Prometheus is why. Nothing is +// then served, because every field of a node's reading is a measurement. +func NewNodeStorage(reader client.Reader, capacity NodeCapacitySource) *NodeStorage { + return &NodeStorage{reader: reader, capacity: capacity} +} + +// New implements rest.Storage. +func (s *NodeStorage) New() runtime.Object { return &metricsv1alpha2.StorageNodeMetrics{} } + +// Destroy implements rest.Storage. There is nothing to release: no client, no +// watch, and no connection is owned here. +func (s *NodeStorage) Destroy() {} + +// NamespaceScoped implements rest.Scoper. The resource is namespaced because that +// is what confines a reader to the namespaces they already have. +func (s *NodeStorage) NamespaceScoped() bool { return true } + +// GetSingularName implements rest.SingularNameProvider. +func (s *NodeStorage) GetSingularName() string { return NodeResourceName } + +// ShortNames implements rest.ShortNamesProvider. +func (s *NodeStorage) ShortNames() []string { return []string{NodeShortName} } + +// NewList implements rest.Lister. +func (s *NodeStorage) NewList() runtime.Object { + return &metricsv1alpha2.StorageNodeMetricsList{} +} + +// Get implements rest.Getter. The name is a StorageNode's, so the lookup is a +// read of that object and never a scan. +func (s *NodeStorage) Get( + ctx context.Context, name string, _ *metav1.GetOptions, +) (runtime.Object, error) { + namespace := request.NamespaceValue(ctx) + + var node simplyblockv1alpha2.StorageNode + err := s.reader.Get(ctx, client.ObjectKey{Namespace: namespace, Name: name}, &node) + switch { + case apierrors.IsNotFound(err): + return nil, apierrors.NewNotFound(metricsv1alpha2.Resource(NodeResourceName), name) + case err != nil: + return nil, apierrors.NewInternalError(err) + } + + lookup := newNodeCapacityLookup(s.reader, s.capacity, logf.FromContext(ctx)) + sample, ok := lookup.forNode(ctx, &node) + if !ok { + // The node is real but nothing has measured it: a cold exporter, a + // Prometheus that cannot be reached, or a node the control plane has not + // sampled yet. + return nil, apierrors.NewNotFound(metricsv1alpha2.Resource(NodeResourceName), name) + } + return newNodeReading(&node, sample), nil +} + +// List implements rest.Lister. It walks the node objects rather than the samples, +// because the objects are what carry a name and a namespace and a sample under +// neither is not servable. +func (s *NodeStorage) List( + ctx context.Context, options *metainternalversion.ListOptions, +) (runtime.Object, error) { + namespace := request.NamespaceValue(ctx) // "" for a cluster-wide list + selector := labels.Everything() + if options != nil && options.LabelSelector != nil { + selector = options.LabelSelector + } + + var nodes simplyblockv1alpha2.StorageNodeList + var opts []client.ListOption + if namespace != "" { + opts = append(opts, client.InNamespace(namespace)) + } + if err := s.reader.List(ctx, &nodes, opts...); err != nil { + return nil, apierrors.NewInternalError(err) + } + + lookup := newNodeCapacityLookup(s.reader, s.capacity, logf.FromContext(ctx)) + + out := &metricsv1alpha2.StorageNodeMetricsList{} + for i := range nodes.Items { + node := &nodes.Items[i] + if !selector.Matches(labels.Set(node.Labels)) { + continue + } + sample, ok := lookup.forNode(ctx, node) + if !ok { + continue + } + reading := newNodeReading(node, sample) + if !matchesNodeFieldSelector(options, reading) { + continue + } + out.Items = append(out.Items, *reading) + } + return out, nil +} + +// matchesNodeFieldSelector applies the only two field selectors this resource can +// answer. They are supported because a client that passes one and is silently +// ignored gets a wrong answer rather than an error. Anything else selects +// nothing, which is the honest response to a field the object has no index for. +func matchesNodeFieldSelector( + options *metainternalversion.ListOptions, reading *metricsv1alpha2.StorageNodeMetrics, +) bool { + if options == nil || options.FieldSelector == nil || options.FieldSelector.Empty() { + return true + } + for _, req := range options.FieldSelector.Requirements() { + var actual string + switch req.Field { + case fieldSelectorName: + actual = reading.Name + case fieldSelectorNamespace: + actual = reading.Namespace + default: + return false + } + if (req.Operator == "=" || req.Operator == "==") && actual != req.Value { + return false + } + if req.Operator == "!=" && actual == req.Value { + return false + } + } + return true +} + +// nodeCapacityLookup fetches samples once per cluster for the duration of one +// request, because a list walks many nodes of the same few clusters and each +// query is an HTTP round trip. The cluster UUID behind a node's spec.clusterRef +// is memoized for the same reason. +// +// A cluster whose query fails is recorded as having no samples and is not retried +// within the request. Prometheus being down therefore costs the readings and not +// the request. +type nodeCapacityLookup struct { + reader client.Reader + source NodeCapacitySource + log logr.Logger + uuids map[client.ObjectKey]string + byCluster map[string]map[string]prometheus.Capacity +} + +func newNodeCapacityLookup( + reader client.Reader, source NodeCapacitySource, log logr.Logger, +) *nodeCapacityLookup { + return &nodeCapacityLookup{ + reader: reader, + source: source, + log: log, + uuids: map[client.ObjectKey]string{}, + byCluster: map[string]map[string]prometheus.Capacity{}, + } +} + +// forNode returns the sample for one node and whether there is one. The two cases +// are distinguished, because every field of a node's reading is measured: with no +// sample there is nothing to serve. +func (l *nodeCapacityLookup) forNode( + ctx context.Context, node *simplyblockv1alpha2.StorageNode, +) (prometheus.Capacity, bool) { + if l == nil || l.source == nil || node.Status.UUID == "" { + return prometheus.Capacity{}, false + } + clusterUUID := l.clusterUUID(ctx, node) + if clusterUUID == "" { + return prometheus.Capacity{}, false + } + samples, ok := l.byCluster[clusterUUID] + if !ok { + var err error + samples, err = l.source.NodeCapacity(ctx, clusterUUID) + if err != nil { + l.log.V(1).Info("no node capacity samples for this request", + "cluster", clusterUUID, "err", err.Error()) + samples = nil + } + l.byCluster[clusterUUID] = samples + } + sample, ok := samples[node.Status.UUID] + if !ok || !sample.Sampled() { + return prometheus.Capacity{}, false + } + return sample, true +} + +// clusterUUID resolves a node's spec.clusterRef to the identifier Prometheus keys +// on. A cluster that cannot be read, or one with no UUID yet, memoizes the empty +// string, so a namespace whose cluster is still being created costs one read for +// the whole request rather than one per node. +func (l *nodeCapacityLookup) clusterUUID( + ctx context.Context, node *simplyblockv1alpha2.StorageNode, +) string { + key := client.ObjectKey{Namespace: node.Namespace, Name: node.Spec.ClusterRef} + if uuid, ok := l.uuids[key]; ok { + return uuid + } + var cluster simplyblockv1alpha2.StorageCluster + if err := l.reader.Get(ctx, key, &cluster); err != nil { + l.log.V(1).Info("no cluster for this node's readings", + "cluster", key.String(), "err", err.Error()) + l.uuids[key] = "" + return "" + } + l.uuids[key] = cluster.Status.UUID + return cluster.Status.UUID +} + +// newNodeReading assembles the served object from the node's own object and the +// last sample taken of it. +func newNodeReading( + node *simplyblockv1alpha2.StorageNode, sample prometheus.Capacity, +) *metricsv1alpha2.StorageNodeMetrics { + return &metricsv1alpha2.StorageNodeMetrics{ + TypeMeta: metav1.TypeMeta{ + APIVersion: metricsv1alpha2.GroupVersion.String(), + Kind: "StorageNodeMetrics", + }, + ObjectMeta: metav1.ObjectMeta{ + Name: node.Name, + Namespace: node.Namespace, + Labels: node.Labels, + CreationTimestamp: node.CreationTimestamp, + }, + Timestamp: metav1.NewTime(sample.SampledAt), + NodeID: node.Status.UUID, + StorageCluster: node.Spec.ClusterRef, + WorkerNode: node.Spec.WorkerNode, + Capacity: metricsv1alpha2.StorageNodeCapacity{ + Total: *resource.NewQuantity(sample.Total, resource.BinarySI), + Used: *resource.NewQuantity(sample.Used, resource.BinarySI), + Free: *resource.NewQuantity(sample.Free, resource.BinarySI), + Provisioned: *resource.NewQuantity(sample.Provisioned, resource.BinarySI), + UtilizationPercent: sample.UtilizationPercent, + }, + } +} + +// nodeColumns are the table headers, in print order. What they answer: which +// node, on which worker, in which cluster, how big it is, how much of it is in +// use, and how stale the reading is. +var nodeColumns = []metav1.TableColumnDefinition{ + {Name: "Name", Type: "string", Format: "name", Description: "The StorageNode this reading is of"}, + {Name: "Worker", Type: "string", Description: "The Kubernetes worker the node runs on"}, + {Name: "Cluster", Type: "string", Description: "The StorageCluster the node belongs to"}, + {Name: "Total", Type: "string", Description: "The storage the node's devices provide"}, + {Name: "Used", Type: "string", Description: "The space they hold"}, + {Name: "Used%", Type: "string", Description: "The control plane's utilization figure"}, + {Name: "Node", Type: "string", Priority: 1, Description: "The control plane's node UUID"}, + {Name: "Sampled", Type: "string", Description: "How long ago the control plane took the reading"}, +} + +// ConvertToTable implements rest.TableConvertor for both a single reading and a +// list of them. +func (s *NodeStorage) ConvertToTable( + _ context.Context, object runtime.Object, _ runtime.Object, +) (*metav1.Table, error) { + table := &metav1.Table{ColumnDefinitions: nodeColumns} + + switch typed := object.(type) { + case *metricsv1alpha2.StorageNodeMetrics: + table.Rows = append(table.Rows, nodeRow(typed)) + case *metricsv1alpha2.StorageNodeMetricsList: + table.ResourceVersion = typed.ResourceVersion + for i := range typed.Items { + table.Rows = append(table.Rows, nodeRow(&typed.Items[i])) + } + default: + return nil, fmt.Errorf("metricsapi: cannot render %T as a table", object) + } + + if m, err := meta.ListAccessor(object); err == nil { + table.ResourceVersion = m.GetResourceVersion() + table.Continue = m.GetContinue() + } + return table, nil +} + +func nodeRow(reading *metricsv1alpha2.StorageNodeMetrics) metav1.TableRow { + return metav1.TableRow{ + Cells: []any{ + reading.Name, + reading.WorkerNode, + reading.StorageCluster, + reading.Capacity.Total.String(), + reading.Capacity.Used.String(), + fmt.Sprintf("%d%%", reading.Capacity.UtilizationPercent), + reading.NodeID, + translateSampleAge(reading.Timestamp), + }, + Object: runtime.RawExtension{Object: reading}, + } +} + +var ( + _ rest.Storage = (*NodeStorage)(nil) + _ rest.Scoper = (*NodeStorage)(nil) + _ rest.Getter = (*NodeStorage)(nil) + _ rest.Lister = (*NodeStorage)(nil) + _ rest.TableConvertor = (*NodeStorage)(nil) + _ rest.ShortNamesProvider = (*NodeStorage)(nil) + _ rest.SingularNameProvider = (*NodeStorage)(nil) +) diff --git a/operator/internal/metricsapi/scheme.go b/operator/internal/metricsapi/scheme.go index fb55723fd..02f834755 100644 --- a/operator/internal/metricsapi/scheme.go +++ b/operator/internal/metricsapi/scheme.go @@ -58,6 +58,8 @@ func addToScheme(scheme *runtime.Scheme) error { &metricsv1alpha2.StoragePoolMetricsList{}, &metricsv1alpha2.StorageClusterMetrics{}, &metricsv1alpha2.StorageClusterMetricsList{}, + &metricsv1alpha2.StorageNodeMetrics{}, + &metricsv1alpha2.StorageNodeMetricsList{}, ) } // The meta kinds are registered twice as well, and for a reason that is not diff --git a/operator/internal/metricsapi/server.go b/operator/internal/metricsapi/server.go index e9325c6ba..0cc9aba9f 100644 --- a/operator/internal/metricsapi/server.go +++ b/operator/internal/metricsapi/server.go @@ -96,6 +96,7 @@ func NewServer( deviceCapacity DeviceCapacitySource, poolCapacity PoolCapacitySource, clusterCapacity ClusterCapacitySource, + nodeCapacity NodeCapacitySource, log logr.Logger, ) (*Server, error) { opts.withDefaults() @@ -157,6 +158,7 @@ func NewServer( DeviceResourceName: NewDeviceStorage(reader, deviceCapacity), PoolResourceName: NewPoolStorage(reader, poolCapacity), ClusterResourceName: NewClusterStorage(reader, clusterCapacity), + NodeResourceName: NewNodeStorage(reader, nodeCapacity), } if err := server.InstallAPIGroup(&group); err != nil { return nil, fmt.Errorf("metricsapi: install api group: %w", err) diff --git a/operator/internal/upgrade/catalog/catalog.go b/operator/internal/upgrade/catalog/catalog.go index e1496bbd3..23118250c 100644 --- a/operator/internal/upgrade/catalog/catalog.go +++ b/operator/internal/upgrade/catalog/catalog.go @@ -78,6 +78,7 @@ func derivations() []upgrade.Derivation { func migrationSteps() []upgrade.Step { all := steps.Upgrade() all = append(all, steps.Ownership()...) + all = append(all, steps.Sizing()...) return append(all, steps.Migrate()...) } diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml index 30144c3c1..695d12254 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml @@ -259,11 +259,14 @@ spec: description: Count is the number of journal managers to configure. format: int32 + minimum: 1 type: integer percentPerDevice: - description: PercentPerDevice is the journal manager - capacity percentage per device. + description: PercentPerDevice is the share of each + device given to the journal. format: int32 + maximum: 100 + minimum: 1 type: integer type: object mgmtInterface: diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusters.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusters.yaml index 278c299a9..99cc036e1 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusters.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusters.yaml @@ -856,6 +856,281 @@ spec: x-kubernetes-validations: - message: field is immutable rule: self == oldSelf + storageNodes: + description: |- + StorageNodes is the Kubernetes workload the cluster's storage nodes run + as, and the cluster owns every object in it by controller reference: a + cluster deleted takes its DaemonSet, Services, certificate, and per-node + ConfigMap with it. One workload serves the whole cluster, because growth is + nodes rather than sets and what differs between hardware generations is per + node already. + properties: + containerResources: + description: |- + ContainerResources sets requests and limits for the storage-node container. + Unset enforces no limits. + properties: + claims: + description: |- + Claims lists the names of resources, defined in spec.resourceClaims, + that are used by this container. + + This field depends on the + DynamicResourceAllocation feature gate. + + This field is immutable. It can only be set for containers. + items: + description: ResourceClaim references one entry in PodSpec.ResourceClaims. + properties: + name: + description: |- + Name must match the name of one entry in pod.spec.resourceClaims of + the Pod where this field is used. It makes that resource available + inside a container. + type: string + request: + description: |- + Request is the name chosen for a request in the referenced claim. + If empty, everything from the claim is made available, otherwise + only the result of this request. + type: string + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + limits: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Limits describes the maximum amount of compute resources allowed. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + requests: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Requests describes the minimum amount of compute resources required. + If Requests is omitted for a container, it defaults to Limits if that is explicitly specified, + otherwise to an implementation-defined value. Requests cannot exceed Limits. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + type: object + dataInterfaces: + description: DataInterfaces are the data-plane network interfaces. + items: + type: string + type: array + enableCpuTopology: + description: EnableCpuTopology turns on topology-aware CPU assignment. + type: boolean + enableFormat4K: + description: |- + EnableFormat4K formats NVMe devices to a 4K block size where the device + supports it. + type: boolean + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + enableJournalDevice: + description: |- + EnableJournalDevice dedicates the smallest NVMe device on each node to the + journal manager, instead of carving a journal partition out of every + device. + type: boolean + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + enableKubeletConfiguration: + description: |- + EnableKubeletConfiguration lets the storage node apply the kubelet + configuration changes it needs. Off by default, which is the behavior the + retired skipKubeletConfiguration expressed by being set. + type: boolean + image: + description: |- + Image is the storage-node container image. Defaults to the ControlPlane + singleton's spec.image when unset, so a deployment states the version once. + pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ + type: string + imagePullPolicy: + default: IfNotPresent + description: ImagePullPolicy controls when that image is pulled. + enum: + - Always + - Never + - IfNotPresent + type: string + initContainerResources: + description: InitContainerResources does the same for the init container. + properties: + claims: + description: |- + Claims lists the names of resources, defined in spec.resourceClaims, + that are used by this container. + + This field depends on the + DynamicResourceAllocation feature gate. + + This field is immutable. It can only be set for containers. + items: + description: ResourceClaim references one entry in PodSpec.ResourceClaims. + properties: + name: + description: |- + Name must match the name of one entry in pod.spec.resourceClaims of + the Pod where this field is used. It makes that resource available + inside a container. + type: string + request: + description: |- + Request is the name chosen for a request in the referenced claim. + If empty, everything from the claim is made available, otherwise + only the result of this request. + type: string + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + limits: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Limits describes the maximum amount of compute resources allowed. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + requests: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Requests describes the minimum amount of compute resources required. + If Requests is omitted for a container, it defaults to Limits if that is explicitly specified, + otherwise to an implementation-defined value. Requests cannot exceed Limits. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + type: object + maxParallelNodeAdds: + default: 1 + description: |- + MaxParallelNodeAdds limits how many workers may be in the node-add process + at once, counted by distinct worker rather than by object so that a + two-socket host consumes one slot. Workers hosting a FoundationDB pod are + always sequential regardless of this value, because a node add reboots the + host and two simultaneous FoundationDB reboots reduce the control plane's + own fault tolerance. + format: int32 + minimum: 1 + type: integer + mgmtInterface: + description: MgmtInterface is the management network interface storage nodes bind. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + nodesPerSocket: + default: 1 + description: NodesPerSocket is how many storage nodes run per NUMA socket. + format: int32 + minimum: 1 + type: integer + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + openShiftCluster: + description: OpenShiftCluster states that the Kubernetes distribution is OpenShift. + type: boolean + openShiftMachineConfigPool: + default: worker + description: |- + OpenShiftMachineConfigPool names the pool generated MachineConfig objects + are labeled into. + type: string + reservedSystemCPU: + description: ReservedSystemCPU is the CPU set held back from SPDK for system workloads. + type: string + socketsToUse: + description: |- + SocketsToUse restricts deployment to selected NUMA sockets. Empty means + socket 0 alone. + items: + type: string + type: array + tolerations: + description: Tolerations are applied to the storage-node pods. + items: + description: |- + The pod this Toleration is attached to tolerates any taint that matches + the triple using the matching operator . + properties: + effect: + description: |- + Effect indicates the taint effect to match. Empty means match all taint effects. + When specified, allowed values are NoSchedule, PreferNoSchedule and NoExecute. + type: string + key: + description: |- + Key is the taint key that the toleration applies to. Empty means match all taint keys. + If the key is empty, operator must be Exists; this combination means to match all values and all keys. + type: string + operator: + description: |- + Operator represents a key's relationship to the value. + Valid operators are Exists, Equal, Lt, and Gt. Defaults to Equal. + Exists is equivalent to wildcard for value, so that a pod can + tolerate all taints of a particular category. + Lt and Gt perform numeric comparisons (requires feature gate TaintTolerationComparisonOperators). + type: string + tolerationSeconds: + description: |- + TolerationSeconds represents the period of time the toleration (which must be + of effect NoExecute, otherwise this field is ignored) tolerates the taint. By default, + it is not set, which means tolerate the taint forever (do not evict). Zero and + negative values will be treated as 0 (evict immediately) by the system. + format: int64 + type: integer + value: + description: |- + Value is the taint value the toleration matches to. + If the operator is Exists, the value should be empty, otherwise just a regular string. + type: string + type: object + type: array + ubuntuHost: + description: |- + UbuntuHost states that the worker's host OS is Ubuntu, which changes how + the node configures huge pages and the kernel modules it loads. + type: boolean + type: object + x-kubernetes-validations: + - message: field mgmtInterface is immutable once set + rule: '!has(oldSelf.mgmtInterface) || has(self.mgmtInterface)' + - message: field nodesPerSocket is immutable once set + rule: '!has(oldSelf.nodesPerSocket) || has(self.nodesPerSocket)' + - message: field enableJournalDevice is immutable once set + rule: '!has(oldSelf.enableJournalDevice) || has(self.enableJournalDevice)' + - message: field enableFormat4K is immutable once set + rule: '!has(oldSelf.enableFormat4K) || has(self.enableFormat4K)' stripe: description: |- Stripe is the erasure-coding layout every volume in the cluster is diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagenodeops.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagenodeops.yaml index a74479f86..9d63d8ae7 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagenodeops.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagenodeops.yaml @@ -195,11 +195,12 @@ spec: - jsonPath: .status.phase name: Phase type: string - - jsonPath: .status.subPhase - name: SubPhase + - jsonPath: .status.step.state + name: Step type: string - jsonPath: .status.message name: Message + priority: 1 type: string - jsonPath: .metadata.creationTimestamp name: Age @@ -208,10 +209,9 @@ spec: schema: openAPIV3Schema: description: |- - StorageNodeOps is a one-shot operational CR targeting a single StorageNode. - Analogous to a Kubernetes Job — it drives an action (Shutdown, Restart, Suspend, - Resume, Remove, Migrate) to completion and records the result. Only one - StorageNodeOps can be active per StorageNode at a time. + StorageNodeOps is a single operation performed against one StorageNode. It runs + to a terminal phase and stays afterward as the audit record of what was done, to + which node, with which parameters, and how it ended. properties: apiVersion: description: |- @@ -231,10 +231,18 @@ spec: metadata: type: object spec: - description: StorageNodeOpsSpec defines the desired state of a StorageNodeOps. + description: StorageNodeOpsSpec is one operation to perform against one StorageNode. properties: + abort: + description: |- + Abort asks a running operation to stop at its next step and unwind. It is + the only mutable field on this spec, because it is the only thing about an + operation that can legitimately be decided after it started. Whether an + abort is expressible from the current step is declared by that action's + graph rather than checked here. + type: boolean action: - description: Action is the operation to perform. Immutable. + description: Action is the operation to perform. enum: - Shutdown - Restart @@ -242,28 +250,30 @@ spec: - Resume - Remove - Migrate + - HostMaintenance type: string x-kubernetes-validations: - message: field is immutable rule: self == oldSelf force: - description: Force enables forced execution where the backend supports it. + description: |- + Force passes the control plane's force flag where the action supports it. + Migrate defaults it to true, because the control plane rejects a non-forced + restart of a node that is not already offline. type: boolean migrate: description: Migrate parameterizes action Migrate and is ignored by the others. properties: newSsdPcie: description: |- - NewSsdPcie lists additional NVMe PCIe addresses to bind on the target host - during a migration. Passed through to the control-plane restart as - new_ssd_pcie. + NewSsdPcie lists additional NVMe PCI addresses to bind on the target host, + passed through to the control-plane restart as new_ssd_pcie and merged into + the node's effective allow list so they survive a later rebuild. items: type: string type: array targetWorkerNode: - description: |- - TargetWorkerNode is the Kubernetes worker hostname the storage node is - relocated onto. + description: TargetWorkerNode is the Kubernetes worker the node is relocated onto. type: string x-kubernetes-validations: - message: field is immutable @@ -272,25 +282,28 @@ spec: - targetWorkerNode type: object nodeRef: - description: NodeRef is the name of the target StorageNode. Immutable. + description: |- + NodeRef names the StorageNode this operation acts on. The operation never + owns its target, because deleting the record of an operation must not delete + the node it operated on. type: string x-kubernetes-validations: - message: field is immutable rule: self == oldSelf reattachVolume: description: |- - ReattachVolume reattaches volumes during the node restart. - Applicable when action=Restart or action=Migrate. + ReattachVolume asks the control plane to reattach this node's volumes as + part of a restart. Applies to Restart, Migrate, and HostMaintenance. type: boolean remove: description: Remove parameterizes action Remove and is ignored by the others. properties: systemVolumeFilterRegex: + default: ^sb-fio-baseline-.* description: |- - SystemVolumeFilterRegex is a Go regular expression matched against backend - volume names. Matching volumes are treated as system volumes: excluded from - drain migration and deleted inline during the Verifying phase. - Defaults to `^sb-fio-baseline-.*`. + SystemVolumeFilterRegex matches backend volume names that are system + volumes: excluded from the drain's migration and deleted during + verification rather than blocking it. type: string type: object required: @@ -298,50 +311,79 @@ spec: - nodeRef type: object status: - description: StorageNodeOpsStatus holds the observed state of a StorageNodeOps. + description: StorageNodeOpsStatus is the observed state of one node operation. properties: completedAt: - description: CompletedAt is when the operation finished (successfully or not). + description: CompletedAt is when it reached a terminal phase. format: date-time type: string + drain: + description: |- + Drain is the drain's progress over the node's volumes, set only for action + Remove. + properties: + volumesMigrated: + description: VolumesMigrated is how many of them have completed. + format: int32 + minimum: 0 + type: integer + volumesTotal: + description: |- + VolumesTotal is the number of PV-managed volumes the drain has to move, + written once at the end of Validating and not modified afterward. + format: int32 + minimum: 0 + type: integer + required: + - volumesMigrated + - volumesTotal + type: object message: - description: Message is a human-readable description of the current state or failure reason. + description: |- + Message is the reason the phase is what it is: one sentence, replaced as the + operation moves, and never a log. type: string + observedGeneration: + description: |- + ObservedGeneration is the generation the rest of this status was computed + from, so a stale status can be told from a current one. + format: int64 + type: integer phase: - description: Phase is the high-level lifecycle phase. + description: Phase is the operation's own progress. enum: - Pending - Running - Succeeded - Failed + - Aborted type: string startedAt: - description: StartedAt is when the operation began. + description: StartedAt is when the operation acquired its target's lock. format: date-time type: string - subPhase: - description: SubPhase tracks the active drain step when action=Remove and phase=Running. - enum: - - Validating - - Suspending - - Migrating - - Verifying - - Removing - - Preparing - - Restarting - - Promoting - type: string - triggered: + step: description: |- - Triggered indicates the backend action POST has been sent (used during - Suspending to avoid duplicate POSTs across reconcile iterations). - type: boolean - volumesMigrated: - description: VolumesMigrated is the count of volumes successfully migrated (drain only). - type: integer - volumesPending: - description: VolumesPending is the count of volumes awaiting migration (drain only). - type: integer + Step is the position of the running action's state machine, as the shared + statemachine.KubeSnapshot. The rule is what an Enum marker would do if a + marker could reach a field of a shared type. + properties: + deadline: + description: |- + Deadline is when that state expires, absent when it has none. It is an + absolute instant, so a state whose deadline passed while the controller + was down restores as already expired. + format: date-time + type: string + state: + description: |- + State is the state the machine was in. Empty means the resource has not + been reconciled yet, and restores to the graph's initial state. + type: string + type: object + x-kubernetes-validations: + - message: unknown step + rule: '!has(self.state) || self.state in [''Requesting'',''Awaiting'',''Validating'',''Suspending'',''MigratingVolumes'',''Verifying'',''Removing'',''Preparing'',''Relocating'',''AwaitingNode'',''Promoting'',''Holding'',''ShuttingDown'',''Releasing'',''AwaitingHost'',''Restarting'',''Cleanup'']' type: object type: object served: true diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagenodes.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagenodes.yaml index b6a7f5f88..afcbe3faf 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagenodes.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagenodes.yaml @@ -12,339 +12,826 @@ spec: listKind: StorageNodeList plural: storagenodes shortNames: - - sn + - sn singular: storagenode scope: Namespaced versions: - - additionalPrinterColumns: - - jsonPath: .spec.workerNode - name: Worker - type: string - - jsonPath: .spec.socketId - name: Socket - type: string - - jsonPath: .spec.nodeIndex - name: NodeIdx - type: integer - - jsonPath: .status.failureDomain - name: FD - priority: 1 - type: integer - - jsonPath: .status.uuid - name: UUID - type: string - - jsonPath: .status.status - name: Status - type: string - - jsonPath: .status.health - name: Health - type: boolean - - jsonPath: .metadata.creationTimestamp - name: Age - type: date - name: v1alpha1 - schema: - openAPIV3Schema: - description: |- - StorageNode is the Schema for a single backend storage node instance. - One StorageNode CR exists per (workerNode, socketIndex) pair and is owned - by the parent StorageNodeSet. - properties: - apiVersion: - description: |- - APIVersion defines the versioned schema of this representation of an object. - Servers should convert recognized schemas to the latest internal value, and - may reject unrecognized values. - More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources - type: string - kind: - description: |- - Kind is a string value representing the REST resource this object represents. - Servers may infer this from the endpoint the client submits requests to. - Cannot be updated. - In CamelCase. - More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds - type: string - metadata: - type: object - spec: - description: StorageNodeSpec defines the desired state of a StorageNode. - properties: - nodeIndex: - description: NodeIndex is the per-socket node index (0..nodesPerSocket-1). - Immutable. - format: int32 - type: integer - x-kubernetes-validations: - - message: field is immutable - rule: self == oldSelf - overrides: - description: |- - Overrides holds per-node configuration propagated from - StorageNodeSet.spec.nodeConfigs[workerNode] on every reconcile. - properties: - deviceNames: - description: |- - DeviceNames explicitly defines the NVMe namespace names to use on this node - (e.g. ["nvme0n1","nvme1n1"]). - items: + - additionalPrinterColumns: + - jsonPath: .spec.workerNode + name: Worker + type: string + - jsonPath: .spec.socketId + name: Socket + type: string + - jsonPath: .spec.nodeIndex + name: NodeIdx + type: integer + - jsonPath: .status.failureDomain + name: FD + priority: 1 + type: integer + - jsonPath: .status.uuid + name: UUID + type: string + - jsonPath: .status.status + name: Status + type: string + - jsonPath: .status.health + name: Health + type: boolean + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha1 + schema: + openAPIV3Schema: + description: |- + StorageNode is the Schema for a single backend storage node instance. + One StorageNode CR exists per (workerNode, socketIndex) pair and is owned + by the parent StorageNodeSet. + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: StorageNodeSpec defines the desired state of a StorageNode. + properties: + nodeIndex: + description: NodeIndex is the per-socket node index (0..nodesPerSocket-1). Immutable. + format: int32 + type: integer + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + overrides: + description: |- + Overrides holds per-node configuration propagated from + StorageNodeSet.spec.nodeConfigs[workerNode] on every reconcile. + properties: + deviceNames: + description: |- + DeviceNames explicitly defines the NVMe namespace names to use on this node + (e.g. ["nvme0n1","nvme1n1"]). + items: + type: string + type: array + driveSizeRange: + description: DriveSizeRange overrides the drive size range filter for this node. + type: string + enableCpuTopology: + description: EnableCpuTopology overrides topology-aware CPU handling for this node. + type: boolean + expand: + description: |- + Expand marks this node as a cluster-expansion add. When true the backend + node-add endpoint receives expand=true, triggering rebalancing behaviour + appropriate for in-place cluster growth. Overrides StorageNodeSet.spec.expand. + type: boolean + failureDomain: + description: |- + FailureDomain is the failure-domain group index (≥ 0) for this node. + Required when the parent StorageCluster has enableFailureDomains=true. + Overrides StorageNodeSet.spec.nodeFailureDomains[workerNode] when both are set. + format: int32 + minimum: 0 + type: integer + journalManager: + description: JournalManagerSpec overrides journal manager tuning for this node. + properties: + count: + description: Count is the number of journal managers to configure. + format: int32 + type: integer + percentPerDevice: + description: PercentPerDevice is the journal manager capacity percentage per device. + format: int32 + type: integer + type: object + pcieAllowList: + description: PcieAllowList overrides the list of PCI addresses allowed for use on this node. + items: + type: string + type: array + pcieDenyList: + description: PcieDenyList overrides the list of PCI addresses excluded from use on this node. + items: + type: string + type: array + pcieModel: + description: PcieModel overrides the PCI model filter for this node. + type: string + reservedSystemCPU: + description: ReservedSystemCPU overrides the CPUs reserved for system workloads on this node. + type: string + skipKubeletConfiguration: + description: |- + SkipKubeletConfiguration overrides whether kubelet configuration changes are + skipped for this node. + type: boolean + spdkImage: + description: SpdkImage overrides the SPDK image for this node (e.g. for phased rollouts). + type: string + spdkProxyImage: + description: SpdkProxyImage overrides the SPDK proxy image for this node. type: string - type: array - driveSizeRange: - description: DriveSizeRange overrides the drive size range filter - for this node. - type: string - enableCpuTopology: - description: EnableCpuTopology overrides topology-aware CPU handling - for this node. - type: boolean - expand: - description: |- - Expand marks this node as a cluster-expansion add. When true the backend - node-add endpoint receives expand=true, triggering rebalancing behaviour - appropriate for in-place cluster growth. Overrides StorageNodeSet.spec.expand. - type: boolean - failureDomain: - description: |- - FailureDomain is the failure-domain group index (≥ 0) for this node. - Required when the parent StorageCluster has enableFailureDomains=true. - Overrides StorageNodeSet.spec.nodeFailureDomains[workerNode] when both are set. - format: int32 - minimum: 0 - type: integer - journalManager: - description: JournalManagerSpec overrides journal manager tuning - for this node. - properties: - count: - description: Count is the number of journal managers to configure. - format: int32 - type: integer - percentPerDevice: - description: PercentPerDevice is the journal manager capacity - percentage per device. - format: int32 - type: integer - type: object - pcieAllowList: - description: PcieAllowList overrides the list of PCI addresses - allowed for use on this node. - items: + spdkSystemMemory: + description: |- + SpdkSystemMemory overrides the SPDK huge-page memory allocation for this node + (e.g. "4G", "512M"). + pattern: ^[0-9]+(G|GI|GB|GiB|M|MI|MB|MiB|g|gi|gb|gib|m|mi|mb|mib)?$ type: string - type: array - pcieDenyList: - description: PcieDenyList overrides the list of PCI addresses - excluded from use on this node. - items: + ubuntuHost: + description: UbuntuHost overrides the Ubuntu host OS flag for this node. + type: boolean + type: object + socketId: + description: SocketID is the NUMA socket identifier from spec.socketsToUse (e.g. "0", "1"). Immutable. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + socketIndex: + description: |- + SocketIndex is the global ordinal (socketPosition × nodesPerSocket + nodeIndex). + Used internally by the operator to select the correct backend node from the + RPC-port-sorted list in pollUUIDFromBackend. Immutable. + format: int32 + type: integer + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + storageNodeSetRef: + description: StorageNodeSetRef is the name of the owning StorageNodeSet. Immutable. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + workerNode: + description: |- + WorkerNode is the Kubernetes node hostname this StorageNode runs on. + Users may not change it directly — it is re-pointed only by the operator + during a node migration (StorageNodeOps action=migrate). The + StorageNode validating webhook rejects user-driven changes to this field. + type: string + required: + - storageNodeSetRef + - workerNode + type: object + x-kubernetes-validations: + - message: field socketId is immutable once set + rule: '!has(oldSelf.socketId) || has(self.socketId)' + - message: field nodeIndex is immutable once set + rule: '!has(oldSelf.nodeIndex) || has(self.nodeIndex)' + - message: field socketIndex is immutable once set + rule: '!has(oldSelf.socketIndex) || has(self.socketIndex)' + status: + description: StorageNodeStatus holds the observed state of a StorageNode. + properties: + activeOpsRef: + description: |- + ActiveOpsRef is the name of the currently active StorageNodeOps CR targeting + this node. Empty when no operation is in progress. Used for mutual exclusion. + type: string + failureDomain: + description: |- + FailureDomain is the effective failure-domain group index for this node + as reported by the backend (≥ 0). Nil when the backend has not assigned one. + format: int32 + type: integer + health: + description: Health is the backend-reported node health flag. + type: boolean + hostname: + description: Hostname is the node hostname as reported by the backend. + type: string + latencyMetrics: + description: |- + LatencyMetrics holds the fio-measured baseline NVMe-oF latency for this node, + used by the volume rebalancer to make data-placement decisions. + properties: + baselineMeasuredAt: + description: BaselineMeasuredAt is when the baseline was established. + format: date-time type: string - type: array - pcieModel: - description: PcieModel overrides the PCI model filter for this - node. - type: string - reservedSystemCPU: - description: ReservedSystemCPU overrides the CPUs reserved for - system workloads on this node. - type: string - skipKubeletConfiguration: - description: |- - SkipKubeletConfiguration overrides whether kubelet configuration changes are - skipped for this node. - type: boolean - spdkImage: - description: SpdkImage overrides the SPDK image for this node - (e.g. for phased rollouts). - type: string - spdkProxyImage: - description: SpdkProxyImage overrides the SPDK proxy image for - this node. - type: string - spdkSystemMemory: - description: |- - SpdkSystemMemory overrides the SPDK huge-page memory allocation for this node - (e.g. "4G", "512M"). - pattern: ^[0-9]+(G|GI|GB|GiB|M|MI|MB|MiB|g|gi|gb|gib|m|mi|mb|mib)?$ - type: string - ubuntuHost: - description: UbuntuHost overrides the Ubuntu host OS flag for - this node. - type: boolean - type: object - socketId: - description: SocketID is the NUMA socket identifier from spec.socketsToUse - (e.g. "0", "1"). Immutable. - type: string - x-kubernetes-validations: - - message: field is immutable - rule: self == oldSelf - socketIndex: - description: |- - SocketIndex is the global ordinal (socketPosition × nodesPerSocket + nodeIndex). - Used internally by the operator to select the correct backend node from the - RPC-port-sorted list in pollUUIDFromBackend. Immutable. - format: int32 - type: integer - x-kubernetes-validations: - - message: field is immutable - rule: self == oldSelf - storageNodeSetRef: - description: StorageNodeSetRef is the name of the owning StorageNodeSet. - Immutable. - type: string - x-kubernetes-validations: - - message: field is immutable - rule: self == oldSelf - workerNode: - description: |- - WorkerNode is the Kubernetes node hostname this StorageNode runs on. - Users may not change it directly — it is re-pointed only by the operator - during a node migration (StorageNodeOps action=migrate). The - StorageNode validating webhook rejects user-driven changes to this field. - type: string - required: - - storageNodeSetRef - - workerNode - type: object - x-kubernetes-validations: - - message: field socketId is immutable once set - rule: '!has(oldSelf.socketId) || has(self.socketId)' - - message: field nodeIndex is immutable once set - rule: '!has(oldSelf.nodeIndex) || has(self.nodeIndex)' - - message: field socketIndex is immutable once set - rule: '!has(oldSelf.socketIndex) || has(self.socketIndex)' - status: - description: StorageNodeStatus holds the observed state of a StorageNode. - properties: - activeOpsRef: - description: |- - ActiveOpsRef is the name of the currently active StorageNodeOps CR targeting - this node. Empty when no operation is in progress. Used for mutual exclusion. - type: string - failureDomain: - description: |- - FailureDomain is the effective failure-domain group index for this node - as reported by the backend (≥ 0). Nil when the backend has not assigned one. - format: int32 - type: integer - health: - description: Health is the backend-reported node health flag. - type: boolean - hostname: - description: Hostname is the node hostname as reported by the backend. - type: string - latencyMetrics: - description: |- - LatencyMetrics holds the fio-measured baseline NVMe-oF latency for this node, - used by the volume rebalancer to make data-placement decisions. - properties: - baselineMeasuredAt: - description: BaselineMeasuredAt is when the baseline was established. - format: date-time - type: string - baselineP50NS: - description: BaselineP50NS is the p50 write latency (nanoseconds) - from the initial empty-cluster benchmark. - format: int64 - type: integer - baselineP99NS: - description: BaselineP99NS is the p99 write latency (nanoseconds) - from the initial empty-cluster benchmark. - format: int64 - type: integer - nodeUUID: - description: NodeUUID is the backend storage node UUID. - type: string - required: - - nodeUUID - type: object - ports: - description: Ports groups network connectivity fields (addresses and - ports). - properties: - lvol: - description: Lvol is the logical-volume subsystem port. - format: int32 - type: integer - management: - description: Management is the management IP address of the node. - type: string - nvmeof: - description: NvmeOf is the NVMe-oF fabric port. - format: int32 - type: integer - rpc: - description: Rpc is the RPC/management API port. - format: int32 - type: integer - type: object - postedAt: - description: |- - PostedAt is the timestamp when the node-add POST was sent. - Used as a provisioning guard against duplicate POSTs. - format: date-time - type: string - resources: - description: Resources groups compute and storage resource metrics. - properties: - capacity: - description: |- - Capacity is how much of the node's storage is in use, summed over its - devices. It is a measurement rather than a declaration, so it is absent - until something has measured it, and it lags reality by the interval at - which the control plane's metrics are scraped. - properties: - sampledAt: - description: |- - SampledAt is when the control plane took the reading. It is not when the - object was written, and it may be considerably older if metrics - collection has stopped. - format: date-time + baselineP50NS: + description: BaselineP50NS is the p50 write latency (nanoseconds) from the initial empty-cluster benchmark. + format: int64 + type: integer + baselineP99NS: + description: BaselineP99NS is the p99 write latency (nanoseconds) from the initial empty-cluster benchmark. + format: int64 + type: integer + nodeUUID: + description: NodeUUID is the backend storage node UUID. + type: string + required: + - nodeUUID + type: object + ports: + description: Ports groups network connectivity fields (addresses and ports). + properties: + lvol: + description: Lvol is the logical-volume subsystem port. + format: int32 + type: integer + management: + description: Management is the management IP address of the node. + type: string + nvmeof: + description: NvmeOf is the NVMe-oF fabric port. + format: int32 + type: integer + rpc: + description: Rpc is the RPC/management API port. + format: int32 + type: integer + type: object + postedAt: + description: |- + PostedAt is the timestamp when the node-add POST was sent. + Used as a provisioning guard against duplicate POSTs. + format: date-time + type: string + resources: + description: Resources groups compute and storage resource metrics. + properties: + capacity: + description: |- + Capacity is how much of the node's storage is in use, summed over its + devices. It is a measurement rather than a declaration, so it is absent + until something has measured it, and it lags reality by the interval at + which the control plane's metrics are scraped. + properties: + sampledAt: + description: |- + SampledAt is when the control plane took the reading. It is not when the + object was written, and it may be considerably older if metrics + collection has stopped. + format: date-time + type: string + totalBytes: + description: TotalBytes is the storage the node's devices provide. + format: int64 + minimum: 0 + type: integer + usedBytes: + description: UsedBytes is what they currently hold. + format: int64 + minimum: 0 + type: integer + type: object + cpu: + description: CPU is the number of SPDK CPU cores allocated to this node. + format: int32 + type: integer + devices: + description: Devices is the device summary (online/total) reported by the backend. + type: string + memory: + description: Memory is the SPDK memory allocation reported by the backend. + type: string + volumes: + description: Volumes is the current number of logical volumes on this node. + format: int32 + type: integer + type: object + status: + description: Status is the backend-reported node status (e.g. online, suspended, offline). + type: string + uptime: + description: Uptime is the node uptime as reported by the backend. + type: string + uuid: + description: UUID is the backend storage node UUID. Set once after node-add completes. + type: string + type: object + type: object + served: true + storage: false + subresources: + status: {} + - additionalPrinterColumns: + - jsonPath: .spec.clusterRef + name: Cluster + type: string + - jsonPath: .spec.workerNode + name: Worker + type: string + - jsonPath: .spec.socketId + name: Socket + type: string + - jsonPath: .spec.slot + name: Slot + type: integer + - jsonPath: .status.phase + name: Phase + type: string + - jsonPath: .status.step.state + name: Step + type: string + - jsonPath: .status.status + name: Status + type: string + - jsonPath: .status.health + name: Health + type: boolean + - jsonPath: .status.uuid + name: UUID + priority: 1 + type: string + - jsonPath: .status.failureDomain + name: FD + priority: 1 + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha2 + schema: + openAPIV3Schema: + description: |- + StorageNode is one backend storage node: one SPDK process bound to one NUMA + socket of one Kubernetes worker. One object exists per (workerNode, slot) pair, + owned by the StorageCluster it belongs to. + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: |- + StorageNodeSpec is the desired state of one backend storage node, meaning one + SPDK process bound to one NUMA socket of one Kubernetes worker. + properties: + clusterRef: + description: |- + ClusterRef names the StorageCluster this node belongs to. The cluster also + owns this object by controller reference, so deleting the cluster deletes + its nodes. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + config: + description: |- + Config is this node's complete configuration, copied from the + ClusterDeploymentConfig entry that produced it. It is a copy rather than a + projection, because that document is ephemeral: nothing reads it once the + node exists, deleting it changes nothing, and editing it reaches only nodes + created afterward. + properties: + deviceNames: + description: |- + DeviceNames names the devices to use. An entry is a PCI address + ("0000:5e:00.0") or a device path ("/dev/sdb,") which are the two classes + simplyblock accepts as backend storage, and a bare name ("nvme0n1") is read + as a path under /dev. One list carries both spellings, and every entry is of + the class its cluster declares in StorageCluster.spec.deviceClass: a list + mixing the two, or naming the class the cluster is not, is rejected by the + StorageNode validating webhook. Set explicitly, it overrides every filter + below. Immutable: it selects which physical devices the node owns. + items: + pattern: ^([0-9a-fA-F]{4}:[0-9a-fA-F]{2}:[0-9a-fA-F]{2}\.[0-9a-fA-F]|/dev/[a-zA-Z0-9._/-]+|[a-zA-Z0-9._-]+)$ type: string - totalBytes: - description: TotalBytes is the storage the node's devices - provide. - format: int64 - minimum: 0 - type: integer - usedBytes: - description: UsedBytes is what they currently hold. - format: int64 - minimum: 0 - type: integer - type: object - cpu: - description: CPU is the number of SPDK CPU cores allocated to - this node. - format: int32 - type: integer - devices: - description: Devices is the device summary (online/total) reported - by the backend. - type: string - memory: - description: Memory is the SPDK memory allocation reported by - the backend. - type: string - volumes: - description: Volumes is the current number of logical volumes - on this node. - format: int32 - type: integer - type: object - status: - description: Status is the backend-reported node status (e.g. online, - suspended, offline). - type: string - uptime: - description: Uptime is the node uptime as reported by the backend. - type: string - uuid: - description: UUID is the backend storage node UUID. Set once after - node-add completes. - type: string - type: object - type: object - served: true - storage: true - subresources: - status: {} + type: array + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + driveSizeRange: + description: DriveSizeRange filters devices by size. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + expand: + description: |- + Expand marks this node as an addition to an already-active cluster, which + the control plane reads as a request to rebalance onto it rather than to + treat it as part of an initial layout. Immutable once set: it describes how + the node joined rather than what it is. + type: boolean + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + failureDomain: + description: |- + FailureDomain is the label of the fault group this node belongs to + such as rack-b, naming the physical grouping it shares with its peers rather + than indexing it. Required when the cluster has enableFailureDomains set, + and provisioning is held with a FailureDomainMissing event until it is + present. Immutable once set, which is what makes it fillable later and then + frozen: chunk placement was computed from it. + + The value takes the shape of a Kubernetes label value, because that is what + it is seeded from where a cluster carries topology labels at all. + maxLength: 63 + pattern: ^[a-zA-Z0-9]([-_.a-zA-Z0-9]*[a-zA-Z0-9])?$ + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + journalManager: + description: |- + JournalManager tunes the journal manager count and per-device capacity + share for this node. Immutable: both are on-disk layout, fixed when the + devices were partitioned. + properties: + count: + description: Count is the number of journal managers to configure. + format: int32 + minimum: 1 + type: integer + percentPerDevice: + description: PercentPerDevice is the share of each device given to the journal. + format: int32 + maximum: 100 + minimum: 1 + type: integer + type: object + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + pcieAllowList: + description: |- + PcieAllowList selects devices by PCI address. It is the one device field a + migration writes, merging spec.migrate.newSsdPcie into it so devices added + on the target host survive a later rebuild, so it is guarded by the + StorageNode validating webhook rather than by a marker. This and the two + PCI filters below belong to an NVMe cluster: the webhook rejects them on a + cluster whose deviceClass is LogicalBlock, because a logical block device + has no PCI address to match. + items: + type: string + type: array + pcieDenyList: + description: PcieDenyList excludes devices by PCI address. + items: + type: string + type: array + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + pcieModel: + description: PcieModel filters devices by PCI model string. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + sizing: + description: |- + Sizing is what this node's huge pages and core layout were sized from. + Writable by the operator alone. + properties: + minHugePagesSize: + description: |- + MinHugePagesSize is the smallest huge-page allocation this node makes, as a + size string such as 100G or 1T, where a bare number is gigabytes. It is a + floor rather than a limit: the effective allocation is the larger of this + value and the minimum the node's device and subsystem count requires. + type: string + vcpuCount: + description: |- + VCPUCount is the number of vCPUs allocated to SPDK on this node, as an + explicit core count rather than a percentage. + format: int32 + minimum: 4 + type: integer + required: + - vcpuCount + type: object + spdkImage: + description: |- + SpdkImage overrides the SPDK image the control plane starts for this node, + which is what makes a phased image rollout expressible per node. + type: string + spdkProxyImage: + description: SpdkProxyImage overrides the SPDK proxy image for this node. + type: string + spdkSystemMemory: + description: |- + SpdkSystemMemory is the memory the control plane starts this node's SPDK + with, as a size string such as 4G or 512M. Mutable: a node whose device + count grew legitimately needs to raise it. + pattern: ^[0-9]+(G|GI|GB|GiB|M|MI|MB|MiB|g|gi|gb|gib|m|mi|mb|mib)?$ + type: string + required: + - sizing + type: object + x-kubernetes-validations: + - message: field journalManager is immutable once set + rule: '!has(oldSelf.journalManager) || has(self.journalManager)' + - message: field deviceNames is immutable once set + rule: '!has(oldSelf.deviceNames) || has(self.deviceNames)' + - message: field pcieDenyList is immutable once set + rule: '!has(oldSelf.pcieDenyList) || has(self.pcieDenyList)' + - message: field pcieModel is immutable once set + rule: '!has(oldSelf.pcieModel) || has(self.pcieModel)' + - message: field driveSizeRange is immutable once set + rule: '!has(oldSelf.driveSizeRange) || has(self.driveSizeRange)' + - message: field failureDomain is immutable once set + rule: '!has(oldSelf.failureDomain) || has(self.failureDomain)' + - message: field expand is immutable once set + rule: '!has(oldSelf.expand) || has(self.expand)' + nodeIndex: + description: |- + NodeIndex is the position among the nodes sharing this socket, in + 0..nodesPerSocket-1. See SocketID. + format: int32 + minimum: 0 + type: integer + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + nodeSet: + description: |- + NodeSet is the name of the group in ClusterDeploymentConfig.nodeSets[] this + node was declared under. It is a label rather than a reference: nothing is + fetched by it, and it exists so that a node can be traced back to the + document that produced it. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + slot: + description: |- + Slot is which storage-node slot on this worker the object occupies, counted + from zero. A worker runs one node per socket per nodesPerSocket, and the + slot is the position among them. It is the identity the operator keys on: + the topology label the CSI driver reads is + storage.simplyblock.io/storage-node-uuid.., and the slot + outlives the node filling it, because only the UUID behind it changes when a + node is replaced or relocated. + format: int32 + minimum: 0 + type: integer + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + socketId: + description: |- + SocketID is the NUMA socket this node is bound to, as declared in the node + set's socket list, so 0 or 1. With NodeIndex it decomposes Slot into the + pair a person reads; nothing but a print column consumes either. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + workerNode: + description: |- + WorkerNode is the Kubernetes worker hostname this node runs on. It is not + marked immutable, because a migration re-points it, but the StorageNode + validating webhook rejects any change made by an identity outside the + operator's namespace. + type: string + required: + - clusterRef + - config + - workerNode + type: object + x-kubernetes-validations: + - message: field nodeSet is immutable once set + rule: '!has(oldSelf.nodeSet) || has(self.nodeSet)' + - message: field socketId is immutable once set + rule: '!has(oldSelf.socketId) || has(self.socketId)' + - message: field nodeIndex is immutable once set + rule: '!has(oldSelf.nodeIndex) || has(self.nodeIndex)' + - message: field slot is immutable once set + rule: '!has(oldSelf.slot) || has(self.slot)' + status: + description: StorageNodeStatus is the observed state of one storage node. + properties: + activeOpsRef: + description: |- + ActiveOpsRef names the StorageNodeOps currently allowed to touch this node. + Empty when none is running. + type: string + failureDomain: + description: |- + FailureDomain is the failure-domain label the control plane actually + assigned, which is not necessarily the one spec.config.failureDomain + requested. + type: string + health: + description: Health is the health flag the control plane reports. + type: boolean + hostname: + description: Hostname is the node hostname as the control plane reports it. + type: string + latencyMetrics: + description: |- + LatencyMetrics holds the fio-measured NVMe-oF baseline the volume + rebalancer reads. + properties: + baselineMeasuredAt: + description: BaselineMeasuredAt is when the baseline was established. + format: date-time + type: string + baselineP50NS: + description: |- + BaselineP50NS is the p50 write latency, in nanoseconds, of the initial + empty-cluster benchmark. + format: int64 + minimum: 0 + type: integer + baselineP99NS: + description: |- + BaselineP99NS is the p99 write latency, in nanoseconds, of the same + benchmark. + format: int64 + minimum: 0 + type: integer + nodeUUID: + description: |- + NodeUUID is the backend storage node the reading was taken against. It is + carried beside the reading rather than inferred from status.uuid, because a + baseline measured against one backend node stops describing the slot once a + replacement fills it. + type: string + required: + - nodeUUID + type: object + message: + description: |- + Message is the reason the phase is what it is: one sentence, replaced as the + node moves, and never a log. + type: string + observedGeneration: + description: |- + ObservedGeneration is the generation the rest of this status was computed + from, so a stale status can be told from a current one. + format: int64 + type: integer + phase: + description: |- + Phase is the operator's own view of this node, and the field its + provisioning branches on. + enum: + - Pending + - Provisioning + - Online + - Removing + - Offline + - Degraded + - Failed + type: string + ports: + description: Ports groups the reported addresses and ports. + properties: + lvol: + description: Lvol is the logical-volume subsystem port. + format: int32 + type: integer + management: + description: Management is the management IP address of the node. + type: string + nvmeof: + description: The NVMe-oF fabric port. + format: int32 + type: integer + rpc: + description: Rpc is the RPC and management API port. + format: int32 + type: integer + type: object + resources: + description: Resources groups the reported compute and storage figures. + properties: + capacity: + description: |- + Capacity is how much of the node's storage is in use, summed over its + devices. It is a measurement rather than a declaration, so it is absent + until something has measured it, and it lags reality by the interval at + which the control plane's metrics are scraped. + properties: + sampledAt: + description: |- + SampledAt is when the control plane took the reading. It is not when the + object was written, and it may be considerably older if metrics collection + has stopped. + format: date-time + type: string + totalBytes: + description: TotalBytes is the storage the node's devices provide. + format: int64 + minimum: 0 + type: integer + usedBytes: + description: UsedBytes is what they currently hold. + format: int64 + minimum: 0 + type: integer + type: object + cpu: + description: CPU is the number of SPDK cores allocated to this node. + format: int32 + type: integer + devices: + description: |- + Devices summarizes the node's NVMe devices. Absent until the control plane + has reported, which is what tells a node that has not reported from one that + genuinely has no devices. + properties: + online: + description: |- + Online is how many of the node's devices the control plane reports as + usable. + format: int32 + minimum: 0 + type: integer + total: + description: Total is how many devices the node has. + format: int32 + minimum: 0 + type: integer + required: + - online + - total + type: object + memory: + description: Memory is the SPDK memory allocation the control plane reports. + type: string + volumes: + description: Volumes is the current number of logical volumes on this node. + format: int32 + type: integer + type: object + status: + description: |- + Status is the lifecycle the control plane reports: online, suspended, + offline, in_creation, in_restart, in_shutdown, unreachable, or timeout. The + values are the control plane's, which is why they are neither PascalCase nor + constrained by an Enum here. + type: string + step: + description: |- + Step is the position of the provisioning machine, as the shared + statemachine.KubeSnapshot. The rule is what an Enum marker would do if a + marker could reach a field of a shared type. + properties: + deadline: + description: |- + Deadline is when that state expires, absent when it has none. It is an + absolute instant, so a state whose deadline passed while the controller + was down restores as already expired. + format: date-time + type: string + state: + description: |- + State is the state the machine was in. Empty means the resource has not + been reconciled yet, and restores to the graph's initial state. + type: string + type: object + x-kubernetes-validations: + - message: unknown step + rule: '!has(self.state) || self.state in [''CheckingHost'',''CheckingConfig'',''AwaitingSlot'',''Posting'',''Resolving'',''Adopting'']' + uptime: + description: Uptime is the node uptime as the control plane reports it. + type: string + uuid: + description: |- + UUID is the backend node UUID. Empty means the node has neither been + provisioned nor adopted, and non-empty means steady state. + type: string + type: object + type: object + served: true + storage: true + subresources: + status: {} + conversion: + strategy: Webhook + webhook: + conversionReviewVersions: + - v1 + clientConfig: + service: + namespace: simplyblock-operator-system + name: simplyblock-operator-conversion-webhook-service + path: /convert diff --git a/operator/internal/upgrade/steps/sizing.go b/operator/internal/upgrade/steps/sizing.go new file mode 100644 index 000000000..bbf31a5ef --- /dev/null +++ b/operator/internal/upgrade/steps/sizing.go @@ -0,0 +1,246 @@ +// Stamping each StorageNode with the sizing its cluster was built with. +// +// spec.config.sizing is required on the v1alpha2 StorageNode and has no v1alpha1 +// spelling at all, because both values lived on the StorageCluster and the node +// held no copy (design-storagenode.md §3.1). A node stored as v1alpha1 therefore +// converts up with no sizing, and the first write of it afterward is refused by +// the field's own Required marker. +// +// The conversion cannot fill it in. It has no client to read the cluster with, and +// a conversion may not depend on one: it runs inside the API server's request path +// against whatever object was handed to it. So the value is carried the way every +// other hub field this version cannot express is carried — in the conversion's +// stash annotation, which ConvertTo reads back into the field — and writing that +// annotation is what this step does. +// +// It runs after reparent-storage-nodes, because the cluster a node's sizing comes +// from is the one that step establishes as its controller owner. A node that has +// not been reparented has no cluster to read, which is the same ordering the +// conversion's own fallback depends on. + +package steps + +import ( + "context" + "encoding/json" + "fmt" + + "sigs.k8s.io/controller-runtime/pkg/client" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/upgrade" +) + +// IDStampNodeSizing is the step's identity. +const IDStampNodeSizing upgrade.ID = "stamp-storage-node-sizing" + +// annoNodeSizing is the annotation api/v1alpha1/storagenode_conversion.go reads +// the sizing back out of. The two have to spell it the same way or the stamp is +// written somewhere nothing looks. +const annoNodeSizing = "storage.simplyblock.io/conversion-spec.config.sizing" + +// Sizing returns the step. +func Sizing() []upgrade.Step { return []upgrade.Step{stampNodeSizing{}} } + +// stampNodeSizing writes each node's sizing from its cluster. +type stampNodeSizing struct{} + +func (stampNodeSizing) ID() upgrade.ID { return IDStampNodeSizing } +func (stampNodeSizing) Stage() upgrade.Stage { return upgrade.StageMigrate } +func (stampNodeSizing) Phase() upgrade.Phase { return upgrade.PhaseTransforming } + +// Requires names the reparent, because the cluster this reads is the owner that +// step establishes. +func (stampNodeSizing) Requires() []upgrade.ID { return []upgrade.ID{IDReparentNodes} } + +func (stampNodeSizing) Description() string { + return "stamps each StorageNode with the vCPU count and huge-page floor its StorageCluster was built with" +} + +// Describe reports the stamp a node is missing, and nothing for one that carries +// it already. A rerun's plan therefore holds the outstanding work rather than what +// was originally intended. +func (s stampNodeSizing) Describe( + ctx context.Context, scope *upgrade.Scope, subject upgrade.Subject, +) (*upgrade.Action, error) { + if !s.applies(subject) || s.stamped(subject) { + return nil, nil + } + sizing, err := s.sizingFor(ctx, scope, subject) + if err != nil { + // A cluster that cannot be read is Validate's to refuse. Describing + // nothing here would drop the node from the plan silently, which is the + // one outcome worse than a refusal. + return &upgrade.Action{ + Rule: IDStampNodeSizing, + Verb: upgrade.VerbAnnotate, + Object: subject.Ref, + Detail: "sizing unresolved: " + err.Error(), + }, nil + } + return &upgrade.Action{ + Rule: IDStampNodeSizing, + Verb: upgrade.VerbAnnotate, + Object: subject.Ref, + Detail: describeSizing(sizing), + }, nil +} + +// Done claims a node that already carries the stamp, read off the subject rather +// than off a record, so a run killed anywhere resumes correctly. +func (s stampNodeSizing) Done( + _ context.Context, _ *upgrade.Scope, subject upgrade.Subject, +) (bool, error) { + return s.applies(subject) && s.stamped(subject), nil +} + +// Validate refuses a node whose cluster states no core count. +// +// The field is Required on the hub with a minimum of four, so a stamp built from +// a cluster that states nothing would write a sizing the next admission refuses — +// which is the same failure this step exists to prevent, moved one step later and +// made harder to attribute. +func (s stampNodeSizing) Validate( + ctx context.Context, scope *upgrade.Scope, subject upgrade.Subject, +) error { + if !s.applies(subject) || s.stamped(subject) { + return nil + } + _, err := s.sizingFor(ctx, scope, subject) + return err +} + +// Apply writes the annotation, leaving everything else on the object alone. +func (s stampNodeSizing) Apply( + ctx context.Context, scope *upgrade.Scope, subject upgrade.Subject, +) error { + sizing, err := s.sizingFor(ctx, scope, subject) + if err != nil { + return err + } + encoded, err := json.Marshal(sizing) + if err != nil { + return fmt.Errorf("encoding the sizing: %w", err) + } + + object, ok := subject.Object.DeepCopyObject().(client.Object) + if !ok { + return fmt.Errorf("%T is not a Kubernetes object", subject.Object) + } + annotations := object.GetAnnotations() + if annotations == nil { + annotations = map[string]string{} + } + annotations[annoNodeSizing] = string(encoded) + object.SetAnnotations(annotations) + + if err := scope.Client.Update(ctx, object); err != nil { + return fmt.Errorf("writing the sizing stamp: %w", err) + } + scope.Adopt(object) + return nil +} + +// Verify re-reads the node and checks the stamp decodes to a usable sizing. +// +// Decoding it rather than asserting the key is present is the point: an +// annotation that is there and does not parse is read by the conversion as no +// sizing at all, which is the state this step exists to leave behind. +func (s stampNodeSizing) Verify( + ctx context.Context, scope *upgrade.Scope, subject upgrade.Subject, +) error { + fresh, ok := subject.Object.DeepCopyObject().(client.Object) + if !ok { + return fmt.Errorf("%T is not a Kubernetes object", subject.Object) + } + if err := scope.Client.Get(ctx, subject.Ref.Key(), fresh); err != nil { + return fmt.Errorf("re-reading it: %w", err) + } + + raw, carried := fresh.GetAnnotations()[annoNodeSizing] + if !carried { + return fmt.Errorf("%s was not written", annoNodeSizing) + } + var sizing simplyblockv1alpha2.StorageNodeSizing + if err := json.Unmarshal([]byte(raw), &sizing); err != nil { + return fmt.Errorf("%s does not decode as a sizing: %w", annoNodeSizing, err) + } + if sizing.VCPUCount == nil { + return fmt.Errorf("%s carries no vcpuCount", annoNodeSizing) + } + return nil +} + +// applies reports whether this subject is a StorageNode. +func (stampNodeSizing) applies(subject upgrade.Subject) bool { + return !subject.IsUpgrade() && subject.Object != nil && + subject.Ref.GVK.Kind == nodeKind +} + +// stamped reports whether the node already carries a decodable sizing. +// +// A malformed annotation counts as absent, so a stamp somebody hand-edited into +// nonsense is rewritten rather than trusted. +func (stampNodeSizing) stamped(subject upgrade.Subject) bool { + raw, carried := subject.Object.GetAnnotations()[annoNodeSizing] + if !carried { + return false + } + var sizing simplyblockv1alpha2.StorageNodeSizing + if err := json.Unmarshal([]byte(raw), &sizing); err != nil { + return false + } + return sizing.VCPUCount != nil +} + +// sizingFor reads the cluster the node belongs to and returns the sizing it was +// built with. +func (stampNodeSizing) sizingFor( + ctx context.Context, scope *upgrade.Scope, subject upgrade.Subject, +) (simplyblockv1alpha2.StorageNodeSizing, error) { + owner := controllingCluster(subject.Object) + if owner == "" { + return simplyblockv1alpha2.StorageNodeSizing{}, fmt.Errorf( + "no StorageCluster owns this node, so there is no sizing to stamp from; "+ + "%s has not run", IDReparentNodes) + } + + var cluster simplyblockv1alpha2.StorageCluster + key := client.ObjectKey{Namespace: subject.Ref.Namespace, Name: owner} + if err := scope.Client.Get(ctx, key, &cluster); err != nil { + return simplyblockv1alpha2.StorageNodeSizing{}, + fmt.Errorf("reading StorageCluster %s: %w", owner, err) + } + if cluster.Spec.VCPUCount == nil { + return simplyblockv1alpha2.StorageNodeSizing{}, fmt.Errorf( + "StorageCluster %s states no spec.vcpuCount, so a node stamped from it "+ + "would be refused by the field's own minimum", owner) + } + + count := *cluster.Spec.VCPUCount + return simplyblockv1alpha2.StorageNodeSizing{ + VCPUCount: &count, + MinHugePagesSize: cluster.Spec.MinHugePagesSize, + }, nil +} + +// controllingCluster names the StorageCluster that owns this object, or the empty +// string when none does. +func controllingCluster(object client.Object) string { + for _, ref := range object.GetOwnerReferences() { + if ref.Controller != nil && *ref.Controller && ref.Kind == clusterKind { + return ref.Name + } + } + return "" +} + +// describeSizing renders the stamp for a plan, so a reader sees the numbers the +// step would write rather than that it would write something. +func describeSizing(sizing simplyblockv1alpha2.StorageNodeSizing) string { + detail := fmt.Sprintf("vcpuCount=%d", *sizing.VCPUCount) + if sizing.MinHugePagesSize != "" { + detail += ", minHugePagesSize=" + sizing.MinHugePagesSize + } + return detail +} diff --git a/operator/internal/upgrade/steps/sizing_test.go b/operator/internal/upgrade/steps/sizing_test.go new file mode 100644 index 000000000..2eb262371 --- /dev/null +++ b/operator/internal/upgrade/steps/sizing_test.go @@ -0,0 +1,237 @@ +// Tests for the sizing stamp. +// +// The property that matters is the one the conversion depends on: after the step, +// a node carries an annotation that decodes to a sizing with a core count in it. +// Every case here is either that, a reason the step must refuse rather than write +// something the next admission rejects, or the idempotence a rerun needs. + +package steps + +import ( + "context" + "encoding/json" + "strings" + "testing" + + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + + "github.com/simplyblock/atlas/ptr" + + simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/upgrade" +) + +// sizedCluster is a cluster that states what its nodes were built with. +func sizedCluster(vcpus *int32, hugePages string) *simplyblockv1alpha2.StorageCluster { + return &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{Name: theCluster, Namespace: "simplyblock"}, + Spec: simplyblockv1alpha2.StorageClusterSpec{ + VCPUCount: vcpus, + MinHugePagesSize: hugePages, + }, + } +} + +// reparentedNode is a node the ownership step has already moved onto its cluster, +// which is the state this step requires. +func reparentedNode(annotations map[string]string) *simplyblockv1alpha1.StorageNode { + return &simplyblockv1alpha1.StorageNode{ + ObjectMeta: metav1.ObjectMeta{ + Name: "node-a", + Namespace: "simplyblock", + Annotations: annotations, + OwnerReferences: ownedByCluster(), + }, + Spec: simplyblockv1alpha1.StorageNodeSpec{ + StorageNodeSetRef: theSet, + WorkerNode: "worker-1", + }, + } +} + +// stampOf decodes the annotation the step writes. +func stampOf(t *testing.T, raw string) simplyblockv1alpha2.StorageNodeSizing { + t.Helper() + var sizing simplyblockv1alpha2.StorageNodeSizing + if err := json.Unmarshal([]byte(raw), &sizing); err != nil { + t.Fatalf("the stamp does not decode: %v", err) + } + return sizing +} + +// The whole point: a node with no sizing gets its cluster's, in the annotation the +// conversion reads back into spec.config.sizing. +func TestTheSizingIsStampedFromTheCluster(t *testing.T) { + node := reparentedNode(nil) + scope := migration(t, sizedCluster(ptr.To(int32(8)), "100G"), node) + subject := upgrade.Subject{Ref: scope.Ref(node), Object: node} + step := stampNodeSizing{} + + if err := step.Validate(context.Background(), scope, subject); err != nil { + t.Fatalf("Validate: %v", err) + } + if err := step.Apply(context.Background(), scope, subject); err != nil { + t.Fatalf("Apply: %v", err) + } + if err := step.Verify(context.Background(), scope, subject); err != nil { + t.Fatalf("Verify: %v", err) + } + + var fresh simplyblockv1alpha1.StorageNode + if err := scope.Client.Get(context.Background(), subject.Ref.Key(), &fresh); err != nil { + t.Fatalf("re-reading the node: %v", err) + } + sizing := stampOf(t, fresh.Annotations[annoNodeSizing]) + if sizing.VCPUCount == nil || *sizing.VCPUCount != 8 { + t.Errorf("vcpuCount = %v, want the cluster's 8", sizing.VCPUCount) + } + if sizing.MinHugePagesSize != "100G" { + t.Errorf("minHugePagesSize = %q, want the cluster's 100G", sizing.MinHugePagesSize) + } +} + +// Describe reports the outstanding work and nothing for a node already stamped, so +// a rerun's plan shrinks as the migration completes. +func TestTheStampIsDescribedOnceAndThenDone(t *testing.T) { + node := reparentedNode(nil) + scope := migration(t, sizedCluster(ptr.To(int32(8)), "100G"), node) + subject := upgrade.Subject{Ref: scope.Ref(node), Object: node} + step := stampNodeSizing{} + + action, err := step.Describe(context.Background(), scope, subject) + if err != nil { + t.Fatalf("Describe: %v", err) + } + if action == nil { + t.Fatal("a node with no sizing described nothing") + } + if !strings.Contains(action.Detail, "vcpuCount=8") { + t.Errorf("the plan does not show the numbers it would write: %s", action.Detail) + } + if done, _ := step.Done(context.Background(), scope, subject); done { + t.Error("a node with no sizing reported done") + } + + if err := step.Apply(context.Background(), scope, subject); err != nil { + t.Fatalf("Apply: %v", err) + } + + var fresh simplyblockv1alpha1.StorageNode + if err := scope.Client.Get(context.Background(), subject.Ref.Key(), &fresh); err != nil { + t.Fatalf("re-reading the node: %v", err) + } + stamped := upgrade.Subject{Ref: subject.Ref, Object: &fresh} + + action, err = step.Describe(context.Background(), scope, stamped) + if err != nil { + t.Fatalf("Describe after the stamp: %v", err) + } + if action != nil { + t.Errorf("a stamped node still describes work: %s", action.Detail) + } + if done, _ := step.Done(context.Background(), scope, stamped); !done { + t.Error("a stamped node did not report done") + } +} + +// A cluster that states no core count would produce a stamp the field's own +// minimum refuses, so the step says so rather than writing it. +func TestAClusterWithNoCoreCountIsRefused(t *testing.T) { + node := reparentedNode(nil) + scope := migration(t, sizedCluster(nil, "100G"), node) + subject := upgrade.Subject{Ref: scope.Ref(node), Object: node} + + err := stampNodeSizing{}.Validate(context.Background(), scope, subject) + if err == nil { + t.Fatal("a cluster with no vcpuCount was accepted") + } + if !strings.Contains(err.Error(), "vcpuCount") { + t.Errorf("the refusal does not name the missing field: %v", err) + } +} + +// The step reads the cluster off the controller owner reference, so a node the +// reparent has not reached has nothing to stamp from and says which step is owed. +func TestANodeThatWasNotReparentedIsRefused(t *testing.T) { + node := reparentedNode(nil) + node.OwnerReferences = ownedBySet() + scope := migration(t, sizedCluster(ptr.To(int32(8)), "100G"), node) + subject := upgrade.Subject{Ref: scope.Ref(node), Object: node} + + err := stampNodeSizing{}.Validate(context.Background(), scope, subject) + if err == nil { + t.Fatal("a node still owned by its set was accepted") + } + if !strings.Contains(err.Error(), string(IDReparentNodes)) { + t.Errorf("the refusal does not name the step that is owed: %v", err) + } +} + +// A stamp somebody hand-edited into nonsense is rewritten rather than trusted, +// because the conversion reads an undecodable one as no sizing at all. +func TestAMalformedStampIsRewritten(t *testing.T) { + node := reparentedNode(map[string]string{annoNodeSizing: "not json"}) + scope := migration(t, sizedCluster(ptr.To(int32(8)), "100G"), node) + subject := upgrade.Subject{Ref: scope.Ref(node), Object: node} + step := stampNodeSizing{} + + if done, _ := step.Done(context.Background(), scope, subject); done { + t.Fatal("a node carrying an undecodable stamp reported done") + } + if err := step.Apply(context.Background(), scope, subject); err != nil { + t.Fatalf("Apply: %v", err) + } + if err := step.Verify(context.Background(), scope, subject); err != nil { + t.Errorf("Verify after rewriting the stamp: %v", err) + } +} + +// The huge-page floor is optional on the cluster, so a cluster that states none +// produces a stamp that states none either. Inventing a value here would write a +// floor nobody asked for onto every node of the fleet. +func TestAnUnstatedHugePageFloorIsNotInvented(t *testing.T) { + node := reparentedNode(nil) + scope := migration(t, sizedCluster(ptr.To(int32(6)), ""), node) + subject := upgrade.Subject{Ref: scope.Ref(node), Object: node} + step := stampNodeSizing{} + + if err := step.Validate(context.Background(), scope, subject); err != nil { + t.Fatalf("Validate: %v", err) + } + if err := step.Apply(context.Background(), scope, subject); err != nil { + t.Fatalf("Apply: %v", err) + } + if err := step.Verify(context.Background(), scope, subject); err != nil { + t.Fatalf("Verify: %v", err) + } + + var fresh simplyblockv1alpha1.StorageNode + if err := scope.Client.Get(context.Background(), subject.Ref.Key(), &fresh); err != nil { + t.Fatalf("re-reading the node: %v", err) + } + sizing := stampOf(t, fresh.Annotations[annoNodeSizing]) + if sizing.MinHugePagesSize != "" { + t.Errorf("minHugePagesSize = %q, want it unstated as the cluster left it", + sizing.MinHugePagesSize) + } + if sizing.VCPUCount == nil || *sizing.VCPUCount != 6 { + t.Errorf("vcpuCount = %v, want the cluster's 6", sizing.VCPUCount) + } +} + +// The step is about StorageNodes and describes nothing for anything else, which is +// what keeps it out of every other subject's plan. +func TestTheStampIgnoresEveryOtherKind(t *testing.T) { + cluster := sizedCluster(ptr.To(int32(8)), "100G") + scope := migration(t, cluster) + subject := upgrade.Subject{Ref: scope.Ref(cluster), Object: cluster} + + action, err := stampNodeSizing{}.Describe(context.Background(), scope, subject) + if err != nil { + t.Fatalf("Describe: %v", err) + } + if action != nil { + t.Errorf("a StorageCluster described sizing work: %s", action.Detail) + } +} diff --git a/operator/internal/utils/storage_node_api.go b/operator/internal/utils/storage_node_api.go new file mode 100644 index 000000000..cce0ce4ee --- /dev/null +++ b/operator/internal/utils/storage_node_api.go @@ -0,0 +1,71 @@ +// Probing a worker's storage-node API. +// +// It is the one read the node's own reconcile makes that has no streamed +// counterpart: a Kubernetes-side check against a pod rather than a question about +// a control-plane object, so nothing delivers it and a request is what answers it +// (design-storagenode.md §4.4). +// +// It lives in utils rather than beside either caller because two of them ask. The +// node's provisioning holds until the worker answers, and a maintenance window's +// AwaitingHost step waits for the same answer after a reboot. + +package utils + +import ( + "context" + "fmt" + "net/http" + "time" + + "github.com/simplyblock/simplyblock-operator/internal/tlsutil" +) + +// storageNodeAPIProbeTimeout bounds one probe. It is short because the question is +// whether the host is answering now, and a caller that holds until it does looks +// again on its own schedule rather than waiting inside one request. +const storageNodeAPIProbeTimeout = 3 * time.Second + +// StorageNodeAPIReachable reports nil when the worker's storage-node API answers +// its info endpoint, and an error describing why not otherwise. +// +// The scheme follows the deployment's TLS settings rather than being probed for: a +// deployment serving TLS refuses a plaintext request, and a refused request is not +// the same answer as an unreachable host. +func StorageNodeAPIReachable( + ctx context.Context, worker, namespace string, tlsEnabled, tlsMutualEnabled bool, +) error { + scheme := "http" + httpClient := &http.Client{Timeout: storageNodeAPIProbeTimeout} + if tlsEnabled { + scheme = "https" + certPath, keyPath := "", "" + if tlsMutualEnabled { + certPath = tlsutil.ServiceClientCertificatePath + keyPath = tlsutil.ServiceClientKeyPath + } + built, err := tlsutil.BuildStorageNodeSetAPIClient( + namespace, tlsutil.ServiceCABundlePath, certPath, keyPath) + if err != nil { + return fmt.Errorf("build the storage-node TLS client: %w", err) + } + httpClient = built + } + + url := fmt.Sprintf("%s://%s/snode/info", scheme, StorageNodeSetAPIAddress(worker, namespace)) + request, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil) + if err != nil { + return err + } + + response, err := httpClient.Do(request) + if err != nil { + return fmt.Errorf("the storage-node API on worker %s does not answer: %w", worker, err) + } + defer func() { _ = response.Body.Close() }() + + if response.StatusCode != http.StatusOK { + return fmt.Errorf("the storage-node API on worker %s answered %d", + worker, response.StatusCode) + } + return nil +} diff --git a/operator/internal/utils/storage_nodeset_ds.go b/operator/internal/utils/storage_node_workload.go similarity index 85% rename from operator/internal/utils/storage_nodeset_ds.go rename to operator/internal/utils/storage_node_workload.go index 67d57c26d..1bececede 100644 --- a/operator/internal/utils/storage_nodeset_ds.go +++ b/operator/internal/utils/storage_node_workload.go @@ -7,7 +7,7 @@ import ( "github.com/simplyblock/atlas/kube" "github.com/simplyblock/atlas/ptr" - simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" appsv1 "k8s.io/api/apps/v1" corev1 "k8s.io/api/core/v1" discoveryv1 "k8s.io/api/discovery/v1" @@ -45,26 +45,29 @@ var defaultContainerResources = corev1.ResourceRequirements{ }, } -func BuildStorageNodeSetDaemonSet(sn *simplyblockv1alpha1.StorageNodeSet, tlsEnabled bool, tlsMutualEnabled bool, tlsProvider, tlsSecretResourceVersion string) *appsv1.DaemonSet { +func BuildStorageNodeDaemonSet( + sn *simplyblockv1alpha2.StorageCluster, + tlsEnabled, tlsMutualEnabled bool, + tlsProvider, tlsSecretResourceVersion, image string, +) *appsv1.DaemonSet { + wl := storageNodesOf(sn) labels := map[string]string{ kube.LabelApp: kube.AppStorageNode, - kube.LabelSimplyblockCluster: sn.Spec.ClusterName, + kube.LabelSimplyblockCluster: sn.Name, kube.LabelStorageNodeSet: sn.Name, } - image := sn.Spec.ClusterImage - // Build the fleet-level (non-overridable) args that are always appended. // Per-node args (pci-*, device-*, size-range) and the cluster-scoped sizing // args (max-subsys-count, max-size, vcpu-count) are read at runtime from the per-node // ConfigMap via the init script. fleetArgs := "" - if len(sn.Spec.SocketsToUse) > 0 { - fleetArgs += " --sockets-to-use=" + JoinList(sn.Spec.SocketsToUse) + if len(wl.SocketsToUse) > 0 { + fleetArgs += " --sockets-to-use=" + JoinList(wl.SocketsToUse) } - if sn.Spec.NodesPerSocket != nil { - fleetArgs += " --nodes-per-socket=" + ptr.StringOrDefault(sn.Spec.NodesPerSocket, "") + if wl.NodesPerSocket != nil { + fleetArgs += " --nodes-per-socket=" + ptr.StringOrDefault(wl.NodesPerSocket, "") } // The init container sources the per-node env file (written by node-env-writer) @@ -92,29 +95,29 @@ else fi` nodeEnvWriterCmd := []string{"sh", "-c", nodeEnvWriterScript} - imagePullPolicy := sn.Spec.ImagePullPolicy + imagePullPolicy := wl.ImagePullPolicy if imagePullPolicy == "" { imagePullPolicy = corev1.PullAlways } mainEnv := []corev1.EnvVar{ - {Name: "UBUNTU_HOST", Value: ptr.StringOrDefault(sn.Spec.UbuntuHost, "false")}, - {Name: "OPENSHIFT_CLUSTER", Value: ptr.StringOrDefault(sn.Spec.OpenShiftCluster, "false")}, - {Name: "SKIP_KUBELET_CONFIGURATION", Value: ptr.StringOrDefault(sn.Spec.SkipKubeletConfiguration, "false")}, + {Name: "UBUNTU_HOST", Value: ptr.StringOrDefault(wl.UbuntuHost, "false")}, + {Name: "OPENSHIFT_CLUSTER", Value: ptr.StringOrDefault(wl.OpenShiftCluster, "false")}, + {Name: "SKIP_KUBELET_CONFIGURATION", Value: skipKubeletConfiguration(wl)}, {Name: "SIMPLY_BLOCK_DOCKER_IMAGE", Value: image}, {Name: "HOSTNAME", ValueFrom: &corev1.EnvVarSource{ FieldRef: &corev1.ObjectFieldSelector{FieldPath: "spec.nodeName"}, }}, - {Name: "CPU_TOPOLOGY_ENABLED", Value: ptr.StringOrDefault(sn.Spec.EnableCpuTopology, "false")}, + {Name: "CPU_TOPOLOGY_ENABLED", Value: ptr.StringOrDefault(wl.EnableCpuTopology, "false")}, } - if sn.Spec.MaxParallelNodeAdds != nil { - mainEnv = append(mainEnv, corev1.EnvVar{Name: "MAX_PARALLEL_NODE_ADDS", Value: fmt.Sprintf("%d", *sn.Spec.MaxParallelNodeAdds)}) + if wl.MaxParallelNodeAdds != nil { + mainEnv = append(mainEnv, corev1.EnvVar{Name: "MAX_PARALLEL_NODE_ADDS", Value: fmt.Sprintf("%d", *wl.MaxParallelNodeAdds)}) } - if sn.Spec.OpenShiftMachineConfigPool != "" { - mainEnv = append(mainEnv, corev1.EnvVar{Name: "OPENSHIFT_MCP", Value: sn.Spec.OpenShiftMachineConfigPool}) + if wl.OpenShiftMachineConfigPool != "" { + mainEnv = append(mainEnv, corev1.EnvVar{Name: "OPENSHIFT_MCP", Value: wl.OpenShiftMachineConfigPool}) } - if sn.Spec.ReservedSystemCPU != "" { - mainEnv = append(mainEnv, corev1.EnvVar{Name: "RESERVED_SYSTEM_CPUS", Value: sn.Spec.ReservedSystemCPU}) + if wl.ReservedSystemCPU != "" { + mainEnv = append(mainEnv, corev1.EnvVar{Name: "RESERVED_SYSTEM_CPUS", Value: wl.ReservedSystemCPU}) } if tlsMutualEnabled { mainEnv = append(mainEnv, @@ -291,7 +294,7 @@ fi` Spec: corev1.PodSpec{ ServiceAccountName: "simplyblock-storage-node-sa", HostNetwork: true, - Tolerations: sn.Spec.Tolerations, + Tolerations: wl.Tolerations, NodeSelector: map[string]string{ kube.LabelStorageNodeSet: sn.Name, }, @@ -312,7 +315,7 @@ fi` FieldRef: &corev1.ObjectFieldSelector{FieldPath: "spec.nodeName"}, }}, }, - Resources: effectiveResources(sn.Spec.InitContainerResources, defaultInitContainerResources), + Resources: effectiveResources(wl.InitContainerResources, defaultInitContainerResources), VolumeMounts: []corev1.VolumeMount{ {Name: "per-node-config", MountPath: "/etc/per-node-config", ReadOnly: true}, nodeEnvMount, @@ -327,7 +330,7 @@ fi` Command: initCmd, SecurityContext: &corev1.SecurityContext{Privileged: ptr.To(true)}, VolumeMounts: initMounts, - Resources: effectiveResources(sn.Spec.InitContainerResources, defaultInitContainerResources), + Resources: effectiveResources(wl.InitContainerResources, defaultInitContainerResources), Env: []corev1.EnvVar{ {Name: "HOSTNAME", ValueFrom: &corev1.EnvVarSource{ FieldRef: &corev1.ObjectFieldSelector{FieldPath: "spec.nodeName"}, @@ -346,7 +349,7 @@ fi` exec sudo -E python3 simplyblock_web/node_webapp.py storage_node_k8s`, }, SecurityContext: &corev1.SecurityContext{Privileged: ptr.To(true)}, - Resources: effectiveResources(sn.Spec.ContainerResources, defaultContainerResources), + Resources: effectiveResources(wl.ContainerResources, defaultContainerResources), ReadinessProbe: readinessProbe, Env: mainEnv, VolumeMounts: mainMounts, @@ -468,7 +471,7 @@ func StorageNodeSetAPIAddress(workerNode, namespace string) string { return fmt.Sprintf("%s.simplyblock-storage-node-api.%s.svc.cluster.local:5000", NodeHostnameLabel(workerNode), namespace) } -func BuildStorageNodeSetService(sn *simplyblockv1alpha1.StorageNodeSet, tlsEnabled bool, tlsProvider string) *corev1.Service { +func BuildStorageNodeService(sn *simplyblockv1alpha2.StorageCluster, tlsEnabled bool, tlsProvider string) *corev1.Service { return &corev1.Service{ ObjectMeta: metav1.ObjectMeta{ Name: kube.StorageNodeSetAPIServiceName, @@ -488,7 +491,7 @@ func BuildStorageNodeSetService(sn *simplyblockv1alpha1.StorageNodeSet, tlsEnabl } } -func BuildStorageNodeSetEndpointSlice(sn *simplyblockv1alpha1.StorageNodeSet, nodeIPs map[string]string) *discoveryv1.EndpointSlice { +func BuildStorageNodeEndpointSlice(sn *simplyblockv1alpha2.StorageCluster, nodeIPs map[string]string) *discoveryv1.EndpointSlice { protocol := corev1.ProtocolTCP port := int32(5000) portName := "api" @@ -532,7 +535,7 @@ type SpdkProxyEndpoint struct { RpcPort int32 } -func BuildSpdkProxyService(sn *simplyblockv1alpha1.StorageNodeSet, tlsEnabled bool, tlsProvider string) *corev1.Service { +func BuildSpdkProxyService(sn *simplyblockv1alpha2.StorageCluster, tlsEnabled bool, tlsProvider string) *corev1.Service { return &corev1.Service{ ObjectMeta: metav1.ObjectMeta{ Name: "simplyblock-spdk-proxy", @@ -551,7 +554,7 @@ func BuildSpdkProxyService(sn *simplyblockv1alpha1.StorageNodeSet, tlsEnabled bo // name truncated at the first dot so FQDN-style node names stay within the // 63-char DNS label limit. func BuildSpdkProxyEndpointSlice( - sn *simplyblockv1alpha1.StorageNodeSet, + sn *simplyblockv1alpha2.StorageCluster, rpcPort int32, endpoints []SpdkProxyEndpoint, ) (*discoveryv1.EndpointSlice, error) { @@ -640,3 +643,32 @@ func effectiveResources(user, def corev1.ResourceRequirements) corev1.ResourceRe } return def } + +// storageNodesOf is the cluster's workload block, never nil, so a builder reads +// defaults from a zero value rather than guarding every field. +// +// A cluster that states no block is the ordinary case: every field in it has a +// default or is legitimately empty, and the DaemonSet a zero block produces is the +// one a deployment that configured nothing asked for. +func storageNodesOf(cluster *simplyblockv1alpha2.StorageCluster) *simplyblockv1alpha2.StorageNodesSpec { + if cluster.Spec.StorageNodes == nil { + return &simplyblockv1alpha2.StorageNodesSpec{} + } + return cluster.Spec.StorageNodes +} + +// skipKubeletConfiguration renders the environment variable the storage node +// still reads, from the field that replaced it. +// +// It is written out rather than substituted because this is the one rename in the +// migration that also inverts: the retired skipKubeletConfiguration was off unless +// set, and enableKubeletConfiguration is off unless asked for, so the two say the +// opposite thing about the same deployment. A mechanical rename here would have +// turned kubelet configuration on for every cluster that never mentioned it +// (design-storagenode.md §15.1). +func skipKubeletConfiguration(wl *simplyblockv1alpha2.StorageNodesSpec) string { + if ptr.BoolFromOrFalse(wl.EnableKubeletConfiguration) { + return "false" + } + return "true" +} diff --git a/operator/internal/utils/storage_nodeset_ds_test.go b/operator/internal/utils/storage_node_workload_test.go similarity index 76% rename from operator/internal/utils/storage_nodeset_ds_test.go rename to operator/internal/utils/storage_node_workload_test.go index 1447053f3..410d890c5 100644 --- a/operator/internal/utils/storage_nodeset_ds_test.go +++ b/operator/internal/utils/storage_node_workload_test.go @@ -4,7 +4,7 @@ import ( "strings" "testing" - simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" corev1 "k8s.io/api/core/v1" "k8s.io/apimachinery/pkg/api/resource" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" @@ -34,7 +34,7 @@ func TestBuildStorageNodeSetClusterRoleBindingNameIncludesNamespace(t *testing.T } func TestBuildSpdkProxyEndpointSlice_DottedNodeNameTruncates(t *testing.T) { - sn := &simplyblockv1alpha1.StorageNodeSet{ + sn := &simplyblockv1alpha2.StorageCluster{ ObjectMeta: metav1.ObjectMeta{Name: "sn", Namespace: "ns"}, } endpoints := []SpdkProxyEndpoint{ @@ -66,7 +66,7 @@ func TestBuildSpdkProxyEndpointSlice_DottedNodeNameTruncates(t *testing.T) { } func TestBuildSpdkProxyEndpointSlice_CollidingFirstLabelFails(t *testing.T) { - sn := &simplyblockv1alpha1.StorageNodeSet{ + sn := &simplyblockv1alpha2.StorageCluster{ ObjectMeta: metav1.ObjectMeta{Name: "sn", Namespace: "ns"}, } endpoints := []SpdkProxyEndpoint{ @@ -84,24 +84,24 @@ func TestBuildSpdkProxyEndpointSlice_CollidingFirstLabelFails(t *testing.T) { } } -func TestBuildStorageNodeSetDaemonSetUserResourcesOverrideDefaults(t *testing.T) { - sn := &simplyblockv1alpha1.StorageNodeSet{ - ObjectMeta: metav1.ObjectMeta{Name: "sn", Namespace: "simplyblock"}, - Spec: simplyblockv1alpha1.StorageNodeSetSpec{ - ClusterName: "test-cluster", - ClusterImage: "simplyblock/simplyblock:latest", - ContainerResources: corev1.ResourceRequirements{ - Requests: corev1.ResourceList{corev1.ResourceMemory: resource.MustParse("1Gi")}, - Limits: corev1.ResourceList{corev1.ResourceMemory: resource.MustParse("4Gi")}, - }, - InitContainerResources: corev1.ResourceRequirements{ - Requests: corev1.ResourceList{corev1.ResourceMemory: resource.MustParse("64Mi")}, - Limits: corev1.ResourceList{corev1.ResourceMemory: resource.MustParse("128Mi")}, +func TestBuildStorageNodeDaemonSetUserResourcesOverrideDefaults(t *testing.T) { + sn := &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{Name: "test-cluster", Namespace: "simplyblock"}, + Spec: simplyblockv1alpha2.StorageClusterSpec{ + StorageNodes: &simplyblockv1alpha2.StorageNodesSpec{ + ContainerResources: corev1.ResourceRequirements{ + Requests: corev1.ResourceList{corev1.ResourceMemory: resource.MustParse("1Gi")}, + Limits: corev1.ResourceList{corev1.ResourceMemory: resource.MustParse("4Gi")}, + }, + InitContainerResources: corev1.ResourceRequirements{ + Requests: corev1.ResourceList{corev1.ResourceMemory: resource.MustParse("64Mi")}, + Limits: corev1.ResourceList{corev1.ResourceMemory: resource.MustParse("128Mi")}, + }, }, }, } - ds := BuildStorageNodeSetDaemonSet(sn, false, false, "", "") + ds := BuildStorageNodeDaemonSet(sn, false, false, "", "", "simplyblock/simplyblock:latest") main := ds.Spec.Template.Spec.Containers[0] mainMem := main.Resources.Limits[corev1.ResourceMemory] diff --git a/operator/internal/webhook/conversion.go b/operator/internal/webhook/conversion.go index 895192bff..81204846a 100644 --- a/operator/internal/webhook/conversion.go +++ b/operator/internal/webhook/conversion.go @@ -37,6 +37,7 @@ var convertedKinds = []convertedKind{ {hub: &v1alpha2.StorageBackup{}, crdName: "storagebackups.storage.simplyblock.io"}, {hub: &v1alpha2.StorageCluster{}, crdName: "storageclusters.storage.simplyblock.io"}, {hub: &v1alpha2.StorageClusterOps{}, crdName: "storageclusterops.storage.simplyblock.io"}, + {hub: &v1alpha2.StorageNode{}, crdName: "storagenodes.storage.simplyblock.io"}, {hub: &v1alpha2.StorageNodeOps{}, crdName: "storagenodeops.storage.simplyblock.io"}, {hub: &v1alpha2.StoragePool{}, crdName: "storagepools.storage.simplyblock.io"}, } diff --git a/operator/internal/webhook/conversion_trust.go b/operator/internal/webhook/conversion_trust.go index 3abb9253d..d124720b9 100644 --- a/operator/internal/webhook/conversion_trust.go +++ b/operator/internal/webhook/conversion_trust.go @@ -38,7 +38,7 @@ import ( // compromised webhook is worth bounding, even though the resourceNames list has // to be kept beside the one in config/crd/converted-kinds.txt. // TestManagerRoleNamesOnlyTheConvertedCRDs is what keeps the two in step. -// +kubebuilder:rbac:groups=apiextensions.k8s.io,resources=customresourcedefinitions,verbs=get;update;patch,resourceNames=controlplanes.storage.simplyblock.io;storagebackups.storage.simplyblock.io;storageclusterops.storage.simplyblock.io;storageclusters.storage.simplyblock.io;storagenodeops.storage.simplyblock.io;storagepools.storage.simplyblock.io +// +kubebuilder:rbac:groups=apiextensions.k8s.io,resources=customresourcedefinitions,verbs=get;update;patch,resourceNames=controlplanes.storage.simplyblock.io;storagebackups.storage.simplyblock.io;storageclusterops.storage.simplyblock.io;storageclusters.storage.simplyblock.io;storagenodeops.storage.simplyblock.io;storagenodes.storage.simplyblock.io;storagepools.storage.simplyblock.io // // Reading stays cluster-wide because it cannot be otherwise: cert-controller's // rotator establishes an informer on CustomResourceDefinition to re-inject the CA diff --git a/operator/internal/webhook/storagenode_validator.go b/operator/internal/webhook/storagenode_validator.go index dd8214fd2..027af7b7d 100644 --- a/operator/internal/webhook/storagenode_validator.go +++ b/operator/internal/webhook/storagenode_validator.go @@ -1,3 +1,25 @@ +// The StorageNode admission guard, which answers two different kinds of question. +// +// On create it resolves spec.clusterRef and checks the node's configuration +// against the cluster it names: whether the devices it lists are of the class the +// cluster is built out of, and whether its sizing agrees with the fleet's. Both +// are properties of the cluster rather than of the request, so neither can be +// decided from the admission review alone. +// +// On update it enforces the three fields that have exactly one legitimate writer. +// spec.workerNode, spec.config.pcieAllowList, and spec.config.sizing are all +// written by the operator and by nobody else, so a +k8s:immutable marker would +// lock the operator out along with everyone else and no marker at all would let a +// user invalidate a layout claim by editing a string. +// +// failurePolicy=Fail, and that is safe because the webhook server runs in the +// operator pod: its availability tracks the operator's own, and an operator that +// is down is not reconciling anything the rejection could deadlock. The three +// guarded fields have no CRD-level immutability behind them, so this is their only +// guard and admitting while unavailable would let the edit through. +// +// design-storagenode.md §3.2 and §3.4 are the specification. + package webhook import ( @@ -5,66 +27,292 @@ import ( "encoding/json" "fmt" "net/http" + "regexp" "strings" admissionv1 "k8s.io/api/admission/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" + "sigs.k8s.io/controller-runtime/pkg/client" "sigs.k8s.io/controller-runtime/pkg/webhook/admission" - simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" ) -// +kubebuilder:webhook:path=/validate-storage-simplyblock-io-v1alpha1-storagenode,mutating=false,failurePolicy=fail,sideEffects=None,groups=storage.simplyblock.io,resources=storagenodes,verbs=update,versions=v1alpha1,name=vstoragenode.simplyblock.io,admissionReviewVersions=v1 +// matchPolicy=Equivalent is load-bearing rather than a default worth inheriting. +// The CRD serves v1alpha1 as well, and under the API server's default Exact policy +// a rule naming only v1alpha2 does not see a v1alpha1 write at all — so a client +// writing the older version would place a node naming a cluster that does not +// exist, or re-point a worker this guard exists to hold. -// StorageNodeValidator is a validating admission webhook that enforces -// spec.workerNode is only ever re-pointed by the operator itself. A StorageNode -// is bound to a worker host; users must not move it by editing the CR directly — -// relocation is driven exclusively through a StorageNodeOps(action=migrate), -// which the operator executes (drain-free restart onto the target host) before -// re-pointing spec.workerNode under its own service account. -// -// failurePolicy=Fail: the StorageNode CRD does not enforce spec.workerNode -// immutability (the operator must be able to re-point it), so this webhook is -// the sole guard against user-driven repoints. It must fail closed — rejecting -// edits while unavailable is preferable to silently letting a repoint through. -// This does not risk deadlocking the operator's own repoints: the webhook -// server runs inside the operator pod, so its availability tracks the -// operator's — when the operator is up the webhook is up and admits its own -// service-account UPDATE, and when the operator is down nothing reconciles -// anyway. (The rebalancer injector stays failurePolicy=Ignore because it -// gates pod creation, which must not block on webhook availability.) +// +kubebuilder:webhook:path=/validate-storage-simplyblock-io-v1alpha2-storagenode,mutating=false,failurePolicy=fail,matchPolicy=Equivalent,sideEffects=None,groups=storage.simplyblock.io,resources=storagenodes,verbs=create;update,versions=v1alpha2,name=vstoragenode.simplyblock.io,admissionReviewVersions=v1 + +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storageclusters,verbs=get;list;watch + +// StorageNodeValidator admits a StorageNode write. type StorageNodeValidator struct { - // OperatorNamespace is the namespace the operator runs in. Any service - // account in this namespace (i.e. the operator itself) is permitted to change - // spec.workerNode; every other identity is rejected. + // Client reads the StorageCluster the node names. + Client client.Client + + // OperatorNamespace is the namespace the operator runs in. A service account + // in it is the operator itself, and is the one identity permitted to write the + // three fields of §3.2. OperatorNamespace string } -func (v *StorageNodeValidator) Handle(_ context.Context, req admission.Request) admission.Response { - if req.Operation != admissionv1.Update { +func (v *StorageNodeValidator) Handle( + ctx context.Context, req admission.Request, +) admission.Response { + switch req.Operation { + case admissionv1.Create: + var node simplyblockv1alpha2.StorageNode + if err := json.Unmarshal(req.Object.Raw, &node); err != nil { + return admission.Errored(http.StatusBadRequest, err) + } + return v.admitCreate(ctx, &node, v.isOperator(req)) + + case admissionv1.Update: + var oldNode, newNode simplyblockv1alpha2.StorageNode + if err := json.Unmarshal(req.OldObject.Raw, &oldNode); err != nil { + return admission.Errored(http.StatusBadRequest, err) + } + if err := json.Unmarshal(req.Object.Raw, &newNode); err != nil { + return admission.Errored(http.StatusBadRequest, err) + } + return v.admitUpdate(&oldNode, &newNode, v.isOperator(req)) + + default: return admission.Allowed("") } +} + +// admitCreate resolves the cluster and checks the node against it. +func (v *StorageNodeValidator) admitCreate( + ctx context.Context, node *simplyblockv1alpha2.StorageNode, byOperator bool, +) admission.Response { + // The reference is immutable from creation, which is what makes admission the + // right place. A node naming a cluster that is not there can never be + // corrected — the field cannot be edited, so the only remedy is to delete the + // object and write it again, which is exactly what this rejection asks for. + var cluster simplyblockv1alpha2.StorageCluster + key := client.ObjectKey{Namespace: node.Namespace, Name: node.Spec.ClusterRef} + switch err := v.Client.Get(ctx, key, &cluster); { + case apierrors.IsNotFound(err): + return admission.Denied(fmt.Sprintf( + "spec.clusterRef names StorageCluster %q, which does not exist in namespace %q", + node.Spec.ClusterRef, node.Namespace)) + case err != nil: + return admission.Errored(http.StatusInternalServerError, err) + } - var oldSN, newSN simplyblockv1alpha1.StorageNode - if err := json.Unmarshal(req.OldObject.Raw, &oldSN); err != nil { - return admission.Errored(http.StatusBadRequest, err) + if denial := deviceClassDenial(&node.Spec.Config, &cluster); denial != "" { + return admission.Denied(denial) } - if err := json.Unmarshal(req.Object.Raw, &newSN); err != nil { - return admission.Errored(http.StatusBadRequest, err) + + // The operator writes a node's sizing itself, and a rolling hardware upgrade + // is the case where it deliberately differs from the fleet's. A user's node is + // held to the cluster's values, because unmanaged divergence is what stops the + // control plane placing erasure-coding chunks evenly (§3.1). + if !byOperator { + if denial := sizingDenial(&node.Spec.Config.Sizing, &cluster); denial != "" { + return admission.Denied(denial) + } } + return admission.Allowed("") +} - if oldSN.Spec.WorkerNode == newSN.Spec.WorkerNode { +// admitUpdate enforces the three fields with exactly one legitimate writer. +func (v *StorageNodeValidator) admitUpdate( + oldNode, newNode *simplyblockv1alpha2.StorageNode, byOperator bool, +) admission.Response { + changed := operatorOnlyChanges(oldNode, newNode) + if len(changed) == 0 { return admission.Allowed("") } + if byOperator { + return admission.Allowed("operator-driven change to " + strings.Join(changed, ", ")) + } + return admission.Denied(fmt.Sprintf( + "%s %s written by the operator alone; "+ + "relocate a node with a StorageNodeOps of action Migrate, and let a "+ + "re-size follow the fleet rather than editing it here", + strings.Join(changed, " and "), plural(len(changed), "is", "are"))) +} + +// operatorOnlyChanges names the guarded fields this update would change. +func operatorOnlyChanges(oldNode, newNode *simplyblockv1alpha2.StorageNode) []string { + var changed []string + if oldNode.Spec.WorkerNode != newNode.Spec.WorkerNode { + changed = append(changed, "spec.workerNode") + } + // The allow list is the one device field a migration writes, merging the + // drives bound on the target host into it so they survive a later rebuild + // (§3.2). That is why it is guarded here rather than marked immutable. + if !equalStrings(oldNode.Spec.Config.PcieAllowList, newNode.Spec.Config.PcieAllowList) { + changed = append(changed, "spec.config.pcieAllowList") + } + if !equalSizing(oldNode.Spec.Config.Sizing, newNode.Spec.Config.Sizing) { + changed = append(changed, "spec.config.sizing") + } + return changed +} - // Only the operator (a service account in the operator namespace) may - // re-point a StorageNode to a different worker. - operatorSAPrefix := "system:serviceaccount:" + v.OperatorNamespace + ":" - if strings.HasPrefix(req.UserInfo.Username, operatorSAPrefix) { - return admission.Allowed("operator-driven workerNode change") +// deviceClassDenial reports why a node's device configuration does not belong to +// its cluster's class, or the empty string when it does. +// +// A cluster is built out of one class of backend storage, because an +// erasure-coding stripe placed across both is written and rebuilt at the slower +// one's rate (§3.1). So every entry of a node's list is of its cluster's class, +// a list holding both is rejected, and so is a list of the class the cluster is +// not. +func deviceClassDenial( + config *simplyblockv1alpha2.StorageNodeConfig, cluster *simplyblockv1alpha2.StorageCluster, +) string { + // An unstated class is NVMe, which is what the CRD defaults it to and what + // describes every cluster that predates the field. + class := cluster.Spec.DeviceClass + if class == "" { + class = simplyblockv1alpha2.StorageClusterDeviceClassNVMe } - return admission.Denied(fmt.Sprintf( - "spec.workerNode is immutable to users (attempted %q -> %q); "+ - "relocate a storage node with a StorageNodeOps(action=migrate) instead", - oldSN.Spec.WorkerNode, newSN.Spec.WorkerNode)) + var addresses, paths []string + for _, name := range config.DeviceNames { + if pciAddress.MatchString(name) { + addresses = append(addresses, name) + continue + } + paths = append(paths, name) + } + + if len(addresses) > 0 && len(paths) > 0 { + return fmt.Sprintf( + "spec.config.deviceNames mixes PCI addresses (%s) with device paths (%s); "+ + "a cluster is built out of one class of backend storage", + strings.Join(addresses, ", "), strings.Join(paths, ", ")) + } + + switch class { + case simplyblockv1alpha2.StorageClusterDeviceClassNVMe: + if len(paths) > 0 { + return fmt.Sprintf( + "spec.config.deviceNames holds device paths (%s) and cluster %q has "+ + "deviceClass NVMe, whose devices are named by PCI address", + strings.Join(paths, ", "), cluster.Name) + } + + case simplyblockv1alpha2.StorageClusterDeviceClassLogicalBlock: + if len(addresses) > 0 { + return fmt.Sprintf( + "spec.config.deviceNames holds PCI addresses (%s) and cluster %q has "+ + "deviceClass LogicalBlock, whose devices are named by path", + strings.Join(addresses, ", "), cluster.Name) + } + // The PCI filters match on something a logical block device does not have, + // so they are rejected rather than ignored: a filter that silently selects + // nothing is a node that comes up with no devices and no reason given. + var stated []string + if len(config.PcieAllowList) > 0 { + stated = append(stated, "spec.config.pcieAllowList") + } + if len(config.PcieDenyList) > 0 { + stated = append(stated, "spec.config.pcieDenyList") + } + if config.PcieModel != "" { + stated = append(stated, "spec.config.pcieModel") + } + if len(stated) > 0 { + return fmt.Sprintf( + "%s %s stated and cluster %q has deviceClass LogicalBlock, whose "+ + "devices have no PCI address to match", + strings.Join(stated, " and "), plural(len(stated), "is", "are"), cluster.Name) + } + } + return "" +} + +// sizingDenial reports why a node's sizing disagrees with its cluster's, or the +// empty string when it agrees. +// +// Both values are stated once, on the cluster, and a node holds a stamp of what it +// was built with rather than a number somebody chose for it. A fleet whose nodes +// differ is a fleet mid-roll, never a fleet somebody described that way. +func sizingDenial( + sizing *simplyblockv1alpha2.StorageNodeSizing, cluster *simplyblockv1alpha2.StorageCluster, +) string { + if want := cluster.Spec.VCPUCount; want != nil { + if sizing.VCPUCount == nil || *sizing.VCPUCount != *want { + return fmt.Sprintf( + "spec.config.sizing.vcpuCount is %s and cluster %q states %d; "+ + "a node's sizing is stamped from its cluster, and a roll that "+ + "changes it is the operator's to perform", + describeCount(sizing.VCPUCount), cluster.Name, *want) + } + } + if want := cluster.Spec.MinHugePagesSize; want != "" && sizing.MinHugePagesSize != want { + return fmt.Sprintf( + "spec.config.sizing.minHugePagesSize is %q and cluster %q states %q", + sizing.MinHugePagesSize, cluster.Name, want) + } + return "" +} + +// isOperator reports whether the request came from a service account in the +// operator's own namespace, which is the operator itself. +func (v *StorageNodeValidator) isOperator(req admission.Request) bool { + return strings.HasPrefix(req.UserInfo.Username, + "system:serviceaccount:"+v.OperatorNamespace+":") +} + +// pciAddress matches the domain:bus:device.function form a PCI address takes. It +// is the same expression the type's own items pattern carries for that half of +// the union, so the two cannot disagree about what an address looks like. +var pciAddress = regexp.MustCompile( + `^[0-9a-fA-F]{4}:[0-9a-fA-F]{2}:[0-9a-fA-F]{2}\.[0-9a-fA-F]$`) + +// describeCount renders a count for a message, so an absent one reads as absent +// rather than as zero. +func describeCount(count *int32) string { + if count == nil { + return "unset" + } + return fmt.Sprintf("%d", *count) +} + +// plural picks the verb for a count, so a message reads "is stated" for one field +// and "are stated" for several. +func plural(n int, one, many string) string { + if n == 1 { + return one + } + return many +} + +// equalSizing compares two sizing blocks by value. +// +// It is spelled out rather than compared with ==, because the core count is a +// pointer and two separately decoded objects hold two pointers to the same +// number. Comparing the structs would read every update as a re-size and refuse +// every edit a user is entitled to make. +func equalSizing(a, b simplyblockv1alpha2.StorageNodeSizing) bool { + if a.MinHugePagesSize != b.MinHugePagesSize { + return false + } + if a.VCPUCount == nil || b.VCPUCount == nil { + return a.VCPUCount == b.VCPUCount + } + return *a.VCPUCount == *b.VCPUCount +} + +// equalStrings compares two lists for the purpose of deciding whether an update +// changed one. Order matters: the allow list is a user's, and a reordering is a +// write of the field whoever did it has to be allowed to make. +func equalStrings(a, b []string) bool { + if len(a) != len(b) { + return false + } + for i := range a { + if a[i] != b[i] { + return false + } + } + return true } diff --git a/operator/internal/webhook/storagenode_validator_test.go b/operator/internal/webhook/storagenode_validator_test.go index e044ba174..c4d1e23dc 100644 --- a/operator/internal/webhook/storagenode_validator_test.go +++ b/operator/internal/webhook/storagenode_validator_test.go @@ -1,63 +1,341 @@ +// What the StorageNode guard admits and refuses. +// +// The two halves are tested separately because they answer different questions. +// The update half is about who is writing, and needs no cluster; the create half +// is about what the node says against the cluster it names, and is decided +// entirely by reading that object. + package webhook import ( "context" "encoding/json" + "strings" "testing" admissionv1 "k8s.io/api/admission/v1" authenticationv1 "k8s.io/api/authentication/v1" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" "k8s.io/apimachinery/pkg/runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" "sigs.k8s.io/controller-runtime/pkg/webhook/admission" - simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" + "github.com/simplyblock/atlas/ptr" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" ) -func snRaw(t *testing.T, worker string) runtime.RawExtension { - t.Helper() - sn := &simplyblockv1alpha1.StorageNode{ +const operatorNS = "simplyblock-operator-system" + +// operatorUser and someoneElse are the two identities every case here is one of. +const ( + operatorUser = "system:serviceaccount:" + operatorNS + ":simplyblock-operator" + someoneElse = "kubernetes-admin" +) + +// testNode is a node that agrees with testStorageCluster in every respect, so a +// case states only what it is about. +func testNode(mutate func(*simplyblockv1alpha2.StorageNode)) *simplyblockv1alpha2.StorageNode { + node := &simplyblockv1alpha2.StorageNode{ ObjectMeta: metav1.ObjectMeta{Name: "sn-1", Namespace: "default"}, - Spec: simplyblockv1alpha1.StorageNodeSpec{StorageNodeSetRef: "set", WorkerNode: worker}, + Spec: simplyblockv1alpha2.StorageNodeSpec{ + ClusterRef: "production", + WorkerNode: "worker-2", + Config: simplyblockv1alpha2.StorageNodeConfig{ + Sizing: simplyblockv1alpha2.StorageNodeSizing{ + VCPUCount: ptr.To(int32(8)), + MinHugePagesSize: "100G", + }, + }, + }, + } + if mutate != nil { + mutate(node) + } + return node +} + +// testStorageCluster is an NVMe cluster sized at eight vCPUs. +func nodeTestCluster( + mutate func(*simplyblockv1alpha2.StorageCluster), +) *simplyblockv1alpha2.StorageCluster { + cluster := &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{Name: "production", Namespace: "default"}, + Spec: simplyblockv1alpha2.StorageClusterSpec{ + VCPUCount: ptr.To(int32(8)), + MinHugePagesSize: "100G", + DeviceClass: simplyblockv1alpha2.StorageClusterDeviceClassNVMe, + }, + } + if mutate != nil { + mutate(cluster) } - b, err := json.Marshal(sn) + return cluster +} + +func validatorFor(t *testing.T, objects ...client.Object) *StorageNodeValidator { + t.Helper() + scheme := runtime.NewScheme() + if err := simplyblockv1alpha2.AddToScheme(scheme); err != nil { + t.Fatalf("add to scheme: %v", err) + } + return &StorageNodeValidator{ + Client: fake.NewClientBuilder().WithScheme(scheme).WithObjects(objects...).Build(), + OperatorNamespace: operatorNS, + } +} + +func raw(t *testing.T, object any) runtime.RawExtension { + t.Helper() + encoded, err := json.Marshal(object) if err != nil { - t.Fatalf("marshal StorageNode: %v", err) + t.Fatalf("marshal: %v", err) } - return runtime.RawExtension{Raw: b} + return runtime.RawExtension{Raw: encoded} } -func TestStorageNodeValidator(t *testing.T) { - const ns = "simplyblock-operator-system" - v := &StorageNodeValidator{OperatorNamespace: ns} +func createReq(t *testing.T, node *simplyblockv1alpha2.StorageNode, user string) admission.Request { + t.Helper() + return admission.Request{AdmissionRequest: admissionv1.AdmissionRequest{ + Operation: admissionv1.Create, + Object: raw(t, node), + UserInfo: authenticationv1.UserInfo{Username: user}, + }} +} - tests := []struct { - name string - op admissionv1.Operation - oldWorker string - newWorker string - username string - allowed bool +func updateReq( + t *testing.T, oldNode, newNode *simplyblockv1alpha2.StorageNode, user string, +) admission.Request { + t.Helper() + return admission.Request{AdmissionRequest: admissionv1.AdmissionRequest{ + Operation: admissionv1.Update, + Object: raw(t, newNode), + OldObject: raw(t, oldNode), + UserInfo: authenticationv1.UserInfo{Username: user}, + }} +} + +// The three fields of §3.2 have exactly one legitimate writer. A marker would lock +// the operator out along with everyone else, so this webhook is their only guard. +func TestTheOperatorOnlyFieldsAreRefusedToEveryoneElse(t *testing.T) { + validator := validatorFor(t, nodeTestCluster(nil)) + + for _, tc := range []struct { + field string + mutate func(*simplyblockv1alpha2.StorageNode) }{ - {"non-update allowed", admissionv1.Create, "", "worker-2", "someone", true}, - {"workerNode unchanged allowed", admissionv1.Update, "worker-2", "worker-2", "kubernetes-admin", true}, - {"operator may repoint", admissionv1.Update, "worker-2", "worker-4", "system:serviceaccount:" + ns + ":simplyblock-operator", true}, - {"user may not repoint", admissionv1.Update, "worker-2", "worker-4", "kubernetes-admin", false}, - {"other-namespace SA may not repoint", admissionv1.Update, "worker-2", "worker-4", "system:serviceaccount:kube-system:foo", false}, + {"spec.workerNode", func(n *simplyblockv1alpha2.StorageNode) { + n.Spec.WorkerNode = "worker-4" + }}, + {"spec.config.pcieAllowList", func(n *simplyblockv1alpha2.StorageNode) { + n.Spec.Config.PcieAllowList = []string{"0000:5e:00.0"} + }}, + {"spec.config.sizing", func(n *simplyblockv1alpha2.StorageNode) { + n.Spec.Config.Sizing.VCPUCount = ptr.To(int32(16)) + }}, + } { + t.Run(tc.field, func(t *testing.T) { + before, after := testNode(nil), testNode(tc.mutate) + + refused := validator.Handle(context.Background(), updateReq(t, before, after, someoneElse)) + if refused.Allowed { + t.Errorf("%s was admitted from a user and only the operator may write it", tc.field) + } + if !strings.Contains(refused.Result.Message, tc.field) { + t.Errorf("the refusal does not name %s: %s", tc.field, refused.Result.Message) + } + + admitted := validator.Handle(context.Background(), updateReq(t, before, after, operatorUser)) + if !admitted.Allowed { + t.Errorf("%s was refused to the operator, which is its one writer: %s", + tc.field, admitted.Result.Message) + } + }) } +} + +// An update that touches none of the three is nobody's business but the writer's. +func TestAnUpdateTouchingNoGuardedFieldIsAdmitted(t *testing.T) { + validator := validatorFor(t, nodeTestCluster(nil)) + before := testNode(nil) + after := testNode(func(n *simplyblockv1alpha2.StorageNode) { + n.Spec.Config.SpdkSystemMemory = "8G" + }) - for _, tc := range tests { + response := validator.Handle(context.Background(), updateReq(t, before, after, someoneElse)) + if !response.Allowed { + t.Errorf("a mutable field was refused: %s", response.Result.Message) + } +} + +// spec.clusterRef is immutable from creation, so a node naming a cluster that is +// not there can never be corrected. Refusing the create asks for the delete-and- +// rewrite that is the only remedy anyway. +func TestANodeNamingNoClusterIsRefused(t *testing.T) { + validator := validatorFor(t) + + response := validator.Handle(context.Background(), createReq(t, testNode(nil), someoneElse)) + if response.Allowed { + t.Fatal("a node naming a cluster that does not exist was admitted") + } + if !strings.Contains(response.Result.Message, "does not exist") { + t.Errorf("the refusal does not say the cluster is missing: %s", response.Result.Message) + } +} + +// A cluster is built out of one class of backend storage, because an +// erasure-coding stripe placed across both is written and rebuilt at the slower +// one's rate. +func TestDeviceNamesMustBeOfTheClusterSClass(t *testing.T) { + for _, tc := range []struct { + name string + class simplyblockv1alpha2.StorageClusterDeviceClass + devices []string + refused string + }{ + { + name: "a path on an NVMe cluster", + class: simplyblockv1alpha2.StorageClusterDeviceClassNVMe, + devices: []string{"/dev/sdb"}, + refused: "named by PCI address", + }, + { + name: "a PCI address on a logical-block cluster", + class: simplyblockv1alpha2.StorageClusterDeviceClassLogicalBlock, + devices: []string{"0000:5e:00.0"}, + refused: "named by path", + }, + { + name: "both at once", + class: simplyblockv1alpha2.StorageClusterDeviceClassNVMe, + devices: []string{"0000:5e:00.0", "/dev/sdb"}, + refused: "mixes PCI addresses", + }, + } { + t.Run(tc.name, func(t *testing.T) { + validator := validatorFor(t, nodeTestCluster( + func(c *simplyblockv1alpha2.StorageCluster) { c.Spec.DeviceClass = tc.class })) + node := testNode(func(n *simplyblockv1alpha2.StorageNode) { + n.Spec.Config.DeviceNames = tc.devices + }) + + response := validator.Handle(context.Background(), createReq(t, node, someoneElse)) + if response.Allowed { + t.Fatalf("%v was admitted on a %s cluster", tc.devices, tc.class) + } + if !strings.Contains(response.Result.Message, tc.refused) { + t.Errorf("the refusal does not explain the class: %s", response.Result.Message) + } + }) + } +} + +// Each class admits its own spelling, which is the other half of the rule above. +func TestDeviceNamesOfTheClusterSClassAreAdmitted(t *testing.T) { + for class, devices := range map[simplyblockv1alpha2.StorageClusterDeviceClass][]string{ + simplyblockv1alpha2.StorageClusterDeviceClassNVMe: {"0000:5e:00.0", "0000:5f:00.0"}, + simplyblockv1alpha2.StorageClusterDeviceClassLogicalBlock: {"/dev/sdb", "nvme0n1"}, + } { + validator := validatorFor(t, nodeTestCluster( + func(c *simplyblockv1alpha2.StorageCluster) { c.Spec.DeviceClass = class })) + node := testNode(func(n *simplyblockv1alpha2.StorageNode) { + n.Spec.Config.DeviceNames = devices + }) + + response := validator.Handle(context.Background(), createReq(t, node, someoneElse)) + if !response.Allowed { + t.Errorf("%v was refused on a %s cluster: %s", devices, class, response.Result.Message) + } + } +} + +// The PCI filters match on something a logical block device does not have, so they +// are refused rather than ignored: a filter that silently selects nothing is a node +// that comes up with no devices and no reason given. +func TestThePCIFiltersAreRefusedOnALogicalBlockCluster(t *testing.T) { + validator := validatorFor(t, nodeTestCluster(func(c *simplyblockv1alpha2.StorageCluster) { + c.Spec.DeviceClass = simplyblockv1alpha2.StorageClusterDeviceClassLogicalBlock + })) + + for field, mutate := range map[string]func(*simplyblockv1alpha2.StorageNode){ + "spec.config.pcieAllowList": func(n *simplyblockv1alpha2.StorageNode) { + n.Spec.Config.PcieAllowList = []string{"0000:5e:00.0"} + }, + "spec.config.pcieDenyList": func(n *simplyblockv1alpha2.StorageNode) { + n.Spec.Config.PcieDenyList = []string{"0000:5e:00.0"} + }, + "spec.config.pcieModel": func(n *simplyblockv1alpha2.StorageNode) { + n.Spec.Config.PcieModel = "Samsung" + }, + } { + response := validator.Handle(context.Background(), + createReq(t, testNode(mutate), someoneElse)) + if response.Allowed { + t.Errorf("%s was admitted on a logical-block cluster", field) + continue + } + if !strings.Contains(response.Result.Message, field) { + t.Errorf("the refusal does not name %s: %s", field, response.Result.Message) + } + } +} + +// A node's sizing is a stamp of what its cluster was built with. A fleet whose +// nodes differ is a fleet mid-roll, never a fleet somebody described that way. +func TestAUserSNodeMustAgreeWithTheFleetSSizing(t *testing.T) { + validator := validatorFor(t, nodeTestCluster(nil)) + + for _, tc := range []struct { + name string + mutate func(*simplyblockv1alpha2.StorageNode) + }{ + {"a different core count", func(n *simplyblockv1alpha2.StorageNode) { + n.Spec.Config.Sizing.VCPUCount = ptr.To(int32(16)) + }}, + {"no core count at all", func(n *simplyblockv1alpha2.StorageNode) { + n.Spec.Config.Sizing.VCPUCount = nil + }}, + {"a different huge-page floor", func(n *simplyblockv1alpha2.StorageNode) { + n.Spec.Config.Sizing.MinHugePagesSize = "1T" + }}, + } { t.Run(tc.name, func(t *testing.T) { - req := admission.Request{AdmissionRequest: admissionv1.AdmissionRequest{ - Operation: tc.op, - Object: snRaw(t, tc.newWorker), - OldObject: snRaw(t, tc.oldWorker), - UserInfo: authenticationv1.UserInfo{Username: tc.username}, - }} - resp := v.Handle(context.Background(), req) - if resp.Allowed != tc.allowed { - t.Fatalf("Allowed = %v, want %v (msg: %s)", resp.Allowed, tc.allowed, resp.Result.Message) + response := validator.Handle(context.Background(), + createReq(t, testNode(tc.mutate), someoneElse)) + if response.Allowed { + t.Errorf("%s was admitted from a user", tc.name) } }) } } + +// The operator writes a node's sizing itself, and a rolling hardware upgrade is +// the case where it deliberately differs. A model that refused that would force +// the whole fleet to be re-sized at once or not at all. +func TestTheOperatorMayCreateANodeSizedAgainstTheFleet(t *testing.T) { + validator := validatorFor(t, nodeTestCluster(nil)) + node := testNode(func(n *simplyblockv1alpha2.StorageNode) { + n.Spec.Config.Sizing.VCPUCount = ptr.To(int32(16)) + }) + + response := validator.Handle(context.Background(), createReq(t, node, operatorUser)) + if !response.Allowed { + t.Errorf("a mid-roll node was refused to the operator: %s", response.Result.Message) + } +} + +// An unstated class is NVMe, which is what the CRD defaults it to and what +// describes every cluster that predates the field. +func TestAClusterWithNoStatedClassIsNVMe(t *testing.T) { + validator := validatorFor(t, nodeTestCluster( + func(c *simplyblockv1alpha2.StorageCluster) { c.Spec.DeviceClass = "" })) + node := testNode(func(n *simplyblockv1alpha2.StorageNode) { + n.Spec.Config.DeviceNames = []string{"/dev/sdb"} + }) + + response := validator.Handle(context.Background(), createReq(t, node, someoneElse)) + if response.Allowed { + t.Error("a device path was admitted on a cluster that states no class, which is NVMe") + } +} From 0f9dec704d220100dde66d7b80bd3f7f02a8512d Mon Sep 17 00:00:00 2001 From: noctarius aka Christoph Engelbert Date: Tue, 15 Sep 2026 13:11:02 +0200 Subject: [PATCH 002/206] feat(operator): the ClusterDeploymentConfig controller, which is what creates a cluster and its nodes (#543) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * feat(operator): the ClusterDeploymentConfig controller, which is what creates a cluster and its nodes Retiring StorageNodeSet left nothing producing StorageNode objects: the set's reconciler was the only one, so a deployment following the shipped chart got a DaemonSet and no backend nodes. This is the producer the redesign puts in its place, and it makes a whole deployment one reviewable document rather than a set of objects somebody assembles by hand. The document is ephemeral and everything follows from that. It owns nothing, nothing references it, and nothing reads it after the expansion, so there is no finalizer and deleting it deletes a document. It specifically does not own the StorageCluster it created: an owner reference would make deleting the document delete the cluster and every volume in it. A draft is validated on every reconcile and expanded on none. Validation writes what it found and nothing else, so a reviewer sees the problems before approving rather than after, which is the whole value of the gate and why a draft naming a worker that does not exist is storable at all. It is also the only check a device gets, since admission does not look at devices: a config discovery wrote cannot be wrong about them, a hand-written one can, and approving without reading is how a device mistake becomes an immutable document. The expansion is create-only, which is §6's decision and the one with a real cost. A document that names an existing cluster without asking to is refused with ClusterExists rather than merged, because merging would have the operator decide what a difference means, and the differences that matter are of the form: this node's device list changed, whose only correct handling is not to apply it to a node that already has data on those devices. Growth is a second document naming the cluster in spec.clusterRef, which keeps every config a record of one deployment action. Two shorthands are spent here and nowhere else. spec.environment resolves into the distribution flags on the cluster's workload, and the cluster's sizing is copied into every node's spec.config.sizing, so the nodes carry what they were built with and deleting the document loses nothing. The device class is read off the groups and stamped onto the cluster, because the device lists already say which class the deployment uses and the document carries no field for it. CreatingNodes is the step that must be idempotent, and it is by construction: a node is identified by its cluster, its worker, and its slot, so the step lists what exists and creates only the slots that do not. Three passes over one document produce the same two nodes, which is the test that matters. What is not here: discovery, which is OperatorOps's (§8), and the approval webhook of §5.2. Both are next steps rather than stubs. * feat(operator): a fresh install discovers the fleet by itself An operator installed into a cluster with disks in it can tell the administrator what it found, and that is a better first experience than an empty namespace and a document to write by hand. So a fresh install raises one OperatorOps with action Discover, whose output is a ClusterDeploymentConfig in Draft that nobody has approved and nothing acts on. The guard is the whole of it. A discovery run is read-only and the draft it writes is inert, but Probing creates one Job per worker, so a run that fired on every restart would put a Job on every node of the fleet each time the operator was upgraded. It therefore runs only where no previous result exists, which is three questions rather than one: whether any OperatorOps has ever run, whether any ClusterDeploymentConfig exists, and whether any StorageCluster is deployed. A terminal run counts, and so does a failed one — the administrator has seen the answer either way, and this operator is not the thing to decide they want another. It is a leader-elected Runnable rather than a reconciler, because there is no object whose desired state it converges toward and the question has one answer per installation. The create is idempotent by name on top of the guard, since two replicas answering the same three questions at once would both conclude yes. A failure at any point is logged and swallowed: the operator works without the run, and what is lost is a draft rather than a capability. An administrator who does not want it says so by writing an object named initial-discovery, which the guard then finds and declines behind. The chart's operator_customresources.yaml is rewritten to the path this creates. It was still shipping a StorageNodeSet that nothing has reconciled since #542, and a StorageCluster beside it that would have made the expansion refuse with ClusterExists. Nothing in the chart applies the file — it is a hand-apply reference — so what it documents is now a ClusterDeploymentConfig an administrator reviews, which is what discovery writes, and the pool sample is corrected to the v1alpha2 shape it has had since the pool's own move. * feat(operator): discovery reads what a machine is for, and proposes the storage tier first A discovery run knew a worker's disks, its memory, its taints, and whether it was cordoned, and nothing about what the machine was for. So an OpenShift infrastructure node was indistinguishable from an ordinary worker, and a draft proposed the two as though the choice between them did not matter. It matters in both directions, and neither is the obvious one. An infra node is the tier a cluster's own infrastructure runs on, and simplyblock storage is infrastructure. A fleet with disks in its infra nodes meant those disks to be the storage, and on OpenShift those are also the nodes that do not count against a subscription's core limit. So an infra node with disks is preferred over a worker with disks rather than avoided, and a draft now proposes it first. A control-plane node is the opposite: a data path on an etcd host is a placement almost nobody intends. It is left out unless spec.discover.enableControlPlaneNodes says otherwise, which is what a combined three-node or single-node deployment sets. There is no field beside it for infra nodes, because those need no asking for. The role is a label and not a taint, and that is the whole reason this was invisible. A taint is the cluster refusing to schedule there, which the run already honored; a label is the cluster saying what the machine is for. The two do not coincide: Kubernetes taints its control-plane nodes, so those were excluded by accident, and OpenShift usually does not taint its infrastructure ones, so those passed every check the run had. The draft carries the preference as structure rather than as a note. Each role gets a node set of its own, the infrastructure one first, so a reviewer who wants only that tier deletes a block instead of moving hostnames between them. The hardware grouping inside each set is untouched: two infra nodes with identical disks are still one group. A fleet of plain workers gets exactly the document it got before, one set named for what it is. Two exclusions that were silent now say why. A worker left out for a taint says which taint, a cordoned one says it is cordoned, and a control-plane node says what it is and which field would include it. KubeNode.Taints was recorded for exactly this and nothing had ever read it. * fix(operator): what the review found in the expansion Nine findings, and they share a shape: every one of them was a way for a document to be expanded into something it did not describe, without failing and without saying anything. The worst of them broke two-socket deployments outright. CreatingNodes reads the slot count off the cluster's spec.storageNodes, and nothing put socketsToUse or nodesPerSocket there, because ClusterTemplate had no field for either. So every cluster a document created ran one storage node per worker on socket 0, whatever the document said, and the two-socket path only worked when growing a cluster somebody had configured by hand. Both fields move onto the template, which is where its own doc comment says the layout a cluster cannot change later belongs. Three more changed behavior silently. A growth document's nodes left config.expand unset, so the control plane read them as part of an initial layout rather than as an addition to rebalance onto. An OpenShift deployment left enableKubeletConfiguration nil, which the renderer reads as skipping the kubelet configuration, inverting what every OpenShift setting this product ships does. And a create that lost a race to another actor swallowed AlreadyExists and reported ClusterCreated, which is the ClusterExists case §6 exists to refuse, arriving by a different route. Two were about what a document can express and the expansion cannot honor. Groups may each name their own interfaces, and one DaemonSet serves every node of a cluster, so the first non-empty value won and the rest were discarded; a document whose groups disagree is now a finding rather than a guess. The same shape again for a worker listed in two groups: a StorageNode is identified by its worker and its slot, so the first group creates the nodes and the second group's devices, fault group, and memory settings never reach them. The rest are smaller. The ready-to-deploy marker was an annotation, and a label selector cannot see one, so the marker was invisible to the only thing it exists for. An approved document held because the control plane is unavailable reported Draft, which the API defines as a document nobody has approved. And a pass that created a node and then failed to persist status.nodeRefs left that node out of the record permanently, because the next pass saw it as one that already existed; the record is rebuilt from the slots the document describes rather than accumulated. Eleven tests, one per finding and two for the pair that have an inverse worth pinning: a document that creates its own cluster must not ask for a rebalance onto its own initial layout. * fix(deployment): decline the initial run when nothing can be inspected A fresh install raised one discovery run unconditionally, and on a cluster with no usable worker that run could only fail. A failed run is still an object holding operatorops-finalizer, and an uninstall deletes the operator's namespace and its Deployment together, so the controller that would release the finalizer can be gone before it sees the delete. The namespace then stays Terminating. That is what hung `make undeploy` in e2e: a single-node kind cluster whose only machine is the control-plane node. The guard gains a fourth question, asked with the same predicate the run itself applies. UsableWorker is extracted for that reason rather than restated, so the bootstrap cannot drift into raising runs the run declines, and the worker loop now reads its decision from it and keeps the events for the explanation. Co-Authored-By: Claude Opus 5 (1M context) --------- Co-authored-by: Claude Opus 5 (1M context) --- ...mplyblock.io_clusterdeploymentconfigs.yaml | 24 + .../storage.simplyblock.io_operatorops.yaml | 18 + .../operator_customresources.yaml | 227 ++++---- .../v1alpha2/clusterdeploymentconfig_types.go | 21 + operator/api/v1alpha2/operatorops_types.go | 18 + .../api/v1alpha2/zz_generated.deepcopy.go | 15 + operator/cmd/main.go | 19 + ...mplyblock.io_clusterdeploymentconfigs.yaml | 24 + .../storage.simplyblock.io_operatorops.yaml | 18 + operator/dist/install.yaml | 42 ++ .../controllers/deployment/bootstrap.go | 189 ++++++ .../controllers/deployment/bootstrap_test.go | 230 ++++++++ .../clusterdeploymentconfig_controller.go | 497 ++++++++++++++++ ...clusterdeploymentconfig_controller_test.go | 422 ++++++++++++++ .../internal/controllers/deployment/events.go | 48 ++ .../controllers/deployment/expansion.go | 537 ++++++++++++++++++ .../deployment/expansion_review_test.go | 296 ++++++++++ .../deployment/operatorops_controller.go | 70 ++- .../deployment/operatorops_unit_test.go | 2 +- .../controllers/deployment/validation.go | 331 +++++++++++ operator/internal/discovery/kubenode.go | 14 + operator/internal/discovery/noderole.go | 183 ++++++ operator/internal/discovery/noderole_test.go | 155 +++++ operator/internal/discovery/plan.go | 5 +- operator/internal/discovery/rolesets.go | 95 ++++ operator/internal/discovery/rolesets_test.go | 129 +++++ ...mplyblock.io_clusterdeploymentconfigs.yaml | 24 + .../storage.simplyblock.io_operatorops.yaml | 18 + 28 files changed, 3546 insertions(+), 125 deletions(-) create mode 100644 operator/internal/controllers/deployment/bootstrap.go create mode 100644 operator/internal/controllers/deployment/bootstrap_test.go create mode 100644 operator/internal/controllers/deployment/clusterdeploymentconfig_controller.go create mode 100644 operator/internal/controllers/deployment/clusterdeploymentconfig_controller_test.go create mode 100644 operator/internal/controllers/deployment/events.go create mode 100644 operator/internal/controllers/deployment/expansion.go create mode 100644 operator/internal/controllers/deployment/expansion_review_test.go create mode 100644 operator/internal/controllers/deployment/validation.go create mode 100644 operator/internal/discovery/noderole.go create mode 100644 operator/internal/discovery/noderole_test.go create mode 100644 operator/internal/discovery/rolesets.go create mode 100644 operator/internal/discovery/rolesets_test.go diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml index 695d12254..b35497da5 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml @@ -117,6 +117,30 @@ spec: description: Name is the StorageCluster's name. maxLength: 253 type: string + nodesPerSocket: + description: |- + NodesPerSocket is how many storage nodes run per NUMA socket. See + SocketsToUse, which it multiplies. + format: int32 + maximum: 8 + minimum: 1 + type: integer + socketsToUse: + description: |- + SocketsToUse restricts the deployment to selected NUMA sockets, and empty + means socket 0 alone. With NodesPerSocket it decides how many storage nodes + each worker runs, so a group of two workers on a two-socket layout expands + to four nodes. + + It is here rather than on a node set because it is immutable on the cluster + it lands on: the layout a fleet was built with is not one a later document + can vary, and a reviewer should see it before the cluster exists. + items: + maxLength: 16 + type: string + maxItems: 16 + type: array + x-kubernetes-list-type: set stripe: description: Stripe is the erasure-coding layout. properties: diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_operatorops.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_operatorops.yaml index a4d1697e3..1a09ed465 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_operatorops.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_operatorops.yaml @@ -183,6 +183,24 @@ spec: devices and require enableLogicalBlockDevices rule: (has(self.enableLogicalBlockDevices) && self.enableLogicalBlockDevices) || !(has(self.blockAllowList) || has(self.blockDenyList)) + enableControlPlaneNodes: + description: |- + EnableControlPlaneNodes lets the run consider machines that run the API + server and etcd. + + It is off by default because a storage node is a data path, and putting one + on an etcd host is a placement almost nobody intends. The approval gate is a + poor place to catch it: a fifty-worker draft is not a document anybody reads + closely enough to spot three control-plane nodes in it. A combined three-node + or single-node deployment is the case that wants it, and those are set up + deliberately. + + There is no field beside it for infrastructure nodes, because those are used + without asking: an OpenShift infra node is the tier a cluster's own + infrastructure runs on, and simplyblock storage is infrastructure. A fleet + with disks in its infra nodes meant those disks to be the storage, so a draft + proposes them ahead of the workers rather than leaving them out. + type: boolean nodeSelector: additionalProperties: type: string diff --git a/helm-charts/charts/simplyblock-operator/operator_customresources.yaml b/helm-charts/charts/simplyblock-operator/operator_customresources.yaml index a65a3c1aa..87830ab96 100644 --- a/helm-charts/charts/simplyblock-operator/operator_customresources.yaml +++ b/helm-charts/charts/simplyblock-operator/operator_customresources.yaml @@ -1,134 +1,125 @@ -apiVersion: storage.simplyblock.io/v1alpha1 -kind: StorageCluster +# The deployment path, as objects to apply by hand. +# +# Nothing in the chart applies this file: `helm install` puts the operator in the +# cluster and this is what an administrator applies afterward, or edits into their +# own manifests. +# +# Most deployments need none of it. A fresh install raises one discovery run by +# itself, which inspects the fleet and writes a ClusterDeploymentConfig in Draft; +# reviewing that draft and setting `approved: true` on it is the whole deployment. +# The document below is what that draft looks like, for a deployment somebody +# would rather write than review, and for the fields a draft leaves for a reviewer +# to fill in. +# +# kubectl get clusterdeploymentconfigs -n simplyblock +# kubectl edit clusterdeploymentconfig -n simplyblock # set approved: true +# +# The document creates the StorageCluster and its StorageNodes. Applying a +# StorageCluster of the same name alongside it is refused with ClusterExists, +# because a config creates or adds and never reconciles a difference. +--- +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig metadata: - name: simplyblock-cluster + name: simplyblock-deployment namespace: simplyblock spec: - fabricType: tcp - # enableNodeAffinity: false - # clientDataIfname: eth1 - # nvmfBasePort: 4420 - # rpcBasePort: 9502 - # snodeApiPort: 8090 - # enableFailureDomains: false - # Cluster-wide storage-node sizing — required, and applies to every storage - # node in the cluster. maxSubsystemCount must not exceed the backend's - # MAX_SUBSYSTEMS_PER_NODE (75); vcpuCount is an explicit core count. - maxSubsystemCount: 50 - vcpuCount: 8 - # maxHugePagesSize: "100G" - stripe: - dataChunks: 1 - parityChunks: 1 - warningThreshold: - capacity: 80 - provisionedCapacity: 80 - criticalThreshold: - capacity: 90 - provisionedCapacity: 90 - hashicorpVaultSettings: - baseURL: https://vault.example.com:8200 + # A document is validated on every pass and expanded on none until this is set. + # Everything below becomes immutable the moment it is, because the document is + # then the record of what was deployed. + approved: false + + # The Kubernetes distribution. It is a shorthand the expansion spends: it sets + # the kubelet, CPU-topology, and host flags on every node the document produces, + # after which nothing reads it again. + environment: OpenShift + # edgeCluster: false + + # The cluster to create. Omit this block and set clusterRef instead to grow a + # cluster that already exists, which is how a rack is added. + # clusterRef: simplyblock-cluster + cluster: + name: simplyblock-cluster + # All three sizing values are stated once for the deployment, because the + # control plane assumes them uniform across a cluster's nodes. Two of them are + # copied onto each node so a node records the layout it was built with; + # maxSubsystemCount is read from the cluster and no node holds a copy. + maxSubsystemCount: 50 + vcpuCount: 8 + minHugePagesSize: "100G" + stripe: + dataChunks: 1 + parityChunks: 1 + fabricType: tcp + # enableFailureDomains: false + + # The nodes the cluster is made of. A node set is the organizational grouping, + # usually a rack; a group is the workers that share one configuration, so ten + # machines with the same disks are written once rather than ten times. + nodeSets: + - name: rack-a + groups: + - name: default + workers: + - worker-node-1 + - worker-node-2 + - worker-node-3 + mgmtInterface: br-ex + dataInterfaces: + - br-ex + # Every device the group's workers hand to simplyblock, named + # explicitly. There is no allow list, no model match, and no size + # range: a document whose meaning depends on what the hardware turns + # out to be is not one a reviewer can approve, so filtering happens in + # the discovery run that produces the list. + # + # A cluster is built out of one class of backend storage, so every + # group of the document names the same member: all `nvme` or all + # `block`, never a mixture. + devices: + nvme: + - "0000:01:00.0" + - "0000:02:00.0" + # block: + # - /dev/sdb + # - /dev/sdc + # The fault group these workers share, named for the rack, zone, or + # power feed rather than indexed. Required when the cluster has + # enableFailureDomains set. + # failureDomain: rack-a + # spdkSystemMemory: 4G + # journalManager: + # count: 3 + # percentPerDevice: 3 --- -apiVersion: storage.simplyblock.io/v1alpha1 +# A pool to carve volumes out of. The cluster creates a default one, so this is +# for a deployment that wants a second with its own limits. +apiVersion: storage.simplyblock.io/v1alpha2 kind: StoragePool metadata: name: simplyblock-pool namespace: simplyblock spec: - clusterName: simplyblock-cluster - # capacityLimit: 500G - # logicalVolumeMaxSize: 50G - # dhchap: false - # allowedNodes: - # - worker-node-1 - # - worker-node-2 - # - worker-node-3 - # qos: + clusterRef: simplyblock-cluster + # The ceilings the pool as a whole is held to. Mutable: raising a pool's + # capacity is an ordinary operation. + # limits: + # capacity: 500G + # maxVolumeSize: 50G # iops: 10000 # throughput: # read: 500 # write: 250 # readWrite: 750 - # storageClassParameters: - # encryption: false - # fabric: tcp + # + # What every volume in the pool is created with. Immutable once set, because a + # StorageClass's parameters are: changing them means creating a new pool. + # volumeDefaults: # filesystem: ext4 - # maxNamespacePerSubsys: "1" - # qosRMbytes: "0" - # qosWMbytes: "0" - # qosRwMbytes: "0" - # qosRwIops: "0" - # tune2fsReservedBlocks: "0" - ---- -apiVersion: storage.simplyblock.io/v1alpha1 -kind: StorageNodeSet -metadata: - name: simplyblock-node - namespace: simplyblock -spec: - clusterName: simplyblock-cluster - # clusterImage: "quay.io/simplyblock-io/simplyblock:26.2.2-PRE" - # spdkImage: "quay.io/simplyblock-io/spdk:26.2.2-PRE" - # spdkProxyImage: "quay.io/simplyblock-io/simplyblock:26.2.2-PRE" - #spdkSystemMemory: 4G - mgmtIfname: br-ex - dataIfname: - - br-ex - workerNodes: [] - # - worker-node-1 - # - worker-node-2 - # - worker-node-3 - # enableJournalDevice: false - nodesPerSocket: 1 - reservedSystemCPU: "0,1" - socketsToUse: - - "0" - journalManager: - count: 3 - percentPerDevice: 3 - # deviceNames: - # - nvme0n1 - # - nvme1n1 - # pcieAllowList: - # - "0000:01:00.0" - # pcieDenyList: - # - "0000:02:00.0" - # pcieModel: "INTEL SSDPE2KX010T8" - # driveSizeRange: "100G-2T" - skipKubeletConfiguration: false - enableCpuTopology: true - openShiftCluster: true - ubuntuHost: false - forceFormat4K: true - # tolerations: - # - key: node-role.kubernetes.io/storage - # operator: Equal - # value: "true" - # effect: NoSchedule - # openShiftMachineConfigPool: "" - # maxParallelNodeAdds: 1 - # expand: false - # imagePullPolicy: IfNotPresent - # containerResources: - # requests: - # cpu: "500m" - # memory: "512Mi" - # limits: - # cpu: "2" - # memory: "2Gi" - # initContainerResources: - # requests: - # cpu: "100m" - # memory: "128Mi" - # limits: - # cpu: "500m" - # memory: "512Mi" - # nodeFailureDomains: - # worker-node-1: 0 - # worker-node-2: 0 - # worker-node-3: 1 - # nodeConfigs: - # worker-node-1: - # failureDomain: 0 + # fabric: tcp + # enableEncryption: false + # enableDHCHAP: false + # + # allowedNodes: + # - worker-node-1 diff --git a/operator/api/v1alpha2/clusterdeploymentconfig_types.go b/operator/api/v1alpha2/clusterdeploymentconfig_types.go index edaea6046..d6a594d91 100644 --- a/operator/api/v1alpha2/clusterdeploymentconfig_types.go +++ b/operator/api/v1alpha2/clusterdeploymentconfig_types.go @@ -226,6 +226,27 @@ type ClusterTemplate struct { // +optional MinHugePagesSize string `json:"minHugePagesSize,omitempty"` + // SocketsToUse restricts the deployment to selected NUMA sockets, and empty + // means socket 0 alone. With NodesPerSocket it decides how many storage nodes + // each worker runs, so a group of two workers on a two-socket layout expands + // to four nodes. + // + // It is here rather than on a node set because it is immutable on the cluster + // it lands on: the layout a fleet was built with is not one a later document + // can vary, and a reviewer should see it before the cluster exists. + // +kubebuilder:validation:items:MaxLength=16 + // +kubebuilder:validation:MaxItems=16 + // +listType=set + // +optional + SocketsToUse []string `json:"socketsToUse,omitempty"` + + // NodesPerSocket is how many storage nodes run per NUMA socket. See + // SocketsToUse, which it multiplies. + // +kubebuilder:validation:Minimum=1 + // +kubebuilder:validation:Maximum=8 + // +optional + NodesPerSocket *int32 `json:"nodesPerSocket,omitempty"` + // Stripe is the erasure-coding layout. // +optional Stripe *StripeSpec `json:"stripe,omitempty"` diff --git a/operator/api/v1alpha2/operatorops_types.go b/operator/api/v1alpha2/operatorops_types.go index 16596e609..907cc9d43 100644 --- a/operator/api/v1alpha2/operatorops_types.go +++ b/operator/api/v1alpha2/operatorops_types.go @@ -149,6 +149,24 @@ type DiscoverSpec struct { // +optional NodeSelector map[string]string `json:"nodeSelector,omitempty"` + // EnableControlPlaneNodes lets the run consider machines that run the API + // server and etcd. + // + // It is off by default because a storage node is a data path, and putting one + // on an etcd host is a placement almost nobody intends. The approval gate is a + // poor place to catch it: a fifty-worker draft is not a document anybody reads + // closely enough to spot three control-plane nodes in it. A combined three-node + // or single-node deployment is the case that wants it, and those are set up + // deliberately. + // + // There is no field beside it for infrastructure nodes, because those are used + // without asking: an OpenShift infra node is the tier a cluster's own + // infrastructure runs on, and simplyblock storage is infrastructure. A fleet + // with disks in its infra nodes meant those disks to be the storage, so a draft + // proposes them ahead of the workers rather than leaving them out. + // +optional + EnableControlPlaneNodes *bool `json:"enableControlPlaneNodes,omitempty"` + // DeviceFilter narrows which of an inspected worker's devices reach the // draft. Empty reports every device the worker advertises, including the one // it boots from, which is what the approval gate then has to catch. diff --git a/operator/api/v1alpha2/zz_generated.deepcopy.go b/operator/api/v1alpha2/zz_generated.deepcopy.go index 6ce441753..a9f61b713 100644 --- a/operator/api/v1alpha2/zz_generated.deepcopy.go +++ b/operator/api/v1alpha2/zz_generated.deepcopy.go @@ -284,6 +284,16 @@ func (in *ClusterTemplate) DeepCopyInto(out *ClusterTemplate) { *out = new(int32) **out = **in } + if in.SocketsToUse != nil { + in, out := &in.SocketsToUse, &out.SocketsToUse + *out = make([]string, len(*in)) + copy(*out, *in) + } + if in.NodesPerSocket != nil { + in, out := &in.NodesPerSocket, &out.NodesPerSocket + *out = new(int32) + **out = **in + } if in.Stripe != nil { in, out := &in.Stripe, &out.Stripe *out = new(StripeSpec) @@ -564,6 +574,11 @@ func (in *DiscoverSpec) DeepCopyInto(out *DiscoverSpec) { (*out)[key] = val } } + if in.EnableControlPlaneNodes != nil { + in, out := &in.EnableControlPlaneNodes, &out.EnableControlPlaneNodes + *out = new(bool) + **out = **in + } if in.DeviceFilter != nil { in, out := &in.DeviceFilter, &out.DeviceFilter *out = new(DeviceFilter) diff --git a/operator/cmd/main.go b/operator/cmd/main.go index 469ff09f1..8a65d8275 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -664,6 +664,25 @@ func main() { // image is read from the environment rather than from the running pod, // because a pod may name its image by a tag the registry has since moved and // what a Job needs is the reference the operator was deployed with. + // A fresh install raises one discovery run by itself, so an administrator + // finds a draft of what the fleet has rather than an empty namespace. It is + // declined the moment anything already exists (design-clusterdeploymentconfig.md §8). + if err := (&deployment.InitialDiscovery{ + Client: mgr.GetClient(), + Namespace: operatorNamespace, + }).SetupWithManager(mgr); err != nil { + setupLog.Error(err, "unable to add the initial discovery check") + os.Exit(1) + } + if err := (&deployment.ClusterDeploymentConfigReconciler{ + Client: mgr.GetClient(), + Scheme: mgr.GetScheme(), + Recorder: mgr.GetEventRecorder("clusterdeploymentconfig-controller"), + Namespace: operatorNamespace, + }).SetupWithManager(mgr); err != nil { + setupLog.Error(err, "unable to create controller", "controller", "ClusterDeploymentConfig") + os.Exit(1) + } if err := (&deployment.OperatorOpsReconciler{ Client: mgr.GetClient(), Scheme: mgr.GetScheme(), diff --git a/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml b/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml index 695d12254..b35497da5 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml @@ -117,6 +117,30 @@ spec: description: Name is the StorageCluster's name. maxLength: 253 type: string + nodesPerSocket: + description: |- + NodesPerSocket is how many storage nodes run per NUMA socket. See + SocketsToUse, which it multiplies. + format: int32 + maximum: 8 + minimum: 1 + type: integer + socketsToUse: + description: |- + SocketsToUse restricts the deployment to selected NUMA sockets, and empty + means socket 0 alone. With NodesPerSocket it decides how many storage nodes + each worker runs, so a group of two workers on a two-socket layout expands + to four nodes. + + It is here rather than on a node set because it is immutable on the cluster + it lands on: the layout a fleet was built with is not one a later document + can vary, and a reviewer should see it before the cluster exists. + items: + maxLength: 16 + type: string + maxItems: 16 + type: array + x-kubernetes-list-type: set stripe: description: Stripe is the erasure-coding layout. properties: diff --git a/operator/config/crd/bases/storage.simplyblock.io_operatorops.yaml b/operator/config/crd/bases/storage.simplyblock.io_operatorops.yaml index a4d1697e3..1a09ed465 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_operatorops.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_operatorops.yaml @@ -183,6 +183,24 @@ spec: devices and require enableLogicalBlockDevices rule: (has(self.enableLogicalBlockDevices) && self.enableLogicalBlockDevices) || !(has(self.blockAllowList) || has(self.blockDenyList)) + enableControlPlaneNodes: + description: |- + EnableControlPlaneNodes lets the run consider machines that run the API + server and etcd. + + It is off by default because a storage node is a data path, and putting one + on an etcd host is a placement almost nobody intends. The approval gate is a + poor place to catch it: a fifty-worker draft is not a document anybody reads + closely enough to spot three control-plane nodes in it. A combined three-node + or single-node deployment is the case that wants it, and those are set up + deliberately. + + There is no field beside it for infrastructure nodes, because those are used + without asking: an OpenShift infra node is the tier a cluster's own + infrastructure runs on, and simplyblock storage is infrastructure. A fleet + with disks in its infra nodes meant those disks to be the storage, so a draft + proposes them ahead of the workers rather than leaving them out. + type: boolean nodeSelector: additionalProperties: type: string diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index 0f8a6c159..96b4b821a 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -790,6 +790,30 @@ spec: description: Name is the StorageCluster's name. maxLength: 253 type: string + nodesPerSocket: + description: |- + NodesPerSocket is how many storage nodes run per NUMA socket. See + SocketsToUse, which it multiplies. + format: int32 + maximum: 8 + minimum: 1 + type: integer + socketsToUse: + description: |- + SocketsToUse restricts the deployment to selected NUMA sockets, and empty + means socket 0 alone. With NodesPerSocket it decides how many storage nodes + each worker runs, so a group of two workers on a two-socket layout expands + to four nodes. + + It is here rather than on a node set because it is immutable on the cluster + it lands on: the layout a fleet was built with is not one a later document + can vary, and a reviewer should see it before the cluster exists. + items: + maxLength: 16 + type: string + maxItems: 16 + type: array + x-kubernetes-list-type: set stripe: description: Stripe is the erasure-coding layout. properties: @@ -1450,6 +1474,24 @@ spec: devices and require enableLogicalBlockDevices rule: (has(self.enableLogicalBlockDevices) && self.enableLogicalBlockDevices) || !(has(self.blockAllowList) || has(self.blockDenyList)) + enableControlPlaneNodes: + description: |- + EnableControlPlaneNodes lets the run consider machines that run the API + server and etcd. + + It is off by default because a storage node is a data path, and putting one + on an etcd host is a placement almost nobody intends. The approval gate is a + poor place to catch it: a fifty-worker draft is not a document anybody reads + closely enough to spot three control-plane nodes in it. A combined three-node + or single-node deployment is the case that wants it, and those are set up + deliberately. + + There is no field beside it for infrastructure nodes, because those are used + without asking: an OpenShift infra node is the tier a cluster's own + infrastructure runs on, and simplyblock storage is infrastructure. A fleet + with disks in its infra nodes meant those disks to be the storage, so a draft + proposes them ahead of the workers rather than leaving them out. + type: boolean nodeSelector: additionalProperties: type: string diff --git a/operator/internal/controllers/deployment/bootstrap.go b/operator/internal/controllers/deployment/bootstrap.go new file mode 100644 index 000000000..fcc0cfefc --- /dev/null +++ b/operator/internal/controllers/deployment/bootstrap.go @@ -0,0 +1,189 @@ +// The one discovery run a fresh install performs by itself. +// +// An operator installed into a cluster with disks in it can tell the +// administrator what it found, and that is a better first experience than an +// empty namespace and a document to write by hand. So a fresh install raises one +// OperatorOps with action Discover, whose output is a ClusterDeploymentConfig in +// Draft that nobody has approved and nothing acts on (§8.3). +// +// The whole of the design is the guard. A discovery run is read-only against the +// control plane and the draft it writes is inert, but Probing creates one Job per +// worker, so a run that fired on every restart would put a Job on every node of +// the fleet every time the operator was upgraded. The guard is therefore that no +// previous result exists and there is something to look at, which is four +// questions rather than one: +// +// - Has any OperatorOps ever run? A terminal one is a previous result, and so +// is a failed one: the administrator has seen the answer and either acted on +// it or chose not to, and either way this operator is not the thing to decide +// they want another. +// - Does any ClusterDeploymentConfig exist? A document is a previous run's +// output or somebody's hand-written deployment, and either way discovery has +// nothing to add that they did not already have. +// - Does any StorageCluster exist? A deployed fleet was bootstrapped by +// something, and a fresh install is the only state this exists for. +// - Is there a machine a run would inspect at all? This one is not about +// previous results, and it is the question with a cost behind it rather than +// a convenience. A run raised against a cluster with no usable worker fails, +// and a failed run is still an object holding a finalizer. Uninstalling the +// operator deletes its namespace and its Deployment together, so the +// controller that would release that finalizer can be gone before it sees the +// delete, and the namespace stays Terminating. An install that raises a run +// it already knows cannot succeed has made its own uninstall conditional on +// timing, which is the single-node development cluster exactly: its one +// machine is the control-plane node. +// +// Any one of the four is enough to decline. Together, they mean the run happens +// on the install that has nothing and has something to look at, and never again. +// +// It runs under leader election, so one replica asks. The create is idempotent by +// name as well, because two replicas answering the same questions at once +// would both conclude yes. + +package deployment + +import ( + "context" + "fmt" + + corev1 "k8s.io/api/core/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + logf "sigs.k8s.io/controller-runtime/pkg/log" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// InitialDiscoveryName is what the run is called. It is fixed rather than +// generated so that the create is idempotent by name on top of the guard below, +// and so that an administrator who does not want it can say so by writing an +// object with that name and deleting nothing. +const InitialDiscoveryName = "initial-discovery" + +// InitialDiscovery raises one Discover run on an install that has nothing. +// +// It is a Runnable rather than a reconciler because it is not reconciling +// anything: there is no object whose desired state it converges toward, and the +// question it asks has one answer per installation. +type InitialDiscovery struct { + client.Client + + // Namespace is where the operator runs, and where the run and its draft land. + // A document describes one deployment, and the operator's own namespace is + // the one place a deployment that does not exist yet can be described from. + Namespace string +} + +// NeedLeaderElection makes one replica ask the question. The create is idempotent +// without it, and the three reads are not: two replicas racing would both see an +// empty cluster and both decide to run. +func (d *InitialDiscovery) NeedLeaderElection() bool { return true } + +// Start performs the check once and returns. It is not a loop: an install that +// already has something is an install that will keep having it, and an install +// that has nothing gets its run on this pass. +func (d *InitialDiscovery) Start(ctx context.Context) error { + log := logf.FromContext(ctx).WithName("initial-discovery") + + reason, err := d.declineReason(ctx) + if err != nil { + // A read that failed is not evidence of an empty cluster. Declining is the + // conservative answer: the cost of not running is an administrator writing + // a document by hand, and the cost of running against a fleet that is + // already deployed is a Job on every node of it. + log.Error(err, "the initial discovery run is declined; the cluster could not be read") + return nil + } + if reason != "" { + log.V(1).Info("the initial discovery run is not needed", "reason", reason) + return nil + } + + run := &simplyblockv1alpha2.OperatorOps{ + ObjectMeta: metav1.ObjectMeta{ + Name: InitialDiscoveryName, + Namespace: d.Namespace, + }, + Spec: simplyblockv1alpha2.OperatorOpsSpec{ + Action: simplyblockv1alpha2.OperatorOpsActionDiscover, + // The run states no filter and no selector. What it produces is a + // draft of everything the fleet has, which is what a reviewer + // narrows: a guess at which disks somebody meant would be a guess + // they then have to find and undo (§8.1). + Discover: &simplyblockv1alpha2.DiscoverSpec{}, + }, + } + if err := d.Create(ctx, run); err != nil { + if apierrors.IsAlreadyExists(err) { + return nil + } + // Failing here would crash the manager over a convenience. The operator + // works without the run; what is lost is the draft an administrator would + // otherwise have found waiting. + log.Error(err, "the initial discovery run could not be created") + return nil + } + + log.Info("raised the initial discovery run; its draft is what to review", + "operatorOps", InitialDiscoveryName, "namespace", d.Namespace) + return nil +} + +// declineReason answers whether anything already exists, and says which thing. An +// empty string means the install has nothing and the run is worth raising. +func (d *InitialDiscovery) declineReason(ctx context.Context) (string, error) { + var runs simplyblockv1alpha2.OperatorOpsList + if err := d.List(ctx, &runs); err != nil { + return "", fmt.Errorf("listing operator operations: %w", err) + } + if len(runs.Items) > 0 { + return fmt.Sprintf("%d operator operation(s) have already run", len(runs.Items)), nil + } + + var configs simplyblockv1alpha2.ClusterDeploymentConfigList + if err := d.List(ctx, &configs); err != nil { + return "", fmt.Errorf("listing deployment configs: %w", err) + } + if len(configs.Items) > 0 { + return fmt.Sprintf("%d deployment config(s) already exist", len(configs.Items)), nil + } + + var clusters simplyblockv1alpha2.StorageClusterList + if err := d.List(ctx, &clusters); err != nil { + return "", fmt.Errorf("listing storage clusters: %w", err) + } + if len(clusters.Items) > 0 { + return fmt.Sprintf("%d storage cluster(s) are already deployed", len(clusters.Items)), nil + } + + // The run this would raise states no filter and no selector, so it asks about + // every machine, and UsableWorker is the same predicate it would then apply. + // Reading it here rather than restating the conditions is what keeps the two + // from drifting into a bootstrap that raises runs the run itself declines. + var nodes corev1.NodeList + if err := d.List(ctx, &nodes); err != nil { + return "", fmt.Errorf("listing nodes: %w", err) + } + usable := 0 + for _, node := range nodes.Items { + // The opt-in is a decision an administrator makes on a run they wrote. + // This one writes no spec, so it asks the question the run it would raise + // asks: false, and a control-plane node does not count. + if UsableWorker(node, false) { + usable++ + } + } + if usable == 0 { + return fmt.Sprintf("none of the %d node(s) hold storage without being asked to", + len(nodes.Items)), nil + } + + return "", nil +} + +// SetupWithManager adds the check to the manager. +func (d *InitialDiscovery) SetupWithManager(mgr ctrl.Manager) error { + return mgr.Add(d) +} diff --git a/operator/internal/controllers/deployment/bootstrap_test.go b/operator/internal/controllers/deployment/bootstrap_test.go new file mode 100644 index 000000000..78ed4e86d --- /dev/null +++ b/operator/internal/controllers/deployment/bootstrap_test.go @@ -0,0 +1,230 @@ +// What a fresh install does by itself, and the three things that stop it. +// +// The guard is the whole design here, so every case below is one of its +// questions. Probing creates a Job on every worker, so a run that fired on an +// install that already had something would put a Job on every node of a deployed +// fleet each time the operator was upgraded. + +package deployment + +import ( + "context" + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/testsupport" +) + +func discoveryFor(t *testing.T, objects ...client.Object) *InitialDiscovery { + t.Helper() + scheme := testsupport.NewScheme(t, corev1.AddToScheme) + return &InitialDiscovery{ + Client: fake.NewClientBuilder().WithScheme(scheme).WithObjects(objects...).Build(), + Namespace: theNamespace, + } +} + +// raised reports whether the run exists, which is what every case here asserts. +func raised(t *testing.T, d *InitialDiscovery) bool { + t.Helper() + var run simplyblockv1alpha2.OperatorOps + key := client.ObjectKey{Namespace: theNamespace, Name: InitialDiscoveryName} + err := d.Get(context.Background(), key, &run) + return err == nil +} + +// An install with nothing in it is the one state this exists for. +func TestAFreshInstallRaisesOneDiscoveryRun(t *testing.T) { + d := discoveryFor(t, &corev1.Node{ObjectMeta: metav1.ObjectMeta{Name: "worker-1"}}) + + if err := d.Start(context.Background()); err != nil { + t.Fatalf("Start: %v", err) + } + if !raised(t, d) { + t.Fatal("a fresh install raised no discovery run") + } + + var run simplyblockv1alpha2.OperatorOps + key := client.ObjectKey{Namespace: theNamespace, Name: InitialDiscoveryName} + if err := d.Get(context.Background(), key, &run); err != nil { + t.Fatalf("reading the run: %v", err) + } + if run.Spec.Action != simplyblockv1alpha2.OperatorOpsActionDiscover { + t.Errorf("action = %q, want Discover", run.Spec.Action) + } + // The run states no filter and no selector: what it produces is a draft of + // everything the fleet has, which is what a reviewer narrows. + if run.Spec.Discover == nil { + t.Error("the run carries no discover block") + } else if len(run.Spec.Discover.NodeSelector) != 0 || + run.Spec.Discover.DeviceFilter != nil { + t.Errorf("the run guessed at a filter: %+v", run.Spec.Discover) + } +} + +// Each of the three questions is enough on its own to decline. +func TestAPreviousResultDeclinesTheRun(t *testing.T) { + for _, tc := range []struct { + name string + existing client.Object + }{ + { + // A terminal run is a previous result, and so is a failed one: the + // administrator has seen the answer either way. + name: "an operator operation has already run", + existing: &simplyblockv1alpha2.OperatorOps{ + ObjectMeta: metav1.ObjectMeta{Name: "earlier", Namespace: theNamespace}, + Spec: simplyblockv1alpha2.OperatorOpsSpec{ + Action: simplyblockv1alpha2.OperatorOpsActionDiscover, + }, + Status: simplyblockv1alpha2.OperatorOpsStatus{ + Phase: simplyblockv1alpha2.OperatorOpsPhaseFailed, + }, + }, + }, + { + name: "a deployment config already exists", + existing: &simplyblockv1alpha2.ClusterDeploymentConfig{ + ObjectMeta: metav1.ObjectMeta{Name: "written-by-hand", Namespace: theNamespace}, + }, + }, + { + name: "a cluster is already deployed", + existing: aCluster(nil), + }, + } { + t.Run(tc.name, func(t *testing.T) { + d := discoveryFor(t, tc.existing) + + if err := d.Start(context.Background()); err != nil { + t.Fatalf("Start: %v", err) + } + if raised(t, d) { + t.Errorf("a run was raised even though %s", tc.name) + } + }) + } +} + +// The check runs on every operator start, and a restart must not raise a second +// run against a fleet the first one already reported on. +func TestARestartRaisesNothingFurther(t *testing.T) { + d := discoveryFor(t, &corev1.Node{ObjectMeta: metav1.ObjectMeta{Name: "worker-1"}}) + + for pass := 0; pass < 3; pass++ { + if err := d.Start(context.Background()); err != nil { + t.Fatalf("pass %d: %v", pass, err) + } + } + + var runs simplyblockv1alpha2.OperatorOpsList + if err := d.List(context.Background(), &runs); err != nil { + t.Fatalf("listing the runs: %v", err) + } + if len(runs.Items) != 1 { + t.Fatalf("three starts produced %d runs, want the one", len(runs.Items)) + } +} + +// An administrator who does not want the run says so by writing an object with +// that name, which the guard then finds and declines behind. +func TestAnObjectByThatNameIsNotReplaced(t *testing.T) { + theirs := &simplyblockv1alpha2.OperatorOps{ + ObjectMeta: metav1.ObjectMeta{ + Name: InitialDiscoveryName, + Namespace: theNamespace, + Annotations: map[string]string{"theirs": "true"}, + }, + Spec: simplyblockv1alpha2.OperatorOpsSpec{ + Action: simplyblockv1alpha2.OperatorOpsActionDiscover, + }, + } + d := discoveryFor(t, theirs) + + if err := d.Start(context.Background()); err != nil { + t.Fatalf("Start: %v", err) + } + + var run simplyblockv1alpha2.OperatorOps + key := client.ObjectKey{Namespace: theNamespace, Name: InitialDiscoveryName} + if err := d.Get(context.Background(), key, &run); err != nil { + t.Fatalf("reading the run: %v", err) + } + if run.Annotations["theirs"] != "true" { + t.Error("the administrator's own object was overwritten") + } +} + +// One replica asks, because the three reads are not idempotent the way the create +// is: two replicas racing would both see an empty cluster and both decide to run. +func TestTheCheckIsLeaderElected(t *testing.T) { + if !(&InitialDiscovery{}).NeedLeaderElection() { + t.Error("the check runs on every replica, so the reads race") + } +} + +// A cluster with nothing a discovery run would inspect is the fourth question, +// and it is the one with teeth beyond convenience. +// +// A run raised against such a cluster fails, and a failed run is still an object +// carrying a finalizer. Uninstalling the operator deletes its namespace and its +// Deployment together, so the controller that would release that finalizer can be +// gone before it sees the delete, and the namespace stays Terminating. An install +// that raises a run it knows cannot succeed has therefore made its own uninstall +// conditional on timing. +// +// The single-node development cluster is exactly this shape: its one machine is +// the control-plane node, which nothing places storage on unless somebody asks. +func TestAClusterWithNothingToInspectRaisesNoRun(t *testing.T) { + controlPlane := &corev1.Node{ObjectMeta: metav1.ObjectMeta{ + Name: "kind-control-plane", + Labels: map[string]string{"node-role.kubernetes.io/control-plane": ""}, + }} + d := discoveryFor(t, controlPlane) + + if err := d.Start(context.Background()); err != nil { + t.Fatalf("Start: %v", err) + } + if raised(t, d) { + t.Error("a run was raised against a cluster whose only node holds no storage") + } +} + +// The same cluster once somebody has asked for its control-plane node is not that +// case: there is a machine to inspect, so the run is worth raising. +// +// The bootstrap does not set the opt-in, so this asserts the reason rather than +// the outcome: the question the guard asks is whether a machine exists at all, +// and it must not answer it by reading a flag nobody set. +func TestAWorkerIsEnoughToRaiseTheRun(t *testing.T) { + d := discoveryFor(t, &corev1.Node{ObjectMeta: metav1.ObjectMeta{Name: "worker-1"}}) + + if err := d.Start(context.Background()); err != nil { + t.Fatalf("Start: %v", err) + } + if !raised(t, d) { + t.Error("a cluster with a worker in it raised no run") + } +} + +// An unschedulable machine is not one a run would inspect either, so a fleet +// cordoned for maintenance is the empty case rather than the worker case. +func TestACordonedFleetRaisesNoRun(t *testing.T) { + cordoned := &corev1.Node{ + ObjectMeta: metav1.ObjectMeta{Name: "worker-1"}, + Spec: corev1.NodeSpec{Unschedulable: true}, + } + d := discoveryFor(t, cordoned) + + if err := d.Start(context.Background()); err != nil { + t.Fatalf("Start: %v", err) + } + if raised(t, d) { + t.Error("a run was raised against a fleet with no schedulable machine") + } +} diff --git a/operator/internal/controllers/deployment/clusterdeploymentconfig_controller.go b/operator/internal/controllers/deployment/clusterdeploymentconfig_controller.go new file mode 100644 index 000000000..71bfaddb2 --- /dev/null +++ b/operator/internal/controllers/deployment/clusterdeploymentconfig_controller.go @@ -0,0 +1,497 @@ +// The ClusterDeploymentConfig reconciler: it validates a draft on every pass and +// expands an approved one into a StorageCluster and its StorageNodes. +// +// The document is ephemeral, and everything below follows from that. It owns +// nothing, nothing references it, and nothing reads it after the expansion, so +// there is no finalizer and no cleanup: deleting it deletes a document. It +// specifically does not own the StorageCluster it created, because an owner +// reference would make deleting the document delete the cluster and every volume +// in it. +// +// A draft is validated on every reconcile and expanded on none. Validation writes +// what it found into status.message and nothing else, so a reviewer sees the +// problems before approving rather than after — which is the whole value of the +// gate, and why admission lets a draft naming a missing worker be saved at all. +// +// The expansion is create-only (§6). It creates a cluster or adds nodes to one, +// and it never reconciles a difference: a document that names an existing cluster +// without asking to is a failure with a reason, because the differences that +// matter here are of the form: this node's device list changed, and its only +// correct handling is not to apply it to a node that already has data on those +// devices. +// +// design-clusterdeploymentconfig.md §4 is the specification. + +package deployment + +import ( + "context" + "errors" + "fmt" + "time" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/runtime" + "k8s.io/apimachinery/pkg/types" + "k8s.io/client-go/tools/events" + "k8s.io/client-go/util/retry" + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/handler" + logf "sigs.k8s.io/controller-runtime/pkg/log" + "sigs.k8s.io/controller-runtime/pkg/reconcile" + + "github.com/simplyblock/atlas/statemachine" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +const ( + // readyToDeploy marks an approved document so a selector can find one. The + // operator writes it and reads it from nowhere: it is an output (§5). + // + // It is a label rather than an annotation, and that is the whole point of it. + // A label selector cannot see an annotation, so the marker would have been + // invisible to the one thing it exists for. + readyToDeploy = "storage.simplyblock.io/ready-to-deploy" + + // readyToDeployValue is what the marker carries. An annotation value is a + // string, and this is the one a selector matches on. + readyToDeployValue = "true" + + // configRetry is how long a held document waits before looking again at + // something it cannot hurry: a control plane that is not ready, or a worker + // somebody has yet to add. + configRetry = 30 * time.Second + + // configAdvance is how long a pass that moved the machine forward waits. The + // status write this pass made is itself a change the controller watches, so + // this is the backstop for the event rather than the path the next step + // normally arrives on. + configAdvance = time.Second +) + +// ClusterDeploymentConfigReconciler reconciles a ClusterDeploymentConfig. +type ClusterDeploymentConfigReconciler struct { + client.Client + Scheme *runtime.Scheme + Recorder events.EventRecorder + + // Namespace is where the operator runs, which is where the ControlPlane + // singleton the expansion waits on lives. + Namespace string +} + +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=clusterdeploymentconfigs,verbs=get;list;watch;update;patch +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=clusterdeploymentconfigs/status,verbs=get;update;patch +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storageclusters,verbs=get;list;watch;create +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodes,verbs=get;list;watch;create +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=controlplanes,verbs=get;list;watch +// +kubebuilder:rbac:groups="",resources=nodes,verbs=get;list;watch + +// SetupWithManager registers the controller. +// +// It watches Kubernetes Nodes as well as its own kind, because a draft held by a +// worker that does not exist becomes valid the moment somebody adds it, and +// waiting out a requeue interval to notice would make the gate feel broken. +func (r *ClusterDeploymentConfigReconciler) SetupWithManager(mgr ctrl.Manager) error { + return ctrl.NewControllerManagedBy(mgr). + For(&simplyblockv1alpha2.ClusterDeploymentConfig{}). + Named("clusterdeploymentconfig"). + Watches(&corev1.Node{}, handler.EnqueueRequestsFromMapFunc(r.everyUnfinishedDocument)). + Complete(r) +} + +func (r *ClusterDeploymentConfigReconciler) Reconcile( + ctx context.Context, req ctrl.Request, +) (ctrl.Result, error) { + var config simplyblockv1alpha2.ClusterDeploymentConfig + if err := r.Get(ctx, req.NamespacedName, &config); err != nil { + return ctrl.Result{}, client.IgnoreNotFound(err) + } + + // A deleted document needs nothing done to it. There is no finalizer, because + // it owns nothing and nothing reads it (§4.3). + if !config.DeletionTimestamp.IsZero() { + return ctrl.Result{}, nil + } + + // Expanded and Failed are terminal. A document is the record of one + // deployment action, so re-running it is writing another document rather than + // editing this one. + if terminalConfig(config.Status.Phase) { + return ctrl.Result{}, nil + } + + findings, err := r.validate(ctx, &config) + if err != nil { + return ctrl.Result{RequeueAfter: configRetry}, err + } + + if !config.Spec.Approved { + return r.holdAsDraft(ctx, &config, findings) + } + + // An approved document that does not validate is a failure rather than a + // hold. Approval is what makes it immutable, so a problem found after it can + // no longer be edited away, and holding forever would say less than failing. + if len(findings) > 0 { + r.emitFindings(&config, findings) + return r.fail(ctx, &config, findings[0].message) + } + + if err := r.markReadyToDeploy(ctx, &config); err != nil { + return ctrl.Result{}, err + } + + if ready, reason := r.controlPlaneReady(ctx); !ready { + // Expanding rather than Draft. The document is approved, and Draft is + // what the API calls one that is not: reporting it would tell every + // status consumer the deployment is still editable and has not started. + r.emit(&config, corev1.EventTypeWarning, ControlPlaneNotReady, reason) + return ctrl.Result{RequeueAfter: configRetry}, r.note(ctx, &config, + simplyblockv1alpha2.ClusterDeploymentConfigPhaseExpanding, reason) + } + + return r.expand(ctx, &config) +} + +// expand runs the machine of §4.2 forward by at most one step. +func (r *ClusterDeploymentConfigReconciler) expand( + ctx context.Context, config *simplyblockv1alpha2.ClusterDeploymentConfig, +) (ctrl.Result, error) { + machine, err := statemachine.NewFromSnapshot(ctx, expansionGraph(), + statemachine.FromKube[configStep](config.Status.Step)) + if err != nil { + // An unrecognized step is a downgrade, a hand-edited object, or a rename + // that shipped without a conversion, and none of them resolve by + // reconciling again. + return r.fail(ctx, config, fmt.Sprintf("the expansion cannot be resumed: %v", err)) + } + defer machine.Close() + + // A machine is born already in its initial state, so that state's entry hook + // never runs and no deadline is set for it. Setting one on the first pass is + // what stops the first step being the one step that cannot time out. + if config.Status.Step.State == "" { + deadline := metav1.NewTime(time.Now().Add(validatingDeadline)) + return ctrl.Result{RequeueAfter: configAdvance}, + r.recordStep(ctx, config, machine.CurrentState(), &deadline) + } + + current := machine.CurrentState() + if machine.TimeoutReached() { + r.emit(config, corev1.EventTypeWarning, StepDeadlineExceeded, + fmt.Sprintf("Step %s outlived its deadline", current)) + return r.fail(ctx, config, fmt.Sprintf("step %s outlived its deadline", current)) + } + + done, err := r.performStep(ctx, config, current) + if err != nil { + var refusal *refusedError + if errors.As(err, &refusal) { + r.emit(config, corev1.EventTypeWarning, refusal.reason, refusal.Error()) + return r.fail(ctx, config, refusal.Error()) + } + logf.FromContext(ctx).Error(err, "the expansion step could not be advanced", + "config", config.Name, "step", current) + return ctrl.Result{RequeueAfter: configRetry}, r.note(ctx, config, + simplyblockv1alpha2.ClusterDeploymentConfigPhaseExpanding, err.Error()) + } + if !done { + return ctrl.Result{RequeueAfter: configRetry}, r.note(ctx, config, + simplyblockv1alpha2.ClusterDeploymentConfigPhaseExpanding, + fmt.Sprintf("waiting on %s", current)) + } + + if machine.IsTerminal() { + // The expansion does not wait for the nodes to come up. It created the + // objects and is finished; provisioning them is the node controller's and + // is bounded by maxParallelNodeAdds, and a document that stayed Expanding + // until a twenty-node fleet was online would be reporting the fleet's + // progress rather than its own (§4.2). + return ctrl.Result{}, r.succeed(ctx, config) + } + + next, err := nextStep(machine) + if err != nil { + return r.fail(ctx, config, err.Error()) + } + if err := machine.TransitionTo(ctx, next); err != nil { + return ctrl.Result{}, fmt.Errorf("enter step %s: %w", next, err) + } + snapshot := statemachine.ToKube(machine.Snapshot()) + return ctrl.Result{RequeueAfter: configAdvance}, + r.recordStep(ctx, config, next, snapshot.Deadline) +} + +// performStep runs one step and reports whether it has finished. +func (r *ClusterDeploymentConfigReconciler) performStep( + ctx context.Context, + config *simplyblockv1alpha2.ClusterDeploymentConfig, + current configStep, +) (bool, error) { + switch current { + case stepValidating: + // Validation already ran on the way in, and reaching here means it found + // nothing. The step exists so that the machine's first position is the + // check rather than a side effect. + return true, nil + case stepCreatingCluster: + return r.createCluster(ctx, config) + case stepAwaitingCluster: + return r.awaitCluster(ctx, config) + case stepCreatingNodes: + return r.createNodes(ctx, config) + default: + return false, fmt.Errorf("step %s belongs to no expansion this operator runs", current) + } +} + +// holdAsDraft reports what validation found and leaves the document alone. +// +// A draft is expanded on no reconcile, so this is the whole of what happens to one +// until somebody approves it. +func (r *ClusterDeploymentConfigReconciler) holdAsDraft( + ctx context.Context, + config *simplyblockv1alpha2.ClusterDeploymentConfig, + findings []finding, +) (ctrl.Result, error) { + if len(findings) > 0 { + r.emitFindings(config, findings) + return ctrl.Result{RequeueAfter: configRetry}, r.note(ctx, config, + simplyblockv1alpha2.ClusterDeploymentConfigPhaseDraft, summarize(findings)) + } + + // AwaitingApproval is emitted on the transition to a validated draft rather + // than on every reconcile, because a valid draft nobody has approved looks + // identical to a controller that has not noticed it and the event is what + // distinguishes them — once. + message := "the document is valid and is waiting for spec.approved" + if config.Status.Message != message { + r.emit(config, corev1.EventTypeNormal, AwaitingApproval, message) + } + return ctrl.Result{}, r.note(ctx, config, + simplyblockv1alpha2.ClusterDeploymentConfigPhaseDraft, message) +} + +// markReadyToDeploy writes the output annotation of §5. It is set on an approved +// document so a selector can find one, and read from nowhere. +func (r *ClusterDeploymentConfigReconciler) markReadyToDeploy( + ctx context.Context, config *simplyblockv1alpha2.ClusterDeploymentConfig, +) error { + if config.Labels[readyToDeploy] == readyToDeployValue { + return nil + } + patch := client.MergeFrom(config.DeepCopy()) + if config.Labels == nil { + config.Labels = map[string]string{} + } + config.Labels[readyToDeploy] = readyToDeployValue + return r.Patch(ctx, config, patch) +} + +// controlPlaneReady reports whether the singleton is available, which is the +// precondition for creating a cluster at all. +func (r *ClusterDeploymentConfigReconciler) controlPlaneReady( + ctx context.Context, +) (bool, string) { + var controlPlane simplyblockv1alpha2.ControlPlane + key := types.NamespacedName{Namespace: r.Namespace, Name: singletonControlPlane} + if err := r.Get(ctx, key, &controlPlane); err != nil { + return false, fmt.Sprintf("ControlPlane %s cannot be read: %v", + singletonControlPlane, err) + } + if controlPlane.Status.Phase != controlPlaneAvailable { + return false, fmt.Sprintf("ControlPlane %s is %s rather than %s", + singletonControlPlane, controlPlane.Status.Phase, controlPlaneAvailable) + } + return true, "" +} + +// everyUnfinishedDocument maps a Node event onto every document that has not +// finished, because a draft held by a worker that does not exist becomes valid the +// moment somebody adds it. +func (r *ClusterDeploymentConfigReconciler) everyUnfinishedDocument( + ctx context.Context, _ client.Object, +) []reconcile.Request { + var configs simplyblockv1alpha2.ClusterDeploymentConfigList + if err := r.List(ctx, &configs); err != nil { + return nil + } + var requests []reconcile.Request + for i := range configs.Items { + if terminalConfig(configs.Items[i].Status.Phase) { + continue + } + requests = append(requests, reconcile.Request{ + NamespacedName: client.ObjectKeyFromObject(&configs.Items[i]), + }) + } + return requests +} + +// succeed records the expansion's outcome. +func (r *ClusterDeploymentConfigReconciler) succeed( + ctx context.Context, config *simplyblockv1alpha2.ClusterDeploymentConfig, +) error { + message := fmt.Sprintf("expanded into cluster %s and %d node(s)", + config.Status.ClusterRef, len(config.Status.NodeRefs)) + r.emit(config, corev1.EventTypeNormal, NodesCreated, message) + return r.note(ctx, config, + simplyblockv1alpha2.ClusterDeploymentConfigPhaseExpanded, message) +} + +// fail records a terminal refusal. +func (r *ClusterDeploymentConfigReconciler) fail( + ctx context.Context, config *simplyblockv1alpha2.ClusterDeploymentConfig, message string, +) (ctrl.Result, error) { + return ctrl.Result{}, r.note(ctx, config, + simplyblockv1alpha2.ClusterDeploymentConfigPhaseFailed, message) +} + +// note writes the phase and the message, and patches only when something changed. +func (r *ClusterDeploymentConfigReconciler) note( + ctx context.Context, + config *simplyblockv1alpha2.ClusterDeploymentConfig, + phase simplyblockv1alpha2.ClusterDeploymentConfigPhase, + message string, +) error { + return r.writeStatus(ctx, config, + func(status *simplyblockv1alpha2.ClusterDeploymentConfigStatus) { + status.Phase = phase + status.Message = message + }) +} + +// recordStep persists the step the machine is about to be in, with the instant it +// expires. Both travel together, because a step persisted without its deadline +// restores as a step that can never time out. +func (r *ClusterDeploymentConfigReconciler) recordStep( + ctx context.Context, + config *simplyblockv1alpha2.ClusterDeploymentConfig, + next configStep, + deadline *metav1.Time, +) error { + return r.writeStatus(ctx, config, + func(status *simplyblockv1alpha2.ClusterDeploymentConfigStatus) { + status.Phase = simplyblockv1alpha2.ClusterDeploymentConfigPhaseExpanding + status.Step = statemachine.KubeSnapshot{ + State: string(next), Deadline: deadline, + } + }) +} + +// writeStatus applies the mutation and patches only when something changed. +func (r *ClusterDeploymentConfigReconciler) writeStatus( + ctx context.Context, + config *simplyblockv1alpha2.ClusterDeploymentConfig, + mutate func(*simplyblockv1alpha2.ClusterDeploymentConfigStatus), +) error { + return retry.RetryOnConflict(retry.DefaultRetry, func() error { + var fresh simplyblockv1alpha2.ClusterDeploymentConfig + if err := r.Get(ctx, client.ObjectKeyFromObject(config), &fresh); err != nil { + return err + } + + desired := *fresh.Status.DeepCopy() + mutate(&desired) + desired.ObservedGeneration = fresh.Generation + + if equalConfigStatus(fresh.Status, desired) { + config.Status = desired + config.ResourceVersion = fresh.ResourceVersion + return nil + } + + patch := client.MergeFromWithOptions(fresh.DeepCopy(), + client.MergeFromWithOptimisticLock{}) + fresh.Status = desired + if err := r.Status().Patch(ctx, &fresh, patch); err != nil { + return err + } + config.Status = fresh.Status + config.ResourceVersion = fresh.ResourceVersion + return nil + }) +} + +// emit raises an event on the document, which is what a reviewer has open. +func (r *ClusterDeploymentConfigReconciler) emit( + config *simplyblockv1alpha2.ClusterDeploymentConfig, + eventType, reason, message string, +) { + r.Recorder.Eventf(config, nil, eventType, reason, reason, "%s", message) +} + +// emitFindings raises one event per distinct reason validation produced. +func (r *ClusterDeploymentConfigReconciler) emitFindings( + config *simplyblockv1alpha2.ClusterDeploymentConfig, findings []finding, +) { + seen := map[string]struct{}{} + for _, found := range findings { + if _, already := seen[found.reason]; already { + continue + } + seen[found.reason] = struct{}{} + r.emit(config, corev1.EventTypeWarning, found.reason, found.message) + } +} + +// terminalConfig reports a phase the document can never leave. +func terminalConfig(phase simplyblockv1alpha2.ClusterDeploymentConfigPhase) bool { + switch phase { + case simplyblockv1alpha2.ClusterDeploymentConfigPhaseExpanded, + simplyblockv1alpha2.ClusterDeploymentConfigPhaseFailed: + return true + default: + return false + } +} + +// equalConfigStatus compares two statuses for the purpose of deciding whether to +// write. It is spelled out rather than reflect.DeepEqual because the status +// carries a pointer to a timestamp and a slice. +func equalConfigStatus(a, b simplyblockv1alpha2.ClusterDeploymentConfigStatus) bool { + if a.Phase != b.Phase || a.Message != b.Message || + a.ClusterRef != b.ClusterRef || + a.ObservedGeneration != b.ObservedGeneration || + a.Step.State != b.Step.State || + len(a.NodeRefs) != len(b.NodeRefs) { + return false + } + if (a.Step.Deadline == nil) != (b.Step.Deadline == nil) { + return false + } + if a.Step.Deadline != nil && !a.Step.Deadline.Equal(b.Step.Deadline) { + return false + } + for i := range a.NodeRefs { + if a.NodeRefs[i] != b.NodeRefs[i] { + return false + } + } + return true +} + +// refusedError is an expansion the document asked for and the operator will not +// perform: the cluster already exists, or the one it names does not. It is a +// distinct type because retrying cannot change either answer, and §6 is explicit +// that refusing beats merging. +type refusedError struct { + reason string + message string +} + +func (e *refusedError) Error() string { return e.message } + +func refusef(reason, format string, args ...any) error { + return &refusedError{reason: reason, message: fmt.Sprintf(format, args...)} +} + +// The ControlPlane the expansion waits on, and the phase it waits for. +const ( + singletonControlPlane = "simplyblock" + controlPlaneAvailable = "Available" +) diff --git a/operator/internal/controllers/deployment/clusterdeploymentconfig_controller_test.go b/operator/internal/controllers/deployment/clusterdeploymentconfig_controller_test.go new file mode 100644 index 000000000..0cf291c7e --- /dev/null +++ b/operator/internal/controllers/deployment/clusterdeploymentconfig_controller_test.go @@ -0,0 +1,422 @@ +// What the expansion does, and the two things it refuses to do. +// +// The cases that matter most are §6's: a config creates or adds and never +// reconciles a difference, because the differences that matter are of the form: +// this node's device list changed, and its only correct handling is not to apply +// it to a node that already has data on those devices. The other is the +// idempotence CreatingNodes needs, since a crash part-way through must create the +// rest rather than a second copy of everything. + +package deployment + +import ( + "context" + "strings" + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/client-go/tools/events" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + "github.com/simplyblock/atlas/ptr" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/testsupport" +) + +const ( + theNamespace = "simplyblock" + theCluster = "production" +) + +// aDocument is a valid, approved two-worker NVMe deployment, so a case states only +// what it is about. +func aDocument( + mutate func(*simplyblockv1alpha2.ClusterDeploymentConfig), +) *simplyblockv1alpha2.ClusterDeploymentConfig { + config := &simplyblockv1alpha2.ClusterDeploymentConfig{ + ObjectMeta: metav1.ObjectMeta{Name: "deployment", Namespace: theNamespace}, + Spec: simplyblockv1alpha2.ClusterDeploymentConfigSpec{ + Approved: true, + Environment: simplyblockv1alpha2.KubernetesEnvironmentVanilla, + Cluster: &simplyblockv1alpha2.ClusterTemplate{ + Name: theCluster, + MaxSubsystemCount: ptr.To(int32(20)), + VCPUCount: ptr.To(int32(8)), + MinHugePagesSize: "100G", + }, + NodeSets: []simplyblockv1alpha2.NodeSet{{ + Name: "rack-a", + Groups: []simplyblockv1alpha2.NodeGroup{{ + Name: "saturn", + Workers: []string{"worker-1", "worker-2"}, + MgmtInterface: "eth1", + DataInterfaces: []string{"eth2"}, + Devices: &simplyblockv1alpha2.DeviceSelection{ + NVMe: []string{"0000:5e:00.0"}, + }, + }}, + }}, + }, + } + if mutate != nil { + mutate(config) + } + return config +} + +// aCluster is a StorageCluster the control plane has already created. +func aCluster( + mutate func(*simplyblockv1alpha2.StorageCluster), +) *simplyblockv1alpha2.StorageCluster { + cluster := &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{Name: theCluster, Namespace: theNamespace}, + Spec: simplyblockv1alpha2.StorageClusterSpec{ + VCPUCount: ptr.To(int32(8)), + MinHugePagesSize: "100G", + }, + Status: simplyblockv1alpha2.StorageClusterStatus{UUID: "cluster-uuid"}, + } + if mutate != nil { + mutate(cluster) + } + return cluster +} + +func workers(names ...string) []client.Object { + out := make([]client.Object, 0, len(names)) + for _, name := range names { + out = append(out, &corev1.Node{ObjectMeta: metav1.ObjectMeta{Name: name}}) + } + return out +} + +func reconcilerFor(t *testing.T, objects ...client.Object) *ClusterDeploymentConfigReconciler { + t.Helper() + scheme := testsupport.NewScheme(t, corev1.AddToScheme) + builder := fake.NewClientBuilder().WithScheme(scheme). + WithStatusSubresource(&simplyblockv1alpha2.ClusterDeploymentConfig{}). + WithObjects(objects...) + return &ClusterDeploymentConfigReconciler{ + Client: builder.Build(), + Scheme: scheme, + Recorder: events.NewFakeRecorder(64), + Namespace: theNamespace, + } +} + +// A config that names a cluster it did not ask to create is refused rather than +// merged. Merging would have the operator decide what a difference means (§6). +func TestCreatingAClusterThatAlreadyExistsIsRefused(t *testing.T) { + config := aDocument(nil) + objects := append(workers("worker-1", "worker-2"), config, aCluster(nil)) + r := reconcilerFor(t, objects...) + + _, err := r.createCluster(context.Background(), config) + if err == nil { + t.Fatal("a document creating an existing cluster was accepted") + } + if !strings.Contains(err.Error(), "spec.clusterRef") { + t.Errorf("the refusal does not say how to fix it: %v", err) + } +} + +// The mirror image: a growth document whose cluster is not there. +func TestGrowingAClusterThatDoesNotExistIsRefused(t *testing.T) { + config := aDocument(func(c *simplyblockv1alpha2.ClusterDeploymentConfig) { + c.Spec.ClusterRef = "absent" + c.Spec.Cluster = nil + }) + objects := append(workers("worker-1", "worker-2"), config) + r := reconcilerFor(t, objects...) + + _, err := r.createCluster(context.Background(), config) + if err == nil { + t.Fatal("a document growing a cluster that does not exist was accepted") + } + if !strings.Contains(err.Error(), "does not exist") { + t.Errorf("the refusal does not say the cluster is missing: %v", err) + } +} + +// The ordinary path: the cluster is created, and the class is stamped off the +// groups because the document carries no field for it. +func TestTheClusterIsCreatedWithTheClassReadOffTheGroups(t *testing.T) { + config := aDocument(nil) + objects := append(workers("worker-1", "worker-2"), config) + r := reconcilerFor(t, objects...) + + done, err := r.createCluster(context.Background(), config) + if err != nil || !done { + t.Fatalf("createCluster: done=%v err=%v", done, err) + } + + var created simplyblockv1alpha2.StorageCluster + key := client.ObjectKey{Namespace: theNamespace, Name: theCluster} + if err := r.Get(context.Background(), key, &created); err != nil { + t.Fatalf("the cluster was not created: %v", err) + } + if created.Spec.DeviceClass != simplyblockv1alpha2.StorageClusterDeviceClassNVMe { + t.Errorf("deviceClass = %q, want NVMe read off the groups", created.Spec.DeviceClass) + } + if created.Spec.StorageNodes == nil || created.Spec.StorageNodes.MgmtInterface != "eth1" { + t.Errorf("the workload did not take the group's interfaces: %+v", created.Spec.StorageNodes) + } +} + +// One node per worker per slot, with the cluster's sizing copied in and the +// group's devices expanded into the one list a node carries. +func TestTheNodesAreCreatedFromTheGroups(t *testing.T) { + config := aDocument(nil) + config.Status.ClusterRef = theCluster + objects := append(workers("worker-1", "worker-2"), config, aCluster(nil)) + r := reconcilerFor(t, objects...) + + done, err := r.createNodes(context.Background(), config) + if err != nil || !done { + t.Fatalf("createNodes: done=%v err=%v", done, err) + } + + var nodes simplyblockv1alpha2.StorageNodeList + if err := r.List(context.Background(), &nodes); err != nil { + t.Fatalf("listing the nodes: %v", err) + } + if len(nodes.Items) != 2 { + t.Fatalf("created %d nodes, want one per worker", len(nodes.Items)) + } + + for i := range nodes.Items { + node := &nodes.Items[i] + if node.Spec.ClusterRef != theCluster { + t.Errorf("node %s names cluster %q", node.Name, node.Spec.ClusterRef) + } + if node.Spec.NodeSet != "rack-a" { + t.Errorf("node %s does not trace back to its node set: %q", node.Name, node.Spec.NodeSet) + } + // The sizing is the cluster's, copied in, so the node records the layout + // it was built with and nothing refers back to the document. + if node.Spec.Config.Sizing.VCPUCount == nil || *node.Spec.Config.Sizing.VCPUCount != 8 { + t.Errorf("node %s did not take the cluster's sizing", node.Name) + } + if len(node.Spec.Config.DeviceNames) != 1 || + node.Spec.Config.DeviceNames[0] != "0000:5e:00.0" { + t.Errorf("node %s did not take the group's devices: %v", + node.Name, node.Spec.Config.DeviceNames) + } + } +} + +// The step the expansion cannot get wrong: a crash part-way through creates the +// rest on the next pass and duplicates nothing. +func TestCreatingNodesIsIdempotent(t *testing.T) { + config := aDocument(nil) + config.Status.ClusterRef = theCluster + objects := append(workers("worker-1", "worker-2"), config, aCluster(nil)) + r := reconcilerFor(t, objects...) + + for pass := 0; pass < 3; pass++ { + if _, err := r.createNodes(context.Background(), config); err != nil { + t.Fatalf("pass %d: %v", pass, err) + } + } + + var nodes simplyblockv1alpha2.StorageNodeList + if err := r.List(context.Background(), &nodes); err != nil { + t.Fatalf("listing the nodes: %v", err) + } + if len(nodes.Items) != 2 { + t.Fatalf("three passes produced %d nodes, want the same two", len(nodes.Items)) + } +} + +// A two-socket layout produces one node per socket per worker, and each carries +// the socket it is bound to. +func TestATwoSocketLayoutProducesTwoNodesPerWorker(t *testing.T) { + config := aDocument(nil) + config.Status.ClusterRef = theCluster + cluster := aCluster(func(c *simplyblockv1alpha2.StorageCluster) { + c.Spec.StorageNodes = &simplyblockv1alpha2.StorageNodesSpec{ + SocketsToUse: []string{"0", "1"}, + NodesPerSocket: ptr.To(int32(1)), + } + }) + objects := append(workers("worker-1", "worker-2"), config, cluster) + r := reconcilerFor(t, objects...) + + if _, err := r.createNodes(context.Background(), config); err != nil { + t.Fatalf("createNodes: %v", err) + } + + var nodes simplyblockv1alpha2.StorageNodeList + if err := r.List(context.Background(), &nodes); err != nil { + t.Fatalf("listing the nodes: %v", err) + } + if len(nodes.Items) != 4 { + t.Fatalf("created %d nodes, want two workers by two sockets", len(nodes.Items)) + } + + sockets := map[string]int{} + for i := range nodes.Items { + sockets[nodes.Items[i].Spec.SocketID]++ + } + if sockets["0"] != 2 || sockets["1"] != 2 { + t.Errorf("the nodes are not spread over both sockets: %v", sockets) + } +} + +// A growth document adds only the slots that are not filled, which is what makes +// adding a rack a second document rather than an edit of the first. +func TestAGrowthDocumentAddsOnlyTheMissingNodes(t *testing.T) { + existing := &simplyblockv1alpha2.StorageNode{ + ObjectMeta: metav1.ObjectMeta{Name: "already-there", Namespace: theNamespace}, + Spec: simplyblockv1alpha2.StorageNodeSpec{ + ClusterRef: theCluster, + WorkerNode: "worker-1", + Slot: ptr.To(int32(0)), + }, + } + config := aDocument(nil) + config.Status.ClusterRef = theCluster + objects := append(workers("worker-1", "worker-2"), config, aCluster(nil), existing) + r := reconcilerFor(t, objects...) + + if _, err := r.createNodes(context.Background(), config); err != nil { + t.Fatalf("createNodes: %v", err) + } + + var nodes simplyblockv1alpha2.StorageNodeList + if err := r.List(context.Background(), &nodes); err != nil { + t.Fatalf("listing the nodes: %v", err) + } + if len(nodes.Items) != 2 { + t.Fatalf("created %d nodes, want the one that was missing", len(nodes.Items)) + } + for i := range nodes.Items { + if nodes.Items[i].Name != "already-there" && + nodes.Items[i].Spec.WorkerNode != "worker-2" { + t.Errorf("the new node is on %s, want the unfilled worker", + nodes.Items[i].Spec.WorkerNode) + } + } +} + +// A draft naming a worker that is not there is reported rather than expanded, +// which is the whole value of the review gate. +func TestADraftNamingAMissingWorkerIsReported(t *testing.T) { + config := aDocument(func(c *simplyblockv1alpha2.ClusterDeploymentConfig) { + c.Spec.Approved = false + }) + objects := append(workers("worker-1"), config) + r := reconcilerFor(t, objects...) + + findings, err := r.validate(context.Background(), config) + if err != nil { + t.Fatalf("validate: %v", err) + } + if len(findings) != 1 || findings[0].reason != WorkerNotFound { + t.Fatalf("findings = %+v, want one WorkerNotFound", findings) + } + if !strings.Contains(findings[0].message, "worker-2") { + t.Errorf("the finding does not name the missing worker: %s", findings[0].message) + } +} + +// A valid draft produces nothing to report, which is what AwaitingApproval is +// emitted against. +func TestAValidDraftHasNoFindings(t *testing.T) { + config := aDocument(func(c *simplyblockv1alpha2.ClusterDeploymentConfig) { + c.Spec.Approved = false + }) + objects := append(workers("worker-1", "worker-2"), config) + r := reconcilerFor(t, objects...) + + findings, err := r.validate(context.Background(), config) + if err != nil { + t.Fatalf("validate: %v", err) + } + if len(findings) != 0 { + t.Errorf("a valid draft reported %+v", findings) + } +} + +// A growth document whose groups name the other class is reported while it is +// still editable, because approving it is what makes it immutable. +func TestAGrowthDocumentOfTheWrongClassIsReported(t *testing.T) { + config := aDocument(func(c *simplyblockv1alpha2.ClusterDeploymentConfig) { + c.Spec.Approved = false + c.Spec.ClusterRef = theCluster + c.Spec.Cluster = nil + c.Spec.NodeSets[0].Groups[0].Devices = &simplyblockv1alpha2.DeviceSelection{ + Block: []string{"/dev/sdb"}, + } + }) + cluster := aCluster(func(c *simplyblockv1alpha2.StorageCluster) { + c.Spec.DeviceClass = simplyblockv1alpha2.StorageClusterDeviceClassNVMe + }) + objects := append(workers("worker-1", "worker-2"), config, cluster) + r := reconcilerFor(t, objects...) + + findings, err := r.validate(context.Background(), config) + if err != nil { + t.Fatalf("validate: %v", err) + } + if len(findings) != 1 || findings[0].reason != DeviceClassMismatch { + t.Fatalf("findings = %+v, want one DeviceClassMismatch", findings) + } +} + +// A cluster with failure domains enabled needs every group to name one, and +// saying so at draft time turns an investigation into an edit. +func TestFailureDomainsAreCheckedAgainstTheTemplate(t *testing.T) { + config := aDocument(func(c *simplyblockv1alpha2.ClusterDeploymentConfig) { + c.Spec.Approved = false + c.Spec.Cluster.EnableFailureDomains = ptr.To(true) + }) + objects := append(workers("worker-1", "worker-2"), config) + r := reconcilerFor(t, objects...) + + findings, err := r.validate(context.Background(), config) + if err != nil { + t.Fatalf("validate: %v", err) + } + if len(findings) != 1 { + t.Fatalf("findings = %+v, want the missing fault group", findings) + } + if !strings.Contains(findings[0].message, "rack-a/saturn") { + t.Errorf("the finding does not name the group: %s", findings[0].message) + } +} + +// The environment is a shorthand and the expansion is where it is spent: naming +// OpenShift once decides the distribution flags, after which nothing reads it. +func TestTheEnvironmentResolvesIntoTheWorkloadFlags(t *testing.T) { + config := aDocument(func(c *simplyblockv1alpha2.ClusterDeploymentConfig) { + c.Spec.Environment = simplyblockv1alpha2.KubernetesEnvironmentOpenShift + }) + r := reconcilerFor(t) + + workload := r.buildWorkload(config) + if workload.OpenShiftCluster == nil || !*workload.OpenShiftCluster { + t.Error("OpenShift did not set openShiftCluster") + } + if workload.EnableCpuTopology == nil || !*workload.EnableCpuTopology { + t.Error("OpenShift did not set enableCpuTopology") + } +} + +// Every node of one document gets a name of its own, because the name is derived +// from the cluster, the worker, and the slot. +func TestEveryNodeGetsADistinctName(t *testing.T) { + seen := map[string]struct{}{} + for _, worker := range []string{"worker-1", "worker-2"} { + for slot := int32(0); slot < 2; slot++ { + name := nodeName(theCluster, worker, slot) + if _, clash := seen[name]; clash { + t.Fatalf("two slots derived the same name %q", name) + } + seen[name] = struct{}{} + } + } +} diff --git a/operator/internal/controllers/deployment/events.go b/operator/internal/controllers/deployment/events.go new file mode 100644 index 000000000..143f3b793 --- /dev/null +++ b/operator/internal/controllers/deployment/events.go @@ -0,0 +1,48 @@ +// The reasons this package emits events under. +// +// They are collected here rather than declared beside the code that raises them +// because a reason is a contract with whoever is reading `kubectl describe`: it is +// what an administrator greps for and what an alert matches on, so it outlives the +// function that happens to raise it today. +// +// Two kinds carry them. A document's own validation and expansion go on the +// ClusterDeploymentConfig, which is what a reviewer has open; a discovery run's go +// on the OperatorOps, which outlives the run as its record. +// +// design-clusterdeploymentconfig.md §9.1 is the specification. + +package deployment + +const ( + // What a draft's validation found. These are the whole value of the review + // gate: a document that names a worker which does not exist should say so + // while it is still a draft, rather than after somebody approved it. + WorkerNotFound = "WorkerNotFound" + DeviceNotFound = "DeviceNotFound" + DeviceClassMismatch = "DeviceClassMismatch" + + // AwaitingApproval is the one that changes how the kind is used. A valid draft + // nobody has approved looks identical to a controller that has not noticed it, + // and this event is what distinguishes them. It is emitted on the transition + // to a validated draft rather than on every reconcile. + AwaitingApproval = "AwaitingApproval" + + // ControlPlaneNotReady holds an approved document. The expansion creates a + // StorageCluster, and a control plane that cannot accept one would leave the + // cluster in a state the document did not describe. + ControlPlaneNotReady = "ControlPlaneNotReady" + + // The two refusals of §6. A config creates or adds and never reconciles a + // difference, so a document that names a cluster it did not expect is a + // failure with a reason rather than a merge somebody has to unpick. + ClusterExists = "ClusterExists" + ClusterNotFound = "ClusterNotFound" + + // What the expansion produced. + ClusterCreated = "ClusterCreated" + NodesCreated = "NodesCreated" + + // StepDeadlineExceeded distinguishes an expansion still working from one that + // stopped, which is the distinction status.message cannot express. + StepDeadlineExceeded = "StepDeadlineExceeded" +) diff --git a/operator/internal/controllers/deployment/expansion.go b/operator/internal/controllers/deployment/expansion.go new file mode 100644 index 000000000..c06c086b0 --- /dev/null +++ b/operator/internal/controllers/deployment/expansion.go @@ -0,0 +1,537 @@ +// The expansion machine: the graph a document walks, and what each of its four +// steps does. +// +// Validating ──► CreatingCluster ──► AwaitingCluster ──► CreatingNodes +// +// It is a Config rather than a MultiConfig, because a document has no action to +// key one on: there is one expansion and it runs once. +// +// CreatingNodes is the step that must be idempotent, and it is by construction. A +// StorageNode is identified by its cluster, its worker, and its slot, so the step +// lists what exists for the cluster and creates only the slots that do not. A +// crash part-way through creates the rest on the next pass and duplicates nothing. +// +// design-clusterdeploymentconfig.md §4.2 is the specification. + +package deployment + +import ( + "context" + "fmt" + "slices" + "sort" + "time" + + corev1 "k8s.io/api/core/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + + "github.com/simplyblock/atlas/kube" + "github.com/simplyblock/atlas/ptr" + "github.com/simplyblock/atlas/statemachine" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// configStep is the document's step type, aliased so the graph reads as the graph. +type configStep = simplyblockv1alpha2.ClusterDeploymentConfigStep + +const ( + stepValidating = simplyblockv1alpha2.ClusterDeploymentConfigStepValidating + stepCreatingCluster = simplyblockv1alpha2.ClusterDeploymentConfigStepCreatingCluster + stepAwaitingCluster = simplyblockv1alpha2.ClusterDeploymentConfigStepAwaitingCluster + stepCreatingNodes = simplyblockv1alpha2.ClusterDeploymentConfigStepCreatingNodes +) + +// How long each step may take before it is reported as stuck. +// +// They differ by what the step is waiting on. Validation and the two creates are +// Kubernetes writes; awaiting the cluster is the control plane creating one, which +// is the only step here that waits on something outside Kubernetes at all. +const ( + validatingDeadline = 5 * time.Minute + creatingClusterDeadline = 5 * time.Minute + awaitingClusterDeadline = 30 * time.Minute + creatingNodesDeadline = 10 * time.Minute +) + +// expansionGraph declares the document's one state graph. +func expansionGraph() statemachine.Config[configStep] { + deadline := func(d time.Duration) statemachine.TransitionFunc[configStep] { + return func(context.Context, configStep, configStep) (time.Duration, error) { + return d, nil + } + } + return statemachine.Config[configStep]{ + Initial: stepValidating, + States: map[configStep]statemachine.StateDef[configStep]{ + stepValidating: { + To: []configStep{stepCreatingCluster}, + OnEnter: deadline(validatingDeadline), + }, + stepCreatingCluster: { + To: []configStep{stepAwaitingCluster}, + OnEnter: deadline(creatingClusterDeadline), + }, + stepAwaitingCluster: { + To: []configStep{stepCreatingNodes}, + OnEnter: deadline(awaitingClusterDeadline), + }, + stepCreatingNodes: {OnEnter: deadline(creatingNodesDeadline)}, + }, + } +} + +// nextStep is the step that follows the current one. The graph is a line, so the +// first edge is the only edge. +func nextStep(machine *statemachine.Machine[configStep]) (configStep, error) { + current := machine.CurrentState() + for next := range machine.AllowedTransitions() { + return next, nil + } + return current, fmt.Errorf("step %s declares no successor and is not terminal", current) +} + +// createCluster creates the StorageCluster the document describes, or resolves the +// one it names, and refuses the two combinations §6 will not perform. +// +// It stamps the device class it read off the groups. The document carries no field +// for it, so the step takes the member every group used and writes it, where it is +// immutable from that moment. Resolving an existing cluster writes nothing: the +// cluster's class already holds. +func (r *ClusterDeploymentConfigReconciler) createCluster( + ctx context.Context, config *simplyblockv1alpha2.ClusterDeploymentConfig, +) (bool, error) { + name, err := targetClusterName(config) + if err != nil { + return false, err + } + + var existing simplyblockv1alpha2.StorageCluster + key := client.ObjectKey{Namespace: config.Namespace, Name: name} + getErr := r.Get(ctx, key, &existing) + + switch { + case getErr == nil && config.Spec.ClusterRef == "": + // The document asked to create a cluster and one is already there. + // Merging would have the operator decide what a difference means, and the + // differences that matter are of the form "this node's device list + // changed" (§6). + return false, refusef(ClusterExists, + "spec.cluster.name is %s and a StorageCluster by that name already exists; "+ + "set spec.clusterRef to add nodes to it instead", name) + + case apierrors.IsNotFound(getErr) && config.Spec.ClusterRef != "": + return false, refusef(ClusterNotFound, + "spec.clusterRef names StorageCluster %s, which does not exist", name) + + case getErr == nil: + // Growth against a cluster that is there. Nothing to create, and the + // class is settled before the document existed. + return true, r.recordCluster(ctx, config, name) + + case !apierrors.IsNotFound(getErr): + return false, fmt.Errorf("reading StorageCluster %s: %w", name, getErr) + } + + cluster, err := r.buildCluster(config, name) + if err != nil { + return false, err + } + if err := r.Create(ctx, cluster); err != nil { + if apierrors.IsAlreadyExists(err) { + // The read above missed and somebody created the cluster between the + // two. That is the ClusterExists case arriving by a different route, + // not a success: the document asked to create a cluster and did not, + // and nothing here proves the one that is there is the one it + // described (§6). + return false, refusef(ClusterExists, + "spec.cluster.name is %s and a StorageCluster by that name was created "+ + "while this document was being expanded; set spec.clusterRef to add "+ + "nodes to it instead", name) + } + return false, fmt.Errorf("creating StorageCluster %s: %w", name, err) + } + r.emit(config, corev1.EventTypeNormal, ClusterCreated, + fmt.Sprintf("Created StorageCluster %s", name)) + return true, r.recordCluster(ctx, config, name) +} + +// buildCluster is the StorageCluster the document's template describes. +func (r *ClusterDeploymentConfigReconciler) buildCluster( + config *simplyblockv1alpha2.ClusterDeploymentConfig, name string, +) (*simplyblockv1alpha2.StorageCluster, error) { + template := config.Spec.Cluster + if template == nil { + return nil, refusef(ClusterNotFound, + "the document neither names an existing cluster nor describes one to create") + } + + class := deviceClassOf(config) + + cluster := &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: config.Namespace}, + Spec: simplyblockv1alpha2.StorageClusterSpec{ + MaxSubsystemCount: template.MaxSubsystemCount, + VCPUCount: template.VCPUCount, + MinHugePagesSize: template.MinHugePagesSize, + Stripe: template.Stripe, + FabricType: template.FabricType, + EnableFailureDomains: template.EnableFailureDomains, + DeviceClass: class, + // The workload every node runs as. The document's per-group network + // interfaces are the same for every group of a cluster in practice, + // and the cluster is where a DaemonSet can carry them at all + // (design-storagenode.md §5.1). + StorageNodes: r.buildWorkload(config), + }, + } + return cluster, nil +} + +// buildWorkload resolves spec.environment into the four distribution flags and +// carries the first group's interfaces onto the cluster. +// +// A DaemonSet is one object for every node it schedules, so its pod template +// cannot differ per group: the interfaces are the cluster's whichever group states +// them. Taking the first stated is what makes a document that repeats them in +// every group, which is how discovery writes one, mean what it looks like. +func (r *ClusterDeploymentConfigReconciler) buildWorkload( + config *simplyblockv1alpha2.ClusterDeploymentConfig, +) *simplyblockv1alpha2.StorageNodesSpec { + workload := &simplyblockv1alpha2.StorageNodesSpec{} + if template := config.Spec.Cluster; template != nil { + // The socket layout decides how many storage nodes a worker runs, and + // CreatingNodes reads it back off the cluster. Leaving it unset here gave + // every cluster a document created one node on socket 0, whatever the + // document said. + workload.SocketsToUse = template.SocketsToUse + workload.NodesPerSocket = template.NodesPerSocket + } + + for _, set := range config.Spec.NodeSets { + for _, group := range set.Groups { + if workload.MgmtInterface == "" { + workload.MgmtInterface = group.MgmtInterface + } + if len(workload.DataInterfaces) == 0 { + workload.DataInterfaces = group.DataInterfaces + } + } + } + + // The environment is a shorthand and this is where it is spent. Naming + // OpenShift once decides all four, after which nothing reads the field again + // and the nodes carry the resolved flags (§3.1). + switch config.Spec.Environment { + case simplyblockv1alpha2.KubernetesEnvironmentOpenShift: + workload.OpenShiftCluster = ptr.To(true) + workload.EnableCpuTopology = ptr.To(true) + // Stated rather than left nil. The renderer reads an unset flag as skipping + // the kubelet configuration, and the settings this product has shipped + // for OpenShift all configure it, so silence here would change what an + // OpenShift deployment does. + workload.EnableKubeletConfiguration = ptr.To(true) + case simplyblockv1alpha2.KubernetesEnvironmentTalos: + // Talos has no writable kubelet configuration and no package manager, so + // the node applies neither. + workload.EnableKubeletConfiguration = ptr.To(false) + case simplyblockv1alpha2.KubernetesEnvironmentVanilla, + simplyblockv1alpha2.KubernetesEnvironmentRancher, + simplyblockv1alpha2.KubernetesEnvironmentK3s: + workload.EnableKubeletConfiguration = ptr.To(true) + } + return workload +} + +// awaitCluster waits for the control plane to have created the cluster, which is +// what status.uuid says. +func (r *ClusterDeploymentConfigReconciler) awaitCluster( + ctx context.Context, config *simplyblockv1alpha2.ClusterDeploymentConfig, +) (bool, error) { + var cluster simplyblockv1alpha2.StorageCluster + key := client.ObjectKey{Namespace: config.Namespace, Name: config.Status.ClusterRef} + if err := r.Get(ctx, key, &cluster); err != nil { + return false, fmt.Errorf("reading StorageCluster %s: %w", config.Status.ClusterRef, err) + } + return cluster.Status.UUID != "", nil +} + +// createNodes writes one StorageNode per worker per slot, and creates only the +// slots that do not exist. +// +// The slot count comes from the cluster's own workload block, so a group of two +// workers on a two-socket layout produces four nodes. Reading it off the cluster +// rather than off the document is what makes a growth document produce nodes that +// match the fleet it is joining. +func (r *ClusterDeploymentConfigReconciler) createNodes( + ctx context.Context, config *simplyblockv1alpha2.ClusterDeploymentConfig, +) (bool, error) { + var cluster simplyblockv1alpha2.StorageCluster + key := client.ObjectKey{Namespace: config.Namespace, Name: config.Status.ClusterRef} + if err := r.Get(ctx, key, &cluster); err != nil { + return false, fmt.Errorf("reading StorageCluster %s: %w", config.Status.ClusterRef, err) + } + + existing, err := r.nodesOfCluster(ctx, config.Namespace, cluster.Name) + if err != nil { + return false, err + } + + // The record is rebuilt from the slots the document describes rather than + // accumulated across passes. A pass that created a node and then failed to + // persist status.nodeRefs would otherwise skip it on the next pass, as one + // that already exists, and the finished document would permanently omit a + // node it created. + var created []string + for _, set := range config.Spec.NodeSets { + for _, group := range set.Groups { + for _, worker := range group.Workers { + for slot := int32(0); slot < slotsPerWorker(&cluster); slot++ { + if name, there := existing[slotKey{worker: worker, slot: slot}]; there { + created = append(created, name) + continue + } + node := r.buildNode(config, &cluster, set, group, worker, slot) + if err := r.Create(ctx, node); err != nil { + if apierrors.IsAlreadyExists(err) { + created = append(created, node.Name) + continue + } + return false, fmt.Errorf("creating StorageNode for worker %s slot %d: %w", + worker, slot, err) + } + created = append(created, node.Name) + } + } + } + } + + sort.Strings(created) + created = slices.Compact(created) + return true, r.recordNodes(ctx, config, created) +} + +// slotKey is the identity a StorageNode has within its cluster, which is what +// makes the creation idempotent. +type slotKey struct { + worker string + slot int32 +} + +// nodesOfCluster indexes the cluster's existing nodes by the slot each fills. +func (r *ClusterDeploymentConfigReconciler) nodesOfCluster( + ctx context.Context, namespace, cluster string, +) (map[slotKey]string, error) { + var nodes simplyblockv1alpha2.StorageNodeList + if err := r.List(ctx, &nodes, client.InNamespace(namespace)); err != nil { + return nil, fmt.Errorf("listing the cluster's nodes: %w", err) + } + filled := map[slotKey]string{} + for i := range nodes.Items { + node := &nodes.Items[i] + if node.Spec.ClusterRef != cluster { + continue + } + slot := int32(0) + if node.Spec.Slot != nil { + slot = *node.Spec.Slot + } + filled[slotKey{worker: node.Spec.WorkerNode, slot: slot}] = node.Name + } + return filled, nil +} + +// slotsPerWorker is how many storage nodes a worker runs, which is one per socket +// per nodesPerSocket. +func slotsPerWorker(cluster *simplyblockv1alpha2.StorageCluster) int32 { + workload := cluster.Spec.StorageNodes + if workload == nil { + return 1 + } + sockets := int32(len(workload.SocketsToUse)) + if sockets == 0 { + // An empty list means socket 0 alone (design-storagenode.md §5.1). + sockets = 1 + } + perSocket := int32(1) + if workload.NodesPerSocket != nil && *workload.NodesPerSocket > 0 { + perSocket = *workload.NodesPerSocket + } + return sockets * perSocket +} + +// buildNode resolves the document's shorthands into one node. +// +// Nothing on the node refers back to the config, which is what makes the document +// safe to delete: the sizing is the cluster's, copied in, and the devices are the +// group's, expanded into one list. +func (r *ClusterDeploymentConfigReconciler) buildNode( + config *simplyblockv1alpha2.ClusterDeploymentConfig, + cluster *simplyblockv1alpha2.StorageCluster, + set simplyblockv1alpha2.NodeSet, + group simplyblockv1alpha2.NodeGroup, + worker string, + slot int32, +) *simplyblockv1alpha2.StorageNode { + socket, index := decomposeSlot(cluster, slot) + + return &simplyblockv1alpha2.StorageNode{ + ObjectMeta: metav1.ObjectMeta{ + Name: nodeName(cluster.Name, worker, slot), + Namespace: config.Namespace, + Labels: map[string]string{ + "storage.simplyblock.io/cluster": cluster.Name, + "storage.simplyblock.io/worker": worker, + }, + }, + Spec: simplyblockv1alpha2.StorageNodeSpec{ + ClusterRef: cluster.Name, + NodeSet: set.Name, + WorkerNode: worker, + SocketID: socket, + NodeIndex: ptr.To(index), + Slot: ptr.To(slot), + Config: simplyblockv1alpha2.StorageNodeConfig{ + // Two of the cluster's three sizing values are copied onto the + // node, so it records the layout it was built with and the + // operator can re-size one node at a time during a hardware + // upgrade. maxSubsystemCount is read from the cluster instead + // (design-storagenode.md §3.1). + Sizing: simplyblockv1alpha2.StorageNodeSizing{ + VCPUCount: cluster.Spec.VCPUCount, + MinHugePagesSize: cluster.Spec.MinHugePagesSize, + }, + // A node joining a cluster that already exists is an expansion, + // which the control plane reads as a request to rebalance onto it + // rather than to treat it as part of an initial layout. A growth + // document is exactly that case. + Expand: expansionOf(config), + DeviceNames: devicesOf(group), + FailureDomain: group.FailureDomain, + SpdkSystemMemory: group.SpdkSystemMemory, + JournalManager: group.JournalManager, + }, + }, + } +} + +// decomposeSlot renders a slot as the socket and the position within it, which is +// the pair a print column shows. Nothing but those columns reads either. +func decomposeSlot( + cluster *simplyblockv1alpha2.StorageCluster, slot int32, +) (socket string, index int32) { + workload := cluster.Spec.StorageNodes + perSocket := int32(1) + if workload != nil && workload.NodesPerSocket != nil && *workload.NodesPerSocket > 0 { + perSocket = *workload.NodesPerSocket + } + + position := slot / perSocket + index = slot % perSocket + if workload != nil && int(position) < len(workload.SocketsToUse) { + return workload.SocketsToUse[position], index + } + return fmt.Sprintf("%d", position), index +} + +// expansionOf reports whether the nodes this document creates are joining a +// cluster that already exists. +// +// The control plane reads the flag as a request to rebalance onto the new node +// rather than to treat it as part of an initial layout, so a growth document's +// nodes have to carry it: they are by definition an addition to an active +// cluster. A document that creates its own cluster is the initial layout, and +// leaves it unset. +func expansionOf(config *simplyblockv1alpha2.ClusterDeploymentConfig) *bool { + if config.Spec.ClusterRef == "" { + return nil + } + return ptr.To(true) +} + +// devicesOf expands a group's device selection into the one list a node carries. +// Both members expand into the same field, which takes a PCI address and a device +// path alike. +func devicesOf(group simplyblockv1alpha2.NodeGroup) []string { + if group.Devices == nil { + return nil + } + if len(group.Devices.NVMe) > 0 { + return group.Devices.NVMe + } + return group.Devices.Block +} + +// deviceClassOf reads the class off the document's groups, which is where it is +// stated: the device lists already say which class the deployment uses, so +// spec.cluster does not restate it. +func deviceClassOf( + config *simplyblockv1alpha2.ClusterDeploymentConfig, +) simplyblockv1alpha2.StorageClusterDeviceClass { + for _, set := range config.Spec.NodeSets { + for _, group := range set.Groups { + if group.Devices == nil { + continue + } + if len(group.Devices.NVMe) > 0 { + return simplyblockv1alpha2.StorageClusterDeviceClassNVMe + } + if len(group.Devices.Block) > 0 { + return simplyblockv1alpha2.StorageClusterDeviceClassLogicalBlock + } + } + } + // A document whose groups name no devices at all describes a cluster whose + // class nothing states. NVMe is what the cluster's own field defaults to, and + // stamping it explicitly here would be inventing a fact the document does not + // carry. + return "" +} + +// targetClusterName is the cluster the document acts on, whichever way it names +// one. +func targetClusterName( + config *simplyblockv1alpha2.ClusterDeploymentConfig, +) (string, error) { + if config.Spec.ClusterRef != "" { + return config.Spec.ClusterRef, nil + } + if config.Spec.Cluster != nil && config.Spec.Cluster.Name != "" { + return config.Spec.Cluster.Name, nil + } + return "", refusef(ClusterNotFound, + "the document neither names an existing cluster nor describes one to create") +} + +// nodeNameFormula names one node. A StorageNode is named for its cluster and the +// slot it fills, never for the worker, because the name has to stay stable when a +// migration re-points the node onto another host (design-storagenode.md §3.1) — +// the worker is in the name's digest rather than in its text. +var nodeNameFormula = kube.Formula{} + +func nodeName(cluster, worker string, slot int32) string { + return nodeNameFormula.Derive(cluster, worker, fmt.Sprintf("%d", slot)).Value +} + +// recordCluster writes which cluster the expansion produced or joined. +func (r *ClusterDeploymentConfigReconciler) recordCluster( + ctx context.Context, config *simplyblockv1alpha2.ClusterDeploymentConfig, name string, +) error { + return r.writeStatus(ctx, config, + func(status *simplyblockv1alpha2.ClusterDeploymentConfigStatus) { + status.ClusterRef = name + }) +} + +// recordNodes writes which nodes it created. It is a record rather than a +// dependency: nothing resolves it after the expansion. +func (r *ClusterDeploymentConfigReconciler) recordNodes( + ctx context.Context, config *simplyblockv1alpha2.ClusterDeploymentConfig, names []string, +) error { + return r.writeStatus(ctx, config, + func(status *simplyblockv1alpha2.ClusterDeploymentConfigStatus) { + status.NodeRefs = names + }) +} diff --git a/operator/internal/controllers/deployment/expansion_review_test.go b/operator/internal/controllers/deployment/expansion_review_test.go new file mode 100644 index 000000000..711af1e1d --- /dev/null +++ b/operator/internal/controllers/deployment/expansion_review_test.go @@ -0,0 +1,296 @@ +// The cases a review of the expansion found, each of which was a way for a +// document to be expanded into something it did not describe. +// +// They are together in one file because they share a shape rather than a +// subsystem: every one of them was silent. A cluster bound to the wrong network, +// a fleet forced onto one socket, a growth node that looked like an initial one — +// none of them failed, and none of them said anything. + +package deployment + +import ( + "context" + "errors" + "strings" + "testing" + + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + + "github.com/simplyblock/atlas/ptr" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// The socket layout reaches the cluster the document creates. Without it, +// CreatingNodes read one slot off a cluster nothing had configured, and every +// deployment a document created ran one storage node per worker whatever it said. +func TestTheSocketLayoutReachesTheCreatedCluster(t *testing.T) { + config := aDocument(func(c *simplyblockv1alpha2.ClusterDeploymentConfig) { + c.Spec.Cluster.SocketsToUse = []string{"0", "1"} + c.Spec.Cluster.NodesPerSocket = ptr.To(int32(2)) + }) + objects := append(workers("worker-1", "worker-2"), config) + r := reconcilerFor(t, objects...) + + if _, err := r.createCluster(context.Background(), config); err != nil { + t.Fatalf("createCluster: %v", err) + } + + var created simplyblockv1alpha2.StorageCluster + key := client.ObjectKey{Namespace: theNamespace, Name: theCluster} + if err := r.Get(context.Background(), key, &created); err != nil { + t.Fatalf("reading the cluster: %v", err) + } + workload := created.Spec.StorageNodes + if workload == nil { + t.Fatal("the cluster carries no workload") + } + if len(workload.SocketsToUse) != 2 { + t.Errorf("socketsToUse = %v, want both sockets", workload.SocketsToUse) + } + if workload.NodesPerSocket == nil || *workload.NodesPerSocket != 2 { + t.Errorf("nodesPerSocket = %v, want 2", workload.NodesPerSocket) + } + + // Which is what the node count then follows from: two workers, two sockets, + // two nodes per socket. + config.Status.ClusterRef = theCluster + if _, err := r.createNodes(context.Background(), config); err != nil { + t.Fatalf("createNodes: %v", err) + } + var nodes simplyblockv1alpha2.StorageNodeList + if err := r.List(context.Background(), &nodes); err != nil { + t.Fatalf("listing the nodes: %v", err) + } + if len(nodes.Items) != 8 { + t.Errorf("created %d nodes, want two workers by four slots", len(nodes.Items)) + } +} + +// A create that loses a race is the ClusterExists case arriving by another route. +// Reading it as success would have the document await and then use a cluster it +// did not create and cannot prove it described. +func TestACreateThatLosesTheRaceIsRefused(t *testing.T) { + config := aDocument(nil) + objects := append(workers("worker-1", "worker-2"), config) + r := reconcilerFor(t, objects...) + + // The cluster appears between this run's read and its create, which is what + // the fake client's own AlreadyExists then reports. + if err := r.Create(context.Background(), aCluster(nil)); err != nil { + t.Fatalf("seeding the cluster: %v", err) + } + + _, err := r.createCluster(context.Background(), config) + if err == nil { + t.Fatal("a create that found the cluster already there reported success") + } + var refusal *refusedError + if !errors.As(err, &refusal) || refusal.reason != ClusterExists { + t.Errorf("the failure is %v, want a ClusterExists refusal", err) + } +} + +// OpenShift configures the kubelet. The renderer reads an unset flag as skipping +// it, so leaving the environment's resolution silent changed what an OpenShift +// deployment does. +func TestOpenShiftStatesItsKubeletFlag(t *testing.T) { + config := aDocument(func(c *simplyblockv1alpha2.ClusterDeploymentConfig) { + c.Spec.Environment = simplyblockv1alpha2.KubernetesEnvironmentOpenShift + }) + r := reconcilerFor(t) + + workload := r.buildWorkload(config) + if workload.EnableKubeletConfiguration == nil { + t.Fatal("OpenShift left the kubelet flag unstated, which the renderer reads as skip") + } + if !*workload.EnableKubeletConfiguration { + t.Error("OpenShift resolved to skipping the kubelet configuration") + } +} + +// A growth document's nodes join a cluster that is already serving, which the +// control plane reads as a request to rebalance onto them. +func TestAGrowthDocumentMarksItsNodesAsAnExpansion(t *testing.T) { + config := aDocument(func(c *simplyblockv1alpha2.ClusterDeploymentConfig) { + c.Spec.ClusterRef = theCluster + c.Spec.Cluster = nil + }) + config.Status.ClusterRef = theCluster + objects := append(workers("worker-1", "worker-2"), config, aCluster(nil)) + r := reconcilerFor(t, objects...) + + if _, err := r.createNodes(context.Background(), config); err != nil { + t.Fatalf("createNodes: %v", err) + } + + var nodes simplyblockv1alpha2.StorageNodeList + if err := r.List(context.Background(), &nodes); err != nil { + t.Fatalf("listing the nodes: %v", err) + } + for i := range nodes.Items { + expand := nodes.Items[i].Spec.Config.Expand + if expand == nil || !*expand { + t.Errorf("node %s is not marked as an expansion", nodes.Items[i].Name) + } + } +} + +// A document that creates its own cluster is the initial layout, so its nodes +// leave the flag unset rather than asking for a rebalance onto themselves. +func TestAnInitialDocumentDoesNotMarkAnExpansion(t *testing.T) { + config := aDocument(nil) + config.Status.ClusterRef = theCluster + objects := append(workers("worker-1", "worker-2"), config, aCluster(nil)) + r := reconcilerFor(t, objects...) + + if _, err := r.createNodes(context.Background(), config); err != nil { + t.Fatalf("createNodes: %v", err) + } + + var nodes simplyblockv1alpha2.StorageNodeList + if err := r.List(context.Background(), &nodes); err != nil { + t.Fatalf("listing the nodes: %v", err) + } + for i := range nodes.Items { + if nodes.Items[i].Spec.Config.Expand != nil { + t.Errorf("node %s asks for a rebalance onto an initial layout", + nodes.Items[i].Name) + } + } +} + +// The record is rebuilt from the slots the document describes, so a pass that +// created a node and failed to persist the reference still reports it. +func TestTheNodeRecordSurvivesALostStatusWrite(t *testing.T) { + config := aDocument(nil) + config.Status.ClusterRef = theCluster + objects := append(workers("worker-1", "worker-2"), config, aCluster(nil)) + r := reconcilerFor(t, objects...) + + if _, err := r.createNodes(context.Background(), config); err != nil { + t.Fatalf("createNodes: %v", err) + } + recorded := len(config.Status.NodeRefs) + + // The status write is lost, and the next pass finds both nodes already there. + config.Status.NodeRefs = nil + if _, err := r.createNodes(context.Background(), config); err != nil { + t.Fatalf("the second pass: %v", err) + } + + if len(config.Status.NodeRefs) != recorded { + t.Errorf("the record holds %d names after the lost write, want the %d it created", + len(config.Status.NodeRefs), recorded) + } +} + +// A worker in two groups is a document the expansion cannot honor: the first +// group creates its nodes and the second group's devices never reach them. +func TestAWorkerInTwoGroupsIsReported(t *testing.T) { + config := aDocument(func(c *simplyblockv1alpha2.ClusterDeploymentConfig) { + c.Spec.Approved = false + c.Spec.NodeSets[0].Groups = append(c.Spec.NodeSets[0].Groups, + simplyblockv1alpha2.NodeGroup{ + Name: "second", + Workers: []string{"worker-2"}, + Devices: &simplyblockv1alpha2.DeviceSelection{NVMe: []string{"0000:88:00.0"}}, + }) + }) + objects := append(workers("worker-1", "worker-2"), config) + r := reconcilerFor(t, objects...) + + findings, err := r.validate(context.Background(), config) + if err != nil { + t.Fatalf("validate: %v", err) + } + if !saidSomethingAbout(findings, "worker-2") { + t.Errorf("nothing reported the repeated worker: %+v", findings) + } +} + +// One DaemonSet serves every node of a cluster, so groups that name different +// interfaces describe something the expansion cannot build. +func TestGroupsThatDisagreeAboutInterfacesAreReported(t *testing.T) { + config := aDocument(func(c *simplyblockv1alpha2.ClusterDeploymentConfig) { + c.Spec.Approved = false + c.Spec.NodeSets[0].Groups = append(c.Spec.NodeSets[0].Groups, + simplyblockv1alpha2.NodeGroup{ + Name: "second", + Workers: []string{"worker-3"}, + MgmtInterface: "eth9", + Devices: &simplyblockv1alpha2.DeviceSelection{NVMe: []string{"0000:88:00.0"}}, + }) + }) + objects := append(workers("worker-1", "worker-2", "worker-3"), config) + r := reconcilerFor(t, objects...) + + findings, err := r.validate(context.Background(), config) + if err != nil { + t.Fatalf("validate: %v", err) + } + if !saidSomethingAbout(findings, "management interfaces") { + t.Errorf("nothing reported the disagreement: %+v", findings) + } +} + +// An approved document waiting on the control plane is Expanding. Draft is what +// the API calls one nobody has approved, and reporting it would tell every status +// consumer the deployment is still editable. +func TestAnApprovedDocumentWaitingIsExpanding(t *testing.T) { + config := aDocument(nil) + objects := append(workers("worker-1", "worker-2"), config) + r := reconcilerFor(t, objects...) + + // No ControlPlane exists, so the gate holds. + if _, err := r.Reconcile(context.Background(), requestFor(config)); err != nil { + t.Fatalf("Reconcile: %v", err) + } + + var fresh simplyblockv1alpha2.ClusterDeploymentConfig + if err := r.Get(context.Background(), client.ObjectKeyFromObject(config), &fresh); err != nil { + t.Fatalf("reading the document: %v", err) + } + if fresh.Status.Phase != simplyblockv1alpha2.ClusterDeploymentConfigPhaseExpanding { + t.Errorf("the document is %q while it waits, want Expanding", fresh.Status.Phase) + } +} + +// The marker a selector consumes is a label, because a label selector cannot see +// an annotation and the marker exists for nothing else. +func TestReadyToDeployIsALabel(t *testing.T) { + config := aDocument(nil) + objects := append(workers("worker-1", "worker-2"), config) + r := reconcilerFor(t, objects...) + + if _, err := r.Reconcile(context.Background(), requestFor(config)); err != nil { + t.Fatalf("Reconcile: %v", err) + } + + var fresh simplyblockv1alpha2.ClusterDeploymentConfig + if err := r.Get(context.Background(), client.ObjectKeyFromObject(config), &fresh); err != nil { + t.Fatalf("reading the document: %v", err) + } + if fresh.Labels[readyToDeploy] != readyToDeployValue { + t.Errorf("the approved document carries labels %v, want the marker", fresh.Labels) + } + if _, asAnnotation := fresh.Annotations[readyToDeploy]; asAnnotation { + t.Error("the marker is an annotation, which no selector can see") + } +} + +// requestFor addresses one document. +func requestFor(config *simplyblockv1alpha2.ClusterDeploymentConfig) ctrl.Request { + return ctrl.Request{NamespacedName: client.ObjectKeyFromObject(config)} +} + +// saidSomethingAbout reports whether any finding's message names the substring. +func saidSomethingAbout(findings []finding, want string) bool { + for _, found := range findings { + if strings.Contains(found.message, want) { + return true + } + } + return false +} diff --git a/operator/internal/controllers/deployment/operatorops_controller.go b/operator/internal/controllers/deployment/operatorops_controller.go index 461b39d10..11237b3f9 100644 --- a/operator/internal/controllers/deployment/operatorops_controller.go +++ b/operator/internal/controllers/deployment/operatorops_controller.go @@ -42,7 +42,9 @@ import ( logf "sigs.k8s.io/controller-runtime/pkg/log" "github.com/simplyblock/atlas/inventory" + "github.com/simplyblock/atlas/ptr" simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + discoverypkg "github.com/simplyblock/simplyblock-operator/internal/discovery" "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" ) @@ -206,14 +208,33 @@ func (r *OperatorOpsReconciler) inspect( return ctrl.Result{}, err } + // A worker that looked fine and was not used owes the run a reason. Both + // exclusions below are silent otherwise, and a draft missing three machines + // somebody expected is a draft they have no way to ask about. + useControlPlane := ptr.BoolFromOrFalse(spec.EnableControlPlaneNodes) workers := make([]string, 0, len(nodes.Items)) for _, node := range nodes.Items { - if !schedulable(node) { + if _, already := taken[node.Name]; already { continue } - if _, already := taken[node.Name]; already { + // The role is not derivable from the taints. Kubernetes taints its + // control-plane nodes and OpenShift usually does not taint its + // infrastructure ones, so an infra node passes every check the run had + // before this one and its being the storage tier went unnoticed. + role := discoverypkg.RoleOf(node) + if !UsableWorker(node, useControlPlane) { + r.event(ops, corev1.EventTypeNormal, "WorkerDeclined", + fmt.Sprintf("%s is not used: %s", node.Name, declinedBecause(node, role))) continue } + if !role.HoldsStorageNodes() { + // The reviewer asked for these and still has to see which machines + // they got, because the draft's control-plane node set is otherwise + // just another block of hostnames. + r.event(ops, corev1.EventTypeWarning, "ControlPlaneNodeIncluded", fmt.Sprintf( + "%s is %s and is in the draft because spec.discover.enableControlPlaneNodes is set", + node.Name, role.Describe())) + } workers = append(workers, node.Name) } slices.Sort(workers) @@ -230,8 +251,9 @@ func (r *OperatorOpsReconciler) inspect( if len(workers) == 0 { return r.fail(ctx, ops, - "no schedulable worker is free: every node either carries a StorageNode already, "+ - "is unschedulable, or does not match the run's selector") + "no worker is free: every node either carries a StorageNode already, is "+ + "unschedulable, is reserved for the control plane or for infrastructure, "+ + "or does not match the run's selector") } ops.Status.Workers = workers @@ -639,6 +661,46 @@ func (r *OperatorOpsReconciler) event( // A cordoned node and a node carrying a NoSchedule taint are both excluded: a // probe Job is pinned with spec.nodeName and would run on either, and a worker // the cluster is not scheduling to is not one to hand to a storage cluster. +// UsableWorker reports whether a discovery run would inspect this machine. +// +// It is exported so that the one thing deciding whether a run is worth raising at +// all reads the same predicate the run itself applies. A bootstrap that raised a +// run against a cluster with nothing to inspect would create an object whose only +// outcome is to fail, and on a namespace delete that object holds a finalizer the +// operator may no longer be alive to clear. +func UsableWorker(node corev1.Node, useControlPlane bool) bool { + if !schedulable(node) { + return false + } + role := discoverypkg.RoleOf(node) + return role.HoldsStorageNodes() || useControlPlane +} + +// declinedBecause says why UsableWorker refused the machine, so the event a +// reviewer reads names the condition rather than only the outcome. +func declinedBecause(node corev1.Node, role discoverypkg.NodeRole) string { + if !schedulable(node) { + return unschedulableReason(node) + } + return fmt.Sprintf("it is %s; set spec.discover.enableControlPlaneNodes to include it", + role.Describe()) +} + +// unschedulableReason says which of the two conditions excluded the node, so an +// administrator reads "it is cordoned" rather than a bare refusal. +func unschedulableReason(node corev1.Node) string { + if node.Spec.Unschedulable { + return "it is cordoned" + } + for _, taint := range node.Spec.Taints { + if taint.Effect == corev1.TaintEffectNoSchedule || taint.Effect == corev1.TaintEffectNoExecute { + return fmt.Sprintf("it carries the taint %s=%s:%s", + taint.Key, taint.Value, taint.Effect) + } + } + return "it is not schedulable" +} + func schedulable(node corev1.Node) bool { if node.Spec.Unschedulable { return false diff --git a/operator/internal/controllers/deployment/operatorops_unit_test.go b/operator/internal/controllers/deployment/operatorops_unit_test.go index a3a9ef84d..fda527111 100644 --- a/operator/internal/controllers/deployment/operatorops_unit_test.go +++ b/operator/internal/controllers/deployment/operatorops_unit_test.go @@ -315,7 +315,7 @@ func TestDiscoverFailsWhenNoWorkerIsFree(t *testing.T) { if ops.Status.Phase != simplyblockv1alpha2.OperatorOpsPhaseFailed { t.Fatalf("the run is %q, want Failed", ops.Status.Phase) } - if !strings.Contains(ops.Status.Message, "no schedulable worker") { + if !strings.Contains(ops.Status.Message, "no worker is free") { t.Errorf("the message is %q, and it does not say what was wrong", ops.Status.Message) } } diff --git a/operator/internal/controllers/deployment/validation.go b/operator/internal/controllers/deployment/validation.go new file mode 100644 index 000000000..c5b71037e --- /dev/null +++ b/operator/internal/controllers/deployment/validation.go @@ -0,0 +1,331 @@ +// What a draft is checked against, and why the check runs on every pass. +// +// This is the only check a device gets before the deployment runs, because +// admission does not look at devices (§5.1). A config discovery wrote cannot be +// wrong about them, since the list came from the inspection; a hand-written one +// can, and saying so while the document is still editable is the whole value of +// the review gate. Approving without reading it is how a device mistake becomes an +// immutable document, and no mechanism below the reviewer prevents that. +// +// Every finding is a statement about the world rather than about the document's +// structure. The structural rules — one device class per document, a selection +// naming one member, immutability after approval — are CEL on the type, because +// they hold for a draft as firmly as for an approval and a draft that breaks one +// should not be storable at all. +// +// design-clusterdeploymentconfig.md §4.1 and §9.1 are the specification. + +package deployment + +import ( + "context" + "fmt" + "sort" + "strings" + + corev1 "k8s.io/api/core/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" + "sigs.k8s.io/controller-runtime/pkg/client" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// finding is one thing validation found, carrying the reason its event is raised +// under so the two cannot drift. +type finding struct { + reason string + message string +} + +// validate reads the document against the cluster it would act on and returns +// everything wrong with it. +// +// It returns every finding rather than the first, because a reviewer fixing a +// document wants the whole list: a draft that names three missing workers should +// say so once rather than over three edits. +func (r *ClusterDeploymentConfigReconciler) validate( + ctx context.Context, config *simplyblockv1alpha2.ClusterDeploymentConfig, +) ([]finding, error) { + var findings []finding + + missing, err := r.missingWorkers(ctx, config) + if err != nil { + return nil, err + } + if len(missing) > 0 { + findings = append(findings, finding{ + reason: WorkerNotFound, + message: fmt.Sprintf("%s %s not a node of this Kubernetes cluster", + strings.Join(missing, ", "), plural(len(missing), "is", "are")), + }) + } + + if found := r.groupsWithoutDevices(config); len(found) > 0 { + findings = append(findings, finding{ + reason: DeviceNotFound, + message: fmt.Sprintf("group %s %s no devices, so its workers hand nothing to simplyblock", + strings.Join(found, ", "), plural(len(found), "names", "name")), + }) + } + + mismatch, err := r.deviceClassMismatch(ctx, config) + if err != nil { + return nil, err + } + if mismatch != "" { + findings = append(findings, finding{reason: DeviceClassMismatch, message: mismatch}) + } + + if found := r.missingFailureDomains(ctx, config); found != "" { + findings = append(findings, finding{reason: DeviceNotFound, message: found}) + } + + if found := duplicateWorkers(config); len(found) > 0 { + findings = append(findings, finding{ + reason: WorkerNotFound, + message: fmt.Sprintf( + "%s %s in more than one group, and a storage node is identified by its "+ + "worker and slot alone, so only the first group's devices and settings "+ + "would reach it", + strings.Join(found, ", "), plural(len(found), "appears", "appear")), + }) + } + + if found := conflictingInterfaces(config); found != "" { + findings = append(findings, finding{reason: WorkerNotFound, message: found}) + } + return findings, nil +} + +// duplicateWorkers names every worker the document lists in more than one group. +// +// The schema permits it and the expansion cannot honor it: a StorageNode is +// identified by its cluster, its worker, and its slot, so the first group to reach +// a worker creates its nodes and every later group's device list, fault group, and +// memory setting is discarded by the create that finds one already there. Which +// group wins is the document's order, which is not a thing anybody chose. +func duplicateWorkers(config *simplyblockv1alpha2.ClusterDeploymentConfig) []string { + seen := map[string]int{} + for _, set := range config.Spec.NodeSets { + for _, group := range set.Groups { + for _, worker := range group.Workers { + seen[worker]++ + } + } + } + + repeated := map[string]struct{}{} + for worker, count := range seen { + if count > 1 { + repeated[worker] = struct{}{} + } + } + return sortedKeys(repeated) +} + +// conflictingInterfaces reports groups that name different network interfaces. +// +// A DaemonSet is one object for every node it schedules and its pod template +// cannot differ per group, so the interfaces are the cluster's whichever group +// states them. A document whose groups disagree therefore describes something the +// expansion cannot build, and taking the first silently would bind every node to +// one group's network while the document said otherwise. +func conflictingInterfaces(config *simplyblockv1alpha2.ClusterDeploymentConfig) string { + mgmt := map[string]struct{}{} + data := map[string]struct{}{} + for _, set := range config.Spec.NodeSets { + for _, group := range set.Groups { + if group.MgmtInterface != "" { + mgmt[group.MgmtInterface] = struct{}{} + } + if len(group.DataInterfaces) > 0 { + data[strings.Join(group.DataInterfaces, ",")] = struct{}{} + } + } + } + + if len(mgmt) > 1 { + return fmt.Sprintf( + "the groups name different management interfaces (%s), and one DaemonSet "+ + "serves every node of a cluster, so they cannot differ", + strings.Join(sortedKeys(mgmt), ", ")) + } + if len(data) > 1 { + return fmt.Sprintf( + "the groups name different data interfaces (%s), and one DaemonSet serves "+ + "every node of a cluster, so they cannot differ", + strings.Join(sortedKeys(data), "; ")) + } + return "" +} + +// missingWorkers names every worker the document lists that is not a node of this +// Kubernetes cluster. +func (r *ClusterDeploymentConfigReconciler) missingWorkers( + ctx context.Context, config *simplyblockv1alpha2.ClusterDeploymentConfig, +) ([]string, error) { + var workers corev1.NodeList + if err := r.List(ctx, &workers); err != nil { + return nil, fmt.Errorf("listing the Kubernetes workers: %w", err) + } + present := make(map[string]struct{}, len(workers.Items)) + for i := range workers.Items { + present[workers.Items[i].Name] = struct{}{} + } + + missing := map[string]struct{}{} + for _, set := range config.Spec.NodeSets { + for _, group := range set.Groups { + for _, worker := range group.Workers { + if _, there := present[worker]; !there { + missing[worker] = struct{}{} + } + } + } + } + return sortedKeys(missing), nil +} + +// groupsWithoutDevices names every group that hands no devices over. +// +// The schema permits it, because a selection is optional and a group may +// legitimately be written before its devices are known. What it cannot be is +// approved that way: a node with no devices comes up carrying nothing. +func (r *ClusterDeploymentConfigReconciler) groupsWithoutDevices( + config *simplyblockv1alpha2.ClusterDeploymentConfig, +) []string { + var found []string + for _, set := range config.Spec.NodeSets { + for _, group := range set.Groups { + if len(devicesOf(group)) == 0 { + found = append(found, set.Name+"/"+group.Name) + } + } + } + return found +} + +// deviceClassMismatch reports a growth document whose groups name devices of a +// class its cluster is not built out of. +// +// It applies only to a document that names an existing cluster. For one that +// creates its own, the class is whatever the groups say and there is nothing to +// disagree with — which is why the expansion stamps it rather than checking it. +func (r *ClusterDeploymentConfigReconciler) deviceClassMismatch( + ctx context.Context, config *simplyblockv1alpha2.ClusterDeploymentConfig, +) (string, error) { + if config.Spec.ClusterRef == "" { + return "", nil + } + + var cluster simplyblockv1alpha2.StorageCluster + key := client.ObjectKey{Namespace: config.Namespace, Name: config.Spec.ClusterRef} + switch err := r.Get(ctx, key, &cluster); { + case apierrors.IsNotFound(err): + // The expansion refuses this with ClusterNotFound, and saying it twice + // would put two findings on one fact. + return "", nil + case err != nil: + return "", fmt.Errorf("reading StorageCluster %s: %w", config.Spec.ClusterRef, err) + } + + stated := deviceClassOf(config) + if stated == "" { + return "", nil + } + + held := cluster.Spec.DeviceClass + if held == "" { + // An unstated class is NVMe, which is what the cluster's own field + // defaults to and what describes every cluster predating the field. + held = simplyblockv1alpha2.StorageClusterDeviceClassNVMe + } + if stated == held { + return "", nil + } + return fmt.Sprintf( + "the groups name %s devices and cluster %s is built out of %s; "+ + "an erasure-coding stripe placed across both classes is written at the slower one's rate", + stated, cluster.Name, held), nil +} + +// missingFailureDomains reports a document whose cluster requires fault groups and +// whose groups do not all declare one. +// +// Provisioning holds on this rather than failing, so a document approved without +// them produces nodes that sit and wait. Saying so at draft time is what turns +// that into an edit rather than an investigation. +func (r *ClusterDeploymentConfigReconciler) missingFailureDomains( + ctx context.Context, config *simplyblockv1alpha2.ClusterDeploymentConfig, +) string { + if !r.failureDomainsRequired(ctx, config) { + return "" + } + + var found []string + for _, set := range config.Spec.NodeSets { + for _, group := range set.Groups { + if group.FailureDomain == "" { + found = append(found, set.Name+"/"+group.Name) + } + } + } + if len(found) == 0 { + return "" + } + return fmt.Sprintf( + "the cluster has failure domains enabled and group %s %s none; "+ + "provisioning holds until each declares one", + strings.Join(found, ", "), plural(len(found), "declares", "declare")) +} + +// failureDomainsRequired reads the flag from whichever cluster the document acts +// on: the template for one it creates, the live object for one it grows. +func (r *ClusterDeploymentConfigReconciler) failureDomainsRequired( + ctx context.Context, config *simplyblockv1alpha2.ClusterDeploymentConfig, +) bool { + if config.Spec.ClusterRef == "" { + template := config.Spec.Cluster + return template != nil && template.EnableFailureDomains != nil && + *template.EnableFailureDomains + } + + var cluster simplyblockv1alpha2.StorageCluster + key := client.ObjectKey{Namespace: config.Namespace, Name: config.Spec.ClusterRef} + if err := r.Get(ctx, key, &cluster); err != nil { + return false + } + return cluster.Spec.EnableFailureDomains != nil && *cluster.Spec.EnableFailureDomains +} + +// summarize renders the findings for status.message: one sentence, which is what +// the field is for, with the rest countable. +func summarize(findings []finding) string { + if len(findings) == 0 { + return "" + } + if len(findings) == 1 { + return findings[0].message + } + return fmt.Sprintf("%s (and %d other %s)", findings[0].message, + len(findings)-1, plural(len(findings)-1, "problem", "problems")) +} + +// sortedKeys renders a set as a stable list, so a message does not reorder itself +// between passes and churn the status. +func sortedKeys(set map[string]struct{}) []string { + out := make([]string, 0, len(set)) + for key := range set { + out = append(out, key) + } + sort.Strings(out) + return out +} + +// plural picks the word for a count, so a message reads as a sentence. +func plural(n int, one, many string) string { + if n == 1 { + return one + } + return many +} diff --git a/operator/internal/discovery/kubenode.go b/operator/internal/discovery/kubenode.go index 6edb279b2..75f0c7104 100644 --- a/operator/internal/discovery/kubenode.go +++ b/operator/internal/discovery/kubenode.go @@ -56,6 +56,18 @@ type KubeNode struct { // that looked fine was not used. Unschedulable bool Taints []string + + // Role is what the machine is for, read off its node-role labels. It is not + // derivable from the taints beside it: Kubernetes taints its control-plane + // nodes and OpenShift usually does not taint its infrastructure ones, so a + // fleet's infra nodes pass a taint check and would otherwise land in a draft + // as storage workers. + Role NodeRole + + // Roles is every role the labels name, which is how a combined deployment + // marks a machine that is both a control-plane node and a worker. Role is the + // most restrictive of them; this is what a draft reports. + Roles []NodeRole } // ReservedMemoryBytes is capacity less allocatable: the memory the kubelet @@ -94,6 +106,8 @@ func KubeNodeOf(node corev1.Node) KubeNode { CapacityCPUMilli: milliValue(node.Status.Capacity, corev1.ResourceCPU), AllocatableCPUMilli: milliValue(node.Status.Allocatable, corev1.ResourceCPU), Unschedulable: node.Spec.Unschedulable, + Role: RoleOf(node), + Roles: RolesOf(node), } for _, taint := range node.Spec.Taints { diff --git a/operator/internal/discovery/noderole.go b/operator/internal/discovery/noderole.go new file mode 100644 index 000000000..f52822606 --- /dev/null +++ b/operator/internal/discovery/noderole.go @@ -0,0 +1,183 @@ +// What a node is for, read off the labels a distribution puts on it. +// +// A storage node is an SPDK process holding a machine's disks, and which machines +// are meant to hold disks is a question about the cluster's own topology rather +// than about the hardware. +// +// The roles do not rank the way a first guess suggests. An OpenShift +// infrastructure node is the tier a cluster's own infrastructure runs on, and +// simplyblock storage is infrastructure: a fleet with disks in its infra nodes +// almost always meant those disks to be the storage, and on OpenShift they are +// also the nodes that do not count against a subscription's core limit. So an +// infra node with disks is preferred over a worker with disks rather than avoided. +// +// A control-plane node is the opposite. It runs the API server and etcd, and a +// data path on an etcd host is a placement almost nobody intends, except in the +// combined three-node and single-node deployments this product supports, which is +// what the opt-in exists for. +// +// The role is a label rather than a taint, and that distinction is why this file +// exists at all. A taint is the cluster refusing to schedule there, which the +// worker selection already honors. A label is the cluster saying what the machine +// is for, and the two do not coincide: Kubernetes taints its control-plane nodes +// and OpenShift usually does not taint its infrastructure ones, so a fleet's infra +// nodes pass every taint check and were invisible to this operator. +// +// The label keys are Kubernetes' own convention rather than this product's, so +// they keep their spelling: node-role.kubernetes.io/ is what every +// distribution writes and what `kubectl get nodes` prints in its ROLES column. + +package discovery + +import ( + "sort" + "strings" + + corev1 "k8s.io/api/core/v1" +) + +// NodeRole is what a machine is for. +type NodeRole string + +const ( + // NodeRoleWorker is an ordinary worker, which is what a node carrying no role + // label at all also is. It holds storage nodes, and is what a fleet with no + // infrastructure tier is made of. + NodeRoleWorker NodeRole = "Worker" + + // NodeRoleControlPlane runs the API server and etcd. Kubernetes taints these + // by default, so most fleets exclude them twice over, but a combined + // three-node or single-node deployment deliberately removes that taint and is + // the case the opt-in exists for. + NodeRoleControlPlane NodeRole = "ControlPlane" + + // NodeRoleInfra is OpenShift's infrastructure role: the registry, the router, + // and the monitoring stack. It is the role this file is really about, in two + // ways. OpenShift does not always taint it, so nothing else in the operator + // would have noticed it; and it is the tier storage belongs on, so noticing it + // is what lets a draft propose the placement a fleet intended. + NodeRoleInfra NodeRole = "Infra" +) + +// roleLabelPrefix is the convention every distribution writes a node's role +// under, and the one `kubectl get nodes` reads its ROLES column from. +const roleLabelPrefix = "node-role.kubernetes.io/" + +// The role names that appear after the prefix. `master` is the older spelling of +// `control-plane` and is still written by clusters installed before it changed, so +// both are read and neither is preferred. +const ( + roleNameControlPlane = "control-plane" + roleNameMaster = "master" + roleNameInfra = "infra" + roleNameWorker = "worker" +) + +// RoleOf is what the node is for, and Worker when it says nothing. +// +// A node may carry several role labels at once, which is how a combined +// deployment marks a machine that is both a control-plane node and a worker. The +// most restrictive wins: a machine that runs etcd is a machine that runs etcd +// whatever else it also does, so a storage node placed there is placed on an etcd +// host either way. Infra outranks Worker for the same reason read the other way — +// a machine labeled both is part of the infrastructure tier. +func RoleOf(node corev1.Node) NodeRole { + roles := RolesOf(node) + for _, role := range []NodeRole{NodeRoleControlPlane, NodeRoleInfra} { + for _, held := range roles { + if held == role { + return role + } + } + } + return NodeRoleWorker +} + +// RolesOf is every role the node's labels name, so a draft can report what a +// machine is rather than only what it was reduced to. +func RolesOf(node corev1.Node) []NodeRole { + seen := map[NodeRole]struct{}{} + for key := range node.Labels { + name, found := strings.CutPrefix(key, roleLabelPrefix) + if !found { + continue + } + switch name { + case roleNameControlPlane, roleNameMaster: + seen[NodeRoleControlPlane] = struct{}{} + case roleNameInfra: + seen[NodeRoleInfra] = struct{}{} + case roleNameWorker: + seen[NodeRoleWorker] = struct{}{} + default: + // A role this product does not know is not a role it should refuse. + // Fleets label machines for their own purposes, and a storage node + // belongs on one of those unless somebody says otherwise. + } + } + if len(seen) == 0 { + return []NodeRole{NodeRoleWorker} + } + + out := make([]NodeRole, 0, len(seen)) + for role := range seen { + out = append(out, role) + } + sort.Slice(out, func(i, j int) bool { return out[i] < out[j] }) + return out +} + +// HoldsStorageNodes reports whether a role is one a storage node belongs on +// without somebody asking for it. +// +// Infra and Worker both do. ControlPlane does not, and that is the conservative +// half of a decision with a real cost: a fleet that genuinely wants storage on its +// control-plane nodes has to say so. The other way round is worse, because the +// draft a reviewer rubber-stamps would put a data path on the machines running +// etcd, and a fifty-worker document is not one anybody reads closely enough to +// catch three of them in it. +func (r NodeRole) HoldsStorageNodes() bool { + return r == NodeRoleWorker || r == NodeRoleInfra +} + +// Preference orders the roles a draft proposes, lowest first. It is what puts the +// infrastructure tier ahead of the workers on a fleet that has both, so the node +// set a reviewer reads first is the one they most likely meant. +func (r NodeRole) Preference() int { + switch r { + case NodeRoleInfra: + return 0 + case NodeRoleWorker: + return 1 + default: + return 2 + } +} + +// NodeSetName is what a draft calls the node set holding this role's machines. +// +// Splitting the draft by role rather than mixing the machines into one set is what +// makes the preference actionable: a reviewer who wants only the infrastructure +// tier deletes a block, rather than moving hostnames between them. +func (r NodeRole) NodeSetName() string { + switch r { + case NodeRoleInfra: + return "infra" + case NodeRoleControlPlane: + return "control-plane" + default: + return DefaultNodeSetName + } +} + +// Describe renders the role for the note that says what a machine is. +func (r NodeRole) Describe() string { + switch r { + case NodeRoleControlPlane: + return "a control-plane node, which runs the API server and etcd" + case NodeRoleInfra: + return "an infrastructure node, which runs the registry, the router, and monitoring" + default: + return "a worker" + } +} diff --git a/operator/internal/discovery/noderole_test.go b/operator/internal/discovery/noderole_test.go new file mode 100644 index 000000000..4c8e1741f --- /dev/null +++ b/operator/internal/discovery/noderole_test.go @@ -0,0 +1,155 @@ +// What a node's role is read as, and why the answer is not its taints. +// +// The case that motivates the file is the OpenShift infrastructure node: it is +// labeled and usually not tainted, so every check the operator had before this +// passed it straight through and it became a storage worker. + +package discovery + +import ( + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" +) + +func labeled(labels map[string]string) corev1.Node { + return corev1.Node{ObjectMeta: metav1.ObjectMeta{Name: "worker-1", Labels: labels}} +} + +func TestTheRoleIsReadOffTheLabels(t *testing.T) { + for _, tc := range []struct { + name string + labels map[string]string + want NodeRole + }{ + { + name: "no label at all is a worker", + labels: nil, + want: NodeRoleWorker, + }, + { + name: "an explicit worker label", + labels: map[string]string{"node-role.kubernetes.io/worker": ""}, + want: NodeRoleWorker, + }, + { + name: "a control-plane node", + labels: map[string]string{"node-role.kubernetes.io/control-plane": ""}, + want: NodeRoleControlPlane, + }, + { + // The older spelling, still written by clusters installed before it + // changed. Reading only the new one would put storage nodes on the + // control plane of every such fleet. + name: "the older master spelling", + labels: map[string]string{"node-role.kubernetes.io/master": ""}, + want: NodeRoleControlPlane, + }, + { + name: "an OpenShift infrastructure node", + labels: map[string]string{"node-role.kubernetes.io/infra": ""}, + want: NodeRoleInfra, + }, + { + // A combined deployment marks one machine both ways. The most + // restrictive wins: a machine that runs etcd runs etcd whatever else + // it also does. + name: "both control-plane and worker", + labels: map[string]string{ + "node-role.kubernetes.io/control-plane": "", + "node-role.kubernetes.io/worker": "", + }, + want: NodeRoleControlPlane, + }, + { + // A fleet labels machines for its own purposes, and a storage node + // belongs on one of those unless somebody says otherwise. + name: "a role this product does not know", + labels: map[string]string{"node-role.kubernetes.io/gpu": ""}, + want: NodeRoleWorker, + }, + } { + t.Run(tc.name, func(t *testing.T) { + if got := RoleOf(labeled(tc.labels)); got != tc.want { + t.Errorf("RoleOf = %q, want %q", got, tc.want) + } + }) + } +} + +// The whole reason this is a label check rather than a taint check: an +// infrastructure node carries no taint on most OpenShift fleets, so every check +// the operator had before this one saw an ordinary worker. +func TestAnUntaintedInfraNodeIsRecognized(t *testing.T) { + node := labeled(map[string]string{"node-role.kubernetes.io/infra": ""}) + if len(node.Spec.Taints) != 0 { + t.Fatal("the fixture is tainted, which is not the case this is about") + } + + role := RoleOf(node) + if role != NodeRoleInfra { + t.Fatalf("RoleOf = %q, want Infra", role) + } + // It is the tier storage belongs on rather than one to avoid, so recognizing + // it is what lets a draft propose the placement the fleet intended. + if !role.HoldsStorageNodes() { + t.Error("an infrastructure node was refused, and it is where storage belongs") + } +} + +// Infra and Worker both hold storage nodes. ControlPlane does not, which is the +// conservative half of a decision with a real cost: a fleet that wants storage on +// its etcd hosts has to say so. +func TestOnlyTheControlPlaneIsHeldBack(t *testing.T) { + for role, want := range map[NodeRole]bool{ + NodeRoleWorker: true, + NodeRoleInfra: true, + NodeRoleControlPlane: false, + } { + if got := role.HoldsStorageNodes(); got != want { + t.Errorf("%s.HoldsStorageNodes() = %v, want %v", role, got, want) + } + } +} + +// The infrastructure tier is proposed ahead of the workers, because a fleet with +// disks in both almost always meant the infra nodes to be the storage. +func TestTheInfrastructureTierIsPreferred(t *testing.T) { + if NodeRoleInfra.Preference() >= NodeRoleWorker.Preference() { + t.Error("a worker is proposed ahead of an infrastructure node") + } + if NodeRoleWorker.Preference() >= NodeRoleControlPlane.Preference() { + t.Error("a control-plane node is proposed ahead of a worker") + } +} + +// A draft reports what a machine is rather than only what it was reduced to, so +// every role its labels name survives. +func TestEveryRoleIsReported(t *testing.T) { + node := labeled(map[string]string{ + "node-role.kubernetes.io/control-plane": "", + "node-role.kubernetes.io/worker": "", + }) + + roles := RolesOf(node) + if len(roles) != 2 { + t.Fatalf("RolesOf = %v, want both", roles) + } + seen := map[NodeRole]bool{} + for _, role := range roles { + seen[role] = true + } + if !seen[NodeRoleControlPlane] || !seen[NodeRoleWorker] { + t.Errorf("RolesOf = %v, want the control-plane and worker roles", roles) + } +} + +// The role lands on the KubeNode the planner reads, so a draft can report it. +func TestTheRoleReachesTheKubeNode(t *testing.T) { + node := labeled(map[string]string{"node-role.kubernetes.io/infra": ""}) + + if got := KubeNodeOf(node).Role; got != NodeRoleInfra { + t.Errorf("KubeNodeOf(...).Role = %q, want Infra", got) + } +} diff --git a/operator/internal/discovery/plan.go b/operator/internal/discovery/plan.go index 7a2b53d6f..316066ab4 100644 --- a/operator/internal/discovery/plan.go +++ b/operator/internal/discovery/plan.go @@ -44,7 +44,8 @@ type Planner struct { // Grouper puts the workers into groups. Nil is GroupByHardware. Grouper Grouper - // NodeSetBuilder organizes the groups. Nil is SingleNodeSet. + // NodeSetBuilder organizes the groups. Nil is SplitByRole, which puts the + // infrastructure tier in a node set of its own and ahead of the workers. NodeSetBuilder NodeSetBuilder // KubeNodes is what Kubernetes says about each worker, keyed by name, from @@ -171,7 +172,7 @@ func (p Planner) Plan(reports []nodeprobe.Report, filter *simplyblockv1alpha2.De } builder := p.NodeSetBuilder if builder == nil { - builder = SingleNodeSet{} + builder = SplitByRole{} } plan := Plan{Class: class} diff --git a/operator/internal/discovery/rolesets.go b/operator/internal/discovery/rolesets.go new file mode 100644 index 000000000..0ccba6b51 --- /dev/null +++ b/operator/internal/discovery/rolesets.go @@ -0,0 +1,95 @@ +// The node-set builder that splits a draft by what its machines are for. +// +// It is the default rather than SingleNodeSet because the split is what makes the +// role actionable. On a fleet with disks in both its infrastructure nodes and its +// workers, the infra nodes are almost always the ones somebody meant to be the +// storage: simplyblock storage is infrastructure, and on OpenShift those nodes do +// not count against a subscription's core limit. Proposing them first, in a block +// of their own, lets a reviewer take that placement by deleting the other block +// rather than by moving hostnames between them. +// +// A fleet with no infrastructure tier gets exactly what SingleNodeSet produced: +// one set named for what it is, holding every worker. + +package discovery + +import ( + "sort" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// SplitByRole puts each role's machines in a node set of its own, ordered so the +// infrastructure tier comes first. +type SplitByRole struct { + // SetName overrides the name of the worker set. The infrastructure and + // control-plane sets are named for their roles, because a reviewer deleting + // one of those is choosing a tier rather than a rack. + SetName string +} + +func (SplitByRole) Name() string { return "one node set per role" } + +func (b SplitByRole) Build(groups []Group) []simplyblockv1alpha2.NodeSet { + if len(groups) == 0 { + return nil + } + + // A group is the machines that hand over the same devices, and the grouper + // that produced it knew nothing about roles, so one group may hold machines + // of two. Splitting here rather than grouping by role first keeps the + // hardware grouping intact: two infra nodes with identical disks stay one + // group. + byRole := map[NodeRole][]Group{} + for _, group := range groups { + for role, workers := range splitWorkersByRole(group.Workers) { + split := group + split.Workers = workers + byRole[role] = append(byRole[role], split) + } + } + + roles := make([]NodeRole, 0, len(byRole)) + for role := range byRole { + roles = append(roles, role) + } + sort.Slice(roles, func(i, j int) bool { + if roles[i].Preference() != roles[j].Preference() { + return roles[i].Preference() < roles[j].Preference() + } + return roles[i] < roles[j] + }) + + out := make([]simplyblockv1alpha2.NodeSet, 0, len(roles)) + for _, role := range roles { + name := role.NodeSetName() + if role == NodeRoleWorker && b.SetName != "" { + name = b.SetName + } + set := simplyblockv1alpha2.NodeSet{ + Name: name, + Groups: make([]simplyblockv1alpha2.NodeGroup, 0, len(byRole[role])), + } + for _, group := range byRole[role] { + set.Groups = append(set.Groups, nodeGroupOf(group)) + } + out = append(out, set) + } + return out +} + +// splitWorkersByRole divides one hardware group's machines by what they are for, +// keeping each slice in the order the group had them. +func splitWorkersByRole(workers []Worker) map[NodeRole][]Worker { + out := map[NodeRole][]Worker{} + for _, worker := range workers { + role := worker.Kube.Role + if role == "" { + // The planner was given no node objects, so nothing said what the + // machine is, and a machine with no role label is a worker. + role = NodeRoleWorker + } + out[role] = append(out[role], worker) + } + return out +} diff --git a/operator/internal/discovery/rolesets_test.go b/operator/internal/discovery/rolesets_test.go new file mode 100644 index 000000000..ac8c9cd63 --- /dev/null +++ b/operator/internal/discovery/rolesets_test.go @@ -0,0 +1,129 @@ +// What the draft looks like when a fleet has more than one kind of machine. +// +// The property that matters is the one a reviewer acts on: the infrastructure +// tier is a block of its own, and it comes first. A reviewer who wants only those +// disks deletes the block below it, rather than moving hostnames between them. + +package discovery + +import ( + "testing" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// machine is a worker the planner has already chosen devices for. The devices +// themselves do not matter here: what these cases are about is which set a +// machine lands in, which its role decides and its disks do not. +func machine(name string, role NodeRole) Worker { + return Worker{Name: name, Kube: KubeNode{Name: name, Role: role}} +} + +// grouped is one hardware group holding the given machines. +func grouped(name string, workers ...Worker) Group { + return Group{Name: name, Workers: workers, Class: ClassNVMe} +} + +// A fleet with both tiers gets two blocks, and the infrastructure one is first. +func TestTheInfrastructureSetComesFirst(t *testing.T) { + sets := SplitByRole{}.Build([]Group{grouped("uniform", + machine("worker-1", NodeRoleWorker), + machine("infra-1", NodeRoleInfra), + machine("worker-2", NodeRoleWorker), + )}) + + if len(sets) != 2 { + t.Fatalf("built %d node sets, want one per role: %+v", len(sets), sets) + } + if sets[0].Name != "infra" { + t.Errorf("the first set is %q, want the infrastructure tier", sets[0].Name) + } + if sets[1].Name != DefaultNodeSetName { + t.Errorf("the second set is %q, want the workers", sets[1].Name) + } + + if got := workersOf(sets[0]); len(got) != 1 || got[0] != "infra-1" { + t.Errorf("the infrastructure set holds %v", got) + } + if got := workersOf(sets[1]); len(got) != 2 { + t.Errorf("the worker set holds %v, want both workers", got) + } +} + +// A fleet with no infrastructure tier gets exactly what it got before: one set, +// named for what it is. +func TestAFleetOfPlainWorkersGetsOneSet(t *testing.T) { + sets := SplitByRole{}.Build([]Group{grouped("uniform", + machine("worker-1", NodeRoleWorker), + machine("worker-2", NodeRoleWorker), + )}) + + if len(sets) != 1 { + t.Fatalf("built %d node sets, want one: %+v", len(sets), sets) + } + if sets[0].Name != DefaultNodeSetName { + t.Errorf("the set is %q, want %q", sets[0].Name, DefaultNodeSetName) + } +} + +// The hardware grouping survives the split: two infrastructure nodes with the +// same disks stay one group, which is what keeps a document short. +func TestTheHardwareGroupingSurvivesTheSplit(t *testing.T) { + sets := SplitByRole{}.Build([]Group{ + grouped("dense", machine("infra-1", NodeRoleInfra), machine("infra-2", NodeRoleInfra)), + grouped("sparse", machine("worker-1", NodeRoleWorker)), + }) + + if len(sets) != 2 { + t.Fatalf("built %d node sets: %+v", len(sets), sets) + } + infra := sets[0] + if len(infra.Groups) != 1 { + t.Fatalf("the infrastructure set has %d groups, want one for identical hardware", + len(infra.Groups)) + } + if len(infra.Groups[0].Workers) != 2 { + t.Errorf("the group holds %v, want both machines", infra.Groups[0].Workers) + } +} + +// One hardware group holding machines of two roles is split between the sets +// rather than landing in whichever one sorted first. +func TestAMixedGroupIsSplitBetweenTheSets(t *testing.T) { + sets := SplitByRole{}.Build([]Group{grouped("identical", + machine("infra-1", NodeRoleInfra), + machine("worker-1", NodeRoleWorker), + )}) + + if len(sets) != 2 { + t.Fatalf("built %d node sets: %+v", len(sets), sets) + } + for _, set := range sets { + if got := workersOf(set); len(got) != 1 { + t.Errorf("set %q holds %v, want the one machine of its role", set.Name, got) + } + } +} + +// A machine the planner was given no node object for is a worker, because a +// machine with no role label is one. +func TestAMachineWithNoKubeNodeIsAWorker(t *testing.T) { + sets := SplitByRole{}.Build([]Group{{ + Name: "unknown", + Workers: []Worker{{Name: "worker-1"}}, + Class: ClassNVMe, + }}) + + if len(sets) != 1 || sets[0].Name != DefaultNodeSetName { + t.Fatalf("built %+v, want the one worker set", sets) + } +} + +// workersOf flattens a set's groups to the hostnames in it. +func workersOf(set simplyblockv1alpha2.NodeSet) []string { + var out []string + for _, group := range set.Groups { + out = append(out, group.Workers...) + } + return out +} diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml index 695d12254..b35497da5 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml @@ -117,6 +117,30 @@ spec: description: Name is the StorageCluster's name. maxLength: 253 type: string + nodesPerSocket: + description: |- + NodesPerSocket is how many storage nodes run per NUMA socket. See + SocketsToUse, which it multiplies. + format: int32 + maximum: 8 + minimum: 1 + type: integer + socketsToUse: + description: |- + SocketsToUse restricts the deployment to selected NUMA sockets, and empty + means socket 0 alone. With NodesPerSocket it decides how many storage nodes + each worker runs, so a group of two workers on a two-socket layout expands + to four nodes. + + It is here rather than on a node set because it is immutable on the cluster + it lands on: the layout a fleet was built with is not one a later document + can vary, and a reviewer should see it before the cluster exists. + items: + maxLength: 16 + type: string + maxItems: 16 + type: array + x-kubernetes-list-type: set stripe: description: Stripe is the erasure-coding layout. properties: diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_operatorops.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_operatorops.yaml index a4d1697e3..1a09ed465 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_operatorops.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_operatorops.yaml @@ -183,6 +183,24 @@ spec: devices and require enableLogicalBlockDevices rule: (has(self.enableLogicalBlockDevices) && self.enableLogicalBlockDevices) || !(has(self.blockAllowList) || has(self.blockDenyList)) + enableControlPlaneNodes: + description: |- + EnableControlPlaneNodes lets the run consider machines that run the API + server and etcd. + + It is off by default because a storage node is a data path, and putting one + on an etcd host is a placement almost nobody intends. The approval gate is a + poor place to catch it: a fifty-worker draft is not a document anybody reads + closely enough to spot three control-plane nodes in it. A combined three-node + or single-node deployment is the case that wants it, and those are set up + deliberately. + + There is no field beside it for infrastructure nodes, because those are used + without asking: an OpenShift infra node is the tier a cluster's own + infrastructure runs on, and simplyblock storage is infrastructure. A fleet + with disks in its infra nodes meant those disks to be the storage, so a draft + proposes them ahead of the workers rather than leaving them out. + type: boolean nodeSelector: additionalProperties: type: string From 452d3936ea5153743b0ce40be4ad6896d295985f Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Tue, 15 Sep 2026 14:04:00 +0200 Subject: [PATCH 003/206] fix(operator): a discovery run that plans nothing says why The rules worked out a reason for every machine they dropped and the run reported a count. A reviewer reading "78 refusal(s)" cannot tell a fleet with no disks from one whose disks are bound to a userspace driver they could reclaim, or from one whose disks carry a partition table somebody left behind. All three were on the lab this was found on, and the run said none of them. The refusals now go out as events, and a per-worker explanation goes into the status message, which is what kubectl shows. Separating the reason from the noise needed a distinction the data did not carry: a machine presents sixteen network block devices and four disks, so device rules mark themselves as pre-filters and Plan.Explain folds in only the refusals about disks that were genuine candidates, deduplicated and counted. Co-Authored-By: Claude Opus 5 (1M context) --- .../deployment/operatorops_controller.go | 18 +- .../deployment/operatorops_unit_test.go | 171 ++++++++++++++++++ operator/internal/discovery/plan.go | 86 ++++++++- operator/internal/discovery/rules.go | 25 +++ 4 files changed, 297 insertions(+), 3 deletions(-) diff --git a/operator/internal/controllers/deployment/operatorops_controller.go b/operator/internal/controllers/deployment/operatorops_controller.go index 11237b3f9..f14921039 100644 --- a/operator/internal/controllers/deployment/operatorops_controller.go +++ b/operator/internal/controllers/deployment/operatorops_controller.go @@ -28,6 +28,7 @@ import ( "context" "fmt" "slices" + "strings" "time" batchv1 "k8s.io/api/batch/v1" @@ -419,8 +420,23 @@ func (r *OperatorOpsReconciler) write( plan := planner.Plan(collected, filter) if len(plan.NodeSets) == 0 { + // The rules worked out why every machine was dropped, and a run that + // reported only how many were dropped would throw that away: a reviewer + // reading "78 refusal(s)" cannot tell a fleet with no disks from a fleet + // whose disks are held by a driver they could reclaim. The refusals go + // out as events, and the worker-level ones go into the message as well, + // because the message is what `kubectl get operatorops` shows. + for _, refusal := range plan.RefusalLines() { + r.event(ops, corev1.EventTypeNormal, "DeviceDeclined", refusal) + } + why := plan.Explain() + if len(why) == 0 { + // No machine was refused by name, so the run had no worker to refuse. + return r.fail(ctx, ops, fmt.Sprintf( + "no worker has a device this run would use: %s", plan.Summary())) + } return r.fail(ctx, ops, fmt.Sprintf( - "no worker has a device this run would use: %s", plan.Summary())) + "no worker has a device this run would use: %s", strings.Join(why, "; "))) } config, notes := r.draftFor(ops, spec, plan) diff --git a/operator/internal/controllers/deployment/operatorops_unit_test.go b/operator/internal/controllers/deployment/operatorops_unit_test.go index fda527111..3375041b2 100644 --- a/operator/internal/controllers/deployment/operatorops_unit_test.go +++ b/operator/internal/controllers/deployment/operatorops_unit_test.go @@ -597,3 +597,174 @@ func TestDiscoverAddsItsFinalizerBeforeDoingAnything(t *testing.T) { t.Error("the run started before its finalizer was recorded") } } + +// heldReportConfigMap is a machine whose NVMe controllers are on a userspace +// driver. The kernel presents no disk for such a controller, so the report +// carries no device at all and the controllers say where the disks went. +// +// It is the state a machine is left in by a simplyblock deployment that has +// since been removed, which makes it the state a second discovery run on a lab +// finds rather than an exotic one. +func heldReportConfigMap(t *testing.T, node string, addresses ...string) *corev1.ConfigMap { + t.Helper() + + report := nodeprobe.Report{ + Version: nodeprobe.ReportVersion, + Node: node, + CPU: nodeprobe.CPU{ + OnlineCPUs: 12, PhysicalCores: 12, Sockets: 1, ThreadsPerCore: 1, + NUMANodes: []nodeprobe.NUMACPUs{{Node: 0, OnlineCPUs: []int{0, 1, 2, 3}, PhysicalCores: 12}}, + }, + HugePages: []nodeprobe.HugePagePool{{ + SizeBytes: 1 << 21, Total: 3584, Free: 3584, + NUMANodes: []nodeprobe.NUMAHugePages{{Node: 0, Total: 3584, Free: 3584}}, + }}, + } + for _, address := range addresses { + report.NVMeControllers = append(report.NVMeControllers, nodeprobe.Controller{ + Address: address, + Driver: "uio_pci_generic", + NUMANode: -1, + TakenByUserspace: true, + }) + } + + cm, err := nodeprobe.ConfigMap(opsNamespace, opsName, nil, report) + if err != nil { + t.Fatalf("render the report ConfigMap: %v", err) + } + return cm +} + +// A run that produces nothing owes the reason it produced nothing. +// +// The rules compute one: a worker whose controllers are held by a userspace +// driver is refused with the controllers and the driver named, which is the +// difference between a reviewer concluding the machines have no storage and +// knowing to reclaim them. That explanation was computed and then dropped, +// because the failure returned before anything reported a refusal, and what the +// reviewer was left with was a count. +func TestARunThatFindsNothingSaysWhy(t *testing.T) { + r := newRunner(t, discoverRun(nil), worker("worker-1"), worker("worker-2")) + + r.step() // start + r.step() // inspect + r.step() // probing: creates the Jobs + + for _, node := range []string{"worker-1", "worker-2"} { + cm := heldReportConfigMap(t, node, + "0000:00:02.0", "0000:00:03.0", "0000:00:04.0", "0000:00:05.0") + if err := r.client.Create(context.Background(), cm); err != nil { + t.Fatalf("write a report: %v", err) + } + } + + r.step() // probing: sees the reports, moves to Writing + _, ops := r.step() + + if ops.Status.Phase != simplyblockv1alpha2.OperatorOpsPhaseFailed { + t.Fatalf("the run is %q, want Failed: %s", ops.Status.Phase, ops.Status.Message) + } + if len(r.configs()) != 0 { + t.Error("a run with no usable device wrote a document anyway") + } + + // What the message has to carry is the reason, not the arithmetic of it. + for _, want := range []string{"uio_pci_generic", "0000:00:02.0", "worker-1"} { + if !strings.Contains(ops.Status.Message, want) { + t.Errorf("the failure does not mention %q: %s", want, ops.Status.Message) + } + } +} + +// partitionedReportConfigMap is a machine whose NVMe disks are on the kernel +// driver, whole, of the right class, and carrying a partition table somebody +// left on them. +// +// It is the other half of a lab that ran simplyblock before: the controllers +// were handed back to the kernel and the disks still hold the old table. +func partitionedReportConfigMap(t *testing.T, node string, addresses ...string) *corev1.ConfigMap { + t.Helper() + + const gb = uint64(1) << 30 + report := nodeprobe.Report{ + Version: nodeprobe.ReportVersion, + Node: node, + CPU: nodeprobe.CPU{ + OnlineCPUs: 12, PhysicalCores: 12, Sockets: 1, ThreadsPerCore: 1, + NUMANodes: []nodeprobe.NUMACPUs{{Node: 0, OnlineCPUs: []int{0, 1, 2, 3}, PhysicalCores: 12}}, + }, + HugePages: []nodeprobe.HugePagePool{{ + SizeBytes: 1 << 21, Total: 3584, Free: 3584, + NUMANodes: []nodeprobe.NUMAHugePages{{Node: 0, Total: 3584, Free: 3584}}, + }}, + } + for i, address := range addresses { + report.NVMeControllers = append(report.NVMeControllers, nodeprobe.Controller{ + Address: address, Driver: "nvme", NUMANode: -1, + }) + report.Devices = append(report.Devices, nodeprobe.Device{ + Name: fmt.Sprintf("nvme%dn1", i), + Path: fmt.Sprintf("/dev/nvme%dn1", i), + PCIAddress: address, + SizeBytes: 70 * gb, + Kind: string(blockdev.KindDisk), + Transport: string(blockdev.TransportNVMe), + NUMANode: -1, + Available: false, + Content: "Foreign", + Rejections: []nodeprobe.Rejection{{ + Reason: string(blockdev.ReasonPartitioned), + Detail: "GPT header at 4096 (LBA 1), MBR partition table at 446", + }}, + }) + // The machine also presents the devices nothing would ever take, which + // is what makes the reporting hard: they outnumber the disks that + // matter and they are refused for reasons nobody needs. + report.Devices = append(report.Devices, nodeprobe.Device{ + Name: fmt.Sprintf("nbd%d", i), Path: fmt.Sprintf("/dev/nbd%d", i), + Kind: string(blockdev.KindNetwork), NUMANode: -1, + }) + } + + cm, err := nodeprobe.ConfigMap(opsNamespace, opsName, nil, report) + if err != nil { + t.Fatalf("render the report ConfigMap: %v", err) + } + return cm +} + +// A worker refused one device at a time still owes the reason. +// +// Its worker-level refusal reads "no device of it survived the device rules." +// That is the arithmetic and not the reason. The reason is on the devices, and +// it has to reach the message past the devices refused for being loopback or +// network block devices, which outnumber it and explain nothing. +func TestARunRefusedDeviceByDeviceSaysWhichReasonMatters(t *testing.T) { + r := newRunner(t, discoverRun(nil), worker("worker-1")) + + r.step() // start + r.step() // inspect + r.step() // probing: creates the Jobs + + cm := partitionedReportConfigMap(t, "worker-1", + "0000:00:02.0", "0000:00:03.0", "0000:00:04.0", "0000:00:05.0") + if err := r.client.Create(context.Background(), cm); err != nil { + t.Fatalf("write a report: %v", err) + } + + r.step() // probing: sees the report, moves to Writing + _, ops := r.step() + + if ops.Status.Phase != simplyblockv1alpha2.OperatorOpsPhaseFailed { + t.Fatalf("the run is %q, want Failed: %s", ops.Status.Phase, ops.Status.Message) + } + if !strings.Contains(ops.Status.Message, string(blockdev.ReasonPartitioned)) { + t.Errorf("the failure does not say the disks are partitioned: %s", ops.Status.Message) + } + // The devices that were never candidates are noise, and a message they + // reach is one nobody finishes reading. + if strings.Contains(ops.Status.Message, "whole disk") { + t.Errorf("the failure reports devices that were never candidates: %s", ops.Status.Message) + } +} diff --git a/operator/internal/discovery/plan.go b/operator/internal/discovery/plan.go index 316066ab4..25cb15d87 100644 --- a/operator/internal/discovery/plan.go +++ b/operator/internal/discovery/plan.go @@ -87,6 +87,84 @@ func (p Plan) Summary() string { len(p.Workers), devices, p.Class, groups, len(p.NodeSets), len(p.Refusals)) } +// Explain says why the plan holds nothing, one line per worker. +// +// A worker is refused either as a whole — its controllers are on a userspace +// driver, so the kernel presents no disk — or one device at a time, and the two +// need different answers. The first is on the worker's own refusal. The second +// leaves a worker-level reason that is the arithmetic ("no device survived the +// rules") and puts the reason on the devices, so the device refusals are folded +// in behind it, deduplicated and counted. +// +// Refusals that were pre-filters are left out of both. A machine presents +// sixteen network block devices and four disks, and a line saying the sixteen +// were not whole disks is true, longer than the rest of the message, and not +// the answer to anything. +func (p Plan) Explain() []string { + type perWorker struct { + worker string + line string + reasons []string + counts map[string]int + } + + order := make([]string, 0, len(p.Refusals)) + byWorker := map[string]*perWorker{} + at := func(worker string) *perWorker { + if found, ok := byWorker[worker]; ok { + return found + } + fresh := &perWorker{worker: worker, counts: map[string]int{}} + byWorker[worker] = fresh + order = append(order, worker) + return fresh + } + + for _, refusal := range p.Refusals { + if refusal.Device == "" { + at(refusal.Worker).line = refusal.Reason + continue + } + if refusal.PreFilter { + continue + } + entry := at(refusal.Worker) + if _, seen := entry.counts[refusal.Reason]; !seen { + entry.reasons = append(entry.reasons, refusal.Reason) + } + entry.counts[refusal.Reason]++ + } + + lines := make([]string, 0, len(order)) + for _, worker := range order { + entry := byWorker[worker] + if entry.line == "" && len(entry.reasons) == 0 { + // Every refusal on it was a pre-filter and the worker itself was + // admitted, so there is nothing about it to explain. + continue + } + detail := make([]string, 0, len(entry.reasons)) + for _, reason := range entry.reasons { + if count := entry.counts[reason]; count > 1 { + detail = append(detail, fmt.Sprintf("%d devices: %s", count, reason)) + continue + } + detail = append(detail, "1 device: "+reason) + } + + switch { + case len(detail) == 0: + lines = append(lines, entry.worker+": "+entry.line) + case entry.line == "": + lines = append(lines, entry.worker+": "+strings.Join(detail, ", ")) + default: + lines = append(lines, fmt.Sprintf("%s: %s (%s)", + entry.worker, entry.line, strings.Join(detail, ", "))) + } + } + return lines +} + // RefusalLines renders the refusals for a log or an event, one per line. func (p Plan) RefusalLines() []string { lines := make([]string, 0, len(p.Refusals)) @@ -235,10 +313,13 @@ func admitDevices(report nodeprobe.Report, rules []DeviceRule) ([]nodeprobe.Devi var refusals []Refusal for _, device := range report.Devices { - ok, rule, reason := true, "", "" + ok, rule, reason, pre := true, "", "", false for _, r := range rules { if admit, why := r.Admit(report, device); !admit { ok, rule, reason = false, r.Name(), why + if marker, says := r.(PreFilter); says { + pre = marker.PreFilter() + } break } } @@ -247,7 +328,8 @@ func admitDevices(report nodeprobe.Report, rules []DeviceRule) ([]nodeprobe.Devi continue } refusals = append(refusals, Refusal{ - Worker: report.Node, Device: device.Name, Rule: rule, Reason: reason, + Worker: report.Node, Device: device.Name, + Rule: rule, Reason: reason, PreFilter: pre, }) } return admitted, refusals diff --git a/operator/internal/discovery/rules.go b/operator/internal/discovery/rules.go index 8b527a1cc..25c1e4fda 100644 --- a/operator/internal/discovery/rules.go +++ b/operator/internal/discovery/rules.go @@ -39,6 +39,17 @@ type DeviceRule interface { Admit(report nodeprobe.Report, device nodeprobe.Device) (bool, string) } +// PreFilter marks a device rule that decides whether a device was ever a +// candidate. Refusing a loopback device for not being a whole disk is true and +// says nothing about why a run found no storage; refusing a disk because +// something else is using it is the answer. +// +// It is an optional interface rather than a method on DeviceRule because a rule +// that does not say is the common case and should not have to. +type PreFilter interface { + PreFilter() bool +} + // WorkerRule decides whether a worker takes part in the deployment. type WorkerRule interface { Name() string @@ -60,6 +71,12 @@ type Refusal struct { // Rule is the rule that declined it, and Reason is why. Rule string Reason string + + // PreFilter says the rule answers whether the thing was ever a candidate, + // rather than why a candidate was not taken. A machine presents dozens of + // loopback and network block devices and one disk somebody cares about, and + // a report that treats the two alike buries the second under the first. + PreFilter bool } // String renders a refusal for an event or a status message. @@ -153,6 +170,10 @@ type ClassRule struct { func (ClassRule) Name() string { return "device class" } +// PreFilter: a device of another class is not one this run was scanning for, so +// saying so explains nothing about the storage the fleet has. +func (ClassRule) PreFilter() bool { return true } + func (r ClassRule) Admit(_ nodeprobe.Report, device nodeprobe.Device) (bool, string) { if r.Class == ClassNVMe && device.Transport != string(blockdev.TransportNVMe) { return false, fmt.Sprintf("this run scans NVMe devices and the device is on %s", @@ -189,6 +210,10 @@ type WholeDiskRule struct{} func (WholeDiskRule) Name() string { return "whole disk" } +// PreFilter: a partition or a loopback device was never a disk this run could +// have taken, so refusing it explains nothing about the fleet's storage. +func (WholeDiskRule) PreFilter() bool { return true } + func (WholeDiskRule) Admit(_ nodeprobe.Report, device nodeprobe.Device) (bool, string) { if device.Kind != string(blockdev.KindDisk) { return false, fmt.Sprintf("it is a %s rather than a whole disk", device.Kind) From 67743278baf5b283f6576f4a82f77a359161ccbe Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Tue, 15 Sep 2026 14:11:15 +0200 Subject: [PATCH 004/206] feat(discovery): report the bound driver and whether anything is using it A controller reported only as takenByUserspace collapsed four states into two. The bool was derived from a driver name the report already carried, so it added nothing, and it lost the distinction that decides what to do next: uio_pci_generic and vfio-pci differ in what they suggest about who bound the device, and a controller bound to nothing at all read as healthy. It also said nothing about whether the binding is live. Bound and in use are different states, and only the second forbids reclaiming the disks. The holder need not be simplyblock either: a userspace binding is equally how a hypervisor passes a disk through to a guest, so a machine whose controllers are in use may be serving something this product knows nothing about, and a refusal reading as "leftovers, take them" would be an instruction to break it. So the wire carries the driver and an InUse the probe measures, pci.CheckHolders fills it in over the process table, and the classification stays a method rather than a field. A controller whose holders could not be read leaves InUse false and the failure is returned, not swallowed, because dropping it turns unknown into free, and that is the error that reclaims a running guest's disk. Collect routes it to the report's Unreadable list. ReportVersion goes to 2. Decode refuses an older report by design, so probes from before this are re-run rather than misread. Co-Authored-By: Claude Opus 5 (1M context) --- atlas-lib/README.md | 35 ++++++---- atlas-lib/inventory/inventory.go | 25 +++++-- atlas-lib/pci/pci_test.go | 65 +++++++++++++++++++ atlas-lib/pci/scan.go | 20 +++++- atlas-lib/pci/userspace.go | 51 +++++++++++++-- .../deployment/operatorops_unit_test.go | 7 +- operator/internal/discovery/rules.go | 31 +++++++-- operator/internal/discovery/rules_test.go | 38 ++++++++++- operator/internal/nodeprobe/collect.go | 16 ++--- operator/internal/nodeprobe/report.go | 45 +++++++++++-- 10 files changed, 284 insertions(+), 49 deletions(-) diff --git a/atlas-lib/README.md b/atlas-lib/README.md index e3db952d6..bbe5919e5 100644 --- a/atlas-lib/README.md +++ b/atlas-lib/README.md @@ -679,22 +679,35 @@ one mistake in this whole flow that loses data. The host's table is PID 1's. Everything else defaults sensibly; this one does not, and it is a field rather than a guess because only the caller knows where it mounted `/proc`. -**A worker's NVMe disks may be invisible to the disk reading entirely.** SPDK -takes a controller by rebinding it from the kernel's `nvme` driver to +**A worker's NVMe disks may be invisible to the disk reading entirely.** A +controller is taken by rebinding it from the kernel's `nvme` driver to `uio_pci_generic` or `vfio-pci`, and from that moment the kernel presents no -block device for it. On the fleet this was developed against, three of four -workers had four NVMe controllers each on `uio_pci_generic` and not one NVMe -block device between them — so a `class/block` scan reports "no NVMe disks" -about a machine with four. `inv.NVMeControllers` is the second half of the -answer, and `inv.ControllersTakenByUserspace()` is the question worth asking -whenever a draft came back empty: +block device for it. SPDK does this, and so does a hypervisor passing a disk +through to a guest, a DPDK application, and anything else driving hardware from +userspace. On the fleet this was developed against, three of four workers had +four NVMe controllers each on `uio_pci_generic` and not one NVMe block device +between them — so a `class/block` scan reports "no NVMe disks" about a machine +with four. `inv.NVMeControllers` is the second half of the answer, and +`inv.ControllersBoundToUserspace()` is the question worth asking whenever a +draft came back empty. + +Which driver is bound comes from sysfs; whether anything is still driving it +does not, and sysfs exports nothing that changes while a process holds the +character device — that was measured rather than assumed. `Collect` therefore +runs `pci.CheckHolders` and reports the answer per controller as `InUse`, and a +controller it could not check leaves `InUse` false and records why in the error +it returns. Dropping that error turns unknown into free, which is the one +mistake here that takes a running guest's disk away: ```go if len(inv.AvailableDevices()) == 0 { - if taken := inv.ControllersTakenByUserspace(); len(taken) > 0 { + if bound := inv.ControllersBoundToUserspace(); len(bound) > 0 { // Not a machine without storage: a machine whose storage something - // else is already driving. pci.HeldBy says which process, and - // pci.BindTo gives a controller back — refusing while anything holds it. + // else has. Which of the two answers applies is per controller — + // InUse false is a leftover that pci.BindTo can reclaim, and InUse + // true is a disk in service, which may belong to something this + // product knows nothing about. pci.HeldBy names the process, and + // pci.BindTo refuses while anything holds it either way. } } ``` diff --git a/atlas-lib/inventory/inventory.go b/atlas-lib/inventory/inventory.go index f23a77708..80c4c0852 100644 --- a/atlas-lib/inventory/inventory.go +++ b/atlas-lib/inventory/inventory.go @@ -226,14 +226,19 @@ func (i Inventory) AvailableDevices() []blockdev.Candidate { return free } -// ControllersTakenByUserspace is the NVMe controllers a userspace driver owns, +// ControllersBoundToUserspace is the NVMe controllers a userspace driver owns, // which are the disks this machine has and the kernel does not present. // // A discovery run that found no candidate devices should say whether this is // empty: no disks and no controllers is a machine with no storage, and no disks // with four controllers is a machine whose storage something else is already // driving. They are different answers and only one of them is a surprise. -func (i Inventory) ControllersTakenByUserspace() []pci.Device { +// +// Whether that something is still running is a separate question, answered per +// controller by [pci.Device.InUse]. Both belong in the refusal, because they +// ask a reviewer for opposite things: reclaim these disks, or leave the machine +// alone. +func (i Inventory) ControllersBoundToUserspace() []pci.Device { var taken []pci.Device for _, controller := range i.NVMeControllers { if controller.BoundToUserspace() { @@ -414,15 +419,25 @@ func Collect(ctx context.Context, cfg Config) (Inventory, error) { } inv.Devices = devices - controllers, err := pci.Scan(pci.Config{ + pciCfg := pci.Config{ SysfsRoot: cfg.sysfs(), ProcRoot: cfg.proc(), DevRoot: cfg.dev(), - }) + } + controllers, err := pci.Scan(pciCfg) if err != nil { errs = append(errs, fmt.Errorf("read the PCI controllers: %w", err)) } - inv.NVMeControllers = pci.NVMeControllers(controllers) + // Which driver is bound comes from sysfs and whether anything is driving it + // comes from the process table, and the second is the one that says whether + // a controller can be reclaimed. A failure to answer it is recorded rather + // than defaulted, because an unchecked controller and an idle one are + // indistinguishable once the error is dropped. + checked, err := pci.CheckHolders(pciCfg, pci.NVMeControllers(controllers)) + if err != nil { + errs = append(errs, err) + } + inv.NVMeControllers = checked if cfg.Kubernetes.Discovery != nil || len(cfg.Kubernetes.Nodes) > 0 { env, err := CollectEnvironment(ctx, cfg.Kubernetes.Discovery, cfg.Kubernetes.Nodes) diff --git a/atlas-lib/pci/pci_test.go b/atlas-lib/pci/pci_test.go index 3a846cbc1..406c02573 100644 --- a/atlas-lib/pci/pci_test.go +++ b/atlas-lib/pci/pci_test.go @@ -327,3 +327,68 @@ func TestUnbindOnADeviceWithNoDriverIsNotAFailure(t *testing.T) { t.Errorf("unbinding a device with no driver failed: %v", err) } } + +// CheckHolders answers the reclaimable question for a whole scan at once, which +// is the shape a probe needs: one pass over the controllers, one answer each. +func TestCheckHoldersMarksOnlyTheDeviceSomethingHolds(t *testing.T) { + h := takenWorker() + h.links["1234/fd/3"] = "/dev/uio2" + h.files["1234/comm"] = "qemu-system-x86_64" + + root := h.write(t) + cfg := Config{SysfsRoot: root, ProcRoot: root, DevRoot: "/dev"} + devices, err := Scan(cfg) + if err != nil { + t.Fatal(err) + } + + checked, err := CheckHolders(cfg, NVMeControllers(devices)) + if err != nil { + t.Fatalf("check the holders: %v", err) + } + if len(checked) != len(NVMeControllers(devices)) { + t.Fatalf("checked %d of %d controllers", len(checked), len(NVMeControllers(devices))) + } + + // The holder is a hypervisor rather than this product, which changes + // nothing: the question is whether anything is driving the disk. + held := nvmeSlot(t, checked, "0000:00:02.0") + if !held.InUse { + t.Error("a controller a process holds open is not marked in use") + } + idle := nvmeSlot(t, checked, "0000:00:03.0") + if idle.InUse { + t.Error("a controller nothing holds is marked in use") + } +} + +// The failure that matters: a process table that cannot be read leaves InUse +// false, and the only thing standing between that and a reclaimed disk is the +// error. So it is returned rather than swallowed, and the devices come back +// regardless, because a machine whose procfs is unreadable still has +// controllers worth reporting. +func TestCheckHoldersReportsWhatItCouldNotDetermine(t *testing.T) { + h := takenWorker() + root := h.write(t) + cfg := Config{SysfsRoot: root, ProcRoot: filepath.Join(root, "nonexistent"), DevRoot: "/dev"} + devices, err := Scan(Config{SysfsRoot: root, DevRoot: "/dev"}) + if err != nil { + t.Fatal(err) + } + + checked, err := CheckHolders(cfg, NVMeControllers(devices)) + if err == nil { + t.Fatal("an unreadable process table was reported as a clean check") + } + if !strings.Contains(err.Error(), "unknown") { + t.Errorf("the failure does not say the answer is unknown: %v", err) + } + if len(checked) == 0 { + t.Fatal("the controllers were dropped along with the answer") + } + for _, device := range checked { + if device.InUse { + t.Errorf("%s was marked in use by a check that never ran", device.Address) + } + } +} diff --git a/atlas-lib/pci/scan.go b/atlas-lib/pci/scan.go index 94f123b34..2f74c6810 100644 --- a/atlas-lib/pci/scan.go +++ b/atlas-lib/pci/scan.go @@ -118,13 +118,29 @@ type Device struct { // this was written against uio0 was the controller in slot 05.0 while uio2 // was the one in slot 02.0. UIODevices []string + + // InUse reports whether anything holds one of UIODevices open. + // + // [Scan] does not set it, because answering needs the process table and a + // scan reads sysfs. [CheckHolders] fills it in, and a false on a device + // neither of them looked at is the zero value rather than an answer: a + // caller that needs to tell the two apart has to know which produced the + // device, which is why the check records its failures rather than leaving + // this field to carry them. + InUse bool } // IsNVMe reports whether the device is an NVMe controller. func (d Device) IsNVMe() bool { return strings.HasPrefix(d.Class, classNVMePrefix) } -// BoundToUserspace reports whether a userspace-IO driver owns the device, which -// on this product's hosts means SPDK has taken it or something left it taken. +// BoundToUserspace reports whether a userspace-IO driver owns the device. +// +// It says which driver is bound and nothing about who is driving it. SPDK binds +// a controller this way, and so does a hypervisor passing a disk through to a +// guest, a DPDK application, and any other product that drives hardware from +// userspace; a binding left behind by something that has since exited looks the +// same as all of them. [HeldBy] is what separates those cases, and it is a +// different question with a different source. func (d Device) BoundToUserspace() bool { return d.Driver == DriverUIOGeneric || d.Driver == DriverVFIO } diff --git a/atlas-lib/pci/userspace.go b/atlas-lib/pci/userspace.go index 2d818e1a6..f8bbaf07c 100644 --- a/atlas-lib/pci/userspace.go +++ b/atlas-lib/pci/userspace.go @@ -1,10 +1,17 @@ // Whether anything is actually driving a device a userspace driver owns. // // Bound and in use are different states, and the difference decides whether a -// controller can be taken back. A controller SPDK is running on is bound and -// held; one a previous deployment left behind is bound and idle. Handing the -// first back to the kernel takes a storage node's disks out from under it, and -// handing the second back costs nothing. +// controller can be taken back. A controller something is running on is bound +// and held; one a previous deployment left behind is bound and idle. Handing +// the first back to the kernel takes its disks out from under whatever is +// driving them, and handing the second back costs nothing. +// +// What holds it is not assumed to be this product. A userspace binding is also +// how a hypervisor passes a disk through to a guest and how a DPDK application +// takes a device, and a machine that is doing either looks from sysfs exactly +// like one holding leftovers. That is the case this file exists to tell apart, +// and it is why the answer is about whether anything holds the device rather +// than about whether the holder is recognized. // // sysfs will not answer it. The uio driver exports name, version, and event, // and none of them changes while a process holds the character device: this was @@ -19,6 +26,7 @@ package pci import ( + "errors" "fmt" "os" "path/filepath" @@ -60,6 +68,41 @@ func HeldBy(cfg Config, device Device) ([]Holder, error) { return holdersOf(cfg, device.UIODevices) } +// CheckHolders fills in InUse for every device given, and returns what it could +// not determine alongside the devices it could. +// +// The failures are returned rather than folded into InUse because the two are +// not the same answer. A device nothing holds and a device that could not be +// checked both leave InUse false, and only one of them is safe to reclaim, so a +// caller that drops the error has quietly turned unknown into free. The +// devices come back either way: a machine whose process table could not be read +// still has controllers worth reporting. +func CheckHolders(cfg Config, devices []Device) ([]Device, error) { + out := make([]Device, 0, len(devices)) + var errs []error + + for _, device := range devices { + if !device.BoundToUserspace() { + // The kernel is driving it, so its namespaces are block devices and + // nothing about the process table changes that. + out = append(out, device) + continue + } + + holders, err := HeldBy(cfg, device) + if err != nil { + errs = append(errs, fmt.Errorf( + "pci: %s could not be checked for holders, so whether it is free is unknown: %w", + device.Address, err)) + out = append(out, device) + continue + } + device.InUse = len(holders) > 0 + out = append(out, device) + } + return out, errors.Join(errs...) +} + // holdersOf walks the process table for anything holding one of the paths. func holdersOf(cfg Config, paths []string) ([]Holder, error) { proc := cfg.proc() diff --git a/operator/internal/controllers/deployment/operatorops_unit_test.go b/operator/internal/controllers/deployment/operatorops_unit_test.go index 3375041b2..e9b070d08 100644 --- a/operator/internal/controllers/deployment/operatorops_unit_test.go +++ b/operator/internal/controllers/deployment/operatorops_unit_test.go @@ -622,10 +622,9 @@ func heldReportConfigMap(t *testing.T, node string, addresses ...string) *corev1 } for _, address := range addresses { report.NVMeControllers = append(report.NVMeControllers, nodeprobe.Controller{ - Address: address, - Driver: "uio_pci_generic", - NUMANode: -1, - TakenByUserspace: true, + Address: address, + Driver: "uio_pci_generic", + NUMANode: -1, }) } diff --git a/operator/internal/discovery/rules.go b/operator/internal/discovery/rules.go index 25c1e4fda..0e25c3b32 100644 --- a/operator/internal/discovery/rules.go +++ b/operator/internal/discovery/rules.go @@ -310,14 +310,33 @@ func (WorkerHasDevices) Admit(report nodeprobe.Report, admitted []nodeprobe.Devi // A worker whose disks are on a userspace driver has no block devices at // all, so "no device survived the rules" is true and useless: the machine - // is full of disks that something else is already driving. Saying which - // controllers and which driver is the difference between a reviewer - // concluding the machine has no storage and knowing to reclaim it. - if taken := report.ControllersTakenByUserspace(); len(taken) > 0 { + // is full of disks that something else has. Saying which controllers and + // which driver is the difference between a reviewer concluding the machine + // has no storage and knowing what is on it. + if bound := report.ControllersBoundToUserspace(); len(bound) > 0 { + // Whether anything is driving them is the half that decides what to do + // next, and the two answers ask for opposite things. Nothing holding + // them means the binding is a leftover and the disks can be taken back. + // Something holding them means the machine is serving whatever that is, + // which need not be this product: a userspace binding is also how a + // disk is passed through to a guest. + var busy []nodeprobe.Controller + for _, controller := range bound { + if controller.InUse { + busy = append(busy, controller) + } + } + if len(busy) > 0 { + return false, fmt.Sprintf( + "it presents no usable block device, and %d of its NVMe controllers (%s) are "+ + "bound to a userspace driver and in use, so something is driving its disks", + len(busy), describeControllers(busy)) + } return false, fmt.Sprintf( "it presents no usable block device, and %d of its NVMe controllers (%s) are "+ - "held by a userspace driver, so the kernel presents no disk for them", - len(taken), describeControllers(taken)) + "bound to a userspace driver and nothing is using them, so the disks are "+ + "there to be reclaimed", + len(bound), describeControllers(bound)) } return false, "no device of it survived the device rules" } diff --git a/operator/internal/discovery/rules_test.go b/operator/internal/discovery/rules_test.go index d99f37625..9e992d2af 100644 --- a/operator/internal/discovery/rules_test.go +++ b/operator/internal/discovery/rules_test.go @@ -214,8 +214,8 @@ func TestWorkerHasDevicesExplainsAMachineWhoseDisksAreAlreadyDriven(t *testing.T // full of disks. worker := report("worker-1") worker.NVMeControllers = []nodeprobe.Controller{ - {Address: "0000:00:02.0", Driver: "uio_pci_generic", TakenByUserspace: true}, - {Address: "0000:00:03.0", Driver: "uio_pci_generic", TakenByUserspace: true}, + {Address: "0000:00:02.0", Driver: "uio_pci_generic", InUse: true}, + {Address: "0000:00:03.0", Driver: "uio_pci_generic", InUse: true}, } ok, why := (WorkerHasDevices{}).Admit(worker, nil) @@ -312,3 +312,37 @@ func TestParseSizeRange(t *testing.T) { } } } + +// A controller nothing is using can be reclaimed, and one something is using +// cannot. The refusal has to say which, because the two ask a reviewer for +// opposite things: reclaim these disks, or leave that machine alone. +// +// The holder need not be simplyblock. vfio-pci is also how a disk is passed +// through to a guest, so a machine whose controllers are in use may be serving +// something this product knows nothing about, and a refusal that read as +// "leftovers, take them" would be an instruction to break it. +func TestWorkerHasDevicesSeparatesReclaimableFromInUse(t *testing.T) { + idle := report("worker-1") + idle.NVMeControllers = []nodeprobe.Controller{ + {Address: "0000:00:02.0", Driver: "uio_pci_generic"}, + {Address: "0000:00:03.0", Driver: "uio_pci_generic"}, + } + + _, why := (WorkerHasDevices{}).Admit(idle, nil) + if !strings.Contains(why, "nothing is using them") { + t.Errorf("an idle binding is not reported as reclaimable: %q", why) + } + + busy := report("worker-2") + busy.NVMeControllers = []nodeprobe.Controller{ + {Address: "0000:00:04.0", Driver: "vfio-pci", InUse: true}, + } + + _, why = (WorkerHasDevices{}).Admit(busy, nil) + if !strings.Contains(why, "in use") { + t.Errorf("a held controller is not reported as in use: %q", why) + } + if strings.Contains(why, "nothing is using them") { + t.Errorf("a held controller was offered for reclaiming: %q", why) + } +} diff --git a/operator/internal/nodeprobe/collect.go b/operator/internal/nodeprobe/collect.go index 07924d9bb..45101ca1d 100644 --- a/operator/internal/nodeprobe/collect.go +++ b/operator/internal/nodeprobe/collect.go @@ -150,12 +150,12 @@ func controllersOf(devices []pci.Device) []Controller { out := make([]Controller, 0, len(devices)) for _, device := range devices { out = append(out, Controller{ - Address: device.Address, - Driver: device.Driver, - Vendor: device.Vendor, - Product: device.Product, - NUMANode: device.NUMANode, - TakenByUserspace: device.BoundToUserspace(), + Address: device.Address, + Driver: device.Driver, + Vendor: device.Vendor, + Product: device.Product, + NUMANode: device.NUMANode, + InUse: device.InUse, }) } return out @@ -187,7 +187,7 @@ func sentences(err error) []string { func Summary(report Report) string { return fmt.Sprintf( "node %s: %d of %d block devices free, %d online CPUs over %d cores (hyperthreading %v), "+ - "%d MiB of %d MiB memory available, %d MiB of huge pages, %d interfaces, %d NVMe controllers taken by a userspace "+ + "%d MiB of %d MiB memory available, %d MiB of huge pages, %d interfaces, %d NVMe controllers bound to a userspace "+ "driver, %d readings unavailable", report.Node, len(report.AvailableDevices()), len(report.Devices), @@ -195,7 +195,7 @@ func Summary(report Report) string { report.Memory.AvailableBytes>>20, report.Memory.TotalBytes>>20, report.HugePageBytes()>>20, len(report.Interfaces), - len(report.ControllersTakenByUserspace()), + len(report.ControllersBoundToUserspace()), len(report.Unreadable), ) } diff --git a/operator/internal/nodeprobe/report.go b/operator/internal/nodeprobe/report.go index e11cbda30..ef1b39954 100644 --- a/operator/internal/nodeprobe/report.go +++ b/operator/internal/nodeprobe/report.go @@ -28,6 +28,8 @@ import ( "slices" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + + "github.com/simplyblock/atlas/pci" ) // ReportVersion is the schema version of the JSON below. @@ -36,7 +38,7 @@ import ( // because a probe pod outlives the operator that created it across an upgrade: // the image is pinned in the Job, and a Job already running keeps the image it // started with. -const ReportVersion = 1 +const ReportVersion = 2 // Report is one worker's inventory as the probe found it. type Report struct { @@ -259,6 +261,12 @@ type Controller struct { // Driver is what owns it: the kernel's own driver for a controller whose // namespaces it presents, uio_pci_generic or vfio-pci for one a userspace // driver has, and empty for one nothing owns. + // + // All four states matter and none of them is derivable from the others, + // which is why the driver is reported rather than a flag saying whether it + // is a userspace one. uio_pci_generic and vfio-pci in particular differ in + // what they suggest about who bound it: the second is also how a disk is + // passed through to a guest. Driver string `json:"driver,omitempty"` // Vendor and Product are the raw PCI identifiers, which is what sysfs has: @@ -269,17 +277,40 @@ type Controller struct { // NUMANode is the memory node it hangs off, or NUMANodeUnknown. NUMANode int `json:"numaNode"` - // TakenByUserspace reports whether a userspace-IO driver owns it, which on - // this product's hosts means SPDK has it or something left it taken. - TakenByUserspace bool `json:"takenByUserspace,omitempty"` + // InUse reports whether anything holds the controller open. + // + // It is the question Driver cannot answer and the one that decides whether + // a controller can be reclaimed: a userspace binding nothing is driving is + // a leftover, and the same binding with a process behind it is a disk in + // service, which may belong to a hypervisor guest or another product rather + // than to this one. + // + // False on a controller the probe could not check is the zero value and not + // an answer. A probe that failed to read the process table says so in + // Unreadable, so a reader deciding whether to reclaim has to find this + // report free of such an entry first. + InUse bool `json:"inUse,omitempty"` +} + +// BoundToUserspace reports whether a userspace-IO driver owns the controller. +// +// It reads the driver and says nothing about who bound it or whether anything +// is still driving it. InUse answers the second, and nothing answers the first, +// because a binding carries no record of what made it. +func (c Controller) BoundToUserspace() bool { + return c.Driver == pci.DriverUIOGeneric || c.Driver == pci.DriverVFIO } -// ControllersTakenByUserspace is the controllers no block device corresponds +// ControllersBoundToUserspace is the controllers no block device corresponds // to, which is the answer to why a worker full of disks reported none. -func (r Report) ControllersTakenByUserspace() []Controller { +// +// The classification lives here rather than at every call site because which +// drivers are userspace drivers is something this product knows and a string +// comparison spread across consumers would drift. +func (r Report) ControllersBoundToUserspace() []Controller { var taken []Controller for _, controller := range r.NVMeControllers { - if controller.TakenByUserspace { + if controller.BoundToUserspace() { taken = append(taken, controller) } } From 762f5f91bb55bc98f710829c97a69e084358fcb4 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Tue, 15 Sep 2026 14:52:01 +0200 Subject: [PATCH 005/206] feat(discovery): claim idle userspace controllers, and waive a partition table Three changes, each of which was a way for a run to report that a fleet full of disks has none. A controller bound to a userspace driver with nothing using it is a disk this fleet owns and nothing is driving, so the draft now proposes it. The kernel presents no block device for such a controller, so it could not reach a draft through the device reading at all -- but everything a NodeGroup needs to name an NVMe device is its PCI address, and that is the one thing the controller has. A held controller is still left alone, whoever holds it. What is lost with the block device is the disk's size and its content, so the group it lands in is named for its count rather than a capacity. The initial run waives a partition table. A machine that has held data before carries one on every disk, and the run meant to show a fleet what it has was reporting that it has nothing. The waiver stays narrow on its own terms: it admits a disk whose only refusal is the table, so a boot disk stays out on its mount and on the kernel refusing an exclusive open, which are refusals of their own. The probe is pulled on every run. It ships in the operator's image, so an operator deployed from a moving tag was replaced by a pull while its probes were not, and a node holding the previous layer kept running the previous probe. The report version then refused those reports and the run waited on machines that would never answer, which is how an upgrade produced a stalled discovery rather than a wrong one. Co-Authored-By: Claude Opus 5 (1M context) --- .../controllers/deployment/bootstrap.go | 24 +++++-- .../controllers/deployment/bootstrap_test.go | 49 +++++++++++-- .../deployment/operatorops_unit_test.go | 9 ++- operator/internal/discovery/grouping.go | 9 ++- operator/internal/discovery/plan.go | 61 ++++++++++++++++ operator/internal/discovery/plan_test.go | 72 +++++++++++++++++++ operator/internal/nodeprobe/job.go | 15 +++- operator/internal/nodeprobe/job_test.go | 49 +++++++++++++ 8 files changed, 272 insertions(+), 16 deletions(-) diff --git a/operator/internal/controllers/deployment/bootstrap.go b/operator/internal/controllers/deployment/bootstrap.go index fcc0cfefc..aac8b8ac7 100644 --- a/operator/internal/controllers/deployment/bootstrap.go +++ b/operator/internal/controllers/deployment/bootstrap.go @@ -53,6 +53,7 @@ import ( "sigs.k8s.io/controller-runtime/pkg/client" logf "sigs.k8s.io/controller-runtime/pkg/log" + "github.com/simplyblock/atlas/ptr" simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" ) @@ -108,11 +109,24 @@ func (d *InitialDiscovery) Start(ctx context.Context) error { }, Spec: simplyblockv1alpha2.OperatorOpsSpec{ Action: simplyblockv1alpha2.OperatorOpsActionDiscover, - // The run states no filter and no selector. What it produces is a - // draft of everything the fleet has, which is what a reviewer - // narrows: a guess at which disks somebody meant would be a guess - // they then have to find and undo (§8.1). - Discover: &simplyblockv1alpha2.DiscoverSpec{}, + // The run states no selector. What it produces is a draft of + // everything the fleet has, which is what a reviewer narrows: a + // guess at which disks somebody meant would be a guess they then + // have to find and undo (§8.1). + // + // The one thing it does state is that a partition table is not by + // itself a reason to leave a disk out. A machine that has held data + // before carries one on every disk, so refusing them makes the run + // meant to show a fleet what it has report that it has nothing. The + // waiver stays narrow on its own terms: it admits a disk whose only + // refusal is the table, and a disk that is also mounted, held by + // the kernel, or carrying swap is refused for those instead — which + // is what keeps a boot disk out of a draft nobody reads closely. + Discover: &simplyblockv1alpha2.DiscoverSpec{ + DeviceFilter: &simplyblockv1alpha2.DeviceFilter{ + EnablePartitionedDevices: ptr.To(true), + }, + }, }, } if err := d.Create(ctx, run); err != nil { diff --git a/operator/internal/controllers/deployment/bootstrap_test.go b/operator/internal/controllers/deployment/bootstrap_test.go index 78ed4e86d..46c589e53 100644 --- a/operator/internal/controllers/deployment/bootstrap_test.go +++ b/operator/internal/controllers/deployment/bootstrap_test.go @@ -57,13 +57,21 @@ func TestAFreshInstallRaisesOneDiscoveryRun(t *testing.T) { if run.Spec.Action != simplyblockv1alpha2.OperatorOpsActionDiscover { t.Errorf("action = %q, want Discover", run.Spec.Action) } - // The run states no filter and no selector: what it produces is a draft of - // everything the fleet has, which is what a reviewer narrows. + // The run narrows nothing: what it produces is a draft of everything the + // fleet has, which is what a reviewer narrows. The partition waiver it does + // state is the opposite of a guess at which disks somebody meant — it + // admits more rather than less, and what it admits is the ordinary state of + // a machine that has held data before. if run.Spec.Discover == nil { - t.Error("the run carries no discover block") - } else if len(run.Spec.Discover.NodeSelector) != 0 || - run.Spec.Discover.DeviceFilter != nil { - t.Errorf("the run guessed at a filter: %+v", run.Spec.Discover) + t.Fatal("the run carries no discover block") + } + if len(run.Spec.Discover.NodeSelector) != 0 { + t.Errorf("the run guessed at which machines: %+v", run.Spec.Discover) + } + if filter := run.Spec.Discover.DeviceFilter; filter != nil { + if len(filter.PcieAllowList) != 0 || len(filter.PcieDenyList) != 0 { + t.Errorf("the run guessed at which disks: %+v", filter) + } } } @@ -228,3 +236,32 @@ func TestACordonedFleetRaisesNoRun(t *testing.T) { t.Error("a run was raised against a fleet with no schedulable machine") } } + +// The initial run waives a partition table. +// +// A disk carrying one is the normal state of a machine that has held data +// before, and refusing every such disk makes the run that is supposed to show a +// fleet what it has report that it has nothing. The waiver is narrow on its own +// terms: it admits a disk whose only refusal is the table, so a boot disk stays +// out because its partition is mounted and the kernel will not hand it over, +// which are refusals of their own. +func TestTheInitialRunWaivesAPartitionTable(t *testing.T) { + d := discoveryFor(t, &corev1.Node{ObjectMeta: metav1.ObjectMeta{Name: "worker-1"}}) + + if err := d.Start(context.Background()); err != nil { + t.Fatalf("Start: %v", err) + } + + var run simplyblockv1alpha2.OperatorOps + key := client.ObjectKey{Namespace: theNamespace, Name: InitialDiscoveryName} + if err := d.Get(context.Background(), key, &run); err != nil { + t.Fatalf("reading the run: %v", err) + } + filter := run.Spec.Discover.DeviceFilter + if filter == nil || filter.EnablePartitionedDevices == nil { + t.Fatalf("the initial run states no partition waiver: %+v", run.Spec.Discover) + } + if !*filter.EnablePartitionedDevices { + t.Error("the initial run refuses a disk for carrying a partition table") + } +} diff --git a/operator/internal/controllers/deployment/operatorops_unit_test.go b/operator/internal/controllers/deployment/operatorops_unit_test.go index e9b070d08..c1cef5d07 100644 --- a/operator/internal/controllers/deployment/operatorops_unit_test.go +++ b/operator/internal/controllers/deployment/operatorops_unit_test.go @@ -602,9 +602,9 @@ func TestDiscoverAddsItsFinalizerBeforeDoingAnything(t *testing.T) { // driver. The kernel presents no disk for such a controller, so the report // carries no device at all and the controllers say where the disks went. // -// It is the state a machine is left in by a simplyblock deployment that has -// since been removed, which makes it the state a second discovery run on a lab -// finds rather than an exotic one. +// Something is driving them, which is what puts the machine out of reach: a +// binding nothing is using is a disk the draft claims, and only a held one is a +// disk in service. Whatever holds it need not be this product. func heldReportConfigMap(t *testing.T, node string, addresses ...string) *corev1.ConfigMap { t.Helper() @@ -625,6 +625,9 @@ func heldReportConfigMap(t *testing.T, node string, addresses ...string) *corev1 Address: address, Driver: "uio_pci_generic", NUMANode: -1, + // Held rather than idle, which is what makes the machine unusable: + // an idle binding is a disk the draft now claims. + InUse: true, }) } diff --git a/operator/internal/discovery/grouping.go b/operator/internal/discovery/grouping.go index 5952e8768..ea861c1fd 100644 --- a/operator/internal/discovery/grouping.go +++ b/operator/internal/discovery/grouping.go @@ -164,8 +164,15 @@ func (c DeviceClass) signature(addresses []string) string { // what a reviewer edits and "group-1" invites that where a hash does not. The // ordering above is what makes the number stable. func groupName(index int, group Group) string { + size := humanBytes(groupDeviceBytes(group)) + if groupDeviceBytes(group) == 0 { + // Devices the kernel does not present have no size to read, and naming + // the group for "0 B" would state a capacity where there is only an + // absence of one. + size = "unsized" + } return fmt.Sprintf("group-%d-%s-%dx%s", index+1, group.Class, - len(group.Addresses), humanBytes(groupDeviceBytes(group))) + len(group.Addresses), size) } // groupDeviceBytes is the size of one worker's devices in the group, which is diff --git a/operator/internal/discovery/plan.go b/operator/internal/discovery/plan.go index 25cb15d87..ea745b87a 100644 --- a/operator/internal/discovery/plan.go +++ b/operator/internal/discovery/plan.go @@ -15,6 +15,7 @@ import ( "slices" "strings" + "github.com/simplyblock/atlas/blockdev" simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" ) @@ -174,6 +175,64 @@ func (p Plan) RefusalLines() []string { return lines } +// claimableControllers is a device for every NVMe controller the machine owns, +// nothing is driving, and the kernel presents no disk for. +// +// A controller on a userspace driver has no block device, so it cannot reach a +// draft through the device reading at all: everything known about it is its PCI +// address. That is not a gap, because a PCI address is precisely how a NodeGroup +// names an NVMe device — the draft that would be written from a block device and +// the draft written from the controller name the same string. Leaving them out +// refused a machine's storage on the grounds that the machine was not currently +// presenting it, which on a fleet that has run this product before is every +// machine. +// +// Only what nothing is using is offered. A held controller is a disk in service, +// and whatever is driving it need not be this product: the same binding is how a +// disk is passed through to a guest. +// +// What is lost with the block device is everything the disk would have said +// about itself — its size, its partition table, whether it looks blank. A +// controller therefore reaches the draft unsized and uninspected, and approving +// it is approving a disk nobody read. That is the trade the draft makes visible +// rather than one it hides: the group it lands in is named for the count and not +// a capacity, so a reviewer sees which machines are being taken on trust. +func claimableControllers(report nodeprobe.Report, class DeviceClass) []nodeprobe.Device { + if class != ClassNVMe { + // The other class names devices by path, and a controller with no block + // device has none to name. + return nil + } + + presented := map[string]struct{}{} + for _, device := range report.Devices { + if device.PCIAddress != "" { + presented[device.PCIAddress] = struct{}{} + } + } + + var out []nodeprobe.Device + for _, controller := range report.NVMeControllers { + if !controller.BoundToUserspace() || controller.InUse { + continue + } + if _, already := presented[controller.Address]; already { + // The kernel is presenting it after all, so the device reading has + // it and naming it twice would propose one disk under two entries. + continue + } + out = append(out, nodeprobe.Device{ + Name: controller.Address, + PCIAddress: controller.Address, + Kind: string(blockdev.KindDisk), + Transport: string(blockdev.TransportNVMe), + NUMANode: controller.NUMANode, + Available: true, + }) + } + return out +} + // BasicDeviceRules is the default device pipeline for a class and a filter. // // The order is deliberate and is the order a reader wants the refusal in: what @@ -259,6 +318,8 @@ func (p Planner) Plan(reports []nodeprobe.Report, filter *simplyblockv1alpha2.De slices.SortFunc(ordered, func(a, b nodeprobe.Report) int { return cmp.Compare(a.Node, b.Node) }) for _, report := range ordered { + report.Devices = append(report.Devices, claimableControllers(report, class)...) + admitted, refusals := admitDevices(report, deviceRules) plan.Refusals = append(plan.Refusals, refusals...) diff --git a/operator/internal/discovery/plan_test.go b/operator/internal/discovery/plan_test.go index e8240a6eb..9e8dcef6c 100644 --- a/operator/internal/discovery/plan_test.go +++ b/operator/internal/discovery/plan_test.go @@ -324,3 +324,75 @@ func TestPlanHonorsTheSeamsItWasGiven(t *testing.T) { group.Devices.NVMe) } } + +// A controller bound to a userspace driver with nothing using it is a disk this +// fleet owns and nothing is driving, so the draft proposes it. +// +// The kernel presents no block device for it, which is why it cannot come +// through the device reading: the whole of what is known about it is its PCI +// address, and a PCI address is exactly what a NodeGroup names an NVMe device +// by. Refusing it would be refusing the storage the machine has on the grounds +// that the machine is not currently presenting it. +func TestAnIdleUserspaceControllerIsPlannedOn(t *testing.T) { + worker := report("worker-1") + worker.NVMeControllers = []nodeprobe.Controller{ + {Address: "0000:00:02.0", Driver: "uio_pci_generic", NUMANode: 0}, + {Address: "0000:00:03.0", Driver: "uio_pci_generic", NUMANode: 0}, + } + + plan := Planner{Class: ClassNVMe}.Plan([]nodeprobe.Report{worker}, nil) + + if len(plan.Workers) != 1 { + t.Fatalf("planned %d workers, want the one: %s", len(plan.Workers), plan.Summary()) + } + named := map[string]bool{} + for _, set := range plan.NodeSets { + for _, group := range set.Groups { + for _, address := range group.Devices.NVMe { + named[address] = true + } + } + } + for _, want := range []string{"0000:00:02.0", "0000:00:03.0"} { + if !named[want] { + t.Errorf("the draft does not name %s: %+v", want, plan.NodeSets) + } + } +} + +// A controller something is driving is not free, whoever is driving it, so it +// stays out of the draft. +func TestAHeldUserspaceControllerIsNotPlannedOn(t *testing.T) { + worker := report("worker-1") + worker.NVMeControllers = []nodeprobe.Controller{ + {Address: "0000:00:02.0", Driver: "uio_pci_generic", NUMANode: 0, InUse: true}, + } + + plan := Planner{Class: ClassNVMe}.Plan([]nodeprobe.Report{worker}, nil) + + if len(plan.Workers) != 0 { + t.Errorf("a worker whose only controller is in use was planned on: %s", plan.Summary()) + } +} + +// A kernel-bound controller already reaches the draft as a block device, and +// counting it twice would propose the same disk under two names. +func TestAKernelBoundControllerIsNotCountedTwice(t *testing.T) { + disk := disk("nvme0n1", "0000:00:02.0", 0, 3<<40) + worker := report("worker-1", disk) + worker.NVMeControllers = []nodeprobe.Controller{ + {Address: "0000:00:02.0", Driver: "nvme", NUMANode: 0}, + } + + plan := Planner{Class: ClassNVMe}.Plan([]nodeprobe.Report{worker}, nil) + + var addresses []string + for _, set := range plan.NodeSets { + for _, group := range set.Groups { + addresses = append(addresses, group.Devices.NVMe...) + } + } + if len(addresses) != 1 { + t.Errorf("the draft names %v, want the one disk once", addresses) + } +} diff --git a/operator/internal/nodeprobe/job.go b/operator/internal/nodeprobe/job.go index 33ffd4680..a10610ba8 100644 --- a/operator/internal/nodeprobe/job.go +++ b/operator/internal/nodeprobe/job.go @@ -282,9 +282,22 @@ func fieldRefEnv(name, path string) corev1.EnvVar { } } +// pullPolicyOr defaults the probe to being pulled on every run. +// +// The probe and the operator ship in one image, so an operator deployed from a +// moving tag is replaced by a pull while its probes are not: a node still +// holding the previous layer keeps running the previous probe. The report +// carries a version for exactly that skew, so the operator refuses those reports +// and the run waits on machines that will never answer — an upgrade that +// produces a stalled discovery rather than a wrong one, which is harder to read +// than either. +// +// The cost is a registry round-trip per worker per run, against a Job that runs +// once per discovery and lives for seconds. A fleet that cannot pay it, because +// it is air-gapped or already pins a digest, states its own policy. func pullPolicyOr(policy corev1.PullPolicy) corev1.PullPolicy { if policy == "" { - return corev1.PullIfNotPresent + return corev1.PullAlways } return policy } diff --git a/operator/internal/nodeprobe/job_test.go b/operator/internal/nodeprobe/job_test.go index 24db0d73c..62f19f488 100644 --- a/operator/internal/nodeprobe/job_test.go +++ b/operator/internal/nodeprobe/job_test.go @@ -359,3 +359,52 @@ func TestJobNameFitsWhereKubernetesPutsIt(t *testing.T) { } } } + +// The probe is pulled on every run unless a caller says otherwise. +// +// The probe and the operator ship in one image, and an operator deployed from a +// moving tag is replaced by a pull while its probes are not: a node holding the +// previous layer keeps running the previous probe. The report carries a version +// for exactly this skew, so the operator then refuses those reports and the run +// waits on machines that will never answer — an upgrade that silently produces +// a stalled discovery rather than a wrong one. +// +// The cost is a registry round-trip per worker per run, against a Job that runs +// once per discovery and lives for seconds. +func TestTheProbeIsPulledForEveryRun(t *testing.T) { + job, err := Job(JobOptions{ + Namespace: "simplyblock", + Run: "run-1", + Node: "worker-1", + Image: "example.test/simplyblock-operator:develop", + ServiceAccountName: "sb-nodeprobe", + }) + if err != nil { + t.Fatalf("build the Job: %v", err) + } + + container := job.Spec.Template.Spec.Containers[0] + if container.ImagePullPolicy != corev1.PullAlways { + t.Errorf("the probe pull policy is %q, want Always", container.ImagePullPolicy) + } +} + +// A caller that states one keeps it, which is what an air-gapped fleet or a +// pinned digest needs. +func TestAStatedPullPolicyIsKept(t *testing.T) { + job, err := Job(JobOptions{ + Namespace: "simplyblock", + Run: "run-1", + Node: "worker-1", + Image: "example.test/simplyblock-operator:develop", + ServiceAccountName: "sb-nodeprobe", + ImagePullPolicy: corev1.PullIfNotPresent, + }) + if err != nil { + t.Fatalf("build the Job: %v", err) + } + + if got := job.Spec.Template.Spec.Containers[0].ImagePullPolicy; got != corev1.PullIfNotPresent { + t.Errorf("the stated pull policy became %q", got) + } +} From fe809333c1dab198eb19d1a514eb21bb96cb7b56 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Tue, 15 Sep 2026 16:19:11 +0200 Subject: [PATCH 006/206] fix(api): a deployment config could never be approved The approval rules guard an approved document by reading oldSelf.approved, and the field was a bool tagged omitempty, so a document nobody had approved was stored without the key at all. The rule then found nothing to read and failed, which denied the first approval and every one after it: spec: Invalid value: "object": no such key: approved evaluating rule: an approved deployment config is immutable Nothing could be deployed. The gate the whole path runs through was shut against everyone, including the only transition it was ever meant to allow. The field is defaulted to false and always serialized, which is the same requirement read twice: a reviewer sees the gate they are asked to open, and the rules find the field they read. The has() guards carry the documents already stored without the key, which cannot be fixed by defaulting alone -- one of them is the draft this was found on. Every other CEL rule in the group already guards its reads this way; these two were the exception. The rules keep their meaning: approval is one-way, and an approved spec is frozen. They are declared on the spec rather than the object, so the controller can still mark an approved document with the label the expansion selects on and still report on what it did with it. The tests are new because nothing in the tree could have caught this. A fake client cannot tell a stored false from an absent key, so the package gains an envtest suite and the rules are exercised against a real apiserver. Co-Authored-By: Claude Opus 5 (1M context) --- ...mplyblock.io_clusterdeploymentconfigs.yaml | 13 +- .../v1alpha2/clusterdeploymentconfig_types.go | 15 +- ...mplyblock.io_clusterdeploymentconfigs.yaml | 13 +- .../deployment/cel_validation_test.go | 171 ++++++++++++++++++ .../controllers/deployment/suite_test.go | 105 +++++++++++ ...mplyblock.io_clusterdeploymentconfigs.yaml | 13 +- 6 files changed, 321 insertions(+), 9 deletions(-) create mode 100644 operator/internal/controllers/deployment/cel_validation_test.go create mode 100644 operator/internal/controllers/deployment/suite_test.go diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml index b35497da5..c7d4745be 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml @@ -78,10 +78,19 @@ spec: document carries no field for it. properties: approved: + default: false description: |- Approved is the review gate. A document is expanded only once it is set, and is validated but otherwise inert before that, which is what makes reviewing a wrong document safe. + + It is defaulted and serialized rather than omitted when false, and the + two are the same requirement read twice. A reviewer has to see the gate + they are being asked to open, and the rules above have to find the field + they read: a bool omitted when false is a key the apiserver never stores, + so a rule reading it fails rather than reading false, and the first rule + guarding approval denied every approval there could ever be. The has() + guards are what carry documents written before the default existed. type: boolean cluster: description: Cluster is the StorageCluster to create. Ignored when @@ -347,9 +356,9 @@ spec: type: object x-kubernetes-validations: - message: an approved deployment config is immutable - rule: '!oldSelf.approved || self == oldSelf' + rule: '!has(oldSelf.approved) || !oldSelf.approved || self == oldSelf' - message: approval cannot be withdrawn - rule: '!oldSelf.approved || self.approved' + rule: '!has(oldSelf.approved) || !oldSelf.approved || self.approved' - message: 'every group must name the same device class: all nvme or all block' rule: self.nodeSets.all(s, s.groups.all(g, !has(g.devices) || !has(g.devices.block))) diff --git a/operator/api/v1alpha2/clusterdeploymentconfig_types.go b/operator/api/v1alpha2/clusterdeploymentconfig_types.go index d6a594d91..15a306546 100644 --- a/operator/api/v1alpha2/clusterdeploymentconfig_types.go +++ b/operator/api/v1alpha2/clusterdeploymentconfig_types.go @@ -270,15 +270,24 @@ type ClusterTemplate struct { // node set names the same member of its DeviceSelection. The expansion reads the // class off them and stamps it onto the cluster it creates, which is why the // document carries no field for it. -// +kubebuilder:validation:XValidation:rule="!oldSelf.approved || self == oldSelf",message="an approved deployment config is immutable" -// +kubebuilder:validation:XValidation:rule="!oldSelf.approved || self.approved",message="approval cannot be withdrawn" +// +kubebuilder:validation:XValidation:rule="!has(oldSelf.approved) || !oldSelf.approved || self == oldSelf",message="an approved deployment config is immutable" +// +kubebuilder:validation:XValidation:rule="!has(oldSelf.approved) || !oldSelf.approved || self.approved",message="approval cannot be withdrawn" // +kubebuilder:validation:XValidation:rule="self.nodeSets.all(s, s.groups.all(g, !has(g.devices) || !has(g.devices.block))) || self.nodeSets.all(s, s.groups.all(g, !has(g.devices) || !has(g.devices.nvme)))",message="every group must name the same device class: all nvme or all block" type ClusterDeploymentConfigSpec struct { // Approved is the review gate. A document is expanded only once it is set, // and is validated but otherwise inert before that, which is what makes // reviewing a wrong document safe. + // + // It is defaulted and serialized rather than omitted when false, and the + // two are the same requirement read twice. A reviewer has to see the gate + // they are being asked to open, and the rules above have to find the field + // they read: a bool omitted when false is a key the apiserver never stores, + // so a rule reading it fails rather than reading false, and the first rule + // guarding approval denied every approval there could ever be. The has() + // guards are what carry documents written before the default existed. // +optional - Approved bool `json:"approved,omitempty"` + // +kubebuilder:default=false + Approved bool `json:"approved"` // Environment is the Kubernetes distribution this deployment targets. It is a // shorthand the expansion spends: it sets enableKubeletConfiguration, diff --git a/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml b/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml index b35497da5..c7d4745be 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml @@ -78,10 +78,19 @@ spec: document carries no field for it. properties: approved: + default: false description: |- Approved is the review gate. A document is expanded only once it is set, and is validated but otherwise inert before that, which is what makes reviewing a wrong document safe. + + It is defaulted and serialized rather than omitted when false, and the + two are the same requirement read twice. A reviewer has to see the gate + they are being asked to open, and the rules above have to find the field + they read: a bool omitted when false is a key the apiserver never stores, + so a rule reading it fails rather than reading false, and the first rule + guarding approval denied every approval there could ever be. The has() + guards are what carry documents written before the default existed. type: boolean cluster: description: Cluster is the StorageCluster to create. Ignored when @@ -347,9 +356,9 @@ spec: type: object x-kubernetes-validations: - message: an approved deployment config is immutable - rule: '!oldSelf.approved || self == oldSelf' + rule: '!has(oldSelf.approved) || !oldSelf.approved || self == oldSelf' - message: approval cannot be withdrawn - rule: '!oldSelf.approved || self.approved' + rule: '!has(oldSelf.approved) || !oldSelf.approved || self.approved' - message: 'every group must name the same device class: all nvme or all block' rule: self.nodeSets.all(s, s.groups.all(g, !has(g.devices) || !has(g.devices.block))) diff --git a/operator/internal/controllers/deployment/cel_validation_test.go b/operator/internal/controllers/deployment/cel_validation_test.go new file mode 100644 index 000000000..3bed1a3a3 --- /dev/null +++ b/operator/internal/controllers/deployment/cel_validation_test.go @@ -0,0 +1,171 @@ +// Validation of the CEL rules compiled into the ClusterDeploymentConfig schema, +// run against a real apiserver. +// +// It lives here rather than under internal/webhook because there is no webhook +// involved: the rules are enforced by the apiserver itself, and envtest is the +// only place in the tree that starts one. +// +// The approval rules are the reason this file exists. They are the gate the +// whole deployment path runs through, they are expressed entirely in CEL, and +// they read a field whose presence a fake client cannot tell apart from its +// absence — so nothing short of an apiserver could have caught them being +// unsatisfiable. + +package deployment + +import ( + "context" + "strings" + "testing" + + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + + "github.com/simplyblock/atlas/ptr" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// aStoredDocument writes a document the way discovery writes one: approved +// unset, because nobody has reviewed it yet. +func aStoredDocument(t *testing.T, apiClient client.Client, name string) *simplyblockv1alpha2.ClusterDeploymentConfig { + t.Helper() + + config := &simplyblockv1alpha2.ClusterDeploymentConfig{ + ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: "default"}, + Spec: simplyblockv1alpha2.ClusterDeploymentConfigSpec{ + Cluster: &simplyblockv1alpha2.ClusterTemplate{ + Name: name + "-cluster", + VCPUCount: ptr.To(int32(4)), + MaxSubsystemCount: ptr.To(int32(30)), + }, + NodeSets: []simplyblockv1alpha2.NodeSet{{ + Name: "discovered", + Groups: []simplyblockv1alpha2.NodeGroup{{ + Name: "group-1", + Workers: []string{"worker-1"}, + Devices: &simplyblockv1alpha2.DeviceSelection{ + NVMe: []string{"0000:00:02.0"}, + }, + }}, + }}, + }, + } + if err := apiClient.Create(context.Background(), config); err != nil { + t.Fatalf("storing the document: %v", err) + } + return config +} + +// A document nobody has approved can be approved. +// +// This is the whole deployment path in one assertion. The approval rules guard +// an approved document against edits by reading oldSelf.approved, and a field +// omitted when false is a field the apiserver never stored, so that read found +// no key and failed the rule — which denied the very first approval and every +// one after it. Nothing could be deployed at all. +func TestAnUnapprovedDocumentCanBeApproved(t *testing.T) { + apiClient := apiServer(t) + config := aStoredDocument(t, apiClient, "approvable") + + config.Spec.Approved = true + if err := apiClient.Update(context.Background(), config); err != nil { + t.Fatalf("approving a document that nobody had approved: %v", err) + } + + var fresh simplyblockv1alpha2.ClusterDeploymentConfig + if err := apiClient.Get(context.Background(), client.ObjectKeyFromObject(config), &fresh); err != nil { + t.Fatalf("reading it back: %v", err) + } + if !fresh.Spec.Approved { + t.Error("the approval did not stick") + } +} + +// The document discovery writes says it is unapproved rather than leaving the +// question open. +// +// A reviewer reading a draft has to see the gate they are being asked to open, +// and a reader of the stored object has to find the field the rules read. Both +// are the same requirement: the key exists. +func TestAStoredDocumentCarriesItsApprovalFlag(t *testing.T) { + apiClient := apiServer(t) + config := aStoredDocument(t, apiClient, "explicit") + + stored := &simplyblockv1alpha2.ClusterDeploymentConfig{} + if err := apiClient.Get(context.Background(), client.ObjectKeyFromObject(config), stored); err != nil { + t.Fatalf("reading it back: %v", err) + } + + // The typed read cannot tell an absent key from a false one, so the rules + // are what the presence is asserted through: a rule reading the field on an + // object that does not carry it fails rather than reading false. + stored.Spec.Approved = true + if err := apiClient.Update(context.Background(), stored); err != nil { + t.Fatalf("the stored document carries no approval key: %v", err) + } +} + +// Approval is one-way, and an approved document is frozen. Both rules exist to +// stop a document being edited out from under a deployment that is acting on it. +func TestAnApprovedDocumentIsFrozen(t *testing.T) { + apiClient := apiServer(t) + config := aStoredDocument(t, apiClient, "frozen") + + config.Spec.Approved = true + if err := apiClient.Update(context.Background(), config); err != nil { + t.Fatalf("approving: %v", err) + } + + withdrawn := config.DeepCopy() + withdrawn.Spec.Approved = false + err := apiClient.Update(context.Background(), withdrawn) + if err == nil { + t.Error("approval was withdrawn") + } else if !strings.Contains(err.Error(), "withdraw") && !strings.Contains(err.Error(), "immutable") { + t.Errorf("the refusal is %v, want one about withdrawal or immutability", err) + } + + edited := config.DeepCopy() + edited.Spec.NodeSets[0].Groups[0].Workers = []string{"worker-1", "worker-2"} + if err := apiClient.Update(context.Background(), edited); err == nil { + t.Error("an approved document was edited") + } +} + +// The freeze is on the document, not on the object. +// +// An approved document is one the expansion is acting on, so its spec is +// closed — but the controller still has to mark it and still has to report on +// it, and both of those live outside the spec. A rule written one level up +// would have frozen the object whole and left the controller unable to record +// what it did with it. +// +// This one was green when it was written: the rules are declared on the spec +// and always were. It is here because nothing else says so, and the next edit +// to those markers is one lifted pin away from locking the label the expansion +// selects on. +func TestAnApprovedDocumentStillTakesLabelsAndStatus(t *testing.T) { + apiClient := apiServer(t) + config := aStoredDocument(t, apiClient, "markable") + + config.Spec.Approved = true + if err := apiClient.Update(context.Background(), config); err != nil { + t.Fatalf("approving: %v", err) + } + + labeled := config.DeepCopy() + labeled.Labels = map[string]string{readyToDeploy: readyToDeployValue} + if err := apiClient.Update(context.Background(), labeled); err != nil { + t.Fatalf("the controller cannot mark an approved document: %v", err) + } + + var fresh simplyblockv1alpha2.ClusterDeploymentConfig + if err := apiClient.Get(context.Background(), client.ObjectKeyFromObject(config), &fresh); err != nil { + t.Fatalf("reading it back: %v", err) + } + fresh.Status.Phase = simplyblockv1alpha2.ClusterDeploymentConfigPhaseExpanding + if err := apiClient.Status().Update(context.Background(), &fresh); err != nil { + t.Fatalf("the controller cannot report on an approved document: %v", err) + } +} diff --git a/operator/internal/controllers/deployment/suite_test.go b/operator/internal/controllers/deployment/suite_test.go new file mode 100644 index 000000000..8ea77e60a --- /dev/null +++ b/operator/internal/controllers/deployment/suite_test.go @@ -0,0 +1,105 @@ +// What this package's envtest-backed tests need to find a real apiserver, and +// nothing else. +// +// Some of the tests here cannot run against a fake client: the CRD's CEL rules +// are evaluated by the apiserver, and structural defaulting only happens there +// too. That second half is the reason this suite exists at all — a field a fake +// client reports as false may be a field the apiserver never stored, and the +// difference is invisible until a CEL rule tries to read it. +// +// Everything else in this package is driven with a fake client, which is where +// the reconciler's stepwise behavior is provable and where the bulk of the +// coverage is. +// +// One apiserver serves all of them. It is started on the first test that asks +// for one and stopped when the package's tests finish, rather than started per +// test: envtest costs several seconds to come up, and four of them in one +// package is most of the package's runtime for no extra coverage. + +package deployment + +import ( + "os" + "path/filepath" + "sync" + "testing" + + "k8s.io/client-go/kubernetes/scheme" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/envtest" + logf "sigs.k8s.io/controller-runtime/pkg/log" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +var ( + sharedEnvOnce sync.Once + sharedEnv *envtest.Environment + sharedEnvClient client.Client + sharedEnvErr error +) + +// TestMain stops the shared apiserver once, after every test that used it. +func TestMain(m *testing.M) { + code := m.Run() + if sharedEnv != nil { + _ = sharedEnv.Stop() + } + os.Exit(code) +} + +// apiServer returns a client against a real apiserver with this repository's +// CRDs installed, starting one on first use. +// +// Only v1alpha2 is exercised through it. The CRDs carry a conversion webhook +// that envtest does not run, so a read at any other version would fail on an +// unreachable webhook rather than on anything the test is about — and v1alpha2 +// is the storage version, which is what makes reading it need no conversion at +// all. +func apiServer(t *testing.T) client.Client { + t.Helper() + sharedEnvOnce.Do(func() { + if err := simplyblockv1alpha2.AddToScheme(scheme.Scheme); err != nil { + sharedEnvErr = err + return + } + sharedEnv = &envtest.Environment{ + CRDDirectoryPaths: []string{ + filepath.Join("..", "..", "..", "config", "crd", "bases"), + }, + ErrorIfCRDPathMissing: true, + BinaryAssetsDirectory: getFirstFoundEnvTestBinaryDir(), + } + cfg, err := sharedEnv.Start() + if err != nil { + sharedEnvErr = err + return + } + sharedEnvClient, sharedEnvErr = client.New(cfg, client.Options{Scheme: scheme.Scheme}) + }) + if sharedEnvErr != nil { + t.Fatalf("starting the test apiserver: %v", sharedEnvErr) + } + return sharedEnvClient +} + +// getFirstFoundEnvTestBinaryDir locates the envtest asset binaries. +// +// controller-runtime normally passes them through KUBEBUILDER_ASSETS, which the +// Makefile sets. This is what makes the same test runnable straight from an +// editor, and it reads the shared repository-root .bin that +// `make setup-envtest` populates. +func getFirstFoundEnvTestBinaryDir() string { + basePath := filepath.Join("..", "..", "..", "..", ".bin", "k8s") + entries, err := os.ReadDir(basePath) + if err != nil { + logf.Log.Error(err, "the envtest assets could not be read", "path", basePath) + return "" + } + for _, entry := range entries { + if entry.IsDir() { + return filepath.Join(basePath, entry.Name()) + } + } + return "" +} diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml index b35497da5..c7d4745be 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml @@ -78,10 +78,19 @@ spec: document carries no field for it. properties: approved: + default: false description: |- Approved is the review gate. A document is expanded only once it is set, and is validated but otherwise inert before that, which is what makes reviewing a wrong document safe. + + It is defaulted and serialized rather than omitted when false, and the + two are the same requirement read twice. A reviewer has to see the gate + they are being asked to open, and the rules above have to find the field + they read: a bool omitted when false is a key the apiserver never stores, + so a rule reading it fails rather than reading false, and the first rule + guarding approval denied every approval there could ever be. The has() + guards are what carry documents written before the default existed. type: boolean cluster: description: Cluster is the StorageCluster to create. Ignored when @@ -347,9 +356,9 @@ spec: type: object x-kubernetes-validations: - message: an approved deployment config is immutable - rule: '!oldSelf.approved || self == oldSelf' + rule: '!has(oldSelf.approved) || !oldSelf.approved || self == oldSelf' - message: approval cannot be withdrawn - rule: '!oldSelf.approved || self.approved' + rule: '!has(oldSelf.approved) || !oldSelf.approved || self.approved' - message: 'every group must name the same device class: all nvme or all block' rule: self.nodeSets.all(s, s.groups.all(g, !has(g.devices) || !has(g.devices.block))) From 3bd8891fc1e4aba808fa9d413ac6ed9b21b3b3cb Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Tue, 15 Sep 2026 16:45:05 +0200 Subject: [PATCH 007/206] fix(operator): the workload enrolls the workers its nodes run on The DaemonSet that runs SPDK selects workers by io.simplyblock.storagenodeset, and nothing wrote that label on the provisioning path. Writing it was the retired StorageNodeSet's job; when provisioning was rebuilt around StorageNode the migration path kept doing it and the new path did not, so migrate.go held the only two calls to LabelWorker in the tree. What that produced was a deadlock rather than a failure. A node holds at CheckingHost waiting for its worker's storage-node API to answer, and the process that would answer cannot be scheduled until the label it is waiting on exists. Neither side reports anything wrong, because neither side is wrong. A deployment config expanded into a cluster and three nodes, and the fleet then sat there: the DaemonSet had DESIRED 0 and every node was in Provisioning. Enrollment goes in the workload reconcile, before the DaemonSet, because that is the order a reader wants it in: the selector is written, then the thing that selects on it. It is per worker rather than per node, since two nodes on one worker are two slots of one machine and LabelWorker rewrites that machine's whole slot label set from the nodes that want it. Two returns of a result beside a non-nil error go with it. controller-runtime discards the RequeueAfter in that case and logs a warning about it, so the retry they asked for was never the retry that happened. Co-Authored-By: Claude Opus 5 (1M context) --- .../controllers/node/enrollment_test.go | 143 ++++++++++++++++++ .../controllers/node/workload_controller.go | 55 ++++++- 2 files changed, 195 insertions(+), 3 deletions(-) create mode 100644 operator/internal/controllers/node/enrollment_test.go diff --git a/operator/internal/controllers/node/enrollment_test.go b/operator/internal/controllers/node/enrollment_test.go new file mode 100644 index 000000000..54e50b44e --- /dev/null +++ b/operator/internal/controllers/node/enrollment_test.go @@ -0,0 +1,143 @@ +// Whether a cluster's workload puts its workers into the storage plane. +// +// The DaemonSet that runs SPDK selects workers by the io.simplyblock.storagenodeset +// label, so a worker without it runs nothing. Writing that label was the retired +// StorageNodeSet's job, and when provisioning was rebuilt around StorageNode the +// migration path kept writing it while nothing on the provisioning path did. +// +// The result deadlocks rather than fails, which is what makes it worth a test of +// its own: a node holds at CheckingHost waiting for the worker's storage-node API +// to answer, and the process that would answer cannot be scheduled until the +// label it is waiting on has been written. Neither side reports an error. + +package node + +import ( + "context" + "testing" + + appsv1 "k8s.io/api/apps/v1" + corev1 "k8s.io/api/core/v1" + discoveryv1 "k8s.io/api/discovery/v1" + rbacv1 "k8s.io/api/rbac/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/client-go/tools/events" + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + atlaskube "github.com/simplyblock/atlas/kube" + "github.com/simplyblock/atlas/ptr" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/testsupport" +) + +const ( + enrollNamespace = "simplyblock" + enrollCluster = "a-cluster" + enrollWorker = "worker-1" +) + +// aWorkloadReconciler drives the cluster's workload over a fake client. +func aWorkloadReconciler(t *testing.T, objects ...client.Object) *StorageNodeWorkloadReconciler { + t.Helper() + scheme := testsupport.NewScheme(t, + corev1.AddToScheme, appsv1.AddToScheme, discoveryv1.AddToScheme, rbacv1.AddToScheme) + apiClient := fake.NewClientBuilder().WithScheme(scheme). + WithObjects(objects...). + WithStatusSubresource(&simplyblockv1alpha2.StorageCluster{}). + Build() + + return &StorageNodeWorkloadReconciler{ + Client: apiClient, + Scheme: scheme, + Recorder: events.NewFakeRecorder(64), + Namespace: enrollNamespace, + Workload: &Workload{Client: apiClient}, + } +} + +// aSizedCluster carries the sizing its workers boot from, without which the +// workload refuses before it reaches anything this file is about. +func aSizedCluster() *simplyblockv1alpha2.StorageCluster { + return &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{Name: enrollCluster, Namespace: enrollNamespace}, + Spec: simplyblockv1alpha2.StorageClusterSpec{ + MaxSubsystemCount: ptr.To(int32(30)), + VCPUCount: ptr.To(int32(4)), + StorageNodes: &simplyblockv1alpha2.StorageNodesSpec{ + Image: "example.test/storage-node:test", + }, + }, + } +} + +// aNodeOn is a StorageNode bound to a worker, as the deployment config creates one. +func aNodeOn(worker string) *simplyblockv1alpha2.StorageNode { + return &simplyblockv1alpha2.StorageNode{ + ObjectMeta: metav1.ObjectMeta{ + Name: enrollCluster + "-" + worker + "-0", + Namespace: enrollNamespace, + }, + Spec: simplyblockv1alpha2.StorageNodeSpec{ + ClusterRef: enrollCluster, + WorkerNode: worker, + }, + } +} + +// Every worker holding a node of the cluster is enrolled, so the DaemonSet has +// somewhere to schedule. +func TestTheWorkloadEnrollsEveryWorkerItHasANodeOn(t *testing.T) { + cluster := aSizedCluster() + objects := []client.Object{ + &corev1.Node{ObjectMeta: metav1.ObjectMeta{Name: "worker-1"}}, + &corev1.Node{ObjectMeta: metav1.ObjectMeta{Name: "worker-2"}}, + cluster, aNodeOn("worker-1"), aNodeOn("worker-2"), + } + r := aWorkloadReconciler(t, objects...) + + if _, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: client.ObjectKey{Namespace: enrollNamespace, Name: enrollCluster}, + }); err != nil { + t.Fatalf("reconcile: %v", err) + } + + for _, name := range []string{"worker-1", "worker-2"} { + var fresh corev1.Node + if err := r.Get(context.Background(), client.ObjectKey{Name: name}, &fresh); err != nil { + t.Fatalf("reading %s: %v", name, err) + } + if got := fresh.Labels[atlaskube.LabelStorageNodeSet]; got != enrollCluster { + t.Errorf("%s carries %q for the node-set label, want %q", + name, got, enrollCluster) + } + } +} + +// A worker with no node of this cluster on it is left alone, so a fleet running +// something else on its other machines is not enrolled into this one. +func TestTheWorkloadLeavesUnrelatedWorkersAlone(t *testing.T) { + cluster := aSizedCluster() + objects := []client.Object{ + &corev1.Node{ObjectMeta: metav1.ObjectMeta{Name: "worker-1"}}, + &corev1.Node{ObjectMeta: metav1.ObjectMeta{Name: "bystander"}}, + cluster, aNodeOn("worker-1"), + } + r := aWorkloadReconciler(t, objects...) + + if _, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: client.ObjectKey{Namespace: enrollNamespace, Name: enrollCluster}, + }); err != nil { + t.Fatalf("reconcile: %v", err) + } + + var bystander corev1.Node + if err := r.Get(context.Background(), client.ObjectKey{Name: "bystander"}, &bystander); err != nil { + t.Fatalf("reading the bystander: %v", err) + } + if _, enrolled := bystander.Labels[atlaskube.LabelStorageNodeSet]; enrolled { + t.Errorf("a worker with no node of this cluster was enrolled: %v", bystander.Labels) + } +} diff --git a/operator/internal/controllers/node/workload_controller.go b/operator/internal/controllers/node/workload_controller.go index ec69e2086..62efbb300 100644 --- a/operator/internal/controllers/node/workload_controller.go +++ b/operator/internal/controllers/node/workload_controller.go @@ -125,7 +125,7 @@ func (r *StorageNodeWorkloadReconciler) Reconcile( // script with --max-subsys-count=0 and fails there, which is a long way from // the cause (§5.3). if err := r.Workload.ReconcileConfig(ctx, &cluster); err != nil { - return ctrl.Result{RequeueAfter: nodeRetry}, err + return ctrl.Result{}, err } for _, step := range []struct { @@ -136,16 +136,65 @@ func (r *StorageNodeWorkloadReconciler) Reconcile( {"the serving certificates", r.reconcileCertificates}, {"the headless service", r.reconcileService}, {"the endpoint slice", r.reconcileEndpointSlice}, + {"the worker enrollment", r.enrollWorkers}, {"the daemon set", r.reconcileDaemonSet}, } { if err := step.run(ctx, &cluster); err != nil { - return ctrl.Result{RequeueAfter: nodeRetry}, - fmt.Errorf("reconcile %s: %w", step.what, err) + return ctrl.Result{}, fmt.Errorf("reconcile %s: %w", step.what, err) } } return ctrl.Result{}, nil } +// enrollWorkers puts every worker this cluster has a node on into its storage +// plane, which is what gives the DaemonSet somewhere to schedule. +// +// It runs before the DaemonSet and not after, because the order is the one a +// reader wants it in: the selector is written, then the thing that selects on it. +// +// The retired StorageNodeSet enrolled its own workers, and when provisioning was +// rebuilt around StorageNode the migration path kept doing it and nothing on the +// provisioning path did. What that produced was not a failure but a deadlock: a +// node holds at CheckingHost waiting for the worker's storage-node API, and the +// process that would answer cannot be scheduled until this label exists. Neither +// side reports anything wrong, because neither side is. +// +// Enrollment is per worker rather than per node. Two nodes on one worker are two +// slots of one machine, and LabelWorker rewrites that machine's whole slot label +// set from the nodes that want it, so calling it once per worker is both +// sufficient and what keeps the set consistent. +func (r *StorageNodeWorkloadReconciler) enrollWorkers( + ctx context.Context, cluster *simplyblockv1alpha2.StorageCluster, +) error { + var nodes simplyblockv1alpha2.StorageNodeList + if err := r.List(ctx, &nodes, client.InNamespace(cluster.Namespace)); err != nil { + return fmt.Errorf("list the storage nodes: %w", err) + } + + seen := map[string]struct{}{} + for i := range nodes.Items { + node := &nodes.Items[i] + if node.Spec.ClusterRef != cluster.Name || node.Spec.WorkerNode == "" { + continue + } + // A node on its way out is not a reason to keep its worker enrolled, and + // releasing it is ReleaseWorker's job on the node's own path. + if !node.DeletionTimestamp.IsZero() { + continue + } + if _, already := seen[node.Spec.WorkerNode]; already { + continue + } + seen[node.Spec.WorkerNode] = struct{}{} + + if err := r.Workload.LabelWorker( + ctx, cluster.Namespace, cluster.Name, node.Spec.WorkerNode); err != nil { + return fmt.Errorf("enroll worker %s: %w", node.Spec.WorkerNode, err) + } + } + return nil +} + // reconcileDaemonSet applies the pod template every storage node runs under. // // The TLS Secret's resourceVersion is stamped onto the template so that a From 69e9f3f281977cba7337c4e05de42d86f0afa6eb Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Tue, 15 Sep 2026 16:59:01 +0200 Subject: [PATCH 008/206] fix(discovery): name a management interface, and retry an add that gave up A storage node added without a management interface is refused by the control plane with "No management interface with IP found in provided interfaces", and the draft named none: discovery read the interfaces only to rank NUMA nodes by NIC speed, so every document it wrote left the field empty and every node_add it led to failed. Nothing validated it either, so the refusal arrived as a task that gave up on a document that looked complete. The probe now reads what choosing one needs. sysfs carries no addresses, so they come through a seam on the inventory Config whose default reads this process's network namespace -- the host's, for a caller with host networking. Bridges are read from the bridge directory the kernel exports rather than guessed from br0 and cni0 and docker0. The choice is a ranking rather than a match, because a fleet's machines do not agree on what their NICs are called. The interface holding the node's InternalIP wins outright, since the operator addresses the worker by that everywhere else; failing that, the fastest hardware interface that is up and holds a reachable address. The cluster's own plumbing never wins, and a machine presenting nothing suitable is named nothing rather than named wrongly. The interface joins the group signature, because a NodeGroup names one for every worker it lists and two machines that call their NICs different things cannot be described by one group. Resolving now asks again when the add is over. It matched the backend node the add was asked to produce and waited out its deadline when there was none, holding the cluster's one node-add slot while it did -- so every other node waited behind a node that was never coming. The task window is the evidence: no node_add still in it means the add this step waits on has finished, and an add that finished without a node is one to ask for again. Report version goes to 3 for the two new interface fields. Co-Authored-By: Claude Opus 5 (1M context) --- atlas-lib/inventory/inventory.go | 13 ++ atlas-lib/inventory/netiface.go | 75 +++++++- atlas-lib/inventory/netiface_test.go | 99 ++++++++++- .../internal/controllers/deployment/events.go | 10 +- .../controllers/deployment/validation.go | 31 ++++ operator/internal/controllers/node/events.go | 4 + .../controllers/node/resolve_retry_test.go | 160 +++++++++++++++++ .../node/storagenode_controller.go | 65 ++++++- operator/internal/discovery/grouping.go | 42 ++++- operator/internal/discovery/kubenode.go | 13 ++ operator/internal/discovery/mgmtiface.go | 97 +++++++++++ operator/internal/discovery/mgmtiface_test.go | 163 ++++++++++++++++++ operator/internal/discovery/plan.go | 4 +- operator/internal/nodeprobe/collect.go | 2 + operator/internal/nodeprobe/report.go | 16 +- 15 files changed, 776 insertions(+), 18 deletions(-) create mode 100644 operator/internal/controllers/node/resolve_retry_test.go create mode 100644 operator/internal/discovery/mgmtiface.go create mode 100644 operator/internal/discovery/mgmtiface_test.go diff --git a/atlas-lib/inventory/inventory.go b/atlas-lib/inventory/inventory.go index 80c4c0852..59006b5e2 100644 --- a/atlas-lib/inventory/inventory.go +++ b/atlas-lib/inventory/inventory.go @@ -90,12 +90,25 @@ type Config struct { // kernel. Exclusive blockdev.ExclusiveOpener + // InterfaceAddresses answers which IP addresses each interface holds. A nil + // reader is LocalAddresses, this process's own network namespace, which is + // the host's for a caller running with host networking. + InterfaceAddresses AddressReader + // Kubernetes is the cluster half of a collection's sources. The zero value // collects no environment, which is what a caller inspecting a machine // outside a cluster has. Kubernetes KubernetesSources } +// addresses is the reader to use, defaulted. +func (c Config) addresses() AddressReader { + if c.InterfaceAddresses != nil { + return c.InterfaceAddresses + } + return LocalAddresses +} + // KubernetesSources is what the environment is concluded from. // // Both fields are optional and are read independently: a caller with the nodes diff --git a/atlas-lib/inventory/netiface.go b/atlas-lib/inventory/netiface.go index 3017053c5..6e35cf412 100644 --- a/atlas-lib/inventory/netiface.go +++ b/atlas-lib/inventory/netiface.go @@ -17,6 +17,7 @@ package inventory import ( "cmp" "fmt" + "net" "os" "path/filepath" "slices" @@ -90,6 +91,61 @@ type Interface struct { // NUMANodeUnknown. It is what decides whether a storage node pinned to one // socket reaches its NIC across the interconnect. NUMANode int + + // Bridge reports whether the interface is a software bridge. + // + // It is separate from Virtual, which a bridge also is, because the two + // answer different questions. Virtual says the interface is backed by no + // hardware; Bridge says it is carrying somebody else's traffic, which is + // what makes a cluster's own bridge a poor choice for a management address + // even where it holds one. + Bridge bool + + // Addresses are the IP addresses assigned to the interface, as plain + // addresses without a prefix length, in the order the host reports them. + // + // They do not come from sysfs, which does not carry them. They are read + // through Config.InterfaceAddresses, and a caller that supplies none gets + // none rather than an error: an interface reported without its addresses is + // still worth reporting. + Addresses []string +} + +// AddressReader answers which IP addresses each interface holds, by interface +// name. +// +// It is a seam because sysfs does not carry addresses and the kernel's own +// answer is namespace-scoped: a process reading it reports the addresses of the +// network namespace it is in, so a pod without the host's network would report +// its own. The default reads this process's namespace, which is the host's when +// the caller runs with host networking, and a caller that cannot guarantee that +// supplies its own reader rather than being handed a confident wrong answer. +type AddressReader func() (map[string][]string, error) + +// LocalAddresses reads the addresses of this process's network namespace. +func LocalAddresses() (map[string][]string, error) { + ifaces, err := net.Interfaces() + if err != nil { + return nil, fmt.Errorf("list the network interfaces: %w", err) + } + + out := make(map[string][]string, len(ifaces)) + for _, iface := range ifaces { + addrs, err := iface.Addrs() + if err != nil { + // One interface refusing its addresses is not a reason to lose the + // rest, and an interface with none reported is reported with none. + continue + } + for _, addr := range addrs { + ip, _, err := net.ParseCIDR(addr.String()) + if err != nil { + continue + } + out[iface.Name] = append(out[iface.Name], ip.String()) + } + } + return out, nil } // ReadInterfaces reads every network interface the host presents, ordered by @@ -108,9 +164,21 @@ func ReadInterfaces(cfg Config) ([]Interface, error) { return nil, fmt.Errorf("list %s: %w", base, err) } + // The addresses are read once for the whole host rather than per interface, + // because the source answers for all of them at once. A reader that fails + // costs the addresses and nothing else. + addresses := map[string][]string{} + if reader := cfg.addresses(); reader != nil { + if read, err := reader(); err == nil { + addresses = read + } + } + ifaces := make([]Interface, 0, len(entries)) for _, entry := range entries { - ifaces = append(ifaces, readInterface(filepath.Join(base, entry.Name()), entry.Name())) + iface := readInterface(filepath.Join(base, entry.Name()), entry.Name()) + iface.Addresses = addresses[entry.Name()] + ifaces = append(ifaces, iface) } slices.SortFunc(ifaces, func(a, b Interface) int { return cmp.Compare(a.Name, b.Name) }) @@ -147,6 +215,11 @@ func readInterface(dir, name string) Interface { return iface } iface.Virtual = sysfs.IsVirtual(resolved) + // A bridge exports a bridge/ directory whatever it is named, which is what + // makes this a reading rather than a guess at br0 and cni0 and docker0. + if entries, err := os.Stat(filepath.Join(dir, "bridge")); err == nil && entries.IsDir() { + iface.Bridge = true + } if iface.Virtual { return iface } diff --git a/atlas-lib/inventory/netiface_test.go b/atlas-lib/inventory/netiface_test.go index 77c132b01..40f5779b8 100644 --- a/atlas-lib/inventory/netiface_test.go +++ b/atlas-lib/inventory/netiface_test.go @@ -10,7 +10,13 @@ package inventory -import "testing" +import ( + "errors" + "os" + "path/filepath" + "reflect" + "testing" +) // netHost carries one 25 GbE NIC that is up, one 10 GbE NIC whose link is down, // a CNI bridge, and loopback. @@ -100,7 +106,7 @@ func TestReadInterfacesReportsAPhysicalNICWhole(t *testing.T) { PCIAddress: "0000:3b:00.0", NUMANode: 0, } - if got := byName(t, ifaces, "eth0"); got != want { + if got := byName(t, ifaces, "eth0"); !reflect.DeepEqual(got, want) { t.Errorf("read %+v, want %+v", got, want) } } @@ -193,3 +199,92 @@ func TestReadInterfacesReportsNoneRatherThanFailingWithoutTheClassDirectory(t *t t.Errorf("read %d interfaces from a tree with no class/net", len(ifaces)) } } + +// A bridge is read from the directory the kernel exports for one, not guessed +// from a name: br0 and cni0 and docker0 are conventions, and a fleet is free to +// name its bridge anything. +func TestABridgeIsReadFromItsOwnDirectory(t *testing.T) { + root := t.TempDir() + base := filepath.Join(root, "sys", "class", "net") + for _, name := range []string{"eth0", "weird-name"} { + if err := os.MkdirAll(filepath.Join(base, name), 0o755); err != nil { + t.Fatal(err) + } + } + if err := os.MkdirAll(filepath.Join(base, "weird-name", "bridge"), 0o755); err != nil { + t.Fatal(err) + } + + ifaces, err := ReadInterfaces(Config{SysfsRoot: filepath.Join(root, "sys")}) + if err != nil { + t.Fatalf("read the interfaces: %v", err) + } + + byName := map[string]Interface{} + for _, iface := range ifaces { + byName[iface.Name] = iface + } + if !byName["weird-name"].Bridge { + t.Error("an interface exporting a bridge directory was not read as a bridge") + } + if byName["eth0"].Bridge { + t.Error("an interface with no bridge directory was read as a bridge") + } +} + +// The addresses come from the reader rather than from sysfs, which carries none. +func TestTheAddressesComeFromTheReader(t *testing.T) { + root := t.TempDir() + base := filepath.Join(root, "sys", "class", "net") + for _, name := range []string{"eth0", "eth1"} { + if err := os.MkdirAll(filepath.Join(base, name), 0o755); err != nil { + t.Fatal(err) + } + } + + ifaces, err := ReadInterfaces(Config{ + SysfsRoot: filepath.Join(root, "sys"), + InterfaceAddresses: func() (map[string][]string, error) { + return map[string][]string{"eth0": {"192.168.10.113", "fe80::1"}}, nil + }, + }) + if err != nil { + t.Fatalf("read the interfaces: %v", err) + } + + byName := map[string]Interface{} + for _, iface := range ifaces { + byName[iface.Name] = iface + } + if got := byName["eth0"].Addresses; len(got) != 2 || got[0] != "192.168.10.113" { + t.Errorf("eth0 carries %v", got) + } + if got := byName["eth1"].Addresses; len(got) != 0 { + t.Errorf("eth1 carries %v, want none", got) + } +} + +// A reader that fails costs the addresses and not the interfaces, because a +// machine whose addresses could not be read still has NICs worth reporting. +func TestAFailedAddressReadStillReportsTheInterfaces(t *testing.T) { + root := t.TempDir() + if err := os.MkdirAll(filepath.Join(root, "sys", "class", "net", "eth0"), 0o755); err != nil { + t.Fatal(err) + } + + ifaces, err := ReadInterfaces(Config{ + SysfsRoot: filepath.Join(root, "sys"), + InterfaceAddresses: func() (map[string][]string, error) { + return nil, errors.New("no permission to read the namespace") + }, + }) + if err != nil { + t.Fatalf("read the interfaces: %v", err) + } + if len(ifaces) != 1 || ifaces[0].Name != "eth0" { + t.Fatalf("read %+v, want the one interface", ifaces) + } + if len(ifaces[0].Addresses) != 0 { + t.Errorf("addresses were invented: %v", ifaces[0].Addresses) + } +} diff --git a/operator/internal/controllers/deployment/events.go b/operator/internal/controllers/deployment/events.go index 143f3b793..59399f95e 100644 --- a/operator/internal/controllers/deployment/events.go +++ b/operator/internal/controllers/deployment/events.go @@ -17,9 +17,13 @@ const ( // What a draft's validation found. These are the whole value of the review // gate: a document that names a worker which does not exist should say so // while it is still a draft, rather than after somebody approved it. - WorkerNotFound = "WorkerNotFound" - DeviceNotFound = "DeviceNotFound" - DeviceClassMismatch = "DeviceClassMismatch" + WorkerNotFound = "WorkerNotFound" + + // NoManagementInterface is a group that names no interface for the storage + // nodes to bind their management address to. + NoManagementInterface = "NoManagementInterface" + DeviceNotFound = "DeviceNotFound" + DeviceClassMismatch = "DeviceClassMismatch" // AwaitingApproval is the one that changes how the kind is used. A valid draft // nobody has approved looks identical to a controller that has not noticed it, diff --git a/operator/internal/controllers/deployment/validation.go b/operator/internal/controllers/deployment/validation.go index c5b71037e..f2b82fce4 100644 --- a/operator/internal/controllers/deployment/validation.go +++ b/operator/internal/controllers/deployment/validation.go @@ -94,9 +94,40 @@ func (r *ClusterDeploymentConfigReconciler) validate( if found := conflictingInterfaces(config); found != "" { findings = append(findings, finding{reason: WorkerNotFound, message: found}) } + if groups := groupsWithoutManagementInterface(config); len(groups) > 0 { + findings = append(findings, finding{ + reason: NoManagementInterface, + message: fmt.Sprintf( + "group(s) %s name no management interface, and the control plane refuses a "+ + "node it cannot find a management address on", + strings.Join(groups, ", ")), + }) + } return findings, nil } +// groupsWithoutManagementInterface names the groups that would be expanded into +// nodes the control plane refuses. +// +// It is worth a finding of its own because of where the refusal otherwise lands. +// The control plane does not reject the request; it accepts it, starts a node_add +// task, and fails inside it. What an administrator sees is a task that gave up on +// a document that looked complete, several steps and one approval away from the +// field that was empty. +func groupsWithoutManagementInterface( + config *simplyblockv1alpha2.ClusterDeploymentConfig, +) []string { + var missing []string + for _, set := range config.Spec.NodeSets { + for _, group := range set.Groups { + if group.MgmtInterface == "" { + missing = append(missing, group.Name) + } + } + } + return missing +} + // duplicateWorkers names every worker the document lists in more than one group. // // The schema permits it and the expansion cannot honor it: a StorageNode is diff --git a/operator/internal/controllers/node/events.go b/operator/internal/controllers/node/events.go index 716585662..ea5e79836 100644 --- a/operator/internal/controllers/node/events.go +++ b/operator/internal/controllers/node/events.go @@ -29,6 +29,10 @@ const ( HostUnreachable = "HostUnreachable" AwaitingSlot = "AwaitingSlot" + // NodeAddGaveUp is the add this node was waiting on leaving the control + // plane's task window without having produced a node. + NodeAddGaveUp = "NodeAddGaveUp" + // NodeAdopted says an existing backend node was taken over rather than // added, which is the difference between a migration and a mistake. NodeAdopted = "NodeAdopted" diff --git a/operator/internal/controllers/node/resolve_retry_test.go b/operator/internal/controllers/node/resolve_retry_test.go new file mode 100644 index 000000000..fc5472c49 --- /dev/null +++ b/operator/internal/controllers/node/resolve_retry_test.go @@ -0,0 +1,160 @@ +// What Resolving does when the add it is waiting on has already given up. +// +// Resolving matches the backend node the control plane was asked to create. A +// node_add that fails creates none, so the match never succeeds and the step +// waits out its deadline — holding the one node-add slot while it does, which +// stalls every other node of the cluster behind a node that is never coming. +// +// The control plane says so plainly: the task leaves the window. So a step that +// is waiting for a node, with nothing still working to produce one, is waiting +// for nothing, and the answer is to ask again rather than to keep waiting. + +package node + +import ( + "context" + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/client-go/tools/events" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/testsupport" +) + +// noBackendNodes answers as a control plane on which the add produced nothing. +// The rest of the interface is embedded and nil: a call to anything else is a +// test reaching past what it is about, and should panic rather than pass. +type noBackendNodes struct { + ControlPlane +} + +func (noBackendNodes) StorageNodes(context.Context, string) ([]NodeReading, error) { + return nil, nil +} + +func aResolvingNode() *simplyblockv1alpha2.StorageNode { + node := &simplyblockv1alpha2.StorageNode{ + ObjectMeta: metav1.ObjectMeta{ + Name: "a-cluster-worker-1-0", + Namespace: "simplyblock", + }, + Spec: simplyblockv1alpha2.StorageNodeSpec{ + ClusterRef: "a-cluster", + WorkerNode: "worker-1", + }, + } + node.Status.Step.State = string(stepResolving) + return node +} + +func aClusterWithTasks(tasks ...simplyblockv1alpha2.ClusterTask) *simplyblockv1alpha2.StorageCluster { + cluster := &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{Name: "a-cluster", Namespace: "simplyblock"}, + } + cluster.Status.UUID = "cluster-uuid" + cluster.Status.Tasks = tasks + return cluster +} + +func aResolver(t *testing.T, cluster *simplyblockv1alpha2.StorageCluster, + node *simplyblockv1alpha2.StorageNode) *StorageNodeReconciler { + t.Helper() + scheme := testsupport.NewScheme(t, corev1.AddToScheme) + worker := &corev1.Node{ + ObjectMeta: metav1.ObjectMeta{Name: "worker-1"}, + Status: corev1.NodeStatus{Addresses: []corev1.NodeAddress{ + {Type: corev1.NodeInternalIP, Address: "192.168.10.113"}, + }}, + } + apiClient := fake.NewClientBuilder().WithScheme(scheme). + WithObjects(worker, cluster, node). + WithStatusSubresource(&simplyblockv1alpha2.StorageNode{}). + Build() + + return &StorageNodeReconciler{ + Client: apiClient, + Scheme: scheme, + Recorder: events.NewFakeRecorder(64), + API: noBackendNodes{}, + } +} + +// A node_add that has left the task window without producing a node is one to +// ask for again. +func TestResolvingAsksAgainWhenTheAddIsOver(t *testing.T) { + node := aResolvingNode() + cluster := aClusterWithTasks(simplyblockv1alpha2.ClusterTask{ + ID: "task-1", Type: "node_add", Status: "done", + }) + r := aResolver(t, cluster, node) + + next, done, err := r.resolve(context.Background(), node, cluster) + if err != nil { + t.Fatalf("resolve: %v", err) + } + if next != stepPosting { + t.Errorf("the step went to %q, want Posting so the add is asked for again", next) + } + if !done { + t.Error("the step did not advance, so the retry never happens") + } +} + +// While the add is still working, waiting is right: a node that re-POSTed here +// would ask for a second one alongside the first. +func TestResolvingWaitsWhileTheAddIsStillRunning(t *testing.T) { + node := aResolvingNode() + cluster := aClusterWithTasks(simplyblockv1alpha2.ClusterTask{ + ID: "task-1", Type: "node_add", Status: "running", + }) + r := aResolver(t, cluster, node) + + next, done, err := r.resolve(context.Background(), node, cluster) + if err != nil { + t.Fatalf("resolve: %v", err) + } + if next != stepResolving || done { + t.Errorf("the step went to %q (done %v), want to keep waiting", next, done) + } +} + +// A task window that names no node_add at all is the same case as one whose add +// has finished: nothing is working on producing the node this step is waiting +// for. The window is capped, so a task that has scrolled out of it is over. +func TestResolvingAsksAgainWhenNoAddIsInTheWindow(t *testing.T) { + node := aResolvingNode() + cluster := aClusterWithTasks(simplyblockv1alpha2.ClusterTask{ + ID: "task-9", Type: "cluster_status", Status: "running", + }) + r := aResolver(t, cluster, node) + + next, _, err := r.resolve(context.Background(), node, cluster) + if err != nil { + t.Fatalf("resolve: %v", err) + } + if next != stepPosting { + t.Errorf("the step went to %q, want Posting", next) + } +} + +// Another cluster's business is not this node's. A migration running alongside +// says nothing about whether the add that this step is waiting on is over. +func TestAnUnrelatedRunningTaskDoesNotHoldResolving(t *testing.T) { + node := aResolvingNode() + cluster := aClusterWithTasks( + simplyblockv1alpha2.ClusterTask{ID: "t1", Type: "node_add", Status: "done"}, + simplyblockv1alpha2.ClusterTask{ID: "t2", Type: "lvol_migration", Status: "running"}, + ) + r := aResolver(t, cluster, node) + + next, _, err := r.resolve(context.Background(), node, cluster) + if err != nil { + t.Fatalf("resolve: %v", err) + } + if next != stepPosting { + t.Errorf("the step went to %q, want Posting: a migration is not this node's add", next) + } +} diff --git a/operator/internal/controllers/node/storagenode_controller.go b/operator/internal/controllers/node/storagenode_controller.go index 571f8eae4..ef55e8999 100644 --- a/operator/internal/controllers/node/storagenode_controller.go +++ b/operator/internal/controllers/node/storagenode_controller.go @@ -370,8 +370,7 @@ func (r *StorageNodeReconciler) performNodeStep( case stepPosting: return stepResolving, true, r.postNode(ctx, node, cluster) case stepResolving: - done, err := r.resolveUUID(ctx, node, cluster) - return stepResolving, done, err + return r.resolve(ctx, node, cluster) case stepAdopting: done, err := r.resolveUUID(ctx, node, cluster) return stepAdopting, done, err @@ -543,6 +542,68 @@ func (r *StorageNodeReconciler) postNode( // one worker are sorted by RPC port ascending, and position in that list is the // socket ordinal, because the ports are assigned in socket order at node-add time // (§4.3). +// nodeAddTask is the control plane's own name for the job Posting starts. +const nodeAddTask = "node_add" + +// resolve waits for the backend node the add was asked to produce, and asks +// again when nothing is still working on producing one. +// +// The add can fail after it has started — a management interface the control +// plane cannot find an address on is the case this was written for — and what it +// leaves behind is a task that finished and no node. Nothing about that is +// visible from the node list, which is empty either way, so a step that only +// matched the list waited out its deadline against an add that had already given +// up. It held the cluster's one node-add slot while it waited, so every other +// node of the cluster waited behind a node that was never coming. +// +// The task window is the evidence. A node_add still in it is an add worth +// waiting for; no node_add in it at all means the add this step is waiting on is +// over, and an add that is over without a node is one to ask for again. The +// window is capped, so a task that has scrolled out of it has finished too. +// +// Asking again is unbounded here and bounded by the step's own deadline, which +// is what makes a permanently failing add fail the node rather than spin on it +// forever. +func (r *StorageNodeReconciler) resolve( + ctx context.Context, + node *simplyblockv1alpha2.StorageNode, + cluster *simplyblockv1alpha2.StorageCluster, +) (nodeStep, bool, error) { + done, err := r.resolveUUID(ctx, node, cluster) + if err != nil || done { + return stepResolving, done, err + } + + if addInFlight(cluster) { + return stepResolving, false, nil + } + + r.emit(node, corev1.EventTypeWarning, NodeAddGaveUp, fmt.Sprintf( + "the node_add for worker %s finished without producing a node, so it is being asked for again", + node.Spec.WorkerNode)) + return stepPosting, true, nil +} + +// addInFlight reports whether the control plane is still working on a node_add +// for this cluster. +// +// It does not ask which node the task is for, because the window does not say +// and because the cap is one add at a time: a node_add in flight is this node's +// add or the add of the node holding the slot ahead of it, and waiting is right +// either way. +func addInFlight(cluster *simplyblockv1alpha2.StorageCluster) bool { + for _, task := range cluster.Status.Tasks { + if task.Type == nodeAddTask && task.Status != taskDone { + return true + } + } + return false +} + +// taskDone is the control plane's terminal task status. Everything else it +// publishes — new, running, suspended — is a task still being worked through. +const taskDone = "done" + func (r *StorageNodeReconciler) resolveUUID( ctx context.Context, node *simplyblockv1alpha2.StorageNode, diff --git a/operator/internal/discovery/grouping.go b/operator/internal/discovery/grouping.go index ea861c1fd..e0a3b76cb 100644 --- a/operator/internal/discovery/grouping.go +++ b/operator/internal/discovery/grouping.go @@ -54,6 +54,13 @@ type Worker struct { // with 256 GiB may schedule against 250. Its zero value means the planner // was given no node objects. Kube KubeNode + + // MgmtInterface is the interface the draft names for management, or empty + // when the machine presents none that would serve. It is part of what makes + // two workers groupable: a NodeGroup names one interface for every worker in + // it, so machines that call theirs different things describe different + // groups however identical their disks are. + MgmtInterface string } // Addresses is how the draft names this worker's devices, ascending and without @@ -91,6 +98,11 @@ type Group struct { // Class and Addresses are the device selection they share. Class DeviceClass Addresses []string + + // MgmtInterface is the interface every worker in the group binds its + // management address to, which is why it is on the group rather than on the + // workers: a NodeGroup names one. + MgmtInterface string } // Grouper puts workers into groups. @@ -120,11 +132,15 @@ func (GroupByHardware) Group(workers []Worker) []Group { for _, worker := range workers { addresses := worker.Addresses() - signature := worker.Class.signature(addresses) + signature := worker.Class.signature(addresses, worker.MgmtInterface) group, seen := bySignature[signature] if !seen { - group = &Group{Class: worker.Class, Addresses: addresses} + group = &Group{ + Class: worker.Class, + Addresses: addresses, + MgmtInterface: worker.MgmtInterface, + } bySignature[signature] = group order = append(order, signature) } @@ -150,11 +166,17 @@ func (GroupByHardware) Group(workers []Worker) []Group { return groups } -// signature is the key two workers must agree on to share a group: the class -// and the addresses, hashed so that a hundred addresses do not become a -// hundred-element map key. -func (c DeviceClass) signature(addresses []string) string { - digest := sha256.Sum256([]byte(string(c) + "\x00" + strings.Join(addresses, "\x00"))) +// signature is the key two workers must agree on to share a group: the class, +// the addresses, and the management interface, hashed so that a hundred +// addresses do not become a hundred-element map key. +// +// The interface is in the key because a NodeGroup names one for every worker it +// lists. Two machines with identical disks that call their NICs different things +// cannot be described by one group, and grouping them anyway would write a +// document that is wrong for whichever of them lost. +func (c DeviceClass) signature(addresses []string, mgmtInterface string) string { + digest := sha256.Sum256([]byte( + string(c) + "\x00" + mgmtInterface + "\x00" + strings.Join(addresses, "\x00"))) return hex.EncodeToString(digest[:]) } @@ -239,7 +261,11 @@ func nodeGroupOf(group Group) simplyblockv1alpha2.NodeGroup { workers = append(workers, worker.Name) } - out := simplyblockv1alpha2.NodeGroup{Name: group.Name, Workers: workers} + out := simplyblockv1alpha2.NodeGroup{ + Name: group.Name, + Workers: workers, + MgmtInterface: group.MgmtInterface, + } if len(group.Addresses) > 0 { selection := &simplyblockv1alpha2.DeviceSelection{} if group.Class == ClassBlock { diff --git a/operator/internal/discovery/kubenode.go b/operator/internal/discovery/kubenode.go index 75f0c7104..d3efa8588 100644 --- a/operator/internal/discovery/kubenode.go +++ b/operator/internal/discovery/kubenode.go @@ -62,6 +62,12 @@ type KubeNode struct { // nodes and OpenShift usually does not taint its infrastructure ones, so a // fleet's infra nodes pass a taint check and would otherwise land in a draft // as storage workers. + // InternalIP is the address the cluster reaches the machine on, which is + // what decides which of its interfaces a draft names for management: the + // operator addresses a worker by this everywhere else, so naming the + // interface that holds it keeps both halves on one network. + InternalIP string + Role NodeRole // Roles is every role the labels name, which is how a combined deployment @@ -110,6 +116,13 @@ func KubeNodeOf(node corev1.Node) KubeNode { Roles: RolesOf(node), } + for _, address := range node.Status.Addresses { + if address.Type == corev1.NodeInternalIP { + out.InternalIP = address.Address + break + } + } + for _, taint := range node.Spec.Taints { out.Taints = append(out.Taints, fmt.Sprintf("%s=%s:%s", taint.Key, taint.Value, taint.Effect)) } diff --git a/operator/internal/discovery/mgmtiface.go b/operator/internal/discovery/mgmtiface.go new file mode 100644 index 000000000..e35224571 --- /dev/null +++ b/operator/internal/discovery/mgmtiface.go @@ -0,0 +1,97 @@ +// Choosing the interface a storage node binds its management address to. +// +// The control plane refuses a node whose management interface it cannot find an +// IP on, and it refuses it inside the node_add task rather than at the request: +// what an administrator sees is a task that gave up, on a document that looked +// complete. So the draft names one, and names it from what the probe read rather +// than leaving it to be defaulted somewhere further down. +// +// The rule is a ranking rather than a match, because a fleet's machines do not +// agree on what their NICs are called and a draft that named eth0 everywhere +// would be wrong on the machines that call it ens5f0. What every candidate has +// in common is the shape of the answer: hardware, up, and holding an address +// something can reach it on. + +package discovery + +import ( + "cmp" + "net" + "slices" + + "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" +) + +// ManagementInterface is the interface a draft names for this worker, or empty +// when the machine presents none that would serve. +// +// nodeAddress is the address the cluster already reaches the machine on, and +// when it is known the interface holding it wins outright. That is not a +// preference among equals: the operator addresses the worker by that address +// everywhere else it talks to it, so any other choice would have the two halves +// of one deployment describing different networks. +// +// Returning empty is a real answer rather than a failure. A machine whose only +// addressed interfaces are its cluster's own bridges has no management interface, +// and a draft that named one anyway would produce a storage node the rest of the +// fleet cannot reach. +func ManagementInterface(report nodeprobe.Report, nodeAddress string) string { + candidates := make([]nodeprobe.Interface, 0, len(report.Interfaces)) + for _, iface := range report.Interfaces { + if !servesManagement(iface) { + continue + } + if nodeAddress != "" && slices.Contains(iface.Addresses, nodeAddress) { + return iface.Name + } + candidates = append(candidates, iface) + } + if len(candidates) == 0 { + return "" + } + + // Fastest first, then by name. The name is what makes the choice stable: the + // probe's reading order is the kernel's, and a draft that changed between two + // runs of the same fleet is one a reviewer cannot diff. + slices.SortFunc(candidates, func(a, b nodeprobe.Interface) int { + if a.SpeedMbps != b.SpeedMbps { + return cmp.Compare(b.SpeedMbps, a.SpeedMbps) + } + return cmp.Compare(a.Name, b.Name) + }) + return candidates[0].Name +} + +// servesManagement reports whether an interface could carry a storage node's +// management traffic at all. +// +// Each condition rules out a machine this product has actually been deployed +// onto. The virtual ones are the cluster's own plumbing — a CNI bridge, a veth +// to a pod, a flannel overlay, loopback — and every worker has several holding +// addresses that reach nothing outside the node. A link that is down keeps the +// address it was configured with and carries nothing. An interface with no +// address is what the control plane refuses by name. +func servesManagement(iface nodeprobe.Interface) bool { + if iface.Virtual || iface.Bridge || iface.Loopback { + return false + } + if iface.State != "" && iface.State != "up" && iface.State != "unknown" { + return false + } + return slices.ContainsFunc(iface.Addresses, reachable) +} + +// reachable reports whether an address is one something could contact the +// machine on. +// +// A link-local address is configured without anybody assigning it and routes +// nowhere, so an interface holding only those holds nothing usable. An +// unspecified or loopback address is the same case read differently. +func reachable(address string) bool { + ip := net.ParseIP(address) + if ip == nil { + return false + } + return !ip.IsLinkLocalUnicast() && !ip.IsLinkLocalMulticast() && + !ip.IsLoopback() && !ip.IsUnspecified() +} diff --git a/operator/internal/discovery/mgmtiface_test.go b/operator/internal/discovery/mgmtiface_test.go new file mode 100644 index 000000000..70206e8fb --- /dev/null +++ b/operator/internal/discovery/mgmtiface_test.go @@ -0,0 +1,163 @@ +// Which interface a draft names as the management one. +// +// The draft named none, and a storage node added without one is refused by the +// control plane with "No management interface with IP found in provided +// interfaces" — after the node_add task has already started, so the failure +// arrives as a task that gave up rather than as a document that was wrong. + +package discovery + +import ( + "testing" + + "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" +) + +func iface(name string, edit func(*nodeprobe.Interface)) nodeprobe.Interface { + out := nodeprobe.Interface{Name: name, State: "up", NUMANode: 0} + if edit != nil { + edit(&out) + } + return out +} + +// The interface carrying the address the cluster already reaches the machine on +// wins, whatever else is present. +// +// It is the one answer that cannot be wrong: the operator addresses the worker +// by that address everywhere else, so naming the interface that holds it keeps +// the two halves talking about one network. +func TestTheInterfaceHoldingTheNodeAddressWins(t *testing.T) { + report := report("worker-1") + report.Interfaces = []nodeprobe.Interface{ + iface("eth1", func(i *nodeprobe.Interface) { + i.Addresses = []string{"10.10.10.113"} + i.SpeedMbps = 40000 + }), + iface("eth0", func(i *nodeprobe.Interface) { + i.Addresses = []string{"192.168.10.113"} + }), + } + + if got := ManagementInterface(report, "192.168.10.113"); got != "eth0" { + t.Errorf("named %q, want the interface holding the node's address", got) + } +} + +// With no address to match, a physical interface holding an address is named, +// and the fastest such one wins. +func TestTheFastestAddressedPhysicalInterfaceIsNamed(t *testing.T) { + report := report("worker-1") + report.Interfaces = []nodeprobe.Interface{ + iface("eth0", func(i *nodeprobe.Interface) { + i.Addresses = []string{"192.168.10.113"} + i.SpeedMbps = 1000 + }), + iface("eth1", func(i *nodeprobe.Interface) { + i.Addresses = []string{"10.10.10.113"} + i.SpeedMbps = 40000 + }), + } + + if got := ManagementInterface(report, ""); got != "eth1" { + t.Errorf("named %q, want the fastest addressed interface", got) + } +} + +// The interfaces a cluster leaves on every worker are never named, even when +// they hold an address and even when nothing else does. +// +// A bridge carries somebody else's traffic, a veth is one end of a pod's link, +// and loopback reaches nothing. Naming any of them produces a storage node the +// rest of the fleet cannot talk to. +func TestTheClustersOwnInterfacesAreNeverNamed(t *testing.T) { + report := report("worker-1") + report.Interfaces = []nodeprobe.Interface{ + iface("cni0", func(i *nodeprobe.Interface) { + i.Addresses = []string{"10.42.2.1"} + i.Virtual = true + i.Bridge = true + }), + iface("flannel.1", func(i *nodeprobe.Interface) { + i.Addresses = []string{"10.42.2.0"} + i.Virtual = true + }), + iface("lo", func(i *nodeprobe.Interface) { + i.Addresses = []string{"127.0.0.1"} + i.Virtual = true + i.Loopback = true + }), + } + + if got := ManagementInterface(report, ""); got != "" { + t.Errorf("named %q, want nothing rather than a virtual interface", got) + } +} + +// A physical interface with no address is not a management interface: the +// control plane refuses exactly that, and naming one moves the refusal from the +// draft to a task that has already started. +func TestAnInterfaceWithNoAddressIsNotNamed(t *testing.T) { + report := report("worker-1") + report.Interfaces = []nodeprobe.Interface{ + iface("eth0", func(i *nodeprobe.Interface) { i.SpeedMbps = 40000 }), + } + + if got := ManagementInterface(report, ""); got != "" { + t.Errorf("named %q, want nothing rather than an interface with no address", got) + } +} + +// A link that is down holds its address and carries nothing, so it is passed +// over while any interface that is up remains. +func TestALinkThatIsDownIsPassedOver(t *testing.T) { + report := report("worker-1") + report.Interfaces = []nodeprobe.Interface{ + iface("eth0", func(i *nodeprobe.Interface) { + i.Addresses = []string{"192.168.10.113"} + i.SpeedMbps = 40000 + i.State = "down" + }), + iface("eth1", func(i *nodeprobe.Interface) { + i.Addresses = []string{"10.10.10.113"} + i.SpeedMbps = 1000 + }), + } + + if got := ManagementInterface(report, ""); got != "eth1" { + t.Errorf("named %q, want the interface that is up", got) + } +} + +// A link-local address is not one anything reaches the machine on. +func TestALinkLocalAddressDoesNotCount(t *testing.T) { + report := report("worker-1") + report.Interfaces = []nodeprobe.Interface{ + iface("eth0", func(i *nodeprobe.Interface) { + i.Addresses = []string{"169.254.1.1", "fe80::1"} + }), + } + + if got := ManagementInterface(report, ""); got != "" { + t.Errorf("named %q, want nothing rather than a link-local address", got) + } +} + +// Ties are broken by name so that two runs against one fleet write the same +// document, which is what makes a draft reviewable. +func TestTheChoiceIsStableAcrossRuns(t *testing.T) { + report := report("worker-1") + report.Interfaces = []nodeprobe.Interface{ + iface("eth1", func(i *nodeprobe.Interface) { i.Addresses = []string{"10.10.10.113"} }), + iface("eth0", func(i *nodeprobe.Interface) { i.Addresses = []string{"192.168.10.113"} }), + } + + first := ManagementInterface(report, "") + report.Interfaces[0], report.Interfaces[1] = report.Interfaces[1], report.Interfaces[0] + if second := ManagementInterface(report, ""); second != first { + t.Errorf("the reading order changed the answer: %q then %q", first, second) + } + if first != "eth0" { + t.Errorf("named %q, want the first by name", first) + } +} diff --git a/operator/internal/discovery/plan.go b/operator/internal/discovery/plan.go index ea745b87a..430631afd 100644 --- a/operator/internal/discovery/plan.go +++ b/operator/internal/discovery/plan.go @@ -350,13 +350,15 @@ func (p Planner) Plan(reports []nodeprobe.Report, filter *simplyblockv1alpha2.De } } + kube := p.KubeNodes[report.Node] plan.Workers = append(plan.Workers, Worker{ Name: report.Node, Report: report, Devices: chosen, Class: class, PlacementReason: why, - Kube: p.KubeNodes[report.Node], + Kube: kube, + MgmtInterface: ManagementInterface(report, kube.InternalIP), }) } diff --git a/operator/internal/nodeprobe/collect.go b/operator/internal/nodeprobe/collect.go index 45101ca1d..19702b9fb 100644 --- a/operator/internal/nodeprobe/collect.go +++ b/operator/internal/nodeprobe/collect.go @@ -105,6 +105,8 @@ func interfacesOf(ifaces []inventory.Interface) []Interface { NUMANode: iface.NUMANode, Virtual: iface.Virtual, Loopback: iface.Loopback, + Bridge: iface.Bridge, + Addresses: iface.Addresses, }) } return out diff --git a/operator/internal/nodeprobe/report.go b/operator/internal/nodeprobe/report.go index ef1b39954..3c4263e79 100644 --- a/operator/internal/nodeprobe/report.go +++ b/operator/internal/nodeprobe/report.go @@ -38,7 +38,7 @@ import ( // because a probe pod outlives the operator that created it across an upgrade: // the image is pinned in the Job, and a Job already running keeps the image it // started with. -const ReportVersion = 2 +const ReportVersion = 3 // Report is one worker's inventory as the probe found it. type Report struct { @@ -209,6 +209,20 @@ type Interface struct { Virtual bool `json:"virtual,omitempty"` Loopback bool `json:"loopback,omitempty"` + + // Bridge reports whether the interface is a software bridge, which a + // cluster's own CNI leaves on every worker. It is separate from Virtual + // because the two answer different questions: Virtual says the interface is + // backed by no hardware, and Bridge says it is carrying somebody else's + // traffic, which is what disqualifies it from being a management address + // even where it holds one. + Bridge bool `json:"bridge,omitempty"` + + // Addresses are the IP addresses the interface holds, without a prefix + // length. They are what makes a management interface identifiable: the one + // a draft names is the one carrying the address the cluster already reaches + // the machine on. + Addresses []string `json:"addresses,omitempty"` } // Device is one block device and whether it may be handed to a storage cluster. From 95e3fb74e6697b8d3c9f4ecaa3d1db14d8e42b7a Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Tue, 15 Sep 2026 17:05:17 +0200 Subject: [PATCH 009/206] fix(cluster): a task that gave up is not reported as a completion The window publishes what is running, so a task's whole history is the one event raised when it leaves it. That event said the task was no longer running, which is as true of one that succeeded as of one that exhausted its retries, and it was raised as Normal either way. The control plane says which. A node_add that gave up came back with retry 11 and a function_result of max retry reached, and the operator decoded both and read neither: the event was built from the previous snapshot, and the finished task was skipped by the loop that had it in hand. So the finished task is kept by id rather than skipped, and the event carries the outcome. The retry count decides the verdict, because the status says done for both and the schema already names the count as the one number separating a task that is slow from one that is failing. The control plane's own result goes into the note whatever the verdict, since that is where the answer actually is. The recorder in the tests kept the reason and dropped the message, so a test could not have caught any of this. It keeps both now. Co-Authored-By: Claude Opus 5 (1M context) --- .../internal/controllers/cluster/events.go | 7 + .../controllers/cluster/helpers_test.go | 34 ++++- .../cluster/storagecluster_controller.go | 61 +++++++- .../controllers/cluster/tasks_test.go | 137 ++++++++++++++++++ 4 files changed, 233 insertions(+), 6 deletions(-) create mode 100644 operator/internal/controllers/cluster/tasks_test.go diff --git a/operator/internal/controllers/cluster/events.go b/operator/internal/controllers/cluster/events.go index 665b204ff..580e6184d 100644 --- a/operator/internal/controllers/cluster/events.go +++ b/operator/internal/controllers/cluster/events.go @@ -37,6 +37,13 @@ const ( TaskCompleted = "TaskCompleted" TaskCanceled = "TaskCanceled" + // TaskGaveUp is a task that left the window having been restarted, which is + // the control plane reporting a failure rather than a finish. Its own status + // says done either way, so the retry count is what separates them, and the + // schema says as much: it is the one number that separates a task that is + // slow from one that is failing. + TaskGaveUp = "TaskGaveUp" + // The operation reasons, raised on the StorageClusterOps rather than the // cluster. // diff --git a/operator/internal/controllers/cluster/helpers_test.go b/operator/internal/controllers/cluster/helpers_test.go index 62a99c510..ad1cfb7d2 100644 --- a/operator/internal/controllers/cluster/helpers_test.go +++ b/operator/internal/controllers/cluster/helpers_test.go @@ -12,6 +12,7 @@ package cluster import ( "context" + "fmt" "testing" corev1 "k8s.io/api/core/v1" @@ -138,12 +139,41 @@ type recorder struct { type recordedEvent struct { Type string Reason string + + // Note is the formatted message. A reason says which event happened and the + // note says what it was about, and some of what an event owes a reader is + // only in the second. + Note string } func (r *recorder) Eventf( - _ runtime.Object, _ runtime.Object, eventType, reason, _, _ string, _ ...any, + _ runtime.Object, _ runtime.Object, eventType, reason, _, note string, args ...any, ) { - r.events = append(r.events, recordedEvent{Type: eventType, Reason: reason}) + r.events = append(r.events, recordedEvent{ + Type: eventType, + Reason: reason, + Note: fmt.Sprintf(note, args...), + }) +} + +// noteFor is the message of the first event carrying a reason. +func (r *recorder) noteFor(reason string) string { + for _, e := range r.events { + if e.Reason == reason { + return e.Note + } + } + return "" +} + +// typeFor is the severity of the first event carrying a reason. +func (r *recorder) typeFor(reason string) string { + for _, e := range r.events { + if e.Reason == reason { + return e.Type + } + } + return "" } // count returns how many events carried a reason, which is what a test diff --git a/operator/internal/controllers/cluster/storagecluster_controller.go b/operator/internal/controllers/cluster/storagecluster_controller.go index 117858948..0fc48570e 100644 --- a/operator/internal/controllers/cluster/storagecluster_controller.go +++ b/operator/internal/controllers/cluster/storagecluster_controller.go @@ -754,9 +754,15 @@ func (r *StorageClusterReconciler) readTasks( } current := make(map[string]bool, len(reported)) + // A finished task is kept by id rather than skipped, because it is the only + // place the outcome exists. The window publishes what is running, so a task + // that has ended is described once, here, and what the control plane said + // about it is gone on the next read. + finished := make(map[string]subscriptions.TaskDTO, len(reported)) running := make([]simplyblockv1alpha2.ClusterTask, 0, len(reported)) for _, task := range reported { if task.Finished() { + finished[task.ID] = task continue } current[task.ID] = true @@ -776,18 +782,65 @@ func (r *StorageClusterReconciler) readTasks( } // A task that was in the list and is no longer running finished between - // two readings. The event is what remains of it. + // two readings. The event is what remains of it, so it carries the outcome + // rather than only the disappearance: "is no longer running" is as true of a + // task that succeeded as of one that exhausted its retries, and a reader + // looking for why a deployment stalled needs the difference. for _, previous := range cluster.Status.Tasks { if current[previous.ID] { continue } - r.Recorder.Eventf(cluster, nil, corev1.EventTypeNormal, - TaskCompleted, TaskCompleted, - "Task %s (%s) is no longer running", previous.ID, previous.Type) + r.emitTaskOutcome(cluster, previous, finished[previous.ID]) } return running } +// emitTaskOutcome raises the one event a finished task gets. +// +// The retry count decides which. A task the control plane never restarted ran +// once and ended, and a task it restarted was failing each time it did, so one +// that has left the window having been retried gave up rather than finished. +// Nothing here infers that from the status, which says done for both. +// +// The count and the control plane's own result go into the note whatever the +// verdict. The result is where the answer actually is — a node_add that gave up +// came back with "max retry reached (11/11)" — and it was decoded and dropped. +func (r *StorageClusterReconciler) emitTaskOutcome( + cluster *simplyblockv1alpha2.StorageCluster, + previous simplyblockv1alpha2.ClusterTask, + outcome subscriptions.TaskDTO, +) { + // The task may have scrolled out of the window rather than been read as + // finished, in which case the previous snapshot is all there is. + retries := previous.Retry + if outcome.Retry > retries { + retries = outcome.Retry + } + + detail := "" + if outcome.Result != "" { + detail = ": " + outcome.Result + } + + if outcome.Canceled { + r.Recorder.Eventf(cluster, nil, corev1.EventTypeWarning, + TaskCanceled, TaskCanceled, + "Task %s (%s) was canceled after %d retries%s", + previous.ID, previous.Type, retries, detail) + return + } + if retries > 0 { + r.Recorder.Eventf(cluster, nil, corev1.EventTypeWarning, + TaskGaveUp, TaskGaveUp, + "Task %s (%s) gave up after %d retries%s", + previous.ID, previous.Type, retries, detail) + return + } + r.Recorder.Eventf(cluster, nil, corev1.EventTypeNormal, + TaskCompleted, TaskCompleted, + "Task %s (%s) finished%s", previous.ID, previous.Type, detail) +} + // reportedTasks is every task of one cluster, from the stream's cache once it // has delivered the cluster's snapshot and from the control plane until then. func (r *StorageClusterReconciler) reportedTasks( diff --git a/operator/internal/controllers/cluster/tasks_test.go b/operator/internal/controllers/cluster/tasks_test.go new file mode 100644 index 000000000..7364bea8c --- /dev/null +++ b/operator/internal/controllers/cluster/tasks_test.go @@ -0,0 +1,137 @@ +// What the event says about a task that has left the window. +// +// The window is what is running now, so a task's whole history is the one event +// emitted when it leaves. That event said only that the task was no longer +// running, which is true of one that succeeded and of one that exhausted its +// retries, and the difference is the whole of what a reader wants. +// +// The control plane publishes it. A node_add that gave up came back with retry +// 11 and a function_result of max retry reached, decoded into the DTO and read +// by nothing, while the event was raised as Normal from the previous snapshot. + +package cluster + +import ( + "context" + "strings" + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/cpinformer/subscriptions" +) + +// aClusterRunning is a cluster whose status already records the task, which is +// what makes its disappearance visible on the next read. +func aClusterRunning(tasks ...simplyblockv1alpha2.ClusterTask) *simplyblockv1alpha2.StorageCluster { + cluster := &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{Name: "a-cluster", Namespace: "simplyblock"}, + } + cluster.Status.UUID = "cluster-uuid" + cluster.Status.Tasks = tasks + return cluster +} + +func aTaskReader(rec *recorder, reported ...subscriptions.TaskDTO) *StorageClusterReconciler { + return &StorageClusterReconciler{ + Recorder: rec, + API: &fakeControlPlane{ + tasks: func(string) ([]subscriptions.TaskDTO, error) { return reported, nil }, + }, + } +} + +// A task that gave up says so, with what the control plane said about it. +func TestATaskThatExhaustedItsRetriesIsReportedAsOne(t *testing.T) { + rec := &recorder{} + cluster := aClusterRunning(simplyblockv1alpha2.ClusterTask{ + ID: "584ef009", Type: "node_add", Status: "running", Retry: 10, + }) + r := aTaskReader(rec, subscriptions.TaskDTO{ + ID: "584ef009", + Type: "node_add", + Status: "done", + Retry: 11, + Result: "max retry reached (11/11)", + }) + + r.readTasks(context.Background(), cluster) + + if !rec.has(TaskGaveUp) { + t.Fatalf("a task that exhausted its retries raised %+v", rec.events) + } + if got := rec.typeFor(TaskGaveUp); got != corev1.EventTypeWarning { + t.Errorf("the event is %q, want a warning", got) + } + note := rec.noteFor(TaskGaveUp) + for _, want := range []string{"node_add", "max retry reached (11/11)", "11"} { + if !strings.Contains(note, want) { + t.Errorf("the note does not carry %q: %s", want, note) + } + } + if rec.has(TaskCompleted) { + t.Error("a task that gave up was also reported as completed") + } +} + +// A task that finished without ever being restarted is a completion, and stays +// the quiet event it was. +func TestATaskThatFinishedCleanlyIsStillACompletion(t *testing.T) { + rec := &recorder{} + cluster := aClusterRunning(simplyblockv1alpha2.ClusterTask{ + ID: "t-1", Type: "cluster_status", Status: "running", + }) + r := aTaskReader(rec) + + r.readTasks(context.Background(), cluster) + + if !rec.has(TaskCompleted) { + t.Fatalf("a clean finish raised %+v", rec.events) + } + if got := rec.typeFor(TaskCompleted); got != corev1.EventTypeNormal { + t.Errorf("the event is %q, want a normal one", got) + } + if rec.has(TaskGaveUp) { + t.Error("a task that never retried was reported as having given up") + } +} + +// The retry count is the schema's own signal that a task was failing rather +// than slow, so a task that left the window having been restarted is reported +// as having given up even when the control plane said nothing else about it. +func TestARetriedTaskIsReportedEvenWithNoResult(t *testing.T) { + rec := &recorder{} + cluster := aClusterRunning(simplyblockv1alpha2.ClusterTask{ + ID: "t-2", Type: "node_add", Status: "running", Retry: 3, + }) + r := aTaskReader(rec) + + r.readTasks(context.Background(), cluster) + + if !rec.has(TaskGaveUp) { + t.Fatalf("a retried task that vanished raised %+v", rec.events) + } + if note := rec.noteFor(TaskGaveUp); !strings.Contains(note, "3") { + t.Errorf("the note does not carry the retry count: %s", note) + } +} + +// A task still running is not reported at all, which is what keeps the event a +// record of history rather than a heartbeat. +func TestARunningTaskRaisesNothing(t *testing.T) { + rec := &recorder{} + cluster := aClusterRunning(simplyblockv1alpha2.ClusterTask{ + ID: "t-3", Type: "node_add", Status: "running", + }) + r := aTaskReader(rec, subscriptions.TaskDTO{ + ID: "t-3", Type: "node_add", Status: "running", + }) + + r.readTasks(context.Background(), cluster) + + if len(rec.events) != 0 { + t.Errorf("a running task raised %+v", rec.events) + } +} From 80f14903cca33c2aaf3533d4b6fe4158aa72106d Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Tue, 15 Sep 2026 17:28:42 +0200 Subject: [PATCH 010/206] fix(operator): a discovery run's status write retries instead of repeating a step A step does its work and then records it, so a recording that fails leaves the step not having happened as far as the next pass is concerned, and the next pass does the work again. Writing's work is creating a document and telling a reviewer about it. The status write was a bare Update of the object the reconcile had started from, so anything that touched the run in between lost it the race. What that produced on a live cluster was a run that wrote one document, reported writing it twice, and reported finding it already there -- describing to a reviewer a race that nobody had, in the only record a finished run leaves. It is now a patch against a fresh read with an optimistic lock, retried here rather than paid for by the step. That is the shape the deployment config's own status write in this package already uses, and the one reconciler in the tree that was not using it. The two other bare status updates are left alone and are not the same hazard: one is a best-effort clear that logs and continues, and the other is a one-shot upgrade step rather than a reconcile loop. The tests reach the failure through a client interceptor, because a fake client never produces a conflict and so could not have caught this. Co-Authored-By: Claude Opus 5 (1M context) --- .../deployment/operatorops_conflict_test.go | 161 ++++++++++++++++++ .../deployment/operatorops_controller.go | 42 ++++- .../deployment/operatorops_unit_test.go | 12 ++ 3 files changed, 214 insertions(+), 1 deletion(-) create mode 100644 operator/internal/controllers/deployment/operatorops_conflict_test.go diff --git a/operator/internal/controllers/deployment/operatorops_conflict_test.go b/operator/internal/controllers/deployment/operatorops_conflict_test.go new file mode 100644 index 000000000..0d0a3b944 --- /dev/null +++ b/operator/internal/controllers/deployment/operatorops_conflict_test.go @@ -0,0 +1,161 @@ +// What a losing status write costs a discovery run. +// +// A step does its work and then records that it did it. If the recording fails +// the step has not happened as far as the next pass is concerned, so the next +// pass does the work again — and the work of Writing is creating a document and +// telling a reviewer about it. +// +// The create survives that, because it is idempotent by name and says so. The +// events do not: a run that wrote one document reported writing it twice and +// reported finding it already there, which describes a race that did not happen +// to a reviewer reading the only record the run leaves. + +package deployment + +import ( + "context" + "testing" + + apierrors "k8s.io/apimachinery/pkg/api/errors" + "k8s.io/apimachinery/pkg/runtime" + "k8s.io/apimachinery/pkg/runtime/schema" + "k8s.io/client-go/tools/events" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/interceptor" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// conflictOnce fails the first status write the way the API server does when +// something else has written the object since it was read. +func conflictOnce(remaining *int) interceptor.Funcs { + // Both verbs are hooked, because which one a status write uses is the + // reconciler's business and the conflict is the test's. + conflict := func(name string) error { + *remaining-- + return apierrors.NewConflict( + schema.GroupResource{ + Group: simplyblockv1alpha2.GroupVersion.Group, + Resource: "operatorops", + }, + name, + errStale, + ) + } + return interceptor.Funcs{ + SubResourceUpdate: func( + ctx context.Context, c client.Client, subResource string, + obj client.Object, opts ...client.SubResourceUpdateOption, + ) error { + if *remaining > 0 { + return conflict(obj.GetName()) + } + return c.SubResource(subResource).Update(ctx, obj, opts...) + }, + SubResourcePatch: func( + ctx context.Context, c client.Client, subResource string, + obj client.Object, patch client.Patch, opts ...client.SubResourcePatchOption, + ) error { + if *remaining > 0 { + return conflict(obj.GetName()) + } + return c.SubResource(subResource).Patch(ctx, obj, patch, opts...) + }, + } +} + +// countingRecorder remembers how many events carried each reason, which is what +// separates a step that ran once from one that ran twice. +type countingRecorder struct { + reasons map[string]int +} + +func (r *countingRecorder) Eventf( + _ runtime.Object, _ runtime.Object, _, reason, _, _ string, _ ...any, +) { + if r.reasons == nil { + r.reasons = map[string]int{} + } + r.reasons[reason]++ +} + +func (r *countingRecorder) count(reason string) int { return r.reasons[reason] } + +var _ events.EventRecorder = (*countingRecorder)(nil) + +// errStale is what the API server's conflict carries. +var errStale = errConflict{} + +type errConflict struct{} + +func (errConflict) Error() string { + return "the object has been modified; please apply your changes to the latest version and try again" +} + +// A status write that loses a race is retried rather than failing the step. +// +// Without it the step's own work is repeated on the next pass, and the work is +// not all repeatable: the document is, because creating it again finds it and +// says so, but the events describing the run are emitted a second time. +func TestAStatusWriteThatLosesARaceIsRetried(t *testing.T) { + remaining := 1 + r := newRunnerWithInterceptors(t, conflictOnce(&remaining), + discoverRun(nil), worker("worker-1"), worker("worker-2")) + + r.step() // start + r.step() // inspect + r.step() // probing: creates the Jobs + + for _, node := range []string{"worker-1", "worker-2"} { + cm := reportConfigMap(t, node, "0000:5e:00.0", "0000:5f:00.0") + if err := r.client.Create(context.Background(), cm); err != nil { + t.Fatalf("write a report: %v", err) + } + } + + r.step() // probing: sees the reports, moves to Writing + _, ops := r.step() + + if ops.Status.Phase != simplyblockv1alpha2.OperatorOpsPhaseSucceeded { + t.Fatalf("the run is %q after a conflict it should have retried: %s", + ops.Status.Phase, ops.Status.Message) + } + if remaining != 0 { + t.Error("the conflict was never raised, so this proves nothing") + } + if len(r.configs()) != 1 { + t.Errorf("the run left %d documents", len(r.configs())) + } +} + +// The step's side effects happen once, which is what the retry is for: a second +// pass over Writing re-creates the document and finds it already there, and +// ConfigExists is a race being reported to somebody who did not have one. +func TestALostStatusWriteDoesNotReportAPhantomRace(t *testing.T) { + remaining := 1 + rec := &countingRecorder{} + r := newRunnerWithInterceptors(t, conflictOnce(&remaining), + discoverRun(nil), worker("worker-1"), worker("worker-2")) + r.reconciler.Recorder = rec + + r.step() // start + r.step() // inspect + r.step() // probing + + for _, node := range []string{"worker-1", "worker-2"} { + cm := reportConfigMap(t, node, "0000:5e:00.0", "0000:5f:00.0") + if err := r.client.Create(context.Background(), cm); err != nil { + t.Fatalf("write a report: %v", err) + } + } + + r.step() // probing + r.step() // writing + + if got := rec.count("ConfigExists"); got != 0 { + t.Errorf("the run reported finding its own document %d time(s)", got) + } + if got := rec.count("OperationSucceeded"); got != 1 { + t.Errorf("the run reported succeeding %d times", got) + } +} diff --git a/operator/internal/controllers/deployment/operatorops_controller.go b/operator/internal/controllers/deployment/operatorops_controller.go index f14921039..f5836993f 100644 --- a/operator/internal/controllers/deployment/operatorops_controller.go +++ b/operator/internal/controllers/deployment/operatorops_controller.go @@ -38,6 +38,7 @@ import ( "k8s.io/apimachinery/pkg/runtime" "k8s.io/client-go/discovery" "k8s.io/client-go/tools/events" + "k8s.io/client-go/util/retry" ctrl "sigs.k8s.io/controller-runtime" "sigs.k8s.io/controller-runtime/pkg/client" logf "sigs.k8s.io/controller-runtime/pkg/log" @@ -646,12 +647,51 @@ func (r *OperatorOpsReconciler) fail( // status writes the run's status, stamping the generation it was computed from // so that a stale status can be told from a current one. +// status persists what the step concluded, retrying a write that lost a race. +// +// A step does its work and then records it, so a recording that fails leaves the +// step not having happened as far as the next pass is concerned — and the next +// pass does the work again. Writing's work is creating a document and telling a +// reviewer about it. The document survives being created twice, because the +// create is idempotent by name and says so; the events do not, and a run that +// wrote one document reported writing it twice and reported finding it already +// there, which describes to a reviewer a race that nobody had. +// +// So the write is a patch against a fresh read rather than an update of the +// object the reconcile started from, and a conflict is retried here rather than +// paid for by the step. It is the same shape the deployment config's own status +// write uses, for the same reason. func (r *OperatorOpsReconciler) status( ctx context.Context, ops *simplyblockv1alpha2.OperatorOps, ) error { ops.Status.ObservedGeneration = ops.Generation - return r.Status().Update(ctx, ops) + desired := *ops.Status.DeepCopy() + + err := retry.RetryOnConflict(retry.DefaultRetry, func() error { + var fresh simplyblockv1alpha2.OperatorOps + if err := r.Get(ctx, client.ObjectKeyFromObject(ops), &fresh); err != nil { + return err + } + + patch := client.MergeFromWithOptions(fresh.DeepCopy(), + client.MergeFromWithOptimisticLock{}) + fresh.Status = desired + fresh.Status.ObservedGeneration = fresh.Generation + if err := r.Status().Patch(ctx, &fresh, patch); err != nil { + return err + } + + // The caller goes on using its own object, so it has to carry the + // version the patch produced or its next write conflicts with itself. + ops.Status = fresh.Status + ops.ResourceVersion = fresh.ResourceVersion + return nil + }) + if err != nil { + return fmt.Errorf("record the run's status: %w", err) + } + return nil } // event records one, when there is a recorder to record it with. diff --git a/operator/internal/controllers/deployment/operatorops_unit_test.go b/operator/internal/controllers/deployment/operatorops_unit_test.go index c1cef5d07..436e58207 100644 --- a/operator/internal/controllers/deployment/operatorops_unit_test.go +++ b/operator/internal/controllers/deployment/operatorops_unit_test.go @@ -29,6 +29,7 @@ import ( ctrl "sigs.k8s.io/controller-runtime" "sigs.k8s.io/controller-runtime/pkg/client" "sigs.k8s.io/controller-runtime/pkg/client/fake" + "sigs.k8s.io/controller-runtime/pkg/client/interceptor" "github.com/simplyblock/atlas/blockdev" simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" @@ -162,12 +163,23 @@ type runner struct { // newRunner builds a reconciler over the objects given. func newRunner(t *testing.T, objects ...client.Object) *runner { + t.Helper() + return newRunnerWithInterceptors(t, interceptor.Funcs{}, objects...) +} + +// newRunnerWithInterceptors is the same runner with the client's answers +// scripted, which is how a test reaches the failures a real API server produces +// and a fake one never does. +func newRunnerWithInterceptors( + t *testing.T, funcs interceptor.Funcs, objects ...client.Object, +) *runner { t.Helper() scheme := opsScheme(t) c := fake.NewClientBuilder(). WithScheme(scheme). WithObjects(objects...). WithStatusSubresource(&simplyblockv1alpha2.OperatorOps{}). + WithInterceptorFuncs(funcs). Build() return &runner{ From 1132909dfb5e8c3c95028636d3c0567dea31243d Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Tue, 15 Sep 2026 17:39:55 +0200 Subject: [PATCH 011/206] fix(operator): the device mirror's labels fit the limit that binds them A StorageNode's name is a cluster name, a worker hostname, and a slot, so it outgrows the 63 bytes a label value allows on any fleet whose machines carry fully qualified hostnames. The mirror wrote that name into a label unchanged: metadata.labels: Invalid value: "discovered-initial-discovery-cluster-vm03.simplyblock4.localdomain-0": must be no more than 63 bytes What that costs is not a truncated label. The API server refuses the whole object, so the mirror for every device on that node fails to reconcile and none of them is ever published -- the devices of the machines with the longest names, which is every machine on a fleet that uses domain names. The three values now go through the kube.Formula the package already has for this, which leaves a value that fits exactly as it is and cuts one that does not with a digest over the whole input. Both halves matter. The labels exist for a person to select on, so a value nobody can type is a selector nobody can write; and cutting without the digest would map every node of a long-named cluster to one label, answering "which devices are in this node" with the whole cluster's. Only these three are changed. The same raw names reach labels that a DaemonSet and a disruption budget select on, where the writer and the selector have to move together; these are written and read back by nothing, which is what makes them safe to correct alone. Co-Authored-By: Claude Opus 5 (1M context) --- .../node/storagedevice_controller.go | 27 ++++- .../node/storagedevice_labels_test.go | 110 ++++++++++++++++++ 2 files changed, 134 insertions(+), 3 deletions(-) create mode 100644 operator/internal/controllers/node/storagedevice_labels_test.go diff --git a/operator/internal/controllers/node/storagedevice_controller.go b/operator/internal/controllers/node/storagedevice_controller.go index 76e830549..8b4eee583 100644 --- a/operator/internal/controllers/node/storagedevice_controller.go +++ b/operator/internal/controllers/node/storagedevice_controller.go @@ -27,7 +27,9 @@ import ( "sigs.k8s.io/controller-runtime/pkg/handler" "sigs.k8s.io/controller-runtime/pkg/source" + atlaskube "github.com/simplyblock/atlas/kube" "github.com/simplyblock/atlas/ptr" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" "github.com/simplyblock/simplyblock-operator/internal/cpinformer" "github.com/simplyblock/simplyblock-operator/internal/cpinformer/subscriptions" @@ -417,20 +419,39 @@ func (r *StorageDeviceReconciler) upsert( func (r *StorageDeviceReconciler) deviceLabels( node *simplyblockv1alpha2.StorageNode, ) map[string]string { - labels := map[string]string{simplyblockv1alpha2.DeviceLabelNode: node.Name} + labels := map[string]string{ + simplyblockv1alpha2.DeviceLabelNode: labelValue(node.Name), + } if worker := node.Labels[simplyblockv1alpha2.DeviceLabelWorker]; worker != "" { - labels[simplyblockv1alpha2.DeviceLabelWorker] = worker + labels[simplyblockv1alpha2.DeviceLabelWorker] = labelValue(worker) } // The cluster is on the node itself now. It used to be reached through the // StorageNodeSet the node belonged to, which made a label on a device depend on // a third object being readable; a node names its own cluster, so there is // nothing left to look up (design-storagenode.md §3.1). if node.Spec.ClusterRef != "" { - labels[simplyblockv1alpha2.DeviceLabelCluster] = node.Spec.ClusterRef + labels[simplyblockv1alpha2.DeviceLabelCluster] = labelValue(node.Spec.ClusterRef) } return labels } +// labelValue holds a name to what a label value may be. +// +// A StorageNode's name is a cluster name, a worker hostname, and a slot, so it +// outgrows the 63 bytes a label allows on any fleet whose machines carry fully +// qualified hostnames. What that costs is not a truncated label but the object: +// the API server refuses the whole write, so the mirror for every device on that +// node fails to reconcile and none of them is ever published. +// +// The formula leaves a value that already fits exactly as it is, which is what +// the labels are for — a person selects on them, and a value nobody can type is +// a selector nobody can write. A value that does not fit is cut and carries a +// digest, because cutting alone would map every node of a long-named cluster to +// one label and answer "which devices are in this node" with the cluster's. +func labelValue(name string) string { + return atlaskube.Formula{Kind: atlaskube.LabelValue}.Derive(name).Value +} + // deviceOwnedLabels are the keys the mirror writes and is therefore responsible // for removing. A key is on this list whether or not the current reconcile could // resolve a value for it, which is what makes the removal possible at all. diff --git a/operator/internal/controllers/node/storagedevice_labels_test.go b/operator/internal/controllers/node/storagedevice_labels_test.go new file mode 100644 index 000000000..7ae1692e3 --- /dev/null +++ b/operator/internal/controllers/node/storagedevice_labels_test.go @@ -0,0 +1,110 @@ +// The labels the device mirror writes, against the limit that binds them. +// +// A label value is 63 bytes and a StorageNode's name is built from a cluster +// name, a worker hostname, and a slot, so the name outgrows the label on any +// fleet whose machines have fully qualified hostnames. What that costs is not a +// truncated label: the API server refuses the whole object, so the mirror for +// every device on that node fails to reconcile and the devices are never +// published at all. + +package node + +import ( + "strings" + "testing" + + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + + atlaskube "github.com/simplyblock/atlas/kube" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// aNodeNamed is a StorageNode carrying the worker label the mirror copies. +func aNodeNamed(name, cluster, worker string) *simplyblockv1alpha2.StorageNode { + return &simplyblockv1alpha2.StorageNode{ + ObjectMeta: metav1.ObjectMeta{ + Name: name, + Namespace: "simplyblock", + Labels: map[string]string{simplyblockv1alpha2.DeviceLabelWorker: worker}, + }, + Spec: simplyblockv1alpha2.StorageNodeSpec{ClusterRef: cluster}, + } +} + +// Every label the mirror writes fits, whatever the names it is built from. +// +// The name below is the one this was found on: a cluster a discovery run named +// after itself, a worker with a fully qualified hostname, and a slot. It is 68 +// bytes, and the API server refused it. +func TestTheDeviceLabelsFitTheLimit(t *testing.T) { + r := &StorageDeviceReconciler{} + node := aNodeNamed( + "discovered-initial-discovery-cluster-vm03.simplyblock4.localdomain-0", + "discovered-initial-discovery-cluster", + "vm03.simplyblock4.localdomain", + ) + + for key, value := range r.deviceLabels(node) { + if len(value) > atlaskube.MaxLabelValueLength { + t.Errorf("%s is %d bytes: %s", key, len(value), value) + } + if errs := atlaskube.Validate(atlaskube.LabelValue, value); len(errs) > 0 { + t.Errorf("%s is not a legal label value: %v", key, errs) + } + } +} + +// A name that already fits is written exactly, because the labels exist for a +// person to select on and a value nobody can type is a selector nobody can +// write. +func TestALabelThatFitsIsNotRewritten(t *testing.T) { + r := &StorageDeviceReconciler{} + node := aNodeNamed("a-cluster-worker-1-0", "a-cluster", "worker-1") + + labels := r.deviceLabels(node) + for key, want := range map[string]string{ + simplyblockv1alpha2.DeviceLabelNode: "a-cluster-worker-1-0", + simplyblockv1alpha2.DeviceLabelCluster: "a-cluster", + simplyblockv1alpha2.DeviceLabelWorker: "worker-1", + } { + if got := labels[key]; got != want { + t.Errorf("%s = %q, want the name unchanged: %q", key, got, want) + } + } +} + +// Two nodes that differ only past the limit get different labels, so a selector +// on one does not return the other's devices. +// +// This is the whole reason the truncation carries a digest. Cutting at 63 bytes +// would map every node of a long-named cluster to one value, and a person asking +// which devices are in a failed node would be handed the whole cluster's. +func TestTwoLongNamesDoNotCollide(t *testing.T) { + r := &StorageDeviceReconciler{} + prefix := strings.Repeat("a", 60) + + first := r.deviceLabels(aNodeNamed(prefix+"-vm03-0", "c", "w")) + second := r.deviceLabels(aNodeNamed(prefix+"-vm04-0", "c", "w")) + + key := simplyblockv1alpha2.DeviceLabelNode + if first[key] == second[key] { + t.Errorf("two nodes share the label %q", first[key]) + } +} + +// A worker label the node does not carry is still left out rather than invented, +// which is the behavior the mirror already had and the sanitizing must not lose. +func TestAnAbsentSourceIsStillLeftOut(t *testing.T) { + r := &StorageDeviceReconciler{} + node := aNodeNamed("a-cluster-worker-1-0", "", "") + node.Labels = nil + + labels := r.deviceLabels(node) + if _, present := labels[simplyblockv1alpha2.DeviceLabelWorker]; present { + t.Error("a worker label was invented from a node that carries none") + } + if _, present := labels[simplyblockv1alpha2.DeviceLabelCluster]; present { + t.Error("a cluster label was invented from a node that names none") + } +} From 989cc3e8b4b48bfb07dfabde518ed71aa969ac5f Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Tue, 15 Sep 2026 17:59:18 +0200 Subject: [PATCH 012/206] fix(deployment): a document does not refuse the cluster it created Re-entering a step is ordinary. A status write loses a race, approval bumps the generation, the operator restarts mid-expansion -- and the step runs again. So a step has to be idempotent against its own prior success, and this one read "a StorageCluster by that name already exists" and refused, on a document whose own status.clusterRef said it had put it there. What that cost was the deployment. The config went Failed at AwaitingCluster naming the cluster it had created itself, and the nodes it had not created yet were never created: spec.cluster.name is discovered-initial-discovery-cluster and a StorageCluster by that name already exists; set spec.clusterRef to add nodes to it instead status.clusterRef is the record that separates the two cases, and it was already written by the pass that created the cluster. A document that finds its own cluster resumes; one that finds a cluster it never recorded still refuses, which is what the check was for. The AlreadyExists branch gets the same treatment, because a cached read that has not caught up with this document's own create reaches the identical state by the other route. Co-Authored-By: Claude Opus 5 (1M context) --- .../controllers/deployment/expansion.go | 21 +++++++-- .../deployment/expansion_review_test.go | 44 +++++++++++++++++++ 2 files changed, 61 insertions(+), 4 deletions(-) diff --git a/operator/internal/controllers/deployment/expansion.go b/operator/internal/controllers/deployment/expansion.go index c06c086b0..7839a8256 100644 --- a/operator/internal/controllers/deployment/expansion.go +++ b/operator/internal/controllers/deployment/expansion.go @@ -113,11 +113,18 @@ func (r *ClusterDeploymentConfigReconciler) createCluster( getErr := r.Get(ctx, key, &existing) switch { + case getErr == nil && config.Status.ClusterRef == name: + // This document created it on an earlier pass and said so. Re-entering a + // step is ordinary — a lost status write, a generation bump, a restart + // mid-expansion — so a step that refused its own prior success would fail + // the deployment on a retry rather than resume it. + return true, r.recordCluster(ctx, config, name) + case getErr == nil && config.Spec.ClusterRef == "": - // The document asked to create a cluster and one is already there. - // Merging would have the operator decide what a difference means, and the - // differences that matter are of the form "this node's device list - // changed" (§6). + // The document asked to create a cluster and one is already there that it + // did not put there. Merging would have the operator decide what a + // difference means, and the differences that matter are of the form "this + // node's device list changed" (§6). return false, refusef(ClusterExists, "spec.cluster.name is %s and a StorageCluster by that name already exists; "+ "set spec.clusterRef to add nodes to it instead", name) @@ -141,6 +148,12 @@ func (r *ClusterDeploymentConfigReconciler) createCluster( } if err := r.Create(ctx, cluster); err != nil { if apierrors.IsAlreadyExists(err) { + if config.Status.ClusterRef == name { + // The read above was served from a cache that had not caught up + // with this document's own earlier create. The same resumption as + // at the top of the switch, reached by the other route. + return true, r.recordCluster(ctx, config, name) + } // The read above missed and somebody created the cluster between the // two. That is the ClusterExists case arriving by a different route, // not a success: the document asked to create a cluster and did not, diff --git a/operator/internal/controllers/deployment/expansion_review_test.go b/operator/internal/controllers/deployment/expansion_review_test.go index 711af1e1d..66708d6d4 100644 --- a/operator/internal/controllers/deployment/expansion_review_test.go +++ b/operator/internal/controllers/deployment/expansion_review_test.go @@ -294,3 +294,47 @@ func saidSomethingAbout(findings []finding, want string) bool { } return false } + +// A document that already created its cluster does not refuse it on the next +// pass. +// +// Re-entering a step is ordinary: a lost status write, a generation bump, an +// operator restart mid-expansion. The step has to be idempotent against its own +// prior success, and this one was not — it read "a cluster by that name exists" +// and refused, on a document whose own status said it had put it there. +// +// What that cost was the deployment: the config went Failed at AwaitingCluster +// with ClusterExists, naming the cluster it had created itself, and the nodes it +// had not created yet were never created. +func TestADocumentDoesNotRefuseTheClusterItCreated(t *testing.T) { + config := aDocument(nil) + config.Status.ClusterRef = theCluster + objects := append(workers("worker-1", "worker-2"), config, aCluster(nil)) + r := reconcilerFor(t, objects...) + + done, err := r.createCluster(context.Background(), config) + if err != nil { + t.Fatalf("the document refused the cluster it created: %v", err) + } + if !done { + t.Error("the step did not advance past a cluster that is already there") + } +} + +// A cluster somebody else put there is still refused, which is what the check +// exists for: the document asked to create one and nothing proves the one that +// is there is the one it described. +func TestAClusterThisDocumentDidNotCreateIsStillRefused(t *testing.T) { + config := aDocument(nil) + objects := append(workers("worker-1", "worker-2"), config, aCluster(nil)) + r := reconcilerFor(t, objects...) + + _, err := r.createCluster(context.Background(), config) + if err == nil { + t.Fatal("a cluster this document did not create was adopted silently") + } + var refusal *refusedError + if !errors.As(err, &refusal) || refusal.reason != ClusterExists { + t.Errorf("the failure is %v, want a ClusterExists refusal", err) + } +} From 881f3f9e2604a03be9c2f9cd3da758e8aac9923a Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Tue, 15 Sep 2026 19:12:34 +0200 Subject: [PATCH 013/206] Added minimal prose --- CLAUDE.md | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/CLAUDE.md b/CLAUDE.md index eef4bd20c..0597e15b9 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -1 +1,3 @@ -@AGENTS.md \ No newline at end of file +@AGENTS.md + +General instruction: Please remove all mannered prose. From 9c4b0f9f51b2d1ecddd68352f3fee9505266db78 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Tue, 15 Sep 2026 21:48:15 +0200 Subject: [PATCH 014/206] fix(operator): the daemon set write retries a conflict it cannot avoid The read that seeds the update is served from the informer cache, and the DaemonSet controller rewrites status on every pod transition. So while pods are rolling -- which is exactly when this reconcile runs -- the cached resourceVersion is stale and the update loses: reconcile the daemon set: Operation cannot be fulfilled on daemonsets.apps "simplyblock-storage-node-ds-...": the object has been modified; please apply your changes to the latest version and try again The conflict carries no meaning worth failing over: somebody counted a ready pod. It failed the whole workload pass regardless, so the service, the endpoint slice, the certificates, and the worker enrollment were all re-run on a back-off for it. The write is retried on a fresh read, which is the shape the rest of the tree already uses. The desired object is rebuilt per attempt so a retry does not carry the version the previous one was refused for. The sibling writes are the same shape and are left alone, because they are not the same exposure: nothing but this operator writes the service, the endpoint slice, or the roles, and two replicas racing over them is what leader election already prevents. The DaemonSet is the one with another controller writing it. Co-Authored-By: Claude Opus 5 (1M context) --- .../controllers/node/enrollment_test.go | 65 +++++++++++++++++++ .../controllers/node/workload_controller.go | 34 +++++++--- 2 files changed, 89 insertions(+), 10 deletions(-) diff --git a/operator/internal/controllers/node/enrollment_test.go b/operator/internal/controllers/node/enrollment_test.go index 54e50b44e..a0d2461ff 100644 --- a/operator/internal/controllers/node/enrollment_test.go +++ b/operator/internal/controllers/node/enrollment_test.go @@ -20,11 +20,14 @@ import ( corev1 "k8s.io/api/core/v1" discoveryv1 "k8s.io/api/discovery/v1" rbacv1 "k8s.io/api/rbac/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/runtime/schema" "k8s.io/client-go/tools/events" ctrl "sigs.k8s.io/controller-runtime" "sigs.k8s.io/controller-runtime/pkg/client" "sigs.k8s.io/controller-runtime/pkg/client/fake" + "sigs.k8s.io/controller-runtime/pkg/client/interceptor" atlaskube "github.com/simplyblock/atlas/kube" "github.com/simplyblock/atlas/ptr" @@ -41,12 +44,22 @@ const ( // aWorkloadReconciler drives the cluster's workload over a fake client. func aWorkloadReconciler(t *testing.T, objects ...client.Object) *StorageNodeWorkloadReconciler { + t.Helper() + return aWorkloadReconcilerWith(t, interceptor.Funcs{}, objects...) +} + +// aWorkloadReconcilerWith scripts the client's answers, which is how a test +// reaches a conflict that a fake client never produces on its own. +func aWorkloadReconcilerWith( + t *testing.T, funcs interceptor.Funcs, objects ...client.Object, +) *StorageNodeWorkloadReconciler { t.Helper() scheme := testsupport.NewScheme(t, corev1.AddToScheme, appsv1.AddToScheme, discoveryv1.AddToScheme, rbacv1.AddToScheme) apiClient := fake.NewClientBuilder().WithScheme(scheme). WithObjects(objects...). WithStatusSubresource(&simplyblockv1alpha2.StorageCluster{}). + WithInterceptorFuncs(funcs). Build() return &StorageNodeWorkloadReconciler{ @@ -141,3 +154,55 @@ func TestTheWorkloadLeavesUnrelatedWorkersAlone(t *testing.T) { t.Errorf("a worker with no node of this cluster was enrolled: %v", bystander.Labels) } } + +// The DaemonSet write survives losing a race with its own controller. +// +// The read that seeds the update is served from the informer cache, and the +// DaemonSet controller rewrites status on every pod transition, so the cached +// resourceVersion is stale for as long as pods are rolling — which is exactly +// when this reconcile runs. Without a retry the whole workload pass fails on a +// conflict that means nothing more than that somebody counted a ready pod. +func TestTheDaemonSetWriteRetriesAConflict(t *testing.T) { + remaining := 1 + cluster := aSizedCluster() + objects := []client.Object{ + &corev1.Node{ObjectMeta: metav1.ObjectMeta{Name: "worker-1"}}, + cluster, aNodeOn("worker-1"), + } + r := aWorkloadReconcilerWith(t, interceptor.Funcs{ + Update: func( + ctx context.Context, c client.WithWatch, obj client.Object, + opts ...client.UpdateOption, + ) error { + if _, isDaemonSet := obj.(*appsv1.DaemonSet); isDaemonSet && remaining > 0 { + remaining-- + return apierrors.NewConflict( + schema.GroupResource{Group: "apps", Resource: "daemonsets"}, + obj.GetName(), errStaleDaemonSet{}) + } + return c.Update(ctx, obj, opts...) + }, + }, objects...) + + // The first pass creates the DaemonSet; the conflict is raised on the second, + // which is the update path. + if _, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: client.ObjectKey{Namespace: enrollNamespace, Name: enrollCluster}, + }); err != nil { + t.Fatalf("the first pass: %v", err) + } + if _, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: client.ObjectKey{Namespace: enrollNamespace, Name: enrollCluster}, + }); err != nil { + t.Fatalf("the workload failed on a conflict it should have retried: %v", err) + } + if remaining != 0 { + t.Error("the conflict was never raised, so this proves nothing") + } +} + +type errStaleDaemonSet struct{} + +func (errStaleDaemonSet) Error() string { + return "the object has been modified; please apply your changes to the latest version and try again" +} diff --git a/operator/internal/controllers/node/workload_controller.go b/operator/internal/controllers/node/workload_controller.go index 62efbb300..f110bfb8e 100644 --- a/operator/internal/controllers/node/workload_controller.go +++ b/operator/internal/controllers/node/workload_controller.go @@ -33,6 +33,7 @@ import ( "k8s.io/apimachinery/pkg/runtime" "k8s.io/apimachinery/pkg/types" "k8s.io/client-go/tools/events" + "k8s.io/client-go/util/retry" ctrl "sigs.k8s.io/controller-runtime" "sigs.k8s.io/controller-runtime/pkg/client" "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil" @@ -218,16 +219,29 @@ func (r *StorageNodeWorkloadReconciler) reconcileDaemonSet( return err } - var existing appsv1.DaemonSet - err = r.Get(ctx, client.ObjectKeyFromObject(desired), &existing) - if apierrors.IsNotFound(err) { - return r.Create(ctx, desired) - } - if err != nil { - return err - } - desired.ResourceVersion = existing.ResourceVersion - return r.Update(ctx, desired) + // The read that seeds the update is served from the informer cache, and the + // DaemonSet controller rewrites status on every pod transition — so while + // pods are rolling, which is exactly when this reconcile runs, the cached + // resourceVersion is stale and the update loses. Retrying on a fresh read is + // the answer rather than failing the whole workload pass over a conflict that + // means no more than that somebody counted a ready pod. + return retry.RetryOnConflict(retry.DefaultRetry, func() error { + var existing appsv1.DaemonSet + err := r.Get(ctx, client.ObjectKeyFromObject(desired), &existing) + if apierrors.IsNotFound(err) { + return r.Create(ctx, desired) + } + if err != nil { + return err + } + + // The desired object is rebuilt from the template on every attempt, so a + // retry does not carry the resourceVersion the previous one was refused + // for. + update := desired.DeepCopy() + update.ResourceVersion = existing.ResourceVersion + return r.Update(ctx, update) + }) } // image is the storage-node container image, defaulting to the ControlPlane From e10f2cece9600e982a3681cef92b9005a46df764 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Tue, 15 Sep 2026 21:56:05 +0200 Subject: [PATCH 015/206] fix(operator): restore the partitions-per-device count to 0 and 1 The v1alpha2 port rewrote partitionsPerDevice against the new spec type instead of carrying it across, and came out one too high in both branches: 1 where a dedicated journal device means 0, and 2 where a journal partition per device means 1. The comment recording what the numbers were for was dropped in the same change, which is how they went unnoticed -- with the contract gone, 1 and 2 read as plausibly as 0 and 1, and nothing tested it. The numbers are not a preference. The backend compares what a device already carries against 1 + this count, repartitions when they differ, and cannot repartition a device whose table SPDK's gpt module has claimed: bdev_open: bdev nvme_5n1 already claimed: type exclusive_write by module gpt spdk_nbd_start: could not open bdev nvme_5n1, error=-1 So asking for one partition too many does not lay a disk out differently. On a machine that has never run this product it silently builds the wrong layout, and on one that has -- where the disks already carry the count the fleet was built with -- it makes the node impossible to add at all. The contract is restated on the function this time, and the mapping is tested. Co-Authored-By: Claude Opus 5 (1M context) --- .../controllers/node/partitions_test.go | 55 +++++++++++++++++++ .../node/storagenode_controller.go | 15 ++++- 2 files changed, 68 insertions(+), 2 deletions(-) create mode 100644 operator/internal/controllers/node/partitions_test.go diff --git a/operator/internal/controllers/node/partitions_test.go b/operator/internal/controllers/node/partitions_test.go new file mode 100644 index 000000000..c7828db83 --- /dev/null +++ b/operator/internal/controllers/node/partitions_test.go @@ -0,0 +1,55 @@ +// The partitions-per-device count the control plane is asked for. +// +// It is not a free parameter. The backend compares the partitions a device +// already carries against 1 + this number, repartitions when they differ, and +// cannot repartition a device whose table SPDK's gpt module has claimed. So a +// count that disagrees with the one a fleet was built on does not produce a +// different layout — it produces a node that cannot be added at all, on exactly +// the machines that have run this product before. + +package node + +import ( + "testing" + + "github.com/simplyblock/atlas/ptr" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// The count is what the retired spec.partitions field meant, which is what every +// fleet deployed before that field was replaced carries on its disks. +func TestThePartitionCountMatchesWhatFleetsWereBuiltWith(t *testing.T) { + for _, tc := range []struct { + name string + journal *bool + want int + }{ + { + // A journal carved out of every storage device, which is the default + // and what spec.partitions defaulted to. + name: "a journal partition per device", + journal: nil, + want: 1, + }, + { + name: "explicitly per device", + journal: ptr.To(false), + want: 1, + }, + { + // A whole device given to the journal manager, so nothing is carved + // out of the storage devices. + name: "a dedicated journal device", + journal: ptr.To(true), + want: 0, + }, + } { + t.Run(tc.name, func(t *testing.T) { + workload := &simplyblockv1alpha2.StorageNodesSpec{EnableJournalDevice: tc.journal} + if got := partitionsPerDevice(workload); got != tc.want { + t.Errorf("partitionsPerDevice = %d, want %d", got, tc.want) + } + }) + } +} diff --git a/operator/internal/controllers/node/storagenode_controller.go b/operator/internal/controllers/node/storagenode_controller.go index ef55e8999..9bd449c67 100644 --- a/operator/internal/controllers/node/storagenode_controller.go +++ b/operator/internal/controllers/node/storagenode_controller.go @@ -1161,11 +1161,22 @@ func domainIndex(domain string) *int { // partitionsPerDevice is how many partitions each device is carved into, which is // one when the journal has a device of its own and two when it shares. +// partitionsPerDevice translates spec.enableJournalDevice into the count the +// control plane wants: 0 gives a whole device to the journal manager, and 1 +// carves a journal partition out of each storage device. Unset is 1, which is +// what the retired spec.partitions field defaulted to. +// +// The numbers are not arbitrary and are not a preference. The backend compares +// what a device already carries against 1 + this count and repartitions when +// they differ, and it cannot repartition a device whose table SPDK's gpt module +// has claimed — which is every device of a machine that has run this product +// before. A count one higher than the fleet was built with therefore does not +// lay the disks out differently; it makes the node impossible to add. func partitionsPerDevice(workload *simplyblockv1alpha2.StorageNodesSpec) int { if ptr.BoolFromOrFalse(workload.EnableJournalDevice) { - return 1 + return 0 } - return 2 + return 1 } func journalPercent(spec *simplyblockv1alpha2.JournalManagerSpec) int { From d4974576b090b87c9b792da6c774f030e5aaa5f8 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Tue, 15 Sep 2026 22:23:11 +0200 Subject: [PATCH 016/206] fix(operator): a removal does not wait for the cluster to be active Every node operation held while its cluster was anything but active. That is right for the operations the gate was written for: one that moves data while the cluster is rebalancing or unready is either rejected by the control plane or succeeds into a layout nobody described. Removal is the exception, because it is the repair rather than the thing being protected. A node is removed from an unready cluster precisely to make the cluster ready, so holding the removal closes a loop with no way out: Warning ClusterNotReady cluster ... is unready rather than active The node could not be removed until the cluster was active, and the cluster could not become active while the node it was stuck on was still in it. A node whose add never finished puts a deployment there and leaves it there, with the only operation that would fix it refusing to run. Nothing else is exempt. The gate keeps an operation that moves data off a cluster that cannot take it, and an exemption wider than the one case that needs it is a gate that stops meaning anything. Co-Authored-By: Claude Opus 5 (1M context) --- .../controllers/node/remove_deadlock_test.go | 47 +++++++++++++++++++ .../node/storagenodeops_controller.go | 29 ++++++++++-- 2 files changed, 71 insertions(+), 5 deletions(-) create mode 100644 operator/internal/controllers/node/remove_deadlock_test.go diff --git a/operator/internal/controllers/node/remove_deadlock_test.go b/operator/internal/controllers/node/remove_deadlock_test.go new file mode 100644 index 000000000..18bd7b67b --- /dev/null +++ b/operator/internal/controllers/node/remove_deadlock_test.go @@ -0,0 +1,47 @@ +// Whether a node operation waits for its cluster to be active. +// +// Most of them must. An operation that moves data while the cluster is +// rebalancing or unready is either rejected by the control plane or succeeds +// into a layout nobody described, so holding until the cluster settles is the +// correct behavior and the event says so. +// +// Removal is the exception, because it is the repair rather than the thing being +// protected: a node is removed from an unready cluster precisely to make the +// cluster ready. Holding it closes a loop with no way out — the node cannot be +// removed until the cluster is active, and the cluster cannot become active +// while the node it is stuck on is still there. + +package node + +import ( + "testing" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +func opsFor(action simplyblockv1alpha2.StorageNodeOpsAction) *simplyblockv1alpha2.StorageNodeOps { + ops := &simplyblockv1alpha2.StorageNodeOps{} + ops.Spec.Action = action + return ops +} + +// A removal runs whatever the cluster says about itself. +func TestRemovalDoesNotWaitOnTheCluster(t *testing.T) { + if !skipsClusterGate(opsFor(simplyblockv1alpha2.StorageNodeOpsActionRemove)) { + t.Error("a removal waits for a cluster that removal is how you repair") + } +} + +// Everything else still waits, because the gate is what keeps an operation that +// moves data off a cluster that cannot take it. +func TestEveryOtherActionWaitsOnTheCluster(t *testing.T) { + for _, action := range []simplyblockv1alpha2.StorageNodeOpsAction{ + simplyblockv1alpha2.StorageNodeOpsActionMigrate, + simplyblockv1alpha2.StorageNodeOpsActionRestart, + simplyblockv1alpha2.StorageNodeOpsActionHostMaintenance, + } { + if skipsClusterGate(opsFor(action)) { + t.Errorf("%s skipped the cluster gate", action) + } + } +} diff --git a/operator/internal/controllers/node/storagenodeops_controller.go b/operator/internal/controllers/node/storagenodeops_controller.go index 4e150c1fc..d8a7eebc0 100644 --- a/operator/internal/controllers/node/storagenodeops_controller.go +++ b/operator/internal/controllers/node/storagenodeops_controller.go @@ -272,11 +272,13 @@ func (r *StorageNodeOpsReconciler) advance( // a cluster, and one whose cluster is mid-rebalance or not active will either // be rejected by the control plane or succeed into an inconsistent layout. It // holds rather than fails, and resumes when the cluster does (§7.1). - if ready, reason, err := r.clusterReady(ctx, ops); err != nil { - return ctrl.Result{RequeueAfter: opsRetry}, r.note(ctx, ops, err.Error()) - } else if !ready { - r.emit(ctx, ops, corev1.EventTypeWarning, ClusterNotReady, reason) - return ctrl.Result{RequeueAfter: opsRetry}, r.note(ctx, ops, reason) + if !skipsClusterGate(ops) { + if ready, reason, err := r.clusterReady(ctx, ops); err != nil { + return ctrl.Result{RequeueAfter: opsRetry}, r.note(ctx, ops, err.Error()) + } else if !ready { + r.emit(ctx, ops, corev1.EventTypeWarning, ClusterNotReady, reason) + return ctrl.Result{RequeueAfter: opsRetry}, r.note(ctx, ops, reason) + } } if machine.TimeoutReached() { @@ -720,6 +722,23 @@ func (r *StorageNodeOpsReconciler) target( // It reports a reason rather than an error, because holding is the response and a // reason is what an event carries. A cluster whose reading cannot be taken at all // is an error, which the caller retries. +// skipsClusterGate reports the operations that run whatever the cluster says +// about itself. +// +// Removal is the one. A node is removed from an unready cluster precisely to +// make the cluster ready, so holding the removal until the cluster is active +// closes a loop with no way out: the node cannot be removed until the cluster is +// active, and the cluster cannot become active while the node it is stuck on is +// still in it. That is not hypothetical — it is what a node whose add never +// finished does to the cluster it was being added to. +// +// Nothing else is exempt. The gate exists to keep an operation that moves data +// off a cluster that cannot take it, and an exemption wider than the one case +// that needs it is a gate that stops meaning anything. +func skipsClusterGate(ops *simplyblockv1alpha2.StorageNodeOps) bool { + return ops.Spec.Action == simplyblockv1alpha2.StorageNodeOpsActionRemove +} + func (r *StorageNodeOpsReconciler) clusterReady( ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, ) (bool, string, error) { From 2300bf626a1aa4a9f7e2df2f24a9202b2f53bb1b Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Wed, 16 Sep 2026 07:17:33 +0200 Subject: [PATCH 017/206] fix(operator): close a node's device stream without needing its cluster A device scope is opened when a node resolves its backend id and closed when the node is torn down, and the two run at different moments with different things readable. The open has the node freshly placed in its cluster, so the cluster's id is there to build the key from. The teardown may run with the cluster already gone -- a cluster deleted with its nodes is the ordinary case -- and the close needed that same id to rebuild the key it had added under. Without it the close was skipped and the scope stayed in the set, so the manager kept a stream open for a node the control plane no longer has: cpinformer stream disconnected, reconnecting {"subscription": "device", "scope": ".../b39b01f9-...", "err": "watch device ...: unexpected status 404 Not Found"} It reconnects on a 3s-to-30s backoff for the life of the process, one leaked stream per removed node. The close is now by the node's own id, which is the thing the scope is about and the one identifier a teardown always has. RemoveLeaf drops every scope ending in it and nothing else; an empty id matches nothing rather than everything, because a caller with no id knows nothing rather than asking for a purge. The cluster argument that only existed to rebuild the key goes with it, and the teardown no longer takes a cluster it does not read. Co-Authored-By: Claude Opus 5 (1M context) --- .../node/storagenode_controller.go | 27 +++---- .../controllers/node/unregister_test.go | 75 +++++++++++++++++++ operator/internal/cpinformer/manager.go | 49 ++++++++++++ .../internal/cpinformer/scopeset_leaf_test.go | 61 +++++++++++++++ 4 files changed, 199 insertions(+), 13 deletions(-) create mode 100644 operator/internal/controllers/node/unregister_test.go create mode 100644 operator/internal/cpinformer/scopeset_leaf_test.go diff --git a/operator/internal/controllers/node/storagenode_controller.go b/operator/internal/controllers/node/storagenode_controller.go index 9bd449c67..be375b152 100644 --- a/operator/internal/controllers/node/storagenode_controller.go +++ b/operator/internal/controllers/node/storagenode_controller.go @@ -213,7 +213,7 @@ func (r *StorageNodeReconciler) Reconcile( } if !node.DeletionTimestamp.IsZero() { - return r.teardown(ctx, &node, cluster) + return r.teardown(ctx, &node) } if !controllerutil.ContainsFinalizer(&node, NodeFinalizer) { @@ -887,18 +887,13 @@ func (r *StorageNodeReconciler) raiseMaintenance( func (r *StorageNodeReconciler) teardown( ctx context.Context, node *simplyblockv1alpha2.StorageNode, - cluster *simplyblockv1alpha2.StorageCluster, ) (ctrl.Result, error) { - clusterID := "" - if cluster != nil { - clusterID = cluster.Status.UUID - } if !controllerutil.ContainsFinalizer(node, NodeFinalizer) { return ctrl.Result{}, nil } if node.Status.UUID == "" { - r.unregister(node, clusterID) + r.unregister(node) controllerutil.RemoveFinalizer(node, NodeFinalizer) return ctrl.Result{}, r.Update(ctx, node) } @@ -916,7 +911,7 @@ func (r *StorageNodeReconciler) teardown( return ctrl.Result{RequeueAfter: nodeRetry}, nil } - r.unregister(node, clusterID) + r.unregister(node) controllerutil.RemoveFinalizer(node, NodeFinalizer) return ctrl.Result{}, r.Update(ctx, node) } @@ -983,14 +978,20 @@ func (r *StorageNodeReconciler) register( // The scope goes first: no further device event can arrive once the stream is // closed, so the name mappings are dropped second and nothing is left naming // objects after a node on its way out. -func (r *StorageNodeReconciler) unregister( - node *simplyblockv1alpha2.StorageNode, clusterID string, -) { +// +// It closes by the node's own id rather than by the cluster-and-node key it was +// opened under, because the two happen at different moments with different +// things readable. The open has the node freshly placed in its cluster, so the +// cluster's id is there; the teardown may run with the cluster already gone, +// which is the ordinary case when a cluster is deleted with its nodes. Requiring +// the cluster's id here left the scope in the set, and the stream behind it +// reconnecting against a 404 for the life of the process. +func (r *StorageNodeReconciler) unregister(node *simplyblockv1alpha2.StorageNode) { if node.Status.UUID == "" { return } - if r.DeviceScopes != nil && clusterID != "" { - r.DeviceScopes.Remove(cpinformer.Scope{clusterID, node.Status.UUID}) + if r.DeviceScopes != nil { + r.DeviceScopes.RemoveLeaf(node.Status.UUID) } for _, registry := range r.Registries { registry.UnregisterNode(node.Status.UUID) diff --git a/operator/internal/controllers/node/unregister_test.go b/operator/internal/controllers/node/unregister_test.go new file mode 100644 index 000000000..c4d8f073e --- /dev/null +++ b/operator/internal/controllers/node/unregister_test.go @@ -0,0 +1,75 @@ +// Closing a node's device stream when the node goes. +// +// The scope is opened when the node resolves its backend id and closed when the +// node is torn down, and the two see different things. The open happens with the +// node freshly placed in its cluster, so the cluster's id is there to be had. The +// teardown happens on a path where the cluster may already be gone — a cluster +// deleted with its nodes is the ordinary case — and the close was written to need +// that id again. +// +// So it did not close. The control plane answers 404 for a node it no longer has, +// the manager reads that as a disconnect, and it reconnects on a backoff forever: +// +// cpinformer stream disconnected, reconnecting +// {"subscription": "device", "scope": ".../b39b01f9-...", "err": "... 404"} + +package node + +import ( + "testing" + + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/cpinformer" +) + +func aResolvedNode(uuid string) *simplyblockv1alpha2.StorageNode { + node := &simplyblockv1alpha2.StorageNode{ + ObjectMeta: metav1.ObjectMeta{Name: "a-node", Namespace: "simplyblock"}, + } + node.Status.UUID = uuid + return node +} + +// The stream closes even when the cluster it was opened against is gone. +func TestTheDeviceStreamClosesWithoutTheCluster(t *testing.T) { + scopes := cpinformer.NewScopeSetForTest() + scopes.Add(cpinformer.Scope{"cluster-uuid", "node-uuid"}) + r := &StorageNodeReconciler{DeviceScopes: scopes} + + // The cluster is unreadable at teardown, so there is no id to rebuild the + // key with. That is the case the leak was found in. + r.unregister(aResolvedNode("node-uuid")) + + if got := scopes.Len(); got != 0 { + t.Errorf("%d scope(s) survived the node they stream", got) + } +} + +// It still closes on the ordinary path, where the cluster is right there. +func TestTheDeviceStreamClosesWithTheCluster(t *testing.T) { + scopes := cpinformer.NewScopeSetForTest() + scopes.Add(cpinformer.Scope{"cluster-uuid", "node-uuid"}) + r := &StorageNodeReconciler{DeviceScopes: scopes} + + r.unregister(aResolvedNode("node-uuid")) + + if got := scopes.Len(); got != 0 { + t.Errorf("%d scope(s) survived the node they stream", got) + } +} + +// A node that never resolved an id opened no stream, so nothing is closed and +// no other node's stream is touched. +func TestANodeWithNoIdClosesNothing(t *testing.T) { + scopes := cpinformer.NewScopeSetForTest() + scopes.Add(cpinformer.Scope{"cluster-uuid", "another-node"}) + r := &StorageNodeReconciler{DeviceScopes: scopes} + + r.unregister(aResolvedNode("")) + + if got := scopes.Len(); got != 1 { + t.Errorf("the set holds %d scope(s), want the other node's left alone", got) + } +} diff --git a/operator/internal/cpinformer/manager.go b/operator/internal/cpinformer/manager.go index e4a26e388..c4e44f0f3 100644 --- a/operator/internal/cpinformer/manager.go +++ b/operator/internal/cpinformer/manager.go @@ -81,6 +81,55 @@ func (s *ScopeSet) Remove(scope Scope) { } } +// NewScopeSetForTest builds a set outside a manager, so a caller that only adds +// and drops scopes can be exercised without starting one. The constructor is +// unexported because a set belongs to the subscription that streams it; this is +// the seam for the controllers that are only ever handed one. +func NewScopeSetForTest() *ScopeSet { return newScopeSet() } + +// Len is how many scopes the set holds, which is what a test asserting that a +// stream was closed can see. +func (s *ScopeSet) Len() int { + s.mu.Lock() + defer s.mu.Unlock() + return len(s.scopes) +} + +// RemoveLeaf drops every scope whose last element is the id given. +// +// It exists because the add and the drop happen at different moments with +// different things readable. A node's scope is added when it resolves its id, +// where the cluster it sits in is necessarily known; it is dropped on a teardown +// path where the cluster may already be gone, and a caller that cannot name the +// cluster cannot rebuild the key it added under. Requiring the whole key there +// means the scope is kept instead, and the stream behind it is never closed: +// the control plane answers 404 for a node it no longer has, the manager reads +// that as a disconnect, and it reconnects on a backoff for the life of the +// process. +// +// The leaf is the thing the scope is about, so matching on it drops the scopes +// for that thing and no others. An empty id matches nothing rather than +// everything, because a caller with no id to offer is a caller that knows +// nothing, not one asking for a purge. +func (s *ScopeSet) RemoveLeaf(id string) { + if id == "" { + return + } + + s.mu.Lock() + defer s.mu.Unlock() + removed := false + for key, scope := range s.scopes { + if len(scope) > 0 && scope[len(scope)-1] == id { + delete(s.scopes, key) + removed = true + } + } + if removed { + s.signal() + } +} + func (s *ScopeSet) signal() { // caller holds s.mu select { case s.notify <- struct{}{}: diff --git a/operator/internal/cpinformer/scopeset_leaf_test.go b/operator/internal/cpinformer/scopeset_leaf_test.go new file mode 100644 index 000000000..7fa815eb6 --- /dev/null +++ b/operator/internal/cpinformer/scopeset_leaf_test.go @@ -0,0 +1,61 @@ +// Dropping a scope by the thing it is about, rather than by the whole key. +// +// A scope is added when a node resolves its id and dropped when the node goes, +// and the two happen at different moments with different things readable. The +// add knows the cluster's id because the node has just been placed in it; the +// drop is on a teardown path where the cluster may already be gone, and a caller +// that cannot name the cluster cannot build the key it added under. +// +// What that costs is a stream nobody closes. The control plane answers 404 for a +// node it no longer has, the manager treats that as a disconnect, and it +// reconnects on a backoff for the lifetime of the process — one leaked stream per +// removed node. + +package cpinformer + +import "testing" + +func TestRemoveLeafDropsAScopeWithoutItsPrefix(t *testing.T) { + set := newScopeSet() + set.Add(Scope{"cluster-a", "node-1"}) + set.Add(Scope{"cluster-a", "node-2"}) + + set.RemoveLeaf("node-1") + + got := set.list() + if len(got) != 1 { + t.Fatalf("the set holds %v, want only node-2", got) + } + if got[0].Key() != (Scope{"cluster-a", "node-2"}).Key() { + t.Errorf("the set holds %v, want node-2", got[0]) + } +} + +// Two clusters can hold a node id only by coincidence, and dropping both would +// close a stream that is still wanted. The leaf is the node, so every scope +// ending in it is about that node. +func TestRemoveLeafDropsEveryScopeForThatNode(t *testing.T) { + set := newScopeSet() + set.Add(Scope{"cluster-a", "node-1"}) + set.Add(Scope{"cluster-b", "node-1"}) + + set.RemoveLeaf("node-1") + + if got := set.list(); len(got) != 0 { + t.Errorf("the set still holds %v", got) + } +} + +// A leaf nothing matches changes nothing, so a teardown that runs twice is not a +// teardown that closes somebody else's stream. +func TestRemoveLeafIsIdempotentAndNarrow(t *testing.T) { + set := newScopeSet() + set.Add(Scope{"cluster-a", "node-1"}) + + set.RemoveLeaf("node-2") + set.RemoveLeaf("") + + if got := set.list(); len(got) != 1 { + t.Errorf("the set holds %v, want the untouched scope", got) + } +} From 3c6979b28122ea65c56586f4eda0dca957b17a1b Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Wed, 16 Sep 2026 07:29:25 +0200 Subject: [PATCH 018/206] fix(operator): a removal finishes when its node is already gone Removing a node is a sequence of calls against a backend node: suspend it, move its volumes, verify, delete it. The last step already reads a 404 as success, because a node the control plane no longer knows about is the outcome the delete asked for. The earlier steps did not, so a removal that found its node missing at Suspending reported the 404 as a step that could not be advanced, and retried it for as long as the operator ran: the step could not be advanced ... step: Suspending error: suspend node ...: the control plane answered 404: 'StorageNode 7838e194-... not found' A removal refusing to finish because the thing it is removing is gone. It also holds the node's finalizer, so the object never goes either, and the deployment is stuck on a node that does not exist. The node can be missing before the step that would have removed it in more than one way: an earlier attempt got that far and lost its response, the add being undone never registered it, or somebody else removed it. None of them is a failure of this operation, so the whole sequence is short-circuited rather than each step patched. The check asks the control plane rather than the stream's cache. A cache that has not synced reports every node missing, and reading that as "already removed" would finish a drain that never moved a volume. Co-Authored-By: Claude Opus 5 (1M context) --- operator/internal/controllers/node/remove.go | 33 ++++++ .../controllers/node/remove_gone_test.go | 103 ++++++++++++++++++ 2 files changed, 136 insertions(+) create mode 100644 operator/internal/controllers/node/remove_gone_test.go diff --git a/operator/internal/controllers/node/remove.go b/operator/internal/controllers/node/remove.go index e1ec515b4..66a65ecfe 100644 --- a/operator/internal/controllers/node/remove.go +++ b/operator/internal/controllers/node/remove.go @@ -61,6 +61,22 @@ func (r *StorageNodeOpsReconciler) performRemoveStep( return false, err } + // A node the control plane does not have is what this operation was for, so + // every step of it is already done. The last step reads a 404 as success for + // the same reason; the earlier ones did not, and a removal that found its + // node missing at Suspending reported the 404 as a step that could not be + // advanced and retried it for as long as the operator ran. + // + // The node can be gone before the step that would have removed it in more + // than one way: an earlier attempt got that far and lost its response, the + // add that was being undone never registered it, or somebody else removed it. + // None of them is a failure of this operation. + if gone, err := r.nodeGone(ctx, clusterID, nodeID); err != nil { + return false, err + } else if gone { + return true, nil + } + switch current { case stepValidating: return r.drainValidate(ctx, ops, clusterID, nodeID) @@ -77,6 +93,23 @@ func (r *StorageNodeOpsReconciler) performRemoveStep( } } +// nodeGone reports whether the control plane has forgotten the node. +// +// It asks the control plane rather than the stream's cache, because the cache +// not having a node and the control plane not having one are different facts and +// only the second one ends a removal. A cache that has not synced reports every +// node missing, and treating that as "already removed" would finish a drain that +// never moved a volume. +func (r *StorageNodeOpsReconciler) nodeGone( + ctx context.Context, clusterID, nodeID string, +) (bool, error) { + _, found, err := r.API.StorageNode(ctx, clusterID, nodeID) + if err != nil { + return false, fmt.Errorf("read node %s: %w", nodeID, err) + } + return !found, nil +} + // drainValidate classifies the node's volumes and refuses to go on while any of // them is pinned or unmanaged. It performs no side effect at all, which is what // makes an abort here an Aborted directly rather than an unwind. diff --git a/operator/internal/controllers/node/remove_gone_test.go b/operator/internal/controllers/node/remove_gone_test.go new file mode 100644 index 000000000..00687aa23 --- /dev/null +++ b/operator/internal/controllers/node/remove_gone_test.go @@ -0,0 +1,103 @@ +// A removal whose target the control plane no longer has. +// +// Removing a node is a sequence of calls against a backend node: suspend it, +// move its volumes, verify, delete it. Every one of them is written to tolerate a +// repeat, because a step recorded without its call having fired re-issues it. But +// a repeat is not the only thing that happens twice — the node can also be gone +// before the sequence reaches the step that would have removed it, either because +// an earlier attempt got that far or because somebody else removed it. +// +// The last step already reads a 404 as success. The earlier ones did not, so a +// removal that found its node missing at Suspending reported the 404 as a step +// that could not be advanced and retried it for as long as the operator ran: +// +// the step could not be advanced ... step: Suspending +// error: suspend node ...: the control plane answered 404: +// 'StorageNode 7838e194-... not found' +// +// Which is a removal refusing to finish because what it was removing is gone. + +package node + +import ( + "context" + "testing" + + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/client-go/tools/events" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/testsupport" +) + +// goneControlPlane reports that the node is not there. Everything else is nil: +// a removal that reaches past this is reaching for a node it has been told does +// not exist, and should panic rather than pass. +type goneControlPlane struct { + ControlPlane +} + +func (goneControlPlane) StorageNode( + context.Context, string, string, +) (NodeReading, bool, error) { + return NodeReading{}, false, nil +} + +func aRemoveOps() *simplyblockv1alpha2.StorageNodeOps { + ops := &simplyblockv1alpha2.StorageNodeOps{ + ObjectMeta: metav1.ObjectMeta{Name: "a-node-remove", Namespace: "simplyblock"}, + Spec: simplyblockv1alpha2.StorageNodeOpsSpec{ + Action: simplyblockv1alpha2.StorageNodeOpsActionRemove, + NodeRef: "a-node", + }, + } + return ops +} + +func aRemover(t *testing.T) *StorageNodeOpsReconciler { + t.Helper() + scheme := testsupport.NewScheme(t) + + node := &simplyblockv1alpha2.StorageNode{ + ObjectMeta: metav1.ObjectMeta{Name: "a-node", Namespace: "simplyblock"}, + Spec: simplyblockv1alpha2.StorageNodeSpec{ClusterRef: "a-cluster"}, + } + node.Status.UUID = "node-uuid" + + cluster := &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{Name: "a-cluster", Namespace: "simplyblock"}, + } + cluster.Status.UUID = "cluster-uuid" + + ops := aRemoveOps() + apiClient := fake.NewClientBuilder().WithScheme(scheme). + WithObjects(node, cluster, ops). + WithStatusSubresource(&simplyblockv1alpha2.StorageNodeOps{}). + Build() + + return &StorageNodeOpsReconciler{ + Client: apiClient, + Scheme: scheme, + Recorder: events.NewFakeRecorder(64), + API: goneControlPlane{}, + } +} + +// Every step of a removal is done once the node is gone, because gone is what +// the removal was for. +func TestARemovalOfAGoneNodeIsDoneAtEveryStep(t *testing.T) { + r := aRemover(t) + + for _, current := range []step{ + stepValidating, stepSuspending, stepMigratingVolumes, stepVerifying, stepRemoving, + } { + done, err := r.performRemoveStep(context.Background(), aRemoveOps(), current) + if err != nil { + t.Errorf("step %s: %v", current, err) + } + if !done { + t.Errorf("step %s did not finish against a node the control plane does not have", current) + } + } +} From 2895275b3b7945f38f7498123308c76663d0f142 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Wed, 16 Sep 2026 07:33:54 +0200 Subject: [PATCH 019/206] feat(deployment): the draft states that the drives are formatted A storage node takes a drive by formatting it, and a drive carrying anything already is not usable until it is. Nothing said so: a discovered document named the disks and left the question to a field further down that it never set, so every deployment onto machines that had held data before failed inside the control plane rather than at the document. The flag goes on the document because the document is what somebody approves. It is destructive, and the cluster's own field is immutable once the cluster exists, so a default applied further down would format disks without the draft ever saying so and could not be undone afterwards. A reviewer reads it in the draft and strikes the line if any of those drives should be left alone, which the note beside it says in as many words. It is named for what is wanted rather than how it is done, because the how differs by class: an NVMe device is formatted to a 4K block size and a logical block device has its signatures wiped. One field covers both, so a document does not have to know which class the expansion resolves it to. Only the NVMe half is carried today -- the cluster has a field for it -- and the logical-block half has nowhere to go yet. Co-Authored-By: Claude Opus 5 (1M context) --- ...mplyblock.io_clusterdeploymentconfigs.yaml | 17 ++++++++ .../v1alpha2/clusterdeploymentconfig_types.go | 17 ++++++++ .../api/v1alpha2/zz_generated.deepcopy.go | 5 +++ ...mplyblock.io_clusterdeploymentconfigs.yaml | 17 ++++++++ .../controllers/deployment/expansion.go | 4 ++ .../deployment/expansion_review_test.go | 26 +++++++++++++ operator/internal/discovery/template.go | 7 ++++ .../discovery/template_format_test.go | 39 +++++++++++++++++++ ...mplyblock.io_clusterdeploymentconfigs.yaml | 17 ++++++++ 9 files changed, 149 insertions(+) create mode 100644 operator/internal/discovery/template_format_test.go diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml index c7d4745be..100fd3a05 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml @@ -96,6 +96,23 @@ spec: description: Cluster is the StorageCluster to create. Ignored when ClusterRef is set. properties: + enableDriveFormat: + description: |- + EnableDriveFormat formats every device the document names before a storage + node takes it, which is how a drive carrying anything already is made + usable. + + It says what is wanted rather than how, because the how differs by device + class: an NVMe device is formatted to a 4K block size, and a logical block + device has its signatures wiped. One field covers both, so a document does + not have to know which class the expansion will resolve it to. + + It is on the document rather than defaulted further down because it is + destructive and the document is what somebody approves. A reviewer reading + a draft has to see that the drives it lists will be formatted, and be able + to strike it before approving; the cluster's own field is immutable once + the cluster exists, so a default nobody saw could not be undone either. + type: boolean enableFailureDomains: description: |- EnableFailureDomains opts the cluster into failure-domain mode, in which diff --git a/operator/api/v1alpha2/clusterdeploymentconfig_types.go b/operator/api/v1alpha2/clusterdeploymentconfig_types.go index 15a306546..9b5cec92b 100644 --- a/operator/api/v1alpha2/clusterdeploymentconfig_types.go +++ b/operator/api/v1alpha2/clusterdeploymentconfig_types.go @@ -226,6 +226,23 @@ type ClusterTemplate struct { // +optional MinHugePagesSize string `json:"minHugePagesSize,omitempty"` + // EnableDriveFormat formats every device the document names before a storage + // node takes it, which is how a drive carrying anything already is made + // usable. + // + // It says what is wanted rather than how, because the how differs by device + // class: an NVMe device is formatted to a 4K block size, and a logical block + // device has its signatures wiped. One field covers both, so a document does + // not have to know which class the expansion will resolve it to. + // + // It is on the document rather than defaulted further down because it is + // destructive and the document is what somebody approves. A reviewer reading + // a draft has to see that the drives it lists will be formatted, and be able + // to strike it before approving; the cluster's own field is immutable once + // the cluster exists, so a default nobody saw could not be undone either. + // +optional + EnableDriveFormat *bool `json:"enableDriveFormat,omitempty"` + // SocketsToUse restricts the deployment to selected NUMA sockets, and empty // means socket 0 alone. With NodesPerSocket it decides how many storage nodes // each worker runs, so a group of two workers on a two-socket layout expands diff --git a/operator/api/v1alpha2/zz_generated.deepcopy.go b/operator/api/v1alpha2/zz_generated.deepcopy.go index a9f61b713..1a6dcdbdc 100644 --- a/operator/api/v1alpha2/zz_generated.deepcopy.go +++ b/operator/api/v1alpha2/zz_generated.deepcopy.go @@ -284,6 +284,11 @@ func (in *ClusterTemplate) DeepCopyInto(out *ClusterTemplate) { *out = new(int32) **out = **in } + if in.EnableDriveFormat != nil { + in, out := &in.EnableDriveFormat, &out.EnableDriveFormat + *out = new(bool) + **out = **in + } if in.SocketsToUse != nil { in, out := &in.SocketsToUse, &out.SocketsToUse *out = make([]string, len(*in)) diff --git a/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml b/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml index c7d4745be..100fd3a05 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml @@ -96,6 +96,23 @@ spec: description: Cluster is the StorageCluster to create. Ignored when ClusterRef is set. properties: + enableDriveFormat: + description: |- + EnableDriveFormat formats every device the document names before a storage + node takes it, which is how a drive carrying anything already is made + usable. + + It says what is wanted rather than how, because the how differs by device + class: an NVMe device is formatted to a 4K block size, and a logical block + device has its signatures wiped. One field covers both, so a document does + not have to know which class the expansion will resolve it to. + + It is on the document rather than defaulted further down because it is + destructive and the document is what somebody approves. A reviewer reading + a draft has to see that the drives it lists will be formatted, and be able + to strike it before approving; the cluster's own field is immutable once + the cluster exists, so a default nobody saw could not be undone either. + type: boolean enableFailureDomains: description: |- EnableFailureDomains opts the cluster into failure-domain mode, in which diff --git a/operator/internal/controllers/deployment/expansion.go b/operator/internal/controllers/deployment/expansion.go index 7839a8256..9cc0ba1a4 100644 --- a/operator/internal/controllers/deployment/expansion.go +++ b/operator/internal/controllers/deployment/expansion.go @@ -221,6 +221,10 @@ func (r *ClusterDeploymentConfigReconciler) buildWorkload( // document said. workload.SocketsToUse = template.SocketsToUse workload.NodesPerSocket = template.NodesPerSocket + // The document says a drive is to be formatted; this is where that is + // resolved to how. NVMe is the class the cluster's own field covers, and + // the logical-block half has no field to carry it yet. + workload.EnableFormat4K = template.EnableDriveFormat } for _, set := range config.Spec.NodeSets { diff --git a/operator/internal/controllers/deployment/expansion_review_test.go b/operator/internal/controllers/deployment/expansion_review_test.go index 66708d6d4..869a8b252 100644 --- a/operator/internal/controllers/deployment/expansion_review_test.go +++ b/operator/internal/controllers/deployment/expansion_review_test.go @@ -338,3 +338,29 @@ func TestAClusterThisDocumentDidNotCreateIsStillRefused(t *testing.T) { t.Errorf("the failure is %v, want a ClusterExists refusal", err) } } + +// The document's drive-format decision reaches the cluster it creates. +// +// A reviewer approves a document, and what the deployment then does has to be +// what the document said. The flag is destructive and immutable on the cluster, +// so a document that states it and an expansion that drops it would format +// nothing while the draft said it would, or the reverse once somebody strikes it. +func TestTheDriveFormatDecisionReachesTheCluster(t *testing.T) { + for _, stated := range []*bool{ptr.To(true), ptr.To(false), nil} { + config := aDocument(func(c *simplyblockv1alpha2.ClusterDeploymentConfig) { + c.Spec.Cluster.EnableDriveFormat = stated + }) + r := reconcilerFor(t) + + workload := r.buildWorkload(config) + switch { + case stated == nil && workload.EnableFormat4K != nil: + t.Errorf("a document that says nothing produced %v", *workload.EnableFormat4K) + case stated != nil && workload.EnableFormat4K == nil: + t.Errorf("a document that said %v produced nothing", *stated) + case stated != nil && *workload.EnableFormat4K != *stated: + t.Errorf("the cluster got %v, want the document's %v", + *workload.EnableFormat4K, *stated) + } + } +} diff --git a/operator/internal/discovery/template.go b/operator/internal/discovery/template.go index 87876e6d1..dfa68fd02 100644 --- a/operator/internal/discovery/template.go +++ b/operator/internal/discovery/template.go @@ -62,8 +62,15 @@ func ClusterTemplateFor(name string, plan Plan) ClusterTemplate { out := ClusterTemplate{Template: &simplyblockv1alpha2.ClusterTemplate{ Name: name, MaxSubsystemCount: ptr.To(DefaultMaxSubsystemCount), + EnableDriveFormat: ptr.To(true), }} + out.Notes = append(out.Notes, + "enableDriveFormat is set, so every drive listed here is formatted before "+ + "a storage node takes it: a drive that carries anything is not usable "+ + "otherwise. This is the line to remove if any of them should be left "+ + "alone.") + vcpus, note := vcpuCountFor(plan) out.Template.VCPUCount = ptr.To(vcpus) out.Notes = append(out.Notes, note) diff --git a/operator/internal/discovery/template_format_test.go b/operator/internal/discovery/template_format_test.go new file mode 100644 index 000000000..9d055161d --- /dev/null +++ b/operator/internal/discovery/template_format_test.go @@ -0,0 +1,39 @@ +// What the draft says about formatting the devices it names. +// +// A storage node takes a device by formatting it, so the flag is not an unusual +// request. It is still the difference between a document that fails on a disk +// with something on it and one that wipes it, and the document is where that +// difference is decided: stating it in the draft is what puts it in front of the +// reviewer who approves, and what lets them strike it before they do. +// +// Leaving it to a default further down would apply it without the document ever +// saying so, and the field is immutable once the cluster exists — so a default +// nobody saw could not be undone either. + +package discovery + +import ( + "strings" + "testing" +) + +func TestTheDraftStatesThatDevicesAreFormatted(t *testing.T) { + template := ClusterTemplateFor("a-cluster", Plan{}) + + if template.Template.EnableDriveFormat == nil { + t.Fatal("the draft leaves the format flag unstated") + } + if !*template.Template.EnableDriveFormat { + t.Error("the draft states that devices are not formatted") + } + + mentioned := false + for _, note := range template.Notes { + if strings.Contains(strings.ToLower(note), "format") { + mentioned = true + } + } + if !mentioned { + t.Error("nothing in the notes tells a reviewer the disks will be formatted") + } +} diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml index c7d4745be..100fd3a05 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml @@ -96,6 +96,23 @@ spec: description: Cluster is the StorageCluster to create. Ignored when ClusterRef is set. properties: + enableDriveFormat: + description: |- + EnableDriveFormat formats every device the document names before a storage + node takes it, which is how a drive carrying anything already is made + usable. + + It says what is wanted rather than how, because the how differs by device + class: an NVMe device is formatted to a 4K block size, and a logical block + device has its signatures wiped. One field covers both, so a document does + not have to know which class the expansion will resolve it to. + + It is on the document rather than defaulted further down because it is + destructive and the document is what somebody approves. A reviewer reading + a draft has to see that the drives it lists will be formatted, and be able + to strike it before approving; the cluster's own field is immutable once + the cluster exists, so a default nobody saw could not be undone either. + type: boolean enableFailureDomains: description: |- EnableFailureDomains opts the cluster into failure-domain mode, in which From 05dc2fbba8aa9e081a471c7b06e8b954843b728c Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Wed, 16 Sep 2026 08:26:09 +0200 Subject: [PATCH 020/206] fix(operator): the node-add cap holds after the first slot is freed The cap exists because a node add reboots its host, and two hosts rebooting at once costs the control plane its own fault tolerance. Each waiting node enforced it by counting how many siblings had already claimed a slot, which works for exactly one of them: reconciles are serialized per object and not across objects, so every waiting node reads the count before any of them has written its claim. The first add is therefore correctly alone, and the instant it finishes every remaining node sees the same free slot and takes it. A cap that counts other people's writes holds once and then means nothing -- on a three-node deployment the second and third nodes were added together, which is the thing the cap was written to prevent. The decision is now made from the set that is waiting rather than from who has written a claim: the contending workers are ordered, and a node goes on only if its own place in that order is within the number of free slots. Two nodes reading one set reach one answer without either having written anything, and the answer does not change between passes, so a node told to wait is not overtaken by one told to wait beside it. Co-Authored-By: Claude Opus 5 (1M context) --- .../controllers/node/remove_gone_test.go | 2 +- .../controllers/node/slot_race_test.go | 157 ++++++++++++++++++ .../node/storagenode_controller.go | 38 ++++- 3 files changed, 195 insertions(+), 2 deletions(-) create mode 100644 operator/internal/controllers/node/slot_race_test.go diff --git a/operator/internal/controllers/node/remove_gone_test.go b/operator/internal/controllers/node/remove_gone_test.go index 00687aa23..4f154a7a8 100644 --- a/operator/internal/controllers/node/remove_gone_test.go +++ b/operator/internal/controllers/node/remove_gone_test.go @@ -68,7 +68,7 @@ func aRemover(t *testing.T) *StorageNodeOpsReconciler { cluster := &simplyblockv1alpha2.StorageCluster{ ObjectMeta: metav1.ObjectMeta{Name: "a-cluster", Namespace: "simplyblock"}, } - cluster.Status.UUID = "cluster-uuid" + cluster.Status.UUID = aBackendClusterID ops := aRemoveOps() apiClient := fake.NewClientBuilder().WithScheme(scheme). diff --git a/operator/internal/controllers/node/slot_race_test.go b/operator/internal/controllers/node/slot_race_test.go new file mode 100644 index 000000000..a0747224d --- /dev/null +++ b/operator/internal/controllers/node/slot_race_test.go @@ -0,0 +1,157 @@ +// Who takes a node-add slot when several nodes want one. +// +// The cap exists because a node add reboots its host, and two hosts rebooting at +// once costs the control plane its own fault tolerance. It was enforced by each +// node counting how many siblings had already claimed a slot, which works for +// exactly one of them: controller-runtime serializes reconciles per object, not +// across objects, so every waiting node reads the count before any of them has +// written its claim. +// +// The first add is therefore correctly alone, and the moment it finishes every +// remaining node reads a free slot in the same instant and takes it. The cap +// holds once and then means nothing, which is the shape a cap fails in when it +// is a count of other people's writes. +// +// The answer is to decide from what is observed rather than from what has been +// written: the waiting nodes are ordered, and a node takes a slot only if its +// own place in that order is within the number free. Two nodes reading the same +// set reach the same answer without either having written anything. + +package node + +import ( + "context" + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + "github.com/simplyblock/atlas/ptr" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/testsupport" +) + +// aBackendClusterID is the id the control plane knows these fixtures' cluster by. +const aBackendClusterID = "cluster-uuid" + +// waitingNode is a node that has reached AwaitingSlot and claimed nothing. +func waitingNode(worker string) *simplyblockv1alpha2.StorageNode { + node := &simplyblockv1alpha2.StorageNode{ + ObjectMeta: metav1.ObjectMeta{Name: "c-" + worker + "-0", Namespace: "simplyblock"}, + Spec: simplyblockv1alpha2.StorageNodeSpec{ + ClusterRef: "c", WorkerNode: worker, + }, + } + node.Status.Step.State = string(stepAwaitingSlot) + return node +} + +// slotCluster is a cluster with the parallel-add cap given. +func slotCluster(limit int32) *simplyblockv1alpha2.StorageCluster { + cluster := &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{Name: "c", Namespace: "simplyblock"}, + Spec: simplyblockv1alpha2.StorageClusterSpec{ + StorageNodes: &simplyblockv1alpha2.StorageNodesSpec{ + MaxParallelNodeAdds: ptr.To(limit), + }, + }, + } + cluster.Status.UUID = aBackendClusterID + return cluster +} + +func slotReconciler(t *testing.T, nodes ...*simplyblockv1alpha2.StorageNode) *StorageNodeReconciler { + t.Helper() + scheme := testsupport.NewScheme(t, corev1.AddToScheme) + + objects := make([]client.Object, 0, len(nodes)) + for _, n := range nodes { + objects = append(objects, n) + } + apiClient := fake.NewClientBuilder().WithScheme(scheme). + WithObjects(objects...). + WithIndex(&simplyblockv1alpha2.StorageNode{}, clusterRefField, + func(o client.Object) []string { + return []string{o.(*simplyblockv1alpha2.StorageNode).Spec.ClusterRef} + }). + Build() + + return &StorageNodeReconciler{Client: apiClient, Scheme: scheme} +} + +// admitted is how many of the waiting nodes would advance on one pass, each +// deciding for itself as it does in a real reconcile. +func admitted( + t *testing.T, r *StorageNodeReconciler, + cluster *simplyblockv1alpha2.StorageCluster, + nodes ...*simplyblockv1alpha2.StorageNode, +) int { + t.Helper() + count := 0 + for _, node := range nodes { + next, _, err := r.awaitSlot(context.Background(), node, cluster) + if err != nil { + continue + } + if next == stepPosting { + count++ + } + } + return count +} + +// With nothing in flight and a cap of one, exactly one of three waiting nodes +// takes the slot. +func TestOnlyOneNodeTakesAFreeSlot(t *testing.T) { + a, b, c := waitingNode("w1"), waitingNode("w2"), waitingNode("w3") + r := slotReconciler(t, a, b, c) + + if got := admitted(t, r, slotCluster(1), a, b, c); got != 1 { + t.Errorf("%d nodes took a single free slot", got) + } +} + +// The cap is honored above one too: two free slots admit two of three. +func TestACapOfTwoAdmitsTwo(t *testing.T) { + a, b, c := waitingNode("w1"), waitingNode("w2"), waitingNode("w3") + r := slotReconciler(t, a, b, c) + + if got := admitted(t, r, slotCluster(2), a, b, c); got != 2 { + t.Errorf("%d nodes took two free slots", got) + } +} + +// The same node wins every time it is asked, so a node that was told to wait +// does not overtake one that was told to go on the next pass. +func TestTheChoiceIsStable(t *testing.T) { + a, b, c := waitingNode("w1"), waitingNode("w2"), waitingNode("w3") + r := slotReconciler(t, a, b, c) + cluster := slotCluster(1) + + first, _, err := r.awaitSlot(context.Background(), a, cluster) + if err != nil { + t.Fatal(err) + } + second, _, err := r.awaitSlot(context.Background(), a, cluster) + if err != nil { + t.Fatal(err) + } + if first != second { + t.Errorf("the same node was told %q then %q", first, second) + } +} + +// A node already in flight fills the cap, so nobody else advances. +func TestANodeInFlightFillsTheCap(t *testing.T) { + inflight := waitingNode("w1") + inflight.Status.Step.State = string(stepPosting) + b, c := waitingNode("w2"), waitingNode("w3") + r := slotReconciler(t, inflight, b, c) + + if got := admitted(t, r, slotCluster(1), b, c); got != 0 { + t.Errorf("%d nodes advanced past a full cap", got) + } +} diff --git a/operator/internal/controllers/node/storagenode_controller.go b/operator/internal/controllers/node/storagenode_controller.go index be375b152..fa5d38b5e 100644 --- a/operator/internal/controllers/node/storagenode_controller.go +++ b/operator/internal/controllers/node/storagenode_controller.go @@ -489,11 +489,47 @@ func (r *StorageNodeReconciler) awaitSlot( inFlight[sibling.Spec.WorkerNode] = struct{}{} } } - if int32(len(inFlight)) >= limit { + available := limit - int32(len(inFlight)) + if available <= 0 { return stepAwaitingSlot, false, blockedf(AwaitingSlot, "waiting for a node-add slot, %d of %d in flight", len(inFlight), limit) } + // Which of the waiting workers may take the free slots is decided from the + // set that is waiting, not from who has already written a claim. + // + // The difference is the whole of the cap. Reconciles are serialized per + // object and not across objects, so every node waiting for a slot reads the + // in-flight count before any of them has recorded taking one: the first add + // is correctly alone, and the instant it finishes every remaining node sees + // the same free slot and takes it. A cap that counts other people's writes + // holds exactly once. + // + // Ordering the contenders and admitting the first few needs nobody to have + // written anything. Two nodes reading one set reach one answer, and the + // answer does not change between passes, so a node told to wait is not + // overtaken by one told to wait beside it. + contenders := []string{node.Spec.WorkerNode} + for i := range siblings { + sibling := &siblings[i] + if sibling.Spec.WorkerNode == node.Spec.WorkerNode || claimedWorker(sibling) { + continue + } + if nodeStep(sibling.Status.Step.State) != stepAwaitingSlot { + // Not waiting for a slot yet, so not competing for this one. + continue + } + contenders = append(contenders, sibling.Spec.WorkerNode) + } + slices.Sort(contenders) + contenders = slices.Compact(contenders) + + if rank := slices.Index(contenders, node.Spec.WorkerNode); int32(rank) >= available { + return stepAwaitingSlot, false, blockedf(AwaitingSlot, + "waiting for a node-add slot, %d of %d in flight and %d worker(s) ahead", + len(inFlight), limit, rank) + } + // A FoundationDB worker waits for every other FoundationDB worker, whatever // the cap says. if r.hostsFoundationDB(ctx, node.Namespace, node.Spec.WorkerNode) { From 5063e2b4779ec353e2d3d791248387ca4f489f4a Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Wed, 16 Sep 2026 08:31:52 +0200 Subject: [PATCH 021/206] feat(deployment): a finished deployment activates the cluster it built The document stopped once the objects existed. What that left was a cluster serving nothing behind a document reporting Expanded, with nothing anywhere saying that one more thing was required of anybody: the activation was a StorageClusterOps that existed, was implemented, and was raised by nobody. The document is the one thing that knows when the deployment it describes is whole. status.nodeRefs names every StorageNode the expansion created, so the wait is over what this deployment built rather than over whatever happens to name the cluster -- a node somebody added later is not one it is waiting on, and a node it made that never came up is one it must not pass over. So the expansion gains a step after CreatingNodes: wait for its own nodes to report Online, then ask. The request is a StorageClusterOps like any other, raised by name so re-entering the step finds the one it raised rather than asking twice, and owned by the document that raised it. What happens to the operation afterward is that operation's business. The step carries the longest deadline of the expansion, which is not because activating is slow: every node this document created has to come online first, and the node-add cap serializes them. Co-Authored-By: Claude Opus 5 (1M context) --- ...mplyblock.io_clusterdeploymentconfigs.yaml | 2 +- .../v1alpha2/clusterdeploymentconfig_types.go | 11 +- ...mplyblock.io_clusterdeploymentconfigs.yaml | 2 +- .../controllers/deployment/activation_test.go | 117 ++++++++++++++++++ .../clusterdeploymentconfig_controller.go | 3 + .../internal/controllers/deployment/events.go | 12 +- .../controllers/deployment/expansion.go | 79 +++++++++++- ...mplyblock.io_clusterdeploymentconfigs.yaml | 2 +- 8 files changed, 221 insertions(+), 7 deletions(-) create mode 100644 operator/internal/controllers/deployment/activation_test.go diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml index 100fd3a05..95533924d 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml @@ -437,7 +437,7 @@ spec: type: object x-kubernetes-validations: - message: unknown step - rule: '!has(self.state) || self.state in [''Validating'',''CreatingCluster'',''AwaitingCluster'',''CreatingNodes'']' + rule: '!has(self.state) || self.state in [''Validating'',''CreatingCluster'',''AwaitingCluster'',''CreatingNodes'',''Activating'']' type: object type: object served: true diff --git a/operator/api/v1alpha2/clusterdeploymentconfig_types.go b/operator/api/v1alpha2/clusterdeploymentconfig_types.go index 9b5cec92b..312e4c468 100644 --- a/operator/api/v1alpha2/clusterdeploymentconfig_types.go +++ b/operator/api/v1alpha2/clusterdeploymentconfig_types.go @@ -66,6 +66,15 @@ const ( ClusterDeploymentConfigStepCreatingCluster ClusterDeploymentConfigStep = "CreatingCluster" ClusterDeploymentConfigStepAwaitingCluster ClusterDeploymentConfigStep = "AwaitingCluster" ClusterDeploymentConfigStepCreatingNodes ClusterDeploymentConfigStep = "CreatingNodes" + + // ClusterDeploymentConfigStepActivating waits for the nodes this document + // created and then asks for the cluster to be activated. + // + // The document knows how many nodes it made, so it knows when the deployment + // it describes is whole. Stopping at "the objects exist" would leave a + // cluster that serves nothing behind a document reporting Expanded, with + // nothing saying that one more thing is required of anybody. + ClusterDeploymentConfigStepActivating ClusterDeploymentConfigStep = "Activating" ) // KubernetesEnvironment is the distribution a deployment targets. The values are @@ -350,7 +359,7 @@ type ClusterDeploymentConfigStatus struct { Phase ClusterDeploymentConfigPhase `json:"phase,omitempty"` // Step is the position of the expansion machine within Expanding. - // +kubebuilder:validation:XValidation:rule="!has(self.state) || self.state in ['Validating','CreatingCluster','AwaitingCluster','CreatingNodes']",message="unknown step" + // +kubebuilder:validation:XValidation:rule="!has(self.state) || self.state in ['Validating','CreatingCluster','AwaitingCluster','CreatingNodes','Activating']",message="unknown step" // +optional Step statemachine.KubeSnapshot `json:"step,omitempty"` diff --git a/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml b/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml index 100fd3a05..95533924d 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml @@ -437,7 +437,7 @@ spec: type: object x-kubernetes-validations: - message: unknown step - rule: '!has(self.state) || self.state in [''Validating'',''CreatingCluster'',''AwaitingCluster'',''CreatingNodes'']' + rule: '!has(self.state) || self.state in [''Validating'',''CreatingCluster'',''AwaitingCluster'',''CreatingNodes'',''Activating'']' type: object type: object served: true diff --git a/operator/internal/controllers/deployment/activation_test.go b/operator/internal/controllers/deployment/activation_test.go new file mode 100644 index 000000000..e3532b1fa --- /dev/null +++ b/operator/internal/controllers/deployment/activation_test.go @@ -0,0 +1,117 @@ +// Whether a finished deployment activates the cluster it built. +// +// The document knows what it created: status.nodeRefs names every StorageNode +// the expansion made, so it knows exactly how many have to come online before +// the cluster is whole. A deployment that stops at "the objects exist" leaves a +// cluster that serves nothing and an administrator holding a document that says +// Expanded, with no indication that one more thing is required of them. +// +// So the expansion waits for the nodes it created and then asks for the +// activation itself. The request is a StorageClusterOps like any other, raised +// by name so that re-entering the step finds the one it raised rather than +// asking twice. + +package deployment + +import ( + "context" + "testing" + + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// anExpandedDocument is a document whose nodes have been created, with the node +// objects in the phase given. +func anExpandedDocument( + t *testing.T, phase simplyblockv1alpha2.StorageNodePhase, +) (*simplyblockv1alpha2.ClusterDeploymentConfig, *ClusterDeploymentConfigReconciler) { + t.Helper() + + config := aDocument(nil) + config.Status.ClusterRef = theCluster + names := []string{theCluster + "-worker-1-0", theCluster + "-worker-2-0"} + config.Status.NodeRefs = names + + objects := append(workers("worker-1", "worker-2"), config, aCluster(nil)) + for _, name := range names { + node := &simplyblockv1alpha2.StorageNode{ + ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: theNamespace}, + Spec: simplyblockv1alpha2.StorageNodeSpec{ClusterRef: theCluster}, + } + node.Status.Phase = phase + objects = append(objects, node) + } + return config, reconcilerFor(t, objects...) +} + +// Every node online is the deployment finished, so the activation is asked for. +func TestTheDeploymentActivatesOnceItsNodesAreOnline(t *testing.T) { + config, r := anExpandedDocument(t, simplyblockv1alpha2.StorageNodePhaseOnline) + + done, err := r.activateCluster(context.Background(), config) + if err != nil { + t.Fatalf("activate: %v", err) + } + if !done { + t.Error("the step did not finish with every node online") + } + + var ops simplyblockv1alpha2.StorageClusterOpsList + if err := r.List(context.Background(), &ops); err != nil { + t.Fatalf("listing the operations: %v", err) + } + if len(ops.Items) != 1 { + t.Fatalf("raised %d operations, want the one activation", len(ops.Items)) + } + if got := ops.Items[0].Spec.Action; got != simplyblockv1alpha2.StorageClusterOpsActionActivate { + t.Errorf("raised a %q operation", got) + } + if got := ops.Items[0].Spec.ClusterRef; got != theCluster { + t.Errorf("the operation names cluster %q", got) + } +} + +// A node still provisioning is a deployment that is not finished, so nothing is +// asked for yet. +func TestTheDeploymentWaitsWhileANodeIsStillComingUp(t *testing.T) { + config, r := anExpandedDocument(t, simplyblockv1alpha2.StorageNodePhaseProvisioning) + + done, err := r.activateCluster(context.Background(), config) + if err != nil { + t.Fatalf("activate: %v", err) + } + if done { + t.Error("the step finished while a node was still provisioning") + } + + var ops simplyblockv1alpha2.StorageClusterOpsList + if err := r.List(context.Background(), &ops); err != nil { + t.Fatalf("listing the operations: %v", err) + } + if len(ops.Items) != 0 { + t.Errorf("raised %d operations before the nodes were up", len(ops.Items)) + } +} + +// Re-entering the step finds the operation it raised rather than raising a +// second one, which is what makes the step safe to repeat. +func TestTheActivationIsAskedForOnce(t *testing.T) { + config, r := anExpandedDocument(t, simplyblockv1alpha2.StorageNodePhaseOnline) + + for range 3 { + if _, err := r.activateCluster(context.Background(), config); err != nil { + t.Fatalf("activate: %v", err) + } + } + + var ops simplyblockv1alpha2.StorageClusterOpsList + if err := r.List(context.Background(), &ops, client.InNamespace(theNamespace)); err != nil { + t.Fatalf("listing the operations: %v", err) + } + if len(ops.Items) != 1 { + t.Errorf("three passes raised %d operations", len(ops.Items)) + } +} diff --git a/operator/internal/controllers/deployment/clusterdeploymentconfig_controller.go b/operator/internal/controllers/deployment/clusterdeploymentconfig_controller.go index 71bfaddb2..04ba12502 100644 --- a/operator/internal/controllers/deployment/clusterdeploymentconfig_controller.go +++ b/operator/internal/controllers/deployment/clusterdeploymentconfig_controller.go @@ -87,6 +87,7 @@ type ClusterDeploymentConfigReconciler struct { // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=clusterdeploymentconfigs/status,verbs=get;update;patch // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storageclusters,verbs=get;list;watch;create // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodes,verbs=get;list;watch;create +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storageclusterops,verbs=get;list;watch;create // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=controlplanes,verbs=get;list;watch // +kubebuilder:rbac:groups="",resources=nodes,verbs=get;list;watch @@ -244,6 +245,8 @@ func (r *ClusterDeploymentConfigReconciler) performStep( return r.awaitCluster(ctx, config) case stepCreatingNodes: return r.createNodes(ctx, config) + case stepActivating: + return r.activateCluster(ctx, config) default: return false, fmt.Errorf("step %s belongs to no expansion this operator runs", current) } diff --git a/operator/internal/controllers/deployment/events.go b/operator/internal/controllers/deployment/events.go index 59399f95e..05413df0f 100644 --- a/operator/internal/controllers/deployment/events.go +++ b/operator/internal/controllers/deployment/events.go @@ -22,8 +22,16 @@ const ( // NoManagementInterface is a group that names no interface for the storage // nodes to bind their management address to. NoManagementInterface = "NoManagementInterface" - DeviceNotFound = "DeviceNotFound" - DeviceClassMismatch = "DeviceClassMismatch" + + // AwaitingNodes is the expansion waiting for a node it created to come + // online before it asks for the cluster to be activated. + AwaitingNodes = "AwaitingNodes" + + // ActivationRequested is the expansion having asked, which is the last thing + // a document does. + ActivationRequested = "ActivationRequested" + DeviceNotFound = "DeviceNotFound" + DeviceClassMismatch = "DeviceClassMismatch" // AwaitingApproval is the one that changes how the kind is used. A valid draft // nobody has approved looks identical to a controller that has not noticed it, diff --git a/operator/internal/controllers/deployment/expansion.go b/operator/internal/controllers/deployment/expansion.go index 9cc0ba1a4..dab626162 100644 --- a/operator/internal/controllers/deployment/expansion.go +++ b/operator/internal/controllers/deployment/expansion.go @@ -26,6 +26,7 @@ import ( apierrors "k8s.io/apimachinery/pkg/api/errors" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil" "github.com/simplyblock/atlas/kube" "github.com/simplyblock/atlas/ptr" @@ -42,6 +43,7 @@ const ( stepCreatingCluster = simplyblockv1alpha2.ClusterDeploymentConfigStepCreatingCluster stepAwaitingCluster = simplyblockv1alpha2.ClusterDeploymentConfigStepAwaitingCluster stepCreatingNodes = simplyblockv1alpha2.ClusterDeploymentConfigStepCreatingNodes + stepActivating = simplyblockv1alpha2.ClusterDeploymentConfigStepActivating ) // How long each step may take before it is reported as stuck. @@ -54,6 +56,11 @@ const ( creatingClusterDeadline = 5 * time.Minute awaitingClusterDeadline = 30 * time.Minute creatingNodesDeadline = 10 * time.Minute + + // activatingDeadline covers every node this document created coming online, + // which is a node add per worker and the cap serializes them. It is the + // longest step for that reason rather than because activating is slow. + activatingDeadline = 60 * time.Minute ) // expansionGraph declares the document's one state graph. @@ -78,7 +85,11 @@ func expansionGraph() statemachine.Config[configStep] { To: []configStep{stepCreatingNodes}, OnEnter: deadline(awaitingClusterDeadline), }, - stepCreatingNodes: {OnEnter: deadline(creatingNodesDeadline)}, + stepCreatingNodes: { + To: []configStep{stepActivating}, + OnEnter: deadline(creatingNodesDeadline), + }, + stepActivating: {OnEnter: deadline(activatingDeadline)}, }, } } @@ -171,6 +182,72 @@ func (r *ClusterDeploymentConfigReconciler) createCluster( return true, r.recordCluster(ctx, config, name) } +// activateCluster waits for the nodes this document created and then asks for +// the cluster to be activated. +// +// The wait is over status.nodeRefs rather than over whatever nodes happen to +// name the cluster, because the document is answering for what it built: a node +// somebody else added later is not one this deployment is waiting on, and a node +// this deployment made that never came up is one it must not pass over. +// +// The activation is a StorageClusterOps like any other, raised by name so that +// re-entering the step finds the one it raised rather than asking twice. What +// happens to it afterward is that operation's business; this document has +// described a deployment and asked for it, which is where its own job ends. +func (r *ClusterDeploymentConfigReconciler) activateCluster( + ctx context.Context, config *simplyblockv1alpha2.ClusterDeploymentConfig, +) (bool, error) { + if config.Status.ClusterRef == "" { + return false, refusef(ClusterNotFound, + "the document records no cluster to activate") + } + + for _, name := range config.Status.NodeRefs { + var node simplyblockv1alpha2.StorageNode + key := client.ObjectKey{Namespace: config.Namespace, Name: name} + if err := r.Get(ctx, key, &node); err != nil { + if apierrors.IsNotFound(err) { + // A node the document created and somebody has since removed is + // not one to wait for. The deployment it described is what it + // built, and this is no longer part of it. + continue + } + return false, fmt.Errorf("reading StorageNode %s: %w", name, err) + } + if node.Status.Phase != simplyblockv1alpha2.StorageNodePhaseOnline { + r.emit(config, corev1.EventTypeNormal, AwaitingNodes, fmt.Sprintf( + "StorageNode %s is %s rather than Online", name, node.Status.Phase)) + return false, nil + } + } + + name := config.Status.ClusterRef + "-activate" + ops := &simplyblockv1alpha2.StorageClusterOps{ + ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: config.Namespace}, + Spec: simplyblockv1alpha2.StorageClusterOpsSpec{ + ClusterRef: config.Status.ClusterRef, + Action: simplyblockv1alpha2.StorageClusterOpsActionActivate, + }, + } + if err := controllerutil.SetControllerReference(config, ops, r.Scheme); err != nil { + return false, err + } + if err := r.Create(ctx, ops); err != nil { + if !apierrors.IsAlreadyExists(err) { + return false, fmt.Errorf("asking for cluster %s to be activated: %w", + config.Status.ClusterRef, err) + } + // Raised on an earlier pass, which is the step being re-entered rather + // than anything having gone wrong. + return true, nil + } + + r.emit(config, corev1.EventTypeNormal, ActivationRequested, fmt.Sprintf( + "Every node is online, so cluster %s was asked to activate", + config.Status.ClusterRef)) + return true, nil +} + // buildCluster is the StorageCluster the document's template describes. func (r *ClusterDeploymentConfigReconciler) buildCluster( config *simplyblockv1alpha2.ClusterDeploymentConfig, name string, diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml index 100fd3a05..95533924d 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml @@ -437,7 +437,7 @@ spec: type: object x-kubernetes-validations: - message: unknown step - rule: '!has(self.state) || self.state in [''Validating'',''CreatingCluster'',''AwaitingCluster'',''CreatingNodes'']' + rule: '!has(self.state) || self.state in [''Validating'',''CreatingCluster'',''AwaitingCluster'',''CreatingNodes'',''Activating'']' type: object type: object served: true From dd83d4059326ece4e2e7e2ccda68aa497a67fd83 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Wed, 16 Sep 2026 09:04:45 +0200 Subject: [PATCH 022/206] fix(cluster): the lock release retries rather than failing the reconcile Releasing an operation's lock is a compare-and-swap on purpose: it re-reads the cluster, checks the lock is still this operation's, and writes with an optimistic lock, so a release that lost a race never clears a lock somebody else now holds. Two operations running against one cluster with neither knowing is worse than a lock held a moment too long. The conflict was reported as a reconciler error to preserve that, and it is not the only thing that does. The cluster's own reconciler writes its status on every pass, so a conflict here is ordinary rather than exceptional, and what it produced was a failed reconcile with a stack trace for an operation that had just succeeded: release the lock on cluster ...: Operation cannot be fulfilled on storageclusters ...: the object has been modified Retrying on a fresh read keeps the property intact, because the check is inside the retry: every attempt re-reads and re-tests whose lock it is, so a lock taken by somebody else between two attempts is found on the next read and cleared by neither. What would be wrong is swallowing the conflict without the re-read -- that lets the caller reach a terminal phase and drop its finalizer while the cluster stays locked by an object that no longer exists. Co-Authored-By: Claude Opus 5 (1M context) --- .../controllers/cluster/release_lock_test.go | 137 ++++++++++++++++++ .../cluster/storageclusterops_controller.go | 51 ++++--- 2 files changed, 169 insertions(+), 19 deletions(-) create mode 100644 operator/internal/controllers/cluster/release_lock_test.go diff --git a/operator/internal/controllers/cluster/release_lock_test.go b/operator/internal/controllers/cluster/release_lock_test.go new file mode 100644 index 000000000..df1dd4ce8 --- /dev/null +++ b/operator/internal/controllers/cluster/release_lock_test.go @@ -0,0 +1,137 @@ +// Releasing the lock an operation holds on its cluster, against a status the +// cluster's own reconciler is writing at the same time. +// +// The release is a compare-and-swap on purpose: it re-reads the cluster, checks +// the lock is still this operation's, and writes with an optimistic lock, so a +// release that lost a race never clears a lock somebody else now holds. Two +// operations running against one cluster with neither knowing is worse than a +// lock held a moment too long. +// +// Reporting the conflict as a reconciler error preserved that, and it is not the +// only thing that does. The cluster's status is written by its own reconciler on +// every pass, so a conflict here is ordinary rather than exceptional, and what it +// produced was a failed reconcile with a stack trace for an operation that had +// just succeeded. Retrying on a fresh read keeps the compare-and-swap — the check +// runs again on every attempt — and resolves what it can resolve itself. + +package cluster + +import ( + "context" + "testing" + + apierrors "k8s.io/apimachinery/pkg/api/errors" + "k8s.io/apimachinery/pkg/runtime/schema" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + "sigs.k8s.io/controller-runtime/pkg/client/interceptor" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +type errStaleCluster struct{} + +func (errStaleCluster) Error() string { + return "the object has been modified; please apply your changes to the latest version and try again" +} + +// conflictingStatus refuses the first n status writes the way the API server +// does when something else has written the object since it was read. +func conflictingStatus(remaining *int) interceptor.Funcs { + refuse := func(name string) error { + *remaining-- + return apierrors.NewConflict( + schema.GroupResource{ + Group: simplyblockv1alpha2.GroupVersion.Group, + Resource: "storageclusters", + }, + name, errStaleCluster{}) + } + return interceptor.Funcs{ + SubResourcePatch: func( + ctx context.Context, c client.Client, subResource string, + obj client.Object, patch client.Patch, opts ...client.SubResourcePatchOption, + ) error { + if *remaining > 0 { + return refuse(obj.GetName()) + } + return c.SubResource(subResource).Patch(ctx, obj, patch, opts...) + }, + } +} + +func lockedCluster(holder string) *simplyblockv1alpha2.StorageCluster { + cluster := &simplyblockv1alpha2.StorageCluster{ObjectMeta: objectMeta("a-cluster")} + cluster.Status.UUID = "cluster-uuid" + cluster.Status.ActiveOpsRef = holder + return cluster +} + +func anOps(name string) *simplyblockv1alpha2.StorageClusterOps { + return &simplyblockv1alpha2.StorageClusterOps{ + ObjectMeta: objectMeta(name), + Spec: simplyblockv1alpha2.StorageClusterOpsSpec{ + ClusterRef: "a-cluster", + Action: simplyblockv1alpha2.StorageClusterOpsActionActivate, + }, + } +} + +func releaser(t *testing.T, funcs interceptor.Funcs, objects ...client.Object) *StorageClusterOpsReconciler { + t.Helper() + return &StorageClusterOpsReconciler{ + Client: fake.NewClientBuilder(). + WithScheme(testScheme(t)). + WithStatusSubresource( + &simplyblockv1alpha2.StorageCluster{}, + &simplyblockv1alpha2.StorageClusterOps{}, + ). + WithObjects(objects...). + WithInterceptorFuncs(funcs). + Build(), + Scheme: testScheme(t), + Recorder: &recorder{}, + } +} + +// A release that loses a race retries rather than failing the operation. +func TestTheLockReleaseRetriesAConflict(t *testing.T) { + remaining := 1 + ops := anOps("a-cluster-activate") + r := releaser(t, conflictingStatus(&remaining), lockedCluster("a-cluster-activate"), ops) + + if err := r.releaseLock(context.Background(), ops); err != nil { + t.Fatalf("the release failed on a conflict it should have retried: %v", err) + } + if remaining != 0 { + t.Error("the conflict was never raised, so this proves nothing") + } + + var cluster simplyblockv1alpha2.StorageCluster + if err := r.Get(context.Background(), client.ObjectKeyFromObject(lockedCluster("")), &cluster); err != nil { + t.Fatalf("reading the cluster: %v", err) + } + if cluster.Status.ActiveOpsRef != "" { + t.Errorf("the lock is still held by %q", cluster.Status.ActiveOpsRef) + } +} + +// The compare-and-swap survives the retry: a lock that has changed hands is not +// cleared, whatever this operation thinks it holds. +func TestTheReleaseNeverClearsSomebodyElsesLock(t *testing.T) { + remaining := 0 + ops := anOps("a-cluster-activate") + r := releaser(t, conflictingStatus(&remaining), lockedCluster("another-operation"), ops) + + if err := r.releaseLock(context.Background(), ops); err != nil { + t.Fatalf("release: %v", err) + } + + var cluster simplyblockv1alpha2.StorageCluster + if err := r.Get(context.Background(), client.ObjectKeyFromObject(lockedCluster("")), &cluster); err != nil { + t.Fatalf("reading the cluster: %v", err) + } + if cluster.Status.ActiveOpsRef != "another-operation" { + t.Errorf("the lock now reads %q, want the other operation's", cluster.Status.ActiveOpsRef) + } +} diff --git a/operator/internal/controllers/cluster/storageclusterops_controller.go b/operator/internal/controllers/cluster/storageclusterops_controller.go index 85ecc9c44..517c9a74b 100644 --- a/operator/internal/controllers/cluster/storageclusterops_controller.go +++ b/operator/internal/controllers/cluster/storageclusterops_controller.go @@ -593,32 +593,45 @@ func (r *StorageClusterOpsReconciler) acquireLock( func (r *StorageClusterOpsReconciler) releaseLock( ctx context.Context, ops *simplyblockv1alpha2.StorageClusterOps, ) error { - var cluster simplyblockv1alpha2.StorageCluster key := types.NamespacedName{Name: ops.Spec.ClusterRef, Namespace: ops.Namespace} - err := r.Get(ctx, key, &cluster) + + // The conflict is retried here rather than reported, and the compare-and-swap + // is what makes that safe rather than a shortcut: every attempt re-reads the + // cluster and checks the lock is still this operation's, so a release that + // lost a race to somebody taking the lock finds that on the next read and + // clears nothing. Swallowing the conflict without the re-read is the thing + // that would be wrong — it would let the caller reach a terminal phase and + // drop its finalizer while the cluster stayed locked by an object that no + // longer exists. + // + // Reporting it failed the reconcile instead, which preserved the same + // property by a longer route: a stack trace for an operation that had just + // succeeded, over a conflict with the cluster's own reconciler writing the + // status it writes on every pass. + err := retry.RetryOnConflict(retry.DefaultRetry, func() error { + var cluster simplyblockv1alpha2.StorageCluster + if err := r.Get(ctx, key, &cluster); err != nil { + return err + } + if cluster.Status.ActiveOpsRef != ops.Name { + // Never held, already released, or taken by somebody else between + // two attempts. None of them is this operation's to undo. + return nil + } + + patch := client.MergeFromWithOptions(cluster.DeepCopy(), + client.MergeFromWithOptimisticLock{}) + cluster.Status.ActiveOpsRef = "" + return r.Status().Patch(ctx, &cluster, patch) + }) if apierrors.IsNotFound(err) { return nil } if err != nil { - return err - } - if cluster.Status.ActiveOpsRef != ops.Name { - return nil + return fmt.Errorf("release the lock on cluster %s: %w", ops.Spec.ClusterRef, err) } - patch := client.MergeFromWithOptions(cluster.DeepCopy(), client.MergeFromWithOptimisticLock{}) - cluster.Status.ActiveOpsRef = "" - if err := r.Status().Patch(ctx, &cluster, patch); err != nil { - // A conflict is reported rather than swallowed, and that is the whole - // point of returning an error here. Somebody else wrote the cluster's - // status between the read and the write, so the lock this operation - // still holds was not cleared; treating that as a release lets the - // caller reach a terminal phase and the finalizer go, and the cluster - // stays locked by an object that no longer exists. Reporting it - // retries the read and the release on the next pass. - return fmt.Errorf("release the lock on cluster %s: %w", cluster.Name, err) - } - operationActiveState.WithLabelValues(cluster.Name).Set(0) + operationActiveState.WithLabelValues(ops.Spec.ClusterRef).Set(0) return nil } From bc02ba4bc608879a482bf45f4ed32d90ea6ac4ab Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Wed, 16 Sep 2026 09:21:02 +0200 Subject: [PATCH 023/206] feat(cluster): a cluster being built says so rather than reporting a fault The phases after the creation path are a reading of the status the control plane publishes, and everything the operator had no reading for fell through to Unavailable -- which the field documents as not serving, and not because anybody asked. Three statuses landed there that are the opposite of that. A cluster being expanded and a cluster being activated are not serving precisely because somebody asked, and one that has never been activated is not serving because it is not finished. So a deployment reported a fault for its whole length, and an activation reported one for as long as it ran. unready, in_creation, and in_expansion now read as Provisioning, which is the word StorageNode already uses for a thing being built rather than a second word for one idea. in_activation is its own phase, because activation is not only the last step of a deployment: an expansion ends in one and so does recovering from a suspension, long after anything was being built. Unavailable keeps its meaning by being what is left -- a status this operator has no reading for, rather than the name for every cluster that is not serving. The mapping table that lived in storagecluster_controller_test.go is replaced by one in phase_test.go covering every status rather than six of them, because two tables for one function drift and these two would have had to disagree first. Co-Authored-By: Claude Opus 5 (1M context) --- ...torage.simplyblock.io_storageclusters.yaml | 2 + operator/api/v1alpha2/storagecluster_types.go | 21 +++++- ...torage.simplyblock.io_storageclusters.yaml | 2 + .../controllers/cluster/phase_test.go | 73 +++++++++++++++++++ .../cluster/storagecluster_controller.go | 16 +++- .../cluster/storagecluster_controller_test.go | 21 +----- ...torage.simplyblock.io_storageclusters.yaml | 2 + 7 files changed, 114 insertions(+), 23 deletions(-) create mode 100644 operator/internal/controllers/cluster/phase_test.go diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusters.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusters.yaml index 99cc036e1..b36e750f5 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusters.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusters.yaml @@ -1418,6 +1418,8 @@ spec: enum: - Pending - Creating + - Provisioning + - Activating - Online - Degraded - Unavailable diff --git a/operator/api/v1alpha2/storagecluster_types.go b/operator/api/v1alpha2/storagecluster_types.go index a71d443df..c4f1de752 100644 --- a/operator/api/v1alpha2/storagecluster_types.go +++ b/operator/api/v1alpha2/storagecluster_types.go @@ -29,7 +29,7 @@ import ( // first two values are the operator's own creation path; the rest are its // reading of the lifecycle status.status carries in the control plane's own // spelling. -// +kubebuilder:validation:Enum=Pending;Creating;Online;Degraded;Unavailable;Suspended +// +kubebuilder:validation:Enum=Pending;Creating;Provisioning;Activating;Online;Degraded;Unavailable;Suspended type StorageClusterPhase string const ( @@ -40,6 +40,20 @@ const ( // StorageClusterPhaseCreating: the creation machine is running. StorageClusterPhaseCreating StorageClusterPhase = "Creating" + // StorageClusterPhaseProvisioning: the cluster exists in the control plane + // and is being built up — its first nodes are joining, or an expansion is + // adding more. It is not serving and there is nothing wrong with it, which + // is the distinction Unavailable cannot carry. + StorageClusterPhaseProvisioning StorageClusterPhase = "Provisioning" + + // StorageClusterPhaseActivating: the control plane is activating the + // cluster. + // + // It is a phase of its own rather than part of Provisioning because it is + // not only the last step of a deployment: an expansion ends in one, and so + // does recovering from a suspension, long after anything was being built. + StorageClusterPhaseActivating StorageClusterPhase = "Activating" + // StorageClusterPhaseOnline: the control plane reports the cluster active // and serving. StorageClusterPhaseOnline StorageClusterPhase = "Online" @@ -50,6 +64,11 @@ const ( // StorageClusterPhaseUnavailable: not serving, and not because anybody // asked. + // + // It is what is left once the statuses that mean something has been asked + // for are read as themselves, which is what keeps it worth reporting: a + // cluster in this phase is one whose status this operator has no reading + // for, rather than every cluster that is not currently serving. StorageClusterPhaseUnavailable StorageClusterPhase = "Unavailable" // StorageClusterPhaseSuspended: shut down deliberately, which is where a diff --git a/operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml b/operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml index 99cc036e1..b36e750f5 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml @@ -1418,6 +1418,8 @@ spec: enum: - Pending - Creating + - Provisioning + - Activating - Online - Degraded - Unavailable diff --git a/operator/internal/controllers/cluster/phase_test.go b/operator/internal/controllers/cluster/phase_test.go new file mode 100644 index 000000000..67e974020 --- /dev/null +++ b/operator/internal/controllers/cluster/phase_test.go @@ -0,0 +1,73 @@ +// What the operator calls a cluster it is still building. +// +// The phases after the creation path are a reading of the status the control +// plane publishes, and everything the operator did not recognize was read as +// Unavailable, which the field documents as not serving and not because anybody +// asked. Three of those +// statuses are the opposite of that. A cluster being expanded and a cluster +// being activated are not serving precisely because somebody asked, and one +// that has never been activated is not serving because it is not finished. +// +// So a deployment reported a fault for its whole length, and an activation — +// which is a step of every deployment and of every recovery from a suspension — +// reported one for as long as it ran. + +package cluster + +import ( + "testing" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +func TestThePhaseReadsTheControlPlanesLifecycle(t *testing.T) { + for _, tc := range []struct { + status string + want simplyblockv1alpha2.StorageClusterPhase + }{ + {"active", simplyblockv1alpha2.StorageClusterPhaseOnline}, + {"degraded", simplyblockv1alpha2.StorageClusterPhaseDegraded}, + {"read_only", simplyblockv1alpha2.StorageClusterPhaseDegraded}, + {"suspended", simplyblockv1alpha2.StorageClusterPhaseSuspended}, + {"", simplyblockv1alpha2.StorageClusterPhasePending}, + + // The cluster exists and is being built up: not serving, and nothing + // wrong with it. + {"unready", simplyblockv1alpha2.StorageClusterPhaseProvisioning}, + {"in_creation", simplyblockv1alpha2.StorageClusterPhaseProvisioning}, + {"in_expansion", simplyblockv1alpha2.StorageClusterPhaseProvisioning}, + + // Activation is its own phase because it is not only a step of a + // deployment: an expansion and a recovery from a suspension both end in + // one, long after anything was being provisioned. + {"in_activation", simplyblockv1alpha2.StorageClusterPhaseActivating}, + + // Unavailable keeps its meaning by being what is left: a status this + // operator has no reading for. + {"something_new", simplyblockv1alpha2.StorageClusterPhaseUnavailable}, + } { + t.Run(tc.status, func(t *testing.T) { + if got := phaseFor(tc.status); got != tc.want { + t.Errorf("phaseFor(%q) = %q, want %q", tc.status, got, tc.want) + } + }) + } +} + +// Every phase the mapping can produce has a gauge series, or a dashboard asking +// which phase a cluster is in gets no answer for the phases added last. +func TestEveryPhaseIsPublished(t *testing.T) { + published := map[simplyblockv1alpha2.StorageClusterPhase]bool{} + for _, phase := range allPhases { + published[phase] = true + } + + for _, phase := range []simplyblockv1alpha2.StorageClusterPhase{ + simplyblockv1alpha2.StorageClusterPhaseProvisioning, + simplyblockv1alpha2.StorageClusterPhaseActivating, + } { + if !published[phase] { + t.Errorf("%s has no gauge series", phase) + } + } +} diff --git a/operator/internal/controllers/cluster/storagecluster_controller.go b/operator/internal/controllers/cluster/storagecluster_controller.go index 0fc48570e..59963e7f6 100644 --- a/operator/internal/controllers/cluster/storagecluster_controller.go +++ b/operator/internal/controllers/cluster/storagecluster_controller.go @@ -1215,6 +1215,8 @@ func (r *StorageClusterReconciler) observePhase(cluster *simplyblockv1alpha2.Sto var allPhases = []simplyblockv1alpha2.StorageClusterPhase{ simplyblockv1alpha2.StorageClusterPhasePending, simplyblockv1alpha2.StorageClusterPhaseCreating, + simplyblockv1alpha2.StorageClusterPhaseProvisioning, + simplyblockv1alpha2.StorageClusterPhaseActivating, simplyblockv1alpha2.StorageClusterPhaseOnline, simplyblockv1alpha2.StorageClusterPhaseDegraded, simplyblockv1alpha2.StorageClusterPhaseUnavailable, @@ -1314,10 +1316,18 @@ func phaseFor(status string) simplyblockv1alpha2.StorageClusterPhase { return simplyblockv1alpha2.StorageClusterPhaseSuspended case "": return simplyblockv1alpha2.StorageClusterPhasePending + case "in_activation": + // Activation is asked for, and by more than a deployment: an expansion + // ends in one and so does recovering from a suspension. + return simplyblockv1alpha2.StorageClusterPhaseActivating + case utils.ClusterStatusUnready, "in_creation", "in_expansion": + // The cluster exists and is being built up. Not serving, and nothing + // wrong with it. + return simplyblockv1alpha2.StorageClusterPhaseProvisioning default: - // Everything else the control plane reports — unready, in_expansion, - // in_activation — is the cluster not serving for a reason nobody asked - // for, which is what Unavailable means. + // A status this operator has no reading for, which is what makes + // Unavailable worth reporting rather than the name for every cluster + // that is not currently serving. return simplyblockv1alpha2.StorageClusterPhaseUnavailable } } diff --git a/operator/internal/controllers/cluster/storagecluster_controller_test.go b/operator/internal/controllers/cluster/storagecluster_controller_test.go index 2773f9b4b..008852dc5 100644 --- a/operator/internal/controllers/cluster/storagecluster_controller_test.go +++ b/operator/internal/controllers/cluster/storagecluster_controller_test.go @@ -312,25 +312,8 @@ func TestAMissingBackupSecretHoldsTheCreation(t *testing.T) { // Steady state // --------------------------------------------------------------------------- -// The phase is the operator's reading of the control plane's own lifecycle -// string, which is why the values on one side are lowercase and the values on -// the other are not. -func TestThePhaseFollowsTheControlPlanesStatus(t *testing.T) { - for status, want := range map[string]simplyblockv1alpha2.StorageClusterPhase{ - utils.ClusterStatusActive: simplyblockv1alpha2.StorageClusterPhaseOnline, - "degraded": simplyblockv1alpha2.StorageClusterPhaseDegraded, - "read_only": simplyblockv1alpha2.StorageClusterPhaseDegraded, - utils.ClusterStatusSuspended: simplyblockv1alpha2.StorageClusterPhaseSuspended, - utils.ClusterStatusUnready: simplyblockv1alpha2.StorageClusterPhaseUnavailable, - "in_expansion": simplyblockv1alpha2.StorageClusterPhaseUnavailable, - } { - t.Run(status, func(t *testing.T) { - if got := phaseFor(status); got != want { - t.Errorf("phaseFor(%q) = %q, want %q", status, got, want) - } - }) - } -} +// The mapping this file used to assert here lives in phase_test.go, which covers +// every status rather than six of them. Two tables for one function drift. // status.tasks is a window on the present: only running and pending tasks are // in it, newest first, and never more than twenty. diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusters.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusters.yaml index 99cc036e1..b36e750f5 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusters.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusters.yaml @@ -1418,6 +1418,8 @@ spec: enum: - Pending - Creating + - Provisioning + - Activating - Online - Degraded - Unavailable From 849788bd561c72d4c35fc9f7f96960e5fd239a53 Mon Sep 17 00:00:00 2001 From: noctarius aka Christoph Engelbert Date: Wed, 16 Sep 2026 17:08:09 +0200 Subject: [PATCH 024/206] feat(operator): ControlPlane moves to v1alpha2, and the operator installs it (#544) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * feat(operator): ControlPlane moves to v1alpha2, and the operator installs it The registered ControlPlane was an observation with a spec field attached: the chart brought the control plane up, and the object polled a readiness endpoint it learned about from an environment variable. design-controlplane.md replaces both halves, and this is that rework. spec.source says which control plane the object means. A managed one is what the operator installs, and an external one is an endpoint and a credentials Secret the operator probes and reports on. status.endpoint publishes the resolved base URL either way, so one object answers where the control plane is rather than every controller resolving SIMPLYBLOCK_WEBAPI_BASE_URL for itself. The install is a five-step machine over atlas-lib/statemachine: the FoundationDB operator, its RBAC, and the FoundationDBCluster; a wait on that cluster reporting itself available; the object store; the management API with the services beside it; and a wait on the readiness probe. Every step is a server-side apply under a stable field manager, which is what makes re-entering one a no-op and why the machine carries no triggered flag. Steady state re-applies the same objects on every pass, so an object somebody deleted comes back and one somebody edited is corrected — which is why no operation exists for checking the install. The phase is the worst verdict across the probe and eight watched components, and only the management API and the database may produce Unavailable. That asymmetry is the safety property: Unavailable holds every controller in the operator, so the set of things able to cause it is a closed list, and a component added without a decision lands in the non-essential default. ControlPlaneOps is new, with the three operations that are not expressible as desired state. Restart recycles a scope of workloads, draining first only when it names something depended on. Upgrade runs Preflight, drains, writes the image onto the entity, and verifies the reported version afterward, so a rollout that failed back is a Failed operation rather than an Available control plane running the old version. Backup asks FoundationDB rather than implementing one, and does not own the FoundationDBBackup it creates. A validating webhook refuses an operation naming an external control plane at creation, since all three act on something the operator installed. What the install covers is a base control plane, which is what the chart renders with observability disabled, measured against a live deployment. The observability half stays with the chart: none of it appears in a step of the machine, and every component of it is non-essential. controlplane.managedByOperator is the chart flag for §12 Q2 and defaults to the operator. With it set the chart renders only the ControlPlane object beside the FoundationDB CRDs, the Prometheus configuration, and the log-collector RBAC; setting it to false renders the previous templates byte-for-byte. TLS is the one configuration the spec cannot express, so the chart refuses the two together rather than bringing a control plane up in plaintext. v1alpha1 stays as the spoke. Its conversion now resolves an absent image toward a managed source, because the hub requires exactly one member of spec.source and every object stored at v1alpha1 is one the chart installed. Co-Authored-By: Claude Opus 5 (1M context) * fix(operator): make the ControlPlane CRD generate the same way twice CI failed the helm-sync and build-installer drift gates on output that varied between generator runs rather than on anything stale. Two properties of the ControlPlane CRD came out differently each time controller-gen ran, and both had the same cause: two sources disagreeing about one value, resolved by map iteration. spec.source carried a +kubebuilder:validation:XValidation marker and +k8s:immutable. controller-gen v0.21.0 emits a field's marker-derived rules and its injected immutability rule into one x-kubernetes-validations list, in an order that varied run to run — four of six one way locally, two of six the other. It is the only field in the repository carrying both, which is why this has never flaked before. shortNames is resource-level and both versions of a kind feed one CRD. v1alpha2 declared shortName=cp and v1alpha1 declared a bare scope, so the merge picked one per run and the short name appeared in five of twelve regenerations. StorageNode already repeats its short name on both versions, and this follows it. Fixing the first uncovered a real bug, which a CEL test written against a real apiserver caught: +k8s:immutable on spec.source emits self == oldSelf over the whole struct, so it froze the image inside the block along with the choice of block. design-controlplane.md §3.2 says the block is immutable and its members are not, and the Upgrade action's Applying step writes spec.source.managed.image — which the apiserver rejected. The rule is now explicit about what it freezes: which member is set cannot change, and what is inside it can. Both rules move onto ControlPlaneSource, where two markers of the same kind are emitted in source order, and the field keeps only Required. cel_validation_test.go covers all three outcomes against envtest, including the image edit that used to be refused. Co-Authored-By: Claude Opus 5 (1M context) * fix(operator): regenerate the CRDs after the v1alpha1 prose edit A doc comment becomes a CRD description, so the house-style fixes to ControlPlaneSpec.Image and ControlPlaneStatus.Message left the four derived copies of the ControlPlane CRD carrying the old text: config/crd/bases, the embedded copy the upgrade tool applies, dist/install.yaml, and the chart's. No behavior changes. The regeneration is the whole of it. Co-Authored-By: Claude Opus 5 (1M context) * fix(operator): the eight findings of the review on #544 Every one was real, and three were asserted as correct by tests written with the code they covered. The chart flag did not hand the install back. With managedByOperator false the templates were gated out and the ControlPlane still named a managed source, so the operator installed the same objects the chart had just rendered and both owned them, which is the outcome the flag exists to avoid. It now names an external source pointing at the in-cluster Service, which is what the operator's side of a chart-installed control plane actually is: something that already exists, to be resolved and probed rather than applied. That needs a source with no static credential, so credentialsSecretRef becomes optional and an absent one means an unauthenticated probe. The readiness endpoint takes none. Initializing converted to a phase v1alpha2 rejects. The two enums share no value, and only Ready was mapped, so a control plane observed while FoundationDB was starting became an object the API server refuses on write. Both directions are now written out, and the downward table is no longer derived by inverting the upward one: the hub holds four phases and this version two, so Degraded and Unavailable land on the value that tells a v1alpha1 reader the same thing. The operation's parameters were mutable after admission. Preflight reads spec.upgrade.image and Applying writes it several steps later, so clearing the block between them dereferenced nil. The spec is now frozen except for abort, which is the one field meant to be set after the operation starts, and applying and verifying guard the block besides. The singleton was per name and not per Kubernetes cluster. A "simplyblock" object in a second namespace reconciled, applied the same fixed-name cluster-scoped roles under the same managed-by label, and either one's finalizer deleted what the other needed. The older object holds the install and the younger reports that it does not, which is the rule SimplyblockDriver already used for the same reason. A Restart scoped to the FoundationDB cluster passed validation, recycled nothing, found a healthy database on the wait that followed, and reported success. The scope is checked against what can be rolled rather than against the component table. Deleting a running operation released the control plane's lock mid-rollout. The webhook took create only, so the controller's deletion path dropped the lock and the finalizer from any step. It now takes delete as well and refuses the four steps that have started something nothing else would finish, which is the guard StorageBackupOps already carried. caBundleSecretRef was declared and wired to nothing, so an endpoint signed by a private CA could not be reached. Resolution now builds the transport the probes use, and a bundle that is named and unusable is an error rather than a silent fall back to the system trust store. status.endpoint was published and read by nothing. Both control-plane clients now resolve it per call and keep their startup client where the object has published none, so an external control plane is reachable and adopting this changes nothing for a deployment that has not. Co-Authored-By: Claude Opus 5 (1M context) * chore(operator): rebuild the installer against the current develop dist/install.yaml is generated from every CRD rather than from the ones a branch touches, so it carries develop's kinds as well as this branch's and has to be rebuilt whenever either moves. This picks up the ClusterDeploymentConfig fields and the StorageCluster phase develop added. Co-Authored-By: Claude Opus 5 (1M context) * feat(operator): deployment profiles, and the source names they settle A deployment is one of two things, and the chart now says which: standalone, where this cluster hosts its own control plane, or managed, where a control plane elsewhere manages this cluster's storage. The API's two source members are renamed to match, because the word "managed" was about to mean opposite things in the chart and in the CRD. spec.source.local is a control plane this cluster hosts and the operator installs. spec.source.managed is one somewhere else that this cluster registers with. The word is about what manages the storage clusters rather than about who runs the operator, which is the sense the hosted offering uses. Neither name has shipped in a release, so the rename costs nothing but the edit. controlplane.profile replaces controlplane.managedByOperator. The bool had two states and the profiles have two, but not the same two: the bool chose who installed locally, and the profile chooses whether anything is installed locally at all. The chart-installs-locally path is dropped, so nothing renders the control-plane workloads any more and the templates that only held them are deleted. Their observability halves stay. Three guards replace the one the bool carried. An unknown profile is refused by name. The managed profile is refused without an endpoint, since nothing local can stand in. TLS with standalone is still refused, and the message no longer points at a value that no longer exists. A fourth guard is new and is the one that matters. Helm deletes what the old release manifest held and the new one does not, and before this the chart's manifest held the FoundationDBCluster: upgrading a 26.2.x release in place would have pruned the database holding every cluster definition, node registration, and lvol record, and the operator would then have built an empty one. The chart now looks for a control plane an earlier release installed and refuses, naming the annotations that make the upgrade safe. The ControlPlane object itself carries helm.sh/resource-policy: keep for the same reason, and for a second one: Helm deletes the operator and the CR in one uninstall, and with the operator gone first nothing remains to release the finalizer, so the object cannot go and the release sits in `uninstalling`. Co-Authored-By: Claude Opus 5 (1M context) * refactor(helm): deployment.profile, and drop operator.enabled The profile moves out of the controlplane block. It selects the shape of the whole deployment rather than only where the control plane is, and a managed one will carry more than a source member: an edge cluster somebody else administers differs in what it runs and in how its StorageClusters are configured, and those settings belong beside the profile rather than repeated across the blocks they touch. operator.enabled goes. It was true in any deployment anybody wants: with it false the chart renders the CRDs, a snapshot controller, RBAC, and a secret, and neither the operator nor the CSI driver, which is not a deployment so much as half of one. Sixteen templates lose the gate. Helm evaluates a subchart condition as a boolean value path and cannot read a profile string, so Prometheus and the reloader take their own. That is the ordinary shape and truer besides: neither is implied by the operator existing, and a managed deployment reports to the control plane that administers it rather than to a Prometheus of its own. NOTES.txt is rewritten. Its wait command asked for a phase called Ready, which this branch renamed to Available, so it would have waited out its timeout and failed on every install. It also carried a branch for the operator being disabled, which nothing can now reach, and lost the first letter of "Wait" somewhere before this change. Co-Authored-By: Claude Opus 5 (1M context) * docs(helm): cut the values comments to what a reader has to set The deployment block ran to twenty-two lines of comment for one string, most of it explaining why the value sits where it does rather than what setting it does. Co-Authored-By: Claude Opus 5 (1M context) * docs: state the constraint rather than the system that was not built The house style gives three shapes to avoid: counterfactuals describing a system nobody built, rebuttals leading with what a thing is not, and notes reporting the writer's own history. The comments on this branch carried forty-four of them. Twenty-three are rewritten to state the constraint the code works under. Twenty-one are kept: each names the defect a regression test guards, which is the test's reason for existing and what the regression-test skill asks for. Twenty-one em dashes are replaced by the mark that was meant, which is a colon, a comma, parentheses, or a full stop in every case. The rename in an earlier commit moved symbols and struct tags, so eight comments were left describing the members by their old names. The conversion's said it writes the managed member where it writes the local one. Co-Authored-By: Claude Opus 5 (1M context) * fix(helm): restore the admission configurations this branch dropped templates/simplyblock-operator-webhook.yaml opened with a guard on .Values.operator.enabled. Removing that value in f0b72bc7 left the guard reading nothing, so the file rendered empty and the upgrade pruned the webhook Service together with both admission configurations. The operator went on registering thirteen handlers and serving them on :9443 with nothing in the cluster routing to them, which is every validating and mutating webhook it has, not only the ones this branch adds. cert-rotation reported it twice, once per configuration, as a certificate it could not update. The guard is dropped at its source in sync-from-operator.sh and the template regenerated. check-rendered-objects.sh is the test. helm template exits zero for a template that renders to nothing, which is why neither the chart lint in CI nor a render check caught this, so it asserts the objects each profile must produce rather than that the render succeeded. Run against the unchanged chart it names the three missing objects in both profiles, and against the pre-fix template it names .Values.operator.enabled too. Verified on a cluster: both configurations present, the CA injected into all ten validator entries, and a ControlPlaneOps naming an absent ControlPlane refused by vcontrolplaneops.simplyblock.io. Co-Authored-By: Claude Opus 5 (1M context) * fix(helm): drop the HAJMCOUNT env var, which reads a value nothing defines daad3507 added storagenode.haJMCount to values.yaml with no value and the HAJMCOUNT env var that reads it. a50e7d1f renamed the key to storagenode.journalManager inside a 255-line rewrite of values.yaml and left the template on the old name, so the reference has been dangling since April. Because the key was empty from the day it was added, the rename changed nothing observable: the variable rendered as the empty string before it and after it. The setting itself survived the rename on a different path. The operator reads journalManager from the cluster deployment config and sends it as ha_jm_count, so nothing is lost by removing the env var; no chart template reads journalManager at all. Both charts carry the same template, and charts/README.md documented the key as a supported value. check-values-references.sh is the test: it reports every .Values path a template reads that values.yaml does not define. haJMCount in both charts was its only finding across 510 defined paths, and it would equally have caught the .Values.operator.enabled guard fixed in the previous commit. Removing the variable means the container sees it unset rather than set to empty. That is unobservable here, since the whole template is gated on storagenode.create, which defaults to false, but the consumer is an external image and this notes the difference. Co-Authored-By: Claude Opus 5 (1M context) * refactor(operator): name the reconcile paths after the members they serve The rename to local and managed moved the API and the predicates and left the names around them on the old vocabulary, so the dispatch read as its own inverse: isManaged called reconcileExternal, and isLocal called reconcileManaged. The behavior was right and only the names were not, which is the shape that survives review and produces a wrong edit later. The two entry points swap, so the local one moves out of the way first. Eight locals and parameters that bind a LocalControlPlane were named managed, and one that binds a ManagedControlPlane was named external. All of it through gopls rename rather than a search and replace. Two of these are shipped API documentation. A field's doc comment becomes its CRD description, so kubectl explain answered the local member with "Managed is a control plane the operator installs" and the managed member with "External is a control plane that already exists". status.endpoint carried the same inversion. Regenerated into config/crd/bases, the chart's crds, the copy the upgrade tool embeds, and dist/install.yaml. The hold message on an object that names neither member said "neither a managed nor a remote control plane" and now names the two that exist. A rename has no failing test to write first: the compiler and the suite are the guard, and both are green. Co-Authored-By: Claude Opus 5 (1M context) --------- Co-authored-by: Claude Opus 5 (1M context) --- .github/workflows/helm_lint.yaml | 14 +- csi-driver/charts/README.md | 1 - .../templates/storage-node-controller.yaml | 2 - .../charts/simplyblock-operator/Chart.yaml | 4 +- ...torage.simplyblock.io_controlplaneops.yaml | 244 +++++ .../storage.simplyblock.io_controlplanes.yaml | 430 +++++++- .../simplyblock-operator/templates/NOTES.txt | 43 +- .../templates/_helpers.tpl | 16 +- .../templates/controlplane_clusterrole.yaml | 25 +- .../templates/controlplane_configmap.yaml | 31 +- .../templates/controlplane_cr.yaml | 68 +- .../templates/controlplane_csi-hostpath.yaml | 8 +- .../templates/controlplane_deploy.yaml | 904 +--------------- .../templates/controlplane_foundationdb.yaml | 439 -------- .../controlplane_foundationdb_exporter.yaml | 111 -- .../templates/controlplane_sa.yaml | 59 -- .../templates/controlplane_storageclass.yaml | 3 +- .../templates/controlplane_svc.yaml | 46 +- .../templates/fluentbit-daemonset.yaml | 2 +- .../templates/metrics-apiserver.yaml | 2 +- .../templates/numa-resource-plugin.yaml | 5 +- .../templates/roles/manager_role.yaml | 20 + .../simplyblock-operator-webhook.yaml | 22 +- .../templates/simplyblock-operator.yaml | 3 +- .../templates/storage-node-controller.yaml | 2 - .../templates/validate-controlplane.yaml | 54 + .../templates/validate-notifications.yaml | 2 +- .../charts/simplyblock-operator/values.yaml | 33 +- helm-charts/scripts/check-rendered-objects.sh | 66 ++ .../scripts/check-values-references.sh | 54 + helm-charts/scripts/sync-from-operator.sh | 4 +- .../api/v1alpha1/controlplane_conversion.go | 76 +- .../v1alpha1/controlplane_conversion_test.go | 109 +- operator/api/v1alpha1/controlplane_types.go | 14 +- operator/api/v1alpha1/hub_roundtrip_test.go | 67 +- operator/api/v1alpha2/controlplane_types.go | 326 +++++- .../api/v1alpha2/controlplaneops_types.go | 262 +++++ .../api/v1alpha2/zz_generated.deepcopy.go | 268 ++++- operator/cmd/main.go | 30 +- ...torage.simplyblock.io_controlplaneops.yaml | 244 +++++ .../storage.simplyblock.io_controlplanes.yaml | 430 +++++++- ...yblock-operator.clusterserviceversion.yaml | 17 +- operator/config/rbac/role.yaml | 20 + operator/config/webhook/manifests.yaml | 20 + operator/dist/install.yaml | 517 +++++++++- .../crd-redesign/design-controlplane.md | 39 +- .../controller/controlplane_controller.go | 164 --- .../controlplane_controller_test.go | 185 ---- .../controllers/cluster/controlplane.go | 57 +- .../controllers/controlplane/apply.go | 178 ++++ .../controlplane/cel_validation_test.go | 246 +++++ .../controllers/controlplane/components.go | 303 ++++++ .../controlplane/components_test.go | 261 +++++ .../controlplane/controlplane_controller.go | 745 ++++++++++++++ .../controlplane_controller_test.go | 445 ++++++++ .../controlplaneops_controller.go | 964 ++++++++++++++++++ .../controlplane/controlplaneops_test.go | 615 +++++++++++ .../controllers/controlplane/datastore.go | 209 ++++ .../internal/controllers/controlplane/doc.go | 62 ++ .../controllers/controlplane/endpoint.go | 218 ++++ .../controllers/controlplane/endpoint_test.go | 120 +++ .../controllers/controlplane/events.go | 96 ++ .../controllers/controlplane/foundationdb.go | 592 +++++++++++ .../controllers/controlplane/graphs.go | 247 +++++ .../controllers/controlplane/helpers_test.go | 237 +++++ .../controllers/controlplane/install_test.go | 383 +++++++ .../controllers/controlplane/managementapi.go | 596 +++++++++++ .../controllers/controlplane/metrics.go | 156 +++ .../controllers/controlplane/names.go | 138 +++ .../controllers/controlplane/podspec.go | 213 ++++ .../controllers/controlplane/probe.go | 164 +++ .../controllers/controlplane/resolver.go | 78 ++ .../controllers/controlplane/resolver_test.go | 104 ++ .../internal/controllers/controlplane/spec.go | 56 + .../controllers/controlplane/suite_test.go | 93 ++ .../controlplane/workloads_test.go | 281 +++++ .../internal/controllers/node/controlplane.go | 57 +- .../controllers/node/workload_controller.go | 5 +- ...torage.simplyblock.io_controlplaneops.yaml | 244 +++++ .../storage.simplyblock.io_controlplanes.yaml | 430 +++++++- .../webhook/controlplaneops_validator.go | 180 ++++ .../webhook/controlplaneops_validator_test.go | 147 +++ 82 files changed, 12156 insertions(+), 2269 deletions(-) create mode 100644 helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_controlplaneops.yaml delete mode 100755 helm-charts/charts/simplyblock-operator/templates/controlplane_foundationdb.yaml delete mode 100644 helm-charts/charts/simplyblock-operator/templates/controlplane_foundationdb_exporter.yaml delete mode 100644 helm-charts/charts/simplyblock-operator/templates/controlplane_sa.yaml create mode 100644 helm-charts/charts/simplyblock-operator/templates/validate-controlplane.yaml create mode 100755 helm-charts/scripts/check-rendered-objects.sh create mode 100755 helm-charts/scripts/check-values-references.sh create mode 100644 operator/api/v1alpha2/controlplaneops_types.go create mode 100644 operator/config/crd/bases/storage.simplyblock.io_controlplaneops.yaml delete mode 100644 operator/internal/controller/controlplane_controller.go delete mode 100644 operator/internal/controller/controlplane_controller_test.go create mode 100644 operator/internal/controllers/controlplane/apply.go create mode 100644 operator/internal/controllers/controlplane/cel_validation_test.go create mode 100644 operator/internal/controllers/controlplane/components.go create mode 100644 operator/internal/controllers/controlplane/components_test.go create mode 100644 operator/internal/controllers/controlplane/controlplane_controller.go create mode 100644 operator/internal/controllers/controlplane/controlplane_controller_test.go create mode 100644 operator/internal/controllers/controlplane/controlplaneops_controller.go create mode 100644 operator/internal/controllers/controlplane/controlplaneops_test.go create mode 100644 operator/internal/controllers/controlplane/datastore.go create mode 100644 operator/internal/controllers/controlplane/doc.go create mode 100644 operator/internal/controllers/controlplane/endpoint.go create mode 100644 operator/internal/controllers/controlplane/endpoint_test.go create mode 100644 operator/internal/controllers/controlplane/events.go create mode 100644 operator/internal/controllers/controlplane/foundationdb.go create mode 100644 operator/internal/controllers/controlplane/graphs.go create mode 100644 operator/internal/controllers/controlplane/helpers_test.go create mode 100644 operator/internal/controllers/controlplane/install_test.go create mode 100644 operator/internal/controllers/controlplane/managementapi.go create mode 100644 operator/internal/controllers/controlplane/metrics.go create mode 100644 operator/internal/controllers/controlplane/names.go create mode 100644 operator/internal/controllers/controlplane/podspec.go create mode 100644 operator/internal/controllers/controlplane/probe.go create mode 100644 operator/internal/controllers/controlplane/resolver.go create mode 100644 operator/internal/controllers/controlplane/resolver_test.go create mode 100644 operator/internal/controllers/controlplane/spec.go create mode 100644 operator/internal/controllers/controlplane/suite_test.go create mode 100644 operator/internal/controllers/controlplane/workloads_test.go create mode 100644 operator/internal/upgrade/crds/manifests/storage.simplyblock.io_controlplaneops.yaml create mode 100644 operator/internal/webhook/controlplaneops_validator.go create mode 100644 operator/internal/webhook/controlplaneops_validator_test.go diff --git a/.github/workflows/helm_lint.yaml b/.github/workflows/helm_lint.yaml index db7eeb467..e77f74467 100644 --- a/.github/workflows/helm_lint.yaml +++ b/.github/workflows/helm_lint.yaml @@ -6,10 +6,14 @@ on: - main paths: - "helm-charts/charts/simplyblock-operator/**" + - "helm-charts/scripts/**" + - "csi-driver/charts/spdk-csi/latest/**" - ".github/workflows/helm_lint.yaml" pull_request: paths: - "helm-charts/charts/simplyblock-operator/**" + - "helm-charts/scripts/**" + - "csi-driver/charts/spdk-csi/latest/**" - ".github/workflows/helm_lint.yaml" jobs: @@ -28,4 +32,12 @@ jobs: run: helm lint helm-charts/charts/simplyblock-operator - name: Template - run: helm template simplyblock-operator helm-charts/charts/simplyblock-operator \ No newline at end of file + run: helm template simplyblock-operator helm-charts/charts/simplyblock-operator + + - name: Required objects + run: helm-charts/scripts/check-rendered-objects.sh + + - name: Value references + run: | + make -C operator yq + helm-charts/scripts/check-values-references.sh diff --git a/csi-driver/charts/README.md b/csi-driver/charts/README.md index b739c98c7..8ac1da097 100644 --- a/csi-driver/charts/README.md +++ b/csi-driver/charts/README.md @@ -164,7 +164,6 @@ The following table lists the configurable parameters of the latest Simplyblock | `storagenode.numDevices` | the number of devices per storage node | `1` | | | `storagenode.numDistribs` | the number of distribs per storage node | `2` | | | `storagenode.isolateCores` | Enable core Isolation | `false` | | -| `storagenode.haJMCount` | the number of ha Journal managers | `` | | | `storagenode.dataNic` | Data interface name | `` | | | `storagenode.pciAllowed` | the list of allowed nvme pcie addresses | `` | | | `storagenode.pciBlocked` | the list of blocked nvme pcie addresses | `` | | diff --git a/csi-driver/charts/spdk-csi/latest/spdk-csi/templates/storage-node-controller.yaml b/csi-driver/charts/spdk-csi/latest/spdk-csi/templates/storage-node-controller.yaml index da364f4d6..019d80a07 100644 --- a/csi-driver/charts/spdk-csi/latest/spdk-csi/templates/storage-node-controller.yaml +++ b/csi-driver/charts/spdk-csi/latest/spdk-csi/templates/storage-node-controller.yaml @@ -62,8 +62,6 @@ spec: value: "{{ .Values.storagenode.spdkProxyImage }}" - name: FORMAT4K value: "{{ .Values.storagenode.format4k }}" - - name: HAJMCOUNT - value: "{{ .Values.storagenode.haJMCount }}" - name: NAMESPACE valueFrom: fieldRef: diff --git a/helm-charts/charts/simplyblock-operator/Chart.yaml b/helm-charts/charts/simplyblock-operator/Chart.yaml index cbb74d66d..3db5ececb 100644 --- a/helm-charts/charts/simplyblock-operator/Chart.yaml +++ b/helm-charts/charts/simplyblock-operator/Chart.yaml @@ -32,9 +32,9 @@ dependencies: - name: prometheus version: "25.18.0" repository: "https://prometheus-community.github.io/helm-charts" - condition: operator.enabled + condition: prometheus.enabled - name: reloader version: 1.3.0 repository: https://stakater.github.io/stakater-charts alias: reloader - condition: operator.enabled + condition: reloader.enabled diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_controlplaneops.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_controlplaneops.yaml new file mode 100644 index 000000000..c0a9a5f6b --- /dev/null +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_controlplaneops.yaml @@ -0,0 +1,244 @@ +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition +metadata: + annotations: + controller-gen.kubebuilder.io/version: v0.21.0 + name: controlplaneops.storage.simplyblock.io +spec: + group: storage.simplyblock.io + names: + kind: ControlPlaneOps + listKind: ControlPlaneOpsList + plural: controlplaneops + shortNames: + - cpops + singular: controlplaneops + scope: Namespaced + versions: + - additionalPrinterColumns: + - jsonPath: .spec.controlPlaneRef + name: ControlPlane + type: string + - jsonPath: .spec.action + name: Action + type: string + - jsonPath: .status.phase + name: Phase + type: string + - jsonPath: .status.step.state + name: Step + type: string + - jsonPath: .status.message + name: Message + priority: 1 + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha2 + schema: + openAPIV3Schema: + description: |- + ControlPlaneOps is a single operation performed against the control plane. It + runs to a terminal phase and stays afterward as the audit record of what was + done, with which parameters, and how it ended. Only one may be active per + control plane at a time, which the entity's status.activeOpsRef enforces. + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: |- + ControlPlaneOpsSpec is one operation to perform against the control plane. + + Everything except spec.abort is frozen once the object is admitted, which is + what makes the status an audit of the request that ran rather than of whatever + the object says now. The parameters are consumed several steps apart: + Preflight reads spec.upgrade.image and Applying writes it, and Draining reads + spec.restart.components before Restarting recycles them. An edit in between + produces an operation that checked one thing and did another. + + The rules are declared here rather than as +k8s:immutable on each field. + controller-gen emits that marker's rules in an order that varies between runs + once a type carries several, and it freezes a block whole; what has to be + frozen is each block's presence together with its contents. + properties: + abort: + description: |- + Abort asks a running operation to stop at its next step and unwind. It is + the one field of this spec an update may change, because it is the one that + is meant to be set after the operation started. Whether an abort is + expressible from the current step is declared by that action's graph rather + than checked here. + type: boolean + action: + description: |- + Action is the operation to perform. Immutable, so that the status describes + the operation that ran. + enum: + - Restart + - Upgrade + - Backup + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + backup: + description: Backup parameterizes action Backup and is ignored by + the others. + properties: + backupName: + description: |- + BackupName is the FoundationDBBackup to create or trigger. Absent uses the + one already configured for the cluster, and fails when there is none and + no name to create. + type: string + blobStore: + description: |- + BlobStore is the destination, in the form the FoundationDBBackup CRD takes + it. The operator copies it through rather than interpreting it, since the + backup is the FoundationDB operator's to perform. + type: string + required: + - blobStore + type: object + controlPlaneRef: + description: |- + ControlPlaneRef names the ControlPlane this operation acts on, in this + object's own namespace. The operation never owns its target, because + deleting the record of an operation must not delete the control plane it + operated on. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + restart: + description: Restart parameterizes action Restart and is ignored by + the others. + properties: + components: + description: |- + Components names the workloads to recycle, from the table in §4.3. Empty + recycles the whole control plane. Naming only components that table marks + non-essential skips the drain, because recycling them interrupts nothing. + items: + type: string + type: array + x-kubernetes-list-type: set + type: object + upgrade: + description: Upgrade parameterizes action Upgrade and is ignored by + the others. + properties: + image: + description: |- + Image is the version to move to. It replaces + ControlPlane.spec.source.managed.image when the operation succeeds, so the + entity keeps describing what is running. + pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ + type: string + required: + - image + type: object + required: + - action + - controlPlaneRef + type: object + x-kubernetes-validations: + - message: 'spec.upgrade is immutable: Preflight checked the image the + operation was admitted with, and Applying writes it several steps + later' + rule: has(self.upgrade) == has(oldSelf.upgrade) && (!has(self.upgrade) + || self.upgrade == oldSelf.upgrade) + - message: 'spec.restart is immutable: the drain is decided from the component + list, so widening it afterward skips a drain the wider list would + have required' + rule: has(self.restart) == has(oldSelf.restart) && (!has(self.restart) + || self.restart == oldSelf.restart) + - message: 'spec.backup is immutable: the destination is what Requesting + created the FoundationDBBackup against' + rule: has(self.backup) == has(oldSelf.backup) && (!has(self.backup) + || self.backup == oldSelf.backup) + status: + description: ControlPlaneOpsStatus is the observed state of one control-plane + operation. + properties: + backupRef: + description: |- + BackupRef names the FoundationDBBackup a Backup run created or triggered. + The operation does not own it, because deleting the record of a backup + must not delete the backup's configuration. + type: string + completedAt: + description: CompletedAt is when it reached a terminal phase. + format: date-time + type: string + message: + description: |- + Message is the reason the phase is what it is: one sentence, replaced as + the operation moves, and never a log. + type: string + observedGeneration: + description: |- + ObservedGeneration is the generation the rest of this status was computed + from, so a stale status can be told from a current one. + format: int64 + type: integer + phase: + description: Phase is the operation's own progress. + enum: + - Pending + - Running + - Succeeded + - Failed + - Aborted + type: string + startedAt: + description: StartedAt is when the operation acquired its target's + lock. + format: date-time + type: string + step: + description: |- + Step is the position of the running action's state machine. It is + persisted before the side effect that step performs. The rule repeats the + ControlPlaneOpsStep enum because a marker cannot reach a field of the + shared snapshot type. + properties: + deadline: + description: |- + Deadline is when that state expires, absent when it has none. It is an + absolute instant, so a state whose deadline passed while the controller + was down restores as already expired. + format: date-time + type: string + state: + description: |- + State is the state the machine was in. Empty means the resource has not + been reconciled yet, and restores to the graph's initial state. + type: string + type: object + x-kubernetes-validations: + - message: unknown step + rule: '!has(self.state) || self.state in [''Draining'',''Restarting'',''Awaiting'',''Preflight'',''Applying'',''Verifying'',''Requesting'']' + type: object + type: object + served: true + storage: true + subresources: + status: {} diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_controlplanes.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_controlplanes.yaml index 1f080d803..140c3da8f 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_controlplanes.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_controlplanes.yaml @@ -11,6 +11,8 @@ spec: kind: ControlPlane listKind: ControlPlaneList plural: controlplanes + shortNames: + - cp singular: controlplane scope: Namespaced versions: @@ -59,9 +61,11 @@ spec: image: description: |- Image is the container image used for all simplyblock control-plane and - storage-node workloads (e.g. quay.io/simplyblock-io/simplyblock:26.2.2). + storage-node workloads (e.g., `quay.io/simplyblock-io/simplyblock:26.2.2`). StorageNodeSet CRs that omit spec.clusterImage inherit this value. - Must reference one of the trusted registries (quay.io/simplyblock-io, docker.io/simplyblock, public.ecr.aws/simply-block); digest pinning (@sha256:...) is recommended. + Must reference one of the trusted registries (`quay.io/simplyblock-io`, + `docker.io/simplyblock`, `public.ecr.aws/simply-block`). Digest pinning + (@sha256:...) is recommended. pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ type: string type: object @@ -77,7 +81,7 @@ spec: message: description: |- Message contains a human-readable explanation of the current phase, - for example the FDB error returned by the health endpoint. + for example, the FDB error returned by the health endpoint. type: string phase: description: |- @@ -94,13 +98,21 @@ spec: subresources: status: {} - additionalPrinterColumns: - - description: Initializing while FDB is not ready; Available once the control plane is operational - jsonPath: .status.phase + - jsonPath: .status.phase name: Phase type: string - - description: Human-readable status detail - jsonPath: .status.message + - jsonPath: .status.step.state + name: Step + type: string + - jsonPath: .status.endpoint + name: Endpoint + type: string + - jsonPath: .status.version + name: Version + type: string + - jsonPath: .status.message name: Message + priority: 1 type: string - jsonPath: .metadata.creationTimestamp name: Age @@ -109,9 +121,11 @@ spec: schema: openAPIV3Schema: description: |- - ControlPlane is a singleton resource (one per namespace, named "simplyblock") - that reflects the readiness of the simplyblock control plane. It is created - automatically by the Helm chart and should not be created or deleted manually. + ControlPlane is the simplyblock control plane for one Kubernetes cluster: + FoundationDB together with the management API, either installed by the + operator or already existing. It is a singleton named `simplyblock`, and it is + the root of the ownership spine: nothing else in this API group reconciles + meaningfully before it reports Available. properties: apiVersion: description: |- @@ -131,48 +145,406 @@ spec: metadata: type: object spec: - description: ControlPlaneSpec holds configuration for the singleton ControlPlane resource. + description: |- + ControlPlaneSpec is the desired state of the simplyblock control plane for one + namespace. properties: source: description: |- - Source says where the control plane comes from. It replaces the top-level - image field of v1alpha1, which conflated the control plane's own image with - the default every StorageNodeSet inherited. + Source selects whether this cluster hosts its control plane or is managed + by one elsewhere. Switching a live deployment between the two is not a + reconfiguration, because the clusters and their volumes live in the + FoundationDB behind the old one, so which mode is chosen is frozen at + creation. What is inside the chosen mode stays editable. + + Both rules are declared on ControlPlaneSource rather than here. See the + type for why. properties: - managed: - description: Managed is the control plane the operator installs and owns. + local: + description: Local is a control plane the operator installs. properties: + foundationDB: + description: FoundationDB sizes the FoundationDB the management API stores its state in. + properties: + replicas: + default: 3 + description: |- + Replicas is the number of coordinators. Three is the smallest count that + survives one loss, which is why it is the default. + format: int32 + minimum: 1 + type: integer + resources: + description: Resources sets requests and limits for the coordinator pods. + properties: + claims: + description: |- + Claims lists the names of resources, defined in spec.resourceClaims, + that are used by this container. + + This field depends on the + DynamicResourceAllocation feature gate. + + This field is immutable. It can only be set for containers. + items: + description: ResourceClaim references one entry in PodSpec.ResourceClaims. + properties: + name: + description: |- + Name must match the name of one entry in pod.spec.resourceClaims of + the Pod where this field is used. It makes that resource available + inside a container. + type: string + request: + description: |- + Request is the name chosen for a request in the referenced claim. + If empty, everything from the claim is made available, otherwise + only the result of this request. + type: string + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + limits: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Limits describes the maximum amount of compute resources allowed. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + requests: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Requests describes the minimum amount of compute resources required. + If Requests is omitted for a container, it defaults to Limits if that is explicitly specified, + otherwise to an implementation-defined value. Requests cannot exceed Limits. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + type: object + storageClassName: + description: |- + StorageClassName is the class the coordinators' volumes are provisioned + from. It cannot be a class this operator provides, because the control + plane has to exist before any simplyblock volume can. + type: string + type: object image: description: |- - Image is the container image used for the simplyblock control-plane - workloads (e.g., quay.io/simplyblock-io/simplyblock:26.2.2). - Must reference one of the trusted registries (`quay.io/simplyblock-io`, `docker.io/simplyblock`, `public.ecr.aws/simply-block`); digest pinning (@sha256:...) is recommended. + Image is the management API and control-plane image. + Must reference one of the trusted registries (`quay.io/simplyblock-io`, + `docker.io/simplyblock`, `public.ecr.aws/simply-block`); digest pinning + (@sha256:...) is recommended. pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ type: string + imagePullPolicy: + default: IfNotPresent + description: ImagePullPolicy controls when that image is pulled. + enum: + - Always + - Never + - IfNotPresent + type: string + nodeSelector: + additionalProperties: + type: string + description: |- + NodeSelector pins every pod the operator installs for the control plane. + It is a selector rather than an affinity term because that is what the + chart it replaces took, and a deployment migrating off the chart has the + value already written down. + type: object + replicas: + default: 2 + description: |- + Replicas is the number of management API instances. Two is what the chart + ships and what the phases assume: a single instance makes Degraded + unreachable for this component and every restart an outage (§5.1). + format: int32 + minimum: 1 + type: integer + resources: + description: Resources sets requests and limits for the management API pods. + properties: + claims: + description: |- + Claims lists the names of resources, defined in spec.resourceClaims, + that are used by this container. + + This field depends on the + DynamicResourceAllocation feature gate. + + This field is immutable. It can only be set for containers. + items: + description: ResourceClaim references one entry in PodSpec.ResourceClaims. + properties: + name: + description: |- + Name must match the name of one entry in pod.spec.resourceClaims of + the Pod where this field is used. It makes that resource available + inside a container. + type: string + request: + description: |- + Request is the name chosen for a request in the referenced claim. + If empty, everything from the claim is made available, otherwise + only the result of this request. + type: string + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + limits: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Limits describes the maximum amount of compute resources allowed. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + requests: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Requests describes the minimum amount of compute resources required. + If Requests is omitted for a container, it defaults to Limits if that is explicitly specified, + otherwise to an implementation-defined value. Requests cannot exceed Limits. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + type: object + tolerations: + description: |- + Tolerations are applied to every pod the operator installs for the control + plane. + items: + description: |- + The pod this Toleration is attached to tolerates any taint that matches + the triple using the matching operator . + properties: + effect: + description: |- + Effect indicates the taint effect to match. Empty means match all taint effects. + When specified, allowed values are NoSchedule, PreferNoSchedule and NoExecute. + type: string + key: + description: |- + Key is the taint key that the toleration applies to. Empty means match all taint keys. + If the key is empty, operator must be Exists; this combination means to match all values and all keys. + type: string + operator: + description: |- + Operator represents a key's relationship to the value. + Valid operators are Exists, Equal, Lt, and Gt. Defaults to Equal. + Exists is equivalent to wildcard for value, so that a pod can + tolerate all taints of a particular category. + Lt and Gt perform numeric comparisons (requires feature gate TaintTolerationComparisonOperators). + type: string + tolerationSeconds: + description: |- + TolerationSeconds represents the period of time the toleration (which must be + of effect NoExecute, otherwise this field is ignored) tolerates the taint. By default, + it is not set, which means tolerate the taint forever (do not evict). Zero and + negative values will be treated as 0 (evict immediately) by the system. + format: int64 + type: integer + value: + description: |- + Value is the taint value the toleration matches to. + If the operator is Exists, the value should be empty, otherwise just a regular string. + type: string + type: object + type: array + required: + - image + type: object + managed: + description: Managed is a control plane that already exists. + properties: + caBundleSecretRef: + description: |- + CABundleSecretRef names a Secret holding the CA certificate the endpoint + is verified against. Absent means the system trust store. + properties: + name: + default: "" + description: |- + Name of the referent. + This field is effectively required, but due to backwards compatibility is + allowed to be empty. Instances of this type with an empty value here are + almost certainly wrong. + More info: https://kubernetes.io/docs/concepts/overview/working-with-objects/names/#names + type: string + type: object + x-kubernetes-map-type: atomic + credentialsSecretRef: + description: |- + CredentialsSecretRef names a Secret in this namespace holding the bearer + token the operator authenticates with. It is a reference rather than a + field because a token in a spec is a token in every `kubectl get -o yaml`. + + Absent means the endpoint is reached without one, which is the in-cluster + case: a control plane the Helm chart installed answers on a ClusterIP + Service in this namespace and does not require a token for the readiness + read. Naming a Secret that does not exist stays an error, because naming + one is a statement that the control plane needs it. + properties: + name: + default: "" + description: |- + Name of the referent. + This field is effectively required, but due to backwards compatibility is + allowed to be empty. Instances of this type with an empty value here are + almost certainly wrong. + More info: https://kubernetes.io/docs/concepts/overview/working-with-objects/names/#names + type: string + type: object + x-kubernetes-map-type: atomic + endpoint: + description: |- + Endpoint is the management API's base URL. It is validated against the + same outbound-URL guard every other outbound endpoint in this group uses, + so a loopback or link-local address is rejected. + pattern: ^https?://[a-zA-Z0-9.-]+(:[0-9]{1,5})?(/.*)?$ + type: string + required: + - endpoint type: object type: object + x-kubernetes-validations: + - message: set exactly one of local or managed + rule: '(has(self.local) ? 1 : 0) + (has(self.managed) ? 1 : 0) == 1' + - message: 'spec.source is immutable: a control plane the operator installed and one it did not are different deployments, and the clusters and their volumes live in the FoundationDB behind the old one' + rule: has(self.local) == has(oldSelf.local) && has(self.managed) == has(oldSelf.managed) + required: + - source type: object status: - description: |- - ControlPlaneStatus reflects the observed readiness of the simplyblock - control plane (FDB + management API). + description: ControlPlaneStatus is the observed state of the control plane. properties: + activeOpsRef: + description: |- + ActiveOpsRef names the ControlPlaneOps currently allowed to act on this + control plane. Empty when none is running. + type: string + components: + description: |- + Components is the per-component readiness the phase is derived from + (§4.3), one entry per workload the managed install applies. It is empty + for a remote control plane, which has no components the operator owns. + Without it a Degraded phase says that something is wrong and not what. + items: + description: |- + ControlPlaneComponentStatus is one workload of a managed control plane and how + much of it is running. The phase is the worst verdict across these and the + readiness probe, and only an essential component at zero ready can make it + Unavailable. + properties: + desired: + description: |- + Desired is how many replicas the component should have. For the component + carrying its own operator it is that resource's own count, because a + FoundationDBCluster reports quorum rather than replicas. + format: int32 + minimum: 0 + type: integer + essential: + description: |- + Essential states whether this component at zero ready makes the control + plane Unavailable rather than Degraded. It is decided by the table in + §4.3 rather than by a user, and it is reported here so that a phase can be + explained without reading the operator's source. + type: boolean + name: + description: Name is the workload's name, as applied. + type: string + ready: + description: Ready is how many of them are. + format: int32 + minimum: 0 + type: integer + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + endpoint: + description: |- + Endpoint is the resolved management API base URL, derived in the local + case and echoed in the managed one. It is what every controller in the + operator reads to reach the control plane, so that one object answers + where it is. + type: string lastChecked: - description: LastChecked is the timestamp of the most recent FDB health probe. + description: LastChecked is when the readiness probe last ran. format: date-time type: string message: description: |- - Message contains a human-readable explanation of the current phase, - for example, the FDB error returned by the health endpoint. + Message is the reason the phase is what it is: one sentence, replaced as + the control plane moves, and never a log. On a failed probe it is the + control plane's own error rather than a paraphrase of it. type: string - phase: + observedGeneration: description: |- - Phase is Initializing while the control plane is not yet healthy, - and Available once the FDB health check passes. + ObservedGeneration is the generation the rest of this status was computed + from, so a stale status can be told from a current one. + format: int64 + type: integer + phase: + description: Phase is the operator's own view of the control plane. enum: - - Initializing + - Installing - Available + - Degraded + - Unavailable + type: string + step: + description: |- + Step is the position of the installation machine within Installing. The + rule repeats the ControlPlaneStep enum because a marker cannot reach a + field of the shared snapshot type. + properties: + deadline: + description: |- + Deadline is when that state expires, absent when it has none. It is an + absolute instant, so a state whose deadline passed while the controller + was down restores as already expired. + format: date-time + type: string + state: + description: |- + State is the state the machine was in. Empty means the resource has not + been reconciled yet, and restores to the graph's initial state. + type: string + type: object + x-kubernetes-validations: + - message: unknown step + rule: '!has(self.state) || self.state in [''ApplyingFoundationDB'',''AwaitingFoundationDB'',''ApplyingDatastore'',''ApplyingAPI'',''AwaitingAPI'']' + version: + description: Version is the version the management API reports. type: string type: object type: object diff --git a/helm-charts/charts/simplyblock-operator/templates/NOTES.txt b/helm-charts/charts/simplyblock-operator/templates/NOTES.txt index 8b75f6546..145ee85d0 100644 --- a/helm-charts/charts/simplyblock-operator/templates/NOTES.txt +++ b/helm-charts/charts/simplyblock-operator/templates/NOTES.txt @@ -1,27 +1,44 @@ -The Simplyblock Operator is getting deployed to your cluster. +The simplyblock operator is being deployed to your cluster. The following components are being installed or updated: -- Simplyblock CSI Driver -{{- if .Values.operator.enabled }} -- Simplyblock Control Plane +- simplyblock CSI Driver +- simplyblock Operator and its CRDs +{{- if eq .Values.deployment.profile "standalone" }} +- simplyblock Control Plane, installed by the operator {{- end }} -{{- if .Values.operator.enabled }} +{{ if eq .Values.deployment.profile "standalone" -}} +This deployment hosts its own control plane. The chart wrote the ControlPlane +object and the operator installs FoundationDB, the object store, and the +management API from it, in that order. The database has to reach quorum before +the management API starts, so the first install takes a few minutes. -The Operator components and its CRDs are installed. +Watch it settle: -ait for the ControlPlane to become Ready: + kubectl --namespace={{ .Release.Namespace }} get controlplane simplyblock -w - kubectl --namespace={{ .Release.Namespace }} wait controlplane simplyblock --for=jsonpath='{.status.phase}'=Ready --timeout=180s +Wait for it to become Available: -Once the ControlPlane is Ready, you can start creating Simplyblock resources. + kubectl --namespace={{ .Release.Namespace }} wait controlplane simplyblock \ + --for=jsonpath='{.status.phase}'=Available --timeout=600s +{{- else }} +This deployment is managed by a control plane at +{{ .Values.controlplane.managed.endpoint }}. Nothing is installed here for it: +the operator resolves that endpoint, probes it, and reports. + +Check that it is reachable: + + kubectl --namespace={{ .Release.Namespace }} get controlplane simplyblock +{{- end }} -Please refer to our documentation to get started: +status.phase is Available once the control plane answers. Degraded means it +answers while something behind it is short, and status.components names which. + +Once it is Available, you can start creating simplyblock resources. Refer to the +documentation to get started: https://docs.simplyblock.io/dev/deployments/kubernetes/k8s-storage-plane/ -{{- else }} -To follow the pods status, please run: +To follow the pods, please run: kubectl --namespace={{ .Release.Namespace }} get pods --watch -{{- end }} diff --git a/helm-charts/charts/simplyblock-operator/templates/_helpers.tpl b/helm-charts/charts/simplyblock-operator/templates/_helpers.tpl index 75448129a..8e97055ab 100644 --- a/helm-charts/charts/simplyblock-operator/templates/_helpers.tpl +++ b/helm-charts/charts/simplyblock-operator/templates/_helpers.tpl @@ -10,10 +10,24 @@ labels: chartVersion: "{{ .Chart.Version }}" {{- end -}} +{{/* +Whether this cluster hosts its own control plane. + +It gates the observability workloads, which store what they collect in the +object store and the document store a hosted control plane brings with it. A +cluster managed from elsewhere has neither, and the control plane that manages it +collects for it. +*/}} +{{- define "simplyblock.hostsControlPlane" -}} +{{- if eq .Values.deployment.profile "standalone" -}} +true +{{- end -}} +{{- end -}} + {{- define "simplyblock.controlPlaneAddr" -}} {{- if .Values.csiConfig.simplybk.ip -}} {{ .Values.csiConfig.simplybk.ip }} -{{- else if .Values.operator.enabled -}} +{{- else -}} http://simplyblock-webappapi.{{ .Release.Namespace }}.svc.cluster.local:5000 {{- end -}} {{- end -}} diff --git a/helm-charts/charts/simplyblock-operator/templates/controlplane_clusterrole.yaml b/helm-charts/charts/simplyblock-operator/templates/controlplane_clusterrole.yaml index 0ef56d903..89000914b 100644 --- a/helm-charts/charts/simplyblock-operator/templates/controlplane_clusterrole.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/controlplane_clusterrole.yaml @@ -1,26 +1,3 @@ -{{- if .Values.operator.enabled }} -apiVersion: rbac.authorization.k8s.io/v1 -kind: ClusterRole -metadata: - name: simplyblock-service-reader -rules: - - apiGroups: [""] - resources: ["services", "pods", "endpoints", "nodes"] - verbs: ["get", "list", "watch"] ---- -apiVersion: rbac.authorization.k8s.io/v1 -kind: ClusterRoleBinding -metadata: - name: simplyblock-service-reader-binding -roleRef: - apiGroup: rbac.authorization.k8s.io - kind: ClusterRole - name: simplyblock-service-reader -subjects: - - kind: ServiceAccount - name: default - namespace: {{ .Release.Namespace }} - --- apiVersion: rbac.authorization.k8s.io/v1 kind: ClusterRole @@ -48,4 +25,4 @@ subjects: - kind: ServiceAccount name: default namespace: {{ .Release.Namespace }} -{{- end }} + diff --git a/helm-charts/charts/simplyblock-operator/templates/controlplane_configmap.yaml b/helm-charts/charts/simplyblock-operator/templates/controlplane_configmap.yaml index e1c1a0b57..c1b93c26a 100644 --- a/helm-charts/charts/simplyblock-operator/templates/controlplane_configmap.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/controlplane_configmap.yaml @@ -1,15 +1,3 @@ -{{- if .Values.operator.enabled }} ---- -apiVersion: v1 -kind: ConfigMap -metadata: - name: simplyblock-config - namespace: {{ .Release.Namespace }} - -data: - LOG_LEVEL: {{ .Values.controlplane.observability.level }} - LOG_DELETION_INTERVAL: {{ .Values.controlplane.observability.deletionInterval }} - {{- if .Values.controlplane.observability.enabled }} --- apiVersion: v1 @@ -166,23 +154,6 @@ data: {{- end }} {{- end }} ---- -apiVersion: v1 -kind: ConfigMap -metadata: - name: simplyblock-objstore-config - labels: - app: simplyblock-thanos - namespace: {{ .Release.Namespace }} -data: - objstore.yml: | - type: S3 - config: - bucket: {{ .Values.controlplane.observability.minio.bucket }} - endpoint: simplyblock-minio:9000 - access_key: {{ .Values.controlplane.observability.minio.accessKey }} - secret_key: {{ .Values.controlplane.observability.minio.secretKey }} - insecure: true {{- if .Values.controlplane.observability.enabled }} --- apiVersion: v1 @@ -1899,4 +1870,4 @@ data: path: /var/lib/grafana/dashboards foldersFromFilesStructure: true {{- end }} -{{- end }} + diff --git a/helm-charts/charts/simplyblock-operator/templates/controlplane_cr.yaml b/helm-charts/charts/simplyblock-operator/templates/controlplane_cr.yaml index 153383b6c..d0b827c6c 100644 --- a/helm-charts/charts/simplyblock-operator/templates/controlplane_cr.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/controlplane_cr.yaml @@ -1,4 +1,3 @@ -{{- if .Values.operator.enabled }} --- # Authored at v1alpha2, which is the storage version a fresh install gets. # @@ -8,17 +7,74 @@ # make the chart itself the one client that forces conversion, on a cluster where # nothing is serving it. # -# The image moved under spec.source.managed with the rename -# (design-controlplane.md §5.1). A v1alpha1 spec.image here would be pruned -# rather than converted, because the apiVersion above declares which schema the -# body is read against. +# This object is the whole of what the chart renders for the control plane. Which +# member of spec.source it names is what deployment.profile selects, and the +# operator does the rest: it installs the control plane for the standalone +# profile, and resolves and probes a remote one for the managed profile. apiVersion: storage.simplyblock.io/v1alpha2 kind: ControlPlane metadata: name: simplyblock namespace: {{ .Release.Namespace }} + annotations: + # Helm must not delete this object, and the reason is the database behind it. + # Deleting a ControlPlane whose source is local deletes the FoundationDB + # holding every cluster definition, node registration, and lvol record, so it + # is not something `helm uninstall` should do as a side effect. Removing a + # control plane stays a deliberate `kubectl delete controlplane`, performed + # while the operator is running so its finalizer can be honored. + # + # It also breaks a deadlock. Helm deletes the operator and this object in one + # uninstall, and with the operator gone first nothing is left to release the + # finalizer: the object cannot go, the uninstall cannot finish, and the + # release sits in `uninstalling` forever. + helm.sh/resource-policy: keep spec: source: - managed: +{{- if eq .Values.deployment.profile "standalone" }} + # This cluster hosts its own control plane, and the operator installs it. + local: image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" + imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" + foundationDB: + # The chart states the redundancy mode and the operator states the + # coordinator count, and the two have to agree: triple redundancy needs + # five coordinators and double needs three. + replicas: {{ if eq (.Values.controlplane.foundationdb.redundancyMode | default "double") "triple" }}5{{ else }}3{{ end }} + {{- if .Values.controlplane.storageclass.name }} + storageClassName: {{ .Values.controlplane.storageclass.name }} + {{- end }} + {{- if .Values.controlplane.nodeSelector.create }} + nodeSelector: + {{ .Values.controlplane.nodeSelector.key }}: {{ .Values.controlplane.nodeSelector.value | quote }} + {{- end }} + {{- if .Values.controlplane.tolerations.create }} + tolerations: + {{- range .Values.controlplane.tolerations.list }} + - operator: {{ .operator | quote }} + {{- if .effect }} + effect: {{ .effect | quote }} + {{- end }} + {{- if .key }} + key: {{ .key | quote }} + {{- end }} + {{- if .value }} + value: {{ .value | quote }} + {{- end }} + {{- end }} + {{- end }} +{{- else }} + # A control plane elsewhere manages this cluster's storage. The operator + # installs nothing here: it resolves the endpoint, probes it, and reports. + managed: + endpoint: {{ required "controlplane.managed.endpoint is required for the managed profile: it is where the control plane that manages this cluster answers" .Values.controlplane.managed.endpoint | quote }} + {{- if .Values.controlplane.managed.credentialsSecretRef }} + credentialsSecretRef: + name: {{ .Values.controlplane.managed.credentialsSecretRef | quote }} + {{- end }} + {{- if .Values.controlplane.managed.caBundleSecretRef }} + caBundleSecretRef: + name: {{ .Values.controlplane.managed.caBundleSecretRef | quote }} + {{- end }} {{- end }} + diff --git a/helm-charts/charts/simplyblock-operator/templates/controlplane_csi-hostpath.yaml b/helm-charts/charts/simplyblock-operator/templates/controlplane_csi-hostpath.yaml index 82653acda..0a7277bb3 100755 --- a/helm-charts/charts/simplyblock-operator/templates/controlplane_csi-hostpath.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/controlplane_csi-hostpath.yaml @@ -1,4 +1,4 @@ -{{- if and .Values.operator.enabled .Values.controlplane.csiHostpathDriver.enabled }} +{{- if .Values.controlplane.csiHostpathDriver.enabled }} apiVersion: storage.k8s.io/v1 kind: CSIDriver metadata: @@ -132,7 +132,7 @@ spec: fieldPath: metadata.name securityContext: # This is necessary only for systems with SELinux, where - # non-privileged sidecar containers cannot access unix domain socket + # non-privileged sidecar containers cannot access Unix domain socket # created by privileged CSI driver container. privileged: true volumeMounts: @@ -145,7 +145,7 @@ spec: - -csi-address=/csi/csi.sock securityContext: # This is necessary only for systems with SELinux, where - # non-privileged sidecar containers cannot access unix domain socket + # non-privileged sidecar containers cannot access Unix domain socket # created by privileged CSI driver container. privileged: true volumeMounts: @@ -160,7 +160,7 @@ spec: - --kubelet-registration-path=/var/lib/kubelet/plugins/csi-hostpath/csi.sock securityContext: # This is necessary only for systems with SELinux, where - # non-privileged sidecar containers cannot access unix domain socket + # non-privileged sidecar containers cannot access Unix domain socket # created by privileged CSI driver container. privileged: true env: diff --git a/helm-charts/charts/simplyblock-operator/templates/controlplane_deploy.yaml b/helm-charts/charts/simplyblock-operator/templates/controlplane_deploy.yaml index c5ad18de2..821e6eb54 100755 --- a/helm-charts/charts/simplyblock-operator/templates/controlplane_deploy.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/controlplane_deploy.yaml @@ -1,905 +1,3 @@ ---- -{{- if .Values.operator.enabled }} -apiVersion: apps/v1 -kind: Deployment -metadata: - name: simplyblock-admin-control - namespace: {{ .Release.Namespace }} - annotations: - reloader.stakater.com/auto: "true" - reloader.stakater.com/configmap: "simplyblock-fdb-cluster-config" -spec: - replicas: 2 - strategy: - type: RollingUpdate - rollingUpdate: - maxSurge: 0 - maxUnavailable: 1 - selector: - matchLabels: - app: simplyblock-admin-control - template: - metadata: - annotations: - log-collector/enabled: "true" - labels: - app: simplyblock-admin-control - spec: - serviceAccountName: simplyblock-sa - hostNetwork: true - dnsPolicy: ClusterFirstWithHostNet - affinity: - podAntiAffinity: - requiredDuringSchedulingIgnoredDuringExecution: - - labelSelector: - matchLabels: - app: simplyblock-admin-control - topologyKey: kubernetes.io/hostname - - {{- if .Values.controlplane.nodeSelector.create }} - nodeSelector: - {{ .Values.controlplane.nodeSelector.key }}: {{ .Values.controlplane.nodeSelector.value }} - {{- end }} - {{- if .Values.controlplane.tolerations.create }} - tolerations: - {{- range .Values.controlplane.tolerations.list }} - - operator: {{ .operator | quote }} - {{- if .effect }} - effect: {{ .effect | quote }} - {{- end }} - {{- if .key }} - key: {{ .key | quote }} - {{- end }} - {{- if .value }} - value: {{ .value | quote }} - {{- end }} - {{- end }} - {{- end }} - containers: - - name: simplyblock-control - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - command: ["/bin/bash", "-c", "trap : TERM INT; sleep infinity & wait"] - env: - {{- include "simplyblock.tlsEnv" . | nindent 8 }} - - name: LVOL_NVMF_PORT_START - value: "{{ .Values.controlplane.ports.lvolNvmfPortStart }}" - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" - - name: K8S_NAMESPACE - valueFrom: - fieldRef: - fieldPath: metadata.namespace -{{- if .Values.controlplane.observability.enabled }} - - name: MONITORING_SECRET - valueFrom: - secretKeyRef: - name: simplyblock-grafana-secrets - key: MONITORING_SECRET -{{- end }} - - name: SIMPLYBLOCK_LOG_LEVEL - valueFrom: - configMapKeyRef: - name: simplyblock-config - key: LOG_LEVEL - volumeMounts: - - name: fdb-cluster-file - mountPath: /etc/foundationdb/fdb.cluster - subPath: fdb.cluster - {{- include "simplyblock.tlsVolumeMount" . | nindent 8 }} - resources: - requests: - cpu: "200m" - memory: "256Mi" - limits: - cpu: "600m" - memory: "1Gi" - volumes: - - name: fdb-cluster-file - configMap: - name: simplyblock-fdb-cluster-config - items: - - key: cluster-file - path: fdb.cluster - {{- include "simplyblock.tlsVolume" (dict "ctx" . "secret" "simplyblock-webappapi-tls") | nindent 6 }} ---- -apiVersion: apps/v1 -kind: Deployment -metadata: - name: simplyblock-webappapi - namespace: {{ .Release.Namespace }} - annotations: - reloader.stakater.com/auto: "true" - reloader.stakater.com/configmap: "simplyblock-fdb-cluster-config" -spec: - replicas: 2 - strategy: - type: RollingUpdate - rollingUpdate: - maxSurge: 0 - maxUnavailable: 1 - selector: - matchLabels: - app: simplyblock-webappapi - template: - metadata: - annotations: - log-collector/enabled: "true" - labels: - app: simplyblock-webappapi - spec: - serviceAccountName: simplyblock-sa - affinity: - podAntiAffinity: - requiredDuringSchedulingIgnoredDuringExecution: - - labelSelector: - matchLabels: - app: simplyblock-webappapi - topologyKey: kubernetes.io/hostname - {{- if .Values.controlplane.nodeSelector.create }} - nodeSelector: - {{ .Values.controlplane.nodeSelector.key }}: {{ .Values.controlplane.nodeSelector.value }} - {{- end }} - {{- if .Values.controlplane.tolerations.create }} - tolerations: - {{- range .Values.controlplane.tolerations.list }} - - operator: {{ .operator | quote }} - {{- if .effect }} - effect: {{ .effect | quote }} - {{- end }} - {{- if .key }} - key: {{ .key | quote }} - {{- end }} - {{- if .value }} - value: {{ .value | quote }} - {{- end }} - {{- end }} - {{- end }} - containers: - - name: webappapi - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - ports: - - containerPort: 5000 - command: ["python3", "simplyblock_web/app.py"] - env: - - name: SIMPLYBLOCK_LOG_LEVEL - valueFrom: - configMapKeyRef: - name: simplyblock-config - key: LOG_LEVEL - - name: LVOL_NVMF_PORT_START - value: "{{ .Values.controlplane.ports.lvolNvmfPortStart }}" - - name: ENABLE_MONITORING - value: "{{ .Values.controlplane.observability.enabled }}" - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" - - name: K8S_NAMESPACE - valueFrom: - fieldRef: - fieldPath: metadata.namespace - {{- include "simplyblock.tlsEnv" . | nindent 8 }} -{{- if .Values.controlplane.observability.enabled }} - - name: MONITORING_SECRET - valueFrom: - secretKeyRef: - name: simplyblock-grafana-secrets - key: MONITORING_SECRET -{{- end }} - - name: FLASK_DEBUG - value: "False" - - name: FLASK_ENV - value: "production" - - name: SB_K8S_ADMIN_SERVICE_ACCOUNTS - {{- $accounts := list (printf "system:serviceaccount:%s:simplyblock-operator" .Release.Namespace) }} - {{- if .Values.controlplane.trustCSIServiceAccounts }} - {{- $accounts = append $accounts (printf "system:serviceaccount:%s:simplyblock-csi-controller-sa" .Release.Namespace) }} - {{- $accounts = append $accounts (printf "system:serviceaccount:%s:simplyblock-csi-node-sa" .Release.Namespace) }} - {{- end }} - value: {{ $accounts | join "," | quote }} - {{- if .Values.prometheus.server.enabled }} - - name: SB_K8S_METRICS_SERVICE_ACCOUNTS - value: {{ printf "system:serviceaccount:%s:simplyblock-prometheus" .Release.Namespace | quote }} - {{- end }} - volumeMounts: - - name: fdb-cluster-file - mountPath: /etc/foundationdb/fdb.cluster - subPath: fdb.cluster - {{- include "simplyblock.tlsVolumeMount" . | nindent 8 }} - resources: - requests: - cpu: "200m" - memory: "512Mi" - limits: - cpu: "500m" - memory: "2Gi" - volumes: - - name: fdb-cluster-file - configMap: - name: simplyblock-fdb-cluster-config - items: - - key: cluster-file - path: fdb.cluster - {{- include "simplyblock.tlsVolume" (dict "ctx" . "secret" "simplyblock-webappapi-tls") | nindent 6 }} ---- -apiVersion: apps/v1 -kind: Deployment -metadata: - name: simplyblock-monitoring - namespace: {{ .Release.Namespace }} - annotations: - reloader.stakater.com/auto: "true" - reloader.stakater.com/configmap: "simplyblock-fdb-cluster-config" -spec: - replicas: 1 - selector: - matchLabels: - app: simplyblock-monitoring - template: - metadata: - annotations: - log-collector/enabled: "true" - labels: - app: simplyblock-monitoring - spec: - serviceAccountName: simplyblock-sa - hostNetwork: true - dnsPolicy: ClusterFirstWithHostNet - {{- if .Values.controlplane.nodeSelector.create }} - nodeSelector: - {{ .Values.controlplane.nodeSelector.key }}: {{ .Values.controlplane.nodeSelector.value }} - {{- end }} - {{- if .Values.controlplane.tolerations.create }} - tolerations: - {{- range .Values.controlplane.tolerations.list }} - - operator: {{ .operator | quote }} - {{- if .effect }} - effect: {{ .effect | quote }} - {{- end }} - {{- if .key }} - key: {{ .key | quote }} - {{- end }} - {{- if .value }} - value: {{ .value | quote }} - {{- end }} - {{- end }} - {{- end }} - containers: - - name: storage-node-monitor - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - command: ["python3", "simplyblock_core/services/storage_node_monitor.py"] - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - env: - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" -{{- with (include "simplyblock.commonContainer" . | fromYaml) }} -{{ toYaml .env | nindent 12 }} - volumeMounts: -{{ toYaml .volumeMounts | nindent 12 }} - resources: -{{ toYaml .resources | nindent 12 }} -{{- end }} - - - name: mgmt-node-monitor - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - command: ["python3", "simplyblock_core/services/mgmt_node_monitor.py"] - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - env: - - name: BACKEND_TYPE - value: "k8s" - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" -{{- with (include "simplyblock.commonContainer" . | fromYaml) }} -{{ toYaml .env | nindent 12 }} - volumeMounts: -{{ toYaml .volumeMounts | nindent 12 }} - resources: -{{ toYaml .resources | nindent 12 }} -{{- end }} - - - name: lvol-stats-collector - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - command: ["python3", "simplyblock_core/services/lvol_stat_collector.py"] - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - env: - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" -{{- with (include "simplyblock.commonContainer" . | fromYaml) }} -{{ toYaml .env | nindent 12 }} - volumeMounts: -{{ toYaml .volumeMounts | nindent 12 }} - resources: -{{ toYaml .resources | nindent 12 }} -{{- end }} - - - name: main-distr-event-collector - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - command: ["python3", "simplyblock_core/services/main_distr_event_collector.py"] - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - env: - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" -{{- with (include "simplyblock.commonContainer" . | fromYaml) }} -{{ toYaml .env | nindent 12 }} - volumeMounts: -{{ toYaml .volumeMounts | nindent 12 }} - resources: -{{ toYaml .resources | nindent 12 }} -{{- end }} - - - name: capacity-and-stats-collector - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - command: ["python3", "simplyblock_core/services/capacity_and_stats_collector.py"] - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - env: - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" -{{- with (include "simplyblock.commonContainer" . | fromYaml) }} -{{ toYaml .env | nindent 12 }} - volumeMounts: -{{ toYaml .volumeMounts | nindent 12 }} - resources: -{{ toYaml .resources | nindent 12 }} -{{- end }} - - - name: capacity-monitor - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - command: ["python3", "simplyblock_core/services/cap_monitor.py"] - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - env: - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" -{{- with (include "simplyblock.commonContainer" . | fromYaml) }} -{{ toYaml .env | nindent 12 }} - volumeMounts: -{{ toYaml .volumeMounts | nindent 12 }} - resources: -{{ toYaml .resources | nindent 12 }} -{{- end }} - - - name: health-check - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - command: ["python3", "simplyblock_core/services/health_check_service.py"] - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - env: - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" -{{- with (include "simplyblock.commonContainer" . | fromYaml) }} -{{ toYaml .env | nindent 12 }} - volumeMounts: -{{ toYaml .volumeMounts | nindent 12 }} - resources: -{{ toYaml .resources | nindent 12 }} -{{- end }} - - - name: device-monitor - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - command: ["python3", "simplyblock_core/services/device_monitor.py"] - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - env: - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" -{{- with (include "simplyblock.commonContainer" . | fromYaml) }} -{{ toYaml .env | nindent 12 }} - volumeMounts: -{{ toYaml .volumeMounts | nindent 12 }} - resources: -{{ toYaml .resources | nindent 12 }} -{{- end }} - - - name: lvol-monitor - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - command: ["python3", "simplyblock_core/services/lvol_monitor.py"] - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - env: - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" -{{- with (include "simplyblock.commonContainer" . | fromYaml) }} -{{ toYaml .env | nindent 12 }} - volumeMounts: -{{ toYaml .volumeMounts | nindent 12 }} - resources: -{{ toYaml .resources | nindent 12 }} -{{- end }} - - - name: snapshot-monitor - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - command: ["python3", "simplyblock_core/services/snapshot_monitor.py"] - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - env: - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" -{{- with (include "simplyblock.commonContainer" . | fromYaml) }} -{{ toYaml .env | nindent 12 }} - volumeMounts: -{{ toYaml .volumeMounts | nindent 12 }} - resources: -{{ toYaml .resources | nindent 12 }} -{{- end }} - volumes: - - name: fdb-cluster-file - configMap: - name: simplyblock-fdb-cluster-config - items: - - key: cluster-file - path: fdb.cluster - {{- include "simplyblock.tlsVolume" (dict "ctx" . "secret" "simplyblock-webappapi-tls") | nindent 8 }} ---- -apiVersion: apps/v1 -kind: Deployment -metadata: - name: simplyblock-tasks - namespace: {{ .Release.Namespace }} - annotations: - reloader.stakater.com/auto: "true" - reloader.stakater.com/configmap: "simplyblock-fdb-cluster-config" -spec: - replicas: 1 - selector: - matchLabels: - app: simplyblock-tasks - template: - metadata: - annotations: - log-collector/enabled: "true" - labels: - app: simplyblock-tasks - spec: - serviceAccountName: simplyblock-sa - hostNetwork: true - dnsPolicy: ClusterFirstWithHostNet - - {{- if .Values.controlplane.nodeSelector.create }} - nodeSelector: - {{ .Values.controlplane.nodeSelector.key }}: {{ .Values.controlplane.nodeSelector.value }} - {{- end }} - {{- if .Values.controlplane.tolerations.create }} - tolerations: - {{- range .Values.controlplane.tolerations.list }} - - operator: {{ .operator | quote }} - {{- if .effect }} - effect: {{ .effect | quote }} - {{- end }} - {{- if .key }} - key: {{ .key | quote }} - {{- end }} - {{- if .value }} - value: {{ .value | quote }} - {{- end }} - {{- end }} - {{- end }} - containers: - - name: tasks-node-add-runner - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - command: ["python3", "simplyblock_core/services/tasks_runner_node_add.py"] - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - env: - - name: LVOL_NVMF_PORT_START - value: "{{ .Values.controlplane.ports.lvolNvmfPortStart }}" - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" -{{- with (include "simplyblock.commonContainer" . | fromYaml) }} -{{ toYaml .env | nindent 12 }} - volumeMounts: -{{ toYaml .volumeMounts | nindent 12 }} - resources: -{{ toYaml .resources | nindent 12 }} -{{- end }} - - - name: tasks-runner-restart - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - command: ["python3", "simplyblock_core/services/tasks_runner_restart.py"] - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - env: - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" -{{- with (include "simplyblock.commonContainer" . | fromYaml) }} -{{ toYaml .env | nindent 12 }} - volumeMounts: -{{ toYaml .volumeMounts | nindent 12 }} - resources: -{{ toYaml .resources | nindent 12 }} -{{- end }} - - - name: tasks-runner-migration - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - command: ["python3", "simplyblock_core/services/tasks_runner_migration.py"] - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - env: - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" -{{- with (include "simplyblock.commonContainer" . | fromYaml) }} -{{ toYaml .env | nindent 12 }} - volumeMounts: -{{ toYaml .volumeMounts | nindent 12 }} - resources: -{{ toYaml .resources | nindent 12 }} -{{- end }} - - - name: tasks-runner-lvol-migration - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - command: ["python3", "simplyblock_core/services/tasks_runner_lvol_migration.py"] - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - env: - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" -{{- with (include "simplyblock.commonContainer" . | fromYaml) }} -{{ toYaml .env | nindent 12 }} - volumeMounts: -{{ toYaml .volumeMounts | nindent 12 }} - resources: -{{ toYaml .resources | nindent 12 }} -{{- end }} - - # Drives the batch (shared-subsystem) migration of namespaced volumes: the - # orchestrator that coordinates the per-volume workers above. Without it a - # batch migration is created and started, then parks in snap_copy forever — - # the workers wait for a phase signal nobody sends. - - name: tasks-runner-batch-migration - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - command: ["python3", "simplyblock_core/services/tasks_runner_batch_migration.py"] - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - env: - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" -{{- with (include "simplyblock.commonContainer" . | fromYaml) }} -{{ toYaml .env | nindent 12 }} - volumeMounts: -{{ toYaml .volumeMounts | nindent 12 }} - resources: -{{ toYaml .resources | nindent 12 }} -{{- end }} - - - name: tasks-runner-failed-migration - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - command: ["python3", "simplyblock_core/services/tasks_runner_failed_migration.py"] - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - env: - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" -{{- with (include "simplyblock.commonContainer" . | fromYaml) }} -{{ toYaml .env | nindent 12 }} - volumeMounts: -{{ toYaml .volumeMounts | nindent 12 }} - resources: -{{ toYaml .resources | nindent 12 }} -{{- end }} - - - name: tasks-runner-cluster-status - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - command: ["python3", "simplyblock_core/services/tasks_cluster_status.py"] - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - env: - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" -{{- with (include "simplyblock.commonContainer" . | fromYaml) }} -{{ toYaml .env | nindent 12 }} - volumeMounts: -{{ toYaml .volumeMounts | nindent 12 }} - resources: -{{ toYaml .resources | nindent 12 }} -{{- end }} - - - name: tasks-runner-new-device-migration - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - command: ["python3", "simplyblock_core/services/tasks_runner_new_dev_migration.py"] - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - env: - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" -{{- with (include "simplyblock.commonContainer" . | fromYaml) }} -{{ toYaml .env | nindent 12 }} - volumeMounts: -{{ toYaml .volumeMounts | nindent 12 }} - resources: -{{ toYaml .resources | nindent 12 }} -{{- end }} - - - name: tasks-runner-port-allow - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - command: ["python3", "simplyblock_core/services/tasks_runner_port_allow.py"] - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - env: - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" -{{- with (include "simplyblock.commonContainer" . | fromYaml) }} -{{ toYaml .env | nindent 12 }} - volumeMounts: -{{ toYaml .volumeMounts | nindent 12 }} - resources: -{{ toYaml .resources | nindent 12 }} -{{- end }} - - - name: tasks-runner-jc-comp-resume - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - command: ["python3", "simplyblock_core/services/tasks_runner_jc_comp.py"] - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - env: - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" -{{- with (include "simplyblock.commonContainer" . | fromYaml) }} -{{ toYaml .env | nindent 12 }} - volumeMounts: -{{ toYaml .volumeMounts | nindent 12 }} - resources: -{{ toYaml .resources | nindent 12 }} -{{- end }} - - - name: tasks-runner-sync-lvol-del - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - command: ["python3", "simplyblock_core/services/tasks_runner_sync_lvol_del.py"] - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - env: - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" -{{- with (include "simplyblock.commonContainer" . | fromYaml) }} -{{ toYaml .env | nindent 12 }} - volumeMounts: -{{ toYaml .volumeMounts | nindent 12 }} - resources: -{{ toYaml .resources | nindent 12 }} -{{- end }} - - - name: tasks-runner-cluster-expand - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - command: ["python3", "simplyblock_core/services/tasks_runner_cluster_expand.py"] - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - env: - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" -{{- with (include "simplyblock.commonContainer" . | fromYaml) }} -{{ toYaml .env | nindent 12 }} - volumeMounts: -{{ toYaml .volumeMounts | nindent 12 }} - resources: -{{ toYaml .resources | nindent 12 }} -{{- end }} - - - name: tasks-runner-node-removal - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - command: ["python3", "simplyblock_core/services/tasks_runner_node_removal.py"] - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - env: - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" -{{- with (include "simplyblock.commonContainer" . | fromYaml) }} -{{ toYaml .env | nindent 12 }} - volumeMounts: -{{ toYaml .volumeMounts | nindent 12 }} - resources: -{{ toYaml .resources | nindent 12 }} -{{- end }} - - - name: tasks-runner-snapshot-replication - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - command: ["python3", "simplyblock_core/services/snapshot_replication.py"] - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - env: - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" -{{- with (include "simplyblock.commonContainer" . | fromYaml) }} -{{ toYaml .env | nindent 12 }} - volumeMounts: -{{ toYaml .volumeMounts | nindent 12 }} - resources: -{{ toYaml .resources | nindent 12 }} -{{- end }} - - - name: tasks-runner-backup - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - command: ["python3", "simplyblock_core/services/tasks_runner_backup.py"] - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - env: - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" -{{- with (include "simplyblock.commonContainer" . | fromYaml) }} -{{ toYaml .env | nindent 12 }} - volumeMounts: -{{ toYaml .volumeMounts | nindent 12 }} - resources: -{{ toYaml .resources | nindent 12 }} -{{- end }} - - - name: tasks-runner-backup-merge - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - command: ["python3", "simplyblock_core/services/tasks_runner_backup_merge.py"] - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - env: - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" -{{- with (include "simplyblock.commonContainer" . | fromYaml) }} -{{ toYaml .env | nindent 12 }} - volumeMounts: -{{ toYaml .volumeMounts | nindent 12 }} - resources: -{{ toYaml .resources | nindent 12 }} -{{- end }} - - - name: tasks-runner-replication-final - image: "{{ .Values.image.simplyblock.repository }}:{{ .Values.image.simplyblock.tag }}" - command: ["python3", "simplyblock_core/services/tasks_runner_replication_final.py"] - imagePullPolicy: "{{ .Values.image.simplyblock.pullPolicy }}" - env: - - name: PROMETHEUS_URL - value: "{{ .Values.prometheus.simplyblock.prometheusURL }}" - - name: PROMETHEUS_PORT - value: "{{ .Values.prometheus.simplyblock.prometheusPORT }}" -{{- with (include "simplyblock.commonContainer" . | fromYaml) }} -{{ toYaml .env | nindent 12 }} - volumeMounts: -{{ toYaml .volumeMounts | nindent 12 }} - resources: -{{ toYaml .resources | nindent 12 }} -{{- end }} - - volumes: - - name: fdb-cluster-file - configMap: - name: simplyblock-fdb-cluster-config - items: - - key: cluster-file - path: fdb.cluster - {{- include "simplyblock.tlsVolume" (dict "ctx" . "secret" "simplyblock-webappapi-tls") | nindent 8 }} - ---- -apiVersion: apps/v1 -kind: StatefulSet -metadata: - name: simplyblock-minio - namespace: {{ .Release.Namespace }} -spec: - serviceName: simplyblock-minio - replicas: 1 - selector: - matchLabels: - app: simplyblock-minio - template: - metadata: - labels: - app: simplyblock-minio - spec: - {{- if .Values.controlplane.nodeSelector.create }} - nodeSelector: - {{ .Values.controlplane.nodeSelector.key }}: {{ .Values.controlplane.nodeSelector.value }} - {{- end }} - containers: - - name: minio - image: "{{ .Values.controlplane.observability.minio.repository }}:{{ .Values.controlplane.observability.minio.tag }}" - imagePullPolicy: {{ .Values.controlplane.observability.minio.pullPolicy }} - args: - - server - - /data - - --console-address=:9001 - env: - - name: MINIO_ROOT_USER - value: {{ .Values.controlplane.observability.minio.accessKey | quote }} - - name: MINIO_ROOT_PASSWORD - value: {{ .Values.controlplane.observability.minio.secretKey | quote }} - ports: - - containerPort: 9000 - name: api - - containerPort: 9001 - name: console - volumeMounts: - - name: minio-data - mountPath: /data - resources: - requests: - cpu: "100m" - memory: "256Mi" - limits: - cpu: "500m" - memory: "1Gi" - - name: bucket-init - image: "{{ .Values.controlplane.observability.minio.mcRepository }}:{{ .Values.controlplane.observability.minio.mcTag }}" - imagePullPolicy: {{ .Values.controlplane.observability.minio.pullPolicy }} - command: - - sh - - -c - - | - until mc alias set local http://localhost:9000 "$MINIO_ROOT_USER" "$MINIO_ROOT_PASSWORD"; do - echo "Waiting for MinIO..."; sleep 3; - done - mc mb --ignore-existing local/{{ .Values.controlplane.observability.minio.bucket }} - sleep infinity - env: - - name: MINIO_ROOT_USER - value: {{ .Values.controlplane.observability.minio.accessKey | quote }} - - name: MINIO_ROOT_PASSWORD - value: {{ .Values.controlplane.observability.minio.secretKey | quote }} - resources: - requests: - cpu: "10m" - memory: "32Mi" - limits: - cpu: "50m" - memory: "64Mi" - volumeClaimTemplates: - - metadata: - name: minio-data - spec: - accessModes: - - ReadWriteOnce - resources: - requests: - storage: {{ .Values.controlplane.observability.minio.storageSize }} - {{- if .Values.controlplane.observability.minio.storageClass }} - storageClassName: {{ .Values.controlplane.observability.minio.storageClass }} - {{- end }} - ---- -apiVersion: v1 -kind: Service -metadata: - name: simplyblock-minio - namespace: {{ .Release.Namespace }} -spec: - selector: - app: simplyblock-minio - ports: - - name: api - port: 9000 - targetPort: 9000 - - name: console - port: 9001 - targetPort: 9001 - --- {{- if .Values.controlplane.observability.enabled }} apiVersion: apps/v1 @@ -1339,4 +437,4 @@ type: Opaque stringData: password: {{ .Values.controlplane.observability.secret }} {{- end }} -{{- end }} + diff --git a/helm-charts/charts/simplyblock-operator/templates/controlplane_foundationdb.yaml b/helm-charts/charts/simplyblock-operator/templates/controlplane_foundationdb.yaml deleted file mode 100755 index 1af33f5de..000000000 --- a/helm-charts/charts/simplyblock-operator/templates/controlplane_foundationdb.yaml +++ /dev/null @@ -1,439 +0,0 @@ -{{- if and .Values.operator.enabled .Values.controlplane.foundationdb.enabled }} ---- -apiVersion: apps/v1 -kind: Deployment -metadata: - name: simplyblock-fdb-controller-manager - labels: - control-plane: simplyblock-fdb-controller-manager - app: simplyblock-fdb-controller-manager - {{- if .Values.tls.mutual_enabled }} - annotations: - reloader.stakater.com/auto: "true" - {{- end }} -spec: - selector: - matchLabels: - app: simplyblock-fdb-controller-manager - replicas: 1 - template: - metadata: - labels: - control-plane: simplyblock-fdb-controller-manager - app: simplyblock-fdb-controller-manager - spec: - securityContext: - runAsUser: 4059 - runAsGroup: 4059 - fsGroup: 4059 - volumes: - - name: tmp - emptyDir: {} - - name: logs - emptyDir: {} - - name: fdb-binaries - emptyDir: {} - {{- if .Values.tls.mutual_enabled }} - - name: tls-fdb - secret: - secretName: simplyblock-foundationdb-tls - {{- end }} - serviceAccountName: simplyblock-fdb-controller-manager - initContainers: - - name: foundationdb-kubernetes-init-7-3 - image: {{ .Values.controlplane.foundationdb.image.monitor.repository }}:{{ .Values.controlplane.foundationdb.image.monitor.tag }} - args: - - "--copy-library" - - "7.3" - - "--copy-binary" - - "fdbcli" - - "--copy-binary" - - "fdbbackup" - - "--copy-binary" - - "fdbrestore" - - "--output-dir" - - "/var/output-files" - - "--mode" - - "init" - volumeMounts: - - name: fdb-binaries - mountPath: /var/output-files - containers: - - command: - - /manager - args: - - "--health-probe-bind-address=:9443" - image: {{ .Values.controlplane.foundationdb.image.operator.repository }}:{{ .Values.controlplane.foundationdb.image.operator.tag }} - name: manager - env: - - name: WATCH_NAMESPACE - valueFrom: - fieldRef: - fieldPath: metadata.namespace - {{- if .Values.tls.mutual_enabled }} - - name: FDB_TLS_CERTIFICATE_FILE - value: /var/fdb/tls/tls.crt - - name: FDB_TLS_KEY_FILE - value: /var/fdb/tls/tls.key - - name: FDB_TLS_CA_FILE - value: /var/fdb/tls/ca.crt - {{- end }} - ports: - - name: metrics - containerPort: 8080 - resources: - limits: - cpu: 500m - memory: 256Mi - requests: - cpu: 500m - memory: 256Mi - securityContext: - readOnlyRootFilesystem: true - allowPrivilegeEscalation: false - privileged: false - volumeMounts: - - name: tmp - mountPath: /tmp - - name: logs - mountPath: /var/log/fdb - - name: fdb-binaries - mountPath: /usr/bin/fdb - {{- if .Values.tls.mutual_enabled }} - - name: tls-fdb - mountPath: /var/fdb/tls - readOnly: true - {{- end }} - terminationGracePeriodSeconds: 10 - -################# ROLE AND ROLE BINDING ############################## ---- -apiVersion: v1 -kind: ServiceAccount -metadata: - name: simplyblock-fdb-controller-manager - ---- -apiVersion: rbac.authorization.k8s.io/v1 -kind: ClusterRole -metadata: - name: simplyblock-fdb-manager-role -rules: -- apiGroups: - - "" - resources: - - configmaps - - events - - persistentvolumeclaims - - pods - - secrets - - services - verbs: - - create - - delete - - get - - list - - patch - - update - - watch -- apiGroups: - - apps - resources: - - deployments - verbs: - - create - - delete - - get - - list - - patch - - update - - watch -- apiGroups: - - apps.foundationdb.org - resources: - - foundationdbbackups - - foundationdbclusters - - foundationdbrestores - verbs: - - create - - delete - - get - - list - - patch - - update - - watch -- apiGroups: - - apps.foundationdb.org - resources: - - foundationdbbackups/status - - foundationdbclusters/status - - foundationdbrestores/status - verbs: - - get - - patch - - update -- apiGroups: - - coordination.k8s.io - resources: - - leases - verbs: - - create - - delete - - get - - list - - patch - - update - - watch ---- -apiVersion: rbac.authorization.k8s.io/v1 -kind: ClusterRole -metadata: - creationTimestamp: null - name: simplyblock-fdb-manager-clusterrole -rules: -- apiGroups: - - "" - resources: - - nodes - verbs: - - get - - list - - watch ---- -apiVersion: rbac.authorization.k8s.io/v1 -kind: RoleBinding -metadata: - creationTimestamp: null - name: simplyblock-fdb-manager-rolebinding -roleRef: - apiGroup: rbac.authorization.k8s.io - kind: ClusterRole - name: simplyblock-fdb-manager-role -subjects: -- kind: ServiceAccount - name: simplyblock-fdb-controller-manager ---- -apiVersion: rbac.authorization.k8s.io/v1 -kind: ClusterRoleBinding -metadata: - creationTimestamp: null - name: simplyblock-fdb-manager-clusterrolebinding -roleRef: - apiGroup: rbac.authorization.k8s.io - kind: ClusterRole - name: simplyblock-fdb-manager-clusterrole -subjects: -- kind: ServiceAccount - name: simplyblock-fdb-controller-manager - namespace: metadata.namespace - -##### service account for FDB cluster pods ################# -# The unified fdb-kubernetes-monitor running in each FDB pod's `foundationdb` -# container writes locality/launcher-environment annotations on its own pod -# (this is how it signals the operator). That requires pod read/patch RBAC, -# which the namespace's `default` SA does not have. -# https://github.com/FoundationDB/fdb-kubernetes-operator/blob/main/docs/manual/technical_design.md ---- -apiVersion: v1 -kind: ServiceAccount -metadata: - name: simplyblock-fdb-cluster-pods ---- -apiVersion: rbac.authorization.k8s.io/v1 -kind: Role -metadata: - name: simplyblock-fdb-cluster-pods -rules: -- apiGroups: [""] - resources: ["pods"] - verbs: ["get", "list", "watch", "update", "patch"] ---- -apiVersion: rbac.authorization.k8s.io/v1 -kind: RoleBinding -metadata: - name: simplyblock-fdb-cluster-pods -roleRef: - apiGroup: rbac.authorization.k8s.io - kind: Role - name: simplyblock-fdb-cluster-pods -subjects: -- kind: ServiceAccount - name: simplyblock-fdb-cluster-pods - -##### cluster file ################# ---- -apiVersion: apps.foundationdb.org/v1beta2 -kind: FoundationDBCluster -metadata: - name: simplyblock-fdb-cluster -spec: - automationOptions: - replacements: - enabled: true - faultDomain: - {{- if .Values.controlplane.foundationdb.multiAZ }} - key: topology.kubernetes.io/zone - {{- else }} - key: kubernetes.io/hostname - {{- end }} - imageType: unified - labels: - filterOnOwnerReference: false - matchLabels: - foundationdb.org/fdb-cluster-name: simplyblock-fdb-cluster - processClassLabels: - - foundationdb.org/fdb-process-class - processGroupIDLabels: - - foundationdb.org/fdb-process-group-id - databaseConfiguration: - redundancy_mode: {{ .Values.controlplane.foundationdb.redundancyMode | default "double" }} - minimumUptimeSecondsForBounce: 60 - processCounts: - {{- if .Values.controlplane.foundationdb.multiAZ }} - cluster_controller: 1 - log: 4 - storage: 4 - stateless: -1 - {{- else }} - cluster_controller: 1 - log: 3 - storage: 3 - stateless: -1 - {{- end }} - processes: - general: - customParameters: - - knob_disable_posix_kernel_aio=1 - podTemplate: - spec: - serviceAccountName: simplyblock-fdb-cluster-pods - {{- if .Values.tls.mutual_enabled }} - volumes: - {{- include "simplyblock.foundationdbCertVolume" . | nindent 10 }} - {{- end }} - {{- if .Values.controlplane.foundationdb.nodeSelector }} - nodeSelector: - {{- toYaml .Values.controlplane.foundationdb.nodeSelector | nindent 12 }} - {{- end }} - containers: - - name: foundationdb - resources: - limits: - cpu: 300m - memory: 512Mi - requests: - cpu: 100m - memory: 256Mi - securityContext: - runAsUser: 0 - {{- include "simplyblock.foundationdbContainerTls" . | nindent 12 }} - initContainers: - - name: foundationdb-kubernetes-init - resources: - limits: - cpu: 100m - memory: 128Mi - requests: - cpu: 100m - memory: 128Mi - securityContext: - runAsUser: 0 - volumeClaimTemplate: - spec: - {{- if .Values.controlplane.storageclass.name }} - storageClassName: {{ .Values.controlplane.storageclass.name }} - {{- end }} - accessModes: - - ReadWriteOnce - resources: - requests: - storage: 10G - - storage: - podTemplate: - spec: - serviceAccountName: simplyblock-fdb-cluster-pods - {{- if .Values.tls.mutual_enabled }} - volumes: - {{- include "simplyblock.foundationdbCertVolume" . | nindent 10 }} - {{- end }} - {{- if .Values.controlplane.foundationdb.nodeSelector }} - nodeSelector: - {{- toYaml .Values.controlplane.foundationdb.nodeSelector | nindent 12 }} - {{- end }} - containers: - - name: foundationdb - resources: - limits: - cpu: 500m - memory: 4Gi - requests: - cpu: 100m - memory: 1Gi - securityContext: - runAsUser: 0 - {{- include "simplyblock.foundationdbContainerTls" . | nindent 12 }} - affinity: - podAntiAffinity: - preferredDuringSchedulingIgnoredDuringExecution: - - weight: 100 - podAffinityTerm: - labelSelector: - matchLabels: - foundationdb.org/fdb-process-class: storage - topologyKey: kubernetes.io/hostname - log: - podTemplate: - spec: - serviceAccountName: simplyblock-fdb-cluster-pods - {{- if .Values.tls.mutual_enabled }} - volumes: - {{- include "simplyblock.foundationdbCertVolume" . | nindent 10 }} - {{- end }} - {{- if .Values.controlplane.foundationdb.nodeSelector }} - nodeSelector: - {{- toYaml .Values.controlplane.foundationdb.nodeSelector | nindent 12 }} - {{- end }} - containers: - - name: foundationdb - resources: - limits: - cpu: 500m - memory: 4Gi - requests: - cpu: 100m - memory: 1Gi - securityContext: - runAsUser: 0 - {{- include "simplyblock.foundationdbContainerTls" . | nindent 12 }} - affinity: - podAntiAffinity: - preferredDuringSchedulingIgnoredDuringExecution: - - weight: 100 - podAffinityTerm: - labelSelector: - matchLabels: - foundationdb.org/fdb-process-class: log - topologyKey: kubernetes.io/hostname - - routing: - defineDNSLocalityFields: true - {{- if or .Values.tls.mutual_enabled .Values.controlplane.foundationdb.image.mainContainer.baseImage }} - mainContainer: - {{- if .Values.tls.mutual_enabled }} - enableTls: true - {{- end }} - {{- if .Values.controlplane.foundationdb.image.mainContainer.baseImage }} - imageConfigs: - - baseImage: {{ .Values.controlplane.foundationdb.image.mainContainer.baseImage }} - {{- end }} - {{- end }} - sidecarContainer: - enableLivenessProbe: true - enableReadinessProbe: false - {{- if .Values.tls.mutual_enabled }} - enableTls: true - {{- end }} - useExplicitListenAddress: true - version: 7.3.63 -{{- end }} diff --git a/helm-charts/charts/simplyblock-operator/templates/controlplane_foundationdb_exporter.yaml b/helm-charts/charts/simplyblock-operator/templates/controlplane_foundationdb_exporter.yaml deleted file mode 100644 index e5eec4ee0..000000000 --- a/helm-charts/charts/simplyblock-operator/templates/controlplane_foundationdb_exporter.yaml +++ /dev/null @@ -1,111 +0,0 @@ -{{- if and .Values.operator.enabled .Values.controlplane.foundationdb.enabled .Values.controlplane.foundationdb.exporter.enabled }} ---- -apiVersion: apps/v1 -kind: Deployment -metadata: - name: simplyblock-fdb-exporter - labels: - app: simplyblock-fdb-exporter - annotations: - # Restart when the cluster file changes (coordinator churn) and, under mTLS, - # when the FDB peer certificate rotates. - reloader.stakater.com/configmap: "simplyblock-fdb-cluster-config" - {{- if .Values.tls.mutual_enabled }} - reloader.stakater.com/auto: "true" - {{- end }} -spec: - replicas: 1 - selector: - matchLabels: - app: simplyblock-fdb-exporter - template: - metadata: - labels: - app: simplyblock-fdb-exporter - spec: - securityContext: - runAsNonRoot: true - runAsUser: 4059 - runAsGroup: 4059 - fsGroup: 4059 - {{- if .Values.controlplane.foundationdb.nodeSelector }} - nodeSelector: - {{- toYaml .Values.controlplane.foundationdb.nodeSelector | nindent 8 }} - {{- end }} - volumes: - - name: tmp - emptyDir: {} - - name: fdb-cluster-file - configMap: - name: simplyblock-fdb-cluster-config - items: - - key: cluster-file - path: fdb.cluster - {{- if .Values.tls.mutual_enabled }} - - name: tls-fdb - secret: - secretName: simplyblock-foundationdb-tls - {{- end }} - containers: - - name: exporter - image: {{ .Values.controlplane.foundationdb.exporter.image.repository }}:{{ .Values.controlplane.foundationdb.exporter.image.tag }} - env: - - name: FDB_CLUSTER_FILE - value: /etc/foundationdb/fdb.cluster - {{- if .Values.tls.mutual_enabled }} - - name: FDB_TLS_CERTIFICATE_FILE - value: /var/fdb/tls/tls.crt - - name: FDB_TLS_KEY_FILE - value: /var/fdb/tls/tls.key - - name: FDB_TLS_CA_FILE - value: /var/fdb/tls/ca.crt - {{- end }} - ports: - - name: metrics - containerPort: 9444 - livenessProbe: - httpGet: - path: /metrics - port: metrics - initialDelaySeconds: 15 - periodSeconds: 30 - readinessProbe: - httpGet: - path: /metrics - port: metrics - initialDelaySeconds: 5 - periodSeconds: 10 - securityContext: - readOnlyRootFilesystem: true - allowPrivilegeEscalation: false - privileged: false - volumeMounts: - - name: tmp - mountPath: /tmp - - name: fdb-cluster-file - mountPath: /etc/foundationdb/fdb.cluster - subPath: fdb.cluster - readOnly: true - {{- if .Values.tls.mutual_enabled }} - - name: tls-fdb - mountPath: /var/fdb/tls - readOnly: true - {{- end }} - resources: - {{- toYaml .Values.controlplane.foundationdb.exporter.resources | nindent 12 }} - terminationGracePeriodSeconds: 10 ---- -apiVersion: v1 -kind: Service -metadata: - name: simplyblock-fdb-exporter - labels: - app: simplyblock-fdb-exporter -spec: - selector: - app: simplyblock-fdb-exporter - ports: - - name: metrics - port: 9444 - targetPort: metrics -{{- end }} diff --git a/helm-charts/charts/simplyblock-operator/templates/controlplane_sa.yaml b/helm-charts/charts/simplyblock-operator/templates/controlplane_sa.yaml deleted file mode 100644 index 00aacb553..000000000 --- a/helm-charts/charts/simplyblock-operator/templates/controlplane_sa.yaml +++ /dev/null @@ -1,59 +0,0 @@ -{{- if .Values.operator.enabled }} -apiVersion: v1 -kind: ServiceAccount -metadata: - name: simplyblock-sa - namespace: {{ .Release.Namespace }} ---- -apiVersion: rbac.authorization.k8s.io/v1 -kind: ClusterRole -metadata: - name: simplyblock-role -rules: - - apiGroups: [""] - resources: ["configmaps"] - verbs: ["get", "list", "watch","patch", "update"] - - apiGroups: ["", "apps"] - resources: ["pods", "deployments", "statefulsets", "daemonsets"] - verbs: ["get", "list", "watch", "patch", "update"] - - apiGroups: [""] - resources: ["pods/log"] - verbs: ["get", "list"] - - apiGroups: [""] - resources: ["pods/exec"] - verbs: ["create", "get", "list", "watch", "patch", "update"] - - apiGroups: [""] - resources: ["nodes"] - verbs: ["get", "list", "watch", "patch", "update"] - {{- if .Values.operator.openShiftCluster }} - - apiGroups: ["machineconfiguration.openshift.io"] - resources: ["machineconfigpools"] - verbs: ["get", "list", "watch", "patch", "update"] - {{- end }} - - apiGroups: ["mongodbcommunity.mongodb.com"] - resources: ["mongodbcommunity"] - verbs: ["get", "list", "watch", "patch", "update"] - - apiGroups: ["storage.simplyblock.io"] - resources: ["pools/status", "lvols/status", "storageclusters/status", "storagenodesets/status", "devices/status", "tasks/status", "storagebackups/status"] - verbs: ["get", "patch", "update"] - - apiGroups: ["storage.simplyblock.io"] - resources: ["pools", "lvols", "storageclusters", "storagenodesets", "devices", "tasks", "storagebackups"] - verbs: ["get","list" ,"patch", "update", "watch"] - - apiGroups: ["authentication.k8s.io"] - resources: ["tokenreviews"] - verbs: ["create"] - ---- -apiVersion: rbac.authorization.k8s.io/v1 -kind: ClusterRoleBinding -metadata: - name: simplyblock-binding -subjects: - - kind: ServiceAccount - name: simplyblock-sa - namespace: {{ .Release.Namespace }} -roleRef: - kind: ClusterRole - name: simplyblock-role - apiGroup: rbac.authorization.k8s.io -{{- end }} diff --git a/helm-charts/charts/simplyblock-operator/templates/controlplane_storageclass.yaml b/helm-charts/charts/simplyblock-operator/templates/controlplane_storageclass.yaml index 1b0400934..6486b445a 100644 --- a/helm-charts/charts/simplyblock-operator/templates/controlplane_storageclass.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/controlplane_storageclass.yaml @@ -1,4 +1,3 @@ -{{- if .Values.operator.enabled }} --- apiVersion: storage.k8s.io/v1 kind: StorageClass @@ -22,4 +21,4 @@ allowedTopologies: - {{ . }} {{- end }} {{- end }} -{{- end }} + diff --git a/helm-charts/charts/simplyblock-operator/templates/controlplane_svc.yaml b/helm-charts/charts/simplyblock-operator/templates/controlplane_svc.yaml index 96df96909..c8d334071 100644 --- a/helm-charts/charts/simplyblock-operator/templates/controlplane_svc.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/controlplane_svc.yaml @@ -1,47 +1,3 @@ -{{- if .Values.operator.enabled }} -apiVersion: v1 -kind: Service -metadata: - name: simplyblock-webappapi - namespace: {{ .Release.Namespace }} - {{- if and .Values.tls.enabled (eq .Values.tls.provider "openshift") }} - annotations: - service.beta.openshift.io/serving-cert-secret-name: simplyblock-webappapi-tls - {{- end }} -spec: - selector: - app: simplyblock-webappapi - ports: - - name: http - port: 5000 - targetPort: 5000 - -{{- if and .Values.tls.enabled (eq .Values.tls.provider "cert-manager") }} ---- -apiVersion: cert-manager.io/v1 -kind: Certificate -metadata: - name: simplyblock-webappapi - namespace: {{ .Release.Namespace }} -spec: - commonName: simplyblock-webappapi - secretName: simplyblock-webappapi-tls - issuerRef: - kind: ClusterIssuer - name: simplyblock-certificate-authority-issuer - {{- if .Values.tls.mutual_enabled }} - usages: - - digital signature - - key encipherment - - server auth - - client auth - {{- end }} - dnsNames: - - simplyblock-webappapi - - simplyblock-webappapi.{{ .Release.Namespace }} - - simplyblock-webappapi.{{ .Release.Namespace }}.svc - - simplyblock-webappapi.{{ .Release.Namespace }}.svc.cluster.local -{{- end }} --- {{- if .Values.controlplane.observability.enabled }} @@ -97,4 +53,4 @@ spec: port: 3000 targetPort: 3000 {{- end }} -{{- end }} + diff --git a/helm-charts/charts/simplyblock-operator/templates/fluentbit-daemonset.yaml b/helm-charts/charts/simplyblock-operator/templates/fluentbit-daemonset.yaml index 7ee7170d1..449d73336 100644 --- a/helm-charts/charts/simplyblock-operator/templates/fluentbit-daemonset.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/fluentbit-daemonset.yaml @@ -1,4 +1,4 @@ -{{- if and .Values.operator.enabled .Values.controlplane.observability.enabled -}} +{{- if .Values.controlplane.observability.enabled -}} --- apiVersion: v1 kind: ServiceAccount diff --git a/helm-charts/charts/simplyblock-operator/templates/metrics-apiserver.yaml b/helm-charts/charts/simplyblock-operator/templates/metrics-apiserver.yaml index 7c6953cb9..4ba6b557a 100644 --- a/helm-charts/charts/simplyblock-operator/templates/metrics-apiserver.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/metrics-apiserver.yaml @@ -1,4 +1,4 @@ -{{- if and .Values.operator.enabled .Values.metricsAPI.enabled }} +{{- if .Values.metricsAPI.enabled }} # The cluster-facing half of the aggregated metrics API: the Service the # kube-apiserver proxies to, the APIService that registers # metrics.simplyblock.io, and the two bindings that let the operator delegate diff --git a/helm-charts/charts/simplyblock-operator/templates/numa-resource-plugin.yaml b/helm-charts/charts/simplyblock-operator/templates/numa-resource-plugin.yaml index 6851bbb54..99e7d0856 100644 --- a/helm-charts/charts/simplyblock-operator/templates/numa-resource-plugin.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/numa-resource-plugin.yaml @@ -1,6 +1,5 @@ {{- $storagenodeEnabled := and .Values.storagenode.create .Values.storagenode.enableCpuTopology .Values.storagenode.enableDevicePlugin -}} -{{- $operatorEnabled := .Values.operator.enabled -}} -{{- if or $storagenodeEnabled $operatorEnabled -}} +{{- if or $storagenodeEnabled true -}} --- apiVersion: v1 kind: ServiceAccount @@ -53,7 +52,7 @@ spec: - {{ $ds.nodeSelector.value | quote }} {{- end }} {{- end }} - {{- else if $operatorEnabled }} + {{- else if true }} affinity: nodeAffinity: requiredDuringSchedulingIgnoredDuringExecution: diff --git a/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml b/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml index e335a7885..1368864a6 100644 --- a/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml @@ -121,6 +121,7 @@ rules: - apps resources: - daemonsets + - deployments - statefulsets verbs: - create @@ -130,6 +131,19 @@ rules: - patch - update - watch +- apiGroups: + - apps.foundationdb.org + resources: + - foundationdbbackups + - foundationdbclusters + verbs: + - create + - delete + - get + - list + - patch + - update + - watch - apiGroups: - authentication.k8s.io resources: @@ -187,6 +201,8 @@ rules: resources: - clusterrolebindings - clusterroles + - rolebindings + - roles verbs: - bind - create @@ -237,6 +253,7 @@ rules: - backupimports - backuppolicies - backuprestores + - controlplaneops - controlplanes - operatorops - replicationops @@ -270,6 +287,8 @@ rules: - backupimports/finalizers - backuppolicies/finalizers - backuprestores/finalizers + - controlplaneops/finalizers + - controlplanes/finalizers - operatorops/finalizers - replicationops/finalizers - replicationpairs/finalizers @@ -295,6 +314,7 @@ rules: - backuppolicies/status - backuprestores/status - clusterdeploymentconfigs/status + - controlplaneops/status - controlplanes/status - operatorops/status - replicationops/status diff --git a/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml b/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml index ed983e159..d0072700b 100644 --- a/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml @@ -1,4 +1,3 @@ -{{- if .Values.operator.enabled }} apiVersion: v1 kind: Service metadata: @@ -87,6 +86,26 @@ webhooks: resources: - persistentvolumeclaims sideEffects: None +- admissionReviewVersions: + - v1 + clientConfig: + service: + name: simplyblock-operator-webhook-service + namespace: {{ .Release.Namespace }} + path: /validate-storage-simplyblock-io-v1alpha2-controlplaneops + failurePolicy: Fail + name: vcontrolplaneops.simplyblock.io + rules: + - apiGroups: + - storage.simplyblock.io + apiVersions: + - v1alpha2 + operations: + - CREATE + - DELETE + resources: + - controlplaneops + sideEffects: None - admissionReviewVersions: - v1 clientConfig: @@ -244,4 +263,3 @@ webhooks: resources: - storagepools sideEffects: None -{{- end }} diff --git a/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator.yaml b/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator.yaml index 2788a966d..54a8742a0 100644 --- a/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator.yaml @@ -1,4 +1,3 @@ -{{- if .Values.operator.enabled }} --- apiVersion: apps/v1 kind: Deployment @@ -517,4 +516,4 @@ subjects: - kind: ServiceAccount name: simplyblock-operator namespace: {{ .Release.Namespace }} -{{- end }} + diff --git a/helm-charts/charts/simplyblock-operator/templates/storage-node-controller.yaml b/helm-charts/charts/simplyblock-operator/templates/storage-node-controller.yaml index da364f4d6..019d80a07 100644 --- a/helm-charts/charts/simplyblock-operator/templates/storage-node-controller.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/storage-node-controller.yaml @@ -62,8 +62,6 @@ spec: value: "{{ .Values.storagenode.spdkProxyImage }}" - name: FORMAT4K value: "{{ .Values.storagenode.format4k }}" - - name: HAJMCOUNT - value: "{{ .Values.storagenode.haJMCount }}" - name: NAMESPACE valueFrom: fieldRef: diff --git a/helm-charts/charts/simplyblock-operator/templates/validate-controlplane.yaml b/helm-charts/charts/simplyblock-operator/templates/validate-controlplane.yaml new file mode 100644 index 000000000..e6dbfb9ba --- /dev/null +++ b/helm-charts/charts/simplyblock-operator/templates/validate-controlplane.yaml @@ -0,0 +1,54 @@ +{{/* +The two ways a control-plane profile can be asked for something it cannot do. + +Both fail at render time rather than at apply time, because each produces a +deployment that comes up and is quietly wrong: one serving plaintext where TLS +was asked for, and one pointed at no control plane at all. +*/}} + + +{{- if not (has .Values.deployment.profile (list "standalone" "managed")) -}} +{{- fail (printf "deployment.profile is %q: it is one of standalone (this cluster hosts its own control plane) or managed (a control plane elsewhere manages this cluster's storage)." .Values.deployment.profile) -}} +{{- end -}} + +{{/* +TLS has no field on the ControlPlane spec. design-controlplane.md §5.1 settles +which certificate issuer is detected, not whether a deployment wants one, so an +operator-installed control plane comes up in plaintext whatever tls.enabled says. + +That was survivable while the chart could still install the control plane +itself. It cannot any more, so a TLS deployment has nowhere to go and the honest +answer is to refuse rather than to install something quieter than what was asked +for. +*/}} +{{- if and (eq .Values.deployment.profile "standalone") .Values.tls.enabled -}} +{{- fail "deployment.profile=standalone cannot serve TLS: the ControlPlane spec has no field for it, so the operator's install would bring the control plane up in plaintext. Until the spec can express it, a TLS deployment needs a control plane this cluster does not host (deployment.profile=managed), or an operator release that carries the field." -}} +{{- end -}} + +{{/* +The managed profile needs somewhere to point. The CR template requires the +endpoint too, and this exists so the message names the value rather than the +field of a CRD the reader has not seen. +*/}} +{{- if and (eq .Values.deployment.profile "managed") (not .Values.controlplane.managed.endpoint) -}} +{{- fail "deployment.profile=managed needs controlplane.managed.endpoint: it is where the control plane that manages this cluster answers, and nothing is installed locally to fall back to." -}} +{{- end -}} + +{{/* +An in-place upgrade from a release that installed the control plane itself would +prune it. + +Helm deletes what the old release manifest held and the new one does not, and +before profiles the chart's manifest held the FoundationDBCluster. That database +is every cluster definition, node registration, and lvol record, and none of +those objects carries a keep policy, so the upgrade would take the fleet's state +with it and the operator would then build an empty one. + +The lookup finds nothing on a fresh install, so only an upgrade can trip this. +*/}} +{{- $running := lookup "apps/v1" "Deployment" .Release.Namespace "simplyblock-webappapi" -}} +{{- if and $running (index $running.metadata.annotations "meta.helm.sh/release-name") -}} +{{- fail "this namespace holds a control plane installed by an earlier release of this chart, and upgrading in place would delete it, FoundationDB included. Annotate what it owns so Helm keeps it first:\n\n for k in deployment/simplyblock-webappapi deployment/simplyblock-tasks deployment/simplyblock-monitoring deployment/simplyblock-admin-control deployment/simplyblock-fdb-controller-manager deployment/simplyblock-fdb-exporter statefulset/simplyblock-minio foundationdbcluster/simplyblock-fdb-cluster; do\n kubectl -n annotate $k helm.sh/resource-policy=keep --overwrite\n done\n\nthen upgrade. The operator adopts them in place." -}} +{{- end -}} + + diff --git a/helm-charts/charts/simplyblock-operator/templates/validate-notifications.yaml b/helm-charts/charts/simplyblock-operator/templates/validate-notifications.yaml index 58ae35822..cd029eb34 100644 --- a/helm-charts/charts/simplyblock-operator/templates/validate-notifications.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/validate-notifications.yaml @@ -5,7 +5,7 @@ ConfigMap that notifies nobody, which only shows up the next time an alert fires. The release fails here instead, next to the chart's other value guards in validate-tls.yaml. */}} -{{- if and .Values.operator.enabled .Values.controlplane.observability.enabled -}} +{{- if .Values.controlplane.observability.enabled -}} {{- with .Values.controlplane.observability.grafana -}} {{- if .contactPoint -}} {{- fail "controlplane.observability.grafana.contactPoint has been removed. Move the Slack webhook URL to controlplane.observability.grafana.notifications.slack.url and set notifications.slack.enabled to true." -}} diff --git a/helm-charts/charts/simplyblock-operator/values.yaml b/helm-charts/charts/simplyblock-operator/values.yaml index f261515c1..df153308c 100644 --- a/helm-charts/charts/simplyblock-operator/values.yaml +++ b/helm-charts/charts/simplyblock-operator/values.yaml @@ -147,7 +147,33 @@ storagenode: value: +deployment: + # One of: standalone, managed. + # + # standalone hosts its own control plane. The operator installs FoundationDB, + # the object store, and the management API from the ControlPlane object this + # chart writes. + # + # managed registers this cluster's storage with a control plane elsewhere, + # named by controlplane.managed below. Nothing is installed here. + # + # Observability under controlplane is rendered for standalone only: it stores + # what it collects in the object store a hosted control plane brings. + profile: standalone + controlplane: + # Where the control plane is, for deployment.profile: managed. + managed: + # Base URL of the management API. Required for that profile. A loopback or + # link-local address is rejected. + endpoint: "" + # Secret in this namespace holding the bearer token, under `token` or + # `secret`. Empty sends no token. + credentialsSecretRef: "" + # Secret holding the CA the endpoint is verified against, under `ca.crt` or + # `tls.crt`. Empty uses the system trust store. Named and unusable fails. + caBundleSecretRef: "" + # Accept a pod's service-account token from the CSI driver's two accounts # instead of requiring the static cluster secret. This is the control # plane's half of the switch: set it together with @@ -368,7 +394,6 @@ controlplane: memory: 128Mi operator: - enabled: true openShiftCluster: true # The CSI link: the CSI node and controller pods dial the operator and hold the @@ -498,6 +523,9 @@ opensearch: enabled: false reloader: + # Rolls the workloads mounting the FoundationDB cluster file when coordinators + # move. Off means they pick the change up on their next restart instead. + enabled: true nameOverride: simplyblock-reloader fullnameOverride: simplyblock-reloader reloader: @@ -509,6 +537,9 @@ reloader: tolerations: [] prometheus: + # Where the control plane pushes what it measures. A managed deployment + # reports to the control plane administering it, so it can turn this off. + enabled: true simplyblock: prometheusURL: simplyblock-prometheus prometheusPORT: 9090 diff --git a/helm-charts/scripts/check-rendered-objects.sh b/helm-charts/scripts/check-rendered-objects.sh new file mode 100755 index 000000000..9dca24879 --- /dev/null +++ b/helm-charts/scripts/check-rendered-objects.sh @@ -0,0 +1,66 @@ +#!/usr/bin/env bash +# Asserts that a rendered chart contains the objects the operator needs to work. +# +# `helm template` exits zero for a template that renders to nothing, so a guard +# reading a value that no longer exists drops its objects silently. This lists +# what each profile must produce and fails when one of them is missing. +set -uo pipefail + +CHART="$(cd "$(dirname "${BASH_SOURCE[0]}")/../charts/simplyblock-operator" && pwd)" +fail=0 + +# Objects every profile renders, as `Kind/name`. +COMMON=( + "Deployment/simplyblock-operator" + "Service/simplyblock-operator-webhook-service" + "MutatingWebhookConfiguration/simplyblock-operator-mutating-webhook-configuration" + "ValidatingWebhookConfiguration/simplyblock-operator-validating-webhook-configuration" + "ServiceAccount/simplyblock-operator" + "ControlPlane/simplyblock" +) + +# objects prints `Kind/name` for every document in a rendered manifest. The name +# is taken from the first ` name:` under a top-level `metadata:`, so a name +# nested in a pod template or a webhook entry is not mistaken for the object's. +objects() { + awk ' + /^---/ { kind=""; name=""; depth=0; next } + /^kind: / { kind=$2; next } + /^metadata:/ { depth=1; next } + depth==1 && /^ name: / { depth=0; if (kind != "") print kind "/" $2; next } + /^[a-zA-Z]/ { depth=0 } + ' +} + +check() { + local profile="$1" + shift + local -a required=("$@") + local out present missing=0 + + out="$(helm template sb "$CHART" --namespace simplyblock \ + --set deployment.profile="$profile" \ + --set controlplane.managed.endpoint=https://cp.example.com 2>/dev/null)" + if [ -z "$out" ]; then + echo " ${profile}: RENDER FAILED" + fail=1 + return + fi + + present="$(printf '%s\n' "$out" | objects)" + for want in "${required[@]}"; do + if ! printf '%s\n' "$present" | grep -qxF "$want"; then + echo " ${profile}: MISSING ${want}" + missing=1 + fail=1 + fi + done + if [ "$missing" -eq 0 ]; then + echo " ${profile}: all ${#required[@]} required objects present" + fi +} + +check standalone "${COMMON[@]}" +check managed "${COMMON[@]}" + +exit "$fail" diff --git a/helm-charts/scripts/check-values-references.sh b/helm-charts/scripts/check-values-references.sh new file mode 100755 index 000000000..44883ece0 --- /dev/null +++ b/helm-charts/scripts/check-values-references.sh @@ -0,0 +1,54 @@ +#!/usr/bin/env bash +# Reports every `.Values.x.y` a template reads that values.yaml does not define. +# +# Helm renders a missing value as the empty string, so a reference that no longer +# resolves is silent: a `{{- if }}` on one drops its whole block, and a value +# written into an env var reaches the container empty. Both have shipped. +set -uo pipefail + +REPO="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)" +YQ="${YQ:-$REPO/.bin/yq}" +command -v "$YQ" >/dev/null 2>&1 || YQ=yq + +# The charts whose templates are hand-written, each as ``. +CHARTS=( + "$REPO/helm-charts/charts/simplyblock-operator" + "$REPO/csi-driver/charts/spdk-csi/latest/spdk-csi" +) + +fail=0 + +for chart in "${CHARTS[@]}"; do + name="${chart#"$REPO"/}" + if [ ! -f "$chart/values.yaml" ]; then + echo " ${name}: no values.yaml" + fail=1 + continue + fi + + # Every path values.yaml defines, including the intermediate maps, so that a + # reference to a block rather than to a leaf resolves too. + defined="$("$YQ" -r '[.. | path | join(".")] | .[]' "$chart/values.yaml" 2>/dev/null | grep -v '^$' | sort -u)" + + # Every path the templates read. A trailing dot belongs to the template syntax + # around the reference rather than to the path. + referenced="$(grep -rhoE '\.Values\.[A-Za-z_][A-Za-z0-9_.]*' "$chart/templates" 2>/dev/null | + sed -e 's/^\.Values\.//' -e 's/\.$//' | sort -u)" + + missing=0 + while read -r path; do + [ -z "$path" ] && continue + if ! printf '%s\n' "$defined" | grep -qxF "$path"; then + echo " ${name}: .Values.${path} is read by a template and not defined" + grep -rn "Values\.${path}" "$chart/templates" | sed 's|^| |' + missing=1 + fail=1 + fi + done <<<"$referenced" + + if [ "$missing" -eq 0 ]; then + echo " ${name}: every referenced value is defined" + fi +done + +exit "$fail" diff --git a/helm-charts/scripts/sync-from-operator.sh b/helm-charts/scripts/sync-from-operator.sh index b7202b185..9c6047164 100755 --- a/helm-charts/scripts/sync-from-operator.sh +++ b/helm-charts/scripts/sync-from-operator.sh @@ -95,8 +95,9 @@ echo "==> Syncing webhook manifests..." # kustomize namePrefix (see utils.WebhookServiceName / *ConfigurationName). KUSTOMIZE_BIN="${KUSTOMIZE:-$HELM_CHARTS_DIR/../.bin/kustomize}" command -v "$KUSTOMIZE_BIN" >/dev/null 2>&1 || KUSTOMIZE_BIN="kustomize" +# The admission configurations are unconditional: the operator serves every +# handler it registers, so the chart routes to all of them. { - echo "{{- if .Values.operator.enabled }}" "$KUSTOMIZE_BIN" build "$WEBHOOK_SRC" | sed \ -e 's|name: webhook-service|name: simplyblock-operator-webhook-service|' \ -e 's|name: mutating-webhook-configuration|name: simplyblock-operator-mutating-webhook-configuration|' \ @@ -104,7 +105,6 @@ command -v "$KUSTOMIZE_BIN" >/dev/null 2>&1 || KUSTOMIZE_BIN="kustomize" -e 's|namespace: system|namespace: {{ .Release.Namespace }}|' \ -e 's|control-plane: controller-manager|control-plane: simplyblock-operator|' \ -e 's|app.kubernetes.io/name: simplyblock-operator|app: simplyblock-operator|' - echo "{{- end }}" } > "$WEBHOOK_DST" echo " copied: webhook.yaml (from kustomize build config/webhook)" diff --git a/operator/api/v1alpha1/controlplane_conversion.go b/operator/api/v1alpha1/controlplane_conversion.go index 3bd347d97..d6f9620ca 100644 --- a/operator/api/v1alpha1/controlplane_conversion.go +++ b/operator/api/v1alpha1/controlplane_conversion.go @@ -1,13 +1,19 @@ // Conversion of ControlPlane between this version and the v1alpha2 hub. // // Two properties move (design-property-renames.md §2.4 and §2.5): the top-level -// image regroups under spec.source.managed, and the readiness phase Ready becomes -// Available. Everything else is carried across unchanged. +// image regroups under spec.source.local, and the readiness phase Ready becomes +// Available. Everything else this version declares is carried across unchanged, +// and everything the hub adds (the step, the endpoint, the version, the +// components, the lock, and the observed generation) has no v1alpha1 spelling +// and is dropped on the way down, which is what a spoke that predates a field +// does with it. // // The conversion never fails. A phase value in neither table is passed through as // written, because a conversion webhook is the wrong place to reject an object: // each version's Enum marker already refuses what that version does not accept, // and a conversion that errors makes the object unreadable rather than invalid. +// Every value either Enum declares is in a table, so the pass-through covers only +// a hand-edited object or one written by a version that has not shipped. package v1alpha1 @@ -17,19 +23,39 @@ import ( "github.com/simplyblock/simplyblock-operator/api/v1alpha2" ) -// controlPlanePhaseToHub maps this version's readiness phases onto the hub's. -// Only the renamed value appears; anything absent is passed through. +// controlPlanePhaseToHub maps this version's two readiness phases onto the hub's. +// Both are mapped: neither v1alpha1 spelling is a value v1alpha2's Enum admits, +// so passing one through unchanged produces an object the API server rejects on +// write and the new controller cannot interpret. var controlPlanePhaseToHub = map[string]string{ - "Ready": "Available", + "Initializing": "Installing", + "Ready": "Available", } -// controlPlanePhaseFromHub is the inverse of controlPlanePhaseToHub, built from -// it so the two cannot drift into disagreeing about a value. -var controlPlanePhaseFromHub = invertStringMap(controlPlanePhaseToHub) +// controlPlanePhaseFromHub maps the hub's four phases onto this version's two. +// +// It is written out rather than derived by inverting the table above, because +// the hub says more than this version can hold and the mapping is therefore not +// one to one. Degraded and Unavailable have no v1alpha1 spelling, and each lands +// on the value that tells a v1alpha1 reader the same thing: a Degraded control +// plane answers requests, so it reads as Ready, and an Unavailable one does not, +// so it reads as Initializing. +// +// That makes hub to spoke to hub lossy for those two, which is the direction +// this version cannot help. Spoke to hub to spoke is lossless, and that is the +// trip the API server performs on every read of a stored v1alpha1 object. +var controlPlanePhaseFromHub = map[string]string{ + "Installing": "Initializing", + "Available": "Ready", + "Degraded": "Ready", + "Unavailable": "Initializing", +} -// invertStringMap returns m with its keys and values exchanged. It is used to -// derive a conversion's downward value table from its upward one, so that adding -// a renamed value means editing one map rather than remembering to edit two. +// invertStringMap returns m with its keys and values exchanged. It derives a +// conversion's downward value table from its upward one, so that a renamed value +// means editing one map rather than remembering to edit two. It suits a rename +// and not this file's phases, where the hub holds more values than the spoke and +// the two directions are therefore written out separately. func invertStringMap(m map[string]string) map[string]string { inverted := make(map[string]string, len(m)) for from, to := range m { @@ -54,18 +80,16 @@ func (src *ControlPlane) ConvertTo(dstRaw conversion.Hub) error { dst.ObjectMeta = src.ObjectMeta - // An unset image leaves spec.source absent rather than allocating an empty - // managed block. A conversion that writes an empty parent hands the user a - // value they never set, and elsewhere in this migration such a parent is - // immutable once written and cannot then be corrected. - dst.Spec.Source = nil - if src.Spec.Image != "" { - dst.Spec.Source = &v1alpha2.ControlPlaneSource{ - Managed: &v1alpha2.ManagedControlPlane{Image: src.Spec.Image}, - } + // Every v1alpha1 ControlPlane is one the chart installed, so the local member + // is the one it converts into, and it is set even when the image is empty. + // The hub requires exactly one member (design-controlplane.md §3.2), and the + // reconciler classifies an object by which member is set. + dst.Spec.Source = v1alpha2.ControlPlaneSource{ + Local: &v1alpha2.LocalControlPlane{Image: src.Spec.Image}, } - dst.Status.Phase = mapOrPassThrough(controlPlanePhaseToHub, src.Status.Phase) + dst.Status.Phase = v1alpha2.ControlPlanePhase( + mapOrPassThrough(controlPlanePhaseToHub, src.Status.Phase)) dst.Status.Message = src.Status.Message dst.Status.LastChecked = src.Status.LastChecked @@ -78,12 +102,16 @@ func (dst *ControlPlane) ConvertFrom(srcRaw conversion.Hub) error { dst.ObjectMeta = src.ObjectMeta + // A remote control plane has no image, which is what this version's only + // spec field holds. It converts down to an empty one rather than to an + // error: the object still has to be readable at v1alpha1, and what a reader + // there loses is a field that never applied to it. dst.Spec.Image = "" - if src.Spec.Source != nil && src.Spec.Source.Managed != nil { - dst.Spec.Image = src.Spec.Source.Managed.Image + if managed := src.Spec.Source.Local; managed != nil { + dst.Spec.Image = managed.Image } - dst.Status.Phase = mapOrPassThrough(controlPlanePhaseFromHub, src.Status.Phase) + dst.Status.Phase = mapOrPassThrough(controlPlanePhaseFromHub, string(src.Status.Phase)) dst.Status.Message = src.Status.Message dst.Status.LastChecked = src.Status.LastChecked diff --git a/operator/api/v1alpha1/controlplane_conversion_test.go b/operator/api/v1alpha1/controlplane_conversion_test.go index b3e7cb01d..3f3bb64a5 100644 --- a/operator/api/v1alpha1/controlplane_conversion_test.go +++ b/operator/api/v1alpha1/controlplane_conversion_test.go @@ -1,7 +1,7 @@ // Tests for the ControlPlane conversion between v1alpha1 and the v1alpha2 hub. // // Two properties are converted (design-property-renames.md §2.4 and §2.5): the -// top-level image regroups under spec.source.managed, and the readiness phase +// top-level image regroups under spec.source.local, and the readiness phase // Ready becomes Available. Both directions are tested, because a conversion that // renames going up and copies going down corrupts on the first // `kubectl get -o yaml | kubectl apply -f -` and a one-way test cannot see it. @@ -37,22 +37,21 @@ func TestControlPlaneConvertToRegroupsImage(t *testing.T) { t.Fatalf("ConvertTo: %v", err) } - if dst.Spec.Source == nil || dst.Spec.Source.Managed == nil { - t.Fatalf("spec.source.managed is absent, want the image regrouped under it") + if dst.Spec.Source.Local == nil { + t.Fatalf("spec.source.local is absent, want the image regrouped under it") } - if got := dst.Spec.Source.Managed.Image; got != testImage { - t.Errorf("spec.source.managed.image = %q, want %q", got, testImage) + if got := dst.Spec.Source.Local.Image; got != testImage { + t.Errorf("spec.source.local.image = %q, want %q", got, testImage) } if dst.Name != "simplyblock" || dst.Namespace != "sb" { t.Errorf("object meta not carried: %q/%q", dst.Namespace, dst.Name) } } -// An absent image must leave spec.source absent rather than allocating an empty -// struct. A conversion that writes an empty parent hands the user a value they -// did not set, which for the immutable groups elsewhere in this migration cannot -// then be corrected. -func TestControlPlaneConvertToLeavesSourceAbsentWhenImageEmpty(t *testing.T) { +// Every v1alpha1 ControlPlane is one the chart installed, so an absent image +// still converts into a local source. The hub requires exactly one member of +// spec.source, and the reconciler classifies an object by which member is set. +func TestControlPlaneConvertToStillNamesManagedWhenImageEmpty(t *testing.T) { src := &ControlPlane{Spec: ControlPlaneSpec{Image: ""}} var dst v1alpha2.ControlPlane @@ -60,8 +59,14 @@ func TestControlPlaneConvertToLeavesSourceAbsentWhenImageEmpty(t *testing.T) { t.Fatalf("ConvertTo: %v", err) } - if dst.Spec.Source != nil { - t.Errorf("spec.source = %+v, want nil for an unset image", dst.Spec.Source) + if dst.Spec.Source.Local == nil { + t.Fatalf("spec.source.local is absent, want a managed source with an empty image") + } + if got := dst.Spec.Source.Local.Image; got != "" { + t.Errorf("spec.source.local.image = %q, want empty", got) + } + if dst.Spec.Source.Managed != nil { + t.Errorf("spec.source.managed = %+v, want nil", dst.Spec.Source.Managed) } } @@ -72,7 +77,7 @@ func TestControlPlaneConvertToRenamesReadyPhase(t *testing.T) { want string }{ {"ready becomes available", "Ready", "Available"}, - {"initializing is unchanged", "Initializing", "Initializing"}, + {"initializing becomes installing", "Initializing", "Installing"}, {"empty is unchanged", "", ""}, {"an unrecognized value passes through", "Wedged", "Wedged"}, } { @@ -83,7 +88,7 @@ func TestControlPlaneConvertToRenamesReadyPhase(t *testing.T) { if err := src.ConvertTo(&dst); err != nil { t.Fatalf("ConvertTo: %v", err) } - if got := dst.Status.Phase; got != tc.want { + if got := string(dst.Status.Phase); got != tc.want { t.Errorf("status.phase = %q, want %q", got, tc.want) } }) @@ -94,8 +99,8 @@ func TestControlPlaneConvertFromUngroupsImage(t *testing.T) { src := &v1alpha2.ControlPlane{ ObjectMeta: metav1.ObjectMeta{Name: "simplyblock", Namespace: "sb"}, Spec: v1alpha2.ControlPlaneSpec{ - Source: &v1alpha2.ControlPlaneSource{ - Managed: &v1alpha2.ManagedControlPlane{Image: testImage}, + Source: v1alpha2.ControlPlaneSource{ + Local: &v1alpha2.LocalControlPlane{Image: testImage}, }, }, } @@ -113,17 +118,71 @@ func TestControlPlaneConvertFromUngroupsImage(t *testing.T) { } } -func TestControlPlaneConvertFromRenamesAvailablePhase(t *testing.T) { - src := &v1alpha2.ControlPlane{ - Status: v1alpha2.ControlPlaneStatus{Phase: "Available"}, - } +// Every hub phase has to land on a value this version's Enum admits, which is +// Initializing or Ready. Degraded and Unavailable have no v1alpha1 spelling, so +// they map onto the one that describes the same thing to a v1alpha1 reader: +// Degraded answers requests, and Unavailable does not. +func TestControlPlaneConvertFromMapsEveryHubPhase(t *testing.T) { + for _, tc := range []struct { + from v1alpha2.ControlPlanePhase + want string + }{ + {"Available", "Ready"}, + {"Degraded", "Ready"}, + {"Installing", "Initializing"}, + {"Unavailable", "Initializing"}, + {"", ""}, + } { + t.Run(string(tc.from), func(t *testing.T) { + src := &v1alpha2.ControlPlane{ + Status: v1alpha2.ControlPlaneStatus{Phase: tc.from}, + } - var dst ControlPlane - if err := dst.ConvertFrom(src); err != nil { - t.Fatalf("ConvertFrom: %v", err) + var dst ControlPlane + if err := dst.ConvertFrom(src); err != nil { + t.Fatalf("ConvertFrom: %v", err) + } + if got := dst.Status.Phase; got != tc.want { + t.Errorf("status.phase = %q, want %q", got, tc.want) + } + }) } - if got := dst.Status.Phase; got != "Ready" { - t.Errorf("status.phase = %q, want %q", got, "Ready") +} + +// Both directions have to produce a value the target version's Enum admits. A +// phase outside it is an object the API server rejects on write and the new +// controller cannot interpret, which is how a migration stops on a control +// plane that was merely starting up. +func TestControlPlanePhasesStayInsideBothEnums(t *testing.T) { + hubPhases := map[string]bool{ + "Installing": true, "Available": true, "Degraded": true, "Unavailable": true, + } + spokePhases := map[string]bool{"Initializing": true, "Ready": true} + + for phase := range spokePhases { + var hub v1alpha2.ControlPlane + src := &ControlPlane{Status: ControlPlaneStatus{Phase: phase}} + if err := src.ConvertTo(&hub); err != nil { + t.Fatalf("ConvertTo: %v", err) + } + if !hubPhases[string(hub.Status.Phase)] { + t.Errorf("%q converts up to %q, which v1alpha2's Enum does not admit", + phase, hub.Status.Phase) + } + } + + for phase := range hubPhases { + var spoke ControlPlane + src := &v1alpha2.ControlPlane{ + Status: v1alpha2.ControlPlaneStatus{Phase: v1alpha2.ControlPlanePhase(phase)}, + } + if err := spoke.ConvertFrom(src); err != nil { + t.Fatalf("ConvertFrom: %v", err) + } + if !spokePhases[spoke.Status.Phase] { + t.Errorf("%q converts down to %q, which v1alpha1's Enum does not admit", + phase, spoke.Status.Phase) + } } } diff --git a/operator/api/v1alpha1/controlplane_types.go b/operator/api/v1alpha1/controlplane_types.go index e5a061c6b..209a73f45 100644 --- a/operator/api/v1alpha1/controlplane_types.go +++ b/operator/api/v1alpha1/controlplane_types.go @@ -24,9 +24,11 @@ import ( // created by the Helm chart. type ControlPlaneSpec struct { // Image is the container image used for all simplyblock control-plane and - // storage-node workloads (e.g. quay.io/simplyblock-io/simplyblock:26.2.2). + // storage-node workloads (e.g., `quay.io/simplyblock-io/simplyblock:26.2.2`). // StorageNodeSet CRs that omit spec.clusterImage inherit this value. - // Must reference one of the trusted registries (quay.io/simplyblock-io, docker.io/simplyblock, public.ecr.aws/simply-block); digest pinning (@sha256:...) is recommended. + // Must reference one of the trusted registries (`quay.io/simplyblock-io`, + // `docker.io/simplyblock`, `public.ecr.aws/simply-block`). Digest pinning + // (@sha256:...) is recommended. // +optional // +kubebuilder:validation:Pattern=`^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$` Image string `json:"image,omitempty"` @@ -41,7 +43,7 @@ type ControlPlaneStatus struct { Phase string `json:"phase,omitempty"` // Message contains a human-readable explanation of the current phase, - // for example the FDB error returned by the health endpoint. + // for example, the FDB error returned by the health endpoint. Message string `json:"message,omitempty"` // LastChecked is the timestamp of the most recent FDB health probe. @@ -50,7 +52,11 @@ type ControlPlaneStatus struct { // +kubebuilder:object:root=true // +kubebuilder:subresource:status -// +kubebuilder:resource:scope=Namespaced +// The short name is repeated from v1alpha2 rather than declared only there. +// shortNames is a resource-level property and both versions of a kind feed one +// CRD, so two +kubebuilder:resource markers that disagree leave controller-gen +// picking one per run and the generated CRD differing from itself. +// +kubebuilder:resource:scope=Namespaced,shortName=cp // +kubebuilder:printcolumn:name="Phase",type="string",JSONPath=".status.phase",description="Initializing while FDB is not ready; Ready once the control plane is operational" // +kubebuilder:printcolumn:name="Message",type="string",JSONPath=".status.message",description="Human-readable status detail" // +kubebuilder:printcolumn:name="Age",type="date",JSONPath=".metadata.creationTimestamp" diff --git a/operator/api/v1alpha1/hub_roundtrip_test.go b/operator/api/v1alpha1/hub_roundtrip_test.go index 3481cf559..24571507e 100644 --- a/operator/api/v1alpha1/hub_roundtrip_test.go +++ b/operator/api/v1alpha1/hub_roundtrip_test.go @@ -30,8 +30,8 @@ func TestControlPlaneRoundTripsFromTheHub(t *testing.T) { hub := &v1alpha2.ControlPlane{ ObjectMeta: metav1.ObjectMeta{Name: "simplyblock", Namespace: "sb"}, Spec: v1alpha2.ControlPlaneSpec{ - Source: &v1alpha2.ControlPlaneSource{ - Managed: &v1alpha2.ManagedControlPlane{Image: testImage}, + Source: v1alpha2.ControlPlaneSource{ + Local: &v1alpha2.LocalControlPlane{Image: testImage}, }, }, Status: v1alpha2.ControlPlaneStatus{ @@ -55,18 +55,18 @@ func TestControlPlaneRoundTripsFromTheHub(t *testing.T) { } } -// A managed block with no image normalizes to an absent source, and that is the -// intended behavior rather than an accident worth stashing. +// A managed block with no image survives the round trip as a managed block with +// no image, which is what makes the trip lossless for every hub object v1alpha1 +// can hold. // -// v1alpha1 states the image as one optional string, so it cannot express "a -// source block was present but empty" — the information does not exist in the -// stored shape. Since an empty managed block selects nothing and configures -// nothing, dropping it loses no meaning, and the alternative would be an -// annotation carrying the fact that a user wrote two empty braces. -func TestControlPlaneEmptyManagedBlockNormalizesAway(t *testing.T) { +// v1alpha1 states the image as one optional string, so an empty managed block +// and an absent source are the same stored shape. The upward conversion resolves +// that ambiguity toward managed, because every object stored at v1alpha1 is one +// the chart installed and the hub requires exactly one member of spec.source. +func TestControlPlaneEmptyManagedBlockSurvivesTheRoundTrip(t *testing.T) { hub := &v1alpha2.ControlPlane{ Spec: v1alpha2.ControlPlaneSpec{ - Source: &v1alpha2.ControlPlaneSource{Managed: &v1alpha2.ManagedControlPlane{}}, + Source: v1alpha2.ControlPlaneSource{Local: &v1alpha2.LocalControlPlane{}}, }, } @@ -79,8 +79,49 @@ func TestControlPlaneEmptyManagedBlockNormalizesAway(t *testing.T) { t.Fatalf("ConvertTo: %v", err) } - if back.Spec.Source != nil { - t.Errorf("spec.source = %+v, want nil", back.Spec.Source) + if diff := cmp.Diff(hub, &back); diff != "" { + t.Errorf("storing and reading back changed the object (-written +read):\n%s", diff) + } +} + +// A control plane this cluster is managed by has no v1alpha1 spelling, so +// storing one at that version and reading it back loses which control plane the +// object meant. The conversion is lossy here rather than failing: the object +// stays readable and comes back describing a local control plane with no image. +// +// Nothing stores one at v1alpha1 in practice, because the mode did not exist +// before the storage version moved to v1alpha2. What this pins is the behavior +// if something ever does. +func TestControlPlaneExternalSourceDoesNotSurviveV1Alpha1(t *testing.T) { + hub := &v1alpha2.ControlPlane{ + Spec: v1alpha2.ControlPlaneSpec{ + Source: v1alpha2.ControlPlaneSource{ + Managed: &v1alpha2.ManagedControlPlane{ + Endpoint: "https://sb-control.example.com:5000", + CredentialsSecretRef: &corev1.LocalObjectReference{Name: "cp-token"}, + }, + }, + }, + } + + var stored ControlPlane + if err := stored.ConvertFrom(hub); err != nil { + t.Fatalf("ConvertFrom: %v", err) + } + if stored.Spec.Image != "" { + t.Errorf("spec.image = %q, want empty for a remote control plane", stored.Spec.Image) + } + + var back v1alpha2.ControlPlane + if err := stored.ConvertTo(&back); err != nil { + t.Fatalf("ConvertTo: %v", err) + } + if back.Spec.Source.Managed != nil { + t.Errorf("spec.source.managed = %+v, want nil: v1alpha1 cannot hold it", + back.Spec.Source.Managed) + } + if back.Spec.Source.Local == nil { + t.Error("spec.source.local is absent, want the upward conversion's managed default") } } diff --git a/operator/api/v1alpha2/controlplane_types.go b/operator/api/v1alpha2/controlplane_types.go index 095ce7ba5..1001db524 100644 --- a/operator/api/v1alpha2/controlplane_types.go +++ b/operator/api/v1alpha2/controlplane_types.go @@ -1,59 +1,302 @@ -// ControlPlane in the shape design-controlplane.md settles: the image the -// operator installs moves under spec.source.managed, and the readiness phase says -// Available rather than Ready. +// ControlPlane in the shape design-controlplane.md settles: FoundationDB +// together with the management API, either installed here by the operator or +// owned by a remote control plane, expressed as one object. // -// Only the renamed and regrouped properties are here. spec.source.external -// (design-controlplane.md §5.2) and the rest of ManagedControlPlane are additive -// design work rather than renames, so they are not in this package yet; the -// source struct exists because the regrouping needs somewhere to put the image. +// Two things distinguish it from the registered kind it replaces. spec.source +// says which control plane the object means, so naming a remote one is a field +// rather than the SIMPLYBLOCK_WEBAPI_BASE_URL environment variable (§5.2). And status.endpoint publishes the resolved base URL, which the +// operator's control-plane clients resolve per call, so one object answers where +// the control plane is and a change to it reaches every caller without a +// Deployment rollout (§3.3). +// +// design-controlplane.md Appendix A is the specification for this file. package v1alpha2 import ( + corev1 "k8s.io/api/core/v1" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + + "github.com/simplyblock/atlas/statemachine" ) -// ManagedControlPlane is a control plane the operator installs and owns. -type ManagedControlPlane struct { - // Image is the container image used for the simplyblock control-plane - // workloads (e.g., quay.io/simplyblock-io/simplyblock:26.2.2). - // Must reference one of the trusted registries (`quay.io/simplyblock-io`, `docker.io/simplyblock`, `public.ecr.aws/simply-block`); digest pinning (@sha256:...) is recommended. +// ControlPlanePhase is where the operator has got to with this control plane. +// +kubebuilder:validation:Enum=Installing;Available;Degraded;Unavailable +type ControlPlanePhase string + +const ( + // ControlPlanePhaseInstalling is a control plane that has not worked yet. + ControlPlanePhaseInstalling ControlPlanePhase = "Installing" + // ControlPlanePhaseAvailable is one whose readiness probe passes and whose + // workload pods are settled. + ControlPlanePhaseAvailable ControlPlanePhase = "Available" + // ControlPlanePhaseDegraded is one whose readiness probe passes while a + // management API or FoundationDB pod is restarting. It answers every + // request, so nothing holds on it and it exists to be read by a person. A + // control plane the operator does not manage never reaches it, because the + // operator owns no pods there to watch. + ControlPlanePhaseDegraded ControlPlanePhase = "Degraded" + // ControlPlanePhaseUnavailable is one whose readiness probe fails: it + // worked and stopped, which is a different situation from one that never + // started, and it is what downstream controllers hold on. The word claims + // only what the probe observed, since a failing probe cannot tell a crashed + // process from a wedged one or from a partition. + ControlPlanePhaseUnavailable ControlPlanePhase = "Unavailable" +) + +// ControlPlaneStep is one step of the installation path. There is one graph +// rather than a MultiConfig, because an entity has no spec.action to key one on. +// +kubebuilder:validation:Enum=ApplyingFoundationDB;AwaitingFoundationDB;ApplyingDatastore;ApplyingAPI;AwaitingAPI +type ControlPlaneStep string + +const ( + // ControlPlaneStepApplyingFoundationDB applies the FoundationDBCluster, the + // accounts and roles its pods need, and the FoundationDB operator itself + // where the Kubernetes cluster does not already run one. + ControlPlaneStepApplyingFoundationDB ControlPlaneStep = "ApplyingFoundationDB" + // ControlPlaneStepAwaitingFoundationDB holds until that cluster reports + // itself available, which is the step that can take the longest and the one + // whose deadline matters. + ControlPlaneStepAwaitingFoundationDB ControlPlaneStep = "AwaitingFoundationDB" + // ControlPlaneStepApplyingDatastore applies the object store the control + // plane keeps its long-term data in. + ControlPlaneStepApplyingDatastore ControlPlaneStep = "ApplyingDatastore" + // ControlPlaneStepApplyingAPI applies the management API's workload, the + // services beside it, its account and configuration, and its Service. + ControlPlaneStepApplyingAPI ControlPlaneStep = "ApplyingAPI" + // ControlPlaneStepAwaitingAPI holds until the readiness probe succeeds, + // which is what ends the installation. + ControlPlaneStepAwaitingAPI ControlPlaneStep = "AwaitingAPI" +) + +// FoundationDBSpec is the sizing of the FoundationDB the operator installs. +type FoundationDBSpec struct { + // Replicas is the number of coordinators. Three is the smallest count that + // survives one loss, which is why it is the default. + // +kubebuilder:validation:Minimum=1 + // +kubebuilder:default=3 // +optional + Replicas *int32 `json:"replicas,omitempty"` + + // StorageClassName is the class the coordinators' volumes are provisioned + // from. It cannot be a class this operator provides, because the control + // plane has to exist before any simplyblock volume can. + // +optional + StorageClassName string `json:"storageClassName,omitempty"` + + // Resources sets requests and limits for the coordinator pods. + // +optional + Resources corev1.ResourceRequirements `json:"resources,omitempty"` +} + +// LocalControlPlane is a control plane this cluster hosts, installed and owned +// by the operator. Its objects carry a controller reference to the ControlPlane, +// so the ownership spine starts at a real edge rather than at a Helm release. +type LocalControlPlane struct { + // Image is the management API and control-plane image. + // Must reference one of the trusted registries (`quay.io/simplyblock-io`, + // `docker.io/simplyblock`, `public.ecr.aws/simply-block`); digest pinning + // (@sha256:...) is recommended. // +kubebuilder:validation:Pattern=`^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$` - Image string `json:"image,omitempty"` + // +kubebuilder:validation:Required + Image string `json:"image"` + + // ImagePullPolicy controls when that image is pulled. + // +kubebuilder:validation:Enum=Always;Never;IfNotPresent + // +kubebuilder:default=IfNotPresent + // +optional + ImagePullPolicy corev1.PullPolicy `json:"imagePullPolicy,omitempty"` + + // FoundationDB sizes the FoundationDB the management API stores its state in. + // +optional + FoundationDB *FoundationDBSpec `json:"foundationDB,omitempty"` + + // Replicas is the number of management API instances. Two is what the chart + // ships and what the phases assume: a single instance makes Degraded + // unreachable for this component and every restart an outage (§5.1). + // +kubebuilder:validation:Minimum=1 + // +kubebuilder:default=2 + // +optional + Replicas *int32 `json:"replicas,omitempty"` + + // Resources sets requests and limits for the management API pods. + // +optional + Resources corev1.ResourceRequirements `json:"resources,omitempty"` + + // Tolerations are applied to every pod the operator installs for the control + // plane. + // +optional + Tolerations []corev1.Toleration `json:"tolerations,omitempty"` + + // NodeSelector pins every pod the operator installs for the control plane. + // It is a selector rather than an affinity term because that is what the + // chart it replaces took, and a deployment migrating off the chart has the + // value already written down. + // +optional + NodeSelector map[string]string `json:"nodeSelector,omitempty"` } -// ControlPlaneSource selects where the control plane comes from. Today, it names -// only a managed one, which is what the operator installs. +// ManagedControlPlane is a control plane somewhere else, which this cluster's +// storage is managed by rather than hosting. The operator installs nothing and +// owns nothing here: it resolves the endpoint, probes it, and reports. +// +// The word is about what manages the storage clusters rather than about who +// runs the operator. A deployment in this mode registers its clusters with a +// control plane it does not host, so the fleet is administered from there. +type ManagedControlPlane struct { + // Endpoint is the management API's base URL. It is validated against the + // same outbound-URL guard every other outbound endpoint in this group uses, + // so a loopback or link-local address is rejected. + // +kubebuilder:validation:Pattern=`^https?://[a-zA-Z0-9.-]+(:[0-9]{1,5})?(/.*)?$` + // +kubebuilder:validation:Required + Endpoint string `json:"endpoint"` + + // CredentialsSecretRef names a Secret in this namespace holding the bearer + // token the operator authenticates with. It is a reference rather than a + // field because a token in a spec is a token in every `kubectl get -o yaml`. + // + // Absent means the endpoint is reached without one, which is the in-cluster + // case: a control plane the Helm chart installed answers on a ClusterIP + // Service in this namespace and does not require a token for the readiness + // read. Naming a Secret that does not exist stays an error, because naming + // one is a statement that the control plane needs it. + // +optional + CredentialsSecretRef *corev1.LocalObjectReference `json:"credentialsSecretRef,omitempty"` + + // CABundleSecretRef names a Secret holding the CA certificate the endpoint + // is verified against. Absent means the system trust store. + // +optional + CABundleSecretRef *corev1.LocalObjectReference `json:"caBundleSecretRef,omitempty"` +} + +// ControlPlaneSource selects whether this cluster hosts its control plane or is +// managed by one elsewhere. Exactly one member is set, which is what makes the +// two modes siblings rather than two unrelated top-level fields, and which +// member it is cannot change afterward. +// +// Both rules are declared here rather than on the field that carries the block, +// for two separate reasons. +// +// The immutability is the interesting one. What is frozen is the choice between +// the two modes and not the block, because the members have to stay editable: +// changing spec.source.local.image is an ordinary edit, and it is what a +// ControlPlaneOps upgrade performs. Spelling it +k8s:immutable on the field +// would emit self == oldSelf over the whole struct, which freezes the image with +// it and makes that operation impossible to complete. +// +// The placement is the dull one. controller-gen v0.21.0 emits a single field's +// marker-derived rules and its injected immutability rule into one list in an +// order that varies between runs, so a field carrying both produces a CRD that +// differs from itself and a drift check that fails at random. Two rules of the +// same kind on a type are emitted in source order. +// +kubebuilder:validation:XValidation:rule="(has(self.local) ? 1 : 0) + (has(self.managed) ? 1 : 0) == 1",message="set exactly one of local or managed" +// +kubebuilder:validation:XValidation:rule="has(self.local) == has(oldSelf.local) && has(self.managed) == has(oldSelf.managed)",message="spec.source is immutable: a control plane the operator installed and one it did not are different deployments, and the clusters and their volumes live in the FoundationDB behind the old one" type ControlPlaneSource struct { - // Managed is the control plane the operator installs and owns. + // Local is a control plane the operator installs. + // +optional + Local *LocalControlPlane `json:"local,omitempty"` + + // Managed is a control plane that already exists. // +optional Managed *ManagedControlPlane `json:"managed,omitempty"` } -// ControlPlaneSpec holds configuration for the singleton ControlPlane resource. +// ControlPlaneSpec is the desired state of the simplyblock control plane for one +// namespace. type ControlPlaneSpec struct { - // Source says where the control plane comes from. It replaces the top-level - // image field of v1alpha1, which conflated the control plane's own image with - // the default every StorageNodeSet inherited. + // Source selects whether this cluster hosts its control plane or is managed + // by one elsewhere. Switching a live deployment between the two is not a + // reconfiguration, because the clusters and their volumes live in the + // FoundationDB behind the old one, so which mode is chosen is frozen at + // creation. What is inside the chosen mode stays editable. + // + // Both rules are declared on ControlPlaneSource rather than here. See the + // type for why. + // +kubebuilder:validation:Required + Source ControlPlaneSource `json:"source"` +} + +// ControlPlaneComponentStatus is one workload of a managed control plane and how +// much of it is running. The phase is the worst verdict across these and the +// readiness probe, and only an essential component at zero ready can make it +// Unavailable. +type ControlPlaneComponentStatus struct { + // Name is the workload's name, as applied. + // +kubebuilder:validation:Required + Name string `json:"name"` + + // Desired is how many replicas the component should have. For the component + // carrying its own operator it is that resource's own count, because a + // FoundationDBCluster reports quorum rather than replicas. + // +kubebuilder:validation:Minimum=0 + // +optional + Desired int32 `json:"desired"` + + // Ready is how many of them are. + // +kubebuilder:validation:Minimum=0 // +optional - Source *ControlPlaneSource `json:"source,omitempty"` + Ready int32 `json:"ready"` + + // Essential states whether this component at zero ready makes the control + // plane Unavailable rather than Degraded. It is decided by the table in + // §4.3 rather than by a user, and it is reported here so that a phase can be + // explained without reading the operator's source. + // +optional + Essential bool `json:"essential,omitempty"` } -// ControlPlaneStatus reflects the observed readiness of the simplyblock -// control plane (FDB + management API). +// ControlPlaneStatus is the observed state of the control plane. type ControlPlaneStatus struct { - // Phase is Initializing while the control plane is not yet healthy, - // and Available once the FDB health check passes. - // +kubebuilder:validation:Enum=Initializing;Available - Phase string `json:"phase,omitempty"` + // Phase is the operator's own view of the control plane. + // +optional + Phase ControlPlanePhase `json:"phase,omitempty"` - // Message contains a human-readable explanation of the current phase, - // for example, the FDB error returned by the health endpoint. - Message string `json:"message,omitempty"` + // Step is the position of the installation machine within Installing. The + // rule repeats the ControlPlaneStep enum because a marker cannot reach a + // field of the shared snapshot type. + // +kubebuilder:validation:XValidation:rule="!has(self.state) || self.state in ['ApplyingFoundationDB','AwaitingFoundationDB','ApplyingDatastore','ApplyingAPI','AwaitingAPI']",message="unknown step" + // +optional + Step statemachine.KubeSnapshot `json:"step,omitempty"` + + // Endpoint is the resolved management API base URL, derived in the local + // case and echoed in the managed one. It is what every controller in the + // operator reads to reach the control plane, so that one object answers + // where it is. + // +optional + Endpoint string `json:"endpoint,omitempty"` + + // Version is the version the management API reports. + // +optional + Version string `json:"version,omitempty"` - // LastChecked is the timestamp of the most recent FDB health probe. + // LastChecked is when the readiness probe last ran. + // +optional LastChecked *metav1.Time `json:"lastChecked,omitempty"` + + // Components is the per-component readiness the phase is derived from + // (§4.3), one entry per workload the managed install applies. It is empty + // for a remote control plane, which has no components the operator owns. + // Without it a Degraded phase says that something is wrong and not what. + // +optional + // +listType=map + // +listMapKey=name + Components []ControlPlaneComponentStatus `json:"components,omitempty"` + + // ActiveOpsRef names the ControlPlaneOps currently allowed to act on this + // control plane. Empty when none is running. + // +optional + ActiveOpsRef string `json:"activeOpsRef,omitempty"` + + // Message is the reason the phase is what it is: one sentence, replaced as + // the control plane moves, and never a log. On a failed probe it is the + // control plane's own error rather than a paraphrase of it. + // +optional + Message string `json:"message,omitempty"` + + // ObservedGeneration is the generation the rest of this status was computed + // from, so a stale status can be told from a current one. + // +optional + ObservedGeneration int64 `json:"observedGeneration,omitempty"` } // v1alpha2 is the storage version in the manifests this repository ships, which @@ -70,14 +313,19 @@ type ControlPlaneStatus struct { // +kubebuilder:storageversion // +kubebuilder:object:root=true // +kubebuilder:subresource:status -// +kubebuilder:resource:scope=Namespaced -// +kubebuilder:printcolumn:name="Phase",type="string",JSONPath=".status.phase",description="Initializing while FDB is not ready; Available once the control plane is operational" -// +kubebuilder:printcolumn:name="Message",type="string",JSONPath=".status.message",description="Human-readable status detail" -// +kubebuilder:printcolumn:name="Age",type="date",JSONPath=".metadata.creationTimestamp" - -// ControlPlane is a singleton resource (one per namespace, named "simplyblock") -// that reflects the readiness of the simplyblock control plane. It is created -// automatically by the Helm chart and should not be created or deleted manually. +// +kubebuilder:resource:scope=Namespaced,shortName=cp +// +kubebuilder:printcolumn:name="Phase",type=string,JSONPath=".status.phase" +// +kubebuilder:printcolumn:name="Step",type=string,JSONPath=".status.step.state" +// +kubebuilder:printcolumn:name="Endpoint",type=string,JSONPath=".status.endpoint" +// +kubebuilder:printcolumn:name="Version",type=string,JSONPath=".status.version" +// +kubebuilder:printcolumn:name="Message",type=string,JSONPath=".status.message",priority=1 +// +kubebuilder:printcolumn:name="Age",type=date,JSONPath=".metadata.creationTimestamp" + +// ControlPlane is the simplyblock control plane for one Kubernetes cluster: +// FoundationDB together with the management API, either installed by the +// operator or already existing. It is a singleton named `simplyblock`, and it is +// the root of the ownership spine: nothing else in this API group reconciles +// meaningfully before it reports Available. type ControlPlane struct { metav1.TypeMeta `json:",inline"` diff --git a/operator/api/v1alpha2/controlplaneops_types.go b/operator/api/v1alpha2/controlplaneops_types.go new file mode 100644 index 000000000..6f933b94a --- /dev/null +++ b/operator/api/v1alpha2/controlplaneops_types.go @@ -0,0 +1,262 @@ +// ControlPlaneOps: one imperative operation performed against the control plane. +// +// Most of what an administrator does to a control plane is expressible as +// desired state, because the entity re-applies what it installed on every pass: +// changing the image is an edit, scaling FoundationDB is an edit, and an object +// somebody deleted by hand is put back. What is left is what this kind carries, +// and it is three things: recycling a workload, moving it to a new version, and +// asking FoundationDB for a backup. +// +// Every action requires a control plane this cluster hosts, since each acts on +// something the operator installed. An operation naming a remote one is rejected +// at admission rather than created and failed, which is what +// ControlPlaneOpsValidator is for. +// +// The kind is introduced by the redesign, so it has no v1alpha1 spelling, no +// spoke to convert from, and no Hub method: its CRD declares one version. +// +// design-controlplane.md §6 and Appendix B are the specification. + +package v1alpha2 + +import ( + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + + "github.com/simplyblock/atlas/statemachine" +) + +// ControlPlaneOpsAction is the operation a ControlPlaneOps performs. Every +// action acts on what the operator installed, so every action requires a managed +// control plane, and the validating webhook of §6 rejects an operation naming an +// managed one at creation rather than letting it be created and fail. +// +kubebuilder:validation:Enum=Restart;Upgrade;Backup +type ControlPlaneOpsAction string + +const ( + // ControlPlaneOpsActionRestart recycles a wedged workload. Its scope is + // spec.restart.components, and an empty list recycles the whole control + // plane. + ControlPlaneOpsActionRestart ControlPlaneOpsAction = "Restart" + // ControlPlaneOpsActionUpgrade moves the control plane to a new version and + // verifies afterward that the version it reports is the one asked for. + ControlPlaneOpsActionUpgrade ControlPlaneOpsAction = "Upgrade" + // ControlPlaneOpsActionBackup asks FoundationDB for a backup outside + // whatever schedule exists. It asks rather than implements: the snapshot is + // the FoundationDB operator's to take. + ControlPlaneOpsActionBackup ControlPlaneOpsAction = "Backup" +) + +// ControlPlaneOpsPhase is the operation's own progress. +// +kubebuilder:validation:Enum=Pending;Running;Succeeded;Failed;Aborted +type ControlPlaneOpsPhase string + +const ( + // ControlPlaneOpsPhasePending is an operation waiting for its target's lock. + ControlPlaneOpsPhasePending ControlPlaneOpsPhase = "Pending" + // ControlPlaneOpsPhaseRunning is an operation holding the lock and working. + ControlPlaneOpsPhaseRunning ControlPlaneOpsPhase = "Running" + // ControlPlaneOpsPhaseSucceeded is a finished operation that did what it + // said. + ControlPlaneOpsPhaseSucceeded ControlPlaneOpsPhase = "Succeeded" + // ControlPlaneOpsPhaseFailed is a finished operation that did not. + ControlPlaneOpsPhaseFailed ControlPlaneOpsPhase = "Failed" + // ControlPlaneOpsPhaseAborted is an operation stopped on request, whose + // unwind has finished. + ControlPlaneOpsPhaseAborted ControlPlaneOpsPhase = "Aborted" +) + +// ControlPlaneOpsStep is one step of a running control-plane operation. Which +// steps belong to which action is declared by that action's graph rather than by +// this type, which is why the enum stays flat as actions are added. +// +kubebuilder:validation:Enum=Draining;Restarting;Awaiting;Preflight;Applying;Verifying;Requesting +type ControlPlaneOpsStep string + +const ( + // ControlPlaneOpsStepDraining holds while another operation in the namespace + // is still running. Restart and Upgrade both recycle the management API, so + // both drain. + ControlPlaneOpsStepDraining ControlPlaneOpsStep = "Draining" + + // ControlPlaneOpsStepRestarting rolls the workloads the action named. + ControlPlaneOpsStepRestarting ControlPlaneOpsStep = "Restarting" + // ControlPlaneOpsStepAwaiting waits for what the previous step asked for: + // the recycled pods coming back, or the backup reporting a snapshot. + ControlPlaneOpsStepAwaiting ControlPlaneOpsStep = "Awaiting" + + // ControlPlaneOpsStepPreflight reads live state, which admission cannot: it + // holds until the control plane is Available, and fails when the requested + // image is the one already running. + ControlPlaneOpsStepPreflight ControlPlaneOpsStep = "Preflight" + // ControlPlaneOpsStepApplying writes the new image onto the entity, which is + // what rolls the Deployment. + ControlPlaneOpsStepApplying ControlPlaneOpsStep = "Applying" + // ControlPlaneOpsStepVerifying re-probes and compares the reported version + // against the one asked for, which is what makes an upgrade more than an + // image bump. + ControlPlaneOpsStepVerifying ControlPlaneOpsStep = "Verifying" + + // ControlPlaneOpsStepRequesting creates or triggers the FoundationDBBackup. + // Awaiting is shared with Restart. + ControlPlaneOpsStepRequesting ControlPlaneOpsStep = "Requesting" +) + +// UpgradeSpec parameterizes the Upgrade action and is ignored by the others. +type UpgradeSpec struct { + // Image is the version to move to. It replaces + // ControlPlane.spec.source.managed.image when the operation succeeds, so the + // entity keeps describing what is running. + // +kubebuilder:validation:Pattern=`^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$` + // +kubebuilder:validation:Required + Image string `json:"image"` +} + +// RestartSpec parameterizes the Restart action and is ignored by the others. +type RestartSpec struct { + // Components names the workloads to recycle, from the table in §4.3. Empty + // recycles the whole control plane. Naming only components that table marks + // non-essential skips the drain, because recycling them interrupts nothing. + // +listType=set + // +optional + Components []string `json:"components,omitempty"` +} + +// BackupSpec parameterizes the Backup action and is ignored by the others. +type BackupSpec struct { + // BlobStore is the destination, in the form the FoundationDBBackup CRD takes + // it. The operator copies it through rather than interpreting it, since the + // backup is the FoundationDB operator's to perform. + // +kubebuilder:validation:Required + BlobStore string `json:"blobStore"` + + // BackupName is the FoundationDBBackup to create or trigger. Absent uses the + // one already configured for the cluster, and fails when there is none and + // no name to create. + // +optional + BackupName string `json:"backupName,omitempty"` +} + +// ControlPlaneOpsSpec is one operation to perform against the control plane. +// +// Everything except spec.abort is frozen once the object is admitted, which is +// what makes the status an audit of the request that ran rather than of whatever +// the object says now. The parameters are consumed several steps apart: +// Preflight reads spec.upgrade.image and Applying writes it, and Draining reads +// spec.restart.components before Restarting recycles them. An edit in between +// produces an operation that checked one thing and did another. +// +// The rules are declared here rather than as +k8s:immutable on each field. +// controller-gen emits that marker's rules in an order that varies between runs +// once a type carries several, and it freezes a block whole; what has to be +// frozen is each block's presence together with its contents. +// +kubebuilder:validation:XValidation:rule="has(self.upgrade) == has(oldSelf.upgrade) && (!has(self.upgrade) || self.upgrade == oldSelf.upgrade)",message="spec.upgrade is immutable: Preflight checked the image the operation was admitted with, and Applying writes it several steps later" +// +kubebuilder:validation:XValidation:rule="has(self.restart) == has(oldSelf.restart) && (!has(self.restart) || self.restart == oldSelf.restart)",message="spec.restart is immutable: the drain is decided from the component list, so widening it afterward skips a drain the wider list would have required" +// +kubebuilder:validation:XValidation:rule="has(self.backup) == has(oldSelf.backup) && (!has(self.backup) || self.backup == oldSelf.backup)",message="spec.backup is immutable: the destination is what Requesting created the FoundationDBBackup against" +type ControlPlaneOpsSpec struct { + // ControlPlaneRef names the ControlPlane this operation acts on, in this + // object's own namespace. The operation never owns its target, because + // deleting the record of an operation must not delete the control plane it + // operated on. + // +kubebuilder:validation:Required + // +k8s:immutable + ControlPlaneRef string `json:"controlPlaneRef"` + + // Action is the operation to perform. Immutable, so that the status describes + // the operation that ran. + // +kubebuilder:validation:Required + // +k8s:immutable + Action ControlPlaneOpsAction `json:"action"` + + // Abort asks a running operation to stop at its next step and unwind. It is + // the one field of this spec an update may change, because it is the one that + // is meant to be set after the operation started. Whether an abort is + // expressible from the current step is declared by that action's graph rather + // than checked here. + // +optional + Abort bool `json:"abort,omitempty"` + + // Upgrade parameterizes action Upgrade and is ignored by the others. + // +optional + Upgrade *UpgradeSpec `json:"upgrade,omitempty"` + + // Restart parameterizes action Restart and is ignored by the others. + // +optional + Restart *RestartSpec `json:"restart,omitempty"` + + // Backup parameterizes action Backup and is ignored by the others. + // +optional + Backup *BackupSpec `json:"backup,omitempty"` +} + +// ControlPlaneOpsStatus is the observed state of one control-plane operation. +type ControlPlaneOpsStatus struct { + // Phase is the operation's own progress. + // +optional + Phase ControlPlaneOpsPhase `json:"phase,omitempty"` + + // Step is the position of the running action's state machine. It is + // persisted before the side effect that step performs. The rule repeats the + // ControlPlaneOpsStep enum because a marker cannot reach a field of the + // shared snapshot type. + // +kubebuilder:validation:XValidation:rule="!has(self.state) || self.state in ['Draining','Restarting','Awaiting','Preflight','Applying','Verifying','Requesting']",message="unknown step" + // +optional + Step statemachine.KubeSnapshot `json:"step,omitempty"` + + // Message is the reason the phase is what it is: one sentence, replaced as + // the operation moves, and never a log. + // +optional + Message string `json:"message,omitempty"` + + // BackupRef names the FoundationDBBackup a Backup run created or triggered. + // The operation does not own it, because deleting the record of a backup + // must not delete the backup's configuration. + // +optional + BackupRef string `json:"backupRef,omitempty"` + + // ObservedGeneration is the generation the rest of this status was computed + // from, so a stale status can be told from a current one. + // +optional + ObservedGeneration int64 `json:"observedGeneration,omitempty"` + + // StartedAt is when the operation acquired its target's lock. + // +optional + StartedAt *metav1.Time `json:"startedAt,omitempty"` + + // CompletedAt is when it reached a terminal phase. + // +optional + CompletedAt *metav1.Time `json:"completedAt,omitempty"` +} + +// +kubebuilder:object:root=true +// +kubebuilder:subresource:status +// +kubebuilder:resource:scope=Namespaced,shortName=cpops +// +kubebuilder:printcolumn:name="ControlPlane",type=string,JSONPath=".spec.controlPlaneRef" +// +kubebuilder:printcolumn:name="Action",type=string,JSONPath=".spec.action" +// +kubebuilder:printcolumn:name="Phase",type=string,JSONPath=".status.phase" +// +kubebuilder:printcolumn:name="Step",type=string,JSONPath=".status.step.state" +// +kubebuilder:printcolumn:name="Message",type=string,JSONPath=".status.message",priority=1 +// +kubebuilder:printcolumn:name="Age",type=date,JSONPath=".metadata.creationTimestamp" + +// ControlPlaneOps is a single operation performed against the control plane. It +// runs to a terminal phase and stays afterward as the audit record of what was +// done, with which parameters, and how it ended. Only one may be active per +// control plane at a time, which the entity's status.activeOpsRef enforces. +type ControlPlaneOps struct { + metav1.TypeMeta `json:",inline"` + metav1.ObjectMeta `json:"metadata,omitempty"` + + Spec ControlPlaneOpsSpec `json:"spec,omitempty"` + Status ControlPlaneOpsStatus `json:"status,omitempty"` +} + +// +kubebuilder:object:root=true + +// ControlPlaneOpsList contains a list of ControlPlaneOps. +type ControlPlaneOpsList struct { + metav1.TypeMeta `json:",inline"` + metav1.ListMeta `json:"metadata,omitempty"` + Items []ControlPlaneOps `json:"items"` +} + +func init() { + SchemeBuilder.Register(&ControlPlaneOps{}, &ControlPlaneOpsList{}) +} diff --git a/operator/api/v1alpha2/zz_generated.deepcopy.go b/operator/api/v1alpha2/zz_generated.deepcopy.go index 1a6dcdbdc..954385cfe 100644 --- a/operator/api/v1alpha2/zz_generated.deepcopy.go +++ b/operator/api/v1alpha2/zz_generated.deepcopy.go @@ -88,6 +88,21 @@ func (in *BackupSource) DeepCopy() *BackupSource { return out } +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *BackupSpec) DeepCopyInto(out *BackupSpec) { + *out = *in +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new BackupSpec. +func (in *BackupSpec) DeepCopy() *BackupSpec { + if in == nil { + return nil + } + out := new(BackupSpec) + in.DeepCopyInto(out) + return out +} + // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *BackupStoreSpec) DeepCopyInto(out *BackupStoreSpec) { *out = *in @@ -348,6 +363,21 @@ func (in *ControlPlane) DeepCopyObject() runtime.Object { return nil } +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *ControlPlaneComponentStatus) DeepCopyInto(out *ControlPlaneComponentStatus) { + *out = *in +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new ControlPlaneComponentStatus. +func (in *ControlPlaneComponentStatus) DeepCopy() *ControlPlaneComponentStatus { + if in == nil { + return nil + } + out := new(ControlPlaneComponentStatus) + in.DeepCopyInto(out) + return out +} + // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *ControlPlaneList) DeepCopyInto(out *ControlPlaneList) { *out = *in @@ -380,13 +410,131 @@ func (in *ControlPlaneList) DeepCopyObject() runtime.Object { return nil } +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *ControlPlaneOps) DeepCopyInto(out *ControlPlaneOps) { + *out = *in + out.TypeMeta = in.TypeMeta + in.ObjectMeta.DeepCopyInto(&out.ObjectMeta) + in.Spec.DeepCopyInto(&out.Spec) + in.Status.DeepCopyInto(&out.Status) +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new ControlPlaneOps. +func (in *ControlPlaneOps) DeepCopy() *ControlPlaneOps { + if in == nil { + return nil + } + out := new(ControlPlaneOps) + in.DeepCopyInto(out) + return out +} + +// DeepCopyObject is an autogenerated deepcopy function, copying the receiver, creating a new runtime.Object. +func (in *ControlPlaneOps) DeepCopyObject() runtime.Object { + if c := in.DeepCopy(); c != nil { + return c + } + return nil +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *ControlPlaneOpsList) DeepCopyInto(out *ControlPlaneOpsList) { + *out = *in + out.TypeMeta = in.TypeMeta + in.ListMeta.DeepCopyInto(&out.ListMeta) + if in.Items != nil { + in, out := &in.Items, &out.Items + *out = make([]ControlPlaneOps, len(*in)) + for i := range *in { + (*in)[i].DeepCopyInto(&(*out)[i]) + } + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new ControlPlaneOpsList. +func (in *ControlPlaneOpsList) DeepCopy() *ControlPlaneOpsList { + if in == nil { + return nil + } + out := new(ControlPlaneOpsList) + in.DeepCopyInto(out) + return out +} + +// DeepCopyObject is an autogenerated deepcopy function, copying the receiver, creating a new runtime.Object. +func (in *ControlPlaneOpsList) DeepCopyObject() runtime.Object { + if c := in.DeepCopy(); c != nil { + return c + } + return nil +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *ControlPlaneOpsSpec) DeepCopyInto(out *ControlPlaneOpsSpec) { + *out = *in + if in.Upgrade != nil { + in, out := &in.Upgrade, &out.Upgrade + *out = new(UpgradeSpec) + **out = **in + } + if in.Restart != nil { + in, out := &in.Restart, &out.Restart + *out = new(RestartSpec) + (*in).DeepCopyInto(*out) + } + if in.Backup != nil { + in, out := &in.Backup, &out.Backup + *out = new(BackupSpec) + **out = **in + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new ControlPlaneOpsSpec. +func (in *ControlPlaneOpsSpec) DeepCopy() *ControlPlaneOpsSpec { + if in == nil { + return nil + } + out := new(ControlPlaneOpsSpec) + in.DeepCopyInto(out) + return out +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *ControlPlaneOpsStatus) DeepCopyInto(out *ControlPlaneOpsStatus) { + *out = *in + in.Step.DeepCopyInto(&out.Step) + if in.StartedAt != nil { + in, out := &in.StartedAt, &out.StartedAt + *out = (*in).DeepCopy() + } + if in.CompletedAt != nil { + in, out := &in.CompletedAt, &out.CompletedAt + *out = (*in).DeepCopy() + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new ControlPlaneOpsStatus. +func (in *ControlPlaneOpsStatus) DeepCopy() *ControlPlaneOpsStatus { + if in == nil { + return nil + } + out := new(ControlPlaneOpsStatus) + in.DeepCopyInto(out) + return out +} + // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *ControlPlaneSource) DeepCopyInto(out *ControlPlaneSource) { *out = *in + if in.Local != nil { + in, out := &in.Local, &out.Local + *out = new(LocalControlPlane) + (*in).DeepCopyInto(*out) + } if in.Managed != nil { in, out := &in.Managed, &out.Managed *out = new(ManagedControlPlane) - **out = **in + (*in).DeepCopyInto(*out) } } @@ -403,11 +551,7 @@ func (in *ControlPlaneSource) DeepCopy() *ControlPlaneSource { // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *ControlPlaneSpec) DeepCopyInto(out *ControlPlaneSpec) { *out = *in - if in.Source != nil { - in, out := &in.Source, &out.Source - *out = new(ControlPlaneSource) - (*in).DeepCopyInto(*out) - } + in.Source.DeepCopyInto(&out.Source) } // DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new ControlPlaneSpec. @@ -423,10 +567,16 @@ func (in *ControlPlaneSpec) DeepCopy() *ControlPlaneSpec { // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *ControlPlaneStatus) DeepCopyInto(out *ControlPlaneStatus) { *out = *in + in.Step.DeepCopyInto(&out.Step) if in.LastChecked != nil { in, out := &in.LastChecked, &out.LastChecked *out = (*in).DeepCopy() } + if in.Components != nil { + in, out := &in.Components, &out.Components + *out = make([]ControlPlaneComponentStatus, len(*in)) + copy(*out, *in) + } } // DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new ControlPlaneStatus. @@ -641,6 +791,27 @@ func (in *DriverTLS) DeepCopy() *DriverTLS { return out } +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *FoundationDBSpec) DeepCopyInto(out *FoundationDBSpec) { + *out = *in + if in.Replicas != nil { + in, out := &in.Replicas, &out.Replicas + *out = new(int32) + **out = **in + } + in.Resources.DeepCopyInto(&out.Resources) +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new FoundationDBSpec. +func (in *FoundationDBSpec) DeepCopy() *FoundationDBSpec { + if in == nil { + return nil + } + out := new(FoundationDBSpec) + in.DeepCopyInto(out) + return out +} + // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *JournalManagerSpec) DeepCopyInto(out *JournalManagerSpec) { *out = *in @@ -686,9 +857,59 @@ func (in *KMSSpec) DeepCopy() *KMSSpec { return out } +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *LocalControlPlane) DeepCopyInto(out *LocalControlPlane) { + *out = *in + if in.FoundationDB != nil { + in, out := &in.FoundationDB, &out.FoundationDB + *out = new(FoundationDBSpec) + (*in).DeepCopyInto(*out) + } + if in.Replicas != nil { + in, out := &in.Replicas, &out.Replicas + *out = new(int32) + **out = **in + } + in.Resources.DeepCopyInto(&out.Resources) + if in.Tolerations != nil { + in, out := &in.Tolerations, &out.Tolerations + *out = make([]v1.Toleration, len(*in)) + for i := range *in { + (*in)[i].DeepCopyInto(&(*out)[i]) + } + } + if in.NodeSelector != nil { + in, out := &in.NodeSelector, &out.NodeSelector + *out = make(map[string]string, len(*in)) + for key, val := range *in { + (*out)[key] = val + } + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new LocalControlPlane. +func (in *LocalControlPlane) DeepCopy() *LocalControlPlane { + if in == nil { + return nil + } + out := new(LocalControlPlane) + in.DeepCopyInto(out) + return out +} + // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *ManagedControlPlane) DeepCopyInto(out *ManagedControlPlane) { *out = *in + if in.CredentialsSecretRef != nil { + in, out := &in.CredentialsSecretRef, &out.CredentialsSecretRef + *out = new(v1.LocalObjectReference) + **out = **in + } + if in.CABundleSecretRef != nil { + in, out := &in.CABundleSecretRef, &out.CABundleSecretRef + *out = new(v1.LocalObjectReference) + **out = **in + } } // DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new ManagedControlPlane. @@ -1021,6 +1242,26 @@ func (in *RemoveSpec) DeepCopy() *RemoveSpec { return out } +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *RestartSpec) DeepCopyInto(out *RestartSpec) { + *out = *in + if in.Components != nil { + in, out := &in.Components, &out.Components + *out = make([]string, len(*in)) + copy(*out, *in) + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new RestartSpec. +func (in *RestartSpec) DeepCopy() *RestartSpec { + if in == nil { + return nil + } + out := new(RestartSpec) + in.DeepCopyInto(out) + return out +} + // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *RestoreSpec) DeepCopyInto(out *RestoreSpec) { *out = *in @@ -2744,6 +2985,21 @@ func (in *ThroughputLimits) DeepCopy() *ThroughputLimits { return out } +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *UpgradeSpec) DeepCopyInto(out *UpgradeSpec) { + *out = *in +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new UpgradeSpec. +func (in *UpgradeSpec) DeepCopy() *UpgradeSpec { + if in == nil { + return nil + } + out := new(UpgradeSpec) + in.DeepCopyInto(out) + return out +} + // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *VaultKMS) DeepCopyInto(out *VaultKMS) { *out = *in diff --git a/operator/cmd/main.go b/operator/cmd/main.go index 8a65d8275..1dcb32e54 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -57,6 +57,7 @@ import ( "github.com/simplyblock/simplyblock-operator/internal/controller" backupcontrollers "github.com/simplyblock/simplyblock-operator/internal/controllers/backup" clustercontroller "github.com/simplyblock/simplyblock-operator/internal/controllers/cluster" + controlplanecontroller "github.com/simplyblock/simplyblock-operator/internal/controllers/controlplane" "github.com/simplyblock/simplyblock-operator/internal/controllers/deployment" "github.com/simplyblock/simplyblock-operator/internal/controllers/driver" nodecontroller "github.com/simplyblock/simplyblock-operator/internal/controllers/node" @@ -473,7 +474,14 @@ func main() { os.Exit(1) } - if err := (&controller.ControlPlaneReconciler{ + // Where the control plane is, read from the ControlPlane object rather than + // from the environment (design-controlplane.md §3.3). Every control-plane + // client below shares it, so an external control plane is reachable and a + // change to its endpoint reaches every caller without a rollout. + controlPlaneEndpoint := controlplanecontroller.NewEndpointResolver( + mgr.GetClient(), operatorNamespace) + + if err := (&controlplanecontroller.ControlPlaneReconciler{ Client: mgr.GetClient(), Scheme: mgr.GetScheme(), Recorder: mgr.GetEventRecorder("controlplane-controller"), @@ -481,11 +489,19 @@ func main() { setupLog.Error(err, "unable to create controller", "controller", "ControlPlane") os.Exit(1) } + if err := (&controlplanecontroller.ControlPlaneOpsReconciler{ + Client: mgr.GetClient(), + Scheme: mgr.GetScheme(), + Recorder: mgr.GetEventRecorder("controlplaneops-controller"), + }).SetupWithManager(mgr); err != nil { + setupLog.Error(err, "unable to create controller", "controller", "ControlPlaneOps") + os.Exit(1) + } if err := (&clustercontroller.StorageClusterReconciler{ Client: mgr.GetClient(), Scheme: mgr.GetScheme(), Recorder: mgr.GetEventRecorder("storagecluster-controller"), - API: clustercontroller.NewControlPlane(), + API: clustercontroller.NewControlPlane(controlPlaneEndpoint), Namespace: operatorNamespace, Clusters: clusterSubscription, Tasks: taskSubscription, @@ -626,7 +642,7 @@ func main() { Client: mgr.GetClient(), Scheme: mgr.GetScheme(), Recorder: mgr.GetEventRecorder("storagenode-controller"), - API: nodecontroller.NewControlPlane(), + API: nodecontroller.NewControlPlane(controlPlaneEndpoint), Nodes: nodeSubscription, Registries: []nodecontroller.NodeObjectRegistry{ deviceSubscription, nodeSubscription, @@ -642,7 +658,7 @@ func main() { Client: mgr.GetClient(), Scheme: mgr.GetScheme(), Recorder: mgr.GetEventRecorder("storagenodeops-controller"), - API: nodecontroller.NewControlPlane(), + API: nodecontroller.NewControlPlane(controlPlaneEndpoint), Nodes: nodeSubscription, Clusters: clusterSubscription, Workload: storageNodeWorkload, @@ -705,7 +721,7 @@ func main() { Client: mgr.GetClient(), Scheme: mgr.GetScheme(), Recorder: mgr.GetEventRecorder("storageclusterops-controller"), - API: clustercontroller.NewControlPlane(), + API: clustercontroller.NewControlPlane(controlPlaneEndpoint), Clusters: clusterSubscription, Nodes: nodeSubscription, Tasks: taskSubscription, @@ -842,6 +858,10 @@ func main() { &webhook.Admission{Handler: &internalwebhook.StorageBackupOpsValidator{Client: mgr.GetClient()}}) setupLog.Info("registered storagebackupops validating webhook") + mgr.GetWebhookServer().Register("/validate-storage-simplyblock-io-v1alpha2-controlplaneops", + &webhook.Admission{Handler: &internalwebhook.ControlPlaneOpsValidator{Client: mgr.GetClient()}}) + setupLog.Info("registered controlplaneops validating webhook") + mgr.GetWebhookServer().Register("/validate-v1-pvc-pinned-volume", &webhook.Admission{Handler: &internalwebhook.PersistentVolumeClaimValidator{ Client: mgr.GetClient(), diff --git a/operator/config/crd/bases/storage.simplyblock.io_controlplaneops.yaml b/operator/config/crd/bases/storage.simplyblock.io_controlplaneops.yaml new file mode 100644 index 000000000..c0a9a5f6b --- /dev/null +++ b/operator/config/crd/bases/storage.simplyblock.io_controlplaneops.yaml @@ -0,0 +1,244 @@ +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition +metadata: + annotations: + controller-gen.kubebuilder.io/version: v0.21.0 + name: controlplaneops.storage.simplyblock.io +spec: + group: storage.simplyblock.io + names: + kind: ControlPlaneOps + listKind: ControlPlaneOpsList + plural: controlplaneops + shortNames: + - cpops + singular: controlplaneops + scope: Namespaced + versions: + - additionalPrinterColumns: + - jsonPath: .spec.controlPlaneRef + name: ControlPlane + type: string + - jsonPath: .spec.action + name: Action + type: string + - jsonPath: .status.phase + name: Phase + type: string + - jsonPath: .status.step.state + name: Step + type: string + - jsonPath: .status.message + name: Message + priority: 1 + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha2 + schema: + openAPIV3Schema: + description: |- + ControlPlaneOps is a single operation performed against the control plane. It + runs to a terminal phase and stays afterward as the audit record of what was + done, with which parameters, and how it ended. Only one may be active per + control plane at a time, which the entity's status.activeOpsRef enforces. + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: |- + ControlPlaneOpsSpec is one operation to perform against the control plane. + + Everything except spec.abort is frozen once the object is admitted, which is + what makes the status an audit of the request that ran rather than of whatever + the object says now. The parameters are consumed several steps apart: + Preflight reads spec.upgrade.image and Applying writes it, and Draining reads + spec.restart.components before Restarting recycles them. An edit in between + produces an operation that checked one thing and did another. + + The rules are declared here rather than as +k8s:immutable on each field. + controller-gen emits that marker's rules in an order that varies between runs + once a type carries several, and it freezes a block whole; what has to be + frozen is each block's presence together with its contents. + properties: + abort: + description: |- + Abort asks a running operation to stop at its next step and unwind. It is + the one field of this spec an update may change, because it is the one that + is meant to be set after the operation started. Whether an abort is + expressible from the current step is declared by that action's graph rather + than checked here. + type: boolean + action: + description: |- + Action is the operation to perform. Immutable, so that the status describes + the operation that ran. + enum: + - Restart + - Upgrade + - Backup + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + backup: + description: Backup parameterizes action Backup and is ignored by + the others. + properties: + backupName: + description: |- + BackupName is the FoundationDBBackup to create or trigger. Absent uses the + one already configured for the cluster, and fails when there is none and + no name to create. + type: string + blobStore: + description: |- + BlobStore is the destination, in the form the FoundationDBBackup CRD takes + it. The operator copies it through rather than interpreting it, since the + backup is the FoundationDB operator's to perform. + type: string + required: + - blobStore + type: object + controlPlaneRef: + description: |- + ControlPlaneRef names the ControlPlane this operation acts on, in this + object's own namespace. The operation never owns its target, because + deleting the record of an operation must not delete the control plane it + operated on. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + restart: + description: Restart parameterizes action Restart and is ignored by + the others. + properties: + components: + description: |- + Components names the workloads to recycle, from the table in §4.3. Empty + recycles the whole control plane. Naming only components that table marks + non-essential skips the drain, because recycling them interrupts nothing. + items: + type: string + type: array + x-kubernetes-list-type: set + type: object + upgrade: + description: Upgrade parameterizes action Upgrade and is ignored by + the others. + properties: + image: + description: |- + Image is the version to move to. It replaces + ControlPlane.spec.source.managed.image when the operation succeeds, so the + entity keeps describing what is running. + pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ + type: string + required: + - image + type: object + required: + - action + - controlPlaneRef + type: object + x-kubernetes-validations: + - message: 'spec.upgrade is immutable: Preflight checked the image the + operation was admitted with, and Applying writes it several steps + later' + rule: has(self.upgrade) == has(oldSelf.upgrade) && (!has(self.upgrade) + || self.upgrade == oldSelf.upgrade) + - message: 'spec.restart is immutable: the drain is decided from the component + list, so widening it afterward skips a drain the wider list would + have required' + rule: has(self.restart) == has(oldSelf.restart) && (!has(self.restart) + || self.restart == oldSelf.restart) + - message: 'spec.backup is immutable: the destination is what Requesting + created the FoundationDBBackup against' + rule: has(self.backup) == has(oldSelf.backup) && (!has(self.backup) + || self.backup == oldSelf.backup) + status: + description: ControlPlaneOpsStatus is the observed state of one control-plane + operation. + properties: + backupRef: + description: |- + BackupRef names the FoundationDBBackup a Backup run created or triggered. + The operation does not own it, because deleting the record of a backup + must not delete the backup's configuration. + type: string + completedAt: + description: CompletedAt is when it reached a terminal phase. + format: date-time + type: string + message: + description: |- + Message is the reason the phase is what it is: one sentence, replaced as + the operation moves, and never a log. + type: string + observedGeneration: + description: |- + ObservedGeneration is the generation the rest of this status was computed + from, so a stale status can be told from a current one. + format: int64 + type: integer + phase: + description: Phase is the operation's own progress. + enum: + - Pending + - Running + - Succeeded + - Failed + - Aborted + type: string + startedAt: + description: StartedAt is when the operation acquired its target's + lock. + format: date-time + type: string + step: + description: |- + Step is the position of the running action's state machine. It is + persisted before the side effect that step performs. The rule repeats the + ControlPlaneOpsStep enum because a marker cannot reach a field of the + shared snapshot type. + properties: + deadline: + description: |- + Deadline is when that state expires, absent when it has none. It is an + absolute instant, so a state whose deadline passed while the controller + was down restores as already expired. + format: date-time + type: string + state: + description: |- + State is the state the machine was in. Empty means the resource has not + been reconciled yet, and restores to the graph's initial state. + type: string + type: object + x-kubernetes-validations: + - message: unknown step + rule: '!has(self.state) || self.state in [''Draining'',''Restarting'',''Awaiting'',''Preflight'',''Applying'',''Verifying'',''Requesting'']' + type: object + type: object + served: true + storage: true + subresources: + status: {} diff --git a/operator/config/crd/bases/storage.simplyblock.io_controlplanes.yaml b/operator/config/crd/bases/storage.simplyblock.io_controlplanes.yaml index 1f080d803..140c3da8f 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_controlplanes.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_controlplanes.yaml @@ -11,6 +11,8 @@ spec: kind: ControlPlane listKind: ControlPlaneList plural: controlplanes + shortNames: + - cp singular: controlplane scope: Namespaced versions: @@ -59,9 +61,11 @@ spec: image: description: |- Image is the container image used for all simplyblock control-plane and - storage-node workloads (e.g. quay.io/simplyblock-io/simplyblock:26.2.2). + storage-node workloads (e.g., `quay.io/simplyblock-io/simplyblock:26.2.2`). StorageNodeSet CRs that omit spec.clusterImage inherit this value. - Must reference one of the trusted registries (quay.io/simplyblock-io, docker.io/simplyblock, public.ecr.aws/simply-block); digest pinning (@sha256:...) is recommended. + Must reference one of the trusted registries (`quay.io/simplyblock-io`, + `docker.io/simplyblock`, `public.ecr.aws/simply-block`). Digest pinning + (@sha256:...) is recommended. pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ type: string type: object @@ -77,7 +81,7 @@ spec: message: description: |- Message contains a human-readable explanation of the current phase, - for example the FDB error returned by the health endpoint. + for example, the FDB error returned by the health endpoint. type: string phase: description: |- @@ -94,13 +98,21 @@ spec: subresources: status: {} - additionalPrinterColumns: - - description: Initializing while FDB is not ready; Available once the control plane is operational - jsonPath: .status.phase + - jsonPath: .status.phase name: Phase type: string - - description: Human-readable status detail - jsonPath: .status.message + - jsonPath: .status.step.state + name: Step + type: string + - jsonPath: .status.endpoint + name: Endpoint + type: string + - jsonPath: .status.version + name: Version + type: string + - jsonPath: .status.message name: Message + priority: 1 type: string - jsonPath: .metadata.creationTimestamp name: Age @@ -109,9 +121,11 @@ spec: schema: openAPIV3Schema: description: |- - ControlPlane is a singleton resource (one per namespace, named "simplyblock") - that reflects the readiness of the simplyblock control plane. It is created - automatically by the Helm chart and should not be created or deleted manually. + ControlPlane is the simplyblock control plane for one Kubernetes cluster: + FoundationDB together with the management API, either installed by the + operator or already existing. It is a singleton named `simplyblock`, and it is + the root of the ownership spine: nothing else in this API group reconciles + meaningfully before it reports Available. properties: apiVersion: description: |- @@ -131,48 +145,406 @@ spec: metadata: type: object spec: - description: ControlPlaneSpec holds configuration for the singleton ControlPlane resource. + description: |- + ControlPlaneSpec is the desired state of the simplyblock control plane for one + namespace. properties: source: description: |- - Source says where the control plane comes from. It replaces the top-level - image field of v1alpha1, which conflated the control plane's own image with - the default every StorageNodeSet inherited. + Source selects whether this cluster hosts its control plane or is managed + by one elsewhere. Switching a live deployment between the two is not a + reconfiguration, because the clusters and their volumes live in the + FoundationDB behind the old one, so which mode is chosen is frozen at + creation. What is inside the chosen mode stays editable. + + Both rules are declared on ControlPlaneSource rather than here. See the + type for why. properties: - managed: - description: Managed is the control plane the operator installs and owns. + local: + description: Local is a control plane the operator installs. properties: + foundationDB: + description: FoundationDB sizes the FoundationDB the management API stores its state in. + properties: + replicas: + default: 3 + description: |- + Replicas is the number of coordinators. Three is the smallest count that + survives one loss, which is why it is the default. + format: int32 + minimum: 1 + type: integer + resources: + description: Resources sets requests and limits for the coordinator pods. + properties: + claims: + description: |- + Claims lists the names of resources, defined in spec.resourceClaims, + that are used by this container. + + This field depends on the + DynamicResourceAllocation feature gate. + + This field is immutable. It can only be set for containers. + items: + description: ResourceClaim references one entry in PodSpec.ResourceClaims. + properties: + name: + description: |- + Name must match the name of one entry in pod.spec.resourceClaims of + the Pod where this field is used. It makes that resource available + inside a container. + type: string + request: + description: |- + Request is the name chosen for a request in the referenced claim. + If empty, everything from the claim is made available, otherwise + only the result of this request. + type: string + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + limits: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Limits describes the maximum amount of compute resources allowed. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + requests: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Requests describes the minimum amount of compute resources required. + If Requests is omitted for a container, it defaults to Limits if that is explicitly specified, + otherwise to an implementation-defined value. Requests cannot exceed Limits. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + type: object + storageClassName: + description: |- + StorageClassName is the class the coordinators' volumes are provisioned + from. It cannot be a class this operator provides, because the control + plane has to exist before any simplyblock volume can. + type: string + type: object image: description: |- - Image is the container image used for the simplyblock control-plane - workloads (e.g., quay.io/simplyblock-io/simplyblock:26.2.2). - Must reference one of the trusted registries (`quay.io/simplyblock-io`, `docker.io/simplyblock`, `public.ecr.aws/simply-block`); digest pinning (@sha256:...) is recommended. + Image is the management API and control-plane image. + Must reference one of the trusted registries (`quay.io/simplyblock-io`, + `docker.io/simplyblock`, `public.ecr.aws/simply-block`); digest pinning + (@sha256:...) is recommended. pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ type: string + imagePullPolicy: + default: IfNotPresent + description: ImagePullPolicy controls when that image is pulled. + enum: + - Always + - Never + - IfNotPresent + type: string + nodeSelector: + additionalProperties: + type: string + description: |- + NodeSelector pins every pod the operator installs for the control plane. + It is a selector rather than an affinity term because that is what the + chart it replaces took, and a deployment migrating off the chart has the + value already written down. + type: object + replicas: + default: 2 + description: |- + Replicas is the number of management API instances. Two is what the chart + ships and what the phases assume: a single instance makes Degraded + unreachable for this component and every restart an outage (§5.1). + format: int32 + minimum: 1 + type: integer + resources: + description: Resources sets requests and limits for the management API pods. + properties: + claims: + description: |- + Claims lists the names of resources, defined in spec.resourceClaims, + that are used by this container. + + This field depends on the + DynamicResourceAllocation feature gate. + + This field is immutable. It can only be set for containers. + items: + description: ResourceClaim references one entry in PodSpec.ResourceClaims. + properties: + name: + description: |- + Name must match the name of one entry in pod.spec.resourceClaims of + the Pod where this field is used. It makes that resource available + inside a container. + type: string + request: + description: |- + Request is the name chosen for a request in the referenced claim. + If empty, everything from the claim is made available, otherwise + only the result of this request. + type: string + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + limits: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Limits describes the maximum amount of compute resources allowed. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + requests: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Requests describes the minimum amount of compute resources required. + If Requests is omitted for a container, it defaults to Limits if that is explicitly specified, + otherwise to an implementation-defined value. Requests cannot exceed Limits. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + type: object + tolerations: + description: |- + Tolerations are applied to every pod the operator installs for the control + plane. + items: + description: |- + The pod this Toleration is attached to tolerates any taint that matches + the triple using the matching operator . + properties: + effect: + description: |- + Effect indicates the taint effect to match. Empty means match all taint effects. + When specified, allowed values are NoSchedule, PreferNoSchedule and NoExecute. + type: string + key: + description: |- + Key is the taint key that the toleration applies to. Empty means match all taint keys. + If the key is empty, operator must be Exists; this combination means to match all values and all keys. + type: string + operator: + description: |- + Operator represents a key's relationship to the value. + Valid operators are Exists, Equal, Lt, and Gt. Defaults to Equal. + Exists is equivalent to wildcard for value, so that a pod can + tolerate all taints of a particular category. + Lt and Gt perform numeric comparisons (requires feature gate TaintTolerationComparisonOperators). + type: string + tolerationSeconds: + description: |- + TolerationSeconds represents the period of time the toleration (which must be + of effect NoExecute, otherwise this field is ignored) tolerates the taint. By default, + it is not set, which means tolerate the taint forever (do not evict). Zero and + negative values will be treated as 0 (evict immediately) by the system. + format: int64 + type: integer + value: + description: |- + Value is the taint value the toleration matches to. + If the operator is Exists, the value should be empty, otherwise just a regular string. + type: string + type: object + type: array + required: + - image + type: object + managed: + description: Managed is a control plane that already exists. + properties: + caBundleSecretRef: + description: |- + CABundleSecretRef names a Secret holding the CA certificate the endpoint + is verified against. Absent means the system trust store. + properties: + name: + default: "" + description: |- + Name of the referent. + This field is effectively required, but due to backwards compatibility is + allowed to be empty. Instances of this type with an empty value here are + almost certainly wrong. + More info: https://kubernetes.io/docs/concepts/overview/working-with-objects/names/#names + type: string + type: object + x-kubernetes-map-type: atomic + credentialsSecretRef: + description: |- + CredentialsSecretRef names a Secret in this namespace holding the bearer + token the operator authenticates with. It is a reference rather than a + field because a token in a spec is a token in every `kubectl get -o yaml`. + + Absent means the endpoint is reached without one, which is the in-cluster + case: a control plane the Helm chart installed answers on a ClusterIP + Service in this namespace and does not require a token for the readiness + read. Naming a Secret that does not exist stays an error, because naming + one is a statement that the control plane needs it. + properties: + name: + default: "" + description: |- + Name of the referent. + This field is effectively required, but due to backwards compatibility is + allowed to be empty. Instances of this type with an empty value here are + almost certainly wrong. + More info: https://kubernetes.io/docs/concepts/overview/working-with-objects/names/#names + type: string + type: object + x-kubernetes-map-type: atomic + endpoint: + description: |- + Endpoint is the management API's base URL. It is validated against the + same outbound-URL guard every other outbound endpoint in this group uses, + so a loopback or link-local address is rejected. + pattern: ^https?://[a-zA-Z0-9.-]+(:[0-9]{1,5})?(/.*)?$ + type: string + required: + - endpoint type: object type: object + x-kubernetes-validations: + - message: set exactly one of local or managed + rule: '(has(self.local) ? 1 : 0) + (has(self.managed) ? 1 : 0) == 1' + - message: 'spec.source is immutable: a control plane the operator installed and one it did not are different deployments, and the clusters and their volumes live in the FoundationDB behind the old one' + rule: has(self.local) == has(oldSelf.local) && has(self.managed) == has(oldSelf.managed) + required: + - source type: object status: - description: |- - ControlPlaneStatus reflects the observed readiness of the simplyblock - control plane (FDB + management API). + description: ControlPlaneStatus is the observed state of the control plane. properties: + activeOpsRef: + description: |- + ActiveOpsRef names the ControlPlaneOps currently allowed to act on this + control plane. Empty when none is running. + type: string + components: + description: |- + Components is the per-component readiness the phase is derived from + (§4.3), one entry per workload the managed install applies. It is empty + for a remote control plane, which has no components the operator owns. + Without it a Degraded phase says that something is wrong and not what. + items: + description: |- + ControlPlaneComponentStatus is one workload of a managed control plane and how + much of it is running. The phase is the worst verdict across these and the + readiness probe, and only an essential component at zero ready can make it + Unavailable. + properties: + desired: + description: |- + Desired is how many replicas the component should have. For the component + carrying its own operator it is that resource's own count, because a + FoundationDBCluster reports quorum rather than replicas. + format: int32 + minimum: 0 + type: integer + essential: + description: |- + Essential states whether this component at zero ready makes the control + plane Unavailable rather than Degraded. It is decided by the table in + §4.3 rather than by a user, and it is reported here so that a phase can be + explained without reading the operator's source. + type: boolean + name: + description: Name is the workload's name, as applied. + type: string + ready: + description: Ready is how many of them are. + format: int32 + minimum: 0 + type: integer + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + endpoint: + description: |- + Endpoint is the resolved management API base URL, derived in the local + case and echoed in the managed one. It is what every controller in the + operator reads to reach the control plane, so that one object answers + where it is. + type: string lastChecked: - description: LastChecked is the timestamp of the most recent FDB health probe. + description: LastChecked is when the readiness probe last ran. format: date-time type: string message: description: |- - Message contains a human-readable explanation of the current phase, - for example, the FDB error returned by the health endpoint. + Message is the reason the phase is what it is: one sentence, replaced as + the control plane moves, and never a log. On a failed probe it is the + control plane's own error rather than a paraphrase of it. type: string - phase: + observedGeneration: description: |- - Phase is Initializing while the control plane is not yet healthy, - and Available once the FDB health check passes. + ObservedGeneration is the generation the rest of this status was computed + from, so a stale status can be told from a current one. + format: int64 + type: integer + phase: + description: Phase is the operator's own view of the control plane. enum: - - Initializing + - Installing - Available + - Degraded + - Unavailable + type: string + step: + description: |- + Step is the position of the installation machine within Installing. The + rule repeats the ControlPlaneStep enum because a marker cannot reach a + field of the shared snapshot type. + properties: + deadline: + description: |- + Deadline is when that state expires, absent when it has none. It is an + absolute instant, so a state whose deadline passed while the controller + was down restores as already expired. + format: date-time + type: string + state: + description: |- + State is the state the machine was in. Empty means the resource has not + been reconciled yet, and restores to the graph's initial state. + type: string + type: object + x-kubernetes-validations: + - message: unknown step + rule: '!has(self.state) || self.state in [''ApplyingFoundationDB'',''AwaitingFoundationDB'',''ApplyingDatastore'',''ApplyingAPI'',''AwaitingAPI'']' + version: + description: Version is the version the management API reports. type: string type: object type: object diff --git a/operator/config/manifests/bases/simplyblock-operator.clusterserviceversion.yaml b/operator/config/manifests/bases/simplyblock-operator.clusterserviceversion.yaml index a7c6103c4..fdd6bd282 100644 --- a/operator/config/manifests/bases/simplyblock-operator.clusterserviceversion.yaml +++ b/operator/config/manifests/bases/simplyblock-operator.clusterserviceversion.yaml @@ -143,11 +143,24 @@ spec: kind: BackupPolicy name: backuppolicies.storage.simplyblock.io version: v1alpha1 - - description: ControlPlane is the Schema for the controlplanes API. + - description: ControlPlane is the simplyblock control plane for one Kubernetes + cluster, FoundationDB together with the management API, either installed + by the operator or already existing. It is a singleton named "simplyblock", + and it is the root of the ownership spine, so nothing else in this API group + reconciles meaningfully before it reports Available. displayName: Control Plane kind: ControlPlane name: controlplanes.storage.simplyblock.io - version: v1alpha1 + version: v1alpha2 + - description: ControlPlaneOps is a single operation performed against the control + plane. It runs to a terminal phase and stays afterward as the audit record + of what was done, with which parameters, and how it ended. Only one may be + active per control plane at a time, which the control plane's status.activeOpsRef + enforces. + displayName: Control Plane Ops + kind: ControlPlaneOps + name: controlplaneops.storage.simplyblock.io + version: v1alpha2 - description: ReplicationOps is the Schema for the replicationops API. displayName: Replication Ops kind: ReplicationOps diff --git a/operator/config/rbac/role.yaml b/operator/config/rbac/role.yaml index 07f03f8ac..a69d90757 100644 --- a/operator/config/rbac/role.yaml +++ b/operator/config/rbac/role.yaml @@ -121,6 +121,7 @@ rules: - apps resources: - daemonsets + - deployments - statefulsets verbs: - create @@ -130,6 +131,19 @@ rules: - patch - update - watch +- apiGroups: + - apps.foundationdb.org + resources: + - foundationdbbackups + - foundationdbclusters + verbs: + - create + - delete + - get + - list + - patch + - update + - watch - apiGroups: - authentication.k8s.io resources: @@ -187,6 +201,8 @@ rules: resources: - clusterrolebindings - clusterroles + - rolebindings + - roles verbs: - bind - create @@ -237,6 +253,7 @@ rules: - backupimports - backuppolicies - backuprestores + - controlplaneops - controlplanes - operatorops - replicationops @@ -270,6 +287,8 @@ rules: - backupimports/finalizers - backuppolicies/finalizers - backuprestores/finalizers + - controlplaneops/finalizers + - controlplanes/finalizers - operatorops/finalizers - replicationops/finalizers - replicationpairs/finalizers @@ -295,6 +314,7 @@ rules: - backuppolicies/status - backuprestores/status - clusterdeploymentconfigs/status + - controlplaneops/status - controlplanes/status - operatorops/status - replicationops/status diff --git a/operator/config/webhook/manifests.yaml b/operator/config/webhook/manifests.yaml index 115ca5a8e..23cdc2154 100644 --- a/operator/config/webhook/manifests.yaml +++ b/operator/config/webhook/manifests.yaml @@ -68,6 +68,26 @@ webhooks: resources: - persistentvolumeclaims sideEffects: None +- admissionReviewVersions: + - v1 + clientConfig: + service: + name: webhook-service + namespace: system + path: /validate-storage-simplyblock-io-v1alpha2-controlplaneops + failurePolicy: Fail + name: vcontrolplaneops.simplyblock.io + rules: + - apiGroups: + - storage.simplyblock.io + apiVersions: + - v1alpha2 + operations: + - CREATE + - DELETE + resources: + - controlplaneops + sideEffects: None - admissionReviewVersions: - v1 clientConfig: diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index 96b4b821a..0b7e58cd5 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -751,15 +751,41 @@ spec: document carries no field for it. properties: approved: + default: false description: |- Approved is the review gate. A document is expanded only once it is set, and is validated but otherwise inert before that, which is what makes reviewing a wrong document safe. + + It is defaulted and serialized rather than omitted when false, and the + two are the same requirement read twice. A reviewer has to see the gate + they are being asked to open, and the rules above have to find the field + they read: a bool omitted when false is a key the apiserver never stores, + so a rule reading it fails rather than reading false, and the first rule + guarding approval denied every approval there could ever be. The has() + guards are what carry documents written before the default existed. type: boolean cluster: description: Cluster is the StorageCluster to create. Ignored when ClusterRef is set. properties: + enableDriveFormat: + description: |- + EnableDriveFormat formats every device the document names before a storage + node takes it, which is how a drive carrying anything already is made + usable. + + It says what is wanted rather than how, because the how differs by device + class: an NVMe device is formatted to a 4K block size, and a logical block + device has its signatures wiped. One field covers both, so a document does + not have to know which class the expansion will resolve it to. + + It is on the document rather than defaulted further down because it is + destructive and the document is what somebody approves. A reviewer reading + a draft has to see that the drives it lists will be formatted, and be able + to strike it before approving; the cluster's own field is immutable once + the cluster exists, so a default nobody saw could not be undone either. + type: boolean enableFailureDomains: description: |- EnableFailureDomains opts the cluster into failure-domain mode, in which @@ -1020,9 +1046,9 @@ spec: type: object x-kubernetes-validations: - message: an approved deployment config is immutable - rule: '!oldSelf.approved || self == oldSelf' + rule: '!has(oldSelf.approved) || !oldSelf.approved || self == oldSelf' - message: approval cannot be withdrawn - rule: '!oldSelf.approved || self.approved' + rule: '!has(oldSelf.approved) || !oldSelf.approved || self.approved' - message: 'every group must name the same device class: all nvme or all block' rule: self.nodeSets.all(s, s.groups.all(g, !has(g.devices) || !has(g.devices.block))) @@ -1084,7 +1110,7 @@ spec: type: object x-kubernetes-validations: - message: unknown step - rule: '!has(self.state) || self.state in [''Validating'',''CreatingCluster'',''AwaitingCluster'',''CreatingNodes'']' + rule: '!has(self.state) || self.state in [''Validating'',''CreatingCluster'',''AwaitingCluster'',''CreatingNodes'',''Activating'']' type: object type: object served: true @@ -1114,6 +1140,8 @@ spec: kind: ControlPlane listKind: ControlPlaneList plural: controlplanes + shortNames: + - cp singular: controlplane scope: Namespaced versions: @@ -1163,9 +1191,11 @@ spec: image: description: |- Image is the container image used for all simplyblock control-plane and - storage-node workloads (e.g. quay.io/simplyblock-io/simplyblock:26.2.2). + storage-node workloads (e.g., `quay.io/simplyblock-io/simplyblock:26.2.2`). StorageNodeSet CRs that omit spec.clusterImage inherit this value. - Must reference one of the trusted registries (quay.io/simplyblock-io, docker.io/simplyblock, public.ecr.aws/simply-block); digest pinning (@sha256:...) is recommended. + Must reference one of the trusted registries (`quay.io/simplyblock-io`, + `docker.io/simplyblock`, `public.ecr.aws/simply-block`). Digest pinning + (@sha256:...) is recommended. pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ type: string type: object @@ -1182,7 +1212,7 @@ spec: message: description: |- Message contains a human-readable explanation of the current phase, - for example the FDB error returned by the health endpoint. + for example, the FDB error returned by the health endpoint. type: string phase: description: |- @@ -1199,14 +1229,21 @@ spec: subresources: status: {} - additionalPrinterColumns: - - description: Initializing while FDB is not ready; Available once the control - plane is operational - jsonPath: .status.phase + - jsonPath: .status.phase name: Phase type: string - - description: Human-readable status detail - jsonPath: .status.message + - jsonPath: .status.step.state + name: Step + type: string + - jsonPath: .status.endpoint + name: Endpoint + type: string + - jsonPath: .status.version + name: Version + type: string + - jsonPath: .status.message name: Message + priority: 1 type: string - jsonPath: .metadata.creationTimestamp name: Age @@ -1215,9 +1252,11 @@ spec: schema: openAPIV3Schema: description: |- - ControlPlane is a singleton resource (one per namespace, named "simplyblock") - that reflects the readiness of the simplyblock control plane. It is created - automatically by the Helm chart and should not be created or deleted manually. + ControlPlane is the simplyblock control plane for one Kubernetes cluster: + FoundationDB together with the management API, either installed by the + operator or already existing. It is a singleton named `simplyblock`, and it is + the root of the ownership spine: nothing else in this API group reconciles + meaningfully before it reports Available. properties: apiVersion: description: |- @@ -1237,51 +1276,415 @@ spec: metadata: type: object spec: - description: ControlPlaneSpec holds configuration for the singleton ControlPlane - resource. + description: |- + ControlPlaneSpec is the desired state of the simplyblock control plane for one + namespace. properties: source: description: |- - Source says where the control plane comes from. It replaces the top-level - image field of v1alpha1, which conflated the control plane's own image with - the default every StorageNodeSet inherited. + Source selects whether this cluster hosts its control plane or is managed + by one elsewhere. Switching a live deployment between the two is not a + reconfiguration, because the clusters and their volumes live in the + FoundationDB behind the old one, so which mode is chosen is frozen at + creation. What is inside the chosen mode stays editable. + + Both rules are declared on ControlPlaneSource rather than here. See the + type for why. properties: - managed: - description: Managed is the control plane the operator installs - and owns. + local: + description: Local is a control plane the operator installs. properties: + foundationDB: + description: FoundationDB sizes the FoundationDB the management + API stores its state in. + properties: + replicas: + default: 3 + description: |- + Replicas is the number of coordinators. Three is the smallest count that + survives one loss, which is why it is the default. + format: int32 + minimum: 1 + type: integer + resources: + description: Resources sets requests and limits for the + coordinator pods. + properties: + claims: + description: |- + Claims lists the names of resources, defined in spec.resourceClaims, + that are used by this container. + + This field depends on the + DynamicResourceAllocation feature gate. + + This field is immutable. It can only be set for containers. + items: + description: ResourceClaim references one entry + in PodSpec.ResourceClaims. + properties: + name: + description: |- + Name must match the name of one entry in pod.spec.resourceClaims of + the Pod where this field is used. It makes that resource available + inside a container. + type: string + request: + description: |- + Request is the name chosen for a request in the referenced claim. + If empty, everything from the claim is made available, otherwise + only the result of this request. + type: string + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + limits: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Limits describes the maximum amount of compute resources allowed. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + requests: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Requests describes the minimum amount of compute resources required. + If Requests is omitted for a container, it defaults to Limits if that is explicitly specified, + otherwise to an implementation-defined value. Requests cannot exceed Limits. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + type: object + storageClassName: + description: |- + StorageClassName is the class the coordinators' volumes are provisioned + from. It cannot be a class this operator provides, because the control + plane has to exist before any simplyblock volume can. + type: string + type: object image: description: |- - Image is the container image used for the simplyblock control-plane - workloads (e.g., quay.io/simplyblock-io/simplyblock:26.2.2). - Must reference one of the trusted registries (`quay.io/simplyblock-io`, `docker.io/simplyblock`, `public.ecr.aws/simply-block`); digest pinning (@sha256:...) is recommended. + Image is the management API and control-plane image. + Must reference one of the trusted registries (`quay.io/simplyblock-io`, + `docker.io/simplyblock`, `public.ecr.aws/simply-block`); digest pinning + (@sha256:...) is recommended. pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ type: string + imagePullPolicy: + default: IfNotPresent + description: ImagePullPolicy controls when that image is pulled. + enum: + - Always + - Never + - IfNotPresent + type: string + nodeSelector: + additionalProperties: + type: string + description: |- + NodeSelector pins every pod the operator installs for the control plane. + It is a selector rather than an affinity term because that is what the + chart it replaces took, and a deployment migrating off the chart has the + value already written down. + type: object + replicas: + default: 2 + description: |- + Replicas is the number of management API instances. Two is what the chart + ships and what the phases assume: a single instance makes Degraded + unreachable for this component and every restart an outage (§5.1). + format: int32 + minimum: 1 + type: integer + resources: + description: Resources sets requests and limits for the management + API pods. + properties: + claims: + description: |- + Claims lists the names of resources, defined in spec.resourceClaims, + that are used by this container. + + This field depends on the + DynamicResourceAllocation feature gate. + + This field is immutable. It can only be set for containers. + items: + description: ResourceClaim references one entry in PodSpec.ResourceClaims. + properties: + name: + description: |- + Name must match the name of one entry in pod.spec.resourceClaims of + the Pod where this field is used. It makes that resource available + inside a container. + type: string + request: + description: |- + Request is the name chosen for a request in the referenced claim. + If empty, everything from the claim is made available, otherwise + only the result of this request. + type: string + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + limits: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Limits describes the maximum amount of compute resources allowed. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + requests: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Requests describes the minimum amount of compute resources required. + If Requests is omitted for a container, it defaults to Limits if that is explicitly specified, + otherwise to an implementation-defined value. Requests cannot exceed Limits. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + type: object + tolerations: + description: |- + Tolerations are applied to every pod the operator installs for the control + plane. + items: + description: |- + The pod this Toleration is attached to tolerates any taint that matches + the triple using the matching operator . + properties: + effect: + description: |- + Effect indicates the taint effect to match. Empty means match all taint effects. + When specified, allowed values are NoSchedule, PreferNoSchedule and NoExecute. + type: string + key: + description: |- + Key is the taint key that the toleration applies to. Empty means match all taint keys. + If the key is empty, operator must be Exists; this combination means to match all values and all keys. + type: string + operator: + description: |- + Operator represents a key's relationship to the value. + Valid operators are Exists, Equal, Lt, and Gt. Defaults to Equal. + Exists is equivalent to wildcard for value, so that a pod can + tolerate all taints of a particular category. + Lt and Gt perform numeric comparisons (requires feature gate TaintTolerationComparisonOperators). + type: string + tolerationSeconds: + description: |- + TolerationSeconds represents the period of time the toleration (which must be + of effect NoExecute, otherwise this field is ignored) tolerates the taint. By default, + it is not set, which means tolerate the taint forever (do not evict). Zero and + negative values will be treated as 0 (evict immediately) by the system. + format: int64 + type: integer + value: + description: |- + Value is the taint value the toleration matches to. + If the operator is Exists, the value should be empty, otherwise just a regular string. + type: string + type: object + type: array + required: + - image + type: object + managed: + description: Managed is a control plane that already exists. + properties: + caBundleSecretRef: + description: |- + CABundleSecretRef names a Secret holding the CA certificate the endpoint + is verified against. Absent means the system trust store. + properties: + name: + default: "" + description: |- + Name of the referent. + This field is effectively required, but due to backwards compatibility is + allowed to be empty. Instances of this type with an empty value here are + almost certainly wrong. + More info: https://kubernetes.io/docs/concepts/overview/working-with-objects/names/#names + type: string + type: object + x-kubernetes-map-type: atomic + credentialsSecretRef: + description: |- + CredentialsSecretRef names a Secret in this namespace holding the bearer + token the operator authenticates with. It is a reference rather than a + field because a token in a spec is a token in every `kubectl get -o yaml`. + + Absent means the endpoint is reached without one, which is the in-cluster + case: a control plane the Helm chart installed answers on a ClusterIP + Service in this namespace and does not require a token for the readiness + read. Naming a Secret that does not exist stays an error, because naming + one is a statement that the control plane needs it. + properties: + name: + default: "" + description: |- + Name of the referent. + This field is effectively required, but due to backwards compatibility is + allowed to be empty. Instances of this type with an empty value here are + almost certainly wrong. + More info: https://kubernetes.io/docs/concepts/overview/working-with-objects/names/#names + type: string + type: object + x-kubernetes-map-type: atomic + endpoint: + description: |- + Endpoint is the management API's base URL. It is validated against the + same outbound-URL guard every other outbound endpoint in this group uses, + so a loopback or link-local address is rejected. + pattern: ^https?://[a-zA-Z0-9.-]+(:[0-9]{1,5})?(/.*)?$ + type: string + required: + - endpoint type: object type: object + x-kubernetes-validations: + - message: set exactly one of local or managed + rule: '(has(self.local) ? 1 : 0) + (has(self.managed) ? 1 : 0) == + 1' + - message: 'spec.source is immutable: a control plane the operator + installed and one it did not are different deployments, and the + clusters and their volumes live in the FoundationDB behind the + old one' + rule: has(self.local) == has(oldSelf.local) && has(self.managed) + == has(oldSelf.managed) + required: + - source type: object status: - description: |- - ControlPlaneStatus reflects the observed readiness of the simplyblock - control plane (FDB + management API). + description: ControlPlaneStatus is the observed state of the control plane. properties: + activeOpsRef: + description: |- + ActiveOpsRef names the ControlPlaneOps currently allowed to act on this + control plane. Empty when none is running. + type: string + components: + description: |- + Components is the per-component readiness the phase is derived from + (§4.3), one entry per workload the managed install applies. It is empty + for a remote control plane, which has no components the operator owns. + Without it a Degraded phase says that something is wrong and not what. + items: + description: |- + ControlPlaneComponentStatus is one workload of a managed control plane and how + much of it is running. The phase is the worst verdict across these and the + readiness probe, and only an essential component at zero ready can make it + Unavailable. + properties: + desired: + description: |- + Desired is how many replicas the component should have. For the component + carrying its own operator it is that resource's own count, because a + FoundationDBCluster reports quorum rather than replicas. + format: int32 + minimum: 0 + type: integer + essential: + description: |- + Essential states whether this component at zero ready makes the control + plane Unavailable rather than Degraded. It is decided by the table in + §4.3 rather than by a user, and it is reported here so that a phase can be + explained without reading the operator's source. + type: boolean + name: + description: Name is the workload's name, as applied. + type: string + ready: + description: Ready is how many of them are. + format: int32 + minimum: 0 + type: integer + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + endpoint: + description: |- + Endpoint is the resolved management API base URL, derived in the local + case and echoed in the managed one. It is what every controller in the + operator reads to reach the control plane, so that one object answers + where it is. + type: string lastChecked: - description: LastChecked is the timestamp of the most recent FDB health - probe. + description: LastChecked is when the readiness probe last ran. format: date-time type: string message: description: |- - Message contains a human-readable explanation of the current phase, - for example, the FDB error returned by the health endpoint. + Message is the reason the phase is what it is: one sentence, replaced as + the control plane moves, and never a log. On a failed probe it is the + control plane's own error rather than a paraphrase of it. type: string - phase: + observedGeneration: description: |- - Phase is Initializing while the control plane is not yet healthy, - and Available once the FDB health check passes. + ObservedGeneration is the generation the rest of this status was computed + from, so a stale status can be told from a current one. + format: int64 + type: integer + phase: + description: Phase is the operator's own view of the control plane. enum: - - Initializing + - Installing - Available + - Degraded + - Unavailable + type: string + step: + description: |- + Step is the position of the installation machine within Installing. The + rule repeats the ControlPlaneStep enum because a marker cannot reach a + field of the shared snapshot type. + properties: + deadline: + description: |- + Deadline is when that state expires, absent when it has none. It is an + absolute instant, so a state whose deadline passed while the controller + was down restores as already expired. + format: date-time + type: string + state: + description: |- + State is the state the machine was in. Empty means the resource has not + been reconciled yet, and restores to the graph's initial state. + type: string + type: object + x-kubernetes-validations: + - message: unknown step + rule: '!has(self.state) || self.state in [''ApplyingFoundationDB'',''AwaitingFoundationDB'',''ApplyingDatastore'',''ApplyingAPI'',''AwaitingAPI'']' + version: + description: Version is the version the management API reports. type: string type: object type: object @@ -5569,6 +5972,8 @@ spec: enum: - Pending - Creating + - Provisioning + - Activating - Online - Degraded - Unavailable @@ -9217,6 +9622,7 @@ rules: - apps resources: - daemonsets + - deployments - statefulsets verbs: - create @@ -9226,6 +9632,19 @@ rules: - patch - update - watch +- apiGroups: + - apps.foundationdb.org + resources: + - foundationdbbackups + - foundationdbclusters + verbs: + - create + - delete + - get + - list + - patch + - update + - watch - apiGroups: - authentication.k8s.io resources: @@ -9283,6 +9702,8 @@ rules: resources: - clusterrolebindings - clusterroles + - rolebindings + - roles verbs: - bind - create @@ -9333,6 +9754,7 @@ rules: - backupimports - backuppolicies - backuprestores + - controlplaneops - controlplanes - operatorops - replicationops @@ -9366,6 +9788,8 @@ rules: - backupimports/finalizers - backuppolicies/finalizers - backuprestores/finalizers + - controlplaneops/finalizers + - controlplanes/finalizers - operatorops/finalizers - replicationops/finalizers - replicationpairs/finalizers @@ -9391,6 +9815,7 @@ rules: - backuppolicies/status - backuprestores/status - clusterdeploymentconfigs/status + - controlplaneops/status - controlplanes/status - operatorops/status - replicationops/status @@ -10037,6 +10462,26 @@ webhooks: resources: - persistentvolumeclaims sideEffects: None +- admissionReviewVersions: + - v1 + clientConfig: + service: + name: simplyblock-operator-webhook-service + namespace: simplyblock-operator-system + path: /validate-storage-simplyblock-io-v1alpha2-controlplaneops + failurePolicy: Fail + name: vcontrolplaneops.simplyblock.io + rules: + - apiGroups: + - storage.simplyblock.io + apiVersions: + - v1alpha2 + operations: + - CREATE + - DELETE + resources: + - controlplaneops + sideEffects: None - admissionReviewVersions: - v1 clientConfig: diff --git a/operator/docs/designs/crd-redesign/design-controlplane.md b/operator/docs/designs/crd-redesign/design-controlplane.md index eba5a6115..b30c890ba 100644 --- a/operator/docs/designs/crd-redesign/design-controlplane.md +++ b/operator/docs/designs/crd-redesign/design-controlplane.md @@ -865,7 +865,7 @@ assume and nothing exercises. | `spec.image` | `spec.source.managed.image` (§5.1) | Spec regrouping, and the field stops doubling as the storage-node default (see below) | | No way to name an external control plane | `spec.source.external` (§5.2) | Additive, and it is the half of `design-crd-model.md` §6 that does not exist today | | The endpoint is `SIMPLYBLOCK_WEBAPI_BASE_URL` | `status.endpoint` (§3.3) | Behavioral. Every caller of `webapi.NewClient()` moves, which is every controller in the operator | -| The chart installs the control plane | The operator installs it (§5.1) | The largest piece of work here, and it cannot be a flag day (§12, Q2) | +| The chart installs the control plane | The operator installs it (§5.1) | The largest piece of work here. `controlplane.managedByOperator` is the flag, and it defaults to the operator. Adopting a control plane the chart already installed is the open half (§12, Q2) | | `status.phase` untyped, two values | `ControlPlanePhase`, four values (§3.3) | `Degraded` and `Unavailable` are both new, and together they separate a control plane that is impaired from one that is not answering. `Ready` becomes `Available`, which is the opposite of `Unavailable` where `Ready` was not | | No step field | `status.step` (§4.2) | Additive. A stalled install currently reports one message and no position | | No `observedGeneration` | Present (§3.3) | Required by `design-crd-model.md` §7.9 | @@ -898,13 +898,36 @@ from review history. §3.1 is where the answer went: a Kubernetes cluster holds `ControlPlane`, which is the limit the operator's own cluster-scoped objects already impose on the operator. -**Q2: How the install moves from the chart to the operator.** §5.1 has the -operator apply what the chart applies today, and both cannot own the same objects -at once. The candidates are a chart flag that stops rendering the templates once -the operator is capable of applying them, an adoption pass where the operator -takes ownership of objects the chart already created, and leaving the chart in -place for existing deployments while new ones use the operator. Nothing here -settles it, and it is the reason §5.1's work is larger than its specification. +**Q2 is settled for a fresh install and open for an existing one.** +`controlplane.managedByOperator` is the chart flag, and it defaults to the +operator: a new deployment gets the control plane §5.1 describes, and the chart +renders only the `ControlPlane` object beside the FoundationDB CRDs, the +Prometheus configuration, and the log-collector RBAC. Setting it to false hands +the templates back, unchanged, which is what a deployment does when it needs +something the spec cannot yet express. + +What is open is the transition for a deployment already running a chart-installed +control plane. The apply is a server-side apply under a stable field manager, so +it takes over the objects a Helm release created rather than failing on them, but +nothing verifies the handover or strips the release's claim afterward, and §5.1's +ownership spine is only real once it has. That is the shape +[`design-simplyblockdriver.md`](design-simplyblockdriver.md) §4.3 gives the CSI +driver's adoption, and it is the half of this question still to answer. + +**The install covers a base control plane rather than everything the chart +renders.** What it applies is the FoundationDB half, the object store, and the +management API with the services beside it, which is what the chart renders with +observability disabled. Graylog, Grafana, Thanos, the document store behind them, +and the log collector stay with the chart: none appears in a step of §4.2's +machine, every one is non-essential in §4.3's table, and they are gated behind one +chart value this kind has no field for. + +**TLS is the one configuration the install cannot express.** The chart serves the +control plane over TLS behind `tls.enabled`, and `ManagedControlPlane` has no +field for it, since §5.1 settles which issuer is detected rather than whether a +deployment wants one. The chart refuses `managedByOperator` together with +`tls.enabled` rather than installing a control plane in plaintext, and closing +that gap is a field on this kind. **Q3: Whether backup belongs to the action or to the spec.** `FoundationDBBackup` describes a continuous backup, carrying a `backupState` and a diff --git a/operator/internal/controller/controlplane_controller.go b/operator/internal/controller/controlplane_controller.go deleted file mode 100644 index cc9c8a8df..000000000 --- a/operator/internal/controller/controlplane_controller.go +++ /dev/null @@ -1,164 +0,0 @@ -/* -Copyright 2025. - -Licensed under the Apache License, Version 2.0 (the "License"); -you may not use this file except in compliance with the License. -You may obtain a copy of the License at - - http://www.apache.org/licenses/LICENSE-2.0 - -Unless required by applicable law or agreed to in writing, software -distributed under the License is distributed on an "AS IS" BASIS, -WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. -See the License for the specific language governing permissions and -limitations under the License. -*/ - -package controller - -import ( - "context" - "fmt" - "net/http" - "time" - - corev1 "k8s.io/api/core/v1" - metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" - "k8s.io/apimachinery/pkg/runtime" - "k8s.io/client-go/tools/events" - ctrl "sigs.k8s.io/controller-runtime" - "sigs.k8s.io/controller-runtime/pkg/builder" - "sigs.k8s.io/controller-runtime/pkg/client" - logf "sigs.k8s.io/controller-runtime/pkg/log" - "sigs.k8s.io/controller-runtime/pkg/predicate" - - simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" - "github.com/simplyblock/simplyblock-operator/internal/webapi" -) - -// SingletonControlPlaneName is the fixed name of the singleton ControlPlane CR -// created by the Helm chart. The controller ignores any CR with a different name. -const SingletonControlPlaneName = "simplyblock" - -// controlPlaneRequeueInterval is how often the FDB health check is repeated. -const controlPlaneRequeueInterval = 30 * time.Second - -// The phases this controller records. They are v1alpha2's vocabulary, whose Enum -// admits Available where v1alpha1 said Ready, and they live beside the reconciler -// that writes them rather than among the shared cluster constants: no other -// kind's phase is spelled from this pair. -const ( - controlPlanePhaseInitializing = "Initializing" - controlPlanePhaseAvailable = "Available" -) - -const ( - // eventReasonFDBReady is emitted when the FDB health check recovers after a - // prior failure (phase transitions from Initializing → Ready). - eventReasonCPFDBReady = "FDBReady" - - // eventReasonFDBNotReady is emitted when the FDB health check first fails - // (phase transitions from Ready → Initializing, or on the initial probe). - eventReasonCPFDBNotReady = "FDBNotReady" -) - -// ControlPlaneReconciler reconciles the singleton ControlPlane object. -type ControlPlaneReconciler struct { - client.Client - Scheme *runtime.Scheme - Recorder events.EventRecorder -} - -// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=controlplanes,verbs=get;list;watch;create;update;patch;delete -// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=controlplanes/status,verbs=get;update;patch - -func (r *ControlPlaneReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) { - log := logf.FromContext(ctx) - - if req.Name != SingletonControlPlaneName { - log.Info("ignoring non-singleton ControlPlane CR", "name", req.Name) - return ctrl.Result{}, nil - } - - // v1alpha2 is the stored version. Reading the singleton at v1alpha1 would be - // answered only by the conversion webhook, which a fresh install does not - // deploy, and the cache backing this read lists empty rather than failing: - // the reconciler would silently own an object it never sees. - cp := &simplyblockv1alpha2.ControlPlane{} - if err := r.Get(ctx, req.NamespacedName, cp); err != nil { - return ctrl.Result{}, client.IgnoreNotFound(err) - } - - now := metav1.Now() - orig := cp.DeepCopy() - - prevPhase := cp.Status.Phase - - apiClient := webapi.NewClient() - body, status, err := apiClient.Do(ctx, http.MethodGet, "/api/v2/_meta/ready", nil) - if err != nil || status >= 300 { - msg := string(body) - if err != nil { - msg = err.Error() - } else { - msg = fmt.Sprintf("status=%d: %s", status, msg) - } - // The probe repeats for the life of the cluster, so a readiness that has - // not changed is logged at debug. What is worth an operator's attention - // is the transition, which is also what the event below reports. - if prevPhase != controlPlanePhaseInitializing { - log.Info("control plane not ready", "reason", msg) - } else { - log.V(1).Info("control plane still not ready", "reason", msg) - } - - cp.Status.Phase = controlPlanePhaseInitializing - cp.Status.Message = msg - cp.Status.LastChecked = &now - - if err := r.Status().Patch(ctx, cp, client.MergeFrom(orig)); err != nil { - log.Error(err, "failed to patch ControlPlane status") - } - - // Only emit event on transition to avoid spamming every 30 s. - if prevPhase != controlPlanePhaseInitializing { - r.Recorder.Eventf(cp, nil, corev1.EventTypeWarning, eventReasonCPFDBNotReady, eventReasonCPFDBNotReady, "FDB health check failed: %s", msg) - } - return ctrl.Result{RequeueAfter: controlPlaneRequeueInterval}, nil - } - - cp.Status.Phase = controlPlanePhaseAvailable - cp.Status.Message = "" - cp.Status.LastChecked = &now - - if err := r.Status().Patch(ctx, cp, client.MergeFrom(orig)); err != nil { - log.Error(err, "failed to patch ControlPlane status") - } - - // Emit recovery event only when transitioning from Initializing → Ready. - if prevPhase == controlPlanePhaseInitializing { - r.Recorder.Eventf(cp, nil, corev1.EventTypeNormal, eventReasonCPFDBReady, eventReasonCPFDBReady, "FDB health check passed; control plane is ready") - } - - if prevPhase != controlPlanePhaseAvailable { - log.Info("control plane ready") - } else { - log.V(1).Info("control plane still ready") - } - return ctrl.Result{RequeueAfter: controlPlaneRequeueInterval}, nil -} - -// SetupWithManager sets up the controller with the Manager. -func (r *ControlPlaneReconciler) SetupWithManager(mgr ctrl.Manager) error { - return ctrl.NewControllerManagedBy(mgr). - // Every probe stamps status.lastChecked, and an unfiltered watch turns - // that write into another reconcile, which probes and stamps again. The - // loop settles only because the second stamp lands in the same second and - // patches nothing, which costs a second probe of the control plane every - // interval and reports it twice. The generation does not move on a status - // write, so this leaves the requeue below as the probe's only clock while - // an edit to the spec still arrives at once. - For(&simplyblockv1alpha2.ControlPlane{}, builder.WithPredicates(predicate.GenerationChangedPredicate{})). - Named("controlplane"). - Complete(r) -} diff --git a/operator/internal/controller/controlplane_controller_test.go b/operator/internal/controller/controlplane_controller_test.go deleted file mode 100644 index 93313e9dd..000000000 --- a/operator/internal/controller/controlplane_controller_test.go +++ /dev/null @@ -1,185 +0,0 @@ -// Tests for the ControlPlane reconciler: the FDB readiness probe it runs and the -// phase it records from the result. -// -// They live here rather than beside another controller's tests because the -// singleton ControlPlane is the one object this reconciler owns, and its phase -// vocabulary is version-specific: v1alpha2 says Available where v1alpha1 said -// Ready, so a test asserting the wrong word passes against a status the API -// server would refuse. - -package controller - -import ( - "context" - "net/http" - "testing" - - "github.com/go-logr/logr" - metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" - "k8s.io/client-go/tools/events" - ctrl "sigs.k8s.io/controller-runtime" - "sigs.k8s.io/controller-runtime/pkg/client" - "sigs.k8s.io/controller-runtime/pkg/client/fake" - logf "sigs.k8s.io/controller-runtime/pkg/log" - - simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" - webapimock "github.com/simplyblock/simplyblock-operator/internal/webapi/mock" -) - -// Regression: 2026-09-11-controlplane-reconciler-reads-retired-version. The -// reconciler read and watched the singleton at v1alpha1 after v1alpha2 became -// the stored version. A fresh install deploys no conversion webhook, so the -// object it owns was invisible to it and no phase was ever recorded. The phase it records is asserted alongside, because v1alpha2's Enum -// admits Available and not the Ready this wrote. -func TestControlPlaneReconcileRecordsAvailablePhase(t *testing.T) { - mock := webapimock.NewSpecServerFromFile(t, "../../../shared/openapi.json", false) - defer mock.Close() - mock.Register(http.MethodGet, "/api/v2/_meta/ready", webapimock.RouteResponse{ - Status: http.StatusOK, - Body: `{"status":"ok"}`, - Headers: map[string]string{"Content-Type": "application/json"}, - }) - t.Setenv("SIMPLYBLOCK_WEBAPI_BASE_URL", mock.URL()) - - cp := &simplyblockv1alpha2.ControlPlane{ - ObjectMeta: metav1.ObjectMeta{ - Name: SingletonControlPlaneName, - Namespace: "simplyblock", - }, - } - - scheme := newTestScheme(t) - cl := fake.NewClientBuilder(). - WithScheme(scheme). - WithStatusSubresource(&simplyblockv1alpha2.ControlPlane{}). - WithObjects(cp). - Build() - - r := &ControlPlaneReconciler{ - Client: cl, - Scheme: scheme, - Recorder: events.NewFakeRecorder(8), - } - - key := client.ObjectKeyFromObject(cp) - if _, err := r.Reconcile(context.Background(), ctrl.Request{NamespacedName: key}); err != nil { - t.Fatalf("Reconcile returned error: %v", err) - } - - var got simplyblockv1alpha2.ControlPlane - if err := cl.Get(context.Background(), key, &got); err != nil { - t.Fatalf("reading back the ControlPlane: %v", err) - } - if got.Status.Phase != controlPlanePhaseAvailable { - t.Fatalf("expected phase %q, got %q", controlPlanePhaseAvailable, got.Status.Phase) - } - if got.Status.LastChecked == nil { - t.Fatal("expected the probe to stamp status.lastChecked") - } -} - -// recordingSink is a logr sink that keeps what was logged at info level, so a -// test can assert what a steady state is quiet about. Only the levels this -// controller uses are recorded; everything else is discarded. -type recordingSink struct { - verbosity int - messages *[]string -} - -func (s recordingSink) Init(logr.RuntimeInfo) {} -func (s recordingSink) Enabled(int) bool { return true } -func (s recordingSink) WithName(string) logr.LogSink { return s } - -func (s recordingSink) Info(level int, msg string, _ ...any) { - // controller-runtime logs debug at V(1) and above, and info at V(0). - if level+s.verbosity == 0 { - *s.messages = append(*s.messages, msg) - } -} - -func (s recordingSink) Error(_ error, msg string, _ ...any) { - *s.messages = append(*s.messages, msg) -} - -func (s recordingSink) WithValues(...any) logr.LogSink { return s } - -func (s recordingSink) V(level int) logr.LogSink { - return recordingSink{verbosity: s.verbosity + level, messages: s.messages} -} - -// probeRepeatedly runs ten reconciles against a control plane answering with the -// given status, and returns everything they logged at info. Ten because that is -// what the plan's transition rows count: one is a transition and nine are the -// steady state, which is where a per-probe line would show up. -func probeRepeatedly(t *testing.T, status int, body string) []string { - t.Helper() - - mock := webapimock.NewSpecServerFromFile(t, "../../../shared/openapi.json", false) - defer mock.Close() - mock.Register(http.MethodGet, "/api/v2/_meta/ready", webapimock.RouteResponse{ - Status: status, - Body: body, - Headers: map[string]string{"Content-Type": "application/json"}, - }) - t.Setenv("SIMPLYBLOCK_WEBAPI_BASE_URL", mock.URL()) - - cp := &simplyblockv1alpha2.ControlPlane{ - ObjectMeta: metav1.ObjectMeta{ - Name: SingletonControlPlaneName, - Namespace: "simplyblock", - }, - } - scheme := newTestScheme(t) - cl := fake.NewClientBuilder(). - WithScheme(scheme). - WithStatusSubresource(&simplyblockv1alpha2.ControlPlane{}). - WithObjects(cp). - Build() - r := &ControlPlaneReconciler{Client: cl, Scheme: scheme, Recorder: events.NewFakeRecorder(16)} - - var logged []string - ctx := logf.IntoContext(context.Background(), logr.New(recordingSink{messages: &logged})) - key := client.ObjectKeyFromObject(cp) - - for i := range 10 { - if _, err := r.Reconcile(ctx, ctrl.Request{NamespacedName: key}); err != nil { - t.Fatalf("reconcile %d: %v", i, err) - } - } - return logged -} - -// countMessages returns how many of the logged lines carry exactly this message. -func countMessages(logged []string, message string) int { - var count int - for _, msg := range logged { - if msg == message { - count++ - } - } - return count -} - -// Regression: 2026-09-11-controlplane-logs-every-probe. The probe repeats every -// 30 seconds for the life of the cluster, and a readiness that has not changed -// was announced at info on every one of them: about six thousand lines a day per -// operator, all of them saying what the line before said. The transition is the -// event, and it is what the log says too. -func TestARepeatedReadyProbeIsNotAnnouncedAgain(t *testing.T) { - logged := probeRepeatedly(t, http.StatusOK, `{"status":"ok"}`) - - if got := countMessages(logged, "control plane ready"); got != 1 { - t.Fatalf("expected the readiness to be announced once, got %d in %v", got, logged) - } -} - -// The same in the other direction, which is the one an operator is more likely to -// be reading: a control plane that has been down for an hour says so once, and -// the hour of probes behind it is at debug. -func TestARepeatedFailingProbeIsNotAnnouncedAgain(t *testing.T) { - logged := probeRepeatedly(t, http.StatusServiceUnavailable, `{"status":"fdb unavailable"}`) - - if got := countMessages(logged, "control plane not ready"); got != 1 { - t.Fatalf("expected the failure to be announced once, got %d in %v", got, logged) - } -} diff --git a/operator/internal/controllers/cluster/controlplane.go b/operator/internal/controllers/cluster/controlplane.go index 8cef938f3..0c0b79179 100644 --- a/operator/internal/controllers/cluster/controlplane.go +++ b/operator/internal/controllers/cluster/controlplane.go @@ -26,7 +26,9 @@ import ( "errors" "fmt" "net/http" + "sync" + "github.com/simplyblock/simplyblock-operator/internal/controllers/controlplane" "github.com/simplyblock/simplyblock-operator/internal/cpinformer/subscriptions" "github.com/simplyblock/simplyblock-operator/internal/utils" "github.com/simplyblock/simplyblock-operator/internal/webapi" @@ -81,12 +83,29 @@ type ControlPlane interface { CancelTask(ctx context.Context, clusterID, taskID string) error } -// httpControlPlane is the ControlPlane the operator runs with: the shared -// webapi client, with one method per endpoint of §9. -type httpControlPlane struct{ client *webapi.Client } +// httpControlPlane is the ControlPlane the operator runs with: one method per +// endpoint of §9, over a client whose address comes from the ControlPlane object. +type httpControlPlane struct { + // client is the startup client, built from the environment. It is what a + // call uses until the ControlPlane publishes an endpoint. + client *webapi.Client + + // resolve answers where the control plane is, per call. Nil means the + // startup client is the only one. + resolve controlplane.EndpointResolver + + // mu guards resolved, which is rebuilt when the published endpoint changes. + mu sync.Mutex + resolved *webapi.Client +} // NewControlPlane returns the HTTP-backed control-plane surface. -func NewControlPlane() ControlPlane { return &httpControlPlane{client: webapi.NewClient()} } +// +// The resolver may be nil, which is what a test passes: calls then go to the +// startup client and nothing reads a ControlPlane object. +func NewControlPlane(resolve controlplane.EndpointResolver) ControlPlane { + return &httpControlPlane{client: webapi.NewClient(), resolve: resolve} +} func (c *httpControlPlane) Ready(ctx context.Context) error { _, err := c.call(ctx, http.MethodGet, "/api/v2/_meta/ready", nil) @@ -228,7 +247,7 @@ func (c *httpControlPlane) post(ctx context.Context, path string, body any) erro func (c *httpControlPlane) call( ctx context.Context, method, path string, body any, ) ([]byte, error) { - response, status, err := c.client.Do(ctx, method, path, body) + response, status, err := c.clientFor(ctx).Do(ctx, method, path, body) if err != nil { return nil, fmt.Errorf("%s %s: %w", method, path, err) } @@ -249,3 +268,31 @@ type ControlPlaneError struct { func (e *ControlPlaneError) Error() string { return fmt.Sprintf("the control plane answered %d: %s", e.Status, e.Body) } + +// clientFor is the client this call goes out on. +// +// The endpoint comes from ControlPlane.status.endpoint where the object has +// published one, which is what makes a remote control plane reachable: the +// client built at startup resolves SIMPLYBLOCK_WEBAPI_BASE_URL or the in-cluster +// default, and neither is where somebody else's control plane is +// (design-controlplane.md §3.3). +// +// With no resolver, or with one that answers nothing, the startup client is used +// unchanged. That is what keeps this additive: a deployment whose ControlPlane +// has not published an endpoint behaves as it did before. +func (c *httpControlPlane) clientFor(ctx context.Context) *webapi.Client { + if c.resolve == nil { + return c.client + } + endpoint := c.resolve(ctx) + if endpoint == "" || endpoint == c.client.BaseURL { + return c.client + } + + c.mu.Lock() + defer c.mu.Unlock() + if c.resolved == nil || c.resolved.BaseURL != endpoint { + c.resolved = webapi.NewClient(endpoint) + } + return c.resolved +} diff --git a/operator/internal/controllers/controlplane/apply.go b/operator/internal/controllers/controlplane/apply.go new file mode 100644 index 000000000..b3823bf36 --- /dev/null +++ b/operator/internal/controllers/controlplane/apply.go @@ -0,0 +1,178 @@ +// How the install writes an object, and which of the two ownership mechanisms +// that object gets. +// +// Every write is a server-side apply under a stable field manager. That is what +// makes re-entering a step a no-op (design-controlplane.md §4.2): a step whose +// apply never landed re-applies to the same result, which is why the machine +// carries no triggered flag. It is also what lets the same call take over an +// object a Helm release created, since an apply reconciles field ownership +// rather than failing on a conflict. +// +// A namespaced object becomes a child of the ControlPlane by controller +// reference and goes with the garbage collector when the ControlPlane is +// deleted. A cluster-scoped one cannot: Kubernetes treats a cluster-scoped +// object owned by a namespaced one as having an owner it cannot resolve and +// never collects it. Those carry storage.simplyblock.io/managed-by instead, and +// the finalizer deletes the ones it marked. That is the split +// design-simplyblockdriver.md §4.1 makes, for the same reason. + +package controlplane + +import ( + "context" + "fmt" + + "k8s.io/apimachinery/pkg/api/errors" + "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" + "k8s.io/apimachinery/pkg/runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/apiutil" + "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil" +) + +const ( + // fieldOwner is the field manager every apply here writes under. Taking + // ownership from a Helm release's manager happens under this name, so it has + // to stay stable across releases. + fieldOwner = client.FieldOwner("simplyblock-operator") + + // managedByLabel marks a cluster-scoped object this operator created without + // owning. It is the one key with that meaning in the group, so a single + // selector finds everything the operator is responsible for. + managedByLabel = "storage.simplyblock.io/managed-by" + + // managedByValue is the lowercased kind of the controller that manages the + // object rather than the object's name, because a label value admits no + // slash and a namespace and a name do not fit in one. + managedByValue = "controlplane" +) + +// applier writes the install's objects. It is the subset of client.Client the +// workload builders need, named so that the apply path can be exercised without +// a manager. +type applier interface { + client.Reader + Apply(ctx context.Context, obj runtime.ApplyConfiguration, opts ...client.ApplyOption) error +} + +// applyAll writes every object of a step in order, stopping at the first one +// that fails. The order within a step is the order the objects depend on each +// other: accounts and configuration before the RBAC that names them, and the +// RBAC before the workloads that run under them. +// +// A step is re-entered on every pass through steady state (§4.3), so this is +// also the mechanism that puts back an object somebody deleted and corrects one +// somebody edited. +func applyAll( + ctx context.Context, c applier, owner client.Object, scheme *runtime.Scheme, objects []client.Object, +) error { + for _, obj := range objects { + if err := applyOne(ctx, c, owner, scheme, obj); err != nil { + return err + } + } + return nil +} + +// applyOne marks an object as this control plane's and writes it. +func applyOne( + ctx context.Context, c applier, owner client.Object, scheme *runtime.Scheme, obj client.Object, +) error { + if err := setOwnership(owner, obj, scheme); err != nil { + return fmt.Errorf("set ownership on %s %s: %w", kindOf(obj, scheme), obj.GetName(), err) + } + cfg, err := applyConfiguration(obj, scheme) + if err != nil { + return fmt.Errorf("encode %s %s for apply: %w", kindOf(obj, scheme), obj.GetName(), err) + } + if err := c.Apply(ctx, cfg, fieldOwner, client.ForceOwnership); err != nil { + return fmt.Errorf("apply %s %s: %w", kindOf(obj, scheme), obj.GetName(), err) + } + return nil +} + +// setOwnership marks obj as this control plane's, by whichever mechanism its +// scope allows. It is additive: an install that meets objects a chart already +// labeled leaves those labels in place, because dropping them would rewrite a +// running object for no reason. +func setOwnership(owner client.Object, obj client.Object, scheme *runtime.Scheme) error { + if obj.GetNamespace() != "" { + return controllerutil.SetControllerReference(owner, obj, scheme) + } + + labels := obj.GetLabels() + if labels == nil { + labels = map[string]string{} + } + labels[managedByLabel] = managedByValue + obj.SetLabels(labels) + return nil +} + +// mayDelete reports whether this controller marked obj and may therefore remove +// it. A cluster-scoped object is shared ground: one carrying another +// controller's value, or none at all, belongs to somebody else and deleting it +// takes away somebody else's RBAC. +func mayDelete(obj client.Object) bool { + return obj.GetLabels()[managedByLabel] == managedByValue +} + +// applyConfiguration turns a built object into the shape a server-side apply +// takes. +// +// The apiVersion and kind have to be on the wire for an apply, and a typed +// object built in Go carries an empty TypeMeta, so the kind is resolved from the +// scheme rather than written by hand at each call site. An unstructured object +// already carries both, which is how the FoundationDB and cert-manager kinds +// this repository has no Go types for are applied. +func applyConfiguration(obj client.Object, scheme *runtime.Scheme) (runtime.ApplyConfiguration, error) { + if u, ok := obj.(*unstructured.Unstructured); ok { + return client.ApplyConfigurationFromUnstructured(u), nil + } + + gvk, err := apiutil.GVKForObject(obj, scheme) + if err != nil { + return nil, err + } + content, err := runtime.DefaultUnstructuredConverter.ToUnstructured(obj) + if err != nil { + return nil, err + } + u := &unstructured.Unstructured{Object: content} + u.SetGroupVersionKind(gvk) + // A built object has no status, and an apply carrying an empty one claims + // ownership of a field this controller does not set. For a Deployment that + // means fighting the Deployment controller over its own replica counts. + unstructured.RemoveNestedField(u.Object, "status") + unstructured.RemoveNestedField(u.Object, "metadata", "creationTimestamp") + return client.ApplyConfigurationFromUnstructured(u), nil +} + +// kindOf names an object for an error message. It falls back to the Go type +// where the scheme does not know the kind, since an error about an unregistered +// type must still say which object it was about. +func kindOf(obj client.Object, scheme *runtime.Scheme) string { + if gvk, err := apiutil.GVKForObject(obj, scheme); err == nil { + return gvk.Kind + } + return fmt.Sprintf("%T", obj) +} + +// deleteIfMarked removes a cluster-scoped object the finalizer is responsible +// for, and leaves alone one this controller did not mark. An object that is not +// there is already in the state the caller wants. +func deleteIfMarked(ctx context.Context, c client.Client, obj client.Object) error { + if err := c.Get(ctx, client.ObjectKeyFromObject(obj), obj); err != nil { + if errors.IsNotFound(err) { + return nil + } + return fmt.Errorf("read %s before deleting it: %w", obj.GetName(), err) + } + if !mayDelete(obj) { + return nil + } + if err := c.Delete(ctx, obj); err != nil && !errors.IsNotFound(err) { + return fmt.Errorf("delete %s: %w", obj.GetName(), err) + } + return nil +} diff --git a/operator/internal/controllers/controlplane/cel_validation_test.go b/operator/internal/controllers/controlplane/cel_validation_test.go new file mode 100644 index 000000000..27c060043 --- /dev/null +++ b/operator/internal/controllers/controlplane/cel_validation_test.go @@ -0,0 +1,246 @@ +// Validation of the CEL rules compiled into the ControlPlane CRD schema, run +// against a real apiserver. It lives here rather than under internal/webhook +// because there is no webhook involved: the rules are enforced by the apiserver +// itself, and envtest is the only place in the tree that starts one. +// +// The rule that one member and only one is set is worth an apiserver test rather +// than a unit test, for two reasons. It is the whole of what makes the two modes +// siblings rather than a convention, so a reconciler asking `isLocal` has to be +// able to assume it. And it is declared on ControlPlaneSource rather than on the +// field that carries it, a placement chosen to work around a generator flake: a +// rule that silently stopped reaching the schema would look exactly like a rule +// that works. +// +// The third case below is the one that earns its keep. The design says the block +// is immutable and its members are not, and the first spelling of that here was +// +k8s:immutable on the field, which freezes the image with the block and makes +// the Upgrade action impossible to complete. + +package controlplane + +import ( + "context" + "strings" + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// Exactly one member of spec.source has to be set. Both is a control plane the +// operator would install and probe somewhere else at the same time, and neither +// is an object nothing downstream can classify. +func TestControlPlaneCELRequiresExactlyOneSource(t *testing.T) { + apiClient := apiServer(t) + + localSource := &simplyblockv1alpha2.LocalControlPlane{Image: testImage} + managedSource := &simplyblockv1alpha2.ManagedControlPlane{ + Endpoint: "https://sb-control.example.com:5000", + CredentialsSecretRef: &corev1.LocalObjectReference{Name: "cp-token"}, + } + + for _, tc := range []struct { + name string + source simplyblockv1alpha2.ControlPlaneSource + wantDenied bool + }{ + { + name: "managed alone is what a fresh deployment writes", + source: simplyblockv1alpha2.ControlPlaneSource{Local: localSource}, + }, + { + name: "managed alone is a control plane that already exists", + source: simplyblockv1alpha2.ControlPlaneSource{Managed: managedSource}, + }, + { + name: "both would install one control plane and probe another", + source: simplyblockv1alpha2.ControlPlaneSource{Local: localSource, Managed: managedSource}, + wantDenied: true, + }, + { + name: "neither is an object nothing downstream can classify", + source: simplyblockv1alpha2.ControlPlaneSource{}, + wantDenied: true, + }, + } { + t.Run(tc.name, func(t *testing.T) { + namespace := freshNamespace(t, apiClient) + + cp := &simplyblockv1alpha2.ControlPlane{ + ObjectMeta: metav1.ObjectMeta{Name: SingletonName, Namespace: namespace}, + Spec: simplyblockv1alpha2.ControlPlaneSpec{Source: tc.source}, + } + + err := apiClient.Create(context.Background(), cp) + switch { + case tc.wantDenied && err == nil: + t.Fatal("the object was admitted, and exactly one of local or managed " + + "has to be set") + case tc.wantDenied: + if !strings.Contains(err.Error(), "set exactly one of local or managed") { + t.Errorf("denied with %q, want the rule's own message", err) + } + case err != nil: + t.Fatalf("a legal source was denied: %v", err) + } + }) + } +} + +// spec.source is immutable. Switching a live deployment between an installed +// control plane and an existing one is not a reconfiguration: the clusters, +// their UUIDs, and their volumes live in the FoundationDB behind the old one. +func TestControlPlaneCELRefusesToChangeTheSource(t *testing.T) { + ctx := context.Background() + apiClient := apiServer(t) + namespace := freshNamespace(t, apiClient) + + cp := &simplyblockv1alpha2.ControlPlane{ + ObjectMeta: metav1.ObjectMeta{Name: SingletonName, Namespace: namespace}, + Spec: simplyblockv1alpha2.ControlPlaneSpec{ + Source: simplyblockv1alpha2.ControlPlaneSource{ + Local: &simplyblockv1alpha2.LocalControlPlane{Image: testImage}, + }, + }, + } + if err := apiClient.Create(ctx, cp); err != nil { + t.Fatalf("create the control plane: %v", err) + } + + cp.Spec.Source = simplyblockv1alpha2.ControlPlaneSource{ + Managed: &simplyblockv1alpha2.ManagedControlPlane{ + Endpoint: "https://sb-control.example.com:5000", + CredentialsSecretRef: &corev1.LocalObjectReference{Name: "cp-token"}, + }, + } + err := apiClient.Update(ctx, cp) + if err == nil { + t.Fatal("the source was changed from local to managed, and the data behind the " + + "old one does not move with it") + } + if !strings.Contains(err.Error(), "immutable") { + t.Errorf("denied with %q, want the immutability rule", err) + } +} + +// Changing the image within a managed source is an ordinary edit, and it is what +// an Upgrade operation performs. The immutability rule covers the block rather +// than its members, so this has to stay admitted. +func TestControlPlaneCELAdmitsAnImageChangeWithinTheSameSource(t *testing.T) { + ctx := context.Background() + apiClient := apiServer(t) + namespace := freshNamespace(t, apiClient) + + cp := &simplyblockv1alpha2.ControlPlane{ + ObjectMeta: metav1.ObjectMeta{Name: SingletonName, Namespace: namespace}, + Spec: simplyblockv1alpha2.ControlPlaneSpec{ + Source: simplyblockv1alpha2.ControlPlaneSource{ + Local: &simplyblockv1alpha2.LocalControlPlane{Image: testImage}, + }, + }, + } + if err := apiClient.Create(ctx, cp); err != nil { + t.Fatalf("create the control plane: %v", err) + } + + cp.Spec.Source.Local.Image = "quay.io/simplyblock-io/simplyblock:26.3.0" + if err := apiClient.Update(ctx, cp); err != nil { + t.Fatalf("an image change was denied, and it is what an Upgrade performs: %v", err) + } +} + +// freshNamespace gives one test case a namespace of its own, so that the +// singleton name can be reused across cases against one shared apiserver. +func freshNamespace(t *testing.T, apiClient client.Client) string { + t.Helper() + ns := &corev1.Namespace{ + ObjectMeta: metav1.ObjectMeta{GenerateName: "cp-cel-"}, + } + if err := apiClient.Create(context.Background(), ns); err != nil { + t.Fatalf("create a namespace: %v", err) + } + return ns.Name +} + +// The operation's parameters are frozen once it is admitted, so the status stays +// an audit of the request that ran. They are consumed several steps apart: +// Preflight reads the image and Applying writes it, so an edit in between +// produces an operation that checked one thing and did another. +func TestControlPlaneOpsCELFreezesTheParametersAfterAdmission(t *testing.T) { + ctx := context.Background() + apiClient := apiServer(t) + namespace := freshNamespace(t, apiClient) + + ops := &simplyblockv1alpha2.ControlPlaneOps{ + ObjectMeta: metav1.ObjectMeta{Name: "an-upgrade", Namespace: namespace}, + Spec: simplyblockv1alpha2.ControlPlaneOpsSpec{ + ControlPlaneRef: SingletonName, + Action: simplyblockv1alpha2.ControlPlaneOpsActionUpgrade, + Upgrade: &simplyblockv1alpha2.UpgradeSpec{ + Image: "quay.io/simplyblock-io/simplyblock:26.3.0", + }, + }, + } + if err := apiClient.Create(ctx, ops); err != nil { + t.Fatalf("create the operation: %v", err) + } + + t.Run("the image cannot be swapped", func(t *testing.T) { + edited := ops.DeepCopy() + edited.Spec.Upgrade.Image = "quay.io/simplyblock-io/simplyblock:26.9.9" + if err := apiClient.Update(ctx, edited); err == nil { + t.Error("the image was changed after admission, so Preflight checked one " + + "version and Applying would write another") + } + }) + + t.Run("the block cannot be cleared", func(t *testing.T) { + edited := ops.DeepCopy() + edited.Spec.Upgrade = nil + if err := apiClient.Update(ctx, edited); err == nil { + t.Error("spec.upgrade was cleared after admission, which is what Applying " + + "would then dereference") + } + }) + + t.Run("abort stays settable", func(t *testing.T) { + edited := ops.DeepCopy() + edited.Spec.Abort = true + if err := apiClient.Update(ctx, edited); err != nil { + t.Errorf("spec.abort could not be set, and it is the one field meant to be "+ + "changed after the operation started: %v", err) + } + }) +} + +// A restart's scope is frozen for the same reason: the drain is decided from the +// component list, so widening it after Draining skips a drain the wider list +// would have required. +func TestControlPlaneOpsCELFreezesTheRestartScope(t *testing.T) { + ctx := context.Background() + apiClient := apiServer(t) + namespace := freshNamespace(t, apiClient) + + ops := &simplyblockv1alpha2.ControlPlaneOps{ + ObjectMeta: metav1.ObjectMeta{Name: "a-restart", Namespace: namespace}, + Spec: simplyblockv1alpha2.ControlPlaneOpsSpec{ + ControlPlaneRef: SingletonName, + Action: simplyblockv1alpha2.ControlPlaneOpsActionRestart, + Restart: &simplyblockv1alpha2.RestartSpec{ + Components: []string{ComponentTasks}, + }, + }, + } + if err := apiClient.Create(ctx, ops); err != nil { + t.Fatalf("create the operation: %v", err) + } + + ops.Spec.Restart.Components = []string{ComponentTasks, ComponentWebAPI} + if err := apiClient.Update(ctx, ops); err == nil { + t.Error("an essential component was added to the scope after admission, which " + + "would recycle the management API without the drain that scope requires") + } +} diff --git a/operator/internal/controllers/controlplane/components.go b/operator/internal/controllers/controlplane/components.go new file mode 100644 index 000000000..f29e8cfe4 --- /dev/null +++ b/operator/internal/controllers/controlplane/components.go @@ -0,0 +1,303 @@ +// The components a managed control plane consists of, and the phase derived +// from them. +// +// design-controlplane.md §4.3 keeps this table in the operator rather than in +// the API, because a user cannot add a component to a managed control plane and +// so has nothing to configure. A spec field here would exist only to let +// somebody mark the management API non-essential, which turns an outage into a +// warning without changing the outage. +// +// # Essential means the work stops, not that it slows down +// +// A component is essential when its absence loses work or stops the control +// plane answering. The task runner is the case that draws the line: its work is +// queued, so a runner at zero defers what is waiting rather than dropping it, +// and the queue is still there when it comes back. That is a control plane doing +// less than it should while remaining correct, which is what Degraded is for. +// +// Only a component marked essential can produce Unavailable, and that asymmetry +// is the safety property: Unavailable holds every controller in the operator, so +// the set of things able to cause it has to be a closed list somebody reviewed. +// A component added to the install without a decision about it lands in the +// default, which is that it can reach Degraded and cannot reach Unavailable. The +// failure that costs is halting a fleet over an exporter, not reporting a +// warning about one. + +package controlplane + +import ( + "context" + "fmt" + + appsv1 "k8s.io/api/apps/v1" + "k8s.io/apimachinery/pkg/api/errors" + "sigs.k8s.io/controller-runtime/pkg/client" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// componentKind is how a component's readiness is read. +type componentKind int + +const ( + // kindDeployment and kindStatefulSet are a ready count against a desired + // count, read off the workload's own status. + kindDeployment componentKind = iota + kindStatefulSet + + // kindFoundationDB is the resource's own report. A FoundationDBCluster at + // two of three coordinators is serving, and a replica count cannot say so. + kindFoundationDB +) + +// component is one workload of a managed control plane. +type component struct { + // name is the object's name, which is also what status.components publishes + // and what a Restart operation names. + name string + + // kind is how its readiness is read. + kind componentKind + + // essential is whether this component at zero ready makes the control plane + // Unavailable rather than Degraded. + essential bool + + // why records the decision behind essential, for the component where it is + // not self-evident. It is read by nothing and exists so the table explains + // itself where it is edited. + why string +} + +// componentTable is every workload the managed install applies, in the order +// status.components reports them: the database first, then what serves, then +// what supports it. +// +// It is exactly the set the install applies, which is what makes it watchable: +// a component the operator does not own is one it cannot report a count for, and +// the observability half the chart still renders is therefore absent from here +// rather than reported as missing. +var componentTable = []component{ + { + name: ComponentFDBCluster, kind: kindFoundationDB, essential: true, + why: "the control plane's entire state is in it; a database that is not " + + "serving is a control plane that answers nothing", + }, + { + name: ComponentWebAPI, kind: kindDeployment, essential: true, + why: "it is the endpoint every controller in this operator calls", + }, + { + name: ComponentFDBOperator, kind: kindDeployment, essential: false, + why: "a database already running keeps running without its operator; " + + "what stops is reconfiguring it and replacing a failed process group", + }, + { + name: ComponentTasks, kind: kindDeployment, essential: false, + why: "its work is queued, so a runner at zero defers what is waiting " + + "rather than dropping it", + }, + { + name: ComponentMonitoring, kind: kindDeployment, essential: false, + why: "the fleet keeps serving while nothing is watching it, and what is " + + "lost is the record of what happened rather than the work", + }, + { + name: ComponentAdminControl, kind: kindDeployment, essential: false, + why: "it runs nothing; it exists to be exec'd into", + }, + { + name: ComponentMinio, kind: kindStatefulSet, essential: false, + why: "design-controlplane.md §12 Q5 leaves this open, and non-essential " + + "is the default a component takes until somebody decides: nothing in " + + "the install path or the data path reads the store", + }, + { + name: ComponentFDBExporter, kind: kindDeployment, essential: false, + why: "it publishes metrics about a database that is serving either way", + }, +} + +// observe reads every component's counts back. A workload that is not there yet +// is not an error: the apply created it and the cache has not caught up, which +// reads as zero ready against zero desired. +func observe( + ctx context.Context, c client.Reader, namespace string, +) ([]simplyblockv1alpha2.ControlPlaneComponentStatus, error) { + out := make([]simplyblockv1alpha2.ControlPlaneComponentStatus, 0, len(componentTable)) + + for _, comp := range componentTable { + status := simplyblockv1alpha2.ControlPlaneComponentStatus{ + Name: comp.name, + Essential: comp.essential, + } + + switch comp.kind { + case kindDeployment: + var d appsv1.Deployment + err := c.Get(ctx, client.ObjectKey{Namespace: namespace, Name: comp.name}, &d) + switch { + case errors.IsNotFound(err): + case err != nil: + return nil, fmt.Errorf("read Deployment %s: %w", comp.name, err) + default: + status.Desired = desiredReplicas(d.Spec.Replicas) + status.Ready = d.Status.ReadyReplicas + } + + case kindStatefulSet: + var s appsv1.StatefulSet + err := c.Get(ctx, client.ObjectKey{Namespace: namespace, Name: comp.name}, &s) + switch { + case errors.IsNotFound(err): + case err != nil: + return nil, fmt.Errorf("read StatefulSet %s: %w", comp.name, err) + default: + status.Desired = desiredReplicas(s.Spec.Replicas) + status.Ready = s.Status.ReadyReplicas + } + + case kindFoundationDB: + health, err := readFoundationDB(ctx, c, namespace) + if err != nil { + return nil, err + } + // The database reports whether it is serving, not how many of its + // processes are. Publishing its process groups as the two counts + // keeps the field meaning the same thing for every component, and + // the health is folded in by reporting zero ready when the cluster + // is not available at all, which is the state that makes this + // component's verdict Unavailable. + status.Desired = health.desired + status.Ready = health.reconciled + if !health.available { + status.Ready = 0 + } + } + + out = append(out, status) + } + return out, nil +} + +// desiredReplicas is a workload's replica count. A nil pointer is Kubernetes's +// default of one rather than zero, which matters because zero desired is what +// this package reads as not there yet. +func desiredReplicas(replicas *int32) int32 { + if replicas == nil { + return 1 + } + return *replicas +} + +// verdict is what one signal says about the control plane. The values are +// ordered by severity so that the phase is the maximum across every signal. +type verdict int + +const ( + verdictAvailable verdict = iota + verdictDegraded + verdictUnavailable +) + +// componentVerdict is what one component's counts say. +// +// A component below its desired count while above zero is Degraded whether or +// not it is essential: the management API at one of two replicas is still +// answering every request, which is exactly the window Degraded exists to name. +// Zero ready is where the two diverge. +func componentVerdict(status simplyblockv1alpha2.ControlPlaneComponentStatus) verdict { + switch { + case status.Desired == 0: + // Nothing is asked for, so nothing is missing. This is the workload the + // apply just created and the cache has not caught up with. + return verdictAvailable + case status.Ready == 0 && status.Essential: + return verdictUnavailable + case status.Ready == 0: + return verdictDegraded + case status.Ready < status.Desired: + return verdictDegraded + default: + return verdictAvailable + } +} + +// derivePhase is the worst verdict among the readiness probe and every +// component, together with the sentence explaining it. +// +// The probe answers whether the control plane responds. The components answer +// whether it is one restart away from not responding. A probe that fails settles +// the phase on its own, because a control plane that does not answer is +// Unavailable whatever its pod counts say. +func derivePhase( + probeOK bool, probeError string, + components []simplyblockv1alpha2.ControlPlaneComponentStatus, +) (simplyblockv1alpha2.ControlPlanePhase, string) { + if !probeOK { + return simplyblockv1alpha2.ControlPlanePhaseUnavailable, probeError + } + + worst := verdictAvailable + reason := "" + for _, status := range components { + v := componentVerdict(status) + if v > worst { + worst = v + reason = fmt.Sprintf("%s has %d of %d replicas ready", + status.Name, status.Ready, status.Desired) + } + } + + switch worst { + case verdictUnavailable: + return simplyblockv1alpha2.ControlPlanePhaseUnavailable, reason + case verdictDegraded: + return simplyblockv1alpha2.ControlPlanePhaseDegraded, reason + default: + return simplyblockv1alpha2.ControlPlanePhaseAvailable, "" + } +} + +// essentialComponents is the set a scoped Restart has to drain before recycling. +// It is derived from the table rather than restated, so a component whose +// classification changes moves both at once. +func essentialComponents() map[string]bool { + out := make(map[string]bool, len(componentTable)) + for _, comp := range componentTable { + if comp.essential { + out[comp.name] = true + } + } + return out +} + +// restartable reports whether a Restart may name this component. +// +// It is deliberately narrower than membership of the table. The FoundationDB +// cluster is a component of the control plane and is not rolled by writing an +// annotation into a pod template, so an operation scoped to it would recycle +// nothing, find a healthy database on the wait that follows, and report success +// having done nothing. Refusing the name is the only answer that does not lie. +func restartable(name string) bool { + for _, comp := range restartableComponents() { + if comp.name == name { + return true + } + } + return false +} + +// restartableComponents are the workloads a Restart can roll: everything in the +// table that is a Deployment or a StatefulSet. The FoundationDBCluster is not +// one of them, because recycling a database is the FoundationDB operator's +// mechanism and not a pod-template annotation. +func restartableComponents() []component { + out := make([]component, 0, len(componentTable)) + for _, comp := range componentTable { + if comp.kind != kindFoundationDB { + out = append(out, comp) + } + } + return out +} diff --git a/operator/internal/controllers/controlplane/components_test.go b/operator/internal/controllers/controlplane/components_test.go new file mode 100644 index 000000000..be753c7d0 --- /dev/null +++ b/operator/internal/controllers/controlplane/components_test.go @@ -0,0 +1,261 @@ +// The phase the component table derives, and the asymmetry that keeps a fleet +// from being halted over an exporter. +// +// The pairs matter more than the individual rows. Only a component the table +// marks essential may produce Unavailable, because Unavailable holds every +// controller in the operator. Each case here is therefore stated twice, once for +// an essential component and once for a non-essential one at the same counts. + +package controlplane + +import ( + "context" + "testing" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// A failing probe settles the phase on its own. A control plane that does not +// answer is Unavailable whatever its pod counts say, and the message is the +// control plane's own rather than a paraphrase of it. +func TestAFailingProbeIsUnavailableWhateverTheComponentsSay(t *testing.T) { + everythingReady := []simplyblockv1alpha2.ControlPlaneComponentStatus{ + componentStatus(ComponentWebAPI, 2, 2, true), + componentStatus(ComponentFDBCluster, 7, 7, true), + } + + phase, message := derivePhase(false, "status=503: fdb unavailable", everythingReady) + + if phase != simplyblockv1alpha2.ControlPlanePhaseUnavailable { + t.Errorf("phase = %s, want Unavailable", phase) + } + if message != "status=503: fdb unavailable" { + t.Errorf("message = %q, want the control plane's own words", message) + } +} + +// An essential component at zero ready is Unavailable; a non-essential one at +// the same counts is Degraded. This is the asymmetry §4.3 calls the safety +// property, and it is the one row of this table worth getting wrong twice. +func TestOnlyAnEssentialComponentAtZeroReachesUnavailable(t *testing.T) { + for _, tc := range []struct { + name string + essential bool + want simplyblockv1alpha2.ControlPlanePhase + }{ + {"an essential component at zero", true, simplyblockv1alpha2.ControlPlanePhaseUnavailable}, + {"a non-essential component at zero", false, simplyblockv1alpha2.ControlPlanePhaseDegraded}, + } { + t.Run(tc.name, func(t *testing.T) { + components := []simplyblockv1alpha2.ControlPlaneComponentStatus{ + componentStatus("a-component", 1, 0, tc.essential), + } + + phase, message := derivePhase(true, "", components) + + if phase != tc.want { + t.Errorf("phase = %s, want %s", phase, tc.want) + } + if message == "" { + t.Error("message is empty; a phase that cannot be explained is one nobody trusts") + } + }) + } +} + +// A component below its desired count while above zero is Degraded whether or +// not it is essential. The management API at one of two replicas is still +// answering every request, which is exactly the window Degraded exists to name. +func TestBelowDesiredButAboveZeroIsDegradedEvenWhenEssential(t *testing.T) { + for _, essential := range []bool{true, false} { + components := []simplyblockv1alpha2.ControlPlaneComponentStatus{ + componentStatus(ComponentWebAPI, 2, 1, essential), + } + + phase, _ := derivePhase(true, "", components) + + if phase != simplyblockv1alpha2.ControlPlanePhaseDegraded { + t.Errorf("essential=%v: phase = %s, want Degraded", essential, phase) + } + } +} + +// The phase is the worst verdict across every component, not the first or the +// last one read. +func TestThePhaseIsTheWorstVerdictAcrossEveryComponent(t *testing.T) { + components := []simplyblockv1alpha2.ControlPlaneComponentStatus{ + componentStatus("healthy", 1, 1, false), + componentStatus("degraded", 2, 1, false), + componentStatus("down", 1, 0, true), + componentStatus("also-healthy", 1, 1, true), + } + + phase, message := derivePhase(true, "", components) + + if phase != simplyblockv1alpha2.ControlPlanePhaseUnavailable { + t.Errorf("phase = %s, want Unavailable", phase) + } + if message != "down has 0 of 1 replicas ready" { + t.Errorf("message = %q, want the component that produced the verdict", message) + } +} + +// A workload the apply just created and the cache has not caught up with reads +// as zero desired, which is nothing asked for rather than something missing. An +// install therefore settles without passing through Unavailable. +func TestZeroDesiredIsNotAnOutage(t *testing.T) { + components := []simplyblockv1alpha2.ControlPlaneComponentStatus{ + componentStatus(ComponentWebAPI, 0, 0, true), + } + + phase, _ := derivePhase(true, "", components) + + if phase != simplyblockv1alpha2.ControlPlanePhaseAvailable { + t.Errorf("phase = %s, want Available: a workload that is not there yet is not an outage", phase) + } +} + +// A remote control plane has no components the operator owns, so the probe is +// the only signal and Degraded is unreachable. +func TestWithNoComponentsThePhaseFollowsTheProbeAlone(t *testing.T) { + for _, tc := range []struct { + probeOK bool + want simplyblockv1alpha2.ControlPlanePhase + }{ + {true, simplyblockv1alpha2.ControlPlanePhaseAvailable}, + {false, simplyblockv1alpha2.ControlPlanePhaseUnavailable}, + } { + phase, _ := derivePhase(tc.probeOK, "unreachable", nil) + if phase != tc.want { + t.Errorf("probeOK=%v: phase = %s, want %s", tc.probeOK, phase, tc.want) + } + } +} + +// The management API and the database are the only components that may halt a +// fleet. This pins the closed list §4.3 says somebody has to have reviewed: a +// component added to the install without a decision about it lands in the +// non-essential default, and this test is what notices when one is added as +// essential instead. +func TestOnlyTheAPIAndTheDatabaseAreEssential(t *testing.T) { + want := map[string]bool{ + ComponentWebAPI: true, + ComponentFDBCluster: true, + } + + got := essentialComponents() + + if len(got) != len(want) { + t.Fatalf("essential components = %v, want exactly %v", got, want) + } + for name := range want { + if !got[name] { + t.Errorf("%s is not essential, and its absence is an outage", name) + } + } +} + +// Every component in the table carries the reason for its classification, so +// that the table explains itself where it is edited. +func TestEveryComponentStatesWhyItIsClassifiedAsItIs(t *testing.T) { + for _, comp := range componentTable { + if comp.why == "" { + t.Errorf("%s states no reason for essential=%v", comp.name, comp.essential) + } + } +} + +// The FoundationDBCluster is not something a Restart rolls: recycling a database +// is the FoundationDB operator's mechanism rather than a pod-template +// annotation, and a restart that silently skipped it would report success having +// done nothing to it. +func TestTheDatabaseIsNotRestartable(t *testing.T) { + for _, comp := range restartableComponents() { + if comp.name == ComponentFDBCluster { + t.Fatalf("%s is restartable, and writing a pod-template annotation onto a "+ + "FoundationDBCluster does nothing", ComponentFDBCluster) + } + } + if restartable(ComponentFDBCluster) { + t.Errorf("%s is accepted as a restart scope, so an operation naming it would recycle "+ + "nothing and report success", ComponentFDBCluster) + } + if !restartable(ComponentTasks) { + t.Errorf("%s is refused as a restart scope, and it is a workload a restart rolls", + ComponentTasks) + } +} + +// Every component the table lists is either restartable or excluded for a stated +// reason. A component that is neither is one a Restart refuses without anything +// saying why. +func TestEveryComponentIsRestartableOrExcludedForAReason(t *testing.T) { + for _, comp := range componentTable { + if comp.kind == kindFoundationDB { + continue + } + if !restartable(comp.name) { + t.Errorf("%s is neither a FoundationDB resource nor restartable", comp.name) + } + } +} + +// A workload that is not in the cluster reads as zero against zero rather than +// as an error, because the apply that creates it and the read that follows are +// two calls with a cache between them. +func TestObserveReadsAnAbsentWorkloadAsNothingAskedFor(t *testing.T) { + c := newClient(t) + + components, err := observe(context.Background(), c, testNamespace) + if err != nil { + t.Fatalf("observe: %v", err) + } + + if len(components) != len(componentTable) { + t.Fatalf("observe reported %d components, want one per table entry (%d)", + len(components), len(componentTable)) + } + for _, status := range components { + if status.Desired != 0 || status.Ready != 0 { + t.Errorf("%s = %d/%d ready, want 0/0 for a workload that does not exist", + status.Name, status.Ready, status.Desired) + } + } +} + +// observe reads the counts off the workloads, and carries each component's +// classification through to the status so a phase can be explained without +// reading the operator's source. +func TestObserveReportsTheCountsAndTheClassification(t *testing.T) { + c := newClient(t, + deployment(ComponentWebAPI, 2, 1), + statefulSet(ComponentMinio, 1, 1), + ) + + components, err := observe(context.Background(), c, testNamespace) + if err != nil { + t.Fatalf("observe: %v", err) + } + + byName := map[string]simplyblockv1alpha2.ControlPlaneComponentStatus{} + for _, status := range components { + byName[status.Name] = status + } + + api := byName[ComponentWebAPI] + if api.Desired != 2 || api.Ready != 1 { + t.Errorf("%s = %d/%d ready, want 1/2", ComponentWebAPI, api.Ready, api.Desired) + } + if !api.Essential { + t.Errorf("%s is reported non-essential", ComponentWebAPI) + } + + store := byName[ComponentMinio] + if store.Desired != 1 || store.Ready != 1 { + t.Errorf("%s = %d/%d ready, want 1/1", ComponentMinio, store.Ready, store.Desired) + } + if store.Essential { + t.Errorf("%s is reported essential, which would let the object store halt a fleet", + ComponentMinio) + } +} diff --git a/operator/internal/controllers/controlplane/controlplane_controller.go b/operator/internal/controllers/controlplane/controlplane_controller.go new file mode 100644 index 000000000..af28868bf --- /dev/null +++ b/operator/internal/controllers/controlplane/controlplane_controller.go @@ -0,0 +1,745 @@ +// The ControlPlane reconciler: the installation machine, the steady state it +// settles into, and the finalizer that refuses while clusters still exist. +// +// Nothing here blocks. A step that is not finished requeues, and the step it is +// on is in the status, so a controller restart resumes rather than restarts. +// +// Two paths leave Reconcile, and they are different in what the operator owns. +// A managed control plane is installed, watched, and re-applied on every pass, +// which is what puts back an object somebody deleted and corrects one somebody +// edited, and it is why no operation exists for checking the install. A remote +// one is resolved, probed, and reported, and the operator touches nothing behind +// its endpoint. +// +// design-controlplane.md §4 is the specification. + +package controlplane + +import ( + "context" + "fmt" + "time" + + appsv1 "k8s.io/api/apps/v1" + corev1 "k8s.io/api/core/v1" + "k8s.io/apimachinery/pkg/api/meta" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/runtime" + "k8s.io/client-go/tools/events" + "k8s.io/client-go/util/retry" + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/builder" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil" + logf "sigs.k8s.io/controller-runtime/pkg/log" + "sigs.k8s.io/controller-runtime/pkg/predicate" + + "github.com/simplyblock/atlas/statemachine" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +const ( + // FinalizerControlPlane is what holds a deletion while any StorageCluster in + // the namespace still exists, and what gives the controller a pass to delete + // the cluster-scoped objects the garbage collector will not. + FinalizerControlPlane = "storage.simplyblock.io/controlplane-finalizer" + + // steadyStateInterval is how often a settled control plane is probed. It is + // also the resolution at which an outage is noticed, which is the reason it + // is not longer. + steadyStateInterval = 30 * time.Second + + // installAdvance is how long a pass that moved the machine forward waits. + // The status write this pass made is itself a change the controller watches, + // so this is the backstop for the event rather than the path the next step + // normally arrives on. + installAdvance = time.Second + + // installRetry is how long a held installation step waits before looking + // again at something it cannot hurry. + installRetry = 15 * time.Second +) + +// ControlPlaneReconciler reconciles the singleton ControlPlane. +type ControlPlaneReconciler struct { + client.Client + Scheme *runtime.Scheme + Recorder events.EventRecorder + + // Prober performs the readiness and version reads. It is an interface so the + // phase branches can be exercised without an HTTP server. + Prober Prober +} + +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=controlplanes,verbs=get;list;watch;create;update;patch;delete +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=controlplanes/status,verbs=get;update;patch +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=controlplanes/finalizers,verbs=update +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storageclusters,verbs=get;list;watch +// +kubebuilder:rbac:groups=apps,resources=deployments;statefulsets,verbs=get;list;watch;create;update;patch;delete +// +kubebuilder:rbac:groups="",resources=serviceaccounts;configmaps;services,verbs=get;list;watch;create;update;patch;delete +// +kubebuilder:rbac:groups="",resources=secrets,verbs=get;list;watch +// +kubebuilder:rbac:groups=rbac.authorization.k8s.io,resources=roles;rolebindings;clusterroles;clusterrolebindings,verbs=get;list;watch;create;update;patch;delete;escalate;bind +// +kubebuilder:rbac:groups=apps.foundationdb.org,resources=foundationdbclusters;foundationdbbackups,verbs=get;list;watch;create;update;patch;delete + +// Reconcile advances the control plane by at most one step. +func (r *ControlPlaneReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) { + log := logf.FromContext(ctx) + + // A ControlPlane under another name is ignored and sits inert, which is the + // singleton enforced by convention rather than by the API server (§3.1). + if req.Name != SingletonName { + log.Info("ignoring a ControlPlane that is not the singleton", "name", req.Name) + return ctrl.Result{}, nil + } + + var cp simplyblockv1alpha2.ControlPlane + if err := r.Get(ctx, req.NamespacedName, &cp); err != nil { + return ctrl.Result{}, client.IgnoreNotFound(err) + } + + if !cp.DeletionTimestamp.IsZero() { + return r.finalize(ctx, &cp) + } + + // The limit is one per Kubernetes cluster and not one per namespace (§3.1), + // and the name check above cannot see that: it admits a "simplyblock" object + // in every namespace. Each of those would apply the same fixed-name + // cluster-scoped roles and bindings under the same managed-by label, so they + // would overwrite each other's and either one's finalizer would delete what + // the other needs. + // + // The check runs before the finalizer is taken, for the reason + // SimplyblockDriver gives for the same ordering: a finalizer on the duplicate + // would delete the holder's cluster-scoped objects when somebody removed the + // duplicate, which is the opposite of what removing a duplicate should do. + holder, err := r.deploymentHolder(ctx) + if err != nil { + return ctrl.Result{}, err + } + // An empty holder means the list came back without the object this reconcile + // just read, which is a stale cache rather than a second control plane. + // Refusing on it would stall an install behind a message naming nobody, so + // the object in hand is taken as the holder and the next pass corrects it. + if holder.Name != "" && holder != client.ObjectKeyFromObject(&cp) { + message := fmt.Sprintf( + "a Kubernetes cluster holds one ControlPlane, and %s/%s holds it", + holder.Namespace, holder.Name) + r.emit(&cp, corev1.EventTypeWarning, DuplicateControlPlane, message) + return ctrl.Result{RequeueAfter: steadyStateInterval}, r.report(ctx, &cp, + simplyblockv1alpha2.ControlPlanePhaseInstalling, message) + } + + if err := r.ensureFinalizer(ctx, &cp); err != nil { + return ctrl.Result{}, err + } + + switch { + case isManaged(&cp): + return r.reconcileManaged(ctx, &cp) + case isLocal(&cp): + return r.reconcileLocal(ctx, &cp) + default: + // The API's CEL rule refuses this at admission, so reaching it means an + // object written before the rule shipped or one a conversion produced. + // It holds and says so rather than picking a mode on the user's behalf. + return ctrl.Result{RequeueAfter: steadyStateInterval}, r.report(ctx, &cp, + simplyblockv1alpha2.ControlPlanePhaseInstalling, + "spec.source names neither a local nor a managed control plane, so there is "+ + "nothing to install and nowhere to probe") + } +} + +// reconcileManaged resolves the endpoint, probes it, and reports. The operator +// installs nothing and owns no components here, so status.components stays empty +// and the probe is the only signal: the phase is Available or Unavailable, and +// Degraded is unreachable (§4.3). +func (r *ControlPlaneReconciler) reconcileManaged( + ctx context.Context, cp *simplyblockv1alpha2.ControlPlane, +) (ctrl.Result, error) { + access, err := resolveManaged(ctx, r.Client, cp) + if err != nil { + var credentials *credentialsError + reason := EndpointUnreachable + if errorsAs(err, &credentials) { + reason = CredentialsError + } + r.emit(cp, corev1.EventTypeWarning, reason, err.Error()) + return ctrl.Result{RequeueAfter: steadyStateInterval}, r.report(ctx, cp, + simplyblockv1alpha2.ControlPlanePhaseUnavailable, err.Error()) + } + + ok, message := r.probe(ctx, cp.Namespace, access) + phase := simplyblockv1alpha2.ControlPlanePhaseAvailable + if !ok { + phase = simplyblockv1alpha2.ControlPlanePhaseUnavailable + } + + r.announce(cp, phase, message) + return ctrl.Result{RequeueAfter: steadyStateInterval}, r.publish(ctx, cp, statusUpdate{ + phase: phase, + message: message, + endpoint: access.endpoint, + version: r.version(ctx, access), + probed: true, + }) +} + +// reconcileLocal runs the installation machine, and then the steady state it +// settles into. +func (r *ControlPlaneReconciler) reconcileLocal( + ctx context.Context, cp *simplyblockv1alpha2.ControlPlane, +) (ctrl.Result, error) { + // The FoundationDB kinds are a prerequisite rather than something to wait + // on: creating a FoundationDBCluster against a group the API server does not + // serve is an error on every attempt, and the CRDs are the chart's to apply. + if served, err := r.foundationDBServed(); err != nil { + return ctrl.Result{}, err + } else if !served { + message := fmt.Sprintf( + "the Kubernetes cluster does not serve %s/%s, so the FoundationDB this control "+ + "plane stores its state in cannot be created", fdbGroup, fdbVersion) + r.emit(cp, corev1.EventTypeWarning, PrerequisiteMissing, message) + return ctrl.Result{RequeueAfter: installRetry}, r.report(ctx, cp, + simplyblockv1alpha2.ControlPlanePhaseInstalling, message) + } + + if cp.Status.Phase != simplyblockv1alpha2.ControlPlanePhaseInstalling && + cp.Status.Phase != "" { + return r.steadyState(ctx, cp) + } + return r.install(ctx, cp) +} + +// install advances the installation machine by at most one step. +func (r *ControlPlaneReconciler) install( + ctx context.Context, cp *simplyblockv1alpha2.ControlPlane, +) (ctrl.Result, error) { + machine, err := statemachine.NewFromSnapshot(ctx, installGraph(), + statemachine.FromKube[installStep](cp.Status.Step)) + if err != nil { + // An unrecognized step is a downgrade, a hand-edited object, or a rename + // that shipped without a conversion, and none of them resolve by + // reconciling again. Restarting the install is the safe answer here and + // not elsewhere in this group: every step is an apply, so re-entering + // the first one re-applies to the same result. + return r.enterStepAt(ctx, cp, stepApplyingFoundationDB) + } + defer machine.Close() + + // A machine is born already in its initial state, so that state's entry hook + // never runs and no deadline is set for it. Setting one on the first pass is + // what stops the first step being the one step that cannot time out. + if cp.Status.Step.State == "" { + return r.enterStepAt(ctx, cp, machine.CurrentState()) + } + + current := machine.CurrentState() + + if machine.TimeoutReached() { + message := fmt.Sprintf("step %s outlived its deadline", current) + r.emit(cp, corev1.EventTypeWarning, StepDeadlineExceeded, message) + // The step is not abandoned. An install that has outrun its budget is + // still the only path to a working control plane, and there is nothing + // to roll back to: the deadline is a detection mechanism rather than a + // recovery one, and what it produces is the event and the message. + return ctrl.Result{RequeueAfter: installRetry}, r.report(ctx, cp, + simplyblockv1alpha2.ControlPlanePhaseInstalling, message) + } + + done, held, err := r.performInstallStep(ctx, cp, current) + if err != nil { + return ctrl.Result{}, err + } + if !done { + r.emit(cp, corev1.EventTypeNormal, AwaitingDependency, held) + return ctrl.Result{RequeueAfter: installRetry}, r.report(ctx, cp, + simplyblockv1alpha2.ControlPlanePhaseInstalling, held) + } + + next, ok := nextStep(machine) + if !ok { + // AwaitingAPI is terminal, and reaching it is what makes the control + // plane Available. The step is left on the object as the record of how + // the install finished. + return r.steadyState(ctx, cp) + } + if err := machine.TransitionTo(ctx, next); err != nil { + return ctrl.Result{}, fmt.Errorf("enter step %s: %w", next, err) + } + r.recordStepDuration(cp, current) + snapshot := statemachine.ToKube(machine.Snapshot()) + return ctrl.Result{RequeueAfter: installAdvance}, r.recordStep(ctx, cp, next, snapshot.Deadline) +} + +// performInstallStep does what one step is for, and reports whether the step is +// finished. A step that is not finished returns what it is waiting on, which is +// read from the thing being waited for so that a stalled install names the +// dependency rather than the operator. +func (r *ControlPlaneReconciler) performInstallStep( + ctx context.Context, cp *simplyblockv1alpha2.ControlPlane, current installStep, +) (done bool, held string, err error) { + switch current { + case stepApplyingFoundationDB: + return true, "", applyAll(ctx, r.Client, cp, r.Scheme, foundationDBObjects(cp)) + + case stepAwaitingFoundationDB: + health, err := readFoundationDB(ctx, r.Client, cp.Namespace) + if err != nil { + return false, "", err + } + if waiting := health.waitingOn(); waiting != "" { + return false, waiting, nil + } + return true, "", nil + + case stepApplyingDatastore: + return true, "", applyAll(ctx, r.Client, cp, r.Scheme, datastoreObjects(cp)) + + case stepApplyingAPI: + return true, "", applyAll(ctx, r.Client, cp, r.Scheme, managementAPIObjects(cp)) + + case stepAwaitingAPI: + ok, message := r.probe(ctx, cp.Namespace, managedAccess{endpoint: localEndpoint(cp.Namespace)}) + if !ok { + return false, fmt.Sprintf("the management API is not answering yet: %s", message), nil + } + return true, "", nil + + default: + return false, "", fmt.Errorf("unknown installation step %q", current) + } +} + +// steadyState re-applies what the install created, probes readiness, and +// republishes the endpoint, the version, and the component counts. +// +// Re-applying is what keeps an object somebody deleted or edited from staying +// that way, and it is why no operation exists for checking the install (§6). +func (r *ControlPlaneReconciler) steadyState( + ctx context.Context, cp *simplyblockv1alpha2.ControlPlane, +) (ctrl.Result, error) { + if err := r.applyEverything(ctx, cp); err != nil { + return ctrl.Result{}, err + } + + // A managed control plane is reached on the Service this install created, so + // there is no token to present and no CA beyond the cluster's own. + access := managedAccess{endpoint: localEndpoint(cp.Namespace)} + ok, message := r.probe(ctx, cp.Namespace, access) + + components, err := observe(ctx, r.Client, cp.Namespace) + if err != nil { + return ctrl.Result{}, err + } + r.recordComponents(cp.Namespace, components) + + phase, reason := derivePhase(ok, message, components) + r.announce(cp, phase, reason) + + return ctrl.Result{RequeueAfter: steadyStateInterval}, r.publish(ctx, cp, statusUpdate{ + phase: phase, + message: reason, + endpoint: access.endpoint, + version: r.version(ctx, access), + components: components, + probed: true, + }) +} + +// applyEverything writes every object of every step, which is the re-apply +// steady state performs. The order is the installation's, because the +// dependencies between the objects do not change once they exist. +func (r *ControlPlaneReconciler) applyEverything( + ctx context.Context, cp *simplyblockv1alpha2.ControlPlane, +) error { + for _, set := range [][]client.Object{ + foundationDBObjects(cp), + datastoreObjects(cp), + managementAPIObjects(cp), + } { + if err := applyAll(ctx, r.Client, cp, r.Scheme, set); err != nil { + return err + } + } + return nil +} + +// probe performs the readiness read and records what it cost. +func (r *ControlPlaneReconciler) probe( + ctx context.Context, namespace string, access managedAccess, +) (bool, string) { + prober := r.Prober + if prober == nil { + prober = &HTTPProber{Token: access.token, Client: access.client} + } + + started := time.Now() + ok, message := prober.Ready(ctx, access.endpoint) + controlPlaneProbeDuration.WithLabelValues(namespace).Observe(time.Since(started).Seconds()) + + if ok { + controlPlaneReadyState.WithLabelValues(namespace).Set(1) + } else { + controlPlaneReadyState.WithLabelValues(namespace).Set(0) + controlPlaneProbeFailures.WithLabelValues(namespace, probeFailureReason(message)).Inc() + } + return ok, message +} + +// version reads what the management API reports. A read that fails publishes +// nothing rather than clearing what was published, because a version the +// operator could not confirm this pass is not a version that changed. +func (r *ControlPlaneReconciler) version(ctx context.Context, access managedAccess) string { + prober := r.Prober + if prober == nil { + prober = &HTTPProber{Token: access.token, Client: access.client} + } + version, err := prober.Version(ctx, access.endpoint) + if err != nil { + return "" + } + return version +} + +// statusUpdate is what one pass concluded, applied to the object as a whole so +// that the phase and the evidence behind it are never written apart. +type statusUpdate struct { + phase simplyblockv1alpha2.ControlPlanePhase + message string + endpoint string + version string + components []simplyblockv1alpha2.ControlPlaneComponentStatus + + // probed says whether this pass ran the readiness probe, which is what makes + // status.lastChecked mean what it says. + probed bool +} + +// publish writes the conclusion of one pass. +func (r *ControlPlaneReconciler) publish( + ctx context.Context, cp *simplyblockv1alpha2.ControlPlane, update statusUpdate, +) error { + return r.writeStatus(ctx, cp, func(status *simplyblockv1alpha2.ControlPlaneStatus) { + status.Phase = update.phase + status.Message = update.message + status.Endpoint = update.endpoint + status.Components = update.components + if update.version != "" { + status.Version = update.version + } + if update.probed { + now := metav1.Now() + status.LastChecked = &now + } + }) +} + +// report writes a phase and its reason and nothing else, which is what a pass +// that reached no conclusion about the endpoint or the components owes. +func (r *ControlPlaneReconciler) report( + ctx context.Context, + cp *simplyblockv1alpha2.ControlPlane, + phase simplyblockv1alpha2.ControlPlanePhase, + message string, +) error { + return r.writeStatus(ctx, cp, func(status *simplyblockv1alpha2.ControlPlaneStatus) { + status.Phase = phase + status.Message = message + }) +} + +// recordStep persists the step and its deadline before the side effect that step +// performs, which is the write-ahead record every multi-step operation in this +// group keeps. +func (r *ControlPlaneReconciler) recordStep( + ctx context.Context, + cp *simplyblockv1alpha2.ControlPlane, + next installStep, + stepDeadline *metav1.Time, +) error { + return r.writeStatus(ctx, cp, func(status *simplyblockv1alpha2.ControlPlaneStatus) { + status.Phase = simplyblockv1alpha2.ControlPlanePhaseInstalling + status.Step = statemachine.KubeSnapshot{State: string(next), Deadline: stepDeadline} + }) +} + +// enterStepAt sets a step's deadline from the graph's budget and persists it. It +// is what the first pass of an install does, and what a status carrying a step +// nothing recognizes is reset to. +func (r *ControlPlaneReconciler) enterStepAt( + ctx context.Context, cp *simplyblockv1alpha2.ControlPlane, step installStep, +) (ctrl.Result, error) { + budget, ok := installStepBudgets[step] + if !ok { + budget = applyingFoundationDBDeadline + } + stepDeadline := metav1.NewTime(time.Now().Add(budget)) + return ctrl.Result{RequeueAfter: installAdvance}, r.recordStep(ctx, cp, step, &stepDeadline) +} + +// nextStep is the step that follows the current one. The installation graph is a +// line, so the first edge is the only edge, and no edge means the step is +// terminal. +func nextStep(machine *statemachine.Machine[installStep]) (installStep, bool) { + for next := range machine.AllowedTransitions() { + return next, true + } + return machine.CurrentState(), false +} + +// recordStepDuration measures how long a step took from the budget it was given +// and the deadline left on it. The deadline is already persisted, so the +// arithmetic costs the API no second timestamp. +func (r *ControlPlaneReconciler) recordStepDuration( + cp *simplyblockv1alpha2.ControlPlane, step installStep, +) { + budget, ok := installStepBudgets[step] + if !ok || cp.Status.Step.Deadline == nil { + return + } + remaining := time.Until(cp.Status.Step.Deadline.Time) + spent := budget - remaining + if spent < 0 { + return + } + controlPlaneInstallStepDuration. + WithLabelValues(cp.Namespace, string(step)).Observe(spent.Seconds()) +} + +// recordComponents publishes each component's two counts. +func (r *ControlPlaneReconciler) recordComponents( + namespace string, components []simplyblockv1alpha2.ControlPlaneComponentStatus, +) { + for _, component := range components { + essential := "false" + if component.Essential { + essential = "true" + } + controlPlaneComponentReady. + WithLabelValues(namespace, component.Name, essential).Set(float64(component.Ready)) + controlPlaneComponentDesired. + WithLabelValues(namespace, component.Name, essential).Set(float64(component.Desired)) + } +} + +// announce emits the phase's event, and only on a transition. +// +// What is worth an event is the arrival at a phase rather than the phase itself, +// which bounds one outage to one event rather than one per thirty-second probe. +func (r *ControlPlaneReconciler) announce( + cp *simplyblockv1alpha2.ControlPlane, + phase simplyblockv1alpha2.ControlPlanePhase, + message string, +) { + if phase == cp.Status.Phase { + return + } + switch phase { + case simplyblockv1alpha2.ControlPlanePhaseUnavailable: + r.emit(cp, corev1.EventTypeWarning, ControlPlaneNotReady, message) + case simplyblockv1alpha2.ControlPlanePhaseDegraded: + r.emit(cp, corev1.EventTypeWarning, ControlPlaneDegraded, message) + case simplyblockv1alpha2.ControlPlanePhaseAvailable: + r.emit(cp, corev1.EventTypeNormal, ControlPlaneReady, + "the control plane's readiness probe passed") + } +} + +// finalize refuses while any StorageCluster in the namespace still exists, then +// removes what the garbage collector will not. +// +// The refusal is the point: deleting a ControlPlane with a managed source +// deletes a database, and the clusters, their UUIDs, and their volumes live in +// it. It is a hold rather than a failure: removing the clusters resolves it, and +// nothing else can. +// +// The same hold applies to a remote control plane, because a namespace whose +// clusters have no control plane to reach is a namespace of objects nothing can +// reconcile. +func (r *ControlPlaneReconciler) finalize( + ctx context.Context, cp *simplyblockv1alpha2.ControlPlane, +) (ctrl.Result, error) { + if !controllerutil.ContainsFinalizer(cp, FinalizerControlPlane) { + return ctrl.Result{}, nil + } + + var clusters simplyblockv1alpha2.StorageClusterList + if err := r.List(ctx, &clusters, client.InNamespace(cp.Namespace)); err != nil { + return ctrl.Result{}, err + } + if len(clusters.Items) > 0 { + message := fmt.Sprintf( + "%d StorageCluster objects still exist in %s, and their data lives behind this "+ + "control plane; remove them first", + len(clusters.Items), cp.Namespace) + r.emit(cp, corev1.EventTypeWarning, ClustersStillPresent, message) + return ctrl.Result{RequeueAfter: steadyStateInterval}, r.report(ctx, cp, cp.Status.Phase, message) + } + + // Deleting a control plane that never held the install must take nothing with + // it. The check is repeated here rather than trusted from the reconcile that + // added the finalizer, because an object may carry one from before this + // ordering existed, and the cost of being wrong is the running deployment's + // RBAC. + holder, err := r.deploymentHolder(ctx) + if err != nil { + return ctrl.Result{}, err + } + heldTheInstall := holder.Name == "" || holder == client.ObjectKeyFromObject(cp) + + // A remote control plane had nothing installed, so there is nothing + // cluster-scoped to remove: deletion takes the object and touches neither + // the endpoint nor its data. + if isLocal(cp) && heldTheInstall { + for _, obj := range append(foundationDBClusterScoped(), managementAPIClusterScoped()...) { + if err := deleteIfMarked(ctx, r.Client, obj); err != nil { + return ctrl.Result{}, err + } + } + } + + base := cp.DeepCopy() + controllerutil.RemoveFinalizer(cp, FinalizerControlPlane) + return ctrl.Result{}, r.Patch(ctx, cp, client.MergeFrom(base)) +} + +func (r *ControlPlaneReconciler) ensureFinalizer( + ctx context.Context, cp *simplyblockv1alpha2.ControlPlane, +) error { + if controllerutil.ContainsFinalizer(cp, FinalizerControlPlane) { + return nil + } + base := cp.DeepCopy() + controllerutil.AddFinalizer(cp, FinalizerControlPlane) + return r.Patch(ctx, cp, client.MergeFrom(base)) +} + +// deploymentHolder is the ControlPlane that owns the install: the oldest in the +// Kubernetes cluster, with namespace and name breaking a tie. Every reconciler +// picks the same one from the same list, so two of them never disagree about +// which object is the second. +func (r *ControlPlaneReconciler) deploymentHolder(ctx context.Context) (client.ObjectKey, error) { + var list simplyblockv1alpha2.ControlPlaneList + if err := r.List(ctx, &list); err != nil { + return client.ObjectKey{}, err + } + + var oldest *simplyblockv1alpha2.ControlPlane + for i := range list.Items { + item := &list.Items[i] + if item.Name != SingletonName { + continue + } + if oldest == nil || olderThan(item, oldest) { + oldest = item + } + } + if oldest == nil { + return client.ObjectKey{}, nil + } + return client.ObjectKey{Namespace: oldest.Namespace, Name: oldest.Name}, nil +} + +// olderThan orders two control planes by creation time, then by namespace and +// name. The tie-break matters because a creation timestamp has one-second +// resolution, so two objects applied together compare equal on it. +func olderThan(a, b *simplyblockv1alpha2.ControlPlane) bool { + if !a.CreationTimestamp.Equal(&b.CreationTimestamp) { + return a.CreationTimestamp.Before(&b.CreationTimestamp) + } + if a.Namespace != b.Namespace { + return a.Namespace < b.Namespace + } + return a.Name < b.Name +} + +// foundationDBServed reports whether the API server knows the FoundationDB +// kinds. It asks the manager's RESTMapper rather than a discovery client, +// because the mapper is already backed by discovery and is what an apply would +// consult anyway. +func (r *ControlPlaneReconciler) foundationDBServed() (bool, error) { + _, err := r.RESTMapper().RESTMapping(fdbClusterGVK.GroupKind(), fdbClusterGVK.Version) + switch { + case err == nil: + return true, nil + case meta.IsNoMatchError(err): + return false, nil + default: + return false, err + } +} + +// writeStatus applies a mutation to the live status, retrying a conflict. The +// observed generation is stamped here rather than by each caller, so that no +// path can write a phase without saying which spec it was computed from. +func (r *ControlPlaneReconciler) writeStatus( + ctx context.Context, + cp *simplyblockv1alpha2.ControlPlane, + mutate func(*simplyblockv1alpha2.ControlPlaneStatus), +) error { + key := client.ObjectKeyFromObject(cp) + return retry.RetryOnConflict(retry.DefaultRetry, func() error { + var current simplyblockv1alpha2.ControlPlane + if err := r.Get(ctx, key, ¤t); err != nil { + return client.IgnoreNotFound(err) + } + mutate(¤t.Status) + current.Status.ObservedGeneration = current.Generation + if err := r.Status().Update(ctx, ¤t); err != nil { + return err + } + cp.Status = current.Status + return nil + }) +} + +// emit records what the reconcile decided, so that a refusal to act is visible +// to somebody reading the object rather than only in a log. +func (r *ControlPlaneReconciler) emit( + object client.Object, eventType, reason, message string, +) { + if r.Recorder == nil { + return + } + // The action is the reason, as every other recorder call in this operator + // passes it. It is a required field, and the phase is empty on the first + // event an object ever emits. + r.Recorder.Eventf(object, nil, eventType, reason, reason, "%s", message) +} + +// errorsAs is errors.As, wrapped so this file does not import the standard +// errors package alongside the Kubernetes one under a renamed identifier. +func errorsAs[T error](err error, target *T) bool { + for err != nil { + if typed, ok := err.(T); ok { + *target = typed + return true + } + unwrapped, ok := err.(interface{ Unwrap() error }) + if !ok { + return false + } + err = unwrapped.Unwrap() + } + return false +} + +// SetupWithManager registers the reconciler. +// +// The workloads are watched as well as the ControlPlane itself: a component's +// ready count changing is what moves the phase between Available and Degraded, +// and the watch is what keeps that within the event rather than the +// steady-state interval. +// +// The ControlPlane's own watch is filtered to generation changes. Every probe +// stamps status.lastChecked, and an unfiltered watch turns that write into +// another reconcile, which probes and stamps again. +func (r *ControlPlaneReconciler) SetupWithManager(mgr ctrl.Manager) error { + return ctrl.NewControllerManagedBy(mgr). + For(&simplyblockv1alpha2.ControlPlane{}, + builder.WithPredicates(predicate.GenerationChangedPredicate{})). + Owns(&appsv1.Deployment{}). + Owns(&appsv1.StatefulSet{}). + Named("controlplane"). + Complete(r) +} diff --git a/operator/internal/controllers/controlplane/controlplane_controller_test.go b/operator/internal/controllers/controlplane/controlplane_controller_test.go new file mode 100644 index 000000000..1408def9b --- /dev/null +++ b/operator/internal/controllers/controlplane/controlplane_controller_test.go @@ -0,0 +1,445 @@ +// The reconciler's branches: the singleton, the two sources, the install's +// progression, and the deletion hold. +// +// The fake client carries no RESTMapper that knows the FoundationDB kinds, so +// the tests that drive a managed install past the prerequisite check call the +// install path directly. What the prerequisite check itself does is tested +// against a reconciler with no mapping, which is the state of a cluster that has +// not been given the CRDs. + +package controlplane + +import ( + "context" + "strings" + "testing" + "time" + + appsv1 "k8s.io/api/apps/v1" + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// A ControlPlane under any other name is ignored and sits inert, which is the +// singleton enforced by convention. Reconciling one would install a second +// control plane beside the first. +func TestAControlPlaneThatIsNotTheSingletonIsIgnored(t *testing.T) { + other := localControlPlane() + other.Name = "a-second-one" + c := newClient(t, other) + r := &ControlPlaneReconciler{Client: c, Scheme: testScheme(t)} + + result, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: client.ObjectKeyFromObject(other), + }) + if err != nil { + t.Fatalf("Reconcile: %v", err) + } + if result.RequeueAfter != 0 { + t.Errorf("RequeueAfter = %s, want 0: an ignored object is not looked at again", + result.RequeueAfter) + } + + var after simplyblockv1alpha2.ControlPlane + if err := c.Get(context.Background(), client.ObjectKeyFromObject(other), &after); err != nil { + t.Fatalf("read the object back: %v", err) + } + if after.Status.Phase != "" { + t.Errorf("status.phase = %s, want nothing written at all", after.Status.Phase) + } + if len(after.Finalizers) != 0 { + t.Errorf("finalizers = %v, want none on an object the controller ignores", after.Finalizers) + } +} + +// A ControlPlane is one per Kubernetes cluster, not one per namespace. Two of +// them reconciling at once apply the same fixed-name cluster-scoped RBAC under +// the same managed-by label, so each overwrites the other's and either one's +// deletion takes away what the other needs. +// +// The older object holds the deployment, and the younger reports that it does +// not. That is the rule SimplyblockDriver uses for the same reason, and the two +// have to agree about which object is the second. +func TestASecondControlPlaneInAnotherNamespaceDoesNotInstall(t *testing.T) { + ctx := context.Background() + + holder := localControlPlane() + holder.CreationTimestamp = metav1.NewTime(time.Now().Add(-time.Hour)) + + second := localControlPlane() + second.Namespace = "simplyblock-second" + second.CreationTimestamp = metav1.NewTime(time.Now()) + + c := newClient(t, holder, second) + recorder := &recordingRecorder{} + r := &ControlPlaneReconciler{ + Client: c, Scheme: testScheme(t), Recorder: recorder, + Prober: &stubProber{ready: true}, + } + + if _, err := r.Reconcile(ctx, ctrl.Request{ + NamespacedName: client.ObjectKeyFromObject(second), + }); err != nil { + t.Fatalf("Reconcile: %v", err) + } + + var deployments appsv1.DeploymentList + if err := c.List(ctx, &deployments, client.InNamespace(second.Namespace)); err != nil { + t.Fatalf("list the second namespace: %v", err) + } + if len(deployments.Items) != 0 { + t.Errorf("the second control plane installed %d workloads, and the first holds the "+ + "cluster-scoped objects they would overwrite", len(deployments.Items)) + } + + var after simplyblockv1alpha2.ControlPlane + if err := c.Get(ctx, client.ObjectKeyFromObject(second), &after); err != nil { + t.Fatalf("read the second control plane back: %v", err) + } + if !strings.Contains(after.Status.Message, holder.Namespace) { + t.Errorf("status.message = %q, want it to name the namespace that holds the "+ + "deployment", after.Status.Message) + } + if recorder.count(DuplicateControlPlane) == 0 { + t.Error("no DuplicateControlPlane event: nothing says why this object does nothing") + } +} + +// The holder itself still installs. A rule that refused both would take a +// working deployment down on the first reconcile after somebody added a second +// object by mistake. +func TestTheOlderControlPlaneStillHoldsTheDeployment(t *testing.T) { + holder := localControlPlane() + holder.CreationTimestamp = metav1.NewTime(time.Now().Add(-time.Hour)) + + second := localControlPlane() + second.Namespace = "simplyblock-second" + second.CreationTimestamp = metav1.NewTime(time.Now()) + + c := newClient(t, holder, second) + r := &ControlPlaneReconciler{Client: c, Scheme: testScheme(t)} + + key, err := r.deploymentHolder(context.Background()) + if err != nil { + t.Fatalf("deploymentHolder: %v", err) + } + if key.Namespace != holder.Namespace { + t.Errorf("the holder is %s/%s, want the older object in %s", + key.Namespace, key.Name, holder.Namespace) + } +} + +// A remote control plane is resolved, probed, and reported, and the operator +// applies nothing. status.endpoint echoes what the spec said, which is the point +// of the field: a reader asks status.endpoint either way and does not have to +// know which mode the deployment is in. +func TestARemoteControlPlaneIsProbedAndNothingIsInstalled(t *testing.T) { + const endpoint = "https://sb-control.example.com:5000" + cp := managedControlPlane(endpoint) + secret := &corev1.Secret{ + ObjectMeta: metav1.ObjectMeta{Name: "cp-token", Namespace: testNamespace}, + Data: map[string][]byte{"token": []byte("a-bearer-token")}, + } + + c := newClient(t, cp, secret) + prober := &stubProber{ready: true, version: "26.2.8"} + r := &ControlPlaneReconciler{Client: c, Scheme: testScheme(t), Prober: prober} + + if _, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: client.ObjectKeyFromObject(cp), + }); err != nil { + t.Fatalf("Reconcile: %v", err) + } + + var after simplyblockv1alpha2.ControlPlane + if err := c.Get(context.Background(), client.ObjectKeyFromObject(cp), &after); err != nil { + t.Fatalf("read the object back: %v", err) + } + + if after.Status.Phase != simplyblockv1alpha2.ControlPlanePhaseAvailable { + t.Errorf("status.phase = %s, want Available", after.Status.Phase) + } + if after.Status.Endpoint != endpoint { + t.Errorf("status.endpoint = %q, want the spec's %q", after.Status.Endpoint, endpoint) + } + if after.Status.Version != "26.2.8" { + t.Errorf("status.version = %q, want what the control plane reported", after.Status.Version) + } + if len(after.Status.Components) != 0 { + t.Errorf("status.components = %v, want empty: the operator owns no pods there", + after.Status.Components) + } + if after.Status.LastChecked == nil { + t.Error("status.lastChecked is unset after a probe ran") + } +} + +// A remote control plane may name no credentials Secret, and that is the +// shape the chart writes when it installs the control plane itself: the endpoint +// is a ClusterIP Service in the same namespace, the readiness probe there is +// unauthenticated, and there is no static token to point at. +// +// It is the field that lets the chart hand the install back: the CR names a +// remote source, and the operator resolves it rather than installing. +func TestARemoteControlPlaneMayNameNoCredentials(t *testing.T) { + ctx := context.Background() + const endpoint = "http://simplyblock-webappapi.simplyblock.svc.cluster.local:5000" + + cp := managedControlPlane(endpoint) + cp.Spec.Source.Managed.CredentialsSecretRef = nil + + c := newClient(t, cp) + prober := &stubProber{ready: true} + r := &ControlPlaneReconciler{Client: c, Scheme: testScheme(t), Prober: prober} + + if _, err := r.Reconcile(ctx, ctrl.Request{ + NamespacedName: client.ObjectKeyFromObject(cp), + }); err != nil { + t.Fatalf("Reconcile: %v", err) + } + + var after simplyblockv1alpha2.ControlPlane + if err := c.Get(ctx, client.ObjectKeyFromObject(cp), &after); err != nil { + t.Fatalf("read the object back: %v", err) + } + if after.Status.Phase != simplyblockv1alpha2.ControlPlanePhaseAvailable { + t.Errorf("status.phase = %s (%q), want Available", after.Status.Phase, after.Status.Message) + } + if after.Status.Endpoint != endpoint { + t.Errorf("status.endpoint = %q, want %q", after.Status.Endpoint, endpoint) + } + if prober.readyCalls == 0 { + t.Error("the endpoint was never probed") + } +} + +// A credentials Secret that is named and absent is still an error. Naming one +// states that the control plane needs it, and the probe is refused until it is +// there. +func TestANamedButMissingCredentialsSecretIsStillAnError(t *testing.T) { + cp := managedControlPlane("https://sb-control.example.com:5000") + c := newClient(t, cp) + r := &ControlPlaneReconciler{Client: c, Scheme: testScheme(t)} + + _, err := resolveManaged(context.Background(), c, cp) + if err == nil { + t.Fatal("a named Secret that does not exist was accepted") + } + var credentials *credentialsError + if !errorsAs(err, &credentials) { + t.Errorf("err = %v, want a credentials error so the event names the right cause", err) + } + _ = r +} + +// A credentials Secret that does not exist is a different problem from an +// endpoint that does not answer, and it is reported as one: the phase is +// Unavailable either way, and the message names the Secret. +func TestAMissingCredentialsSecretIsReportedAsSuch(t *testing.T) { + cp := managedControlPlane("https://sb-control.example.com:5000") + c := newClient(t, cp) + prober := &stubProber{ready: true} + r := &ControlPlaneReconciler{Client: c, Scheme: testScheme(t), Prober: prober} + + if _, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: client.ObjectKeyFromObject(cp), + }); err != nil { + t.Fatalf("Reconcile: %v", err) + } + + var after simplyblockv1alpha2.ControlPlane + if err := c.Get(context.Background(), client.ObjectKeyFromObject(cp), &after); err != nil { + t.Fatalf("read the object back: %v", err) + } + if after.Status.Phase != simplyblockv1alpha2.ControlPlanePhaseUnavailable { + t.Errorf("status.phase = %s, want Unavailable", after.Status.Phase) + } + if !strings.Contains(after.Status.Message, "cp-token") { + t.Errorf("status.message = %q, want it to name the Secret", after.Status.Message) + } + if prober.readyCalls != 0 { + t.Error("the endpoint was probed although there was no token to probe it with") + } +} + +// An endpoint that resolves inside the operator's own pod is refused. The spec's +// pattern admits it, so this is the guard that stops an operator being pointed +// at itself. +func TestALoopbackEndpointIsRefused(t *testing.T) { + for _, endpoint := range []string{ + "http://localhost:5000", + "http://127.0.0.1:5000", + "https://169.254.169.254/latest", + } { + if err := validateEndpoint(endpoint); err == nil { + t.Errorf("%s was admitted, and it does not name a control plane", endpoint) + } + } + for _, endpoint := range []string{ + "https://sb-control.example.com:5000", + "http://10.0.0.5:5000", + } { + if err := validateEndpoint(endpoint); err != nil { + t.Errorf("%s was refused: %v", endpoint, err) + } + } +} + +// A managed install holds with a named prerequisite where the FoundationDB kinds +// are not served, rather than failing against the API server on every pass. The +// CRDs are the chart's to apply, and creating a FoundationDBCluster against a +// group the API server does not know is an error that reconciling does not fix. +func TestAMissingFoundationDBGroupHoldsTheInstallWithAReason(t *testing.T) { + cp := localControlPlane() + c := newClient(t, cp) + r := &ControlPlaneReconciler{Client: c, Scheme: testScheme(t)} + + result, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: client.ObjectKeyFromObject(cp), + }) + if err != nil { + t.Fatalf("Reconcile: %v", err) + } + if result.RequeueAfter == 0 { + t.Error("the install was not requeued, so it would never notice the CRDs arriving") + } + + var after simplyblockv1alpha2.ControlPlane + if err := c.Get(context.Background(), client.ObjectKeyFromObject(cp), &after); err != nil { + t.Fatalf("read the object back: %v", err) + } + if after.Status.Phase != simplyblockv1alpha2.ControlPlanePhaseInstalling { + t.Errorf("status.phase = %s, want Installing", after.Status.Phase) + } + if !strings.Contains(after.Status.Message, fdbGroup) { + t.Errorf("status.message = %q, want it to name the group that is missing", + after.Status.Message) + } +} + +// The finalizer is taken before anything is installed, so that an object deleted +// mid-install still goes through the hold rather than being collected while its +// database is coming up. +func TestTheFinalizerIsTakenBeforeAnythingIsInstalled(t *testing.T) { + cp := localControlPlane() + c := newClient(t, cp) + r := &ControlPlaneReconciler{Client: c, Scheme: testScheme(t)} + + if _, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: client.ObjectKeyFromObject(cp), + }); err != nil { + t.Fatalf("Reconcile: %v", err) + } + + var after simplyblockv1alpha2.ControlPlane + if err := c.Get(context.Background(), client.ObjectKeyFromObject(cp), &after); err != nil { + t.Fatalf("read the object back: %v", err) + } + if !containsString(after.Finalizers, FinalizerControlPlane) { + t.Errorf("finalizers = %v, want %s", after.Finalizers, FinalizerControlPlane) + } +} + +// Deleting a ControlPlane while clusters still exist is held, not failed: +// removing the clusters resolves it, and nothing else can. The data those +// clusters describe lives in the FoundationDB this object's deletion would take +// with it. +func TestDeletionIsHeldWhileClustersStillExist(t *testing.T) { + cp := localControlPlane() + cp.Finalizers = []string{FinalizerControlPlane} + now := metav1.Now() + cp.DeletionTimestamp = &now + + cluster := &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{Name: "production", Namespace: testNamespace}, + } + c := newClient(t, cp, cluster) + r := &ControlPlaneReconciler{Client: c, Scheme: testScheme(t)} + + result, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: client.ObjectKeyFromObject(cp), + }) + if err != nil { + t.Fatalf("Reconcile: %v", err) + } + if result.RequeueAfter == 0 { + t.Error("the hold was not requeued, so removing the cluster would not release it") + } + + var after simplyblockv1alpha2.ControlPlane + if err := c.Get(context.Background(), client.ObjectKeyFromObject(cp), &after); err != nil { + t.Fatalf("the object was collected while a cluster still existed: %v", err) + } + if !containsString(after.Finalizers, FinalizerControlPlane) { + t.Errorf("finalizers = %v, want the hold still in place", after.Finalizers) + } + if !strings.Contains(after.Status.Message, "StorageCluster") { + t.Errorf("status.message = %q, want it to say what is holding the deletion", + after.Status.Message) + } +} + +// With no clusters left, the hold releases and the cluster-scoped objects this +// controller marked go with it. +func TestDeletionReleasesOnceTheClustersAreGone(t *testing.T) { + ctx := context.Background() + cp := localControlPlane() + cp.Finalizers = []string{FinalizerControlPlane} + now := metav1.Now() + cp.DeletionTimestamp = &now + + // A cluster-scoped object this controller marked, and one it did not. + ours := fdbManagerClusterRoleObject() + ours.Labels = map[string]string{managedByLabel: managedByValue} + theirs := serviceReaderClusterRole() + theirs.Labels = map[string]string{managedByLabel: "somebody-else"} + + c := newClient(t, cp, ours, theirs) + r := &ControlPlaneReconciler{Client: c, Scheme: testScheme(t)} + + if _, err := r.Reconcile(ctx, ctrl.Request{ + NamespacedName: client.ObjectKeyFromObject(cp), + }); err != nil { + t.Fatalf("Reconcile: %v", err) + } + + var after simplyblockv1alpha2.ControlPlane + err := c.Get(ctx, client.ObjectKeyFromObject(cp), &after) + if err == nil && containsString(after.Finalizers, FinalizerControlPlane) { + t.Error("the finalizer is still held with no clusters left") + } +} + +// The event marks the arrival at a phase rather than the phase itself, which +// bounds one outage to one event rather than one per thirty-second probe. +func TestTheEventMarksTheTransitionRatherThanTheState(t *testing.T) { + recorder := &recordingRecorder{} + cp := localControlPlane() + cp.Status.Phase = simplyblockv1alpha2.ControlPlanePhaseUnavailable + r := &ControlPlaneReconciler{Recorder: recorder} + + r.announce(cp, simplyblockv1alpha2.ControlPlanePhaseUnavailable, "still down") + if len(recorder.reasons) != 0 { + t.Errorf("an event was emitted for a phase that did not change: %v", recorder.reasons) + } + + r.announce(cp, simplyblockv1alpha2.ControlPlanePhaseAvailable, "") + if len(recorder.reasons) != 1 || recorder.reasons[0] != ControlPlaneReady { + t.Errorf("reasons = %v, want one %s on the recovery", recorder.reasons, ControlPlaneReady) + } +} + +// containsString is slices.Contains for a finalizer list, spelled here so the +// assertions read as what they check. +func containsString(list []string, want string) bool { + for _, item := range list { + if item == want { + return true + } + } + return false +} diff --git a/operator/internal/controllers/controlplane/controlplaneops_controller.go b/operator/internal/controllers/controlplane/controlplaneops_controller.go new file mode 100644 index 000000000..99746a374 --- /dev/null +++ b/operator/internal/controllers/controlplane/controlplaneops_controller.go @@ -0,0 +1,964 @@ +// The ControlPlaneOps reconciler: the lock, the three actions, and the finalizer +// that releases the lock even when the operation is deleted while holding it. +// +// One operation per control plane is the strongest form of the limit in this +// group, because a control-plane operation interrupts every controller rather +// than one cluster or one node. A second operation is admitted, sits at Pending, +// and runs when the lock frees. +// +// Every action requires a managed control plane, since each acts on something +// the operator installed. The webhook refuses one naming a managed control +// plane at creation; this reconciler repeats the check, because an object may +// have been created while the webhook was not serving. +// +// design-controlplane.md §6 and §7 are the specification. + +package controlplane + +import ( + "context" + "fmt" + "slices" + "strings" + "time" + + appsv1 "k8s.io/api/apps/v1" + corev1 "k8s.io/api/core/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" + "k8s.io/apimachinery/pkg/runtime" + "k8s.io/client-go/tools/events" + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil" + + "github.com/simplyblock/atlas/statemachine" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +const ( + // FinalizerControlPlaneOps is what guarantees the lock is released even when + // the operation is deleted while it holds one, so that the control plane is + // never left locked by an object that no longer exists. + FinalizerControlPlaneOps = "storage.simplyblock.io/controlplaneops-finalizer" + + // opsRetry is how long an operation waits before looking again at something + // it cannot hurry: a lock another operation holds, or work in flight it is + // draining behind. + opsRetry = 15 * time.Second + + // opsAdvance is how long a pass that moved the machine forward waits. + opsAdvance = time.Second + + // restartedAtAnnotation is what recycles a workload. Writing a new value + // into the pod template is what a rolling restart is: the Deployment + // controller sees a changed template and rolls it, which is the same + // mechanism `kubectl rollout restart` uses. + restartedAtAnnotation = "storage.simplyblock.io/restartedAt" +) + +// ControlPlaneOpsReconciler reconciles a ControlPlaneOps. +type ControlPlaneOpsReconciler struct { + client.Client + Scheme *runtime.Scheme + Recorder events.EventRecorder + + // Prober is what Verifying re-reads the version with. It is the same + // interface the entity's reconciler uses, for the same reason. + Prober Prober +} + +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=controlplaneops,verbs=get;list;watch;create;update;patch;delete +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=controlplaneops/status,verbs=get;update;patch +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=controlplaneops/finalizers,verbs=update +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storageclusterops;storagenodeops;storagepoolops,verbs=get;list;watch + +// Reconcile advances one operation by at most one step. +func (r *ControlPlaneOpsReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) { + var ops simplyblockv1alpha2.ControlPlaneOps + if err := r.Get(ctx, req.NamespacedName, &ops); err != nil { + return ctrl.Result{}, client.IgnoreNotFound(err) + } + + if !ops.DeletionTimestamp.IsZero() { + return r.reconcileDeletion(ctx, &ops) + } + + // A terminal operation is a record rather than a task. Re-reconciling one + // does nothing at all, including nothing to its target: an operation that + // finished has already released the lock, and taking it again to release it + // again is how an unrelated operation loses one it legitimately holds. + if terminalOps(ops.Status.Phase) { + return ctrl.Result{}, r.ensureFinalizer(ctx, &ops) + } + if err := r.ensureFinalizer(ctx, &ops); err != nil { + return ctrl.Result{}, err + } + + target := &simplyblockv1alpha2.ControlPlane{} + key := client.ObjectKey{Namespace: ops.Namespace, Name: ops.Spec.ControlPlaneRef} + switch err := r.Get(ctx, key, target); { + case apierrors.IsNotFound(err): + return r.finish(ctx, &ops, nil, simplyblockv1alpha2.ControlPlaneOpsPhaseFailed, + fmt.Sprintf("no ControlPlane %q in namespace %s", ops.Spec.ControlPlaneRef, ops.Namespace)) + case err != nil: + return ctrl.Result{}, err + } + + // The webhook refuses this at creation. It is repeated here because an + // object may carry a source the webhook was not serving to check, and + // because spec.source is immutable so the answer cannot have changed since. + if !isLocal(target) { + return r.failAndRelease(ctx, &ops, target, + fmt.Sprintf("ControlPlane %q names a control plane this cluster does not host, "+ + "and every action of this kind acts on something the operator installed", + target.Name)) + } + + if ops.Status.Phase != simplyblockv1alpha2.ControlPlaneOpsPhaseRunning { + acquired, result, err := r.acquireLock(ctx, &ops, target) + if err != nil || !acquired { + return result, err + } + } + + return r.advance(ctx, &ops, target) +} + +// acquireLock takes the target's status.activeOpsRef, which is the mutual +// exclusion every Ops kind in this group uses. An operation that finds the lock +// held by another stays Pending and asks again: the other operation finishes, +// and the wait is what keeps the outcome independent of the order two objects +// were applied in. +func (r *ControlPlaneOpsReconciler) acquireLock( + ctx context.Context, + ops *simplyblockv1alpha2.ControlPlaneOps, + target *simplyblockv1alpha2.ControlPlane, +) (bool, ctrl.Result, error) { + if held := target.Status.ActiveOpsRef; held != "" && held != ops.Name { + if ops.Status.Phase != simplyblockv1alpha2.ControlPlaneOpsPhasePending { + if err := r.setPhase(ctx, ops, simplyblockv1alpha2.ControlPlaneOpsPhasePending, + fmt.Sprintf("waiting for operation %q to release the control plane", held)); err != nil { + return false, ctrl.Result{}, err + } + } + r.emit(ops, corev1.EventTypeNormal, OperationQueued, + fmt.Sprintf("control plane %q is held by operation %q", target.Name, held)) + return false, ctrl.Result{RequeueAfter: opsRetry}, nil + } + + // The optimistic lock is what makes this a lock at all: two reconcilers that + // both saw a free field patch the same resourceVersion, and one of them gets + // a 409 and comes back to find the field taken. + base := target.DeepCopy() + target.Status.ActiveOpsRef = ops.Name + patch := client.MergeFromWithOptions(base, client.MergeFromWithOptimisticLock{}) + if err := r.Status().Patch(ctx, target, patch); err != nil { + return false, ctrl.Result{RequeueAfter: opsRetry}, nil //nolint:nilerr // a lost race is retried, not failed + } + + now := metav1.Now() + opsBase := ops.DeepCopy() + ops.Status.Phase = simplyblockv1alpha2.ControlPlaneOpsPhaseRunning + ops.Status.StartedAt = &now + ops.Status.Message = "" + ops.Status.ObservedGeneration = ops.Generation + if err := r.Status().Patch(ctx, ops, client.MergeFrom(opsBase)); err != nil { + return false, ctrl.Result{}, err + } + r.emit(ops, corev1.EventTypeNormal, OperationStarted, + fmt.Sprintf("acquired control plane %q and started %s", target.Name, ops.Spec.Action)) + return true, ctrl.Result{Requeue: true}, nil +} + +// advance runs the action's machine forward by at most one step. +func (r *ControlPlaneOpsReconciler) advance( + ctx context.Context, + ops *simplyblockv1alpha2.ControlPlaneOps, + target *simplyblockv1alpha2.ControlPlane, +) (ctrl.Result, error) { + machine, err := opsGraphs().FromSnapshot(ctx, action(ops.Spec.Action), + statemachine.FromKube[opsStep](ops.Status.Step)) + if err != nil { + // An unrecognized step or action is a downgrade, a hand-edited object, + // or a rename that shipped without a conversion, and none of them + // resolve by reconciling again. + return r.failAndRelease(ctx, ops, target, + fmt.Sprintf("the operation cannot be resumed: %v", err)) + } + defer machine.Close() + + if ops.Status.Step.State == "" { + return r.enterInitialStep(ctx, ops, machine) + } + + current := machine.CurrentState() + + if ops.Spec.Abort { + return r.unwind(ctx, ops, target, current) + } + + if machine.TimeoutReached() { + return r.failAndRelease(ctx, ops, target, + fmt.Sprintf("step %s outlived its deadline", current)) + } + + done, held, err := r.perform(ctx, ops, target, current) + switch { + case err != nil: + var fatal *terminalStepError + if errorsAs(err, &fatal) { + return r.failAndRelease(ctx, ops, target, fatal.Error()) + } + return ctrl.Result{}, err + case !done: + return ctrl.Result{RequeueAfter: opsRetry}, r.note(ctx, ops, held) + } + + next, ok := nextOpsStep(machine) + if !ok { + if err := r.releaseLock(ctx, ops, target); err != nil { + return ctrl.Result{}, err + } + return r.finish(ctx, ops, target, simplyblockv1alpha2.ControlPlaneOpsPhaseSucceeded, + fmt.Sprintf("%s finished", ops.Spec.Action)) + } + if err := machine.TransitionTo(ctx, next); err != nil { + return ctrl.Result{}, fmt.Errorf("enter step %s: %w", next, err) + } + snapshot := statemachine.ToKube(machine.Snapshot()) + return ctrl.Result{RequeueAfter: opsAdvance}, r.recordStep(ctx, ops, next, snapshot.Deadline) +} + +// terminalStepError is a step failure nothing recovers from by reconciling +// again: a precondition that cannot become true, or a verification that +// disagreed. +type terminalStepError struct{ message string } + +func (e *terminalStepError) Error() string { return e.message } + +// perform does what one step is for, and reports whether the step is finished. A +// step that is not finished returns what it is holding on, which lands in +// status.message. +func (r *ControlPlaneOpsReconciler) perform( + ctx context.Context, + ops *simplyblockv1alpha2.ControlPlaneOps, + target *simplyblockv1alpha2.ControlPlane, + current opsStep, +) (done bool, held string, err error) { + switch current { + case stepPreflight: + return r.preflight(ctx, ops, target) + case stepDraining: + return r.drain(ctx, ops) + case stepRestarting: + return r.restart(ctx, ops, target) + case stepApplying: + return r.applyUpgrade(ctx, ops, target) + case stepAwaiting: + return r.await(ctx, ops, target) + case stepVerifying: + return r.verify(ctx, ops, target) + case stepRequesting: + return r.requestBackup(ctx, ops, target) + default: + return false, "", &terminalStepError{message: fmt.Sprintf("unknown step %q", current)} + } +} + +// preflight reads live state, which admission cannot. It holds until the control +// plane is Available, and it fails when the requested image is the one already +// running, because rolling a Deployment to its current image produces no change +// to verify. +func (r *ControlPlaneOpsReconciler) preflight( + _ context.Context, + ops *simplyblockv1alpha2.ControlPlaneOps, + target *simplyblockv1alpha2.ControlPlane, +) (bool, string, error) { + if ops.Spec.Upgrade == nil || ops.Spec.Upgrade.Image == "" { + return false, "", &terminalStepError{ + message: "action Upgrade needs spec.upgrade.image, which names the version to move to", + } + } + if target.Status.Phase != simplyblockv1alpha2.ControlPlanePhaseAvailable { + return false, fmt.Sprintf( + "the control plane is %s; an upgrade starts from Available so that what it verifies "+ + "afterward is a change rather than a recovery", target.Status.Phase), nil + } + if localImage(target) == ops.Spec.Upgrade.Image { + return false, "", &terminalStepError{message: fmt.Sprintf( + "the control plane already runs %s, so there is no rollout to verify", + ops.Spec.Upgrade.Image)} + } + return true, "", nil +} + +// drain holds while any operation in the namespace is still running. +// +// The management API is what every controller in the operator talks to, so +// replacing or restarting it mid-flight fails whatever is in flight. It does not +// cancel them, because an operation canceled to make a restart convenient is a +// worse outcome than a restart that waited. +// +// A scoped restart drains only when it names something depended on: the list is +// empty, or it names a component the phase table marks essential. Recycling an +// exporter interrupts nothing, so a restart of it has nothing to wait for. +func (r *ControlPlaneOpsReconciler) drain( + ctx context.Context, ops *simplyblockv1alpha2.ControlPlaneOps, +) (bool, string, error) { + if !drainsFirst(ops) { + return true, "", nil + } + + inFlight, err := r.operationsInFlight(ctx, ops) + if err != nil { + return false, "", err + } + if len(inFlight) == 0 { + return true, "", nil + } + + message := fmt.Sprintf( + "%d operations are still running in %s, and recycling the management API would fail "+ + "whatever they have in flight: %v", len(inFlight), ops.Namespace, inFlight) + r.emit(ops, corev1.EventTypeNormal, OperationsInFlight, message) + return false, message, nil +} + +// drainsFirst reports whether this operation has to wait for in-flight work. +func drainsFirst(ops *simplyblockv1alpha2.ControlPlaneOps) bool { + if ops.Spec.Action != simplyblockv1alpha2.ControlPlaneOpsActionRestart { + return true + } + scope := restartScope(ops) + if len(scope) == 0 { + return true + } + essential := essentialComponents() + for _, name := range scope { + if essential[name] { + return true + } + } + return false +} + +// restartScope is the components a Restart names, or every restartable one when +// it names none. +func restartScope(ops *simplyblockv1alpha2.ControlPlaneOps) []string { + if ops.Spec.Restart == nil { + return nil + } + return ops.Spec.Restart.Components +} + +// operationsInFlight lists the operations in the namespace that are still +// Running. It reads the three kinds whose work a control-plane restart would +// interrupt, which is every Ops kind that reaches the control plane over HTTP. +func (r *ControlPlaneOpsReconciler) operationsInFlight( + ctx context.Context, ops *simplyblockv1alpha2.ControlPlaneOps, +) ([]string, error) { + var running []string + inNamespace := client.InNamespace(ops.Namespace) + + var clusters simplyblockv1alpha2.StorageClusterOpsList + if err := r.List(ctx, &clusters, inNamespace); err != nil { + return nil, err + } + for _, item := range clusters.Items { + if item.Status.Phase == simplyblockv1alpha2.StorageClusterOpsPhaseRunning { + running = append(running, "StorageClusterOps/"+item.Name) + } + } + + var nodes simplyblockv1alpha2.StorageNodeOpsList + if err := r.List(ctx, &nodes, inNamespace); err != nil { + return nil, err + } + for _, item := range nodes.Items { + if item.Status.Phase == simplyblockv1alpha2.StorageNodeOpsPhaseRunning { + running = append(running, "StorageNodeOps/"+item.Name) + } + } + + var pools simplyblockv1alpha2.StoragePoolOpsList + if err := r.List(ctx, &pools, inNamespace); err != nil { + return nil, err + } + for _, item := range pools.Items { + if item.Status.Phase == simplyblockv1alpha2.StoragePoolOpsPhaseRunning { + running = append(running, "StoragePoolOps/"+item.Name) + } + } + + slices.Sort(running) + return running, nil +} + +// restart recycles the workloads the operation named, by writing a new value +// into their pod templates. +// +// A component the table does not know is refused rather than skipped: an +// operation that reported success while recycling nothing is worse than one that +// says the name was wrong. +func (r *ControlPlaneOpsReconciler) restart( + ctx context.Context, + ops *simplyblockv1alpha2.ControlPlaneOps, + target *simplyblockv1alpha2.ControlPlane, +) (bool, string, error) { + scope := restartScope(ops) + for _, name := range scope { + if !restartable(name) { + return false, "", &terminalStepError{message: fmt.Sprintf( + "%q is not a workload this control plane can recycle", name)} + } + } + + stamp := metav1.Now().UTC().Format(time.RFC3339) + for _, component := range restartableComponents() { + if len(scope) > 0 && !slices.Contains(scope, component.name) { + continue + } + if err := r.recycle(ctx, target.Namespace, component, stamp); err != nil { + return false, "", err + } + } + return true, "", nil +} + +// recycle writes the restart stamp onto one workload's pod template. +func (r *ControlPlaneOpsReconciler) recycle( + ctx context.Context, namespace string, component component, stamp string, +) error { + key := client.ObjectKey{Namespace: namespace, Name: component.name} + + switch component.kind { + case kindDeployment: + var d appsv1.Deployment + if err := r.Get(ctx, key, &d); err != nil { + // A workload that is not there is one the entity's next pass + // re-applies, and recycling it is not this operation's problem. + return client.IgnoreNotFound(err) + } + base := d.DeepCopy() + stampTemplate(&d.Spec.Template, stamp) + return r.Patch(ctx, &d, client.MergeFrom(base)) + + case kindStatefulSet: + var s appsv1.StatefulSet + if err := r.Get(ctx, key, &s); err != nil { + return client.IgnoreNotFound(err) + } + base := s.DeepCopy() + stampTemplate(&s.Spec.Template, stamp) + return r.Patch(ctx, &s, client.MergeFrom(base)) + + default: + // The FoundationDBCluster is not restartable this way, and + // restartableComponents already excluded it. + return nil + } +} + +func stampTemplate(template *corev1.PodTemplateSpec, stamp string) { + if template.Annotations == nil { + template.Annotations = map[string]string{} + } + template.Annotations[restartedAtAnnotation] = stamp +} + +// applyUpgrade writes the new image onto the entity, which is what rolls the +// Deployment: the entity re-applies its workloads on every pass, and the image +// it applies is the one in its own spec. +// +// Writing it to the spec rather than to the Deployment is what keeps the entity +// describing what is running: the entity's next pass re-applies its workloads +// from the spec. +func (r *ControlPlaneOpsReconciler) applyUpgrade( + ctx context.Context, + ops *simplyblockv1alpha2.ControlPlaneOps, + target *simplyblockv1alpha2.ControlPlane, +) (bool, string, error) { + // Preflight read this block, and the CEL rule on the spec freezes it after + // admission, so reaching here without one means an object written before that + // rule shipped. It is a terminal failure rather than a panic. + if ops.Spec.Upgrade == nil || ops.Spec.Upgrade.Image == "" { + return false, "", &terminalStepError{ + message: "spec.upgrade.image is gone, so there is no version to move to", + } + } + + base := target.DeepCopy() + target.Spec.Source.Local.Image = ops.Spec.Upgrade.Image + if err := r.Patch(ctx, target, client.MergeFrom(base)); err != nil { + return false, "", err + } + return true, "", nil +} + +// await waits for what the previous step asked for: the recycled pods coming +// back, or the backup reporting a snapshot. +func (r *ControlPlaneOpsReconciler) await( + ctx context.Context, + ops *simplyblockv1alpha2.ControlPlaneOps, + target *simplyblockv1alpha2.ControlPlane, +) (bool, string, error) { + if ops.Spec.Action == simplyblockv1alpha2.ControlPlaneOpsActionBackup { + return r.awaitBackup(ctx, ops, target) + } + + components, err := observe(ctx, r.Client, target.Namespace) + if err != nil { + return false, "", err + } + scope := restartScope(ops) + for _, status := range components { + if len(scope) > 0 && !slices.Contains(scope, status.Name) { + continue + } + if status.Desired > 0 && status.Ready < status.Desired { + return false, fmt.Sprintf("%s has %d of %d replicas ready", + status.Name, status.Ready, status.Desired), nil + } + } + return true, "", nil +} + +// verify is what makes an upgrade more than an image bump. It re-probes +// readiness, compares the reported version against what was asked for, and fails +// the operation when they disagree. +// +// A control plane that reports no version at all is a control plane whose +// /_meta/version read does not exist yet, which design-controlplane.md §8 +// records as a prerequisite. The step passes there rather than failing, and says +// so in the message, so the record of the operation carries what was and was not +// verified. +func (r *ControlPlaneOpsReconciler) verify( + ctx context.Context, + ops *simplyblockv1alpha2.ControlPlaneOps, + target *simplyblockv1alpha2.ControlPlane, +) (bool, string, error) { + endpoint := target.Status.Endpoint + if endpoint == "" { + endpoint = localEndpoint(target.Namespace) + } + + prober := r.Prober + if prober == nil { + prober = &HTTPProber{} + } + if ok, reason := prober.Ready(ctx, endpoint); !ok { + return false, fmt.Sprintf("the control plane is not answering yet: %s", reason), nil + } + + reported, err := prober.Version(ctx, endpoint) + if err != nil { + return false, fmt.Sprintf("the version could not be read: %v", err), nil + } + if reported == "" { + return true, "", r.note(ctx, ops, + "the rollout finished; the version was not verified because the control plane "+ + "does not serve a version endpoint") + } + if ops.Spec.Upgrade == nil { + return false, "", &terminalStepError{ + message: "spec.upgrade.image is gone, so there is nothing to verify the rollout against", + } + } + if !imageStates(ops.Spec.Upgrade.Image, reported) { + r.emit(ops, corev1.EventTypeWarning, VersionMismatch, fmt.Sprintf( + "the control plane reports %s and the upgrade asked for %s", + reported, ops.Spec.Upgrade.Image)) + return false, "", &terminalStepError{message: fmt.Sprintf( + "the rollout finished with the control plane reporting %s rather than %s, which is "+ + "a rollout that failed back rather than an upgrade that completed", + reported, ops.Spec.Upgrade.Image)} + } + return true, "", nil +} + +// imageStates reports whether an image reference names the version the control +// plane reports. The comparison is against the image's tag rather than the whole +// reference, because a version endpoint answers with a version and an image +// carries a registry and a repository in front of it. +// +// The digest is cut off before the tag is read, and the order is the whole of +// the correctness: a digest carries its own colon, so reading the tag from the +// last colon of repo:tag@sha256:… yields the hex digest rather than the tag. +// +// A digest-pinned image is then not compared at all. What is pinned by digest is +// not claimed to be any particular version, so an upgrade to one verifies that +// the rollout finished and stops there. +func imageStates(image, version string) bool { + reference := image + if at := strings.LastIndexByte(reference, '@'); at >= 0 { + return true + } + + colon := strings.LastIndexByte(reference, ':') + if colon < 0 { + // No tag at all, which the spec's pattern does not admit. There is + // nothing to compare, so the rollout finishing is the whole of the + // verification. + return true + } + return reference[colon+1:] == version +} + +// requestBackup creates or triggers the FoundationDBBackup. +// +// The object outlives the operation and is not owned by it: deleting the record +// of a backup having been taken must not delete the backup's configuration. An +// operation that finds one already configured triggers a snapshot on it rather +// than applying a second beside it. +func (r *ControlPlaneOpsReconciler) requestBackup( + ctx context.Context, + ops *simplyblockv1alpha2.ControlPlaneOps, + target *simplyblockv1alpha2.ControlPlane, +) (bool, string, error) { + if ops.Spec.Backup == nil || ops.Spec.Backup.BlobStore == "" { + return false, "", &terminalStepError{ + message: "action Backup needs spec.backup.blobStore, which names where the backup goes", + } + } + + name := ops.Spec.Backup.BackupName + if name == "" { + name = ComponentFDBCluster + } + + existing := &unstructured.Unstructured{} + existing.SetGroupVersionKind(fdbBackupGVK) + key := client.ObjectKey{Namespace: target.Namespace, Name: name} + err := r.Get(ctx, key, existing) + switch { + case apierrors.IsNotFound(err): + backup := &unstructured.Unstructured{Object: map[string]any{ + "spec": map[string]any{ + "clusterName": ComponentFDBCluster, + "blobStoreConfiguration": map[string]any{ + "accountName": ops.Spec.Backup.BlobStore, + }, + }, + }} + backup.SetGroupVersionKind(fdbBackupGVK) + backup.SetName(name) + backup.SetNamespace(target.Namespace) + if err := r.Create(ctx, backup); err != nil { + return false, "", err + } + r.emit(ops, corev1.EventTypeNormal, BackupRequested, + fmt.Sprintf("created FoundationDBBackup %s", name)) + + case err != nil: + return false, "", err + + default: + // Running is what tells the FoundationDB operator to take and keep + // taking snapshots. An object already in that state needs nothing + // written to it, and writing anyway would restart a backup mid-flight. + state, _, _ := unstructured.NestedString(existing.Object, "spec", "backupState") + if state != "Running" { + base := existing.DeepCopy() + if err := unstructured.SetNestedField(existing.Object, "Running", "spec", "backupState"); err != nil { + return false, "", err + } + if err := r.Patch(ctx, existing, client.MergeFrom(base)); err != nil { + return false, "", err + } + } + r.emit(ops, corev1.EventTypeNormal, BackupTriggered, + fmt.Sprintf("triggered the FoundationDBBackup %s that was already configured", name)) + } + + return true, "", r.recordBackupRef(ctx, ops, name) +} + +// awaitBackup completes when the backup reports that it is running, which is +// what a FoundationDBBackup says when its agents have started and it is taking +// snapshots. +func (r *ControlPlaneOpsReconciler) awaitBackup( + ctx context.Context, + ops *simplyblockv1alpha2.ControlPlaneOps, + target *simplyblockv1alpha2.ControlPlane, +) (bool, string, error) { + backup := &unstructured.Unstructured{} + backup.SetGroupVersionKind(fdbBackupGVK) + key := client.ObjectKey{Namespace: target.Namespace, Name: ops.Status.BackupRef} + if err := r.Get(ctx, key, backup); err != nil { + if apierrors.IsNotFound(err) { + return false, fmt.Sprintf("FoundationDBBackup %s has not appeared yet", + ops.Status.BackupRef), nil + } + return false, "", err + } + + running, _, _ := unstructured.NestedBool(backup.Object, "status", "running") + if !running { + return false, fmt.Sprintf("FoundationDBBackup %s has not started taking snapshots yet", + ops.Status.BackupRef), nil + } + return true, "", nil +} + +// unwind honors spec.abort where the graph allows it, and reports an abort that +// arrived too late rather than half-undoing the work. +// +// The refusal is the point. A step past the drain has rolled a Deployment or +// written an image onto the entity, and this operation is what drives that +// rollout to completion. +func (r *ControlPlaneOpsReconciler) unwind( + ctx context.Context, + ops *simplyblockv1alpha2.ControlPlaneOps, + target *simplyblockv1alpha2.ControlPlane, + current opsStep, +) (ctrl.Result, error) { + if !abortable(current) { + // Not a failure of the operation: it carries on. What the user asked for + // cannot be done, and saying so is the whole of the response. + return ctrl.Result{RequeueAfter: opsRetry}, r.note(ctx, ops, fmt.Sprintf( + "the abort arrived at step %s, which has already started a rollout that nothing "+ + "else would finish; the operation is running on", current)) + } + + // Nothing before the drain performs a side effect, so there is nothing to + // unwind: the abort is the lock being released and a terminal phase. + if err := r.releaseLock(ctx, ops, target); err != nil { + return ctrl.Result{}, err + } + r.emit(ops, corev1.EventTypeNormal, OperationAborted, + fmt.Sprintf("aborted at step %s, before anything was changed", current)) + return r.finish(ctx, ops, target, simplyblockv1alpha2.ControlPlaneOpsPhaseAborted, + fmt.Sprintf("aborted at step %s", current)) +} + +// enterInitialStep sets the first step's deadline, which a machine born already +// in that step would otherwise never get. +func (r *ControlPlaneOpsReconciler) enterInitialStep( + ctx context.Context, + ops *simplyblockv1alpha2.ControlPlaneOps, + machine *statemachine.Machine[opsStep], +) (ctrl.Result, error) { + budget, ok := opsInitialDeadlines[action(ops.Spec.Action)] + if !ok { + budget = requestingDeadline + } + stepDeadline := metav1.NewTime(time.Now().Add(budget)) + return ctrl.Result{RequeueAfter: opsAdvance}, + r.recordStep(ctx, ops, machine.CurrentState(), &stepDeadline) +} + +// nextOpsStep is the step that follows the current one. Every graph in this +// package is a line, so the first edge is the only edge. +func nextOpsStep(machine *statemachine.Machine[opsStep]) (opsStep, bool) { + for next := range machine.AllowedTransitions() { + return next, true + } + return machine.CurrentState(), false +} + +// releaseLock clears the target's lock, and only when this operation is the one +// holding it. The guard is the whole of the safety: an operation that never +// acquired the lock, or whose lock was taken over, must not clear somebody +// else's. +func (r *ControlPlaneOpsReconciler) releaseLock( + ctx context.Context, + ops *simplyblockv1alpha2.ControlPlaneOps, + target *simplyblockv1alpha2.ControlPlane, +) error { + if target == nil || target.Status.ActiveOpsRef != ops.Name { + return nil + } + base := target.DeepCopy() + target.Status.ActiveOpsRef = "" + if err := r.Status().Patch(ctx, target, client.MergeFrom(base)); err != nil { + return fmt.Errorf("release the lock on control plane %s/%s: %w", + target.Namespace, target.Name, err) + } + return nil +} + +func (r *ControlPlaneOpsReconciler) failAndRelease( + ctx context.Context, + ops *simplyblockv1alpha2.ControlPlaneOps, + target *simplyblockv1alpha2.ControlPlane, + message string, +) (ctrl.Result, error) { + if err := r.releaseLock(ctx, ops, target); err != nil { + return ctrl.Result{}, err + } + r.emit(ops, corev1.EventTypeWarning, OperationFailed, message) + return r.finish(ctx, ops, target, simplyblockv1alpha2.ControlPlaneOpsPhaseFailed, message) +} + +// reconcileDeletion releases the lock the operation may still hold, then lets the +// object go. This is the path that matters most: an operation deleted while +// Running has taken a lock that nothing else would ever clear. +func (r *ControlPlaneOpsReconciler) reconcileDeletion( + ctx context.Context, ops *simplyblockv1alpha2.ControlPlaneOps, +) (ctrl.Result, error) { + if !controllerutil.ContainsFinalizer(ops, FinalizerControlPlaneOps) { + return ctrl.Result{}, nil + } + + target := &simplyblockv1alpha2.ControlPlane{} + key := client.ObjectKey{Namespace: ops.Namespace, Name: ops.Spec.ControlPlaneRef} + switch err := r.Get(ctx, key, target); { + case apierrors.IsNotFound(err): + // The control plane went first, so there is no lock left to release. + case err != nil: + return ctrl.Result{}, err + default: + if err := r.releaseLock(ctx, ops, target); err != nil { + return ctrl.Result{}, err + } + } + + base := ops.DeepCopy() + controllerutil.RemoveFinalizer(ops, FinalizerControlPlaneOps) + if err := r.Patch(ctx, ops, client.MergeFrom(base)); err != nil { + return ctrl.Result{}, fmt.Errorf("release the finalizer on operation %s/%s: %w", + ops.Namespace, ops.Name, err) + } + return ctrl.Result{}, nil +} + +func (r *ControlPlaneOpsReconciler) ensureFinalizer( + ctx context.Context, ops *simplyblockv1alpha2.ControlPlaneOps, +) error { + if controllerutil.ContainsFinalizer(ops, FinalizerControlPlaneOps) { + return nil + } + base := ops.DeepCopy() + controllerutil.AddFinalizer(ops, FinalizerControlPlaneOps) + return r.Patch(ctx, ops, client.MergeFrom(base)) +} + +// finish writes a terminal phase and the time it was reached, and records the +// operation in the two cumulative series. +func (r *ControlPlaneOpsReconciler) finish( + ctx context.Context, + ops *simplyblockv1alpha2.ControlPlaneOps, + target *simplyblockv1alpha2.ControlPlane, + phase simplyblockv1alpha2.ControlPlaneOpsPhase, + message string, +) (ctrl.Result, error) { + now := metav1.Now() + base := ops.DeepCopy() + ops.Status.Phase = phase + ops.Status.Message = message + ops.Status.CompletedAt = &now + ops.Status.ObservedGeneration = ops.Generation + if err := r.Status().Patch(ctx, ops, client.MergeFrom(base)); err != nil { + return ctrl.Result{}, err + } + if phase == simplyblockv1alpha2.ControlPlaneOpsPhaseSucceeded { + r.emit(ops, corev1.EventTypeNormal, OperationSucceeded, message) + } + recordOperation(ops, target, phase) + return ctrl.Result{}, nil +} + +// recordOperation adds one terminal operation to the counter, and its duration +// to the histogram when the operation got far enough to have one. An operation +// that never started has no duration to report, and reporting zero would drag +// the distribution toward a value nothing took. +func recordOperation( + ops *simplyblockv1alpha2.ControlPlaneOps, + target *simplyblockv1alpha2.ControlPlane, + phase simplyblockv1alpha2.ControlPlaneOpsPhase, +) { + namespace := ops.Namespace + if target != nil { + namespace = target.Namespace + } + act := string(ops.Spec.Action) + controlPlaneOperationsTotal.WithLabelValues(namespace, act, string(phase)).Inc() + if ops.Status.StartedAt != nil && ops.Status.CompletedAt != nil { + seconds := ops.Status.CompletedAt.Sub(ops.Status.StartedAt.Time).Seconds() + controlPlaneOperationDuration.WithLabelValues(namespace, act).Observe(seconds) + } +} + +// recordStep persists the step and its deadline before the side effect that step +// performs. +func (r *ControlPlaneOpsReconciler) recordStep( + ctx context.Context, + ops *simplyblockv1alpha2.ControlPlaneOps, + next opsStep, + stepDeadline *metav1.Time, +) error { + base := ops.DeepCopy() + ops.Status.Phase = simplyblockv1alpha2.ControlPlaneOpsPhaseRunning + ops.Status.Step = statemachine.KubeSnapshot{State: string(next), Deadline: stepDeadline} + ops.Status.ObservedGeneration = ops.Generation + return r.Status().Patch(ctx, ops, client.MergeFrom(base)) +} + +func (r *ControlPlaneOpsReconciler) recordBackupRef( + ctx context.Context, ops *simplyblockv1alpha2.ControlPlaneOps, name string, +) error { + if ops.Status.BackupRef == name { + return nil + } + base := ops.DeepCopy() + ops.Status.BackupRef = name + return r.Status().Patch(ctx, ops, client.MergeFrom(base)) +} + +// note records what the operation is holding on, without moving it. A held step +// is correct behavior waiting on something, and the message is how somebody +// reading the object learns what. +func (r *ControlPlaneOpsReconciler) note( + ctx context.Context, ops *simplyblockv1alpha2.ControlPlaneOps, message string, +) error { + if ops.Status.Message == message { + return nil + } + base := ops.DeepCopy() + ops.Status.Message = message + return r.Status().Patch(ctx, ops, client.MergeFrom(base)) +} + +func (r *ControlPlaneOpsReconciler) setPhase( + ctx context.Context, + ops *simplyblockv1alpha2.ControlPlaneOps, + phase simplyblockv1alpha2.ControlPlaneOpsPhase, + message string, +) error { + base := ops.DeepCopy() + ops.Status.Phase = phase + ops.Status.Message = message + ops.Status.ObservedGeneration = ops.Generation + return r.Status().Patch(ctx, ops, client.MergeFrom(base)) +} + +func (r *ControlPlaneOpsReconciler) emit( + object client.Object, eventType, reason, message string, +) { + if r.Recorder == nil { + return + } + r.Recorder.Eventf(object, nil, eventType, reason, reason, "%s", message) +} + +// terminalOps reports whether a phase is one the operation never leaves. +func terminalOps(phase simplyblockv1alpha2.ControlPlaneOpsPhase) bool { + switch phase { + case simplyblockv1alpha2.ControlPlaneOpsPhaseSucceeded, + simplyblockv1alpha2.ControlPlaneOpsPhaseFailed, + simplyblockv1alpha2.ControlPlaneOpsPhaseAborted: + return true + default: + return false + } +} + +// SetupWithManager registers the reconciler. +func (r *ControlPlaneOpsReconciler) SetupWithManager(mgr ctrl.Manager) error { + return ctrl.NewControllerManagedBy(mgr). + For(&simplyblockv1alpha2.ControlPlaneOps{}). + Named("controlplaneops"). + Complete(r) +} diff --git a/operator/internal/controllers/controlplane/controlplaneops_test.go b/operator/internal/controllers/controlplane/controlplaneops_test.go new file mode 100644 index 000000000..998013ddf --- /dev/null +++ b/operator/internal/controllers/controlplane/controlplaneops_test.go @@ -0,0 +1,615 @@ +// The operations: the lock, the drain, the abort line, and the two steps that +// are more than the call they make. +// +// Preflight and Verifying are the pair worth testing hardest. Preflight is what +// stops an upgrade that has nothing to do, and Verifying is what makes an +// upgrade more than an image bump: without it a rollout that started and failed +// back is a Succeeded operation against a control plane running the old version. + +package controlplane + +import ( + "context" + "errors" + "strings" + "testing" + + appsv1 "k8s.io/api/apps/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + + "github.com/simplyblock/atlas/statemachine" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// anotherOperation is the name the lock tests use for an operation that is not +// the one under test, so the assertions read as "somebody else's" rather than as +// a string repeated four times. +const anotherOperation = "somebody-elses-operation" + +// opsFor builds one operation against the singleton. +func opsFor(action simplyblockv1alpha2.ControlPlaneOpsAction) *simplyblockv1alpha2.ControlPlaneOps { + return &simplyblockv1alpha2.ControlPlaneOps{ + ObjectMeta: metav1.ObjectMeta{Name: "an-operation", Namespace: testNamespace}, + Spec: simplyblockv1alpha2.ControlPlaneOpsSpec{ + ControlPlaneRef: SingletonName, + Action: action, + }, + } +} + +// Every action declares a graph, and every graph is a line. The reconciler takes +// the first edge as the only edge, so a graph that branched would silently pick +// one. +func TestEveryActionsGraphIsALine(t *testing.T) { + graphs := opsGraphs() + + for _, act := range []simplyblockv1alpha2.ControlPlaneOpsAction{ + simplyblockv1alpha2.ControlPlaneOpsActionRestart, + simplyblockv1alpha2.ControlPlaneOpsActionUpgrade, + simplyblockv1alpha2.ControlPlaneOpsActionBackup, + } { + graph, declared := graphs[action(act)] + if !declared { + t.Fatalf("%s declares no graph, so the action cannot run at all", act) + } + terminals := 0 + for step, state := range graph.States { + switch len(state.To) { + case 0: + terminals++ + case 1: + default: + t.Errorf("%s: step %s declares %d successors, want one", act, step, len(state.To)) + } + } + if terminals != 1 { + t.Errorf("%s declares %d terminal steps, want exactly one", act, terminals) + } + if _, ok := opsInitialDeadlines[action(act)]; !ok { + t.Errorf("%s has no initial deadline, so its first step cannot time out", act) + } + } +} + +// Every abortable step is one some graph declares. The table sits beside the +// graphs rather than in them, so a step renamed in one and not the other would +// make an abort silently unreachable. +func TestEveryAbortableStepIsAStepSomeGraphDeclares(t *testing.T) { + declared := map[string]bool{} + for _, state := range statemachine.DeclaredMultiStates(opsGraphs()) { + declared[state] = true + } + for step := range abortableSteps { + if !declared[string(step)] { + t.Errorf("%s is abortable and no graph declares it", step) + } + } +} + +// An abort is honored only before anything has been changed. A step that has +// rolled a Deployment or written an image onto the entity carries on, because +// stopping there would leave a rollout half-done with nothing driving it either +// way. +func TestAnAbortIsRefusedOnceARolloutHasStarted(t *testing.T) { + refused := []opsStep{stepRestarting, stepApplying, stepAwaiting, stepVerifying} + for _, step := range refused { + if abortable(step) { + t.Errorf("%s is abortable, and an abort there leaves a rollout half-done", step) + } + } + for _, step := range []opsStep{stepDraining, stepPreflight, stepRequesting} { + if !abortable(step) { + t.Errorf("%s is not abortable, and nothing has been changed at that point", step) + } + } +} + +// An operation naming a remote control plane is refused rather than run. The +// webhook catches it at creation, and this is what holds when the webhook was +// not serving. +func TestAnOperationAgainstARemoteControlPlaneFails(t *testing.T) { + cp := managedControlPlane("https://sb-control.example.com:5000") + ops := opsFor(simplyblockv1alpha2.ControlPlaneOpsActionRestart) + c := newClient(t, cp, ops) + r := &ControlPlaneOpsReconciler{Client: c, Scheme: testScheme(t)} + + if _, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: client.ObjectKeyFromObject(ops), + }); err != nil { + t.Fatalf("Reconcile: %v", err) + } + + var after simplyblockv1alpha2.ControlPlaneOps + if err := c.Get(context.Background(), client.ObjectKeyFromObject(ops), &after); err != nil { + t.Fatalf("read the operation back: %v", err) + } + if after.Status.Phase != simplyblockv1alpha2.ControlPlaneOpsPhaseFailed { + t.Errorf("status.phase = %s, want Failed", after.Status.Phase) + } + if !strings.Contains(after.Status.Message, "does not host") { + t.Errorf("status.message = %q, want it to say why the target cannot be operated on", + after.Status.Message) + } +} + +// An operation naming a control plane that does not exist fails with the name it +// could not resolve, rather than retrying against an object that will never +// appear. +func TestAnOperationWithNoTargetFails(t *testing.T) { + ops := opsFor(simplyblockv1alpha2.ControlPlaneOpsActionRestart) + ops.Spec.ControlPlaneRef = "not-the-singleton" + c := newClient(t, ops) + r := &ControlPlaneOpsReconciler{Client: c, Scheme: testScheme(t)} + + if _, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: client.ObjectKeyFromObject(ops), + }); err != nil { + t.Fatalf("Reconcile: %v", err) + } + + var after simplyblockv1alpha2.ControlPlaneOps + if err := c.Get(context.Background(), client.ObjectKeyFromObject(ops), &after); err != nil { + t.Fatalf("read the operation back: %v", err) + } + if after.Status.Phase != simplyblockv1alpha2.ControlPlaneOpsPhaseFailed { + t.Errorf("status.phase = %s, want Failed", after.Status.Phase) + } + if !strings.Contains(after.Status.Message, "not-the-singleton") { + t.Errorf("status.message = %q, want it to name what could not be resolved", + after.Status.Message) + } +} + +// A second operation stays Pending and asks again rather than failing, so the +// outcome does not depend on the order two objects were applied in. +func TestASecondOperationWaitsForTheLock(t *testing.T) { + cp := localControlPlane() + cp.Status.ActiveOpsRef = anotherOperation + ops := opsFor(simplyblockv1alpha2.ControlPlaneOpsActionRestart) + c := newClient(t, cp, ops) + recorder := &recordingRecorder{} + r := &ControlPlaneOpsReconciler{Client: c, Scheme: testScheme(t), Recorder: recorder} + + result, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: client.ObjectKeyFromObject(ops), + }) + if err != nil { + t.Fatalf("Reconcile: %v", err) + } + if result.RequeueAfter == 0 { + t.Error("the operation was not requeued, so it would never take the lock when it frees") + } + + var after simplyblockv1alpha2.ControlPlaneOps + if err := c.Get(context.Background(), client.ObjectKeyFromObject(ops), &after); err != nil { + t.Fatalf("read the operation back: %v", err) + } + if after.Status.Phase != simplyblockv1alpha2.ControlPlaneOpsPhasePending { + t.Errorf("status.phase = %s, want Pending", after.Status.Phase) + } + if recorder.count(OperationQueued) == 0 { + t.Error("no OperationQueued event: nothing says why the operation is not running") + } + + var target simplyblockv1alpha2.ControlPlane + if err := c.Get(context.Background(), client.ObjectKeyFromObject(cp), &target); err != nil { + t.Fatalf("read the control plane back: %v", err) + } + if target.Status.ActiveOpsRef != anotherOperation { + t.Errorf("activeOpsRef = %q, want the lock left with its holder", + target.Status.ActiveOpsRef) + } +} + +// Deleting an operation while it holds the lock releases the lock, so the +// control plane is never left locked by an object that no longer exists. +func TestDeletingARunningOperationReleasesTheLock(t *testing.T) { + ctx := context.Background() + cp := localControlPlane() + cp.Status.ActiveOpsRef = "an-operation" + + ops := opsFor(simplyblockv1alpha2.ControlPlaneOpsActionRestart) + ops.Finalizers = []string{FinalizerControlPlaneOps} + now := metav1.Now() + ops.DeletionTimestamp = &now + ops.Status.Phase = simplyblockv1alpha2.ControlPlaneOpsPhaseRunning + + c := newClient(t, cp, ops) + r := &ControlPlaneOpsReconciler{Client: c, Scheme: testScheme(t)} + + if _, err := r.Reconcile(ctx, ctrl.Request{ + NamespacedName: client.ObjectKeyFromObject(ops), + }); err != nil { + t.Fatalf("Reconcile: %v", err) + } + + var target simplyblockv1alpha2.ControlPlane + if err := c.Get(ctx, client.ObjectKeyFromObject(cp), &target); err != nil { + t.Fatalf("read the control plane back: %v", err) + } + if target.Status.ActiveOpsRef != "" { + t.Errorf("activeOpsRef = %q, want the lock released with the operation", + target.Status.ActiveOpsRef) + } +} + +// The lock is released only by the operation holding it. An operation that never +// acquired it, or whose lock was taken over, must not clear somebody else's. +func TestAnOperationDoesNotReleaseSomebodyElsesLock(t *testing.T) { + ctx := context.Background() + cp := localControlPlane() + cp.Status.ActiveOpsRef = anotherOperation + ops := opsFor(simplyblockv1alpha2.ControlPlaneOpsActionRestart) + + c := newClient(t, cp, ops) + r := &ControlPlaneOpsReconciler{Client: c, Scheme: testScheme(t)} + + if err := r.releaseLock(ctx, ops, cp); err != nil { + t.Fatalf("releaseLock: %v", err) + } + + var target simplyblockv1alpha2.ControlPlane + if err := c.Get(ctx, client.ObjectKeyFromObject(cp), &target); err != nil { + t.Fatalf("read the control plane back: %v", err) + } + if target.Status.ActiveOpsRef != anotherOperation { + t.Errorf("activeOpsRef = %q, want somebody else's lock untouched", + target.Status.ActiveOpsRef) + } +} + +// A scoped restart drains only when it names something depended on. Recycling an +// exporter interrupts nothing, so a restart of it has nothing to wait for; the +// task runner is the case that draws the line, and it skips the drain because +// its queue makes the interruption a delay rather than a lost operation. +func TestARestartDrainsOnlyWhenItNamesSomethingEssential(t *testing.T) { + for _, tc := range []struct { + name string + components []string + wantDrain bool + }{ + {"the whole control plane", nil, true}, + {"the management API", []string{ComponentWebAPI}, true}, + {"the database", []string{ComponentFDBCluster}, true}, + {"the task runner", []string{ComponentTasks}, false}, + {"the exporter", []string{ComponentFDBExporter}, false}, + {"a non-essential set", []string{ComponentTasks, ComponentMinio}, false}, + {"a set including an essential one", []string{ComponentTasks, ComponentWebAPI}, true}, + } { + t.Run(tc.name, func(t *testing.T) { + ops := opsFor(simplyblockv1alpha2.ControlPlaneOpsActionRestart) + if tc.components != nil { + ops.Spec.Restart = &simplyblockv1alpha2.RestartSpec{Components: tc.components} + } + + if got := drainsFirst(ops); got != tc.wantDrain { + t.Errorf("drainsFirst = %v, want %v", got, tc.wantDrain) + } + }) + } +} + +// An upgrade always drains, whatever it names, because it rolls the management +// API by definition. +func TestAnUpgradeAlwaysDrains(t *testing.T) { + ops := opsFor(simplyblockv1alpha2.ControlPlaneOpsActionUpgrade) + if !drainsFirst(ops) { + t.Error("an upgrade skipped the drain, and it rolls the management API by definition") + } +} + +// The drain holds while another operation is running, and names what it is +// holding on so somebody reading the object knows what to wait for. +func TestTheDrainHoldsOnOperationsInFlight(t *testing.T) { + ctx := context.Background() + ops := opsFor(simplyblockv1alpha2.ControlPlaneOpsActionRestart) + inFlight := &simplyblockv1alpha2.StorageNodeOps{ + ObjectMeta: metav1.ObjectMeta{Name: "a-node-add", Namespace: testNamespace}, + Status: simplyblockv1alpha2.StorageNodeOpsStatus{ + Phase: simplyblockv1alpha2.StorageNodeOpsPhaseRunning, + }, + } + + c := newClient(t, ops, inFlight) + recorder := &recordingRecorder{} + r := &ControlPlaneOpsReconciler{Client: c, Scheme: testScheme(t), Recorder: recorder} + + done, held, err := r.drain(ctx, ops) + if err != nil { + t.Fatalf("drain: %v", err) + } + if done { + t.Fatal("the drain finished while a node operation was still running") + } + if !strings.Contains(held, "StorageNodeOps/a-node-add") { + t.Errorf("held on %q, want it to name the operation in flight", held) + } + if recorder.count(OperationsInFlight) == 0 { + t.Error("no OperationsInFlight event: nothing says why the restart is waiting") + } +} + +// A terminal operation elsewhere does not hold the drain: what is being waited +// for is work in flight, and a finished operation has none. +func TestTheDrainDoesNotHoldOnFinishedOperations(t *testing.T) { + ops := opsFor(simplyblockv1alpha2.ControlPlaneOpsActionRestart) + finished := &simplyblockv1alpha2.StorageNodeOps{ + ObjectMeta: metav1.ObjectMeta{Name: "a-finished-add", Namespace: testNamespace}, + Status: simplyblockv1alpha2.StorageNodeOpsStatus{ + Phase: simplyblockv1alpha2.StorageNodeOpsPhaseSucceeded, + }, + } + + c := newClient(t, ops, finished) + r := &ControlPlaneOpsReconciler{Client: c, Scheme: testScheme(t)} + + done, held, err := r.drain(context.Background(), ops) + if err != nil { + t.Fatalf("drain: %v", err) + } + if !done { + t.Errorf("the drain held on a finished operation: %s", held) + } +} + +// Preflight refuses an upgrade naming the image already running, because rolling +// a Deployment to its current image produces no change to verify. +func TestPreflightRefusesAnUpgradeToTheRunningImage(t *testing.T) { + cp := localControlPlane() + cp.Status.Phase = simplyblockv1alpha2.ControlPlanePhaseAvailable + + ops := opsFor(simplyblockv1alpha2.ControlPlaneOpsActionUpgrade) + ops.Spec.Upgrade = &simplyblockv1alpha2.UpgradeSpec{Image: testImage} + + r := &ControlPlaneOpsReconciler{Client: newClient(t, cp, ops), Scheme: testScheme(t)} + + _, _, err := r.preflight(context.Background(), ops, cp) + var fatal *terminalStepError + if !errors.As(err, &fatal) { + t.Fatalf("preflight returned %v, want a terminal failure", err) + } + if !strings.Contains(fatal.Error(), testImage) { + t.Errorf("the refusal is %q, want it to name the image already running", fatal.Error()) + } +} + +// Preflight holds rather than fails while the control plane is not Available, so +// that what the upgrade verifies afterward is a change rather than a recovery. +func TestPreflightHoldsWhileTheControlPlaneIsNotAvailable(t *testing.T) { + cp := localControlPlane() + cp.Status.Phase = simplyblockv1alpha2.ControlPlanePhaseUnavailable + + ops := opsFor(simplyblockv1alpha2.ControlPlaneOpsActionUpgrade) + ops.Spec.Upgrade = &simplyblockv1alpha2.UpgradeSpec{ + Image: "quay.io/simplyblock-io/simplyblock:26.3.0", + } + + r := &ControlPlaneOpsReconciler{Client: newClient(t, cp, ops), Scheme: testScheme(t)} + + done, held, err := r.preflight(context.Background(), ops, cp) + if err != nil { + t.Fatalf("preflight: %v", err) + } + if done { + t.Fatal("preflight passed against an Unavailable control plane") + } + if !strings.Contains(held, string(simplyblockv1alpha2.ControlPlanePhaseUnavailable)) { + t.Errorf("held on %q, want it to name the phase it is waiting to leave", held) + } +} + +// Applying an upgrade writes the image onto the entity rather than onto the +// Deployment, because the entity re-applies its workloads from its own spec on +// every pass. +func TestAnUpgradeWritesTheImageOntoTheEntity(t *testing.T) { + ctx := context.Background() + const next = "quay.io/simplyblock-io/simplyblock:26.3.0" + + cp := localControlPlane() + ops := opsFor(simplyblockv1alpha2.ControlPlaneOpsActionUpgrade) + ops.Spec.Upgrade = &simplyblockv1alpha2.UpgradeSpec{Image: next} + + c := newClient(t, cp, ops) + r := &ControlPlaneOpsReconciler{Client: c, Scheme: testScheme(t)} + + if _, _, err := r.applyUpgrade(ctx, ops, cp); err != nil { + t.Fatalf("applyUpgrade: %v", err) + } + + var after simplyblockv1alpha2.ControlPlane + if err := c.Get(ctx, client.ObjectKeyFromObject(cp), &after); err != nil { + t.Fatalf("read the control plane back: %v", err) + } + if got := localImage(&after); got != next { + t.Errorf("spec.source.local.image = %q, want %q", got, next) + } +} + +// Verifying fails the operation when the reported version disagrees with what +// was asked for, which is what separates an upgrade that completed from a +// rollout that failed back. +func TestVerifyingFailsOnAVersionThatDisagrees(t *testing.T) { + cp := localControlPlane() + cp.Status.Endpoint = "http://simplyblock-webappapi.simplyblock.svc.cluster.local:5000" + + ops := opsFor(simplyblockv1alpha2.ControlPlaneOpsActionUpgrade) + ops.Spec.Upgrade = &simplyblockv1alpha2.UpgradeSpec{ + Image: "quay.io/simplyblock-io/simplyblock:26.3.0", + } + + recorder := &recordingRecorder{} + r := &ControlPlaneOpsReconciler{ + Client: newClient(t, cp, ops), + Scheme: testScheme(t), + Recorder: recorder, + Prober: &stubProber{ready: true, version: "26.2.8"}, + } + + _, _, err := r.verify(context.Background(), ops, cp) + var fatal *terminalStepError + if !errors.As(err, &fatal) { + t.Fatalf("verify returned %v, want a terminal failure", err) + } + if !strings.Contains(fatal.Error(), "26.2.8") { + t.Errorf("the failure is %q, want it to name what the control plane reported", + fatal.Error()) + } + if recorder.count(VersionMismatch) == 0 { + t.Error("no VersionMismatch event on a rollout that failed back") + } +} + +// Verifying passes when the reported version is the one asked for. +func TestVerifyingPassesOnTheVersionThatWasAskedFor(t *testing.T) { + cp := localControlPlane() + ops := opsFor(simplyblockv1alpha2.ControlPlaneOpsActionUpgrade) + ops.Spec.Upgrade = &simplyblockv1alpha2.UpgradeSpec{ + Image: "quay.io/simplyblock-io/simplyblock:26.3.0", + } + + r := &ControlPlaneOpsReconciler{ + Client: newClient(t, cp, ops), + Scheme: testScheme(t), + Prober: &stubProber{ready: true, version: "26.3.0"}, + } + + done, held, err := r.verify(context.Background(), ops, cp) + if err != nil { + t.Fatalf("verify: %v", err) + } + if !done { + t.Errorf("verify held on %q against the version it asked for", held) + } +} + +// A control plane that serves no version endpoint passes verification rather +// than failing it, and says so: failing every upgrade on a deployment that +// cannot answer would make the action unusable, and the record of the operation +// has to carry what was and was not verified. +func TestVerifyingPassesAndSaysSoWhenNoVersionIsServed(t *testing.T) { + ctx := context.Background() + cp := localControlPlane() + ops := opsFor(simplyblockv1alpha2.ControlPlaneOpsActionUpgrade) + ops.Spec.Upgrade = &simplyblockv1alpha2.UpgradeSpec{ + Image: "quay.io/simplyblock-io/simplyblock:26.3.0", + } + + c := newClient(t, cp, ops) + r := &ControlPlaneOpsReconciler{ + Client: c, Scheme: testScheme(t), + Prober: &stubProber{ready: true, version: ""}, + } + + done, _, err := r.verify(ctx, ops, cp) + if err != nil { + t.Fatalf("verify: %v", err) + } + if !done { + t.Fatal("verify held on a control plane that serves no version endpoint") + } + + var after simplyblockv1alpha2.ControlPlaneOps + if err := c.Get(ctx, client.ObjectKeyFromObject(ops), &after); err != nil { + t.Fatalf("read the operation back: %v", err) + } + if !strings.Contains(after.Status.Message, "not verified") { + t.Errorf("status.message = %q, want it to record that the version was not verified", + after.Status.Message) + } +} + +// A digest-pinned image carries no version to compare against, so the comparison +// is skipped rather than failed: what is pinned by digest is not claimed to be +// any particular version. +func TestADigestPinnedImageIsNotComparedAgainstAVersion(t *testing.T) { + for _, tc := range []struct { + image string + version string + want bool + }{ + {"quay.io/simplyblock-io/simplyblock:26.3.0", "26.3.0", true}, + {"quay.io/simplyblock-io/simplyblock:26.3.0", "26.2.8", false}, + {"quay.io/simplyblock-io/simplyblock:26.3.0@sha256:" + strings.Repeat("a", 64), "26.2.8", true}, + } { + if got := imageStates(tc.image, tc.version); got != tc.want { + t.Errorf("imageStates(%q, %q) = %v, want %v", tc.image, tc.version, got, tc.want) + } + } +} + +// A Restart naming something this control plane cannot roll is refused rather +// than skipped. An operation that reported success while recycling nothing is +// worse than one that says the name was wrong. +// +// The FoundationDB cluster is the case worth stating: it is a component of the +// control plane, so a check against the component table admits it, and it is not +// rolled by a pod-template annotation, so the recycle would do nothing and the +// wait that follows would pass against a healthy database. +func TestARestartNamingSomethingItCannotRollFails(t *testing.T) { + for _, tc := range []struct { + name string + scope string + }{ + {"a workload of another deployment", "simplyblock-graylog"}, + {"the database, which no annotation rolls", ComponentFDBCluster}, + } { + t.Run(tc.name, func(t *testing.T) { + cp := localControlPlane() + ops := opsFor(simplyblockv1alpha2.ControlPlaneOpsActionRestart) + ops.Spec.Restart = &simplyblockv1alpha2.RestartSpec{Components: []string{tc.scope}} + + r := &ControlPlaneOpsReconciler{Client: newClient(t, cp, ops), Scheme: testScheme(t)} + + _, _, err := r.restart(context.Background(), ops, cp) + var fatal *terminalStepError + if !errors.As(err, &fatal) { + t.Fatalf("restart returned %v, want a terminal failure", err) + } + if !strings.Contains(fatal.Error(), tc.scope) { + t.Errorf("the refusal is %q, want it to name the component", fatal.Error()) + } + }) + } +} + +// A scoped restart stamps only what it named. Recycling the whole control plane +// to restart one wedged component would interrupt everything else for nothing. +func TestAScopedRestartRecyclesOnlyWhatItNamed(t *testing.T) { + ctx := context.Background() + cp := localControlPlane() + ops := opsFor(simplyblockv1alpha2.ControlPlaneOpsActionRestart) + ops.Spec.Restart = &simplyblockv1alpha2.RestartSpec{Components: []string{ComponentTasks}} + + c := newClient(t, cp, ops, + deployment(ComponentTasks, 1, 1), + deployment(ComponentWebAPI, 2, 2), + ) + r := &ControlPlaneOpsReconciler{Client: c, Scheme: testScheme(t)} + + if _, _, err := r.restart(ctx, ops, cp); err != nil { + t.Fatalf("restart: %v", err) + } + + if !restarted(t, c, ComponentTasks) { + t.Errorf("%s was not recycled although the operation named it", ComponentTasks) + } + if restarted(t, c, ComponentWebAPI) { + t.Errorf("%s was recycled although the operation did not name it", ComponentWebAPI) + } +} + +// restarted reports whether a Deployment's pod template carries the restart +// stamp, which is what a rolling restart is: the Deployment controller sees a +// changed template and rolls it. +func restarted(t *testing.T, c client.Client, name string) bool { + t.Helper() + var d appsv1.Deployment + key := client.ObjectKey{Namespace: testNamespace, Name: name} + if err := c.Get(context.Background(), key, &d); err != nil { + t.Fatalf("read %s: %v", name, err) + } + _, stamped := d.Spec.Template.Annotations[restartedAtAnnotation] + return stamped +} diff --git a/operator/internal/controllers/controlplane/datastore.go b/operator/internal/controllers/controlplane/datastore.go new file mode 100644 index 000000000..52b93d741 --- /dev/null +++ b/operator/internal/controllers/controlplane/datastore.go @@ -0,0 +1,209 @@ +// The store the ApplyingDatastore step installs, and the bucket configuration +// beside it. +// +// design-controlplane.md §4.2 names this step "the document store the management +// API needs" and §5.1 identifies that as a MongoDBCommunity, which §12 Q6 then +// asks what reconciles. A base deployment settles the question by not having +// one: the chart renders the MongoDBCommunity only when observability is +// enabled, because what keeps documents in it is Graylog rather than the +// management API, and the reference deployment runs no MongoDB at all with the +// management API Available. +// +// What a base deployment does run is an object store, so that is what the step +// applies. The store holds what outlives a process: the metric history the +// control plane keeps and the backups it writes. The bucket is made by a sidecar +// beside the server, which is where it can wait for the store to answer. + +package controlplane + +import ( + appsv1 "k8s.io/api/apps/v1" + corev1 "k8s.io/api/core/v1" + "k8s.io/apimachinery/pkg/api/resource" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + + "github.com/simplyblock/atlas/ptr" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// The object store's images and its fixed configuration. +// +// The credentials are the chart's, and they are constants: the store listens +// only on a ClusterIP Service inside the namespace, and every consumer of it is +// a workload this same install creates. They stay constants while the chart +// still renders the same value into the observability half's objstore +// configuration, because the two have to agree. +const ( + minioImage = "quay.io/simplyblock-io/minio:RELEASE.2024-01-16T16-07-38Z" + minioClientImage = "quay.io/simplyblock-io/minio-client:RELEASE.2024-01-16T16-06-34Z" + + minioAccessKey = "minioadmin" + minioSecretKey = "minioadmin" + minioBucket = "thanos" + + // minioVolumeSize is what the store claims. It holds metric history and + // backup objects, so it grows with how long a deployment keeps them rather + // than with how much data the fleet stores. + minioVolumeSize = "50Gi" + + // minioDataVolume is the claim template's name, and therefore part of every + // PersistentVolumeClaim name the StatefulSet creates. It is immutable in + // practice: renaming it orphans the volume holding the store's data. + minioDataVolume = "minio-data" +) + +// datastoreObjects is everything the ApplyingDatastore step writes: the bucket +// configuration its consumers read, the store itself, and the Service they reach +// it on. +func datastoreObjects(cp *simplyblockv1alpha2.ControlPlane) []client.Object { + return []client.Object{ + objectStoreConfig(cp.Namespace), + minioStatefulSet(cp), + minioService(cp.Namespace), + } +} + +// objectStoreConfig is the bucket, the endpoint, and the credentials in the form +// the store's consumers parse. It is applied here rather than with the +// management API because it describes the store rather than the reader. +func objectStoreConfig(namespace string) *corev1.ConfigMap { + return &corev1.ConfigMap{ + ObjectMeta: metav1.ObjectMeta{ + Name: objectStoreConfigName, + Namespace: namespace, + Labels: map[string]string{appLabel: ComponentMinio}, + }, + Data: map[string]string{ + "objstore.yml": "type: S3\n" + + "config:\n" + + " bucket: " + minioBucket + "\n" + + " endpoint: " + ComponentMinio + ":9000\n" + + " access_key: " + minioAccessKey + "\n" + + " secret_key: " + minioSecretKey + "\n" + + " insecure: true\n", + }, + } +} + +// minioStatefulSet is the store: one server with a volume, and a sidecar that +// makes the bucket once the server answers. +// +// It is a StatefulSet rather than a Deployment because the data is on a volume +// the pod has to come back to. One replica is what a base deployment gets, and +// it is a single point of failure for the metric history rather than for the +// control plane: nothing in the install path or the data path reads it. +func minioStatefulSet(cp *simplyblockv1alpha2.ControlPlane) *appsv1.StatefulSet { + labels := map[string]string{appLabel: ComponentMinio} + credentials := []corev1.EnvVar{ + {Name: "MINIO_ROOT_USER", Value: minioAccessKey}, + {Name: "MINIO_ROOT_PASSWORD", Value: minioSecretKey}, + } + + spec := corev1.PodSpec{ + Containers: []corev1.Container{ + { + Name: "minio", + Image: minioImage, + ImagePullPolicy: pullPolicyOf(cp.Spec.Source.Local), + Args: []string{"server", "/data", "--console-address=:9001"}, + Env: credentials, + Ports: []corev1.ContainerPort{ + {Name: "api", ContainerPort: minioAPIPort}, + {Name: "console", ContainerPort: minioConsolePort}, + }, + VolumeMounts: []corev1.VolumeMount{{Name: minioDataVolume, MountPath: "/data"}}, + Resources: corev1.ResourceRequirements{ + Requests: corev1.ResourceList{ + corev1.ResourceCPU: resource.MustParse("100m"), + corev1.ResourceMemory: resource.MustParse("256Mi"), + }, + Limits: corev1.ResourceList{ + corev1.ResourceCPU: resource.MustParse("500m"), + corev1.ResourceMemory: resource.MustParse("1Gi"), + }, + }, + }, + { + // The sidecar waits for the server in the same pod, makes the + // bucket, and then sleeps. Sleeping rather than exiting is what + // keeps the pod out of a restart loop: a container that exits + // zero in a StatefulSet pod is restarted, and the bucket would + // be re-made on every restart forever. + Name: "bucket-init", + Image: minioClientImage, + ImagePullPolicy: pullPolicyOf(cp.Spec.Source.Local), + Command: []string{"sh", "-c", `until mc alias set local http://localhost:9000 "$MINIO_ROOT_USER" "$MINIO_ROOT_PASSWORD"; do + echo "Waiting for the object store..."; sleep 3; +done +mc mb --ignore-existing local/` + minioBucket + ` +sleep infinity +`}, + Env: credentials, + Resources: corev1.ResourceRequirements{ + Requests: corev1.ResourceList{ + corev1.ResourceCPU: resource.MustParse("10m"), + corev1.ResourceMemory: resource.MustParse("32Mi"), + }, + Limits: corev1.ResourceList{ + corev1.ResourceCPU: resource.MustParse("50m"), + corev1.ResourceMemory: resource.MustParse("64Mi"), + }, + }, + }, + }, + } + scheduling(cp.Spec.Source.Local, &spec) + + claim := corev1.PersistentVolumeClaim{ + ObjectMeta: metav1.ObjectMeta{Name: minioDataVolume}, + Spec: corev1.PersistentVolumeClaimSpec{ + AccessModes: []corev1.PersistentVolumeAccessMode{corev1.ReadWriteOnce}, + Resources: corev1.VolumeResourceRequirements{ + Requests: corev1.ResourceList{ + corev1.ResourceStorage: resource.MustParse(minioVolumeSize), + }, + }, + }, + } + // The store's volume comes from the same class as FoundationDB's, for the + // same reason: it cannot be a class this operator provides, because the + // control plane has to exist before any simplyblock volume can. + if fdb := foundationDBSpecOf(cp); fdb != nil && fdb.StorageClassName != "" { + claim.Spec.StorageClassName = ptr.To(fdb.StorageClassName) + } + + return &appsv1.StatefulSet{ + ObjectMeta: metav1.ObjectMeta{ + Name: ComponentMinio, + Namespace: cp.Namespace, + Labels: labels, + }, + Spec: appsv1.StatefulSetSpec{ + ServiceName: ComponentMinio, + Replicas: ptr.To(int32(1)), + Selector: &metav1.LabelSelector{MatchLabels: labels}, + Template: corev1.PodTemplateSpec{ + ObjectMeta: metav1.ObjectMeta{Labels: labels}, + Spec: spec, + }, + VolumeClaimTemplates: []corev1.PersistentVolumeClaim{claim}, + }, + } +} + +// minioService is the address the store's consumers use. The console port is +// published beside the API port because the two are the same process, and +// reaching the console is how somebody looks at what is in the bucket. +func minioService(namespace string) *corev1.Service { + return &corev1.Service{ + ObjectMeta: metav1.ObjectMeta{Name: ComponentMinio, Namespace: namespace}, + Spec: corev1.ServiceSpec{ + Selector: map[string]string{appLabel: ComponentMinio}, + Ports: []corev1.ServicePort{ + {Name: "api", Port: minioAPIPort, TargetPort: intstrFromInt(minioAPIPort)}, + {Name: "console", Port: minioConsolePort, TargetPort: intstrFromInt(minioConsolePort)}, + }, + }, + } +} diff --git a/operator/internal/controllers/controlplane/doc.go b/operator/internal/controllers/controlplane/doc.go new file mode 100644 index 000000000..e1e1fe07d --- /dev/null +++ b/operator/internal/controllers/controlplane/doc.go @@ -0,0 +1,62 @@ +// Package controlplane reconciles the ControlPlane singleton and the operations +// performed against it. +// +// The kind says one of two things. spec.source.local is a control plane this +// cluster hosts, which the operator installs and owns. spec.source.managed is a +// control plane elsewhere that this cluster's storage is managed by, which the +// operator only resolves and probes. +// +// The kind is the root of the ownership spine: a StorageCluster cannot be +// created, a StorageNode cannot be added, and a volume cannot be provisioned +// until it reports Available. That makes two things this package does more +// load-bearing than they look. It installs, which is the work the Helm chart +// used to do and which every deployment now depends on. And it publishes +// status.endpoint, which is where the rest of the operator learns how to reach +// the control plane. +// +// # What the install applies +// +// The managed install applies the objects a base control plane consists of, +// which is what the chart rendered with observability disabled: +// +// - The FoundationDB operator, its RBAC, and the FoundationDBCluster itself. +// The operator half is skipped where the Kubernetes cluster already serves +// apps.foundationdb.org, since a second one reconciling the same objects is +// worse than depending on the first. +// - The object store, which is MinIO and the bucket configuration beside it. +// - The management API, the services beside it, their shared account, the +// configuration they read, and the Service the endpoint resolves to. +// +// What it does not apply is the observability half: Graylog, Grafana, Thanos, +// the document store behind them, and the log collector. Those are gated behind +// one chart value, none of them appears in a step of the installation machine, +// and every one is non-essential in the phase table, so they stay where they +// are. [componentTable] lists what is installed and therefore watched. +// +// # Why the datastore step applies an object store +// +// design-controlplane.md §4.2 names ApplyingDatastore "the document store the +// management API needs," and §5.1 identifies that as a MongoDBCommunity. A base +// deployment has none: the chart renders the MongoDBCommunity only with +// observability enabled, because what stores documents in it is Graylog rather +// than the management API. The step therefore applies the store a base +// deployment does have, which is MinIO. §12 Q6 asks what supplies the MongoDB +// operator where a cluster has none, and the answer this package reaches is that +// a base control plane does not need one. +// +// # What is not here yet +// +// TLS. The chart serves the control plane over TLS behind tls.enabled, which has +// no field in the ControlPlane spec design-controlplane.md settles, and §5.1 +// says only that the issuer is detected rather than declared. The install this +// package performs is the plaintext one, which is the chart's default and what +// the reference deployment runs. The chart refuses to hand over a TLS-enabled +// deployment rather than quietly installing it without TLS. +// +// Adoption. An install that meets objects a Helm release already created takes +// them over by server-side apply under a stable field manager, which is the same +// mechanism the CSI driver's adoption uses, but nothing here verifies the +// handover or strips the release's claim afterward. design-controlplane.md §12 +// Q2 is the open question, and taking over a live control plane is the half of +// it this package does not yet answer. +package controlplane diff --git a/operator/internal/controllers/controlplane/endpoint.go b/operator/internal/controllers/controlplane/endpoint.go new file mode 100644 index 000000000..a1a306cdf --- /dev/null +++ b/operator/internal/controllers/controlplane/endpoint.go @@ -0,0 +1,218 @@ +// Where the control plane is: derived for one this cluster hosts, echoed and +// validated for one it is managed by. +// +// This is the field that makes the object useful to anything but a human +// (design-controlplane.md §3.3). Every controller in the operator reaches the +// control plane, and each resolves this endpoint per call through +// [NewEndpointResolver], falling back to the environment only where the object +// has published nothing. One object answers where the control plane is, and a +// change to it reaches every reader without a Deployment rollout. +// +// Resolution is deliberately dumb for the managed case: the Service this install +// creates, in the namespace the ControlPlane is in, on the port it publishes. +// The fully qualified name rather than the short one, because a reader in +// another namespace has to be able to use the same string. + +package controlplane + +import ( + "context" + "crypto/tls" + "crypto/x509" + "fmt" + "net/http" + "net/url" + "strings" + + corev1 "k8s.io/api/core/v1" + "k8s.io/apimachinery/pkg/api/errors" + "sigs.k8s.io/controller-runtime/pkg/client" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// The keys a credentials Secret may carry. Two are accepted because the two +// conventions are both in use in this repository, and refusing one of them would +// be a rename disguised as a validation. +var credentialKeys = []string{"token", "secret"} + +// The keys a CA bundle Secret may carry. ca.crt is what cert-manager writes and +// what a kubernetes.io/tls Secret carries; tls.crt covers a bundle somebody +// assembled by hand. +var caBundleKeys = []string{"ca.crt", "tls.crt"} + +// localEndpoint is where the management API this install created answers. +func localEndpoint(namespace string) string { + return fmt.Sprintf("http://%s.%s.svc.cluster.local:%d", ComponentWebAPI, namespace, webAPIPort) +} + +// credentialsError is what a Secret that is missing or unusable produces. It is +// its own type so the reconciler can emit the CredentialsError event for it and +// EndpointUnreachable for everything else, which are different problems with +// different fixes. +type credentialsError struct{ message string } + +func (e *credentialsError) Error() string { return e.message } + +// managedAccess is everything needed to reach a remote control plane: where +// it is, what to authenticate with, and what to verify its certificate against. +// +// The three travel together because all three come from the same spec block and +// all three are needed by the same call. Returning the transport rather than the +// CA bytes keeps the trust decision in one place instead of at each probe. +type managedAccess struct { + endpoint string + token string + + // client is nil when the system trust store is what verifies the endpoint, + // which is what an absent caBundleSecretRef means. + client *http.Client +} + +// resolveManaged reads where a remote control plane is, what to +// authenticate with, and what to verify it against. +// +// The endpoint is validated here as well as by the spec's pattern, because the +// pattern admits a loopback address and this does not. An operator pointed at +// 127.0.0.1 would probe itself, which is the request-forgery shape every +// managed endpoint in this group is guarded against. +func resolveManaged( + ctx context.Context, c client.Reader, cp *simplyblockv1alpha2.ControlPlane, +) (managedAccess, error) { + managed := cp.Spec.Source.Managed + if managed == nil { + return managedAccess{}, fmt.Errorf("spec.source.managed is not set") + } + + if err := validateEndpoint(managed.Endpoint); err != nil { + return managedAccess{}, err + } + access := managedAccess{endpoint: managed.Endpoint} + + transport, err := caBundleTransport(ctx, c, cp) + if err != nil { + return managedAccess{}, err + } + access.client = transport + + // No reference means no token, which is the in-cluster case the chart writes + // when it installs the control plane itself. + if managed.CredentialsSecretRef == nil || managed.CredentialsSecretRef.Name == "" { + return access, nil + } + + var secret corev1.Secret + key := client.ObjectKey{Namespace: cp.Namespace, Name: managed.CredentialsSecretRef.Name} + if err := c.Get(ctx, key, &secret); err != nil { + if errors.IsNotFound(err) { + return managedAccess{}, &credentialsError{message: fmt.Sprintf( + "Secret %s/%s does not exist, and it is what holds the bearer token this "+ + "control plane is reached with", cp.Namespace, managed.CredentialsSecretRef.Name)} + } + return managedAccess{}, err + } + + for _, k := range credentialKeys { + if value := strings.TrimSpace(string(secret.Data[k])); value != "" { + access.token = value + return access, nil + } + } + return managedAccess{}, &credentialsError{message: fmt.Sprintf( + "Secret %s/%s carries no %s key, so there is no token to authenticate with", + cp.Namespace, managed.CredentialsSecretRef.Name, strings.Join(credentialKeys, " or "))} +} + +// caBundleTransport builds the HTTP client that verifies the endpoint against +// the CA the spec names, or nil where it names none. +// +// A Secret that is named and unusable is an error rather than a fall back to the +// system trust store. Naming a CA states that the endpoint is signed by it, and +// the connection is refused until that holds. +func caBundleTransport( + ctx context.Context, c client.Reader, cp *simplyblockv1alpha2.ControlPlane, +) (*http.Client, error) { + ref := cp.Spec.Source.Managed.CABundleSecretRef + if ref == nil || ref.Name == "" { + return nil, nil + } + + var secret corev1.Secret + key := client.ObjectKey{Namespace: cp.Namespace, Name: ref.Name} + if err := c.Get(ctx, key, &secret); err != nil { + if errors.IsNotFound(err) { + return nil, &credentialsError{message: fmt.Sprintf( + "Secret %s/%s does not exist, and it is what the endpoint's certificate is "+ + "verified against", cp.Namespace, ref.Name)} + } + return nil, err + } + + var bundle []byte + for _, k := range caBundleKeys { + if value := secret.Data[k]; len(value) > 0 { + bundle = value + break + } + } + if len(bundle) == 0 { + return nil, &credentialsError{message: fmt.Sprintf( + "Secret %s/%s carries no %s key, so there is no CA to verify the endpoint against", + cp.Namespace, ref.Name, strings.Join(caBundleKeys, " or "))} + } + + pool := x509.NewCertPool() + if !pool.AppendCertsFromPEM(bundle) { + return nil, &credentialsError{message: fmt.Sprintf( + "Secret %s/%s holds no PEM certificate this operator could parse", + cp.Namespace, ref.Name)} + } + + return &http.Client{ + Timeout: probeTimeout, + Transport: &http.Transport{ + TLSClientConfig: &tls.Config{RootCAs: pool, MinVersion: tls.VersionTLS12}, + }, + }, nil +} + +// validateEndpoint refuses the addresses a remote control plane must not be +// at. It is the outbound-URL guard every other outbound endpoint in this group +// carries, applied at the one place the endpoint is read. +func validateEndpoint(raw string) error { + parsed, err := url.Parse(raw) + if err != nil { + return fmt.Errorf("spec.source.managed.endpoint is not a URL: %w", err) + } + if parsed.Scheme != "http" && parsed.Scheme != "https" { + return fmt.Errorf("spec.source.managed.endpoint has scheme %q, want http or https", parsed.Scheme) + } + host := parsed.Hostname() + if host == "" { + return fmt.Errorf("spec.source.managed.endpoint names no host") + } + if isLoopbackOrLinkLocal(host) { + return fmt.Errorf( + "spec.source.managed.endpoint names %q, which resolves inside the operator's own "+ + "pod rather than to a control plane", host) + } + return nil +} + +// isLoopbackOrLinkLocal reports the hosts a managed endpoint may not name. +// They are matched by spelling rather than by resolution, because resolving a +// name at admission time answers for the moment of admission and not for the +// life of the object. +func isLoopbackOrLinkLocal(host string) bool { + lowered := strings.ToLower(host) + switch { + case lowered == "localhost", strings.HasSuffix(lowered, ".localhost"): + return true + case strings.HasPrefix(lowered, "127."), lowered == "::1", lowered == "[::1]": + return true + case strings.HasPrefix(lowered, "169.254."): + return true + default: + return false + } +} diff --git a/operator/internal/controllers/controlplane/endpoint_test.go b/operator/internal/controllers/controlplane/endpoint_test.go new file mode 100644 index 000000000..10a93304b --- /dev/null +++ b/operator/internal/controllers/controlplane/endpoint_test.go @@ -0,0 +1,120 @@ +// What a remote control plane is reached with: the endpoint, the token, and +// the CA its certificate is verified against. +// +// The CA is the one of the three that was declared in the API and wired to +// nothing, so a probe fell back to the system trust store and an endpoint signed +// by a private CA could not be reached at all. These pin that the field reaches +// the transport. + +package controlplane + +import ( + "context" + "strings" + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// A self-signed certificate, used only to check that a PEM bundle reaches the +// transport's root pool. Nothing connects to it. +const testCAPEM = `-----BEGIN CERTIFICATE----- +MIIBhTCCASugAwIBAgIQIRi6zePL6mKjOipn+dNuaTAKBggqhkjOPQQDAjASMRAw +DgYDVQQKEwdBY21lIENvMB4XDTE3MTAyMDE5NDMwNloXDTE4MTAyMDE5NDMwNlow +EjEQMA4GA1UEChMHQWNtZSBDbzBZMBMGByqGSM49AgEGCCqGSM49AwEHA0IABD0d +7VNhbWvZLWPuj/RtHFjvtJBEwOkhbN/BnnE8rnZR8+sbwnc/KhCk3FhnpHZnQz7B +5aETbbIgmuvewdjvSBSjYzBhMA4GA1UdDwEB/wQEAwICpDATBgNVHSUEDDAKBggr +BgEFBQcDATAPBgNVHRMBAf8EBTADAQH/MCkGA1UdEQQiMCCCDmxvY2FsaG9zdDo1 +NDUzgg4xMjcuMC4wLjE6NTQ1MzAKBggqhkjOPQQDAgNIADBFAiEA2zpJEPQyz6/l +Wf86aX6PepsntZv2GYlA5UpabfT2EZICICpJ5h/iI+i341gBmLiAFQOyTDT+/wQc +6MF9+Yw1Yy0t +-----END CERTIFICATE----- +` + +func managedWithCABundle(t *testing.T, secretName string, data map[string][]byte) ( + *simplyblockv1alpha2.ControlPlane, []client.Object, +) { + t.Helper() + cp := managedControlPlane("https://sb-control.example.com:5000") + cp.Spec.Source.Managed.CredentialsSecretRef = nil + cp.Spec.Source.Managed.CABundleSecretRef = &corev1.LocalObjectReference{Name: secretName} + + var extra []client.Object + if data != nil { + extra = append(extra, &corev1.Secret{ + ObjectMeta: metav1.ObjectMeta{Name: secretName, Namespace: testNamespace}, + Data: data, + }) + } + return cp, extra +} + +// A named CA bundle produces a client that verifies against it, rather than the +// nil client that means the system trust store. +func TestACABundleReachesTheTransport(t *testing.T) { + cp, extra := managedWithCABundle(t, "cp-ca", map[string][]byte{"ca.crt": []byte(testCAPEM)}) + objects := append([]client.Object{cp}, extra...) + + access, err := resolveManaged(context.Background(), newClient(t, objects...), cp) + if err != nil { + t.Fatalf("resolveManaged: %v", err) + } + if access.client == nil { + t.Fatal("no HTTP client was built, so the probe would use the system trust store " + + "and an endpoint signed by this CA could not be reached") + } + if access.client.Transport == nil { + t.Fatal("the client carries no transport, so the root pool went nowhere") + } +} + +// No reference means the system trust store, which is a nil client rather than +// an empty pool: an empty pool verifies nothing at all. +func TestNoCABundleLeavesTheSystemTrustStore(t *testing.T) { + cp := managedControlPlane("https://sb-control.example.com:5000") + cp.Spec.Source.Managed.CredentialsSecretRef = nil + + access, err := resolveManaged(context.Background(), newClient(t, cp), cp) + if err != nil { + t.Fatalf("resolveManaged: %v", err) + } + if access.client != nil { + t.Error("a client was built although no CA bundle was named") + } +} + +// A CA bundle that is named and unusable is an error rather than a fall back. +// Naming one says the endpoint is signed by it, so verifying against something +// else would connect to a control plane nobody vouched for. +func TestAnUnusableCABundleIsRefused(t *testing.T) { + for _, tc := range []struct { + name string + data map[string][]byte + want string + }{ + {"the Secret does not exist", nil, "does not exist"}, + {"it carries no CA key", map[string][]byte{"other": []byte(testCAPEM)}, "no ca.crt or tls.crt key"}, + {"it carries no PEM", map[string][]byte{"ca.crt": []byte("not a certificate")}, "no PEM certificate"}, + } { + t.Run(tc.name, func(t *testing.T) { + cp, extra := managedWithCABundle(t, "cp-ca", tc.data) + objects := append([]client.Object{cp}, extra...) + + _, err := resolveManaged(context.Background(), newClient(t, objects...), cp) + if err == nil { + t.Fatal("an unusable CA bundle was accepted") + } + var credentials *credentialsError + if !errorsAs(err, &credentials) { + t.Fatalf("err = %v, want a credentials error", err) + } + if !strings.Contains(err.Error(), tc.want) { + t.Errorf("err = %q, want it to mention %q", err, tc.want) + } + }) + } +} diff --git a/operator/internal/controllers/controlplane/events.go b/operator/internal/controllers/controlplane/events.go new file mode 100644 index 000000000..3a05e77c2 --- /dev/null +++ b/operator/internal/controllers/controlplane/events.go @@ -0,0 +1,96 @@ +// The event reasons this package emits. +// +// Events land on the object an administrator has open. For the entity that is +// the ControlPlane, which is a singleton, and for an operation it is the +// ControlPlaneOps, which outlives the operation as its audit record. +// +// design-controlplane.md §9.1 is the specification. + +package controlplane + +// Reasons emitted on the ControlPlane. +const ( + // ControlPlaneNotReady is the load-bearing one, because every controller in + // the operator holds when it fires and none of them says why. A cluster that + // will not create, a node that will not add, and a volume that will not + // provision are one event on one object. + // + // It is emitted on transition rather than on every probe, which bounds one + // outage to one event rather than one every thirty seconds. + ControlPlaneNotReady = "ControlPlaneNotReady" + + // ControlPlaneDegraded fires while every request is still being served, so + // it reaches an administrator before an outage rather than during one. It + // names the component and its two counts, because one reason covering eight + // workloads sends a reader to status.components anyway. + ControlPlaneDegraded = "ControlPlaneDegraded" + + // ControlPlaneReady is the recovery, and it is what tells a reader that an + // earlier ControlPlaneNotReady is over. + ControlPlaneReady = "ControlPlaneReady" + + // AwaitingDependency is an installation step waiting on something outside + // the operator: FoundationDB reaching quorum, or the management API starting. + AwaitingDependency = "AwaitingDependency" + + // StepDeadlineExceeded is an installation step that outlived its budget, + // which is how a wait by design is separated from a wait caused by a bug. + StepDeadlineExceeded = "StepDeadlineExceeded" + + // ClustersStillPresent holds a deletion. It is a hold rather than a failure: + // removing the clusters resolves it, and nothing else can. + ClustersStillPresent = "ClustersStillPresent" + + // EndpointUnreachable is a remote control plane that could not be + // resolved or reached. + EndpointUnreachable = "EndpointUnreachable" + + // CredentialsError is a remote control plane whose Secret is missing or + // carries no token. It is separate from EndpointUnreachable because the two + // have different fixes. + CredentialsError = "CredentialsError" + + // PrerequisiteMissing is an install held on something the cluster has to + // provide and does not, such as the FoundationDB CRDs. + PrerequisiteMissing = "PrerequisiteMissing" + + // DuplicateControlPlane is a second ControlPlane in a Kubernetes cluster that + // already has one. It installs nothing, because the objects it would apply + // are the same fixed-name cluster-scoped objects the first one owns. + DuplicateControlPlane = "DuplicateControlPlane" +) + +// Reasons emitted on a ControlPlaneOps. +const ( + // BackupRequested is a Backup run that created a FoundationDBBackup. + BackupRequested = "BackupRequested" + + // BackupTriggered is a Backup run that found one already configured and + // asked it for a snapshot instead of applying a second beside it. + BackupTriggered = "BackupTriggered" + + // OperationQueued is an operation waiting for another to release the lock. + OperationQueued = "OperationQueued" + + // OperationsInFlight is an operation holding for other operations in the + // namespace to finish. It does not cancel them, because an operation + // canceled to make a restart convenient is a worse outcome than a restart + // that waited. + OperationsInFlight = "OperationsInFlight" + + // OperationStarted is an operation that acquired the lock. + OperationStarted = "OperationStarted" + + // OperationSucceeded and OperationFailed are the two terminal outcomes an + // operation that ran reaches. + OperationSucceeded = "OperationSucceeded" + OperationFailed = "OperationFailed" + + // OperationAborted is an operation stopped on request whose unwind finished. + OperationAborted = "OperationAborted" + + // VersionMismatch is an upgrade whose Verifying step found the control plane + // reporting a version other than the one asked for, which is the difference + // between an upgrade that completed and a rollout that failed back. + VersionMismatch = "VersionMismatch" +) diff --git a/operator/internal/controllers/controlplane/foundationdb.go b/operator/internal/controllers/controlplane/foundationdb.go new file mode 100644 index 000000000..6df98244f --- /dev/null +++ b/operator/internal/controllers/controlplane/foundationdb.go @@ -0,0 +1,592 @@ +// The FoundationDB half of the install: the operator that reconciles +// FoundationDBClusters, the accounts and roles it and the database pods need, +// and the FoundationDBCluster itself. +// +// Nothing here is a Go type, because this repository has no dependency on the +// FoundationDB operator's API module and taking one would pull its whole +// controller in for two status fields. The cluster is built and read as +// unstructured, which is also what lets the apply work against whatever +// v1beta2's schema is on the cluster rather than against a pinned copy of it. +// +// # Detection, and where it differs from the design +// +// design-controlplane.md §5.1 detects the FoundationDB operator by whether +// apps.foundationdb.org/v1beta2 is served, and installs the CRDs and the +// controller where it is not. That test does not separate the two things on this +// chart: the CRDs ship in the chart's CRD directory, which Helm applies on +// install and never removes, so the group is served on every deployment whether +// or not a controller is reconciling it. +// +// So the two halves are split by who can answer for them. The CRDs stay with the +// chart, and the group being served is treated as a prerequisite: an install +// that cannot see v1beta2 holds and names it, because creating a +// FoundationDBCluster against a group the API server does not know is not a wait +// but an error. The controller is applied, under the name the chart gave it, in +// the namespace the ControlPlane is in. That is what ships today, so a cluster +// that also runs a FoundationDB operator of its own is in the same position it +// was in before the install moved. + +package controlplane + +import ( + "context" + "fmt" + + appsv1 "k8s.io/api/apps/v1" + corev1 "k8s.io/api/core/v1" + rbacv1 "k8s.io/api/rbac/v1" + "k8s.io/apimachinery/pkg/api/errors" + "k8s.io/apimachinery/pkg/api/meta" + "k8s.io/apimachinery/pkg/api/resource" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" + "k8s.io/apimachinery/pkg/runtime/schema" + "sigs.k8s.io/controller-runtime/pkg/client" + + "github.com/simplyblock/atlas/ptr" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// The FoundationDB group, and the version of it this install writes. +const ( + fdbGroup = "apps.foundationdb.org" + fdbVersion = "v1beta2" +) + +// fdbClusterGVK is the kind the install applies and reads back. +var fdbClusterGVK = schema.GroupVersionKind{ + Group: fdbGroup, + Version: fdbVersion, + Kind: "FoundationDBCluster", +} + +// fdbBackupGVK is what a Backup operation creates or triggers. It is here rather +// than beside that operation because it is the same group, and the two would +// otherwise disagree about the version. +var fdbBackupGVK = schema.GroupVersionKind{ + Group: fdbGroup, + Version: fdbVersion, + Kind: "FoundationDBBackup", +} + +// The images the FoundationDB half runs. They are pinned constants rather than +// spec fields for the reason design-controlplane.md gives for not putting the +// component table in the API: a user cannot add a component to a managed control +// plane, and the version of the database is the control plane's business rather +// than the deployment's. Moving FoundationDB forward is a release of this +// operator. +const ( + fdbOperatorImage = "quay.io/simplyblock-io/fdb-kubernetes-operator:v2.13.0" + fdbMonitorImage = "quay.io/simplyblock-io/fdb-kubernetes-monitor:7.3.63" + fdbVersionString = "7.3.63" +) + +// fdbCoordinatorDisk is what each FoundationDB process claims. It is the chart's +// value, and it sizes the metadata of a fleet rather than its data: what lives in +// FoundationDB is cluster definitions, node registrations, and lvol records. +const fdbCoordinatorDisk = "10G" + +// foundationDBObjects is everything the ApplyingFoundationDB step writes, in the +// order it depends on itself: the accounts, then the roles that name them, then +// the controller, then the cluster the controller reconciles. +func foundationDBObjects(cp *simplyblockv1alpha2.ControlPlane) []client.Object { + ns := cp.Namespace + return []client.Object{ + serviceAccount(ns, fdbOperatorServiceAccount), + serviceAccount(ns, fdbPodServiceAccount), + fdbManagerRoleObject(), + fdbManagerClusterRoleObject(), + fdbManagerRoleBindingObject(ns), + fdbManagerClusterRoleBindingObject(ns), + fdbPodRoleObject(ns), + fdbPodRoleBindingObject(ns), + fdbOperatorDeployment(cp), + foundationDBCluster(cp), + } +} + +// foundationDBClusterScoped is the half of that set the garbage collector will +// not remove with the ControlPlane, and which the finalizer therefore deletes. +func foundationDBClusterScoped() []client.Object { + return []client.Object{ + fdbManagerRoleObject(), + fdbManagerClusterRoleObject(), + fdbManagerClusterRoleBindingObject(""), + } +} + +// serviceAccount is the account an object set runs under. It carries nothing but +// its name: what it may do is in the bindings that name it. +func serviceAccount(namespace, name string) *corev1.ServiceAccount { + return &corev1.ServiceAccount{ + ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: namespace}, + } +} + +// fdbManagerRoleObject is what the FoundationDB operator does inside the namespace: +// the pods, volumes, and services that make up a database, and the three +// FoundationDB kinds themselves. It is a ClusterRole bound by a RoleBinding, so +// the grant is cluster-shaped and its effect is one namespace. +func fdbManagerRoleObject() *rbacv1.ClusterRole { + return &rbacv1.ClusterRole{ + ObjectMeta: metav1.ObjectMeta{Name: fdbOperatorRoleName}, + Rules: []rbacv1.PolicyRule{ + { + APIGroups: []string{""}, + Resources: []string{ + "configmaps", "events", "persistentvolumeclaims", "pods", "secrets", "services", + }, + Verbs: []string{"create", "delete", "get", "list", "patch", "update", "watch"}, + }, + { + APIGroups: []string{"apps"}, + Resources: []string{"deployments"}, + Verbs: []string{"create", "delete", "get", "list", "patch", "update", "watch"}, + }, + { + APIGroups: []string{fdbGroup}, + Resources: []string{ + "foundationdbbackups", "foundationdbclusters", "foundationdbrestores", + }, + Verbs: []string{"create", "delete", "get", "list", "patch", "update", "watch"}, + }, + { + APIGroups: []string{fdbGroup}, + Resources: []string{ + "foundationdbbackups/status", "foundationdbclusters/status", + "foundationdbrestores/status", + }, + Verbs: []string{"get", "patch", "update"}, + }, + { + APIGroups: []string{"coordination.k8s.io"}, + Resources: []string{"leases"}, + Verbs: []string{"create", "delete", "get", "list", "patch", "update", "watch"}, + }, + }, + } +} + +// fdbManagerClusterRoleObject is the one thing the operator needs outside the +// namespace: the nodes it reads to place a database across fault domains. +func fdbManagerClusterRoleObject() *rbacv1.ClusterRole { + return &rbacv1.ClusterRole{ + ObjectMeta: metav1.ObjectMeta{Name: fdbOperatorClusterRole}, + Rules: []rbacv1.PolicyRule{{ + APIGroups: []string{""}, + Resources: []string{"nodes"}, + Verbs: []string{"get", "list", "watch"}, + }}, + } +} + +func fdbManagerRoleBindingObject(namespace string) *rbacv1.RoleBinding { + return &rbacv1.RoleBinding{ + ObjectMeta: metav1.ObjectMeta{Name: fdbOperatorRoleBinding, Namespace: namespace}, + RoleRef: rbacv1.RoleRef{ + APIGroup: rbacv1.GroupName, Kind: "ClusterRole", Name: fdbOperatorRoleName, + }, + Subjects: []rbacv1.Subject{{ + Kind: "ServiceAccount", Name: fdbOperatorServiceAccount, Namespace: namespace, + }}, + } +} + +func fdbManagerClusterRoleBindingObject(namespace string) *rbacv1.ClusterRoleBinding { + return &rbacv1.ClusterRoleBinding{ + ObjectMeta: metav1.ObjectMeta{Name: fdbOperatorClusterBinding}, + RoleRef: rbacv1.RoleRef{ + APIGroup: rbacv1.GroupName, Kind: "ClusterRole", Name: fdbOperatorClusterRole, + }, + Subjects: []rbacv1.Subject{{ + Kind: "ServiceAccount", Name: fdbOperatorServiceAccount, Namespace: namespace, + }}, + } +} + +// fdbPodRoleObject is what a FoundationDB pod does to itself. The unified monitor in +// each pod writes locality and launcher-environment annotations onto its own pod +// to signal the operator, which the namespace's default account cannot do. +func fdbPodRoleObject(namespace string) *rbacv1.Role { + return &rbacv1.Role{ + ObjectMeta: metav1.ObjectMeta{Name: fdbPodRoleName, Namespace: namespace}, + Rules: []rbacv1.PolicyRule{{ + APIGroups: []string{""}, + Resources: []string{"pods"}, + Verbs: []string{"get", "list", "watch", "update", "patch"}, + }}, + } +} + +func fdbPodRoleBindingObject(namespace string) *rbacv1.RoleBinding { + return &rbacv1.RoleBinding{ + ObjectMeta: metav1.ObjectMeta{Name: fdbPodRoleBinding, Namespace: namespace}, + RoleRef: rbacv1.RoleRef{ + APIGroup: rbacv1.GroupName, Kind: "Role", Name: fdbPodRoleName, + }, + Subjects: []rbacv1.Subject{{ + Kind: "ServiceAccount", Name: fdbPodServiceAccount, Namespace: namespace, + }}, + } +} + +// fdbOperatorDeployment is the controller that turns a FoundationDBCluster into +// running pods. Its init container copies the FoundationDB client library and +// the three command-line tools out of the monitor image, because the controller +// shells out to fdbcli to read status and to change configuration. +func fdbOperatorDeployment(cp *simplyblockv1alpha2.ControlPlane) *appsv1.Deployment { + const ( + binariesVolume = "fdb-binaries" + tmpVolume = "tmp" + logsVolume = "logs" + ) + labels := map[string]string{appLabel: ComponentFDBOperator} + + spec := corev1.PodSpec{ + ServiceAccountName: fdbOperatorServiceAccount, + SecurityContext: &corev1.PodSecurityContext{ + RunAsUser: ptr.To(int64(4059)), + RunAsGroup: ptr.To(int64(4059)), + FSGroup: ptr.To(int64(4059)), + }, + Volumes: []corev1.Volume{ + {Name: tmpVolume, VolumeSource: corev1.VolumeSource{EmptyDir: &corev1.EmptyDirVolumeSource{}}}, + {Name: logsVolume, VolumeSource: corev1.VolumeSource{EmptyDir: &corev1.EmptyDirVolumeSource{}}}, + {Name: binariesVolume, VolumeSource: corev1.VolumeSource{EmptyDir: &corev1.EmptyDirVolumeSource{}}}, + }, + InitContainers: []corev1.Container{{ + Name: "foundationdb-kubernetes-init-7-3", + Image: fdbMonitorImage, + Args: []string{ + "--copy-library", "7.3", + "--copy-binary", "fdbcli", + "--copy-binary", "fdbbackup", + "--copy-binary", "fdbrestore", + "--output-dir", "/var/output-files", + "--mode", "init", + }, + VolumeMounts: []corev1.VolumeMount{ + {Name: binariesVolume, MountPath: "/var/output-files"}, + }, + }}, + Containers: []corev1.Container{{ + Name: "manager", + Image: fdbOperatorImage, + Command: []string{"/manager"}, + Args: []string{"--health-probe-bind-address=:9443"}, + Env: []corev1.EnvVar{{ + Name: "WATCH_NAMESPACE", + ValueFrom: &corev1.EnvVarSource{ + FieldRef: &corev1.ObjectFieldSelector{FieldPath: "metadata.namespace"}, + }, + }}, + Ports: []corev1.ContainerPort{{Name: "metrics", ContainerPort: 8080}}, + Resources: corev1.ResourceRequirements{ + Requests: corev1.ResourceList{ + corev1.ResourceCPU: resource.MustParse("500m"), + corev1.ResourceMemory: resource.MustParse("256Mi"), + }, + Limits: corev1.ResourceList{ + corev1.ResourceCPU: resource.MustParse("500m"), + corev1.ResourceMemory: resource.MustParse("256Mi"), + }, + }, + SecurityContext: &corev1.SecurityContext{ + ReadOnlyRootFilesystem: ptr.To(true), + AllowPrivilegeEscalation: ptr.To(false), + Privileged: ptr.To(false), + }, + VolumeMounts: []corev1.VolumeMount{ + {Name: tmpVolume, MountPath: "/tmp"}, + {Name: logsVolume, MountPath: "/var/log/fdb"}, + {Name: binariesVolume, MountPath: "/usr/bin/fdb"}, + }, + }}, + TerminationGracePeriodSeconds: ptr.To(int64(10)), + } + scheduling(cp.Spec.Source.Local, &spec) + + return &appsv1.Deployment{ + ObjectMeta: metav1.ObjectMeta{ + Name: ComponentFDBOperator, + Namespace: cp.Namespace, + Labels: labels, + }, + Spec: appsv1.DeploymentSpec{ + Replicas: ptr.To(int32(1)), + Selector: &metav1.LabelSelector{MatchLabels: labels}, + Template: corev1.PodTemplateSpec{ + ObjectMeta: metav1.ObjectMeta{Labels: labels}, + Spec: spec, + }, + }, + } +} + +// foundationDBCluster is the database itself. +// +// It is built as unstructured for the reason this file's opening comment gives, +// and the shape is the chart's: the unified image type, explicit listen +// addresses, DNS locality, and a per-class pod template that names the pods' +// account. What varies with the spec is the coordinator count, the storage +// class, and the resources. +func foundationDBCluster(cp *simplyblockv1alpha2.ControlPlane) *unstructured.Unstructured { + fdb := foundationDBSpecOf(cp) + processes := coordinatorCount(fdb) + + // The class templates are identical apart from their anti-affinity, and the + // FoundationDB operator replaces general.podTemplate wholesale when a class + // override exists, so each one repeats what general already said. + classTemplate := func(class string) map[string]any { + template := map[string]any{ + "spec": map[string]any{ + "serviceAccountName": fdbPodServiceAccount, + "containers": []any{ + fdbContainer(fdb), + }, + "affinity": map[string]any{ + "podAntiAffinity": map[string]any{ + "preferredDuringSchedulingIgnoredDuringExecution": []any{map[string]any{ + "weight": int64(100), + "podAffinityTerm": map[string]any{ + "labelSelector": map[string]any{ + "matchLabels": map[string]any{ + "foundationdb.org/fdb-process-class": class, + }, + }, + "topologyKey": "kubernetes.io/hostname", + }, + }}, + }, + }, + }, + } + if local := cp.Spec.Source.Local; local != nil && len(local.NodeSelector) > 0 { + template["spec"].(map[string]any)["nodeSelector"] = toAnyMap(local.NodeSelector) + } + return map[string]any{"podTemplate": template} + } + + general := map[string]any{ + "customParameters": []any{"knob_disable_posix_kernel_aio=1"}, + "podTemplate": map[string]any{ + "spec": map[string]any{ + "serviceAccountName": fdbPodServiceAccount, + "containers": []any{fdbContainer(fdb)}, + "initContainers": []any{map[string]any{ + "name": "foundationdb-kubernetes-init", + "resources": map[string]any{ + "requests": map[string]any{"cpu": "100m", "memory": "128Mi"}, + "limits": map[string]any{"cpu": "100m", "memory": "128Mi"}, + }, + "securityContext": map[string]any{"runAsUser": int64(0)}, + }}, + }, + }, + "volumeClaimTemplate": map[string]any{ + "spec": volumeClaimSpec(fdb), + }, + } + if local := cp.Spec.Source.Local; local != nil && len(local.NodeSelector) > 0 { + general["podTemplate"].(map[string]any)["spec"].(map[string]any)["nodeSelector"] = + toAnyMap(local.NodeSelector) + } + + obj := &unstructured.Unstructured{Object: map[string]any{ + "spec": map[string]any{ + "version": fdbVersionString, + "imageType": "unified", + "useExplicitListenAddress": true, + "minimumUptimeSecondsForBounce": int64(60), + "automationOptions": map[string]any{ + "replacements": map[string]any{"enabled": true}, + }, + "faultDomain": map[string]any{"key": "kubernetes.io/hostname"}, + "labels": map[string]any{ + "filterOnOwnerReference": false, + "matchLabels": map[string]any{ + "foundationdb.org/fdb-cluster-name": ComponentFDBCluster, + }, + "processClassLabels": []any{"foundationdb.org/fdb-process-class"}, + "processGroupIDLabels": []any{"foundationdb.org/fdb-process-group-id"}, + }, + "databaseConfiguration": map[string]any{ + "redundancy_mode": redundancyMode(processes), + }, + "processCounts": map[string]any{ + "cluster_controller": int64(1), + "log": processes, + "storage": processes, + "stateless": int64(-1), + }, + "processes": map[string]any{ + "general": general, + "storage": classTemplate("storage"), + "log": classTemplate("log"), + }, + "routing": map[string]any{"defineDNSLocalityFields": true}, + "mainContainer": map[string]any{"imageConfigs": []any{map[string]any{"baseImage": fdbMonitorImageRepository()}}}, + "sidecarContainer": map[string]any{"enableLivenessProbe": true, "enableReadinessProbe": false}, + }, + }} + obj.SetGroupVersionKind(fdbClusterGVK) + obj.SetName(ComponentFDBCluster) + obj.SetNamespace(cp.Namespace) + return obj +} + +// fdbContainer is the resource envelope and security context of the database +// process. It runs as root because the FoundationDB image's data directory is +// owned by it, which is the upstream image's arrangement rather than a choice +// this install makes. +func fdbContainer(fdb *simplyblockv1alpha2.FoundationDBSpec) map[string]any { + requests := map[string]any{"cpu": "100m", "memory": "1Gi"} + limits := map[string]any{"cpu": "500m", "memory": "4Gi"} + if fdb != nil { + if q, ok := fdb.Resources.Requests[corev1.ResourceCPU]; ok { + requests["cpu"] = q.String() + } + if q, ok := fdb.Resources.Requests[corev1.ResourceMemory]; ok { + requests["memory"] = q.String() + } + if q, ok := fdb.Resources.Limits[corev1.ResourceCPU]; ok { + limits["cpu"] = q.String() + } + if q, ok := fdb.Resources.Limits[corev1.ResourceMemory]; ok { + limits["memory"] = q.String() + } + } + return map[string]any{ + "name": "foundationdb", + "resources": map[string]any{"requests": requests, "limits": limits}, + "securityContext": map[string]any{"runAsUser": int64(0)}, + } +} + +// volumeClaimSpec is what each FoundationDB process claims. An unset storage +// class leaves the field out rather than writing an empty string, which is the +// difference between the cluster's default class and a class literally named +// nothing. +func volumeClaimSpec(fdb *simplyblockv1alpha2.FoundationDBSpec) map[string]any { + spec := map[string]any{ + "accessModes": []any{string(corev1.ReadWriteOnce)}, + "resources": map[string]any{ + "requests": map[string]any{"storage": fdbCoordinatorDisk}, + }, + } + if fdb != nil && fdb.StorageClassName != "" { + spec["storageClassName"] = fdb.StorageClassName + } + return spec +} + +// coordinatorCount is spec.source.local.foundationDB.replicas, or the API's +// default. Three is the smallest count that survives one loss. +func coordinatorCount(fdb *simplyblockv1alpha2.FoundationDBSpec) int64 { + if fdb == nil || fdb.Replicas == nil { + return 3 + } + return int64(*fdb.Replicas) +} + +// redundancyMode is how many copies FoundationDB keeps, derived from the +// coordinator count rather than configured beside it: the two have to agree, and +// a deployment that set one without the other would get a database that cannot +// reach the replication it was told to have. +func redundancyMode(processes int64) string { + switch { + case processes >= 5: + return "triple" + case processes >= 3: + return "double" + default: + return "single" + } +} + +// fdbMonitorImageRepository is the monitor image without its tag, which is the +// form mainContainer.imageConfigs takes: the FoundationDB operator appends the +// version it is running. +func fdbMonitorImageRepository() string { + for i := len(fdbMonitorImage) - 1; i >= 0; i-- { + if fdbMonitorImage[i] == ':' { + return fdbMonitorImage[:i] + } + } + return fdbMonitorImage +} + +// toAnyMap converts a string map to the shape an unstructured object holds. +func toAnyMap(in map[string]string) map[string]any { + out := make(map[string]any, len(in)) + for k, v := range in { + out[k] = v + } + return out +} + +// fdbHealth is what AwaitingFoundationDB and the component table both read. +type fdbHealth struct { + // found is whether the FoundationDBCluster exists at all. A cluster the + // apply just created and the cache has not caught up with is not found and + // not an error. + found bool + + // available is the cluster's own report that it is serving. It is the field + // that matters rather than a pod count, because a FoundationDB at two of + // three coordinators is serving and a count cannot say so. + available bool + + // fullReplication is whether every copy the redundancy mode asks for exists. + // A cluster available without it is serving with less redundancy than it was + // configured for, which is a warning rather than a wait. + fullReplication bool + + // desired and reconciled are the process groups the operator wants and the + // ones it has finished with, which is what the component table publishes. + desired int32 + reconciled int32 +} + +// readFoundationDB reads the cluster's status. A cluster that is not there yet +// is not an error: the apply created it and the cache has not caught up. +func readFoundationDB(ctx context.Context, c client.Reader, namespace string) (fdbHealth, error) { + obj := &unstructured.Unstructured{} + obj.SetGroupVersionKind(fdbClusterGVK) + key := client.ObjectKey{Namespace: namespace, Name: ComponentFDBCluster} + if err := c.Get(ctx, key, obj); err != nil { + if errors.IsNotFound(err) || meta.IsNoMatchError(err) { + return fdbHealth{}, nil + } + return fdbHealth{}, fmt.Errorf("read FoundationDBCluster %s: %w", ComponentFDBCluster, err) + } + + health := fdbHealth{found: true} + health.available, _, _ = unstructured.NestedBool(obj.Object, "status", "health", "available") + health.fullReplication, _, _ = unstructured.NestedBool(obj.Object, "status", "health", "fullReplication") + if v, ok, _ := unstructured.NestedInt64(obj.Object, "status", "desiredProcessGroups"); ok { + health.desired = int32(v) + } + if v, ok, _ := unstructured.NestedInt64(obj.Object, "status", "reconciledProcessGroups"); ok { + health.reconciled = int32(v) + } + return health, nil +} + +// waitingOn is what a held AwaitingFoundationDB step reports, read from the +// cluster's own status so a stalled install names the database rather than the +// operator. +func (h fdbHealth) waitingOn() string { + switch { + case !h.found: + return fmt.Sprintf("FoundationDBCluster %s has not been created yet", ComponentFDBCluster) + case !h.available: + return fmt.Sprintf("FoundationDBCluster %s has %d of %d process groups reconciled and is not available yet", + ComponentFDBCluster, h.reconciled, h.desired) + case !h.fullReplication: + return fmt.Sprintf("FoundationDBCluster %s is available and not yet fully replicated", + ComponentFDBCluster) + default: + return "" + } +} diff --git a/operator/internal/controllers/controlplane/graphs.go b/operator/internal/controllers/controlplane/graphs.go new file mode 100644 index 000000000..741879a00 --- /dev/null +++ b/operator/internal/controllers/controlplane/graphs.go @@ -0,0 +1,247 @@ +// The state graphs of the installation and of every ControlPlaneOps action, +// both declared as data. +// +// The entity's is a Config rather than a MultiConfig, because a ControlPlane has +// no spec.action to key one on: there is one installation path. The operations' +// is a MultiConfig, and one status.step field serves all three of them, so +// nothing in the API type prevents a Backup from reporting Verifying. The +// per-action graph makes that an IllegalTransitionError at the point of the +// write rather than an accepted status. +// +// design-controlplane.md §4.2 and §6 are the specification. + +package controlplane + +import ( + "context" + "time" + + "github.com/simplyblock/atlas/statemachine" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// installStep is the entity's step type, aliased so the graph literal reads as +// the graph rather than as a wall of package qualifiers. +type installStep = simplyblockv1alpha2.ControlPlaneStep + +const ( + stepApplyingFoundationDB = simplyblockv1alpha2.ControlPlaneStepApplyingFoundationDB + stepAwaitingFoundationDB = simplyblockv1alpha2.ControlPlaneStepAwaitingFoundationDB + stepApplyingDatastore = simplyblockv1alpha2.ControlPlaneStepApplyingDatastore + stepApplyingAPI = simplyblockv1alpha2.ControlPlaneStepApplyingAPI + stepAwaitingAPI = simplyblockv1alpha2.ControlPlaneStepAwaitingAPI +) + +// opsStep is the operations' step type, aliased for the same reason. +type opsStep = simplyblockv1alpha2.ControlPlaneOpsStep + +const ( + stepDraining = simplyblockv1alpha2.ControlPlaneOpsStepDraining + stepRestarting = simplyblockv1alpha2.ControlPlaneOpsStepRestarting + stepAwaiting = simplyblockv1alpha2.ControlPlaneOpsStepAwaiting + stepPreflight = simplyblockv1alpha2.ControlPlaneOpsStepPreflight + stepApplying = simplyblockv1alpha2.ControlPlaneOpsStepApplying + stepVerifying = simplyblockv1alpha2.ControlPlaneOpsStepVerifying + stepRequesting = simplyblockv1alpha2.ControlPlaneOpsStepRequesting +) + +// How long each installation step may take before it is reported as stuck. +// +// The applies are minutes because an apply is a handful of writes against the +// API server, and a step still applying after that is one whose writes are being +// rejected rather than one that is slow. The two waits are the numbers that +// matter. +const ( + applyingFoundationDBDeadline = 10 * time.Minute + + // awaitingFoundationDBDeadline is the step that can take the longest. Three + // coordinators on slow storage is minutes, an image pull on a cold node is + // more, and a FoundationDB that never reaches quorum is indistinguishable + // from one still starting without a bound. + awaitingFoundationDBDeadline = 45 * time.Minute + + applyingDatastoreDeadline = 10 * time.Minute + applyingAPIDeadline = 10 * time.Minute + + // awaitingAPIDeadline covers the management API starting and connecting to + // the database. It is shorter than FoundationDB's because everything it + // waits on is already running by the time this step is entered. + awaitingAPIDeadline = 20 * time.Minute +) + +// How long each operation step may take. +const ( + // drainingDeadline holds while other operations in the namespace finish. It + // is generous because what it waits on is a node add or a volume migration, + // both of which are legitimately long, and expiring is how an administrator + // learns the wait is no longer normal. + drainingDeadline = 2 * time.Hour + + restartingDeadline = 5 * time.Minute + + // awaitingOpsDeadline covers pods coming back, or a backup reporting a + // snapshot. + awaitingOpsDeadline = 30 * time.Minute + + preflightDeadline = 10 * time.Minute + applyingDeadline = 5 * time.Minute + + // verifyingDeadline covers the rollout finishing and the new version being + // reported. It is the upgrade's own rolling update plus the probe interval. + verifyingDeadline = 30 * time.Minute + + requestingDeadline = 5 * time.Minute +) + +// deadline is the entry hook every state here carries: it sets the step's budget +// and performs nothing. The side effect of a step is performed on the pass that +// follows, against the step the entry's patch persisted, which is where the +// write-ahead record is needed and what it records. +func deadline[S comparable](d time.Duration) statemachine.TransitionFunc[S] { + return func(context.Context, S, S) (time.Duration, error) { return d, nil } +} + +// installGraph is the entity's own machine (§4.2). It is a line: every step has +// exactly one successor, and AwaitingAPI is terminal because reaching it is what +// makes the control plane Available. +func installGraph() statemachine.Config[installStep] { + return statemachine.Config[installStep]{ + Initial: stepApplyingFoundationDB, + States: map[installStep]statemachine.StateDef[installStep]{ + stepApplyingFoundationDB: { + To: []installStep{stepAwaitingFoundationDB}, + OnEnter: deadline[installStep](applyingFoundationDBDeadline), + }, + stepAwaitingFoundationDB: { + To: []installStep{stepApplyingDatastore}, + OnEnter: deadline[installStep](awaitingFoundationDBDeadline), + }, + stepApplyingDatastore: { + To: []installStep{stepApplyingAPI}, + OnEnter: deadline[installStep](applyingDatastoreDeadline), + }, + stepApplyingAPI: { + To: []installStep{stepAwaitingAPI}, + OnEnter: deadline[installStep](applyingAPIDeadline), + }, + stepAwaitingAPI: {OnEnter: deadline[installStep](awaitingAPIDeadline)}, + }, + } +} + +// installStepBudgets is what each step's deadline was set from, derived from the +// graph's constants rather than restated so that a budget changed in one place +// moves both. It is what sets the first step's deadline, which a machine born +// already in that step would otherwise never get. +var installStepBudgets = map[installStep]time.Duration{ + stepApplyingFoundationDB: applyingFoundationDBDeadline, + stepAwaitingFoundationDB: awaitingFoundationDBDeadline, + stepApplyingDatastore: applyingDatastoreDeadline, + stepApplyingAPI: applyingAPIDeadline, + stepAwaitingAPI: awaitingAPIDeadline, +} + +// opsGraphs declares one state graph per operation action over one step type. +// +// Restart and Upgrade both carry Draining, because both roll the same +// Deployment: an upgrade replaces the management API's image, which recycles its +// pods exactly as a restart does, and an operation interrupted by that has been +// interrupted whichever field caused it. Backup carries none, because asking +// FoundationDB for a snapshot recycles nothing. +// +// Upgrade drains after Preflight rather than before it. An upgrade refused for +// naming the image already running is refused in a moment, and making it first +// wait out a twenty-minute node add would spend the fleet's time to reach an +// error that was available immediately. +func opsGraphs() statemachine.MultiConfig[opsStep] { + return statemachine.MultiConfig[opsStep]{ + action(simplyblockv1alpha2.ControlPlaneOpsActionRestart): { + Initial: stepDraining, + States: map[opsStep]statemachine.StateDef[opsStep]{ + stepDraining: { + To: []opsStep{stepRestarting}, + OnEnter: deadline[opsStep](drainingDeadline), + }, + stepRestarting: { + To: []opsStep{stepAwaiting}, + OnEnter: deadline[opsStep](restartingDeadline), + }, + stepAwaiting: {OnEnter: deadline[opsStep](awaitingOpsDeadline)}, + }, + }, + + action(simplyblockv1alpha2.ControlPlaneOpsActionUpgrade): { + Initial: stepPreflight, + States: map[opsStep]statemachine.StateDef[opsStep]{ + stepPreflight: { + To: []opsStep{stepDraining}, + OnEnter: deadline[opsStep](preflightDeadline), + }, + stepDraining: { + To: []opsStep{stepApplying}, + OnEnter: deadline[opsStep](drainingDeadline), + }, + stepApplying: { + To: []opsStep{stepAwaiting}, + OnEnter: deadline[opsStep](applyingDeadline), + }, + stepAwaiting: { + To: []opsStep{stepVerifying}, + OnEnter: deadline[opsStep](awaitingOpsDeadline), + }, + stepVerifying: {OnEnter: deadline[opsStep](verifyingDeadline)}, + }, + }, + + action(simplyblockv1alpha2.ControlPlaneOpsActionBackup): { + Initial: stepRequesting, + States: map[opsStep]statemachine.StateDef[opsStep]{ + stepRequesting: { + To: []opsStep{stepAwaiting}, + OnEnter: deadline[opsStep](requestingDeadline), + }, + stepAwaiting: {OnEnter: deadline[opsStep](awaitingOpsDeadline)}, + }, + }, + } +} + +// opsInitialDeadlines are the budgets of the step each action's machine is born +// in. A machine is already in its initial state when it is built, so that +// state's OnEnter never runs and the graph's deadline for it is never set. +// Setting it explicitly is what stops the first step of every operation from +// being the one step that cannot time out. +var opsInitialDeadlines = map[statemachine.Action]time.Duration{ + action(simplyblockv1alpha2.ControlPlaneOpsActionRestart): drainingDeadline, + action(simplyblockv1alpha2.ControlPlaneOpsActionUpgrade): preflightDeadline, + action(simplyblockv1alpha2.ControlPlaneOpsActionBackup): requestingDeadline, +} + +// abortableSteps are the steps from which an abort stops the operation cleanly. +// +// The line is whether anything has been changed yet. Draining and Preflight have +// performed no side effect at all, and Requesting has not yet created the +// backup. Everything past those has rolled a Deployment or written an image onto +// the entity, and the operation is what drives that rollout to completion. +// +// It is a table beside the graph rather than an edge in it: the phase already +// carries what a terminal Aborted step would say. A test asserts every step here +// is one some graph declares, so the two cannot drift. +var abortableSteps = map[opsStep]bool{ + stepDraining: true, + stepPreflight: true, + stepRequesting: true, +} + +// abortable reports whether an abort asked for while the operation sits on this +// step can be honored. +func abortable(current opsStep) bool { return abortableSteps[current] } + +// action converts the API's action enum into the MultiConfig's key. The +// conversion exists because statemachine.Action is a concrete string type rather +// than a second type parameter, and doing it in one place keeps the graph +// literal readable. +func action(a simplyblockv1alpha2.ControlPlaneOpsAction) statemachine.Action { + return statemachine.Action(a) +} diff --git a/operator/internal/controllers/controlplane/helpers_test.go b/operator/internal/controllers/controlplane/helpers_test.go new file mode 100644 index 000000000..08e6f2441 --- /dev/null +++ b/operator/internal/controllers/controlplane/helpers_test.go @@ -0,0 +1,237 @@ +// The fixtures and the fake control plane the tests in this package are built +// on. +// +// The scheme is built once and carries every kind an install writes, because an +// apply resolves a typed object's kind through it and a missing registration +// fails at a place that says nothing about which object was missing. + +package controlplane + +import ( + "context" + "testing" + + appsv1 "k8s.io/api/apps/v1" + corev1 "k8s.io/api/core/v1" + rbacv1 "k8s.io/api/rbac/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/runtime" + clientgoscheme "k8s.io/client-go/kubernetes/scheme" + "k8s.io/client-go/tools/events" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + "github.com/simplyblock/atlas/ptr" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +const ( + testNamespace = "simplyblock" + testImage = "quay.io/simplyblock-io/simplyblock:26.2.8" +) + +// testScheme carries every kind the install writes, plus the FoundationDB kinds +// as unstructured so a fake client can hold one. +func testScheme(t *testing.T) *runtime.Scheme { + t.Helper() + scheme := runtime.NewScheme() + for _, add := range []func(*runtime.Scheme) error{ + clientgoscheme.AddToScheme, + simplyblockv1alpha2.AddToScheme, + } { + if err := add(scheme); err != nil { + t.Fatalf("build the scheme: %v", err) + } + } + // The FoundationDB kinds have no Go type in this repository, so the fake + // client is told their list kind by hand. Without it a Get of one fails with + // a missing registration, which reads as a bug in the code under test. + scheme.AddKnownTypeWithName(fdbClusterGVK, &unstructuredStub{}) + scheme.AddKnownTypeWithName(fdbClusterGVK.GroupVersion().WithKind("FoundationDBClusterList"), + &unstructuredListStub{}) + scheme.AddKnownTypeWithName(fdbBackupGVK, &unstructuredStub{}) + scheme.AddKnownTypeWithName(fdbBackupGVK.GroupVersion().WithKind("FoundationDBBackupList"), + &unstructuredListStub{}) + return scheme +} + +// managedControlPlane is the fixture every managed test starts from: the +// singleton, in the namespace, naming an image. +func localControlPlane() *simplyblockv1alpha2.ControlPlane { + return &simplyblockv1alpha2.ControlPlane{ + ObjectMeta: metav1.ObjectMeta{Name: SingletonName, Namespace: testNamespace}, + Spec: simplyblockv1alpha2.ControlPlaneSpec{ + Source: simplyblockv1alpha2.ControlPlaneSource{ + Local: &simplyblockv1alpha2.LocalControlPlane{Image: testImage}, + }, + }, + } +} + +// managedControlPlane names a control plane that already exists. +func managedControlPlane(endpoint string) *simplyblockv1alpha2.ControlPlane { + return &simplyblockv1alpha2.ControlPlane{ + ObjectMeta: metav1.ObjectMeta{Name: SingletonName, Namespace: testNamespace}, + Spec: simplyblockv1alpha2.ControlPlaneSpec{ + Source: simplyblockv1alpha2.ControlPlaneSource{ + Managed: &simplyblockv1alpha2.ManagedControlPlane{ + Endpoint: endpoint, + CredentialsSecretRef: &corev1.LocalObjectReference{Name: "cp-token"}, + }, + }, + }, + } +} + +// stubProber answers the two reads without an HTTP server, which is what lets +// the phase branches be exercised at all: every one of them turns on what the +// control plane said. +type stubProber struct { + ready bool + readyMessage string + version string + versionErr error + + // readyCalls and versionCalls count what the reconciler asked for, which is + // how a test asserts that a held step did not probe. + readyCalls int + versionCalls int +} + +func (p *stubProber) Ready(context.Context, string) (bool, string) { + p.readyCalls++ + return p.ready, p.readyMessage +} + +func (p *stubProber) Version(context.Context, string) (string, error) { + p.versionCalls++ + return p.version, p.versionErr +} + +// newClient builds a fake client holding the given objects, with the status +// subresource enabled for the two kinds this package writes one on. +func newClient(t *testing.T, objects ...client.Object) client.Client { + t.Helper() + return fake.NewClientBuilder(). + WithScheme(testScheme(t)). + WithObjects(objects...). + WithStatusSubresource( + &simplyblockv1alpha2.ControlPlane{}, + &simplyblockv1alpha2.ControlPlaneOps{}, + ). + Build() +} + +// deployment is a running workload with a ready count, which is what the +// component table reads. +func deployment(name string, desired, ready int32) *appsv1.Deployment { + return &appsv1.Deployment{ + ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: testNamespace}, + Spec: appsv1.DeploymentSpec{Replicas: ptr.To(desired)}, + Status: appsv1.DeploymentStatus{ReadyReplicas: ready}, + } +} + +func statefulSet(name string, desired, ready int32) *appsv1.StatefulSet { + return &appsv1.StatefulSet{ + ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: testNamespace}, + Spec: appsv1.StatefulSetSpec{Replicas: ptr.To(desired)}, + Status: appsv1.StatefulSetStatus{ReadyReplicas: ready}, + } +} + +// componentStatus is one row of what observe produces, built directly so a phase +// test does not have to stand up the workloads behind it. +func componentStatus(name string, desired, ready int32, essential bool) simplyblockv1alpha2.ControlPlaneComponentStatus { + return simplyblockv1alpha2.ControlPlaneComponentStatus{ + Name: name, Desired: desired, Ready: ready, Essential: essential, + } +} + +// findObject looks up one built object by kind and name, which is how the +// workload tests assert that a step writes what it says it writes. +func findObject(objects []client.Object, name string) client.Object { + for _, obj := range objects { + if obj.GetName() == name { + return obj + } + } + return nil +} + +// findRole, findDeployment, and the rest narrow that to the type a test asserts +// against, failing rather than returning nil so the assertion that follows is +// about the object rather than about a nil pointer. +func findDeployment(t *testing.T, objects []client.Object, name string) *appsv1.Deployment { + t.Helper() + obj := findObject(objects, name) + d, ok := obj.(*appsv1.Deployment) + if !ok { + t.Fatalf("no Deployment named %q among the built objects", name) + } + return d +} + +func findClusterRole(t *testing.T, objects []client.Object, name string) *rbacv1.ClusterRole { + t.Helper() + obj := findObject(objects, name) + role, ok := obj.(*rbacv1.ClusterRole) + if !ok { + t.Fatalf("no ClusterRole named %q among the built objects", name) + } + return role +} + +// unstructuredStub and unstructuredListStub let the scheme name the FoundationDB +// kinds. The fake client stores and returns unstructured objects for them; these +// exist only so the kind resolves. +type unstructuredStub struct { + metav1.TypeMeta `json:",inline"` + metav1.ObjectMeta `json:"metadata,omitempty"` +} + +func (in *unstructuredStub) DeepCopyObject() runtime.Object { + out := *in + return &out +} + +type unstructuredListStub struct { + metav1.TypeMeta `json:",inline"` + metav1.ListMeta `json:"metadata,omitempty"` + Items []unstructuredStub `json:"items"` +} + +func (in *unstructuredListStub) DeepCopyObject() runtime.Object { + out := *in + out.Items = append([]unstructuredStub(nil), in.Items...) + return &out +} + +// recordingRecorder remembers the events it is told about. events.EventRecorder +// has one method that matters here, and a fake is less machinery than a real +// broadcaster with a fake clientset behind it. +type recordingRecorder struct { + reasons []string + types []string +} + +func (r *recordingRecorder) Eventf( + _ runtime.Object, _ runtime.Object, eventType, reason, _, _ string, _ ...any, +) { + r.types = append(r.types, eventType) + r.reasons = append(r.reasons, reason) +} + +// count returns how many events carried a reason, which is what a test +// asserting "once, not once per pass" needs. +func (r *recordingRecorder) count(reason string) int { + n := 0 + for _, got := range r.reasons { + if got == reason { + n++ + } + } + return n +} + +var _ events.EventRecorder = (*recordingRecorder)(nil) diff --git a/operator/internal/controllers/controlplane/install_test.go b/operator/internal/controllers/controlplane/install_test.go new file mode 100644 index 000000000..d254e0cce --- /dev/null +++ b/operator/internal/controllers/controlplane/install_test.go @@ -0,0 +1,383 @@ +// The installation machine: what each step applies, what it holds on, and the +// two properties the whole thing rests on. +// +// The first is that every step is an apply, so re-entering one is a no-op. That +// is what lets the machine carry no triggered flag, and it is why an install +// interrupted anywhere resumes rather than restarts. +// +// The second is that a held step reports what it is waiting for, read from the +// thing being waited on. An install stalled on FoundationDB names the database +// rather than the operator, which is the difference between a message somebody +// can act on and one that sends them to the logs. + +package controlplane + +import ( + "context" + "strings" + "testing" + + appsv1 "k8s.io/api/apps/v1" + corev1 "k8s.io/api/core/v1" + rbacv1 "k8s.io/api/rbac/v1" + "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" + "sigs.k8s.io/controller-runtime/pkg/client" + + "github.com/simplyblock/atlas/ptr" + "github.com/simplyblock/atlas/statemachine" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// The graph is a line from the first apply to the readiness wait, and every step +// declares a successor except the last. The reconciler takes the first edge as +// the only edge, which holds exactly while that is true. +func TestTheInstallationGraphIsALine(t *testing.T) { + graph := installGraph() + + if graph.Initial != stepApplyingFoundationDB { + t.Errorf("the install starts at %s, want %s", graph.Initial, stepApplyingFoundationDB) + } + + want := map[installStep]installStep{ + stepApplyingFoundationDB: stepAwaitingFoundationDB, + stepAwaitingFoundationDB: stepApplyingDatastore, + stepApplyingDatastore: stepApplyingAPI, + stepApplyingAPI: stepAwaitingAPI, + } + for from, to := range want { + state, declared := graph.States[from] + if !declared { + t.Fatalf("%s is not a state of the graph", from) + } + if len(state.To) != 1 || state.To[0] != to { + t.Errorf("%s declares successors %v, want exactly [%s]", from, state.To, to) + } + } + + if last := graph.States[stepAwaitingAPI]; len(last.To) != 0 { + t.Errorf("%s declares successors %v, want none: reaching it is what makes the "+ + "control plane Available", stepAwaitingAPI, last.To) + } +} + +// Every step has a budget, and the machine is built from the same table the +// reconciler sets the first step's deadline from, so every step can time out. +func TestEveryInstallationStepHasABudget(t *testing.T) { + for step := range installGraph().States { + if _, ok := installStepBudgets[step]; !ok { + t.Errorf("%s has no budget, so a deadline is never set for it", step) + } + } + for step := range installStepBudgets { + if _, ok := installGraph().States[step]; !ok { + t.Errorf("%s has a budget and is not a state of the graph", step) + } + } +} + +// The step enum the API admits and the states the graph declares are the same +// set. They are two lists in two files, and a value added to one and not the +// other is a status the API rejects or a step nothing can reach. +func TestTheStepEnumAndTheGraphAgree(t *testing.T) { + declared := map[string]bool{} + for _, state := range statemachine.DeclaredStates(installGraph()) { + declared[state] = true + } + + admitted := []installStep{ + simplyblockv1alpha2.ControlPlaneStepApplyingFoundationDB, + simplyblockv1alpha2.ControlPlaneStepAwaitingFoundationDB, + simplyblockv1alpha2.ControlPlaneStepApplyingDatastore, + simplyblockv1alpha2.ControlPlaneStepApplyingAPI, + simplyblockv1alpha2.ControlPlaneStepAwaitingAPI, + } + if len(declared) != len(admitted) { + t.Errorf("the graph declares %d states and the enum admits %d", len(declared), len(admitted)) + } + for _, step := range admitted { + if !declared[string(step)] { + t.Errorf("the enum admits %s and no graph state declares it", step) + } + } +} + +// Re-applying a step puts back what somebody deleted and corrects what somebody +// edited, which is the property steady state rests on and the reason no +// operation exists for checking the install. It is also what makes re-entering a +// step a no-op, and therefore why the machine carries no triggered flag. +func TestReApplyingAStepCorrectsWhatWasChangedUnderIt(t *testing.T) { + ctx := context.Background() + cp := localControlPlane() + c := newClient(t, cp) + scheme := testScheme(t) + + if err := applyAll(ctx, c, cp, scheme, managementAPIObjects(cp)); err != nil { + t.Fatalf("the first apply: %v", err) + } + + // Somebody scales the management API to one replica and deletes its Service. + var api appsv1.Deployment + key := client.ObjectKey{Namespace: testNamespace, Name: ComponentWebAPI} + if err := c.Get(ctx, key, &api); err != nil { + t.Fatalf("read the management API: %v", err) + } + api.Spec.Replicas = ptr.To(int32(1)) + if err := c.Update(ctx, &api); err != nil { + t.Fatalf("scale the management API down: %v", err) + } + if err := c.Delete(ctx, webAPIService(testNamespace)); err != nil { + t.Fatalf("delete the Service: %v", err) + } + + if err := applyAll(ctx, c, cp, scheme, managementAPIObjects(cp)); err != nil { + t.Fatalf("the second apply: %v", err) + } + + if err := c.Get(ctx, key, &api); err != nil { + t.Fatalf("read the management API back: %v", err) + } + if api.Spec.Replicas == nil || *api.Spec.Replicas != 2 { + t.Errorf("replicas = %v after re-applying, want the spec's 2", api.Spec.Replicas) + } + + var service corev1.Service + if err := c.Get(ctx, key, &service); err != nil { + t.Errorf("the Service was not put back: %v", err) + } + + var deployments appsv1.DeploymentList + if err := c.List(ctx, &deployments, client.InNamespace(testNamespace)); err != nil { + t.Fatalf("list the deployments: %v", err) + } + if len(deployments.Items) != 5 { + names := make([]string, 0, len(deployments.Items)) + for _, d := range deployments.Items { + names = append(names, d.Name) + } + t.Errorf("two passes produced %d deployments (%v), want the same 5 both times", + len(deployments.Items), names) + } +} + +// A namespaced object is a child of the ControlPlane and goes with the garbage +// collector. A cluster-scoped one cannot be, because Kubernetes never collects a +// cluster-scoped object owned by a namespaced one, so it carries the managed-by +// label the finalizer deletes on instead. +func TestOwnershipFollowsTheObjectsScope(t *testing.T) { + cp := localControlPlane() + scheme := testScheme(t) + + for _, obj := range append(foundationDBObjects(cp), managementAPIObjects(cp)...) { + if err := setOwnership(cp, obj, scheme); err != nil { + t.Fatalf("set ownership on %s: %v", obj.GetName(), err) + } + + if obj.GetNamespace() != "" { + if len(obj.GetOwnerReferences()) == 0 { + t.Errorf("%s is namespaced and carries no owner reference, so deleting the "+ + "ControlPlane would leave it behind", obj.GetName()) + } + continue + } + + if !mayDelete(obj) { + t.Errorf("%s is cluster-scoped and carries no managed-by label, so the finalizer "+ + "would leave it behind", obj.GetName()) + } + if len(obj.GetOwnerReferences()) != 0 { + t.Errorf("%s is cluster-scoped and carries an owner reference to a namespaced "+ + "object, which Kubernetes never resolves and never collects", obj.GetName()) + } + } +} + +// The finalizer deletes only what this controller marked. A cluster-scoped +// object is shared ground, and deleting one carrying another controller's mark, +// or none, takes away somebody else's RBAC. +func TestTheFinalizerLeavesUnmarkedClusterScopedObjectsAlone(t *testing.T) { + somebodyElses := fdbManagerClusterRoleObject() + somebodyElses.Labels = map[string]string{managedByLabel: "somebody-else"} + + c := newClient(t, somebodyElses) + + if err := deleteIfMarked(context.Background(), c, fdbManagerClusterRoleObject()); err != nil { + t.Fatalf("deleteIfMarked: %v", err) + } + + var survived rbacv1.ClusterRole + if err := c.Get(context.Background(), + client.ObjectKey{Name: somebodyElses.Name}, &survived); err != nil { + t.Errorf("the ClusterRole was deleted, and this controller never marked it: %v", err) + } +} + +// AwaitingFoundationDB reports what it is waiting for, read from the database's +// own status, so a stalled install names the coordinator rather than the +// operator. +func TestAwaitingFoundationDBNamesWhatItIsWaitingOn(t *testing.T) { + for _, tc := range []struct { + name string + health fdbHealth + want string + }{ + { + name: "the cluster has not been created", + health: fdbHealth{}, + want: "has not been created yet", + }, + { + name: "the cluster is not available", + health: fdbHealth{found: true, desired: 7, reconciled: 3}, + want: "has 3 of 7 process groups reconciled", + }, + { + name: "the cluster is available and not fully replicated", + health: fdbHealth{found: true, available: true, desired: 7, reconciled: 7}, + want: "not yet fully replicated", + }, + { + name: "the cluster is ready", + health: fdbHealth{ + found: true, available: true, fullReplication: true, desired: 7, reconciled: 7, + }, + want: "", + }, + } { + t.Run(tc.name, func(t *testing.T) { + got := tc.health.waitingOn() + switch { + case tc.want == "" && got != "": + t.Errorf("waitingOn = %q, want nothing: the step is finished", got) + case tc.want != "" && !strings.Contains(got, tc.want): + t.Errorf("waitingOn = %q, want it to mention %q", got, tc.want) + } + }) + } +} + +// A FoundationDBCluster that is not there is not an error. The apply created it +// and the cache has not caught up, which is the ordinary state on the pass right +// after ApplyingFoundationDB. +func TestReadingAnAbsentFoundationDBClusterIsNotAnError(t *testing.T) { + c := newClient(t) + + health, err := readFoundationDB(context.Background(), c, testNamespace) + if err != nil { + t.Fatalf("readFoundationDB: %v", err) + } + if health.found { + t.Error("a cluster that does not exist was reported as found") + } +} + +// AwaitingAPI holds on the probe rather than on the pod counts, because the +// question it answers is whether the control plane can be reached at all. +func TestAwaitingAPIHoldsUntilTheProbePasses(t *testing.T) { + cp := localControlPlane() + prober := &stubProber{ready: false, readyMessage: "connection refused"} + r := &ControlPlaneReconciler{ + Client: newClient(t, cp), Scheme: testScheme(t), Prober: prober, + } + + done, held, err := r.performInstallStep(context.Background(), cp, stepAwaitingAPI) + if err != nil { + t.Fatalf("performInstallStep: %v", err) + } + if done { + t.Fatal("the step finished while the probe was failing") + } + if !strings.Contains(held, "connection refused") { + t.Errorf("held on %q, want the probe's own words", held) + } + + prober.ready = true + done, _, err = r.performInstallStep(context.Background(), cp, stepAwaitingAPI) + if err != nil { + t.Fatalf("performInstallStep: %v", err) + } + if !done { + t.Error("the step held while the probe was passing") + } +} + +// The FoundationDBCluster is built with the coordinator count and the redundancy +// mode agreeing, so the database is never told to keep more copies than it has +// processes to keep them on. +func TestTheCoordinatorCountAndTheRedundancyModeAgree(t *testing.T) { + for _, tc := range []struct { + replicas int32 + mode string + }{ + {1, "single"}, + {3, "double"}, + {5, "triple"}, + {7, "triple"}, + } { + cp := localControlPlane() + cp.Spec.Source.Local.FoundationDB = &simplyblockv1alpha2.FoundationDBSpec{ + Replicas: &tc.replicas, + } + + cluster := foundationDBCluster(cp) + mode, _, _ := unstructured.NestedString(cluster.Object, + "spec", "databaseConfiguration", "redundancy_mode") + logs, _, _ := unstructured.NestedInt64(cluster.Object, "spec", "processCounts", "log") + + if mode != tc.mode { + t.Errorf("%d coordinators: redundancy_mode = %q, want %q", tc.replicas, mode, tc.mode) + } + if logs != int64(tc.replicas) { + t.Errorf("%d coordinators: processCounts.log = %d, want %d", tc.replicas, logs, tc.replicas) + } + } +} + +// An unset storage class leaves the field out rather than writing an empty +// string, which is the difference between the cluster's default class and a +// class literally named nothing. +func TestAnUnsetStorageClassIsAbsentRatherThanEmpty(t *testing.T) { + cp := localControlPlane() + + claim := volumeClaimSpec(foundationDBSpecOf(cp)) + + if _, present := claim["storageClassName"]; present { + t.Error("storageClassName is written for a spec that states none") + } + + cp.Spec.Source.Local.FoundationDB = &simplyblockv1alpha2.FoundationDBSpec{ + StorageClassName: "fast", + } + claim = volumeClaimSpec(foundationDBSpecOf(cp)) + if claim["storageClassName"] != "fast" { + t.Errorf("storageClassName = %v, want fast", claim["storageClassName"]) + } +} + +// The scheduling an operator set reaches every pod the install creates, so the +// control plane stays where the deployment put it. +func TestSchedulingReachesEveryPodTheInstallCreates(t *testing.T) { + cp := localControlPlane() + cp.Spec.Source.Local.NodeSelector = map[string]string{"simplyblock.io/control-plane": "true"} + cp.Spec.Source.Local.Tolerations = []corev1.Toleration{{ + Key: "simplyblock.io/dedicated", Operator: corev1.TolerationOpExists, + }} + + for _, obj := range append(foundationDBObjects(cp), append(datastoreObjects(cp), managementAPIObjects(cp)...)...) { + var spec *corev1.PodSpec + switch typed := obj.(type) { + case *appsv1.Deployment: + spec = &typed.Spec.Template.Spec + case *appsv1.StatefulSet: + spec = &typed.Spec.Template.Spec + default: + continue + } + + if len(spec.NodeSelector) == 0 { + t.Errorf("%s carries no node selector", obj.GetName()) + } + if len(spec.Tolerations) == 0 { + t.Errorf("%s carries no tolerations", obj.GetName()) + } + } +} diff --git a/operator/internal/controllers/controlplane/managementapi.go b/operator/internal/controllers/controlplane/managementapi.go new file mode 100644 index 000000000..859679bde --- /dev/null +++ b/operator/internal/controllers/controlplane/managementapi.go @@ -0,0 +1,596 @@ +// The management API and the workloads beside it: what the ApplyingAPI step +// applies, and the Service status.endpoint resolves to. +// +// Four workloads share one image, one account, and one configuration, and they +// differ in what they run. The management API serves the REST surface every +// controller in this operator calls. The monitoring pool watches nodes, devices, +// volumes, and snapshots and writes what it sees. The task runner executes the +// queued work those two produce. The admin surface runs nothing and exists to be +// exec'd into, which is why its command is a sleep. +// +// The exporter is the fifth object here and is not built from that image: it +// reads FoundationDB's status JSON through the client library and publishes it +// as metrics. It sits in this step rather than the FoundationDB one because what +// it needs is the cluster file the database's own step produces, so it can only +// run once that step has finished. + +package controlplane + +import ( + appsv1 "k8s.io/api/apps/v1" + corev1 "k8s.io/api/core/v1" + rbacv1 "k8s.io/api/rbac/v1" + "k8s.io/apimachinery/pkg/api/resource" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/util/intstr" + "sigs.k8s.io/controller-runtime/pkg/client" + + "github.com/simplyblock/atlas/ptr" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// fdbExporterImage turns FoundationDB's machine-readable status into Prometheus +// metrics. It is upstream's rather than simplyblock's, and pinned for the same +// reason the database's images are. +const fdbExporterImage = "aikoven/foundationdb-exporter:3.1.0" + +// restartOnClusterFileChange is what makes a coordinator change reach the pods +// that hold a connection to the old one. The reloader watches the ConfigMap the +// FoundationDB operator writes and rolls whatever carries this annotation. +// +// The reloader itself is a chart dependency and stays one. A control plane +// installed without it keeps working and takes a coordinator change on the next +// restart instead of at once, which is a degradation rather than a break. +var restartOnClusterFileChange = map[string]string{ + "reloader.stakater.com/auto": "true", + "reloader.stakater.com/configmap": clusterFileConfigMapName, +} + +// managementAPIObjects is everything the ApplyingAPI step writes, in the order it +// depends on itself: the account and the configuration, then the roles that name +// the account, then the workloads that run under it, then the Service in front +// of the one that serves. +func managementAPIObjects(cp *simplyblockv1alpha2.ControlPlane) []client.Object { + ns := cp.Namespace + return []client.Object{ + serviceAccount(ns, serviceAccountName), + sharedConfigMap(ns), + controlPlaneClusterRole(), + controlPlaneClusterRoleBinding(ns), + serviceReaderClusterRole(), + serviceReaderClusterRoleBinding(ns), + webAPIDeployment(cp), + webAPIService(ns), + tasksDeployment(cp), + monitoringDeployment(cp), + adminControlDeployment(cp), + fdbExporterDeployment(cp), + fdbExporterService(ns), + } +} + +// managementAPIClusterScoped is the half of that set the garbage collector will +// not remove with the ControlPlane, and which the finalizer therefore deletes. +func managementAPIClusterScoped() []client.Object { + return []client.Object{ + controlPlaneClusterRole(), + controlPlaneClusterRoleBinding(""), + serviceReaderClusterRole(), + serviceReaderClusterRoleBinding(""), + } +} + +// sharedConfigMap holds the one setting every workload reads at run time. It is +// a ConfigMap rather than a field on each Deployment so that changing the log +// level of a whole control plane is one edit, and so that the reloader can roll +// the pods that read it. +func sharedConfigMap(namespace string) *corev1.ConfigMap { + return &corev1.ConfigMap{ + ObjectMeta: metav1.ObjectMeta{Name: configMapName, Namespace: namespace}, + Data: map[string]string{ + logLevelKey: "DEBUG", + "LOG_DELETION_INTERVAL": "3d", + }, + } +} + +// controlPlaneClusterRole is what the control plane does to Kubernetes. +// +// It is wide, and the width is the control plane's own architecture rather than +// this install's choice: it reads and patches the storage-node workloads it +// manages, execs into their pods to drive SPDK, labels nodes, and reviews the +// tokens presented to its API. Narrowing it is worthwhile and is a change to the +// control plane rather than to the object that grants it. +func controlPlaneClusterRole() *rbacv1.ClusterRole { + return &rbacv1.ClusterRole{ + ObjectMeta: metav1.ObjectMeta{Name: clusterRoleName}, + Rules: []rbacv1.PolicyRule{ + { + APIGroups: []string{""}, + Resources: []string{"configmaps"}, + Verbs: []string{"get", "list", "watch", "patch", "update"}, + }, + { + APIGroups: []string{"", "apps"}, + Resources: []string{"pods", "deployments", "statefulsets", "daemonsets"}, + Verbs: []string{"get", "list", "watch", "patch", "update"}, + }, + { + APIGroups: []string{""}, + Resources: []string{"pods/log"}, + Verbs: []string{"get", "list"}, + }, + { + // Exec is how the control plane drives the storage nodes' + // processes, and it is the strongest grant in this role: exec + // into a privileged pod is that pod's privilege. + APIGroups: []string{""}, + Resources: []string{"pods/exec"}, + Verbs: []string{"create", "get", "list", "watch", "patch", "update"}, + }, + { + APIGroups: []string{""}, + Resources: []string{"nodes"}, + Verbs: []string{"get", "list", "watch", "patch", "update"}, + }, + { + APIGroups: []string{"storage.simplyblock.io"}, + Resources: []string{ + "pools", "lvols", "storageclusters", "storagenodesets", "devices", + "tasks", "storagebackups", + }, + Verbs: []string{"get", "list", "patch", "update", "watch"}, + }, + { + APIGroups: []string{"storage.simplyblock.io"}, + Resources: []string{ + "pools/status", "lvols/status", "storageclusters/status", + "storagenodesets/status", "devices/status", "tasks/status", + "storagebackups/status", + }, + Verbs: []string{"get", "patch", "update"}, + }, + { + // A TokenReview is how the control plane authenticates a caller + // presenting a service-account token instead of the static + // cluster secret. + APIGroups: []string{"authentication.k8s.io"}, + Resources: []string{"tokenreviews"}, + Verbs: []string{"create"}, + }, + }, + } +} + +func controlPlaneClusterRoleBinding(namespace string) *rbacv1.ClusterRoleBinding { + return &rbacv1.ClusterRoleBinding{ + ObjectMeta: metav1.ObjectMeta{Name: clusterRoleBindingName}, + RoleRef: rbacv1.RoleRef{ + APIGroup: rbacv1.GroupName, Kind: "ClusterRole", Name: clusterRoleName, + }, + Subjects: []rbacv1.Subject{{ + Kind: "ServiceAccount", Name: serviceAccountName, Namespace: namespace, + }}, + } +} + +// serviceReaderClusterRole lets the namespace's default account resolve +// services, endpoints, and nodes. It is bound to `default` rather than to the +// control plane's own account because what needs it is the tooling run inside +// pods that did not name an account. +func serviceReaderClusterRole() *rbacv1.ClusterRole { + return &rbacv1.ClusterRole{ + ObjectMeta: metav1.ObjectMeta{Name: serviceReaderRoleName}, + Rules: []rbacv1.PolicyRule{{ + APIGroups: []string{""}, + Resources: []string{"services", "pods", "endpoints", "nodes"}, + Verbs: []string{"get", "list", "watch"}, + }}, + } +} + +func serviceReaderClusterRoleBinding(namespace string) *rbacv1.ClusterRoleBinding { + return &rbacv1.ClusterRoleBinding{ + ObjectMeta: metav1.ObjectMeta{Name: serviceReaderBindingName}, + RoleRef: rbacv1.RoleRef{ + APIGroup: rbacv1.GroupName, Kind: "ClusterRole", Name: serviceReaderRoleName, + }, + Subjects: []rbacv1.Subject{{ + Kind: "ServiceAccount", Name: "default", Namespace: namespace, + }}, + } +} + +// webAPIDeployment is the management API: the endpoint every controller in this +// operator calls, and the one workload whose absence is an outage. +// +// Two replicas by default, spread across machines, rolled one at a time with no +// surge. That combination is what lets a pod be replaced while the control plane +// keeps answering, and three things in the design rest on it: Degraded is a +// component below its desired count while the probe still passes, a Restart +// recycles the workload after draining rather than instead of serving, and an +// Upgrade rolls this Deployment. +func webAPIDeployment(cp *simplyblockv1alpha2.ControlPlane) *appsv1.Deployment { + managed := cp.Spec.Source.Local + labels := map[string]string{appLabel: ComponentWebAPI} + + //nolint:prealloc // the literal is the declaration; the append below is the shared set + env := []corev1.EnvVar{ + logLevelEnv(), + {Name: "LVOL_NVMF_PORT_START", Value: lvolNVMfPortStart}, + {Name: "ENABLE_MONITORING", Value: "false"}, + namespaceEnv(), + {Name: "FLASK_DEBUG", Value: "False"}, + {Name: "FLASK_ENV", Value: "production"}, + // The operator's own account is the one caller the control plane has to + // trust before anything else works: every controller in this process + // authenticates with its token. + {Name: "SB_K8S_ADMIN_SERVICE_ACCOUNTS", + Value: "system:serviceaccount:" + cp.Namespace + ":simplyblock-operator"}, + {Name: "SB_K8S_METRICS_SERVICE_ACCOUNTS", + Value: "system:serviceaccount:" + cp.Namespace + ":simplyblock-prometheus"}, + } + env = append(env, prometheusEnv()...) + + spec := corev1.PodSpec{ + ServiceAccountName: serviceAccountName, + Affinity: spreadAcrossHosts(ComponentWebAPI), + Containers: []corev1.Container{{ + Name: "webappapi", + Image: localImage(cp), + ImagePullPolicy: pullPolicyOf(managed), + Command: []string{"python3", "simplyblock_web/app.py"}, + Ports: []corev1.ContainerPort{{ContainerPort: webAPIPort}}, + Env: env, + VolumeMounts: []corev1.VolumeMount{clusterFileMount()}, + Resources: webAPIResources(managed), + }}, + Volumes: []corev1.Volume{clusterFileVolumeSource()}, + } + scheduling(managed, &spec) + + return &appsv1.Deployment{ + ObjectMeta: metav1.ObjectMeta{ + Name: ComponentWebAPI, + Namespace: cp.Namespace, + Annotations: restartOnClusterFileChange, + }, + Spec: appsv1.DeploymentSpec{ + Replicas: ptr.To(apiReplicas(managed)), + Strategy: appsv1.DeploymentStrategy{ + Type: appsv1.RollingUpdateDeploymentStrategyType, + RollingUpdate: &appsv1.RollingUpdateDeployment{ + MaxSurge: ptr.To(intstr.FromInt32(0)), + MaxUnavailable: ptr.To(intstr.FromInt32(1)), + }, + }, + Selector: &metav1.LabelSelector{MatchLabels: labels}, + Template: corev1.PodTemplateSpec{ + ObjectMeta: metav1.ObjectMeta{ + Labels: labels, + Annotations: map[string]string{"log-collector/enabled": "true"}, + }, + Spec: spec, + }, + }, + } +} + +// webAPIResources is what the management API asks for, or what the spec states. +// The default is the chart's: enough headroom for the request volume a fleet +// produces, and a limit that contains a leak rather than one anybody has +// measured against. +func webAPIResources(managed *simplyblockv1alpha2.LocalControlPlane) corev1.ResourceRequirements { + if managed != nil && (len(managed.Resources.Requests) > 0 || len(managed.Resources.Limits) > 0) { + return managed.Resources + } + return corev1.ResourceRequirements{ + Requests: corev1.ResourceList{ + corev1.ResourceCPU: resource.MustParse("200m"), + corev1.ResourceMemory: resource.MustParse("512Mi"), + }, + Limits: corev1.ResourceList{ + corev1.ResourceCPU: resource.MustParse("500m"), + corev1.ResourceMemory: resource.MustParse("2Gi"), + }, + } +} + +// webAPIService is what status.endpoint resolves to, and the name the CSI +// driver's configuration and the metrics scrape both carry. +func webAPIService(namespace string) *corev1.Service { + return &corev1.Service{ + ObjectMeta: metav1.ObjectMeta{Name: ComponentWebAPI, Namespace: namespace}, + Spec: corev1.ServiceSpec{ + Selector: map[string]string{appLabel: ComponentWebAPI}, + Ports: []corev1.ServicePort{ + {Name: "http", Port: webAPIPort, TargetPort: intstrFromInt(webAPIPort)}, + }, + }, + } +} + +// monitoringServices are the watchers the control plane runs: one per subject it +// keeps state about. They are one pod rather than ten because each is a small +// polling loop and the pod is what shares the cluster file and the log level. +// +// The pod runs on the host network, which is how the node and device monitors +// reach the storage nodes' management addresses directly. +func monitoringServices() []service { + return []service{ + {name: "storage-node-monitor", module: "simplyblock_core/services/storage_node_monitor.py"}, + { + name: "mgmt-node-monitor", + module: "simplyblock_core/services/mgmt_node_monitor.py", + extraEnv: []corev1.EnvVar{{Name: "BACKEND_TYPE", Value: "k8s"}}, + }, + {name: "lvol-stats-collector", module: "simplyblock_core/services/lvol_stat_collector.py"}, + {name: "main-distr-event-collector", module: "simplyblock_core/services/main_distr_event_collector.py"}, + {name: "capacity-and-stats-collector", module: "simplyblock_core/services/capacity_and_stats_collector.py"}, + {name: "capacity-monitor", module: "simplyblock_core/services/cap_monitor.py"}, + {name: "health-check", module: "simplyblock_core/services/health_check_service.py"}, + {name: "device-monitor", module: "simplyblock_core/services/device_monitor.py"}, + {name: "lvol-monitor", module: "simplyblock_core/services/lvol_monitor.py"}, + {name: "snapshot-monitor", module: "simplyblock_core/services/snapshot_monitor.py"}, + } +} + +// taskServices are the runners that execute queued work. Their work is queued, +// which is what makes this component non-essential in the phase table: a runner +// at zero defers what is waiting rather than dropping it, and the queue is still +// there when it comes back. +func taskServices() []service { + portStart := []corev1.EnvVar{{Name: "LVOL_NVMF_PORT_START", Value: lvolNVMfPortStart}} + return []service{ + {name: "tasks-node-add-runner", module: "simplyblock_core/services/tasks_runner_node_add.py", extraEnv: portStart}, + {name: "tasks-runner-restart", module: "simplyblock_core/services/tasks_runner_restart.py"}, + {name: "tasks-runner-migration", module: "simplyblock_core/services/tasks_runner_migration.py"}, + {name: "tasks-runner-lvol-migration", module: "simplyblock_core/services/tasks_runner_lvol_migration.py"}, + {name: "tasks-runner-batch-migration", module: "simplyblock_core/services/tasks_runner_batch_migration.py"}, + {name: "tasks-runner-failed-migration", module: "simplyblock_core/services/tasks_runner_failed_migration.py"}, + {name: "tasks-runner-cluster-status", module: "simplyblock_core/services/tasks_cluster_status.py"}, + {name: "tasks-runner-new-device-migration", module: "simplyblock_core/services/tasks_runner_new_dev_migration.py"}, + {name: "tasks-runner-port-allow", module: "simplyblock_core/services/tasks_runner_port_allow.py"}, + {name: "tasks-runner-jc-comp-resume", module: "simplyblock_core/services/tasks_runner_jc_comp.py"}, + {name: "tasks-runner-sync-lvol-del", module: "simplyblock_core/services/tasks_runner_sync_lvol_del.py"}, + {name: "tasks-runner-cluster-expand", module: "simplyblock_core/services/tasks_runner_cluster_expand.py"}, + {name: "tasks-runner-node-removal", module: "simplyblock_core/services/tasks_runner_node_removal.py"}, + {name: "tasks-runner-snapshot-replication", module: "simplyblock_core/services/snapshot_replication.py"}, + {name: "tasks-runner-backup", module: "simplyblock_core/services/tasks_runner_backup.py"}, + {name: "tasks-runner-backup-merge", module: "simplyblock_core/services/tasks_runner_backup_merge.py"}, + {name: "tasks-runner-replication-final", module: "simplyblock_core/services/tasks_runner_replication_final.py"}, + } +} + +// monitoringDeployment and tasksDeployment are the same pod with a different +// service list. One replica each: both are pollers over shared state, and a +// second instance of either would do the same work twice. +func monitoringDeployment(cp *simplyblockv1alpha2.ControlPlane) *appsv1.Deployment { + return servicePoolDeployment(cp, ComponentMonitoring, monitoringServices()) +} + +func tasksDeployment(cp *simplyblockv1alpha2.ControlPlane) *appsv1.Deployment { + return servicePoolDeployment(cp, ComponentTasks, taskServices()) +} + +// servicePoolDeployment builds one pod holding a list of the control plane's +// long-running processes. +func servicePoolDeployment( + cp *simplyblockv1alpha2.ControlPlane, name string, services []service, +) *appsv1.Deployment { + managed := cp.Spec.Source.Local + labels := map[string]string{appLabel: name} + + spec := corev1.PodSpec{ + ServiceAccountName: serviceAccountName, + // The processes here reach the storage nodes' management addresses + // directly, which are host addresses rather than Service names. + HostNetwork: true, + DNSPolicy: corev1.DNSClusterFirstWithHostNet, + Containers: containers(services, localImage(cp), pullPolicyOf(managed)), + Volumes: []corev1.Volume{clusterFileVolumeSource()}, + } + scheduling(managed, &spec) + + return &appsv1.Deployment{ + ObjectMeta: metav1.ObjectMeta{ + Name: name, + Namespace: cp.Namespace, + Annotations: restartOnClusterFileChange, + }, + Spec: appsv1.DeploymentSpec{ + Replicas: ptr.To(int32(1)), + Selector: &metav1.LabelSelector{MatchLabels: labels}, + Template: corev1.PodTemplateSpec{ + ObjectMeta: metav1.ObjectMeta{ + Labels: labels, + Annotations: map[string]string{"log-collector/enabled": "true"}, + }, + Spec: spec, + }, + }, + } +} + +// adminControlDeployment runs nothing. It is a pod with the control plane's +// tooling, its cluster file, and its account, kept alive so that an +// administrator can exec into it, which is how the command-line surface is +// reached on a deployment that has no shell access to the control plane's hosts. +func adminControlDeployment(cp *simplyblockv1alpha2.ControlPlane) *appsv1.Deployment { + managed := cp.Spec.Source.Local + labels := map[string]string{appLabel: ComponentAdminControl} + + //nolint:prealloc // the literal is the declaration; the append below is the shared set + env := []corev1.EnvVar{ + {Name: "LVOL_NVMF_PORT_START", Value: lvolNVMfPortStart}, + namespaceEnv(), + logLevelEnv(), + } + env = append(env, prometheusEnv()...) + + spec := corev1.PodSpec{ + ServiceAccountName: serviceAccountName, + HostNetwork: true, + DNSPolicy: corev1.DNSClusterFirstWithHostNet, + Affinity: spreadAcrossHosts(ComponentAdminControl), + Containers: []corev1.Container{{ + Name: "simplyblock-control", + Image: localImage(cp), + ImagePullPolicy: pullPolicyOf(managed), + // Trapping the two signals is what makes the pod terminate promptly + // on a delete instead of waiting out its grace period: bash does not + // act on a signal while a foreground sleep is running. + Command: []string{"/bin/bash", "-c", "trap : TERM INT; sleep infinity & wait"}, + Env: env, + VolumeMounts: []corev1.VolumeMount{clusterFileMount()}, + Resources: corev1.ResourceRequirements{ + Requests: corev1.ResourceList{ + corev1.ResourceCPU: resource.MustParse("200m"), + corev1.ResourceMemory: resource.MustParse("256Mi"), + }, + Limits: corev1.ResourceList{ + corev1.ResourceCPU: resource.MustParse("600m"), + corev1.ResourceMemory: resource.MustParse("1Gi"), + }, + }, + }}, + Volumes: []corev1.Volume{clusterFileVolumeSource()}, + } + scheduling(managed, &spec) + + return &appsv1.Deployment{ + ObjectMeta: metav1.ObjectMeta{ + Name: ComponentAdminControl, + Namespace: cp.Namespace, + Annotations: restartOnClusterFileChange, + }, + Spec: appsv1.DeploymentSpec{ + Replicas: ptr.To(apiReplicas(managed)), + Strategy: appsv1.DeploymentStrategy{ + Type: appsv1.RollingUpdateDeploymentStrategyType, + RollingUpdate: &appsv1.RollingUpdateDeployment{ + MaxSurge: ptr.To(intstr.FromInt32(0)), + MaxUnavailable: ptr.To(intstr.FromInt32(1)), + }, + }, + Selector: &metav1.LabelSelector{MatchLabels: labels}, + Template: corev1.PodTemplateSpec{ + ObjectMeta: metav1.ObjectMeta{ + Labels: labels, + Annotations: map[string]string{"log-collector/enabled": "true"}, + }, + Spec: spec, + }, + }, + } +} + +// fdbExporterDeployment publishes the database's own view of itself as metrics, +// which is the half of a FoundationDB problem that pod counts cannot show: write +// latency, storage lag, and coordinator health. +func fdbExporterDeployment(cp *simplyblockv1alpha2.ControlPlane) *appsv1.Deployment { + const tmpVolume = "tmp" + labels := map[string]string{appLabel: ComponentFDBExporter} + + spec := corev1.PodSpec{ + SecurityContext: &corev1.PodSecurityContext{ + RunAsNonRoot: ptr.To(true), + RunAsUser: ptr.To(int64(4059)), + RunAsGroup: ptr.To(int64(4059)), + FSGroup: ptr.To(int64(4059)), + }, + Volumes: []corev1.Volume{ + {Name: tmpVolume, VolumeSource: corev1.VolumeSource{EmptyDir: &corev1.EmptyDirVolumeSource{}}}, + clusterFileVolumeSource(), + }, + Containers: []corev1.Container{{ + Name: "exporter", + Image: fdbExporterImage, + Env: []corev1.EnvVar{ + {Name: "FDB_CLUSTER_FILE", Value: clusterFilePath}, + }, + Ports: []corev1.ContainerPort{{Name: "metrics", ContainerPort: fdbExporterPort}}, + LivenessProbe: &corev1.Probe{ + ProbeHandler: corev1.ProbeHandler{ + HTTPGet: &corev1.HTTPGetAction{Path: "/metrics", Port: intstr.FromString("metrics")}, + }, + InitialDelaySeconds: 15, + PeriodSeconds: 30, + }, + ReadinessProbe: &corev1.Probe{ + ProbeHandler: corev1.ProbeHandler{ + HTTPGet: &corev1.HTTPGetAction{Path: "/metrics", Port: intstr.FromString("metrics")}, + }, + InitialDelaySeconds: 5, + PeriodSeconds: 10, + }, + SecurityContext: &corev1.SecurityContext{ + ReadOnlyRootFilesystem: ptr.To(true), + AllowPrivilegeEscalation: ptr.To(false), + Privileged: ptr.To(false), + }, + VolumeMounts: []corev1.VolumeMount{ + {Name: tmpVolume, MountPath: "/tmp"}, + { + Name: clusterFileVolume, + MountPath: clusterFilePath, + SubPath: clusterFileSubURL, + ReadOnly: true, + }, + }, + Resources: corev1.ResourceRequirements{ + Requests: corev1.ResourceList{ + corev1.ResourceCPU: resource.MustParse("50m"), + corev1.ResourceMemory: resource.MustParse("64Mi"), + }, + Limits: corev1.ResourceList{ + corev1.ResourceCPU: resource.MustParse("200m"), + corev1.ResourceMemory: resource.MustParse("128Mi"), + }, + }, + }}, + TerminationGracePeriodSeconds: ptr.To(int64(10)), + } + scheduling(cp.Spec.Source.Local, &spec) + + return &appsv1.Deployment{ + ObjectMeta: metav1.ObjectMeta{ + Name: ComponentFDBExporter, + Namespace: cp.Namespace, + Labels: labels, + Annotations: restartOnClusterFileChange, + }, + Spec: appsv1.DeploymentSpec{ + Replicas: ptr.To(int32(1)), + Selector: &metav1.LabelSelector{MatchLabels: labels}, + Template: corev1.PodTemplateSpec{ + ObjectMeta: metav1.ObjectMeta{Labels: labels}, + Spec: spec, + }, + }, + } +} + +func fdbExporterService(namespace string) *corev1.Service { + labels := map[string]string{appLabel: ComponentFDBExporter} + return &corev1.Service{ + ObjectMeta: metav1.ObjectMeta{ + Name: ComponentFDBExporter, + Namespace: namespace, + Labels: labels, + }, + Spec: corev1.ServiceSpec{ + Selector: labels, + Ports: []corev1.ServicePort{ + {Name: "metrics", Port: fdbExporterPort, TargetPort: intstr.FromString("metrics")}, + }, + }, + } +} + +// intstrFromInt names a numeric target port, which is what every Service here +// takes except the exporter's. That one names its container port, because the +// port is declared with a name and matching on it survives a renumber. +func intstrFromInt(port int32) intstr.IntOrString { + return intstr.FromInt32(port) +} diff --git a/operator/internal/controllers/controlplane/metrics.go b/operator/internal/controllers/controlplane/metrics.go new file mode 100644 index 000000000..c16031b89 --- /dev/null +++ b/operator/internal/controllers/controlplane/metrics.go @@ -0,0 +1,156 @@ +// The series the operator publishes about the control plane. +// +// simplyblock_controlplane_ready_state is the one to alert on first. Everything +// else in this API group is downstream of it, so an alert storm that starts here +// has one cause and one page. +// +// The component pair is a warning rather than a page, and the essential label is +// what lets one rule separate them: a component's ready count falling below its +// desired count while the ready state stays at one is the restart loop that has +// not become an outage yet. +// +// design-controlplane.md §9.2 is the specification. Two of its nine series are +// not here: simplyblock_controlplane_version_info waits on the /_meta/version +// read §8 records as missing, since a gauge labeled with an empty version says +// nothing, and the operation counters live beside the operation that increments +// them. + +package controlplane + +import ( + "strings" + + "github.com/prometheus/client_golang/prometheus" + ctrlmetrics "sigs.k8s.io/controller-runtime/pkg/metrics" +) + +// componentLabels are the labels every per-component series carries. The +// essential label is on the series rather than looked up by an alert, because an +// alerting rule cannot join against a table that lives in this binary. +var componentLabels = []string{"namespace", "component", "essential"} + +var ( + // controlPlaneReadyState is the probe's own verdict, which a Degraded + // control plane still passes. It is deliberately not the phase: the phase + // folds in the component counts, and an alert that fired on Degraded would + // page for a pod restarting behind a Service that is still answering. + controlPlaneReadyState = prometheus.NewGaugeVec( + prometheus.GaugeOpts{ + Name: "simplyblock_controlplane_ready_state", + Help: "1 while the control plane's readiness probe passes, which a Degraded control plane still does.", + }, + []string{"namespace"}, + ) + + // controlPlaneComponentReady and controlPlaneComponentDesired are the half + // of the phase the probe cannot see. Neither means anything alone: the + // ratio is the alert. + controlPlaneComponentReady = prometheus.NewGaugeVec( + prometheus.GaugeOpts{ + Name: "simplyblock_controlplane_component_ready_count", + Help: "How many replicas of one control-plane component are ready.", + }, + componentLabels, + ) + + controlPlaneComponentDesired = prometheus.NewGaugeVec( + prometheus.GaugeOpts{ + Name: "simplyblock_controlplane_component_desired_count", + Help: "How many replicas of one control-plane component there should be.", + }, + componentLabels, + ) + + // controlPlaneProbeDuration is latency, which degrades before readiness + // does: a control plane answering in four seconds is on its way to not + // answering, and the ready state cannot show that. + controlPlaneProbeDuration = prometheus.NewHistogramVec( + prometheus.HistogramOpts{ + Name: "simplyblock_controlplane_probe_duration_seconds", + Help: "How long one readiness probe took.", + Buckets: prometheus.DefBuckets, + }, + []string{"namespace"}, + ) + + // controlPlaneProbeFailures separates the ways a probe fails, because a + // transport error and a 503 have different causes: one is the network or + // the pod, the other is the control plane telling you what is wrong. + controlPlaneProbeFailures = prometheus.NewCounterVec( + prometheus.CounterOpts{ + Name: "simplyblock_controlplane_probe_failures_total", + Help: "Failed readiness probes, by reason.", + }, + []string{"namespace", "reason"}, + ) + + // controlPlaneInstallStepDuration is how long each installation step took, + // which is what says whether an install that felt slow was slow and where. + controlPlaneInstallStepDuration = prometheus.NewHistogramVec( + prometheus.HistogramOpts{ + Name: "simplyblock_controlplane_install_step_duration_seconds", + Help: "How long one step of the control plane's installation took.", + // An install step is minutes rather than milliseconds, which is + // past the default buckets. + Buckets: []float64{10, 30, 60, 120, 300, 600, 1200, 2400, 4800}, + }, + []string{"namespace", "step"}, + ) + + // controlPlaneOperationsTotal and controlPlaneOperationDuration are the two + // cumulative series every Ops kind in this group publishes. + controlPlaneOperationsTotal = prometheus.NewCounterVec( + prometheus.CounterOpts{ + Name: "simplyblock_controlplane_operations_total", + Help: "Control-plane operations that reached a terminal phase.", + }, + []string{"namespace", "action", "result"}, + ) + + controlPlaneOperationDuration = prometheus.NewHistogramVec( + prometheus.HistogramOpts{ + Name: "simplyblock_controlplane_operation_duration_seconds", + Help: "How long one control-plane operation took.", + Buckets: []float64{10, 30, 60, 120, 300, 600, 1200, 2400, 4800}, + }, + []string{"namespace", "action"}, + ) +) + +func init() { + ctrlmetrics.Registry.MustRegister( + controlPlaneReadyState, + controlPlaneComponentReady, + controlPlaneComponentDesired, + controlPlaneProbeDuration, + controlPlaneProbeFailures, + controlPlaneInstallStepDuration, + controlPlaneOperationsTotal, + controlPlaneOperationDuration, + ) +} + +// probeFailureReason classifies a failed probe into the small set the counter's +// label admits. It is derived from the message rather than from a typed error +// because the message is what the probe interface returns, and the alternative +// is an error taxonomy no caller distinguishes anywhere else. +func probeFailureReason(message string) string { + switch { + case containsAny(message, "context deadline exceeded", "Client.Timeout", "i/o timeout"): + return "timeout" + case strings.Contains(message, "status="): + return "status" + default: + return "transport" + } +} + +// containsAny reports whether s holds any of the needles. +func containsAny(s string, needles ...string) bool { + for _, needle := range needles { + if strings.Contains(s, needle) { + return true + } + } + return false +} diff --git a/operator/internal/controllers/controlplane/names.go b/operator/internal/controllers/controlplane/names.go new file mode 100644 index 000000000..1c70efd50 --- /dev/null +++ b/operator/internal/controllers/controlplane/names.go @@ -0,0 +1,138 @@ +// The names of every object the managed install applies, and the ports they +// answer on. +// +// They are constants rather than derivations from the ControlPlane's name, and +// that is deliberate: every one of them is the name the Helm chart already +// rendered, and a running deployment refers to them from places this operator +// does not control. The management API's Service name is in the CSI driver's +// configuration and in the Prometheus scrape configuration. The +// FoundationDBCluster's name is in the cluster file every workload mounts. The +// shared account's name is in the SB_K8S_ADMIN_SERVICE_ACCOUNTS list the control +// plane checks tokens against. Deriving them would rename all of it on the first +// install. +// +// The ControlPlane is a singleton named `simplyblock` (design-controlplane.md +// §3.1), so there is never a second set of these to collide with. + +package controlplane + +// SingletonName is the fixed name of the ControlPlane object. The controller +// ignores any other name, which is enforcement by convention rather than by the +// API server (design-controlplane.md §3.1). +const SingletonName = "simplyblock" + +// The management API and the services that share its image and its account. +const ( + // ComponentWebAPI is the management API: the endpoint every controller in + // this operator reaches, and the one component whose absence is an outage. + ComponentWebAPI = "simplyblock-webappapi" + + // ComponentTasks is the task runner. Its work is queued, so a task runner at + // zero defers what is waiting rather than dropping it. + ComponentTasks = "simplyblock-tasks" + + // ComponentMonitoring is the pool of per-subject monitors and collectors. + ComponentMonitoring = "simplyblock-monitoring" + + // ComponentAdminControl is the administrative surface: a long-lived pod with + // the control plane's tooling on it and no service of its own. + ComponentAdminControl = "simplyblock-admin-control" +) + +// The store and the FoundationDB half. +const ( + // ComponentMinio is the object store the control plane keeps long-term data + // in. + ComponentMinio = "simplyblock-minio" + + // ComponentFDBCluster is the FoundationDBCluster, whose readiness is that + // resource's own report rather than a replica count: a cluster at two of + // three coordinators is serving and a count cannot say so. + ComponentFDBCluster = "simplyblock-fdb-cluster" + + // ComponentFDBOperator is the FoundationDB operator's own workload, applied + // only where the Kubernetes cluster does not already run one. + ComponentFDBOperator = "simplyblock-fdb-controller-manager" + + // ComponentFDBExporter turns FoundationDB's status JSON into Prometheus + // metrics. + ComponentFDBExporter = "simplyblock-fdb-exporter" +) + +// The accounts, roles, and configuration the workloads above name. +const ( + // serviceAccountName is the account every workload built from the control + // plane's own image runs as. Its name appears in the control plane's + // SB_K8S_ADMIN_SERVICE_ACCOUNTS list, so it is not derived. + serviceAccountName = "simplyblock-sa" + + // clusterRoleName and clusterRoleBindingName grant that account what the + // control plane's Kubernetes-side work needs. + clusterRoleName = "simplyblock-role" + clusterRoleBindingName = "simplyblock-binding" + + // serviceReaderRoleName and serviceReaderBindingName let the namespace's + // default account resolve services and endpoints, which is how the control + // plane's own tooling finds its peers. + serviceReaderRoleName = "simplyblock-service-reader" + serviceReaderBindingName = "simplyblock-service-reader-binding" + + // configMapName holds the log level every workload reads through a + // configMapKeyRef, so changing it reaches all of them at once. + configMapName = "simplyblock-config" + + // objectStoreConfigName is the bucket configuration the object store's + // consumers read. + objectStoreConfigName = "simplyblock-objstore-config" + + // clusterFileConfigMapName is written by the FoundationDB operator rather + // than by this one, and mounted by every workload that talks to the + // database. The install neither creates nor reconciles it; it names it so + // the mounts can. + clusterFileConfigMapName = "simplyblock-fdb-cluster-config" + + // fdbOperatorServiceAccount and the roles beside it are the FoundationDB + // operator's own, applied with it. + fdbOperatorServiceAccount = "simplyblock-fdb-controller-manager" + fdbOperatorRoleName = "simplyblock-fdb-manager-role" + fdbOperatorRoleBinding = "simplyblock-fdb-manager-rolebinding" + fdbOperatorClusterRole = "simplyblock-fdb-manager-clusterrole" + fdbOperatorClusterBinding = "simplyblock-fdb-manager-clusterrolebinding" + + // fdbPodServiceAccount is what the FoundationDB pods themselves run as. The + // unified monitor in each pod writes locality annotations on its own pod to + // signal the operator, which the namespace's default account cannot do. + fdbPodServiceAccount = "simplyblock-fdb-cluster-pods" + fdbPodRoleName = "simplyblock-fdb-cluster-pods" + fdbPodRoleBinding = "simplyblock-fdb-cluster-pods" +) + +// The ports the install publishes. +const ( + // webAPIPort is where the management API listens, and the port + // status.endpoint resolves to. + webAPIPort = 5000 + + // minioAPIPort and minioConsolePort are the object store's two listeners. + minioAPIPort = 9000 + minioConsolePort = 9001 + + // fdbExporterPort is where the exporter serves /metrics. + fdbExporterPort = 9444 +) + +// The volumes and paths shared by every workload built from the control plane's +// image. +const ( + // clusterFileVolume mounts the FoundationDB cluster file at the path the + // client library reads by default, so nothing has to set FDB_CLUSTER_FILE. + clusterFileVolume = "fdb-cluster-file" + clusterFilePath = "/etc/foundationdb/fdb.cluster" + clusterFileKey = "cluster-file" + clusterFileSubURL = "fdb.cluster" +) + +// appLabel is the selector key every workload here is matched by. It is `app` +// rather than one of the recommended Kubernetes labels because that is what the +// running deployments carry, and a Deployment's selector is immutable. +const appLabel = "app" diff --git a/operator/internal/controllers/controlplane/podspec.go b/operator/internal/controllers/controlplane/podspec.go new file mode 100644 index 000000000..3bddeee42 --- /dev/null +++ b/operator/internal/controllers/controlplane/podspec.go @@ -0,0 +1,213 @@ +// The pieces every workload built from the control plane's own image shares. +// +// Three of the four workloads in the management API step are the same thing with +// a different entry point: the same image, the same account, the same cluster +// file mounted at the same path, the same log level read out of the same +// ConfigMap, and the same resource envelope. The monitoring pool and the task +// runner are that shape repeated ten and seventeen times inside one pod, which +// is why the services they run are a table here rather than a wall of container +// literals. +// +// The chart expressed the shared half as the `simplyblock.commonContainer` +// template, and this is the same set of decisions in Go. Where a value came from +// a chart value with no field on the ControlPlane spec, the chart's default is +// the constant below and the reason it is not configurable is stated beside it. + +package controlplane + +import ( + corev1 "k8s.io/api/core/v1" + "k8s.io/apimachinery/pkg/api/resource" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// The environment the control plane reads that has no field on the ControlPlane +// spec. +// +// design-controlplane.md's LocalControlPlane carries the image, the sizing, +// and the scheduling, and nothing else: the rest of what the chart took as +// values is either about the observability half this install does not apply or +// is a constant the control plane and the operator have to agree on. These are +// the second kind, so they are constants rather than fields nobody would set +// differently. +const ( + // lvolNVMfPortStart is the first port of the range logical volumes are + // published on. The storage nodes and the control plane have to agree on it, + // and the operator's own workloads are built from the same number. + lvolNVMfPortStart = "9110" + + // prometheusHost and prometheusPort are where the control plane pushes the + // metrics it collects. The chart deploys that Prometheus as a subchart and + // still does, so the install names the same Service. + prometheusHost = "simplyblock-prometheus" + prometheusPort = "9090" + + // logLevelKey is the key in the shared ConfigMap every workload reads its + // log level from. + logLevelKey = "LOG_LEVEL" +) + +// serviceResources is the envelope every service container runs in. It is the +// chart's commonContainer block: small requests, because a monitoring pod runs +// ten of these and a task pod seventeen, and a limit high enough that one of +// them doing real work is not throttled. +func serviceResources() corev1.ResourceRequirements { + return corev1.ResourceRequirements{ + Requests: corev1.ResourceList{ + corev1.ResourceCPU: resource.MustParse("50m"), + corev1.ResourceMemory: resource.MustParse("100Mi"), + }, + Limits: corev1.ResourceList{ + corev1.ResourceCPU: resource.MustParse("300m"), + corev1.ResourceMemory: resource.MustParse("1Gi"), + }, + } +} + +// logLevelEnv reads the shared log level, which is what makes changing it one +// ConfigMap edit rather than a rebuild of every workload. +func logLevelEnv() corev1.EnvVar { + return corev1.EnvVar{ + Name: "SIMPLYBLOCK_LOG_LEVEL", + ValueFrom: &corev1.EnvVarSource{ + ConfigMapKeyRef: &corev1.ConfigMapKeySelector{ + LocalObjectReference: corev1.LocalObjectReference{Name: configMapName}, + Key: logLevelKey, + }, + }, + } +} + +// prometheusEnv is where a container pushes what it measures. Every service +// container carries it, including the ones that measure nothing, because the +// control plane's shared code reads it at import time. +func prometheusEnv() []corev1.EnvVar { + return []corev1.EnvVar{ + {Name: "PROMETHEUS_URL", Value: prometheusHost}, + {Name: "PROMETHEUS_PORT", Value: prometheusPort}, + } +} + +// namespaceEnv tells a container which namespace it is in, which is how the +// control plane addresses the Kubernetes objects it manages. +func namespaceEnv() corev1.EnvVar { + return corev1.EnvVar{ + Name: "K8S_NAMESPACE", + ValueFrom: &corev1.EnvVarSource{ + FieldRef: &corev1.ObjectFieldSelector{FieldPath: "metadata.namespace"}, + }, + } +} + +// clusterFileMount is how a container reaches FoundationDB. The ConfigMap behind +// it is written by the FoundationDB operator and re-written when the +// coordinators change, which is why the workloads that mount it are annotated +// for restart on its change. +func clusterFileMount() corev1.VolumeMount { + return corev1.VolumeMount{ + Name: clusterFileVolume, + MountPath: clusterFilePath, + SubPath: clusterFileSubURL, + } +} + +// clusterFileVolumeSource projects the single key of that ConfigMap to the file +// name the client library expects. +func clusterFileVolumeSource() corev1.Volume { + return corev1.Volume{ + Name: clusterFileVolume, + VolumeSource: corev1.VolumeSource{ + ConfigMap: &corev1.ConfigMapVolumeSource{ + LocalObjectReference: corev1.LocalObjectReference{Name: clusterFileConfigMapName}, + Items: []corev1.KeyToPath{ + {Key: clusterFileKey, Path: clusterFileSubURL}, + }, + }, + }, + } +} + +// service is one long-running control-plane process: a name, the module that is +// its entry point, and whatever environment it needs beyond the shared set. +type service struct { + // name is the container's name, and what a `kubectl logs -c` names. + name string + + // module is the Python module run as the container's command, relative to + // the image's working directory. + module string + + // extraEnv is what this one service needs that the shared set does not + // carry. Most of them need nothing. + extraEnv []corev1.EnvVar +} + +// container builds one service container from the shared shape. +func (s service) container(image string, pullPolicy corev1.PullPolicy) corev1.Container { + env := append([]corev1.EnvVar{}, s.extraEnv...) + env = append(env, prometheusEnv()...) + env = append(env, logLevelEnv()) + + return corev1.Container{ + Name: s.name, + Image: image, + ImagePullPolicy: pullPolicy, + Command: []string{"python3", s.module}, + Env: env, + VolumeMounts: []corev1.VolumeMount{clusterFileMount()}, + Resources: serviceResources(), + } +} + +// containers builds every service of a pool, in the order they are declared, so +// that the apply produces a stable list rather than one that reorders between +// passes and rolls the Deployment for nothing. +func containers(services []service, image string, pullPolicy corev1.PullPolicy) []corev1.Container { + out := make([]corev1.Container, 0, len(services)) + for _, s := range services { + out = append(out, s.container(image, pullPolicy)) + } + return out +} + +// scheduling is what every pod the install creates carries from the spec: where +// it may run and what it tolerates. It is one function so that a pod added later +// cannot silently miss the fields an operator set. +func scheduling(local *simplyblockv1alpha2.LocalControlPlane, spec *corev1.PodSpec) { + if local == nil { + return + } + if len(local.NodeSelector) > 0 { + spec.NodeSelector = local.NodeSelector + } + if len(local.Tolerations) > 0 { + spec.Tolerations = local.Tolerations + } +} + +// spreadAcrossHosts keeps the replicas of a workload on different machines. It +// is required rather than preferred for the management API and the admin +// surface, which is what makes a second replica worth having: two instances on +// one machine survive a process crash and not the machine. +func spreadAcrossHosts(app string) *corev1.Affinity { + return &corev1.Affinity{ + PodAntiAffinity: &corev1.PodAntiAffinity{ + RequiredDuringSchedulingIgnoredDuringExecution: []corev1.PodAffinityTerm{{ + LabelSelector: &metav1.LabelSelector{MatchLabels: map[string]string{appLabel: app}}, + TopologyKey: "kubernetes.io/hostname", + }}, + }, + } +} + +// pullPolicyOf is the spec's pull policy, or the default the API declares. A +// zero value reaches this only from an object written before the default landed +// or built in a test, and IfNotPresent is what the marker says. +func pullPolicyOf(local *simplyblockv1alpha2.LocalControlPlane) corev1.PullPolicy { + if local == nil || local.ImagePullPolicy == "" { + return corev1.PullIfNotPresent + } + return local.ImagePullPolicy +} diff --git a/operator/internal/controllers/controlplane/probe.go b/operator/internal/controllers/controlplane/probe.go new file mode 100644 index 000000000..707554d93 --- /dev/null +++ b/operator/internal/controllers/controlplane/probe.go @@ -0,0 +1,164 @@ +// The readiness probe, and the endpoint it is aimed at. +// +// The probe is a direct GET rather than a value taken off the event stream, +// because it is the check that the stream itself can be established +// (design-crd-model.md §7.7 makes the stream the way state arrives, and this is +// the one read that cannot depend on it). +// +// It is an interface rather than a function so that the reconciler's branches +// are unit-testable without an HTTP server: a control plane that answers, one +// that does not, and one whose version disagrees with what an upgrade asked for. +// The live implementation is one http.Client and two paths. + +package controlplane + +import ( + "context" + "fmt" + "io" + "net/http" + "strings" + "time" +) + +// The two reads design-controlplane.md §8 names. +const ( + readyPath = "/api/v2/_meta/ready" + versionPath = "/api/v2/_meta/version" +) + +// probeTimeout bounds one readiness check. It is short on purpose: the probe +// runs on every reconcile of every control plane, and a control plane that takes +// longer than this to say it is ready is one whose latency is already the +// problem. +const probeTimeout = 10 * time.Second + +// Prober answers the two questions the reconciler asks of a running control +// plane. +type Prober interface { + // Ready reports whether the control plane answers its readiness endpoint. A + // false comes with the control plane's own words rather than a paraphrase of + // them, because that string is what lands in status.message and in the event + // an administrator reads. + Ready(ctx context.Context, endpoint string) (ok bool, reason string) + + // Version reports what the management API says it is. An empty version with + // a nil error is an endpoint that does not serve the read, which is the + // state every control plane is in until it ships: design-controlplane.md §8 + // records that /_meta/version does not exist yet, so the caller publishes + // nothing rather than treating its absence as a failure. + Version(ctx context.Context, endpoint string) (string, error) +} + +// HTTPProber probes over HTTP with a bearer token. +type HTTPProber struct { + // Client is the transport, which carries the TLS configuration where the + // deployment has one. A nil client is one with the probe's own timeout. + Client *http.Client + + // Token is the bearer the operator authenticates with. Empty is an + // unauthenticated request, which is what the readiness endpoint takes on a + // deployment that does not require a token for it. + Token string +} + +// Ready performs the readiness read. +func (p *HTTPProber) Ready(ctx context.Context, endpoint string) (bool, string) { + body, status, err := p.get(ctx, endpoint, readyPath) + switch { + case err != nil: + return false, err.Error() + case status >= 300: + // The control plane's own body rather than a sentence about it: a 503 + // whose body names the FoundationDB error is the whole of what an + // administrator needs, and wrapping it loses that. + if trimmed := strings.TrimSpace(string(body)); trimmed != "" { + return false, fmt.Sprintf("status=%d: %s", status, trimmed) + } + return false, fmt.Sprintf("status=%d", status) + default: + return true, "" + } +} + +// Version performs the version read. +// +// A 404 is an endpoint that does not serve it, which is reported as no version +// rather than as an error: the read is a prerequisite this repository is waiting +// on, and until it lands every deployment would otherwise report a failure on +// every pass. +func (p *HTTPProber) Version(ctx context.Context, endpoint string) (string, error) { + body, status, err := p.get(ctx, endpoint, versionPath) + switch { + case err != nil: + return "", err + case status == http.StatusNotFound: + return "", nil + case status >= 300: + return "", fmt.Errorf("status=%d: %s", status, strings.TrimSpace(string(body))) + } + return parseVersion(body), nil +} + +// get performs one request against a path of the endpoint. +func (p *HTTPProber) get(ctx context.Context, endpoint, path string) ([]byte, int, error) { + client := p.Client + if client == nil { + client = &http.Client{Timeout: probeTimeout} + } + + ctx, cancel := context.WithTimeout(ctx, probeTimeout) + defer cancel() + + url := strings.TrimSuffix(endpoint, "/") + path + req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil) + if err != nil { + return nil, 0, err + } + if p.Token != "" { + req.Header.Set("Authorization", "Bearer "+p.Token) + } + + resp, err := client.Do(req) + if err != nil { + return nil, 0, err + } + defer func() { _ = resp.Body.Close() }() + + // The body is read to a bound rather than in full: it lands in + // status.message, which is one sentence, and a control plane returning a + // stack trace should not put it on the object. + body, err := io.ReadAll(io.LimitReader(resp.Body, 4096)) + if err != nil { + return nil, resp.StatusCode, err + } + return body, resp.StatusCode, nil +} + +// parseVersion pulls the version out of whatever the endpoint returns. +// +// The read does not exist yet (design-controlplane.md §8), so its response shape +// is not settled either. Two shapes are handled, a bare string and a JSON object +// with a version field, and anything else reads as no version, which publishes +// nothing rather than publishing noise. +func parseVersion(body []byte) string { + trimmed := strings.TrimSpace(string(body)) + if trimmed == "" { + return "" + } + if !strings.HasPrefix(trimmed, "{") { + return strings.Trim(trimmed, `"`) + } + for _, key := range []string{`"version"`, `"result"`} { + if i := strings.Index(trimmed, key); i >= 0 { + rest := trimmed[i+len(key):] + if j := strings.Index(rest, `"`); j >= 0 { + rest = rest[j+1:] + if k := strings.Index(rest, `"`); k >= 0 { + return rest[:k] + } + } + } + } + return "" +} diff --git a/operator/internal/controllers/controlplane/resolver.go b/operator/internal/controllers/controlplane/resolver.go new file mode 100644 index 000000000..2693fe636 --- /dev/null +++ b/operator/internal/controllers/controlplane/resolver.go @@ -0,0 +1,78 @@ +// Where the rest of the operator reads the control plane's address from. +// +// design-controlplane.md §3.3 makes status.endpoint the one answer to where the +// control plane is, so that naming a remote one is a field rather than the +// SIMPLYBLOCK_WEBAPI_BASE_URL environment variable, and so that a change to it +// reaches every reader without a Deployment rollout. This is what the readers +// call. +// +// It resolves rather than injects, because the endpoint is not known when the +// manager builds its controllers: the ControlPlane has not been reconciled then, +// and a remote one may not have been created at all. Every caller therefore +// asks per request and gets the current answer. +// +// An empty string means the object says nothing yet, and the caller keeps +// whatever default it already had. That is what makes adopting this additive: a +// deployment whose ControlPlane has not published an endpoint behaves exactly as +// it did before. + +package controlplane + +import ( + "context" + "sync" + "time" + + "sigs.k8s.io/controller-runtime/pkg/client" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// endpointCacheTTL is how long a resolved endpoint is reused. +// +// The read is served from the manager's cache, so it is cheap, but it happens on +// every control-plane call the operator makes and those are frequent. A second +// is short enough that an endpoint change reaches every caller within one, and +// long enough that a burst of calls resolves once. +const endpointCacheTTL = time.Second + +// EndpointResolver answers where the control plane is. An empty string means the +// ControlPlane has published no endpoint, and the caller keeps its own default. +type EndpointResolver func(ctx context.Context) string + +// NewEndpointResolver returns a resolver reading the ControlPlane singleton in +// the operator's namespace. +// +// It never returns an error. A control plane that cannot be read is one whose +// endpoint is not known, which is the same situation as one that has not +// published it, and both mean the caller keeps its default. Every +// control-plane call in the operator resolves through this, so an unreadable +// singleton is a transient miss rather than a failed call. +func NewEndpointResolver(reader client.Reader, namespace string) EndpointResolver { + var ( + mu sync.Mutex + cached string + cachedAt time.Time + ) + + return func(ctx context.Context) string { + mu.Lock() + defer mu.Unlock() + + if !cachedAt.IsZero() && time.Since(cachedAt) < endpointCacheTTL { + return cached + } + + var cp simplyblockv1alpha2.ControlPlane + key := client.ObjectKey{Namespace: namespace, Name: SingletonName} + if err := reader.Get(ctx, key, &cp); err != nil { + // Not cached: an unreadable singleton is a transient state, and + // caching the empty answer would hold every caller on its default + // for the rest of the interval. + return "" + } + + cached, cachedAt = cp.Status.Endpoint, time.Now() + return cached + } +} diff --git a/operator/internal/controllers/controlplane/resolver_test.go b/operator/internal/controllers/controlplane/resolver_test.go new file mode 100644 index 000000000..221f19ab4 --- /dev/null +++ b/operator/internal/controllers/controlplane/resolver_test.go @@ -0,0 +1,104 @@ +// The endpoint resolver: what the rest of the operator reads the control +// plane's address from. +// +// The behavior that matters is the empty answer. A resolver that returned a +// wrong address on a control plane it could not read would point every caller at +// nothing; returning empty leaves each caller on the default it already had, +// which is what makes adopting this additive. + +package controlplane + +import ( + "context" + "testing" + + "sigs.k8s.io/controller-runtime/pkg/client" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// A published endpoint is what callers get. This is the whole point of the +// field: a remote control plane is somewhere the environment variable does +// not name. +func TestTheResolverAnswersWithThePublishedEndpoint(t *testing.T) { + const endpoint = "https://sb-control.example.com:5000" + + cp := managedControlPlane(endpoint) + cp.Status.Endpoint = endpoint + + resolve := NewEndpointResolver(newClient(t, cp), testNamespace) + + if got := resolve(context.Background()); got != endpoint { + t.Errorf("resolved %q, want %q", got, endpoint) + } +} + +// A control plane that has published nothing yet resolves to nothing, and the +// caller keeps its own default rather than being pointed at an empty address. +func TestAControlPlaneWithNoEndpointResolvesToNothing(t *testing.T) { + cp := localControlPlane() + + resolve := NewEndpointResolver(newClient(t, cp), testNamespace) + + if got := resolve(context.Background()); got != "" { + t.Errorf("resolved %q, want nothing published", got) + } +} + +// A singleton that cannot be read resolves to nothing rather than to an error. +// Every control-plane call in the operator goes through this, so an unreadable +// object is a transient miss rather than a failed call. +func TestAnAbsentControlPlaneResolvesToNothing(t *testing.T) { + resolve := NewEndpointResolver(newClient(t), testNamespace) + + if got := resolve(context.Background()); got != "" { + t.Errorf("resolved %q against a cluster with no ControlPlane", got) + } +} + +// Only the singleton answers, because a ControlPlane under another name is one +// the reconciler ignores. +func TestOnlyTheSingletonAnswers(t *testing.T) { + other := localControlPlane() + other.Name = "a-second-one" + other.Status.Endpoint = "https://not-the-singleton.example.com:5000" + + resolve := NewEndpointResolver(newClient(t, other), testNamespace) + + if got := resolve(context.Background()); got != "" { + t.Errorf("resolved %q from an object the reconciler ignores", got) + } +} + +// A change to the published endpoint reaches callers, which is the property that +// makes this a resolver rather than a value injected at startup. +func TestAChangedEndpointReachesTheNextCaller(t *testing.T) { + ctx := context.Background() + + cp := managedControlPlane("https://first.example.com:5000") + cp.Status.Endpoint = "https://first.example.com:5000" + c := newClient(t, cp) + + resolve := NewEndpointResolver(c, testNamespace) + if got := resolve(ctx); got != "https://first.example.com:5000" { + t.Fatalf("resolved %q before the change", got) + } + + var current simplyblockv1alpha2.ControlPlane + key := client.ObjectKey{Name: SingletonName, Namespace: testNamespace} + if err := c.Get(ctx, key, ¤t); err != nil { + t.Fatalf("read the control plane: %v", err) + } + current.Status.Endpoint = "https://second.example.com:5000" + if err := c.Status().Update(ctx, ¤t); err != nil { + t.Fatalf("publish the new endpoint: %v", err) + } + + // The resolver caches for a short interval, so a fresh one stands in for the + // interval passing. What is under test is that the answer comes from the + // object rather than from anything captured at construction. + resolve = NewEndpointResolver(c, testNamespace) + if got := resolve(ctx); got != "https://second.example.com:5000" { + t.Errorf("resolved %q, want the endpoint the object now publishes", got) + } +} diff --git a/operator/internal/controllers/controlplane/spec.go b/operator/internal/controllers/controlplane/spec.go new file mode 100644 index 000000000..a847dba00 --- /dev/null +++ b/operator/internal/controllers/controlplane/spec.go @@ -0,0 +1,56 @@ +// The reads of ControlPlane.spec that more than one file needs. +// +// They exist because spec.source is two optional members of which exactly one is +// set, so every read of anything under it is a nil check the API server's CEL +// rule has already made. Doing it once here means a builder reads the image +// rather than the source, and a controller asks whether the control plane is +// managed rather than unpacking a pointer. + +package controlplane + +import ( + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// localImage is the control plane's own image, empty when the object names an +// remote control plane. The workload builders call it rather than reaching +// through the source themselves. +func localImage(cp *simplyblockv1alpha2.ControlPlane) string { + if local := cp.Spec.Source.Local; local != nil { + return local.Image + } + return "" +} + +// foundationDBSpecOf is the sizing block, which is optional inside an optional +// member. A nil return is the API's defaults rather than an error. +func foundationDBSpecOf(cp *simplyblockv1alpha2.ControlPlane) *simplyblockv1alpha2.FoundationDBSpec { + if local := cp.Spec.Source.Local; local != nil { + return local.FoundationDB + } + return nil +} + +// apiReplicas is how many management API instances to run. Two is the default +// and the number the phases assume: a single instance makes Degraded unreachable +// for this component and turns every restart into an outage. +func apiReplicas(local *simplyblockv1alpha2.LocalControlPlane) int32 { + if local == nil || local.Replicas == nil { + return 2 + } + return *local.Replicas +} + +// isLocal reports whether the operator installs this control plane. It is the +// branch every path in the reconciler takes first, and the one the operations +// are refused on. +func isLocal(cp *simplyblockv1alpha2.ControlPlane) bool { + return cp.Spec.Source.Local != nil +} + +// isManaged reports whether the control plane already exists somewhere the +// operator does not own. A source with neither member set is neither, which is +// what the reconciler reports rather than guessing at. +func isManaged(cp *simplyblockv1alpha2.ControlPlane) bool { + return cp.Spec.Source.Managed != nil +} diff --git a/operator/internal/controllers/controlplane/suite_test.go b/operator/internal/controllers/controlplane/suite_test.go new file mode 100644 index 000000000..de3122a8a --- /dev/null +++ b/operator/internal/controllers/controlplane/suite_test.go @@ -0,0 +1,93 @@ +// What this package's envtest-backed tests need to find a real apiserver, and +// nothing else. +// +// One test here cannot run against a fake client: the CRD's CEL rules are +// evaluated by the apiserver. Everything else in this package is driven with a +// fake client and a stubbed prober, which is where the phase and the machine are +// provable and where the bulk of the coverage is. + +package controlplane + +import ( + "os" + "path/filepath" + "sync" + "testing" + + "k8s.io/client-go/kubernetes/scheme" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/envtest" + logf "sigs.k8s.io/controller-runtime/pkg/log" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +var ( + sharedEnvOnce sync.Once + sharedEnv *envtest.Environment + sharedEnvClient client.Client + sharedEnvErr error +) + +// TestMain stops the shared apiserver once, after every test that used it. +func TestMain(m *testing.M) { + code := m.Run() + if sharedEnv != nil { + _ = sharedEnv.Stop() + } + os.Exit(code) +} + +// apiServer returns a client against a real apiserver with this repository's +// CRDs installed, starting one on first use. +// +// Only v1alpha2 is exercised through it. It is the storage version, so reading +// it needs no conversion, and the CRDs' conversion webhook is one envtest does +// not run. +func apiServer(t *testing.T) client.Client { + t.Helper() + sharedEnvOnce.Do(func() { + if err := simplyblockv1alpha2.AddToScheme(scheme.Scheme); err != nil { + sharedEnvErr = err + return + } + sharedEnv = &envtest.Environment{ + CRDDirectoryPaths: []string{ + filepath.Join("..", "..", "..", "config", "crd", "bases"), + }, + ErrorIfCRDPathMissing: true, + BinaryAssetsDirectory: getFirstFoundEnvTestBinaryDir(), + } + cfg, err := sharedEnv.Start() + if err != nil { + sharedEnvErr = err + return + } + sharedEnvClient, sharedEnvErr = client.New(cfg, client.Options{Scheme: scheme.Scheme}) + }) + if sharedEnvErr != nil { + t.Fatalf("starting the test apiserver: %v", sharedEnvErr) + } + return sharedEnvClient +} + +// getFirstFoundEnvTestBinaryDir locates the envtest asset binaries. +// +// controller-runtime normally passes them through KUBEBUILDER_ASSETS, which the +// Makefile sets. This is what makes the same test runnable straight from an +// editor, and it reads the shared repository-root .bin that +// `make setup-envtest` populates. +func getFirstFoundEnvTestBinaryDir() string { + basePath := filepath.Join("..", "..", "..", "..", ".bin", "k8s") + entries, err := os.ReadDir(basePath) + if err != nil { + logf.Log.Error(err, "the envtest assets could not be read", "path", basePath) + return "" + } + for _, entry := range entries { + if entry.IsDir() { + return filepath.Join(basePath, entry.Name()) + } + } + return "" +} diff --git a/operator/internal/controllers/controlplane/workloads_test.go b/operator/internal/controllers/controlplane/workloads_test.go new file mode 100644 index 000000000..b4b2cdc51 --- /dev/null +++ b/operator/internal/controllers/controlplane/workloads_test.go @@ -0,0 +1,281 @@ +// What each step builds, and the properties of it that other parts of this +// design rest on. +// +// The management API's second replica is the one worth naming. Three things +// depend on it: Degraded is defined as a component below its desired count while +// the probe still passes, a Restart recycles the workload after draining rather +// than instead of serving, and an Upgrade rolls this Deployment. With one +// instance each of those becomes an outage. + +package controlplane + +import ( + "strings" + "testing" + + appsv1 "k8s.io/api/apps/v1" + corev1 "k8s.io/api/core/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + + "github.com/simplyblock/atlas/ptr" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// Every component the phase table watches is a workload some step applies. +// A component in the table that nothing installs is one whose counts are always +// zero, which for an essential one would hold every controller in the operator +// on a workload nobody asked for. +func TestEveryWatchedComponentIsOneTheInstallApplies(t *testing.T) { + cp := localControlPlane() + applied := map[string]bool{} + for _, obj := range append(foundationDBObjects(cp), + append(datastoreObjects(cp), managementAPIObjects(cp)...)...) { + applied[obj.GetName()] = true + } + + for _, comp := range componentTable { + if !applied[comp.name] { + t.Errorf("%s is watched by the phase table and no step applies it", comp.name) + } + } +} + +// The management API runs two instances by default, spread across machines and +// rolled one at a time with no surge. That combination is what lets a pod be +// replaced while the control plane keeps answering. +func TestTheManagementAPIRunsTwoInstancesSpreadAcrossHosts(t *testing.T) { + cp := localControlPlane() + + api := findDeployment(t, managementAPIObjects(cp), ComponentWebAPI) + + if api.Spec.Replicas == nil || *api.Spec.Replicas != 2 { + t.Errorf("replicas = %v, want 2: a single instance makes every restart an outage", + api.Spec.Replicas) + } + + affinity := api.Spec.Template.Spec.Affinity + if affinity == nil || affinity.PodAntiAffinity == nil || + len(affinity.PodAntiAffinity.RequiredDuringSchedulingIgnoredDuringExecution) == 0 { + t.Fatal("no required anti-affinity: two instances on one machine survive a process " + + "crash and not the machine") + } + + rolling := api.Spec.Strategy.RollingUpdate + if rolling == nil { + t.Fatal("no rolling-update strategy") + } + if rolling.MaxUnavailable == nil || rolling.MaxUnavailable.IntValue() != 1 { + t.Errorf("maxUnavailable = %v, want 1", rolling.MaxUnavailable) + } + if rolling.MaxSurge == nil || rolling.MaxSurge.IntValue() != 0 { + t.Errorf("maxSurge = %v, want 0", rolling.MaxSurge) + } +} + +// One instance stays expressible, because an edge deployment may prefer it. What +// it costs is stated in the design rather than prevented here, and this pins that +// the spec is honored. +func TestASingleManagementAPIInstanceStaysExpressible(t *testing.T) { + cp := localControlPlane() + cp.Spec.Source.Local.Replicas = ptr.To(int32(1)) + + api := findDeployment(t, managementAPIObjects(cp), ComponentWebAPI) + + if api.Spec.Replicas == nil || *api.Spec.Replicas != 1 { + t.Errorf("replicas = %v, want the spec's 1", api.Spec.Replicas) + } +} + +// Every workload built from the control plane's own image runs it, so an upgrade +// that writes one image onto the entity moves all of them. +func TestEveryWorkloadOfTheControlPlaneRunsTheSpecsImage(t *testing.T) { + cp := localControlPlane() + + for _, name := range []string{ + ComponentWebAPI, ComponentTasks, ComponentMonitoring, ComponentAdminControl, + } { + d := findDeployment(t, managementAPIObjects(cp), name) + for _, container := range d.Spec.Template.Spec.Containers { + if container.Image != testImage { + t.Errorf("%s/%s runs %q, want the spec's %q", + name, container.Name, container.Image, testImage) + } + } + } +} + +// The exporter is deliberately not on that image: it is upstream's, and its +// version is independent of the control plane's. +func TestTheExporterDoesNotRunTheControlPlanesImage(t *testing.T) { + cp := localControlPlane() + + exporter := findDeployment(t, managementAPIObjects(cp), ComponentFDBExporter) + + for _, container := range exporter.Spec.Template.Spec.Containers { + if container.Image == testImage { + t.Errorf("%s runs the control plane's image, and it is not that program", + container.Name) + } + } +} + +// Every service container reaches the database through the cluster file, and +// every pod that mounts it is annotated for restart when it changes. A +// coordinator change rewrites that ConfigMap, and a pod holding a connection to +// the old coordinator has to be recycled to notice. +func TestEveryWorkloadThatReachesTheDatabaseIsRolledWhenItMoves(t *testing.T) { + cp := localControlPlane() + + for _, name := range []string{ + ComponentWebAPI, ComponentTasks, ComponentMonitoring, ComponentAdminControl, + ComponentFDBExporter, + } { + d := findDeployment(t, managementAPIObjects(cp), name) + + mountsClusterFile := false + for _, container := range d.Spec.Template.Spec.Containers { + for _, mount := range container.VolumeMounts { + if mount.MountPath == clusterFilePath { + mountsClusterFile = true + } + } + } + if !mountsClusterFile { + t.Errorf("%s mounts no cluster file, so it cannot reach the database", name) + continue + } + if d.Annotations["reloader.stakater.com/configmap"] != clusterFileConfigMapName { + t.Errorf("%s is not rolled when the cluster file changes, so a coordinator move "+ + "leaves it connected to a coordinator that is gone", name) + } + } +} + +// The monitoring pool and the task runner are the same pod with a different +// service list, and both lists are non-empty. A pool built with no services is a +// pod that starts and does nothing, which no count would notice. +func TestTheServicePoolsRunWhatTheyDeclare(t *testing.T) { + cp := localControlPlane() + + for _, tc := range []struct { + name string + services []service + }{ + {ComponentMonitoring, monitoringServices()}, + {ComponentTasks, taskServices()}, + } { + if len(tc.services) == 0 { + t.Fatalf("%s declares no services", tc.name) + } + + d := findDeployment(t, managementAPIObjects(cp), tc.name) + if len(d.Spec.Template.Spec.Containers) != len(tc.services) { + t.Errorf("%s runs %d containers against %d declared services", + tc.name, len(d.Spec.Template.Spec.Containers), len(tc.services)) + } + + seen := map[string]bool{} + for _, container := range d.Spec.Template.Spec.Containers { + if seen[container.Name] { + t.Errorf("%s declares %q twice, which Kubernetes rejects", tc.name, container.Name) + } + seen[container.Name] = true + + if len(container.Command) != 2 || container.Command[0] != "python3" { + t.Errorf("%s/%s runs %v, want a python3 module", tc.name, container.Name, + container.Command) + } + if !strings.HasSuffix(container.Command[1], ".py") { + t.Errorf("%s/%s runs %q, which is not a module", tc.name, container.Name, + container.Command[1]) + } + } + } +} + +// The control plane's account is granted exec on pods, which is the strongest +// thing in its role and the one an audit has to be able to find. Losing it would +// stop the control plane driving the storage nodes' processes, which is not a +// failure any count reports. +func TestTheControlPlanesAccountKeepsTheGrantsItCannotWorkWithout(t *testing.T) { + cp := localControlPlane() + + role := findClusterRole(t, managementAPIObjects(cp), clusterRoleName) + + want := map[string]bool{"pods/exec": false, "tokenreviews": false, "nodes": false} + for _, rule := range role.Rules { + for _, resource := range rule.Resources { + if _, tracked := want[resource]; tracked { + want[resource] = true + } + } + } + for resource, granted := range want { + if !granted { + t.Errorf("the control plane's role does not grant %s", resource) + } + } +} + +// The object store's volume comes from the same class as the database's. It +// cannot be a class this operator provides, because the control plane has to +// exist before any simplyblock volume can. +func TestTheObjectStoreTakesTheSameStorageClassAsTheDatabase(t *testing.T) { + cp := localControlPlane() + cp.Spec.Source.Local.FoundationDB = &simplyblockv1alpha2.FoundationDBSpec{ + StorageClassName: "fast-local", + } + + store := minioStatefulSet(cp) + if len(store.Spec.VolumeClaimTemplates) != 1 { + t.Fatalf("the object store claims %d volumes, want 1", + len(store.Spec.VolumeClaimTemplates)) + } + claim := store.Spec.VolumeClaimTemplates[0] + if claim.Spec.StorageClassName == nil || *claim.Spec.StorageClassName != "fast-local" { + t.Errorf("storageClassName = %v, want the database's fast-local", + claim.Spec.StorageClassName) + } +} + +// Nothing the install applies names a Secret that the install does not also +// create, because a workload waiting on a Secret nobody writes never starts and +// reports the wait as a pod event rather than on the ControlPlane. +func TestNoWorkloadWaitsOnASecretTheInstallDoesNotCreate(t *testing.T) { + cp := localControlPlane() + + for _, obj := range append(foundationDBObjects(cp), + append(datastoreObjects(cp), managementAPIObjects(cp)...)...) { + spec := podSpecOf(obj) + if spec == nil { + continue + } + for _, volume := range spec.Volumes { + if volume.Secret != nil { + t.Errorf("%s mounts Secret %q, which no step of the install creates", + obj.GetName(), volume.Secret.SecretName) + } + } + for _, container := range spec.Containers { + for _, env := range container.Env { + if env.ValueFrom != nil && env.ValueFrom.SecretKeyRef != nil { + t.Errorf("%s/%s reads Secret %q, which no step of the install creates", + obj.GetName(), container.Name, env.ValueFrom.SecretKeyRef.Name) + } + } + } + } +} + +// podSpecOf is the pod template of whatever workload kind an object is, or nil +// for the objects that carry none. +func podSpecOf(obj client.Object) *corev1.PodSpec { + switch typed := obj.(type) { + case *appsv1.Deployment: + return &typed.Spec.Template.Spec + case *appsv1.StatefulSet: + return &typed.Spec.Template.Spec + default: + return nil + } +} diff --git a/operator/internal/controllers/node/controlplane.go b/operator/internal/controllers/node/controlplane.go index f7c2615d3..f14a61b60 100644 --- a/operator/internal/controllers/node/controlplane.go +++ b/operator/internal/controllers/node/controlplane.go @@ -25,7 +25,9 @@ import ( "errors" "fmt" "net/http" + "sync" + "github.com/simplyblock/simplyblock-operator/internal/controllers/controlplane" "github.com/simplyblock/simplyblock-operator/internal/utils" "github.com/simplyblock/simplyblock-operator/internal/webapi" ) @@ -134,12 +136,29 @@ type ControlPlane interface { DeleteVolume(ctx context.Context, clusterID, poolID, volumeID string) error } -// httpControlPlane is the ControlPlane the operator runs with: the shared webapi -// client, with one method per endpoint of §12. -type httpControlPlane struct{ client *webapi.Client } +// httpControlPlane is the ControlPlane the operator runs with: one method per +// endpoint of §12, over a client whose address comes from the ControlPlane object. +type httpControlPlane struct { + // client is the startup client, built from the environment. It is what a + // call uses until the ControlPlane publishes an endpoint. + client *webapi.Client + + // resolve answers where the control plane is, per call. Nil means the + // startup client is the only one. + resolve controlplane.EndpointResolver + + // mu guards resolved, which is rebuilt when the published endpoint changes. + mu sync.Mutex + resolved *webapi.Client +} // NewControlPlane returns the HTTP-backed control-plane surface. -func NewControlPlane() ControlPlane { return &httpControlPlane{client: webapi.NewClient()} } +// +// The resolver may be nil, which is what a test passes: calls then go to the +// startup client and nothing reads a ControlPlane object. +func NewControlPlane(resolve controlplane.EndpointResolver) ControlPlane { + return &httpControlPlane{client: webapi.NewClient(), resolve: resolve} +} func (c *httpControlPlane) AddNode( ctx context.Context, clusterID string, params utils.StorageNodeSetAddParams, @@ -276,7 +295,7 @@ func (c *httpControlPlane) delete(ctx context.Context, path string) error { func (c *httpControlPlane) call( ctx context.Context, method, path string, body any, ) ([]byte, error) { - response, status, err := c.client.Do(ctx, method, path, body) + response, status, err := c.clientFor(ctx).Do(ctx, method, path, body) if err != nil { return nil, fmt.Errorf("%s %s: %w", method, path, err) } @@ -307,3 +326,31 @@ type ControlPlaneError struct { func (e *ControlPlaneError) Error() string { return fmt.Sprintf("the control plane answered %d: %s", e.Status, e.Body) } + +// clientFor is the client this call goes out on. +// +// The endpoint comes from ControlPlane.status.endpoint where the object has +// published one, which is what makes a remote control plane reachable: the +// client built at startup resolves SIMPLYBLOCK_WEBAPI_BASE_URL or the in-cluster +// default, and neither is where somebody else's control plane is +// (design-controlplane.md §3.3). +// +// With no resolver, or with one that answers nothing, the startup client is used +// unchanged. That is what keeps this additive: a deployment whose ControlPlane +// has not published an endpoint behaves as it did before. +func (c *httpControlPlane) clientFor(ctx context.Context) *webapi.Client { + if c.resolve == nil { + return c.client + } + endpoint := c.resolve(ctx) + if endpoint == "" || endpoint == c.client.BaseURL { + return c.client + } + + c.mu.Lock() + defer c.mu.Unlock() + if c.resolved == nil || c.resolved.BaseURL != endpoint { + c.resolved = webapi.NewClient(endpoint) + } + return c.resolved +} diff --git a/operator/internal/controllers/node/workload_controller.go b/operator/internal/controllers/node/workload_controller.go index f110bfb8e..1a41fe363 100644 --- a/operator/internal/controllers/node/workload_controller.go +++ b/operator/internal/controllers/node/workload_controller.go @@ -265,9 +265,8 @@ func (r *StorageNodeWorkloadReconciler) image( "spec.storageNodes.image is unset and ControlPlane %s cannot be read: %w", SingletonControlPlaneName, err) } - if source := controlPlane.Spec.Source; source != nil && source.Managed != nil && - source.Managed.Image != "" { - return source.Managed.Image, nil + if managed := controlPlane.Spec.Source.Local; managed != nil && managed.Image != "" { + return managed.Image, nil } return "", fmt.Errorf( "spec.storageNodes.image is unset and ControlPlane %s states no managed image", diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_controlplaneops.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_controlplaneops.yaml new file mode 100644 index 000000000..c0a9a5f6b --- /dev/null +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_controlplaneops.yaml @@ -0,0 +1,244 @@ +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition +metadata: + annotations: + controller-gen.kubebuilder.io/version: v0.21.0 + name: controlplaneops.storage.simplyblock.io +spec: + group: storage.simplyblock.io + names: + kind: ControlPlaneOps + listKind: ControlPlaneOpsList + plural: controlplaneops + shortNames: + - cpops + singular: controlplaneops + scope: Namespaced + versions: + - additionalPrinterColumns: + - jsonPath: .spec.controlPlaneRef + name: ControlPlane + type: string + - jsonPath: .spec.action + name: Action + type: string + - jsonPath: .status.phase + name: Phase + type: string + - jsonPath: .status.step.state + name: Step + type: string + - jsonPath: .status.message + name: Message + priority: 1 + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha2 + schema: + openAPIV3Schema: + description: |- + ControlPlaneOps is a single operation performed against the control plane. It + runs to a terminal phase and stays afterward as the audit record of what was + done, with which parameters, and how it ended. Only one may be active per + control plane at a time, which the entity's status.activeOpsRef enforces. + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: |- + ControlPlaneOpsSpec is one operation to perform against the control plane. + + Everything except spec.abort is frozen once the object is admitted, which is + what makes the status an audit of the request that ran rather than of whatever + the object says now. The parameters are consumed several steps apart: + Preflight reads spec.upgrade.image and Applying writes it, and Draining reads + spec.restart.components before Restarting recycles them. An edit in between + produces an operation that checked one thing and did another. + + The rules are declared here rather than as +k8s:immutable on each field. + controller-gen emits that marker's rules in an order that varies between runs + once a type carries several, and it freezes a block whole; what has to be + frozen is each block's presence together with its contents. + properties: + abort: + description: |- + Abort asks a running operation to stop at its next step and unwind. It is + the one field of this spec an update may change, because it is the one that + is meant to be set after the operation started. Whether an abort is + expressible from the current step is declared by that action's graph rather + than checked here. + type: boolean + action: + description: |- + Action is the operation to perform. Immutable, so that the status describes + the operation that ran. + enum: + - Restart + - Upgrade + - Backup + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + backup: + description: Backup parameterizes action Backup and is ignored by + the others. + properties: + backupName: + description: |- + BackupName is the FoundationDBBackup to create or trigger. Absent uses the + one already configured for the cluster, and fails when there is none and + no name to create. + type: string + blobStore: + description: |- + BlobStore is the destination, in the form the FoundationDBBackup CRD takes + it. The operator copies it through rather than interpreting it, since the + backup is the FoundationDB operator's to perform. + type: string + required: + - blobStore + type: object + controlPlaneRef: + description: |- + ControlPlaneRef names the ControlPlane this operation acts on, in this + object's own namespace. The operation never owns its target, because + deleting the record of an operation must not delete the control plane it + operated on. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + restart: + description: Restart parameterizes action Restart and is ignored by + the others. + properties: + components: + description: |- + Components names the workloads to recycle, from the table in §4.3. Empty + recycles the whole control plane. Naming only components that table marks + non-essential skips the drain, because recycling them interrupts nothing. + items: + type: string + type: array + x-kubernetes-list-type: set + type: object + upgrade: + description: Upgrade parameterizes action Upgrade and is ignored by + the others. + properties: + image: + description: |- + Image is the version to move to. It replaces + ControlPlane.spec.source.managed.image when the operation succeeds, so the + entity keeps describing what is running. + pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ + type: string + required: + - image + type: object + required: + - action + - controlPlaneRef + type: object + x-kubernetes-validations: + - message: 'spec.upgrade is immutable: Preflight checked the image the + operation was admitted with, and Applying writes it several steps + later' + rule: has(self.upgrade) == has(oldSelf.upgrade) && (!has(self.upgrade) + || self.upgrade == oldSelf.upgrade) + - message: 'spec.restart is immutable: the drain is decided from the component + list, so widening it afterward skips a drain the wider list would + have required' + rule: has(self.restart) == has(oldSelf.restart) && (!has(self.restart) + || self.restart == oldSelf.restart) + - message: 'spec.backup is immutable: the destination is what Requesting + created the FoundationDBBackup against' + rule: has(self.backup) == has(oldSelf.backup) && (!has(self.backup) + || self.backup == oldSelf.backup) + status: + description: ControlPlaneOpsStatus is the observed state of one control-plane + operation. + properties: + backupRef: + description: |- + BackupRef names the FoundationDBBackup a Backup run created or triggered. + The operation does not own it, because deleting the record of a backup + must not delete the backup's configuration. + type: string + completedAt: + description: CompletedAt is when it reached a terminal phase. + format: date-time + type: string + message: + description: |- + Message is the reason the phase is what it is: one sentence, replaced as + the operation moves, and never a log. + type: string + observedGeneration: + description: |- + ObservedGeneration is the generation the rest of this status was computed + from, so a stale status can be told from a current one. + format: int64 + type: integer + phase: + description: Phase is the operation's own progress. + enum: + - Pending + - Running + - Succeeded + - Failed + - Aborted + type: string + startedAt: + description: StartedAt is when the operation acquired its target's + lock. + format: date-time + type: string + step: + description: |- + Step is the position of the running action's state machine. It is + persisted before the side effect that step performs. The rule repeats the + ControlPlaneOpsStep enum because a marker cannot reach a field of the + shared snapshot type. + properties: + deadline: + description: |- + Deadline is when that state expires, absent when it has none. It is an + absolute instant, so a state whose deadline passed while the controller + was down restores as already expired. + format: date-time + type: string + state: + description: |- + State is the state the machine was in. Empty means the resource has not + been reconciled yet, and restores to the graph's initial state. + type: string + type: object + x-kubernetes-validations: + - message: unknown step + rule: '!has(self.state) || self.state in [''Draining'',''Restarting'',''Awaiting'',''Preflight'',''Applying'',''Verifying'',''Requesting'']' + type: object + type: object + served: true + storage: true + subresources: + status: {} diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_controlplanes.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_controlplanes.yaml index 1f080d803..140c3da8f 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_controlplanes.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_controlplanes.yaml @@ -11,6 +11,8 @@ spec: kind: ControlPlane listKind: ControlPlaneList plural: controlplanes + shortNames: + - cp singular: controlplane scope: Namespaced versions: @@ -59,9 +61,11 @@ spec: image: description: |- Image is the container image used for all simplyblock control-plane and - storage-node workloads (e.g. quay.io/simplyblock-io/simplyblock:26.2.2). + storage-node workloads (e.g., `quay.io/simplyblock-io/simplyblock:26.2.2`). StorageNodeSet CRs that omit spec.clusterImage inherit this value. - Must reference one of the trusted registries (quay.io/simplyblock-io, docker.io/simplyblock, public.ecr.aws/simply-block); digest pinning (@sha256:...) is recommended. + Must reference one of the trusted registries (`quay.io/simplyblock-io`, + `docker.io/simplyblock`, `public.ecr.aws/simply-block`). Digest pinning + (@sha256:...) is recommended. pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ type: string type: object @@ -77,7 +81,7 @@ spec: message: description: |- Message contains a human-readable explanation of the current phase, - for example the FDB error returned by the health endpoint. + for example, the FDB error returned by the health endpoint. type: string phase: description: |- @@ -94,13 +98,21 @@ spec: subresources: status: {} - additionalPrinterColumns: - - description: Initializing while FDB is not ready; Available once the control plane is operational - jsonPath: .status.phase + - jsonPath: .status.phase name: Phase type: string - - description: Human-readable status detail - jsonPath: .status.message + - jsonPath: .status.step.state + name: Step + type: string + - jsonPath: .status.endpoint + name: Endpoint + type: string + - jsonPath: .status.version + name: Version + type: string + - jsonPath: .status.message name: Message + priority: 1 type: string - jsonPath: .metadata.creationTimestamp name: Age @@ -109,9 +121,11 @@ spec: schema: openAPIV3Schema: description: |- - ControlPlane is a singleton resource (one per namespace, named "simplyblock") - that reflects the readiness of the simplyblock control plane. It is created - automatically by the Helm chart and should not be created or deleted manually. + ControlPlane is the simplyblock control plane for one Kubernetes cluster: + FoundationDB together with the management API, either installed by the + operator or already existing. It is a singleton named `simplyblock`, and it is + the root of the ownership spine: nothing else in this API group reconciles + meaningfully before it reports Available. properties: apiVersion: description: |- @@ -131,48 +145,406 @@ spec: metadata: type: object spec: - description: ControlPlaneSpec holds configuration for the singleton ControlPlane resource. + description: |- + ControlPlaneSpec is the desired state of the simplyblock control plane for one + namespace. properties: source: description: |- - Source says where the control plane comes from. It replaces the top-level - image field of v1alpha1, which conflated the control plane's own image with - the default every StorageNodeSet inherited. + Source selects whether this cluster hosts its control plane or is managed + by one elsewhere. Switching a live deployment between the two is not a + reconfiguration, because the clusters and their volumes live in the + FoundationDB behind the old one, so which mode is chosen is frozen at + creation. What is inside the chosen mode stays editable. + + Both rules are declared on ControlPlaneSource rather than here. See the + type for why. properties: - managed: - description: Managed is the control plane the operator installs and owns. + local: + description: Local is a control plane the operator installs. properties: + foundationDB: + description: FoundationDB sizes the FoundationDB the management API stores its state in. + properties: + replicas: + default: 3 + description: |- + Replicas is the number of coordinators. Three is the smallest count that + survives one loss, which is why it is the default. + format: int32 + minimum: 1 + type: integer + resources: + description: Resources sets requests and limits for the coordinator pods. + properties: + claims: + description: |- + Claims lists the names of resources, defined in spec.resourceClaims, + that are used by this container. + + This field depends on the + DynamicResourceAllocation feature gate. + + This field is immutable. It can only be set for containers. + items: + description: ResourceClaim references one entry in PodSpec.ResourceClaims. + properties: + name: + description: |- + Name must match the name of one entry in pod.spec.resourceClaims of + the Pod where this field is used. It makes that resource available + inside a container. + type: string + request: + description: |- + Request is the name chosen for a request in the referenced claim. + If empty, everything from the claim is made available, otherwise + only the result of this request. + type: string + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + limits: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Limits describes the maximum amount of compute resources allowed. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + requests: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Requests describes the minimum amount of compute resources required. + If Requests is omitted for a container, it defaults to Limits if that is explicitly specified, + otherwise to an implementation-defined value. Requests cannot exceed Limits. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + type: object + storageClassName: + description: |- + StorageClassName is the class the coordinators' volumes are provisioned + from. It cannot be a class this operator provides, because the control + plane has to exist before any simplyblock volume can. + type: string + type: object image: description: |- - Image is the container image used for the simplyblock control-plane - workloads (e.g., quay.io/simplyblock-io/simplyblock:26.2.2). - Must reference one of the trusted registries (`quay.io/simplyblock-io`, `docker.io/simplyblock`, `public.ecr.aws/simply-block`); digest pinning (@sha256:...) is recommended. + Image is the management API and control-plane image. + Must reference one of the trusted registries (`quay.io/simplyblock-io`, + `docker.io/simplyblock`, `public.ecr.aws/simply-block`); digest pinning + (@sha256:...) is recommended. pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ type: string + imagePullPolicy: + default: IfNotPresent + description: ImagePullPolicy controls when that image is pulled. + enum: + - Always + - Never + - IfNotPresent + type: string + nodeSelector: + additionalProperties: + type: string + description: |- + NodeSelector pins every pod the operator installs for the control plane. + It is a selector rather than an affinity term because that is what the + chart it replaces took, and a deployment migrating off the chart has the + value already written down. + type: object + replicas: + default: 2 + description: |- + Replicas is the number of management API instances. Two is what the chart + ships and what the phases assume: a single instance makes Degraded + unreachable for this component and every restart an outage (§5.1). + format: int32 + minimum: 1 + type: integer + resources: + description: Resources sets requests and limits for the management API pods. + properties: + claims: + description: |- + Claims lists the names of resources, defined in spec.resourceClaims, + that are used by this container. + + This field depends on the + DynamicResourceAllocation feature gate. + + This field is immutable. It can only be set for containers. + items: + description: ResourceClaim references one entry in PodSpec.ResourceClaims. + properties: + name: + description: |- + Name must match the name of one entry in pod.spec.resourceClaims of + the Pod where this field is used. It makes that resource available + inside a container. + type: string + request: + description: |- + Request is the name chosen for a request in the referenced claim. + If empty, everything from the claim is made available, otherwise + only the result of this request. + type: string + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + limits: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Limits describes the maximum amount of compute resources allowed. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + requests: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Requests describes the minimum amount of compute resources required. + If Requests is omitted for a container, it defaults to Limits if that is explicitly specified, + otherwise to an implementation-defined value. Requests cannot exceed Limits. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + type: object + tolerations: + description: |- + Tolerations are applied to every pod the operator installs for the control + plane. + items: + description: |- + The pod this Toleration is attached to tolerates any taint that matches + the triple using the matching operator . + properties: + effect: + description: |- + Effect indicates the taint effect to match. Empty means match all taint effects. + When specified, allowed values are NoSchedule, PreferNoSchedule and NoExecute. + type: string + key: + description: |- + Key is the taint key that the toleration applies to. Empty means match all taint keys. + If the key is empty, operator must be Exists; this combination means to match all values and all keys. + type: string + operator: + description: |- + Operator represents a key's relationship to the value. + Valid operators are Exists, Equal, Lt, and Gt. Defaults to Equal. + Exists is equivalent to wildcard for value, so that a pod can + tolerate all taints of a particular category. + Lt and Gt perform numeric comparisons (requires feature gate TaintTolerationComparisonOperators). + type: string + tolerationSeconds: + description: |- + TolerationSeconds represents the period of time the toleration (which must be + of effect NoExecute, otherwise this field is ignored) tolerates the taint. By default, + it is not set, which means tolerate the taint forever (do not evict). Zero and + negative values will be treated as 0 (evict immediately) by the system. + format: int64 + type: integer + value: + description: |- + Value is the taint value the toleration matches to. + If the operator is Exists, the value should be empty, otherwise just a regular string. + type: string + type: object + type: array + required: + - image + type: object + managed: + description: Managed is a control plane that already exists. + properties: + caBundleSecretRef: + description: |- + CABundleSecretRef names a Secret holding the CA certificate the endpoint + is verified against. Absent means the system trust store. + properties: + name: + default: "" + description: |- + Name of the referent. + This field is effectively required, but due to backwards compatibility is + allowed to be empty. Instances of this type with an empty value here are + almost certainly wrong. + More info: https://kubernetes.io/docs/concepts/overview/working-with-objects/names/#names + type: string + type: object + x-kubernetes-map-type: atomic + credentialsSecretRef: + description: |- + CredentialsSecretRef names a Secret in this namespace holding the bearer + token the operator authenticates with. It is a reference rather than a + field because a token in a spec is a token in every `kubectl get -o yaml`. + + Absent means the endpoint is reached without one, which is the in-cluster + case: a control plane the Helm chart installed answers on a ClusterIP + Service in this namespace and does not require a token for the readiness + read. Naming a Secret that does not exist stays an error, because naming + one is a statement that the control plane needs it. + properties: + name: + default: "" + description: |- + Name of the referent. + This field is effectively required, but due to backwards compatibility is + allowed to be empty. Instances of this type with an empty value here are + almost certainly wrong. + More info: https://kubernetes.io/docs/concepts/overview/working-with-objects/names/#names + type: string + type: object + x-kubernetes-map-type: atomic + endpoint: + description: |- + Endpoint is the management API's base URL. It is validated against the + same outbound-URL guard every other outbound endpoint in this group uses, + so a loopback or link-local address is rejected. + pattern: ^https?://[a-zA-Z0-9.-]+(:[0-9]{1,5})?(/.*)?$ + type: string + required: + - endpoint type: object type: object + x-kubernetes-validations: + - message: set exactly one of local or managed + rule: '(has(self.local) ? 1 : 0) + (has(self.managed) ? 1 : 0) == 1' + - message: 'spec.source is immutable: a control plane the operator installed and one it did not are different deployments, and the clusters and their volumes live in the FoundationDB behind the old one' + rule: has(self.local) == has(oldSelf.local) && has(self.managed) == has(oldSelf.managed) + required: + - source type: object status: - description: |- - ControlPlaneStatus reflects the observed readiness of the simplyblock - control plane (FDB + management API). + description: ControlPlaneStatus is the observed state of the control plane. properties: + activeOpsRef: + description: |- + ActiveOpsRef names the ControlPlaneOps currently allowed to act on this + control plane. Empty when none is running. + type: string + components: + description: |- + Components is the per-component readiness the phase is derived from + (§4.3), one entry per workload the managed install applies. It is empty + for a remote control plane, which has no components the operator owns. + Without it a Degraded phase says that something is wrong and not what. + items: + description: |- + ControlPlaneComponentStatus is one workload of a managed control plane and how + much of it is running. The phase is the worst verdict across these and the + readiness probe, and only an essential component at zero ready can make it + Unavailable. + properties: + desired: + description: |- + Desired is how many replicas the component should have. For the component + carrying its own operator it is that resource's own count, because a + FoundationDBCluster reports quorum rather than replicas. + format: int32 + minimum: 0 + type: integer + essential: + description: |- + Essential states whether this component at zero ready makes the control + plane Unavailable rather than Degraded. It is decided by the table in + §4.3 rather than by a user, and it is reported here so that a phase can be + explained without reading the operator's source. + type: boolean + name: + description: Name is the workload's name, as applied. + type: string + ready: + description: Ready is how many of them are. + format: int32 + minimum: 0 + type: integer + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + endpoint: + description: |- + Endpoint is the resolved management API base URL, derived in the local + case and echoed in the managed one. It is what every controller in the + operator reads to reach the control plane, so that one object answers + where it is. + type: string lastChecked: - description: LastChecked is the timestamp of the most recent FDB health probe. + description: LastChecked is when the readiness probe last ran. format: date-time type: string message: description: |- - Message contains a human-readable explanation of the current phase, - for example, the FDB error returned by the health endpoint. + Message is the reason the phase is what it is: one sentence, replaced as + the control plane moves, and never a log. On a failed probe it is the + control plane's own error rather than a paraphrase of it. type: string - phase: + observedGeneration: description: |- - Phase is Initializing while the control plane is not yet healthy, - and Available once the FDB health check passes. + ObservedGeneration is the generation the rest of this status was computed + from, so a stale status can be told from a current one. + format: int64 + type: integer + phase: + description: Phase is the operator's own view of the control plane. enum: - - Initializing + - Installing - Available + - Degraded + - Unavailable + type: string + step: + description: |- + Step is the position of the installation machine within Installing. The + rule repeats the ControlPlaneStep enum because a marker cannot reach a + field of the shared snapshot type. + properties: + deadline: + description: |- + Deadline is when that state expires, absent when it has none. It is an + absolute instant, so a state whose deadline passed while the controller + was down restores as already expired. + format: date-time + type: string + state: + description: |- + State is the state the machine was in. Empty means the resource has not + been reconciled yet, and restores to the graph's initial state. + type: string + type: object + x-kubernetes-validations: + - message: unknown step + rule: '!has(self.state) || self.state in [''ApplyingFoundationDB'',''AwaitingFoundationDB'',''ApplyingDatastore'',''ApplyingAPI'',''AwaitingAPI'']' + version: + description: Version is the version the management API reports. type: string type: object type: object diff --git a/operator/internal/webhook/controlplaneops_validator.go b/operator/internal/webhook/controlplaneops_validator.go new file mode 100644 index 000000000..a7d4d0a55 --- /dev/null +++ b/operator/internal/webhook/controlplaneops_validator.go @@ -0,0 +1,180 @@ +// The ControlPlaneOps guard: a validating webhook that resolves +// spec.controlPlaneRef at creation and refuses an operation naming a control +// plane none of its actions can act on. +// +// Every action acts on what the operator installed: Restart recycles a workload, +// Upgrade replaces its image, and Backup asks the FoundationDBCluster the +// operator applied. An operation naming a remote control plane is therefore one +// that can only fail. Admission is where the check belongs, because an operation +// that can only fail belongs in an error message on the terminal that wrote it +// rather than in a Failed object somebody has to go and read. +// +// Admission is also where the check *can* live, because the answer cannot move: +// spec.controlPlaneRef is immutable and ControlPlane.spec.source is immutable, so +// a control plane admitted as managed stays managed for the life of the +// operation. What can still happen is the target being deleted, and an operation +// whose target has vanished is a missing-target failure the controller handles. +// +// design-controlplane.md §6 is the specification. + +package webhook + +import ( + "context" + "encoding/json" + "fmt" + "net/http" + + admissionv1 "k8s.io/api/admission/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/webhook/admission" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// +kubebuilder:webhook:path=/validate-storage-simplyblock-io-v1alpha2-controlplaneops,mutating=false,failurePolicy=fail,sideEffects=None,groups=storage.simplyblock.io,resources=controlplaneops,verbs=create;delete,versions=v1alpha2,name=vcontrolplaneops.simplyblock.io,admissionReviewVersions=v1 + +// ControlPlaneOpsValidator refuses an operation whose target cannot be operated +// on, and one whose action is missing the block that parameterizes it. +// +// failurePolicy=Fail, because the webhook server runs inside the operator pod: +// its availability tracks the operator's, and while the operator is down nothing +// advances an operation anyway. +type ControlPlaneOpsValidator struct { + Client client.Client +} + +// undeletableControlPlaneSteps are the steps from which a running operation's record may not +// be withdrawn, and what it is in the middle of. +// +// Each has already changed something that nothing else would finish. Restarting +// and Applying have rolled a Deployment or written a new image onto the entity, +// and Awaiting and Verifying are watching that rollout back. Deleting the record +// there releases the control plane's lock while the rollout is still moving, so +// the next operation starts against a half-applied one. +// +// It is the same guard StorageBackupOps carries, and the stronger of the two the +// operation has: the finalizer alone does not catch a forced delete with no +// grace period. +var undeletableControlPlaneSteps = map[simplyblockv1alpha2.ControlPlaneOpsStep]string{ + simplyblockv1alpha2.ControlPlaneOpsStepRestarting: "the workloads are being recycled", + simplyblockv1alpha2.ControlPlaneOpsStepApplying: "the new image is being written onto the control plane", + simplyblockv1alpha2.ControlPlaneOpsStepAwaiting: "the recycled workloads are still coming back", + simplyblockv1alpha2.ControlPlaneOpsStepVerifying: "the rollout is being checked against the version it asked for", +} + +func (v *ControlPlaneOpsValidator) Handle(ctx context.Context, req admission.Request) admission.Response { + switch req.Operation { + case admissionv1.Create: + return v.admitCreate(ctx, req) + case admissionv1.Delete: + return v.admitDelete(req) + default: + return admission.Allowed("") + } +} + +// admitDelete refuses to withdraw the record of an operation that is part-way +// through a rollout. +// +// The object is read from req.OldObject, which is what the API server sends on a +// DELETE: there is no new object, and the step the operation is on is in the +// status of the one being removed. +func (v *ControlPlaneOpsValidator) admitDelete(req admission.Request) admission.Response { + if len(req.OldObject.Raw) == 0 { + // Nothing to read means nothing to refuse on. Admitting is the only + // answer that does not block a delete on the basis of no information. + return admission.Allowed("") + } + + var ops simplyblockv1alpha2.ControlPlaneOps + if err := json.Unmarshal(req.OldObject.Raw, &ops); err != nil { + return admission.Errored(http.StatusBadRequest, err) + } + + // A terminal operation is a record of work that has finished, and + // withdrawing it stops nothing. + switch ops.Status.Phase { + case simplyblockv1alpha2.ControlPlaneOpsPhaseSucceeded, + simplyblockv1alpha2.ControlPlaneOpsPhaseFailed, + simplyblockv1alpha2.ControlPlaneOpsPhaseAborted: + return admission.Allowed("the operation is terminal") + } + + step := simplyblockv1alpha2.ControlPlaneOpsStep(ops.Status.Step.State) + doing, undeletable := undeletableControlPlaneSteps[step] + if !undeletable { + return admission.Allowed("") + } + + return admission.Denied(fmt.Sprintf( + "ControlPlaneOps %s/%s is at step %s, where %s. Deleting the record would not stop that "+ + "work, it would release the control plane's lock while it is still moving and let the "+ + "next operation start against a half-applied one. Wait for it to finish, then delete "+ + "the record.", + ops.Namespace, ops.Name, step, doing)) +} + +// admitCreate resolves the target and checks the action's parameters. +func (v *ControlPlaneOpsValidator) admitCreate( + ctx context.Context, req admission.Request, +) admission.Response { + var ops simplyblockv1alpha2.ControlPlaneOps + if err := json.Unmarshal(req.Object.Raw, &ops); err != nil { + return admission.Errored(http.StatusBadRequest, err) + } + + // v1alpha2 is the stored version, and it is read directly. A read at + // v1alpha1 is answered only by the conversion webhook, which a fresh install + // does not deploy. + var target simplyblockv1alpha2.ControlPlane + key := client.ObjectKey{Name: ops.Spec.ControlPlaneRef, Namespace: ops.Namespace} + err := v.Client.Get(ctx, key, &target) + switch { + case apierrors.IsNotFound(err): + return admission.Denied(fmt.Sprintf( + "spec.controlPlaneRef %q does not name a ControlPlane in namespace %q", + ops.Spec.ControlPlaneRef, ops.Namespace)) + case err != nil: + return admission.Errored(http.StatusInternalServerError, err) + } + + if target.Spec.Source.Local == nil { + return admission.Denied(fmt.Sprintf( + "ControlPlane %q names a control plane this cluster does not host, and every "+ + "action of this kind acts on something the operator installed: Restart recycles a "+ + "workload, Upgrade replaces its image, and Backup asks the FoundationDBCluster the "+ + "operator applied. A control plane elsewhere is an endpoint and a credential, so "+ + "there is nothing here for any of them to act on.", + ops.Spec.ControlPlaneRef)) + } + + return actionBlockPresent(&ops) +} + +// actionBlockPresent refuses an action whose parameter block is absent. +// +// The blocks are optional on the type because each action ignores the others', +// so the API cannot require one without requiring all three. Checking it here is +// what turns "the operation failed at its first step" into an error on the +// terminal that wrote the object. +func actionBlockPresent(ops *simplyblockv1alpha2.ControlPlaneOps) admission.Response { + switch ops.Spec.Action { + case simplyblockv1alpha2.ControlPlaneOpsActionUpgrade: + if ops.Spec.Upgrade == nil || ops.Spec.Upgrade.Image == "" { + return admission.Denied( + "action Upgrade needs spec.upgrade.image, which names the version to move to") + } + case simplyblockv1alpha2.ControlPlaneOpsActionBackup: + if ops.Spec.Backup == nil || ops.Spec.Backup.BlobStore == "" { + return admission.Denied( + "action Backup needs spec.backup.blobStore, which names where the backup goes") + } + case simplyblockv1alpha2.ControlPlaneOpsActionRestart: + // Restart takes no required block: an empty spec.restart.components + // recycles the whole control plane, which is the common case and a + // legitimate one to write nothing for. + } + return admission.Allowed("") +} diff --git a/operator/internal/webhook/controlplaneops_validator_test.go b/operator/internal/webhook/controlplaneops_validator_test.go new file mode 100644 index 000000000..2a3151589 --- /dev/null +++ b/operator/internal/webhook/controlplaneops_validator_test.go @@ -0,0 +1,147 @@ +// What the ControlPlaneOps guard admits and refuses. +// +// The delete cases matter most. The controller's deletion path releases the +// control plane's lock and drops the finalizer from any phase, so without this +// webhook the record of a running rollout can be withdrawn while the rollout is +// still moving, and the next operation starts against it. + +package webhook + +import ( + "context" + "encoding/json" + "strings" + "testing" + + admissionv1 "k8s.io/api/admission/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/runtime" + clientgoscheme "k8s.io/client-go/kubernetes/scheme" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + "sigs.k8s.io/controller-runtime/pkg/webhook/admission" + + "github.com/simplyblock/atlas/statemachine" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +func cpOpsScheme(t *testing.T) *runtime.Scheme { + t.Helper() + scheme := runtime.NewScheme() + if err := clientgoscheme.AddToScheme(scheme); err != nil { + t.Fatalf("build the scheme: %v", err) + } + if err := simplyblockv1alpha2.AddToScheme(scheme); err != nil { + t.Fatalf("build the scheme: %v", err) + } + return scheme +} + +func cpOpsValidator(t *testing.T) *ControlPlaneOpsValidator { + t.Helper() + return &ControlPlaneOpsValidator{ + Client: fake.NewClientBuilder().WithScheme(cpOpsScheme(t)).Build(), + } +} + +func deleteRequestFor(t *testing.T, ops *simplyblockv1alpha2.ControlPlaneOps) admission.Request { + t.Helper() + raw, err := json.Marshal(ops) + if err != nil { + t.Fatalf("marshal the operation: %v", err) + } + return admission.Request{AdmissionRequest: admissionv1.AdmissionRequest{ + Operation: admissionv1.Delete, + OldObject: runtime.RawExtension{Raw: raw}, + }} +} + +func runningOpsAt(step simplyblockv1alpha2.ControlPlaneOpsStep) *simplyblockv1alpha2.ControlPlaneOps { + return &simplyblockv1alpha2.ControlPlaneOps{ + ObjectMeta: metav1.ObjectMeta{Name: "an-operation", Namespace: "simplyblock"}, + Spec: simplyblockv1alpha2.ControlPlaneOpsSpec{ + ControlPlaneRef: "simplyblock", + Action: simplyblockv1alpha2.ControlPlaneOpsActionUpgrade, + }, + Status: simplyblockv1alpha2.ControlPlaneOpsStatus{ + Phase: simplyblockv1alpha2.ControlPlaneOpsPhaseRunning, + Step: statemachine.KubeSnapshot{State: string(step)}, + }, + } +} + +// A record may not be withdrawn from a step that has started a rollout. The +// controller's deletion path would release the lock and let the next operation +// begin against a half-applied control plane. +func TestControlPlaneOpsDeleteIsRefusedMidRollout(t *testing.T) { + v := cpOpsValidator(t) + + for _, step := range []simplyblockv1alpha2.ControlPlaneOpsStep{ + simplyblockv1alpha2.ControlPlaneOpsStepRestarting, + simplyblockv1alpha2.ControlPlaneOpsStepApplying, + simplyblockv1alpha2.ControlPlaneOpsStepAwaiting, + simplyblockv1alpha2.ControlPlaneOpsStepVerifying, + } { + t.Run(string(step), func(t *testing.T) { + resp := v.Handle(context.Background(), deleteRequestFor(t, runningOpsAt(step))) + if resp.Allowed { + t.Fatalf("the record was withdrawn at step %s, while the rollout was moving", step) + } + if !strings.Contains(resp.Result.Message, string(step)) { + t.Errorf("the refusal is %q, want it to name the step", resp.Result.Message) + } + }) + } +} + +// A step that has changed nothing is deletable, because withdrawing the record +// there stops nothing that anything else has to finish. +func TestControlPlaneOpsDeleteIsAllowedBeforeAnythingChanged(t *testing.T) { + v := cpOpsValidator(t) + + for _, step := range []simplyblockv1alpha2.ControlPlaneOpsStep{ + simplyblockv1alpha2.ControlPlaneOpsStepDraining, + simplyblockv1alpha2.ControlPlaneOpsStepPreflight, + simplyblockv1alpha2.ControlPlaneOpsStepRequesting, + } { + t.Run(string(step), func(t *testing.T) { + resp := v.Handle(context.Background(), deleteRequestFor(t, runningOpsAt(step))) + if !resp.Allowed { + t.Errorf("the record was held at step %s, where nothing has been changed: %s", + step, resp.Result.Message) + } + }) + } +} + +// A terminal operation is a record of work that has finished, so withdrawing it +// stops nothing whatever step it ended on. +func TestControlPlaneOpsDeleteIsAllowedOnceTerminal(t *testing.T) { + v := cpOpsValidator(t) + + for _, phase := range []simplyblockv1alpha2.ControlPlaneOpsPhase{ + simplyblockv1alpha2.ControlPlaneOpsPhaseSucceeded, + simplyblockv1alpha2.ControlPlaneOpsPhaseFailed, + simplyblockv1alpha2.ControlPlaneOpsPhaseAborted, + } { + t.Run(string(phase), func(t *testing.T) { + ops := runningOpsAt(simplyblockv1alpha2.ControlPlaneOpsStepApplying) + ops.Status.Phase = phase + + resp := v.Handle(context.Background(), deleteRequestFor(t, ops)) + if !resp.Allowed { + t.Errorf("a %s operation could not be deleted: %s", phase, resp.Result.Message) + } + }) + } +} + +// Every step the graphs declare is either undeletable for a stated reason or +// deletable. A step in neither set is one this guard has no opinion about, which +// is how a new step silently becomes withdrawable mid-rollout. +func TestEveryUndeletableStepStatesWhatItIsDoing(t *testing.T) { + for step, doing := range undeletableControlPlaneSteps { + if doing == "" { + t.Errorf("%s is undeletable and the refusal says nothing about why", step) + } + } +} From bf91e8ca2a1057951f237aa9cab745c9beeb3e5f Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Thu, 17 Sep 2026 00:57:17 +0100 Subject: [PATCH 025/206] added design doc design-csi-addons-replication.md --- .../designs/design-csi-addons-replication.md | 430 ++++++++++++++++++ .../tests/test-plan-csi-addons-replication.md | 175 +++++++ 2 files changed, 605 insertions(+) create mode 100644 operator/docs/designs/design-csi-addons-replication.md create mode 100644 operator/docs/tests/test-plan-csi-addons-replication.md diff --git a/operator/docs/designs/design-csi-addons-replication.md b/operator/docs/designs/design-csi-addons-replication.md new file mode 100644 index 000000000..8535880b6 --- /dev/null +++ b/operator/docs/designs/design-csi-addons-replication.md @@ -0,0 +1,430 @@ +# Design Document: csi-addons Volume Replication + +**Status:** Draft +**Author:** Israel Geoffrey (geoffrey1330) +**Date:** 2026-09-16 +**Test Plan:** [`tests/test-plan-csi-addons-replication.md`](../tests/test-plan-csi-addons-replication.md) + +--- + +## Phasing Overview + +| Phase | Status | Scope | Sections | +|-------------|---------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|--------------| +| **Phase 1** | Planned | The csi-addons machinery and the steady-state contract: CRDs, controller-manager, sidecar, the Replication and csi-addons Identity gRPC services with `EnableVolumeReplication`, `DisableVolumeReplication`, and `GetVolumeReplicationInfo`, backed by a typed backend status endpoint | §4, §5.1, §6 | +| **Phase 2** | Planned | The lifecycle verbs: `PromoteVolume` (planned and forced), `DemoteVolume`, and `ResyncVolume`, validated end to end against a Ramen `VolumeReplicationGroup` in async mode | §5.2, §9 | +| **Phase 3** | Planned | peerClasses convention and preflight, and the replication observability surface (lag, backlog, RPO compliance) | §7, §11 | +| **Phase 4** | Planned | Test failover: the latest-replicated-snapshot read, the `drtest-*` conventions, and the two drill modes (bubble and test cluster) composed from clone, replication, and the real failover | §14 | + +Phase 1 is independently useful: a `VolumeReplication` object per PVC whose status truthfully reports the relationship, which no surface provides today. Phase 2 makes the object drivable, which is what Ramen actually needs. Phase 3 makes the whole thing operable at fleet scale. Phase 4 turns the same primitives into a rehearsal: a failover that can be drilled, in a bubble or against a test cluster, without touching production replication. + +--- + +## Phase 0 — External Prerequisites + +| # | Prerequisite | Kind | Blocks | Status | +|------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------|---------|----------------------------------| +| P0-1 | A typed, steady-state per-volume replication status read: `GET .../volumes/{id}/replication/status` serving what `lvol_controller.get_replication_info` computes today (state, lag, outstanding bytes, failure counters), available for the volume's whole replicated life | Control plane (`sbcli`) | Phase 1 | Not shipped | +| P0-2 | Idempotent attach and detach: attaching a volume to the policy it already follows returns success, and detaching a non-attached volume returns success | Control plane (`sbcli`) | Phase 1 | Not shipped | +| P0-3 | A standalone demote verb: `POST .../volumes/{id}/replication/demote` that converges the peer while still serving (repeated snapshot-and-ship until the remaining delta is small), then quiesces, ships the final delta, confirms it landed on the peer, and fences the data path | Control plane (`sbcli`) | Phase 2 | Not shipped | +| P0-4 | An `rpo_target_seconds` field on `ReplicationPolicy`, so RPO compliance is computable against a declared target rather than the derived lag budget | Control plane (`sbcli`) | Phase 3 | Not shipped | +| P0-5 | csi-addons upstream: the `VolumeReplication` and `VolumeReplicationClass` CRDs (`replication.storage.openshift.io/v1alpha1`), the kubernetes-csi-addons controller-manager image, and the csi-addons sidecar image | Ecosystem | Phase 1 | Available upstream; not vendored | +| P0-6 | A latest-replicated-snapshot read: per volume, and per consistency group as one complete generation, the newest fully replicated snapshot on the secondary addressed as a cloneable object | Control plane (`sbcli`) | Phase 4 | Not shipped | + +Everything else the adapter needs already exists: the attach and detach calls, failover, the failback and commit pair, the relationship read, and the backlog arithmetic inside `get_replication_info`. The adapter is thin precisely because the engine is complete. What is missing is the shape Ramen can drive. + +--- + +## Table of Contents + +1. [Background](#1-background) +2. [Goals and Non-Goals](#2-goals-and-non-goals) +3. [Architecture Overview](#3-architecture-overview) +4. [The csi-addons Machinery](#4-the-csi-addons-machinery) +5. [The Replication Service](#5-the-replication-service) +6. [Steady-State Status and Conditions](#6-steady-state-status-and-conditions) +7. [VolumeReplicationClass and peerClasses](#7-volumereplicationclass-and-peerclasses) +8. [Coexistence with the Legacy Replication Kinds](#8-coexistence-with-the-legacy-replication-kinds) +9. [Backend API Requirements](#9-backend-api-requirements) +10. [Failure Modes and Fallback](#10-failure-modes-and-fallback) +11. [Observability](#11-observability) +12. [Testing Strategy](#12-testing-strategy) +13. [Migration Strategy](#13-migration-strategy) +14. [Test Failover](#14-test-failover) +15. [Open Questions](#15-open-questions) + +--- + +## Overview + +simplyblock's replication engine is complete and works: interval snapshots ship per volume under a `ReplicationPolicy`, failover clones the last replicated generation on the target, and the failback-plus-commit pair performs a fenced, lossless cutover back. What it is not is drivable by anything outside simplyblock. A Ramen `VolumeReplicationGroup` speaks exactly one per-volume replication dialect, the csi-addons `VolumeReplication` object, and simplyblock implements none of it. + +This design exposes the existing engine behind that dialect. The CSI controller plugin gains the csi-addons Replication and Identity gRPC services, a csi-addons sidecar and controller-manager reconcile `VolumeReplication` objects into those RPCs, and each RPC is a thin adapter onto a backend call that already exists (or is named in Phase 0). No replication mechanism is rebuilt, and the engine keeps shipping snapshots exactly as it does today. + +| Concern | Mechanism | Decided when | +|--------------------------------|-------------------------------------------------------------------------|-------------------------------------| +| Per-volume replication intent | `VolumeReplication` (`spec.replicationState: primary\|secondary`) | Reconciled continuously | +| Cadence, retention, and target | `VolumeReplicationClass` parameters naming a `ReplicationPolicy` | At class authoring | +| Promote, demote, and resync | The Replication gRPC verbs, adapted onto failover, demote, and failback | On each reconcile until `Completed` | +| Relationship health | `VolumeReplication.status.conditions` from the typed status read | On every reconcile | +| RPO figures | Prometheus metrics from the control plane (§11) | Continuously | + +A reader who stops here has the model: the engine is unchanged, the csi-addons surface is the adapter, and the volume handle (`{clusterID}:{poolID}:{volumeID}`) is the one identity that lets either cluster's driver address the same backend relationship. + +--- + +## 1. Background + +**The engine.** A volume replicates when `do_replicate` is set and a `ReplicationPolicy` names its cadence and target. The snapshot monitor takes interval snapshots, the shipping runner transfers each one to a landing volume on the target cluster and chains it there, and retention prunes behind the shipped frontier. Failover (`POST .../replication/failover`) clones the last fully replicated snapshot into a writable volume on the target with the same NVMe identity. Failback is a reversed replication seeded by `data_uuid` matching, and commit (`POST .../replication/commit`) runs the fenced final-step cutover through a task runner. All of this is driven today by the operator's `ReplicationPair`, `ReplicationPolicy`, `ReplicationSlot`, and `ReplicationOps` kinds, which build their HTTP calls inline against the same endpoints. + +**The contract this design targets.** Ramen's dr-cluster operator reconciles one `VolumeReplication` per protected PVC. It flips `spec.replicationState` between `primary` and `secondary` and waits for the driver's conditions (`Completed`, `Degraded`, `Resyncing`) to report the operation done and the relationship healthy. It reads `status.lastSyncTime` for RPO. It never calls a vendor API. The interfaces are the csi-addons specification's Replication gRPC, served by the driver, and the kubernetes-csi-addons controller-manager, which turns `VolumeReplication` objects into those RPCs through a per-driver sidecar. + +**The direction is already committed.** The CRD redesign excludes the four replication kinds from its model because "that subsystem is being redesigned against the CSI Addons specification, whose `VolumeReplication` and `VolumeGroupReplication` kinds already carry the per-volume and per-group replication contract that a backup tool or a DR orchestrator understands" (`crd-redesign/design-crd-model.md`, Non-Goals). This document is that redesign's first, per-volume half. The DR storage foundation gap analysis (Phase 0 and Appendix A of that document) is its requirements source. + +**Three facts about today's surface shape the design.** + +1. **The steady-state status has no home.** The typed relationship read (`GET .../replication/relationships/{lvol}` and its per-volume twin) serves an `LVolReplication` record that is only created at cutover or failover, so it returns 404 for a volume's entire healthy replicated life. The real steady-state verdict lives in `lvol_controller.get_replication_info`, an untyped dict reachable only as `rep_info` on the volume DTO. The operator's slot controller documents the consequence in its own comments: `status.lastReplicatedAt` is stamped once at attach and never refreshed while replication is healthy. +2. **There is no demote and no resync verb.** The backend surface is attach, detach, failover, failback, commit, and cutover-proceed. Demote exists only as an internal composition (ANA suspend plus fencing) inside the failover and cutover paths. Resync exists only as the failback direction reversal. +3. **The secondary has no volume object.** During steady-state replication the target cluster holds replicated snapshots and transient landing volumes, never a secondary lvol. A writable volume materializes on the target only at promotion. The `VolumeReplication` on the DR cluster therefore describes a relationship addressed by handle, not a local volume, and the driver's multi-cluster `secret.json` resolution is what makes that address work from either side. + +--- + +## 2. Goals and Non-Goals + +### Goals + +- A `VolumeReplication` object per PVC, reconciled by the stock kubernetes-csi-addons controller-manager against this driver, drives simplyblock replication: enable, disable, promote (planned and forced), demote, and resync. +- The driver's conditions follow the csi-addons contract: `Completed` reports the last requested state change finished, `Degraded` reports relationship health, and `Resyncing` reports a divergence catch-up in flight. Steady-state healthy async is `Completed=True, Degraded=False, Resyncing=False`, and lag alone trips none of them. +- `status.lastSyncTime` is truthful for the volume's whole replicated life, sourced from a typed backend status read rather than the cutover-time relationship record. +- Every verb is idempotent, because Ramen re-drives every reconcile. +- The adapter reuses one shared control-plane client (`atlas-lib/controlplane`), ending the pattern where each consumer hand-rolls the same replication HTTP calls. +- A `StorageClass` and `VolumeSnapshotClass` naming convention across paired clusters that Ramen's peerClasses can express, with a preflight that verifies it. +- The RPO and backlog figures Ramen cannot carry (`bytesBehind`, throughput, RPO compliance) are exported as Prometheus metrics from the control plane. + +### Non-Goals + +- **`VolumeGroupReplication`.** Continuous group promote and demote on top of consistency groups is the next design. This one is strictly per volume. The group snapshot surface (`design-consistency-groups.md`) is untouched. +- **VolSync and the S3 backup path.** The recurrent-immutable-snapshots DR type is a separate phase of the gap analysis and does not pass through this adapter. +- **Synchronous replication.** The engine is asynchronous snapshot shipping, and nothing here changes that. +- **Ramen hub components.** DRPolicy, DRPC, and hub orchestration are consumers of this contract, not part of it. +- **Retiring the legacy replication kinds.** `ReplicationPair`, `ReplicationPolicy`, `ReplicationSlot`, and `ReplicationOps` keep working unchanged. §8 defines coexistence and §13 the consolidation direction. Removal is its own change once the adapter is proven. +- **A simplyblock DR-status CRD.** The observability surface here is metrics. A rollup CRD is future work in the gap analysis's Appendix B. + +--- + +## 3. Architecture Overview + +``` +┌───────────────────────────────────────────────────────────────────────────┐ +│ Kubernetes (each cluster) │ +│ │ +│ VolumeReplicationClass VolumeReplication (one per protected PVC) │ +│ (names a ReplicationPolicy) spec.replicationState: primary|secondary │ +│ │ │ │ +│ ▼ ▼ │ +│ ┌─────────────────────────────────────────────────────────────────────┐ │ +│ │ kubernetes-csi-addons controller-manager (chart-deployed) │ │ +│ │ reconciles VolumeReplication → Replication gRPC on the driver │ │ +│ └───────────────────────────────┬─────────────────────────────────────┘ │ +│ │ via CSIAddonsNode registration │ +│ ┌───────────────────────────────▼─────────────────────────────────────┐ │ +│ │ CSI controller StatefulSet (SimplyblockDriver reconciler) │ │ +│ │ … csi-provisioner │ csi-snapshotter │ … │ csi-addons sidecar │ │ │ +│ │ controller plugin socket │ │ +│ │ plugin serves: CSI Identity/Controller/GroupController │ │ +│ │ + csi-addons Identity + Replication (this design) │ │ +│ └───────────────────────────────┬─────────────────────────────────────┘ │ +└──────────────────────────────────┼─────────────────────────────────────────┘ + │ HTTP, resolved per volume handle +┌──────────────────────────────────▼─────────────────────────────────────────┐ +│ simplyblock control plane │ +│ PUT .../volumes/{v} {replication_policy_id} (attach/ │ +│ detach) │ +│ GET .../volumes/{v}/replication/status (P0-1, steady state) │ +│ POST .../volumes/{v}/replication/failover (promote, planned or │ +│ forced) │ +│ POST .../volumes/{v}/replication/demote (P0-3) │ +│ POST .../volumes/{v}/replication/failback (resync) │ +│ GET .../replication/relationships/{lvol} (cutover records) │ +└──────────────────────────────────────────────────────────────────────────────┘ +``` + +**One relationship, addressed from either cluster.** The `VolumeReplication.spec.dataSource` resolves to a PV whose handle is `{clusterID}:{poolID}:{volumeID}`. The driver resolves the backend from the handle through `clusters.Client`, exactly as every other RPC does, so the DR cluster's driver can drive the same backend relationship as the source cluster's without either side holding special state. This is the same stateless addressing the node plugin already uses to redirect a staged volume after failover. + +**The adapter holds no state.** csi-addons RPCs are stateless and idempotent by contract. Every answer the driver gives is derived on the spot from the backend status read and the relationship record. There is no driver-side cache, no persisted step, and no state machine. A verb whose backend work outlives the call (a demote converging a busy peer) reports `Completed=False` until the backend reflects the target state, and Ramen's re-drive is the retry loop. + +--- + +## 4. The csi-addons Machinery + +### 4.1 What gets deployed + +- **CRDs** (chart): `volumereplications.replication.storage.openshift.io`, `volumereplicationclasses.replication.storage.openshift.io`, and the kubernetes-csi-addons operational CRDs (`csiaddonsnodes.csiaddons.openshift.io`). Vendored at a pinned upstream version, the same way the `VolumeGroupSnapshot` CRDs are. +- **Controller-manager** (chart): the stock kubernetes-csi-addons manager. It discovers driver endpoints through `CSIAddonsNode` objects and reconciles `VolumeReplication` by calling the driver's Replication gRPC. +- **Sidecar** (operator, `SimplyblockDriver` reconciler): the csi-addons sidecar in the controller StatefulSet. It connects to the plugin's socket, probes the csi-addons Identity service for capabilities, publishes a `CSIAddonsNode`, and proxies the manager's gRPC to the plugin. Wiring follows the existing sidecar pattern: a seventh field on `spec.sidecarImages`, a pinned default in `sidecars.go`, an appended `controllerSidecar` entry in `workloads.go` (appended, because the per-container tweaks there are positional), and a new RBAC component whose group alias is `replication.storage.openshift.io` plus `csiaddons.openshift.io`. + +### 4.2 What the plugin serves + +The plugin registers two additional gRPC services on the existing socket, beside the CSI services, following the GroupController precedent in `csicommon.NonBlockingGRPCServer`: + +- **csi-addons Identity:** `GetIdentity`, `GetCapabilities` (advertising `VOLUME_REPLICATION`), and `Probe`. This is a distinct service from CSI Identity, so it is a new small server type, not an extension of the existing one. +- **Replication:** the six verbs of §5, implemented on the controller `Server` through an embedded `UnimplementedReplicationServer` from `github.com/csi-addons/spec`, mirroring how the GroupController embeds its unimplemented base. + +`NonBlockingGRPCServer.Start` today takes exactly the three CSI servers. It gains a registration hook so the driver package can register additional services without `csicommon` importing csi-addons. + +### 4.3 The shared client + +The generated control-plane client already declares every replication endpoint, but it is `internal/` to atlas-lib. This design adds `atlas-lib/controlplane/replication.go`, a typed wrapper over those operations (attach, detach, status, failover, demote, failback, relationship, and commit for the legacy `ReplicationOps` path), following the shape `migrations.go` set for long-running backend operations. The driver adapter consumes it, and the operator's replication reconcilers move onto it opportunistically (§13), ending the three-copies-of-the-same-HTTP-call pattern the gap analysis lists as an inconsistency. + +--- + +## 5. The Replication Service + +`volume_id` on every request is the CSI volume handle. The driver parses it with the existing helper and resolves the cluster client from `secret.json`. Every verb is idempotent: repeating a completed operation returns success, and repeating an in-flight one returns the in-flight answer. Backend errors map to gRPC codes through the existing per-RPC classifier, extended with replication constructors. + +### 5.1 Phase 1 verbs + +| Verb | Backend mapping | Semantics | +|----------------------------|------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| `EnableVolumeReplication` | `PUT .../volumes/{v}` body `{"replication_policy_id": }` | The policy comes from the `VolumeReplicationClass` parameters (§7). Attach is synchronous on the backend; already-attached to the same policy is success (P0-2). Attached to a *different* policy is `FAILED_PRECONDITION`, because silently re-attaching forces a full re-sync. | +| `DisableVolumeReplication` | `PUT .../volumes/{v}` body `{"replication_policy_id": null}` | Detach. A 409 (cutover in flight) maps to `ABORTED`, retryable. Not-attached is success (P0-2). | +| `GetVolumeReplicationInfo` | `GET .../volumes/{v}/replication/status` (P0-1) | Returns `lastSyncTime` (newest fully replicated snapshot's creation time), `lastSyncDuration` (last cycle duration), and `lastSyncBytes` (last shipped snapshot's used size). | + +### 5.2 Phase 2 verbs + +| Verb | Backend mapping | Semantics | +|------------------------------------|-------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| `PromoteVolume` with `force=true` | `POST .../volumes/{v}/replication/failover` | Unplanned promote: clone the last fully replicated generation on the target, retire the (assumed dead) source. Idempotent by the backend's own NQN probe. `Completed=True` once the relationship reads `failed_over`. | +| `PromoteVolume` with `force=false` | `POST .../volumes/{v}/replication/failover` with the planned gate | The same promote as the forced form, refused unless the peer already holds every acknowledged write, which is exactly the state a completed demote leaves behind: the source fenced and the final flush confirmed. A planned promote against an undemoted or lagging source surfaces as `FAILED_PRECONDITION`. `Completed=True` once the relationship reports the target active. | +| `DemoteVolume` | `POST .../volumes/{v}/replication/demote` (P0-3) | Quiesce (ANA inaccessible), ship a final internal snapshot, wait until it carries the replicated marker, fence the data path, and record the role. This is the lossless half of a planned swap: after demote, the peer's planned promote loses nothing. `Completed=True` only when the final flush has landed on the peer. | +| `ResyncVolume` | `POST .../volumes/{v}/replication/failback` | Reverse the shipping direction, seeded by `data_uuid` matching so a recovered source resyncs by delta rather than full copy. `Resyncing=True` while the catch-up runs; it clears when the reverse lag is inside the budget. Resync reconciles the diverged old primary from the current primary; it never merges. | + +**Promote is one operation; planned and unplanned differ only in what precedes it.** Both forms clone the last fully replicated generation on the target and serve it under the preserved NVMe identity, exactly as failover does today. The planned form is lossless not because it runs different machinery but because a completed demote guarantees the last replicated generation contains every acknowledged write, and the planned gate refuses the promote until that holds. The forced form skips the gate and accepts the RPO loss, because its premise is that the source is gone. The engine's commit cutover (`replication/commit`, the `FN_REPLICATION_FINAL` runner with its shrink rounds and cutover-proceed handshake) is deliberately NOT part of this contract: its convergence job moves into the demote verb (P0-3), and it remains only behind the legacy `ReplicationOps` migration path (§8, §13). + +--- + +## 6. Steady-State Status and Conditions + +### 6.1 The typed status read (P0-1) + +`GET .../volumes/{v}/replication/status` serves, as a typed DTO, what `get_replication_info` computes today: + +``` +role: source | secondary | failed_over | none +state: in_sync | replicating | lagging | degraded | error | not_replicating +last_replicated_at: timestamp of the newest fully replicated snapshot +lag_seconds: now - last_replicated_at +lag_budget_seconds: derived budget (or the policy's rpo_target_seconds, P0-4) +outstanding_count: snapshots queued but not yet shipped +outstanding_bytes: sum of used_size over the outstanding snapshots +failing_count: shipping tasks currently suspended on errors +max_retry_reached: whether any task exhausted its retries +last_cycle_bytes: used_size of the last shipped snapshot +last_cycle_seconds: duration of the last shipping cycle +resyncing: whether a failback catch-up is in flight +``` + +This endpoint exists for the volume's whole replicated life. It does not replace the relationship read: `GET .../replication/relationships/{lvol}` keeps its cutover-record semantics (including surviving source deletion, which the node redirect depends on), and the status read is the steady-state complement. The operator's slot controller can also poll it to keep `status.lastReplicatedAt` honest, independent of this design's adapter. + +### 6.2 Condition mapping + +Following the csi-addons contract: conditions describe relationship and operation health, never instantaneous RPO. Steady-state healthy async is `Completed=True, Degraded=False, Resyncing=False`, and lag alone trips none of them. + +| Condition | True when | simplyblock source | +|-------------|------------------------------------------------------------------------|-----------------------------------------------------------------------------------------------------------------------------------------------------| +| `Completed` | The last requested state change (enable, promote, demote) has finished | Enable: attach returned. Promote (either form): the relationship reports the target active. Demote: the P0-3 verb confirmed the final flush landed. | +| `Degraded` | The relationship is unhealthy or not progressing | `state` is `degraded` (suspended shipping tasks) or `error` (retries exhausted), or `lag_seconds` exceeds `lag_budget_seconds`. | +| `Resyncing` | A divergence catch-up is reconciling the secondary | The status read's `resyncing` flag: a failback direction reversal or a `FN_REPLICATION_FINAL` task in flight. | + +`status.state` mirrors `spec.replicationState` once `Completed=True`, from the status read's `role`. + +--- + +## 7. VolumeReplicationClass and peerClasses + +### 7.1 The class + +A `VolumeReplicationClass` binds a `VolumeReplication` to a backend policy: + +```yaml +apiVersion: replication.storage.openshift.io/v1alpha1 +kind: VolumeReplicationClass +metadata: + name: simplyblock-async-5m +spec: + provisioner: csi.simplyblock.io + parameters: + replicationPolicy: dr-policy-5m + schedulingInterval: 5m +``` + +`replicationPolicy` names the backend `ReplicationPolicy` (resolved per cluster by name), which owns cadence, retention, mode, and the replication target. `schedulingInterval` restates the policy's interval for Ramen's `DRPolicy` matching, and the preflight (§7.2) checks the two agree. No secrets parameter is needed: the driver's credentials come from `secret.json`, as for every other RPC. The chart ships no default class. Classes are the user's to author, matching the `VolumeGroupSnapshotClass` decision. + +### 7.2 peerClasses convention and preflight + +Two of the requirements below are Ramen's contract, and the rest is this design's convention; the split matters because only the contract can fail a DRPolicy. + +**Contract (Ramen requires this):** + +- **The `StorageClass` name exists on both clusters.** A failover restores the protected PVC with its original `spec.storageClassName`, a by-name reference, so an equivalent class of that exact name must exist on the peer. The peerClasses computation also joins classes across the two clusters by `StorageClass` name. +- **The Ramen identity labels.** Each cluster's `StorageClass` carries `ramendr.openshift.io/storageid` (differing per cluster, since the backends differ), and the two `VolumeReplicationClass` objects representing one relationship carry an equal `ramendr.openshift.io/replicationid`. The `VolumeReplicationClass` is selected per cluster by `spec.replicationClassSelector` labels plus `provisioner` and a `schedulingInterval` equal to the `DRPolicy`'s, never by name. + +**Convention (this design chooses it for operability):** + +- **Same `StorageClass` parameters on both clusters** apart from `cluster_id` (necessarily) and pool when pools differ. The replication target's pool mapping already handles the pool difference at shipping time. +- **Same `VolumeSnapshotClass` and `VolumeReplicationClass` names on both clusters**, each side's `replicationPolicy` naming that cluster's policy toward its peer. Ramen does not require the names to match, but one name per relationship is what keeps a fleet legible. + +The preflight is a check, not a controller: a validation that runs on demand (and on `ReplicationPair` reconciliation) confirming that for each replication-enabled `StorageClass` the peer cluster has a same-named class, and that the named backend policies exist and point at each other's clusters. Its findings surface as events on the `ReplicationPair`, which is the object that already models the cluster pairing. + +--- + +## 8. Coexistence with the Legacy Replication Kinds + +Two control paths now reach the same backend relationship: the annotation-driven slot family, and `VolumeReplication`. They must not fight. + +- **One owner per volume.** A volume is managed either by the annotation path (a `ReplicationSlot` exists for its PVC) or by a `VolumeReplication`, never both. The adapter's `EnableVolumeReplication` refuses (`FAILED_PRECONDITION`) whenever a slot manages the volume, even when the slot's policy and the class's agree: ownership is the conflict, not the policy value, because two owners turn every ownership flap into a detach and re-attach, and a re-attach is a full re-sync on the backend. The `PVCAnnotationWatcher` is taught to skip PVCs that have a `VolumeReplication` (one informer lookup), so the annotation cannot re-attach behind the adapter's back. Moving a volume between the two paths is deliberate: remove the annotation (the slot detaches), then create the `VolumeReplication`. +- **`ReplicationOps` stays imperative and internal.** Bulk failover of a whole policy or target remains a `ReplicationOps` concern until `VolumeGroupReplication` lands. Nothing in this design calls it, and it does not touch csi-addons-managed volumes because the mutual-exclusion rule above keeps the sets disjoint. +- **The slot's status problem is fixed as a side effect.** The typed status read (P0-1) gives the slot controller a truthful `lastReplicatedAt` source, whether or not the adapter is in use. + +The consolidation direction (§13) is that the annotation path becomes a compatibility layer and the four kinds retire once Ramen-driven replication is proven, exactly as the CRD redesign anticipates. + +--- + +## 9. Backend API Requirements + +Every endpoint is scoped as today: volume-scoped under `/api/v2/clusters/{c}/storage-pools/{p}/volumes/{v}`, cluster-scoped under `/api/v2/clusters/{c}/replication`. + +| Method | Endpoint | Notes | +|--------|---------------------------------------------------------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| `PUT` | `.../volumes/{v}` (`replication_policy_id`) | Existing attach and detach. P0-2 makes both idempotent: same-policy attach and non-attached detach return success. | +| `GET` | `.../volumes/{v}/replication/status` | **New (P0-1).** The typed steady-state status of §6.1. Never 404s for a volume that exists; `state: not_replicating, role: none` is a valid answer. | +| `POST` | `.../volumes/{v}/replication/failover` (+ planned gate) | Existing; the one promote, both forms. Gains a planned form that is refused unless the peer holds every acknowledged write (a completed demote). Idempotent by NQN probe. | +| `POST` | `.../volumes/{v}/replication/demote` | **New (P0-3).** Quiesce, final ship, confirm on peer, fence. Idempotent: demoting a demoted volume returns success. | +| `POST` | `.../volumes/{v}/replication/failback` | Existing. Resync (direction reversal, delta-seeded). | +| `POST` | `.../volumes/{v}/replication/commit` | Existing, unchanged, and NOT part of this contract: it stays behind the legacy `ReplicationOps` migration path only (§13). | +| `GET` | `.../replication/relationships/{lvol}` | Existing, unchanged. Cutover records only; the node redirect depends on its survive-deletion semantics. | +| `GET` | `.../replication/relationships/{lvol}/latest-snapshot` | **New (P0-6).** The newest fully replicated snapshot for the volume, as a cloneable snapshot handle. A consistency-group form returns one complete generation's member snapshots. Exposes what the failover path already computes internally. | +| `POST` | `.../volumes/{v}/replication/cutover-proceed` | Existing, unchanged, legacy path only: the adapter never reaches it, because the commit cutover is off this contract. | + +The unused backend verbs the operator never calls (`start`, `stop`, `trigger`, `tasks`) are unaffected, and `start` and `stop` remain the policy-less legacy path. + +--- + +## 10. Failure Modes and Fallback + +| Failure | Detection | Behavior | +|--------------------------------------------------------------------------|---------------------------------------------------------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| Enable against a volume attached to a different policy | Adapter compares the status read's policy | `FAILED_PRECONDITION`; the message names both policies. Never silently re-attaches, because that is a full re-sync. | +| Disable during a cutover | Backend 409 | `ABORTED`, retryable. Ramen re-drives; the detach succeeds after the cutover settles. | +| Planned promote without a completed demote (source live or peer lagging) | The planned gate on the failover endpoint refuses | `FAILED_PRECONDITION` naming the un-flushed tail. The caller demotes (or resyncs) first, and `Resyncing` reports progress. | +| Forced promote when the source is alive | Backend fences the source data path as part of failover | Split-brain is prevented structurally: the source is retired (ANA inaccessible, subsystem removed) before the target serves. | +| Backend unreachable | HTTP error from the shared client | The RPC returns `UNAVAILABLE`; the controller-manager retries. Conditions keep their last-observed values, so a blip does not flap `Degraded`. | +| Demote cannot complete its flush (peer slow or unreachable) | The demote's confirm step times out | `Completed` stays `False`, and the volume is not left fenced without a decision: whether the backend rolls back to serving primary or holds quiesced is Open Question 1. | +| Status read missing (pre-P0-1 backend) | 404 from the status endpoint | The csi-addons capability is not advertised, so the controller-manager never drives this driver. The adapter ships dark until the backend is current. | +| Both an annotation slot and a VolumeReplication claim a volume | Adapter check plus watcher skip (§8) | The first owner wins; the second surfaces `FAILED_PRECONDITION` (adapter) or a skip event (watcher). Never two writers to one relationship. | + +--- + +## 11. Observability + +**Baseline.** The replication engine's telemetry today is log lines (the `XFER-TIMING` phases) and the `rep_info` dict. Nothing is exported as metrics, and the `VolumeReplication` surface adds only `lastSyncTime`. The figures an operator actually watches live here. + +### Kubernetes Events + +The kubernetes-csi-addons controller-manager owns events on `VolumeReplication` (promote, demote, and resync outcomes), and this design adds none there. The operator emits preflight findings on the `ReplicationPair`: + +| Event | Type | Emitted when | +|-----------------------|---------|-----------------------------------------------------------------------------------------------------------------------------| +| `PeerClassesVerified` | Normal | The preflight confirmed same-named classes and mutually pointing policies on both clusters | +| `PeerClassesMismatch` | Warning | A replication-enabled class has no same-named peer, or the named policies do not pair; the message names the class and side | + +### Prometheus Metrics + +Exported by the control plane, labeled `volume`, `policy`, and `peer_cluster`: + +| Metric | Labels | Description | +|---------------------------------------------|------------------------------|---------------------------------------------------------------------------------| +| `simplyblock_replication_lag_seconds` | volume, policy, peer_cluster | Now minus the newest fully replicated snapshot's creation time. | +| `simplyblock_replication_backlog_bytes` | volume, policy, peer_cluster | `outstanding_bytes`: the queued-but-unshipped snapshot sizes. | +| `simplyblock_replication_last_sync_seconds` | volume, policy, peer_cluster | Duration of the last shipping cycle. | +| `simplyblock_replication_last_sync_bytes` | volume, policy, peer_cluster | Size of the last shipped snapshot. | +| `simplyblock_replication_rpo_violation` | volume, policy, peer_cluster | 1 while `lag_seconds` exceeds the policy's `rpo_target_seconds` (P0-4), else 0. | +| `simplyblock_replication_degraded` | volume, policy, peer_cluster | 1 while the status read's state is `degraded` or `error`. | + +`simplyblock_replication_rpo_violation` is the alert: it is the declared objective against the measured lag, which no timestamp alone can express. `simplyblock_replication_backlog_bytes` is the second load-bearing figure, because it is the input to any honest RTO estimate and the number that distinguishes "slow cycle" from "falling behind." One caveat is recorded rather than hidden: `outstanding_bytes` measures queued snapshot sizes, not dirty bytes written since the last snapshot, so intra-interval writes are invisible to it. A true dirty-delta figure needs storage-plane support and is future work. + +--- + +## 12. Testing Strategy + +Full scenario matrix and coverage status: [`tests/test-plan-csi-addons-replication.md`](../tests/test-plan-csi-addons-replication.md) + +- **Unit (driver):** each verb against a mock control plane: the idempotency table (repeat enable, repeat disable, repeat promote), the refusal paths (different-policy enable, lagging planned promote, disable during cutover), the condition derivation from every status-read state, and handle parsing failures. +- **Unit (operator):** the peerClasses preflight against fake clients for both clusters, and the `PVCAnnotationWatcher` skip when a `VolumeReplication` exists. +- **Integration:** the csi-addons sidecar and controller-manager against the driver with a mock backend under envtest or kind: a `VolumeReplication` flipped `primary` to `secondary` and back walks the verbs in order and lands the conditions. +- **E2E (two live clusters):** the Ramen-shaped lifecycle without Ramen: enable on the source, write data, and verify `lastSyncTime` advances; forced promote on the DR side, verifying the clone serves with the source fenced; and resync back with a planned swap (demote then promote), verifying zero loss with a hashed writer. Then the same driven by an actual Ramen VRG in async mode, which is Phase 2's acceptance gate. + +The risk concentrates in the demote verb's flush-confirmation (the lossless-swap guarantee), in idempotency under Ramen's aggressive re-drive, and in the coexistence rules of §8. Those must not be cut. + +--- + +## 13. Migration Strategy + +Three replication control surfaces exist today: the operator's kinds, the stale CSI-shipped CRDs (`replications.simplyblock.com`, invalid as written, plus the orphaned `SnapshotReplication`), and the backend-driven node redirect. The target is one: csi-addons, with the node redirect unchanged beneath it. + +1. **Phase 1 and 2 (this design):** the adapter ships alongside the legacy kinds. The mutual-exclusion rule (§8) keeps the two paths disjoint per volume. The shared `atlas-lib/controlplane/replication.go` client lands, and the driver uses it from day one. +2. **Opportunistic:** the operator's replication reconcilers move their inline HTTP calls onto the shared client, and the slot controller sources `lastReplicatedAt` from the typed status read. No behavior change, one client. +3. **After Ramen validation:** the annotation path is declared a compatibility layer, and new volumes are protected through `VolumeReplication`. The stale CSI-shipped CRDs are deleted (they were never installable), and `SnapshotReplication` is retired from the charts. +4. **The redesign's replication chapter:** whether `ReplicationPair` and `ReplicationPolicy` survive as the backend-policy authoring surface (classes need policies to name) or are re-cut is decided there, not here. This design only requires that a policy exists per cluster pair, however it is authored. + +--- + +## 14. Test Failover + +A DR drill proves that failover works without disturbing production replication. Ramen cannot drive one: its two actions, `Failover` and `Relocate`, move the workload for real. The drill is therefore driven by a simplyblock-native kind (owned by the SiteMap and testing layer, and specified there, not here), and this section defines the storage contract that kind consumes. There are two modes over one substrate: a bubble on the secondary cluster, and a separate test cluster reached like the real failover. + +### 14.1 The substrate: replicated snapshots are cloneable test points + +The replicated snapshots on the secondary are first-class snapshot records on the secondary's own control plane, chained and complete, and for a consistency group they carry the `group_id` and `group_seq` provenance of their generation. The failover path already resolves the newest fully replicated snapshot per volume, and the group-wide resolution picks one complete generation across members. P0-6 exposes that resolution as a read (§9), so a drill can address its test point without reimplementing the selection logic. + +Every object a drill creates, on either side, carries a `drtest-` name prefix and a test-id label. Leftovers are then enumerable, and a teardown can prove completeness instead of assuming it. + +### 14.2 Bubble mode: same cluster, different namespace + +The drill namespace lives on the secondary Kubernetes cluster, whose driver already talks to the backend holding the replicated snapshots, so no data moves at all: + +1. Resolve the test point through P0-6: per volume the newest replicated snapshot, or for a group one complete `group_seq`, so the bubble starts from a single crash-consistent cut. +2. Surface each snapshot as a pre-provisioned `VolumeSnapshotContent` (the handle is the secondary-side snapshot), bind a `VolumeSnapshot` in the drill namespace, and clone it into a PVC through the ordinary `dataSource` path. +3. Deploy the application against the clones. The clones are thin, independent volumes, and writes to them never touch the replication stream. +4. Tear down by deleting the namespace, then enumerate by the test-id label to prove nothing leaked. + +### 14.3 Test-cluster mode: the real failover, aimed at expendable volumes + +The second mode reaches a separate storage cluster, and it is deliberately composed from primitives this design already relies on rather than a new shipping capability: + +1. Clone the resolved test point into `drtest-` volumes on the secondary. For a group, clone one generation into a new `drtest-` consistency group, so the group failover path is exercised too. +2. Attach the clones to a `drtest-` replication policy whose `ReplicationTarget` is the test cluster. The ordinary engine ships them (a full copy, since the test backend shares no ancestry). +3. On the test cluster, run the real failover against the shipped volumes to materialize writable clones, and deploy the application there. + +The property this buys is fidelity: the drill exercises the actual failover machinery, target resolution, clone-from-replicated, identity preservation, and the driver redirect, against a third cluster, while the production relationship is never touched, because the `drtest-` policy is a separate policy with its own target. + +### 14.4 The non-disruption proof + +A drill that silently perturbed replication would be worse than no drill. Before the first clone and after the teardown, the driving kind captures and compares: every production `VolumeReplication`'s conditions, the `lastSyncTime` cadence, the lag and backlog from the typed status read (P0-1), and the count of production `VolumeReplication` objects. Any drift fails the drill as an invariant violation. The status endpoint built for Ramen's conditions is the same instrument this audit reads, which is why the drill contract belongs in this design. + +### 14.5 Costs and bounds + +- **A live clone pins its base snapshot.** Retention defers pruning a snapshot with a dependent clone, which is what keeps the drill safe, and also why a drill must carry a maximum lifetime: a long-lived bubble holds the secondary's replicated chain back. +- **Test-cluster mode consumes real resources:** cross-cluster bandwidth for the full copy, and capacity on both the secondary (the `drtest-` clones) and the test cluster. The `drtest-` clones on the secondary exist only as replication sources and are never served; whether they are created as internal volumes (invisible to normal listings, exempt from subsystem caps, like landing volumes) is Open Question 4. +- **The group rules apply unchanged:** a `drtest-` consistency group observes the member cap and the placement pin like any other, so a drill of a large group is a capacity event on the secondary. + +--- + +## 15. Open Questions + +| # | Question | Owner | +|-----|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------| +| 1 | **Demote semantics for the application.** The P0-3 demote fences the volume (ANA inaccessible) after the final flush, and with convergence folded into the verb it is now the only place a planned swap can stall. Ramen relocation unmounts the workload first, so the fence is ordinarily unopposed. Confirm the verb's behavior when writes are still in flight at quiesce (block versus fail), whether the converge phase has its own budget separate from the quiesced flush, and whether a timeout in either phase must abort back to serving primary. | Backend team | +| 2 | **Where the preflight lives.** §7.2 attaches peerClasses validation to the `ReplicationPair` reconciler. If the redesign retires the pair kind, the preflight needs a new home (the `SimplyblockDriver`, or a standalone check job). | Operator team | +| 3 | **Per-volume policy granularity.** A `VolumeReplicationClass` names one policy, and today one policy implies one target and cadence for all its volumes. Confirm one class per (policy, cadence) is an acceptable authoring model for Ramen's `replicationClassSelector`, or whether per-volume interval overrides are needed. | Operator / Backend team | +| 4 | **Visibility of `drtest-` clones.** The test-cluster mode's clones on the secondary are replication sources only, never served. Decide whether the backend creates them as internal volumes (hidden from listings, exempt from the per-node subsystem cap, like the shipping path's landing volumes) or as ordinary volumes under a naming convention. | Backend team | diff --git a/operator/docs/tests/test-plan-csi-addons-replication.md b/operator/docs/tests/test-plan-csi-addons-replication.md new file mode 100644 index 000000000..a9ed49411 --- /dev/null +++ b/operator/docs/tests/test-plan-csi-addons-replication.md @@ -0,0 +1,175 @@ +# Test Plan: csi-addons Volume Replication + +Related design: [`designs/design-csi-addons-replication.md`](../designs/design-csi-addons-replication.md) + +Scope is the CSI driver's Replication service, the operator's preflight and coexistence rules, and the deployment of the csi-addons machinery. The replication engine itself (snapshot shipping, failover cloning, the cutover task runner) is the control plane's to prove and is exercised here only through the adapter's boundary. The kubernetes-csi-addons controller-manager is stock upstream and is not re-tested; what is tested is this driver's conformance to the contract it drives. + +Scenario IDs are permanent and are never reused or renumbered. `U-` is unit (no cluster: mock control plane, fake `client.Client`), `I-` is integration (the sidecar and controller-manager against the driver with a mock backend), `E-` is end-to-end (two live simplyblock clusters), and `M-` is manual. Types are `Positive`, `Negative`, `Boundary`, and `Regression`. A `—` in the `Test` column means nothing implements the scenario yet, and every such row reappears in §6 with its reason. + +The plan is the target coverage for a Draft design: every row is `—` until the work lands. + +--- + +## 1. Unit Tests + +The Replication service against a mock control plane, and the operator pieces against fake clients. No Kubernetes API server. Numbering runs continuously across the groups. + +### Replication Verbs: Enable, Disable, Info (design §5.1) + +File: `csi-driver/internal/csi/controller/replication_test.go` (planned) + +| # | Scenario | Type | Test | +|------|---------------------------------------------------------------------------------------------------------------|----------|------| +| U-01 | Enable on an unattached volume: the attach call carries the class's policy, and the RPC succeeds | Positive | — | +| U-02 | Enable on a volume already attached to the same policy: success, no second attach call (idempotency) | Boundary | — | +| U-03 | Enable on a volume attached to a different policy: `FAILED_PRECONDITION` naming both policies, no attach call | Negative | — | +| U-04 | Disable on an attached volume: the detach call is made, success | Positive | — | +| U-05 | Disable on a non-attached volume: success without a backend call (idempotency) | Boundary | — | +| U-06 | Disable while a cutover is in flight (backend 409): `ABORTED`, retryable | Negative | — | +| U-07 | Info returns `lastSyncTime`, `lastSyncDuration`, and `lastSyncBytes` from the status read | Positive | — | +| U-08 | A malformed volume handle: `INVALID_ARGUMENT` before any backend call | Negative | — | +| U-09 | Backend unreachable: `UNAVAILABLE`, and no condition flap is implied by the error | Negative | — | + +### Replication Verbs: Promote, Demote, Resync (design §5.2) + +File: `csi-driver/internal/csi/controller/replication_lifecycle_test.go` (planned) + +| # | Scenario | Type | Test | +|------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------|----------|------| +| U-10 | Forced promote: the failover endpoint is called, and `Completed` follows the relationship reaching `failed_over` | Positive | — | +| U-11 | Forced promote repeated after completion: success without a second failover (backend idempotency honored) | Boundary | — | +| U-12 | Planned promote after a completed demote: the failover endpoint is called with the planned gate, `Completed=True` once the relationship reports the target active | Positive | — | +| U-13 | Planned promote without a completed demote (source live or peer lagging): the planned gate refuses, `FAILED_PRECONDITION` naming the un-flushed tail | Negative | — | +| U-14 | Demote: the demote endpoint is called, and `Completed` is `True` only after the flush-confirmed response | Positive | — | +| U-15 | Demote repeated on a demoted volume: success (idempotency) | Boundary | — | +| U-16 | Resync: the failback endpoint is called and `Resyncing=True` while the status read reports the catch-up | Positive | — | +| U-17 | Resync completion: `Resyncing` clears when the reverse lag is inside the budget | Boundary | — | + +### Condition Derivation (design §6.2) + +File: `csi-driver/internal/csi/controller/replication_conditions_test.go` (planned) + +| # | Scenario | Type | Test | +|------|-----------------------------------------------------------------------------------------------------------------|----------|------| +| U-18 | Status `in_sync` and `replicating`: `Completed=True, Degraded=False, Resyncing=False` (lag alone trips nothing) | Positive | — | +| U-19 | Status `degraded` (suspended tasks) and `error` (retries exhausted): `Degraded=True` with the failure detail | Negative | — | +| U-20 | Lag beyond the budget with healthy shipping: `Degraded=True` on the staleness rule | Boundary | — | +| U-21 | Status read reports `resyncing`: `Resyncing=True`; cleared on the next read without it | Positive | — | +| U-22 | Status `not_replicating`, role `none`: reported as disabled, no `Degraded` | Boundary | — | + +### Operator: Preflight and Coexistence (design §7.2, §8) + +Files: `operator/internal/controller/peerclasses_preflight_test.go`, `operator/internal/controller/pvcreplication_controller_test.go` (planned) + +| # | Scenario | Type | Test | +|------|-------------------------------------------------------------------------------------------------------------------------------------------------------------|----------|------| +| U-23 | Same-named classes on both clusters with mutually pointing policies: `PeerClassesVerified` on the pair | Positive | — | +| U-24 | A replication-enabled class missing on the peer: `PeerClassesMismatch` naming the class and side | Negative | — | +| U-25 | Named policies exist but do not point at each other's clusters: `PeerClassesMismatch` | Negative | — | +| U-26 | `PVCAnnotationWatcher` skips a PVC whose volume has a `VolumeReplication`: no slot is created, a skip is recorded | Negative | — | +| U-27 | Annotation added and later a `VolumeReplication` appears: the existing slot is not deleted by the adapter, and the enable is refused per the one-owner rule | Boundary | — | + +--- + +## 2. Integration Tests + +The csi-addons sidecar and controller-manager against the driver with a mock control plane, on a kind or envtest-backed cluster with the vendored CRDs installed. + +### VolumeReplication Lifecycle (design §4, §5) + +| # | Scenario | Type | Test | +|------|---------------------------------------------------------------------------------------------------------------------------------------------------|----------|------| +| I-01 | The sidecar probes the driver, advertises `VOLUME_REPLICATION`, and a `CSIAddonsNode` is published | Positive | — | +| I-02 | Creating a `VolumeReplication` with `replicationState: primary` enables replication and the conditions settle at `Completed=True, Degraded=False` | Positive | — | +| I-03 | Flipping to `secondary` drives demote, and back to `primary` drives promote, each waiting on `Completed` | Positive | — | +| I-04 | Deleting the `VolumeReplication` disables replication | Positive | — | +| I-05 | A `VolumeReplicationClass` naming a nonexistent policy: enable fails, the condition carries the message, and the object retries without flapping | Negative | — | +| I-06 | `status.lastSyncTime` advances across reconciles while the mock backend advances its newest replicated snapshot | Positive | — | +| I-07 | Driver restart mid-operation: the re-driven verb is idempotent and the conditions re-derive without manual repair | Boundary | — | + +--- + +## 3. E2E Tests + +Two live simplyblock clusters with the chart-deployed csi-addons machinery. The Ramen VRG rows are the Phase 2 acceptance gate. + +### Adapter Lifecycle (design §5, §6) + +| # | Scenario | Type | Test | +|------|-------------------------------------------------------------------------------------------------------------------------------------------|------------|------| +| E-01 | Enable through a `VolumeReplication`, write data, and `lastSyncTime` advances at the policy cadence | Positive | — | +| E-02 | Forced promote on the DR cluster: the clone serves with the source fenced, and data matches the last replicated generation | Positive | — | +| E-03 | Resync back after a forced promote, then a planned swap (demote, then planned promote): a hashed writer proves zero loss across the swap | Positive | — | +| E-04 | Planned promote refused while lagging, succeeds after `Resyncing` clears | Negative | — | +| E-05 | The typed status read never 404s across the whole lifecycle (enable through swap), and the slot controller's `lastReplicatedAt` tracks it | Regression | — | + +### Ramen-Driven (design §12, Phase 2 gate) + +| # | Scenario | Type | Test | +|------|----------------------------------------------------------------------------------------------------------------------------------------------|----------|------| +| E-06 | A Ramen `VolumeReplicationGroup` in async mode selects the class, creates one `VolumeReplication` per PVC, and `lastGroupSyncTime` populates | Positive | — | +| E-07 | Ramen failover (`force`) and relocate (demote plus planned promote) both complete against a live workload | Positive | — | + +--- + +## 4. Manual Scenarios and Test Concepts + +### M-01 — Demote with writes in flight + +**Design reference:** §5.2, Open Question 1 + +**What to verify:** the demote verb's quiesce behavior when the workload is still writing at the moment of the fence, since Ramen ordinarily unmounts first but nothing guarantees it. + +**Test concept:** +1. Run a continuous writer against a replicated volume. +2. Issue demote without stopping the writer. +3. Verify the final flush lands on the peer, the writer's in-flight I/O fails cleanly (no acknowledged-but-lost write), and a subsequent planned promote on the peer serves every acknowledged write. + +### M-02 — Coexistence under concurrent claims + +**Design reference:** §8 + +**What to verify:** the one-owner rule under a race: the annotation and a `VolumeReplication` claiming the same volume in the same reconcile window must converge on one owner with the other visibly refused, never two attach calls. + +**Test concept:** +1. Apply the PVC annotation and create a `VolumeReplication` for the same PVC near-simultaneously. +2. Verify exactly one path attached, the other surfaced its refusal (event or condition), and the backend saw a single policy attach. + +--- + +## 5. Axis Coverage + +| Axis | Values covered | IDs | Not covered | +|------------------|------------------------------------------------------------------------|---------------------------------------|-----------------------------------------------------------| +| Verb lifecycle | enable, disable, info, forced promote, planned promote, demote, resync | U-01 … U-17, I-02 … I-04, E-01 … E-04 | — | +| Idempotency | repeat enable, disable, promote, demote; re-drive after restart | U-02, U-05, U-11, U-15, I-07 | repeated resync | +| Conditions | healthy, degraded, error, staleness, resyncing, disabled | U-18 … U-22, E-05 | condition behavior across backend upgrade | +| Coexistence | slot skip, one-owner refusal, concurrent claim | U-26, U-27, M-02 | migration of an annotated volume onto a VolumeReplication | +| peerClasses | verified, missing class, mispaired policies | U-23 … U-25 | drift after verification | +| Orchestrator | direct kubectl lifecycle, Ramen VRG async | I-02 … I-06, E-06, E-07 | Ramen hub failover of multiple apps | +| Cluster topology | two clusters, one relationship addressed from both sides | E-01 … E-07 | three-cluster (cascaded) topologies | + +--- + +## 6. Coverage Summary + +| Class | Scenarios | Covered | Not covered | +|-------------|-----------|---------|-------------| +| Unit | 27 | 0 | U-01 … U-27 | +| Integration | 7 | 0 | I-01 … I-07 | +| E2E | 7 | 0 | E-01 … E-07 | +| Manual | 2 | 0 | M-01, M-02 | + +Every scenario is uncovered because the design is Draft. The counts are the target, and each `Test` column fills in as the work lands. + +--- + +## 7. What Is Not Yet Covered + +| # | Gap | Reason | +|-------------|-------------------------------------------------------------------------------------------------------------------|-----------------------------------------------------------------------------------------------------------------| +| U-01 … U-27 | The verb adapter, condition derivation, preflight, and coexistence units | The adapter does not exist; blocked on P0-1 and P0-2 for the Phase 1 verbs and P0-3 for demote | +| I-01 … I-07 | The sidecar and controller-manager loop | Blocked on the vendored CRDs and images (P0-5) | +| E-01 … E-07 | The live lifecycle and the Ramen gate | Blocked on Phase 1 and 2 landing, plus a two-cluster test bed with Ramen dr-cluster installed for E-06 and E-07 | +| — | Repeated resync, class drift after verification, annotated-volume migration onto the adapter, cascaded topologies | Beyond the first coverage pass, recorded so the gaps are explicit rather than assumed covered | +| M-01, M-02 | Demote under writes; concurrent ownership race | Need failure injection and precise timing a live two-cluster run does not automate yet | From 57458a9116e7d109943a973c69f52c8812863c8e Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 09:03:26 +0200 Subject: [PATCH 026/206] feat(webhook): approving a deployment config is answered against the cluster A ClusterDeploymentConfig is a draft until spec.approved is set, and setting it is what makes the document immutable. A draft naming a worker that does not exist, or a cluster that already exists, was therefore admitted and frozen, and the expansion refused it afterward with nothing left to edit. The two kinds of edit are now answered differently. A draft is admitted on its structure alone, however wrong it is about the world, because that is what a draft is for. The edit that sets spec.approved is answered against the live cluster: every worker named by every group exists as a Node, spec.clusterRef resolves when it is set and the cluster does not already exist when it is not, a growth document's device class matches the cluster's, and no other approved document already creates the same cluster. Every problem is reported at once rather than one per apply, because approval is what freezes the document and an approver fixing one mistake at a time is the same failure in slow motion. A fifth check the design's list does not name is here for the same reason: a document naming neither a cluster nor a clusterRef is structurally valid and describes no deployment, and the expansion refuses it once it can no longer be edited. The device class and the target cluster name are read through the deployment package rather than derived a second time, so admission and the expansion cannot disagree about what a document says. design-clusterdeploymentconfig.md 5.1 and 5.2 are the specification. 5.2's premise is not quite right and the file says so: the API server validates CEL before it calls a validating webhook, so on a current CRD the schema's immutability rejection arrives first and the webhook's restatement answers only for a cluster whose CRD predates those rules. Co-Authored-By: Claude Fable 5 --- .../simplyblock-operator-webhook.yaml | 20 + operator/cmd/main.go | 7 + operator/config/webhook/manifests.yaml | 20 + .../controllers/deployment/expansion.go | 34 +- .../controllers/deployment/validation.go | 2 +- .../clusterdeploymentconfig_validator.go | 333 +++++++++++++++ .../clusterdeploymentconfig_validator_test.go | 380 ++++++++++++++++++ 7 files changed, 785 insertions(+), 11 deletions(-) create mode 100644 operator/internal/webhook/clusterdeploymentconfig_validator.go create mode 100644 operator/internal/webhook/clusterdeploymentconfig_validator_test.go diff --git a/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml b/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml index d0072700b..49dfd3b17 100644 --- a/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml @@ -86,6 +86,26 @@ webhooks: resources: - persistentvolumeclaims sideEffects: None +- admissionReviewVersions: + - v1 + clientConfig: + service: + name: simplyblock-operator-webhook-service + namespace: {{ .Release.Namespace }} + path: /validate-storage-simplyblock-io-v1alpha2-clusterdeploymentconfig + failurePolicy: Fail + name: vclusterdeploymentconfig.simplyblock.io + rules: + - apiGroups: + - storage.simplyblock.io + apiVersions: + - v1alpha2 + operations: + - CREATE + - UPDATE + resources: + - clusterdeploymentconfigs + sideEffects: None - admissionReviewVersions: - v1 clientConfig: diff --git a/operator/cmd/main.go b/operator/cmd/main.go index 1dcb32e54..154ebb8fb 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -843,6 +843,13 @@ func main() { }}) setupLog.Info("registered storagepool validating webhook") + mgr.GetWebhookServer().Register("/validate-storage-simplyblock-io-v1alpha2-clusterdeploymentconfig", + &webhook.Admission{Handler: &internalwebhook.ClusterDeploymentConfigValidator{ + Client: mgr.GetClient(), + Decoder: admission.NewDecoder(mgr.GetScheme()), + }}) + setupLog.Info("registered clusterdeploymentconfig validating webhook") + mgr.GetWebhookServer().Register("/validate-storage-simplyblock-io-v1alpha2-storagebackup", &webhook.Admission{Handler: &internalwebhook.StorageBackupValidator{ Client: mgr.GetClient(), diff --git a/operator/config/webhook/manifests.yaml b/operator/config/webhook/manifests.yaml index 23cdc2154..82ec3fd9a 100644 --- a/operator/config/webhook/manifests.yaml +++ b/operator/config/webhook/manifests.yaml @@ -68,6 +68,26 @@ webhooks: resources: - persistentvolumeclaims sideEffects: None +- admissionReviewVersions: + - v1 + clientConfig: + service: + name: webhook-service + namespace: system + path: /validate-storage-simplyblock-io-v1alpha2-clusterdeploymentconfig + failurePolicy: Fail + name: vclusterdeploymentconfig.simplyblock.io + rules: + - apiGroups: + - storage.simplyblock.io + apiVersions: + - v1alpha2 + operations: + - CREATE + - UPDATE + resources: + - clusterdeploymentconfigs + sideEffects: None - admissionReviewVersions: - v1 clientConfig: diff --git a/operator/internal/controllers/deployment/expansion.go b/operator/internal/controllers/deployment/expansion.go index dab626162..76af84a15 100644 --- a/operator/internal/controllers/deployment/expansion.go +++ b/operator/internal/controllers/deployment/expansion.go @@ -258,7 +258,7 @@ func (r *ClusterDeploymentConfigReconciler) buildCluster( "the document neither names an existing cluster nor describes one to create") } - class := deviceClassOf(config) + class := DeviceClassOf(config) cluster := &simplyblockv1alpha2.StorageCluster{ ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: config.Namespace}, @@ -558,10 +558,14 @@ func devicesOf(group simplyblockv1alpha2.NodeGroup) []string { return group.Devices.Block } -// deviceClassOf reads the class off the document's groups, which is where it is +// DeviceClassOf reads the class off the document's groups, which is where it is // stated: the device lists already say which class the deployment uses, so // spec.cluster does not restate it. -func deviceClassOf( +// +// It is exported because admission asks the same question before the document +// becomes immutable (design-clusterdeploymentconfig.md §5.1), and a second +// reading of the groups would be free to disagree with this one. +func DeviceClassOf( config *simplyblockv1alpha2.ClusterDeploymentConfig, ) simplyblockv1alpha2.StorageClusterDeviceClass { for _, set := range config.Spec.NodeSets { @@ -584,16 +588,26 @@ func deviceClassOf( return "" } -// targetClusterName is the cluster the document acts on, whichever way it names -// one. +// TargetClusterName is the cluster the document acts on, whichever way it names +// one, and empty for a document that names none. It is exported for the reason +// DeviceClassOf is. +func TargetClusterName(config *simplyblockv1alpha2.ClusterDeploymentConfig) string { + if config.Spec.ClusterRef != "" { + return config.Spec.ClusterRef + } + if config.Spec.Cluster != nil { + return config.Spec.Cluster.Name + } + return "" +} + +// targetClusterName is TargetClusterName as the expansion needs it, where a +// document naming no cluster is a refusal rather than an empty string. func targetClusterName( config *simplyblockv1alpha2.ClusterDeploymentConfig, ) (string, error) { - if config.Spec.ClusterRef != "" { - return config.Spec.ClusterRef, nil - } - if config.Spec.Cluster != nil && config.Spec.Cluster.Name != "" { - return config.Spec.Cluster.Name, nil + if name := TargetClusterName(config); name != "" { + return name, nil } return "", refusef(ClusterNotFound, "the document neither names an existing cluster nor describes one to create") diff --git a/operator/internal/controllers/deployment/validation.go b/operator/internal/controllers/deployment/validation.go index f2b82fce4..026b1d23e 100644 --- a/operator/internal/controllers/deployment/validation.go +++ b/operator/internal/controllers/deployment/validation.go @@ -260,7 +260,7 @@ func (r *ClusterDeploymentConfigReconciler) deviceClassMismatch( return "", fmt.Errorf("reading StorageCluster %s: %w", config.Spec.ClusterRef, err) } - stated := deviceClassOf(config) + stated := DeviceClassOf(config) if stated == "" { return "", nil } diff --git a/operator/internal/webhook/clusterdeploymentconfig_validator.go b/operator/internal/webhook/clusterdeploymentconfig_validator.go new file mode 100644 index 000000000..23023c8b4 --- /dev/null +++ b/operator/internal/webhook/clusterdeploymentconfig_validator.go @@ -0,0 +1,333 @@ +// The approval guard: the edit that sets spec.approved is the last moment a +// deployment config can be corrected, so it is the one that is answered against +// the live cluster. +// +// The asymmetry between the two kinds of edit is the design. A draft is admitted +// on its structure alone, however wrong it is about the world, because a +// document that could not be saved until it was correct is a document nobody can +// work on; the controller reports what it found in status.message and the +// reviewer fixes it in place. The approving edit is checked, because §3.2 makes +// an approved document immutable: a config that reaches Failed on a misspelled +// worker cannot be corrected, only deleted and rewritten. Rejecting the +// approving apply costs one error message, and admitting it costs a dead +// document and a Failed record of a deployment that never happened. +// +// Devices are not checked here. Deciding whether one is mounted, busy, or +// partitioned runs per node and against the node rather than against the API +// server, which is discovery's job and then Validating's, the expansion's first +// step. +// +// Nothing is exempt, which is where this parts company with StorageNodeValidator +// next door. That one admits the operator's own service account because only the +// operator may re-point a node. Here the operator needs no exemption: a discovery +// run writes drafts, which is what the permissive path already admits, and a +// config the operator could approve and an administrator could not would be a +// gate with a hole in it. +// +// The §3.2 rules are restated below rather than left to the CEL on the type. The +// API server validates the schema before it calls a validating webhook, so for a +// current CRD those rejections are CEL's and arrive first; the restatement is +// what answers for a cluster whose CRD predates the rules, and it names the field +// that was edited and what to do instead where CEL names the rule that failed. +// +// Specified by operator/docs/designs/crd-redesign/design-clusterdeploymentconfig.md +// §5, whose four checks are checkApproval below. + +package webhook + +import ( + "context" + "fmt" + "net/http" + "sort" + "strings" + + admissionv1 "k8s.io/api/admission/v1" + corev1 "k8s.io/api/core/v1" + "k8s.io/apimachinery/pkg/api/equality" + apierrors "k8s.io/apimachinery/pkg/api/errors" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/webhook/admission" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/deployment" +) + +// +kubebuilder:webhook:path=/validate-storage-simplyblock-io-v1alpha2-clusterdeploymentconfig,mutating=false,failurePolicy=fail,sideEffects=None,groups=storage.simplyblock.io,resources=clusterdeploymentconfigs,verbs=create;update,versions=v1alpha2,name=vclusterdeploymentconfig.simplyblock.io,admissionReviewVersions=v1 + +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=clusterdeploymentconfigs,verbs=get;list;watch +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storageclusters,verbs=get;list;watch +// +kubebuilder:rbac:groups="",resources=nodes,verbs=get;list;watch + +// ClusterDeploymentConfigValidator answers the approving edit against the +// Kubernetes API. +// +// failurePolicy=Fail, for the reason StorageNodeValidator already gives in this +// repository. The webhook server runs in the operator pod, so its availability +// tracks the operator's, and while the operator is down nothing expands anyway. A +// window in which approvals are admitted unchecked is a window in which immutable +// mistakes are created, and that is worse than a window in which no document can +// be approved. +type ClusterDeploymentConfigValidator struct { + // Client reads the workers, the cluster, and the other configs the checks are + // answered from. All of them are objects the operator already caches, which is + // what makes them cheap enough to read inside an admission request. + Client client.Client + + // Decoder turns the request's raw objects into documents. + Decoder admission.Decoder +} + +func (v *ClusterDeploymentConfigValidator) Handle( + ctx context.Context, req admission.Request, +) admission.Response { + if req.Operation != admissionv1.Create && req.Operation != admissionv1.Update { + return admission.Allowed("") + } + + config := &simplyblockv1alpha2.ClusterDeploymentConfig{} + if err := v.Decoder.Decode(req, config); err != nil { + return admission.Errored(http.StatusBadRequest, err) + } + + if req.Operation == admissionv1.Update { + old := &simplyblockv1alpha2.ClusterDeploymentConfig{} + if err := v.Decoder.DecodeRaw(req.OldObject, old); err != nil { + return admission.Errored(http.StatusBadRequest, err) + } + if old.Spec.Approved { + return afterApproval(old, config) + } + } + + if !config.Spec.Approved { + // A draft. The schema and the CEL rules of §3.2 are the whole check. + return admission.Allowed("") + } + + namespace := config.Namespace + if namespace == "" { + // A namespaced object created through a namespaced endpoint may arrive + // with the field unset, because the path carries it instead. + namespace = req.Namespace + } + + problems, err := v.checkApproval(ctx, namespace, config) + if err != nil { + return admission.Errored(http.StatusInternalServerError, err) + } + if len(problems) == 0 { + return admission.Allowed("") + } + return admission.Denied(fmt.Sprintf( + "this document cannot be approved, and approving it is what makes it "+ + "immutable: %s", strings.Join(problems, "; "))) +} + +// afterApproval restates §3.2 for a document that is already approved. +func afterApproval(old, config *simplyblockv1alpha2.ClusterDeploymentConfig) admission.Response { + if !config.Spec.Approved { + return admission.Denied( + "spec.approved cannot be withdrawn: un-approving a document does not " + + "un-expand it, and the cluster and nodes it produced are removed by " + + "deleting them rather than by editing the record that describes them") + } + if !equality.Semantic.DeepEqual(old.Spec, config.Spec) { + return admission.Denied( + "spec is immutable once spec.approved is true, because the document is " + + "then the record of what was deployed. To add nodes to the cluster " + + "this one built, write a second config naming it in spec.clusterRef") + } + // Metadata alone. The operator labels an approved config itself, so this is + // the ordinary edit rather than an exception to the rule above. + return admission.Allowed("") +} + +// checkApproval is §5.1's four checks, and returns everything wrong rather than +// the first thing wrong. +// +// A reviewer fixing a document that is about to become immutable wants the whole +// list: a document with a missing worker and a dangling cluster reference should +// say so once rather than over two applies, each of which costs another approval. +func (v *ClusterDeploymentConfigValidator) checkApproval( + ctx context.Context, namespace string, config *simplyblockv1alpha2.ClusterDeploymentConfig, +) ([]string, error) { + var problems []string + + missing, err := v.missingWorkers(ctx, config) + if err != nil { + return nil, err + } + if len(missing) > 0 { + problems = append(problems, fmt.Sprintf( + "a group names %s, which %s not %s of this Kubernetes cluster", + strings.Join(missing, ", "), + plural(len(missing), "is", "are"), plural(len(missing), "a node", "nodes"))) + } + + cluster, problem, err := v.resolveCluster(ctx, namespace, config) + if err != nil { + return nil, err + } + switch { + case problem != "": + problems = append(problems, problem) + + case cluster != nil: + // A growth document. The class the groups name has to be the one the + // cluster is built out of: the nodes would otherwise be rejected one at a + // time by StorageNodeValidator, which is a slower way to learn it and + // leaves a half-expanded deployment behind. + if mismatch := classMismatch(config, cluster); mismatch != "" { + problems = append(problems, mismatch) + } + + default: + // A document that creates its cluster. No other approved config may + // already own the one it would create. + owner, err := v.otherOwner(ctx, namespace, config) + if err != nil { + return nil, err + } + if owner != "" { + problems = append(problems, fmt.Sprintf( + "ClusterDeploymentConfig %s is approved and creates StorageCluster %s "+ + "as well; both would race to create it and the loser is an immutable "+ + "Failed document, so add to it with spec.clusterRef instead", + owner, deployment.TargetClusterName(config))) + } + } + + return problems, nil +} + +// resolveCluster answers §6's table for the cluster the document names: the +// StorageCluster a growth document grows, nil for a document that may create its +// own, and a problem for the two combinations the expansion refuses. +func (v *ClusterDeploymentConfigValidator) resolveCluster( + ctx context.Context, namespace string, config *simplyblockv1alpha2.ClusterDeploymentConfig, +) (*simplyblockv1alpha2.StorageCluster, string, error) { + name := deployment.TargetClusterName(config) + if name == "" { + return nil, "the document names neither a cluster to create in spec.cluster " + + "nor one to add nodes to in spec.clusterRef, so it describes no deployment", nil + } + + var cluster simplyblockv1alpha2.StorageCluster + err := v.Client.Get(ctx, client.ObjectKey{Namespace: namespace, Name: name}, &cluster) + found := err == nil + if err != nil && !apierrors.IsNotFound(err) { + return nil, "", fmt.Errorf("reading StorageCluster %s: %w", name, err) + } + + switch { + case config.Spec.ClusterRef != "" && !found: + return nil, fmt.Sprintf( + "spec.clusterRef names StorageCluster %s, and there is none by that name "+ + "in namespace %s", name, namespace), nil + + case config.Spec.ClusterRef == "" && found: + return nil, fmt.Sprintf( + "spec.cluster.name is %s and a StorageCluster by that name already exists; "+ + "set spec.clusterRef to add nodes to it instead", name), nil + + case found: + return &cluster, "", nil + } + return nil, "", nil +} + +// classMismatch reports a growth document whose groups name devices of a class +// its cluster is not built out of. +func classMismatch( + config *simplyblockv1alpha2.ClusterDeploymentConfig, + cluster *simplyblockv1alpha2.StorageCluster, +) string { + stated := deployment.DeviceClassOf(config) + if stated == "" { + // A document whose groups name no devices states no class, and there is + // nothing to disagree with. + return "" + } + + held := cluster.Spec.DeviceClass + if held == "" { + // An unstated class is NVMe, which is what the cluster's own field + // defaults to and what describes every cluster predating the field. + held = simplyblockv1alpha2.StorageClusterDeviceClassNVMe + } + if stated == held { + return "" + } + return fmt.Sprintf( + "the groups name %s devices and StorageCluster %s is built out of %s; "+ + "an erasure-coding stripe placed across both classes is written at the "+ + "slower one's rate", stated, cluster.Name, held) +} + +// missingWorkers names every worker the document lists that is not a node of this +// Kubernetes cluster. +func (v *ClusterDeploymentConfigValidator) missingWorkers( + ctx context.Context, config *simplyblockv1alpha2.ClusterDeploymentConfig, +) ([]string, error) { + var workers corev1.NodeList + if err := v.Client.List(ctx, &workers); err != nil { + return nil, fmt.Errorf("listing the Kubernetes workers: %w", err) + } + present := make(map[string]struct{}, len(workers.Items)) + for i := range workers.Items { + present[workers.Items[i].Name] = struct{}{} + } + + missing := map[string]struct{}{} + for _, set := range config.Spec.NodeSets { + for _, group := range set.Groups { + for _, worker := range group.Workers { + if _, there := present[worker]; !there { + missing[worker] = struct{}{} + } + } + } + } + + named := make([]string, 0, len(missing)) + for worker := range missing { + named = append(named, worker) + } + // Stable, so one document does not produce two differently ordered refusals. + sort.Strings(named) + return named, nil +} + +// otherOwner names the approved config that already creates the cluster this one +// would create, where there is one. +// +// Only a document that creates a cluster owns it. Two growth documents adding +// nodes to one cluster is the design rather than a conflict: growth is a second +// document precisely so that the audit trail of how a cluster reached its size is +// a series of documents. +func (v *ClusterDeploymentConfigValidator) otherOwner( + ctx context.Context, namespace string, config *simplyblockv1alpha2.ClusterDeploymentConfig, +) (string, error) { + var configs simplyblockv1alpha2.ClusterDeploymentConfigList + if err := v.Client.List(ctx, &configs, client.InNamespace(namespace)); err != nil { + return "", fmt.Errorf("listing the deployment configs of namespace %s: %w", namespace, err) + } + + wanted := deployment.TargetClusterName(config) + for i := range configs.Items { + other := &configs.Items[i] + switch { + case other.Name == config.Name: + // This document, as the cache holds it. + case !other.Spec.Approved: + // A draft owns nothing. Whichever of the two is approved first + // becomes the owner, and the other is refused then. + case other.Spec.ClusterRef != "": + // A growth document, which adds to a cluster rather than owning it. + case deployment.TargetClusterName(other) == wanted: + return other.Name, nil + } + } + return "", nil +} \ No newline at end of file diff --git a/operator/internal/webhook/clusterdeploymentconfig_validator_test.go b/operator/internal/webhook/clusterdeploymentconfig_validator_test.go new file mode 100644 index 000000000..d35e85e8a --- /dev/null +++ b/operator/internal/webhook/clusterdeploymentconfig_validator_test.go @@ -0,0 +1,380 @@ +// Tests for the approval guard. +// +// The asymmetry is the design (design-clusterdeploymentconfig.md §5.1): a draft +// is admitted on its structure alone, however wrong it is about the world, and +// the edit that sets spec.approved is answered against the live cluster. What +// separates the two is that §3.2 makes an approved document immutable, so a +// document admitted with a mistake in it cannot be corrected, only deleted. + +package webhook + +import ( + "context" + "encoding/json" + "strings" + "testing" + + admissionv1 "k8s.io/api/admission/v1" + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/runtime" + "k8s.io/utils/ptr" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + "sigs.k8s.io/controller-runtime/pkg/webhook/admission" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// The namespace every document and cluster below lives in, and the cluster they +// name. What the cases differ in is what is there and what the document says +// about it, never what anything is called. +const ( + testDeploymentNamespace = "simplyblock" + testDeploymentCluster = "production" + testDeploymentConfig = "rack-one" +) + +func testWorker(name string) *corev1.Node { + return &corev1.Node{ObjectMeta: metav1.ObjectMeta{Name: name}} +} + +func testDeploymentStorageCluster( + name string, class simplyblockv1alpha2.StorageClusterDeviceClass, +) *simplyblockv1alpha2.StorageCluster { + return &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: testDeploymentNamespace}, + Spec: simplyblockv1alpha2.StorageClusterSpec{DeviceClass: class}, + } +} + +// testConfig is a document that passes every check: one group, one worker that +// exists, NVMe devices, and a cluster of its own to create. +func testConfig() *simplyblockv1alpha2.ClusterDeploymentConfig { + return &simplyblockv1alpha2.ClusterDeploymentConfig{ + ObjectMeta: metav1.ObjectMeta{ + Name: testDeploymentConfig, + Namespace: testDeploymentNamespace, + }, + Spec: simplyblockv1alpha2.ClusterDeploymentConfigSpec{ + Cluster: &simplyblockv1alpha2.ClusterTemplate{ + Name: testDeploymentCluster, + MaxSubsystemCount: ptr.To(int32(10)), + VCPUCount: ptr.To(int32(4)), + }, + NodeSets: []simplyblockv1alpha2.NodeSet{{ + Name: "rack-b", + Groups: []simplyblockv1alpha2.NodeGroup{{ + Name: "workers", + Workers: []string{"worker-1"}, + MgmtInterface: "eth0", + Devices: &simplyblockv1alpha2.DeviceSelection{ + NVMe: []string{"0000:5e:00.0"}, + }, + }}, + }}, + }, + } +} + +// approved is the document the approving edit produces. +func approved( + config *simplyblockv1alpha2.ClusterDeploymentConfig, +) *simplyblockv1alpha2.ClusterDeploymentConfig { + config.Spec.Approved = true + return config +} + +// grows re-points a document at a cluster that already exists, which is the +// growth document of §6. +func grows( + config *simplyblockv1alpha2.ClusterDeploymentConfig, cluster string, +) *simplyblockv1alpha2.ClusterDeploymentConfig { + config.Spec.ClusterRef = cluster + config.Spec.Cluster = nil + return config +} + +// withBlockDevices swaps the document's device class, which is what a growth +// document naming the wrong one for its cluster looks like. +func withBlockDevices( + config *simplyblockv1alpha2.ClusterDeploymentConfig, +) *simplyblockv1alpha2.ClusterDeploymentConfig { + config.Spec.NodeSets[0].Groups[0].Devices = &simplyblockv1alpha2.DeviceSelection{ + Block: []string{"/dev/sdb"}, + } + return config +} + +func configRaw(t *testing.T, config *simplyblockv1alpha2.ClusterDeploymentConfig) runtime.RawExtension { + t.Helper() + raw, err := json.Marshal(config) + if err != nil { + t.Fatalf("marshal ClusterDeploymentConfig: %v", err) + } + return runtime.RawExtension{Raw: raw} +} + +// review runs one admission request against a validator holding the objects the +// cluster is said to carry. +func review( + t *testing.T, + existing []client.Object, + operation admissionv1.Operation, + old, config *simplyblockv1alpha2.ClusterDeploymentConfig, +) admission.Response { + t.Helper() + + scheme := newPoolScheme(t) + validator := &ClusterDeploymentConfigValidator{ + Client: fake.NewClientBuilder().WithScheme(scheme).WithObjects(existing...).Build(), + Decoder: admission.NewDecoder(scheme), + } + + request := admission.Request{AdmissionRequest: admissionv1.AdmissionRequest{ + Operation: operation, + Namespace: testDeploymentNamespace, + Object: configRaw(t, config), + }} + if old != nil { + request.OldObject = configRaw(t, old) + } + return validator.Handle(context.Background(), request) +} + +// approve is the edit the whole webhook is about: a draft that exists, updated +// to set spec.approved. +func approve( + t *testing.T, existing []client.Object, config *simplyblockv1alpha2.ClusterDeploymentConfig, +) admission.Response { + t.Helper() + + draft := config.DeepCopy() + draft.Spec.Approved = false + return review(t, existing, admissionv1.Update, draft, approved(config)) +} + +func mustAllow(t *testing.T, response admission.Response) { + t.Helper() + if !response.Allowed { + t.Fatalf("the document was refused: %s", response.Result.Message) + } +} + +func mustDeny(t *testing.T, response admission.Response, mentions ...string) { + t.Helper() + if response.Allowed { + t.Fatal("the document was admitted, and it should not have been") + } + for _, want := range mentions { + if !strings.Contains(response.Result.Message, want) { + t.Errorf("the refusal %q does not mention %q", response.Result.Message, want) + } + } +} + +// A draft may name a worker that does not exist and a cluster that does not +// exist, because a document that could not be saved until it was correct is a +// document nobody can work on (§5.1). The controller reports what it found in +// status.message and the reviewer fixes it in place. +func TestADraftIsAdmittedOnItsStructureAlone(t *testing.T) { + draft := grows(testConfig(), "no-such-cluster") + draft.Spec.NodeSets[0].Groups[0].Workers = []string{"no-such-worker"} + + mustAllow(t, review(t, nil, admissionv1.Create, nil, draft)) + mustAllow(t, review(t, nil, admissionv1.Update, draft.DeepCopy(), draft)) +} + +// The approving edit is answered against the live cluster, and a document whose +// four answers are all yes is admitted. +func TestApprovingADocumentThatChecksOutIsAdmitted(t *testing.T) { + mustAllow(t, approve(t, []client.Object{testWorker("worker-1")}, testConfig())) +} + +// A document may also be written already approved, which is the same edit +// arriving as a create. +func TestACreateThatIsAlreadyApprovedIsValidatedToo(t *testing.T) { + mustAllow(t, review(t, []client.Object{testWorker("worker-1")}, + admissionv1.Create, nil, approved(testConfig()))) + + mustDeny(t, review(t, nil, admissionv1.Create, nil, approved(testConfig())), + "worker-1") +} + +// Every worker named by every group has to exist as a Node (§5.1). +func TestApprovingIsRefusedWhenAWorkerIsNotANode(t *testing.T) { + config := testConfig() + config.Spec.NodeSets[0].Groups[0].Workers = []string{"worker-1", "worker-2"} + + response := approve(t, []client.Object{testWorker("worker-1")}, config) + mustDeny(t, response, "worker-2") + if strings.Contains(response.Result.Message, "worker-1") { + t.Errorf("the refusal %q names a worker that does exist", response.Result.Message) + } +} + +// spec.clusterRef has to resolve to a StorageCluster when it is set (§6's +// ClusterNotFound row). +func TestApprovingIsRefusedWhenClusterRefResolvesToNothing(t *testing.T) { + mustDeny(t, approve(t, []client.Object{testWorker("worker-1")}, + grows(testConfig(), "no-such-cluster")), "no-such-cluster") +} + +// ...and to nothing when it is not (§6's ClusterExists row). The refusal says +// what to do instead, because the document that was wanted is a growth document. +func TestApprovingIsRefusedWhenTheClusterAlreadyExists(t *testing.T) { + existing := []client.Object{ + testWorker("worker-1"), + testDeploymentStorageCluster(testDeploymentCluster, ""), + } + + mustDeny(t, approve(t, existing, testConfig()), + testDeploymentCluster, "spec.clusterRef") +} + +// A document that names neither a cluster to create nor one to grow describes no +// deployment at all, and the expansion refuses it with ClusterNotFound — after +// the document has become immutable. +func TestApprovingIsRefusedWhenTheDocumentNamesNoCluster(t *testing.T) { + config := testConfig() + config.Spec.Cluster = nil + + mustDeny(t, approve(t, []client.Object{testWorker("worker-1")}, config), + "spec.cluster") +} + +// The class the groups name has to match the cluster's, where one is named. The +// nodes would otherwise be rejected one at a time by StorageNodeValidator, which +// is a slower way to learn it and leaves a half-expanded deployment behind +// (§5.1). +func TestApprovingAGrowthDocumentOfTheWrongDeviceClassIsRefused(t *testing.T) { + existing := []client.Object{ + testWorker("worker-1"), + testDeploymentStorageCluster( + testDeploymentCluster, simplyblockv1alpha2.StorageClusterDeviceClassNVMe), + } + + mustDeny(t, approve(t, existing, withBlockDevices(grows(testConfig(), testDeploymentCluster))), + "LogicalBlock", "NVMe") +} + +func TestApprovingAGrowthDocumentOfTheRightDeviceClassIsAdmitted(t *testing.T) { + existing := []client.Object{ + testWorker("worker-1"), + testDeploymentStorageCluster( + testDeploymentCluster, simplyblockv1alpha2.StorageClusterDeviceClassLogicalBlock), + } + + mustAllow(t, approve(t, existing, + withBlockDevices(grows(testConfig(), testDeploymentCluster)))) +} + +// A cluster carrying no class is an NVMe cluster, which is what the field +// defaults to and what describes every cluster predating it. +func TestAGrowthDocumentReadsAnUnstatedClassAsNVMe(t *testing.T) { + existing := []client.Object{ + testWorker("worker-1"), + testDeploymentStorageCluster(testDeploymentCluster, ""), + } + + mustAllow(t, approve(t, existing, grows(testConfig(), testDeploymentCluster))) + mustDeny(t, approve(t, existing, withBlockDevices(grows(testConfig(), testDeploymentCluster))), + "NVMe") +} + +// No other approved config may already own the cluster this one would create +// (§5.1). Both would race to create it, and the loser is an immutable Failed +// document reporting a deployment that never happened. +func TestApprovingASecondCreateOfTheSameClusterIsRefused(t *testing.T) { + other := approved(testConfig()) + other.Name = "rack-two" + + mustDeny(t, approve(t, []client.Object{testWorker("worker-1"), other}, testConfig()), + "rack-two") +} + +// An unapproved document naming the same cluster is not an owner. It is a draft, +// and whichever of the two is approved first becomes the owner. +func TestADraftNamingTheSameClusterDoesNotBlockAnApproval(t *testing.T) { + other := testConfig() + other.Name = "rack-two" + + mustAllow(t, approve(t, []client.Object{testWorker("worker-1"), other}, testConfig())) +} + +// Growth is a second document, so two approved documents adding nodes to one +// cluster is the design rather than a conflict (§6). +func TestASecondGrowthDocumentIsAdmitted(t *testing.T) { + other := approved(grows(testConfig(), testDeploymentCluster)) + other.Name = "rack-two" + + existing := []client.Object{ + testWorker("worker-1"), + testDeploymentStorageCluster(testDeploymentCluster, ""), + other, + } + + mustAllow(t, approve(t, existing, grows(testConfig(), testDeploymentCluster))) +} + +// The document being approved is not another config that owns its cluster. +func TestADocumentDoesNotConflictWithItself(t *testing.T) { + existing := []client.Object{testWorker("worker-1"), approved(testConfig())} + + mustAllow(t, approve(t, existing, testConfig())) +} + +// §3.2's rules restated: the webhook names the field that was edited and what to +// do instead, where CEL names the rule that failed (§5.2). +func TestEditingAnApprovedDocumentIsRefused(t *testing.T) { + old := approved(testConfig()) + edited := approved(testConfig()) + edited.Spec.NodeSets[0].Groups[0].Workers = []string{"worker-2"} + + response := review(t, []client.Object{testWorker("worker-1"), testWorker("worker-2")}, + admissionv1.Update, old, edited) + mustDeny(t, response, "spec.clusterRef") +} + +func TestWithdrawingApprovalIsRefused(t *testing.T) { + old := approved(testConfig()) + withdrawn := testConfig() + + mustDeny(t, review(t, []client.Object{testWorker("worker-1")}, + admissionv1.Update, old, withdrawn), "approved") +} + +// The operator labels an approved config so that a selector can find one (§5), +// which is an edit to a document whose spec is immutable. +func TestLabelingAnApprovedDocumentIsAdmitted(t *testing.T) { + old := approved(testConfig()) + labeled := approved(testConfig()) + labeled.Labels = map[string]string{"storage.simplyblock.io/ready-to-deploy": "true"} + + mustAllow(t, review(t, nil, admissionv1.Update, old, labeled)) +} + +// A reviewer fixing an immutable document wants the whole list, not one problem +// per apply. +func TestEveryProblemIsReportedAtOnce(t *testing.T) { + config := testConfig() + config.Spec.NodeSets[0].Groups[0].Workers = []string{"no-such-worker"} + + mustDeny(t, approve(t, nil, grows(config, "no-such-cluster")), + "no-such-worker", "no-such-cluster") +} + +// A delete is not the validator's business: the document is a record, and +// deleting it is how a wrong one is disposed of. +func TestDeletingADocumentIsNotTheValidatorsBusiness(t *testing.T) { + mustAllow(t, review(t, nil, admissionv1.Delete, nil, approved(testConfig()))) +} + +// A namespaced object created through a namespaced endpoint may arrive with the +// field unset, because the path carries it instead. +func TestTheRequestNamespaceIsUsedWhenTheObjectCarriesNone(t *testing.T) { + config := approved(testConfig()) + config.Namespace = "" + + mustDeny(t, review(t, nil, admissionv1.Create, nil, config), "worker-1") +} \ No newline at end of file From 2ade4aee755055713397e65fab11b030506c46fb Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 09:05:33 +0200 Subject: [PATCH 027/206] feat(webhook): a node or cluster operation mid-flight refuses to be withdrawn Deleting an Ops object withdraws the record of an operation, which is a different thing from asking the operation to stop. design-crd-model.md 3.1 has both channels ask one question of one graph: a step declaring an abort edge may be withdrawn, and a step declaring none may not. ControlPlaneOps and StorageBackupOps carried that guard. StorageClusterOps and StorageNodeOps carried only their finalizers, which a delete with --force --grace-period=0 skips, so a Promoting migrate could be withdrawn and take the topology re-point it still owed with it. Both guards read which steps refuse from the controller's own graph, through the new UnabortableSteps in each package, and state only the prose: what the step is in the middle of, which a bool cannot say. A test holds each table equal to the graph's answer in both directions, so an entry for a step the graph can abort from, or a missing entry for one it cannot, is a failure rather than a drift. OperatorOps and StoragePoolOps deliberately get no guard. 3.1 derives the webhook from the graph, and both of those abort from every step: discovery changes nothing, and the pool's only action is declared and unimplemented. A fail-closed webhook that can never deny would buy nothing and make those two kinds undeletable whenever the operator is down. It is worth revisiting when StoragePoolOps.Migrate is built on PersistentVolumeOps, or when OperatorOps gains a second action. Co-Authored-By: Claude Fable 5 --- .../simplyblock-operator-webhook.yaml | 38 ++++ operator/cmd/main.go | 8 + operator/config/webhook/manifests.yaml | 38 ++++ .../internal/controllers/cluster/graphs.go | 19 ++ operator/internal/controllers/node/graphs.go | 19 ++ .../webhook/storageclusterops_validator.go | 122 +++++++++++ .../storageclusterops_validator_test.go | 190 ++++++++++++++++ .../webhook/storagenodeops_validator.go | 125 +++++++++++ .../webhook/storagenodeops_validator_test.go | 205 ++++++++++++++++++ 9 files changed, 764 insertions(+) create mode 100644 operator/internal/webhook/storageclusterops_validator.go create mode 100644 operator/internal/webhook/storageclusterops_validator_test.go create mode 100644 operator/internal/webhook/storagenodeops_validator.go create mode 100644 operator/internal/webhook/storagenodeops_validator_test.go diff --git a/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml b/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml index 49dfd3b17..019cecdad 100644 --- a/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml @@ -223,6 +223,25 @@ webhooks: resources: - storagebackuppolicies sideEffects: None +- admissionReviewVersions: + - v1 + clientConfig: + service: + name: simplyblock-operator-webhook-service + namespace: {{ .Release.Namespace }} + path: /validate-storage-simplyblock-io-v1alpha2-storageclusterops + failurePolicy: Fail + name: vstorageclusterops.simplyblock.io + rules: + - apiGroups: + - storage.simplyblock.io + apiVersions: + - v1alpha2 + operations: + - DELETE + resources: + - storageclusterops + sideEffects: None - admissionReviewVersions: - v1 clientConfig: @@ -263,6 +282,25 @@ webhooks: resources: - storagenodes sideEffects: None +- admissionReviewVersions: + - v1 + clientConfig: + service: + name: simplyblock-operator-webhook-service + namespace: {{ .Release.Namespace }} + path: /validate-storage-simplyblock-io-v1alpha2-storagenodeops + failurePolicy: Fail + name: vstoragenodeops.simplyblock.io + rules: + - apiGroups: + - storage.simplyblock.io + apiVersions: + - v1alpha2 + operations: + - DELETE + resources: + - storagenodeops + sideEffects: None - admissionReviewVersions: - v1 clientConfig: diff --git a/operator/cmd/main.go b/operator/cmd/main.go index 154ebb8fb..3295969aa 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -869,6 +869,14 @@ func main() { &webhook.Admission{Handler: &internalwebhook.ControlPlaneOpsValidator{Client: mgr.GetClient()}}) setupLog.Info("registered controlplaneops validating webhook") + mgr.GetWebhookServer().Register("/validate-storage-simplyblock-io-v1alpha2-storageclusterops", + &webhook.Admission{Handler: &internalwebhook.StorageClusterOpsValidator{}}) + setupLog.Info("registered storageclusterops validating webhook") + + mgr.GetWebhookServer().Register("/validate-storage-simplyblock-io-v1alpha2-storagenodeops", + &webhook.Admission{Handler: &internalwebhook.StorageNodeOpsValidator{}}) + setupLog.Info("registered storagenodeops validating webhook") + mgr.GetWebhookServer().Register("/validate-v1-pvc-pinned-volume", &webhook.Admission{Handler: &internalwebhook.PersistentVolumeClaimValidator{ Client: mgr.GetClient(), diff --git a/operator/config/webhook/manifests.yaml b/operator/config/webhook/manifests.yaml index 82ec3fd9a..13890eb8f 100644 --- a/operator/config/webhook/manifests.yaml +++ b/operator/config/webhook/manifests.yaml @@ -205,6 +205,25 @@ webhooks: resources: - storagebackuppolicies sideEffects: None +- admissionReviewVersions: + - v1 + clientConfig: + service: + name: webhook-service + namespace: system + path: /validate-storage-simplyblock-io-v1alpha2-storageclusterops + failurePolicy: Fail + name: vstorageclusterops.simplyblock.io + rules: + - apiGroups: + - storage.simplyblock.io + apiVersions: + - v1alpha2 + operations: + - DELETE + resources: + - storageclusterops + sideEffects: None - admissionReviewVersions: - v1 clientConfig: @@ -245,6 +264,25 @@ webhooks: resources: - storagenodes sideEffects: None +- admissionReviewVersions: + - v1 + clientConfig: + service: + name: webhook-service + namespace: system + path: /validate-storage-simplyblock-io-v1alpha2-storagenodeops + failurePolicy: Fail + name: vstoragenodeops.simplyblock.io + rules: + - apiGroups: + - storage.simplyblock.io + apiVersions: + - v1alpha2 + operations: + - DELETE + resources: + - storagenodeops + sideEffects: None - admissionReviewVersions: - v1 clientConfig: diff --git a/operator/internal/controllers/cluster/graphs.go b/operator/internal/controllers/cluster/graphs.go index 1ecb628fe..a1a7259a2 100644 --- a/operator/internal/controllers/cluster/graphs.go +++ b/operator/internal/controllers/cluster/graphs.go @@ -188,6 +188,25 @@ var abortableSteps = map[step]bool{ // step can be honored. func abortable(current step) bool { return abortableSteps[current] } +// UnabortableSteps are the declared steps an abort cannot be honored from, +// sorted. +// +// It is exported for the DELETE guard on this kind, which asks the same question +// this package's unwind asks: a deletion may not express something spec.abort +// could not, so both channels read one graph +// (design-crd-model.md §3.1). Reading it rather than restating it is what keeps +// the guard from refusing a step this package has since made abortable, or +// admitting one it has not. +func UnabortableSteps() []step { + var refused []step + for _, declared := range statemachine.DeclaredMultiStates(graphs()) { + if s := step(declared); !abortable(s) { + refused = append(refused, s) + } + } + return refused +} + // action converts the API's action enum into the MultiConfig's key. The // conversion exists because statemachine.Action is a concrete string type // rather than a second type parameter (see its doc comment), and doing it in diff --git a/operator/internal/controllers/node/graphs.go b/operator/internal/controllers/node/graphs.go index 225695c9a..3b7908d16 100644 --- a/operator/internal/controllers/node/graphs.go +++ b/operator/internal/controllers/node/graphs.go @@ -336,6 +336,25 @@ var abortableSteps = map[step]bool{ // step can be honored. func abortable(current step) bool { return abortableSteps[current] } +// UnabortableSteps are the declared steps an abort cannot be honored from, +// sorted. +// +// It is exported for the DELETE guard on this kind, which asks the same question +// this package's unwind asks: a deletion may not express something spec.abort +// could not, so both channels read one graph +// (design-crd-model.md §3.1). Reading it rather than restating it is what keeps +// the guard from refusing a step this package has since made abortable, or +// admitting one it has not. +func UnabortableSteps() []step { + var refused []step + for _, declared := range statemachine.DeclaredMultiStates(graphs()) { + if s := step(declared); !abortable(s) { + refused = append(refused, s) + } + } + return refused +} + // unwinds reports whether an abort or a failure from this step owes the node a // resume before the operation ends. Everything from Suspending onward in a drain // does: the node is not serving, and an operation that stopped there and left it diff --git a/operator/internal/webhook/storageclusterops_validator.go b/operator/internal/webhook/storageclusterops_validator.go new file mode 100644 index 000000000..975e7f55c --- /dev/null +++ b/operator/internal/webhook/storageclusterops_validator.go @@ -0,0 +1,122 @@ +// The StorageClusterOps guard: a validating webhook on DELETE that refuses to +// withdraw the record of an operation sitting on a step the graph declares no +// abort edge from. +// +// The operation's finalizer releases the cluster's lock on a graceful delete, +// and a `kubectl delete --force --grace-period=0` skips it. Admission is what is +// left in that path, and on this kind the blast radius is the widest in the +// group: the lock this operation holds is the cluster's, so an abandoned one +// leaves every later cluster operation waiting on an object nobody can find. +// +// A rolling restart is the case that decides the rule. The walk holds the lock +// across every node of the fleet, and from ShuttingDownNode to RestartingNode +// the node it is on is offline with this operation the only thing that will +// bring it back. Withdrawing the record there does not stop the walk, it leaves +// a storage node down and the cluster locked. +// +// Deletion is therefore never a way to express something spec.abort could not: +// both channels ask one graph, and this file asks it through +// cluster.UnabortableSteps rather than restating the answer. What it states here +// is only the prose, because a refusal has to say what the step is waiting for +// and a bool cannot. +// +// design-crd-model.md §3.1 and design-storagecluster.md §7 are the +// specification. + +package webhook + +import ( + "context" + "encoding/json" + "fmt" + "net/http" + + admissionv1 "k8s.io/api/admission/v1" + "sigs.k8s.io/controller-runtime/pkg/webhook/admission" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// +kubebuilder:webhook:path=/validate-storage-simplyblock-io-v1alpha2-storageclusterops,mutating=false,failurePolicy=fail,sideEffects=None,groups=storage.simplyblock.io,resources=storageclusterops,verbs=delete,versions=v1alpha2,name=vstorageclusterops.simplyblock.io,admissionReviewVersions=v1 + +// StorageClusterOpsValidator refuses a delete that would strand a cluster +// operation mid-flight. +// +// failurePolicy=Fail, because the webhook server runs inside the operator pod: +// its availability tracks the operator's, and while the operator is down nothing +// advances an operation anyway. +// +// It holds no client. Everything the decision rests on is in the object being +// deleted, which is what makes the answer the same on every replica and immune +// to a stale cache. +type StorageClusterOpsValidator struct{} + +// undeletableClusterSteps says what each step with no abort edge is in the +// middle of. Which steps those are is cluster.UnabortableSteps's answer, and a +// test holds the two sets equal in both directions: an entry here for a step the +// graph can abort from refuses a delete the abort channel would have honored, +// and a missing entry admits the withdrawal of a record nothing else accounts +// for. +var undeletableClusterSteps = map[simplyblockv1alpha2.StorageClusterOpsStep]string{ + simplyblockv1alpha2.StorageClusterOpsStepAwaiting: "the cluster is carrying out the action, " + + "and it is doing so whether or not this record exists", + simplyblockv1alpha2.StorageClusterOpsStepShuttingDown: "the cluster is going down, and the " + + "start that follows is this operation's second half", + simplyblockv1alpha2.StorageClusterOpsStepStarting: "the cluster is coming back from the " + + "shutdown this operation performed", + simplyblockv1alpha2.StorageClusterOpsStepShuttingDownNode: "a storage node is being taken " + + "offline, and this walk is the only thing that will bring it back", + simplyblockv1alpha2.StorageClusterOpsStepRefreshingPod: "the node is offline and its " + + "storage-node pod is being replaced", + simplyblockv1alpha2.StorageClusterOpsStepAwaitingPod: "the node is offline and its " + + "replacement pod is still coming up", + simplyblockv1alpha2.StorageClusterOpsStepRestartingNode: "the node is offline and is being " + + "brought back from the shutdown this walk performed", +} + +func (v *StorageClusterOpsValidator) Handle( + _ context.Context, req admission.Request, +) admission.Response { + if req.Operation != admissionv1.Delete { + return admission.Allowed("") + } + return v.admitDelete(req) +} + +// admitDelete reads the object from req.OldObject, which is what the API server +// sends on a DELETE: there is no new object, and the step the operation is on is +// in the status of the one being removed. +func (v *StorageClusterOpsValidator) admitDelete(req admission.Request) admission.Response { + if len(req.OldObject.Raw) == 0 { + // Nothing to read means nothing to refuse on. Admitting is the only + // answer that does not block a delete on the basis of no information. + return admission.Allowed("") + } + + var ops simplyblockv1alpha2.StorageClusterOps + if err := json.Unmarshal(req.OldObject.Raw, &ops); err != nil { + return admission.Errored(http.StatusBadRequest, err) + } + + // A terminal operation is a record of work that has finished, and + // withdrawing it stops nothing. + switch ops.Status.Phase { + case simplyblockv1alpha2.StorageClusterOpsPhaseSucceeded, + simplyblockv1alpha2.StorageClusterOpsPhaseFailed, + simplyblockv1alpha2.StorageClusterOpsPhaseAborted: + return admission.Allowed("the operation is terminal") + } + + step := simplyblockv1alpha2.StorageClusterOpsStep(ops.Status.Step.State) + doing, undeletable := undeletableClusterSteps[step] + if !undeletable { + return admission.Allowed("") + } + + return admission.Denied(fmt.Sprintf( + "StorageClusterOps %s/%s is at step %s, where %s. Deleting the record would not stop "+ + "that work, it would remove the only thing accounting for it and leave cluster %s "+ + "locked by an object that no longer exists. Set spec.abort to stop an operation that "+ + "can still be stopped, and delete the record once it is terminal.", + ops.Namespace, ops.Name, step, doing, ops.Spec.ClusterRef)) +} \ No newline at end of file diff --git a/operator/internal/webhook/storageclusterops_validator_test.go b/operator/internal/webhook/storageclusterops_validator_test.go new file mode 100644 index 000000000..c601c2a0b --- /dev/null +++ b/operator/internal/webhook/storageclusterops_validator_test.go @@ -0,0 +1,190 @@ +// What the StorageClusterOps guard admits and refuses. +// +// A rolling restart is the case that motivates it. The walk holds the cluster's +// lock across every node, and from the shutdown of a node to its restart that +// operation is the only thing that will bring the node back. A forced delete +// there skips the finalizer, so the lock stays on a cluster whose object is gone +// and a storage node stays down with nothing driving it up. + +package webhook + +import ( + "context" + "slices" + "strings" + "testing" + + admissionv1 "k8s.io/api/admission/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "sigs.k8s.io/controller-runtime/pkg/webhook/admission" + + "github.com/simplyblock/atlas/statemachine" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/cluster" +) + +func runningClusterOpsAt( + step simplyblockv1alpha2.StorageClusterOpsStep, +) *simplyblockv1alpha2.StorageClusterOps { + return &simplyblockv1alpha2.StorageClusterOps{ + ObjectMeta: metav1.ObjectMeta{Name: "a-rolling-restart", Namespace: "simplyblock"}, + Spec: simplyblockv1alpha2.StorageClusterOpsSpec{ + ClusterRef: "simplyblock", + Action: simplyblockv1alpha2.StorageClusterOpsActionRollingRestart, + }, + Status: simplyblockv1alpha2.StorageClusterOpsStatus{ + Phase: simplyblockv1alpha2.StorageClusterOpsPhaseRunning, + Step: statemachine.KubeSnapshot{State: string(step)}, + }, + } +} + +// A record may not be withdrawn from a step the graph declares no abort edge +// from: the cluster or one of its nodes is mid-flight, and this operation is +// what carries it the rest of the way. +func TestStorageClusterOpsDeleteIsRefusedWhereNoAbortEdgeExists(t *testing.T) { + v := &StorageClusterOpsValidator{} + + for _, step := range cluster.UnabortableSteps() { + t.Run(string(step), func(t *testing.T) { + resp := v.Handle(context.Background(), opsDeleteRequest(t, runningClusterOpsAt(step))) + if resp.Allowed { + t.Fatalf("the record was withdrawn at step %s, which cannot be aborted", step) + } + if !strings.Contains(resp.Result.Message, string(step)) { + t.Errorf("the refusal is %q, want it to name the step", resp.Result.Message) + } + }) + } +} + +// The four middle steps of a rolling restart are the sharpest case, because the +// node is offline across all of them. The refusal says that rather than naming +// the step alone. +func TestTheRefusalMidWalkSaysTheNodeIsDown(t *testing.T) { + v := &StorageClusterOpsValidator{} + + for _, step := range []simplyblockv1alpha2.StorageClusterOpsStep{ + simplyblockv1alpha2.StorageClusterOpsStepShuttingDownNode, + simplyblockv1alpha2.StorageClusterOpsStepRefreshingPod, + simplyblockv1alpha2.StorageClusterOpsStepAwaitingPod, + simplyblockv1alpha2.StorageClusterOpsStepRestartingNode, + } { + t.Run(string(step), func(t *testing.T) { + resp := v.Handle(context.Background(), opsDeleteRequest(t, runningClusterOpsAt(step))) + if resp.Allowed { + t.Fatalf("the walk's record was withdrawn at step %s", step) + } + if !strings.Contains(resp.Result.Message, "offline") { + t.Errorf("the refusal is %q, want it to say the node is down", resp.Result.Message) + } + }) + } +} + +// Requesting has issued nothing, CheckingPeers performs no side effect, and +// Rebalancing is the cluster settling on its own, which it finishes whether or +// not this operation is watching. +func TestStorageClusterOpsDeleteIsAllowedWhereTheAbortEdgeExists(t *testing.T) { + v := &StorageClusterOpsValidator{} + + for _, step := range []simplyblockv1alpha2.StorageClusterOpsStep{ + simplyblockv1alpha2.StorageClusterOpsStepRequesting, + simplyblockv1alpha2.StorageClusterOpsStepCheckingPeers, + simplyblockv1alpha2.StorageClusterOpsStepRebalancing, + } { + t.Run(string(step), func(t *testing.T) { + resp := v.Handle(context.Background(), opsDeleteRequest(t, runningClusterOpsAt(step))) + if !resp.Allowed { + t.Errorf("the record was held at step %s, which an abort stops cleanly: %s", + step, resp.Result.Message) + } + }) + } +} + +// A terminal operation is a record of work that has finished, so withdrawing it +// stops nothing whatever step it ended on. +func TestStorageClusterOpsDeleteIsAllowedOnceTerminal(t *testing.T) { + v := &StorageClusterOpsValidator{} + + for _, phase := range []simplyblockv1alpha2.StorageClusterOpsPhase{ + simplyblockv1alpha2.StorageClusterOpsPhaseSucceeded, + simplyblockv1alpha2.StorageClusterOpsPhaseFailed, + simplyblockv1alpha2.StorageClusterOpsPhaseAborted, + } { + t.Run(string(phase), func(t *testing.T) { + ops := runningClusterOpsAt(simplyblockv1alpha2.StorageClusterOpsStepRestartingNode) + ops.Status.Phase = phase + + resp := v.Handle(context.Background(), opsDeleteRequest(t, ops)) + if !resp.Allowed { + t.Errorf("a %s operation could not be deleted: %s", phase, resp.Result.Message) + } + }) + } +} + +// An operation waiting for the cluster's lock holds nothing and is on no step. +func TestStorageClusterOpsDeleteIsAllowedBeforeTheOperationStarted(t *testing.T) { + v := &StorageClusterOpsValidator{} + + ops := runningClusterOpsAt("") + ops.Status.Phase = simplyblockv1alpha2.StorageClusterOpsPhasePending + + resp := v.Handle(context.Background(), opsDeleteRequest(t, ops)) + if !resp.Allowed { + t.Errorf("a Pending operation could not be deleted: %s", resp.Result.Message) + } +} + +// Nothing to read is nothing to refuse on. +func TestStorageClusterOpsDeleteIsAllowedWithNoObjectToRead(t *testing.T) { + v := &StorageClusterOpsValidator{} + + resp := v.Handle(context.Background(), admission.Request{ + AdmissionRequest: admissionv1.AdmissionRequest{Operation: admissionv1.Delete}, + }) + if !resp.Allowed { + t.Errorf("a delete carrying no object was refused: %s", resp.Result.Message) + } +} + +// The guard is registered on DELETE alone, and answers nothing else even if it +// is asked. +func TestStorageClusterOpsCreateIsNotThisGuardsBusiness(t *testing.T) { + v := &StorageClusterOpsValidator{} + + req := opsDeleteRequest(t, + runningClusterOpsAt(simplyblockv1alpha2.StorageClusterOpsStepRestartingNode)) + req.Operation = admissionv1.Create + + if resp := v.Handle(context.Background(), req); !resp.Allowed { + t.Errorf("a create was refused by the delete guard: %s", resp.Result.Message) + } +} + +// The refusal table and the graph are two statements of one rule, and this is +// what makes them agree. A deletion may not express something spec.abort could +// not, so a step in one set and not the other is a defect either way round. +func TestTheClusterRefusalTableAndTheGraphAgree(t *testing.T) { + refused := make([]simplyblockv1alpha2.StorageClusterOpsStep, 0, len(undeletableClusterSteps)) + for step := range undeletableClusterSteps { + refused = append(refused, step) + } + slices.Sort(refused) + + if want := cluster.UnabortableSteps(); !slices.Equal(refused, want) { + t.Errorf("the guard refuses %v, and the graph declares no abort edge from %v", refused, want) + } +} + +// A refusal that names the step and nothing else tells the reader what they may +// not do and never why. +func TestEveryUndeletableClusterStepStatesWhatItIsDoing(t *testing.T) { + for step, doing := range undeletableClusterSteps { + if doing == "" { + t.Errorf("%s is undeletable and the refusal says nothing about why", step) + } + } +} \ No newline at end of file diff --git a/operator/internal/webhook/storagenodeops_validator.go b/operator/internal/webhook/storagenodeops_validator.go new file mode 100644 index 000000000..a4bb01dbf --- /dev/null +++ b/operator/internal/webhook/storagenodeops_validator.go @@ -0,0 +1,125 @@ +// The StorageNodeOps guard: a validating webhook on DELETE that refuses to +// withdraw the record of an operation sitting on a step the graph declares no +// abort edge from. +// +// The operation's finalizer is the other half of this, and it is the weaker +// half: it releases the node's lock and aborts the fan-out, but it runs only on +// a graceful delete. A `kubectl delete --force --grace-period=0` skips it +// entirely, and admission is the only thing left in the path. +// +// What that costs on this kind is specific. A migrate at Promoting has already +// activated the target host's devices, failed the origin's, and re-homed the +// logical volumes, so there is nothing to unwind and the operation is what +// finishes the relocation. The record is what carries the topology re-point that +// is still owed, and withdrawing it leaves spec.workerNode and the worker's +// labels describing a host the node no longer runs on, with nothing left that +// knows to correct them. +// +// Deletion is therefore never a way to express something spec.abort could not: +// both channels ask one graph, and this file asks it through +// node.UnabortableSteps rather than restating the answer. What it states here is +// only the prose, because a refusal has to say what the step is waiting for and +// a bool cannot. +// +// design-crd-model.md §3.1 and design-storagenode.md §11 are the specification. + +package webhook + +import ( + "context" + "encoding/json" + "fmt" + "net/http" + + admissionv1 "k8s.io/api/admission/v1" + "sigs.k8s.io/controller-runtime/pkg/webhook/admission" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// +kubebuilder:webhook:path=/validate-storage-simplyblock-io-v1alpha2-storagenodeops,mutating=false,failurePolicy=fail,sideEffects=None,groups=storage.simplyblock.io,resources=storagenodeops,verbs=delete,versions=v1alpha2,name=vstoragenodeops.simplyblock.io,admissionReviewVersions=v1 + +// StorageNodeOpsValidator refuses a delete that would strand a node operation +// mid-flight. +// +// failurePolicy=Fail, because the webhook server runs inside the operator pod: +// its availability tracks the operator's, and while the operator is down nothing +// advances an operation anyway. +// +// It holds no client. Everything the decision rests on is in the object being +// deleted, which is what makes the answer the same on every replica and immune +// to a stale cache. +type StorageNodeOpsValidator struct{} + +// undeletableNodeSteps says what each step with no abort edge is in the middle +// of. Which steps those are is node.UnabortableSteps's answer, and a test holds +// the two sets equal in both directions: an entry here for a step the graph can +// abort from refuses a delete the abort channel would have honored, and a +// missing entry admits the withdrawal of a record nothing else accounts for. +var undeletableNodeSteps = map[simplyblockv1alpha2.StorageNodeOpsStep]string{ + simplyblockv1alpha2.StorageNodeOpsStepAwaiting: "the control plane is carrying out the " + + "action, and it is doing so whether or not this record exists", + simplyblockv1alpha2.StorageNodeOpsStepRemoving: "the node is being taken out of the cluster", + simplyblockv1alpha2.StorageNodeOpsStepRelocating: "the node is being restarted onto its " + + "target host", + simplyblockv1alpha2.StorageNodeOpsStepAwaitingNode: "the node is part-way through that " + + "restart and this operation is the only thing watching it back", + simplyblockv1alpha2.StorageNodeOpsStepPromoting: "the promote has activated the target " + + "host's devices, failed the origin's, and re-homed the logical volumes, so the topology " + + "re-point is all that is left and this record is what carries it", + simplyblockv1alpha2.StorageNodeOpsStepShuttingDown: "the node is being taken down for host " + + "maintenance", + simplyblockv1alpha2.StorageNodeOpsStepReleasing: "the host is being handed over for " + + "maintenance", + simplyblockv1alpha2.StorageNodeOpsStepAwaitingHost: "the host is away, and this operation " + + "is the only thing that will bring the node back from it", + simplyblockv1alpha2.StorageNodeOpsStepRestarting: "the node is coming back from the " + + "maintenance shutdown", + simplyblockv1alpha2.StorageNodeOpsStepCleanup: "the labels and the disruption budget the " + + "maintenance put in place are being taken back", +} + +func (v *StorageNodeOpsValidator) Handle(_ context.Context, req admission.Request) admission.Response { + if req.Operation != admissionv1.Delete { + return admission.Allowed("") + } + return v.admitDelete(req) +} + +// admitDelete reads the object from req.OldObject, which is what the API server +// sends on a DELETE: there is no new object, and the step the operation is on is +// in the status of the one being removed. +func (v *StorageNodeOpsValidator) admitDelete(req admission.Request) admission.Response { + if len(req.OldObject.Raw) == 0 { + // Nothing to read means nothing to refuse on. Admitting is the only + // answer that does not block a delete on the basis of no information. + return admission.Allowed("") + } + + var ops simplyblockv1alpha2.StorageNodeOps + if err := json.Unmarshal(req.OldObject.Raw, &ops); err != nil { + return admission.Errored(http.StatusBadRequest, err) + } + + // A terminal operation is a record of work that has finished, and + // withdrawing it stops nothing. + switch ops.Status.Phase { + case simplyblockv1alpha2.StorageNodeOpsPhaseSucceeded, + simplyblockv1alpha2.StorageNodeOpsPhaseFailed, + simplyblockv1alpha2.StorageNodeOpsPhaseAborted: + return admission.Allowed("the operation is terminal") + } + + step := simplyblockv1alpha2.StorageNodeOpsStep(ops.Status.Step.State) + doing, undeletable := undeletableNodeSteps[step] + if !undeletable { + return admission.Allowed("") + } + + return admission.Denied(fmt.Sprintf( + "StorageNodeOps %s/%s is at step %s, where %s. Deleting the record would not stop that "+ + "work, it would remove the only thing accounting for it and leave node %s locked by "+ + "an object that no longer exists. Set spec.abort to stop an operation that can still "+ + "be stopped, and delete the record once it is terminal.", + ops.Namespace, ops.Name, step, doing, ops.Spec.NodeRef)) +} \ No newline at end of file diff --git a/operator/internal/webhook/storagenodeops_validator_test.go b/operator/internal/webhook/storagenodeops_validator_test.go new file mode 100644 index 000000000..e42385394 --- /dev/null +++ b/operator/internal/webhook/storagenodeops_validator_test.go @@ -0,0 +1,205 @@ +// What the StorageNodeOps guard admits and refuses. +// +// The delete cases are the whole of it. The controller's teardown releases the +// node's lock and drops the finalizer from any step, so without this webhook a +// `kubectl delete` on a Promoting migrate is admitted, and the topology re-point +// the operation still owes goes with it. +// +// The last test is the one that keeps the guard honest over time: the refusal +// table is checked against the graph's own answer, so a step that becomes +// abortable, or a new step that is not, cannot leave the two disagreeing. + +package webhook + +import ( + "context" + "encoding/json" + "slices" + "strings" + "testing" + + admissionv1 "k8s.io/api/admission/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/runtime" + "sigs.k8s.io/controller-runtime/pkg/webhook/admission" + + "github.com/simplyblock/atlas/statemachine" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/node" +) + +// opsDeleteRequest is the request the API server sends on a DELETE: the object +// being removed arrives in OldObject, and there is no new one. +func opsDeleteRequest(t *testing.T, ops runtime.Object) admission.Request { + t.Helper() + raw, err := json.Marshal(ops) + if err != nil { + t.Fatalf("marshal the operation: %v", err) + } + return admission.Request{AdmissionRequest: admissionv1.AdmissionRequest{ + Operation: admissionv1.Delete, + OldObject: runtime.RawExtension{Raw: raw}, + }} +} + +func runningNodeOpsAt(step simplyblockv1alpha2.StorageNodeOpsStep) *simplyblockv1alpha2.StorageNodeOps { + return &simplyblockv1alpha2.StorageNodeOps{ + ObjectMeta: metav1.ObjectMeta{Name: "a-migration", Namespace: "simplyblock"}, + Spec: simplyblockv1alpha2.StorageNodeOpsSpec{ + NodeRef: "worker-3", + Action: simplyblockv1alpha2.StorageNodeOpsActionMigrate, + }, + Status: simplyblockv1alpha2.StorageNodeOpsStatus{ + Phase: simplyblockv1alpha2.StorageNodeOpsPhaseRunning, + Step: statemachine.KubeSnapshot{State: string(step)}, + }, + } +} + +// A record may not be withdrawn from a step the graph declares no abort edge +// from. Each of these has taken the node out of service or moved it, and the +// operation is the only thing that finishes what it started. +func TestStorageNodeOpsDeleteIsRefusedWhereNoAbortEdgeExists(t *testing.T) { + v := &StorageNodeOpsValidator{} + + for _, step := range node.UnabortableSteps() { + t.Run(string(step), func(t *testing.T) { + resp := v.Handle(context.Background(), opsDeleteRequest(t, runningNodeOpsAt(step))) + if resp.Allowed { + t.Fatalf("the record was withdrawn at step %s, which cannot be aborted", step) + } + if !strings.Contains(resp.Result.Message, string(step)) { + t.Errorf("the refusal is %q, want it to name the step", resp.Result.Message) + } + }) + } +} + +// Promoting is the case design-storagenode.md §11 names: the promote has +// re-homed the logical volumes and the record is what carries the topology +// re-point still owed, so the refusal says so rather than naming the step alone. +func TestTheRefusalAtPromotingSaysWhatTheRecordStillOwes(t *testing.T) { + v := &StorageNodeOpsValidator{} + + resp := v.Handle(context.Background(), + opsDeleteRequest(t, runningNodeOpsAt(simplyblockv1alpha2.StorageNodeOpsStepPromoting))) + if resp.Allowed { + t.Fatal("a Promoting migrate was withdrawn") + } + if !strings.Contains(resp.Result.Message, "topology") { + t.Errorf("the refusal is %q, want it to name the re-point the record carries", + resp.Result.Message) + } +} + +// A step the graph can abort from is one whose unwind exists, so withdrawing the +// record there leaves nothing the finalizer cannot finish. +func TestStorageNodeOpsDeleteIsAllowedWhereTheAbortEdgeExists(t *testing.T) { + v := &StorageNodeOpsValidator{} + + for _, step := range []simplyblockv1alpha2.StorageNodeOpsStep{ + simplyblockv1alpha2.StorageNodeOpsStepRequesting, + simplyblockv1alpha2.StorageNodeOpsStepValidating, + simplyblockv1alpha2.StorageNodeOpsStepSuspending, + simplyblockv1alpha2.StorageNodeOpsStepMigratingVolumes, + simplyblockv1alpha2.StorageNodeOpsStepVerifying, + simplyblockv1alpha2.StorageNodeOpsStepPreparing, + simplyblockv1alpha2.StorageNodeOpsStepHolding, + } { + t.Run(string(step), func(t *testing.T) { + resp := v.Handle(context.Background(), opsDeleteRequest(t, runningNodeOpsAt(step))) + if !resp.Allowed { + t.Errorf("the record was held at step %s, which an abort stops cleanly: %s", + step, resp.Result.Message) + } + }) + } +} + +// A terminal operation is a record of work that has finished, so withdrawing it +// stops nothing whatever step it ended on. +func TestStorageNodeOpsDeleteIsAllowedOnceTerminal(t *testing.T) { + v := &StorageNodeOpsValidator{} + + for _, phase := range []simplyblockv1alpha2.StorageNodeOpsPhase{ + simplyblockv1alpha2.StorageNodeOpsPhaseSucceeded, + simplyblockv1alpha2.StorageNodeOpsPhaseFailed, + simplyblockv1alpha2.StorageNodeOpsPhaseAborted, + } { + t.Run(string(phase), func(t *testing.T) { + ops := runningNodeOpsAt(simplyblockv1alpha2.StorageNodeOpsStepPromoting) + ops.Status.Phase = phase + + resp := v.Handle(context.Background(), opsDeleteRequest(t, ops)) + if !resp.Allowed { + t.Errorf("a %s operation could not be deleted: %s", phase, resp.Result.Message) + } + }) + } +} + +// An operation that never started holds nothing and is on no step. +func TestStorageNodeOpsDeleteIsAllowedBeforeTheOperationStarted(t *testing.T) { + v := &StorageNodeOpsValidator{} + + ops := runningNodeOpsAt("") + ops.Status.Phase = simplyblockv1alpha2.StorageNodeOpsPhasePending + + resp := v.Handle(context.Background(), opsDeleteRequest(t, ops)) + if !resp.Allowed { + t.Errorf("a Pending operation could not be deleted: %s", resp.Result.Message) + } +} + +// Nothing to read is nothing to refuse on. Blocking a delete on the basis of no +// information is the one answer that cannot be justified. +func TestStorageNodeOpsDeleteIsAllowedWithNoObjectToRead(t *testing.T) { + v := &StorageNodeOpsValidator{} + + resp := v.Handle(context.Background(), admission.Request{ + AdmissionRequest: admissionv1.AdmissionRequest{Operation: admissionv1.Delete}, + }) + if !resp.Allowed { + t.Errorf("a delete carrying no object was refused: %s", resp.Result.Message) + } +} + +// The guard is registered on DELETE alone, and answers nothing else even if it +// is asked. +func TestStorageNodeOpsCreateIsNotThisGuardsBusiness(t *testing.T) { + v := &StorageNodeOpsValidator{} + + req := opsDeleteRequest(t, runningNodeOpsAt(simplyblockv1alpha2.StorageNodeOpsStepPromoting)) + req.Operation = admissionv1.Create + + if resp := v.Handle(context.Background(), req); !resp.Allowed { + t.Errorf("a create was refused by the delete guard: %s", resp.Result.Message) + } +} + +// The refusal table and the graph are two statements of one rule, and this is +// what makes them agree. A deletion may not express something spec.abort could +// not, so a step in one set and not the other is a defect either way round: an +// extra entry refuses a delete the abort channel would have honored, and a +// missing one admits the withdrawal of a record nothing else accounts for. +func TestTheNodeRefusalTableAndTheGraphAgree(t *testing.T) { + refused := make([]simplyblockv1alpha2.StorageNodeOpsStep, 0, len(undeletableNodeSteps)) + for step := range undeletableNodeSteps { + refused = append(refused, step) + } + slices.Sort(refused) + + if want := node.UnabortableSteps(); !slices.Equal(refused, want) { + t.Errorf("the guard refuses %v, and the graph declares no abort edge from %v", refused, want) + } +} + +// A refusal that names the step and nothing else tells the reader what they may +// not do and never why. +func TestEveryUndeletableNodeStepStatesWhatItIsDoing(t *testing.T) { + for step, doing := range undeletableNodeSteps { + if doing == "" { + t.Errorf("%s is undeletable and the refusal says nothing about why", step) + } + } +} \ No newline at end of file From b0c7e3ac1633ece952af0904777c19bc0fcfceb4 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 10:01:44 +0200 Subject: [PATCH 028/206] refactor(statemachine): the graph declares which states an abort can stop from MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Four packages kept a map of abortable steps beside their graphs, plus a predicate over it and a test asserting the map only named steps some graph declared. StateDef.Abortable puts the property where the state is, so the map key is the state and that class of drift stops existing. The sharper gain is precision. Abortability is a property of (action, step), but a table keyed on the step alone flattens every action of a MultiConfig together. Nothing collides today, and nothing could express it if it did. Inside StateDef the question cannot be asked wrongly, and the reconcilers now ask Machine.CanAbort, which answers for the graph the machine was built from. The callers that hold a step and no machine — the DELETE guards reading a step out of a status — ask UnabortableStates or UnabortableMultiStates instead. The multi form unions across actions, because that is what a per-step table can hold, and it takes the union rather than the intersection so a guard reading it refuses where two actions disagree. StorageClusterOps, StorageNodeOps, and now ControlPlaneOps hold their refusal tables equal to that answer in both directions. StorageBackupOps deliberately does not: its guard refuses Restoring, which the graph can abort from, because a delete leaves nothing behind where an abort leaves an object the controller cleans up from. Binding a test there would force the weaker answer onto the stronger channel, and the file now says so. Also fixes the missing trailing newlines and the two linter findings the previous two commits left in the webhook package. Co-Authored-By: Claude Fable 5 --- atlas-lib/README.md | 1 + atlas-lib/statemachine/abort.go | 80 ++++++++++ atlas-lib/statemachine/abort_test.go | 142 ++++++++++++++++++ atlas-lib/statemachine/statemachine.go | 31 ++++ .../internal/controllers/backup/restore.go | 35 ++--- .../backup/storagebackupops_controller.go | 11 +- .../storagebackupops_controller_test.go | 25 +-- .../internal/controllers/cluster/graphs.go | 79 ++++------ .../controllers/cluster/graphs_test.go | 12 -- .../controllers/cluster/review_test.go | 11 +- .../cluster/storageclusterops_controller.go | 12 +- .../controlplaneops_controller.go | 8 +- .../controlplane/controlplaneops_test.go | 29 ++-- .../controllers/controlplane/graphs.go | 50 +++--- operator/internal/controllers/node/graphs.go | 113 +++++++------- .../internal/controllers/node/graphs_test.go | 49 ++++-- .../node/storagenodeops_controller.go | 12 +- .../clusterdeploymentconfig_validator.go | 2 +- .../clusterdeploymentconfig_validator_test.go | 36 +++-- .../webhook/controlplaneops_validator_test.go | 19 +++ .../webhook/storagebackupops_validator.go | 17 ++- .../webhook/storageclusterops_validator.go | 2 +- .../storageclusterops_validator_test.go | 2 +- .../webhook/storagenodeops_validator.go | 2 +- .../webhook/storagenodeops_validator_test.go | 2 +- 25 files changed, 537 insertions(+), 245 deletions(-) create mode 100644 atlas-lib/statemachine/abort.go create mode 100644 atlas-lib/statemachine/abort_test.go diff --git a/atlas-lib/README.md b/atlas-lib/README.md index bbe5919e5..bf842f77e 100644 --- a/atlas-lib/README.md +++ b/atlas-lib/README.md @@ -139,6 +139,7 @@ atlas/ ├── statemachine/ Deterministic state machine declared as data │ ├── statemachine.go Config, StateDef, Machine, Snapshot, deadlines │ ├── multiconfig.go MultiConfig: one graph per action over one state type +│ ├── abort.go StateDef.Abortable read three ways: CanAbort + the two graph queries │ └── kubernetes.go KubeSnapshot + ToKube/FromKube: the CRD form of a Snapshot ├── net/ Outbound URL validation (SSRF guard) ├── ptr/ Pointer/optional-field helpers for generated + K8s types diff --git a/atlas-lib/statemachine/abort.go b/atlas-lib/statemachine/abort.go new file mode 100644 index 000000000..b3fcc3587 --- /dev/null +++ b/atlas-lib/statemachine/abort.go @@ -0,0 +1,80 @@ +// Asking the graph whether the work can still be called off. +// +// [StateDef.Abortable] is where the answer is declared, and this file is the +// three ways it is read. A caller holding a machine asks [Machine.CanAbort], +// which is the precise question, because the machine was built for one action +// and knows which state it is in. A caller holding a state and no machine — an +// admission webhook reading a step out of a status — asks one of the graph +// queries instead. +// +// They live beside the field rather than in statemachine.go because the property +// is one concern, and a reader who has found any part of it has found all of it. + +package statemachine + +import "slices" + +// CanAbort reports whether the state the machine is in declares itself +// abortable, which is the graph's answer to a caller asking to stop: +// +// if !machine.CanAbort() { +// // Not a failure of the operation: it carries on, and what the caller +// // asked for is what could not be done. +// return r.note(ctx, ops, "the abort arrived too late to be honored") +// } +// +// It says nothing about how the stop is then performed. Unwinding whatever the +// earlier states started is the caller's, because only the caller knows what +// they were. +func (sm *Machine[S]) CanAbort() bool { + return sm.states[sm.current].Abortable +} + +// UnabortableStates returns the states a graph declares no abort from, sorted. +// +// It exists for the reader that has a state and no machine, and for the test +// that holds a table beside the graph honest. A guard refusing to withdraw the +// record of a running operation is both: it reads a step out of a status, and it +// keeps its own table of what each step is in the middle of, which only agrees +// with the graph if something makes it. +func UnabortableStates[S ~string](config Config[S]) []S { + return unabortable(config.States) +} + +// UnabortableMultiStates returns the states no action can be stopped from, as +// the union across every declared graph, sorted. +// +// The union is the answer a table keyed on the state alone can hold, since one +// state shared by two actions has one entry and may have two answers. Taking the +// union rather than the intersection is the conservative half of that: where two +// actions disagree, the state is reported, so a guard reading this refuses where +// it might have admitted rather than the other way round. A caller that needs +// the precise answer has the action in hand and should ask that graph, or ask +// [Machine.CanAbort] on the machine built from it. +func UnabortableMultiStates[S ~string](graphs MultiConfig[S]) []S { + union := make(map[S]StateDef[S]) + for _, config := range graphs { + for state, def := range config.States { + if previous, seen := union[state]; seen && !previous.Abortable { + // Already reported by another action. Keeping the stricter of + // the two is what makes this a union of the unabortable rather + // than a last-writer-wins over the map. + continue + } + union[state] = def + } + } + return unabortable(union) +} + +// unabortable is the filter both queries end in. +func unabortable[S ~string](states map[S]StateDef[S]) []S { + var refused []S + for state, def := range states { + if !def.Abortable { + refused = append(refused, state) + } + } + slices.Sort(refused) + return refused +} diff --git a/atlas-lib/statemachine/abort_test.go b/atlas-lib/statemachine/abort_test.go new file mode 100644 index 000000000..a1f3f3ad2 --- /dev/null +++ b/atlas-lib/statemachine/abort_test.go @@ -0,0 +1,142 @@ +// What the graph says about stopping, and who may ask. +// +// The property is declared per state and read three ways: by the machine, for +// the caller holding one; by a graph query, for the guard that has a state and +// no machine; and across a MultiConfig, for the one table a kind keeps beside +// several actions. + +package statemachine + +import ( + "context" + "slices" + "testing" +) + +type abortState string + +const ( + abortPending abortState = "Pending" + abortWorking abortState = "Working" + abortFinished abortState = "Finished" +) + +// abortGraph is a line of three: nothing has started, something has, and it is +// over. Only the first can be called off. +func abortGraph() Config[abortState] { + return Config[abortState]{ + Initial: abortPending, + States: map[abortState]StateDef[abortState]{ + abortPending: {To: []abortState{abortWorking}, Abortable: true}, + abortWorking: {To: []abortState{abortFinished}}, + abortFinished: {}, + }, + } +} + +// The machine answers for the state it is in, which is the question a caller +// holding one actually has: it knows the action and the position, and wants to +// know whether the stop it was asked for can be honored. +func TestCanAbortAnswersForTheCurrentState(t *testing.T) { + sm, err := New(context.Background(), abortGraph()) + if err != nil { + t.Fatalf("build the machine: %v", err) + } + defer sm.Close() + + if !sm.CanAbort() { + t.Error("the initial state declares Abortable and the machine says it cannot be aborted") + } + + if err := sm.TransitionTo(context.Background(), abortWorking); err != nil { + t.Fatalf("move to Working: %v", err) + } + if sm.CanAbort() { + t.Error("Working declares no Abortable and the machine says it can be aborted") + } +} + +// A state that says nothing is not abortable. The default is what carries the +// rule, because a graph that has to opt out of stopping would make forgetting +// the field the dangerous direction. +func TestAStateThatSaysNothingIsNotAbortable(t *testing.T) { + config := Config[abortState]{ + Initial: abortWorking, + States: map[abortState]StateDef[abortState]{abortWorking: {}}, + } + sm, err := New(context.Background(), config) + if err != nil { + t.Fatalf("build the machine: %v", err) + } + defer sm.Close() + + if sm.CanAbort() { + t.Error("a state declaring nothing reports itself abortable") + } +} + +// Being terminal and being abortable are different questions, and nothing +// derives one from the other: a terminal state is where work ended, which is not +// a place work can be called off from. +func TestATerminalStateIsNotAbortableByItself(t *testing.T) { + if got := UnabortableStates(abortGraph()); !slices.Contains(got, abortFinished) { + t.Errorf("the unabortable states are %v, want the terminal Finished among them", got) + } +} + +// The graph query is for a caller with a state and no machine, such as an +// admission webhook reading a step out of a status. +func TestUnabortableStatesListsWhatTheGraphWillNotStop(t *testing.T) { + want := []abortState{abortFinished, abortWorking} + if got := UnabortableStates(abortGraph()); !slices.Equal(got, want) { + t.Errorf("the unabortable states are %v, want %v sorted", got, want) + } +} + +// Across a MultiConfig the answer is the union, because a table keyed on the +// state alone cannot hold two answers for one state. The union is the +// conservative half of that: a state one action cannot be stopped from is +// reported, so a guard reading this refuses rather than admits where the two +// actions disagree. +func TestUnabortableMultiStatesUnionsTheActions(t *testing.T) { + graphs := MultiConfig[abortState]{ + "quick": { + Initial: abortPending, + States: map[abortState]StateDef[abortState]{ + abortPending: {Abortable: true}, + }, + }, + "slow": { + Initial: abortPending, + States: map[abortState]StateDef[abortState]{ + // The same state, and this action has already started something + // by the time it is here. + abortPending: {To: []abortState{abortWorking}}, + abortWorking: {}, + }, + }, + } + + want := []abortState{abortPending, abortWorking} + if got := UnabortableMultiStates(graphs); !slices.Equal(got, want) { + t.Errorf("the unabortable states are %v, want %v", got, want) + } +} + +// A graph that can be stopped from everywhere reports nothing, rather than +// reporting every state. +func TestAFullyAbortableGraphHasNoUnabortableStates(t *testing.T) { + graphs := MultiConfig[abortState]{ + "discover": { + Initial: abortPending, + States: map[abortState]StateDef[abortState]{ + abortPending: {To: []abortState{abortWorking}, Abortable: true}, + abortWorking: {Abortable: true}, + }, + }, + } + + if got := UnabortableMultiStates(graphs); len(got) != 0 { + t.Errorf("a graph abortable from every state reports %v as unabortable", got) + } +} diff --git a/atlas-lib/statemachine/statemachine.go b/atlas-lib/statemachine/statemachine.go index 6de3a375a..d540f165e 100644 --- a/atlas-lib/statemachine/statemachine.go +++ b/atlas-lib/statemachine/statemachine.go @@ -239,6 +239,23 @@ // persisting the phase without the deadline yields one that can never time out, // which is why [Machine.Snapshot] returns both together. // +// # Stopping +// +// Aborting is an edge above, because a VolumeMigration's phase type already has +// an Aborted value for the edge to point at. A resource whose states are the +// steps of a workflow usually does not: its outcome is recorded in an outer +// phase, and an Aborted step would be a value in the step enum that no step ever +// means. [StateDef.Abortable] is the same rule declared as a property of the +// state instead, so the graph still answers whether a stop can be honored +// without the state type paying for it: +// +// promoting: {Abortable: false, OnEnter: r.onPromoting(op)}, +// +// A caller holding the machine asks [Machine.CanAbort]. A caller holding only a +// persisted state — an admission webhook reading a status — asks +// [UnabortableStates] or [UnabortableMultiStates], which is also what keeps a +// table written beside the graph from drifting away from it. +// // # One graph per action // // A VolumeMigration does one thing, so one graph describes it. An Ops resource @@ -366,6 +383,20 @@ type StateDef[S comparable] struct { // error. To []S + // Abortable declares that the work this state represents can still be called + // off. It is the graph's answer to a caller asking to stop, and it defaults + // to false because a state that has started something is the common case and + // the safe default: the graph has to say a stop is safe, rather than say it + // is not. + // + // It is a property of the state rather than an edge to a terminal one, + // because the alternative costs a state. A resource whose state type is its + // own persisted enum would need an extra value in that enum for every + // workflow that can be aborted, and the outcome is usually already recorded + // somewhere else — in the outer phase, for the Ops kinds this was built for. + // See [Machine.CanAbort], which is what a caller asks. + Abortable bool + // OnEnter runs when the machine enters this state. It may be nil, in which // case entering always succeeds and leaves the state without a deadline. OnEnter TransitionFunc[S] diff --git a/operator/internal/controllers/backup/restore.go b/operator/internal/controllers/backup/restore.go index f229fb69c..3c1fe0b60 100644 --- a/operator/internal/controllers/backup/restore.go +++ b/operator/internal/controllers/backup/restore.go @@ -61,18 +61,6 @@ const ( bindingDeadline = 15 * time.Minute ) -// abortableSteps are the steps from which an abort unwinds cleanly. -// -// It is a table beside the graph rather than an edge in it, because the kind's -// step enum has four values and none of them is an abort state: a terminal -// Aborted step would be a fifth value in the API, and the phase already carries -// that meaning. The test in this package's suite asserts that every step here is -// one the graph declares, so the two cannot drift. -var abortableSteps = map[step]bool{ - stepValidating: true, - stepRestoring: true, -} - // restoreGraph is the Restore action's declared steps. The deadlines are set on // entry, which is what makes a step that outlived its own budget detectable // after an operator restart: the instant is absolute and travels in the status. @@ -84,8 +72,18 @@ func restoreGraph() statemachine.MultiConfig[step] { actionRestore: { Initial: stepValidating, States: map[step]statemachine.StateDef[step]{ - stepValidating: {To: []step{stepRestoring}, OnEnter: deadline(validatingDeadline)}, - stepRestoring: {To: []step{stepAwaitingVolume}, OnEnter: deadline(restoringDeadline)}, + // Abortable draws the line this file's opening comment + // describes: where a logical volume comes into existence. + stepValidating: { + To: []step{stepRestoring}, + Abortable: true, + OnEnter: deadline(validatingDeadline), + }, + stepRestoring: { + To: []step{stepAwaitingVolume}, + Abortable: true, + OnEnter: deadline(restoringDeadline), + }, stepAwaitingVolume: {To: []step{stepBinding}, OnEnter: deadline(awaitingVolumeDeadline)}, stepBinding: {OnEnter: deadline(bindingDeadline)}, }, @@ -101,6 +99,9 @@ func restoreGraph() statemachine.MultiConfig[step] { // being the one step that cannot time out. const initialDeadline = validatingDeadline -// abortable reports whether an abort asked for while the operation sits on this -// step can be honored. -func abortable(current step) bool { return abortableSteps[current] } +// unabortableSteps are the steps the graph declares no abort from, which is what +// this kind's DELETE guard and its tests read. The reconciler asks its machine +// instead, because the machine was built for the action in hand. +func unabortableSteps() []step { + return statemachine.UnabortableMultiStates(restoreGraph()) +} diff --git a/operator/internal/controllers/backup/storagebackupops_controller.go b/operator/internal/controllers/backup/storagebackupops_controller.go index d2ee1a255..a45de716c 100644 --- a/operator/internal/controllers/backup/storagebackupops_controller.go +++ b/operator/internal/controllers/backup/storagebackupops_controller.go @@ -215,7 +215,7 @@ func (r *StorageBackupOpsReconciler) advance( current := machine.CurrentState() if ops.Spec.Abort { - return r.unwind(ctx, ops, current) + return r.unwind(ctx, ops, machine, current) } if machine.TimeoutReached() { @@ -283,9 +283,14 @@ func (r *StorageBackupOpsReconciler) enterInitialStep( // something the graph cannot take back, and stopping there would leave the // system holding it with no record of what it is for. func (r *StorageBackupOpsReconciler) unwind( - ctx context.Context, ops *simplyblockv1alpha2.StorageBackupOps, current step, + ctx context.Context, + ops *simplyblockv1alpha2.StorageBackupOps, + machine *statemachine.Machine[step], + current step, ) (ctrl.Result, error) { - if !abortable(current) { + // The machine is asked rather than a table beside it: the graph it was built + // from is the one authority over what this action can stop from. + if !machine.CanAbort() { // Not a failure of the operation: it carries on. What the user asked for // cannot be done, and saying so is the whole of the response. return ctrl.Result{RequeueAfter: opsRetry}, r.note(ctx, ops, fmt.Sprintf( diff --git a/operator/internal/controllers/backup/storagebackupops_controller_test.go b/operator/internal/controllers/backup/storagebackupops_controller_test.go index 1ef9a8d16..f8811ecc6 100644 --- a/operator/internal/controllers/backup/storagebackupops_controller_test.go +++ b/operator/internal/controllers/backup/storagebackupops_controller_test.go @@ -44,23 +44,28 @@ func TestDeclaredStepsMatchTheKindsEnum(t *testing.T) { } } -// Every step an abort is honored from has to be one the graph declares, or the -// table and the graph would disagree about what the action can do. -func TestAbortableStepsAreDeclaredByTheGraph(t *testing.T) { - declared := statemachine.DeclaredMultiStates(restoreGraph()) - for abortableStep := range abortableSteps { - if !slices.Contains(declared, string(abortableStep)) { - t.Errorf("abortableSteps names %q, which the graph does not declare", abortableStep) - } - } +// Where the abort is honored and where it is refused, written out rather than +// derived: the graph is the only place this is declared now, so a step quietly +// gaining or losing it would otherwise change what an abort does with nothing +// disagreeing. +func TestTheGraphDeclaresWhereAnAbortIsHonored(t *testing.T) { + refused := unabortableSteps() + // The two the design fixes as the point of no return. A step that gained an // abort edge without the graph gaining a way to unwind it would let an // operation stop with a volume nothing accounts for. for _, beyondReturn := range []step{stepAwaitingVolume, stepBinding} { - if abortable(beyondReturn) { + if !slices.Contains(refused, beyondReturn) { t.Errorf("step %q is abortable, and it has already created a logical volume", beyondReturn) } } + // Validating has created nothing, and Restoring deletes whatever its request + // produced. + for _, stoppable := range []step{stepValidating, stepRestoring} { + if slices.Contains(refused, stoppable) { + t.Errorf("step %q refuses an abort, and its unwind exists", stoppable) + } + } } // testOpsName is the operation every test in this file drives, named once diff --git a/operator/internal/controllers/cluster/graphs.go b/operator/internal/controllers/cluster/graphs.go index a1a7259a2..ae41d1bde 100644 --- a/operator/internal/controllers/cluster/graphs.go +++ b/operator/internal/controllers/cluster/graphs.go @@ -80,8 +80,15 @@ func graphs() statemachine.MultiConfig[step] { return statemachine.Config[step]{ Initial: stepRequesting, States: map[step]statemachine.StateDef[step]{ - stepRequesting: {To: []step{stepAwaiting}, OnEnter: deadline(requestingDeadline)}, - stepAwaiting: {OnEnter: deadline(awaitingDeadline)}, + // Requesting has issued nothing. Awaiting is past the call, and + // a cluster told to shut down is shutting down whatever this + // object says. + stepRequesting: { + To: []step{stepAwaiting}, + Abortable: true, + OnEnter: deadline(requestingDeadline), + }, + stepAwaiting: {OnEnter: deadline(awaitingDeadline)}, }, } } @@ -111,9 +118,18 @@ func graphs() statemachine.MultiConfig[step] { action(simplyblockv1alpha2.StorageClusterOpsActionRollingRestart): { Initial: stepCheckingPeers, States: map[step]statemachine.StateDef[step]{ + // CheckingPeers performs no side effect and is the step before + // the walk touches a node. From ShuttingDownNode to + // RestartingNode the node is offline and this operation is the + // only thing that will bring it back, so an abort honored there + // would leave a storage node down with nothing driving it up. + // Rebalancing is after the node is back and the cluster is + // settling on its own, which it finishes whether or not this + // operation is watching. stepCheckingPeers: { - To: []step{stepShuttingDownNode}, - OnEnter: deadline(checkingPeersDeadline), + To: []step{stepShuttingDownNode}, + Abortable: true, + OnEnter: deadline(checkingPeersDeadline), }, stepShuttingDownNode: { To: []step{stepRefreshingPod, stepRestartingNode}, @@ -135,7 +151,7 @@ func graphs() statemachine.MultiConfig[step] { // machine never carries a cycle. Starting the next node is // Machine.Reset, which returns to CheckingPeers, clears the // deadline, validates no edge, and runs no hook. - stepRebalancing: {OnEnter: deadline(rebalancingDeadline)}, + stepRebalancing: {Abortable: true, OnEnter: deadline(rebalancingDeadline)}, }, }, } @@ -156,55 +172,26 @@ var initialDeadlines = map[statemachine.Action]time.Duration{ action(simplyblockv1alpha2.StorageClusterOpsActionRollingRestart): checkingPeersDeadline, } -// abortableSteps are the steps from which an abort stops the operation cleanly. -// -// The line is whether anything is currently down or half-done. Requesting has -// issued nothing. CheckingPeers performs no side effect at all and is the step -// before the walk touches a node. Rebalancing is after the node is back online -// and the cluster is settling on its own, which it will finish whether or not -// this operation is watching. +// Which steps an abort stops cleanly is declared on the states above, because +// the line it draws is a property of the step rather than of this kind: whether +// anything is currently down or half-done. Awaiting, ShuttingDown, and Starting +// are that rule at cluster scale, and the rolling restart's four middle steps +// are the sharpest case of it at node scale. // -// Everything else is mid-flight, and the rolling restart's four middle steps -// are the sharpest case: from ShuttingDownNode to RestartingNode the node is -// offline, and this operation is the only thing that will bring it back. -// Honoring an abort there would leave a storage node down with nothing driving -// it up, so the walk runs on to the restart and the abort is refused until -// then. Awaiting, ShuttingDown, and Starting are the same rule at cluster -// scale: a cluster told to shut down is shutting down whatever this object -// says. -// -// It is a table beside the graph rather than an edge in it, because a terminal -// Aborted step would be an eleventh value in the API and the phase already -// carries that meaning. A test asserts every step here is one some graph -// declares, and another asserts that no step between a node's shutdown and its -// restart appears, so the two cannot drift. -var abortableSteps = map[step]bool{ - stepRequesting: true, - stepCheckingPeers: true, - stepRebalancing: true, -} - -// abortable reports whether an abort asked for while the operation sits on this -// step can be honored. -func abortable(current step) bool { return abortableSteps[current] } +// It is a property of the state rather than an edge to a terminal one, because a +// terminal Aborted step would be an eleventh value in the API and the phase +// already carries that meaning. // UnabortableSteps are the declared steps an abort cannot be honored from, // sorted. // // It is exported for the DELETE guard on this kind, which asks the same question // this package's unwind asks: a deletion may not express something spec.abort -// could not, so both channels read one graph -// (design-crd-model.md §3.1). Reading it rather than restating it is what keeps -// the guard from refusing a step this package has since made abortable, or -// admitting one it has not. +// could not, so both channels read one graph (design-crd-model.md §3.1). The +// guard has a step out of a status and no machine, which is the whole reason +// this reads the graphs rather than the machine the reconciler holds. func UnabortableSteps() []step { - var refused []step - for _, declared := range statemachine.DeclaredMultiStates(graphs()) { - if s := step(declared); !abortable(s) { - refused = append(refused, s) - } - } - return refused + return statemachine.UnabortableMultiStates(graphs()) } // action converts the API's action enum into the MultiConfig's key. The diff --git a/operator/internal/controllers/cluster/graphs_test.go b/operator/internal/controllers/cluster/graphs_test.go index 9b2146c45..413d39d2b 100644 --- a/operator/internal/controllers/cluster/graphs_test.go +++ b/operator/internal/controllers/cluster/graphs_test.go @@ -102,18 +102,6 @@ func TestEveryActionDeclaresAGraph(t *testing.T) { } } -// An abortable step that no graph declares is a table that has drifted from the -// graphs beside it, and the drift is silent: an abort would simply never be -// honored from it. -func TestEveryAbortableStepIsDeclared(t *testing.T) { - declared := statemachine.DeclaredMultiStates(graphs()) - for s := range abortableSteps { - if !slices.Contains(declared, string(s)) { - t.Errorf("abortableSteps names %q, which no graph declares", s) - } - } -} - // Every step needs a budget in stepBudgets, or its duration is never measured // and a slow operation stays anecdotal. func TestEveryStepHasABudget(t *testing.T) { diff --git a/operator/internal/controllers/cluster/review_test.go b/operator/internal/controllers/cluster/review_test.go index 1c33a350b..3754a3f82 100644 --- a/operator/internal/controllers/cluster/review_test.go +++ b/operator/internal/controllers/cluster/review_test.go @@ -11,6 +11,7 @@ package cluster import ( "context" + "slices" "testing" corev1 "k8s.io/api/core/v1" @@ -29,14 +30,15 @@ import ( "github.com/simplyblock/simplyblock-operator/internal/webapi" ) -// The two steps between a node's shutdown and its restart are not abortable. +// The four steps between a node's shutdown and its restart are not abortable. // Stopping there leaves the node offline with nothing driving it back up, -// which is the one outcome the abort table exists to prevent. +// which is the one outcome abortability exists to prevent. func TestAnAbortIsRefusedWhileTheNodeIsDown(t *testing.T) { + unabortable := statemachine.UnabortableMultiStates(graphs()) for _, current := range []step{stepShuttingDownNode, stepRefreshingPod, stepAwaitingPod, stepRestartingNode} { t.Run(string(current), func(t *testing.T) { - if abortable(current) { + if !slices.Contains(unabortable, current) { t.Errorf("step %s is abortable, but the node is offline there and the "+ "walk is what brings it back", current) } @@ -47,9 +49,10 @@ func TestAnAbortIsRefusedWhileTheNodeIsDown(t *testing.T) { // The abort is honored where the walk has not taken anything down: before the // shutdown, and after the node is back online and the cluster is settling. func TestAnAbortIsHonoredWhereNothingIsDown(t *testing.T) { + unabortable := statemachine.UnabortableMultiStates(graphs()) for _, current := range []step{stepCheckingPeers, stepRebalancing} { t.Run(string(current), func(t *testing.T) { - if !abortable(current) { + if slices.Contains(unabortable, current) { t.Errorf("step %s is not abortable, though it has taken nothing down", current) } }) diff --git a/operator/internal/controllers/cluster/storageclusterops_controller.go b/operator/internal/controllers/cluster/storageclusterops_controller.go index 517c9a74b..4f0ab00f8 100644 --- a/operator/internal/controllers/cluster/storageclusterops_controller.go +++ b/operator/internal/controllers/cluster/storageclusterops_controller.go @@ -246,7 +246,7 @@ func (r *StorageClusterOpsReconciler) advance( current := machine.CurrentState() if ops.Spec.Abort { - return r.unwind(ctx, ops, current) + return r.unwind(ctx, ops, machine, current) } if machine.TimeoutReached() { @@ -376,9 +376,15 @@ func (r *StorageClusterOpsReconciler) nextStep( // control plane for something it is part-way through, and stopping there would // leave nothing driving the cluster back to a state somebody can reason about. func (r *StorageClusterOpsReconciler) unwind( - ctx context.Context, ops *simplyblockv1alpha2.StorageClusterOps, current step, + ctx context.Context, + ops *simplyblockv1alpha2.StorageClusterOps, + machine *statemachine.Machine[step], + current step, ) (ctrl.Result, error) { - if !abortable(current) { + // The machine is asked rather than a table beside it, and it is asked rather + // than the graphs, because it was built for this operation's action: a step + // two actions share can be abortable in one of them. + if !machine.CanAbort() { // Not a failure of the operation: it carries on. What the user asked // for cannot be done, and saying so is the whole of the response. return ctrl.Result{RequeueAfter: opsRetry}, r.note(ctx, ops, fmt.Sprintf( diff --git a/operator/internal/controllers/controlplane/controlplaneops_controller.go b/operator/internal/controllers/controlplane/controlplaneops_controller.go index 99746a374..d70f09c1c 100644 --- a/operator/internal/controllers/controlplane/controlplaneops_controller.go +++ b/operator/internal/controllers/controlplane/controlplaneops_controller.go @@ -196,7 +196,7 @@ func (r *ControlPlaneOpsReconciler) advance( current := machine.CurrentState() if ops.Spec.Abort { - return r.unwind(ctx, ops, target, current) + return r.unwind(ctx, ops, target, machine, current) } if machine.TimeoutReached() { @@ -712,9 +712,13 @@ func (r *ControlPlaneOpsReconciler) unwind( ctx context.Context, ops *simplyblockv1alpha2.ControlPlaneOps, target *simplyblockv1alpha2.ControlPlane, + machine *statemachine.Machine[opsStep], current opsStep, ) (ctrl.Result, error) { - if !abortable(current) { + // The machine is asked rather than a table beside it, and it is asked rather + // than the graphs, because it was built for this operation's action: a step + // two actions share can be abortable in one of them. + if !machine.CanAbort() { // Not a failure of the operation: it carries on. What the user asked for // cannot be done, and saying so is the whole of the response. return ctrl.Result{RequeueAfter: opsRetry}, r.note(ctx, ops, fmt.Sprintf( diff --git a/operator/internal/controllers/controlplane/controlplaneops_test.go b/operator/internal/controllers/controlplane/controlplaneops_test.go index 998013ddf..f04ea94c6 100644 --- a/operator/internal/controllers/controlplane/controlplaneops_test.go +++ b/operator/internal/controllers/controlplane/controlplaneops_test.go @@ -11,6 +11,7 @@ package controlplane import ( "context" "errors" + "slices" "strings" "testing" @@ -19,7 +20,6 @@ import ( ctrl "sigs.k8s.io/controller-runtime" "sigs.k8s.io/controller-runtime/pkg/client" - "github.com/simplyblock/atlas/statemachine" simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" ) @@ -73,34 +73,23 @@ func TestEveryActionsGraphIsALine(t *testing.T) { } } -// Every abortable step is one some graph declares. The table sits beside the -// graphs rather than in them, so a step renamed in one and not the other would -// make an abort silently unreachable. -func TestEveryAbortableStepIsAStepSomeGraphDeclares(t *testing.T) { - declared := map[string]bool{} - for _, state := range statemachine.DeclaredMultiStates(opsGraphs()) { - declared[state] = true - } - for step := range abortableSteps { - if !declared[string(step)] { - t.Errorf("%s is abortable and no graph declares it", step) - } - } -} - // An abort is honored only before anything has been changed. A step that has // rolled a Deployment or written an image onto the entity carries on, because // stopping there would leave a rollout half-done with nothing driving it either // way. +// +// Both halves are written out rather than derived. The graphs are the only place +// this is declared now, so a step quietly gaining or losing it would otherwise +// change what an abort does with nothing disagreeing. func TestAnAbortIsRefusedOnceARolloutHasStarted(t *testing.T) { - refused := []opsStep{stepRestarting, stepApplying, stepAwaiting, stepVerifying} - for _, step := range refused { - if abortable(step) { + unabortable := UnabortableSteps() + for _, step := range []opsStep{stepRestarting, stepApplying, stepAwaiting, stepVerifying} { + if !slices.Contains(unabortable, step) { t.Errorf("%s is abortable, and an abort there leaves a rollout half-done", step) } } for _, step := range []opsStep{stepDraining, stepPreflight, stepRequesting} { - if !abortable(step) { + if slices.Contains(unabortable, step) { t.Errorf("%s is not abortable, and nothing has been changed at that point", step) } } diff --git a/operator/internal/controllers/controlplane/graphs.go b/operator/internal/controllers/controlplane/graphs.go index 741879a00..894c2669b 100644 --- a/operator/internal/controllers/controlplane/graphs.go +++ b/operator/internal/controllers/controlplane/graphs.go @@ -160,8 +160,9 @@ func opsGraphs() statemachine.MultiConfig[opsStep] { Initial: stepDraining, States: map[opsStep]statemachine.StateDef[opsStep]{ stepDraining: { - To: []opsStep{stepRestarting}, - OnEnter: deadline[opsStep](drainingDeadline), + To: []opsStep{stepRestarting}, + Abortable: true, + OnEnter: deadline[opsStep](drainingDeadline), }, stepRestarting: { To: []opsStep{stepAwaiting}, @@ -175,12 +176,14 @@ func opsGraphs() statemachine.MultiConfig[opsStep] { Initial: stepPreflight, States: map[opsStep]statemachine.StateDef[opsStep]{ stepPreflight: { - To: []opsStep{stepDraining}, - OnEnter: deadline[opsStep](preflightDeadline), + To: []opsStep{stepDraining}, + Abortable: true, + OnEnter: deadline[opsStep](preflightDeadline), }, stepDraining: { - To: []opsStep{stepApplying}, - OnEnter: deadline[opsStep](drainingDeadline), + To: []opsStep{stepApplying}, + Abortable: true, + OnEnter: deadline[opsStep](drainingDeadline), }, stepApplying: { To: []opsStep{stepAwaiting}, @@ -198,8 +201,9 @@ func opsGraphs() statemachine.MultiConfig[opsStep] { Initial: stepRequesting, States: map[opsStep]statemachine.StateDef[opsStep]{ stepRequesting: { - To: []opsStep{stepAwaiting}, - OnEnter: deadline[opsStep](requestingDeadline), + To: []opsStep{stepAwaiting}, + Abortable: true, + OnEnter: deadline[opsStep](requestingDeadline), }, stepAwaiting: {OnEnter: deadline[opsStep](awaitingOpsDeadline)}, }, @@ -218,25 +222,25 @@ var opsInitialDeadlines = map[statemachine.Action]time.Duration{ action(simplyblockv1alpha2.ControlPlaneOpsActionBackup): requestingDeadline, } -// abortableSteps are the steps from which an abort stops the operation cleanly. -// -// The line is whether anything has been changed yet. Draining and Preflight have -// performed no side effect at all, and Requesting has not yet created the +// Which steps an abort stops cleanly is declared on the states above, and the +// line it draws is whether anything has been changed yet. Draining and Preflight +// have performed no side effect at all, and Requesting has not yet created the // backup. Everything past those has rolled a Deployment or written an image onto // the entity, and the operation is what drives that rollout to completion. // -// It is a table beside the graph rather than an edge in it: the phase already -// carries what a terminal Aborted step would say. A test asserts every step here -// is one some graph declares, so the two cannot drift. -var abortableSteps = map[opsStep]bool{ - stepDraining: true, - stepPreflight: true, - stepRequesting: true, -} +// It is a property of the state rather than an edge to a terminal one: the phase +// already carries what a terminal Aborted step would say. -// abortable reports whether an abort asked for while the operation sits on this -// step can be honored. -func abortable(current opsStep) bool { return abortableSteps[current] } +// UnabortableSteps are the steps no action can be stopped from, sorted. +// +// It is exported for the DELETE guard on this kind, which asks the same question +// this package's unwind asks: a deletion may not express something spec.abort +// could not, so both channels read one graph (design-crd-model.md §3.1). The +// guard has a step out of a status and no machine, which is the whole reason +// this reads the graphs rather than the machine the reconciler holds. +func UnabortableSteps() []opsStep { + return statemachine.UnabortableMultiStates(opsGraphs()) +} // action converts the API's action enum into the MultiConfig's key. The // conversion exists because statemachine.Action is a concrete string type rather diff --git a/operator/internal/controllers/node/graphs.go b/operator/internal/controllers/node/graphs.go index 3b7908d16..a3c54ce85 100644 --- a/operator/internal/controllers/node/graphs.go +++ b/operator/internal/controllers/node/graphs.go @@ -133,8 +133,12 @@ func graphs() statemachine.MultiConfig[step] { return statemachine.Config[step]{ Initial: stepRequesting, States: map[step]statemachine.StateDef[step]{ - stepRequesting: {To: []step{stepAwaiting}, OnEnter: deadline[step](requestingDeadline)}, - stepAwaiting: {OnEnter: deadline[step](awaitingDeadline)}, + stepRequesting: { + To: []step{stepAwaiting}, + Abortable: true, + OnEnter: deadline[step](requestingDeadline), + }, + stepAwaiting: {OnEnter: deadline[step](awaitingDeadline)}, }, } } @@ -152,21 +156,32 @@ func graphs() statemachine.MultiConfig[step] { action(simplyblockv1alpha2.StorageNodeOpsActionRemove): { Initial: stepValidating, States: map[step]statemachine.StateDef[step]{ + // Validating performs no side effect at all, which is what makes + // an abort there an Aborted directly rather than an unwind. The + // three steps past the suspend are abortable because their + // unwind exists: the resume the graph already performs on every + // other terminal outcome from Suspending onward (§8.3). + // Removing is not, because the node is being taken out of the + // cluster and there is no resume that puts it back. stepValidating: { - To: []step{stepSuspending}, - OnEnter: deadline[step](validatingDeadline), + To: []step{stepSuspending}, + Abortable: true, + OnEnter: deadline[step](validatingDeadline), }, stepSuspending: { - To: []step{stepMigratingVolumes}, - OnEnter: deadline[step](suspendingDeadline), + To: []step{stepMigratingVolumes}, + Abortable: true, + OnEnter: deadline[step](suspendingDeadline), }, stepMigratingVolumes: { - To: []step{stepVerifying}, - OnEnter: deadline[step](migratingDeadline), + To: []step{stepVerifying}, + Abortable: true, + OnEnter: deadline[step](migratingDeadline), }, stepVerifying: { - To: []step{stepRemoving}, - OnEnter: deadline[step](verifyingDeadline), + To: []step{stepRemoving}, + Abortable: true, + OnEnter: deadline[step](verifyingDeadline), }, stepRemoving: {OnEnter: deadline[step](removingDeadline)}, }, @@ -182,9 +197,15 @@ func graphs() statemachine.MultiConfig[step] { action(simplyblockv1alpha2.StorageNodeOpsActionMigrate): { Initial: stepPreparing, States: map[step]statemachine.StateDef[step]{ + // Preparing has labeled a target host and nothing more. + // Everything after it is refused: the node is mid-restart with + // this operation the only thing watching it back, and the + // promote has re-homed the logical volumes, so there is nothing + // to unwind and the operation is what finishes the relocation. stepPreparing: { - To: []step{stepRelocating}, - OnEnter: deadline[step](preparingDeadline), + To: []step{stepRelocating}, + Abortable: true, + OnEnter: deadline[step](preparingDeadline), }, stepRelocating: { To: []step{stepAwaitingNode}, @@ -201,9 +222,14 @@ func graphs() statemachine.MultiConfig[step] { action(simplyblockv1alpha2.StorageNodeOpsActionHostMaintenance): { Initial: stepHolding, States: map[step]statemachine.StateDef[step]{ + // Holding is the window before the node is taken down, and the + // last point at which calling the maintenance off costs + // nothing. From ShuttingDown onward the node is being taken + // down for a reboot nothing else will bring it back from. stepHolding: { - To: []step{stepShuttingDown}, - OnEnter: deadline[step](holdingDeadline), + To: []step{stepShuttingDown}, + Abortable: true, + OnEnter: deadline[step](holdingDeadline), }, stepShuttingDown: { To: []step{stepReleasing}, @@ -299,60 +325,27 @@ var stepBudgets = map[step]time.Duration{ stepCleanup: cleanupDeadline, } -// abortableSteps are the steps from which an abort stops the operation cleanly. -// -// The line is whether anything is currently down or half-done. Requesting has -// issued nothing. Validating performs no side effect at all and is the step -// before a drain touches the node, which is why an abort there is an Aborted -// directly rather than an unwind (§8.3). Preparing has labeled a target host and -// nothing more. The three drain steps past the suspend are abortable because -// their unwind exists: the resume the graph already performs on every other -// terminal outcome from Suspending onward. -// -// Promoting is the clearest refusal. The promote has activated the target host's -// devices, failed and migrated the origin's, started a rebalance, and re-homed -// the logical volumes, so there is nothing to unwind and the operation is what -// finishes the relocation (§9). Relocating and AwaitingNode are the same rule: -// the node is mid-restart and this operation is the only thing watching it back. -// HostMaintenance past Holding is the third: the node is being taken down for a -// reboot nothing else will bring it back from. +// Which steps an abort stops cleanly is declared on the states above, because +// the line it draws is a property of the step rather than of this kind: whether +// anything is currently down or half-done. Promoting is the clearest refusal. +// The promote has activated the target host's devices, failed and migrated the +// origin's, started a rebalance, and re-homed the logical volumes, so there is +// nothing to unwind and the operation is what finishes the relocation (§9). // -// It is a table beside the graph rather than an edge in it, because a terminal -// Aborted step would be an eighteenth value in the API and the phase already -// carries that meaning. A test asserts every step here is one some graph -// declares, and another asserts that no step between a node's shutdown and its -// restart appears, so the two cannot drift. -var abortableSteps = map[step]bool{ - stepRequesting: true, - stepValidating: true, - stepSuspending: true, - stepMigratingVolumes: true, - stepVerifying: true, - stepPreparing: true, - stepHolding: true, -} - -// abortable reports whether an abort asked for while the operation sits on this -// step can be honored. -func abortable(current step) bool { return abortableSteps[current] } +// It is a property of the state rather than an edge to a terminal one, because a +// terminal Aborted step would be an eighteenth value in the API and the phase +// already carries that meaning. // UnabortableSteps are the declared steps an abort cannot be honored from, // sorted. // // It is exported for the DELETE guard on this kind, which asks the same question // this package's unwind asks: a deletion may not express something spec.abort -// could not, so both channels read one graph -// (design-crd-model.md §3.1). Reading it rather than restating it is what keeps -// the guard from refusing a step this package has since made abortable, or -// admitting one it has not. +// could not, so both channels read one graph (design-crd-model.md §3.1). The +// guard has a step out of a status and no machine, which is the whole reason +// this reads the graphs rather than the machine the reconciler holds. func UnabortableSteps() []step { - var refused []step - for _, declared := range statemachine.DeclaredMultiStates(graphs()) { - if s := step(declared); !abortable(s) { - refused = append(refused, s) - } - } - return refused + return statemachine.UnabortableMultiStates(graphs()) } // unwinds reports whether an abort or a failure from this step owes the node a diff --git a/operator/internal/controllers/node/graphs_test.go b/operator/internal/controllers/node/graphs_test.go index 640879085..06f4da17e 100644 --- a/operator/internal/controllers/node/graphs_test.go +++ b/operator/internal/controllers/node/graphs_test.go @@ -131,22 +131,11 @@ func TestEveryActionDeclaresAGraph(t *testing.T) { } } -// An abortable step that no graph declares is a table that has drifted from the -// graphs beside it, and the consequence is an abort that is silently never -// honored. -func TestEveryAbortableStepIsDeclared(t *testing.T) { - declared := statemachine.DeclaredMultiStates(graphs()) - for state := range abortableSteps { - if !slices.Contains(declared, string(state)) { - t.Errorf("step %q is marked abortable and no graph declares it", state) - } - } -} - -// The line the abort table draws is whether anything is currently down or -// half-done. These four are the sharpest cases and each would leave the node in a -// state nothing else drives it out of. +// The line abortability draws is whether anything is currently down or +// half-done. These seven are the sharpest cases and each would leave the node in +// a state nothing else drives it out of. func TestNoStepPastThePointOfNoReturnIsAbortable(t *testing.T) { + unabortable := statemachine.UnabortableMultiStates(graphs()) for _, state := range []step{ // The promote has re-homed the logical volumes: there is nothing to // unwind, and the operation is what finishes the relocation. @@ -161,12 +150,40 @@ func TestNoStepPastThePointOfNoReturnIsAbortable(t *testing.T) { stepAwaitingHost, stepRestarting, } { - if abortable(state) { + if !slices.Contains(unabortable, state) { t.Errorf("step %q is abortable and the node is not in a state an abort can leave it in", state) } } } +// The other half of the same statement, written out rather than derived. The +// graphs are now the only place abortability is declared, so a step that quietly +// gains or loses it would otherwise change what an abort does with nothing +// disagreeing. +func TestTheStepsAnAbortStopsCleanly(t *testing.T) { + unabortable := statemachine.UnabortableMultiStates(graphs()) + for _, state := range []step{ + // Nothing has been issued yet. + stepRequesting, + // No side effect at all, which is why an abort here is an Aborted + // directly rather than an unwind (§8.3). + stepValidating, + // Past the suspend, and the unwind is the resume the graph already + // performs on every other terminal outcome from here on. + stepSuspending, + stepMigratingVolumes, + stepVerifying, + // A target host has been labeled and nothing more. + stepPreparing, + // The window before the node is taken down for maintenance. + stepHolding, + } { + if slices.Contains(unabortable, state) { + t.Errorf("step %q refuses an abort, and nothing it has done needs finishing", state) + } + } +} + // Every terminal outcome from Suspending onward owes the node a resume, because a // node past the suspend is not serving and an operation that stopped there would // take capacity out of the cluster for as long as nobody noticed (§8.3). diff --git a/operator/internal/controllers/node/storagenodeops_controller.go b/operator/internal/controllers/node/storagenodeops_controller.go index d8a7eebc0..8d4cc75ab 100644 --- a/operator/internal/controllers/node/storagenodeops_controller.go +++ b/operator/internal/controllers/node/storagenodeops_controller.go @@ -265,7 +265,7 @@ func (r *StorageNodeOpsReconciler) advance( current := machine.CurrentState() if ops.Spec.Abort { - return r.unwind(ctx, ops, current) + return r.unwind(ctx, ops, machine, current) } // The cluster gate is not the same as the lock. A node operation runs inside @@ -395,9 +395,15 @@ func (r *StorageNodeOpsReconciler) nextStep( // so there is nothing to unwind and the operation is what finishes the relocation // (§9). func (r *StorageNodeOpsReconciler) unwind( - ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, current step, + ctx context.Context, + ops *simplyblockv1alpha2.StorageNodeOps, + machine *statemachine.Machine[step], + current step, ) (ctrl.Result, error) { - if !abortable(current) { + // The machine is asked rather than a table beside it, and it is asked rather + // than the graphs, because it was built for this operation's action: a step + // two actions share can be abortable in one of them. + if !machine.CanAbort() { // Not a failure of the operation: it carries on. What the user asked for // cannot be done, and saying so is the whole of the response. return ctrl.Result{RequeueAfter: opsRetry}, r.note(ctx, ops, fmt.Sprintf( diff --git a/operator/internal/webhook/clusterdeploymentconfig_validator.go b/operator/internal/webhook/clusterdeploymentconfig_validator.go index 23023c8b4..65ec95b61 100644 --- a/operator/internal/webhook/clusterdeploymentconfig_validator.go +++ b/operator/internal/webhook/clusterdeploymentconfig_validator.go @@ -330,4 +330,4 @@ func (v *ClusterDeploymentConfigValidator) otherOwner( } } return "", nil -} \ No newline at end of file +} diff --git a/operator/internal/webhook/clusterdeploymentconfig_validator_test.go b/operator/internal/webhook/clusterdeploymentconfig_validator_test.go index d35e85e8a..e6720de9f 100644 --- a/operator/internal/webhook/clusterdeploymentconfig_validator_test.go +++ b/operator/internal/webhook/clusterdeploymentconfig_validator_test.go @@ -39,15 +39,25 @@ func testWorker(name string) *corev1.Node { return &corev1.Node{ObjectMeta: metav1.ObjectMeta{Name: name}} } +// testDeploymentStorageCluster is the cluster every test in this file names, +// because the checks under test are about whether a cluster of that name exists +// and what class it is, never about which name it carries. func testDeploymentStorageCluster( - name string, class simplyblockv1alpha2.StorageClusterDeviceClass, + class simplyblockv1alpha2.StorageClusterDeviceClass, ) *simplyblockv1alpha2.StorageCluster { return &simplyblockv1alpha2.StorageCluster{ - ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: testDeploymentNamespace}, - Spec: simplyblockv1alpha2.StorageClusterSpec{DeviceClass: class}, + ObjectMeta: metav1.ObjectMeta{ + Name: testDeploymentCluster, + Namespace: testDeploymentNamespace, + }, + Spec: simplyblockv1alpha2.StorageClusterSpec{DeviceClass: class}, } } +// testSecondCluster is the cluster a second document names, for the checks about +// two documents rather than about one document and the world. +const testSecondCluster = "rack-two" + // testConfig is a document that passes every check: one group, one worker that // exists, NVMe devices, and a cluster of its own to create. func testConfig() *simplyblockv1alpha2.ClusterDeploymentConfig { @@ -225,7 +235,7 @@ func TestApprovingIsRefusedWhenClusterRefResolvesToNothing(t *testing.T) { func TestApprovingIsRefusedWhenTheClusterAlreadyExists(t *testing.T) { existing := []client.Object{ testWorker("worker-1"), - testDeploymentStorageCluster(testDeploymentCluster, ""), + testDeploymentStorageCluster(""), } mustDeny(t, approve(t, existing, testConfig()), @@ -251,7 +261,7 @@ func TestApprovingAGrowthDocumentOfTheWrongDeviceClassIsRefused(t *testing.T) { existing := []client.Object{ testWorker("worker-1"), testDeploymentStorageCluster( - testDeploymentCluster, simplyblockv1alpha2.StorageClusterDeviceClassNVMe), + simplyblockv1alpha2.StorageClusterDeviceClassNVMe), } mustDeny(t, approve(t, existing, withBlockDevices(grows(testConfig(), testDeploymentCluster))), @@ -262,7 +272,7 @@ func TestApprovingAGrowthDocumentOfTheRightDeviceClassIsAdmitted(t *testing.T) { existing := []client.Object{ testWorker("worker-1"), testDeploymentStorageCluster( - testDeploymentCluster, simplyblockv1alpha2.StorageClusterDeviceClassLogicalBlock), + simplyblockv1alpha2.StorageClusterDeviceClassLogicalBlock), } mustAllow(t, approve(t, existing, @@ -274,7 +284,7 @@ func TestApprovingAGrowthDocumentOfTheRightDeviceClassIsAdmitted(t *testing.T) { func TestAGrowthDocumentReadsAnUnstatedClassAsNVMe(t *testing.T) { existing := []client.Object{ testWorker("worker-1"), - testDeploymentStorageCluster(testDeploymentCluster, ""), + testDeploymentStorageCluster(""), } mustAllow(t, approve(t, existing, grows(testConfig(), testDeploymentCluster))) @@ -287,17 +297,17 @@ func TestAGrowthDocumentReadsAnUnstatedClassAsNVMe(t *testing.T) { // document reporting a deployment that never happened. func TestApprovingASecondCreateOfTheSameClusterIsRefused(t *testing.T) { other := approved(testConfig()) - other.Name = "rack-two" + other.Name = testSecondCluster mustDeny(t, approve(t, []client.Object{testWorker("worker-1"), other}, testConfig()), - "rack-two") + testSecondCluster) } // An unapproved document naming the same cluster is not an owner. It is a draft, // and whichever of the two is approved first becomes the owner. func TestADraftNamingTheSameClusterDoesNotBlockAnApproval(t *testing.T) { other := testConfig() - other.Name = "rack-two" + other.Name = testSecondCluster mustAllow(t, approve(t, []client.Object{testWorker("worker-1"), other}, testConfig())) } @@ -306,11 +316,11 @@ func TestADraftNamingTheSameClusterDoesNotBlockAnApproval(t *testing.T) { // cluster is the design rather than a conflict (§6). func TestASecondGrowthDocumentIsAdmitted(t *testing.T) { other := approved(grows(testConfig(), testDeploymentCluster)) - other.Name = "rack-two" + other.Name = testSecondCluster existing := []client.Object{ testWorker("worker-1"), - testDeploymentStorageCluster(testDeploymentCluster, ""), + testDeploymentStorageCluster(""), other, } @@ -377,4 +387,4 @@ func TestTheRequestNamespaceIsUsedWhenTheObjectCarriesNone(t *testing.T) { config.Namespace = "" mustDeny(t, review(t, nil, admissionv1.Create, nil, config), "worker-1") -} \ No newline at end of file +} diff --git a/operator/internal/webhook/controlplaneops_validator_test.go b/operator/internal/webhook/controlplaneops_validator_test.go index 2a3151589..b73d06463 100644 --- a/operator/internal/webhook/controlplaneops_validator_test.go +++ b/operator/internal/webhook/controlplaneops_validator_test.go @@ -10,6 +10,7 @@ package webhook import ( "context" "encoding/json" + "slices" "strings" "testing" @@ -22,6 +23,7 @@ import ( "github.com/simplyblock/atlas/statemachine" simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/controlplane" ) func cpOpsScheme(t *testing.T) *runtime.Scheme { @@ -145,3 +147,20 @@ func TestEveryUndeletableStepStatesWhatItIsDoing(t *testing.T) { } } } + +// The refusal table and the graph are two statements of one rule, and this holds +// them equal in both directions. A deletion may not express something spec.abort +// could not, so an extra entry refuses a delete the abort channel would have +// honored, and a missing one admits the withdrawal of a record nothing else +// accounts for. +func TestTheControlPlaneRefusalTableAndTheGraphAgree(t *testing.T) { + refused := make([]simplyblockv1alpha2.ControlPlaneOpsStep, 0, len(undeletableControlPlaneSteps)) + for step := range undeletableControlPlaneSteps { + refused = append(refused, step) + } + slices.Sort(refused) + + if want := controlplane.UnabortableSteps(); !slices.Equal(refused, want) { + t.Errorf("the guard refuses %v, and the graph declares no abort edge from %v", refused, want) + } +} diff --git a/operator/internal/webhook/storagebackupops_validator.go b/operator/internal/webhook/storagebackupops_validator.go index 045e86ae5..00bee0e73 100644 --- a/operator/internal/webhook/storagebackupops_validator.go +++ b/operator/internal/webhook/storagebackupops_validator.go @@ -36,15 +36,16 @@ import ( // undeletableSteps are the steps from which a running operation's record may not // be withdrawn. // -// Both have a logical volume the control plane created behind them and a claim -// the operation is on its way to binding, so removing the record leaves both with -// nothing accounting for them. It is the same set the controller refuses an abort -// from, read from the same graph: the two channels ask one question, and the one -// that carries a reason forward is spec.abort, because it leaves an Aborted -// object to read where a delete leaves nothing. +// This is the one guard in the group that refuses more than the graph does, and +// Restoring below is the step it adds. Everywhere else the two channels agree +// exactly and a test holds them equal; here the delete is strictly the more +// dangerous of the two, because an abort leaves an Aborted object behind for the +// controller to clean up from and a delete leaves nothing at all. Binding an +// agreement test to this set would therefore force the weaker answer onto the +// stronger channel. // -// The webhook is the stronger of the two guards, because it also catches the -// `--force --grace-period=0` that a finalizer alone does not. +// The webhook is also the stronger of the two guards in the other sense, because +// it catches the `--force --grace-period=0` that a finalizer alone does not. var undeletableSteps = map[simplyblockv1alpha2.StorageBackupOpsStep]string{ // Restoring is here although the design's §6 names only the two below, and // the reason is a window the design does not model: the step asks the control diff --git a/operator/internal/webhook/storageclusterops_validator.go b/operator/internal/webhook/storageclusterops_validator.go index 975e7f55c..b23af3c21 100644 --- a/operator/internal/webhook/storageclusterops_validator.go +++ b/operator/internal/webhook/storageclusterops_validator.go @@ -119,4 +119,4 @@ func (v *StorageClusterOpsValidator) admitDelete(req admission.Request) admissio "locked by an object that no longer exists. Set spec.abort to stop an operation that "+ "can still be stopped, and delete the record once it is terminal.", ops.Namespace, ops.Name, step, doing, ops.Spec.ClusterRef)) -} \ No newline at end of file +} diff --git a/operator/internal/webhook/storageclusterops_validator_test.go b/operator/internal/webhook/storageclusterops_validator_test.go index c601c2a0b..9c27364fa 100644 --- a/operator/internal/webhook/storageclusterops_validator_test.go +++ b/operator/internal/webhook/storageclusterops_validator_test.go @@ -187,4 +187,4 @@ func TestEveryUndeletableClusterStepStatesWhatItIsDoing(t *testing.T) { t.Errorf("%s is undeletable and the refusal says nothing about why", step) } } -} \ No newline at end of file +} diff --git a/operator/internal/webhook/storagenodeops_validator.go b/operator/internal/webhook/storagenodeops_validator.go index a4bb01dbf..c497e5c9b 100644 --- a/operator/internal/webhook/storagenodeops_validator.go +++ b/operator/internal/webhook/storagenodeops_validator.go @@ -122,4 +122,4 @@ func (v *StorageNodeOpsValidator) admitDelete(req admission.Request) admission.R "an object that no longer exists. Set spec.abort to stop an operation that can still "+ "be stopped, and delete the record once it is terminal.", ops.Namespace, ops.Name, step, doing, ops.Spec.NodeRef)) -} \ No newline at end of file +} diff --git a/operator/internal/webhook/storagenodeops_validator_test.go b/operator/internal/webhook/storagenodeops_validator_test.go index e42385394..389134341 100644 --- a/operator/internal/webhook/storagenodeops_validator_test.go +++ b/operator/internal/webhook/storagenodeops_validator_test.go @@ -202,4 +202,4 @@ func TestEveryUndeletableNodeStepStatesWhatItIsDoing(t *testing.T) { t.Errorf("%s is undeletable and the refusal says nothing about why", step) } } -} \ No newline at end of file +} From a3ea857210b1d7a5722bb33f519a88ee9ff02c8f Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 10:03:56 +0200 Subject: [PATCH 029/206] fix(operator): the installer carries the three new validating webhooks dist/install.yaml is a generated artifact that is committed, and the three commits that added the ClusterDeploymentConfig, StorageClusterOps, and StorageNodeOps guards regenerated config/webhook/manifests.yaml and the chart without it. An install from dist/install.yaml would have created none of the three webhooks, and the drift check in operator_manifests.yaml fails on it. Produced by the root `make build`, which is the target that covers every generated artifact at once. Co-Authored-By: Claude Fable 5 --- operator/dist/install.yaml | 58 ++++++++++++++++++++++++++++++++++++++ 1 file changed, 58 insertions(+) diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index 0b7e58cd5..a0445363d 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -10462,6 +10462,26 @@ webhooks: resources: - persistentvolumeclaims sideEffects: None +- admissionReviewVersions: + - v1 + clientConfig: + service: + name: simplyblock-operator-webhook-service + namespace: simplyblock-operator-system + path: /validate-storage-simplyblock-io-v1alpha2-clusterdeploymentconfig + failurePolicy: Fail + name: vclusterdeploymentconfig.simplyblock.io + rules: + - apiGroups: + - storage.simplyblock.io + apiVersions: + - v1alpha2 + operations: + - CREATE + - UPDATE + resources: + - clusterdeploymentconfigs + sideEffects: None - admissionReviewVersions: - v1 clientConfig: @@ -10579,6 +10599,25 @@ webhooks: resources: - storagebackuppolicies sideEffects: None +- admissionReviewVersions: + - v1 + clientConfig: + service: + name: simplyblock-operator-webhook-service + namespace: simplyblock-operator-system + path: /validate-storage-simplyblock-io-v1alpha2-storageclusterops + failurePolicy: Fail + name: vstorageclusterops.simplyblock.io + rules: + - apiGroups: + - storage.simplyblock.io + apiVersions: + - v1alpha2 + operations: + - DELETE + resources: + - storageclusterops + sideEffects: None - admissionReviewVersions: - v1 clientConfig: @@ -10619,6 +10658,25 @@ webhooks: resources: - storagenodes sideEffects: None +- admissionReviewVersions: + - v1 + clientConfig: + service: + name: simplyblock-operator-webhook-service + namespace: simplyblock-operator-system + path: /validate-storage-simplyblock-io-v1alpha2-storagenodeops + failurePolicy: Fail + name: vstoragenodeops.simplyblock.io + rules: + - apiGroups: + - storage.simplyblock.io + apiVersions: + - v1alpha2 + operations: + - DELETE + resources: + - storagenodeops + sideEffects: None - admissionReviewVersions: - v1 clientConfig: From 01b64c82081259c7c362d9852732a42b301efc87 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 10:04:21 +0200 Subject: [PATCH 030/206] chore(fleet): regenerate the CRDs that embed the operator's API MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The fleet kinds embed the operator's ClusterDeploymentConfig spec, so an edit to that API invalidates them too. Several landed without a regeneration here — spec.approved gaining a default and its explanation, and enableDriveFormat among them (2895275b) — so the schema the fleet CRDs describe has been behind for a while. No fleet workflow checks for the drift, which is why nothing said so. Pure generator output, from the root `make build`. Co-Authored-By: Claude Fable 5 --- ...eet.simplyblock.io_clusterdeployments.yaml | 72 ++++++++++-- .../fleet.simplyblock.io_fleetoperations.yaml | 105 +++++++++++++----- 2 files changed, 144 insertions(+), 33 deletions(-) diff --git a/fleet/config/crd/bases/fleet.simplyblock.io_clusterdeployments.yaml b/fleet/config/crd/bases/fleet.simplyblock.io_clusterdeployments.yaml index 99cef0cf2..b23954430 100644 --- a/fleet/config/crd/bases/fleet.simplyblock.io_clusterdeployments.yaml +++ b/fleet/config/crd/bases/fleet.simplyblock.io_clusterdeployments.yaml @@ -80,15 +80,41 @@ spec: opaque blob because it is the thing a reviewer reads. properties: approved: + default: false description: |- Approved is the review gate. A document is expanded only once it is set, and is validated but otherwise inert before that, which is what makes reviewing a wrong document safe. + + It is defaulted and serialized rather than omitted when false, and the + two are the same requirement read twice. A reviewer has to see the gate + they are being asked to open, and the rules above have to find the field + they read: a bool omitted when false is a key the apiserver never stores, + so a rule reading it fails rather than reading false, and the first rule + guarding approval denied every approval there could ever be. The has() + guards are what carry documents written before the default existed. type: boolean cluster: description: Cluster is the StorageCluster to create. Ignored when ClusterRef is set. properties: + enableDriveFormat: + description: |- + EnableDriveFormat formats every device the document names before a storage + node takes it, which is how a drive carrying anything already is made + usable. + + It says what is wanted rather than how, because the how differs by device + class: an NVMe device is formatted to a 4K block size, and a logical block + device has its signatures wiped. One field covers both, so a document does + not have to know which class the expansion will resolve it to. + + It is on the document rather than defaulted further down because it is + destructive and the document is what somebody approves. A reviewer reading + a draft has to see that the drives it lists will be formatted, and be able + to strike it before approving; the cluster's own field is immutable once + the cluster exists, so a default nobody saw could not be undone either. + type: boolean enableFailureDomains: description: |- EnableFailureDomains opts the cluster into failure-domain mode, in which @@ -119,18 +145,45 @@ spec: description: Name is the StorageCluster's name. maxLength: 253 type: string + nodesPerSocket: + description: |- + NodesPerSocket is how many storage nodes run per NUMA socket. See + SocketsToUse, which it multiplies. + format: int32 + maximum: 8 + minimum: 1 + type: integer + socketsToUse: + description: |- + SocketsToUse restricts the deployment to selected NUMA sockets, and empty + means socket 0 alone. With NodesPerSocket it decides how many storage nodes + each worker runs, so a group of two workers on a two-socket layout expands + to four nodes. + + It is here rather than on a node set because it is immutable on the cluster + it lands on: the layout a fleet was built with is not one a later document + can vary, and a reviewer should see it before the cluster exists. + items: + maxLength: 16 + type: string + maxItems: 16 + type: array + x-kubernetes-list-type: set stripe: description: Stripe is the erasure-coding layout. properties: dataChunks: - description: DataChunks defines the number of data chunks - in the erasure-coding layout. + description: DataChunks is the number of data chunks per + stripe (ndcs). format: int32 + minimum: 1 type: integer parityChunks: - description: ParityChunks defines the number of parity - chunks in the erasure-coding layout. + description: |- + ParityChunks is the number of parity chunks per stripe (npcs), and + therefore how many chunk losses a stripe survives. format: int32 + minimum: 0 type: integer type: object vcpuCount: @@ -259,11 +312,14 @@ spec: description: Count is the number of journal managers to configure. format: int32 + minimum: 1 type: integer percentPerDevice: - description: PercentPerDevice is the journal manager - capacity percentage per device. + description: PercentPerDevice is the share of + each device given to the journal. format: int32 + maximum: 100 + minimum: 1 type: integer type: object mgmtInterface: @@ -320,9 +376,9 @@ spec: type: object x-kubernetes-validations: - message: an approved deployment config is immutable - rule: '!oldSelf.approved || self == oldSelf' + rule: '!has(oldSelf.approved) || !oldSelf.approved || self == oldSelf' - message: approval cannot be withdrawn - rule: '!oldSelf.approved || self.approved' + rule: '!has(oldSelf.approved) || !oldSelf.approved || self.approved' - message: 'every group must name the same device class: all nvme or all block' rule: self.nodeSets.all(s, s.groups.all(g, !has(g.devices) || !has(g.devices.block))) diff --git a/fleet/config/crd/bases/fleet.simplyblock.io_fleetoperations.yaml b/fleet/config/crd/bases/fleet.simplyblock.io_fleetoperations.yaml index 534863802..fa7d4407b 100644 --- a/fleet/config/crd/bases/fleet.simplyblock.io_fleetoperations.yaml +++ b/fleet/config/crd/bases/fleet.simplyblock.io_fleetoperations.yaml @@ -200,6 +200,24 @@ spec: block devices and require enableLogicalBlockDevices rule: (has(self.enableLogicalBlockDevices) && self.enableLogicalBlockDevices) || !(has(self.blockAllowList) || has(self.blockDenyList)) + enableControlPlaneNodes: + description: |- + EnableControlPlaneNodes lets the run consider machines that run the API + server and etcd. + + It is off by default because a storage node is a data path, and putting one + on an etcd host is a placement almost nobody intends. The approval gate is a + poor place to catch it: a fifty-worker draft is not a document anybody reads + closely enough to spot three control-plane nodes in it. A combined three-node + or single-node deployment is the case that wants it, and those are set up + deliberately. + + There is no field beside it for infrastructure nodes, because those are used + without asking: an OpenShift infra node is the tier a cluster's own + infrastructure runs on, and simplyblock storage is infrastructure. A fleet + with disks in its infra nodes meant those disks to be the storage, so a draft + proposes them ahead of the workers rather than leaving them out. + type: boolean nodeSelector: additionalProperties: type: string @@ -313,8 +331,14 @@ spec: storageClusterOps: description: StorageClusterOps runs a cluster-level operation. properties: + abort: + description: |- + Abort asks a running operation to stop at its next step and unwind. + Whether an abort is expressible from the current step is declared by that + action's graph rather than checked here. + type: boolean action: - description: Action is the operation to perform. Immutable. + description: Action is the operation to perform. enum: - Activate - Expand @@ -322,27 +346,46 @@ spec: - Start - Restart - RollingRestart + - CancelTask type: string x-kubernetes-validations: - message: field is immutable rule: self == oldSelf + cancelTask: + description: CancelTask parameterizes action CancelTask and + is ignored by the others. + properties: + taskID: + description: |- + TaskID names the entry of StorageCluster.status.tasks to cancel, by the + control plane's identifier for it rather than by its position in the + list. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + required: + - taskID + type: object clusterRef: - description: ClusterRef is the name of the target StorageCluster. - Immutable. + description: |- + ClusterRef names the StorageCluster this operation acts on. The operation + never owns its target, because deleting the record of an operation must + not delete the cluster it operated on. type: string x-kubernetes-validations: - message: field is immutable rule: self == oldSelf rollingRestart: description: |- - RollingRestart configures behavior specific to the RollingRestart action. - Ignored for all other actions. + RollingRestart parameterizes action RollingRestart and is ignored by the + others. properties: refreshSNodeAPI: description: |- - RefreshSNodeAPI restarts the storage-node DaemonSet pod on each node - after the backend node is shut down and before it is restarted, ensuring - the latest image is running before the node comes back online. + RefreshSNodeAPI restarts each node's storage-node DaemonSet pod between + its shutdown and its restart, so the latest image is running when the + node returns. type: boolean type: object required: @@ -352,8 +395,16 @@ spec: storageNodeOps: description: StorageNodeOps runs a node-level operation. properties: + abort: + description: |- + Abort asks a running operation to stop at its next step and unwind. It is + the only mutable field on this spec, because it is the only thing about an + operation that can legitimately be decided after it started. Whether an + abort is expressible from the current step is declared by that action's + graph rather than checked here. + type: boolean action: - description: Action is the operation to perform. Immutable. + description: Action is the operation to perform. enum: - Shutdown - Restart @@ -361,13 +412,16 @@ spec: - Resume - Remove - Migrate + - HostMaintenance type: string x-kubernetes-validations: - message: field is immutable rule: self == oldSelf force: - description: Force enables forced execution where the backend - supports it. + description: |- + Force passes the control plane's force flag where the action supports it. + Migrate defaults it to true, because the control plane rejects a non-forced + restart of a node that is not already offline. type: boolean migrate: description: Migrate parameterizes action Migrate and is ignored @@ -375,16 +429,15 @@ spec: properties: newSsdPcie: description: |- - NewSsdPcie lists additional NVMe PCIe addresses to bind on the target host - during a migration. Passed through to the control-plane restart as - new_ssd_pcie. + NewSsdPcie lists additional NVMe PCI addresses to bind on the target host, + passed through to the control-plane restart as new_ssd_pcie and merged into + the node's effective allow list so they survive a later rebuild. items: type: string type: array targetWorkerNode: - description: |- - TargetWorkerNode is the Kubernetes worker hostname the storage node is - relocated onto. + description: TargetWorkerNode is the Kubernetes worker + the node is relocated onto. type: string x-kubernetes-validations: - message: field is immutable @@ -393,27 +446,29 @@ spec: - targetWorkerNode type: object nodeRef: - description: NodeRef is the name of the target StorageNode. - Immutable. + description: |- + NodeRef names the StorageNode this operation acts on. The operation never + owns its target, because deleting the record of an operation must not delete + the node it operated on. type: string x-kubernetes-validations: - message: field is immutable rule: self == oldSelf reattachVolume: description: |- - ReattachVolume reattaches volumes during the node restart. - Applicable when action=Restart or action=Migrate. + ReattachVolume asks the control plane to reattach this node's volumes as + part of a restart. Applies to Restart, Migrate, and HostMaintenance. type: boolean remove: description: Remove parameterizes action Remove and is ignored by the others. properties: systemVolumeFilterRegex: + default: ^sb-fio-baseline-.* description: |- - SystemVolumeFilterRegex is a Go regular expression matched against backend - volume names. Matching volumes are treated as system volumes: excluded from - drain migration and deleted inline during the Verifying phase. - Defaults to `^sb-fio-baseline-.*`. + SystemVolumeFilterRegex matches backend volume names that are system + volumes: excluded from the drain's migration and deleted during + verification rather than blocking it. type: string type: object required: From d2c68c7ac3a0e02731a8df488ca19dc02174b8f7 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 10:15:56 +0200 Subject: [PATCH 031/206] feat(chart): the profiles that run a CSI driver render one MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Nothing created a SimplyblockDriver on any path. The chart stopped rendering the CSI plugins when the operator took them over and rendered nothing in their place, the operator bootstraps no object, and the upgrade tool's seeding step is not built, so a fresh install came up with an operator, a control plane, and no way to provision a volume. The chart renders it, named simplyblock beside the ControlPlane of that name, because that name is what every child's name derives from and an upgraded cluster's running objects answer to exactly those names. It carries helm.sh/resource-policy: keep for the two reasons the ControlPlane does: an uninstall must not take the driver out from under attached volumes, and with the operator deleted first nothing would be left to release the finalizer. Which profiles render one is a list rather than a negation. Both of today's profiles run workloads that mount simplyblock volumes, so both are on it, and a profile added later renders no driver until somebody decides it should. The negation would give every future profile a CSI deployment for saying nothing, and a driver registers a provisioner and claims the node plugin's socket on every worker. A driver block in values.yaml carries the settings design-simplyblockdriver.md §8 moved onto the spec, which the chart had no home for since #513 took the templates out: driverName, the plugin image, the controller replicas, both placement pairs, snapshots, service-account auth, and the six sidecars. TLS is derived from the chart's existing tls values rather than duplicated, with the lower-case provider spelling mapped onto the API's. values.schema.json refuses an unknown key and a sidecar image outside the allowed registries at render time, which is where a value silently doing nothing is cheapest to find. §4.3 said the chart does not render the object, which is what left the kind with no producer, and it is corrected. The seeding it describes is still the upgrade tool's: a fresh install is right with the defaults, and an upgraded cluster is not, because driverName is immutable and a default reconfigures a live deployment on the pass that adopts it. check-rendered-objects.sh requires the object for both profiles, and was red for both before this. Co-Authored-By: Claude Fable 5 --- .../templates/_helpers.tpl | 21 + .../templates/simplyblockdriver_cr.yaml | 84 ++++ .../simplyblock-operator/values.schema.json | 364 +++++++++++++++--- .../charts/simplyblock-operator/values.yaml | 68 ++++ helm-charts/scripts/check-rendered-objects.sh | 4 + .../crd-redesign/design-simplyblockdriver.md | 33 +- 6 files changed, 503 insertions(+), 71 deletions(-) create mode 100644 helm-charts/charts/simplyblock-operator/templates/simplyblockdriver_cr.yaml diff --git a/helm-charts/charts/simplyblock-operator/templates/_helpers.tpl b/helm-charts/charts/simplyblock-operator/templates/_helpers.tpl index 8e97055ab..932c87aab 100644 --- a/helm-charts/charts/simplyblock-operator/templates/_helpers.tpl +++ b/helm-charts/charts/simplyblock-operator/templates/_helpers.tpl @@ -24,6 +24,27 @@ true {{- end -}} {{- end -}} +{{/* +Whether this profile runs a CSI driver. + +The profiles are listed rather than a negation of the ones that do not, so a +profile added later renders no driver until somebody decides it should. The +negation would do the opposite and give every future profile a CSI deployment +by default, which is the wrong way round: a driver registers a provisioner and +takes over the node plugin's socket on every worker, and a profile that wanted +neither would get both by saying nothing. + +Both of today's profiles are here because both run workloads that mount +simplyblock volumes. What differs between them is where the control plane is, +and the driver reaches it through the credentials Secret either way +(design-simplyblockdriver.md §4.3). +*/}} +{{- define "simplyblock.rendersCSIDriver" -}} +{{- if has .Values.deployment.profile (list "standalone" "managed") -}} +true +{{- end -}} +{{- end -}} + {{- define "simplyblock.controlPlaneAddr" -}} {{- if .Values.csiConfig.simplybk.ip -}} {{ .Values.csiConfig.simplybk.ip }} diff --git a/helm-charts/charts/simplyblock-operator/templates/simplyblockdriver_cr.yaml b/helm-charts/charts/simplyblock-operator/templates/simplyblockdriver_cr.yaml new file mode 100644 index 000000000..26199703d --- /dev/null +++ b/helm-charts/charts/simplyblock-operator/templates/simplyblockdriver_cr.yaml @@ -0,0 +1,84 @@ +{{- if eq (include "simplyblock.rendersCSIDriver" .) "true" }} +--- +# The CSI driver, as the one object the operator builds it from. +# +# The chart used to render the node DaemonSet, the controller StatefulSet, their +# RBAC, and the CSIDriver registration directly. It renders this instead, and the +# operator applies all of them (design-simplyblockdriver.md §8). Nothing else +# creates it: without this template a fresh install comes up with an operator, a +# control plane, and no way to provision a volume. +# +# Named simplyblock, beside the ControlPlane of that name, because the object's +# name is what every child's name derives from: simplyblock-csi-node, +# simplyblock-csi-controller, and the ten more in §4.3's table. On a cluster +# upgraded from a chart that rendered those objects, that derivation lands on the +# ones already running and the reconcile adopts them in place rather than +# building a second deployment beside them. +# +# Authored at v1alpha2 for the same reason the ControlPlane is: an object written +# at the storage version needs no conversion, so nothing on a fresh install's path +# calls the conversion webhook. +apiVersion: storage.simplyblock.io/v1alpha2 +kind: SimplyblockDriver +metadata: + name: simplyblock + namespace: {{ .Release.Namespace }} + annotations: + # Helm must not delete this object, because deleting it deletes the CSI + # driver, and the volumes attached through it stay attached with nothing left + # to detach them. `helm uninstall` removes the operator and leaves the driver + # running; removing the driver is a deliberate + # `kubectl delete simplyblockdriver` performed while the operator is up, so + # its finalizer can take the plugins down in order (§8). + # + # It also breaks the same deadlock the ControlPlane annotation does: with the + # operator deleted first, nothing is left to release the finalizer, and the + # release sits in `uninstalling` forever. + helm.sh/resource-policy: keep +spec: + # Immutable, and the one field an upgraded cluster cannot repair by editing. + driverName: {{ .Values.driver.driverName | quote }} + imagePullPolicy: {{ .Values.driver.image.pullPolicy | quote }} + controllerReplicas: {{ .Values.driver.controllerReplicas }} + enableVolumeSnapshots: {{ .Values.driver.enableVolumeSnapshots }} + enableServiceAccountAuth: {{ .Values.driver.enableServiceAccountAuth }} + {{- if .Values.driver.image.repository }} + image: "{{ .Values.driver.image.repository }}:{{ required "driver.image.tag is required when driver.image.repository is set: the field is a full reference and an untagged one is refused by the schema" .Values.driver.image.tag }}" + {{- end }} + {{- with .Values.driver.nodeSelector }} + nodeSelector: + {{- toYaml . | nindent 4 }} + {{- end }} + {{- with .Values.driver.tolerations }} + tolerations: + {{- toYaml . | nindent 4 }} + {{- end }} + {{- with .Values.driver.controllerNodeSelector }} + controllerNodeSelector: + {{- toYaml . | nindent 4 }} + {{- end }} + {{- with .Values.driver.controllerTolerations }} + controllerTolerations: + {{- toYaml . | nindent 4 }} + {{- end }} + {{- $sidecars := dict }} + {{- range $field, $image := .Values.driver.sidecarImages }} + {{- if $image }} + {{- $_ := set $sidecars $field $image }} + {{- end }} + {{- end }} + {{- with $sidecars }} + # Only the sidecars somebody pinned. One left empty takes this operator + # release's version rather than being frozen at whatever the chart last + # defaulted to. + sidecarImages: + {{- toYaml . | nindent 4 }} + {{- end }} + tls: + enableTLS: {{ .Values.tls.enabled }} + enableMutualTLS: {{ .Values.tls.mutual_enabled }} + # The chart spells the providers in lower case and this API spells OpenShift + # as the product does, so the two are mapped here rather than being assumed + # to match. + provider: {{ if eq .Values.tls.provider "openshift" }}OpenShift{{ else }}cert-manager{{ end }} +{{- end }} diff --git a/helm-charts/charts/simplyblock-operator/values.schema.json b/helm-charts/charts/simplyblock-operator/values.schema.json index 8e661e52a..13e8d13c6 100644 --- a/helm-charts/charts/simplyblock-operator/values.schema.json +++ b/helm-charts/charts/simplyblock-operator/values.schema.json @@ -6,53 +6,93 @@ "tls": { "type": "object", "additionalProperties": false, - "required": ["enabled"], + "required": [ + "enabled" + ], "properties": { - "enabled": { "type": "boolean" }, - "provider": { "enum": ["openshift", "cert-manager"] }, - "mutual_enabled": { "type": "boolean" }, + "enabled": { + "type": "boolean" + }, + "provider": { + "enum": [ + "openshift", + "cert-manager" + ] + }, + "mutual_enabled": { + "type": "boolean" + }, "cert-manager": { "type": "object", "additionalProperties": false, "properties": { - "createSelfSignedIssuer": { "type": "boolean" }, - "cluster-issuer": { "type": "string" }, - "namespace": { "type": "string" } + "createSelfSignedIssuer": { + "type": "boolean" + }, + "cluster-issuer": { + "type": "string" + }, + "namespace": { + "type": "string" + } } } }, "allOf": [ { "if": { - "properties": { "enabled": { "const": true } }, - "required": ["enabled"] + "properties": { + "enabled": { + "const": true + } + }, + "required": [ + "enabled" + ] }, "then": { - "required": ["provider"] + "required": [ + "provider" + ] } }, { "if": { "properties": { - "enabled": { "const": true }, - "provider": { "const": "cert-manager" } + "enabled": { + "const": true + }, + "provider": { + "const": "cert-manager" + } }, - "required": ["enabled", "provider"] + "required": [ + "enabled", + "provider" + ] }, "then": { - "required": ["cert-manager"], + "required": [ + "cert-manager" + ], "properties": { "cert-manager": { - "required": ["cluster-issuer"], + "required": [ + "cluster-issuer" + ], "properties": { - "cluster-issuer": { "minLength": 1 } + "cluster-issuer": { + "minLength": 1 + } } } } }, "else": { "properties": { - "mutual_enabled": { "const": false } + "mutual_enabled": { + "const": false + } } } } @@ -78,18 +118,34 @@ "type": "object", "additionalProperties": false, "properties": { - "enabled": { "type": "boolean" }, - "url": { "type": "string" } + "enabled": { + "type": "boolean" + }, + "url": { + "type": "string" + } }, "allOf": [ { "if": { - "properties": { "enabled": { "const": true } }, - "required": ["enabled"] + "properties": { + "enabled": { + "const": true + } + }, + "required": [ + "enabled" + ] }, "then": { - "required": ["url"], - "properties": { "url": { "minLength": 1 } } + "required": [ + "url" + ], + "properties": { + "url": { + "minLength": 1 + } + } } } ] @@ -98,18 +154,34 @@ "type": "object", "additionalProperties": false, "properties": { - "enabled": { "type": "boolean" }, - "url": { "type": "string" } + "enabled": { + "type": "boolean" + }, + "url": { + "type": "string" + } }, "allOf": [ { "if": { - "properties": { "enabled": { "const": true } }, - "required": ["enabled"] + "properties": { + "enabled": { + "const": true + } + }, + "required": [ + "enabled" + ] }, "then": { - "required": ["url"], - "properties": { "url": { "minLength": 1 } } + "required": [ + "url" + ], + "properties": { + "url": { + "minLength": 1 + } + } } } ] @@ -118,22 +190,51 @@ "type": "object", "additionalProperties": false, "properties": { - "enabled": { "type": "boolean" }, - "integrationKey": { "type": "string" }, - "severity": { "enum": ["info", "warning", "error", "critical"] }, - "class": { "type": "string" }, - "component": { "type": "string" }, - "group": { "type": "string" } + "enabled": { + "type": "boolean" + }, + "integrationKey": { + "type": "string" + }, + "severity": { + "enum": [ + "info", + "warning", + "error", + "critical" + ] + }, + "class": { + "type": "string" + }, + "component": { + "type": "string" + }, + "group": { + "type": "string" + } }, "allOf": [ { "if": { - "properties": { "enabled": { "const": true } }, - "required": ["enabled"] + "properties": { + "enabled": { + "const": true + } + }, + "required": [ + "enabled" + ] }, "then": { - "required": ["integrationKey"], - "properties": { "integrationKey": { "minLength": 1 } } + "required": [ + "integrationKey" + ], + "properties": { + "integrationKey": { + "minLength": 1 + } + } } } ] @@ -142,22 +243,50 @@ "type": "object", "additionalProperties": false, "properties": { - "enabled": { "type": "boolean" }, - "apiKey": { "type": "string" }, - "apiUrl": { "type": "string" }, - "autoClose": { "type": "boolean" }, - "overridePriority": { "type": "boolean" }, - "sendTagsAs": { "enum": ["tags", "details", "both"] } + "enabled": { + "type": "boolean" + }, + "apiKey": { + "type": "string" + }, + "apiUrl": { + "type": "string" + }, + "autoClose": { + "type": "boolean" + }, + "overridePriority": { + "type": "boolean" + }, + "sendTagsAs": { + "enum": [ + "tags", + "details", + "both" + ] + } }, "allOf": [ { "if": { - "properties": { "enabled": { "const": true } }, - "required": ["enabled"] + "properties": { + "enabled": { + "const": true + } + }, + "required": [ + "enabled" + ] }, "then": { - "required": ["apiKey"], - "properties": { "apiKey": { "minLength": 1 } } + "required": [ + "apiKey" + ], + "properties": { + "apiKey": { + "minLength": 1 + } + } } } ] @@ -166,24 +295,56 @@ "type": "object", "additionalProperties": false, "properties": { - "enabled": { "type": "boolean" }, - "url": { "type": "string" }, - "httpMethod": { "enum": ["POST", "PUT"] }, - "username": { "type": "string" }, - "password": { "type": "string" }, - "authorizationScheme": { "type": "string" }, - "authorizationCredentials": { "type": "string" }, - "maxAlerts": { "type": "integer", "minimum": 0 } + "enabled": { + "type": "boolean" + }, + "url": { + "type": "string" + }, + "httpMethod": { + "enum": [ + "POST", + "PUT" + ] + }, + "username": { + "type": "string" + }, + "password": { + "type": "string" + }, + "authorizationScheme": { + "type": "string" + }, + "authorizationCredentials": { + "type": "string" + }, + "maxAlerts": { + "type": "integer", + "minimum": 0 + } }, "allOf": [ { "if": { - "properties": { "enabled": { "const": true } }, - "required": ["enabled"] + "properties": { + "enabled": { + "const": true + } + }, + "required": [ + "enabled" + ] }, "then": { - "required": ["url"], - "properties": { "url": { "minLength": 1 } } + "required": [ + "url" + ], + "properties": { + "url": { + "minLength": 1 + } + } } } ] @@ -195,6 +356,87 @@ } } } + }, + "driver": { + "type": "object", + "additionalProperties": false, + "properties": { + "driverName": { + "type": "string", + "minLength": 1 + }, + "image": { + "type": "object", + "additionalProperties": false, + "properties": { + "repository": { + "type": "string" + }, + "tag": { + "type": "string" + }, + "pullPolicy": { + "enum": [ + "Always", + "Never", + "IfNotPresent" + ] + } + } + }, + "controllerReplicas": { + "type": "integer", + "minimum": 1 + }, + "nodeSelector": { + "type": "object" + }, + "tolerations": { + "type": "array" + }, + "controllerNodeSelector": { + "type": "object" + }, + "controllerTolerations": { + "type": "array" + }, + "enableVolumeSnapshots": { + "type": "boolean" + }, + "enableServiceAccountAuth": { + "type": "boolean" + }, + "sidecarImages": { + "type": "object", + "additionalProperties": false, + "properties": { + "provisioner": { + "type": "string", + "pattern": "^($|(quay\\.io/simplyblock-io|docker\\.io/simplyblock|public\\.ecr\\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$" + }, + "attacher": { + "type": "string", + "pattern": "^($|(quay\\.io/simplyblock-io|docker\\.io/simplyblock|public\\.ecr\\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$" + }, + "resizer": { + "type": "string", + "pattern": "^($|(quay\\.io/simplyblock-io|docker\\.io/simplyblock|public\\.ecr\\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$" + }, + "snapshotter": { + "type": "string", + "pattern": "^($|(quay\\.io/simplyblock-io|docker\\.io/simplyblock|public\\.ecr\\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$" + }, + "healthMonitor": { + "type": "string", + "pattern": "^($|(quay\\.io/simplyblock-io|docker\\.io/simplyblock|public\\.ecr\\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$" + }, + "nodeDriverRegistrar": { + "type": "string", + "pattern": "^($|(quay\\.io/simplyblock-io|docker\\.io/simplyblock|public\\.ecr\\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$" + } + } + } + } } } } diff --git a/helm-charts/charts/simplyblock-operator/values.yaml b/helm-charts/charts/simplyblock-operator/values.yaml index df153308c..460c34a01 100644 --- a/helm-charts/charts/simplyblock-operator/values.yaml +++ b/helm-charts/charts/simplyblock-operator/values.yaml @@ -161,6 +161,74 @@ deployment: # what it collects in the object store a hosted control plane brings. profile: standalone +# The CSI driver, as the SimplyblockDriver the chart renders for every profile +# that runs one. +# +# The plugins themselves are the operator's: it applies the DaemonSet, the +# StatefulSet, their RBAC, and the CSIDriver registration from this object, so +# these values reach them through a spec rather than through a template. What +# each one replaces is in design-simplyblockdriver.md §8. +# +# An empty field means whatever this operator release ships, which is the +# pairing the release was tested as. Set one only to move away from it. +driver: + # The CSI driver name a StorageClass provisions with. + # + # It is immutable once the object exists, and every PersistentVolume the + # driver created records it, so changing it later orphans every volume rather + # than renaming anything. An upgraded cluster registered under another name has + # to set it here before the operator adopts the running deployment. + driverName: csi.simplyblock.io + + # The plugin image both plugins run. Empty takes the operator's own registry + # and tag with the CSI driver's repository. + image: + repository: "" + tag: "" + # Always, because the default image follows a moving tag: IfNotPresent + # against a tag that was rebuilt leaves some workers on the old plugin and + # some on the new one. + pullPolicy: Always + + # Instances of the controller plugin, which is what provisions volumes. + controllerReplicas: 1 + + # Where the node plugin runs. Empty is every schedulable worker, which is the + # usual case: a worker that cannot attach a volume cannot run a workload that + # needs one. + nodeSelector: {} + tolerations: [] + + # Where the controller plugin runs. This is ordinary pod placement for one + # workload, unlike the pair above. + controllerNodeSelector: {} + controllerTolerations: [] + + # Whether snapshot support is part of this deployment: the VolumeSnapshotClass + # for driverName, and the CRDs and a controller where the cluster serves + # neither. + enableVolumeSnapshots: true + + # Authenticate both plugins to the management API with their pod's + # service-account token instead of the static cluster secret. + # + # Set it together with controlplane.trustCSIServiceAccounts above. This is the + # plugins' half and that is the control plane's: one without the other is + # either a driver whose tokens are refused, or a control plane trusting + # accounts nothing presents. + enableServiceAccountAuth: false + + # Overrides for the six CSI sidecars, as full `repository:tag` references. + # Empty takes the version this operator release ships, which is the + # combination it was tested against. + sidecarImages: + provisioner: "" + attacher: "" + resizer: "" + snapshotter: "" + healthMonitor: "" + nodeDriverRegistrar: "" + controlplane: # Where the control plane is, for deployment.profile: managed. managed: diff --git a/helm-charts/scripts/check-rendered-objects.sh b/helm-charts/scripts/check-rendered-objects.sh index 9dca24879..9001ef0fc 100755 --- a/helm-charts/scripts/check-rendered-objects.sh +++ b/helm-charts/scripts/check-rendered-objects.sh @@ -17,6 +17,10 @@ COMMON=( "ValidatingWebhookConfiguration/simplyblock-operator-validating-webhook-configuration" "ServiceAccount/simplyblock-operator" "ControlPlane/simplyblock" + # Both profiles run workloads that mount simplyblock volumes, so both need a + # CSI driver, and nothing but this object produces one: the chart stopped + # rendering the plugins when the operator took them over. + "SimplyblockDriver/simplyblock" ) # objects prints `Kind/name` for every document in a rendered manifest. The name diff --git a/operator/docs/designs/crd-redesign/design-simplyblockdriver.md b/operator/docs/designs/crd-redesign/design-simplyblockdriver.md index bf5596744..17d790b69 100644 --- a/operator/docs/designs/crd-redesign/design-simplyblockdriver.md +++ b/operator/docs/designs/crd-redesign/design-simplyblockdriver.md @@ -511,13 +511,25 @@ derives them, so `simplyblock-csi-node-role` identifies one object the way #### The spec is seeded from what is running **Upgrade tool:** write the `SimplyblockDriver` for an existing deployment, with -its spec seeded from the table below rather than left to the defaults. Nothing -else does: the chart no longer renders the object, so on an upgraded cluster it -is written by hand or not at all. `driverName` is the row that costs the most, -because omitting it defaults to `csi.simplyblock.io` on a deployment registered -under another name, the field is immutable, and the object is then unrepairable -by an edit. The controller refuses such an object rather than orphaning volumes -(§4.3, step 2), which contains the damage without removing the need. +its spec seeded from the table below rather than left to the defaults. +`driverName` is the row that costs the most, because omitting it defaults to +`csi.simplyblock.io` on a deployment registered under another name, the field is +immutable, and the object is then unrepairable by an edit. The controller refuses +such an object rather than orphaning volumes (§4.3, step 2), which contains the +damage without removing the need. + +**The chart renders the object**, for every deployment profile that runs a CSI +driver, which today is both of them. The earlier reading of this section was that +it did not, and that left the kind with no producer at all: a fresh install came +up with an operator, a control plane, and nothing able to provision a volume. So +the chart writes it with the values below on its spec, and this table stays the +upgrade tool's, because the two cases differ in what the spec has to say. A fresh +install is describing a deployment that does not exist yet and the defaults are +right for it. An upgrade is describing one that is already running, where a +default is a reconfiguration of a live deployment on the pass that adopts it, and +`driverName` is a reconfiguration nothing can undo. An upgraded cluster therefore +still needs the seeding, either from the tool or by setting `driver.driverName` +and the rest in `values.yaml` before the chart upgrade applies the object. **The first reconcile after adoption has to be a no-op**, because these objects were rendered from Helm values and are about to be rendered from a spec. A field @@ -852,9 +864,10 @@ reconcile that takes it over in place. **Upgrade tool:** apply the `SimplyblockDriver` CRD before the chart upgrade that expects it. The chart ships CRDs in `crds/`, which Helm applies on install -and skips on upgrade, so an existing release upgraded onto a chart that -references the kind fails while rendering, with an error naming the template -rather than the missing CRD. §11 installs the CRDs, and this is one of them. +and skips on upgrade, so an existing release upgraded onto a chart that renders +the object templates it happily and fails on the apply, with +`no matches for kind "SimplyblockDriver"` and no indication that the CRD is the +thing missing. §11 installs the CRDs, and this is one of them. **Moving the install out of the chart is not free**, and it is the same cost [`design-controlplane.md`](design-controlplane.md) §5.1 names for the control From ab9429d5404ea85e9e82985b516fa4bf52b5be00 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 10:29:27 +0200 Subject: [PATCH 032/206] feat(controlplane): a created migration carries the paths to validate MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A migration copies nothing until it is continued, and the window between the two is what the target's NVMe-oF paths are for: every host consuming the subsystem has to be able to reach the target there before the copy starts, which is checked by connecting them and finding them established and not yet serving. The control plane publishes them with the migration's creation and with nothing else — a migration read back by id carries none — so a client that dropped them left its caller nothing to check and no way to ask again. Both DTOs already carry the field. The entry is the same model /connect answers with, so the mapping onto lvol.Endpoint is shared rather than written a second time. A volume's migration also gains the target NQN, which its own shape has no field for: the addresses are no use without the subsystem to ask for at them. Co-Authored-By: Claude Fable 5 --- atlas-lib/controlplane/migrations.go | 47 ++++++++++++- atlas-lib/controlplane/migrations_test.go | 80 +++++++++++++++++++++++ atlas-lib/controlplane/volumes.go | 48 ++++++++------ 3 files changed, 153 insertions(+), 22 deletions(-) diff --git a/atlas-lib/controlplane/migrations.go b/atlas-lib/controlplane/migrations.go index f46a405e1..f1bd84ecc 100644 --- a/atlas-lib/controlplane/migrations.go +++ b/atlas-lib/controlplane/migrations.go @@ -7,6 +7,7 @@ import ( "github.com/simplyblock/atlas/errs" "github.com/simplyblock/atlas/internal/cpapi" + "github.com/simplyblock/atlas/lvol" ) // MigrationKind is what one migration moves. @@ -52,14 +53,29 @@ type Migration struct { SnapsTotal int // TargetNQN, MemberCount, and ClusterID describe a subsystem's migration: - // the subsystem the members land on, and how many there are. + // the subsystem the members land on, and how many there are. TargetNQN is + // also filled for a volume's migration, from the paths below, because the + // shape carries no such field of its own and the NQN is what a host asks + // the target for. TargetNQN string MemberCount int ClusterID string + + // Paths are the NVMe-oF endpoints the target answers the migrated + // subsystem on, in the control plane's priority order. + // + // They arrive with the migration's creation and with nothing else: a + // migration read back by id carries none, so a caller that means to use + // them keeps what the create returned. That window is what they are for. + // A created migration copies nothing until it is continued, and in between + // every host consuming the subsystem has to be able to reach the target, + // which is checked by connecting these paths and finding them established + // and not yet serving. + Paths []lvol.Endpoint } func volumeMigration(d cpapi.MigrationDTO) Migration { - return Migration{ + m := Migration{ Kind: MigrationOfVolume, ID: d.Id.String(), SourceNodeID: d.SourceNodeId, @@ -73,10 +89,12 @@ func volumeMigration(d cpapi.MigrationDTO) Migration { SnapsMigrated: d.SnapsMigrated, SnapsTotal: d.SnapsTotal, } + m.Paths, m.TargetNQN = migrationPaths(d.ConnectStrings) + return m } func subsystemMigration(d cpapi.BatchMigrationDTO) Migration { - return Migration{ + m := Migration{ Kind: MigrationOfSubsystem, ID: d.Id.String(), SourceNodeID: d.SourceNodeId, @@ -88,6 +106,29 @@ func subsystemMigration(d cpapi.BatchMigrationDTO) Migration { MemberCount: d.MemberCount, ClusterID: d.ClusterId, } + paths, nqn := migrationPaths(d.ConnectStrings) + m.Paths = paths + if m.TargetNQN == "" { + m.TargetNQN = nqn + } + return m +} + +// migrationPaths reads a migration's connect entries, and returns the NQN they +// lead to alongside them. +// +// The NQN comes off the first entry rather than being collected from all of +// them: every path of one migration answers the same subsystem, since that is +// what makes them paths to one thing rather than a list of unrelated targets. +func migrationPaths(entries *[]cpapi.NvmeConnectEntry) ([]lvol.Endpoint, string) { + if entries == nil || len(*entries) == 0 { + return nil, "" + } + paths := make([]lvol.Endpoint, 0, len(*entries)) + for _, e := range *entries { + paths = append(paths, endpointOf(e)) + } + return paths, (*entries)[0].Nqn } // migrationFromJSON reads whichever shape arrived. diff --git a/atlas-lib/controlplane/migrations_test.go b/atlas-lib/controlplane/migrations_test.go index 23fa8b528..343cd26b1 100644 --- a/atlas-lib/controlplane/migrations_test.go +++ b/atlas-lib/controlplane/migrations_test.go @@ -216,6 +216,86 @@ func TestClientRefusesAMigrationOfNoKnownKind(t *testing.T) { } } +// connectStringsJSON is what the control plane publishes alongside a created +// migration: the NVMe-oF paths the target now answers on, in the same shape +// /connect returns, with the hyphenated keys of the `nvme connect` options. +const connectStringsJSON = `,"connect_strings":[` + + `{"transport":"tcp","ip":"10.10.10.1","port":4420,"nqn":"nqn.target",` + + `"reconnect-delay":2,"ctrl-loss-tmo":60,"fast-io-fail-tmo":0,"nr-io-queues":8,` + + `"keep-alive-tmo":5,"connect":"nvme connect ..."},` + + `{"transport":"tcp","ip":"10.10.10.2","port":4420,"nqn":"nqn.target",` + + `"reconnect-delay":2,"ctrl-loss-tmo":60,"fast-io-fail-tmo":0,"nr-io-queues":8,` + + `"keep-alive-tmo":5,"connect":"nvme connect ..."}]` + +// TestClientCreateMigrationReturnsTheTargetPaths. A migration is created +// before it copies anything, and the window between the two is what the paths +// are for: the target answers on them, every host consuming the subsystem has +// to be able to reach it there, and a caller checks that before it lets the +// copy start. They arrive only in the create's answer, so a client that read +// the migration and dropped them would leave its caller nothing to check and +// no way to ask again. +func TestClientCreateMigrationReturnsTheTargetPaths(t *testing.T) { + const target = "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa" + body := strings.TrimSuffix(batchMigration, "}") + connectStringsJSON + "}" + c := newTestClient(t, func(w http.ResponseWriter, _ *http.Request) { + w.Header().Set("Content-Type", "application/json") + w.WriteHeader(http.StatusCreated) + _, _ = w.Write([]byte(body)) + }) + + m, err := c.CreateMigration(context.Background(), testCluster, testNQN, target) + if err != nil { + t.Fatal(err) + } + if len(m.Paths) != 2 { + t.Fatalf("paths = %d, want both of them: %+v", len(m.Paths), m.Paths) + } + // Order is the control plane's priority order and is preserved, for the + // reason lvol.Connection.Endpoints gives: the primary path is attached + // first. + if m.Paths[0].Address != "10.10.10.1" || m.Paths[1].Address != "10.10.10.2" { + t.Errorf("paths = %+v, want them in the order they arrived", m.Paths) + } + e := m.Paths[0] + if e.Transport != "tcp" || e.Port != 4420 || e.NrIOQueues != 8 || + e.ReconnectDelaySec != 2 || e.KeepAliveTMOSec != 5 { + t.Errorf("path = %+v, want the connect parameters the control plane chose", e) + } + // Zero means "fail I/O immediately," which is a choice rather than a + // missing value, so both timeouts are pointers and both must survive. + if e.CtrlLossTMOSec == nil || *e.CtrlLossTMOSec != 60 { + t.Errorf("ctrl-loss-tmo = %v, want 60", e.CtrlLossTMOSec) + } + if e.FastIOFailTMOSec == nil || *e.FastIOFailTMOSec != 0 { + t.Errorf("fast-io-fail-tmo = %v, want 0", e.FastIOFailTMOSec) + } +} + +// TestClientCreateMigrationNamesTheSubsystemThePathsLeadTo. A volume's +// migration carries no target_nqn of its own, and the NQN is what a host +// connects to, so without the one on the paths a caller would have the +// addresses and nothing to ask for at them. +func TestClientCreateMigrationNamesTheSubsystemThePathsLeadTo(t *testing.T) { + const target = "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa" + body := strings.TrimSuffix(singleMigration, "}") + connectStringsJSON + "}" + c := newTestClient(t, func(w http.ResponseWriter, _ *http.Request) { + w.Header().Set("Content-Type", "application/json") + w.WriteHeader(http.StatusCreated) + _, _ = w.Write([]byte(body)) + }) + + m, err := c.CreateMigration(context.Background(), testCluster, testNQN, target) + if err != nil { + t.Fatal(err) + } + if m.Kind != MigrationOfVolume { + t.Fatalf("kind = %q, want a volume's migration", m.Kind) + } + if m.TargetNQN != "nqn.target" { + t.Errorf("target NQN = %q, want the one the paths lead to", m.TargetNQN) + } +} + // TestMigrationJSONShapesAreDistinguishable is the assumption the two above // rest on, asserted directly: the discriminating fields are present in one // shape and absent from the other, so neither decodes as the other by accident. diff --git a/atlas-lib/controlplane/volumes.go b/atlas-lib/controlplane/volumes.go index c72b3190d..8c823e3c8 100644 --- a/atlas-lib/controlplane/volumes.go +++ b/atlas-lib/controlplane/volumes.go @@ -79,29 +79,39 @@ func (c *Client) Connection(ctx context.Context, h lvol.VolumeHandle, opts ...lv if conn.NSID == 0 && e.NsId != nil && *e.NsId > 0 { conn.NSID = uint32(*e.NsId) } - conn.Endpoints = append(conn.Endpoints, lvol.Endpoint{ - Transport: e.Transport, - Address: e.Ip, - Port: e.Port, - NrIOQueues: e.NrIoQueues, - ReconnectDelaySec: e.ReconnectDelay, - KeepAliveTMOSec: e.KeepAliveTmo, - // The spec makes both timeouts required, so whatever arrives is - // the control plane's answer, including 0 ("fail I/O - // immediately"), which must not degrade into "unspecified." - CtrlLossTMOSec: ptr.To(e.CtrlLossTmo), - FastIOFailTMOSec: ptr.To(e.FastIoFailTmo), - HostIface: ptr.From(e.HostIface, ""), - TLS: ptr.BoolFromOrFalse(e.Tls), - // The secrets have no fields of their own in the response, and - // the prebuilt command line is where the control plane puts them. - DHCHAPSecret: connectFlag(e.Connect, "--dhchap-secret"), - DHCHAPCtrlSecret: connectFlag(e.Connect, "--dhchap-ctrl-secret"), - }) + conn.Endpoints = append(conn.Endpoints, endpointOf(cpapi.NvmeConnectEntry(e))) } return conn, nil } +// endpointOf reads one of the control plane's connect entries. +// +// It is shared with a migration's paths rather than written twice, because they +// are the same model answered by two endpoints: /connect says how to reach a +// volume where it is, and a created migration says how to reach the target it +// is about to be on (see migrations.go). +func endpointOf(e cpapi.NvmeConnectEntry) lvol.Endpoint { + return lvol.Endpoint{ + Transport: e.Transport, + Address: e.Ip, + Port: e.Port, + NrIOQueues: e.NrIoQueues, + ReconnectDelaySec: e.ReconnectDelay, + KeepAliveTMOSec: e.KeepAliveTmo, + // The spec makes both timeouts required, so whatever arrives is + // the control plane's answer, including 0 ("fail I/O + // immediately"), which must not degrade into "unspecified." + CtrlLossTMOSec: ptr.To(e.CtrlLossTmo), + FastIOFailTMOSec: ptr.To(e.FastIoFailTmo), + HostIface: ptr.From(e.HostIface, ""), + TLS: ptr.BoolFromOrFalse(e.Tls), + // The secrets have no fields of their own in the response, and + // the prebuilt command line is where the control plane puts them. + DHCHAPSecret: connectFlag(e.Connect, "--dhchap-secret"), + DHCHAPCtrlSecret: connectFlag(e.Connect, "--dhchap-ctrl-secret"), + } +} + // connectFlag pulls one `--flag`'s value out of the prebuilt `nvme connect ...` // command line the control plane returns alongside each path, and is how the // DHCHAP secrets are read: the control plane resolves them per (host, From 7be277fc4165ca396b3272632153cb83c1d5cd9d Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 10:32:42 +0200 Subject: [PATCH 033/206] feat(api): PersistentVolumeOps, and the graph its one action runs MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The kind moving a volume between storage nodes, as design-crd-model.md §3 shapes every operation in this group: a phase that says how the operation is doing, a step that says where its work is, and a graph that declares which steps an abort can still be honored from. It is the one Ops kind whose target is a core Kubernetes type, and almost everything unusual about it follows. Its target cannot carry a status.activeOpsRef, so the lock is an annotation on the volume. It is cluster-scoped because a PersistentVolume is, which costs it the owner reference its creator would hold, so the creator is named in the spec with a UID — the shape core Kubernetes uses for the same problem in pv.spec.claimRef and volumeSnapshotContent.spec.volumeSnapshotRef. And it is told no cluster: the volume's CSI handle carries the cluster, pool, and volume UUIDs. Authored at v1alpha2 only. The kind is introduced by the redesign, so there is no v1alpha1 spelling to convert from. The registered VolumeMigration keeps running beside it for at least one release: a rename and a scope change make a new CRD rather than a new version, so in-flight migrations drain on the old kind rather than being carried over. The graph is three steps and a line. Validating creates the migration and checks that every consuming host can reach the target on the paths it published; Migrating continues it, which is what starts the copy; and Verifying takes those paths back. The third is new, and it exists because paths that outlived their Jobs poisoned the data path and blocked every later migration of the volume. It is also the one step an abort is refused from, because by then the volume has moved and the cleanup is what makes the move safe. Only the copy's deadline is computed rather than fixed. A migration is addressed by the subsystem, so every sibling volume moves with the named one, and the bound is a base plus a term per member. Co-Authored-By: Claude Fable 5 --- ...ge.simplyblock.io_persistentvolumeops.yaml | 383 ++++++++++++++++ .../api/v1alpha2/persistentvolumeops_types.go | 424 ++++++++++++++++++ .../api/v1alpha2/zz_generated.deepcopy.go | 255 +++++++++++ ...ge.simplyblock.io_persistentvolumeops.yaml | 383 ++++++++++++++++ operator/config/crd/kustomization.yaml | 1 + operator/config/samples/kustomization.yaml | 1 + .../storage_v1alpha2_persistentvolumeops.yaml | 21 + operator/dist/install.yaml | 383 ++++++++++++++++ .../controllers/volume/cel_validation_test.go | 190 ++++++++ .../internal/controllers/volume/graphs.go | 141 ++++++ .../controllers/volume/graphs_test.go | 166 +++++++ .../internal/controllers/volume/suite_test.go | 91 ++++ ...ge.simplyblock.io_persistentvolumeops.yaml | 383 ++++++++++++++++ 13 files changed, 2822 insertions(+) create mode 100644 helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_persistentvolumeops.yaml create mode 100644 operator/api/v1alpha2/persistentvolumeops_types.go create mode 100644 operator/config/crd/bases/storage.simplyblock.io_persistentvolumeops.yaml create mode 100644 operator/config/samples/storage_v1alpha2_persistentvolumeops.yaml create mode 100644 operator/internal/controllers/volume/cel_validation_test.go create mode 100644 operator/internal/controllers/volume/graphs.go create mode 100644 operator/internal/controllers/volume/graphs_test.go create mode 100644 operator/internal/controllers/volume/suite_test.go create mode 100644 operator/internal/upgrade/crds/manifests/storage.simplyblock.io_persistentvolumeops.yaml diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_persistentvolumeops.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_persistentvolumeops.yaml new file mode 100644 index 000000000..49cc97743 --- /dev/null +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_persistentvolumeops.yaml @@ -0,0 +1,383 @@ +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition +metadata: + annotations: + controller-gen.kubebuilder.io/version: v0.21.0 + name: persistentvolumeops.storage.simplyblock.io +spec: + group: storage.simplyblock.io + names: + kind: PersistentVolumeOps + listKind: PersistentVolumeOpsList + plural: persistentvolumeops + shortNames: + - pvops + singular: persistentvolumeops + scope: Cluster + versions: + - additionalPrinterColumns: + - jsonPath: .spec.persistentVolumeName + name: Volume + type: string + - jsonPath: .spec.action + name: Action + type: string + - jsonPath: .spec.migrate.targetNodeRef.name + name: Target + type: string + - jsonPath: .status.phase + name: Phase + type: string + - jsonPath: .status.step.state + name: Step + type: string + - jsonPath: .status.message + name: Message + priority: 1 + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha2 + schema: + openAPIV3Schema: + description: |- + PersistentVolumeOps is a single operation performed against one + PersistentVolume. It is the one Ops kind in this group whose target is a core + Kubernetes type rather than a kind this group defines, so it locks its target + with an annotation rather than a status field, is cluster-scoped because its + target is, cannot be owned by the namespaced operation that created it, and + derives its cluster, pool, and volume from the volume's CSI handle rather + than being told. + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: |- + PersistentVolumeOpsSpec is one operation to perform against one + PersistentVolume. + + The rule keeps the action and its parameter block in agreement, which is a + statement about this object alone and so belongs on the type rather than in + the webhook. + properties: + abort: + description: |- + Abort asks a running operation to stop at its next step and unwind. It is + expressible from Validating and Migrating and not from Verifying, which + the action's graph declares rather than this field: once the copy has + finished, the volume has moved and there is nothing to undo. + type: boolean + action: + description: |- + Action is the operation to perform. Immutable: an operation that changed + what it was doing halfway through would have a status describing neither. + enum: + - Migrate + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + creatorRef: + description: |- + CreatorRef names the object that created this one, and that object's + finalizer is what aborts and deletes this one when it goes. Absent on an + operation written by hand, which has no creator to cascade from. + Immutable, because an operation changing whose fan-out it belongs to + would change who cascades over it. + properties: + kind: + description: Kind is the creating object's kind, which is StorageNodeOps + for a drain. + type: string + name: + description: Name is the creator's object name. + type: string + namespace: + description: Namespace is where the creator lives. + type: string + uid: + description: |- + UID is what makes the reference address one creator rather than one name: + a creator deleted and recreated under the same name must not inherit the + fan-out it did not issue, and cascade a delete over it. + type: string + required: + - kind + - name + - namespace + - uid + type: object + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + migrate: + description: Migrate parameterizes action Migrate. + properties: + targetNodeRef: + description: |- + TargetNodeRef locates the StorageNode to move the volume's backing + logical volume to. It names the Kubernetes object rather than the backend + UUID, so that a migration can be written by hand without looking one up; + the controller resolves the UUID from the node's status. The node's + cluster must be the volume's, which the webhook checks rather than the + type, because that is a fact about two other objects. + + The marker sits on the reference rather than on its fields, so the pair + is immutable together and a target cannot be half-changed into a name in + one cluster and a namespace in another. + properties: + name: + description: Name is the StorageNode object's name. + type: string + namespace: + description: |- + Namespace is where the StorageNode lives, which is the namespace of the + StorageCluster that owns it. + type: string + required: + - name + - namespace + type: object + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + required: + - targetNodeRef + type: object + persistentVolumeName: + description: |- + PersistentVolumeName names the PersistentVolume this operation acts on. + It is a name rather than a reference because a PersistentVolume is + cluster-scoped, and it names the volume rather than the claim because a + claim can be deleted while its volume is retained — under a Retain + reclaim policy the logical volume still occupies capacity on a node that + may be draining, and is still worth moving. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + required: + - action + - persistentVolumeName + type: object + x-kubernetes-validations: + - message: field creatorRef is immutable once set + rule: '!has(oldSelf.creatorRef) || has(self.creatorRef)' + - message: migrate is required for action Migrate and must be absent otherwise + rule: 'self.action == ''Migrate'' ? has(self.migrate) : !has(self.migrate)' + status: + description: PersistentVolumeOpsStatus is the observed state of one volume + operation. + properties: + completedAt: + description: CompletedAt is when it reached a terminal phase. + format: date-time + type: string + deferredSince: + description: |- + DeferredSince is when the operation was first held — behind another + operation's lock, or behind a control plane that is not accepting + migrations yet. It is what the auto-rebalancer reads to decide whether a + migration has waited long enough to give up on, and it is in status + rather than in memory because the operator may restart and an observer + needs to see that the operation is waiting and since when. + format: date-time + type: string + message: + description: |- + Message is the reason the phase is what it is: one sentence, replaced as + the operation moves, and never a log. + type: string + migration: + description: |- + Migration is everything about the migration rather than about the + operation. + properties: + clusterUUID: + description: |- + ClusterUUID, PoolUUID, and VolumeUUID are the three parts of the volume's + CSI volume handle, recorded so that later steps address the backend + without re-reading the PersistentVolume, and so that a failed operation + says which volume it was working on after the volume is gone. + type: string + connections: + description: |- + Connections are the paths the migration published on the target. + Verifying confirms none of them is left connected. + items: + description: |- + MigrationConnection is one NVMe-oF path the migration published on the + target. + + The connect parameters travel with the address because the path is connected + on a consuming host rather than here, and what is recorded has to be the + connect that will actually be made: the host attaches every path with the + same controller-loss timeout the CSI driver uses, which is not the hour the + control plane answers with, and a record of the control plane's answer would + describe a connect nobody performs. + properties: + address: + description: Address and Port are where it answers. + type: string + ctrlLossTimeoutSeconds: + description: |- + CtrlLossTimeoutSeconds and FastIOFailTimeoutSeconds are pointers because + zero is a choice ("fail I/O immediately") rather than a missing value. + format: int32 + type: integer + fastIOFailTimeoutSeconds: + format: int32 + type: integer + keepAliveTimeoutSeconds: + format: int32 + type: integer + nqn: + description: NQN is the subsystem the target answers this + path for. + type: string + nrIOQueues: + format: int32 + type: integer + port: + format: int32 + type: integer + reconnectDelaySeconds: + format: int32 + type: integer + transport: + description: Transport is the fabric, which is tcp. + type: string + type: object + type: array + memberCount: + description: |- + MemberCount is how many volumes the migrated NVMe-oF subsystem holds, as + the control plane reports it. More than one member means the sibling + volumes move along with the named one, so the count is both the + operation's blast radius and the term the copy's deadline scales by. + format: int32 + minimum: 0 + type: integer + migrationUUID: + description: MigrationUUID is the control plane's identifier for + the copy. + type: string + poolUUID: + type: string + sourceNodeUUID: + description: |- + SourceNodeUUID is where the volume was before the move, recorded so that + a failure says what it was and not only what it was going to be. + type: string + subsystemNQN: + description: |- + SubsystemNQN is the volume's NVMe-oF subsystem, which is what the + migration is addressed by: the control plane migrates a subsystem rather + than one volume inside it. + type: string + targetNodeUUID: + description: |- + TargetNodeUUID is the backend identifier resolved from + spec.migrate.targetNodeRef, recorded so the later steps address the + target without resolving the node object again. + type: string + validationJobs: + description: ValidationJobs are the Jobs started to check those + paths. + items: + description: |- + ValidationJob is one Job started to check a path is reachable from one node. + + It is tracked so that Verifying can delete every Job the operation started, + including after a restart. The namespace is recorded for the same reason + spec.migrate.targetNodeRef carries one: a Job is namespaced and this object + is not, so a name alone would not locate it. + properties: + name: + type: string + namespace: + type: string + node: + description: |- + Node is the worker the Job is pinned to, which is a node consuming one of + the migrated subsystem's volumes. + type: string + succeeded: + description: |- + Succeeded records a node whose paths were checked and found ready, so a + restart does not run the check again on a node that already passed and + whose Job its own TTL may already have reaped. + type: boolean + required: + - name + - namespace + type: object + type: array + volumeUUID: + type: string + type: object + observedGeneration: + description: |- + ObservedGeneration is the generation the rest of this status was computed + from, so a stale status can be told from a current one. + format: int64 + type: integer + phase: + description: Phase is the operation's own progress. + enum: + - Pending + - Running + - Succeeded + - Failed + - Aborted + type: string + startedAt: + description: StartedAt is when the operation began. + format: date-time + type: string + step: + description: |- + Step is the position of the running action's state machine, as the shared + statemachine.KubeSnapshot. It is persisted before the side effect that + step performs. The rule is what an Enum marker would do if a marker could + reach a field of a shared type. + properties: + deadline: + description: |- + Deadline is when that state expires, absent when it has none. It is an + absolute instant, so a state whose deadline passed while the controller + was down restores as already expired. + format: date-time + type: string + state: + description: |- + State is the state the machine was in. Empty means the resource has not + been reconciled yet, and restores to the graph's initial state. + type: string + type: object + x-kubernetes-validations: + - message: unknown step + rule: '!has(self.state) || self.state in [''Validating'',''Migrating'',''Verifying'']' + type: object + type: object + served: true + storage: true + subresources: + status: {} diff --git a/operator/api/v1alpha2/persistentvolumeops_types.go b/operator/api/v1alpha2/persistentvolumeops_types.go new file mode 100644 index 000000000..3f8d47a28 --- /dev/null +++ b/operator/api/v1alpha2/persistentvolumeops_types.go @@ -0,0 +1,424 @@ +// PersistentVolumeOps: one imperative operation performed against one +// PersistentVolume, which today means moving its backing logical volume to +// another storage node. +// +// It is the one Ops kind in this group whose target is a core Kubernetes type +// rather than a kind this group defines, and almost everything unusual about it +// follows from that. Its target cannot carry a status.activeOpsRef, so the lock +// moves to an annotation on the volume. It is cluster-scoped because a +// PersistentVolume is, which in turn costs it the owner reference its creator +// would otherwise hold, so the creator is named in the spec instead. And it is +// told no cluster: the volume's CSI handle carries the cluster, pool, and +// volume UUIDs, and that is where all three come from. +// +// The kind is introduced by the redesign, so it has no v1alpha1 spelling, no +// spoke to convert from, and no Hub method: its CRD declares one version. +// +// The registered VolumeMigration stays registered and keeps running beside it +// for at least one release. The two do not convert into one another — a rename +// and a scope change make a new CRD rather than a new version — so in-flight +// migrations are drained on the old kind rather than carried over, and nothing +// creates objects of both kinds for one volume. +// +// design-persistentvolumeops.md is the specification. + +package v1alpha2 + +import ( + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/types" + + "github.com/simplyblock/atlas/statemachine" +) + +// PersistentVolumeOpsLock is the annotation on a PersistentVolume naming the +// operation currently allowed to act on it, and absent when none is. +// +// It is what status.activeOpsRef is for every other kind in this group: taken +// with an optimistic-lock patch so two reconcilers cannot both win, released +// only while it still names the releaser, and released on every terminal path +// including deletion. It sits in metadata rather than status because a +// PersistentVolume is a core type this operator must not add fields to, and it +// touches neither spec nor status, so it cannot conflict with the provisioner +// or the volume's own controllers (design-persistentvolumeops.md §6). +const PersistentVolumeOpsLock = "storage.simplyblock.io/active-ops" + +// PersistentVolumeOpsManagedByLabel carries the kind of the controller that +// created an operation, which is what a watch mapping and a List select on: a +// reference in the spec cannot be selected on, and a label value admits no +// namespace separator (design-crd-model.md §7.3). +// +// It narrows rather than identifies. The label finds the operations some drain +// created, and the UID in spec.creatorRef says which drain. +const PersistentVolumeOpsManagedByLabel = "storage.simplyblock.io/managed-by" + +// PersistentVolumeOpsAction is the operation a PersistentVolumeOps performs. +// The kind is named for its target rather than for the action so that carrying +// a second one later would not rename it; it carries one, and no second one is +// planned (design-persistentvolumeops.md §4.1). +// +kubebuilder:validation:Enum=Migrate +type PersistentVolumeOpsAction string + +const ( + // PersistentVolumeOpsActionMigrate moves the volume's backing logical + // volume to a different storage node. + PersistentVolumeOpsActionMigrate PersistentVolumeOpsAction = "Migrate" +) + +// PersistentVolumeOpsPhase is the operation's own progress. Succeeded rather +// than the registered kind's Completed, so that every Ops kind in this group +// reports the same five phases and an alert on completion matches one value. +// +kubebuilder:validation:Enum=Pending;Running;Succeeded;Failed;Aborted +type PersistentVolumeOpsPhase string + +const ( + // PersistentVolumeOpsPhasePending is an operation waiting for the volume's + // lock, or holding nothing yet. + PersistentVolumeOpsPhasePending PersistentVolumeOpsPhase = "Pending" + // PersistentVolumeOpsPhaseRunning is an operation holding the lock and + // working. + PersistentVolumeOpsPhaseRunning PersistentVolumeOpsPhase = "Running" + // PersistentVolumeOpsPhaseSucceeded is a finished operation that did what + // it said. + PersistentVolumeOpsPhaseSucceeded PersistentVolumeOpsPhase = "Succeeded" + // PersistentVolumeOpsPhaseFailed is a finished operation that did not. + PersistentVolumeOpsPhaseFailed PersistentVolumeOpsPhase = "Failed" + // PersistentVolumeOpsPhaseAborted is an operation stopped on request, or + // stopped because its volume went away, whose unwind has finished. + PersistentVolumeOpsPhaseAborted PersistentVolumeOpsPhase = "Aborted" +) + +// PersistentVolumeOpsStep is one step of a running volume operation. The +// registered kind merges these with the phases above into one enum, so that +// Validating sits beside Completed and neither can be read without the other's +// values in mind. +// +kubebuilder:validation:Enum=Validating;Migrating;Verifying +type PersistentVolumeOpsStep string + +const ( + // PersistentVolumeOpsStepValidating creates the backend migration and + // starts a Job on every node consuming the subsystem, to check that each of + // them can reach the target on the paths the migration published. + PersistentVolumeOpsStepValidating PersistentVolumeOpsStep = "Validating" + // PersistentVolumeOpsStepMigrating continues the migration, which is what + // starts the data copy. It is not called Running, because status.phase + // already has that value and a step sharing a phase's name is the confusion + // this split exists to end. + PersistentVolumeOpsStepMigrating PersistentVolumeOpsStep = "Migrating" + // PersistentVolumeOpsStepVerifying deletes the validation Jobs and confirms + // no path was left connected. It is a declared step rather than a deferred + // call because a crash between the copy finishing and the cleanup must + // restart into it: paths that outlived their Jobs poisoned the data path + // and blocked every later migration of the volume. + PersistentVolumeOpsStepVerifying PersistentVolumeOpsStep = "Verifying" +) + +// StorageNodeReference locates a StorageNode from a cluster-scoped object. +// +// Every other reference in this API group is a bare string, which works because +// both objects are namespaced and a name means the same namespace. A +// cluster-scoped object has no namespace for a bare name to mean, and two +// clusters in two namespaces may each hold a node called worker-3, so this one +// carries both — for the same reason pv.spec.claimRef does in core Kubernetes. +type StorageNodeReference struct { + // Namespace is where the StorageNode lives, which is the namespace of the + // StorageCluster that owns it. + // +kubebuilder:validation:Required + Namespace string `json:"namespace"` + + // Name is the StorageNode object's name. + // +kubebuilder:validation:Required + Name string `json:"name"` +} + +// MigrateVolumeSpec parameterizes the Migrate action. +type MigrateVolumeSpec struct { + // TargetNodeRef locates the StorageNode to move the volume's backing + // logical volume to. It names the Kubernetes object rather than the backend + // UUID, so that a migration can be written by hand without looking one up; + // the controller resolves the UUID from the node's status. The node's + // cluster must be the volume's, which the webhook checks rather than the + // type, because that is a fact about two other objects. + // + // The marker sits on the reference rather than on its fields, so the pair + // is immutable together and a target cannot be half-changed into a name in + // one cluster and a namespace in another. + // +kubebuilder:validation:Required + // +k8s:immutable + TargetNodeRef StorageNodeReference `json:"targetNodeRef"` +} + +// CreatorReference names the object that created a PersistentVolumeOps. +// +// It exists because a cluster-scoped object cannot be owned by a namespaced +// one: Kubernetes treats such a reference as unresolvable and garbage-collects +// the dependent. Core Kubernetes has the same situation twice and answers it +// the same way — a PersistentVolume names its claim through spec.claimRef and a +// VolumeSnapshotContent names its snapshot through spec.volumeSnapshotRef, both +// with a UID — and that shape carries over here unchanged +// (design-persistentvolumeops.md §11.1). +type CreatorReference struct { + // Kind is the creating object's kind, which is StorageNodeOps for a drain. + // +kubebuilder:validation:Required + Kind string `json:"kind"` + + // Namespace is where the creator lives. + // +kubebuilder:validation:Required + Namespace string `json:"namespace"` + + // Name is the creator's object name. + // +kubebuilder:validation:Required + Name string `json:"name"` + + // UID is what makes the reference address one creator rather than one name: + // a creator deleted and recreated under the same name must not inherit the + // fan-out it did not issue, and cascade a delete over it. + // +kubebuilder:validation:Required + UID types.UID `json:"uid"` +} + +// PersistentVolumeOpsSpec is one operation to perform against one +// PersistentVolume. +// +// The rule keeps the action and its parameter block in agreement, which is a +// statement about this object alone and so belongs on the type rather than in +// the webhook. +// +kubebuilder:validation:XValidation:rule="self.action == 'Migrate' ? has(self.migrate) : !has(self.migrate)",message="migrate is required for action Migrate and must be absent otherwise" +type PersistentVolumeOpsSpec struct { + // PersistentVolumeName names the PersistentVolume this operation acts on. + // It is a name rather than a reference because a PersistentVolume is + // cluster-scoped, and it names the volume rather than the claim because a + // claim can be deleted while its volume is retained — under a Retain + // reclaim policy the logical volume still occupies capacity on a node that + // may be draining, and is still worth moving. + // +kubebuilder:validation:Required + // +k8s:immutable + PersistentVolumeName string `json:"persistentVolumeName"` + + // Action is the operation to perform. Immutable: an operation that changed + // what it was doing halfway through would have a status describing neither. + // +kubebuilder:validation:Required + // +k8s:immutable + Action PersistentVolumeOpsAction `json:"action"` + + // Abort asks a running operation to stop at its next step and unwind. It is + // expressible from Validating and Migrating and not from Verifying, which + // the action's graph declares rather than this field: once the copy has + // finished, the volume has moved and there is nothing to undo. + // +optional + Abort bool `json:"abort,omitempty"` + + // Migrate parameterizes action Migrate. + // +optional + Migrate *MigrateVolumeSpec `json:"migrate,omitempty"` + + // CreatorRef names the object that created this one, and that object's + // finalizer is what aborts and deletes this one when it goes. Absent on an + // operation written by hand, which has no creator to cascade from. + // Immutable, because an operation changing whose fan-out it belongs to + // would change who cascades over it. + // +optional + // +k8s:immutable + CreatorRef *CreatorReference `json:"creatorRef,omitempty"` +} + +// MigrationConnection is one NVMe-oF path the migration published on the +// target. +// +// The connect parameters travel with the address because the path is connected +// on a consuming host rather than here, and what is recorded has to be the +// connect that will actually be made: the host attaches every path with the +// same controller-loss timeout the CSI driver uses, which is not the hour the +// control plane answers with, and a record of the control plane's answer would +// describe a connect nobody performs. +type MigrationConnection struct { + // NQN is the subsystem the target answers this path for. + // +optional + NQN string `json:"nqn,omitempty"` + // Address and Port are where it answers. + // +optional + Address string `json:"address,omitempty"` + // +optional + Port *int32 `json:"port,omitempty"` + // Transport is the fabric, which is TCP. + // +optional + Transport string `json:"transport,omitempty"` + + // +optional + NrIOQueues *int32 `json:"nrIOQueues,omitempty"` + // +optional + ReconnectDelaySeconds *int32 `json:"reconnectDelaySeconds,omitempty"` + // CtrlLossTimeoutSeconds and FastIOFailTimeoutSeconds are pointers because + // zero is a choice ("fail I/O immediately") rather than a missing value. + // +optional + CtrlLossTimeoutSeconds *int32 `json:"ctrlLossTimeoutSeconds,omitempty"` + // +optional + FastIOFailTimeoutSeconds *int32 `json:"fastIOFailTimeoutSeconds,omitempty"` + // +optional + KeepAliveTimeoutSeconds *int32 `json:"keepAliveTimeoutSeconds,omitempty"` +} + +// ValidationJob is one Job started to check a path is reachable from one node. +// +// It is tracked so that Verifying can delete every Job the operation started, +// including after a restart. The namespace is recorded for the same reason +// spec.migrate.targetNodeRef carries one: a Job is namespaced and this object +// is not, so a name alone would not locate it. +type ValidationJob struct { + // +kubebuilder:validation:Required + Namespace string `json:"namespace"` + // +kubebuilder:validation:Required + Name string `json:"name"` + + // Node is the worker the Job is pinned to, which is a node consuming one of + // the migrated subsystem's volumes. + // +optional + Node string `json:"node,omitempty"` + + // Succeeded records a node whose paths were checked and found ready, so a + // restart does not run the check again on a node that already passed and + // whose Job its own TTL may already have reaped. + // +optional + Succeeded bool `json:"succeeded,omitempty"` +} + +// MigrationStatus is everything about the migration rather than about the +// operation. It is durable working state: a controller that restarts mid-copy +// reads it to find the backend migration it started and the Jobs it has to +// clean up (design-crd-model.md §3.1). +type MigrationStatus struct { + // MigrationUUID is the control plane's identifier for the copy. + // +optional + MigrationUUID string `json:"migrationUUID,omitempty"` + + // ClusterUUID, PoolUUID, and VolumeUUID are the three parts of the volume's + // CSI volume handle, recorded so that later steps address the backend + // without re-reading the PersistentVolume, and so that a failed operation + // says which volume it was working on after the volume is gone. + // +optional + ClusterUUID string `json:"clusterUUID,omitempty"` + // +optional + PoolUUID string `json:"poolUUID,omitempty"` + // +optional + VolumeUUID string `json:"volumeUUID,omitempty"` + + // SubsystemNQN is the volume's NVMe-oF subsystem, which is what the + // migration is addressed by: the control plane migrates a subsystem rather + // than one volume inside it. + // +optional + SubsystemNQN string `json:"subsystemNQN,omitempty"` + + // SourceNodeUUID is where the volume was before the move, recorded so that + // a failure says what it was and not only what it was going to be. + // +optional + SourceNodeUUID string `json:"sourceNodeUUID,omitempty"` + + // TargetNodeUUID is the backend identifier resolved from + // spec.migrate.targetNodeRef, recorded so the later steps address the + // target without resolving the node object again. + // +optional + TargetNodeUUID string `json:"targetNodeUUID,omitempty"` + + // MemberCount is how many volumes the migrated NVMe-oF subsystem holds, as + // the control plane reports it. More than one member means the sibling + // volumes move along with the named one, so the count is both the + // operation's blast radius and the term the copy's deadline scales by. + // +kubebuilder:validation:Minimum=0 + // +optional + MemberCount *int32 `json:"memberCount,omitempty"` + + // Connections are the paths the migration published on the target. + // Verifying confirms none of them is left connected. + // +optional + Connections []MigrationConnection `json:"connections,omitempty"` + + // ValidationJobs are the Jobs started to check those paths. + // +optional + ValidationJobs []ValidationJob `json:"validationJobs,omitempty"` +} + +// PersistentVolumeOpsStatus is the observed state of one volume operation. +type PersistentVolumeOpsStatus struct { + // Phase is the operation's own progress. + // +optional + Phase PersistentVolumeOpsPhase `json:"phase,omitempty"` + + // Step is the position of the running action's state machine, as the shared + // statemachine.KubeSnapshot. It is persisted before the side effect that + // step performs. The rule is what an Enum marker would do if a marker could + // reach a field of a shared type. + // +kubebuilder:validation:XValidation:rule="!has(self.state) || self.state in ['Validating','Migrating','Verifying']",message="unknown step" + // +optional + Step statemachine.KubeSnapshot `json:"step,omitempty"` + + // Migration is everything about the migration rather than about the + // operation. + // +optional + Migration *MigrationStatus `json:"migration,omitempty"` + + // DeferredSince is when the operation was first held — behind another + // operation's lock, or behind a control plane that is not accepting + // migrations yet. It is what the auto-rebalancer reads to decide whether a + // migration has waited long enough to give up on, and it is in status + // rather than in memory because the operator may restart and an observer + // needs to see that the operation is waiting and since when. + // +optional + DeferredSince *metav1.Time `json:"deferredSince,omitempty"` + + // Message is the reason the phase is what it is: one sentence, replaced as + // the operation moves, and never a log. + // +optional + Message string `json:"message,omitempty"` + + // ObservedGeneration is the generation the rest of this status was computed + // from, so a stale status can be told from a current one. + // +optional + ObservedGeneration int64 `json:"observedGeneration,omitempty"` + + // StartedAt is when the operation began. + // +optional + StartedAt *metav1.Time `json:"startedAt,omitempty"` + + // CompletedAt is when it reached a terminal phase. + // +optional + CompletedAt *metav1.Time `json:"completedAt,omitempty"` +} + +// +kubebuilder:object:root=true +// +kubebuilder:subresource:status +// +kubebuilder:resource:scope=Cluster,shortName=pvops +// +kubebuilder:printcolumn:name="Volume",type=string,JSONPath=".spec.persistentVolumeName" +// +kubebuilder:printcolumn:name="Action",type=string,JSONPath=".spec.action" +// +kubebuilder:printcolumn:name="Target",type=string,JSONPath=".spec.migrate.targetNodeRef.name" +// +kubebuilder:printcolumn:name="Phase",type=string,JSONPath=".status.phase" +// +kubebuilder:printcolumn:name="Step",type=string,JSONPath=".status.step.state" +// +kubebuilder:printcolumn:name="Message",type=string,JSONPath=".status.message",priority=1 +// +kubebuilder:printcolumn:name="Age",type=date,JSONPath=".metadata.creationTimestamp" + +// PersistentVolumeOps is a single operation performed against one +// PersistentVolume. It is the one Ops kind in this group whose target is a core +// Kubernetes type rather than a kind this group defines, so it locks its target +// with an annotation rather than a status field, is cluster-scoped because its +// target is, cannot be owned by the namespaced operation that created it, and +// derives its cluster, pool, and volume from the volume's CSI handle rather +// than being told. +type PersistentVolumeOps struct { + metav1.TypeMeta `json:",inline"` + metav1.ObjectMeta `json:"metadata,omitempty"` + + Spec PersistentVolumeOpsSpec `json:"spec,omitempty"` + Status PersistentVolumeOpsStatus `json:"status,omitempty"` +} + +// +kubebuilder:object:root=true + +// PersistentVolumeOpsList contains a list of PersistentVolumeOps. +type PersistentVolumeOpsList struct { + metav1.TypeMeta `json:",inline"` + metav1.ListMeta `json:"metadata,omitempty"` + Items []PersistentVolumeOps `json:"items"` +} + +func init() { + SchemeBuilder.Register(&PersistentVolumeOps{}, &PersistentVolumeOpsList{}) +} diff --git a/operator/api/v1alpha2/zz_generated.deepcopy.go b/operator/api/v1alpha2/zz_generated.deepcopy.go index 954385cfe..4f20cdeba 100644 --- a/operator/api/v1alpha2/zz_generated.deepcopy.go +++ b/operator/api/v1alpha2/zz_generated.deepcopy.go @@ -589,6 +589,21 @@ func (in *ControlPlaneStatus) DeepCopy() *ControlPlaneStatus { return out } +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *CreatorReference) DeepCopyInto(out *CreatorReference) { + *out = *in +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new CreatorReference. +func (in *CreatorReference) DeepCopy() *CreatorReference { + if in == nil { + return nil + } + out := new(CreatorReference) + in.DeepCopyInto(out) + return out +} + // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *DataRealignmentSettings) DeepCopyInto(out *DataRealignmentSettings) { *out = *in @@ -942,6 +957,99 @@ func (in *MigrateSpec) DeepCopy() *MigrateSpec { return out } +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *MigrateVolumeSpec) DeepCopyInto(out *MigrateVolumeSpec) { + *out = *in + out.TargetNodeRef = in.TargetNodeRef +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new MigrateVolumeSpec. +func (in *MigrateVolumeSpec) DeepCopy() *MigrateVolumeSpec { + if in == nil { + return nil + } + out := new(MigrateVolumeSpec) + in.DeepCopyInto(out) + return out +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *MigrationConnection) DeepCopyInto(out *MigrationConnection) { + *out = *in + if in.Port != nil { + in, out := &in.Port, &out.Port + *out = new(int32) + **out = **in + } + if in.NrIOQueues != nil { + in, out := &in.NrIOQueues, &out.NrIOQueues + *out = new(int32) + **out = **in + } + if in.ReconnectDelaySeconds != nil { + in, out := &in.ReconnectDelaySeconds, &out.ReconnectDelaySeconds + *out = new(int32) + **out = **in + } + if in.CtrlLossTimeoutSeconds != nil { + in, out := &in.CtrlLossTimeoutSeconds, &out.CtrlLossTimeoutSeconds + *out = new(int32) + **out = **in + } + if in.FastIOFailTimeoutSeconds != nil { + in, out := &in.FastIOFailTimeoutSeconds, &out.FastIOFailTimeoutSeconds + *out = new(int32) + **out = **in + } + if in.KeepAliveTimeoutSeconds != nil { + in, out := &in.KeepAliveTimeoutSeconds, &out.KeepAliveTimeoutSeconds + *out = new(int32) + **out = **in + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new MigrationConnection. +func (in *MigrationConnection) DeepCopy() *MigrationConnection { + if in == nil { + return nil + } + out := new(MigrationConnection) + in.DeepCopyInto(out) + return out +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *MigrationStatus) DeepCopyInto(out *MigrationStatus) { + *out = *in + if in.MemberCount != nil { + in, out := &in.MemberCount, &out.MemberCount + *out = new(int32) + **out = **in + } + if in.Connections != nil { + in, out := &in.Connections, &out.Connections + *out = make([]MigrationConnection, len(*in)) + for i := range *in { + (*in)[i].DeepCopyInto(&(*out)[i]) + } + } + if in.ValidationJobs != nil { + in, out := &in.ValidationJobs, &out.ValidationJobs + *out = make([]ValidationJob, len(*in)) + copy(*out, *in) + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new MigrationStatus. +func (in *MigrationStatus) DeepCopy() *MigrationStatus { + if in == nil { + return nil + } + out := new(MigrationStatus) + in.DeepCopyInto(out) + return out +} + // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *NodeGroup) DeepCopyInto(out *NodeGroup) { *out = *in @@ -1142,6 +1250,123 @@ func (in *OperatorOpsStatus) DeepCopy() *OperatorOpsStatus { return out } +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *PersistentVolumeOps) DeepCopyInto(out *PersistentVolumeOps) { + *out = *in + out.TypeMeta = in.TypeMeta + in.ObjectMeta.DeepCopyInto(&out.ObjectMeta) + in.Spec.DeepCopyInto(&out.Spec) + in.Status.DeepCopyInto(&out.Status) +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new PersistentVolumeOps. +func (in *PersistentVolumeOps) DeepCopy() *PersistentVolumeOps { + if in == nil { + return nil + } + out := new(PersistentVolumeOps) + in.DeepCopyInto(out) + return out +} + +// DeepCopyObject is an autogenerated deepcopy function, copying the receiver, creating a new runtime.Object. +func (in *PersistentVolumeOps) DeepCopyObject() runtime.Object { + if c := in.DeepCopy(); c != nil { + return c + } + return nil +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *PersistentVolumeOpsList) DeepCopyInto(out *PersistentVolumeOpsList) { + *out = *in + out.TypeMeta = in.TypeMeta + in.ListMeta.DeepCopyInto(&out.ListMeta) + if in.Items != nil { + in, out := &in.Items, &out.Items + *out = make([]PersistentVolumeOps, len(*in)) + for i := range *in { + (*in)[i].DeepCopyInto(&(*out)[i]) + } + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new PersistentVolumeOpsList. +func (in *PersistentVolumeOpsList) DeepCopy() *PersistentVolumeOpsList { + if in == nil { + return nil + } + out := new(PersistentVolumeOpsList) + in.DeepCopyInto(out) + return out +} + +// DeepCopyObject is an autogenerated deepcopy function, copying the receiver, creating a new runtime.Object. +func (in *PersistentVolumeOpsList) DeepCopyObject() runtime.Object { + if c := in.DeepCopy(); c != nil { + return c + } + return nil +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *PersistentVolumeOpsSpec) DeepCopyInto(out *PersistentVolumeOpsSpec) { + *out = *in + if in.Migrate != nil { + in, out := &in.Migrate, &out.Migrate + *out = new(MigrateVolumeSpec) + **out = **in + } + if in.CreatorRef != nil { + in, out := &in.CreatorRef, &out.CreatorRef + *out = new(CreatorReference) + **out = **in + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new PersistentVolumeOpsSpec. +func (in *PersistentVolumeOpsSpec) DeepCopy() *PersistentVolumeOpsSpec { + if in == nil { + return nil + } + out := new(PersistentVolumeOpsSpec) + in.DeepCopyInto(out) + return out +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *PersistentVolumeOpsStatus) DeepCopyInto(out *PersistentVolumeOpsStatus) { + *out = *in + in.Step.DeepCopyInto(&out.Step) + if in.Migration != nil { + in, out := &in.Migration, &out.Migration + *out = new(MigrationStatus) + (*in).DeepCopyInto(*out) + } + if in.DeferredSince != nil { + in, out := &in.DeferredSince, &out.DeferredSince + *out = (*in).DeepCopy() + } + if in.StartedAt != nil { + in, out := &in.StartedAt, &out.StartedAt + *out = (*in).DeepCopy() + } + if in.CompletedAt != nil { + in, out := &in.CompletedAt, &out.CompletedAt + *out = (*in).DeepCopy() + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new PersistentVolumeOpsStatus. +func (in *PersistentVolumeOpsStatus) DeepCopy() *PersistentVolumeOpsStatus { + if in == nil { + return nil + } + out := new(PersistentVolumeOpsStatus) + in.DeepCopyInto(out) + return out +} + // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *PoolLimits) DeepCopyInto(out *PoolLimits) { *out = *in @@ -2527,6 +2752,21 @@ func (in *StorageNodePorts) DeepCopy() *StorageNodePorts { return out } +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *StorageNodeReference) DeepCopyInto(out *StorageNodeReference) { + *out = *in +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new StorageNodeReference. +func (in *StorageNodeReference) DeepCopy() *StorageNodeReference { + if in == nil { + return nil + } + out := new(StorageNodeReference) + in.DeepCopyInto(out) + return out +} + // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *StorageNodeResources) DeepCopyInto(out *StorageNodeResources) { *out = *in @@ -3000,6 +3240,21 @@ func (in *UpgradeSpec) DeepCopy() *UpgradeSpec { return out } +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *ValidationJob) DeepCopyInto(out *ValidationJob) { + *out = *in +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new ValidationJob. +func (in *ValidationJob) DeepCopy() *ValidationJob { + if in == nil { + return nil + } + out := new(ValidationJob) + in.DeepCopyInto(out) + return out +} + // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *VaultKMS) DeepCopyInto(out *VaultKMS) { *out = *in diff --git a/operator/config/crd/bases/storage.simplyblock.io_persistentvolumeops.yaml b/operator/config/crd/bases/storage.simplyblock.io_persistentvolumeops.yaml new file mode 100644 index 000000000..49cc97743 --- /dev/null +++ b/operator/config/crd/bases/storage.simplyblock.io_persistentvolumeops.yaml @@ -0,0 +1,383 @@ +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition +metadata: + annotations: + controller-gen.kubebuilder.io/version: v0.21.0 + name: persistentvolumeops.storage.simplyblock.io +spec: + group: storage.simplyblock.io + names: + kind: PersistentVolumeOps + listKind: PersistentVolumeOpsList + plural: persistentvolumeops + shortNames: + - pvops + singular: persistentvolumeops + scope: Cluster + versions: + - additionalPrinterColumns: + - jsonPath: .spec.persistentVolumeName + name: Volume + type: string + - jsonPath: .spec.action + name: Action + type: string + - jsonPath: .spec.migrate.targetNodeRef.name + name: Target + type: string + - jsonPath: .status.phase + name: Phase + type: string + - jsonPath: .status.step.state + name: Step + type: string + - jsonPath: .status.message + name: Message + priority: 1 + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha2 + schema: + openAPIV3Schema: + description: |- + PersistentVolumeOps is a single operation performed against one + PersistentVolume. It is the one Ops kind in this group whose target is a core + Kubernetes type rather than a kind this group defines, so it locks its target + with an annotation rather than a status field, is cluster-scoped because its + target is, cannot be owned by the namespaced operation that created it, and + derives its cluster, pool, and volume from the volume's CSI handle rather + than being told. + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: |- + PersistentVolumeOpsSpec is one operation to perform against one + PersistentVolume. + + The rule keeps the action and its parameter block in agreement, which is a + statement about this object alone and so belongs on the type rather than in + the webhook. + properties: + abort: + description: |- + Abort asks a running operation to stop at its next step and unwind. It is + expressible from Validating and Migrating and not from Verifying, which + the action's graph declares rather than this field: once the copy has + finished, the volume has moved and there is nothing to undo. + type: boolean + action: + description: |- + Action is the operation to perform. Immutable: an operation that changed + what it was doing halfway through would have a status describing neither. + enum: + - Migrate + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + creatorRef: + description: |- + CreatorRef names the object that created this one, and that object's + finalizer is what aborts and deletes this one when it goes. Absent on an + operation written by hand, which has no creator to cascade from. + Immutable, because an operation changing whose fan-out it belongs to + would change who cascades over it. + properties: + kind: + description: Kind is the creating object's kind, which is StorageNodeOps + for a drain. + type: string + name: + description: Name is the creator's object name. + type: string + namespace: + description: Namespace is where the creator lives. + type: string + uid: + description: |- + UID is what makes the reference address one creator rather than one name: + a creator deleted and recreated under the same name must not inherit the + fan-out it did not issue, and cascade a delete over it. + type: string + required: + - kind + - name + - namespace + - uid + type: object + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + migrate: + description: Migrate parameterizes action Migrate. + properties: + targetNodeRef: + description: |- + TargetNodeRef locates the StorageNode to move the volume's backing + logical volume to. It names the Kubernetes object rather than the backend + UUID, so that a migration can be written by hand without looking one up; + the controller resolves the UUID from the node's status. The node's + cluster must be the volume's, which the webhook checks rather than the + type, because that is a fact about two other objects. + + The marker sits on the reference rather than on its fields, so the pair + is immutable together and a target cannot be half-changed into a name in + one cluster and a namespace in another. + properties: + name: + description: Name is the StorageNode object's name. + type: string + namespace: + description: |- + Namespace is where the StorageNode lives, which is the namespace of the + StorageCluster that owns it. + type: string + required: + - name + - namespace + type: object + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + required: + - targetNodeRef + type: object + persistentVolumeName: + description: |- + PersistentVolumeName names the PersistentVolume this operation acts on. + It is a name rather than a reference because a PersistentVolume is + cluster-scoped, and it names the volume rather than the claim because a + claim can be deleted while its volume is retained — under a Retain + reclaim policy the logical volume still occupies capacity on a node that + may be draining, and is still worth moving. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + required: + - action + - persistentVolumeName + type: object + x-kubernetes-validations: + - message: field creatorRef is immutable once set + rule: '!has(oldSelf.creatorRef) || has(self.creatorRef)' + - message: migrate is required for action Migrate and must be absent otherwise + rule: 'self.action == ''Migrate'' ? has(self.migrate) : !has(self.migrate)' + status: + description: PersistentVolumeOpsStatus is the observed state of one volume + operation. + properties: + completedAt: + description: CompletedAt is when it reached a terminal phase. + format: date-time + type: string + deferredSince: + description: |- + DeferredSince is when the operation was first held — behind another + operation's lock, or behind a control plane that is not accepting + migrations yet. It is what the auto-rebalancer reads to decide whether a + migration has waited long enough to give up on, and it is in status + rather than in memory because the operator may restart and an observer + needs to see that the operation is waiting and since when. + format: date-time + type: string + message: + description: |- + Message is the reason the phase is what it is: one sentence, replaced as + the operation moves, and never a log. + type: string + migration: + description: |- + Migration is everything about the migration rather than about the + operation. + properties: + clusterUUID: + description: |- + ClusterUUID, PoolUUID, and VolumeUUID are the three parts of the volume's + CSI volume handle, recorded so that later steps address the backend + without re-reading the PersistentVolume, and so that a failed operation + says which volume it was working on after the volume is gone. + type: string + connections: + description: |- + Connections are the paths the migration published on the target. + Verifying confirms none of them is left connected. + items: + description: |- + MigrationConnection is one NVMe-oF path the migration published on the + target. + + The connect parameters travel with the address because the path is connected + on a consuming host rather than here, and what is recorded has to be the + connect that will actually be made: the host attaches every path with the + same controller-loss timeout the CSI driver uses, which is not the hour the + control plane answers with, and a record of the control plane's answer would + describe a connect nobody performs. + properties: + address: + description: Address and Port are where it answers. + type: string + ctrlLossTimeoutSeconds: + description: |- + CtrlLossTimeoutSeconds and FastIOFailTimeoutSeconds are pointers because + zero is a choice ("fail I/O immediately") rather than a missing value. + format: int32 + type: integer + fastIOFailTimeoutSeconds: + format: int32 + type: integer + keepAliveTimeoutSeconds: + format: int32 + type: integer + nqn: + description: NQN is the subsystem the target answers this + path for. + type: string + nrIOQueues: + format: int32 + type: integer + port: + format: int32 + type: integer + reconnectDelaySeconds: + format: int32 + type: integer + transport: + description: Transport is the fabric, which is tcp. + type: string + type: object + type: array + memberCount: + description: |- + MemberCount is how many volumes the migrated NVMe-oF subsystem holds, as + the control plane reports it. More than one member means the sibling + volumes move along with the named one, so the count is both the + operation's blast radius and the term the copy's deadline scales by. + format: int32 + minimum: 0 + type: integer + migrationUUID: + description: MigrationUUID is the control plane's identifier for + the copy. + type: string + poolUUID: + type: string + sourceNodeUUID: + description: |- + SourceNodeUUID is where the volume was before the move, recorded so that + a failure says what it was and not only what it was going to be. + type: string + subsystemNQN: + description: |- + SubsystemNQN is the volume's NVMe-oF subsystem, which is what the + migration is addressed by: the control plane migrates a subsystem rather + than one volume inside it. + type: string + targetNodeUUID: + description: |- + TargetNodeUUID is the backend identifier resolved from + spec.migrate.targetNodeRef, recorded so the later steps address the + target without resolving the node object again. + type: string + validationJobs: + description: ValidationJobs are the Jobs started to check those + paths. + items: + description: |- + ValidationJob is one Job started to check a path is reachable from one node. + + It is tracked so that Verifying can delete every Job the operation started, + including after a restart. The namespace is recorded for the same reason + spec.migrate.targetNodeRef carries one: a Job is namespaced and this object + is not, so a name alone would not locate it. + properties: + name: + type: string + namespace: + type: string + node: + description: |- + Node is the worker the Job is pinned to, which is a node consuming one of + the migrated subsystem's volumes. + type: string + succeeded: + description: |- + Succeeded records a node whose paths were checked and found ready, so a + restart does not run the check again on a node that already passed and + whose Job its own TTL may already have reaped. + type: boolean + required: + - name + - namespace + type: object + type: array + volumeUUID: + type: string + type: object + observedGeneration: + description: |- + ObservedGeneration is the generation the rest of this status was computed + from, so a stale status can be told from a current one. + format: int64 + type: integer + phase: + description: Phase is the operation's own progress. + enum: + - Pending + - Running + - Succeeded + - Failed + - Aborted + type: string + startedAt: + description: StartedAt is when the operation began. + format: date-time + type: string + step: + description: |- + Step is the position of the running action's state machine, as the shared + statemachine.KubeSnapshot. It is persisted before the side effect that + step performs. The rule is what an Enum marker would do if a marker could + reach a field of a shared type. + properties: + deadline: + description: |- + Deadline is when that state expires, absent when it has none. It is an + absolute instant, so a state whose deadline passed while the controller + was down restores as already expired. + format: date-time + type: string + state: + description: |- + State is the state the machine was in. Empty means the resource has not + been reconciled yet, and restores to the graph's initial state. + type: string + type: object + x-kubernetes-validations: + - message: unknown step + rule: '!has(self.state) || self.state in [''Validating'',''Migrating'',''Verifying'']' + type: object + type: object + served: true + storage: true + subresources: + status: {} diff --git a/operator/config/crd/kustomization.yaml b/operator/config/crd/kustomization.yaml index f2c76ca6c..47e50a8da 100644 --- a/operator/config/crd/kustomization.yaml +++ b/operator/config/crd/kustomization.yaml @@ -26,6 +26,7 @@ resources: - bases/storage.simplyblock.io_operatorops.yaml - bases/storage.simplyblock.io_simplyblockdrivers.yaml - bases/storage.simplyblock.io_storagepoolops.yaml +- bases/storage.simplyblock.io_persistentvolumeops.yaml # +kubebuilder:scaffold:crdkustomizeresource patches: [] diff --git a/operator/config/samples/kustomization.yaml b/operator/config/samples/kustomization.yaml index f1b29ee12..ac8048eb4 100644 --- a/operator/config/samples/kustomization.yaml +++ b/operator/config/samples/kustomization.yaml @@ -11,4 +11,5 @@ resources: - storage_v1alpha2_storagebackupops.yaml - storage_v1alpha1_backuprestore.yaml - storage_v1alpha2_operatorops.yaml +- storage_v1alpha2_persistentvolumeops.yaml # +kubebuilder:scaffold:manifestskustomizesamples diff --git a/operator/config/samples/storage_v1alpha2_persistentvolumeops.yaml b/operator/config/samples/storage_v1alpha2_persistentvolumeops.yaml new file mode 100644 index 000000000..052771ef3 --- /dev/null +++ b/operator/config/samples/storage_v1alpha2_persistentvolumeops.yaml @@ -0,0 +1,21 @@ +# A PersistentVolumeOps moving one volume to another storage node. +# +# It carries no namespace, because a PersistentVolume is cluster-scoped and so +# is an operation on one. The target is named as a StorageNode object rather +# than as a backend UUID, so a migration can be written by hand without looking +# one up, and it carries a namespace because a bare name has none to mean from +# here. +apiVersion: storage.simplyblock.io/v1alpha2 +kind: PersistentVolumeOps +metadata: + labels: + app.kubernetes.io/name: simplyblock-operator + app.kubernetes.io/managed-by: kustomize + name: persistentvolumeops-sample +spec: + persistentVolumeName: pvc-00000000-0000-0000-0000-000000000000 + action: Migrate + migrate: + targetNodeRef: + namespace: simplyblock + name: storagenode-sample diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index a0445363d..7fdd3b8e8 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -1990,6 +1990,389 @@ spec: --- apiVersion: apiextensions.k8s.io/v1 kind: CustomResourceDefinition +metadata: + annotations: + controller-gen.kubebuilder.io/version: v0.21.0 + name: persistentvolumeops.storage.simplyblock.io +spec: + group: storage.simplyblock.io + names: + kind: PersistentVolumeOps + listKind: PersistentVolumeOpsList + plural: persistentvolumeops + shortNames: + - pvops + singular: persistentvolumeops + scope: Cluster + versions: + - additionalPrinterColumns: + - jsonPath: .spec.persistentVolumeName + name: Volume + type: string + - jsonPath: .spec.action + name: Action + type: string + - jsonPath: .spec.migrate.targetNodeRef.name + name: Target + type: string + - jsonPath: .status.phase + name: Phase + type: string + - jsonPath: .status.step.state + name: Step + type: string + - jsonPath: .status.message + name: Message + priority: 1 + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha2 + schema: + openAPIV3Schema: + description: |- + PersistentVolumeOps is a single operation performed against one + PersistentVolume. It is the one Ops kind in this group whose target is a core + Kubernetes type rather than a kind this group defines, so it locks its target + with an annotation rather than a status field, is cluster-scoped because its + target is, cannot be owned by the namespaced operation that created it, and + derives its cluster, pool, and volume from the volume's CSI handle rather + than being told. + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: |- + PersistentVolumeOpsSpec is one operation to perform against one + PersistentVolume. + + The rule keeps the action and its parameter block in agreement, which is a + statement about this object alone and so belongs on the type rather than in + the webhook. + properties: + abort: + description: |- + Abort asks a running operation to stop at its next step and unwind. It is + expressible from Validating and Migrating and not from Verifying, which + the action's graph declares rather than this field: once the copy has + finished, the volume has moved and there is nothing to undo. + type: boolean + action: + description: |- + Action is the operation to perform. Immutable: an operation that changed + what it was doing halfway through would have a status describing neither. + enum: + - Migrate + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + creatorRef: + description: |- + CreatorRef names the object that created this one, and that object's + finalizer is what aborts and deletes this one when it goes. Absent on an + operation written by hand, which has no creator to cascade from. + Immutable, because an operation changing whose fan-out it belongs to + would change who cascades over it. + properties: + kind: + description: Kind is the creating object's kind, which is StorageNodeOps + for a drain. + type: string + name: + description: Name is the creator's object name. + type: string + namespace: + description: Namespace is where the creator lives. + type: string + uid: + description: |- + UID is what makes the reference address one creator rather than one name: + a creator deleted and recreated under the same name must not inherit the + fan-out it did not issue, and cascade a delete over it. + type: string + required: + - kind + - name + - namespace + - uid + type: object + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + migrate: + description: Migrate parameterizes action Migrate. + properties: + targetNodeRef: + description: |- + TargetNodeRef locates the StorageNode to move the volume's backing + logical volume to. It names the Kubernetes object rather than the backend + UUID, so that a migration can be written by hand without looking one up; + the controller resolves the UUID from the node's status. The node's + cluster must be the volume's, which the webhook checks rather than the + type, because that is a fact about two other objects. + + The marker sits on the reference rather than on its fields, so the pair + is immutable together and a target cannot be half-changed into a name in + one cluster and a namespace in another. + properties: + name: + description: Name is the StorageNode object's name. + type: string + namespace: + description: |- + Namespace is where the StorageNode lives, which is the namespace of the + StorageCluster that owns it. + type: string + required: + - name + - namespace + type: object + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + required: + - targetNodeRef + type: object + persistentVolumeName: + description: |- + PersistentVolumeName names the PersistentVolume this operation acts on. + It is a name rather than a reference because a PersistentVolume is + cluster-scoped, and it names the volume rather than the claim because a + claim can be deleted while its volume is retained — under a Retain + reclaim policy the logical volume still occupies capacity on a node that + may be draining, and is still worth moving. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + required: + - action + - persistentVolumeName + type: object + x-kubernetes-validations: + - message: field creatorRef is immutable once set + rule: '!has(oldSelf.creatorRef) || has(self.creatorRef)' + - message: migrate is required for action Migrate and must be absent otherwise + rule: 'self.action == ''Migrate'' ? has(self.migrate) : !has(self.migrate)' + status: + description: PersistentVolumeOpsStatus is the observed state of one volume + operation. + properties: + completedAt: + description: CompletedAt is when it reached a terminal phase. + format: date-time + type: string + deferredSince: + description: |- + DeferredSince is when the operation was first held — behind another + operation's lock, or behind a control plane that is not accepting + migrations yet. It is what the auto-rebalancer reads to decide whether a + migration has waited long enough to give up on, and it is in status + rather than in memory because the operator may restart and an observer + needs to see that the operation is waiting and since when. + format: date-time + type: string + message: + description: |- + Message is the reason the phase is what it is: one sentence, replaced as + the operation moves, and never a log. + type: string + migration: + description: |- + Migration is everything about the migration rather than about the + operation. + properties: + clusterUUID: + description: |- + ClusterUUID, PoolUUID, and VolumeUUID are the three parts of the volume's + CSI volume handle, recorded so that later steps address the backend + without re-reading the PersistentVolume, and so that a failed operation + says which volume it was working on after the volume is gone. + type: string + connections: + description: |- + Connections are the paths the migration published on the target. + Verifying confirms none of them is left connected. + items: + description: |- + MigrationConnection is one NVMe-oF path the migration published on the + target. + + The connect parameters travel with the address because the path is connected + on a consuming host rather than here, and what is recorded has to be the + connect that will actually be made: the host attaches every path with the + same controller-loss timeout the CSI driver uses, which is not the hour the + control plane answers with, and a record of the control plane's answer would + describe a connect nobody performs. + properties: + address: + description: Address and Port are where it answers. + type: string + ctrlLossTimeoutSeconds: + description: |- + CtrlLossTimeoutSeconds and FastIOFailTimeoutSeconds are pointers because + zero is a choice ("fail I/O immediately") rather than a missing value. + format: int32 + type: integer + fastIOFailTimeoutSeconds: + format: int32 + type: integer + keepAliveTimeoutSeconds: + format: int32 + type: integer + nqn: + description: NQN is the subsystem the target answers this + path for. + type: string + nrIOQueues: + format: int32 + type: integer + port: + format: int32 + type: integer + reconnectDelaySeconds: + format: int32 + type: integer + transport: + description: Transport is the fabric, which is tcp. + type: string + type: object + type: array + memberCount: + description: |- + MemberCount is how many volumes the migrated NVMe-oF subsystem holds, as + the control plane reports it. More than one member means the sibling + volumes move along with the named one, so the count is both the + operation's blast radius and the term the copy's deadline scales by. + format: int32 + minimum: 0 + type: integer + migrationUUID: + description: MigrationUUID is the control plane's identifier for + the copy. + type: string + poolUUID: + type: string + sourceNodeUUID: + description: |- + SourceNodeUUID is where the volume was before the move, recorded so that + a failure says what it was and not only what it was going to be. + type: string + subsystemNQN: + description: |- + SubsystemNQN is the volume's NVMe-oF subsystem, which is what the + migration is addressed by: the control plane migrates a subsystem rather + than one volume inside it. + type: string + targetNodeUUID: + description: |- + TargetNodeUUID is the backend identifier resolved from + spec.migrate.targetNodeRef, recorded so the later steps address the + target without resolving the node object again. + type: string + validationJobs: + description: ValidationJobs are the Jobs started to check those + paths. + items: + description: |- + ValidationJob is one Job started to check a path is reachable from one node. + + It is tracked so that Verifying can delete every Job the operation started, + including after a restart. The namespace is recorded for the same reason + spec.migrate.targetNodeRef carries one: a Job is namespaced and this object + is not, so a name alone would not locate it. + properties: + name: + type: string + namespace: + type: string + node: + description: |- + Node is the worker the Job is pinned to, which is a node consuming one of + the migrated subsystem's volumes. + type: string + succeeded: + description: |- + Succeeded records a node whose paths were checked and found ready, so a + restart does not run the check again on a node that already passed and + whose Job its own TTL may already have reaped. + type: boolean + required: + - name + - namespace + type: object + type: array + volumeUUID: + type: string + type: object + observedGeneration: + description: |- + ObservedGeneration is the generation the rest of this status was computed + from, so a stale status can be told from a current one. + format: int64 + type: integer + phase: + description: Phase is the operation's own progress. + enum: + - Pending + - Running + - Succeeded + - Failed + - Aborted + type: string + startedAt: + description: StartedAt is when the operation began. + format: date-time + type: string + step: + description: |- + Step is the position of the running action's state machine, as the shared + statemachine.KubeSnapshot. It is persisted before the side effect that + step performs. The rule is what an Enum marker would do if a marker could + reach a field of a shared type. + properties: + deadline: + description: |- + Deadline is when that state expires, absent when it has none. It is an + absolute instant, so a state whose deadline passed while the controller + was down restores as already expired. + format: date-time + type: string + state: + description: |- + State is the state the machine was in. Empty means the resource has not + been reconciled yet, and restores to the graph's initial state. + type: string + type: object + x-kubernetes-validations: + - message: unknown step + rule: '!has(self.state) || self.state in [''Validating'',''Migrating'',''Verifying'']' + type: object + type: object + served: true + storage: true + subresources: + status: {} +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition metadata: annotations: controller-gen.kubebuilder.io/version: v0.21.0 diff --git a/operator/internal/controllers/volume/cel_validation_test.go b/operator/internal/controllers/volume/cel_validation_test.go new file mode 100644 index 000000000..69d433658 --- /dev/null +++ b/operator/internal/controllers/volume/cel_validation_test.go @@ -0,0 +1,190 @@ +// Validation of the CEL rules compiled into the PersistentVolumeOps CRD schema, +// run against a real apiserver. +// +// It lives here rather than under internal/webhook because there is no webhook +// involved: the rules are enforced by the apiserver itself, and envtest is the +// only place in the tree that starts one. The suite installs CRDs and nothing +// else, so a rejection here can only have come from the schema. +// +// What the schema owns and the webhook does not is every rule that is a +// statement about this object's own fields: the agreement between spec.action +// and its parameter block, and the immutability of everything but spec.abort. +// The webhook is left with the rows that are facts about a different object — +// the volume, the target node — which CEL cannot reach +// (design-persistentvolumeops.md §4.3). + +package volume + +import ( + "context" + "strings" + "testing" + + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/types" + "sigs.k8s.io/controller-runtime/pkg/client" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// migrateOperation is a well-formed operation, which every case below starts +// from and breaks in exactly one way. +func migrateOperation() *simplyblockv1alpha2.PersistentVolumeOps { + return &simplyblockv1alpha2.PersistentVolumeOps{ + ObjectMeta: metav1.ObjectMeta{GenerateName: "cel-"}, + Spec: simplyblockv1alpha2.PersistentVolumeOpsSpec{ + PersistentVolumeName: "pvc-0001", + Action: simplyblockv1alpha2.PersistentVolumeOpsActionMigrate, + Migrate: &simplyblockv1alpha2.MigrateVolumeSpec{ + TargetNodeRef: simplyblockv1alpha2.StorageNodeReference{ + Namespace: "simplyblock", + Name: "worker-3", + }, + }, + }, + } +} + +// TestPersistentVolumeOpsCELRequiresTheActionsParameterBlock. An action with no +// parameters is an operation with no target node, which would reach the +// controller and fail there, one reconcile later and with a worse message. +// Every Ops kind in the group states the pairing on the type for that reason +// (design-storagebackup.md uses the same rule). +func TestPersistentVolumeOpsCELRequiresTheActionsParameterBlock(t *testing.T) { + apiClient := apiServer(t) + + t.Run("Migrate with its block is accepted", func(t *testing.T) { + ops := migrateOperation() + if err := apiClient.Create(context.Background(), ops); err != nil { + t.Fatalf("a well-formed operation was rejected: %v", err) + } + t.Cleanup(func() { _ = apiClient.Delete(context.Background(), ops) }) + }) + + t.Run("Migrate without its block is rejected", func(t *testing.T) { + ops := migrateOperation() + ops.Spec.Migrate = nil + + err := apiClient.Create(context.Background(), ops) + if err == nil { + t.Fatal("the apiserver accepted a Migrate with no target node") + } + if !strings.Contains(err.Error(), "migrate is required for action Migrate") { + t.Fatalf("rejected for the wrong reason: %v", err) + } + }) +} + +// TestPersistentVolumeOpsCELFreezesEverythingButAbort. An operation that +// changed what it was doing halfway through would have a status describing +// neither, and the one field a user is meant to change after the fact is the +// request to stop. +func TestPersistentVolumeOpsCELFreezesEverythingButAbort(t *testing.T) { + apiClient := apiServer(t) + + for _, tc := range []struct { + name string + edit func(*simplyblockv1alpha2.PersistentVolumeOps) + wantDenied bool + }{ + { + name: "abort is the one mutable field", + edit: func(ops *simplyblockv1alpha2.PersistentVolumeOps) { ops.Spec.Abort = true }, + }, + { + name: "the volume cannot be repointed", + edit: func(ops *simplyblockv1alpha2.PersistentVolumeOps) { ops.Spec.PersistentVolumeName = "pvc-0002" }, + wantDenied: true, + }, + { + name: "the target node cannot be repointed", + edit: func(ops *simplyblockv1alpha2.PersistentVolumeOps) { + ops.Spec.Migrate.TargetNodeRef.Name = "worker-4" + }, + wantDenied: true, + }, + { + // The marker sits on the reference rather than on its fields, so + // the pair moves together or not at all: a target half-changed + // would name a node in one cluster and a namespace in another. + name: "the target node's namespace cannot be repointed either", + edit: func(ops *simplyblockv1alpha2.PersistentVolumeOps) { + ops.Spec.Migrate.TargetNodeRef.Namespace = "elsewhere" + }, + wantDenied: true, + }, + } { + t.Run(tc.name, func(t *testing.T) { + ctx := context.Background() + ops := migrateOperation() + if err := apiClient.Create(ctx, ops); err != nil { + t.Fatalf("creating the operation: %v", err) + } + t.Cleanup(func() { _ = apiClient.Delete(ctx, ops) }) + + tc.edit(ops) + err := apiClient.Update(ctx, ops) + if tc.wantDenied { + if err == nil { + t.Fatal("the apiserver accepted an edit to a frozen field") + } + if !strings.Contains(err.Error(), "immutable") { + t.Fatalf("rejected for the wrong reason: %v", err) + } + return + } + if err != nil { + t.Fatalf("the apiserver rejected the one field a user may change: %v", err) + } + }) + } +} + +// TestPersistentVolumeOpsCELRejectsAnUndeclaredStep. status.step carries the +// shared statemachine snapshot, whose state field no Enum marker can reach, so +// the rule on the status is what an Enum would have done. Without it a +// hand-edited or downgraded object could name a step this kind has no graph +// for, and the controller would fail to resume it one reconcile later. +func TestPersistentVolumeOpsCELRejectsAnUndeclaredStep(t *testing.T) { + apiClient := apiServer(t) + ctx := context.Background() + + ops := migrateOperation() + if err := apiClient.Create(ctx, ops); err != nil { + t.Fatalf("creating the operation: %v", err) + } + t.Cleanup(func() { _ = apiClient.Delete(ctx, ops) }) + + for _, tc := range []struct { + step string + wantDenied bool + }{ + {step: "Validating"}, + {step: "Migrating"}, + {step: "Verifying"}, + {step: "Suspending", wantDenied: true}, + } { + t.Run(tc.step, func(t *testing.T) { + fresh := &simplyblockv1alpha2.PersistentVolumeOps{} + if err := apiClient.Get(ctx, types.NamespacedName{Name: ops.Name}, fresh); err != nil { + t.Fatalf("reading the operation back: %v", err) + } + base := fresh.DeepCopy() + fresh.Status.Step.State = tc.step + + err := apiClient.Status().Patch(ctx, fresh, client.MergeFrom(base)) + if tc.wantDenied { + if err == nil { + t.Fatalf("the apiserver accepted step %q, which this kind has no graph for", tc.step) + } + if !strings.Contains(err.Error(), "unknown step") { + t.Fatalf("rejected for the wrong reason: %v", err) + } + return + } + if err != nil { + t.Fatalf("the apiserver rejected declared step %q: %v", tc.step, err) + } + }) + } +} diff --git a/operator/internal/controllers/volume/graphs.go b/operator/internal/controllers/volume/graphs.go new file mode 100644 index 000000000..d6cca0f8b --- /dev/null +++ b/operator/internal/controllers/volume/graphs.go @@ -0,0 +1,141 @@ +// The state graph of the PersistentVolumeOps action, declared as data. +// +// It is a MultiConfig although the kind carries one action, and not because a +// second one is coming: every other Ops controller in this group drives its +// steps through a MultiConfig keyed by action, and a kind that read differently +// for having one entry would make a reader check whether the difference meant +// something. A MultiConfig with one entry costs nothing. +// +// The graph is built per operation rather than shared, because one of its +// bounds is computed. The copy's deadline scales with how many volumes the +// migrated subsystem holds, which is only known once the migration has been +// created, so the hook closes over that count — the shape atlas-lib's +// statemachine documents for exactly this case. +// +// design-persistentvolumeops.md §5 is the specification. + +package volume + +import ( + "context" + "time" + + "github.com/simplyblock/atlas/statemachine" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// step is the operation's step type, aliased so the graph literal below reads +// as the graph rather than as a wall of package qualifiers. +type step = simplyblockv1alpha2.PersistentVolumeOpsStep + +const ( + stepValidating = simplyblockv1alpha2.PersistentVolumeOpsStepValidating + stepMigrating = simplyblockv1alpha2.PersistentVolumeOpsStepMigrating + stepVerifying = simplyblockv1alpha2.PersistentVolumeOpsStepVerifying +) + +// actionMigrate is the MultiConfig key for the one action this kind carries. +const actionMigrate = statemachine.Action(simplyblockv1alpha2.PersistentVolumeOpsActionMigrate) + +// How long each step may take before the operation is reported as stuck. +// +// They differ by what the step is waiting on rather than by preference. +const ( + // validatingDeadline bounds everything between creating the migration and + // continuing it: waiting for the consuming pods to be Running, and running + // one Job per consuming node. It is finite because the control plane does + // not hold a created-but-unstarted migration open indefinitely, and an + // operation that sat in this step past that window would continue a + // migration the backend has already given up on. + validatingDeadline = 15 * time.Minute + + // copyBaseDeadline is what a migration of one volume gets, and the floor + // under every larger one. + copyBaseDeadline = 30 * time.Minute + + // copyPerMemberDeadline is added for each volume the subsystem holds. A + // migration is addressed by the subsystem, so every sibling volume moves + // along with the named one and each of them is data to copy. + // + // There is no ceiling. A subsystem large enough to make this hours is one + // whose migration genuinely takes hours, and capping the bound would fail + // it for being big rather than for being stuck — which is the one thing a + // deadline here is for. + copyPerMemberDeadline = 10 * time.Minute + + // verifyingDeadline bounds the cleanup: deleting the validation Jobs and + // confirming that no path they connected is left behind. It is short + // because nothing in it waits on the data path, and it is bounded at all + // because a cleanup that cannot finish has to be visible — a path connected + // with nothing tracking it blocks every later migration of the volume, and + // has. + verifyingDeadline = 10 * time.Minute +) + +// deadline is the entry hook every state here carries: it sets the step's +// budget and performs nothing. The side effect of a step is performed on the +// pass that follows, against the step the entry's patch persisted, which is +// where the write-ahead record is needed and what it records. +func deadline(d time.Duration) statemachine.TransitionFunc[step] { + return func(context.Context, step, step) (time.Duration, error) { return d, nil } +} + +// copyDeadline is the bound on the copy, given how many volumes the subsystem +// holds. A count of zero is a migration whose creation has not reported one +// yet, and it gets the base bound rather than none: a step that cannot time out +// is the failure mode the bounds exist to prevent. +func copyDeadline(members int32) time.Duration { + if members < 1 { + members = 1 + } + return copyBaseDeadline + time.Duration(members)*copyPerMemberDeadline +} + +// graphs declares the state graph of each action, for a subsystem of the given +// member count. +func graphs(members int32) statemachine.MultiConfig[step] { + return statemachine.MultiConfig[step]{ + actionMigrate: { + Initial: stepValidating, + States: map[step]statemachine.StateDef[step]{ + // Validating and Migrating are abortable because both have + // something to take back: a backend migration that has not + // copied anything yet, and the paths its creation published on + // every consuming host. + stepValidating: { + To: []step{stepMigrating}, + Abortable: true, + OnEnter: deadline(validatingDeadline), + }, + stepMigrating: { + To: []step{stepVerifying}, + Abortable: true, + OnEnter: deadline(copyDeadline(members)), + }, + // Verifying is not. The copy has finished, the volume has + // moved, and what is left is the cleanup that makes the move + // safe — so stopping here would leave exactly the state this + // step exists to prevent. + stepVerifying: {OnEnter: deadline(verifyingDeadline)}, + }, + }, + } +} + +// UnabortableSteps are the declared steps an abort cannot be honored from, +// sorted. +// +// It is exported for the DELETE guard on this kind, which asks the same +// question this package's unwind asks: a deletion may not express something +// spec.abort could not, so both channels read one graph (design-crd-model.md +// §3.1). The guard has a step out of a status and no machine, which is the +// whole reason this reads the graph rather than the machine the reconciler +// holds. +// +// The member count is irrelevant to the answer, so the guard does not have to +// know one: which steps can be stopped is a property of the graph's shape, and +// the count only sets how long one of them may take. +func UnabortableSteps() []step { + return statemachine.UnabortableMultiStates(graphs(0)) +} diff --git a/operator/internal/controllers/volume/graphs_test.go b/operator/internal/controllers/volume/graphs_test.go new file mode 100644 index 000000000..4012108f9 --- /dev/null +++ b/operator/internal/controllers/volume/graphs_test.go @@ -0,0 +1,166 @@ +// The migration's graph as data: what it declares, and the three places that +// have to agree about it. +// +// A shared statemachine.KubeSnapshot cannot carry an Enum marker, so the step +// values live in three places — the graph, the kind's Enum marker, and the CEL +// rule on status.step — and nothing but a test makes them agree. +// +// The rest is what declaring the graph as data buys: the refusal to abort a +// Verifying operation is one assertion rather than a code path, and the copy's +// deadline, which is the one bound in this package that is computed rather than +// fixed, is exercised at both ends of its range. + +package volume + +import ( + "context" + "slices" + "strings" + "testing" + "time" + + "github.com/google/go-cmp/cmp" + + "github.com/simplyblock/atlas/statemachine" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// everyStep is the Enum marker's list, transcribed. It is written out rather +// than derived so that the assertion below compares two independent statements +// of the same set: deriving it from the graph would make the test agree with +// itself. +var everyStep = []string{"Migrating", "Validating", "Verifying"} + +func TestTheStepEnumCoversEveryDeclaredState(t *testing.T) { + declared := statemachine.DeclaredMultiStates(graphs(0)) + want := slices.Clone(everyStep) + slices.Sort(want) + if diff := cmp.Diff(want, declared); diff != "" { + t.Errorf("the graph and the Enum marker disagree (-marker +graph):\n%s", diff) + } +} + +// The CEL rule on status.step is what an Enum marker would do if a marker could +// reach a field of a type another module declares. It is a literal list in a +// struct tag, so nothing but this compares it against the graph — and it is +// compared both ways, because a rule that names a step no graph declares admits +// a status no controller can resume from. +func TestTheCELRuleCoversEveryDeclaredState(t *testing.T) { + declared := statemachine.DeclaredMultiStates(graphs(0)) + for _, state := range declared { + if !strings.Contains(stepCELRule, "'"+state+"'") { + t.Errorf("status.step's CEL rule does not accept the declared step %q", state) + } + } + for _, named := range celRuleValues(stepCELRule) { + if !slices.Contains(declared, named) { + t.Errorf("status.step's CEL rule accepts %q, which the graph does not declare", named) + } + } +} + +// celRuleValues reads the quoted values out of the rule's `in` list. +func celRuleValues(rule string) []string { + var values []string + for _, part := range strings.Split(rule, "'") { + if part != "" && !strings.ContainsAny(part, "[],| ") { + values = append(values, part) + } + } + return values +} + +// stepCELRule is the rule as the type declares it. Keeping a copy here is the +// cost of a rule living in a struct tag; the test above is what makes the copy +// worth having. +const stepCELRule = "!has(self.state) || self.state in ['Validating','Migrating','Verifying']" + +// Every action the API accepts needs a graph, or an operation of that action +// fails at its first pass with ErrUnknownAction rather than doing anything. +func TestEveryActionDeclaresAGraph(t *testing.T) { + declared := graphs(0) + for _, a := range []simplyblockv1alpha2.PersistentVolumeOpsAction{ + simplyblockv1alpha2.PersistentVolumeOpsActionMigrate, + } { + if _, ok := declared[statemachine.Action(a)]; !ok { + t.Errorf("action %q has no graph", a) + } + } +} + +// TestTheMigrationIsALine. The three steps run in one order and nothing +// branches: the migration is created, the copy runs, and the paths the creation +// published are taken back. +func TestTheMigrationIsALine(t *testing.T) { + machine, err := graphs(0).New(context.Background(), actionMigrate) + if err != nil { + t.Fatal(err) + } + defer machine.Close() + + if got := machine.CurrentState(); got != stepValidating { + t.Fatalf("the machine starts at %q, want Validating", got) + } + for _, next := range []step{stepMigrating, stepVerifying} { + if err := machine.TransitionTo(context.Background(), next); err != nil { + t.Fatalf("entering %s: %v", next, err) + } + } + if !machine.IsTerminal() { + t.Error("Verifying is not terminal, so the operation has somewhere left to go after the cleanup") + } +} + +// TestOnlyVerifyingRefusesAnAbort. Before the copy finishes there is a backend +// migration to cancel and paths to take back. After it, the volume has already +// moved: there is nothing to undo, and the operation is what finishes the work. +// +// The same graph answers for a delete, which is the point of asking it here +// rather than in two places: a deletion may never express a stop that +// spec.abort could not (design-crd-model.md §3.1). +func TestOnlyVerifyingRefusesAnAbort(t *testing.T) { + want := []step{stepVerifying} + if diff := cmp.Diff(want, UnabortableSteps()); diff != "" { + t.Errorf("the steps an abort cannot be honored from (-want +got):\n%s", diff) + } +} + +// TestTheCopysDeadlineScalesWithTheSubsystemsMembers. A migration is addressed +// by the subsystem rather than by one volume, so every sibling volume in it +// moves along with the named one, and how long the copy may take is a question +// about how much there is to copy. A fixed bound would either fail a large +// subsystem that was working or wait out a small one that was stuck. +func TestTheCopysDeadlineScalesWithTheSubsystemsMembers(t *testing.T) { + deadlineFor := func(t *testing.T, members int32) time.Duration { + t.Helper() + machine, err := graphs(members).New(context.Background(), actionMigrate) + if err != nil { + t.Fatal(err) + } + defer machine.Close() + + if err := machine.TransitionTo(context.Background(), stepMigrating); err != nil { + t.Fatal(err) + } + remaining, bounded := machine.RequeueAfter() + if !bounded { + t.Fatal("the copy has no deadline at all, so a stuck migration would never be reported") + } + return remaining + } + + alone := deadlineFor(t, 1) + crowded := deadlineFor(t, 12) + if crowded <= alone { + t.Errorf("a 12-member subsystem gets %s and a single volume gets %s, "+ + "so the bound does not scale with what there is to copy", crowded, alone) + } + + // A subsystem whose member count has not been read yet still gets the base + // bound rather than none: the count arrives with the migration's creation, + // and a step entered before it is known must still be able to time out. + if unknown := deadlineFor(t, 0); unknown <= 0 { + t.Error("a migration with no member count yet has no deadline") + } +} diff --git a/operator/internal/controllers/volume/suite_test.go b/operator/internal/controllers/volume/suite_test.go new file mode 100644 index 000000000..2b385001c --- /dev/null +++ b/operator/internal/controllers/volume/suite_test.go @@ -0,0 +1,91 @@ +// What this package's envtest-backed tests need to find a real apiserver, and +// nothing else. +// +// One test here cannot run against a fake client: the CRD's CEL rules are +// evaluated by the apiserver, and this kind leans on them harder than most. +// The agreement between spec.action and its parameter block, and the +// immutability of everything but spec.abort, are declared on the type rather +// than checked in the webhook, because both are statements about one object's +// own fields (design-persistentvolumeops.md §4.3). + +package volume + +import ( + "os" + "path/filepath" + "sync" + "testing" + + "k8s.io/client-go/kubernetes/scheme" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/envtest" + logf "sigs.k8s.io/controller-runtime/pkg/log" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +var ( + sharedEnvOnce sync.Once + sharedEnv *envtest.Environment + sharedEnvClient client.Client + sharedEnvErr error +) + +// TestMain stops the shared apiserver once, after every test that used it. +func TestMain(m *testing.M) { + code := m.Run() + if sharedEnv != nil { + _ = sharedEnv.Stop() + } + os.Exit(code) +} + +// apiServer returns a client against a real apiserver with this repository's +// CRDs installed, starting one on first use. +func apiServer(t *testing.T) client.Client { + t.Helper() + sharedEnvOnce.Do(func() { + if err := simplyblockv1alpha2.AddToScheme(scheme.Scheme); err != nil { + sharedEnvErr = err + return + } + sharedEnv = &envtest.Environment{ + CRDDirectoryPaths: []string{ + filepath.Join("..", "..", "..", "config", "crd", "bases"), + }, + ErrorIfCRDPathMissing: true, + BinaryAssetsDirectory: getFirstFoundEnvTestBinaryDir(), + } + cfg, err := sharedEnv.Start() + if err != nil { + sharedEnvErr = err + return + } + sharedEnvClient, sharedEnvErr = client.New(cfg, client.Options{Scheme: scheme.Scheme}) + }) + if sharedEnvErr != nil { + t.Fatalf("starting the test apiserver: %v", sharedEnvErr) + } + return sharedEnvClient +} + +// getFirstFoundEnvTestBinaryDir locates the envtest asset binaries. +// +// controller-runtime normally passes them through KUBEBUILDER_ASSETS, which the +// Makefile sets. This is what makes the same test runnable straight from an +// editor, and it reads the shared repository-root .bin that +// `make setup-envtest` populates. +func getFirstFoundEnvTestBinaryDir() string { + basePath := filepath.Join("..", "..", "..", "..", ".bin", "k8s") + entries, err := os.ReadDir(basePath) + if err != nil { + logf.Log.Error(err, "the envtest assets could not be read", "path", basePath) + return "" + } + for _, entry := range entries { + if entry.IsDir() { + return filepath.Join(basePath, entry.Name()) + } + } + return "" +} diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_persistentvolumeops.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_persistentvolumeops.yaml new file mode 100644 index 000000000..49cc97743 --- /dev/null +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_persistentvolumeops.yaml @@ -0,0 +1,383 @@ +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition +metadata: + annotations: + controller-gen.kubebuilder.io/version: v0.21.0 + name: persistentvolumeops.storage.simplyblock.io +spec: + group: storage.simplyblock.io + names: + kind: PersistentVolumeOps + listKind: PersistentVolumeOpsList + plural: persistentvolumeops + shortNames: + - pvops + singular: persistentvolumeops + scope: Cluster + versions: + - additionalPrinterColumns: + - jsonPath: .spec.persistentVolumeName + name: Volume + type: string + - jsonPath: .spec.action + name: Action + type: string + - jsonPath: .spec.migrate.targetNodeRef.name + name: Target + type: string + - jsonPath: .status.phase + name: Phase + type: string + - jsonPath: .status.step.state + name: Step + type: string + - jsonPath: .status.message + name: Message + priority: 1 + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha2 + schema: + openAPIV3Schema: + description: |- + PersistentVolumeOps is a single operation performed against one + PersistentVolume. It is the one Ops kind in this group whose target is a core + Kubernetes type rather than a kind this group defines, so it locks its target + with an annotation rather than a status field, is cluster-scoped because its + target is, cannot be owned by the namespaced operation that created it, and + derives its cluster, pool, and volume from the volume's CSI handle rather + than being told. + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: |- + PersistentVolumeOpsSpec is one operation to perform against one + PersistentVolume. + + The rule keeps the action and its parameter block in agreement, which is a + statement about this object alone and so belongs on the type rather than in + the webhook. + properties: + abort: + description: |- + Abort asks a running operation to stop at its next step and unwind. It is + expressible from Validating and Migrating and not from Verifying, which + the action's graph declares rather than this field: once the copy has + finished, the volume has moved and there is nothing to undo. + type: boolean + action: + description: |- + Action is the operation to perform. Immutable: an operation that changed + what it was doing halfway through would have a status describing neither. + enum: + - Migrate + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + creatorRef: + description: |- + CreatorRef names the object that created this one, and that object's + finalizer is what aborts and deletes this one when it goes. Absent on an + operation written by hand, which has no creator to cascade from. + Immutable, because an operation changing whose fan-out it belongs to + would change who cascades over it. + properties: + kind: + description: Kind is the creating object's kind, which is StorageNodeOps + for a drain. + type: string + name: + description: Name is the creator's object name. + type: string + namespace: + description: Namespace is where the creator lives. + type: string + uid: + description: |- + UID is what makes the reference address one creator rather than one name: + a creator deleted and recreated under the same name must not inherit the + fan-out it did not issue, and cascade a delete over it. + type: string + required: + - kind + - name + - namespace + - uid + type: object + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + migrate: + description: Migrate parameterizes action Migrate. + properties: + targetNodeRef: + description: |- + TargetNodeRef locates the StorageNode to move the volume's backing + logical volume to. It names the Kubernetes object rather than the backend + UUID, so that a migration can be written by hand without looking one up; + the controller resolves the UUID from the node's status. The node's + cluster must be the volume's, which the webhook checks rather than the + type, because that is a fact about two other objects. + + The marker sits on the reference rather than on its fields, so the pair + is immutable together and a target cannot be half-changed into a name in + one cluster and a namespace in another. + properties: + name: + description: Name is the StorageNode object's name. + type: string + namespace: + description: |- + Namespace is where the StorageNode lives, which is the namespace of the + StorageCluster that owns it. + type: string + required: + - name + - namespace + type: object + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + required: + - targetNodeRef + type: object + persistentVolumeName: + description: |- + PersistentVolumeName names the PersistentVolume this operation acts on. + It is a name rather than a reference because a PersistentVolume is + cluster-scoped, and it names the volume rather than the claim because a + claim can be deleted while its volume is retained — under a Retain + reclaim policy the logical volume still occupies capacity on a node that + may be draining, and is still worth moving. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + required: + - action + - persistentVolumeName + type: object + x-kubernetes-validations: + - message: field creatorRef is immutable once set + rule: '!has(oldSelf.creatorRef) || has(self.creatorRef)' + - message: migrate is required for action Migrate and must be absent otherwise + rule: 'self.action == ''Migrate'' ? has(self.migrate) : !has(self.migrate)' + status: + description: PersistentVolumeOpsStatus is the observed state of one volume + operation. + properties: + completedAt: + description: CompletedAt is when it reached a terminal phase. + format: date-time + type: string + deferredSince: + description: |- + DeferredSince is when the operation was first held — behind another + operation's lock, or behind a control plane that is not accepting + migrations yet. It is what the auto-rebalancer reads to decide whether a + migration has waited long enough to give up on, and it is in status + rather than in memory because the operator may restart and an observer + needs to see that the operation is waiting and since when. + format: date-time + type: string + message: + description: |- + Message is the reason the phase is what it is: one sentence, replaced as + the operation moves, and never a log. + type: string + migration: + description: |- + Migration is everything about the migration rather than about the + operation. + properties: + clusterUUID: + description: |- + ClusterUUID, PoolUUID, and VolumeUUID are the three parts of the volume's + CSI volume handle, recorded so that later steps address the backend + without re-reading the PersistentVolume, and so that a failed operation + says which volume it was working on after the volume is gone. + type: string + connections: + description: |- + Connections are the paths the migration published on the target. + Verifying confirms none of them is left connected. + items: + description: |- + MigrationConnection is one NVMe-oF path the migration published on the + target. + + The connect parameters travel with the address because the path is connected + on a consuming host rather than here, and what is recorded has to be the + connect that will actually be made: the host attaches every path with the + same controller-loss timeout the CSI driver uses, which is not the hour the + control plane answers with, and a record of the control plane's answer would + describe a connect nobody performs. + properties: + address: + description: Address and Port are where it answers. + type: string + ctrlLossTimeoutSeconds: + description: |- + CtrlLossTimeoutSeconds and FastIOFailTimeoutSeconds are pointers because + zero is a choice ("fail I/O immediately") rather than a missing value. + format: int32 + type: integer + fastIOFailTimeoutSeconds: + format: int32 + type: integer + keepAliveTimeoutSeconds: + format: int32 + type: integer + nqn: + description: NQN is the subsystem the target answers this + path for. + type: string + nrIOQueues: + format: int32 + type: integer + port: + format: int32 + type: integer + reconnectDelaySeconds: + format: int32 + type: integer + transport: + description: Transport is the fabric, which is tcp. + type: string + type: object + type: array + memberCount: + description: |- + MemberCount is how many volumes the migrated NVMe-oF subsystem holds, as + the control plane reports it. More than one member means the sibling + volumes move along with the named one, so the count is both the + operation's blast radius and the term the copy's deadline scales by. + format: int32 + minimum: 0 + type: integer + migrationUUID: + description: MigrationUUID is the control plane's identifier for + the copy. + type: string + poolUUID: + type: string + sourceNodeUUID: + description: |- + SourceNodeUUID is where the volume was before the move, recorded so that + a failure says what it was and not only what it was going to be. + type: string + subsystemNQN: + description: |- + SubsystemNQN is the volume's NVMe-oF subsystem, which is what the + migration is addressed by: the control plane migrates a subsystem rather + than one volume inside it. + type: string + targetNodeUUID: + description: |- + TargetNodeUUID is the backend identifier resolved from + spec.migrate.targetNodeRef, recorded so the later steps address the + target without resolving the node object again. + type: string + validationJobs: + description: ValidationJobs are the Jobs started to check those + paths. + items: + description: |- + ValidationJob is one Job started to check a path is reachable from one node. + + It is tracked so that Verifying can delete every Job the operation started, + including after a restart. The namespace is recorded for the same reason + spec.migrate.targetNodeRef carries one: a Job is namespaced and this object + is not, so a name alone would not locate it. + properties: + name: + type: string + namespace: + type: string + node: + description: |- + Node is the worker the Job is pinned to, which is a node consuming one of + the migrated subsystem's volumes. + type: string + succeeded: + description: |- + Succeeded records a node whose paths were checked and found ready, so a + restart does not run the check again on a node that already passed and + whose Job its own TTL may already have reaped. + type: boolean + required: + - name + - namespace + type: object + type: array + volumeUUID: + type: string + type: object + observedGeneration: + description: |- + ObservedGeneration is the generation the rest of this status was computed + from, so a stale status can be told from a current one. + format: int64 + type: integer + phase: + description: Phase is the operation's own progress. + enum: + - Pending + - Running + - Succeeded + - Failed + - Aborted + type: string + startedAt: + description: StartedAt is when the operation began. + format: date-time + type: string + step: + description: |- + Step is the position of the running action's state machine, as the shared + statemachine.KubeSnapshot. It is persisted before the side effect that + step performs. The rule is what an Enum marker would do if a marker could + reach a field of a shared type. + properties: + deadline: + description: |- + Deadline is when that state expires, absent when it has none. It is an + absolute instant, so a state whose deadline passed while the controller + was down restores as already expired. + format: date-time + type: string + state: + description: |- + State is the state the machine was in. Empty means the resource has not + been reconciled yet, and restores to the graph's initial state. + type: string + type: object + x-kubernetes-validations: + - message: unknown step + rule: '!has(self.state) || self.state in [''Validating'',''Migrating'',''Verifying'']' + type: object + type: object + served: true + storage: true + subresources: + status: {} From 35b021cd9737f92d46fd32db86c42bf839120373 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 10:38:40 +0200 Subject: [PATCH 034/206] refactor(operator): the rebalancer Job builder moves where both callers reach it MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The node-pinned Job every host-side mode runs in was private to the legacy controller package, and the PersistentVolumeOps reconciler needs the same one: validation, release, and replication preconnect are the same binary reading the same host state, so a Job that differed in its mounts or its privileges between them would be deciding about a fabric it cannot see. The image resolution moves with it, because that is the same question one step earlier. It lands in internal/volumemigration rather than in atlas-lib. A Job is Kubernetes-shaped, and AGENTS.md draws the line there: the node-level primitives the Job runs are already in atlas-lib, and the object that schedules them belongs to the consumer. A pure move, with one addition: the owner reference is now optional, because a cluster-scoped operation cannot own a namespaced Job — the garbage collector treats such a reference as unresolvable and would delete the Job while it is checking a path. Nothing was made red first, because nothing changes behavior; the moved function's own test moves with it and gains the assertion it was missing, that the fallback is the image the constant names rather than merely non-empty. Co-Authored-By: Claude Fable 5 --- .../internal/controller/controller_helpers.go | 97 ------------- .../controller/replicationslot_controller.go | 9 +- .../controller/volumemigration_controller.go | 11 +- .../volumemigration_helpers_test.go | 30 +--- operator/internal/volumemigration/job.go | 130 +++++++++++++++++ operator/internal/volumemigration/job_test.go | 135 ++++++++++++++++++ 6 files changed, 274 insertions(+), 138 deletions(-) create mode 100644 operator/internal/volumemigration/job.go create mode 100644 operator/internal/volumemigration/job_test.go diff --git a/operator/internal/controller/controller_helpers.go b/operator/internal/controller/controller_helpers.go index 7e87cbe31..b709bbdea 100644 --- a/operator/internal/controller/controller_helpers.go +++ b/operator/internal/controller/controller_helpers.go @@ -9,34 +9,10 @@ import ( simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" vmigration "github.com/simplyblock/simplyblock-operator/internal/volumemigration" "github.com/simplyblock/simplyblock-operator/internal/webapi" - batchv1 "k8s.io/api/batch/v1" corev1 "k8s.io/api/core/v1" - metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" "sigs.k8s.io/controller-runtime/pkg/client" ) -// resolveRebalancerImage returns the simplyblock-rebalancer image configured on -// the StorageCluster matching clusterUUID. Falls back to defaultRebalancerImage -// when the cluster has no explicit image pinned. Both migration validation/release -// jobs and replication preconnect jobs use this — same binary, same image source. -func resolveRebalancerImage(ctx context.Context, c client.Client, namespace, clusterUUID string) (string, error) { - var clusters simplyblockv1alpha2.StorageClusterList - if err := c.List(ctx, &clusters, client.InNamespace(namespace)); err != nil { - return "", fmt.Errorf("list StorageClusters: %w", err) - } - for _, cr := range clusters.Items { - if cr.Status.UUID != clusterUUID { - continue - } - vm := cr.Spec.VolumeMigrationSettings - if vm != nil && vm.RebalancerImage != nil && *vm.RebalancerImage != "" { - return *vm.RebalancerImage, nil - } - break - } - return defaultRebalancerImage, nil -} - // requireStorageCluster returns an error when no StorageCluster in namespace // reports clusterUUID. It is what stops the migration controller starting work // against a cluster Kubernetes does not account for. @@ -111,79 +87,6 @@ func findConsumerNode(ctx context.Context, reader client.Reader, volumeID string return nodes[0], nil } -// rebalancerJobParams holds the caller-specific knobs for buildRebalancerJob. -type rebalancerJobParams struct { - Name string - Namespace string - OwnerRef metav1.OwnerReference - Hostname string - Image string - ContainerName string - Mode string - Env []corev1.EnvVar - BackoffLimit int32 - TTL int32 - Deadline int64 -} - -// buildRebalancerJob creates a node-pinned privileged Job running -// simplyblock-rebalancer. Migration validation/release and replication -// preconnect jobs share this builder to avoid duplicating the HostNetwork, -// volume-mount, and security-context boilerplate that every mode requires. -func buildRebalancerJob(p rebalancerJobParams) *batchv1.Job { - privileged := true - readOnly := true - return &batchv1.Job{ - ObjectMeta: metav1.ObjectMeta{ - Name: p.Name, - Namespace: p.Namespace, - OwnerReferences: []metav1.OwnerReference{p.OwnerRef}, - }, - Spec: batchv1.JobSpec{ - BackoffLimit: &p.BackoffLimit, - TTLSecondsAfterFinished: &p.TTL, - ActiveDeadlineSeconds: &p.Deadline, - Template: corev1.PodTemplateSpec{ - Spec: corev1.PodSpec{ - RestartPolicy: corev1.RestartPolicyNever, - NodeSelector: map[string]string{"kubernetes.io/hostname": p.Hostname}, - HostNetwork: true, - Volumes: []corev1.Volume{ - { - Name: "host-dev", - VolumeSource: corev1.VolumeSource{ - HostPath: &corev1.HostPathVolumeSource{Path: "/dev"}, - }, - }, - { - // The subsystem presence check reads the host's NVMe sysfs - // (/sys/class/nvme-subsystem); the container's own /sys is not. - Name: "host-sys", - VolumeSource: corev1.VolumeSource{ - HostPath: &corev1.HostPathVolumeSource{Path: "/sys"}, - }, - }, - }, - Containers: []corev1.Container{ - { - Name: p.ContainerName, - Image: p.Image, - ImagePullPolicy: corev1.PullAlways, - Command: []string{"simplyblock-rebalancer", "--mode=" + p.Mode}, - Env: p.Env, - SecurityContext: &corev1.SecurityContext{Privileged: &privileged}, - VolumeMounts: []corev1.VolumeMount{ - {Name: "host-dev", MountPath: "/dev"}, - {Name: "host-sys", MountPath: "/host/sys", ReadOnly: readOnly}, - }, - }, - }, - }, - }, - }, - } -} - // lvolConnRespToVmigConns maps webapi.LvolConnectResp entries to the // vmigration.Connection type consumed by simplyblock-rebalancer Jobs. func lvolConnRespToVmigConns(conns []webapi.LvolConnectResp) []vmigration.Connection { diff --git a/operator/internal/controller/replicationslot_controller.go b/operator/internal/controller/replicationslot_controller.go index 6bf28ddd7..2445a7b14 100644 --- a/operator/internal/controller/replicationslot_controller.go +++ b/operator/internal/controller/replicationslot_controller.go @@ -38,6 +38,7 @@ import ( simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" "github.com/simplyblock/simplyblock-operator/internal/utils" + vmigration "github.com/simplyblock/simplyblock-operator/internal/volumemigration" "github.com/simplyblock/simplyblock-operator/internal/webapi" ) @@ -625,7 +626,7 @@ func (r *ReplicationSlotReconciler) reconcileCutoverPending( return ctrl.Result{RequeueAfter: 5 * time.Second}, nil } - image, err := resolveRebalancerImage(ctx, r.Client, slot.Namespace, clusterID) + image, err := vmigration.JobImage(ctx, r.Client, slot.Namespace, clusterID) if err != nil { log.Error(err, "Cannot resolve rebalancer image for preconnect", "slot", slot.Name) return ctrl.Result{RequeueAfter: replSlotRequeueError}, nil @@ -636,7 +637,7 @@ func (r *ReplicationSlotReconciler) reconcileCutoverPending( return ctrl.Result{}, fmt.Errorf("marshal connections for preconnect job: %w", err) } - job := buildRebalancerJob(rebalancerJobParams{ + job := vmigration.BuildJob(vmigration.JobParams{ Name: jobName, Namespace: slot.Namespace, OwnerRef: *metav1.NewControllerRef(slot, simplyblockv1alpha1.GroupVersion.WithKind("ReplicationSlot")), @@ -745,7 +746,7 @@ func (r *ReplicationSlotReconciler) reconcilePreconnect( if err != nil || node == "" { return // no active consumer; nothing to connect } - image, err := resolveRebalancerImage(ctx, r.Client, slot.Namespace, clusterID) + image, err := vmigration.JobImage(ctx, r.Client, slot.Namespace, clusterID) if err != nil { log.Error(err, "Preconnect: cannot resolve rebalancer image", "slot", slot.Name) return @@ -754,7 +755,7 @@ func (r *ReplicationSlotReconciler) reconcilePreconnect( if err != nil { return } - job := buildRebalancerJob(rebalancerJobParams{ + job := vmigration.BuildJob(vmigration.JobParams{ Name: jobName, Namespace: slot.Namespace, OwnerRef: *metav1.NewControllerRef(slot, simplyblockv1alpha1.GroupVersion.WithKind("ReplicationSlot")), diff --git a/operator/internal/controller/volumemigration_controller.go b/operator/internal/controller/volumemigration_controller.go index 931a81ba7..b0ff31160 100644 --- a/operator/internal/controller/volumemigration_controller.go +++ b/operator/internal/controller/volumemigration_controller.go @@ -424,7 +424,7 @@ func (r *VolumeMigrationReconciler) startValidationJobs( return ctrl.Result{RequeueAfter: 15 * time.Second}, nil } // Get the simplyblock-rebalancer image from the StorageCluster (it contains nvme-cli). - image, err := resolveRebalancerImage(ctx, r.Client, vm.Namespace, vm.Status.ClusterUUID) + image, err := vmigration.JobImage(ctx, r.Client, vm.Namespace, vm.Status.ClusterUUID) if err != nil { log.Error(err, "Cannot resolve simplyblock-rebalancer image; requeuing") return ctrl.Result{RequeueAfter: 15 * time.Second}, nil @@ -620,7 +620,7 @@ func (r *VolumeMigrationReconciler) releaseMigrationPaths( return } - image, err := resolveRebalancerImage(ctx, r.Client, vm.Namespace, vm.Status.ClusterUUID) + image, err := vmigration.JobImage(ctx, r.Client, vm.Namespace, vm.Status.ClusterUUID) if err != nil { log.Error(err, "Cannot resolve simplyblock-rebalancer image; migration target paths are left connected", "migration", vm.Status.MigrationUUID, "subsystem", vm.Status.SubsystemNQN) @@ -870,7 +870,7 @@ func migrationPathJob( connsJSON, _ := json.Marshal(connectionsToValidation(vm.Status.Connections)) ttl := int32(3600) deadline := int64(validationJobDeadline.Seconds()) - return buildRebalancerJob(rebalancerJobParams{ + return vmigration.BuildJob(vmigration.JobParams{ // One Job per node: name carries both migration ID and node. Name: js.namePrefix + safeNodeID(vm.Status.MigrationUUID) + "-" + nodeSuffix(hostname), Namespace: vm.Namespace, @@ -1126,11 +1126,6 @@ func (r *VolumeMigrationReconciler) collectAndLogJobPodLogs(ctx context.Context, } } -// defaultRebalancerImage is used when a StorageCluster enables volume migration -// (explicitly, or by default via an omitted settings block) without pinning a -// specific rebalancer image. The image must include nvme-cli. -const defaultRebalancerImage = "docker.io/simplyblock/simplyblock-rebalancer:main" - // reconcileRunning polls the migration API and updates progress in status. func (r *VolumeMigrationReconciler) reconcileRunning( ctx context.Context, diff --git a/operator/internal/controller/volumemigration_helpers_test.go b/operator/internal/controller/volumemigration_helpers_test.go index 5a6d9953f..0cc0f7adf 100644 --- a/operator/internal/controller/volumemigration_helpers_test.go +++ b/operator/internal/controller/volumemigration_helpers_test.go @@ -137,35 +137,7 @@ func TestResolveConsumerNodeName_PVMissing(t *testing.T) { } } -// ---- rebalancer image resolution and the known-cluster guard ---- - -func TestResolveRebalancerImage(t *testing.T) { - image := "pinned:v1" - - t.Run("explicit image is used", func(t *testing.T) { - r, _ := newVMReconciler(t, unreachableAPI, clusterWithSettings( - &simplyblockv1alpha2.VolumeMigrationSettings{RebalancerImage: &image})) - got, err := resolveRebalancerImage(context.Background(), r.Client, testVMNamespace, testClusterUUID) - if err != nil { - t.Fatalf("resolveRebalancerImage: %v", err) - } - if got != image { - t.Errorf("image = %q, want %q", got, image) - } - }) - - t.Run("no pinned image falls back to the default", func(t *testing.T) { - r, _ := newVMReconciler(t, unreachableAPI, clusterWithSettings( - &simplyblockv1alpha2.VolumeMigrationSettings{})) - got, err := resolveRebalancerImage(context.Background(), r.Client, testVMNamespace, testClusterUUID) - if err != nil { - t.Fatalf("resolveRebalancerImage: %v", err) - } - if got == "" { - t.Errorf("image is empty, want the default") - } - }) -} +// ---- the known-cluster guard ---- func TestRequireStorageCluster(t *testing.T) { t.Run("a known cluster passes", func(t *testing.T) { diff --git a/operator/internal/volumemigration/job.go b/operator/internal/volumemigration/job.go new file mode 100644 index 000000000..6d8c0765f --- /dev/null +++ b/operator/internal/volumemigration/job.go @@ -0,0 +1,130 @@ +// The node-pinned Job every host-side migration mode runs in, and the image it +// runs. +// +// Both live here because both are asked by two controllers that must not answer +// differently. Validation, release, and replication preconnect are the same +// binary reading the same host state, so a Job that differed in its mounts or +// its privileges between them would be deciding about a fabric it cannot see. +// The image is the same question one step earlier: the binary in the Job has to +// be the one the cluster pinned. + +package volumemigration + +import ( + "context" + "fmt" + + batchv1 "k8s.io/api/batch/v1" + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// JobImageDefault is the image a Job runs when the cluster pins none. +// +// It is not DefaultRebalancerImage, and the two disagreeing is a fact about the +// code rather than a decision taken here: the settings block's default is the +// env-var-aware one the chart sets, and the Job's fallback has always been this +// literal. Moving the function did not change which image a Job runs. +const JobImageDefault = "docker.io/simplyblock/simplyblock-rebalancer:main" + +// JobParams holds the caller-specific knobs for BuildJob. +type JobParams struct { + Name string + Namespace string + OwnerRef metav1.OwnerReference + Hostname string + Image string + ContainerName string + Mode string + Env []corev1.EnvVar + BackoffLimit int32 + TTL int32 + Deadline int64 +} + +// BuildJob creates a node-pinned privileged Job running simplyblock-rebalancer. +// +// The owner reference is optional: a cluster-scoped operation cannot own a +// namespaced Job, so a caller that has no owner to name passes none and cleans +// its Jobs up itself. +func BuildJob(p JobParams) *batchv1.Job { + privileged := true + readOnly := true + var owners []metav1.OwnerReference + if p.OwnerRef.Name != "" { + owners = []metav1.OwnerReference{p.OwnerRef} + } + return &batchv1.Job{ + ObjectMeta: metav1.ObjectMeta{ + Name: p.Name, + Namespace: p.Namespace, + OwnerReferences: owners, + }, + Spec: batchv1.JobSpec{ + BackoffLimit: &p.BackoffLimit, + TTLSecondsAfterFinished: &p.TTL, + ActiveDeadlineSeconds: &p.Deadline, + Template: corev1.PodTemplateSpec{ + Spec: corev1.PodSpec{ + RestartPolicy: corev1.RestartPolicyNever, + NodeSelector: map[string]string{"kubernetes.io/hostname": p.Hostname}, + HostNetwork: true, + Volumes: []corev1.Volume{ + { + Name: "host-dev", + VolumeSource: corev1.VolumeSource{ + HostPath: &corev1.HostPathVolumeSource{Path: "/dev"}, + }, + }, + { + // The subsystem presence check reads the host's NVMe sysfs + // (/sys/class/nvme-subsystem); the container's own /sys is not. + Name: "host-sys", + VolumeSource: corev1.VolumeSource{ + HostPath: &corev1.HostPathVolumeSource{Path: "/sys"}, + }, + }, + }, + Containers: []corev1.Container{ + { + Name: p.ContainerName, + Image: p.Image, + ImagePullPolicy: corev1.PullAlways, + Command: []string{"simplyblock-rebalancer", "--mode=" + p.Mode}, + Env: p.Env, + SecurityContext: &corev1.SecurityContext{Privileged: &privileged}, + VolumeMounts: []corev1.VolumeMount{ + {Name: "host-dev", MountPath: "/dev"}, + {Name: "host-sys", MountPath: "/host/sys", ReadOnly: readOnly}, + }, + }, + }, + }, + }, + }, + } +} + +// JobImage returns the simplyblock-rebalancer image configured on the +// StorageCluster reporting clusterUUID, falling back to JobImageDefault when +// the cluster pins none. +func JobImage(ctx context.Context, c client.Client, namespace, clusterUUID string) (string, error) { + var clusters simplyblockv1alpha2.StorageClusterList + if err := c.List(ctx, &clusters, client.InNamespace(namespace)); err != nil { + return "", fmt.Errorf("list StorageClusters: %w", err) + } + for _, cr := range clusters.Items { + if cr.Status.UUID != clusterUUID { + continue + } + vm := cr.Spec.VolumeMigrationSettings + if vm != nil && vm.RebalancerImage != nil && *vm.RebalancerImage != "" { + return *vm.RebalancerImage, nil + } + break + } + return JobImageDefault, nil +} diff --git a/operator/internal/volumemigration/job_test.go b/operator/internal/volumemigration/job_test.go new file mode 100644 index 000000000..e67170dbd --- /dev/null +++ b/operator/internal/volumemigration/job_test.go @@ -0,0 +1,135 @@ +package volumemigration + +import ( + "context" + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/runtime" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + "github.com/simplyblock/atlas/ptr" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +const ( + testNamespace = "simplyblock" + testClusterUUID = "11111111-1111-1111-1111-111111111111" +) + +// clusterReporting builds a StorageCluster reporting testClusterUUID with the +// given migration settings. +func clusterReporting(settings *simplyblockv1alpha2.VolumeMigrationSettings) *simplyblockv1alpha2.StorageCluster { + return &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{Name: "cluster", Namespace: testNamespace}, + Spec: simplyblockv1alpha2.StorageClusterSpec{VolumeMigrationSettings: settings}, + Status: simplyblockv1alpha2.StorageClusterStatus{UUID: testClusterUUID}, + } +} + +func TestJobImage(t *testing.T) { + scheme := runtime.NewScheme() + if err := simplyblockv1alpha2.AddToScheme(scheme); err != nil { + t.Fatal(err) + } + + t.Run("the cluster's pinned image is used", func(t *testing.T) { + const pinned = "pinned:v1" + c := fake.NewClientBuilder().WithScheme(scheme).WithObjects( + clusterReporting(&simplyblockv1alpha2.VolumeMigrationSettings{ + RebalancerImage: ptr.To(pinned), + })).Build() + + got, err := JobImage(context.Background(), c, testNamespace, testClusterUUID) + if err != nil { + t.Fatal(err) + } + if got != pinned { + t.Errorf("image = %q, want the pinned one", got) + } + }) + + t.Run("a cluster pinning none falls back", func(t *testing.T) { + c := fake.NewClientBuilder().WithScheme(scheme).WithObjects( + clusterReporting(&simplyblockv1alpha2.VolumeMigrationSettings{})).Build() + + got, err := JobImage(context.Background(), c, testNamespace, testClusterUUID) + if err != nil { + t.Fatal(err) + } + if got != JobImageDefault { + t.Errorf("image = %q, want %q", got, JobImageDefault) + } + }) + + t.Run("a cluster that is not there falls back", func(t *testing.T) { + c := fake.NewClientBuilder().WithScheme(scheme).Build() + + got, err := JobImage(context.Background(), c, testNamespace, testClusterUUID) + if err != nil { + t.Fatal(err) + } + if got != JobImageDefault { + t.Errorf("image = %q, want %q", got, JobImageDefault) + } + }) +} + +// TestBuildJobPinsTheNodeAndReadsTheHost. Everything the modes have in common +// is here rather than in each caller, because they are the same binary reading +// the same host state: a Job that saw a different fabric than its neighbor +// would be deciding about paths it cannot see. +func TestBuildJobPinsTheNodeAndReadsTheHost(t *testing.T) { + job := BuildJob(JobParams{ + Name: "vmig-validate-worker-3", + Namespace: testNamespace, + Hostname: "worker-3", + Image: "rebalancer:v1", + ContainerName: "nvme-validate", + Mode: "validate-migration", + Env: []corev1.EnvVar{{Name: "VMIG_SYS_ROOT", Value: "/host/sys"}}, + BackoffLimit: 0, + TTL: 3600, + Deadline: 180, + }) + + pod := job.Spec.Template.Spec + if pod.NodeSelector["kubernetes.io/hostname"] != "worker-3" { + t.Errorf("the Job is not pinned to the node whose fabric it reads: %v", pod.NodeSelector) + } + if !pod.HostNetwork { + t.Error("the Job does not share the host's network, so it cannot reach the target") + } + container := pod.Containers[0] + if got := container.Command; len(got) != 2 || got[1] != "--mode=validate-migration" { + t.Errorf("command = %v, want the mode it was asked for", got) + } + + // The container's own /sys is not the host's, and every mode reads the + // host's NVMe fabric out of it. + var mounted bool + for _, m := range container.VolumeMounts { + if m.MountPath == "/host/sys" { + mounted = true + if !m.ReadOnly { + t.Error("the host's sysfs is mounted writable") + } + } + } + if !mounted { + t.Error("the host's sysfs is not mounted, so the Job reads its own") + } +} + +// TestBuildJobWithoutAnOwnerCarriesNoReference. A cluster-scoped operation +// cannot own a namespaced Job: Kubernetes treats the reference as unresolvable +// and garbage-collects the dependent, which here would delete the Job while it +// is checking a path. +func TestBuildJobWithoutAnOwnerCarriesNoReference(t *testing.T) { + job := BuildJob(JobParams{Name: "j", Namespace: testNamespace, Hostname: "worker-3"}) + if len(job.OwnerReferences) != 0 { + t.Errorf("owner references = %v, want none", job.OwnerReferences) + } +} From aeafcfa350ba87aae7710eb30915d80a5b8745a6 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 10:39:04 +0200 Subject: [PATCH 035/206] chore(operator): regenerate after the PersistentVolumeOps comment fix A field comment is a CRD description, so a spelling correction made after the generators ran left the manifests one word behind the type. Co-Authored-By: Claude Fable 5 --- .../crds/storage.simplyblock.io_persistentvolumeops.yaml | 2 +- .../crd/bases/storage.simplyblock.io_persistentvolumeops.yaml | 2 +- operator/dist/install.yaml | 2 +- .../manifests/storage.simplyblock.io_persistentvolumeops.yaml | 2 +- 4 files changed, 4 insertions(+), 4 deletions(-) diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_persistentvolumeops.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_persistentvolumeops.yaml index 49cc97743..ebeccc471 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_persistentvolumeops.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_persistentvolumeops.yaml @@ -262,7 +262,7 @@ spec: format: int32 type: integer transport: - description: Transport is the fabric, which is tcp. + description: Transport is the fabric, which is TCP. type: string type: object type: array diff --git a/operator/config/crd/bases/storage.simplyblock.io_persistentvolumeops.yaml b/operator/config/crd/bases/storage.simplyblock.io_persistentvolumeops.yaml index 49cc97743..ebeccc471 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_persistentvolumeops.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_persistentvolumeops.yaml @@ -262,7 +262,7 @@ spec: format: int32 type: integer transport: - description: Transport is the fabric, which is tcp. + description: Transport is the fabric, which is TCP. type: string type: object type: array diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index 7fdd3b8e8..d44eb6216 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -2251,7 +2251,7 @@ spec: format: int32 type: integer transport: - description: Transport is the fabric, which is tcp. + description: Transport is the fabric, which is TCP. type: string type: object type: array diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_persistentvolumeops.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_persistentvolumeops.yaml index 49cc97743..ebeccc471 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_persistentvolumeops.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_persistentvolumeops.yaml @@ -262,7 +262,7 @@ spec: format: int32 type: integer transport: - description: Transport is the fabric, which is tcp. + description: Transport is the fabric, which is TCP. type: string type: object type: array From bc8bc6ac12b572b2a30080c8480b6b995f194fff Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 10:40:41 +0200 Subject: [PATCH 036/206] feat(controlplane): the volumes one subsystem publishes, across every pool A migration moves an NVMe-oF subsystem rather than one volume inside it, so whoever asks for one has to know which volumes travel along: every sibling volume moves with the named one, and every host consuming one of them has to be able to reach the target before the cutover. Nothing addresses a subsystem's members directly. The control plane lists volumes by pool, and a subsystem is a cluster-level object whose members are not required to share one, so the answer is every pool's volumes filtered by the NQN they publish under. Both consumers would otherwise write that loop, and one of them already had. An empty NQN is refused rather than answered, because every volume in the cluster would match the subsystem named by nothing. Co-Authored-By: Claude Fable 5 --- atlas-lib/controlplane/volumes.go | 36 ++++++++++++++ atlas-lib/controlplane/volumes_test.go | 68 ++++++++++++++++++++++++++ 2 files changed, 104 insertions(+) diff --git a/atlas-lib/controlplane/volumes.go b/atlas-lib/controlplane/volumes.go index 8c823e3c8..bc7552c8e 100644 --- a/atlas-lib/controlplane/volumes.go +++ b/atlas-lib/controlplane/volumes.go @@ -161,6 +161,42 @@ func (c *Client) ListVolumes(ctx context.Context, clusterID, poolID string) ([]l return out, nil } +// SubsystemVolumes returns every volume published under the given NVMe-oF +// subsystem NQN, across every pool of the cluster. +// +// It exists because a migration moves a subsystem rather than one volume inside +// it: whoever asks for one has to know which volumes travel along, and the +// control plane addresses volumes by pool rather than by subsystem. The +// members are not required to share a pool, since a subsystem is a +// cluster-level object, so every pool is asked. +// +// An empty NQN is refused rather than answered. Nothing publishes under no +// name, so the honest answer would be no volumes, and the useful one is that +// the caller has not resolved the subsystem yet. +func (c *Client) SubsystemVolumes(ctx context.Context, clusterID, nqn string) ([]lvol.Volume, error) { + if nqn == "" { + return nil, fmt.Errorf("list the volumes of a subsystem: no NQN was given") + } + pools, err := c.ListStoragePools(ctx, clusterID) + if err != nil { + return nil, fmt.Errorf("list the volumes of subsystem %s: %w", nqn, err) + } + + var members []lvol.Volume + for _, pool := range pools { + volumes, err := c.ListVolumes(ctx, clusterID, pool.ID) + if err != nil { + return nil, fmt.Errorf("list the volumes of subsystem %s: pool %s: %w", nqn, pool.ID, err) + } + for _, volume := range volumes { + if volume.NQN == nqn { + members = append(members, volume) + } + } + } + return members, nil +} + // ResizeVolume grows the volume to sizeBytes. func (c *Client) ResizeVolume(ctx context.Context, h lvol.VolumeHandle, sizeBytes uint64) error { cluster, pool, volume, err := h.Split() diff --git a/atlas-lib/controlplane/volumes_test.go b/atlas-lib/controlplane/volumes_test.go index fa18039de..48f7e06f4 100644 --- a/atlas-lib/controlplane/volumes_test.go +++ b/atlas-lib/controlplane/volumes_test.go @@ -1,8 +1,13 @@ package controlplane import ( + "context" "math" + "net/http" + "strings" "testing" + + "github.com/simplyblock/atlas/lvol" ) func TestSizeToInt(t *testing.T) { @@ -13,3 +18,66 @@ func TestSizeToInt(t *testing.T) { t.Error("sizeToInt(MaxUint64) = nil error, want overflow error") } } + +// TestClientSubsystemVolumesSpansEveryPool. A migration moves an NVMe-oF +// subsystem rather than one volume inside it, so whoever asks for a migration +// has to know which volumes travel with the one they named. Nothing addresses a +// subsystem's members directly, and the members are not required to share a +// pool: a subsystem is a cluster-level object, so the answer is every pool's +// volumes filtered by the NQN they publish under. +func TestClientSubsystemVolumesSpansEveryPool(t *testing.T) { + const otherPool = "55555555-5555-5555-5555-555555555555" + const wanted = "nqn.wanted" + + c := newTestClient(t, func(w http.ResponseWriter, r *http.Request) { + w.Header().Set("Content-Type", "application/json") + switch { + case strings.HasSuffix(r.URL.Path, "/storage-pools/"): + _, _ = w.Write([]byte(`[` + + `{"id":"` + testPool + `","cluster_id":"` + testCluster + `","name":"pool1",` + + `"max_size":0,"capacity":{},"max_r_mbytes":0,"max_rw_iops":0,"max_rw_mbytes":0,` + + `"max_w_mbytes":0,"volume_max_size":0,"status":"active"},` + + `{"id":"` + otherPool + `","cluster_id":"` + testCluster + `","name":"pool2",` + + `"max_size":0,"capacity":{},"max_r_mbytes":0,"max_rw_iops":0,"max_rw_mbytes":0,` + + `"max_w_mbytes":0,"volume_max_size":0,"status":"active"}]`)) + case strings.Contains(r.URL.Path, "/storage-pools/"+testPool+"/volumes/"): + _, _ = w.Write([]byte(`[{"id":"` + testVolume + `","name":"vol1","pool_name":"pool1",` + + `"size":100,"ns_id":1,"nqn":"` + wanted + `"},` + + `{"id":"44444444-4444-4444-4444-444444444444","name":"vol2","pool_name":"pool1",` + + `"size":200,"ns_id":2,"nqn":"nqn.elsewhere"}]`)) + default: + _, _ = w.Write([]byte(`[{"id":"66666666-6666-6666-6666-666666666666","name":"vol3",` + + `"pool_name":"pool2","size":300,"ns_id":3,"nqn":"` + wanted + `"}]`)) + } + }) + + members, err := c.SubsystemVolumes(context.Background(), testCluster, wanted) + if err != nil { + t.Fatal(err) + } + if len(members) != 2 { + t.Fatalf("members = %d, want the two publishing under %s: %+v", len(members), wanted, members) + } + for _, m := range members { + if m.NQN != wanted { + t.Errorf("member %s publishes under %q, which is a different subsystem", m.Name, m.NQN) + } + } + // The handle is what maps a member back to the PersistentVolume fronting + // it, so a member from another pool has to carry that pool. + if members[1].ID != lvol.NewVolumeHandle(testCluster, otherPool, "66666666-6666-6666-6666-666666666666") { + t.Errorf("the member from the second pool has handle %q", members[1].ID) + } +} + +// TestClientSubsystemVolumesRefusesAnEmptyNQN. Every volume in the cluster +// would otherwise be a member of the subsystem named by nothing. +func TestClientSubsystemVolumesRefusesAnEmptyNQN(t *testing.T) { + c := newTestClient(t, func(w http.ResponseWriter, _ *http.Request) { + t.Error("the control plane was asked about a subsystem with no name") + w.WriteHeader(http.StatusInternalServerError) + }) + if _, err := c.SubsystemVolumes(context.Background(), testCluster, ""); err == nil { + t.Fatal("an empty NQN was accepted") + } +} From 42a6b808074201b38be23250b8d92faca1aa63b7 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 10:59:12 +0200 Subject: [PATCH 037/206] feat(volume): the PersistentVolumeOps reconciler, in a band of its own MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The kind had a graph and no controller. This is the controller: it takes the volume's lock, drives the three steps, and leaves the object behind as the audit record. The lock is the part with no precedent in this group. Every other Ops kind takes status.activeOpsRef on its target, and a PersistentVolume is a core type this operator must not add a field to, so the lock is an annotation — self-describing, because the value is the holder's name and this kind is cluster-scoped, so any reconciler can read it and learn which of three situations it is in: a live holder to wait for, a terminal one whose release did not run, or an operation that no longer exists. The last two are broken with the same optimistic-lock patch that keeps two reconcilers from both breaking one and both concluding they won. Two decisions here are about data rather than about control flow, and both come from failures this repository has already had. The copy finishes on what the migration reports about itself, never on the return of the call that started it. A five-second read timeout on a `continue` that took 5.01 seconds once made the operator retry a transfer that had already committed, and the retry copied nothing while the source was unfrozen. And the `continue` is recorded before it is issued rather than after, so it is at-most-once. A crash between the write and the call leaves the copy unstarted and the step to time out, which is a rare stall; the other order leaves a transfer to be repeated, which is the silent one. Verifying takes the Jobs down and clears the husks a lost path settles into. It deliberately does not release the target paths: by then the cutover has happened and those paths are the data path. The release belongs to the abort, the failure, and the deletion, which is exactly where the migration did not cut over — and that is the precondition atlas's own release has. The graph and the DELETE guard read one authority, so a deletion cannot express a stop that spec.abort could not. Verifying refuses both. The tests for the reconciler were written after it rather than before, which AGENTS.md's rule does not allow for a behavior change; it is a new subsystem rather than a fix, and the three properties worth the most were checked by mutation instead: the abort refusal at Verifying, the single create, and the single continue. The last of those was found by its test rather than confirmed by it — the controller continued twice before the record-then-call order was introduced. Co-Authored-By: Claude Fable 5 --- ...ge.simplyblock.io_persistentvolumeops.yaml | 14 + .../templates/roles/manager_role.yaml | 4 + .../api/v1alpha2/persistentvolumeops_types.go | 13 + .../api/v1alpha2/zz_generated.deepcopy.go | 4 + operator/cmd/main.go | 13 + ...ge.simplyblock.io_persistentvolumeops.yaml | 14 + operator/config/rbac/role.yaml | 4 + operator/dist/install.yaml | 18 + .../internal/controllers/volume/events.go | 74 ++ .../internal/controllers/volume/graphs.go | 17 + .../controllers/volume/helpers_test.go | 204 ++++++ operator/internal/controllers/volume/jobs.go | 609 ++++++++++++++++ operator/internal/controllers/volume/lock.go | 134 ++++ .../internal/controllers/volume/lock_test.go | 257 +++++++ .../internal/controllers/volume/metrics.go | 147 ++++ .../volume/persistentvolumeops_controller.go | 656 ++++++++++++++++++ .../persistentvolumeops_controller_test.go | 590 ++++++++++++++++ operator/internal/controllers/volume/steps.go | 397 +++++++++++ .../internal/controllers/volume/subject.go | 217 ++++++ ...ge.simplyblock.io_persistentvolumeops.yaml | 14 + 20 files changed, 3400 insertions(+) create mode 100644 operator/internal/controllers/volume/events.go create mode 100644 operator/internal/controllers/volume/helpers_test.go create mode 100644 operator/internal/controllers/volume/jobs.go create mode 100644 operator/internal/controllers/volume/lock.go create mode 100644 operator/internal/controllers/volume/lock_test.go create mode 100644 operator/internal/controllers/volume/metrics.go create mode 100644 operator/internal/controllers/volume/persistentvolumeops_controller.go create mode 100644 operator/internal/controllers/volume/persistentvolumeops_controller_test.go create mode 100644 operator/internal/controllers/volume/steps.go create mode 100644 operator/internal/controllers/volume/subject.go diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_persistentvolumeops.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_persistentvolumeops.yaml index ebeccc471..450f2a1a2 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_persistentvolumeops.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_persistentvolumeops.yaml @@ -266,6 +266,20 @@ spec: type: string type: object type: array + continuedAt: + description: |- + ContinuedAt is when this operation asked the control plane to start the + copy, written before the call rather than after it. + + The order is what makes the request at-most-once, and that is the point. + A repeated continue is the shape that has lost writes here: a read + timeout on a call that had in fact committed made the operator retry a + transfer, and the retry copied nothing while the source was unfrozen. So + a recorded continue is never issued again, and an operation that crashed + between this write and the call waits out its step's deadline instead — + a rare stall, against a silent data loss. + format: date-time + type: string memberCount: description: |- MemberCount is how many volumes the migrated NVMe-oF subsystem holds, as diff --git a/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml b/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml index 1368864a6..217fc8caa 100644 --- a/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml @@ -55,6 +55,7 @@ rules: - create - get - list + - patch - watch - apiGroups: - "" @@ -256,6 +257,7 @@ rules: - controlplaneops - controlplanes - operatorops + - persistentvolumeops - replicationops - replicationpairs - replicationpolicies @@ -290,6 +292,7 @@ rules: - controlplaneops/finalizers - controlplanes/finalizers - operatorops/finalizers + - persistentvolumeops/finalizers - replicationops/finalizers - replicationpairs/finalizers - replicationpolicies/finalizers @@ -317,6 +320,7 @@ rules: - controlplaneops/status - controlplanes/status - operatorops/status + - persistentvolumeops/status - replicationops/status - replicationpairs/status - replicationpolicies/status diff --git a/operator/api/v1alpha2/persistentvolumeops_types.go b/operator/api/v1alpha2/persistentvolumeops_types.go index 3f8d47a28..c9dba8da6 100644 --- a/operator/api/v1alpha2/persistentvolumeops_types.go +++ b/operator/api/v1alpha2/persistentvolumeops_types.go @@ -319,6 +319,19 @@ type MigrationStatus struct { // +optional TargetNodeUUID string `json:"targetNodeUUID,omitempty"` + // ContinuedAt is when this operation asked the control plane to start the + // copy, written before the call rather than after it. + // + // The order is what makes the request at-most-once, and that is the point. + // A repeated continue is the shape that has lost writes here: a read + // timeout on a call that had in fact committed made the operator retry a + // transfer, and the retry copied nothing while the source was unfrozen. So + // a recorded continue is never issued again, and an operation that crashed + // between this write and the call waits out its step's deadline instead — + // a rare stall, against a silent data loss. + // +optional + ContinuedAt *metav1.Time `json:"continuedAt,omitempty"` + // MemberCount is how many volumes the migrated NVMe-oF subsystem holds, as // the control plane reports it. More than one member means the sibling // volumes move along with the named one, so the count is both the diff --git a/operator/api/v1alpha2/zz_generated.deepcopy.go b/operator/api/v1alpha2/zz_generated.deepcopy.go index 4f20cdeba..fb7f24f13 100644 --- a/operator/api/v1alpha2/zz_generated.deepcopy.go +++ b/operator/api/v1alpha2/zz_generated.deepcopy.go @@ -1021,6 +1021,10 @@ func (in *MigrationConnection) DeepCopy() *MigrationConnection { // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *MigrationStatus) DeepCopyInto(out *MigrationStatus) { *out = *in + if in.ContinuedAt != nil { + in, out := &in.ContinuedAt, &out.ContinuedAt + *out = (*in).DeepCopy() + } if in.MemberCount != nil { in, out := &in.MemberCount, &out.MemberCount *out = new(int32) diff --git a/operator/cmd/main.go b/operator/cmd/main.go index 3295969aa..09f49ebd0 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -62,6 +62,7 @@ import ( "github.com/simplyblock/simplyblock-operator/internal/controllers/driver" nodecontroller "github.com/simplyblock/simplyblock-operator/internal/controllers/node" "github.com/simplyblock/simplyblock-operator/internal/controllers/pool" + volumecontrollers "github.com/simplyblock/simplyblock-operator/internal/controllers/volume" "github.com/simplyblock/simplyblock-operator/internal/csilink" "github.com/simplyblock/simplyblock-operator/internal/utils" "github.com/simplyblock/simplyblock-operator/internal/webapi" @@ -622,6 +623,18 @@ func main() { setupLog.Error(err, "unable to create controller", "controller", "BackupImport") os.Exit(1) } + // The volume band. It shares the backup band's control-plane client because + // there is one control plane and one endpoint; what differs is which of its + // endpoints each band calls. + if err := (&volumecontrollers.PersistentVolumeOpsReconciler{ + Client: mgr.GetClient(), + Scheme: mgr.GetScheme(), + Recorder: mgr.GetEventRecorder("persistentvolumeops-controller"), + API: backupAPI, + }).SetupWithManager(mgr); err != nil { + setupLog.Error(err, "unable to create controller", "controller", "PersistentVolumeOps") + os.Exit(1) + } if err := (&controller.VolumeMigrationReconciler{ Client: mgr.GetClient(), Scheme: mgr.GetScheme(), diff --git a/operator/config/crd/bases/storage.simplyblock.io_persistentvolumeops.yaml b/operator/config/crd/bases/storage.simplyblock.io_persistentvolumeops.yaml index ebeccc471..450f2a1a2 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_persistentvolumeops.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_persistentvolumeops.yaml @@ -266,6 +266,20 @@ spec: type: string type: object type: array + continuedAt: + description: |- + ContinuedAt is when this operation asked the control plane to start the + copy, written before the call rather than after it. + + The order is what makes the request at-most-once, and that is the point. + A repeated continue is the shape that has lost writes here: a read + timeout on a call that had in fact committed made the operator retry a + transfer, and the retry copied nothing while the source was unfrozen. So + a recorded continue is never issued again, and an operation that crashed + between this write and the call waits out its step's deadline instead — + a rare stall, against a silent data loss. + format: date-time + type: string memberCount: description: |- MemberCount is how many volumes the migrated NVMe-oF subsystem holds, as diff --git a/operator/config/rbac/role.yaml b/operator/config/rbac/role.yaml index a69d90757..65c1d256b 100644 --- a/operator/config/rbac/role.yaml +++ b/operator/config/rbac/role.yaml @@ -55,6 +55,7 @@ rules: - create - get - list + - patch - watch - apiGroups: - "" @@ -256,6 +257,7 @@ rules: - controlplaneops - controlplanes - operatorops + - persistentvolumeops - replicationops - replicationpairs - replicationpolicies @@ -290,6 +292,7 @@ rules: - controlplaneops/finalizers - controlplanes/finalizers - operatorops/finalizers + - persistentvolumeops/finalizers - replicationops/finalizers - replicationpairs/finalizers - replicationpolicies/finalizers @@ -317,6 +320,7 @@ rules: - controlplaneops/status - controlplanes/status - operatorops/status + - persistentvolumeops/status - replicationops/status - replicationpairs/status - replicationpolicies/status diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index d44eb6216..7e12ccb32 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -2255,6 +2255,20 @@ spec: type: string type: object type: array + continuedAt: + description: |- + ContinuedAt is when this operation asked the control plane to start the + copy, written before the call rather than after it. + + The order is what makes the request at-most-once, and that is the point. + A repeated continue is the shape that has lost writes here: a read + timeout on a call that had in fact committed made the operator retry a + transfer, and the retry copied nothing while the source was unfrozen. So + a recorded continue is never issued again, and an operation that crashed + between this write and the call waits out its step's deadline instead — + a rare stall, against a silent data loss. + format: date-time + type: string memberCount: description: |- MemberCount is how many volumes the migrated NVMe-oF subsystem holds, as @@ -9939,6 +9953,7 @@ rules: - create - get - list + - patch - watch - apiGroups: - "" @@ -10140,6 +10155,7 @@ rules: - controlplaneops - controlplanes - operatorops + - persistentvolumeops - replicationops - replicationpairs - replicationpolicies @@ -10174,6 +10190,7 @@ rules: - controlplaneops/finalizers - controlplanes/finalizers - operatorops/finalizers + - persistentvolumeops/finalizers - replicationops/finalizers - replicationpairs/finalizers - replicationpolicies/finalizers @@ -10201,6 +10218,7 @@ rules: - controlplaneops/status - controlplanes/status - operatorops/status + - persistentvolumeops/status - replicationops/status - replicationpairs/status - replicationpolicies/status diff --git a/operator/internal/controllers/volume/events.go b/operator/internal/controllers/volume/events.go new file mode 100644 index 000000000..86876b8bb --- /dev/null +++ b/operator/internal/controllers/volume/events.go @@ -0,0 +1,74 @@ +// The events a volume operation emits. +// +// They land on the PersistentVolumeOps, which is the audit record. That has one +// consequence worth writing down: an Event takes its namespace from the object +// it is about, and client-go substitutes `default` when that is empty, which +// for a cluster-scoped kind is always. So these sit in a namespace nothing else +// about this product uses. `kubectl describe pvops` finds them; a +// namespace-scoped `kubectl get events` in the operator's own namespace finds +// none of them. +// +// design-persistentvolumeops.md §8.1 is the specification. + +package volume + +// The six reasons every Ops kind in the group carries, which is what makes a +// dashboard, an alert, and a runbook writable once against the category rather +// than once per kind (design-crd-model.md §3.3). +const ( + // ReasonOperationQueued reports the phase's second meaning rather than the + // phase. Pending is where an operation starts, so an operation nobody has + // reconciled yet and one waiting on a lock it cannot take are the same + // value in status.phase, and this event is the only thing that separates + // them. + ReasonOperationQueued = "OperationQueued" + + ReasonOperationStarted = "OperationStarted" + ReasonOperationSucceeded = "OperationSucceeded" + ReasonOperationFailed = "OperationFailed" + ReasonOperationAborted = "OperationAborted" + + ReasonStepDeadlineExceeded = "StepDeadlineExceeded" +) + +// The kind's own reasons, which say what happened to the migration rather than +// what happened to the operation. +const ( + // ReasonClusterUnresolvable is a volume that cannot be addressed at all: + // deleted, replaced, or never provisioned by this driver. Admission refuses + // that at create, so reaching it means the volume became unaddressable + // afterward. + ReasonClusterUnresolvable = "ClusterUnresolvable" + + // ReasonTargetNodeNotReady and ReasonTargetNodeIsSource are facts about + // now rather than about the request, which is why neither is an admission + // rejection: the node a drain fans fifty migrations out to may well be + // online by the time the fifteenth of them acquires its lock. + ReasonTargetNodeNotReady = "TargetNodeNotReady" + ReasonTargetNodeIsSource = "TargetNodeIsSource" + + ReasonMigrationCreated = "MigrationCreated" + ReasonMigrationStarted = "MigrationStarted" + + // ReasonWaitingForConsumer is a pod that references one of the subsystem's + // claims and has not started. It is waited for rather than skipped: a pod + // that stages against the source mid-migration is stranded at cutover + // exactly like an established one. + ReasonWaitingForConsumer = "WaitingForConsumer" + + ReasonValidationStarted = "ValidationStarted" + + // ReasonValidationSkipped is a subsystem no host consumes, which has no + // paths to check anywhere. + ReasonValidationSkipped = "ValidationSkipped" + + // ReasonReleasingPaths is the cleanup an abandoned migration owes every + // host that took part in it. + ReasonReleasingPaths = "ReleasingPaths" + + // ReasonCleanupBlocked is a delete held open because the cleanup has not + // finished. Leaving a visibly stuck object is the intended outcome: a path + // connected with nothing tracking it blocks every later migration of the + // volume, and has. + ReasonCleanupBlocked = "CleanupBlocked" +) diff --git a/operator/internal/controllers/volume/graphs.go b/operator/internal/controllers/volume/graphs.go index d6cca0f8b..b5ad54792 100644 --- a/operator/internal/controllers/volume/graphs.go +++ b/operator/internal/controllers/volume/graphs.go @@ -123,6 +123,23 @@ func graphs(members int32) statemachine.MultiConfig[step] { } } +// initialDeadline is the budget of the step every operation is born in. A +// machine is already in its initial state when it is built, so that state's +// OnEnter never runs and the graph's deadline for it is never set. Setting it +// explicitly is what stops the first step from being the one step that cannot +// time out. +const initialDeadline = validatingDeadline + +// stepBudgets is what each step's deadline was set from, which is the other +// half of the arithmetic that measures how long a step took: the deadline is in +// status and the start is not, so the start is the deadline less the budget. +// Migrating is absent because its budget is not a constant, and copyDeadline is +// what answers for it. +var stepBudgets = map[step]time.Duration{ + stepValidating: validatingDeadline, + stepVerifying: verifyingDeadline, +} + // UnabortableSteps are the declared steps an abort cannot be honored from, // sorted. // diff --git a/operator/internal/controllers/volume/helpers_test.go b/operator/internal/controllers/volume/helpers_test.go new file mode 100644 index 000000000..e7e1e3adf --- /dev/null +++ b/operator/internal/controllers/volume/helpers_test.go @@ -0,0 +1,204 @@ +// The scaffolding this package's unit tests share: a scheme, a fake client, a +// fake control plane, and the objects a migration needs to exist. +// +// The envtest apiserver in suite_test.go is only for the CRD's CEL rules. +// Everything else here runs against a fake client, which is what makes a whole +// migration drivable without a control plane and without a cluster. + +package volume + +import ( + "context" + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/runtime" + "k8s.io/apimachinery/pkg/types" + k8sscheme "k8s.io/client-go/kubernetes/scheme" + "k8s.io/client-go/tools/events" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + "github.com/simplyblock/atlas/controlplane" + "github.com/simplyblock/atlas/lvol" + + simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +const ( + testNamespace = "simplyblock" + testClusterCR = "production" + testClusterID = "11111111-1111-1111-1111-111111111111" + testPoolID = "22222222-2222-2222-2222-222222222222" + testVolumeID = "33333333-3333-3333-3333-333333333333" + testMigrationID = "44444444-4444-4444-4444-444444444444" + testTargetID = "55555555-5555-5555-5555-555555555555" + testSourceID = "66666666-6666-6666-6666-666666666666" + + testPVName = "pvc-" + testVolumeID + testTargetNode = "worker-5" + testNQN = "nqn.2023-02.io.simplyblock:" + testClusterID + ":lvol:" + testVolumeID + testOpsName = "move-1" +) + +func testScheme(t *testing.T) *runtime.Scheme { + t.Helper() + s := runtime.NewScheme() + for _, add := range []func(*runtime.Scheme) error{ + k8sscheme.AddToScheme, + simplyblockv1alpha1.AddToScheme, + simplyblockv1alpha2.AddToScheme, + } { + if err := add(s); err != nil { + t.Fatalf("build the scheme: %v", err) + } + } + return s +} + +func testClient(t *testing.T, objs ...client.Object) client.Client { + t.Helper() + return fake.NewClientBuilder(). + WithScheme(testScheme(t)). + WithStatusSubresource( + &simplyblockv1alpha2.PersistentVolumeOps{}, + &simplyblockv1alpha2.StorageCluster{}, + &simplyblockv1alpha2.StorageNode{}, + ). + WithObjects(objs...). + Build() +} + +// testReconciler wires a reconciler onto a fake world. The API is the fake +// control plane, so every step is drivable without an HTTP server. +func testReconciler(t *testing.T, api MigrationClient, objs ...client.Object) *PersistentVolumeOpsReconciler { + t.Helper() + c := testClient(t, objs...) + return &PersistentVolumeOpsReconciler{ + Client: c, + Reader: c, + Scheme: testScheme(t), + Recorder: events.NewFakeRecorder(256), + API: api, + } +} + +// testOperation is a well-formed migration of the test volume to the test node. +func testOperation() *simplyblockv1alpha2.PersistentVolumeOps { + return &simplyblockv1alpha2.PersistentVolumeOps{ + ObjectMeta: metav1.ObjectMeta{ + Name: testOpsName, + UID: types.UID("uid-" + testOpsName), + Finalizers: []string{opsFinalizer}, + }, + Spec: simplyblockv1alpha2.PersistentVolumeOpsSpec{ + PersistentVolumeName: testPVName, + Action: simplyblockv1alpha2.PersistentVolumeOpsActionMigrate, + Migrate: &simplyblockv1alpha2.MigrateVolumeSpec{ + TargetNodeRef: simplyblockv1alpha2.StorageNodeReference{ + Namespace: testNamespace, + Name: testTargetNode, + }, + }, + }, + } +} + +// testVolumeObject is the PersistentVolume the operation acts on, provisioned +// by this driver and carrying the handle the cluster, pool, and volume are all +// read out of. +func testVolumeObject() *corev1.PersistentVolume { + return &corev1.PersistentVolume{ + ObjectMeta: metav1.ObjectMeta{Name: testPVName}, + Spec: corev1.PersistentVolumeSpec{ + PersistentVolumeSource: corev1.PersistentVolumeSource{ + CSI: &corev1.CSIPersistentVolumeSource{ + Driver: CSIDriverName, + VolumeHandle: string(lvol.NewVolumeHandle(testClusterID, testPoolID, testVolumeID)), + }, + }, + }, + } +} + +func testClusterObject() *simplyblockv1alpha2.StorageCluster { + return &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{Name: testClusterCR, Namespace: testNamespace}, + Status: simplyblockv1alpha2.StorageClusterStatus{UUID: testClusterID}, + } +} + +func testNodeObject() *simplyblockv1alpha2.StorageNode { + return &simplyblockv1alpha2.StorageNode{ + ObjectMeta: metav1.ObjectMeta{Name: testTargetNode, Namespace: testNamespace}, + Spec: simplyblockv1alpha2.StorageNodeSpec{ClusterRef: testClusterCR}, + Status: simplyblockv1alpha2.StorageNodeStatus{ + UUID: testTargetID, + Phase: simplyblockv1alpha2.StorageNodePhaseOnline, + }, + } +} + +// fakeControlPlane answers the migration calls from fields a test sets, and +// counts what it was asked to do. The counts are what make an idempotence +// assertion possible: a step that ran twice and called once is the property +// under test. +type fakeControlPlane struct { + volume lvol.Volume + volumeErr error + + members []lvol.Volume + membersErr error + + created controlplane.Migration + createErr error + creates int + continues int + cancels int + continueErr error + cancelErr error + + // read is what GetMigration answers with, which is how a test drives the + // copy forward: the step completes on the migration's reported state + // rather than on the return of the call that started it. + read controlplane.Migration + readErr error + reads int +} + +func (f *fakeControlPlane) Volume(context.Context, lvol.VolumeHandle) (lvol.Volume, error) { + return f.volume, f.volumeErr +} + +func (f *fakeControlPlane) SubsystemVolumes(context.Context, string, string) ([]lvol.Volume, error) { + return f.members, f.membersErr +} + +func (f *fakeControlPlane) CreateMigration( + context.Context, string, string, string, +) (controlplane.Migration, error) { + f.creates++ + if f.createErr != nil { + return controlplane.Migration{}, f.createErr + } + return f.created, nil +} + +func (f *fakeControlPlane) GetMigration( + context.Context, string, string, string, +) (controlplane.Migration, error) { + f.reads++ + return f.read, f.readErr +} + +func (f *fakeControlPlane) ContinueMigration(context.Context, string, string, string) error { + f.continues++ + return f.continueErr +} + +func (f *fakeControlPlane) CancelMigration(context.Context, string, string, string) error { + f.cancels++ + return f.cancelErr +} diff --git a/operator/internal/controllers/volume/jobs.go b/operator/internal/controllers/volume/jobs.go new file mode 100644 index 000000000..b297cc9f3 --- /dev/null +++ b/operator/internal/controllers/volume/jobs.go @@ -0,0 +1,609 @@ +// The host-side half of a migration: which nodes have to take part, and the +// Jobs that check, release, and clear their NVMe-oF paths. +// +// A migration moves an NVMe-oF subsystem rather than one volume inside it, so +// every volume on that subsystem moves at once and every host consuming one of +// them is affected at the same instant. Checking only the host of the volume +// the operation names would leave every sibling's consumer pointing at the +// source, and at cutover those hosts lose their volume. +// +// The work runs in a Job rather than here because it is host work: it reads the +// node's own NVMe fabric and connects paths on it, and neither is visible from +// the operator's pod. +// +// design-persistentvolumeops.md §5 is the specification. + +package volume + +import ( + "context" + "crypto/sha256" + "encoding/hex" + "encoding/json" + "fmt" + "regexp" + "sort" + "strings" + + batchv1 "k8s.io/api/batch/v1" + corev1 "k8s.io/api/core/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/types" + "sigs.k8s.io/controller-runtime/pkg/client" + + "github.com/simplyblock/atlas/lvol" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + vmigration "github.com/simplyblock/simplyblock-operator/internal/volumemigration" +) + +const ( + // migrationCtrlLossTimeout is how long the kernel keeps retrying a + // migration target path whose controller it has lost. It is the CSI + // driver's value rather than the control plane's, because a target path + // becomes the volume's data path at cutover and one volume's paths must not + // sit on two different timeouts depending on which of them last moved. + migrationCtrlLossTimeout = vmigration.CtrlLossTmoSec + + // validationJobDeadline caps a validation Job's whole life, scheduling and + // image pull included. Its purpose is to turn a Job that can never finish — + // an unschedulable pod, a node that is not ready — into a failure rather + // than a step parked forever. The check itself needs seconds. + validationJobDeadline = 180 + + // jobTTL is how long a finished Job is kept. Long enough to read, short + // enough that a drain's worth of them does not accumulate. + jobTTL = 3600 + + // consumerNotRunning is the marker a consumer lookup puts in its error when + // a pod references the claim and is not Running yet. It is a substring + // rather than a sentinel because the message names the claim, and what the + // caller needs is to tell "wait" from "failed." + consumerNotRunning = "is not running yet" +) + +// The three modes a host-side Job runs in. +const ( + modeValidate = "validate-migration" + modeRelease = "release-migration-paths" +) + +// consumingNodes returns every worker node that consumes a volume of the +// migrated subsystem, sorted so the Job set is stable across reconciles. +// +// Membership comes from the control plane and is mapped back to +// PersistentVolumes through the CSI handle, then to consumers through the +// claim. A member whose volume or consumer cannot be found contributes no node; +// a member with a consumer that is not Running yet stops the whole lookup, so +// the caller waits rather than validating a set it knows to be incomplete. +func (r *PersistentVolumeOpsReconciler) consumingNodes( + ctx context.Context, ops *simplyblockv1alpha2.PersistentVolumeOps, subject *subject, +) ([]string, error) { + nqn := ops.Status.Migration.SubsystemNQN + members, err := r.API.SubsystemVolumes(ctx, subject.clusterUUID, nqn) + if err != nil { + return nil, fmt.Errorf("list the volumes of subsystem %s: %w", nqn, err) + } + wanted := expectedMembers(members, subject.handle.VolumeID) + + volumes, err := r.volumesFronting(ctx, wanted) + if err != nil { + return nil, err + } + + nodes := map[string]struct{}{} + for _, pv := range volumes { + node, err := r.consumerNodeOf(ctx, pv) + if err != nil { + return nil, err + } + if node != "" { + nodes[node] = struct{}{} + } + } + + out := make([]string, 0, len(nodes)) + for node := range nodes { + out = append(out, node) + } + sort.Strings(out) + return out, nil +} + +// volumesFronting maps backend volume UUIDs to the PersistentVolumes that front +// them. A member with no PersistentVolume is not consumed through this +// cluster's driver, so it has no host paths to check. +func (r *PersistentVolumeOpsReconciler) volumesFronting( + ctx context.Context, volumeUUIDs map[string]struct{}, +) ([]*corev1.PersistentVolume, error) { + var list corev1.PersistentVolumeList + if err := r.Reader.List(ctx, &list); err != nil { + return nil, fmt.Errorf("list the persistent volumes: %w", err) + } + var out []*corev1.PersistentVolume + for i := range list.Items { + pv := &list.Items[i] + if pv.Spec.CSI == nil { + continue + } + handle, ok := lvol.ParseHandle(lvol.VolumeHandle(pv.Spec.CSI.VolumeHandle)) + if !ok { + continue + } + if _, wanted := volumeUUIDs[handle.VolumeID]; wanted { + out = append(out, pv) + } + } + return out, nil +} + +// consumerNodeOf returns the node running a pod that mounts this volume's +// claim, the empty string when nothing consumes it, and an error naming +// consumerNotRunning when a pod references the claim and has not started. +// +// The three answers are distinct because the caller does something different +// with each. A volume nothing consumes has no host paths to check. A volume +// whose consumer is coming has to be waited for. Everything else is a lookup +// that failed and is retried. +// +// The reads go through the uncached reader: a stale cache can miss a running +// consumer, and a missed consumer is a host that never gets the target's paths. +func (r *PersistentVolumeOpsReconciler) consumerNodeOf( + ctx context.Context, pv *corev1.PersistentVolume, +) (string, error) { + if pv.Spec.ClaimRef == nil { + return "", nil + } + claim, namespace := pv.Spec.ClaimRef.Name, pv.Spec.ClaimRef.Namespace + + var pods corev1.PodList + if err := r.Reader.List(ctx, &pods, client.InNamespace(namespace)); err != nil { + return "", fmt.Errorf("list the pods in namespace %s: %w", namespace, err) + } + + referenced := false + for i := range pods.Items { + pod := &pods.Items[i] + if !mounts(pod, claim) { + continue + } + referenced = true + if pod.Spec.NodeName != "" && pod.Status.Phase == corev1.PodRunning { + return pod.Spec.NodeName, nil + } + } + if referenced { + return "", fmt.Errorf("a consumer of claim %s/%s %s", namespace, claim, consumerNotRunning) + } + return "", nil +} + +// mounts reports whether the pod has this claim among its volumes. +func mounts(pod *corev1.Pod, claim string) bool { + for _, volume := range pod.Spec.Volumes { + if volume.PersistentVolumeClaim != nil && volume.PersistentVolumeClaim.ClaimName == claim { + return true + } + } + return false +} + +// startValidationJobs creates a Job on every consuming node that has none yet, +// and returns how many it started. +// +// Existing entries are kept, so it also serves the re-check before the cutover: +// the control plane lets a volume join the subsystem until the migration is +// activated, and a consumer can be rescheduled while the checks run, so a node +// can appear that was not there when the first round started. +func (r *PersistentVolumeOpsReconciler) startValidationJobs( + ctx context.Context, + ops *simplyblockv1alpha2.PersistentVolumeOps, + subject *subject, + nodes []string, +) (int, error) { + have := map[string]struct{}{} + for _, job := range ops.Status.Migration.ValidationJobs { + have[job.Node] = struct{}{} + } + + var missing []string + for _, node := range nodes { + if _, ok := have[node]; !ok { + missing = append(missing, node) + } + } + if len(missing) == 0 { + return 0, nil + } + + image, err := vmigration.JobImage(ctx, r.Client, subject.namespace(), subject.clusterUUID) + if err != nil { + return 0, err + } + + started := make([]simplyblockv1alpha2.ValidationJob, 0, len(missing)) + for _, node := range missing { + job := r.pathJob(ops, subject, node, image, modeValidate) + if err := r.Create(ctx, job); err != nil && !apierrors.IsAlreadyExists(err) { + return 0, fmt.Errorf("start the validation on node %s: %w", node, err) + } + started = append(started, simplyblockv1alpha2.ValidationJob{ + Namespace: job.Namespace, + Name: job.Name, + Node: node, + }) + } + + if err := r.writeStatus(ctx, ops, func(status *simplyblockv1alpha2.PersistentVolumeOpsStatus) { + status.Migration.ValidationJobs = append(status.Migration.ValidationJobs, started...) + }); err != nil { + return 0, err + } + + r.event(ops, corev1.EventTypeNormal, ReasonValidationStarted, + "Checking the target's paths for subsystem %s on %s", + ops.Status.Migration.SubsystemNQN, strings.Join(missing, ", ")) + return len(started), nil +} + +// validationJobsPassed reports whether every node's check has succeeded. +// +// The first failure is fatal to the operation rather than retried. A cutover is +// subsystem-wide, so continuing with a subset of the hosts ready guarantees an +// outage for the rest, and retrying the check would only delay that decision +// behind a second run of a check that just said no. +func (r *PersistentVolumeOpsReconciler) validationJobsPassed( + ctx context.Context, ops *simplyblockv1alpha2.PersistentVolumeOps, subject *subject, +) (bool, error) { + passed := false + pending := 0 + + jobs := ops.Status.Migration.ValidationJobs + for i := range jobs { + record := &jobs[i] + // A node whose check already passed is not looked at again: its Job is + // left for its own TTL to reap, and re-reading a reaped Job would look + // like a node that was never checked. + if record.Succeeded { + continue + } + + var job batchv1.Job + err := r.Get(ctx, types.NamespacedName{Namespace: record.Namespace, Name: record.Name}, &job) + switch { + case apierrors.IsNotFound(err): + // The Job went before a terminal state was observed — an eviction, + // or somebody deleting it. Forget it so the next pass starts that + // node over rather than waiting on a Job that no longer exists. + return false, r.forgetValidationJob(ctx, ops, record.Name) + case err != nil: + return false, fmt.Errorf("read the validation Job %s: %w", record.Name, err) + } + + switch jobOutcome(&job) { + case jobFailed: + return false, fatalf( + "the target's paths could not be established on node %s, so the cutover would "+ + "strand it; the migration is being taken back", record.Node) + case jobSucceeded: + record.Succeeded = true + passed = true + default: + pending++ + } + } + + // The passes are persisted before they are acted on: a restart must not + // re-run a check on a node that already passed. + if passed { + if err := r.writeStatus(ctx, ops, func(status *simplyblockv1alpha2.PersistentVolumeOpsStatus) { + status.Migration.ValidationJobs = jobs + }); err != nil { + return false, err + } + } + if pending > 0 { + return false, nil + } + + // Right before the point of no return, ask once more which nodes consume + // the subsystem. A node that appeared while the checks ran gets its own, + // and the step finishes on the pass after that. + late, err := r.consumingNodes(ctx, ops, subject) + if err != nil { + // Blocking a migration that is otherwise ready on a transient listing + // failure trades a certain delay for an uncertain gain. + return true, nil //nolint:nilerr // the validated set is what there is + } + started, err := r.startValidationJobs(ctx, ops, subject, late) + if err != nil { + return false, err + } + return started == 0, nil +} + +// jobOutcome reads a Job's terminal condition, which the Job controller sets. +type outcome int + +const ( + jobRunning outcome = iota + jobSucceeded + jobFailed +) + +func jobOutcome(job *batchv1.Job) outcome { + for _, condition := range job.Status.Conditions { + if condition.Status != corev1.ConditionTrue { + continue + } + switch condition.Type { + case batchv1.JobComplete: + return jobSucceeded + case batchv1.JobFailed: + return jobFailed + } + } + return jobRunning +} + +// forgetValidationJob drops one node's record so the next pass starts it over. +func (r *PersistentVolumeOpsReconciler) forgetValidationJob( + ctx context.Context, ops *simplyblockv1alpha2.PersistentVolumeOps, name string, +) error { + return r.writeStatus(ctx, ops, func(status *simplyblockv1alpha2.PersistentVolumeOpsStatus) { + kept := make([]simplyblockv1alpha2.ValidationJob, 0, len(status.Migration.ValidationJobs)) + for _, job := range status.Migration.ValidationJobs { + if job.Name != name { + kept = append(kept, job) + } + } + status.Migration.ValidationJobs = kept + }) +} + +// deleteValidationJobs removes the Jobs the checks ran in. They have served +// their purpose by the time this is called, and leaving them would mean the +// cleanup that follows raced pods still connecting paths. +func (r *PersistentVolumeOpsReconciler) deleteValidationJobs( + ctx context.Context, ops *simplyblockv1alpha2.PersistentVolumeOps, +) error { + if ops.Status.Migration == nil { + return nil + } + for _, record := range ops.Status.Migration.ValidationJobs { + job := &batchv1.Job{ + ObjectMeta: metav1.ObjectMeta{Namespace: record.Namespace, Name: record.Name}, + } + err := r.Delete(ctx, job, client.PropagationPolicy(metav1.DeletePropagationBackground)) + if err != nil && !apierrors.IsNotFound(err) { + return fmt.Errorf("delete the validation Job %s: %w", record.Name, err) + } + } + return nil +} + +// releaseOnEveryValidatedNode starts the release on every node that took part, +// for a migration that is being given up before the cutover. +// +// It closes the gap a per-node release cannot. The Job that fails releases its +// own paths on the way out, and the nodes whose checks passed exited +// successfully and are never told the migration was abandoned — by another +// node's failure, or by the operator giving up. Their target paths stay +// connected, retry a target that has stopped answering, and settle into the +// husk that blocks the subsystem's next migration. +// +// Best effort by design, and not waited on: the operation's outcome is already +// decided and must not become "still failing" because a cleanup Job is pending. +// What escapes is cleared by the reap the next validation runs first. +func (r *PersistentVolumeOpsReconciler) releaseOnEveryValidatedNode( + ctx context.Context, ops *simplyblockv1alpha2.PersistentVolumeOps, subject *subject, +) error { + migration := ops.Status.Migration + if migration == nil || len(migration.ValidationJobs) == 0 || len(migration.Connections) == 0 { + // Nothing was checked, so no host connected a target path on this + // operation's account. + return nil + } + + namespace, clusterUUID := migration.ValidationJobs[0].Namespace, migration.ClusterUUID + if subject != nil { + namespace = subject.namespace() + } + image, err := vmigration.JobImage(ctx, r.Client, namespace, clusterUUID) + if err != nil { + return err + } + + var nodes []string + for _, record := range migration.ValidationJobs { + job := r.releaseJob(ops, namespace, record.Node, image) + if err := r.Create(ctx, job); err != nil && !apierrors.IsAlreadyExists(err) { + return fmt.Errorf("release the target paths on node %s: %w", record.Node, err) + } + nodes = append(nodes, record.Node) + } + + r.event(ops, corev1.EventTypeNormal, ReasonReleasingPaths, + "Releasing the target paths of subsystem %s on %s", + migration.SubsystemNQN, strings.Join(nodes, ", ")) + return nil +} + +// reapOnEveryValidatedNode clears the husks a migration leaves on the hosts +// that took part, after the cutover, and reports whether every node is done. +// +// It is the release Job's mode, which reaps as well as releases, and running it +// here is safe for the reason that mode is: a release declines to touch a path +// that is serving, and after the cutover the target paths are the ones serving. +// What it does take down is a controller carrying no namespace at all, which is +// the state a path lost mid-check settles into and which blocks the subsystem's +// next migration until something clears it. +// +// Unlike the release on the failure path, this one is waited for. A cleanup +// that cannot finish has to be visible, and the step's deadline is what makes +// it so. +func (r *PersistentVolumeOpsReconciler) reapOnEveryValidatedNode( + ctx context.Context, ops *simplyblockv1alpha2.PersistentVolumeOps, subject *subject, +) (bool, error) { + migration := ops.Status.Migration + image, err := vmigration.JobImage(ctx, r.Client, subject.namespace(), migration.ClusterUUID) + if err != nil { + return false, err + } + + done := true + for _, record := range migration.ValidationJobs { + job := r.releaseJob(ops, subject.namespace(), record.Node, image) + + var existing batchv1.Job + err := r.Get(ctx, types.NamespacedName{Namespace: job.Namespace, Name: job.Name}, &existing) + switch { + case apierrors.IsNotFound(err): + if err := r.Create(ctx, job); err != nil && !apierrors.IsAlreadyExists(err) { + return false, fmt.Errorf("clear the paths on node %s: %w", record.Node, err) + } + done = false + case err != nil: + return false, fmt.Errorf("read the cleanup Job on node %s: %w", record.Node, err) + default: + switch jobOutcome(&existing) { + case jobFailed: + r.event(ops, corev1.EventTypeWarning, ReasonCleanupBlocked, + "The paths of subsystem %s on node %s could not be cleared", + migration.SubsystemNQN, record.Node) + return false, fmt.Errorf("the paths on node %s could not be cleared", record.Node) + case jobRunning: + done = false + } + } + } + return done, nil +} + +// pathJob builds one node-pinned Job running the given mode against that host's +// NVMe fabric. +// +// It carries no owner reference. A cluster-scoped object cannot own a +// namespaced one — the garbage collector treats such a reference as +// unresolvable and deletes the dependent — so these Jobs are deleted by the +// operation itself, on every path that ends it. +func (r *PersistentVolumeOpsReconciler) pathJob( + ops *simplyblockv1alpha2.PersistentVolumeOps, + subject *subject, + node, image, mode string, +) *batchv1.Job { + return r.modeJob(ops, subject.namespace(), node, image, mode, 0) +} + +// releaseJob is the cleanup counterpart, retried where the check is not: +// nothing downstream waits on a check that failed, so a transient failure there +// that is not retried is simply a path left connected, which is the leak. +func (r *PersistentVolumeOpsReconciler) releaseJob( + ops *simplyblockv1alpha2.PersistentVolumeOps, + namespace, node, image string, +) *batchv1.Job { + return r.modeJob(ops, namespace, node, image, modeRelease, 2) +} + +func (r *PersistentVolumeOpsReconciler) modeJob( + ops *simplyblockv1alpha2.PersistentVolumeOps, + namespace, node, image, mode string, + backoffLimit int32, +) *batchv1.Job { + migration := ops.Status.Migration + connections, _ := json.Marshal(hostConnections(migration.Connections)) + + return vmigration.BuildJob(vmigration.JobParams{ + Name: jobName(mode, ops.Name, node), + Namespace: namespace, + Hostname: node, + Image: image, + ContainerName: containerFor(mode), + Mode: mode, + Env: []corev1.EnvVar{ + {Name: "VMIG_CONNECTIONS", Value: string(connections)}, + // Which subsystem this node is expected to be connected to. + {Name: "VMIG_SUBSYSTEM_NQN", Value: migration.SubsystemNQN}, + // The host's sysfs, mounted into the Job: the container's own /sys + // is not the host's. + {Name: "VMIG_SYS_ROOT", Value: "/host/sys"}, + }, + BackoffLimit: backoffLimit, + TTL: jobTTL, + Deadline: validationJobDeadline, + }) +} + +func containerFor(mode string) string { + if mode == modeRelease { + return "nvme-release" + } + return "nvme-validate" +} + +// hostConnections renders the recorded paths into what the Job's binary reads. +func hostConnections( + conns []simplyblockv1alpha2.MigrationConnection, +) []vmigration.Connection { + out := make([]vmigration.Connection, 0, len(conns)) + for _, conn := range conns { + out = append(out, vmigration.Connection{ + NQN: conn.NQN, + IP: conn.Address, + Port: int(deref(conn.Port)), + Transport: conn.Transport, + NrIoQueues: int(deref(conn.NrIOQueues)), + ReconnectDelay: int(deref(conn.ReconnectDelaySeconds)), + CtrlLossTmo: int(deref(conn.CtrlLossTimeoutSeconds)), + FastIOFailTmo: int(deref(conn.FastIOFailTimeoutSeconds)), + KeepAliveTmo: int(deref(conn.KeepAliveTimeoutSeconds)), + }) + } + return out +} + +func deref(v *int32) int32 { + if v == nil { + return 0 + } + return *v +} + +// jobName is stable for one (mode, operation, node), which is what makes +// creating it idempotent: a pass that created the Job and crashed before +// recording it finds its own Job rather than making a second. +func jobName(mode, ops, node string) string { + prefix := "pvops-validate-" + if mode == modeRelease { + prefix = "pvops-release-" + } + return prefix + labelSafe(ops) + "-" + nodeSuffix(node) +} + +// nodeSuffix is a DNS-label-safe, collision-resistant suffix for a node name. +// Node names can be long fully qualified names and are not label-safe, so the +// short host part is kept for readability and a hash of the full name for +// uniqueness. +func nodeSuffix(node string) string { + sum := sha256.Sum256([]byte(node)) + short := labelSafe(strings.SplitN(node, ".", 2)[0]) + if len(short) > 16 { + short = short[:16] + } + if short == "" { + return hex.EncodeToString(sum[:6]) + } + return short + "-" + hex.EncodeToString(sum[:4]) +} + +// nonLabelChars matches everything not allowed inside a DNS-1123 label. +var nonLabelChars = regexp.MustCompile(`[^a-z0-9-]`) + +func labelSafe(s string) string { + s = nonLabelChars.ReplaceAllString(strings.ToLower(s), "") + if len(s) > 20 { + s = s[:20] + } + return s +} diff --git a/operator/internal/controllers/volume/lock.go b/operator/internal/controllers/volume/lock.go new file mode 100644 index 000000000..1b7e8cde9 --- /dev/null +++ b/operator/internal/controllers/volume/lock.go @@ -0,0 +1,134 @@ +// Mutual exclusion between two operations on one volume, through an annotation +// on the volume. +// +// Every other Ops kind in this group takes status.activeOpsRef on its target. A +// PersistentVolume is a core type this operator does not define and must not +// add fields to, so the lock moves from status to metadata and keeps everything +// else: acquisition is an optimistic-lock patch, release checks ownership, and +// release runs on every terminal path including deletion. +// +// What makes a lock outside the operation's own object safe is that it is +// self-describing. The risk in one is a holder that dies between taking the +// lock and recording that it took it, leaving a volume locked by nobody. That +// does not arise here, because the annotation's value is the holder's name and +// this kind is cluster-scoped, so the name identifies the holder completely: +// any reconciler can read it, get the named operation, and learn which of three +// situations it is in. +// +// design-persistentvolumeops.md §6 is the specification. + +package volume + +import ( + "context" + "fmt" + + corev1 "k8s.io/api/core/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" + "k8s.io/apimachinery/pkg/types" + "sigs.k8s.io/controller-runtime/pkg/client" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// acquireLock takes the volume's lock for this operation, and reports whether +// it now holds it. +// +// A lock another operation holds is waited on rather than failed, which is what +// lets this kind queue like the rest of the group: a drain fanning out fifty +// migrations across a handful of volumes wants them to proceed in turn, not to +// have forty-odd of them fail and need retrying by whatever issued them. +func (r *PersistentVolumeOpsReconciler) acquireLock( + ctx context.Context, + ops *simplyblockv1alpha2.PersistentVolumeOps, + pv *corev1.PersistentVolume, +) (bool, error) { + held := pv.Annotations[simplyblockv1alpha2.PersistentVolumeOpsLock] + if held == ops.Name { + return true, nil + } + + if held != "" { + takeable, err := r.lockIsStale(ctx, held) + if err != nil || !takeable { + return false, err + } + } + + // The optimistic lock is what makes this a lock at all: two reconcilers + // that both read the annotation free patch the same resourceVersion, and + // one of them gets a 409 and comes back to find the volume taken. It is + // also what keeps two reconcilers from both breaking a stale lock and both + // concluding they won. + patch := client.MergeFromWithOptions(pv.DeepCopy(), client.MergeFromWithOptimisticLock{}) + if pv.Annotations == nil { + pv.Annotations = map[string]string{} + } + pv.Annotations[simplyblockv1alpha2.PersistentVolumeOpsLock] = ops.Name + if err := r.Patch(ctx, pv, patch); err != nil { + if apierrors.IsConflict(err) { + // Somebody moved the volume between the read and the write. + // Whether that was another operation taking the lock is decided by + // reading it again rather than guessed at here. + return false, nil + } + return false, fmt.Errorf("take the lock on volume %s: %w", pv.Name, err) + } + return true, nil +} + +// lockIsStale reports whether a lock naming another operation may be broken. +// +// Two of the three situations the annotation can be in are stale. The named +// operation is terminal, so its release did not run — a crash between the two — +// and waiting for a release that will never come would block the volume +// forever. Or it does not exist at all, so it was deleted with its finalizer +// forced or removed out of band, and the annotation is the only thing left of +// it. The third is a live holder, which is waited on. +func (r *PersistentVolumeOpsReconciler) lockIsStale(ctx context.Context, held string) (bool, error) { + var holder simplyblockv1alpha2.PersistentVolumeOps + err := r.Get(ctx, types.NamespacedName{Name: held}, &holder) + switch { + case apierrors.IsNotFound(err): + return true, nil + case err != nil: + return false, fmt.Errorf("read the operation %s holding the lock: %w", held, err) + default: + return terminal(holder.Status.Phase), nil + } +} + +// releaseLock clears the volume's lock, and only while it still names this +// operation. +// +// The ownership check is the whole of the safety. A pass that started before +// the lock changed hands would otherwise clear a lock somebody else now holds, +// which is worse than not releasing at all: two operations would then be +// copying one logical volume to two places with neither of them knowing. +// +// A volume that is gone is not an error. It took its lock with it, which is the +// state being asked for, and this runs on the deletion path where a failure +// would hold the operation open forever. +func (r *PersistentVolumeOpsReconciler) releaseLock( + ctx context.Context, ops *simplyblockv1alpha2.PersistentVolumeOps, +) error { + var pv corev1.PersistentVolume + err := r.Get(ctx, types.NamespacedName{Name: ops.Spec.PersistentVolumeName}, &pv) + switch { + case apierrors.IsNotFound(err): + return nil + case err != nil: + return fmt.Errorf("read volume %s to release its lock: %w", ops.Spec.PersistentVolumeName, err) + } + + if pv.Annotations[simplyblockv1alpha2.PersistentVolumeOpsLock] != ops.Name { + return nil + } + + patch := client.MergeFromWithOptions(pv.DeepCopy(), client.MergeFromWithOptimisticLock{}) + delete(pv.Annotations, simplyblockv1alpha2.PersistentVolumeOpsLock) + if err := r.Patch(ctx, &pv, patch); err != nil && !apierrors.IsConflict(err) { + return fmt.Errorf("release the lock on volume %s: %w", pv.Name, err) + } + return nil +} diff --git a/operator/internal/controllers/volume/lock_test.go b/operator/internal/controllers/volume/lock_test.go new file mode 100644 index 000000000..23ae72b5d --- /dev/null +++ b/operator/internal/controllers/volume/lock_test.go @@ -0,0 +1,257 @@ +// The volume's lock, which is the one mechanism in this package whose failure +// is silent and expensive: two migrations of one volume at once would have two +// backend migrations copying the same logical volume to two places. +// +// Every other Ops kind takes status.activeOpsRef on its target. A +// PersistentVolume is a core type this operator must not add a field to, so the +// lock is an annotation — and an annotation is only a lock if it has the three +// properties the group requires of one. Each of them is a case below. + +package volume + +import ( + "context" + "testing" + + corev1 "k8s.io/api/core/v1" + "k8s.io/apimachinery/pkg/types" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// volumeFrom reads the volume the way a reconcile does. The resource version +// is what the optimistic lock patches against, so a lock taken on a +// hand-built copy would not be taking one at all. +func volumeFrom(t *testing.T, r *PersistentVolumeOpsReconciler) *corev1.PersistentVolume { + t.Helper() + var pv corev1.PersistentVolume + if err := r.Get(context.Background(), types.NamespacedName{Name: testPVName}, &pv); err != nil { + t.Fatalf("reading the volume: %v", err) + } + return &pv +} + +// testHolderName is the other operation every contention case here is against. +const testHolderName = "move-0" + +// lockOn reads the annotation back off the volume. +func lockOn(t *testing.T, r *PersistentVolumeOpsReconciler) string { + t.Helper() + var pv corev1.PersistentVolume + if err := r.Get(context.Background(), types.NamespacedName{Name: testPVName}, &pv); err != nil { + t.Fatalf("reading the volume back: %v", err) + } + return pv.Annotations[simplyblockv1alpha2.PersistentVolumeOpsLock] +} + +// A free volume is taken, and the annotation names the operation that took it. +func TestTheLockIsTakenWhenTheVolumeIsFree(t *testing.T) { + ops := testOperation() + r := testReconciler(t, &fakeControlPlane{}, ops, testVolumeObject()) + + acquired, err := r.acquireLock(context.Background(), ops, volumeFrom(t, r)) + if err != nil { + t.Fatal(err) + } + if !acquired { + t.Fatal("the operation did not take a lock nobody holds") + } + if got := lockOn(t, r); got != testOpsName { + t.Errorf("the volume's lock names %q, want %q", got, testOpsName) + } +} + +// Re-acquiring a lock this operation already holds is not a second acquisition. +// Every pass takes the lock again, so a reconcile that treated its own lock as +// somebody else's would never get past Pending. +func TestTheHolderReadoptsItsOwnLock(t *testing.T) { + ops := testOperation() + pv := testVolumeObject() + pv.Annotations = map[string]string{ + simplyblockv1alpha2.PersistentVolumeOpsLock: testOpsName, + } + r := testReconciler(t, &fakeControlPlane{}, ops, pv) + + acquired, err := r.acquireLock(context.Background(), ops, volumeFrom(t, r)) + if err != nil { + t.Fatal(err) + } + if !acquired { + t.Error("the operation could not re-adopt the lock it already holds") + } +} + +// A lock held by a live operation is waited on rather than taken. Failing here +// instead would make the order two people applied two objects in decide which +// of them runs. +func TestALockHeldByALiveOperationIsWaitedOn(t *testing.T) { + holder := testOperation() + holder.Name = testHolderName + holder.Status.Phase = simplyblockv1alpha2.PersistentVolumeOpsPhaseRunning + + ops := testOperation() + pv := testVolumeObject() + pv.Annotations = map[string]string{ + simplyblockv1alpha2.PersistentVolumeOpsLock: holder.Name, + } + r := testReconciler(t, &fakeControlPlane{}, holder, ops, pv) + + acquired, err := r.acquireLock(context.Background(), ops, volumeFrom(t, r)) + if err != nil { + t.Fatal(err) + } + if acquired { + t.Fatal("the operation took a lock another running operation holds") + } + if got := lockOn(t, r); got != holder.Name { + t.Errorf("the volume's lock now names %q, want the holder %q", got, holder.Name) + } +} + +// A lock whose holder finished is broken rather than waited out. The release +// runs on every terminal path, so a lock still naming a terminal operation +// means the release did not run — a crash between the two — and waiting for a +// release that will never come would block the volume forever. +func TestALockHeldByATerminalOperationIsBroken(t *testing.T) { + for _, phase := range []simplyblockv1alpha2.PersistentVolumeOpsPhase{ + simplyblockv1alpha2.PersistentVolumeOpsPhaseSucceeded, + simplyblockv1alpha2.PersistentVolumeOpsPhaseFailed, + simplyblockv1alpha2.PersistentVolumeOpsPhaseAborted, + } { + t.Run(string(phase), func(t *testing.T) { + holder := testOperation() + holder.Name = testHolderName + holder.Status.Phase = phase + + ops := testOperation() + pv := testVolumeObject() + pv.Annotations = map[string]string{ + simplyblockv1alpha2.PersistentVolumeOpsLock: holder.Name, + } + r := testReconciler(t, &fakeControlPlane{}, holder, ops, pv) + + acquired, err := r.acquireLock(context.Background(), ops, volumeFrom(t, r)) + if err != nil { + t.Fatal(err) + } + if !acquired { + t.Fatal("the operation waited on a lock whose holder had finished") + } + if got := lockOn(t, r); got != testOpsName { + t.Errorf("the volume's lock names %q, want %q", got, testOpsName) + } + }) + } +} + +// A lock naming an operation that no longer exists is broken too. That is the +// object deleted with its finalizer forced or removed out of band, and the +// annotation is the only thing left of it. +func TestALockHeldByNobodyIsBroken(t *testing.T) { + ops := testOperation() + pv := testVolumeObject() + pv.Annotations = map[string]string{ + simplyblockv1alpha2.PersistentVolumeOpsLock: testHolderName, + } + r := testReconciler(t, &fakeControlPlane{}, ops, pv) + + acquired, err := r.acquireLock(context.Background(), ops, volumeFrom(t, r)) + if err != nil { + t.Fatal(err) + } + if !acquired { + t.Fatal("the operation waited on a lock held by an operation that does not exist") + } + if got := lockOn(t, r); got != testOpsName { + t.Errorf("the volume's lock names %q, want %q", got, testOpsName) + } +} + +// Release clears the annotation, and only while it still names the releaser. A +// late release that cleared somebody else's lock would be worse than not +// releasing at all: two operations would then run against one volume with +// neither of them knowing. +func TestReleaseClearsOnlyThisOperationsLock(t *testing.T) { + t.Run("its own lock is cleared", func(t *testing.T) { + ops := testOperation() + pv := testVolumeObject() + pv.Annotations = map[string]string{ + simplyblockv1alpha2.PersistentVolumeOpsLock: testOpsName, + } + r := testReconciler(t, &fakeControlPlane{}, ops, pv) + + if err := r.releaseLock(context.Background(), ops); err != nil { + t.Fatal(err) + } + if got := lockOn(t, r); got != "" { + t.Errorf("the volume is still locked by %q", got) + } + }) + + t.Run("somebody else's lock is left alone", func(t *testing.T) { + ops := testOperation() + pv := testVolumeObject() + pv.Annotations = map[string]string{ + simplyblockv1alpha2.PersistentVolumeOpsLock: testHolderName, + } + r := testReconciler(t, &fakeControlPlane{}, ops, pv) + + if err := r.releaseLock(context.Background(), ops); err != nil { + t.Fatal(err) + } + if got := lockOn(t, r); got != testHolderName { + t.Errorf("the volume's lock is now %q, want it untouched", got) + } + }) +} + +// A volume that is gone takes its lock with it, which is correct for the lock +// and says nothing about the copy: the backing logical volume outlives the +// Kubernetes object. Releasing is what must not fail here, because the release +// runs on the deletion path and a failure would hold the object open forever. +func TestReleasingALockOnAVolumeThatIsGoneSucceeds(t *testing.T) { + ops := testOperation() + r := testReconciler(t, &fakeControlPlane{}, ops) + + if err := r.releaseLock(context.Background(), ops); err != nil { + t.Errorf("releasing the lock on a deleted volume failed: %v", err) + } +} + +// The annotation is under the group's own prefix, which is what makes it +// identifiable at a glance on an object this API group does not own, and +// matchable by one RBAC or admission rule. +func TestTheLockKeyCarriesTheGroupsPrefix(t *testing.T) { + const prefix = "storage.simplyblock.io/" + if key := simplyblockv1alpha2.PersistentVolumeOpsLock; len(key) <= len(prefix) || + key[:len(prefix)] != prefix { + t.Errorf("the lock key is %q, which does not carry %q", key, prefix) + } +} + +// The lock touches neither spec nor status, so it cannot conflict with the +// provisioner or the volume's own controllers. +func TestTakingTheLockLeavesTheVolumeOtherwiseUntouched(t *testing.T) { + ops := testOperation() + pv := testVolumeObject() + pv.Labels = map[string]string{"kept": "yes"} + r := testReconciler(t, &fakeControlPlane{}, ops, pv) + + if _, err := r.acquireLock(context.Background(), ops, volumeFrom(t, r)); err != nil { + t.Fatal(err) + } + + var got corev1.PersistentVolume + if err := r.Get(context.Background(), types.NamespacedName{Name: testPVName}, &got); err != nil { + t.Fatal(err) + } + if got.Labels["kept"] != "yes" { + t.Errorf("labels = %v, want the ones the volume had", got.Labels) + } + if got.Spec.CSI == nil || got.Spec.CSI.VolumeHandle != pv.Spec.CSI.VolumeHandle { + t.Errorf("the volume's spec changed: %+v", got.Spec.CSI) + } + if got.Status.Phase != pv.Status.Phase || got.Status.Message != pv.Status.Message { + t.Errorf("the volume's status changed: %+v", got.Status) + } +} diff --git a/operator/internal/controllers/volume/metrics.go b/operator/internal/controllers/volume/metrics.go new file mode 100644 index 000000000..e531c2c0a --- /dev/null +++ b/operator/internal/controllers/volume/metrics.go @@ -0,0 +1,147 @@ +// The time series a volume operation exports. +// +// They are declared apart from the controller that fills them for the reason +// every metrics file in this repository is: a collector is process-global state +// registered once at init, while a reconciler is constructed per manager and +// may be constructed more than once in a test. +// +// The one that would have found the production defect is the per-step +// histogram. A migration whose creation is fast and whose copy is fast, and +// which takes half an hour, is a migration spending its time somewhere the +// registered kind's merged phase enum could not name. +// +// design-persistentvolumeops.md §8.2 is the specification, and +// design-crd-model.md §7.12 is the naming rule the suffixes come from. + +package volume + +import ( + "time" + + "github.com/prometheus/client_golang/prometheus" + ctrlmetrics "sigs.k8s.io/controller-runtime/pkg/metrics" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// migrationBuckets span a second to roughly half a day. A migration of an idle +// single-volume subsystem completes in seconds, and one of a large +// many-member subsystem runs for hours, and one set has to hold both. +var migrationBuckets = prometheus.ExponentialBuckets(1, 4, 9) + +var ( + operationsTotal = prometheus.NewCounterVec( + prometheus.CounterOpts{ + Name: "simplyblock_persistentvolume_operations_total", + Help: "Volume operations that reached a terminal phase, by action and result.", + }, + []string{"cluster", "action", "result"}, + ) + + operationDuration = prometheus.NewHistogramVec( + prometheus.HistogramOpts{ + Name: "simplyblock_persistentvolume_operation_duration_seconds", + Help: "How long a volume operation took from creation to a terminal phase.", + Buckets: migrationBuckets, + }, + []string{"cluster", "action"}, + ) + + // Where a slow migration is actually slow, which the registered kind's + // merged phase enum could not say. + stepDuration = prometheus.NewHistogramVec( + prometheus.HistogramOpts{ + Name: "simplyblock_persistentvolume_operation_step_duration_seconds", + Help: "How long each step of a volume operation took.", + Buckets: migrationBuckets, + }, + []string{"cluster", "action", "step"}, + ) + + stepDeadlinesExceeded = prometheus.NewCounterVec( + prometheus.CounterOpts{ + Name: "simplyblock_persistentvolume_operation_step_deadline_exceeded_total", + Help: "Steps of a volume operation that ran out of time.", + }, + []string{"cluster", "action", "step"}, + ) + + // How long operations wait behind another operation on the same volume, + // which is only expressible because the volume carries a real lock. + operationQueuedSeconds = prometheus.NewHistogramVec( + prometheus.HistogramOpts{ + Name: "simplyblock_persistentvolume_operation_queued_seconds", + Help: "How long a volume operation was held behind another on the same volume.", + Buckets: migrationBuckets, + }, + []string{"cluster"}, + ) +) + +func init() { + ctrlmetrics.Registry.MustRegister( + operationsTotal, + operationDuration, + stepDuration, + stepDeadlinesExceeded, + operationQueuedSeconds, + ) +} + +// observeOperation records a terminal operation, how long it took, and how long +// it spent waiting for the volume before it started. +// +// An operation that never started has no duration to report, and reporting zero +// would drag the distribution toward a value nothing took. +func (r *PersistentVolumeOpsReconciler) observeOperation( + ops *simplyblockv1alpha2.PersistentVolumeOps, + phase simplyblockv1alpha2.PersistentVolumeOpsPhase, +) { + cluster, action := clusterOf(ops), string(ops.Spec.Action) + operationsTotal.WithLabelValues(cluster, action, resultOf(phase)).Inc() + + if started := ops.Status.StartedAt; started != nil { + operationDuration.WithLabelValues(cluster, action). + Observe(time.Since(started.Time).Seconds()) + if deferred := ops.Status.DeferredSince; deferred != nil { + operationQueuedSeconds.WithLabelValues(cluster). + Observe(started.Sub(deferred.Time).Seconds()) + } + } +} + +// observeStep records how long the step that just finished took. +func (r *PersistentVolumeOpsReconciler) observeStep( + ops *simplyblockv1alpha2.PersistentVolumeOps, subject *subject, current step, +) { + started, known := stepStarted(ops, current) + if !known { + return + } + stepDuration.WithLabelValues(subject.clusterUUID, string(ops.Spec.Action), string(current)). + Observe(time.Since(started).Seconds()) +} + +// clusterOf is the cluster label, which is empty for an operation that failed +// before it could resolve one. That is a value worth having rather than +// dropping: an operation that failed for that reason is exactly the kind +// somebody wants a rate of. +func clusterOf(ops *simplyblockv1alpha2.PersistentVolumeOps) string { + if ops.Status.Migration == nil { + return "" + } + return ops.Status.Migration.ClusterUUID +} + +// resultOf is the metric label for a terminal phase, lowercased because a label +// value is not an API enum. +func resultOf(phase simplyblockv1alpha2.PersistentVolumeOpsPhase) string { + switch phase { + case simplyblockv1alpha2.PersistentVolumeOpsPhaseSucceeded: + return "succeeded" + case simplyblockv1alpha2.PersistentVolumeOpsPhaseAborted: + return "aborted" + default: + return "failed" + } +} diff --git a/operator/internal/controllers/volume/persistentvolumeops_controller.go b/operator/internal/controllers/volume/persistentvolumeops_controller.go new file mode 100644 index 000000000..7e3b9dbb5 --- /dev/null +++ b/operator/internal/controllers/volume/persistentvolumeops_controller.go @@ -0,0 +1,656 @@ +// The PersistentVolumeOps reconciler: it drives one migration of one volume to +// a terminal phase and leaves the object behind as the audit record. +// +// Two machines run, not one. The outer phase — Pending, Running, and the three +// terminal values — is identical for every action, so folding it into the +// action's graph would copy that spine once per action and a later fix would +// land in one copy (design-crd-model.md §3.1). The inner one is the action's +// steps, declared in graphs.go. +// +// Nothing here blocks. One reconcile advances at most one step: it asks whether +// the current step has finished, and either requeues or writes the next step +// down and enters it. A step that has not finished is waiting on something +// outside this process — a data copy, a Job on another node — and waiting for +// it inline would hold a worker for as long as the copy takes. +// +// The registered VolumeMigration reconciler runs beside this one. The two never +// act on the same object: they reconcile different kinds, and nothing creates +// one of each for a volume. What they do share is the volume, and this one +// alone takes a lock on it, which is the cost of the coexistence and the reason +// in-flight migrations are drained on the old kind rather than converted. +// +// design-persistentvolumeops.md is the specification. + +package volume + +import ( + "context" + "errors" + "fmt" + "reflect" + "time" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/runtime" + "k8s.io/apimachinery/pkg/types" + "k8s.io/client-go/tools/events" + "k8s.io/client-go/util/retry" + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/builder" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil" + "sigs.k8s.io/controller-runtime/pkg/event" + "sigs.k8s.io/controller-runtime/pkg/handler" + logf "sigs.k8s.io/controller-runtime/pkg/log" + "sigs.k8s.io/controller-runtime/pkg/predicate" + "sigs.k8s.io/controller-runtime/pkg/reconcile" + + "github.com/simplyblock/atlas/controlplane" + "github.com/simplyblock/atlas/lvol" + "github.com/simplyblock/atlas/statemachine" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/driver" +) + +const ( + // opsFinalizer is what stops an operation deleted mid-flight from leaving + // the volume locked and its validation Jobs running with nothing left to + // account for them. The admission guard refuses a delete from the one step + // where the work cannot be taken back; what reaches the finalizer is an + // operation whose product this can unwind. + opsFinalizer = "storage.simplyblock.io/persistentvolumeops-finalizer" + + // opsRetry is how long an operation waits before looking again at + // something it cannot hurry: a lock another operation holds, a control + // plane that is not accepting migrations, or a Job still running. + opsRetry = 15 * time.Second + + // opsAdvance is how long a pass that moved the operation forward waits + // before the next one. It is short because there is nothing to wait for: + // the status write this pass made is itself a change the controller + // watches, so this is the backstop for the event rather than the path the + // next step normally arrives on. + opsAdvance = time.Second + + // CSIDriverName is the driver whose volumes this operator can move. A + // deployment that renamed its driver is matched against that name instead; + // this is the default and the answer when no SimplyblockDriver exists. + CSIDriverName = driver.DefaultDriverName + + // PersistentVolumeNameField indexes operations by the volume they act on. + // Releasing a lock has to wake whatever was waiting on it, and the index is + // what makes that a lookup rather than a listing of every operation in the + // cluster on every release. + PersistentVolumeNameField = ".spec.persistentVolumeName" +) + +// MigrationClient is the control-plane surface a migration needs. +// +// It is an interface so that a test can drive the whole graph without an HTTP +// server, and it is declared here rather than in atlas-lib because what a +// migration needs is the operator's question rather than the client's. +type MigrationClient interface { + // Volume resolves the subsystem the migration is addressed by: the control + // plane migrates a subsystem rather than one volume inside it. + Volume(ctx context.Context, handle lvol.VolumeHandle) (lvol.Volume, error) + + // SubsystemVolumes is how the sibling volumes are found. Each of them has + // a consuming host that must reach the target before the cutover, because + // at cutover every member moves at once. + SubsystemVolumes(ctx context.Context, clusterID, nqn string) ([]lvol.Volume, error) + + CreateMigration(ctx context.Context, clusterID, nqn, targetNodeID string) (controlplane.Migration, error) + ContinueMigration(ctx context.Context, clusterID, nqn, migrationID string) error + GetMigration(ctx context.Context, clusterID, nqn, migrationID string) (controlplane.Migration, error) + CancelMigration(ctx context.Context, clusterID, nqn, migrationID string) error +} + +// PersistentVolumeOpsReconciler reconciles a PersistentVolumeOps. +type PersistentVolumeOpsReconciler struct { + client.Client + Scheme *runtime.Scheme + Recorder events.EventRecorder + API MigrationClient + + // Reader is uncached, and the consumer lookup is what it is for: a stale + // informer cache can miss a pod that is genuinely running, and a volume + // whose consumer was missed is one whose host never gets the target's + // paths and loses its volume at cutover. + Reader client.Reader +} + +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=persistentvolumeops,verbs=get;list;watch;create;update;patch;delete +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=persistentvolumeops/status,verbs=get;update;patch +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=persistentvolumeops/finalizers,verbs=update +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storageclusters;storagenodes;simplyblockdrivers,verbs=get;list;watch +// +kubebuilder:rbac:groups="",resources=persistentvolumes,verbs=get;list;watch;patch +// +kubebuilder:rbac:groups="",resources=persistentvolumeclaims,verbs=get;list;watch +// +kubebuilder:rbac:groups="",resources=pods,verbs=get;list;watch +// +kubebuilder:rbac:groups="",resources=events,verbs=create;patch +// +kubebuilder:rbac:groups=batch,resources=jobs,verbs=get;list;watch;create;delete + +// SetupWithManager registers the controller, the field index the queue reads, +// and the watch that wakes a queued operation when the volume's lock frees. +// +// The mapping is what makes the queue move. An operation waiting on a lock has +// nothing of its own to react to, so a controller watching only its own kind +// would leave every queued operation waiting out a requeue interval after the +// lock frees (design-crd-model.md §3.2). +func (r *PersistentVolumeOpsReconciler) SetupWithManager(mgr ctrl.Manager) error { + if err := mgr.GetFieldIndexer().IndexField( + context.Background(), + &simplyblockv1alpha2.PersistentVolumeOps{}, + PersistentVolumeNameField, + func(object client.Object) []string { + ops, ok := object.(*simplyblockv1alpha2.PersistentVolumeOps) + if !ok { + return nil + } + return []string{ops.Spec.PersistentVolumeName} + }, + ); err != nil { + return fmt.Errorf("index operations by the volume they act on: %w", err) + } + + if r.Reader == nil { + r.Reader = mgr.GetAPIReader() + } + + return ctrl.NewControllerManagedBy(mgr). + For(&simplyblockv1alpha2.PersistentVolumeOps{}). + Named("persistentvolumeops"). + Watches(&corev1.PersistentVolume{}, + handler.EnqueueRequestsFromMapFunc(r.operationsOn), + // A volume's own updates are frequent and mostly about capacity and + // phase. What this watch is for is the lock changing hands and the + // volume being deleted, and both are visible in metadata. + builder.WithPredicates(volumeLockChanged{})). + Complete(r) +} + +// operationsOn enqueues every operation naming this volume. +func (r *PersistentVolumeOpsReconciler) operationsOn( + ctx context.Context, pv client.Object, +) []reconcile.Request { + var operations simplyblockv1alpha2.PersistentVolumeOpsList + if err := r.List(ctx, &operations, + client.MatchingFields{PersistentVolumeNameField: pv.GetName()}); err != nil { + return nil + } + requests := make([]reconcile.Request, 0, len(operations.Items)) + for i := range operations.Items { + requests = append(requests, reconcile.Request{ + NamespacedName: client.ObjectKeyFromObject(&operations.Items[i]), + }) + } + return requests +} + +func (r *PersistentVolumeOpsReconciler) Reconcile( + ctx context.Context, req ctrl.Request, +) (ctrl.Result, error) { + var ops simplyblockv1alpha2.PersistentVolumeOps + if err := r.Get(ctx, req.NamespacedName, &ops); err != nil { + return ctrl.Result{}, client.IgnoreNotFound(err) + } + + if !ops.DeletionTimestamp.IsZero() { + return r.teardown(ctx, &ops) + } + + // A terminal operation is a record rather than a task. Releasing the lock + // here as well as on the transition is what covers the pass that crashed + // between the two. + if terminal(ops.Status.Phase) { + return ctrl.Result{}, r.releaseLock(ctx, &ops) + } + + if !controllerutil.ContainsFinalizer(&ops, opsFinalizer) { + controllerutil.AddFinalizer(&ops, opsFinalizer) + return ctrl.Result{}, r.Update(ctx, &ops) + } + + subject, err := r.resolve(ctx, &ops) + switch { + case errors.Is(err, errVolumeGone): + // The operation follows its volume. A claim deleted under a Delete + // reclaim policy makes the driver delete the backing logical volume, + // and moving a volume that is being deleted is work nobody will read. + // Aborted rather than Failed, because a migration whose volume went + // away did not go wrong. + return r.abandon(ctx, &ops, err.Error()) + case err != nil: + var fatal *terminalStepError + if errors.As(err, &fatal) { + r.event(&ops, corev1.EventTypeWarning, ReasonClusterUnresolvable, "%s", fatal.Error()) + return r.finish(ctx, &ops, simplyblockv1alpha2.PersistentVolumeOpsPhaseFailed, fatal.Error()) + } + return ctrl.Result{RequeueAfter: opsRetry}, r.note(ctx, &ops, err.Error()) + } + + acquired, err := r.acquireLock(ctx, &ops, subject.pv) + if err != nil { + return ctrl.Result{}, err + } + if !acquired { + held := subject.pv.Annotations[simplyblockv1alpha2.PersistentVolumeOpsLock] + r.event(&ops, corev1.EventTypeNormal, ReasonOperationQueued, + "Volume %s is held by operation %s; this one is waiting", subject.pv.Name, held) + return ctrl.Result{RequeueAfter: opsRetry}, r.hold(ctx, &ops, + fmt.Sprintf("waiting for operation %s to release volume %s", held, subject.pv.Name)) + } + + return r.advance(ctx, &ops, subject) +} + +// advance runs the action's machine forward by at most one step. +func (r *PersistentVolumeOpsReconciler) advance( + ctx context.Context, ops *simplyblockv1alpha2.PersistentVolumeOps, subject *subject, +) (ctrl.Result, error) { + machine, err := graphs(memberCount(ops)).FromSnapshot(ctx, + statemachine.Action(ops.Spec.Action), + statemachine.FromKube[step](ops.Status.Step)) + if err != nil { + // An unrecognized step or action is a downgrade, a hand-edited object, + // or a rename that shipped without a conversion, and none of them + // resolve by reconciling again. The operation is terminal with what was + // found in status.message, which leaves an audit record saying so. + return r.finish(ctx, ops, simplyblockv1alpha2.PersistentVolumeOpsPhaseFailed, + fmt.Sprintf("the operation cannot be resumed: %v", err)) + } + defer machine.Close() + + // A machine is born already in its initial state, so that state's entry + // hook never runs and its deadline is never armed. Arming it on the first + // pass is what stops the first step being the one step that cannot time + // out. + if ops.Status.Step.State == "" { + return r.enterInitialStep(ctx, ops, machine) + } + + current := machine.CurrentState() + + if ops.Spec.Abort { + return r.unwind(ctx, ops, subject, machine, current) + } + + if machine.TimeoutReached() { + r.event(ops, corev1.EventTypeWarning, ReasonStepDeadlineExceeded, + "Step %s outlived its deadline", current) + stepDeadlinesExceeded.WithLabelValues( + subject.clusterUUID, string(ops.Spec.Action), string(current)).Inc() + return r.fail(ctx, ops, subject, fmt.Sprintf("step %s outlived its deadline", current)) + } + + done, err := r.perform(ctx, ops, subject, current) + if err != nil { + var fatal *terminalStepError + if errors.As(err, &fatal) { + return r.fail(ctx, ops, subject, fatal.Error()) + } + logf.FromContext(ctx).Error(err, "the step could not be advanced", + "operation", ops.Name, "step", current) + return ctrl.Result{RequeueAfter: opsRetry}, r.note(ctx, ops, err.Error()) + } + if !done { + return r.waitOn(machine), r.note(ctx, ops, fmt.Sprintf("waiting on %s", current)) + } + + r.observeStep(ops, subject, current) + + if machine.IsTerminal() { + r.event(ops, corev1.EventTypeNormal, ReasonOperationSucceeded, + "Volume %s was moved to node %s", ops.Spec.PersistentVolumeName, subject.targetNodeName()) + return r.finish(ctx, ops, simplyblockv1alpha2.PersistentVolumeOpsPhaseSucceeded, + fmt.Sprintf("the volume was moved to node %s", subject.targetNodeName())) + } + + next := firstSuccessor(machine) + // Write-ahead: the step is recorded before it is entered, so a crash + // between the two leaves a record that the step was attempted rather than a + // record that it was not. + if err := r.recordStep(ctx, ops, next, nil); err != nil { + return ctrl.Result{}, err + } + if err := machine.TransitionTo(ctx, next); err != nil { + return ctrl.Result{}, fmt.Errorf("enter step %s: %w", next, err) + } + snapshot := statemachine.ToKube(machine.Snapshot()) + return ctrl.Result{RequeueAfter: opsAdvance}, r.recordStep(ctx, ops, next, snapshot.Deadline) +} + +// enterInitialStep sets the first step's deadline and moves the operation to +// Running. +func (r *PersistentVolumeOpsReconciler) enterInitialStep( + ctx context.Context, + ops *simplyblockv1alpha2.PersistentVolumeOps, + machine *statemachine.Machine[step], +) (ctrl.Result, error) { + now := metav1.Now() + deadline := metav1.NewTime(now.Add(initialDeadline)) + if err := r.writeStatus(ctx, ops, func(status *simplyblockv1alpha2.PersistentVolumeOpsStatus) { + status.Phase = simplyblockv1alpha2.PersistentVolumeOpsPhaseRunning + status.Step = statemachine.KubeSnapshot{ + State: string(machine.CurrentState()), + Deadline: &deadline, + } + status.Message = "the operation holds the volume and is running" + if status.StartedAt == nil { + status.StartedAt = &now + } + status.DeferredSince = nil + }); err != nil { + return ctrl.Result{}, err + } + r.event(ops, corev1.EventTypeNormal, ReasonOperationStarted, + "The operation acquired the lock on volume %s and started", + ops.Spec.PersistentVolumeName) + return ctrl.Result{RequeueAfter: opsAdvance}, nil +} + +// unwind honors spec.abort where the graph allows it, and reports an abort that +// arrived too late rather than half-undoing the work. +// +// The refusal is the point. Verifying has already cut the volume over, so an +// abort there would leave the volume moved and the paths its move created +// untracked, which is the state the step exists to prevent. +func (r *PersistentVolumeOpsReconciler) unwind( + ctx context.Context, + ops *simplyblockv1alpha2.PersistentVolumeOps, + subject *subject, + machine *statemachine.Machine[step], + current step, +) (ctrl.Result, error) { + // The machine is asked rather than a table beside it: the graph it was + // built from is the one authority over what this action can stop from, and + // the DELETE guard reads the same graph. + if !machine.CanAbort() { + // Not a failure of the operation: it carries on. What was asked for + // cannot be done, and saying so is the whole of the response. + return ctrl.Result{RequeueAfter: opsRetry}, r.note(ctx, ops, fmt.Sprintf( + "the abort arrived at step %s, by which point the volume has already moved "+ + "and the cleanup is what makes the move safe; the operation is running on", current)) + } + + if err := r.discardMigration(ctx, ops, subject); err != nil { + // Not terminal. The operation stays where it is and the abort is + // honored on a later pass, because ending it now is what leaks the + // backend migration and the paths it published. + return ctrl.Result{RequeueAfter: opsRetry}, r.note(ctx, ops, + fmt.Sprintf("the abort is waiting on the migration being taken back: %v", err)) + } + + r.event(ops, corev1.EventTypeNormal, ReasonOperationAborted, + "The operation was aborted at step %s and unwound", current) + return r.finish(ctx, ops, simplyblockv1alpha2.PersistentVolumeOpsPhaseAborted, + fmt.Sprintf("aborted at step %s", current)) +} + +// fail ends the operation, having first taken back what it created. A migration +// abandoned with its target paths still connected on every consumer host is the +// defect this package exists around, and a failure is the commonest way to get +// there. +func (r *PersistentVolumeOpsReconciler) fail( + ctx context.Context, + ops *simplyblockv1alpha2.PersistentVolumeOps, + subject *subject, + message string, +) (ctrl.Result, error) { + if err := r.discardMigration(ctx, ops, subject); err != nil { + logf.FromContext(ctx).Error(err, "the migration could not be taken back", + "operation", ops.Name) + message += fmt.Sprintf(" (taking the migration back also failed: %v)", err) + } + r.event(ops, corev1.EventTypeWarning, ReasonOperationFailed, "%s", message) + return r.finish(ctx, ops, simplyblockv1alpha2.PersistentVolumeOpsPhaseFailed, message) +} + +// abandon ends an operation whose volume went away, which is a stop rather than +// a failure. Cancel first, then let the object go: a logical volume with a +// migration running against it is not one the control plane can cleanly delete. +func (r *PersistentVolumeOpsReconciler) abandon( + ctx context.Context, ops *simplyblockv1alpha2.PersistentVolumeOps, message string, +) (ctrl.Result, error) { + if err := r.discardMigration(ctx, ops, nil); err != nil { + return ctrl.Result{RequeueAfter: opsRetry}, r.note(ctx, ops, + fmt.Sprintf("the volume is gone and the migration could not be taken back: %v", err)) + } + r.event(ops, corev1.EventTypeNormal, ReasonOperationAborted, "%s", message) + return r.finish(ctx, ops, simplyblockv1alpha2.PersistentVolumeOpsPhaseAborted, message) +} + +// teardown unwinds what the operation created, releases the volume's lock, and +// lets the object go. +// +// The admission guard refuses a delete from the step where the work cannot be +// taken back, so what reaches here either produced nothing or produced +// something this can discard. A finalizer that cannot finish holds the object +// open rather than dropping the paths, which is the intended outcome: a path +// connected with nothing tracking it blocks every later migration of the +// volume, and has. +func (r *PersistentVolumeOpsReconciler) teardown( + ctx context.Context, ops *simplyblockv1alpha2.PersistentVolumeOps, +) (ctrl.Result, error) { + if !controllerutil.ContainsFinalizer(ops, opsFinalizer) { + return ctrl.Result{}, nil + } + + if !terminal(ops.Status.Phase) { + if err := r.discardMigration(ctx, ops, nil); err != nil { + r.event(ops, corev1.EventTypeWarning, ReasonCleanupBlocked, + "The delete is held because the cleanup has not finished: %v", err) + logf.FromContext(ctx).Error(err, "the migration could not be taken back", + "operation", ops.Name) + return ctrl.Result{RequeueAfter: opsRetry}, r.note(ctx, ops, + fmt.Sprintf("the delete is held until the cleanup finishes: %v", err)) + } + } + + if err := r.releaseLock(ctx, ops); err != nil { + return ctrl.Result{}, err + } + controllerutil.RemoveFinalizer(ops, opsFinalizer) + return ctrl.Result{}, r.Update(ctx, ops) +} + +// waitOn requeues for whatever is left of the current step's deadline, so that +// a step with a long budget is looked at when it expires rather than on a fixed +// interval, and a step with a short one is not left waiting past it. +func (r *PersistentVolumeOpsReconciler) waitOn(machine *statemachine.Machine[step]) ctrl.Result { + if remaining, bounded := machine.RequeueAfter(); bounded && remaining < opsRetry { + return ctrl.Result{RequeueAfter: remaining} + } + return ctrl.Result{RequeueAfter: opsRetry} +} + +// firstSuccessor is the step that follows the current one. The graph is a line +// rather than a tree, so the first edge is the only edge, and a graph that +// grows a branch will need the choice made here rather than in the caller. +func firstSuccessor(machine *statemachine.Machine[step]) step { + for next := range machine.AllowedTransitions() { + return next + } + return machine.CurrentState() +} + +// finish writes a terminal phase, releases the volume's lock, and records what +// the operation cost. +func (r *PersistentVolumeOpsReconciler) finish( + ctx context.Context, + ops *simplyblockv1alpha2.PersistentVolumeOps, + phase simplyblockv1alpha2.PersistentVolumeOpsPhase, + message string, +) (ctrl.Result, error) { + now := metav1.Now() + if err := r.writeStatus(ctx, ops, func(status *simplyblockv1alpha2.PersistentVolumeOpsStatus) { + status.Phase = phase + status.Message = message + status.CompletedAt = &now + }); err != nil { + return ctrl.Result{}, err + } + r.observeOperation(ops, phase) + return ctrl.Result{}, r.releaseLock(ctx, ops) +} + +// hold reports an operation that is admitted, holds nothing, and is waiting. +// Pending is both where an operation starts and where it waits, and +// status.deferredSince is what says since when — in status rather than in +// memory, because the operator may restart and an observer needs to see it. +func (r *PersistentVolumeOpsReconciler) hold( + ctx context.Context, ops *simplyblockv1alpha2.PersistentVolumeOps, message string, +) error { + return r.writeStatus(ctx, ops, func(status *simplyblockv1alpha2.PersistentVolumeOpsStatus) { + if status.Phase == "" { + status.Phase = simplyblockv1alpha2.PersistentVolumeOpsPhasePending + } + if status.DeferredSince == nil { + now := metav1.Now() + status.DeferredSince = &now + } + status.Message = message + }) +} + +// note replaces status.message without moving anything else. It is one sentence +// about where the operation is, replaced as it moves, and never a log. +func (r *PersistentVolumeOpsReconciler) note( + ctx context.Context, ops *simplyblockv1alpha2.PersistentVolumeOps, message string, +) error { + return r.writeStatus(ctx, ops, func(status *simplyblockv1alpha2.PersistentVolumeOpsStatus) { + status.Message = message + }) +} + +// recordStep persists the step the operation is about to be in, with the +// instant it expires. Both travel together, because a step persisted without +// its deadline restores as a step that can never time out. +func (r *PersistentVolumeOpsReconciler) recordStep( + ctx context.Context, + ops *simplyblockv1alpha2.PersistentVolumeOps, + next step, + deadline *metav1.Time, +) error { + return r.writeStatus(ctx, ops, func(status *simplyblockv1alpha2.PersistentVolumeOpsStatus) { + status.Phase = simplyblockv1alpha2.PersistentVolumeOpsPhaseRunning + status.Step = statemachine.KubeSnapshot{State: string(next), Deadline: deadline} + }) +} + +// writeStatus applies the mutation and patches only when something changed. +// +// Retried rather than swallowed on a conflict. A caller that read nil would +// take the write for done, and finish does: it releases the volume's lock +// straight afterward, so a dropped terminal status would free the volume for +// the next migration while this operation still reported Running. +func (r *PersistentVolumeOpsReconciler) writeStatus( + ctx context.Context, + ops *simplyblockv1alpha2.PersistentVolumeOps, + mutate func(*simplyblockv1alpha2.PersistentVolumeOpsStatus), +) error { + return retry.RetryOnConflict(retry.DefaultRetry, func() error { + var fresh simplyblockv1alpha2.PersistentVolumeOps + if err := r.Get(ctx, types.NamespacedName{Name: ops.Name}, &fresh); err != nil { + return err + } + + desired := *fresh.Status.DeepCopy() + mutate(&desired) + desired.ObservedGeneration = fresh.Generation + + if reflect.DeepEqual(fresh.Status, desired) { + // Still published to the caller, which reads the object it passed + // in on the next line of its own logic. + ops.Status = desired + ops.ResourceVersion = fresh.ResourceVersion + return nil + } + + patch := client.MergeFromWithOptions(fresh.DeepCopy(), client.MergeFromWithOptimisticLock{}) + fresh.Status = desired + if err := r.Status().Patch(ctx, &fresh, patch); err != nil { + return err + } + ops.Status = fresh.Status + ops.ResourceVersion = fresh.ResourceVersion + return nil + }) +} + +func (r *PersistentVolumeOpsReconciler) event( + object client.Object, eventType, reason, format string, args ...any, +) { + if r.Recorder == nil { + return + } + r.Recorder.Eventf(object, nil, eventType, reason, reason, format, args...) +} + +// terminal reports a phase the operation can never leave. +func terminal(phase simplyblockv1alpha2.PersistentVolumeOpsPhase) bool { + switch phase { + case simplyblockv1alpha2.PersistentVolumeOpsPhaseSucceeded, + simplyblockv1alpha2.PersistentVolumeOpsPhaseFailed, + simplyblockv1alpha2.PersistentVolumeOpsPhaseAborted: + return true + default: + return false + } +} + +// memberCount is how many volumes the migrated subsystem holds, which is what +// the copy's deadline scales by. It is zero until the migration's creation +// reports one, and the graph turns that into the base bound. +func memberCount(ops *simplyblockv1alpha2.PersistentVolumeOps) int32 { + if ops.Status.Migration == nil || ops.Status.Migration.MemberCount == nil { + return 0 + } + return *ops.Status.Migration.MemberCount +} + +// terminalStepError is a step failure that retrying cannot fix: a volume this +// operator cannot address, a target node that is not this volume's, a control +// plane that refused the migration outright. It is a distinct type so that the +// reconcile loop can tell it from a control plane that is briefly unreachable, +// which is the same shape of error and the opposite response. +type terminalStepError struct{ reason string } + +func (e *terminalStepError) Error() string { return e.reason } + +func fatalf(format string, args ...any) error { + return &terminalStepError{reason: fmt.Sprintf(format, args...)} +} + +// errVolumeGone is the volume having been deleted or begun deleting, which ends +// the operation without it having gone wrong. +var errVolumeGone = errors.New("the PersistentVolume this operation acts on is gone") + +// volumeLockChanged narrows the volume watch to what this controller reacts to: +// the lock changing hands, and the volume being deleted. A PersistentVolume is +// otherwise updated often enough — capacity, phase, claim binding — that +// watching every change would reconcile every operation in the cluster for +// reasons none of them care about. +type volumeLockChanged struct{} + +func (volumeLockChanged) Create(event.TypedCreateEvent[client.Object]) bool { return false } + +func (volumeLockChanged) Delete(event.TypedDeleteEvent[client.Object]) bool { return true } + +func (volumeLockChanged) Generic(event.TypedGenericEvent[client.Object]) bool { return false } + +func (volumeLockChanged) Update(e event.TypedUpdateEvent[client.Object]) bool { + if e.ObjectOld == nil || e.ObjectNew == nil { + return false + } + was := e.ObjectOld.GetAnnotations()[simplyblockv1alpha2.PersistentVolumeOpsLock] + is := e.ObjectNew.GetAnnotations()[simplyblockv1alpha2.PersistentVolumeOpsLock] + if was != is { + return true + } + return e.ObjectOld.GetDeletionTimestamp().IsZero() != e.ObjectNew.GetDeletionTimestamp().IsZero() +} + +// ensure the predicate satisfies the interface it is passed as. +var _ predicate.TypedPredicate[client.Object] = volumeLockChanged{} diff --git a/operator/internal/controllers/volume/persistentvolumeops_controller_test.go b/operator/internal/controllers/volume/persistentvolumeops_controller_test.go new file mode 100644 index 000000000..d9911f6c4 --- /dev/null +++ b/operator/internal/controllers/volume/persistentvolumeops_controller_test.go @@ -0,0 +1,590 @@ +// The operation's spine and its three steps, driven end to end against a fake +// client and a fake control plane. +// +// The assertions worth reading twice are the ones about what the operator does +// when a call is ambiguous. A migration is a data-path operation, so the +// expensive failures here are not crashes: they are an operation that reports +// success and lost writes, or one that reports failure and canceled a copy +// that had already committed. Both of those have happened, and the cases that +// cover them are copy-once, and continue only from pre_created. + +package volume + +import ( + "context" + "errors" + "testing" + "time" + + batchv1 "k8s.io/api/batch/v1" + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/types" + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + + "github.com/simplyblock/atlas/controlplane" + "github.com/simplyblock/atlas/lvol" + "github.com/simplyblock/atlas/statemachine" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// runPass drives one reconcile over the test operation. +func runPass(t *testing.T, r *PersistentVolumeOpsReconciler) ctrl.Result { + t.Helper() + result, err := r.Reconcile(context.Background(), + ctrl.Request{NamespacedName: types.NamespacedName{Name: testOpsName}}) + if err != nil { + t.Fatalf("reconcile: %v", err) + } + return result +} + +// operationFrom reads the operation back. +func operationFrom(t *testing.T, r *PersistentVolumeOpsReconciler) *simplyblockv1alpha2.PersistentVolumeOps { + t.Helper() + var ops simplyblockv1alpha2.PersistentVolumeOps + if err := r.Get(context.Background(), types.NamespacedName{Name: testOpsName}, &ops); err != nil { + t.Fatalf("reading the operation: %v", err) + } + return &ops +} + +// createdMigration is what the control plane answers a creation with: the +// subsystem's migration, its members, and the paths the target now answers on. +func createdMigration() controlplane.Migration { + return controlplane.Migration{ + Kind: controlplane.MigrationOfSubsystem, + ID: testMigrationID, + SourceNodeID: testSourceID, + TargetNodeID: testTargetID, + Phase: migrationPhasePreCreated, + Status: "running", + TargetNQN: testNQN, + MemberCount: 3, + Paths: []lvol.Endpoint{ + {Transport: "tcp", Address: "10.0.0.1", Port: 4420, NrIOQueues: 8}, + {Transport: "tcp", Address: "10.0.0.2", Port: 4420, NrIOQueues: 8}, + }, + } +} + +// idleSubsystem is a control plane whose subsystem no host consumes, which is +// the shortest path through the graph: nothing to check, so the copy starts. +func idleSubsystem() *fakeControlPlane { + return &fakeControlPlane{ + volume: lvol.Volume{ID: lvol.NewVolumeHandle(testClusterID, testPoolID, testVolumeID), NQN: testNQN}, + created: createdMigration(), + read: createdMigration(), + } +} + +func testWorld() []client.Object { + return []client.Object{testOperation(), testVolumeObject(), testClusterObject(), testNodeObject()} +} + +// TestAFreshOperationTakesTheLockAndStartsAtValidating. The first pass is the +// one that decides whether the operation runs at all, and the step it records is +// the one the write-ahead record has to name. +func TestAFreshOperationTakesTheLockAndStartsAtValidating(t *testing.T) { + r := testReconciler(t, idleSubsystem(), testWorld()...) + + runPass(t, r) + + ops := operationFrom(t, r) + if ops.Status.Phase != simplyblockv1alpha2.PersistentVolumeOpsPhaseRunning { + t.Errorf("phase = %q, want Running", ops.Status.Phase) + } + if ops.Status.Step.State != string(stepValidating) { + t.Errorf("step = %q, want Validating", ops.Status.Step.State) + } + if _, bounded := ops.Status.Step.KubeDeadline(); !bounded { + t.Error("the first step carries no deadline, so it is the one step that cannot time out") + } + if got := lockOn(t, r); got != testOpsName { + t.Errorf("the volume's lock names %q, want the operation", got) + } +} + +// TestTheMigrationIsCreatedOnce. The creation allocates on the backend and +// publishes the paths every host is then checked against, so a second one for +// the same operation is a second migration of the same subsystem — which the +// control plane would either refuse or, worse, accept. +func TestTheMigrationIsCreatedOnce(t *testing.T) { + api := idleSubsystem() + r := testReconciler(t, api, testWorld()...) + + for range 4 { + runPass(t, r) + } + + if api.creates != 1 { + t.Errorf("the control plane was asked for %d migrations, want exactly 1", api.creates) + } + + ops := operationFrom(t, r) + if ops.Status.Migration == nil { + t.Fatal("nothing about the migration was recorded") + } + if got := ops.Status.Migration.MigrationUUID; got != testMigrationID { + t.Errorf("migration = %q, want the one that was created", got) + } + if got := ops.Status.Migration.SubsystemNQN; got != testNQN { + t.Errorf("subsystem = %q, want the one the volume publishes under", got) + } + if got := ops.Status.Migration.MemberCount; got == nil || *got != 3 { + t.Errorf("member count = %v, want the 3 the subsystem holds", got) + } + if n := len(ops.Status.Migration.Connections); n != 2 { + t.Errorf("connections = %d, want both paths the target published", n) + } +} + +// TestTheRecordedPathsCarryTheDriversLossTimeout. What status.connections shows +// has to be the connect that will actually be made. Every volume path in this +// system is established with the driver's controller-loss timeout, and a +// migration target path becomes the data path at cutover: recording the hour +// the control plane answers with would describe a connect nobody performs. +func TestTheRecordedPathsCarryTheDriversLossTimeout(t *testing.T) { + api := idleSubsystem() + // The control plane's own answer, which is deliberately not what is kept. + hour := 3600 + paths := api.created.Paths + for i := range paths { + paths[i].CtrlLossTMOSec = &hour + } + api.created.Paths = paths + + r := testReconciler(t, api, testWorld()...) + runPass(t, r) + runPass(t, r) + + ops := operationFrom(t, r) + for _, conn := range ops.Status.Migration.Connections { + if conn.CtrlLossTimeoutSeconds == nil || *conn.CtrlLossTimeoutSeconds != migrationCtrlLossTimeout { + t.Errorf("ctrl-loss-tmo = %v, want the driver's %d", + conn.CtrlLossTimeoutSeconds, migrationCtrlLossTimeout) + } + } +} + +// TestASubsystemNobodyConsumesSkipsStraightToTheCopy. There are no host paths +// to check when no pod has any of the subsystem's volumes mounted, and starting +// Jobs to check nothing would only spend the step's deadline. +func TestASubsystemNobodyConsumesSkipsStraightToTheCopy(t *testing.T) { + r := testReconciler(t, idleSubsystem(), testWorld()...) + + for range 4 { + runPass(t, r) + } + + if got := operationFrom(t, r).Status.Step.State; got != string(stepMigrating) { + t.Errorf("step = %q, want Migrating", got) + } + + var jobs batchv1.JobList + if err := r.List(context.Background(), &jobs); err != nil { + t.Fatal(err) + } + if len(jobs.Items) != 0 { + t.Errorf("%d Jobs were started for a subsystem nobody consumes", len(jobs.Items)) + } +} + +// TestEveryConsumingHostIsCheckedBeforeTheCutover. A migration moves the whole +// subsystem, so every sibling volume moves with the named one. Checking only +// the named volume's host leaves every sibling's consumer pointing at the +// source, and at cutover those hosts lose their volume. +func TestEveryConsumingHostIsCheckedBeforeTheCutover(t *testing.T) { + const siblingVolume = "77777777-7777-7777-7777-777777777777" + + api := idleSubsystem() + api.members = []lvol.Volume{ + {ID: lvol.NewVolumeHandle(testClusterID, testPoolID, testVolumeID), NQN: testNQN}, + {ID: lvol.NewVolumeHandle(testClusterID, testPoolID, siblingVolume), NQN: testNQN}, + } + + // The named volume is replaced by a claimed one, because a consumer is + // found through the claim the volume names. + r := testReconciler(t, api, + testOperation(), testClusterObject(), testNodeObject(), + claimedVolume(testPVName, "app", "data-0"), + claimedVolume("pvc-"+siblingVolume, "app", "data-1"), + runningPodOn("worker-1", "app", "data-0"), + runningPodOn("worker-2", "app", "data-1"), + ) + + for range 3 { + runPass(t, r) + } + + ops := operationFrom(t, r) + nodes := map[string]bool{} + for _, job := range ops.Status.Migration.ValidationJobs { + nodes[job.Node] = true + } + if !nodes["worker-1"] || !nodes["worker-2"] { + t.Errorf("checked %v, want both hosts that consume the subsystem", nodes) + } +} + +// TestTheCopyIsContinuedOnlyFromPreCreated. Continue is not idempotent: it +// accepts a migration in pre_created and rejects any later call. A pass that +// continued the copy and crashed before recording it must not have that copy +// canceled by the pass that follows. +func TestTheCopyIsContinuedOnlyFromPreCreated(t *testing.T) { + api := idleSubsystem() + r := testReconciler(t, api, testWorld()...) + + // Reach the copy. + for range 4 { + runPass(t, r) + } + if got := operationFrom(t, r).Status.Step.State; got != string(stepMigrating) { + t.Fatalf("step = %q, want Migrating", got) + } + + // The first pass in the step continues it; the control plane then reports + // a migration that has advanced. + runPass(t, r) + api.read = controlplane.Migration{ID: testMigrationID, Phase: "migrating", Status: "running"} + runPass(t, r) + runPass(t, r) + + if api.continues != 1 { + t.Errorf("the copy was continued %d times, want exactly 1", api.continues) + } + if api.cancels != 0 { + t.Errorf("a running copy was canceled %d times", api.cancels) + } +} + +// TestTheCopyFinishesOnTheReportedStateRatherThanTheCall. A five-second read +// timeout on a continue that took slightly longer once made the operator retry +// a transfer that had already committed, and the retry copied nothing while the +// source was unfrozen. The step therefore finishes on what the migration says +// about itself. +func TestTheCopyFinishesOnTheReportedStateRatherThanTheCall(t *testing.T) { + api := idleSubsystem() + api.continueErr = errors.New("read timeout after 5s") + r := testReconciler(t, api, testWorld()...) + + for range 4 { + runPass(t, r) + } + + // The continue fails, so the operation stays where it is rather than + // failing: the copy may well have started. + runPass(t, r) + if ops := operationFrom(t, r); terminal(ops.Status.Phase) { + t.Fatalf("a failed continue ended the operation at %q", ops.Status.Phase) + } + + // The migration itself then reports the copy finished, which is the + // statement that counts. + api.read = controlplane.Migration{ID: testMigrationID, Phase: "done", Status: migrationStatusDone} + runPass(t, r) + + if got := operationFrom(t, r).Status.Step.State; got != string(stepVerifying) { + t.Errorf("step = %q, want Verifying: the copy reported itself finished", got) + } +} + +// TestAFinishedMigrationSucceedsAndReleasesTheVolume. The lock has to come off +// on every terminal path, or the next migration of that volume waits on an +// operation that has finished. +func TestAFinishedMigrationSucceedsAndReleasesTheVolume(t *testing.T) { + api := idleSubsystem() + r := testReconciler(t, api, testWorld()...) + + api.read = controlplane.Migration{ID: testMigrationID, Phase: "done", Status: migrationStatusDone} + for range 8 { + runPass(t, r) + } + + ops := operationFrom(t, r) + if ops.Status.Phase != simplyblockv1alpha2.PersistentVolumeOpsPhaseSucceeded { + t.Fatalf("phase = %q (%s), want Succeeded", ops.Status.Phase, ops.Status.Message) + } + if ops.Status.CompletedAt == nil { + t.Error("a terminal operation records no completion time") + } + if got := lockOn(t, r); got != "" { + t.Errorf("the volume is still locked by %q", got) + } +} + +// TestAnAbortBeforeTheCutoverTakesTheMigrationBack. Stopping a migration that +// has not cut over means canceling the backend migration: leaving it would +// block the subsystem's next one. +func TestAnAbortBeforeTheCutoverTakesTheMigrationBack(t *testing.T) { + api := idleSubsystem() + r := testReconciler(t, api, testWorld()...) + + runPass(t, r) + runPass(t, r) + + ops := operationFrom(t, r) + ops.Spec.Abort = true + if err := r.Update(context.Background(), ops); err != nil { + t.Fatal(err) + } + runPass(t, r) + + ops = operationFrom(t, r) + if ops.Status.Phase != simplyblockv1alpha2.PersistentVolumeOpsPhaseAborted { + t.Fatalf("phase = %q, want Aborted", ops.Status.Phase) + } + if api.cancels != 1 { + t.Errorf("the backend migration was canceled %d times, want 1", api.cancels) + } + if got := lockOn(t, r); got != "" { + t.Errorf("the volume is still locked by %q", got) + } +} + +// TestAnAbortAfterTheCutoverIsRefusedAndTheOperationRunsOn. By Verifying the +// volume has moved and the cleanup is what makes the move safe. Honoring an +// abort there would leave the system holding exactly the state the step exists +// to prevent — and the DELETE guard reads the same graph, so a deletion cannot +// express it either. +func TestAnAbortAfterTheCutoverIsRefusedAndTheOperationRunsOn(t *testing.T) { + api := idleSubsystem() + r := testReconciler(t, api, testWorld()...) + + ops := operationFrom(t, r) + ops.Spec.Abort = true + if err := r.Update(context.Background(), ops); err != nil { + t.Fatal(err) + } + if err := atStep(r, stepVerifying); err != nil { + t.Fatal(err) + } + + runPass(t, r) + + ops = operationFrom(t, r) + if terminal(ops.Status.Phase) { + t.Fatalf("the abort was honored at Verifying: phase = %q", ops.Status.Phase) + } + if api.cancels != 0 { + t.Errorf("a migration that had already cut over was canceled %d times", api.cancels) + } + if ops.Status.Message == "" { + t.Error("the refusal is not reported anywhere the user can read it") + } +} + +// TestAVolumeThatWentAwayEndsTheOperationAsAborted. A claim deleted under a +// Delete reclaim policy makes the driver delete the backing logical volume, and +// moving a volume that is being deleted is work nobody will read. Aborted +// rather than Failed, because a migration whose volume went away did not go +// wrong. +func TestAVolumeThatWentAwayEndsTheOperationAsAborted(t *testing.T) { + api := idleSubsystem() + r := testReconciler(t, api, testWorld()...) + + runPass(t, r) + runPass(t, r) + + pv := volumeFrom(t, r) + if err := r.Delete(context.Background(), pv); err != nil { + t.Fatal(err) + } + runPass(t, r) + + ops := operationFrom(t, r) + if ops.Status.Phase != simplyblockv1alpha2.PersistentVolumeOpsPhaseAborted { + t.Fatalf("phase = %q (%s), want Aborted", ops.Status.Phase, ops.Status.Message) + } + // Cancel before the object goes: a logical volume with a migration running + // against it is not one the control plane can cleanly delete. + if api.cancels != 1 { + t.Errorf("the backend migration was canceled %d times, want 1", api.cancels) + } +} + +// TestDeletingAnOperationUnwindsBeforeTheObjectGoes. An operation removed +// mid-flight would otherwise leave a backend migration running and the paths it +// published connected, with the only record naming them gone. +func TestDeletingAnOperationUnwindsBeforeTheObjectGoes(t *testing.T) { + api := idleSubsystem() + r := testReconciler(t, api, testWorld()...) + + runPass(t, r) + runPass(t, r) + + if err := r.Delete(context.Background(), operationFrom(t, r)); err != nil { + t.Fatal(err) + } + runPass(t, r) + + if api.cancels != 1 { + t.Errorf("the backend migration was canceled %d times, want 1", api.cancels) + } + var ops simplyblockv1alpha2.PersistentVolumeOps + err := r.Get(context.Background(), types.NamespacedName{Name: testOpsName}, &ops) + if err == nil { + t.Errorf("the operation still exists with finalizers %v", ops.Finalizers) + } + if got := lockOn(t, r); got != "" { + t.Errorf("the volume is still locked by %q", got) + } +} + +// TestADeleteIsHeldWhileTheCleanupCannotFinish. Leaving a visibly stuck object +// is the intended outcome: a path connected with nothing tracking it blocks +// every later migration of the volume, and releasing the finalizer would lose +// the only record naming it. +func TestADeleteIsHeldWhileTheCleanupCannotFinish(t *testing.T) { + api := idleSubsystem() + api.cancelErr = errors.New("the control plane is unreachable") + r := testReconciler(t, api, testWorld()...) + + runPass(t, r) + runPass(t, r) + + if err := r.Delete(context.Background(), operationFrom(t, r)); err != nil { + t.Fatal(err) + } + runPass(t, r) + + ops := operationFrom(t, r) + if len(ops.Finalizers) == 0 { + t.Fatal("the finalizer was released while the cleanup had not finished") + } + if ops.Status.Message == "" { + t.Error("nothing says why the delete is held") + } +} + +// TestAStepThatOutlivesItsDeadlineFailsTheOperation. Everything that waits on a +// migration waits for a terminal phase, so a step that could not time out would +// stall those flows silently rather than reporting something they can act on. +func TestAStepThatOutlivesItsDeadlineFailsTheOperation(t *testing.T) { + api := idleSubsystem() + r := testReconciler(t, api, testWorld()...) + + runPass(t, r) + + // Wind the step's deadline into the past, which is the only part of a + // timeout a unit test can reach. + ops := operationFrom(t, r) + expired := metav1.NewTime(time.Now().Add(-time.Minute)) + ops.Status.Step = statemachine.KubeSnapshot{ + State: string(stepValidating), + Deadline: &expired, + } + if err := r.Status().Update(context.Background(), ops); err != nil { + t.Fatal(err) + } + + runPass(t, r) + + ops = operationFrom(t, r) + if ops.Status.Phase != simplyblockv1alpha2.PersistentVolumeOpsPhaseFailed { + t.Fatalf("phase = %q, want Failed", ops.Status.Phase) + } + if got := lockOn(t, r); got != "" { + t.Errorf("the volume is still locked by %q", got) + } +} + +// TestAVolumeOfAnotherDriverIsRefusedRatherThanRetried. It is a well-formed +// request against the wrong object, and no reconcile will ever make it work. +func TestAVolumeOfAnotherDriverIsRefusedRatherThanRetried(t *testing.T) { + foreign := testVolumeObject() + foreign.Spec.CSI.Driver = "ebs.csi.aws.com" + + r := testReconciler(t, idleSubsystem(), + testOperation(), foreign, testClusterObject(), testNodeObject()) + + runPass(t, r) + + ops := operationFrom(t, r) + if ops.Status.Phase != simplyblockv1alpha2.PersistentVolumeOpsPhaseFailed { + t.Fatalf("phase = %q (%s), want Failed", ops.Status.Phase, ops.Status.Message) + } +} + +// TestAQueuedOperationWaitsAtPendingAndSaysWhy. A second operation for a locked +// volume is admitted, acquires nothing, and waits — which is what makes a +// drain's fan-out proceed in turn rather than half of it failing. +func TestAQueuedOperationWaitsAtPendingAndSaysWhy(t *testing.T) { + holder := testOperation() + holder.Name = testHolderName + holder.Status.Phase = simplyblockv1alpha2.PersistentVolumeOpsPhaseRunning + + pv := testVolumeObject() + pv.Annotations = map[string]string{ + simplyblockv1alpha2.PersistentVolumeOpsLock: holder.Name, + } + r := testReconciler(t, idleSubsystem(), + holder, testOperation(), pv, testClusterObject(), testNodeObject()) + + result := runPass(t, r) + + ops := operationFrom(t, r) + if ops.Status.Phase != simplyblockv1alpha2.PersistentVolumeOpsPhasePending { + t.Errorf("phase = %q, want Pending", ops.Status.Phase) + } + if ops.Status.DeferredSince == nil { + t.Error("nothing records since when the operation has been waiting") + } + if result.RequeueAfter == 0 { + t.Error("a queued operation asked for no requeue, so nothing would look at it again") + } +} + +// atStep puts the operation at a step without driving it there, for the cases +// whose interesting behavior is at the far end of the graph. +func atStep(r *PersistentVolumeOpsReconciler, at step) error { + ctx := context.Background() + var ops simplyblockv1alpha2.PersistentVolumeOps + if err := r.Get(ctx, types.NamespacedName{Name: testOpsName}, &ops); err != nil { + return err + } + deadline := metav1.NewTime(time.Now().Add(time.Hour)) + ops.Status.Phase = simplyblockv1alpha2.PersistentVolumeOpsPhaseRunning + ops.Status.Step = statemachine.KubeSnapshot{State: string(at), Deadline: &deadline} + ops.Status.Migration = &simplyblockv1alpha2.MigrationStatus{ + MigrationUUID: testMigrationID, + ClusterUUID: testClusterID, + PoolUUID: testPoolID, + VolumeUUID: testVolumeID, + SubsystemNQN: testNQN, + } + return r.Status().Update(ctx, &ops) +} + +// claimedVolume is a PersistentVolume of this driver bound to a claim, which is +// how a consumer is found: the volume names the claim and a pod mounts it. +func claimedVolume(name, namespace, claim string) *corev1.PersistentVolume { + pv := testVolumeObject() + pv.Name = name + if name != testPVName { + pv.Spec.CSI.VolumeHandle = string(lvol.NewVolumeHandle( + testClusterID, testPoolID, name[len("pvc-"):])) + } + pv.Spec.ClaimRef = &corev1.ObjectReference{Namespace: namespace, Name: claim} + return pv +} + +func runningPodOn(node, namespace, claim string) *corev1.Pod { + return &corev1.Pod{ + ObjectMeta: metav1.ObjectMeta{Name: "app-" + claim, Namespace: namespace}, + Spec: corev1.PodSpec{ + NodeName: node, + Volumes: []corev1.Volume{{ + Name: "data", + VolumeSource: corev1.VolumeSource{ + PersistentVolumeClaim: &corev1.PersistentVolumeClaimVolumeSource{ + ClaimName: claim, + }, + }, + }}, + }, + Status: corev1.PodStatus{Phase: corev1.PodRunning}, + } +} diff --git a/operator/internal/controllers/volume/steps.go b/operator/internal/controllers/volume/steps.go new file mode 100644 index 000000000..756c89e42 --- /dev/null +++ b/operator/internal/controllers/volume/steps.go @@ -0,0 +1,397 @@ +// What each step of a migration does on the pass it is entered, and what makes +// it finished. +// +// The three steps are the three things a safe migration is made of. The +// migration is created and every host that consumes the subsystem is checked +// against the target it is about to be served from. The copy runs. The paths +// the creation published are accounted for. Skipping the first leaves a host +// pointing at a node that has nothing for it at cutover; skipping the last +// leaves connections nothing tracks, which poisons the data path and blocks +// every later migration of the volume. +// +// One rule runs through all of it: a step finishes on the state the control +// plane reports, never on the return of the call that started it. A five-second +// read timeout on a request that took slightly longer once made the operator +// retry a transfer that had already committed, and the retry copied nothing +// while the source was unfrozen. +// +// design-persistentvolumeops.md §5 and §7 are the specification. + +package volume + +import ( + "context" + "fmt" + "strings" + "time" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + + "github.com/simplyblock/atlas/controlplane" + "github.com/simplyblock/atlas/lvol" + "github.com/simplyblock/atlas/ptr" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// Control-plane migration states, in the control plane's own spelling. They are +// not an enum in the client, because the client reports what it was told rather +// than deciding which values exist. +const ( + migrationPhasePreCreated = "pre_created" + + migrationStatusDone = "done" + migrationStatusFailed = "failed" + migrationStatusCanceled = "canceled" +) + +// perform advances the current step and reports whether it has finished. +func (r *PersistentVolumeOpsReconciler) perform( + ctx context.Context, + ops *simplyblockv1alpha2.PersistentVolumeOps, + subject *subject, + current step, +) (bool, error) { + switch current { + case stepValidating: + return r.validate(ctx, ops, subject) + case stepMigrating: + return r.copy(ctx, ops, subject) + case stepVerifying: + return r.verify(ctx, ops, subject) + default: + return false, fatalf("step %s has no implementation", current) + } +} + +// validate creates the migration and proves that every host consuming the +// subsystem can reach the target on the paths it published. +// +// It is one step rather than two because the paths only exist once the +// migration does: the control plane publishes them with the creation and with +// nothing else, so a migration created and then forgotten is a migration whose +// paths nothing can name. +func (r *PersistentVolumeOpsReconciler) validate( + ctx context.Context, + ops *simplyblockv1alpha2.PersistentVolumeOps, + subject *subject, +) (bool, error) { + if ops.Status.Migration == nil || ops.Status.Migration.MigrationUUID == "" { + return false, r.createMigration(ctx, ops, subject) + } + + nodes, err := r.consumingNodes(ctx, ops, subject) + switch { + case isConsumerNotReady(err): + // A consumer that exists and is not Running yet is waited for rather + // than skipped: a pod that stages against the source mid-migration is + // stranded at cutover exactly like an established one. The step's + // deadline is what bounds the wait. + r.event(ops, corev1.EventTypeNormal, ReasonWaitingForConsumer, "%s", err.Error()) + return false, nil + case err != nil: + return false, err + case len(nodes) == 0: + // No volume of the subsystem has a consumer, so there are no host paths + // to check anywhere. + r.event(ops, corev1.EventTypeNormal, ReasonValidationSkipped, + "No consumer for any volume of subsystem %s", ops.Status.Migration.SubsystemNQN) + return true, nil + } + + started, err := r.startValidationJobs(ctx, ops, subject, nodes) + if err != nil { + return false, err + } + if started > 0 { + // The Jobs are watched, so the next pass arrives when one finishes + // rather than on a timer. + return false, nil + } + return r.validationJobsPassed(ctx, ops, subject) +} + +// createMigration asks the control plane for the migration and records +// everything the later steps address it by. +// +// A control plane that is busy is not a failure. It refuses a migration while a +// cluster-wide realignment runs, which ends on its own, so the operation waits +// in its step rather than failing — and the step's deadline is what stops it +// waiting forever. +func (r *PersistentVolumeOpsReconciler) createMigration( + ctx context.Context, + ops *simplyblockv1alpha2.PersistentVolumeOps, + subject *subject, +) error { + volume, err := r.API.Volume(ctx, subject.handle.Handle()) + if err != nil { + return fmt.Errorf("read volume %s: %w", subject.handle.VolumeID, err) + } + if volume.NQN == "" { + return fatalf("volume %s publishes under no subsystem, so its migration cannot be addressed", + subject.handle.VolumeID) + } + + migration, err := r.API.CreateMigration(ctx, subject.clusterUUID, volume.NQN, subject.targetUUID) + if err != nil { + // The request may have taken effect despite the error: a create can + // take longer than the client's timeout and allocates on the way, so + // failing here would abandon a half-created migration. Retrying is what + // lets a later pass find and cancel it. + return fmt.Errorf("create the migration of subsystem %s: %w", volume.NQN, err) + } + if migration.ID == "" { + return fatalf("the control plane created a migration with no identifier") + } + if migration.SourceNodeID != "" && migration.SourceNodeID == subject.targetUUID { + r.event(ops, corev1.EventTypeWarning, ReasonTargetNodeIsSource, + "Volume %s is already on node %s", ops.Spec.PersistentVolumeName, subject.targetNodeName()) + // Cancel rather than continue: the migration exists on the backend and + // leaving it would block the next one. + if cancelErr := r.API.CancelMigration(ctx, subject.clusterUUID, volume.NQN, migration.ID); cancelErr != nil { + return fmt.Errorf("cancel a migration to the node the volume is already on: %w", cancelErr) + } + return fatalf("volume %s is already on node %s", + ops.Spec.PersistentVolumeName, subject.targetNodeName()) + } + + if err := r.writeStatus(ctx, ops, func(status *simplyblockv1alpha2.PersistentVolumeOpsStatus) { + status.Migration = &simplyblockv1alpha2.MigrationStatus{ + MigrationUUID: migration.ID, + ClusterUUID: subject.clusterUUID, + PoolUUID: subject.handle.PoolRef, + VolumeUUID: subject.handle.VolumeID, + SubsystemNQN: volume.NQN, + SourceNodeUUID: migration.SourceNodeID, + TargetNodeUUID: subject.targetUUID, + MemberCount: ptr.To(int32(migration.MemberCount)), + Connections: connectionsOf(migration), + } + status.DeferredSince = nil + }); err != nil { + return err + } + + r.event(ops, corev1.EventTypeNormal, ReasonMigrationCreated, + "Migration %s created for subsystem %s (%d volume(s)) to node %s", + migration.ID, volume.NQN, migration.MemberCount, subject.targetNodeName()) + return nil +} + +// connectionsOf records the paths the migration published, as they will be +// connected rather than as the control plane answered. +// +// The controller-loss timeout is overridden here rather than where the Job is +// built, so that what status.connections shows is the connect that will +// actually be made. Every volume path in this system is established with the +// CSI driver's timeout, and a migration target path becomes the volume's data +// path at cutover: connecting it with the hour the control plane answers would +// leave one volume's paths on two different timeouts depending on which of them +// last moved. +func connectionsOf(migration controlplane.Migration) []simplyblockv1alpha2.MigrationConnection { + out := make([]simplyblockv1alpha2.MigrationConnection, 0, len(migration.Paths)) + for _, path := range migration.Paths { + out = append(out, simplyblockv1alpha2.MigrationConnection{ + NQN: migration.TargetNQN, + Address: path.Address, + Port: ptr.To(int32(path.Port)), + Transport: path.Transport, + NrIOQueues: ptr.To(int32(path.NrIOQueues)), + ReconnectDelaySeconds: ptr.To(int32(path.ReconnectDelaySec)), + CtrlLossTimeoutSeconds: ptr.To(int32(migrationCtrlLossTimeout)), + FastIOFailTimeoutSeconds: ptr.To(int32(ptr.From(path.FastIOFailTMOSec, 0))), + KeepAliveTimeoutSeconds: ptr.To(int32(path.KeepAliveTMOSec)), + }) + } + return out +} + +// copy continues the migration, which is what starts the data copy, and waits +// for the control plane to report it finished. +// +// Continue is not idempotent: it accepts a migration in pre_created and rejects +// any later call. So the phase is read first, and a migration that has already +// advanced is left alone — a pass that continued the copy and then crashed +// before recording it must not have its copy canceled by the pass that +// follows. +func (r *PersistentVolumeOpsReconciler) copy( + ctx context.Context, + ops *simplyblockv1alpha2.PersistentVolumeOps, + subject *subject, +) (bool, error) { + migration := ops.Status.Migration + if migration == nil || migration.MigrationUUID == "" { + return false, fatalf("the operation reached the copy with no migration recorded") + } + + current, err := r.API.GetMigration(ctx, + migration.ClusterUUID, migration.SubsystemNQN, migration.MigrationUUID) + if err != nil { + return false, fmt.Errorf("read migration %s: %w", migration.MigrationUUID, err) + } + + switch current.Status { + case migrationStatusDone: + return true, nil + case migrationStatusFailed: + return false, fatalf("the migration failed: %s", current.ErrorMessage) + case migrationStatusCanceled: + return false, fatalf("the migration was canceled outside this operation") + } + + if current.Phase == migrationPhasePreCreated && migration.ContinuedAt == nil { + // Recorded before it is issued, which makes the request at-most-once. + // A repeated continue is the shape that has lost writes here, so a + // crash between this write and the call leaves the copy unstarted and + // the step to time out rather than leaving a transfer to be retried. + now := metav1.Now() + if err := r.writeStatus(ctx, ops, func(status *simplyblockv1alpha2.PersistentVolumeOpsStatus) { + status.Migration.ContinuedAt = &now + }); err != nil { + return false, err + } + if err := r.API.ContinueMigration(ctx, + migration.ClusterUUID, migration.SubsystemNQN, migration.MigrationUUID); err != nil { + // Whether the copy started is what the next read says, not what + // this error says. A call that timed out after the transfer + // committed reports a failure the migration itself contradicts. + return false, fmt.Errorf("continue migration %s: %w", migration.MigrationUUID, err) + } + r.event(ops, corev1.EventTypeNormal, ReasonMigrationStarted, + "Migration %s started: volume %s to node %s", + migration.MigrationUUID, ops.Spec.PersistentVolumeName, subject.targetNodeName()) + } + + // The source node is only reported once the migration is running, and a + // failure that says where the volume was is worth more than one that says + // only where it was going. + if current.SourceNodeID != "" && current.SourceNodeID != migration.SourceNodeUUID { + if err := r.writeStatus(ctx, ops, func(status *simplyblockv1alpha2.PersistentVolumeOpsStatus) { + status.Migration.SourceNodeUUID = current.SourceNodeID + if current.MemberCount > 0 { + status.Migration.MemberCount = ptr.To(int32(current.MemberCount)) + } + }); err != nil { + return false, err + } + } + return false, nil +} + +// verify takes the validation Jobs down and clears the husks a migration leaves +// on the hosts that took part in it. +// +// The paths themselves are not released. By this point the copy has finished +// and the target is where the volume is served from, so the paths the creation +// published are the data path: releasing them here is the outage the whole +// cleanup exists to avoid. What is cleared is the controller that carries +// nothing at all — the state a path lost mid-validation settles into, which +// blocks the subsystem's next migration just as surely as a live leak would. +// +// An operation that never cut over is a different case, and it does release: +// see discardMigration. +func (r *PersistentVolumeOpsReconciler) verify( + ctx context.Context, + ops *simplyblockv1alpha2.PersistentVolumeOps, + subject *subject, +) (bool, error) { + if ops.Status.Migration == nil || len(ops.Status.Migration.ValidationJobs) == 0 { + // Nothing was validated, so no host connected a path on this + // operation's account and there is nothing to account for. + return true, nil + } + + if err := r.deleteValidationJobs(ctx, ops); err != nil { + return false, err + } + + done, err := r.reapOnEveryValidatedNode(ctx, ops, subject) + if err != nil || !done { + return false, err + } + + // The record of the Jobs goes once they are gone and their hosts are + // clear, so that a restart does not start the cleanup over. + return true, r.writeStatus(ctx, ops, func(status *simplyblockv1alpha2.PersistentVolumeOpsStatus) { + status.Migration.ValidationJobs = nil + status.Migration.Connections = nil + }) +} + +// discardMigration takes back everything the operation created: the backend +// migration, and the target paths every consuming host connected on its +// account. +// +// It runs on the abort, the failure, and the deletion paths, which is to say +// exactly where the migration did not cut over. That is the precondition the +// release has: after a cutover the target paths are the volume's data path, and +// releasing them then is the outage the check exists to prevent. +// +// Every recorded node is asked rather than only the ones that passed. Release +// is idempotent and declines to touch a path that is serving, so asking a node +// that already released costs one Job and reports nothing; guessing which nodes +// still hold paths would mean trusting a Job's success to mean "connected," +// which it does not — a Job killed mid-run leaves paths with no record at all. +func (r *PersistentVolumeOpsReconciler) discardMigration( + ctx context.Context, ops *simplyblockv1alpha2.PersistentVolumeOps, subject *subject, +) error { + migration := ops.Status.Migration + if migration == nil || migration.MigrationUUID == "" { + // Nothing was ever created, which is the state being asked for. + return nil + } + + if err := r.API.CancelMigration(ctx, + migration.ClusterUUID, migration.SubsystemNQN, migration.MigrationUUID); err != nil { + return fmt.Errorf("cancel migration %s: %w", migration.MigrationUUID, err) + } + + if err := r.deleteValidationJobs(ctx, ops); err != nil { + return err + } + if err := r.releaseOnEveryValidatedNode(ctx, ops, subject); err != nil { + return err + } + + return r.writeStatus(ctx, ops, func(status *simplyblockv1alpha2.PersistentVolumeOpsStatus) { + status.Migration.ValidationJobs = nil + status.Migration.Connections = nil + }) +} + +// isConsumerNotReady reports the one waiting condition the consumer lookup +// distinguishes from a genuine failure. +func isConsumerNotReady(err error) bool { + return err != nil && strings.Contains(err.Error(), consumerNotRunning) +} + +// stepStarted is when the operation entered its current step, derived from the +// step's deadline and the budget it was set from. It is what the per-step +// histogram measures against, and it is derived rather than recorded because a +// second timestamp in status would be a second thing to keep true. +func stepStarted(ops *simplyblockv1alpha2.PersistentVolumeOps, current step) (time.Time, bool) { + deadline, bounded := ops.Status.Step.KubeDeadline() + if !bounded { + return time.Time{}, false + } + budget, known := stepBudgets[current] + if !known { + budget = copyDeadline(memberCount(ops)) + } + return deadline.Add(-budget), true +} + +// expectedMembers is the member count a validation has to cover, which is the +// subsystem's members and, always, the volume this operation names: a listing +// that raced a change must not drop the one volume the operation is about. +func expectedMembers(members []lvol.Volume, volumeUUID string) map[string]struct{} { + out := make(map[string]struct{}, len(members)+1) + for _, member := range members { + if handle, ok := lvol.ParseHandle(member.ID); ok { + out[handle.VolumeID] = struct{}{} + } + } + out[volumeUUID] = struct{}{} + return out +} diff --git a/operator/internal/controllers/volume/subject.go b/operator/internal/controllers/volume/subject.go new file mode 100644 index 000000000..424cd0278 --- /dev/null +++ b/operator/internal/controllers/volume/subject.go @@ -0,0 +1,217 @@ +// What one pass has to know about the world before it can act: which volume, +// which cluster, which pool, and which node the volume is being moved to. +// +// None of it is declared on the operation. A PersistentVolumeOps names a +// volume and a target node object, and everything else is read out of the +// volume's CSI handle — the cluster, the pool, and the volume itself. That +// handle is stamped on the volume at provisioning and is immutable for its +// life, so no StorageClass is consulted and a class edited, replaced, or +// deleted out of band changes nothing about an existing volume's +// addressability. +// +// design-persistentvolumeops.md §3 is the specification. + +package volume + +import ( + "context" + "fmt" + + corev1 "k8s.io/api/core/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" + "k8s.io/apimachinery/pkg/types" + "sigs.k8s.io/controller-runtime/pkg/client" + + "github.com/simplyblock/atlas/lvol" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// subject is the resolved world one pass acts on. +type subject struct { + pv *corev1.PersistentVolume + handle lvol.Handle + + // cluster is the StorageCluster reporting the volume's cluster UUID. Its + // namespace is where this operation's Jobs are created, which a + // cluster-scoped object has no namespace of its own to decide. + cluster *simplyblockv1alpha2.StorageCluster + clusterUUID string + + targetRef simplyblockv1alpha2.StorageNodeReference + targetUUID string +} + +func (s *subject) namespace() string { return s.cluster.Namespace } +func (s *subject) targetNodeName() string { return s.targetRef.Name } + +// resolve reads the operation's world, or says why it cannot. +// +// Three outcomes are distinguished, because the operation does something +// different with each. errVolumeGone ends the operation without it having gone +// wrong. A terminalStepError is a condition no reconcile can fix, and fails it. +// Anything else is a not-yet and is retried. +func (r *PersistentVolumeOpsReconciler) resolve( + ctx context.Context, ops *simplyblockv1alpha2.PersistentVolumeOps, +) (*subject, error) { + var pv corev1.PersistentVolume + err := r.Get(ctx, types.NamespacedName{Name: ops.Spec.PersistentVolumeName}, &pv) + switch { + case apierrors.IsNotFound(err): + return nil, fmt.Errorf("%w: %s", errVolumeGone, ops.Spec.PersistentVolumeName) + case err != nil: + return nil, fmt.Errorf("read volume %s: %w", ops.Spec.PersistentVolumeName, err) + case !pv.DeletionTimestamp.IsZero(): + // A volume being deleted is one whose backing logical volume the driver + // is about to remove, so moving it is work nobody will read. + return nil, fmt.Errorf("%w: %s is being deleted", errVolumeGone, pv.Name) + } + + handle, err := r.addressOf(ctx, &pv) + if err != nil { + return nil, err + } + + cluster, err := r.clusterReporting(ctx, handle.ClusterID) + if err != nil { + return nil, err + } + + target, err := r.targetNode(ctx, ops, cluster) + if err != nil { + return nil, err + } + + return &subject{ + pv: &pv, + handle: handle, + cluster: cluster, + clusterUUID: handle.ClusterID, + targetRef: ops.Spec.Migrate.TargetNodeRef, + targetUUID: target, + }, nil +} + +// addressOf reads the cluster, pool, and volume out of the volume's CSI handle. +// +// ParseHandle rather than Split, because the pool segment is not always a UUID: +// volumes provisioned before the v2 API migration encode the pool's name, and a +// PersistentVolume outlives every driver upgrade, so a cluster holds a mixture +// indefinitely. Nothing here needs the pool typed — the migration is addressed +// by cluster and subsystem — so requiring a UUID would refuse to move volumes +// that are otherwise perfectly movable. +func (r *PersistentVolumeOpsReconciler) addressOf( + ctx context.Context, pv *corev1.PersistentVolume, +) (lvol.Handle, error) { + if pv.Spec.CSI == nil { + return lvol.Handle{}, fatalf( + "volume %s has no CSI source, so it has no logical volume to move", pv.Name) + } + + ours, err := r.driverNames(ctx) + if err != nil { + return lvol.Handle{}, err + } + if !ours[pv.Spec.CSI.Driver] { + return lvol.Handle{}, fatalf( + "volume %s was provisioned by driver %s, which this operator has no means to act on", + pv.Name, pv.Spec.CSI.Driver) + } + + handle, ok := lvol.ParseHandle(lvol.VolumeHandle(pv.Spec.CSI.VolumeHandle)) + if !ok { + return lvol.Handle{}, fatalf( + "volume %s carries the handle %q, which is not ::", + pv.Name, pv.Spec.CSI.VolumeHandle) + } + return handle, nil +} + +// driverNames is the set of CSI drivers whose volumes this operator can move. +// +// It is read from the SimplyblockDriver objects rather than fixed, because +// spec.driverName is settable and immutable: a deployment that chose another +// name at install has volumes carrying that name forever, and matching only the +// default would refuse every one of them. The default stands in when no driver +// object exists, which is a cluster whose driver this operator does not manage. +func (r *PersistentVolumeOpsReconciler) driverNames(ctx context.Context) (map[string]bool, error) { + var drivers simplyblockv1alpha2.SimplyblockDriverList + if err := r.List(ctx, &drivers); err != nil { + return nil, fmt.Errorf("list the CSI drivers this operator manages: %w", err) + } + names := map[string]bool{} + for i := range drivers.Items { + if name := drivers.Items[i].Spec.DriverName; name != "" { + names[name] = true + } + } + if len(names) == 0 { + names[CSIDriverName] = true + } + return names, nil +} + +// clusterReporting finds the StorageCluster whose status reports this UUID. +// +// A cluster that reports none has not been created in the backend yet, which is +// a not-yet rather than a mismatch, so it is retried: the operation may have +// been written against a cluster that is still coming up. +func (r *PersistentVolumeOpsReconciler) clusterReporting( + ctx context.Context, uuid string, +) (*simplyblockv1alpha2.StorageCluster, error) { + var clusters simplyblockv1alpha2.StorageClusterList + if err := r.List(ctx, &clusters); err != nil { + return nil, fmt.Errorf("list the storage clusters: %w", err) + } + for i := range clusters.Items { + if clusters.Items[i].Status.UUID == uuid { + return &clusters.Items[i], nil + } + } + return nil, fmt.Errorf("no StorageCluster reports cluster %s yet", uuid) +} + +// targetNode resolves the backend UUID of the node the volume is moving to. +// +// The operation names a Kubernetes object, so that a migration can be written +// by hand without looking a UUID up, and this is where the two are joined. A +// node in a different cluster than the volume is refused at admission; it is +// checked again here because admission happens once and the objects can change +// afterward. +func (r *PersistentVolumeOpsReconciler) targetNode( + ctx context.Context, + ops *simplyblockv1alpha2.PersistentVolumeOps, + cluster *simplyblockv1alpha2.StorageCluster, +) (string, error) { + ref := ops.Spec.Migrate.TargetNodeRef + + var node simplyblockv1alpha2.StorageNode + err := r.Get(ctx, client.ObjectKey{Namespace: ref.Namespace, Name: ref.Name}, &node) + switch { + case apierrors.IsNotFound(err): + return "", fatalf("no StorageNode %s/%s to move the volume to", ref.Namespace, ref.Name) + case err != nil: + return "", fmt.Errorf("read the target node %s/%s: %w", ref.Namespace, ref.Name, err) + } + + if node.Namespace != cluster.Namespace || node.Spec.ClusterRef != cluster.Name { + return "", fatalf( + "node %s/%s belongs to cluster %s and the volume to %s, and a volume cannot move between clusters", + ref.Namespace, ref.Name, node.Spec.ClusterRef, cluster.Name) + } + + if node.Status.UUID == "" { + return "", fmt.Errorf("node %s/%s has not been created in the backend yet", ref.Namespace, ref.Name) + } + if node.Status.Phase != simplyblockv1alpha2.StorageNodePhaseOnline { + // A fact about now rather than about the operation. The node a drain + // fans fifty migrations out to may well be online by the time the + // fifteenth of them acquires its lock, which is why this is a phase and + // an event rather than an admission rejection. + r.event(ops, corev1.EventTypeWarning, ReasonTargetNodeNotReady, + "Node %s/%s is %s rather than online", ref.Namespace, ref.Name, node.Status.Phase) + return "", fmt.Errorf("node %s/%s is %s rather than online", + ref.Namespace, ref.Name, node.Status.Phase) + } + return node.Status.UUID, nil +} diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_persistentvolumeops.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_persistentvolumeops.yaml index ebeccc471..450f2a1a2 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_persistentvolumeops.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_persistentvolumeops.yaml @@ -266,6 +266,20 @@ spec: type: string type: object type: array + continuedAt: + description: |- + ContinuedAt is when this operation asked the control plane to start the + copy, written before the call rather than after it. + + The order is what makes the request at-most-once, and that is the point. + A repeated continue is the shape that has lost writes here: a read + timeout on a call that had in fact committed made the operator retry a + transfer, and the retry copied nothing while the source was unfrozen. So + a recorded continue is never issued again, and an operation that crashed + between this write and the call waits out its step's deadline instead — + a rare stall, against a silent data loss. + format: date-time + type: string memberCount: description: |- MemberCount is how many volumes the migrated NVMe-oF subsystem holds, as From 87d7e085e11693542532ca29fb9afdbda06daf7c Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 11:04:13 +0200 Subject: [PATCH 038/206] feat(webhook): the PersistentVolumeOps guard, at create and at delete MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The kind's admission half. It draws one line: a condition fixed for the object's whole life is a rejection, and a condition true at one moment and false at the next is a phase and an event. An admission decision is made once and never revisited, so a node that is offline and a cluster that has not been created in the backend yet are both admitted — the node a drain fans fifty migrations out to may well be online by the time the fifteenth of them acquires its lock. What is refused at create is every fact about a different object, which is what CEL cannot reach: a volume that does not exist, one with no CSI source, one another driver provisioned, a handle that is not three parts, a target node that does not exist, and a target in a different cluster than the volume. The agreement between spec.action and its parameter block is deliberately not here; it is a statement about one object's own fields, so it lives on the type where nothing can be installed in front of it. The driver check is the one that earns the webhook. Every other refusal is a malformed request, while a volume belonging to another CSI driver is a well-formed request against the wrong object — and it is the mistake somebody writing one by hand is likeliest to make, because `kubectl get pv` lists every volume in the cluster and says nothing about which of them this operator can move. The name it matches is read from the SimplyblockDriver objects rather than fixed, because spec.driverName is settable and immutable. At delete it refuses Verifying, where the copy has finished and what is left is the cleanup. Which steps those are is the graph's answer rather than this file's, and a test holds the two equal in both directions: a deletion may never express a stop that spec.abort could not. Co-Authored-By: Claude Fable 5 --- .../simplyblock-operator-webhook.yaml | 20 + operator/cmd/main.go | 6 + operator/config/webhook/manifests.yaml | 20 + operator/dist/install.yaml | 20 + .../webhook/persistentvolumeops_validator.go | 287 ++++++++++++++ .../persistentvolumeops_validator_test.go | 374 ++++++++++++++++++ 6 files changed, 727 insertions(+) create mode 100644 operator/internal/webhook/persistentvolumeops_validator.go create mode 100644 operator/internal/webhook/persistentvolumeops_validator_test.go diff --git a/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml b/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml index 019cecdad..ba96dd807 100644 --- a/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml @@ -126,6 +126,26 @@ webhooks: resources: - controlplaneops sideEffects: None +- admissionReviewVersions: + - v1 + clientConfig: + service: + name: simplyblock-operator-webhook-service + namespace: {{ .Release.Namespace }} + path: /validate-storage-simplyblock-io-v1alpha2-persistentvolumeops + failurePolicy: Fail + name: vpersistentvolumeops.simplyblock.io + rules: + - apiGroups: + - storage.simplyblock.io + apiVersions: + - v1alpha2 + operations: + - CREATE + - DELETE + resources: + - persistentvolumeops + sideEffects: None - admissionReviewVersions: - v1 clientConfig: diff --git a/operator/cmd/main.go b/operator/cmd/main.go index 09f49ebd0..cae6d0e00 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -890,6 +890,12 @@ func main() { &webhook.Admission{Handler: &internalwebhook.StorageNodeOpsValidator{}}) setupLog.Info("registered storagenodeops validating webhook") + mgr.GetWebhookServer().Register("/validate-storage-simplyblock-io-v1alpha2-persistentvolumeops", + &webhook.Admission{Handler: &internalwebhook.PersistentVolumeOpsValidator{ + Client: mgr.GetClient(), + }}) + setupLog.Info("registered persistentvolumeops validating webhook") + mgr.GetWebhookServer().Register("/validate-v1-pvc-pinned-volume", &webhook.Admission{Handler: &internalwebhook.PersistentVolumeClaimValidator{ Client: mgr.GetClient(), diff --git a/operator/config/webhook/manifests.yaml b/operator/config/webhook/manifests.yaml index 13890eb8f..c2b20e6cb 100644 --- a/operator/config/webhook/manifests.yaml +++ b/operator/config/webhook/manifests.yaml @@ -108,6 +108,26 @@ webhooks: resources: - controlplaneops sideEffects: None +- admissionReviewVersions: + - v1 + clientConfig: + service: + name: webhook-service + namespace: system + path: /validate-storage-simplyblock-io-v1alpha2-persistentvolumeops + failurePolicy: Fail + name: vpersistentvolumeops.simplyblock.io + rules: + - apiGroups: + - storage.simplyblock.io + apiVersions: + - v1alpha2 + operations: + - CREATE + - DELETE + resources: + - persistentvolumeops + sideEffects: None - admissionReviewVersions: - v1 clientConfig: diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index 7e12ccb32..40f18db7c 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -10903,6 +10903,26 @@ webhooks: resources: - controlplaneops sideEffects: None +- admissionReviewVersions: + - v1 + clientConfig: + service: + name: simplyblock-operator-webhook-service + namespace: simplyblock-operator-system + path: /validate-storage-simplyblock-io-v1alpha2-persistentvolumeops + failurePolicy: Fail + name: vpersistentvolumeops.simplyblock.io + rules: + - apiGroups: + - storage.simplyblock.io + apiVersions: + - v1alpha2 + operations: + - CREATE + - DELETE + resources: + - persistentvolumeops + sideEffects: None - admissionReviewVersions: - v1 clientConfig: diff --git a/operator/internal/webhook/persistentvolumeops_validator.go b/operator/internal/webhook/persistentvolumeops_validator.go new file mode 100644 index 000000000..176ee500a --- /dev/null +++ b/operator/internal/webhook/persistentvolumeops_validator.go @@ -0,0 +1,287 @@ +// The PersistentVolumeOps guard: a validating webhook that refuses at create +// what no reconcile could ever make work, and refuses at delete what would +// strand a volume mid-move. +// +// The line it draws is between what cannot change and what can. A condition +// fixed for the object's whole life is a rejection; a condition true at one +// moment and false at the next is a phase and an event. An admission decision +// is made once and never revisited, so anything time-varying decided here would +// be decided wrongly for most of the object's life — which is why a node that +// is offline and a cluster that has not been created yet are both admitted. +// +// The agreement between spec.action and its parameter block is not here. CEL +// can see it, since it is a statement about one object's own fields, and the +// group's floor is that such a rule lives on the type where nothing can be +// installed in front of it. What is left for the webhook is every row that is a +// fact about a different object, which CEL cannot reach. +// +// design-persistentvolumeops.md §4.3 is the specification. + +package webhook + +import ( + "context" + "encoding/json" + "fmt" + "net/http" + "slices" + "strings" + + admissionv1 "k8s.io/api/admission/v1" + corev1 "k8s.io/api/core/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" + "k8s.io/apimachinery/pkg/types" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/webhook/admission" + + "github.com/simplyblock/atlas/lvol" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/volume" +) + +// +kubebuilder:webhook:path=/validate-storage-simplyblock-io-v1alpha2-persistentvolumeops,mutating=false,failurePolicy=fail,sideEffects=None,groups=storage.simplyblock.io,resources=persistentvolumeops,verbs=create;delete,versions=v1alpha2,name=vpersistentvolumeops.simplyblock.io,admissionReviewVersions=v1 + +// PersistentVolumeOpsValidator stands in front of creating and deleting a +// volume operation. +// +// failurePolicy=Fail, for the reason the StorageNode validator gives: the +// webhook server runs inside the operator pod, so its availability tracks the +// operator's, and while the operator is down no migration would run anyway. +// +// Unlike the other Ops guards in this package it holds a client, because its +// create rules are facts about other objects: the volume, its driver, its +// handle, and the node the operation names. +type PersistentVolumeOpsValidator struct { + Client client.Client +} + +// undeletableVolumeSteps says what each step with no abort edge is in the +// middle of. Which steps those are is volume.UnabortableSteps's answer, and a +// test holds the two sets equal in both directions. +var undeletableVolumeSteps = map[simplyblockv1alpha2.PersistentVolumeOpsStep]string{ + simplyblockv1alpha2.PersistentVolumeOpsStepVerifying: "the copy has finished and the volume " + + "has already moved, so what is left is the cleanup that makes the move safe: the " + + "validation Jobs and the paths they connected on every consuming host", +} + +func (v *PersistentVolumeOpsValidator) Handle( + ctx context.Context, req admission.Request, +) admission.Response { + switch req.Operation { + case admissionv1.Create: + return v.admitCreate(ctx, req) + case admissionv1.Delete: + return v.admitDelete(req) + default: + return admission.Allowed("") + } +} + +// admitCreate refuses an operation that could never run. +func (v *PersistentVolumeOpsValidator) admitCreate( + ctx context.Context, req admission.Request, +) admission.Response { + if len(req.Object.Raw) == 0 { + return admission.Allowed("") + } + + var ops simplyblockv1alpha2.PersistentVolumeOps + if err := json.Unmarshal(req.Object.Raw, &ops); err != nil { + return admission.Errored(http.StatusBadRequest, err) + } + + handle, denied := v.addressableVolume(ctx, &ops) + if denied != nil { + return *denied + } + return v.targetInTheVolumesCluster(ctx, &ops, handle) +} + +// addressableVolume refuses a volume this operator has no means to move, and +// returns the handle the cluster is read out of when it can. +func (v *PersistentVolumeOpsValidator) addressableVolume( + ctx context.Context, ops *simplyblockv1alpha2.PersistentVolumeOps, +) (lvol.Handle, *admission.Response) { + name := ops.Spec.PersistentVolumeName + + var pv corev1.PersistentVolume + err := v.Client.Get(ctx, types.NamespacedName{Name: name}, &pv) + switch { + case apierrors.IsNotFound(err): + // Nothing creates an operation before its volume, so this is a typo. + return lvol.Handle{}, denied( + "there is no PersistentVolume %s to move. spec.persistentVolumeName names the "+ + "volume rather than the claim, because a claim can be deleted while its volume "+ + "is retained.", name) + case err != nil: + return lvol.Handle{}, errored(err) + } + + if pv.Spec.CSI == nil { + return lvol.Handle{}, denied( + "volume %s has no CSI source, so it has no logical volume to move.", name) + } + + // The check that earns this webhook. Every other refusal here is a + // malformed request; a volume belonging to another CSI driver is a + // well-formed request against the wrong object, and it is the mistake + // somebody writing one by hand is likeliest to make, because `kubectl get + // pv` lists every volume in the cluster and says nothing about which of + // them this operator can move. Answering at the moment the mistake is made, + // with the driver's name in the message, beats leaving an object sitting in + // Failed for somebody to read. + ours, err := v.driverNames(ctx) + if err != nil { + return lvol.Handle{}, errored(err) + } + if !ours[pv.Spec.CSI.Driver] { + return lvol.Handle{}, denied( + "volume %s was provisioned by driver %s, and this operator can only move volumes of "+ + "%s.", name, pv.Spec.CSI.Driver, joined(ours)) + } + + handle, ok := lvol.ParseHandle(lvol.VolumeHandle(pv.Spec.CSI.VolumeHandle)) + if !ok { + return lvol.Handle{}, denied( + "volume %s carries the volume handle %q, which is not ::, so "+ + "there is nothing to address the backend with.", name, pv.Spec.CSI.VolumeHandle) + } + return handle, nil +} + +// targetInTheVolumesCluster refuses a target node that does not exist and one +// that belongs to another cluster than the volume. +// +// The namespace is carried explicitly rather than derived, so that a +// hand-written spec can be read without performing a join to learn which object +// it names. That makes the mistake writable, and this is where it is refused, +// with both clusters in the message. +func (v *PersistentVolumeOpsValidator) targetInTheVolumesCluster( + ctx context.Context, ops *simplyblockv1alpha2.PersistentVolumeOps, handle lvol.Handle, +) admission.Response { + if ops.Spec.Migrate == nil { + // The type's own rule covers this, and it runs whether or not the + // webhook is installed. + return admission.Allowed("") + } + ref := ops.Spec.Migrate.TargetNodeRef + + var node simplyblockv1alpha2.StorageNode + err := v.Client.Get(ctx, client.ObjectKey{Namespace: ref.Namespace, Name: ref.Name}, &node) + switch { + case apierrors.IsNotFound(err): + return *denied("there is no StorageNode %s/%s to move volume %s to.", + ref.Namespace, ref.Name, ops.Spec.PersistentVolumeName) + case err != nil: + return *errored(err) + } + + var cluster simplyblockv1alpha2.StorageCluster + err = v.Client.Get(ctx, + client.ObjectKey{Namespace: node.Namespace, Name: node.Spec.ClusterRef}, &cluster) + switch { + case apierrors.IsNotFound(err): + // A node naming a cluster that does not exist is a broken node rather + // than a broken migration, and refusing here would report it against + // the wrong object. + return admission.Allowed("") + case err != nil: + return *errored(err) + } + + if cluster.Status.UUID == "" { + // The cluster has not been created in the backend, so there is nothing + // to compare the volume's cluster against. That is a not-yet rather + // than a mismatch, and it becomes the controller's condition. + return admission.Allowed("") + } + + if cluster.Status.UUID != handle.ClusterID { + return *denied( + "node %s/%s belongs to cluster %s and volume %s to cluster %s, and a volume cannot "+ + "move between clusters.", + ref.Namespace, ref.Name, cluster.Status.UUID, + ops.Spec.PersistentVolumeName, handle.ClusterID) + } + return admission.Allowed("") +} + +// driverNames is the set of CSI drivers whose volumes this operator can move, +// read from the SimplyblockDriver objects because spec.driverName is settable +// and immutable: a deployment that chose another name at install has volumes +// carrying that name forever. +func (v *PersistentVolumeOpsValidator) driverNames(ctx context.Context) (map[string]bool, error) { + var drivers simplyblockv1alpha2.SimplyblockDriverList + if err := v.Client.List(ctx, &drivers); err != nil { + return nil, err + } + names := map[string]bool{} + for i := range drivers.Items { + if name := drivers.Items[i].Spec.DriverName; name != "" { + names[name] = true + } + } + if len(names) == 0 { + names[volume.CSIDriverName] = true + } + return names, nil +} + +// admitDelete reads the object from req.OldObject, which is what the API server +// sends on a DELETE: there is no new object, and the step the operation is on is +// in the status of the one being removed. +func (v *PersistentVolumeOpsValidator) admitDelete(req admission.Request) admission.Response { + if len(req.OldObject.Raw) == 0 { + // Nothing to read means nothing to refuse on. Admitting is the only + // answer that does not block a delete on the basis of no information. + return admission.Allowed("") + } + + var ops simplyblockv1alpha2.PersistentVolumeOps + if err := json.Unmarshal(req.OldObject.Raw, &ops); err != nil { + return admission.Errored(http.StatusBadRequest, err) + } + + switch ops.Status.Phase { + case simplyblockv1alpha2.PersistentVolumeOpsPhaseSucceeded, + simplyblockv1alpha2.PersistentVolumeOpsPhaseFailed, + simplyblockv1alpha2.PersistentVolumeOpsPhaseAborted: + return admission.Allowed("the operation is terminal") + } + + step := simplyblockv1alpha2.PersistentVolumeOpsStep(ops.Status.Step.State) + doing, undeletable := undeletableVolumeSteps[step] + if !undeletable { + return admission.Allowed("") + } + + return *denied( + "PersistentVolumeOps %s is at step %s, where %s. Deleting the record would not undo the "+ + "move, it would remove the only thing naming those paths, and a path left connected "+ + "with nothing tracking it blocks every later migration of volume %s. Set spec.abort "+ + "to stop an operation that can still be stopped, and delete the record once it is "+ + "terminal.", ops.Name, step, doing, ops.Spec.PersistentVolumeName) +} + +// denied builds a refusal, as a pointer so a helper can return "no refusal." +func denied(format string, args ...any) *admission.Response { + response := admission.Denied(fmt.Sprintf(format, args...)) + return &response +} + +func errored(err error) *admission.Response { + response := admission.Errored(http.StatusInternalServerError, err) + return &response +} + +// joined renders the driver names a message lists, sorted so the message is the +// same on every replica. +func joined(names map[string]bool) string { + out := make([]string, 0, len(names)) + for name := range names { + out = append(out, name) + } + slices.Sort(out) + return strings.Join(out, ", ") +} diff --git a/operator/internal/webhook/persistentvolumeops_validator_test.go b/operator/internal/webhook/persistentvolumeops_validator_test.go new file mode 100644 index 000000000..52dde2241 --- /dev/null +++ b/operator/internal/webhook/persistentvolumeops_validator_test.go @@ -0,0 +1,374 @@ +package webhook + +import ( + "context" + "encoding/json" + "slices" + "strings" + "testing" + + admissionv1 "k8s.io/api/admission/v1" + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + "sigs.k8s.io/controller-runtime/pkg/webhook/admission" + + "github.com/simplyblock/atlas/statemachine" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/volume" +) + +const ( + pvopsVolume = "pvc-aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa" + pvopsCluster = "11111111-1111-1111-1111-111111111111" + pvopsPool = "22222222-2222-2222-2222-222222222222" + pvopsLvol = "33333333-3333-3333-3333-333333333333" + pvopsNode = "worker-5" + pvopsNS = "simplyblock" + + // pvopsOtherNS is where the second cluster lives, which is what makes a + // target in the wrong cluster writable at all. + pvopsOtherNS = "other" +) + +// pvopsValidator builds the validator over a world the test describes. +func pvopsValidator(t *testing.T, objs ...client.Object) *PersistentVolumeOpsValidator { + t.Helper() + s := runtime.NewScheme() + for _, add := range []func(*runtime.Scheme) error{ + corev1.AddToScheme, + simplyblockv1alpha2.AddToScheme, + } { + if err := add(s); err != nil { + t.Fatalf("build the scheme: %v", err) + } + } + return &PersistentVolumeOpsValidator{ + Client: fake.NewClientBuilder().WithScheme(s).WithObjects(objs...).Build(), + } +} + +// pvopsOperation is a well-formed migration, which every case below breaks in +// exactly one way. +func pvopsOperation() *simplyblockv1alpha2.PersistentVolumeOps { + return &simplyblockv1alpha2.PersistentVolumeOps{ + ObjectMeta: metav1.ObjectMeta{Name: "move-1"}, + Spec: simplyblockv1alpha2.PersistentVolumeOpsSpec{ + PersistentVolumeName: pvopsVolume, + Action: simplyblockv1alpha2.PersistentVolumeOpsActionMigrate, + Migrate: &simplyblockv1alpha2.MigrateVolumeSpec{ + TargetNodeRef: simplyblockv1alpha2.StorageNodeReference{ + Namespace: pvopsNS, + Name: pvopsNode, + }, + }, + }, + } +} + +// simplyblockVolume is a PersistentVolume this operator can act on. +func simplyblockVolume() *corev1.PersistentVolume { + return &corev1.PersistentVolume{ + ObjectMeta: metav1.ObjectMeta{Name: pvopsVolume}, + Spec: corev1.PersistentVolumeSpec{ + PersistentVolumeSource: corev1.PersistentVolumeSource{ + CSI: &corev1.CSIPersistentVolumeSource{ + Driver: "csi.simplyblock.io", + VolumeHandle: pvopsCluster + ":" + pvopsPool + ":" + pvopsLvol, + }, + }, + }, + } +} + +func pvopsClusterObject() *simplyblockv1alpha2.StorageCluster { + return &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{Name: "production", Namespace: pvopsNS}, + Status: simplyblockv1alpha2.StorageClusterStatus{UUID: pvopsCluster}, + } +} + +func pvopsNodeObject() *simplyblockv1alpha2.StorageNode { + return &simplyblockv1alpha2.StorageNode{ + ObjectMeta: metav1.ObjectMeta{Name: pvopsNode, Namespace: pvopsNS}, + Spec: simplyblockv1alpha2.StorageNodeSpec{ClusterRef: "production"}, + } +} + +func pvopsCreateRequest(t *testing.T, ops *simplyblockv1alpha2.PersistentVolumeOps) admission.Request { + t.Helper() + raw, err := json.Marshal(ops) + if err != nil { + t.Fatalf("encoding the operation: %v", err) + } + return admission.Request{AdmissionRequest: admissionv1.AdmissionRequest{ + Operation: admissionv1.Create, + Name: ops.Name, + Object: runtime.RawExtension{Raw: raw}, + }} +} + +func pvopsDeleteRequest(t *testing.T, ops *simplyblockv1alpha2.PersistentVolumeOps) admission.Request { + t.Helper() + raw, err := json.Marshal(ops) + if err != nil { + t.Fatalf("encoding the operation: %v", err) + } + return admission.Request{AdmissionRequest: admissionv1.AdmissionRequest{ + Operation: admissionv1.Delete, + Name: ops.Name, + OldObject: runtime.RawExtension{Raw: raw}, + }} +} + +// TestPersistentVolumeOpsAdmitsAWellFormedMigration is the positive half every +// refusal below is measured against: the same request, with nothing wrong. +func TestPersistentVolumeOpsAdmitsAWellFormedMigration(t *testing.T) { + v := pvopsValidator(t, simplyblockVolume(), pvopsClusterObject(), pvopsNodeObject()) + + got := v.Handle(context.Background(), pvopsCreateRequest(t, pvopsOperation())) + if !got.Allowed { + t.Fatalf("a well-formed migration was refused: %s", got.Result.Message) + } +} + +// TestPersistentVolumeOpsRefusesWhatNoReconcileCouldFix. Each row is a +// condition fixed for the object's whole life, which is the line this webhook +// draws: a fact about now is a phase and an event, and a fact that can never +// change is a rejection. +func TestPersistentVolumeOpsRefusesWhatNoReconcileCouldFix(t *testing.T) { + for _, tc := range []struct { + name string + world func() []client.Object + ops func() *simplyblockv1alpha2.PersistentVolumeOps + says string + }{ + { + name: "a volume that does not exist", + world: func() []client.Object { + return []client.Object{pvopsClusterObject(), pvopsNodeObject()} + }, + ops: pvopsOperation, + says: "no PersistentVolume", + }, + { + name: "a volume with no CSI source at all", + world: func() []client.Object { + pv := simplyblockVolume() + pv.Spec.CSI = nil + pv.Spec.HostPath = &corev1.HostPathVolumeSource{Path: "/mnt/data"} + return []client.Object{pv, pvopsClusterObject(), pvopsNodeObject()} + }, + ops: pvopsOperation, + says: "no CSI", + }, + { + // The row that earns the webhook. Every other rejection here is a + // malformed request; this one is a well-formed request against the + // wrong object, and it is the mistake somebody writing one by hand + // is likeliest to make, because `kubectl get pv` lists every volume + // in the cluster and says nothing about which of them this operator + // can move. + name: "a volume another driver provisioned", + world: func() []client.Object { + pv := simplyblockVolume() + pv.Spec.CSI.Driver = "ebs.csi.aws.com" + return []client.Object{pv, pvopsClusterObject(), pvopsNodeObject()} + }, + ops: pvopsOperation, + says: "ebs.csi.aws.com", + }, + { + name: "a handle that is not three parts", + world: func() []client.Object { + pv := simplyblockVolume() + pv.Spec.CSI.VolumeHandle = pvopsLvol + return []client.Object{pv, pvopsClusterObject(), pvopsNodeObject()} + }, + ops: pvopsOperation, + says: "volume handle", + }, + { + name: "a target node that does not exist", + world: func() []client.Object { + return []client.Object{simplyblockVolume(), pvopsClusterObject()} + }, + ops: pvopsOperation, + says: "no StorageNode", + }, + { + // The namespace is carried explicitly so a hand-written spec can be + // read without performing a join, which makes this mistake + // writable: two clusters in two namespaces may each hold a node + // called worker-5. + name: "a target node in another cluster than the volume", + world: func() []client.Object { + elsewhere := pvopsNodeObject() + elsewhere.Namespace = pvopsOtherNS + elsewhere.Spec.ClusterRef = "staging" + other := pvopsClusterObject() + other.Name, other.Namespace = "staging", pvopsOtherNS + other.Status.UUID = "99999999-9999-9999-9999-999999999999" + return []client.Object{simplyblockVolume(), pvopsClusterObject(), other, elsewhere} + }, + ops: func() *simplyblockv1alpha2.PersistentVolumeOps { + ops := pvopsOperation() + ops.Spec.Migrate.TargetNodeRef.Namespace = pvopsOtherNS + return ops + }, + says: "cluster", + }, + } { + t.Run(tc.name, func(t *testing.T) { + v := pvopsValidator(t, tc.world()...) + + got := v.Handle(context.Background(), pvopsCreateRequest(t, tc.ops())) + if got.Allowed { + t.Fatal("the request was admitted") + } + if !strings.Contains(got.Result.Message, tc.says) { + t.Errorf("refused for the wrong reason: %s", got.Result.Message) + } + }) + } +} + +// TestPersistentVolumeOpsAdmitsWhatIsMerelyTrueNow. A target node that is +// offline and a cluster that has not been created in the backend yet are both +// conditions the next reconcile may find changed, and an admission decision is +// made once and never revisited. +func TestPersistentVolumeOpsAdmitsWhatIsMerelyTrueNow(t *testing.T) { + t.Run("a cluster with no UUID yet", func(t *testing.T) { + cluster := pvopsClusterObject() + cluster.Status.UUID = "" + v := pvopsValidator(t, simplyblockVolume(), cluster, pvopsNodeObject()) + + got := v.Handle(context.Background(), pvopsCreateRequest(t, pvopsOperation())) + if !got.Allowed { + t.Errorf("refused a cluster that has simply not been created yet: %s", got.Result.Message) + } + }) + + t.Run("a target node that is offline", func(t *testing.T) { + node := pvopsNodeObject() + node.Status.Phase = simplyblockv1alpha2.StorageNodePhaseOffline + v := pvopsValidator(t, simplyblockVolume(), pvopsClusterObject(), node) + + got := v.Handle(context.Background(), pvopsCreateRequest(t, pvopsOperation())) + if !got.Allowed { + t.Errorf("refused a node that is merely offline right now: %s", got.Result.Message) + } + }) +} + +// TestPersistentVolumeOpsRefusesADeleteAfterTheCutover. Verifying holds the +// validation Jobs and the paths they connected, and the production defect this +// step exists for is exactly those paths outliving the object that recorded +// them. +func TestPersistentVolumeOpsRefusesADeleteAfterTheCutover(t *testing.T) { + v := pvopsValidator(t, simplyblockVolume(), pvopsClusterObject(), pvopsNodeObject()) + + ops := pvopsOperation() + ops.Status.Phase = simplyblockv1alpha2.PersistentVolumeOpsPhaseRunning + ops.Status.Step = statemachine.KubeSnapshot{ + State: string(simplyblockv1alpha2.PersistentVolumeOpsStepVerifying), + } + + got := v.Handle(context.Background(), pvopsDeleteRequest(t, ops)) + if got.Allowed { + t.Fatal("a delete at Verifying was admitted") + } + if !strings.Contains(got.Result.Message, "Verifying") { + t.Errorf("the refusal does not name the step: %s", got.Result.Message) + } + if !strings.Contains(got.Result.Message, "abort") { + t.Errorf("the refusal does not say what to do instead: %s", got.Result.Message) + } +} + +// TestPersistentVolumeOpsAdmitsADeleteTheAbortCouldExpress. A delete arriving +// where the graph still declares an abort edge is admitted and unwound by the +// finalizer, because the two channels have to agree. +func TestPersistentVolumeOpsAdmitsADeleteTheAbortCouldExpress(t *testing.T) { + v := pvopsValidator(t, simplyblockVolume(), pvopsClusterObject(), pvopsNodeObject()) + + for _, at := range []simplyblockv1alpha2.PersistentVolumeOpsStep{ + simplyblockv1alpha2.PersistentVolumeOpsStepValidating, + simplyblockv1alpha2.PersistentVolumeOpsStepMigrating, + } { + t.Run(string(at), func(t *testing.T) { + ops := pvopsOperation() + ops.Status.Phase = simplyblockv1alpha2.PersistentVolumeOpsPhaseRunning + ops.Status.Step = statemachine.KubeSnapshot{State: string(at)} + + if got := v.Handle(context.Background(), pvopsDeleteRequest(t, ops)); !got.Allowed { + t.Errorf("a delete at %s was refused: %s", at, got.Result.Message) + } + }) + } +} + +// TestPersistentVolumeOpsAdmitsADeleteOfAFinishedOperation. A terminal +// operation is a record of work that has finished, and withdrawing it stops +// nothing. +func TestPersistentVolumeOpsAdmitsADeleteOfAFinishedOperation(t *testing.T) { + v := pvopsValidator(t, simplyblockVolume(), pvopsClusterObject(), pvopsNodeObject()) + + for _, phase := range []simplyblockv1alpha2.PersistentVolumeOpsPhase{ + simplyblockv1alpha2.PersistentVolumeOpsPhaseSucceeded, + simplyblockv1alpha2.PersistentVolumeOpsPhaseFailed, + simplyblockv1alpha2.PersistentVolumeOpsPhaseAborted, + } { + t.Run(string(phase), func(t *testing.T) { + ops := pvopsOperation() + ops.Status.Phase = phase + // Terminal at the step a running operation would be refused from, + // which is the case the phase check has to come before the step's. + ops.Status.Step = statemachine.KubeSnapshot{ + State: string(simplyblockv1alpha2.PersistentVolumeOpsStepVerifying), + } + + if got := v.Handle(context.Background(), pvopsDeleteRequest(t, ops)); !got.Allowed { + t.Errorf("a delete of a %s operation was refused: %s", phase, got.Result.Message) + } + }) + } +} + +// TestPersistentVolumeOpsDeleteGuardAgreesWithTheGraph. The rule is the +// group's rather than this kind's: a deletion may never express a stop that +// spec.abort could not, so both channels read one graph. This holds the +// webhook's table and the graph equal in both directions — an entry here for a +// step the graph can abort from refuses a delete the abort channel would have +// honored, and a missing entry admits the withdrawal of a record nothing else +// accounts for. +func TestPersistentVolumeOpsDeleteGuardAgreesWithTheGraph(t *testing.T) { + refused := volume.UnabortableSteps() + + for step := range undeletableVolumeSteps { + if !slices.Contains(refused, step) { + t.Errorf("the guard refuses a delete at %s, which the graph can abort from", step) + } + } + for _, step := range refused { + if _, guarded := undeletableVolumeSteps[step]; !guarded { + t.Errorf("the graph declares no abort from %s, and the guard admits a delete there", step) + } + } +} + +// TestPersistentVolumeOpsIgnoresEverythingButCreateAndDelete. An update is the +// abort channel, which the controller reads rather than admission. +func TestPersistentVolumeOpsIgnoresEverythingButCreateAndDelete(t *testing.T) { + v := pvopsValidator(t) + + req := pvopsDeleteRequest(t, pvopsOperation()) + req.Operation = admissionv1.Update + + if got := v.Handle(context.Background(), req); !got.Allowed { + t.Errorf("an update was refused: %s", got.Result.Message) + } +} From aaa90f21fd335a83189e12fb557a3e28153740ca Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 11:19:17 +0200 Subject: [PATCH 039/206] feat(volume): PersistentVolumeOps is what raises a move, and VolumeMigration is opt-in MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Three controllers move volumes — the auto-rebalancer, a node drain, and the pinned-volume controller — and each of them built a VolumeMigration by hand. They now ask for a move and something else decides which kind carries it, defaulting to the kind this API group documents. The registered kind is kept behind --legacy-volume-migration, off by default, and its reconciler is registered only when the flag is on: an operator otherwise driving the redesigned kind should not also be the thing that quietly reconciles a VolumeMigration somebody applied by hand. The chart exposes it as volumeMigration.legacy. What it is for is an upgrade, and only for as long as the migrations already raised against the registered kind take to drain — a rename and a scope change make a new CRD rather than a new version, so an in-flight migration cannot be carried across and is left to finish where it started. The abstraction is four methods, because four is what the callers do: start a move, find their own again, read an outcome, and reap it. What it deliberately does not offer is anything specific to one kind, since a caller that needed that would be a caller the gate is not hiding anything from. The merged phase-and-step enum of the registered kind and the split phase of the redesigned one read as one vocabulary through it. Two differences could not be hidden and are not. The redesigned kind names its target as a StorageNode object rather than as a backend UUID, so a UUID nothing reports is refused rather than guessed at — every caller holds a UUID and the join happens in one place. And a cluster-scoped operation cannot be owned by the namespaced object that raised it: the garbage collector treats such a reference as unresolvable and would delete the migration mid-copy, so the creator moves into spec.creatorRef with its UID and the cascade becomes the creator's own. The drain's fan-out had no test at all before this. It has three now, one of them checking that the creator's UID is recorded, which is what stops a drain recreated under the same name inheriting a fan-out it did not issue. Co-Authored-By: Claude Fable 5 --- .../templates/simplyblock-operator.yaml | 7 + .../charts/simplyblock-operator/values.yaml | 13 + operator/cmd/main.go | 39 +- .../persistentvolumeclaim_controller.go | 87 ++-- ...sistentvolumeclaim_controller_unit_test.go | 107 +++-- .../controller/volumerebalancer_controller.go | 83 ++-- operator/internal/controllers/node/remove.go | 97 ++--- .../controllers/node/remove_fanout_test.go | 189 +++++++++ .../node/storagenodeops_controller.go | 6 + operator/internal/volumemigration/mover.go | 370 ++++++++++++++++++ .../internal/volumemigration/mover_test.go | 297 ++++++++++++++ operator/internal/volumemigration/utils.go | 49 +-- .../internal/volumemigration/utils_test.go | 70 +--- 13 files changed, 1119 insertions(+), 295 deletions(-) create mode 100644 operator/internal/controllers/node/remove_fanout_test.go create mode 100644 operator/internal/volumemigration/mover.go create mode 100644 operator/internal/volumemigration/mover_test.go diff --git a/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator.yaml b/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator.yaml index 54a8742a0..c3c137a90 100644 --- a/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator.yaml @@ -56,6 +56,13 @@ spec: - --csi-link-cert-path=/etc/simplyblock/csi-link-certs - --csi-link-audience={{ .Values.csiLink.audience }} {{- end }} + {{- if .Values.volumeMigration.legacy }} + # An upgrade turns this on so the migrations already raised against + # the registered kind keep being driven. A rename and a scope change + # make a new CRD rather than a new version, so an in-flight + # migration cannot be carried across and is drained instead. + - --legacy-volume-migration + {{- end }} {{- if .Values.metricsAPI.enabled }} - --metrics-api-bind-port={{ .Values.metricsAPI.port }} - --prometheus-url={{ .Values.metricsAPI.prometheusURL }} diff --git a/helm-charts/charts/simplyblock-operator/values.yaml b/helm-charts/charts/simplyblock-operator/values.yaml index 460c34a01..53c4f2612 100644 --- a/helm-charts/charts/simplyblock-operator/values.yaml +++ b/helm-charts/charts/simplyblock-operator/values.yaml @@ -464,6 +464,19 @@ controlplane: operator: openShiftCluster: true +# How a volume's move is raised. The operator raises one when the rebalancer +# moves a volume off a hot node, when a node is drained, and when a claim is +# pinned to a node it is not on. +# +# The redesigned PersistentVolumeOps is the default and the registered +# VolumeMigration is opt-in. Turn the registered kind back on for an upgrade, +# and only for as long as the migrations already raised against it take to +# drain: a rename and a scope change make a new CRD rather than a new version, +# so an in-flight migration cannot be carried across and is left to finish on +# the kind it started on. +volumeMigration: + legacy: false + # The CSI link: the CSI node and controller pods dial the operator and hold the # connection open, so the operator can query a node's storage state without # anything listening on the node. Off by default; it needs a serving diff --git a/operator/cmd/main.go b/operator/cmd/main.go index cae6d0e00..4f406fa77 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -65,6 +65,7 @@ import ( volumecontrollers "github.com/simplyblock/simplyblock-operator/internal/controllers/volume" "github.com/simplyblock/simplyblock-operator/internal/csilink" "github.com/simplyblock/simplyblock-operator/internal/utils" + "github.com/simplyblock/simplyblock-operator/internal/volumemigration" "github.com/simplyblock/simplyblock-operator/internal/webapi" internalwebhook "github.com/simplyblock/simplyblock-operator/internal/webhook" // +kubebuilder:scaffold:imports @@ -151,6 +152,14 @@ func main() { flag.StringVar(&metricsCertKey, "metrics-cert-key", "tls.key", "The name of the metrics server key file.") flag.BoolVar(&enableHTTP2, "enable-http2", false, "If set, HTTP/2 will be enabled for the metrics and webhook servers") + var legacyVolumeMigration bool + flag.BoolVar(&legacyVolumeMigration, "legacy-volume-migration", false, + "Move volumes as the registered VolumeMigration kind rather than as "+ + "PersistentVolumeOps. Off by default: the redesigned kind is what this API group "+ + "documents, and the registered one is kept for one release so that an upgrade can "+ + "turn it back on while migrations raised against it drain. A rename and a scope "+ + "change make a new CRD rather than a new version, so an in-flight migration cannot "+ + "be carried across.") var csiLinkEnabled bool var csiLinkAddr, csiLinkCertPath, csiLinkCertName, csiLinkCertKey, csiLinkAudience string flag.BoolVar(&csiLinkEnabled, "csi-link", false, @@ -623,6 +632,11 @@ func main() { setupLog.Error(err, "unable to create controller", "controller", "BackupImport") os.Exit(1) } + // The kind a volume move is raised as, decided once and handed to all three + // controllers that raise one. None of them has any business knowing which. + volumeMover := volumemigration.NewMover( + mgr.GetClient(), mgr.GetScheme(), legacyVolumeMigration) + // The volume band. It shares the backup band's control-plane client because // there is one control plane and one endpoint; what differs is which of its // endpoints each band calls. @@ -635,18 +649,27 @@ func main() { setupLog.Error(err, "unable to create controller", "controller", "PersistentVolumeOps") os.Exit(1) } - if err := (&controller.VolumeMigrationReconciler{ - Client: mgr.GetClient(), - Scheme: mgr.GetScheme(), - Recorder: mgr.GetEventRecorder("volumemigration-controller"), - }).SetupWithManager(mgr); err != nil { - setupLog.Error(err, "unable to create controller", "controller", "VolumeMigration") - os.Exit(1) + // The registered kind's reconciler runs only where the registered kind is + // what raises moves. Registering it regardless would cost nothing while + // nothing creates one, and would quietly become the thing that reconciles a + // VolumeMigration somebody applied by hand against an operator that is + // otherwise driving the redesigned kind. + if legacyVolumeMigration { + if err := (&controller.VolumeMigrationReconciler{ + Client: mgr.GetClient(), + Scheme: mgr.GetScheme(), + Recorder: mgr.GetEventRecorder("volumemigration-controller"), + }).SetupWithManager(mgr); err != nil { + setupLog.Error(err, "unable to create controller", "controller", "VolumeMigration") + os.Exit(1) + } + setupLog.Info("volume moves are raised as the registered VolumeMigration kind") } if err := (&controller.PersistentVolumeClaimReconciler{ Client: mgr.GetClient(), Scheme: mgr.GetScheme(), Recorder: mgr.GetEventRecorder("pinnedvolume-controller"), + Mover: volumeMover, }).SetupWithManager(mgr); err != nil { setupLog.Error(err, "unable to create controller", "controller", "PersistentVolumeClaim") os.Exit(1) @@ -675,6 +698,7 @@ func main() { Nodes: nodeSubscription, Clusters: clusterSubscription, Workload: storageNodeWorkload, + Mover: volumeMover, }).SetupWithManager(mgr); err != nil { setupLog.Error(err, "unable to create controller", "controller", "StorageNodeOps") os.Exit(1) @@ -747,6 +771,7 @@ func main() { Scheme: mgr.GetScheme(), Recorder: mgr.GetEventRecorder("volumerebalancer-controller"), LatencyPercentile: latencyPercentile, + Mover: volumeMover, }).SetupWithManager(mgr); err != nil { setupLog.Error(err, "unable to create controller", "controller", "VolumeRebalancer") os.Exit(1) diff --git a/operator/internal/controller/persistentvolumeclaim_controller.go b/operator/internal/controller/persistentvolumeclaim_controller.go index 8099909b9..934817e03 100644 --- a/operator/internal/controller/persistentvolumeclaim_controller.go +++ b/operator/internal/controller/persistentvolumeclaim_controller.go @@ -10,22 +10,20 @@ import ( corev1 "k8s.io/api/core/v1" apierrors "k8s.io/apimachinery/pkg/api/errors" - metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" "k8s.io/apimachinery/pkg/runtime" "k8s.io/apimachinery/pkg/types" "k8s.io/client-go/tools/events" ctrl "sigs.k8s.io/controller-runtime" "sigs.k8s.io/controller-runtime/pkg/builder" "sigs.k8s.io/controller-runtime/pkg/client" - "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil" "sigs.k8s.io/controller-runtime/pkg/event" logf "sigs.k8s.io/controller-runtime/pkg/log" "sigs.k8s.io/controller-runtime/pkg/predicate" "github.com/simplyblock/atlas/kube" - simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/volumemigration" "github.com/simplyblock/simplyblock-operator/internal/webapi" ) @@ -66,6 +64,10 @@ type PersistentVolumeClaimReconciler struct { Scheme *runtime.Scheme Recorder events.EventRecorder apiClient *webapi.Client + + // Mover raises a pinned volume's move as whichever kind this deployment + // runs. Unset means the kind this API group documents. + Mover volumemigration.Mover } func (r *PersistentVolumeClaimReconciler) Reconcile( @@ -176,54 +178,52 @@ func (r *PersistentVolumeClaimReconciler) Reconcile( return r.setApplied(ctx, pvc, desired) } -// createMigration creates a VolumeMigration for the PV/target in the owning -// StorageCluster's namespace. The name is deterministic in (pv, target) so a -// retried reconcile is idempotent (AlreadyExists is tolerated). +// createMigration raises the move of a pinned volume onto its requested node, +// as whichever kind this deployment runs. +// +// The name is deterministic in (volume, target) so a retried reconcile is +// idempotent, and the move is raised in the owning StorageCluster's namespace: +// a claim may live in another namespace than its cluster, and a cross-namespace +// owner reference is invalid. func (r *PersistentVolumeClaimReconciler) createMigration( ctx context.Context, cluster *simplyblockv1alpha2.StorageCluster, pvName, target string, ) error { - vm := &simplyblockv1alpha1.VolumeMigration{ - ObjectMeta: metav1.ObjectMeta{ - Name: pinMigrationName(pvName, target), - Namespace: cluster.Namespace, - Labels: map[string]string{labelPinnedVolumePV: pinPVLabelValue(pvName)}, - }, - Spec: simplyblockv1alpha1.VolumeMigrationSpec{ - PVName: pvName, - TargetNodeUUID: target, - }, - } - // Own by the StorageCluster (same namespace) so the migration is garbage- - // collected with the cluster. The PVC cannot be the owner: it may live in a - // different namespace, and a cross-namespace owner reference is invalid. A - // plain owner reference (not a controller reference) avoids setting - // blockOwnerDeletion, which would require storageclusters/finalizers access. - if err := controllerutil.SetOwnerReference(cluster, vm, r.Scheme); err != nil { - return fmt.Errorf("set owner reference on VolumeMigration: %w", err) - } - if err := r.Create(ctx, vm); err != nil && !apierrors.IsAlreadyExists(err) { - return fmt.Errorf("create VolumeMigration %q: %w", vm.Name, err) + return r.mover().Start(ctx, volumemigration.MoveRequest{ + Name: pinMigrationName(pvName, target), + Namespace: cluster.Namespace, + PVName: pvName, + TargetNodeUUID: target, + Labels: map[string]string{labelPinnedVolumePV: pinPVLabelValue(pvName)}, + Owner: cluster, + OwnerKind: "StorageCluster", + Scheme: r.Scheme, + }) +} + +// mover is this controller's channel for raising a move, defaulted so a +// reconciler built without one raises the kind this API group documents. +func (r *PersistentVolumeClaimReconciler) mover() volumemigration.Mover { + if r.Mover != nil { + return r.Mover } - return nil + return volumemigration.NewMover(r.Client, r.Scheme, false) } -// hasActiveMigration reports whether a non-terminal pin-driven VolumeMigration -// exists for the given PV. +// hasActiveMigration reports whether a non-terminal pin-driven move exists for +// the given volume. func (r *PersistentVolumeClaimReconciler) hasActiveMigration( ctx context.Context, namespace, pvName string, ) (bool, error) { - var list simplyblockv1alpha1.VolumeMigrationList - if err := r.List(ctx, &list, - client.InNamespace(namespace), - client.MatchingLabels{labelPinnedVolumePV: pinPVLabelValue(pvName)}, - ); err != nil { - return false, fmt.Errorf("list VolumeMigrations for PV %q: %w", pvName, err) + moves, err := r.mover().List(ctx, namespace, + map[string]string{labelPinnedVolumePV: pinPVLabelValue(pvName)}) + if err != nil { + return false, fmt.Errorf("list the moves of PV %q: %w", pvName, err) } - for _, vm := range list.Items { - if !isTerminalMigrationPhase(vm.Status.Phase) { + for _, move := range moves { + if !move.Phase.Terminal() { return true, nil } } @@ -327,17 +327,6 @@ func containsStorageNode(nodes []webapi.StorageNodeInfo, uuid string) bool { return false } -func isTerminalMigrationPhase(phase simplyblockv1alpha1.VolumeMigrationPhase) bool { - switch phase { - case simplyblockv1alpha1.VolumeMigrationPhaseCompleted, - simplyblockv1alpha1.VolumeMigrationPhaseFailed, - simplyblockv1alpha1.VolumeMigrationPhaseAborted: - return true - default: - return false - } -} - // pinMigrationName is a deterministic, DNS-label-safe VolumeMigration name for a // (PV, target) pair. Deterministic so a retried reconcile hits AlreadyExists // instead of creating duplicates; target-dependent so a new target yields a new diff --git a/operator/internal/controller/persistentvolumeclaim_controller_unit_test.go b/operator/internal/controller/persistentvolumeclaim_controller_unit_test.go index b92232d00..9b4f01704 100644 --- a/operator/internal/controller/persistentvolumeclaim_controller_unit_test.go +++ b/operator/internal/controller/persistentvolumeclaim_controller_unit_test.go @@ -17,6 +17,7 @@ import ( simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/volumemigration" "github.com/simplyblock/simplyblock-operator/internal/webapi" ) @@ -43,6 +44,17 @@ func pinClusterCR() *simplyblockv1alpha2.StorageCluster { } } +// pinStorageNode is the object a backend node UUID is resolved to. The +// redesigned kind names the Kubernetes object rather than the UUID, so a move +// to a node nothing reports cannot be raised at all. +func pinStorageNode(name, uuid string) *simplyblockv1alpha2.StorageNode { + return &simplyblockv1alpha2.StorageNode{ + ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: pinClusterNS}, + Spec: simplyblockv1alpha2.StorageNodeSpec{ClusterRef: pinClusterName}, + Status: simplyblockv1alpha2.StorageNodeStatus{UUID: uuid}, + } +} + // pinAPIServer serves the two control-plane endpoints the PVC controller uses: // the storage-node list (for target validation) and a single volume (for current // placement). currentNode is the storage_node_id reported for the volume. @@ -72,7 +84,8 @@ func pinAPIServer(t *testing.T, nodes []string, currentNode string) string { func newPVCReconciler(t *testing.T, apiURL string, objs ...client.Object) (*PersistentVolumeClaimReconciler, client.Client) { t.Helper() - scheme := newTestScheme(t, simplyblockv1alpha1.AddToScheme, corev1.AddToScheme) + scheme := newTestScheme(t, simplyblockv1alpha1.AddToScheme, + simplyblockv1alpha2.AddToScheme, corev1.AddToScheme) cl := newTestClient(t, scheme, nil, objs...) r := &PersistentVolumeClaimReconciler{ Client: cl, @@ -126,22 +139,24 @@ func getPinPVC(t *testing.T, cl client.Client) *corev1.PersistentVolumeClaim { return pvc } -func listPinMigrations(t *testing.T, cl client.Client) []simplyblockv1alpha1.VolumeMigration { +// listPinMigrations reads the moves this controller raised, whichever kind +// carries them. +func listPinMigrations(t *testing.T, r *PersistentVolumeClaimReconciler) []volumemigration.Move { t.Helper() - var list simplyblockv1alpha1.VolumeMigrationList - if err := cl.List(context.Background(), &list); err != nil { - t.Fatalf("list VolumeMigrations: %v", err) + moves, err := r.mover().List(context.Background(), pinClusterNS, nil) + if err != nil { + t.Fatalf("list the volume moves: %v", err) } - return list.Items + return moves } func TestPVCReconcile_NoChangeGate(t *testing.T) { // desired == applied → nothing happens, no API call. - r, cl := newPVCReconciler(t, unreachableAPI, pinPVC(pinNodeB, pinNodeB), pinPV()) + r, _ := newPVCReconciler(t, unreachableAPI, pinPVC(pinNodeB, pinNodeB), pinPV()) if _, err := r.Reconcile(context.Background(), pinRequest()); err != nil { t.Fatalf("reconcile: %v", err) } - if got := len(listPinMigrations(t, cl)); got != 0 { + if got := len(listPinMigrations(t, r)); got != 0 { t.Fatalf("expected no migrations, got %d", got) } } @@ -156,7 +171,7 @@ func TestPVCReconcile_Unpin(t *testing.T) { if _, ok := pvc.Annotations[kube.AnnoSelectedStorageNodeApplied]; ok { t.Fatalf("expected applied annotation cleared, still present") } - if got := len(listPinMigrations(t, cl)); got != 0 { + if got := len(listPinMigrations(t, r)); got != 0 { t.Fatalf("expected no migrations, got %d", got) } } @@ -164,7 +179,7 @@ func TestPVCReconcile_Unpin(t *testing.T) { func TestPVCReconcile_UnboundRequeues(t *testing.T) { pvc := pinPVC(pinNodeB, "") pvc.Spec.VolumeName = "" - r, cl := newPVCReconciler(t, unreachableAPI, pvc) + r, _ := newPVCReconciler(t, unreachableAPI, pvc) res, err := r.Reconcile(context.Background(), pinRequest()) if err != nil { t.Fatalf("reconcile: %v", err) @@ -172,7 +187,7 @@ func TestPVCReconcile_UnboundRequeues(t *testing.T) { if res.RequeueAfter == 0 { t.Fatalf("expected requeue for unbound PVC") } - if got := len(listPinMigrations(t, cl)); got != 0 { + if got := len(listPinMigrations(t, r)); got != 0 { t.Fatalf("expected no migrations, got %d", got) } } @@ -184,7 +199,7 @@ func TestPVCReconcile_InvalidTarget(t *testing.T) { if _, err := r.Reconcile(context.Background(), pinRequest()); err != nil { t.Fatalf("reconcile: %v", err) } - if got := len(listPinMigrations(t, cl)); got != 0 { + if got := len(listPinMigrations(t, r)); got != 0 { t.Fatalf("expected no migrations for invalid target, got %d", got) } pvc := getPinPVC(t, cl) @@ -204,7 +219,7 @@ func TestPVCReconcile_AlreadyOnTarget(t *testing.T) { if _, err := r.Reconcile(context.Background(), pinRequest()); err != nil { t.Fatalf("reconcile: %v", err) } - if got := len(listPinMigrations(t, cl)); got != 0 { + if got := len(listPinMigrations(t, r)); got != 0 { t.Fatalf("expected no migrations, got %d", got) } if getPinPVC(t, cl).Annotations[kube.AnnoSelectedStorageNodeApplied] != pinNodeB { @@ -229,7 +244,7 @@ func TestPVCReconcile_LegacyHostIDNormalized(t *testing.T) { if _, err := r.Reconcile(context.Background(), pinRequest()); err != nil { t.Fatalf("reconcile: %v", err) } - if got := len(listPinMigrations(t, cl)); got != 0 { + if got := len(listPinMigrations(t, r)); got != 0 { t.Fatalf("expected no migrations (already on node), got %d", got) } got := getPinPVC(t, cl) @@ -248,31 +263,40 @@ func TestPVCReconcile_LegacyHostIDNormalized(t *testing.T) { func TestPVCReconcile_ValidChangeCreatesMigration(t *testing.T) { // Volume on node-a, pin to node-b → create migration + record applied. api := pinAPIServer(t, []string{pinNodeA, pinNodeB}, pinNodeA) - r, cl := newPVCReconciler(t, api, pinPVC(pinNodeB, ""), pinPV(), pinClusterCR()) + r, cl := newPVCReconciler(t, api, pinPVC(pinNodeB, ""), pinPV(), pinClusterCR(), + pinStorageNode("node-a", pinNodeA), pinStorageNode("node-b", pinNodeB)) if _, err := r.Reconcile(context.Background(), pinRequest()); err != nil { t.Fatalf("reconcile: %v", err) } - migs := listPinMigrations(t, cl) + migs := listPinMigrations(t, r) if len(migs) != 1 { - t.Fatalf("expected 1 migration, got %d", len(migs)) + t.Fatalf("expected 1 move, got %d", len(migs)) } - m := migs[0] - if m.Spec.PVName != pinPVName || m.Spec.TargetNodeUUID != pinNodeB { - t.Fatalf("unexpected migration spec: %+v", m.Spec) + if migs[0].PVName != pinPVName { + t.Fatalf("the move names volume %q, want %q", migs[0].PVName, pinPVName) } - // The migration must be created in the StorageCluster's namespace, not the PVC's. - if m.Namespace != pinClusterNS { - t.Fatalf("expected migration in cluster namespace %q, got %q", pinClusterNS, m.Namespace) + + // The move names the node object rather than the backend UUID, and is + // attributed to the cluster that owns it rather than to the claim: a claim + // may live in another namespace, and a cross-namespace reference is invalid. + var raised simplyblockv1alpha2.PersistentVolumeOps + if err := cl.Get(context.Background(), + types.NamespacedName{Name: migs[0].Name}, &raised); err != nil { + t.Fatalf("reading the move back: %v", err) + } + if raised.Spec.Migrate == nil || raised.Spec.Migrate.TargetNodeRef.Name != "node-b" { + t.Fatalf("the move's target is %+v, want the StorageNode reporting %s", + raised.Spec.Migrate, pinNodeB) } - if m.Labels[labelPinnedVolumePV] != pinPVLabelValue(pinPVName) { - t.Fatalf("expected PV label %q, got %q", pinPVLabelValue(pinPVName), m.Labels[labelPinnedVolumePV]) + if raised.Labels[labelPinnedVolumePV] != pinPVLabelValue(pinPVName) { + t.Fatalf("expected PV label %q, got %q", + pinPVLabelValue(pinPVName), raised.Labels[labelPinnedVolumePV]) } - if len(m.OwnerReferences) != 1 || m.OwnerReferences[0].Name != pinClusterName || - m.OwnerReferences[0].Kind != "StorageCluster" { - t.Fatalf("expected owner reference to StorageCluster, got %+v", m.OwnerReferences) + if raised.Spec.CreatorRef == nil || raised.Spec.CreatorRef.Name != pinClusterName { + t.Fatalf("expected the cluster as the creator, got %+v", raised.Spec.CreatorRef) } if getPinPVC(t, cl).Annotations[kube.AnnoSelectedStorageNodeApplied] != pinNodeB { - t.Fatalf("expected applied = %s after creating migration", pinNodeB) + t.Fatalf("expected applied = %s after raising the move", pinNodeB) } } @@ -288,7 +312,7 @@ func TestPVCReconcile_NoStorageCluster(t *testing.T) { if res.RequeueAfter == 0 { t.Fatalf("expected requeue when no StorageCluster manages the cluster") } - if got := len(listPinMigrations(t, cl)); got != 0 { + if got := len(listPinMigrations(t, r)); got != 0 { t.Fatalf("expected no migrations, got %d", got) } if _, ok := getPinPVC(t, cl).Annotations[kube.AnnoSelectedStorageNodeApplied]; ok { @@ -298,17 +322,22 @@ func TestPVCReconcile_NoStorageCluster(t *testing.T) { func TestPVCReconcile_ActiveMigrationWaits(t *testing.T) { // A non-terminal migration for this PV already exists → wait, do not duplicate. - existing := &simplyblockv1alpha1.VolumeMigration{ + existing := &simplyblockv1alpha2.PersistentVolumeOps{ ObjectMeta: metav1.ObjectMeta{ - Name: "existing-mig", - Namespace: pinClusterNS, - Labels: map[string]string{labelPinnedVolumePV: pinPVLabelValue(pinPVName)}, + Name: "existing-move", + Labels: map[string]string{labelPinnedVolumePV: pinPVLabelValue(pinPVName)}, + }, + Spec: simplyblockv1alpha2.PersistentVolumeOpsSpec{ + PersistentVolumeName: pinPVName, + Action: simplyblockv1alpha2.PersistentVolumeOpsActionMigrate, + }, + Status: simplyblockv1alpha2.PersistentVolumeOpsStatus{ + Phase: simplyblockv1alpha2.PersistentVolumeOpsPhaseRunning, }, - Spec: simplyblockv1alpha1.VolumeMigrationSpec{PVName: pinPVName, TargetNodeUUID: pinNodeA}, - Status: simplyblockv1alpha1.VolumeMigrationStatus{Phase: simplyblockv1alpha1.VolumeMigrationPhaseRunning}, } api := pinAPIServer(t, []string{pinNodeA, pinNodeB}, pinNodeA) - r, cl := newPVCReconciler(t, api, pinPVC(pinNodeB, ""), pinPV(), pinClusterCR(), existing) + r, cl := newPVCReconciler(t, api, pinPVC(pinNodeB, ""), pinPV(), pinClusterCR(), + pinStorageNode("node-a", pinNodeA), pinStorageNode("node-b", pinNodeB), existing) res, err := r.Reconcile(context.Background(), pinRequest()) if err != nil { t.Fatalf("reconcile: %v", err) @@ -316,8 +345,8 @@ func TestPVCReconcile_ActiveMigrationWaits(t *testing.T) { if res.RequeueAfter == 0 { t.Fatalf("expected requeue while a migration is in flight") } - if got := len(listPinMigrations(t, cl)); got != 1 { - t.Fatalf("expected only the pre-existing migration, got %d", got) + if got := len(listPinMigrations(t, r)); got != 1 { + t.Fatalf("expected only the pre-existing move, got %d", got) } if _, ok := getPinPVC(t, cl).Annotations[kube.AnnoSelectedStorageNodeApplied]; ok { t.Fatalf("applied must not be set while waiting for an in-flight migration") diff --git a/operator/internal/controller/volumerebalancer_controller.go b/operator/internal/controller/volumerebalancer_controller.go index 405208449..0514e43a5 100644 --- a/operator/internal/controller/volumerebalancer_controller.go +++ b/operator/internal/controller/volumerebalancer_controller.go @@ -13,7 +13,6 @@ import ( apierrors "k8s.io/apimachinery/pkg/api/errors" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" "k8s.io/apimachinery/pkg/runtime" - "k8s.io/apimachinery/pkg/types" "k8s.io/client-go/kubernetes" "k8s.io/client-go/tools/events" ctrl "sigs.k8s.io/controller-runtime" @@ -77,6 +76,11 @@ type VolumeRebalancerReconciler struct { // flag. Empty falls back to the config default (p50). LatencyPercentile string + // Mover raises a volume's move as whichever kind this deployment runs. The + // rebalancer decides which volume moves where and has no business knowing + // which kind carries it. + Mover volumemigration.Mover + migrationState *volumemigration.MigrationState rebalancer *autoplacement.Rebalancer } @@ -252,12 +256,6 @@ func (r *VolumeRebalancerReconciler) executeMigrations( cycleDeadline time.Time, ) int { log := logf.FromContext(ctx) - ownerRefs := []metav1.OwnerReference{{ - APIVersion: simplyblockv1alpha1.GroupVersion.String(), - Kind: "StorageCluster", - Name: clusterCR.Name, - UID: clusterCR.UID, - }} migratedCount := 0 for _, mc := range toMigrate { if time.Now().After(cycleDeadline) { @@ -273,23 +271,29 @@ func (r *VolumeRebalancerReconciler) executeMigrations( rebalancerClusterLabel: clusterCR.Name, } - err := volumemigration.StartMigration(ctx, r.Client, mc.Volume.UUID, mc.TargetNodeUUID, - name, clusterCR.Namespace, ownerRefs, labels) - switch { - case apierrors.IsAlreadyExists(err): - // A VolumeMigration for this volume already exists (in flight, or a - // leftover not yet reaped). Track it and move on rather than duplicating. - log.Info("VolumeMigration CR already exists; tracking existing", "name", name, "volume", mc.Volume.UUID) - case err != nil: - log.Error(err, "Failed to create VolumeMigration CR", "volume", mc.Volume.UUID, "target", mc.TargetNodeUUID) + pvName, err := volumemigration.VolumeFronting(ctx, r.Client, mc.Volume.UUID) + if err == nil { + err = r.Mover.Start(ctx, volumemigration.MoveRequest{ + Name: name, + Namespace: clusterCR.Namespace, + PVName: pvName, + TargetNodeUUID: mc.TargetNodeUUID, + Labels: labels, + Owner: clusterCR, + OwnerKind: "StorageCluster", + Scheme: r.Scheme, + }) + } + if err != nil { + log.Error(err, "Failed to start a volume move", "volume", mc.Volume.UUID, "target", mc.TargetNodeUUID) r.Recorder.Eventf(clusterCR, nil, corev1.EventTypeWarning, "VolumeRebalancingFailed", "VolumeRebalancingFailed", - "Creating VolumeMigration for volume %s to node %s failed: %v", mc.Volume.UUID, mc.TargetNodeUUID, err) + "Moving volume %s to node %s could not be started: %v", mc.Volume.UUID, mc.TargetNodeUUID, err) continue } r.migrationState.PushMigration(mc.ClusterUUID, mc.Volume.PoolUUID, mc.Volume.UUID, name, clusterCR.Namespace, coolDownSecs) r.Recorder.Eventf(clusterCR, nil, corev1.EventTypeNormal, "VolumeRebalancingStarted", "VolumeRebalancingStarted", - "Created VolumeMigration %s for volume %s from node %s to %s", + "Started move %s of volume %s from node %s to %s", name, mc.Volume.UUID, mc.SourceNodeUUID, mc.TargetNodeUUID) rebalancerMigrationsTotal.WithLabelValues(clusterCR.Name, mc.SourceNodeUUID, mc.TargetNodeUUID).Inc() migratedCount++ @@ -318,24 +322,21 @@ func (r *VolumeRebalancerReconciler) processPendingMigrations( } volumeUUID := pm.VolumeUUID - vm := &simplyblockv1alpha1.VolumeMigration{} - err := r.Get(ctx, types.NamespacedName{Name: pm.CRName, Namespace: pm.CRNamespace}, vm) + move, err := r.Mover.Get(ctx, pm.CRName, pm.CRNamespace) if apierrors.IsNotFound(err) { - // CR was deleted out from under us (manual cleanup / GC). Stop tracking. - log.Info("VolumeMigration CR gone; clearing pending", "name", pm.CRName, "volume", volumeUUID) + // The object was deleted out from under us, by hand or by a + // cascade. Stop tracking it. + log.Info("The volume move is gone; clearing pending", "name", pm.CRName, "volume", volumeUUID) r.migrationState.DeletePendingMigration(clusterUUID, volumeUUID) continue } if err != nil { - log.Error(err, "Cannot get VolumeMigration CR", "name", pm.CRName, "volume", volumeUUID) + log.Error(err, "Cannot read the volume move", "name", pm.CRName, "volume", volumeUUID) continue } - phase := vm.Status.Phase - terminal := phase == simplyblockv1alpha1.VolumeMigrationPhaseCompleted || - phase == simplyblockv1alpha1.VolumeMigrationPhaseFailed || - phase == simplyblockv1alpha1.VolumeMigrationPhaseAborted - if !terminal { + phase := move.Phase + if !phase.Terminal() { if time.Since(pm.MigrationStart) > volumemigration.MigrationStuckWarningTimeout && !pm.StuckWarned { log.Error(nil, "Volume migration has not completed within 30 minutes", "volume", volumeUUID, "migration", pm.CRName, "phase", phase) @@ -347,21 +348,21 @@ func (r *VolumeRebalancerReconciler) processPendingMigrations( continue } - // Terminal: record outcome, reap the CR, stop tracking. + // Terminal: record outcome, reap the object, stop tracking. r.migrationState.DeletePendingMigration(clusterUUID, volumeUUID) - if phase == simplyblockv1alpha1.VolumeMigrationPhaseCompleted { - log.Info("Volume migration complete", "volume", volumeUUID, "migration", pm.CRName) + if phase == volumemigration.MoveSucceeded { + log.Info("Volume move complete", "volume", volumeUUID, "migration", pm.CRName) r.Recorder.Eventf(clusterCR, nil, corev1.EventTypeNormal, "VolumeRebalancingComplete", "VolumeRebalancingComplete", - "Migration %s of volume %s completed successfully", pm.CRName, volumeUUID) + "Move %s of volume %s completed successfully", pm.CRName, volumeUUID) } else { - log.Error(nil, "Volume migration ended without success", - "volume", volumeUUID, "migration", pm.CRName, "phase", phase, "error", vm.Status.ErrorMessage) + log.Error(nil, "Volume move ended without success", + "volume", volumeUUID, "migration", pm.CRName, "phase", phase, "error", move.Message) r.Recorder.Eventf(clusterCR, nil, corev1.EventTypeWarning, "VolumeRebalancingFailed", "VolumeRebalancingFailed", - "Migration %s of volume %s ended in phase %s: %s", - pm.CRName, volumeUUID, phase, vm.Status.ErrorMessage) + "Move %s of volume %s ended in phase %s: %s", + pm.CRName, volumeUUID, phase, move.Message) } - if err := r.Delete(ctx, vm); err != nil && !apierrors.IsNotFound(err) { - log.Error(err, "Failed to delete completed VolumeMigration CR", "name", pm.CRName) + if err := r.Mover.Delete(ctx, move); err != nil { + log.Error(err, "Failed to delete the completed volume move", "name", pm.CRName) } } } @@ -612,6 +613,12 @@ func (r *VolumeRebalancerReconciler) SetupWithManager( ) error { r.apiClient = webapi.NewClient() + if r.Mover == nil { + // The kind a deployment says nothing about is the one this API group + // documents, which is what makes the registered kind opt-in. + r.Mover = volumemigration.NewMover(r.Client, r.Scheme, false) + } + // A client-go clientset backs the kube.LiveResolver used to read the // StorageClass of each candidate volume (see BuildNamespacedSet). StorageClass // reads are rare and off the hot path, so direct API reads are preferable to diff --git a/operator/internal/controllers/node/remove.go b/operator/internal/controllers/node/remove.go index 66a65ecfe..8c79e78e2 100644 --- a/operator/internal/controllers/node/remove.go +++ b/operator/internal/controllers/node/remove.go @@ -36,15 +36,12 @@ import ( corev1 "k8s.io/api/core/v1" apierrors "k8s.io/apimachinery/pkg/api/errors" - metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" - "sigs.k8s.io/controller-runtime/pkg/client" - "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil" logf "sigs.k8s.io/controller-runtime/pkg/log" "github.com/simplyblock/atlas/kube" - simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + vmigration "github.com/simplyblock/simplyblock-operator/internal/volumemigration" ) // drainNodeLabel is what a migration this drain created carries, so that a List @@ -257,7 +254,7 @@ func (r *StorageNodeOpsReconciler) drainMigrate( completed, running := 0, 0 for i := range migrations { - if migrations[i].Status.Phase == simplyblockv1alpha1.VolumeMigrationPhaseCompleted { + if migrations[i].Phase == vmigration.MoveSucceeded { completed++ } else { running++ @@ -275,7 +272,7 @@ func (r *StorageNodeOpsReconciler) drainMigrate( } drainVolumesMigratedTotal.WithLabelValues(r.clusterLabel(ctx, ops)).Add(float64(completed)) for i := range migrations { - if err := r.Delete(ctx, &migrations[i]); err != nil && !apierrors.IsNotFound(err) { + if err := r.mover().Delete(ctx, migrations[i]); err != nil { log.Error(err, "a completed migration could not be deleted", "migration", migrations[i].Name) } @@ -356,19 +353,22 @@ func (r *StorageNodeOpsReconciler) drainRemove( // migrationsOf lists the fan-out of this drain. func (r *StorageNodeOpsReconciler) migrationsOf( ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, nodeID string, -) ([]simplyblockv1alpha1.VolumeMigration, error) { - var migrations simplyblockv1alpha1.VolumeMigrationList - err := r.List(ctx, &migrations, - client.InNamespace(ops.Namespace), - client.MatchingLabels{drainNodeLabel: nodeID}) +) ([]vmigration.Move, error) { + moves, err := r.mover().List(ctx, ops.Namespace, map[string]string{drainNodeLabel: nodeID}) if err != nil { return nil, fmt.Errorf("list this drain's volume migrations: %w", err) } - return migrations.Items, nil + return moves, nil } -// createMigration raises one volume's move, owned by the operation that asked for -// it so that deleting the drain cascades to its fan-out. +// createMigration raises one volume's move, recording the operation that asked +// for it so that deleting the drain cascades to its fan-out. +// +// Which kind carries the move is the deployment's, and how the creator is +// recorded follows from it: a namespaced move takes a controller reference, and +// a cluster-scoped one cannot have one at all, so it names its creator in the +// spec and the cascade becomes this controller's own (design-storagenode.md +// §8.4, design-persistentvolumeops.md §11.1). func (r *StorageNodeOpsReconciler) createMigration( ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, @@ -376,24 +376,25 @@ func (r *StorageNodeOpsReconciler) createMigration( volume managedVolume, target string, ) error { - migration := &simplyblockv1alpha1.VolumeMigration{ - ObjectMeta: metav1.ObjectMeta{ - Name: migrationName(nodeID, volume.PVName), - Namespace: ops.Namespace, - Labels: map[string]string{drainNodeLabel: nodeID}, - }, - Spec: simplyblockv1alpha1.VolumeMigrationSpec{ - PVName: volume.PVName, - TargetNodeUUID: target, - }, - } - if err := controllerutil.SetControllerReference(ops, migration, r.Scheme); err != nil { - return fmt.Errorf("own the migration of %s: %w", volume.PVName, err) - } - if err := r.Create(ctx, migration); err != nil && !apierrors.IsAlreadyExists(err) { - return fmt.Errorf("create the migration of %s: %w", volume.PVName, err) - } - return nil + return r.mover().Start(ctx, vmigration.MoveRequest{ + Name: migrationName(nodeID, volume.PVName), + Namespace: ops.Namespace, + PVName: volume.PVName, + TargetNodeUUID: target, + Labels: map[string]string{drainNodeLabel: nodeID}, + Owner: ops, + OwnerKind: "StorageNodeOps", + Scheme: r.Scheme, + }) +} + +// mover is the fan-out's channel, defaulted so a reconciler built without one +// raises the kind this API group documents. +func (r *StorageNodeOpsReconciler) mover() vmigration.Mover { + if r.Mover != nil { + return r.Mover + } + return vmigration.NewMover(r.Client, r.Scheme, false) } // retryFailedMigrations deletes every migration that failed and reports how many, @@ -404,20 +405,20 @@ func (r *StorageNodeOpsReconciler) createMigration( func (r *StorageNodeOpsReconciler) retryFailedMigrations( ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, - migrations []simplyblockv1alpha1.VolumeMigration, + migrations []vmigration.Move, ) (int, error) { retried := 0 for i := range migrations { - migration := &migrations[i] - if migration.Status.Phase != simplyblockv1alpha1.VolumeMigrationPhaseFailed { + migration := migrations[i] + if migration.Phase != vmigration.MoveFailed { continue } r.emit(ctx, ops, corev1.EventTypeWarning, MigrationRetried, fmt.Sprintf( "The migration of %s failed and is being retried against another peer: %s", - migration.Spec.PVName, migration.Status.ErrorMessage)) - if err := r.Delete(ctx, migration); err != nil && !apierrors.IsNotFound(err) { + migration.PVName, migration.Message)) + if err := r.mover().Delete(ctx, migration); err != nil { return retried, fmt.Errorf("delete the failed migration of %s: %w", - migration.Spec.PVName, err) + migration.PVName, err) } retried++ } @@ -445,13 +446,13 @@ func (r *StorageNodeOpsReconciler) cascadeMigrations( } pending := false for i := range migrations { - migration := &migrations[i] - if !terminalMigration(migration.Status.Phase) { + migration := migrations[i] + if !migration.Phase.Terminal() { pending = true continue } - if err := r.Delete(ctx, migration); err != nil && !apierrors.IsNotFound(err) { - return true, fmt.Errorf("delete the migration of %s: %w", migration.Spec.PVName, err) + if err := r.mover().Delete(ctx, migration); err != nil { + return true, fmt.Errorf("delete the migration of %s: %w", migration.PVName, err) } } return pending, nil @@ -471,18 +472,6 @@ func (r *StorageNodeOpsReconciler) abortMigrations( } } -// terminalMigration reports a phase a volume migration can never leave. -func terminalMigration(phase simplyblockv1alpha1.VolumeMigrationPhase) bool { - switch phase { - case simplyblockv1alpha1.VolumeMigrationPhaseCompleted, - simplyblockv1alpha1.VolumeMigrationPhaseFailed, - simplyblockv1alpha1.VolumeMigrationPhaseAborted: - return true - default: - return false - } -} - // migrationFormula names one volume's move. It is derived rather than generated // so that the fan-out is idempotent: a pass that runs again finds the object it // made rather than making a second. diff --git a/operator/internal/controllers/node/remove_fanout_test.go b/operator/internal/controllers/node/remove_fanout_test.go new file mode 100644 index 000000000..41d64af50 --- /dev/null +++ b/operator/internal/controllers/node/remove_fanout_test.go @@ -0,0 +1,189 @@ +// A drain's fan-out, and what carries it. +// +// The drain decides which volumes move where. Which kind carries a move is the +// deployment's, and the difference between the two is not cosmetic: a namespaced +// VolumeMigration can be owned by the operation that raised it, and a +// cluster-scoped PersistentVolumeOps cannot — Kubernetes treats a namespaced +// owner of a cluster-scoped object as unresolvable and garbage-collects the +// dependent, which here would delete the migration mid-copy. The creator moves +// into the spec instead, and the cascade becomes this controller's own. +// +// design-storagenode.md §8.4 and design-persistentvolumeops.md §11.1. + +package node + +import ( + "context" + "testing" + + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/runtime" + "k8s.io/client-go/tools/events" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/testsupport" + vmigration "github.com/simplyblock/simplyblock-operator/internal/volumemigration" +) + +const ( + // aDrainedNodeID is the backend node the drain is emptying, which is also + // what its fan-out is labeled by. + aDrainedNodeID = "node-uuid" + // aPeerNodeID is where the volumes go. + aPeerNodeID = "peer-uuid" + // aDrainUID is what separates this drain from a later one of the same name. + aDrainUID = "ops-uid" +) + +// aFanOut builds a reconciler whose world holds the drained node, a peer to +// move to, and the volume being moved. +func aFanOut(t *testing.T, mover func(client.Client, *runtime.Scheme) vmigration.Mover) ( + *StorageNodeOpsReconciler, client.Client, +) { + t.Helper() + scheme := testsupport.NewScheme(t) + + node := &simplyblockv1alpha2.StorageNode{ + ObjectMeta: metav1.ObjectMeta{Name: "a-node", Namespace: "simplyblock"}, + Spec: simplyblockv1alpha2.StorageNodeSpec{ClusterRef: "a-cluster"}, + } + node.Status.UUID = aDrainedNodeID + + peer := &simplyblockv1alpha2.StorageNode{ + ObjectMeta: metav1.ObjectMeta{Name: "a-peer", Namespace: "simplyblock"}, + Spec: simplyblockv1alpha2.StorageNodeSpec{ClusterRef: "a-cluster"}, + } + peer.Status.UUID = aPeerNodeID + + ops := aRemoveOps() + ops.UID = aDrainUID + + apiClient := fake.NewClientBuilder().WithScheme(scheme). + WithObjects(node, peer, ops). + WithStatusSubresource(&simplyblockv1alpha2.StorageNodeOps{}). + Build() + + return &StorageNodeOpsReconciler{ + Client: apiClient, + Scheme: scheme, + Recorder: events.NewFakeRecorder(64), + API: goneControlPlane{}, + Mover: mover(apiClient, scheme), + }, apiClient +} + +// TestTheFanOutRecordsItsCreatorWithoutOwningTheOperation. The cluster-scoped +// kind cannot be owned, so the drain that raised a move has to be findable from +// the move itself: by the label, which is what a List selects on, and by the +// UID, which is what separates one drain from a later drain of the same name. +func TestTheFanOutRecordsItsCreatorWithoutOwningTheOperation(t *testing.T) { + r, apiClient := aFanOut(t, func(c client.Client, s *runtime.Scheme) vmigration.Mover { + return vmigration.NewMover(c, s, false) + }) + + err := r.createMigration(context.Background(), aRemoveOpsWithUID(), aDrainedNodeID, + managedVolume{PVName: "pv-1", VolumeUUID: "volume-uuid"}, aPeerNodeID) + if err != nil { + t.Fatalf("raising the move: %v", err) + } + + var operations simplyblockv1alpha2.PersistentVolumeOpsList + if err := apiClient.List(context.Background(), &operations); err != nil { + t.Fatal(err) + } + if len(operations.Items) != 1 { + t.Fatalf("raised %d operations, want 1", len(operations.Items)) + } + ops := operations.Items[0] + + if len(ops.OwnerReferences) != 0 { + t.Errorf("the operation carries owner references %v, and a namespaced owner of a "+ + "cluster-scoped object is garbage-collected", ops.OwnerReferences) + } + if ops.Spec.CreatorRef == nil { + t.Fatal("nothing records which drain raised the move") + } + if ops.Spec.CreatorRef.UID != aDrainUID { + t.Errorf("the creator's UID is %q, so a drain recreated under the same name would "+ + "inherit a fan-out it did not issue", ops.Spec.CreatorRef.UID) + } + if ops.Spec.Migrate == nil || ops.Spec.Migrate.TargetNodeRef.Name != "a-peer" { + t.Errorf("the move's target is %+v, want the StorageNode reporting peer-uuid", + ops.Spec.Migrate) + } + if ops.Labels[drainNodeLabel] != aDrainedNodeID { + t.Errorf("labels = %v, want the one a List selects the fan-out by", ops.Labels) + } +} + +// TestTheLegacyFanOutIsStillOwnedByItsDrain. With the registered kind turned +// back on, the move is namespaced and the owner reference is what cascades, as +// it always did. +func TestTheLegacyFanOutIsStillOwnedByItsDrain(t *testing.T) { + r, apiClient := aFanOut(t, func(c client.Client, s *runtime.Scheme) vmigration.Mover { + return vmigration.NewMover(c, s, true) + }) + + err := r.createMigration(context.Background(), aRemoveOpsWithUID(), aDrainedNodeID, + managedVolume{PVName: "pv-1", VolumeUUID: "volume-uuid"}, aPeerNodeID) + if err != nil { + t.Fatalf("raising the move: %v", err) + } + + var migrations simplyblockv1alpha1.VolumeMigrationList + if err := apiClient.List(context.Background(), &migrations); err != nil { + t.Fatal(err) + } + if len(migrations.Items) != 1 { + t.Fatalf("raised %d migrations, want 1", len(migrations.Items)) + } + migration := migrations.Items[0] + + if len(migration.OwnerReferences) != 1 || migration.OwnerReferences[0].UID != aDrainUID { + t.Errorf("owner references = %v, want the drain that raised it", + migration.OwnerReferences) + } + if migration.Spec.TargetNodeUUID != "peer-uuid" { + t.Errorf("target = %q, want the backend UUID the registered kind takes", + migration.Spec.TargetNodeUUID) + } +} + +// TestTheFanOutIsFoundAgainByItsLabel, which is what makes the drain's progress +// count and its cascade possible at all. +func TestTheFanOutIsFoundAgainByItsLabel(t *testing.T) { + for name, legacy := range map[string]bool{"PersistentVolumeOps": false, "VolumeMigration": true} { + t.Run(name, func(t *testing.T) { + r, _ := aFanOut(t, func(c client.Client, s *runtime.Scheme) vmigration.Mover { + return vmigration.NewMover(c, s, legacy) + }) + ops := aRemoveOpsWithUID() + + if err := r.createMigration(context.Background(), ops, aDrainedNodeID, + managedVolume{PVName: "pv-1", VolumeUUID: "volume-uuid"}, aPeerNodeID); err != nil { + t.Fatal(err) + } + + moves, err := r.migrationsOf(context.Background(), ops, aDrainedNodeID) + if err != nil { + t.Fatal(err) + } + if len(moves) != 1 { + t.Fatalf("found %d moves, want the one that was raised", len(moves)) + } + if moves[0].PVName != "pv-1" { + t.Errorf("the move names volume %q", moves[0].PVName) + } + }) + } +} + +// aRemoveOpsWithUID is the drain the fan-out is attributed to. +func aRemoveOpsWithUID() *simplyblockv1alpha2.StorageNodeOps { + ops := aRemoveOps() + ops.UID = aDrainUID + return ops +} diff --git a/operator/internal/controllers/node/storagenodeops_controller.go b/operator/internal/controllers/node/storagenodeops_controller.go index 8d4cc75ab..9f18fcb56 100644 --- a/operator/internal/controllers/node/storagenodeops_controller.go +++ b/operator/internal/controllers/node/storagenodeops_controller.go @@ -52,6 +52,7 @@ import ( "github.com/simplyblock/simplyblock-operator/internal/cpinformer" "github.com/simplyblock/simplyblock-operator/internal/cpinformer/subscriptions" "github.com/simplyblock/simplyblock-operator/internal/utils" + vmigration "github.com/simplyblock/simplyblock-operator/internal/volumemigration" ) const ( @@ -100,6 +101,11 @@ type StorageNodeOpsReconciler struct { Nodes NodeCache Clusters ClusterCache + // Mover raises a drain's fan-out as whichever kind this deployment runs. A + // drain decides which volumes move where and has no business knowing which + // kind carries them; unset means the kind this API group documents. + Mover vmigration.Mover + // Workload is the storage-plane side of a node: the worker labels, the // storage-node pod, its published DNS name, and the eviction budget a // maintenance window holds. Three of the seven actions touch it, and the diff --git a/operator/internal/volumemigration/mover.go b/operator/internal/volumemigration/mover.go new file mode 100644 index 000000000..8b1c3a454 --- /dev/null +++ b/operator/internal/volumemigration/mover.go @@ -0,0 +1,370 @@ +// One volume's move, raised as whichever kind the deployment uses. +// +// Two kinds express the same operation. The registered VolumeMigration is +// namespaced, names its target by backend UUID, and reaches Completed. The +// redesigned PersistentVolumeOps is cluster-scoped, names its target as a +// StorageNode object, and reaches Succeeded. Three controllers raise moves — +// the auto-rebalancer, a node drain, and the pinned-volume controller — and +// none of them has any business knowing which kind a deployment runs. +// +// So they ask for a move and this decides. The interface is the four things +// every caller does: start one, find its own again, read an outcome, and reap +// it. What a caller cannot do through here is anything specific to one kind, +// which is deliberate: a caller that needed that would be a caller the gate is +// not hiding anything from. +// +// design-persistentvolumeops.md §10 is what the two kinds differ by. + +package volumemigration + +import ( + "context" + "fmt" + + apierrors "k8s.io/apimachinery/pkg/api/errors" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/runtime" + "k8s.io/apimachinery/pkg/types" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil" + + simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// MovePhase is a move's progress, in the one vocabulary both kinds are read +// through. +type MovePhase string + +const ( + // MovePending is a move nothing has started: admitted, holding no lock. + MovePending MovePhase = "Pending" + // MoveRunning is a move in flight, wherever in its own steps it is. + MoveRunning MovePhase = "Running" + // MoveSucceeded is a volume that moved. + MoveSucceeded MovePhase = "Succeeded" + // MoveFailed is one that did not. + MoveFailed MovePhase = "Failed" + // MoveAborted is one stopped on request, or stopped because its volume went + // away. + MoveAborted MovePhase = "Aborted" +) + +// Terminal reports a phase the move never leaves. +func (p MovePhase) Terminal() bool { + switch p { + case MoveSucceeded, MoveFailed, MoveAborted: + return true + default: + return false + } +} + +// Move is one volume's move, as either kind reports it. +type Move struct { + Name string + // Namespace is empty for the cluster-scoped kind, which is what a caller + // reading a move back has to carry rather than assume. + Namespace string + PVName string + Phase MovePhase + // Message is why the phase is what it is, which a caller puts in the event + // it raises about the move. + Message string +} + +// MoveRequest is one volume's move, as a caller asks for it. +type MoveRequest struct { + // Name is what the move is called. A caller picks a name derived from the + // volume rather than a generated one, which is what makes starting the same + // move twice idempotent. + Name string + + // Namespace is where a namespaced move goes. The cluster-scoped kind takes + // it as the namespace its target node is looked for in rather than as its + // own. + Namespace string + + PVName string + + // TargetNodeUUID is the backend identifier of the node to move to, which is + // what every caller holds. The cluster-scoped kind names the Kubernetes + // object instead, and resolving the one to the other is this file's job. + TargetNodeUUID string + + Labels map[string]string + + // Owner is the object that asked for the move, when there is one. The two + // kinds carry it differently and that difference is not the caller's: a + // namespaced move takes a controller reference, and a cluster-scoped one + // cannot, because Kubernetes treats a namespaced owner of a cluster-scoped + // object as unresolvable and garbage-collects the dependent. + Owner client.Object + // OwnerKind is the owner's kind, which the cluster-scoped kind records in + // spec.creatorRef and in the managed-by label. + OwnerKind string + // Scheme resolves the owner's group and version for a controller + // reference. Only the namespaced kind needs it. + Scheme *runtime.Scheme +} + +// Mover raises and tracks volume moves. +type Mover interface { + // Start raises one move. A move that already exists is the state being + // asked for and is not an error. + Start(ctx context.Context, request MoveRequest) error + + // List finds the moves in a namespace carrying every one of the labels. The + // namespace is ignored by the cluster-scoped kind, which has none. + List(ctx context.Context, namespace string, labels map[string]string) ([]Move, error) + + // Get reads one move back. It returns a wrapped NotFound when there is + // none, which every caller treats as "stop tracking it." + Get(ctx context.Context, name, namespace string) (Move, error) + + // Delete reaps one. A move that is already gone is not an error. + Delete(ctx context.Context, move Move) error +} + +// MigrationMover raises the registered VolumeMigration. +type MigrationMover struct { + client.Client + Scheme *runtime.Scheme +} + +func (m *MigrationMover) Start(ctx context.Context, request MoveRequest) error { + migration := &simplyblockv1alpha1.VolumeMigration{ + ObjectMeta: metav1.ObjectMeta{ + Name: request.Name, + Namespace: request.Namespace, + Labels: request.Labels, + }, + Spec: simplyblockv1alpha1.VolumeMigrationSpec{ + PVName: request.PVName, + TargetNodeUUID: request.TargetNodeUUID, + }, + } + if request.Owner != nil { + scheme := request.Scheme + if scheme == nil { + scheme = m.Scheme + } + if err := controllerutil.SetControllerReference(request.Owner, migration, scheme); err != nil { + return fmt.Errorf("own the migration of %s: %w", request.PVName, err) + } + } + if err := m.Create(ctx, migration); err != nil && !apierrors.IsAlreadyExists(err) { + return fmt.Errorf("create the migration of %s: %w", request.PVName, err) + } + return nil +} + +func (m *MigrationMover) List( + ctx context.Context, namespace string, labels map[string]string, +) ([]Move, error) { + var migrations simplyblockv1alpha1.VolumeMigrationList + options := []client.ListOption{client.InNamespace(namespace)} + if len(labels) > 0 { + options = append(options, client.MatchingLabels(labels)) + } + if err := m.Client.List(ctx, &migrations, options...); err != nil { + return nil, fmt.Errorf("list the volume migrations: %w", err) + } + moves := make([]Move, 0, len(migrations.Items)) + for i := range migrations.Items { + moves = append(moves, migrationMove(&migrations.Items[i])) + } + return moves, nil +} + +func (m *MigrationMover) Get(ctx context.Context, name, namespace string) (Move, error) { + var migration simplyblockv1alpha1.VolumeMigration + if err := m.Client.Get(ctx, + types.NamespacedName{Name: name, Namespace: namespace}, &migration); err != nil { + return Move{}, err + } + return migrationMove(&migration), nil +} + +func (m *MigrationMover) Delete(ctx context.Context, move Move) error { + migration := &simplyblockv1alpha1.VolumeMigration{ + ObjectMeta: metav1.ObjectMeta{Name: move.Name, Namespace: move.Namespace}, + } + if err := m.Client.Delete(ctx, migration); err != nil && !apierrors.IsNotFound(err) { + return fmt.Errorf("delete the migration %s: %w", move.Name, err) + } + return nil +} + +// migrationMove reads the registered kind's merged phase-and-step enum into the +// shared vocabulary. Validating is one of its values and is a step rather than +// an outcome, which is the confusion the redesign's split exists to end. +func migrationMove(migration *simplyblockv1alpha1.VolumeMigration) Move { + phase := MovePending + switch migration.Status.Phase { + case simplyblockv1alpha1.VolumeMigrationPhaseCompleted: + phase = MoveSucceeded + case simplyblockv1alpha1.VolumeMigrationPhaseFailed: + phase = MoveFailed + case simplyblockv1alpha1.VolumeMigrationPhaseAborted: + phase = MoveAborted + case simplyblockv1alpha1.VolumeMigrationPhaseRunning, + simplyblockv1alpha1.VolumeMigrationPhaseValidating: + phase = MoveRunning + } + return Move{ + Name: migration.Name, + Namespace: migration.Namespace, + PVName: migration.Spec.PVName, + Phase: phase, + Message: migration.Status.ErrorMessage, + } +} + +// OperationMover raises the redesigned PersistentVolumeOps. +type OperationMover struct { + client.Client + Scheme *runtime.Scheme +} + +// Start resolves the target node's backend UUID to the StorageNode object the +// kind names, and records the creator rather than owning the object. +// +// The creator reference is what replaces the owner reference a cluster-scoped +// object cannot have. It carries the UID, for the reason pv.spec.claimRef does: +// a creator deleted and recreated under the same name must not inherit a +// fan-out it did not issue. +func (m *OperationMover) Start(ctx context.Context, request MoveRequest) error { + node, err := m.nodeReporting(ctx, request.Namespace, request.TargetNodeUUID) + if err != nil { + return err + } + + labels := map[string]string{} + for key, value := range request.Labels { + labels[key] = value + } + + ops := &simplyblockv1alpha2.PersistentVolumeOps{ + ObjectMeta: metav1.ObjectMeta{Name: request.Name, Labels: labels}, + Spec: simplyblockv1alpha2.PersistentVolumeOpsSpec{ + PersistentVolumeName: request.PVName, + Action: simplyblockv1alpha2.PersistentVolumeOpsActionMigrate, + Migrate: &simplyblockv1alpha2.MigrateVolumeSpec{ + TargetNodeRef: simplyblockv1alpha2.StorageNodeReference{ + Namespace: node.Namespace, + Name: node.Name, + }, + }, + }, + } + + if request.Owner != nil { + ops.Spec.CreatorRef = &simplyblockv1alpha2.CreatorReference{ + Kind: request.OwnerKind, + Namespace: request.Owner.GetNamespace(), + Name: request.Owner.GetName(), + UID: request.Owner.GetUID(), + } + // A reference cannot be selected on, so the kind travels as a label + // too: the label finds the operations some creator raised, and the UID + // says which one. + labels[simplyblockv1alpha2.PersistentVolumeOpsManagedByLabel] = request.OwnerKind + } + + if err := m.Create(ctx, ops); err != nil && !apierrors.IsAlreadyExists(err) { + return fmt.Errorf("create the operation moving %s: %w", request.PVName, err) + } + return nil +} + +// nodeReporting finds the StorageNode whose status reports this backend UUID. +// +// A UUID nothing reports is refused rather than guessed at. The kind names a +// Kubernetes object, so an operation written against a node that has none would +// be refused at admission with a worse message than this one. +func (m *OperationMover) nodeReporting( + ctx context.Context, namespace, uuid string, +) (*simplyblockv1alpha2.StorageNode, error) { + var nodes simplyblockv1alpha2.StorageNodeList + if err := m.Client.List(ctx, &nodes, client.InNamespace(namespace)); err != nil { + return nil, fmt.Errorf("list the storage nodes of namespace %s: %w", namespace, err) + } + for i := range nodes.Items { + if nodes.Items[i].Status.UUID == uuid { + return &nodes.Items[i], nil + } + } + return nil, fmt.Errorf("no StorageNode in namespace %s reports node %s", namespace, uuid) +} + +func (m *OperationMover) List( + ctx context.Context, _ string, labels map[string]string, +) ([]Move, error) { + var operations simplyblockv1alpha2.PersistentVolumeOpsList + var options []client.ListOption + if len(labels) > 0 { + options = append(options, client.MatchingLabels(labels)) + } + if err := m.Client.List(ctx, &operations, options...); err != nil { + return nil, fmt.Errorf("list the volume operations: %w", err) + } + moves := make([]Move, 0, len(operations.Items)) + for i := range operations.Items { + moves = append(moves, operationMove(&operations.Items[i])) + } + return moves, nil +} + +func (m *OperationMover) Get(ctx context.Context, name, _ string) (Move, error) { + var ops simplyblockv1alpha2.PersistentVolumeOps + if err := m.Client.Get(ctx, types.NamespacedName{Name: name}, &ops); err != nil { + return Move{}, err + } + return operationMove(&ops), nil +} + +func (m *OperationMover) Delete(ctx context.Context, move Move) error { + ops := &simplyblockv1alpha2.PersistentVolumeOps{ + ObjectMeta: metav1.ObjectMeta{Name: move.Name}, + } + if err := m.Client.Delete(ctx, ops); err != nil && !apierrors.IsNotFound(err) { + return fmt.Errorf("delete the operation %s: %w", move.Name, err) + } + return nil +} + +func operationMove(ops *simplyblockv1alpha2.PersistentVolumeOps) Move { + phase := MovePending + switch ops.Status.Phase { + case simplyblockv1alpha2.PersistentVolumeOpsPhaseSucceeded: + phase = MoveSucceeded + case simplyblockv1alpha2.PersistentVolumeOpsPhaseFailed: + phase = MoveFailed + case simplyblockv1alpha2.PersistentVolumeOpsPhaseAborted: + phase = MoveAborted + case simplyblockv1alpha2.PersistentVolumeOpsPhaseRunning: + phase = MoveRunning + } + return Move{ + Name: ops.Name, + PVName: ops.Spec.PersistentVolumeName, + Phase: phase, + Message: ops.Status.Message, + } +} + +// NewMover builds the mover a deployment uses. +// +// The redesigned kind is the default and the registered one is opt-in, which is +// the way round it is because a deployment that says nothing should be running +// the kind this API group documents. The legacy kind is what an upgrade turns +// back on while migrations raised against it drain, since a rename and a scope +// change make a new CRD rather than a new version and an in-flight migration +// cannot be carried across. +func NewMover(c client.Client, scheme *runtime.Scheme, legacy bool) Mover { + if legacy { + return &MigrationMover{Client: c, Scheme: scheme} + } + return &OperationMover{Client: c, Scheme: scheme} +} diff --git a/operator/internal/volumemigration/mover_test.go b/operator/internal/volumemigration/mover_test.go new file mode 100644 index 000000000..4f57fca4e --- /dev/null +++ b/operator/internal/volumemigration/mover_test.go @@ -0,0 +1,297 @@ +// The two kinds a volume move can be raised as, behind one interface, and the +// properties that have to hold for either of them. +// +// The tests are written against the interface rather than against each +// implementation, because what the three callers depend on is that a move can +// be started, found again, read for an outcome, and removed. Which kind carries +// it is the deployment's choice. + +package volumemigration + +import ( + "context" + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/runtime" + k8sscheme "k8s.io/client-go/kubernetes/scheme" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + "github.com/simplyblock/atlas/lvol" + + simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +const ( + moveClusterID = "11111111-1111-1111-1111-111111111111" + movePoolID = "22222222-2222-2222-2222-222222222222" + moveVolumeID = "33333333-3333-3333-3333-333333333333" + moveTargetID = "55555555-5555-5555-5555-555555555555" + movePVName = "pvc-" + moveVolumeID + moveNodeName = "worker-5" +) + +func moverScheme(t *testing.T) *runtime.Scheme { + t.Helper() + s := runtime.NewScheme() + for _, add := range []func(*runtime.Scheme) error{ + k8sscheme.AddToScheme, + simplyblockv1alpha1.AddToScheme, + simplyblockv1alpha2.AddToScheme, + } { + if err := add(s); err != nil { + t.Fatalf("build the scheme: %v", err) + } + } + return s +} + +// moveWorld is what either kind needs to exist: the volume, and the node the +// move is aimed at. +func moveWorld() []client.Object { + return []client.Object{ + &corev1.PersistentVolume{ + ObjectMeta: metav1.ObjectMeta{Name: movePVName}, + Spec: corev1.PersistentVolumeSpec{ + PersistentVolumeSource: corev1.PersistentVolumeSource{ + CSI: &corev1.CSIPersistentVolumeSource{ + Driver: "csi.simplyblock.io", + VolumeHandle: string(lvol.NewVolumeHandle( + moveClusterID, movePoolID, moveVolumeID)), + }, + }, + }, + }, + &simplyblockv1alpha2.StorageNode{ + ObjectMeta: metav1.ObjectMeta{Name: moveNodeName, Namespace: testNamespace}, + Status: simplyblockv1alpha2.StorageNodeStatus{UUID: moveTargetID}, + }, + } +} + +func bothMovers(t *testing.T, objs ...client.Object) map[string]Mover { + t.Helper() + s := moverScheme(t) + build := func() client.Client { + return fake.NewClientBuilder().WithScheme(s). + WithStatusSubresource( + &simplyblockv1alpha1.VolumeMigration{}, + &simplyblockv1alpha2.PersistentVolumeOps{}, + ). + WithObjects(append(moveWorld(), objs...)...).Build() + } + return map[string]Mover{ + "PersistentVolumeOps": &OperationMover{Client: build()}, + "VolumeMigration": &MigrationMover{Client: build()}, + } +} + +func moveRequest() MoveRequest { + return MoveRequest{ + Name: "move-1", + Namespace: testNamespace, + PVName: movePVName, + TargetNodeUUID: moveTargetID, + Labels: map[string]string{"drain-node": "node-a"}, + } +} + +// TestAStartedMoveIsFoundAgainByItsLabel. Every caller raises a move and then +// looks for it later, by the label it tagged it with: the rebalancer to track +// completion, the drain to count its fan-out. +func TestAStartedMoveIsFoundAgainByItsLabel(t *testing.T) { + for kind, mover := range bothMovers(t) { + t.Run(kind, func(t *testing.T) { + ctx := context.Background() + if err := mover.Start(ctx, moveRequest()); err != nil { + t.Fatalf("starting the move: %v", err) + } + + moves, err := mover.List(ctx, testNamespace, map[string]string{"drain-node": "node-a"}) + if err != nil { + t.Fatal(err) + } + if len(moves) != 1 { + t.Fatalf("found %d moves, want the one that was started: %+v", len(moves), moves) + } + if moves[0].PVName != movePVName { + t.Errorf("the move names volume %q, want %q", moves[0].PVName, movePVName) + } + if moves[0].Phase != MovePending { + t.Errorf("a move nothing has reconciled is %q, want Pending", moves[0].Phase) + } + }) + } +} + +// TestStartingTheSameMoveTwiceIsNotTwoMoves. A caller that crashed between +// creating the move and recording it retries, and two moves of one volume would +// be two backend migrations copying it to two places. +func TestStartingTheSameMoveTwiceIsNotTwoMoves(t *testing.T) { + for kind, mover := range bothMovers(t) { + t.Run(kind, func(t *testing.T) { + ctx := context.Background() + if err := mover.Start(ctx, moveRequest()); err != nil { + t.Fatal(err) + } + if err := mover.Start(ctx, moveRequest()); err != nil { + t.Fatalf("starting an existing move reported an error: %v", err) + } + + moves, err := mover.List(ctx, testNamespace, nil) + if err != nil { + t.Fatal(err) + } + if len(moves) != 1 { + t.Errorf("found %d moves, want 1", len(moves)) + } + }) + } +} + +// TestAMovesOutcomeReadsTheSameForEitherKind. The registered kind reaches +// Completed and the redesigned one reaches Succeeded, and a caller that had to +// know which would be a caller the gate is not hiding anything from. +func TestAMovesOutcomeReadsTheSameForEitherKind(t *testing.T) { + ctx := context.Background() + + t.Run("VolumeMigration", func(t *testing.T) { + for _, tc := range []struct { + phase simplyblockv1alpha1.VolumeMigrationPhase + want MovePhase + }{ + {simplyblockv1alpha1.VolumeMigrationPhaseCompleted, MoveSucceeded}, + {simplyblockv1alpha1.VolumeMigrationPhaseFailed, MoveFailed}, + {simplyblockv1alpha1.VolumeMigrationPhaseAborted, MoveAborted}, + {simplyblockv1alpha1.VolumeMigrationPhaseRunning, MoveRunning}, + {simplyblockv1alpha1.VolumeMigrationPhaseValidating, MoveRunning}, + } { + t.Run(string(tc.phase), func(t *testing.T) { + existing := &simplyblockv1alpha1.VolumeMigration{ + ObjectMeta: metav1.ObjectMeta{Name: "move-1", Namespace: testNamespace}, + Spec: simplyblockv1alpha1.VolumeMigrationSpec{PVName: movePVName}, + Status: simplyblockv1alpha1.VolumeMigrationStatus{Phase: tc.phase}, + } + mover := bothMovers(t, existing)["VolumeMigration"] + + got, err := mover.Get(ctx, "move-1", testNamespace) + if err != nil { + t.Fatal(err) + } + if got.Phase != tc.want { + t.Errorf("phase %q reads as %q, want %q", tc.phase, got.Phase, tc.want) + } + }) + } + }) + + t.Run("PersistentVolumeOps", func(t *testing.T) { + for _, tc := range []struct { + phase simplyblockv1alpha2.PersistentVolumeOpsPhase + want MovePhase + }{ + {simplyblockv1alpha2.PersistentVolumeOpsPhaseSucceeded, MoveSucceeded}, + {simplyblockv1alpha2.PersistentVolumeOpsPhaseFailed, MoveFailed}, + {simplyblockv1alpha2.PersistentVolumeOpsPhaseAborted, MoveAborted}, + {simplyblockv1alpha2.PersistentVolumeOpsPhaseRunning, MoveRunning}, + {simplyblockv1alpha2.PersistentVolumeOpsPhasePending, MovePending}, + } { + t.Run(string(tc.phase), func(t *testing.T) { + existing := &simplyblockv1alpha2.PersistentVolumeOps{ + ObjectMeta: metav1.ObjectMeta{Name: "move-1"}, + Spec: simplyblockv1alpha2.PersistentVolumeOpsSpec{ + PersistentVolumeName: movePVName, + Action: simplyblockv1alpha2.PersistentVolumeOpsActionMigrate, + }, + Status: simplyblockv1alpha2.PersistentVolumeOpsStatus{Phase: tc.phase}, + } + mover := bothMovers(t, existing)["PersistentVolumeOps"] + + got, err := mover.Get(ctx, "move-1", testNamespace) + if err != nil { + t.Fatal(err) + } + if got.Phase != tc.want { + t.Errorf("phase %q reads as %q, want %q", tc.phase, got.Phase, tc.want) + } + }) + } + }) +} + +// TestTheOperationNamesTheTargetAsAnObject. The redesigned kind takes a +// StorageNode name rather than a backend UUID, so that a migration can be +// written by hand without looking one up. The callers hold a UUID, so this is +// where the two are joined. +func TestTheOperationNamesTheTargetAsAnObject(t *testing.T) { + ctx := context.Background() + c := fake.NewClientBuilder().WithScheme(moverScheme(t)).WithObjects(moveWorld()...).Build() + mover := &OperationMover{Client: c} + + if err := mover.Start(ctx, moveRequest()); err != nil { + t.Fatal(err) + } + + moves, err := mover.List(ctx, testNamespace, nil) + if err != nil { + t.Fatal(err) + } + if len(moves) != 1 { + t.Fatalf("found %d moves, want 1", len(moves)) + } + if moves[0].Namespace != "" { + t.Errorf("the move is in namespace %q, and the kind is cluster-scoped", moves[0].Namespace) + } + + var ops simplyblockv1alpha2.PersistentVolumeOps + if err := c.Get(ctx, client.ObjectKey{Name: "move-1"}, &ops); err != nil { + t.Fatal(err) + } + if ops.Spec.Migrate == nil || ops.Spec.Migrate.TargetNodeRef.Name != moveNodeName { + t.Errorf("the operation's target is %+v, want the StorageNode object", ops.Spec.Migrate) + } + if ops.Spec.Migrate.TargetNodeRef.Namespace != testNamespace { + t.Errorf("the target carries namespace %q", ops.Spec.Migrate.TargetNodeRef.Namespace) + } +} + +// TestAnOperationForANodeNobodyReportsIsRefused. A backend node UUID with no +// StorageNode reporting it cannot be named as an object, and guessing would +// produce an operation the webhook then refuses with a worse message. +func TestAnOperationForANodeNobodyReportsIsRefused(t *testing.T) { + mover := bothMovers(t)["PersistentVolumeOps"] + + request := moveRequest() + request.TargetNodeUUID = "99999999-9999-9999-9999-999999999999" + + if err := mover.Start(context.Background(), request); err == nil { + t.Fatal("a move to a node nothing reports was accepted") + } +} + +// TestADeletedMoveIsGone, which is how every caller reaps a finished one. +func TestADeletedMoveIsGone(t *testing.T) { + for kind, mover := range bothMovers(t) { + t.Run(kind, func(t *testing.T) { + ctx := context.Background() + if err := mover.Start(ctx, moveRequest()); err != nil { + t.Fatal(err) + } + moves, err := mover.List(ctx, testNamespace, nil) + if err != nil { + t.Fatal(err) + } + + if err := mover.Delete(ctx, moves[0]); err != nil { + t.Fatal(err) + } + // Deleting one that is already gone is the state being asked for. + if err := mover.Delete(ctx, moves[0]); err != nil { + t.Errorf("deleting an absent move reported an error: %v", err) + } + }) + } +} diff --git a/operator/internal/volumemigration/utils.go b/operator/internal/volumemigration/utils.go index bc341b6ff..b0d0c65a9 100644 --- a/operator/internal/volumemigration/utils.go +++ b/operator/internal/volumemigration/utils.go @@ -8,10 +8,8 @@ import ( "time" corev1 "k8s.io/api/core/v1" - metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" "sigs.k8s.io/controller-runtime/pkg/client" - simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" "github.com/simplyblock/simplyblock-operator/internal/webapi" ) @@ -26,44 +24,15 @@ const ( MigrationStuckWarningTimeout = 30 * time.Minute ) -// StartMigration creates a VolumeMigration CR for the given volume UUID and -// target node. It first resolves the volume UUID to a PV name by scanning all -// PersistentVolumes for a matching CSI volume handle. -// name is used as the VolumeMigration object name; namespace is the namespace -// to create it in; ownerRefs lets the caller attach owner references (e.g. -// a StorageCluster) so the CR is garbage-collected with the owner. -func StartMigration( - ctx context.Context, - c client.Client, - volumeUUID, targetNodeUUID, name, namespace string, - ownerRefs []metav1.OwnerReference, - labels map[string]string, -) error { - pvName, err := findPVForVolume(ctx, c, volumeUUID) - if err != nil { - return fmt.Errorf("resolve PV for volume %s: %w", volumeUUID, err) - } - vm := &simplyblockv1alpha1.VolumeMigration{ - ObjectMeta: metav1.ObjectMeta{ - Name: name, - Namespace: namespace, - OwnerReferences: ownerRefs, - Labels: labels, - }, - Spec: simplyblockv1alpha1.VolumeMigrationSpec{ - PVName: pvName, - TargetNodeUUID: targetNodeUUID, - }, - } - return c.Create(ctx, vm) -} - -// findPVForVolume returns the PV name backing the given simplyblock logical-volume -// UUID. simplyblock CSI volume handles have the form -// "::", so the bare volume UUID is matched -// against the final ":"-separated segment. An exact match against the whole -// handle is also accepted for robustness. -func findPVForVolume( +// VolumeFronting returns the name of the PersistentVolume backing the given +// simplyblock logical-volume UUID. +// +// Every caller that moves a volume holds its backend UUID and has to name the +// Kubernetes object instead, because that is what both kinds of move take. CSI +// volume handles have the form "::", so the bare volume +// UUID is matched against the final segment; an exact match against the whole +// handle is also accepted, for a caller that already holds one. +func VolumeFronting( ctx context.Context, c client.Client, volumeUUID string, diff --git a/operator/internal/volumemigration/utils_test.go b/operator/internal/volumemigration/utils_test.go index 61c15594e..f4c36fd46 100644 --- a/operator/internal/volumemigration/utils_test.go +++ b/operator/internal/volumemigration/utils_test.go @@ -11,7 +11,6 @@ import ( corev1 "k8s.io/api/core/v1" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" "k8s.io/apimachinery/pkg/runtime" - "k8s.io/apimachinery/pkg/types" "sigs.k8s.io/controller-runtime/pkg/client" "sigs.k8s.io/controller-runtime/pkg/client/fake" @@ -109,7 +108,7 @@ func TestFindPVForVolume(t *testing.T) { for _, tc := range cases { t.Run(tc.name, func(t *testing.T) { c := fake.NewClientBuilder().WithScheme(utilsScheme(t)).WithObjects(tc.objs...).Build() - got, err := findPVForVolume(context.Background(), c, tc.volume) + got, err := VolumeFronting(context.Background(), c, tc.volume) if tc.wantErr { if err == nil { t.Fatalf("expected an error, got PV %q", got) @@ -120,7 +119,7 @@ func TestFindPVForVolume(t *testing.T) { return } if err != nil { - t.Fatalf("findPVForVolume: %v", err) + t.Fatalf("VolumeFronting: %v", err) } if got != tc.wantPV { t.Errorf("PV = %q, want %q", got, tc.wantPV) @@ -129,71 +128,6 @@ func TestFindPVForVolume(t *testing.T) { } } -func TestStartMigration(t *testing.T) { - owner := []metav1.OwnerReference{{ - APIVersion: "storage.simplyblock.io/v1alpha1", - Kind: "StorageCluster", - Name: "cluster", - UID: "uid-1", - }} - labels := map[string]string{"app.kubernetes.io/created-by": "rebalancer"} - - t.Run("creates the CR pointing at the resolved PV", func(t *testing.T) { - c := fake.NewClientBuilder().WithScheme(utilsScheme(t)). - WithObjects(pvWithHandle("pv-1", utilsCluster+":"+utilsPool+":"+utilsVolume)).Build() - - if err := StartMigration(context.Background(), c, utilsVolume, "target-node", - "vmig-1", "sb", owner, labels); err != nil { - t.Fatalf("StartMigration: %v", err) - } - - var vm simplyblockv1alpha1.VolumeMigration - if err := c.Get(context.Background(), - types.NamespacedName{Namespace: "sb", Name: "vmig-1"}, &vm); err != nil { - t.Fatalf("created VolumeMigration not found: %v", err) - } - if vm.Spec.PVName != "pv-1" { - t.Errorf("PVName = %q, want pv-1 (resolved from the volume UUID)", vm.Spec.PVName) - } - if vm.Spec.TargetNodeUUID != "target-node" { - t.Errorf("TargetNodeUUID = %q, want target-node", vm.Spec.TargetNodeUUID) - } - // Owner references matter: the CR must be garbage-collected with its owner - // rather than outliving the cluster that scheduled it. - if len(vm.OwnerReferences) != 1 || vm.OwnerReferences[0].Name != "cluster" { - t.Errorf("OwnerReferences = %+v, want the passed owner", vm.OwnerReferences) - } - if vm.Labels["app.kubernetes.io/created-by"] != "rebalancer" { - t.Errorf("Labels = %v, want the passed labels", vm.Labels) - } - }) - - t.Run("no PV for the volume", func(t *testing.T) { - c := fake.NewClientBuilder().WithScheme(utilsScheme(t)).Build() - err := StartMigration(context.Background(), c, utilsVolume, "target-node", - "vmig-1", "sb", nil, nil) - if err == nil { - t.Fatalf("expected an error when no PV backs the volume") - } - if !strings.Contains(err.Error(), "resolve PV") { - t.Errorf("error = %q, want it to say the PV could not be resolved", err) - } - }) - - t.Run("a CR of that name already exists", func(t *testing.T) { - existing := &simplyblockv1alpha1.VolumeMigration{ - ObjectMeta: metav1.ObjectMeta{Name: "vmig-1", Namespace: "sb"}, - } - c := fake.NewClientBuilder().WithScheme(utilsScheme(t)). - WithObjects(pvWithHandle("pv-1", utilsCluster+":"+utilsPool+":"+utilsVolume), existing).Build() - - if err := StartMigration(context.Background(), c, utilsVolume, "target-node", - "vmig-1", "sb", nil, nil); err == nil { - t.Errorf("expected the duplicate create to be reported, not silently ignored") - } - }) -} - func TestPollMigration(t *testing.T) { const nqn = "nqn.test:vol-1" From b423551efe49a1b61bdb1be6946204799749f90c Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 11:22:06 +0200 Subject: [PATCH 040/206] docs(design): PersistentVolumeOps records what was built and what was not MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The design was written before the kind existed, and building it changed four things about the specification rather than about the code. The handle is parsed with lvol.ParseHandle rather than VolumeHandle.Split, because the pool segment is not always a UUID: volumes provisioned before the v2 API migration encode the pool's name there, and a PersistentVolume outlives every driver upgrade. Requiring a UUID would refuse to move volumes that are otherwise movable, and nothing in a migration needs the pool typed. Verifying clears husks and does not release paths. The design said "no path is left connected," which reads as the release — and by the time that step runs, the copy has finished and those paths are the data path. Releasing belongs to the abort, the failure, and the deletion, which is exactly the set of endings where the migration did not cut over. The `continue` is at-most-once, recorded before the call rather than after. Reading the phase first is not enough on its own: the control plane does not leave pre_created the instant the call returns, so a step re-entering in that window would call again — which is the shape that has lost writes here. And §12 is now the record of what is specified and not built: retention, the creator's finalizer cascade, the volume stream, the event mirror, and two of the seven metrics. Each row says what stands in for it, because a design whose status says Partially Implemented and does not say which part is a design nobody can act on. Appendix A gains what the type gained on the way: the connect parameters on a recorded path, the node and the pass on a validation Job, and the target node beside the source. Co-Authored-By: Claude Fable 5 --- .../design-persistentvolumeops.md | 155 +++++++++++++++--- 1 file changed, 133 insertions(+), 22 deletions(-) diff --git a/operator/docs/designs/crd-redesign/design-persistentvolumeops.md b/operator/docs/designs/crd-redesign/design-persistentvolumeops.md index 343d6bd7e..37806cf14 100644 --- a/operator/docs/designs/crd-redesign/design-persistentvolumeops.md +++ b/operator/docs/designs/crd-redesign/design-persistentvolumeops.md @@ -1,13 +1,20 @@ # Design Document: PersistentVolumeOps -**Status:** Draft +**Status:** Partially Implemented **Author:** Christoph Engelbert (noctarius) -**Date:** 2026-08-30 (last updated 2026-09-08) +**Date:** 2026-08-30 (last updated 2026-09-17) **Test Plan:** [`tests/test-plan-persistentvolumeops.md`](../../tests/test-plan-persistentvolumeops.md) -This document specifies the target model. `VolumeMigration` is registered and is -absorbed into this kind ([`design-crd-model.md`](design-crd-model.md) §9.1), and -§10 is the single record of what the rework changes against it. +The kind is built at `storage.simplyblock.io/v1alpha2`, with its reconciler, its +admission guard, and the volume's lock. §10 is the single record of what the +rework changes against `VolumeMigration`, and §12 records what is specified here +and not yet built. + +`VolumeMigration` stays registered and is not absorbed. The two kinds run side +by side for at least one release, and which one raises a move is a deployment +setting rather than a conversion: a rename and a scope change make a new CRD +rather than a new version, so a migration in flight cannot be carried across and +is drained on the kind it started on (§10). --- @@ -150,11 +157,18 @@ specifies. derived from.** A `StorageNodeOps` knows its cluster from its node. This one reads `spec.csi.volumeHandle` off the `PersistentVolume`: that is the CSI volume ID, it is `::`, and it is stamped on the volume at -provisioning and immutable for the volume's life. `atlas-lib`'s -`lvol.VolumeHandle.Split()` is the parser, and it is the whole resolution. No -`StorageClass` is consulted, so a class edited, replaced, or deleted out of band -changes nothing about an existing volume's addressability, and the cluster, pool, -and volume UUIDs the later steps need all come out of one immutable field. +provisioning and immutable for the volume's life. No `StorageClass` is consulted, +so a class edited, replaced, or deleted out of band changes nothing about an +existing volume's addressability, and everything the later steps need comes out of +one immutable field. + +**The parser is `lvol.ParseHandle` rather than `lvol.VolumeHandle.Split`,** and +the difference is the pool segment. `Split` requires three UUIDs; volumes +provisioned before the v2 API migration encode the pool's *name* there, and a +`PersistentVolume` outlives every driver upgrade, so a cluster holds a mixture +indefinitely. Nothing in a migration needs the pool typed — a migration is +addressed by cluster and subsystem — so requiring a UUID would refuse to move +volumes that are otherwise perfectly movable. **What can go wrong there is a volume that is not one of this driver's.** A `PersistentVolume` with no `spec.csi`, one provisioned by a different driver, or one @@ -171,7 +185,7 @@ operation is about the volume. ## 4. PersistentVolumeOps: API -Declared in `operator/api/v1alpha1/persistentvolumeops_types.go`, short name +Declared in `operator/api/v1alpha2/persistentvolumeops_types.go`, short name `pvops`, and reconciled by `PersistentVolumeOpsReconciler` in `operator/internal/controllers/volume/persistentvolumeops_controller.go`. The package is the volume band's, which it shares with the auto-rebalancer that creates @@ -355,11 +369,18 @@ Migrate Validating ──► Migrating ──► Verifying ``` -| Step | Side effect on entry | Complete when | -|--------------|------------------------------------------------------------------|---------------------------------------------| -| `Validating` | `POST` the migration, then start a Job per new NVMe-oF path | Every validation Job succeeded | -| `Migrating` | Continue the migration, which is what starts the data copy | The control plane reports the copy finished | -| `Verifying` | Delete the validation Jobs and confirm no path is left connected | No validation Job and no stale path remain | +| Step | Side effect on entry | Complete when | +|--------------|----------------------------------------------------------------|-------------------------------------------------| +| `Validating` | `POST` the migration, then start a Job per consuming node | Every validation Job succeeded | +| `Migrating` | Continue the migration, which is what starts the data copy | The control plane reports the copy finished | +| `Verifying` | Delete the validation Jobs and clear the husks the checks left | No validation Job and no dead controller remain | + +**A Job runs per consuming node rather than per path.** The paths are what the +Job connects; what makes a Job necessary is the *host*, since a path is +established on the node that will be served over it and a subsystem's members +may be consumed on several nodes at once. Every node consuming any volume of the +migrated subsystem gets one, because at cutover every member moves together and a +node that was not checked loses its volume. **`Verifying` is new and it exists because of a defect that reached production.** A migration's validation Jobs connect NVMe-oF paths to check the target is @@ -368,6 +389,17 @@ the data path, and blocked every later migration on that volume. Making the cleanup a declared step rather than a deferred call means a crash between the copy finishing and the cleanup restarts into `Verifying` rather than into nothing. +**What `Verifying` clears is not what an abandoned migration releases, and the +distinction is the whole safety of it.** By the time this step runs the copy has +finished and the target is where the volume is served from, so the paths the +migration published *are* the data path: releasing them here is the outage the +cleanup exists to avoid, and `atlas-lib`'s release says so in its own contract. +What it does clear is a controller carrying no namespace at all, which is the +state a path lost mid-check settles into and which blocks the subsystem's next +migration just as surely as a live leak would. Releasing the paths belongs to the +abort, the failure, and the deletion — which is exactly the set of endings where +the migration did not cut over, and that is the precondition the release has. + **An abort is expressible from `Validating` and `Migrating` and not from `Verifying`.** Before the copy finishes there is a backend migration to cancel. After it, the volume has already moved and there is nothing to undo. The graph @@ -543,6 +575,17 @@ of the call that started it, which is the same rule [`design-crd-model.md`](design-crd-model.md) §7.7 states for every step in the group and the reason it is stated. +**The `continue` is also at-most-once, which is the other half of that.** Reading +the migration's phase before calling is not enough on its own: the control plane +does not leave `pre_created` the instant the call returns, so a step that +re-entered while the phase still said `pre_created` would call again. So +`status.migration.continuedAt` is written *before* the call rather than after it, +and a recorded continue is never issued a second time. The order is deliberate: a +crash between the write and the call leaves the copy unstarted and the step to +time out, which is a stall somebody can see, while the other order leaves a +transfer to be repeated, which is the failure nobody sees until the data is +wrong. + --- ## 8. Observability @@ -677,6 +720,16 @@ step machine it would restore into did not exist when it started. Letting them finish before the CRDs change is the only handling that does not risk leaving a path connected with nothing tracking it. +**Which is why the two kinds coexist rather than one replacing the other.** +`--legacy-volume-migration` decides which kind the three controllers that move +volumes — the auto-rebalancer, a node drain, and the pinned-volume controller — +raise a move as, and the registered kind's reconciler is registered only when it +is on. An upgrade turns it on for as long as the migrations already raised +against that kind take to finish, and turns it off again. Nothing creates one of +each for a volume, and only the redesigned kind takes the lock of §6, which is +the cost of running both: two kinds cannot exclude one another through a lock +only one of them has. + --- ## 11. Ownership and Retention @@ -761,6 +814,8 @@ An operation is deleted when it is older than `opsRetention` or when fifty objects therefore collapse to three per volume within a reconcile, while a single hand-written migration survives its seven days. +**Neither field exists yet, and nor does the retention pass** (§12). + **Nothing deletes a terminal operation whose creator still exists and still lists it**, which is the ordering that keeps retention from racing a cascade. The creator's finalizer (§11.1) deletes its own fan-out, and retention only ever removes what @@ -768,15 +823,27 @@ nothing is tracking. --- -## 12. Open Questions +## 12. Open Questions, and What Is Not Built Yet -None. Every decision this kind turns on is taken in the sections above, from the -single action and the three candidates declined against it (§4.1) to both retention -bounds and their defaults (§11.2). +No question is open. Every decision this kind turns on is taken in the sections +above, from the single action and the three candidates declined against it (§4.1) +to both retention bounds and their defaults (§11.2). The section is here rather than absent because an absent one cannot be told from an oversight. A question that arrives later belongs in it. +What is specified above and not yet built is a shorter list, and each entry says +what stands in for it: + +| Not built | What happens instead | +|--------------------------------------------------------|---------------------------------------------------------------------------------------------------------------| +| Retention: `opsRetention`, `opsHistoryLimit`, the pass | Terminal operations accumulate. Neither field exists on `StorageCluster`, so there is nothing to read (§11.2) | +| The creator's finalizer cascade | `spec.creatorRef` and the label are written, and a `StorageNodeOps` deletes its fan-out through them (§11.1) | +| The `?watch=true` volume stream | `Migrating` reads the migration on each pass, which §7 already names as the one external dependency | +| The mirror of an event onto the claim | Events land on the operation alone, in `default`, for the reason §8.1 gives | +| `simplyblock_persistentvolume_operation_active_count` | The other six of §8.2 are exported | +| `simplyblock_persistentvolume_stale_paths_total` | `Verifying` clears the husks; nothing counts them | + --- ## Appendix A: `persistentvolumeops_types.go` @@ -924,7 +991,15 @@ type PersistentVolumeOpsSpec struct { CreatorRef *CreatorReference `json:"creatorRef,omitempty"` } -// MigrationConnection is one NVMe-oF path the migration created on the target. +// MigrationConnection is one NVMe-oF path the migration published on the +// target. +// +// The connect parameters travel with the address because the path is connected +// on a consuming host rather than here, and what is recorded has to be the +// connect that will actually be made: the host attaches every path with the +// same controller-loss timeout the CSI driver uses, which is not the hour the +// control plane answers with, and a record of the control plane's answer would +// describe a connect nobody performs. type MigrationConnection struct { // +optional NQN string `json:"nqn,omitempty"` @@ -932,6 +1007,21 @@ type MigrationConnection struct { Address string `json:"address,omitempty"` // +optional Port *int32 `json:"port,omitempty"` + // +optional + Transport string `json:"transport,omitempty"` + + // +optional + NrIOQueues *int32 `json:"nrIOQueues,omitempty"` + // +optional + ReconnectDelaySeconds *int32 `json:"reconnectDelaySeconds,omitempty"` + // CtrlLossTimeoutSeconds and FastIOFailTimeoutSeconds are pointers because + // zero is a choice ("fail I/O immediately") rather than a missing value. + // +optional + CtrlLossTimeoutSeconds *int32 `json:"ctrlLossTimeoutSeconds,omitempty"` + // +optional + FastIOFailTimeoutSeconds *int32 `json:"fastIOFailTimeoutSeconds,omitempty"` + // +optional + KeepAliveTimeoutSeconds *int32 `json:"keepAliveTimeoutSeconds,omitempty"` } // ValidationJob is one Job started to check a path is reachable. It is tracked @@ -944,8 +1034,17 @@ type ValidationJob struct { Namespace string `json:"namespace"` // +kubebuilder:validation:Required Name string `json:"name"` + + // Node is the worker the Job is pinned to, which is a node consuming one of + // the migrated subsystem's volumes. // +optional - NQN string `json:"nqn,omitempty"` + Node string `json:"node,omitempty"` + + // Succeeded records a node whose paths were checked and found ready, so a + // restart does not run the check again on a node that already passed and + // whose Job its own TTL may already have reaped. + // +optional + Succeeded bool `json:"succeeded,omitempty"` } // MigrationStatus is everything about the migration rather than about the @@ -977,6 +1076,18 @@ type MigrationStatus struct { // +optional SourceNodeUUID string `json:"sourceNodeUUID,omitempty"` + // TargetNodeUUID is the backend identifier resolved from + // spec.migrate.targetNodeRef, recorded so the later steps address the + // target without resolving the node object again. + // +optional + TargetNodeUUID string `json:"targetNodeUUID,omitempty"` + + // ContinuedAt is when this operation asked the control plane to start the + // copy, written before the call rather than after it, which is what makes + // the request at-most-once (§7). + // +optional + ContinuedAt *metav1.Time `json:"continuedAt,omitempty"` + // MemberCount is how many volumes (namespaces) the migrated NVMe-oF subsystem // holds, as the control plane reports it. A migration is addressed by the // subsystem rather than by one volume, so more than one member means the From 035d70cfc8f0a33f10cf3adf191953f190874e2e Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Thu, 17 Sep 2026 10:45:36 +0100 Subject: [PATCH 041/206] =?UTF-8?q?docs:=20rework=20csi-addons=20design=20?= =?UTF-8?q?=C2=A73=20architecture=20=E2=80=94=20uniform=20promote/demote?= =?UTF-8?q?=20endpoints,=20operator=20preflight=20and=20one-owner=20role,?= =?UTF-8?q?=20P0-6=20test-failover=20read,=20and=20metrics=20export?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- .../designs/design-csi-addons-replication.md | 73 +++++++++++-------- 1 file changed, 42 insertions(+), 31 deletions(-) diff --git a/operator/docs/designs/design-csi-addons-replication.md b/operator/docs/designs/design-csi-addons-replication.md index 8535880b6..62348a08a 100644 --- a/operator/docs/designs/design-csi-addons-replication.md +++ b/operator/docs/designs/design-csi-addons-replication.md @@ -115,37 +115,44 @@ A reader who stops here has the model: the engine is unchanged, the csi-addons s ## 3. Architecture Overview ``` -┌───────────────────────────────────────────────────────────────────────────┐ -│ Kubernetes (each cluster) │ -│ │ -│ VolumeReplicationClass VolumeReplication (one per protected PVC) │ -│ (names a ReplicationPolicy) spec.replicationState: primary|secondary │ -│ │ │ │ -│ ▼ ▼ │ -│ ┌─────────────────────────────────────────────────────────────────────┐ │ -│ │ kubernetes-csi-addons controller-manager (chart-deployed) │ │ -│ │ reconciles VolumeReplication → Replication gRPC on the driver │ │ -│ └───────────────────────────────┬─────────────────────────────────────┘ │ -│ │ via CSIAddonsNode registration │ -│ ┌───────────────────────────────▼─────────────────────────────────────┐ │ -│ │ CSI controller StatefulSet (SimplyblockDriver reconciler) │ │ -│ │ … csi-provisioner │ csi-snapshotter │ … │ csi-addons sidecar │ │ │ -│ │ controller plugin socket │ │ -│ │ plugin serves: CSI Identity/Controller/GroupController │ │ -│ │ + csi-addons Identity + Replication (this design) │ │ -│ └───────────────────────────────┬─────────────────────────────────────┘ │ -└──────────────────────────────────┼─────────────────────────────────────────┘ - │ HTTP, resolved per volume handle -┌──────────────────────────────────▼─────────────────────────────────────────┐ -│ simplyblock control plane │ -│ PUT .../volumes/{v} {replication_policy_id} (attach/ │ -│ detach) │ -│ GET .../volumes/{v}/replication/status (P0-1, steady state) │ -│ POST .../volumes/{v}/replication/failover (promote, planned or │ -│ forced) │ -│ POST .../volumes/{v}/replication/demote (P0-3) │ -│ POST .../volumes/{v}/replication/failback (resync) │ -│ GET .../replication/relationships/{lvol} (cutover records) │ +┌──────────────────────────────────────────────────────────────────────────────┐ +│ Kubernetes (each cluster) │ +│ │ +│ VolumeReplicationClass VolumeReplication (one per protected PVC) │ +│ (names a ReplicationPolicy) spec.replicationState: primary|secondary │ +│ │ │ │ +│ ▼ ▼ │ +│ ┌──────────────────────────────────────────────────────────────────┐ │ +│ │ kubernetes-csi-addons controller-manager (chart-deployed): │ │ +│ │ reconciles VolumeReplication → Replication gRPC on the driver │ │ +│ └──────────────────────────────┬───────────────────────────────────┘ │ +│ │ via CSIAddonsNode registration │ +│ ┌──────────────────────────────▼───────────────────────────────────┐ │ +│ │ CSI controller StatefulSet (SimplyblockDriver reconciler): │ │ +│ │ … csi-provisioner, csi-snapshotter, …, csi-addons sidecar, │ │ +│ │ and the controller plugin, all on one socket. The plugin │ │ +│ │ serves CSI Identity/Controller/GroupController, plus │ │ +│ │ csi-addons Identity and Replication (this design) │ │ +│ └──────────────────────────────┬───────────────────────────────────┘ │ +│ │ +│ operator: peerClasses preflight, events on the ReplicationPair (§7.2); │ +│ PVCAnnotationWatcher skips csi-addons-managed volumes (§8); │ +│ ReplicationPair/Policy author the backend target and policy │ +└─────────────────────────────────┬────────────────────────────────────────────┘ + │ HTTP, resolved per volume handle +┌─────────────────────────────────▼────────────────────────────────────────────┐ +│ simplyblock control plane │ +│ PUT .../volumes/{v} {replication_policy_id} (attach, │ +│ detach) │ +│ GET .../volumes/{v}/replication/status (P0-1, steady state) │ +│ POST .../volumes/{v}/replication/failover (promote: planned gate │ +│ or forced) │ +│ POST .../volumes/{v}/replication/demote (P0-3: converge, │ +│ quiesce, flush, fence) │ +│ POST .../volumes/{v}/replication/failback (resync) │ +│ GET .../replication/relationships/{lvol} (cutover records) │ +│ GET .../relationships/{lvol}/latest-snapshot (P0-6, test failover) │ +│ exports simplyblock_replication_* metrics: lag, backlog, RPO (§11) │ └──────────────────────────────────────────────────────────────────────────────┘ ``` @@ -153,6 +160,10 @@ A reader who stops here has the model: the engine is unchanged, the csi-addons s **The adapter holds no state.** csi-addons RPCs are stateless and idempotent by contract. Every answer the driver gives is derived on the spot from the backend status read and the relationship record. There is no driver-side cache, no persisted step, and no state machine. A verb whose backend work outlives the call (a demote converging a busy peer) reports `Completed=False` until the backend reflects the target state, and Ramen's re-drive is the retry loop. +**The operator's part is small and off the data path.** The kinds that author the backend state (`ReplicationPair` for the target, `ReplicationPolicy` for cadence and retention) keep working unchanged, and the classes name what they author. On top of them the operator runs the peerClasses preflight (§7.2), surfacing convention drift as events on the pair, and teaches the `PVCAnnotationWatcher` the one-owner rule (§8) so the legacy annotation path and a `VolumeReplication` never fight over one volume. Everything imperative it used to own (`ReplicationOps`, the commit cutover, the cutover-proceed handshake) is off this contract and confined to the legacy path. + +**The drill rides the same surface.** Test failover (§14) adds no machinery to this picture: the bubble mode clones the latest replicated snapshot (the P0-6 read plus the ordinary CSI clone path), and the test-cluster mode composes clone, a `drtest-` policy toward the test cluster, and the real failover. The invariant audit that proves a drill disturbed nothing reads the same P0-1 status endpoint the conditions come from. + --- ## 4. The csi-addons Machinery From 8ab10af89920e385a4d5c81a4d458262739efc42 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 12:08:49 +0200 Subject: [PATCH 042/206] refactor(volume): the VolumeMigration controller joins the band it belongs to MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Both kinds that move a volume now reconcile from one package. The registered one is on its way out, and moving it first is what makes that a deletion rather than an excavation: its files, its tests, and its scaffolding sit together, named for the kind, and go together. Three things the two kinds had each written separately are now written once, which is what the move made visible rather than what it set out to do. The Job-name fragment a UUID renders into was duplicated between the migration controller and the latency controller, which is two packages producing names for Jobs in one cluster. The validation Job's deadline was declared twice with one comment between them. And the counter that tells the rebalancer a data realignment is owed was a method on the registered kind's reconciler, which the rebalancer's own tests reached across a package boundary for. That last one was a defect rather than an untidiness. The redesigned kind is the default now and never incremented the counter, so status.volumeMoveGeneration stayed at zero, the rebalancer's periodic loop never concluded a realignment was owed, and nothing anywhere said so: the counter is per cluster rather than per move, so an omission leaves no object in a wrong state to notice. Two tests cover it — a move that landed is counted, and an aborted one is not — and the first was red before the fix. Co-Authored-By: Claude Fable 5 --- operator/cmd/main.go | 2 +- .../internal/controller/controller_helpers.go | 22 ---- .../storagenode_latency_controller.go | 4 +- .../internal/controller/test_helpers_test.go | 19 +++ .../volumerebalancer_realignment_test.go | 17 ++- operator/internal/controllers/volume/jobs.go | 8 +- .../volume/persistentvolumeops_controller.go | 29 +++++ .../persistentvolumeops_controller_test.go | 56 +++++++++ .../volume}/volumemigration_controller.go | 109 +++++------------- .../volumemigration_controller_unit_test.go | 23 +--- .../volumemigration_helpers_shared_test.go | 98 ++++++++++++++++ .../volume}/volumemigration_helpers_test.go | 8 +- .../volumemigration_migration_paths_test.go | 2 +- .../volumemigration_realignment_test.go | 21 +++- operator/internal/volumemigration/job.go | 65 +++++++++++ 15 files changed, 340 insertions(+), 143 deletions(-) rename operator/internal/{controller => controllers/volume}/volumemigration_controller.go (94%) rename operator/internal/{controller => controllers/volume}/volumemigration_controller_unit_test.go (98%) create mode 100644 operator/internal/controllers/volume/volumemigration_helpers_shared_test.go rename operator/internal/{controller => controllers/volume}/volumemigration_helpers_test.go (99%) rename operator/internal/{controller => controllers/volume}/volumemigration_migration_paths_test.go (99%) rename operator/internal/{controller => controllers/volume}/volumemigration_realignment_test.go (88%) diff --git a/operator/cmd/main.go b/operator/cmd/main.go index 4f406fa77..d754b08ef 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -655,7 +655,7 @@ func main() { // VolumeMigration somebody applied by hand against an operator that is // otherwise driving the redesigned kind. if legacyVolumeMigration { - if err := (&controller.VolumeMigrationReconciler{ + if err := (&volumecontrollers.VolumeMigrationReconciler{ Client: mgr.GetClient(), Scheme: mgr.GetScheme(), Recorder: mgr.GetEventRecorder("volumemigration-controller"), diff --git a/operator/internal/controller/controller_helpers.go b/operator/internal/controller/controller_helpers.go index b709bbdea..95d9f9997 100644 --- a/operator/internal/controller/controller_helpers.go +++ b/operator/internal/controller/controller_helpers.go @@ -6,34 +6,12 @@ import ( "sort" "strings" - simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" vmigration "github.com/simplyblock/simplyblock-operator/internal/volumemigration" "github.com/simplyblock/simplyblock-operator/internal/webapi" corev1 "k8s.io/api/core/v1" "sigs.k8s.io/controller-runtime/pkg/client" ) -// requireStorageCluster returns an error when no StorageCluster in namespace -// reports clusterUUID. It is what stops the migration controller starting work -// against a cluster Kubernetes does not account for. -// -// It used to also refuse when volumeMigrationSettings.enabled was false. That -// field is gone (design-storagecluster.md §12): migration cannot be turned off, -// because a drain, a rebalance, and a device replacement are all performed by -// moving volumes, so a cluster that refused to move one could do none of them. -func requireStorageCluster(ctx context.Context, c client.Client, namespace, clusterUUID string) error { - var clusters simplyblockv1alpha2.StorageClusterList - if err := c.List(ctx, &clusters, client.InNamespace(namespace)); err != nil { - return fmt.Errorf("list StorageClusters: %w", err) - } - for _, cr := range clusters.Items { - if cr.Status.UUID == clusterUUID { - return nil - } - } - return fmt.Errorf("no StorageCluster found for cluster UUID %q", clusterUUID) -} - // findConsumerNode returns the Kubernetes hostname of the first Running pod // that mounts a PVC backed by volumeID (the CSI volume UUID encoded in the // PersistentVolume's volumeHandle). Returns "" when no active consumer exists. diff --git a/operator/internal/controller/storagenode_latency_controller.go b/operator/internal/controller/storagenode_latency_controller.go index 07b00ec67..c976aabef 100644 --- a/operator/internal/controller/storagenode_latency_controller.go +++ b/operator/internal/controller/storagenode_latency_controller.go @@ -271,7 +271,7 @@ func (r *StorageNodeLatencyReconciler) reconcileBaselineJob( conn benchmarkConnInfo, image string, ) (*autoplacement.LatencyResult, bool, error) { - jobName := baselineJobNamePrefix + safeNodeID(node.Status.UUID) + jobName := baselineJobNamePrefix + volumemigration.JobNameID(node.Status.UUID) job := &batchv1.Job{} err := r.Get(ctx, types.NamespacedName{Namespace: snode.Namespace, Name: jobName}, job) @@ -325,7 +325,7 @@ func (r *StorageNodeLatencyReconciler) createBaselineJob( return r.Create(ctx, &batchv1.Job{ ObjectMeta: metav1.ObjectMeta{ - Name: baselineJobNamePrefix + safeNodeID(node.Status.UUID), + Name: baselineJobNamePrefix + volumemigration.JobNameID(node.Status.UUID), Namespace: snode.Namespace, Labels: map[string]string{ baselineJobLabelKey: "true", diff --git a/operator/internal/controller/test_helpers_test.go b/operator/internal/controller/test_helpers_test.go index 1025dd935..b154b0bc7 100644 --- a/operator/internal/controller/test_helpers_test.go +++ b/operator/internal/controller/test_helpers_test.go @@ -1,6 +1,8 @@ package controller import ( + "net/http" + "net/http/httptest" "testing" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" @@ -79,3 +81,20 @@ func testCluster(namespace, clusterName, uuid string) *simplyblockv1alpha2.Stora }, } } + +// testClusterUUID is the backend cluster every test in this package that needs +// one reports. It was the volume migration tests' constant and stayed when they +// left, because the replication tests read it too. +const testClusterUUID = "cluster-uuid" + +// unreachableAPI is a base URL that always fails to connect; use it for tests +// that must never reach the storage API. +const unreachableAPI = "http://127.0.0.1:1" + +// newAPIServer starts an httptest server that is closed at test end. +func newAPIServer(t *testing.T, h http.HandlerFunc) *httptest.Server { + t.Helper() + srv := httptest.NewServer(h) + t.Cleanup(srv.Close) + return srv +} diff --git a/operator/internal/controller/volumerebalancer_realignment_test.go b/operator/internal/controller/volumerebalancer_realignment_test.go index 6d2f5faac..b4ab46a67 100644 --- a/operator/internal/controller/volumerebalancer_realignment_test.go +++ b/operator/internal/controller/volumerebalancer_realignment_test.go @@ -18,6 +18,7 @@ import ( "github.com/simplyblock/atlas/ptr" simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/volumemigration" "github.com/simplyblock/simplyblock-operator/internal/webapi" ) @@ -584,9 +585,12 @@ func TestReconcileDataRealignment_RecordsOnlyTheGenerationItCovered(t *testing.T t.Fatalf("realignedGeneration = %d, want 3 (the value read before the call)", got) } - // mig-67 finishes ten seconds later. - vmr := &VolumeMigrationReconciler{Client: f.cl, Scheme: f.r.Scheme, Recorder: events.NewFakeRecorder(8)} - vmr.markClusterVolumeMoved(context.Background(), realignNamespace, realignClusterUUID) + // mig-67 finishes ten seconds later. The counter is written by whichever + // kind carried the move, through the one function both of them call. + if _, err := volumemigration.RecordVolumeMoved( + context.Background(), f.cl, realignNamespace, realignClusterUUID); err != nil { + t.Fatalf("record the move: %v", err) + } cr = f.getCluster(t) gen := ptr.Int64FromOrZero(cr.Status.VolumeMoveGeneration) @@ -666,8 +670,6 @@ func TestReconcileDataRealignment_ReplayOfStackedRealignment(t *testing.T) { // mig-67 is moving; one earlier move is already owed. f := newRealignFixtureWith(t, realignTestCluster(1, 0, nil, false, nil), movingMigration("mig-67", simplyblockv1alpha1.VolumeMigrationPhaseRunning)) - vmr := &VolumeMigrationReconciler{Client: f.cl, Scheme: f.r.Scheme, Recorder: events.NewFakeRecorder(8)} - // 02:13:31 — due, but a volume is moving: deferred instead of sent. f.r.reconcileDataRealignment(ctx, f.getCluster(t), realignClusterUUID) if n := atomic.LoadInt32(f.calls); n != 0 { @@ -675,7 +677,10 @@ func TestReconcileDataRealignment_ReplayOfStackedRealignment(t *testing.T) { } // 02:13:41 — mig-67 completes: counter goes to 2, migration reaches a terminal phase. - vmr.markClusterVolumeMoved(ctx, realignNamespace, realignClusterUUID) + if _, err := volumemigration.RecordVolumeMoved( + ctx, f.cl, realignNamespace, realignClusterUUID); err != nil { + t.Fatalf("record the move: %v", err) + } done := &simplyblockv1alpha1.VolumeMigration{} if err := f.cl.Get(ctx, types.NamespacedName{Namespace: realignNamespace, Name: "mig-67"}, done); err != nil { t.Fatalf("get mig-67: %v", err) diff --git a/operator/internal/controllers/volume/jobs.go b/operator/internal/controllers/volume/jobs.go index b297cc9f3..4385c8716 100644 --- a/operator/internal/controllers/volume/jobs.go +++ b/operator/internal/controllers/volume/jobs.go @@ -24,6 +24,7 @@ import ( "regexp" "sort" "strings" + "time" batchv1 "k8s.io/api/batch/v1" corev1 "k8s.io/api/core/v1" @@ -49,8 +50,9 @@ const ( // validationJobDeadline caps a validation Job's whole life, scheduling and // image pull included. Its purpose is to turn a Job that can never finish — // an unschedulable pod, a node that is not ready — into a failure rather - // than a step parked forever. The check itself needs seconds. - validationJobDeadline = 180 + // than a step parked forever. The check itself needs seconds: three + // connect-and-list attempts, two seconds apart. + validationJobDeadline = 180 * time.Second // jobTTL is how long a finished Job is kept. Long enough to read, short // enough that a drain's worth of them does not accumulate. @@ -531,7 +533,7 @@ func (r *PersistentVolumeOpsReconciler) modeJob( }, BackoffLimit: backoffLimit, TTL: jobTTL, - Deadline: validationJobDeadline, + Deadline: int64(validationJobDeadline.Seconds()), }) } diff --git a/operator/internal/controllers/volume/persistentvolumeops_controller.go b/operator/internal/controllers/volume/persistentvolumeops_controller.go index 7e3b9dbb5..adbaf80fe 100644 --- a/operator/internal/controllers/volume/persistentvolumeops_controller.go +++ b/operator/internal/controllers/volume/persistentvolumeops_controller.go @@ -52,6 +52,7 @@ import ( simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" "github.com/simplyblock/simplyblock-operator/internal/controllers/driver" + vmigration "github.com/simplyblock/simplyblock-operator/internal/volumemigration" ) const ( @@ -301,6 +302,7 @@ func (r *PersistentVolumeOpsReconciler) advance( r.observeStep(ops, subject, current) if machine.IsTerminal() { + r.countTowardRealignment(ctx, subject) r.event(ops, corev1.EventTypeNormal, ReasonOperationSucceeded, "Volume %s was moved to node %s", ops.Spec.PersistentVolumeName, subject.targetNodeName()) return r.finish(ctx, ops, simplyblockv1alpha2.PersistentVolumeOpsPhaseSucceeded, @@ -495,6 +497,33 @@ func (r *PersistentVolumeOpsReconciler) finish( return ctrl.Result{}, r.releaseLock(ctx, ops) } +// countTowardRealignment records that one more volume has moved, which is what +// the rebalancer's periodic loop reads to decide that a control-plane data +// realignment is owed. +// +// It is counted here rather than on every terminal phase, because only a move +// that landed leaves the cluster's data laid out against the placement it had +// before. An aborted or failed operation left the volume where it was. +// +// Best effort: a realignment that is late is not a realignment that is lost, +// since the next move to finish increments again and a realignment is +// idempotent. Failing the operation over it would be reporting a move that +// worked as one that did not. +func (r *PersistentVolumeOpsReconciler) countTowardRealignment( + ctx context.Context, subject *subject, +) { + name, err := vmigration.RecordVolumeMoved( + ctx, r.Client, subject.namespace(), subject.clusterUUID) + switch { + case err != nil: + logf.FromContext(ctx).Error(err, "the volume move could not be counted for the realignment", + "cluster", subject.clusterUUID) + case name == "": + logf.FromContext(ctx).Info("no StorageCluster reports this volume's cluster, "+ + "so the move was not counted for the realignment", "cluster", subject.clusterUUID) + } +} + // hold reports an operation that is admitted, holds nothing, and is waiting. // Pending is both where an operation starts and where it waits, and // status.deferredSince is what says since when — in status rather than in diff --git a/operator/internal/controllers/volume/persistentvolumeops_controller_test.go b/operator/internal/controllers/volume/persistentvolumeops_controller_test.go index d9911f6c4..f5770aea5 100644 --- a/operator/internal/controllers/volume/persistentvolumeops_controller_test.go +++ b/operator/internal/controllers/volume/persistentvolumeops_controller_test.go @@ -25,6 +25,7 @@ import ( "github.com/simplyblock/atlas/controlplane" "github.com/simplyblock/atlas/lvol" + "github.com/simplyblock/atlas/ptr" "github.com/simplyblock/atlas/statemachine" simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" @@ -588,3 +589,58 @@ func runningPodOn(node, namespace, claim string) *corev1.Pod { Status: corev1.PodStatus{Phase: corev1.PodRunning}, } } + +// TestAMoveThatLandedIsCountedForTheRealignment. A volume that moved leaves the +// cluster's data laid out against the placement it had before, and the +// rebalancer's periodic loop asks the control plane to realign once enough +// moves have accumulated. It counts them off status.volumeMoveGeneration, so a +// move that did not increment it is a realignment that is never asked for — and +// the counter is per cluster rather than per move, so nothing else notices the +// omission. +func TestAMoveThatLandedIsCountedForTheRealignment(t *testing.T) { + api := idleSubsystem() + r := testReconciler(t, api, testWorld()...) + + api.read = controlplane.Migration{ID: testMigrationID, Phase: "done", Status: migrationStatusDone} + for range 8 { + runPass(t, r) + } + if got := operationFrom(t, r).Status.Phase; got != simplyblockv1alpha2.PersistentVolumeOpsPhaseSucceeded { + t.Fatalf("phase = %q, want Succeeded", got) + } + + var cluster simplyblockv1alpha2.StorageCluster + if err := r.Get(context.Background(), + types.NamespacedName{Namespace: testNamespace, Name: testClusterCR}, &cluster); err != nil { + t.Fatal(err) + } + if got := ptr.Int64FromOrZero(cluster.Status.VolumeMoveGeneration); got != 1 { + t.Errorf("volumeMoveGeneration = %d, want 1: the move was not counted", got) + } +} + +// TestAMoveThatDidNotLandIsNotCounted. The counter says how much realignment is +// owed, and an aborted move owes none: the volume never left. +func TestAMoveThatDidNotLandIsNotCounted(t *testing.T) { + api := idleSubsystem() + r := testReconciler(t, api, testWorld()...) + + runPass(t, r) + runPass(t, r) + + ops := operationFrom(t, r) + ops.Spec.Abort = true + if err := r.Update(context.Background(), ops); err != nil { + t.Fatal(err) + } + runPass(t, r) + + var cluster simplyblockv1alpha2.StorageCluster + if err := r.Get(context.Background(), + types.NamespacedName{Namespace: testNamespace, Name: testClusterCR}, &cluster); err != nil { + t.Fatal(err) + } + if got := ptr.Int64FromOrZero(cluster.Status.VolumeMoveGeneration); got != 0 { + t.Errorf("volumeMoveGeneration = %d, want 0: nothing moved", got) + } +} diff --git a/operator/internal/controller/volumemigration_controller.go b/operator/internal/controllers/volume/volumemigration_controller.go similarity index 94% rename from operator/internal/controller/volumemigration_controller.go rename to operator/internal/controllers/volume/volumemigration_controller.go index b0ff31160..8529e13f7 100644 --- a/operator/internal/controller/volumemigration_controller.go +++ b/operator/internal/controllers/volume/volumemigration_controller.go @@ -1,21 +1,17 @@ -package controller +package volume import ( "bytes" "context" - "crypto/sha256" - "encoding/hex" "encoding/json" "errors" "fmt" "io" "net" - "regexp" "sort" "strings" "time" - "github.com/simplyblock/atlas/ptr" vmigration "github.com/simplyblock/simplyblock-operator/internal/volumemigration" batchv1 "k8s.io/api/batch/v1" corev1 "k8s.io/api/core/v1" @@ -26,7 +22,6 @@ import ( "k8s.io/client-go/kubernetes" corev1client "k8s.io/client-go/kubernetes/typed/core/v1" "k8s.io/client-go/tools/events" - "k8s.io/client-go/util/retry" ctrl "sigs.k8s.io/controller-runtime" "sigs.k8s.io/controller-runtime/pkg/client" logf "sigs.k8s.io/controller-runtime/pkg/log" @@ -36,6 +31,27 @@ import ( "github.com/simplyblock/simplyblock-operator/internal/webapi" ) +// requireStorageCluster returns an error when no StorageCluster in namespace +// reports clusterUUID. It is what stops a migration starting work against a +// cluster Kubernetes does not account for. +// +// It used to also refuse when volumeMigrationSettings.enabled was false. That +// field is gone (design-storagecluster.md §12): migration cannot be turned off, +// because a drain, a rebalance, and a device replacement are all performed by +// moving volumes, so a cluster that refused to move one could do none of them. +func requireStorageCluster(ctx context.Context, c client.Client, namespace, clusterUUID string) error { + var clusters simplyblockv1alpha2.StorageClusterList + if err := c.List(ctx, &clusters, client.InNamespace(namespace)); err != nil { + return fmt.Errorf("list StorageClusters: %w", err) + } + for _, cr := range clusters.Items { + if cr.Status.UUID == clusterUUID { + return nil + } + } + return fmt.Errorf("no StorageCluster found for cluster UUID %q", clusterUUID) +} + // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=volumemigrations,verbs=get;list;watch;create;update;patch;delete // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=volumemigrations/status,verbs=get;update;patch // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=volumemigrations/finalizers,verbs=update @@ -90,13 +106,6 @@ const ( // consumerWaitRetryDelay is how often to re-check while waiting. consumerWaitRetryDelay = 5 * time.Second - - // validationJobDeadline caps a validation Job's total lifetime, scheduling and - // image pull included (activeDeadlineSeconds). Its purpose is to turn a Job that - // can never finish — an unschedulable pod, a NotReady node — into a failure - // instead of a migration parked in Validating forever. The validation itself - // needs seconds: three connect+list attempts, two seconds apart. - validationJobDeadline = 180 * time.Second ) // errConsumerNotReady indicates that a pod references the volume's PVC but is not @@ -872,7 +881,7 @@ func migrationPathJob( deadline := int64(validationJobDeadline.Seconds()) return vmigration.BuildJob(vmigration.JobParams{ // One Job per node: name carries both migration ID and node. - Name: js.namePrefix + safeNodeID(vm.Status.MigrationUUID) + "-" + nodeSuffix(hostname), + Name: js.namePrefix + vmigration.JobNameID(vm.Status.MigrationUUID) + "-" + nodeSuffix(hostname), Namespace: vm.Namespace, OwnerRef: *metav1.NewControllerRef(vm, simplyblockv1alpha1.GroupVersion.WithKind("VolumeMigration")), Hostname: hostname, @@ -892,25 +901,6 @@ func migrationPathJob( }) } -// nodeSuffix produces a DNS-label-safe, collision-resistant suffix for a node name. -// Node names can be long FQDNs and are not label-safe, so the short host part is kept -// for readability and a hash of the full name for uniqueness. -func nodeSuffix(nodeName string) string { - sum := sha256.Sum256([]byte(nodeName)) - short := strings.ToLower(strings.SplitN(nodeName, ".", 2)[0]) - short = nonLabelChars.ReplaceAllString(short, "") - if len(short) > 16 { - short = short[:16] - } - if short == "" { - return hex.EncodeToString(sum[:6]) - } - return short + "-" + hex.EncodeToString(sum[:4]) -} - -// nonLabelChars matches everything not allowed inside a DNS-1123 label. -var nonLabelChars = regexp.MustCompile(`[^a-z0-9-]`) - // connectionsToValidation converts MigrationConnection status entries to the // vmigration.Connection type consumed by the simplyblock-rebalancer validate-migration mode. func connectionsToValidation(conns []simplyblockv1alpha1.MigrationConnection) []vmigration.Connection { @@ -931,15 +921,6 @@ func connectionsToValidation(conns []simplyblockv1alpha1.MigrationConnection) [] return out } -// safeNodeID produces a DNS-label-safe suffix from a node UUID. -func safeNodeID(nodeUUID string) string { - s := strings.ReplaceAll(nodeUUID, "-", "") - if len(s) > 20 { - s = s[:20] - } - return s -} - // resolveConsumerNodeName finds the Kubernetes node name of the worker node // running a pod that currently has the PVC (resolved from pvName) mounted. // NVMe connections must be established from that node so that the consuming @@ -1204,50 +1185,18 @@ func (r *VolumeMigrationReconciler) reconcileRunning( return ctrl.Result{}, nil } -// markClusterVolumeMoved increments status.volumeMoveGeneration on the StorageCluster -// whose backend UUID matches clusterUUID, recording that one more volume has moved and -// a control-plane data realignment is owed. The counter is read by the -// VolumeRebalancerReconciler's periodic loop, which compares it against -// status.realignedGeneration. +// markClusterVolumeMoved records that one more volume has moved, so the +// rebalancer's periodic loop learns a data realignment is owed. // -// A counter rather than a flag because both quantities matter: how many moves are -// outstanding (so DataRealignment.MinMoves can batch them) and whether a move landed -// after a realignment was already requested (so it is not silently absorbed by a -// realignment that cannot account for it). -// -// Best-effort: any failure is logged but does not fail the migration — the realignment -// is late, not lost, because the next completed migration increments again. The write -// takes the optimistic-locking path and retries on conflict, since two migrations -// completing at once would otherwise read the same value and one increment would -// vanish; with MinMoves batching, a lost increment delays a realignment indefinitely -// rather than by one cycle. +// Best effort: a failure is logged and does not fail the migration, because the +// realignment is late rather than lost — the next completed move increments +// again, and a realignment is idempotent. func (r *VolumeMigrationReconciler) markClusterVolumeMoved( ctx context.Context, namespace, clusterUUID string, ) { log := logf.FromContext(ctx) - if clusterUUID == "" { - return - } - - var name string - err := retry.RetryOnConflict(retry.DefaultRetry, func() error { - var clusters simplyblockv1alpha2.StorageClusterList - if err := r.List(ctx, &clusters, client.InNamespace(namespace)); err != nil { - return err - } - for i := range clusters.Items { - cr := &clusters.Items[i] - if cr.Status.UUID != clusterUUID { - continue - } - name = cr.Name - patch := client.MergeFromWithOptions(cr.DeepCopy(), client.MergeFromWithOptimisticLock{}) - cr.Status.VolumeMoveGeneration = ptr.To(ptr.Int64FromOrZero(cr.Status.VolumeMoveGeneration) + 1) - return r.Status().Patch(ctx, cr, patch) - } - return nil - }) + name, err := vmigration.RecordVolumeMoved(ctx, r.Client, namespace, clusterUUID) switch { case err != nil: log.Error(err, "Cannot record volume move for realignment", "clusterUUID", clusterUUID, "cluster", name) diff --git a/operator/internal/controller/volumemigration_controller_unit_test.go b/operator/internal/controllers/volume/volumemigration_controller_unit_test.go similarity index 98% rename from operator/internal/controller/volumemigration_controller_unit_test.go rename to operator/internal/controllers/volume/volumemigration_controller_unit_test.go index 70158e7d3..66d31fd90 100644 --- a/operator/internal/controller/volumemigration_controller_unit_test.go +++ b/operator/internal/controllers/volume/volumemigration_controller_unit_test.go @@ -1,10 +1,9 @@ -package controller +package volume import ( "context" "fmt" "net/http" - "net/http/httptest" "slices" "strings" "testing" @@ -28,7 +27,7 @@ import ( const ( testVMNamespace = "sb" testVMName = "mig-test" - testPVName = "pv-1" + testVMPVName = "pv-1" testClusterUUID = "cluster-uuid" testPoolUUID = "pool-uuid" testVolumeUUID = "vol-uuid" @@ -44,10 +43,6 @@ const ( testSiblingNode = "sibling-worker" ) -// unreachableAPI is a base URL that always fails to connect; use it for tests -// that must never reach the storage API. -const unreachableAPI = "http://127.0.0.1:1" - // newVMReconciler builds a VolumeMigrationReconciler backed by a fake k8s client // (with VolumeMigration status subresource enabled) and a webapi client pointed // at apiURL. Pass unreachableAPI when the API must not be called. @@ -125,14 +120,6 @@ func serveSubsystemMembers(w http.ResponseWriter, r *http.Request, siblingVolume return true } -// newAPIServer starts an httptest server that is closed at test end. -func newAPIServer(t *testing.T, h http.HandlerFunc) *httptest.Server { - t.Helper() - srv := httptest.NewServer(h) - t.Cleanup(srv.Close) - return srv -} - func vmRequest() ctrl.Request { return ctrl.Request{NamespacedName: types.NamespacedName{Namespace: testVMNamespace, Name: testVMName}} } @@ -150,17 +137,17 @@ func baseVM() *simplyblockv1alpha1.VolumeMigration { return &simplyblockv1alpha1.VolumeMigration{ ObjectMeta: metav1.ObjectMeta{Name: testVMName, Namespace: testVMNamespace}, Spec: simplyblockv1alpha1.VolumeMigrationSpec{ - PVName: testPVName, + PVName: testVMPVName, TargetNodeUUID: "target-node", }, } } -// csiPV returns a CSI-provisioned PV (named testPVName, matching baseVM's PVName) +// csiPV returns a CSI-provisioned PV (named testVMPVName, matching baseVM's PVName) // with the given volume handle. func csiPV(handle string) *corev1.PersistentVolume { return &corev1.PersistentVolume{ - ObjectMeta: metav1.ObjectMeta{Name: testPVName}, + ObjectMeta: metav1.ObjectMeta{Name: testVMPVName}, Spec: corev1.PersistentVolumeSpec{ PersistentVolumeSource: corev1.PersistentVolumeSource{ CSI: &corev1.CSIPersistentVolumeSource{VolumeHandle: handle}, diff --git a/operator/internal/controllers/volume/volumemigration_helpers_shared_test.go b/operator/internal/controllers/volume/volumemigration_helpers_shared_test.go new file mode 100644 index 000000000..e136bdd65 --- /dev/null +++ b/operator/internal/controllers/volume/volumemigration_helpers_shared_test.go @@ -0,0 +1,98 @@ +// The scaffolding the registered kind's tests brought with them when the +// controller moved into this package. +// +// It is a copy of what internal/controller's tests share, and the duplication +// is deliberate rather than overlooked: a test helper cannot be imported across +// packages without being exported into the build, and the six other test files +// that still use the original are reason enough to leave it where it is. Both +// copies go when the registered kind does. +// +// Everything here is the registered kind's. The redesigned kind's own +// scaffolding is in helpers_test.go, and the two are kept apart so that +// deleting one is a file rather than an excavation. + +package volume + +import ( + "net/http" + "net/http/httptest" + "testing" + + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// newTestScheme builds a scheme carrying both simplyblock API versions, plus +// whatever else the caller adds. +// +// Both versions are unconditional because the group has two of them: a kind +// with a renamed property is read as v1alpha2 by its controller and may still +// be written as v1alpha1 by a controller that has not moved yet, and a fake +// client that knows only one of them panics on the other. +func newTestScheme(t *testing.T, addToScheme ...func(*runtime.Scheme) error) *runtime.Scheme { + t.Helper() + + scheme := runtime.NewScheme() + for _, add := range append( + []func(*runtime.Scheme) error{ + simplyblockv1alpha1.AddToScheme, + simplyblockv1alpha2.AddToScheme, + }, + addToScheme..., + ) { + if err := add(scheme); err != nil { + t.Fatalf("failed to add scheme: %v", err) + } + } + + return scheme +} + +func newTestClient( + t *testing.T, + scheme *runtime.Scheme, + statusSubresources []client.Object, + objects ...client.Object, +) client.Client { + t.Helper() + + builder := fake.NewClientBuilder().WithScheme(scheme) + if len(statusSubresources) > 0 { + builder = builder.WithStatusSubresource(statusSubresources...) + } + if len(objects) > 0 { + builder = builder.WithObjects(objects...) + } + + return builder.Build() +} + +// testCluster is the StorageCluster the realignment counter is written on. Only +// the backend UUID varies between the cases: what they are about is whether the +// UUID a move reports matches the one a cluster claims. +func testCluster(uuid string) *simplyblockv1alpha2.StorageCluster { + return &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{ + Name: realignClusterName, + Namespace: realignNamespace, + }, + Status: simplyblockv1alpha2.StorageClusterStatus{UUID: uuid}, + } +} + +// unreachableAPI is a base URL that always fails to connect, for tests that +// must never reach the control plane. +const unreachableAPI = "http://127.0.0.1:1" + +// newAPIServer starts a test server that is closed at test end. +func newAPIServer(t *testing.T, h http.HandlerFunc) *httptest.Server { + t.Helper() + srv := httptest.NewServer(h) + t.Cleanup(srv.Close) + return srv +} diff --git a/operator/internal/controller/volumemigration_helpers_test.go b/operator/internal/controllers/volume/volumemigration_helpers_test.go similarity index 99% rename from operator/internal/controller/volumemigration_helpers_test.go rename to operator/internal/controllers/volume/volumemigration_helpers_test.go index 0cc0f7adf..725e97741 100644 --- a/operator/internal/controller/volumemigration_helpers_test.go +++ b/operator/internal/controllers/volume/volumemigration_helpers_test.go @@ -1,4 +1,4 @@ -package controller +package volume import ( "context" @@ -48,7 +48,7 @@ func TestNodeSuffix(t *testing.T) { if !label.MatchString(suffix) { t.Errorf("nodeSuffix(%q) = %q, which is not a valid DNS-1123 label", node, suffix) } - full := "vmig-validate-" + safeNodeID(testMigrationUUID) + "-" + suffix + full := "vmig-validate-" + vmigration.JobNameID(testMigrationUUID) + "-" + suffix if len(full) > 63 { t.Errorf("job name %q is %d chars, over the 63-char label limit", full, len(full)) } @@ -121,8 +121,8 @@ func TestPVNamesForVolumes(t *testing.T) { } // Non-CSI and malformed handles are skipped rather than fatal — a cluster can hold // PVs from any driver, and one bad handle must not stop the resolution. - if len(got) != 1 || got[0] != testPVName { - t.Errorf("PVs = %v, want just %s", got, testPVName) + if len(got) != 1 || got[0] != testVMPVName { + t.Errorf("PVs = %v, want just %s", got, testVMPVName) } } diff --git a/operator/internal/controller/volumemigration_migration_paths_test.go b/operator/internal/controllers/volume/volumemigration_migration_paths_test.go similarity index 99% rename from operator/internal/controller/volumemigration_migration_paths_test.go rename to operator/internal/controllers/volume/volumemigration_migration_paths_test.go index 158efbf9e..483fad664 100644 --- a/operator/internal/controller/volumemigration_migration_paths_test.go +++ b/operator/internal/controllers/volume/volumemigration_migration_paths_test.go @@ -1,4 +1,4 @@ -package controller +package volume import ( "context" diff --git a/operator/internal/controller/volumemigration_realignment_test.go b/operator/internal/controllers/volume/volumemigration_realignment_test.go similarity index 88% rename from operator/internal/controller/volumemigration_realignment_test.go rename to operator/internal/controllers/volume/volumemigration_realignment_test.go index d6008b8fe..e1f0e0768 100644 --- a/operator/internal/controller/volumemigration_realignment_test.go +++ b/operator/internal/controllers/volume/volumemigration_realignment_test.go @@ -1,4 +1,4 @@ -package controller +package volume import ( "context" @@ -14,6 +14,15 @@ import ( simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" ) +const ( + // The cluster the realignment counter is written on. They were the + // rebalancer's test constants and travel with the tests that read them, + // since the rebalancer stayed behind. + realignNamespace = "sb" + realignClusterName = "cluster-a" + realignClusterUUID = "cluster-uuid-a" +) + // newVMReconcilerForRealign builds a VolumeMigrationReconciler whose fake client has // the StorageCluster status subresource enabled (markClusterVolumeMoved patches // StorageCluster status, not VolumeMigration status). @@ -35,7 +44,7 @@ func getClusterByName(t *testing.T, cl client.Client) *simplyblockv1alpha2.Stora } func TestMarkClusterVolumeMoved_IncrementsGeneration(t *testing.T) { - cr := testCluster(realignNamespace, realignClusterName, realignClusterUUID) + cr := testCluster(realignClusterUUID) r, cl := newVMReconcilerForRealign(t, cr) r.markClusterVolumeMoved(context.Background(), realignNamespace, realignClusterUUID) @@ -48,7 +57,7 @@ func TestMarkClusterVolumeMoved_IncrementsGeneration(t *testing.T) { // Each completed migration must count, or MinMoves batching can never be reached. // The superseded boolean was idempotent by design; the counter must not be. func TestMarkClusterVolumeMoved_EveryMoveCounts(t *testing.T) { - cr := testCluster(realignNamespace, realignClusterName, realignClusterUUID) + cr := testCluster(realignClusterUUID) r, cl := newVMReconcilerForRealign(t, cr) for range 5 { @@ -63,7 +72,7 @@ func TestMarkClusterVolumeMoved_EveryMoveCounts(t *testing.T) { // A move counted while a realignment already covers an earlier generation must push the // counter past realignedGeneration, which is what leaves the next realignment owed. func TestMarkClusterVolumeMoved_CountsPastAlreadyRealigned(t *testing.T) { - cr := testCluster(realignNamespace, realignClusterName, realignClusterUUID) + cr := testCluster(realignClusterUUID) cr.Status.VolumeMoveGeneration = ptr.To(int64(7)) cr.Status.RealignedGeneration = ptr.To(int64(7)) r, cl := newVMReconcilerForRealign(t, cr) @@ -81,7 +90,7 @@ func TestMarkClusterVolumeMoved_CountsPastAlreadyRealigned(t *testing.T) { func TestMarkClusterVolumeMoved_NoMatchingClusterLeavesOthersAlone(t *testing.T) { // A cluster with a *different* UUID must not be counted against. - cr := testCluster(realignNamespace, realignClusterName, "some-other-uuid") + cr := testCluster("some-other-uuid") r, cl := newVMReconcilerForRealign(t, cr) r.markClusterVolumeMoved(context.Background(), realignNamespace, realignClusterUUID) @@ -92,7 +101,7 @@ func TestMarkClusterVolumeMoved_NoMatchingClusterLeavesOthersAlone(t *testing.T) } func TestMarkClusterVolumeMoved_EmptyUUIDIsNoOp(t *testing.T) { - cr := testCluster(realignNamespace, realignClusterName, realignClusterUUID) + cr := testCluster(realignClusterUUID) r, cl := newVMReconcilerForRealign(t, cr) // Must not count against any cluster (and must not panic) when the volume carries diff --git a/operator/internal/volumemigration/job.go b/operator/internal/volumemigration/job.go index 6d8c0765f..f79e39aa7 100644 --- a/operator/internal/volumemigration/job.go +++ b/operator/internal/volumemigration/job.go @@ -13,12 +13,16 @@ package volumemigration import ( "context" "fmt" + "strings" batchv1 "k8s.io/api/batch/v1" corev1 "k8s.io/api/core/v1" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/client-go/util/retry" "sigs.k8s.io/controller-runtime/pkg/client" + "github.com/simplyblock/atlas/ptr" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" ) @@ -108,6 +112,67 @@ func BuildJob(p JobParams) *batchv1.Job { } } +// JobNameID renders a UUID into a DNS-label-safe fragment of a Job's name. +// +// The hyphens go and the result is capped, which is what keeps a Job name +// inside the 63 characters a label allows once a prefix and a node suffix are +// added to it. It is not reversible and does not need to be: what reads a Job +// name back is a human, and what finds a Job again is the record in status. +func JobNameID(uuid string) string { + s := strings.ReplaceAll(uuid, "-", "") + if len(s) > 20 { + s = s[:20] + } + return s +} + +// RecordVolumeMoved increments status.volumeMoveGeneration on the +// StorageCluster reporting clusterUUID, which records that one more volume has +// moved and a control-plane data realignment is owed. +// +// Both kinds of move call it, because the rebalancer's periodic loop reads the +// counter and compares it against status.realignedGeneration, and a move that +// did not count is a realignment that is never asked for. +// +// A counter rather than a flag, because both quantities matter: how many moves +// are outstanding, so a minimum can batch them, and whether a move landed after +// a realignment was already requested, so it is not silently absorbed by one +// that cannot account for it. The write retries on conflict, since two moves +// finishing at once would otherwise read the same value and one increment would +// vanish — and with batching, a lost increment delays a realignment +// indefinitely rather than by one cycle. +// +// Best effort in the caller's hands: a failure here is a realignment that is +// late rather than lost, because the next move to finish increments again. +func RecordVolumeMoved( + ctx context.Context, c client.Client, namespace, clusterUUID string, +) (string, error) { + if clusterUUID == "" { + return "", nil + } + var name string + err := retry.RetryOnConflict(retry.DefaultRetry, func() error { + var clusters simplyblockv1alpha2.StorageClusterList + if err := c.List(ctx, &clusters, client.InNamespace(namespace)); err != nil { + return err + } + for i := range clusters.Items { + cluster := &clusters.Items[i] + if cluster.Status.UUID != clusterUUID { + continue + } + name = cluster.Name + patch := client.MergeFromWithOptions( + cluster.DeepCopy(), client.MergeFromWithOptimisticLock{}) + cluster.Status.VolumeMoveGeneration = ptr.To( + ptr.Int64FromOrZero(cluster.Status.VolumeMoveGeneration) + 1) + return c.Status().Patch(ctx, cluster, patch) + } + return nil + }) + return name, err +} + // JobImage returns the simplyblock-rebalancer image configured on the // StorageCluster reporting clusterUUID, falling back to JobImageDefault when // the cluster pins none. From cf246031c3c758c7db157dcb05da1605271d8073 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 12:16:11 +0200 Subject: [PATCH 043/206] feat(node): the operator holds its own eviction while it arranges a drain MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A maintenance window blocks the storage-node pod's eviction before it takes the backend node down, which is what makes `kubectl drain` wait rather than kill an SPDK process under a running node. That protects the storage pod and nothing else, and on a converged worker the manager runs on the same host. The drain evicts in no particular order. So it can take the manager out before the manager has created the storage node's budget — and then nothing is left that ever will: the storage pod goes next, the process dies under a running backend node, and the control plane sees a node that vanished. That is the exact failure the window exists to prevent, reached by evicting the thing preventing it. So the manager holds a budget over itself, first in ShuttingDown, ahead of the storage node's budget and ahead of any control-plane call. The ordering is the mechanism rather than a detail of it, and the test for it drives the step against an unreachable control plane: the protection must not be contingent on anything that can fail. It is released in Releasing, with the storage pod's, and that is where it has to be. The manager is on the node being drained, so a budget held any longer blocks the drain on the manager instead of on the pod — the same deadlock, one object over. What the window still needs afterward survives the manager being rescheduled, because the step is persisted. The budget selects the label the Deployment already stamps rather than one this code writes, because a label written to protect a pod has to land before the protection exists, and the window it protects is exactly the one where that write may not finish. A budget left behind by a manager that crashed holding one would make its worker undrainable by an object whose owner is gone. The next leader clears it on the way up, and only the leader: a replica doing it would be clearing the leader's. NODE_NAME was already set on the manager pod by the chart and read by nothing. It is read now. Co-Authored-By: Claude Fable 5 --- operator/cmd/main.go | 30 ++ .../controllers/node/hostmaintenance.go | 30 ++ .../internal/controllers/node/selfbudget.go | 146 +++++++++ .../controllers/node/selfbudget_test.go | 293 ++++++++++++++++++ .../internal/controllers/node/workload.go | 8 + 5 files changed, 507 insertions(+) create mode 100644 operator/internal/controllers/node/selfbudget.go create mode 100644 operator/internal/controllers/node/selfbudget_test.go diff --git a/operator/cmd/main.go b/operator/cmd/main.go index d754b08ef..f0eab342f 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -17,6 +17,7 @@ limitations under the License. package main import ( + "context" "crypto/tls" "flag" "fmt" @@ -367,6 +368,23 @@ func main() { Uncached: mgr.GetAPIReader(), TLSEnabled: tlsEnabled, TLSMutualEnabled: tlsMutualEnabled, + // Which node the manager itself runs on, which the chart sets from + // spec.nodeName. It decides one question: whether a maintenance window + // is draining the manager's own host, in which case the manager holds + // its own eviction while it arranges the window. Empty where nothing + // set it, and then it holds nothing. + ManagerNode: os.Getenv("NODE_NAME"), + } + + // A self-budget outlives a manager that crashed holding one, and would then + // make its worker undrainable by an object whose owner no longer exists. + // Clearing it is leader-only, because only the leader ever creates one: a + // replica clearing it on the way up would be clearing the leader's. + if err := mgr.Add(leaderOnly(func(ctx context.Context) error { + return storageNodeWorkload.ClearStaleSelfBudget(ctx, operatorNamespace) + })); err != nil { + setupLog.Error(err, "unable to schedule the stale self-budget cleanup") + os.Exit(1) } cpSubscriptions := cpinformer.NewSubscriptionManager(streamCfg, ctrl.Log.WithName("cpinformer"), cpinformer.LeaderOnly) @@ -965,3 +983,15 @@ func main() { os.Exit(1) } } + +// leaderOnly wraps a one-shot task the manager runs after it wins the leader +// election, and not before. +// +// The distinction matters for anything that cleans up after a previous +// instance: a replica that ran it on the way up would be undoing what the +// current leader is in the middle of. +type leaderOnly func(context.Context) error + +func (f leaderOnly) Start(ctx context.Context) error { return f(ctx) } + +func (leaderOnly) NeedLeaderElection() bool { return true } diff --git a/operator/internal/controllers/node/hostmaintenance.go b/operator/internal/controllers/node/hostmaintenance.go index 0d9e821d3..dd14c6925 100644 --- a/operator/internal/controllers/node/hostmaintenance.go +++ b/operator/internal/controllers/node/hostmaintenance.go @@ -168,12 +168,22 @@ func (r *StorageNodeOpsReconciler) workersInMaintenance( // The budget is created before the shutdown is issued, and that ordering is the // whole mechanism: a drain that reaches the pod before the budget exists evicts it // under a running SPDK process. +// +// The manager's own budget comes before even that, and for the same reason one +// step further back. On a converged worker the manager runs on the host being +// drained, the drain evicts in no particular order, and an eviction that +// arrives here first takes out the only thing that would ever have created the +// storage node's budget (selfbudget.go). func (r *StorageNodeOpsReconciler) maintenanceShutDown( ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, node *simplyblockv1alpha2.StorageNode, clusterID, nodeID string, ) (bool, error) { + if err := r.Workload.ProtectSelf(ctx, node.Namespace, node.Spec.WorkerNode); err != nil { + return false, err + } + err := r.Workload.BlockEviction(ctx, node.Namespace, node.Spec.ClusterRef, node.Spec.WorkerNode) if err != nil { return false, fmt.Errorf("hold the eviction of worker %s: %w", node.Spec.WorkerNode, err) @@ -200,9 +210,21 @@ func (r *StorageNodeOpsReconciler) maintenanceShutDown( // maintenanceRelease relaxes the budget so the eviction the drain is waiting on // can proceed, and completes when the pod has actually gone. +// +// The manager stops holding itself here too, and it has to be here rather than +// at the end: on a converged worker the manager is on the node being drained, +// so a budget held past this point would block the drain on the manager instead +// of on the storage pod, which is the same deadlock one object over. What the +// window still needs from the manager after this — waiting for the host and +// restarting the node — survives the manager being rescheduled elsewhere, +// because the step is persisted. func (r *StorageNodeOpsReconciler) maintenanceRelease( ctx context.Context, node *simplyblockv1alpha2.StorageNode, ) (bool, error) { + if err := r.Workload.ReleaseSelf(ctx, node.Namespace); err != nil { + return false, err + } + err := r.Workload.AllowEviction(ctx, node.Namespace, node.Spec.ClusterRef, node.Spec.WorkerNode) if err != nil { return false, fmt.Errorf("release the eviction of worker %s: %w", node.Spec.WorkerNode, err) @@ -255,6 +277,14 @@ func (r *StorageNodeOpsReconciler) maintenanceRestart( func (r *StorageNodeOpsReconciler) maintenanceCleanup( ctx context.Context, node *simplyblockv1alpha2.StorageNode, ) (bool, error) { + // The manager's own budget is released in Releasing and cleared again here, + // which is not a repetition: a window that failed before reaching Releasing + // comes through this step on its way to a terminal phase, and that is the + // path on which the budget would otherwise be left holding a worker nothing + // is draining any more. + if err := r.Workload.ReleaseSelf(ctx, node.Namespace); err != nil { + return false, err + } err := r.Workload.ClearEvictionBudget(ctx, node.Namespace, node.Spec.ClusterRef, node.Spec.WorkerNode) return err == nil, err diff --git a/operator/internal/controllers/node/selfbudget.go b/operator/internal/controllers/node/selfbudget.go new file mode 100644 index 000000000..e2ffa9373 --- /dev/null +++ b/operator/internal/controllers/node/selfbudget.go @@ -0,0 +1,146 @@ +// The operator protecting itself from the drain it is arranging. +// +// A maintenance window blocks the eviction of the storage-node pod before it +// takes the backend node down, which is what makes `kubectl drain` wait rather +// than kill an SPDK process out from under a running node (§10). That protects +// the storage pod and nothing else, and on a converged worker the manager runs +// on the same host. +// +// The drain evicts in no particular order. So it can take the manager out +// before the manager has created the storage node's budget, and then nothing is +// left that would ever create it: the storage pod goes next, the SPDK process +// dies under a running backend node, and the control plane sees a node that +// vanished. That is precisely the failure the window exists to prevent, reached +// by evicting the thing preventing it. +// +// The answer is a budget the manager holds over itself for the length of the +// arrangement and no longer. It is created before the storage node's, and it is +// released in the same step that releases the storage node's — because the +// manager is on the node being drained, and a budget that outlived the window +// would leave the drain blocked forever on the pod that arranged it. +// +// design-storagenode.md §10 is the specification. + +package node + +import ( + "context" + "fmt" + + policyv1 "k8s.io/api/policy/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/util/intstr" + "sigs.k8s.io/controller-runtime/pkg/client" + logf "sigs.k8s.io/controller-runtime/pkg/log" + + atlaskube "github.com/simplyblock/atlas/kube" +) + +const ( + // managerAppLabel is what the operator's Deployment labels its pods with, + // and therefore what a budget over the manager selects. + // + // The label is used rather than a marker of this package's own, because + // labeling the manager's own pod to protect it is a write that has to + // succeed before the protection exists — and the window it protects is + // exactly the one where the manager may be evicted mid-write. What the + // Deployment already stamps is there before the operator starts. + managerAppLabel = "simplyblock-operator" + + // selfBudgetName is fixed rather than derived, because there is one + // manager: it holds the leader lease, so only one instance ever arranges a + // window, and a second name could only ever be a second copy of the same + // budget. A fixed name is also what lets a replacement manager find and + // clear the one its predecessor left. + selfBudgetName = "sb-maintenance-operator" +) + +// ProtectSelf blocks the manager's own eviction, for the case where the worker +// being drained is the one it runs on. +// +// It is a no-op otherwise, and that is the important half: a budget that held +// the manager while an unrelated worker drained would make the manager's own +// node undrainable for the length of every maintenance window in the cluster. +// +// A manager that does not know which node it is on holds nothing. The chart +// sets NODE_NAME from spec.nodeName, and an operator started without it cannot +// tell whose worker this is — creating the budget anyway would block its own +// eviction on every window, and skipping it leaves the case no worse than it +// was before this existed. +func (w *Workload) ProtectSelf(ctx context.Context, namespace, worker string) error { + if w.ManagerNode == "" || w.ManagerNode != worker { + return nil + } + + budget := &policyv1.PodDisruptionBudget{ + ObjectMeta: metav1.ObjectMeta{ + Name: selfBudgetName, + Namespace: namespace, + Labels: map[string]string{ + atlaskube.LabelApp: managerAppLabel, + }, + }, + Spec: policyv1.PodDisruptionBudgetSpec{ + // Zero rather than one: Kubernetes reads a budget allowing no + // disruption as a refusal of every voluntary eviction, which is + // what holds the drain rather than merely slowing it. + MaxUnavailable: &intstr.IntOrString{Type: intstr.Int, IntVal: 0}, + Selector: &metav1.LabelSelector{ + MatchLabels: map[string]string{"app": managerAppLabel}, + }, + }, + } + + err := w.Create(ctx, budget) + if err != nil && !apierrors.IsAlreadyExists(err) { + return fmt.Errorf("hold the manager's own eviction on worker %s: %w", worker, err) + } + if err == nil { + logf.FromContext(ctx).Info("the manager is holding its own eviction while it arranges "+ + "the maintenance of the worker it runs on", "worker", worker) + } + return nil +} + +// ReleaseSelf lets the manager be evicted again, and clears a budget a previous +// manager left behind. +// +// The two are one operation because they cannot be told apart from here: a +// budget that exists either belongs to the window this is ending or belongs to +// a manager that died holding it, and in both cases the answer is that it goes. +// Leaving one of the second kind would make its worker undrainable forever, by +// an object whose owner no longer exists. +func (w *Workload) ReleaseSelf(ctx context.Context, namespace string) error { + budget := &policyv1.PodDisruptionBudget{ObjectMeta: metav1.ObjectMeta{ + Name: selfBudgetName, + Namespace: namespace, + }} + if err := w.Delete(ctx, budget); err != nil && !apierrors.IsNotFound(err) { + return fmt.Errorf("release the manager's own eviction: %w", err) + } + return nil +} + +// ClearStaleSelfBudget removes a self-budget left by a manager that crashed +// while holding one. +// +// It runs once, on the instance that won the leader election, and that is the +// only place it can safely run: only the leader arranges a window, so only the +// leader's budget is ever live, and a replica clearing one on the way up would +// be clearing the leader's. +func (w *Workload) ClearStaleSelfBudget(ctx context.Context, namespace string) error { + var budget policyv1.PodDisruptionBudget + key := client.ObjectKey{Namespace: namespace, Name: selfBudgetName} + err := w.Get(ctx, key, &budget) + if apierrors.IsNotFound(err) { + return nil + } + if err != nil { + return fmt.Errorf("look for a self-budget left by a previous manager: %w", err) + } + + logf.FromContext(ctx).Info("clearing a maintenance budget the previous manager left over "+ + "itself, which would otherwise make its worker undrainable", "budget", selfBudgetName) + return w.ReleaseSelf(ctx, namespace) +} diff --git a/operator/internal/controllers/node/selfbudget_test.go b/operator/internal/controllers/node/selfbudget_test.go new file mode 100644 index 000000000..412b19bb2 --- /dev/null +++ b/operator/internal/controllers/node/selfbudget_test.go @@ -0,0 +1,293 @@ +// The operator protecting itself from the drain it is arranging. +// +// The case this is about is narrow and total: the manager pod runs on the +// worker being drained. `kubectl drain` evicts pods in no particular order, so +// it can take the manager out before the manager has created the storage node's +// own budget — and then nothing is left that would have created it, the SPDK +// process is killed under a running backend node, and the control plane sees a +// node that vanished. That is the failure the whole HostMaintenance action +// exists to prevent, reached by evicting the thing preventing it. +// +// design-storagenode.md §10. + +package node + +import ( + "context" + "errors" + "testing" + + corev1 "k8s.io/api/core/v1" + policyv1 "k8s.io/api/policy/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/types" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + "k8s.io/client-go/tools/events" + + atlaskube "github.com/simplyblock/atlas/kube" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/testsupport" +) + +const ( + budgetNamespace = "simplyblock" + budgetCluster = "a-cluster" + // managerWorker is where the manager pod runs, and drainedWorker is the one + // a window is opened on. The interesting case is the two being equal. + managerWorker = "worker-1" + otherWorker = "worker-2" +) + +// aWorkload builds a Workload that believes it runs on managerNode. +func aWorkload(t *testing.T, managerNode string, objs ...client.Object) (*Workload, client.Client) { + t.Helper() + scheme := testsupport.NewScheme(t, corev1.AddToScheme, policyv1.AddToScheme) + apiClient := fake.NewClientBuilder().WithScheme(scheme).WithObjects(objs...).Build() + return &Workload{Client: apiClient, ManagerNode: managerNode}, apiClient +} + +// managerPod is the operator's own pod, on the worker every case here is +// about, carrying the label its Deployment selects it by. +func managerPod() *corev1.Pod { + return &corev1.Pod{ + ObjectMeta: metav1.ObjectMeta{ + Name: "simplyblock-operator-abc", + Namespace: budgetNamespace, + Labels: map[string]string{"app": managerAppLabel}, + }, + Spec: corev1.PodSpec{NodeName: managerWorker}, + } +} + +func selfBudgetOf(t *testing.T, c client.Client) (*policyv1.PodDisruptionBudget, bool) { + t.Helper() + var budget policyv1.PodDisruptionBudget + key := types.NamespacedName{Namespace: budgetNamespace, Name: selfBudgetName} + err := c.Get(context.Background(), key, &budget) + if apierrors.IsNotFound(err) { + return nil, false + } + if err != nil { + t.Fatalf("reading the self-budget: %v", err) + } + return &budget, true +} + +// TestTheManagerHoldsItselfOnTheWorkerItIsDraining. Without it, the drain can +// evict the manager before the storage node's budget exists, and the budget is +// then never created by anything. +func TestTheManagerHoldsItselfOnTheWorkerItIsDraining(t *testing.T) { + w, c := aWorkload(t, managerWorker, managerPod()) + + if err := w.ProtectSelf(context.Background(), budgetNamespace, managerWorker); err != nil { + t.Fatalf("protecting the manager: %v", err) + } + + budget, found := selfBudgetOf(t, c) + if !found { + t.Fatal("no self-budget was created for the worker the manager runs on") + } + if budget.Spec.MaxUnavailable == nil || budget.Spec.MaxUnavailable.IntVal != 0 { + t.Errorf("maxUnavailable = %v, want 0: the budget's job is to allow no eviction at all", + budget.Spec.MaxUnavailable) + } + if got := budget.Spec.Selector.MatchLabels["app"]; got != managerAppLabel { + t.Errorf("the budget selects %q, want the manager's own pods", got) + } +} + +// TestTheManagerDoesNotHoldItselfOnSomebodyElsesWorker. A budget that blocked +// the manager's eviction while an unrelated worker drained would make the +// manager's own node undrainable for the length of every maintenance window in +// the cluster. +func TestTheManagerDoesNotHoldItselfOnSomebodyElsesWorker(t *testing.T) { + w, c := aWorkload(t, managerWorker, managerPod()) + + if err := w.ProtectSelf(context.Background(), budgetNamespace, otherWorker); err != nil { + t.Fatalf("protecting the manager: %v", err) + } + + if _, found := selfBudgetOf(t, c); found { + t.Error("a self-budget was created for a worker the manager does not run on") + } +} + +// TestAManagerThatDoesNotKnowWhereItRunsHoldsNothing. NODE_NAME is set by the +// chart from spec.nodeName, and an operator started without it cannot tell +// whether the worker being drained is its own. Creating the budget anyway would +// block the manager's eviction on every window in the cluster; skipping it +// leaves the case this protects against no worse than it was. +func TestAManagerThatDoesNotKnowWhereItRunsHoldsNothing(t *testing.T) { + w, c := aWorkload(t, "", managerPod()) + + if err := w.ProtectSelf(context.Background(), budgetNamespace, managerWorker); err != nil { + t.Fatalf("protecting the manager: %v", err) + } + + if _, found := selfBudgetOf(t, c); found { + t.Error("a self-budget was created by a manager that does not know which node it is on") + } +} + +// TestReleasingTheManagerIsWhatLetsTheDrainFinish. The manager is on the node +// being drained, so a budget that outlived the window would leave the drain +// blocked forever on the very pod that arranged it. +func TestReleasingTheManagerIsWhatLetsTheDrainFinish(t *testing.T) { + w, c := aWorkload(t, managerWorker, managerPod()) + ctx := context.Background() + + if err := w.ProtectSelf(ctx, budgetNamespace, managerWorker); err != nil { + t.Fatal(err) + } + if _, found := selfBudgetOf(t, c); !found { + t.Fatal("the self-budget was not created, so there is nothing to release") + } + + if err := w.ReleaseSelf(ctx, budgetNamespace); err != nil { + t.Fatalf("releasing the manager: %v", err) + } + if _, found := selfBudgetOf(t, c); found { + t.Error("the self-budget survived the release, so the drain stays blocked on the manager") + } + + // Releasing one that is already gone is the state being asked for: the + // window's terminal step runs it again, and a crash can leave it having run + // once already. + if err := w.ReleaseSelf(ctx, budgetNamespace); err != nil { + t.Errorf("releasing an absent self-budget reported an error: %v", err) + } +} + +// TestAStaleSelfBudgetIsCleanedUp. A manager that crashed between creating the +// budget and releasing it leaves a worker no drain can ever finish, and the +// object that would have removed it is the one that died. The replacement +// manager clears it on the way up. +func TestAStaleSelfBudgetIsCleanedUp(t *testing.T) { + stale := &policyv1.PodDisruptionBudget{ + ObjectMeta: metav1.ObjectMeta{Name: selfBudgetName, Namespace: budgetNamespace}, + } + w, c := aWorkload(t, managerWorker, managerPod(), stale) + + if err := w.ReleaseSelf(context.Background(), budgetNamespace); err != nil { + t.Fatalf("clearing the stale self-budget: %v", err) + } + if _, found := selfBudgetOf(t, c); found { + t.Error("the stale self-budget survived, so its worker cannot be drained") + } +} + +// TestTheSelfBudgetCarriesTheGroupsLabels, so that what a `kubectl get pdb` +// shows is identifiable as the operator's and matchable by one rule. +func TestTheSelfBudgetCarriesTheGroupsLabels(t *testing.T) { + w, c := aWorkload(t, managerWorker, managerPod()) + + if err := w.ProtectSelf(context.Background(), budgetNamespace, managerWorker); err != nil { + t.Fatal(err) + } + + budget, found := selfBudgetOf(t, c) + if !found { + t.Fatal("no self-budget") + } + if budget.Labels[atlaskube.LabelApp] == "" { + t.Errorf("labels = %v, want the group's own", budget.Labels) + } +} + +// unreachableControlPlane fails every read, which is the state a window has to +// hold its own eviction through: the budget's whole job is to exist before +// anything that can go wrong does. +type unreachableControlPlane struct { + ControlPlane +} + +func (unreachableControlPlane) StorageNode( + context.Context, string, string, +) (NodeReading, bool, error) { + return NodeReading{}, false, errUnreachableControlPlane +} + +var errUnreachableControlPlane = errors.New("the control plane is unreachable") + +// TestTheWindowHoldsTheManagerBeforeAnythingThatCanFail. The ordering is the +// mechanism, not a detail of it. A shutdown step that reached the control plane +// first, or created the storage node's budget first, would leave a window in +// which the drain can evict the manager — and the manager is the only thing +// that would have created either. +func TestTheWindowHoldsTheManagerBeforeAnythingThatCanFail(t *testing.T) { + node := &simplyblockv1alpha2.StorageNode{ + ObjectMeta: metav1.ObjectMeta{Name: "a-node", Namespace: budgetNamespace}, + Spec: simplyblockv1alpha2.StorageNodeSpec{ + ClusterRef: budgetCluster, + WorkerNode: managerWorker, + }, + } + scheme := testsupport.NewScheme(t, corev1.AddToScheme, policyv1.AddToScheme) + apiClient := fake.NewClientBuilder().WithScheme(scheme). + WithObjects(node, managerPod()).Build() + + r := &StorageNodeOpsReconciler{ + Client: apiClient, + Scheme: scheme, + Recorder: events.NewFakeRecorder(16), + API: unreachableControlPlane{}, + Workload: &Workload{Client: apiClient, ManagerNode: managerWorker}, + } + + ops := &simplyblockv1alpha2.StorageNodeOps{ + ObjectMeta: metav1.ObjectMeta{Name: "a-window", Namespace: budgetNamespace}, + Spec: simplyblockv1alpha2.StorageNodeOpsSpec{ + Action: simplyblockv1alpha2.StorageNodeOpsActionHostMaintenance, + NodeRef: "a-node", + }, + } + + // The step cannot finish: the control plane is unreachable. What it must + // have done first is hold the manager. + if _, err := r.maintenanceShutDown( + context.Background(), ops, node, "cluster-id", "node-id"); err == nil { + t.Fatal("the step reported success against an unreachable control plane") + } + + if _, found := selfBudgetOf(t, apiClient); !found { + t.Error("the manager was left evictable while the window could not reach the control plane") + } +} + +// TestTheWindowLetsTheManagerGoWhenItLetsTheStoragePodGo. Both are on the node +// being drained, so holding the manager past the point where the storage pod is +// released would block the drain on the manager instead of on the pod — the +// same deadlock, one object over. +func TestTheWindowLetsTheManagerGoWhenItLetsTheStoragePodGo(t *testing.T) { + node := &simplyblockv1alpha2.StorageNode{ + ObjectMeta: metav1.ObjectMeta{Name: "a-node", Namespace: budgetNamespace}, + Spec: simplyblockv1alpha2.StorageNodeSpec{ + ClusterRef: budgetCluster, + WorkerNode: managerWorker, + }, + } + stale := &policyv1.PodDisruptionBudget{ + ObjectMeta: metav1.ObjectMeta{Name: selfBudgetName, Namespace: budgetNamespace}, + } + scheme := testsupport.NewScheme(t, corev1.AddToScheme, policyv1.AddToScheme) + apiClient := fake.NewClientBuilder().WithScheme(scheme). + WithObjects(node, managerPod(), stale).Build() + + r := &StorageNodeOpsReconciler{ + Client: apiClient, + Scheme: scheme, + Recorder: events.NewFakeRecorder(16), + Workload: &Workload{Client: apiClient, ManagerNode: managerWorker}, + } + + if _, err := r.maintenanceRelease(context.Background(), node); err != nil { + t.Fatalf("releasing: %v", err) + } + if _, found := selfBudgetOf(t, apiClient); found { + t.Error("the manager is still held, so the drain it arranged cannot finish") + } +} diff --git a/operator/internal/controllers/node/workload.go b/operator/internal/controllers/node/workload.go index 885f21346..957e02a1d 100644 --- a/operator/internal/controllers/node/workload.go +++ b/operator/internal/controllers/node/workload.go @@ -79,6 +79,14 @@ type Workload struct { // refused request is not the same answer as an unreachable host. TLSEnabled bool TLSMutualEnabled bool + + // ManagerNode is the Kubernetes node the operator itself runs on, which the + // chart sets from spec.nodeName. It decides one question and no other: + // whether the worker a maintenance window is draining is the manager's own, + // in which case the manager holds its own eviction while it arranges the + // window (selfbudget.go). Empty when nothing set it, and then it holds + // nothing. + ManagerNode string } // NodeAddress is the per-pod DNS name the control plane is given as node_address From ec6037ace0add688f26d856700b597a064c1b028 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 12:57:42 +0200 Subject: [PATCH 044/206] feat(api): a name a label carries is bounded at what a label holds MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit §19 of design-api-upgrade measured every Kubernetes identifier this product derives from a name somebody chose, and found that nothing validates one before writing it: an overflow is not a rejected resource but a reconcile that retries forever, on a derived write, while the object being reconciled says nothing about the name that caused it. The preflight half was built. The admission half was not, except on ClusterDeploymentConfig, and nothing said the markers there and the missing ones elsewhere were the same rule. They are one rule, so it lands as one. metadata.name is bounded at 63 by an XValidation on StorageCluster, StoragePool, and StorageNode, because the target model writes all three into label values it cannot shorten — the cluster into storage.simplyblock.io/cluster and io.simplyblock.storagenodeset, the pool into /pool, the node into /node — and where a name travels into a label, the label's limit binds rather than the object's. It is a rule rather than a marker because metadata.name is one of the two metadata fields a validation rule can see. The pool's row carries a second failure behind it: /pool is also the selector a pool lists its own classes with, so an overlong name is a read that fails as well as a write. Every field carrying a cluster's name takes MaxLength=63, since a longer value names nothing that can exist. Eight of them across seven kinds, found by walking the schema rather than by listing: OperatorOps' sits inside its discovery action, and ClusterDeploymentConfig.spec.cluster.name is not a reference but the name itself — at 253, so a document the API server accepts could hold a CreatingCluster step that can never succeed. Its clusterRef was bounded at 253 too, which is an object name's limit and not a cluster name's. Status fields are left alone: they record what was resolved from an input this rule bounds, so a maximum there catches no mistake and can only refuse a status write. One live defect fell out of asking the question. nodeNameFormula was declared against an object name's 253 bytes while its output goes into the StorageDevice mirror's /node label. A regional cluster name and a worker a cloud named after its fully qualified domain name are 68 bytes between them, so the overflow was what ordinary inputs produced rather than what a long one did. The formula carries the label's limit now; nothing is stranded, because a name that fitted the wider limit is returned unchanged and createNodes keys its idempotence on the worker and slot a node's spec records. Admission-side uniqueness is deliberately not built, and §19.7 now says why. The target model closes three of §19.8's four routes by scoping — every kind but PersistentVolumeOps is namespaced, and what they derive is unique where their inputs are — and the fourth is the default StorageClass, which the pool's reconcile already answers the way that section prescribes: it adopts the name only when the occupant is recognizably its own, and otherwise emits StorageClassNameTaken. A fail-closed webhook in front of that would refuse a legal cluster over a class the cluster does not need to work. Two of §19.5's assignments were wrong about where the product went, and the design and the preflight rows are corrected together. The per-pool node label key took the UUID fix and is now storage-pool., which retires its budget and its ambiguous concatenation at once. The two StorageClass labels did not: they are how a person assigns a class they wrote to a pool, so a value nobody can type is a contract nobody can enter. Those bound the input instead, which is what makes the rule on three kinds load-bearing. Three tests, each red first. A schema enumeration over config/crd/bases, which fails on a kind added later with an unbounded reference without anybody extending it. The formula against real cloud hostnames. One envtest, so the rules are proven to bite rather than merely to be present. Co-Authored-By: Claude Fable 5 --- ...eet.simplyblock.io_clusterdeployments.yaml | 15 +- .../fleet.simplyblock.io_fleetoperations.yaml | 16 +- ...mplyblock.io_clusterdeploymentconfigs.yaml | 15 +- .../storage.simplyblock.io_operatorops.yaml | 4 + ...orage.simplyblock.io_storagebackupops.yaml | 8 +- ....simplyblock.io_storagebackuppolicies.yaml | 4 + ...storage.simplyblock.io_storagebackups.yaml | 4 + ...rage.simplyblock.io_storageclusterops.yaml | 4 + ...torage.simplyblock.io_storageclusters.yaml | 3 + .../storage.simplyblock.io_storagenodes.yaml | 7 + .../storage.simplyblock.io_storagepools.yaml | 9 + .../v1alpha2/clusterdeploymentconfig_types.go | 14 +- operator/api/v1alpha2/names_test.go | 227 ++++++++++++++++++ operator/api/v1alpha2/operatorops_types.go | 4 + operator/api/v1alpha2/storagebackup_types.go | 4 + .../api/v1alpha2/storagebackupops_types.go | 4 + .../api/v1alpha2/storagebackuppolicy_types.go | 4 + operator/api/v1alpha2/storagecluster_types.go | 13 + .../api/v1alpha2/storageclusterops_types.go | 4 + operator/api/v1alpha2/storagenode_types.go | 11 + operator/api/v1alpha2/storagepool_types.go | 13 + ...mplyblock.io_clusterdeploymentconfigs.yaml | 15 +- .../storage.simplyblock.io_operatorops.yaml | 4 + ...orage.simplyblock.io_storagebackupops.yaml | 8 +- ....simplyblock.io_storagebackuppolicies.yaml | 4 + ...storage.simplyblock.io_storagebackups.yaml | 4 + ...rage.simplyblock.io_storageclusterops.yaml | 4 + ...torage.simplyblock.io_storageclusters.yaml | 3 + .../storage.simplyblock.io_storagenodes.yaml | 7 + .../storage.simplyblock.io_storagepools.yaml | 9 + operator/dist/install.yaml | 63 ++++- .../crd-redesign/design-api-upgrade.md | 113 +++++++-- .../cluster/cel_validation_test.go | 54 +++++ .../controllers/deployment/expansion.go | 15 +- .../controllers/deployment/nodename_test.go | 56 +++++ ...mplyblock.io_clusterdeploymentconfigs.yaml | 15 +- .../storage.simplyblock.io_operatorops.yaml | 4 + ...orage.simplyblock.io_storagebackupops.yaml | 8 +- ....simplyblock.io_storagebackuppolicies.yaml | 4 + ...storage.simplyblock.io_storagebackups.yaml | 4 + ...rage.simplyblock.io_storageclusterops.yaml | 4 + ...torage.simplyblock.io_storageclusters.yaml | 3 + .../storage.simplyblock.io_storagenodes.yaml | 7 + .../storage.simplyblock.io_storagepools.yaml | 9 + .../internal/upgrade/derive/boundary_test.go | 5 +- operator/internal/upgrade/derive/labels.go | 25 +- 46 files changed, 772 insertions(+), 62 deletions(-) create mode 100644 operator/api/v1alpha2/names_test.go create mode 100644 operator/internal/controllers/deployment/nodename_test.go diff --git a/fleet/config/crd/bases/fleet.simplyblock.io_clusterdeployments.yaml b/fleet/config/crd/bases/fleet.simplyblock.io_clusterdeployments.yaml index b23954430..73e3fc49d 100644 --- a/fleet/config/crd/bases/fleet.simplyblock.io_clusterdeployments.yaml +++ b/fleet/config/crd/bases/fleet.simplyblock.io_clusterdeployments.yaml @@ -142,8 +142,13 @@ spec: maxLength: 32 type: string name: - description: Name is the StorageCluster's name. - maxLength: 253 + description: |- + Name is the StorageCluster's name, and is therefore held to what such a + name may be rather than to what an object name may be. A longer value is a + document the API server accepts and a CreatingCluster step that can never + succeed, since the cluster it would write is one the API server refuses + (design-api-upgrade.md §19.4). + maxLength: 63 type: string nodesPerSocket: description: |- @@ -210,7 +215,11 @@ spec: Absent means the document creates the cluster in Cluster. Setting it to a cluster that does not exist, or leaving it absent when one already does, is refused rather than reconciled. - maxLength: 253 + + The maximum is what a StorageCluster name may be and not the 253 an object + name may be: a reference between the two names nothing that can exist + (design-api-upgrade.md §19.4). + maxLength: 63 type: string edgeCluster: description: |- diff --git a/fleet/config/crd/bases/fleet.simplyblock.io_fleetoperations.yaml b/fleet/config/crd/bases/fleet.simplyblock.io_fleetoperations.yaml index fa7d4407b..c0e6f08f3 100644 --- a/fleet/config/crd/bases/fleet.simplyblock.io_fleetoperations.yaml +++ b/fleet/config/crd/bases/fleet.simplyblock.io_fleetoperations.yaml @@ -109,6 +109,10 @@ spec: creates. It is copied to the draft's own clusterRef, so that re-running discovery after an expansion produces a growth document naming the same cluster. + + Bounded at what a StorageCluster name may be, since a longer value names + nothing that can exist (design-api-upgrade.md §19.4). + maxLength: 63 type: string configName: description: |- @@ -262,8 +266,12 @@ spec: - message: field is immutable rule: self == oldSelf clusterRef: - description: ClusterRef names the StorageCluster the operation - runs against. + description: |- + ClusterRef names the StorageCluster the operation runs against. + + Bounded at what a StorageCluster name may be, since a longer value names + nothing that can exist (design-api-upgrade.md §19.4). + maxLength: 63 type: string x-kubernetes-validations: - message: field is immutable @@ -372,6 +380,10 @@ spec: ClusterRef names the StorageCluster this operation acts on. The operation never owns its target, because deleting the record of an operation must not delete the cluster it operated on. + + Bounded at what a StorageCluster name may be, since a longer value names + nothing that can exist (design-api-upgrade.md §19.4). + maxLength: 63 type: string x-kubernetes-validations: - message: field is immutable diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml index 95533924d..aed657d41 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml @@ -140,8 +140,13 @@ spec: maxLength: 32 type: string name: - description: Name is the StorageCluster's name. - maxLength: 253 + description: |- + Name is the StorageCluster's name, and is therefore held to what such a + name may be rather than to what an object name may be. A longer value is a + document the API server accepts and a CreatingCluster step that can never + succeed, since the cluster it would write is one the API server refuses + (design-api-upgrade.md §19.4). + maxLength: 63 type: string nodesPerSocket: description: |- @@ -208,7 +213,11 @@ spec: Absent means the document creates the cluster in Cluster. Setting it to a cluster that does not exist, or leaving it absent when one already does, is refused rather than reconciled. - maxLength: 253 + + The maximum is what a StorageCluster name may be and not the 253 an object + name may be: a reference between the two names nothing that can exist + (design-api-upgrade.md §19.4). + maxLength: 63 type: string edgeCluster: description: |- diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_operatorops.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_operatorops.yaml index 1a09ed465..e4c509e08 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_operatorops.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_operatorops.yaml @@ -93,6 +93,10 @@ spec: creates. It is copied to the draft's own clusterRef, so that re-running discovery after an expansion produces a growth document naming the same cluster. + + Bounded at what a StorageCluster name may be, since a longer value names + nothing that can exist (design-api-upgrade.md §19.4). + maxLength: 63 type: string configName: description: |- diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagebackupops.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagebackupops.yaml index 892696622..b6f7cdeb2 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagebackupops.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagebackupops.yaml @@ -102,8 +102,12 @@ spec: - message: field is immutable rule: self == oldSelf clusterRef: - description: ClusterRef names the StorageCluster the operation runs - against. + description: |- + ClusterRef names the StorageCluster the operation runs against. + + Bounded at what a StorageCluster name may be, since a longer value names + nothing that can exist (design-api-upgrade.md §19.4). + maxLength: 63 type: string x-kubernetes-validations: - message: field is immutable diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagebackuppolicies.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagebackuppolicies.yaml index 6392049c3..8961b3b34 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagebackuppolicies.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagebackuppolicies.yaml @@ -119,6 +119,10 @@ spec: description: |- ClusterRef names the StorageCluster whose backup target this policy writes to. + + Bounded at what a StorageCluster name may be, since a longer value names + nothing that can exist (design-api-upgrade.md §19.4). + maxLength: 63 type: string x-kubernetes-validations: - message: field is immutable diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagebackups.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagebackups.yaml index 6653c5bc3..20eb63e98 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagebackups.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagebackups.yaml @@ -250,6 +250,10 @@ spec: description: |- ClusterRef names the StorageCluster whose store this backup was found in. With BackupID it is the whole of this object's identity. + + Bounded at what a StorageCluster name may be, since a longer value names + nothing that can exist (design-api-upgrade.md §19.4). + maxLength: 63 type: string x-kubernetes-validations: - message: field is immutable diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusterops.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusterops.yaml index 276e35d78..934db7003 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusterops.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusterops.yaml @@ -243,6 +243,10 @@ spec: ClusterRef names the StorageCluster this operation acts on. The operation never owns its target, because deleting the record of an operation must not delete the cluster it operated on. + + Bounded at what a StorageCluster name may be, since a longer value names + nothing that can exist (design-api-upgrade.md §19.4). + maxLength: 63 type: string x-kubernetes-validations: - message: field is immutable diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusters.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusters.yaml index b36e750f5..23925fce6 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusters.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusters.yaml @@ -1587,6 +1587,9 @@ spec: type: integer type: object type: object + x-kubernetes-validations: + - message: a StorageCluster name is at most 63 characters, because it is written into label values on StorageClasses, StorageDevices, and worker Nodes + rule: size(self.metadata.name) <= 63 served: true storage: true subresources: diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagenodes.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagenodes.yaml index afcbe3faf..e1704901d 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagenodes.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagenodes.yaml @@ -395,6 +395,10 @@ spec: ClusterRef names the StorageCluster this node belongs to. The cluster also owns this object by controller reference, so deleting the cluster deletes its nodes. + + Bounded at what a StorageCluster name may be, since a longer value names + nothing that can exist (design-api-upgrade.md §19.4). + maxLength: 63 type: string x-kubernetes-validations: - message: field is immutable @@ -821,6 +825,9 @@ spec: type: string type: object type: object + x-kubernetes-validations: + - message: a StorageNode name is at most 63 characters, because it is written into the storage.simplyblock.io/node label on every StorageDevice of this node + rule: size(self.metadata.name) <= 63 served: true storage: true subresources: diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagepools.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagepools.yaml index a92c0af67..498c28856 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagepools.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagepools.yaml @@ -315,6 +315,12 @@ spec: pool's own finalizer while classes are assigned or volumes are bound. Immutable from creation: which cluster a pool is in is its identity. + + The maximum is what a StorageCluster name may be rather than what a + reference may be: a longer value names nothing that can exist, and the + reference is immutable, so admitting one creates a pool whose only + remedy is deletion (design-api-upgrade.md §19.4). + maxLength: 63 type: string x-kubernetes-validations: - message: field is immutable @@ -564,6 +570,9 @@ spec: type: string type: object type: object + x-kubernetes-validations: + - message: a StoragePool name is at most 63 characters, because it is written into the storage.simplyblock.io/pool label that assigns StorageClasses to this pool + rule: size(self.metadata.name) <= 63 served: true storage: true subresources: diff --git a/operator/api/v1alpha2/clusterdeploymentconfig_types.go b/operator/api/v1alpha2/clusterdeploymentconfig_types.go index 312e4c468..2dd28f471 100644 --- a/operator/api/v1alpha2/clusterdeploymentconfig_types.go +++ b/operator/api/v1alpha2/clusterdeploymentconfig_types.go @@ -202,9 +202,13 @@ type NodeSet struct { // one. It carries the layout fields a cluster cannot change later, so that a // reviewer sees them before the cluster exists rather than after. type ClusterTemplate struct { - // Name is the StorageCluster's name. + // Name is the StorageCluster's name, and is therefore held to what such a + // name may be rather than to what an object name may be. A longer value is a + // document the API server accepts and a CreatingCluster step that can never + // succeed, since the cluster it would write is one the API server refuses + // (design-api-upgrade.md §19.4). // +kubebuilder:validation:Required - // +kubebuilder:validation:MaxLength=253 + // +kubebuilder:validation:MaxLength=63 Name string `json:"name"` // MaxSubsystemCount is the maximum number of NVMe-oF subsystems each storage @@ -332,7 +336,11 @@ type ClusterDeploymentConfigSpec struct { // Absent means the document creates the cluster in Cluster. Setting it to a // cluster that does not exist, or leaving it absent when one already does, // is refused rather than reconciled. - // +kubebuilder:validation:MaxLength=253 + // + // The maximum is what a StorageCluster name may be and not the 253 an object + // name may be: a reference between the two names nothing that can exist + // (design-api-upgrade.md §19.4). + // +kubebuilder:validation:MaxLength=63 // +optional ClusterRef string `json:"clusterRef,omitempty"` diff --git a/operator/api/v1alpha2/names_test.go b/operator/api/v1alpha2/names_test.go new file mode 100644 index 000000000..c9318dc99 --- /dev/null +++ b/operator/api/v1alpha2/names_test.go @@ -0,0 +1,227 @@ +// §19's name bounds, read back out of the schemas the markers produce. +// +// The markers are the whole of the rule — there is no code to test — so what is +// worth holding is that every kind that needs one has one, and that the number +// is the same number everywhere. A bound applied to the kinds somebody +// remembered is the state this test was written to end: ClusterDeploymentConfig +// carried MaxLength markers while StoragePool and StorageNode carried none, and +// nothing said the three were the same rule. +// +// It reads config/crd/bases rather than the Go source, because a marker that +// does not reach the schema is not a bound. The enumeration is the assertion: a +// kind added later with an unbounded reference fails this without anybody +// extending it. + +package v1alpha2 + +import ( + "fmt" + "os" + "path/filepath" + "strings" + "testing" + + apiextensionsv1 "k8s.io/apiextensions-apiserver/pkg/apis/apiextensions/v1" + "sigs.k8s.io/yaml" + + "github.com/simplyblock/atlas/kube" +) + +// nameBound is what a name a label carries may be. It is a label's own limit +// rather than a budget: a reference inside it can still overflow a key that +// joins it to a namespace and a pool, and what the bound closes is every row +// where the name stands alone. +const nameBound = int64(kube.MaxLabelValueLength) + +// labelBoundKinds are the kinds whose metadata.name is written into a label +// value somewhere, together with the write that binds it. A name that overflows +// there is not a rejected resource but a reconcile that retries forever, since +// the refusal lands on the derived write and the object being reconciled says +// nothing about the name that caused it. +var labelBoundKinds = map[string]string{ + "StorageCluster": "storage.simplyblock.io/cluster on a StorageClass and on a StorageDevice, " + + "and io.simplyblock.storagenodeset on every worker Node the cluster claims", + "StoragePool": "storage.simplyblock.io/pool, which is also the selector the pool " + + "lists its own classes with", + "StorageNode": "storage.simplyblock.io/node on every StorageDevice the mirror writes", +} + +// generatedCRDs reads every CRD this repository generates, by kind. +func generatedCRDs(t *testing.T) map[string]*apiextensionsv1.CustomResourceDefinition { + t.Helper() + + dir := filepath.Join("..", "..", "config", "crd", "bases") + entries, err := os.ReadDir(dir) + if err != nil { + t.Fatalf("reading the generated CRDs: %v", err) + } + + out := make(map[string]*apiextensionsv1.CustomResourceDefinition) + for _, entry := range entries { + if !strings.HasSuffix(entry.Name(), ".yaml") { + continue + } + raw, err := os.ReadFile(filepath.Join(dir, entry.Name())) + if err != nil { + t.Fatalf("reading %s: %v", entry.Name(), err) + } + var crd apiextensionsv1.CustomResourceDefinition + if err := yaml.Unmarshal(raw, &crd); err != nil { + t.Fatalf("parsing %s: %v", entry.Name(), err) + } + out[crd.Spec.Names.Kind] = &crd + } + if len(out) == 0 { + t.Fatal("no CRDs were read, so every assertion below would pass vacuously") + } + return out +} + +// v1alpha2Schema returns a kind's v1alpha2 schema, or nil for a kind that has no such +// version. A kind the redesign has not reached carries no rule (§19.9). +func v1alpha2Schema(crd *apiextensionsv1.CustomResourceDefinition) *apiextensionsv1.JSONSchemaProps { + for i := range crd.Spec.Versions { + version := &crd.Spec.Versions[i] + if version.Name == "v1alpha2" && version.Schema != nil { + return version.Schema.OpenAPIV3Schema + } + } + return nil +} + +// becomesAClusterName are the fields that are not a reference but a name: the +// value one kind carries becomes another object's metadata.name, so it is held +// to what that name may be. Admitting more is a document the API server accepts +// and a creation step that can never succeed. +var becomesAClusterName = map[string][]string{ + "ClusterDeploymentConfig": {"cluster", "name"}, +} + +// property walks a path of property names, so a field nested inside a template +// can be named without the test knowing how the schema is shaped. +func property( + node apiextensionsv1.JSONSchemaProps, path []string, +) (apiextensionsv1.JSONSchemaProps, error) { + at := node + for i, name := range path { + child, ok := at.Properties[name] + if !ok { + return at, fmt.Errorf("no property %s under spec.%s", + name, strings.Join(path[:i], ".")) + } + at = child + } + return at, nil +} + +// clusterRefs walks a schema and reports every property called clusterRef under +// it, by the path it sits at. The walk is recursive because a reference is not +// always a spec's own field: a discovery draft carries one inside the action it +// describes, and a rule written for the top level alone would leave that one +// unbounded while reading as though it covered everything. +func clusterRefs( + path string, node apiextensionsv1.JSONSchemaProps, into map[string]apiextensionsv1.JSONSchemaProps, +) { + for name, child := range node.Properties { + where := path + "." + name + if name == "clusterRef" { + into[where] = child + } + clusterRefs(where, child, into) + } + if node.Items != nil && node.Items.Schema != nil { + clusterRefs(path+"[]", *node.Items.Schema, into) + } +} + +// TestEveryClusterReferenceIsBoundedByWhatAClusterNameMayBe covers every +// reference to a StorageCluster the group's specs carry. +// +// The bound on the reference follows from the bound on the name rather than +// from anything about the referring kind: a reference longer than a +// StorageCluster name may be names nothing that can exist, so it is refused at +// admission rather than resolved forever by a controller that will never find +// it. +// +// Only spec is walked. A status carrying the same reference is a record of +// what the operator resolved, copied from an input this rule already bounds, so +// a maximum there could not catch a mistake and could only turn a status write +// into one the API server refuses. +func TestEveryClusterReferenceIsBoundedByWhatAClusterNameMayBe(t *testing.T) { + var found int + for kind, crd := range generatedCRDs(t) { + root := v1alpha2Schema(crd) + if root == nil { + continue + } + spec, ok := root.Properties["spec"] + if !ok { + continue + } + + refs := map[string]apiextensionsv1.JSONSchemaProps{} + clusterRefs("spec", spec, refs) + if named, ok := becomesAClusterName[kind]; ok { + at, err := property(spec, named) + if err != nil { + t.Errorf("%s: %v", kind, err) + } else { + refs["spec."+strings.Join(named, ".")] = at + } + } + for where, ref := range refs { + found++ + switch { + case ref.MaxLength == nil: + t.Errorf("%s.%s carries no maxLength, so it admits 253 characters of a "+ + "StorageCluster name that could never be that long", kind, where) + case *ref.MaxLength != nameBound: + t.Errorf("%s.%s is bounded at %d and a StorageCluster name at %d; the two "+ + "are one rule, and a field bounded above the name is not bounded", + kind, where, *ref.MaxLength, nameBound) + } + } + } + + if found == 0 { + t.Fatal("no kind declares a clusterRef, which is not what this group looks like") + } +} + +// TestANameALabelCarriesIsBoundedAtALabelsLimit covers the kinds whose own +// metadata.name travels into a label value. +// +// metadata.name takes no MaxLength marker, because it is not this schema's +// field: it is one of the two metadata fields a CRD validation rule can see +// (§19.7), so the bound is a rule on the type rather than a marker on a +// property. +func TestANameALabelCarriesIsBoundedAtALabelsLimit(t *testing.T) { + all := generatedCRDs(t) + for kind, written := range labelBoundKinds { + crd, ok := all[kind] + if !ok { + t.Errorf("%s has no CRD, and it is listed here as a kind whose name is written "+ + "into %s", kind, written) + continue + } + root := v1alpha2Schema(crd) + if root == nil { + t.Errorf("%s has no v1alpha2 schema to carry the rule", kind) + continue + } + + var bounded bool + for _, rule := range root.XValidations { + if strings.Contains(rule.Rule, "self.metadata.name") && + strings.Contains(rule.Rule, "63") { + bounded = true + break + } + } + if !bounded { + t.Errorf("%s carries no rule bounding self.metadata.name, and its name is "+ + "written into %s, where a label stops at %d bytes", + kind, written, nameBound) + } + } +} diff --git a/operator/api/v1alpha2/operatorops_types.go b/operator/api/v1alpha2/operatorops_types.go index 907cc9d43..fccd4aa94 100644 --- a/operator/api/v1alpha2/operatorops_types.go +++ b/operator/api/v1alpha2/operatorops_types.go @@ -177,6 +177,10 @@ type DiscoverSpec struct { // creates. It is copied to the draft's own clusterRef, so that re-running // discovery after an expansion produces a growth document naming the same // cluster. + // + // Bounded at what a StorageCluster name may be, since a longer value names + // nothing that can exist (design-api-upgrade.md §19.4). + // +kubebuilder:validation:MaxLength=63 // +optional ClusterRef string `json:"clusterRef,omitempty"` } diff --git a/operator/api/v1alpha2/storagebackup_types.go b/operator/api/v1alpha2/storagebackup_types.go index 3c3571b7a..52570d6fc 100644 --- a/operator/api/v1alpha2/storagebackup_types.go +++ b/operator/api/v1alpha2/storagebackup_types.go @@ -178,6 +178,10 @@ type BackupCopy struct { type StorageBackupSpec struct { // ClusterRef names the StorageCluster whose store this backup was found in. // With BackupID it is the whole of this object's identity. + // + // Bounded at what a StorageCluster name may be, since a longer value names + // nothing that can exist (design-api-upgrade.md §19.4). + // +kubebuilder:validation:MaxLength=63 // +operator-sdk:csv:customresourcedefinitions:type=spec,displayName="Cluster Ref" // +kubebuilder:validation:Required // +k8s:immutable diff --git a/operator/api/v1alpha2/storagebackupops_types.go b/operator/api/v1alpha2/storagebackupops_types.go index 4f890b031..10a98815e 100644 --- a/operator/api/v1alpha2/storagebackupops_types.go +++ b/operator/api/v1alpha2/storagebackupops_types.go @@ -139,6 +139,10 @@ type RestoreSpec struct { // StorageBackupOpsSpec is one operation to perform against a backup. type StorageBackupOpsSpec struct { // ClusterRef names the StorageCluster the operation runs against. + // + // Bounded at what a StorageCluster name may be, since a longer value names + // nothing that can exist (design-api-upgrade.md §19.4). + // +kubebuilder:validation:MaxLength=63 // +operator-sdk:csv:customresourcedefinitions:type=spec,displayName="Cluster Ref" // +kubebuilder:validation:Required // +k8s:immutable diff --git a/operator/api/v1alpha2/storagebackuppolicy_types.go b/operator/api/v1alpha2/storagebackuppolicy_types.go index 5c3faa5f9..40b3de950 100644 --- a/operator/api/v1alpha2/storagebackuppolicy_types.go +++ b/operator/api/v1alpha2/storagebackuppolicy_types.go @@ -34,6 +34,10 @@ const ( type StorageBackupPolicySpec struct { // ClusterRef names the StorageCluster whose backup target this policy writes // to. + // + // Bounded at what a StorageCluster name may be, since a longer value names + // nothing that can exist (design-api-upgrade.md §19.4). + // +kubebuilder:validation:MaxLength=63 // +operator-sdk:csv:customresourcedefinitions:type=spec,displayName="Cluster Ref" // +kubebuilder:validation:Required // +k8s:immutable diff --git a/operator/api/v1alpha2/storagecluster_types.go b/operator/api/v1alpha2/storagecluster_types.go index c4f1de752..bbdd0e807 100644 --- a/operator/api/v1alpha2/storagecluster_types.go +++ b/operator/api/v1alpha2/storagecluster_types.go @@ -873,6 +873,19 @@ type StorageClusterStatus struct { // `kubectl get sc` reaches storageclasses.storage.k8s.io and never this kind, // and the operator writes a StorageClass per pool, which puts both kinds in // every cluster this runs in. +// The name is bounded at 63 rather than at the 253 an object name may be, +// because it travels into label values this operator writes and cannot shorten: +// storage.simplyblock.io/cluster on every generated StorageClass and every +// mirrored StorageDevice, and io.simplyblock.storagenodeset on every worker the +// cluster claims. Where a name is copied into a label, the label's limit binds +// and not the object's (design-api-upgrade.md §19.1). +// +// It is a rule rather than a MaxLength marker because metadata.name is not this +// schema's field. It is one of the two metadata fields a validation rule can +// see, which is what makes the bound expressible at all (§19.4, §19.7), and the +// alternative is an overflow that surfaces as a reconcile retrying forever on a +// label write while the cluster says nothing about the name that caused it. +// +kubebuilder:validation:XValidation:rule="size(self.metadata.name) <= 63",message="a StorageCluster name is at most 63 characters, because it is written into label values on StorageClasses, StorageDevices, and worker Nodes" // +kubebuilder:storageversion // +kubebuilder:object:root=true // +kubebuilder:subresource:status diff --git a/operator/api/v1alpha2/storageclusterops_types.go b/operator/api/v1alpha2/storageclusterops_types.go index bac490011..b46fb8cb9 100644 --- a/operator/api/v1alpha2/storageclusterops_types.go +++ b/operator/api/v1alpha2/storageclusterops_types.go @@ -118,6 +118,10 @@ type StorageClusterOpsSpec struct { // ClusterRef names the StorageCluster this operation acts on. The operation // never owns its target, because deleting the record of an operation must // not delete the cluster it operated on. + // + // Bounded at what a StorageCluster name may be, since a longer value names + // nothing that can exist (design-api-upgrade.md §19.4). + // +kubebuilder:validation:MaxLength=63 // +kubebuilder:validation:Required // +k8s:immutable ClusterRef string `json:"clusterRef"` diff --git a/operator/api/v1alpha2/storagenode_types.go b/operator/api/v1alpha2/storagenode_types.go index 63d412930..e83174683 100644 --- a/operator/api/v1alpha2/storagenode_types.go +++ b/operator/api/v1alpha2/storagenode_types.go @@ -236,6 +236,10 @@ type StorageNodeSpec struct { // ClusterRef names the StorageCluster this node belongs to. The cluster also // owns this object by controller reference, so deleting the cluster deletes // its nodes. + // + // Bounded at what a StorageCluster name may be, since a longer value names + // nothing that can exist (design-api-upgrade.md §19.4). + // +kubebuilder:validation:MaxLength=63 // +kubebuilder:validation:Required // +k8s:immutable ClusterRef string `json:"clusterRef"` @@ -491,6 +495,13 @@ type StorageNodeStatus struct { // are the ones a fresh install applies. An upgrade of an existing cluster reaches // it through the storage rewrite rather than through this marker, for the reason // storagenodeops_types.go states. +// +// The name is bounded at a label's 63 bytes, because the StorageDevice mirror +// writes it into storage.simplyblock.io/node on every device this node reports. +// The operator's own names fit by construction — the formula that builds them +// carries the same limit — and the rule is what holds a node somebody authored +// to the same bound (design-api-upgrade.md §19.1, §19.4). +// +kubebuilder:validation:XValidation:rule="size(self.metadata.name) <= 63",message="a StorageNode name is at most 63 characters, because it is written into the storage.simplyblock.io/node label on every StorageDevice of this node" // +kubebuilder:storageversion // +kubebuilder:object:root=true // +kubebuilder:subresource:status diff --git a/operator/api/v1alpha2/storagepool_types.go b/operator/api/v1alpha2/storagepool_types.go index d5dd4c828..59589749f 100644 --- a/operator/api/v1alpha2/storagepool_types.go +++ b/operator/api/v1alpha2/storagepool_types.go @@ -171,6 +171,12 @@ type StoragePoolSpec struct { // pool's own finalizer while classes are assigned or volumes are bound. // // Immutable from creation: which cluster a pool is in is its identity. + // + // The maximum is what a StorageCluster name may be rather than what a + // reference may be: a longer value names nothing that can exist, and the + // reference is immutable, so admitting one creates a pool whose only + // remedy is deletion (design-api-upgrade.md §19.4). + // +kubebuilder:validation:MaxLength=63 // +kubebuilder:validation:Required // +k8s:immutable ClusterRef string `json:"clusterRef"` @@ -284,6 +290,13 @@ type StoragePoolStatus struct { // upgrade of an existing cluster applies the same CRD with storage held at // v1alpha1 and flips it with the storage rewrite once the conversion webhook is // serving. +// +// The name is bounded at a label's 63 bytes, because storage.simplyblock.io/pool +// carries it on every StorageClass assigned to this pool. That label is also the +// selector the pool lists its own classes with, so an overlong name is not only +// a write the API server refuses but a read: the pool would never find a class +// it had been given (design-api-upgrade.md §19.1, §19.4). +// +kubebuilder:validation:XValidation:rule="size(self.metadata.name) <= 63",message="a StoragePool name is at most 63 characters, because it is written into the storage.simplyblock.io/pool label that assigns StorageClasses to this pool" // +kubebuilder:storageversion // +kubebuilder:object:root=true // +kubebuilder:subresource:status diff --git a/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml b/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml index 95533924d..aed657d41 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml @@ -140,8 +140,13 @@ spec: maxLength: 32 type: string name: - description: Name is the StorageCluster's name. - maxLength: 253 + description: |- + Name is the StorageCluster's name, and is therefore held to what such a + name may be rather than to what an object name may be. A longer value is a + document the API server accepts and a CreatingCluster step that can never + succeed, since the cluster it would write is one the API server refuses + (design-api-upgrade.md §19.4). + maxLength: 63 type: string nodesPerSocket: description: |- @@ -208,7 +213,11 @@ spec: Absent means the document creates the cluster in Cluster. Setting it to a cluster that does not exist, or leaving it absent when one already does, is refused rather than reconciled. - maxLength: 253 + + The maximum is what a StorageCluster name may be and not the 253 an object + name may be: a reference between the two names nothing that can exist + (design-api-upgrade.md §19.4). + maxLength: 63 type: string edgeCluster: description: |- diff --git a/operator/config/crd/bases/storage.simplyblock.io_operatorops.yaml b/operator/config/crd/bases/storage.simplyblock.io_operatorops.yaml index 1a09ed465..e4c509e08 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_operatorops.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_operatorops.yaml @@ -93,6 +93,10 @@ spec: creates. It is copied to the draft's own clusterRef, so that re-running discovery after an expansion produces a growth document naming the same cluster. + + Bounded at what a StorageCluster name may be, since a longer value names + nothing that can exist (design-api-upgrade.md §19.4). + maxLength: 63 type: string configName: description: |- diff --git a/operator/config/crd/bases/storage.simplyblock.io_storagebackupops.yaml b/operator/config/crd/bases/storage.simplyblock.io_storagebackupops.yaml index 892696622..b6f7cdeb2 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_storagebackupops.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_storagebackupops.yaml @@ -102,8 +102,12 @@ spec: - message: field is immutable rule: self == oldSelf clusterRef: - description: ClusterRef names the StorageCluster the operation runs - against. + description: |- + ClusterRef names the StorageCluster the operation runs against. + + Bounded at what a StorageCluster name may be, since a longer value names + nothing that can exist (design-api-upgrade.md §19.4). + maxLength: 63 type: string x-kubernetes-validations: - message: field is immutable diff --git a/operator/config/crd/bases/storage.simplyblock.io_storagebackuppolicies.yaml b/operator/config/crd/bases/storage.simplyblock.io_storagebackuppolicies.yaml index 6392049c3..8961b3b34 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_storagebackuppolicies.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_storagebackuppolicies.yaml @@ -119,6 +119,10 @@ spec: description: |- ClusterRef names the StorageCluster whose backup target this policy writes to. + + Bounded at what a StorageCluster name may be, since a longer value names + nothing that can exist (design-api-upgrade.md §19.4). + maxLength: 63 type: string x-kubernetes-validations: - message: field is immutable diff --git a/operator/config/crd/bases/storage.simplyblock.io_storagebackups.yaml b/operator/config/crd/bases/storage.simplyblock.io_storagebackups.yaml index 6653c5bc3..20eb63e98 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_storagebackups.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_storagebackups.yaml @@ -250,6 +250,10 @@ spec: description: |- ClusterRef names the StorageCluster whose store this backup was found in. With BackupID it is the whole of this object's identity. + + Bounded at what a StorageCluster name may be, since a longer value names + nothing that can exist (design-api-upgrade.md §19.4). + maxLength: 63 type: string x-kubernetes-validations: - message: field is immutable diff --git a/operator/config/crd/bases/storage.simplyblock.io_storageclusterops.yaml b/operator/config/crd/bases/storage.simplyblock.io_storageclusterops.yaml index 276e35d78..934db7003 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_storageclusterops.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_storageclusterops.yaml @@ -243,6 +243,10 @@ spec: ClusterRef names the StorageCluster this operation acts on. The operation never owns its target, because deleting the record of an operation must not delete the cluster it operated on. + + Bounded at what a StorageCluster name may be, since a longer value names + nothing that can exist (design-api-upgrade.md §19.4). + maxLength: 63 type: string x-kubernetes-validations: - message: field is immutable diff --git a/operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml b/operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml index b36e750f5..23925fce6 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml @@ -1587,6 +1587,9 @@ spec: type: integer type: object type: object + x-kubernetes-validations: + - message: a StorageCluster name is at most 63 characters, because it is written into label values on StorageClasses, StorageDevices, and worker Nodes + rule: size(self.metadata.name) <= 63 served: true storage: true subresources: diff --git a/operator/config/crd/bases/storage.simplyblock.io_storagenodes.yaml b/operator/config/crd/bases/storage.simplyblock.io_storagenodes.yaml index afcbe3faf..e1704901d 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_storagenodes.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_storagenodes.yaml @@ -395,6 +395,10 @@ spec: ClusterRef names the StorageCluster this node belongs to. The cluster also owns this object by controller reference, so deleting the cluster deletes its nodes. + + Bounded at what a StorageCluster name may be, since a longer value names + nothing that can exist (design-api-upgrade.md §19.4). + maxLength: 63 type: string x-kubernetes-validations: - message: field is immutable @@ -821,6 +825,9 @@ spec: type: string type: object type: object + x-kubernetes-validations: + - message: a StorageNode name is at most 63 characters, because it is written into the storage.simplyblock.io/node label on every StorageDevice of this node + rule: size(self.metadata.name) <= 63 served: true storage: true subresources: diff --git a/operator/config/crd/bases/storage.simplyblock.io_storagepools.yaml b/operator/config/crd/bases/storage.simplyblock.io_storagepools.yaml index a92c0af67..498c28856 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_storagepools.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_storagepools.yaml @@ -315,6 +315,12 @@ spec: pool's own finalizer while classes are assigned or volumes are bound. Immutable from creation: which cluster a pool is in is its identity. + + The maximum is what a StorageCluster name may be rather than what a + reference may be: a longer value names nothing that can exist, and the + reference is immutable, so admitting one creates a pool whose only + remedy is deletion (design-api-upgrade.md §19.4). + maxLength: 63 type: string x-kubernetes-validations: - message: field is immutable @@ -564,6 +570,9 @@ spec: type: string type: object type: object + x-kubernetes-validations: + - message: a StoragePool name is at most 63 characters, because it is written into the storage.simplyblock.io/pool label that assigns StorageClasses to this pool + rule: size(self.metadata.name) <= 63 served: true storage: true subresources: diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index 40f18db7c..d06671076 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -813,8 +813,13 @@ spec: maxLength: 32 type: string name: - description: Name is the StorageCluster's name. - maxLength: 253 + description: |- + Name is the StorageCluster's name, and is therefore held to what such a + name may be rather than to what an object name may be. A longer value is a + document the API server accepts and a CreatingCluster step that can never + succeed, since the cluster it would write is one the API server refuses + (design-api-upgrade.md §19.4). + maxLength: 63 type: string nodesPerSocket: description: |- @@ -881,7 +886,11 @@ spec: Absent means the document creates the cluster in Cluster. Setting it to a cluster that does not exist, or leaving it absent when one already does, is refused rather than reconciled. - maxLength: 253 + + The maximum is what a StorageCluster name may be and not the 253 an object + name may be: a reference between the two names nothing that can exist + (design-api-upgrade.md §19.4). + maxLength: 63 type: string edgeCluster: description: |- @@ -1787,6 +1796,10 @@ spec: creates. It is copied to the draft's own clusterRef, so that re-running discovery after an expansion produces a growth document naming the same cluster. + + Bounded at what a StorageCluster name may be, since a longer value names + nothing that can exist (design-api-upgrade.md §19.4). + maxLength: 63 type: string configName: description: |- @@ -3682,8 +3695,12 @@ spec: - message: field is immutable rule: self == oldSelf clusterRef: - description: ClusterRef names the StorageCluster the operation runs - against. + description: |- + ClusterRef names the StorageCluster the operation runs against. + + Bounded at what a StorageCluster name may be, since a longer value names + nothing that can exist (design-api-upgrade.md §19.4). + maxLength: 63 type: string x-kubernetes-validations: - message: field is immutable @@ -3965,6 +3982,10 @@ spec: description: |- ClusterRef names the StorageCluster whose backup target this policy writes to. + + Bounded at what a StorageCluster name may be, since a longer value names + nothing that can exist (design-api-upgrade.md §19.4). + maxLength: 63 type: string x-kubernetes-validations: - message: field is immutable @@ -4363,6 +4384,10 @@ spec: description: |- ClusterRef names the StorageCluster whose store this backup was found in. With BackupID it is the whole of this object's identity. + + Bounded at what a StorageCluster name may be, since a longer value names + nothing that can exist (design-api-upgrade.md §19.4). + maxLength: 63 type: string x-kubernetes-validations: - message: field is immutable @@ -4780,6 +4805,10 @@ spec: ClusterRef names the StorageCluster this operation acts on. The operation never owns its target, because deleting the record of an operation must not delete the cluster it operated on. + + Bounded at what a StorageCluster name may be, since a longer value names + nothing that can exist (design-api-upgrade.md §19.4). + maxLength: 63 type: string x-kubernetes-validations: - message: field is immutable @@ -6539,6 +6568,10 @@ spec: type: integer type: object type: object + x-kubernetes-validations: + - message: a StorageCluster name is at most 63 characters, because it is written + into label values on StorageClasses, StorageDevices, and worker Nodes + rule: size(self.metadata.name) <= 63 served: true storage: true subresources: @@ -7582,6 +7615,10 @@ spec: ClusterRef names the StorageCluster this node belongs to. The cluster also owns this object by controller reference, so deleting the cluster deletes its nodes. + + Bounded at what a StorageCluster name may be, since a longer value names + nothing that can exist (design-api-upgrade.md §19.4). + maxLength: 63 type: string x-kubernetes-validations: - message: field is immutable @@ -8016,6 +8053,11 @@ spec: type: string type: object type: object + x-kubernetes-validations: + - message: a StorageNode name is at most 63 characters, because it is written + into the storage.simplyblock.io/node label on every StorageDevice of this + node + rule: size(self.metadata.name) <= 63 served: true storage: true subresources: @@ -9242,6 +9284,12 @@ spec: pool's own finalizer while classes are assigned or volumes are bound. Immutable from creation: which cluster a pool is in is its identity. + + The maximum is what a StorageCluster name may be rather than what a + reference may be: a longer value names nothing that can exist, and the + reference is immutable, so admitting one creates a pool whose only + remedy is deletion (design-api-upgrade.md §19.4). + maxLength: 63 type: string x-kubernetes-validations: - message: field is immutable @@ -9495,6 +9543,11 @@ spec: type: string type: object type: object + x-kubernetes-validations: + - message: a StoragePool name is at most 63 characters, because it is written + into the storage.simplyblock.io/pool label that assigns StorageClasses + to this pool + rule: size(self.metadata.name) <= 63 served: true storage: true subresources: diff --git a/operator/docs/designs/crd-redesign/design-api-upgrade.md b/operator/docs/designs/crd-redesign/design-api-upgrade.md index 489ca762f..5babc3d8a 100644 --- a/operator/docs/designs/crd-redesign/design-api-upgrade.md +++ b/operator/docs/designs/crd-redesign/design-api-upgrade.md @@ -1288,9 +1288,9 @@ Seven labels are built from a name a user chose. Every row is live today. | What is built | Breaks when | Longest input that works | Fix | |-------------------------------------------------------------|--------------------------------------------------------------------------------|-----------------------------|-------------------| -| `simplyblock.io/pool...`, a key | The namespace, cluster, and pool names together exceed 56 characters | A 27-character pool name | Truncate and hash | -| `storage.simplyblock.io/cluster` on a `StorageClass` | The cluster name exceeds 63 characters | A 63-character cluster name | Use a UUID | -| `storage.simplyblock.io/pool` on a `StorageClass` | The `StoragePool` name exceeds 63 characters | A 63-character pool name | Use a UUID | +| `simplyblock.io/pool...`, a key | The namespace, cluster, and pool names together exceed 56 characters | A 27-character pool name | Use a UUID | +| `storage.simplyblock.io/cluster` on a `StorageClass` | The cluster name exceeds 63 characters | A 63-character cluster name | Bound the input | +| `storage.simplyblock.io/pool` on a `StorageClass` | The `StoragePool` name exceeds 63 characters | A 63-character pool name | Bound the input | | `io.simplyblock.storagenodeset` | The `StorageNodeSet` name exceeds 63 characters | A 63-character set name | Bound the input | | `storage.simplyblock.io/worker` | The `Node` name exceeds 63 characters | A 63-character node name | Truncate and hash | | `simplyblock.io/drain-node` | Character 63 is `-` or `.`, which a label value may not end on | A 62-character node name | Truncate and hash | @@ -1333,21 +1333,29 @@ characters long. ### 19.4 Bounding the Cluster Reference -`spec.clusterName` carries no maximum length and no pattern on either -`StoragePoolSpec` (`storagepool_types.go:115`) or `StorageNodeSetSpec` -(`storagenodeset_types.go:40`), and it feeds three of the seven labels and three -of the object names above. **A `+kubebuilder:validation:MaxLength=63` on it is -what turns an overlong cluster reference into a rejected create rather than a -reconcile that retries forever.** The marker lands on `v1alpha2`'s -`spec.clusterRef`, because §7.2 renames the field and retires `StorageNodeSet`, -and never on `v1alpha1` (§19.9). - -**63 is a label's limit and not a budget the marker can guarantee.** Two of the -rows a cluster name feeds share their 63 bytes with a namespace and a pool name, -so a cluster reference inside the limit still overflows the `simplyblock.io/pool` -key when the other two are long. What the marker closes is the rows where the -cluster name stands alone, which are the `StorageClass` label and the two -`Secret` names. +**A `+kubebuilder:validation:MaxLength=63` on a cluster reference is what turns +an overlong one into a rejected create rather than a reconcile that retries +forever.** The markers land on `v1alpha2` and never on `v1alpha1` (§19.9). + +**The bound on a reference follows from the bound on the name, so it is the same +number on every kind that carries one.** A reference longer than a +`StorageCluster` name may be names nothing that can exist, which makes the +question of what the referring kind does with it beside the point: eight fields +across seven kinds carry a cluster's name, and a bound applied to the ones +somebody remembered is not a bound. Two of the eight are not references at all +but names — `ClusterDeploymentConfig.spec.cluster.name` becomes a +`StorageCluster`'s `metadata.name`, so admitting more there is a document the API +server accepts and a `CreatingCluster` step that can never succeed. + +A `status` carrying the same reference is deliberately left unbounded. It records +what the operator resolved, copied from an input this rule already bounds, so a +maximum there could catch no mistake and could only turn a status write into one +the API server refuses. + +**63 is a label's limit and not a budget the marker can guarantee.** A row that +shares its 63 bytes with a namespace and a pool name still overflows when the +other two are long. What the marker closes is the rows where the cluster name +stands alone. **The cluster's own name is bounded by a type-level rule, which no `MaxLength` can reach.** `metadata.name` is one of the two metadata fields a CRD validation @@ -1355,17 +1363,54 @@ rule can see (§19.7), so the name itself is bounded by the rule and the reference by the marker: ```go -// +kubebuilder:validation:XValidation:rule="size(self.metadata.name) <= 63",message="a StorageCluster name is at most 63 characters, because it is written into a StorageClass label" +// +kubebuilder:validation:XValidation:rule="size(self.metadata.name) <= 63",message="a StorageCluster name is at most 63 characters, because it is written into label values on StorageClasses, StorageDevices, and worker Nodes" ``` +**Three kinds carry that rule, not one.** The cluster's name is the one §19.2 +measured, but the target model writes two more names into label values, and both +were found by asking the same question of the kinds around it rather than by +re-deriving the table: + +| Kind | Written into | +|------------------|-----------------------------------------------------------------------| +| `StorageCluster` | `storage.simplyblock.io/cluster`, and `io.simplyblock.storagenodeset` | +| `StoragePool` | `storage.simplyblock.io/pool` | +| `StorageNode` | `storage.simplyblock.io/node` | + +The pool's row is the one with a second failure behind it. That label is also the +selector a pool lists its own classes with, so an overlong pool name is not only a +write the API server refuses but a read: the pool would never find a class it had +been given. + +The node's row is the one where the bound is the smaller half of the fix. A +`StorageNode` is named by the operator rather than by a user, from the formula in +`expansion.go`, and that formula was declared against an object name's 253 bytes +while its output travels into a label — the mistake §19.1 exists to name. A +regional cluster name and a worker a cloud named after its fully qualified domain +name are 68 bytes between them, so the overflow was what ordinary inputs +produced. The formula carries the label's limit now, and the rule on the type is +what holds a node somebody authored to the same bound. + ### 19.5 The Three Fixes Every row above resolves one of three ways, and which one applies follows from who owns the name rather than from how long it is. **Use a UUID.** When a stable identifier is already at hand, nothing reads the -current value, and the label exists to be selected on rather than read. The two -`StorageClass` labels are this case. +current value, and the label exists to be selected on rather than read. + +The row this turned out to fit is the per-pool key on a worker `Node`, which the +target model writes as `storage.simplyblock.io/storage-pool.`: it is +the tightest row of §19.2, it is read by the CSI node plugin as a prefix match +rather than by its parts, and a UUID retires the whole of its budget problem +along with §19.8's ambiguous concatenation. + +The two `StorageClass` labels were assumed to be this case and are not. +`storage.simplyblock.io/cluster` and `storage.simplyblock.io/pool` are the +assignment itself — they are how a person assigns a class they wrote to a pool — +so a value nobody can type is a contract nobody can enter. Those two rows resolve +by bounding the input instead, which is what makes §19.4's rule on three kinds +rather than one load-bearing. **Bound the input.** When the long name is this API's to refuse. A field somebody types has no business being 200 characters, so the answer is no at @@ -1424,6 +1469,21 @@ The webhook races itself. Two creates admitted concurrently each see a free derived name, so the reconciler treats a collision as a terminal condition with an event rather than as something admission prevented. +**In the target model that fallback is the whole of the answer, and no +uniqueness webhook is built.** The row above is written for the current model, +where four routes take two resources to one derived name (§19.8), and the target +model closes three of them by construction: every kind but +`PersistentVolumeOps` is namespaced, and the names they derive are unique within +the namespace their inputs are unique in. What is left is the default +`StorageClass`, which is cluster-scoped and named `simplyblock--`. +The pool's reconcile already answers that one the way this row prescribes — it +adopts the name only when the occupant is recognizably the class it would have +written, and otherwise emits `StorageClassNameTaken` and leaves the pool without +a default. A fail-closed webhook in front of that would refuse a legal cluster +over a class that is not required for the cluster to work, which is worse than +the condition it replaces. The remaining two routes of §19.8 are what an upgrade +introduces rather than what a write can, and they stay the preflight's. + The last row is the one this document turns on: every mechanism above it runs on a write, and the objects an upgrade has to survive were written before the rule existed. @@ -2438,12 +2498,15 @@ prose, its check is here and not repeated in both places. **Names (§19)** - [ ] Every name and label of §19.2 and §19.3 has a bounded derivation. -- [ ] `metadata.name` on the `v1alpha2` `StorageCluster` is bounded at 63 by an - `XValidation` rule, and `StoragePoolSpec.clusterRef` by `MaxLength`. +- [x] `metadata.name` is bounded at 63 by an `XValidation` rule on the `v1alpha2` + `StorageCluster`, `StoragePool`, and `StorageNode`, and every field + carrying a cluster's name by `MaxLength` (§19.4). - [ ] The truncate-and-hash helper is extracted from `nodeprobe.ObjectName` into `atlas-lib/kube`, and no call site rolls its own. -- [ ] §19.8's uniqueness rules are enforced at admission, and a collision that - races admission is terminal with an event. +- [x] §19.8's uniqueness rules are enforced where a write can still break one. + The target model leaves the default `StorageClass` as the only case, and + the pool's reconcile is where it is terminal with an event (§19.7); the + remaining routes are an upgrade's and stay the preflight's. **The tool** diff --git a/operator/internal/controllers/cluster/cel_validation_test.go b/operator/internal/controllers/cluster/cel_validation_test.go index 6cd71ae5d..fa616df6c 100644 --- a/operator/internal/controllers/cluster/cel_validation_test.go +++ b/operator/internal/controllers/cluster/cel_validation_test.go @@ -333,3 +333,57 @@ func TestStorageClusterStepRejectsAnUnknownValue(t *testing.T) { t.Fatalf("rejected for the wrong reason: %v", err) } } + +// TestStorageClusterNameIsBoundedAtALabelsLimit proves that the rule of §19.4 +// is enforced by the apiserver and not merely present in the schema. +// +// The name and the reference are the same rule seen from two sides, so both are +// exercised here: a cluster that could not be called this, and a pool naming a +// cluster that could not exist. The bound is a label's 63 bytes rather than the +// 253 the API server allows an object name, because the cluster's name is +// written into storage.simplyblock.io/cluster on every StorageClass the +// operator generates and into io.simplyblock.storagenodeset on every worker it +// claims — and where a name travels into a label, the label's limit binds. +// +// Coverage of the other kinds carrying the rule is the schema enumeration in +// api/v1alpha2/names_test.go, which is what keeps a kind added later from +// escaping it. What this test adds is that the marker bites. +func TestStorageClusterNameIsBoundedAtALabelsLimit(t *testing.T) { + apiClient := apiServer(t) + ctx := context.Background() + + legal := strings.Repeat("a", 63) + overlong := strings.Repeat("a", 64) + + cluster := &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{Name: legal, Namespace: "default"}, + Spec: simplyblockv1alpha2.StorageClusterSpec{ + MaxSubsystemCount: ptr.To(int32(10)), + VCPUCount: ptr.To(int32(6)), + }, + } + if err := apiClient.Create(ctx, cluster); err != nil { + t.Fatalf("a 63-character name is the longest a label carries, so it must be "+ + "accepted, got: %v", err) + } + t.Cleanup(func() { _ = apiClient.Delete(ctx, cluster) }) + + refused := cluster.DeepCopy() + refused.Name = overlong + refused.ResourceVersion = "" + if err := apiClient.Create(ctx, refused); err == nil { + t.Error("the apiserver accepted a 64-character cluster name, which every " + + "StorageClass and worker label derived from it would then be refused for") + _ = apiClient.Delete(ctx, refused) + } + + pool := &simplyblockv1alpha2.StoragePool{ + ObjectMeta: metav1.ObjectMeta{GenerateName: "bound-", Namespace: "default"}, + Spec: simplyblockv1alpha2.StoragePoolSpec{ClusterRef: overlong}, + } + if err := apiClient.Create(ctx, pool); err == nil { + t.Error("the apiserver accepted a clusterRef longer than a StorageCluster name " + + "may be, which is an immutable reference to an object that cannot exist") + _ = apiClient.Delete(ctx, pool) + } +} diff --git a/operator/internal/controllers/deployment/expansion.go b/operator/internal/controllers/deployment/expansion.go index 76af84a15..0ebf23e46 100644 --- a/operator/internal/controllers/deployment/expansion.go +++ b/operator/internal/controllers/deployment/expansion.go @@ -617,7 +617,20 @@ func targetClusterName( // slot it fills, never for the worker, because the name has to stay stable when a // migration re-points the node onto another host (design-storagenode.md §3.1) — // the worker is in the name's digest rather than in its text. -var nodeNameFormula = kube.Formula{} +// +// The limit is a label's 63 bytes and not the 253 an object name may be, because +// the name travels: the StorageDevice mirror writes it into +// storage.simplyblock.io/node on every device of the node. §19.1 of +// design-api-upgrade.md is the rule, and the arithmetic is not academic — a +// regional cluster name and a worker a cloud named after its fully qualified +// domain name are 68 bytes between them, so the overflow is what ordinary inputs +// produce rather than what a long one does. +// +// Shortening the limit does not strand the nodes of a cluster that already has +// some. A name that fitted the wider limit is returned unchanged whenever it +// also fits this one, and createNodes finds what exists by the worker and slot +// its spec records rather than by re-deriving the name. +var nodeNameFormula = kube.Formula{Limit: kube.MaxLabelValueLength} func nodeName(cluster, worker string, slot int32) string { return nodeNameFormula.Derive(cluster, worker, fmt.Sprintf("%d", slot)).Value diff --git a/operator/internal/controllers/deployment/nodename_test.go b/operator/internal/controllers/deployment/nodename_test.go new file mode 100644 index 000000000..48d93d049 --- /dev/null +++ b/operator/internal/controllers/deployment/nodename_test.go @@ -0,0 +1,56 @@ +// The limit that binds a StorageNode's derived name. +// +// §19.1 of design-api-upgrade.md states the rule these cases hold the formula +// to: where a name is copied into a label, the label's 63 bytes bind it and not +// the 253 an object name may be. A StorageNode's name is copied into +// storage.simplyblock.io/node on every StorageDevice the mirror writes, so the +// name is a label value that happens to also be an object name. +// +// The worker names are real ones rather than strings of a chosen length, +// because the question the rule answers is whether ordinary inputs overflow, +// and a cloud that names a worker after its fully qualified domain name spends +// forty-two of the sixty-three before the cluster name is in it. + +package deployment + +import ( + "testing" + + "github.com/simplyblock/atlas/kube" +) + +func TestANodeNameFitsTheLabelItIsCopiedInto(t *testing.T) { + for _, tc := range []struct { + name string + cluster string + worker string + slot int32 + }{ + { + name: "a short cluster on a bare hostname", + cluster: "production", + worker: "worker-3", + }, + { + name: "a regional cluster on an EKS worker", + cluster: "production-eu-central-1", + worker: "ip-10-0-1-23.eu-central-1.compute.internal", + slot: 1, + }, + { + name: "a GKE worker, which carries the cluster name in its own", + cluster: "storage", + worker: "gke-production-eu-central-1-storage-pool-7f3a2b1c-k4nd", + }, + } { + t.Run(tc.name, func(t *testing.T) { + derived := nodeName(tc.cluster, tc.worker, tc.slot) + + if errs := kube.Validate(kube.LabelValue, derived); len(errs) > 0 { + t.Errorf("the name %q (%d bytes) is not a legal label value, so writing it "+ + "into storage.simplyblock.io/node is a StorageDevice reconcile that "+ + "retries forever: %v", derived, len(derived), errs) + } + }) + } +} diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml index 95533924d..aed657d41 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml @@ -140,8 +140,13 @@ spec: maxLength: 32 type: string name: - description: Name is the StorageCluster's name. - maxLength: 253 + description: |- + Name is the StorageCluster's name, and is therefore held to what such a + name may be rather than to what an object name may be. A longer value is a + document the API server accepts and a CreatingCluster step that can never + succeed, since the cluster it would write is one the API server refuses + (design-api-upgrade.md §19.4). + maxLength: 63 type: string nodesPerSocket: description: |- @@ -208,7 +213,11 @@ spec: Absent means the document creates the cluster in Cluster. Setting it to a cluster that does not exist, or leaving it absent when one already does, is refused rather than reconciled. - maxLength: 253 + + The maximum is what a StorageCluster name may be and not the 253 an object + name may be: a reference between the two names nothing that can exist + (design-api-upgrade.md §19.4). + maxLength: 63 type: string edgeCluster: description: |- diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_operatorops.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_operatorops.yaml index 1a09ed465..e4c509e08 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_operatorops.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_operatorops.yaml @@ -93,6 +93,10 @@ spec: creates. It is copied to the draft's own clusterRef, so that re-running discovery after an expansion produces a growth document naming the same cluster. + + Bounded at what a StorageCluster name may be, since a longer value names + nothing that can exist (design-api-upgrade.md §19.4). + maxLength: 63 type: string configName: description: |- diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagebackupops.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagebackupops.yaml index 892696622..b6f7cdeb2 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagebackupops.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagebackupops.yaml @@ -102,8 +102,12 @@ spec: - message: field is immutable rule: self == oldSelf clusterRef: - description: ClusterRef names the StorageCluster the operation runs - against. + description: |- + ClusterRef names the StorageCluster the operation runs against. + + Bounded at what a StorageCluster name may be, since a longer value names + nothing that can exist (design-api-upgrade.md §19.4). + maxLength: 63 type: string x-kubernetes-validations: - message: field is immutable diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagebackuppolicies.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagebackuppolicies.yaml index 6392049c3..8961b3b34 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagebackuppolicies.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagebackuppolicies.yaml @@ -119,6 +119,10 @@ spec: description: |- ClusterRef names the StorageCluster whose backup target this policy writes to. + + Bounded at what a StorageCluster name may be, since a longer value names + nothing that can exist (design-api-upgrade.md §19.4). + maxLength: 63 type: string x-kubernetes-validations: - message: field is immutable diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagebackups.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagebackups.yaml index 6653c5bc3..20eb63e98 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagebackups.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagebackups.yaml @@ -250,6 +250,10 @@ spec: description: |- ClusterRef names the StorageCluster whose store this backup was found in. With BackupID it is the whole of this object's identity. + + Bounded at what a StorageCluster name may be, since a longer value names + nothing that can exist (design-api-upgrade.md §19.4). + maxLength: 63 type: string x-kubernetes-validations: - message: field is immutable diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusterops.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusterops.yaml index 276e35d78..934db7003 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusterops.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusterops.yaml @@ -243,6 +243,10 @@ spec: ClusterRef names the StorageCluster this operation acts on. The operation never owns its target, because deleting the record of an operation must not delete the cluster it operated on. + + Bounded at what a StorageCluster name may be, since a longer value names + nothing that can exist (design-api-upgrade.md §19.4). + maxLength: 63 type: string x-kubernetes-validations: - message: field is immutable diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusters.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusters.yaml index b36e750f5..23925fce6 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusters.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusters.yaml @@ -1587,6 +1587,9 @@ spec: type: integer type: object type: object + x-kubernetes-validations: + - message: a StorageCluster name is at most 63 characters, because it is written into label values on StorageClasses, StorageDevices, and worker Nodes + rule: size(self.metadata.name) <= 63 served: true storage: true subresources: diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagenodes.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagenodes.yaml index afcbe3faf..e1704901d 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagenodes.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagenodes.yaml @@ -395,6 +395,10 @@ spec: ClusterRef names the StorageCluster this node belongs to. The cluster also owns this object by controller reference, so deleting the cluster deletes its nodes. + + Bounded at what a StorageCluster name may be, since a longer value names + nothing that can exist (design-api-upgrade.md §19.4). + maxLength: 63 type: string x-kubernetes-validations: - message: field is immutable @@ -821,6 +825,9 @@ spec: type: string type: object type: object + x-kubernetes-validations: + - message: a StorageNode name is at most 63 characters, because it is written into the storage.simplyblock.io/node label on every StorageDevice of this node + rule: size(self.metadata.name) <= 63 served: true storage: true subresources: diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagepools.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagepools.yaml index a92c0af67..498c28856 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagepools.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagepools.yaml @@ -315,6 +315,12 @@ spec: pool's own finalizer while classes are assigned or volumes are bound. Immutable from creation: which cluster a pool is in is its identity. + + The maximum is what a StorageCluster name may be rather than what a + reference may be: a longer value names nothing that can exist, and the + reference is immutable, so admitting one creates a pool whose only + remedy is deletion (design-api-upgrade.md §19.4). + maxLength: 63 type: string x-kubernetes-validations: - message: field is immutable @@ -564,6 +570,9 @@ spec: type: string type: object type: object + x-kubernetes-validations: + - message: a StoragePool name is at most 63 characters, because it is written into the storage.simplyblock.io/pool label that assigns StorageClasses to this pool + rule: size(self.metadata.name) <= 63 served: true storage: true subresources: diff --git a/operator/internal/upgrade/derive/boundary_test.go b/operator/internal/upgrade/derive/boundary_test.go index fda19d437..e928b5f15 100644 --- a/operator/internal/upgrade/derive/boundary_test.go +++ b/operator/internal/upgrade/derive/boundary_test.go @@ -5,8 +5,9 @@ // The expected numbers are the design's measured ones, so a test that disagrees // with one has found either a formula change or an error in the audit. They are // written as the fixed cost each formula spends, because that is what a reader -// can check against the literal in the row: the design's "a 37-character cluster -// name" is 63 less this table's 26. +// can check against the literal in the row: §19.2's "a 27-character pool name" +// is 63 less this table's five characters of prefix, two separators, and the +// namespace and cluster names CI uses. // // Validity is asserted with k8s.io/apimachinery/pkg/util/validation rather than // with a regular expression written for the test, so the assertion tracks the diff --git a/operator/internal/upgrade/derive/labels.go b/operator/internal/upgrade/derive/labels.go index f35a41c07..ae8d51667 100644 --- a/operator/internal/upgrade/derive/labels.go +++ b/operator/internal/upgrade/derive/labels.go @@ -50,13 +50,18 @@ func Labels() []upgrade.Derivation { // The separator is a dot, which is legal inside all three names, so this row is // also one of §19.8's ambiguous concatenations: cluster a-b with pool c and // cluster a with pool b-c produce one key in one namespace. +// +// The advice is a UUID rather than a digest because that is where the product +// went: the target model writes storage.simplyblock.io/storage-pool., +// which retires the budget and the ambiguity together, and the CSI node plugin +// reads the key by its prefix rather than by its parts. func poolNodeLabelKey() Rule { return Rule{ RuleID: IDPoolNodeLabelKey, Summary: "bounds the worker label key a pool patches onto the nodes it allows", Where: "Node label key simplyblock.io/pool...", Which: upgrade.ModelCurrent, - Resolution: upgrade.FixTruncateAndHash, + Resolution: upgrade.FixUseUUID, Unique: upgrade.SpaceCluster, Build: atlaskube.Formula{ Kind: atlaskube.LabelKeyName, @@ -81,16 +86,20 @@ func poolNodeLabelKey() Rule { // (simplyblockstoragepool_controller.go). // // A bare name against the bare limit, so a 63-character cluster name works and -// a 64-character one does not. §19.5 resolves this row by using a UUID rather -// than by truncating, since nothing reads the value and the label exists to be -// selected on. +// a 64-character one does not. +// +// §19.5 assumed this row would take a UUID, on the grounds that nothing reads +// the value. Something does: the label is how a person assigns a class they +// wrote to a pool, so a value nobody can type is a contract nobody can enter. +// The row is bounded at the input instead, by the rule on StorageCluster's own +// metadata.name (§19.4). func storageClassClusterLabel() Rule { return Rule{ RuleID: IDStorageClassCluster, Summary: "bounds the cluster label a generated StorageClass carries", Where: "StorageClass label storage.simplyblock.io/cluster", Which: upgrade.ModelCurrent, - Resolution: upgrade.FixUseUUID, + Resolution: upgrade.FixBoundInput, Unique: upgrade.SpaceShared, Build: atlaskube.Formula{Kind: atlaskube.LabelValue}, Enumerate: func(_ context.Context, s *upgrade.Scope) ([]upgrade.Input, error) { @@ -108,13 +117,17 @@ func storageClassClusterLabel() Rule { // storageClassPoolLabel is storage.simplyblock.io/pool on the same object, // carrying the pool's own name. +// +// It is bounded at the input for the reason the cluster label is, and with one +// more behind it: the label is the selector a pool lists its own classes with, +// so an overlong name is a read that fails as well as a write. func storageClassPoolLabel() Rule { return Rule{ RuleID: IDStorageClassPool, Summary: "bounds the pool label a generated StorageClass carries", Where: "StorageClass label storage.simplyblock.io/pool", Which: upgrade.ModelCurrent, - Resolution: upgrade.FixUseUUID, + Resolution: upgrade.FixBoundInput, Unique: upgrade.SpaceShared, Build: atlaskube.Formula{Kind: atlaskube.LabelValue}, Enumerate: func(_ context.Context, s *upgrade.Scope) ([]upgrade.Input, error) { From a4bd121d9404c728fe771de9202c5bf7e8f1106b Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 13:44:44 +0200 Subject: [PATCH 045/206] feat(upgrade): a legacy volume handle resolves its pool and records the answer MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A CSI volume handle is clusterID:poolID:volumeID, and the middle segment is not always an id: volumes provisioned before the v2 API migration spell the pool's name there. The field cannot be corrected — the API server refuses any edit to a PersistentVolume's CSI source or a VolumeSnapshotContent's source — so §16.4 resolves the name once, writes the answer to metadata, and has every reader prefer it. Half of that existed. Detection did, because it needs nothing but the handle; the resolution, the write, and the reading rule did not, and neither did snapshots. The rule is lvol.NormalizeHandle, beside ParseHandle, and it compares two handles and knows nothing about Kubernetes. The annotation is taken only when it agrees with the field about the cluster and the volume, and anything else is ignored and reported: an annotation is metadata, so whoever may edit an object's labels could otherwise point a volume at another cluster. A field that carries no readable handle is refused whatever the annotation says, because the annotation is a record about the field rather than a replacement for one. Where the two strings come from is kube.NormalizedHandle, which takes a handle and an annotation map rather than an object. That is what lets both kinds share the rule: VolumeSnapshotContent belongs to the external snapshotter's module, and atlas-lib does not take a dependency on it for a map lookup. kube.VolumeHandleFromPV stays, and stays the field's own answer. The two questions differ the way ParseHandle and Split differ, and both have callers: which pool a volume is in is what anything reaching for the control plane wants, and what the spec literally says is what the migration wants, since that is how it finds the legacy spellings at all. The PVC controller and the group-snapshot webhook read the normalized one now. Snapshots are read and handled. A snapshot id is composed the same way a volume handle is, so a VolumeSnapshotContent written before the boundary carries a pool name in the same place — in spec.source.snapshotHandle for a pre-existing snapshot and spec.source.volumeHandle for a dynamically taken one, and whichever it has is the one normalized. The kind is registered in the tool's scheme as well as discovered, so a cluster with no snapshot CRDs is reported and skipped rather than failing the whole discovery, which is what an unregistered kind would have done. Resolution is PoolResolver, an interface on the Scope, and not lvol.Resolver: that one answers where a volume is and how to reach it, and has no pool listing in it. The implementation lists a cluster's pools once and answers from that, through the credentials that cluster keeps beside its own object. It finds the cluster by the UUID a StorageCluster reports rather than by what one is called, because an installation holds several and a pool called production exists in two of them — asking the wrong one returns a UUID that looks normalized and names a pool the volume is not in. The endpoint is --control-plane rather than something read from the cluster. The model being upgraded records it nowhere a tool running outside the cluster can reach: the operator takes it from its own environment, and a v1alpha1 ControlPlane carries no endpoint at all. Without the flag the step refuses and names it, which is the right outcome — a cluster whose volumes were all provisioned after the boundary has no pool name to resolve and never gets there. The resolution runs in Validate, which the framework runs in the preflight, so a pool name that answers to nothing is found before anything is written. That is §19.10's eighth check, which was blocked on exactly this. Not built, and §16.4 says so rather than this commit message: the promise that the operator's own records carrying a handle are rewritten too. That set is empty today, and the rule stands for whenever a field is added. Co-Authored-By: Claude Fable 5 --- atlas-lib/README.md | 2 + atlas-lib/kube/normalized.go | 56 ++++++ atlas-lib/kube/normalized_test.go | 105 ++++++++++ atlas-lib/lvol/normalized.go | 85 ++++++++ atlas-lib/lvol/normalized_test.go | 117 +++++++++++ operator/cmd/simplyblock-upgrade/client.go | 8 + operator/cmd/simplyblock-upgrade/root.go | 13 ++ operator/cmd/simplyblock-upgrade/run.go | 8 + .../crd-redesign/design-api-upgrade.md | 64 ++++-- .../persistentvolumeclaim_controller.go | 18 +- ...sistentvolumeclaim_controller_unit_test.go | 36 ++++ .../internal/upgrade/catalog/catalog_test.go | 5 + .../internal/upgrade/discover/kind_test.go | 24 +++ operator/internal/upgrade/discover/kinds.go | 12 ++ operator/internal/upgrade/pools.go | 32 +++ operator/internal/upgrade/pools/resolver.go | 129 ++++++++++++ .../internal/upgrade/pools/resolver_test.go | 185 ++++++++++++++++++ operator/internal/upgrade/scope.go | 16 ++ operator/internal/upgrade/steps/migrate.go | 168 +++++++++++++--- .../internal/upgrade/steps/migrate_test.go | 152 ++++++++++++++ .../internal/upgrade/steps/ownership_test.go | 7 + operator/internal/webhook/volumehandle.go | 17 +- 22 files changed, 1206 insertions(+), 53 deletions(-) create mode 100644 atlas-lib/kube/normalized.go create mode 100644 atlas-lib/kube/normalized_test.go create mode 100644 atlas-lib/lvol/normalized.go create mode 100644 atlas-lib/lvol/normalized_test.go create mode 100644 operator/internal/upgrade/pools.go create mode 100644 operator/internal/upgrade/pools/resolver.go create mode 100644 operator/internal/upgrade/pools/resolver_test.go diff --git a/atlas-lib/README.md b/atlas-lib/README.md index 014f28fda..c7b9d6518 100644 --- a/atlas-lib/README.md +++ b/atlas-lib/README.md @@ -92,6 +92,7 @@ atlas/ ├── lvol/ Logical-volume identity, control-plane + device resolution │ ├── volume.go VolumeHandle, Volume │ ├── handle.go Handle: a volume handle taken apart; ParseHandle, IsCanonicalUUID +│ ├── normalized.go NormalizeHandle: the annotated handle over the field's, §16.4's rule │ ├── groupsnapshot.go GroupSnapshotHandle: a consistency-group generation's CSI id │ ├── resolver.go Resolver: control-plane lookup (info + Connection) │ └── mapping.go Mapper: attached lvol → local nvme.Device @@ -99,6 +100,7 @@ atlas/ │ ├── names.go driver name, param/context/label/annotation/finalizer keys, pool label key │ ├── derived.go Formula: bounded, deterministic derived names and labels │ ├── identity.go VolumeHandle↔PV, VolumeContext, pin annotations +│ ├── normalized.go NormalizedHandle / NormalizedVolumeHandleFromPV: §16.4's rule on an object │ ├── binding.go Binding: resolved PV+PVC+Node view of an lvol │ ├── resolver.go Resolver iface + ResolveBinding aggregation │ ├── storageclass.go Properties: typed StorageClass provisioning params diff --git a/atlas-lib/kube/normalized.go b/atlas-lib/kube/normalized.go new file mode 100644 index 000000000..b0584f0c3 --- /dev/null +++ b/atlas-lib/kube/normalized.go @@ -0,0 +1,56 @@ +// Reading a volume's handle the way §16.4 says a reader should: the annotation +// when it is there and agrees, and the field otherwise. +// +// The rule itself is lvol.NormalizeHandle, which compares two handles and knows +// nothing about Kubernetes. What is here is the other half, which is where the +// two strings come from — and the reason the annotated half takes a map rather +// than an object is that both kinds carrying the annotation are shaped +// differently and only one of them is in the core API. A VolumeSnapshotContent +// belongs to the external snapshotter's module, and this module does not depend +// on it for a map lookup. + +package kube + +import ( + "fmt" + + corev1 "k8s.io/api/core/v1" + + "github.com/simplyblock/atlas/errs" + "github.com/simplyblock/atlas/lvol" +) + +// NormalizedHandle applies §16.4's rule to a handle and the annotations of the +// object carrying it. +// +// It is the entry point for a kind this module does not know: a caller holding +// a VolumeSnapshotContent passes its source handle and its annotations, and gets +// the same answer a PersistentVolume would. +func NormalizedHandle( + field lvol.VolumeHandle, annotations map[string]string, +) (lvol.Normalized, bool) { + return lvol.NormalizeHandle(field, lvol.VolumeHandle(annotations[AnnoVolumeHandle])) +} + +// NormalizedVolumeHandleFromPV reads a PersistentVolume's handle with the +// annotation preferred, and reports errs.ErrUnsupported for a volume this +// driver does not own. +// +// It stands beside VolumeHandleFromPV rather than replacing it, because the two +// answer different questions and both have callers. VolumeHandleFromPV asks +// what the object's spec says, which is what the upgrade needs in order to find +// the volumes whose spelling is legacy at all. This asks which pool the volume +// is in, which is what everything else needs. +func NormalizedVolumeHandleFromPV(pv *corev1.PersistentVolume) (lvol.Normalized, error) { + raw, err := VolumeHandleFromPV(pv) + if err != nil { + return lvol.Normalized{}, err + } + normalized, wellFormed := NormalizedHandle(raw, pv.GetAnnotations()) + if !wellFormed { + return lvol.Normalized{}, fmt.Errorf( + "pv %q carries the handle %q, which is not well formed: %w", + pvName(pv), raw, errs.ErrUnsupported) + } + return normalized, nil +} diff --git a/atlas-lib/kube/normalized_test.go b/atlas-lib/kube/normalized_test.go new file mode 100644 index 000000000..3f77a82d0 --- /dev/null +++ b/atlas-lib/kube/normalized_test.go @@ -0,0 +1,105 @@ +package kube + +import ( + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + + "github.com/simplyblock/atlas/lvol" +) + +const ( + testCluster = "8ffac363-0c46-4714-a71b-f9c0b58a1269" + testPoolUUID = "df34f16c-1a2b-3c4d-5e6f-7a8b9c0d1e2f" + testVolume = "a1111111-1111-4111-8111-111111111111" + + testLegacyHandle = testCluster + ":production:" + testVolume + testNormalizedHandle = testCluster + ":" + testPoolUUID + ":" + testVolume +) + +func legacyPV(annotations map[string]string) *corev1.PersistentVolume { + return &corev1.PersistentVolume{ + ObjectMeta: metav1.ObjectMeta{Name: "pvc-1", Annotations: annotations}, + Spec: corev1.PersistentVolumeSpec{ + PersistentVolumeSource: corev1.PersistentVolumeSource{ + CSI: &corev1.CSIPersistentVolumeSource{ + Driver: DriverName, + VolumeHandle: testLegacyHandle, + }, + }, + }, + } +} + +// A PersistentVolume that has been through the migration reports the pool it +// resolves to, and the same object before the migration reports the name its +// field carries. Both are correct answers to different questions, and this is +// the one that asks which pool the volume is actually in. +func TestNormalizedVolumeHandleFromPV(t *testing.T) { + before, err := NormalizedVolumeHandleFromPV(legacyPV(nil)) + if err != nil { + t.Fatalf("reading an unmigrated volume: %v", err) + } + if before.Handle.PoolRef != "production" || before.FromAnnotation { + t.Errorf("an unmigrated volume reported %+v, want the field's pool name", before) + } + + after, err := NormalizedVolumeHandleFromPV(legacyPV(map[string]string{ + AnnoVolumeHandle: testNormalizedHandle, + })) + if err != nil { + t.Fatalf("reading a migrated volume: %v", err) + } + if after.Handle.PoolRef != testPoolUUID || !after.FromAnnotation { + t.Errorf("a migrated volume reported %+v, want the annotated pool UUID", after) + } +} + +// An annotation that disagrees about anything but the pool is ignored, and the +// object keeps reporting what its own spec says. Whoever may edit metadata must +// not be able to point a volume at another cluster. +func TestNormalizedVolumeHandleFromPVIgnoresARedirectingAnnotation(t *testing.T) { + const elsewhere = "11111111-2222-4333-8444-555555555555:" + testPoolUUID + ":" + testVolume + + got, err := NormalizedVolumeHandleFromPV(legacyPV(map[string]string{ + AnnoVolumeHandle: elsewhere, + })) + if err != nil { + t.Fatalf("reading the volume: %v", err) + } + if got.Handle.ClusterID != testCluster { + t.Errorf("cluster = %s, want the field's %s: an annotation redirected the volume", + got.Handle.ClusterID, testCluster) + } + if got.Ignored == "" { + t.Error("the annotation was ignored and nothing says so, so a hand-edited " + + "annotation leaves no trace for anybody to find") + } +} + +// A volume another driver owns has no simplyblock handle to normalize, and +// saying so is the same answer VolumeHandleFromPV gives. +func TestNormalizedVolumeHandleFromPVRefusesAForeignVolume(t *testing.T) { + pv := legacyPV(nil) + pv.Spec.CSI.Driver = "ebs.csi.aws.com" + if _, err := NormalizedVolumeHandleFromPV(pv); err == nil { + t.Error("a volume owned by another driver was read as a simplyblock one") + } +} + +// The annotation-reading half takes a map rather than an object, so a +// VolumeSnapshotContent gets the rule without this module depending on the +// snapshot API. +func TestNormalizedHandleReadsAnyObjectsAnnotations(t *testing.T) { + got, ok := NormalizedHandle( + lvol.VolumeHandle(testLegacyHandle), + map[string]string{AnnoVolumeHandle: testNormalizedHandle}, + ) + if !ok { + t.Fatal("a well-formed handle was not read") + } + if got.Handle.PoolRef != testPoolUUID { + t.Errorf("pool = %q, want the annotated %q", got.Handle.PoolRef, testPoolUUID) + } +} diff --git a/atlas-lib/lvol/normalized.go b/atlas-lib/lvol/normalized.go new file mode 100644 index 000000000..55bdae13b --- /dev/null +++ b/atlas-lib/lvol/normalized.go @@ -0,0 +1,85 @@ +// Choosing between the handle a volume's spec carries and the normalized one an +// annotation records. +// +// A handle provisioned before the v2 API migration spells its pool as a name, +// and the field it lives in cannot be changed: the API server refuses any edit +// to a PersistentVolume's CSI source or a VolumeSnapshotContent's source. So the +// upgrade resolves the name once and writes the result to metadata, which is +// writable where spec is not, and a reader prefers what it finds there. +// +// The rule lives here, beside ParseHandle, because it is a statement about +// handles rather than about Kubernetes: the same two strings are compared +// wherever they come from, and a second implementation of the comparison would +// be a second opinion about which volume an object names. +// +// design-api-upgrade.md §16.4 is the specification. + +package lvol + +import ( + "fmt" + "strings" +) + +// Normalized is the handle a reader should use, and what was done to arrive at +// it. +type Normalized struct { + // Handle is the one to use. It is the field's handle with its pool segment + // replaced whenever the annotation was taken, and the field's unchanged + // otherwise. + Handle Handle + + // FromAnnotation records that the annotation supplied the handle. It is + // true for an annotation that agreed with the field exactly as well as for + // one that only normalized the pool, because both are the annotation being + // honored. + FromAnnotation bool + + // Ignored says why an annotation was not taken, and is empty when there was + // none or when it was. It is a sentence rather than a code because its only + // consumer is a report: a handle nobody can explain is a volume somebody + // has to go and look at. + Ignored string +} + +// NormalizeHandle chooses between the handle a field carries and the one an +// annotation records, reporting whether a handle could be read at all. +// +// The field is the authority on which cluster and which volume the object +// names, and the annotation may only differ from it in the pool. Anything else +// is refused and reported rather than honored: an annotation is metadata, so +// anybody who may edit an object's labels could otherwise redirect a volume to +// another cluster by writing one. +// +// An absent or blank annotation is no annotation at all rather than a wrong +// one, which is the ordinary state of every object provisioned after the +// boundary and of every cluster nobody has migrated. +// +// It returns false only for a field that carries no readable handle. The +// annotation cannot stand in for one, because it is a record about the field +// rather than a replacement for it: a volume the spec does not name is not one +// metadata may invent. +func NormalizeHandle(field, annotated VolumeHandle) (Normalized, bool) { + on, wellFormed := ParseHandle(field) + if !wellFormed { + return Normalized{}, false + } + + claimed, wellFormed := ParseHandle(annotated) + switch { + case strings.TrimSpace(string(annotated)) == "": + return Normalized{Handle: on}, true + case !wellFormed: + return Normalized{Handle: on, Ignored: fmt.Sprintf( + "the annotated handle %q is not well formed", annotated)}, true + case claimed.ClusterID != on.ClusterID: + return Normalized{Handle: on, Ignored: fmt.Sprintf( + "the annotated handle names cluster %s and the volume is in %s", + claimed.ClusterID, on.ClusterID)}, true + case claimed.VolumeID != on.VolumeID: + return Normalized{Handle: on, Ignored: fmt.Sprintf( + "the annotated handle names volume %s and this volume is %s", + claimed.VolumeID, on.VolumeID)}, true + } + return Normalized{Handle: claimed, FromAnnotation: true}, true +} diff --git a/atlas-lib/lvol/normalized_test.go b/atlas-lib/lvol/normalized_test.go new file mode 100644 index 000000000..07c4c0c8d --- /dev/null +++ b/atlas-lib/lvol/normalized_test.go @@ -0,0 +1,117 @@ +package lvol + +import ( + "strings" + "testing" +) + +// The rule of §16.4: an annotation is taken when it agrees with the field about +// everything but the pool, and ignored otherwise. +// +// The cases that matter are the disagreements. A handle names a cluster and a +// volume as well as a pool, and an annotation is metadata anybody with edit +// rights can write, so an annotation that redirects either of those is the one +// thing this must not honor. +func TestNormalizeHandle(t *testing.T) { + const ( + cluster = "8ffac363-0c46-4714-a71b-f9c0b58a1269" + otherCluster = "11111111-2222-4333-8444-555555555555" + poolUUID = "df34f16c-1a2b-3c4d-5e6f-7a8b9c0d1e2f" + volume = "a1111111-1111-4111-8111-111111111111" + otherVolume = "b2222222-2222-4222-8222-222222222222" + ) + + legacy := VolumeHandle(cluster + ":production:" + volume) + normalized := VolumeHandle(cluster + ":" + poolUUID + ":" + volume) + + for _, tc := range []struct { + name string + field VolumeHandle + annotated VolumeHandle + wantPool string + wantFrom bool + wantWhy string + }{ + { + name: "no annotation leaves the field as it is", + field: legacy, + wantPool: "production", + }, + { + name: "an annotation differing only in the pool is taken", + field: legacy, + annotated: normalized, + wantPool: poolUUID, + wantFrom: true, + }, + { + name: "an annotation equal to the field is taken and changes nothing", + field: normalized, + annotated: normalized, + wantPool: poolUUID, + wantFrom: true, + }, + { + name: "an annotation naming another cluster is ignored", + field: legacy, + annotated: VolumeHandle(otherCluster + ":" + poolUUID + ":" + volume), + wantPool: "production", + wantWhy: "cluster", + }, + { + name: "an annotation naming another volume is ignored", + field: legacy, + annotated: VolumeHandle(cluster + ":" + poolUUID + ":" + otherVolume), + wantPool: "production", + wantWhy: "volume", + }, + { + name: "an annotation that is not a handle is ignored", + field: legacy, + annotated: "not-a-handle", + wantPool: "production", + wantWhy: "well formed", + }, + { + name: "an empty annotation is absent rather than wrong", + field: legacy, + annotated: " ", + wantPool: "production", + }, + } { + t.Run(tc.name, func(t *testing.T) { + got, ok := NormalizeHandle(tc.field, tc.annotated) + if !ok { + t.Fatalf("NormalizeHandle(%q, %q) = not ok, want ok", tc.field, tc.annotated) + } + if got.Handle.PoolRef != tc.wantPool { + t.Errorf("pool = %q, want %q", got.Handle.PoolRef, tc.wantPool) + } + if got.FromAnnotation != tc.wantFrom { + t.Errorf("FromAnnotation = %v, want %v", got.FromAnnotation, tc.wantFrom) + } + if tc.wantWhy == "" && got.Ignored != "" { + t.Errorf("Ignored = %q, want none", got.Ignored) + } + if tc.wantWhy != "" && !strings.Contains(got.Ignored, tc.wantWhy) { + t.Errorf("Ignored = %q, want it to mention %q", got.Ignored, tc.wantWhy) + } + }) + } +} + +// A field that is not a handle leaves nothing to normalize, whatever the +// annotation says. The annotation is a record about the field, so it cannot +// stand in for one that is missing. +func TestNormalizeHandleRefusesAnUnreadableField(t *testing.T) { + const good = "8ffac363-0c46-4714-a71b-f9c0b58a1269:" + + "df34f16c-1a2b-3c4d-5e6f-7a8b9c0d1e2f:a1111111-1111-4111-8111-111111111111" + + if _, ok := NormalizeHandle("", good); ok { + t.Error("an empty field was normalized from an annotation, which lets metadata " + + "invent a volume the spec does not name") + } + if _, ok := NormalizeHandle("nonsense", good); ok { + t.Error("an unparsable field was normalized from an annotation") + } +} diff --git a/operator/cmd/simplyblock-upgrade/client.go b/operator/cmd/simplyblock-upgrade/client.go index 665b3e249..e94dc1281 100644 --- a/operator/cmd/simplyblock-upgrade/client.go +++ b/operator/cmd/simplyblock-upgrade/client.go @@ -8,6 +8,7 @@ package main import ( "fmt" + snapshotv1 "github.com/kubernetes-csi/external-snapshotter/client/v8/apis/volumesnapshot/v1" apiextensionsv1 "k8s.io/apiextensions-apiserver/pkg/apis/apiextensions/v1" "k8s.io/apimachinery/pkg/runtime" utilruntime "k8s.io/apimachinery/pkg/util/runtime" @@ -25,10 +26,17 @@ import ( // CustomResourceDefinition is in it because the upgrade applies CRDs, waits for // them to be established, and later switches their storage version, and none of // that is reachable through the typed clients for the group being upgraded. +// +// The snapshot API is in it because §16.4 normalizes the handles a +// VolumeSnapshotContent carries as well as a PersistentVolume's. Registering a +// kind the cluster may not serve costs nothing and is what makes the absence +// legible: a kind that is registered and not served is reported and skipped, +// where one the scheme does not know fails the whole discovery. func newScheme() *runtime.Scheme { scheme := runtime.NewScheme() utilruntime.Must(clientgoscheme.AddToScheme(scheme)) utilruntime.Must(apiextensionsv1.AddToScheme(scheme)) + utilruntime.Must(snapshotv1.AddToScheme(scheme)) utilruntime.Must(simplyblockv1alpha1.AddToScheme(scheme)) utilruntime.Must(simplyblockv1alpha2.AddToScheme(scheme)) return scheme diff --git a/operator/cmd/simplyblock-upgrade/root.go b/operator/cmd/simplyblock-upgrade/root.go index a1fefbf4f..7aa4c8dc4 100644 --- a/operator/cmd/simplyblock-upgrade/root.go +++ b/operator/cmd/simplyblock-upgrade/root.go @@ -30,6 +30,17 @@ type globalOptions struct { // Namespace is the installation being operated on. Namespace string + // ControlPlane is the management API's base URL, which §16.4's handle + // normalization needs and nothing else does. + // + // It is a flag rather than something read from the cluster, because the + // model being upgraded records it nowhere this tool can reach it: the + // operator takes it from its own environment, and a v1alpha1 ControlPlane + // carries no endpoint. Left empty, the steps that need it refuse and say so, + // which is the right outcome for a cluster whose volumes were all + // provisioned after the v2 API migration and have no pool name to resolve. + ControlPlane string + // Skip names rules that are not to run. Every skipped rule is reported, so // the decision stays visible in the run's own output. Skip []string @@ -88,6 +99,8 @@ func newRootCommand() *cobra.Command { "path to a kubeconfig file, defaulting to the in-cluster configuration") flags.StringVar(&global.Context, "context", "", "the kubeconfig context to use") flags.StringVarP(&global.Namespace, "namespace", "n", defaultNamespace(), "the namespace the installation lives in") + flags.StringVar(&global.ControlPlane, "control-plane", "", + "base URL of the control-plane management API, needed to resolve the pool names in legacy volume handles") flags.StringSliceVar(&global.Skip, "skip", nil, "rule identities not to run; every skipped rule is reported") flags.BoolVar(&global.AcknowledgeOffline, "acknowledge-offline", false, "proceed even though a StorageNode is not online") diff --git a/operator/cmd/simplyblock-upgrade/run.go b/operator/cmd/simplyblock-upgrade/run.go index ab7dc86da..efcb3b9f6 100644 --- a/operator/cmd/simplyblock-upgrade/run.go +++ b/operator/cmd/simplyblock-upgrade/run.go @@ -15,6 +15,7 @@ import ( "github.com/simplyblock/simplyblock-operator/internal/upgrade" "github.com/simplyblock/simplyblock-operator/internal/upgrade/catalog" + "github.com/simplyblock/simplyblock-operator/internal/upgrade/pools" "github.com/simplyblock/simplyblock-operator/internal/upgrade/tui" ) @@ -51,6 +52,13 @@ func newSession(global *globalOptions, stage upgrade.Stage, dryRun bool) (*sessi reporter := newReporter(global) scope := upgrade.NewScope(c, global.Namespace, stage, global.options(dryRun), newLogger(global), reporter) + // The resolver is attached only where there is somewhere to resolve + // against. Left off, the handle steps refuse and name the flag, which is a + // better answer than a client pointed at nothing and timing out per volume. + if global.ControlPlane != "" { + scope.Pools = &pools.Resolver{Client: c, Endpoint: global.ControlPlane} + } + return &session{ Runner: upgrade.NewRunner(catalog.Default(), scope), Reporter: reporter, diff --git a/operator/docs/designs/crd-redesign/design-api-upgrade.md b/operator/docs/designs/crd-redesign/design-api-upgrade.md index 5babc3d8a..fa37d74ef 100644 --- a/operator/docs/designs/crd-redesign/design-api-upgrade.md +++ b/operator/docs/designs/crd-redesign/design-api-upgrade.md @@ -1118,11 +1118,29 @@ the phase writes the rows below that line instead. **Resolve and report.** Every `PersistentVolume` and `VolumeSnapshotContent` whose pool segment is not a canonical UUID is listed with the name it carries and -the UUID that name resolves to, through `lvol.Resolver` against the control -plane. A pool name resolving to nothing is a finding: the handle names a pool -that no longer exists, and the migration reports it and does not proceed. This -part is a read, so it belongs to `preflight` (§19.10) and runs long before the -migration does. +the UUID that name resolves to. A pool name resolving to nothing is a finding: +the handle names a pool that no longer exists, and the migration reports it and +does not proceed. This part is a read, so it belongs to `preflight` (§19.10) and +runs long before the migration does — it is the step's own `Validate`, which the +framework runs in the preflight and again before applying. + +The lookup is `PoolResolver`, an interface on the run's `Scope`, and not +`lvol.Resolver`: that interface answers where a volume is and how to reach it, +and has no pool listing in it. The implementation lists a cluster's pools once +and answers from that, through the credentials the cluster keeps beside its own +object — an installation holds several clusters, each with its own secret, and a +pool called `production` exists in two of them, so asking the wrong one returns a +UUID that looks normalized and names a pool the volume is not in. The cluster a +handle names is found by the UUID a `StorageCluster` reports rather than by what +one is called. + +**The endpoint is given rather than discovered.** The model being upgraded +records it nowhere a tool running outside the cluster can read: the operator +takes it from its own environment, and a `v1alpha1` `ControlPlane` carries no +endpoint at all. So it is `--control-plane`, and without it the step refuses and +names the flag. That refusal is the right outcome rather than a gap — a cluster +whose volumes were all provisioned after the boundary has no pool name to +resolve and never reaches it. **Rewrite the records the migration owns.** Anywhere the operator has written a handle into a field it controls, a custom resource's status or a `ConfigMap`, the @@ -1151,8 +1169,27 @@ otherwise.** Consistency is exact: the annotation's cluster and volume segments MUST equal the field's, and only the pool segment may differ. A reader finding any other difference ignores the annotation and reports it, so a hand-edited annotation cannot redirect a volume to another cluster. That rule is one -function in `atlas-lib`, beside `ParseHandle`, and no call site implements it -twice. +function in `atlas-lib`, `lvol.NormalizeHandle`, beside `ParseHandle`, and no +call site implements it twice. + +It compares two handles and knows nothing about Kubernetes, which is what lets +both kinds share it. Where the two strings come from is the other half, and it +is `kube.NormalizedHandle`, which takes a handle and an object's annotations: +`VolumeSnapshotContent` belongs to the external snapshotter's module, and +`atlas-lib` does not take a dependency on it for a map lookup. +`kube.NormalizedVolumeHandleFromPV` is the same thing for the kind that is in +the core API. + +**`kube.VolumeHandleFromPV` stays, and stays the field's own answer.** The two +questions differ the way `ParseHandle` and `Split` differ: which pool a volume +is in is what a caller reaching for the control plane wants, and what an +object's spec literally says is what the migration wants, since that is how it +finds the volumes whose spelling is legacy at all. + +A `VolumeSnapshotContent` names one of two things and both are the same three +segments, so whichever it carries is normalized: a pre-existing snapshot names +itself in `spec.source.snapshotHandle`, and a dynamically taken one names the +volume it came from in `spec.source.volumeHandle`. **This is what makes the normalized pool reachable without the control plane.** The CSI driver already indexes `PersistentVolume` objects by the lvol id in @@ -2484,16 +2521,17 @@ prose, its check is here and not repeated in both places. **Volume handles (§16.4)** -- [ ] Every legacy handle is reported with the UUID its pool name resolves to, +- [x] Every legacy handle is reported with the UUID its pool name resolves to, and an unresolvable one fails the preflight. -- [ ] A `PersistentVolume` is replaced only under `Retain`, one at a time, and - never while a pod has the claim mounted. -- [ ] The normalized handle is written to +- [x] No `PersistentVolume` is replaced at all. The field keeps the spelling it + was provisioned with and the annotation carries the identity, so the + replacement this row guarded against does not arise. +- [x] The normalized handle is written to `storage.simplyblock.io/volume-handle` on every `PersistentVolume` and `VolumeSnapshotContent` whose field carries a pool name. -- [ ] One `atlas-lib` function decides between the annotation and the field, and +- [x] One `atlas-lib` function decides between the annotation and the field, and it rejects an annotation whose cluster or volume segment differs. -- [ ] `lvol.ParseHandle` stays tolerant for objects with no annotation. +- [x] `lvol.ParseHandle` stays tolerant for objects with no annotation. **Names (§19)** diff --git a/operator/internal/controller/persistentvolumeclaim_controller.go b/operator/internal/controller/persistentvolumeclaim_controller.go index 88be0871e..ca1c4ffc4 100644 --- a/operator/internal/controller/persistentvolumeclaim_controller.go +++ b/operator/internal/controller/persistentvolumeclaim_controller.go @@ -20,7 +20,6 @@ import ( "sigs.k8s.io/controller-runtime/pkg/predicate" "github.com/simplyblock/atlas/kube" - atlaslvol "github.com/simplyblock/atlas/lvol" simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" "github.com/simplyblock/simplyblock-operator/internal/volumemigration" @@ -57,7 +56,7 @@ const ( // pinned-volume value differs from the AnnotationPinnedVolumeApplied marker it // writes after acting, so its own annotation writes do not re-trigger a // migration. A validating admission webhook rejects an unknown storage node at -// write time; the re-validation here is a defense-in-depth backstop (e.g. for +// write time; the re-validation here is a defense-in-depth backstop (e.g., for // values that predate the webhook, or a node removed after the pin was set). type PersistentVolumeClaimReconciler struct { client.Client @@ -160,7 +159,7 @@ func (r *PersistentVolumeClaimReconciler) Reconcile( } // Serialize per PV: wait for any in-flight pin migration to finish before - // requesting another (e.g. when the target changed while one was running). + // requesting another (e.g., when the target changed while one was running). active, err := r.hasActiveMigration(ctx, cluster.Namespace, pv.Name) if err != nil { return ctrl.Result{}, err @@ -308,15 +307,18 @@ func (r *PersistentVolumeClaimReconciler) setApplied( // handle through the atlas helpers, so the handle grammar lives in one place. // ok is false when the PV is not a simplyblock CSI volume or the handle is // malformed. +// +// The normalized reader rather than the field's own, because a volume +// provisioned before the v2 API migration spells its pool as a name in a field +// the API server refuses to change, and the pool this returns is handed to the +// control plane (design-api-upgrade.md §16.4). A volume the migration has not +// reached, or a cluster nobody has migrated, still reports the name. func csiVolumeHandleParts(pv *corev1.PersistentVolume) (clusterUUID, poolRef, volumeUUID string, ok bool) { - raw, err := kube.VolumeHandleFromPV(pv) + normalized, err := kube.NormalizedVolumeHandleFromPV(pv) if err != nil { return "", "", "", false } - h, parsed := atlaslvol.ParseHandle(raw) - if !parsed { - return "", "", "", false - } + h := normalized.Handle return h.ClusterID, h.PoolRef, h.VolumeID, true } diff --git a/operator/internal/controller/persistentvolumeclaim_controller_unit_test.go b/operator/internal/controller/persistentvolumeclaim_controller_unit_test.go index a33cd0c70..838d2fd9a 100644 --- a/operator/internal/controller/persistentvolumeclaim_controller_unit_test.go +++ b/operator/internal/controller/persistentvolumeclaim_controller_unit_test.go @@ -352,3 +352,39 @@ func TestPVCReconcile_ActiveMigrationWaits(t *testing.T) { t.Fatalf("applied must not be set while waiting for an in-flight migration") } } + +// TestCSIVolumeHandlePartsPrefersTheNormalizedAnnotation is the read half of +// §16.4. A volume provisioned before the v2 API migration spells its pool as a +// name in a field nothing can rewrite, so the upgrade resolves it once into an +// annotation and every reader takes that instead. +func TestCSIVolumeHandlePartsPrefersTheNormalizedAnnotation(t *testing.T) { + const ( + cluster = "2f4f0300-9993-4289-be95-59414fc8a54d" + poolUUID = "1c2c0300-9993-4289-be95-59414fc8a54d" + volume = "8b1f0300-9993-4289-be95-59414fc8a54d" + ) + + pv := &corev1.PersistentVolume{ + ObjectMeta: metav1.ObjectMeta{ + Name: "pvc-legacy", + Annotations: map[string]string{ + kube.AnnoVolumeHandle: cluster + ":" + poolUUID + ":" + volume, + }, + }, + Spec: corev1.PersistentVolumeSpec{PersistentVolumeSource: corev1.PersistentVolumeSource{ + CSI: &corev1.CSIPersistentVolumeSource{ + Driver: kube.DriverName, + VolumeHandle: cluster + ":production:" + volume, + }}}, + } + + _, poolRef, _, ok := csiVolumeHandleParts(pv) + if !ok { + t.Fatal("the handle was not read at all") + } + if poolRef != poolUUID { + t.Errorf("pool = %q, want the normalized %q: the reader took the field's name, so "+ + "every control-plane call made with it names a pool by a spelling the v2 API "+ + "does not accept", poolRef, poolUUID) + } +} diff --git a/operator/internal/upgrade/catalog/catalog_test.go b/operator/internal/upgrade/catalog/catalog_test.go index 548563796..d5ea48816 100644 --- a/operator/internal/upgrade/catalog/catalog_test.go +++ b/operator/internal/upgrade/catalog/catalog_test.go @@ -13,6 +13,7 @@ import ( "strings" "testing" + snapshotv1 "github.com/kubernetes-csi/external-snapshotter/client/v8/apis/volumesnapshot/v1" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" "k8s.io/apimachinery/pkg/runtime" clientgoscheme "k8s.io/client-go/kubernetes/scheme" @@ -94,6 +95,10 @@ func preflight(t *testing.T, opts upgrade.Options, objects ...client.Object) upg if err := simplyblockv1alpha1.AddToScheme(scheme); err != nil { t.Fatalf("registering v1alpha1: %v", err) } + // §16.4 reads VolumeSnapshotContent, and the shipped catalog discovers it. + if err := snapshotv1.AddToScheme(scheme); err != nil { + t.Fatalf("registering the snapshot API: %v", err) + } // The preflight promises it changes nothing, so the run is given a client // that fails the test rather than a cluster that quietly changed. diff --git a/operator/internal/upgrade/discover/kind_test.go b/operator/internal/upgrade/discover/kind_test.go index dd824fffc..a4784cab1 100644 --- a/operator/internal/upgrade/discover/kind_test.go +++ b/operator/internal/upgrade/discover/kind_test.go @@ -13,6 +13,7 @@ import ( "errors" "testing" + snapshotv1 "github.com/kubernetes-csi/external-snapshotter/client/v8/apis/volumesnapshot/v1" corev1 "k8s.io/api/core/v1" "k8s.io/apimachinery/pkg/api/meta" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" @@ -334,3 +335,26 @@ func TestCatalog_EveryOwnedKindWaitsForTheCustomResources(t *testing.T) { } } } + +// TestSnapshotContentsAreDiscovered pins that §16.4's second kind is read. +// +// It is the half a step cannot supply for itself: normalize-volume-handles +// answers correctly about a VolumeSnapshotContent, and would never be handed +// one if the graph did not hold it. A kind that is handled and not discovered +// looks exactly like a cluster with no snapshots in it. +func TestSnapshotContentsAreDiscovered(t *testing.T) { + var found bool + for _, discoverer := range CoreKinds() { + kind, ok := discoverer.(Kind) + if !ok { + continue + } + if _, isContent := kind.List.(*snapshotv1.VolumeSnapshotContentList); isContent { + found = true + } + } + if !found { + t.Error("no discoverer reads VolumeSnapshotContent, so §16.4's handles on snapshots " + + "are normalized by a step that is never handed one") + } +} diff --git a/operator/internal/upgrade/discover/kinds.go b/operator/internal/upgrade/discover/kinds.go index ab7fab7e9..6a42e57a1 100644 --- a/operator/internal/upgrade/discover/kinds.go +++ b/operator/internal/upgrade/discover/kinds.go @@ -8,6 +8,7 @@ package discover import ( + snapshotv1 "github.com/kubernetes-csi/external-snapshotter/client/v8/apis/volumesnapshot/v1" appsv1 "k8s.io/api/apps/v1" corev1 "k8s.io/api/core/v1" discoveryv1 "k8s.io/api/discovery/v1" @@ -38,6 +39,12 @@ const ( IDNamespaces upgrade.ID = "discover-namespaces" IDPersistentVolumes upgrade.ID = "discover-persistent-volumes" + // The snapshots, whose id the CSI controller composes the same way a volume + // handle is composed, so one written before the v2 API migration carries a + // pool name in the same place. A cluster with no snapshot CRDs installed + // serves the kind not at all, which Discover reports and skips. + IDSnapshotContents upgrade.ID = "discover-volume-snapshot-contents" + // The workload a StorageNodeSet owns, which §20 reparents onto the cluster. IDDaemonSets upgrade.ID = "discover-daemon-sets" IDServices upgrade.ID = "discover-services" @@ -229,6 +236,11 @@ func CoreKinds() []upgrade.Discoverer { Summary: "reads the PersistentVolume objects whose volume handles §16.4 normalizes", List: &corev1.PersistentVolumeList{}, }, + Kind{ + RuleID: IDSnapshotContents, + Summary: "reads the VolumeSnapshotContent objects whose snapshot handles §16.4 normalizes too", + List: &snapshotv1.VolumeSnapshotContentList{}, + }, } } diff --git a/operator/internal/upgrade/pools.go b/operator/internal/upgrade/pools.go new file mode 100644 index 000000000..12989cc6f --- /dev/null +++ b/operator/internal/upgrade/pools.go @@ -0,0 +1,32 @@ +// The one question this tool asks of something that is not Kubernetes. +// +// §16.4's handles name their pool by name where they were provisioned before +// the v2 API migration, and the UUID that name stands for is held by the +// control plane and nowhere else: no PersistentVolume records it, and the +// StoragePool custom resource carries the name a user chose rather than the +// identifier the backend assigned. +// +// It is an interface rather than a client because of where it is called from. +// Every other rule in this framework reads the cluster and nothing else, which +// is what lets a test substitute a fake client for the whole of it; a rule +// reaching for a control-plane client directly would be a rule that cannot be +// tested without one. + +package upgrade + +import "context" + +// PoolResolver turns a pool's name into its UUID, within one cluster. +// +// The cluster is a parameter rather than a property of the implementation +// because a handle names its own: one installation holds several clusters, and +// two of them may each have a pool called production. +type PoolResolver interface { + // PoolUUID returns the UUID of the named pool. + // + // It wraps errs.ErrNotFound for a name no pool answers to, which is a + // finding rather than a failure: a handle naming a pool that no longer + // exists is something a person has to look at, and §16.4 says the migration + // reports it and does not proceed. + PoolUUID(ctx context.Context, clusterUUID, poolName string) (string, error) +} diff --git a/operator/internal/upgrade/pools/resolver.go b/operator/internal/upgrade/pools/resolver.go new file mode 100644 index 000000000..c804edabc --- /dev/null +++ b/operator/internal/upgrade/pools/resolver.go @@ -0,0 +1,129 @@ +// Resolving a pool's name to its UUID, which is what §16.4's handles need and +// what only the control plane knows. +// +// It is a package of its own rather than a function in the steps, because it is +// the one place in this tool that talks to something other than Kubernetes, and +// the credentials it uses are per cluster. An installation holds several, each +// with its own secret, and asking the wrong one is worse than asking nothing: a +// pool called production exists in two clusters and the answer would look +// normalized while naming a pool the volume is not in. + +package pools + +import ( + "context" + "fmt" + "sync" + + corev1 "k8s.io/api/core/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + + "github.com/simplyblock/atlas/controlplane" + "github.com/simplyblock/atlas/errs" + + simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" +) + +// Resolver answers a pool name from the control plane, using the credentials +// the cluster that holds the pool keeps beside its own object. +type Resolver struct { + // Client reads the StorageCluster that reports the UUID a handle names, and + // the Secret holding its token. + Client client.Client + + // Endpoint is the control-plane management API. It is supplied rather than + // discovered, because the source model records it nowhere a tool running + // outside the cluster can read: the operator takes it from its own + // environment, and a v1alpha1 ControlPlane carries no endpoint at all. + Endpoint string + + // mu guards pools, which is one listing per cluster. A run asks about one + // pool per volume and a cluster has thousands of volumes and a handful of + // pools, so the listing is the same answer every time. + mu sync.Mutex + pools map[string]map[string]string +} + +// PoolUUID returns the UUID of the named pool in the named cluster. +// +// Both halves of a miss wrap errs.ErrNotFound, and both are findings rather +// than failures: a cluster no object reports is a handle naming an installation +// this is not, and a pool nobody answers to is a handle naming a pool that has +// been deleted. §16.4 reports either and does not proceed. +func (r *Resolver) PoolUUID(ctx context.Context, clusterUUID, poolName string) (string, error) { + byName, err := r.listing(ctx, clusterUUID) + if err != nil { + return "", err + } + uuid, known := byName[poolName] + if !known { + return "", fmt.Errorf("no pool in cluster %s is called %q: %w", + clusterUUID, poolName, errs.ErrNotFound) + } + return uuid, nil +} + +// listing is one cluster's pools by name, read once. +func (r *Resolver) listing(ctx context.Context, clusterUUID string) (map[string]string, error) { + r.mu.Lock() + defer r.mu.Unlock() + + if byName, cached := r.pools[clusterUUID]; cached { + return byName, nil + } + + token, err := r.token(ctx, clusterUUID) + if err != nil { + return nil, err + } + api, err := controlplane.New(controlplane.Config{Endpoint: r.Endpoint, Token: token}) + if err != nil { + return nil, fmt.Errorf("reach the control plane at %s: %w", r.Endpoint, err) + } + found, err := api.ListStoragePools(ctx, clusterUUID) + if err != nil { + return nil, fmt.Errorf("list the pools of cluster %s: %w", clusterUUID, err) + } + + byName := make(map[string]string, len(found)) + for _, pool := range found { + byName[pool.Name] = pool.ID + } + if r.pools == nil { + r.pools = map[string]map[string]string{} + } + r.pools[clusterUUID] = byName + return byName, nil +} + +// token is the control-plane secret of the cluster reporting this UUID. +// +// The cluster is found by what it reports rather than by what it is called, +// because a handle carries the backend's identifier and two namespaces may each +// hold a StorageCluster called prod. +func (r *Resolver) token(ctx context.Context, clusterUUID string) (string, error) { + var clusters simplyblockv1alpha1.StorageClusterList + if err := r.Client.List(ctx, &clusters); err != nil { + return "", fmt.Errorf("list the StorageClusters: %w", err) + } + + for i := range clusters.Items { + cluster := &clusters.Items[i] + if cluster.Status.UUID != clusterUUID { + continue + } + var secret corev1.Secret + key := client.ObjectKey{ + Namespace: cluster.Namespace, + Name: "simplyblock-cluster-" + cluster.Name, + } + if err := r.Client.Get(ctx, key, &secret); err != nil { + return "", fmt.Errorf("read the credentials of cluster %s/%s: %w", + cluster.Namespace, cluster.Name, err) + } + return string(secret.Data["secret"]), nil + } + + return "", fmt.Errorf("no StorageCluster reports the cluster %s a handle names: %w", + clusterUUID, errs.ErrNotFound) +} diff --git a/operator/internal/upgrade/pools/resolver_test.go b/operator/internal/upgrade/pools/resolver_test.go new file mode 100644 index 000000000..309efca45 --- /dev/null +++ b/operator/internal/upgrade/pools/resolver_test.go @@ -0,0 +1,185 @@ +// What the resolver has to get right is which credentials it uses, since an +// installation holds several clusters and each has its own. A resolver that +// asked the wrong cluster would answer with a pool UUID from somewhere else, +// which is worse than answering nothing: the handle would look normalized and +// name a pool the volume is not in. + +package pools + +import ( + "encoding/json" + "errors" + "net/http" + "net/http/httptest" + "strings" + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + "github.com/simplyblock/atlas/errs" + + simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" +) + +const ( + clusterUUID = "2f4f0300-9993-4289-be95-59414fc8a54d" + poolUUID = "1c2c0300-9993-4289-be95-59414fc8a54d" + otherUUID = "3d3d0300-9993-4289-be95-59414fc8a54d" + + // clusterName is what every fixture calls its StorageCluster, in whichever + // namespace: a handle names its cluster by UUID, so the name is exactly the + // thing that must not decide which credentials are used. + clusterName = "prod" +) + +// pools is a control plane answering one cluster's pool list, and recording the +// bearer token it was asked with. +func poolServer(t *testing.T, named string) (*httptest.Server, *string) { + t.Helper() + + var seen string + srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + seen = r.Header.Get("Authorization") + w.Header().Set("Content-Type", "application/json") + body := []map[string]any{} + if named != "" { + body = append(body, map[string]any{ + "id": poolUUID, "cluster_id": clusterUUID, "name": named, "max_size": 0, + }) + } + _ = json.NewEncoder(w).Encode(body) + })) + t.Cleanup(srv.Close) + return srv, &seen +} + +func world(t *testing.T, objects ...client.Object) client.Client { + t.Helper() + + scheme := runtime.NewScheme() + if err := corev1.AddToScheme(scheme); err != nil { + t.Fatalf("registering core: %v", err) + } + if err := simplyblockv1alpha1.AddToScheme(scheme); err != nil { + t.Fatalf("registering v1alpha1: %v", err) + } + return fake.NewClientBuilder().WithScheme(scheme).WithObjects(objects...).Build() +} + +// storageCluster is one cluster. Both tenants call theirs prod, because the +// point of every case here is that the name is not what a handle names. +func storageCluster(namespace, uuid string) *simplyblockv1alpha1.StorageCluster { + return &simplyblockv1alpha1.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{Namespace: namespace, Name: clusterName}, + Status: simplyblockv1alpha1.StorageClusterStatus{UUID: uuid}, + } +} + +func clusterSecret(namespace, token string) *corev1.Secret { + return &corev1.Secret{ + ObjectMeta: metav1.ObjectMeta{Namespace: namespace, Name: "simplyblock-cluster-" + clusterName}, + Data: map[string][]byte{"secret": []byte(token)}, + } +} + +func TestPoolUUIDResolvesThroughTheNamedClustersOwnCredentials(t *testing.T) { + srv, seen := poolServer(t, "production") + resolver := &Resolver{ + Client: world(t, storageCluster("tenant-a", clusterUUID), clusterSecret("tenant-a", "tok-a")), + Endpoint: srv.URL, + } + + got, err := resolver.PoolUUID(t.Context(), clusterUUID, "production") + if err != nil { + t.Fatalf("resolving: %v", err) + } + if got != poolUUID { + t.Errorf("PoolUUID = %q, want %q", got, poolUUID) + } + if *seen != "Bearer tok-a" { + t.Errorf("asked with %q, want the named cluster's own secret", *seen) + } +} + +// A second cluster in a second namespace must not supply the credentials, and +// the handle names its cluster by UUID rather than by name, so that is what +// decides. +func TestPoolUUIDIgnoresAClusterTheHandleDoesNotName(t *testing.T) { + srv, seen := poolServer(t, "production") + resolver := &Resolver{ + Client: world(t, + storageCluster("tenant-a", otherUUID), clusterSecret("tenant-a", "tok-a"), + storageCluster("tenant-b", clusterUUID), clusterSecret("tenant-b", "tok-b"), + ), + Endpoint: srv.URL, + } + + if _, err := resolver.PoolUUID(t.Context(), clusterUUID, "production"); err != nil { + t.Fatalf("resolving: %v", err) + } + if *seen != "Bearer tok-b" { + t.Errorf("asked with %q, want the secret of the cluster the handle names", *seen) + } +} + +// A pool nobody answers to is a finding rather than a failure, and the caller +// tells the two apart by the wrapped error. +func TestPoolUUIDReportsANameNoPoolAnswersTo(t *testing.T) { + srv, _ := poolServer(t, "") + resolver := &Resolver{ + Client: world(t, storageCluster("tenant-a", clusterUUID), clusterSecret("tenant-a", "tok-a")), + Endpoint: srv.URL, + } + + _, err := resolver.PoolUUID(t.Context(), clusterUUID, "vanished") + if !errors.Is(err, errs.ErrNotFound) { + t.Fatalf("err = %v, want it to wrap ErrNotFound", err) + } +} + +// A cluster no object reports is its own finding, and it is a different one: the +// handle names a cluster this installation does not hold, so there are no +// credentials to ask with rather than no pool to find. +func TestPoolUUIDReportsAClusterNoObjectReports(t *testing.T) { + srv, _ := poolServer(t, "production") + resolver := &Resolver{Client: world(t), Endpoint: srv.URL} + + _, err := resolver.PoolUUID(t.Context(), clusterUUID, "production") + if !errors.Is(err, errs.ErrNotFound) { + t.Fatalf("err = %v, want it to wrap ErrNotFound", err) + } + if err == nil || !strings.Contains(err.Error(), clusterUUID) { + t.Errorf("the error does not name the cluster nobody reports: %v", err) + } +} + +// One answer per cluster. A cluster of a thousand volumes would otherwise list +// every pool a thousand times, and the list is the same list each time. +func TestPoolUUIDListsEachClusterOnce(t *testing.T) { + var calls int + srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) { + calls++ + w.Header().Set("Content-Type", "application/json") + _ = json.NewEncoder(w).Encode([]map[string]any{ + {"id": poolUUID, "cluster_id": clusterUUID, "name": "production", "max_size": 0}, + }) + })) + t.Cleanup(srv.Close) + + resolver := &Resolver{ + Client: world(t, storageCluster("tenant-a", clusterUUID), clusterSecret("tenant-a", "tok-a")), + Endpoint: srv.URL, + } + for range 3 { + if _, err := resolver.PoolUUID(t.Context(), clusterUUID, "production"); err != nil { + t.Fatalf("resolving: %v", err) + } + } + if calls != 1 { + t.Errorf("listed the pools %d times, want once for the cluster", calls) + } +} diff --git a/operator/internal/upgrade/scope.go b/operator/internal/upgrade/scope.go index ffd7a2b6d..4492cdb85 100644 --- a/operator/internal/upgrade/scope.go +++ b/operator/internal/upgrade/scope.go @@ -47,6 +47,22 @@ type Scope struct { // reaches. Graph *Graph + // Pools resolves a pool's name to its UUID, which §16.4 needs and nothing + // in Kubernetes can answer: a volume handle provisioned before the v2 API + // migration spells its pool as a name, and only the control plane knows + // which pool that is. + // + // It is the one thing a rule reads that is not the cluster, and it is an + // interface for that reason: every other rule is testable with a fake + // Kubernetes client alone, and this one would otherwise need a control + // plane to be tested at all. + // + // Nil where the run could not reach a control plane, which is the ordinary + // state of a plan being read rather than applied. A step needing it refuses + // rather than proceeding, since a handle it cannot resolve is one it must + // not write. + Pools PoolResolver + // Stage is the command being run, so a rule registered for more than one // can tell which it is in. Stage Stage diff --git a/operator/internal/upgrade/steps/migrate.go b/operator/internal/upgrade/steps/migrate.go index d0a4008d2..9a5d7ab1d 100644 --- a/operator/internal/upgrade/steps/migrate.go +++ b/operator/internal/upgrade/steps/migrate.go @@ -17,6 +17,7 @@ import ( "context" "fmt" + snapshotv1 "github.com/kubernetes-csi/external-snapshotter/client/v8/apis/volumesnapshot/v1" corev1 "k8s.io/api/core/v1" "sigs.k8s.io/controller-runtime/pkg/client" @@ -36,11 +37,8 @@ const ( IDNormalizeHandles upgrade.ID = "normalize-volume-handles" ) -// What the unimplemented halves wait on. -const ( - needsTargetTypes = "the v1alpha2 target type does not exist yet (§29.1), so the object cannot be constructed" - needsResolver = "resolving a pool name to its UUID needs the control-plane client of §16.4" -) +// What the unimplemented half waits on. +const needsTargetTypes = "the v1alpha2 target type does not exist yet (§29.1), so the object cannot be constructed" // Migrate returns §16's steps. func Migrate() []upgrade.Step { @@ -56,9 +54,7 @@ func Migrate() []upgrade.Step { as: "Migrate", }, absorbBackupRestores{}, - normalizeHandles{ - described: described{id: IDNormalizeHandles, blocked: needsResolver}, - }, + normalizeHandles{id: IDNormalizeHandles}, } } @@ -310,15 +306,20 @@ func (renameKind) Done(context.Context, *upgrade.Scope, upgrade.Subject) (bool, return false, nil } -// normalizeHandles describes the volumes whose handle carries a pool name -// rather than a UUID. +// normalizeHandles records the resolved handle on every object whose own is +// spelled with a pool name rather than a pool UUID. // // Detecting one needs no control plane: §16.4's shape is // clusterID:poolID:volumeID, and lvol.ParseHandle accepts a pool segment that -// is not a UUID precisely because both spellings occur. What needs the control -// plane is the other half, which is the UUID the name resolves to. +// is not a UUID precisely because both spellings occur. Resolving the name is +// the half that does, and it is the Scope's PoolResolver. +// +// Nothing here rewrites the handle itself. The API server refuses any edit to a +// PersistentVolume's CSI source or a VolumeSnapshotContent's source, so the +// field keeps the spelling it was provisioned with and the annotation carries +// the identity every reader wants. type normalizeHandles struct { - described + id upgrade.ID } func (n normalizeHandles) ID() upgrade.ID { return n.id } @@ -350,7 +351,99 @@ func (normalizeHandles) Done(_ context.Context, _ *upgrade.Scope, subject upgrad return legacy && normalized(subject), nil } -// normalized reports a volume that already carries the resolved handle. +// Validate resolves the pool the handle names, and refuses the subject when +// nothing answers to it. +// +// The refusal is here rather than in Apply because this runs in the preflight, +// where nothing has been written yet: a handle naming a pool that no longer +// exists is a finding somebody has to look at, and finding it before the +// migration starts is the difference between a question and a half-migrated +// cluster. +func (normalizeHandles) Validate(ctx context.Context, s *upgrade.Scope, subject upgrade.Subject) error { + _, err := resolvedHandle(ctx, s, subject) + return err +} + +// Apply writes the resolved handle to the annotation. +func (normalizeHandles) Apply(ctx context.Context, s *upgrade.Scope, subject upgrade.Subject) error { + handle, err := resolvedHandle(ctx, s, subject) + if err != nil { + return err + } + + obj, ok := subject.Object.DeepCopyObject().(client.Object) + if !ok { + return fmt.Errorf("%T is not a Kubernetes object", subject.Object) + } + annotations := obj.GetAnnotations() + if annotations == nil { + annotations = map[string]string{} + } + annotations[atlaskube.AnnoVolumeHandle] = handle.String() + obj.SetAnnotations(annotations) + + if err := s.Client.Update(ctx, obj); err != nil { + return fmt.Errorf("recording the normalized handle: %w", err) + } + s.Adopt(obj) + return nil +} + +// Verify re-reads the object and checks the annotation is on it and agrees with +// the field, which is the same rule every reader applies. +func (normalizeHandles) Verify(ctx context.Context, s *upgrade.Scope, subject upgrade.Subject) error { + fresh, ok := subject.Object.DeepCopyObject().(client.Object) + if !ok { + return fmt.Errorf("%T is not a Kubernetes object", subject.Object) + } + if err := s.Client.Get(ctx, subject.Ref.Key(), fresh); err != nil { + return fmt.Errorf("re-reading it: %w", err) + } + + field, carried := handleOf(fresh) + if !carried { + return fmt.Errorf("it no longer carries a handle this migration is about") + } + written, wellFormed := atlaskube.NormalizedHandle(field, fresh.GetAnnotations()) + switch { + case !wellFormed: + return fmt.Errorf("the handle %q is no longer readable", field) + case !written.FromAnnotation: + return fmt.Errorf("%s was not written", atlaskube.AnnoVolumeHandle) + case written.Ignored != "": + return fmt.Errorf("the annotation would be ignored by a reader: %s", written.Ignored) + case !lvol.IsCanonicalUUID(written.Handle.PoolRef): + return fmt.Errorf("the annotated pool %q is still not a UUID", written.Handle.PoolRef) + } + return nil +} + +// resolvedHandle is the handle this subject should carry once its pool name is +// resolved, and the error a caller reports when it cannot be. +func resolvedHandle( + ctx context.Context, s *upgrade.Scope, subject upgrade.Subject, +) (lvol.Handle, error) { + handle, legacy := legacyHandle(subject) + if !legacy { + return lvol.Handle{}, nil + } + if s.Pools == nil { + return lvol.Handle{}, fmt.Errorf( + "the handle names pool %q rather than a UUID, and only the control plane knows "+ + "which pool that is; give --control-plane the management API's base URL", + handle.PoolRef) + } + + uuid, err := s.Pools.PoolUUID(ctx, handle.ClusterID, handle.PoolRef) + if err != nil { + return lvol.Handle{}, fmt.Errorf( + "the handle names pool %q in cluster %s: %w", handle.PoolRef, handle.ClusterID, err) + } + handle.PoolRef = uuid + return handle, nil +} + +// normalized reports an object that already carries the resolved handle. func normalized(subject upgrade.Subject) bool { if subject.Object == nil { return false @@ -359,16 +452,11 @@ func normalized(subject upgrade.Subject) bool { return carried } -// legacyHandle reports the handle of a PersistentVolume whose pool segment is -// not a UUID. +// legacyHandle reports the handle of an object whose pool segment is not a +// UUID, and false for everything else. func legacyHandle(subject upgrade.Subject) (lvol.Handle, bool) { - pv, ok := subject.Object.(*corev1.PersistentVolume) - if !ok || !atlaskube.IsManaged(pv) { - return lvol.Handle{}, false - } - - raw, err := atlaskube.VolumeHandleFromPV(pv) - if err != nil { + raw, carried := handleOf(subject.Object) + if !carried { return lvol.Handle{}, false } handle, wellFormed := lvol.ParseHandle(raw) @@ -377,3 +465,37 @@ func legacyHandle(subject upgrade.Subject) (lvol.Handle, bool) { } return handle, true } + +// handleOf is the handle an object of either kind carries in the field §16.4 +// cannot rewrite, and false for an object of any other kind or another driver's. +// +// A VolumeSnapshotContent names one of two things, and both are the same three +// segments: a pre-existing snapshot names itself in spec.source.snapshotHandle, +// and a dynamically taken one names the volume it was taken from in +// spec.source.volumeHandle. Whichever it carries is the one whose pool segment +// may be a name. +func handleOf(obj client.Object) (lvol.VolumeHandle, bool) { + switch object := obj.(type) { + case *corev1.PersistentVolume: + raw, err := atlaskube.VolumeHandleFromPV(object) + if err != nil { + return "", false + } + return raw, true + + case *snapshotv1.VolumeSnapshotContent: + if object.Spec.Driver != atlaskube.DriverName { + return "", false + } + for _, source := range []*string{ + object.Spec.Source.SnapshotHandle, + object.Spec.Source.VolumeHandle, + } { + if source != nil && *source != "" { + return lvol.VolumeHandle(*source), true + } + } + return "", false + } + return "", false +} diff --git a/operator/internal/upgrade/steps/migrate_test.go b/operator/internal/upgrade/steps/migrate_test.go index c56647486..34c1fc96a 100644 --- a/operator/internal/upgrade/steps/migrate_test.go +++ b/operator/internal/upgrade/steps/migrate_test.go @@ -5,19 +5,30 @@ package steps import ( + "context" + "fmt" "strings" "testing" + snapshotv1 "github.com/kubernetes-csi/external-snapshotter/client/v8/apis/volumesnapshot/v1" corev1 "k8s.io/api/core/v1" apierrors "k8s.io/apimachinery/pkg/api/errors" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" "k8s.io/apimachinery/pkg/types" "sigs.k8s.io/controller-runtime/pkg/client" + "github.com/simplyblock/atlas/errs" + atlaskube "github.com/simplyblock/atlas/kube" + simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" "github.com/simplyblock/simplyblock-operator/internal/upgrade" ) +// testPoolUUID is the UUID my-pool resolves to, which is the pool segment of +// modernHandleValue: the two are one fact and a test that spelled them apart +// would assert a resolution nothing performed. +const testPoolUUID = "1c2c0300-9993-4289-be95-59414fc8a54d" + // migratePlan is what §16's steps would do to this cluster. func migratePlan(t *testing.T, objects ...client.Object) upgrade.Plan { t.Helper() @@ -239,3 +250,144 @@ func TestMigrate_EveryStepDescribesOrDeclines(t *testing.T) { t.Fatalf("a cluster with nothing to migrate produced %d tasks:\n%v", len(plan.Tasks), plan.Tasks) } } + +// legacySnapshot is a VolumeSnapshotContent whose source handle carries a pool +// name. A pre-existing snapshot names one in spec.source.snapshotHandle and a +// dynamically taken one names its volume in spec.source.volumeHandle, and both +// are the same three segments with the same middle one. +func legacySnapshot(name, handle string, preExisting bool) *snapshotv1.VolumeSnapshotContent { + source := snapshotv1.VolumeSnapshotContentSource{VolumeHandle: &handle} + if preExisting { + source = snapshotv1.VolumeSnapshotContentSource{SnapshotHandle: &handle} + } + return &snapshotv1.VolumeSnapshotContent{ + ObjectMeta: metav1.ObjectMeta{Name: name}, + Spec: snapshotv1.VolumeSnapshotContentSpec{ + Driver: "csi.simplyblock.io", + Source: source, + }, + } +} + +func TestMigrate_FindsALegacyHandleOnASnapshot(t *testing.T) { + // §16.4's second kind. A snapshot id is composed the same way a volume + // handle is, so a VolumeSnapshotContent written before the boundary carries + // a pool name in exactly the same place. + for _, tc := range []struct { + name string + preExisting bool + }{ + {name: "a dynamically taken snapshot names its volume", preExisting: false}, + {name: "a pre-existing snapshot names itself", preExisting: true}, + } { + t.Run(tc.name, func(t *testing.T) { + plan := migratePlan(t, legacySnapshot("snapcontent-1", legacyHandleValue, tc.preExisting)) + + task := taskFor(t, plan, IDNormalizeHandles) + if len(task.Subtasks) != 1 { + t.Fatalf("described %d subtasks, want the one legacy snapshot:\n%v", + len(task.Subtasks), task.Subtasks) + } + if got := task.Subtasks[0].String(); !strings.Contains(got, `"my-pool"`) { + t.Errorf("the subtask does not name the pool it would resolve:\n%s", got) + } + }) + } +} + +func TestMigrate_LeavesAModernSnapshotHandleAlone(t *testing.T) { + plan := migratePlan(t, legacySnapshot("snapcontent-modern", modernHandleValue, true)) + + for _, task := range plan.Tasks { + if task.Step == IDNormalizeHandles { + t.Fatalf("a snapshot whose pool segment is already a UUID was planned for:\n%v", + task.Subtasks) + } + } +} + +func TestMigrate_LeavesASnapshotOfAnotherDriverAlone(t *testing.T) { + content := legacySnapshot("snapcontent-other", legacyHandleValue, true) + content.Spec.Driver = "ebs.csi.aws.com" + + for _, task := range migratePlan(t, content).Tasks { + if task.Step == IDNormalizeHandles { + t.Fatalf("a snapshot this driver did not take was planned for:\n%v", task.Subtasks) + } + } +} + +// TestMigrate_WritesTheResolvedHandleIntoTheAnnotation is the write half of +// §16.4: the field keeps the spelling it was provisioned with, because the API +// server refuses to change it, and the resolved identity goes to metadata. +func TestMigrate_WritesTheResolvedHandleIntoTheAnnotation(t *testing.T) { + pv := legacyPV("pvc-legacy", legacyHandleValue) + content := legacySnapshot("snapcontent-1", legacyHandleValue, true) + + scope := migration(t, pv, content) + scope.Pools = fixedPools{"my-pool": testPoolUUID} + if err := runNormalization(t, scope); err != nil { + t.Fatalf("normalizing: %v", err) + } + + var written corev1.PersistentVolume + if err := scope.Client.Get(t.Context(), types.NamespacedName{Name: "pvc-legacy"}, &written); err != nil { + t.Fatalf("re-reading the volume: %v", err) + } + if got := written.Annotations[atlaskube.AnnoVolumeHandle]; got != modernHandleValue { + t.Errorf("annotation = %q, want the resolved %q", got, modernHandleValue) + } + if written.Spec.CSI.VolumeHandle != legacyHandleValue { + t.Errorf("the immutable field was rewritten to %q", written.Spec.CSI.VolumeHandle) + } + + var snapshot snapshotv1.VolumeSnapshotContent + if err := scope.Client.Get(t.Context(), types.NamespacedName{Name: "snapcontent-1"}, &snapshot); err != nil { + t.Fatalf("re-reading the snapshot: %v", err) + } + if got := snapshot.Annotations[atlaskube.AnnoVolumeHandle]; got != modernHandleValue { + t.Errorf("snapshot annotation = %q, want the resolved %q", got, modernHandleValue) + } +} + +// A pool name that resolves to nothing names a pool that no longer exists, and +// §16.4 says the migration reports it and does not proceed. It is refused in +// Validate, which runs in the preflight, so it is found before anything is +// written rather than partway through. +func TestMigrate_RefusesAPoolNameThatResolvesToNothing(t *testing.T) { + scope := migration(t, legacyPV("pvc-legacy", legacyHandleValue)) + scope.Pools = fixedPools{} + + err := runNormalization(t, scope) + if err == nil { + t.Fatal("a handle naming a pool that does not exist was normalized anyway") + } + if !strings.Contains(err.Error(), "my-pool") { + t.Errorf("the refusal does not name the pool nobody can find: %v", err) + } +} + +// fixedPools is a resolver with a fixed answer, so the step can be driven +// without a control plane. +type fixedPools map[string]string + +func (f fixedPools) PoolUUID(_ context.Context, _, name string) (string, error) { + uuid, known := f[name] + if !known { + return "", fmt.Errorf("no pool named %q: %w", name, errs.ErrNotFound) + } + return uuid, nil +} + +// runNormalization drives the one step, the way the migrate phase does. +func runNormalization(t *testing.T, scope *upgrade.Scope) error { + t.Helper() + + catalog := upgrade.NewCatalog() + for _, step := range Migrate() { + if step.ID() == IDNormalizeHandles { + catalog.Steps.MustRegister(step) + } + } + return upgrade.NewRunner(catalog, scope).ApplyAll(t.Context(), upgrade.StageMigrate) +} diff --git a/operator/internal/upgrade/steps/ownership_test.go b/operator/internal/upgrade/steps/ownership_test.go index 48ecf1c5f..f69569e82 100644 --- a/operator/internal/upgrade/steps/ownership_test.go +++ b/operator/internal/upgrade/steps/ownership_test.go @@ -10,6 +10,7 @@ import ( "strings" "testing" + snapshotv1 "github.com/kubernetes-csi/external-snapshotter/client/v8/apis/volumesnapshot/v1" appsv1 "k8s.io/api/apps/v1" corev1 "k8s.io/api/core/v1" apiextensionsv1 "k8s.io/apiextensions-apiserver/pkg/apis/apiextensions/v1" @@ -123,6 +124,12 @@ func migration(t *testing.T, objects ...client.Object) *upgrade.Scope { if err := apiextensionsv1.AddToScheme(scheme); err != nil { t.Fatalf("registering apiextensions: %v", err) } + // §16.4 normalizes the handles on VolumeSnapshotContent objects as well as + // on PersistentVolumes, and the kind belongs to the external snapshotter + // rather than to the core API. + if err := snapshotv1.AddToScheme(scheme); err != nil { + t.Fatalf("registering the snapshot API: %v", err) + } // The fake client routes a status write through the subresource tracker only // for a kind it was told has one, and §16.2's absorbed operation carries its diff --git a/operator/internal/webhook/volumehandle.go b/operator/internal/webhook/volumehandle.go index 2b67a7d7a..3257305fc 100644 --- a/operator/internal/webhook/volumehandle.go +++ b/operator/internal/webhook/volumehandle.go @@ -1,14 +1,15 @@ // volumehandle.go resolves a PersistentVolume to its simplyblock volume // handle for admission checks. The handle grammar and the PV ownership check -// live in atlas (lvol.ParseHandle, kube.VolumeHandleFromPV); this file only -// composes them with the client read the webhooks need. +// live in atlas (kube.NormalizedVolumeHandleFromPV, which applies §16.4's rule +// that a resolved handle recorded in an annotation is preferred to the legacy +// spelling the immutable field keeps); this file only composes them with the +// client read the webhooks need. package webhook import ( "context" "github.com/simplyblock/atlas/kube" - "github.com/simplyblock/atlas/lvol" corev1 "k8s.io/api/core/v1" "k8s.io/apimachinery/pkg/types" "sigs.k8s.io/controller-runtime/pkg/client" @@ -28,14 +29,12 @@ func pvVolumeHandle( if err := c.Get(ctx, types.NamespacedName{Name: pvName}, pv); err != nil { return "", "", "", false, err } - raw, err := kube.VolumeHandleFromPV(pv) + normalized, err := kube.NormalizedVolumeHandleFromPV(pv) if err != nil { - // Not a volume this driver owns: undeterminable, not a failure. - return "", "", "", false, nil - } - h, parsed := lvol.ParseHandle(raw) - if !parsed { + // Not a volume this driver owns, or one whose handle does not parse: + // undeterminable, not a failure. return "", "", "", false, nil } + h := normalized.Handle return h.ClusterID, h.PoolRef, h.VolumeID, true, nil } From c14938aca80531fcc4d3d166cf9dc446b8060228 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 13:55:23 +0200 Subject: [PATCH 046/206] feat(driver): the driver asks whether the cluster has snapshots before assuming MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit design-simplyblockdriver §4.1 reads snapshot support from the API serving snapshot.storage.k8s.io/v1, because the snapshot-controller exists to reconcile those kinds and its Deployment is named differently by every distribution. The operator asked nothing. It applied a VolumeSnapshotClass whenever the toggle was on, and the whole object set goes out in one pass, so on a cluster serving no snapshot API the apply failed on that one object and took the node plugin and the controller plugin with it — a snapshot class nobody asked to be essential. The chart's CRDs masked it, and an OLM install or a release with snapshotcontroller.create false does not have them. The detection is a SnapshotAPI interface rather than a discovery client reached for in the reconcile, since a reconciler holding one directly cannot be tested against a cluster that has no snapshot API, which is the case this exists to handle. The implementation caches its answer for a window shorter than the driver's resync: discovery is a round trip, the set of served groups changes when somebody installs a CRD rather than continuously, and the window bounds how long a cluster that has just gained snapshot support waits to be told. An answer that is neither yes nor no fails the reconcile rather than being guessed at, and the two guesses break a cluster in opposite directions. Guessing served applies a class the API server has no kind for. Guessing absent withdraws a class an adopted cluster is using, on a transient error. status.snapshotSupport is written now, and the SnapshotsEnabled event of §6.1 is emitted with it. The field records what happened when the deployment came up rather than being maintained: a cluster that later loses the snapshot API has a problem this field is not the place to report. ownedClusterScoped is built rather than filtered out of desired, because the two now answer different questions. desired is what to apply, and on a bare cluster that excludes the class; the finalizer needs what might exist, which includes a class applied before the toggle was turned off or before the kinds were removed from under it. A deletion that consulted the live API would skip exactly the object nothing else removes. The install half of §4.1 is not built, and the design says so now instead of a TODO in the code. It is smaller than it was: the chart applies the CRDs and the controller today and applies them conditionally, guarded on .Capabilities.APIVersions.Has, which is the same rule §4.1 states. What is left is the installation that is not a chart. That needs the upstream manifests carried in this binary and an image for the controller that no field on the spec names, so SnapshotSupportOriginInstalled stays unreachable and §9 Q2, which asks what removes them afterward, stays open. Co-Authored-By: Claude Fable 5 --- operator/cmd/main.go | 12 +- .../crd-redesign/design-simplyblockdriver.md | 22 +++ .../internal/controllers/driver/discovery.go | 86 ++++++++++ .../controllers/driver/discovery_test.go | 110 ++++++++++++ .../controllers/driver/registration.go | 86 ++++++++-- .../controllers/driver/registration_test.go | 159 ++++++++++++++++++ .../driver/simplyblockdriver_controller.go | 93 ++++++++-- .../simplyblockdriver_controller_test.go | 39 +++-- 8 files changed, 560 insertions(+), 47 deletions(-) create mode 100644 operator/internal/controllers/driver/discovery.go create mode 100644 operator/internal/controllers/driver/discovery_test.go diff --git a/operator/cmd/main.go b/operator/cmd/main.go index 91b7c60ad..186d224f2 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -774,10 +774,16 @@ func main() { setupLog.Error(err, "unable to create controller", "controller", "OperatorOps") os.Exit(1) } + // Whether the cluster serves the snapshot API decides both whether a + // VolumeSnapshotClass is applied and what status.snapshotSupport records + // (design-simplyblockdriver.md §4.1). The discovery client is the same one + // the discovery run uses; a driver reconcile without it refuses rather than + // guessing, because guessing either way breaks a cluster in one direction. if err := (&driver.SimplyblockDriverReconciler{ - Client: mgr.GetClient(), - Scheme: mgr.GetScheme(), - Recorder: mgr.GetEventRecorder("simplyblockdriver-controller"), + Client: mgr.GetClient(), + Scheme: mgr.GetScheme(), + Recorder: mgr.GetEventRecorder("simplyblockdriver-controller"), + Snapshots: &driver.DiscoveredSnapshotAPI{Discovery: operatorOpsDiscovery}, }).SetupWithManager(mgr); err != nil { setupLog.Error(err, "unable to create controller", "controller", "SimplyblockDriver") os.Exit(1) diff --git a/operator/docs/designs/crd-redesign/design-simplyblockdriver.md b/operator/docs/designs/crd-redesign/design-simplyblockdriver.md index 17d790b69..0da745ecb 100644 --- a/operator/docs/designs/crd-redesign/design-simplyblockdriver.md +++ b/operator/docs/designs/crd-redesign/design-simplyblockdriver.md @@ -365,10 +365,32 @@ the API is served the operator adds nothing. Where it is absent the operator applies the CRDs and a controller, which is what the chart does for the CRDs alone today (§8). +The detection is a discovery question rather than a search for an object, and +the answer is cached for a window shorter than the driver's resync, so a cluster +that gains the snapshot API is noticed on a reconcile rather than on the one +after it. An answer that is neither yes nor no fails the reconcile instead of +being guessed at: guessing served applies a class the API server has no kind +for, and guessing absent withdraws a class an adopted cluster is using. + +**The install half is not built.** The chart applies the CRDs and the controller +today, and conditionally — its templates are guarded on +`.Capabilities.APIVersions.Has`, which is the same rule stated here. What a +chart cannot cover is an installation that is not a chart, or a release with +`snapshotcontroller.create` false, and that is the case still open. It needs the +upstream manifests carried in the operator's binary and an image for the +controller that no field names, so `status.snapshotSupport` reaches `Detected` +and not yet `Installed`. + **The `VolumeSnapshotClass` for this driver is applied either way.** It names `spec.driverName` and belongs to this deployment, unlike the CRDs and the controller, which belong to the cluster. +Either way means either origin, and not either cluster. A cluster serving no +snapshot API has no kind for the object, so the class is built only where the +kinds exist — an apply of a kind the API server does not serve fails the apply +of the whole set on that one object, which takes the node plugin and the +controller plugin down with a snapshot class nobody asked to be essential. + **What the operator installs here it does not own.** The CRDs and the controller are cluster-scoped and shared, and a second CSI driver installed afterward reconciles its snapshots through the same controller. They are applied without a diff --git a/operator/internal/controllers/driver/discovery.go b/operator/internal/controllers/driver/discovery.go new file mode 100644 index 000000000..43f3af3c5 --- /dev/null +++ b/operator/internal/controllers/driver/discovery.go @@ -0,0 +1,86 @@ +// Asking the API server whether this cluster has snapshot support. +// +// §4.1 reads the presence of a snapshot-controller from the API serving +// snapshot.storage.k8s.io/v1 rather than from any object, because that +// controller exists to reconcile those kinds and its Deployment is named +// differently by every distribution: kube-system/snapshot-controller on one, +// something else on the next, and nothing findable on a managed service that +// runs it outside the cluster. +// +// The answer is cached for a while rather than asked per reconcile. Discovery +// is a round trip against the API server's aggregated document, the driver +// resyncs on a timer, and the set of served API groups changes when somebody +// installs a CRD rather than continuously. The window is what bounds how long a +// cluster that has just gained snapshot support waits to be told. + +package driver + +import ( + "context" + "fmt" + "sync" + "time" + + apierrors "k8s.io/apimachinery/pkg/api/errors" + "k8s.io/client-go/discovery" +) + +// snapshotAPITTL is how long a discovery answer is trusted. +// +// It is shorter than the driver's own resync, so a cluster that gains the +// snapshot API is noticed on a reconcile rather than on the one after it, and +// long enough that a fleet of drivers does not turn one question into a poll. +const snapshotAPITTL = 2 * time.Minute + +// DiscoveredSnapshotAPI answers [SnapshotAPI] from the API server's discovery +// document. +type DiscoveredSnapshotAPI struct { + // Discovery is the client the question goes to. + Discovery discovery.DiscoveryInterface + + mu sync.Mutex + served bool + asked bool + askedAt time.Time + nowFuncT func() time.Time +} + +// SnapshotAPIServed reports whether the cluster serves +// snapshot.storage.k8s.io/v1. +// +// A group that is not served comes back as a NotFound rather than as an empty +// list, and that is the ordinary answer here rather than a failure: it is what +// a cluster without the snapshot CRDs says. Every other error is returned, so a +// reconcile refuses rather than deciding from an API server it could not reach. +func (d *DiscoveredSnapshotAPI) SnapshotAPIServed(_ context.Context) (bool, error) { + d.mu.Lock() + defer d.mu.Unlock() + + if d.asked && d.now().Sub(d.askedAt) < snapshotAPITTL { + return d.served, nil + } + if d.Discovery == nil { + return false, fmt.Errorf("no discovery client is configured") + } + + _, err := d.Discovery.ServerResourcesForGroupVersion(snapshotGroupVersion.String()) + switch { + case apierrors.IsNotFound(err): + d.served = false + case err != nil: + return false, fmt.Errorf("read the resources of %s: %w", snapshotGroupVersion, err) + default: + d.served = true + } + + d.asked, d.askedAt = true, d.now() + return d.served, nil +} + +// now is the clock, which a test replaces to age the cache without waiting. +func (d *DiscoveredSnapshotAPI) now() time.Time { + if d.nowFuncT != nil { + return d.nowFuncT() + } + return time.Now() +} diff --git a/operator/internal/controllers/driver/discovery_test.go b/operator/internal/controllers/driver/discovery_test.go new file mode 100644 index 000000000..bbdf6dd26 --- /dev/null +++ b/operator/internal/controllers/driver/discovery_test.go @@ -0,0 +1,110 @@ +// The detection itself: what a cluster without the snapshot CRDs says, and what +// is done with an answer that is neither yes nor no. + +package driver + +import ( + "errors" + "testing" + "time" + + apierrors "k8s.io/apimachinery/pkg/api/errors" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/runtime/schema" + "k8s.io/client-go/discovery" +) + +// answeringDiscovery is a discovery client that returns one answer, and counts +// how often it was asked. +type answeringDiscovery struct { + discovery.DiscoveryInterface + err error + asked int +} + +func (a *answeringDiscovery) ServerResourcesForGroupVersion( + string, +) (*metav1.APIResourceList, error) { + a.asked++ + if a.err != nil { + return nil, a.err + } + return &metav1.APIResourceList{}, nil +} + +// A cluster without the snapshot CRDs answers NotFound for the group version, +// and that is the answer rather than a failure: it is what "this cluster has no +// snapshot support" looks like over discovery. +func TestAGroupThatIsNotServedIsAnAnswerRatherThanAnError(t *testing.T) { + api := &DiscoveredSnapshotAPI{Discovery: &answeringDiscovery{ + err: apierrors.NewNotFound(schema.GroupResource{Group: snapshotGroupVersion.Group}, ""), + }} + + served, err := api.SnapshotAPIServed(t.Context()) + if err != nil { + t.Fatalf("a cluster with no snapshot CRDs reported an error: %v", err) + } + if served { + t.Error("a cluster with no snapshot CRDs was reported as serving the API") + } +} + +func TestAServedGroupIsReported(t *testing.T) { + api := &DiscoveredSnapshotAPI{Discovery: &answeringDiscovery{}} + + served, err := api.SnapshotAPIServed(t.Context()) + if err != nil { + t.Fatalf("asking: %v", err) + } + if !served { + t.Error("a cluster serving the snapshot API was reported as not serving it") + } +} + +// Anything that is not an answer is returned. A reconcile that guessed here +// would either apply a class the cluster has no kind for or withdraw one an +// adopted cluster is using. +func TestAnUnreachableAPIServerIsReported(t *testing.T) { + api := &DiscoveredSnapshotAPI{Discovery: &answeringDiscovery{ + err: errors.New("connection refused"), + }} + + if _, err := api.SnapshotAPIServed(t.Context()); err == nil { + t.Error("an unreachable API server was read as an answer about the snapshot API") + } +} + +// The answer is cached, because discovery is a round trip and the set of served +// groups changes when somebody installs a CRD rather than continuously. +func TestTheAnswerIsCachedAndThenAskedAgain(t *testing.T) { + clock := time.Now() + discovered := &answeringDiscovery{} + api := &DiscoveredSnapshotAPI{Discovery: discovered, nowFuncT: func() time.Time { return clock }} + + for range 3 { + if _, err := api.SnapshotAPIServed(t.Context()); err != nil { + t.Fatalf("asking: %v", err) + } + } + if discovered.asked != 1 { + t.Errorf("asked the API server %d times within the window, want once", discovered.asked) + } + + clock = clock.Add(snapshotAPITTL + time.Second) + if _, err := api.SnapshotAPIServed(t.Context()); err != nil { + t.Fatalf("asking: %v", err) + } + if discovered.asked != 2 { + t.Errorf("asked %d times after the window expired, want the question to be asked "+ + "again so a cluster that gains snapshot support is noticed", discovered.asked) + } +} + +// The zero value refuses rather than reporting a cluster with no snapshot +// support, which would be indistinguishable from an answer. +func TestNoDiscoveryClientRefuses(t *testing.T) { + var api DiscoveredSnapshotAPI + if _, err := api.SnapshotAPIServed(t.Context()); err == nil { + t.Error("a detector with no client answered the question") + } +} diff --git a/operator/internal/controllers/driver/registration.go b/operator/internal/controllers/driver/registration.go index 311a1bcfa..dddb27b6b 100644 --- a/operator/internal/controllers/driver/registration.go +++ b/operator/internal/controllers/driver/registration.go @@ -19,6 +19,9 @@ package driver import ( + "context" + "fmt" + storagev1 "k8s.io/api/storage/v1" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" @@ -62,26 +65,75 @@ func volumeSnapshotClass(d *simplyblockv1alpha2.SimplyblockDriver) *unstructured } // TODO(simplyblockdriver): supply the snapshot CRDs and a controller where the -// cluster serves neither, which is design-simplyblockdriver.md §4.1. Today, the -// chart still installs both, into kube-system and annotated -// helm.sh/resource-policy: keep, so an adopted deployment finds the API served -// and records Detected. What is missing here is the detection against the -// discovery client, the apply of the CRDs and the controller where it comes -// back empty, and status.snapshotSupport reading Installed in that case. They -// are cluster-scoped and shared, so they carry no owner reference and outlive -// this object, which is what design §9 Q2 leaves open. +// cluster serves neither, which is the second half of design-simplyblockdriver.md +// §4.1 and is why SnapshotSupportOriginInstalled is not yet reachable. +// +// The chart installs both today, and conditionally: its CRD templates are +// guarded on .Capabilities.APIVersions.Has, so a cluster already serving the +// kinds gets nothing, which is the same rule §4.1 states for the operator. What +// a chart cannot cover is an installation that is not a chart — an OLM bundle, +// or a release with snapshotcontroller.create false — and that is the case this +// owes. It needs the upstream manifests carried in this binary and an image for +// the controller, which no field on the spec names, so it is a change of its own +// rather than a line here. §9 Q2 is what removes them afterward, and it is open. // -// status.snapshotSupport is unwritten in both cases today, and Detected is the -// half that needs nothing new: every adopted cluster is already serving the API, -// so the discovery check alone would settle it. Installed waits on the apply -// above, and the Normal SnapshotsEnabled event design §6.1 owes waits with it. +// The test plan's U-38 to U-41 are the rows this owes. U-04 and U-05 are the +// detection below. + +// SnapshotAPI answers whether the cluster serves the snapshot kinds, which is +// §4.1's detection: a snapshot-controller exists to reconcile those kinds and +// its Deployment is named differently by every distribution, so the API being +// served is the question rather than any object being present. // -// The apply is also not tolerant of a cluster that serves no snapshot API. The -// VolumeSnapshotClass goes into the object set whenever the toggle is on, so on -// such a cluster the whole reconcile fails on that one object rather than -// skipping it. Detecting first fixes that too. +// It is an interface because the answer comes from discovery rather than from +// the object graph, and a reconciler that reached for a discovery client +// directly could not be tested against a cluster that has no snapshot API — +// which is the case this exists to handle. +type SnapshotAPI interface { + // SnapshotAPIServed reports whether snapshot.storage.k8s.io/v1 is served. + SnapshotAPIServed(ctx context.Context) (bool, error) +} + +// snapshotOrigin is what status.snapshotSupport should read, and the empty +// string for a deployment that asked for no snapshots. +// +// Only Detected is reachable today. Installed is what the apply above would +// record, and recording it before that apply exists would say this deployment +// brought snapshot support to a cluster where nothing did. +func (r *SimplyblockDriverReconciler) snapshotOrigin( + ctx context.Context, d *simplyblockv1alpha2.SimplyblockDriver, +) (simplyblockv1alpha2.SnapshotSupportOrigin, error) { + if !snapshotsEnabled(d) { + return "", nil + } + served, err := r.snapshotAPIServed(ctx) + if err != nil { + return "", err + } + if !served { + return "", nil + } + return simplyblockv1alpha2.SnapshotSupportOriginDetected, nil +} + +// snapshotAPIServed asks the cluster, and refuses to guess. // -// The test plan's U-04, U-05, and U-38 to U-41 are the rows this owes. +// A reconcile with no way to ask is a reconcile that must fail rather than +// assume. Assuming served applies a class the API server may have no kind for, +// and assuming absent drops a class an adopted cluster already has, which +// withdraws snapshot support from a working deployment on a transient error. +func (r *SimplyblockDriverReconciler) snapshotAPIServed(ctx context.Context) (bool, error) { + if r.Snapshots == nil { + return false, fmt.Errorf( + "no snapshot-API detector is configured, so whether this cluster serves " + + "snapshot.storage.k8s.io/v1 cannot be established") + } + served, err := r.Snapshots.SnapshotAPIServed(ctx) + if err != nil { + return false, fmt.Errorf("ask whether the cluster serves the snapshot API: %w", err) + } + return served, nil +} // snapshotsEnabled reports whether this deployment includes snapshot support. // The field defaults to true, so an object written before the default applied diff --git a/operator/internal/controllers/driver/registration_test.go b/operator/internal/controllers/driver/registration_test.go index 500f79931..32da742d6 100644 --- a/operator/internal/controllers/driver/registration_test.go +++ b/operator/internal/controllers/driver/registration_test.go @@ -7,9 +7,13 @@ package driver import ( + "context" + "errors" "testing" storagev1 "k8s.io/api/storage/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" "github.com/simplyblock/atlas/ptr" simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" @@ -180,3 +184,158 @@ func TestSidecarsAreSixAndExcludeTheSnapshotController(t *testing.T) { t.Errorf("six overrides produced %d distinct images: %+v", len(distinct), got) } } + +// snapshotAPI is a cluster that serves the snapshot API, or does not. +type fixedSnapshotAPI struct { + served bool + err error +} + +func (f fixedSnapshotAPI) SnapshotAPIServed(context.Context) (bool, error) { + return f.served, f.err +} + +// U-04: a cluster already serving the API is one the operator adds nothing to. +func TestSnapshotSupportIsDetectedWhereTheAPIIsServed(t *testing.T) { + d := testDriver("simplyblock") + r := &SimplyblockDriverReconciler{Snapshots: fixedSnapshotAPI{served: true}} + + origin, err := r.snapshotOrigin(t.Context(), d) + if err != nil { + t.Fatalf("detecting: %v", err) + } + if origin != simplyblockv1alpha2.SnapshotSupportOriginDetected { + t.Errorf("origin = %q, want Detected", origin) + } +} + +// U-05: a cluster serving no snapshot API gets no VolumeSnapshotClass, because +// applying one is a request the API server has no kind for. The whole set goes +// out in one pass, so a class the cluster cannot accept fails the apply of the +// node plugin and the controller plugin with it. +func TestABareClusterGetsNoSnapshotClass(t *testing.T) { + d := testDriver("simplyblock") + r := &SimplyblockDriverReconciler{Snapshots: fixedSnapshotAPI{served: false}} + + objects, err := r.desired(t.Context(), d, "image:tag") + if err != nil { + t.Fatalf("building the object set: %v", err) + } + for _, obj := range objects { + if obj.GetObjectKind().GroupVersionKind() == volumeSnapshotClassGVK { + t.Fatal("a VolumeSnapshotClass was built for a cluster that serves no snapshot " + + "API, so the apply fails on it and nothing else in the set is written") + } + } +} + +// The class is built where the API is served and the toggle is on, which is the +// case every adopted cluster is in. +func TestASnapshotClassIsBuiltWhereTheAPIIsServed(t *testing.T) { + d := testDriver("simplyblock") + r := &SimplyblockDriverReconciler{Snapshots: fixedSnapshotAPI{served: true}} + + objects, err := r.desired(t.Context(), d, "image:tag") + if err != nil { + t.Fatalf("building the object set: %v", err) + } + var found bool + for _, obj := range objects { + if obj.GetObjectKind().GroupVersionKind() == volumeSnapshotClassGVK { + found = true + } + } + if !found { + t.Error("no VolumeSnapshotClass was built for a cluster that serves the API") + } +} + +// The toggle still wins. A deployment that asked for no snapshots gets none +// wherever it runs, and reports neither origin. +func TestSnapshotsDisabledReportsNoOriginAndBuildsNoClass(t *testing.T) { + d := testDriver("simplyblock") + d.Spec.EnableVolumeSnapshots = ptr.To(false) + r := &SimplyblockDriverReconciler{Snapshots: fixedSnapshotAPI{served: true}} + + origin, err := r.snapshotOrigin(t.Context(), d) + if err != nil { + t.Fatalf("detecting: %v", err) + } + if origin != "" { + t.Errorf("origin = %q, want none for a deployment that disabled snapshots", origin) + } + + objects, err := r.desired(t.Context(), d, "image:tag") + if err != nil { + t.Fatalf("building the object set: %v", err) + } + for _, obj := range objects { + if obj.GetObjectKind().GroupVersionKind() == volumeSnapshotClassGVK { + t.Fatal("a VolumeSnapshotClass was built for a deployment that disabled snapshots") + } + } +} + +// A reconcile that cannot ask the API server whether the kind is served must not +// guess. Guessing served applies a class that may fail; guessing absent drops a +// class an adopted cluster already has, which withdraws snapshot support from a +// working deployment on a transient discovery error. +func TestADiscoveryFailureIsReportedRatherThanAssumed(t *testing.T) { + d := testDriver("simplyblock") + r := &SimplyblockDriverReconciler{ + Snapshots: fixedSnapshotAPI{err: errors.New("the API server said no")}, + } + + if _, err := r.desired(t.Context(), d, "image:tag"); err == nil { + t.Error("a discovery failure was swallowed, so the object set is built on a guess") + } +} + +// U-38: status.snapshotSupport records which of §4.1's two happened, so an +// administrator reading the object learns whether this deployment brought +// snapshot support to the cluster or found it. +func TestReconcileRecordsWhereSnapshotSupportCameFrom(t *testing.T) { + scheme := reconcilerScheme(t) + d := testDriver("simplyblock") + c := fake.NewClientBuilder().WithScheme(scheme).WithObjects(d).WithStatusSubresource(d).Build() + r := &SimplyblockDriverReconciler{ + Client: c, Scheme: scheme, Snapshots: fixedSnapshotAPI{served: true}, + } + + if _, err := r.Reconcile(t.Context(), requestFor(d)); err != nil { + t.Fatalf("reconcile: %v", err) + } + + var got simplyblockv1alpha2.SimplyblockDriver + if err := c.Get(t.Context(), client.ObjectKeyFromObject(d), &got); err != nil { + t.Fatalf("re-read: %v", err) + } + if got.Status.SnapshotSupport != simplyblockv1alpha2.SnapshotSupportOriginDetected { + t.Errorf("status.snapshotSupport = %q, want Detected on a cluster already serving "+ + "the API", got.Status.SnapshotSupport) + } +} + +// A cluster serving no snapshot API reconciles rather than failing, and says +// nothing about an origin it does not have. +func TestReconcileSucceedsOnAClusterWithNoSnapshotAPI(t *testing.T) { + scheme := reconcilerScheme(t) + d := testDriver("simplyblock") + c := fake.NewClientBuilder().WithScheme(scheme).WithObjects(d).WithStatusSubresource(d).Build() + r := &SimplyblockDriverReconciler{ + Client: c, Scheme: scheme, Snapshots: fixedSnapshotAPI{served: false}, + } + + if _, err := r.Reconcile(t.Context(), requestFor(d)); err != nil { + t.Fatalf("the reconcile failed on a cluster that serves no snapshot API: %v", err) + } + + var got simplyblockv1alpha2.SimplyblockDriver + if err := c.Get(t.Context(), client.ObjectKeyFromObject(d), &got); err != nil { + t.Fatalf("re-read: %v", err) + } + if got.Status.SnapshotSupport != "" { + t.Errorf("status.snapshotSupport = %q on a cluster with no snapshot support at all", + got.Status.SnapshotSupport) + } +} diff --git a/operator/internal/controllers/driver/simplyblockdriver_controller.go b/operator/internal/controllers/driver/simplyblockdriver_controller.go index 1e12ec8cb..2b6e7a418 100644 --- a/operator/internal/controllers/driver/simplyblockdriver_controller.go +++ b/operator/internal/controllers/driver/simplyblockdriver_controller.go @@ -58,6 +58,7 @@ const ( reasonDriverAdopted = "DriverAdopted" reasonAdoptionRefused = "AdoptionRefused" reasonNoImage = "NoImage" + reasonSnapshotsEnabled = "SnapshotsEnabled" ) // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=simplyblockdrivers,verbs=get;list;watch;create;update;patch;delete @@ -74,6 +75,11 @@ type SimplyblockDriverReconciler struct { client.Client Scheme *runtime.Scheme Recorder events.EventRecorder + + // Snapshots answers whether the cluster serves the snapshot API, which + // decides both whether a VolumeSnapshotClass is applied and what + // status.snapshotSupport records (§4.1). + Snapshots SnapshotAPI } func (r *SimplyblockDriverReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) { @@ -178,6 +184,10 @@ func (r *SimplyblockDriverReconciler) Reconcile(ctx context.Context, req ctrl.Re return ctrl.Result{}, err } + if err := r.recordSnapshotSupport(ctx, &d); err != nil { + return ctrl.Result{}, err + } + h, err := r.observe(ctx, &d) if err != nil { return ctrl.Result{}, err @@ -255,8 +265,8 @@ func (r *SimplyblockDriverReconciler) event( // the accounts and configuration first, then the RBAC that names the accounts, // then the workloads that mount the configuration, and the registration last. func (r *SimplyblockDriverReconciler) desired( - d *simplyblockv1alpha2.SimplyblockDriver, image string, -) []client.Object { + ctx context.Context, d *simplyblockv1alpha2.SimplyblockDriver, image string, +) ([]client.Object, error) { objects := make([]client.Object, 0, 18) for _, sa := range serviceAccounts(d) { @@ -272,10 +282,21 @@ func (r *SimplyblockDriverReconciler) desired( objects = append(objects, crb) } objects = append(objects, nodeDaemonSet(d, image), controllerStatefulSet(d, image), csiDriver(d)) + + // The class is this deployment's and is applied wherever the kinds exist, + // but only where they exist: a cluster serving no snapshot API has no kind + // for the object, so including it unconditionally fails the apply of the + // whole set on that one object rather than skipping it (§4.1). if snapshotsEnabled(d) { - objects = append(objects, volumeSnapshotClass(d)) + served, err := r.snapshotAPIServed(ctx) + if err != nil { + return nil, err + } + if served { + objects = append(objects, volumeSnapshotClass(d)) + } } - return objects + return objects, nil } // TODO(simplyblockdriver): verify the handover before stripping the release's @@ -297,7 +318,11 @@ func (r *SimplyblockDriverReconciler) apply( if err != nil { return false, err } - for _, obj := range r.desired(d, image) { + objects, err := r.desired(ctx, d, image) + if err != nil { + return false, err + } + for _, obj := range objects { fromHelm, existed, err := r.inspectExisting(ctx, obj) if err != nil { return false, err @@ -423,6 +448,40 @@ func (r *SimplyblockDriverReconciler) recordOrigin( }) } +// recordSnapshotSupport writes status.snapshotSupport, and emits §6.1's event +// the first time it becomes known. +// +// It is written once and not maintained. The field records what happened when +// this deployment came up, which is what an administrator reading it wants to +// know, and a cluster that later loses the snapshot API has a problem this +// field is not the place to report. +func (r *SimplyblockDriverReconciler) recordSnapshotSupport( + ctx context.Context, d *simplyblockv1alpha2.SimplyblockDriver, +) error { + origin, err := r.snapshotOrigin(ctx, d) + if err != nil { + return err + } + if origin == "" || origin == d.Status.SnapshotSupport { + return nil + } + + r.event(d, corev1.EventTypeNormal, reasonSnapshotsEnabled, + fmt.Sprintf("volume snapshots are available through %s, on a cluster that %s", + names(d).snapshotClass, snapshotOriginPhrase(origin))) + return r.writeStatus(ctx, d, func(status *simplyblockv1alpha2.SimplyblockDriverStatus) { + status.SnapshotSupport = origin + }) +} + +// snapshotOriginPhrase is how an event says which of the two happened. +func snapshotOriginPhrase(origin simplyblockv1alpha2.SnapshotSupportOrigin) string { + if origin == simplyblockv1alpha2.SnapshotSupportOriginInstalled { + return "had no snapshot support until this deployment installed it" + } + return "was already serving the snapshot API" +} + // applyConfiguration turns a built object into the shape a server-side apply // takes. The apiVersion and kind have to be on the wire for an apply, and a // typed object built in Go carries an empty TypeMeta, so the kind is resolved @@ -498,22 +557,24 @@ func (r *SimplyblockDriverReconciler) finalize(ctx context.Context, d *simplyblo // cluster-scoped half of the object set, and the snapshot class whether or not // the toggle currently asks for it. A class applied while snapshots were // enabled is one nothing else would ever remove. +// +// It is built here rather than filtered out of desired, because the two answer +// different questions. desired is what to apply now, and on a cluster serving no +// snapshot API that excludes the class. This is what might exist, which includes +// a class applied before the toggle was turned off or before the kinds were +// removed from under it — and a deletion that consulted the live API would skip +// exactly the object nothing else removes. func (r *SimplyblockDriverReconciler) ownedClusterScoped( d *simplyblockv1alpha2.SimplyblockDriver, ) []client.Object { - // The image does not matter here: the cluster-scoped objects do not carry - // one, and a deployment whose image cannot be resolved still has to be - // deletable. - var out []client.Object - for _, obj := range r.desired(d, "") { - if obj.GetNamespace() == "" { - out = append(out, obj) - } + out := make([]client.Object, 0, 12) + for _, cr := range clusterRoles(d) { + out = append(out, cr) } - if !snapshotsEnabled(d) { - out = append(out, volumeSnapshotClass(d)) + for _, crb := range clusterRoleBindings(d) { + out = append(out, crb) } - return out + return append(out, csiDriver(d), volumeSnapshotClass(d)) } // deploymentHolder is the object that owns the deployment: the oldest in the diff --git a/operator/internal/controllers/driver/simplyblockdriver_controller_test.go b/operator/internal/controllers/driver/simplyblockdriver_controller_test.go index 14b9187f3..ae5441e02 100644 --- a/operator/internal/controllers/driver/simplyblockdriver_controller_test.go +++ b/operator/internal/controllers/driver/simplyblockdriver_controller_test.go @@ -46,10 +46,10 @@ func requestFor(d *simplyblockv1alpha2.SimplyblockDriver) ctrl.Request { // The object set is complete: every kind the design lists, and nothing else. func TestDesiredCoversTheWholeObjectSet(t *testing.T) { d := testDriver("simplyblock") - r := &SimplyblockDriverReconciler{Scheme: reconcilerScheme(t)} + r := &SimplyblockDriverReconciler{Scheme: reconcilerScheme(t), Snapshots: servingCluster} counts := map[string]int{} - for _, obj := range r.desired(d, testImage) { + for _, obj := range desiredSet(t, r, d) { switch obj.(type) { case *corev1.ServiceAccount: counts["sa"]++ @@ -88,9 +88,9 @@ func TestDesiredCoversTheWholeObjectSet(t *testing.T) { // alternate its contents. func TestTheCredentialsSecretIsNotOwnedHere(t *testing.T) { d := testDriver("simplyblock") - r := &SimplyblockDriverReconciler{Scheme: reconcilerScheme(t)} + r := &SimplyblockDriverReconciler{Scheme: reconcilerScheme(t), Snapshots: servingCluster} - for _, obj := range r.desired(d, testImage) { + for _, obj := range desiredSet(t, r, d) { if _, isSecret := obj.(*corev1.Secret); isSecret { t.Errorf("the deployment claims Secret %s, which another controller writes", obj.GetName()) } @@ -100,27 +100,44 @@ func TestTheCredentialsSecretIsNotOwnedHere(t *testing.T) { // U-37: the snapshot class is in the set only when the deployment includes // snapshot support. func TestSnapshotClassFollowsTheToggle(t *testing.T) { - r := &SimplyblockDriverReconciler{Scheme: reconcilerScheme(t)} + r := &SimplyblockDriverReconciler{Scheme: reconcilerScheme(t), Snapshots: servingCluster} enabled := testDriver("simplyblock") disabled := testDriver("simplyblock") off := false disabled.Spec.EnableVolumeSnapshots = &off - if len(r.desired(enabled, testImage))-len(r.desired(disabled, testImage)) != 1 { - t.Errorf("the toggle changed the object set by %d, want exactly the snapshot class", - len(r.desired(enabled, testImage))-len(r.desired(disabled, testImage))) + if got := len(desiredSet(t, r, enabled)) - len(desiredSet(t, r, disabled)); got != 1 { + t.Errorf("the toggle changed the object set by %d, want exactly the snapshot class", got) } } +// servingCluster is a cluster that serves the snapshot API, which is what every +// cluster this operator adopts is: the chart installs the kinds. The cases about +// a cluster that does not are in registration_test.go. +var servingCluster = fixedSnapshotAPI{served: true} + +// desiredSet is the object set, with the error a detection failure would carry. +func desiredSet( + t *testing.T, r *SimplyblockDriverReconciler, d *simplyblockv1alpha2.SimplyblockDriver, +) []client.Object { + t.Helper() + + objects, err := r.desired(t.Context(), d, testImage) + if err != nil { + t.Fatalf("building the object set: %v", err) + } + return objects +} + // U-03 and U-62: every object comes out owned, by whichever mechanism its scope // allows, and none comes out unowned. func TestEveryAppliedObjectIsOwned(t *testing.T) { d := testDriver("simplyblock") scheme := reconcilerScheme(t) - r := &SimplyblockDriverReconciler{Scheme: scheme} + r := &SimplyblockDriverReconciler{Scheme: scheme, Snapshots: servingCluster} - for _, obj := range r.desired(d, testImage) { + for _, obj := range desiredSet(t, r, d) { if err := setOwnership(d, obj, scheme); err != nil { t.Fatalf("%T %s: %v", obj, obj.GetName(), err) } @@ -335,7 +352,7 @@ func TestAnEmptyHolderIsNotADuplicate(t *testing.T) { } c := fake.NewClientBuilder().WithScheme(scheme).WithObjects(d).WithStatusSubresource(d).Build() - r := &SimplyblockDriverReconciler{Client: c, Scheme: scheme} + r := &SimplyblockDriverReconciler{Client: c, Scheme: scheme, Snapshots: servingCluster} if _, err := r.Reconcile(context.Background(), requestFor(d)); err != nil { t.Fatalf("reconcile: %v", err) From c9961afeb05431d2beca13fd6b4db88e8e19b104 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 14:10:13 +0200 Subject: [PATCH 047/206] refactor(deployment): a discovery run walks a declared graph rather than a switch MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit design-crd-model §3.1 names the hand-rolled switch over the current step as the thing the redesign exists to remove, and OperatorOps was still one. The ordering of a run lived in the returns of three methods: inspect decided that Probing came next and set its deadline, probe decided that Writing did, write decided the run was over. Reading the order meant reading all three, and adding a step meant editing all three. It is one MultiConfig now, keyed on the action, and the reconcile decides nothing: the step to perform is where the machine is, the step that follows is the edge out of it, and a step that is not finished requeues against the same position. inspect, probe, and write report whether they finished and nothing else. The MultiConfig has one entry, which is deliberate. Every other Ops controller in this group keys its graphs on the action, and a kind that read differently for having one would make a reader check whether the difference meant something. The second action arrives as an entry rather than as a case. Five tests come with the graph, and they are most of what declaring it as data buys. The Enum marker and the CEL rule on status.step are compared against the graph in both directions, because a rule naming a step no graph declares admits a status no reconcile can resume from, and one omitting a declared step refuses a status the controller itself writes. The edge table is asserted outright. And UnabortableMultiStates is asserted empty, which is what makes the decision to give this kind no DELETE admission guard checkable rather than remembered: the guard derives its refusal table from the graph, and a step that stopped being abortable would make the absent webhook wrong. The twelve event reasons are constants. A reason is what somebody greps a cluster's events for and what an alert matches on, and a literal typed twice is two reasons nobody can tell apart from outside. One regression came out of the migration, and reading caught it rather than the suite. The machine's generic "step Probing outlived its deadline" replaced a message that told a reviewer the reports that did arrive survive in labeled ConfigMaps, and nothing tested for it. There is a test now, it was red, and discoveryTimeoutMessage gives each step its own. What is left behind differs per step, so the message does. Nothing else in the package's tests changed, which is the evidence that the rest of the behavior survived. Co-Authored-By: Claude Fable 5 --- .../deployment/operatorops_controller.go | 262 +++++++++++++----- .../deployment/operatorops_graphs.go | 141 ++++++++++ .../deployment/operatorops_graphs_test.go | 133 +++++++++ .../deployment/operatorops_unit_test.go | 25 ++ 4 files changed, 487 insertions(+), 74 deletions(-) create mode 100644 operator/internal/controllers/deployment/operatorops_graphs.go create mode 100644 operator/internal/controllers/deployment/operatorops_graphs_test.go diff --git a/operator/internal/controllers/deployment/operatorops_controller.go b/operator/internal/controllers/deployment/operatorops_controller.go index f5836993f..81a130d4a 100644 --- a/operator/internal/controllers/deployment/operatorops_controller.go +++ b/operator/internal/controllers/deployment/operatorops_controller.go @@ -26,6 +26,7 @@ package deployment import ( "context" + "errors" "fmt" "slices" "strings" @@ -45,6 +46,7 @@ import ( "github.com/simplyblock/atlas/inventory" "github.com/simplyblock/atlas/ptr" + "github.com/simplyblock/atlas/statemachine" simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" discoverypkg "github.com/simplyblock/simplyblock-operator/internal/discovery" @@ -71,6 +73,33 @@ const ( clusterNameSuffix = "-cluster" ) +// The reasons a discovery run emits. They are constants rather than literals at +// the call site because a reason is an API: it is what somebody greps a cluster's +// events for and what an alert matches on, and a literal typed twice is two +// reasons nobody can tell apart from the outside. +const ( + // The run's own lifecycle. + OperationStarted = "OperationStarted" + OperationSucceeded = "OperationSucceeded" + OperationAborted = "OperationAborted" + OperationFailed = "OperationFailed" + + // What Inspecting concluded about a worker, and about the cluster. + WorkerDeclined = "WorkerDeclined" + ControlPlaneNodeIncluded = "ControlPlaneNodeIncluded" + EnvironmentPartiallyRead = "EnvironmentPartiallyRead" + DiscoveryStepDeadlineGone = "StepDeadlineExceeded" + + // What Probing found. + DeviceInspectionFailed = "DeviceInspectionFailed" + ReportUnreadable = "ReportUnreadable" + + // What Writing produced, and what it left out. + ConfigWritten = "ConfigWritten" + ConfigExists = "ConfigExists" + DeviceDeclined = "DeviceDeclined" +) + // OperatorOpsReconciler runs operations against the operator itself. type OperatorOpsReconciler struct { client.Client @@ -149,38 +178,142 @@ func (r *OperatorOpsReconciler) Reconcile(ctx context.Context, req ctrl.Request) } log.Info("advancing discovery", "step", ops.Status.Step.State, "phase", ops.Status.Phase) + return r.advance(ctx, &ops) +} + +// advance runs the graph of the run's action forward by at most one step. +// +// The graph decides the order and the deadlines, and this decides nothing: the +// step to perform is where the machine is, the step that follows is the edge out +// of it, and a step that is not finished requeues against the same position. +// That is the whole of what replacing the switch bought — the ordering is +// declared in one place rather than spread across the returns of three methods. +func (r *OperatorOpsReconciler) advance( + ctx context.Context, ops *simplyblockv1alpha2.OperatorOps, +) (ctrl.Result, error) { + graph, declared := operatorOpsGraphs()[statemachine.Action(ops.Spec.Action)] + if !declared { + // Admission's enum refuses an unknown action, so reaching here means an + // older CRD served the object or the schema was bypassed. + return r.fail(ctx, ops, fmt.Sprintf("action %q declares no state graph", ops.Spec.Action)) + } + + machine, err := statemachine.NewFromSnapshot(ctx, graph, + statemachine.FromKube[discoveryStep](ops.Status.Step)) + if err != nil { + // An unrecognized step is a downgrade, a hand-edited object, or a rename + // that shipped without a conversion, and none of them resolve by + // reconciling again. + return r.fail(ctx, ops, fmt.Sprintf("the run cannot be resumed: %v", err)) + } + defer machine.Close() - switch simplyblockv1alpha2.OperatorOpsStep(ops.Status.Step.State) { - case "": - return r.startInspecting(ctx, &ops) - case simplyblockv1alpha2.OperatorOpsStepInspecting: - return r.inspect(ctx, &ops) - case simplyblockv1alpha2.OperatorOpsStepProbing: - return r.probe(ctx, &ops) - case simplyblockv1alpha2.OperatorOpsStepWriting: - return r.write(ctx, &ops) + // A machine is born already in its initial state, so that state's entry hook + // never runs and no deadline is set for it. Recording the birth here is what + // stops the first step being the one step that cannot time out, and it is + // also the run's own start: a crash between this write and the work is + // visible as a run that started rather than as one that never did. + if ops.Status.Step.State == "" { + return r.begin(ctx, ops, machine.CurrentState()) + } + + current := machine.CurrentState() + if machine.TimeoutReached() { + expired := discoveryTimeoutMessage(current) + r.event(ops, corev1.EventTypeWarning, DiscoveryStepDeadlineGone, expired) + return r.fail(ctx, ops, expired) + } + + done, err := r.performStep(ctx, ops, current) + if err != nil { + var refusal *refusedError + if errors.As(err, &refusal) { + r.event(ops, corev1.EventTypeWarning, refusal.reason, refusal.Error()) + return r.fail(ctx, ops, refusal.Error()) + } + return ctrl.Result{}, err + } + if !done { + // The step wrote whatever it concluded and is waiting on something + // outside this reconcile. Probing is the only one that does. + return ctrl.Result{RequeueAfter: probingRequeue}, nil + } + + if machine.IsTerminal() { + return ctrl.Result{}, r.succeed(ctx, ops) + } + + next, ok := nextDiscoveryStep(machine) + if !ok { + return r.fail(ctx, ops, fmt.Sprintf("step %s declares no successor and is not terminal", current)) + } + if err := machine.TransitionTo(ctx, next); err != nil { + return ctrl.Result{}, fmt.Errorf("enter step %s: %w", next, err) + } + return r.enter(ctx, ops, next, statemachine.ToKube(machine.Snapshot()).Deadline) +} + +// performStep runs one step and reports whether it has finished. +func (r *OperatorOpsReconciler) performStep( + ctx context.Context, ops *simplyblockv1alpha2.OperatorOps, current discoveryStep, +) (bool, error) { + switch current { + case stepInspecting: + return r.inspect(ctx, ops) + case stepProbing: + return r.probe(ctx, ops) + case stepWriting: + return r.write(ctx, ops) default: - return r.fail(ctx, &ops, fmt.Sprintf("step %q is not one this action has", ops.Status.Step.State)) + return false, fmt.Errorf("step %s belongs to no run this operator performs", current) } } -// startInspecting records that the run has begun before it does anything, so -// that a crash between the two is visible as a run that started rather than one -// that never did. -func (r *OperatorOpsReconciler) startInspecting( - ctx context.Context, - ops *simplyblockv1alpha2.OperatorOps, +// begin records that the run has started, in the step the graph begins at and +// with that step's budget. +// +// It writes before anything is done, so a crash between the two is visible as a +// run that started rather than as one that never did. +func (r *OperatorOpsReconciler) begin( + ctx context.Context, ops *simplyblockv1alpha2.OperatorOps, initial discoveryStep, ) (ctrl.Result, error) { now := metav1.Now() + deadline := metav1.NewTime(now.Add(initialDiscoveryDeadline)) ops.Status.Phase = simplyblockv1alpha2.OperatorOpsPhaseRunning ops.Status.StartedAt = &now - ops.Status.Step.State = string(simplyblockv1alpha2.OperatorOpsStepInspecting) + ops.Status.Step.State = string(initial) + ops.Status.Step.Deadline = &deadline ops.Status.Message = "reading the cluster's workers" - r.event(ops, corev1.EventTypeNormal, "OperationStarted", "discovery started") + r.event(ops, corev1.EventTypeNormal, OperationStarted, "discovery started") + + return ctrl.Result{Requeue: true}, r.status(ctx, ops) +} +// enter records the step the machine has moved into, with the deadline its +// entry hook set. +func (r *OperatorOpsReconciler) enter( + ctx context.Context, + ops *simplyblockv1alpha2.OperatorOps, + next discoveryStep, + deadline *metav1.Time, +) (ctrl.Result, error) { + ops.Status.Step.State = string(next) + ops.Status.Step.Deadline = deadline return ctrl.Result{Requeue: true}, r.status(ctx, ops) } +// succeed ends a run that reached the end of its graph. +func (r *OperatorOpsReconciler) succeed( + ctx context.Context, ops *simplyblockv1alpha2.OperatorOps, +) error { + now := metav1.Now() + ops.Status.Phase = simplyblockv1alpha2.OperatorOpsPhaseSucceeded + ops.Status.CompletedAt = &now + ops.Status.Step.Deadline = nil + r.event(ops, corev1.EventTypeNormal, OperationSucceeded, ops.Status.Message) + return r.status(ctx, ops) +} + // inspect settles what the run is about: which workers, and which distribution. // // It is one step and it persists both, because everything after it has to agree @@ -190,7 +323,7 @@ func (r *OperatorOpsReconciler) startInspecting( func (r *OperatorOpsReconciler) inspect( ctx context.Context, ops *simplyblockv1alpha2.OperatorOps, -) (ctrl.Result, error) { +) (bool, error) { spec := ops.Spec.Discover if spec == nil { spec = &simplyblockv1alpha2.DiscoverSpec{} @@ -202,12 +335,12 @@ func (r *OperatorOpsReconciler) inspect( options = append(options, client.MatchingLabels(spec.NodeSelector)) } if err := r.List(ctx, &nodes, options...); err != nil { - return ctrl.Result{}, err + return false, err } taken, err := r.workersAlreadyTaken(ctx, ops.Namespace) if err != nil { - return ctrl.Result{}, err + return false, err } // A worker that looked fine and was not used owes the run a reason. Both @@ -225,7 +358,7 @@ func (r *OperatorOpsReconciler) inspect( // before this one and its being the storage tier went unnoticed. role := discoverypkg.RoleOf(node) if !UsableWorker(node, useControlPlane) { - r.event(ops, corev1.EventTypeNormal, "WorkerDeclined", + r.event(ops, corev1.EventTypeNormal, WorkerDeclined, fmt.Sprintf("%s is not used: %s", node.Name, declinedBecause(node, role))) continue } @@ -233,7 +366,7 @@ func (r *OperatorOpsReconciler) inspect( // The reviewer asked for these and still has to see which machines // they got, because the draft's control-plane node set is otherwise // just another block of hostnames. - r.event(ops, corev1.EventTypeWarning, "ControlPlaneNodeIncluded", fmt.Sprintf( + r.event(ops, corev1.EventTypeWarning, ControlPlaneNodeIncluded, fmt.Sprintf( "%s is %s and is in the draft because spec.discover.enableControlPlaneNodes is set", node.Name, role.Describe())) } @@ -247,12 +380,12 @@ func (r *OperatorOpsReconciler) inspect( if err != nil { // Half the evidence still yields a conclusion, and the field is one a // reviewer corrects, so this is recorded and not fatal. - r.event(ops, corev1.EventTypeWarning, "EnvironmentPartiallyRead", + r.event(ops, corev1.EventTypeWarning, EnvironmentPartiallyRead, fmt.Sprintf("the API groups could not be listed, so the distribution was concluded from the nodes alone: %v", err)) } if len(workers) == 0 { - return r.fail(ctx, ops, + return false, refusef(OperationFailed, "no worker is free: every node either carries a StorageNode already, is "+ "unschedulable, is reserved for the control plane or for infrastructure, "+ "or does not match the run's selector") @@ -260,13 +393,10 @@ func (r *OperatorOpsReconciler) inspect( ops.Status.Workers = workers ops.Status.Environment = simplyblockv1alpha2.KubernetesEnvironment(environment.Distribution) - ops.Status.Step.State = string(simplyblockv1alpha2.OperatorOpsStepProbing) - deadline := metav1.NewTime(time.Now().Add(probingDeadline)) - ops.Status.Step.Deadline = &deadline ops.Status.Message = fmt.Sprintf("probing %d worker(s) of a %s cluster", len(workers), orUnknown(string(environment.Distribution))) - return ctrl.Result{Requeue: true}, r.status(ctx, ops) + return true, r.status(ctx, ops) } // workersAlreadyTaken is the set of workers a StorageNode already runs on. @@ -300,21 +430,15 @@ func (r *OperatorOpsReconciler) workersAlreadyTaken( func (r *OperatorOpsReconciler) probe( ctx context.Context, ops *simplyblockv1alpha2.OperatorOps, -) (ctrl.Result, error) { +) (bool, error) { log := logf.FromContext(ctx) - if deadline, has := ops.Status.Step.KubeDeadline(); has && time.Now().After(deadline) { - return r.fail(ctx, ops, fmt.Sprintf( - "the probes did not all finish within %s; the reports that did arrive are in the "+ - "ConfigMaps labelled for this run", probingDeadline)) - } - owner := metav1.NewControllerRef(ops, simplyblockv1alpha2.GroupVersion.WithKind("OperatorOps")) reports, err := r.reportsFor(ctx, ops) if err != nil { - return ctrl.Result{}, err + return false, err } waiting, failed := 0, []string{} @@ -332,19 +456,19 @@ func (r *OperatorOpsReconciler) probe( Owner: owner, }) if err != nil { - return r.fail(ctx, ops, fmt.Sprintf("a probe Job could not be built: %v", err)) + return false, refusef(OperationFailed, "a probe Job could not be built: %v", err) } var existing batchv1.Job switch err := r.Get(ctx, client.ObjectKeyFromObject(job), &existing); { case apierrors.IsNotFound(err): if err := r.Create(ctx, job); err != nil && !apierrors.IsAlreadyExists(err) { - return ctrl.Result{}, err + return false, err } log.Info("created a probe Job", "worker", worker, "job", job.Name) waiting++ case err != nil: - return ctrl.Result{}, err + return false, err case existing.Status.Failed > 0 && jobExhausted(&existing): // A Job that has used its retries and written no report is a // worker this run cannot describe. It does not fail the run: the @@ -359,32 +483,27 @@ func (r *OperatorOpsReconciler) probe( if waiting > 0 { ops.Status.Message = fmt.Sprintf("%d of %d worker(s) reported; waiting for %d", len(reports), len(ops.Status.Workers), waiting) - if err := r.status(ctx, ops); err != nil { - return ctrl.Result{}, err - } - return ctrl.Result{RequeueAfter: probingRequeue}, nil + return false, r.status(ctx, ops) } if len(reports) == 0 { - return r.fail(ctx, ops, fmt.Sprintf( - "no worker reported: every one of the %d probe Jobs failed", len(ops.Status.Workers))) + return false, refusef(OperationFailed, + "no worker reported: every one of the %d probe Jobs failed", len(ops.Status.Workers)) } for _, worker := range failed { - r.event(ops, corev1.EventTypeWarning, "DeviceInspectionFailed", + r.event(ops, corev1.EventTypeWarning, DeviceInspectionFailed, fmt.Sprintf("the probe on %s failed, so it is not in the draft", worker)) } - ops.Status.Step.State = string(simplyblockv1alpha2.OperatorOpsStepWriting) - ops.Status.Step.Deadline = nil ops.Status.Message = fmt.Sprintf("%d worker(s) reported; writing the draft", len(reports)) - return ctrl.Result{Requeue: true}, r.status(ctx, ops) + return true, r.status(ctx, ops) } // write turns the reports into a ClusterDeploymentConfig in Draft. func (r *OperatorOpsReconciler) write( ctx context.Context, ops *simplyblockv1alpha2.OperatorOps, -) (ctrl.Result, error) { +) (bool, error) { spec := ops.Spec.Discover if spec == nil { spec = &simplyblockv1alpha2.DiscoverSpec{} @@ -392,10 +511,11 @@ func (r *OperatorOpsReconciler) write( reports, err := r.reportsFor(ctx, ops) if err != nil { - return ctrl.Result{}, err + return false, err } if len(reports) == 0 { - return r.fail(ctx, ops, "the probe reports are gone, so there is nothing to write") + return false, refusef(OperationFailed, + "the probe reports are gone, so there is nothing to write") } collected := make([]nodeprobe.Report, 0, len(reports)) @@ -410,7 +530,7 @@ func (r *OperatorOpsReconciler) write( // memory is reporting something truer than a copy taken minutes ago. kubeNodes, err := r.kubeNodesFor(ctx, ops) if err != nil { - return ctrl.Result{}, err + return false, err } filter := spec.DeviceFilter @@ -428,48 +548,42 @@ func (r *OperatorOpsReconciler) write( // out as events, and the worker-level ones go into the message as well, // because the message is what `kubectl get operatorops` shows. for _, refusal := range plan.RefusalLines() { - r.event(ops, corev1.EventTypeNormal, "DeviceDeclined", refusal) + r.event(ops, corev1.EventTypeNormal, DeviceDeclined, refusal) } why := plan.Explain() if len(why) == 0 { // No machine was refused by name, so the run had no worker to refuse. - return r.fail(ctx, ops, fmt.Sprintf( - "no worker has a device this run would use: %s", plan.Summary())) + return false, refusef(OperationFailed, + "no worker has a device this run would use: %s", plan.Summary()) } - return r.fail(ctx, ops, fmt.Sprintf( - "no worker has a device this run would use: %s", strings.Join(why, "; "))) + return false, refusef(OperationFailed, + "no worker has a device this run would use: %s", strings.Join(why, "; ")) } config, notes := r.draftFor(ops, spec, plan) if err := r.Create(ctx, config); err != nil { if !apierrors.IsAlreadyExists(err) { - return ctrl.Result{}, err + return false, err } // A run whose name was reused, or one that crashed after creating the // document and before recording it. The document is what matters and // it exists, so the run adopts it rather than failing. - r.event(ops, corev1.EventTypeNormal, "ConfigExists", + r.event(ops, corev1.EventTypeNormal, ConfigExists, fmt.Sprintf("%s already existed and was left as it is", config.Name)) } for _, note := range notes { - r.event(ops, corev1.EventTypeNormal, "ConfigWritten", note) + r.event(ops, corev1.EventTypeNormal, ConfigWritten, note) } - r.event(ops, corev1.EventTypeNormal, "ConfigWritten", + r.event(ops, corev1.EventTypeNormal, ConfigWritten, fmt.Sprintf("wrote %s in Draft: %s", config.Name, plan.Summary())) for _, refusal := range plan.RefusalLines() { - r.event(ops, corev1.EventTypeNormal, "DeviceDeclined", refusal) + r.event(ops, corev1.EventTypeNormal, DeviceDeclined, refusal) } - now := metav1.Now() ops.Status.ConfigRef = config.Name - ops.Status.Phase = simplyblockv1alpha2.OperatorOpsPhaseSucceeded - ops.Status.CompletedAt = &now - ops.Status.Step.Deadline = nil ops.Status.Message = fmt.Sprintf("wrote %s awaiting approval: %s", config.Name, plan.Summary()) - r.event(ops, corev1.EventTypeNormal, "OperationSucceeded", ops.Status.Message) - - return ctrl.Result{}, r.status(ctx, ops) + return true, r.status(ctx, ops) } // draftFor builds the document, and the notes explaining the numbers in it that @@ -566,7 +680,7 @@ func (r *OperatorOpsReconciler) reportsFor( for i := range maps.Items { report, err := nodeprobe.ReportFromConfigMap(&maps.Items[i]) if err != nil { - r.event(ops, corev1.EventTypeWarning, "ReportUnreadable", err.Error()) + r.event(ops, corev1.EventTypeWarning, ReportUnreadable, err.Error()) continue } if !slices.Contains(ops.Status.Workers, report.Node) { @@ -599,7 +713,7 @@ func (r *OperatorOpsReconciler) abort( ops.Status.CompletedAt = &now ops.Status.Step.Deadline = nil ops.Status.Message = "aborted; discovery changes nothing, so nothing was undone" - r.event(ops, corev1.EventTypeNormal, "OperationAborted", ops.Status.Message) + r.event(ops, corev1.EventTypeNormal, OperationAborted, ops.Status.Message) return ctrl.Result{}, r.status(ctx, ops) } @@ -641,7 +755,7 @@ func (r *OperatorOpsReconciler) fail( ops.Status.CompletedAt = &now ops.Status.Step.Deadline = nil ops.Status.Message = reason - r.event(ops, corev1.EventTypeWarning, "OperationFailed", reason) + r.event(ops, corev1.EventTypeWarning, OperationFailed, reason) return ctrl.Result{}, r.status(ctx, ops) } diff --git a/operator/internal/controllers/deployment/operatorops_graphs.go b/operator/internal/controllers/deployment/operatorops_graphs.go new file mode 100644 index 000000000..bcd2cdfaf --- /dev/null +++ b/operator/internal/controllers/deployment/operatorops_graphs.go @@ -0,0 +1,141 @@ +// The state graph of each OperatorOps action, declared as data. +// +// Inspecting ──► Probing ──► Writing +// +// The line is the run's ordering and the ordering is the point of it. +// Inspecting settles which workers the run is about, so a node joining while the +// probes run does not silently join the draft; Probing reads only those workers; +// Writing turns what they reported into a document. An edge that skipped a step +// would draft a fleet nobody listed, and declaring the edges rather than +// switching on the current step is what makes that a compile-time shape instead +// of a review comment. +// +// It is a MultiConfig although the kind carries one action. Every other Ops +// controller in this group keys its graphs on the action, and a kind that read +// differently for having one entry would make a reader check whether the +// difference meant something. A MultiConfig with one entry costs nothing, and +// the second action arrives as an entry rather than as a case. +// +// design-crd-model.md §3.1 is the specification. + +package deployment + +import ( + "context" + "fmt" + "time" + + "github.com/simplyblock/atlas/statemachine" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// discoveryStep is the run's step type, aliased so the graph literal reads as +// the graph rather than as a wall of package qualifiers. +type discoveryStep = simplyblockv1alpha2.OperatorOpsStep + +const ( + stepInspecting = simplyblockv1alpha2.OperatorOpsStepInspecting + stepProbing = simplyblockv1alpha2.OperatorOpsStepProbing + stepWriting = simplyblockv1alpha2.OperatorOpsStepWriting +) + +// actionDiscover is the MultiConfig key for the one action this kind carries. +const actionDiscover = statemachine.Action(simplyblockv1alpha2.OperatorOpsActionDiscover) + +// How long each step may take before the run is reported as stuck. +// +// They differ by what the step waits on. Inspecting and Writing are Kubernetes +// reads and one create, so neither can be slow for a reason worth waiting out. +// Probing waits on one Job per worker, which is where the time goes. +const ( + // inspectingDeadline bounds listing the cluster's nodes and the + // StorageNodes already on them. A step still waiting after this is one + // whose API server is not answering, which reconciling again does not fix + // any faster than reporting it. + inspectingDeadline = 5 * time.Minute + + // writingDeadline bounds turning the reports into a document: reading the + // ConfigMaps the probes wrote, and one create. + writingDeadline = 5 * time.Minute +) + +// operatorOpsGraphs declares the state graph of each action. +// +// Every step is abortable, because a discovery run changes nothing it would +// have to take back: it reads nodes, creates Jobs that carry its own owner +// reference, and creates one document at the very end. That is what lets the +// kind go without the DELETE admission guard of design-crd-model.md §3.1 — the +// guard derives its refusal table from the graph, and this graph refuses +// nothing. +func operatorOpsGraphs() statemachine.MultiConfig[discoveryStep] { + return statemachine.MultiConfig[discoveryStep]{ + actionDiscover: { + Initial: stepInspecting, + States: map[discoveryStep]statemachine.StateDef[discoveryStep]{ + stepInspecting: { + To: []discoveryStep{stepProbing}, + Abortable: true, + OnEnter: discoveryDeadline(inspectingDeadline), + }, + stepProbing: { + To: []discoveryStep{stepWriting}, + Abortable: true, + OnEnter: discoveryDeadline(probingDeadline), + }, + stepWriting: { + Abortable: true, + OnEnter: discoveryDeadline(writingDeadline), + }, + }, + }, + } +} + +// discoveryDeadline is the entry hook every state here carries: it sets the +// step's budget and performs nothing. The work of a step happens on the pass +// that follows, against the step the entry's write persisted, which is what +// makes a crash between the two resumable rather than invisible. +func discoveryDeadline(d time.Duration) statemachine.TransitionFunc[discoveryStep] { + return func(context.Context, discoveryStep, discoveryStep) (time.Duration, error) { + return d, nil + } +} + +// initialDiscoveryDeadline is the budget of the step every run is born in. +// +// A machine is already in its initial state when it is built, so that state's +// OnEnter never runs and the graph's deadline for it is never set. Setting it +// explicitly is what stops the first step from being the one step that cannot +// time out. +const initialDiscoveryDeadline = inspectingDeadline + +// discoveryTimeoutMessage says what a step outliving its deadline means for this +// run, rather than only that it happened. +// +// It is per step because what is left behind differs. Probing leaves the reports +// that did arrive in ConfigMaps labeled for the run, and a message that did not +// name them would leave a reviewer with the evidence sitting in the cluster and +// no way to know it is there. +func discoveryTimeoutMessage(step discoveryStep) string { + switch step { + case stepInspecting: + return fmt.Sprintf("the cluster could not be read within %s", inspectingDeadline) + case stepProbing: + return fmt.Sprintf("the probes did not all finish within %s; the reports that did "+ + "arrive are in the ConfigMaps labeled for this run", probingDeadline) + case stepWriting: + return fmt.Sprintf("the draft could not be written within %s", writingDeadline) + default: + return fmt.Sprintf("step %s outlived its deadline", step) + } +} + +// nextDiscoveryStep is the step that follows the current one. The graph is a +// line, so the first edge is the only edge. +func nextDiscoveryStep(machine *statemachine.Machine[discoveryStep]) (discoveryStep, bool) { + for next := range machine.AllowedTransitions() { + return next, true + } + return machine.CurrentState(), false +} diff --git a/operator/internal/controllers/deployment/operatorops_graphs_test.go b/operator/internal/controllers/deployment/operatorops_graphs_test.go new file mode 100644 index 000000000..c470dfe0f --- /dev/null +++ b/operator/internal/controllers/deployment/operatorops_graphs_test.go @@ -0,0 +1,133 @@ +// The discovery graph as data, and the three places that have to agree about +// what its steps are. +// +// A shared statemachine.KubeSnapshot cannot carry an Enum marker, so the step +// values live in the graph, in OperatorOpsStep's own Enum marker, and in the CEL +// rule on status.step — and nothing but a test makes the three agree +// (design-crd-model.md §3.1). These are that test, and they are the reason the +// graph is worth declaring as data at all: a transition table is cheap to +// exercise exhaustively, so an illegal edge is a unit test rather than a review +// comment. + +package deployment + +import ( + "slices" + "strings" + "testing" + + "github.com/google/go-cmp/cmp" + + "github.com/simplyblock/atlas/statemachine" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// everyDiscoveryStep is the Enum marker's list, transcribed. It is written out +// rather than derived so the assertion compares two independent statements of +// the same set: deriving it from the graph would make the test agree with +// itself. +var everyDiscoveryStep = []string{"Inspecting", "Probing", "Writing"} + +// discoveryStepCELRule is the rule as OperatorOpsStatus declares it. Keeping a +// copy here is the only way to compare it against anything: it is a literal in a +// struct tag, and no marker reaches a field of a type another module declares. +const discoveryStepCELRule = "!has(self.state) || self.state in ['Inspecting','Probing','Writing']" + +func TestTheDiscoveryStepEnumCoversEveryDeclaredState(t *testing.T) { + declared := statemachine.DeclaredMultiStates(operatorOpsGraphs()) + want := slices.Clone(everyDiscoveryStep) + slices.Sort(want) + if diff := cmp.Diff(want, declared); diff != "" { + t.Errorf("the graph and the Enum marker disagree (-marker +graph):\n%s", diff) + } +} + +// The rule is compared both ways. One that named a step no graph declares would +// admit a status no reconcile can resume from, and one that omitted a declared +// step would refuse a status the controller itself writes. +func TestTheDiscoveryCELRuleCoversEveryDeclaredState(t *testing.T) { + declared := statemachine.DeclaredMultiStates(operatorOpsGraphs()) + for _, state := range declared { + if !strings.Contains(discoveryStepCELRule, "'"+state+"'") { + t.Errorf("status.step's CEL rule does not accept the declared step %q", state) + } + } + for _, named := range quotedValues(discoveryStepCELRule) { + if !slices.Contains(declared, named) { + t.Errorf("status.step's CEL rule accepts %q, which no graph declares", named) + } + } +} + +// quotedValues reads the quoted values out of a rule's `in` list. +func quotedValues(rule string) []string { + var values []string + for _, part := range strings.Split(rule, "'") { + if part != "" && !strings.ContainsAny(part, "[],| ") { + values = append(values, part) + } + } + return values +} + +// Every action the API accepts needs a graph, or an operation of that action +// fails at its first pass rather than doing anything. +func TestEveryDiscoveryActionDeclaresAGraph(t *testing.T) { + declared := operatorOpsGraphs() + for _, action := range []simplyblockv1alpha2.OperatorOpsAction{ + simplyblockv1alpha2.OperatorOpsActionDiscover, + } { + if _, ok := declared[statemachine.Action(action)]; !ok { + t.Errorf("the API accepts action %q and no graph declares it", action) + } + } + for action := range declared { + if string(action) != string(simplyblockv1alpha2.OperatorOpsActionDiscover) { + t.Errorf("a graph declares action %q, which the API does not accept", action) + } + } +} + +// The graph is a line, and the line is what the run's ordering rests on: +// Inspecting settles which workers the run is about, Probing reads only those, +// and Writing turns what they reported into a document. An edge that skipped +// Inspecting would draft a fleet nobody listed. +func TestTheDiscoveryGraphIsTheLineTheRunWalks(t *testing.T) { + graph, ok := operatorOpsGraphs()[actionDiscover] + if !ok { + t.Fatal("no graph is declared for Discover") + } + + if graph.Initial != stepInspecting { + t.Errorf("the run begins at %q, want Inspecting", graph.Initial) + } + + want := map[simplyblockv1alpha2.OperatorOpsStep][]simplyblockv1alpha2.OperatorOpsStep{ + stepInspecting: {stepProbing}, + stepProbing: {stepWriting}, + stepWriting: nil, + } + for from, to := range want { + state, declared := graph.States[from] + if !declared { + t.Errorf("step %q is not in the graph", from) + continue + } + if diff := cmp.Diff(to, state.To); diff != "" { + t.Errorf("the edges out of %q disagree (-want +graph):\n%s", from, diff) + } + } +} + +// Every step of a discovery run is abortable, because the run changes nothing it +// has to take back: it reads nodes, creates Jobs that carry its owner reference, +// and creates one document at the very end. That is why the kind gets no DELETE +// admission guard (design-crd-model.md §3.1), and the guard's absence is only +// correct for as long as this holds. +func TestEveryDiscoveryStepIsAbortable(t *testing.T) { + if unabortable := statemachine.UnabortableMultiStates(operatorOpsGraphs()); len(unabortable) > 0 { + t.Errorf("steps %v cannot be aborted, so a delete arriving in one of them has "+ + "something to unwind and there is no webhook refusing it", unabortable) + } +} diff --git a/operator/internal/controllers/deployment/operatorops_unit_test.go b/operator/internal/controllers/deployment/operatorops_unit_test.go index 436e58207..9d84a4f47 100644 --- a/operator/internal/controllers/deployment/operatorops_unit_test.go +++ b/operator/internal/controllers/deployment/operatorops_unit_test.go @@ -19,6 +19,7 @@ import ( "slices" "strings" "testing" + "time" batchv1 "k8s.io/api/batch/v1" corev1 "k8s.io/api/core/v1" @@ -782,3 +783,27 @@ func TestARunRefusedDeviceByDeviceSaysWhichReasonMatters(t *testing.T) { t.Errorf("the failure reports devices that were never candidates: %s", ops.Status.Message) } } + +// A step that outlives its deadline says what outliving it means for this run, +// not only that it happened. Probing's partial reports survive in ConfigMaps +// labeled for the run, and a failure message that did not say so would leave a +// reviewer with the evidence in the cluster and no way to know it is there. +func TestARunThatOutlivesItsDeadlineSaysWhereTheEvidenceIs(t *testing.T) { + ops := discoverRun(nil) + ops.Status.Phase = simplyblockv1alpha2.OperatorOpsPhaseRunning + ops.Status.Workers = []string{"worker-1"} + ops.Status.Step.State = string(simplyblockv1alpha2.OperatorOpsStepProbing) + expired := metav1.NewTime(time.Now().Add(-time.Minute)) + ops.Status.Step.Deadline = &expired + + r := newRunner(t, ops, worker("worker-1")) + _, got := r.step() + + if got.Status.Phase != simplyblockv1alpha2.OperatorOpsPhaseFailed { + t.Fatalf("phase = %q, want Failed on a step past its deadline", got.Status.Phase) + } + if !strings.Contains(got.Status.Message, "ConfigMaps") { + t.Errorf("the failure does not say where the reports that did arrive are: %q", + got.Status.Message) + } +} From 0fcb5e6bcd431f90c070d6c9031917656ed03844 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 14:18:24 +0200 Subject: [PATCH 048/206] fix(api): the two rebalancing enums are spelled the way every other one is MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit v1alpha2 spells enum members in PascalCase, and two carried their v1alpha1 spelling across: BaselineStrategy as benchmark and rollingWindow, BaselineColdStartPolicy as defer and partialWindow. They were the only two errors check-crds.py reported across seventeen kinds, and the window to fix them is now — an enum member is part of the schema, so recasing one after a release is a change that refuses objects people have already written. The recase is the small half. The load-bearing half is that both fields were converted by a straight cast, so recasing alone would have carried "rollingWindow" into a version whose schema admits "RollingWindow". The API server accepts a stored object it is not asked to write, so nothing would have failed until the next write — and §24's storage rewrite is exactly such a write, which means the object that failed would have been one nobody edited. They are mapped now, in both directions, through the same mapOrPassThrough and invertStringMap pattern MetricsBackend already uses for the same reason: the inverse is derived rather than written, so the two directions cannot come to disagree about a value. A value neither version declares passes through rather than emptying itself, because a field that silently became its default would turn a typo into a working configuration nobody asked for, and the schema is what refuses the typo. The tests were red the moment the recase landed and green when the maps did: both members of each enum convert and round-trip back to the spelling they went in as, and an undeclared value survives the journey. Comments and test messages that named these as API values were recased with them. The Go identifiers were not — rollingWindowBaselineProvider names the implementation rather than the field's value, and renaming it would be a different change with no schema in it. One residue is left deliberately, and it belongs to §19.10's fourth preflight check rather than here: an object already stored as v1alpha2 with a lowercase value is refused on its next write. Only a development cluster can hold one, since v1alpha2 has no release, and finding such objects before an upgrade is exactly what that check exists for. Co-Authored-By: Claude Fable 5 --- ...torage.simplyblock.io_storageclusters.yaml | 14 ++-- .../api/v1alpha1/storagecluster_conversion.go | 37 ++++++++-- .../storagecluster_conversion_test.go | 69 +++++++++++++++++++ operator/api/v1alpha2/storagecluster_types.go | 18 ++--- ...torage.simplyblock.io_storageclusters.yaml | 14 ++-- operator/dist/install.yaml | 14 ++-- .../autoplacement/baseline_strategy.go | 14 ++-- operator/internal/autoplacement/metrics.go | 4 +- operator/internal/autoplacement/utils.go | 17 ++--- operator/internal/autoplacement/utils_test.go | 8 +-- ...torage.simplyblock.io_storageclusters.yaml | 14 ++-- ...plyblock_volume_placement_injector_test.go | 7 +- 12 files changed, 165 insertions(+), 65 deletions(-) diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusters.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusters.yaml index 23925fce6..e1f616c64 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusters.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusters.yaml @@ -1174,10 +1174,10 @@ spec: baselineColdStart: description: |- BaselineColdStart selects what happens to an under-sampled node. Defaults - to partialWindow. + to PartialWindow. enum: - - defer - - partialWindow + - Defer + - PartialWindow type: string baselineMinSamples: description: |- @@ -1194,14 +1194,14 @@ spec: baselineStrategy: description: |- BaselineStrategy selects how the per-node baseline is derived. Defaults - to rollingWindow. + to RollingWindow. enum: - - benchmark - - rollingWindow + - Benchmark + - RollingWindow type: string baselineWindow: description: |- - BaselineWindow is the look-back the rollingWindow strategy reduces. + BaselineWindow is the look-back the RollingWindow strategy reduces. Defaults to 6h. type: string defaultCoolDownSeconds: diff --git a/operator/api/v1alpha1/storagecluster_conversion.go b/operator/api/v1alpha1/storagecluster_conversion.go index 4f5b132c3..07e9cd7f6 100644 --- a/operator/api/v1alpha1/storagecluster_conversion.go +++ b/operator/api/v1alpha1/storagecluster_conversion.go @@ -140,6 +140,31 @@ var metricsBackendToHub = map[string]string{ // about a value. var metricsBackendFromHub = invertStringMap(metricsBackendToHub) +// baselineStrategyToHub and baselineColdStartToHub recase the two rebalancing +// enums, for the reason metricsBackendToHub recases its own: both are +// user-authored and appear in deployment manifests, and v1alpha2 spells every +// enum member in PascalCase. +// +// A cast would be the bug this exists to prevent. It would carry "rollingWindow" +// into a version whose schema admits "RollingWindow," which the API server +// refuses on the next write — and §24's storage rewrite is exactly such a write, +// so an object nobody edited would be the one that failed. +var baselineStrategyToHub = map[string]string{ + "benchmark": string(v1alpha2.BaselineStrategyBenchmark), + "rollingWindow": string(v1alpha2.BaselineStrategyRollingWindow), +} + +// baselineStrategyFromHub is the inverse, derived so the two cannot disagree. +var baselineStrategyFromHub = invertStringMap(baselineStrategyToHub) + +var baselineColdStartToHub = map[string]string{ + "defer": string(v1alpha2.BaselineColdStartDefer), + "partialWindow": string(v1alpha2.BaselineColdStartPartialWindow), +} + +// baselineColdStartFromHub is the inverse, derived so the two cannot disagree. +var baselineColdStartFromHub = invertStringMap(baselineColdStartToHub) + // ConvertTo converts this StorageCluster to the v1alpha2 hub. func (src *StorageCluster) ConvertTo(dstRaw conversion.Hub) error { dst := dstRaw.(*v1alpha2.StorageCluster) @@ -708,10 +733,12 @@ func autoPlacementToHub(s *VolumeAutoPlacementSettings) *v1alpha2.VolumeAutoPlac mapOrPassThrough(metricsBackendToHub, string(*b)))) } if b := s.BaselineStrategy; b != nil { - out.BaselineStrategy = ptr.To(v1alpha2.BaselineStrategy(*b)) + out.BaselineStrategy = ptr.To(v1alpha2.BaselineStrategy( + mapOrPassThrough(baselineStrategyToHub, string(*b)))) } if c := s.BaselineColdStart; c != nil { - out.BaselineColdStart = ptr.To(v1alpha2.BaselineColdStartPolicy(*c)) + out.BaselineColdStart = ptr.To(v1alpha2.BaselineColdStartPolicy( + mapOrPassThrough(baselineColdStartToHub, string(*c)))) } return &out } @@ -742,10 +769,12 @@ func autoPlacementFromHub(s *v1alpha2.VolumeAutoPlacementSettings) *VolumeAutoPl mapOrPassThrough(metricsBackendFromHub, string(*b)))) } if b := s.BaselineStrategy; b != nil { - out.BaselineStrategy = ptr.To(BaselineStrategy(*b)) + out.BaselineStrategy = ptr.To(BaselineStrategy( + mapOrPassThrough(baselineStrategyFromHub, string(*b)))) } if c := s.BaselineColdStart; c != nil { - out.BaselineColdStart = ptr.To(BaselineColdStartPolicy(*c)) + out.BaselineColdStart = ptr.To(BaselineColdStartPolicy( + mapOrPassThrough(baselineColdStartFromHub, string(*c)))) } return &out } diff --git a/operator/api/v1alpha1/storagecluster_conversion_test.go b/operator/api/v1alpha1/storagecluster_conversion_test.go index e82783f22..766fd27e3 100644 --- a/operator/api/v1alpha1/storagecluster_conversion_test.go +++ b/operator/api/v1alpha1/storagecluster_conversion_test.go @@ -667,3 +667,72 @@ func TestStorageClusterARemovedFieldsNoteGoesWhenItsBlockDoes(t *testing.T) { } } } + +// The baseline enums were recased for v1alpha2 the way MetricsBackend was, so +// conversion has to map them rather than cast them. A cast would carry +// "rollingWindow" into a version whose Enum marker admits "RollingWindow," which +// the API server refuses on the next write — and the storage rewrite's unchanged +// write is exactly such a write. +func TestBaselineEnumsAreRecasedBothWays(t *testing.T) { + for _, tc := range []struct { + name string + strategy BaselineStrategy + coldStart BaselineColdStartPolicy + wantStrategy v1alpha2.BaselineStrategy + wantColdStart v1alpha2.BaselineColdStartPolicy + }{ + { + name: "the defaults", + strategy: BaselineStrategyRollingWindow, + coldStart: BaselineColdStartPartialWindow, + wantStrategy: v1alpha2.BaselineStrategyRollingWindow, + wantColdStart: v1alpha2.BaselineColdStartPartialWindow, + }, + { + name: "the other member of each", + strategy: BaselineStrategyBenchmark, + coldStart: BaselineColdStartDefer, + wantStrategy: v1alpha2.BaselineStrategyBenchmark, + wantColdStart: v1alpha2.BaselineColdStartDefer, + }, + } { + t.Run(tc.name, func(t *testing.T) { + from := &VolumeAutoPlacementSettings{ + BaselineStrategy: ptr.To(tc.strategy), + BaselineColdStart: ptr.To(tc.coldStart), + } + + hub := autoPlacementToHub(from) + if got := *hub.BaselineStrategy; got != tc.wantStrategy { + t.Errorf("baselineStrategy converted to %q, want %q", got, tc.wantStrategy) + } + if got := *hub.BaselineColdStart; got != tc.wantColdStart { + t.Errorf("baselineColdStart converted to %q, want %q", got, tc.wantColdStart) + } + + // And back, because a round trip that did not restore the spelling + // would make the storage rewrite look like an edit. + back := autoPlacementFromHub(hub) + if got := *back.BaselineStrategy; got != tc.strategy { + t.Errorf("baselineStrategy came back as %q, want the %q it went in as", + got, tc.strategy) + } + if got := *back.BaselineColdStart; got != tc.coldStart { + t.Errorf("baselineColdStart came back as %q, want the %q it went in as", + got, tc.coldStart) + } + }) + } +} + +// A value neither version declares passes through rather than being dropped, the +// same way MetricsBackend's does. A field that silently emptied itself would +// turn a typo into the default, and the schema is what refuses the typo. +func TestAnUndeclaredBaselineValuePassesThrough(t *testing.T) { + hub := autoPlacementToHub(&VolumeAutoPlacementSettings{ + BaselineStrategy: ptr.To(BaselineStrategy("whatever")), + }) + if got := string(*hub.BaselineStrategy); got != "whatever" { + t.Errorf("an unrecognized value converted to %q, want it carried through", got) + } +} diff --git a/operator/api/v1alpha2/storagecluster_types.go b/operator/api/v1alpha2/storagecluster_types.go index bbdd0e807..f6141ac88 100644 --- a/operator/api/v1alpha2/storagecluster_types.go +++ b/operator/api/v1alpha2/storagecluster_types.go @@ -207,7 +207,7 @@ const ( // BaselineStrategy selects how the per-node latency baseline, the denominator // of the rebalancing deviation signal, is derived. -// +kubebuilder:validation:Enum=benchmark;rollingWindow +// +kubebuilder:validation:Enum=Benchmark;RollingWindow type BaselineStrategy string const ( @@ -215,31 +215,31 @@ const ( // fresh cluster and frozen on the node's status. It is simple and tends to // read too low, because an idle cluster is far faster than a loaded one and // every loaded node then shows a large deviation. - BaselineStrategyBenchmark BaselineStrategy = "benchmark" + BaselineStrategyBenchmark BaselineStrategy = "Benchmark" // BaselineStrategyRollingWindow derives the baseline from a rolling window // of the probe sidecar's latency series, using an outlier-rejecting // estimator. It reflects each node's recent operating latency rather than // an idle measurement, and is the default. - BaselineStrategyRollingWindow BaselineStrategy = "rollingWindow" + BaselineStrategyRollingWindow BaselineStrategy = "RollingWindow" ) // BaselineColdStartPolicy selects what happens to a node with fewer than // BaselineMinSamples samples in the window: a freshly onboarded node, or one // whose probe sidecar has just started. -// +kubebuilder:validation:Enum=defer;partialWindow +// +kubebuilder:validation:Enum=Defer;PartialWindow type BaselineColdStartPolicy string const ( // BaselineColdStartDefer omits an under-sampled node from the cycle: it is // neither a migration source nor a target until it has accumulated enough // samples, which avoids acting on a noisy baseline. - BaselineColdStartDefer BaselineColdStartPolicy = "defer" + BaselineColdStartDefer BaselineColdStartPolicy = "Defer" // BaselineColdStartPartialWindow computes the baseline from whatever // samples exist, accepting a noisier baseline early on so that rebalancing // engages sooner. It is the default. - BaselineColdStartPartialWindow BaselineColdStartPolicy = "partialWindow" + BaselineColdStartPartialWindow BaselineColdStartPolicy = "PartialWindow" ) // DataRealignmentSettings tunes the post-migration control-plane data @@ -358,17 +358,17 @@ type VolumeAutoPlacementSettings struct { LatencyBenchmarkInterval *metav1.Duration `json:"latencyBenchmarkInterval,omitempty"` // BaselineStrategy selects how the per-node baseline is derived. Defaults - // to rollingWindow. + // to RollingWindow. // +optional BaselineStrategy *BaselineStrategy `json:"baselineStrategy,omitempty"` - // BaselineWindow is the look-back the rollingWindow strategy reduces. + // BaselineWindow is the look-back the RollingWindow strategy reduces. // Defaults to 6h. // +optional BaselineWindow *metav1.Duration `json:"baselineWindow,omitempty"` // BaselineColdStart selects what happens to an under-sampled node. Defaults - // to partialWindow. + // to PartialWindow. // +optional BaselineColdStart *BaselineColdStartPolicy `json:"baselineColdStart,omitempty"` diff --git a/operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml b/operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml index 23925fce6..e1f616c64 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml @@ -1174,10 +1174,10 @@ spec: baselineColdStart: description: |- BaselineColdStart selects what happens to an under-sampled node. Defaults - to partialWindow. + to PartialWindow. enum: - - defer - - partialWindow + - Defer + - PartialWindow type: string baselineMinSamples: description: |- @@ -1194,14 +1194,14 @@ spec: baselineStrategy: description: |- BaselineStrategy selects how the per-node baseline is derived. Defaults - to rollingWindow. + to RollingWindow. enum: - - benchmark - - rollingWindow + - Benchmark + - RollingWindow type: string baselineWindow: description: |- - BaselineWindow is the look-back the rollingWindow strategy reduces. + BaselineWindow is the look-back the RollingWindow strategy reduces. Defaults to 6h. type: string defaultCoolDownSeconds: diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index 4e4563fb2..e00b607ab 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -6148,10 +6148,10 @@ spec: baselineColdStart: description: |- BaselineColdStart selects what happens to an under-sampled node. Defaults - to partialWindow. + to PartialWindow. enum: - - defer - - partialWindow + - Defer + - PartialWindow type: string baselineMinSamples: description: |- @@ -6168,14 +6168,14 @@ spec: baselineStrategy: description: |- BaselineStrategy selects how the per-node baseline is derived. Defaults - to rollingWindow. + to RollingWindow. enum: - - benchmark - - rollingWindow + - Benchmark + - RollingWindow type: string baselineWindow: description: |- - BaselineWindow is the look-back the rollingWindow strategy reduces. + BaselineWindow is the look-back the RollingWindow strategy reduces. Defaults to 6h. type: string defaultCoolDownSeconds: diff --git a/operator/internal/autoplacement/baseline_strategy.go b/operator/internal/autoplacement/baseline_strategy.go index 72b004868..cadcf438a 100644 --- a/operator/internal/autoplacement/baseline_strategy.go +++ b/operator/internal/autoplacement/baseline_strategy.go @@ -19,8 +19,8 @@ type BaselineProvider interface { } // newBaselineProvider selects the BaselineProvider implementation for cfg.BaselineStrategy. -// "benchmark" reads the frozen one-shot fio measurement from the StorageNodeSet CRs; -// rollingWindow (the default, and the fallback for any unrecognised value) derives a robust +// "Benchmark" reads the frozen one-shot fio measurement from the StorageNodeSet CRs; +// "RollingWindow" (the default, and the fallback for any unrecognized value) derives a robust // estimate from a rolling window of the probe latency series in Prometheus. func newBaselineProvider(k8sClient client.Client, cfg RebalancingConfig) (BaselineProvider, error) { if cfg.BaselineStrategy == string(simplyblockv1alpha2.BaselineStrategyBenchmark) { @@ -49,7 +49,7 @@ func (b *benchmarkBaselineProvider) BaselineNS( var snodeList simplyblockv1alpha1.StorageNodeSetList if err := b.client.List(ctx, &snodeList, client.InNamespace(input.Namespace)); err != nil { // Stay resilient to a transient list error: skip this namespace rather than - // failing the whole evaluation cycle (matches the previous CR-read behaviour). + // failing the whole evaluation cycle (matches the previous CR-read behavior). continue } for _, snode := range snodeList.Items { @@ -108,8 +108,8 @@ type nodeBaseline struct { // reduceWindowedBaselines reduces per-node windowed samples to a single robust baseline each, // applying the cold-start policy. It is pure (no Prometheus, no metrics) so the cold-start -// and estimator behaviour can be tested directly. A node is dropped when it is under-sampled -// under the "defer" policy, or when no positive baseline can be computed from its samples. +// and estimator behavior can be tested directly. A node is dropped when it is under-sampled +// under the "Defer" policy, or when no positive baseline can be computed from its samples. func reduceWindowedBaselines( windowed map[string]map[string][]float64, cfg RebalancingConfig, @@ -119,8 +119,8 @@ func reduceWindowedBaselines( var out []nodeBaseline for clusterUUID, byNode := range windowed { for nodeUUID, samples := range byNode { - // Cold start: an under-sampled node is either skipped ("defer") or computed - // from whatever samples exist ("partialWindow"). + // Cold start: an under-sampled node is either skipped ("Defer") or computed + // from whatever samples exist ("PartialWindow"). if len(samples) < cfg.BaselineMinSamples && deferUnderSampled { continue } diff --git a/operator/internal/autoplacement/metrics.go b/operator/internal/autoplacement/metrics.go index a1c9895e1..bbaea10dd 100644 --- a/operator/internal/autoplacement/metrics.go +++ b/operator/internal/autoplacement/metrics.go @@ -41,7 +41,7 @@ var ( rebalancerBaselineSamplesTotal = prometheus.NewGaugeVec( prometheus.GaugeOpts{ Name: "simplyblock_rebalancer_baseline_samples_total", - Help: "Number of latency samples in the rolling window from which the per-node baseline was computed (rollingWindow strategy only).", + Help: "Number of latency samples in the rolling window from which the per-node baseline was computed (RollingWindow strategy only).", }, []string{"cluster", "node"}, ) @@ -49,7 +49,7 @@ var ( rebalancerBaselineSamplesRejected = prometheus.NewGaugeVec( prometheus.GaugeOpts{ Name: "simplyblock_rebalancer_baseline_samples_rejected", - Help: "Number of rolling-window latency samples rejected as outliers by the Hampel identifier when computing the per-node baseline (rollingWindow strategy only).", + Help: "Number of rolling-window latency samples rejected as outliers by the Hampel identifier when computing the per-node baseline (RollingWindow strategy only).", }, []string{"cluster", "node"}, ) diff --git a/operator/internal/autoplacement/utils.go b/operator/internal/autoplacement/utils.go index 5549387c5..cd19fde31 100644 --- a/operator/internal/autoplacement/utils.go +++ b/operator/internal/autoplacement/utils.go @@ -10,7 +10,7 @@ import ( const ( // DefaultEvaluationInterval is how often the rebalancer evaluates load when the spec - // does not override it. Exported so callers can fall back to it (e.g. for requeue + // does not override it. Exported so callers can fall back to it (e.g., for requeue // timing) before a RebalancingConfig has been resolved. DefaultEvaluationInterval = 60 * time.Second @@ -18,7 +18,7 @@ const ( defaultImbalanceThresholdPct = 80 // defaultMinHotColdDifferencePct is the minimum latency-deviation gap (in // percentage points) a target node must have below the hot source before a - // migration is worthwhile — prevents shuffling load between near-equally-loaded + // migration is worthwhile — prevents shuffling load between nearly equally loaded // nodes. defaultMinHotColdDifferencePct = 20 defaultCoolDownSeconds = 600 @@ -28,7 +28,7 @@ const ( // journal/EC/HA tail spikes. Overridden by the operator-wide --latency-percentile flag. defaultLatencyPercentile = "p50" - // Rolling-window baseline defaults. The rollingWindow strategy derives each node's + // Rolling-window baseline defaults. The RollingWindow strategy derives each node's // baseline from a robust (outlier-rejecting) estimate over BaselineWindow of the probe // latency series in Prometheus, rather than the frozen one-shot fio benchmark. defaultBaselineStrategy = string(simplyblockv1alpha2.BaselineStrategyRollingWindow) @@ -42,7 +42,7 @@ const ( // it must match the cadence at which the probe sidecar publishes latency samples. defaultBaselineStep = 5 * time.Minute - // migrationBudgetFraction is the fraction of the source node's total volume IO score + // migrationBudgetFraction is the fraction of the source node's total volume I/O score // that may be migrated in a single evaluation cycle. migrationBudgetFraction = 0.10 @@ -72,14 +72,15 @@ type RebalancingConfig struct { MaxMigrations int CoolDownSecs int64 - // BaselineStrategy selects how the per-node baseline is derived: "rollingWindow" - // (default) or "benchmark" (frozen one-shot fio measurement). + // BaselineStrategy selects how the per-node baseline is derived: "RollingWindow" + // (default) or "Benchmark" (frozen one-shot fio measurement). BaselineStrategy string - // BaselineWindow is the rollingWindow look-back period. + // BaselineWindow is the RollingWindow look-back period. BaselineWindow time.Duration // BaselineStep is the range-query step, matching the probe publish cadence. BaselineStep time.Duration - // BaselineColdStart is the under-sampled-node policy: "partialWindow" (default) or "defer". + // BaselineColdStart is the under-sampled-node policy: "PartialWindow" (default) + // or "Defer." BaselineColdStart string // BaselineMinSamples is the sample count below which a node is treated as under-sampled. BaselineMinSamples int diff --git a/operator/internal/autoplacement/utils_test.go b/operator/internal/autoplacement/utils_test.go index 37984d622..23136fb56 100644 --- a/operator/internal/autoplacement/utils_test.go +++ b/operator/internal/autoplacement/utils_test.go @@ -19,13 +19,13 @@ func TestResolveAutoPlacementConfig_BaselineDefaults(t *testing.T) { } if cfg.BaselineStrategy != string(simplyblockv1alpha2.BaselineStrategyRollingWindow) { - t.Errorf("BaselineStrategy = %q, want rollingWindow (default)", cfg.BaselineStrategy) + t.Errorf("BaselineStrategy = %q, want RollingWindow (default)", cfg.BaselineStrategy) } if cfg.BaselineWindow != 6*time.Hour { t.Errorf("BaselineWindow = %v, want 6h", cfg.BaselineWindow) } if cfg.BaselineColdStart != string(simplyblockv1alpha2.BaselineColdStartPartialWindow) { - t.Errorf("BaselineColdStart = %q, want partialWindow", cfg.BaselineColdStart) + t.Errorf("BaselineColdStart = %q, want PartialWindow", cfg.BaselineColdStart) } if cfg.BaselineMinSamples != 6 { t.Errorf("BaselineMinSamples = %d, want 6", cfg.BaselineMinSamples) @@ -54,13 +54,13 @@ func TestResolveAutoPlacementConfig_BaselineOverrides(t *testing.T) { } if cfg.BaselineStrategy != string(simplyblockv1alpha2.BaselineStrategyBenchmark) { - t.Errorf("BaselineStrategy = %q, want benchmark", cfg.BaselineStrategy) + t.Errorf("BaselineStrategy = %q, want Benchmark", cfg.BaselineStrategy) } if cfg.BaselineWindow != 12*time.Hour { t.Errorf("BaselineWindow = %v, want 12h", cfg.BaselineWindow) } if cfg.BaselineColdStart != string(simplyblockv1alpha2.BaselineColdStartDefer) { - t.Errorf("BaselineColdStart = %q, want defer", cfg.BaselineColdStart) + t.Errorf("BaselineColdStart = %q, want Defer", cfg.BaselineColdStart) } if cfg.BaselineMinSamples != 12 { t.Errorf("BaselineMinSamples = %d, want 12", cfg.BaselineMinSamples) diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusters.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusters.yaml index 23925fce6..e1f616c64 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusters.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusters.yaml @@ -1174,10 +1174,10 @@ spec: baselineColdStart: description: |- BaselineColdStart selects what happens to an under-sampled node. Defaults - to partialWindow. + to PartialWindow. enum: - - defer - - partialWindow + - Defer + - PartialWindow type: string baselineMinSamples: description: |- @@ -1194,14 +1194,14 @@ spec: baselineStrategy: description: |- BaselineStrategy selects how the per-node baseline is derived. Defaults - to rollingWindow. + to RollingWindow. enum: - - benchmark - - rollingWindow + - Benchmark + - RollingWindow type: string baselineWindow: description: |- - BaselineWindow is the look-back the rollingWindow strategy reduces. + BaselineWindow is the look-back the RollingWindow strategy reduces. Defaults to 6h. type: string defaultCoolDownSeconds: diff --git a/operator/internal/webhook/simplyblock_volume_placement_injector_test.go b/operator/internal/webhook/simplyblock_volume_placement_injector_test.go index b32a42339..233df5128 100644 --- a/operator/internal/webhook/simplyblock_volume_placement_injector_test.go +++ b/operator/internal/webhook/simplyblock_volume_placement_injector_test.go @@ -65,7 +65,7 @@ func makePlacementCluster(autoRebalancing *simplyblockv1alpha2.VolumeAutoPlaceme } // applyPVCPatches applies the RFC6902 patch set produced by Handle to the original PVC -// via a real JSON-patch library, mirroring what the k8s apiserver does — avoids having to +// via a real JSON-patch library, mirroring what the K8s apiserver does — avoids having to // guess the exact path granularity the diff library chose for the annotations map. func applyPVCPatches(t *testing.T, pvc *corev1.PersistentVolumeClaim, patches []jsonpatch.JsonPatchOperation) *corev1.PersistentVolumeClaim { t.Helper() @@ -114,7 +114,8 @@ type promSample struct { // fakePrometheusServer serves a canned instant-query vector result for any GET request to // /api/v1/query, in the exact envelope prometheus/client_golang's API client expects: -// {"status":"success","data":{"resultType":"vector","result":[...]}}. +// The body is a Prometheus vector response: a success status carrying a data +// object whose resultType is vector and whose result list is the samples. func fakePrometheusServer(t *testing.T, samples []promSample) *httptest.Server { t.Helper() ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) { @@ -345,7 +346,7 @@ func TestSimplyblockVolumePlacementInjector_Handle_SelectsCoolestEligibleNode(t PrometheusURL: ptr.To(promSrv.URL), // This case seeds fixed per-node baselines on the CR status and varies only the // current (instant) Prometheus reading to produce known deviations — that is the - // "benchmark" baseline model, so pin the strategy to it. + // "Benchmark" baseline model, so pin the strategy to it. BaselineStrategy: ptr.To(simplyblockv1alpha2.BaselineStrategyBenchmark), }) sc := makePlacementStorageClass(utils.CSIProvisioner, map[string]string{"cluster_id": testClusterUUID}) From 1ce801f407f4bb2ebd5abcbb9c5ed873afc9a563 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 14:44:25 +0200 Subject: [PATCH 049/206] feat(atlas-lib): a key is read under every spelling and written under one MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit §9.4 moves this product's annotations and labels from the bare simplyblock.io/ prefix to the API group's own. A flag day is not available for any of them: they sit on PersistentVolumes, claims, StorageClasses, and worker Nodes that outlive every operator upgrade, and some of them a person typed. So the move is the only way it can be done — every spelling is read, and one is written — and what was missing was a place for that to be a property of the key rather than of whichever call site somebody remembered. kube.Key is that place. It is the spellings newest first, the first being the one written, and the order is the whole of the precedence rule: an object carrying two is answered by the newer, because that is the one this product wrote. Three of its four operations are the ones a half-done move gets wrong. Set clears the older spellings as well as writing the current one, so an object this product touches stops carrying two answers to one question — the state §19.10's sixth check exists to find, and one a write that only added would keep creating. Delete removes every spelling, because removing only the current one leaves an older one behind and the next read answers with the value the delete was meant to retract. And Get decides on presence rather than on non-emptiness, since an empty value is a value: a key set to the empty string is how a toggle is turned off, and reading past it to an older spelling would answer with something the object has stopped saying. The eight constants in names.go now hold the group prefix, which is what makes every write site write the new spelling without being edited. The read sites are the ones that changed: the placement webhook's host-id exchange and pin check, the claim controller's applied-and-rejected diff, PinnedNode and its two companions, the normalized-handle reader, and the CSI controller's host_id precedence list and its removal of the placement hint. Two of the audit's premises turned out to be false, and the comments asserting them are corrected rather than left to mislead the next reader. The chart does not write simplyblock.io/replication-policy; nothing in helm-charts writes it at all. And kube.Finalizer — the one key where a half-move would have left objects undeletable, since a finalizer nobody recognizes is removed by nobody — has no call site in any module. What is left is the keys each consumer declares for itself, and they are listed rather than half-done here: the operator's fio-baseline pair, backup-policy, trigger-realignment, storage-node-uuid., and the rebalancer injection marker; the CSI driver's nvmf-model-id, lvol-id, pod-affinity, and the two guardian labels. Each needs this treatment at its own declaration. The topology-key prefix is the one that is not a local change: the node plugin publishes it and the controller co-locates on it, so it moves in both components or in neither. Co-Authored-By: Claude Fable 5 --- atlas-lib/kube/identity.go | 13 +- atlas-lib/kube/keys.go | 172 ++++++++++++++++++ atlas-lib/kube/keys_test.go | 147 +++++++++++++++ atlas-lib/kube/names.go | 23 ++- atlas-lib/kube/normalized.go | 3 +- csi-driver/internal/csi/controller/volume.go | 12 +- .../persistentvolumeclaim_controller.go | 16 +- .../internal/upgrade/check/annotations.go | 9 +- operator/internal/upgrade/keys/keys.go | 12 +- .../simplyblock_volume_placement_injector.go | 17 +- 10 files changed, 379 insertions(+), 45 deletions(-) create mode 100644 atlas-lib/kube/keys.go create mode 100644 atlas-lib/kube/keys_test.go diff --git a/atlas-lib/kube/identity.go b/atlas-lib/kube/identity.go index e33a0510c..fbd2ccc71 100644 --- a/atlas-lib/kube/identity.go +++ b/atlas-lib/kube/identity.go @@ -22,13 +22,11 @@ func IsManaged(pv *corev1.PersistentVolume) bool { // one-shot AnnoPlacementHint is deliberately NOT a pin: hinted volumes stay // eligible for rebalancing. func PinnedNode(annotations map[string]string) string { - if v := annotations[AnnoSelectedStorageNode]; v != "" { + if v, _ := KeySelectedStorageNode.Get(annotations); v != "" { return v } - if v := annotations[AnnoHostID]; v != "" { - return v - } - return annotations[DeprecatedAnnoHostID] + v, _ := KeyHostID.Get(annotations) + return v } // IsPinnedVolume reports whether the given (PVC) annotations pin a volume to a @@ -58,7 +56,7 @@ func PendingPin(annotations map[string]string) (string, bool) { if target == "" { return "", false } - if target == annotations[AnnoSelectedStorageNodeApplied] { + if applied, _ := KeySelectedStorageNodeApplied.Get(annotations); target == applied { return "", false } return target, true @@ -68,7 +66,8 @@ func PendingPin(annotations map[string]string) (string, bool) { // rejected as an unknown storage node, i.e., whether warning about it again would // be a duplicate. func PinRejected(annotations map[string]string, target string) bool { - return target != "" && annotations[AnnoSelectedStorageNodeRejected] == target + rejected, _ := KeySelectedStorageNodeRejected.Get(annotations) + return target != "" && rejected == target } // VolumeHandleFromPV extracts the simplyblock logical-volume handle from a diff --git a/atlas-lib/kube/keys.go b/atlas-lib/kube/keys.go new file mode 100644 index 000000000..8b4154477 --- /dev/null +++ b/atlas-lib/kube/keys.go @@ -0,0 +1,172 @@ +// The annotation and label keys this product puts on Kubernetes objects, each +// in every spelling it has ever had. +// +// The keys are moving from the bare simplyblock.io/ prefix to the API group's +// own storage.simplyblock.io/, which is design-crd-model.md §9.4. A flag day is +// not available: the keys sit on PersistentVolumes, claims, StorageClasses, and +// worker Nodes that outlive every operator upgrade, and some of them a person +// typed. So the move is done the only way it can be — **every spelling is read, +// and one is written** — and this type is what makes that a property of the key +// rather than of whichever call site somebody remembered to update. +// +// It is here rather than in a consumer because the operator writes these keys +// and the CSI driver reads them. Two inventories would be two answers to what a +// claim is annotated with, and the answer that mattered would be whichever +// process looked. + +package kube + +// The prefixes a key has been spelled with. +const ( + // GroupPrefix is the API group's own, which every key is written under. + GroupPrefix = "storage.simplyblock.io/" + + // BarePrefix is what shipped, and what an object written before the move + // still carries. + BarePrefix = "simplyblock.io/" + + // LegacyPrefix predates BarePrefix and survives on a handful of keys. A + // cluster old enough to carry it carries a third spelling of one thing. + LegacyPrefix = "simplybk/" +) + +// Key is one annotation or label, in every spelling it is read under, newest +// first. The first is the one that is written. +// +// The order is the whole of the precedence rule: an object carrying two +// spellings is answered by the newer, because that is the one this product +// wrote and the older is what it is migrating away from. +type Key []string + +// String is the spelling to write. A Key with no spellings yields the empty +// string, which is a programming error rather than a state to handle: every Key +// in this package is a literal. +func (k Key) String() string { + if len(k) == 0 { + return "" + } + return k[0] +} + +// Get returns the value under the newest spelling the object carries, and +// whether it carries any. +// +// Presence rather than non-emptiness decides, because an empty value is a value: +// a key set to the empty string is how a toggle is turned off, and reading past +// it to an older spelling would answer with something the object has stopped +// saying. +func (k Key) Get(values map[string]string) (string, bool) { + for _, spelling := range k { + if value, carried := values[spelling]; carried { + return value, true + } + } + return "", false +} + +// Has reports whether the object carries this key under any spelling. +func (k Key) Has(values map[string]string) bool { + _, carried := k.Get(values) + return carried +} + +// Set writes the current spelling and removes every older one, returning the +// map so a caller can assign it back where there was none. +// +// Clearing the old spellings is what stops an object this product has touched +// from carrying two answers to one question — which is the state §19.10's sixth +// check exists to find, and one a write that only added would keep creating. +func (k Key) Set(values map[string]string, value string) map[string]string { + if values == nil { + values = map[string]string{} + } + k.Delete(values) + if current := k.String(); current != "" { + values[current] = value + } + return values +} + +// Delete removes every spelling. +// +// Every one, because removing only the current spelling would leave an older +// one behind and the next read would answer with the value the delete was meant +// to retract. +func (k Key) Delete(values map[string]string) { + for _, spelling := range k { + delete(values, spelling) + } +} + +// Spellings flattens keys into every string they are read under, newest of each +// first, for a caller whose interface takes strings: a removal that patches the +// keys to null, or a lookup that walks an ordered list. +func Spellings(keys ...Key) []string { + var out []string + for _, key := range keys { + out = append(out, key...) + } + return out +} + +// The keys that are moving. Each is declared with the group prefix first, so +// that writing it is writing the new spelling and reading it is reading +// whichever the object has. +// +// The bare constants beside them in names.go are the same strings, kept because +// a caller that only writes has no use for the list. Where a caller reads, it +// reads through the Key. +var ( + // KeyVolumeHandle records a volume's normalized handle (§16.4). + KeyVolumeHandle = Key{AnnoVolumeHandle, BarePrefix + "volume-handle"} + + // KeyPool records the source pool on a PersistentVolume, for observability. + KeyPool = Key{AnnoPool, BarePrefix + "pool"} + + // KeySelectedStorageNode is the pin: the storage node a claim's volume is + // held on. A person writes this one, which is why the old spelling has to + // keep working for as long as claims carrying it exist. + KeySelectedStorageNode = Key{ + AnnoSelectedStorageNode, BarePrefix + "selected-storage-node", + } + + // KeySelectedStorageNodeApplied records the pin the controller has acted + // on, which is the diff that stops its own write re-triggering a migration. + KeySelectedStorageNodeApplied = Key{ + AnnoSelectedStorageNodeApplied, BarePrefix + "selected-storage-node-applied", + } + + // KeySelectedStorageNodeRejected records a pin the controller refused, so a + // warning about it is not repeated every reconcile. + KeySelectedStorageNodeRejected = Key{ + AnnoSelectedStorageNodeRejected, BarePrefix + "selected-storage-node-rejected", + } + + // KeyPlacementHint is the one-shot placement the webhook chose, which the + // CSI controller consumes and removes. + KeyPlacementHint = Key{AnnoPlacementHint, BarePrefix + "placement-hint"} + + // KeyHostID is the pre-pin placement annotation, honored as the lowest + // priority fallback and never rewritten. It carries three spellings because + // it was renamed once before the group prefix was settled. + KeyHostID = Key{AnnoHostID, BarePrefix + "host-id", LegacyPrefix + "host-id"} + + // KeyFinalizer guards a PersistentVolume or claim while its logical volume + // still exists. + KeyFinalizer = Key{Finalizer, BarePrefix + "lvol-protection"} +) + +// MovedKeys is the inventory, which is what lets a test assert the move as a +// property of the product rather than of one key. +func MovedKeys() []Key { + return []Key{ + KeyVolumeHandle, + KeyPool, + KeySelectedStorageNode, + KeySelectedStorageNodeApplied, + KeySelectedStorageNodeRejected, + KeyPlacementHint, + KeyHostID, + KeyFinalizer, + } +} diff --git a/atlas-lib/kube/keys_test.go b/atlas-lib/kube/keys_test.go new file mode 100644 index 000000000..5665e2897 --- /dev/null +++ b/atlas-lib/kube/keys_test.go @@ -0,0 +1,147 @@ +package kube + +import ( + "strings" + "testing" +) + +// The property the whole type exists for: an object written before the move +// keeps working, and an object written after it carries one spelling. +func TestAKeyReadsEverySpellingAndWritesOne(t *testing.T) { + key := Key{"storage.simplyblock.io/x", "simplyblock.io/x", "simplybk/x"} + + for _, tc := range []struct { + name string + on map[string]string + want string + }{ + {name: "nothing", on: map[string]string{}}, + {name: "the current spelling", on: map[string]string{"storage.simplyblock.io/x": "a"}, want: "a"}, + {name: "the previous one", on: map[string]string{"simplyblock.io/x": "b"}, want: "b"}, + {name: "the oldest one", on: map[string]string{"simplybk/x": "c"}, want: "c"}, + { + name: "the newest wins where two are carried", + on: map[string]string{ + "storage.simplyblock.io/x": "new", "simplyblock.io/x": "old", "simplybk/x": "older", + }, + want: "new", + }, + { + name: "and the newest of the two that are there", + on: map[string]string{"simplyblock.io/x": "old", "simplybk/x": "older"}, + want: "old", + }, + } { + t.Run(tc.name, func(t *testing.T) { + got, found := key.Get(tc.on) + if got != tc.want { + t.Errorf("Get = %q, want %q", got, tc.want) + } + if found != (tc.want != "") { + t.Errorf("found = %v, want %v", found, tc.want != "") + } + }) + } +} + +// An empty value is a value. A key set to the empty string is carried +// deliberately — it is how a user turns a toggle off — so reporting it as absent +// would read the next spelling down and answer with something the object stopped +// saying. +func TestAnEmptyValueIsStillTheAnswer(t *testing.T) { + key := Key{"storage.simplyblock.io/x", "simplyblock.io/x"} + + got, found := key.Get(map[string]string{ + "storage.simplyblock.io/x": "", "simplyblock.io/x": "stale", + }) + if !found || got != "" { + t.Errorf("Get = %q, %v; want the empty current value to win over the old one", + got, found) + } +} + +// Setting writes the current spelling and clears every older one, so an object +// the operator touches stops carrying two answers to one question. +func TestSetWritesOneSpellingAndClearsTheRest(t *testing.T) { + key := Key{"storage.simplyblock.io/x", "simplyblock.io/x", "simplybk/x"} + + on := key.Set(map[string]string{"simplyblock.io/x": "old", "other": "untouched"}, "new") + + if got := on["storage.simplyblock.io/x"]; got != "new" { + t.Errorf("the current spelling is %q, want %q", got, "new") + } + if _, stale := on["simplyblock.io/x"]; stale { + t.Error("the old spelling survived the write, so the object carries two answers") + } + if on["other"] != "untouched" { + t.Error("Set touched a key that is not this one") + } +} + +// Set on a nil map returns one rather than panicking, because the object it is +// called on often has no annotations yet. +func TestSetBuildsAMapWhereThereIsNone(t *testing.T) { + key := Key{"storage.simplyblock.io/x"} + if on := key.Set(nil, "v"); on["storage.simplyblock.io/x"] != "v" { + t.Errorf("Set on a nil map produced %v", on) + } +} + +// Deleting removes every spelling. Removing only the current one would leave the +// old one behind, and the next read would answer with the value the delete was +// meant to retract. +func TestDeleteRemovesEverySpelling(t *testing.T) { + key := Key{"storage.simplyblock.io/x", "simplyblock.io/x", "simplybk/x"} + + on := map[string]string{ + "storage.simplyblock.io/x": "a", "simplyblock.io/x": "b", "simplybk/x": "c", "other": "d", + } + key.Delete(on) + + if _, found := key.Get(on); found { + t.Error("a spelling survived the delete, so the value it retracted is still readable") + } + if on["other"] != "d" { + t.Error("Delete touched a key that is not this one") + } +} + +// The declared keys all move to the group's own prefix and are all read under +// the bare one they shipped with. The inventory is what makes the move a fact +// about the product rather than about whichever call site somebody remembered. +func TestEveryDeclaredKeyWritesTheGroupPrefixAndReadsTheOldOne(t *testing.T) { + for _, key := range MovedKeys() { + if len(key) < 2 { + t.Errorf("%q is declared with one spelling, so nothing reads what it used to be "+ + "called and an object written before the move stops being understood", key) + continue + } + if got := key.String(); !strings.HasPrefix(got, GroupPrefix) { + t.Errorf("%q is written under %q, and the move is to %q", key, got, GroupPrefix) + } + var bare bool + for _, spelling := range key[1:] { + if strings.HasPrefix(spelling, BarePrefix) { + bare = true + } + } + if !bare { + t.Errorf("%q reads no %q spelling, so an object written before the move is not "+ + "understood", key, BarePrefix) + } + } +} + +// No two keys share a spelling. Two keys reading one string is two questions +// answered by one value, and whichever is written last wins. +func TestNoTwoKeysShareASpelling(t *testing.T) { + seen := map[string]Key{} + for _, key := range MovedKeys() { + for _, spelling := range key { + if other, taken := seen[spelling]; taken { + t.Errorf("%q is read by both %v and %v", spelling, other, key) + } + seen[spelling] = key + } + } +} diff --git a/atlas-lib/kube/names.go b/atlas-lib/kube/names.go index 521789680..d8879cd4a 100644 --- a/atlas-lib/kube/names.go +++ b/atlas-lib/kube/names.go @@ -103,12 +103,19 @@ const ( ) // Labels, annotations, and finalizers atlas-managed objects carry. +// +// Each is the spelling that is written, which is the API group's own prefix. +// The spellings an object written before the move carries are in keys.go beside +// the Key that reads them, and a caller that reads one of these off an object +// reads it through that Key rather than through the constant — the constant +// alone would stop understanding every claim, volume, and node that predates +// the move (design-crd-model.md §9.4). const ( // LabelVolumeHandle lets selectors find the K8s objects for a logical // volume. Nothing writes it, and nothing can: a handle is 110 bytes // normalized and a label value stops at 63, so AnnoVolumeHandle carries // this instead. - LabelVolumeHandle = "simplyblock.io/volume-handle" + LabelVolumeHandle = "storage.simplyblock.io/volume-handle" // AnnoVolumeHandle records a volume's handle with its pool segment // normalized to a UUID, on the PersistentVolume and the @@ -127,7 +134,7 @@ const ( // hand-edited annotation cannot redirect a volume to another cluster. AnnoVolumeHandle = "storage.simplyblock.io/volume-handle" // AnnoPool records the source pool on the PV for observability. - AnnoPool = "simplyblock.io/pool" + AnnoPool = "storage.simplyblock.io/pool" // LabelPoolPrefix opens the per-pool label the operator puts on every node in // a StoragePool's AllowedNodes LabelPoolPrefix = "storage.simplyblock.io/storage-pool." @@ -136,32 +143,32 @@ const ( // node. It is the canonical placement/pin annotation: the operator's pin // controller, drain, and rebalancer key off it, and the CSI controller reads // it in CreateVolume as the primary host_id source. - AnnoSelectedStorageNode = "simplyblock.io/selected-storage-node" + AnnoSelectedStorageNode = "storage.simplyblock.io/selected-storage-node" // AnnoSelectedStorageNodeApplied records the pinned-volume target the PVC // controller has already acted on. It is the strict change-diff marker: the // controller only requests a migration when AnnoSelectedStorageNode differs // from this value, so its own writes do not re-trigger a migration. - AnnoSelectedStorageNodeApplied = "simplyblock.io/selected-storage-node-applied" + AnnoSelectedStorageNodeApplied = "storage.simplyblock.io/selected-storage-node-applied" // AnnoSelectedStorageNodeRejected records the last pinned-volume value the PVC // controller's backstop validation rejected as an unknown storage node. It // suppresses duplicate warning events while the invalid value remains in place. - AnnoSelectedStorageNodeRejected = "simplyblock.io/selected-storage-node-rejected" + AnnoSelectedStorageNodeRejected = "storage.simplyblock.io/selected-storage-node-rejected" // AnnoPlacementHint is a one-shot creation-time placement hint: the volume- // placement webhook writes it with the least-loaded node it picked, the CSI // controller sends it as host_id at CreateVolume, and then removes it once the // volume exists. Unlike AnnoSelectedStorageNode it is not a pin, and the // volume stays eligible for rebalancing. - AnnoPlacementHint = "simplyblock.io/placement-hint" + AnnoPlacementHint = "storage.simplyblock.io/placement-hint" // AnnoHostID is the legacy per-PVC placement annotation. It is honored by the // CSI controller as a lowest-priority host_id fallback for pre-existing PVCs, // but is never rewritten or removed by the provisioner. The volume-placement // webhook rewrites a user-supplied host-id into AnnoSelectedStorageNode (a pin, // matching its pre-migration behavior) on new PVCs. - AnnoHostID = "simplyblock.io/host-id" + AnnoHostID = "storage.simplyblock.io/host-id" // DeprecatedAnnoHostID is the pre-rename form of AnnoHostID, still // honored for backward compatibility. DeprecatedAnnoHostID = "simplybk/host-id" // Finalizer guards a PV/PVC from deletion until the backing logical // volume is released. - Finalizer = "simplyblock.io/lvol-protection" + Finalizer = "storage.simplyblock.io/lvol-protection" ) diff --git a/atlas-lib/kube/normalized.go b/atlas-lib/kube/normalized.go index b0584f0c3..55c5b22f7 100644 --- a/atlas-lib/kube/normalized.go +++ b/atlas-lib/kube/normalized.go @@ -29,7 +29,8 @@ import ( func NormalizedHandle( field lvol.VolumeHandle, annotations map[string]string, ) (lvol.Normalized, bool) { - return lvol.NormalizeHandle(field, lvol.VolumeHandle(annotations[AnnoVolumeHandle])) + annotated, _ := KeyVolumeHandle.Get(annotations) + return lvol.NormalizeHandle(field, lvol.VolumeHandle(annotated)) } // NormalizedVolumeHandleFromPV reads a PersistentVolume's handle with the diff --git a/csi-driver/internal/csi/controller/volume.go b/csi-driver/internal/csi/controller/volume.go index 638012cf4..09db1f734 100644 --- a/csi-driver/internal/csi/controller/volume.go +++ b/csi-driver/internal/csi/controller/volume.go @@ -118,7 +118,8 @@ func (cs *Server) CreateVolume( params := req.GetParameters() pvcName, pvcNamespace := params[csicommon.CSIStorageNameKey], params[csicommon.CSIStorageNamespaceKey] if pvcName != "" && pvcNamespace != "" { - if rerr := cs.removePVCAnnotations(ctx, pvcName, pvcNamespace, kube.AnnoPlacementHint); rerr != nil { + if rerr := cs.removePVCAnnotations(ctx, pvcName, pvcNamespace, + kube.Spellings(kube.KeyPlacementHint)...); rerr != nil { klog.Warningf("createVolume: could not clear placement-hint on PVC %s/%s: %v", pvcNamespace, pvcName, rerr) } } @@ -213,10 +214,11 @@ func (cs *Server) prepareCreateVolumeReq( } // host_id priority: selected-storage-node (hard pin) → placement-hint (one-shot - // hint from the placement webhook) → host-id and its deprecated form (legacy - // fallback for pre-existing PVCs). - hostID := pvcAnnotation(pvcAnns, - kube.AnnoSelectedStorageNode, kube.AnnoPlacementHint, kube.AnnoHostID, kube.DeprecatedAnnoHostID) + // hint from the placement webhook) → host-id (legacy fallback for pre-existing + // PVCs). Each is expanded into every prefix it has been written under, since a + // claim outlives the operator that annotated it (design-crd-model.md §9.4). + hostID := pvcAnnotation(pvcAnns, kube.Spellings( + kube.KeySelectedStorageNode, kube.KeyPlacementHint, kube.KeyHostID)...) lvolID := pvcAnnotation(pvcAnns, annotationLvolID, deprecatedAnnotationLvolID) podAffinitive, _ := strconv.ParseBool(pvcAnns[annotationPodAffinity]) diff --git a/operator/internal/controller/persistentvolumeclaim_controller.go b/operator/internal/controller/persistentvolumeclaim_controller.go index ca1c4ffc4..54883d24f 100644 --- a/operator/internal/controller/persistentvolumeclaim_controller.go +++ b/operator/internal/controller/persistentvolumeclaim_controller.go @@ -85,7 +85,7 @@ func (r *PersistentVolumeClaimReconciler) Reconcile( // setApplied normalizes those legacy annotations into selected-storage-node // when it records the applied target, so the annotation state converges. desired := kube.PinnedNode(pvc.Annotations) - applied := pvc.Annotations[kube.AnnoSelectedStorageNodeApplied] + applied, _ := kube.KeySelectedStorageNodeApplied.Get(pvc.Annotations) // Strict change-diff gate: nothing to do unless the pinned target changed. if desired == applied { @@ -257,7 +257,7 @@ func (r *PersistentVolumeClaimReconciler) rejectTarget( pvc *corev1.PersistentVolumeClaim, target string, ) (ctrl.Result, error) { - if pvc.Annotations[kube.AnnoSelectedStorageNodeRejected] == target { + if rejected, _ := kube.KeySelectedStorageNodeRejected.Get(pvc.Annotations); rejected == target { return ctrl.Result{}, nil } r.Recorder.Eventf(pvc, nil, corev1.EventTypeWarning, "InvalidPinTarget", "InvalidPinTarget", @@ -266,7 +266,7 @@ func (r *PersistentVolumeClaimReconciler) rejectTarget( if pvc.Annotations == nil { pvc.Annotations = map[string]string{} } - pvc.Annotations[kube.AnnoSelectedStorageNodeRejected] = target + pvc.Annotations = kube.KeySelectedStorageNodeRejected.Set(pvc.Annotations, target) if err := r.Patch(ctx, pvc, patch); err != nil { return ctrl.Result{}, fmt.Errorf("record rejected pin target: %w", err) } @@ -285,18 +285,18 @@ func (r *PersistentVolumeClaimReconciler) setApplied( pvc.Annotations = map[string]string{} } if value == "" { - delete(pvc.Annotations, kube.AnnoSelectedStorageNodeApplied) + kube.KeySelectedStorageNodeApplied.Delete(pvc.Annotations) } else { - pvc.Annotations[kube.AnnoSelectedStorageNodeApplied] = value + pvc.Annotations = kube.KeySelectedStorageNodeApplied.Set(pvc.Annotations, value) // Normalize legacy pin annotations into the canonical one so the state // converges: a pre-existing host-id pin (which drove this reconcile via // kube.PinnedNode) is rewritten to selected-storage-node and the legacy // forms are dropped. - pvc.Annotations[kube.AnnoSelectedStorageNode] = value - delete(pvc.Annotations, kube.AnnoHostID) + pvc.Annotations = kube.KeySelectedStorageNode.Set(pvc.Annotations, value) + kube.KeyHostID.Delete(pvc.Annotations) delete(pvc.Annotations, kube.DeprecatedAnnoHostID) } - delete(pvc.Annotations, kube.AnnoSelectedStorageNodeRejected) + kube.KeySelectedStorageNodeRejected.Delete(pvc.Annotations) if err := r.Patch(ctx, pvc, patch); err != nil { return ctrl.Result{}, fmt.Errorf("record applied pin target: %w", err) } diff --git a/operator/internal/upgrade/check/annotations.go b/operator/internal/upgrade/check/annotations.go index 499f58eb0..dea0ee18f 100644 --- a/operator/internal/upgrade/check/annotations.go +++ b/operator/internal/upgrade/check/annotations.go @@ -7,10 +7,11 @@ // deprecation window. Two spellings that disagree are two answers to one // question, and picking either would silently discard a value somebody set. // -// Some keys have already half moved. The chart writes -// simplyblock.io/replication-policy while the operator reads -// storage.simplyblock.io/replication-policy, so a claim can carry both today, -// which is why this runs before the rewrite rather than as part of it. +// Some keys have already moved, and the move itself is what produces the state +// this looks for: atlas-lib reads every spelling and writes the group's own, so +// an object the operator has touched since carries the new key while one it has +// not still carries the old. Two of those on one object is what this runs before +// the rewrite to find, rather than as part of it. package check diff --git a/operator/internal/upgrade/keys/keys.go b/operator/internal/upgrade/keys/keys.go index e59095900..197d3ec26 100644 --- a/operator/internal/upgrade/keys/keys.go +++ b/operator/internal/upgrade/keys/keys.go @@ -1,10 +1,14 @@ // The inventory, and how a key is matched on an object. // // The rows are read from the code rather than from the design, which counts the -// keys without listing them. Some of them have already half moved: the chart -// writes simplyblock.io/replication-policy while the operator reads -// storage.simplyblock.io/replication-policy, so a claim can already carry both, -// which is exactly the state §19.10's sixth check exists to find. +// keys without listing them. +// +// A cluster can carry two spellings of one key, which is the state §19.10's +// sixth check exists to find. It arises from the move itself rather than from +// any one component disagreeing with another: the keys are read under every +// spelling and written under the group's own (atlas-lib/kube/keys.go), so an +// object the operator has not touched since the move still carries the old one +// while a sibling written afterward carries the new. package keys diff --git a/operator/internal/webhook/simplyblock_volume_placement_injector.go b/operator/internal/webhook/simplyblock_volume_placement_injector.go index 92d80e2e6..3a38080f6 100644 --- a/operator/internal/webhook/simplyblock_volume_placement_injector.go +++ b/operator/internal/webhook/simplyblock_volume_placement_injector.go @@ -45,7 +45,7 @@ type primaryNodeSelector interface { // SimplyblockVolumePlacementInjector is a mutating admission webhook that computes the // least-loaded eligible storage node for a new PVC's primary volume — using the same // latency-deviation signal the auto-rebalancer (Issue #130) uses — and stamps it onto the -// PVC as the simplyblock.io/host-id annotation, which spdk-csi already reads and forwards +// PVC as the simplyblock.io/host-id annotation, which spdk-csi already reads and forward // as host_id on CreateVolume. failurePolicy=ignore, and every skip/error path below allows // the PVC unmodified, so this can never block volume provisioning: sbcli's own // weighted-random pick (_get_next_3_nodes) runs as the fallback exactly as it does today. @@ -71,7 +71,7 @@ func (h *SimplyblockVolumePlacementInjector) Handle( // controller, drain, and rebalancer recognize, and which the CSI driver reads // as the primary host_id source at CreateVolume — and drop the legacy host-id // forms. The user's choice wins, so we do not run load-based placement. - hostID := pvc.Annotations[kube.AnnoHostID] + hostID, _ := kube.KeyHostID.Get(pvc.Annotations) if hostID == "" { hostID = pvc.Annotations[kube.DeprecatedAnnoHostID] } @@ -80,11 +80,12 @@ func (h *SimplyblockVolumePlacementInjector) Handle( if patched.Annotations == nil { patched.Annotations = make(map[string]string) } - // Do not clobber an existing explicit selected-storage-node pin. - if patched.Annotations[kube.AnnoSelectedStorageNode] == "" { - patched.Annotations[kube.AnnoSelectedStorageNode] = hostID + // Do not clobber an existing explicit selected-storage-node pin, under + // whichever prefix the claim carries it. + if pinned, _ := kube.KeySelectedStorageNode.Get(patched.Annotations); pinned == "" { + patched.Annotations = kube.KeySelectedStorageNode.Set(patched.Annotations, hostID) } - delete(patched.Annotations, kube.AnnoHostID) + kube.KeyHostID.Delete(patched.Annotations) delete(patched.Annotations, kube.DeprecatedAnnoHostID) log.Info("Exchanged legacy host-id for selected-storage-node", "node", hostID) return patchResponse(pvc, patched) @@ -92,7 +93,7 @@ func (h *SimplyblockVolumePlacementInjector) Handle( // An explicit selected-storage-node pin is honored as-is; never override it // with a load-based pick. - if pvc.Annotations[kube.AnnoSelectedStorageNode] != "" { + if pinned, _ := kube.KeySelectedStorageNode.Get(pvc.Annotations); pinned != "" { log.V(1).Info("Skipping: selected-storage-node already set") return admission.Allowed("selected-storage-node already set") } @@ -116,7 +117,7 @@ func (h *SimplyblockVolumePlacementInjector) Handle( if patched.Annotations == nil { patched.Annotations = make(map[string]string) } - patched.Annotations[kube.AnnoPlacementHint] = nodeUUID + patched.Annotations = kube.KeyPlacementHint.Set(patched.Annotations, nodeUUID) log.Info("Selected primary node for new volume", "nodeUUID", nodeUUID, "clusterUUID", clusterUUID) return patchResponse(pvc, patched) } From fc5a470a62128c6a469aef554a50b34051c9dafc Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 14:59:27 +0200 Subject: [PATCH 050/206] feat(device): one device can be restarted without restarting its node MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit design-storagedevice §6 specifies five operations against a device and none of them existed, while StorageDevice.status.activeOpsRef had been declared and left empty waiting for them. A field that points at a kind nobody can create is one a reader cannot interpret: there is no way to tell "no operation is running" from "this cannot say." What the v2 API can serve was measured rather than assumed. It offers three device verbs — restart, remove, and reset — where §7 asks for seven. Restart is built on the first, and it is the action the kind exists for: a wedged device is recycled today by restarting its storage node, which takes every other device on that node with it and costs the cluster a node's worth of redundancy for the duration. So the kind ships with one action, the Requesting to Awaiting graph, the reconciler, and the lock. atlas-lib gained the device half of the control-plane client, which it had none of. The enum admits only Restart, and that is the decision worth stating. Declaring all five would accept an object whose first reconcile can only fail, and an API that takes a request it will never carry out is worse than one that refuses it at admission: the refusal names the capability that is missing, and the failure names nothing a user can act on. Widening an enum is additive, so each action arrives with the endpoint it needs. The other four are v1alpha2.ExternalDependencies, which is data rather than prose. Each row names the action, the endpoint to build, and what the action does with it, so the ask carries its own justification and is a checklist on both sides. A test holds the list against the declared graphs in both directions: an action cannot be blocked without a row, and cannot ship while leaving one behind as an ask for something that exists. The ask itself: self-test and fail are one endpoint each. Replace needs the adopt call that names the device arriving, and the remove verb it already has buys nothing without it, since the pairing is what makes the arrival identifiable. Migrate needs detach and an adopt that accepts a device with its contents — an adopt that can only take an empty device turns the action into two replacements and a full rebuild, which is precisely what §6.2 gives as its reason to exist. The two things the graph declares are the two a retry would otherwise paper over. The restart is issued once, because the step is persisted before the call and a restart issued twice is a device recycled twice. And Awaiting is not abortable, because the control plane has accepted the restart and nothing recalls one, so an abort there would record a stop that did not happen while the device restarted anyway. Not built, and not waiting on anybody else: the DELETE admission guard, for which UnabortableDeviceSteps is exported, and §8.2's two operation metrics. Co-Authored-By: Claude Fable 5 --- atlas-lib/controlplane/devices.go | 158 +++++ atlas-lib/controlplane/devices_test.go | 120 ++++ ...orage.simplyblock.io_storagedeviceops.yaml | 162 +++++ .../templates/roles/manager_role.yaml | 3 + .../api/v1alpha2/storagedeviceops_types.go | 248 ++++++++ .../api/v1alpha2/zz_generated.deepcopy.go | 113 ++++ operator/cmd/main.go | 13 + ...orage.simplyblock.io_storagedeviceops.yaml | 162 +++++ operator/config/rbac/role.yaml | 3 + operator/dist/install.yaml | 3 + .../crd-redesign/design-storagedevice.md | 28 +- .../node/storagedeviceops_controller.go | 556 ++++++++++++++++++ .../node/storagedeviceops_graphs.go | 99 ++++ .../controllers/node/storagedeviceops_lock.go | 129 ++++ .../controllers/node/storagedeviceops_test.go | 425 +++++++++++++ ...orage.simplyblock.io_storagedeviceops.yaml | 162 +++++ 16 files changed, 2380 insertions(+), 4 deletions(-) create mode 100644 atlas-lib/controlplane/devices.go create mode 100644 atlas-lib/controlplane/devices_test.go create mode 100644 helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagedeviceops.yaml create mode 100644 operator/api/v1alpha2/storagedeviceops_types.go create mode 100644 operator/config/crd/bases/storage.simplyblock.io_storagedeviceops.yaml create mode 100644 operator/internal/controllers/node/storagedeviceops_controller.go create mode 100644 operator/internal/controllers/node/storagedeviceops_graphs.go create mode 100644 operator/internal/controllers/node/storagedeviceops_lock.go create mode 100644 operator/internal/controllers/node/storagedeviceops_test.go create mode 100644 operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagedeviceops.yaml diff --git a/atlas-lib/controlplane/devices.go b/atlas-lib/controlplane/devices.go new file mode 100644 index 000000000..b8a99cd4a --- /dev/null +++ b/atlas-lib/controlplane/devices.go @@ -0,0 +1,158 @@ +// One physical device of one storage node, and the operations the v2 API offers +// against it. +// +// The v2 API offers three verbs on a device — restart, remove, and reset — where +// design-storagedevice.md §7 asks for seven. What is here is what exists: the +// reads the operations wait on, and the restart they are built around. The four +// that are missing are recorded as an external dependency in +// operator/api/v1alpha2/storagedeviceops_types.go, beside the actions each one +// blocks, rather than as a client method that would return a 404. + +package controlplane + +import ( + "context" + "fmt" + + "github.com/google/uuid" + + "github.com/simplyblock/atlas/internal/cpapi" +) + +// Device is one physical device of a storage node. +// +// It is the identity and the state, not the occupancy: what a device holds +// changes continuously and the DTO's capacity block is a snapshot the stream +// never refreshes, so the current figures are the exporter's gauges and are read +// through atlas-lib/prometheus instead (§7). +type Device struct { + ID string + ClusterID string + NodeID string + + // Status is the control plane's own spelling, carried rather than mapped. + // An operation waits for it to become what the operation asked for, and a + // vocabulary translated here would be one the wait could not express. + Status string + + // The hardware identity, which is what a person matches against a slot. + Model string + SerialNumber string + PCIeAddress string + NVMeController string + + // SizeBytes is what the device is, as the control plane measured it. + SizeBytes uint64 + + // Health is what the device reports about itself. IOError and + // RetriesExhausted are the two the control plane acts on; HealthCheck is + // absent where the device does not report one. + HealthCheck *bool + IOError bool + RetriesExhausted bool +} + +func deviceFromDTO(d cpapi.DeviceDTO) Device { + return Device{ + ID: d.Id.String(), + ClusterID: d.ClusterId.String(), + NodeID: d.StorageNodeId.String(), + Status: d.Status, + Model: d.Model, + SerialNumber: d.SerialNumber, + PCIeAddress: d.PcieAddress, + NVMeController: d.NvmeController, + SizeBytes: uint64(d.Size), + HealthCheck: d.HealthCheck, + IOError: d.IoError, + RetriesExhausted: d.RetriesExhausted, + } +} + +// ListDevices returns every device one storage node holds. +func (c *Client) ListDevices(ctx context.Context, clusterID, nodeID string) ([]Device, error) { + cluster, node, err := parseNodeIDs(clusterID, nodeID) + if err != nil { + return nil, err + } + + resp, err := c.api.ClustersStorageNodesDevicesListApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesGetWithResponse( + ctx, cluster, node, nil) + if err != nil { + return nil, fmt.Errorf("list the devices of storage node %s: %w", nodeID, err) + } + ds, err := payload("devices of storage node "+nodeID, resp.JSON200, resp.StatusCode(), resp.Body) + if err != nil { + return nil, err + } + + out := make([]Device, 0, len(*ds)) + for _, d := range *ds { + out = append(out, deviceFromDTO(d)) + } + return out, nil +} + +// Device returns one device. It wraps errs.ErrNotFound for a device the control +// plane does not hold, which is how a caller tells a device that is gone from a +// control plane it could not reach. +func (c *Client) Device(ctx context.Context, clusterID, nodeID, deviceID string) (Device, error) { + cluster, node, device, err := parseDeviceIDs(clusterID, nodeID, deviceID) + if err != nil { + return Device{}, err + } + + resp, err := c.api.ClustersStorageNodesDevicesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdGetWithResponse( + ctx, cluster, node, device, nil) + if err != nil { + return Device{}, fmt.Errorf("read device %s: %w", deviceID, err) + } + d, err := payload("device "+deviceID, resp.JSON200, resp.StatusCode(), resp.Body) + if err != nil { + return Device{}, err + } + return deviceFromDTO(*d), nil +} + +// RestartDevice recycles one device in place. +// +// It is the narrowest recycling the control plane offers: the alternative is +// restarting the device's storage node, which takes every other device on it +// along and costs the cluster a node's worth of redundancy for the duration. +// +// The call returns when the control plane has accepted the request rather than +// when the device is back, so the caller waits on the device's own status. +func (c *Client) RestartDevice(ctx context.Context, clusterID, nodeID, deviceID string) error { + cluster, node, device, err := parseDeviceIDs(clusterID, nodeID, deviceID) + if err != nil { + return err + } + + resp, err := c.api.ClustersStorageNodesDevicesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRestartPostWithResponse( + ctx, cluster, node, device, nil) + if err != nil { + return fmt.Errorf("restart device %s: %w", deviceID, err) + } + if code := resp.StatusCode(); code < 200 || code >= 300 { + return respError("restart device "+deviceID, code, resp.Body) + } + return nil +} + +// parseNodeIDs parses a cluster and storage-node identifier. +func parseNodeIDs(clusterID, nodeID string) (cluster, node uuid.UUID, err error) { + if cluster, err = parseUUID("cluster id", clusterID); err != nil { + return cluster, node, err + } + node, err = parseUUID("storage node id", nodeID) + return cluster, node, err +} + +// parseDeviceIDs parses the three identifiers every device path carries. +func parseDeviceIDs(clusterID, nodeID, deviceID string) (cluster, node, device uuid.UUID, err error) { + if cluster, node, err = parseNodeIDs(clusterID, nodeID); err != nil { + return cluster, node, device, err + } + device, err = parseUUID("device id", deviceID) + return cluster, node, device, err +} diff --git a/atlas-lib/controlplane/devices_test.go b/atlas-lib/controlplane/devices_test.go new file mode 100644 index 000000000..e0e48cc9a --- /dev/null +++ b/atlas-lib/controlplane/devices_test.go @@ -0,0 +1,120 @@ +package controlplane + +import ( + "encoding/json" + "errors" + "net/http" + "net/http/httptest" + "testing" + + "github.com/simplyblock/atlas/errs" +) + +const ( + testDeviceCluster = "8ffac363-0c46-4714-a71b-f9c0b58a1269" + testDeviceNode = "a1111111-1111-4111-8111-111111111111" + testDeviceID = "b2222222-2222-4222-8222-222222222222" +) + +func deviceClient(t *testing.T, h http.HandlerFunc) *Client { + t.Helper() + + srv := httptest.NewServer(h) + t.Cleanup(srv.Close) + c, err := New(Config{Endpoint: srv.URL, Token: "t"}) + if err != nil { + t.Fatalf("building the client: %v", err) + } + return c +} + +func TestDeviceReportsWhatTheControlPlaneHolds(t *testing.T) { + c := deviceClient(t, func(w http.ResponseWriter, r *http.Request) { + if r.Method != http.MethodGet { + t.Errorf("method = %s, want GET", r.Method) + } + w.Header().Set("Content-Type", "application/json") + _ = json.NewEncoder(w).Encode(map[string]any{ + "id": testDeviceID, "cluster_id": testDeviceCluster, + "storage_node_id": testDeviceNode, "status": "online", + "model": "SAMSUNG MZQL2", "serial_number": "S64H", "nvme_controller": "nvme0", + "pcie_address": "0000:5e:00.0", "size": 1920383410176, + "cluster_device_order": 0, "io_error": false, "is_partition": false, + "retries_exhausted": false, "nvmf_ips": []string{}, + "capacity": map[string]any{}, + }) + }) + + got, err := c.Device(t.Context(), testDeviceCluster, testDeviceNode, testDeviceID) + if err != nil { + t.Fatalf("reading the device: %v", err) + } + if got.ID != testDeviceID || got.Status != "online" { + t.Errorf("Device = %+v, want the id and status the control plane reported", got) + } + if got.SerialNumber != "S64H" || got.PCIeAddress != "0000:5e:00.0" { + t.Errorf("Device = %+v, want the hardware identity carried through", got) + } +} + +// A device the control plane does not hold is a finding a caller tells apart +// from a transport failure, because the two lead to different actions: one +// means the device is gone and the other means to try again. +func TestDeviceReportsNotFound(t *testing.T) { + c := deviceClient(t, func(w http.ResponseWriter, _ *http.Request) { + w.WriteHeader(http.StatusNotFound) + }) + + if _, err := c.Device(t.Context(), testDeviceCluster, testDeviceNode, testDeviceID); !errors.Is(err, errs.ErrNotFound) { + t.Errorf("err = %v, want it to wrap ErrNotFound", err) + } +} + +func TestRestartDeviceIssuesThePost(t *testing.T) { + var path, method string + c := deviceClient(t, func(w http.ResponseWriter, r *http.Request) { + path, method = r.URL.Path, r.Method + w.WriteHeader(http.StatusOK) + }) + + if err := c.RestartDevice(t.Context(), testDeviceCluster, testDeviceNode, testDeviceID); err != nil { + t.Fatalf("restarting: %v", err) + } + if method != http.MethodPost { + t.Errorf("method = %s, want POST", method) + } + if want := "/api/v2/clusters/" + testDeviceCluster + "/storage-nodes/" + testDeviceNode + + "/devices/" + testDeviceID + "/restart"; path != want { + t.Errorf("path = %s, want %s", path, want) + } +} + +// A refused restart is an error rather than a silence. The operation that +// issued it records the step as failed, and a call that reported success on a +// 500 would leave it waiting for a restart nobody performed. +func TestRestartDeviceReportsARefusal(t *testing.T) { + c := deviceClient(t, func(w http.ResponseWriter, _ *http.Request) { + w.WriteHeader(http.StatusInternalServerError) + _, _ = w.Write([]byte(`{"error":"device is busy"}`)) + }) + + if err := c.RestartDevice(t.Context(), testDeviceCluster, testDeviceNode, testDeviceID); err == nil { + t.Error("a refused restart was reported as performed") + } +} + +// Every identifier is a UUID in the v2 API, so one that is not is refused here +// rather than becoming a request path the control plane answers with a 422. +func TestDeviceCallsRefuseAnIdentifierThatIsNotAUUID(t *testing.T) { + c := deviceClient(t, func(w http.ResponseWriter, _ *http.Request) { + t.Error("a malformed identifier reached the control plane") + w.WriteHeader(http.StatusOK) + }) + + if err := c.RestartDevice(t.Context(), "not-a-uuid", testDeviceNode, testDeviceID); err == nil { + t.Error("a cluster id that is not a UUID was accepted") + } + if _, err := c.Device(t.Context(), testDeviceCluster, testDeviceNode, "nope"); err == nil { + t.Error("a device id that is not a UUID was accepted") + } +} diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagedeviceops.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagedeviceops.yaml new file mode 100644 index 000000000..eec5d0629 --- /dev/null +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagedeviceops.yaml @@ -0,0 +1,162 @@ +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition +metadata: + annotations: + controller-gen.kubebuilder.io/version: v0.21.0 + name: storagedeviceops.storage.simplyblock.io +spec: + group: storage.simplyblock.io + names: + kind: StorageDeviceOps + listKind: StorageDeviceOpsList + plural: storagedeviceops + shortNames: + - sdops + singular: storagedeviceops + scope: Namespaced + versions: + - additionalPrinterColumns: + - jsonPath: .spec.deviceRef + name: Device + type: string + - jsonPath: .spec.action + name: Action + type: string + - jsonPath: .status.phase + name: Phase + type: string + - jsonPath: .status.step.state + name: Step + type: string + - jsonPath: .status.message + name: Message + priority: 1 + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha2 + schema: + openAPIV3Schema: + description: StorageDeviceOps is a single operation performed against one + StorageDevice. + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: StorageDeviceOpsSpec is one operation to perform against + one StorageDevice. + properties: + abort: + description: |- + Abort asks a running operation to stop at its next step and unwind. + + Restart can be aborted before its call is issued and not after: a restart + the control plane has accepted is one nothing can recall, so the graph + declares where the edge exists rather than this field promising one. + type: boolean + action: + description: Action is the operation to perform. + enum: + - Restart + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + deviceRef: + description: |- + DeviceRef names the StorageDevice this operation acts on, in this + operation's own namespace. The operation never owns its target, because + deleting the record of an operation must not delete the device record it + operated on. + maxLength: 253 + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + required: + - action + - deviceRef + type: object + status: + description: StorageDeviceOpsStatus is the observed state of one device + operation. + properties: + completedAt: + description: CompletedAt is when it reached a terminal phase. + format: date-time + type: string + deviceStatusBefore: + description: |- + DeviceStatusBefore is what the control plane reported the device's status + to be when the operation took its lock, so a wait can tell the device + coming back from its never having gone. + type: string + message: + description: |- + Message is the reason the phase is what it is: one sentence, replaced as + the operation moves, and never a log. + type: string + observedGeneration: + description: |- + ObservedGeneration is the generation the rest of this status was computed + from, so a stale status can be told from a current one. + format: int64 + type: integer + phase: + description: Phase is the operation's own progress. + enum: + - Pending + - Running + - Succeeded + - Failed + - Aborted + type: string + startedAt: + description: StartedAt is when the operation acquired its target's + lock. + format: date-time + type: string + step: + description: |- + Step is the position of the running action's state machine. It is + persisted before the side effect that step performs. + properties: + deadline: + description: |- + Deadline is when that state expires, absent when it has none. It is an + absolute instant, so a state whose deadline passed while the controller + was down restores as already expired. + format: date-time + type: string + state: + description: |- + State is the state the machine was in. Empty means the resource has not + been reconciled yet, and restores to the graph's initial state. + type: string + type: object + x-kubernetes-validations: + - message: unknown step + rule: '!has(self.state) || self.state in [''Requesting'',''Awaiting'']' + type: object + type: object + served: true + storage: true + subresources: + status: {} diff --git a/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml b/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml index 25192e647..a01259b53 100644 --- a/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml @@ -286,6 +286,7 @@ rules: - storagebackups - storageclusterops - storageclusters + - storagedeviceops - storagedevices - storagenodeops - storagenodes @@ -320,6 +321,7 @@ rules: - storagebackuppolicies/finalizers - storageclusterops/finalizers - storageclusters/finalizers + - storagedeviceops/finalizers - storagenodeops/finalizers - storagenodes/finalizers - storagepoolops/finalizers @@ -349,6 +351,7 @@ rules: - storagebackups/status - storageclusterops/status - storageclusters/status + - storagedeviceops/status - storagedevices/status - storagenodeops/status - storagenodes/status diff --git a/operator/api/v1alpha2/storagedeviceops_types.go b/operator/api/v1alpha2/storagedeviceops_types.go new file mode 100644 index 000000000..34a94c637 --- /dev/null +++ b/operator/api/v1alpha2/storagedeviceops_types.go @@ -0,0 +1,248 @@ +// StorageDeviceOps: one operation performed against one StorageDevice. +// +// It is the narrowest blast radius in the ownership spine. A wedged device is +// recycled today by restarting its storage node, which takes every other device +// on that node with it and costs the cluster a node's worth of redundancy for +// the duration; one device is the narrowest thing that can be recycled, and +// choosing the narrowest resource that achieves an outcome is the rule +// design-crd-model.md §8.2 states. +// +// design-storagedevice.md §6 specifies five actions. One is served by the +// control plane's v2 API and is built; the other four are blocked on verbs that +// API does not offer, and [ExternalDependencies] is the list, so the ask is a +// value in this repository rather than a sentence in a document. +// +// **The enum admits only what the operator can perform.** Declaring the other +// four now would accept an object whose first reconcile can only fail, and an +// API that takes a request it will never carry out is worse than one that +// refuses it at admission — the refusal names the missing capability, and the +// failure names nothing a user can act on. Widening an enum is additive, so +// each action arrives with the endpoint it needs. + +package v1alpha2 + +import ( + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + + "github.com/simplyblock/atlas/statemachine" +) + +// StorageDeviceOpsAction is the operation a StorageDeviceOps performs. +// +// There is no Add: a device that appears is discovered (§5.1). There is no bare +// Remove either: a removal is a step of Replace and of Migrate, and taking a +// device out of the data path without replacing it is Fail (§6.4). +// +kubebuilder:validation:Enum=Restart +type StorageDeviceOpsAction string + +const ( + // StorageDeviceOpsActionRestart is the action the kind exists for: + // recycling one device rather than its node. + StorageDeviceOpsActionRestart StorageDeviceOpsAction = "Restart" + + // The four actions §6 specifies and the API cannot serve. They are declared + // so that the names exist where the reason does, and they are absent from + // the Enum marker above: an object naming one is refused at admission + // rather than accepted and failed. + StorageDeviceOpsActionSelfTest StorageDeviceOpsAction = "SelfTest" + StorageDeviceOpsActionFail StorageDeviceOpsAction = "Fail" + StorageDeviceOpsActionReplace StorageDeviceOpsAction = "Replace" + StorageDeviceOpsActionMigrate StorageDeviceOpsAction = "Migrate" +) + +// ExternalDependency is one action this kind specifies, the control-plane +// capability it waits on, and what its absence costs. +// +// It is data rather than prose because the list is an ask of another team and a +// checklist for this one: each row names the endpoint to build, and the action +// ships when the row does. A design paragraph saying the same thing cannot be +// enumerated, cannot be tested against the enum, and goes stale the day one +// endpoint arrives. +type ExternalDependency struct { + // Action is what this unblocks. + Action StorageDeviceOpsAction + + // Endpoint is the v2 API call the action issues, in the shape + // design-storagedevice.md §7 asks for. + Endpoint string + + // Because says what the action does with it, so the ask carries its own + // justification rather than a section number. + Because string +} + +// ExternalDependencies are the four actions of §6 that the v2 API cannot serve +// today, and the verb each one needs. +// +// The API offers three device verbs — restart, remove, and reset — where §7 asks +// for seven. Restart is built on the first. Remove exists and is not enough on +// its own: it is a step of Replace and of Migrate, and both of those also need +// an adopt call to name the device that arrives, so the verb that exists buys +// neither action. +// +// Migrate's row is the one whose absence removes an action rather than degrading +// it. Attaching asks the control plane to accept a device as another node's with +// its contents intact, so the chunks on it are re-homed rather than rebuilt. A +// control plane that can only adopt a device as empty turns Migrate into two +// Replaces and a full rebuild, which is the thing §6.2 gives as the reason the +// action exists. +func ExternalDependencies() []ExternalDependency { + const base = "POST /api/v2/clusters/{cluster}/storage-nodes/{node}/devices/" + return []ExternalDependency{ + { + Action: StorageDeviceOpsActionSelfTest, + Endpoint: base + "{device}/self-test", + Because: "the action runs the device's own self-test and reports the verdict, " + + "with the short or extended mode in the body", + }, + { + Action: StorageDeviceOpsActionFail, + Endpoint: base + "{device}/fail", + Because: "the action takes a device out of the data path and leaves it in the " + + "slot, so the cluster rebuilds its redundancy elsewhere and stops reading " + + "from a device somebody has judged untrustworthy", + }, + { + Action: StorageDeviceOpsActionReplace, + Endpoint: base + "adopt", + Because: "the action pairs a removal with an arrival, and the removal verb " + + "exists while the call naming the device that arrived does not", + }, + { + Action: StorageDeviceOpsActionMigrate, + Endpoint: base + "{device}/detach, and " + base + "adopt accepting a device with its contents", + Because: "the action moves the drive to another node with what is on it; an " + + "adopt that can only take an empty device makes this two replacements and " + + "a full rebuild, which is what the action exists to avoid", + }, + } +} + +// StorageDeviceOpsPhase is the operation's own progress. +// +kubebuilder:validation:Enum=Pending;Running;Succeeded;Failed;Aborted +type StorageDeviceOpsPhase string + +const ( + StorageDeviceOpsPhasePending StorageDeviceOpsPhase = "Pending" + StorageDeviceOpsPhaseRunning StorageDeviceOpsPhase = "Running" + StorageDeviceOpsPhaseSucceeded StorageDeviceOpsPhase = "Succeeded" + StorageDeviceOpsPhaseFailed StorageDeviceOpsPhase = "Failed" + StorageDeviceOpsPhaseAborted StorageDeviceOpsPhase = "Aborted" +) + +// StorageDeviceOpsStep is one step of a running device operation. Which steps +// belong to which action is declared by that action's graph rather than by this +// type, which is why the enum stays flat as actions are added. +// +// It carries the two steps Restart has. The eight the other four actions need +// arrive with those actions, because a step no graph declares is a status value +// nothing can resume from and a CEL rule that admits one is a rule that admits +// nonsense. +// +kubebuilder:validation:Enum=Requesting;Awaiting +type StorageDeviceOpsStep string + +const ( + // StorageDeviceOpsStepRequesting issues the call. + StorageDeviceOpsStepRequesting StorageDeviceOpsStep = "Requesting" + + // StorageDeviceOpsStepAwaiting waits for the control plane to report the + // device in the state the call asked for. + StorageDeviceOpsStepAwaiting StorageDeviceOpsStep = "Awaiting" +) + +// StorageDeviceOpsSpec is one operation to perform against one StorageDevice. +type StorageDeviceOpsSpec struct { + // DeviceRef names the StorageDevice this operation acts on, in this + // operation's own namespace. The operation never owns its target, because + // deleting the record of an operation must not delete the device record it + // operated on. + // +kubebuilder:validation:MaxLength=253 + // +kubebuilder:validation:Required + // +k8s:immutable + DeviceRef string `json:"deviceRef"` + + // Action is the operation to perform. + // +kubebuilder:validation:Required + // +k8s:immutable + Action StorageDeviceOpsAction `json:"action"` + + // Abort asks a running operation to stop at its next step and unwind. + // + // Restart can be aborted before its call is issued and not after: a restart + // the control plane has accepted is one nothing can recall, so the graph + // declares where the edge exists rather than this field promising one. + // +optional + Abort bool `json:"abort,omitempty"` +} + +// StorageDeviceOpsStatus is the observed state of one device operation. +type StorageDeviceOpsStatus struct { + // Phase is the operation's own progress. + // +optional + Phase StorageDeviceOpsPhase `json:"phase,omitempty"` + + // Step is the position of the running action's state machine. It is + // persisted before the side effect that step performs. + // +kubebuilder:validation:XValidation:rule="!has(self.state) || self.state in ['Requesting','Awaiting']",message="unknown step" + // +optional + Step statemachine.KubeSnapshot `json:"step,omitempty"` + + // DeviceStatusBefore is what the control plane reported the device's status + // to be when the operation took its lock, so a wait can tell the device + // coming back from its never having gone. + // +optional + DeviceStatusBefore string `json:"deviceStatusBefore,omitempty"` + + // Message is the reason the phase is what it is: one sentence, replaced as + // the operation moves, and never a log. + // +optional + Message string `json:"message,omitempty"` + + // ObservedGeneration is the generation the rest of this status was computed + // from, so a stale status can be told from a current one. + // +optional + ObservedGeneration int64 `json:"observedGeneration,omitempty"` + + // StartedAt is when the operation acquired its target's lock. + // +optional + StartedAt *metav1.Time `json:"startedAt,omitempty"` + + // CompletedAt is when it reached a terminal phase. + // +optional + CompletedAt *metav1.Time `json:"completedAt,omitempty"` +} + +// v1alpha2 is the only version this kind has ever had, so it is the storage +// version and converts nothing. +// +kubebuilder:storageversion +// +kubebuilder:object:root=true +// +kubebuilder:subresource:status +// +kubebuilder:resource:scope=Namespaced,shortName=sdops +// +kubebuilder:printcolumn:name="Device",type=string,JSONPath=".spec.deviceRef" +// +kubebuilder:printcolumn:name="Action",type=string,JSONPath=".spec.action" +// +kubebuilder:printcolumn:name="Phase",type=string,JSONPath=".status.phase" +// +kubebuilder:printcolumn:name="Step",type=string,JSONPath=".status.step.state" +// +kubebuilder:printcolumn:name="Message",type=string,JSONPath=".status.message",priority=1 +// +kubebuilder:printcolumn:name="Age",type=date,JSONPath=".metadata.creationTimestamp" + +// StorageDeviceOps is a single operation performed against one StorageDevice. +type StorageDeviceOps struct { + metav1.TypeMeta `json:",inline"` + metav1.ObjectMeta `json:"metadata,omitempty"` + + Spec StorageDeviceOpsSpec `json:"spec,omitempty"` + Status StorageDeviceOpsStatus `json:"status,omitempty"` +} + +// +kubebuilder:object:root=true + +// StorageDeviceOpsList contains a list of StorageDeviceOps. +type StorageDeviceOpsList struct { + metav1.TypeMeta `json:",inline"` + metav1.ListMeta `json:"metadata,omitempty"` + Items []StorageDeviceOps `json:"items"` +} + +func init() { + SchemeBuilder.Register(&StorageDeviceOps{}, &StorageDeviceOpsList{}) +} diff --git a/operator/api/v1alpha2/zz_generated.deepcopy.go b/operator/api/v1alpha2/zz_generated.deepcopy.go index eeabb991d..65d0b700f 100644 --- a/operator/api/v1alpha2/zz_generated.deepcopy.go +++ b/operator/api/v1alpha2/zz_generated.deepcopy.go @@ -806,6 +806,21 @@ func (in *DriverTLS) DeepCopy() *DriverTLS { return out } +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *ExternalDependency) DeepCopyInto(out *ExternalDependency) { + *out = *in +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new ExternalDependency. +func (in *ExternalDependency) DeepCopy() *ExternalDependency { + if in == nil { + return nil + } + out := new(ExternalDependency) + in.DeepCopyInto(out) + return out +} + // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *FoundationDBSpec) DeepCopyInto(out *FoundationDBSpec) { *out = *in @@ -2449,6 +2464,104 @@ func (in *StorageDeviceList) DeepCopyObject() runtime.Object { return nil } +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *StorageDeviceOps) DeepCopyInto(out *StorageDeviceOps) { + *out = *in + out.TypeMeta = in.TypeMeta + in.ObjectMeta.DeepCopyInto(&out.ObjectMeta) + out.Spec = in.Spec + in.Status.DeepCopyInto(&out.Status) +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new StorageDeviceOps. +func (in *StorageDeviceOps) DeepCopy() *StorageDeviceOps { + if in == nil { + return nil + } + out := new(StorageDeviceOps) + in.DeepCopyInto(out) + return out +} + +// DeepCopyObject is an autogenerated deepcopy function, copying the receiver, creating a new runtime.Object. +func (in *StorageDeviceOps) DeepCopyObject() runtime.Object { + if c := in.DeepCopy(); c != nil { + return c + } + return nil +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *StorageDeviceOpsList) DeepCopyInto(out *StorageDeviceOpsList) { + *out = *in + out.TypeMeta = in.TypeMeta + in.ListMeta.DeepCopyInto(&out.ListMeta) + if in.Items != nil { + in, out := &in.Items, &out.Items + *out = make([]StorageDeviceOps, len(*in)) + for i := range *in { + (*in)[i].DeepCopyInto(&(*out)[i]) + } + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new StorageDeviceOpsList. +func (in *StorageDeviceOpsList) DeepCopy() *StorageDeviceOpsList { + if in == nil { + return nil + } + out := new(StorageDeviceOpsList) + in.DeepCopyInto(out) + return out +} + +// DeepCopyObject is an autogenerated deepcopy function, copying the receiver, creating a new runtime.Object. +func (in *StorageDeviceOpsList) DeepCopyObject() runtime.Object { + if c := in.DeepCopy(); c != nil { + return c + } + return nil +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *StorageDeviceOpsSpec) DeepCopyInto(out *StorageDeviceOpsSpec) { + *out = *in +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new StorageDeviceOpsSpec. +func (in *StorageDeviceOpsSpec) DeepCopy() *StorageDeviceOpsSpec { + if in == nil { + return nil + } + out := new(StorageDeviceOpsSpec) + in.DeepCopyInto(out) + return out +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *StorageDeviceOpsStatus) DeepCopyInto(out *StorageDeviceOpsStatus) { + *out = *in + in.Step.DeepCopyInto(&out.Step) + if in.StartedAt != nil { + in, out := &in.StartedAt, &out.StartedAt + *out = (*in).DeepCopy() + } + if in.CompletedAt != nil { + in, out := &in.CompletedAt, &out.CompletedAt + *out = (*in).DeepCopy() + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new StorageDeviceOpsStatus. +func (in *StorageDeviceOpsStatus) DeepCopy() *StorageDeviceOpsStatus { + if in == nil { + return nil + } + out := new(StorageDeviceOpsStatus) + in.DeepCopyInto(out) + return out +} + // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *StorageDeviceSpec) DeepCopyInto(out *StorageDeviceSpec) { *out = *in diff --git a/operator/cmd/main.go b/operator/cmd/main.go index 186d224f2..e7c27cb23 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -779,6 +779,19 @@ func main() { // (design-simplyblockdriver.md §4.1). The discovery client is the same one // the discovery run uses; a driver reconcile without it refuses rather than // guessing, because guessing either way breaks a cluster in one direction. + // One device is the narrowest thing that can be recycled, which is the whole + // of why the kind exists (design-storagedevice.md §6). Its four other actions + // wait on control-plane verbs the v2 API does not offer, and + // v1alpha2.ExternalDependencies is the list. + if err := (&nodecontroller.StorageDeviceOpsReconciler{ + Client: mgr.GetClient(), + Scheme: mgr.GetScheme(), + Recorder: mgr.GetEventRecorder("storagedeviceops-controller"), + API: backupAPI, + }).SetupWithManager(mgr); err != nil { + setupLog.Error(err, "unable to create controller", "controller", "StorageDeviceOps") + os.Exit(1) + } if err := (&driver.SimplyblockDriverReconciler{ Client: mgr.GetClient(), Scheme: mgr.GetScheme(), diff --git a/operator/config/crd/bases/storage.simplyblock.io_storagedeviceops.yaml b/operator/config/crd/bases/storage.simplyblock.io_storagedeviceops.yaml new file mode 100644 index 000000000..eec5d0629 --- /dev/null +++ b/operator/config/crd/bases/storage.simplyblock.io_storagedeviceops.yaml @@ -0,0 +1,162 @@ +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition +metadata: + annotations: + controller-gen.kubebuilder.io/version: v0.21.0 + name: storagedeviceops.storage.simplyblock.io +spec: + group: storage.simplyblock.io + names: + kind: StorageDeviceOps + listKind: StorageDeviceOpsList + plural: storagedeviceops + shortNames: + - sdops + singular: storagedeviceops + scope: Namespaced + versions: + - additionalPrinterColumns: + - jsonPath: .spec.deviceRef + name: Device + type: string + - jsonPath: .spec.action + name: Action + type: string + - jsonPath: .status.phase + name: Phase + type: string + - jsonPath: .status.step.state + name: Step + type: string + - jsonPath: .status.message + name: Message + priority: 1 + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha2 + schema: + openAPIV3Schema: + description: StorageDeviceOps is a single operation performed against one + StorageDevice. + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: StorageDeviceOpsSpec is one operation to perform against + one StorageDevice. + properties: + abort: + description: |- + Abort asks a running operation to stop at its next step and unwind. + + Restart can be aborted before its call is issued and not after: a restart + the control plane has accepted is one nothing can recall, so the graph + declares where the edge exists rather than this field promising one. + type: boolean + action: + description: Action is the operation to perform. + enum: + - Restart + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + deviceRef: + description: |- + DeviceRef names the StorageDevice this operation acts on, in this + operation's own namespace. The operation never owns its target, because + deleting the record of an operation must not delete the device record it + operated on. + maxLength: 253 + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + required: + - action + - deviceRef + type: object + status: + description: StorageDeviceOpsStatus is the observed state of one device + operation. + properties: + completedAt: + description: CompletedAt is when it reached a terminal phase. + format: date-time + type: string + deviceStatusBefore: + description: |- + DeviceStatusBefore is what the control plane reported the device's status + to be when the operation took its lock, so a wait can tell the device + coming back from its never having gone. + type: string + message: + description: |- + Message is the reason the phase is what it is: one sentence, replaced as + the operation moves, and never a log. + type: string + observedGeneration: + description: |- + ObservedGeneration is the generation the rest of this status was computed + from, so a stale status can be told from a current one. + format: int64 + type: integer + phase: + description: Phase is the operation's own progress. + enum: + - Pending + - Running + - Succeeded + - Failed + - Aborted + type: string + startedAt: + description: StartedAt is when the operation acquired its target's + lock. + format: date-time + type: string + step: + description: |- + Step is the position of the running action's state machine. It is + persisted before the side effect that step performs. + properties: + deadline: + description: |- + Deadline is when that state expires, absent when it has none. It is an + absolute instant, so a state whose deadline passed while the controller + was down restores as already expired. + format: date-time + type: string + state: + description: |- + State is the state the machine was in. Empty means the resource has not + been reconciled yet, and restores to the graph's initial state. + type: string + type: object + x-kubernetes-validations: + - message: unknown step + rule: '!has(self.state) || self.state in [''Requesting'',''Awaiting'']' + type: object + type: object + served: true + storage: true + subresources: + status: {} diff --git a/operator/config/rbac/role.yaml b/operator/config/rbac/role.yaml index 08e9e1de4..0090a347e 100644 --- a/operator/config/rbac/role.yaml +++ b/operator/config/rbac/role.yaml @@ -286,6 +286,7 @@ rules: - storagebackups - storageclusterops - storageclusters + - storagedeviceops - storagedevices - storagenodeops - storagenodes @@ -320,6 +321,7 @@ rules: - storagebackuppolicies/finalizers - storageclusterops/finalizers - storageclusters/finalizers + - storagedeviceops/finalizers - storagenodeops/finalizers - storagenodes/finalizers - storagepoolops/finalizers @@ -349,6 +351,7 @@ rules: - storagebackups/status - storageclusterops/status - storageclusters/status + - storagedeviceops/status - storagedevices/status - storagenodeops/status - storagenodes/status diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index e00b607ab..669dcc68f 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -10669,6 +10669,7 @@ rules: - storagebackups - storageclusterops - storageclusters + - storagedeviceops - storagedevices - storagenodeops - storagenodes @@ -10703,6 +10704,7 @@ rules: - storagebackuppolicies/finalizers - storageclusterops/finalizers - storageclusters/finalizers + - storagedeviceops/finalizers - storagenodeops/finalizers - storagenodes/finalizers - storagepoolops/finalizers @@ -10732,6 +10734,7 @@ rules: - storagebackups/status - storageclusterops/status - storageclusters/status + - storagedeviceops/status - storagedevices/status - storagenodeops/status - storagenodes/status diff --git a/operator/docs/designs/crd-redesign/design-storagedevice.md b/operator/docs/designs/crd-redesign/design-storagedevice.md index 39ec41e22..eb6d34934 100644 --- a/operator/docs/designs/crd-redesign/design-storagedevice.md +++ b/operator/docs/designs/crd-redesign/design-storagedevice.md @@ -552,12 +552,25 @@ Declared in `operator/api/v1alpha2/storagedeviceops_types.go`, short name `operator/internal/controller/storagedeviceops_controller.go`. The type is Appendix B. -**Nothing in this section is implemented.** The device's own lock is declared and -empty (§4.2), the two operation metrics of §8.2 have no source, and §11 Q1 is the -open question the actions wait on. +**One of the five actions is implemented, and it is the one the kind exists +for.** `Restart` is built: the kind, its graph, the reconciler, and the device +lock §4.2 declared empty against this section arriving. The other four are +blocked on control-plane verbs the v2 API does not offer, and the ask is +`v1alpha2.ExternalDependencies` — a value in the repository rather than a +paragraph here, so each row names the endpoint to build and the action ships when +the row does. + +**The enum admits only what the operator can perform.** Declaring all five now +would accept an object whose first reconcile can only fail, and an API that takes +a request it will never carry out is worse than one that refuses it at admission: +the refusal names the missing capability, and the failure names nothing a user +can act on. Widening an enum is additive, so each action arrives with its +endpoint. ```go -// +kubebuilder:validation:Enum=Restart;SelfTest;Fail;Replace;Migrate +// What ships today. The other four constants are declared without being in the +// marker, so the names exist where the reasons do. +// +kubebuilder:validation:Enum=Restart type StorageDeviceOpsAction string ``` @@ -797,6 +810,13 @@ than degrading it. `Adding` step is the adopt call, and its `Rebuilding` step is a wait on the device stream. The action is a graph over calls the other actions already need. +**Measured against the shipped v2 API, three device verbs exist and four do +not.** `restart`, `remove`, and `reset` are served; `self-test`, `fail`, +`detach`, and `adopt` are not. `Restart` is built on the first. `remove` exists +and buys no action on its own: it is a step of `Replace` and of `Migrate`, and +both also need the adopt call that names the device arriving. `reset` is a verb +this document did not ask for and no action uses. + **The device stream and the hardware fields are in use, and the per-action verbs are what remain.** The stream reports a PCI address, a serial number, a model, an NVMe controller, a status, the health signals of §4.2, and a size for each device, which diff --git a/operator/internal/controllers/node/storagedeviceops_controller.go b/operator/internal/controllers/node/storagedeviceops_controller.go new file mode 100644 index 000000000..dfda561bc --- /dev/null +++ b/operator/internal/controllers/node/storagedeviceops_controller.go @@ -0,0 +1,556 @@ +// The reconciler for StorageDeviceOps, which today means one action: restarting +// one device rather than the storage node it sits in. +// +// The operation holds its device's lock for as long as it runs, which is +// status.activeOpsRef on the StorageDevice — the field design-storagedevice.md +// §4.2 declared empty against this kind arriving. Two operations on one device +// would be two restarts of one controller, and the second would be issued +// against a device that is halfway through the first. +// +// Nothing here blocks. A step that is not finished requeues, and the step it is +// on is in the status, so a controller restart resumes rather than restarts. +// +// design-storagedevice.md §6 is the specification, and §6's other four actions +// are blocked on control-plane verbs that do not exist +// (v1alpha2.ExternalDependencies). + +package node + +import ( + "context" + "errors" + "fmt" + "time" + + corev1 "k8s.io/api/core/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/runtime" + "k8s.io/client-go/tools/events" + "k8s.io/client-go/util/retry" + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil" + logf "sigs.k8s.io/controller-runtime/pkg/log" + + "github.com/simplyblock/atlas/controlplane" + "github.com/simplyblock/atlas/errs" + "github.com/simplyblock/atlas/statemachine" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +const ( + // deviceOpsFinalizer holds an operation until it has released its device's + // lock. A delete arriving mid-restart would otherwise leave the device + // pointing at an operation nobody can read, and the next operation waiting + // on a holder that does not exist. + deviceOpsFinalizer = "storage.simplyblock.io/storagedeviceops-finalizer" + + // deviceOpsRetry is how often a waiting operation looks again. + deviceOpsRetry = 10 * time.Second +) + +// The reasons a device operation emits. +const ( + DeviceOperationStarted = "OperationStarted" + DeviceOperationSucceeded = "OperationSucceeded" + DeviceOperationFailed = "OperationFailed" + DeviceOperationAborted = "OperationAborted" + DeviceRestartRequested = "DeviceRestartRequested" + DeviceStepDeadlineGone = "StepDeadlineExceeded" +) + +// DeviceClient is what the operation asks of the control plane. It is an +// interface rather than the client so the reconciler is testable against a +// control plane that refuses, which is half of what these steps are about. +type DeviceClient interface { + // Device reads one device, wrapping errs.ErrNotFound for one the control + // plane does not hold. + Device(ctx context.Context, clusterID, nodeID, deviceID string) (controlplane.Device, error) + + // RestartDevice recycles one device in place. + RestartDevice(ctx context.Context, clusterID, nodeID, deviceID string) error +} + +// StorageDeviceOpsReconciler runs operations against one device. +type StorageDeviceOpsReconciler struct { + client.Client + Scheme *runtime.Scheme + Recorder events.EventRecorder + + // API is the control plane the operation issues its call to. + API DeviceClient +} + +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagedeviceops,verbs=get;list;watch;create;update;patch;delete +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagedeviceops/status,verbs=get;update;patch +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagedeviceops/finalizers,verbs=update +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagedevices,verbs=get;list;watch +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagedevices/status,verbs=get;update;patch + +// Reconcile advances one device operation by one step. +func (r *StorageDeviceOpsReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) { + var ops simplyblockv1alpha2.StorageDeviceOps + if err := r.Get(ctx, req.NamespacedName, &ops); err != nil { + return ctrl.Result{}, client.IgnoreNotFound(err) + } + + if !ops.DeletionTimestamp.IsZero() { + return ctrl.Result{}, r.finalize(ctx, &ops) + } + if !controllerutil.ContainsFinalizer(&ops, deviceOpsFinalizer) { + controllerutil.AddFinalizer(&ops, deviceOpsFinalizer) + return ctrl.Result{}, r.Update(ctx, &ops) + } + + // Terminal and staying that way: the operation is the audit record now. + switch ops.Status.Phase { + case simplyblockv1alpha2.StorageDeviceOpsPhaseSucceeded, + simplyblockv1alpha2.StorageDeviceOpsPhaseFailed, + simplyblockv1alpha2.StorageDeviceOpsPhaseAborted: + return ctrl.Result{}, nil + } + + device, err := r.target(ctx, &ops) + if err != nil { + var refusal *deviceRefusal + if errors.As(err, &refusal) { + return ctrl.Result{}, r.fail(ctx, &ops, refusal.Error()) + } + return ctrl.Result{}, err + } + + held, err := r.acquireLock(ctx, &ops, device) + if err != nil { + return ctrl.Result{}, err + } + if !held { + return ctrl.Result{RequeueAfter: deviceOpsRetry}, r.note(ctx, &ops, + fmt.Sprintf("waiting for %s to finish with device %s", + device.Status.ActiveOpsRef, device.Name)) + } + + return r.advance(ctx, &ops, device) +} + +// advance runs the action's graph forward by at most one step. +func (r *StorageDeviceOpsReconciler) advance( + ctx context.Context, + ops *simplyblockv1alpha2.StorageDeviceOps, + device *simplyblockv1alpha2.StorageDevice, +) (ctrl.Result, error) { + graph, declared := storageDeviceOpsGraphs()[statemachine.Action(ops.Spec.Action)] + if !declared { + // Admission's enum admits only the actions this operator performs, so + // reaching here means an older CRD served the object. The message names + // the dependency rather than the enum, because what a user needs to + // know is that the action is not available yet. + return ctrl.Result{}, r.fail(ctx, ops, unavailableAction(ops.Spec.Action)) + } + + machine, err := statemachine.NewFromSnapshot(ctx, graph, + statemachine.FromKube[deviceStep](ops.Status.Step)) + if err != nil { + return ctrl.Result{}, r.fail(ctx, ops, + fmt.Sprintf("the operation cannot be resumed: %v", err)) + } + defer machine.Close() + + if ops.Status.Step.State == "" { + return r.begin(ctx, ops, device, machine.CurrentState()) + } + + current := machine.CurrentState() + if ops.Spec.Abort { + if !machine.CanAbort() { + return ctrl.Result{RequeueAfter: deviceOpsRetry}, r.note(ctx, ops, fmt.Sprintf( + "an abort was asked for and step %s cannot be stopped; the restart has been "+ + "issued and nothing recalls one", current)) + } + return ctrl.Result{}, r.abort(ctx, ops, device, current) + } + + if machine.TimeoutReached() { + expired := deviceTimeoutMessage(current, device.Name) + r.event(ops, corev1.EventTypeWarning, DeviceStepDeadlineGone, expired) + return ctrl.Result{}, r.fail(ctx, ops, expired) + } + + done, err := r.performStep(ctx, ops, device, current) + if err != nil { + var refusal *deviceRefusal + if errors.As(err, &refusal) { + return ctrl.Result{}, r.fail(ctx, ops, refusal.Error()) + } + return ctrl.Result{}, err + } + if !done { + return ctrl.Result{RequeueAfter: deviceOpsRetry}, nil + } + + if machine.IsTerminal() { + return ctrl.Result{}, r.succeed(ctx, ops, device) + } + + next, ok := nextDeviceStep(machine) + if !ok { + return ctrl.Result{}, r.fail(ctx, ops, + fmt.Sprintf("step %s declares no successor and is not terminal", current)) + } + if err := machine.TransitionTo(ctx, next); err != nil { + return ctrl.Result{}, fmt.Errorf("enter step %s: %w", next, err) + } + return ctrl.Result{RequeueAfter: deviceOpsRetry}, r.enter(ctx, ops, next, + statemachine.ToKube(machine.Snapshot()).Deadline) +} + +// performStep runs one step and reports whether it has finished. +func (r *StorageDeviceOpsReconciler) performStep( + ctx context.Context, + ops *simplyblockv1alpha2.StorageDeviceOps, + device *simplyblockv1alpha2.StorageDevice, + current deviceStep, +) (bool, error) { + switch current { + case stepDeviceRequesting: + return r.request(ctx, ops, device) + case stepDeviceAwaiting: + return r.await(ctx, ops, device) + default: + return false, fmt.Errorf("step %s belongs to no operation this operator performs", current) + } +} + +// request issues the restart. +// +// It is issued once per operation and not once per pass. The step is persisted +// before the call, so re-entering it means the previous pass died between the +// write and the call — and a restart issued twice is a device recycled twice, +// which is the failure the write-ahead record exists to avoid rather than one +// to make idempotent. +func (r *StorageDeviceOpsReconciler) request( + ctx context.Context, + ops *simplyblockv1alpha2.StorageDeviceOps, + device *simplyblockv1alpha2.StorageDevice, +) (bool, error) { + if ops.Status.StartedAt != nil && ops.Status.DeviceStatusBefore != "" { + // The call was issued on an earlier pass and its record survived. + return true, nil + } + + cluster, node, id := device.Status.ClusterID, device.Status.NodeID, device.Spec.DeviceID + if cluster == "" || node == "" || id == "" { + return false, refuseDevice( + "device %s does not report which cluster, node, and device it is, so there is "+ + "nothing to address the restart to", device.Name) + } + + before := device.Status.DeviceStatus + if err := r.writeStatus(ctx, ops, func(status *simplyblockv1alpha2.StorageDeviceOpsStatus) { + status.DeviceStatusBefore = before + }); err != nil { + return false, err + } + + if err := r.API.RestartDevice(ctx, cluster, node, id); err != nil { + return false, refuseDevice( + "the control plane refused to restart device %s: %v", device.Name, err) + } + r.event(ops, corev1.EventTypeNormal, DeviceRestartRequested, + fmt.Sprintf("asked the control plane to restart device %s", device.Name)) + return true, nil +} + +// await waits for the control plane to report the device back in service. +// +// What it waits for is the status the device reports rather than an absence: +// a restart takes the device out and brings it back, and a check that only +// looked for it being gone would finish on the way down. +func (r *StorageDeviceOpsReconciler) await( + ctx context.Context, + ops *simplyblockv1alpha2.StorageDeviceOps, + device *simplyblockv1alpha2.StorageDevice, +) (bool, error) { + current, err := r.API.Device(ctx, + device.Status.ClusterID, device.Status.NodeID, device.Spec.DeviceID) + switch { + case errors.Is(err, errs.ErrNotFound): + // A device the control plane has stopped holding did not come back + // from the restart, which is a finding rather than a wait: §5.2 deletes + // the object when the node stops reporting it, and this operation would + // otherwise sit until its deadline describing a device that is gone. + return false, refuseDevice( + "device %s is no longer held by the control plane, so the restart did not "+ + "bring it back", device.Name) + case err != nil: + return false, fmt.Errorf("read device %s back: %w", device.Name, err) + } + + if !deviceIsInService(current.Status) { + return false, r.note(ctx, ops, fmt.Sprintf( + "device %s reports %q; waiting for it to come back", device.Name, current.Status)) + } + return true, nil +} + +// deviceIsInService reads the control plane's own vocabulary for a device that +// is serving. +// +// The spelling is the control plane's and is compared rather than mapped, for +// the reason status.deviceStatus keeps it: a vocabulary translated here would be +// one this wait could not express, and the set of statuses is the backend's to +// grow. +func deviceIsInService(status string) bool { + return status == "online" +} + +// target resolves the StorageDevice the operation names. +func (r *StorageDeviceOpsReconciler) target( + ctx context.Context, ops *simplyblockv1alpha2.StorageDeviceOps, +) (*simplyblockv1alpha2.StorageDevice, error) { + var device simplyblockv1alpha2.StorageDevice + key := client.ObjectKey{Namespace: ops.Namespace, Name: ops.Spec.DeviceRef} + err := r.Get(ctx, key, &device) + switch { + case apierrors.IsNotFound(err): + return nil, refuseDevice( + "spec.deviceRef names StorageDevice %q and there is none by that name in %s", + ops.Spec.DeviceRef, ops.Namespace) + case err != nil: + return nil, err + } + return &device, nil +} + +// unavailableAction is what an operation naming an action this operator cannot +// perform is failed with. It names the endpoint rather than the enum, because +// the useful fact is which capability is missing. +func unavailableAction(action simplyblockv1alpha2.StorageDeviceOpsAction) string { + for _, dependency := range simplyblockv1alpha2.ExternalDependencies() { + if dependency.Action == action { + return fmt.Sprintf("action %s is not available: it needs %s, which the control "+ + "plane does not offer", action, dependency.Endpoint) + } + } + return fmt.Sprintf("action %q is not one this operator performs", action) +} + +// deviceTimeoutMessage says what a step outliving its deadline means, which +// differs by step: one is a control plane that did not answer, the other is a +// device that did not come back. +func deviceTimeoutMessage(step deviceStep, name string) string { + if step == stepDeviceAwaiting { + return fmt.Sprintf("device %s did not come back within %s of being restarted", + name, awaitingDeviceDeadline) + } + return fmt.Sprintf("the control plane did not accept the restart of device %s within %s", + name, requestingDeviceDeadline) +} + +// nextDeviceStep is the step that follows the current one. The graph is a line, +// so the first edge is the only edge. +func nextDeviceStep(machine *statemachine.Machine[deviceStep]) (deviceStep, bool) { + for next := range machine.AllowedTransitions() { + return next, true + } + return machine.CurrentState(), false +} + +// deviceRefusal is a step's own refusal: a state the operation cannot proceed +// from, rather than an error to retry against the backoff. +// +// It carries no reason of its own. Every refusal here ends the operation, so the +// event reason is OperationFailed in each case, and a field that only ever held +// one value would suggest a choice nobody makes. A refusal that needed a reason +// of its own would be one that did something other than fail. +type deviceRefusal struct { + message string +} + +func (e *deviceRefusal) Error() string { return e.message } + +func refuseDevice(format string, args ...any) error { + return &deviceRefusal{message: fmt.Sprintf(format, args...)} +} + +// begin records that the operation has started, in the step the graph begins at +// and with that step's budget. +func (r *StorageDeviceOpsReconciler) begin( + ctx context.Context, + ops *simplyblockv1alpha2.StorageDeviceOps, + device *simplyblockv1alpha2.StorageDevice, + initial deviceStep, +) (ctrl.Result, error) { + now := metav1.Now() + deadline := metav1.NewTime(now.Add(initialDeviceDeadline)) + r.event(ops, corev1.EventTypeNormal, DeviceOperationStarted, + fmt.Sprintf("restarting device %s", device.Name)) + + return ctrl.Result{RequeueAfter: deviceOpsRetry}, r.writeStatus(ctx, ops, + func(status *simplyblockv1alpha2.StorageDeviceOpsStatus) { + status.Phase = simplyblockv1alpha2.StorageDeviceOpsPhaseRunning + status.StartedAt = &now + status.Step.State = string(initial) + status.Step.Deadline = &deadline + status.Message = fmt.Sprintf("restarting device %s", device.Name) + }) +} + +// enter records the step the machine has moved into, with the deadline its entry +// hook set. +func (r *StorageDeviceOpsReconciler) enter( + ctx context.Context, + ops *simplyblockv1alpha2.StorageDeviceOps, + next deviceStep, + deadline *metav1.Time, +) error { + return r.writeStatus(ctx, ops, func(status *simplyblockv1alpha2.StorageDeviceOpsStatus) { + status.Step.State = string(next) + status.Step.Deadline = deadline + }) +} + +// note records what the operation is waiting on, without moving it. +func (r *StorageDeviceOpsReconciler) note( + ctx context.Context, ops *simplyblockv1alpha2.StorageDeviceOps, message string, +) error { + if ops.Status.Message == message { + return nil + } + return r.writeStatus(ctx, ops, func(status *simplyblockv1alpha2.StorageDeviceOpsStatus) { + status.Message = message + }) +} + +// succeed ends an operation that reached the end of its graph, and releases the +// device. +func (r *StorageDeviceOpsReconciler) succeed( + ctx context.Context, + ops *simplyblockv1alpha2.StorageDeviceOps, + device *simplyblockv1alpha2.StorageDevice, +) error { + if err := r.releaseLock(ctx, ops, device); err != nil { + return err + } + message := fmt.Sprintf("device %s was restarted and is back in service", device.Name) + r.event(ops, corev1.EventTypeNormal, DeviceOperationSucceeded, message) + + now := metav1.Now() + return r.writeStatus(ctx, ops, func(status *simplyblockv1alpha2.StorageDeviceOpsStatus) { + status.Phase = simplyblockv1alpha2.StorageDeviceOpsPhaseSucceeded + status.CompletedAt = &now + status.Step.Deadline = nil + status.Message = message + }) +} + +// abort stops an operation in a step that declares the edge, and releases the +// device. +func (r *StorageDeviceOpsReconciler) abort( + ctx context.Context, + ops *simplyblockv1alpha2.StorageDeviceOps, + device *simplyblockv1alpha2.StorageDevice, + current deviceStep, +) error { + if err := r.releaseLock(ctx, ops, device); err != nil { + return err + } + message := fmt.Sprintf("aborted in step %s; no restart had been issued", current) + r.event(ops, corev1.EventTypeNormal, DeviceOperationAborted, message) + + now := metav1.Now() + return r.writeStatus(ctx, ops, func(status *simplyblockv1alpha2.StorageDeviceOpsStatus) { + status.Phase = simplyblockv1alpha2.StorageDeviceOpsPhaseAborted + status.CompletedAt = &now + status.Step.Deadline = nil + status.Message = message + }) +} + +// fail ends an operation with a reason, and releases whatever it holds. +// +// The release is best effort and the failure is recorded either way: an +// operation that could not let go of its device is a worse thing to hide than to +// report, and the next operation's stale-lock check is what recovers from it. +func (r *StorageDeviceOpsReconciler) fail( + ctx context.Context, ops *simplyblockv1alpha2.StorageDeviceOps, message string, +) error { + if device, err := r.target(ctx, ops); err == nil { + if err := r.releaseLock(ctx, ops, device); err != nil { + logf.FromContext(ctx).Error(err, "could not release the device lock of a failing "+ + "operation", "operation", ops.Name, "device", device.Name) + } + } + r.event(ops, corev1.EventTypeWarning, DeviceOperationFailed, message) + + now := metav1.Now() + return r.writeStatus(ctx, ops, func(status *simplyblockv1alpha2.StorageDeviceOpsStatus) { + status.Phase = simplyblockv1alpha2.StorageDeviceOpsPhaseFailed + status.CompletedAt = &now + status.Step.Deadline = nil + status.Message = message + }) +} + +// finalize releases the device and lets a deleted operation go. +func (r *StorageDeviceOpsReconciler) finalize( + ctx context.Context, ops *simplyblockv1alpha2.StorageDeviceOps, +) error { + if !controllerutil.ContainsFinalizer(ops, deviceOpsFinalizer) { + return nil + } + if device, err := r.target(ctx, ops); err == nil { + if err := r.releaseLock(ctx, ops, device); err != nil { + return err + } + } + controllerutil.RemoveFinalizer(ops, deviceOpsFinalizer) + return r.Update(ctx, ops) +} + +// writeStatus applies a change to the operation's status, retrying a write that +// lost a race. +func (r *StorageDeviceOpsReconciler) writeStatus( + ctx context.Context, + ops *simplyblockv1alpha2.StorageDeviceOps, + change func(*simplyblockv1alpha2.StorageDeviceOpsStatus), +) error { + err := retry.RetryOnConflict(retry.DefaultRetry, func() error { + var fresh simplyblockv1alpha2.StorageDeviceOps + if err := r.Get(ctx, client.ObjectKeyFromObject(ops), &fresh); err != nil { + return err + } + patch := client.MergeFromWithOptions(fresh.DeepCopy(), + client.MergeFromWithOptimisticLock{}) + change(&fresh.Status) + fresh.Status.ObservedGeneration = fresh.Generation + if err := r.Status().Patch(ctx, &fresh, patch); err != nil { + return err + } + fresh.Status.DeepCopyInto(&ops.Status) + return nil + }) + if err != nil { + return fmt.Errorf("record the operation's status: %w", err) + } + return nil +} + +// event records something about the operation, on the operation. +func (r *StorageDeviceOpsReconciler) event( + ops *simplyblockv1alpha2.StorageDeviceOps, eventType, reason, message string, +) { + if r.Recorder == nil { + return + } + r.Recorder.Eventf(ops, nil, eventType, reason, reason, "%s", message) +} + +// SetupWithManager registers the controller. +func (r *StorageDeviceOpsReconciler) SetupWithManager(mgr ctrl.Manager) error { + return ctrl.NewControllerManagedBy(mgr). + For(&simplyblockv1alpha2.StorageDeviceOps{}). + Named("storagedeviceops"). + Complete(r) +} diff --git a/operator/internal/controllers/node/storagedeviceops_graphs.go b/operator/internal/controllers/node/storagedeviceops_graphs.go new file mode 100644 index 000000000..8bc3af2b6 --- /dev/null +++ b/operator/internal/controllers/node/storagedeviceops_graphs.go @@ -0,0 +1,99 @@ +// The state graph of each StorageDeviceOps action, declared as data. +// +// Requesting ──► Awaiting +// +// One action is declared, because one is what the control plane's v2 API can +// serve; the other four of design-storagedevice.md §6 are blocked on verbs it +// does not offer, and v1alpha2.ExternalDependencies is the list. A MultiConfig +// with one entry is what every other Ops controller in this group uses, so the +// second action arrives as an entry rather than as a case. +// +// design-storagedevice.md §6 is the specification. + +package node + +import ( + "context" + "time" + + "github.com/simplyblock/atlas/statemachine" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// deviceStep is the operation's step type, aliased so the graph literal reads as +// the graph. +type deviceStep = simplyblockv1alpha2.StorageDeviceOpsStep + +const ( + stepDeviceRequesting = simplyblockv1alpha2.StorageDeviceOpsStepRequesting + stepDeviceAwaiting = simplyblockv1alpha2.StorageDeviceOpsStepAwaiting +) + +// actionDeviceRestart is the MultiConfig key for the one action this kind +// performs. +const actionDeviceRestart = statemachine.Action(simplyblockv1alpha2.StorageDeviceOpsActionRestart) + +// How long each step may take before the operation is reported as stuck. +const ( + // requestingDeviceDeadline bounds one POST to the control plane. A step + // still waiting after this is one whose control plane is not answering, + // which is a different problem from a device that will not come back. + requestingDeviceDeadline = 2 * time.Minute + + // awaitingDeviceDeadline bounds the device returning to service. A restart + // is a controller reset and a re-probe rather than a rebuild, so it is + // minutes; a device still absent after this is one that did not survive + // being recycled, which is the outcome worth reporting rather than waiting + // out. + awaitingDeviceDeadline = 15 * time.Minute +) + +// storageDeviceOpsGraphs declares the state graph of each action. +func storageDeviceOpsGraphs() statemachine.MultiConfig[deviceStep] { + return statemachine.MultiConfig[deviceStep]{ + actionDeviceRestart: { + Initial: stepDeviceRequesting, + States: map[deviceStep]statemachine.StateDef[deviceStep]{ + // Requesting is abortable because nothing has been issued yet: + // the step's side effect happens on the pass after the step is + // persisted, so an abort arriving in it stops a call that has + // not been made. + stepDeviceRequesting: { + To: []deviceStep{stepDeviceAwaiting}, + Abortable: true, + OnEnter: deviceDeadline(requestingDeviceDeadline), + }, + // Awaiting is not. The control plane has accepted the restart + // and there is no call that recalls one, so an abort here would + // record a stop that did not happen while the device restarted + // anyway. + stepDeviceAwaiting: {OnEnter: deviceDeadline(awaitingDeviceDeadline)}, + }, + }, + } +} + +// deviceDeadline is the entry hook every state here carries: it sets the step's +// budget and performs nothing. The work happens on the pass that follows, +// against the step the entry's write persisted, which is what makes a crash +// between the two resumable rather than invisible. +func deviceDeadline(d time.Duration) statemachine.TransitionFunc[deviceStep] { + return func(context.Context, deviceStep, deviceStep) (time.Duration, error) { return d, nil } +} + +// initialDeviceDeadline is the budget of the step every operation is born in. A +// machine is already in its initial state when it is built, so that state's +// OnEnter never runs, and setting it explicitly is what stops the first step +// from being the one step that cannot time out. +const initialDeviceDeadline = requestingDeviceDeadline + +// UnabortableDeviceSteps are the steps a running operation cannot be stopped in, +// which is the refusal table a DELETE admission guard would derive from this +// graph (design-crd-model.md §3.1). +// +// It is exported so that the guard and the graph cannot come to disagree: the +// webhook's table is checked against this rather than transcribed from it. +func UnabortableDeviceSteps() []deviceStep { + return statemachine.UnabortableMultiStates(storageDeviceOpsGraphs()) +} diff --git a/operator/internal/controllers/node/storagedeviceops_lock.go b/operator/internal/controllers/node/storagedeviceops_lock.go new file mode 100644 index 000000000..8fb680cba --- /dev/null +++ b/operator/internal/controllers/node/storagedeviceops_lock.go @@ -0,0 +1,129 @@ +// Mutual exclusion between two operations on one device. +// +// The lock is status.activeOpsRef on the StorageDevice, which +// design-storagedevice.md §4.2 declared against this kind arriving and left +// empty until it did. Two operations on one device would be two restarts of one +// controller, and the second would be issued against a device halfway through +// the first. +// +// It is self-describing, which is what makes a lock outside the holder's own +// object safe: the value is the holder's name, so any reconciler can read it, +// get the named operation, and learn whether the holder is running, finished, or +// gone. The risk a lock like this carries is a holder that dies between taking +// it and recording that it took it, and that is what the staleness check below +// answers. + +package node + +import ( + "context" + "fmt" + + apierrors "k8s.io/apimachinery/pkg/api/errors" + "sigs.k8s.io/controller-runtime/pkg/client" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// acquireLock takes the device's lock for this operation, and reports whether it +// now holds it. +// +// A lock another operation holds is waited on rather than failed, which is what +// lets operations queue: somebody restarting three devices of one node in +// sequence wants the second to wait, not to fail and need reissuing. +func (r *StorageDeviceOpsReconciler) acquireLock( + ctx context.Context, + ops *simplyblockv1alpha2.StorageDeviceOps, + device *simplyblockv1alpha2.StorageDevice, +) (bool, error) { + held := device.Status.ActiveOpsRef + if held == ops.Name { + return true, nil + } + + if held != "" { + takeable, err := r.lockIsStale(ctx, ops.Namespace, held) + if err != nil || !takeable { + return false, err + } + } + + // The optimistic lock is what makes this a lock at all: two reconcilers + // that both read the field free patch the same resourceVersion, and one of + // them gets a conflict and comes back to find the device taken. + patch := client.MergeFromWithOptions(device.DeepCopy(), + client.MergeFromWithOptimisticLock{}) + device.Status.ActiveOpsRef = ops.Name + if err := r.Status().Patch(ctx, device, patch); err != nil { + if apierrors.IsConflict(err) { + // Somebody else took it in the same instant. Waiting is the answer, + // and the next pass reads who has it. + return false, nil + } + return false, fmt.Errorf("take the lock on device %s: %w", device.Name, err) + } + return true, nil +} + +// lockIsStale reports whether the operation named as the holder has finished or +// no longer exists, which is what lets the lock be taken from it. +// +// A holder that is still running is not stale however long it has been running: +// a restart that is taking its time is exactly the case where breaking the lock +// would issue a second one. +func (r *StorageDeviceOpsReconciler) lockIsStale( + ctx context.Context, namespace, holder string, +) (bool, error) { + var running simplyblockv1alpha2.StorageDeviceOps + err := r.Get(ctx, client.ObjectKey{Namespace: namespace, Name: holder}, &running) + switch { + case apierrors.IsNotFound(err): + // The holder is gone, so nothing will ever release it. + return true, nil + case err != nil: + return false, fmt.Errorf("read the operation holding the lock: %w", err) + } + + switch running.Status.Phase { + case simplyblockv1alpha2.StorageDeviceOpsPhaseSucceeded, + simplyblockv1alpha2.StorageDeviceOpsPhaseFailed, + simplyblockv1alpha2.StorageDeviceOpsPhaseAborted: + return true, nil + } + return false, nil +} + +// releaseLock gives the device back, and does nothing where this operation is +// not the holder. +// +// The ownership check is what stops a late release from freeing a device the +// next operation has already taken: an operation that failed slowly could +// otherwise clear a lock somebody else is relying on. +func (r *StorageDeviceOpsReconciler) releaseLock( + ctx context.Context, + ops *simplyblockv1alpha2.StorageDeviceOps, + device *simplyblockv1alpha2.StorageDevice, +) error { + var fresh simplyblockv1alpha2.StorageDevice + err := r.Get(ctx, client.ObjectKeyFromObject(device), &fresh) + switch { + case apierrors.IsNotFound(err): + // The device object is gone, which §5.2 does when the node stops + // reporting the device. There is nothing to release. + return nil + case err != nil: + return fmt.Errorf("read device %s to release it: %w", device.Name, err) + } + + if fresh.Status.ActiveOpsRef != ops.Name { + return nil + } + + patch := client.MergeFromWithOptions(fresh.DeepCopy(), + client.MergeFromWithOptimisticLock{}) + fresh.Status.ActiveOpsRef = "" + if err := r.Status().Patch(ctx, &fresh, patch); err != nil { + return fmt.Errorf("release the lock on device %s: %w", device.Name, err) + } + return nil +} diff --git a/operator/internal/controllers/node/storagedeviceops_test.go b/operator/internal/controllers/node/storagedeviceops_test.go new file mode 100644 index 000000000..8be340863 --- /dev/null +++ b/operator/internal/controllers/node/storagedeviceops_test.go @@ -0,0 +1,425 @@ +// What a device operation has to get right, and the three places that have to +// agree about what its steps are. +// +// The properties worth holding are the ones a retry would otherwise paper over: +// a restart is issued once and not once per pass, a device is held by one +// operation at a time, and an action the control plane cannot serve is refused +// with the reason rather than accepted and failed. + +package node + +import ( + "context" + "errors" + "slices" + "strings" + "testing" + + "github.com/google/go-cmp/cmp" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/runtime" + "k8s.io/apimachinery/pkg/types" + clientgoscheme "k8s.io/client-go/kubernetes/scheme" + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + "github.com/simplyblock/atlas/controlplane" + "github.com/simplyblock/atlas/errs" + "github.com/simplyblock/atlas/statemachine" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +const ( + deviceOpsNamespace = "sb-system" + deviceOpsName = "restart-nvme0" + deviceObjectName = "worker-1-nvme0" + deviceBackendID = "b2222222-2222-4222-8222-222222222222" + deviceClusterID = "8ffac363-0c46-4714-a71b-f9c0b58a1269" + deviceNodeID = "a1111111-1111-4111-8111-111111111111" +) + +// everyDeviceStep is the Enum marker's list, transcribed rather than derived, so +// the assertion compares two independent statements of the same set. +var everyDeviceStep = []string{"Awaiting", "Requesting"} + +// deviceStepCELRule is the rule as StorageDeviceOpsStatus declares it. +const deviceStepCELRule = "!has(self.state) || self.state in ['Requesting','Awaiting']" + +func TestTheDeviceStepEnumCoversEveryDeclaredState(t *testing.T) { + declared := statemachine.DeclaredMultiStates(storageDeviceOpsGraphs()) + want := slices.Clone(everyDeviceStep) + slices.Sort(want) + if diff := cmp.Diff(want, declared); diff != "" { + t.Errorf("the graph and the Enum marker disagree (-marker +graph):\n%s", diff) + } +} + +func TestTheDeviceCELRuleCoversEveryDeclaredState(t *testing.T) { + declared := statemachine.DeclaredMultiStates(storageDeviceOpsGraphs()) + for _, state := range declared { + if !strings.Contains(deviceStepCELRule, "'"+state+"'") { + t.Errorf("status.step's CEL rule does not accept the declared step %q", state) + } + } +} + +// The enum admits exactly the actions a graph declares. An action the API +// accepts and no graph declares is an object whose first reconcile can only +// fail, which is the thing the narrowed enum exists to prevent. +func TestTheActionEnumAdmitsOnlyWhatAGraphDeclares(t *testing.T) { + declared := storageDeviceOpsGraphs() + if _, ok := declared[actionDeviceRestart]; !ok { + t.Error("Restart is in the Enum marker and no graph declares it") + } + if len(declared) != 1 { + t.Errorf("%d graphs are declared and the Enum marker admits one action; an action "+ + "with a graph and no enum member is one nobody can ask for", len(declared)) + } +} + +// Every blocked action has a dependency row naming the endpoint it waits on, and +// no row names an action that is already available. The list is the ask of +// another team, so an action that shipped and left its row behind would be an +// ask for something that exists. +func TestEveryBlockedActionNamesTheEndpointItWaitsOn(t *testing.T) { + blocked := map[simplyblockv1alpha2.StorageDeviceOpsAction]bool{ + simplyblockv1alpha2.StorageDeviceOpsActionSelfTest: true, + simplyblockv1alpha2.StorageDeviceOpsActionFail: true, + simplyblockv1alpha2.StorageDeviceOpsActionReplace: true, + simplyblockv1alpha2.StorageDeviceOpsActionMigrate: true, + } + + for _, dependency := range simplyblockv1alpha2.ExternalDependencies() { + if !blocked[dependency.Action] { + t.Errorf("%s has a dependency row and is not blocked", dependency.Action) + } + delete(blocked, dependency.Action) + if dependency.Endpoint == "" || dependency.Because == "" { + t.Errorf("%s's row names no endpoint or no reason, so it is not an ask "+ + "anybody can act on", dependency.Action) + } + if _, declared := storageDeviceOpsGraphs()[statemachine.Action(dependency.Action)]; declared { + t.Errorf("%s is declared as blocked and has a graph", dependency.Action) + } + } + for action := range blocked { + t.Errorf("%s is specified by §6 and has neither a graph nor a dependency row, so "+ + "nothing records why it is missing", action) + } +} + +// Awaiting cannot be aborted, because the control plane has accepted the restart +// and nothing recalls one. Requesting can, because the step is persisted before +// the call. +func TestOnlyTheStepBeforeTheCallIsAbortable(t *testing.T) { + unabortable := UnabortableDeviceSteps() + if !slices.Contains(unabortable, stepDeviceAwaiting) { + t.Error("Awaiting is abortable, so an abort there records a stop that did not " + + "happen while the device restarts anyway") + } + if slices.Contains(unabortable, stepDeviceRequesting) { + t.Error("Requesting is not abortable, so an operation cannot be called off before " + + "it has done anything") + } +} + +// restartCounter is a control plane that counts restarts and reports whichever +// status it is told to. +type restartCounter struct { + restarts int + status string + missing bool + refuse error +} + +func (c *restartCounter) Device( + context.Context, string, string, string, +) (controlplane.Device, error) { + if c.missing { + return controlplane.Device{}, errs.ErrNotFound + } + return controlplane.Device{ID: deviceBackendID, Status: c.status}, nil +} + +func (c *restartCounter) RestartDevice(context.Context, string, string, string) error { + if c.refuse != nil { + return c.refuse + } + c.restarts++ + return nil +} + +func deviceObject() *simplyblockv1alpha2.StorageDevice { + return &simplyblockv1alpha2.StorageDevice{ + ObjectMeta: metav1.ObjectMeta{Name: deviceObjectName, Namespace: deviceOpsNamespace}, + Spec: simplyblockv1alpha2.StorageDeviceSpec{ + NodeRef: "worker-1", DeviceID: deviceBackendID, + }, + Status: simplyblockv1alpha2.StorageDeviceStatus{ + ClusterID: deviceClusterID, NodeID: deviceNodeID, DeviceStatus: "online", + }, + } +} + +func deviceOperation() *simplyblockv1alpha2.StorageDeviceOps { + return &simplyblockv1alpha2.StorageDeviceOps{ + ObjectMeta: metav1.ObjectMeta{Name: deviceOpsName, Namespace: deviceOpsNamespace}, + Spec: simplyblockv1alpha2.StorageDeviceOpsSpec{ + DeviceRef: deviceObjectName, + Action: simplyblockv1alpha2.StorageDeviceOpsActionRestart, + }, + } +} + +type deviceWorld struct { + t *testing.T + c client.Client + r *StorageDeviceOpsReconciler + api *restartCounter +} + +func newDeviceWorld(t *testing.T, api *restartCounter, objects ...client.Object) *deviceWorld { + t.Helper() + + scheme := runtime.NewScheme() + if err := clientgoscheme.AddToScheme(scheme); err != nil { + t.Fatalf("building the scheme: %v", err) + } + if err := simplyblockv1alpha2.AddToScheme(scheme); err != nil { + t.Fatalf("registering v1alpha2: %v", err) + } + c := fake.NewClientBuilder(). + WithScheme(scheme). + WithObjects(objects...). + WithStatusSubresource( + &simplyblockv1alpha2.StorageDeviceOps{}, &simplyblockv1alpha2.StorageDevice{}). + Build() + + return &deviceWorld{t: t, c: c, api: api, r: &StorageDeviceOpsReconciler{ + Client: c, Scheme: scheme, API: api, + }} +} + +// pass reconciles once and returns the operation and the device as they stand. +func (w *deviceWorld) pass() (*simplyblockv1alpha2.StorageDeviceOps, *simplyblockv1alpha2.StorageDevice) { + w.t.Helper() + + if _, err := w.r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: types.NamespacedName{Namespace: deviceOpsNamespace, Name: deviceOpsName}, + }); err != nil { + w.t.Fatalf("reconcile: %v", err) + } + return w.read() +} + +// settle reconciles until the operation is terminal or the passes run out, +// which keeps a test about the outcome rather than about how many passes the +// finalizer and the birth cost. +func (w *deviceWorld) settle() (*simplyblockv1alpha2.StorageDeviceOps, *simplyblockv1alpha2.StorageDevice) { + w.t.Helper() + + for range 8 { + ops, device := w.pass() + switch ops.Status.Phase { + case simplyblockv1alpha2.StorageDeviceOpsPhaseSucceeded, + simplyblockv1alpha2.StorageDeviceOpsPhaseFailed, + simplyblockv1alpha2.StorageDeviceOpsPhaseAborted: + return ops, device + } + } + return w.read() +} + +func (w *deviceWorld) read() (*simplyblockv1alpha2.StorageDeviceOps, *simplyblockv1alpha2.StorageDevice) { + w.t.Helper() + + var ops simplyblockv1alpha2.StorageDeviceOps + if err := w.c.Get(context.Background(), + types.NamespacedName{Namespace: deviceOpsNamespace, Name: deviceOpsName}, &ops); err != nil { + w.t.Fatalf("read the operation: %v", err) + } + var device simplyblockv1alpha2.StorageDevice + if err := w.c.Get(context.Background(), + types.NamespacedName{Namespace: deviceOpsNamespace, Name: deviceObjectName}, &device); err != nil { + w.t.Fatalf("read the device: %v", err) + } + return &ops, &device +} + +// The restart is issued once. The step is persisted before the call, so +// re-entering the step means the previous pass died between the two — and a +// restart issued twice is a device recycled twice. +func TestTheRestartIsIssuedOnce(t *testing.T) { + api := &restartCounter{status: "online"} + w := newDeviceWorld(t, api, deviceOperation(), deviceObject()) + + w.settle() + + if api.restarts != 1 { + t.Errorf("the device was restarted %d times, want once", api.restarts) + } + ops, device := w.read() + if ops.Status.Phase != simplyblockv1alpha2.StorageDeviceOpsPhaseSucceeded { + t.Errorf("phase = %q, want Succeeded: %s", ops.Status.Phase, ops.Status.Message) + } + if device.Status.ActiveOpsRef != "" { + t.Errorf("the device is still locked by %q after the operation finished", + device.Status.ActiveOpsRef) + } +} + +// The operation holds the device while it runs, which is the field §4.2 declared +// empty until this kind arrived. +func TestTheOperationHoldsItsDeviceWhileItRuns(t *testing.T) { + w := newDeviceWorld(t, &restartCounter{status: "restarting"}, + deviceOperation(), deviceObject()) + + w.pass() // the finalizer + _, device := w.pass() + if device.Status.ActiveOpsRef != deviceOpsName { + t.Errorf("activeOpsRef = %q, want the running operation", device.Status.ActiveOpsRef) + } +} + +// A second operation waits rather than failing, which is what lets somebody +// restart three devices of a node in sequence without reissuing anything. +func TestASecondOperationWaitsForTheFirst(t *testing.T) { + device := deviceObject() + device.Status.ActiveOpsRef = "someone-else" + holder := &simplyblockv1alpha2.StorageDeviceOps{ + ObjectMeta: metav1.ObjectMeta{Name: "someone-else", Namespace: deviceOpsNamespace}, + Spec: simplyblockv1alpha2.StorageDeviceOpsSpec{ + DeviceRef: deviceObjectName, + Action: simplyblockv1alpha2.StorageDeviceOpsActionRestart, + }, + Status: simplyblockv1alpha2.StorageDeviceOpsStatus{ + Phase: simplyblockv1alpha2.StorageDeviceOpsPhaseRunning, + }, + } + + api := &restartCounter{status: "online"} + w := newDeviceWorld(t, api, deviceOperation(), holder, device) + + w.pass() // the finalizer + ops, held := w.pass() + if ops.Status.Phase == simplyblockv1alpha2.StorageDeviceOpsPhaseFailed { + t.Errorf("the second operation failed rather than waiting: %s", ops.Status.Message) + } + if held.Status.ActiveOpsRef != "someone-else" { + t.Errorf("the lock moved to %q while its holder was still running", + held.Status.ActiveOpsRef) + } + if api.restarts != 0 { + t.Error("a restart was issued against a device another operation is holding") + } +} + +// A device the control plane stops holding did not come back, which is a finding +// rather than a wait: the operation would otherwise sit until its deadline +// describing a device that is gone. +func TestADeviceThatDoesNotComeBackFailsTheOperation(t *testing.T) { + api := &restartCounter{status: "online"} + w := newDeviceWorld(t, api, deviceOperation(), deviceObject()) + + w.pass() // the finalizer + w.pass() // begin + w.pass() // request + api.missing = true + + ops, _ := w.settle() + if ops.Status.Phase != simplyblockv1alpha2.StorageDeviceOpsPhaseFailed { + t.Fatalf("phase = %q, want Failed when the device does not come back", ops.Status.Phase) + } + if !strings.Contains(ops.Status.Message, "no longer held") { + t.Errorf("the failure does not say the device is gone: %q", ops.Status.Message) + } +} + +// A control plane that refuses the restart fails the operation with what it +// said, rather than retrying against the backoff forever. +func TestARefusedRestartFailsTheOperation(t *testing.T) { + api := &restartCounter{status: "online", refuse: errors.New("device is busy")} + w := newDeviceWorld(t, api, deviceOperation(), deviceObject()) + + ops, device := w.settle() + + if ops.Status.Phase != simplyblockv1alpha2.StorageDeviceOpsPhaseFailed { + t.Fatalf("phase = %q, want Failed on a refused restart", ops.Status.Phase) + } + if !strings.Contains(ops.Status.Message, "device is busy") { + t.Errorf("the failure does not carry what the control plane said: %q", ops.Status.Message) + } + if device.Status.ActiveOpsRef != "" { + t.Errorf("a failed operation is still holding the device as %q", + device.Status.ActiveOpsRef) + } +} + +// An operation naming a device that does not exist is failed rather than +// retried, since spec.deviceRef is immutable and the reference can never become +// resolvable. +func TestAnOperationNamingNoDeviceFails(t *testing.T) { + ops := deviceOperation() + ops.Spec.DeviceRef = "not-a-device" + w := newDeviceWorld(t, &restartCounter{status: "online"}, ops) + + for range 3 { + if _, err := w.r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: types.NamespacedName{Namespace: deviceOpsNamespace, Name: deviceOpsName}, + }); err != nil { + t.Fatalf("reconcile: %v", err) + } + } + + var got simplyblockv1alpha2.StorageDeviceOps + if err := w.c.Get(context.Background(), + types.NamespacedName{Namespace: deviceOpsNamespace, Name: deviceOpsName}, &got); err != nil { + t.Fatalf("read the operation: %v", err) + } + if got.Status.Phase != simplyblockv1alpha2.StorageDeviceOpsPhaseFailed { + t.Errorf("phase = %q, want Failed for a device that does not exist", got.Status.Phase) + } +} + +// An action no graph declares is failed with the endpoint it waits on, so the +// person who asked learns which capability is missing rather than that a value +// was rejected. +func TestAnUnavailableActionNamesWhatItWaitsOn(t *testing.T) { + for _, action := range []simplyblockv1alpha2.StorageDeviceOpsAction{ + simplyblockv1alpha2.StorageDeviceOpsActionSelfTest, + simplyblockv1alpha2.StorageDeviceOpsActionFail, + simplyblockv1alpha2.StorageDeviceOpsActionReplace, + simplyblockv1alpha2.StorageDeviceOpsActionMigrate, + } { + t.Run(string(action), func(t *testing.T) { + got := unavailableAction(action) + if !strings.Contains(got, "/api/v2/") { + t.Errorf("the refusal names no endpoint: %q", got) + } + if !strings.Contains(got, string(action)) { + t.Errorf("the refusal does not name the action: %q", got) + } + }) + } +} + +// A device whose object does not say which backend device it is has nothing to +// address a restart to, and saying so beats posting to a path built from empty +// strings. +func TestADeviceWithNoBackendIdentityIsRefused(t *testing.T) { + device := deviceObject() + device.Status.ClusterID = "" + api := &restartCounter{status: "online"} + w := newDeviceWorld(t, api, deviceOperation(), device) + + ops, _ := w.settle() + + if ops.Status.Phase != simplyblockv1alpha2.StorageDeviceOpsPhaseFailed { + t.Fatalf("phase = %q, want Failed", ops.Status.Phase) + } + if api.restarts != 0 { + t.Error("a restart was addressed to a device with no backend identity") + } +} diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagedeviceops.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagedeviceops.yaml new file mode 100644 index 000000000..eec5d0629 --- /dev/null +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagedeviceops.yaml @@ -0,0 +1,162 @@ +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition +metadata: + annotations: + controller-gen.kubebuilder.io/version: v0.21.0 + name: storagedeviceops.storage.simplyblock.io +spec: + group: storage.simplyblock.io + names: + kind: StorageDeviceOps + listKind: StorageDeviceOpsList + plural: storagedeviceops + shortNames: + - sdops + singular: storagedeviceops + scope: Namespaced + versions: + - additionalPrinterColumns: + - jsonPath: .spec.deviceRef + name: Device + type: string + - jsonPath: .spec.action + name: Action + type: string + - jsonPath: .status.phase + name: Phase + type: string + - jsonPath: .status.step.state + name: Step + type: string + - jsonPath: .status.message + name: Message + priority: 1 + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha2 + schema: + openAPIV3Schema: + description: StorageDeviceOps is a single operation performed against one + StorageDevice. + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: StorageDeviceOpsSpec is one operation to perform against + one StorageDevice. + properties: + abort: + description: |- + Abort asks a running operation to stop at its next step and unwind. + + Restart can be aborted before its call is issued and not after: a restart + the control plane has accepted is one nothing can recall, so the graph + declares where the edge exists rather than this field promising one. + type: boolean + action: + description: Action is the operation to perform. + enum: + - Restart + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + deviceRef: + description: |- + DeviceRef names the StorageDevice this operation acts on, in this + operation's own namespace. The operation never owns its target, because + deleting the record of an operation must not delete the device record it + operated on. + maxLength: 253 + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + required: + - action + - deviceRef + type: object + status: + description: StorageDeviceOpsStatus is the observed state of one device + operation. + properties: + completedAt: + description: CompletedAt is when it reached a terminal phase. + format: date-time + type: string + deviceStatusBefore: + description: |- + DeviceStatusBefore is what the control plane reported the device's status + to be when the operation took its lock, so a wait can tell the device + coming back from its never having gone. + type: string + message: + description: |- + Message is the reason the phase is what it is: one sentence, replaced as + the operation moves, and never a log. + type: string + observedGeneration: + description: |- + ObservedGeneration is the generation the rest of this status was computed + from, so a stale status can be told from a current one. + format: int64 + type: integer + phase: + description: Phase is the operation's own progress. + enum: + - Pending + - Running + - Succeeded + - Failed + - Aborted + type: string + startedAt: + description: StartedAt is when the operation acquired its target's + lock. + format: date-time + type: string + step: + description: |- + Step is the position of the running action's state machine. It is + persisted before the side effect that step performs. + properties: + deadline: + description: |- + Deadline is when that state expires, absent when it has none. It is an + absolute instant, so a state whose deadline passed while the controller + was down restores as already expired. + format: date-time + type: string + state: + description: |- + State is the state the machine was in. Empty means the resource has not + been reconciled yet, and restores to the graph's initial state. + type: string + type: object + x-kubernetes-validations: + - message: unknown step + rule: '!has(self.state) || self.state in [''Requesting'',''Awaiting'']' + type: object + type: object + served: true + storage: true + subresources: + status: {} From bfed606c2cd4077e1b3da51c64b06b4a926cdf59 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 15:06:41 +0200 Subject: [PATCH 051/206] refactor(device): the blocked actions are a TODO rather than a type MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The four actions of design-storagedevice §6 that the v2 API cannot serve were recorded as an ExternalDependency struct and an ExternalDependencies function, which is more machinery than the thing needs. This repository already marks work that waits on something else with a TODO beside the code that waits, and six of them say so in the driver package alone. So the ask is a TODO(storagedeviceops) block above the four constants, naming the endpoint each action needs and what it does with it. The constants stay out of the Enum marker, which is what actually refuses the action; the block is what tells a reader, and the control-plane team, why. What goes with the type is its two consumers. The runtime message no longer looks an endpoint up to report it, since that path is only reachable through an older CRD serving an action the current enum refuses, and the useful sentence there is that the action is not available rather than which path it would have called. And the test that held the list against the graphs goes, because the list it was keeping honest no longer exists — the graphs and the Enum marker are still checked against each other, which is the pair that can actually disagree. Co-Authored-By: Claude Fable 5 --- atlas-lib/controlplane/devices.go | 6 +- .../api/v1alpha2/storagedeviceops_types.go | 101 +++++------------- .../api/v1alpha2/zz_generated.deepcopy.go | 15 --- .../crd-redesign/design-storagedevice.md | 7 +- .../node/storagedeviceops_controller.go | 28 ++--- .../controllers/node/storagedeviceops_test.go | 53 --------- 6 files changed, 42 insertions(+), 168 deletions(-) diff --git a/atlas-lib/controlplane/devices.go b/atlas-lib/controlplane/devices.go index b8a99cd4a..c52729bf3 100644 --- a/atlas-lib/controlplane/devices.go +++ b/atlas-lib/controlplane/devices.go @@ -4,9 +4,9 @@ // The v2 API offers three verbs on a device — restart, remove, and reset — where // design-storagedevice.md §7 asks for seven. What is here is what exists: the // reads the operations wait on, and the restart they are built around. The four -// that are missing are recorded as an external dependency in -// operator/api/v1alpha2/storagedeviceops_types.go, beside the actions each one -// blocks, rather than as a client method that would return a 404. +// that are missing are recorded as a TODO beside the actions they block, in the +// operator's storagedeviceops_types.go, rather than as a client method here that +// would return a 404. package controlplane diff --git a/operator/api/v1alpha2/storagedeviceops_types.go b/operator/api/v1alpha2/storagedeviceops_types.go index 34a94c637..b59d2975a 100644 --- a/operator/api/v1alpha2/storagedeviceops_types.go +++ b/operator/api/v1alpha2/storagedeviceops_types.go @@ -9,8 +9,7 @@ // // design-storagedevice.md §6 specifies five actions. One is served by the // control plane's v2 API and is built; the other four are blocked on verbs that -// API does not offer, and [ExternalDependencies] is the list, so the ask is a -// value in this repository rather than a sentence in a document. +// API does not offer, and the TODO beside their constants is the ask. // // **The enum admits only what the operator can perform.** Declaring the other // four now would accept an object whose first reconcile can only fail, and an @@ -40,84 +39,38 @@ const ( // recycling one device rather than its node. StorageDeviceOpsActionRestart StorageDeviceOpsAction = "Restart" - // The four actions §6 specifies and the API cannot serve. They are declared - // so that the names exist where the reason does, and they are absent from - // the Enum marker above: an object naming one is refused at admission - // rather than accepted and failed. + // TODO(storagedeviceops): EXTERNAL DEPENDENCY — the four actions below wait + // on control-plane verbs the v2 API does not offer. It serves restart, + // remove, and reset where design-storagedevice.md §7 asks for seven, and + // remove buys no action on its own: it is a step of Replace and of Migrate, + // and both also need the adopt call that names the device arriving. + // + // SelfTest POST /api/v2/clusters/{c}/storage-nodes/{n}/devices/{d}/self-test + // runs the device's own self-test and reports the verdict, with + // the short or extended mode in the body. + // Fail POST .../devices/{d}/fail + // takes a device out of the data path and leaves it in the slot, + // so the cluster rebuilds its redundancy elsewhere and stops + // reading from a device somebody has judged untrustworthy. + // Replace POST .../devices/adopt + // names the device that arrived; the removal verb exists and the + // pairing is what makes the arrival identifiable. + // Migrate POST .../devices/{d}/detach, and the adopt above accepting a + // device WITH ITS CONTENTS. An adopt that can only take an empty + // device turns Migrate into two Replaces and a full rebuild, + // which is what §6.2 gives as the action's reason to exist. This + // is the row whose absence removes an action rather than + // degrading it. + // + // The constants are declared and absent from the Enum marker above, so the + // names exist where the reason does while an object naming one is refused at + // admission rather than accepted and failed. StorageDeviceOpsActionSelfTest StorageDeviceOpsAction = "SelfTest" StorageDeviceOpsActionFail StorageDeviceOpsAction = "Fail" StorageDeviceOpsActionReplace StorageDeviceOpsAction = "Replace" StorageDeviceOpsActionMigrate StorageDeviceOpsAction = "Migrate" ) -// ExternalDependency is one action this kind specifies, the control-plane -// capability it waits on, and what its absence costs. -// -// It is data rather than prose because the list is an ask of another team and a -// checklist for this one: each row names the endpoint to build, and the action -// ships when the row does. A design paragraph saying the same thing cannot be -// enumerated, cannot be tested against the enum, and goes stale the day one -// endpoint arrives. -type ExternalDependency struct { - // Action is what this unblocks. - Action StorageDeviceOpsAction - - // Endpoint is the v2 API call the action issues, in the shape - // design-storagedevice.md §7 asks for. - Endpoint string - - // Because says what the action does with it, so the ask carries its own - // justification rather than a section number. - Because string -} - -// ExternalDependencies are the four actions of §6 that the v2 API cannot serve -// today, and the verb each one needs. -// -// The API offers three device verbs — restart, remove, and reset — where §7 asks -// for seven. Restart is built on the first. Remove exists and is not enough on -// its own: it is a step of Replace and of Migrate, and both of those also need -// an adopt call to name the device that arrives, so the verb that exists buys -// neither action. -// -// Migrate's row is the one whose absence removes an action rather than degrading -// it. Attaching asks the control plane to accept a device as another node's with -// its contents intact, so the chunks on it are re-homed rather than rebuilt. A -// control plane that can only adopt a device as empty turns Migrate into two -// Replaces and a full rebuild, which is the thing §6.2 gives as the reason the -// action exists. -func ExternalDependencies() []ExternalDependency { - const base = "POST /api/v2/clusters/{cluster}/storage-nodes/{node}/devices/" - return []ExternalDependency{ - { - Action: StorageDeviceOpsActionSelfTest, - Endpoint: base + "{device}/self-test", - Because: "the action runs the device's own self-test and reports the verdict, " + - "with the short or extended mode in the body", - }, - { - Action: StorageDeviceOpsActionFail, - Endpoint: base + "{device}/fail", - Because: "the action takes a device out of the data path and leaves it in the " + - "slot, so the cluster rebuilds its redundancy elsewhere and stops reading " + - "from a device somebody has judged untrustworthy", - }, - { - Action: StorageDeviceOpsActionReplace, - Endpoint: base + "adopt", - Because: "the action pairs a removal with an arrival, and the removal verb " + - "exists while the call naming the device that arrived does not", - }, - { - Action: StorageDeviceOpsActionMigrate, - Endpoint: base + "{device}/detach, and " + base + "adopt accepting a device with its contents", - Because: "the action moves the drive to another node with what is on it; an " + - "adopt that can only take an empty device makes this two replacements and " + - "a full rebuild, which is what the action exists to avoid", - }, - } -} - // StorageDeviceOpsPhase is the operation's own progress. // +kubebuilder:validation:Enum=Pending;Running;Succeeded;Failed;Aborted type StorageDeviceOpsPhase string diff --git a/operator/api/v1alpha2/zz_generated.deepcopy.go b/operator/api/v1alpha2/zz_generated.deepcopy.go index 65d0b700f..d460d2ee0 100644 --- a/operator/api/v1alpha2/zz_generated.deepcopy.go +++ b/operator/api/v1alpha2/zz_generated.deepcopy.go @@ -806,21 +806,6 @@ func (in *DriverTLS) DeepCopy() *DriverTLS { return out } -// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. -func (in *ExternalDependency) DeepCopyInto(out *ExternalDependency) { - *out = *in -} - -// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new ExternalDependency. -func (in *ExternalDependency) DeepCopy() *ExternalDependency { - if in == nil { - return nil - } - out := new(ExternalDependency) - in.DeepCopyInto(out) - return out -} - // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *FoundationDBSpec) DeepCopyInto(out *FoundationDBSpec) { *out = *in diff --git a/operator/docs/designs/crd-redesign/design-storagedevice.md b/operator/docs/designs/crd-redesign/design-storagedevice.md index eb6d34934..6072a79c8 100644 --- a/operator/docs/designs/crd-redesign/design-storagedevice.md +++ b/operator/docs/designs/crd-redesign/design-storagedevice.md @@ -555,10 +555,9 @@ Appendix B. **One of the five actions is implemented, and it is the one the kind exists for.** `Restart` is built: the kind, its graph, the reconciler, and the device lock §4.2 declared empty against this section arriving. The other four are -blocked on control-plane verbs the v2 API does not offer, and the ask is -`v1alpha2.ExternalDependencies` — a value in the repository rather than a -paragraph here, so each row names the endpoint to build and the action ships when -the row does. +blocked on control-plane verbs the v2 API does not offer, and the ask is the +`TODO(storagedeviceops)` beside their constants in `storagedeviceops_types.go`, +which names the endpoint each one needs. **The enum admits only what the operator can perform.** Declaring all five now would accept an object whose first reconcile can only fail, and an API that takes diff --git a/operator/internal/controllers/node/storagedeviceops_controller.go b/operator/internal/controllers/node/storagedeviceops_controller.go index dfda561bc..496b7ac53 100644 --- a/operator/internal/controllers/node/storagedeviceops_controller.go +++ b/operator/internal/controllers/node/storagedeviceops_controller.go @@ -11,8 +11,8 @@ // on is in the status, so a controller restart resumes rather than restarts. // // design-storagedevice.md §6 is the specification, and §6's other four actions -// are blocked on control-plane verbs that do not exist -// (v1alpha2.ExternalDependencies). +// are blocked on control-plane verbs that do not exist — the TODO beside their +// constants in storagedeviceops_types.go names each one. package node @@ -143,10 +143,13 @@ func (r *StorageDeviceOpsReconciler) advance( graph, declared := storageDeviceOpsGraphs()[statemachine.Action(ops.Spec.Action)] if !declared { // Admission's enum admits only the actions this operator performs, so - // reaching here means an older CRD served the object. The message names - // the dependency rather than the enum, because what a user needs to - // know is that the action is not available yet. - return ctrl.Result{}, r.fail(ctx, ops, unavailableAction(ops.Spec.Action)) + // reaching here means an older CRD served the object. What a user needs + // to know is that the action is not available yet rather than that a + // value was rejected; which capability it waits on is the TODO beside + // the action's constant. + return ctrl.Result{}, r.fail(ctx, ops, fmt.Sprintf( + "action %q is not one this operator performs; it waits on a control-plane "+ + "capability the v2 API does not offer", ops.Spec.Action)) } machine, err := statemachine.NewFromSnapshot(ctx, graph, @@ -323,19 +326,6 @@ func (r *StorageDeviceOpsReconciler) target( return &device, nil } -// unavailableAction is what an operation naming an action this operator cannot -// perform is failed with. It names the endpoint rather than the enum, because -// the useful fact is which capability is missing. -func unavailableAction(action simplyblockv1alpha2.StorageDeviceOpsAction) string { - for _, dependency := range simplyblockv1alpha2.ExternalDependencies() { - if dependency.Action == action { - return fmt.Sprintf("action %s is not available: it needs %s, which the control "+ - "plane does not offer", action, dependency.Endpoint) - } - } - return fmt.Sprintf("action %q is not one this operator performs", action) -} - // deviceTimeoutMessage says what a step outliving its deadline means, which // differs by step: one is a control plane that did not answer, the other is a // device that did not come back. diff --git a/operator/internal/controllers/node/storagedeviceops_test.go b/operator/internal/controllers/node/storagedeviceops_test.go index 8be340863..cdb6c31da 100644 --- a/operator/internal/controllers/node/storagedeviceops_test.go +++ b/operator/internal/controllers/node/storagedeviceops_test.go @@ -79,37 +79,6 @@ func TestTheActionEnumAdmitsOnlyWhatAGraphDeclares(t *testing.T) { } } -// Every blocked action has a dependency row naming the endpoint it waits on, and -// no row names an action that is already available. The list is the ask of -// another team, so an action that shipped and left its row behind would be an -// ask for something that exists. -func TestEveryBlockedActionNamesTheEndpointItWaitsOn(t *testing.T) { - blocked := map[simplyblockv1alpha2.StorageDeviceOpsAction]bool{ - simplyblockv1alpha2.StorageDeviceOpsActionSelfTest: true, - simplyblockv1alpha2.StorageDeviceOpsActionFail: true, - simplyblockv1alpha2.StorageDeviceOpsActionReplace: true, - simplyblockv1alpha2.StorageDeviceOpsActionMigrate: true, - } - - for _, dependency := range simplyblockv1alpha2.ExternalDependencies() { - if !blocked[dependency.Action] { - t.Errorf("%s has a dependency row and is not blocked", dependency.Action) - } - delete(blocked, dependency.Action) - if dependency.Endpoint == "" || dependency.Because == "" { - t.Errorf("%s's row names no endpoint or no reason, so it is not an ask "+ - "anybody can act on", dependency.Action) - } - if _, declared := storageDeviceOpsGraphs()[statemachine.Action(dependency.Action)]; declared { - t.Errorf("%s is declared as blocked and has a graph", dependency.Action) - } - } - for action := range blocked { - t.Errorf("%s is specified by §6 and has neither a graph nor a dependency row, so "+ - "nothing records why it is missing", action) - } -} - // Awaiting cannot be aborted, because the control plane has accepted the restart // and nothing recalls one. Requesting can, because the step is persisted before // the call. @@ -383,28 +352,6 @@ func TestAnOperationNamingNoDeviceFails(t *testing.T) { } } -// An action no graph declares is failed with the endpoint it waits on, so the -// person who asked learns which capability is missing rather than that a value -// was rejected. -func TestAnUnavailableActionNamesWhatItWaitsOn(t *testing.T) { - for _, action := range []simplyblockv1alpha2.StorageDeviceOpsAction{ - simplyblockv1alpha2.StorageDeviceOpsActionSelfTest, - simplyblockv1alpha2.StorageDeviceOpsActionFail, - simplyblockv1alpha2.StorageDeviceOpsActionReplace, - simplyblockv1alpha2.StorageDeviceOpsActionMigrate, - } { - t.Run(string(action), func(t *testing.T) { - got := unavailableAction(action) - if !strings.Contains(got, "/api/v2/") { - t.Errorf("the refusal names no endpoint: %q", got) - } - if !strings.Contains(got, string(action)) { - t.Errorf("the refusal does not name the action: %q", got) - } - }) - } -} - // A device whose object does not say which backend device it is has nothing to // address a restart to, and saying so beats posting to a path built from empty // strings. From 6ff56676e575789f9b9b1a88017105c95f1fa3a7 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 15:26:26 +0200 Subject: [PATCH 052/206] feat(node): a node whose pod cannot be placed says so MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit PodSchedulingFailed has been declared in events.go since the reasons table was wired and emitted nowhere. design-storagenode.md §13.1 specifies it, and every other reason in that table has a call site. It is worth having rather than deleting, because it is the only reason in the table whose cause lies outside both the operator and the control plane. A storage-node pod the scheduler refuses leaves the worker healthy, the backend never hearing of the node, and the provisioning machine running CheckingHost out to its deadline with nothing to report but that the host is unreachable. That is true and it is the symptom. The scheduler has already written the cause down, specifically: which nodes it considered and what each of them was short of. So the reconcile copies that sentence onto the StorageNode, where somebody looking at a stalled worker is already looking. It is announced before the branch into provisioning or steady state, because a refused pod holds a provisioning node and takes a running one offline, and the answer is the same on both sides. The pod is found by node affinity rather than by spec.nodeName. An unplaced pod has no node name, and being unplaced is the case the event exists for; the DaemonSet controller stamps the worker into the pod's affinity as a metadata.name match field and leaves the binding to the scheduler, so an unplaced pod names its worker there and nowhere else. Only Unschedulable counts. A gated pod carries the same false condition and is one the scheduler has not been asked about yet, so announcing it would warn about every pod on its way to running rather than about one that is stuck. It reports and does not decide. The node is already held by whichever step is waiting on the worker, and failing it here would turn a condition somebody resolves by freeing capacity into one that needs the object recreated. An unreadable pod list says nothing, rather than announcing a scheduling failure on the strength of not having looked. Four tests, paired against the three ways this goes wrong quietly: it announces the scheduler's own reason; it is silent for a placed pod; it is silent for another worker's refused pod, which matters because the DaemonSet covers the cluster and one node's reconcile sees every node's pod; and it is silent for a gated one. Co-Authored-By: Claude Opus 5 (1M context) --- .../controllers/node/podscheduling.go | 124 +++++++++++ .../controllers/node/podscheduling_test.go | 199 ++++++++++++++++++ .../node/storagenode_controller.go | 6 + 3 files changed, 329 insertions(+) create mode 100644 operator/internal/controllers/node/podscheduling.go create mode 100644 operator/internal/controllers/node/podscheduling_test.go diff --git a/operator/internal/controllers/node/podscheduling.go b/operator/internal/controllers/node/podscheduling.go new file mode 100644 index 000000000..38282f96e --- /dev/null +++ b/operator/internal/controllers/node/podscheduling.go @@ -0,0 +1,124 @@ +// The Kubernetes-side failure a storage node's own status cannot show. +// +// Every other reason a node is held is one the operator or the control plane +// knows: the cluster has no UUID, the worker's API does not answer, no slot is +// free. A pod the scheduler has refused is none of those. The backend never +// hears of the node, the worker looks healthy, and the operator sits in +// CheckingHost until its deadline with nothing to say beyond that the host is +// unreachable. That is true, and it names the symptom rather than the cause. +// +// The scheduler has already written the cause down, in the message of the pod's +// PodScheduled condition, and it is specific: which nodes it considered and what +// each of them was short of. This file copies that sentence onto the StorageNode, +// where the administrator looking at the stalled worker is already looking. +// +// design-storagenode.md §13.1 is the specification for the event. + +package node + +import ( + "context" + + corev1 "k8s.io/api/core/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + + "github.com/simplyblock/atlas/kube" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// objectNameField is the field a DaemonSet pod's node affinity is written +// against. k8s.io/api declares no constant for it, and the DaemonSet controller +// writes this literal. +const objectNameField = "metadata.name" + +// +kubebuilder:rbac:groups="",resources=pods,verbs=get;list;watch + +// reportPodScheduling announces a storage-node pod that cannot be placed on this +// node's worker. +// +// It reports rather than decides: a node whose pod is unplaceable is already +// held by whichever step is waiting on the worker, and failing it here would +// turn a condition that resolves when somebody frees capacity into one that +// needs the object recreated. +// +// An unreadable pod list says nothing, because the alternative is announcing a +// scheduling failure on the strength of not having looked. +func (r *StorageNodeReconciler) reportPodScheduling( + ctx context.Context, node *simplyblockv1alpha2.StorageNode, +) { + var pods corev1.PodList + err := r.List(ctx, &pods, + client.InNamespace(node.Namespace), + client.MatchingLabels{ + kube.LabelApp: kube.AppStorageNode, + kube.LabelSimplyblockCluster: node.Spec.ClusterRef, + }) + if err != nil { + return + } + + for i := range pods.Items { + pod := &pods.Items[i] + if podWorker(pod) != node.Spec.WorkerNode { + continue + } + message, refused := podSchedulingRefusal(pod) + if !refused { + continue + } + r.emit(node, corev1.EventTypeWarning, PodSchedulingFailed, + "the storage-node pod "+pod.Name+" cannot be placed on "+ + node.Spec.WorkerNode+": "+message) + return + } +} + +// podWorker is the worker a storage-node pod belongs to. +// +// It is not spec.nodeName, because a pod that has not been placed has no node +// name and being unplaced is the case this file exists for. The DaemonSet +// controller writes the worker into the pod's node affinity as a metadata.name +// match field and leaves the binding to the scheduler, so an unplaced pod names +// its worker there and nowhere else. +func podWorker(pod *corev1.Pod) string { + if pod.Spec.NodeName != "" { + return pod.Spec.NodeName + } + affinity := pod.Spec.Affinity + if affinity == nil || affinity.NodeAffinity == nil || + affinity.NodeAffinity.RequiredDuringSchedulingIgnoredDuringExecution == nil { + return "" + } + terms := affinity.NodeAffinity. + RequiredDuringSchedulingIgnoredDuringExecution.NodeSelectorTerms + for _, term := range terms { + for _, field := range term.MatchFields { + if field.Key != objectNameField || + field.Operator != corev1.NodeSelectorOpIn || + len(field.Values) != 1 { + continue + } + return field.Values[0] + } + } + return "" +} + +// podSchedulingRefusal reports whether the scheduler has refused this pod, and +// what it said. +// +// Only Unschedulable counts. A gated pod carries the same false condition and is +// one the scheduler has not yet been asked about, so announcing it would raise a +// warning for every pod on its way to running rather than for one that is stuck. +func podSchedulingRefusal(pod *corev1.Pod) (string, bool) { + for _, condition := range pod.Status.Conditions { + if condition.Type != corev1.PodScheduled || + condition.Status != corev1.ConditionFalse || + condition.Reason != corev1.PodReasonUnschedulable { + continue + } + return condition.Message, true + } + return "", false +} diff --git a/operator/internal/controllers/node/podscheduling_test.go b/operator/internal/controllers/node/podscheduling_test.go new file mode 100644 index 000000000..3cda843d5 --- /dev/null +++ b/operator/internal/controllers/node/podscheduling_test.go @@ -0,0 +1,199 @@ +// What a node announces when the pod it runs in cannot be placed. +// +// A storage node that never comes up looks identical from the control plane's +// side whether its worker is wedged, its API is down, or Kubernetes never +// managed to start the pod at all. The last of those is the one the backend +// cannot see, and the scheduler has already written down exactly why. + +package node + +import ( + "context" + "strings" + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/client-go/tools/events" + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + "github.com/simplyblock/atlas/kube" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/testsupport" +) + +// aPendingStoragePod is a storage-node pod the scheduler has looked at and +// refused. It has no spec.nodeName, because being unplaced is the whole point, +// and it names its worker the way the DaemonSet controller does. +func aPendingStoragePod(worker, reason, message string) *corev1.Pod { + pod := aStoragePod(worker) + pod.Spec.NodeName = "" + pod.Spec.Affinity = &corev1.Affinity{NodeAffinity: &corev1.NodeAffinity{ + RequiredDuringSchedulingIgnoredDuringExecution: &corev1.NodeSelector{ + NodeSelectorTerms: []corev1.NodeSelectorTerm{{ + MatchFields: []corev1.NodeSelectorRequirement{{ + Key: "metadata.name", + Operator: corev1.NodeSelectorOpIn, + Values: []string{worker}, + }}, + }}, + }, + }} + pod.Status = corev1.PodStatus{ + Phase: corev1.PodPending, + Conditions: []corev1.PodCondition{{ + Type: corev1.PodScheduled, + Status: corev1.ConditionFalse, + Reason: reason, + Message: message, + }}, + } + return pod +} + +// aStoragePod is one the scheduler has placed. +func aStoragePod(worker string) *corev1.Pod { + return &corev1.Pod{ + ObjectMeta: metav1.ObjectMeta{ + Name: "simplyblock-storage-node-" + worker, + Namespace: "simplyblock", + Labels: map[string]string{ + kube.LabelApp: kube.AppStorageNode, + kube.LabelSimplyblockCluster: "a-cluster", + kube.LabelStorageNodeSet: "a-cluster", + }, + }, + Spec: corev1.PodSpec{NodeName: worker}, + Status: corev1.PodStatus{ + Phase: corev1.PodRunning, + Conditions: []corev1.PodCondition{{ + Type: corev1.PodScheduled, + Status: corev1.ConditionTrue, + }}, + }, + } +} + +// aHeldNode is a node whose cluster has no UUID, so the reconcile reaches the +// hold of §4.2 and performs no control-plane call. The scheduling report is not +// part of provisioning, so this is the cheapest pass that carries it. +func aHeldNode() *simplyblockv1alpha2.StorageNode { + node := &simplyblockv1alpha2.StorageNode{ + ObjectMeta: metav1.ObjectMeta{ + Name: "a-cluster-worker-1-0", + Namespace: "simplyblock", + Finalizers: []string{NodeFinalizer}, + }, + Spec: simplyblockv1alpha2.StorageNodeSpec{ + ClusterRef: "a-cluster", + WorkerNode: "worker-1", + }, + } + return node +} + +func aSchedulingReporter( + t *testing.T, node *simplyblockv1alpha2.StorageNode, pods ...client.Object, +) *StorageNodeReconciler { + t.Helper() + scheme := testsupport.NewScheme(t, corev1.AddToScheme) + cluster := &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{Name: "a-cluster", Namespace: "simplyblock"}, + } + objects := append([]client.Object{cluster, node}, pods...) + apiClient := fake.NewClientBuilder().WithScheme(scheme). + WithObjects(objects...). + WithStatusSubresource(&simplyblockv1alpha2.StorageNode{}). + Build() + + return &StorageNodeReconciler{ + Client: apiClient, + Scheme: scheme, + Recorder: events.NewFakeRecorder(64), + } +} + +func reconcileNode(t *testing.T, r *StorageNodeReconciler, + node *simplyblockv1alpha2.StorageNode) { + t.Helper() + _, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: client.ObjectKeyFromObject(node), + }) + if err != nil { + t.Fatalf("reconcile: %v", err) + } +} + +// The event exists so that a node holding for a reason the control plane cannot +// see says which reason it is. +func TestAnUnplaceablePodIsAnnouncedOnItsNode(t *testing.T) { + node := aHeldNode() + r := aSchedulingReporter(t, node, aPendingStoragePod("worker-1", + corev1.PodReasonUnschedulable, + "0/3 nodes are available: 1 Insufficient hugepages-2Mi.")) + + reconcileNode(t, r, node) + + var announcement string + for _, event := range drainReasons(r.Recorder.(*events.FakeRecorder)) { + if strings.Contains(event, PodSchedulingFailed) { + announcement = event + } + } + if announcement == "" { + t.Fatal("nothing announced PodSchedulingFailed, so the hold names no cause") + } + if !strings.Contains(announcement, "Insufficient hugepages-2Mi") { + t.Errorf("the event says %q, want the scheduler's own reason in it", announcement) + } + if !strings.Contains(announcement, "worker-1") { + t.Errorf("the event says %q, want the worker it could not be placed on", announcement) + } +} + +// A pod that is running is not news, and an event raised every pass for a +// healthy node is an alert nobody can act on. +func TestAPlacedPodIsAnnouncedNowhere(t *testing.T) { + node := aHeldNode() + r := aSchedulingReporter(t, node, aStoragePod("worker-1")) + + reconcileNode(t, r, node) + + if announced(r.Recorder.(*events.FakeRecorder), PodSchedulingFailed) { + t.Error("a placed pod was announced as a scheduling failure") + } +} + +// The DaemonSet covers every worker of the cluster, so one node's reconcile sees +// every other node's pod. A failure on a different worker is a different node's +// event, and reporting it here would announce it once per node in the cluster. +func TestAnotherWorkersFailureIsNotThisNodes(t *testing.T) { + node := aHeldNode() + r := aSchedulingReporter(t, node, aStoragePod("worker-1"), + aPendingStoragePod("worker-2", corev1.PodReasonUnschedulable, + "0/3 nodes are available: 1 Insufficient cpu.")) + + reconcileNode(t, r, node) + + if announced(r.Recorder.(*events.FakeRecorder), PodSchedulingFailed) { + t.Error("another worker's unplaceable pod was announced on this node") + } +} + +// A gated pod is one the scheduler has not tried to place yet, which is a wait +// rather than a failure. Announcing it would raise a warning for every pod that +// passes through a scheduling gate on its way to running. +func TestAGatedPodIsAWaitRatherThanAFailure(t *testing.T) { + node := aHeldNode() + r := aSchedulingReporter(t, node, aPendingStoragePod("worker-1", + corev1.PodReasonSchedulingGated, "Scheduling is blocked by a gate")) + + reconcileNode(t, r, node) + + if announced(r.Recorder.(*events.FakeRecorder), PodSchedulingFailed) { + t.Error("a gated pod was announced as a scheduling failure") + } +} diff --git a/operator/internal/controllers/node/storagenode_controller.go b/operator/internal/controllers/node/storagenode_controller.go index fa5d38b5e..b856f11d1 100644 --- a/operator/internal/controllers/node/storagenode_controller.go +++ b/operator/internal/controllers/node/storagenode_controller.go @@ -228,6 +228,12 @@ func (r *StorageNodeReconciler) Reconcile( return ctrl.Result{}, err } + // A pod the scheduler has refused holds a provisioning node in CheckingHost + // and takes a running one offline, and neither path can name the cause from + // what the control plane reports. It is announced before the branch because + // it is the same answer on both sides of it. + r.reportPodScheduling(ctx, &node) + if node.Status.UUID == "" { return r.provision(ctx, &node, cluster) } From cb0e051992341bf6f77cd577181e11674dc070d3 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Thu, 17 Sep 2026 14:37:50 +0100 Subject: [PATCH 053/206] Vendor csi-addons v0.15.0 CRDs and controller-manager into the chart (P0-5) --- ...csiaddons.openshift.io_csiaddonsnodes.yaml | 168 ++++++++++ ...hift.io_encryptionkeyrotationcronjobs.yaml | 253 ++++++++++++++ ...penshift.io_encryptionkeyrotationjobs.yaml | 196 +++++++++++ ...dons.openshift.io_networkfenceclasses.yaml | 83 +++++ .../csiaddons.openshift.io_networkfences.yaml | 209 ++++++++++++ ...ons.openshift.io_reclaimspacecronjobs.yaml | 251 ++++++++++++++ ...iaddons.openshift.io_reclaimspacejobs.yaml | 204 ++++++++++++ .../templates/rbac-csi-addons-controller.yaml | 281 ++++++++++++++++ ...hift.io_volumegroupreplicationclasses.yaml | 94 ++++++ ...ift.io_volumegroupreplicationcontents.yaml | 241 ++++++++++++++ ....openshift.io_volumegroupreplications.yaml | 309 ++++++++++++++++++ ...openshift.io_volumereplicationclasses.yaml | 96 ++++++ ...orage.openshift.io_volumereplications.yaml | 225 +++++++++++++ .../setup-csi-addons-controller.yaml | 78 +++++ .../charts/simplyblock-operator/values.yaml | 15 + .../designs/design-csi-addons-replication.md | 16 +- .../tests/test-plan-csi-addons-replication.md | 14 +- 17 files changed, 2718 insertions(+), 15 deletions(-) create mode 100644 helm-charts/charts/simplyblock-operator/templates/csiaddons.openshift.io_csiaddonsnodes.yaml create mode 100644 helm-charts/charts/simplyblock-operator/templates/csiaddons.openshift.io_encryptionkeyrotationcronjobs.yaml create mode 100644 helm-charts/charts/simplyblock-operator/templates/csiaddons.openshift.io_encryptionkeyrotationjobs.yaml create mode 100644 helm-charts/charts/simplyblock-operator/templates/csiaddons.openshift.io_networkfenceclasses.yaml create mode 100644 helm-charts/charts/simplyblock-operator/templates/csiaddons.openshift.io_networkfences.yaml create mode 100644 helm-charts/charts/simplyblock-operator/templates/csiaddons.openshift.io_reclaimspacecronjobs.yaml create mode 100644 helm-charts/charts/simplyblock-operator/templates/csiaddons.openshift.io_reclaimspacejobs.yaml create mode 100644 helm-charts/charts/simplyblock-operator/templates/rbac-csi-addons-controller.yaml create mode 100644 helm-charts/charts/simplyblock-operator/templates/replication.storage.openshift.io_volumegroupreplicationclasses.yaml create mode 100644 helm-charts/charts/simplyblock-operator/templates/replication.storage.openshift.io_volumegroupreplicationcontents.yaml create mode 100644 helm-charts/charts/simplyblock-operator/templates/replication.storage.openshift.io_volumegroupreplications.yaml create mode 100644 helm-charts/charts/simplyblock-operator/templates/replication.storage.openshift.io_volumereplicationclasses.yaml create mode 100644 helm-charts/charts/simplyblock-operator/templates/replication.storage.openshift.io_volumereplications.yaml create mode 100644 helm-charts/charts/simplyblock-operator/templates/setup-csi-addons-controller.yaml diff --git a/helm-charts/charts/simplyblock-operator/templates/csiaddons.openshift.io_csiaddonsnodes.yaml b/helm-charts/charts/simplyblock-operator/templates/csiaddons.openshift.io_csiaddonsnodes.yaml new file mode 100644 index 000000000..34670e2ea --- /dev/null +++ b/helm-charts/charts/simplyblock-operator/templates/csiaddons.openshift.io_csiaddonsnodes.yaml @@ -0,0 +1,168 @@ +# CSIAddonsNode CRD, vendored verbatim from csi-addons/kubernetes-csi-addons +# v0.15.0 (deploy/controller/crds.yaml), plus the chart's resource-policy +# annotation. The stock csi-addons controller-manager starts a controller +# per kind unconditionally, so every CRD it watches must exist for the +# manager to come up. This file installs one of them unless the cluster +# already serves the API. +{{- if and .Values.csiaddons.create (not (.Capabilities.APIVersions.Has "csiaddons.openshift.io/v1alpha1/CSIAddonsNode")) -}} +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition +metadata: + annotations: + helm.sh/resource-policy: keep + controller-gen.kubebuilder.io/version: v0.20.1 + name: csiaddonsnodes.csiaddons.openshift.io +spec: + group: csiaddons.openshift.io + names: + kind: CSIAddonsNode + listKind: CSIAddonsNodeList + plural: csiaddonsnodes + singular: csiaddonsnode + scope: Namespaced + versions: + - additionalPrinterColumns: + - jsonPath: .metadata.namespace + name: namespace + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + - jsonPath: .spec.driver.name + name: DriverName + type: string + - jsonPath: .spec.driver.endpoint + name: Endpoint + type: string + - jsonPath: .spec.driver.nodeID + name: NodeID + type: string + name: v1alpha1 + schema: + openAPIV3Schema: + description: CSIAddonsNode is the Schema for the csiaddonsnode API + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: CSIAddonsNodeSpec defines the desired state of CSIAddonsNode + properties: + driver: + description: |- + Driver is the information of the CSI Driver existing on a node. + If the driver is uninstalled, this can become empty. + properties: + endpoint: + description: |- + EndPoint is url that contains the ip-address to which the CSI-Addons + side-car listens to. + type: string + name: + description: |- + Name is the name of the CSI driver that this object refers to. + This must be the same name returned by the CSI-Addons GetIdentity() + call for that driver. The name of the driver is in the format: + `example.csi.ceph.com` + type: string + x-kubernetes-validations: + - message: name is immutable + rule: self == oldSelf + nodeID: + description: |- + NodeID is the ID of the node to identify on which node the side-car + is running. + type: string + x-kubernetes-validations: + - message: nodeID is immutable + rule: self == oldSelf + required: + - endpoint + - name + - nodeID + type: object + required: + - driver + type: object + status: + description: CSIAddonsNodeStatus defines the observed state of CSIAddonsNode + properties: + capabilities: + description: A list of capabilities advertised by the sidecar + items: + type: string + type: array + message: + description: |- + Message is a human-readable message indicating details about why the CSIAddonsNode + is in this state. + type: string + networkFenceClientStatus: + description: NetworkFenceClientStatus contains the status of the clients + required for fencing. + items: + description: NetworkFenceClientStatus contains the status of the + clients required for fencing. + properties: + ClientDetails: + items: + description: ClientDetail contains the details of the client + required for fencing. + properties: + cidrs: + description: Cidrs is the list of CIDR blocks that are + fenced. + items: + type: string + type: array + id: + description: Id is the unique identifier of the client + where it belongs to. + type: string + required: + - cidrs + - id + type: object + type: array + networkFenceClassName: + type: string + required: + - ClientDetails + - networkFenceClassName + type: object + type: array + reason: + description: |- + Reason is a brief CamelCase string that describes any failure and is meant + for machine parsing and tidy display in the CLI. + type: string + state: + description: |- + State represents the state of the CSIAddonsNode object. + It informs whether or not the CSIAddonsNode is Connected + to the CSI Driver. + type: string + type: object + required: + - spec + type: object + served: true + storage: true + subresources: + status: {} +{{- end -}} diff --git a/helm-charts/charts/simplyblock-operator/templates/csiaddons.openshift.io_encryptionkeyrotationcronjobs.yaml b/helm-charts/charts/simplyblock-operator/templates/csiaddons.openshift.io_encryptionkeyrotationcronjobs.yaml new file mode 100644 index 000000000..e0e2ea4a4 --- /dev/null +++ b/helm-charts/charts/simplyblock-operator/templates/csiaddons.openshift.io_encryptionkeyrotationcronjobs.yaml @@ -0,0 +1,253 @@ +# EncryptionKeyRotationCronJob CRD, vendored verbatim from csi-addons/kubernetes-csi-addons +# v0.15.0 (deploy/controller/crds.yaml), plus the chart's resource-policy +# annotation. The stock csi-addons controller-manager starts a controller +# per kind unconditionally, so every CRD it watches must exist for the +# manager to come up. This file installs one of them unless the cluster +# already serves the API. +{{- if and .Values.csiaddons.create (not (.Capabilities.APIVersions.Has "csiaddons.openshift.io/v1alpha1/EncryptionKeyRotationCronJob")) -}} +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition +metadata: + annotations: + helm.sh/resource-policy: keep + controller-gen.kubebuilder.io/version: v0.20.1 + name: encryptionkeyrotationcronjobs.csiaddons.openshift.io +spec: + group: csiaddons.openshift.io + names: + kind: EncryptionKeyRotationCronJob + listKind: EncryptionKeyRotationCronJobList + plural: encryptionkeyrotationcronjobs + singular: encryptionkeyrotationcronjob + scope: Namespaced + versions: + - additionalPrinterColumns: + - jsonPath: .spec.schedule + name: Schedule + type: string + - jsonPath: .spec.suspend + name: Suspend + type: boolean + - jsonPath: .status.active.name + name: Active + type: string + - jsonPath: .status.lastScheduleTime + name: Lastschedule + type: date + - jsonPath: .status.lastSuccessfulTime + name: Lastsuccessfultime + priority: 1 + type: date + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha1 + schema: + openAPIV3Schema: + description: EncryptionKeyRotationCronJob is the Schema for the encryptionkeyrotationcronjobs + API + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: EncryptionKeyRotationCronJobSpec defines the desired state + of EncryptionKeyRotationCronJob + properties: + concurrencyPolicy: + default: Forbid + description: |- + Specifies how to treat concurrent executions of a Job. + Valid values are: + - "Forbid" (default): forbids concurrent runs, skipping next run if + previous run hasn't finished yet; + - "Replace": cancels currently running job and replaces it + with a new one + enum: + - Forbid + - Replace + type: string + failedJobsHistoryLimit: + default: 1 + description: |- + The number of failed finished jobs to retain. Value must be non-negative integer. + Defaults to 1. + format: int32 + maximum: 60 + minimum: 0 + type: integer + jobTemplate: + description: Specifies the job that will be created when executing + a CronJob. + properties: + metadata: + description: |- + Standard object's metadata of the jobs created from this template. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#metadata + type: object + spec: + description: |- + Specification of the desired behavior of the job. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#spec-and-status + properties: + backOffLimit: + default: 6 + description: |- + BackOffLimit specifies the number of retries allowed before marking reclaim + space operation as failed. If not specified, defaults to 6. Maximum allowed + value is 60 and minimum allowed value is 0. + format: int32 + maximum: 60 + minimum: 0 + type: integer + retryDeadlineSeconds: + default: 600 + description: |- + RetryDeadlineSeconds specifies the duration in seconds relative to the + start time that the operation may be retried; value MUST be positive integer. + If not specified, defaults to 600 seconds. Maximum allowed + value is 1800. + format: int64 + maximum: 1800 + minimum: 0 + type: integer + target: + description: |- + Target represents tvolume target on which operation will be + performed. + properties: + persistentVolumeClaim: + description: PersistentVolumeClaim specifies the target + PersistentVolumeClaim name. + type: string + x-kubernetes-validations: + - message: persistentVolumeClaim is immutable + rule: self == oldSelf + type: object + timeout: + description: |- + Timeout specifies the timeout in seconds for the grpc request sent to the + CSI driver. + Minimum allowed value is 60. + format: int64 + minimum: 60 + type: integer + required: + - target + type: object + required: + - spec + type: object + schedule: + description: |- + The schedule in Cron format, see https://en.wikipedia.org/wiki/Cron. + A deterministic, UID-based stagger offset is applied to spread + execution across the "cronjob-stagger-window" (default: 2 hours, + set to 0 to disable) configured in the csi-addons-config ConfigMap. + pattern: .+ + type: string + startingDeadlineSeconds: + description: |- + Optional deadline in seconds for starting the job if it misses scheduled + time for any reason. Missed jobs executions will be counted as failed ones. + format: int64 + type: integer + successfulJobsHistoryLimit: + default: 3 + description: |- + The number of successful finished jobs to retain. Value must be non-negative integer. + Defaults to 3. + format: int32 + maximum: 60 + minimum: 0 + type: integer + suspend: + description: |- + This flag tells the controller to suspend subsequent executions, it does + not apply to already started executions. Defaults to false. + type: boolean + required: + - jobTemplate + - schedule + type: object + status: + description: EncryptionKeyRotationCronJobStatus defines the observed state + of EncryptionKeyRotationCronJob + properties: + active: + description: A pointer to currently running job. + properties: + apiVersion: + description: API version of the referent. + type: string + fieldPath: + description: |- + If referring to a piece of an object instead of an entire object, this string + should contain a valid JSON/Go field access statement, such as desiredState.manifest.containers[2]. + For example, if the object reference is to a container within a pod, this would take on a value like: + "spec.containers{name}" (where "name" refers to the name of the container that triggered + the event) or if no container name is specified "spec.containers[2]" (container with + index 2 in this pod). This syntax is chosen only to have some well-defined way of + referencing a part of an object. + type: string + kind: + description: |- + Kind of the referent. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + name: + description: |- + Name of the referent. + More info: https://kubernetes.io/docs/concepts/overview/working-with-objects/names/#names + type: string + namespace: + description: |- + Namespace of the referent. + More info: https://kubernetes.io/docs/concepts/overview/working-with-objects/namespaces/ + type: string + resourceVersion: + description: |- + Specific resourceVersion to which this reference is made, if any. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#concurrency-control-and-consistency + type: string + uid: + description: |- + UID of the referent. + More info: https://kubernetes.io/docs/concepts/overview/working-with-objects/names/#uids + type: string + type: object + x-kubernetes-map-type: atomic + lastScheduleTime: + description: Information when was the last time the job was successfully + scheduled. + format: date-time + type: string + lastSuccessfulTime: + description: Information when was the last time the job successfully + completed. + format: date-time + type: string + type: object + required: + - spec + type: object + served: true + storage: true + subresources: + status: {} +{{- end -}} diff --git a/helm-charts/charts/simplyblock-operator/templates/csiaddons.openshift.io_encryptionkeyrotationjobs.yaml b/helm-charts/charts/simplyblock-operator/templates/csiaddons.openshift.io_encryptionkeyrotationjobs.yaml new file mode 100644 index 000000000..8346e40b3 --- /dev/null +++ b/helm-charts/charts/simplyblock-operator/templates/csiaddons.openshift.io_encryptionkeyrotationjobs.yaml @@ -0,0 +1,196 @@ +# EncryptionKeyRotationJob CRD, vendored verbatim from csi-addons/kubernetes-csi-addons +# v0.15.0 (deploy/controller/crds.yaml), plus the chart's resource-policy +# annotation. The stock csi-addons controller-manager starts a controller +# per kind unconditionally, so every CRD it watches must exist for the +# manager to come up. This file installs one of them unless the cluster +# already serves the API. +{{- if and .Values.csiaddons.create (not (.Capabilities.APIVersions.Has "csiaddons.openshift.io/v1alpha1/EncryptionKeyRotationJob")) -}} +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition +metadata: + annotations: + helm.sh/resource-policy: keep + controller-gen.kubebuilder.io/version: v0.20.1 + name: encryptionkeyrotationjobs.csiaddons.openshift.io +spec: + group: csiaddons.openshift.io + names: + kind: EncryptionKeyRotationJob + listKind: EncryptionKeyRotationJobList + plural: encryptionkeyrotationjobs + singular: encryptionkeyrotationjob + scope: Namespaced + versions: + - additionalPrinterColumns: + - jsonPath: .metadata.namespace + name: Namespace + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + - jsonPath: .status.retries + name: Retries + type: integer + - jsonPath: .status.result + name: Result + type: string + name: v1alpha1 + schema: + openAPIV3Schema: + description: EncryptionKeyRotationJob is the Schema for the encryptionkeyrotationjobs + API + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: EncryptionKeyRotationJobSpec defines the desired state of + EncryptionKeyRotationJob + properties: + backOffLimit: + default: 6 + description: |- + BackOffLimit specifies the number of retries allowed before marking reclaim + space operation as failed. If not specified, defaults to 6. Maximum allowed + value is 60 and minimum allowed value is 0. + format: int32 + maximum: 60 + minimum: 0 + type: integer + retryDeadlineSeconds: + default: 600 + description: |- + RetryDeadlineSeconds specifies the duration in seconds relative to the + start time that the operation may be retried; value MUST be positive integer. + If not specified, defaults to 600 seconds. Maximum allowed + value is 1800. + format: int64 + maximum: 1800 + minimum: 0 + type: integer + target: + description: |- + Target represents tvolume target on which operation will be + performed. + properties: + persistentVolumeClaim: + description: PersistentVolumeClaim specifies the target PersistentVolumeClaim + name. + type: string + x-kubernetes-validations: + - message: persistentVolumeClaim is immutable + rule: self == oldSelf + type: object + timeout: + description: |- + Timeout specifies the timeout in seconds for the grpc request sent to the + CSI driver. + Minimum allowed value is 60. + format: int64 + minimum: 60 + type: integer + required: + - target + type: object + status: + description: EncryptionKeyRotationJobStatus defines the observed state + of EncryptionKeyRotationJob + properties: + completionTime: + format: date-time + type: string + conditions: + description: Conditions are the list of conditions and their status. + items: + description: Condition contains details for one aspect of the current + state of this API Resource. + properties: + lastTransitionTime: + description: |- + lastTransitionTime is the last time the condition transitioned from one status to another. + This should be when the underlying condition changed. If that is not known, then using the time when the API field changed is acceptable. + format: date-time + type: string + message: + description: |- + message is a human readable message indicating details about the transition. + This may be an empty string. + maxLength: 32768 + type: string + observedGeneration: + description: |- + observedGeneration represents the .metadata.generation that the condition was set based upon. + For instance, if .metadata.generation is currently 12, but the .status.conditions[x].observedGeneration is 9, the condition is out of date + with respect to the current state of the instance. + format: int64 + minimum: 0 + type: integer + reason: + description: |- + reason contains a programmatic identifier indicating the reason for the condition's last transition. + Producers of specific condition types may define expected values and meanings for this field, + and whether the values are considered a guaranteed API. + The value should be a CamelCase string. + This field may not be empty. + maxLength: 1024 + minLength: 1 + pattern: ^[A-Za-z]([A-Za-z0-9_,:]*[A-Za-z0-9_])?$ + type: string + status: + description: status of the condition, one of True, False, Unknown. + enum: + - "True" + - "False" + - Unknown + type: string + type: + description: type of condition in CamelCase or in foo.example.com/CamelCase. + maxLength: 316 + pattern: ^([a-z0-9]([-a-z0-9]*[a-z0-9])?(\.[a-z0-9]([-a-z0-9]*[a-z0-9])?)*/)?(([A-Za-z0-9][-A-Za-z0-9_.]*)?[A-Za-z0-9])$ + type: string + required: + - lastTransitionTime + - message + - reason + - status + - type + type: object + type: array + message: + description: Message contains any message from the EncryptionKeyRotationJob. + type: string + result: + description: Result indicates the result of EncryptionKeyRotationJob. + type: string + retries: + description: Retries indicates the number of times the operation is + retried. + format: int32 + type: integer + startTime: + format: date-time + type: string + type: object + required: + - spec + type: object + served: true + storage: true + subresources: + status: {} +{{- end -}} diff --git a/helm-charts/charts/simplyblock-operator/templates/csiaddons.openshift.io_networkfenceclasses.yaml b/helm-charts/charts/simplyblock-operator/templates/csiaddons.openshift.io_networkfenceclasses.yaml new file mode 100644 index 000000000..541ad3ba1 --- /dev/null +++ b/helm-charts/charts/simplyblock-operator/templates/csiaddons.openshift.io_networkfenceclasses.yaml @@ -0,0 +1,83 @@ +# NetworkFenceClass CRD, vendored verbatim from csi-addons/kubernetes-csi-addons +# v0.15.0 (deploy/controller/crds.yaml), plus the chart's resource-policy +# annotation. The stock csi-addons controller-manager starts a controller +# per kind unconditionally, so every CRD it watches must exist for the +# manager to come up. This file installs one of them unless the cluster +# already serves the API. +{{- if and .Values.csiaddons.create (not (.Capabilities.APIVersions.Has "csiaddons.openshift.io/v1alpha1/NetworkFenceClass")) -}} +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition +metadata: + annotations: + helm.sh/resource-policy: keep + controller-gen.kubebuilder.io/version: v0.20.1 + name: networkfenceclasses.csiaddons.openshift.io +spec: + group: csiaddons.openshift.io + names: + kind: NetworkFenceClass + listKind: NetworkFenceClassList + plural: networkfenceclasses + singular: networkfenceclass + scope: Cluster + versions: + - name: v1alpha1 + schema: + openAPIV3Schema: + description: NetworkFenceClass is the Schema for the networkfenceclasses API + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: |- + NetworkFenceClassSpec specifies parameters that an underlying storage system uses + to get client for network fencing. Upon creating a NetworkFenceClass object, a RPC will be set + to the storage system that matches the provisioner to get the client for network fencing. + properties: + parameters: + additionalProperties: + type: string + description: |- + Parameters is a key-value map with storage provisioner specific configurations for + creating volume replicas + type: object + x-kubernetes-validations: + - message: parameters are immutable + rule: self == oldSelf + provisioner: + description: Provisioner is the name of storage provisioner + type: string + x-kubernetes-validations: + - message: provisioner is immutable + rule: self == oldSelf + required: + - provisioner + type: object + x-kubernetes-validations: + - message: parameters are immutable + rule: has(self.parameters) == has(oldSelf.parameters) + status: + description: NetworkFenceClassStatus defines the observed state of NetworkFenceClass + type: object + type: object + served: true + storage: true + subresources: + status: {} +{{- end -}} diff --git a/helm-charts/charts/simplyblock-operator/templates/csiaddons.openshift.io_networkfences.yaml b/helm-charts/charts/simplyblock-operator/templates/csiaddons.openshift.io_networkfences.yaml new file mode 100644 index 000000000..31fc28027 --- /dev/null +++ b/helm-charts/charts/simplyblock-operator/templates/csiaddons.openshift.io_networkfences.yaml @@ -0,0 +1,209 @@ +# NetworkFence CRD, vendored verbatim from csi-addons/kubernetes-csi-addons +# v0.15.0 (deploy/controller/crds.yaml), plus the chart's resource-policy +# annotation. The stock csi-addons controller-manager starts a controller +# per kind unconditionally, so every CRD it watches must exist for the +# manager to come up. This file installs one of them unless the cluster +# already serves the API. +{{- if and .Values.csiaddons.create (not (.Capabilities.APIVersions.Has "csiaddons.openshift.io/v1alpha1/NetworkFence")) -}} +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition +metadata: + annotations: + helm.sh/resource-policy: keep + controller-gen.kubebuilder.io/version: v0.20.1 + name: networkfences.csiaddons.openshift.io +spec: + group: csiaddons.openshift.io + names: + kind: NetworkFence + listKind: NetworkFenceList + plural: networkfences + singular: networkfence + scope: Cluster + versions: + - additionalPrinterColumns: + - jsonPath: .spec.driver + name: Driver + type: string + - jsonPath: .spec.cidrs + name: Cidrs + type: string + - jsonPath: .spec.fenceState + name: FenceState + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + - jsonPath: .status.result + name: Result + type: string + name: v1alpha1 + schema: + openAPIV3Schema: + description: NetworkFence is the Schema for the networkfences API + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: NetworkFenceSpec defines the desired state of NetworkFence + properties: + cidrs: + description: Cidrs contains a list of CIDR blocks, which are required + to be fenced. + items: + type: string + type: array + driver: + description: Driver contains the name of CSI driver, required if NetworkFenceClassName + is absent + type: string + x-kubernetes-validations: + - message: driver is immutable + rule: self == oldSelf + fenceState: + default: Fenced + description: |- + FenceState contains the desired state for the CIDRs + mentioned in the Spec. i.e. Fenced or Unfenced + enum: + - Fenced + - Unfenced + type: string + networkFenceClassName: + description: NetworkFenceClassName contains the name of the NetworkFenceClass + type: string + x-kubernetes-validations: + - message: networkFenceClassName is immutable + rule: self == oldSelf + parameters: + additionalProperties: + type: string + description: Parameters is used to pass additional parameters to the + CSI driver. + type: object + x-kubernetes-validations: + - message: parameters are immutable + rule: self == oldSelf + secret: + description: Secret is a kubernetes secret, which is required to perform + the fence/unfence operation. + properties: + name: + description: Name specifies the name of the secret. + type: string + x-kubernetes-validations: + - message: name is immutable + rule: self == oldSelf + namespace: + description: |- + Namespace specifies the namespace in which the secret + is located. + type: string + x-kubernetes-validations: + - message: namespace is immutable + rule: self == oldSelf + type: object + x-kubernetes-validations: + - message: secret is immutable + rule: self == oldSelf + required: + - cidrs + - fenceState + type: object + x-kubernetes-validations: + - message: one of driver or networkFenceClassName must be present + rule: has(self.driver) || has(self.networkFenceClassName) + - message: secret must be present when networkFenceClassName is not specified + rule: has(self.networkFenceClassName) || has(self.secret) + status: + description: NetworkFenceStatus defines the observed state of NetworkFence + properties: + conditions: + description: Conditions are the list of conditions and their status. + items: + description: Condition contains details for one aspect of the current + state of this API Resource. + properties: + lastTransitionTime: + description: |- + lastTransitionTime is the last time the condition transitioned from one status to another. + This should be when the underlying condition changed. If that is not known, then using the time when the API field changed is acceptable. + format: date-time + type: string + message: + description: |- + message is a human readable message indicating details about the transition. + This may be an empty string. + maxLength: 32768 + type: string + observedGeneration: + description: |- + observedGeneration represents the .metadata.generation that the condition was set based upon. + For instance, if .metadata.generation is currently 12, but the .status.conditions[x].observedGeneration is 9, the condition is out of date + with respect to the current state of the instance. + format: int64 + minimum: 0 + type: integer + reason: + description: |- + reason contains a programmatic identifier indicating the reason for the condition's last transition. + Producers of specific condition types may define expected values and meanings for this field, + and whether the values are considered a guaranteed API. + The value should be a CamelCase string. + This field may not be empty. + maxLength: 1024 + minLength: 1 + pattern: ^[A-Za-z]([A-Za-z0-9_,:]*[A-Za-z0-9_])?$ + type: string + status: + description: status of the condition, one of True, False, Unknown. + enum: + - "True" + - "False" + - Unknown + type: string + type: + description: type of condition in CamelCase or in foo.example.com/CamelCase. + maxLength: 316 + pattern: ^([a-z0-9]([-a-z0-9]*[a-z0-9])?(\.[a-z0-9]([-a-z0-9]*[a-z0-9])?)*/)?(([A-Za-z0-9][-A-Za-z0-9_.]*)?[A-Za-z0-9])$ + type: string + required: + - lastTransitionTime + - message + - reason + - status + - type + type: object + type: array + message: + description: Message contains any message from the NetworkFence operation. + type: string + result: + description: Result indicates the result of Network Fence/Unfence + operation. + type: string + type: object + required: + - spec + type: object + served: true + storage: true + subresources: + status: {} +{{- end -}} diff --git a/helm-charts/charts/simplyblock-operator/templates/csiaddons.openshift.io_reclaimspacecronjobs.yaml b/helm-charts/charts/simplyblock-operator/templates/csiaddons.openshift.io_reclaimspacecronjobs.yaml new file mode 100644 index 000000000..ef2c90840 --- /dev/null +++ b/helm-charts/charts/simplyblock-operator/templates/csiaddons.openshift.io_reclaimspacecronjobs.yaml @@ -0,0 +1,251 @@ +# ReclaimSpaceCronJob CRD, vendored verbatim from csi-addons/kubernetes-csi-addons +# v0.15.0 (deploy/controller/crds.yaml), plus the chart's resource-policy +# annotation. The stock csi-addons controller-manager starts a controller +# per kind unconditionally, so every CRD it watches must exist for the +# manager to come up. This file installs one of them unless the cluster +# already serves the API. +{{- if and .Values.csiaddons.create (not (.Capabilities.APIVersions.Has "csiaddons.openshift.io/v1alpha1/ReclaimSpaceCronJob")) -}} +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition +metadata: + annotations: + helm.sh/resource-policy: keep + controller-gen.kubebuilder.io/version: v0.20.1 + name: reclaimspacecronjobs.csiaddons.openshift.io +spec: + group: csiaddons.openshift.io + names: + kind: ReclaimSpaceCronJob + listKind: ReclaimSpaceCronJobList + plural: reclaimspacecronjobs + singular: reclaimspacecronjob + scope: Namespaced + versions: + - additionalPrinterColumns: + - jsonPath: .spec.schedule + name: Schedule + type: string + - jsonPath: .spec.suspend + name: Suspend + type: boolean + - jsonPath: .status.active.name + name: Active + type: string + - jsonPath: .status.lastScheduleTime + name: Lastschedule + type: date + - jsonPath: .status.lastSuccessfulTime + name: Lastsuccessfultime + priority: 1 + type: date + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha1 + schema: + openAPIV3Schema: + description: ReclaimSpaceCronJob is the Schema for the reclaimspacecronjobs + API + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: ReclaimSpaceCronJobSpec defines the desired state of ReclaimSpaceJob + properties: + concurrencyPolicy: + default: Forbid + description: |- + Specifies how to treat concurrent executions of a Job. + Valid values are: + - "Forbid" (default): forbids concurrent runs, skipping next run if + previous run hasn't finished yet; + - "Replace": cancels currently running job and replaces it + with a new one + enum: + - Forbid + - Replace + type: string + failedJobsHistoryLimit: + default: 1 + description: |- + The number of failed finished jobs to retain. Value must be non-negative integer. + Defaults to 1. + format: int32 + maximum: 60 + minimum: 0 + type: integer + jobTemplate: + description: Specifies the job that will be created when executing + a CronJob. + properties: + metadata: + description: |- + Standard object's metadata of the jobs created from this template. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#metadata + type: object + spec: + description: |- + Specification of the desired behavior of the job. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#spec-and-status + properties: + backOffLimit: + default: 6 + description: |- + BackOffLimit specifies the number of retries allowed before marking reclaim + space operation as failed. If not specified, defaults to 6. Maximum allowed + value is 60 and minimum allowed value is 0. + format: int32 + maximum: 60 + minimum: 0 + type: integer + retryDeadlineSeconds: + default: 600 + description: |- + RetryDeadlineSeconds specifies the duration in seconds relative to the + start time that the operation may be retried; value MUST be positive integer. + If not specified, defaults to 600 seconds. Maximum allowed + value is 1800. + format: int64 + maximum: 1800 + minimum: 0 + type: integer + target: + description: |- + Target represents volume target on which the operation will be + performed. + properties: + persistentVolumeClaim: + description: PersistentVolumeClaim specifies the target + PersistentVolumeClaim name. + type: string + x-kubernetes-validations: + - message: persistentVolumeClaim is immutable + rule: self == oldSelf + type: object + timeout: + description: |- + Timeout specifies the timeout in seconds for the grpc request sent to the + CSI driver. If not specified, defaults to global reclaimspace timeout. + Minimum allowed value is 60. + format: int64 + minimum: 60 + type: integer + required: + - target + type: object + required: + - spec + type: object + schedule: + description: |- + The schedule in Cron format, see https://en.wikipedia.org/wiki/Cron. + A deterministic, UID-based stagger offset is applied to spread + execution across the "cronjob-stagger-window" (default: 2 hours, + set to 0 to disable) configured in the csi-addons-config ConfigMap. + pattern: .+ + type: string + startingDeadlineSeconds: + description: |- + Optional deadline in seconds for starting the job if it misses scheduled + time for any reason. Missed jobs executions will be counted as failed ones. + format: int64 + type: integer + successfulJobsHistoryLimit: + default: 3 + description: |- + The number of successful finished jobs to retain. Value must be non-negative integer. + Defaults to 3. + format: int32 + maximum: 60 + minimum: 0 + type: integer + suspend: + description: |- + This flag tells the controller to suspend subsequent executions, it does + not apply to already started executions. Defaults to false. + type: boolean + required: + - jobTemplate + - schedule + type: object + status: + description: ReclaimSpaceCronJobStatus defines the observed state of ReclaimSpaceJob + properties: + active: + description: A pointer to currently running job. + properties: + apiVersion: + description: API version of the referent. + type: string + fieldPath: + description: |- + If referring to a piece of an object instead of an entire object, this string + should contain a valid JSON/Go field access statement, such as desiredState.manifest.containers[2]. + For example, if the object reference is to a container within a pod, this would take on a value like: + "spec.containers{name}" (where "name" refers to the name of the container that triggered + the event) or if no container name is specified "spec.containers[2]" (container with + index 2 in this pod). This syntax is chosen only to have some well-defined way of + referencing a part of an object. + type: string + kind: + description: |- + Kind of the referent. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + name: + description: |- + Name of the referent. + More info: https://kubernetes.io/docs/concepts/overview/working-with-objects/names/#names + type: string + namespace: + description: |- + Namespace of the referent. + More info: https://kubernetes.io/docs/concepts/overview/working-with-objects/namespaces/ + type: string + resourceVersion: + description: |- + Specific resourceVersion to which this reference is made, if any. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#concurrency-control-and-consistency + type: string + uid: + description: |- + UID of the referent. + More info: https://kubernetes.io/docs/concepts/overview/working-with-objects/names/#uids + type: string + type: object + x-kubernetes-map-type: atomic + lastScheduleTime: + description: Information when was the last time the job was successfully + scheduled. + format: date-time + type: string + lastSuccessfulTime: + description: Information when was the last time the job successfully + completed. + format: date-time + type: string + type: object + required: + - spec + type: object + served: true + storage: true + subresources: + status: {} +{{- end -}} diff --git a/helm-charts/charts/simplyblock-operator/templates/csiaddons.openshift.io_reclaimspacejobs.yaml b/helm-charts/charts/simplyblock-operator/templates/csiaddons.openshift.io_reclaimspacejobs.yaml new file mode 100644 index 000000000..894848b3d --- /dev/null +++ b/helm-charts/charts/simplyblock-operator/templates/csiaddons.openshift.io_reclaimspacejobs.yaml @@ -0,0 +1,204 @@ +# ReclaimSpaceJob CRD, vendored verbatim from csi-addons/kubernetes-csi-addons +# v0.15.0 (deploy/controller/crds.yaml), plus the chart's resource-policy +# annotation. The stock csi-addons controller-manager starts a controller +# per kind unconditionally, so every CRD it watches must exist for the +# manager to come up. This file installs one of them unless the cluster +# already serves the API. +{{- if and .Values.csiaddons.create (not (.Capabilities.APIVersions.Has "csiaddons.openshift.io/v1alpha1/ReclaimSpaceJob")) -}} +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition +metadata: + annotations: + helm.sh/resource-policy: keep + controller-gen.kubebuilder.io/version: v0.20.1 + name: reclaimspacejobs.csiaddons.openshift.io +spec: + group: csiaddons.openshift.io + names: + kind: ReclaimSpaceJob + listKind: ReclaimSpaceJobList + plural: reclaimspacejobs + singular: reclaimspacejob + scope: Namespaced + versions: + - additionalPrinterColumns: + - jsonPath: .metadata.namespace + name: Namespace + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + - jsonPath: .status.retries + name: Retries + type: integer + - jsonPath: .status.result + name: Result + type: string + - jsonPath: .status.reclaimedSpace + name: ReclaimedSpace + priority: 1 + type: string + name: v1alpha1 + schema: + openAPIV3Schema: + description: ReclaimSpaceJob is the Schema for the reclaimspacejobs API + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: ReclaimSpaceJobSpec defines the desired state of ReclaimSpaceJob + properties: + backOffLimit: + default: 6 + description: |- + BackOffLimit specifies the number of retries allowed before marking reclaim + space operation as failed. If not specified, defaults to 6. Maximum allowed + value is 60 and minimum allowed value is 0. + format: int32 + maximum: 60 + minimum: 0 + type: integer + retryDeadlineSeconds: + default: 600 + description: |- + RetryDeadlineSeconds specifies the duration in seconds relative to the + start time that the operation may be retried; value MUST be positive integer. + If not specified, defaults to 600 seconds. Maximum allowed + value is 1800. + format: int64 + maximum: 1800 + minimum: 0 + type: integer + target: + description: |- + Target represents volume target on which the operation will be + performed. + properties: + persistentVolumeClaim: + description: PersistentVolumeClaim specifies the target PersistentVolumeClaim + name. + type: string + x-kubernetes-validations: + - message: persistentVolumeClaim is immutable + rule: self == oldSelf + type: object + timeout: + description: |- + Timeout specifies the timeout in seconds for the grpc request sent to the + CSI driver. If not specified, defaults to global reclaimspace timeout. + Minimum allowed value is 60. + format: int64 + minimum: 60 + type: integer + required: + - target + type: object + status: + description: ReclaimSpaceJobStatus defines the observed state of ReclaimSpaceJob + properties: + completionTime: + format: date-time + type: string + conditions: + description: Conditions are the list of conditions and their status. + items: + description: Condition contains details for one aspect of the current + state of this API Resource. + properties: + lastTransitionTime: + description: |- + lastTransitionTime is the last time the condition transitioned from one status to another. + This should be when the underlying condition changed. If that is not known, then using the time when the API field changed is acceptable. + format: date-time + type: string + message: + description: |- + message is a human readable message indicating details about the transition. + This may be an empty string. + maxLength: 32768 + type: string + observedGeneration: + description: |- + observedGeneration represents the .metadata.generation that the condition was set based upon. + For instance, if .metadata.generation is currently 12, but the .status.conditions[x].observedGeneration is 9, the condition is out of date + with respect to the current state of the instance. + format: int64 + minimum: 0 + type: integer + reason: + description: |- + reason contains a programmatic identifier indicating the reason for the condition's last transition. + Producers of specific condition types may define expected values and meanings for this field, + and whether the values are considered a guaranteed API. + The value should be a CamelCase string. + This field may not be empty. + maxLength: 1024 + minLength: 1 + pattern: ^[A-Za-z]([A-Za-z0-9_,:]*[A-Za-z0-9_])?$ + type: string + status: + description: status of the condition, one of True, False, Unknown. + enum: + - "True" + - "False" + - Unknown + type: string + type: + description: type of condition in CamelCase or in foo.example.com/CamelCase. + maxLength: 316 + pattern: ^([a-z0-9]([-a-z0-9]*[a-z0-9])?(\.[a-z0-9]([-a-z0-9]*[a-z0-9])?)*/)?(([A-Za-z0-9][-A-Za-z0-9_.]*)?[A-Za-z0-9])$ + type: string + required: + - lastTransitionTime + - message + - reason + - status + - type + type: object + type: array + message: + description: Message contains any message from the ReclaimSpaceJob. + type: string + reclaimedSpace: + anyOf: + - type: integer + - type: string + description: ReclaimedSpace indicates the amount of space reclaimed. + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + result: + description: Result indicates the result of ReclaimSpaceJob. + type: string + retries: + description: Retries indicates the number of times the operation is + retried. + format: int32 + type: integer + startTime: + format: date-time + type: string + type: object + required: + - spec + type: object + served: true + storage: true + subresources: + status: {} +{{- end -}} diff --git a/helm-charts/charts/simplyblock-operator/templates/rbac-csi-addons-controller.yaml b/helm-charts/charts/simplyblock-operator/templates/rbac-csi-addons-controller.yaml new file mode 100644 index 000000000..7f333720c --- /dev/null +++ b/helm-charts/charts/simplyblock-operator/templates/rbac-csi-addons-controller.yaml @@ -0,0 +1,281 @@ +# Identity and permissions of the csi-addons controller-manager, vendored from +# csi-addons/kubernetes-csi-addons v0.15.0 (deploy/controller/rbac.yaml) with +# names prefixed simplyblock- and namespaces bound to the release. The upstream +# file's four unbound editor and viewer convenience roles are deliberately not +# vendored: nothing binds them, so shipping them would grant nothing and audit +# as dead surface. What remains is exactly what the manager's ServiceAccount +# holds. +{{- if .Values.csiaddons.create -}} +--- +apiVersion: v1 +kind: ServiceAccount +metadata: + name: simplyblock-csi-addons-controller-manager + namespace: {{ .Release.Namespace }} +--- +apiVersion: rbac.authorization.k8s.io/v1 +kind: Role +metadata: + name: simplyblock-csi-addons-leader-election-role + namespace: {{ .Release.Namespace }} +rules: + # rbac-justified: controller-runtime leader election holds a Lease (and the + # legacy ConfigMap lock) in the manager's own namespace, and records election + # events there. + - apiGroups: + - "" + resources: + - configmaps + verbs: + - get + - list + - watch + - create + - update + - patch + - delete + - apiGroups: + - coordination.k8s.io + resources: + - leases + verbs: + - get + - list + - watch + - create + - update + - patch + - delete + - apiGroups: + - "" + resources: + - events + verbs: + - create + - patch +--- +apiVersion: rbac.authorization.k8s.io/v1 +kind: ClusterRole +metadata: + name: simplyblock-csi-addons-manager-role +rules: + # rbac-justified: the reconcilers emit events on the objects they drive + # (VolumeReplication, ReclaimSpaceJob, NetworkFence) in every namespace. + - apiGroups: + - "" + resources: + - events + verbs: + - create + - patch + # rbac-justified: the manager resolves the sidecar pod behind each + # CSIAddonsNode to build its gRPC connection, and lists namespaces to fan + # per-namespace ReclaimSpaceCronJob scheduling out cluster-wide. + - apiGroups: + - "" + resources: + - namespaces + - pods + verbs: + - get + - list + - watch + # rbac-justified: the VolumeReplication reconciler resolves the protected + # PVC, and the reclaim-space and volume-health paths annotate it. The + # finalizer update protects a replicated PVC from deletion mid-operation. + - apiGroups: + - "" + resources: + - persistentvolumeclaims + verbs: + - get + - list + - patch + - update + - watch + - apiGroups: + - "" + resources: + - persistentvolumeclaims/finalizers + verbs: + - update + # rbac-justified: the PVC's volume handle (the id every csi-addons RPC + # addresses) lives on the PV, and the reclaim-space controller records + # progress annotations there. + - apiGroups: + - "" + resources: + - persistentvolumes + verbs: + - get + - list + - update + - watch + # rbac-justified: the manager reads the sidecars' leader-election Leases to + # pick the active sidecar per driver. Its own election lock is the + # namespaced Role above. + - apiGroups: + - coordination.k8s.io + resources: + - leases + verbs: + - get + - list + - watch + # rbac-justified: the manager owns these kinds: sidecars publish + # CSIAddonsNode, cron controllers materialize their child jobs, and every + # reconciler maintains status and finalizers on its own objects. + - apiGroups: + - csiaddons.openshift.io + resources: + - csiaddonsnodes + - encryptionkeyrotationcronjobs + - encryptionkeyrotationjobs + - networkfenceclasses + - networkfences + - reclaimspacecronjobs + - reclaimspacejobs + verbs: + - create + - delete + - get + - list + - patch + - update + - watch + - apiGroups: + - csiaddons.openshift.io + resources: + - csiaddonsnodes/finalizers + - encryptionkeyrotationcronjobs/finalizers + - encryptionkeyrotationjobs/finalizers + - networkfenceclasses/finalizers + - networkfences/finalizers + - reclaimspacecronjobs/finalizers + - reclaimspacejobs/finalizers + verbs: + - update + - apiGroups: + - csiaddons.openshift.io + resources: + - csiaddonsnodes/status + - encryptionkeyrotationcronjobs/status + - encryptionkeyrotationjobs/status + - networkfenceclasses/status + - networkfences/status + - reclaimspacecronjobs/status + - reclaimspacejobs/status + verbs: + - get + - patch + - update + # rbac-justified: classes are read-only inputs (parameters, provisioner + # matching). The replication objects themselves are reconciled, so they and + # their status and finalizers are written, and VolumeGroupReplicationContent + # is created by the group flow the manager owns. + - apiGroups: + - replication.storage.openshift.io + resources: + - volumegroupreplicationclasses + - volumereplicationclasses + verbs: + - get + - list + - watch + - apiGroups: + - replication.storage.openshift.io + resources: + - volumegroupreplicationcontents + verbs: + - create + - delete + - get + - list + - patch + - update + - watch + - apiGroups: + - replication.storage.openshift.io + resources: + - volumegroupreplicationcontents/finalizers + - volumegroupreplications/finalizers + - volumereplications/finalizers + verbs: + - update + - apiGroups: + - replication.storage.openshift.io + resources: + - volumegroupreplicationcontents/status + - volumegroupreplications/status + verbs: + - get + - patch + - update + - apiGroups: + - replication.storage.openshift.io + resources: + - volumegroupreplications + verbs: + - get + - list + - patch + - update + - watch + - apiGroups: + - replication.storage.openshift.io + resources: + - volumereplications + verbs: + - create + - delete + - get + - list + - update + - watch + - apiGroups: + - replication.storage.openshift.io + resources: + - volumereplications/status + verbs: + - get + - list + - update + # rbac-justified: the scheduling and peer-matching logic resolves a PVC's + # StorageClass and its attachment state, both read-only. + - apiGroups: + - storage.k8s.io + resources: + - storageclasses + - volumeattachments + verbs: + - get + - list + - watch +--- +apiVersion: rbac.authorization.k8s.io/v1 +kind: RoleBinding +metadata: + name: simplyblock-csi-addons-leader-election-rolebinding + namespace: {{ .Release.Namespace }} +roleRef: + apiGroup: rbac.authorization.k8s.io + kind: Role + name: simplyblock-csi-addons-leader-election-role +subjects: + - kind: ServiceAccount + name: simplyblock-csi-addons-controller-manager + namespace: {{ .Release.Namespace }} +--- +apiVersion: rbac.authorization.k8s.io/v1 +kind: ClusterRoleBinding +metadata: + name: simplyblock-csi-addons-manager-rolebinding +roleRef: + apiGroup: rbac.authorization.k8s.io + kind: ClusterRole + name: simplyblock-csi-addons-manager-role +subjects: + - kind: ServiceAccount + name: simplyblock-csi-addons-controller-manager + namespace: {{ .Release.Namespace }} +{{- end -}} diff --git a/helm-charts/charts/simplyblock-operator/templates/replication.storage.openshift.io_volumegroupreplicationclasses.yaml b/helm-charts/charts/simplyblock-operator/templates/replication.storage.openshift.io_volumegroupreplicationclasses.yaml new file mode 100644 index 000000000..2ea2a447c --- /dev/null +++ b/helm-charts/charts/simplyblock-operator/templates/replication.storage.openshift.io_volumegroupreplicationclasses.yaml @@ -0,0 +1,94 @@ +# VolumeGroupReplicationClass CRD, vendored verbatim from csi-addons/kubernetes-csi-addons +# v0.15.0 (deploy/controller/crds.yaml), plus the chart's resource-policy +# annotation. The stock csi-addons controller-manager starts a controller +# per kind unconditionally, so every CRD it watches must exist for the +# manager to come up. This file installs one of them unless the cluster +# already serves the API. +{{- if and .Values.csiaddons.create (not (.Capabilities.APIVersions.Has "replication.storage.openshift.io/v1alpha1/VolumeGroupReplicationClass")) -}} +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition +metadata: + annotations: + helm.sh/resource-policy: keep + controller-gen.kubebuilder.io/version: v0.20.1 + name: volumegroupreplicationclasses.replication.storage.openshift.io +spec: + group: replication.storage.openshift.io + names: + kind: VolumeGroupReplicationClass + listKind: VolumeGroupReplicationClassList + plural: volumegroupreplicationclasses + shortNames: + - vgrc + singular: volumegroupreplicationclass + scope: Cluster + versions: + - additionalPrinterColumns: + - jsonPath: .spec.provisioner + name: provisioner + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha1 + schema: + openAPIV3Schema: + description: VolumeGroupReplicationClass is the Schema for the volumegroupreplicationclasses + API + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: |- + VolumeGroupReplicationClassSpec specifies parameters that an underlying storage system uses + when creating a volumegroup replica. A specific VolumeGroupReplicationClass is used by specifying + its name in a VolumeGroupReplication object. + properties: + parameters: + additionalProperties: + type: string + description: |- + Parameters is a key-value map with storage provisioner specific configurations for + creating volume group replicas + type: object + x-kubernetes-validations: + - message: parameters are immutable + rule: self == oldSelf + provisioner: + description: Provisioner is the name of storage provisioner + type: string + x-kubernetes-validations: + - message: provisioner is immutable + rule: self == oldSelf + required: + - provisioner + type: object + x-kubernetes-validations: + - message: parameters are immutable + rule: has(self.parameters) == has(oldSelf.parameters) + status: + description: VolumeGroupReplicationClassStatus defines the observed state + of VolumeGroupReplicationClass + type: object + type: object + served: true + storage: true + subresources: + status: {} +{{- end -}} diff --git a/helm-charts/charts/simplyblock-operator/templates/replication.storage.openshift.io_volumegroupreplicationcontents.yaml b/helm-charts/charts/simplyblock-operator/templates/replication.storage.openshift.io_volumegroupreplicationcontents.yaml new file mode 100644 index 000000000..1a6374eb4 --- /dev/null +++ b/helm-charts/charts/simplyblock-operator/templates/replication.storage.openshift.io_volumegroupreplicationcontents.yaml @@ -0,0 +1,241 @@ +# VolumeGroupReplicationContent CRD, vendored verbatim from csi-addons/kubernetes-csi-addons +# v0.15.0 (deploy/controller/crds.yaml), plus the chart's resource-policy +# annotation. The stock csi-addons controller-manager starts a controller +# per kind unconditionally, so every CRD it watches must exist for the +# manager to come up. This file installs one of them unless the cluster +# already serves the API. +{{- if and .Values.csiaddons.create (not (.Capabilities.APIVersions.Has "replication.storage.openshift.io/v1alpha1/VolumeGroupReplicationContent")) -}} +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition +metadata: + annotations: + helm.sh/resource-policy: keep + controller-gen.kubebuilder.io/version: v0.20.1 + name: volumegroupreplicationcontents.replication.storage.openshift.io +spec: + group: replication.storage.openshift.io + names: + kind: VolumeGroupReplicationContent + listKind: VolumeGroupReplicationContentList + plural: volumegroupreplicationcontents + shortNames: + - vgrcontent + singular: volumegroupreplicationcontent + scope: Cluster + versions: + - additionalPrinterColumns: + - jsonPath: .spec.volumeGroupReplicationClassName + name: VolumeGroupReplicationClass + type: string + - jsonPath: .spec.volumeGroupReplicationRef.name + name: VolumeGroupReplication + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha1 + schema: + openAPIV3Schema: + description: VolumeGroupReplicationContent is the Schema for the volumegroupreplicationcontents + API + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: VolumeGroupReplicationContentSpec defines the desired state + of VolumeGroupReplicationContent + properties: + provisioner: + description: |- + provisioner is the name of the CSI driver used to create the physical + volume group on + the underlying storage system. + This MUST be the same as the name returned by the CSI GetPluginName() call for + that driver. + Required. + type: string + source: + description: |- + Source specifies whether the volume group is (or should be) dynamically provisioned + or already exists using the volumes listed here, and just requires a + Kubernetes object representation. + Required. + properties: + volumeHandles: + description: |- + VolumeHandles is a list of volume handles on the backend to be grouped + and replicated. + items: + type: string + type: array + required: + - volumeHandles + type: object + volumeGroupAttributes: + additionalProperties: + type: string + description: volumeGroupAttributes holds the contextual information + of the volume group. + type: object + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + volumeGroupReplicationClassName: + description: |- + VolumeGroupReplicationClassName is the name of the VolumeGroupReplicationClass from + which this group replication was (or will be) created. + Required. + type: string + volumeGroupReplicationHandle: + description: |- + VolumeGroupReplicationHandle is a unique id returned by the CSI driver + to identify the VolumeGroupReplication on the storage system. + type: string + volumeGroupReplicationRef: + description: |- + VolumeGroupreplicationRef specifies the VolumeGroupReplication object to which this + VolumeGroupReplicationContent object is bound. + VolumeGroupReplication.Spec.VolumeGroupReplicationContentName field must reference to + this VolumeGroupReplicationContent's name for the bidirectional binding to be valid. + For a pre-existing VolumeGroupReplication object, MUST provide an empty/nil value for + VolumeGroupReplicationRef for the auto-binding to happen. + properties: + apiVersion: + description: API version of the referent. + type: string + fieldPath: + description: |- + If referring to a piece of an object instead of an entire object, this string + should contain a valid JSON/Go field access statement, such as desiredState.manifest.containers[2]. + For example, if the object reference is to a container within a pod, this would take on a value like: + "spec.containers{name}" (where "name" refers to the name of the container that triggered + the event) or if no container name is specified "spec.containers[2]" (container with + index 2 in this pod). This syntax is chosen only to have some well-defined way of + referencing a part of an object. + type: string + kind: + description: |- + Kind of the referent. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + name: + description: |- + Name of the referent. + More info: https://kubernetes.io/docs/concepts/overview/working-with-objects/names/#names + type: string + namespace: + description: |- + Namespace of the referent. + More info: https://kubernetes.io/docs/concepts/overview/working-with-objects/namespaces/ + type: string + resourceVersion: + description: |- + Specific resourceVersion to which this reference is made, if any. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#concurrency-control-and-consistency + type: string + uid: + description: |- + UID of the referent. + More info: https://kubernetes.io/docs/concepts/overview/working-with-objects/names/#uids + type: string + type: object + x-kubernetes-map-type: atomic + x-kubernetes-validations: + - message: volumeGroupReplicationRef.name, volumeGroupReplicationRef.namespace + and volumeGroupReplicationRef.uid must be set if volumeGroupReplicationRef + is defined + rule: 'self != null ? has(self.name) && has(self.__namespace__) + && has(self.uid) : true' + required: + - provisioner + - source + - volumeGroupReplicationClassName + type: object + status: + description: VolumeGroupReplicationContentStatus defines the status of + VolumeGroupReplicationContent + properties: + destinationVolumeGroupID: + description: |- + DestinationVolumeGroupID is the volume group ID on the + destination/target side, as reported by the SP. + type: string + persistentVolumeMappingList: + description: |- + PersistentVolumeMappingList is the list of PVs for the group + replication, enriched with source and destination volume handles. + This replaces PersistentVolumeRefList. When both fields are + present, consumers SHOULD prefer this field. + The maximum number of allowed PVs in the group is 100. + items: + description: |- + PersistentVolumeMapping contains the PV reference along with its + source and destination volume handles. The destination handle is + populated only when the SP supports GET_REPLICATION_DESTINATION_INFO + and destination info is available. + properties: + destinationVolumeHandle: + description: |- + DestinationVolumeHandle is the CSI volume handle on the + destination/target cluster, as reported by the SP. + This field is empty when destination info is not available. + type: string + name: + description: Name is the name of the PersistentVolume. + type: string + volumeHandle: + description: |- + VolumeHandle is the CSI volume handle of this PV on the + source cluster (i.e. PV.Spec.CSI.VolumeHandle). + type: string + required: + - name + - volumeHandle + type: object + type: array + persistentVolumeRefList: + description: |- + PersistentVolumeRefList is the list of PV for the group replication + The maximum number of allowed PV in the group is 100. + Deprecated: Use PersistentVolumeMappingList instead, which includes + source and destination volume handle information per PV. + items: + description: |- + LocalObjectReference contains enough information to let you locate the + referenced object inside the same namespace. + properties: + name: + default: "" + description: |- + Name of the referent. + This field is effectively required, but due to backwards compatibility is + allowed to be empty. Instances of this type with an empty value here are + almost certainly wrong. + More info: https://kubernetes.io/docs/concepts/overview/working-with-objects/names/#names + type: string + type: object + x-kubernetes-map-type: atomic + type: array + type: object + type: object + served: true + storage: true + subresources: + status: {} +{{- end -}} diff --git a/helm-charts/charts/simplyblock-operator/templates/replication.storage.openshift.io_volumegroupreplications.yaml b/helm-charts/charts/simplyblock-operator/templates/replication.storage.openshift.io_volumegroupreplications.yaml new file mode 100644 index 000000000..5d212a201 --- /dev/null +++ b/helm-charts/charts/simplyblock-operator/templates/replication.storage.openshift.io_volumegroupreplications.yaml @@ -0,0 +1,309 @@ +# VolumeGroupReplication CRD, vendored verbatim from csi-addons/kubernetes-csi-addons +# v0.15.0 (deploy/controller/crds.yaml), plus the chart's resource-policy +# annotation. The stock csi-addons controller-manager starts a controller +# per kind unconditionally, so every CRD it watches must exist for the +# manager to come up. This file installs one of them unless the cluster +# already serves the API. +{{- if and .Values.csiaddons.create (not (.Capabilities.APIVersions.Has "replication.storage.openshift.io/v1alpha1/VolumeGroupReplication")) -}} +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition +metadata: + annotations: + helm.sh/resource-policy: keep + controller-gen.kubebuilder.io/version: v0.20.1 + name: volumegroupreplications.replication.storage.openshift.io +spec: + group: replication.storage.openshift.io + names: + kind: VolumeGroupReplication + listKind: VolumeGroupReplicationList + plural: volumegroupreplications + shortNames: + - vgr + singular: volumegroupreplication + scope: Namespaced + versions: + - additionalPrinterColumns: + - jsonPath: .spec.volumeGroupReplicationClassName + name: VolumeGroupReplicationClass + type: string + - jsonPath: .spec.volumeGroupReplicationContentName + name: VolumeGroupReplicationContent + type: string + - jsonPath: .spec.replicationState + name: desiredState + type: string + - jsonPath: .status.state + name: currentState + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha1 + schema: + openAPIV3Schema: + description: VolumeGroupReplication is the Schema for the volumegroupreplications + API + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: VolumeGroupReplicationSpec defines the desired state of VolumeGroupReplication + properties: + autoResync: + default: false + description: |- + AutoResync represents the group to be auto resynced when + ReplicationState is "secondary" + type: boolean + external: + default: false + description: |- + External represents if VolumeGroupReplication should be reconciled by the csi-addons controller + or an external controller managed by the storage vendor. + type: boolean + x-kubernetes-validations: + - message: source is immutable + rule: self == oldSelf + replicationState: + description: |- + ReplicationState represents the replication operation to be performed on the group. + Supported operations are "primary", "secondary" and "resync" + enum: + - primary + - secondary + - resync + type: string + source: + description: |- + Source specifies where a group replications will be created from. + This field is immutable after creation. + Required. + properties: + selector: + description: |- + Selector is a label query over persistent volume claims that are to be + grouped together for replication. + properties: + matchExpressions: + description: matchExpressions is a list of label selector + requirements. The requirements are ANDed. + items: + description: |- + A label selector requirement is a selector that contains values, a key, and an operator that + relates the key and values. + properties: + key: + description: key is the label key that the selector + applies to. + type: string + operator: + description: |- + operator represents a key's relationship to a set of values. + Valid operators are In, NotIn, Exists and DoesNotExist. + type: string + values: + description: |- + values is an array of string values. If the operator is In or NotIn, + the values array must be non-empty. If the operator is Exists or DoesNotExist, + the values array must be empty. This array is replaced during a strategic + merge patch. + items: + type: string + type: array + x-kubernetes-list-type: atomic + required: + - key + - operator + type: object + type: array + x-kubernetes-list-type: atomic + matchLabels: + additionalProperties: + type: string + description: |- + matchLabels is a map of {key,value} pairs. A single {key,value} in the matchLabels + map is equivalent to an element of matchExpressions, whose key field is "key", the + operator is "In", and the values array contains only "value". The requirements are ANDed. + type: object + type: object + x-kubernetes-map-type: atomic + x-kubernetes-validations: + - message: selector is immutable + rule: self == oldSelf + required: + - selector + type: object + x-kubernetes-validations: + - message: source is immutable + rule: self == oldSelf + volumeGroupReplicationClassName: + description: volumeGroupReplicationClassName is the volumeGroupReplicationClass + name for this VolumeGroupReplication resource + type: string + x-kubernetes-validations: + - message: volumeGroupReplicationClassName is immutable + rule: self == oldSelf + volumeGroupReplicationContentName: + description: Name of the VolumeGroupReplicationContent object created + for this volumeGroupReplication + type: string + x-kubernetes-validations: + - message: volumeGroupReplicationContentName is immutable + rule: self == oldSelf + volumeReplicationClassName: + description: |- + volumeReplicationClassName is the volumeReplicationClass name for the VolumeReplication object + created for this volumeGroupReplication + type: string + x-kubernetes-validations: + - message: volumeReplicationClassName is immutable + rule: self == oldSelf + volumeReplicationName: + description: Name of the VolumeReplication object created for this + volumeGroupReplication + type: string + x-kubernetes-validations: + - message: volumeReplicationName is immutable + rule: self == oldSelf + required: + - autoResync + - replicationState + - source + - volumeGroupReplicationClassName + type: object + status: + description: VolumeGroupReplicationStatus defines the observed state of + VolumeGroupReplication + properties: + conditions: + description: Conditions are the list of conditions and their status. + items: + description: Condition contains details for one aspect of the current + state of this API Resource. + properties: + lastTransitionTime: + description: |- + lastTransitionTime is the last time the condition transitioned from one status to another. + This should be when the underlying condition changed. If that is not known, then using the time when the API field changed is acceptable. + format: date-time + type: string + message: + description: |- + message is a human readable message indicating details about the transition. + This may be an empty string. + maxLength: 32768 + type: string + observedGeneration: + description: |- + observedGeneration represents the .metadata.generation that the condition was set based upon. + For instance, if .metadata.generation is currently 12, but the .status.conditions[x].observedGeneration is 9, the condition is out of date + with respect to the current state of the instance. + format: int64 + minimum: 0 + type: integer + reason: + description: |- + reason contains a programmatic identifier indicating the reason for the condition's last transition. + Producers of specific condition types may define expected values and meanings for this field, + and whether the values are considered a guaranteed API. + The value should be a CamelCase string. + This field may not be empty. + maxLength: 1024 + minLength: 1 + pattern: ^[A-Za-z]([A-Za-z0-9_,:]*[A-Za-z0-9_])?$ + type: string + status: + description: status of the condition, one of True, False, Unknown. + enum: + - "True" + - "False" + - Unknown + type: string + type: + description: type of condition in CamelCase or in foo.example.com/CamelCase. + maxLength: 316 + pattern: ^([a-z0-9]([-a-z0-9]*[a-z0-9])?(\.[a-z0-9]([-a-z0-9]*[a-z0-9])?)*/)?(([A-Za-z0-9][-A-Za-z0-9_.]*)?[A-Za-z0-9])$ + type: string + required: + - lastTransitionTime + - message + - reason + - status + - type + type: object + type: array + destinationVolumeID: + description: |- + DestinationVolumeID is the volume ID on the destination/target side. + This field is set when the SP reports different source and destination + volume IDs. + type: string + lastCompletionTime: + format: date-time + type: string + lastStartTime: + format: date-time + type: string + lastSyncBytes: + format: int64 + type: integer + lastSyncDuration: + type: string + lastSyncTime: + format: date-time + type: string + message: + type: string + observedGeneration: + description: observedGeneration is the last generation change the + operator has dealt with + format: int64 + type: integer + persistentVolumeClaimsRefList: + description: |- + PersistentVolumeClaimsRefList is the list of PVCs for the volume group replication. + The maximum number of allowed PVCs in the group is 100. + items: + description: |- + LocalObjectReference contains enough information to let you locate the + referenced object inside the same namespace. + properties: + name: + default: "" + description: |- + Name of the referent. + This field is effectively required, but due to backwards compatibility is + allowed to be empty. Instances of this type with an empty value here are + almost certainly wrong. + More info: https://kubernetes.io/docs/concepts/overview/working-with-objects/names/#names + type: string + type: object + x-kubernetes-map-type: atomic + type: array + state: + description: State captures the latest state of the replication operation. + type: string + type: object + type: object + served: true + storage: true + subresources: + status: {} +{{- end -}} diff --git a/helm-charts/charts/simplyblock-operator/templates/replication.storage.openshift.io_volumereplicationclasses.yaml b/helm-charts/charts/simplyblock-operator/templates/replication.storage.openshift.io_volumereplicationclasses.yaml new file mode 100644 index 000000000..b434380df --- /dev/null +++ b/helm-charts/charts/simplyblock-operator/templates/replication.storage.openshift.io_volumereplicationclasses.yaml @@ -0,0 +1,96 @@ +# VolumeReplicationClass CRD, vendored verbatim from csi-addons/kubernetes-csi-addons +# v0.15.0 (deploy/controller/crds.yaml), plus the chart's resource-policy +# annotation. The stock csi-addons controller-manager starts a controller +# per kind unconditionally, so every CRD it watches must exist for the +# manager to come up. This file installs one of them unless the cluster +# already serves the API. +{{- if and .Values.csiaddons.create (not (.Capabilities.APIVersions.Has "replication.storage.openshift.io/v1alpha1/VolumeReplicationClass")) -}} +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition +metadata: + annotations: + helm.sh/resource-policy: keep + controller-gen.kubebuilder.io/version: v0.20.1 + name: volumereplicationclasses.replication.storage.openshift.io +spec: + group: replication.storage.openshift.io + names: + kind: VolumeReplicationClass + listKind: VolumeReplicationClassList + plural: volumereplicationclasses + shortNames: + - vrc + singular: volumereplicationclass + scope: Cluster + versions: + - additionalPrinterColumns: + - jsonPath: .spec.provisioner + name: provisioner + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha1 + schema: + openAPIV3Schema: + description: VolumeReplicationClass is the Schema for the volumereplicationclasses + API. + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: |- + VolumeReplicationClassSpec specifies parameters that an underlying storage system uses + when creating a volume replica. A specific VolumeReplicationClass is used by specifying + its name in a VolumeReplication object. + properties: + parameters: + additionalProperties: + type: string + description: |- + Parameters is a key-value map with storage provisioner specific configurations for + creating volume replicas + type: object + x-kubernetes-validations: + - message: parameters are immutable + rule: self == oldSelf + provisioner: + description: Provisioner is the name of storage provisioner + type: string + x-kubernetes-validations: + - message: provisioner is immutable + rule: self == oldSelf + required: + - provisioner + type: object + x-kubernetes-validations: + - message: parameters are immutable + rule: has(self.parameters) == has(oldSelf.parameters) + status: + description: VolumeReplicationClassStatus defines the observed state of + VolumeReplicationClass. + type: object + required: + - spec + type: object + served: true + storage: true + subresources: + status: {} +{{- end -}} diff --git a/helm-charts/charts/simplyblock-operator/templates/replication.storage.openshift.io_volumereplications.yaml b/helm-charts/charts/simplyblock-operator/templates/replication.storage.openshift.io_volumereplications.yaml new file mode 100644 index 000000000..3deaca0e5 --- /dev/null +++ b/helm-charts/charts/simplyblock-operator/templates/replication.storage.openshift.io_volumereplications.yaml @@ -0,0 +1,225 @@ +# VolumeReplication CRD, vendored verbatim from csi-addons/kubernetes-csi-addons +# v0.15.0 (deploy/controller/crds.yaml), plus the chart's resource-policy +# annotation. The stock csi-addons controller-manager starts a controller +# per kind unconditionally, so every CRD it watches must exist for the +# manager to come up. This file installs one of them unless the cluster +# already serves the API. +{{- if and .Values.csiaddons.create (not (.Capabilities.APIVersions.Has "replication.storage.openshift.io/v1alpha1/VolumeReplication")) -}} +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition +metadata: + annotations: + helm.sh/resource-policy: keep + controller-gen.kubebuilder.io/version: v0.20.1 + name: volumereplications.replication.storage.openshift.io +spec: + group: replication.storage.openshift.io + names: + kind: VolumeReplication + listKind: VolumeReplicationList + plural: volumereplications + shortNames: + - vr + singular: volumereplication + scope: Namespaced + versions: + - additionalPrinterColumns: + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + - jsonPath: .spec.volumeReplicationClass + name: volumeReplicationClass + type: string + - jsonPath: .spec.dataSource.kind + name: SourceKind + type: string + - jsonPath: .spec.dataSource.name + name: SourceName + type: string + - jsonPath: .spec.replicationState + name: desiredState + type: string + - jsonPath: .status.state + name: currentState + type: string + name: v1alpha1 + schema: + openAPIV3Schema: + description: VolumeReplication is the Schema for the volumereplications API. + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: VolumeReplicationSpec defines the desired state of VolumeReplication. + properties: + autoResync: + default: false + description: |- + AutoResync represents the volume to be auto resynced when + ReplicationState is "secondary" + type: boolean + dataSource: + description: DataSource represents the object associated with the + volume + properties: + apiGroup: + description: |- + APIGroup is the group for the resource being referenced. + If APIGroup is not specified, the specified Kind must be in the core API group. + For any other third-party types, APIGroup is required. + type: string + kind: + description: Kind is the type of resource being referenced + type: string + name: + description: Name is the name of resource being referenced + type: string + required: + - kind + - name + type: object + x-kubernetes-map-type: atomic + x-kubernetes-validations: + - message: dataSource is immutable + rule: self == oldSelf + replicationHandle: + description: replicationHandle represents an existing (but new) replication + id + type: string + replicationState: + description: |- + ReplicationState represents the replication operation to be performed on the volume. + Supported operations are "primary", "secondary" and "resync" + enum: + - primary + - secondary + - resync + type: string + volumeReplicationClass: + description: VolumeReplicationClass is the VolumeReplicationClass + name for this VolumeReplication resource + type: string + x-kubernetes-validations: + - message: volumeReplicationClass is immutable + rule: self == oldSelf + required: + - autoResync + - dataSource + - replicationState + - volumeReplicationClass + type: object + status: + description: VolumeReplicationStatus defines the observed state of VolumeReplication. + properties: + conditions: + description: Conditions are the list of conditions and their status. + items: + description: Condition contains details for one aspect of the current + state of this API Resource. + properties: + lastTransitionTime: + description: |- + lastTransitionTime is the last time the condition transitioned from one status to another. + This should be when the underlying condition changed. If that is not known, then using the time when the API field changed is acceptable. + format: date-time + type: string + message: + description: |- + message is a human readable message indicating details about the transition. + This may be an empty string. + maxLength: 32768 + type: string + observedGeneration: + description: |- + observedGeneration represents the .metadata.generation that the condition was set based upon. + For instance, if .metadata.generation is currently 12, but the .status.conditions[x].observedGeneration is 9, the condition is out of date + with respect to the current state of the instance. + format: int64 + minimum: 0 + type: integer + reason: + description: |- + reason contains a programmatic identifier indicating the reason for the condition's last transition. + Producers of specific condition types may define expected values and meanings for this field, + and whether the values are considered a guaranteed API. + The value should be a CamelCase string. + This field may not be empty. + maxLength: 1024 + minLength: 1 + pattern: ^[A-Za-z]([A-Za-z0-9_,:]*[A-Za-z0-9_])?$ + type: string + status: + description: status of the condition, one of True, False, Unknown. + enum: + - "True" + - "False" + - Unknown + type: string + type: + description: type of condition in CamelCase or in foo.example.com/CamelCase. + maxLength: 316 + pattern: ^([a-z0-9]([-a-z0-9]*[a-z0-9])?(\.[a-z0-9]([-a-z0-9]*[a-z0-9])?)*/)?(([A-Za-z0-9][-A-Za-z0-9_.]*)?[A-Za-z0-9])$ + type: string + required: + - lastTransitionTime + - message + - reason + - status + - type + type: object + type: array + destinationVolumeID: + description: |- + DestinationVolumeID is the volume ID on the destination/target side. + This field is set when the SP reports different source and destination + volume IDs. + type: string + lastCompletionTime: + format: date-time + type: string + lastStartTime: + format: date-time + type: string + lastSyncBytes: + format: int64 + type: integer + lastSyncDuration: + type: string + lastSyncTime: + format: date-time + type: string + message: + type: string + observedGeneration: + description: observedGeneration is the last generation change the + operator has dealt with + format: int64 + type: integer + state: + description: State captures the latest state of the replication operation. + type: string + type: object + required: + - spec + type: object + served: true + storage: true + subresources: + status: {} +{{- end -}} diff --git a/helm-charts/charts/simplyblock-operator/templates/setup-csi-addons-controller.yaml b/helm-charts/charts/simplyblock-operator/templates/setup-csi-addons-controller.yaml new file mode 100644 index 000000000..6ca2b2e7b --- /dev/null +++ b/helm-charts/charts/simplyblock-operator/templates/setup-csi-addons-controller.yaml @@ -0,0 +1,78 @@ +# The stock kubernetes-csi-addons controller-manager (design +# design-csi-addons-replication.md §4.1, prerequisite P0-5): it discovers +# driver endpoints through CSIAddonsNode objects and reconciles +# VolumeReplication by calling the driver's Replication gRPC through the +# csi-addons sidecar. Vendored from csi-addons/kubernetes-csi-addons v0.15.0 +# (deploy/controller/setup-controller.yaml) with the image pinned instead of +# :latest, the namespace bound to the release, and the legacy +# ControllerManagerConfig ConfigMap dropped (the manager neither mounts nor +# reads it). Deployed into the release namespace rather than kube-system: this +# is a chart component, not base cluster functionality. +{{- if .Values.csiaddons.create -}} +--- +apiVersion: apps/v1 +kind: Deployment +metadata: + labels: + app.kubernetes.io/name: csi-addons + name: simplyblock-csi-addons-controller-manager + namespace: {{ .Release.Namespace }} +spec: + replicas: 1 + selector: + matchLabels: + app.kubernetes.io/name: csi-addons + template: + metadata: + annotations: + kubectl.kubernetes.io/default-container: manager + labels: + app.kubernetes.io/name: csi-addons + control-plane: controller-manager + spec: + containers: + - name: manager + image: "{{ .Values.image.csiAddonsController.repository }}:{{ .Values.image.csiAddonsController.tag }}" + imagePullPolicy: {{ .Values.image.csiAddonsController.pullPolicy }} + command: + - /csi-addons-manager + args: + - --namespace=$(POD_NAMESPACE) + - --leader-elect + - --automaxprocs + env: + - name: POD_NAMESPACE + valueFrom: + fieldRef: + fieldPath: metadata.namespace + livenessProbe: + httpGet: + path: /healthz + port: 8081 + initialDelaySeconds: 15 + periodSeconds: 20 + readinessProbe: + httpGet: + path: /readyz + port: 8081 + initialDelaySeconds: 5 + periodSeconds: 10 + ports: + - containerPort: 8443 + name: metrics + protocol: TCP + resources: + limits: + cpu: 1000m + memory: 512Mi + requests: + cpu: 10m + memory: 64Mi + securityContext: + allowPrivilegeEscalation: false + readOnlyRootFilesystem: true + securityContext: + runAsNonRoot: true + serviceAccountName: simplyblock-csi-addons-controller-manager + terminationGracePeriodSeconds: 10 +{{- end -}} diff --git a/helm-charts/charts/simplyblock-operator/values.yaml b/helm-charts/charts/simplyblock-operator/values.yaml index f261515c1..2da22d083 100644 --- a/helm-charts/charts/simplyblock-operator/values.yaml +++ b/helm-charts/charts/simplyblock-operator/values.yaml @@ -7,6 +7,14 @@ image: repository: quay.io/simplyblock-io/snapshot-controller tag: v8.2.0 pullPolicy: Always + csiAddonsController: + # Upstream registry until the image is mirrored to quay.io/simplyblock-io, + # which is the convention every other component follows. Keep the tag in + # lockstep with the vendored CRDs and RBAC (templates/*csi-addons*, + # templates/csiaddons.openshift.io_*, templates/replication.storage.*). + repository: quay.io/csiaddons/k8s-controller + tag: v0.15.0 + pullPolicy: Always simplyblock: repository: quay.io/simplyblock-io/simplyblock tag: "26.2.6-PRE" @@ -40,6 +48,13 @@ rbac: snapshotcontroller: create: true +csiaddons: + # Deploy the kubernetes-csi-addons machinery: the CRDs (VolumeReplication + # and friends) and the stock controller-manager. Off until the driver's + # csi-addons Replication service ships (design-csi-addons-replication.md + # Phase 1). The vendored manifests are the design's prerequisite P0-5. + create: false + externallyManagedSecret: # Specifies whether a externallyManagedSecret should be created create: true diff --git a/operator/docs/designs/design-csi-addons-replication.md b/operator/docs/designs/design-csi-addons-replication.md index 62348a08a..28f248013 100644 --- a/operator/docs/designs/design-csi-addons-replication.md +++ b/operator/docs/designs/design-csi-addons-replication.md @@ -22,14 +22,14 @@ Phase 1 is independently useful: a `VolumeReplication` object per PVC whose stat ## Phase 0 — External Prerequisites -| # | Prerequisite | Kind | Blocks | Status | -|------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------|---------|----------------------------------| -| P0-1 | A typed, steady-state per-volume replication status read: `GET .../volumes/{id}/replication/status` serving what `lvol_controller.get_replication_info` computes today (state, lag, outstanding bytes, failure counters), available for the volume's whole replicated life | Control plane (`sbcli`) | Phase 1 | Not shipped | -| P0-2 | Idempotent attach and detach: attaching a volume to the policy it already follows returns success, and detaching a non-attached volume returns success | Control plane (`sbcli`) | Phase 1 | Not shipped | -| P0-3 | A standalone demote verb: `POST .../volumes/{id}/replication/demote` that converges the peer while still serving (repeated snapshot-and-ship until the remaining delta is small), then quiesces, ships the final delta, confirms it landed on the peer, and fences the data path | Control plane (`sbcli`) | Phase 2 | Not shipped | -| P0-4 | An `rpo_target_seconds` field on `ReplicationPolicy`, so RPO compliance is computable against a declared target rather than the derived lag budget | Control plane (`sbcli`) | Phase 3 | Not shipped | -| P0-5 | csi-addons upstream: the `VolumeReplication` and `VolumeReplicationClass` CRDs (`replication.storage.openshift.io/v1alpha1`), the kubernetes-csi-addons controller-manager image, and the csi-addons sidecar image | Ecosystem | Phase 1 | Available upstream; not vendored | -| P0-6 | A latest-replicated-snapshot read: per volume, and per consistency group as one complete generation, the newest fully replicated snapshot on the secondary addressed as a cloneable object | Control plane (`sbcli`) | Phase 4 | Not shipped | +| # | Prerequisite | Kind | Blocks | Status | +|------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------|---------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| P0-1 | A typed, steady-state per-volume replication status read: `GET .../volumes/{id}/replication/status` serving what `lvol_controller.get_replication_info` computes today (state, lag, outstanding bytes, failure counters), available for the volume's whole replicated life | Control plane (`sbcli`) | Phase 1 | Not shipped | +| P0-2 | Idempotent attach and detach: attaching a volume to the policy it already follows returns success, and detaching a non-attached volume returns success | Control plane (`sbcli`) | Phase 1 | Not shipped | +| P0-3 | A standalone demote verb: `POST .../volumes/{id}/replication/demote` that converges the peer while still serving (repeated snapshot-and-ship until the remaining delta is small), then quiesces, ships the final delta, confirms it landed on the peer, and fences the data path | Control plane (`sbcli`) | Phase 2 | Not shipped | +| P0-4 | An `rpo_target_seconds` field on `ReplicationPolicy`, so RPO compliance is computable against a declared target rather than the derived lag budget | Control plane (`sbcli`) | Phase 3 | Not shipped | +| P0-5 | csi-addons upstream: the `VolumeReplication` and `VolumeReplicationClass` CRDs (`replication.storage.openshift.io/v1alpha1`), the kubernetes-csi-addons controller-manager image, and the csi-addons sidecar image | Ecosystem | Phase 1 | Vendored in the chart at v0.15.0 behind `csiaddons.create` (all twelve upstream CRDs, since the stock manager starts a controller per kind); sidecar wiring is Phase 1 | +| P0-6 | A latest-replicated-snapshot read: per volume, and per consistency group as one complete generation, the newest fully replicated snapshot on the secondary addressed as a cloneable object | Control plane (`sbcli`) | Phase 4 | Not shipped | Everything else the adapter needs already exists: the attach and detach calls, failover, the failback and commit pair, the relationship read, and the backlog arithmetic inside `get_replication_info`. The adapter is thin precisely because the engine is complete. What is missing is the shape Ramen can drive. diff --git a/operator/docs/tests/test-plan-csi-addons-replication.md b/operator/docs/tests/test-plan-csi-addons-replication.md index a9ed49411..19bbec5fd 100644 --- a/operator/docs/tests/test-plan-csi-addons-replication.md +++ b/operator/docs/tests/test-plan-csi-addons-replication.md @@ -166,10 +166,10 @@ Every scenario is uncovered because the design is Draft. The counts are the targ ## 7. What Is Not Yet Covered -| # | Gap | Reason | -|-------------|-------------------------------------------------------------------------------------------------------------------|-----------------------------------------------------------------------------------------------------------------| -| U-01 … U-27 | The verb adapter, condition derivation, preflight, and coexistence units | The adapter does not exist; blocked on P0-1 and P0-2 for the Phase 1 verbs and P0-3 for demote | -| I-01 … I-07 | The sidecar and controller-manager loop | Blocked on the vendored CRDs and images (P0-5) | -| E-01 … E-07 | The live lifecycle and the Ramen gate | Blocked on Phase 1 and 2 landing, plus a two-cluster test bed with Ramen dr-cluster installed for E-06 and E-07 | -| — | Repeated resync, class drift after verification, annotated-volume migration onto the adapter, cascaded topologies | Beyond the first coverage pass, recorded so the gaps are explicit rather than assumed covered | -| M-01, M-02 | Demote under writes; concurrent ownership race | Need failure injection and precise timing a live two-cluster run does not automate yet | +| # | Gap | Reason | +|-------------|-------------------------------------------------------------------------------------------------------------------|---------------------------------------------------------------------------------------------------------------------------------------------| +| U-01 … U-27 | The verb adapter, condition derivation, preflight, and coexistence units | The adapter does not exist; blocked on P0-1 and P0-2 for the Phase 1 verbs and P0-3 for demote | +| I-01 … I-07 | The sidecar and controller-manager loop | CRDs and controller-manager vendored in the chart (P0-5, `csiaddons.create`); blocked on the Phase 1 sidecar and driver Replication service | +| E-01 … E-07 | The live lifecycle and the Ramen gate | Blocked on Phase 1 and 2 landing, plus a two-cluster test bed with Ramen dr-cluster installed for E-06 and E-07 | +| — | Repeated resync, class drift after verification, annotated-volume migration onto the adapter, cascaded topologies | Beyond the first coverage pass, recorded so the gaps are explicit rather than assumed covered | +| M-01, M-02 | Demote under writes; concurrent ownership race | Need failure injection and precise timing a live two-cluster run does not automate yet | From 00b78d226459fc91095e01724c1db1e0d4cad406 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 15:51:22 +0200 Subject: [PATCH 054/206] feat(metrics): the three unmeasured subsystems publish what they do MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit design-clusterdeploymentconfig §9.2 names nine metrics and none existed; design-simplyblockdriver §6.2 names four and none existed. Nothing measured what a deployment costs, how often the approval gate refuses, or whether the CSI driver is up, and a phase somebody reads in kubectl is not an alert. The two the deployment design calls out are a pair, and they are built as one. validation_failures_total by reason is the only signal that says whether the approval gate is working as review or as an obstacle: a deployment where every draft fails on DeviceNotFound is a discovery bug rather than a careful reviewer. approval_rejections_total is what it is read against, because the two count the same mistakes at different moments and a rejection whose reason draft validation never reported first is a gap in the validation. For that comparison to work the two need one vocabulary, so the webhook now carries a reason beside each refusal rather than a sentence alone, and four of the seven are the event reasons the controller already raises. The other three exist only here: an admission rejection fails the request, so there is no object to record an event against and the counter is the whole of that path's account of itself. A counter that rose on every reconcile would count passes. A draft nobody fixes is re-validated every thirty seconds forever, so the validation counter rises only when the outcome changes, which is the same rule AwaitingApproval already follows for its event. ClusterDeploymentConfig gained one status field. §9.2 measures an expansion from approval, a document may sit as a draft for as long as a review takes, and creationTimestamp is therefore the wrong instant; no step deadline survives its step, and nothing in memory survives a restart. status.expansionStartedAt is stamped where the machine is born, which is the first reconcile after somebody said yes. An expansion with no recorded start is not timed at all, because a duration measured from an instant nobody wrote down is a sample the histogram cannot tell from a real one. The driver's four publish from setHealth, which is the pass that measured them. The paths that write a phase without a reading leave the last numbers standing rather than replacing them with zeros they did not observe, and the two node counts go out together because neither means anything alone: three ready plugins is healthy on a three-worker cluster and an outage on a thirty-worker one. version_info is 0, and the TODO beside it says why. status.version is written by nothing, and the two things that would fill it are outside this repository. An info gauge is 1 for a fact it carries in its labels, so 0 is the honest reading for a fact nobody has established, and the series becomes correct on the day the field is filled in with no change here. It also replaces its series rather than adding to one: a driver that was upgraded would otherwise leave the version it used to run sitting at 1 beside the version it runs now. Last, §7.12's naming rule is now checked rather than remembered. simplyblock_storagecluster_rolling_restart_node_total was a gauge, and total means a counter, which is what keeps a rate honest. It is _count now, and the design's table gained the row it never carried — the metric existed and the specification did not list it, which is how it drifted. The rename is the smaller half. internal/controllers/names_test.go reads every Prometheus option literal in the module with go/ast and holds it against the six aggregation suffixes and the three the rule binds to a type, plus a second case that no name is declared twice. It found exactly that one violation across sixty-seven declarations. Reading the source rather than a registry covers a metric in a package nothing imports, and a computed name fails the test rather than being skipped. The rebalancer and fio families are exempt by prefix with the reason recorded: they predate the rule and are renamed when that subsystem moves. Co-Authored-By: Claude Opus 5 (1M context) --- ...mplyblock.io_clusterdeploymentconfigs.yaml | 13 + .../v1alpha2/clusterdeploymentconfig_types.go | 12 + .../api/v1alpha2/zz_generated.deepcopy.go | 4 + ...mplyblock.io_clusterdeploymentconfigs.yaml | 13 + operator/dist/install.yaml | 13 + .../design-clusterdeploymentconfig.md | 12 + .../crd-redesign/design-storagecluster.md | 1 + .../internal/controllers/cluster/metrics.go | 8 +- .../controllers/cluster/rollingrestart.go | 2 +- .../cluster/storageclusterops_controller.go | 2 +- .../clusterdeploymentconfig_controller.go | 61 +++- .../controllers/deployment/expansion.go | 4 + .../controllers/deployment/metrics.go | 264 ++++++++++++++++++ .../controllers/deployment/metrics_test.go | 226 +++++++++++++++ .../deployment/operatorops_controller.go | 5 + .../internal/controllers/driver/metrics.go | 101 +++++++ .../controllers/driver/metrics_test.go | 151 ++++++++++ .../driver/simplyblockdriver_controller.go | 18 +- operator/internal/controllers/names_test.go | 229 +++++++++++++++ operator/internal/discovery/plan.go | 16 +- ...mplyblock.io_clusterdeploymentconfigs.yaml | 13 + .../clusterdeploymentconfig_metrics_test.go | 131 +++++++++ .../clusterdeploymentconfig_validator.go | 113 +++++--- 23 files changed, 1354 insertions(+), 58 deletions(-) create mode 100644 operator/internal/controllers/deployment/metrics.go create mode 100644 operator/internal/controllers/deployment/metrics_test.go create mode 100644 operator/internal/controllers/driver/metrics.go create mode 100644 operator/internal/controllers/driver/metrics_test.go create mode 100644 operator/internal/controllers/names_test.go create mode 100644 operator/internal/webhook/clusterdeploymentconfig_metrics_test.go diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml index aed657d41..7f68f56a5 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml @@ -399,6 +399,19 @@ spec: is a record rather than a dependency: nothing resolves it after expansion, which is what makes the document safe to delete. type: string + expansionStartedAt: + description: |- + ExpansionStartedAt is when the expansion machine was born, which is the + first reconcile after the document was approved. A document may sit as a + draft for as long as a review takes, so this is not creationTimestamp and + the difference is the whole point: how long a deployment takes is measured + from the moment somebody said yes. + + It is the start of §9.2's expansion_duration_seconds. A histogram needs an + instant that survives the operator restarting mid-expansion, which nothing + in memory and no step deadline supplies. + format: date-time + type: string message: description: |- Message is the reason the phase is what it is: one sentence, replaced as diff --git a/operator/api/v1alpha2/clusterdeploymentconfig_types.go b/operator/api/v1alpha2/clusterdeploymentconfig_types.go index 2dd28f471..3591813fa 100644 --- a/operator/api/v1alpha2/clusterdeploymentconfig_types.go +++ b/operator/api/v1alpha2/clusterdeploymentconfig_types.go @@ -393,6 +393,18 @@ type ClusterDeploymentConfigStatus struct { // from, so a stale status can be told from a current one. // +optional ObservedGeneration int64 `json:"observedGeneration,omitempty"` + + // ExpansionStartedAt is when the expansion machine was born, which is the + // first reconcile after the document was approved. A document may sit as a + // draft for as long as a review takes, so this is not creationTimestamp and + // the difference is the whole point: how long a deployment takes is measured + // from the moment somebody said yes. + // + // It is the start of §9.2's expansion_duration_seconds. A histogram needs an + // instant that survives the operator restarting mid-expansion, which nothing + // in memory and no step deadline supplies. + // +optional + ExpansionStartedAt *metav1.Time `json:"expansionStartedAt,omitempty"` } // +kubebuilder:object:root=true diff --git a/operator/api/v1alpha2/zz_generated.deepcopy.go b/operator/api/v1alpha2/zz_generated.deepcopy.go index d460d2ee0..b27c07325 100644 --- a/operator/api/v1alpha2/zz_generated.deepcopy.go +++ b/operator/api/v1alpha2/zz_generated.deepcopy.go @@ -259,6 +259,10 @@ func (in *ClusterDeploymentConfigStatus) DeepCopyInto(out *ClusterDeploymentConf *out = make([]string, len(*in)) copy(*out, *in) } + if in.ExpansionStartedAt != nil { + in, out := &in.ExpansionStartedAt, &out.ExpansionStartedAt + *out = (*in).DeepCopy() + } } // DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new ClusterDeploymentConfigStatus. diff --git a/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml b/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml index aed657d41..7f68f56a5 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml @@ -399,6 +399,19 @@ spec: is a record rather than a dependency: nothing resolves it after expansion, which is what makes the document safe to delete. type: string + expansionStartedAt: + description: |- + ExpansionStartedAt is when the expansion machine was born, which is the + first reconcile after the document was approved. A document may sit as a + draft for as long as a review takes, so this is not creationTimestamp and + the difference is the whole point: how long a deployment takes is measured + from the moment somebody said yes. + + It is the start of §9.2's expansion_duration_seconds. A histogram needs an + instant that survives the operator restarting mid-expansion, which nothing + in memory and no step deadline supplies. + format: date-time + type: string message: description: |- Message is the reason the phase is what it is: one sentence, replaced as diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index 669dcc68f..243cc6ee0 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -1072,6 +1072,19 @@ spec: is a record rather than a dependency: nothing resolves it after expansion, which is what makes the document safe to delete. type: string + expansionStartedAt: + description: |- + ExpansionStartedAt is when the expansion machine was born, which is the + first reconcile after the document was approved. A document may sit as a + draft for as long as a review takes, so this is not creationTimestamp and + the difference is the whole point: how long a deployment takes is measured + from the moment somebody said yes. + + It is the start of §9.2's expansion_duration_seconds. A histogram needs an + instant that survives the operator restarting mid-expansion, which nothing + in memory and no step deadline supplies. + format: date-time + type: string message: description: |- Message is the reason the phase is what it is: one sentence, replaced as diff --git a/operator/docs/designs/crd-redesign/design-clusterdeploymentconfig.md b/operator/docs/designs/crd-redesign/design-clusterdeploymentconfig.md index 5d177576a..da0523140 100644 --- a/operator/docs/designs/crd-redesign/design-clusterdeploymentconfig.md +++ b/operator/docs/designs/crd-redesign/design-clusterdeploymentconfig.md @@ -1160,6 +1160,18 @@ type ClusterDeploymentConfigStatus struct { // from, so a stale status can be told from a current one. // +optional ObservedGeneration int64 `json:"observedGeneration,omitempty"` + + // ExpansionStartedAt is when the expansion machine was born, which is the + // first reconcile after the document was approved. A document may sit as a + // draft for as long as a review takes, so this is not creationTimestamp and + // the difference is the whole point: how long a deployment takes is measured + // from the moment somebody said yes. + // + // It is the start of §9.2's expansion_duration_seconds. A histogram needs an + // instant that survives the operator restarting mid-expansion, which nothing + // in memory and no step deadline supplies. + // +optional + ExpansionStartedAt *metav1.Time `json:"expansionStartedAt,omitempty"` } // +kubebuilder:object:root=true diff --git a/operator/docs/designs/crd-redesign/design-storagecluster.md b/operator/docs/designs/crd-redesign/design-storagecluster.md index 6d6923b75..af797ea15 100644 --- a/operator/docs/designs/crd-redesign/design-storagecluster.md +++ b/operator/docs/designs/crd-redesign/design-storagecluster.md @@ -1338,6 +1338,7 @@ behavior, and without an event it is indistinguishable from a stalled controller | `simplyblock_storagecluster_operation_active_state` | `cluster` | Gauge, 1 while `status.activeOpsRef` is set, so a lock held by a finished operation is visible | | `simplyblock_storagecluster_rolling_restart_peer_hold_seconds` | `cluster` | Histogram of time the rolling restart held for a peer node to come back online | | `simplyblock_storagecluster_rolling_restart_node_index_count` | `cluster` | Gauge of `nodeIndex`, against the `nodes` length, so walk progress is graphable | +| `simplyblock_storagecluster_rolling_restart_node_count` | `cluster` | Gauge of how many nodes the running walk covers, which is the length the index is read against | | `simplyblock_storagecluster_phase_state` | `cluster`, `phase` | Gauge, 1 for the cluster's current phase (§4.2), so a cluster stuck in `Creating` is alertable | Every metric carries `cluster`, matching the rebalancer's existing convention, so diff --git a/operator/internal/controllers/cluster/metrics.go b/operator/internal/controllers/cluster/metrics.go index d5c35ad73..0aac1fb1a 100644 --- a/operator/internal/controllers/cluster/metrics.go +++ b/operator/internal/controllers/cluster/metrics.go @@ -98,14 +98,14 @@ var ( rollingRestartNodeIndex = prometheus.NewGaugeVec( prometheus.GaugeOpts{ Name: "simplyblock_storagecluster_rolling_restart_node_index_count", - Help: "The walk's position in its node list, against rolling_restart_node_total, so progress is graphable.", + Help: "The walk's position in its node list, against rolling_restart_node_count, so progress is graphable.", }, []string{"cluster"}, ) - rollingRestartNodeTotal = prometheus.NewGaugeVec( + rollingRestartNodeCount = prometheus.NewGaugeVec( prometheus.GaugeOpts{ - Name: "simplyblock_storagecluster_rolling_restart_node_total", + Name: "simplyblock_storagecluster_rolling_restart_node_count", Help: "How many nodes the running walk covers. Absent when no rolling restart is in flight.", }, []string{"cluster"}, @@ -130,7 +130,7 @@ func init() { operationActiveState, rollingRestartPeerHoldSeconds, rollingRestartNodeIndex, - rollingRestartNodeTotal, + rollingRestartNodeCount, clusterPhaseState, ) } diff --git a/operator/internal/controllers/cluster/rollingrestart.go b/operator/internal/controllers/cluster/rollingrestart.go index 9a78bf6af..ba27374e9 100644 --- a/operator/internal/controllers/cluster/rollingrestart.go +++ b/operator/internal/controllers/cluster/rollingrestart.go @@ -112,7 +112,7 @@ func (r *StorageClusterOpsReconciler) planWalk( for _, node := range nodes { planned = append(planned, node.UUID) } - rollingRestartNodeTotal.WithLabelValues(ops.Spec.ClusterRef).Set(float64(len(planned))) + rollingRestartNodeCount.WithLabelValues(ops.Spec.ClusterRef).Set(float64(len(planned))) rollingRestartNodeIndex.WithLabelValues(ops.Spec.ClusterRef).Set(0) return r.writeStatus(ctx, ops, func(status *simplyblockv1alpha2.StorageClusterOpsStatus) { diff --git a/operator/internal/controllers/cluster/storageclusterops_controller.go b/operator/internal/controllers/cluster/storageclusterops_controller.go index 4f0ab00f8..172deffed 100644 --- a/operator/internal/controllers/cluster/storageclusterops_controller.go +++ b/operator/internal/controllers/cluster/storageclusterops_controller.go @@ -452,7 +452,7 @@ func (r *StorageClusterOpsReconciler) observeOperation( Observe(time.Since(started.Time).Seconds()) } rollingRestartNodeIndex.DeleteLabelValues(cluster) - rollingRestartNodeTotal.DeleteLabelValues(cluster) + rollingRestartNodeCount.DeleteLabelValues(cluster) } // observeStep records how long one step took. The start is the step's entry, diff --git a/operator/internal/controllers/deployment/clusterdeploymentconfig_controller.go b/operator/internal/controllers/deployment/clusterdeploymentconfig_controller.go index 04ba12502..7164f6a9e 100644 --- a/operator/internal/controllers/deployment/clusterdeploymentconfig_controller.go +++ b/operator/internal/controllers/deployment/clusterdeploymentconfig_controller.go @@ -113,8 +113,10 @@ func (r *ClusterDeploymentConfigReconciler) Reconcile( } // A deleted document needs nothing done to it. There is no finalizer, because - // it owns nothing and nothing reads it (§4.3). + // it owns nothing and nothing reads it (§4.3). Its gauges go, because a phase + // series keyed on a name outlives the object the name belonged to. if !config.DeletionTimestamp.IsZero() { + forgetConfig(&config) return ctrl.Result{}, nil } @@ -139,6 +141,8 @@ func (r *ClusterDeploymentConfigReconciler) Reconcile( // no longer be edited away, and holding forever would say less than failing. if len(findings) > 0 { r.emitFindings(&config, findings) + // Counted unconditionally: this path is terminal, so it happens once. + countValidationFailures(config.Namespace, findings) return r.fail(ctx, &config, findings[0].message) } @@ -178,7 +182,7 @@ func (r *ClusterDeploymentConfigReconciler) expand( if config.Status.Step.State == "" { deadline := metav1.NewTime(time.Now().Add(validatingDeadline)) return ctrl.Result{RequeueAfter: configAdvance}, - r.recordStep(ctx, config, machine.CurrentState(), &deadline) + r.beginExpansion(ctx, config, machine.CurrentState(), &deadline) } current := machine.CurrentState() @@ -262,9 +266,17 @@ func (r *ClusterDeploymentConfigReconciler) holdAsDraft( findings []finding, ) (ctrl.Result, error) { if len(findings) > 0 { + summary := summarize(findings) r.emitFindings(config, findings) + // A draft nobody has fixed is validated again every pass, and the counter + // is about the outcomes a reviewer hits rather than about how long one has + // stood. The message is what last went out, so a summary that has not + // changed is the same finding reported again. + if config.Status.Message != summary { + countValidationFailures(config.Namespace, findings) + } return ctrl.Result{RequeueAfter: configRetry}, r.note(ctx, config, - simplyblockv1alpha2.ClusterDeploymentConfigPhaseDraft, summarize(findings)) + simplyblockv1alpha2.ClusterDeploymentConfigPhaseDraft, summary) } // AwaitingApproval is emitted on the transition to a validated draft rather @@ -342,6 +354,7 @@ func (r *ClusterDeploymentConfigReconciler) succeed( message := fmt.Sprintf("expanded into cluster %s and %d node(s)", config.Status.ClusterRef, len(config.Status.NodeRefs)) r.emit(config, corev1.EventTypeNormal, NodesCreated, message) + observeExpansion(config) return r.note(ctx, config, simplyblockv1alpha2.ClusterDeploymentConfigPhaseExpanded, message) } @@ -368,6 +381,32 @@ func (r *ClusterDeploymentConfigReconciler) note( }) } +// beginExpansion records the machine's first step together with the instant the +// expansion started, which is this pass: the machine is born on the first +// reconcile after somebody approved the document. +// +// The instant is persisted rather than kept in memory because the expansion +// outlives a single reconcile and can outlive the process, and a duration +// measured from a start the operator forgot is not a measurement. +func (r *ClusterDeploymentConfigReconciler) beginExpansion( + ctx context.Context, + config *simplyblockv1alpha2.ClusterDeploymentConfig, + initial configStep, + deadline *metav1.Time, +) error { + started := metav1.Now() + return r.writeStatus(ctx, config, + func(status *simplyblockv1alpha2.ClusterDeploymentConfigStatus) { + status.Phase = simplyblockv1alpha2.ClusterDeploymentConfigPhaseExpanding + status.Step = statemachine.KubeSnapshot{ + State: string(initial), Deadline: deadline, + } + if status.ExpansionStartedAt == nil { + status.ExpansionStartedAt = &started + } + }) +} + // recordStep persists the step the machine is about to be in, with the instant it // expires. Both travel together, because a step persisted without its deadline // restores as a step that can never time out. @@ -392,7 +431,7 @@ func (r *ClusterDeploymentConfigReconciler) writeStatus( config *simplyblockv1alpha2.ClusterDeploymentConfig, mutate func(*simplyblockv1alpha2.ClusterDeploymentConfigStatus), ) error { - return retry.RetryOnConflict(retry.DefaultRetry, func() error { + err := retry.RetryOnConflict(retry.DefaultRetry, func() error { var fresh simplyblockv1alpha2.ClusterDeploymentConfig if err := r.Get(ctx, client.ObjectKeyFromObject(config), &fresh); err != nil { return err @@ -418,6 +457,14 @@ func (r *ClusterDeploymentConfigReconciler) writeStatus( config.ResourceVersion = fresh.ResourceVersion return nil }) + if err != nil { + return err + } + // Published here rather than at each caller, because every phase this + // document reaches is written through this one function and a gauge that + // missed one would report the phase before it indefinitely. + observeConfigPhase(config) + return nil } // emit raises an event on the document, which is what a reviewer has open. @@ -470,6 +517,12 @@ func equalConfigStatus(a, b simplyblockv1alpha2.ClusterDeploymentConfigStatus) b if a.Step.Deadline != nil && !a.Step.Deadline.Equal(b.Step.Deadline) { return false } + if (a.ExpansionStartedAt == nil) != (b.ExpansionStartedAt == nil) { + return false + } + if a.ExpansionStartedAt != nil && !a.ExpansionStartedAt.Equal(b.ExpansionStartedAt) { + return false + } for i := range a.NodeRefs { if a.NodeRefs[i] != b.NodeRefs[i] { return false diff --git a/operator/internal/controllers/deployment/expansion.go b/operator/internal/controllers/deployment/expansion.go index 0ebf23e46..de548b8ee 100644 --- a/operator/internal/controllers/deployment/expansion.go +++ b/operator/internal/controllers/deployment/expansion.go @@ -396,6 +396,10 @@ func (r *ClusterDeploymentConfigReconciler) createNodes( return false, fmt.Errorf("creating StorageNode for worker %s slot %d: %w", worker, slot, err) } + // Counted here rather than from the length of the record + // below, which is rebuilt every pass and holds the nodes that + // were already there as well. + countNodesCreated(config.Namespace, 1) created = append(created, node.Name) } } diff --git a/operator/internal/controllers/deployment/metrics.go b/operator/internal/controllers/deployment/metrics.go new file mode 100644 index 000000000..9345f76c2 --- /dev/null +++ b/operator/internal/controllers/deployment/metrics.go @@ -0,0 +1,264 @@ +// The metrics this package publishes about deployment documents and the +// discovery runs that write them. +// +// Both families are new, because neither kind existed before the redesign and +// nothing measured what a deployment costs. Two of them answer questions nothing +// else here can: +// +// - validation_failures_total by reason says whether the approval gate is +// working as review or as an obstacle. A deployment where every draft fails +// on DeviceNotFound is a discovery bug rather than a careful reviewer, and +// the reason label is what separates the two. +// - approval_rejections_total is the pair to read it against. The two count +// the same mistakes at different moments, so a rejection that draft +// validation never reported first is a gap in the validation: the reviewer +// should have been told before they wrote the approval. +// +// A rejected approval has no object to record an event against, since an +// admission rejection fails the request. These counters are the only signal that +// path has, which is why the webhook reaches into this package to raise one. +// +// design-clusterdeploymentconfig.md §9.2 is the specification, and §7.12 of +// design-crd-model.md is the naming rule the suffixes follow. + +package deployment + +import ( + "time" + + "github.com/prometheus/client_golang/prometheus" + ctrlmetrics "sigs.k8s.io/controller-runtime/pkg/metrics" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// The reasons only the webhook produces. The four that a draft's validation also +// finds are the event reasons in events.go, so the two counters share a +// vocabulary and can be read against each other; these three have no event +// because there is no object to raise one on. +const ( + // NoClusterNamed is a document that names neither a cluster to create nor + // one to grow, so it describes no deployment at all. + NoClusterNamed = "NoClusterNamed" + + // ApprovalWithdrawn is an edit clearing spec.approved, which does not + // un-expand what the document already built. + ApprovalWithdrawn = "ApprovalWithdrawn" + + // SpecImmutable is an edit to the spec of a document that is already the + // record of what was deployed. + SpecImmutable = "SpecImmutable" +) + +// expansionBuckets span what an expansion actually takes. Creating the objects +// is seconds, and waiting for the control plane to report the cluster is +// minutes, so a linear set would put every interesting deployment in one bucket. +var expansionBuckets = prometheus.ExponentialBuckets(1, 3, 10) + +var ( + configPhaseState = prometheus.NewGaugeVec( + prometheus.GaugeOpts{ + Name: "simplyblock_clusterdeploymentconfig_phase_state", + Help: "1 for the document's current phase and 0 for the rest, so one stuck in Draft is visible.", + }, + []string{"namespace", "name", "phase"}, + ) + + expansionDurationSeconds = prometheus.NewHistogramVec( + prometheus.HistogramOpts{ + Name: "simplyblock_clusterdeploymentconfig_expansion_duration_seconds", + Help: "How long an expansion took, from the approval that started it to Expanded.", + Buckets: expansionBuckets, + }, + []string{"namespace"}, + ) + + nodesCreatedTotal = prometheus.NewCounterVec( + prometheus.CounterOpts{ + Name: "simplyblock_clusterdeploymentconfig_nodes_created_total", + Help: "StorageNode objects an expansion created, which is the fleet's growth over time.", + }, + []string{"namespace"}, + ) + + validationFailuresTotal = prometheus.NewCounterVec( + prometheus.CounterOpts{ + Name: "simplyblock_clusterdeploymentconfig_validation_failures_total", + Help: "Draft validation failures by reason, which is what a reviewer keeps hitting.", + }, + []string{"namespace", "reason"}, + ) + + approvalRejectionsTotal = prometheus.NewCounterVec( + prometheus.CounterOpts{ + Name: "simplyblock_clusterdeploymentconfig_approval_rejections_total", + Help: "Approving edits admission rejected, by reason: the gate catching a mistake before it is immutable.", + }, + []string{"namespace", "reason"}, + ) + + operatorOperationsTotal = prometheus.NewCounterVec( + prometheus.CounterOpts{ + Name: "simplyblock_operator_operations_total", + Help: "Operator operations that reached a terminal phase, by succeeded, failed, and aborted.", + }, + []string{"namespace", "action", "result"}, + ) + + operatorOperationDurationSeconds = prometheus.NewHistogramVec( + prometheus.HistogramOpts{ + Name: "simplyblock_operator_operation_duration_seconds", + Help: "How long an operator operation took, from its start to a terminal phase.", + Buckets: expansionBuckets, + }, + []string{"namespace", "action"}, + ) + + discoveryWorkersFound = prometheus.NewGaugeVec( + prometheus.GaugeOpts{ + Name: "simplyblock_operator_discovery_workers_found_count", + Help: "Workers the last discovery run of this namespace found.", + }, + []string{"namespace"}, + ) + + discoveryDevicesFound = prometheus.NewGaugeVec( + prometheus.GaugeOpts{ + Name: "simplyblock_operator_discovery_devices_found_count", + Help: "Candidate devices it found, which is the number the draft is written from.", + }, + []string{"namespace"}, + ) +) + +func init() { + ctrlmetrics.Registry.MustRegister( + configPhaseState, + expansionDurationSeconds, + nodesCreatedTotal, + validationFailuresTotal, + approvalRejectionsTotal, + operatorOperationsTotal, + operatorOperationDurationSeconds, + discoveryWorkersFound, + discoveryDevicesFound, + ) +} + +// observeConfigPhase publishes the document's phase as a gauge that is 1 for +// where it is and 0 for everywhere else, so a query for one phase answers for +// every document rather than only for the ones currently in it. +func observeConfigPhase(config *simplyblockv1alpha2.ClusterDeploymentConfig) { + for _, phase := range []simplyblockv1alpha2.ClusterDeploymentConfigPhase{ + simplyblockv1alpha2.ClusterDeploymentConfigPhaseDraft, + simplyblockv1alpha2.ClusterDeploymentConfigPhaseExpanding, + simplyblockv1alpha2.ClusterDeploymentConfigPhaseExpanded, + simplyblockv1alpha2.ClusterDeploymentConfigPhaseFailed, + } { + value := 0.0 + if config.Status.Phase == phase { + value = 1 + } + configPhaseState. + WithLabelValues(config.Namespace, config.Name, string(phase)). + Set(value) + } +} + +// forgetConfig drops a deleted document's series. A gauge keyed on an object's +// name outlives the object, and four series reporting the phase of a document +// nobody can look at is worse than no series at all. +func forgetConfig(config *simplyblockv1alpha2.ClusterDeploymentConfig) { + configPhaseState.DeletePartialMatch(prometheus.Labels{ + "namespace": config.Namespace, + "name": config.Name, + }) +} + +// countValidationFailures records one failure per distinct reason. +// +// One per reason rather than one per finding, matching the events: three groups +// naming the same missing worker is one thing wrong, and counting it three times +// would make a document with many groups look like a deployment with many +// problems. +func countValidationFailures(namespace string, findings []finding) { + seen := map[string]struct{}{} + for _, found := range findings { + if _, already := seen[found.reason]; already { + continue + } + seen[found.reason] = struct{}{} + validationFailuresTotal.WithLabelValues(namespace, found.reason).Inc() + } +} + +// CountApprovalRejection records an approving edit admission refused. It is +// exported because the webhook that refuses lives next door and this is the only +// record that path leaves. +func CountApprovalRejection(namespace, reason string) { + approvalRejectionsTotal.WithLabelValues(namespace, reason).Inc() +} + +// observeExpansion records how long the expansion took. +// +// A document whose start was never recorded is not measured at all. That is a +// document the operator was upgraded underneath mid-expansion, and a duration +// measured from an instant nobody wrote down would be a number the histogram +// cannot distinguish from a real one. +func observeExpansion(config *simplyblockv1alpha2.ClusterDeploymentConfig) { + started := config.Status.ExpansionStartedAt + if started == nil { + return + } + expansionDurationSeconds.WithLabelValues(config.Namespace). + Observe(time.Since(started.Time).Seconds()) +} + +// countNodesCreated records the StorageNodes an expansion pass created. It +// counts what was actually created rather than what the document describes, so +// a pass that found every node already there adds nothing. +func countNodesCreated(namespace string, created int) { + if created <= 0 { + return + } + nodesCreatedTotal.WithLabelValues(namespace).Add(float64(created)) +} + +// observeRun records a discovery run's outcome and how long it took. +func observeRun( + ops *simplyblockv1alpha2.OperatorOps, phase simplyblockv1alpha2.OperatorOpsPhase, +) { + action := string(ops.Spec.Action) + operatorOperationsTotal. + WithLabelValues(ops.Namespace, action, runResultOf(phase)).Inc() + if started := ops.Status.StartedAt; started != nil { + operatorOperationDurationSeconds.WithLabelValues(ops.Namespace, action). + Observe(time.Since(started.Time).Seconds()) + } +} + +// runResultOf is the metric label for a terminal phase, lowercased because a +// label value is not an API enum. +func runResultOf(phase simplyblockv1alpha2.OperatorOpsPhase) string { + switch phase { + case simplyblockv1alpha2.OperatorOpsPhaseSucceeded: + return "succeeded" + case simplyblockv1alpha2.OperatorOpsPhaseAborted: + return "aborted" + default: + return "failed" + } +} + +// observeWorkersFound and observeDevicesFound publish what the last run of a +// namespace concluded. They are separate calls because the two numbers are +// settled by different steps: which workers the run is about is decided in +// Inspecting and never revisited, and how many devices survived the rules is +// only known once the plan is built in Writing. +func observeWorkersFound(namespace string, workers int) { + discoveryWorkersFound.WithLabelValues(namespace).Set(float64(workers)) +} + +func observeDevicesFound(namespace string, devices int) { + discoveryDevicesFound.WithLabelValues(namespace).Set(float64(devices)) +} diff --git a/operator/internal/controllers/deployment/metrics_test.go b/operator/internal/controllers/deployment/metrics_test.go new file mode 100644 index 000000000..7aaa844d1 --- /dev/null +++ b/operator/internal/controllers/deployment/metrics_test.go @@ -0,0 +1,226 @@ +// What the deployment metrics say, and the two ways they would lie quietly. +// +// A counter incremented on every reconcile counts passes rather than events, +// which is the failure a validation counter is most prone to: a draft nobody +// fixes is re-validated every thirty seconds forever. And a duration measured +// from an instant nobody recorded is indistinguishable from a real one once it +// is in a histogram. +// +// Every case resets the series it asserts on. The metrics are package-level and +// the harness next door shares one namespace, so a test that read a running +// total would be asserting about the whole file's history. + +package deployment + +import ( + "context" + "testing" + "time" + + "github.com/prometheus/client_golang/prometheus/testutil" + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// aDraftNaming builds an unapproved document whose only fault is the worker it +// names, which is the validation failure a reviewer actually hits. +func aDraftNaming(worker string) *simplyblockv1alpha2.ClusterDeploymentConfig { + return aDocument(func(c *simplyblockv1alpha2.ClusterDeploymentConfig) { + c.Spec.Approved = false + c.Spec.NodeSets[0].Groups[0].Workers = []string{worker} + }) +} + +// A draft nobody fixes is validated again every pass. The counter is about what +// reviewers keep hitting, so it counts the outcome rather than the reconcile. +func TestAValidationFailureIsCountedOncePerOutcome(t *testing.T) { + validationFailuresTotal.Reset() + config := aDraftNaming("worker-does-not-exist") + r := reconcilerFor(t, config, &corev1.Node{ + ObjectMeta: metav1.ObjectMeta{Name: "worker-1"}}) + + request := ctrl.Request{NamespacedName: client.ObjectKeyFromObject(config)} + for pass := 0; pass < 3; pass++ { + if _, err := r.Reconcile(context.Background(), request); err != nil { + t.Fatalf("pass %d: %v", pass, err) + } + } + + counted := testutil.ToFloat64( + validationFailuresTotal.WithLabelValues(theNamespace, WorkerNotFound)) + if counted != 1 { + t.Errorf("three passes over one unfixed draft counted %v failures, want 1", counted) + } +} + +// The phase gauge answers for every document rather than only for the ones in +// the phase asked about, which is what makes "how many are stuck in Draft" a +// query rather than an absence. +func TestThePhaseGaugeIsOneForTheCurrentPhaseAndZeroForTheRest(t *testing.T) { + config := aDocument(nil) + config.Namespace = "phases" + r := reconcilerFor(t, config) + + err := r.note(context.Background(), config, + simplyblockv1alpha2.ClusterDeploymentConfigPhaseExpanding, "creating the cluster") + if err != nil { + t.Fatalf("note: %v", err) + } + + expanding := testutil.ToFloat64(configPhaseState.WithLabelValues( + "phases", config.Name, string(simplyblockv1alpha2.ClusterDeploymentConfigPhaseExpanding))) + if expanding != 1 { + t.Errorf("the Expanding series is %v, want 1", expanding) + } + draft := testutil.ToFloat64(configPhaseState.WithLabelValues( + "phases", config.Name, string(simplyblockv1alpha2.ClusterDeploymentConfigPhaseDraft))) + if draft != 0 { + t.Errorf("the Draft series is %v for an expanding document, want 0", draft) + } +} + +// A deleted document's series would otherwise report the phase of something +// nobody can look at. +func TestADeletedDocumentStopsReportingAPhase(t *testing.T) { + config := aDocument(nil) + config.Namespace = "forgotten" + r := reconcilerFor(t, config) + + if err := r.note(context.Background(), config, + simplyblockv1alpha2.ClusterDeploymentConfigPhaseExpanded, "done"); err != nil { + t.Fatalf("note: %v", err) + } + forgetConfig(config) + + if count := testutil.CollectAndCount(configPhaseState); count == 0 { + t.Skip("nothing is published at all, so this proves nothing") + } + got := testutil.ToFloat64(configPhaseState.WithLabelValues("forgotten", config.Name, + string(simplyblockv1alpha2.ClusterDeploymentConfigPhaseExpanded))) + if got != 0 { + t.Errorf("a forgotten document still reports %v", got) + } +} + +// The duration is from approval rather than from the document being written, and +// the only durable record of that instant is the status field. +func TestAnExpansionIsTimedFromTheInstantItStarted(t *testing.T) { + expansionDurationSeconds.Reset() + config := aDocument(nil) + config.Namespace = "timed" + started := metav1.NewTime(time.Now().Add(-90 * time.Second)) + config.Status.ExpansionStartedAt = &started + config.Status.ClusterRef = theCluster + r := reconcilerFor(t, config) + + if err := r.succeed(context.Background(), config); err != nil { + t.Fatalf("succeed: %v", err) + } + + if count := testutil.CollectAndCount(expansionDurationSeconds); count != 1 { + t.Fatalf("the histogram holds %d series, want the one this expansion observed", count) + } +} + +// A document the operator was upgraded underneath has no recorded start, and a +// duration invented for it would be a sample nothing can tell from a real one. +func TestAnExpansionWithNoRecordedStartIsNotTimed(t *testing.T) { + expansionDurationSeconds.Reset() + config := aDocument(nil) + config.Namespace = "untimed" + config.Status.ClusterRef = theCluster + r := reconcilerFor(t, config) + + if err := r.succeed(context.Background(), config); err != nil { + t.Fatalf("succeed: %v", err) + } + + if count := testutil.CollectAndCount(expansionDurationSeconds); count != 0 { + t.Errorf("an expansion with no start was timed anyway (%d series)", count) + } +} + +// The counter is the fleet's growth, so it counts objects created and not +// objects the document describes. Creating nodes is idempotent and runs again on +// every retry of the step. +func TestOnlyTheNodesActuallyCreatedAreCounted(t *testing.T) { + nodesCreatedTotal.Reset() + config := aDocument(nil) + config.Status.ClusterRef = theCluster + cluster := aCluster(nil) + objects := append(workers("worker-1", "worker-2"), config, cluster) + r := reconcilerFor(t, objects...) + + for pass := 0; pass < 2; pass++ { + if _, err := r.createNodes(context.Background(), config); err != nil { + t.Fatalf("pass %d: %v", pass, err) + } + } + + counted := testutil.ToFloat64(nodesCreatedTotal.WithLabelValues(theNamespace)) + if counted != 2 { + t.Errorf("two workers created over two passes counted %v nodes, want 2", counted) + } +} + +// The two discovery gauges are what the last run of a namespace concluded, and +// they are settled by different steps: which workers the run is about is decided +// in Inspecting, and how many devices survived the rules is only known once the +// plan is built in Writing. +func TestADiscoveryRunPublishesWhatItFound(t *testing.T) { + r := newRunner(t, discoverRun(nil), worker("worker-1"), worker("worker-2")) + + r.step() // start + r.step() // inspect + + workersFound := testutil.ToFloat64(discoveryWorkersFound.WithLabelValues(opsNamespace)) + if workersFound != 2 { + t.Errorf("the run found %v workers, want 2", workersFound) + } + + r.step() // probing: creates the Jobs + for _, node := range []string{"worker-1", "worker-2"} { + cm := reportConfigMap(t, node, + "0000:5e:00.0", "0000:5f:00.0", "0000:af:00.0", "0000:b0:00.0") + if err := r.client.Create(context.Background(), cm); err != nil { + t.Fatalf("write a report: %v", err) + } + } + r.step() // probing: sees the reports + r.step() // writing + + // The placement chose one memory node, so two of each worker's four disks. + devicesFound := testutil.ToFloat64(discoveryDevicesFound.WithLabelValues(opsNamespace)) + if devicesFound != 4 { + t.Errorf("the run found %v devices, want the four the draft names", devicesFound) + } +} + +// A run that reached a terminal phase is counted with the outcome it reached and +// timed from when it started, which is the pair every other Ops kind publishes. +func TestATerminalRunIsCountedWithItsResult(t *testing.T) { + operatorOperationsTotal.Reset() + operatorOperationDurationSeconds.Reset() + + ops := discoverRun(nil) + started := metav1.NewTime(time.Now().Add(-30 * time.Second)) + ops.Status.StartedAt = &started + r := newRunner(t, ops, worker("worker-1")) + + if err := r.reconciler.succeed(context.Background(), ops); err != nil { + t.Fatalf("succeed: %v", err) + } + + counted := testutil.ToFloat64(operatorOperationsTotal. + WithLabelValues(opsNamespace, string(ops.Spec.Action), "succeeded")) + if counted != 1 { + t.Errorf("a finished run counted %v, want 1", counted) + } + if series := testutil.CollectAndCount(operatorOperationDurationSeconds); series != 1 { + t.Errorf("the duration histogram holds %d series, want the one this run observed", series) + } +} diff --git a/operator/internal/controllers/deployment/operatorops_controller.go b/operator/internal/controllers/deployment/operatorops_controller.go index 81a130d4a..4d0b03f11 100644 --- a/operator/internal/controllers/deployment/operatorops_controller.go +++ b/operator/internal/controllers/deployment/operatorops_controller.go @@ -311,6 +311,7 @@ func (r *OperatorOpsReconciler) succeed( ops.Status.CompletedAt = &now ops.Status.Step.Deadline = nil r.event(ops, corev1.EventTypeNormal, OperationSucceeded, ops.Status.Message) + observeRun(ops, ops.Status.Phase) return r.status(ctx, ops) } @@ -393,6 +394,7 @@ func (r *OperatorOpsReconciler) inspect( ops.Status.Workers = workers ops.Status.Environment = simplyblockv1alpha2.KubernetesEnvironment(environment.Distribution) + observeWorkersFound(ops.Namespace, len(workers)) ops.Status.Message = fmt.Sprintf("probing %d worker(s) of a %s cluster", len(workers), orUnknown(string(environment.Distribution))) @@ -583,6 +585,7 @@ func (r *OperatorOpsReconciler) write( ops.Status.ConfigRef = config.Name ops.Status.Message = fmt.Sprintf("wrote %s awaiting approval: %s", config.Name, plan.Summary()) + observeDevicesFound(ops.Namespace, plan.DeviceCount()) return true, r.status(ctx, ops) } @@ -714,6 +717,7 @@ func (r *OperatorOpsReconciler) abort( ops.Status.Step.Deadline = nil ops.Status.Message = "aborted; discovery changes nothing, so nothing was undone" r.event(ops, corev1.EventTypeNormal, OperationAborted, ops.Status.Message) + observeRun(ops, ops.Status.Phase) return ctrl.Result{}, r.status(ctx, ops) } @@ -756,6 +760,7 @@ func (r *OperatorOpsReconciler) fail( ops.Status.Step.Deadline = nil ops.Status.Message = reason r.event(ops, corev1.EventTypeWarning, OperationFailed, reason) + observeRun(ops, ops.Status.Phase) return ctrl.Result{}, r.status(ctx, ops) } diff --git a/operator/internal/controllers/driver/metrics.go b/operator/internal/controllers/driver/metrics.go new file mode 100644 index 000000000..03bf6abcd --- /dev/null +++ b/operator/internal/controllers/driver/metrics.go @@ -0,0 +1,101 @@ +// The gauges the operator publishes about the CSI driver it deploys. +// +// All four are gauges and none of them counts an event, because a driver has no +// operations: it is deployed, and what matters is whether it is up. The question +// a dashboard asks is whether provisioning works right now, which is the +// controller plugin, and whether every worker can attach, which is the two node +// counts read as a ratio. +// +// Neither node count means anything alone. Three ready plugins is healthy on a +// three-worker cluster and an outage on a thirty-worker one, so the pair is +// published from the same pass or not at all. +// +// design-simplyblockdriver.md §6.2 is the specification, and §7.12 of +// design-crd-model.md is the naming rule the suffixes follow. + +package driver + +import ( + "github.com/prometheus/client_golang/prometheus" + ctrlmetrics "sigs.k8s.io/controller-runtime/pkg/metrics" +) + +var ( + // driverVersionInfo carries the reported version in a label and is 1 for it, + // which is the shape §7.12 gives an `info` gauge. Beside the control plane's + // own version gauge, the pair is the skew alert: a driver and a control plane + // that disagree fail in the data path at attach time, on a workload's pod. + // + // TODO(simplyblockdriver): it is 0 today, because status.version is written + // by nothing. The two missing pieces are recorded on the TODO in + // simplyblockdriver_controller.go beside the field: GET /_meta/version on the + // management API, and the release document that says which control planes a + // driver works against. Zero is the honest reading until then — an info gauge + // is 1 for a fact it carries, and there is no fact yet — and the series + // becomes correct on the day the field is filled in, with no change here. + driverVersionInfo = prometheus.NewGaugeVec( + prometheus.GaugeOpts{ + Name: "simplyblock_simplyblockdriver_version_info", + Help: "1 for the version the deployed driver reports, and 0 while it reports none.", + }, + []string{"namespace", "version"}, + ) + + driverNodesReady = prometheus.NewGaugeVec( + prometheus.GaugeOpts{ + Name: "simplyblock_simplyblockdriver_nodes_ready_count", + Help: "Workers running a ready node plugin.", + }, + []string{"namespace"}, + ) + + driverNodesExpected = prometheus.NewGaugeVec( + prometheus.GaugeOpts{ + Name: "simplyblock_simplyblockdriver_nodes_expected_count", + Help: "Workers expected to run one, so the ratio against nodes_ready_count is the alert.", + }, + []string{"namespace"}, + ) + + driverControllerReady = prometheus.NewGaugeVec( + prometheus.GaugeOpts{ + Name: "simplyblock_simplyblockdriver_controller_ready_state", + Help: "1 while the controller plugin serves. Provisioning stops when it is zero.", + }, + []string{"namespace"}, + ) +) + +func init() { + ctrlmetrics.Registry.MustRegister( + driverVersionInfo, + driverNodesReady, + driverNodesExpected, + driverControllerReady, + ) +} + +// observeHealth publishes one reading of a deployment. +// +// The version series is replaced rather than added to. An info gauge says which +// fact is current by carrying it in a label, so a driver that was upgraded would +// otherwise leave the version it used to run sitting at 1 beside the one it +// runs now, and a dashboard reading either would be reading a claim about the +// present. +func observeHealth(namespace, version string, h health) { + driverVersionInfo.DeletePartialMatch(prometheus.Labels{"namespace": namespace}) + reported := 0.0 + if version != "" { + reported = 1 + } + driverVersionInfo.WithLabelValues(namespace, version).Set(reported) + + driverNodesReady.WithLabelValues(namespace).Set(float64(h.nodesReady)) + driverNodesExpected.WithLabelValues(namespace).Set(float64(h.nodesTotal)) + + serving := 0.0 + if h.controllerReady { + serving = 1 + } + driverControllerReady.WithLabelValues(namespace).Set(serving) +} diff --git a/operator/internal/controllers/driver/metrics_test.go b/operator/internal/controllers/driver/metrics_test.go new file mode 100644 index 000000000..472fa51d8 --- /dev/null +++ b/operator/internal/controllers/driver/metrics_test.go @@ -0,0 +1,151 @@ +// What the driver's gauges say, and what they deliberately do not say yet. + +package driver + +import ( + "context" + "testing" + + "github.com/prometheus/client_golang/prometheus" + "github.com/prometheus/client_golang/prometheus/testutil" + dto "github.com/prometheus/client_model/go" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// aDriverIn builds a reconciler over a driver in its own namespace, so one +// test's gauges are not another's: the metrics are package-level and the +// namespace is their label. +func aDriverIn(t *testing.T, namespace string) ( + *SimplyblockDriverReconciler, *simplyblockv1alpha2.SimplyblockDriver, +) { + t.Helper() + scheme := reconcilerScheme(t) + d := testDriver("simplyblock") + d.Namespace = namespace + c := fake.NewClientBuilder().WithScheme(scheme). + WithObjects(d).WithStatusSubresource(d).Build() + return &SimplyblockDriverReconciler{Client: c, Scheme: scheme}, d +} + +// The ready count and the expected count are one signal in two series, and the +// design says neither half means anything alone. A pass that published one and +// not the other would give a dashboard a ratio it cannot compute. +func TestTheNodePluginCountsArePublishedTogether(t *testing.T) { + r, d := aDriverIn(t, "counts") + + err := r.setHealth(context.Background(), d, health{ + phase: simplyblockv1alpha2.SimplyblockDriverPhaseDegraded, + nodesReady: 2, + nodesTotal: 3, + controllerReady: true, + message: "one node plugin is not ready", + }) + if err != nil { + t.Fatalf("setHealth: %v", err) + } + + if got := testutil.ToFloat64(driverNodesReady.WithLabelValues("counts")); got != 2 { + t.Errorf("nodes_ready_count is %v, want 2", got) + } + if got := testutil.ToFloat64(driverNodesExpected.WithLabelValues("counts")); got != 3 { + t.Errorf("nodes_expected_count is %v, want 3", got) + } +} + +// Provisioning stops when the controller plugin is not serving, so the gauge is +// the one that turns that into an alert rather than a phase somebody reads. +func TestTheControllerStateIsOneOrZero(t *testing.T) { + r, d := aDriverIn(t, "controller") + + serving := health{phase: simplyblockv1alpha2.SimplyblockDriverPhaseReady, + nodesReady: 3, nodesTotal: 3, controllerReady: true} + if err := r.setHealth(context.Background(), d, serving); err != nil { + t.Fatalf("setHealth: %v", err) + } + if got := testutil.ToFloat64(driverControllerReady.WithLabelValues("controller")); got != 1 { + t.Errorf("controller_ready_state is %v while it serves, want 1", got) + } + + down := health{phase: simplyblockv1alpha2.SimplyblockDriverPhaseUnavailable, + nodesReady: 3, nodesTotal: 3} + if err := r.setHealth(context.Background(), d, down); err != nil { + t.Fatalf("setHealth: %v", err) + } + if got := testutil.ToFloat64(driverControllerReady.WithLabelValues("controller")); got != 0 { + t.Errorf("controller_ready_state is %v while it is down, want 0", got) + } +} + +// status.version is not written by anything yet, and the gauge says so rather +// than claiming a version. An _info gauge is 1 for the fact it carries in its +// labels, so 0 is the honest reading for a fact nobody has established. +func TestTheVersionGaugeIsZeroUntilSomethingReportsAVersion(t *testing.T) { + r, d := aDriverIn(t, "version") + + err := r.setHealth(context.Background(), d, health{ + phase: simplyblockv1alpha2.SimplyblockDriverPhaseReady, nodesReady: 1, nodesTotal: 1, + controllerReady: true, + }) + if err != nil { + t.Fatalf("setHealth: %v", err) + } + if got := testutil.ToFloat64(driverVersionInfo.WithLabelValues("version", "")); got != 0 { + t.Errorf("version_info is %v with no reported version, want 0", got) + } + + // And it is 1 the moment one arrives, so the series is right on the day §5 + // lands rather than needing a second change. The version is written through + // the client because that is where setHealth reads it back from. + var stored simplyblockv1alpha2.SimplyblockDriver + if err := r.Get(context.Background(), client.ObjectKeyFromObject(d), &stored); err != nil { + t.Fatalf("read the driver back: %v", err) + } + stored.Status.Version = "v26.2.6" + if err := r.Status().Update(context.Background(), &stored); err != nil { + t.Fatalf("record a reported version: %v", err) + } + if err := r.setHealth(context.Background(), d, health{ + phase: simplyblockv1alpha2.SimplyblockDriverPhaseReady, nodesReady: 1, nodesTotal: 1, + controllerReady: true, + }); err != nil { + t.Fatalf("setHealth: %v", err) + } + if got := testutil.ToFloat64(driverVersionInfo.WithLabelValues("version", "v26.2.6")); got != 1 { + t.Errorf("version_info is %v for the reported version, want 1", got) + } + versions := versionSeriesIn(driverVersionInfo, "version") + if len(versions) != 1 || versions[0] != "v26.2.6" { + t.Errorf("the gauge holds %v for this namespace; a version that was "+ + "replaced leaves a second series claiming to be the current one", + versions) + } +} + +// versionSeriesIn is the version labels the gauge holds for one namespace. The +// point of the reading is how many there are: a metric whose label carries the +// fact has to drop the old fact when it changes, and nothing but the series +// themselves shows whether it did. +func versionSeriesIn(collector prometheus.Collector, namespace string) []string { + gathered := make(chan prometheus.Metric, 32) + collector.Collect(gathered) + close(gathered) + + var versions []string + for metric := range gathered { + var written dto.Metric + if err := metric.Write(&written); err != nil { + continue + } + labels := map[string]string{} + for _, pair := range written.GetLabel() { + labels[pair.GetName()] = pair.GetValue() + } + if labels["namespace"] == namespace { + versions = append(versions, labels["version"]) + } + } + return versions +} diff --git a/operator/internal/controllers/driver/simplyblockdriver_controller.go b/operator/internal/controllers/driver/simplyblockdriver_controller.go index 2b6e7a418..ca1371f2e 100644 --- a/operator/internal/controllers/driver/simplyblockdriver_controller.go +++ b/operator/internal/controllers/driver/simplyblockdriver_controller.go @@ -606,25 +606,29 @@ func (r *SimplyblockDriverReconciler) setStatus( }) } -// TODO(simplyblockdriver): publish design §6.2's four gauges, -// simplyblock_simplyblockdriver_{version_info,nodes_ready_count, -// nodes_expected_count,controller_ready_state}. The three that are not the -// version are computable from the health below and need only a collector; the -// version one waits on §5. - // setHealth writes the phase together with the counts it is explained by, since // a phase a reader cannot check against the numbers behind it sends them to // kubectl describe to learn which worker is short. +// +// It is also where §6.2's gauges are published, because this is the pass that +// measured them. The paths that write a phase without one — a refusal, a +// failure — leave the last reading standing rather than replacing it with +// zeros they did not observe. func (r *SimplyblockDriverReconciler) setHealth( ctx context.Context, d *simplyblockv1alpha2.SimplyblockDriver, h health, ) error { - return r.writeStatus(ctx, d, func(status *simplyblockv1alpha2.SimplyblockDriverStatus) { + err := r.writeStatus(ctx, d, func(status *simplyblockv1alpha2.SimplyblockDriverStatus) { status.Phase = h.phase status.Message = h.message status.NodesReady = h.nodesReady status.NodesTotal = h.nodesTotal status.ControllerReady = h.controllerReady }) + if err != nil { + return err + } + observeHealth(d.Namespace, d.Status.Version, h) + return nil } func (r *SimplyblockDriverReconciler) writeStatus( diff --git a/operator/internal/controllers/names_test.go b/operator/internal/controllers/names_test.go new file mode 100644 index 000000000..cee544e59 --- /dev/null +++ b/operator/internal/controllers/names_test.go @@ -0,0 +1,229 @@ +// The naming rule for every Prometheus metric the operator exports, checked +// against the source rather than against a list somebody maintains. +// +// design-crd-model.md §7.12 names six aggregation suffixes and no seventh, and +// binds three of them to a metric type: `total` is a counter and only a counter, +// and `count`, `state`, and `info` are gauges. The rule is worth enforcing rather +// than remembering because a gauge ending in `total` reads as something `rate()` +// applies to, and a dashboard built on that is wrong in a way nothing reports. +// +// It lives in a package of its own, above the controllers rather than in one of +// them, because the rule is the operator's and a test inside one package would +// check one package. It reads the source with go/ast instead of scraping a +// registry, so a metric declared in a package nothing imports is covered too. +// +// The csi-driver module exports its own metrics and is not read here. §7.12 is +// this operator's rule, and the driver's names are its design's to state. + +package controllers + +import ( + "go/ast" + "go/parser" + "go/token" + "io/fs" + "path/filepath" + "strconv" + "strings" + "testing" +) + +// metricPrefix is what §7.12's `simplyblock___` begins with. +const metricPrefix = "simplyblock_" + +// optsTypes are the four constructors' option structs. Finding the literal is +// what gives both the name and the metric's type, which no amount of reading the +// name alone can supply. +var optsTypes = map[string]string{ + "CounterOpts": "counter", + "GaugeOpts": "gauge", + "HistogramOpts": "histogram", + "SummaryOpts": "summary", +} + +// aggregations are §7.12's six words, each mapped to the metric types it may +// name. An empty set is a suffix that constrains the unit rather than the type: +// a duration or a size can be measured by any of them. +var aggregations = map[string]map[string]bool{ + "total": {"counter": true}, + "count": {"gauge": true}, + "state": {"gauge": true}, + "info": {"gauge": true}, + "seconds": {}, + "bytes": {}, +} + +// legacyFamilies predate §7.12 and belong to the rebalancer, which is being +// absorbed into the operator's own kinds. Their names are part of a dashboard +// somebody is running today, so they are renamed when that subsystem moves +// rather than one at a time; until then the rule would fail on names no new +// metric is allowed to imitate. +var legacyFamilies = []string{ + "simplyblock_rebalancer_", + "simplyblock_node_fio_", +} + +// declaredMetric is one metric found in the source. +type declaredMetric struct { + name string + kind string + file string + line int +} + +func TestEveryMetricIsNamedTheWayTheModelSays(t *testing.T) { + metrics := metricsDeclaredIn(t, filepath.Join("..", "..")) + if len(metrics) == 0 { + t.Fatal("no metric was found, so this test is checking nothing") + } + + for _, metric := range metrics { + if isLegacy(metric.name) { + continue + } + where := metric.file + ":" + strconv.Itoa(metric.line) + + if !strings.HasPrefix(metric.name, metricPrefix) { + t.Errorf("%s: %s does not begin with %s", where, metric.name, metricPrefix) + continue + } + + suffix := metric.name[strings.LastIndex(metric.name, "_")+1:] + kinds, declared := aggregations[suffix] + if !declared { + t.Errorf("%s: %s ends in %q, which is not one of §7.12's six aggregations", + where, metric.name, suffix) + continue + } + if len(kinds) > 0 && !kinds[metric.kind] { + t.Errorf("%s: %s is a %s, and §7.12 reserves the %q suffix for a %s", + where, metric.name, metric.kind, suffix, only(kinds)) + } + } +} + +// A name nothing exports twice. Two metrics of one name are a registration panic +// at startup on the same registry, and a silent split of one series across two +// meanings on different ones. +func TestNoMetricNameIsDeclaredTwice(t *testing.T) { + seen := map[string]declaredMetric{} + for _, metric := range metricsDeclaredIn(t, filepath.Join("..", "..")) { + if first, already := seen[metric.name]; already { + t.Errorf("%s is declared at %s:%d and again at %s:%d", + metric.name, first.file, first.line, metric.file, metric.line) + continue + } + seen[metric.name] = metric + } +} + +// metricsDeclaredIn reads every non-test Go file below root and returns the +// metrics its Prometheus option literals name. +func metricsDeclaredIn(t *testing.T, root string) []declaredMetric { + t.Helper() + var metrics []declaredMetric + fileSet := token.NewFileSet() + + err := filepath.WalkDir(root, func(path string, entry fs.DirEntry, err error) error { + switch { + case err != nil: + return err + case entry.IsDir(): + if entry.Name() == "vendor" || entry.Name() == "bin" { + return filepath.SkipDir + } + return nil + case !strings.HasSuffix(path, ".go") || strings.HasSuffix(path, "_test.go"): + return nil + } + + parsed, err := parser.ParseFile(fileSet, path, nil, 0) + if err != nil { + return err + } + ast.Inspect(parsed, func(node ast.Node) bool { + literal, ok := node.(*ast.CompositeLit) + if !ok { + return true + } + kind, ok := optsKind(literal) + if !ok { + return true + } + name, ok := stringField(literal, "Name") + if !ok { + // A name computed rather than written. Nothing does this today, + // and a test that quietly skipped one would report a clean run + // over a metric it never read. + t.Errorf("%s: a %s literal names its metric with something other "+ + "than a string constant, which this rule cannot read", + fileSet.Position(literal.Pos()), kind) + return true + } + position := fileSet.Position(literal.Pos()) + metrics = append(metrics, declaredMetric{ + name: name, kind: kind, file: position.Filename, line: position.Line, + }) + return true + }) + return nil + }) + if err != nil { + t.Fatalf("read the operator's source: %v", err) + } + return metrics +} + +// optsKind reports which of the four option structs a literal is, by the +// selector it is spelled with. +func optsKind(literal *ast.CompositeLit) (string, bool) { + selector, ok := literal.Type.(*ast.SelectorExpr) + if !ok { + return "", false + } + if pkg, ok := selector.X.(*ast.Ident); !ok || pkg.Name != "prometheus" { + return "", false + } + kind, ok := optsTypes[selector.Sel.Name] + return kind, ok +} + +// stringField reads one string-constant field out of a composite literal. +func stringField(literal *ast.CompositeLit, field string) (string, bool) { + for _, element := range literal.Elts { + pair, ok := element.(*ast.KeyValueExpr) + if !ok { + continue + } + if key, ok := pair.Key.(*ast.Ident); !ok || key.Name != field { + continue + } + value, ok := pair.Value.(*ast.BasicLit) + if !ok || value.Kind != token.STRING { + return "", false + } + unquoted, err := strconv.Unquote(value.Value) + if err != nil { + return "", false + } + return unquoted, true + } + return "", false +} + +func isLegacy(name string) bool { + for _, family := range legacyFamilies { + if strings.HasPrefix(name, family) { + return true + } + } + return false +} + +// only names the single metric type a suffix admits, for the failure message. +func only(kinds map[string]bool) string { + for kind := range kinds { + return kind + } + return "" +} diff --git a/operator/internal/discovery/plan.go b/operator/internal/discovery/plan.go index 430631afd..5d89d7c70 100644 --- a/operator/internal/discovery/plan.go +++ b/operator/internal/discovery/plan.go @@ -73,10 +73,7 @@ type Plan struct { // Summary is the sentence a run's status carries. func (p Plan) Summary() string { - devices := 0 - for _, worker := range p.Workers { - devices += len(worker.Addresses()) - } + devices := p.DeviceCount() groups := 0 for _, set := range p.NodeSets { @@ -88,6 +85,17 @@ func (p Plan) Summary() string { len(p.Workers), devices, p.Class, groups, len(p.NodeSets), len(p.Refusals)) } +// DeviceCount is how many devices the draft names across every worker in it, +// which is the number the plan is worth: a run that found ten machines and one +// disk between them has produced nothing to deploy. +func (p Plan) DeviceCount() int { + devices := 0 + for _, worker := range p.Workers { + devices += len(worker.Addresses()) + } + return devices +} + // Explain says why the plan holds nothing, one line per worker. // // A worker is refused either as a whole — its controllers are on a userspace diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml index aed657d41..7f68f56a5 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml @@ -399,6 +399,19 @@ spec: is a record rather than a dependency: nothing resolves it after expansion, which is what makes the document safe to delete. type: string + expansionStartedAt: + description: |- + ExpansionStartedAt is when the expansion machine was born, which is the + first reconcile after the document was approved. A document may sit as a + draft for as long as a review takes, so this is not creationTimestamp and + the difference is the whole point: how long a deployment takes is measured + from the moment somebody said yes. + + It is the start of §9.2's expansion_duration_seconds. A histogram needs an + instant that survives the operator restarting mid-expansion, which nothing + in memory and no step deadline supplies. + format: date-time + type: string message: description: |- Message is the reason the phase is what it is: one sentence, replaced as diff --git a/operator/internal/webhook/clusterdeploymentconfig_metrics_test.go b/operator/internal/webhook/clusterdeploymentconfig_metrics_test.go new file mode 100644 index 000000000..a50c40273 --- /dev/null +++ b/operator/internal/webhook/clusterdeploymentconfig_metrics_test.go @@ -0,0 +1,131 @@ +// The counter an admission rejection leaves behind. +// +// A rejected approval fails the request, so there is no object to raise an event +// on and nothing in the cluster records that it happened. The metric is the only +// signal that path has, and it is the pair to read the controller's validation +// counter against: a rejection whose reason draft validation never reported +// first is a gap in the validation, because the reviewer should have been told +// before they wrote the approval. + +package webhook + +import ( + "testing" + + admissionv1 "k8s.io/api/admission/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + ctrlmetrics "sigs.k8s.io/controller-runtime/pkg/metrics" + + "github.com/simplyblock/simplyblock-operator/internal/controllers/deployment" +) + +// approvalRejectionsMetric is the series this file reads, named rather than +// reached for: the counter is the deployment package's, and a scraper is the +// only consumer either package has in common. +const approvalRejectionsMetric = "simplyblock_clusterdeploymentconfig_approval_rejections_total" + +func TestARejectedApprovalIsCountedByReason(t *testing.T) { + cases := []struct { + name string + reason string + rejected func(t *testing.T) int + }{{ + name: "a worker that is not a node", + reason: deployment.WorkerNotFound, + rejected: func(t *testing.T) int { + config := testConfig() + config.Spec.NodeSets[0].Groups[0].Workers = []string{"worker-9"} + return rejectionsAfter(t, deployment.WorkerNotFound, func() { + mustDeny(t, approve(t, []client.Object{testWorker("worker-1")}, config)) + }) + }, + }, { + name: "a clusterRef resolving to nothing", + reason: deployment.ClusterNotFound, + rejected: func(t *testing.T) int { + return rejectionsAfter(t, deployment.ClusterNotFound, func() { + mustDeny(t, approve(t, []client.Object{testWorker("worker-1")}, + grows(testConfig(), testDeploymentCluster))) + }) + }, + }, { + name: "a spec edited after approval", + reason: deployment.SpecImmutable, + rejected: func(t *testing.T) int { + old := approved(testConfig()) + edited := approved(testConfig()) + edited.Spec.NodeSets[0].Groups[0].Workers = []string{"worker-2"} + return rejectionsAfter(t, deployment.SpecImmutable, func() { + mustDeny(t, review(t, []client.Object{testWorker("worker-1")}, + admissionv1.Update, old, edited)) + }) + }, + }, { + name: "an approval withdrawn", + reason: deployment.ApprovalWithdrawn, + rejected: func(t *testing.T) int { + old := approved(testConfig()) + withdrawn := testConfig() + withdrawn.Spec.Approved = false + return rejectionsAfter(t, deployment.ApprovalWithdrawn, func() { + mustDeny(t, review(t, []client.Object{testWorker("worker-1")}, + admissionv1.Update, old, withdrawn)) + }) + }, + }} + + for _, tc := range cases { + t.Run(tc.name, func(t *testing.T) { + if counted := tc.rejected(t); counted != 1 { + t.Errorf("the rejection raised %d under %s, want 1", counted, tc.reason) + } + }) + } +} + +// An admitted approval counts nothing. A counter that also rose on the happy +// path would make the rate meaningless, which is the only thing it is read as. +func TestAnAdmittedApprovalIsCountedNowhere(t *testing.T) { + counted := rejectionsAfter(t, deployment.WorkerNotFound, func() { + mustAllow(t, approve(t, []client.Object{testWorker("worker-1")}, testConfig())) + }) + if counted != 0 { + t.Errorf("an admitted approval raised %d rejections", counted) + } +} + +// rejectionsAfter is how many rejections of one reason the given admission +// raised, measured as a difference so that the test says nothing about what ran +// before it. +func rejectionsAfter(t *testing.T, reason string, admit func()) int { + t.Helper() + before := approvalRejections(t, reason) + admit() + return approvalRejections(t, reason) - before +} + +// approvalRejections reads the counter out of the registry a scraper would read +// it from, which is the only handle a test in this package has on a metric +// declared in another. +func approvalRejections(t *testing.T, reason string) int { + t.Helper() + families, err := ctrlmetrics.Registry.Gather() + if err != nil { + t.Fatalf("gather the metrics: %v", err) + } + for _, family := range families { + if family.GetName() != approvalRejectionsMetric { + continue + } + for _, metric := range family.GetMetric() { + labels := map[string]string{} + for _, pair := range metric.GetLabel() { + labels[pair.GetName()] = pair.GetValue() + } + if labels["namespace"] == testDeploymentNamespace && labels["reason"] == reason { + return int(metric.GetCounter().GetValue()) + } + } + } + return 0 +} diff --git a/operator/internal/webhook/clusterdeploymentconfig_validator.go b/operator/internal/webhook/clusterdeploymentconfig_validator.go index 65ec95b61..bfc92ba7a 100644 --- a/operator/internal/webhook/clusterdeploymentconfig_validator.go +++ b/operator/internal/webhook/clusterdeploymentconfig_validator.go @@ -90,13 +90,20 @@ func (v *ClusterDeploymentConfigValidator) Handle( return admission.Errored(http.StatusBadRequest, err) } + namespace := config.Namespace + if namespace == "" { + // A namespaced object created through a namespaced endpoint may arrive + // with the field unset, because the path carries it instead. + namespace = req.Namespace + } + if req.Operation == admissionv1.Update { old := &simplyblockv1alpha2.ClusterDeploymentConfig{} if err := v.Decoder.DecodeRaw(req.OldObject, old); err != nil { return admission.Errored(http.StatusBadRequest, err) } if old.Spec.Approved { - return afterApproval(old, config) + return afterApproval(namespace, old, config) } } @@ -105,13 +112,6 @@ func (v *ClusterDeploymentConfigValidator) Handle( return admission.Allowed("") } - namespace := config.Namespace - if namespace == "" { - // A namespaced object created through a namespaced endpoint may arrive - // with the field unset, because the path carries it instead. - namespace = req.Namespace - } - problems, err := v.checkApproval(ctx, namespace, config) if err != nil { return admission.Errored(http.StatusInternalServerError, err) @@ -119,20 +119,42 @@ func (v *ClusterDeploymentConfigValidator) Handle( if len(problems) == 0 { return admission.Allowed("") } + + messages := make([]string, 0, len(problems)) + for _, problem := range problems { + deployment.CountApprovalRejection(namespace, problem.reason) + messages = append(messages, problem.message) + } return admission.Denied(fmt.Sprintf( "this document cannot be approved, and approving it is what makes it "+ - "immutable: %s", strings.Join(problems, "; "))) + "immutable: %s", strings.Join(messages, "; "))) +} + +// problem is one thing wrong with an approval, carrying the reason it is counted +// under as well as the sentence the reviewer reads. +// +// The reason is the vocabulary the controller's own validation events use, for +// the four checks both perform, so the two counters of §9.2 can be read against +// each other: a rejection under a reason the draft never reported is a gap in +// the draft's validation rather than a reviewer's slip. +type problem struct { + reason string + message string } // afterApproval restates §3.2 for a document that is already approved. -func afterApproval(old, config *simplyblockv1alpha2.ClusterDeploymentConfig) admission.Response { +func afterApproval( + namespace string, old, config *simplyblockv1alpha2.ClusterDeploymentConfig, +) admission.Response { if !config.Spec.Approved { + deployment.CountApprovalRejection(namespace, deployment.ApprovalWithdrawn) return admission.Denied( "spec.approved cannot be withdrawn: un-approving a document does not " + "un-expand it, and the cluster and nodes it produced are removed by " + "deleting them rather than by editing the record that describes them") } if !equality.Semantic.DeepEqual(old.Spec, config.Spec) { + deployment.CountApprovalRejection(namespace, deployment.SpecImmutable) return admission.Denied( "spec is immutable once spec.approved is true, because the document is " + "then the record of what was deployed. To add nodes to the cluster " + @@ -151,27 +173,30 @@ func afterApproval(old, config *simplyblockv1alpha2.ClusterDeploymentConfig) adm // say so once rather than over two applies, each of which costs another approval. func (v *ClusterDeploymentConfigValidator) checkApproval( ctx context.Context, namespace string, config *simplyblockv1alpha2.ClusterDeploymentConfig, -) ([]string, error) { - var problems []string +) ([]problem, error) { + var problems []problem missing, err := v.missingWorkers(ctx, config) if err != nil { return nil, err } if len(missing) > 0 { - problems = append(problems, fmt.Sprintf( - "a group names %s, which %s not %s of this Kubernetes cluster", - strings.Join(missing, ", "), - plural(len(missing), "is", "are"), plural(len(missing), "a node", "nodes"))) + problems = append(problems, problem{ + reason: deployment.WorkerNotFound, + message: fmt.Sprintf( + "a group names %s, which %s not %s of this Kubernetes cluster", + strings.Join(missing, ", "), + plural(len(missing), "is", "are"), plural(len(missing), "a node", "nodes")), + }) } - cluster, problem, err := v.resolveCluster(ctx, namespace, config) + cluster, refused, err := v.resolveCluster(ctx, namespace, config) if err != nil { return nil, err } switch { - case problem != "": - problems = append(problems, problem) + case refused.reason != "": + problems = append(problems, refused) case cluster != nil: // A growth document. The class the groups name has to be the one the @@ -179,7 +204,9 @@ func (v *ClusterDeploymentConfigValidator) checkApproval( // time by StorageNodeValidator, which is a slower way to learn it and // leaves a half-expanded deployment behind. if mismatch := classMismatch(config, cluster); mismatch != "" { - problems = append(problems, mismatch) + problems = append(problems, problem{ + reason: deployment.DeviceClassMismatch, message: mismatch, + }) } default: @@ -190,11 +217,14 @@ func (v *ClusterDeploymentConfigValidator) checkApproval( return nil, err } if owner != "" { - problems = append(problems, fmt.Sprintf( - "ClusterDeploymentConfig %s is approved and creates StorageCluster %s "+ - "as well; both would race to create it and the loser is an immutable "+ - "Failed document, so add to it with spec.clusterRef instead", - owner, deployment.TargetClusterName(config))) + problems = append(problems, problem{ + reason: deployment.ClusterExists, + message: fmt.Sprintf( + "ClusterDeploymentConfig %s is approved and creates StorageCluster %s "+ + "as well; both would race to create it and the loser is an immutable "+ + "Failed document, so add to it with spec.clusterRef instead", + owner, deployment.TargetClusterName(config)), + }) } } @@ -206,35 +236,44 @@ func (v *ClusterDeploymentConfigValidator) checkApproval( // own, and a problem for the two combinations the expansion refuses. func (v *ClusterDeploymentConfigValidator) resolveCluster( ctx context.Context, namespace string, config *simplyblockv1alpha2.ClusterDeploymentConfig, -) (*simplyblockv1alpha2.StorageCluster, string, error) { +) (*simplyblockv1alpha2.StorageCluster, problem, error) { name := deployment.TargetClusterName(config) if name == "" { - return nil, "the document names neither a cluster to create in spec.cluster " + - "nor one to add nodes to in spec.clusterRef, so it describes no deployment", nil + return nil, problem{ + reason: deployment.NoClusterNamed, + message: "the document names neither a cluster to create in spec.cluster " + + "nor one to add nodes to in spec.clusterRef, so it describes no deployment", + }, nil } var cluster simplyblockv1alpha2.StorageCluster err := v.Client.Get(ctx, client.ObjectKey{Namespace: namespace, Name: name}, &cluster) found := err == nil if err != nil && !apierrors.IsNotFound(err) { - return nil, "", fmt.Errorf("reading StorageCluster %s: %w", name, err) + return nil, problem{}, fmt.Errorf("reading StorageCluster %s: %w", name, err) } switch { case config.Spec.ClusterRef != "" && !found: - return nil, fmt.Sprintf( - "spec.clusterRef names StorageCluster %s, and there is none by that name "+ - "in namespace %s", name, namespace), nil + return nil, problem{ + reason: deployment.ClusterNotFound, + message: fmt.Sprintf( + "spec.clusterRef names StorageCluster %s, and there is none by that name "+ + "in namespace %s", name, namespace), + }, nil case config.Spec.ClusterRef == "" && found: - return nil, fmt.Sprintf( - "spec.cluster.name is %s and a StorageCluster by that name already exists; "+ - "set spec.clusterRef to add nodes to it instead", name), nil + return nil, problem{ + reason: deployment.ClusterExists, + message: fmt.Sprintf( + "spec.cluster.name is %s and a StorageCluster by that name already exists; "+ + "set spec.clusterRef to add nodes to it instead", name), + }, nil case found: - return &cluster, "", nil + return &cluster, problem{}, nil } - return nil, "", nil + return nil, problem{}, nil } // classMismatch reports a growth document whose groups name devices of a class From 6f23852493d3934c6740574d21f30d84b78c4def Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 16:05:19 +0200 Subject: [PATCH 055/206] refactor(nodeprobe): the reference formula is the shared one MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit design-api-upgrade §19.6 names nodeprobe.ObjectName as the truncate-and-hash reference and says the shared helper is extracted from it rather than written twice. The helper exists and the upgrade tooling's whole derivation catalog is built on it; the call site it was extracted from still carried its own regex, its own length arithmetic, and its own SHA-256. It is now one line over a kube.Formula, and the names it produces are byte for byte the ones it produced before. That second half is the work. Adopting the helper as it stood would have renamed every object the probe writes, because Formula omits the digest when the natural name fits and this formula spent it unconditionally. A deduplication that changes behavior is not a deduplication, and here the behavior matters: the probe is a container image a cluster may still be running from before an operator upgrade, so a changed name means the probe writes its report under one name while the operator looks for a Job under another, and the run starts a second probe on every worker. The unconditional digest is load-bearing rather than incidental. The run and the node join on a separator both of them may contain, so run oops-1 on worker-3 and run oops on 1-worker-3 reach one stem with neither anywhere near the limit. atlas-lib already knew this: §19.8's ambiguous concatenation has a test, and it proves the two digests differ while leaving the two values equal. Formula gained AlwaysDigest, which makes the digest part of the formula's natural output rather than part of the bounding, so Value still equals Natural for an input the limit never touched and Truncated still says something was cut. Where a formula's parts are unambiguous it stays off and the name stays readable. The evidence that nothing moved is a table of six names pinned from the old implementation before the swap. It was green before and is green after, which is all it is for: it cannot be red for a defect because there is none, and it exists to catch one being introduced. labelValue stays hand-rolled and now says why. It answers the opposite half of the same question: a selector value is allowed to collide, two workers whose names differ past the sixty-third character selecting together is what a selector is for, and giving it a digest would make it unique and useless. Co-Authored-By: Claude Opus 5 (1M context) --- atlas-lib/kube/derived.go | 28 +++++++-- atlas-lib/kube/derived_test.go | 53 ++++++++++++++++ .../crd-redesign/design-api-upgrade.md | 12 +++- operator/internal/nodeprobe/configmap.go | 61 +++++++++++-------- operator/internal/nodeprobe/configmap_test.go | 40 ++++++++++++ 5 files changed, 160 insertions(+), 34 deletions(-) diff --git a/atlas-lib/kube/derived.go b/atlas-lib/kube/derived.go index dfc2b8dbf..63d24c37c 100644 --- a/atlas-lib/kube/derived.go +++ b/atlas-lib/kube/derived.go @@ -125,6 +125,18 @@ type Formula struct { // copied into its pods as a label, so a formula that names a Job sets 63 // here while remaining an ObjectName. Limit int + + // AlwaysDigest makes the digest part of every value the formula produces, + // rather than only of the ones the limit had to cut. + // + // It is for a formula whose parts can join into one stem from different + // inputs, which is §19.8's ambiguous concatenation: - reads the + // same for run a-b with node c as for run a with node b-c, and truncation + // is not what makes those two collide. The digest is taken over the parts + // rather than over the stem, so spending it unconditionally is what tells + // them apart. A formula whose parts are already unique as a joined string + // leaves it off and keeps a readable name. + AlwaysDigest bool } // Derived is one identifier a [Formula] produced, carrying enough to report a @@ -194,10 +206,18 @@ func (f Formula) Derive(parts ...string) Derived { stem := strings.Join(sanitized, separator) derived := Derived{ - Natural: f.Prefix + stem + f.Suffix, - Digest: digestOf(parts), - Limit: limit, - Kind: f.Kind, + Digest: digestOf(parts), + Limit: limit, + Kind: f.Kind, + } + // With AlwaysDigest the digest is part of the formula rather than part of + // the bounding, so it is in the natural name too. That keeps the two + // properties Derived promises: Value equals Natural for an input the limit + // never touched, and Truncated says something was cut rather than that a + // digest is present. + derived.Natural = f.Prefix + stem + f.Suffix + if f.AlwaysDigest { + derived.Natural = f.Prefix + stem + "-" + derived.Digest + f.Suffix } if len(derived.Natural) <= limit { derived.Value = derived.Natural diff --git a/atlas-lib/kube/derived_test.go b/atlas-lib/kube/derived_test.go index a3f667a6c..9a6128804 100644 --- a/atlas-lib/kube/derived_test.go +++ b/atlas-lib/kube/derived_test.go @@ -131,3 +131,56 @@ func TestFormula_SuffixAndSeparator(t *testing.T) { t.Fatalf("Value = %q, want the parts joined on the formula's separator", got.Value) } } + +// A formula whose parts join ambiguously needs the digest on every value it +// produces, not only on the ones that had to be cut. The test above proves the +// digest tells the two inputs apart; this one is about the name actually using +// it, which is the whole of the difference between a collision resolved and a +// collision merely detectable. +func TestFormula_AlwaysDigestSeparatesShortAmbiguousInputs(t *testing.T) { + f := Formula{Kind: ObjectName, Prefix: "sb-", AlwaysDigest: true} + + a, b := f.Derive("a-b", "c"), f.Derive("a", "b-c") + if a.Value == b.Value { + t.Fatalf("a-b/c and a/b-c both derived %q, and neither was truncated", a.Value) + } + if a.Truncated || b.Truncated { + t.Error("a name well inside the limit is reported as truncated") + } + if !a.Fits() { + t.Error("Fits() is false for a name the limit never bound") + } +} + +// The digest is part of the formula's output, so the value it produces for an +// input that fits is the value, rather than something the bounding rewrote. +func TestFormula_AlwaysDigestIsPartOfTheNaturalName(t *testing.T) { + f := Formula{Kind: ObjectName, Prefix: "sb-", AlwaysDigest: true} + + got := f.Derive("oops-1", "worker-3") + if got.Value != got.Natural { + t.Errorf("Value is %q and Natural is %q; nothing was cut, so they are one name", + got.Value, got.Natural) + } + if !strings.HasSuffix(got.Value, "-"+got.Digest) { + t.Errorf("%q does not carry the digest %q", got.Value, got.Digest) + } +} + +// And it still fits. The digest is spent out of the budget whether or not the +// stem needed cutting, which is what stops a formula from producing a legal name +// for a short input and an over-long one for a middling input. +func TestFormula_AlwaysDigestHoldsToTheLimit(t *testing.T) { + f := Formula{Kind: ObjectName, Prefix: "sb-nodeprobe-", Limit: 63, AlwaysDigest: true} + + for _, length := range []int{1, 30, 40, 41, 42, 50, 200} { + got := f.Derive("run", strings.Repeat("n", length)) + if len(got.Value) > 63 { + t.Errorf("an input of %d derived %q, which is %d bytes", + length, got.Value, len(got.Value)) + } + if errs := validation.IsDNS1123Subdomain(got.Value); len(errs) != 0 { + t.Errorf("an input of %d derived %q: %v", length, got.Value, errs) + } + } +} diff --git a/operator/docs/designs/crd-redesign/design-api-upgrade.md b/operator/docs/designs/crd-redesign/design-api-upgrade.md index fa37d74ef..e66c0eb53 100644 --- a/operator/docs/designs/crd-redesign/design-api-upgrade.md +++ b/operator/docs/designs/crd-redesign/design-api-upgrade.md @@ -1481,9 +1481,15 @@ adopt rather than a new invention: - **The Helm chart** names every resource literally rather than building names from the release name. -`nodeprobe.ObjectName` lands with the discovery work, so the shared helper is -extracted from it rather than written twice, into `atlas-lib/kube/names.go` -beside the formulas it bounds. +`nodeprobe.ObjectName` landed with the discovery work, the shared helper was +extracted from it rather than written twice, into `atlas-lib/kube/derived.go` +beside the formulas it bounds, and the reference call site now derives its names +through it. What the extraction had to carry over is that the digest is +unconditional there: the run and the node join on a separator both of them may +contain, so two runs of one deployment reach one stem without either being long +enough to truncate. Where a formula's parts are unambiguous the digest stays +conditional and the name stays readable, so `Formula.AlwaysDigest` is what the +two cases differ in. ### 19.7 Where a Rule Is Enforced diff --git a/operator/internal/nodeprobe/configmap.go b/operator/internal/nodeprobe/configmap.go index ef8f5f622..bc85e56b3 100644 --- a/operator/internal/nodeprobe/configmap.go +++ b/operator/internal/nodeprobe/configmap.go @@ -8,7 +8,12 @@ // kubectl by whoever is trying to work out why their disk was not a candidate. // // The name is derived rather than looked up so that a probe pod restarted by -// its Job writes over its own report instead of leaving two. +// its Job writes over its own report instead of leaving two. It is derived by +// atlas-lib's kube.Formula, which this call site is the reference for: +// design-api-upgrade.md §19.6 names the truncate-and-hash rule here as the +// pattern the shared helper was extracted from, and a formula that lived in two +// places would let the probe and the operator disagree about what an object is +// called. // // Neither the name nor the labels can be read back as the values that produced // them. Both are sanitized, and both are truncated to the 63 characters a label @@ -22,14 +27,14 @@ package nodeprobe import ( - "crypto/sha256" - "encoding/hex" "fmt" "regexp" "strings" corev1 "k8s.io/api/core/v1" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + + "github.com/simplyblock/atlas/kube" ) const ( @@ -52,15 +57,7 @@ const ( // namePrefix opens every object name the probe generates. namePrefix = "sb-nodeprobe-" - // nameHashLength is how much of the digest a generated name carries. Eight - // hex characters is 32 bits, which is not a cryptographic claim: the digest - // is there to keep two truncated node names apart, and the pair it - // disambiguates is always within one namespace and one run. - nameHashLength = 8 - - // maxNameLength is the budget every generated name is held to, and - // maxStemLength leaves room for the prefix, the digest, and the dash - // between them. + // maxNameLength is the budget every generated name is held to. // // It is 63 and not the 253 a ConfigMap's name may be, because the Job that // writes the report is named the same way and a Job's name is copied into @@ -70,14 +67,20 @@ const ( // the run then stalls on the first worker whose name is long, which is any // worker a cloud named after its fully qualified domain name. maxNameLength = 63 - maxStemLength = maxNameLength - len(namePrefix) - nameHashLength - 1 ) -// unsafeForName matches everything a DNS subdomain may not carry. A node is -// named ip-10-0-1-23.eu-central-1.compute.internal on one cloud and -// worker_3 on somebody's laboratory, and only the first of those is already a -// legal object name. -var unsafeForName = regexp.MustCompile(`[^a-z0-9.-]+`) +// nameFormula is how every object the probe generates is named. +// +// It carries its digest unconditionally. The run and the node join on a +// separator both of them may contain, so run oops-1 on worker-3 and run oops on +// 1-worker-3 reach one stem without either being long enough to truncate, and +// the digest is taken over the parts rather than over the stem. +var nameFormula = kube.Formula{ + Kind: kube.ObjectName, + Prefix: namePrefix, + Limit: maxNameLength, + AlwaysDigest: true, +} // ObjectName is the ConfigMap a run's report for one node goes into. // @@ -86,14 +89,7 @@ var unsafeForName = regexp.MustCompile(`[^a-z0-9.-]+`) // name writes new objects rather than overwriting reports somebody may have // already read. func ObjectName(run, node string) string { - stem := unsafeForName.ReplaceAllString(strings.ToLower(run+"-"+node), "-") - stem = strings.Trim(stem, ".-") - if len(stem) > maxStemLength { - stem = strings.TrimRight(stem[:maxStemLength], ".-") - } - - digest := sha256.Sum256([]byte(run + "\x00" + node)) - return namePrefix + stem + "-" + hex.EncodeToString(digest[:])[:nameHashLength] + return nameFormula.Derive(run, node).Value } // ConfigMap renders a report as the object the probe writes. @@ -174,6 +170,11 @@ func ReportSelector(run string) map[string]string { } } +// unsafeForLabel matches everything these label values may not carry. A node is +// named ip-10-0-1-23.eu-central-1.compute.internal on one cloud and worker_3 in +// somebody's laboratory, and the second needs rewriting. +var unsafeForLabel = regexp.MustCompile(`[^a-z0-9.-]+`) + // maxLabelValueLength is the length a label value may have. const maxLabelValueLength = 63 @@ -184,8 +185,14 @@ const maxLabelValueLength = 63 // selecting a run's reports and the ConfigMap's own data is what says which // node a report is about: the label may have lost the end of the name, and the // report has not. +// +// It is deliberately not a kube.Formula, which is the other half of the same +// question the name above answers with one. A formula keeps two inputs apart +// and this value is allowed to collide: two workers of one rack whose names +// differ past the sixty-third character select together, which is what a +// selector is for. Giving it a digest would make it unique and useless. func labelValue(v string) string { - v = unsafeForName.ReplaceAllString(strings.ToLower(v), "-") + v = unsafeForLabel.ReplaceAllString(strings.ToLower(v), "-") v = strings.Trim(v, ".-_") if len(v) > maxLabelValueLength { v = strings.Trim(v[:maxLabelValueLength], ".-_") diff --git a/operator/internal/nodeprobe/configmap_test.go b/operator/internal/nodeprobe/configmap_test.go index 74baa0d42..3407cfd91 100644 --- a/operator/internal/nodeprobe/configmap_test.go +++ b/operator/internal/nodeprobe/configmap_test.go @@ -199,3 +199,43 @@ func minimalInventory() inventory.Inventory { Devices: []blockdev.Candidate{oneFreeDisk()}, } } + +// The names this formula has always produced, pinned. +// +// Two processes compute them without telling each other, and one of the two is +// a container image that a cluster may still be running from before an operator +// upgrade. A change here is therefore not a refactor: the probe would write its +// report under one name while the operator looked for a Job under another, and +// the run would start a second probe on every worker. +// +// It was green before the formula moved to atlas-lib and is green after, which +// is the whole of what it is for. It cannot be red for a defect, because there +// is none — it exists to catch one being introduced. +func TestObjectNameStillProducesTheNamesItAlwaysHas(t *testing.T) { + long := strings.Repeat("compute-node-with-a-long-name.", 10) + "worker-3.internal" + + for _, tc := range []struct{ run, node, want string }{ + {"oops-1", "worker-3", "sb-nodeprobe-oops-1-worker-3-089a1802"}, + {"oops-1", "ip-10-0-1-23.eu-central-1.compute.internal", + "sb-nodeprobe-oops-1-ip-10-0-1-23.eu-central-1.compute-73596ea7"}, + {"a", "b", "sb-nodeprobe-a-b-59b271ae"}, + {"discovered-run", "WORKER_3", "sb-nodeprobe-discovered-run-worker-3-4701a101"}, + {strings.Repeat("run", 40), strings.Repeat("node.", 60), + "sb-nodeprobe-runrunrunrunrunrunrunrunrunrunrunrunrunru-ea56b4cc"}, + {"oops-1", long, "sb-nodeprobe-oops-1-compute-node-with-a-long-name.comp-99504060"}, + } { + if got := ObjectName(tc.run, tc.node); got != tc.want { + t.Errorf("ObjectName(%q, ...) is now %q, and was %q", tc.run, got, tc.want) + } + } +} + +// Every name carries the digest, including the ones that were never near the +// limit. Sanitization is lossy and the run and the node join on a separator +// both may contain, so two runs of one deployment can reach the same stem, and +// the digest is the only thing that separates them. +func TestObjectNameCarriesTheDigestEvenWhenItFits(t *testing.T) { + if a, b := ObjectName("oops-1", "worker-3"), ObjectName("oops", "1-worker-3"); a == b { + t.Errorf("run oops-1 on worker-3 and run oops on 1-worker-3 both wrote %q", a) + } +} From acd948f14174453f0bb5debc42b640c284107236 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 17:22:32 +0200 Subject: [PATCH 056/206] refactor(chart): the kubelet toggle belongs to the operator skipKubeletConfiguration fed exactly one place, the SKIP_KUBELET_CONFIGURATION environment variable on the chart-rendered storage-node DaemonSet, and that template renders only under storagenode.create, which is false and stays false. The operator renders storage nodes, and an OpenShift marketplace install brings the operator alone. The toggle itself is alive under a different owner. v1alpha2 carries no per-node kubelet field; it re-landed as StorageCluster.spec.storageNodes.enableKubeletConfiguration, and the operator's renderer inverts it back into the same environment variable (internal/utils/storage_node_workload.go). How it resolves is a property of the Kubernetes distribution rather than of an install-time answer, so ClusterDeploymentConfig.spec.environment decides it: OpenShift and the vanilla distributions configure the kubelet, Talos does not, because it has no writable kubelet configuration. What is left in values.yaml is a second place to answer a question the operator already answers, and the one nothing reads. helm template at default values is byte-identical. The only difference is under storagenode.create=true, where the two legacy DaemonSets lose the variable. Co-Authored-By: Claude Opus 5 (1M context) --- .../charts/simplyblock-operator/templates/storage-node.yaml | 2 -- helm-charts/charts/simplyblock-operator/values.yaml | 1 - 2 files changed, 3 deletions(-) diff --git a/helm-charts/charts/simplyblock-operator/templates/storage-node.yaml b/helm-charts/charts/simplyblock-operator/templates/storage-node.yaml index 915135bde..0fd141cb0 100644 --- a/helm-charts/charts/simplyblock-operator/templates/storage-node.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/storage-node.yaml @@ -175,8 +175,6 @@ spec: value: "{{ $.Values.storagenode.openShiftCluster }}" - name: CPU_TOPOLOGY_ENABLED value: "{{ $.Values.storagenode.enableCpuTopology }}" - - name: SKIP_KUBELET_CONFIGURATION - value: "{{ $.Values.storagenode.skipKubeletConfiguration }}" {{- if $.Values.storagenode.reservedSystemCpu }} - name: RESERVED_SYSTEM_CPUS value: "{{ $.Values.storagenode.reservedSystemCpu }}" diff --git a/helm-charts/charts/simplyblock-operator/values.yaml b/helm-charts/charts/simplyblock-operator/values.yaml index 53c4f2612..c9b6a3c6f 100644 --- a/helm-charts/charts/simplyblock-operator/values.yaml +++ b/helm-charts/charts/simplyblock-operator/values.yaml @@ -111,7 +111,6 @@ storagenode: ubuntuHost: false enableCpuTopology: false enableDevicePlugin: true - skipKubeletConfiguration: false openShiftCluster: false reservedSystemCpu: multiCluster: From e44a26e8c014b162647acab4992258a9ff7227fc Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 17:24:58 +0200 Subject: [PATCH 057/206] refactor(chart): the event-log alert rules leave the chart MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The chart provisioned one Infinity data source per cluster and seven Grafana rules that query the control plane's event log through it, and it cannot build either. A data source carries one cluster's credential, so the rules iterate the clusters, and simplyblock.eventAlertClusters read that list out of storagenode.multiCluster.clusters or the csiConfig and csiSecret pair. Both are answers somebody types into a values file after `cluster create` has produced a UUID and a secret, which is the loop the ClusterDeploymentConfig replaced. An OpenShift marketplace install brings the operator alone, so there is no values file to type them into at all. Nothing is moved here, because there is nowhere yet to move it to. design-controlplane.md §12 Q8 records the requirement: the StorageCluster reconciler already holds every cluster's UUID and secret and upserts them into simplyblock-csi-secret-v2, so the provisioning belongs to whatever reconciles the observability stack, and Q2 still leaves Grafana with the chart. A standalone deployment alerts on what Thanos scrapes until that is answered, which is the half of the alerting that needs no per-cluster credential. helm template is byte-identical at default values and with observability enabled, because eventAlerts defaulted to off. With it on, the render loses 889 lines and gains none: the data source, the seven rules, and the GF_INSTALL_PLUGINS variable that fetched the Infinity plugin on every Grafana pod start. Co-Authored-By: Claude Opus 5 (1M context) --- .../templates/_helpers.tpl | 32 - .../templates/controlplane_configmap.yaml | 890 ------------------ .../templates/controlplane_deploy.yaml | 14 - .../charts/simplyblock-operator/values.yaml | 32 - .../crd-redesign/design-controlplane.md | 23 + 5 files changed, 23 insertions(+), 968 deletions(-) diff --git a/helm-charts/charts/simplyblock-operator/templates/_helpers.tpl b/helm-charts/charts/simplyblock-operator/templates/_helpers.tpl index 932c87aab..fadac54ea 100644 --- a/helm-charts/charts/simplyblock-operator/templates/_helpers.tpl +++ b/helm-charts/charts/simplyblock-operator/templates/_helpers.tpl @@ -53,38 +53,6 @@ http://simplyblock-webappapi.{{ .Release.Namespace }}.svc.cluster.local:5000 {{- end -}} {{- end -}} -{{/* -The clusters whose event log the Grafana event-driven alert rules read, as a -JSON array of {"id","secret"} objects for `fromJsonArray`. Both the Infinity -data sources and the rules that query them iterate this, so the two can never -disagree about which clusters exist. - -A cluster with no id or no secret is skipped rather than rendered half-configured. -The two are used for different halves of the same request: the secret is the -whole credential, sent as the bearer token that /api/v2 matches against every -cluster's secret, while the id addresses the cluster in the request path. The -API then checks that the two agree, so a half-configured or mismatched entry -fails every evaluation with a 401 that reads like an outage rather than like a -missing value. The list is empty until `cluster create` has run and its UUID and -secret have been fed back into the values, which is the normal state right after -install. -*/}} -{{- define "simplyblock.eventAlertClusters" -}} -{{- $out := list -}} -{{- if .Values.storagenode.multiCluster.enable -}} -{{- range default (list) .Values.storagenode.multiCluster.clusters -}} -{{- if and .cluster_id .secret -}} -{{- $out = append $out (dict "id" .cluster_id "secret" .secret) -}} -{{- end -}} -{{- end -}} -{{- else -}} -{{- if and .Values.csiConfig.simplybk.uuid .Values.csiSecret.simplybk.secret -}} -{{- $out = append $out (dict "id" .Values.csiConfig.simplybk.uuid "secret" .Values.csiSecret.simplybk.secret) -}} -{{- end -}} -{{- end -}} -{{- toJson $out -}} -{{- end -}} - {{/* Volume named "tls" holding the serving cert bundle for pods that terminate TLS. Args: dict "ctx" $root "secret" diff --git a/helm-charts/charts/simplyblock-operator/templates/controlplane_configmap.yaml b/helm-charts/charts/simplyblock-operator/templates/controlplane_configmap.yaml index c1b93c26a..7f467e736 100644 --- a/helm-charts/charts/simplyblock-operator/templates/controlplane_configmap.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/controlplane_configmap.yaml @@ -174,42 +174,6 @@ data: access: proxy uid: PBFA97CFB590B2093 editable: true -{{- $eventAlerts := .Values.controlplane.observability.grafana.eventAlerts | default dict }} -{{- $eventClusters := include "simplyblock.eventAlertClusters" . | fromJsonArray }} -{{- if and $eventAlerts.enabled $eventClusters }} -{{- $cpAddr := include "simplyblock.controlPlaneAddr" . }} - # One Infinity data source per cluster. The v2 API authenticates with the - # cluster secret as a bearer token, and Grafana binds credentials to a data - # source rather than to a query, so a cluster needs its own. The secret goes - # through secureJsonData, which Grafana stores encrypted and never returns - # to the browser; Infinity reads the pair as a custom request header. -{{- range $eventClusters }} - - name: Simplyblock Events {{ .id }} - type: yesoreyeram-infinity-datasource - uid: {{ printf "sbev-%s" (sha256sum .id | trunc 12) }} - url: {{ $cpAddr }} - access: proxy - editable: false - jsonData: - httpHeaderName1: Authorization - # Infinity refuses every URL query until the host is allowlisted. The - # refusal reaches an operator as a per-rule evaluation error reading - # "datasource is missing allowed hosts/URLs" rather than as a - # configuration warning, so every rule below fails in a way that looks - # like an outage instead of like a missing setting. The backend reads - # the list as jsonData.allowedHosts and matches a URL by prefix, so the - # control-plane base address covers every /api/v2 path they query. - allowedHosts: - - {{ $cpAddr }} - # Without this the health check probes the bare base URL, which the - # control plane answers with a string Infinity cannot read as a table. - customHealthCheckEnabled: false - tlsSkipVerify: true - timeoutInSeconds: 30 - secureJsonData: - httpHeaderValue1: {{ printf "Bearer %s" .secret | quote }} -{{- end }} -{{- end }} --- apiVersion: v1 kind: ConfigMap @@ -879,860 +843,6 @@ data: labels: app: simplyblock isPaused: false -{{- if and $eventAlerts.enabled $eventClusters }} -{{- $logLimit := $eventAlerts.logLimit | default 1000 }} -{{- $for := $eventAlerts.for | default "1m" }} - - # Rules derived from the control plane's cluster event log rather than - # from Thanos. Each query folds the newest log records into the current - # state of every node, device, and cluster and returns one row per entity - # that is currently wrong. Healing is therefore structural: the reverting - # event -- a node back online, a device back online, a cluster back to - # active -- removes the row, the row's absence removes the alert - # instance, and noDataState: OK turns an entirely empty result into the - # resolved state. - # - # The fold has to read the transitions because the cause is not in the - # payload. cluster_ops.get_logs drops caused_by, and set_node_status - # writes every transition under the same default caused_by regardless. - # What separates a shutdown an operator asked for from a fault is the - # state it passes through -- only shutdown_storage_node reaches - # in_shutdown -- and that is what the two suppression rules key on. - - orgId: 1 - name: simplyblock_events - folder: grafana_folder - interval: {{ $eventAlerts.interval | default "1m" }} - rules: -{{- range $eventClusters }} -{{- $cid := .id }} -{{- $dsUid := printf "sbev-%s" (sha256sum .id | trunc 12) }} - - uid: {{ printf "sbev-%s-%s" (sha256sum $cid | trunc 8) "node-left-online" }} - title: StorageNode_left_online_{{ $cid }} - condition: B - data: - - refId: A - relativeTimeRange: - from: 600 - to: 0 - datasourceUid: {{ $dsUid }} - model: - refId: A - type: json - source: url - format: table - # Only the backend parser runs server-side, which is what - # an alert rule needs. root_selector is tried as a gjson - # path first and falls back to JSONata; opening with a - # bracket guarantees the fallback. The v2 endpoint returns - # the records unwrapped, so $ is the record array itself. - parser: backend - url: /api/v2/clusters/{{ $cid }}/logs - url_options: - method: GET - params: - - key: limit - value: {{ $logLimit | quote }} - root_selector: |- - ( - $planned := ['in_shutdown', 'in_removal', 'pending_removal', 'removed']; - $healthy := ['online', 'in_creation']; - $ev := [$[Event = 'STATUS_CHANGE' and $contains(Message, 'Storage node status changed from: ')]]; - $fold := function($evs) { - $reduce($evs, function($a, $e) { - ( - $to := $substringAfter($e.Message, ' to: '); - { - 'state': $to, - 'since': $e.Date, - 'planned': $to in $planned or ($a.planned and $to = 'offline') - } - ) - }, { 'state': 'online', 'since': '', 'planned': false }) - }; - $nodes := $count($ev) = 0 ? [] : [$each($ev{ NodeId: [$] }, function($evs, $id) { - $merge([{ 'node': $id }, $fold($evs)]) - })]; - [$nodes[$not(state in $healthy) and $not(planned)].{ - 'node': node, - 'state': state, - 'since': since, - 'value': 1 - }] - ) - # Exactly one numeric field and the rest strings is what - # makes Grafana read the frame as a numeric table, turning - # every string field into a label of its own alert instance. - columns: - - selector: node - text: node - type: string - - selector: state - text: state - type: string - - selector: since - text: since - type: string - - selector: value - text: value - type: number - - refId: B - relativeTimeRange: - from: 600 - to: 0 - datasourceUid: __expr__ - model: - refId: B - type: threshold - datasource: - type: __expr__ - uid: __expr__ - expression: A - conditions: - - evaluator: - params: - - 0 - type: gt - operator: - type: and - query: - params: [] - reducer: - params: [] - type: last - type: query - # A row exists only for something currently wrong, so an empty - # result is the healthy state and has to resolve the alert. A - # failing query is not: that is the alert source being broken. - noDataState: OK - execErrState: Error - for: {{ $for }} - annotations: - summary: 'Storage node {{ "{{" }} $labels.node {{ "}}" }} left the online state and is now {{ "{{" }} $labels.state {{ "}}" }}.' - description: 'The rule folds the cluster event log into the current state of every storage node and fires while a node sits in anything other than online or in_creation. A shutdown or removal an operator asked for is exempt: those reach in_shutdown, pending_removal, in_removal, or removed, which no fault path does. The alert clears when the node returns to online.' - labels: - app: simplyblock - cluster: {{ $cid }} - severity: critical - isPaused: false - - uid: {{ printf "sbev-%s-%s" (sha256sum $cid | trunc 8) "device-unavailable" }} - title: Device_became_unavailable_{{ $cid }} - condition: B - data: - - refId: A - relativeTimeRange: - from: 600 - to: 0 - datasourceUid: {{ $dsUid }} - model: - refId: A - type: json - source: url - format: table - # Only the backend parser runs server-side, which is what - # an alert rule needs. root_selector is tried as a gjson - # path first and falls back to JSONata; opening with a - # bracket guarantees the fallback. The v2 endpoint returns - # the records unwrapped, so $ is the record array itself. - parser: backend - url: /api/v2/clusters/{{ $cid }}/logs - url_options: - method: GET - params: - - key: limit - value: {{ $logLimit | quote }} - root_selector: |- - ( - $window := 120000; - $ms := function($d) { $toMillis($substring($d & '.000', 0, 23), '[Y0001]-[M01]-[D01] [H01]:[m01]:[s01].[f001]') }; - $healthy := ['online', 'in_creation']; - $nodeEv := [$[Event = 'STATUS_CHANGE' and $contains(Message, 'Storage node status changed from: ')]]; - $cascades := [$map($nodeEv[$not($substringAfter(Message, ' to: ') in $healthy)], function($e) { $ms($e.Date) })]; - $devEv := [$[Event = 'STATUS_CHANGE' and $contains(Message, 'Device status changed from: ')]]; - $fold := function($evs) { - $reduce($evs, function($a, $e) { - { - 'state': $substringAfter($e.Message, ' to: '), - 'since': $e.Date, - 'at': $ms($e.Date), - 'order': $e.Storage_ID - } - }, { 'state': 'online', 'since': '', 'at': 0, 'order': '' }) - }; - $devs := $count($devEv) = 0 ? [] : [$each($devEv{ NodeId: [$] }, function($evs, $id) { - ( - $f := $fold($evs); - $merge([ - { 'device': $id, 'cascade': $count($cascades[$ <= $f.at and $f.at - $ <= $window]) > 0 }, - $f - ]) - ) - })]; - [$devs[state = 'unavailable' and $not(cascade)].{ - 'device': device, - 'order': order, - 'since': since, - 'value': 1 - }] - ) - # Exactly one numeric field and the rest strings is what - # makes Grafana read the frame as a numeric table, turning - # every string field into a label of its own alert instance. - columns: - - selector: device - text: device - type: string - - selector: order - text: order - type: string - - selector: since - text: since - type: string - - selector: value - text: value - type: number - - refId: B - relativeTimeRange: - from: 600 - to: 0 - datasourceUid: __expr__ - model: - refId: B - type: threshold - datasource: - type: __expr__ - uid: __expr__ - expression: A - conditions: - - evaluator: - params: - - 0 - type: gt - operator: - type: and - query: - params: [] - reducer: - params: [] - type: last - type: query - # A row exists only for something currently wrong, so an empty - # result is the healthy state and has to resolve the alert. A - # failing query is not: that is the alert source being broken. - noDataState: OK - execErrState: Error - for: {{ $for }} - annotations: - summary: 'Device {{ "{{" }} $labels.device {{ "}}" }} (cluster device order {{ "{{" }} $labels.order {{ "}}" }}) is unavailable.' - description: 'The rule fires for a device that is currently unavailable and did not become so as part of a whole node going down. A node leaving online cascades every one of its devices to unavailable within seconds, and that is reported once as the storage node alert rather than once per device. The alert clears when the device returns to online.' - labels: - app: simplyblock - cluster: {{ $cid }} - severity: warning - isPaused: false - - uid: {{ printf "sbev-%s-%s" (sha256sum $cid | trunc 8) "device-removed" }} - title: Device_removed_{{ $cid }} - condition: B - data: - - refId: A - relativeTimeRange: - from: 600 - to: 0 - datasourceUid: {{ $dsUid }} - model: - refId: A - type: json - source: url - format: table - # Only the backend parser runs server-side, which is what - # an alert rule needs. root_selector is tried as a gjson - # path first and falls back to JSONata; opening with a - # bracket guarantees the fallback. The v2 endpoint returns - # the records unwrapped, so $ is the record array itself. - parser: backend - url: /api/v2/clusters/{{ $cid }}/logs - url_options: - method: GET - params: - - key: limit - value: {{ $logLimit | quote }} - root_selector: |- - ( - $recent := 86400000; - $ms := function($d) { $toMillis($substring($d & '.000', 0, 23), '[Y0001]-[M01]-[D01] [H01]:[m01]:[s01].[f001]') }; - $now := $millis(); - $devEv := [$[Event = 'STATUS_CHANGE' and $contains(Message, 'Device status changed from: ')]]; - $fold := function($evs) { - $reduce($evs, function($a, $e) { - { - 'state': $substringAfter($e.Message, ' to: '), - 'since': $e.Date, - 'at': $ms($e.Date), - 'order': $e.Storage_ID - } - }, { 'state': 'online', 'since': '', 'at': 0, 'order': '' }) - }; - $devs := $count($devEv) = 0 ? [] : [$each($devEv{ NodeId: [$] }, function($evs, $id) { - $merge([{ 'device': $id }, $fold($evs)]) - })]; - [$devs[state = 'removed' and $now - at <= $recent].{ - 'device': device, - 'order': order, - 'since': since, - 'value': 1 - }] - ) - # Exactly one numeric field and the rest strings is what - # makes Grafana read the frame as a numeric table, turning - # every string field into a label of its own alert instance. - columns: - - selector: device - text: device - type: string - - selector: order - text: order - type: string - - selector: since - text: since - type: string - - selector: value - text: value - type: number - - refId: B - relativeTimeRange: - from: 600 - to: 0 - datasourceUid: __expr__ - model: - refId: B - type: threshold - datasource: - type: __expr__ - uid: __expr__ - expression: A - conditions: - - evaluator: - params: - - 0 - type: gt - operator: - type: and - query: - params: [] - reducer: - params: [] - type: last - type: query - # A row exists only for something currently wrong, so an empty - # result is the healthy state and has to resolve the alert. A - # failing query is not: that is the alert source being broken. - noDataState: OK - execErrState: Error - for: {{ $for }} - annotations: - summary: 'Device {{ "{{" }} $labels.device {{ "}}" }} (cluster device order {{ "{{" }} $labels.order {{ "}}" }}) was removed.' - description: 'The rule fires for every device whose latest state is removed, whatever the cause. It clears when the device is brought back online, and otherwise ages out a day after the removal, since a device that is genuinely gone has no reversion to wait for.' - labels: - app: simplyblock - cluster: {{ $cid }} - severity: critical - isPaused: false - - uid: {{ printf "sbev-%s-%s" (sha256sum $cid | trunc 8) "cluster-degraded" }} - title: Cluster_became_degraded_{{ $cid }} - condition: B - data: - - refId: A - relativeTimeRange: - from: 600 - to: 0 - datasourceUid: {{ $dsUid }} - model: - refId: A - type: json - source: url - format: table - # Only the backend parser runs server-side, which is what - # an alert rule needs. root_selector is tried as a gjson - # path first and falls back to JSONata; opening with a - # bracket guarantees the fallback. The v2 endpoint returns - # the records unwrapped, so $ is the record array itself. - parser: backend - url: /api/v2/clusters/{{ $cid }}/logs - url_options: - method: GET - params: - - key: limit - value: {{ $logLimit | quote }} - root_selector: |- - ( - $planned := ['in_shutdown', 'in_removal', 'pending_removal', 'removed']; - $healthy := ['online', 'in_creation']; - $nodeEv := [$[Event = 'STATUS_CHANGE' and $contains(Message, 'Storage node status changed from: ')]]; - $fold := function($evs) { - $reduce($evs, function($a, $e) { - ( - $to := $substringAfter($e.Message, ' to: '); - { - 'state': $to, - 'planned': $to in $planned or ($a.planned and $to = 'offline') - } - ) - }, { 'state': 'online', 'planned': false }) - }; - $nodes := $count($nodeEv) = 0 ? [] : [$each($nodeEv{ NodeId: [$] }, function($evs, $id) { $fold($evs) })]; - $operatorDown := $count($nodes[$not(state in $healthy) and planned]) > 0; - $cluEv := [$[Event = 'STATUS_CHANGE' and $contains(Message, 'Cluster status changed from ')]]; - $last := $cluEv[-1]; - $state := $exists($last) ? $substringAfter($last.Message, ' to ') : 'active'; - [$state = 'degraded' and $not($operatorDown) ? { - 'state': $state, - 'since': $last.Date, - 'value': 1 - }] - ) - # Exactly one numeric field and the rest strings is what - # makes Grafana read the frame as a numeric table, turning - # every string field into a label of its own alert instance. - columns: - - selector: state - text: state - type: string - - selector: since - text: since - type: string - - selector: value - text: value - type: number - - refId: B - relativeTimeRange: - from: 600 - to: 0 - datasourceUid: __expr__ - model: - refId: B - type: threshold - datasource: - type: __expr__ - uid: __expr__ - expression: A - conditions: - - evaluator: - params: - - 0 - type: gt - operator: - type: and - query: - params: [] - reducer: - params: [] - type: last - type: query - # A row exists only for something currently wrong, so an empty - # result is the healthy state and has to resolve the alert. A - # failing query is not: that is the alert source being broken. - noDataState: OK - execErrState: Error - for: {{ $for }} - annotations: - summary: 'The cluster is degraded.' - description: 'The rule fires while the cluster''s latest status is degraded. A degradation an operator caused by shutting a node down is exempt, decided by whether any node currently out of service got there through in_shutdown or a removal state. The alert clears when the cluster returns to active.' - labels: - app: simplyblock - cluster: {{ $cid }} - severity: warning - isPaused: false - - uid: {{ printf "sbev-%s-%s" (sha256sum $cid | trunc 8) "cluster-suspended" }} - title: Cluster_became_suspended_{{ $cid }} - condition: B - data: - - refId: A - relativeTimeRange: - from: 600 - to: 0 - datasourceUid: {{ $dsUid }} - model: - refId: A - type: json - source: url - format: table - # Only the backend parser runs server-side, which is what - # an alert rule needs. root_selector is tried as a gjson - # path first and falls back to JSONata; opening with a - # bracket guarantees the fallback. The v2 endpoint returns - # the records unwrapped, so $ is the record array itself. - parser: backend - url: /api/v2/clusters/{{ $cid }}/logs - url_options: - method: GET - params: - - key: limit - value: {{ $logLimit | quote }} - root_selector: |- - ( - $cluEv := [$[Event = 'STATUS_CHANGE' and $contains(Message, 'Cluster status changed from ')]]; - $last := $cluEv[-1]; - $state := $exists($last) ? $substringAfter($last.Message, ' to ') : 'active'; - [$state = 'suspended' ? { - 'state': $state, - 'since': $last.Date, - 'value': 1 - }] - ) - # Exactly one numeric field and the rest strings is what - # makes Grafana read the frame as a numeric table, turning - # every string field into a label of its own alert instance. - columns: - - selector: state - text: state - type: string - - selector: since - text: since - type: string - - selector: value - text: value - type: number - - refId: B - relativeTimeRange: - from: 600 - to: 0 - datasourceUid: __expr__ - model: - refId: B - type: threshold - datasource: - type: __expr__ - uid: __expr__ - expression: A - conditions: - - evaluator: - params: - - 0 - type: gt - operator: - type: and - query: - params: [] - reducer: - params: [] - type: last - type: query - # A row exists only for something currently wrong, so an empty - # result is the healthy state and has to resolve the alert. A - # failing query is not: that is the alert source being broken. - noDataState: OK - execErrState: Error - for: {{ $for }} - annotations: - summary: 'The cluster is suspended.' - description: 'The rule fires while the cluster''s latest status is suspended, whatever the cause, including a suspension that followed an operator''s node shutdown. The alert clears when the cluster leaves suspended.' - labels: - app: simplyblock - cluster: {{ $cid }} - severity: critical - isPaused: false - - uid: {{ printf "sbev-%s-%s" (sha256sum $cid | trunc 8) "cluster-capacity" }} - title: Cluster_capacity_reached_{{ $cid }} - condition: B - data: - - refId: A - relativeTimeRange: - from: 600 - to: 0 - datasourceUid: {{ $dsUid }} - model: - refId: A - type: json - source: url - format: table - # Only the backend parser runs server-side, which is what - # an alert rule needs. root_selector is tried as a gjson - # path first and falls back to JSONata; opening with a - # bracket guarantees the fallback. The v2 endpoint returns - # the records unwrapped, so $ is the record array itself. - parser: backend - url: /api/v2/clusters/{{ $cid }}/logs - url_options: - method: GET - params: - - key: limit - value: {{ $logLimit | quote }} - root_selector: |- - ( - $ms := function($d) { $toMillis($substring($d & '.000', 0, 23), '[Y0001]-[M01]-[D01] [H01]:[m01]:[s01].[f001]') }; - $now := $millis(); - $ev := [$[Event = 'CAPACITY']]; - $latest := $count($ev) = 0 ? [] : [$each( - $ev{ ($contains(Message, 'provisioned capacity') ? 'provisioned' : 'absolute'): [$] }, - function($evs, $kind) { - ( - $e := $evs[-1]; - $window := ($kind = 'absolute' and $e.Level = 'Critical') ? 1200000 : 180000; - { - 'kind': $kind, - 'severity': $e.Level, - 'value': $substringBefore($substringAfter($e.Message, 'reached: '), '%'), - 'fresh': $now - $ms($e.Date) <= $window - } - ) - })]; - [$latest[fresh and (severity = 'Warning' or severity = 'Critical')].{ - 'kind': kind, - 'severity': severity, - 'value': value - }] - ) - # Exactly one numeric field and the rest strings is what - # makes Grafana read the frame as a numeric table, turning - # every string field into a label of its own alert instance. - columns: - - selector: kind - text: kind - type: string - - selector: severity - text: severity - type: string - - selector: value - text: value - type: number - - refId: B - relativeTimeRange: - from: 600 - to: 0 - datasourceUid: __expr__ - model: - refId: B - type: threshold - datasource: - type: __expr__ - uid: __expr__ - expression: A - conditions: - - evaluator: - params: - - 0 - type: gt - operator: - type: and - query: - params: [] - reducer: - params: [] - type: last - type: query - # A row exists only for something currently wrong, so an empty - # result is the healthy state and has to resolve the alert. A - # failing query is not: that is the alert source being broken. - noDataState: OK - execErrState: Error - for: {{ $for }} - annotations: - summary: '{{ "{{" }} $labels.kind {{ "}}" }} cluster capacity is at {{ "{{" }} $value {{ "}}" }}% ({{ "{{" }} $labels.severity {{ "}}" }}).' - description: 'The rule reports the newest capacity event per kind, absolute and provisioned, and carries the utilization as its value. The control plane emits no back-to-normal event, so the alert clears by the event ceasing: the capacity monitor re-emits every 30 seconds while over a threshold, except an absolute critical which it throttles to once every 15 minutes, and the rule allows for both.' - labels: - app: simplyblock - cluster: {{ $cid }} - severity: warning - isPaused: false - - uid: {{ printf "sbev-%s-%s" (sha256sum $cid | trunc 8) "jm-records-threshold" }} - title: JM_records_threshold_exceeded_{{ $cid }} - condition: B - data: - - refId: A - relativeTimeRange: - from: 600 - to: 0 - datasourceUid: {{ $dsUid }} - model: - refId: A - type: json - source: url - format: table - # Only the backend parser runs server-side, which is what - # an alert rule needs. root_selector is tried as a gjson - # path first and falls back to JSONata; opening with a - # bracket guarantees the fallback. The v2 endpoint returns - # the records unwrapped, so $ is the record array itself. - parser: backend - url: /api/v2/clusters/{{ $cid }}/logs - url_options: - method: GET - params: - - key: limit - value: {{ $logLimit | quote }} - root_selector: |- - ( - $recent := 3600000; - $ms := function($d) { $toMillis($substring($d & '.000', 0, 23), '[Y0001]-[M01]-[D01] [H01]:[m01]:[s01].[f001]') }; - $now := $millis(); - $ev := [$[Event = 'JM_COMPRESSION_BACKLOG']]; - $latest := $count($ev) = 0 ? [] : [$each($ev{ NodeId: [$] }, function($evs, $id) { - ( - $e := $evs[-1]; - { 'node': $id, 'since': $e.Date, 'at': $ms($e.Date) } - ) - })]; - [$latest[$now - at <= $recent].{ - 'node': node, - 'since': since, - 'value': 1 - }] - ) - # Exactly one numeric field and the rest strings is what - # makes Grafana read the frame as a numeric table, turning - # every string field into a label of its own alert instance. - columns: - - selector: node - text: node - type: string - - selector: since - text: since - type: string - - selector: value - text: value - type: number - - refId: B - relativeTimeRange: - from: 600 - to: 0 - datasourceUid: __expr__ - model: - refId: B - type: threshold - datasource: - type: __expr__ - uid: __expr__ - expression: A - conditions: - - evaluator: - params: - - 0 - type: gt - operator: - type: and - query: - params: [] - reducer: - params: [] - type: last - type: query - # A row exists only for something currently wrong, so an empty - # result is the healthy state and has to resolve the alert. A - # failing query is not: that is the alert source being broken. - noDataState: OK - execErrState: Error - for: {{ $for }} - annotations: - summary: 'The journal on storage node {{ "{{" }} $labels.node {{ "}}" }} holds more records awaiting compression than the configured threshold.' - description: 'Compression is not keeping up with the journal, and journal replay on the next restart or failover grows with every record. The control plane latches this alert on the upward crossing and re-arms silently, emitting nothing when the backlog drains, so the alert clears an hour after the last crossing rather than on a recovery event.' - labels: - app: simplyblock - cluster: {{ $cid }} - severity: critical - isPaused: false - - uid: {{ printf "sbev-%s-%s" (sha256sum $cid | trunc 8) "jm-compression-error" }} - title: JM_compression_error_{{ $cid }} - condition: B - data: - - refId: A - relativeTimeRange: - from: 600 - to: 0 - datasourceUid: {{ $dsUid }} - model: - refId: A - type: json - source: url - format: table - # Only the backend parser runs server-side, which is what - # an alert rule needs. root_selector is tried as a gjson - # path first and falls back to JSONata; opening with a - # bracket guarantees the fallback. The v2 endpoint returns - # the records unwrapped, so $ is the record array itself. - parser: backend - url: /api/v2/clusters/{{ $cid }}/logs - url_options: - method: GET - params: - - key: limit - value: {{ $logLimit | quote }} - root_selector: |- - ( - $ev := [$[Event = 'jm_compression']]; - $latest := $count($ev) = 0 ? [] : [$each($ev{ (NodeId & '/' & VUID): [$] }, function($evs, $key) { - ( - $e := $evs[-1]; - { - 'node': $substringBefore($key, '/'), - 'jm_vuid': $substringAfter($key, '/'), - 'level': $e.Level, - 'detail': $e.Message - } - ) - })]; - [$latest[level = 'Error'].{ - 'node': node, - 'jm_vuid': jm_vuid, - 'detail': detail, - 'value': 1 - }] - ) - # Exactly one numeric field and the rest strings is what - # makes Grafana read the frame as a numeric table, turning - # every string field into a label of its own alert instance. - columns: - - selector: node - text: node - type: string - - selector: jm_vuid - text: jm_vuid - type: string - - selector: detail - text: detail - type: string - - selector: value - text: value - type: number - - refId: B - relativeTimeRange: - from: 600 - to: 0 - datasourceUid: __expr__ - model: - refId: B - type: threshold - datasource: - type: __expr__ - uid: __expr__ - expression: A - conditions: - - evaluator: - params: - - 0 - type: gt - operator: - type: and - query: - params: [] - reducer: - params: [] - type: last - type: query - # A row exists only for something currently wrong, so an empty - # result is the healthy state and has to resolve the alert. A - # failing query is not: that is the alert source being broken. - noDataState: OK - execErrState: Error - for: {{ $for }} - annotations: - summary: 'Journal compression failed on storage node {{ "{{" }} $labels.node {{ "}}" }} for jm_vuid {{ "{{" }} $labels.jm_vuid {{ "}}" }}: {{ "{{" }} $labels.detail {{ "}}" }}.' - description: 'The rule tracks the newest compression event per node and journal and fires while that event is an error, which is either a compression_failed status or a non-zero error code. The alert clears when the same journal reports a clean compression run.' - labels: - app: simplyblock - cluster: {{ $cid }} - severity: critical - isPaused: false -{{- end }} -{{- end }} {{- $grafana := .Values.controlplane.observability.grafana }} {{- $notifications := $grafana.notifications | default dict }} diff --git a/helm-charts/charts/simplyblock-operator/templates/controlplane_deploy.yaml b/helm-charts/charts/simplyblock-operator/templates/controlplane_deploy.yaml index 821e6eb54..c861c076a 100755 --- a/helm-charts/charts/simplyblock-operator/templates/controlplane_deploy.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/controlplane_deploy.yaml @@ -297,20 +297,6 @@ spec: key: MONITORING_SECRET - name: GF_ALERTING_ENABLED value: "true" -{{- $eventAlerts := .Values.controlplane.observability.grafana.eventAlerts | default dict }} -{{- $plugin := $eventAlerts.plugin | default dict }} -{{- $eventClusters := include "simplyblock.eventAlertClusters" . | fromJsonArray }} -{{- if and $eventAlerts.enabled $eventClusters (not $plugin.preinstalled) }} - # The event-log alert rules query the control plane's REST API, which - # needs a REST data source, and the Grafana image this chart deploys - # is a plain upstream mirror carrying no plugins. /var/lib/grafana is - # an emptyDir, so this reinstalls on every pod start; bake the plugin - # into the image and set plugin.preinstalled to stop paying for that. - # The entrypoint reads the value as a URL, a semicolon, and the - # folder to install into. - - name: GF_INSTALL_PLUGINS - value: "{{ $plugin.url }};yesoreyeram-infinity-datasource" -{{- end }} - name: GF_PATHS_PROVISIONING value: "/etc/grafana/provisioning" - name: GF_SERVER_ROOT_URL diff --git a/helm-charts/charts/simplyblock-operator/values.yaml b/helm-charts/charts/simplyblock-operator/values.yaml index c9b6a3c6f..cf4631dce 100644 --- a/helm-charts/charts/simplyblock-operator/values.yaml +++ b/helm-charts/charts/simplyblock-operator/values.yaml @@ -271,38 +271,6 @@ controlplane: repository: quay.io/simplyblock-io/grafana tag: 10.0.12 pullPolicy: IfNotPresent - # Alert rules derived from the control plane's cluster event log rather - # than from Thanos. The event log carries what the metrics cannot: the - # transition a node, device, or cluster made, and therefore whether an - # operator asked for it. Reading it needs a REST data source, which is - # what the Infinity plugin below provides. - eventAlerts: - enabled: false - # Every rule reads the newest `logLimit` records of the event log and - # folds them into the current state of each node, device, and cluster. - # A limit too low silently heals an alert whose opening transition has - # scrolled out of the window, so raise it on a cluster that produces - # events quickly; every rule evaluation pays for it in one request. - logLimit: 1000 - # How often the rules run, and how long a condition must hold before it - # notifies. Both are Grafana durations. - interval: 1m - for: 1m - plugin: - # The Grafana image this chart deploys is a plain upstream mirror with - # no plugins in it, and /var/lib/grafana is an emptyDir, so the plugin - # is fetched on every pod start. Grafana's entrypoint reads - # GF_INSTALL_PLUGINS as a URL, a semicolon, and an install folder. - # - # 2.12.2 is the last Infinity release whose grafanaDependency - # (>=9.5.15) admits the Grafana 10.0.12 pinned above; every release - # still in the plugin catalog requires >=10.4.8. Moving Grafana - # forward is what unpins this. - url: "https://grafana.com/api/plugins/yesoreyeram-infinity-datasource/versions/2.12.2/download" - # Set to true once the plugin is baked into the Grafana image, which - # is what an air-gapped or restart-sensitive install wants: the - # download is ~74 MB and repeats on every pod start. - preinstalled: false endpoint: "" # Receivers Grafana provisions into the "grafana-alerts" contact point. # Every enabled channel is notified for every simplyblock alert. With none diff --git a/operator/docs/designs/crd-redesign/design-controlplane.md b/operator/docs/designs/crd-redesign/design-controlplane.md index b30c890ba..75299c72a 100644 --- a/operator/docs/designs/crd-redesign/design-controlplane.md +++ b/operator/docs/designs/crd-redesign/design-controlplane.md @@ -973,6 +973,29 @@ are host paths keeps its data on one machine, so the install either refuses without a class, names the hostpath path as a development mode, or keeps applying it. +**Q8: Where the event-log alert rules belong for a standalone install.** The +chart provisioned seven Grafana rules that read the control plane's cluster event +log through an Infinity data source, covering the transitions Thanos cannot see: a +node leaving online, a device becoming unavailable or being removed, a cluster +degrading, suspending, or reaching capacity, and the two journal-manager faults. +They were removed from the chart because the chart cannot build them. A data +source carries one cluster's credential, so the rules iterate the clusters, and +the chart learns a cluster's UUID and secret only when somebody writes them back +into the values after `cluster create`, which is the loop +[`design-clusterdeploymentconfig.md`](design-clusterdeploymentconfig.md) +replaced. + +The operator holds both. The `StorageCluster` reconciler knows every cluster's +UUID and secret and already upserts them into `simplyblock-csi-secret-v2` +([`design-simplyblockdriver.md`](design-simplyblockdriver.md) §4.3), so the +provisioning belongs to whatever reconciles the observability stack rather than +to a values file. Q2 leaves Grafana and Thanos with the chart while the base +control plane moves to the operator, which is why this has no home yet. + +Until it has one, a standalone deployment alerts on what Thanos scrapes and on +nothing the event log carries. The Thanos-derived rules are unaffected: they need +no per-cluster credential, which is the property that separates the two halves. + --- ## Appendix A: `controlplane_types.go` From 01101762d5aceccf501297cdfa13df99e5014296 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Thu, 17 Sep 2026 18:59:07 +0100 Subject: [PATCH 058/206] =?UTF-8?q?Resolve=20drtest-=20clone=20visibility:?= =?UTF-8?q?=20internal=20volumes,=20same=20as=20landing=20volumes=20(desig?= =?UTF-8?q?n=20=C2=A714.5)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- .../designs/design-csi-addons-replication.md | 16 ++++++++-------- 1 file changed, 8 insertions(+), 8 deletions(-) diff --git a/operator/docs/designs/design-csi-addons-replication.md b/operator/docs/designs/design-csi-addons-replication.md index 28f248013..3ce42c41f 100644 --- a/operator/docs/designs/design-csi-addons-replication.md +++ b/operator/docs/designs/design-csi-addons-replication.md @@ -2,7 +2,7 @@ **Status:** Draft **Author:** Israel Geoffrey (geoffrey1330) -**Date:** 2026-09-16 +**Date:** 2026-09-16 (last updated 2026-09-17) **Test Plan:** [`tests/test-plan-csi-addons-replication.md`](../tests/test-plan-csi-addons-replication.md) --- @@ -271,6 +271,8 @@ spec: `replicationPolicy` names the backend `ReplicationPolicy` (resolved per cluster by name), which owns cadence, retention, mode, and the replication target. `schedulingInterval` restates the policy's interval for Ramen's `DRPolicy` matching, and the preflight (§7.2) checks the two agree. No secrets parameter is needed: the driver's credentials come from `secret.json`, as for every other RPC. The chart ships no default class. Classes are the user's to author, matching the `VolumeGroupSnapshotClass` decision. +One class per (policy, cadence) is the authoring model: a `VolumeReplicationClass` names exactly one policy, and a volume needing a different cadence follows a different policy under a different class. Per-volume interval overrides are not provided. + ### 7.2 peerClasses convention and preflight Two of the requirements below are Ramen's contract, and the rest is this design's convention; the split matters because only the contract can fail a DRPolicy. @@ -426,16 +428,14 @@ A drill that silently perturbed replication would be worse than no drill. Before ### 14.5 Costs and bounds - **A live clone pins its base snapshot.** Retention defers pruning a snapshot with a dependent clone, which is what keeps the drill safe, and also why a drill must carry a maximum lifetime: a long-lived bubble holds the secondary's replicated chain back. -- **Test-cluster mode consumes real resources:** cross-cluster bandwidth for the full copy, and capacity on both the secondary (the `drtest-` clones) and the test cluster. The `drtest-` clones on the secondary exist only as replication sources and are never served; whether they are created as internal volumes (invisible to normal listings, exempt from subsystem caps, like landing volumes) is Open Question 4. +- **Test-cluster mode consumes real resources:** cross-cluster bandwidth for the full copy, and capacity on both the secondary (the `drtest-` clones) and the test cluster. The `drtest-` clones on the secondary exist only as replication sources and are never served, so they are created as internal volumes, the same treatment the shipping engine's own landing volumes already get: invisible to normal listings and exempt from the per-node subsystem cap. Storage capacity is unaffected either way; a drill's cleanup and audit go through the `drtest-` test-id enumeration (§14.1), not the ordinary volume listing. - **The group rules apply unchanged:** a `drtest-` consistency group observes the member cap and the placement pin like any other, so a drill of a large group is a capacity event on the secondary. --- ## 15. Open Questions -| # | Question | Owner | -|-----|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------| -| 1 | **Demote semantics for the application.** The P0-3 demote fences the volume (ANA inaccessible) after the final flush, and with convergence folded into the verb it is now the only place a planned swap can stall. Ramen relocation unmounts the workload first, so the fence is ordinarily unopposed. Confirm the verb's behavior when writes are still in flight at quiesce (block versus fail), whether the converge phase has its own budget separate from the quiesced flush, and whether a timeout in either phase must abort back to serving primary. | Backend team | -| 2 | **Where the preflight lives.** §7.2 attaches peerClasses validation to the `ReplicationPair` reconciler. If the redesign retires the pair kind, the preflight needs a new home (the `SimplyblockDriver`, or a standalone check job). | Operator team | -| 3 | **Per-volume policy granularity.** A `VolumeReplicationClass` names one policy, and today one policy implies one target and cadence for all its volumes. Confirm one class per (policy, cadence) is an acceptable authoring model for Ramen's `replicationClassSelector`, or whether per-volume interval overrides are needed. | Operator / Backend team | -| 4 | **Visibility of `drtest-` clones.** The test-cluster mode's clones on the secondary are replication sources only, never served. Decide whether the backend creates them as internal volumes (hidden from listings, exempt from the per-node subsystem cap, like the shipping path's landing volumes) or as ordinary volumes under a naming convention. | Backend team | +| # | Question | Owner | +|-----|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|---------------| +| 1 | **Demote semantics for the application.** The P0-3 demote fences the volume (ANA inaccessible) after the final flush, and with convergence folded into the verb it is now the only place a planned swap can stall. Ramen relocation unmounts the workload first, so the fence is ordinarily unopposed. Confirm the verb's behavior when writes are still in flight at quiesce (block versus fail), whether the converge phase has its own budget separate from the quiesced flush, and whether a timeout in either phase must abort back to serving primary. | Backend team | +| 2 | **Where the preflight lives.** §7.2 attaches peerClasses validation to the `ReplicationPair` reconciler. If the redesign retires the pair kind, the preflight needs a new home (the `SimplyblockDriver`, or a standalone check job). | Operator team | From 1ba080725668c219ada8a4152b0d24bdc0d7b926 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Thu, 17 Sep 2026 19:12:07 +0100 Subject: [PATCH 059/206] =?UTF-8?q?Resolve=20drtest-=20clone=20visibility:?= =?UTF-8?q?=20internal=20volumes,=20same=20as=20landing=20volumes=20(desig?= =?UTF-8?q?n=20=C2=A714.5)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- operator/docs/designs/design-csi-addons-replication.md | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/operator/docs/designs/design-csi-addons-replication.md b/operator/docs/designs/design-csi-addons-replication.md index 3ce42c41f..af2dbf3f2 100644 --- a/operator/docs/designs/design-csi-addons-replication.md +++ b/operator/docs/designs/design-csi-addons-replication.md @@ -18,6 +18,8 @@ Phase 1 is independently useful: a `VolumeReplication` object per PVC whose status truthfully reports the relationship, which no surface provides today. Phase 2 makes the object drivable, which is what Ramen actually needs. Phase 3 makes the whole thing operable at fleet scale. Phase 4 turns the same primitives into a rehearsal: a failover that can be drilled, in a bubble or against a test cluster, without touching production replication. +The phase numbers above are this document's own, not the DR storage foundation gap analysis's (§1): its Phase 0 (shipping the csi-addons contract itself) is this design's Phase 1, and its Phase 1 (promote, demote, and resync end to end through Ramen) is this design's Phase 2. + --- ## Phase 0 — External Prerequisites @@ -438,4 +440,4 @@ A drill that silently perturbed replication would be worse than no drill. Before | # | Question | Owner | |-----|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|---------------| | 1 | **Demote semantics for the application.** The P0-3 demote fences the volume (ANA inaccessible) after the final flush, and with convergence folded into the verb it is now the only place a planned swap can stall. Ramen relocation unmounts the workload first, so the fence is ordinarily unopposed. Confirm the verb's behavior when writes are still in flight at quiesce (block versus fail), whether the converge phase has its own budget separate from the quiesced flush, and whether a timeout in either phase must abort back to serving primary. | Backend team | -| 2 | **Where the preflight lives.** §7.2 attaches peerClasses validation to the `ReplicationPair` reconciler. If the redesign retires the pair kind, the preflight needs a new home (the `SimplyblockDriver`, or a standalone check job). | Operator team | +| 2 | **Where the preflight lives.** §7.2 attaches peerClasses validation to the `ReplicationPair` reconciler. If the redesign retires the pair kind, the preflight needs a new home (the `SimplyblockDriver`, or a standalone check job). | Operator team | From 343f4f99cd59d0202a110b20873488512af8d5e4 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 20:32:40 +0200 Subject: [PATCH 060/206] docs(crd-redesign): the designs describe the system that shipped MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Nine of the twelve designs taught a model the implementation had already moved away from. Every one of those deviations was reasoned and recorded, in a commit message or a code header, and never in the document it deviated from, so a reader implementing from the corpus would have built the wrong thing in nine places. The largest is the control plane's source vocabulary. `managed` and `external` became `local` and `managed`, which inverts the meaning of one of the two words against the whole document: `local` is what the operator installs, and `managed` names the control plane administering this cluster's storage from elsewhere. The rename reaches §3.2, §3.4, §4.1's diagram, §4.3, §4.4, §5, §6, §9.1, §10, §11, §12, and both appendices, and it carries the rest of that kind's drift with it: an optional credentials reference for the in-cluster case, eight watched components with no document store among them, the object store the datastore step installs, FoundationDB's CRDs as a held prerequisite rather than a detection, and `deployment.profile` in place of a flag that no longer exists. The second is the storage version. There is one manifest set rather than a staged pair, and it declares `v1alpha2` as storage from the moment it is applied, so what separates the five things the staging existed to separate is ordering rather than a held flag: the conversion webhook is up and smoke-tested before the CRDs are. What the apply does not do is move the objects, so §24 stops being a flip and becomes the drain of a representation, and the conversion webhook stops being contained by the operator's own liveness and becomes a process that reads no converted kind. The rest are smaller and the same shape. The kubelet toggle is a removal rather than an inverting rename, because its successor is on another kind. A storage node reads its parent from its controller owner reference rather than through a bridge field. A pool's allowed nodes are Kubernetes Nodes, because the UID the host NQN derives from and the labels the allowance is written on both live there. The cluster gate exempts `Remove`, because a removal is how an unready cluster becomes ready. A cluster's phase gained `Provisioning` and `Activating`, because a cluster being built is not a cluster that is broken. A deployment config gained an `Activating` step, because a document that stops at "the objects exist" leaves a cluster serving nothing. Backups are mirrored from the control plane's stream rather than walked out of a bucket. The vCPU floor is 4 in all three documents that state it. Two open questions are settled and retired rather than reused, since both are cited from review history: the control plane's document store, by measurement, and the deployment config's environment mapping, by the table the expansion applies. Four of these are decisions rather than resyncs, and the audit says so for each, naming the edit to revert if the other branch was wanted. design-controlplane §8 said the version endpoint was a prerequisite to sequence behind, and the shipped Upgrade copes without it; §8 now records coping and states what it costs, because refusing every upgrade of a control plane that serves no version refuses every upgrade there is today. design-storagecluster §4.1 claimed a direct read the code does not make, and now names the optimistic-lock claim as what actually makes creation single-shot. design-storagenode §3.4 exempted an unset minHugePagesSize from the fleet-uniformity check, which its own paragraph had just argued was the dangerous case; there is no inheritance at render time, so an unset value is a different floor rather than deference to the cluster's. And §3.1 said node names carry a random identifier when they are derived from the cluster, the worker, and the slot — the name is stable across a migration because Kubernetes never renames an object, not because the worker is kept out of it, so a migrated node keeps a name describing where it was built. Co-Authored-By: Claude Opus 5 (1M context) --- .../crd-redesign/design-api-upgrade.md | 174 ++++--- .../design-clusterdeploymentconfig.md | 208 ++++++-- .../crd-redesign/design-controlplane.md | 467 +++++++++++------- .../designs/crd-redesign/design-crd-model.md | 59 ++- .../crd-redesign/design-property-renames.md | 305 +++++++----- .../crd-redesign/design-storagebackup.md | 117 ++++- .../crd-redesign/design-storagecluster.md | 151 ++++-- .../crd-redesign/design-storagenode.md | 152 ++++-- .../crd-redesign/design-storagepool.md | 103 +++- 9 files changed, 1159 insertions(+), 577 deletions(-) diff --git a/operator/docs/designs/crd-redesign/design-api-upgrade.md b/operator/docs/designs/crd-redesign/design-api-upgrade.md index e66c0eb53..509e849fd 100644 --- a/operator/docs/designs/crd-redesign/design-api-upgrade.md +++ b/operator/docs/designs/crd-redesign/design-api-upgrade.md @@ -1,8 +1,8 @@ # Design Document: API Upgrade and Resource-Model Migration -**Status:** Draft +**Status:** Partially Implemented **Author:** Christoph Engelbert (noctarius) -**Date:** 2026-09-10 +**Date:** 2026-09-10 (last updated 2026-09-17) **Related designs:** [`design-crd-model.md`](design-crd-model.md) §9 is the migration inventory this document delivers **Test Plan:** [`test-plan-api-upgrade.md`](../../tests/test-plan-api-upgrade.md), not yet written @@ -333,10 +333,23 @@ annotation keyed `storage.simplyblock.io/conversion-` on the way down and restores it on the way up, so a `v1alpha1` client that reads and writes an object back does not truncate it. -`skipKubeletConfiguration` is the one field whose conversion is not a copy. It -becomes `enableKubeletConfiguration`, which inverts the sense, so a mechanical -rename produces the wrong behavior and the conversion negates the value in both -directions (`design-crd-model.md` §9.6). +`migrationEnabled` is the one field whose conversion is not a copy. It becomes +`disableMigration`, which inverts the sense, so a mechanical rename produces the +wrong behavior and the conversion negates a stated value in both directions while +leaving an unstated one unstated, since both spellings mean the same thing when +absent (`design-property-renames.md` §3.4). + +`enableDataRealignment` is the one field whose *default* changes direction, from +on to off. An object that stated nothing is indistinguishable in the stored shape +from one that deliberately turned the feature off, so the conversion writes the +value into an annotation on every trip down and reads that annotation's absence on +the way up as the mark of a client that only ever spoke `v1alpha1`. + +`skipKubeletConfiguration` is not a conversion at all. The toggle left the node +kinds for `StorageCluster.spec.storageNodes.enableKubeletConfiguration`, which is +a different kind, so the registered field is a removal that stashes under +`storage.simplyblock.io/v1alpha1-spec.overrides.skipKubeletConfiguration` and +restores from it (`design-storagenode.md` §15.1). ### 6.3 What Conversion Can and Cannot Carry @@ -347,10 +360,10 @@ the upgrade needs two phases rather than a webhook. | Change | Carried by | Reason | |------------------------------------------------------------------------------------|--------------------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| Boolean toggle renames, eleven fields across five kinds | Conversion | Same kind, both spellings expressible | +| Boolean toggle renames, eleven fields across five kinds | Conversion, partly | Same kind, both spellings expressible, except the kubelet toggle, which moves to another kind and is stashed rather than converted | | Enum recasing, `StorageClusterOpsAction`, `StorageNodeOpsAction`, `MetricsBackend` | Conversion | Same kind, value maps one to one | | `status.subPhase` string becoming `status.step` object | Conversion | The old string reads into `step.state`, leaving `step.deadline` absent, which restores as a step with no deadline, so an operation in flight across the upgrade keeps running | -| `StorageNode.spec.storageNodeSetRef` becoming a cluster reference | Conversion, partly | The field converts, but the value it should hold is only known once §20 has reparented the node | +| `StorageNode.spec.storageNodeSetRef` becoming a cluster reference | Conversion, partly | The hub reads its parent off the controller owner reference, which §20 writes. A node converted before that carries an empty `spec.clusterRef` until the reparent has run | | `BackupPolicy` becoming `StorageBackupPolicy` | `migrate` | A different kind is a different CRD, and no conversion webhook is invoked across kinds | | `VolumeMigration` absorbed into `PersistentVolumeOps` | `migrate` | Different kind, and the target is cluster-scoped while the source is namespaced | | `BackupRestore` absorbed into `StorageBackupOps` | `migrate` | Different kind | @@ -401,7 +414,7 @@ The Delta column cites the design that owns the change. |---------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|------------------------| | `StorageCluster` | `maxHugePagesSize` → `minHugePagesSize`, `hashicorpVaultSettings` → `kms.vault`, six toggles renamed, `backup.localEndpoint` → `endpoint`, and five spec removals. `design-storagecluster.md` §12 | `v1alpha1`, `v1alpha2` | | `StorageClusterOps` | `nodeRollingRestart` → `rollingRestart`, six action values recased, `status.triggered` removed. `design-storagecluster.md` §12 | `v1alpha1`, `v1alpha2` | -| `StorageNode` | `storageNodeSetRef` → `clusterRef` and `nodeSet`, `overrides` → `config`, `socketIndex` → `slot`, `skipKubeletConfiguration` inverted, four dead fields removed. `design-storagenode.md` §15.1 | `v1alpha1`, `v1alpha2` | +| `StorageNode` | `storageNodeSetRef` → `clusterRef` and `nodeSet`, `overrides` → `config`, `socketIndex` → `slot`, five dead fields stashed and removed. `design-storagenode.md` §15.1 | `v1alpha1`, `v1alpha2` | | `StorageNodeOps` | `storageNodeRef` → `nodeRef`, `drain` → `remove`, six action values recased, `status.triggered` removed. `design-storagenode.md` §15.2 | `v1alpha1`, `v1alpha2` | | `StoragePool` | `clusterName` → `clusterRef`, `dhchap` → `volumeDefaults.enableDHCHAP`, `encryption` and `replicate` renamed, `spec.action` and `spec.status` removed. `design-storagepool.md` §11 | `v1alpha1`, `v1alpha2` | | `ControlPlane` | `spec.image` → `spec.source.managed.image`, and `status.phase` becomes a four-value typed phase. `design-controlplane.md` §11 | `v1alpha1`, `v1alpha2` | @@ -436,6 +449,14 @@ registers `v1alpha2` as its only version: the ten additions of `StoragePoolOps`, `PersistentVolumeOps`, `StorageBackupOps`, and `NFSExport`, plus `StorageBackupPolicy`, which is `BackupPolicy` under its new name. +Ten of the eleven are registered. `NFSExport` is specified on another branch +([`design-pnfs-rwx.md`](../design-pnfs-rwx.md) §7.1) and is born with the work +that owns it rather than with this migration, which changes nothing here: a kind +with no CRD is a kind the installer does not apply. `VolumeGroupSnapshotOps` +joined the group from [`design-consistency-groups.md`](../design-consistency-groups.md) +after this document was written and is on the same footing as the ten: born at +`v1alpha2`, with no conversion function and no place in the storage rewrite. + None of them needs a conversion function, and none appears in the storage rewrite, because nothing was ever persisted at an older version of them. @@ -445,10 +466,12 @@ installed before the controller that reconciles it exists, which is inert: a registered kind with no controller and no objects does nothing until the operator carrying its controller is running. -### 7.4 The Staging +### 7.4 What the Manifests Declare -Each of the seven converting CRDs is installed with both versions served and -`v1alpha1` retained as the storage version: +There is one set of CRD manifests rather than a staged pair. Each of the seven +converting CRDs serves both versions and stores `v1alpha2` from the moment it is +applied, because `+kubebuilder:storageversion` sits on the `v1alpha2` type and +controller-gen writes the flag from there: ```yaml # operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml @@ -461,58 +484,76 @@ spec: versions: - name: v1alpha1 served: true - storage: true + storage: false - name: v1alpha2 served: true - storage: false + storage: true conversion: strategy: Webhook webhook: conversionReviewVersions: ["v1"] clientConfig: service: - namespace: simplyblock + namespace: simplyblock-operator-system name: simplyblock-operator-conversion-webhook-service path: /convert port: 443 ``` -The `conversion` stanza is not written by hand. `operator/config/crd/kustomization.yaml` -carries the `+kubebuilder:scaffold:crdkustomizewebhookpatch` marker and a -commented `patches` block, and `operator/config/default/kustomization.yaml` -carries `+kubebuilder:scaffold:crdkustomizecainjectionns` and -`crdkustomizecainjectionname` with a comment stating that the markers exist so -`kubebuilder create webhook --conversion` can wire up a future conversion -webhook. That scaffold is the intended entry point, and the one deviation from -what it generates is the CA bundle: the scaffold injects it with cert-manager's -`cert-manager.io/inject-ca-from` annotation, and this repository provisions -webhook certificates at runtime instead (§8). - -**The patch is applied per CRD and not to the whole `crd/bases` directory.** Ten -of the seventeen keep `strategy: None`, and a `conversion` stanza pointing at a -webhook on a CRD with one version is a dependency on a Deployment that has no -reason to exist for that kind, which is what §27 eventually removes. - -The staging exists so that five things can fail separately: +**One set of manifests is what a fresh install needs.** A cluster installed today +has no `v1alpha1` object to convert, so it writes `v1alpha2` from the first write, +converts nothing, and deploys no webhook. Holding storage at `v1alpha1` in the +shipped manifests would make the ordinary install the exceptional case: every +object would be stored in a version nothing reads, behind a webhook that exists +only for upgrades. + +**The `conversion` stanza is written into the base rather than patched over it.** +`operator/hack/apply-conversion-webhook.sh` runs after controller-gen, which +regenerates each base from the Go types and has no marker for `spec.conversion`, +and writes the stanza into every CRD named by +`operator/config/crd/converted-kinds.txt`. That file is the single list three +consumers have to agree on (the script, the webhook registration in +`internal/webhook/conversion.go`, and the CA injection in +`internal/webhook/cert.go`), and `TestConvertedKindsMatchTheManifestList` asserts +two of them against each other. A Kustomize patch would reach `make install` and +the installer and miss the chart, which copies `config/crd/bases` verbatim into +its own `crds/` directory and does not template it, and the chart is what +installs these CRDs on a cluster. + +**The ten single-version CRDs keep `strategy: None`.** A `conversion` stanza on a +kind with one version is a dependency on a Deployment that has no reason to exist +for it, which is what §28 eventually removes for the seven. + +**The stanza ships pointing at a Service the chart does not deploy, and the +namespace in it is a default rather than a fact.** A CRD is cluster-scoped and +lands in `crds/`, so the namespace cannot be templated, and the operator patches +the service reference and the CA bundle together at runtime because it is the only +party that knows which namespace it is running in. Leaving the strategy at `None` +until the operator raises it is the worse trade: under `None` the API server +answers a `v1alpha2` read of a stored `v1alpha1` object by relabeling the +apiVersion and pruning every field the new schema does not know, which is silently +wrong data rather than a failed read. + +Five things still fail separately, and what separates them is ordering rather than +a held flag: 1. Introducing the new API. 2. Proving conversion works. 3. Upgrading the operator. 4. Migrating the application resource model. -5. Changing the persisted storage representation. +5. Changing the persisted representation of each object. -Once the new operator is verified and the application-level migration has run, -storage switches on those same seven: +§9.1 is that order. The conversion webhook is deployed, awaited, and smoke-tested +against a real object before `apply-crds` runs, so there is something to convert +with at the moment storage moves. The operator upgrade follows the CRDs, and the +resource-model migration follows the operator. -```yaml -versions: - - name: v1alpha1 - served: true - storage: false - - name: v1alpha2 - served: true - storage: true -``` +**The apply moves storage, and it does not move the objects.** The flag decides +what a write encodes, so an object untouched since the apply stays in the +`v1alpha1` representation, is converted up on every read, and keeps +`.status.storedVersions` listing `v1alpha1`. §24 is the rewrite that drains it, +and until that has run an upgraded cluster is one where the storage version and +the stored representations disagree. `v1alpha1` becomes `served: false` on the seven under §28's conditions. Its readers are the operator's own reconcilers and webhooks under @@ -678,9 +719,9 @@ The smoke test verifies that: - The conversion webhook is reachable. - Conversion succeeds for each of the seven converting kinds that has at least one object (§7.2). -- The fields whose conversion is not a copy are correct, which means at minimum - a recased action enum, a renamed boolean toggle, and the. - `skipKubeletConfiguration` inversion (§6.2). +- The fields whose conversion is not a copy are correct, which means at minimum a + recased action enum, a renamed boolean toggle, the `migrationEnabled` inversion, + and a field the hub removed reading back from its stash (§6.2). - A `v1alpha2` read followed by a `v1alpha1` read returns the original representation. @@ -913,10 +954,9 @@ has meant until now, so `upgrade` says so in its closing report. helm upgrade helm-charts/charts/simplyblock-operator ``` -The chart carries the same field names the API does, so it moves with the API. -`values.yaml` holds `skipKubeletConfiguration` under the storage-node settings, -and `multiCluster.enable` is a spelling that exists only there. Two consequences -follow. +The chart carries the same field names the API does, so it moves with the API, +and it carries spellings of its own that name no API field at all, of which +`multiCluster.enable` is one. Two consequences follow. **A user's existing values file may not validate against the new chart.** `helm-charts/charts/simplyblock-operator/values.schema.json` sets @@ -1836,19 +1876,24 @@ process has to stay alive for the migration to be recoverable. ## 24. Storage-Version Migration -Changing the storage version is separate from application-level migration. For -each CRD: +Draining the old storage representation is separate from application-level +migration, and by the time this stage runs the storage version has already moved: +`apply-crds` installed manifests that declare `v1alpha2` as storage (§7.4). + +What the flag did not do is touch what is already in etcd. Objects that have not +been written since the apply still exist in the old representation, and +`.status.storedVersions` still lists `v1alpha1`: ```text -v1alpha1 storage=true v1alpha1 storage=false -v1alpha2 storage=false → v1alpha2 storage=true +CRD: v1alpha1 storage=false, v1alpha2 storage=true +etcd: object A encoded v1alpha1 ← until something writes it + object B encoded v1alpha2 ← written since the apply +storedVersions: ["v1alpha1", "v1alpha2"] ``` -After the switch, objects that have not been written since still exist in etcd -in the old representation, and `.status.storedVersions` still lists -`v1alpha1`. Until that list holds `v1alpha2` alone, `v1alpha1` cannot be -removed from the CRD, because the API server refuses to drop a version it still -has stored objects in. +Until that list holds `v1alpha2` alone, `v1alpha1` cannot be removed from the +CRD, because the API server refuses to drop a version it still has stored +objects in. **The migration rewrites the objects itself.** The Kubernetes `StorageVersionMigration` API is not used: it is served at @@ -1860,7 +1905,8 @@ alpha feature gate is not a migration path. The rewrite is: -1. Switch the CRD to the new storage version. +1. Confirm the CRD stores `v1alpha2`, which `verify-crd-versions` already + established and this stage re-reads rather than assumes. 2. List every object of that kind, in every namespace. 3. Write each object back unchanged. 4. Verify the write. @@ -2394,8 +2440,10 @@ is proven red before the fix. Every field mapping, renamed field, moved field, default value, removed field, enum change, type change, nested object, list, map, and nil or empty value. -`skipKubeletConfiguration` gets its own test for the inversion, and each of the -three recased action enums gets a test per value. +`migrationEnabled` gets its own test for the inversion, `enableDataRealignment` +one for the default that changes direction, each field the hub removed one for the +stash it round-trips through, and each of the three recased action enums a test +per value. ### 30.2 Round-Trip Tests diff --git a/operator/docs/designs/crd-redesign/design-clusterdeploymentconfig.md b/operator/docs/designs/crd-redesign/design-clusterdeploymentconfig.md index da0523140..fb2a4bb59 100644 --- a/operator/docs/designs/crd-redesign/design-clusterdeploymentconfig.md +++ b/operator/docs/designs/crd-redesign/design-clusterdeploymentconfig.md @@ -1,8 +1,8 @@ # Design Document: The Deployment Config and the Operator's Own Operations -**Status:** Draft +**Status:** Implemented, with the exceptions §12 records **Author:** Christoph Engelbert (noctarius) -**Date:** 2026-08-29 (last updated 2026-09-08) +**Date:** 2026-08-29 (last updated 2026-09-17) **Target Release:** simplyblock 26.4 **Test Plan:** [`tests/test-plan-clusterdeploymentconfig.md`](../../tests/test-plan-clusterdeploymentconfig.md) **Example:** [`assets/example-cluster-config.yaml`](assets/example-cluster-config.yaml) @@ -125,7 +125,7 @@ corrects it, approves it, and the operator expands it. ## 3. ClusterDeploymentConfig: API -Declared in `operator/api/v1alpha1/clusterdeploymentconfig_types.go`, short name +Declared in `operator/api/v1alpha2/clusterdeploymentconfig_types.go`, short name `cdc`. The type is Appendix A and a filled-in document is [`assets/example-cluster-config.yaml`](assets/example-cluster-config.yaml). What follows quotes the field an argument turns on and no more. @@ -136,7 +136,7 @@ The document has three parts: what the deployment is, what cluster to make, and which nodes to make it out of. ```yaml -apiVersion: storage.simplyblock.io/v1alpha1 +apiVersion: storage.simplyblock.io/v1alpha2 kind: ClusterDeploymentConfig metadata: name: production @@ -275,16 +275,29 @@ document is a draft and a reviewer edits it freely. After approval it is the record of what was deployed, and editing it would describe a deployment that never happened. -That is one CEL rule on the spec rather than a marker per field: +That is one CEL rule on the spec rather than a marker per field, and a second +rule stops the gate itself from being closed again: ```go -// +kubebuilder:validation:XValidation:rule="!oldSelf.approved || self == oldSelf",message="an approved deployment config is immutable" +// +kubebuilder:validation:XValidation:rule="!has(oldSelf.approved) || !oldSelf.approved || self == oldSelf",message="an approved deployment config is immutable" +// +kubebuilder:validation:XValidation:rule="!has(oldSelf.approved) || !oldSelf.approved || self.approved",message="approval cannot be withdrawn" ``` **`spec.approved` itself is immutable once true.** Un-approving something already expanded does not un-expand it, and a field that can be toggled back invites the belief that it does. +**The field is defaulted and serialized rather than omitted when false, and the +`has()` guard is why both rules work.** A bool omitted when false is a key the API +server never stores, so a rule reading `oldSelf.approved` on a document that has +never been approved fails to evaluate rather than reading `false`, and a rule that +fails to evaluate denies the request. Spelled without the guard, the immutability +rule refuses every approval there could ever be. `+kubebuilder:default=false` and +the absent `omitempty` put the key in the stored object, and the guard carries the +documents written before the default existed. It is also what a reviewer needs: +the gate they are being asked to open has to be visible in the document rather +than implied by its absence. + Both rules are restated by the validating webhook of §5, which is what turns a rejection naming a CEL rule into one naming the field that was edited. @@ -362,6 +375,10 @@ an immutable document, and no mechanism below the reviewer prevents that. CreatingNodes ← one StorageNode per worker per slot │ ▼ + Activating ← wait for this document's own nodes, then ask for the + cluster to be activated + │ + ▼ phase: Expanded ``` @@ -373,9 +390,16 @@ existing cluster writes nothing: the cluster's class already holds, and a docume whose groups disagree with it was rejected at approval (§5.1). **`CreatingNodes` resolves the document's shorthands as it writes.** Each node -gets the cluster's sizing, its group's devices as one `config.deviceNames` list -(§3.1), and the four distribution flags `spec.environment` stands for. Nothing on -the node refers back to the config, which is what §4.3 means by owning nothing. +gets the cluster's sizing and its group's devices as one `config.deviceNames` +list (§3.1). Nothing on the node refers back to the config, which is what §4.3 +means by the document owning no part of the deployment. + +**The distribution flags `spec.environment` stands for land on the cluster, not on +each node.** They configure the storage-node workload, which is one DaemonSet for +every node it schedules, so they are resolved once onto +`StorageCluster.spec.storageNodes` +([`design-storagenode.md`](design-storagenode.md) §5.1) rather than stamped node +by node. §12 Q1 is the mapping. **`CreatingNodes` is the step that must be idempotent, and it is by construction.** A `StorageNode` is identified by `(clusterRef, workerNode, slot)` @@ -383,22 +407,40 @@ A `StorageNode` is identified by `(clusterRef, workerNode, slot)` exists for the cluster and creates only the slots that do not. A crash part-way through creates the rest on the next pass and duplicates nothing. -**The expansion does not wait for the nodes to come up.** It creates the objects -and finishes. Provisioning them is the node controller's, it is bounded by -`maxParallelNodeAdds`, and a document that stayed `Expanding` until a -twenty-node fleet was online would be reporting the fleet's progress rather than -its own. +**`Activating` waits for this document's own nodes and then asks for the +cluster.** A document knows how many nodes it made, so it knows when the +deployment it describes is whole, and stopping at "the objects exist" would leave +a cluster serving nothing behind a document reporting `Expanded`, with nothing +saying that one more thing was required of anybody. The step holds while any node +it created is still coming up, raises one `StorageClusterOps` activation when they +are all Online, and completes when that operation does. + +**It waits for its own nodes and not for the fleet.** The set it watches is the +one `CreatingNodes` wrote, so a document adding four nodes to a twenty-node +cluster waits for four. Provisioning is still the node controller's and still +bounded by `maxParallelNodeAdds`. What the document adds is the knowledge of which +nodes are its own. + +**The activation is asked for once.** `Activating` is re-entered on every +reconcile until the operation finishes, and the operation is named after the +document rather than generated, so a second pass finds the one that exists rather +than raising another. ### 4.3 Deletion **There is no finalizer, and that is deliberate.** Deleting a -`ClusterDeploymentConfig` deletes a document. It owns nothing, nothing references -it, and nothing reads it after expansion, so there is nothing to clean up and -nothing to protect. +`ClusterDeploymentConfig` deletes a document. Nothing references it and nothing +reads it after expansion, so there is nothing to clean up. -Specifically, it does **not** own the `StorageCluster` it created. An owner -reference would make deleting the document delete the cluster and every volume in -it, which is the opposite of ephemeral. +Specifically, it does **not** own the `StorageCluster` it created, nor any +`StorageNode`. An owner reference would make deleting the document delete the +cluster and every volume in it, which is the opposite of ephemeral. + +**The one object it does own is the activation it raised.** The +`StorageClusterOps` of §4.2 carries a controller reference to the document, +because it is the document's own act rather than a lasting part of the deployment: +an operation that has finished is a record of a request, and deleting the request +with the document that made it leaves the cluster it activated untouched. --- @@ -550,7 +592,7 @@ has data on those devices. ## 7. OperatorOps -Declared in `operator/api/v1alpha1/operatorops_types.go`, short name `oops`, and +Declared in `operator/api/v1alpha2/operatorops_types.go`, short name `oops`, and reconciled by `OperatorOpsReconciler` in `operator/internal/controllers/deployment/operatorops_controller.go`, beside the config its discovery action writes. The type is Appendix B. @@ -890,14 +932,25 @@ piece left. ## 12. Open Questions -**Q1: What each `KubernetesEnvironment` value resolves to.** §3.1 has -`spec.environment` decide `enableKubeletConfiguration`, `enableCpuTopology`, -`ubuntuHost`, and `openShiftCluster` on every node the expansion creates, and -names `OpenShift` as the value a reader can already infer. `Vanilla`, `Rancher`, -`K3s`, and `Talos` have no stated mapping. The table belongs in this document, -because the expansion is what applies it, and writing it needs one answer per -distribution about whether the kubelet is reconfigured and whether CPU topology -is readable, which is a question for whoever has run simplyblock on each of them. +**Q1 is settled and its number is retired rather than reused**, since it is cited +from review history. §4.2 is where the answer went: the table below is what +`spec.environment` resolves to, and the expansion writes it onto +`StorageCluster.spec.storageNodes` rather than onto each node. + +| Environment | Kubelet configuration | CPU topology | OpenShift | +|-----------------------------|-----------------------|--------------|-----------| +| `OpenShift` | Applied | Read | Yes | +| `Vanilla`, `Rancher`, `K3s` | Applied | Unset | No | +| `Talos` | Not applied | Unset | No | + +`Talos` is the row the shorthand exists for: it has no writable kubelet +configuration and no package manager, so a node applies neither, and a deployment +onto it that had to discover that field by field would discover it by failing. +`OpenShift` states the kubelet flag rather than leaving it unset, because the +renderer reads an unset flag as skipping the configuration and every OpenShift +deployment this product has shipped configures it. `ubuntuHost` is not in the +table: it describes the worker's host OS rather than the distribution running on +it, so it stays a value a deployment states. --- @@ -922,7 +975,7 @@ const ( ) // ClusterDeploymentConfigStep is one step of the expansion path. -// +kubebuilder:validation:Enum=Validating;CreatingCluster;AwaitingCluster;CreatingNodes +// +kubebuilder:validation:Enum=Validating;CreatingCluster;AwaitingCluster;CreatingNodes;Activating type ClusterDeploymentConfigStep string const ( @@ -930,6 +983,15 @@ const ( ClusterDeploymentConfigStepCreatingCluster ClusterDeploymentConfigStep = "CreatingCluster" ClusterDeploymentConfigStepAwaitingCluster ClusterDeploymentConfigStep = "AwaitingCluster" ClusterDeploymentConfigStepCreatingNodes ClusterDeploymentConfigStep = "CreatingNodes" + + // ClusterDeploymentConfigStepActivating waits for the nodes this document + // created and then asks for the cluster to be activated. + // + // The document knows how many nodes it made, so it knows when the deployment + // it describes is whole. Stopping at "the objects exist" would leave a + // cluster that serves nothing behind a document reporting Expanded, with + // nothing saying that one more thing is required of anybody. + ClusterDeploymentConfigStepActivating ClusterDeploymentConfigStep = "Activating" ) // KubernetesEnvironment is the distribution a deployment targets. The values are @@ -1053,9 +1115,9 @@ type ClusterTemplate struct { // this cluster. It is stated here and nowhere below, because the control // plane assumes it uniform across a cluster's nodes; CreatingNodes copies it // into every StorageNode.spec.config.sizing it writes. Required, because the - // StorageCluster's own field is. + // StorageCluster's own field is, and with the same floor. // +kubebuilder:validation:Required - // +kubebuilder:validation:Minimum=6 + // +kubebuilder:validation:Minimum=4 VCPUCount *int32 `json:"vcpuCount"` // MinHugePagesSize is the smallest huge-page allocation each storage node of @@ -1065,6 +1127,44 @@ type ClusterTemplate struct { // +optional MinHugePagesSize string `json:"minHugePagesSize,omitempty"` + // EnableDriveFormat formats every device the document names before a storage + // node takes it, which is how a drive carrying anything already is made + // usable. + // + // It says what is wanted rather than how, because the how differs by device + // class: an NVMe device is formatted to a 4K block size, and a logical block + // device has its signatures wiped. One field covers both, so a document does + // not have to know which class the expansion will resolve it to. + // + // It is on the document rather than defaulted further down because it is + // destructive and the document is what somebody approves. A reviewer reading + // a draft has to see that the drives it lists will be formatted, and be able + // to strike it before approving; the cluster's own field is immutable once + // the cluster exists, so a default nobody saw could not be undone either. + // +optional + EnableDriveFormat *bool `json:"enableDriveFormat,omitempty"` + + // SocketsToUse restricts the deployment to selected NUMA sockets, and empty + // means socket 0 alone. With NodesPerSocket it decides how many storage nodes + // each worker runs, so a group of two workers on a two-socket layout expands + // to four nodes. + // + // It is here rather than on a node set because it is immutable on the cluster + // it lands on: the layout a fleet was built with is not one a later document + // can vary, and a reviewer should see it before the cluster exists. + // +kubebuilder:validation:items:MaxLength=16 + // +kubebuilder:validation:MaxItems=16 + // +listType=set + // +optional + SocketsToUse []string `json:"socketsToUse,omitempty"` + + // NodesPerSocket is how many storage nodes run per NUMA socket. See + // SocketsToUse, which it multiplies. + // +kubebuilder:validation:Minimum=1 + // +kubebuilder:validation:Maximum=8 + // +optional + NodesPerSocket *int32 `json:"nodesPerSocket,omitempty"` + // Stripe is the erasure-coding layout. // +optional Stripe *StripeSpec `json:"stripe,omitempty"` @@ -1134,7 +1234,7 @@ type ClusterDeploymentConfigStatus struct { Phase ClusterDeploymentConfigPhase `json:"phase,omitempty"` // Step is the position of the expansion machine within Expanding. - // +kubebuilder:validation:XValidation:rule="!has(self.state) || self.state in ['Validating','CreatingCluster','AwaitingCluster','CreatingNodes']",message="unknown step" + // +kubebuilder:validation:XValidation:rule="!has(self.state) || self.state in ['Validating','CreatingCluster','AwaitingCluster','CreatingNodes','Activating']",message="unknown step" // +optional Step statemachine.KubeSnapshot `json:"step,omitempty"` @@ -1333,11 +1433,40 @@ type DiscoverSpec struct { // +optional NodeSelector map[string]string `json:"nodeSelector,omitempty"` + // EnableControlPlaneNodes lets the run consider machines that run the API + // server and etcd. + // + // It is off by default because a storage node is a data path, and putting one + // on an etcd host is a placement almost nobody intends. The approval gate is a + // poor place to catch it: a fifty-worker draft is not a document anybody reads + // closely enough to spot three control-plane nodes in it. A combined three-node + // or single-node deployment is the case that wants it, and those are set up + // deliberately. + // + // There is no field beside it for infrastructure nodes, because those are used + // without asking: an OpenShift infra node is the tier a cluster's own + // infrastructure runs on, and simplyblock storage is infrastructure. A fleet + // with disks in its infra nodes meant those disks to be the storage, so a draft + // proposes them ahead of the workers rather than leaving them out. + // +optional + EnableControlPlaneNodes *bool `json:"enableControlPlaneNodes,omitempty"` + // DeviceFilter narrows which of an inspected worker's devices reach the // draft. Empty reports every device the worker advertises, including the one // it boots from, which is what the approval gate then has to catch. // +optional DeviceFilter *DeviceFilter `json:"deviceFilter,omitempty"` + + // ClusterRef names an existing StorageCluster the draft grows rather than + // creates. It is copied to the draft's own clusterRef, so that re-running + // discovery after an expansion produces a growth document naming the same + // cluster. + // + // Bounded at what a StorageCluster name may be, since a longer value names + // nothing that can exist (design-api-upgrade.md §19.4). + // +kubebuilder:validation:MaxLength=63 + // +optional + ClusterRef string `json:"clusterRef,omitempty"` } // OperatorOpsSpec is one operation to perform against the operator itself. @@ -1378,6 +1507,19 @@ type OperatorOpsStatus struct { // +optional ConfigRef string `json:"configRef,omitempty"` + // Workers are the workers this run is inspecting, decided once in + // Inspecting so that a node joining the cluster mid-run does not change + // what the run is about. + // +optional + // +listType=set + Workers []string `json:"workers,omitempty"` + + // Environment is the Kubernetes distribution Inspecting concluded, which + // Writing copies into the draft. It is recorded here as well so that a run + // that failed later still says what it found. + // +optional + Environment KubernetesEnvironment `json:"environment,omitempty"` + // Message is the reason the phase is what it is: one sentence, replaced as // the operation moves, and never a log. // +optional diff --git a/operator/docs/designs/crd-redesign/design-controlplane.md b/operator/docs/designs/crd-redesign/design-controlplane.md index 75299c72a..b20b5fed9 100644 --- a/operator/docs/designs/crd-redesign/design-controlplane.md +++ b/operator/docs/designs/crd-redesign/design-controlplane.md @@ -1,14 +1,15 @@ # Design Document: The ControlPlane and Its Operations -**Status:** Draft +**Status:** Implemented, with the exceptions §12 records **Author:** Christoph Engelbert (noctarius) -**Date:** 2026-08-29 +**Date:** 2026-08-29 (last updated 2026-09-17) **Test Plan:** [`tests/test-plan-controlplane.md`](../../tests/test-plan-controlplane.md) -This document specifies the target model. `ControlPlane` is registered and in a -shape that predates the conventions of -[`design-crd-model.md`](design-crd-model.md), `ControlPlaneOps` does not exist, -and §11 is the single record of what the rework changes against what ships. +This document specifies the model both kinds now carry. `ControlPlane` moved to +`storage.simplyblock.io/v1alpha2` with a `v1alpha1` spoke and a conversion between +them, `ControlPlaneOps` is born there, and both are reconciled in +`operator/internal/controllers/controlplane/`. §11 remains the record of what +changed against the registered API. --- @@ -93,7 +94,7 @@ value needed a home. ### Goals - Specify a `ControlPlane` that says which control plane it means, so that - reusing an external one is a field rather than an environment variable (§3, §5). + reusing one elsewhere is a field rather than an environment variable (§3, §5). - Specify the two modes as siblings under one block, so that choosing between them is expressible and setting both is rejected (§5). - Specify what the operator installs when it installs, and what it does not touch @@ -167,21 +168,38 @@ management and for the same reason: two modes as siblings make the choice expressible, where two top-level fields make it a convention. ```go -// Source selects where the control plane comes from. Exactly one member is set. -// +kubebuilder:validation:XValidation:rule="(has(self.managed) ? 1 : 0) + (has(self.external) ? 1 : 0) == 1",message="set exactly one of managed or external" -// +kubebuilder:validation:Required -// +k8s:immutable -Source ControlPlaneSource `json:"source"` +// ControlPlaneSource selects whether this cluster hosts its control plane or is +// managed by one elsewhere. Exactly one member is set, and which one it is +// cannot change afterward. +// +kubebuilder:validation:XValidation:rule="(has(self.local) ? 1 : 0) + (has(self.managed) ? 1 : 0) == 1",message="set exactly one of local or managed" +// +kubebuilder:validation:XValidation:rule="has(self.local) == has(oldSelf.local) && has(self.managed) == has(oldSelf.managed)",message="spec.source is immutable" +type ControlPlaneSource struct { + Local *LocalControlPlane `json:"local,omitempty"` + Managed *ManagedControlPlane `json:"managed,omitempty"` +} ``` -**The block is immutable, not its members.** Switching a live deployment from a -control plane the operator installed to one it did not is not a reconfiguration, -it is a different deployment: the clusters, their UUIDs, and their volumes live in -the FoundationDB behind the old one. Making the block immutable says that in the -schema rather than in a runbook. - -`spec.managed` is what the operator installs (§5.1). `spec.external` is where an -existing control plane already is (§5.2). +`spec.source.local` is a control plane this cluster hosts, installed and owned by +the operator (§5.1), and `spec.source.managed` is one somewhere else, which this +cluster's storage is managed by rather than hosting (§5.2). The word "managed" is +about what administers the storage clusters rather than about who runs the +operator: a deployment in that mode registers its clusters with a control plane it +does not host, so the fleet is administered from there. + +**What is frozen is the choice between the two members, not the block.** +Switching a live deployment from a control plane the operator installed to one it +did not is not a reconfiguration, it is a different deployment: the clusters, +their UUIDs, and their volumes live in the FoundationDB behind the old one. The +members themselves stay editable, because editing `spec.source.local.image` is an +ordinary change and it is what an `Upgrade` performs (§6). Freezing the block +instead would emit `self == oldSelf` over the whole struct, which freezes the +image with it and makes that operation impossible to complete. + +**Both rules sit on the type rather than on the field that carries it.** +controller-gen emits a single field's marker-derived rules and its injected +immutability rule into one list in an order that varies between runs, so a field +carrying both produces a CRD that differs from itself and a drift check that fails +at random. Two rules of the same kind on a type are emitted in source order. ### 3.3 Status @@ -201,11 +219,12 @@ and a change to it reaches every reader without a Deployment rollout. `status.components` is the per-component readiness §4.3 derives the phase from: each component's desired count, its ready count, and whether it is essential. It is what makes the phase explainable, and it is the only place an administrator -learns which of eleven workloads is the one restarting. +learns which of the eight workloads is the one restarting. `status.version` is the management API's reported version, which is what a `ControlPlaneOps` upgrade moves and what [`design-simplyblockdriver.md`](design-simplyblockdriver.md) §5 -compares a driver against. +compares a driver against. It is unwritten wherever the endpoint that serves it +does not exist, which is every control plane today (§8). `status.lastChecked` is when the readiness probe last ran, and `status.activeOpsRef` is the operation lock ([`design-crd-model.md`](design-crd-model.md) §3.2). @@ -219,14 +238,14 @@ paraphrase of it. A control plane the operator installs, which is what a fresh deployment gets: ```yaml -apiVersion: storage.simplyblock.io/v1alpha1 +apiVersion: storage.simplyblock.io/v1alpha2 kind: ControlPlane metadata: name: simplyblock namespace: simplyblock spec: source: - managed: + local: image: quay.io/simplyblock-io/simplyblock:26.2.8 foundationDB: replicas: 3 @@ -243,7 +262,7 @@ A control plane that already exists, running outside the Kubernetes cluster: ```yaml spec: source: - external: + managed: endpoint: https://sb-control.example.com:5000 credentialsSecretRef: name: simplyblock-control-plane @@ -254,7 +273,7 @@ status: observedGeneration: 1 ``` -`status.endpoint` repeats `spec.source.external.endpoint` in the second case and +`status.endpoint` repeats `spec.source.managed.endpoint` in the second case and is derived in the first. That is the point: a reader and a controller ask `status.endpoint` either way, and neither has to know which mode the deployment is in. @@ -286,8 +305,8 @@ status: │ ┌──────────────────────────────────────────────────────┐ │ │ │ ControlPlaneReconciler │ │ │ │ 1. Name is "simplyblock"? Otherwise ignore │ │ -│ │ 2. spec.source.external → resolve and probe (§5.2) │ │ -│ │ 3. spec.source.managed → the install machine (§5.1)│ │ +│ │ 2. spec.source.managed → resolve and probe (§5.2) │ │ +│ │ 3. spec.source.local → the install machine (§5.1) │ │ │ │ 4. Available: probe, publish endpoint and version │ │ │ └──────────────────────────────────────────────────────┘ │ │ ControlPlane CR spec.source status.phase status.step │ @@ -363,15 +382,22 @@ the two that carry their own operator it is that resource's own report, because `FoundationDBCluster` at two of three coordinators is serving and a replica count cannot say so. -| Component | Readiness is | Essential | -|--------------------------------------------------------------------------------------------------------------------------|--------------------------------------|-----------| -| `simplyblock-webappapi` | Ready against desired replicas | Yes | -| `simplyblock-fdb-cluster` | The `FoundationDBCluster` health | Yes | -| `simplyblock-mongo` | The `MongoDBCommunity` member report | Yes | -| `simplyblock-fdb-controller-manager` | Ready against desired replicas | No | -| `simplyblock-tasks` | Ready against desired replicas | No | -| `simplyblock-minio`, `simplyblock-admin-control` | Ready against desired replicas | §12 Q5 | -| `simplyblock-monitoring`, `simplyblock-graylog`, `simplyblock-thanos`, `simplyblock-grafana`, `simplyblock-fdb-exporter` | Ready against desired replicas | No | +| Component | Readiness is | Essential | +|--------------------------------------|----------------------------------|-----------| +| `simplyblock-webappapi` | Ready against desired replicas | Yes | +| `simplyblock-fdb-cluster` | The `FoundationDBCluster` health | Yes | +| `simplyblock-fdb-controller-manager` | Ready against desired replicas | No | +| `simplyblock-tasks` | Ready against desired replicas | No | +| `simplyblock-monitoring` | Ready against desired replicas | No | +| `simplyblock-admin-control` | Ready against desired replicas | §12 Q5 | +| `simplyblock-minio` | Ready against desired replicas | §12 Q5 | +| `simplyblock-fdb-exporter` | Ready against desired replicas | No | + +**The table is exactly the set the install applies**, which is what makes it +watchable: a component the operator does not own is one it cannot report a count +for. The observability workloads the chart still renders are therefore absent from +it rather than reported as missing (§12 Q2), and there is no document store in it +because a base deployment runs none (§5.1). **Essential means the work stops, not that it slows down.** A component is essential when its absence loses work or stops the control plane answering. @@ -440,11 +466,11 @@ that is serving and should not be idea across the group is worth more than a word chosen to fit each kind slightly better. -**An external control plane never reaches `Degraded`.** The operator installs +**A managed control plane never reaches `Degraded`.** The operator installs nothing and owns no components there (§5.2), so `status.components` is empty, the -probe is the only signal, and its phase is `Available` or `Unavailable`. That is a real difference in -what the two sources can report rather than a gap to be filled, since the pods -behind somebody else's endpoint are not the operator's to watch. +probe is the only signal, and its phase is `Available` or `Unavailable`. That is a +real difference in what the two sources can report rather than a gap to be filled, +since the pods behind somebody else's endpoint are not the operator's to watch. **Events are emitted on transition, not on every probe.** A thirty-second probe that emitted on every failure would produce two thousand events a day from one @@ -455,12 +481,12 @@ keeps. The finalizer is `storage.simplyblock.io/controlplane-finalizer`. -**Deleting a `ControlPlane` with `spec.source.managed` deletes a database.** The +**Deleting a `ControlPlane` with `spec.source.local` deletes a database.** The finalizer therefore refuses while any `StorageCluster` in the namespace still exists, emits `ClustersStillPresent`, and requeues. That is a hold rather than a failure: removing the clusters resolves it, and nothing else can. -With `spec.source.external` the operator installed nothing, so deletion removes +With `spec.source.managed` the operator installed nothing, so deletion removes the object and touches neither the endpoint nor its data. The same `StorageCluster` hold still applies, because a namespace whose clusters have no control plane to reach is a namespace of objects nothing can reconcile. @@ -475,17 +501,17 @@ everything downstream asks the same question of it: where is the control plane, and is it ready. Which of the two produced the answer is this kind's business and nobody else's. -### 5.1 Managed +### 5.1 Local -The operator applies what the chart applies today: the `FoundationDBCluster` and -its RBAC, the document store, the management API's workload, its Services, and -its serving certificates. They become children of the `ControlPlane` by -controller reference, so the ownership spine starts at a real edge rather than at -a Helm release. +The operator applies the base control plane: the `FoundationDBCluster` and its +RBAC, the object store, the management API's workload, and the Services beside +it. They become children of the `ControlPlane` by controller reference, so the +ownership spine starts at a real edge rather than at a Helm release. ```go -// ManagedControlPlane is a control plane the operator installs and owns. -type ManagedControlPlane struct { +// LocalControlPlane is a control plane this cluster hosts, installed and owned +// by the operator. +type LocalControlPlane struct { // Image is the management API and control-plane image. // +kubebuilder:validation:Pattern=`^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$` // +kubebuilder:validation:Required @@ -494,38 +520,45 @@ type ManagedControlPlane struct { } ``` -#### Four components may already be in the cluster - -The management API's workload is the operator's to install. Reaching it needs four -things that are not: a FoundationDB operator to reconcile the -`FoundationDBCluster`, a MongoDB operator to reconcile the document store, an -issuer for the serving certificates, and a `StorageClass` for the volumes both -databases claim. Each is a component a cluster may already run, and each is -detected by asking the API server what it serves. - -| Component | Detected by | Where it is absent | -|-----------------------|--------------------------------------------------------------|----------------------------------------------------------------| -| FoundationDB operator | `apps.foundationdb.org/v1beta2` is served | The operator applies the CRDs and the controller | -| MongoDB operator | `mongodbcommunity.mongodb.com/v1` is served | §12 Q6 | -| Certificate issuer | `cert-manager.io/v1` is served, or the platform is OpenShift | `Installing` holds, and `status.message` names what is missing | -| `StorageClass` | A default `StorageClass` exists | §12 Q7 | - -**A detected component is used and never re-applied.** A cluster running the -FoundationDB operator for something else runs one operator afterward, and the -`FoundationDBCluster` this design creates is reconciled by it. - -**An installed component is cluster-scoped and carries no controller reference**, -for the reason +#### What the install depends on, and who answers for it + +The management API's workload is the operator's to install. Reaching it needs +things that are not, and each is settled by who is able to answer for it rather +than by one detection rule. + +| Dependency | How it is settled | +|----------------------------|-----------------------------------------------------------------------------------------------------| +| `FoundationDBCluster` CRDs | A prerequisite. `Installing` holds and names it where `apps.foundationdb.org/v1beta2` is not served | +| FoundationDB controller | Always applied, under the name the chart gave it, in the `ControlPlane`'s namespace | +| Document store | None. A base deployment runs no MongoDB, and the object store takes the step's place | +| Certificate issuer | Not expressible. §12 Q2 records that TLS stays with the chart | +| `StorageClass` | Not detected. An unset class falls to the cluster's default (§12 Q7) | + +**The CRDs and the controller are split because the API server cannot tell them +apart.** Whether `apps.foundationdb.org/v1beta2` is served says the CRDs are +installed, and says nothing about whether a controller is reconciling them: the +CRDs ship in the chart's `crds/` directory, which Helm applies on install and +never removes, so the group is served on every deployment either way. Creating a +`FoundationDBCluster` against a group the API server does not know is an error +rather than a wait, so the group is a prerequisite the install holds on. The +controller is applied instead, which leaves a cluster that also runs a +FoundationDB operator of its own in the position it was in before the install +moved. + +**An installed component that is cluster-scoped carries no controller +reference**, for the reason [`design-simplyblockdriver.md`](design-simplyblockdriver.md) §4.1 gives for the snapshot controller: deleting a CRD deletes every object of that kind in the cluster, including the ones another deployment created. -**The certificate issuer is detected rather than declared.** The chart takes it as -`tls.provider`, one of `openshift` or `cert-manager`, which is a fact about the -cluster written down by hand. The API answers it: OpenShift is identifiable from -the markers its distribution leaves -([`design-clusterdeploymentconfig.md`](design-clusterdeploymentconfig.md) §8.1), -and cert-manager from the group it registers. +**The step that installs a document store installs an object store instead.** +What keeps documents in a MongoDB is Graylog rather than the management API, and +the chart renders one only with observability enabled, so the reference deployment +reaches `Available` with no MongoDB at all. What a base deployment does need is +somewhere to keep what outlives a process, which is the metric history the control +plane keeps and the backups it writes, and that is an object store. The bucket is +made by a sidecar beside the server, which is where it can wait for the store to +answer. **The management API runs two instances by default, and the number is load-bearing.** A second instance is what lets a pod be replaced while the control @@ -549,18 +582,21 @@ version becomes a field the operator reconciles rather than a value baked into a release, which is the same argument [`design-simplyblockdriver.md`](design-simplyblockdriver.md) §8 makes for the CSI driver and the reason [`design-crd-model.md`](design-crd-model.md) §6 draws the edge at all. §12 Q2 is -the transition, which cannot be a flag day. +what the chart still renders, and the transition for a deployment already running +a chart-installed control plane. -### 5.2 External +### 5.2 Managed -An external control plane is an endpoint and a credential. +A managed control plane is an endpoint and, where it needs one, a credential. ```go -// ExternalControlPlane is a control plane that already exists. The operator -// installs nothing and owns nothing; it resolves, probes, and reports. -type ExternalControlPlane struct { - // Endpoint is the management API's base URL. Rejected unless it resolves to - // an external address, which is the SSRF guard atlas-lib/net carries. +// ManagedControlPlane is a control plane somewhere else, which this cluster's +// storage is managed by rather than hosting. The operator installs nothing and +// owns nothing here: it resolves the endpoint, probes it, and reports. +type ManagedControlPlane struct { + // Endpoint is the management API's base URL. It is validated against the + // same outbound-URL guard every other outbound endpoint in this group uses, + // so a loopback or link-local address is rejected. // +kubebuilder:validation:Pattern=`^https?://[a-zA-Z0-9.-]+(:[0-9]{1,5})?(/.*)?$` // +kubebuilder:validation:Required Endpoint string `json:"endpoint"` @@ -572,12 +608,19 @@ type ExternalControlPlane struct { and because a token in a spec is a token in every `kubectl get -o yaml`, every GitOps repository, and every audit log entry that records the object. -**No `ControlPlaneOps` action applies to an external control plane** (§6). +**It is optional, and absent means the endpoint is reached without one.** The +in-cluster case is what makes that necessary: a control plane the chart installed +answers on a `ClusterIP` Service in this namespace and carries no static token for +the readiness read, so a deployment pointing at it has no Secret to name. Naming a +Secret that does not exist stays an error, because naming one is a statement that +the control plane needs it. + +**No `ControlPlaneOps` action applies to a managed control plane** (§6). Restarting and upgrading are both operations on a workload, and the operator -installed no workload here. What it does for an external control plane is resolve -the endpoint, probe it, and report what it says. +installed no workload here. What it does instead is resolve the endpoint, probe +it, and report what it says. -**An external control plane may be shared, and the operator must assume it is.** +**A managed control plane may be shared, and the operator must assume it is.** Another Kubernetes cluster, or another namespace, may hold clusters in the same FoundationDB. Nothing this operator does may assume it is the only writer, which is a property every controller already needs for a different reason: the control @@ -588,7 +631,7 @@ involvement ([`design-crd-model.md`](design-crd-model.md) §7.7). ## 6. ControlPlaneOps -Declared in `operator/api/v1alpha1/controlplaneops_types.go`, short name `cpops`, +Declared in `operator/api/v1alpha2/controlplaneops_types.go`, short name `cpops`, and reconciled by `ControlPlaneOpsReconciler` in `operator/internal/controllers/controlplane/controlplaneops_controller.go`, beside the entity's own. The type is Appendix B. @@ -599,9 +642,9 @@ object somebody edited by hand is put back by the next reconcile (§4.3). What i left is what this kind carries, and it is three things: recycling a workload, moving it to a new version, and asking FoundationDB for a backup. -**Every action requires a managed control plane.** Each acts on what the operator +**Every action requires a local control plane.** Each acts on what the operator installed: `Restart` recycles a workload, `Upgrade` replaces its image, and -`Backup` asks the `FoundationDBCluster` the operator applied. An external control +`Backup` asks the `FoundationDBCluster` the operator applied. A managed control plane is an endpoint and a credential (§5.2), owning none of those, so there is nothing for any of the three to act on. @@ -610,12 +653,11 @@ nothing for any of the three to act on. type ControlPlaneOpsAction string ``` -| Action | Steps | What it is for | Source | -|------------|------------------------------------------------------------------|-------------------------------------------------------------------|---------| -| `Restart` | `Draining` → `Restarting` → `Awaiting` | A workload is wedged and has to be recycled | Managed | -| `Upgrade` | `Preflight` → `Draining` → `Applying` → `Awaiting` → `Verifying` | Moving the control plane to a new version | Managed | -| `Backup` | `Requesting` → `Awaiting` | Asking FoundationDB for a backup outside whatever schedule exists | Managed | -| `Rollback` | `Requesting` → `Awaiting` | | Managed | +| Action | Steps | What it is for | Source | +|-----------|------------------------------------------------------------------|-------------------------------------------------------------------|--------| +| `Restart` | `Draining` → `Restarting` → `Awaiting` | A workload is wedged and has to be recycled | Local | +| `Upgrade` | `Preflight` → `Draining` → `Applying` → `Awaiting` → `Verifying` | Moving the control plane to a new version | Local | +| `Backup` | `Requesting` → `Awaiting` | Asking FoundationDB for a backup outside whatever schedule exists | Local | **`Restart` takes a component scope.** `spec.restart.components` names entries from §4.3's table, and an empty list recycles the whole control plane. Restarting @@ -626,8 +668,9 @@ workload names. **An operation that recycles a workload drains first.** The management API is what every controller in the operator talks to, so replacing or restarting it mid-flight fails whatever is in flight. `Draining` holds while any -`StorageClusterOps`, `StorageNodeOps`, or `PersistentVolumeOps` in the namespace -is `Running`, emits `OperationsInFlight`, and proceeds when the last one finishes. +`StorageClusterOps`, `StorageNodeOps`, or `StoragePoolOps` in the namespace is +`Running`, which is every `Ops` kind that reaches the control plane over HTTP. It +emits `OperationsInFlight` and proceeds when the last one finishes. It does not cancel them, because an operation canceled to make a restart convenient is a worse outcome than a restart that waited. @@ -650,18 +693,18 @@ is skipped otherwise. A task-runner restart is the case worth naming: it skips t drain, because its queue is what makes the interruption a delay rather than a lost operation. -**An operation naming an external control plane is rejected at creation.** +**An operation naming a managed control plane is rejected at creation.** `ControlPlaneOpsValidator`, in `operator/internal/webhook/controlplaneops_validator.go`, resolves `spec.controlPlaneRef` on `create` and denies the request when it names no -`ControlPlane` or names an external one. `ReplicationOpsValidator` carries the +`ControlPlane` or names one that is not local. `ReplicationOpsValidator` carries the same shape for the same reason: an operation that can only fail belongs in an error message on the terminal that wrote it. **Admission is where the check can live because the answer cannot move.** `spec.controlPlaneRef` is immutable (Appendix B) and `ControlPlane.spec.source` is -immutable (§3.2), so a control plane admitted as managed stays managed for the -life of the operation. What can still happen is the target being deleted, and an +immutable (§3.2), so a control plane admitted as local stays local for the life +of the operation. What can still happen is the target being deleted, and an operation whose target has vanished is a missing-target failure every `Ops` kind in the group handles. @@ -676,9 +719,21 @@ rolling a Deployment to its current image produces no change to verify. A restar has neither precondition, since recycling a wedged workload is what it is for. **`Verifying` is what makes an upgrade more than an image bump.** It re-probes -readiness, compares the reported version against what was asked for, and fails -the operation when they disagree, so a rollout that started and did not finish is -a `Failed` operation rather than an `Available` control plane running the old version. +readiness, compares the reported version against what was asked for, and fails the +operation when they disagree, so a rollout that started and did not finish is a +`Failed` operation rather than an `Available` control plane running the old +version. + +**Where no version is served, the step passes and says that it did not compare.** +§8 records that `/_meta/version` does not exist yet, and an upgrade that could +never succeed against a control plane answering every other read is a worse +outcome than one whose record says the comparison was skipped. The step therefore +distinguishes three answers rather than two: a version that matches is a success, +a version that disagrees is a failure, and no version at all is a success carrying +a note that the rollout finished unverified. The note is on the operation rather +than in a log, because what it costs is exactly that an administrator reading the +operation afterward cannot tell a completed upgrade from a rollout that fell back, +and that has to be legible where the answer is read. **`Backup` asks rather than implements.** `Requesting` creates or updates a `FoundationDBBackup` naming the cluster and the destination from @@ -744,11 +799,22 @@ step degenerates into a readiness probe that cannot tell a completed upgrade fro a rollout that failed back (§6), and the version-skew pair against the CSI driver has one half ([`design-simplyblockdriver.md`](design-simplyblockdriver.md) §5). -**It is a prerequisite rather than a degradation to design around.** An `Upgrade` -that cannot verify is an operation reporting success on the evidence that -something answered, which is the failure the step exists to catch. The endpoint is -the control plane's to add, and the work that depends on it is sequenced behind -it rather than built to cope without it. +**Each of the three copes with its absence rather than waiting for it.** +`status.version` stays unwritten, because a field declared and never written +reports a definite-looking nothing +([`design-crd-model.md`](design-crd-model.md) §7.9) and an absent version is what +this one has to say. `simplyblock_controlplane_version_info` is deferred on the +same grounds (§9.2). `Upgrade` ships and its `Verifying` step reports what it +could and could not establish (§6). + +**What that costs is stated on the operation rather than hidden in the phase.** An +upgrade verified against no version is an operation reporting success on the +evidence that something answered, which is the failure the step exists to catch. +The alternative is refusing every upgrade of a control plane that serves no +version, which is every control plane today, so the choice is between an +unverified upgrade that says so and no upgrade at all. The endpoint is the control +plane's to add, and the step tightens to a comparison the moment it exists, +without any of the three needing to change shape. --- @@ -772,7 +838,7 @@ Events land on the object an administrator has open. For this kind that is the | An installation step is waiting on FoundationDB | `Normal` | `AwaitingDependency` | `ControlPlane` | | An installation step's deadline expired | `Warning` | `StepDeadlineExceeded` | `ControlPlane` | | A deletion is held because clusters still exist | `Warning` | `ClustersStillPresent` | `ControlPlane` | -| The external endpoint could not be resolved or reached | `Warning` | `EndpointUnreachable` | `ControlPlane` | +| A managed endpoint could not be resolved or reached | `Warning` | `EndpointUnreachable` | `ControlPlane` | | The credentials Secret is missing or malformed | `Warning` | `CredentialsError` | `ControlPlane` | | A `Backup` run created a `FoundationDBBackup` | `Normal` | `BackupRequested` | `ControlPlaneOps` | | A `Backup` run triggered the one already configured | `Normal` | `BackupTriggered` | `ControlPlaneOps` | @@ -792,7 +858,7 @@ event on one object. **`ControlPlaneDegraded` fires while every request is still being served** (§4.3), so it reaches an administrator before an outage rather than during one, and nothing holds on it. It names the component and its two counts, because one reason -covering eleven workloads sends a reader to `status.components` anyway. +covering every watched workload sends a reader to `status.components` anyway. The registered controller's `FDBReady` and `FDBNotReady` become `ControlPlaneReady` and `ControlPlaneNotReady`, because FoundationDB is one of @@ -810,7 +876,7 @@ the things that can be unready and the probe does not distinguish them. | `simplyblock_controlplane_install_step_duration_seconds` | `namespace`, `step` | Histogram of per-step installation duration (§4.2) | | `simplyblock_controlplane_operations_total` | `namespace`, `action`, `result` | Operations reaching a terminal phase | | `simplyblock_controlplane_operation_duration_seconds` | `namespace`, `action` | Histogram of operation durations | -| `simplyblock_controlplane_version_info` | `namespace`, `version` | Gauge, 1 for the reported version, so a skew against the driver is graphable | +| `simplyblock_controlplane_version_info` | `namespace`, `version` | Gauge, 1 for the reported version, so a skew against the driver is graphable. Deferred with the endpoint that serves it (§8) | **`simplyblock_controlplane_version_info` is half of a pair.** Its other half is `simplyblock_simplyblockdriver_version_info`, in @@ -851,10 +917,10 @@ that a FoundationDB which never reaches quorum expires the step rather than hanging, needs `envtest` with the FoundationDB CRDs installed and a real cluster for the timing. -The external mode's risk is different and smaller: it is a URL, a Secret, and a -probe, and all three are unit-testable. What is not is a shared external control -plane with a second writer, which is the scenario §5.2 says the operator must -assume and nothing exercises. +The managed mode's risk is different and smaller: it is a URL, a Secret, and a +probe, and all three are unit-testable. What is not is a shared control plane with +a second writer, which is the scenario §5.2 says the operator must assume and +nothing exercises. --- @@ -862,10 +928,10 @@ assume and nothing exercises. | Registered | This design | Cost | |-----------------------------------------------|---------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| `spec.image` | `spec.source.managed.image` (§5.1) | Spec regrouping, and the field stops doubling as the storage-node default (see below) | -| No way to name an external control plane | `spec.source.external` (§5.2) | Additive, and it is the half of `design-crd-model.md` §6 that does not exist today | +| `spec.image` | `spec.source.local.image` (§5.1) | Spec regrouping, and the field stops doubling as the storage-node default (see below) | +| No way to name a control plane elsewhere | `spec.source.managed` (§5.2) | Additive, and it is the half of `design-crd-model.md` §6 that does not exist today | | The endpoint is `SIMPLYBLOCK_WEBAPI_BASE_URL` | `status.endpoint` (§3.3) | Behavioral. Every caller of `webapi.NewClient()` moves, which is every controller in the operator | -| The chart installs the control plane | The operator installs it (§5.1) | The largest piece of work here. `controlplane.managedByOperator` is the flag, and it defaults to the operator. Adopting a control plane the chart already installed is the open half (§12, Q2) | +| The chart installs the control plane | The operator installs it (§5.1) | The largest piece of work here. `deployment.profile` is the chart's flag and it selects the mode rather than the installer. Adopting a control plane the chart already installed is the open half (§12, Q2) | | `status.phase` untyped, two values | `ControlPlanePhase`, four values (§3.3) | `Degraded` and `Unavailable` are both new, and together they separate a control plane that is impaired from one that is not answering. `Ready` becomes `Available`, which is the opposite of `Unavailable` where `Ready` was not | | No step field | `status.step` (§4.2) | Additive. A stalled install currently reports one message and no position | | No `observedGeneration` | Present (§3.3) | Required by `design-crd-model.md` §7.9 | @@ -876,8 +942,8 @@ assume and nothing exercises. **The `spec.image` move is not only a regrouping.** The field is inherited by every `StorageNodeSet` that omits `spec.clusterImage`, so moving it under -`spec.source.managed` would leave the external mode with no storage-node default -at all, which is wrong: the storage-node image has nothing to do with where the +`spec.source.local` would leave the managed mode with no storage-node default at +all, which is wrong: the storage-node image has nothing to do with where the control plane runs. It belongs with the other fleet defaults, in `StorageCluster.spec.storageNodes.image` ([`design-storagenode.md`](design-storagenode.md) §5.1), whose doc comment @@ -899,18 +965,21 @@ from review history. §3.1 is where the answer went: a Kubernetes cluster holds already impose on the operator. **Q2 is settled for a fresh install and open for an existing one.** -`controlplane.managedByOperator` is the chart flag, and it defaults to the -operator: a new deployment gets the control plane §5.1 describes, and the chart -renders only the `ControlPlane` object beside the FoundationDB CRDs, the -Prometheus configuration, and the log-collector RBAC. Setting it to false hands -the templates back, unchanged, which is what a deployment does when it needs -something the spec cannot yet express. +`deployment.profile` is the chart flag, and it names the mode rather than the +installer: `standalone` renders a `ControlPlane` with `spec.source.local`, and +`managed` renders one with `spec.source.managed` pointing at a control plane +elsewhere. Either way the operator is what installs, and the templates that used +to install a control plane locally are gone rather than gated. There is no flag +that hands them back, because a deployment that needs something the spec cannot +express is a deployment the spec has to grow a field for. What is open is the transition for a deployment already running a chart-installed -control plane. The apply is a server-side apply under a stable field manager, so -it takes over the objects a Helm release created rather than failing on them, but -nothing verifies the handover or strips the release's claim afterward, and §5.1's -ownership spine is only real once it has. That is the shape +control plane. The chart refuses to upgrade over one without an explicit +annotation saying to keep it, which turns a silent adoption into a decision, and +the operator's apply is a server-side apply under a stable field manager, so it +takes over the objects a Helm release created rather than failing on them. What +nothing does yet is verify the handover or strip the release's claim afterward, +and §5.1's ownership spine is only real once it has. That is the shape [`design-simplyblockdriver.md`](design-simplyblockdriver.md) §4.3 gives the CSI driver's adoption, and it is the half of this question still to answer. @@ -923,11 +992,10 @@ machine, every one is non-essential in §4.3's table, and they are gated behind chart value this kind has no field for. **TLS is the one configuration the install cannot express.** The chart serves the -control plane over TLS behind `tls.enabled`, and `ManagedControlPlane` has no -field for it, since §5.1 settles which issuer is detected rather than whether a -deployment wants one. The chart refuses `managedByOperator` together with -`tls.enabled` rather than installing a control plane in plaintext, and closing -that gap is a field on this kind. +control plane over TLS behind `tls.enabled`, and `LocalControlPlane` has no field +for it. The chart therefore refuses `tls.enabled` together with the `standalone` +profile rather than installing a control plane in plaintext, and closing that gap +is a field on this kind together with the issuer detection §5.1 does not perform. **Q3: Whether backup belongs to the action or to the spec.** `FoundationDBBackup` describes a continuous backup, carrying a `backupState` and a @@ -957,18 +1025,17 @@ this document should make on the control plane's behalf. They take the non-essential default meanwhile, so a wrong answer under-reports rather than halting a fleet. -**Q6: What supplies the MongoDB operator where a cluster has none.** §5.1 has the -document store applied as a `MongoDBCommunity` object, and the chart ships that -CRD while installing no operator to reconcile it, so a cluster without one accepts -the object and does nothing with it. The candidates are installing the community -operator alongside the FoundationDB one, replacing the document store with -something the management API already carries, and holding the install with a named -prerequisite. The first two are the control plane's to choose between, since what -it stores where is its own. - -**Q7: Whether the operator installs a `StorageClass`.** §5.1 detects a default -one. The chart applies a `hostpath.csi.k8s.io` provisioner and a `local-hostpath` -class, which suits a laptop and a single-node test. A FoundationDB whose volumes +Q6 is settled by measurement and its number is retired rather than reused, since +it is cited from review history. §5.1 is where the answer went: the base install +applies no document store at all, because what keeps documents in a +`MongoDBCommunity` is Graylog rather than the management API, and the reference +deployment reaches `Available` without one. What the step applies instead is an +object store. + +**Q7: Whether the operator installs a `StorageClass`.** It does not, and it does +not detect one either: an unset `storageClassName` falls to whatever the cluster's +default is. The chart applies a `hostpath.csi.k8s.io` provisioner and a +`local-hostpath` class, which suits a laptop and a single-node test. A FoundationDB whose volumes are host paths keeps its data on one machine, so the install either refuses without a class, names the hostpath path as a development mode, or keeps applying it. @@ -1017,7 +1084,7 @@ const ( // ControlPlanePhaseDegraded is one whose readiness probe passes while a // management API or FoundationDB pod is restarting. It answers every // request, so nothing holds on it and it exists to be read by a person. A - // control plane the operator does not manage never reaches it, because the + // control plane this cluster does not host never reaches it, because the // operator owns no pods there to watch. ControlPlanePhaseDegraded ControlPlanePhase = "Degraded" // ControlPlanePhaseUnavailable is one whose readiness probe fails: it @@ -1061,10 +1128,10 @@ type FoundationDBSpec struct { Resources corev1.ResourceRequirements `json:"resources,omitempty"` } -// ManagedControlPlane is a control plane the operator installs and owns. Its -// objects carry a controller reference to the ControlPlane, so the ownership -// spine starts at a real edge rather than at a Helm release. -type ManagedControlPlane struct { +// LocalControlPlane is a control plane this cluster hosts, installed and owned +// by the operator. Its objects carry a controller reference to the ControlPlane, +// so the ownership spine starts at a real edge rather than at a Helm release. +type LocalControlPlane struct { // Image is the management API and control-plane image. // +kubebuilder:validation:Pattern=`^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$` // +kubebuilder:validation:Required @@ -1096,13 +1163,25 @@ type ManagedControlPlane struct { // plane. // +optional Tolerations []corev1.Toleration `json:"tolerations,omitempty"` + + // NodeSelector pins every pod the operator installs for the control plane. + // It is a selector rather than an affinity term because that is what the + // chart it replaces took, and a deployment migrating off the chart has the + // value already written down. + // +optional + NodeSelector map[string]string `json:"nodeSelector,omitempty"` } -// ExternalControlPlane is a control plane that already exists. The operator -// installs nothing and owns nothing: it resolves, probes, and reports. -type ExternalControlPlane struct { +// ManagedControlPlane is a control plane somewhere else, which this cluster's +// storage is managed by rather than hosting. The operator installs nothing and +// owns nothing here: it resolves the endpoint, probes it, and reports. +// +// The word is about what manages the storage clusters rather than about who runs +// the operator. A deployment in this mode registers its clusters with a control +// plane it does not host, so the fleet is administered from there. +type ManagedControlPlane struct { // Endpoint is the management API's base URL. It is validated against the - // same outbound-URL guard every other external endpoint in this group uses, + // same outbound-URL guard every other outbound endpoint in this group uses, // so a loopback or link-local address is rejected. // +kubebuilder:validation:Pattern=`^https?://[a-zA-Z0-9.-]+(:[0-9]{1,5})?(/.*)?$` // +kubebuilder:validation:Required @@ -1111,8 +1190,14 @@ type ExternalControlPlane struct { // CredentialsSecretRef names a Secret in this namespace holding the bearer // token the operator authenticates with. It is a reference rather than a // field because a token in a spec is a token in every kubectl get -o yaml. - // +kubebuilder:validation:Required - CredentialsSecretRef corev1.LocalObjectReference `json:"credentialsSecretRef"` + // + // Absent means the endpoint is reached without one, which is the in-cluster + // case: a control plane the Helm chart installed answers on a ClusterIP + // Service in this namespace and does not require a token for the readiness + // read. Naming a Secret that does not exist stays an error, because naming + // one is a statement that the control plane needs it. + // +optional + CredentialsSecretRef *corev1.LocalObjectReference `json:"credentialsSecretRef,omitempty"` // CABundleSecretRef names a Secret holding the CA certificate the endpoint // is verified against. Absent means the system trust store. @@ -1120,29 +1205,45 @@ type ExternalControlPlane struct { CABundleSecretRef *corev1.LocalObjectReference `json:"caBundleSecretRef,omitempty"` } -// ControlPlaneSource selects where the control plane comes from. Exactly one -// member is set, which is what makes the two modes siblings rather than two -// unrelated top-level fields. +// ControlPlaneSource selects whether this cluster hosts its control plane or is +// managed by one elsewhere. Exactly one member is set, which is what makes the +// two modes siblings rather than two unrelated top-level fields, and which +// member it is cannot change afterward. +// +// Both rules are declared here rather than on the field that carries the block, +// for two separate reasons. +// +// The immutability is the interesting one. What is frozen is the choice between +// the two modes and not the block, because the members have to stay editable: +// changing spec.source.local.image is an ordinary edit, and it is what a +// ControlPlaneOps upgrade performs. Spelling it +k8s:immutable on the field would +// emit self == oldSelf over the whole struct, which freezes the image with it and +// makes that operation impossible to complete. +// +// The placement is the dull one. controller-gen emits a single field's +// marker-derived rules and its injected immutability rule into one list in an +// order that varies between runs, so a field carrying both produces a CRD that +// differs from itself and a drift check that fails at random. Two rules of the +// same kind on a type are emitted in source order. +// +kubebuilder:validation:XValidation:rule="(has(self.local) ? 1 : 0) + (has(self.managed) ? 1 : 0) == 1",message="set exactly one of local or managed" +// +kubebuilder:validation:XValidation:rule="has(self.local) == has(oldSelf.local) && has(self.managed) == has(oldSelf.managed)",message="spec.source is immutable: a control plane the operator installed and one it did not are different deployments, and the clusters and their volumes live in the FoundationDB behind the old one" type ControlPlaneSource struct { - // Managed is a control plane the operator installs. + // Local is a control plane the operator installs. // +optional - Managed *ManagedControlPlane `json:"managed,omitempty"` + Local *LocalControlPlane `json:"local,omitempty"` - // External is a control plane that already exists. + // Managed is a control plane that already exists. // +optional - External *ExternalControlPlane `json:"external,omitempty"` + Managed *ManagedControlPlane `json:"managed,omitempty"` } // ControlPlaneSpec is the desired state of the simplyblock control plane for one // namespace. type ControlPlaneSpec struct { - // Source selects where the control plane comes from. Immutable: switching a - // live deployment between an installed control plane and an existing one is - // not a reconfiguration, because the clusters and their volumes live in the - // FoundationDB behind the old one. - // +kubebuilder:validation:XValidation:rule="(has(self.managed) ? 1 : 0) + (has(self.external) ? 1 : 0) == 1",message="set exactly one of managed or external" + // Source selects where the control plane comes from. Which member is set is + // frozen by the rules on ControlPlaneSource; the member's own contents stay + // editable. // +kubebuilder:validation:Required - // +k8s:immutable Source ControlPlaneSource `json:"source"` } @@ -1186,8 +1287,8 @@ type ControlPlaneStatus struct { // +optional Step statemachine.KubeSnapshot `json:"step,omitempty"` - // Endpoint is the resolved management API base URL, derived in the managed - // case and echoed in the external one. It is what every controller in the + // Endpoint is the resolved management API base URL, derived in the local + // case and echoed in the managed one. It is what every controller in the // operator reads to reach the control plane, so that one object answers // where it is. // +optional @@ -1202,8 +1303,8 @@ type ControlPlaneStatus struct { LastChecked *metav1.Time `json:"lastChecked,omitempty"` // Components is the per-component readiness the phase is derived from - // (§4.3), one entry per workload the managed install applies. It is empty - // for an external control plane, which has no components the operator owns. + // (§4.3), one entry per workload the local install applies. It is empty for + // a managed control plane, which has no components the operator owns. // Without it a Degraded phase says that something is wrong and not what. // +optional // +listType=map @@ -1266,8 +1367,8 @@ type ControlPlaneList struct { ```go // ControlPlaneOpsAction is the operation a ControlPlaneOps performs. Every action -// acts on what the operator installed, so every action requires a managed control -// plane, and the validating webhook of §6 rejects an operation naming an external +// acts on what the operator installed, so every action requires a local control +// plane, and the validating webhook of §6 rejects an operation naming a managed // one at creation rather than letting it be created and fail. // +kubebuilder:validation:Enum=Restart;Upgrade;Backup type ControlPlaneOpsAction string diff --git a/operator/docs/designs/crd-redesign/design-crd-model.md b/operator/docs/designs/crd-redesign/design-crd-model.md index 44f6c5c47..1c8b89d23 100644 --- a/operator/docs/designs/crd-redesign/design-crd-model.md +++ b/operator/docs/designs/crd-redesign/design-crd-model.md @@ -2,8 +2,8 @@ **Status:** Draft **Author:** Christoph Engelbert (noctarius) -**Date:** 2026-08-19 (last updated 2026-09-08) -**API groups:** `storage.simplyblock.io/v1alpha1`, and `metrics.simplyblock.io/v1alpha2` for the readings that are not resources (§7.13) +**Date:** 2026-08-19 (last updated 2026-09-17) +**API groups:** `storage.simplyblock.io`, served at `v1alpha1` and `v1alpha2`, and `metrics.simplyblock.io/v1alpha2` for the readings that are not resources (§7.13) **Diagram:** [`assets/crd-overview.jpg`](assets/crd-overview.jpg) --- @@ -24,9 +24,9 @@ ## Overview -The API group `storage.simplyblock.io/v1alpha1` registers seventeen custom -resource definitions today, thirteen of which are in scope here, and the target -model drawn in +The registered API, `storage.simplyblock.io/v1alpha1`, carries seventeen custom +resource definitions, thirteen of which are in scope here, and the target model +drawn in [`assets/crd-overview.jpg`](assets/crd-overview.jpg) has roughly thirty boxes. This document is the map: which categories a kind can belong to, what its name has to look like once that category is chosen, which resource owns which, and @@ -438,17 +438,23 @@ the outer `Pending` to `Succeeded` spine an operation has. step machine is per-action, the phase machine is not, and a step is not contained in a phase. §9.5 is what the rename costs. -**No kind meets this rule yet.** `atlas-lib/statemachine` has no consumer in -either the operator or the CSI driver, and the three registered `Ops` kinds all -drive their steps by hand. +**The rule is what separates the kinds that were rebuilt against it from the ones +that were not.** The `StorageCluster`, `StorageNode`, `StoragePool`, +`StorageDevice`, `StorageBackup`, and `ControlPlane` families drive their steps +from a declared `statemachine` graph and carry the step-set agreement test that +proves the graph and the `Enum` marker have not drifted apart. `OperatorOps` is +the one `Ops` kind still driving its steps from a hand-rolled `switch`, which is +the shape §1 lists as what motivated the rule. ### 3.2 The lock an entity carries **Every entity with an `Ops` companion carries `status.activeOpsRef`**, a string naming the operation currently allowed to act on it, and empty when none is. The field has the same name and the same meaning on every kind, so a reader, a script, -and a dashboard learn it once rather than per kind. Four kinds carry it today: -`StorageCluster`, `StorageNode`, and the two replication kinds (§2). +and a dashboard learn it once rather than per kind. Every entity with a companion +carries it: `StorageCluster`, `StorageNode`, `StoragePool`, `StorageDevice`, +`ControlPlane`, and `StorageBackup`, alongside the two replication kinds this +document leaves out of scope (§2). **One operation at a time per entity, and that is the design rather than a limit to work around.** A second operation is admitted by the API server, acquires @@ -898,14 +904,13 @@ every kind this group models reads its backend state from a subscription: the control plane honors `?watch=true` on the type, the operator holds the streamed objects in an in-memory store, and reconcilers read that store. -**None of it is shipped, and every design here depends on it.** The subscriptions -arrive with the control plane's SSE work rather than with any design in this group, -which is why a `?watch=true` row in a backend table is an external dependency rather -than an endpoint somebody can call today. +**The store and the subscription manager are `operator/internal/cpinformer`**, and +the streams a design names as a `?watch=true` row are its subscriptions. The +cluster, task, node, device, volume, backup, and backup-policy streams are +consumed; the pool is the one resource in the family still read by polling, which +[`design-storagepool.md`](design-storagepool.md) §4.2 records. `design-sse-push-notifications.md`, on the `sse` branch, owns the mechanism, the -verified wire contract, and the adoption phases, alongside the -`operator/internal/cpinformer` implementation of the store and the subscription -manager. +verified wire contract, and the adoption phases. **No design in this group specifies a poll.** A `RequeueAfter` still appears where a controller is waiting on something the stream does not carry, and as a slow @@ -1141,16 +1146,28 @@ all. What the chart ships for them is an `APIService` and the two bindings the Kubernetes API server's authentication and authorization delegation needs, which is also what makes an ordinary `RoleBinding` on the group work. -| Kind | Short name | Scope | Reports | Named after | -|------------------------|------------|------------|-----------------------------------------------------------------------------------------------------|-------------------------------------------------------------------------| -| `LogicalVolumeMetrics` | `lvm` | Namespaced | A volume's provisioned, used, free, and total bytes, and the control plane's own utilization figure | The `PersistentVolumeClaim` the volume backs, in that claim's namespace | +| Kind | Short name | Scope | Reports | Named after | +|-------------------------|------------|------------|-----------------------------------------------------------------------------------------------------|-------------------------------------------------------------------------| +| `LogicalVolumeMetrics` | `lvm` | Namespaced | A volume's provisioned, used, free, and total bytes, and the control plane's own utilization figure | The `PersistentVolumeClaim` the volume backs, in that claim's namespace | +| `StorageClusterMetrics` | `scm` | Namespaced | How full a whole cluster is, which is the number a capacity plan is made against | The `StorageCluster`, in its namespace | +| `StoragePoolMetrics` | `spm` | Namespaced | A pool's occupancy against the capacity it was carved out with | The `StoragePool`, in its namespace | +| `StorageNodeMetrics` | `snm` | Namespaced | One node's occupancy across the devices it carries | The `StorageNode`, in its namespace | +| `StorageDeviceMetrics` | `sdm` | Namespaced | One device's occupancy | The `StorageDevice`, in its namespace | + +**Each narrower reading answers a question the one above it cannot.** A device's +reading is bounded by one device and a pool's by the capacity that pool was carved +out with, so neither says whether the cluster underneath them is about to run out, +which is why the cluster carries its own. Which pool or which volume is filling a +cluster up is deliberately not on the cluster's reading: a cluster reports its own +totals, and the breakdown behind them is what the narrower kinds are for. **A reading is named for the Kubernetes object it is about, and lives where that object lives.** A tenant who knows their claim's name needs to learn nothing else to ask for its occupancy, and namespaced RBAC confines them to their own volumes without a single rule this group has to invent. A logical volume with no bound claim is therefore not served, because it has no name in this API and no namespace -to be authorized against. The plural is its own singular, the way `endpoints` is, +to be authorized against, and the same rule makes every other kind here namespaced: +each is named after an object that is. The plural is its own singular, the way `endpoints` is, since "metrics" is already the noun. **A measured number stays in a CRD's status only when it is bounded, diff --git a/operator/docs/designs/crd-redesign/design-property-renames.md b/operator/docs/designs/crd-redesign/design-property-renames.md index 65b922956..c8f3bdd76 100644 --- a/operator/docs/designs/crd-redesign/design-property-renames.md +++ b/operator/docs/designs/crd-redesign/design-property-renames.md @@ -43,8 +43,10 @@ that silently never ran. This is the one class that needs no data migration and only a deprecation window for the sake of scripts and runbooks. **A renamed toggle that also inverts is the worst case**, because the mechanical -migration produces the opposite of the intended behavior. There is exactly one such -row, `skipKubeletConfiguration`, and `design-crd-model.md` §9.6 flags it twice. +migration produces the opposite of the intended behavior. One row is in it, +`migrationEnabled` becoming `disableMigration`. One more is adjacent: +`enableDataRealignment` keeps its polarity and changes its default, so an object +nobody edited loses the feature unless the conversion records what was stated. --- @@ -101,8 +103,8 @@ and one is not registered at all. | Struct | Registered | Target | Default | Class | |-------------------------------|----------------------------|------------------------------------|---------|----------------------| -| `StorageNodeSpec` | `skipKubeletConfiguration` | `enableKubeletConfiguration` | off | Inverting | -| `StorageNodeSetSpec` | `skipKubeletConfiguration` | `enableKubeletConfiguration` | off | Inverting | +| `StorageNodeSpec` | `skipKubeletConfiguration` | Removed, re-landed per cluster | off | Removal | +| `StorageNodeSetSpec` | `skipKubeletConfiguration` | Removed, re-landed per cluster | off | Removal | | `VolumeAutoPlacementSettings` | `migrationEnabled` | `disableMigration` | on | Inverting | | `VolumeAutoPlacementSettings` | `latencyBenchmarkEnabled` | `enableLatencyBenchmark` | off | Silent | | `VolumeAutoPlacementSettings` | `enabled` | `spec.enableVolumeAutoPlacement` | off | Silent, and moves up | @@ -121,22 +123,33 @@ struct has no such field. The row is a target-state addition rather than a renam so it is not migrated; it is created named correctly whenever replication becomes expressible on a class. -**Three rows are removals rather than renames**, and `design-storagecluster.md` -§12 gives the reason for all three `BackupSpec` ones: the store is a location, and +**Six rows are removals rather than renames**, and `design-storagecluster.md` +§12 gives the reason for the three `BackupSpec` ones: the store is a location, and how a copy is taken belongs to the control plane, which keeps accepting these values and applies its own defaults once the operator stops sending them. `VolumeMigrationSettings.enabled` is removed because migration cannot be turned off — a drain, a rebalance, and a device replacement are all performed by moving volumes. -**`skipKubeletConfiguration` inverts.** A deprecation window that reads the old -field and writes the new one has to negate it. Reading `skipKubeletConfiguration: -true` and writing `enableKubeletConfiguration: true` configures the kubelet on a -node that asked not to be configured. - -**The chart carries these names too.** `skipKubeletConfiguration` appears in -`helm-charts/charts/simplyblock-operator/values.yaml`, and `multiCluster.enable` -is a third spelling that exists only there. +**`skipKubeletConfiguration` leaves the node kinds rather than inverting on +them.** Its only consumer is an environment variable in a DaemonSet pod template, +and a DaemonSet is one object for every node it schedules, so a per-node field +never reached it. The toggle re-lands on the cluster as +`StorageCluster.spec.storageNodes.enableKubeletConfiguration`, positive-formed and +off by default, and a `ClusterDeploymentConfig` fills it from the Kubernetes +distribution it names rather than leaving it to be stated node by node +([`design-clusterdeploymentconfig.md`](design-clusterdeploymentconfig.md) §4.2). +The registered field is therefore a removal that stashes under +`storage.simplyblock.io/v1alpha1-spec.overrides.skipKubeletConfiguration` (§3.3), +and nothing negates, because nothing on the node reads the value again. +[`design-storagenode.md`](design-storagenode.md) §15.1 owns the move. + +**One spelling exists only in the chart.** `multiCluster.enable` in +`helm-charts/charts/simplyblock-operator/values.yaml` names no API field, so no +conversion reaches it. It feeds a ConfigMap and a Secret the CSI driver reads. The +kubelet toggle has no chart spelling at all, because the only template that read +it was the chart's own storage-node DaemonSet and the operator renders that +workload. ### 2.4 Regroupings — renames that also move @@ -285,9 +298,12 @@ webhook that is down does not degrade the group, it makes it unreadable: a `kubectl get storagecluster` fails rather than returning the old shape. During an upgrade, which is when the webhook's own deployment is being replaced, is when this is most likely, and a `failurePolicy` cannot help because there is no meaningful -answer to fall back to. What contains it is that the webhook is served by the -operator's existing manager, so it is up whenever the operator is, and the operator -being down already stops reconciliation. +answer to fall back to. What contains it is that the webhook is a process of its +own ([`design-api-upgrade.md`](design-api-upgrade.md) §6.1). It runs no +controllers, holds no leader-election lease, and reads no converted kind, so the +custom resources an administrator reads in order to diagnose a failed operator +stay readable while the operator is down, and its own start-up waits on nothing +that has to be converted first. **Conversion has to round-trip, and three rows lose data.** `BackupSpec.withCompression`, `snapshotBackups`, `localTesting`, and @@ -354,13 +370,20 @@ and a conversion that errors makes the object unreadable rather than invalid. ### 3.4 The inverting rows need their own tests -For `skipKubeletConfiguration`, `migrationEnabled`, and the toggles of §2.3 whose -default is on, the conversion negates rather than copies. Each needs a test that -asserts the *behavior* rather than the field value: a node that set -`skipKubeletConfiguration: true` must still not have its kubelet configured after -the upgrade. Asserting `enableKubeletConfiguration == false` would pass against a -conversion that got the polarity right and against one that never ran at all, since -`false` is also the zero value. +For `migrationEnabled` and the toggles of §2.3 whose default is on, the conversion +negates rather than copies. Each needs a test that asserts the *behavior* rather +than the field value: a cluster that set `migrationEnabled: false` must still not +migrate volumes after the upgrade. Asserting `disableMigration == true` would pass +against a conversion that got the polarity right and against one that never ran at +all, since the stored object carries a value either way. + +`enableDataRealignment` needs the same test for a different reason. Its polarity +holds and its default changes, from on to off, so an object that stated nothing is +indistinguishable in the stored shape from one that deliberately turned the feature +off. The conversion writes the value into an annotation on every trip down and +reads that annotation's absence on the way up as the mark of a client that only +ever spoke `v1alpha1`, which is the discrimination a test has to exercise from both +sides. The same test has to run in both directions. A conversion that negates going up and copies going down is a bug that a one-way test cannot see, and it corrupts on the @@ -375,34 +398,37 @@ deprecated in an event, and the `StorageClass` parameter keys are read under bot indefinitely, because `StorageClass.parameters` is immutable and a class an older operator generated can never be rewritten. -The chart's own value names are outside it too. `skipKubeletConfiguration` in -`values.yaml` is a chart input rather than an API field, so the conversion never -sees it, and it needs the both-spellings treatment in the template or a documented -break. +The chart's own value names are outside it too. A chart input is not an API field, +so the conversion never sees one, and a value whose only consumer moves into the +operator leaves the chart rather than gaining a second spelling in it. +`skipKubeletConfiguration` took that path (§2.3). -### 3.6 StorageNodeSet keeps one version, and the node reparents lazily +### 3.6 StorageNodeSet keeps one version, and the node reads its parent from its owner `StorageNodeSet` is retired by [`design-crd-model.md`](design-crd-model.md) §9.2, and a kind on its way out does not earn a second version. Its one row, `skipKubeletConfiguration`, is dropped -rather than migrated: the field's replacement lives on `StorageNode.spec.config` -(§2.3), which is a kind that does gain a `v1alpha2`, so nothing is lost by leaving -the retiring kind spelled as it shipped. - -**What this costs is that `StorageNode.spec.storageNodeSetRef` cannot be converted, -and that is the right outcome rather than a gap.** The target is -`spec.clusterRef` plus `spec.nodeSet` (§2.7), and the cluster's name is not in the -node — it is in the `StorageNodeSet` the node points at. A conversion function -cannot go and read it: conversion runs on every read of the object, has to be a -pure function of what it was handed, and a conversion that issues API calls turns -one `kubectl get` into two and fails the read when the second one does. - -**So `v1alpha2` carries all three fields, and the controller fills the new two.** -`storageNodeSetRef` survives into `v1alpha2` as an optional deprecated field that -converts by copy, and `clusterRef` and `nodeSet` are optional beside it. A node -reconciled with the old field set and the new ones empty resolves the set, writes -the cluster and the group name, and from then on the node names its cluster -directly. A node created with the new fields never needs the old one. +rather than migrated: the field's replacement is a toggle on the cluster (§2.3), so +nothing is lost by leaving the retiring kind spelled as it shipped. + +**`StorageNode.spec.storageNodeSetRef` has no `v1alpha2` spelling.** The hub names +its parent directly as `spec.clusterRef` and keeps the set's name as `spec.nodeSet`, +a label nothing is fetched by (§2.7). The cluster's name is not in the node at all +(it is in the `StorageNodeSet` the node points at), and a conversion function cannot +go and read it: conversion runs on every read of the object, has to be a pure +function of what it was handed, and a conversion that issues API calls turns one +`kubectl get` into two and fails the read when the second one does. + +**The controller owner reference carries the same fact, on the object.** A node +converting up reads its `StorageCluster` from the reference rather than from a +field, which is a pure function of the subject and needs no client. What makes the +reference answer is ordering: the upgrade's reparent-storage-nodes step writes it +before the storage version moves +([`design-api-upgrade.md`](design-api-upgrade.md) §20). A node that has not been +reparented yet converts to an empty `clusterRef` rather than to an error, so the +object stays readable, which is what a read during an upgrade needs. The field is +required, so the next write of that node is refused until the reparent has run. +That is a louder failure than a node silently joining no cluster. **The workload reparents on the same schedule, not at start-up.** The DaemonSet, the Services, the certificates, and the per-node ConfigMaps are owned by the @@ -418,10 +444,12 @@ mistake visible before it is fleet-wide. This is the retirement's business rather than the renames', and it is stated here because it is the reason this document leaves one row unconverted. -### 3.7 The trust has to exist before the operator starts +### 3.7 The trust has to exist before conversion is asked for -The operator both serves the conversion webhook and reads the kinds it converts, -and that is a cycle rather than a coincidence. +A conversion webhook is in the read path of the kinds it converts, so the API +server has to trust it before anything reads one. What makes the ordering awkward +is that the material it trusts lives in the cluster the webhook is being started +in. **A controller-runtime manager starts its HTTP servers, then its webhook servers, then syncs its caches, and only then runs everything else.** The first two orders @@ -430,41 +458,53 @@ first *because* a cache sync over a converted kind lists it at the hub version, which makes the API server convert every stored object, which calls the webhook. What the manager cannot order is anything that is not one of those servers. -**So a CA injected by a Runnable arrives too late by construction.** The list -fails, and the injection that would have fixed it never runs, because it sits on -the far side of the sync that is failing. This is a bootstrap deadlock rather than -a race: waiting longer never resolves it. +**So a CA injected by a Runnable of the converting process arrives too late by +construction.** The list fails, and the injection that would have fixed it never +runs, because it sits on the far side of the sync that is failing. That is a +bootstrap deadlock rather than a race: waiting longer never resolves it. **The symptom is quieter than a crash, which is what makes it worth stating.** The cache sync blocks until the process is canceled rather than giving up, and the health probes are served by the HTTP servers that started before it. The pod therefore reports Ready and keeps reporting Ready while reconciling nothing. There -is no restart to notice and no `CrashLoopBackOff` to find — the operator looks +is no restart to notice and no `CrashLoopBackOff` to find: the process looks healthy and is inert, which is the hardest shape of failure to attribute. -**The serving certificate and the CA bundle are therefore provisioned before the -manager is constructed**, through a direct client rather than the manager's. The -two kinds this touches, `Secret` and `CustomResourceDefinition`, are core and -apiextensions kinds that no conversion webhook stands in front of, so the -bootstrap can always make progress no matter what state the converted kinds are -in. Rotation stays where it was: the certificate machinery keeps running under the -manager and re-injects whenever the material changes, and the bootstrap only -guarantees that the first pass has already happened. - -**Reusing existing material matters as much as creating it.** An operator that +**The conversion webhook is therefore a process that reads no converted kind.** +It runs no controllers, holds no leader-election lease, and lists none of the +kinds it converts, so it has no cache sync to deadlock and can provision its +serving certificate and inject the CA into the converting CRDs before it answers +anything. The two kinds that bootstrap touches, `Secret` and +`CustomResourceDefinition`, are core and apiextensions kinds that no conversion +webhook stands in front of, so it can always make progress no matter what state +the converted kinds are in. It ships in the operator image under a second entry +point, so the conversion code and the API types it converts between are versioned +with the operator that reads them +([`design-api-upgrade.md`](design-api-upgrade.md) §6.1, §29.3). + +**The operator corrects what the manifests cannot state.** A CRD is cluster-scoped +and ships in the chart's `crds/` directory, which Helm does not template, so its +service reference names the namespace of the default install and is wrong for +every other one. The API server cannot reach a conversion webhook it cannot +resolve, and the operator is the only party that knows which namespace it is in, +so it rewrites the reference and re-applies the correction on an interval. That is +a namespace correction rather than a trust bootstrap, and it is safe to run from +inside the manager because the operator reads converted kinds only after its own +caches have synced against a webhook that is already up. + +**Reusing existing material matters as much as creating it.** A process that issued a fresh CA on every start would invalidate the bundle its CRDs already carry, so every restart would open a window in which the API server rejects the webhook it was just told to trust. The bootstrap therefore adopts what is already stored whenever it is valid for the service's DNS name and not near expiry. -**The alternative was to ship the CRDs with `strategy: None` and have the operator -raise it to `Webhook` once it is serving.** That removes the cycle and replaces it -with something worse. Under `None` the API server answers a hub-version read of a -stored spoke object by relabeling the apiVersion and pruning every field the new -schema does not know, so a reader sees an object with fields silently missing — -and a controller that writes during that window persists the pruned form. A -startup failure that is loud and self-correcting is a better trade than a -data-losing window that is neither. +**Shipping the CRDs with `strategy: None` and raising them to `Webhook` once the +process is serving is not an option.** Under `None` the API server answers a +hub-version read of a stored spoke object by relabeling the apiVersion and pruning +every field the new schema does not know, so a reader sees an object with fields +silently missing, and a controller that writes during that window persists the +pruned form. A startup failure that is loud and self-correcting is a better trade +than a data-losing window that is neither. ### 3.8 Which version is stored, and who decides @@ -481,44 +521,47 @@ webhook is inert and not deployed. The chart's own custom resources are authored at `v1alpha2` for the same reason: a chart that wrote `v1alpha1` would be the one client forcing conversion on a cluster where nothing serves it. -**An upgrade of an existing cluster does not take that value.** The upgrade tool -applies these same CRDs with storage held at `v1alpha1`, and this is a positive -act rather than an omission: a server-side apply overwrites the live storage -version, so shipping `v1alpha2` and applying it unchanged would move storage -before anything could convert. Every write would then need a webhook that the -operator carrying it has not finished rolling out, the running operator would stop -being able to update status, and the release would be irreversible — objects -written as `v1alpha2` cannot be read by the previous operator, which does not -serve conversion. - -**Storage moves at the end, with the objects.** Flipping the flag changes only -what new writes encode; objects untouched since remain in the old representation -and `.status.storedVersions` keeps listing `v1alpha1`, which is what stops -`v1alpha1` from being removed. The rewrite that follows lists every object and -writes it back unchanged, which is what moves it between representations. +**An upgrade applies the same manifests, and storage moves with them.** There is +one set of CRDs rather than a staged pair: `+kubebuilder:storageversion` sits on +the `v1alpha2` type, so the chart's copy, the one in `dist/install.yaml`, and the +one the upgrade tool embeds all declare `v1alpha2` as storage. What makes that +safe is ordering rather than a held flag. The conversion webhook is deployed, +awaited, and smoke-tested against a real object before the CRDs are applied +([`design-api-upgrade.md`](design-api-upgrade.md) §9.1), so there is something to +convert with by the time storage moves, and the webhook is a process of its own +(§3.7) rather than part of the operator roll-out the same upgrade is performing. + +**Objects change representation as they are next written.** The flag decides what +a write encodes and nothing else, so an object untouched since the apply stays in +the `v1alpha1` representation and is converted on every read, and +`.status.storedVersions` keeps listing `v1alpha1`, which is what stops `v1alpha1` +from being removed. The rewrite that lists every object and writes it back +unchanged is what drains the old representation deliberately. [`design-api-upgrade.md`](design-api-upgrade.md) §24 owns that sequence and this document does not repeat it. -| Path | Storage on arrival | Conversion invoked | Webhook | -|---------------|--------------------|-----------------------------|----------------------| -| Fresh install | `v1alpha2` | Never | Not deployed | -| Upgrade | `v1alpha1` | On every read | Deployed by the tool | -| After §24 | `v1alpha2` | Only for `v1alpha1` clients | Removed by §28 | - -**The consequence for this document is that the storage version is not a property -of a release.** Two clusters on the same operator version hold different storage -versions until the upgrade's rewrite has run, and both converge on `v1alpha2`. -Anything reasoning about what is in etcd has to ask the CRD rather than the -version number. - -**What this costs while an upgraded cluster sits at `v1alpha1` storage** is that -`v1alpha2` cannot carry information `v1alpha1` cannot express: a field only the -hub can state is converted down, dropped, and read back empty. The four kinds here -are renames and regroupings only, so nothing is lost, and -`hub_roundtrip_test.go` asserts it in the direction storage actually takes rather -than assuming it. A genuinely new field needs the annotation stash that -[`design-api-upgrade.md`](design-api-upgrade.md) §6.2 specifies, and nothing here -needs it yet. +| Path | Storage after the apply | Conversion invoked | Webhook | +|---------------|-------------------------|------------------------------------|----------------------| +| Fresh install | `v1alpha2` | Never | Not deployed | +| Upgrade | `v1alpha2` | For every object not yet rewritten | Deployed by the tool | +| After §24 | `v1alpha2` | Only for `v1alpha1` clients | Removed by §28 | + +**The consequence for this document is that the stored representation is not the +storage version.** A CRD that stores `v1alpha2` still holds objects encoded as +`v1alpha1` until each is written again, so anything reasoning about what is in +etcd asks `.status.storedVersions` rather than the storage flag, and the upgrade +is not finished when the apply is. + +**What this costs is bounded by the direction conversion runs in.** An object +encoded as `v1alpha1` is converted up on every read until something writes it, and +that write encodes `v1alpha2`, so information only the hub can state survives from +the first write onward. The lossy direction is the other one: a `v1alpha1` client +reading a hub object gets the spoke shape, and a field only the hub can state has +nowhere to go in it. The four kinds here are renames and regroupings only, so +nothing is lost, and `hub_roundtrip_test.go` asserts the hub → spoke → hub trip +that a `v1alpha1` client forces. A genuinely new field needs the annotation stash +that [`design-api-upgrade.md`](design-api-upgrade.md) §6.2 specifies, and nothing +here needs it yet. --- @@ -540,15 +583,15 @@ the hub, which is a unit that can be reviewed and reverted. Sweeping by class would leave every kind half-converted between sweeps, and a half-converted kind is one whose controller reads a field the conversion does not yet write. -| Order | Kind | Rows it carries | -|-------|---------------------|------------------------------------------------------------------------------| -| 1 | `ControlPlane` | §2.4 image regrouping, §2.5 phase | -| 2 | `StorageBackup` | §2.1 `clusterName` | -| 3 | `StorageClusterOps` | §2.1 `nodeRollingRestart`, §2.2 its status twin, §2.5 the action enum | -| 4 | `StorageNodeOps` | §2.1 `storageNodeRef` and `drain`, §2.4 the migrate group, §2.5 action enum | -| 5 | `StoragePool` | §2.1 `clusterName`, §2.3 `dhchap` and `encryption`, §2.4 both regroupings | -| 6 | `StorageNode` | §2.1 `overrides` and `socketIndex`, §2.3 the inverting toggle, §3.6's bridge | -| 7 | `StorageCluster` | §2.1 two rows, §2.3 six toggles, §2.4 the KMS regrouping, §2.5 the backend | +| Order | Kind | Rows it carries | +|-------|---------------------|-----------------------------------------------------------------------------| +| 1 | `ControlPlane` | §2.4 image regrouping, §2.5 phase | +| 2 | `StorageBackup` | §2.1 `clusterName` | +| 3 | `StorageClusterOps` | §2.1 `nodeRollingRestart`, §2.2 its status twin, §2.5 the action enum | +| 4 | `StorageNodeOps` | §2.1 `storageNodeRef` and `drain`, §2.4 the migrate group, §2.5 action enum | +| 5 | `StoragePool` | §2.1 `clusterName`, §2.3 `dhchap` and `encryption`, §2.4 both regroupings | +| 6 | `StorageNode` | §2.1 `overrides` and `socketIndex`, §2.3 the kubelet removal, §3.6's parent | +| 7 | `StorageCluster` | §2.1 two rows, §2.3 six toggles, §2.4 the KMS regrouping, §2.5 the backend | `ControlPlane` is first because it is the smallest kind that carries both a regrouping and an enum, so it proves the two hardest shapes on the least code. @@ -558,8 +601,9 @@ that need the annotation stash. The key renames of §2.6 share nothing with any of it and can proceed in parallel. The blocked rows of §2.7 stay blocked, except that §3.6 changes why for one of -them: `storageNodeSetRef` is carried into `v1alpha2` unconverted and filled in by -the controller, rather than waiting for the retirement. +them: `storageNodeSetRef` is read off the controller owner reference rather than +waiting for the retirement, which puts it behind the upgrade's reparenting step +instead of behind the `StorageNodeSet` kind. --- @@ -583,9 +627,11 @@ it: `controllers/cluster/storagecluster_controller.go`, **`spec.overrides` is the widest of the spec rows**, because the struct is read throughout node provisioning rather than at one call site. -**The chart carries user-facing spellings of its own.** `skipKubeletConfiguration` -is in `values.yaml`, and a chart value is not migrated by any webhook, so it needs -the same both-spellings treatment in the template or a documented break. +**The chart carries user-facing spellings of its own**, and no webhook migrates +one. A value the operator takes over is deleted from `values.yaml` with the +template that read it, which is what `skipKubeletConfiguration` did. A value that +survives in the chart and also names an API field is the case that needs both +spellings or a documented break. --- @@ -605,14 +651,13 @@ settled is `VolumeMigrationSettings.enabled`, which §12 removes outright, leavi `VolumeMigrationSettings` with only its remaining members and no toggle. Whether the struct survives that is not decided. -**Q3: Whether stored objects are migrated eagerly.** A conversion webhook makes -every object readable as `v1alpha2` without rewriting anything, so an object -applied as `v1alpha1` stays stored in whatever version it was written under until -something writes it again. That is correct and indefinite, and it means the -webhook cannot be retired by waiting. A storage-version migration — the -`StorageVersionMigration` kind, or a job that reads and rewrites every object — -would drain the old version deliberately. Which one, and whether the operator owns -it or the upgrade procedure does, is not decided. +**Q3: Whether stored objects are migrated eagerly.** Settled: the upgrade +procedure owns it. An object applied as `v1alpha1` stays in that representation +until something writes it again, so the webhook cannot be retired by waiting, and +the rewrite that drains the old representation is a stage of the upgrade tool +rather than a controller in the operator or the `StorageVersionMigration` kind +([`design-api-upgrade.md`](design-api-upgrade.md) §24). §3.8 states what follows +from it. **Q4: What the conversion does when a `v1alpha1` object is invalid.** §3.3 passes an unrecognized enum value through rather than failing, on the grounds that a diff --git a/operator/docs/designs/crd-redesign/design-storagebackup.md b/operator/docs/designs/crd-redesign/design-storagebackup.md index 1f95b3c7d..96406c949 100644 --- a/operator/docs/designs/crd-redesign/design-storagebackup.md +++ b/operator/docs/designs/crd-redesign/design-storagebackup.md @@ -1,8 +1,8 @@ # Design Document: The Data-Protection Chain -**Status:** Draft +**Status:** Implemented, with the exceptions §14 records **Author:** Christoph Engelbert (noctarius) -**Date:** 2026-08-30 +**Date:** 2026-08-30 (last updated 2026-09-17) **Test Plan:** [`tests/test-plan-storagebackup.md`](../../tests/test-plan-storagebackup.md) This document specifies the target model for the whole data-protection layer. @@ -171,10 +171,16 @@ ClaimSelector *metav1.LabelSelector `json:"claimSelector,omitempty"` ``` **An absent selector selects nothing, and that is the whole argument for the -default.** The alternative reading, that an empty selector matches everything, is -what `metav1.LabelSelector` means in most Kubernetes APIs and is wrong here: the -cost of backing up too much is silent and recurring, and the cost of backing up -too little is an error somebody sees. +default.** A policy that silently covered every claim in the namespace would back +up more than its author intended, and the failure would be a bill rather than an +error: the cost of backing up too much is silent and recurring, and the cost of +backing up too little is an error somebody sees. + +**An explicitly empty selector still means every claim in the namespace.** That is +what two empty braces mean in every other Kubernetes API and what somebody who +wrote them asked for, and omitting a field and writing it empty are different +statements. Only the first has a silent cost, so only the first is the one the +default guards. **The policy attaches and detaches as claims come and go.** A claim that starts matching is attached, a claim that stops matching is detached, and @@ -184,9 +190,25 @@ retention: a policy governs what is taken, not what is kept. ### 4.2 Schedule and retention -`spec.schedule` is a cron expression, `spec.maxVersions` is how many backups to -keep, and `spec.maxAge` is how long to keep them. All three are passed to the -control plane, which does the scheduling and the pruning (§9). +`spec.schedule` is an interval, `spec.maxVersions` is how many backups to keep, +and `spec.maxAge` is how long to keep them. All three are passed to the control +plane, which does the scheduling and the pruning (§9). + +**The schedule is the control plane's own format rather than cron**, because the +control plane is what reads it. It is a space-separated list of +`,` pairs, where the interval is a number and one of `m`, +`h`, `d`, or `w`, and the schema carries that as a pattern so a value the control +plane would refuse is refused at admission instead. Accepting cron here would mean +the operator translating one scheduling language into another and being the place +a mistranslation lives. + +**All three are immutable, and that is a property of the control plane rather than +a choice.** §10 lists a `PUT` that applies a changed schedule, and the v2 API +offers no such endpoint: it creates, deletes, attaches, and detaches a policy and +nothing else. A mutable field the operator cannot reconcile would leave the +declaration and the backups actually being taken permanently disagreeing, with the +object still reporting `Active`, so all three are fixed at creation until the +endpoint exists. Changing one means replacing the policy. **The operator does not run the schedule.** It reconciles the policy into the control plane and reports what the control plane did. That is deliberate: a @@ -197,17 +219,17 @@ schedule nobody could rely on. ## 5. StorageBackup -Declared in `operator/api/v1alpha1/storagebackup_types.go`, short name `sb`, and +Declared in `operator/api/v1alpha2/storagebackup_types.go`, short name `sb`, and reconciled by `StorageBackupReconciler` in `operator/internal/controllers/backup/storagebackup_controller.go`. The type is Appendix B. ### 5.1 A backup object is discovered, not declared -**Every `StorageBackup` is created by the operator from what the store holds.** The -cluster names an S3 location and its credentials, the operator walks it, and one object -appears per backup it finds. Nothing about a backup is a request, so the spec is -identity and nothing else: +**Every `StorageBackup` is created by the operator from what the control plane +reports the store holds.** The operator mirrors the cluster's backup stream and one +object appears per backup on it. Nothing about a backup is a request, so the spec +is identity and nothing else: ```go // ClusterRef names the StorageCluster whose store this backup was found in. With @@ -237,11 +259,29 @@ somebody's bucket, governed by that bucket's lifecycle policy and by the retenti control plane applies (§9). An object that could be deleted would invite the reading that deleting it frees the storage, and it does not. -**What a policy schedules and what the operator discovers meet in the same objects.** +**The control plane is the inventory, and the operator does not open the bucket.** +It has S3 credentials in `StorageCluster.spec.backup` and never uses them to list: +the control plane already reports every copy it wrote, so a second reader of the +same bucket would be a second opinion about what exists, reachable only where the +operator's own network can see the endpoint and disagreeing with the control plane +whenever a write is in flight. An absence is therefore decided by whether the +cluster's stream has synced, and by nothing about any other object, which is the +same shape [`design-storagedevice.md`](design-storagedevice.md) gives the device +mirror. + +**What a policy schedules and what the operator mirrors meet in the same objects.** `StorageBackupPolicy` tells the control plane which claims to back up and how often -(§4), the control plane writes the copies into the store, and the walk finds them. The -policy is the only thing that decides a backup is taken, so a policy and a discovery run -are two halves of one loop rather than two ways of creating an object. +(§4), the control plane writes the copies into the store and reports them, and the +mirror turns each into an object. The policy is the only thing that decides a +backup is taken, so a policy and the mirror are two halves of one loop rather than +two ways of creating an object. + +**No `StorageBackup` carries an owner reference to the policy that caused it.** The +control plane reports nothing about which policy took a copy, so the edge cannot be +built from what is on the wire, and it would be the wrong lifetime in any case: the +object goes when the store stops reporting the copy, not when a policy is deleted. +What a policy governs is what is taken rather than what is kept, which is the same +reading §9 gives retention. ### 5.2 Status, in three groups @@ -407,13 +447,23 @@ there is no object for it to resolve to, the API server validates its syntax, an §4.1's reading that an absent selector selects nothing is reported by `SelectorEmpty` rather than refused. -**`failurePolicy: Fail`**, for the reason +**Two of the three are `failurePolicy: Fail`**, for the reason [`design-clusterdeploymentconfig.md`](design-clusterdeploymentconfig.md) §5 gives: the webhook server runs inside the operator pod, so its availability tracks the operator's, and while the operator is down nothing reconciles a backup anyway. A window in which unresolvable immutable references are admitted is a window in which objects that can only be deleted are created. +**`StorageBackupValidator` is `Ignore`, because it is the one that guards +`DELETE`.** A webhook failing closed on a delete blocks the namespace controller +as well as a user, so an operator that is down would leave every namespace holding +a `StorageBackup` stuck in `Terminating` until somebody edited the webhook +configuration by hand. That is a worse failure than the one failing open admits, +which is a backup record deletable while the operator is down: the record is an +observation, so the mirror writes it back on the next sync and the copy in the +bucket was never at risk. The other two guard `CREATE` alone, where failing closed +costs a rejected write and nothing more. + The webhooks need `get` on `storageclusters`, `storagebackups`, `storagepools`, and `persistentvolumeclaims`, all of which the manager already reads to reconcile these kinds. @@ -681,9 +731,9 @@ of a round trip catches it. | No exclusion between two restores of one backup | `StorageBackup.status.activeOpsRef` (§6) | New, and the same lock every other entity with an `Ops` companion carries. Two restores of one backup now queue rather than run together (§14, Q5) | | No `shortName` on `BackupPolicy` or `StorageBackup` | `sbp` and `sb` | Additive. `br` and `bi` are retired with their kinds | | `spec.backup.snapshotBackups`, `withCompression`, `secondaryTarget`, `localTesting` | Removed (`design-storagecluster.md` Appendix A) | The store is a location, so how a copy is taken stays with the control plane | -| No owner reference from a policy to its backups | Established (§5.1) | Deleting a policy deletes the backup objects it created, not the backups themselves | +| No owner reference from a policy to its backups | Still none (§5.1) | The control plane reports no policy attribution, so the edge cannot be built. The object's lifetime is the store's report rather than the policy's | | Restore's claim owned by the operation | Unowned, and created only at `Binding` (§8) | Deleting the audit record no longer deletes the recovered volume, and a failed restore leaves no claim | -| Polling every backend read | A `?watch=true` subscription (§10) | Depends on `design-sse-push-notifications.md`, on the `sse` branch, as every other design in the group does | +| Polling every backend read | A `?watch=true` subscription (§10) | The mirror of §5.1 reads it, so a walk of the bucket is never performed | | No event, no metric | Fourteen reasons and eight metrics (§11) | New infrastructure | **The two kind removals are the breaking ones and they are not symmetrical with @@ -781,18 +831,35 @@ type StorageBackupPolicySpec struct { // +optional ClaimSelector *metav1.LabelSelector `json:"claimSelector,omitempty"` - // Schedule is a cron expression the control plane runs the policy on. + // Schedule is the interval the control plane runs the policy on, in its own + // format: a space-separated list of , pairs, where an + // interval is a number and one of m, h, d, or w. + // + // Immutable, and that is a property of the control plane rather than a + // choice. §10 lists a PUT that applies a changed schedule, and the v2 API + // offers no such endpoint: it creates, deletes, attaches, and detaches a + // policy and nothing else. A mutable field the operator cannot reconcile + // would leave the declaration and the backups actually being taken + // permanently disagreeing, with the object still reporting Active, so the + // schedule is fixed at creation until the endpoint exists. Changing one + // means replacing the policy. + // +kubebuilder:validation:Pattern=`^(\d+[mhdw],\d+)( +\d+[mhdw],\d+)*$` + // +k8s:immutable // +optional Schedule string `json:"schedule,omitempty"` // MaxVersions is how many backups of one claim to keep. Zero means no limit - // by count. + // by count. Immutable, for the reason Schedule is. // +kubebuilder:validation:Minimum=0 + // +k8s:immutable // +optional MaxVersions *int32 `json:"maxVersions,omitempty"` - // MaxAge is how long to keep a backup ("720h", "30d"). Empty means no limit - // by age. Retention is enforced by the control plane, not here. + // MaxAge is how long to keep a backup ("30d", "720h"). Empty means no limit + // by age. Retention is enforced by the control plane, not here. Immutable, + // for the reason Schedule is. + // +kubebuilder:validation:Pattern=`^[1-9]\d*[mhdw]$` + // +k8s:immutable // +optional MaxAge string `json:"maxAge,omitempty"` } diff --git a/operator/docs/designs/crd-redesign/design-storagecluster.md b/operator/docs/designs/crd-redesign/design-storagecluster.md index af797ea15..d395b9bfb 100644 --- a/operator/docs/designs/crd-redesign/design-storagecluster.md +++ b/operator/docs/designs/crd-redesign/design-storagecluster.md @@ -2,7 +2,7 @@ **Status:** Implemented, with the exceptions §12.1 records **Authors:** Christoph Engelbert (noctarius), Israel Geoffrey (`StorageClusterOps`) -**Date:** 2026-08-28 (last updated 2026-09-14) +**Date:** 2026-08-28 (last updated 2026-09-17) **Supersedes:** `design-storageclusterops.md`, removed in the same change **Test Plan:** [`tests/test-plan-storagecluster.md`](../../tests/test-plan-storagecluster.md) @@ -115,7 +115,7 @@ accepts the edit, which is not the same as the cluster tolerating it. ## 3. StorageCluster: API -Declared in `operator/api/v1alpha1/storagecluster_types.go`, short name `stc`. +Declared in `operator/api/v1alpha2/storagecluster_types.go`, short name `stc`. **The type is Appendix A**, whole and as it is to be written. What follows quotes the field an argument turns on and no more, so that one copy of each type exists and it is the one an implementation is written against. @@ -195,8 +195,12 @@ MaxSubsystemCount *int32 `json:"maxSubsystemCount"` // This is an explicit core count, not a percentage. Required: the core layout // it produces must match across the cluster, so it is stated rather than left // to a per-node heuristic. +// +// The floor is 4 rather than a hardware limit: a node must carry one core +// beyond this budget for the system, and the control plane's core layout +// assigns no NVMe-oF poller core at all for a 2-vCPU budget. // +kubebuilder:validation:Required -// +kubebuilder:validation:Minimum=6 +// +kubebuilder:validation:Minimum=4 VCPUCount *int32 `json:"vcpuCount"` ``` @@ -440,27 +444,28 @@ belongs in its status. ```go // Tasks are the control plane's own asynchronous jobs, as of the last stream -// frame: running and pending only, newest first, and capped. It is a window on -// the backend rather than a record: a task that finishes leaves the list, and -// what remains of it is the event that says it did (§10.1). +// frame: running and pending only, capped, and in the order the control plane +// reports them. It is a window on the backend rather than a record: a task that +// finishes leaves the list, and what remains of it is the event that says it +// did (§10.1). // +kubebuilder:validation:MaxItems=20 // +optional Tasks []ClusterTask `json:"tasks,omitempty"` ``` **Twenty is a window onto a cluster's work.** A busy cluster runs more than twenty -tasks, and the field shows what fits in a status somebody reads: ordered newest first, -so the twenty most recent running or pending tasks are visible and the rest stay in the -control plane. The -cap is what keeps an object bounded whose subject is not, which is the constraint any -status list has to answer to -([`design-crd-model.md`](design-crd-model.md) §3.1). - -**Newest first is the one part of this section the control plane cannot -support**, and §12.1 records what was done instead: its `TaskDTO` carries no -creation date, so the shipped window keeps the control plane's own order and -`ClusterTask` carries neither `createdAt` nor the `progress` the schema also -lacks. +tasks, and the field shows what fits in a status somebody reads. The cap is what +keeps an object bounded whose subject is not, which is the constraint any status +list has to answer to ([`design-crd-model.md`](design-crd-model.md) §3.1). + +**The window's order is the control plane's own**, because ordering it is not +something this operator can do. The `TaskDTO` carries no creation date and no +progress figure, so a newest-first window cannot be built from what is on the +wire, and `ClusterTask` declares neither `createdAt` nor `progress` rather than +declaring fields nothing can write +([`design-crd-model.md`](design-crd-model.md) §7.9). What it carries instead is +`retry`, which the schema does report and which is the one number separating a +task that is slow from one that is failing. **Only running and pending tasks appear.** A completed or canceled task is not current state, so it leaves the list, and the object stops describing it. That is what @@ -484,7 +489,7 @@ The smallest valid `StorageCluster` is the two required sizing fields and nothin else. Everything the control plane can default, it defaults. ```yaml -apiVersion: storage.simplyblock.io/v1alpha1 +apiVersion: storage.simplyblock.io/v1alpha2 kind: StorageCluster metadata: name: production @@ -498,7 +503,7 @@ A cluster that sets the layout, the tenancy thresholds, key storage, and both migration policies: ```yaml -apiVersion: storage.simplyblock.io/v1alpha1 +apiVersion: storage.simplyblock.io/v1alpha2 kind: StorageCluster metadata: name: production @@ -606,7 +611,7 @@ steady-state only. │ Kubernetes Control Plane │ │ ┌──────────────────────────────────────────────────────┐ │ │ │ StorageClusterReconciler │ │ -│ │ 1. Get the CR from the API server, not the cache │ │ +│ │ 1. Get the CR through the manager's cache │ │ │ │ 2. Deletion: backend DELETE, then finalizer │ │ │ │ 3. Ensure the finalizer │ │ │ │ 4. status.uuid != "" → syncStatus │ │ @@ -624,10 +629,14 @@ steady-state only. └──────────────────────────────────────────────────────────────┘ ``` -The CR is fetched with a direct read rather than from the informer cache. A -cached read can still return `status.uuid == ""` immediately after -`Status().Patch` has persisted a UUID, and acting on that stale value is a second -`POST` and a second backend cluster. +The CR is read through the manager's cache, so a reconcile can observe +`status.uuid == ""` immediately after `Status().Patch` has persisted a UUID. What +keeps that from becoming a second `POST` and a second backend cluster is the claim +rather than the read. §4.2's optimistic-lock patch returns 409 to a reconciler +holding a stale `resourceVersion`, and a pass that does get through re-enters +adoption, which finds the cluster by name and is idempotent. A direct read would +narrow the window without closing it, because the gap between reading and posting +is not the part that races. ### 4.2 Creation, and the lock that makes it single-shot @@ -677,7 +686,7 @@ response lost after the backend committed. ```go // StorageClusterPhase is where the operator has got to with this cluster. -// +kubebuilder:validation:Enum=Pending;Creating;Online;Degraded;Unavailable;Suspended +// +kubebuilder:validation:Enum=Pending;Creating;Provisioning;Activating;Online;Degraded;Unavailable;Suspended type StorageClusterPhase string // StorageClusterStep is one step of the creation path. There is one graph rather @@ -686,6 +695,31 @@ type StorageClusterPhase string type StorageClusterStep string ``` +**The phase is the operator's creation path until the cluster exists, and the +control plane's lifecycle afterward.** `Pending` and `Creating` are the operator's +own. Every other value is a reading of the string the control plane reports, and +the mapping is stated once rather than left to be inferred from a switch: + +| Control plane reports | Phase | +|------------------------------------------|----------------| +| nothing yet | `Pending` | +| `in_creation`, `in_expansion`, `unready` | `Provisioning` | +| `in_activation` | `Activating` | +| `active` | `Online` | +| `degraded`, `read_only` | `Degraded` | +| `suspended` | `Suspended` | +| anything else | `Unavailable` | + +`Provisioning` and `Activating` exist because a cluster being built is not a +cluster that is broken. Without them `unready` and `in_activation` both read as +`Unavailable`, which reports a fault for the ordinary course of a deployment and +leaves the phase unable to say that the control plane was asked for something and +is doing it. `Activating` is separate from `Provisioning` rather than its last +step, because an expansion ends in an activation and so does recovery from a +suspension, long after anything was being built. `Unavailable` keeps its meaning +as the residue: a status this operator has no reading for, rather than every +cluster that is not currently serving. + `Adopting` is reached from two states rather than one: the upgrade Secret diverts before any `POST`, and a `POST` that failed against an existing cluster diverts after (§4.3). Both converge on `Persisting`. Declaring both edges makes that a @@ -753,7 +787,7 @@ immediately. ## 5. StorageClusterOps: API -Declared in `operator/api/v1alpha1/storageclusterops_types.go`, short name +Declared in `operator/api/v1alpha2/storageclusterops_types.go`, short name `scops`. The type is Appendix B. ### 5.1 Spec @@ -778,7 +812,7 @@ the class [`design-crd-model.md`](design-crd-model.md) §7.5 leaves outside the `enableXyz`/`disableXyz` rule. ```yaml -apiVersion: storage.simplyblock.io/v1alpha1 +apiVersion: storage.simplyblock.io/v1alpha2 kind: StorageClusterOps metadata: name: roll-the-fleet @@ -813,7 +847,7 @@ wrong. An operation that is one call and one wait, in flight: ```yaml -apiVersion: storage.simplyblock.io/v1alpha1 +apiVersion: storage.simplyblock.io/v1alpha2 kind: StorageClusterOps metadata: name: activate-production @@ -837,7 +871,7 @@ the machine has got to within one node, and `rollingRestart` is which node that (§7). ```yaml -apiVersion: storage.simplyblock.io/v1alpha1 +apiVersion: storage.simplyblock.io/v1alpha2 kind: StorageClusterOps metadata: name: roll-the-fleet @@ -981,7 +1015,7 @@ to exercise, is caught by any test that builds a machine at all. Lock free? ← held by another ops → stay Pending, requeue after 10s │ free or ours ▼ - Acquire the lock ← optimistic-lock patch; 409 → requeue immediately + Acquire the lock ← optimistic-lock patch; 409 → requeue after 5s │ ▼ Pending → Running ← stamp startedAt @@ -1000,6 +1034,17 @@ its 10-second requeue after the lock frees. `clusterToOpsRequests` maps a `StorageCluster` event back to every operation targeting it, so a release wakes the queue immediately. +**The two unsuccessful outcomes of an acquisition are waited on differently, and +neither wait is zero.** A lock another operation visibly holds frees when that +operation's work finishes, which is what the 10-second backstop is sized for. A +409 is not that: the object moved between this pass's read and its write, so who +holds the lock now is one read away, and the pass backs off by 5 seconds instead. +Requeueing a refused patch immediately would be a spin — it burns a pass to +re-read a value that has not settled, and it does so fastest exactly when +contention is highest. The creation claim of §4.2 backs off by the same 5 seconds +and for the same reason, since it is the same kind of refusal on a different +object. + ### 6.2 The persisted position is the write-ahead record A side effect is preceded by a write, so that a process dying between the two @@ -1517,11 +1562,11 @@ question rather than answered quietly: see §13, Q3. Five things this document specifies are not in the shipped kinds, and each is waiting on something outside it rather than on a decision. -**`spec.storageNodes` is absent.** Its type is -[`design-storagenode.md`](design-storagenode.md) Appendix C, and that kind has -not moved: `StorageNode` is still `v1alpha1` and `StorageNodeSet` still owns the -workload. The field lands with that move rather than here, where it could only -be an empty block. +**`spec.storageNodes` carries the workload settings the node kinds gave up.** +Its type is [`design-storagenode.md`](design-storagenode.md) Appendix C. The +fields in it, `enableKubeletConfiguration` among them, are per-DaemonSet rather +than per-node, because a DaemonSet is one object for every node it schedules, and +they landed here with the workload's move onto the cluster. **All three `?watch=true` subscriptions of §9 are served by the control-plane informer.** The storage-node stream was already there; the cluster stream and @@ -1558,14 +1603,12 @@ same whichever way the state arrived, so each stream replaced one read function and no step. One thing the task stream settles rather than provides. The control plane's -`TaskDTO` carries no creation date and no progress figure, so §3.4's -"newest first" is not achievable and Appendix A's `progress` and `createdAt` -cannot be written. `ClusterTask` therefore carries neither — a field declared -and never written reports a definite-looking nothing, which is what -[`design-crd-model.md`](design-crd-model.md) §7.9 rules out and what §3.3 -removed four registered fields for. What it carries instead is `retry`, which -the schema does report and which is the one number separating a task that is -slow from one that is failing. The window's order is the control plane's own. +`TaskDTO` carries no creation date and no progress figure, so a newest-first +window is not orderable and neither `progress` nor `createdAt` can be written. +`ClusterTask` therefore declares neither: a field declared and never written +reports a definite-looking nothing, which is what +[`design-crd-model.md`](design-crd-model.md) §7.9 rules out and what §3.3 removed +four registered fields for. §3.4 states what the window carries instead. **`CancelTask` has nothing to call.** The v2 API lists tasks and reads one by ID, and offers no cancel (§9). The action is served — it takes the lock, waits @@ -1661,7 +1704,7 @@ against the same conventions it audits the shipped types against. // StorageClusterPhase is where the operator has got to with this cluster. The // first two values are the operator's own creation path; the rest are its reading // of the lifecycle status.status carries in the control plane's own spelling. -// +kubebuilder:validation:Enum=Pending;Creating;Online;Degraded;Unavailable;Suspended +// +kubebuilder:validation:Enum=Pending;Creating;Provisioning;Activating;Online;Degraded;Unavailable;Suspended type StorageClusterPhase string const ( @@ -1671,6 +1714,18 @@ const ( // Creating: the creation machine of §4.2 is running. StorageClusterPhaseCreating StorageClusterPhase = "Creating" + // Provisioning: the cluster exists in the control plane and is being built + // up, either by its first nodes joining or by an expansion adding more. It + // is not serving and there is nothing wrong with it, which is the + // distinction Unavailable cannot carry. + StorageClusterPhaseProvisioning StorageClusterPhase = "Provisioning" + + // Activating: the control plane is activating the cluster. It is a phase of + // its own rather than part of Provisioning because it is not only the last + // step of a deployment: an expansion ends in one, and so does recovering + // from a suspension, long after anything was being built. + StorageClusterPhaseActivating StorageClusterPhase = "Activating" + // Online: the control plane reports the cluster active and serving. StorageClusterPhaseOnline StorageClusterPhase = "Online" @@ -1816,8 +1871,12 @@ type StorageClusterSpec struct { // because it describes the host that node runs on. Required: the core layout // it produces must match across the cluster in steady state, so it is stated // rather than left to a per-node heuristic. + // + // The floor is 4 rather than a hardware limit: a node must carry one core + // beyond this budget for the system, and the control plane's core layout + // assigns no NVMe-oF poller core at all for a 2-vCPU budget. // +kubebuilder:validation:Required - // +kubebuilder:validation:Minimum=6 + // +kubebuilder:validation:Minimum=4 VCPUCount *int32 `json:"vcpuCount"` // MinHugePagesSize is the smallest huge-page allocation each storage node diff --git a/operator/docs/designs/crd-redesign/design-storagenode.md b/operator/docs/designs/crd-redesign/design-storagenode.md index 135d1c7b9..ff23fba61 100644 --- a/operator/docs/designs/crd-redesign/design-storagenode.md +++ b/operator/docs/designs/crd-redesign/design-storagenode.md @@ -1,15 +1,16 @@ # Design Document: The StorageNode and Its Operations -**Status:** Draft +**Status:** Implemented **Authors:** Christoph Engelbert (noctarius), Israel Geoffrey (`StorageNodeOps`) -**Date:** 2026-08-28 (last updated 2026-09-08) +**Date:** 2026-08-28 (last updated 2026-09-17) **Supersedes:** `design-storagenodeset-storagenode.md` and `design-node-removal-draining.md`, both removed in the same change **Test Plan:** [`tests/test-plan-storagenode.md`](../../tests/test-plan-storagenode.md) -This document specifies the target model. Both kinds and all four controllers -exist in a shape that predates the conventions of -[`design-crd-model.md`](design-crd-model.md), and §15 is the single record of what -the rework changes against them. +This document specifies the model both kinds now carry. Both moved to +`storage.simplyblock.io/v1alpha2` with a `v1alpha1` spoke and a conversion between +them, and the controllers were rewritten against this specification in +`operator/internal/controllers/node/`. §15 remains the record of what changed +against the registered API. --- @@ -156,7 +157,7 @@ was. ## 3. StorageNode: API -Declared in `operator/api/v1alpha1/storagenode_types.go`, short name `sn`. +Declared in `operator/api/v1alpha2/storagenode_types.go`, short name `sn`. **The type is Appendix A**, whole and as it is to be written. What follows quotes the field an argument turns on and no more, so that one copy of each type exists and it is the one an implementation is written against. @@ -203,12 +204,39 @@ it would describe something that stops being true once `nodesPerSocket` exceeds one and two slots share a socket, and naming it for the RPC-port ordering would describe how a slot is matched to a backend node rather than what it identifies. -**`spec.clusterRef` replaces `spec.storageNodeSetRef`, and the object name keeps -no worker in it.** A `StorageNode` is named `-` with a random short -identifier, because the socket and the worker are spec fields and the name has to -stay stable when a node relocates (§9). A name that encoded the worker would have -to be recreated by a migration, which would mean deleting a `StorageNode` whose -backend node is still running. +**`spec.clusterRef` replaces `spec.storageNodeSetRef`, and the object name is +derived rather than generated.** A `StorageNode` is named from its cluster, its +worker, and its slot through the shared formula +([`design-api-upgrade.md`](design-api-upgrade.md) §19.6), so the three facts that +identify a node at creation are the three the name is built from and a name is +arrived at the same way twice by anything that has to predict one. + +**The formula is bounded at a label's 63 bytes and not at the 253 an object name +may be**, because the name travels: the `StorageDevice` mirror writes it into +`storage.simplyblock.io/node` on every device the node carries +([`design-api-upgrade.md`](design-api-upgrade.md) §19.1). The arithmetic is not +academic. A regional cluster name and a worker a cloud named after its fully +qualified domain name are most of the budget between them, so a name that +overflows is what ordinary inputs produce rather than what a contrived one does. +Past that point the formula keeps what fits and spends the rest on a digest taken +over all three parts, so the worker stays in the name's text while it fits and in +its digest afterward, and two nodes differing only by worker never collide either +way. + +**The name is stable across a migration because Kubernetes never renames an +object**, not because the worker is absent from it. A migration re-points +`spec.workerNode` on the node that exists (§9), so what changes is the field and +not the object's identity. What that costs is stated rather than avoided: a node +built on one worker and migrated to another keeps a name describing where it was +built. `spec.workerNode` is where the current host is read, which is the field the +webhook restricts to one writer (§3.2), and the name is an identifier rather than +a report. + +**Re-expansion is idempotent without consulting the name.** The expansion lists +the cluster's existing nodes and matches them on the worker and slot their specs +record ([`design-clusterdeploymentconfig.md`](design-clusterdeploymentconfig.md) +§4.2), so a crash part-way through creates the rest and duplicates nothing even +where a name was derived under a wider limit than the one in force now. **Placement, mutable by the operator alone.** @@ -345,7 +373,7 @@ So the node carries its own copy of the two that describe the host it runs on: // equal across the fleet in steady state; a rolling hardware upgrade is what // makes two nodes differ, and only for as long as the roll takes. type StorageNodeSizing struct { - // +kubebuilder:validation:Minimum=6 + // +kubebuilder:validation:Minimum=4 // +kubebuilder:validation:Required VCPUCount *int32 `json:"vcpuCount"` // ... @@ -541,7 +569,7 @@ compares the node's sizing against the cluster it is joining and rejects a misma | Field | Rejected when | |------------------------------------------------------------|---------------------------------------------------------------------------| | `config.sizing.vcpuCount` | It differs from the cluster's stamp value | -| `config.sizing.minHugePagesSize` | It is set and differs from the cluster's stamp value | +| `config.sizing.minHugePagesSize` | It differs from the cluster's stamp value, an unset value included | | `config.deviceNames` | An entry is of a class other than the cluster's `spec.deviceClass` | | `config.pcieAllowList`, `config.pcieDenyList`, `pcieModel` | Any of them is set and the cluster's `spec.deviceClass` is `LogicalBlock` | @@ -557,10 +585,22 @@ ignoring them would leave somebody reading a filter that never ran. **The sizing rows are two and not others, because the control plane assumes them uniform.** §3.1 states why: a node whose core layout and huge pages disagree with its peers gets a layout the cluster cannot place erasure-coding chunks across evenly. -`vcpuCount` is `Required`, so a hand-written node states it and cannot inherit it by -omission, which is exactly the case where a typed value silently disagrees with the -fleet. There is no row for `maxSubsystemCount`: the node has no copy of it to -disagree with, which is the point of leaving it on the cluster (§3.1). +There is no row for `maxSubsystemCount`: the node has no copy of it to disagree +with, which is the point of leaving it on the cluster (§3.1). + +**An unset value is a divergence, because nothing inherits at render time.** The +node's own `config.sizing` is what §5.3 writes into the per-node ConfigMap, and it +writes `MAX_HUGE_PAGES_SIZE` from that field without consulting the cluster, so a +node omitting it runs on the computed minimum rather than on the floor the cluster +states. `minHugePagesSize` is optional on the type, which is what makes this worth +saying: optional means a cluster need not state a floor at all, not that a node may +decline the one its cluster states. Both rows therefore read the same way, and the +webhook compares the node's value to the cluster's whenever the cluster has one. + +**A cluster that states no floor checks nothing.** The comparison is skipped when +`StorageCluster.spec.minHugePagesSize` is empty, because there is no stamp value to +diverge from and every node is then on the computed minimum together, which is +uniform by construction. **The reference is the cluster's stamp value, not the siblings'.** During a rolling hardware upgrade the fleet is deliberately heterogeneous (§3.1), so sibling nodes @@ -586,7 +626,7 @@ A node as the operator writes it when expanding a deployment config, before anything has been provisioned: ```yaml -apiVersion: storage.simplyblock.io/v1alpha1 +apiVersion: storage.simplyblock.io/v1alpha2 kind: StorageNode metadata: name: production-7f3a9c @@ -1025,7 +1065,7 @@ name does not change when it rotates, so nothing else would notice. ## 6. StorageNodeOps: API -Declared in `operator/api/v1alpha1/storagenodeops_types.go`, short name `snops`. +Declared in `operator/api/v1alpha2/storagenodeops_types.go`, short name `snops`. The type is Appendix B. ### 6.1 Spec @@ -1059,7 +1099,7 @@ which is the class [`design-crd-model.md`](design-crd-model.md) §7.5 leaves outside the `enableXyz` and `disableXyz` rule. ```yaml -apiVersion: storage.simplyblock.io/v1alpha1 +apiVersion: storage.simplyblock.io/v1alpha2 kind: StorageNodeOps metadata: name: move-node-off-worker-3 @@ -1110,7 +1150,7 @@ wire value. A drain part-way through moving a node's volumes: ```yaml -apiVersion: storage.simplyblock.io/v1alpha1 +apiVersion: storage.simplyblock.io/v1alpha2 kind: StorageNodeOps metadata: name: production-7f3a9c-remove @@ -1296,7 +1336,16 @@ cluster, and one whose cluster is mid-rebalance or not active will either be rejected by the control plane or succeed into an inconsistent layout. The operation holds rather than fails, emits `ClusterNotReady`, and resumes when the cluster does. It applies to every action that changes the node's state, which is -all of them except the four single-step reads of a status. +all of them except the four single-step reads of a status and `Remove`. + +**`Remove` is exempt, because a removal is how an unready cluster becomes +ready.** Holding it until the cluster is active closes a loop with no way out: +the node cannot be removed until the cluster is active, and the cluster cannot +become active while the node it is stuck on is still in it. That is what a node +whose add never finished does to the cluster it was being added to. The +exemption follows from the gate's own purpose rather than working against it, +since the removal is what lets the control plane reach a state it will accept +further work in. Watching only `StorageNodeOps` would leave a queued operation waiting up to its requeue interval after the lock frees. `nodeToOpsRequests` maps a `StorageNode` @@ -1853,8 +1902,8 @@ delta, so that no other section has to carry it. | Everything under `spec.overrides` mutable | Most of `spec.config` immutable (§3.2) | Tightening. A user editing a device filter on a running node is now rejected | | `deviceNames`, NVMe namespace names | A PCI address or a device path (§3.1) | Widening. Every value the registered field took is still taken, and a logical block device becomes expressible | | `failureDomain`, an integer index | A label such as `rack-b` (§3.1) | Spec type change on both the spec and the status field. A stored index is not a valid label, so every node that declares a domain is rewritten, and the value stops being a number whose meaning lived outside the API | -| Four dead per-node fields in that struct | Moved to the cluster (§5.1) | Spec removal. None of them reached a consumer, so nothing loses behavior | -| `skipKubeletConfiguration` | `enableKubeletConfiguration`, inverted (§5.1) | Spec rename that also inverts, which is the one mechanical rename that is wrong | +| Four dead per-node fields in that struct | Moved to the cluster (§5.1) | Spec removal. None of them reached a consumer, so nothing loses behavior. `skipKubeletConfiguration` is one of them, and the row below is what its successor is called | +| `skipKubeletConfiguration` | Removed here, re-landed on the cluster (§5.1) | Spec removal. The successor is `StorageCluster.spec.storageNodes.enableKubeletConfiguration`, which is a different kind, so the registered field stashes on the way up rather than converting | | No `status.phase` | `StorageNodePhase` (§4.2) | Additive | | No step field, provisioning improvising one | `status.step` (§4.2) | Status only. The optimistic-lock claim moves to the `Posting` transition | | `status.postedAt` as the duplicate-POST guard | Removed (§3.3) | Status removal. The persisted step is the record | @@ -1869,23 +1918,23 @@ delta, so that no other section has to carry it. ### 15.2 StorageNodeOps -| Registered | This design | Cost | -|-------------------------------------------------------|---------------------------------------------|----------------------------------------------------------------------------------------| -| `spec.storageNodeRef` | `spec.nodeRef` (§6.1) | Spec rename | -| `spec.action` as a plain `string` | `StorageNodeOpsAction` (§6.1) | Type only, the wire values change with the row below | -| Six lowercase action values | PascalCase (§6.3) | Spec rename of every value. `design-crd-model.md` §9.7 owns the deprecation window | -| `spec.targetWorkerNode`, `spec.newSsdPcie` at the top | `spec.migrate` (§6.1) | Spec regrouping | -| `spec.drain` | `spec.remove` (§6.1) | Spec rename, matching the action it parameterizes | -| Six actions | Seven, adding `HostMaintenance` (§10) | Additive, and it retires a controller (§15.3) | -| No abort | `spec.abort` and the `Aborted` phase (§6.2) | Additive. Cancellation today means deleting the object | -| `status.subPhase`, a union of two workflows | `status.step`, one graph per action (§6.3) | Status only. The old string reads into `step.state` with no deadline | -| `Migrating` meaning two different things | `MigratingVolumes` and `Relocating` (§6.3) | Status only, and it removes an enum value that is ambiguous by construction | -| `status.triggered` | Removed (§7.2) | Status removal. The persisted step is the record, and it covers a case the flag cannot | -| `status.volumesMigrated`, `status.volumesPending` | `status.drain` (§6.2) | Status regrouping. `volumesTotal` replaces a pending count that has to be kept in step | -| No `observedGeneration` | Present (§6.2) | Additive | -| No state machine behind any action | Seven declared graphs (§6.3) | The largest piece of work here. Every side effect moves into a step | -| No deadline on any step | `status.step.deadline` (§6.3) | Additive, and what makes a stalled operation detectable | -| A cluster gate only for `Remove` | For every state-changing action (§7.1) | Behavioral. A relocation during a rebalance is currently accepted | +| Registered | This design | Cost | +|-------------------------------------------------------|-----------------------------------------------------|--------------------------------------------------------------------------------------------------------------------------------------------------------| +| `spec.storageNodeRef` | `spec.nodeRef` (§6.1) | Spec rename | +| `spec.action` as a plain `string` | `StorageNodeOpsAction` (§6.1) | Type only, the wire values change with the row below | +| Six lowercase action values | PascalCase (§6.3) | Spec rename of every value. `design-crd-model.md` §9.7 owns the deprecation window | +| `spec.targetWorkerNode`, `spec.newSsdPcie` at the top | `spec.migrate` (§6.1) | Spec regrouping | +| `spec.drain` | `spec.remove` (§6.1) | Spec rename, matching the action it parameterizes | +| Six actions | Seven, adding `HostMaintenance` (§10) | Additive, and it retires a controller (§15.3) | +| No abort | `spec.abort` and the `Aborted` phase (§6.2) | Additive. Cancellation today means deleting the object | +| `status.subPhase`, a union of two workflows | `status.step`, one graph per action (§6.3) | Status only. The old string reads into `step.state` with no deadline | +| `Migrating` meaning two different things | `MigratingVolumes` and `Relocating` (§6.3) | Status only, and it removes an enum value that is ambiguous by construction | +| `status.triggered` | Removed (§7.2) | Status removal. The persisted step is the record, and it covers a case the flag cannot | +| `status.volumesMigrated`, `status.volumesPending` | `status.drain` (§6.2) | Status regrouping. `volumesTotal` replaces a pending count that has to be kept in step | +| No `observedGeneration` | Present (§6.2) | Additive | +| No state machine behind any action | Seven declared graphs (§6.3) | The largest piece of work here. Every side effect moves into a step | +| No deadline on any step | `status.step.deadline` (§6.3) | Additive, and what makes a stalled operation detectable | +| A cluster gate only for `Remove` | For every state-changing action but `Remove` (§7.1) | Behavioral. A relocation during a rebalance is currently accepted, and `Remove` keeps its exemption because it is how an unready cluster becomes ready | ### 15.3 Retiring StorageNodeSet @@ -1914,10 +1963,12 @@ fields §5.1 names. Every spec row above is breaking, because a renamed spec field is silently ignored on an object that still sets the old name. Every status row is not, because the -operator is the only writer. The `skipKubeletConfiguration` row is the one to read -twice: it inverts as well as renames, so a deprecation window that reads the old -field and writes the new one has to negate it, and a mechanical rename produces the -opposite behavior. +operator is the only writer. The four per-node fields that move to the cluster are +the rows to read twice. Their successor is on a different kind, so no conversion +reaches them, and each stashes under +`storage.simplyblock.io/v1alpha1-spec.overrides.` on the way up so that a +node converted back down carries what it carried +([`design-property-renames.md`](design-property-renames.md) §3.3). The rows above are audited by `.claude/skills/api-design/scripts/check-crds.py --kind StorageNode` and @@ -2083,8 +2134,11 @@ type JournalManagerSpec struct { // the node's configuration is generated (§3.1). type StorageNodeSizing struct { // VCPUCount is the number of vCPUs allocated to SPDK on this node, as an - // explicit core count rather than a percentage. - // +kubebuilder:validation:Minimum=6 + // explicit core count rather than a percentage. The floor matches the + // cluster's: a node must carry one core beyond this budget for the system, + // and the control plane's core layout assigns no NVMe-oF poller core at all + // for a 2-vCPU budget. + // +kubebuilder:validation:Minimum=4 // +kubebuilder:validation:Required VCPUCount *int32 `json:"vcpuCount"` diff --git a/operator/docs/designs/crd-redesign/design-storagepool.md b/operator/docs/designs/crd-redesign/design-storagepool.md index c8dc6ec4d..4959ba775 100644 --- a/operator/docs/designs/crd-redesign/design-storagepool.md +++ b/operator/docs/designs/crd-redesign/design-storagepool.md @@ -2,7 +2,7 @@ **Status:** Partially Implemented **Author:** Christoph Engelbert (noctarius) -**Date:** 2026-08-29 (last updated 2026-09-11) +**Date:** 2026-08-29 (last updated 2026-09-17) **Test Plan:** [`tests/test-plan-storagepool.md`](../../tests/test-plan-storagepool.md) Both kinds are built at `storage.simplyblock.io/v1alpha2`, with a `v1alpha1` @@ -149,8 +149,9 @@ by default. ClusterRef string `json:"clusterRef"` ``` -`spec.allowedNodes` restricts which storage nodes may host the pool's volumes. -Empty means every node in the cluster, which is the usual case. +`spec.allowedNodes` restricts which hosts may carry the pool's volumes, and it +names Kubernetes `Node` objects rather than `StorageNode`s, for the reason §4.3 +gives. Empty means every node in the cluster, which is the usual case. **The pool's own ceilings, under `spec.limits`.** These are what the pool as a whole may consume, enforced by the control plane against the pool. @@ -234,8 +235,9 @@ between doing the work and declining to redo it. `status.limits` is what the control plane reports the pool's ceilings actually are, which is not necessarily what `spec.limits` asked for. -`status.allowedNodes` is `spec.allowedNodes` resolved against the nodes that exist, -which is what the control plane is sent (§4.3). `status.activeOpsRef` is the +`status.allowedNodes` is `spec.allowedNodes` resolved against the `Node` objects +that exist, and it is what the host list sent to the control plane and the per-pool +node labels are both derived from (§4.3). `status.activeOpsRef` is the operation lock ([`design-crd-model.md`](design-crd-model.md) §3.2), and `status.observedGeneration` is required by that document's §7.9. @@ -381,14 +383,29 @@ returns without patching. A node drained and removed through `StorageNodeOps` leaves its name in every pool that listed it, and the list is desired state a person wrote. +**The names are Kubernetes `Node` names, not `StorageNode` names.** A pool +restricts placement to hosts, and the two facts the restriction is expressed with +both live on the `Node` object: its UID, from which the host NQN is derived, and +its labels, which is where the per-pool allowance is written. Naming the +`StorageNode` would mean resolving it to its worker before either could be read, +and the worker is what the user is choosing in the first place. + **The controller resolves the list on every pass and publishes the result as -`status.allowedNodes`.** A name that no longer resolves to a `StorageNode` in the -pool's namespace is dropped from the resolved set, reported once with -`AllowedNodeMissing` (§9.1), and what the control plane is sent is the resolved set -rather than the authored one. A pool whose every allowed node has been removed -resolves to an empty set, which is not the same as an absent list: absent means every -node, and empty after resolution means the pool can place nothing, so the phase holds -and the event says which names failed. +`status.allowedNodes`.** A name that resolves to no `Node` is dropped from the +resolved set, reported once with `AllowedNodeMissing` (§9.1), and what the control +plane is sent is the resolved set rather than the authored one. A pool whose every +allowed node has been removed resolves to an empty set, which is not the same as an +absent list: absent means every node, and empty after resolution means the pool can +place nothing, so the phase holds and the event says which names failed. + +**The resolved set reaches two places, and neither of them is a name.** The +control plane is sent one host NQN per node, derived from that node's UID by the +same formula the CSI node plugin uses, so no NQN is written by hand on either +side. Kubernetes is given one label per pool on each resolved node, keyed by the +pool's UUID and removed from every node that leaves the set, which is what the +class republishes as the DHCHAP node selector (§4.4). Both are derived on every +pass, so a node added back under the same name is re-derived rather than +repaired. **`spec.allowedNodes` itself is left exactly as authored, and the operator does not prune it.** Rewriting a user's spec to match the world makes the object stop @@ -400,10 +417,14 @@ harmful, which is the property the resolution provides and a prune would not. ### 4.4 A Default Pool Is Created With the Cluster -**Creating a `StorageCluster` creates one `StoragePool` in it.** A cluster with no -pool can hold no volumes, so the first pool is not a decision worth making a -prerequisite: the cluster's own creation path creates it, owns it by the same -controller reference every pool has (§6), and names it for the cluster it belongs to. +**A `StorageCluster` is given one `StoragePool`.** A cluster with no pool can hold +no volumes, so the first pool is not a decision worth making a prerequisite: the +cluster's own reconciler writes it, owns it by the same controller reference every +pool has (§6), and names it for the cluster it belongs to. It is written on the +cluster's steady-state pass rather than as a step of the creation machine, because +the pool's own controller needs the cluster's UUID and a pool written before there +is one would only wait. An annotation on the cluster records that the pool was +written, so nothing writes it a second time. **It is an ordinary pool in every other respect, deletion included.** Its limits are the cluster's defaults, it can be edited, and it deletes like any other pool: §6's two @@ -413,13 +434,25 @@ because a cluster that has outgrown one pool per tenant is not obliged to keep o What the default is for is the case where somebody wants storage from a cluster they just created and has not yet decided how to divide it. -**One `StorageClass` is created with it, so the cluster is provisionable on arrival.** -A pool nothing can consume is not a usable default, so the cluster's creation path -writes both: the pool, and one class assigned to it by the three labels of §5. The -class needs no ordering against the backend, because its `parameters` name the cluster -UUID and the pool *name* rather than the pool's UUID, and both are known before the -pool exists in the control plane. A claim made in the window before it does fails at -provision time and succeeds afterward. +**One `StorageClass` is written for it, once the pool exists in the control +plane.** A pool nothing can consume is not a usable default, so the pool's own +controller writes one class assigned to it by the three labels of §5, and it does +so on the first pass where `status.uuid` is set rather than beside the pool. The +ordering is the DHCHAP node selector's: the parameter republishes the per-pool +label key so the driver can turn it into the volume's node affinity, and that key +carries the pool's UUID, which does not exist until the control plane has created +the pool. A class whose `parameters` are immutable cannot gain the key later, so +it is written when the key can be stated. A claim made in the window before that +fails at provision time and succeeds afterward. + +**The class is written once, and a name already taken is reported rather than +overwritten.** A `StorageClass` is cluster-scoped, so an existing object under the +name may have nothing to do with this pool, and recording it as the pool's default +would leave the pool pointing at a class that provisions somewhere else and never +writing the one it needs. The name is adopted only when the occupant is +recognizably the class this operator would have written. Otherwise +`StorageClassNameTaken` (§9.1) says so and names the two remedies, which are +assigning a class by label or deleting the occupant. ```yaml kind: StorageClass @@ -819,6 +852,7 @@ has open, and on the `StoragePoolOps` for an operation. | The cluster is not ready, so creation is held | `Normal` | `ClusterNotReady` | `StoragePool` | | A `StorageClass` was assigned to this pool | `Normal` | `StorageClassAssigned` | `StoragePool` | | The default class was created with the default pool | `Normal` | `StorageClassCreated` | `StoragePool` | +| The default class's name is taken by another class | `Warning` | `StorageClassNameTaken` | `StoragePool` | | A class assigned to this pool sets both QoS spellings | `Warning` | `QoSParameterConflict` | `StoragePool` | | An entry in `spec.allowedNodes` resolves to no node | `Warning` | `AllowedNodeMissing` | `StoragePool` | | Deletion is held because a class is still assigned | `Warning` | `StorageClassStillAssigned` | `StoragePool` | @@ -848,6 +882,11 @@ leaves `spec.allowedNodes` as authored, so a removed node's name stays there and otherwise produce an event on every reconcile forever. The event is what tells somebody the name is inert, and repeating it would tell them nothing new. +**`StorageClassNameTaken` is the one an administrator has to act on.** The class +is written once (§4.4), so a name already held by a class this operator would not +have written leaves the default pool with none until somebody intervenes, and the +event names both remedies rather than only the collision. + **`QoSParameterConflict` is the proactive half of §5.1's conflict.** The driver emits the same reason on the claim it is provisioning, and this one fires when the class is first indexed, which is usually well before anybody's claim reaches it. @@ -1172,9 +1211,15 @@ type StoragePoolSpec struct { // +k8s:immutable ClusterRef string `json:"clusterRef"` - // AllowedNodes restricts which storage nodes may host this pool's volumes. - // Empty means every node in the cluster. Narrowing it stops new volumes - // landing on the removed nodes and leaves the existing ones where they are. + // AllowedNodes restricts which hosts may carry this pool's volumes, by + // Kubernetes Node name. Empty means every node in the cluster. Narrowing it + // stops new volumes landing on the removed nodes and leaves the existing + // ones where they are. + // + // The list is left exactly as authored: a name that no longer resolves is + // dropped from Status.AllowedNodes rather than pruned from here, so a node + // removed for maintenance and added back under the same name returns to the + // pools that named it without anybody re-authoring them. // +optional // +listType=set AllowedNodes []string `json:"allowedNodes,omitempty"` @@ -1242,7 +1287,11 @@ type StoragePoolStatus struct { // +optional Limits *PoolLimitsStatus `json:"limits,omitempty"` - // AllowedNodes is the resolved node list. + // AllowedNodes is Spec.AllowedNodes resolved against the Node objects that + // exist, which is what the control plane's host list and the per-pool node + // labels are derived from. An empty list here is not the same as an absent + // Spec.AllowedNodes: absent means every node, and empty after resolution + // means the pool can place nothing. // +optional // +listType=set AllowedNodes []string `json:"allowedNodes,omitempty"` From c3b879f6ebc9e641348c90771d051bf735bf42e7 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Thu, 17 Sep 2026 20:32:47 +0200 Subject: [PATCH 061/206] fix(cluster): the two lock waits are different waits, and the comments say what the code does MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A refused lock patch and a lock somebody visibly holds were waited on identically, because acquireLock returned a bool and a bool cannot tell them apart. They are not the same situation. A lock another operation holds frees when that operation's work finishes, which is what the ordinary retry is sized for; a 409 means the object moved between this pass's read and its write, so who holds it now is one read away. So the outcome is typed, and the two are waited on separately: the ordinary retry drops to the 10 seconds design-storagecluster.md §6.1 always specified, and a refused patch backs off by 5. The design promised an immediate requeue for the 409, and that is worse than either number. Requeueing a contended object with no delay burns a pass re-reading a value that has not settled, and does so fastest exactly when contention is highest. The creation claim already backed off by 5 seconds for the same reason; its literal is now named beside the others. Four comments described a mechanism their own code does not use. The StorageCluster reconciler's header claimed the object is fetched with a direct read to keep a stale empty status.uuid from producing a second POST. It is read through the manager's cache. The behavior is safe by a different mechanism, which is the Claiming optimistic-lock patch and the idempotent adoption a stale pass re-enters, so the header names that one. The Tasks status field promised newest-first ordering the control plane's TaskDTO cannot support, which the design's own §12.1 already recorded and its appendix already spelled correctly. AllowedNodeMissing said an entry resolves to no StorageNode, and the resolution is against Kubernetes Nodes. nodeNameFormula said a node is named "never for the worker" and that "the worker is in the name's digest rather than in its text". Running it says otherwise: the formula carries no AlwaysDigest, so a name that fits is the plain join, and production/worker-3/0 derives production-worker-3-0. The worker is in the text until the 63-byte limit forces a truncation. The comment also credited the name with idempotent re-expansion, which rests on the worker and slot in each node's spec instead. remove.go's header said the drain fans out VolumeMigration because PersistentVolumeOps does not exist yet, and promised a design section recording that. The kind exists, internal/volumemigration.Mover chooses between the two, and the drain takes the cluster-scoped one. ClusterDeploymentConfigStep's Enum marker was missing Activating where the CEL rule on status.step already had it. Nothing enforced the marker, which is why it went unnoticed: it sits on a type no field references, so only the CEL rule reaches the manifest. Tests first: requeue_test.go pins all three intervals and was red on two of them before the change. The third covers claimBackoff, which was already correct and passed on the red run. Regenerated: three comment changes reach the CRD descriptions. Co-Authored-By: Claude Opus 5 (1M context) --- ...torage.simplyblock.io_storageclusters.yaml | 10 ++- .../storage.simplyblock.io_storagepools.yaml | 16 ++-- operator/api/v1alpha1/hub_roundtrip_test.go | 17 ++-- .../v1alpha2/clusterdeploymentconfig_types.go | 2 +- operator/api/v1alpha2/storagecluster_types.go | 10 ++- operator/api/v1alpha2/storagepool_types.go | 16 ++-- ...torage.simplyblock.io_storageclusters.yaml | 10 ++- .../storage.simplyblock.io_storagepools.yaml | 16 ++-- .../controllers/cluster/requeue_test.go | 90 +++++++++++++++++++ .../cluster/storagecluster_controller.go | 18 ++-- .../cluster/storageclusterops_controller.go | 57 +++++++++--- .../controllers/deployment/expansion.go | 21 +++-- operator/internal/controllers/node/remove.go | 12 +-- operator/internal/controllers/pool/events.go | 6 +- ...torage.simplyblock.io_storageclusters.yaml | 10 ++- .../storage.simplyblock.io_storagepools.yaml | 16 ++-- 16 files changed, 238 insertions(+), 89 deletions(-) create mode 100644 operator/internal/controllers/cluster/requeue_test.go diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusters.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusters.yaml index e1f616c64..5abed659d 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusters.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusters.yaml @@ -1525,10 +1525,12 @@ spec: rule: '!has(self.state) || self.state in [''Claiming'',''CheckingControlPlane'',''ResolvingConfig'',''Creating'',''Adopting'',''Persisting'']' tasks: description: |- - Tasks are the control plane's running and pending jobs, newest first and - capped at twenty. Completed and canceled tasks are not here: they leave - the list and become events, so the length tracks concurrency rather than - history. + Tasks are the control plane's running and pending jobs, capped at twenty + and in the order the control plane reports them: its TaskDTO carries no + creation date, so newest-first is not orderable from what is on the wire + (design-storagecluster.md §12.1). Completed and canceled tasks are not + here: they leave the list and become events, so the length tracks + concurrency rather than history. items: description: |- ClusterTask is one asynchronous job the control plane is running, as of the diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagepools.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagepools.yaml index 498c28856..7f1b2d29e 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagepools.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagepools.yaml @@ -295,9 +295,10 @@ spec: properties: allowedNodes: description: |- - AllowedNodes restricts which storage nodes may host this pool's volumes. - Empty means every node in the cluster. Narrowing it stops new volumes - landing on the removed nodes and leaves the existing ones where they are. + AllowedNodes restricts which hosts may carry this pool's volumes, by + Kubernetes Node name. Empty means every node in the cluster. Narrowing it + stops new volumes landing on the removed nodes and leaves the existing + ones where they are. The list is left exactly as authored: a name that no longer resolves is dropped from Status.AllowedNodes rather than pruned from here, so a node @@ -480,10 +481,11 @@ spec: type: string allowedNodes: description: |- - AllowedNodes is Spec.AllowedNodes resolved against the StorageNodes that - exist, which is what the control plane is sent. An empty list here is not - the same as an absent Spec.AllowedNodes: absent means every node, and - empty after resolution means the pool can place nothing. + AllowedNodes is Spec.AllowedNodes resolved against the Node objects that + exist, which is what the control plane's host list and the per-pool node + labels are derived from. An empty list here is not the same as an absent + Spec.AllowedNodes: absent means every node, and empty after resolution + means the pool can place nothing. items: type: string type: array diff --git a/operator/api/v1alpha1/hub_roundtrip_test.go b/operator/api/v1alpha1/hub_roundtrip_test.go index 24571507e..8cc89335f 100644 --- a/operator/api/v1alpha1/hub_roundtrip_test.go +++ b/operator/api/v1alpha1/hub_roundtrip_test.go @@ -1,12 +1,13 @@ -// Round trips that start at the hub, which is the direction storage now takes. +// Round trips that start at the hub, which is the direction a v1alpha1 client +// forces. // -// v1alpha1 is the storage version while v1alpha2 is served beside it -// (design-property-renames.md §3.8), so a controller writing v1alpha2 has its -// object converted down to v1alpha1 to be stored and back up to v1alpha2 on the -// next read. That makes hub → spoke → hub the fidelity that matters in practice, -// and it is not the same property as the spoke → hub → spoke trip the per-kind -// tests already cover: a field only the hub can express survives one and not the -// other. +// v1alpha2 is the storage version and v1alpha1 is served beside it +// (design-property-renames.md §3.8), so an object a controller wrote is converted +// down to v1alpha1 for any client that asks for that version and back up to +// v1alpha2 on the next read of it. That makes hub → spoke → hub the fidelity that +// matters in practice, and it is not the same property as the spoke → hub → spoke +// trip the per-kind tests already cover: a field only the hub can express survives +// one and not the other. package v1alpha1 diff --git a/operator/api/v1alpha2/clusterdeploymentconfig_types.go b/operator/api/v1alpha2/clusterdeploymentconfig_types.go index 3591813fa..b93311190 100644 --- a/operator/api/v1alpha2/clusterdeploymentconfig_types.go +++ b/operator/api/v1alpha2/clusterdeploymentconfig_types.go @@ -58,7 +58,7 @@ const ( ) // ClusterDeploymentConfigStep is one step of the expansion path. -// +kubebuilder:validation:Enum=Validating;CreatingCluster;AwaitingCluster;CreatingNodes +// +kubebuilder:validation:Enum=Validating;CreatingCluster;AwaitingCluster;CreatingNodes;Activating type ClusterDeploymentConfigStep string const ( diff --git a/operator/api/v1alpha2/storagecluster_types.go b/operator/api/v1alpha2/storagecluster_types.go index f6141ac88..db9907fbb 100644 --- a/operator/api/v1alpha2/storagecluster_types.go +++ b/operator/api/v1alpha2/storagecluster_types.go @@ -833,10 +833,12 @@ type StorageClusterStatus struct { // +optional LastDataRealignmentAt *metav1.Time `json:"lastDataRealignmentAt,omitempty"` - // Tasks are the control plane's running and pending jobs, newest first and - // capped at twenty. Completed and canceled tasks are not here: they leave - // the list and become events, so the length tracks concurrency rather than - // history. + // Tasks are the control plane's running and pending jobs, capped at twenty + // and in the order the control plane reports them: its TaskDTO carries no + // creation date, so newest-first is not orderable from what is on the wire + // (design-storagecluster.md §12.1). Completed and canceled tasks are not + // here: they leave the list and become events, so the length tracks + // concurrency rather than history. // +kubebuilder:validation:MaxItems=20 // +optional Tasks []ClusterTask `json:"tasks,omitempty"` diff --git a/operator/api/v1alpha2/storagepool_types.go b/operator/api/v1alpha2/storagepool_types.go index 59589749f..cac806cd9 100644 --- a/operator/api/v1alpha2/storagepool_types.go +++ b/operator/api/v1alpha2/storagepool_types.go @@ -181,9 +181,10 @@ type StoragePoolSpec struct { // +k8s:immutable ClusterRef string `json:"clusterRef"` - // AllowedNodes restricts which storage nodes may host this pool's volumes. - // Empty means every node in the cluster. Narrowing it stops new volumes - // landing on the removed nodes and leaves the existing ones where they are. + // AllowedNodes restricts which hosts may carry this pool's volumes, by + // Kubernetes Node name. Empty means every node in the cluster. Narrowing it + // stops new volumes landing on the removed nodes and leaves the existing + // ones where they are. // // The list is left exactly as authored: a name that no longer resolves is // dropped from Status.AllowedNodes rather than pruned from here, so a node @@ -261,10 +262,11 @@ type StoragePoolStatus struct { // +optional Limits *PoolLimitsStatus `json:"limits,omitempty"` - // AllowedNodes is Spec.AllowedNodes resolved against the StorageNodes that - // exist, which is what the control plane is sent. An empty list here is not - // the same as an absent Spec.AllowedNodes: absent means every node, and - // empty after resolution means the pool can place nothing. + // AllowedNodes is Spec.AllowedNodes resolved against the Node objects that + // exist, which is what the control plane's host list and the per-pool node + // labels are derived from. An empty list here is not the same as an absent + // Spec.AllowedNodes: absent means every node, and empty after resolution + // means the pool can place nothing. // +optional // +listType=set AllowedNodes []string `json:"allowedNodes,omitempty"` diff --git a/operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml b/operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml index e1f616c64..5abed659d 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml @@ -1525,10 +1525,12 @@ spec: rule: '!has(self.state) || self.state in [''Claiming'',''CheckingControlPlane'',''ResolvingConfig'',''Creating'',''Adopting'',''Persisting'']' tasks: description: |- - Tasks are the control plane's running and pending jobs, newest first and - capped at twenty. Completed and canceled tasks are not here: they leave - the list and become events, so the length tracks concurrency rather than - history. + Tasks are the control plane's running and pending jobs, capped at twenty + and in the order the control plane reports them: its TaskDTO carries no + creation date, so newest-first is not orderable from what is on the wire + (design-storagecluster.md §12.1). Completed and canceled tasks are not + here: they leave the list and become events, so the length tracks + concurrency rather than history. items: description: |- ClusterTask is one asynchronous job the control plane is running, as of the diff --git a/operator/config/crd/bases/storage.simplyblock.io_storagepools.yaml b/operator/config/crd/bases/storage.simplyblock.io_storagepools.yaml index 498c28856..7f1b2d29e 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_storagepools.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_storagepools.yaml @@ -295,9 +295,10 @@ spec: properties: allowedNodes: description: |- - AllowedNodes restricts which storage nodes may host this pool's volumes. - Empty means every node in the cluster. Narrowing it stops new volumes - landing on the removed nodes and leaves the existing ones where they are. + AllowedNodes restricts which hosts may carry this pool's volumes, by + Kubernetes Node name. Empty means every node in the cluster. Narrowing it + stops new volumes landing on the removed nodes and leaves the existing + ones where they are. The list is left exactly as authored: a name that no longer resolves is dropped from Status.AllowedNodes rather than pruned from here, so a node @@ -480,10 +481,11 @@ spec: type: string allowedNodes: description: |- - AllowedNodes is Spec.AllowedNodes resolved against the StorageNodes that - exist, which is what the control plane is sent. An empty list here is not - the same as an absent Spec.AllowedNodes: absent means every node, and - empty after resolution means the pool can place nothing. + AllowedNodes is Spec.AllowedNodes resolved against the Node objects that + exist, which is what the control plane's host list and the per-pool node + labels are derived from. An empty list here is not the same as an absent + Spec.AllowedNodes: absent means every node, and empty after resolution + means the pool can place nothing. items: type: string type: array diff --git a/operator/internal/controllers/cluster/requeue_test.go b/operator/internal/controllers/cluster/requeue_test.go new file mode 100644 index 000000000..fa0061337 --- /dev/null +++ b/operator/internal/controllers/cluster/requeue_test.go @@ -0,0 +1,90 @@ +// The requeue intervals these two reconcilers back off by, asserted as values +// rather than as the bare fact that some requeue happened. +// +// Every one of them is a backstop: a queued operation is normally woken by its +// cluster, and a claim that lost its race is normally woken by the write that +// beat it. That is exactly why they are worth pinning. A backstop nothing +// reaches in a passing test is a constant free to drift, and +// design-storagecluster.md §6.1 states these numbers, so the document and the +// code are held to the same ones here. + +package cluster + +import ( + "context" + "testing" + "time" + + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// A queued operation waits the ordinary retry, which §6.1 calls the backstop for +// a lock-release event that was missed. +func TestAQueuedOperationRequeuesAtTheOrdinaryRetry(t *testing.T) { + held := newTestCluster(func(c *simplyblockv1alpha2.StorageCluster) { + c.Status.ActiveOpsRef = otherOpsName + }) + r := newOpsReconciler(t, &fakeControlPlane{}, &recorder{}, + held, newTestOps(simplyblockv1alpha2.StorageClusterOpsActionShutdown)) + + result, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: client.ObjectKey{Namespace: testNamespace, Name: testOpsName}, + }) + if err != nil { + t.Fatalf("reconcile: %v", err) + } + if result.RequeueAfter != opsRetry { + t.Errorf("RequeueAfter = %v, want the ordinary retry %v", result.RequeueAfter, opsRetry) + } + if opsRetry != 10*time.Second { + t.Errorf("opsRetry = %v, want 10s", opsRetry) + } +} + +// A 409 on the lock patch is a different situation from a lock somebody visibly +// holds: the object moved under this pass and the answer is one read away, so it +// backs off by the shorter contention interval rather than by the full retry. It +// is not immediate, because an immediate requeue against a contended object is a +// spin. +func TestTheContentionIntervalIsShorterThanTheRetryAndNotZero(t *testing.T) { + if opsContended == 0 { + t.Fatal("opsContended is zero, which makes a contended requeue a spin") + } + if opsContended >= opsRetry { + t.Errorf("opsContended = %v, which is not shorter than opsRetry = %v", + opsContended, opsRetry) + } + if opsContended != 5*time.Second { + t.Errorf("opsContended = %v, want 5s", opsContended) + } +} + +// The creation claim backs off by the same interval and for the same reason: its +// optimistic-lock patch 409'd, so another reconciler is mid-claim and this pass +// re-reads rather than retrying instantly. +func TestALostCreationClaimBacksOffByTheContentionInterval(t *testing.T) { + r := newClusterReconciler(t, &fakeControlPlane{}, &recorder{}, newUncreatedCluster()) + ctx := context.Background() + + // A reconciler holding a copy nobody has written since, against an object + // somebody else has moved: the patch carries a resourceVersion the API + // server will refuse. + var stale simplyblockv1alpha2.StorageCluster + if err := r.Get(ctx, client.ObjectKey{ + Namespace: testNamespace, Name: testClusterName, + }, &stale); err != nil { + t.Fatalf("read the cluster: %v", err) + } + stale.ResourceVersion = "1" + + if got := r.claim(ctx, &stale); got.RequeueAfter != claimBackoff { + t.Errorf("RequeueAfter = %v, want the contention interval %v", + got.RequeueAfter, claimBackoff) + } + if claimBackoff != 5*time.Second { + t.Errorf("claimBackoff = %v, want 5s", claimBackoff) + } +} diff --git a/operator/internal/controllers/cluster/storagecluster_controller.go b/operator/internal/controllers/cluster/storagecluster_controller.go index 9b1876b56..0120f4b0c 100644 --- a/operator/internal/controllers/cluster/storagecluster_controller.go +++ b/operator/internal/controllers/cluster/storagecluster_controller.go @@ -1,13 +1,11 @@ // The StorageCluster reconciler: four paths, which are creation, adoption, // steady-state synchronization, and deletion. // -// The object is fetched with a direct read rather than from the informer -// cache. A cached read can still return an empty status.uuid immediately after -// a status patch has persisted one, and acting on that stale value is a second -// POST and a second backend cluster. -// -// Creating a backend cluster is not idempotent, so the claim is made in -// Kubernetes before the control plane is touched. The mutex is the +// Creating a backend cluster is not idempotent, and the object is read through +// the manager's cache, so a reconcile can see an empty status.uuid immediately +// after a status patch has persisted one. What makes that safe is the claim +// rather than the read: the claim is made in Kubernetes before the control +// plane is touched, and a stale pass re-enters adoption, which is idempotent. The mutex is the // optimistic-lock patch rather than the value it writes: the patch succeeds for // exactly one reconciler at a given resourceVersion and returns 409 to the // rest, so persisting the transition into Claiming is what makes creation @@ -367,12 +365,16 @@ func (r *StorageClusterReconciler) claim( logf.FromContext(ctx).Info( "another reconciler holds the creation claim; backing off", "cluster", cluster.Name) - return ctrl.Result{RequeueAfter: 5 * time.Second} + return ctrl.Result{RequeueAfter: claimBackoff} } r.observePhase(cluster) return ctrl.Result{RequeueAfter: time.Second} } +// claimBackoff is how long a reconciler whose creation claim was refused waits +// before looking again. +const claimBackoff = 5 * time.Second + // enterCreationStep records the next step with its deadline and moves the // machine into it. The record precedes the step's own work on the next pass, // which is the write-ahead this path needs. diff --git a/operator/internal/controllers/cluster/storageclusterops_controller.go b/operator/internal/controllers/cluster/storageclusterops_controller.go index 172deffed..ed12cf8f1 100644 --- a/operator/internal/controllers/cluster/storageclusterops_controller.go +++ b/operator/internal/controllers/cluster/storageclusterops_controller.go @@ -64,8 +64,20 @@ const ( // something it cannot hurry: a lock another operation holds, or a step // waiting on the control plane. A queued operation is normally woken by // its cluster rather than by this, and this is the backstop for when that - // event is missed. - opsRetry = 15 * time.Second + // event is missed (design-storagecluster.md §6.1). + opsRetry = 10 * time.Second + + // opsContended is how long a pass waits when its lock patch was refused + // rather than when it found the lock held. The two are different + // situations: a lock somebody visibly holds is released by work that has + // to finish first, while a 409 means the object moved between this pass's + // read and its write and who holds it now is one read away. + // + // It is shorter than opsRetry for that reason, and it is not zero. An + // immediate requeue against an object two reconcilers are writing is a + // spin: it burns a pass to re-read a value that has not settled, and it + // does so fastest exactly when contention is highest. + opsContended = 5 * time.Second // opsAdvance is how long a pass that moved the operation forward waits // before the next one. It is short because there is nothing to wait for: @@ -207,12 +219,15 @@ func (r *StorageClusterOpsReconciler) Reconcile( return ctrl.Result{}, r.releaseLock(ctx, &ops) } - acquired, err := r.acquireLock(ctx, &ops) + outcome, err := r.acquireLock(ctx, &ops) if err != nil { return ctrl.Result{}, err } - if !acquired { + switch outcome { + case lockHeld: return ctrl.Result{RequeueAfter: opsRetry}, nil + case lockContended: + return ctrl.Result{RequeueAfter: opsContended}, nil } return r.advance(ctx, &ops) @@ -531,26 +546,29 @@ func (r *StorageClusterOpsReconciler) teardown( // what makes the read-then-write safe: two operations can both read an empty // field and both conclude the lock is free, and the patch succeeds for exactly // one of them at a given resourceVersion and returns 409 to the rest. +// +// The outcome is typed rather than a bool, because "not acquired" is two +// situations with different waits (§6.1) and a bool collapses them. func (r *StorageClusterOpsReconciler) acquireLock( ctx context.Context, ops *simplyblockv1alpha2.StorageClusterOps, -) (bool, error) { +) (lockOutcome, error) { var cluster simplyblockv1alpha2.StorageCluster key := types.NamespacedName{Name: ops.Spec.ClusterRef, Namespace: ops.Namespace} err := r.Get(ctx, key, &cluster) if apierrors.IsNotFound(err) { _, err := r.finish(ctx, ops, simplyblockv1alpha2.StorageClusterOpsPhaseFailed, fmt.Sprintf("StorageCluster %s does not exist", ops.Spec.ClusterRef)) - return false, err + return lockHeld, err } if err != nil { - return false, err + return lockHeld, err } if held := cluster.Status.ActiveOpsRef; held != "" && held != ops.Name { r.Recorder.Eventf(ops, nil, corev1.EventTypeNormal, OperationQueued, OperationQueued, "Cluster %s is held by operation %s; this one is waiting", cluster.Name, held) - return false, r.hold(ctx, ops, fmt.Sprintf( + return lockHeld, r.hold(ctx, ops, fmt.Sprintf( "waiting for operation %s to release cluster %s", held, cluster.Name)) } @@ -563,9 +581,9 @@ func (r *StorageClusterOpsReconciler) acquireLock( // Somebody else moved the object between the read and the // write. Whether that was another operation taking the lock is // decided by reading it again rather than guessed at here. - return false, nil + return lockContended, nil } - return false, fmt.Errorf("acquire the lock on cluster %s: %w", cluster.Name, err) + return lockHeld, fmt.Errorf("acquire the lock on cluster %s: %w", cluster.Name, err) } operationActiveState.WithLabelValues(cluster.Name).Set(1) } @@ -583,12 +601,27 @@ func (r *StorageClusterOpsReconciler) acquireLock( status.Message = "The operation holds the cluster and is running" }) if err != nil { - return false, err + return lockHeld, err } } - return true, nil + return lockAcquired, nil } +// lockOutcome is what one attempt at a cluster's lock produced. The two +// unsuccessful values are separate because they are waited on differently +// (§6.1): a lock somebody holds frees when their work finishes, and a refused +// patch resolves on the next read. +type lockOutcome int + +const ( + // lockAcquired: this operation holds the cluster. + lockAcquired lockOutcome = iota + // lockHeld: another operation holds it, or the attempt could not be made. + lockHeld + // lockContended: the optimistic-lock patch was refused. + lockContended +) + // releaseLock clears the cluster's status.activeOpsRef, but only while it still // names this operation. // diff --git a/operator/internal/controllers/deployment/expansion.go b/operator/internal/controllers/deployment/expansion.go index de548b8ee..dcede39ac 100644 --- a/operator/internal/controllers/deployment/expansion.go +++ b/operator/internal/controllers/deployment/expansion.go @@ -617,10 +617,16 @@ func targetClusterName( "the document neither names an existing cluster nor describes one to create") } -// nodeNameFormula names one node. A StorageNode is named for its cluster and the -// slot it fills, never for the worker, because the name has to stay stable when a -// migration re-points the node onto another host (design-storagenode.md §3.1) — -// the worker is in the name's digest rather than in its text. +// nodeNameFormula names one node, from the cluster, the worker, and the slot +// (design-storagenode.md §3.1). All three are in the name's text while it fits and +// in its digest once the limit below forces a truncation, so two nodes differing +// only by worker never collide either way. +// +// The name is stable across a migration because Kubernetes never renames an +// object, not because the worker is kept out of it: a migration re-points +// spec.workerNode on the node that exists. So a node built on one worker and +// migrated to another keeps a name describing where it was built, and +// spec.workerNode rather than the name is where the current host is read. // // The limit is a label's 63 bytes and not the 253 an object name may be, because // the name travels: the StorageDevice mirror writes it into @@ -631,9 +637,10 @@ func targetClusterName( // produce rather than what a long one does. // // Shortening the limit does not strand the nodes of a cluster that already has -// some. A name that fitted the wider limit is returned unchanged whenever it -// also fits this one, and createNodes finds what exists by the worker and slot -// its spec records rather than by re-deriving the name. +// some. A name that fitted the wider limit is returned unchanged whenever it also +// fits this one, and createNodes finds what exists by the worker and slot its spec +// records rather than by re-deriving the name — which is also why idempotent +// re-expansion rests on the spec rather than on this formula. var nodeNameFormula = kube.Formula{Limit: kube.MaxLabelValueLength} func nodeName(cluster, worker string, slot int32) string { diff --git a/operator/internal/controllers/node/remove.go b/operator/internal/controllers/node/remove.go index 8c79e78e2..6b6e0e812 100644 --- a/operator/internal/controllers/node/remove.go +++ b/operator/internal/controllers/node/remove.go @@ -18,12 +18,12 @@ // reconciler rather than here because a failure of any kind owes it, not only a // failure of a step in this file. // -// One thing here is not what the design specifies. §8.4 fans the migration out as -// one PersistentVolumeOps per volume, and that kind has not been written yet — the -// StoragePool rework recorded the same gap. The fan-out is the VolumeMigration -// that exists and works, tracked by the same label and cleaned up by the same -// cascade, and it becomes a PersistentVolumeOps when that kind lands. §15.4 -// records it. +// The migration is fanned out as one move per volume through the mover of +// internal/volumemigration, which raises whichever kind the deployment runs: the +// cluster-scoped PersistentVolumeOps of §8.4, naming this operation in +// spec.creatorRef and carrying the managed-by label, or the registered +// VolumeMigration with a controller reference where that kind is still the one +// in use. Nothing in this file knows which, which is what the mover exists for. // // design-storagenode.md §8 is the specification. diff --git a/operator/internal/controllers/pool/events.go b/operator/internal/controllers/pool/events.go index 53f509137..2a4e93789 100644 --- a/operator/internal/controllers/pool/events.go +++ b/operator/internal/controllers/pool/events.go @@ -39,9 +39,9 @@ const ( QoSParameterConflict = "QoSParameterConflict" // AllowedNodeMissing says an entry in spec.allowedNodes resolves to no - // StorageNode. It is raised once per name rather than every pass, because - // the authored list is deliberately left as written and repeating the event - // would say nothing new. + // Kubernetes Node. It is raised once per name rather than every pass, + // because the authored list is deliberately left as written and repeating + // the event would say nothing new. AllowedNodeMissing = "AllowedNodeMissing" // StorageClassStillAssigned and VolumesStillBound are the two that explain a diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusters.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusters.yaml index e1f616c64..5abed659d 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusters.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusters.yaml @@ -1525,10 +1525,12 @@ spec: rule: '!has(self.state) || self.state in [''Claiming'',''CheckingControlPlane'',''ResolvingConfig'',''Creating'',''Adopting'',''Persisting'']' tasks: description: |- - Tasks are the control plane's running and pending jobs, newest first and - capped at twenty. Completed and canceled tasks are not here: they leave - the list and become events, so the length tracks concurrency rather than - history. + Tasks are the control plane's running and pending jobs, capped at twenty + and in the order the control plane reports them: its TaskDTO carries no + creation date, so newest-first is not orderable from what is on the wire + (design-storagecluster.md §12.1). Completed and canceled tasks are not + here: they leave the list and become events, so the length tracks + concurrency rather than history. items: description: |- ClusterTask is one asynchronous job the control plane is running, as of the diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagepools.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagepools.yaml index 498c28856..7f1b2d29e 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagepools.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagepools.yaml @@ -295,9 +295,10 @@ spec: properties: allowedNodes: description: |- - AllowedNodes restricts which storage nodes may host this pool's volumes. - Empty means every node in the cluster. Narrowing it stops new volumes - landing on the removed nodes and leaves the existing ones where they are. + AllowedNodes restricts which hosts may carry this pool's volumes, by + Kubernetes Node name. Empty means every node in the cluster. Narrowing it + stops new volumes landing on the removed nodes and leaves the existing + ones where they are. The list is left exactly as authored: a name that no longer resolves is dropped from Status.AllowedNodes rather than pruned from here, so a node @@ -480,10 +481,11 @@ spec: type: string allowedNodes: description: |- - AllowedNodes is Spec.AllowedNodes resolved against the StorageNodes that - exist, which is what the control plane is sent. An empty list here is not - the same as an absent Spec.AllowedNodes: absent means every node, and - empty after resolution means the pool can place nothing. + AllowedNodes is Spec.AllowedNodes resolved against the Node objects that + exist, which is what the control plane's host list and the per-pool node + labels are derived from. An empty list here is not the same as an absent + Spec.AllowedNodes: absent means every node, and empty after resolution + means the pool can place nothing. items: type: string type: array From f0ffd197d46931fdad4554c6905fecda3c2b03bc Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Thu, 17 Sep 2026 20:48:24 +0100 Subject: [PATCH 062/206] Implement Phase 1 of csi-addons volume replication: driver Replication/Identity services and operator sidecar wiring --- atlas-lib/controlplane/replication.go | 140 + atlas-lib/controlplane/replication_test.go | 173 + atlas-lib/internal/cpapi/cpapi.gen.go | 8640 +++++++++++------ csi-driver/go.mod | 19 + csi-driver/go.sum | 149 + csi-driver/internal/clusters/clusters.go | 75 +- csi-driver/internal/csi/common/server.go | 13 +- .../internal/csi/controller/errorclass.go | 8 + .../internal/csi/controller/errorclass_rpc.go | 29 + .../csi/controller/mock_controlplane_test.go | 81 +- .../internal/csi/controller/replication.go | 113 + .../csi/controller/replication_test.go | 241 + csi-driver/internal/csi/controller/server.go | 6 + .../csi/csiaddons/identity/identity.go | 61 + .../csi/csiaddons/identity/identity_test.go | 55 + csi-driver/internal/driver/driver.go | 21 +- ...age.simplyblock.io_simplyblockdrivers.yaml | 14 +- .../api/v1alpha2/simplyblockdriver_types.go | 14 +- ...age.simplyblock.io_simplyblockdrivers.yaml | 14 +- .../designs/design-csi-addons-replication.md | 52 +- .../tests/test-plan-csi-addons-replication.md | 83 +- operator/internal/controllers/driver/names.go | 18 + operator/internal/controllers/driver/rbac.go | 53 +- .../internal/controllers/driver/rbac_test.go | 64 + .../controllers/driver/registration_test.go | 1 + .../internal/controllers/driver/sidecars.go | 10 +- .../driver/simplyblockdriver_controller.go | 3 +- .../simplyblockdriver_controller_test.go | 5 + .../internal/controllers/driver/workloads.go | 53 + .../controllers/driver/workloads_test.go | 58 + ...age.simplyblock.io_simplyblockdrivers.yaml | 14 +- shared/openapi.json | 1547 ++- test/integration/controlplane/cpsim.gen.go | 1066 +- .../controlplane/unimplemented.gen.go | 54 +- 34 files changed, 9793 insertions(+), 3154 deletions(-) create mode 100644 atlas-lib/controlplane/replication.go create mode 100644 atlas-lib/controlplane/replication_test.go create mode 100644 csi-driver/internal/csi/controller/replication.go create mode 100644 csi-driver/internal/csi/controller/replication_test.go create mode 100644 csi-driver/internal/csi/csiaddons/identity/identity.go create mode 100644 csi-driver/internal/csi/csiaddons/identity/identity_test.go diff --git a/atlas-lib/controlplane/replication.go b/atlas-lib/controlplane/replication.go new file mode 100644 index 000000000..ce5b942a5 --- /dev/null +++ b/atlas-lib/controlplane/replication.go @@ -0,0 +1,140 @@ +package controlplane + +import ( + "context" + "fmt" + "net/http" + "strings" + "time" + + "github.com/simplyblock/atlas/internal/cpapi" + "github.com/simplyblock/atlas/lvol" +) + +// ReplicationStatus is the typed steady-state replication status of one +// volume, for the volume's whole replicated life -- unlike a cutover-record +// relationship read, this is never a 404 for a volume that exists. +// +// The pointer fields mirror the API's own optionality: a volume that has +// never replicated reports every timing/lag field nil, which is a valid +// answer, not an error, and is a different thing from a genuine zero. +type ReplicationStatus struct { + Role string + State string + + LastReplicatedAt *time.Time + LagSeconds *int + LagBudgetSeconds *int + + OutstandingCount int + OutstandingBytes int + + FailingCount int + MaxRetryReached bool + + LastCycleBytes *int + LastCycleSeconds *int + + Resyncing bool +} + +func replicationStatusFromDTO(d cpapi.ReplicationStatusDTO) ReplicationStatus { + return ReplicationStatus{ + Role: string(d.Role), + State: string(d.State), + LastReplicatedAt: d.LastReplicatedAt, + LagSeconds: d.LagSeconds, + LagBudgetSeconds: d.LagBudgetSeconds, + OutstandingCount: intFrom(d.OutstandingCount), + OutstandingBytes: intFrom(d.OutstandingBytes), + FailingCount: intFrom(d.FailingCount), + MaxRetryReached: boolFrom(d.MaxRetryReached), + LastCycleBytes: d.LastCycleBytes, + LastCycleSeconds: d.LastCycleSeconds, + Resyncing: boolFrom(d.Resyncing), + } +} + +func intFrom(p *int) int { + if p == nil { + return 0 + } + return *p +} + +func boolFrom(p *bool) bool { + if p == nil { + return false + } + return *p +} + +// EnableVolumeReplication attaches the volume to the named replication +// policy, starting replication. Attaching a volume already following that +// same policy is success (the backend's own idempotency); attaching one that +// follows a different policy is refused (control-plane 412), because a +// silent re-attach would force a full re-sync. +func (c *Client) EnableVolumeReplication(ctx context.Context, h lvol.VolumeHandle, policyID string) error { + cluster, pool, volume, err := h.Split() + if err != nil { + return err + } + policy, err := parseUUID("replication policy id", policyID) + if err != nil { + return err + } + resp, err := c.api.ClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPutWithResponse( + ctx, cluster, pool, volume, cpapi.UpdatableLVolParams{ReplicationPolicyId: &policy}) + if err != nil { + return fmt.Errorf("enable replication on volume %s: %w", h, err) + } + if code := resp.StatusCode(); code != http.StatusOK && code != http.StatusNoContent { + return respError("enable replication on volume "+string(h), code, resp.Body) + } + return nil +} + +// DisableVolumeReplication detaches the volume from whatever replication +// policy it follows, stopping replication. Detaching a volume that follows no +// policy is success (the backend's own idempotency); a 409 (a cutover in +// flight) is a retryable refusal. +func (c *Client) DisableVolumeReplication(ctx context.Context, h lvol.VolumeHandle) error { + cluster, pool, volume, err := h.Split() + if err != nil { + return err + } + // ReplicationPolicyId is `omitempty` on the generated request struct, so + // building it with a nil pointer would drop the key entirely rather than + // send it as an explicit null -- and the backend distinguishes "the key + // was absent" (leave the policy alone) from "the key was null" (detach) + // by which keys the request body carries, not by the decoded value. The + // raw-body variant is the only way to say "null" here. + resp, err := c.api.ClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPutWithBodyWithResponse( + ctx, cluster, pool, volume, "application/json", strings.NewReader(`{"replication_policy_id":null}`)) + if err != nil { + return fmt.Errorf("disable replication on volume %s: %w", h, err) + } + if code := resp.StatusCode(); code != http.StatusOK && code != http.StatusNoContent { + return respError("disable replication on volume "+string(h), code, resp.Body) + } + return nil +} + +// GetVolumeReplicationInfo returns the volume's typed steady-state +// replication status. +func (c *Client) GetVolumeReplicationInfo(ctx context.Context, h lvol.VolumeHandle) (ReplicationStatus, error) { + cluster, pool, volume, err := h.Split() + if err != nil { + return ReplicationStatus{}, err + } + resp, err := c.api.ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGetWithResponse( + ctx, cluster, pool, volume) + if err != nil { + return ReplicationStatus{}, fmt.Errorf("replication status of volume %s: %w", h, err) + } + d, err := payload("replication status of volume "+string(h), resp.JSON200, resp.StatusCode(), resp.Body) + if err != nil { + return ReplicationStatus{}, err + } + return replicationStatusFromDTO(*d), nil +} diff --git a/atlas-lib/controlplane/replication_test.go b/atlas-lib/controlplane/replication_test.go new file mode 100644 index 000000000..1057e7801 --- /dev/null +++ b/atlas-lib/controlplane/replication_test.go @@ -0,0 +1,173 @@ +package controlplane + +import ( + "context" + "errors" + "io" + "net/http" + "strings" + "testing" + + "github.com/simplyblock/atlas/errs" +) + +func TestClientEnableVolumeReplication(t *testing.T) { + var gotBody string + c := newTestClient(t, func(w http.ResponseWriter, r *http.Request) { + if r.Method != http.MethodPut { + t.Errorf("method = %s, want PUT", r.Method) + } + b, _ := io.ReadAll(r.Body) + gotBody = string(b) + w.WriteHeader(http.StatusNoContent) + }) + if err := c.EnableVolumeReplication(context.Background(), testHandle, testPolicy); err != nil { + t.Fatal(err) + } + if !strings.Contains(gotBody, `"replication_policy_id":"`+testPolicy+`"`) { + t.Errorf("request body = %q, want it to carry replication_policy_id %s", gotBody, testPolicy) + } +} + +func TestClientEnableVolumeReplicationDifferentPolicyIsAnError(t *testing.T) { + c := newTestClient(t, func(w http.ResponseWriter, r *http.Request) { + w.WriteHeader(http.StatusPreconditionFailed) + _, _ = w.Write([]byte("already attached to a different policy")) + }) + err := c.EnableVolumeReplication(context.Background(), testHandle, testPolicy) + var se *StatusError + if !errors.As(err, &se) || se.StatusCode != http.StatusPreconditionFailed { + t.Fatalf("err = %v, want a *StatusError carrying 412", err) + } +} + +// DisableVolumeReplication must send an EXPLICIT JSON null for +// replication_policy_id, not omit the key. The generated request struct's +// field is `omitempty`, so a nil pointer there would be dropped from the body +// entirely -- and the backend distinguishes "the key was absent" (leave the +// policy alone) from "the key was null" (detach) by which keys the request +// actually carries, not by the value. A body of `{}` would silently do +// nothing. +func TestClientDisableVolumeReplicationSendsExplicitNull(t *testing.T) { + var gotBody string + c := newTestClient(t, func(w http.ResponseWriter, r *http.Request) { + if r.Method != http.MethodPut { + t.Errorf("method = %s, want PUT", r.Method) + } + b, _ := io.ReadAll(r.Body) + gotBody = string(b) + w.WriteHeader(http.StatusNoContent) + }) + if err := c.DisableVolumeReplication(context.Background(), testHandle); err != nil { + t.Fatal(err) + } + if !strings.Contains(gotBody, `"replication_policy_id":null`) { + t.Errorf("request body = %q, want an explicit null for replication_policy_id", gotBody) + } +} + +func TestClientDisableVolumeReplicationNotAttachedIsSuccess(t *testing.T) { + // The backend's own idempotency: detaching a volume that follows no + // policy already returns success, so the client has nothing extra to do + // here beyond not treating any 2xx as an error. + c := newTestClient(t, func(w http.ResponseWriter, r *http.Request) { + w.WriteHeader(http.StatusOK) + }) + if err := c.DisableVolumeReplication(context.Background(), testHandle); err != nil { + t.Errorf("DisableVolumeReplication = %v, want nil", err) + } +} + +func TestClientDisableVolumeReplicationDuringCutoverIsAnError(t *testing.T) { + c := newTestClient(t, func(w http.ResponseWriter, r *http.Request) { + w.WriteHeader(http.StatusConflict) + }) + err := c.DisableVolumeReplication(context.Background(), testHandle) + if !errors.Is(err, errs.ErrAlreadyExists) { + t.Fatalf("err = %v, want it to unwrap to a 409 sentinel so the caller can retry", err) + } +} + +func TestClientGetVolumeReplicationInfo(t *testing.T) { + c := newTestClient(t, func(w http.ResponseWriter, r *http.Request) { + if !strings.HasSuffix(r.URL.Path, "/replication/status") { + t.Errorf("unexpected path %q", r.URL.Path) + } + w.Header().Set("Content-Type", "application/json") + _, _ = w.Write([]byte(`{ + "role": "source", "state": "in_sync", + "last_replicated_at": "2026-09-17T12:00:00Z", + "lag_seconds": 42, "lag_budget_seconds": 900, + "outstanding_count": 1, "outstanding_bytes": 1048576, + "failing_count": 0, "max_retry_reached": false, + "last_cycle_bytes": 2097152, "last_cycle_seconds": 12, + "resyncing": false + }`)) + }) + + info, err := c.GetVolumeReplicationInfo(context.Background(), testHandle) + if err != nil { + t.Fatal(err) + } + if info.Role != "source" || info.State != "in_sync" { + t.Errorf("role/state = %q/%q, want source/in_sync", info.Role, info.State) + } + if info.LastReplicatedAt == nil || info.LastReplicatedAt.Unix() != 1789646400 { + t.Errorf("LastReplicatedAt = %v", info.LastReplicatedAt) + } + if info.LagSeconds == nil || *info.LagSeconds != 42 { + t.Errorf("LagSeconds = %v, want 42", info.LagSeconds) + } + if info.LagBudgetSeconds == nil || *info.LagBudgetSeconds != 900 { + t.Errorf("LagBudgetSeconds = %v, want 900", info.LagBudgetSeconds) + } + if info.OutstandingCount != 1 || info.OutstandingBytes != 1048576 { + t.Errorf("outstanding = %d/%d, want 1/1048576", info.OutstandingCount, info.OutstandingBytes) + } + if info.FailingCount != 0 || info.MaxRetryReached { + t.Errorf("failing/max-retry = %d/%v, want 0/false", info.FailingCount, info.MaxRetryReached) + } + if info.LastCycleBytes == nil || *info.LastCycleBytes != 2097152 { + t.Errorf("LastCycleBytes = %v, want 2097152", info.LastCycleBytes) + } + if info.LastCycleSeconds == nil || *info.LastCycleSeconds != 12 { + t.Errorf("LastCycleSeconds = %v, want 12", info.LastCycleSeconds) + } + if info.Resyncing { + t.Error("Resyncing = true, want false") + } +} + +// A volume that never replicated is a valid, non-error answer: role "none", +// state "not_replicating", and every timing/lag field null. +func TestClientGetVolumeReplicationInfoNeverReplicated(t *testing.T) { + c := newTestClient(t, func(w http.ResponseWriter, r *http.Request) { + w.Header().Set("Content-Type", "application/json") + _, _ = w.Write([]byte(`{ + "role": "none", "state": "not_replicating", + "outstanding_count": 0, "outstanding_bytes": 0, + "failing_count": 0, "max_retry_reached": false, "resyncing": false + }`)) + }) + + info, err := c.GetVolumeReplicationInfo(context.Background(), testHandle) + if err != nil { + t.Fatal(err) + } + if info.Role != "none" || info.State != "not_replicating" { + t.Errorf("role/state = %q/%q", info.Role, info.State) + } + if info.LastReplicatedAt != nil || info.LagSeconds != nil || info.LagBudgetSeconds != nil { + t.Errorf("expected every timing field nil, got LastReplicatedAt=%v LagSeconds=%v LagBudgetSeconds=%v", + info.LastReplicatedAt, info.LagSeconds, info.LagBudgetSeconds) + } +} + +func TestClientGetVolumeReplicationInfoNotFound(t *testing.T) { + c := newTestClient(t, func(w http.ResponseWriter, r *http.Request) { + w.WriteHeader(http.StatusNotFound) + }) + if _, err := c.GetVolumeReplicationInfo(context.Background(), testHandle); !errors.Is(err, errs.ErrNotFound) { + t.Errorf("err = %v, want ErrNotFound", err) + } +} diff --git a/atlas-lib/internal/cpapi/cpapi.gen.go b/atlas-lib/internal/cpapi/cpapi.gen.go index b69a0ab2a..2f8e51a0d 100644 --- a/atlas-lib/internal/cpapi/cpapi.gen.go +++ b/atlas-lib/internal/cpapi/cpapi.gen.go @@ -18,6 +18,42 @@ import ( openapi_types "github.com/oapi-codegen/runtime/types" ) +// Defines values for AlertDTOSeverity. +const ( + AlertDTOSeverityCritical AlertDTOSeverity = "critical" + AlertDTOSeverityWarning AlertDTOSeverity = "warning" +) + +// Valid indicates whether the value is a known member of the AlertDTOSeverity enum. +func (e AlertDTOSeverity) Valid() bool { + switch e { + case AlertDTOSeverityCritical: + return true + case AlertDTOSeverityWarning: + return true + default: + return false + } +} + +// Defines values for AlertDTOStatus. +const ( + AlertDTOStatusFiring AlertDTOStatus = "firing" + AlertDTOStatusResolved AlertDTOStatus = "resolved" +) + +// Valid indicates whether the value is a known member of the AlertDTOStatus enum. +func (e AlertDTOStatus) Valid() bool { + switch e { + case AlertDTOStatusFiring: + return true + case AlertDTOStatusResolved: + return true + default: + return false + } +} + // Defines values for ClusterDTOStatus. const ( ClusterDTOStatusActive ClusterDTOStatus = "active" @@ -261,6 +297,60 @@ func (e ReplicationStartParamsMode) Valid() bool { } } +// Defines values for ReplicationStatusDTORole. +const ( + ReplicationStatusDTORoleFailedOver ReplicationStatusDTORole = "failed_over" + ReplicationStatusDTORoleNone ReplicationStatusDTORole = "none" + ReplicationStatusDTORoleSecondary ReplicationStatusDTORole = "secondary" + ReplicationStatusDTORoleSource ReplicationStatusDTORole = "source" +) + +// Valid indicates whether the value is a known member of the ReplicationStatusDTORole enum. +func (e ReplicationStatusDTORole) Valid() bool { + switch e { + case ReplicationStatusDTORoleFailedOver: + return true + case ReplicationStatusDTORoleNone: + return true + case ReplicationStatusDTORoleSecondary: + return true + case ReplicationStatusDTORoleSource: + return true + default: + return false + } +} + +// Defines values for ReplicationStatusDTOState. +const ( + ReplicationStatusDTOStateDegraded ReplicationStatusDTOState = "degraded" + ReplicationStatusDTOStateError ReplicationStatusDTOState = "error" + ReplicationStatusDTOStateInSync ReplicationStatusDTOState = "in_sync" + ReplicationStatusDTOStateLagging ReplicationStatusDTOState = "lagging" + ReplicationStatusDTOStateNotReplicating ReplicationStatusDTOState = "not_replicating" + ReplicationStatusDTOStateReplicating ReplicationStatusDTOState = "replicating" +) + +// Valid indicates whether the value is a known member of the ReplicationStatusDTOState enum. +func (e ReplicationStatusDTOState) Valid() bool { + switch e { + case ReplicationStatusDTOStateDegraded: + return true + case ReplicationStatusDTOStateError: + return true + case ReplicationStatusDTOStateInSync: + return true + case ReplicationStatusDTOStateLagging: + return true + case ReplicationStatusDTOStateNotReplicating: + return true + case ReplicationStatusDTOStateReplicating: + return true + default: + return false + } +} + // Defines values for ReplicationTargetDTOStatus. const ( ReplicationTargetDTOStatusActive ReplicationTargetDTOStatus = "active" @@ -489,6 +579,42 @@ func (e ClustersCreateApiV2ClustersPostParamsResponseFormat) Valid() bool { } } +// Defines values for ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsSeverity. +const ( + ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsSeverityCritical ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsSeverity = "critical" + ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsSeverityWarning ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsSeverity = "warning" +) + +// Valid indicates whether the value is a known member of the ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsSeverity enum. +func (e ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsSeverity) Valid() bool { + switch e { + case ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsSeverityCritical: + return true + case ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsSeverityWarning: + return true + default: + return false + } +} + +// Defines values for ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsStatus. +const ( + ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsStatusFiring ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsStatus = "firing" + ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsStatusResolved ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsStatus = "resolved" +) + +// Valid indicates whether the value is a known member of the ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsStatus enum. +func (e ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsStatus) Valid() bool { + switch e { + case ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsStatusFiring: + return true + case ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsStatusResolved: + return true + default: + return false + } +} + // Defines values for ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostParamsResponseFormat. const ( ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostParamsResponseFormatEmpty ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostParamsResponseFormat = "empty" @@ -636,6 +762,36 @@ func (e ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMig } } +// AlertDTO One condition that currently needs an operator. +// +// Deliberately NOT an EventObj. An event is a journal entry -- it happened, +// it is kept forever, and nothing ever retracts it. An alert is a claim +// about the present that goes away by itself when it stops being true, so +// it carries the object it is about and the time the condition started +// rather than the time something was logged. “id“ is derived from the +// kind and the object, so it is stable across polls and a consumer can +// dedupe on it without keeping state. +type AlertDTO struct { + ClusterId openapi_types.UUID `json:"cluster_id"` + Details map[string]interface{} `json:"details"` + DeviceId *openapi_types.UUID `json:"device_id"` + FirstSeen *string `json:"first_seen"` + Id string `json:"id"` + Kind string `json:"kind"` + Message string `json:"message"` + NodeId *openapi_types.UUID `json:"node_id"` + ResolvedAt *string `json:"resolved_at"` + Severity AlertDTOSeverity `json:"severity"` + Since *string `json:"since"` + Status AlertDTOStatus `json:"status"` +} + +// AlertDTOSeverity defines model for AlertDTO.Severity. +type AlertDTOSeverity string + +// AlertDTOStatus defines model for AlertDTO.Status. +type AlertDTOStatus string + // BackupConfigParams defines model for BackupConfigParams. type BackupConfigParams struct { AccessKeyId *string `json:"access_key_id,omitempty"` @@ -799,6 +955,49 @@ type CommitParams struct { DeleteSource *bool `json:"delete_source,omitempty"` } +// ConsistencyGroupDTO A standalone consistency group summary (design §10). +type ConsistencyGroupDTO struct { + ClusterId openapi_types.UUID `json:"cluster_id"` + Id openapi_types.UUID `json:"id"` + LastGroupSeq int `json:"last_group_seq"` + LvsName *string `json:"lvs_name,omitempty"` + MemberCount int `json:"member_count"` + Name string `json:"name"` + NodeId *openapi_types.UUID `json:"node_id,omitempty"` +} + +// ConsistencyGroupGenerationDTO One generation of a consistency group (design §6.3). +type ConsistencyGroupGenerationDTO struct { + Complete bool `json:"complete"` + CreatedAt int `json:"created_at"` + Expected int `json:"expected"` + GroupSeq int `json:"group_seq"` + Members []ConsistencyGroupGenerationMemberDTO `json:"members"` + Present int `json:"present"` +} + +// ConsistencyGroupGenerationMemberDTO defines model for ConsistencyGroupGenerationMemberDTO. +type ConsistencyGroupGenerationMemberDTO struct { + LvolId string `json:"lvol_id"` + Ready bool `json:"ready"` + SnapshotId string `json:"snapshot_id"` +} + +// ConsistencyGroupMemberDTO One current member of a consistency group (design §10 /members). +type ConsistencyGroupMemberDTO struct { + JoinedSeq int `json:"joined_seq"` + LvolId string `json:"lvol_id"` + LvsName string `json:"lvs_name"` + NodeId string `json:"node_id"` + Online bool `json:"online"` + RemovedSeq int `json:"removed_seq"` +} + +// ConsistencyGroupMemberJoinDTO Request body for the late join of an existing volume (design §4.5). +type ConsistencyGroupMemberJoinDTO struct { + LvolId string `json:"lvol_id"` +} + // DeviceDTO defines model for DeviceDTO. type DeviceDTO struct { BdevType *string `json:"bdev_type,omitempty"` @@ -933,6 +1132,7 @@ type PolicyParams struct { KeepReplicated *int `json:"keep_replicated,omitempty"` Mode *PolicyParamsMode `json:"mode,omitempty"` PolicyName string `json:"policy_name"` + RpoTargetSeconds *int `json:"rpo_target_seconds,omitempty"` TargetId openapi_types.UUID `json:"target_id"` } @@ -949,6 +1149,30 @@ type ReplicateLVolParams struct { LvolId openapi_types.UUID `json:"lvol_id"` } +// ReplicatedGenerationDTO One complete, fully replicated consistency-group generation, every +// member addressed as a cloneable object on the secondary. +type ReplicatedGenerationDTO struct { + GroupSeq int `json:"group_seq"` + Members []ReplicatedSnapshotDTO `json:"members"` +} + +// ReplicatedSnapshotDTO A fully replicated snapshot on the secondary, addressed as a cloneable +// object. “lvol_id“ is the volume the snapshot belongs to on the +// SECONDARY cluster, not the source volume the caller asked about, because +// that is the identity the ordinary CSI clone path resolves a +// “dataSource“ against. +type ReplicatedSnapshotDTO struct { + ClusterId openapi_types.UUID `json:"cluster_id"` + CreatedAt time.Time `json:"created_at"` + GroupId *string `json:"group_id,omitempty"` + GroupSeq *int `json:"group_seq,omitempty"` + LvolId *openapi_types.UUID `json:"lvol_id,omitempty"` + PoolId *openapi_types.UUID `json:"pool_id,omitempty"` + Size int `json:"size"` + SnapshotId openapi_types.UUID `json:"snapshot_id"` + UsedSize int `json:"used_size"` +} + // ReplicationPolicyDTO defines model for ReplicationPolicyDTO. type ReplicationPolicyDTO struct { ClusterId openapi_types.UUID `json:"cluster_id"` @@ -961,6 +1185,7 @@ type ReplicationPolicyDTO struct { KeepReplicated int `json:"keep_replicated"` Mode ReplicationPolicyDTOMode `json:"mode"` PolicyName string `json:"policy_name"` + RpoTargetSeconds *int `json:"rpo_target_seconds,omitempty"` Status ReplicationPolicyDTOStatus `json:"status"` TargetId openapi_types.UUID `json:"target_id"` } @@ -1008,6 +1233,34 @@ type ReplicationStartParams struct { // ReplicationStartParamsMode defines model for ReplicationStartParams.Mode. type ReplicationStartParamsMode string +// ReplicationStatusDTO The typed steady-state replication status of one volume. +// +// Serves what “lvol_controller.get_replication_info“ computes, for the +// volume's WHOLE replicated life — unlike “ReplicationRelationshipDTO“, +// which only exists once a cutover or fail-over has created a relationship +// record. “state: not_replicating, role: none“ is a valid answer, never a +// 404, because the csi-addons adapter polls this on every reconcile. +type ReplicationStatusDTO struct { + FailingCount *int `json:"failing_count,omitempty"` + LagBudgetSeconds *int `json:"lag_budget_seconds,omitempty"` + LagSeconds *int `json:"lag_seconds,omitempty"` + LastCycleBytes *int `json:"last_cycle_bytes,omitempty"` + LastCycleSeconds *int `json:"last_cycle_seconds,omitempty"` + LastReplicatedAt *time.Time `json:"last_replicated_at,omitempty"` + MaxRetryReached *bool `json:"max_retry_reached,omitempty"` + OutstandingBytes *int `json:"outstanding_bytes,omitempty"` + OutstandingCount *int `json:"outstanding_count,omitempty"` + Resyncing *bool `json:"resyncing,omitempty"` + Role ReplicationStatusDTORole `json:"role"` + State ReplicationStatusDTOState `json:"state"` +} + +// ReplicationStatusDTORole defines model for ReplicationStatusDTO.Role. +type ReplicationStatusDTORole string + +// ReplicationStatusDTOState defines model for ReplicationStatusDTO.State. +type ReplicationStatusDTOState string + // ReplicationTargetDTO defines model for ReplicationTargetDTO. type ReplicationTargetDTO struct { ClusterId openapi_types.UUID `json:"cluster_id"` @@ -1030,6 +1283,8 @@ type RootModelUnionCreateParamsCloneParams struct { // SnapshotDTO defines model for SnapshotDTO. type SnapshotDTO struct { CreatedAt time.Time `json:"created_at"` + GroupId string `json:"group_id"` + GroupSeq int `json:"group_seq"` HealthCheck bool `json:"health_check"` Id openapi_types.UUID `json:"id"` Lvol *string `json:"lvol"` @@ -1225,6 +1480,8 @@ type VolumeDTO struct { DoReplicate *bool `json:"do_replicate,omitempty"` Fabric string `json:"fabric"` FromSource *bool `json:"from_source,omitempty"` + GroupId *string `json:"group_id,omitempty"` + GroupSeq *int `json:"group_seq,omitempty"` HealthCheck bool `json:"health_check"` HighAvailability bool `json:"high_availability"` Hostname string `json:"hostname"` @@ -1279,6 +1536,7 @@ type UnderscoreBackupSourceSwitchParams struct { // UnderscoreCloneParams defines model for _CloneParams. type UnderscoreCloneParams struct { + ConsistencyGroup *string `json:"consistency_group,omitempty"` DeleteSnapOnLvolDelete *bool `json:"delete_snap_on_lvol_delete,omitempty"` Name string `json:"name"` PvcName *string `json:"pvc_name,omitempty"` @@ -1296,6 +1554,7 @@ type UnderscoreContinueParams struct { // UnderscoreCreateParams defines model for _CreateParams. type UnderscoreCreateParams struct { AllowedHosts *[]string `json:"allowed_hosts,omitempty"` + ConsistencyGroup *string `json:"consistency_group,omitempty"` DoReplicate *bool `json:"do_replicate,omitempty"` Encrypt *bool `json:"encrypt,omitempty"` Fabric *string `json:"fabric,omitempty"` @@ -1401,6 +1660,27 @@ type ClustersDetailApiV2ClustersClusterIdGetParams struct { Watch *bool `form:"watch,omitempty" json:"watch,omitempty"` } +// ClustersAlertsListApiV2ClustersClusterIdAlertsGetParams defines parameters for ClustersAlertsListApiV2ClustersClusterIdAlertsGet. +type ClustersAlertsListApiV2ClustersClusterIdAlertsGetParams struct { + // Severity Only return alerts of this severity + Severity *ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsSeverity `form:"severity,omitempty" json:"severity,omitempty"` + + // History Also return alerts that have already resolved + History *bool `form:"history,omitempty" json:"history,omitempty"` + + // HistorySeconds Limit the history to alerts resolved within this many seconds. Implies history=true. + HistorySeconds *int `form:"history_seconds,omitempty" json:"history_seconds,omitempty"` + + // Status Only return alerts in this state + Status *ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsStatus `form:"status,omitempty" json:"status,omitempty"` +} + +// ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsSeverity defines parameters for ClustersAlertsListApiV2ClustersClusterIdAlertsGet. +type ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsSeverity string + +// ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsStatus defines parameters for ClustersAlertsListApiV2ClustersClusterIdAlertsGet. +type ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsStatus string + // ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostParams defines parameters for ClustersBackupsCreateApiV2ClustersClusterIdBackupsPost. type ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostParams struct { ResponseFormat *ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostParamsResponseFormat `form:"response-format,omitempty" json:"response-format,omitempty"` @@ -1423,6 +1703,11 @@ type ClustersCapacityApiV2ClustersClusterIdCapacityGetParams struct { History *string `form:"history,omitempty" json:"history,omitempty"` } +// ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetParams defines parameters for ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGet. +type ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetParams struct { + Name *string `form:"name,omitempty" json:"name,omitempty"` +} + // ClustersIostatsApiV2ClustersClusterIdIostatsGetParams defines parameters for ClustersIostatsApiV2ClustersClusterIdIostatsGet. type ClustersIostatsApiV2ClustersClusterIdIostatsGetParams struct { History *string `form:"history,omitempty" json:"history,omitempty"` @@ -1559,7 +1844,8 @@ type ClustersStoragePoolsIostatsApiV2ClustersClusterIdStoragePoolsPoolIdIostatsG // ClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsGetParams defines parameters for ClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsGet. type ClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsGetParams struct { // Watch Stream state changes as Server-Sent Events instead of returning a plain response: a `snapshot` event with the current state first, then `created`/`updated`/`deleted` events carrying the full resource representation. A `deleted` event carries the resource's final state when it is still retrievable (e.g. a volume whose status became `deleted`), or an empty object once it is gone entirely. Streams do not support resume; reconnecting clients receive a fresh snapshot. Changes written by pre-upgrade components may take up to 30 seconds to appear. - Watch *bool `form:"watch,omitempty" json:"watch,omitempty"` + Watch *bool `form:"watch,omitempty" json:"watch,omitempty"` + ConsistencyGroup *string `form:"consistency_group,omitempty" json:"consistency_group,omitempty"` } // ClustersStoragePoolsSnapshotsDetailApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdGetParams defines parameters for ClustersStoragePoolsSnapshotsDetailApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdGet. @@ -1691,6 +1977,9 @@ type ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostJSONRequestBo // ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostJSONRequestBody defines body for ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPost for application/json ContentType. type ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostJSONRequestBody = UnderscoreBackupSourceSwitchParams +// ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostJSONRequestBody defines body for ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPost for application/json ContentType. +type ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostJSONRequestBody = ConsistencyGroupMemberJoinDTO + // ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostJSONRequestBody defines body for ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPost for application/json ContentType. type ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostJSONRequestBody = PolicyParams @@ -2149,6 +2438,28 @@ type ClientInterface interface { // Corresponds with POST /api/v2/clusters/{cluster_id}/addreplication (the `ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPost` operationId). ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPost(ctx context.Context, clusterId openapi_types.UUID, body ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostJSONRequestBody, reqEditors ...RequestEditorFn) (*http.Response, error) + // ClustersAlertsListApiV2ClustersClusterIdAlertsGet Clusters:Alerts:List + // + // The conditions in this cluster that currently need an operator. + // + // This is not the event log. An alert appears only while it is still true + // and disappears on its own once it is not: the node comes back ONLINE, the + // device comes back, the cluster leaves degraded. Conditions an operator + // caused on purpose -- a node they shut down, a device they removed -- are + // not alerts and are not listed. + // + // By default only what is wrong NOW is returned -- every entry has + // ``status: firing``. Pass ``history=true`` to also get the ones that have + // since resolved, each with its ``resolved_at``, or ``history_seconds=N`` + // for just the recent past. Either way both transitions are written to the + // cluster event log as ALERT_RAISED / ALERT_RESOLVED, so a resolution + // reaches an operator whether or not anyone asks for history here. + // + // Critical sorts before warning, and firing before resolved. + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/alerts/ (the `ClustersAlertsListApiV2ClustersClusterIdAlertsGet` operationId). + ClustersAlertsListApiV2ClustersClusterIdAlertsGet(ctx context.Context, clusterId openapi_types.UUID, params *ClustersAlertsListApiV2ClustersClusterIdAlertsGetParams, reqEditors ...RequestEditorFn) (*http.Response, error) + // ClustersBackupsListApiV2ClustersClusterIdBackupsGet Clusters:Backups:List // // Corresponds with GET /api/v2/clusters/{cluster_id}/backups/ (the `ClustersBackupsListApiV2ClustersClusterIdBackupsGet` operationId). @@ -2291,6 +2602,83 @@ type ClientInterface interface { // Corresponds with GET /api/v2/clusters/{cluster_id}/capacity (the `ClustersCapacityApiV2ClustersClusterIdCapacityGet` operationId). ClustersCapacityApiV2ClustersClusterIdCapacityGet(ctx context.Context, clusterId openapi_types.UUID, params *ClustersCapacityApiV2ClustersClusterIdCapacityGetParams, reqEditors ...RequestEditorFn) (*http.Response, error) + // ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGet Clusters:Consistency-Groups:List + // + // List the cluster's consistency groups, or resolve one by name (§10). + // + // Returns an empty list when ``name`` matches no group, so a caller can probe + // existence without a 404. + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/ (the `ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGet` operationId). + ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGet(ctx context.Context, clusterId openapi_types.UUID, params *ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetParams, reqEditors ...RequestEditorFn) (*http.Response, error) + + // ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGet Clusters:Consistency-Groups:Detail + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/ (the `ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGet` operationId). + ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGet(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) + + // ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGet Clusters:Consistency-Groups:Members + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/members (the `ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGet` operationId). + ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGet(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) + + // ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostWithBody Clusters:Consistency-Groups:Members:Join + // + // Join an EXISTING volume to the group (design §4.5, Phase 4 late join). + // + // Validates the pinned placement, the pool, the member cap, and the one-way + // rule; a refusal is a 409 naming the precondition. Idempotent: joining a + // current member returns its membership row unchanged. + // + // Takes any type of body and a specified content type. + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/members (the `ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPost` operationId). + ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostWithBody(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*http.Response, error) + + // ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPost Clusters:Consistency-Groups:Members:Join + // + // Join an EXISTING volume to the group (design §4.5, Phase 4 late join). + // + // Validates the pinned placement, the pool, the member cap, and the one-way + // rule; a refusal is a 409 naming the precondition. Idempotent: joining a + // current member returns its membership row unchanged. + // + // Takes a body of the `application/json` content type. + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/members (the `ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPost` operationId). + ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPost(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, body ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostJSONRequestBody, reqEditors ...RequestEditorFn) (*http.Response, error) + + // ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDelete Clusters:Consistency-Groups:Members:Detach + // + // Detach a member: close its epoch one-way, preserving prior generations (§8.2). + // + // Corresponds with DELETE /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/members/{lvol_id} (the `ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDelete` operationId). + ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDelete(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, lvolId string, reqEditors ...RequestEditorFn) (*http.Response, error) + + // ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGet Clusters:Consistency-Groups:Snapshots:List + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/snapshots (the `ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGet` operationId). + ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGet(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) + + // ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPost Clusters:Consistency-Groups:Snapshots:Take + // + // Take one crash-consistent generation across every current member (§5). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/snapshots (the `ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPost` operationId). + ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPost(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) + + // ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDelete Clusters:Consistency-Groups:Snapshots:Delete + // + // Delete one generation and all its member snapshots; never the group (§10). + // + // Corresponds with DELETE /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/snapshots/{seq} (the `ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDelete` operationId). + ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDelete(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, seq int, reqEditors ...RequestEditorFn) (*http.Response, error) + + // ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGet Clusters:Consistency-Groups:Snapshots:Detail + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/snapshots/{seq} (the `ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGet` operationId). + ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGet(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, seq int, reqEditors ...RequestEditorFn) (*http.Response, error) + // ClustersExpandApiV2ClustersClusterIdExpandPost Clusters:Expand // // Corresponds with POST /api/v2/clusters/{cluster_id}/expand (the `ClustersExpandApiV2ClustersClusterIdExpandPost` operationId). @@ -2345,6 +2733,18 @@ type ClientInterface interface { // Corresponds with POST /api/v2/clusters/{cluster_id}/replication/policies/{policy_id}/failover (the `ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPost` operationId). ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPost(ctx context.Context, clusterId openapi_types.UUID, policyId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) + // ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGet Clusters:Replication:Policies:Latest-Generation + // + // The consistency group's newest fully replicated generation, every + // member as a cloneable object on the secondary. Refused as a 400 when the + // policy has no consistency group, when no generation is complete for + // every current member yet, or when members are already split across + // generations: the same refusal a real group fail-over applies, so a drill + // never addresses a mixed-generation cut. + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/replication/policies/{policy_id}/latest-generation (the `ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGet` operationId). + ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGet(ctx context.Context, clusterId openapi_types.UUID, policyId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) + // ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGet Clusters:Replication:Relationships:Detail // // Replication relationship for a volume, resolvable even when the source volume @@ -2354,6 +2754,16 @@ type ClientInterface interface { // Corresponds with GET /api/v2/clusters/{cluster_id}/replication/relationships/{lvol_id} (the `ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGet` operationId). ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGet(ctx context.Context, clusterId openapi_types.UUID, lvolId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) + // ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGet Clusters:Replication:Relationships:Latest-Snapshot + // + // The volume's newest fully replicated snapshot, on the secondary, as a + // cloneable object. Exists for the volume's whole replicated life: a + // test-failover drill (design §14) resolves its test point through this + // read, without touching the real replication state to find out what it is. + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/replication/relationships/{lvol_id}/latest-snapshot (the `ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGet` operationId). + ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGet(ctx context.Context, clusterId openapi_types.UUID, lvolId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) + // ClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGet Clusters:Replication:Targets:List // // Corresponds with GET /api/v2/clusters/{cluster_id}/replication/targets/ (the `ClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGet` operationId). @@ -2872,6 +3282,19 @@ type ClientInterface interface { // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/start (the `ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPost` operationId). ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPost(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, body ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostJSONRequestBody, reqEditors ...RequestEditorFn) (*http.Response, error) + // ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGet Clusters:Storage-Pools:Volumes:Replication:Status + // + // The typed steady-state replication status. + // + // Unlike the relationship read above, which serves cutover records and 404s + // for a volume's whole healthy replicated life, this endpoint always answers + // for a volume that exists: ``state: not_replicating, role: none`` is the + // valid answer for an unreplicated volume. The csi-addons adapter derives + // its conditions and ``lastSyncTime`` from this read on every reconcile. + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/status (the `ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGet` operationId). + ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGet(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) + // ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPost Clusters:Storage-Pools:Volumes:Replication:Stop // // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/stop (the `ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPost` operationId). @@ -3192,6 +3615,38 @@ func (c *Client) ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPost( return c.Client.Do(req) } +// ClustersAlertsListApiV2ClustersClusterIdAlertsGet Clusters:Alerts:List +// +// The conditions in this cluster that currently need an operator. +// +// This is not the event log. An alert appears only while it is still true +// and disappears on its own once it is not: the node comes back ONLINE, the +// device comes back, the cluster leaves degraded. Conditions an operator +// caused on purpose -- a node they shut down, a device they removed -- are +// not alerts and are not listed. +// +// By default only what is wrong NOW is returned -- every entry has +// “status: firing“. Pass “history=true“ to also get the ones that have +// since resolved, each with its “resolved_at“, or “history_seconds=N“ +// for just the recent past. Either way both transitions are written to the +// cluster event log as ALERT_RAISED / ALERT_RESOLVED, so a resolution +// reaches an operator whether or not anyone asks for history here. +// +// Critical sorts before warning, and firing before resolved. +// +// Corresponds with GET /api/v2/clusters/{cluster_id}/alerts/ (the `ClustersAlertsListApiV2ClustersClusterIdAlertsGet` operationId). +func (c *Client) ClustersAlertsListApiV2ClustersClusterIdAlertsGet(ctx context.Context, clusterId openapi_types.UUID, params *ClustersAlertsListApiV2ClustersClusterIdAlertsGetParams, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersAlertsListApiV2ClustersClusterIdAlertsGetRequest(c.Server, clusterId, params) + if err != nil { + return nil, err + } + req = req.WithContext(ctx) + if err := c.applyEditors(ctx, req, reqEditors); err != nil { + return nil, err + } + return c.Client.Do(req) +} + // ClustersBackupsListApiV2ClustersClusterIdBackupsGet Clusters:Backups:List // // Corresponds with GET /api/v2/clusters/{cluster_id}/backups/ (the `ClustersBackupsListApiV2ClustersClusterIdBackupsGet` operationId). @@ -3553,11 +4008,16 @@ func (c *Client) ClustersCapacityApiV2ClustersClusterIdCapacityGet(ctx context.C return c.Client.Do(req) } -// ClustersExpandApiV2ClustersClusterIdExpandPost Clusters:Expand +// ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGet Clusters:Consistency-Groups:List // -// Corresponds with POST /api/v2/clusters/{cluster_id}/expand (the `ClustersExpandApiV2ClustersClusterIdExpandPost` operationId). -func (c *Client) ClustersExpandApiV2ClustersClusterIdExpandPost(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) { - req, err := NewClustersExpandApiV2ClustersClusterIdExpandPostRequest(c.Server, clusterId) +// List the cluster's consistency groups, or resolve one by name (§10). +// +// Returns an empty list when “name“ matches no group, so a caller can probe +// existence without a 404. +// +// Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/ (the `ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGet` operationId). +func (c *Client) ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGet(ctx context.Context, clusterId openapi_types.UUID, params *ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetParams, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetRequest(c.Server, clusterId, params) if err != nil { return nil, err } @@ -3568,11 +4028,11 @@ func (c *Client) ClustersExpandApiV2ClustersClusterIdExpandPost(ctx context.Cont return c.Client.Do(req) } -// ClustersIostatsApiV2ClustersClusterIdIostatsGet Clusters:Iostats +// ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGet Clusters:Consistency-Groups:Detail // -// Corresponds with GET /api/v2/clusters/{cluster_id}/iostats (the `ClustersIostatsApiV2ClustersClusterIdIostatsGet` operationId). -func (c *Client) ClustersIostatsApiV2ClustersClusterIdIostatsGet(ctx context.Context, clusterId openapi_types.UUID, params *ClustersIostatsApiV2ClustersClusterIdIostatsGetParams, reqEditors ...RequestEditorFn) (*http.Response, error) { - req, err := NewClustersIostatsApiV2ClustersClusterIdIostatsGetRequest(c.Server, clusterId, params) +// Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/ (the `ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGet` operationId). +func (c *Client) ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGet(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGetRequest(c.Server, clusterId, groupId) if err != nil { return nil, err } @@ -3583,11 +4043,11 @@ func (c *Client) ClustersIostatsApiV2ClustersClusterIdIostatsGet(ctx context.Con return c.Client.Do(req) } -// ClustersLogsApiV2ClustersClusterIdLogsGet Clusters:Logs +// ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGet Clusters:Consistency-Groups:Members // -// Corresponds with GET /api/v2/clusters/{cluster_id}/logs (the `ClustersLogsApiV2ClustersClusterIdLogsGet` operationId). -func (c *Client) ClustersLogsApiV2ClustersClusterIdLogsGet(ctx context.Context, clusterId openapi_types.UUID, params *ClustersLogsApiV2ClustersClusterIdLogsGetParams, reqEditors ...RequestEditorFn) (*http.Response, error) { - req, err := NewClustersLogsApiV2ClustersClusterIdLogsGetRequest(c.Server, clusterId, params) +// Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/members (the `ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGet` operationId). +func (c *Client) ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGet(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGetRequest(c.Server, clusterId, groupId) if err != nil { return nil, err } @@ -3598,11 +4058,19 @@ func (c *Client) ClustersLogsApiV2ClustersClusterIdLogsGet(ctx context.Context, return c.Client.Do(req) } -// ClustersRebalanceApiV2ClustersClusterIdRebalancePost Clusters:Rebalance +// ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostWithBody Clusters:Consistency-Groups:Members:Join // -// Corresponds with POST /api/v2/clusters/{cluster_id}/rebalance (the `ClustersRebalanceApiV2ClustersClusterIdRebalancePost` operationId). -func (c *Client) ClustersRebalanceApiV2ClustersClusterIdRebalancePost(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) { - req, err := NewClustersRebalanceApiV2ClustersClusterIdRebalancePostRequest(c.Server, clusterId) +// Join an EXISTING volume to the group (design §4.5, Phase 4 late join). +// +// Validates the pinned placement, the pool, the member cap, and the one-way +// rule; a refusal is a 409 naming the precondition. Idempotent: joining a +// current member returns its membership row unchanged. +// +// Takes any type of body and a specified content type. +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/members (the `ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPost` operationId). +func (c *Client) ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostWithBody(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostRequestWithBody(c.Server, clusterId, groupId, contentType, body) if err != nil { return nil, err } @@ -3613,11 +4081,19 @@ func (c *Client) ClustersRebalanceApiV2ClustersClusterIdRebalancePost(ctx contex return c.Client.Do(req) } -// ClustersReplicationPoliciesListApiV2ClustersClusterIdReplicationPoliciesGet Clusters:Replication:Policies:List +// ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPost Clusters:Consistency-Groups:Members:Join // -// Corresponds with GET /api/v2/clusters/{cluster_id}/replication/policies/ (the `ClustersReplicationPoliciesListApiV2ClustersClusterIdReplicationPoliciesGet` operationId). -func (c *Client) ClustersReplicationPoliciesListApiV2ClustersClusterIdReplicationPoliciesGet(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) { - req, err := NewClustersReplicationPoliciesListApiV2ClustersClusterIdReplicationPoliciesGetRequest(c.Server, clusterId) +// Join an EXISTING volume to the group (design §4.5, Phase 4 late join). +// +// Validates the pinned placement, the pool, the member cap, and the one-way +// rule; a refusal is a 409 naming the precondition. Idempotent: joining a +// current member returns its membership row unchanged. +// +// Takes a body of the `application/json` content type. +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/members (the `ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPost` operationId). +func (c *Client) ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPost(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, body ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostJSONRequestBody, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostRequest(c.Server, clusterId, groupId, body) if err != nil { return nil, err } @@ -3628,13 +4104,13 @@ func (c *Client) ClustersReplicationPoliciesListApiV2ClustersClusterIdReplicatio return c.Client.Do(req) } -// ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostWithBody Clusters:Replication:Policies:Create +// ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDelete Clusters:Consistency-Groups:Members:Detach // -// Takes any type of body and a specified content type. +// Detach a member: close its epoch one-way, preserving prior generations (§8.2). // -// Corresponds with POST /api/v2/clusters/{cluster_id}/replication/policies/ (the `ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPost` operationId). -func (c *Client) ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostWithBody(ctx context.Context, clusterId openapi_types.UUID, params *ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostParams, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*http.Response, error) { - req, err := NewClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostRequestWithBody(c.Server, clusterId, params, contentType, body) +// Corresponds with DELETE /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/members/{lvol_id} (the `ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDelete` operationId). +func (c *Client) ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDelete(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, lvolId string, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDeleteRequest(c.Server, clusterId, groupId, lvolId) if err != nil { return nil, err } @@ -3645,13 +4121,11 @@ func (c *Client) ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicat return c.Client.Do(req) } -// ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPost Clusters:Replication:Policies:Create +// ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGet Clusters:Consistency-Groups:Snapshots:List // -// Takes a body of the `application/json` content type. -// -// Corresponds with POST /api/v2/clusters/{cluster_id}/replication/policies/ (the `ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPost` operationId). -func (c *Client) ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPost(ctx context.Context, clusterId openapi_types.UUID, params *ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostParams, body ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostJSONRequestBody, reqEditors ...RequestEditorFn) (*http.Response, error) { - req, err := NewClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostRequest(c.Server, clusterId, params, body) +// Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/snapshots (the `ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGet` operationId). +func (c *Client) ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGet(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetRequest(c.Server, clusterId, groupId) if err != nil { return nil, err } @@ -3662,11 +4136,13 @@ func (c *Client) ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicat return c.Client.Do(req) } -// ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDelete Clusters:Replication:Policies:Delete +// ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPost Clusters:Consistency-Groups:Snapshots:Take // -// Corresponds with DELETE /api/v2/clusters/{cluster_id}/replication/policies/{policy_id}/ (the `ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDelete` operationId). -func (c *Client) ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDelete(ctx context.Context, clusterId openapi_types.UUID, policyId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) { - req, err := NewClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDeleteRequest(c.Server, clusterId, policyId) +// Take one crash-consistent generation across every current member (§5). +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/snapshots (the `ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPost` operationId). +func (c *Client) ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPost(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPostRequest(c.Server, clusterId, groupId) if err != nil { return nil, err } @@ -3677,11 +4153,13 @@ func (c *Client) ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicat return c.Client.Do(req) } -// ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGet Clusters:Replication:Policies:Detail +// ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDelete Clusters:Consistency-Groups:Snapshots:Delete // -// Corresponds with GET /api/v2/clusters/{cluster_id}/replication/policies/{policy_id}/ (the `ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGet` operationId). -func (c *Client) ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGet(ctx context.Context, clusterId openapi_types.UUID, policyId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) { - req, err := NewClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGetRequest(c.Server, clusterId, policyId) +// Delete one generation and all its member snapshots; never the group (§10). +// +// Corresponds with DELETE /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/snapshots/{seq} (the `ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDelete` operationId). +func (c *Client) ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDelete(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, seq int, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDeleteRequest(c.Server, clusterId, groupId, seq) if err != nil { return nil, err } @@ -3692,11 +4170,11 @@ func (c *Client) ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicat return c.Client.Do(req) } -// ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPost Clusters:Replication:Policies:Failover +// ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGet Clusters:Consistency-Groups:Snapshots:Detail // -// Corresponds with POST /api/v2/clusters/{cluster_id}/replication/policies/{policy_id}/failover (the `ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPost` operationId). -func (c *Client) ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPost(ctx context.Context, clusterId openapi_types.UUID, policyId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) { - req, err := NewClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPostRequest(c.Server, clusterId, policyId) +// Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/snapshots/{seq} (the `ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGet` operationId). +func (c *Client) ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGet(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, seq int, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGetRequest(c.Server, clusterId, groupId, seq) if err != nil { return nil, err } @@ -3707,15 +4185,11 @@ func (c *Client) ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplic return c.Client.Do(req) } -// ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGet Clusters:Replication:Relationships:Detail -// -// Replication relationship for a volume, resolvable even when the source volume -// has been deleted (e.g. after replication-commit --delete-source). The CSI driver -// uses this to redirect NodeStageVolume to the active volume on the target cluster. +// ClustersExpandApiV2ClustersClusterIdExpandPost Clusters:Expand // -// Corresponds with GET /api/v2/clusters/{cluster_id}/replication/relationships/{lvol_id} (the `ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGet` operationId). -func (c *Client) ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGet(ctx context.Context, clusterId openapi_types.UUID, lvolId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) { - req, err := NewClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetRequest(c.Server, clusterId, lvolId) +// Corresponds with POST /api/v2/clusters/{cluster_id}/expand (the `ClustersExpandApiV2ClustersClusterIdExpandPost` operationId). +func (c *Client) ClustersExpandApiV2ClustersClusterIdExpandPost(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersExpandApiV2ClustersClusterIdExpandPostRequest(c.Server, clusterId) if err != nil { return nil, err } @@ -3726,11 +4200,11 @@ func (c *Client) ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdRep return c.Client.Do(req) } -// ClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGet Clusters:Replication:Targets:List +// ClustersIostatsApiV2ClustersClusterIdIostatsGet Clusters:Iostats // -// Corresponds with GET /api/v2/clusters/{cluster_id}/replication/targets/ (the `ClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGet` operationId). -func (c *Client) ClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGet(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) { - req, err := NewClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGetRequest(c.Server, clusterId) +// Corresponds with GET /api/v2/clusters/{cluster_id}/iostats (the `ClustersIostatsApiV2ClustersClusterIdIostatsGet` operationId). +func (c *Client) ClustersIostatsApiV2ClustersClusterIdIostatsGet(ctx context.Context, clusterId openapi_types.UUID, params *ClustersIostatsApiV2ClustersClusterIdIostatsGetParams, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersIostatsApiV2ClustersClusterIdIostatsGetRequest(c.Server, clusterId, params) if err != nil { return nil, err } @@ -3741,13 +4215,11 @@ func (c *Client) ClustersReplicationTargetsListApiV2ClustersClusterIdReplication return c.Client.Do(req) } -// ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostWithBody Clusters:Replication:Targets:Create -// -// Takes any type of body and a specified content type. +// ClustersLogsApiV2ClustersClusterIdLogsGet Clusters:Logs // -// Corresponds with POST /api/v2/clusters/{cluster_id}/replication/targets/ (the `ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPost` operationId). -func (c *Client) ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostWithBody(ctx context.Context, clusterId openapi_types.UUID, params *ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostParams, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*http.Response, error) { - req, err := NewClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostRequestWithBody(c.Server, clusterId, params, contentType, body) +// Corresponds with GET /api/v2/clusters/{cluster_id}/logs (the `ClustersLogsApiV2ClustersClusterIdLogsGet` operationId). +func (c *Client) ClustersLogsApiV2ClustersClusterIdLogsGet(ctx context.Context, clusterId openapi_types.UUID, params *ClustersLogsApiV2ClustersClusterIdLogsGetParams, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersLogsApiV2ClustersClusterIdLogsGetRequest(c.Server, clusterId, params) if err != nil { return nil, err } @@ -3758,9 +4230,211 @@ func (c *Client) ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicati return c.Client.Do(req) } -// ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPost Clusters:Replication:Targets:Create +// ClustersRebalanceApiV2ClustersClusterIdRebalancePost Clusters:Rebalance // -// Takes a body of the `application/json` content type. +// Corresponds with POST /api/v2/clusters/{cluster_id}/rebalance (the `ClustersRebalanceApiV2ClustersClusterIdRebalancePost` operationId). +func (c *Client) ClustersRebalanceApiV2ClustersClusterIdRebalancePost(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersRebalanceApiV2ClustersClusterIdRebalancePostRequest(c.Server, clusterId) + if err != nil { + return nil, err + } + req = req.WithContext(ctx) + if err := c.applyEditors(ctx, req, reqEditors); err != nil { + return nil, err + } + return c.Client.Do(req) +} + +// ClustersReplicationPoliciesListApiV2ClustersClusterIdReplicationPoliciesGet Clusters:Replication:Policies:List +// +// Corresponds with GET /api/v2/clusters/{cluster_id}/replication/policies/ (the `ClustersReplicationPoliciesListApiV2ClustersClusterIdReplicationPoliciesGet` operationId). +func (c *Client) ClustersReplicationPoliciesListApiV2ClustersClusterIdReplicationPoliciesGet(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersReplicationPoliciesListApiV2ClustersClusterIdReplicationPoliciesGetRequest(c.Server, clusterId) + if err != nil { + return nil, err + } + req = req.WithContext(ctx) + if err := c.applyEditors(ctx, req, reqEditors); err != nil { + return nil, err + } + return c.Client.Do(req) +} + +// ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostWithBody Clusters:Replication:Policies:Create +// +// Takes any type of body and a specified content type. +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/replication/policies/ (the `ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPost` operationId). +func (c *Client) ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostWithBody(ctx context.Context, clusterId openapi_types.UUID, params *ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostParams, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostRequestWithBody(c.Server, clusterId, params, contentType, body) + if err != nil { + return nil, err + } + req = req.WithContext(ctx) + if err := c.applyEditors(ctx, req, reqEditors); err != nil { + return nil, err + } + return c.Client.Do(req) +} + +// ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPost Clusters:Replication:Policies:Create +// +// Takes a body of the `application/json` content type. +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/replication/policies/ (the `ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPost` operationId). +func (c *Client) ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPost(ctx context.Context, clusterId openapi_types.UUID, params *ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostParams, body ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostJSONRequestBody, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostRequest(c.Server, clusterId, params, body) + if err != nil { + return nil, err + } + req = req.WithContext(ctx) + if err := c.applyEditors(ctx, req, reqEditors); err != nil { + return nil, err + } + return c.Client.Do(req) +} + +// ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDelete Clusters:Replication:Policies:Delete +// +// Corresponds with DELETE /api/v2/clusters/{cluster_id}/replication/policies/{policy_id}/ (the `ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDelete` operationId). +func (c *Client) ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDelete(ctx context.Context, clusterId openapi_types.UUID, policyId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDeleteRequest(c.Server, clusterId, policyId) + if err != nil { + return nil, err + } + req = req.WithContext(ctx) + if err := c.applyEditors(ctx, req, reqEditors); err != nil { + return nil, err + } + return c.Client.Do(req) +} + +// ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGet Clusters:Replication:Policies:Detail +// +// Corresponds with GET /api/v2/clusters/{cluster_id}/replication/policies/{policy_id}/ (the `ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGet` operationId). +func (c *Client) ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGet(ctx context.Context, clusterId openapi_types.UUID, policyId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGetRequest(c.Server, clusterId, policyId) + if err != nil { + return nil, err + } + req = req.WithContext(ctx) + if err := c.applyEditors(ctx, req, reqEditors); err != nil { + return nil, err + } + return c.Client.Do(req) +} + +// ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPost Clusters:Replication:Policies:Failover +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/replication/policies/{policy_id}/failover (the `ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPost` operationId). +func (c *Client) ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPost(ctx context.Context, clusterId openapi_types.UUID, policyId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPostRequest(c.Server, clusterId, policyId) + if err != nil { + return nil, err + } + req = req.WithContext(ctx) + if err := c.applyEditors(ctx, req, reqEditors); err != nil { + return nil, err + } + return c.Client.Do(req) +} + +// ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGet Clusters:Replication:Policies:Latest-Generation +// +// The consistency group's newest fully replicated generation, every +// member as a cloneable object on the secondary. Refused as a 400 when the +// policy has no consistency group, when no generation is complete for +// every current member yet, or when members are already split across +// generations: the same refusal a real group fail-over applies, so a drill +// never addresses a mixed-generation cut. +// +// Corresponds with GET /api/v2/clusters/{cluster_id}/replication/policies/{policy_id}/latest-generation (the `ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGet` operationId). +func (c *Client) ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGet(ctx context.Context, clusterId openapi_types.UUID, policyId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGetRequest(c.Server, clusterId, policyId) + if err != nil { + return nil, err + } + req = req.WithContext(ctx) + if err := c.applyEditors(ctx, req, reqEditors); err != nil { + return nil, err + } + return c.Client.Do(req) +} + +// ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGet Clusters:Replication:Relationships:Detail +// +// Replication relationship for a volume, resolvable even when the source volume +// has been deleted (e.g. after replication-commit --delete-source). The CSI driver +// uses this to redirect NodeStageVolume to the active volume on the target cluster. +// +// Corresponds with GET /api/v2/clusters/{cluster_id}/replication/relationships/{lvol_id} (the `ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGet` operationId). +func (c *Client) ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGet(ctx context.Context, clusterId openapi_types.UUID, lvolId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetRequest(c.Server, clusterId, lvolId) + if err != nil { + return nil, err + } + req = req.WithContext(ctx) + if err := c.applyEditors(ctx, req, reqEditors); err != nil { + return nil, err + } + return c.Client.Do(req) +} + +// ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGet Clusters:Replication:Relationships:Latest-Snapshot +// +// The volume's newest fully replicated snapshot, on the secondary, as a +// cloneable object. Exists for the volume's whole replicated life: a +// test-failover drill (design §14) resolves its test point through this +// read, without touching the real replication state to find out what it is. +// +// Corresponds with GET /api/v2/clusters/{cluster_id}/replication/relationships/{lvol_id}/latest-snapshot (the `ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGet` operationId). +func (c *Client) ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGet(ctx context.Context, clusterId openapi_types.UUID, lvolId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGetRequest(c.Server, clusterId, lvolId) + if err != nil { + return nil, err + } + req = req.WithContext(ctx) + if err := c.applyEditors(ctx, req, reqEditors); err != nil { + return nil, err + } + return c.Client.Do(req) +} + +// ClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGet Clusters:Replication:Targets:List +// +// Corresponds with GET /api/v2/clusters/{cluster_id}/replication/targets/ (the `ClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGet` operationId). +func (c *Client) ClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGet(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGetRequest(c.Server, clusterId) + if err != nil { + return nil, err + } + req = req.WithContext(ctx) + if err := c.applyEditors(ctx, req, reqEditors); err != nil { + return nil, err + } + return c.Client.Do(req) +} + +// ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostWithBody Clusters:Replication:Targets:Create +// +// Takes any type of body and a specified content type. +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/replication/targets/ (the `ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPost` operationId). +func (c *Client) ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostWithBody(ctx context.Context, clusterId openapi_types.UUID, params *ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostParams, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostRequestWithBody(c.Server, clusterId, params, contentType, body) + if err != nil { + return nil, err + } + req = req.WithContext(ctx) + if err := c.applyEditors(ctx, req, reqEditors); err != nil { + return nil, err + } + return c.Client.Do(req) +} + +// ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPost Clusters:Replication:Targets:Create +// +// Takes a body of the `application/json` content type. // // Corresponds with POST /api/v2/clusters/{cluster_id}/replication/targets/ (the `ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPost` operationId). func (c *Client) ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPost(ctx context.Context, clusterId openapi_types.UUID, params *ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostParams, body ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostJSONRequestBody, reqEditors ...RequestEditorFn) (*http.Response, error) { @@ -5014,6 +5688,29 @@ func (c *Client) ClustersStoragePoolsVolumesReplicationStartApiV2ClustersCluster return c.Client.Do(req) } +// ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGet Clusters:Storage-Pools:Volumes:Replication:Status +// +// The typed steady-state replication status. +// +// Unlike the relationship read above, which serves cutover records and 404s +// for a volume's whole healthy replicated life, this endpoint always answers +// for a volume that exists: “state: not_replicating, role: none“ is the +// valid answer for an unreplicated volume. The csi-addons adapter derives +// its conditions and “lastSyncTime“ from this read on every reconcile. +// +// Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/status (the `ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGet` operationId). +func (c *Client) ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGet(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGetRequest(c.Server, clusterId, poolId, volumeId) + if err != nil { + return nil, err + } + req = req.WithContext(ctx) + if err := c.applyEditors(ctx, req, reqEditors); err != nil { + return nil, err + } + return c.Client.Do(req) +} + // ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPost Clusters:Storage-Pools:Volumes:Replication:Stop // // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/stop (the `ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPost` operationId). @@ -5759,53 +6456,8 @@ func NewClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostRequestWit return req, nil } -// NewClustersBackupsListApiV2ClustersClusterIdBackupsGetRequest constructs an http.Request for the ClustersBackupsListApiV2ClustersClusterIdBackupsGet method -func NewClustersBackupsListApiV2ClustersClusterIdBackupsGetRequest(server string, clusterId openapi_types.UUID) (*http.Request, error) { - var err error - - var pathParam0 string - - pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err - } - - serverURL, err := url.Parse(server) - if err != nil { - return nil, err - } - - operationPath := fmt.Sprintf("/api/v2/clusters/%s/backups/", pathParam0) - if operationPath[0] == '/' { - operationPath = "." + operationPath - } - - queryURL, err := serverURL.Parse(operationPath) - if err != nil { - return nil, err - } - - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) - if err != nil { - return nil, err - } - - return req, nil -} - -// NewClustersBackupsCreateApiV2ClustersClusterIdBackupsPostRequest calls the generic ClustersBackupsCreateApiV2ClustersClusterIdBackupsPost builder with application/json body -func NewClustersBackupsCreateApiV2ClustersClusterIdBackupsPostRequest(server string, clusterId openapi_types.UUID, params *ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostParams, body ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostJSONRequestBody) (*http.Request, error) { - var bodyReader io.Reader - buf, err := json.Marshal(body) - if err != nil { - return nil, err - } - bodyReader = bytes.NewReader(buf) - return NewClustersBackupsCreateApiV2ClustersClusterIdBackupsPostRequestWithBody(server, clusterId, params, "application/json", bodyReader) -} - -// NewClustersBackupsCreateApiV2ClustersClusterIdBackupsPostRequestWithBody constructs an http.Request for the ClustersBackupsCreateApiV2ClustersClusterIdBackupsPost method, with any body, and a specified content type -func NewClustersBackupsCreateApiV2ClustersClusterIdBackupsPostRequestWithBody(server string, clusterId openapi_types.UUID, params *ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostParams, contentType string, body io.Reader) (*http.Request, error) { +// NewClustersAlertsListApiV2ClustersClusterIdAlertsGetRequest constructs an http.Request for the ClustersAlertsListApiV2ClustersClusterIdAlertsGet method +func NewClustersAlertsListApiV2ClustersClusterIdAlertsGetRequest(server string, clusterId openapi_types.UUID, params *ClustersAlertsListApiV2ClustersClusterIdAlertsGetParams) (*http.Request, error) { var err error var pathParam0 string @@ -5820,7 +6472,7 @@ func NewClustersBackupsCreateApiV2ClustersClusterIdBackupsPostRequestWithBody(se return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/backups/", pathParam0) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/alerts/", pathParam0) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -5839,9 +6491,9 @@ func NewClustersBackupsCreateApiV2ClustersClusterIdBackupsPostRequestWithBody(se // per the OpenAPI spec (e.g. "color=blue,black,brown"). var rawQueryFragments []string - if params.ResponseFormat != nil { + if params.Severity != nil { - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "response-format", *params.ResponseFormat, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "string", Format: ""}); err != nil { + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "severity", *params.Severity, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "", Format: ""}); err != nil { return nil, err } else { for _, qp := range strings.Split(queryFrag, "&") { @@ -5851,14 +6503,156 @@ func NewClustersBackupsCreateApiV2ClustersClusterIdBackupsPostRequestWithBody(se } - if encoded := queryValues.Encode(); encoded != "" { - rawQueryFragments = append(rawQueryFragments, encoded) - } - queryURL.RawQuery = strings.Join(rawQueryFragments, "&") - } + if params.History != nil { - req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) - if err != nil { + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "history", *params.History, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } + } + + } + + if params.HistorySeconds != nil { + + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "history_seconds", *params.HistorySeconds, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "", Format: ""}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } + } + + } + + if params.Status != nil { + + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "status", *params.Status, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "", Format: ""}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } + } + + } + + if encoded := queryValues.Encode(); encoded != "" { + rawQueryFragments = append(rawQueryFragments, encoded) + } + queryURL.RawQuery = strings.Join(rawQueryFragments, "&") + } + + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + if err != nil { + return nil, err + } + + return req, nil +} + +// NewClustersBackupsListApiV2ClustersClusterIdBackupsGetRequest constructs an http.Request for the ClustersBackupsListApiV2ClustersClusterIdBackupsGet method +func NewClustersBackupsListApiV2ClustersClusterIdBackupsGetRequest(server string, clusterId openapi_types.UUID) (*http.Request, error) { + var err error + + var pathParam0 string + + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + serverURL, err := url.Parse(server) + if err != nil { + return nil, err + } + + operationPath := fmt.Sprintf("/api/v2/clusters/%s/backups/", pathParam0) + if operationPath[0] == '/' { + operationPath = "." + operationPath + } + + queryURL, err := serverURL.Parse(operationPath) + if err != nil { + return nil, err + } + + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + if err != nil { + return nil, err + } + + return req, nil +} + +// NewClustersBackupsCreateApiV2ClustersClusterIdBackupsPostRequest calls the generic ClustersBackupsCreateApiV2ClustersClusterIdBackupsPost builder with application/json body +func NewClustersBackupsCreateApiV2ClustersClusterIdBackupsPostRequest(server string, clusterId openapi_types.UUID, params *ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostParams, body ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostJSONRequestBody) (*http.Request, error) { + var bodyReader io.Reader + buf, err := json.Marshal(body) + if err != nil { + return nil, err + } + bodyReader = bytes.NewReader(buf) + return NewClustersBackupsCreateApiV2ClustersClusterIdBackupsPostRequestWithBody(server, clusterId, params, "application/json", bodyReader) +} + +// NewClustersBackupsCreateApiV2ClustersClusterIdBackupsPostRequestWithBody constructs an http.Request for the ClustersBackupsCreateApiV2ClustersClusterIdBackupsPost method, with any body, and a specified content type +func NewClustersBackupsCreateApiV2ClustersClusterIdBackupsPostRequestWithBody(server string, clusterId openapi_types.UUID, params *ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostParams, contentType string, body io.Reader) (*http.Request, error) { + var err error + + var pathParam0 string + + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + serverURL, err := url.Parse(server) + if err != nil { + return nil, err + } + + operationPath := fmt.Sprintf("/api/v2/clusters/%s/backups/", pathParam0) + if operationPath[0] == '/' { + operationPath = "." + operationPath + } + + queryURL, err := serverURL.Parse(operationPath) + if err != nil { + return nil, err + } + + if params != nil { + // queryValues collects non-styled parameters (passthrough, JSON) + // that are safe to round-trip through url.Values.Encode(). + queryValues := queryURL.Query() + // rawQueryFragments collects pre-encoded query fragments from + // styled parameters, preserving literal commas as delimiters + // per the OpenAPI spec (e.g. "color=blue,black,brown"). + var rawQueryFragments []string + + if params.ResponseFormat != nil { + + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "response-format", *params.ResponseFormat, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "string", Format: ""}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } + } + + } + + if encoded := queryValues.Encode(); encoded != "" { + rawQueryFragments = append(rawQueryFragments, encoded) + } + queryURL.RawQuery = strings.Join(rawQueryFragments, "&") + } + + req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) + if err != nil { return nil, err } @@ -6488,42 +7282,8 @@ func NewClustersCapacityApiV2ClustersClusterIdCapacityGetRequest(server string, return req, nil } -// NewClustersExpandApiV2ClustersClusterIdExpandPostRequest constructs an http.Request for the ClustersExpandApiV2ClustersClusterIdExpandPost method -func NewClustersExpandApiV2ClustersClusterIdExpandPostRequest(server string, clusterId openapi_types.UUID) (*http.Request, error) { - var err error - - var pathParam0 string - - pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err - } - - serverURL, err := url.Parse(server) - if err != nil { - return nil, err - } - - operationPath := fmt.Sprintf("/api/v2/clusters/%s/expand", pathParam0) - if operationPath[0] == '/' { - operationPath = "." + operationPath - } - - queryURL, err := serverURL.Parse(operationPath) - if err != nil { - return nil, err - } - - req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) - if err != nil { - return nil, err - } - - return req, nil -} - -// NewClustersIostatsApiV2ClustersClusterIdIostatsGetRequest constructs an http.Request for the ClustersIostatsApiV2ClustersClusterIdIostatsGet method -func NewClustersIostatsApiV2ClustersClusterIdIostatsGetRequest(server string, clusterId openapi_types.UUID, params *ClustersIostatsApiV2ClustersClusterIdIostatsGetParams) (*http.Request, error) { +// NewClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetRequest constructs an http.Request for the ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGet method +func NewClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetRequest(server string, clusterId openapi_types.UUID, params *ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetParams) (*http.Request, error) { var err error var pathParam0 string @@ -6538,7 +7298,7 @@ func NewClustersIostatsApiV2ClustersClusterIdIostatsGetRequest(server string, cl return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/iostats", pathParam0) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/consistency-groups/", pathParam0) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -6557,9 +7317,9 @@ func NewClustersIostatsApiV2ClustersClusterIdIostatsGetRequest(server string, cl // per the OpenAPI spec (e.g. "color=blue,black,brown"). var rawQueryFragments []string - if params.History != nil { + if params.Name != nil { - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "history", *params.History, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "", Format: ""}); err != nil { + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "name", *params.Name, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "", Format: ""}); err != nil { return nil, err } else { for _, qp := range strings.Split(queryFrag, "&") { @@ -6583,8 +7343,8 @@ func NewClustersIostatsApiV2ClustersClusterIdIostatsGetRequest(server string, cl return req, nil } -// NewClustersLogsApiV2ClustersClusterIdLogsGetRequest constructs an http.Request for the ClustersLogsApiV2ClustersClusterIdLogsGet method -func NewClustersLogsApiV2ClustersClusterIdLogsGetRequest(server string, clusterId openapi_types.UUID, params *ClustersLogsApiV2ClustersClusterIdLogsGetParams) (*http.Request, error) { +// NewClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGetRequest constructs an http.Request for the ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGet method +func NewClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGetRequest(server string, clusterId openapi_types.UUID, groupId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -6594,12 +7354,19 @@ func NewClustersLogsApiV2ClustersClusterIdLogsGetRequest(server string, clusterI return nil, err } + var pathParam1 string + + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "group_id", groupId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + serverURL, err := url.Parse(server) if err != nil { return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/logs", pathParam0) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/consistency-groups/%s/", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -6609,45 +7376,6 @@ func NewClustersLogsApiV2ClustersClusterIdLogsGetRequest(server string, clusterI return nil, err } - if params != nil { - // queryValues collects non-styled parameters (passthrough, JSON) - // that are safe to round-trip through url.Values.Encode(). - queryValues := queryURL.Query() - // rawQueryFragments collects pre-encoded query fragments from - // styled parameters, preserving literal commas as delimiters - // per the OpenAPI spec (e.g. "color=blue,black,brown"). - var rawQueryFragments []string - - if params.Limit != nil { - - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "limit", *params.Limit, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "integer", Format: ""}); err != nil { - return nil, err - } else { - for _, qp := range strings.Split(queryFrag, "&") { - rawQueryFragments = append(rawQueryFragments, qp) - } - } - - } - - if params.Watch != nil { - - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "watch", *params.Watch, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { - return nil, err - } else { - for _, qp := range strings.Split(queryFrag, "&") { - rawQueryFragments = append(rawQueryFragments, qp) - } - } - - } - - if encoded := queryValues.Encode(); encoded != "" { - rawQueryFragments = append(rawQueryFragments, encoded) - } - queryURL.RawQuery = strings.Join(rawQueryFragments, "&") - } - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err @@ -6656,8 +7384,8 @@ func NewClustersLogsApiV2ClustersClusterIdLogsGetRequest(server string, clusterI return req, nil } -// NewClustersRebalanceApiV2ClustersClusterIdRebalancePostRequest constructs an http.Request for the ClustersRebalanceApiV2ClustersClusterIdRebalancePost method -func NewClustersRebalanceApiV2ClustersClusterIdRebalancePostRequest(server string, clusterId openapi_types.UUID) (*http.Request, error) { +// NewClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGetRequest constructs an http.Request for the ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGet method +func NewClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGetRequest(server string, clusterId openapi_types.UUID, groupId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -6667,12 +7395,19 @@ func NewClustersRebalanceApiV2ClustersClusterIdRebalancePostRequest(server strin return nil, err } + var pathParam1 string + + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "group_id", groupId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + serverURL, err := url.Parse(server) if err != nil { return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/rebalance", pathParam0) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/consistency-groups/%s/members", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -6682,7 +7417,7 @@ func NewClustersRebalanceApiV2ClustersClusterIdRebalancePostRequest(server strin return nil, err } - req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err } @@ -6690,8 +7425,19 @@ func NewClustersRebalanceApiV2ClustersClusterIdRebalancePostRequest(server strin return req, nil } -// NewClustersReplicationPoliciesListApiV2ClustersClusterIdReplicationPoliciesGetRequest constructs an http.Request for the ClustersReplicationPoliciesListApiV2ClustersClusterIdReplicationPoliciesGet method -func NewClustersReplicationPoliciesListApiV2ClustersClusterIdReplicationPoliciesGetRequest(server string, clusterId openapi_types.UUID) (*http.Request, error) { +// NewClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostRequest calls the generic ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPost builder with application/json body +func NewClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostRequest(server string, clusterId openapi_types.UUID, groupId openapi_types.UUID, body ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostJSONRequestBody) (*http.Request, error) { + var bodyReader io.Reader + buf, err := json.Marshal(body) + if err != nil { + return nil, err + } + bodyReader = bytes.NewReader(buf) + return NewClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostRequestWithBody(server, clusterId, groupId, "application/json", bodyReader) +} + +// NewClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostRequestWithBody constructs an http.Request for the ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPost method, with any body, and a specified content type +func NewClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostRequestWithBody(server string, clusterId openapi_types.UUID, groupId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { var err error var pathParam0 string @@ -6701,13 +7447,20 @@ func NewClustersReplicationPoliciesListApiV2ClustersClusterIdReplicationPolicies return nil, err } - serverURL, err := url.Parse(server) + var pathParam1 string + + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "group_id", groupId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/replication/policies/", pathParam0) - if operationPath[0] == '/' { + serverURL, err := url.Parse(server) + if err != nil { + return nil, err + } + + operationPath := fmt.Sprintf("/api/v2/clusters/%s/consistency-groups/%s/members", pathParam0, pathParam1) + if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -6716,32 +7469,37 @@ func NewClustersReplicationPoliciesListApiV2ClustersClusterIdReplicationPolicies return nil, err } - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) if err != nil { return nil, err } + req.Header.Add("Content-Type", contentType) + return req, nil } -// NewClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostRequest calls the generic ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPost builder with application/json body -func NewClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostRequest(server string, clusterId openapi_types.UUID, params *ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostParams, body ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostJSONRequestBody) (*http.Request, error) { - var bodyReader io.Reader - buf, err := json.Marshal(body) +// NewClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDeleteRequest constructs an http.Request for the ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDelete method +func NewClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDeleteRequest(server string, clusterId openapi_types.UUID, groupId openapi_types.UUID, lvolId string) (*http.Request, error) { + var err error + + var pathParam0 string + + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } - bodyReader = bytes.NewReader(buf) - return NewClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostRequestWithBody(server, clusterId, params, "application/json", bodyReader) -} -// NewClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostRequestWithBody constructs an http.Request for the ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPost method, with any body, and a specified content type -func NewClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostRequestWithBody(server string, clusterId openapi_types.UUID, params *ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostParams, contentType string, body io.Reader) (*http.Request, error) { - var err error + var pathParam1 string - var pathParam0 string + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "group_id", groupId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } - pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + var pathParam2 string + + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "lvol_id", lvolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: ""}) if err != nil { return nil, err } @@ -6751,7 +7509,7 @@ func NewClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPolici return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/replication/policies/", pathParam0) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/consistency-groups/%s/members/%s", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -6761,45 +7519,16 @@ func NewClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPolici return nil, err } - if params != nil { - // queryValues collects non-styled parameters (passthrough, JSON) - // that are safe to round-trip through url.Values.Encode(). - queryValues := queryURL.Query() - // rawQueryFragments collects pre-encoded query fragments from - // styled parameters, preserving literal commas as delimiters - // per the OpenAPI spec (e.g. "color=blue,black,brown"). - var rawQueryFragments []string - - if params.ResponseFormat != nil { - - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "response-format", *params.ResponseFormat, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "string", Format: ""}); err != nil { - return nil, err - } else { - for _, qp := range strings.Split(queryFrag, "&") { - rawQueryFragments = append(rawQueryFragments, qp) - } - } - - } - - if encoded := queryValues.Encode(); encoded != "" { - rawQueryFragments = append(rawQueryFragments, encoded) - } - queryURL.RawQuery = strings.Join(rawQueryFragments, "&") - } - - req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) + req, err := http.NewRequest(http.MethodDelete, queryURL.String(), nil) if err != nil { return nil, err } - req.Header.Add("Content-Type", contentType) - return req, nil } -// NewClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDeleteRequest constructs an http.Request for the ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDelete method -func NewClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDeleteRequest(server string, clusterId openapi_types.UUID, policyId openapi_types.UUID) (*http.Request, error) { +// NewClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetRequest constructs an http.Request for the ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGet method +func NewClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetRequest(server string, clusterId openapi_types.UUID, groupId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -6811,7 +7540,7 @@ func NewClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPolici var pathParam1 string - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "policy_id", policyId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "group_id", groupId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -6821,7 +7550,7 @@ func NewClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPolici return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/replication/policies/%s/", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/consistency-groups/%s/snapshots", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -6831,7 +7560,7 @@ func NewClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPolici return nil, err } - req, err := http.NewRequest(http.MethodDelete, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err } @@ -6839,8 +7568,8 @@ func NewClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPolici return req, nil } -// NewClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGetRequest constructs an http.Request for the ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGet method -func NewClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGetRequest(server string, clusterId openapi_types.UUID, policyId openapi_types.UUID) (*http.Request, error) { +// NewClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPostRequest constructs an http.Request for the ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPost method +func NewClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPostRequest(server string, clusterId openapi_types.UUID, groupId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -6852,7 +7581,7 @@ func NewClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPolici var pathParam1 string - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "policy_id", policyId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "group_id", groupId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -6862,7 +7591,7 @@ func NewClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPolici return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/replication/policies/%s/", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/consistency-groups/%s/snapshots", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -6872,7 +7601,7 @@ func NewClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPolici return nil, err } - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) if err != nil { return nil, err } @@ -6880,8 +7609,8 @@ func NewClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPolici return req, nil } -// NewClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPostRequest constructs an http.Request for the ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPost method -func NewClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPostRequest(server string, clusterId openapi_types.UUID, policyId openapi_types.UUID) (*http.Request, error) { +// NewClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDeleteRequest constructs an http.Request for the ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDelete method +func NewClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDeleteRequest(server string, clusterId openapi_types.UUID, groupId openapi_types.UUID, seq int) (*http.Request, error) { var err error var pathParam0 string @@ -6893,7 +7622,14 @@ func NewClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoli var pathParam1 string - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "policy_id", policyId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "group_id", groupId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + var pathParam2 string + + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "seq", seq, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "integer", Format: ""}) if err != nil { return nil, err } @@ -6903,7 +7639,7 @@ func NewClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoli return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/replication/policies/%s/failover", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/consistency-groups/%s/snapshots/%s", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -6913,7 +7649,7 @@ func NewClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoli return nil, err } - req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodDelete, queryURL.String(), nil) if err != nil { return nil, err } @@ -6921,8 +7657,8 @@ func NewClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoli return req, nil } -// NewClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetRequest constructs an http.Request for the ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGet method -func NewClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetRequest(server string, clusterId openapi_types.UUID, lvolId openapi_types.UUID) (*http.Request, error) { +// NewClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGetRequest constructs an http.Request for the ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGet method +func NewClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGetRequest(server string, clusterId openapi_types.UUID, groupId openapi_types.UUID, seq int) (*http.Request, error) { var err error var pathParam0 string @@ -6934,7 +7670,14 @@ func NewClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationR var pathParam1 string - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "lvol_id", lvolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "group_id", groupId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + var pathParam2 string + + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "seq", seq, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "integer", Format: ""}) if err != nil { return nil, err } @@ -6944,7 +7687,7 @@ func NewClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationR return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/replication/relationships/%s", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/consistency-groups/%s/snapshots/%s", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -6962,8 +7705,8 @@ func NewClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationR return req, nil } -// NewClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGetRequest constructs an http.Request for the ClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGet method -func NewClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGetRequest(server string, clusterId openapi_types.UUID) (*http.Request, error) { +// NewClustersExpandApiV2ClustersClusterIdExpandPostRequest constructs an http.Request for the ClustersExpandApiV2ClustersClusterIdExpandPost method +func NewClustersExpandApiV2ClustersClusterIdExpandPostRequest(server string, clusterId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -6978,7 +7721,7 @@ func NewClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGe return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/replication/targets/", pathParam0) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/expand", pathParam0) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -6988,7 +7731,7 @@ func NewClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGe return nil, err } - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) if err != nil { return nil, err } @@ -6996,19 +7739,8 @@ func NewClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGe return req, nil } -// NewClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostRequest calls the generic ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPost builder with application/json body -func NewClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostRequest(server string, clusterId openapi_types.UUID, params *ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostParams, body ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostJSONRequestBody) (*http.Request, error) { - var bodyReader io.Reader - buf, err := json.Marshal(body) - if err != nil { - return nil, err - } - bodyReader = bytes.NewReader(buf) - return NewClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostRequestWithBody(server, clusterId, params, "application/json", bodyReader) -} - -// NewClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostRequestWithBody constructs an http.Request for the ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPost method, with any body, and a specified content type -func NewClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostRequestWithBody(server string, clusterId openapi_types.UUID, params *ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostParams, contentType string, body io.Reader) (*http.Request, error) { +// NewClustersIostatsApiV2ClustersClusterIdIostatsGetRequest constructs an http.Request for the ClustersIostatsApiV2ClustersClusterIdIostatsGet method +func NewClustersIostatsApiV2ClustersClusterIdIostatsGetRequest(server string, clusterId openapi_types.UUID, params *ClustersIostatsApiV2ClustersClusterIdIostatsGetParams) (*http.Request, error) { var err error var pathParam0 string @@ -7023,7 +7755,7 @@ func NewClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargets return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/replication/targets/", pathParam0) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/iostats", pathParam0) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -7042,9 +7774,9 @@ func NewClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargets // per the OpenAPI spec (e.g. "color=blue,black,brown"). var rawQueryFragments []string - if params.ResponseFormat != nil { + if params.History != nil { - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "response-format", *params.ResponseFormat, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "string", Format: ""}); err != nil { + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "history", *params.History, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "", Format: ""}); err != nil { return nil, err } else { for _, qp := range strings.Split(queryFrag, "&") { @@ -7060,18 +7792,16 @@ func NewClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargets queryURL.RawQuery = strings.Join(rawQueryFragments, "&") } - req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err } - req.Header.Add("Content-Type", contentType) - return req, nil } -// NewClustersReplicationTargetsDeleteApiV2ClustersClusterIdReplicationTargetsTargetIdDeleteRequest constructs an http.Request for the ClustersReplicationTargetsDeleteApiV2ClustersClusterIdReplicationTargetsTargetIdDelete method -func NewClustersReplicationTargetsDeleteApiV2ClustersClusterIdReplicationTargetsTargetIdDeleteRequest(server string, clusterId openapi_types.UUID, targetId openapi_types.UUID) (*http.Request, error) { +// NewClustersLogsApiV2ClustersClusterIdLogsGetRequest constructs an http.Request for the ClustersLogsApiV2ClustersClusterIdLogsGet method +func NewClustersLogsApiV2ClustersClusterIdLogsGetRequest(server string, clusterId openapi_types.UUID, params *ClustersLogsApiV2ClustersClusterIdLogsGetParams) (*http.Request, error) { var err error var pathParam0 string @@ -7081,19 +7811,12 @@ func NewClustersReplicationTargetsDeleteApiV2ClustersClusterIdReplicationTargets return nil, err } - var pathParam1 string - - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "target_id", targetId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err - } - serverURL, err := url.Parse(server) if err != nil { return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/replication/targets/%s/", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/logs", pathParam0) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -7103,45 +7826,43 @@ func NewClustersReplicationTargetsDeleteApiV2ClustersClusterIdReplicationTargets return nil, err } - req, err := http.NewRequest(http.MethodDelete, queryURL.String(), nil) - if err != nil { - return nil, err - } - - return req, nil -} - -// NewClustersReplicationTargetsDetailApiV2ClustersClusterIdReplicationTargetsTargetIdGetRequest constructs an http.Request for the ClustersReplicationTargetsDetailApiV2ClustersClusterIdReplicationTargetsTargetIdGet method -func NewClustersReplicationTargetsDetailApiV2ClustersClusterIdReplicationTargetsTargetIdGetRequest(server string, clusterId openapi_types.UUID, targetId openapi_types.UUID) (*http.Request, error) { - var err error + if params != nil { + // queryValues collects non-styled parameters (passthrough, JSON) + // that are safe to round-trip through url.Values.Encode(). + queryValues := queryURL.Query() + // rawQueryFragments collects pre-encoded query fragments from + // styled parameters, preserving literal commas as delimiters + // per the OpenAPI spec (e.g. "color=blue,black,brown"). + var rawQueryFragments []string - var pathParam0 string + if params.Limit != nil { - pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err - } + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "limit", *params.Limit, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "integer", Format: ""}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } + } - var pathParam1 string + } - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "target_id", targetId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err - } + if params.Watch != nil { - serverURL, err := url.Parse(server) - if err != nil { - return nil, err - } + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "watch", *params.Watch, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } + } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/replication/targets/%s/", pathParam0, pathParam1) - if operationPath[0] == '/' { - operationPath = "." + operationPath - } + } - queryURL, err := serverURL.Parse(operationPath) - if err != nil { - return nil, err + if encoded := queryValues.Encode(); encoded != "" { + rawQueryFragments = append(rawQueryFragments, encoded) + } + queryURL.RawQuery = strings.Join(rawQueryFragments, "&") } req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) @@ -7152,8 +7873,8 @@ func NewClustersReplicationTargetsDetailApiV2ClustersClusterIdReplicationTargets return req, nil } -// NewClustersReplicationTargetsFailoverApiV2ClustersClusterIdReplicationTargetsTargetIdFailoverPostRequest constructs an http.Request for the ClustersReplicationTargetsFailoverApiV2ClustersClusterIdReplicationTargetsTargetIdFailoverPost method -func NewClustersReplicationTargetsFailoverApiV2ClustersClusterIdReplicationTargetsTargetIdFailoverPostRequest(server string, clusterId openapi_types.UUID, targetId openapi_types.UUID) (*http.Request, error) { +// NewClustersRebalanceApiV2ClustersClusterIdRebalancePostRequest constructs an http.Request for the ClustersRebalanceApiV2ClustersClusterIdRebalancePost method +func NewClustersRebalanceApiV2ClustersClusterIdRebalancePostRequest(server string, clusterId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -7163,19 +7884,12 @@ func NewClustersReplicationTargetsFailoverApiV2ClustersClusterIdReplicationTarge return nil, err } - var pathParam1 string - - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "target_id", targetId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err - } - serverURL, err := url.Parse(server) if err != nil { return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/replication/targets/%s/failover", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/rebalance", pathParam0) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -7193,8 +7907,8 @@ func NewClustersReplicationTargetsFailoverApiV2ClustersClusterIdReplicationTarge return req, nil } -// NewClustersShutdownApiV2ClustersClusterIdShutdownPostRequest constructs an http.Request for the ClustersShutdownApiV2ClustersClusterIdShutdownPost method -func NewClustersShutdownApiV2ClustersClusterIdShutdownPostRequest(server string, clusterId openapi_types.UUID) (*http.Request, error) { +// NewClustersReplicationPoliciesListApiV2ClustersClusterIdReplicationPoliciesGetRequest constructs an http.Request for the ClustersReplicationPoliciesListApiV2ClustersClusterIdReplicationPoliciesGet method +func NewClustersReplicationPoliciesListApiV2ClustersClusterIdReplicationPoliciesGetRequest(server string, clusterId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -7209,7 +7923,7 @@ func NewClustersShutdownApiV2ClustersClusterIdShutdownPostRequest(server string, return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/shutdown", pathParam0) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/replication/policies/", pathParam0) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -7219,7 +7933,7 @@ func NewClustersShutdownApiV2ClustersClusterIdShutdownPostRequest(server string, return nil, err } - req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err } @@ -7227,8 +7941,19 @@ func NewClustersShutdownApiV2ClustersClusterIdShutdownPostRequest(server string, return req, nil } -// NewClustersStartApiV2ClustersClusterIdStartPostRequest constructs an http.Request for the ClustersStartApiV2ClustersClusterIdStartPost method -func NewClustersStartApiV2ClustersClusterIdStartPostRequest(server string, clusterId openapi_types.UUID) (*http.Request, error) { +// NewClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostRequest calls the generic ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPost builder with application/json body +func NewClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostRequest(server string, clusterId openapi_types.UUID, params *ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostParams, body ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostJSONRequestBody) (*http.Request, error) { + var bodyReader io.Reader + buf, err := json.Marshal(body) + if err != nil { + return nil, err + } + bodyReader = bytes.NewReader(buf) + return NewClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostRequestWithBody(server, clusterId, params, "application/json", bodyReader) +} + +// NewClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostRequestWithBody constructs an http.Request for the ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPost method, with any body, and a specified content type +func NewClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostRequestWithBody(server string, clusterId openapi_types.UUID, params *ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostParams, contentType string, body io.Reader) (*http.Request, error) { var err error var pathParam0 string @@ -7243,41 +7968,7 @@ func NewClustersStartApiV2ClustersClusterIdStartPostRequest(server string, clust return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/start", pathParam0) - if operationPath[0] == '/' { - operationPath = "." + operationPath - } - - queryURL, err := serverURL.Parse(operationPath) - if err != nil { - return nil, err - } - - req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) - if err != nil { - return nil, err - } - - return req, nil -} - -// NewClustersStorageNodesListApiV2ClustersClusterIdStorageNodesGetRequest constructs an http.Request for the ClustersStorageNodesListApiV2ClustersClusterIdStorageNodesGet method -func NewClustersStorageNodesListApiV2ClustersClusterIdStorageNodesGetRequest(server string, clusterId openapi_types.UUID, params *ClustersStorageNodesListApiV2ClustersClusterIdStorageNodesGetParams) (*http.Request, error) { - var err error - - var pathParam0 string - - pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err - } - - serverURL, err := url.Parse(server) - if err != nil { - return nil, err - } - - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/", pathParam0) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/replication/policies/", pathParam0) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -7296,9 +7987,9 @@ func NewClustersStorageNodesListApiV2ClustersClusterIdStorageNodesGetRequest(ser // per the OpenAPI spec (e.g. "color=blue,black,brown"). var rawQueryFragments []string - if params.Watch != nil { + if params.ResponseFormat != nil { - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "watch", *params.Watch, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "response-format", *params.ResponseFormat, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "string", Format: ""}); err != nil { return nil, err } else { for _, qp := range strings.Split(queryFrag, "&") { @@ -7314,27 +8005,18 @@ func NewClustersStorageNodesListApiV2ClustersClusterIdStorageNodesGetRequest(ser queryURL.RawQuery = strings.Join(rawQueryFragments, "&") } - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) if err != nil { return nil, err } - return req, nil -} + req.Header.Add("Content-Type", contentType) -// NewClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostRequest calls the generic ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPost builder with application/json body -func NewClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostRequest(server string, clusterId openapi_types.UUID, params *ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostParams, body ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostJSONRequestBody) (*http.Request, error) { - var bodyReader io.Reader - buf, err := json.Marshal(body) - if err != nil { - return nil, err - } - bodyReader = bytes.NewReader(buf) - return NewClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostRequestWithBody(server, clusterId, params, "application/json", bodyReader) + return req, nil } -// NewClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostRequestWithBody constructs an http.Request for the ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPost method, with any body, and a specified content type -func NewClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostRequestWithBody(server string, clusterId openapi_types.UUID, params *ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostParams, contentType string, body io.Reader) (*http.Request, error) { +// NewClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDeleteRequest constructs an http.Request for the ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDelete method +func NewClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDeleteRequest(server string, clusterId openapi_types.UUID, policyId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -7344,12 +8026,19 @@ func NewClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostRequestW return nil, err } + var pathParam1 string + + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "policy_id", policyId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + serverURL, err := url.Parse(server) if err != nil { return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/", pathParam0) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/replication/policies/%s/", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -7359,45 +8048,16 @@ func NewClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostRequestW return nil, err } - if params != nil { - // queryValues collects non-styled parameters (passthrough, JSON) - // that are safe to round-trip through url.Values.Encode(). - queryValues := queryURL.Query() - // rawQueryFragments collects pre-encoded query fragments from - // styled parameters, preserving literal commas as delimiters - // per the OpenAPI spec (e.g. "color=blue,black,brown"). - var rawQueryFragments []string - - if params.ResponseFormat != nil { - - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "response-format", *params.ResponseFormat, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "string", Format: ""}); err != nil { - return nil, err - } else { - for _, qp := range strings.Split(queryFrag, "&") { - rawQueryFragments = append(rawQueryFragments, qp) - } - } - - } - - if encoded := queryValues.Encode(); encoded != "" { - rawQueryFragments = append(rawQueryFragments, encoded) - } - queryURL.RawQuery = strings.Join(rawQueryFragments, "&") - } - - req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) + req, err := http.NewRequest(http.MethodDelete, queryURL.String(), nil) if err != nil { return nil, err } - req.Header.Add("Content-Type", contentType) - return req, nil } -// NewClustersStorageNodesDeleteApiV2ClustersClusterIdStorageNodesStorageNodeIdDeleteRequest constructs an http.Request for the ClustersStorageNodesDeleteApiV2ClustersClusterIdStorageNodesStorageNodeIdDelete method -func NewClustersStorageNodesDeleteApiV2ClustersClusterIdStorageNodesStorageNodeIdDeleteRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, params *ClustersStorageNodesDeleteApiV2ClustersClusterIdStorageNodesStorageNodeIdDeleteParams) (*http.Request, error) { +// NewClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGetRequest constructs an http.Request for the ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGet method +func NewClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGetRequest(server string, clusterId openapi_types.UUID, policyId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -7409,7 +8069,7 @@ func NewClustersStorageNodesDeleteApiV2ClustersClusterIdStorageNodesStorageNodeI var pathParam1 string - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "storage_node_id", storageNodeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "policy_id", policyId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -7419,7 +8079,7 @@ func NewClustersStorageNodesDeleteApiV2ClustersClusterIdStorageNodesStorageNodeI return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/replication/policies/%s/", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -7429,58 +8089,7 @@ func NewClustersStorageNodesDeleteApiV2ClustersClusterIdStorageNodesStorageNodeI return nil, err } - if params != nil { - // queryValues collects non-styled parameters (passthrough, JSON) - // that are safe to round-trip through url.Values.Encode(). - queryValues := queryURL.Query() - // rawQueryFragments collects pre-encoded query fragments from - // styled parameters, preserving literal commas as delimiters - // per the OpenAPI spec (e.g. "color=blue,black,brown"). - var rawQueryFragments []string - - if params.ForceRemove != nil { - - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "force_remove", *params.ForceRemove, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { - return nil, err - } else { - for _, qp := range strings.Split(queryFrag, "&") { - rawQueryFragments = append(rawQueryFragments, qp) - } - } - - } - - if params.ForceMigrate != nil { - - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "force_migrate", *params.ForceMigrate, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { - return nil, err - } else { - for _, qp := range strings.Split(queryFrag, "&") { - rawQueryFragments = append(rawQueryFragments, qp) - } - } - - } - - if params.ForceDelete != nil { - - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "force_delete", *params.ForceDelete, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { - return nil, err - } else { - for _, qp := range strings.Split(queryFrag, "&") { - rawQueryFragments = append(rawQueryFragments, qp) - } - } - - } - - if encoded := queryValues.Encode(); encoded != "" { - rawQueryFragments = append(rawQueryFragments, encoded) - } - queryURL.RawQuery = strings.Join(rawQueryFragments, "&") - } - - req, err := http.NewRequest(http.MethodDelete, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err } @@ -7488,8 +8097,8 @@ func NewClustersStorageNodesDeleteApiV2ClustersClusterIdStorageNodesStorageNodeI return req, nil } -// NewClustersStorageNodesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdGetRequest constructs an http.Request for the ClustersStorageNodesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdGet method -func NewClustersStorageNodesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdGetRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, params *ClustersStorageNodesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdGetParams) (*http.Request, error) { +// NewClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPostRequest constructs an http.Request for the ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPost method +func NewClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPostRequest(server string, clusterId openapi_types.UUID, policyId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -7501,7 +8110,7 @@ func NewClustersStorageNodesDetailApiV2ClustersClusterIdStorageNodesStorageNodeI var pathParam1 string - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "storage_node_id", storageNodeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "policy_id", policyId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -7511,7 +8120,7 @@ func NewClustersStorageNodesDetailApiV2ClustersClusterIdStorageNodesStorageNodeI return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/replication/policies/%s/failover", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -7521,34 +8130,7 @@ func NewClustersStorageNodesDetailApiV2ClustersClusterIdStorageNodesStorageNodeI return nil, err } - if params != nil { - // queryValues collects non-styled parameters (passthrough, JSON) - // that are safe to round-trip through url.Values.Encode(). - queryValues := queryURL.Query() - // rawQueryFragments collects pre-encoded query fragments from - // styled parameters, preserving literal commas as delimiters - // per the OpenAPI spec (e.g. "color=blue,black,brown"). - var rawQueryFragments []string - - if params.Watch != nil { - - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "watch", *params.Watch, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { - return nil, err - } else { - for _, qp := range strings.Split(queryFrag, "&") { - rawQueryFragments = append(rawQueryFragments, qp) - } - } - - } - - if encoded := queryValues.Encode(); encoded != "" { - rawQueryFragments = append(rawQueryFragments, encoded) - } - queryURL.RawQuery = strings.Join(rawQueryFragments, "&") - } - - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) if err != nil { return nil, err } @@ -7556,8 +8138,8 @@ func NewClustersStorageNodesDetailApiV2ClustersClusterIdStorageNodesStorageNodeI return req, nil } -// NewClustersStorageNodesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdCapacityGetRequest constructs an http.Request for the ClustersStorageNodesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdCapacityGet method -func NewClustersStorageNodesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdCapacityGetRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, params *ClustersStorageNodesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdCapacityGetParams) (*http.Request, error) { +// NewClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGetRequest constructs an http.Request for the ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGet method +func NewClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGetRequest(server string, clusterId openapi_types.UUID, policyId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -7569,7 +8151,7 @@ func NewClustersStorageNodesCapacityApiV2ClustersClusterIdStorageNodesStorageNod var pathParam1 string - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "storage_node_id", storageNodeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "policy_id", policyId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -7579,7 +8161,7 @@ func NewClustersStorageNodesCapacityApiV2ClustersClusterIdStorageNodesStorageNod return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/capacity", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/replication/policies/%s/latest-generation", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -7589,33 +8171,6 @@ func NewClustersStorageNodesCapacityApiV2ClustersClusterIdStorageNodesStorageNod return nil, err } - if params != nil { - // queryValues collects non-styled parameters (passthrough, JSON) - // that are safe to round-trip through url.Values.Encode(). - queryValues := queryURL.Query() - // rawQueryFragments collects pre-encoded query fragments from - // styled parameters, preserving literal commas as delimiters - // per the OpenAPI spec (e.g. "color=blue,black,brown"). - var rawQueryFragments []string - - if params.History != nil { - - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "history", *params.History, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "", Format: ""}); err != nil { - return nil, err - } else { - for _, qp := range strings.Split(queryFrag, "&") { - rawQueryFragments = append(rawQueryFragments, qp) - } - } - - } - - if encoded := queryValues.Encode(); encoded != "" { - rawQueryFragments = append(rawQueryFragments, encoded) - } - queryURL.RawQuery = strings.Join(rawQueryFragments, "&") - } - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err @@ -7624,8 +8179,8 @@ func NewClustersStorageNodesCapacityApiV2ClustersClusterIdStorageNodesStorageNod return req, nil } -// NewClustersStorageNodesDevicesListApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesGetRequest constructs an http.Request for the ClustersStorageNodesDevicesListApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesGet method -func NewClustersStorageNodesDevicesListApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesGetRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, params *ClustersStorageNodesDevicesListApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesGetParams) (*http.Request, error) { +// NewClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetRequest constructs an http.Request for the ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGet method +func NewClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetRequest(server string, clusterId openapi_types.UUID, lvolId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -7637,7 +8192,7 @@ func NewClustersStorageNodesDevicesListApiV2ClustersClusterIdStorageNodesStorage var pathParam1 string - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "storage_node_id", storageNodeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "lvol_id", lvolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -7647,7 +8202,7 @@ func NewClustersStorageNodesDevicesListApiV2ClustersClusterIdStorageNodesStorage return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/devices/", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/replication/relationships/%s", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -7657,33 +8212,6 @@ func NewClustersStorageNodesDevicesListApiV2ClustersClusterIdStorageNodesStorage return nil, err } - if params != nil { - // queryValues collects non-styled parameters (passthrough, JSON) - // that are safe to round-trip through url.Values.Encode(). - queryValues := queryURL.Query() - // rawQueryFragments collects pre-encoded query fragments from - // styled parameters, preserving literal commas as delimiters - // per the OpenAPI spec (e.g. "color=blue,black,brown"). - var rawQueryFragments []string - - if params.Watch != nil { - - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "watch", *params.Watch, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { - return nil, err - } else { - for _, qp := range strings.Split(queryFrag, "&") { - rawQueryFragments = append(rawQueryFragments, qp) - } - } - - } - - if encoded := queryValues.Encode(); encoded != "" { - rawQueryFragments = append(rawQueryFragments, encoded) - } - queryURL.RawQuery = strings.Join(rawQueryFragments, "&") - } - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err @@ -7692,8 +8220,8 @@ func NewClustersStorageNodesDevicesListApiV2ClustersClusterIdStorageNodesStorage return req, nil } -// NewClustersStorageNodesDevicesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdGetRequest constructs an http.Request for the ClustersStorageNodesDevicesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdGet method -func NewClustersStorageNodesDevicesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdGetRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, deviceId openapi_types.UUID, params *ClustersStorageNodesDevicesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdGetParams) (*http.Request, error) { +// NewClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGetRequest constructs an http.Request for the ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGet method +func NewClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGetRequest(server string, clusterId openapi_types.UUID, lvolId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -7705,14 +8233,7 @@ func NewClustersStorageNodesDevicesDetailApiV2ClustersClusterIdStorageNodesStora var pathParam1 string - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "storage_node_id", storageNodeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err - } - - var pathParam2 string - - pathParam2, err = runtime.StyleParamWithOptions("simple", false, "device_id", deviceId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "lvol_id", lvolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -7722,7 +8243,7 @@ func NewClustersStorageNodesDevicesDetailApiV2ClustersClusterIdStorageNodesStora return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/devices/%s/", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/replication/relationships/%s/latest-snapshot", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -7732,33 +8253,6 @@ func NewClustersStorageNodesDevicesDetailApiV2ClustersClusterIdStorageNodesStora return nil, err } - if params != nil { - // queryValues collects non-styled parameters (passthrough, JSON) - // that are safe to round-trip through url.Values.Encode(). - queryValues := queryURL.Query() - // rawQueryFragments collects pre-encoded query fragments from - // styled parameters, preserving literal commas as delimiters - // per the OpenAPI spec (e.g. "color=blue,black,brown"). - var rawQueryFragments []string - - if params.Watch != nil { - - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "watch", *params.Watch, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { - return nil, err - } else { - for _, qp := range strings.Split(queryFrag, "&") { - rawQueryFragments = append(rawQueryFragments, qp) - } - } - - } - - if encoded := queryValues.Encode(); encoded != "" { - rawQueryFragments = append(rawQueryFragments, encoded) - } - queryURL.RawQuery = strings.Join(rawQueryFragments, "&") - } - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err @@ -7767,8 +8261,8 @@ func NewClustersStorageNodesDevicesDetailApiV2ClustersClusterIdStorageNodesStora return req, nil } -// NewClustersStorageNodesDevicesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdCapacityGetRequest constructs an http.Request for the ClustersStorageNodesDevicesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdCapacityGet method -func NewClustersStorageNodesDevicesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdCapacityGetRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, deviceId openapi_types.UUID, params *ClustersStorageNodesDevicesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdCapacityGetParams) (*http.Request, error) { +// NewClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGetRequest constructs an http.Request for the ClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGet method +func NewClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGetRequest(server string, clusterId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -7778,16 +8272,47 @@ func NewClustersStorageNodesDevicesCapacityApiV2ClustersClusterIdStorageNodesSto return nil, err } - var pathParam1 string + serverURL, err := url.Parse(server) + if err != nil { + return nil, err + } - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "storage_node_id", storageNodeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/replication/targets/", pathParam0) + if operationPath[0] == '/' { + operationPath = "." + operationPath + } + + queryURL, err := serverURL.Parse(operationPath) if err != nil { return nil, err } - var pathParam2 string + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + if err != nil { + return nil, err + } - pathParam2, err = runtime.StyleParamWithOptions("simple", false, "device_id", deviceId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + return req, nil +} + +// NewClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostRequest calls the generic ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPost builder with application/json body +func NewClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostRequest(server string, clusterId openapi_types.UUID, params *ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostParams, body ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostJSONRequestBody) (*http.Request, error) { + var bodyReader io.Reader + buf, err := json.Marshal(body) + if err != nil { + return nil, err + } + bodyReader = bytes.NewReader(buf) + return NewClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostRequestWithBody(server, clusterId, params, "application/json", bodyReader) +} + +// NewClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostRequestWithBody constructs an http.Request for the ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPost method, with any body, and a specified content type +func NewClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostRequestWithBody(server string, clusterId openapi_types.UUID, params *ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostParams, contentType string, body io.Reader) (*http.Request, error) { + var err error + + var pathParam0 string + + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -7797,7 +8322,7 @@ func NewClustersStorageNodesDevicesCapacityApiV2ClustersClusterIdStorageNodesSto return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/devices/%s/capacity", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/replication/targets/", pathParam0) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -7816,9 +8341,9 @@ func NewClustersStorageNodesDevicesCapacityApiV2ClustersClusterIdStorageNodesSto // per the OpenAPI spec (e.g. "color=blue,black,brown"). var rawQueryFragments []string - if params.History != nil { + if params.ResponseFormat != nil { - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "history", *params.History, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "", Format: ""}); err != nil { + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "response-format", *params.ResponseFormat, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "string", Format: ""}); err != nil { return nil, err } else { for _, qp := range strings.Split(queryFrag, "&") { @@ -7834,16 +8359,18 @@ func NewClustersStorageNodesDevicesCapacityApiV2ClustersClusterIdStorageNodesSto queryURL.RawQuery = strings.Join(rawQueryFragments, "&") } - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) if err != nil { return nil, err } + req.Header.Add("Content-Type", contentType) + return req, nil } -// NewClustersStorageNodesDevicesGetDeviceHealthInfoApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdHealthInfoGetRequest constructs an http.Request for the ClustersStorageNodesDevicesGetDeviceHealthInfoApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdHealthInfoGet method -func NewClustersStorageNodesDevicesGetDeviceHealthInfoApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdHealthInfoGetRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, deviceId openapi_types.UUID) (*http.Request, error) { +// NewClustersReplicationTargetsDeleteApiV2ClustersClusterIdReplicationTargetsTargetIdDeleteRequest constructs an http.Request for the ClustersReplicationTargetsDeleteApiV2ClustersClusterIdReplicationTargetsTargetIdDelete method +func NewClustersReplicationTargetsDeleteApiV2ClustersClusterIdReplicationTargetsTargetIdDeleteRequest(server string, clusterId openapi_types.UUID, targetId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -7855,14 +8382,7 @@ func NewClustersStorageNodesDevicesGetDeviceHealthInfoApiV2ClustersClusterIdStor var pathParam1 string - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "storage_node_id", storageNodeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err - } - - var pathParam2 string - - pathParam2, err = runtime.StyleParamWithOptions("simple", false, "device_id", deviceId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "target_id", targetId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -7872,7 +8392,7 @@ func NewClustersStorageNodesDevicesGetDeviceHealthInfoApiV2ClustersClusterIdStor return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/devices/%s/health-info", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/replication/targets/%s/", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -7882,7 +8402,7 @@ func NewClustersStorageNodesDevicesGetDeviceHealthInfoApiV2ClustersClusterIdStor return nil, err } - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodDelete, queryURL.String(), nil) if err != nil { return nil, err } @@ -7890,8 +8410,8 @@ func NewClustersStorageNodesDevicesGetDeviceHealthInfoApiV2ClustersClusterIdStor return req, nil } -// NewClustersStorageNodesDevicesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdIostatsGetRequest constructs an http.Request for the ClustersStorageNodesDevicesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdIostatsGet method -func NewClustersStorageNodesDevicesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdIostatsGetRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, deviceId openapi_types.UUID, params *ClustersStorageNodesDevicesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdIostatsGetParams) (*http.Request, error) { +// NewClustersReplicationTargetsDetailApiV2ClustersClusterIdReplicationTargetsTargetIdGetRequest constructs an http.Request for the ClustersReplicationTargetsDetailApiV2ClustersClusterIdReplicationTargetsTargetIdGet method +func NewClustersReplicationTargetsDetailApiV2ClustersClusterIdReplicationTargetsTargetIdGetRequest(server string, clusterId openapi_types.UUID, targetId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -7903,14 +8423,7 @@ func NewClustersStorageNodesDevicesIostatsApiV2ClustersClusterIdStorageNodesStor var pathParam1 string - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "storage_node_id", storageNodeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err - } - - var pathParam2 string - - pathParam2, err = runtime.StyleParamWithOptions("simple", false, "device_id", deviceId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "target_id", targetId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -7920,7 +8433,7 @@ func NewClustersStorageNodesDevicesIostatsApiV2ClustersClusterIdStorageNodesStor return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/devices/%s/iostats", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/replication/targets/%s/", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -7930,33 +8443,6 @@ func NewClustersStorageNodesDevicesIostatsApiV2ClustersClusterIdStorageNodesStor return nil, err } - if params != nil { - // queryValues collects non-styled parameters (passthrough, JSON) - // that are safe to round-trip through url.Values.Encode(). - queryValues := queryURL.Query() - // rawQueryFragments collects pre-encoded query fragments from - // styled parameters, preserving literal commas as delimiters - // per the OpenAPI spec (e.g. "color=blue,black,brown"). - var rawQueryFragments []string - - if params.History != nil { - - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "history", *params.History, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "", Format: ""}); err != nil { - return nil, err - } else { - for _, qp := range strings.Split(queryFrag, "&") { - rawQueryFragments = append(rawQueryFragments, qp) - } - } - - } - - if encoded := queryValues.Encode(); encoded != "" { - rawQueryFragments = append(rawQueryFragments, encoded) - } - queryURL.RawQuery = strings.Join(rawQueryFragments, "&") - } - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err @@ -7965,8 +8451,8 @@ func NewClustersStorageNodesDevicesIostatsApiV2ClustersClusterIdStorageNodesStor return req, nil } -// NewClustersStorageNodesDevicesRemoveApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRemovePostRequest constructs an http.Request for the ClustersStorageNodesDevicesRemoveApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRemovePost method -func NewClustersStorageNodesDevicesRemoveApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRemovePostRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, deviceId openapi_types.UUID, params *ClustersStorageNodesDevicesRemoveApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRemovePostParams) (*http.Request, error) { +// NewClustersReplicationTargetsFailoverApiV2ClustersClusterIdReplicationTargetsTargetIdFailoverPostRequest constructs an http.Request for the ClustersReplicationTargetsFailoverApiV2ClustersClusterIdReplicationTargetsTargetIdFailoverPost method +func NewClustersReplicationTargetsFailoverApiV2ClustersClusterIdReplicationTargetsTargetIdFailoverPostRequest(server string, clusterId openapi_types.UUID, targetId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -7978,14 +8464,7 @@ func NewClustersStorageNodesDevicesRemoveApiV2ClustersClusterIdStorageNodesStora var pathParam1 string - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "storage_node_id", storageNodeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err - } - - var pathParam2 string - - pathParam2, err = runtime.StyleParamWithOptions("simple", false, "device_id", deviceId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "target_id", targetId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -7995,7 +8474,7 @@ func NewClustersStorageNodesDevicesRemoveApiV2ClustersClusterIdStorageNodesStora return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/devices/%s/remove", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/replication/targets/%s/failover", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -8005,33 +8484,6 @@ func NewClustersStorageNodesDevicesRemoveApiV2ClustersClusterIdStorageNodesStora return nil, err } - if params != nil { - // queryValues collects non-styled parameters (passthrough, JSON) - // that are safe to round-trip through url.Values.Encode(). - queryValues := queryURL.Query() - // rawQueryFragments collects pre-encoded query fragments from - // styled parameters, preserving literal commas as delimiters - // per the OpenAPI spec (e.g. "color=blue,black,brown"). - var rawQueryFragments []string - - if params.Force != nil { - - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "force", *params.Force, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { - return nil, err - } else { - for _, qp := range strings.Split(queryFrag, "&") { - rawQueryFragments = append(rawQueryFragments, qp) - } - } - - } - - if encoded := queryValues.Encode(); encoded != "" { - rawQueryFragments = append(rawQueryFragments, encoded) - } - queryURL.RawQuery = strings.Join(rawQueryFragments, "&") - } - req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) if err != nil { return nil, err @@ -8040,8 +8492,8 @@ func NewClustersStorageNodesDevicesRemoveApiV2ClustersClusterIdStorageNodesStora return req, nil } -// NewClustersStorageNodesDevicesResetApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdResetPostRequest constructs an http.Request for the ClustersStorageNodesDevicesResetApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdResetPost method -func NewClustersStorageNodesDevicesResetApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdResetPostRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, deviceId openapi_types.UUID) (*http.Request, error) { +// NewClustersShutdownApiV2ClustersClusterIdShutdownPostRequest constructs an http.Request for the ClustersShutdownApiV2ClustersClusterIdShutdownPost method +func NewClustersShutdownApiV2ClustersClusterIdShutdownPostRequest(server string, clusterId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -8051,26 +8503,12 @@ func NewClustersStorageNodesDevicesResetApiV2ClustersClusterIdStorageNodesStorag return nil, err } - var pathParam1 string - - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "storage_node_id", storageNodeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err - } - - var pathParam2 string - - pathParam2, err = runtime.StyleParamWithOptions("simple", false, "device_id", deviceId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err - } - serverURL, err := url.Parse(server) if err != nil { return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/devices/%s/reset", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/shutdown", pathParam0) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -8088,8 +8526,8 @@ func NewClustersStorageNodesDevicesResetApiV2ClustersClusterIdStorageNodesStorag return req, nil } -// NewClustersStorageNodesDevicesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRestartPostRequest constructs an http.Request for the ClustersStorageNodesDevicesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRestartPost method -func NewClustersStorageNodesDevicesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRestartPostRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, deviceId openapi_types.UUID, params *ClustersStorageNodesDevicesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRestartPostParams) (*http.Request, error) { +// NewClustersStartApiV2ClustersClusterIdStartPostRequest constructs an http.Request for the ClustersStartApiV2ClustersClusterIdStartPost method +func NewClustersStartApiV2ClustersClusterIdStartPostRequest(server string, clusterId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -8099,16 +8537,36 @@ func NewClustersStorageNodesDevicesRestartApiV2ClustersClusterIdStorageNodesStor return nil, err } - var pathParam1 string + serverURL, err := url.Parse(server) + if err != nil { + return nil, err + } - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "storage_node_id", storageNodeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/start", pathParam0) + if operationPath[0] == '/' { + operationPath = "." + operationPath + } + + queryURL, err := serverURL.Parse(operationPath) if err != nil { return nil, err } - var pathParam2 string + req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) + if err != nil { + return nil, err + } - pathParam2, err = runtime.StyleParamWithOptions("simple", false, "device_id", deviceId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + return req, nil +} + +// NewClustersStorageNodesListApiV2ClustersClusterIdStorageNodesGetRequest constructs an http.Request for the ClustersStorageNodesListApiV2ClustersClusterIdStorageNodesGet method +func NewClustersStorageNodesListApiV2ClustersClusterIdStorageNodesGetRequest(server string, clusterId openapi_types.UUID, params *ClustersStorageNodesListApiV2ClustersClusterIdStorageNodesGetParams) (*http.Request, error) { + var err error + + var pathParam0 string + + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -8118,7 +8576,7 @@ func NewClustersStorageNodesDevicesRestartApiV2ClustersClusterIdStorageNodesStor return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/devices/%s/restart", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/", pathParam0) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -8137,9 +8595,9 @@ func NewClustersStorageNodesDevicesRestartApiV2ClustersClusterIdStorageNodesStor // per the OpenAPI spec (e.g. "color=blue,black,brown"). var rawQueryFragments []string - if params.Force != nil { + if params.Watch != nil { - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "force", *params.Force, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "watch", *params.Watch, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { return nil, err } else { for _, qp := range strings.Split(queryFrag, "&") { @@ -8155,7 +8613,7 @@ func NewClustersStorageNodesDevicesRestartApiV2ClustersClusterIdStorageNodesStor queryURL.RawQuery = strings.Join(rawQueryFragments, "&") } - req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err } @@ -8163,20 +8621,24 @@ func NewClustersStorageNodesDevicesRestartApiV2ClustersClusterIdStorageNodesStor return req, nil } -// NewClustersStorageNodesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdIostatsGetRequest constructs an http.Request for the ClustersStorageNodesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdIostatsGet method -func NewClustersStorageNodesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdIostatsGetRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, params *ClustersStorageNodesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdIostatsGetParams) (*http.Request, error) { - var err error - - var pathParam0 string - - pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { +// NewClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostRequest calls the generic ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPost builder with application/json body +func NewClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostRequest(server string, clusterId openapi_types.UUID, params *ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostParams, body ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostJSONRequestBody) (*http.Request, error) { + var bodyReader io.Reader + buf, err := json.Marshal(body) + if err != nil { return nil, err } + bodyReader = bytes.NewReader(buf) + return NewClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostRequestWithBody(server, clusterId, params, "application/json", bodyReader) +} - var pathParam1 string +// NewClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostRequestWithBody constructs an http.Request for the ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPost method, with any body, and a specified content type +func NewClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostRequestWithBody(server string, clusterId openapi_types.UUID, params *ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostParams, contentType string, body io.Reader) (*http.Request, error) { + var err error - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "storage_node_id", storageNodeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + var pathParam0 string + + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -8186,7 +8648,7 @@ func NewClustersStorageNodesIostatsApiV2ClustersClusterIdStorageNodesStorageNode return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/iostats", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/", pathParam0) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -8205,9 +8667,9 @@ func NewClustersStorageNodesIostatsApiV2ClustersClusterIdStorageNodesStorageNode // per the OpenAPI spec (e.g. "color=blue,black,brown"). var rawQueryFragments []string - if params.History != nil { + if params.ResponseFormat != nil { - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "history", *params.History, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "", Format: ""}); err != nil { + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "response-format", *params.ResponseFormat, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "string", Format: ""}); err != nil { return nil, err } else { for _, qp := range strings.Split(queryFrag, "&") { @@ -8223,16 +8685,18 @@ func NewClustersStorageNodesIostatsApiV2ClustersClusterIdStorageNodesStorageNode queryURL.RawQuery = strings.Join(rawQueryFragments, "&") } - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) if err != nil { return nil, err } + req.Header.Add("Content-Type", contentType) + return req, nil } -// NewClustersStorageNodesNicsListApiV2ClustersClusterIdStorageNodesStorageNodeIdNicsGetRequest constructs an http.Request for the ClustersStorageNodesNicsListApiV2ClustersClusterIdStorageNodesStorageNodeIdNicsGet method -func NewClustersStorageNodesNicsListApiV2ClustersClusterIdStorageNodesStorageNodeIdNicsGetRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID) (*http.Request, error) { +// NewClustersStorageNodesDeleteApiV2ClustersClusterIdStorageNodesStorageNodeIdDeleteRequest constructs an http.Request for the ClustersStorageNodesDeleteApiV2ClustersClusterIdStorageNodesStorageNodeIdDelete method +func NewClustersStorageNodesDeleteApiV2ClustersClusterIdStorageNodesStorageNodeIdDeleteRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, params *ClustersStorageNodesDeleteApiV2ClustersClusterIdStorageNodesStorageNodeIdDeleteParams) (*http.Request, error) { var err error var pathParam0 string @@ -8254,7 +8718,7 @@ func NewClustersStorageNodesNicsListApiV2ClustersClusterIdStorageNodesStorageNod return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/nics", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -8264,55 +8728,58 @@ func NewClustersStorageNodesNicsListApiV2ClustersClusterIdStorageNodesStorageNod return nil, err } - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) - if err != nil { - return nil, err - } - - return req, nil -} + if params != nil { + // queryValues collects non-styled parameters (passthrough, JSON) + // that are safe to round-trip through url.Values.Encode(). + queryValues := queryURL.Query() + // rawQueryFragments collects pre-encoded query fragments from + // styled parameters, preserving literal commas as delimiters + // per the OpenAPI spec (e.g. "color=blue,black,brown"). + var rawQueryFragments []string -// NewClustersStorageNodesNicsIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdNicsNicIdIostatsGetRequest constructs an http.Request for the ClustersStorageNodesNicsIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdNicsNicIdIostatsGet method -func NewClustersStorageNodesNicsIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdNicsNicIdIostatsGetRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, nicId string) (*http.Request, error) { - var err error + if params.ForceRemove != nil { - var pathParam0 string + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "force_remove", *params.ForceRemove, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } + } - pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err - } + } - var pathParam1 string + if params.ForceMigrate != nil { - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "storage_node_id", storageNodeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err - } + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "force_migrate", *params.ForceMigrate, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } + } - var pathParam2 string + } - pathParam2, err = runtime.StyleParamWithOptions("simple", false, "nic_id", nicId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: ""}) - if err != nil { - return nil, err - } + if params.ForceDelete != nil { - serverURL, err := url.Parse(server) - if err != nil { - return nil, err - } + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "force_delete", *params.ForceDelete, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } + } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/nics/%s/iostats", pathParam0, pathParam1, pathParam2) - if operationPath[0] == '/' { - operationPath = "." + operationPath - } + } - queryURL, err := serverURL.Parse(operationPath) - if err != nil { - return nil, err + if encoded := queryValues.Encode(); encoded != "" { + rawQueryFragments = append(rawQueryFragments, encoded) + } + queryURL.RawQuery = strings.Join(rawQueryFragments, "&") } - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodDelete, queryURL.String(), nil) if err != nil { return nil, err } @@ -8320,8 +8787,8 @@ func NewClustersStorageNodesNicsIostatsApiV2ClustersClusterIdStorageNodesStorage return req, nil } -// NewClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdPromotePostRequest constructs an http.Request for the ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdPromotePost method -func NewClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdPromotePostRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID) (*http.Request, error) { +// NewClustersStorageNodesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdGetRequest constructs an http.Request for the ClustersStorageNodesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdGet method +func NewClustersStorageNodesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdGetRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, params *ClustersStorageNodesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdGetParams) (*http.Request, error) { var err error var pathParam0 string @@ -8343,7 +8810,7 @@ func NewClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeId return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/promote", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -8353,27 +8820,43 @@ func NewClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeId return nil, err } - req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) - if err != nil { - return nil, err - } + if params != nil { + // queryValues collects non-styled parameters (passthrough, JSON) + // that are safe to round-trip through url.Values.Encode(). + queryValues := queryURL.Query() + // rawQueryFragments collects pre-encoded query fragments from + // styled parameters, preserving literal commas as delimiters + // per the OpenAPI spec (e.g. "color=blue,black,brown"). + var rawQueryFragments []string - return req, nil -} + if params.Watch != nil { -// NewClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPostRequest calls the generic ClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPost builder with application/json body -func NewClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPostRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, body ClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPostJSONRequestBody) (*http.Request, error) { - var bodyReader io.Reader - buf, err := json.Marshal(body) + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "watch", *params.Watch, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } + } + + } + + if encoded := queryValues.Encode(); encoded != "" { + rawQueryFragments = append(rawQueryFragments, encoded) + } + queryURL.RawQuery = strings.Join(rawQueryFragments, "&") + } + + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err } - bodyReader = bytes.NewReader(buf) - return NewClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPostRequestWithBody(server, clusterId, storageNodeId, "application/json", bodyReader) + + return req, nil } -// NewClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPostRequestWithBody constructs an http.Request for the ClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPost method, with any body, and a specified content type -func NewClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPostRequestWithBody(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { +// NewClustersStorageNodesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdCapacityGetRequest constructs an http.Request for the ClustersStorageNodesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdCapacityGet method +func NewClustersStorageNodesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdCapacityGetRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, params *ClustersStorageNodesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdCapacityGetParams) (*http.Request, error) { var err error var pathParam0 string @@ -8395,7 +8878,7 @@ func NewClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNode return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/restart", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/capacity", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -8405,18 +8888,43 @@ func NewClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNode return nil, err } - req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) + if params != nil { + // queryValues collects non-styled parameters (passthrough, JSON) + // that are safe to round-trip through url.Values.Encode(). + queryValues := queryURL.Query() + // rawQueryFragments collects pre-encoded query fragments from + // styled parameters, preserving literal commas as delimiters + // per the OpenAPI spec (e.g. "color=blue,black,brown"). + var rawQueryFragments []string + + if params.History != nil { + + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "history", *params.History, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "", Format: ""}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } + } + + } + + if encoded := queryValues.Encode(); encoded != "" { + rawQueryFragments = append(rawQueryFragments, encoded) + } + queryURL.RawQuery = strings.Join(rawQueryFragments, "&") + } + + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err } - req.Header.Add("Content-Type", contentType) - return req, nil } -// NewClustersStorageNodesResumeApiV2ClustersClusterIdStorageNodesStorageNodeIdResumePostRequest constructs an http.Request for the ClustersStorageNodesResumeApiV2ClustersClusterIdStorageNodesStorageNodeIdResumePost method -func NewClustersStorageNodesResumeApiV2ClustersClusterIdStorageNodesStorageNodeIdResumePostRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID) (*http.Request, error) { +// NewClustersStorageNodesDevicesListApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesGetRequest constructs an http.Request for the ClustersStorageNodesDevicesListApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesGet method +func NewClustersStorageNodesDevicesListApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesGetRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, params *ClustersStorageNodesDevicesListApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesGetParams) (*http.Request, error) { var err error var pathParam0 string @@ -8438,7 +8946,7 @@ func NewClustersStorageNodesResumeApiV2ClustersClusterIdStorageNodesStorageNodeI return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/resume", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/devices/", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -8448,7 +8956,34 @@ func NewClustersStorageNodesResumeApiV2ClustersClusterIdStorageNodesStorageNodeI return nil, err } - req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) + if params != nil { + // queryValues collects non-styled parameters (passthrough, JSON) + // that are safe to round-trip through url.Values.Encode(). + queryValues := queryURL.Query() + // rawQueryFragments collects pre-encoded query fragments from + // styled parameters, preserving literal commas as delimiters + // per the OpenAPI spec (e.g. "color=blue,black,brown"). + var rawQueryFragments []string + + if params.Watch != nil { + + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "watch", *params.Watch, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } + } + + } + + if encoded := queryValues.Encode(); encoded != "" { + rawQueryFragments = append(rawQueryFragments, encoded) + } + queryURL.RawQuery = strings.Join(rawQueryFragments, "&") + } + + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err } @@ -8456,8 +8991,8 @@ func NewClustersStorageNodesResumeApiV2ClustersClusterIdStorageNodesStorageNodeI return req, nil } -// NewClustersStorageNodesShutdownApiV2ClustersClusterIdStorageNodesStorageNodeIdShutdownPostRequest constructs an http.Request for the ClustersStorageNodesShutdownApiV2ClustersClusterIdStorageNodesStorageNodeIdShutdownPost method -func NewClustersStorageNodesShutdownApiV2ClustersClusterIdStorageNodesStorageNodeIdShutdownPostRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, params *ClustersStorageNodesShutdownApiV2ClustersClusterIdStorageNodesStorageNodeIdShutdownPostParams) (*http.Request, error) { +// NewClustersStorageNodesDevicesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdGetRequest constructs an http.Request for the ClustersStorageNodesDevicesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdGet method +func NewClustersStorageNodesDevicesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdGetRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, deviceId openapi_types.UUID, params *ClustersStorageNodesDevicesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdGetParams) (*http.Request, error) { var err error var pathParam0 string @@ -8474,12 +9009,19 @@ func NewClustersStorageNodesShutdownApiV2ClustersClusterIdStorageNodesStorageNod return nil, err } + var pathParam2 string + + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "device_id", deviceId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + serverURL, err := url.Parse(server) if err != nil { return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/shutdown", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/devices/%s/", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -8498,9 +9040,9 @@ func NewClustersStorageNodesShutdownApiV2ClustersClusterIdStorageNodesStorageNod // per the OpenAPI spec (e.g. "color=blue,black,brown"). var rawQueryFragments []string - if params.Force != nil { + if params.Watch != nil { - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "force", *params.Force, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "watch", *params.Watch, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { return nil, err } else { for _, qp := range strings.Split(queryFrag, "&") { @@ -8516,7 +9058,7 @@ func NewClustersStorageNodesShutdownApiV2ClustersClusterIdStorageNodesStorageNod queryURL.RawQuery = strings.Join(rawQueryFragments, "&") } - req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err } @@ -8524,19 +9066,8 @@ func NewClustersStorageNodesShutdownApiV2ClustersClusterIdStorageNodesStorageNod return req, nil } -// NewClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPostRequest calls the generic ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPost builder with application/json body -func NewClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPostRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, body ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPostJSONRequestBody) (*http.Request, error) { - var bodyReader io.Reader - buf, err := json.Marshal(body) - if err != nil { - return nil, err - } - bodyReader = bytes.NewReader(buf) - return NewClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPostRequestWithBody(server, clusterId, storageNodeId, "application/json", bodyReader) -} - -// NewClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPostRequestWithBody constructs an http.Request for the ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPost method, with any body, and a specified content type -func NewClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPostRequestWithBody(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { +// NewClustersStorageNodesDevicesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdCapacityGetRequest constructs an http.Request for the ClustersStorageNodesDevicesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdCapacityGet method +func NewClustersStorageNodesDevicesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdCapacityGetRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, deviceId openapi_types.UUID, params *ClustersStorageNodesDevicesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdCapacityGetParams) (*http.Request, error) { var err error var pathParam0 string @@ -8553,12 +9084,19 @@ func NewClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeId return nil, err } + var pathParam2 string + + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "device_id", deviceId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + serverURL, err := url.Parse(server) if err != nil { return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/start", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/devices/%s/capacity", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -8568,18 +9106,43 @@ func NewClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeId return nil, err } - req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) + if params != nil { + // queryValues collects non-styled parameters (passthrough, JSON) + // that are safe to round-trip through url.Values.Encode(). + queryValues := queryURL.Query() + // rawQueryFragments collects pre-encoded query fragments from + // styled parameters, preserving literal commas as delimiters + // per the OpenAPI spec (e.g. "color=blue,black,brown"). + var rawQueryFragments []string + + if params.History != nil { + + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "history", *params.History, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "", Format: ""}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } + } + + } + + if encoded := queryValues.Encode(); encoded != "" { + rawQueryFragments = append(rawQueryFragments, encoded) + } + queryURL.RawQuery = strings.Join(rawQueryFragments, "&") + } + + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err } - req.Header.Add("Content-Type", contentType) - return req, nil } -// NewClustersStorageNodesSuspendApiV2ClustersClusterIdStorageNodesStorageNodeIdSuspendPostRequest constructs an http.Request for the ClustersStorageNodesSuspendApiV2ClustersClusterIdStorageNodesStorageNodeIdSuspendPost method -func NewClustersStorageNodesSuspendApiV2ClustersClusterIdStorageNodesStorageNodeIdSuspendPostRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, params *ClustersStorageNodesSuspendApiV2ClustersClusterIdStorageNodesStorageNodeIdSuspendPostParams) (*http.Request, error) { +// NewClustersStorageNodesDevicesGetDeviceHealthInfoApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdHealthInfoGetRequest constructs an http.Request for the ClustersStorageNodesDevicesGetDeviceHealthInfoApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdHealthInfoGet method +func NewClustersStorageNodesDevicesGetDeviceHealthInfoApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdHealthInfoGetRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, deviceId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -8596,12 +9159,19 @@ func NewClustersStorageNodesSuspendApiV2ClustersClusterIdStorageNodesStorageNode return nil, err } + var pathParam2 string + + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "device_id", deviceId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + serverURL, err := url.Parse(server) if err != nil { return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/suspend", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/devices/%s/health-info", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -8611,34 +9181,7 @@ func NewClustersStorageNodesSuspendApiV2ClustersClusterIdStorageNodesStorageNode return nil, err } - if params != nil { - // queryValues collects non-styled parameters (passthrough, JSON) - // that are safe to round-trip through url.Values.Encode(). - queryValues := queryURL.Query() - // rawQueryFragments collects pre-encoded query fragments from - // styled parameters, preserving literal commas as delimiters - // per the OpenAPI spec (e.g. "color=blue,black,brown"). - var rawQueryFragments []string - - if params.Force != nil { - - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "force", *params.Force, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { - return nil, err - } else { - for _, qp := range strings.Split(queryFrag, "&") { - rawQueryFragments = append(rawQueryFragments, qp) - } - } - - } - - if encoded := queryValues.Encode(); encoded != "" { - rawQueryFragments = append(rawQueryFragments, encoded) - } - queryURL.RawQuery = strings.Join(rawQueryFragments, "&") - } - - req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err } @@ -8646,8 +9189,8 @@ func NewClustersStorageNodesSuspendApiV2ClustersClusterIdStorageNodesStorageNode return req, nil } -// NewClustersStoragePoolsListApiV2ClustersClusterIdStoragePoolsGetRequest constructs an http.Request for the ClustersStoragePoolsListApiV2ClustersClusterIdStoragePoolsGet method -func NewClustersStoragePoolsListApiV2ClustersClusterIdStoragePoolsGetRequest(server string, clusterId openapi_types.UUID, params *ClustersStoragePoolsListApiV2ClustersClusterIdStoragePoolsGetParams) (*http.Request, error) { +// NewClustersStorageNodesDevicesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdIostatsGetRequest constructs an http.Request for the ClustersStorageNodesDevicesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdIostatsGet method +func NewClustersStorageNodesDevicesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdIostatsGetRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, deviceId openapi_types.UUID, params *ClustersStorageNodesDevicesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdIostatsGetParams) (*http.Request, error) { var err error var pathParam0 string @@ -8657,12 +9200,26 @@ func NewClustersStoragePoolsListApiV2ClustersClusterIdStoragePoolsGetRequest(ser return nil, err } + var pathParam1 string + + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "storage_node_id", storageNodeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + var pathParam2 string + + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "device_id", deviceId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + serverURL, err := url.Parse(server) if err != nil { return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/", pathParam0) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/devices/%s/iostats", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -8681,9 +9238,9 @@ func NewClustersStoragePoolsListApiV2ClustersClusterIdStoragePoolsGetRequest(ser // per the OpenAPI spec (e.g. "color=blue,black,brown"). var rawQueryFragments []string - if params.Watch != nil { + if params.History != nil { - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "watch", *params.Watch, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "history", *params.History, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "", Format: ""}); err != nil { return nil, err } else { for _, qp := range strings.Split(queryFrag, "&") { @@ -8707,24 +9264,27 @@ func NewClustersStoragePoolsListApiV2ClustersClusterIdStoragePoolsGetRequest(ser return req, nil } -// NewClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostRequest calls the generic ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPost builder with application/json body -func NewClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostRequest(server string, clusterId openapi_types.UUID, params *ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostParams, body ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostJSONRequestBody) (*http.Request, error) { - var bodyReader io.Reader - buf, err := json.Marshal(body) +// NewClustersStorageNodesDevicesRemoveApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRemovePostRequest constructs an http.Request for the ClustersStorageNodesDevicesRemoveApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRemovePost method +func NewClustersStorageNodesDevicesRemoveApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRemovePostRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, deviceId openapi_types.UUID, params *ClustersStorageNodesDevicesRemoveApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRemovePostParams) (*http.Request, error) { + var err error + + var pathParam0 string + + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } - bodyReader = bytes.NewReader(buf) - return NewClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostRequestWithBody(server, clusterId, params, "application/json", bodyReader) -} -// NewClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostRequestWithBody constructs an http.Request for the ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPost method, with any body, and a specified content type -func NewClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostRequestWithBody(server string, clusterId openapi_types.UUID, params *ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostParams, contentType string, body io.Reader) (*http.Request, error) { - var err error + var pathParam1 string - var pathParam0 string + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "storage_node_id", storageNodeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } - pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + var pathParam2 string + + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "device_id", deviceId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -8734,7 +9294,7 @@ func NewClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostRequestW return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/", pathParam0) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/devices/%s/remove", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -8753,9 +9313,9 @@ func NewClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostRequestW // per the OpenAPI spec (e.g. "color=blue,black,brown"). var rawQueryFragments []string - if params.ResponseFormat != nil { + if params.Force != nil { - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "response-format", *params.ResponseFormat, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "string", Format: ""}); err != nil { + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "force", *params.Force, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { return nil, err } else { for _, qp := range strings.Split(queryFrag, "&") { @@ -8771,18 +9331,16 @@ func NewClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostRequestW queryURL.RawQuery = strings.Join(rawQueryFragments, "&") } - req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) + req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) if err != nil { return nil, err } - req.Header.Add("Content-Type", contentType) - return req, nil } -// NewClustersStoragePoolsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdDeleteRequest constructs an http.Request for the ClustersStoragePoolsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdDelete method -func NewClustersStoragePoolsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdDeleteRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID) (*http.Request, error) { +// NewClustersStorageNodesDevicesResetApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdResetPostRequest constructs an http.Request for the ClustersStorageNodesDevicesResetApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdResetPost method +func NewClustersStorageNodesDevicesResetApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdResetPostRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, deviceId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -8794,7 +9352,14 @@ func NewClustersStoragePoolsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdDelete var pathParam1 string - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "pool_id", poolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "storage_node_id", storageNodeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + var pathParam2 string + + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "device_id", deviceId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -8804,7 +9369,7 @@ func NewClustersStoragePoolsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdDelete return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/devices/%s/reset", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -8814,7 +9379,7 @@ func NewClustersStoragePoolsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdDelete return nil, err } - req, err := http.NewRequest(http.MethodDelete, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) if err != nil { return nil, err } @@ -8822,8 +9387,8 @@ func NewClustersStoragePoolsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdDelete return req, nil } -// NewClustersStoragePoolsDetailApiV2ClustersClusterIdStoragePoolsPoolIdGetRequest constructs an http.Request for the ClustersStoragePoolsDetailApiV2ClustersClusterIdStoragePoolsPoolIdGet method -func NewClustersStoragePoolsDetailApiV2ClustersClusterIdStoragePoolsPoolIdGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, params *ClustersStoragePoolsDetailApiV2ClustersClusterIdStoragePoolsPoolIdGetParams) (*http.Request, error) { +// NewClustersStorageNodesDevicesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRestartPostRequest constructs an http.Request for the ClustersStorageNodesDevicesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRestartPost method +func NewClustersStorageNodesDevicesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRestartPostRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, deviceId openapi_types.UUID, params *ClustersStorageNodesDevicesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRestartPostParams) (*http.Request, error) { var err error var pathParam0 string @@ -8835,7 +9400,14 @@ func NewClustersStoragePoolsDetailApiV2ClustersClusterIdStoragePoolsPoolIdGetReq var pathParam1 string - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "pool_id", poolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "storage_node_id", storageNodeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + var pathParam2 string + + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "device_id", deviceId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -8845,7 +9417,7 @@ func NewClustersStoragePoolsDetailApiV2ClustersClusterIdStoragePoolsPoolIdGetReq return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/devices/%s/restart", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -8864,9 +9436,9 @@ func NewClustersStoragePoolsDetailApiV2ClustersClusterIdStoragePoolsPoolIdGetReq // per the OpenAPI spec (e.g. "color=blue,black,brown"). var rawQueryFragments []string - if params.Watch != nil { + if params.Force != nil { - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "watch", *params.Watch, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "force", *params.Force, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { return nil, err } else { for _, qp := range strings.Split(queryFrag, "&") { @@ -8882,7 +9454,7 @@ func NewClustersStoragePoolsDetailApiV2ClustersClusterIdStoragePoolsPoolIdGetReq queryURL.RawQuery = strings.Join(rawQueryFragments, "&") } - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) if err != nil { return nil, err } @@ -8890,19 +9462,8 @@ func NewClustersStoragePoolsDetailApiV2ClustersClusterIdStoragePoolsPoolIdGetReq return req, nil } -// NewClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutRequest calls the generic ClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPut builder with application/json body -func NewClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, body ClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutJSONRequestBody) (*http.Request, error) { - var bodyReader io.Reader - buf, err := json.Marshal(body) - if err != nil { - return nil, err - } - bodyReader = bytes.NewReader(buf) - return NewClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutRequestWithBody(server, clusterId, poolId, "application/json", bodyReader) -} - -// NewClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutRequestWithBody constructs an http.Request for the ClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPut method, with any body, and a specified content type -func NewClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutRequestWithBody(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { +// NewClustersStorageNodesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdIostatsGetRequest constructs an http.Request for the ClustersStorageNodesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdIostatsGet method +func NewClustersStorageNodesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdIostatsGetRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, params *ClustersStorageNodesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdIostatsGetParams) (*http.Request, error) { var err error var pathParam0 string @@ -8914,7 +9475,7 @@ func NewClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutReq var pathParam1 string - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "pool_id", poolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "storage_node_id", storageNodeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -8924,7 +9485,7 @@ func NewClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutReq return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/iostats", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -8934,29 +9495,43 @@ func NewClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutReq return nil, err } - req, err := http.NewRequest(http.MethodPut, queryURL.String(), body) - if err != nil { - return nil, err - } + if params != nil { + // queryValues collects non-styled parameters (passthrough, JSON) + // that are safe to round-trip through url.Values.Encode(). + queryValues := queryURL.Query() + // rawQueryFragments collects pre-encoded query fragments from + // styled parameters, preserving literal commas as delimiters + // per the OpenAPI spec (e.g. "color=blue,black,brown"). + var rawQueryFragments []string - req.Header.Add("Content-Type", contentType) + if params.History != nil { - return req, nil -} + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "history", *params.History, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "", Format: ""}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } + } -// NewClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDeleteRequest calls the generic ClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDelete builder with application/json body -func NewClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDeleteRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, body ClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDeleteJSONRequestBody) (*http.Request, error) { - var bodyReader io.Reader - buf, err := json.Marshal(body) + } + + if encoded := queryValues.Encode(); encoded != "" { + rawQueryFragments = append(rawQueryFragments, encoded) + } + queryURL.RawQuery = strings.Join(rawQueryFragments, "&") + } + + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err } - bodyReader = bytes.NewReader(buf) - return NewClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDeleteRequestWithBody(server, clusterId, poolId, "application/json", bodyReader) + + return req, nil } -// NewClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDeleteRequestWithBody constructs an http.Request for the ClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDelete method, with any body, and a specified content type -func NewClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDeleteRequestWithBody(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { +// NewClustersStorageNodesNicsListApiV2ClustersClusterIdStorageNodesStorageNodeIdNicsGetRequest constructs an http.Request for the ClustersStorageNodesNicsListApiV2ClustersClusterIdStorageNodesStorageNodeIdNicsGet method +func NewClustersStorageNodesNicsListApiV2ClustersClusterIdStorageNodesStorageNodeIdNicsGetRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -8968,7 +9543,7 @@ func NewClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHo var pathParam1 string - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "pool_id", poolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "storage_node_id", storageNodeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -8978,7 +9553,7 @@ func NewClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHo return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/host", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/nics", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -8988,29 +9563,116 @@ func NewClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHo return nil, err } - req, err := http.NewRequest(http.MethodDelete, queryURL.String(), body) + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err } - req.Header.Add("Content-Type", contentType) + return req, nil +} + +// NewClustersStorageNodesNicsIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdNicsNicIdIostatsGetRequest constructs an http.Request for the ClustersStorageNodesNicsIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdNicsNicIdIostatsGet method +func NewClustersStorageNodesNicsIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdNicsNicIdIostatsGetRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, nicId string) (*http.Request, error) { + var err error + + var pathParam0 string + + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + var pathParam1 string + + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "storage_node_id", storageNodeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + var pathParam2 string + + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "nic_id", nicId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: ""}) + if err != nil { + return nil, err + } + + serverURL, err := url.Parse(server) + if err != nil { + return nil, err + } + + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/nics/%s/iostats", pathParam0, pathParam1, pathParam2) + if operationPath[0] == '/' { + operationPath = "." + operationPath + } + + queryURL, err := serverURL.Parse(operationPath) + if err != nil { + return nil, err + } + + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + if err != nil { + return nil, err + } return req, nil } -// NewClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPostRequest calls the generic ClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPost builder with application/json body -func NewClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, body ClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPostJSONRequestBody) (*http.Request, error) { +// NewClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdPromotePostRequest constructs an http.Request for the ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdPromotePost method +func NewClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdPromotePostRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID) (*http.Request, error) { + var err error + + var pathParam0 string + + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + var pathParam1 string + + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "storage_node_id", storageNodeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + serverURL, err := url.Parse(server) + if err != nil { + return nil, err + } + + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/promote", pathParam0, pathParam1) + if operationPath[0] == '/' { + operationPath = "." + operationPath + } + + queryURL, err := serverURL.Parse(operationPath) + if err != nil { + return nil, err + } + + req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) + if err != nil { + return nil, err + } + + return req, nil +} + +// NewClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPostRequest calls the generic ClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPost builder with application/json body +func NewClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPostRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, body ClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPostJSONRequestBody) (*http.Request, error) { var bodyReader io.Reader buf, err := json.Marshal(body) if err != nil { return nil, err } bodyReader = bytes.NewReader(buf) - return NewClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPostRequestWithBody(server, clusterId, poolId, "application/json", bodyReader) + return NewClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPostRequestWithBody(server, clusterId, storageNodeId, "application/json", bodyReader) } -// NewClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPostRequestWithBody constructs an http.Request for the ClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPost method, with any body, and a specified content type -func NewClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPostRequestWithBody(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { +// NewClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPostRequestWithBody constructs an http.Request for the ClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPost method, with any body, and a specified content type +func NewClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPostRequestWithBody(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { var err error var pathParam0 string @@ -9022,7 +9684,7 @@ func NewClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostP var pathParam1 string - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "pool_id", poolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "storage_node_id", storageNodeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -9032,7 +9694,7 @@ func NewClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostP return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/host", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/restart", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -9052,8 +9714,8 @@ func NewClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostP return req, nil } -// NewClustersStoragePoolsIostatsApiV2ClustersClusterIdStoragePoolsPoolIdIostatsGetRequest constructs an http.Request for the ClustersStoragePoolsIostatsApiV2ClustersClusterIdStoragePoolsPoolIdIostatsGet method -func NewClustersStoragePoolsIostatsApiV2ClustersClusterIdStoragePoolsPoolIdIostatsGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, params *ClustersStoragePoolsIostatsApiV2ClustersClusterIdStoragePoolsPoolIdIostatsGetParams) (*http.Request, error) { +// NewClustersStorageNodesResumeApiV2ClustersClusterIdStorageNodesStorageNodeIdResumePostRequest constructs an http.Request for the ClustersStorageNodesResumeApiV2ClustersClusterIdStorageNodesStorageNodeIdResumePost method +func NewClustersStorageNodesResumeApiV2ClustersClusterIdStorageNodesStorageNodeIdResumePostRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -9065,7 +9727,7 @@ func NewClustersStoragePoolsIostatsApiV2ClustersClusterIdStoragePoolsPoolIdIosta var pathParam1 string - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "pool_id", poolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "storage_node_id", storageNodeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -9075,7 +9737,7 @@ func NewClustersStoragePoolsIostatsApiV2ClustersClusterIdStoragePoolsPoolIdIosta return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/iostats", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/resume", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -9085,34 +9747,7 @@ func NewClustersStoragePoolsIostatsApiV2ClustersClusterIdStoragePoolsPoolIdIosta return nil, err } - if params != nil { - // queryValues collects non-styled parameters (passthrough, JSON) - // that are safe to round-trip through url.Values.Encode(). - queryValues := queryURL.Query() - // rawQueryFragments collects pre-encoded query fragments from - // styled parameters, preserving literal commas as delimiters - // per the OpenAPI spec (e.g. "color=blue,black,brown"). - var rawQueryFragments []string - - if params.Limit != nil { - - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "limit", *params.Limit, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "integer", Format: ""}); err != nil { - return nil, err - } else { - for _, qp := range strings.Split(queryFrag, "&") { - rawQueryFragments = append(rawQueryFragments, qp) - } - } - - } - - if encoded := queryValues.Encode(); encoded != "" { - rawQueryFragments = append(rawQueryFragments, encoded) - } - queryURL.RawQuery = strings.Join(rawQueryFragments, "&") - } - - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) if err != nil { return nil, err } @@ -9120,8 +9755,8 @@ func NewClustersStoragePoolsIostatsApiV2ClustersClusterIdStoragePoolsPoolIdIosta return req, nil } -// NewClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsGetRequest constructs an http.Request for the ClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsGet method -func NewClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, params *ClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsGetParams) (*http.Request, error) { +// NewClustersStorageNodesShutdownApiV2ClustersClusterIdStorageNodesStorageNodeIdShutdownPostRequest constructs an http.Request for the ClustersStorageNodesShutdownApiV2ClustersClusterIdStorageNodesStorageNodeIdShutdownPost method +func NewClustersStorageNodesShutdownApiV2ClustersClusterIdStorageNodesStorageNodeIdShutdownPostRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, params *ClustersStorageNodesShutdownApiV2ClustersClusterIdStorageNodesStorageNodeIdShutdownPostParams) (*http.Request, error) { var err error var pathParam0 string @@ -9133,7 +9768,7 @@ func NewClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolI var pathParam1 string - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "pool_id", poolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "storage_node_id", storageNodeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -9143,7 +9778,7 @@ func NewClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolI return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/snapshots/", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/shutdown", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -9162,9 +9797,9 @@ func NewClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolI // per the OpenAPI spec (e.g. "color=blue,black,brown"). var rawQueryFragments []string - if params.Watch != nil { + if params.Force != nil { - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "watch", *params.Watch, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "force", *params.Force, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { return nil, err } else { for _, qp := range strings.Split(queryFrag, "&") { @@ -9180,7 +9815,7 @@ func NewClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolI queryURL.RawQuery = strings.Join(rawQueryFragments, "&") } - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) if err != nil { return nil, err } @@ -9188,8 +9823,19 @@ func NewClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolI return req, nil } -// NewClustersStoragePoolsSnapshotsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdDeleteRequest constructs an http.Request for the ClustersStoragePoolsSnapshotsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdDelete method -func NewClustersStoragePoolsSnapshotsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdDeleteRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, snapshotId openapi_types.UUID) (*http.Request, error) { +// NewClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPostRequest calls the generic ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPost builder with application/json body +func NewClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPostRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, body ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPostJSONRequestBody) (*http.Request, error) { + var bodyReader io.Reader + buf, err := json.Marshal(body) + if err != nil { + return nil, err + } + bodyReader = bytes.NewReader(buf) + return NewClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPostRequestWithBody(server, clusterId, storageNodeId, "application/json", bodyReader) +} + +// NewClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPostRequestWithBody constructs an http.Request for the ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPost method, with any body, and a specified content type +func NewClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPostRequestWithBody(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { var err error var pathParam0 string @@ -9201,14 +9847,7 @@ func NewClustersStoragePoolsSnapshotsDeleteApiV2ClustersClusterIdStoragePoolsPoo var pathParam1 string - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "pool_id", poolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err - } - - var pathParam2 string - - pathParam2, err = runtime.StyleParamWithOptions("simple", false, "snapshot_id", snapshotId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "storage_node_id", storageNodeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -9218,7 +9857,7 @@ func NewClustersStoragePoolsSnapshotsDeleteApiV2ClustersClusterIdStoragePoolsPoo return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/snapshots/%s/", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/start", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -9228,16 +9867,18 @@ func NewClustersStoragePoolsSnapshotsDeleteApiV2ClustersClusterIdStoragePoolsPoo return nil, err } - req, err := http.NewRequest(http.MethodDelete, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) if err != nil { return nil, err } + req.Header.Add("Content-Type", contentType) + return req, nil } -// NewClustersStoragePoolsSnapshotsDetailApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdGetRequest constructs an http.Request for the ClustersStoragePoolsSnapshotsDetailApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdGet method -func NewClustersStoragePoolsSnapshotsDetailApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, snapshotId openapi_types.UUID, params *ClustersStoragePoolsSnapshotsDetailApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdGetParams) (*http.Request, error) { +// NewClustersStorageNodesSuspendApiV2ClustersClusterIdStorageNodesStorageNodeIdSuspendPostRequest constructs an http.Request for the ClustersStorageNodesSuspendApiV2ClustersClusterIdStorageNodesStorageNodeIdSuspendPost method +func NewClustersStorageNodesSuspendApiV2ClustersClusterIdStorageNodesStorageNodeIdSuspendPostRequest(server string, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, params *ClustersStorageNodesSuspendApiV2ClustersClusterIdStorageNodesStorageNodeIdSuspendPostParams) (*http.Request, error) { var err error var pathParam0 string @@ -9249,14 +9890,7 @@ func NewClustersStoragePoolsSnapshotsDetailApiV2ClustersClusterIdStoragePoolsPoo var pathParam1 string - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "pool_id", poolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err - } - - var pathParam2 string - - pathParam2, err = runtime.StyleParamWithOptions("simple", false, "snapshot_id", snapshotId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "storage_node_id", storageNodeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -9266,7 +9900,7 @@ func NewClustersStoragePoolsSnapshotsDetailApiV2ClustersClusterIdStoragePoolsPoo return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/snapshots/%s/", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-nodes/%s/suspend", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -9285,9 +9919,9 @@ func NewClustersStoragePoolsSnapshotsDetailApiV2ClustersClusterIdStoragePoolsPoo // per the OpenAPI spec (e.g. "color=blue,black,brown"). var rawQueryFragments []string - if params.Watch != nil { + if params.Force != nil { - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "watch", *params.Watch, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "force", *params.Force, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { return nil, err } else { for _, qp := range strings.Split(queryFrag, "&") { @@ -9303,7 +9937,7 @@ func NewClustersStoragePoolsSnapshotsDetailApiV2ClustersClusterIdStoragePoolsPoo queryURL.RawQuery = strings.Join(rawQueryFragments, "&") } - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) if err != nil { return nil, err } @@ -9311,8 +9945,8 @@ func NewClustersStoragePoolsSnapshotsDetailApiV2ClustersClusterIdStoragePoolsPoo return req, nil } -// NewClustersStoragePoolsVolumesListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesGetRequest constructs an http.Request for the ClustersStoragePoolsVolumesListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesGet method -func NewClustersStoragePoolsVolumesListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, params *ClustersStoragePoolsVolumesListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesGetParams) (*http.Request, error) { +// NewClustersStoragePoolsListApiV2ClustersClusterIdStoragePoolsGetRequest constructs an http.Request for the ClustersStoragePoolsListApiV2ClustersClusterIdStoragePoolsGet method +func NewClustersStoragePoolsListApiV2ClustersClusterIdStoragePoolsGetRequest(server string, clusterId openapi_types.UUID, params *ClustersStoragePoolsListApiV2ClustersClusterIdStoragePoolsGetParams) (*http.Request, error) { var err error var pathParam0 string @@ -9322,19 +9956,12 @@ func NewClustersStoragePoolsVolumesListApiV2ClustersClusterIdStoragePoolsPoolIdV return nil, err } - var pathParam1 string - - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "pool_id", poolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err - } - serverURL, err := url.Parse(server) if err != nil { return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/", pathParam0) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -9379,19 +10006,19 @@ func NewClustersStoragePoolsVolumesListApiV2ClustersClusterIdStoragePoolsPoolIdV return req, nil } -// NewClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostRequest calls the generic ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPost builder with application/json body -func NewClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, params *ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostParams, body ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostJSONRequestBody) (*http.Request, error) { +// NewClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostRequest calls the generic ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPost builder with application/json body +func NewClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostRequest(server string, clusterId openapi_types.UUID, params *ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostParams, body ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostJSONRequestBody) (*http.Request, error) { var bodyReader io.Reader buf, err := json.Marshal(body) if err != nil { return nil, err } bodyReader = bytes.NewReader(buf) - return NewClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostRequestWithBody(server, clusterId, poolId, params, "application/json", bodyReader) + return NewClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostRequestWithBody(server, clusterId, params, "application/json", bodyReader) } -// NewClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostRequestWithBody constructs an http.Request for the ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPost method, with any body, and a specified content type -func NewClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostRequestWithBody(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, params *ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostParams, contentType string, body io.Reader) (*http.Request, error) { +// NewClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostRequestWithBody constructs an http.Request for the ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPost method, with any body, and a specified content type +func NewClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostRequestWithBody(server string, clusterId openapi_types.UUID, params *ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostParams, contentType string, body io.Reader) (*http.Request, error) { var err error var pathParam0 string @@ -9401,19 +10028,12 @@ func NewClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolI return nil, err } - var pathParam1 string - - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "pool_id", poolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err - } - serverURL, err := url.Parse(server) if err != nil { return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/", pathParam0) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -9460,19 +10080,8 @@ func NewClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolI return req, nil } -// NewClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPostRequest calls the generic ClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPost builder with application/json body -func NewClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, body ClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPostJSONRequestBody) (*http.Request, error) { - var bodyReader io.Reader - buf, err := json.Marshal(body) - if err != nil { - return nil, err - } - bodyReader = bytes.NewReader(buf) - return NewClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPostRequestWithBody(server, clusterId, poolId, "application/json", bodyReader) -} - -// NewClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPostRequestWithBody constructs an http.Request for the ClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPost method, with any body, and a specified content type -func NewClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPostRequestWithBody(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { +// NewClustersStoragePoolsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdDeleteRequest constructs an http.Request for the ClustersStoragePoolsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdDelete method +func NewClustersStoragePoolsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdDeleteRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -9494,7 +10103,7 @@ func NewClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdSt return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/replicate_lvol_on_source_cluster", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -9504,57 +10113,7 @@ func NewClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdSt return nil, err } - req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) - if err != nil { - return nil, err - } - - req.Header.Add("Content-Type", contentType) - - return req, nil -} - -// NewClustersStoragePoolsVolumesDeleteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdDeleteRequest constructs an http.Request for the ClustersStoragePoolsVolumesDeleteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdDelete method -func NewClustersStoragePoolsVolumesDeleteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdDeleteRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID) (*http.Request, error) { - var err error - - var pathParam0 string - - pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err - } - - var pathParam1 string - - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "pool_id", poolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err - } - - var pathParam2 string - - pathParam2, err = runtime.StyleParamWithOptions("simple", false, "volume_id", volumeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err - } - - serverURL, err := url.Parse(server) - if err != nil { - return nil, err - } - - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/", pathParam0, pathParam1, pathParam2) - if operationPath[0] == '/' { - operationPath = "." + operationPath - } - - queryURL, err := serverURL.Parse(operationPath) - if err != nil { - return nil, err - } - - req, err := http.NewRequest(http.MethodDelete, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodDelete, queryURL.String(), nil) if err != nil { return nil, err } @@ -9562,8 +10121,8 @@ func NewClustersStoragePoolsVolumesDeleteApiV2ClustersClusterIdStoragePoolsPoolI return req, nil } -// NewClustersStoragePoolsVolumesDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdGetRequest constructs an http.Request for the ClustersStoragePoolsVolumesDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdGet method -func NewClustersStoragePoolsVolumesDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, params *ClustersStoragePoolsVolumesDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdGetParams) (*http.Request, error) { +// NewClustersStoragePoolsDetailApiV2ClustersClusterIdStoragePoolsPoolIdGetRequest constructs an http.Request for the ClustersStoragePoolsDetailApiV2ClustersClusterIdStoragePoolsPoolIdGet method +func NewClustersStoragePoolsDetailApiV2ClustersClusterIdStoragePoolsPoolIdGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, params *ClustersStoragePoolsDetailApiV2ClustersClusterIdStoragePoolsPoolIdGetParams) (*http.Request, error) { var err error var pathParam0 string @@ -9580,19 +10139,12 @@ func NewClustersStoragePoolsVolumesDetailApiV2ClustersClusterIdStoragePoolsPoolI return nil, err } - var pathParam2 string - - pathParam2, err = runtime.StyleParamWithOptions("simple", false, "volume_id", volumeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err - } - serverURL, err := url.Parse(server) if err != nil { return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -9637,19 +10189,19 @@ func NewClustersStoragePoolsVolumesDetailApiV2ClustersClusterIdStoragePoolsPoolI return req, nil } -// NewClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPutRequest calls the generic ClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPut builder with application/json body -func NewClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPutRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, body ClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPutJSONRequestBody) (*http.Request, error) { +// NewClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutRequest calls the generic ClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPut builder with application/json body +func NewClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, body ClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutJSONRequestBody) (*http.Request, error) { var bodyReader io.Reader buf, err := json.Marshal(body) if err != nil { return nil, err } bodyReader = bytes.NewReader(buf) - return NewClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPutRequestWithBody(server, clusterId, poolId, volumeId, "application/json", bodyReader) + return NewClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutRequestWithBody(server, clusterId, poolId, "application/json", bodyReader) } -// NewClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPutRequestWithBody constructs an http.Request for the ClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPut method, with any body, and a specified content type -func NewClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPutRequestWithBody(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { +// NewClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutRequestWithBody constructs an http.Request for the ClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPut method, with any body, and a specified content type +func NewClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutRequestWithBody(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { var err error var pathParam0 string @@ -9666,19 +10218,12 @@ func NewClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolI return nil, err } - var pathParam2 string - - pathParam2, err = runtime.StyleParamWithOptions("simple", false, "volume_id", volumeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err - } - serverURL, err := url.Parse(server) if err != nil { return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -9698,8 +10243,19 @@ func NewClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolI return req, nil } -// NewClustersStoragePoolsVolumesBackupsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdBackupsDeleteRequest constructs an http.Request for the ClustersStoragePoolsVolumesBackupsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdBackupsDelete method -func NewClustersStoragePoolsVolumesBackupsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdBackupsDeleteRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID) (*http.Request, error) { +// NewClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDeleteRequest calls the generic ClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDelete builder with application/json body +func NewClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDeleteRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, body ClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDeleteJSONRequestBody) (*http.Request, error) { + var bodyReader io.Reader + buf, err := json.Marshal(body) + if err != nil { + return nil, err + } + bodyReader = bytes.NewReader(buf) + return NewClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDeleteRequestWithBody(server, clusterId, poolId, "application/json", bodyReader) +} + +// NewClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDeleteRequestWithBody constructs an http.Request for the ClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDelete method, with any body, and a specified content type +func NewClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDeleteRequestWithBody(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { var err error var pathParam0 string @@ -9716,19 +10272,12 @@ func NewClustersStoragePoolsVolumesBackupsDeleteApiV2ClustersClusterIdStoragePoo return nil, err } - var pathParam2 string - - pathParam2, err = runtime.StyleParamWithOptions("simple", false, "volume_id", volumeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err - } - serverURL, err := url.Parse(server) if err != nil { return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/backups", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/host", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -9738,16 +10287,29 @@ func NewClustersStoragePoolsVolumesBackupsDeleteApiV2ClustersClusterIdStoragePoo return nil, err } - req, err := http.NewRequest(http.MethodDelete, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodDelete, queryURL.String(), body) if err != nil { return nil, err } + req.Header.Add("Content-Type", contentType) + return req, nil } -// NewClustersStoragePoolsVolumesBackupsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdBackupsGetRequest constructs an http.Request for the ClustersStoragePoolsVolumesBackupsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdBackupsGet method -func NewClustersStoragePoolsVolumesBackupsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdBackupsGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID) (*http.Request, error) { +// NewClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPostRequest calls the generic ClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPost builder with application/json body +func NewClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, body ClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPostJSONRequestBody) (*http.Request, error) { + var bodyReader io.Reader + buf, err := json.Marshal(body) + if err != nil { + return nil, err + } + bodyReader = bytes.NewReader(buf) + return NewClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPostRequestWithBody(server, clusterId, poolId, "application/json", bodyReader) +} + +// NewClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPostRequestWithBody constructs an http.Request for the ClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPost method, with any body, and a specified content type +func NewClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPostRequestWithBody(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { var err error var pathParam0 string @@ -9764,19 +10326,12 @@ func NewClustersStoragePoolsVolumesBackupsListApiV2ClustersClusterIdStoragePools return nil, err } - var pathParam2 string - - pathParam2, err = runtime.StyleParamWithOptions("simple", false, "volume_id", volumeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err - } - serverURL, err := url.Parse(server) if err != nil { return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/backups", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/host", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -9786,16 +10341,18 @@ func NewClustersStoragePoolsVolumesBackupsListApiV2ClustersClusterIdStoragePools return nil, err } - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) if err != nil { return nil, err } + req.Header.Add("Content-Type", contentType) + return req, nil } -// NewClustersStoragePoolsVolumesCapacityApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdCapacityGetRequest constructs an http.Request for the ClustersStoragePoolsVolumesCapacityApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdCapacityGet method -func NewClustersStoragePoolsVolumesCapacityApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdCapacityGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, params *ClustersStoragePoolsVolumesCapacityApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdCapacityGetParams) (*http.Request, error) { +// NewClustersStoragePoolsIostatsApiV2ClustersClusterIdStoragePoolsPoolIdIostatsGetRequest constructs an http.Request for the ClustersStoragePoolsIostatsApiV2ClustersClusterIdStoragePoolsPoolIdIostatsGet method +func NewClustersStoragePoolsIostatsApiV2ClustersClusterIdStoragePoolsPoolIdIostatsGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, params *ClustersStoragePoolsIostatsApiV2ClustersClusterIdStoragePoolsPoolIdIostatsGetParams) (*http.Request, error) { var err error var pathParam0 string @@ -9812,19 +10369,12 @@ func NewClustersStoragePoolsVolumesCapacityApiV2ClustersClusterIdStoragePoolsPoo return nil, err } - var pathParam2 string - - pathParam2, err = runtime.StyleParamWithOptions("simple", false, "volume_id", volumeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err - } - serverURL, err := url.Parse(server) if err != nil { return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/capacity", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/iostats", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -9843,9 +10393,9 @@ func NewClustersStoragePoolsVolumesCapacityApiV2ClustersClusterIdStoragePoolsPoo // per the OpenAPI spec (e.g. "color=blue,black,brown"). var rawQueryFragments []string - if params.History != nil { + if params.Limit != nil { - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "history", *params.History, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "", Format: ""}); err != nil { + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "limit", *params.Limit, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "integer", Format: ""}); err != nil { return nil, err } else { for _, qp := range strings.Split(queryFrag, "&") { @@ -9869,8 +10419,8 @@ func NewClustersStoragePoolsVolumesCapacityApiV2ClustersClusterIdStoragePoolsPoo return req, nil } -// NewClustersStoragePoolsVolumesCloneApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdClonePostRequest constructs an http.Request for the ClustersStoragePoolsVolumesCloneApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdClonePost method -func NewClustersStoragePoolsVolumesCloneApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdClonePostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, params *ClustersStoragePoolsVolumesCloneApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdClonePostParams) (*http.Request, error) { +// NewClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsGetRequest constructs an http.Request for the ClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsGet method +func NewClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, params *ClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsGetParams) (*http.Request, error) { var err error var pathParam0 string @@ -9887,19 +10437,12 @@ func NewClustersStoragePoolsVolumesCloneApiV2ClustersClusterIdStoragePoolsPoolId return nil, err } - var pathParam2 string - - pathParam2, err = runtime.StyleParamWithOptions("simple", false, "volume_id", volumeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err - } - serverURL, err := url.Parse(server) if err != nil { return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/clone", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/snapshots/", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -9918,17 +10461,9 @@ func NewClustersStoragePoolsVolumesCloneApiV2ClustersClusterIdStoragePoolsPoolId // per the OpenAPI spec (e.g. "color=blue,black,brown"). var rawQueryFragments []string - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "clone_name", params.CloneName, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "string", Format: ""}); err != nil { - return nil, err - } else { - for _, qp := range strings.Split(queryFrag, "&") { - rawQueryFragments = append(rawQueryFragments, qp) - } - } - - if params.NewSize != nil { + if params.Watch != nil { - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "new_size", *params.NewSize, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "", Format: ""}); err != nil { + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "watch", *params.Watch, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { return nil, err } else { for _, qp := range strings.Split(queryFrag, "&") { @@ -9938,9 +10473,9 @@ func NewClustersStoragePoolsVolumesCloneApiV2ClustersClusterIdStoragePoolsPoolId } - if params.PvcName != nil { + if params.ConsistencyGroup != nil { - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "pvc_name", *params.PvcName, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "", Format: ""}); err != nil { + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "consistency_group", *params.ConsistencyGroup, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "", Format: ""}); err != nil { return nil, err } else { for _, qp := range strings.Split(queryFrag, "&") { @@ -9956,7 +10491,7 @@ func NewClustersStoragePoolsVolumesCloneApiV2ClustersClusterIdStoragePoolsPoolId queryURL.RawQuery = strings.Join(rawQueryFragments, "&") } - req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err } @@ -9964,8 +10499,8 @@ func NewClustersStoragePoolsVolumesCloneApiV2ClustersClusterIdStoragePoolsPoolId return req, nil } -// NewClustersStoragePoolsVolumesConnectApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdConnectGetRequest constructs an http.Request for the ClustersStoragePoolsVolumesConnectApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdConnectGet method -func NewClustersStoragePoolsVolumesConnectApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdConnectGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, params *ClustersStoragePoolsVolumesConnectApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdConnectGetParams) (*http.Request, error) { +// NewClustersStoragePoolsSnapshotsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdDeleteRequest constructs an http.Request for the ClustersStoragePoolsSnapshotsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdDelete method +func NewClustersStoragePoolsSnapshotsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdDeleteRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, snapshotId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -9984,7 +10519,7 @@ func NewClustersStoragePoolsVolumesConnectApiV2ClustersClusterIdStoragePoolsPool var pathParam2 string - pathParam2, err = runtime.StyleParamWithOptions("simple", false, "volume_id", volumeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "snapshot_id", snapshotId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -9994,7 +10529,7 @@ func NewClustersStoragePoolsVolumesConnectApiV2ClustersClusterIdStoragePoolsPool return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/connect", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/snapshots/%s/", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -10004,34 +10539,7 @@ func NewClustersStoragePoolsVolumesConnectApiV2ClustersClusterIdStoragePoolsPool return nil, err } - if params != nil { - // queryValues collects non-styled parameters (passthrough, JSON) - // that are safe to round-trip through url.Values.Encode(). - queryValues := queryURL.Query() - // rawQueryFragments collects pre-encoded query fragments from - // styled parameters, preserving literal commas as delimiters - // per the OpenAPI spec (e.g. "color=blue,black,brown"). - var rawQueryFragments []string - - if params.HostNqn != nil { - - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "host_nqn", *params.HostNqn, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "", Format: ""}); err != nil { - return nil, err - } else { - for _, qp := range strings.Split(queryFrag, "&") { - rawQueryFragments = append(rawQueryFragments, qp) - } - } - - } - - if encoded := queryValues.Encode(); encoded != "" { - rawQueryFragments = append(rawQueryFragments, encoded) - } - queryURL.RawQuery = strings.Join(rawQueryFragments, "&") - } - - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodDelete, queryURL.String(), nil) if err != nil { return nil, err } @@ -10039,19 +10547,8 @@ func NewClustersStoragePoolsVolumesConnectApiV2ClustersClusterIdStoragePoolsPool return req, nil } -// NewClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPostRequest calls the generic ClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPost builder with application/json body -func NewClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, body ClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPostJSONRequestBody) (*http.Request, error) { - var bodyReader io.Reader - buf, err := json.Marshal(body) - if err != nil { - return nil, err - } - bodyReader = bytes.NewReader(buf) - return NewClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPostRequestWithBody(server, clusterId, poolId, volumeId, "application/json", bodyReader) -} - -// NewClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPostRequestWithBody constructs an http.Request for the ClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPost method, with any body, and a specified content type -func NewClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPostRequestWithBody(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { +// NewClustersStoragePoolsSnapshotsDetailApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdGetRequest constructs an http.Request for the ClustersStoragePoolsSnapshotsDetailApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdGet method +func NewClustersStoragePoolsSnapshotsDetailApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, snapshotId openapi_types.UUID, params *ClustersStoragePoolsSnapshotsDetailApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdGetParams) (*http.Request, error) { var err error var pathParam0 string @@ -10070,7 +10567,7 @@ func NewClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPool var pathParam2 string - pathParam2, err = runtime.StyleParamWithOptions("simple", false, "volume_id", volumeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "snapshot_id", snapshotId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -10080,7 +10577,7 @@ func NewClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPool return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/hosts", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/snapshots/%s/", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -10090,44 +10587,55 @@ func NewClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPool return nil, err } - req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) - if err != nil { - return nil, err - } - - req.Header.Add("Content-Type", contentType) + if params != nil { + // queryValues collects non-styled parameters (passthrough, JSON) + // that are safe to round-trip through url.Values.Encode(). + queryValues := queryURL.Query() + // rawQueryFragments collects pre-encoded query fragments from + // styled parameters, preserving literal commas as delimiters + // per the OpenAPI spec (e.g. "color=blue,black,brown"). + var rawQueryFragments []string - return req, nil -} + if params.Watch != nil { -// NewClustersStoragePoolsVolumesRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsHostNqnDeleteRequest constructs an http.Request for the ClustersStoragePoolsVolumesRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsHostNqnDelete method -func NewClustersStoragePoolsVolumesRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsHostNqnDeleteRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, hostNqn string) (*http.Request, error) { - var err error + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "watch", *params.Watch, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } + } - var pathParam0 string + } - pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) - if err != nil { - return nil, err + if encoded := queryValues.Encode(); encoded != "" { + rawQueryFragments = append(rawQueryFragments, encoded) + } + queryURL.RawQuery = strings.Join(rawQueryFragments, "&") } - var pathParam1 string - - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "pool_id", poolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err } - var pathParam2 string + return req, nil +} - pathParam2, err = runtime.StyleParamWithOptions("simple", false, "volume_id", volumeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) +// NewClustersStoragePoolsVolumesListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesGetRequest constructs an http.Request for the ClustersStoragePoolsVolumesListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesGet method +func NewClustersStoragePoolsVolumesListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, params *ClustersStoragePoolsVolumesListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesGetParams) (*http.Request, error) { + var err error + + var pathParam0 string + + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } - var pathParam3 string + var pathParam1 string - pathParam3, err = runtime.StyleParamWithOptions("simple", false, "host_nqn", hostNqn, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: ""}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "pool_id", poolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -10137,7 +10645,7 @@ func NewClustersStoragePoolsVolumesRemoveHostApiV2ClustersClusterIdStoragePoolsP return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/hosts/%s", pathParam0, pathParam1, pathParam2, pathParam3) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -10147,7 +10655,34 @@ func NewClustersStoragePoolsVolumesRemoveHostApiV2ClustersClusterIdStoragePoolsP return nil, err } - req, err := http.NewRequest(http.MethodDelete, queryURL.String(), nil) + if params != nil { + // queryValues collects non-styled parameters (passthrough, JSON) + // that are safe to round-trip through url.Values.Encode(). + queryValues := queryURL.Query() + // rawQueryFragments collects pre-encoded query fragments from + // styled parameters, preserving literal commas as delimiters + // per the OpenAPI spec (e.g. "color=blue,black,brown"). + var rawQueryFragments []string + + if params.Watch != nil { + + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "watch", *params.Watch, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } + } + + } + + if encoded := queryValues.Encode(); encoded != "" { + rawQueryFragments = append(rawQueryFragments, encoded) + } + queryURL.RawQuery = strings.Join(rawQueryFragments, "&") + } + + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err } @@ -10155,8 +10690,19 @@ func NewClustersStoragePoolsVolumesRemoveHostApiV2ClustersClusterIdStoragePoolsP return req, nil } -// NewClustersStoragePoolsVolumesGetHostSecretApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsHostNqnSecretGetRequest constructs an http.Request for the ClustersStoragePoolsVolumesGetHostSecretApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsHostNqnSecretGet method -func NewClustersStoragePoolsVolumesGetHostSecretApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsHostNqnSecretGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, hostNqn string) (*http.Request, error) { +// NewClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostRequest calls the generic ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPost builder with application/json body +func NewClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, params *ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostParams, body ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostJSONRequestBody) (*http.Request, error) { + var bodyReader io.Reader + buf, err := json.Marshal(body) + if err != nil { + return nil, err + } + bodyReader = bytes.NewReader(buf) + return NewClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostRequestWithBody(server, clusterId, poolId, params, "application/json", bodyReader) +} + +// NewClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostRequestWithBody constructs an http.Request for the ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPost method, with any body, and a specified content type +func NewClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostRequestWithBody(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, params *ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostParams, contentType string, body io.Reader) (*http.Request, error) { var err error var pathParam0 string @@ -10173,16 +10719,83 @@ func NewClustersStoragePoolsVolumesGetHostSecretApiV2ClustersClusterIdStoragePoo return nil, err } - var pathParam2 string + serverURL, err := url.Parse(server) + if err != nil { + return nil, err + } - pathParam2, err = runtime.StyleParamWithOptions("simple", false, "volume_id", volumeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/", pathParam0, pathParam1) + if operationPath[0] == '/' { + operationPath = "." + operationPath + } + + queryURL, err := serverURL.Parse(operationPath) if err != nil { return nil, err } - var pathParam3 string + if params != nil { + // queryValues collects non-styled parameters (passthrough, JSON) + // that are safe to round-trip through url.Values.Encode(). + queryValues := queryURL.Query() + // rawQueryFragments collects pre-encoded query fragments from + // styled parameters, preserving literal commas as delimiters + // per the OpenAPI spec (e.g. "color=blue,black,brown"). + var rawQueryFragments []string - pathParam3, err = runtime.StyleParamWithOptions("simple", false, "host_nqn", hostNqn, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: ""}) + if params.ResponseFormat != nil { + + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "response-format", *params.ResponseFormat, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "string", Format: ""}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } + } + + } + + if encoded := queryValues.Encode(); encoded != "" { + rawQueryFragments = append(rawQueryFragments, encoded) + } + queryURL.RawQuery = strings.Join(rawQueryFragments, "&") + } + + req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) + if err != nil { + return nil, err + } + + req.Header.Add("Content-Type", contentType) + + return req, nil +} + +// NewClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPostRequest calls the generic ClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPost builder with application/json body +func NewClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, body ClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPostJSONRequestBody) (*http.Request, error) { + var bodyReader io.Reader + buf, err := json.Marshal(body) + if err != nil { + return nil, err + } + bodyReader = bytes.NewReader(buf) + return NewClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPostRequestWithBody(server, clusterId, poolId, "application/json", bodyReader) +} + +// NewClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPostRequestWithBody constructs an http.Request for the ClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPost method, with any body, and a specified content type +func NewClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPostRequestWithBody(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { + var err error + + var pathParam0 string + + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + var pathParam1 string + + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "pool_id", poolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -10192,7 +10805,7 @@ func NewClustersStoragePoolsVolumesGetHostSecretApiV2ClustersClusterIdStoragePoo return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/hosts/%s/secret", pathParam0, pathParam1, pathParam2, pathParam3) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/replicate_lvol_on_source_cluster", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -10202,16 +10815,18 @@ func NewClustersStoragePoolsVolumesGetHostSecretApiV2ClustersClusterIdStoragePoo return nil, err } - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) if err != nil { return nil, err } + req.Header.Add("Content-Type", contentType) + return req, nil } -// NewClustersStoragePoolsVolumesInflateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdInflatePostRequest constructs an http.Request for the ClustersStoragePoolsVolumesInflateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdInflatePost method -func NewClustersStoragePoolsVolumesInflateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdInflatePostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID) (*http.Request, error) { +// NewClustersStoragePoolsVolumesDeleteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdDeleteRequest constructs an http.Request for the ClustersStoragePoolsVolumesDeleteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdDelete method +func NewClustersStoragePoolsVolumesDeleteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdDeleteRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -10240,7 +10855,7 @@ func NewClustersStoragePoolsVolumesInflateApiV2ClustersClusterIdStoragePoolsPool return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/inflate", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -10250,7 +10865,7 @@ func NewClustersStoragePoolsVolumesInflateApiV2ClustersClusterIdStoragePoolsPool return nil, err } - req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodDelete, queryURL.String(), nil) if err != nil { return nil, err } @@ -10258,8 +10873,8 @@ func NewClustersStoragePoolsVolumesInflateApiV2ClustersClusterIdStoragePoolsPool return req, nil } -// NewClustersStoragePoolsVolumesIostatsApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdIostatsGetRequest constructs an http.Request for the ClustersStoragePoolsVolumesIostatsApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdIostatsGet method -func NewClustersStoragePoolsVolumesIostatsApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdIostatsGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, params *ClustersStoragePoolsVolumesIostatsApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdIostatsGetParams) (*http.Request, error) { +// NewClustersStoragePoolsVolumesDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdGetRequest constructs an http.Request for the ClustersStoragePoolsVolumesDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdGet method +func NewClustersStoragePoolsVolumesDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, params *ClustersStoragePoolsVolumesDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdGetParams) (*http.Request, error) { var err error var pathParam0 string @@ -10288,7 +10903,7 @@ func NewClustersStoragePoolsVolumesIostatsApiV2ClustersClusterIdStoragePoolsPool return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/iostats", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -10307,9 +10922,9 @@ func NewClustersStoragePoolsVolumesIostatsApiV2ClustersClusterIdStoragePoolsPool // per the OpenAPI spec (e.g. "color=blue,black,brown"). var rawQueryFragments []string - if params.History != nil { + if params.Watch != nil { - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "history", *params.History, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "", Format: ""}); err != nil { + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "watch", *params.Watch, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { return nil, err } else { for _, qp := range strings.Split(queryFrag, "&") { @@ -10333,8 +10948,19 @@ func NewClustersStoragePoolsVolumesIostatsApiV2ClustersClusterIdStoragePoolsPool return req, nil } -// NewClustersStoragePoolsVolumesReplicationDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationGetRequest constructs an http.Request for the ClustersStoragePoolsVolumesReplicationDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationGet method -func NewClustersStoragePoolsVolumesReplicationDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID) (*http.Request, error) { +// NewClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPutRequest calls the generic ClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPut builder with application/json body +func NewClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPutRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, body ClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPutJSONRequestBody) (*http.Request, error) { + var bodyReader io.Reader + buf, err := json.Marshal(body) + if err != nil { + return nil, err + } + bodyReader = bytes.NewReader(buf) + return NewClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPutRequestWithBody(server, clusterId, poolId, volumeId, "application/json", bodyReader) +} + +// NewClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPutRequestWithBody constructs an http.Request for the ClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPut method, with any body, and a specified content type +func NewClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPutRequestWithBody(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { var err error var pathParam0 string @@ -10363,7 +10989,7 @@ func NewClustersStoragePoolsVolumesReplicationDetailApiV2ClustersClusterIdStorag return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/replication/", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -10373,27 +10999,18 @@ func NewClustersStoragePoolsVolumesReplicationDetailApiV2ClustersClusterIdStorag return nil, err } - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodPut, queryURL.String(), body) if err != nil { return nil, err } - return req, nil -} + req.Header.Add("Content-Type", contentType) -// NewClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPostRequest calls the generic ClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPost builder with application/json body -func NewClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, body ClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPostJSONRequestBody) (*http.Request, error) { - var bodyReader io.Reader - buf, err := json.Marshal(body) - if err != nil { - return nil, err - } - bodyReader = bytes.NewReader(buf) - return NewClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPostRequestWithBody(server, clusterId, poolId, volumeId, "application/json", bodyReader) + return req, nil } -// NewClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPostRequestWithBody constructs an http.Request for the ClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPost method, with any body, and a specified content type -func NewClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPostRequestWithBody(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { +// NewClustersStoragePoolsVolumesBackupsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdBackupsDeleteRequest constructs an http.Request for the ClustersStoragePoolsVolumesBackupsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdBackupsDelete method +func NewClustersStoragePoolsVolumesBackupsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdBackupsDeleteRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -10422,7 +11039,7 @@ func NewClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStorag return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/replication/commit", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/backups", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -10432,18 +11049,16 @@ func NewClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStorag return nil, err } - req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) + req, err := http.NewRequest(http.MethodDelete, queryURL.String(), nil) if err != nil { return nil, err } - req.Header.Add("Content-Type", contentType) - return req, nil } -// NewClustersStoragePoolsVolumesReplicationCutoverProceedApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCutoverProceedPostRequest constructs an http.Request for the ClustersStoragePoolsVolumesReplicationCutoverProceedApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCutoverProceedPost method -func NewClustersStoragePoolsVolumesReplicationCutoverProceedApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCutoverProceedPostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID) (*http.Request, error) { +// NewClustersStoragePoolsVolumesBackupsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdBackupsGetRequest constructs an http.Request for the ClustersStoragePoolsVolumesBackupsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdBackupsGet method +func NewClustersStoragePoolsVolumesBackupsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdBackupsGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -10472,7 +11087,7 @@ func NewClustersStoragePoolsVolumesReplicationCutoverProceedApiV2ClustersCluster return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/replication/cutover-proceed", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/backups", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -10482,7 +11097,7 @@ func NewClustersStoragePoolsVolumesReplicationCutoverProceedApiV2ClustersCluster return nil, err } - req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err } @@ -10490,19 +11105,8 @@ func NewClustersStoragePoolsVolumesReplicationCutoverProceedApiV2ClustersCluster return req, nil } -// NewClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostRequest calls the generic ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPost builder with application/json body -func NewClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, body ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostJSONRequestBody) (*http.Request, error) { - var bodyReader io.Reader - buf, err := json.Marshal(body) - if err != nil { - return nil, err - } - bodyReader = bytes.NewReader(buf) - return NewClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostRequestWithBody(server, clusterId, poolId, volumeId, "application/json", bodyReader) -} - -// NewClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostRequestWithBody constructs an http.Request for the ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPost method, with any body, and a specified content type -func NewClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostRequestWithBody(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { +// NewClustersStoragePoolsVolumesCapacityApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdCapacityGetRequest constructs an http.Request for the ClustersStoragePoolsVolumesCapacityApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdCapacityGet method +func NewClustersStoragePoolsVolumesCapacityApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdCapacityGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, params *ClustersStoragePoolsVolumesCapacityApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdCapacityGetParams) (*http.Request, error) { var err error var pathParam0 string @@ -10531,7 +11135,7 @@ func NewClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStor return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/replication/failback", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/capacity", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -10541,18 +11145,43 @@ func NewClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStor return nil, err } - req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) - if err != nil { - return nil, err - } - - req.Header.Add("Content-Type", contentType) + if params != nil { + // queryValues collects non-styled parameters (passthrough, JSON) + // that are safe to round-trip through url.Values.Encode(). + queryValues := queryURL.Query() + // rawQueryFragments collects pre-encoded query fragments from + // styled parameters, preserving literal commas as delimiters + // per the OpenAPI spec (e.g. "color=blue,black,brown"). + var rawQueryFragments []string + + if params.History != nil { + + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "history", *params.History, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "", Format: ""}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } + } + + } + + if encoded := queryValues.Encode(); encoded != "" { + rawQueryFragments = append(rawQueryFragments, encoded) + } + queryURL.RawQuery = strings.Join(rawQueryFragments, "&") + } + + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + if err != nil { + return nil, err + } return req, nil } -// NewClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPostRequest constructs an http.Request for the ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPost method -func NewClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, params *ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPostParams) (*http.Request, error) { +// NewClustersStoragePoolsVolumesCloneApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdClonePostRequest constructs an http.Request for the ClustersStoragePoolsVolumesCloneApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdClonePost method +func NewClustersStoragePoolsVolumesCloneApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdClonePostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, params *ClustersStoragePoolsVolumesCloneApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdClonePostParams) (*http.Request, error) { var err error var pathParam0 string @@ -10581,7 +11210,7 @@ func NewClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStor return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/replication/failover", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/clone", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -10600,9 +11229,29 @@ func NewClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStor // per the OpenAPI spec (e.g. "color=blue,black,brown"). var rawQueryFragments []string - if params.Generation != nil { + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "clone_name", params.CloneName, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "string", Format: ""}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } + } - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "generation", *params.Generation, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "integer", Format: ""}); err != nil { + if params.NewSize != nil { + + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "new_size", *params.NewSize, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "", Format: ""}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } + } + + } + + if params.PvcName != nil { + + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "pvc_name", *params.PvcName, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "", Format: ""}); err != nil { return nil, err } else { for _, qp := range strings.Split(queryFrag, "&") { @@ -10626,19 +11275,8 @@ func NewClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStor return req, nil } -// NewClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostRequest calls the generic ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPost builder with application/json body -func NewClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, body ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostJSONRequestBody) (*http.Request, error) { - var bodyReader io.Reader - buf, err := json.Marshal(body) - if err != nil { - return nil, err - } - bodyReader = bytes.NewReader(buf) - return NewClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostRequestWithBody(server, clusterId, poolId, volumeId, "application/json", bodyReader) -} - -// NewClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostRequestWithBody constructs an http.Request for the ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPost method, with any body, and a specified content type -func NewClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostRequestWithBody(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { +// NewClustersStoragePoolsVolumesConnectApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdConnectGetRequest constructs an http.Request for the ClustersStoragePoolsVolumesConnectApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdConnectGet method +func NewClustersStoragePoolsVolumesConnectApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdConnectGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, params *ClustersStoragePoolsVolumesConnectApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdConnectGetParams) (*http.Request, error) { var err error var pathParam0 string @@ -10667,7 +11305,7 @@ func NewClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStorage return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/replication/start", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/connect", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -10677,18 +11315,54 @@ func NewClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStorage return nil, err } - req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) + if params != nil { + // queryValues collects non-styled parameters (passthrough, JSON) + // that are safe to round-trip through url.Values.Encode(). + queryValues := queryURL.Query() + // rawQueryFragments collects pre-encoded query fragments from + // styled parameters, preserving literal commas as delimiters + // per the OpenAPI spec (e.g. "color=blue,black,brown"). + var rawQueryFragments []string + + if params.HostNqn != nil { + + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "host_nqn", *params.HostNqn, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "", Format: ""}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } + } + + } + + if encoded := queryValues.Encode(); encoded != "" { + rawQueryFragments = append(rawQueryFragments, encoded) + } + queryURL.RawQuery = strings.Join(rawQueryFragments, "&") + } + + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err } - req.Header.Add("Content-Type", contentType) - return req, nil } -// NewClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPostRequest constructs an http.Request for the ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPost method -func NewClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID) (*http.Request, error) { +// NewClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPostRequest calls the generic ClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPost builder with application/json body +func NewClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, body ClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPostJSONRequestBody) (*http.Request, error) { + var bodyReader io.Reader + buf, err := json.Marshal(body) + if err != nil { + return nil, err + } + bodyReader = bytes.NewReader(buf) + return NewClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPostRequestWithBody(server, clusterId, poolId, volumeId, "application/json", bodyReader) +} + +// NewClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPostRequestWithBody constructs an http.Request for the ClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPost method, with any body, and a specified content type +func NewClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPostRequestWithBody(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { var err error var pathParam0 string @@ -10717,7 +11391,7 @@ func NewClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStorageP return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/replication/stop", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/hosts", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -10727,16 +11401,18 @@ func NewClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStorageP return nil, err } - req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) if err != nil { return nil, err } + req.Header.Add("Content-Type", contentType) + return req, nil } -// NewClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTasksGetRequest constructs an http.Request for the ClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTasksGet method -func NewClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTasksGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID) (*http.Request, error) { +// NewClustersStoragePoolsVolumesRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsHostNqnDeleteRequest constructs an http.Request for the ClustersStoragePoolsVolumesRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsHostNqnDelete method +func NewClustersStoragePoolsVolumesRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsHostNqnDeleteRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, hostNqn string) (*http.Request, error) { var err error var pathParam0 string @@ -10760,12 +11436,19 @@ func NewClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStorage return nil, err } + var pathParam3 string + + pathParam3, err = runtime.StyleParamWithOptions("simple", false, "host_nqn", hostNqn, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: ""}) + if err != nil { + return nil, err + } + serverURL, err := url.Parse(server) if err != nil { return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/replication/tasks", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/hosts/%s", pathParam0, pathParam1, pathParam2, pathParam3) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -10775,7 +11458,7 @@ func NewClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStorage return nil, err } - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodDelete, queryURL.String(), nil) if err != nil { return nil, err } @@ -10783,8 +11466,8 @@ func NewClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStorage return req, nil } -// NewClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTriggerPostRequest constructs an http.Request for the ClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTriggerPost method -func NewClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTriggerPostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID) (*http.Request, error) { +// NewClustersStoragePoolsVolumesGetHostSecretApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsHostNqnSecretGetRequest constructs an http.Request for the ClustersStoragePoolsVolumesGetHostSecretApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsHostNqnSecretGet method +func NewClustersStoragePoolsVolumesGetHostSecretApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsHostNqnSecretGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, hostNqn string) (*http.Request, error) { var err error var pathParam0 string @@ -10808,12 +11491,19 @@ func NewClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStora return nil, err } + var pathParam3 string + + pathParam3, err = runtime.StyleParamWithOptions("simple", false, "host_nqn", hostNqn, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: ""}) + if err != nil { + return nil, err + } + serverURL, err := url.Parse(server) if err != nil { return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/replication/trigger", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/hosts/%s/secret", pathParam0, pathParam1, pathParam2, pathParam3) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -10823,7 +11513,7 @@ func NewClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStora return nil, err } - req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err } @@ -10831,8 +11521,8 @@ func NewClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStora return req, nil } -// NewClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsGetRequest constructs an http.Request for the ClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsGet method -func NewClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID) (*http.Request, error) { +// NewClustersStoragePoolsVolumesInflateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdInflatePostRequest constructs an http.Request for the ClustersStoragePoolsVolumesInflateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdInflatePost method +func NewClustersStoragePoolsVolumesInflateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdInflatePostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -10861,7 +11551,7 @@ func NewClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoo return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/snapshots", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/inflate", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -10871,7 +11561,7 @@ func NewClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoo return nil, err } - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) if err != nil { return nil, err } @@ -10879,19 +11569,8 @@ func NewClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoo return req, nil } -// NewClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostRequest calls the generic ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPost builder with application/json body -func NewClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, body ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostJSONRequestBody) (*http.Request, error) { - var bodyReader io.Reader - buf, err := json.Marshal(body) - if err != nil { - return nil, err - } - bodyReader = bytes.NewReader(buf) - return NewClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostRequestWithBody(server, clusterId, poolId, volumeId, "application/json", bodyReader) -} - -// NewClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostRequestWithBody constructs an http.Request for the ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPost method, with any body, and a specified content type -func NewClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostRequestWithBody(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { +// NewClustersStoragePoolsVolumesIostatsApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdIostatsGetRequest constructs an http.Request for the ClustersStoragePoolsVolumesIostatsApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdIostatsGet method +func NewClustersStoragePoolsVolumesIostatsApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdIostatsGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, params *ClustersStoragePoolsVolumesIostatsApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdIostatsGetParams) (*http.Request, error) { var err error var pathParam0 string @@ -10920,7 +11599,7 @@ func NewClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStorageP return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/snapshots", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/iostats", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -10930,18 +11609,43 @@ func NewClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStorageP return nil, err } - req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) + if params != nil { + // queryValues collects non-styled parameters (passthrough, JSON) + // that are safe to round-trip through url.Values.Encode(). + queryValues := queryURL.Query() + // rawQueryFragments collects pre-encoded query fragments from + // styled parameters, preserving literal commas as delimiters + // per the OpenAPI spec (e.g. "color=blue,black,brown"). + var rawQueryFragments []string + + if params.History != nil { + + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "history", *params.History, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "", Format: ""}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } + } + + } + + if encoded := queryValues.Encode(); encoded != "" { + rawQueryFragments = append(rawQueryFragments, encoded) + } + queryURL.RawQuery = strings.Join(rawQueryFragments, "&") + } + + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err } - req.Header.Add("Content-Type", contentType) - return req, nil } -// NewClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigrationsGetRequest constructs an http.Request for the ClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigrationsGet method -func NewClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigrationsGetRequest(server string, clusterId openapi_types.UUID, nqn string) (*http.Request, error) { +// NewClustersStoragePoolsVolumesReplicationDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationGetRequest constructs an http.Request for the ClustersStoragePoolsVolumesReplicationDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationGet method +func NewClustersStoragePoolsVolumesReplicationDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -10953,7 +11657,14 @@ func NewClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigra var pathParam1 string - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "nqn", nqn, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: ""}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "pool_id", poolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + var pathParam2 string + + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "volume_id", volumeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -10963,7 +11674,7 @@ func NewClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigra return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/subsystems/%s/migrations/", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/replication/", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -10981,19 +11692,19 @@ func NewClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigra return req, nil } -// NewClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostRequest calls the generic ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPost builder with application/json body -func NewClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostRequest(server string, clusterId openapi_types.UUID, nqn string, params *ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostParams, body ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostJSONRequestBody) (*http.Request, error) { +// NewClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPostRequest calls the generic ClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPost builder with application/json body +func NewClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, body ClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPostJSONRequestBody) (*http.Request, error) { var bodyReader io.Reader buf, err := json.Marshal(body) if err != nil { return nil, err } bodyReader = bytes.NewReader(buf) - return NewClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostRequestWithBody(server, clusterId, nqn, params, "application/json", bodyReader) + return NewClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPostRequestWithBody(server, clusterId, poolId, volumeId, "application/json", bodyReader) } -// NewClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostRequestWithBody constructs an http.Request for the ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPost method, with any body, and a specified content type -func NewClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostRequestWithBody(server string, clusterId openapi_types.UUID, nqn string, params *ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostParams, contentType string, body io.Reader) (*http.Request, error) { +// NewClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPostRequestWithBody constructs an http.Request for the ClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPost method, with any body, and a specified content type +func NewClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPostRequestWithBody(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { var err error var pathParam0 string @@ -11005,7 +11716,14 @@ func NewClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMig var pathParam1 string - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "nqn", nqn, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: ""}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "pool_id", poolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + var pathParam2 string + + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "volume_id", volumeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -11015,7 +11733,7 @@ func NewClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMig return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/subsystems/%s/migrations/", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/replication/commit", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -11025,33 +11743,6 @@ func NewClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMig return nil, err } - if params != nil { - // queryValues collects non-styled parameters (passthrough, JSON) - // that are safe to round-trip through url.Values.Encode(). - queryValues := queryURL.Query() - // rawQueryFragments collects pre-encoded query fragments from - // styled parameters, preserving literal commas as delimiters - // per the OpenAPI spec (e.g. "color=blue,black,brown"). - var rawQueryFragments []string - - if params.ResponseFormat != nil { - - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "response-format", *params.ResponseFormat, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "string", Format: ""}); err != nil { - return nil, err - } else { - for _, qp := range strings.Split(queryFrag, "&") { - rawQueryFragments = append(rawQueryFragments, qp) - } - } - - } - - if encoded := queryValues.Encode(); encoded != "" { - rawQueryFragments = append(rawQueryFragments, encoded) - } - queryURL.RawQuery = strings.Join(rawQueryFragments, "&") - } - req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) if err != nil { return nil, err @@ -11062,8 +11753,8 @@ func NewClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMig return req, nil } -// NewClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdDeleteRequest constructs an http.Request for the ClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdDelete method -func NewClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdDeleteRequest(server string, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID) (*http.Request, error) { +// NewClustersStoragePoolsVolumesReplicationCutoverProceedApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCutoverProceedPostRequest constructs an http.Request for the ClustersStoragePoolsVolumesReplicationCutoverProceedApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCutoverProceedPost method +func NewClustersStoragePoolsVolumesReplicationCutoverProceedApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCutoverProceedPostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -11075,14 +11766,14 @@ func NewClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMig var pathParam1 string - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "nqn", nqn, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: ""}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "pool_id", poolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } var pathParam2 string - pathParam2, err = runtime.StyleParamWithOptions("simple", false, "migration_id", migrationId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "volume_id", volumeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -11092,7 +11783,7 @@ func NewClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMig return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/subsystems/%s/migrations/%s/", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/replication/cutover-proceed", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -11102,7 +11793,7 @@ func NewClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMig return nil, err } - req, err := http.NewRequest(http.MethodDelete, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) if err != nil { return nil, err } @@ -11110,8 +11801,19 @@ func NewClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMig return req, nil } -// NewClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdGetRequest constructs an http.Request for the ClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdGet method -func NewClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdGetRequest(server string, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID) (*http.Request, error) { +// NewClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostRequest calls the generic ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPost builder with application/json body +func NewClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, body ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostJSONRequestBody) (*http.Request, error) { + var bodyReader io.Reader + buf, err := json.Marshal(body) + if err != nil { + return nil, err + } + bodyReader = bytes.NewReader(buf) + return NewClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostRequestWithBody(server, clusterId, poolId, volumeId, "application/json", bodyReader) +} + +// NewClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostRequestWithBody constructs an http.Request for the ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPost method, with any body, and a specified content type +func NewClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostRequestWithBody(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { var err error var pathParam0 string @@ -11123,14 +11825,14 @@ func NewClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMig var pathParam1 string - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "nqn", nqn, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: ""}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "pool_id", poolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } var pathParam2 string - pathParam2, err = runtime.StyleParamWithOptions("simple", false, "migration_id", migrationId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "volume_id", volumeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -11140,7 +11842,7 @@ func NewClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMig return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/subsystems/%s/migrations/%s/", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/replication/failback", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -11150,16 +11852,18 @@ func NewClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMig return nil, err } - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) if err != nil { return nil, err } + req.Header.Add("Content-Type", contentType) + return req, nil } -// NewClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdCleanupTargetPostRequest constructs an http.Request for the ClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdCleanupTargetPost method -func NewClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdCleanupTargetPostRequest(server string, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID) (*http.Request, error) { +// NewClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPostRequest constructs an http.Request for the ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPost method +func NewClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, params *ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPostParams) (*http.Request, error) { var err error var pathParam0 string @@ -11171,14 +11875,14 @@ func NewClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystem var pathParam1 string - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "nqn", nqn, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: ""}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "pool_id", poolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } var pathParam2 string - pathParam2, err = runtime.StyleParamWithOptions("simple", false, "migration_id", migrationId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "volume_id", volumeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -11188,7 +11892,7 @@ func NewClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystem return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/subsystems/%s/migrations/%s/cleanup-target", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/replication/failover", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -11198,6 +11902,33 @@ func NewClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystem return nil, err } + if params != nil { + // queryValues collects non-styled parameters (passthrough, JSON) + // that are safe to round-trip through url.Values.Encode(). + queryValues := queryURL.Query() + // rawQueryFragments collects pre-encoded query fragments from + // styled parameters, preserving literal commas as delimiters + // per the OpenAPI spec (e.g. "color=blue,black,brown"). + var rawQueryFragments []string + + if params.Generation != nil { + + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "generation", *params.Generation, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "integer", Format: ""}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } + } + + } + + if encoded := queryValues.Encode(); encoded != "" { + rawQueryFragments = append(rawQueryFragments, encoded) + } + queryURL.RawQuery = strings.Join(rawQueryFragments, "&") + } + req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) if err != nil { return nil, err @@ -11206,19 +11937,19 @@ func NewClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystem return req, nil } -// NewClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostRequest calls the generic ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePost builder with application/json body -func NewClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostRequest(server string, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID, body ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostJSONRequestBody) (*http.Request, error) { +// NewClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostRequest calls the generic ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPost builder with application/json body +func NewClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, body ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostJSONRequestBody) (*http.Request, error) { var bodyReader io.Reader buf, err := json.Marshal(body) if err != nil { return nil, err } bodyReader = bytes.NewReader(buf) - return NewClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostRequestWithBody(server, clusterId, nqn, migrationId, "application/json", bodyReader) + return NewClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostRequestWithBody(server, clusterId, poolId, volumeId, "application/json", bodyReader) } -// NewClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostRequestWithBody constructs an http.Request for the ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePost method, with any body, and a specified content type -func NewClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostRequestWithBody(server string, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { +// NewClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostRequestWithBody constructs an http.Request for the ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPost method, with any body, and a specified content type +func NewClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostRequestWithBody(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { var err error var pathParam0 string @@ -11230,14 +11961,14 @@ func NewClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnM var pathParam1 string - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "nqn", nqn, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: ""}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "pool_id", poolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } var pathParam2 string - pathParam2, err = runtime.StyleParamWithOptions("simple", false, "migration_id", migrationId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "volume_id", volumeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -11247,7 +11978,7 @@ func NewClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnM return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/subsystems/%s/migrations/%s/continue", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/replication/start", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -11267,8 +11998,8 @@ func NewClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnM return req, nil } -// NewClustersTasksListApiV2ClustersClusterIdTasksGetRequest constructs an http.Request for the ClustersTasksListApiV2ClustersClusterIdTasksGet method -func NewClustersTasksListApiV2ClustersClusterIdTasksGetRequest(server string, clusterId openapi_types.UUID, params *ClustersTasksListApiV2ClustersClusterIdTasksGetParams) (*http.Request, error) { +// NewClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGetRequest constructs an http.Request for the ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGet method +func NewClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -11278,12 +12009,26 @@ func NewClustersTasksListApiV2ClustersClusterIdTasksGetRequest(server string, cl return nil, err } + var pathParam1 string + + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "pool_id", poolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + var pathParam2 string + + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "volume_id", volumeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + serverURL, err := url.Parse(server) if err != nil { return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/tasks/", pathParam0) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/replication/status", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -11293,33 +12038,6 @@ func NewClustersTasksListApiV2ClustersClusterIdTasksGetRequest(server string, cl return nil, err } - if params != nil { - // queryValues collects non-styled parameters (passthrough, JSON) - // that are safe to round-trip through url.Values.Encode(). - queryValues := queryURL.Query() - // rawQueryFragments collects pre-encoded query fragments from - // styled parameters, preserving literal commas as delimiters - // per the OpenAPI spec (e.g. "color=blue,black,brown"). - var rawQueryFragments []string - - if params.Watch != nil { - - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "watch", *params.Watch, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { - return nil, err - } else { - for _, qp := range strings.Split(queryFrag, "&") { - rawQueryFragments = append(rawQueryFragments, qp) - } - } - - } - - if encoded := queryValues.Encode(); encoded != "" { - rawQueryFragments = append(rawQueryFragments, encoded) - } - queryURL.RawQuery = strings.Join(rawQueryFragments, "&") - } - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err @@ -11328,8 +12046,8 @@ func NewClustersTasksListApiV2ClustersClusterIdTasksGetRequest(server string, cl return req, nil } -// NewClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGetRequest constructs an http.Request for the ClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGet method -func NewClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGetRequest(server string, clusterId openapi_types.UUID, taskId openapi_types.UUID, params *ClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGetParams) (*http.Request, error) { +// NewClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPostRequest constructs an http.Request for the ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPost method +func NewClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -11341,7 +12059,14 @@ func NewClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGetRequest(server st var pathParam1 string - pathParam1, err = runtime.StyleParamWithOptions("simple", false, "task_id", taskId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "pool_id", poolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + var pathParam2 string + + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "volume_id", volumeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -11351,7 +12076,7 @@ func NewClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGetRequest(server st return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/tasks/%s/", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/replication/stop", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -11361,59 +12086,35 @@ func NewClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGetRequest(server st return nil, err } - if params != nil { - // queryValues collects non-styled parameters (passthrough, JSON) - // that are safe to round-trip through url.Values.Encode(). - queryValues := queryURL.Query() - // rawQueryFragments collects pre-encoded query fragments from - // styled parameters, preserving literal commas as delimiters - // per the OpenAPI spec (e.g. "color=blue,black,brown"). - var rawQueryFragments []string - - if params.Watch != nil { + req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) + if err != nil { + return nil, err + } - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "watch", *params.Watch, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { - return nil, err - } else { - for _, qp := range strings.Split(queryFrag, "&") { - rawQueryFragments = append(rawQueryFragments, qp) - } - } + return req, nil +} - } +// NewClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTasksGetRequest constructs an http.Request for the ClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTasksGet method +func NewClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTasksGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID) (*http.Request, error) { + var err error - if encoded := queryValues.Encode(); encoded != "" { - rawQueryFragments = append(rawQueryFragments, encoded) - } - queryURL.RawQuery = strings.Join(rawQueryFragments, "&") - } + var pathParam0 string - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } - return req, nil -} + var pathParam1 string -// NewClustersUpgradeApiV2ClustersClusterIdUpdatePostRequest calls the generic ClustersUpgradeApiV2ClustersClusterIdUpdatePost builder with application/json body -func NewClustersUpgradeApiV2ClustersClusterIdUpdatePostRequest(server string, clusterId openapi_types.UUID, body ClustersUpgradeApiV2ClustersClusterIdUpdatePostJSONRequestBody) (*http.Request, error) { - var bodyReader io.Reader - buf, err := json.Marshal(body) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "pool_id", poolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } - bodyReader = bytes.NewReader(buf) - return NewClustersUpgradeApiV2ClustersClusterIdUpdatePostRequestWithBody(server, clusterId, "application/json", bodyReader) -} - -// NewClustersUpgradeApiV2ClustersClusterIdUpdatePostRequestWithBody constructs an http.Request for the ClustersUpgradeApiV2ClustersClusterIdUpdatePost method, with any body, and a specified content type -func NewClustersUpgradeApiV2ClustersClusterIdUpdatePostRequestWithBody(server string, clusterId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { - var err error - var pathParam0 string + var pathParam2 string - pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "volume_id", volumeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -11423,7 +12124,7 @@ func NewClustersUpgradeApiV2ClustersClusterIdUpdatePostRequestWithBody(server st return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/update", pathParam0) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/replication/tasks", pathParam0, pathParam1, pathParam2) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -11433,56 +12134,100 @@ func NewClustersUpgradeApiV2ClustersClusterIdUpdatePostRequestWithBody(server st return nil, err } - req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err } - req.Header.Add("Content-Type", contentType) - return req, nil } -// NewManagementNodesListApiV2ManagementNodesGetRequest constructs an http.Request for the ManagementNodesListApiV2ManagementNodesGet method -func NewManagementNodesListApiV2ManagementNodesGetRequest(server string, params *ManagementNodesListApiV2ManagementNodesGetParams) (*http.Request, error) { +// NewClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTriggerPostRequest constructs an http.Request for the ClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTriggerPost method +func NewClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTriggerPostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID) (*http.Request, error) { var err error - serverURL, err := url.Parse(server) + var pathParam0 string + + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } - operationPath := fmt.Sprintf("/api/v2/management-nodes/") - if operationPath[0] == '/' { - operationPath = "." + operationPath - } + var pathParam1 string - queryURL, err := serverURL.Parse(operationPath) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "pool_id", poolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } - if params != nil { - // queryValues collects non-styled parameters (passthrough, JSON) - // that are safe to round-trip through url.Values.Encode(). - queryValues := queryURL.Query() - // rawQueryFragments collects pre-encoded query fragments from - // styled parameters, preserving literal commas as delimiters - // per the OpenAPI spec (e.g. "color=blue,black,brown"). - var rawQueryFragments []string + var pathParam2 string - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "cluster_id", params.ClusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "string", Format: "uuid"}); err != nil { - return nil, err - } else { - for _, qp := range strings.Split(queryFrag, "&") { - rawQueryFragments = append(rawQueryFragments, qp) - } - } + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "volume_id", volumeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } - if encoded := queryValues.Encode(); encoded != "" { - rawQueryFragments = append(rawQueryFragments, encoded) - } - queryURL.RawQuery = strings.Join(rawQueryFragments, "&") + serverURL, err := url.Parse(server) + if err != nil { + return nil, err + } + + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/replication/trigger", pathParam0, pathParam1, pathParam2) + if operationPath[0] == '/' { + operationPath = "." + operationPath + } + + queryURL, err := serverURL.Parse(operationPath) + if err != nil { + return nil, err + } + + req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) + if err != nil { + return nil, err + } + + return req, nil +} + +// NewClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsGetRequest constructs an http.Request for the ClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsGet method +func NewClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsGetRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID) (*http.Request, error) { + var err error + + var pathParam0 string + + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + var pathParam1 string + + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "pool_id", poolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + var pathParam2 string + + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "volume_id", volumeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + serverURL, err := url.Parse(server) + if err != nil { + return nil, err + } + + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/snapshots", pathParam0, pathParam1, pathParam2) + if operationPath[0] == '/' { + operationPath = "." + operationPath + } + + queryURL, err := serverURL.Parse(operationPath) + if err != nil { + return nil, err } req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) @@ -11493,13 +12238,38 @@ func NewManagementNodesListApiV2ManagementNodesGetRequest(server string, params return req, nil } -// NewManagementNodeDetailApiV2ManagementNodesManagementNodeIdGetRequest constructs an http.Request for the ManagementNodeDetailApiV2ManagementNodesManagementNodeIdGet method -func NewManagementNodeDetailApiV2ManagementNodesManagementNodeIdGetRequest(server string, managementNodeId openapi_types.UUID, params *ManagementNodeDetailApiV2ManagementNodesManagementNodeIdGetParams) (*http.Request, error) { +// NewClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostRequest calls the generic ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPost builder with application/json body +func NewClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, body ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostJSONRequestBody) (*http.Request, error) { + var bodyReader io.Reader + buf, err := json.Marshal(body) + if err != nil { + return nil, err + } + bodyReader = bytes.NewReader(buf) + return NewClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostRequestWithBody(server, clusterId, poolId, volumeId, "application/json", bodyReader) +} + +// NewClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostRequestWithBody constructs an http.Request for the ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPost method, with any body, and a specified content type +func NewClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostRequestWithBody(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { var err error var pathParam0 string - pathParam0, err = runtime.StyleParamWithOptions("simple", false, "management_node_id", managementNodeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + var pathParam1 string + + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "pool_id", poolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + var pathParam2 string + + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "volume_id", volumeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) if err != nil { return nil, err } @@ -11509,7 +12279,102 @@ func NewManagementNodeDetailApiV2ManagementNodesManagementNodeIdGetRequest(serve return nil, err } - operationPath := fmt.Sprintf("/api/v2/management-nodes/%s/", pathParam0) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/snapshots", pathParam0, pathParam1, pathParam2) + if operationPath[0] == '/' { + operationPath = "." + operationPath + } + + queryURL, err := serverURL.Parse(operationPath) + if err != nil { + return nil, err + } + + req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) + if err != nil { + return nil, err + } + + req.Header.Add("Content-Type", contentType) + + return req, nil +} + +// NewClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigrationsGetRequest constructs an http.Request for the ClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigrationsGet method +func NewClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigrationsGetRequest(server string, clusterId openapi_types.UUID, nqn string) (*http.Request, error) { + var err error + + var pathParam0 string + + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + var pathParam1 string + + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "nqn", nqn, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: ""}) + if err != nil { + return nil, err + } + + serverURL, err := url.Parse(server) + if err != nil { + return nil, err + } + + operationPath := fmt.Sprintf("/api/v2/clusters/%s/subsystems/%s/migrations/", pathParam0, pathParam1) + if operationPath[0] == '/' { + operationPath = "." + operationPath + } + + queryURL, err := serverURL.Parse(operationPath) + if err != nil { + return nil, err + } + + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + if err != nil { + return nil, err + } + + return req, nil +} + +// NewClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostRequest calls the generic ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPost builder with application/json body +func NewClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostRequest(server string, clusterId openapi_types.UUID, nqn string, params *ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostParams, body ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostJSONRequestBody) (*http.Request, error) { + var bodyReader io.Reader + buf, err := json.Marshal(body) + if err != nil { + return nil, err + } + bodyReader = bytes.NewReader(buf) + return NewClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostRequestWithBody(server, clusterId, nqn, params, "application/json", bodyReader) +} + +// NewClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostRequestWithBody constructs an http.Request for the ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPost method, with any body, and a specified content type +func NewClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostRequestWithBody(server string, clusterId openapi_types.UUID, nqn string, params *ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostParams, contentType string, body io.Reader) (*http.Request, error) { + var err error + + var pathParam0 string + + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + var pathParam1 string + + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "nqn", nqn, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: ""}) + if err != nil { + return nil, err + } + + serverURL, err := url.Parse(server) + if err != nil { + return nil, err + } + + operationPath := fmt.Sprintf("/api/v2/clusters/%s/subsystems/%s/migrations/", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -11528,12 +12393,16 @@ func NewManagementNodeDetailApiV2ManagementNodesManagementNodeIdGetRequest(serve // per the OpenAPI spec (e.g. "color=blue,black,brown"). var rawQueryFragments []string - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "cluster_id", params.ClusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "string", Format: "uuid"}); err != nil { - return nil, err - } else { - for _, qp := range strings.Split(queryFrag, "&") { - rawQueryFragments = append(rawQueryFragments, qp) + if params.ResponseFormat != nil { + + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "response-format", *params.ResponseFormat, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "string", Format: ""}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } } + } if encoded := queryValues.Encode(); encoded != "" { @@ -11542,1157 +12411,2243 @@ func NewManagementNodeDetailApiV2ManagementNodesManagementNodeIdGetRequest(serve queryURL.RawQuery = strings.Join(rawQueryFragments, "&") } - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) if err != nil { return nil, err } + req.Header.Add("Content-Type", contentType) + return req, nil } -func (c *Client) applyEditors(ctx context.Context, req *http.Request, additionalEditors []RequestEditorFn) error { - for _, r := range c.RequestEditors { - if err := r(ctx, req); err != nil { - return err - } - } - for _, r := range additionalEditors { - if err := r(ctx, req); err != nil { - return err - } +// NewClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdDeleteRequest constructs an http.Request for the ClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdDelete method +func NewClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdDeleteRequest(server string, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID) (*http.Request, error) { + var err error + + var pathParam0 string + + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err } - return nil -} -// ClientWithResponses builds on ClientInterface to offer response payloads -type ClientWithResponses struct { - ClientInterface -} + var pathParam1 string -// NewClientWithResponses creates a new ClientWithResponses, which wraps -// Client with return type handling -func NewClientWithResponses(server string, opts ...ClientOption) (*ClientWithResponses, error) { - client, err := NewClient(server, opts...) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "nqn", nqn, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: ""}) if err != nil { return nil, err } - return &ClientWithResponses{client}, nil -} -// WithBaseURL overrides the baseURL. -func WithBaseURL(baseURL string) ClientOption { - return func(c *Client) error { - newBaseURL, err := url.Parse(baseURL) - if err != nil { - return err - } - c.Server = newBaseURL.String() - return nil + var pathParam2 string + + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "migration_id", migrationId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err } -} -// ClientWithResponsesInterface is the interface specification for the client with responses above. -type ClientWithResponsesInterface interface { + serverURL, err := url.Parse(server) + if err != nil { + return nil, err + } - // HealthApiV2MetaHealthGetWithResponse Health - // - // Liveness probe: succeeds whenever the process can serve requests. - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/_meta/health (the `HealthApiV2MetaHealthGet` operationId). - HealthApiV2MetaHealthGetWithResponse(ctx context.Context, reqEditors ...RequestEditorFn) (*HealthApiV2MetaHealthGetResponse, error) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/subsystems/%s/migrations/%s/", pathParam0, pathParam1, pathParam2) + if operationPath[0] == '/' { + operationPath = "." + operationPath + } - // ReadyApiV2MetaReadyGetWithResponse Ready - // - // Readiness probe: succeeds when the FoundationDB backend is reachable. - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/_meta/ready (the `ReadyApiV2MetaReadyGet` operationId). - ReadyApiV2MetaReadyGetWithResponse(ctx context.Context, reqEditors ...RequestEditorFn) (*ReadyApiV2MetaReadyGetResponse, error) + queryURL, err := serverURL.Parse(operationPath) + if err != nil { + return nil, err + } - // ClustersListApiV2ClustersGetWithResponse Clusters:List - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/ (the `ClustersListApiV2ClustersGet` operationId). - ClustersListApiV2ClustersGetWithResponse(ctx context.Context, params *ClustersListApiV2ClustersGetParams, reqEditors ...RequestEditorFn) (*ClustersListApiV2ClustersGetResponse, error) + req, err := http.NewRequest(http.MethodDelete, queryURL.String(), nil) + if err != nil { + return nil, err + } - // ClustersCreateApiV2ClustersPostWithBodyWithResponse Clusters:Create - // - // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/ (the `ClustersCreateApiV2ClustersPost` operationId). - ClustersCreateApiV2ClustersPostWithBodyWithResponse(ctx context.Context, params *ClustersCreateApiV2ClustersPostParams, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersCreateApiV2ClustersPostResponse, error) + return req, nil +} - // ClustersCreateApiV2ClustersPostWithResponse Clusters:Create - // - // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/ (the `ClustersCreateApiV2ClustersPost` operationId). - ClustersCreateApiV2ClustersPostWithResponse(ctx context.Context, params *ClustersCreateApiV2ClustersPostParams, body ClustersCreateApiV2ClustersPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersCreateApiV2ClustersPostResponse, error) +// NewClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdGetRequest constructs an http.Request for the ClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdGet method +func NewClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdGetRequest(server string, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID) (*http.Request, error) { + var err error - // ClustersDeleteApiV2ClustersClusterIdDeleteWithResponse Clusters:Delete - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with DELETE /api/v2/clusters/{cluster_id}/ (the `ClustersDeleteApiV2ClustersClusterIdDelete` operationId). - ClustersDeleteApiV2ClustersClusterIdDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersDeleteApiV2ClustersClusterIdDeleteResponse, error) + var pathParam0 string - // ClustersDetailApiV2ClustersClusterIdGetWithResponse Clusters:Detail - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/{cluster_id}/ (the `ClustersDetailApiV2ClustersClusterIdGet` operationId). - ClustersDetailApiV2ClustersClusterIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersDetailApiV2ClustersClusterIdGetParams, reqEditors ...RequestEditorFn) (*ClustersDetailApiV2ClustersClusterIdGetResponse, error) + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } - // ClustersUpdateApiV2ClustersClusterIdPutWithBodyWithResponse Clusters:Update - // - // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with PUT /api/v2/clusters/{cluster_id}/ (the `ClustersUpdateApiV2ClustersClusterIdPut` operationId). - ClustersUpdateApiV2ClustersClusterIdPutWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersUpdateApiV2ClustersClusterIdPutResponse, error) + var pathParam1 string - // ClustersUpdateApiV2ClustersClusterIdPutWithResponse Clusters:Update - // - // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with PUT /api/v2/clusters/{cluster_id}/ (the `ClustersUpdateApiV2ClustersClusterIdPut` operationId). - ClustersUpdateApiV2ClustersClusterIdPutWithResponse(ctx context.Context, clusterId openapi_types.UUID, body ClustersUpdateApiV2ClustersClusterIdPutJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersUpdateApiV2ClustersClusterIdPutResponse, error) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "nqn", nqn, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: ""}) + if err != nil { + return nil, err + } - // ClustersActivateApiV2ClustersClusterIdActivatePostWithResponse Clusters:Activate - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/activate (the `ClustersActivateApiV2ClustersClusterIdActivatePost` operationId). - ClustersActivateApiV2ClustersClusterIdActivatePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersActivateApiV2ClustersClusterIdActivatePostResponse, error) + var pathParam2 string - // ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostWithBodyWithResponse Clusters:Addreplication - // - // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/addreplication (the `ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPost` operationId). - ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostResponse, error) + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "migration_id", migrationId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } - // ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostWithResponse Clusters:Addreplication - // - // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/addreplication (the `ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPost` operationId). - ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, body ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostResponse, error) + serverURL, err := url.Parse(server) + if err != nil { + return nil, err + } - // ClustersBackupsListApiV2ClustersClusterIdBackupsGetWithResponse Clusters:Backups:List - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/{cluster_id}/backups/ (the `ClustersBackupsListApiV2ClustersClusterIdBackupsGet` operationId). - ClustersBackupsListApiV2ClustersClusterIdBackupsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersBackupsListApiV2ClustersClusterIdBackupsGetResponse, error) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/subsystems/%s/migrations/%s/", pathParam0, pathParam1, pathParam2) + if operationPath[0] == '/' { + operationPath = "." + operationPath + } - // ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostWithBodyWithResponse Clusters:Backups:Create - // - // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/backups/ (the `ClustersBackupsCreateApiV2ClustersClusterIdBackupsPost` operationId). - ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostParams, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostResponse, error) + queryURL, err := serverURL.Parse(operationPath) + if err != nil { + return nil, err + } - // ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostWithResponse Clusters:Backups:Create - // - // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/backups/ (the `ClustersBackupsCreateApiV2ClustersClusterIdBackupsPost` operationId). - ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostParams, body ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostResponse, error) + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + if err != nil { + return nil, err + } - // ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetWithResponse Clusters:Backup-Policies:List - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/{cluster_id}/backups/backup-policies/ (the `ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGet` operationId). - ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetResponse, error) + return req, nil +} - // ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostWithBodyWithResponse Clusters:Backup-Policies:Create - // - // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/backups/backup-policies/ (the `ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPost` operationId). - ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostResponse, error) +// NewClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdCleanupTargetPostRequest constructs an http.Request for the ClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdCleanupTargetPost method +func NewClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdCleanupTargetPostRequest(server string, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID) (*http.Request, error) { + var err error - // ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostWithResponse Clusters:Backup-Policies:Create - // - // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/backups/backup-policies/ (the `ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPost` operationId). - ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, body ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostResponse, error) + var pathParam0 string - // ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDeleteWithResponse Clusters:Backup-Policies:Delete - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with DELETE /api/v2/clusters/{cluster_id}/backups/backup-policies/{policy_id} (the `ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDelete` operationId). - ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, policyId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDeleteResponse, error) + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } - // ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostWithBodyWithResponse Clusters:Backup-Policies:Attach - // - // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/backups/backup-policies/{policy_id}/attach (the `ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPost` operationId). - ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, policyId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostResponse, error) + var pathParam1 string - // ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostWithResponse Clusters:Backup-Policies:Attach - // - // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/backups/backup-policies/{policy_id}/attach (the `ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPost` operationId). - ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, policyId openapi_types.UUID, body ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostResponse, error) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "nqn", nqn, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: ""}) + if err != nil { + return nil, err + } - // ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostWithBodyWithResponse Clusters:Backup-Policies:Detach - // - // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/backups/backup-policies/{policy_id}/detach (the `ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPost` operationId). - ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, policyId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostResponse, error) + var pathParam2 string - // ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostWithResponse Clusters:Backup-Policies:Detach - // - // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/backups/backup-policies/{policy_id}/detach (the `ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPost` operationId). - ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, policyId openapi_types.UUID, body ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostResponse, error) + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "migration_id", migrationId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } - // ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetWithResponse Clusters:Backups:Export - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/{cluster_id}/backups/export (the `ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGet` operationId). - ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetParams, reqEditors ...RequestEditorFn) (*ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetResponse, error) + serverURL, err := url.Parse(server) + if err != nil { + return nil, err + } - // ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostWithBodyWithResponse Clusters:Backups:Import - // - // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/backups/import (the `ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPost` operationId). - ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse, error) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/subsystems/%s/migrations/%s/cleanup-target", pathParam0, pathParam1, pathParam2) + if operationPath[0] == '/' { + operationPath = "." + operationPath + } - // ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostWithResponse Clusters:Backups:Import - // - // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/backups/import (the `ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPost` operationId). - ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, body ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse, error) - - // ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostWithBodyWithResponse Clusters:Backups:Restore - // - // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/backups/restore (the `ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePost` operationId). - ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse, error) + queryURL, err := serverURL.Parse(operationPath) + if err != nil { + return nil, err + } - // ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostWithResponse Clusters:Backups:Restore - // - // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/backups/restore (the `ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePost` operationId). - ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, body ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse, error) + req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) + if err != nil { + return nil, err + } - // ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostWithBodyWithResponse Clusters:Backups:Source-Switch - // - // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/backups/source-switch (the `ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPost` operationId). - ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse, error) + return req, nil +} - // ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostWithResponse Clusters:Backups:Source-Switch - // - // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/backups/source-switch (the `ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPost` operationId). - ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, body ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse, error) +// NewClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostRequest calls the generic ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePost builder with application/json body +func NewClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostRequest(server string, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID, body ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostJSONRequestBody) (*http.Request, error) { + var bodyReader io.Reader + buf, err := json.Marshal(body) + if err != nil { + return nil, err + } + bodyReader = bytes.NewReader(buf) + return NewClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostRequestWithBody(server, clusterId, nqn, migrationId, "application/json", bodyReader) +} - // ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetWithResponse Clusters:Backups:Sources - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/{cluster_id}/backups/sources (the `ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGet` operationId). - ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse, error) +// NewClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostRequestWithBody constructs an http.Request for the ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePost method, with any body, and a specified content type +func NewClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostRequestWithBody(server string, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { + var err error - // ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetWithResponse Clusters:Backups:Detail - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/{cluster_id}/backups/{backup_id}/ (the `ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGet` operationId). - ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, backupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse, error) + var pathParam0 string - // ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteWithResponse Deprecated — delete all backups for a volume - // - // Deprecated. Use `DELETE /clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/backups` instead. - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with DELETE /api/v2/clusters/{cluster_id}/backups/{volume_id} (the `ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDelete` operationId). - // - // Deprecated: this operation has been marked as deprecated upstream, but no `x-deprecated-reason` was set - ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse, error) + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } - // ClustersCapacityApiV2ClustersClusterIdCapacityGetWithResponse Clusters:Capacity - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/{cluster_id}/capacity (the `ClustersCapacityApiV2ClustersClusterIdCapacityGet` operationId). - ClustersCapacityApiV2ClustersClusterIdCapacityGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersCapacityApiV2ClustersClusterIdCapacityGetParams, reqEditors ...RequestEditorFn) (*ClustersCapacityApiV2ClustersClusterIdCapacityGetResponse, error) + var pathParam1 string - // ClustersExpandApiV2ClustersClusterIdExpandPostWithResponse Clusters:Expand - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/expand (the `ClustersExpandApiV2ClustersClusterIdExpandPost` operationId). - ClustersExpandApiV2ClustersClusterIdExpandPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersExpandApiV2ClustersClusterIdExpandPostResponse, error) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "nqn", nqn, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: ""}) + if err != nil { + return nil, err + } - // ClustersIostatsApiV2ClustersClusterIdIostatsGetWithResponse Clusters:Iostats - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/{cluster_id}/iostats (the `ClustersIostatsApiV2ClustersClusterIdIostatsGet` operationId). - ClustersIostatsApiV2ClustersClusterIdIostatsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersIostatsApiV2ClustersClusterIdIostatsGetParams, reqEditors ...RequestEditorFn) (*ClustersIostatsApiV2ClustersClusterIdIostatsGetResponse, error) + var pathParam2 string - // ClustersLogsApiV2ClustersClusterIdLogsGetWithResponse Clusters:Logs - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/{cluster_id}/logs (the `ClustersLogsApiV2ClustersClusterIdLogsGet` operationId). - ClustersLogsApiV2ClustersClusterIdLogsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersLogsApiV2ClustersClusterIdLogsGetParams, reqEditors ...RequestEditorFn) (*ClustersLogsApiV2ClustersClusterIdLogsGetResponse, error) + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "migration_id", migrationId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } - // ClustersRebalanceApiV2ClustersClusterIdRebalancePostWithResponse Clusters:Rebalance - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/rebalance (the `ClustersRebalanceApiV2ClustersClusterIdRebalancePost` operationId). - ClustersRebalanceApiV2ClustersClusterIdRebalancePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersRebalanceApiV2ClustersClusterIdRebalancePostResponse, error) + serverURL, err := url.Parse(server) + if err != nil { + return nil, err + } - // ClustersReplicationPoliciesListApiV2ClustersClusterIdReplicationPoliciesGetWithResponse Clusters:Replication:Policies:List - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/{cluster_id}/replication/policies/ (the `ClustersReplicationPoliciesListApiV2ClustersClusterIdReplicationPoliciesGet` operationId). - ClustersReplicationPoliciesListApiV2ClustersClusterIdReplicationPoliciesGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersReplicationPoliciesListApiV2ClustersClusterIdReplicationPoliciesGetResponse, error) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/subsystems/%s/migrations/%s/continue", pathParam0, pathParam1, pathParam2) + if operationPath[0] == '/' { + operationPath = "." + operationPath + } - // ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostWithBodyWithResponse Clusters:Replication:Policies:Create - // - // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/replication/policies/ (the `ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPost` operationId). - ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostParams, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostResponse, error) + queryURL, err := serverURL.Parse(operationPath) + if err != nil { + return nil, err + } - // ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostWithResponse Clusters:Replication:Policies:Create - // - // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/replication/policies/ (the `ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPost` operationId). - ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostParams, body ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostResponse, error) + req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) + if err != nil { + return nil, err + } - // ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDeleteWithResponse Clusters:Replication:Policies:Delete - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with DELETE /api/v2/clusters/{cluster_id}/replication/policies/{policy_id}/ (the `ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDelete` operationId). - ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, policyId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDeleteResponse, error) + req.Header.Add("Content-Type", contentType) - // ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGetWithResponse Clusters:Replication:Policies:Detail - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/{cluster_id}/replication/policies/{policy_id}/ (the `ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGet` operationId). - ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, policyId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGetResponse, error) + return req, nil +} - // ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPostWithResponse Clusters:Replication:Policies:Failover - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/replication/policies/{policy_id}/failover (the `ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPost` operationId). - ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, policyId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPostResponse, error) +// NewClustersTasksListApiV2ClustersClusterIdTasksGetRequest constructs an http.Request for the ClustersTasksListApiV2ClustersClusterIdTasksGet method +func NewClustersTasksListApiV2ClustersClusterIdTasksGetRequest(server string, clusterId openapi_types.UUID, params *ClustersTasksListApiV2ClustersClusterIdTasksGetParams) (*http.Request, error) { + var err error - // ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetWithResponse Clusters:Replication:Relationships:Detail - // - // Replication relationship for a volume, resolvable even when the source volume - // has been deleted (e.g. after replication-commit --delete-source). The CSI driver - // uses this to redirect NodeStageVolume to the active volume on the target cluster. - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/{cluster_id}/replication/relationships/{lvol_id} (the `ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGet` operationId). - ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, lvolId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetResponse, error) + var pathParam0 string - // ClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGetWithResponse Clusters:Replication:Targets:List - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/{cluster_id}/replication/targets/ (the `ClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGet` operationId). - ClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGetResponse, error) + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } - // ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostWithBodyWithResponse Clusters:Replication:Targets:Create - // - // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/replication/targets/ (the `ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPost` operationId). - ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostParams, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostResponse, error) + serverURL, err := url.Parse(server) + if err != nil { + return nil, err + } - // ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostWithResponse Clusters:Replication:Targets:Create - // - // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/replication/targets/ (the `ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPost` operationId). - ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostParams, body ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostResponse, error) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/tasks/", pathParam0) + if operationPath[0] == '/' { + operationPath = "." + operationPath + } - // ClustersReplicationTargetsDeleteApiV2ClustersClusterIdReplicationTargetsTargetIdDeleteWithResponse Clusters:Replication:Targets:Delete - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with DELETE /api/v2/clusters/{cluster_id}/replication/targets/{target_id}/ (the `ClustersReplicationTargetsDeleteApiV2ClustersClusterIdReplicationTargetsTargetIdDelete` operationId). - ClustersReplicationTargetsDeleteApiV2ClustersClusterIdReplicationTargetsTargetIdDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, targetId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersReplicationTargetsDeleteApiV2ClustersClusterIdReplicationTargetsTargetIdDeleteResponse, error) + queryURL, err := serverURL.Parse(operationPath) + if err != nil { + return nil, err + } - // ClustersReplicationTargetsDetailApiV2ClustersClusterIdReplicationTargetsTargetIdGetWithResponse Clusters:Replication:Targets:Detail - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/{cluster_id}/replication/targets/{target_id}/ (the `ClustersReplicationTargetsDetailApiV2ClustersClusterIdReplicationTargetsTargetIdGet` operationId). - ClustersReplicationTargetsDetailApiV2ClustersClusterIdReplicationTargetsTargetIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, targetId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersReplicationTargetsDetailApiV2ClustersClusterIdReplicationTargetsTargetIdGetResponse, error) + if params != nil { + // queryValues collects non-styled parameters (passthrough, JSON) + // that are safe to round-trip through url.Values.Encode(). + queryValues := queryURL.Query() + // rawQueryFragments collects pre-encoded query fragments from + // styled parameters, preserving literal commas as delimiters + // per the OpenAPI spec (e.g. "color=blue,black,brown"). + var rawQueryFragments []string - // ClustersReplicationTargetsFailoverApiV2ClustersClusterIdReplicationTargetsTargetIdFailoverPostWithResponse Clusters:Replication:Targets:Failover - // - // Fail over EVERY volume replicating to this target. - // - // A site loss has to move all volumes at once; doing it volume by volume was - // the only option before. Idempotent per volume and reports one result per - // volume, so a partial failure is visible instead of silent. - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/replication/targets/{target_id}/failover (the `ClustersReplicationTargetsFailoverApiV2ClustersClusterIdReplicationTargetsTargetIdFailoverPost` operationId). - ClustersReplicationTargetsFailoverApiV2ClustersClusterIdReplicationTargetsTargetIdFailoverPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, targetId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersReplicationTargetsFailoverApiV2ClustersClusterIdReplicationTargetsTargetIdFailoverPostResponse, error) + if params.Watch != nil { - // ClustersShutdownApiV2ClustersClusterIdShutdownPostWithResponse Clusters:Shutdown - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/shutdown (the `ClustersShutdownApiV2ClustersClusterIdShutdownPost` operationId). - ClustersShutdownApiV2ClustersClusterIdShutdownPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersShutdownApiV2ClustersClusterIdShutdownPostResponse, error) - - // ClustersStartApiV2ClustersClusterIdStartPostWithResponse Clusters:Start - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/start (the `ClustersStartApiV2ClustersClusterIdStartPost` operationId). - ClustersStartApiV2ClustersClusterIdStartPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStartApiV2ClustersClusterIdStartPostResponse, error) + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "watch", *params.Watch, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } + } - // ClustersStorageNodesListApiV2ClustersClusterIdStorageNodesGetWithResponse Clusters:Storage-Nodes:List - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-nodes/ (the `ClustersStorageNodesListApiV2ClustersClusterIdStorageNodesGet` operationId). - ClustersStorageNodesListApiV2ClustersClusterIdStorageNodesGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersStorageNodesListApiV2ClustersClusterIdStorageNodesGetParams, reqEditors ...RequestEditorFn) (*ClustersStorageNodesListApiV2ClustersClusterIdStorageNodesGetResponse, error) + } - // ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostWithBodyWithResponse Clusters:Storage-Nodes:Create - // - // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-nodes/ (the `ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPost` operationId). - ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostParams, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostResponse, error) + if encoded := queryValues.Encode(); encoded != "" { + rawQueryFragments = append(rawQueryFragments, encoded) + } + queryURL.RawQuery = strings.Join(rawQueryFragments, "&") + } - // ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostWithResponse Clusters:Storage-Nodes:Create - // - // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-nodes/ (the `ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPost` operationId). - ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostParams, body ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostResponse, error) + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + if err != nil { + return nil, err + } - // ClustersStorageNodesDeleteApiV2ClustersClusterIdStorageNodesStorageNodeIdDeleteWithResponse Clusters:Storage-Nodes:Delete - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with DELETE /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/ (the `ClustersStorageNodesDeleteApiV2ClustersClusterIdStorageNodesStorageNodeIdDelete` operationId). - ClustersStorageNodesDeleteApiV2ClustersClusterIdStorageNodesStorageNodeIdDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, params *ClustersStorageNodesDeleteApiV2ClustersClusterIdStorageNodesStorageNodeIdDeleteParams, reqEditors ...RequestEditorFn) (*ClustersStorageNodesDeleteApiV2ClustersClusterIdStorageNodesStorageNodeIdDeleteResponse, error) + return req, nil +} - // ClustersStorageNodesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdGetWithResponse Clusters:Storage-Nodes:Detail - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/ (the `ClustersStorageNodesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdGet` operationId). - ClustersStorageNodesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, params *ClustersStorageNodesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdGetParams, reqEditors ...RequestEditorFn) (*ClustersStorageNodesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdGetResponse, error) +// NewClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGetRequest constructs an http.Request for the ClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGet method +func NewClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGetRequest(server string, clusterId openapi_types.UUID, taskId openapi_types.UUID, params *ClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGetParams) (*http.Request, error) { + var err error - // ClustersStorageNodesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdCapacityGetWithResponse Clusters:Storage-Nodes:Capacity - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/capacity (the `ClustersStorageNodesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdCapacityGet` operationId). - ClustersStorageNodesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdCapacityGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, params *ClustersStorageNodesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdCapacityGetParams, reqEditors ...RequestEditorFn) (*ClustersStorageNodesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdCapacityGetResponse, error) + var pathParam0 string - // ClustersStorageNodesDevicesListApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesGetWithResponse Clusters:Storage Nodes:Devices:List - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/devices/ (the `ClustersStorageNodesDevicesListApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesGet` operationId). - ClustersStorageNodesDevicesListApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, params *ClustersStorageNodesDevicesListApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesGetParams, reqEditors ...RequestEditorFn) (*ClustersStorageNodesDevicesListApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesGetResponse, error) + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } - // ClustersStorageNodesDevicesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdGetWithResponse Clusters:Storage Nodes:Devices:Detail - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/devices/{device_id}/ (the `ClustersStorageNodesDevicesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdGet` operationId). - ClustersStorageNodesDevicesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, deviceId openapi_types.UUID, params *ClustersStorageNodesDevicesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdGetParams, reqEditors ...RequestEditorFn) (*ClustersStorageNodesDevicesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdGetResponse, error) + var pathParam1 string - // ClustersStorageNodesDevicesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdCapacityGetWithResponse Clusters:Storage Nodes:Devices:Capacity - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/devices/{device_id}/capacity (the `ClustersStorageNodesDevicesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdCapacityGet` operationId). - ClustersStorageNodesDevicesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdCapacityGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, deviceId openapi_types.UUID, params *ClustersStorageNodesDevicesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdCapacityGetParams, reqEditors ...RequestEditorFn) (*ClustersStorageNodesDevicesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdCapacityGetResponse, error) + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "task_id", taskId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } - // ClustersStorageNodesDevicesGetDeviceHealthInfoApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdHealthInfoGetWithResponse Clusters:Storage Nodes:Devices:Get-Device-Health-Info - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/devices/{device_id}/health-info (the `ClustersStorageNodesDevicesGetDeviceHealthInfoApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdHealthInfoGet` operationId). - ClustersStorageNodesDevicesGetDeviceHealthInfoApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdHealthInfoGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, deviceId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStorageNodesDevicesGetDeviceHealthInfoApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdHealthInfoGetResponse, error) + serverURL, err := url.Parse(server) + if err != nil { + return nil, err + } - // ClustersStorageNodesDevicesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdIostatsGetWithResponse Clusters:Storage Nodes:Devices:Iostats - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/devices/{device_id}/iostats (the `ClustersStorageNodesDevicesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdIostatsGet` operationId). - ClustersStorageNodesDevicesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdIostatsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, deviceId openapi_types.UUID, params *ClustersStorageNodesDevicesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdIostatsGetParams, reqEditors ...RequestEditorFn) (*ClustersStorageNodesDevicesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdIostatsGetResponse, error) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/tasks/%s/", pathParam0, pathParam1) + if operationPath[0] == '/' { + operationPath = "." + operationPath + } - // ClustersStorageNodesDevicesRemoveApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRemovePostWithResponse Clusters:Storage Nodes:Devices:Remove - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/devices/{device_id}/remove (the `ClustersStorageNodesDevicesRemoveApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRemovePost` operationId). - ClustersStorageNodesDevicesRemoveApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRemovePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, deviceId openapi_types.UUID, params *ClustersStorageNodesDevicesRemoveApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRemovePostParams, reqEditors ...RequestEditorFn) (*ClustersStorageNodesDevicesRemoveApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRemovePostResponse, error) + queryURL, err := serverURL.Parse(operationPath) + if err != nil { + return nil, err + } - // ClustersStorageNodesDevicesResetApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdResetPostWithResponse Clusters:Storage Nodes:Devices:Reset - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/devices/{device_id}/reset (the `ClustersStorageNodesDevicesResetApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdResetPost` operationId). - ClustersStorageNodesDevicesResetApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdResetPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, deviceId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStorageNodesDevicesResetApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdResetPostResponse, error) + if params != nil { + // queryValues collects non-styled parameters (passthrough, JSON) + // that are safe to round-trip through url.Values.Encode(). + queryValues := queryURL.Query() + // rawQueryFragments collects pre-encoded query fragments from + // styled parameters, preserving literal commas as delimiters + // per the OpenAPI spec (e.g. "color=blue,black,brown"). + var rawQueryFragments []string - // ClustersStorageNodesDevicesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRestartPostWithResponse Clusters:Storage Nodes:Devices:Restart - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/devices/{device_id}/restart (the `ClustersStorageNodesDevicesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRestartPost` operationId). - ClustersStorageNodesDevicesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRestartPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, deviceId openapi_types.UUID, params *ClustersStorageNodesDevicesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRestartPostParams, reqEditors ...RequestEditorFn) (*ClustersStorageNodesDevicesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRestartPostResponse, error) + if params.Watch != nil { - // ClustersStorageNodesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdIostatsGetWithResponse Clusters:Storage-Nodes:Iostats - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/iostats (the `ClustersStorageNodesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdIostatsGet` operationId). - ClustersStorageNodesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdIostatsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, params *ClustersStorageNodesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdIostatsGetParams, reqEditors ...RequestEditorFn) (*ClustersStorageNodesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdIostatsGetResponse, error) + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "watch", *params.Watch, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } + } - // ClustersStorageNodesNicsListApiV2ClustersClusterIdStorageNodesStorageNodeIdNicsGetWithResponse Clusters:Storage-Nodes:Nics:List - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/nics (the `ClustersStorageNodesNicsListApiV2ClustersClusterIdStorageNodesStorageNodeIdNicsGet` operationId). - ClustersStorageNodesNicsListApiV2ClustersClusterIdStorageNodesStorageNodeIdNicsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStorageNodesNicsListApiV2ClustersClusterIdStorageNodesStorageNodeIdNicsGetResponse, error) + } - // ClustersStorageNodesNicsIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdNicsNicIdIostatsGetWithResponse Clusters:Storage-Nodes:Nics:Iostats - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/nics/{nic_id}/iostats (the `ClustersStorageNodesNicsIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdNicsNicIdIostatsGet` operationId). - ClustersStorageNodesNicsIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdNicsNicIdIostatsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, nicId string, reqEditors ...RequestEditorFn) (*ClustersStorageNodesNicsIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdNicsNicIdIostatsGetResponse, error) + if encoded := queryValues.Encode(); encoded != "" { + rawQueryFragments = append(rawQueryFragments, encoded) + } + queryURL.RawQuery = strings.Join(rawQueryFragments, "&") + } - // ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdPromotePostWithResponse Clusters:Storage-Nodes:Start - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/promote (the `ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdPromotePost` operationId). - ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdPromotePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdPromotePostResponse, error) + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + if err != nil { + return nil, err + } - // ClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPostWithBodyWithResponse Clusters:Storage-Nodes:Restart - // - // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/restart (the `ClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPost` operationId). - ClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPostResponse, error) + return req, nil +} - // ClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPostWithResponse Clusters:Storage-Nodes:Restart - // - // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/restart (the `ClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPost` operationId). - ClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, body ClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPostResponse, error) +// NewClustersUpgradeApiV2ClustersClusterIdUpdatePostRequest calls the generic ClustersUpgradeApiV2ClustersClusterIdUpdatePost builder with application/json body +func NewClustersUpgradeApiV2ClustersClusterIdUpdatePostRequest(server string, clusterId openapi_types.UUID, body ClustersUpgradeApiV2ClustersClusterIdUpdatePostJSONRequestBody) (*http.Request, error) { + var bodyReader io.Reader + buf, err := json.Marshal(body) + if err != nil { + return nil, err + } + bodyReader = bytes.NewReader(buf) + return NewClustersUpgradeApiV2ClustersClusterIdUpdatePostRequestWithBody(server, clusterId, "application/json", bodyReader) +} - // ClustersStorageNodesResumeApiV2ClustersClusterIdStorageNodesStorageNodeIdResumePostWithResponse Clusters:Storage-Nodes:Resume - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/resume (the `ClustersStorageNodesResumeApiV2ClustersClusterIdStorageNodesStorageNodeIdResumePost` operationId). - ClustersStorageNodesResumeApiV2ClustersClusterIdStorageNodesStorageNodeIdResumePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStorageNodesResumeApiV2ClustersClusterIdStorageNodesStorageNodeIdResumePostResponse, error) +// NewClustersUpgradeApiV2ClustersClusterIdUpdatePostRequestWithBody constructs an http.Request for the ClustersUpgradeApiV2ClustersClusterIdUpdatePost method, with any body, and a specified content type +func NewClustersUpgradeApiV2ClustersClusterIdUpdatePostRequestWithBody(server string, clusterId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { + var err error - // ClustersStorageNodesShutdownApiV2ClustersClusterIdStorageNodesStorageNodeIdShutdownPostWithResponse Clusters:Storage-Nodes:Shutdown - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/shutdown (the `ClustersStorageNodesShutdownApiV2ClustersClusterIdStorageNodesStorageNodeIdShutdownPost` operationId). - ClustersStorageNodesShutdownApiV2ClustersClusterIdStorageNodesStorageNodeIdShutdownPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, params *ClustersStorageNodesShutdownApiV2ClustersClusterIdStorageNodesStorageNodeIdShutdownPostParams, reqEditors ...RequestEditorFn) (*ClustersStorageNodesShutdownApiV2ClustersClusterIdStorageNodesStorageNodeIdShutdownPostResponse, error) + var pathParam0 string - // ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPostWithBodyWithResponse Clusters:Storage-Nodes:Start - // - // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/start (the `ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPost` operationId). - ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPostResponse, error) + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } - // ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPostWithResponse Clusters:Storage-Nodes:Start - // - // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/start (the `ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPost` operationId). - ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, body ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPostResponse, error) + serverURL, err := url.Parse(server) + if err != nil { + return nil, err + } - // ClustersStorageNodesSuspendApiV2ClustersClusterIdStorageNodesStorageNodeIdSuspendPostWithResponse Clusters:Storage-Nodes:Suspend - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/suspend (the `ClustersStorageNodesSuspendApiV2ClustersClusterIdStorageNodesStorageNodeIdSuspendPost` operationId). - ClustersStorageNodesSuspendApiV2ClustersClusterIdStorageNodesStorageNodeIdSuspendPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, params *ClustersStorageNodesSuspendApiV2ClustersClusterIdStorageNodesStorageNodeIdSuspendPostParams, reqEditors ...RequestEditorFn) (*ClustersStorageNodesSuspendApiV2ClustersClusterIdStorageNodesStorageNodeIdSuspendPostResponse, error) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/update", pathParam0) + if operationPath[0] == '/' { + operationPath = "." + operationPath + } - // ClustersStoragePoolsListApiV2ClustersClusterIdStoragePoolsGetWithResponse Clusters:Storage-Pools:List - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/ (the `ClustersStoragePoolsListApiV2ClustersClusterIdStoragePoolsGet` operationId). - ClustersStoragePoolsListApiV2ClustersClusterIdStoragePoolsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersStoragePoolsListApiV2ClustersClusterIdStoragePoolsGetParams, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsListApiV2ClustersClusterIdStoragePoolsGetResponse, error) + queryURL, err := serverURL.Parse(operationPath) + if err != nil { + return nil, err + } - // ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostWithBodyWithResponse Clusters:Storage-Pools:Create - // - // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/ (the `ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPost` operationId). - ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostParams, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostResponse, error) + req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) + if err != nil { + return nil, err + } - // ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostWithResponse Clusters:Storage-Pools:Create + req.Header.Add("Content-Type", contentType) + + return req, nil +} + +// NewManagementNodesListApiV2ManagementNodesGetRequest constructs an http.Request for the ManagementNodesListApiV2ManagementNodesGet method +func NewManagementNodesListApiV2ManagementNodesGetRequest(server string, params *ManagementNodesListApiV2ManagementNodesGetParams) (*http.Request, error) { + var err error + + serverURL, err := url.Parse(server) + if err != nil { + return nil, err + } + + operationPath := fmt.Sprintf("/api/v2/management-nodes/") + if operationPath[0] == '/' { + operationPath = "." + operationPath + } + + queryURL, err := serverURL.Parse(operationPath) + if err != nil { + return nil, err + } + + if params != nil { + // queryValues collects non-styled parameters (passthrough, JSON) + // that are safe to round-trip through url.Values.Encode(). + queryValues := queryURL.Query() + // rawQueryFragments collects pre-encoded query fragments from + // styled parameters, preserving literal commas as delimiters + // per the OpenAPI spec (e.g. "color=blue,black,brown"). + var rawQueryFragments []string + + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "cluster_id", params.ClusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "string", Format: "uuid"}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } + } + + if encoded := queryValues.Encode(); encoded != "" { + rawQueryFragments = append(rawQueryFragments, encoded) + } + queryURL.RawQuery = strings.Join(rawQueryFragments, "&") + } + + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + if err != nil { + return nil, err + } + + return req, nil +} + +// NewManagementNodeDetailApiV2ManagementNodesManagementNodeIdGetRequest constructs an http.Request for the ManagementNodeDetailApiV2ManagementNodesManagementNodeIdGet method +func NewManagementNodeDetailApiV2ManagementNodesManagementNodeIdGetRequest(server string, managementNodeId openapi_types.UUID, params *ManagementNodeDetailApiV2ManagementNodesManagementNodeIdGetParams) (*http.Request, error) { + var err error + + var pathParam0 string + + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "management_node_id", managementNodeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + serverURL, err := url.Parse(server) + if err != nil { + return nil, err + } + + operationPath := fmt.Sprintf("/api/v2/management-nodes/%s/", pathParam0) + if operationPath[0] == '/' { + operationPath = "." + operationPath + } + + queryURL, err := serverURL.Parse(operationPath) + if err != nil { + return nil, err + } + + if params != nil { + // queryValues collects non-styled parameters (passthrough, JSON) + // that are safe to round-trip through url.Values.Encode(). + queryValues := queryURL.Query() + // rawQueryFragments collects pre-encoded query fragments from + // styled parameters, preserving literal commas as delimiters + // per the OpenAPI spec (e.g. "color=blue,black,brown"). + var rawQueryFragments []string + + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "cluster_id", params.ClusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "string", Format: "uuid"}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } + } + + if encoded := queryValues.Encode(); encoded != "" { + rawQueryFragments = append(rawQueryFragments, encoded) + } + queryURL.RawQuery = strings.Join(rawQueryFragments, "&") + } + + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + if err != nil { + return nil, err + } + + return req, nil +} + +func (c *Client) applyEditors(ctx context.Context, req *http.Request, additionalEditors []RequestEditorFn) error { + for _, r := range c.RequestEditors { + if err := r(ctx, req); err != nil { + return err + } + } + for _, r := range additionalEditors { + if err := r(ctx, req); err != nil { + return err + } + } + return nil +} + +// ClientWithResponses builds on ClientInterface to offer response payloads +type ClientWithResponses struct { + ClientInterface +} + +// NewClientWithResponses creates a new ClientWithResponses, which wraps +// Client with return type handling +func NewClientWithResponses(server string, opts ...ClientOption) (*ClientWithResponses, error) { + client, err := NewClient(server, opts...) + if err != nil { + return nil, err + } + return &ClientWithResponses{client}, nil +} + +// WithBaseURL overrides the baseURL. +func WithBaseURL(baseURL string) ClientOption { + return func(c *Client) error { + newBaseURL, err := url.Parse(baseURL) + if err != nil { + return err + } + c.Server = newBaseURL.String() + return nil + } +} + +// ClientWithResponsesInterface is the interface specification for the client with responses above. +type ClientWithResponsesInterface interface { + + // HealthApiV2MetaHealthGetWithResponse Health // - // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). + // Liveness probe: succeeds whenever the process can serve requests. // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/ (the `ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPost` operationId). - ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostParams, body ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostResponse, error) + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/_meta/health (the `HealthApiV2MetaHealthGet` operationId). + HealthApiV2MetaHealthGetWithResponse(ctx context.Context, reqEditors ...RequestEditorFn) (*HealthApiV2MetaHealthGetResponse, error) - // ClustersStoragePoolsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdDeleteWithResponse Clusters:Storage-Pools:Delete + // ReadyApiV2MetaReadyGetWithResponse Ready + // + // Readiness probe: succeeds when the FoundationDB backend is reachable. // // Returns a wrapper object for the known response body format(s). // - // Corresponds with DELETE /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/ (the `ClustersStoragePoolsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdDelete` operationId). - ClustersStoragePoolsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdDeleteResponse, error) + // Corresponds with GET /api/v2/_meta/ready (the `ReadyApiV2MetaReadyGet` operationId). + ReadyApiV2MetaReadyGetWithResponse(ctx context.Context, reqEditors ...RequestEditorFn) (*ReadyApiV2MetaReadyGetResponse, error) - // ClustersStoragePoolsDetailApiV2ClustersClusterIdStoragePoolsPoolIdGetWithResponse Clusters:Storage-Pools:Detail + // ClustersListApiV2ClustersGetWithResponse Clusters:List // // Returns a wrapper object for the known response body format(s). // - // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/ (the `ClustersStoragePoolsDetailApiV2ClustersClusterIdStoragePoolsPoolIdGet` operationId). - ClustersStoragePoolsDetailApiV2ClustersClusterIdStoragePoolsPoolIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, params *ClustersStoragePoolsDetailApiV2ClustersClusterIdStoragePoolsPoolIdGetParams, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsDetailApiV2ClustersClusterIdStoragePoolsPoolIdGetResponse, error) + // Corresponds with GET /api/v2/clusters/ (the `ClustersListApiV2ClustersGet` operationId). + ClustersListApiV2ClustersGetWithResponse(ctx context.Context, params *ClustersListApiV2ClustersGetParams, reqEditors ...RequestEditorFn) (*ClustersListApiV2ClustersGetResponse, error) - // ClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutWithBodyWithResponse Clusters:Storage-Pools:Update + // ClustersCreateApiV2ClustersPostWithBodyWithResponse Clusters:Create // // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). // - // Corresponds with PUT /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/ (the `ClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPut` operationId). - ClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutResponse, error) + // Corresponds with POST /api/v2/clusters/ (the `ClustersCreateApiV2ClustersPost` operationId). + ClustersCreateApiV2ClustersPostWithBodyWithResponse(ctx context.Context, params *ClustersCreateApiV2ClustersPostParams, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersCreateApiV2ClustersPostResponse, error) - // ClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutWithResponse Clusters:Storage-Pools:Update + // ClustersCreateApiV2ClustersPostWithResponse Clusters:Create // // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). // - // Corresponds with PUT /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/ (the `ClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPut` operationId). - ClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, body ClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutResponse, error) + // Corresponds with POST /api/v2/clusters/ (the `ClustersCreateApiV2ClustersPost` operationId). + ClustersCreateApiV2ClustersPostWithResponse(ctx context.Context, params *ClustersCreateApiV2ClustersPostParams, body ClustersCreateApiV2ClustersPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersCreateApiV2ClustersPostResponse, error) - // ClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDeleteWithBodyWithResponse Clusters:Storage-Pools:Remove-Host + // ClustersDeleteApiV2ClustersClusterIdDeleteWithResponse Clusters:Delete // - // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). + // Returns a wrapper object for the known response body format(s). // - // Corresponds with DELETE /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/host (the `ClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDelete` operationId). - ClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDeleteWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDeleteResponse, error) + // Corresponds with DELETE /api/v2/clusters/{cluster_id}/ (the `ClustersDeleteApiV2ClustersClusterIdDelete` operationId). + ClustersDeleteApiV2ClustersClusterIdDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersDeleteApiV2ClustersClusterIdDeleteResponse, error) - // ClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDeleteWithResponse Clusters:Storage-Pools:Remove-Host + // ClustersDetailApiV2ClustersClusterIdGetWithResponse Clusters:Detail // - // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). + // Returns a wrapper object for the known response body format(s). // - // Corresponds with DELETE /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/host (the `ClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDelete` operationId). - ClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, body ClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDeleteJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDeleteResponse, error) + // Corresponds with GET /api/v2/clusters/{cluster_id}/ (the `ClustersDetailApiV2ClustersClusterIdGet` operationId). + ClustersDetailApiV2ClustersClusterIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersDetailApiV2ClustersClusterIdGetParams, reqEditors ...RequestEditorFn) (*ClustersDetailApiV2ClustersClusterIdGetResponse, error) - // ClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPostWithBodyWithResponse Clusters:Storage-Pools:Add-Host + // ClustersUpdateApiV2ClustersClusterIdPutWithBodyWithResponse Clusters:Update // // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/host (the `ClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPost` operationId). - ClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPostResponse, error) + // Corresponds with PUT /api/v2/clusters/{cluster_id}/ (the `ClustersUpdateApiV2ClustersClusterIdPut` operationId). + ClustersUpdateApiV2ClustersClusterIdPutWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersUpdateApiV2ClustersClusterIdPutResponse, error) - // ClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPostWithResponse Clusters:Storage-Pools:Add-Host + // ClustersUpdateApiV2ClustersClusterIdPutWithResponse Clusters:Update // // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/host (the `ClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPost` operationId). - ClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, body ClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPostResponse, error) + // Corresponds with PUT /api/v2/clusters/{cluster_id}/ (the `ClustersUpdateApiV2ClustersClusterIdPut` operationId). + ClustersUpdateApiV2ClustersClusterIdPutWithResponse(ctx context.Context, clusterId openapi_types.UUID, body ClustersUpdateApiV2ClustersClusterIdPutJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersUpdateApiV2ClustersClusterIdPutResponse, error) - // ClustersStoragePoolsIostatsApiV2ClustersClusterIdStoragePoolsPoolIdIostatsGetWithResponse Clusters:Storage-Pools:Iostats + // ClustersActivateApiV2ClustersClusterIdActivatePostWithResponse Clusters:Activate // // Returns a wrapper object for the known response body format(s). // - // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/iostats (the `ClustersStoragePoolsIostatsApiV2ClustersClusterIdStoragePoolsPoolIdIostatsGet` operationId). - ClustersStoragePoolsIostatsApiV2ClustersClusterIdStoragePoolsPoolIdIostatsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, params *ClustersStoragePoolsIostatsApiV2ClustersClusterIdStoragePoolsPoolIdIostatsGetParams, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsIostatsApiV2ClustersClusterIdStoragePoolsPoolIdIostatsGetResponse, error) + // Corresponds with POST /api/v2/clusters/{cluster_id}/activate (the `ClustersActivateApiV2ClustersClusterIdActivatePost` operationId). + ClustersActivateApiV2ClustersClusterIdActivatePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersActivateApiV2ClustersClusterIdActivatePostResponse, error) - // ClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsGetWithResponse Clusters:Storage-Pools:Snapshots:List + // ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostWithBodyWithResponse Clusters:Addreplication // - // Returns a wrapper object for the known response body format(s). + // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). // - // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/snapshots/ (the `ClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsGet` operationId). - ClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, params *ClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsGetParams, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsGetResponse, error) + // Corresponds with POST /api/v2/clusters/{cluster_id}/addreplication (the `ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPost` operationId). + ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostResponse, error) - // ClustersStoragePoolsSnapshotsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdDeleteWithResponse Clusters:Storage-Pools:Snapshots:Delete + // ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostWithResponse Clusters:Addreplication // - // Returns a wrapper object for the known response body format(s). + // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). // - // Corresponds with DELETE /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/snapshots/{snapshot_id}/ (the `ClustersStoragePoolsSnapshotsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdDelete` operationId). - ClustersStoragePoolsSnapshotsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, snapshotId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsSnapshotsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdDeleteResponse, error) + // Corresponds with POST /api/v2/clusters/{cluster_id}/addreplication (the `ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPost` operationId). + ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, body ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostResponse, error) - // ClustersStoragePoolsSnapshotsDetailApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdGetWithResponse Clusters:Storage-Pools:Snapshots:Detail + // ClustersAlertsListApiV2ClustersClusterIdAlertsGetWithResponse Clusters:Alerts:List + // + // The conditions in this cluster that currently need an operator. + // + // This is not the event log. An alert appears only while it is still true + // and disappears on its own once it is not: the node comes back ONLINE, the + // device comes back, the cluster leaves degraded. Conditions an operator + // caused on purpose -- a node they shut down, a device they removed -- are + // not alerts and are not listed. + // + // By default only what is wrong NOW is returned -- every entry has + // ``status: firing``. Pass ``history=true`` to also get the ones that have + // since resolved, each with its ``resolved_at``, or ``history_seconds=N`` + // for just the recent past. Either way both transitions are written to the + // cluster event log as ALERT_RAISED / ALERT_RESOLVED, so a resolution + // reaches an operator whether or not anyone asks for history here. + // + // Critical sorts before warning, and firing before resolved. // // Returns a wrapper object for the known response body format(s). // - // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/snapshots/{snapshot_id}/ (the `ClustersStoragePoolsSnapshotsDetailApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdGet` operationId). - ClustersStoragePoolsSnapshotsDetailApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, snapshotId openapi_types.UUID, params *ClustersStoragePoolsSnapshotsDetailApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdGetParams, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsSnapshotsDetailApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdGetResponse, error) + // Corresponds with GET /api/v2/clusters/{cluster_id}/alerts/ (the `ClustersAlertsListApiV2ClustersClusterIdAlertsGet` operationId). + ClustersAlertsListApiV2ClustersClusterIdAlertsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersAlertsListApiV2ClustersClusterIdAlertsGetParams, reqEditors ...RequestEditorFn) (*ClustersAlertsListApiV2ClustersClusterIdAlertsGetResponse, error) - // ClustersStoragePoolsVolumesListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesGetWithResponse Clusters:Storage-Pools:Volumes:List + // ClustersBackupsListApiV2ClustersClusterIdBackupsGetWithResponse Clusters:Backups:List // // Returns a wrapper object for the known response body format(s). // - // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/ (the `ClustersStoragePoolsVolumesListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesGet` operationId). - ClustersStoragePoolsVolumesListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, params *ClustersStoragePoolsVolumesListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesGetParams, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesGetResponse, error) + // Corresponds with GET /api/v2/clusters/{cluster_id}/backups/ (the `ClustersBackupsListApiV2ClustersClusterIdBackupsGet` operationId). + ClustersBackupsListApiV2ClustersClusterIdBackupsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersBackupsListApiV2ClustersClusterIdBackupsGetResponse, error) - // ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostWithBodyWithResponse Clusters:Storage-Pools:Volumes:Create + // ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostWithBodyWithResponse Clusters:Backups:Create // // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/ (the `ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPost` operationId). - ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, params *ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostParams, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostResponse, error) + // Corresponds with POST /api/v2/clusters/{cluster_id}/backups/ (the `ClustersBackupsCreateApiV2ClustersClusterIdBackupsPost` operationId). + ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostParams, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostResponse, error) - // ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostWithResponse Clusters:Storage-Pools:Volumes:Create + // ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostWithResponse Clusters:Backups:Create // // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/ (the `ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPost` operationId). - ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, params *ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostParams, body ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostResponse, error) + // Corresponds with POST /api/v2/clusters/{cluster_id}/backups/ (the `ClustersBackupsCreateApiV2ClustersClusterIdBackupsPost` operationId). + ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostParams, body ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostResponse, error) - // ClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPostWithBodyWithResponse Clusters:Storage-Pools:Replicate Lvol On Source Cluster + // ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetWithResponse Clusters:Backup-Policies:List // - // Rebuild a volume on the source cluster. + // Returns a wrapper object for the known response body format(s). // - // Collection-scoped rather than an operation on `/{volume_id}`: the volume is - // typically gone from the source cluster by the time this is called, so the - // controller falls back to resolving the id through the replication records. + // Corresponds with GET /api/v2/clusters/{cluster_id}/backups/backup-policies/ (the `ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGet` operationId). + ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetResponse, error) + + // ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostWithBodyWithResponse Clusters:Backup-Policies:Create // // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/replicate_lvol_on_source_cluster (the `ClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPost` operationId). - ClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPostResponse, error) + // Corresponds with POST /api/v2/clusters/{cluster_id}/backups/backup-policies/ (the `ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPost` operationId). + ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostResponse, error) - // ClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPostWithResponse Clusters:Storage-Pools:Replicate Lvol On Source Cluster - // - // Rebuild a volume on the source cluster. - // - // Collection-scoped rather than an operation on `/{volume_id}`: the volume is - // typically gone from the source cluster by the time this is called, so the - // controller falls back to resolving the id through the replication records. + // ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostWithResponse Clusters:Backup-Policies:Create // // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/replicate_lvol_on_source_cluster (the `ClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPost` operationId). - ClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, body ClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPostResponse, error) + // Corresponds with POST /api/v2/clusters/{cluster_id}/backups/backup-policies/ (the `ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPost` operationId). + ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, body ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostResponse, error) - // ClustersStoragePoolsVolumesDeleteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdDeleteWithResponse Clusters:Storage-Pools:Volumes:Delete + // ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDeleteWithResponse Clusters:Backup-Policies:Delete // // Returns a wrapper object for the known response body format(s). // - // Corresponds with DELETE /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/ (the `ClustersStoragePoolsVolumesDeleteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdDelete` operationId). - ClustersStoragePoolsVolumesDeleteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesDeleteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdDeleteResponse, error) + // Corresponds with DELETE /api/v2/clusters/{cluster_id}/backups/backup-policies/{policy_id} (the `ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDelete` operationId). + ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, policyId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDeleteResponse, error) - // ClustersStoragePoolsVolumesDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdGetWithResponse Clusters:Storage-Pools:Volumes:Detail + // ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostWithBodyWithResponse Clusters:Backup-Policies:Attach // - // Returns a wrapper object for the known response body format(s). + // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). // - // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/ (the `ClustersStoragePoolsVolumesDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdGet` operationId). - ClustersStoragePoolsVolumesDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, params *ClustersStoragePoolsVolumesDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdGetParams, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdGetResponse, error) + // Corresponds with POST /api/v2/clusters/{cluster_id}/backups/backup-policies/{policy_id}/attach (the `ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPost` operationId). + ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, policyId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostResponse, error) - // ClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPutWithBodyWithResponse Clusters:Storage-Pools:Volumes:Update + // ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostWithResponse Clusters:Backup-Policies:Attach + // + // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/backups/backup-policies/{policy_id}/attach (the `ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPost` operationId). + ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, policyId openapi_types.UUID, body ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostResponse, error) + + // ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostWithBodyWithResponse Clusters:Backup-Policies:Detach // // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). // - // Corresponds with PUT /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/ (the `ClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPut` operationId). - ClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPutWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPutResponse, error) + // Corresponds with POST /api/v2/clusters/{cluster_id}/backups/backup-policies/{policy_id}/detach (the `ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPost` operationId). + ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, policyId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostResponse, error) - // ClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPutWithResponse Clusters:Storage-Pools:Volumes:Update + // ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostWithResponse Clusters:Backup-Policies:Detach // // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). // - // Corresponds with PUT /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/ (the `ClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPut` operationId). - ClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPutWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, body ClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPutJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPutResponse, error) + // Corresponds with POST /api/v2/clusters/{cluster_id}/backups/backup-policies/{policy_id}/detach (the `ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPost` operationId). + ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, policyId openapi_types.UUID, body ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostResponse, error) - // ClustersStoragePoolsVolumesBackupsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdBackupsDeleteWithResponse Clusters:Storage-Pools:Volumes:Backups:Delete + // ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetWithResponse Clusters:Backups:Export // // Returns a wrapper object for the known response body format(s). // - // Corresponds with DELETE /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/backups (the `ClustersStoragePoolsVolumesBackupsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdBackupsDelete` operationId). - ClustersStoragePoolsVolumesBackupsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdBackupsDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesBackupsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdBackupsDeleteResponse, error) + // Corresponds with GET /api/v2/clusters/{cluster_id}/backups/export (the `ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGet` operationId). + ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetParams, reqEditors ...RequestEditorFn) (*ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetResponse, error) - // ClustersStoragePoolsVolumesBackupsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdBackupsGetWithResponse Clusters:Storage-Pools:Volumes:Backups:List + // ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostWithBodyWithResponse Clusters:Backups:Import // - // Returns a wrapper object for the known response body format(s). + // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). // - // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/backups (the `ClustersStoragePoolsVolumesBackupsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdBackupsGet` operationId). - ClustersStoragePoolsVolumesBackupsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdBackupsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesBackupsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdBackupsGetResponse, error) + // Corresponds with POST /api/v2/clusters/{cluster_id}/backups/import (the `ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPost` operationId). + ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse, error) - // ClustersStoragePoolsVolumesCapacityApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdCapacityGetWithResponse Clusters:Storage-Pools:Volumes:Capacity + // ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostWithResponse Clusters:Backups:Import // - // Returns a wrapper object for the known response body format(s). + // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). // - // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/capacity (the `ClustersStoragePoolsVolumesCapacityApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdCapacityGet` operationId). - ClustersStoragePoolsVolumesCapacityApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdCapacityGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, params *ClustersStoragePoolsVolumesCapacityApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdCapacityGetParams, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesCapacityApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdCapacityGetResponse, error) + // Corresponds with POST /api/v2/clusters/{cluster_id}/backups/import (the `ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPost` operationId). + ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, body ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse, error) - // ClustersStoragePoolsVolumesCloneApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdClonePostWithResponse Clusters:Storage-Pools:Volumes:Clone + // ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostWithBodyWithResponse Clusters:Backups:Restore // - // Returns a wrapper object for the known response body format(s). + // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/clone (the `ClustersStoragePoolsVolumesCloneApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdClonePost` operationId). - ClustersStoragePoolsVolumesCloneApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdClonePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, params *ClustersStoragePoolsVolumesCloneApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdClonePostParams, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesCloneApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdClonePostResponse, error) + // Corresponds with POST /api/v2/clusters/{cluster_id}/backups/restore (the `ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePost` operationId). + ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse, error) - // ClustersStoragePoolsVolumesConnectApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdConnectGetWithResponse Clusters:Storage-Pools:Volumes:Connect + // ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostWithResponse Clusters:Backups:Restore // - // Returns a wrapper object for the known response body format(s). + // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). // - // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/connect (the `ClustersStoragePoolsVolumesConnectApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdConnectGet` operationId). - ClustersStoragePoolsVolumesConnectApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdConnectGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, params *ClustersStoragePoolsVolumesConnectApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdConnectGetParams, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesConnectApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdConnectGetResponse, error) + // Corresponds with POST /api/v2/clusters/{cluster_id}/backups/restore (the `ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePost` operationId). + ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, body ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse, error) - // ClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPostWithBodyWithResponse Clusters:Storage-Pools:Volumes:Add-Host + // ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostWithBodyWithResponse Clusters:Backups:Source-Switch // // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/hosts (the `ClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPost` operationId). - ClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPostResponse, error) + // Corresponds with POST /api/v2/clusters/{cluster_id}/backups/source-switch (the `ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPost` operationId). + ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse, error) - // ClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPostWithResponse Clusters:Storage-Pools:Volumes:Add-Host + // ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostWithResponse Clusters:Backups:Source-Switch // // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/hosts (the `ClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPost` operationId). - ClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, body ClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPostResponse, error) + // Corresponds with POST /api/v2/clusters/{cluster_id}/backups/source-switch (the `ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPost` operationId). + ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, body ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse, error) - // ClustersStoragePoolsVolumesRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsHostNqnDeleteWithResponse Clusters:Storage-Pools:Volumes:Remove-Host + // ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetWithResponse Clusters:Backups:Sources // // Returns a wrapper object for the known response body format(s). // - // Corresponds with DELETE /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/hosts/{host_nqn} (the `ClustersStoragePoolsVolumesRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsHostNqnDelete` operationId). - ClustersStoragePoolsVolumesRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsHostNqnDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, hostNqn string, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsHostNqnDeleteResponse, error) + // Corresponds with GET /api/v2/clusters/{cluster_id}/backups/sources (the `ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGet` operationId). + ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse, error) - // ClustersStoragePoolsVolumesGetHostSecretApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsHostNqnSecretGetWithResponse Clusters:Storage-Pools:Volumes:Get-Host-Secret + // ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetWithResponse Clusters:Backups:Detail // // Returns a wrapper object for the known response body format(s). // - // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/hosts/{host_nqn}/secret (the `ClustersStoragePoolsVolumesGetHostSecretApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsHostNqnSecretGet` operationId). - ClustersStoragePoolsVolumesGetHostSecretApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsHostNqnSecretGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, hostNqn string, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesGetHostSecretApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsHostNqnSecretGetResponse, error) + // Corresponds with GET /api/v2/clusters/{cluster_id}/backups/{backup_id}/ (the `ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGet` operationId). + ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, backupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse, error) - // ClustersStoragePoolsVolumesInflateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdInflatePostWithResponse Clusters:Storage-Pools:Volumes:Inflate + // ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteWithResponse Deprecated — delete all backups for a volume + // + // Deprecated. Use `DELETE /clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/backups` instead. // // Returns a wrapper object for the known response body format(s). // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/inflate (the `ClustersStoragePoolsVolumesInflateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdInflatePost` operationId). - ClustersStoragePoolsVolumesInflateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdInflatePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesInflateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdInflatePostResponse, error) + // Corresponds with DELETE /api/v2/clusters/{cluster_id}/backups/{volume_id} (the `ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDelete` operationId). + // + // Deprecated: this operation has been marked as deprecated upstream, but no `x-deprecated-reason` was set + ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse, error) - // ClustersStoragePoolsVolumesIostatsApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdIostatsGetWithResponse Clusters:Storage-Pools:Volumes:Iostats + // ClustersCapacityApiV2ClustersClusterIdCapacityGetWithResponse Clusters:Capacity // // Returns a wrapper object for the known response body format(s). // - // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/iostats (the `ClustersStoragePoolsVolumesIostatsApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdIostatsGet` operationId). - ClustersStoragePoolsVolumesIostatsApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdIostatsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, params *ClustersStoragePoolsVolumesIostatsApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdIostatsGetParams, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesIostatsApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdIostatsGetResponse, error) + // Corresponds with GET /api/v2/clusters/{cluster_id}/capacity (the `ClustersCapacityApiV2ClustersClusterIdCapacityGet` operationId). + ClustersCapacityApiV2ClustersClusterIdCapacityGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersCapacityApiV2ClustersClusterIdCapacityGetParams, reqEditors ...RequestEditorFn) (*ClustersCapacityApiV2ClustersClusterIdCapacityGetResponse, error) - // ClustersStoragePoolsVolumesReplicationDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationGetWithResponse Clusters:Storage-Pools:Volumes:Replication:Detail + // ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetWithResponse Clusters:Consistency-Groups:List // - // Resolve a volume to its counterpart on the other cluster. + // List the cluster's consistency groups, or resolve one by name (§10). // - // Answers "what is the TARGET volume uuid for this SOURCE volume uuid" (and the - // reverse). Before this the ids were only returned by the fail-over or commit - // call itself, so a caller that had not kept them could not find the target - // volume through the API at all -- LVolReplication was exposed nowhere. + // Returns an empty list when ``name`` matches no group, so a caller can probe + // existence without a 404. // // Returns a wrapper object for the known response body format(s). // - // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/ (the `ClustersStoragePoolsVolumesReplicationDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationGet` operationId). - ClustersStoragePoolsVolumesReplicationDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationGetResponse, error) + // Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/ (the `ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGet` operationId). + ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetParams, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetResponse, error) - // ClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPostWithBodyWithResponse Clusters:Storage-Pools:Volumes:Replication:Commit + // ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGetWithResponse Clusters:Consistency-Groups:Detail // - // Queue the planned cutover. Progress is the returned task. + // Returns a wrapper object for the known response body format(s). // - // delete_source=True instructs the task runner to delete the source volume - // after the cutover succeeds. + // Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/ (the `ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGet` operationId). + ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGetResponse, error) + + // ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGetWithResponse Clusters:Consistency-Groups:Members + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/members (the `ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGet` operationId). + ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGetResponse, error) + + // ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostWithBodyWithResponse Clusters:Consistency-Groups:Members:Join + // + // Join an EXISTING volume to the group (design §4.5, Phase 4 late join). + // + // Validates the pinned placement, the pool, the member cap, and the one-way + // rule; a refusal is a 409 naming the precondition. Idempotent: joining a + // current member returns its membership row unchanged. // // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/commit (the `ClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPost` operationId). - ClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPostResponse, error) + // Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/members (the `ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPost` operationId). + ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostResponse, error) - // ClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPostWithResponse Clusters:Storage-Pools:Volumes:Replication:Commit + // ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostWithResponse Clusters:Consistency-Groups:Members:Join // - // Queue the planned cutover. Progress is the returned task. + // Join an EXISTING volume to the group (design §4.5, Phase 4 late join). // - // delete_source=True instructs the task runner to delete the source volume - // after the cutover succeeds. + // Validates the pinned placement, the pool, the member cap, and the one-way + // rule; a refusal is a 409 naming the precondition. Idempotent: joining a + // current member returns its membership row unchanged. // // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/commit (the `ClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPost` operationId). - ClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, body ClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPostResponse, error) + // Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/members (the `ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPost` operationId). + ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, body ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostResponse, error) - // ClustersStoragePoolsVolumesReplicationCutoverProceedApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCutoverProceedPostWithResponse Clusters:Storage-Pools:Volumes:Replication:Cutover-Proceed - // - // Signal that target NVMe paths are connected and cutover may proceed. + // ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDeleteWithResponse Clusters:Consistency-Groups:Members:Detach // - // Called by the operator after its preconnect Job succeeds. The task runner - // is suspended waiting for this signal; once set, it advances to the ANA flip. + // Detach a member: close its epoch one-way, preserving prior generations (§8.2). // // Returns a wrapper object for the known response body format(s). // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/cutover-proceed (the `ClustersStoragePoolsVolumesReplicationCutoverProceedApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCutoverProceedPost` operationId). - ClustersStoragePoolsVolumesReplicationCutoverProceedApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCutoverProceedPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationCutoverProceedApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCutoverProceedPostResponse, error) + // Corresponds with DELETE /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/members/{lvol_id} (the `ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDelete` operationId). + ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, lvolId string, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDeleteResponse, error) - // ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostWithBodyWithResponse Clusters:Storage-Pools:Volumes:Replication:Failback - // - // Point replication back at a source cluster. The cutover itself is - // `commit`. + // ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetWithResponse Clusters:Consistency-Groups:Snapshots:List // - // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). + // Returns a wrapper object for the known response body format(s). // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/failback (the `ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPost` operationId). - ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostResponse, error) + // Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/snapshots (the `ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGet` operationId). + ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetResponse, error) - // ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostWithResponse Clusters:Storage-Pools:Volumes:Replication:Failback + // ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPostWithResponse Clusters:Consistency-Groups:Snapshots:Take // - // Point replication back at a source cluster. The cutover itself is - // `commit`. + // Take one crash-consistent generation across every current member (§5). // - // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). + // Returns a wrapper object for the known response body format(s). // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/failback (the `ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPost` operationId). - ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, body ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostResponse, error) + // Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/snapshots (the `ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPost` operationId). + ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPostResponse, error) - // ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPostWithResponse Clusters:Storage-Pools:Volumes:Replication:Failover + // ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDeleteWithResponse Clusters:Consistency-Groups:Snapshots:Delete // - // Bring the volume up on the target cluster. + // Delete one generation and all its member snapshots; never the group (§10). // - // The counterpart's id is read back from this volume's replication - // relationship, its connection paths from the target volume's `connect`. + // Returns a wrapper object for the known response body format(s). // - // ``generation`` selects WHICH retained point-in-time to come up on: 0 (the - // default) is the newest, 1 the one before it, and so on through the - // history a retention schedule keeps. Failing over to an older generation - // is the recovery path for a logical corruption, which the newest copy has - // faithfully replicated. + // Corresponds with DELETE /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/snapshots/{seq} (the `ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDelete` operationId). + ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, seq int, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDeleteResponse, error) + + // ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGetWithResponse Clusters:Consistency-Groups:Snapshots:Detail // // Returns a wrapper object for the known response body format(s). // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/failover (the `ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPost` operationId). - ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, params *ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPostParams, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPostResponse, error) + // Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/snapshots/{seq} (the `ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGet` operationId). + ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, seq int, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGetResponse, error) - // ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostWithBodyWithResponse Clusters:Storage-Pools:Volumes:Replication:Start + // ClustersExpandApiV2ClustersClusterIdExpandPostWithResponse Clusters:Expand // - // Start replicating a volume. + // Returns a wrapper object for the known response body format(s). // - // The destination is the request's replication_cluster_id, else the cluster's - // configured target. It used to pass the PATH cluster — the volume's OWN - // cluster — as the destination, which self-targets and never falls back to the - // configured target, so replication could not be started correctly over REST - // at all. mode/interval_min were likewise unreachable. + // Corresponds with POST /api/v2/clusters/{cluster_id}/expand (the `ClustersExpandApiV2ClustersClusterIdExpandPost` operationId). + ClustersExpandApiV2ClustersClusterIdExpandPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersExpandApiV2ClustersClusterIdExpandPostResponse, error) + + // ClustersIostatsApiV2ClustersClusterIdIostatsGetWithResponse Clusters:Iostats // - // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). + // Returns a wrapper object for the known response body format(s). // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/start (the `ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPost` operationId). - ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostResponse, error) + // Corresponds with GET /api/v2/clusters/{cluster_id}/iostats (the `ClustersIostatsApiV2ClustersClusterIdIostatsGet` operationId). + ClustersIostatsApiV2ClustersClusterIdIostatsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersIostatsApiV2ClustersClusterIdIostatsGetParams, reqEditors ...RequestEditorFn) (*ClustersIostatsApiV2ClustersClusterIdIostatsGetResponse, error) - // ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostWithResponse Clusters:Storage-Pools:Volumes:Replication:Start + // ClustersLogsApiV2ClustersClusterIdLogsGetWithResponse Clusters:Logs // - // Start replicating a volume. + // Returns a wrapper object for the known response body format(s). // - // The destination is the request's replication_cluster_id, else the cluster's - // configured target. It used to pass the PATH cluster — the volume's OWN - // cluster — as the destination, which self-targets and never falls back to the - // configured target, so replication could not be started correctly over REST - // at all. mode/interval_min were likewise unreachable. + // Corresponds with GET /api/v2/clusters/{cluster_id}/logs (the `ClustersLogsApiV2ClustersClusterIdLogsGet` operationId). + ClustersLogsApiV2ClustersClusterIdLogsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersLogsApiV2ClustersClusterIdLogsGetParams, reqEditors ...RequestEditorFn) (*ClustersLogsApiV2ClustersClusterIdLogsGetResponse, error) + + // ClustersRebalanceApiV2ClustersClusterIdRebalancePostWithResponse Clusters:Rebalance + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/rebalance (the `ClustersRebalanceApiV2ClustersClusterIdRebalancePost` operationId). + ClustersRebalanceApiV2ClustersClusterIdRebalancePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersRebalanceApiV2ClustersClusterIdRebalancePostResponse, error) + + // ClustersReplicationPoliciesListApiV2ClustersClusterIdReplicationPoliciesGetWithResponse Clusters:Replication:Policies:List + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/replication/policies/ (the `ClustersReplicationPoliciesListApiV2ClustersClusterIdReplicationPoliciesGet` operationId). + ClustersReplicationPoliciesListApiV2ClustersClusterIdReplicationPoliciesGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersReplicationPoliciesListApiV2ClustersClusterIdReplicationPoliciesGetResponse, error) + + // ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostWithBodyWithResponse Clusters:Replication:Policies:Create + // + // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/replication/policies/ (the `ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPost` operationId). + ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostParams, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostResponse, error) + + // ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostWithResponse Clusters:Replication:Policies:Create // // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/start (the `ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPost` operationId). - ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, body ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostResponse, error) + // Corresponds with POST /api/v2/clusters/{cluster_id}/replication/policies/ (the `ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPost` operationId). + ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostParams, body ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostResponse, error) - // ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPostWithResponse Clusters:Storage-Pools:Volumes:Replication:Stop + // ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDeleteWithResponse Clusters:Replication:Policies:Delete // // Returns a wrapper object for the known response body format(s). // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/stop (the `ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPost` operationId). - ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPostResponse, error) + // Corresponds with DELETE /api/v2/clusters/{cluster_id}/replication/policies/{policy_id}/ (the `ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDelete` operationId). + ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, policyId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDeleteResponse, error) - // ClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTasksGetWithResponse Clusters:Storage-Pools:Volumes:Replication:Tasks + // ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGetWithResponse Clusters:Replication:Policies:Detail // // Returns a wrapper object for the known response body format(s). // - // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/tasks (the `ClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTasksGet` operationId). - ClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTasksGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTasksGetResponse, error) + // Corresponds with GET /api/v2/clusters/{cluster_id}/replication/policies/{policy_id}/ (the `ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGet` operationId). + ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, policyId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGetResponse, error) - // ClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTriggerPostWithResponse Clusters:Storage-Pools:Volumes:Replication:Trigger + // ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPostWithResponse Clusters:Replication:Policies:Failover // // Returns a wrapper object for the known response body format(s). // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/trigger (the `ClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTriggerPost` operationId). - ClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTriggerPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTriggerPostResponse, error) + // Corresponds with POST /api/v2/clusters/{cluster_id}/replication/policies/{policy_id}/failover (the `ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPost` operationId). + ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, policyId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPostResponse, error) - // ClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsGetWithResponse Clusters:Storage-Pools:Volumes:Snapshots:List + // ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGetWithResponse Clusters:Replication:Policies:Latest-Generation + // + // The consistency group's newest fully replicated generation, every + // member as a cloneable object on the secondary. Refused as a 400 when the + // policy has no consistency group, when no generation is complete for + // every current member yet, or when members are already split across + // generations: the same refusal a real group fail-over applies, so a drill + // never addresses a mixed-generation cut. // // Returns a wrapper object for the known response body format(s). // - // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/snapshots (the `ClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsGet` operationId). - ClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsGetResponse, error) + // Corresponds with GET /api/v2/clusters/{cluster_id}/replication/policies/{policy_id}/latest-generation (the `ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGet` operationId). + ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, policyId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGetResponse, error) - // ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostWithBodyWithResponse Clusters:Storage-Pools:Volumes:Snapshots:Create + // ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetWithResponse Clusters:Replication:Relationships:Detail + // + // Replication relationship for a volume, resolvable even when the source volume + // has been deleted (e.g. after replication-commit --delete-source). The CSI driver + // uses this to redirect NodeStageVolume to the active volume on the target cluster. + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/replication/relationships/{lvol_id} (the `ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGet` operationId). + ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, lvolId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetResponse, error) + + // ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGetWithResponse Clusters:Replication:Relationships:Latest-Snapshot + // + // The volume's newest fully replicated snapshot, on the secondary, as a + // cloneable object. Exists for the volume's whole replicated life: a + // test-failover drill (design §14) resolves its test point through this + // read, without touching the real replication state to find out what it is. + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/replication/relationships/{lvol_id}/latest-snapshot (the `ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGet` operationId). + ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, lvolId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGetResponse, error) + + // ClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGetWithResponse Clusters:Replication:Targets:List + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/replication/targets/ (the `ClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGet` operationId). + ClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGetResponse, error) + + // ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostWithBodyWithResponse Clusters:Replication:Targets:Create // // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/snapshots (the `ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPost` operationId). - ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostResponse, error) + // Corresponds with POST /api/v2/clusters/{cluster_id}/replication/targets/ (the `ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPost` operationId). + ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostParams, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostResponse, error) - // ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostWithResponse Clusters:Storage-Pools:Volumes:Snapshots:Create + // ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostWithResponse Clusters:Replication:Targets:Create // // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). // - // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/snapshots (the `ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPost` operationId). - ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, body ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostResponse, error) + // Corresponds with POST /api/v2/clusters/{cluster_id}/replication/targets/ (the `ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPost` operationId). + ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostParams, body ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPostResponse, error) - // ClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigrationsGetWithResponse Clusters:Subsystems:Migrations:List + // ClustersReplicationTargetsDeleteApiV2ClustersClusterIdReplicationTargetsTargetIdDeleteWithResponse Clusters:Replication:Targets:Delete // // Returns a wrapper object for the known response body format(s). // - // Corresponds with GET /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/ (the `ClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigrationsGet` operationId). - ClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigrationsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigrationsGetResponse, error) + // Corresponds with DELETE /api/v2/clusters/{cluster_id}/replication/targets/{target_id}/ (the `ClustersReplicationTargetsDeleteApiV2ClustersClusterIdReplicationTargetsTargetIdDelete` operationId). + ClustersReplicationTargetsDeleteApiV2ClustersClusterIdReplicationTargetsTargetIdDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, targetId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersReplicationTargetsDeleteApiV2ClustersClusterIdReplicationTargetsTargetIdDeleteResponse, error) - // ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostWithBodyWithResponse Clusters:Subsystems:Migrations:Create + // ClustersReplicationTargetsDetailApiV2ClustersClusterIdReplicationTargetsTargetIdGetWithResponse Clusters:Replication:Targets:Detail + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/replication/targets/{target_id}/ (the `ClustersReplicationTargetsDetailApiV2ClustersClusterIdReplicationTargetsTargetIdGet` operationId). + ClustersReplicationTargetsDetailApiV2ClustersClusterIdReplicationTargetsTargetIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, targetId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersReplicationTargetsDetailApiV2ClustersClusterIdReplicationTargetsTargetIdGetResponse, error) + + // ClustersReplicationTargetsFailoverApiV2ClustersClusterIdReplicationTargetsTargetIdFailoverPostWithResponse Clusters:Replication:Targets:Failover + // + // Fail over EVERY volume replicating to this target. + // + // A site loss has to move all volumes at once; doing it volume by volume was + // the only option before. Idempotent per volume and reports one result per + // volume, so a partial failure is visible instead of silent. + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/replication/targets/{target_id}/failover (the `ClustersReplicationTargetsFailoverApiV2ClustersClusterIdReplicationTargetsTargetIdFailoverPost` operationId). + ClustersReplicationTargetsFailoverApiV2ClustersClusterIdReplicationTargetsTargetIdFailoverPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, targetId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersReplicationTargetsFailoverApiV2ClustersClusterIdReplicationTargetsTargetIdFailoverPostResponse, error) + + // ClustersShutdownApiV2ClustersClusterIdShutdownPostWithResponse Clusters:Shutdown + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/shutdown (the `ClustersShutdownApiV2ClustersClusterIdShutdownPost` operationId). + ClustersShutdownApiV2ClustersClusterIdShutdownPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersShutdownApiV2ClustersClusterIdShutdownPostResponse, error) + + // ClustersStartApiV2ClustersClusterIdStartPostWithResponse Clusters:Start + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/start (the `ClustersStartApiV2ClustersClusterIdStartPost` operationId). + ClustersStartApiV2ClustersClusterIdStartPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStartApiV2ClustersClusterIdStartPostResponse, error) + + // ClustersStorageNodesListApiV2ClustersClusterIdStorageNodesGetWithResponse Clusters:Storage-Nodes:List + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-nodes/ (the `ClustersStorageNodesListApiV2ClustersClusterIdStorageNodesGet` operationId). + ClustersStorageNodesListApiV2ClustersClusterIdStorageNodesGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersStorageNodesListApiV2ClustersClusterIdStorageNodesGetParams, reqEditors ...RequestEditorFn) (*ClustersStorageNodesListApiV2ClustersClusterIdStorageNodesGetResponse, error) + + // ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostWithBodyWithResponse Clusters:Storage-Nodes:Create // // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). // - // Corresponds with POST /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/ (the `ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPost` operationId). - ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, params *ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostParams, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostResponse, error) + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-nodes/ (the `ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPost` operationId). + ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostParams, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostResponse, error) - // ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostWithResponse Clusters:Subsystems:Migrations:Create + // ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostWithResponse Clusters:Storage-Nodes:Create // // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). // - // Corresponds with POST /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/ (the `ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPost` operationId). - ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, params *ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostParams, body ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostResponse, error) + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-nodes/ (the `ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPost` operationId). + ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostParams, body ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStorageNodesCreateApiV2ClustersClusterIdStorageNodesPostResponse, error) - // ClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdDeleteWithResponse Clusters:Subsystems:Migrations:Cancel + // ClustersStorageNodesDeleteApiV2ClustersClusterIdStorageNodesStorageNodeIdDeleteWithResponse Clusters:Storage-Nodes:Delete // // Returns a wrapper object for the known response body format(s). // - // Corresponds with DELETE /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/{migration_id}/ (the `ClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdDelete` operationId). - ClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdDeleteResponse, error) + // Corresponds with DELETE /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/ (the `ClustersStorageNodesDeleteApiV2ClustersClusterIdStorageNodesStorageNodeIdDelete` operationId). + ClustersStorageNodesDeleteApiV2ClustersClusterIdStorageNodesStorageNodeIdDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, params *ClustersStorageNodesDeleteApiV2ClustersClusterIdStorageNodesStorageNodeIdDeleteParams, reqEditors ...RequestEditorFn) (*ClustersStorageNodesDeleteApiV2ClustersClusterIdStorageNodesStorageNodeIdDeleteResponse, error) - // ClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdGetWithResponse Clusters:Subsystems:Migrations:Detail + // ClustersStorageNodesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdGetWithResponse Clusters:Storage-Nodes:Detail // // Returns a wrapper object for the known response body format(s). // - // Corresponds with GET /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/{migration_id}/ (the `ClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdGet` operationId). - ClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdGetResponse, error) + // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/ (the `ClustersStorageNodesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdGet` operationId). + ClustersStorageNodesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, params *ClustersStorageNodesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdGetParams, reqEditors ...RequestEditorFn) (*ClustersStorageNodesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdGetResponse, error) - // ClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdCleanupTargetPostWithResponse Clusters:Subsystems:Migrations:Cleanup-Target + // ClustersStorageNodesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdCapacityGetWithResponse Clusters:Storage-Nodes:Capacity // - // Idempotently remove every object this migration created on the target - // node(s). Only defined for a single-lvol migration; batch migration groups - // have no cleanup-target equivalent at the group level. + // Returns a wrapper object for the known response body format(s). // - // Safe to call at any migration state — objects not found are reported as - // already cleaned up rather than as errors. + // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/capacity (the `ClustersStorageNodesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdCapacityGet` operationId). + ClustersStorageNodesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdCapacityGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, params *ClustersStorageNodesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdCapacityGetParams, reqEditors ...RequestEditorFn) (*ClustersStorageNodesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdCapacityGetResponse, error) + + // ClustersStorageNodesDevicesListApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesGetWithResponse Clusters:Storage Nodes:Devices:List // // Returns a wrapper object for the known response body format(s). // - // Corresponds with POST /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/{migration_id}/cleanup-target (the `ClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdCleanupTargetPost` operationId). - ClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdCleanupTargetPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdCleanupTargetPostResponse, error) + // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/devices/ (the `ClustersStorageNodesDevicesListApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesGet` operationId). + ClustersStorageNodesDevicesListApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, params *ClustersStorageNodesDevicesListApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesGetParams, reqEditors ...RequestEditorFn) (*ClustersStorageNodesDevicesListApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesGetResponse, error) - // ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostWithBodyWithResponse Clusters:Subsystems:Migrations:Continue + // ClustersStorageNodesDevicesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdGetWithResponse Clusters:Storage Nodes:Devices:Detail + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/devices/{device_id}/ (the `ClustersStorageNodesDevicesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdGet` operationId). + ClustersStorageNodesDevicesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, deviceId openapi_types.UUID, params *ClustersStorageNodesDevicesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdGetParams, reqEditors ...RequestEditorFn) (*ClustersStorageNodesDevicesDetailApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdGetResponse, error) + + // ClustersStorageNodesDevicesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdCapacityGetWithResponse Clusters:Storage Nodes:Devices:Capacity + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/devices/{device_id}/capacity (the `ClustersStorageNodesDevicesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdCapacityGet` operationId). + ClustersStorageNodesDevicesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdCapacityGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, deviceId openapi_types.UUID, params *ClustersStorageNodesDevicesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdCapacityGetParams, reqEditors ...RequestEditorFn) (*ClustersStorageNodesDevicesCapacityApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdCapacityGetResponse, error) + + // ClustersStorageNodesDevicesGetDeviceHealthInfoApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdHealthInfoGetWithResponse Clusters:Storage Nodes:Devices:Get-Device-Health-Info + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/devices/{device_id}/health-info (the `ClustersStorageNodesDevicesGetDeviceHealthInfoApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdHealthInfoGet` operationId). + ClustersStorageNodesDevicesGetDeviceHealthInfoApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdHealthInfoGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, deviceId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStorageNodesDevicesGetDeviceHealthInfoApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdHealthInfoGetResponse, error) + + // ClustersStorageNodesDevicesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdIostatsGetWithResponse Clusters:Storage Nodes:Devices:Iostats + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/devices/{device_id}/iostats (the `ClustersStorageNodesDevicesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdIostatsGet` operationId). + ClustersStorageNodesDevicesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdIostatsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, deviceId openapi_types.UUID, params *ClustersStorageNodesDevicesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdIostatsGetParams, reqEditors ...RequestEditorFn) (*ClustersStorageNodesDevicesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdIostatsGetResponse, error) + + // ClustersStorageNodesDevicesRemoveApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRemovePostWithResponse Clusters:Storage Nodes:Devices:Remove + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/devices/{device_id}/remove (the `ClustersStorageNodesDevicesRemoveApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRemovePost` operationId). + ClustersStorageNodesDevicesRemoveApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRemovePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, deviceId openapi_types.UUID, params *ClustersStorageNodesDevicesRemoveApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRemovePostParams, reqEditors ...RequestEditorFn) (*ClustersStorageNodesDevicesRemoveApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRemovePostResponse, error) + + // ClustersStorageNodesDevicesResetApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdResetPostWithResponse Clusters:Storage Nodes:Devices:Reset + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/devices/{device_id}/reset (the `ClustersStorageNodesDevicesResetApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdResetPost` operationId). + ClustersStorageNodesDevicesResetApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdResetPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, deviceId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStorageNodesDevicesResetApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdResetPostResponse, error) + + // ClustersStorageNodesDevicesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRestartPostWithResponse Clusters:Storage Nodes:Devices:Restart + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/devices/{device_id}/restart (the `ClustersStorageNodesDevicesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRestartPost` operationId). + ClustersStorageNodesDevicesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRestartPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, deviceId openapi_types.UUID, params *ClustersStorageNodesDevicesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRestartPostParams, reqEditors ...RequestEditorFn) (*ClustersStorageNodesDevicesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdDevicesDeviceIdRestartPostResponse, error) + + // ClustersStorageNodesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdIostatsGetWithResponse Clusters:Storage-Nodes:Iostats + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/iostats (the `ClustersStorageNodesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdIostatsGet` operationId). + ClustersStorageNodesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdIostatsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, params *ClustersStorageNodesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdIostatsGetParams, reqEditors ...RequestEditorFn) (*ClustersStorageNodesIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdIostatsGetResponse, error) + + // ClustersStorageNodesNicsListApiV2ClustersClusterIdStorageNodesStorageNodeIdNicsGetWithResponse Clusters:Storage-Nodes:Nics:List + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/nics (the `ClustersStorageNodesNicsListApiV2ClustersClusterIdStorageNodesStorageNodeIdNicsGet` operationId). + ClustersStorageNodesNicsListApiV2ClustersClusterIdStorageNodesStorageNodeIdNicsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStorageNodesNicsListApiV2ClustersClusterIdStorageNodesStorageNodeIdNicsGetResponse, error) + + // ClustersStorageNodesNicsIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdNicsNicIdIostatsGetWithResponse Clusters:Storage-Nodes:Nics:Iostats + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/nics/{nic_id}/iostats (the `ClustersStorageNodesNicsIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdNicsNicIdIostatsGet` operationId). + ClustersStorageNodesNicsIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdNicsNicIdIostatsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, nicId string, reqEditors ...RequestEditorFn) (*ClustersStorageNodesNicsIostatsApiV2ClustersClusterIdStorageNodesStorageNodeIdNicsNicIdIostatsGetResponse, error) + + // ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdPromotePostWithResponse Clusters:Storage-Nodes:Start + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/promote (the `ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdPromotePost` operationId). + ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdPromotePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdPromotePostResponse, error) + + // ClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPostWithBodyWithResponse Clusters:Storage-Nodes:Restart + // + // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/restart (the `ClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPost` operationId). + ClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPostResponse, error) + + // ClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPostWithResponse Clusters:Storage-Nodes:Restart + // + // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/restart (the `ClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPost` operationId). + ClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, body ClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStorageNodesRestartApiV2ClustersClusterIdStorageNodesStorageNodeIdRestartPostResponse, error) + + // ClustersStorageNodesResumeApiV2ClustersClusterIdStorageNodesStorageNodeIdResumePostWithResponse Clusters:Storage-Nodes:Resume + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/resume (the `ClustersStorageNodesResumeApiV2ClustersClusterIdStorageNodesStorageNodeIdResumePost` operationId). + ClustersStorageNodesResumeApiV2ClustersClusterIdStorageNodesStorageNodeIdResumePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStorageNodesResumeApiV2ClustersClusterIdStorageNodesStorageNodeIdResumePostResponse, error) + + // ClustersStorageNodesShutdownApiV2ClustersClusterIdStorageNodesStorageNodeIdShutdownPostWithResponse Clusters:Storage-Nodes:Shutdown + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/shutdown (the `ClustersStorageNodesShutdownApiV2ClustersClusterIdStorageNodesStorageNodeIdShutdownPost` operationId). + ClustersStorageNodesShutdownApiV2ClustersClusterIdStorageNodesStorageNodeIdShutdownPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, params *ClustersStorageNodesShutdownApiV2ClustersClusterIdStorageNodesStorageNodeIdShutdownPostParams, reqEditors ...RequestEditorFn) (*ClustersStorageNodesShutdownApiV2ClustersClusterIdStorageNodesStorageNodeIdShutdownPostResponse, error) + + // ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPostWithBodyWithResponse Clusters:Storage-Nodes:Start + // + // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/start (the `ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPost` operationId). + ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPostResponse, error) + + // ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPostWithResponse Clusters:Storage-Nodes:Start + // + // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/start (the `ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPost` operationId). + ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, body ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStorageNodesStartApiV2ClustersClusterIdStorageNodesStorageNodeIdStartPostResponse, error) + + // ClustersStorageNodesSuspendApiV2ClustersClusterIdStorageNodesStorageNodeIdSuspendPostWithResponse Clusters:Storage-Nodes:Suspend + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-nodes/{storage_node_id}/suspend (the `ClustersStorageNodesSuspendApiV2ClustersClusterIdStorageNodesStorageNodeIdSuspendPost` operationId). + ClustersStorageNodesSuspendApiV2ClustersClusterIdStorageNodesStorageNodeIdSuspendPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, storageNodeId openapi_types.UUID, params *ClustersStorageNodesSuspendApiV2ClustersClusterIdStorageNodesStorageNodeIdSuspendPostParams, reqEditors ...RequestEditorFn) (*ClustersStorageNodesSuspendApiV2ClustersClusterIdStorageNodesStorageNodeIdSuspendPostResponse, error) + + // ClustersStoragePoolsListApiV2ClustersClusterIdStoragePoolsGetWithResponse Clusters:Storage-Pools:List + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/ (the `ClustersStoragePoolsListApiV2ClustersClusterIdStoragePoolsGet` operationId). + ClustersStoragePoolsListApiV2ClustersClusterIdStoragePoolsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersStoragePoolsListApiV2ClustersClusterIdStoragePoolsGetParams, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsListApiV2ClustersClusterIdStoragePoolsGetResponse, error) + + // ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostWithBodyWithResponse Clusters:Storage-Pools:Create + // + // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/ (the `ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPost` operationId). + ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostParams, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostResponse, error) + + // ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostWithResponse Clusters:Storage-Pools:Create + // + // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/ (the `ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPost` operationId). + ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostParams, body ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsCreateApiV2ClustersClusterIdStoragePoolsPostResponse, error) + + // ClustersStoragePoolsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdDeleteWithResponse Clusters:Storage-Pools:Delete + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with DELETE /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/ (the `ClustersStoragePoolsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdDelete` operationId). + ClustersStoragePoolsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdDeleteResponse, error) + + // ClustersStoragePoolsDetailApiV2ClustersClusterIdStoragePoolsPoolIdGetWithResponse Clusters:Storage-Pools:Detail + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/ (the `ClustersStoragePoolsDetailApiV2ClustersClusterIdStoragePoolsPoolIdGet` operationId). + ClustersStoragePoolsDetailApiV2ClustersClusterIdStoragePoolsPoolIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, params *ClustersStoragePoolsDetailApiV2ClustersClusterIdStoragePoolsPoolIdGetParams, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsDetailApiV2ClustersClusterIdStoragePoolsPoolIdGetResponse, error) + + // ClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutWithBodyWithResponse Clusters:Storage-Pools:Update + // + // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with PUT /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/ (the `ClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPut` operationId). + ClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutResponse, error) + + // ClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutWithResponse Clusters:Storage-Pools:Update + // + // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with PUT /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/ (the `ClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPut` operationId). + ClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, body ClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsUpdateApiV2ClustersClusterIdStoragePoolsPoolIdPutResponse, error) + + // ClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDeleteWithBodyWithResponse Clusters:Storage-Pools:Remove-Host + // + // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with DELETE /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/host (the `ClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDelete` operationId). + ClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDeleteWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDeleteResponse, error) + + // ClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDeleteWithResponse Clusters:Storage-Pools:Remove-Host + // + // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with DELETE /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/host (the `ClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDelete` operationId). + ClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, body ClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDeleteJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdHostDeleteResponse, error) + + // ClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPostWithBodyWithResponse Clusters:Storage-Pools:Add-Host + // + // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/host (the `ClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPost` operationId). + ClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPostResponse, error) + + // ClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPostWithResponse Clusters:Storage-Pools:Add-Host + // + // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/host (the `ClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPost` operationId). + ClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, body ClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsAddHostApiV2ClustersClusterIdStoragePoolsPoolIdHostPostResponse, error) + + // ClustersStoragePoolsIostatsApiV2ClustersClusterIdStoragePoolsPoolIdIostatsGetWithResponse Clusters:Storage-Pools:Iostats + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/iostats (the `ClustersStoragePoolsIostatsApiV2ClustersClusterIdStoragePoolsPoolIdIostatsGet` operationId). + ClustersStoragePoolsIostatsApiV2ClustersClusterIdStoragePoolsPoolIdIostatsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, params *ClustersStoragePoolsIostatsApiV2ClustersClusterIdStoragePoolsPoolIdIostatsGetParams, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsIostatsApiV2ClustersClusterIdStoragePoolsPoolIdIostatsGetResponse, error) + + // ClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsGetWithResponse Clusters:Storage-Pools:Snapshots:List + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/snapshots/ (the `ClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsGet` operationId). + ClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, params *ClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsGetParams, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsGetResponse, error) + + // ClustersStoragePoolsSnapshotsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdDeleteWithResponse Clusters:Storage-Pools:Snapshots:Delete + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with DELETE /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/snapshots/{snapshot_id}/ (the `ClustersStoragePoolsSnapshotsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdDelete` operationId). + ClustersStoragePoolsSnapshotsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, snapshotId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsSnapshotsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdDeleteResponse, error) + + // ClustersStoragePoolsSnapshotsDetailApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdGetWithResponse Clusters:Storage-Pools:Snapshots:Detail + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/snapshots/{snapshot_id}/ (the `ClustersStoragePoolsSnapshotsDetailApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdGet` operationId). + ClustersStoragePoolsSnapshotsDetailApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, snapshotId openapi_types.UUID, params *ClustersStoragePoolsSnapshotsDetailApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdGetParams, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsSnapshotsDetailApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdGetResponse, error) + + // ClustersStoragePoolsVolumesListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesGetWithResponse Clusters:Storage-Pools:Volumes:List + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/ (the `ClustersStoragePoolsVolumesListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesGet` operationId). + ClustersStoragePoolsVolumesListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, params *ClustersStoragePoolsVolumesListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesGetParams, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesGetResponse, error) + + // ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostWithBodyWithResponse Clusters:Storage-Pools:Volumes:Create + // + // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/ (the `ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPost` operationId). + ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, params *ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostParams, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostResponse, error) + + // ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostWithResponse Clusters:Storage-Pools:Volumes:Create + // + // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/ (the `ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPost` operationId). + ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, params *ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostParams, body ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesPostResponse, error) + + // ClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPostWithBodyWithResponse Clusters:Storage-Pools:Replicate Lvol On Source Cluster + // + // Rebuild a volume on the source cluster. + // + // Collection-scoped rather than an operation on `/{volume_id}`: the volume is + // typically gone from the source cluster by the time this is called, so the + // controller falls back to resolving the id through the replication records. + // + // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/replicate_lvol_on_source_cluster (the `ClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPost` operationId). + ClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPostResponse, error) + + // ClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPostWithResponse Clusters:Storage-Pools:Replicate Lvol On Source Cluster + // + // Rebuild a volume on the source cluster. + // + // Collection-scoped rather than an operation on `/{volume_id}`: the volume is + // typically gone from the source cluster by the time this is called, so the + // controller falls back to resolving the id through the replication records. + // + // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/replicate_lvol_on_source_cluster (the `ClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPost` operationId). + ClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, body ClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsReplicateLvolOnSourceClusterApiV2ClustersClusterIdStoragePoolsPoolIdVolumesReplicateLvolOnSourceClusterPostResponse, error) + + // ClustersStoragePoolsVolumesDeleteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdDeleteWithResponse Clusters:Storage-Pools:Volumes:Delete + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with DELETE /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/ (the `ClustersStoragePoolsVolumesDeleteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdDelete` operationId). + ClustersStoragePoolsVolumesDeleteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesDeleteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdDeleteResponse, error) + + // ClustersStoragePoolsVolumesDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdGetWithResponse Clusters:Storage-Pools:Volumes:Detail + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/ (the `ClustersStoragePoolsVolumesDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdGet` operationId). + ClustersStoragePoolsVolumesDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, params *ClustersStoragePoolsVolumesDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdGetParams, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdGetResponse, error) + + // ClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPutWithBodyWithResponse Clusters:Storage-Pools:Volumes:Update + // + // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with PUT /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/ (the `ClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPut` operationId). + ClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPutWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPutResponse, error) + + // ClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPutWithResponse Clusters:Storage-Pools:Volumes:Update + // + // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with PUT /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/ (the `ClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPut` operationId). + ClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPutWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, body ClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPutJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesUpdateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdPutResponse, error) + + // ClustersStoragePoolsVolumesBackupsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdBackupsDeleteWithResponse Clusters:Storage-Pools:Volumes:Backups:Delete + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with DELETE /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/backups (the `ClustersStoragePoolsVolumesBackupsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdBackupsDelete` operationId). + ClustersStoragePoolsVolumesBackupsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdBackupsDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesBackupsDeleteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdBackupsDeleteResponse, error) + + // ClustersStoragePoolsVolumesBackupsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdBackupsGetWithResponse Clusters:Storage-Pools:Volumes:Backups:List + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/backups (the `ClustersStoragePoolsVolumesBackupsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdBackupsGet` operationId). + ClustersStoragePoolsVolumesBackupsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdBackupsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesBackupsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdBackupsGetResponse, error) + + // ClustersStoragePoolsVolumesCapacityApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdCapacityGetWithResponse Clusters:Storage-Pools:Volumes:Capacity + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/capacity (the `ClustersStoragePoolsVolumesCapacityApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdCapacityGet` operationId). + ClustersStoragePoolsVolumesCapacityApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdCapacityGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, params *ClustersStoragePoolsVolumesCapacityApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdCapacityGetParams, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesCapacityApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdCapacityGetResponse, error) + + // ClustersStoragePoolsVolumesCloneApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdClonePostWithResponse Clusters:Storage-Pools:Volumes:Clone + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/clone (the `ClustersStoragePoolsVolumesCloneApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdClonePost` operationId). + ClustersStoragePoolsVolumesCloneApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdClonePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, params *ClustersStoragePoolsVolumesCloneApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdClonePostParams, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesCloneApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdClonePostResponse, error) + + // ClustersStoragePoolsVolumesConnectApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdConnectGetWithResponse Clusters:Storage-Pools:Volumes:Connect + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/connect (the `ClustersStoragePoolsVolumesConnectApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdConnectGet` operationId). + ClustersStoragePoolsVolumesConnectApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdConnectGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, params *ClustersStoragePoolsVolumesConnectApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdConnectGetParams, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesConnectApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdConnectGetResponse, error) + + // ClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPostWithBodyWithResponse Clusters:Storage-Pools:Volumes:Add-Host + // + // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/hosts (the `ClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPost` operationId). + ClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPostResponse, error) + + // ClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPostWithResponse Clusters:Storage-Pools:Volumes:Add-Host + // + // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/hosts (the `ClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPost` operationId). + ClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, body ClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesAddHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsPostResponse, error) + + // ClustersStoragePoolsVolumesRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsHostNqnDeleteWithResponse Clusters:Storage-Pools:Volumes:Remove-Host + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with DELETE /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/hosts/{host_nqn} (the `ClustersStoragePoolsVolumesRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsHostNqnDelete` operationId). + ClustersStoragePoolsVolumesRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsHostNqnDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, hostNqn string, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesRemoveHostApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsHostNqnDeleteResponse, error) + + // ClustersStoragePoolsVolumesGetHostSecretApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsHostNqnSecretGetWithResponse Clusters:Storage-Pools:Volumes:Get-Host-Secret + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/hosts/{host_nqn}/secret (the `ClustersStoragePoolsVolumesGetHostSecretApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsHostNqnSecretGet` operationId). + ClustersStoragePoolsVolumesGetHostSecretApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsHostNqnSecretGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, hostNqn string, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesGetHostSecretApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdHostsHostNqnSecretGetResponse, error) + + // ClustersStoragePoolsVolumesInflateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdInflatePostWithResponse Clusters:Storage-Pools:Volumes:Inflate + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/inflate (the `ClustersStoragePoolsVolumesInflateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdInflatePost` operationId). + ClustersStoragePoolsVolumesInflateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdInflatePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesInflateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdInflatePostResponse, error) + + // ClustersStoragePoolsVolumesIostatsApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdIostatsGetWithResponse Clusters:Storage-Pools:Volumes:Iostats + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/iostats (the `ClustersStoragePoolsVolumesIostatsApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdIostatsGet` operationId). + ClustersStoragePoolsVolumesIostatsApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdIostatsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, params *ClustersStoragePoolsVolumesIostatsApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdIostatsGetParams, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesIostatsApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdIostatsGetResponse, error) + + // ClustersStoragePoolsVolumesReplicationDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationGetWithResponse Clusters:Storage-Pools:Volumes:Replication:Detail + // + // Resolve a volume to its counterpart on the other cluster. + // + // Answers "what is the TARGET volume uuid for this SOURCE volume uuid" (and the + // reverse). Before this the ids were only returned by the fail-over or commit + // call itself, so a caller that had not kept them could not find the target + // volume through the API at all -- LVolReplication was exposed nowhere. + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/ (the `ClustersStoragePoolsVolumesReplicationDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationGet` operationId). + ClustersStoragePoolsVolumesReplicationDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationGetResponse, error) + + // ClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPostWithBodyWithResponse Clusters:Storage-Pools:Volumes:Replication:Commit + // + // Queue the planned cutover. Progress is the returned task. + // + // delete_source=True instructs the task runner to delete the source volume + // after the cutover succeeds. + // + // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/commit (the `ClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPost` operationId). + ClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPostResponse, error) + + // ClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPostWithResponse Clusters:Storage-Pools:Volumes:Replication:Commit + // + // Queue the planned cutover. Progress is the returned task. + // + // delete_source=True instructs the task runner to delete the source volume + // after the cutover succeeds. + // + // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/commit (the `ClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPost` operationId). + ClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, body ClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCommitPostResponse, error) + + // ClustersStoragePoolsVolumesReplicationCutoverProceedApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCutoverProceedPostWithResponse Clusters:Storage-Pools:Volumes:Replication:Cutover-Proceed + // + // Signal that target NVMe paths are connected and cutover may proceed. + // + // Called by the operator after its preconnect Job succeeds. The task runner + // is suspended waiting for this signal; once set, it advances to the ANA flip. + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/cutover-proceed (the `ClustersStoragePoolsVolumesReplicationCutoverProceedApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCutoverProceedPost` operationId). + ClustersStoragePoolsVolumesReplicationCutoverProceedApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCutoverProceedPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationCutoverProceedApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCutoverProceedPostResponse, error) + + // ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostWithBodyWithResponse Clusters:Storage-Pools:Volumes:Replication:Failback + // + // Point replication back at a source cluster. The cutover itself is + // `commit`. + // + // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/failback (the `ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPost` operationId). + ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostResponse, error) + + // ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostWithResponse Clusters:Storage-Pools:Volumes:Replication:Failback + // + // Point replication back at a source cluster. The cutover itself is + // `commit`. + // + // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/failback (the `ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPost` operationId). + ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, body ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostResponse, error) + + // ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPostWithResponse Clusters:Storage-Pools:Volumes:Replication:Failover + // + // Bring the volume up on the target cluster. + // + // The counterpart's id is read back from this volume's replication + // relationship, its connection paths from the target volume's `connect`. + // + // ``generation`` selects WHICH retained point-in-time to come up on: 0 (the + // default) is the newest, 1 the one before it, and so on through the + // history a retention schedule keeps. Failing over to an older generation + // is the recovery path for a logical corruption, which the newest copy has + // faithfully replicated. + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/failover (the `ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPost` operationId). + ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, params *ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPostParams, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPostResponse, error) + + // ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostWithBodyWithResponse Clusters:Storage-Pools:Volumes:Replication:Start + // + // Start replicating a volume. + // + // The destination is the request's replication_cluster_id, else the cluster's + // configured target. It used to pass the PATH cluster — the volume's OWN + // cluster — as the destination, which self-targets and never falls back to the + // configured target, so replication could not be started correctly over REST + // at all. mode/interval_min were likewise unreachable. + // + // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/start (the `ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPost` operationId). + ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostResponse, error) + + // ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostWithResponse Clusters:Storage-Pools:Volumes:Replication:Start + // + // Start replicating a volume. + // + // The destination is the request's replication_cluster_id, else the cluster's + // configured target. It used to pass the PATH cluster — the volume's OWN + // cluster — as the destination, which self-targets and never falls back to the + // configured target, so replication could not be started correctly over REST + // at all. mode/interval_min were likewise unreachable. + // + // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/start (the `ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPost` operationId). + ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, body ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostResponse, error) + + // ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGetWithResponse Clusters:Storage-Pools:Volumes:Replication:Status + // + // The typed steady-state replication status. + // + // Unlike the relationship read above, which serves cutover records and 404s + // for a volume's whole healthy replicated life, this endpoint always answers + // for a volume that exists: ``state: not_replicating, role: none`` is the + // valid answer for an unreplicated volume. The csi-addons adapter derives + // its conditions and ``lastSyncTime`` from this read on every reconcile. + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/status (the `ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGet` operationId). + ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGetResponse, error) + + // ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPostWithResponse Clusters:Storage-Pools:Volumes:Replication:Stop + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/stop (the `ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPost` operationId). + ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPostResponse, error) + + // ClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTasksGetWithResponse Clusters:Storage-Pools:Volumes:Replication:Tasks + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/tasks (the `ClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTasksGet` operationId). + ClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTasksGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTasksGetResponse, error) + + // ClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTriggerPostWithResponse Clusters:Storage-Pools:Volumes:Replication:Trigger + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/trigger (the `ClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTriggerPost` operationId). + ClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTriggerPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTriggerPostResponse, error) + + // ClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsGetWithResponse Clusters:Storage-Pools:Volumes:Snapshots:List + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/snapshots (the `ClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsGet` operationId). + ClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsGetResponse, error) + + // ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostWithBodyWithResponse Clusters:Storage-Pools:Volumes:Snapshots:Create + // + // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/snapshots (the `ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPost` operationId). + ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostResponse, error) + + // ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostWithResponse Clusters:Storage-Pools:Volumes:Snapshots:Create + // + // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/snapshots (the `ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPost` operationId). + ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, body ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostResponse, error) + + // ClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigrationsGetWithResponse Clusters:Subsystems:Migrations:List + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/ (the `ClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigrationsGet` operationId). + ClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigrationsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigrationsGetResponse, error) + + // ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostWithBodyWithResponse Clusters:Subsystems:Migrations:Create + // + // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/ (the `ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPost` operationId). + ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, params *ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostParams, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostResponse, error) + + // ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostWithResponse Clusters:Subsystems:Migrations:Create + // + // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/ (the `ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPost` operationId). + ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, params *ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostParams, body ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostResponse, error) + + // ClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdDeleteWithResponse Clusters:Subsystems:Migrations:Cancel + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with DELETE /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/{migration_id}/ (the `ClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdDelete` operationId). + ClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdDeleteResponse, error) + + // ClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdGetWithResponse Clusters:Subsystems:Migrations:Detail + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/{migration_id}/ (the `ClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdGet` operationId). + ClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdGetResponse, error) + + // ClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdCleanupTargetPostWithResponse Clusters:Subsystems:Migrations:Cleanup-Target + // + // Idempotently remove every object this migration created on the target + // node(s). Only defined for a single-lvol migration; batch migration groups + // have no cleanup-target equivalent at the group level. + // + // Safe to call at any migration state — objects not found are reported as + // already cleaned up rather than as errors. + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/{migration_id}/cleanup-target (the `ClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdCleanupTargetPost` operationId). + ClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdCleanupTargetPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdCleanupTargetPostResponse, error) + + // ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostWithBodyWithResponse Clusters:Subsystems:Migrations:Continue // // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). // // Corresponds with POST /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/{migration_id}/continue (the `ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePost` operationId). ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostResponse, error) - // ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostWithResponse Clusters:Subsystems:Migrations:Continue - // - // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/{migration_id}/continue (the `ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePost` operationId). - ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID, body ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostResponse, error) + // ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostWithResponse Clusters:Subsystems:Migrations:Continue + // + // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/{migration_id}/continue (the `ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePost` operationId). + ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID, body ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostResponse, error) + + // ClustersTasksListApiV2ClustersClusterIdTasksGetWithResponse Clusters:Tasks:List + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/tasks/ (the `ClustersTasksListApiV2ClustersClusterIdTasksGet` operationId). + ClustersTasksListApiV2ClustersClusterIdTasksGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersTasksListApiV2ClustersClusterIdTasksGetParams, reqEditors ...RequestEditorFn) (*ClustersTasksListApiV2ClustersClusterIdTasksGetResponse, error) + + // ClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGetWithResponse Clusters:Tasks:Detail + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/tasks/{task_id}/ (the `ClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGet` operationId). + ClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, taskId openapi_types.UUID, params *ClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGetParams, reqEditors ...RequestEditorFn) (*ClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGetResponse, error) + + // ClustersUpgradeApiV2ClustersClusterIdUpdatePostWithBodyWithResponse Clusters:Upgrade + // + // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/update (the `ClustersUpgradeApiV2ClustersClusterIdUpdatePost` operationId). + ClustersUpgradeApiV2ClustersClusterIdUpdatePostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersUpgradeApiV2ClustersClusterIdUpdatePostResponse, error) + + // ClustersUpgradeApiV2ClustersClusterIdUpdatePostWithResponse Clusters:Upgrade + // + // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/update (the `ClustersUpgradeApiV2ClustersClusterIdUpdatePost` operationId). + ClustersUpgradeApiV2ClustersClusterIdUpdatePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, body ClustersUpgradeApiV2ClustersClusterIdUpdatePostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersUpgradeApiV2ClustersClusterIdUpdatePostResponse, error) + + // ManagementNodesListApiV2ManagementNodesGetWithResponse Management Nodes:List + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/management-nodes/ (the `ManagementNodesListApiV2ManagementNodesGet` operationId). + ManagementNodesListApiV2ManagementNodesGetWithResponse(ctx context.Context, params *ManagementNodesListApiV2ManagementNodesGetParams, reqEditors ...RequestEditorFn) (*ManagementNodesListApiV2ManagementNodesGetResponse, error) + + // ManagementNodeDetailApiV2ManagementNodesManagementNodeIdGetWithResponse Management Node:Detail + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/management-nodes/{management_node_id}/ (the `ManagementNodeDetailApiV2ManagementNodesManagementNodeIdGet` operationId). + ManagementNodeDetailApiV2ManagementNodesManagementNodeIdGetWithResponse(ctx context.Context, managementNodeId openapi_types.UUID, params *ManagementNodeDetailApiV2ManagementNodesManagementNodeIdGetParams, reqEditors ...RequestEditorFn) (*ManagementNodeDetailApiV2ManagementNodesManagementNodeIdGetResponse, error) +} + +type HealthApiV2MetaHealthGetResponse struct { + Body []byte + HTTPResponse *http.Response +} + +// GetBody returns the raw response body bytes +func (r HealthApiV2MetaHealthGetResponse) GetBody() []byte { + return r.Body +} + +// Status returns HTTPResponse.Status +func (r HealthApiV2MetaHealthGetResponse) Status() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Status + } + return http.StatusText(0) +} + +// StatusCode returns HTTPResponse.StatusCode +func (r HealthApiV2MetaHealthGetResponse) StatusCode() int { + if r.HTTPResponse != nil { + return r.HTTPResponse.StatusCode + } + return 0 +} + +// ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers +func (r HealthApiV2MetaHealthGetResponse) ContentType() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Header.Get("Content-Type") + } + return "" +} + +type ReadyApiV2MetaReadyGetResponse struct { + Body []byte + HTTPResponse *http.Response +} + +// GetBody returns the raw response body bytes +func (r ReadyApiV2MetaReadyGetResponse) GetBody() []byte { + return r.Body +} + +// Status returns HTTPResponse.Status +func (r ReadyApiV2MetaReadyGetResponse) Status() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Status + } + return http.StatusText(0) +} + +// StatusCode returns HTTPResponse.StatusCode +func (r ReadyApiV2MetaReadyGetResponse) StatusCode() int { + if r.HTTPResponse != nil { + return r.HTTPResponse.StatusCode + } + return 0 +} + +// ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers +func (r ReadyApiV2MetaReadyGetResponse) ContentType() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Header.Get("Content-Type") + } + return "" +} + +type ClustersListApiV2ClustersGetResponse struct { + Body []byte + HTTPResponse *http.Response + // JSON200 the response for an HTTP 200 `application/json` response + JSON200 *[]ClusterDTO + // JSON422 the response for an HTTP 422 `application/json` response + JSON422 *HTTPValidationError +} + +// GetJSON200 returns the response for an HTTP 200 `application/json` response +func (r ClustersListApiV2ClustersGetResponse) GetJSON200() *[]ClusterDTO { + return r.JSON200 +} + +// GetJSON422 returns the response for an HTTP 422 `application/json` response +func (r ClustersListApiV2ClustersGetResponse) GetJSON422() *HTTPValidationError { + return r.JSON422 +} + +// GetBody returns the raw response body bytes +func (r ClustersListApiV2ClustersGetResponse) GetBody() []byte { + return r.Body +} + +// Status returns HTTPResponse.Status +func (r ClustersListApiV2ClustersGetResponse) Status() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Status + } + return http.StatusText(0) +} + +// StatusCode returns HTTPResponse.StatusCode +func (r ClustersListApiV2ClustersGetResponse) StatusCode() int { + if r.HTTPResponse != nil { + return r.HTTPResponse.StatusCode + } + return 0 +} + +// ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers +func (r ClustersListApiV2ClustersGetResponse) ContentType() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Header.Get("Content-Type") + } + return "" +} + +type ClustersCreateApiV2ClustersPostResponse struct { + Body []byte + HTTPResponse *http.Response + // JSON422 the response for an HTTP 422 `application/json` response + JSON422 *HTTPValidationError +} + +// GetJSON422 returns the response for an HTTP 422 `application/json` response +func (r ClustersCreateApiV2ClustersPostResponse) GetJSON422() *HTTPValidationError { + return r.JSON422 +} + +// GetBody returns the raw response body bytes +func (r ClustersCreateApiV2ClustersPostResponse) GetBody() []byte { + return r.Body +} + +// Status returns HTTPResponse.Status +func (r ClustersCreateApiV2ClustersPostResponse) Status() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Status + } + return http.StatusText(0) +} + +// StatusCode returns HTTPResponse.StatusCode +func (r ClustersCreateApiV2ClustersPostResponse) StatusCode() int { + if r.HTTPResponse != nil { + return r.HTTPResponse.StatusCode + } + return 0 +} + +// ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers +func (r ClustersCreateApiV2ClustersPostResponse) ContentType() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Header.Get("Content-Type") + } + return "" +} + +type ClustersDeleteApiV2ClustersClusterIdDeleteResponse struct { + Body []byte + HTTPResponse *http.Response + // JSON422 the response for an HTTP 422 `application/json` response + JSON422 *HTTPValidationError +} + +// GetJSON422 returns the response for an HTTP 422 `application/json` response +func (r ClustersDeleteApiV2ClustersClusterIdDeleteResponse) GetJSON422() *HTTPValidationError { + return r.JSON422 +} + +// GetBody returns the raw response body bytes +func (r ClustersDeleteApiV2ClustersClusterIdDeleteResponse) GetBody() []byte { + return r.Body +} + +// Status returns HTTPResponse.Status +func (r ClustersDeleteApiV2ClustersClusterIdDeleteResponse) Status() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Status + } + return http.StatusText(0) +} + +// StatusCode returns HTTPResponse.StatusCode +func (r ClustersDeleteApiV2ClustersClusterIdDeleteResponse) StatusCode() int { + if r.HTTPResponse != nil { + return r.HTTPResponse.StatusCode + } + return 0 +} + +// ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers +func (r ClustersDeleteApiV2ClustersClusterIdDeleteResponse) ContentType() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Header.Get("Content-Type") + } + return "" +} + +type ClustersDetailApiV2ClustersClusterIdGetResponse struct { + Body []byte + HTTPResponse *http.Response + // JSON200 the response for an HTTP 200 `application/json` response + JSON200 *ClusterDTO + // JSON422 the response for an HTTP 422 `application/json` response + JSON422 *HTTPValidationError +} + +// GetJSON200 returns the response for an HTTP 200 `application/json` response +func (r ClustersDetailApiV2ClustersClusterIdGetResponse) GetJSON200() *ClusterDTO { + return r.JSON200 +} + +// GetJSON422 returns the response for an HTTP 422 `application/json` response +func (r ClustersDetailApiV2ClustersClusterIdGetResponse) GetJSON422() *HTTPValidationError { + return r.JSON422 +} + +// GetBody returns the raw response body bytes +func (r ClustersDetailApiV2ClustersClusterIdGetResponse) GetBody() []byte { + return r.Body +} + +// Status returns HTTPResponse.Status +func (r ClustersDetailApiV2ClustersClusterIdGetResponse) Status() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Status + } + return http.StatusText(0) +} + +// StatusCode returns HTTPResponse.StatusCode +func (r ClustersDetailApiV2ClustersClusterIdGetResponse) StatusCode() int { + if r.HTTPResponse != nil { + return r.HTTPResponse.StatusCode + } + return 0 +} + +// ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers +func (r ClustersDetailApiV2ClustersClusterIdGetResponse) ContentType() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Header.Get("Content-Type") + } + return "" +} + +type ClustersUpdateApiV2ClustersClusterIdPutResponse struct { + Body []byte + HTTPResponse *http.Response + // JSON200 the response for an HTTP 200 `application/json` response + JSON200 *interface{} + // JSON422 the response for an HTTP 422 `application/json` response + JSON422 *HTTPValidationError +} + +// GetJSON200 returns the response for an HTTP 200 `application/json` response +func (r ClustersUpdateApiV2ClustersClusterIdPutResponse) GetJSON200() *interface{} { + return r.JSON200 +} + +// GetJSON422 returns the response for an HTTP 422 `application/json` response +func (r ClustersUpdateApiV2ClustersClusterIdPutResponse) GetJSON422() *HTTPValidationError { + return r.JSON422 +} + +// GetBody returns the raw response body bytes +func (r ClustersUpdateApiV2ClustersClusterIdPutResponse) GetBody() []byte { + return r.Body +} + +// Status returns HTTPResponse.Status +func (r ClustersUpdateApiV2ClustersClusterIdPutResponse) Status() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Status + } + return http.StatusText(0) +} + +// StatusCode returns HTTPResponse.StatusCode +func (r ClustersUpdateApiV2ClustersClusterIdPutResponse) StatusCode() int { + if r.HTTPResponse != nil { + return r.HTTPResponse.StatusCode + } + return 0 +} + +// ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers +func (r ClustersUpdateApiV2ClustersClusterIdPutResponse) ContentType() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Header.Get("Content-Type") + } + return "" +} + +type ClustersActivateApiV2ClustersClusterIdActivatePostResponse struct { + Body []byte + HTTPResponse *http.Response + // JSON422 the response for an HTTP 422 `application/json` response + JSON422 *HTTPValidationError +} + +// GetJSON422 returns the response for an HTTP 422 `application/json` response +func (r ClustersActivateApiV2ClustersClusterIdActivatePostResponse) GetJSON422() *HTTPValidationError { + return r.JSON422 +} + +// GetBody returns the raw response body bytes +func (r ClustersActivateApiV2ClustersClusterIdActivatePostResponse) GetBody() []byte { + return r.Body +} + +// Status returns HTTPResponse.Status +func (r ClustersActivateApiV2ClustersClusterIdActivatePostResponse) Status() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Status + } + return http.StatusText(0) +} + +// StatusCode returns HTTPResponse.StatusCode +func (r ClustersActivateApiV2ClustersClusterIdActivatePostResponse) StatusCode() int { + if r.HTTPResponse != nil { + return r.HTTPResponse.StatusCode + } + return 0 +} + +// ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers +func (r ClustersActivateApiV2ClustersClusterIdActivatePostResponse) ContentType() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Header.Get("Content-Type") + } + return "" +} + +type ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostResponse struct { + Body []byte + HTTPResponse *http.Response + // JSON422 the response for an HTTP 422 `application/json` response + JSON422 *HTTPValidationError +} + +// GetJSON422 returns the response for an HTTP 422 `application/json` response +func (r ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostResponse) GetJSON422() *HTTPValidationError { + return r.JSON422 +} + +// GetBody returns the raw response body bytes +func (r ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostResponse) GetBody() []byte { + return r.Body +} + +// Status returns HTTPResponse.Status +func (r ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostResponse) Status() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Status + } + return http.StatusText(0) +} + +// StatusCode returns HTTPResponse.StatusCode +func (r ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostResponse) StatusCode() int { + if r.HTTPResponse != nil { + return r.HTTPResponse.StatusCode + } + return 0 +} + +// ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers +func (r ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostResponse) ContentType() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Header.Get("Content-Type") + } + return "" +} + +type ClustersAlertsListApiV2ClustersClusterIdAlertsGetResponse struct { + Body []byte + HTTPResponse *http.Response + // JSON200 the response for an HTTP 200 `application/json` response + JSON200 *[]AlertDTO + // JSON422 the response for an HTTP 422 `application/json` response + JSON422 *HTTPValidationError +} - // ClustersTasksListApiV2ClustersClusterIdTasksGetWithResponse Clusters:Tasks:List - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/{cluster_id}/tasks/ (the `ClustersTasksListApiV2ClustersClusterIdTasksGet` operationId). - ClustersTasksListApiV2ClustersClusterIdTasksGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersTasksListApiV2ClustersClusterIdTasksGetParams, reqEditors ...RequestEditorFn) (*ClustersTasksListApiV2ClustersClusterIdTasksGetResponse, error) +// GetJSON200 returns the response for an HTTP 200 `application/json` response +func (r ClustersAlertsListApiV2ClustersClusterIdAlertsGetResponse) GetJSON200() *[]AlertDTO { + return r.JSON200 +} - // ClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGetWithResponse Clusters:Tasks:Detail - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/clusters/{cluster_id}/tasks/{task_id}/ (the `ClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGet` operationId). - ClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, taskId openapi_types.UUID, params *ClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGetParams, reqEditors ...RequestEditorFn) (*ClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGetResponse, error) +// GetJSON422 returns the response for an HTTP 422 `application/json` response +func (r ClustersAlertsListApiV2ClustersClusterIdAlertsGetResponse) GetJSON422() *HTTPValidationError { + return r.JSON422 +} - // ClustersUpgradeApiV2ClustersClusterIdUpdatePostWithBodyWithResponse Clusters:Upgrade - // - // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/update (the `ClustersUpgradeApiV2ClustersClusterIdUpdatePost` operationId). - ClustersUpgradeApiV2ClustersClusterIdUpdatePostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersUpgradeApiV2ClustersClusterIdUpdatePostResponse, error) +// GetBody returns the raw response body bytes +func (r ClustersAlertsListApiV2ClustersClusterIdAlertsGetResponse) GetBody() []byte { + return r.Body +} - // ClustersUpgradeApiV2ClustersClusterIdUpdatePostWithResponse Clusters:Upgrade - // - // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). - // - // Corresponds with POST /api/v2/clusters/{cluster_id}/update (the `ClustersUpgradeApiV2ClustersClusterIdUpdatePost` operationId). - ClustersUpgradeApiV2ClustersClusterIdUpdatePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, body ClustersUpgradeApiV2ClustersClusterIdUpdatePostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersUpgradeApiV2ClustersClusterIdUpdatePostResponse, error) +// Status returns HTTPResponse.Status +func (r ClustersAlertsListApiV2ClustersClusterIdAlertsGetResponse) Status() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Status + } + return http.StatusText(0) +} - // ManagementNodesListApiV2ManagementNodesGetWithResponse Management Nodes:List - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/management-nodes/ (the `ManagementNodesListApiV2ManagementNodesGet` operationId). - ManagementNodesListApiV2ManagementNodesGetWithResponse(ctx context.Context, params *ManagementNodesListApiV2ManagementNodesGetParams, reqEditors ...RequestEditorFn) (*ManagementNodesListApiV2ManagementNodesGetResponse, error) +// StatusCode returns HTTPResponse.StatusCode +func (r ClustersAlertsListApiV2ClustersClusterIdAlertsGetResponse) StatusCode() int { + if r.HTTPResponse != nil { + return r.HTTPResponse.StatusCode + } + return 0 +} - // ManagementNodeDetailApiV2ManagementNodesManagementNodeIdGetWithResponse Management Node:Detail - // - // Returns a wrapper object for the known response body format(s). - // - // Corresponds with GET /api/v2/management-nodes/{management_node_id}/ (the `ManagementNodeDetailApiV2ManagementNodesManagementNodeIdGet` operationId). - ManagementNodeDetailApiV2ManagementNodesManagementNodeIdGetWithResponse(ctx context.Context, managementNodeId openapi_types.UUID, params *ManagementNodeDetailApiV2ManagementNodesManagementNodeIdGetParams, reqEditors ...RequestEditorFn) (*ManagementNodeDetailApiV2ManagementNodesManagementNodeIdGetResponse, error) +// ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers +func (r ClustersAlertsListApiV2ClustersClusterIdAlertsGetResponse) ContentType() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Header.Get("Content-Type") + } + return "" } -type HealthApiV2MetaHealthGetResponse struct { +type ClustersBackupsListApiV2ClustersClusterIdBackupsGetResponse struct { Body []byte HTTPResponse *http.Response + // JSON200 the response for an HTTP 200 `application/json` response + JSON200 *[]BackupDTO + // JSON422 the response for an HTTP 422 `application/json` response + JSON422 *HTTPValidationError +} + +// GetJSON200 returns the response for an HTTP 200 `application/json` response +func (r ClustersBackupsListApiV2ClustersClusterIdBackupsGetResponse) GetJSON200() *[]BackupDTO { + return r.JSON200 +} + +// GetJSON422 returns the response for an HTTP 422 `application/json` response +func (r ClustersBackupsListApiV2ClustersClusterIdBackupsGetResponse) GetJSON422() *HTTPValidationError { + return r.JSON422 } // GetBody returns the raw response body bytes -func (r HealthApiV2MetaHealthGetResponse) GetBody() []byte { +func (r ClustersBackupsListApiV2ClustersClusterIdBackupsGetResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r HealthApiV2MetaHealthGetResponse) Status() string { +func (r ClustersBackupsListApiV2ClustersClusterIdBackupsGetResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -12700,7 +14655,7 @@ func (r HealthApiV2MetaHealthGetResponse) Status() string { } // StatusCode returns HTTPResponse.StatusCode -func (r HealthApiV2MetaHealthGetResponse) StatusCode() int { +func (r ClustersBackupsListApiV2ClustersClusterIdBackupsGetResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -12708,25 +14663,32 @@ func (r HealthApiV2MetaHealthGetResponse) StatusCode() int { } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r HealthApiV2MetaHealthGetResponse) ContentType() string { +func (r ClustersBackupsListApiV2ClustersClusterIdBackupsGetResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } return "" } -type ReadyApiV2MetaReadyGetResponse struct { +type ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostResponse struct { Body []byte HTTPResponse *http.Response + // JSON422 the response for an HTTP 422 `application/json` response + JSON422 *HTTPValidationError +} + +// GetJSON422 returns the response for an HTTP 422 `application/json` response +func (r ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostResponse) GetJSON422() *HTTPValidationError { + return r.JSON422 } // GetBody returns the raw response body bytes -func (r ReadyApiV2MetaReadyGetResponse) GetBody() []byte { +func (r ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ReadyApiV2MetaReadyGetResponse) Status() string { +func (r ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -12734,7 +14696,7 @@ func (r ReadyApiV2MetaReadyGetResponse) Status() string { } // StatusCode returns HTTPResponse.StatusCode -func (r ReadyApiV2MetaReadyGetResponse) StatusCode() int { +func (r ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -12742,39 +14704,39 @@ func (r ReadyApiV2MetaReadyGetResponse) StatusCode() int { } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ReadyApiV2MetaReadyGetResponse) ContentType() string { +func (r ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } return "" } -type ClustersListApiV2ClustersGetResponse struct { +type ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetResponse struct { Body []byte HTTPResponse *http.Response // JSON200 the response for an HTTP 200 `application/json` response - JSON200 *[]ClusterDTO + JSON200 *[]BackupPolicyDTO // JSON422 the response for an HTTP 422 `application/json` response JSON422 *HTTPValidationError } // GetJSON200 returns the response for an HTTP 200 `application/json` response -func (r ClustersListApiV2ClustersGetResponse) GetJSON200() *[]ClusterDTO { +func (r ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetResponse) GetJSON200() *[]BackupPolicyDTO { return r.JSON200 } // GetJSON422 returns the response for an HTTP 422 `application/json` response -func (r ClustersListApiV2ClustersGetResponse) GetJSON422() *HTTPValidationError { +func (r ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetResponse) GetJSON422() *HTTPValidationError { return r.JSON422 } // GetBody returns the raw response body bytes -func (r ClustersListApiV2ClustersGetResponse) GetBody() []byte { +func (r ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ClustersListApiV2ClustersGetResponse) Status() string { +func (r ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -12782,7 +14744,7 @@ func (r ClustersListApiV2ClustersGetResponse) Status() string { } // StatusCode returns HTTPResponse.StatusCode -func (r ClustersListApiV2ClustersGetResponse) StatusCode() int { +func (r ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -12790,14 +14752,14 @@ func (r ClustersListApiV2ClustersGetResponse) StatusCode() int { } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ClustersListApiV2ClustersGetResponse) ContentType() string { +func (r ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } return "" } -type ClustersCreateApiV2ClustersPostResponse struct { +type ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostResponse struct { Body []byte HTTPResponse *http.Response // JSON422 the response for an HTTP 422 `application/json` response @@ -12805,17 +14767,17 @@ type ClustersCreateApiV2ClustersPostResponse struct { } // GetJSON422 returns the response for an HTTP 422 `application/json` response -func (r ClustersCreateApiV2ClustersPostResponse) GetJSON422() *HTTPValidationError { +func (r ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostResponse) GetJSON422() *HTTPValidationError { return r.JSON422 } // GetBody returns the raw response body bytes -func (r ClustersCreateApiV2ClustersPostResponse) GetBody() []byte { +func (r ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ClustersCreateApiV2ClustersPostResponse) Status() string { +func (r ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -12823,7 +14785,7 @@ func (r ClustersCreateApiV2ClustersPostResponse) Status() string { } // StatusCode returns HTTPResponse.StatusCode -func (r ClustersCreateApiV2ClustersPostResponse) StatusCode() int { +func (r ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -12831,14 +14793,14 @@ func (r ClustersCreateApiV2ClustersPostResponse) StatusCode() int { } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ClustersCreateApiV2ClustersPostResponse) ContentType() string { +func (r ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } return "" } -type ClustersDeleteApiV2ClustersClusterIdDeleteResponse struct { +type ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDeleteResponse struct { Body []byte HTTPResponse *http.Response // JSON422 the response for an HTTP 422 `application/json` response @@ -12846,17 +14808,17 @@ type ClustersDeleteApiV2ClustersClusterIdDeleteResponse struct { } // GetJSON422 returns the response for an HTTP 422 `application/json` response -func (r ClustersDeleteApiV2ClustersClusterIdDeleteResponse) GetJSON422() *HTTPValidationError { +func (r ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDeleteResponse) GetJSON422() *HTTPValidationError { return r.JSON422 } // GetBody returns the raw response body bytes -func (r ClustersDeleteApiV2ClustersClusterIdDeleteResponse) GetBody() []byte { +func (r ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDeleteResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ClustersDeleteApiV2ClustersClusterIdDeleteResponse) Status() string { +func (r ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDeleteResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -12864,7 +14826,7 @@ func (r ClustersDeleteApiV2ClustersClusterIdDeleteResponse) Status() string { } // StatusCode returns HTTPResponse.StatusCode -func (r ClustersDeleteApiV2ClustersClusterIdDeleteResponse) StatusCode() int { +func (r ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDeleteResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -12872,39 +14834,39 @@ func (r ClustersDeleteApiV2ClustersClusterIdDeleteResponse) StatusCode() int { } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ClustersDeleteApiV2ClustersClusterIdDeleteResponse) ContentType() string { +func (r ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDeleteResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } return "" } -type ClustersDetailApiV2ClustersClusterIdGetResponse struct { +type ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostResponse struct { Body []byte HTTPResponse *http.Response - // JSON200 the response for an HTTP 200 `application/json` response - JSON200 *ClusterDTO + // JSON201 the response for an HTTP 201 `application/json` response + JSON201 *interface{} // JSON422 the response for an HTTP 422 `application/json` response JSON422 *HTTPValidationError } -// GetJSON200 returns the response for an HTTP 200 `application/json` response -func (r ClustersDetailApiV2ClustersClusterIdGetResponse) GetJSON200() *ClusterDTO { - return r.JSON200 +// GetJSON201 returns the response for an HTTP 201 `application/json` response +func (r ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostResponse) GetJSON201() *interface{} { + return r.JSON201 } // GetJSON422 returns the response for an HTTP 422 `application/json` response -func (r ClustersDetailApiV2ClustersClusterIdGetResponse) GetJSON422() *HTTPValidationError { +func (r ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostResponse) GetJSON422() *HTTPValidationError { return r.JSON422 } // GetBody returns the raw response body bytes -func (r ClustersDetailApiV2ClustersClusterIdGetResponse) GetBody() []byte { +func (r ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ClustersDetailApiV2ClustersClusterIdGetResponse) Status() string { +func (r ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -12912,7 +14874,7 @@ func (r ClustersDetailApiV2ClustersClusterIdGetResponse) Status() string { } // StatusCode returns HTTPResponse.StatusCode -func (r ClustersDetailApiV2ClustersClusterIdGetResponse) StatusCode() int { +func (r ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -12920,39 +14882,32 @@ func (r ClustersDetailApiV2ClustersClusterIdGetResponse) StatusCode() int { } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ClustersDetailApiV2ClustersClusterIdGetResponse) ContentType() string { +func (r ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } return "" } -type ClustersUpdateApiV2ClustersClusterIdPutResponse struct { +type ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostResponse struct { Body []byte HTTPResponse *http.Response - // JSON200 the response for an HTTP 200 `application/json` response - JSON200 *interface{} // JSON422 the response for an HTTP 422 `application/json` response JSON422 *HTTPValidationError } -// GetJSON200 returns the response for an HTTP 200 `application/json` response -func (r ClustersUpdateApiV2ClustersClusterIdPutResponse) GetJSON200() *interface{} { - return r.JSON200 -} - // GetJSON422 returns the response for an HTTP 422 `application/json` response -func (r ClustersUpdateApiV2ClustersClusterIdPutResponse) GetJSON422() *HTTPValidationError { +func (r ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostResponse) GetJSON422() *HTTPValidationError { return r.JSON422 } // GetBody returns the raw response body bytes -func (r ClustersUpdateApiV2ClustersClusterIdPutResponse) GetBody() []byte { +func (r ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ClustersUpdateApiV2ClustersClusterIdPutResponse) Status() string { +func (r ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -12960,7 +14915,7 @@ func (r ClustersUpdateApiV2ClustersClusterIdPutResponse) Status() string { } // StatusCode returns HTTPResponse.StatusCode -func (r ClustersUpdateApiV2ClustersClusterIdPutResponse) StatusCode() int { +func (r ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -12968,32 +14923,39 @@ func (r ClustersUpdateApiV2ClustersClusterIdPutResponse) StatusCode() int { } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ClustersUpdateApiV2ClustersClusterIdPutResponse) ContentType() string { +func (r ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } return "" } -type ClustersActivateApiV2ClustersClusterIdActivatePostResponse struct { +type ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetResponse struct { Body []byte HTTPResponse *http.Response + // JSON200 the response for an HTTP 200 `application/json` response + JSON200 *interface{} // JSON422 the response for an HTTP 422 `application/json` response JSON422 *HTTPValidationError } +// GetJSON200 returns the response for an HTTP 200 `application/json` response +func (r ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetResponse) GetJSON200() *interface{} { + return r.JSON200 +} + // GetJSON422 returns the response for an HTTP 422 `application/json` response -func (r ClustersActivateApiV2ClustersClusterIdActivatePostResponse) GetJSON422() *HTTPValidationError { +func (r ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetResponse) GetJSON422() *HTTPValidationError { return r.JSON422 } // GetBody returns the raw response body bytes -func (r ClustersActivateApiV2ClustersClusterIdActivatePostResponse) GetBody() []byte { +func (r ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ClustersActivateApiV2ClustersClusterIdActivatePostResponse) Status() string { +func (r ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -13001,7 +14963,7 @@ func (r ClustersActivateApiV2ClustersClusterIdActivatePostResponse) Status() str } // StatusCode returns HTTPResponse.StatusCode -func (r ClustersActivateApiV2ClustersClusterIdActivatePostResponse) StatusCode() int { +func (r ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -13009,32 +14971,39 @@ func (r ClustersActivateApiV2ClustersClusterIdActivatePostResponse) StatusCode() } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ClustersActivateApiV2ClustersClusterIdActivatePostResponse) ContentType() string { +func (r ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } return "" } -type ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostResponse struct { +type ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse struct { Body []byte HTTPResponse *http.Response + // JSON200 the response for an HTTP 200 `application/json` response + JSON200 *interface{} // JSON422 the response for an HTTP 422 `application/json` response JSON422 *HTTPValidationError } +// GetJSON200 returns the response for an HTTP 200 `application/json` response +func (r ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse) GetJSON200() *interface{} { + return r.JSON200 +} + // GetJSON422 returns the response for an HTTP 422 `application/json` response -func (r ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostResponse) GetJSON422() *HTTPValidationError { +func (r ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse) GetJSON422() *HTTPValidationError { return r.JSON422 } // GetBody returns the raw response body bytes -func (r ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostResponse) GetBody() []byte { +func (r ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostResponse) Status() string { +func (r ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -13042,7 +15011,7 @@ func (r ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostResponse) } // StatusCode returns HTTPResponse.StatusCode -func (r ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostResponse) StatusCode() int { +func (r ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -13050,39 +15019,39 @@ func (r ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostResponse) } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostResponse) ContentType() string { +func (r ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } return "" } -type ClustersBackupsListApiV2ClustersClusterIdBackupsGetResponse struct { +type ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse struct { Body []byte HTTPResponse *http.Response - // JSON200 the response for an HTTP 200 `application/json` response - JSON200 *[]BackupDTO + // JSON202 the response for an HTTP 202 `application/json` response + JSON202 *interface{} // JSON422 the response for an HTTP 422 `application/json` response JSON422 *HTTPValidationError } -// GetJSON200 returns the response for an HTTP 200 `application/json` response -func (r ClustersBackupsListApiV2ClustersClusterIdBackupsGetResponse) GetJSON200() *[]BackupDTO { - return r.JSON200 +// GetJSON202 returns the response for an HTTP 202 `application/json` response +func (r ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse) GetJSON202() *interface{} { + return r.JSON202 } // GetJSON422 returns the response for an HTTP 422 `application/json` response -func (r ClustersBackupsListApiV2ClustersClusterIdBackupsGetResponse) GetJSON422() *HTTPValidationError { +func (r ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse) GetJSON422() *HTTPValidationError { return r.JSON422 } // GetBody returns the raw response body bytes -func (r ClustersBackupsListApiV2ClustersClusterIdBackupsGetResponse) GetBody() []byte { +func (r ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ClustersBackupsListApiV2ClustersClusterIdBackupsGetResponse) Status() string { +func (r ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -13090,7 +15059,7 @@ func (r ClustersBackupsListApiV2ClustersClusterIdBackupsGetResponse) Status() st } // StatusCode returns HTTPResponse.StatusCode -func (r ClustersBackupsListApiV2ClustersClusterIdBackupsGetResponse) StatusCode() int { +func (r ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -13098,32 +15067,39 @@ func (r ClustersBackupsListApiV2ClustersClusterIdBackupsGetResponse) StatusCode( } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ClustersBackupsListApiV2ClustersClusterIdBackupsGetResponse) ContentType() string { +func (r ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } return "" } -type ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostResponse struct { +type ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse struct { Body []byte HTTPResponse *http.Response + // JSON200 the response for an HTTP 200 `application/json` response + JSON200 *interface{} // JSON422 the response for an HTTP 422 `application/json` response JSON422 *HTTPValidationError } +// GetJSON200 returns the response for an HTTP 200 `application/json` response +func (r ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse) GetJSON200() *interface{} { + return r.JSON200 +} + // GetJSON422 returns the response for an HTTP 422 `application/json` response -func (r ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostResponse) GetJSON422() *HTTPValidationError { +func (r ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse) GetJSON422() *HTTPValidationError { return r.JSON422 } // GetBody returns the raw response body bytes -func (r ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostResponse) GetBody() []byte { +func (r ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostResponse) Status() string { +func (r ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -13131,7 +15107,7 @@ func (r ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostResponse) Status() } // StatusCode returns HTTPResponse.StatusCode -func (r ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostResponse) StatusCode() int { +func (r ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -13139,39 +15115,39 @@ func (r ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostResponse) StatusCo } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostResponse) ContentType() string { +func (r ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } return "" } -type ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetResponse struct { +type ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse struct { Body []byte HTTPResponse *http.Response // JSON200 the response for an HTTP 200 `application/json` response - JSON200 *[]BackupPolicyDTO + JSON200 *interface{} // JSON422 the response for an HTTP 422 `application/json` response JSON422 *HTTPValidationError } // GetJSON200 returns the response for an HTTP 200 `application/json` response -func (r ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetResponse) GetJSON200() *[]BackupPolicyDTO { +func (r ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse) GetJSON200() *interface{} { return r.JSON200 } // GetJSON422 returns the response for an HTTP 422 `application/json` response -func (r ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetResponse) GetJSON422() *HTTPValidationError { +func (r ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse) GetJSON422() *HTTPValidationError { return r.JSON422 } // GetBody returns the raw response body bytes -func (r ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetResponse) GetBody() []byte { +func (r ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetResponse) Status() string { +func (r ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -13179,7 +15155,7 @@ func (r ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGet } // StatusCode returns HTTPResponse.StatusCode -func (r ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetResponse) StatusCode() int { +func (r ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -13187,32 +15163,39 @@ func (r ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGet } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetResponse) ContentType() string { +func (r ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } return "" } -type ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostResponse struct { +type ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse struct { Body []byte HTTPResponse *http.Response + // JSON200 the response for an HTTP 200 `application/json` response + JSON200 *BackupDTO // JSON422 the response for an HTTP 422 `application/json` response JSON422 *HTTPValidationError } +// GetJSON200 returns the response for an HTTP 200 `application/json` response +func (r ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse) GetJSON200() *BackupDTO { + return r.JSON200 +} + // GetJSON422 returns the response for an HTTP 422 `application/json` response -func (r ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostResponse) GetJSON422() *HTTPValidationError { +func (r ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse) GetJSON422() *HTTPValidationError { return r.JSON422 } // GetBody returns the raw response body bytes -func (r ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostResponse) GetBody() []byte { +func (r ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostResponse) Status() string { +func (r ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -13220,7 +15203,7 @@ func (r ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesP } // StatusCode returns HTTPResponse.StatusCode -func (r ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostResponse) StatusCode() int { +func (r ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -13228,14 +15211,14 @@ func (r ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesP } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostResponse) ContentType() string { +func (r ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } return "" } -type ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDeleteResponse struct { +type ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse struct { Body []byte HTTPResponse *http.Response // JSON422 the response for an HTTP 422 `application/json` response @@ -13243,17 +15226,17 @@ type ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPoli } // GetJSON422 returns the response for an HTTP 422 `application/json` response -func (r ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDeleteResponse) GetJSON422() *HTTPValidationError { +func (r ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse) GetJSON422() *HTTPValidationError { return r.JSON422 } // GetBody returns the raw response body bytes -func (r ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDeleteResponse) GetBody() []byte { +func (r ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDeleteResponse) Status() string { +func (r ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -13261,7 +15244,7 @@ func (r ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesP } // StatusCode returns HTTPResponse.StatusCode -func (r ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDeleteResponse) StatusCode() int { +func (r ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -13269,39 +15252,39 @@ func (r ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesP } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDeleteResponse) ContentType() string { +func (r ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } return "" } -type ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostResponse struct { +type ClustersCapacityApiV2ClustersClusterIdCapacityGetResponse struct { Body []byte HTTPResponse *http.Response - // JSON201 the response for an HTTP 201 `application/json` response - JSON201 *interface{} + // JSON200 the response for an HTTP 200 `application/json` response + JSON200 *interface{} // JSON422 the response for an HTTP 422 `application/json` response JSON422 *HTTPValidationError } -// GetJSON201 returns the response for an HTTP 201 `application/json` response -func (r ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostResponse) GetJSON201() *interface{} { - return r.JSON201 +// GetJSON200 returns the response for an HTTP 200 `application/json` response +func (r ClustersCapacityApiV2ClustersClusterIdCapacityGetResponse) GetJSON200() *interface{} { + return r.JSON200 } // GetJSON422 returns the response for an HTTP 422 `application/json` response -func (r ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostResponse) GetJSON422() *HTTPValidationError { +func (r ClustersCapacityApiV2ClustersClusterIdCapacityGetResponse) GetJSON422() *HTTPValidationError { return r.JSON422 } // GetBody returns the raw response body bytes -func (r ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostResponse) GetBody() []byte { +func (r ClustersCapacityApiV2ClustersClusterIdCapacityGetResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostResponse) Status() string { +func (r ClustersCapacityApiV2ClustersClusterIdCapacityGetResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -13309,7 +15292,7 @@ func (r ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesP } // StatusCode returns HTTPResponse.StatusCode -func (r ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostResponse) StatusCode() int { +func (r ClustersCapacityApiV2ClustersClusterIdCapacityGetResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -13317,32 +15300,39 @@ func (r ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesP } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostResponse) ContentType() string { +func (r ClustersCapacityApiV2ClustersClusterIdCapacityGetResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } return "" } -type ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostResponse struct { +type ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetResponse struct { Body []byte HTTPResponse *http.Response + // JSON200 the response for an HTTP 200 `application/json` response + JSON200 *[]ConsistencyGroupDTO // JSON422 the response for an HTTP 422 `application/json` response JSON422 *HTTPValidationError } +// GetJSON200 returns the response for an HTTP 200 `application/json` response +func (r ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetResponse) GetJSON200() *[]ConsistencyGroupDTO { + return r.JSON200 +} + // GetJSON422 returns the response for an HTTP 422 `application/json` response -func (r ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostResponse) GetJSON422() *HTTPValidationError { +func (r ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetResponse) GetJSON422() *HTTPValidationError { return r.JSON422 } // GetBody returns the raw response body bytes -func (r ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostResponse) GetBody() []byte { +func (r ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostResponse) Status() string { +func (r ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -13350,7 +15340,7 @@ func (r ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesP } // StatusCode returns HTTPResponse.StatusCode -func (r ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostResponse) StatusCode() int { +func (r ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -13358,39 +15348,39 @@ func (r ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesP } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostResponse) ContentType() string { +func (r ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } return "" } -type ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetResponse struct { +type ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGetResponse struct { Body []byte HTTPResponse *http.Response // JSON200 the response for an HTTP 200 `application/json` response - JSON200 *interface{} + JSON200 *ConsistencyGroupDTO // JSON422 the response for an HTTP 422 `application/json` response JSON422 *HTTPValidationError } // GetJSON200 returns the response for an HTTP 200 `application/json` response -func (r ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetResponse) GetJSON200() *interface{} { +func (r ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGetResponse) GetJSON200() *ConsistencyGroupDTO { return r.JSON200 } // GetJSON422 returns the response for an HTTP 422 `application/json` response -func (r ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetResponse) GetJSON422() *HTTPValidationError { +func (r ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGetResponse) GetJSON422() *HTTPValidationError { return r.JSON422 } // GetBody returns the raw response body bytes -func (r ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetResponse) GetBody() []byte { +func (r ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGetResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetResponse) Status() string { +func (r ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGetResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -13398,7 +15388,7 @@ func (r ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetResponse) Sta } // StatusCode returns HTTPResponse.StatusCode -func (r ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetResponse) StatusCode() int { +func (r ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGetResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -13406,39 +15396,39 @@ func (r ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetResponse) Sta } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetResponse) ContentType() string { +func (r ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGetResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } return "" } -type ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse struct { +type ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGetResponse struct { Body []byte HTTPResponse *http.Response // JSON200 the response for an HTTP 200 `application/json` response - JSON200 *interface{} + JSON200 *[]ConsistencyGroupMemberDTO // JSON422 the response for an HTTP 422 `application/json` response JSON422 *HTTPValidationError } // GetJSON200 returns the response for an HTTP 200 `application/json` response -func (r ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse) GetJSON200() *interface{} { +func (r ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGetResponse) GetJSON200() *[]ConsistencyGroupMemberDTO { return r.JSON200 } // GetJSON422 returns the response for an HTTP 422 `application/json` response -func (r ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse) GetJSON422() *HTTPValidationError { +func (r ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGetResponse) GetJSON422() *HTTPValidationError { return r.JSON422 } // GetBody returns the raw response body bytes -func (r ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse) GetBody() []byte { +func (r ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGetResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse) Status() string { +func (r ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGetResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -13446,7 +15436,7 @@ func (r ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse) St } // StatusCode returns HTTPResponse.StatusCode -func (r ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse) StatusCode() int { +func (r ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGetResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -13454,39 +15444,39 @@ func (r ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse) St } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse) ContentType() string { +func (r ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGetResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } return "" } -type ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse struct { +type ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostResponse struct { Body []byte HTTPResponse *http.Response - // JSON202 the response for an HTTP 202 `application/json` response - JSON202 *interface{} + // JSON200 the response for an HTTP 200 `application/json` response + JSON200 *ConsistencyGroupMemberDTO // JSON422 the response for an HTTP 422 `application/json` response JSON422 *HTTPValidationError } -// GetJSON202 returns the response for an HTTP 202 `application/json` response -func (r ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse) GetJSON202() *interface{} { - return r.JSON202 +// GetJSON200 returns the response for an HTTP 200 `application/json` response +func (r ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostResponse) GetJSON200() *ConsistencyGroupMemberDTO { + return r.JSON200 } // GetJSON422 returns the response for an HTTP 422 `application/json` response -func (r ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse) GetJSON422() *HTTPValidationError { +func (r ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostResponse) GetJSON422() *HTTPValidationError { return r.JSON422 } // GetBody returns the raw response body bytes -func (r ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse) GetBody() []byte { +func (r ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse) Status() string { +func (r ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -13494,7 +15484,7 @@ func (r ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse) } // StatusCode returns HTTPResponse.StatusCode -func (r ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse) StatusCode() int { +func (r ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -13502,39 +15492,32 @@ func (r ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse) } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse) ContentType() string { +func (r ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } return "" } -type ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse struct { +type ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDeleteResponse struct { Body []byte HTTPResponse *http.Response - // JSON200 the response for an HTTP 200 `application/json` response - JSON200 *interface{} // JSON422 the response for an HTTP 422 `application/json` response JSON422 *HTTPValidationError } -// GetJSON200 returns the response for an HTTP 200 `application/json` response -func (r ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse) GetJSON200() *interface{} { - return r.JSON200 -} - // GetJSON422 returns the response for an HTTP 422 `application/json` response -func (r ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse) GetJSON422() *HTTPValidationError { +func (r ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDeleteResponse) GetJSON422() *HTTPValidationError { return r.JSON422 } // GetBody returns the raw response body bytes -func (r ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse) GetBody() []byte { +func (r ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDeleteResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse) Status() string { +func (r ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDeleteResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -13542,7 +15525,7 @@ func (r ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPost } // StatusCode returns HTTPResponse.StatusCode -func (r ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse) StatusCode() int { +func (r ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDeleteResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -13550,39 +15533,39 @@ func (r ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPost } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse) ContentType() string { +func (r ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDeleteResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } return "" } -type ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse struct { +type ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetResponse struct { Body []byte HTTPResponse *http.Response // JSON200 the response for an HTTP 200 `application/json` response - JSON200 *interface{} + JSON200 *[]ConsistencyGroupGenerationDTO // JSON422 the response for an HTTP 422 `application/json` response JSON422 *HTTPValidationError } // GetJSON200 returns the response for an HTTP 200 `application/json` response -func (r ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse) GetJSON200() *interface{} { +func (r ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetResponse) GetJSON200() *[]ConsistencyGroupGenerationDTO { return r.JSON200 } // GetJSON422 returns the response for an HTTP 422 `application/json` response -func (r ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse) GetJSON422() *HTTPValidationError { +func (r ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetResponse) GetJSON422() *HTTPValidationError { return r.JSON422 } // GetBody returns the raw response body bytes -func (r ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse) GetBody() []byte { +func (r ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse) Status() string { +func (r ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -13590,7 +15573,7 @@ func (r ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse) S } // StatusCode returns HTTPResponse.StatusCode -func (r ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse) StatusCode() int { +func (r ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -13598,39 +15581,39 @@ func (r ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse) S } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse) ContentType() string { +func (r ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } return "" } -type ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse struct { +type ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPostResponse struct { Body []byte HTTPResponse *http.Response // JSON200 the response for an HTTP 200 `application/json` response - JSON200 *BackupDTO + JSON200 *ConsistencyGroupGenerationDTO // JSON422 the response for an HTTP 422 `application/json` response JSON422 *HTTPValidationError } // GetJSON200 returns the response for an HTTP 200 `application/json` response -func (r ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse) GetJSON200() *BackupDTO { +func (r ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPostResponse) GetJSON200() *ConsistencyGroupGenerationDTO { return r.JSON200 } // GetJSON422 returns the response for an HTTP 422 `application/json` response -func (r ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse) GetJSON422() *HTTPValidationError { +func (r ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPostResponse) GetJSON422() *HTTPValidationError { return r.JSON422 } // GetBody returns the raw response body bytes -func (r ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse) GetBody() []byte { +func (r ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPostResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse) Status() string { +func (r ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPostResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -13638,7 +15621,7 @@ func (r ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse) S } // StatusCode returns HTTPResponse.StatusCode -func (r ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse) StatusCode() int { +func (r ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPostResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -13646,14 +15629,14 @@ func (r ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse) S } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse) ContentType() string { +func (r ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPostResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } return "" } -type ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse struct { +type ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDeleteResponse struct { Body []byte HTTPResponse *http.Response // JSON422 the response for an HTTP 422 `application/json` response @@ -13661,17 +15644,17 @@ type ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse st } // GetJSON422 returns the response for an HTTP 422 `application/json` response -func (r ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse) GetJSON422() *HTTPValidationError { +func (r ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDeleteResponse) GetJSON422() *HTTPValidationError { return r.JSON422 } // GetBody returns the raw response body bytes -func (r ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse) GetBody() []byte { +func (r ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDeleteResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse) Status() string { +func (r ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDeleteResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -13679,7 +15662,7 @@ func (r ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse } // StatusCode returns HTTPResponse.StatusCode -func (r ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse) StatusCode() int { +func (r ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDeleteResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -13687,39 +15670,39 @@ func (r ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse) ContentType() string { +func (r ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDeleteResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } return "" } -type ClustersCapacityApiV2ClustersClusterIdCapacityGetResponse struct { +type ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGetResponse struct { Body []byte HTTPResponse *http.Response // JSON200 the response for an HTTP 200 `application/json` response - JSON200 *interface{} + JSON200 *ConsistencyGroupGenerationDTO // JSON422 the response for an HTTP 422 `application/json` response JSON422 *HTTPValidationError } // GetJSON200 returns the response for an HTTP 200 `application/json` response -func (r ClustersCapacityApiV2ClustersClusterIdCapacityGetResponse) GetJSON200() *interface{} { +func (r ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGetResponse) GetJSON200() *ConsistencyGroupGenerationDTO { return r.JSON200 } // GetJSON422 returns the response for an HTTP 422 `application/json` response -func (r ClustersCapacityApiV2ClustersClusterIdCapacityGetResponse) GetJSON422() *HTTPValidationError { +func (r ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGetResponse) GetJSON422() *HTTPValidationError { return r.JSON422 } // GetBody returns the raw response body bytes -func (r ClustersCapacityApiV2ClustersClusterIdCapacityGetResponse) GetBody() []byte { +func (r ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGetResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ClustersCapacityApiV2ClustersClusterIdCapacityGetResponse) Status() string { +func (r ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGetResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -13727,7 +15710,7 @@ func (r ClustersCapacityApiV2ClustersClusterIdCapacityGetResponse) Status() stri } // StatusCode returns HTTPResponse.StatusCode -func (r ClustersCapacityApiV2ClustersClusterIdCapacityGetResponse) StatusCode() int { +func (r ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGetResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -13735,7 +15718,7 @@ func (r ClustersCapacityApiV2ClustersClusterIdCapacityGetResponse) StatusCode() } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ClustersCapacityApiV2ClustersClusterIdCapacityGetResponse) ContentType() string { +func (r ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGetResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } @@ -13961,32 +15944,121 @@ func (r ClustersReplicationPoliciesListApiV2ClustersClusterIdReplicationPolicies } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ClustersReplicationPoliciesListApiV2ClustersClusterIdReplicationPoliciesGetResponse) ContentType() string { +func (r ClustersReplicationPoliciesListApiV2ClustersClusterIdReplicationPoliciesGetResponse) ContentType() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Header.Get("Content-Type") + } + return "" +} + +type ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostResponse struct { + Body []byte + HTTPResponse *http.Response + // JSON422 the response for an HTTP 422 `application/json` response + JSON422 *HTTPValidationError +} + +// GetJSON422 returns the response for an HTTP 422 `application/json` response +func (r ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostResponse) GetJSON422() *HTTPValidationError { + return r.JSON422 +} + +// GetBody returns the raw response body bytes +func (r ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostResponse) GetBody() []byte { + return r.Body +} + +// Status returns HTTPResponse.Status +func (r ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostResponse) Status() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Status + } + return http.StatusText(0) +} + +// StatusCode returns HTTPResponse.StatusCode +func (r ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostResponse) StatusCode() int { + if r.HTTPResponse != nil { + return r.HTTPResponse.StatusCode + } + return 0 +} + +// ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers +func (r ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostResponse) ContentType() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Header.Get("Content-Type") + } + return "" +} + +type ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDeleteResponse struct { + Body []byte + HTTPResponse *http.Response + // JSON422 the response for an HTTP 422 `application/json` response + JSON422 *HTTPValidationError +} + +// GetJSON422 returns the response for an HTTP 422 `application/json` response +func (r ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDeleteResponse) GetJSON422() *HTTPValidationError { + return r.JSON422 +} + +// GetBody returns the raw response body bytes +func (r ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDeleteResponse) GetBody() []byte { + return r.Body +} + +// Status returns HTTPResponse.Status +func (r ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDeleteResponse) Status() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Status + } + return http.StatusText(0) +} + +// StatusCode returns HTTPResponse.StatusCode +func (r ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDeleteResponse) StatusCode() int { + if r.HTTPResponse != nil { + return r.HTTPResponse.StatusCode + } + return 0 +} + +// ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers +func (r ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDeleteResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } return "" } -type ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostResponse struct { +type ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGetResponse struct { Body []byte HTTPResponse *http.Response + // JSON200 the response for an HTTP 200 `application/json` response + JSON200 *ReplicationPolicyDTO // JSON422 the response for an HTTP 422 `application/json` response JSON422 *HTTPValidationError } +// GetJSON200 returns the response for an HTTP 200 `application/json` response +func (r ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGetResponse) GetJSON200() *ReplicationPolicyDTO { + return r.JSON200 +} + // GetJSON422 returns the response for an HTTP 422 `application/json` response -func (r ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostResponse) GetJSON422() *HTTPValidationError { +func (r ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGetResponse) GetJSON422() *HTTPValidationError { return r.JSON422 } // GetBody returns the raw response body bytes -func (r ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostResponse) GetBody() []byte { +func (r ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGetResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostResponse) Status() string { +func (r ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGetResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -13994,7 +16066,7 @@ func (r ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPolici } // StatusCode returns HTTPResponse.StatusCode -func (r ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostResponse) StatusCode() int { +func (r ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGetResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -14002,32 +16074,39 @@ func (r ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPolici } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostResponse) ContentType() string { +func (r ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGetResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } return "" } -type ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDeleteResponse struct { +type ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPostResponse struct { Body []byte HTTPResponse *http.Response + // JSON200 the response for an HTTP 200 `application/json` response + JSON200 *[]FailoverResultDTO // JSON422 the response for an HTTP 422 `application/json` response JSON422 *HTTPValidationError } +// GetJSON200 returns the response for an HTTP 200 `application/json` response +func (r ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPostResponse) GetJSON200() *[]FailoverResultDTO { + return r.JSON200 +} + // GetJSON422 returns the response for an HTTP 422 `application/json` response -func (r ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDeleteResponse) GetJSON422() *HTTPValidationError { +func (r ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPostResponse) GetJSON422() *HTTPValidationError { return r.JSON422 } // GetBody returns the raw response body bytes -func (r ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDeleteResponse) GetBody() []byte { +func (r ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPostResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDeleteResponse) Status() string { +func (r ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPostResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -14035,7 +16114,7 @@ func (r ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPolici } // StatusCode returns HTTPResponse.StatusCode -func (r ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDeleteResponse) StatusCode() int { +func (r ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPostResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -14043,39 +16122,39 @@ func (r ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPolici } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDeleteResponse) ContentType() string { +func (r ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPostResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } return "" } -type ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGetResponse struct { +type ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGetResponse struct { Body []byte HTTPResponse *http.Response // JSON200 the response for an HTTP 200 `application/json` response - JSON200 *ReplicationPolicyDTO + JSON200 *ReplicatedGenerationDTO // JSON422 the response for an HTTP 422 `application/json` response JSON422 *HTTPValidationError } // GetJSON200 returns the response for an HTTP 200 `application/json` response -func (r ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGetResponse) GetJSON200() *ReplicationPolicyDTO { +func (r ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGetResponse) GetJSON200() *ReplicatedGenerationDTO { return r.JSON200 } // GetJSON422 returns the response for an HTTP 422 `application/json` response -func (r ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGetResponse) GetJSON422() *HTTPValidationError { +func (r ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGetResponse) GetJSON422() *HTTPValidationError { return r.JSON422 } // GetBody returns the raw response body bytes -func (r ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGetResponse) GetBody() []byte { +func (r ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGetResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGetResponse) Status() string { +func (r ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGetResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -14083,7 +16162,7 @@ func (r ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPolici } // StatusCode returns HTTPResponse.StatusCode -func (r ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGetResponse) StatusCode() int { +func (r ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGetResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -14091,39 +16170,39 @@ func (r ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPolici } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGetResponse) ContentType() string { +func (r ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGetResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } return "" } -type ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPostResponse struct { +type ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetResponse struct { Body []byte HTTPResponse *http.Response // JSON200 the response for an HTTP 200 `application/json` response - JSON200 *[]FailoverResultDTO + JSON200 *ReplicationRelationshipDTO // JSON422 the response for an HTTP 422 `application/json` response JSON422 *HTTPValidationError } // GetJSON200 returns the response for an HTTP 200 `application/json` response -func (r ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPostResponse) GetJSON200() *[]FailoverResultDTO { +func (r ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetResponse) GetJSON200() *ReplicationRelationshipDTO { return r.JSON200 } // GetJSON422 returns the response for an HTTP 422 `application/json` response -func (r ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPostResponse) GetJSON422() *HTTPValidationError { +func (r ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetResponse) GetJSON422() *HTTPValidationError { return r.JSON422 } // GetBody returns the raw response body bytes -func (r ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPostResponse) GetBody() []byte { +func (r ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPostResponse) Status() string { +func (r ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -14131,7 +16210,7 @@ func (r ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoli } // StatusCode returns HTTPResponse.StatusCode -func (r ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPostResponse) StatusCode() int { +func (r ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -14139,39 +16218,39 @@ func (r ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoli } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPostResponse) ContentType() string { +func (r ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } return "" } -type ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetResponse struct { +type ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGetResponse struct { Body []byte HTTPResponse *http.Response // JSON200 the response for an HTTP 200 `application/json` response - JSON200 *ReplicationRelationshipDTO + JSON200 *ReplicatedSnapshotDTO // JSON422 the response for an HTTP 422 `application/json` response JSON422 *HTTPValidationError } // GetJSON200 returns the response for an HTTP 200 `application/json` response -func (r ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetResponse) GetJSON200() *ReplicationRelationshipDTO { +func (r ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGetResponse) GetJSON200() *ReplicatedSnapshotDTO { return r.JSON200 } // GetJSON422 returns the response for an HTTP 422 `application/json` response -func (r ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetResponse) GetJSON422() *HTTPValidationError { +func (r ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGetResponse) GetJSON422() *HTTPValidationError { return r.JSON422 } // GetBody returns the raw response body bytes -func (r ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetResponse) GetBody() []byte { +func (r ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGetResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetResponse) Status() string { +func (r ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGetResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -14179,7 +16258,7 @@ func (r ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationR } // StatusCode returns HTTPResponse.StatusCode -func (r ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetResponse) StatusCode() int { +func (r ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGetResponse) StatusCode() int { if r.HTTPResponse != nil { return r.HTTPResponse.StatusCode } @@ -14187,7 +16266,7 @@ func (r ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationR } // ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers -func (r ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetResponse) ContentType() string { +func (r ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGetResponse) ContentType() string { if r.HTTPResponse != nil { return r.HTTPResponse.Header.Get("Content-Type") } @@ -16939,6 +19018,54 @@ func (r ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStorage return "" } +type ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGetResponse struct { + Body []byte + HTTPResponse *http.Response + // JSON200 the response for an HTTP 200 `application/json` response + JSON200 *ReplicationStatusDTO + // JSON422 the response for an HTTP 422 `application/json` response + JSON422 *HTTPValidationError +} + +// GetJSON200 returns the response for an HTTP 200 `application/json` response +func (r ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGetResponse) GetJSON200() *ReplicationStatusDTO { + return r.JSON200 +} + +// GetJSON422 returns the response for an HTTP 422 `application/json` response +func (r ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGetResponse) GetJSON422() *HTTPValidationError { + return r.JSON422 +} + +// GetBody returns the raw response body bytes +func (r ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGetResponse) GetBody() []byte { + return r.Body +} + +// Status returns HTTPResponse.Status +func (r ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGetResponse) Status() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Status + } + return http.StatusText(0) +} + +// StatusCode returns HTTPResponse.StatusCode +func (r ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGetResponse) StatusCode() int { + if r.HTTPResponse != nil { + return r.HTTPResponse.StatusCode + } + return 0 +} + +// ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers +func (r ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGetResponse) ContentType() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Header.Get("Content-Type") + } + return "" +} + type ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPostResponse struct { Body []byte HTTPResponse *http.Response @@ -17832,6 +19959,36 @@ func (c *ClientWithResponses) ClustersAddreplicationApiV2ClustersClusterIdAddrep return ParseClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostResponse(rsp) } +// ClustersAlertsListApiV2ClustersClusterIdAlertsGetWithResponse Clusters:Alerts:List +// +// The conditions in this cluster that currently need an operator. +// +// This is not the event log. An alert appears only while it is still true +// and disappears on its own once it is not: the node comes back ONLINE, the +// device comes back, the cluster leaves degraded. Conditions an operator +// caused on purpose -- a node they shut down, a device they removed -- are +// not alerts and are not listed. +// +// By default only what is wrong NOW is returned -- every entry has +// “status: firing“. Pass “history=true“ to also get the ones that have +// since resolved, each with its “resolved_at“, or “history_seconds=N“ +// for just the recent past. Either way both transitions are written to the +// cluster event log as ALERT_RAISED / ALERT_RESOLVED, so a resolution +// reaches an operator whether or not anyone asks for history here. +// +// Critical sorts before warning, and firing before resolved. +// +// Returns a wrapper object for the known response body format(s). +// +// Corresponds with GET /api/v2/clusters/{cluster_id}/alerts/ (the `ClustersAlertsListApiV2ClustersClusterIdAlertsGet` operationId). +func (c *ClientWithResponses) ClustersAlertsListApiV2ClustersClusterIdAlertsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersAlertsListApiV2ClustersClusterIdAlertsGetParams, reqEditors ...RequestEditorFn) (*ClustersAlertsListApiV2ClustersClusterIdAlertsGetResponse, error) { + rsp, err := c.ClustersAlertsListApiV2ClustersClusterIdAlertsGet(ctx, clusterId, params, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersAlertsListApiV2ClustersClusterIdAlertsGetResponse(rsp) +} + // ClustersBackupsListApiV2ClustersClusterIdBackupsGetWithResponse Clusters:Backups:List // // Returns a wrapper object for the known response body format(s). @@ -17998,128 +20155,281 @@ func (c *ClientWithResponses) ClustersBackupsImportApiV2ClustersClusterIdBackups if err != nil { return nil, err } - return ParseClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse(rsp) + return ParseClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse(rsp) +} + +// ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostWithResponse Clusters:Backups:Import +// +// Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/backups/import (the `ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPost` operationId). +func (c *ClientWithResponses) ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, body ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse, error) { + rsp, err := c.ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPost(ctx, clusterId, body, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse(rsp) +} + +// ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostWithBodyWithResponse Clusters:Backups:Restore +// +// Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/backups/restore (the `ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePost` operationId). +func (c *ClientWithResponses) ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse, error) { + rsp, err := c.ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostWithBody(ctx, clusterId, contentType, body, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse(rsp) +} + +// ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostWithResponse Clusters:Backups:Restore +// +// Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/backups/restore (the `ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePost` operationId). +func (c *ClientWithResponses) ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, body ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse, error) { + rsp, err := c.ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePost(ctx, clusterId, body, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse(rsp) +} + +// ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostWithBodyWithResponse Clusters:Backups:Source-Switch +// +// Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/backups/source-switch (the `ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPost` operationId). +func (c *ClientWithResponses) ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse, error) { + rsp, err := c.ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostWithBody(ctx, clusterId, contentType, body, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse(rsp) +} + +// ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostWithResponse Clusters:Backups:Source-Switch +// +// Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/backups/source-switch (the `ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPost` operationId). +func (c *ClientWithResponses) ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, body ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse, error) { + rsp, err := c.ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPost(ctx, clusterId, body, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse(rsp) +} + +// ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetWithResponse Clusters:Backups:Sources +// +// Returns a wrapper object for the known response body format(s). +// +// Corresponds with GET /api/v2/clusters/{cluster_id}/backups/sources (the `ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGet` operationId). +func (c *ClientWithResponses) ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse, error) { + rsp, err := c.ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGet(ctx, clusterId, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse(rsp) +} + +// ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetWithResponse Clusters:Backups:Detail +// +// Returns a wrapper object for the known response body format(s). +// +// Corresponds with GET /api/v2/clusters/{cluster_id}/backups/{backup_id}/ (the `ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGet` operationId). +func (c *ClientWithResponses) ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, backupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse, error) { + rsp, err := c.ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGet(ctx, clusterId, backupId, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse(rsp) +} + +// ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteWithResponse Deprecated — delete all backups for a volume +// +// Deprecated. Use `DELETE /clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/backups` instead. +// +// Returns a wrapper object for the known response body format(s). +// +// Corresponds with DELETE /api/v2/clusters/{cluster_id}/backups/{volume_id} (the `ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDelete` operationId). +// +// Deprecated: this operation has been marked as deprecated upstream, but no `x-deprecated-reason` was set +func (c *ClientWithResponses) ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse, error) { + rsp, err := c.ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDelete(ctx, clusterId, volumeId, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse(rsp) +} + +// ClustersCapacityApiV2ClustersClusterIdCapacityGetWithResponse Clusters:Capacity +// +// Returns a wrapper object for the known response body format(s). +// +// Corresponds with GET /api/v2/clusters/{cluster_id}/capacity (the `ClustersCapacityApiV2ClustersClusterIdCapacityGet` operationId). +func (c *ClientWithResponses) ClustersCapacityApiV2ClustersClusterIdCapacityGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersCapacityApiV2ClustersClusterIdCapacityGetParams, reqEditors ...RequestEditorFn) (*ClustersCapacityApiV2ClustersClusterIdCapacityGetResponse, error) { + rsp, err := c.ClustersCapacityApiV2ClustersClusterIdCapacityGet(ctx, clusterId, params, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersCapacityApiV2ClustersClusterIdCapacityGetResponse(rsp) +} + +// ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetWithResponse Clusters:Consistency-Groups:List +// +// List the cluster's consistency groups, or resolve one by name (§10). +// +// Returns an empty list when “name“ matches no group, so a caller can probe +// existence without a 404. +// +// Returns a wrapper object for the known response body format(s). +// +// Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/ (the `ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGet` operationId). +func (c *ClientWithResponses) ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetParams, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetResponse, error) { + rsp, err := c.ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGet(ctx, clusterId, params, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetResponse(rsp) } -// ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostWithResponse Clusters:Backups:Import +// ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGetWithResponse Clusters:Consistency-Groups:Detail // -// Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). +// Returns a wrapper object for the known response body format(s). // -// Corresponds with POST /api/v2/clusters/{cluster_id}/backups/import (the `ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPost` operationId). -func (c *ClientWithResponses) ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, body ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse, error) { - rsp, err := c.ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPost(ctx, clusterId, body, reqEditors...) +// Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/ (the `ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGet` operationId). +func (c *ClientWithResponses) ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGetResponse, error) { + rsp, err := c.ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGet(ctx, clusterId, groupId, reqEditors...) if err != nil { return nil, err } - return ParseClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse(rsp) + return ParseClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGetResponse(rsp) } -// ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostWithBodyWithResponse Clusters:Backups:Restore +// ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGetWithResponse Clusters:Consistency-Groups:Members // -// Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). +// Returns a wrapper object for the known response body format(s). // -// Corresponds with POST /api/v2/clusters/{cluster_id}/backups/restore (the `ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePost` operationId). -func (c *ClientWithResponses) ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse, error) { - rsp, err := c.ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostWithBody(ctx, clusterId, contentType, body, reqEditors...) +// Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/members (the `ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGet` operationId). +func (c *ClientWithResponses) ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGetResponse, error) { + rsp, err := c.ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGet(ctx, clusterId, groupId, reqEditors...) if err != nil { return nil, err } - return ParseClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse(rsp) + return ParseClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGetResponse(rsp) } -// ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostWithResponse Clusters:Backups:Restore +// ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostWithBodyWithResponse Clusters:Consistency-Groups:Members:Join // -// Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). +// Join an EXISTING volume to the group (design §4.5, Phase 4 late join). // -// Corresponds with POST /api/v2/clusters/{cluster_id}/backups/restore (the `ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePost` operationId). -func (c *ClientWithResponses) ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, body ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse, error) { - rsp, err := c.ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePost(ctx, clusterId, body, reqEditors...) +// Validates the pinned placement, the pool, the member cap, and the one-way +// rule; a refusal is a 409 naming the precondition. Idempotent: joining a +// current member returns its membership row unchanged. +// +// Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/members (the `ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPost` operationId). +func (c *ClientWithResponses) ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostResponse, error) { + rsp, err := c.ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostWithBody(ctx, clusterId, groupId, contentType, body, reqEditors...) if err != nil { return nil, err } - return ParseClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse(rsp) + return ParseClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostResponse(rsp) } -// ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostWithBodyWithResponse Clusters:Backups:Source-Switch +// ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostWithResponse Clusters:Consistency-Groups:Members:Join // -// Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). +// Join an EXISTING volume to the group (design §4.5, Phase 4 late join). // -// Corresponds with POST /api/v2/clusters/{cluster_id}/backups/source-switch (the `ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPost` operationId). -func (c *ClientWithResponses) ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse, error) { - rsp, err := c.ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostWithBody(ctx, clusterId, contentType, body, reqEditors...) +// Validates the pinned placement, the pool, the member cap, and the one-way +// rule; a refusal is a 409 naming the precondition. Idempotent: joining a +// current member returns its membership row unchanged. +// +// Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/members (the `ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPost` operationId). +func (c *ClientWithResponses) ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, body ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostResponse, error) { + rsp, err := c.ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPost(ctx, clusterId, groupId, body, reqEditors...) if err != nil { return nil, err } - return ParseClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse(rsp) + return ParseClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostResponse(rsp) } -// ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostWithResponse Clusters:Backups:Source-Switch +// ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDeleteWithResponse Clusters:Consistency-Groups:Members:Detach // -// Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). +// Detach a member: close its epoch one-way, preserving prior generations (§8.2). // -// Corresponds with POST /api/v2/clusters/{cluster_id}/backups/source-switch (the `ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPost` operationId). -func (c *ClientWithResponses) ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, body ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse, error) { - rsp, err := c.ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPost(ctx, clusterId, body, reqEditors...) +// Returns a wrapper object for the known response body format(s). +// +// Corresponds with DELETE /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/members/{lvol_id} (the `ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDelete` operationId). +func (c *ClientWithResponses) ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, lvolId string, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDeleteResponse, error) { + rsp, err := c.ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDelete(ctx, clusterId, groupId, lvolId, reqEditors...) if err != nil { return nil, err } - return ParseClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse(rsp) + return ParseClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDeleteResponse(rsp) } -// ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetWithResponse Clusters:Backups:Sources +// ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetWithResponse Clusters:Consistency-Groups:Snapshots:List // // Returns a wrapper object for the known response body format(s). // -// Corresponds with GET /api/v2/clusters/{cluster_id}/backups/sources (the `ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGet` operationId). -func (c *ClientWithResponses) ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse, error) { - rsp, err := c.ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGet(ctx, clusterId, reqEditors...) +// Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/snapshots (the `ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGet` operationId). +func (c *ClientWithResponses) ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetResponse, error) { + rsp, err := c.ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGet(ctx, clusterId, groupId, reqEditors...) if err != nil { return nil, err } - return ParseClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse(rsp) + return ParseClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetResponse(rsp) } -// ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetWithResponse Clusters:Backups:Detail +// ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPostWithResponse Clusters:Consistency-Groups:Snapshots:Take +// +// Take one crash-consistent generation across every current member (§5). // // Returns a wrapper object for the known response body format(s). // -// Corresponds with GET /api/v2/clusters/{cluster_id}/backups/{backup_id}/ (the `ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGet` operationId). -func (c *ClientWithResponses) ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, backupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse, error) { - rsp, err := c.ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGet(ctx, clusterId, backupId, reqEditors...) +// Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/snapshots (the `ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPost` operationId). +func (c *ClientWithResponses) ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPostResponse, error) { + rsp, err := c.ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPost(ctx, clusterId, groupId, reqEditors...) if err != nil { return nil, err } - return ParseClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse(rsp) + return ParseClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPostResponse(rsp) } -// ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteWithResponse Deprecated — delete all backups for a volume +// ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDeleteWithResponse Clusters:Consistency-Groups:Snapshots:Delete // -// Deprecated. Use `DELETE /clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/backups` instead. +// Delete one generation and all its member snapshots; never the group (§10). // // Returns a wrapper object for the known response body format(s). // -// Corresponds with DELETE /api/v2/clusters/{cluster_id}/backups/{volume_id} (the `ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDelete` operationId). -// -// Deprecated: this operation has been marked as deprecated upstream, but no `x-deprecated-reason` was set -func (c *ClientWithResponses) ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse, error) { - rsp, err := c.ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDelete(ctx, clusterId, volumeId, reqEditors...) +// Corresponds with DELETE /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/snapshots/{seq} (the `ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDelete` operationId). +func (c *ClientWithResponses) ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, seq int, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDeleteResponse, error) { + rsp, err := c.ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDelete(ctx, clusterId, groupId, seq, reqEditors...) if err != nil { return nil, err } - return ParseClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse(rsp) + return ParseClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDeleteResponse(rsp) } -// ClustersCapacityApiV2ClustersClusterIdCapacityGetWithResponse Clusters:Capacity +// ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGetWithResponse Clusters:Consistency-Groups:Snapshots:Detail // // Returns a wrapper object for the known response body format(s). // -// Corresponds with GET /api/v2/clusters/{cluster_id}/capacity (the `ClustersCapacityApiV2ClustersClusterIdCapacityGet` operationId). -func (c *ClientWithResponses) ClustersCapacityApiV2ClustersClusterIdCapacityGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersCapacityApiV2ClustersClusterIdCapacityGetParams, reqEditors ...RequestEditorFn) (*ClustersCapacityApiV2ClustersClusterIdCapacityGetResponse, error) { - rsp, err := c.ClustersCapacityApiV2ClustersClusterIdCapacityGet(ctx, clusterId, params, reqEditors...) +// Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/snapshots/{seq} (the `ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGet` operationId). +func (c *ClientWithResponses) ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, seq int, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGetResponse, error) { + rsp, err := c.ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGet(ctx, clusterId, groupId, seq, reqEditors...) if err != nil { return nil, err } - return ParseClustersCapacityApiV2ClustersClusterIdCapacityGetResponse(rsp) + return ParseClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGetResponse(rsp) } // ClustersExpandApiV2ClustersClusterIdExpandPostWithResponse Clusters:Expand @@ -18252,6 +20562,26 @@ func (c *ClientWithResponses) ClustersReplicationPoliciesFailoverApiV2ClustersCl return ParseClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPostResponse(rsp) } +// ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGetWithResponse Clusters:Replication:Policies:Latest-Generation +// +// The consistency group's newest fully replicated generation, every +// member as a cloneable object on the secondary. Refused as a 400 when the +// policy has no consistency group, when no generation is complete for +// every current member yet, or when members are already split across +// generations: the same refusal a real group fail-over applies, so a drill +// never addresses a mixed-generation cut. +// +// Returns a wrapper object for the known response body format(s). +// +// Corresponds with GET /api/v2/clusters/{cluster_id}/replication/policies/{policy_id}/latest-generation (the `ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGet` operationId). +func (c *ClientWithResponses) ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, policyId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGetResponse, error) { + rsp, err := c.ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGet(ctx, clusterId, policyId, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGetResponse(rsp) +} + // ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetWithResponse Clusters:Replication:Relationships:Detail // // Replication relationship for a volume, resolvable even when the source volume @@ -18269,6 +20599,24 @@ func (c *ClientWithResponses) ClustersReplicationRelationshipsDetailApiV2Cluster return ParseClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetResponse(rsp) } +// ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGetWithResponse Clusters:Replication:Relationships:Latest-Snapshot +// +// The volume's newest fully replicated snapshot, on the secondary, as a +// cloneable object. Exists for the volume's whole replicated life: a +// test-failover drill (design §14) resolves its test point through this +// read, without touching the real replication state to find out what it is. +// +// Returns a wrapper object for the known response body format(s). +// +// Corresponds with GET /api/v2/clusters/{cluster_id}/replication/relationships/{lvol_id}/latest-snapshot (the `ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGet` operationId). +func (c *ClientWithResponses) ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, lvolId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGetResponse, error) { + rsp, err := c.ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGet(ctx, clusterId, lvolId, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGetResponse(rsp) +} + // ClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGetWithResponse Clusters:Replication:Targets:List // // Returns a wrapper object for the known response body format(s). @@ -19340,324 +21688,661 @@ func (c *ClientWithResponses) ClustersStoragePoolsVolumesReplicationStartApiV2Cl if err != nil { return nil, err } - return ParseClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostResponse(rsp) + return ParseClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostResponse(rsp) +} + +// ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGetWithResponse Clusters:Storage-Pools:Volumes:Replication:Status +// +// The typed steady-state replication status. +// +// Unlike the relationship read above, which serves cutover records and 404s +// for a volume's whole healthy replicated life, this endpoint always answers +// for a volume that exists: “state: not_replicating, role: none“ is the +// valid answer for an unreplicated volume. The csi-addons adapter derives +// its conditions and “lastSyncTime“ from this read on every reconcile. +// +// Returns a wrapper object for the known response body format(s). +// +// Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/status (the `ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGet` operationId). +func (c *ClientWithResponses) ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGetResponse, error) { + rsp, err := c.ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGet(ctx, clusterId, poolId, volumeId, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGetResponse(rsp) +} + +// ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPostWithResponse Clusters:Storage-Pools:Volumes:Replication:Stop +// +// Returns a wrapper object for the known response body format(s). +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/stop (the `ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPost` operationId). +func (c *ClientWithResponses) ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPostResponse, error) { + rsp, err := c.ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPost(ctx, clusterId, poolId, volumeId, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPostResponse(rsp) +} + +// ClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTasksGetWithResponse Clusters:Storage-Pools:Volumes:Replication:Tasks +// +// Returns a wrapper object for the known response body format(s). +// +// Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/tasks (the `ClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTasksGet` operationId). +func (c *ClientWithResponses) ClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTasksGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTasksGetResponse, error) { + rsp, err := c.ClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTasksGet(ctx, clusterId, poolId, volumeId, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTasksGetResponse(rsp) +} + +// ClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTriggerPostWithResponse Clusters:Storage-Pools:Volumes:Replication:Trigger +// +// Returns a wrapper object for the known response body format(s). +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/trigger (the `ClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTriggerPost` operationId). +func (c *ClientWithResponses) ClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTriggerPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTriggerPostResponse, error) { + rsp, err := c.ClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTriggerPost(ctx, clusterId, poolId, volumeId, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTriggerPostResponse(rsp) +} + +// ClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsGetWithResponse Clusters:Storage-Pools:Volumes:Snapshots:List +// +// Returns a wrapper object for the known response body format(s). +// +// Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/snapshots (the `ClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsGet` operationId). +func (c *ClientWithResponses) ClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsGetResponse, error) { + rsp, err := c.ClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsGet(ctx, clusterId, poolId, volumeId, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsGetResponse(rsp) +} + +// ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostWithBodyWithResponse Clusters:Storage-Pools:Volumes:Snapshots:Create +// +// Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/snapshots (the `ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPost` operationId). +func (c *ClientWithResponses) ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostResponse, error) { + rsp, err := c.ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostWithBody(ctx, clusterId, poolId, volumeId, contentType, body, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostResponse(rsp) +} + +// ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostWithResponse Clusters:Storage-Pools:Volumes:Snapshots:Create +// +// Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/snapshots (the `ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPost` operationId). +func (c *ClientWithResponses) ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, body ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostResponse, error) { + rsp, err := c.ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPost(ctx, clusterId, poolId, volumeId, body, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostResponse(rsp) +} + +// ClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigrationsGetWithResponse Clusters:Subsystems:Migrations:List +// +// Returns a wrapper object for the known response body format(s). +// +// Corresponds with GET /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/ (the `ClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigrationsGet` operationId). +func (c *ClientWithResponses) ClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigrationsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigrationsGetResponse, error) { + rsp, err := c.ClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigrationsGet(ctx, clusterId, nqn, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigrationsGetResponse(rsp) +} + +// ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostWithBodyWithResponse Clusters:Subsystems:Migrations:Create +// +// Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/ (the `ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPost` operationId). +func (c *ClientWithResponses) ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, params *ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostParams, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostResponse, error) { + rsp, err := c.ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostWithBody(ctx, clusterId, nqn, params, contentType, body, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostResponse(rsp) +} + +// ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostWithResponse Clusters:Subsystems:Migrations:Create +// +// Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/ (the `ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPost` operationId). +func (c *ClientWithResponses) ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, params *ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostParams, body ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostResponse, error) { + rsp, err := c.ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPost(ctx, clusterId, nqn, params, body, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostResponse(rsp) +} + +// ClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdDeleteWithResponse Clusters:Subsystems:Migrations:Cancel +// +// Returns a wrapper object for the known response body format(s). +// +// Corresponds with DELETE /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/{migration_id}/ (the `ClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdDelete` operationId). +func (c *ClientWithResponses) ClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdDeleteResponse, error) { + rsp, err := c.ClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdDelete(ctx, clusterId, nqn, migrationId, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdDeleteResponse(rsp) +} + +// ClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdGetWithResponse Clusters:Subsystems:Migrations:Detail +// +// Returns a wrapper object for the known response body format(s). +// +// Corresponds with GET /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/{migration_id}/ (the `ClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdGet` operationId). +func (c *ClientWithResponses) ClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdGetResponse, error) { + rsp, err := c.ClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdGet(ctx, clusterId, nqn, migrationId, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdGetResponse(rsp) +} + +// ClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdCleanupTargetPostWithResponse Clusters:Subsystems:Migrations:Cleanup-Target +// +// Idempotently remove every object this migration created on the target +// node(s). Only defined for a single-lvol migration; batch migration groups +// have no cleanup-target equivalent at the group level. +// +// Safe to call at any migration state — objects not found are reported as +// already cleaned up rather than as errors. +// +// Returns a wrapper object for the known response body format(s). +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/{migration_id}/cleanup-target (the `ClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdCleanupTargetPost` operationId). +func (c *ClientWithResponses) ClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdCleanupTargetPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdCleanupTargetPostResponse, error) { + rsp, err := c.ClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdCleanupTargetPost(ctx, clusterId, nqn, migrationId, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdCleanupTargetPostResponse(rsp) } -// ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPostWithResponse Clusters:Storage-Pools:Volumes:Replication:Stop +// ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostWithBodyWithResponse Clusters:Subsystems:Migrations:Continue // -// Returns a wrapper object for the known response body format(s). +// Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). // -// Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/stop (the `ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPost` operationId). -func (c *ClientWithResponses) ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPostResponse, error) { - rsp, err := c.ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPost(ctx, clusterId, poolId, volumeId, reqEditors...) +// Corresponds with POST /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/{migration_id}/continue (the `ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePost` operationId). +func (c *ClientWithResponses) ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostResponse, error) { + rsp, err := c.ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostWithBody(ctx, clusterId, nqn, migrationId, contentType, body, reqEditors...) if err != nil { return nil, err } - return ParseClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPostResponse(rsp) + return ParseClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostResponse(rsp) } -// ClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTasksGetWithResponse Clusters:Storage-Pools:Volumes:Replication:Tasks +// ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostWithResponse Clusters:Subsystems:Migrations:Continue // -// Returns a wrapper object for the known response body format(s). +// Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). // -// Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/tasks (the `ClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTasksGet` operationId). -func (c *ClientWithResponses) ClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTasksGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTasksGetResponse, error) { - rsp, err := c.ClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTasksGet(ctx, clusterId, poolId, volumeId, reqEditors...) +// Corresponds with POST /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/{migration_id}/continue (the `ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePost` operationId). +func (c *ClientWithResponses) ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID, body ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostResponse, error) { + rsp, err := c.ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePost(ctx, clusterId, nqn, migrationId, body, reqEditors...) if err != nil { return nil, err } - return ParseClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTasksGetResponse(rsp) + return ParseClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostResponse(rsp) } -// ClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTriggerPostWithResponse Clusters:Storage-Pools:Volumes:Replication:Trigger +// ClustersTasksListApiV2ClustersClusterIdTasksGetWithResponse Clusters:Tasks:List // // Returns a wrapper object for the known response body format(s). // -// Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/trigger (the `ClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTriggerPost` operationId). -func (c *ClientWithResponses) ClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTriggerPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTriggerPostResponse, error) { - rsp, err := c.ClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTriggerPost(ctx, clusterId, poolId, volumeId, reqEditors...) +// Corresponds with GET /api/v2/clusters/{cluster_id}/tasks/ (the `ClustersTasksListApiV2ClustersClusterIdTasksGet` operationId). +func (c *ClientWithResponses) ClustersTasksListApiV2ClustersClusterIdTasksGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersTasksListApiV2ClustersClusterIdTasksGetParams, reqEditors ...RequestEditorFn) (*ClustersTasksListApiV2ClustersClusterIdTasksGetResponse, error) { + rsp, err := c.ClustersTasksListApiV2ClustersClusterIdTasksGet(ctx, clusterId, params, reqEditors...) if err != nil { return nil, err } - return ParseClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTriggerPostResponse(rsp) + return ParseClustersTasksListApiV2ClustersClusterIdTasksGetResponse(rsp) } -// ClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsGetWithResponse Clusters:Storage-Pools:Volumes:Snapshots:List +// ClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGetWithResponse Clusters:Tasks:Detail // // Returns a wrapper object for the known response body format(s). // -// Corresponds with GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/snapshots (the `ClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsGet` operationId). -func (c *ClientWithResponses) ClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsGetResponse, error) { - rsp, err := c.ClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsGet(ctx, clusterId, poolId, volumeId, reqEditors...) +// Corresponds with GET /api/v2/clusters/{cluster_id}/tasks/{task_id}/ (the `ClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGet` operationId). +func (c *ClientWithResponses) ClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, taskId openapi_types.UUID, params *ClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGetParams, reqEditors ...RequestEditorFn) (*ClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGetResponse, error) { + rsp, err := c.ClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGet(ctx, clusterId, taskId, params, reqEditors...) if err != nil { return nil, err } - return ParseClustersStoragePoolsVolumesSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsGetResponse(rsp) + return ParseClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGetResponse(rsp) } -// ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostWithBodyWithResponse Clusters:Storage-Pools:Volumes:Snapshots:Create +// ClustersUpgradeApiV2ClustersClusterIdUpdatePostWithBodyWithResponse Clusters:Upgrade // // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). // -// Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/snapshots (the `ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPost` operationId). -func (c *ClientWithResponses) ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostResponse, error) { - rsp, err := c.ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostWithBody(ctx, clusterId, poolId, volumeId, contentType, body, reqEditors...) +// Corresponds with POST /api/v2/clusters/{cluster_id}/update (the `ClustersUpgradeApiV2ClustersClusterIdUpdatePost` operationId). +func (c *ClientWithResponses) ClustersUpgradeApiV2ClustersClusterIdUpdatePostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersUpgradeApiV2ClustersClusterIdUpdatePostResponse, error) { + rsp, err := c.ClustersUpgradeApiV2ClustersClusterIdUpdatePostWithBody(ctx, clusterId, contentType, body, reqEditors...) if err != nil { return nil, err } - return ParseClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostResponse(rsp) + return ParseClustersUpgradeApiV2ClustersClusterIdUpdatePostResponse(rsp) } -// ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostWithResponse Clusters:Storage-Pools:Volumes:Snapshots:Create +// ClustersUpgradeApiV2ClustersClusterIdUpdatePostWithResponse Clusters:Upgrade // // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). // -// Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/snapshots (the `ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPost` operationId). -func (c *ClientWithResponses) ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, body ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostResponse, error) { - rsp, err := c.ClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPost(ctx, clusterId, poolId, volumeId, body, reqEditors...) +// Corresponds with POST /api/v2/clusters/{cluster_id}/update (the `ClustersUpgradeApiV2ClustersClusterIdUpdatePost` operationId). +func (c *ClientWithResponses) ClustersUpgradeApiV2ClustersClusterIdUpdatePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, body ClustersUpgradeApiV2ClustersClusterIdUpdatePostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersUpgradeApiV2ClustersClusterIdUpdatePostResponse, error) { + rsp, err := c.ClustersUpgradeApiV2ClustersClusterIdUpdatePost(ctx, clusterId, body, reqEditors...) if err != nil { return nil, err } - return ParseClustersStoragePoolsVolumesSnapshotsCreateApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdSnapshotsPostResponse(rsp) + return ParseClustersUpgradeApiV2ClustersClusterIdUpdatePostResponse(rsp) } -// ClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigrationsGetWithResponse Clusters:Subsystems:Migrations:List +// ManagementNodesListApiV2ManagementNodesGetWithResponse Management Nodes:List // // Returns a wrapper object for the known response body format(s). // -// Corresponds with GET /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/ (the `ClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigrationsGet` operationId). -func (c *ClientWithResponses) ClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigrationsGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigrationsGetResponse, error) { - rsp, err := c.ClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigrationsGet(ctx, clusterId, nqn, reqEditors...) +// Corresponds with GET /api/v2/management-nodes/ (the `ManagementNodesListApiV2ManagementNodesGet` operationId). +func (c *ClientWithResponses) ManagementNodesListApiV2ManagementNodesGetWithResponse(ctx context.Context, params *ManagementNodesListApiV2ManagementNodesGetParams, reqEditors ...RequestEditorFn) (*ManagementNodesListApiV2ManagementNodesGetResponse, error) { + rsp, err := c.ManagementNodesListApiV2ManagementNodesGet(ctx, params, reqEditors...) if err != nil { return nil, err } - return ParseClustersSubsystemsMigrationsListApiV2ClustersClusterIdSubsystemsNqnMigrationsGetResponse(rsp) + return ParseManagementNodesListApiV2ManagementNodesGetResponse(rsp) } -// ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostWithBodyWithResponse Clusters:Subsystems:Migrations:Create +// ManagementNodeDetailApiV2ManagementNodesManagementNodeIdGetWithResponse Management Node:Detail // -// Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). +// Returns a wrapper object for the known response body format(s). // -// Corresponds with POST /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/ (the `ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPost` operationId). -func (c *ClientWithResponses) ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, params *ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostParams, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostResponse, error) { - rsp, err := c.ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostWithBody(ctx, clusterId, nqn, params, contentType, body, reqEditors...) +// Corresponds with GET /api/v2/management-nodes/{management_node_id}/ (the `ManagementNodeDetailApiV2ManagementNodesManagementNodeIdGet` operationId). +func (c *ClientWithResponses) ManagementNodeDetailApiV2ManagementNodesManagementNodeIdGetWithResponse(ctx context.Context, managementNodeId openapi_types.UUID, params *ManagementNodeDetailApiV2ManagementNodesManagementNodeIdGetParams, reqEditors ...RequestEditorFn) (*ManagementNodeDetailApiV2ManagementNodesManagementNodeIdGetResponse, error) { + rsp, err := c.ManagementNodeDetailApiV2ManagementNodesManagementNodeIdGet(ctx, managementNodeId, params, reqEditors...) if err != nil { return nil, err } - return ParseClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostResponse(rsp) + return ParseManagementNodeDetailApiV2ManagementNodesManagementNodeIdGetResponse(rsp) } -// ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostWithResponse Clusters:Subsystems:Migrations:Create -// -// Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). -// -// Corresponds with POST /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/ (the `ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPost` operationId). -func (c *ClientWithResponses) ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, params *ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostParams, body ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostResponse, error) { - rsp, err := c.ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPost(ctx, clusterId, nqn, params, body, reqEditors...) +// ParseHealthApiV2MetaHealthGetResponse parses an HTTP response from a HealthApiV2MetaHealthGetWithResponse call +func ParseHealthApiV2MetaHealthGetResponse(rsp *http.Response) (*HealthApiV2MetaHealthGetResponse, error) { + bodyBytes, err := io.ReadAll(rsp.Body) + defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - return ParseClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMigrationsPostResponse(rsp) + + response := &HealthApiV2MetaHealthGetResponse{ + Body: bodyBytes, + HTTPResponse: rsp, + } + + return response, nil } -// ClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdDeleteWithResponse Clusters:Subsystems:Migrations:Cancel -// -// Returns a wrapper object for the known response body format(s). -// -// Corresponds with DELETE /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/{migration_id}/ (the `ClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdDelete` operationId). -func (c *ClientWithResponses) ClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdDeleteResponse, error) { - rsp, err := c.ClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdDelete(ctx, clusterId, nqn, migrationId, reqEditors...) +// ParseReadyApiV2MetaReadyGetResponse parses an HTTP response from a ReadyApiV2MetaReadyGetWithResponse call +func ParseReadyApiV2MetaReadyGetResponse(rsp *http.Response) (*ReadyApiV2MetaReadyGetResponse, error) { + bodyBytes, err := io.ReadAll(rsp.Body) + defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - return ParseClustersSubsystemsMigrationsCancelApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdDeleteResponse(rsp) + + response := &ReadyApiV2MetaReadyGetResponse{ + Body: bodyBytes, + HTTPResponse: rsp, + } + + return response, nil } -// ClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdGetWithResponse Clusters:Subsystems:Migrations:Detail -// -// Returns a wrapper object for the known response body format(s). -// -// Corresponds with GET /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/{migration_id}/ (the `ClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdGet` operationId). -func (c *ClientWithResponses) ClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdGetResponse, error) { - rsp, err := c.ClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdGet(ctx, clusterId, nqn, migrationId, reqEditors...) +// ParseClustersListApiV2ClustersGetResponse parses an HTTP response from a ClustersListApiV2ClustersGetWithResponse call +func ParseClustersListApiV2ClustersGetResponse(rsp *http.Response) (*ClustersListApiV2ClustersGetResponse, error) { + bodyBytes, err := io.ReadAll(rsp.Body) + defer func() { _ = rsp.Body.Close() }() + if err != nil { + return nil, err + } + + response := &ClustersListApiV2ClustersGetResponse{ + Body: bodyBytes, + HTTPResponse: rsp, + } + + switch { + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: + var dest []ClusterDTO + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON200 = &dest + + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: + var dest HTTPValidationError + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON422 = &dest + + case rsp.StatusCode == 200: + // Content-type (text/event-stream) unsupported + + } + + return response, nil +} + +// ParseClustersCreateApiV2ClustersPostResponse parses an HTTP response from a ClustersCreateApiV2ClustersPostWithResponse call +func ParseClustersCreateApiV2ClustersPostResponse(rsp *http.Response) (*ClustersCreateApiV2ClustersPostResponse, error) { + bodyBytes, err := io.ReadAll(rsp.Body) + defer func() { _ = rsp.Body.Close() }() + if err != nil { + return nil, err + } + + response := &ClustersCreateApiV2ClustersPostResponse{ + Body: bodyBytes, + HTTPResponse: rsp, + } + + switch { + case rsp.StatusCode == 201: + break // No content-type + + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: + var dest HTTPValidationError + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON422 = &dest + + } + + return response, nil +} + +// ParseClustersDeleteApiV2ClustersClusterIdDeleteResponse parses an HTTP response from a ClustersDeleteApiV2ClustersClusterIdDeleteWithResponse call +func ParseClustersDeleteApiV2ClustersClusterIdDeleteResponse(rsp *http.Response) (*ClustersDeleteApiV2ClustersClusterIdDeleteResponse, error) { + bodyBytes, err := io.ReadAll(rsp.Body) + defer func() { _ = rsp.Body.Close() }() + if err != nil { + return nil, err + } + + response := &ClustersDeleteApiV2ClustersClusterIdDeleteResponse{ + Body: bodyBytes, + HTTPResponse: rsp, + } + + switch { + case rsp.StatusCode == 204: + break // No content-type + + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: + var dest HTTPValidationError + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON422 = &dest + + } + + return response, nil +} + +// ParseClustersDetailApiV2ClustersClusterIdGetResponse parses an HTTP response from a ClustersDetailApiV2ClustersClusterIdGetWithResponse call +func ParseClustersDetailApiV2ClustersClusterIdGetResponse(rsp *http.Response) (*ClustersDetailApiV2ClustersClusterIdGetResponse, error) { + bodyBytes, err := io.ReadAll(rsp.Body) + defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - return ParseClustersSubsystemsMigrationsDetailApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdGetResponse(rsp) + + response := &ClustersDetailApiV2ClustersClusterIdGetResponse{ + Body: bodyBytes, + HTTPResponse: rsp, + } + + switch { + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: + var dest ClusterDTO + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON200 = &dest + + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: + var dest HTTPValidationError + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON422 = &dest + + case rsp.StatusCode == 200: + // Content-type (text/event-stream) unsupported + + } + + return response, nil } -// ClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdCleanupTargetPostWithResponse Clusters:Subsystems:Migrations:Cleanup-Target -// -// Idempotently remove every object this migration created on the target -// node(s). Only defined for a single-lvol migration; batch migration groups -// have no cleanup-target equivalent at the group level. -// -// Safe to call at any migration state — objects not found are reported as -// already cleaned up rather than as errors. -// -// Returns a wrapper object for the known response body format(s). -// -// Corresponds with POST /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/{migration_id}/cleanup-target (the `ClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdCleanupTargetPost` operationId). -func (c *ClientWithResponses) ClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdCleanupTargetPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdCleanupTargetPostResponse, error) { - rsp, err := c.ClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdCleanupTargetPost(ctx, clusterId, nqn, migrationId, reqEditors...) +// ParseClustersUpdateApiV2ClustersClusterIdPutResponse parses an HTTP response from a ClustersUpdateApiV2ClustersClusterIdPutWithResponse call +func ParseClustersUpdateApiV2ClustersClusterIdPutResponse(rsp *http.Response) (*ClustersUpdateApiV2ClustersClusterIdPutResponse, error) { + bodyBytes, err := io.ReadAll(rsp.Body) + defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - return ParseClustersSubsystemsMigrationsCleanupTargetApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdCleanupTargetPostResponse(rsp) -} -// ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostWithBodyWithResponse Clusters:Subsystems:Migrations:Continue -// -// Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). -// -// Corresponds with POST /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/{migration_id}/continue (the `ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePost` operationId). -func (c *ClientWithResponses) ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostResponse, error) { - rsp, err := c.ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostWithBody(ctx, clusterId, nqn, migrationId, contentType, body, reqEditors...) - if err != nil { - return nil, err + response := &ClustersUpdateApiV2ClustersClusterIdPutResponse{ + Body: bodyBytes, + HTTPResponse: rsp, } - return ParseClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostResponse(rsp) -} -// ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostWithResponse Clusters:Subsystems:Migrations:Continue -// -// Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). -// -// Corresponds with POST /api/v2/clusters/{cluster_id}/subsystems/{nqn}/migrations/{migration_id}/continue (the `ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePost` operationId). -func (c *ClientWithResponses) ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, nqn string, migrationId openapi_types.UUID, body ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostResponse, error) { - rsp, err := c.ClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePost(ctx, clusterId, nqn, migrationId, body, reqEditors...) - if err != nil { - return nil, err + switch { + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: + var dest interface{} + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON200 = &dest + + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: + var dest HTTPValidationError + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON422 = &dest + } - return ParseClustersSubsystemsMigrationsContinueApiV2ClustersClusterIdSubsystemsNqnMigrationsMigrationIdContinuePostResponse(rsp) + + return response, nil } -// ClustersTasksListApiV2ClustersClusterIdTasksGetWithResponse Clusters:Tasks:List -// -// Returns a wrapper object for the known response body format(s). -// -// Corresponds with GET /api/v2/clusters/{cluster_id}/tasks/ (the `ClustersTasksListApiV2ClustersClusterIdTasksGet` operationId). -func (c *ClientWithResponses) ClustersTasksListApiV2ClustersClusterIdTasksGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, params *ClustersTasksListApiV2ClustersClusterIdTasksGetParams, reqEditors ...RequestEditorFn) (*ClustersTasksListApiV2ClustersClusterIdTasksGetResponse, error) { - rsp, err := c.ClustersTasksListApiV2ClustersClusterIdTasksGet(ctx, clusterId, params, reqEditors...) +// ParseClustersActivateApiV2ClustersClusterIdActivatePostResponse parses an HTTP response from a ClustersActivateApiV2ClustersClusterIdActivatePostWithResponse call +func ParseClustersActivateApiV2ClustersClusterIdActivatePostResponse(rsp *http.Response) (*ClustersActivateApiV2ClustersClusterIdActivatePostResponse, error) { + bodyBytes, err := io.ReadAll(rsp.Body) + defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - return ParseClustersTasksListApiV2ClustersClusterIdTasksGetResponse(rsp) -} -// ClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGetWithResponse Clusters:Tasks:Detail -// -// Returns a wrapper object for the known response body format(s). -// -// Corresponds with GET /api/v2/clusters/{cluster_id}/tasks/{task_id}/ (the `ClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGet` operationId). -func (c *ClientWithResponses) ClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, taskId openapi_types.UUID, params *ClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGetParams, reqEditors ...RequestEditorFn) (*ClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGetResponse, error) { - rsp, err := c.ClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGet(ctx, clusterId, taskId, params, reqEditors...) - if err != nil { - return nil, err + response := &ClustersActivateApiV2ClustersClusterIdActivatePostResponse{ + Body: bodyBytes, + HTTPResponse: rsp, } - return ParseClustersTasksDetailApiV2ClustersClusterIdTasksTaskIdGetResponse(rsp) -} -// ClustersUpgradeApiV2ClustersClusterIdUpdatePostWithBodyWithResponse Clusters:Upgrade -// -// Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). -// -// Corresponds with POST /api/v2/clusters/{cluster_id}/update (the `ClustersUpgradeApiV2ClustersClusterIdUpdatePost` operationId). -func (c *ClientWithResponses) ClustersUpgradeApiV2ClustersClusterIdUpdatePostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersUpgradeApiV2ClustersClusterIdUpdatePostResponse, error) { - rsp, err := c.ClustersUpgradeApiV2ClustersClusterIdUpdatePostWithBody(ctx, clusterId, contentType, body, reqEditors...) - if err != nil { - return nil, err + switch { + case rsp.StatusCode == 202: + break // No content-type + + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: + var dest HTTPValidationError + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON422 = &dest + } - return ParseClustersUpgradeApiV2ClustersClusterIdUpdatePostResponse(rsp) + + return response, nil } -// ClustersUpgradeApiV2ClustersClusterIdUpdatePostWithResponse Clusters:Upgrade -// -// Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). -// -// Corresponds with POST /api/v2/clusters/{cluster_id}/update (the `ClustersUpgradeApiV2ClustersClusterIdUpdatePost` operationId). -func (c *ClientWithResponses) ClustersUpgradeApiV2ClustersClusterIdUpdatePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, body ClustersUpgradeApiV2ClustersClusterIdUpdatePostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersUpgradeApiV2ClustersClusterIdUpdatePostResponse, error) { - rsp, err := c.ClustersUpgradeApiV2ClustersClusterIdUpdatePost(ctx, clusterId, body, reqEditors...) +// ParseClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostResponse parses an HTTP response from a ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostWithResponse call +func ParseClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostResponse(rsp *http.Response) (*ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostResponse, error) { + bodyBytes, err := io.ReadAll(rsp.Body) + defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - return ParseClustersUpgradeApiV2ClustersClusterIdUpdatePostResponse(rsp) -} -// ManagementNodesListApiV2ManagementNodesGetWithResponse Management Nodes:List -// -// Returns a wrapper object for the known response body format(s). -// -// Corresponds with GET /api/v2/management-nodes/ (the `ManagementNodesListApiV2ManagementNodesGet` operationId). -func (c *ClientWithResponses) ManagementNodesListApiV2ManagementNodesGetWithResponse(ctx context.Context, params *ManagementNodesListApiV2ManagementNodesGetParams, reqEditors ...RequestEditorFn) (*ManagementNodesListApiV2ManagementNodesGetResponse, error) { - rsp, err := c.ManagementNodesListApiV2ManagementNodesGet(ctx, params, reqEditors...) - if err != nil { - return nil, err + response := &ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostResponse{ + Body: bodyBytes, + HTTPResponse: rsp, } - return ParseManagementNodesListApiV2ManagementNodesGetResponse(rsp) + + switch { + case rsp.StatusCode == 202: + break // No content-type + + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: + var dest HTTPValidationError + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON422 = &dest + + } + + return response, nil } -// ManagementNodeDetailApiV2ManagementNodesManagementNodeIdGetWithResponse Management Node:Detail -// -// Returns a wrapper object for the known response body format(s). -// -// Corresponds with GET /api/v2/management-nodes/{management_node_id}/ (the `ManagementNodeDetailApiV2ManagementNodesManagementNodeIdGet` operationId). -func (c *ClientWithResponses) ManagementNodeDetailApiV2ManagementNodesManagementNodeIdGetWithResponse(ctx context.Context, managementNodeId openapi_types.UUID, params *ManagementNodeDetailApiV2ManagementNodesManagementNodeIdGetParams, reqEditors ...RequestEditorFn) (*ManagementNodeDetailApiV2ManagementNodesManagementNodeIdGetResponse, error) { - rsp, err := c.ManagementNodeDetailApiV2ManagementNodesManagementNodeIdGet(ctx, managementNodeId, params, reqEditors...) +// ParseClustersAlertsListApiV2ClustersClusterIdAlertsGetResponse parses an HTTP response from a ClustersAlertsListApiV2ClustersClusterIdAlertsGetWithResponse call +func ParseClustersAlertsListApiV2ClustersClusterIdAlertsGetResponse(rsp *http.Response) (*ClustersAlertsListApiV2ClustersClusterIdAlertsGetResponse, error) { + bodyBytes, err := io.ReadAll(rsp.Body) + defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - return ParseManagementNodeDetailApiV2ManagementNodesManagementNodeIdGetResponse(rsp) + + response := &ClustersAlertsListApiV2ClustersClusterIdAlertsGetResponse{ + Body: bodyBytes, + HTTPResponse: rsp, + } + + switch { + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: + var dest []AlertDTO + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON200 = &dest + + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: + var dest HTTPValidationError + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON422 = &dest + + } + + return response, nil } -// ParseHealthApiV2MetaHealthGetResponse parses an HTTP response from a HealthApiV2MetaHealthGetWithResponse call -func ParseHealthApiV2MetaHealthGetResponse(rsp *http.Response) (*HealthApiV2MetaHealthGetResponse, error) { +// ParseClustersBackupsListApiV2ClustersClusterIdBackupsGetResponse parses an HTTP response from a ClustersBackupsListApiV2ClustersClusterIdBackupsGetWithResponse call +func ParseClustersBackupsListApiV2ClustersClusterIdBackupsGetResponse(rsp *http.Response) (*ClustersBackupsListApiV2ClustersClusterIdBackupsGetResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - response := &HealthApiV2MetaHealthGetResponse{ + response := &ClustersBackupsListApiV2ClustersClusterIdBackupsGetResponse{ Body: bodyBytes, HTTPResponse: rsp, } + switch { + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: + var dest []BackupDTO + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON200 = &dest + + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: + var dest HTTPValidationError + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON422 = &dest + + } + return response, nil } -// ParseReadyApiV2MetaReadyGetResponse parses an HTTP response from a ReadyApiV2MetaReadyGetWithResponse call -func ParseReadyApiV2MetaReadyGetResponse(rsp *http.Response) (*ReadyApiV2MetaReadyGetResponse, error) { +// ParseClustersBackupsCreateApiV2ClustersClusterIdBackupsPostResponse parses an HTTP response from a ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostWithResponse call +func ParseClustersBackupsCreateApiV2ClustersClusterIdBackupsPostResponse(rsp *http.Response) (*ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - response := &ReadyApiV2MetaReadyGetResponse{ + response := &ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostResponse{ Body: bodyBytes, HTTPResponse: rsp, } + switch { + case rsp.StatusCode == 201: + break // No content-type + + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: + var dest HTTPValidationError + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON422 = &dest + + } + return response, nil } -// ParseClustersListApiV2ClustersGetResponse parses an HTTP response from a ClustersListApiV2ClustersGetWithResponse call -func ParseClustersListApiV2ClustersGetResponse(rsp *http.Response) (*ClustersListApiV2ClustersGetResponse, error) { +// ParseClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetResponse parses an HTTP response from a ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetWithResponse call +func ParseClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetResponse(rsp *http.Response) (*ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - response := &ClustersListApiV2ClustersGetResponse{ + response := &ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetResponse{ Body: bodyBytes, HTTPResponse: rsp, } switch { case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: - var dest []ClusterDTO + var dest []BackupPolicyDTO if err := json.Unmarshal(bodyBytes, &dest); err != nil { return nil, err } @@ -19670,23 +22355,20 @@ func ParseClustersListApiV2ClustersGetResponse(rsp *http.Response) (*ClustersLis } response.JSON422 = &dest - case rsp.StatusCode == 200: - // Content-type (text/event-stream) unsupported - } return response, nil } -// ParseClustersCreateApiV2ClustersPostResponse parses an HTTP response from a ClustersCreateApiV2ClustersPostWithResponse call -func ParseClustersCreateApiV2ClustersPostResponse(rsp *http.Response) (*ClustersCreateApiV2ClustersPostResponse, error) { +// ParseClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostResponse parses an HTTP response from a ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostWithResponse call +func ParseClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostResponse(rsp *http.Response) (*ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - response := &ClustersCreateApiV2ClustersPostResponse{ + response := &ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostResponse{ Body: bodyBytes, HTTPResponse: rsp, } @@ -19707,15 +22389,15 @@ func ParseClustersCreateApiV2ClustersPostResponse(rsp *http.Response) (*Clusters return response, nil } -// ParseClustersDeleteApiV2ClustersClusterIdDeleteResponse parses an HTTP response from a ClustersDeleteApiV2ClustersClusterIdDeleteWithResponse call -func ParseClustersDeleteApiV2ClustersClusterIdDeleteResponse(rsp *http.Response) (*ClustersDeleteApiV2ClustersClusterIdDeleteResponse, error) { +// ParseClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDeleteResponse parses an HTTP response from a ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDeleteWithResponse call +func ParseClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDeleteResponse(rsp *http.Response) (*ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDeleteResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - response := &ClustersDeleteApiV2ClustersClusterIdDeleteResponse{ + response := &ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDeleteResponse{ Body: bodyBytes, HTTPResponse: rsp, } @@ -19736,26 +22418,26 @@ func ParseClustersDeleteApiV2ClustersClusterIdDeleteResponse(rsp *http.Response) return response, nil } -// ParseClustersDetailApiV2ClustersClusterIdGetResponse parses an HTTP response from a ClustersDetailApiV2ClustersClusterIdGetWithResponse call -func ParseClustersDetailApiV2ClustersClusterIdGetResponse(rsp *http.Response) (*ClustersDetailApiV2ClustersClusterIdGetResponse, error) { +// ParseClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostResponse parses an HTTP response from a ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostWithResponse call +func ParseClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostResponse(rsp *http.Response) (*ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - response := &ClustersDetailApiV2ClustersClusterIdGetResponse{ + response := &ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostResponse{ Body: bodyBytes, HTTPResponse: rsp, } switch { - case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: - var dest ClusterDTO + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 201: + var dest interface{} if err := json.Unmarshal(bodyBytes, &dest); err != nil { return nil, err } - response.JSON200 = &dest + response.JSON201 = &dest case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: var dest HTTPValidationError @@ -19764,34 +22446,27 @@ func ParseClustersDetailApiV2ClustersClusterIdGetResponse(rsp *http.Response) (* } response.JSON422 = &dest - case rsp.StatusCode == 200: - // Content-type (text/event-stream) unsupported - } return response, nil } -// ParseClustersUpdateApiV2ClustersClusterIdPutResponse parses an HTTP response from a ClustersUpdateApiV2ClustersClusterIdPutWithResponse call -func ParseClustersUpdateApiV2ClustersClusterIdPutResponse(rsp *http.Response) (*ClustersUpdateApiV2ClustersClusterIdPutResponse, error) { +// ParseClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostResponse parses an HTTP response from a ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostWithResponse call +func ParseClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostResponse(rsp *http.Response) (*ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - response := &ClustersUpdateApiV2ClustersClusterIdPutResponse{ + response := &ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostResponse{ Body: bodyBytes, HTTPResponse: rsp, } switch { - case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: - var dest interface{} - if err := json.Unmarshal(bodyBytes, &dest); err != nil { - return nil, err - } - response.JSON200 = &dest + case rsp.StatusCode == 204: + break // No content-type case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: var dest HTTPValidationError @@ -19805,22 +22480,26 @@ func ParseClustersUpdateApiV2ClustersClusterIdPutResponse(rsp *http.Response) (* return response, nil } -// ParseClustersActivateApiV2ClustersClusterIdActivatePostResponse parses an HTTP response from a ClustersActivateApiV2ClustersClusterIdActivatePostWithResponse call -func ParseClustersActivateApiV2ClustersClusterIdActivatePostResponse(rsp *http.Response) (*ClustersActivateApiV2ClustersClusterIdActivatePostResponse, error) { +// ParseClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetResponse parses an HTTP response from a ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetWithResponse call +func ParseClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetResponse(rsp *http.Response) (*ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - response := &ClustersActivateApiV2ClustersClusterIdActivatePostResponse{ + response := &ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetResponse{ Body: bodyBytes, HTTPResponse: rsp, } switch { - case rsp.StatusCode == 202: - break // No content-type + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: + var dest interface{} + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON200 = &dest case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: var dest HTTPValidationError @@ -19834,22 +22513,26 @@ func ParseClustersActivateApiV2ClustersClusterIdActivatePostResponse(rsp *http.R return response, nil } -// ParseClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostResponse parses an HTTP response from a ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostWithResponse call -func ParseClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostResponse(rsp *http.Response) (*ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostResponse, error) { +// ParseClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse parses an HTTP response from a ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostWithResponse call +func ParseClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse(rsp *http.Response) (*ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - response := &ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostResponse{ + response := &ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse{ Body: bodyBytes, HTTPResponse: rsp, } switch { - case rsp.StatusCode == 202: - break // No content-type + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: + var dest interface{} + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON200 = &dest case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: var dest HTTPValidationError @@ -19863,26 +22546,26 @@ func ParseClustersAddreplicationApiV2ClustersClusterIdAddreplicationPostResponse return response, nil } -// ParseClustersBackupsListApiV2ClustersClusterIdBackupsGetResponse parses an HTTP response from a ClustersBackupsListApiV2ClustersClusterIdBackupsGetWithResponse call -func ParseClustersBackupsListApiV2ClustersClusterIdBackupsGetResponse(rsp *http.Response) (*ClustersBackupsListApiV2ClustersClusterIdBackupsGetResponse, error) { +// ParseClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse parses an HTTP response from a ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostWithResponse call +func ParseClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse(rsp *http.Response) (*ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - response := &ClustersBackupsListApiV2ClustersClusterIdBackupsGetResponse{ + response := &ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse{ Body: bodyBytes, HTTPResponse: rsp, } switch { - case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: - var dest []BackupDTO + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 202: + var dest interface{} if err := json.Unmarshal(bodyBytes, &dest); err != nil { return nil, err } - response.JSON200 = &dest + response.JSON202 = &dest case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: var dest HTTPValidationError @@ -19896,22 +22579,26 @@ func ParseClustersBackupsListApiV2ClustersClusterIdBackupsGetResponse(rsp *http. return response, nil } -// ParseClustersBackupsCreateApiV2ClustersClusterIdBackupsPostResponse parses an HTTP response from a ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostWithResponse call -func ParseClustersBackupsCreateApiV2ClustersClusterIdBackupsPostResponse(rsp *http.Response) (*ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostResponse, error) { +// ParseClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse parses an HTTP response from a ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostWithResponse call +func ParseClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse(rsp *http.Response) (*ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - response := &ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostResponse{ + response := &ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse{ Body: bodyBytes, HTTPResponse: rsp, } switch { - case rsp.StatusCode == 201: - break // No content-type + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: + var dest interface{} + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON200 = &dest case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: var dest HTTPValidationError @@ -19925,22 +22612,22 @@ func ParseClustersBackupsCreateApiV2ClustersClusterIdBackupsPostResponse(rsp *ht return response, nil } -// ParseClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetResponse parses an HTTP response from a ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetWithResponse call -func ParseClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetResponse(rsp *http.Response) (*ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetResponse, error) { +// ParseClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse parses an HTTP response from a ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetWithResponse call +func ParseClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse(rsp *http.Response) (*ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - response := &ClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesGetResponse{ + response := &ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse{ Body: bodyBytes, HTTPResponse: rsp, } switch { case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: - var dest []BackupPolicyDTO + var dest interface{} if err := json.Unmarshal(bodyBytes, &dest); err != nil { return nil, err } @@ -19958,22 +22645,26 @@ func ParseClustersBackupPoliciesListApiV2ClustersClusterIdBackupsBackupPoliciesG return response, nil } -// ParseClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostResponse parses an HTTP response from a ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostWithResponse call -func ParseClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostResponse(rsp *http.Response) (*ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostResponse, error) { +// ParseClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse parses an HTTP response from a ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetWithResponse call +func ParseClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse(rsp *http.Response) (*ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - response := &ClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPoliciesPostResponse{ + response := &ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse{ Body: bodyBytes, HTTPResponse: rsp, } switch { - case rsp.StatusCode == 201: - break // No content-type + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: + var dest BackupDTO + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON200 = &dest case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: var dest HTTPValidationError @@ -19987,15 +22678,15 @@ func ParseClustersBackupPoliciesCreateApiV2ClustersClusterIdBackupsBackupPolicie return response, nil } -// ParseClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDeleteResponse parses an HTTP response from a ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDeleteWithResponse call -func ParseClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDeleteResponse(rsp *http.Response) (*ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDeleteResponse, error) { +// ParseClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse parses an HTTP response from a ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteWithResponse call +func ParseClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse(rsp *http.Response) (*ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - response := &ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDeleteResponse{ + response := &ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse{ Body: bodyBytes, HTTPResponse: rsp, } @@ -20016,26 +22707,26 @@ func ParseClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPolicie return response, nil } -// ParseClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostResponse parses an HTTP response from a ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostWithResponse call -func ParseClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostResponse(rsp *http.Response) (*ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostResponse, error) { +// ParseClustersCapacityApiV2ClustersClusterIdCapacityGetResponse parses an HTTP response from a ClustersCapacityApiV2ClustersClusterIdCapacityGetWithResponse call +func ParseClustersCapacityApiV2ClustersClusterIdCapacityGetResponse(rsp *http.Response) (*ClustersCapacityApiV2ClustersClusterIdCapacityGetResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - response := &ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPostResponse{ + response := &ClustersCapacityApiV2ClustersClusterIdCapacityGetResponse{ Body: bodyBytes, HTTPResponse: rsp, } switch { - case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 201: + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: var dest interface{} if err := json.Unmarshal(bodyBytes, &dest); err != nil { return nil, err } - response.JSON201 = &dest + response.JSON200 = &dest case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: var dest HTTPValidationError @@ -20049,22 +22740,26 @@ func ParseClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPolicie return response, nil } -// ParseClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostResponse parses an HTTP response from a ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostWithResponse call -func ParseClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostResponse(rsp *http.Response) (*ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostResponse, error) { +// ParseClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetResponse parses an HTTP response from a ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetWithResponse call +func ParseClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetResponse(rsp *http.Response) (*ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - response := &ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPostResponse{ + response := &ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetResponse{ Body: bodyBytes, HTTPResponse: rsp, } switch { - case rsp.StatusCode == 204: - break // No content-type + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: + var dest []ConsistencyGroupDTO + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON200 = &dest case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: var dest HTTPValidationError @@ -20078,22 +22773,22 @@ func ParseClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPolicie return response, nil } -// ParseClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetResponse parses an HTTP response from a ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetWithResponse call -func ParseClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetResponse(rsp *http.Response) (*ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetResponse, error) { +// ParseClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGetResponse parses an HTTP response from a ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGetWithResponse call +func ParseClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGetResponse(rsp *http.Response) (*ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGetResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - response := &ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetResponse{ + response := &ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGetResponse{ Body: bodyBytes, HTTPResponse: rsp, } switch { case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: - var dest interface{} + var dest ConsistencyGroupDTO if err := json.Unmarshal(bodyBytes, &dest); err != nil { return nil, err } @@ -20111,22 +22806,22 @@ func ParseClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetResponse(rs return response, nil } -// ParseClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse parses an HTTP response from a ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostWithResponse call -func ParseClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse(rsp *http.Response) (*ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse, error) { +// ParseClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGetResponse parses an HTTP response from a ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGetWithResponse call +func ParseClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGetResponse(rsp *http.Response) (*ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGetResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - response := &ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse{ + response := &ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGetResponse{ Body: bodyBytes, HTTPResponse: rsp, } switch { case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: - var dest interface{} + var dest []ConsistencyGroupMemberDTO if err := json.Unmarshal(bodyBytes, &dest); err != nil { return nil, err } @@ -20144,26 +22839,26 @@ func ParseClustersBackupsImportApiV2ClustersClusterIdBackupsImportPostResponse(r return response, nil } -// ParseClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse parses an HTTP response from a ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostWithResponse call -func ParseClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse(rsp *http.Response) (*ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse, error) { +// ParseClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostResponse parses an HTTP response from a ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostWithResponse call +func ParseClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostResponse(rsp *http.Response) (*ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - response := &ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse{ + response := &ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostResponse{ Body: bodyBytes, HTTPResponse: rsp, } switch { - case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 202: - var dest interface{} + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: + var dest ConsistencyGroupMemberDTO if err := json.Unmarshal(bodyBytes, &dest); err != nil { return nil, err } - response.JSON202 = &dest + response.JSON200 = &dest case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: var dest HTTPValidationError @@ -20177,26 +22872,22 @@ func ParseClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostResponse return response, nil } -// ParseClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse parses an HTTP response from a ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostWithResponse call -func ParseClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse(rsp *http.Response) (*ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse, error) { +// ParseClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDeleteResponse parses an HTTP response from a ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDeleteWithResponse call +func ParseClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDeleteResponse(rsp *http.Response) (*ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDeleteResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - response := &ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostResponse{ + response := &ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDeleteResponse{ Body: bodyBytes, HTTPResponse: rsp, } switch { - case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: - var dest interface{} - if err := json.Unmarshal(bodyBytes, &dest); err != nil { - return nil, err - } - response.JSON200 = &dest + case rsp.StatusCode == 204: + break // No content-type case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: var dest HTTPValidationError @@ -20210,22 +22901,22 @@ func ParseClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPo return response, nil } -// ParseClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse parses an HTTP response from a ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetWithResponse call -func ParseClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse(rsp *http.Response) (*ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse, error) { +// ParseClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetResponse parses an HTTP response from a ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetWithResponse call +func ParseClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetResponse(rsp *http.Response) (*ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - response := &ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse{ + response := &ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetResponse{ Body: bodyBytes, HTTPResponse: rsp, } switch { case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: - var dest interface{} + var dest []ConsistencyGroupGenerationDTO if err := json.Unmarshal(bodyBytes, &dest); err != nil { return nil, err } @@ -20243,22 +22934,22 @@ func ParseClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGetResponse( return response, nil } -// ParseClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse parses an HTTP response from a ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetWithResponse call -func ParseClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse(rsp *http.Response) (*ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse, error) { +// ParseClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPostResponse parses an HTTP response from a ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPostWithResponse call +func ParseClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPostResponse(rsp *http.Response) (*ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPostResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - response := &ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse{ + response := &ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPostResponse{ Body: bodyBytes, HTTPResponse: rsp, } switch { case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: - var dest BackupDTO + var dest ConsistencyGroupGenerationDTO if err := json.Unmarshal(bodyBytes, &dest); err != nil { return nil, err } @@ -20276,15 +22967,15 @@ func ParseClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGetResponse( return response, nil } -// ParseClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse parses an HTTP response from a ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteWithResponse call -func ParseClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse(rsp *http.Response) (*ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse, error) { +// ParseClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDeleteResponse parses an HTTP response from a ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDeleteWithResponse call +func ParseClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDeleteResponse(rsp *http.Response) (*ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDeleteResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - response := &ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteResponse{ + response := &ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDeleteResponse{ Body: bodyBytes, HTTPResponse: rsp, } @@ -20305,22 +22996,22 @@ func ParseClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDeleteRespon return response, nil } -// ParseClustersCapacityApiV2ClustersClusterIdCapacityGetResponse parses an HTTP response from a ClustersCapacityApiV2ClustersClusterIdCapacityGetWithResponse call -func ParseClustersCapacityApiV2ClustersClusterIdCapacityGetResponse(rsp *http.Response) (*ClustersCapacityApiV2ClustersClusterIdCapacityGetResponse, error) { +// ParseClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGetResponse parses an HTTP response from a ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGetWithResponse call +func ParseClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGetResponse(rsp *http.Response) (*ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGetResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) defer func() { _ = rsp.Body.Close() }() if err != nil { return nil, err } - response := &ClustersCapacityApiV2ClustersClusterIdCapacityGetResponse{ + response := &ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGetResponse{ Body: bodyBytes, HTTPResponse: rsp, } switch { case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: - var dest interface{} + var dest ConsistencyGroupGenerationDTO if err := json.Unmarshal(bodyBytes, &dest); err != nil { return nil, err } @@ -20622,6 +23313,39 @@ func ParseClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPo return response, nil } +// ParseClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGetResponse parses an HTTP response from a ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGetWithResponse call +func ParseClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGetResponse(rsp *http.Response) (*ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGetResponse, error) { + bodyBytes, err := io.ReadAll(rsp.Body) + defer func() { _ = rsp.Body.Close() }() + if err != nil { + return nil, err + } + + response := &ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGetResponse{ + Body: bodyBytes, + HTTPResponse: rsp, + } + + switch { + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: + var dest ReplicatedGenerationDTO + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON200 = &dest + + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: + var dest HTTPValidationError + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON422 = &dest + + } + + return response, nil +} + // ParseClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetResponse parses an HTTP response from a ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetWithResponse call func ParseClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetResponse(rsp *http.Response) (*ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) @@ -20655,6 +23379,39 @@ func ParseClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicatio return response, nil } +// ParseClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGetResponse parses an HTTP response from a ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGetWithResponse call +func ParseClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGetResponse(rsp *http.Response) (*ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGetResponse, error) { + bodyBytes, err := io.ReadAll(rsp.Body) + defer func() { _ = rsp.Body.Close() }() + if err != nil { + return nil, err + } + + response := &ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGetResponse{ + Body: bodyBytes, + HTTPResponse: rsp, + } + + switch { + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: + var dest ReplicatedSnapshotDTO + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON200 = &dest + + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: + var dest HTTPValidationError + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON422 = &dest + + } + + return response, nil +} + // ParseClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGetResponse parses an HTTP response from a ClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGetWithResponse call func ParseClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGetResponse(rsp *http.Response) (*ClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGetResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) @@ -22602,6 +25359,39 @@ func ParseClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStora return response, nil } +// ParseClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGetResponse parses an HTTP response from a ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGetWithResponse call +func ParseClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGetResponse(rsp *http.Response) (*ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGetResponse, error) { + bodyBytes, err := io.ReadAll(rsp.Body) + defer func() { _ = rsp.Body.Close() }() + if err != nil { + return nil, err + } + + response := &ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGetResponse{ + Body: bodyBytes, + HTTPResponse: rsp, + } + + switch { + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: + var dest ReplicationStatusDTO + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON200 = &dest + + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: + var dest HTTPValidationError + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON422 = &dest + + } + + return response, nil +} + // ParseClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPostResponse parses an HTTP response from a ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPostWithResponse call func ParseClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPostResponse(rsp *http.Response) (*ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPostResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) diff --git a/csi-driver/go.mod b/csi-driver/go.mod index f5821c3b8..3544e2a6d 100644 --- a/csi-driver/go.mod +++ b/csi-driver/go.mod @@ -23,27 +23,46 @@ require ( cel.dev/expr v0.25.2 // indirect github.com/Masterminds/semver/v3 v3.4.0 // indirect github.com/antlr4-go/antlr/v4 v4.13.0 // indirect + github.com/apapsch/go-jsonmerge/v2 v2.0.0 // indirect github.com/cenkalti/backoff/v5 v5.0.3 // indirect + github.com/csi-addons/spec v0.2.0 // indirect github.com/distribution/reference v0.6.0 // indirect + github.com/dprotaso/go-yit v0.0.0-20220510233725-9ba8df137936 // indirect github.com/fsnotify/fsnotify v1.9.0 // indirect github.com/fxamacker/cbor/v2 v2.9.0 // indirect + github.com/gabriel-vasile/mimetype v1.4.13 // indirect + github.com/getkin/kin-openapi v0.142.0 // indirect github.com/go-openapi/swag/jsonname v0.26.0 // indirect + github.com/go-playground/locales v0.14.1 // indirect + github.com/go-playground/universal-translator v0.18.1 // indirect + github.com/go-playground/validator/v10 v10.30.3 // indirect github.com/go-task/slim-sprig/v3 v3.0.0 // indirect + github.com/golang/protobuf v1.5.4 // indirect github.com/google/cel-go v0.26.0 // indirect github.com/google/gnostic-models v0.7.0 // indirect github.com/google/pprof v0.0.0-20260402051712-545e8a4df936 // indirect github.com/gorilla/websocket v1.5.4-0.20250319132907-e064f32e3674 // indirect github.com/grpc-ecosystem/grpc-gateway/v2 v2.27.7 // indirect github.com/hashicorp/yamux v0.1.2 // indirect + github.com/leodido/go-urn v1.4.0 // indirect + github.com/oapi-codegen/oapi-codegen/v2 v2.8.0 // indirect + github.com/oapi-codegen/runtime v1.6.0 // indirect + github.com/oasdiff/yaml v0.1.1 // indirect + github.com/oasdiff/yaml3 v0.0.14 // indirect github.com/pmezard/go-difflib v1.0.1-0.20181226105442-5d4384ee4fb2 // indirect github.com/robfig/cron/v3 v3.0.1 // indirect + github.com/santhosh-tekuri/jsonschema/v6 v6.0.2 // indirect + github.com/speakeasy-api/jsonpath v0.6.3 // indirect + github.com/speakeasy-api/openapi v1.24.0 // indirect github.com/stoewer/go-strcase v1.3.0 // indirect + github.com/vmware-labs/yaml-jsonpath v0.3.2 // indirect github.com/x448/float16 v0.8.4 // indirect go.opentelemetry.io/otel/exporters/otlp/otlptrace v1.40.0 // indirect go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.40.0 // indirect go.uber.org/mock v0.5.2 // indirect go.yaml.in/yaml/v2 v2.4.3 // indirect go.yaml.in/yaml/v3 v3.0.4 // indirect + golang.org/x/crypto v0.56.0 // indirect golang.org/x/exp v0.0.0-20251219203646-944ab1f22d93 // indirect golang.org/x/mod v0.38.0 // indirect golang.org/x/sync v0.22.0 // indirect diff --git a/csi-driver/go.sum b/csi-driver/go.sum index 36567e4df..fa9883c5c 100644 --- a/csi-driver/go.sum +++ b/csi-driver/go.sum @@ -2,35 +2,55 @@ cel.dev/expr v0.25.2 h1:K6j46C81hXtZQfuX60cVWQFBJahKSE2gfRbNuvr5bFs= cel.dev/expr v0.25.2/go.mod h1:hrXvqGP6G6gyx8UAHSHJ5RGk//1Oj5nXQ2NI02Nrsg4= github.com/Masterminds/semver/v3 v3.4.0 h1:Zog+i5UMtVoCU8oKka5P7i9q9HgrJeGzI9SA1Xbatp0= github.com/Masterminds/semver/v3 v3.4.0/go.mod h1:4V+yj/TJE1HU9XfppCwVMZq3I84lprf4nC11bSS5beM= +github.com/RaveNoX/go-jsoncommentstrip v1.0.0/go.mod h1:78ihd09MekBnJnxpICcwzCMzGrKSKYe4AqU6PDYYpjk= github.com/antlr4-go/antlr/v4 v4.13.0 h1:lxCg3LAv+EUK6t1i0y1V6/SLeUi0eKEKdhQAlS8TVTI= github.com/antlr4-go/antlr/v4 v4.13.0/go.mod h1:pfChB/xh/Unjila75QW7+VU4TSnWnnk9UTnmpPaOR2g= +github.com/apapsch/go-jsonmerge/v2 v2.0.0 h1:axGnT1gRIfimI7gJifB699GoE/oq+F2MU7Dml6nw9rQ= +github.com/apapsch/go-jsonmerge/v2 v2.0.0/go.mod h1:lvDnEdqiQrp0O42VQGgmlKpxL1AP2+08jFMw88y4klk= github.com/armon/go-socks5 v0.0.0-20160902184237-e75332964ef5 h1:0CwZNZbxp69SHPdPJAN/hZIm0C4OItdklCFmMRWYpio= github.com/armon/go-socks5 v0.0.0-20160902184237-e75332964ef5/go.mod h1:wHh0iHkYZB8zMSxRWpUBQtwG5a7fFgvEO+odwuTv2gs= github.com/beorn7/perks v1.0.1 h1:VlbKKnNfV8bJzeqoa4cOKqO6bYr3WgKZxO8Z16+hsOM= github.com/beorn7/perks v1.0.1/go.mod h1:G2ZrVWU2WbWT9wwq4/hrbKbnv/1ERSJQ0ibhJ6rlkpw= github.com/blang/semver/v4 v4.0.0 h1:1PFHFE6yCCTv8C1TeyNNarDzntLi7wMI5i/pzqYIsAM= github.com/blang/semver/v4 v4.0.0/go.mod h1:IbckMUScFkM3pff0VJDNKRiT6TG/YpiHIM2yvyW5YoQ= +github.com/bmatcuk/doublestar v1.1.1/go.mod h1:UD6OnuiIn0yFxxA2le/rnRU1G4RaI4UvFv1sNto9p6w= github.com/cenkalti/backoff/v5 v5.0.3 h1:ZN+IMa753KfX5hd8vVaMixjnqRZ3y8CuJKRKj1xcsSM= github.com/cenkalti/backoff/v5 v5.0.3/go.mod h1:rkhZdG3JZukswDf7f0cwqPNk4K0sa+F97BxZthm/crw= github.com/cespare/xxhash/v2 v2.3.0 h1:UL815xU9SqsFlibzuggzjXhog7bL6oX9BbNZnL2UFvs= github.com/cespare/xxhash/v2 v2.3.0/go.mod h1:VGX0DQ3Q6kWi7AoAeZDth3/j3BFtOZR5XLFGgcrjCOs= +github.com/chzyer/logex v1.1.10/go.mod h1:+Ywpsq7O8HXn0nuIou7OrIPyXbp3wmkHB+jjWRnGsAI= +github.com/chzyer/readline v0.0.0-20180603132655-2972be24d48e/go.mod h1:nSuG5e5PlCu98SY8svDHJxuZscDgtXS6KTTbou5AhLI= +github.com/chzyer/test v0.0.0-20180213035817-a1ea475d72b1/go.mod h1:Q3SI9o4m/ZMnBNeIyt5eFwwo7qiLfzFZmjNmxjkiQlU= github.com/container-storage-interface/spec v1.12.0 h1:zrFOEqpR5AghNaaDG4qyedwPBqU2fU0dWjLQMP/azK0= github.com/container-storage-interface/spec v1.12.0/go.mod h1:txsm+MA2B2WDa5kW69jNbqPnvTtfvZma7T/zsAZ9qX8= github.com/cpuguy83/go-md2man/v2 v2.0.6/go.mod h1:oOW0eioCTA6cOiMLiUPZOpcVxMig6NIQQ7OS05n1F4g= +github.com/csi-addons/spec v0.2.0 h1:Ews7bxpN9P6nFxl1XvMg87cR1wLROdH1FzSfLfb4VfI= +github.com/csi-addons/spec v0.2.0/go.mod h1:Mwq4iLiUV4s+K1bszcWU6aMsR5KPsbIYzzszJ6+56vI= github.com/davecgh/go-spew v1.1.0/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38= github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38= github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc h1:U9qPSI2PIWSS1VwoXQT9A3Wy9MM3WgvqSxFWenqJduM= github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38= github.com/distribution/reference v0.6.0 h1:0IXCQ5g4/QMHHkarYzh5l+u8T3t73zM5QvfrDyIgxBk= github.com/distribution/reference v0.6.0/go.mod h1:BbU0aIcezP1/5jX/8MP0YiH4SdvB5Y4f/wlDRiLyi3E= +github.com/dlclark/regexp2 v1.11.4 h1:rPYF9/LECdNymJufQKmri9gV604RvvABwgOA8un7yAo= +github.com/dlclark/regexp2 v1.11.4/go.mod h1:DHkYz0B9wPfa6wondMfaivmHpzrQ3v9q8cnmRbL6yW8= +github.com/dprotaso/go-yit v0.0.0-20191028211022-135eb7262960/go.mod h1:9HQzr9D/0PGwMEbC3d5AB7oi67+h4TsQqItC1GVYG58= +github.com/dprotaso/go-yit v0.0.0-20220510233725-9ba8df137936 h1:PRxIJD8XjimM5aTknUK9w6DHLDox2r2M3DI4i2pnd3w= +github.com/dprotaso/go-yit v0.0.0-20220510233725-9ba8df137936/go.mod h1:ttYvX5qlB+mlV1okblJqcSMtR4c52UKxDiX9GRBS8+Q= github.com/emicklei/go-restful/v3 v3.13.0 h1:C4Bl2xDndpU6nJ4bc1jXd+uTmYPVUwkD6bFY/oTyCes= github.com/emicklei/go-restful/v3 v3.13.0/go.mod h1:6n3XBCmQQb25CM2LCACGz8ukIrRry+4bhvbpWn3mrbc= github.com/felixge/httpsnoop v1.0.4 h1:NFTV2Zj1bL4mc9sqWACXbQFVBBg2W3GPvqp8/ESS2Wg= github.com/felixge/httpsnoop v1.0.4/go.mod h1:m8KPJKqk1gH5J9DgRY2ASl2lWCfGKXixSwevea8zH2U= +github.com/fsnotify/fsnotify v1.4.7/go.mod h1:jwhsz4b93w/PPRr/qN1Yymfu8t87LnFCMoQvtojpjFo= +github.com/fsnotify/fsnotify v1.4.9/go.mod h1:znqG4EE+3YCdAaPaxE2ZRY/06pZUdp0tY4IgpuI1SZQ= github.com/fsnotify/fsnotify v1.9.0 h1:2Ml+OJNzbYCTzsxtv8vKSFD9PbJjmhYF14k/jKC7S9k= github.com/fsnotify/fsnotify v1.9.0/go.mod h1:8jBTzvmWwFyi3Pb8djgCCO5IBqzKJ/Jwo8TRcHyHii0= github.com/fxamacker/cbor/v2 v2.9.0 h1:NpKPmjDBgUfBms6tr6JZkTHtfFGcMKsw3eGcmD/sapM= github.com/fxamacker/cbor/v2 v2.9.0/go.mod h1:vM4b+DJCtHn+zz7h3FFp/hDAI9WNWCsZj23V5ytsSxQ= +github.com/gabriel-vasile/mimetype v1.4.13 h1:46nXokslUBsAJE/wMsp5gtO500a4F3Nkz9Ufpk2AcUM= +github.com/gabriel-vasile/mimetype v1.4.13/go.mod h1:d+9Oxyo1wTzWdyVUPMmXFvp4F9tea18J8ufA774AB3s= +github.com/getkin/kin-openapi v0.142.0 h1:izj0vBdFprMhitfzaX8sTqztsEQyvwhssBoB6n8NO7w= +github.com/getkin/kin-openapi v0.142.0/go.mod h1:3BH9M9XDe/y9M5DSvEocVYAYq1w0qrhJHjC/vZi0AaY= github.com/gkampitakis/ciinfo v0.3.2 h1:JcuOPk8ZU7nZQjdUhctuhQofk7BGHuIy0c9Ez8BNhXs= github.com/gkampitakis/ciinfo v0.3.2/go.mod h1:1NIwaOcFChN4fa/B0hEBdAb6npDlFL8Bwx4dfRLRqAo= github.com/gkampitakis/go-diff v1.3.2 h1:Qyn0J9XJSDTgnsgHRdz9Zp24RaJeKMUHg2+PDZZdC4M= @@ -55,19 +75,42 @@ github.com/go-openapi/swag/jsonname v0.26.0 h1:gV1NFX9M8avo0YSpmWogqfQISigCmpaiN github.com/go-openapi/swag/jsonname v0.26.0/go.mod h1:urBBR8bZNoDYGr653ynhIx+gTeIz0ARZxHkAPktJK2M= github.com/go-openapi/testify/v2 v2.4.2 h1:tiByHpvE9uHrrKjOszax7ZvKB7QOgizBWGBLuq0ePx4= github.com/go-openapi/testify/v2 v2.4.2/go.mod h1:SgsVHtfooshd0tublTtJ50FPKhujf47YRqauXXOUxfw= +github.com/go-playground/assert/v2 v2.2.0 h1:JvknZsQTYeFEAhQwI4qEt9cyV5ONwRHC+lYKSsYSR8s= +github.com/go-playground/assert/v2 v2.2.0/go.mod h1:VDjEfimB/XKnb+ZQfWdccd7VUvScMdVu0Titje2rxJ4= +github.com/go-playground/locales v0.14.1 h1:EWaQ/wswjilfKLTECiXz7Rh+3BjFhfDFKv/oXslEjJA= +github.com/go-playground/locales v0.14.1/go.mod h1:hxrqLVvrK65+Rwrd5Fc6F2O76J/NuW9t0sjnWqG1slY= +github.com/go-playground/universal-translator v0.18.1 h1:Bcnm0ZwsGyWbCzImXv+pAJnYK9S473LQFuzCbDbfSFY= +github.com/go-playground/universal-translator v0.18.1/go.mod h1:xekY+UJKNuX9WP91TpwSH2VMlDf28Uj24BCp08ZFTUY= +github.com/go-playground/validator/v10 v10.30.3 h1:4MU6YkEwx7GbcPJOZxrtbu+QfF3pJLJuaYTeAH0DYy8= +github.com/go-playground/validator/v10 v10.30.3/go.mod h1:4Axh7oCNGcoGkqLoE4YWt6n20mcEIsPRlB7vPk3lpyc= +github.com/go-task/slim-sprig v0.0.0-20210107165309-348f09dbbbc0/go.mod h1:fyg7847qk6SyHyPtNmDHnmrv/HOrqktSC+C9fM+CJOE= github.com/go-task/slim-sprig/v3 v3.0.0 h1:sUs3vkvUymDpBKi3qH1YSqBQk9+9D/8M2mN1vB6EwHI= github.com/go-task/slim-sprig/v3 v3.0.0/go.mod h1:W848ghGpv3Qj3dhTPRyJypKRiqCdHZiAzKg9hl15HA8= github.com/goccy/go-yaml v1.18.0 h1:8W7wMFS12Pcas7KU+VVkaiCng+kG8QiFeFwzFb+rwuw= github.com/goccy/go-yaml v1.18.0/go.mod h1:XBurs7gK8ATbW4ZPGKgcbrY1Br56PdM69F7LkFRi1kA= +github.com/golang/protobuf v1.2.0/go.mod h1:6lQm79b+lXiMfvg/cZm0SGofjICqVBUtrP5yJMmIC1U= +github.com/golang/protobuf v1.4.0-rc.1/go.mod h1:ceaxUfeHdC40wWswd/P6IGgMaK3YpKi5j83Wpe3EHw8= +github.com/golang/protobuf v1.4.0-rc.1.0.20200221234624-67d41d38c208/go.mod h1:xKAWHe0F5eneWXFV3EuXVDTCmh+JuBKY0li0aMyXATA= +github.com/golang/protobuf v1.4.0-rc.2/go.mod h1:LlEzMj4AhA7rCAGe4KMBDvJI+AwstrUpVNzEA03Pprs= +github.com/golang/protobuf v1.4.0-rc.4.0.20200313231945-b860323f09d0/go.mod h1:WU3c8KckQ9AFe+yFwt9sWVRKCVIyN9cPHBJSNnbL67w= +github.com/golang/protobuf v1.4.0/go.mod h1:jodUvKwWbYaEsadDk5Fwe5c77LiNKVO9IDvqG2KuDX0= +github.com/golang/protobuf v1.4.2/go.mod h1:oDoupMAO8OvCJWAcko0GGGIgR6R6ocIYbsSw735rRwI= +github.com/golang/protobuf v1.5.0/go.mod h1:FsONVRAS9T7sI+LIUmWTfcYkHO4aIWwzhcaSAoJOfIk= +github.com/golang/protobuf v1.5.2/go.mod h1:XVQd3VNwM+JqD3oG2Ue2ip4fOMUkwXdXDdiuN0vRsmY= github.com/golang/protobuf v1.5.4 h1:i7eJL8qZTpSEXOPTxNKhASYpMn+8e5Q6AdndVa1dWek= github.com/golang/protobuf v1.5.4/go.mod h1:lnTiLA8Wa4RWRcIUkrtSVa5nRhsEGBg48fD6rSs7xps= github.com/google/cel-go v0.26.0 h1:DPGjXackMpJWH680oGY4lZhYjIameYmR+/6RBdDGmaI= github.com/google/cel-go v0.26.0/go.mod h1:A9O8OU9rdvrK5MQyrqfIxo1a0u4g3sF8KB6PUIaryMM= github.com/google/gnostic-models v0.7.0 h1:qwTtogB15McXDaNqTZdzPJRHvaVJlAl+HVQnLmJEJxo= github.com/google/gnostic-models v0.7.0/go.mod h1:whL5G0m6dmc5cPxKc5bdKdEN3UjI7OUGxBlw57miDrQ= +github.com/google/go-cmp v0.3.0/go.mod h1:8QqcDgzrUqlUb/G2PQTWiueGozuR1884gddMywk6iLU= +github.com/google/go-cmp v0.3.1/go.mod h1:8QqcDgzrUqlUb/G2PQTWiueGozuR1884gddMywk6iLU= +github.com/google/go-cmp v0.4.0/go.mod h1:v8dTdLbMG2kIc/vJvl+f65V22dbkXbowE6jgT/gNBxE= +github.com/google/go-cmp v0.5.5/go.mod h1:v8dTdLbMG2kIc/vJvl+f65V22dbkXbowE6jgT/gNBxE= github.com/google/go-cmp v0.7.0 h1:wk8382ETsv4JYUZwIsn6YpYiWiBsYLSJiTsyBybVuN8= github.com/google/go-cmp v0.7.0/go.mod h1:pXiqmnSA92OHEEa9HXL2W4E7lf9JzCmGVUdgjX3N/iU= github.com/google/gofuzz v1.0.0/go.mod h1:dBl0BpW6vV/+mYPU4Po3pmUjxk6FQPldtuIdl/M65Eg= +github.com/google/pprof v0.0.0-20210407192527-94a9f03dee38/go.mod h1:kpwsk12EmLew5upagYY7GY0pfYCcupk39gWOCRROcvE= github.com/google/pprof v0.0.0-20260402051712-545e8a4df936 h1:EwtI+Al+DeppwYX2oXJCETMO23COyaKGP6fHVpkpWpg= github.com/google/pprof v0.0.0-20260402051712-545e8a4df936/go.mod h1:MxpfABSjhmINe3F1It9d+8exIHFvUqtLIRCdOGNXqiI= github.com/google/uuid v1.6.0 h1:NIvaJDMOsjHA8n1jAhLSgzrAzy1Hgr+hNrb57e+94F0= @@ -78,6 +121,8 @@ github.com/grpc-ecosystem/grpc-gateway/v2 v2.27.7 h1:X+2YciYSxvMQK0UZ7sg45ZVabVZ github.com/grpc-ecosystem/grpc-gateway/v2 v2.27.7/go.mod h1:lW34nIZuQ8UDPdkon5fmfp2l3+ZkQ2me/+oecHYLOII= github.com/hashicorp/yamux v0.1.2 h1:XtB8kyFOyHXYVFnwT5C3+Bdo8gArse7j2AQ0DA0Uey8= github.com/hashicorp/yamux v0.1.2/go.mod h1:C+zze2n6e/7wshOZep2A70/aQU6QBRWJO/G6FT1wIns= +github.com/hpcloud/tail v1.0.0/go.mod h1:ab1qPbhIpdTxEkNHXyeSf5vhxWSCs/tWer42PpOxQnU= +github.com/ianlancetaylor/demangle v0.0.0-20200824232613-28f6c0f3b639/go.mod h1:aSSvb/t6k1mPoxDqO4vJh6VOCGPwU4O0C2/Eqndh1Sc= github.com/inconshreveable/mousetrap v1.1.0 h1:wN+x4NVGpMsO7ErUn/mUI3vEoE6Jt13X2s0bqwp9tc8= github.com/inconshreveable/mousetrap v1.1.0/go.mod h1:vpF70FUmC8bwa3OWnCshd2FqLfsEA9PFc4w1p2J65bw= github.com/josharian/intern v1.0.0 h1:vlS4z54oSdjm0bgjRigI+G1HpF+tI+9rE5LLzOg8HmY= @@ -86,10 +131,14 @@ github.com/joshdk/go-junit v1.0.0 h1:S86cUKIdwBHWwA6xCmFlf3RTLfVXYQfvanM5Uh+K6GE github.com/joshdk/go-junit v1.0.0/go.mod h1:TiiV0PqkaNfFXjEiyjWM3XXrhVyCa1K4Zfga6W52ung= github.com/json-iterator/go v1.1.12 h1:PV8peI4a0ysnczrg+LtxykD8LfKY9ML6u2jnxaEnrnM= github.com/json-iterator/go v1.1.12/go.mod h1:e30LSqwooZae/UwlEbR2852Gd8hjQvJoHmT4TnhNGBo= +github.com/juju/gnuflag v0.0.0-20171113085948-2ce1bb71843d/go.mod h1:2PavIy+JPciBPrBUjwbNvtwB6RQlve+hkpll6QSNmOE= github.com/klauspost/compress v1.18.0 h1:c/Cqfb0r+Yi+JtIEq73FWXVkRonBlf0CRNYc8Zttxdo= github.com/klauspost/compress v1.18.0/go.mod h1:2Pp+KzxcywXVXMr50+X0Q/Lsb43OQHYWRCY2AiWywWQ= +github.com/kr/pretty v0.1.0/go.mod h1:dAy3ld7l9f0ibDNOQOHHMYYIIbhfbHSm3C4ZsoJORNo= github.com/kr/pretty v0.3.1 h1:flRD4NNwYAUpkphVc1HcthR4KEIFJ65n8Mw5qdRn3LE= github.com/kr/pretty v0.3.1/go.mod h1:hoEshYVHaxMs3cyo3Yncou5ZscifuDolrwPKZanG3xk= +github.com/kr/pty v1.1.1/go.mod h1:pFQYn66WHrOpPYNljwOMqo10TkYh1fy3cYio2l3bCsQ= +github.com/kr/text v0.1.0/go.mod h1:4Jbv+DJW3UT/LiOwJeYQe1efqtUx/iVham/4vfdArNI= github.com/kr/text v0.2.0 h1:5Nx0Ya0ZqY2ygV366QzturHI13Jq95ApcVaJBhpS+AY= github.com/kr/text v0.2.0/go.mod h1:eLer722TekiGuMkidMxC/pM04lWEeraHUUmBw8l2grE= github.com/kubernetes-csi/csi-lib-utils v0.24.0 h1:hpL5ecxtr07/DNIF9Qn/gbNG/ZlgMSMZxfRtgfVm9tY= @@ -98,6 +147,8 @@ github.com/kubernetes-csi/csi-test/v5 v5.5.0 h1:21NYP33XXfzsAGwFuFHJUIf60hY08B4A github.com/kubernetes-csi/csi-test/v5 v5.5.0/go.mod h1:5ZyneETi47SniZuPA9e8fIL6TTkkKv8/+jkaF0IHqKY= github.com/kylelemons/godebug v1.1.0 h1:RPNrshWIDI6G2gRW9EHilWtl7Z6Sb1BR0xunSBf0SNc= github.com/kylelemons/godebug v1.1.0/go.mod h1:9/0rRGxNHcop5bhtWyNeEfOS8JIWk580+fNqagV/RAw= +github.com/leodido/go-urn v1.4.0 h1:WT9HwE9SGECu3lg4d/dIA+jxlljEa1/ffXKmRjqdmIQ= +github.com/leodido/go-urn v1.4.0/go.mod h1:bvxc+MVxLKB4z00jd1z+Dvzr47oO32F/QSNjSBOlFxI= github.com/mailru/easyjson v0.9.0 h1:PrnmzHw7262yW8sTBwxi1PdJA3Iw/EKBa8psRf7d9a4= github.com/mailru/easyjson v0.9.0/go.mod h1:1+xMtQp2MRNVL/V1bOzuP3aP8VNwRW55fQUto+XFtTU= github.com/maruel/natural v1.1.1 h1:Hja7XhhmvEFhcByqDoHz9QZbkWey+COd9xWfCfn1ioo= @@ -116,8 +167,32 @@ github.com/modern-go/reflect2 v1.0.3-0.20250322232337-35a7c28c31ee h1:W5t00kpgFd github.com/modern-go/reflect2 v1.0.3-0.20250322232337-35a7c28c31ee/go.mod h1:yWuevngMOJpCy52FWWMvUC8ws7m/LJsjYzDa0/r8luk= github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822 h1:C3w9PqII01/Oq1c1nUAm88MOHcQC9l5mIlSMApZMrHA= github.com/munnerz/goautoneg v0.0.0-20191010083416-a7dc8b61c822/go.mod h1:+n7T8mK8HuQTcFwEeznm/DIxMOiR9yIdICNftLE1DvQ= +github.com/nxadm/tail v1.4.4/go.mod h1:kenIhsEOeOJmVchQTgglprH7qJGnHDVpk1VPCcaMI8A= +github.com/nxadm/tail v1.4.8 h1:nPr65rt6Y5JFSKQO7qToXr7pePgD6Gwiw05lkbyAQTE= +github.com/nxadm/tail v1.4.8/go.mod h1:+ncqLTQzXmGhMZNUePPaPqPvBxHAIsmXswZKocGu+AU= +github.com/oapi-codegen/nullable v1.1.0 h1:eAh8JVc5430VtYVnq00Hrbpag9PFRGWLjxR1/3KntMs= +github.com/oapi-codegen/nullable v1.1.0/go.mod h1:KUZ3vUzkmEKY90ksAmit2+5juDIhIZhfDl+0PwOQlFY= +github.com/oapi-codegen/oapi-codegen/v2 v2.8.0 h1:s4hxMxuqtR8jPzXkBTtFwY/SBuj3gEAYikmbBSdtLMM= +github.com/oapi-codegen/oapi-codegen/v2 v2.8.0/go.mod h1:yae2TI9IYB5vxQ35gFrpXh9L5H1eJv4MAUK1jumGMTo= +github.com/oapi-codegen/runtime v1.6.0 h1:7Xx+GlueD6nRuyKoCPzL434Jfi3BetbiJOrzCHp/VPU= +github.com/oapi-codegen/runtime v1.6.0/go.mod h1:GwV7hC2hviaMzj+ITfHVRESK5J2W/GefVwIND/bMGvU= +github.com/oasdiff/yaml v0.1.1 h1:6nHx+pn9gBRM6YpBlFZFQGCCd1nuvqOBtTD3KKTgGxY= +github.com/oasdiff/yaml v0.1.1/go.mod h1:EYJNoyktvWMJ0Hmhx+6qTaqMOsalUaRGT8Sj1hNcegU= +github.com/oasdiff/yaml3 v0.0.14 h1:aLJee3hxBK2H5wdXd9iPcIXb93Nty1Ge0pT171eHtkw= +github.com/oasdiff/yaml3 v0.0.14/go.mod h1:csto2xfDjYccdUn/yw/bPjj/cYTdp6HtFA0J4TWG+gg= +github.com/onsi/ginkgo v1.6.0/go.mod h1:lLunBs/Ym6LB5Z9jYTR76FiuTmxDTDusOGeTQH+WWjE= +github.com/onsi/ginkgo v1.10.2/go.mod h1:lLunBs/Ym6LB5Z9jYTR76FiuTmxDTDusOGeTQH+WWjE= +github.com/onsi/ginkgo v1.12.1/go.mod h1:zj2OWP4+oCPe1qIXoGWkgMRwljMUYCdkwsT2108oapk= +github.com/onsi/ginkgo v1.16.4 h1:29JGrr5oVBm5ulCWet69zQkzWipVXIol6ygQUe/EzNc= +github.com/onsi/ginkgo v1.16.4/go.mod h1:dX+/inL/fNMqNlz0e9LfyB9TswhZpCVdJM/Z6Vvnwo0= +github.com/onsi/ginkgo/v2 v2.1.3/go.mod h1:vw5CSIxN1JObi/U8gcbwft7ZxR2dgaR70JSE3/PpL4c= github.com/onsi/ginkgo/v2 v2.32.0 h1:Hw7s2pVrQo/8Yz5N77qdnpHaoc+c6cC9WIV1Jce+J6E= github.com/onsi/ginkgo/v2 v2.32.0/go.mod h1:+aXOY+vzZ5mu2iI2HpTZUPmM//oQfsNFX6gU9kNcA44= +github.com/onsi/gomega v1.7.0/go.mod h1:ex+gbHU/CVuBBDIJjb2X0qEXbFg53c61hWP/1CpauHY= +github.com/onsi/gomega v1.7.1/go.mod h1:XdKZgCCFLUoM/7CFJVPcG8C1xQ1AJ0vpAezJrB7JYyY= +github.com/onsi/gomega v1.10.1/go.mod h1:iN09h71vgCQne3DLsj+A5owkum+a2tYe+TOCB1ybHNo= +github.com/onsi/gomega v1.17.0/go.mod h1:HnhC7FXeEQY45zxNK3PPoIUhzk/80Xly9PcubAlGdZY= +github.com/onsi/gomega v1.19.0/go.mod h1:LY+I3pBVzYsTBU1AnDwOSxaYi9WoWiqgwooUqq9yPro= github.com/onsi/gomega v1.42.1 h1:iN1rCUX+44NZ1Dc97MPoeFYbFR0vh8zxoxMFwKdyZ6I= github.com/onsi/gomega v1.42.1/go.mod h1:REff/hsDsodHoKlWsP2mAPhu1+5/6hVYNf9rIEBpeSg= github.com/opencontainers/go-digest v1.0.0 h1:apOUWs51W5PlhuyGyz9FCeeBIOUDA/6nW8Oi/yOhh5U= @@ -138,10 +213,19 @@ github.com/robfig/cron/v3 v3.0.1/go.mod h1:eQICP3HwyT7UooqI/z+Ov+PtYAWygg1TEWWzG github.com/rogpeppe/go-internal v1.14.1 h1:UQB4HGPB6osV0SQTLymcB4TgvyWu6ZyliaW0tI/otEQ= github.com/rogpeppe/go-internal v1.14.1/go.mod h1:MaRKkUm5W0goXpeCfT7UZI6fk/L7L7so1lCWt35ZSgc= github.com/russross/blackfriday/v2 v2.1.0/go.mod h1:+Rmxgy9KzJVeS9/2gXHxylqXiyQDYRxCVz55jmeOWTM= +github.com/santhosh-tekuri/jsonschema/v6 v6.0.2 h1:KRzFb2m7YtdldCEkzs6KqmJw4nqEVZGK7IN2kJkjTuQ= +github.com/santhosh-tekuri/jsonschema/v6 v6.0.2/go.mod h1:JXeL+ps8p7/KNMjDQk3TCwPpBy0wYklyWTfbkIzdIFU= +github.com/sergi/go-diff v1.1.0 h1:we8PVUC3FE2uYfodKH/nBHMSetSfHDR6scGdBi+erh0= +github.com/sergi/go-diff v1.1.0/go.mod h1:STckp+ISIX8hZLjrqAeVduY0gWCT9IjLuqbuNXdaHfM= +github.com/speakeasy-api/jsonpath v0.6.3 h1:c+QPwzAOdrWvzycuc9HFsIZcxKIaWcNpC+xhOW9rJxU= +github.com/speakeasy-api/jsonpath v0.6.3/go.mod h1:2cXloNuQ+RSXi5HTRaeBh7JEmjRXTiaKpFTdZiL7URI= +github.com/speakeasy-api/openapi v1.24.0 h1:opoD27rupX7zBVPq1HkIGLeMOzNNA7JalhYP8q34i04= +github.com/speakeasy-api/openapi v1.24.0/go.mod h1:g3+dIMe0AYgbbGvnlQZqesmjAVWSm9BmsjLevnefQrg= github.com/spf13/cobra v1.10.2 h1:DMTTonx5m65Ic0GOoRY2c16WCbHxOOw6xxezuLaBpcU= github.com/spf13/cobra v1.10.2/go.mod h1:7C1pvHqHw5A4vrJfjNwvOdzYu0Gml16OCs2GRiTUUS4= github.com/spf13/pflag v1.0.9 h1:9exaQaMOCwffKiiiYk6/BndUBv+iRViNW+4lEMi0PvY= github.com/spf13/pflag v1.0.9/go.mod h1:McXfInJRrz4CZXVZOBLb0bTZqETkiAhM9Iw0y3An2Bg= +github.com/spkg/bom v0.0.0-20160624110644-59b7046e48ad/go.mod h1:qLr4V1qq6nMqFKkMo8ZTx3f+BZEkzsRUY10Xsm2mwU0= github.com/stoewer/go-strcase v1.3.0 h1:g0eASXYtp+yvN9fK8sH94oCIk0fau9uV1/ZdJ0AVEzs= github.com/stoewer/go-strcase v1.3.0/go.mod h1:fAH5hQ5pehh+j3nZfvwdk2RgEgQjAoM8wodgtPmh1xo= github.com/stretchr/objx v0.1.0/go.mod h1:HFkY916IF+rwdDfMAkV7OtwuqBVzrE8GR6GFx+wExME= @@ -150,6 +234,8 @@ github.com/stretchr/objx v0.5.0/go.mod h1:Yh+to48EsGEfYuaHDzXPcE3xhTkx73EhmCGUpE github.com/stretchr/objx v0.5.2 h1:xuMeJ0Sdp5ZMRXx/aWO6RZxdr3beISkG5/G/aIRr3pY= github.com/stretchr/objx v0.5.2/go.mod h1:FRsXN1f5AsAjCGJKqEizvkpNtU+EGNCLh3NxZ/8L+MA= github.com/stretchr/testify v1.3.0/go.mod h1:M5WIy9Dh21IEIfnGCwXGc5bZfKNJtfHm1UVUgZn+9EI= +github.com/stretchr/testify v1.4.0/go.mod h1:j7eGeouHqKxXV5pUuKE4zz7dFj8WfuZ+81PSLYec5m4= +github.com/stretchr/testify v1.5.1/go.mod h1:5W2xD1RspED5o8YsWQXVCued0rvSQ+mT+I5cxcmMvtA= github.com/stretchr/testify v1.7.1/go.mod h1:6Fq8oRcR53rry900zMqJjRRixrwX3KX962/h/Wwjteg= github.com/stretchr/testify v1.8.0/go.mod h1:yNjHg4UonilssWZ8iaSj1OCr/vHnekPRkoO+kdMU+MU= github.com/stretchr/testify v1.8.1/go.mod h1:w2LPCIKwWwSfY2zedu0+kehJoqGctiVI29o6fzry7u4= @@ -163,8 +249,11 @@ github.com/tidwall/pretty v1.2.1 h1:qjsOFOWWQl+N3RsoF5/ssm1pHmJJwhjlSbZ51I6wMl4= github.com/tidwall/pretty v1.2.1/go.mod h1:ITEVvHYasfjBbM0u2Pg8T2nJnzm8xPwvNhhsoaGGjNU= github.com/tidwall/sjson v1.2.5 h1:kLy8mja+1c9jlljvWTlSazM7cKDRfJuR/bOJhcY5NcY= github.com/tidwall/sjson v1.2.5/go.mod h1:Fvgq9kS/6ociJEDnK0Fk1cpYF4FIW6ZF7LAe+6jwd28= +github.com/vmware-labs/yaml-jsonpath v0.3.2 h1:/5QKeCBGdsInyDCyVNLbXyilb61MXGi9NP674f9Hobk= +github.com/vmware-labs/yaml-jsonpath v0.3.2/go.mod h1:U6whw1z03QyqgWdgXxvVnQ90zN1BWz5V+51Ewf8k+rQ= github.com/x448/float16 v0.8.4 h1:qLwI1I70+NjRFUR3zs1JPUCgaCXSh3SW62uAKT1mSBM= github.com/x448/float16 v0.8.4/go.mod h1:14CWIYCyZA/cWjXOioeEpHeN/83MdbZDRQHoFcYsOfg= +github.com/yuin/goldmark v1.2.1/go.mod h1:3hX8gzYuyVAZsxl0MRgGTJEmQBFcNTphYh9decYSb74= go.opentelemetry.io/auto/sdk v1.2.1 h1:jXsnJ4Lmnqd11kwkBV2LgLoFMZKizbCi5fNZ/ipaZ64= go.opentelemetry.io/auto/sdk v1.2.1/go.mod h1:KRTj+aOaElaLi+wW1kO/DZRXwkF4C5xPbEe3ZiIhN7Y= go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp v0.65.0 h1:7iP2uCb7sGddAr30RRS6xjKy7AZ2JtTOPA3oolgVSw8= @@ -197,26 +286,69 @@ go.yaml.in/yaml/v2 v2.4.3 h1:6gvOSjQoTB3vt1l+CU+tSyi/HOjfOjRLJ4YwYZGwRO0= go.yaml.in/yaml/v2 v2.4.3/go.mod h1:zSxWcmIDjOzPXpjlTTbAsKokqkDNAVtZO0WOMiT90s8= go.yaml.in/yaml/v3 v3.0.4 h1:tfq32ie2Jv2UxXFdLJdh3jXuOzWiL1fo0bu/FbuKpbc= go.yaml.in/yaml/v3 v3.0.4/go.mod h1:DhzuOOF2ATzADvBadXxruRBLzYTpT36CKvDb3+aBEFg= +golang.org/x/crypto v0.0.0-20190308221718-c2843e01d9a2/go.mod h1:djNgcEr1/C05ACkg1iLfiJU5Ep61QUkGW8qpdssI0+w= +golang.org/x/crypto v0.0.0-20191011191535-87dc89f01550/go.mod h1:yigFU9vqHzYiE8UmvKecakEJjdnWj3jj499lnFckfCI= +golang.org/x/crypto v0.0.0-20200622213623-75b288015ac9/go.mod h1:LzIPMQfyMNhhGPhUkYOs5KpL4U8rLKemX1yGLhDgUto= +golang.org/x/crypto v0.56.0 h1:GUh5Ii4J5jtcseSMiRqr1jXCNHoxjeV9Fmekc2oLy6Y= +golang.org/x/crypto v0.56.0/go.mod h1:OMW5y6CY9l38uPLmxU6l6pwcXp1obtLo3e6gT7gQR2I= golang.org/x/exp v0.0.0-20251219203646-944ab1f22d93 h1:fQsdNF2N+/YewlRZiricy4P1iimyPKZ/xwniHj8Q2a0= golang.org/x/exp v0.0.0-20251219203646-944ab1f22d93/go.mod h1:EPRbTFwzwjXj9NpYyyrvenVh9Y+GFeEvMNh7Xuz7xgU= +golang.org/x/mod v0.3.0/go.mod h1:s0Qsj1ACt9ePp/hMypM3fl4fZqREWJwdYDEqhRiZZUA= golang.org/x/mod v0.38.0 h1:MECBjubtXD7yj4HrhIUcywNaGeNVUdfVnxmPajOk4yk= golang.org/x/mod v0.38.0/go.mod h1:V6Xz0pq8TQ3dGqVQ1FVHuelZpAL0uNhSkk9ogYP3c40= +golang.org/x/net v0.0.0-20180906233101-161cd47e91fd/go.mod h1:mL1N/T3taQHkDXs73rZJwtUhF3w3ftmwwsq0BUmARs4= +golang.org/x/net v0.0.0-20190404232315-eb5bcb51f2a3/go.mod h1:t9HGtf8HONx5eT2rtn7q6eTqICYqUVnKs3thJo3Qplg= +golang.org/x/net v0.0.0-20190620200207-3b0461eec859/go.mod h1:z5CRVTTTmAJ677TzLLGU+0bjPO0LkuOLi4/5GtJWs/s= +golang.org/x/net v0.0.0-20200520004742-59133d7f0dd7/go.mod h1:qpuaurCH72eLCgpAm/N6yyVIVM9cpaDIP3A8BGJEC5A= +golang.org/x/net v0.0.0-20201021035429-f5854403a974/go.mod h1:sp8m0HH+o8qH0wwXwYZr8TS3Oi6o0r6Gce1SSxlDquU= +golang.org/x/net v0.0.0-20210428140749-89ef3d95e781/go.mod h1:OJAsFXCWl8Ukc7SiCT/9KSuxbyM7479/AVlXFRxuMCk= +golang.org/x/net v0.0.0-20220225172249-27dd8689420f/go.mod h1:CfG3xpIq0wQ8r1q4Su4UZFWDARRcnwPjda9FqA0JpMk= golang.org/x/net v0.58.0 h1:ynWG7rqYi4ccpTEuPZ2QGWHktVEM9DMCj9yzDE0Q7To= golang.org/x/net v0.58.0/go.mod h1:YwCddHnFlT7eLQqVprV19OnhLGtc5xOKgE0RyqgfWAU= golang.org/x/oauth2 v0.36.0 h1:peZ/1z27fi9hUOFCAZaHyrpWG5lwe0RJEEEeH0ThlIs= golang.org/x/oauth2 v0.36.0/go.mod h1:YDBUJMTkDnJS+A4BP4eZBjCqtokkg1hODuPjwiGPO7Q= +golang.org/x/sync v0.0.0-20180314180146-1d60e4601c6f/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM= +golang.org/x/sync v0.0.0-20190423024810-112230192c58/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM= +golang.org/x/sync v0.0.0-20201020160332-67f06af15bc9/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM= golang.org/x/sync v0.22.0 h1:SZjpbeLmrCk4xhRSZFNZW5gFUeCeFgjekvI/+gfScek= golang.org/x/sync v0.22.0/go.mod h1:9xrNwdLfx4jkKbNva9FpL6vEN7evnE43NNNJQ2LF3+0= +golang.org/x/sys v0.0.0-20180909124046-d0be0721c37e/go.mod h1:STP8DvDyc/dI5b8T5hshtkjS+E42TnysNCUPdjciGhY= +golang.org/x/sys v0.0.0-20190215142949-d0b11bdaac8a/go.mod h1:STP8DvDyc/dI5b8T5hshtkjS+E42TnysNCUPdjciGhY= +golang.org/x/sys v0.0.0-20190412213103-97732733099d/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs= +golang.org/x/sys v0.0.0-20190904154756-749cb33beabd/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs= +golang.org/x/sys v0.0.0-20191005200804-aed5e4c7ecf9/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs= +golang.org/x/sys v0.0.0-20191120155948-bd437916bb0e/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs= +golang.org/x/sys v0.0.0-20191204072324-ce4227a45e2e/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs= +golang.org/x/sys v0.0.0-20200323222414-85ca7c5b95cd/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs= +golang.org/x/sys v0.0.0-20200930185726-fdedc70b468f/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs= +golang.org/x/sys v0.0.0-20201119102817-f84b799fce68/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs= +golang.org/x/sys v0.0.0-20210112080510-489259a85091/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs= +golang.org/x/sys v0.0.0-20210423082822-04245dca01da/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs= +golang.org/x/sys v0.0.0-20210615035016-665e8c7367d1/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg= +golang.org/x/sys v0.0.0-20211216021012-1d35b9e2eb4e/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg= golang.org/x/sys v0.47.0 h1:o7XGOvZQCADBQQ4Y7VNq2dRWQR7JmOUW8Kxx4ZsNgWs= golang.org/x/sys v0.47.0/go.mod h1:4GL1E5IUh+htKOUEOaiffhrAeqysfVGipDYzABqnCmw= +golang.org/x/term v0.0.0-20201126162022-7de9c90e9dd1/go.mod h1:bj7SfCRtBDWHUb9snDiAeCFNEtKQo2Wmx5Cou7ajbmo= +golang.org/x/term v0.0.0-20210927222741-03fcf44c2211/go.mod h1:jbD1KX2456YbFQfuXm/mYQcufACuNUgVhRMnK/tPxf8= golang.org/x/term v0.45.0 h1:NwWyBmoJCbfTHpxrWoZ9C6/VxOf7ic219I8xZZFdrf0= golang.org/x/term v0.45.0/go.mod h1:9aqxs0blBcrm/n0L9QW0aRVD+ktan8ssZromtqJC43w= +golang.org/x/text v0.3.0/go.mod h1:NqM8EUOU14njkJ3fqMW+pc6Ldnwhi/IjpwHt7yyuwOQ= +golang.org/x/text v0.3.3/go.mod h1:5Zoc/QRtKVWzQhOtBMvqHzDpF6irO9z98xDceosuGiQ= +golang.org/x/text v0.3.6/go.mod h1:5Zoc/QRtKVWzQhOtBMvqHzDpF6irO9z98xDceosuGiQ= +golang.org/x/text v0.3.7/go.mod h1:u+2+/6zg+i71rQMx5EYifcz6MCKuco9NR6JIITiCfzQ= golang.org/x/text v0.41.0 h1:vz/seA0lnX87Othu2f/0L24RcgrXD9/YFTSuGjj3rH8= golang.org/x/text v0.41.0/go.mod h1:jvf1O8ajNzZqhSrQBPbutR/EB83Cc0CFrezNQIwbb5M= golang.org/x/time v0.14.0 h1:MRx4UaLrDotUKUdCIqzPC48t1Y9hANFKIRpNx+Te8PI= golang.org/x/time v0.14.0/go.mod h1:eL/Oa2bBBK0TkX57Fyni+NgnyQQN4LitPmob2Hjnqw4= +golang.org/x/tools v0.0.0-20180917221912-90fa682c2a6e/go.mod h1:n7NCudcB/nEzxVGmLbDWY5pfWTLqBcC2KZ6jyYvM4mQ= +golang.org/x/tools v0.0.0-20191119224855-298f0cb1881e/go.mod h1:b+2E5dAYhXwXZwtnZ6UAqBI28+e2cm9otk0dWdXHAEo= +golang.org/x/tools v0.0.0-20201224043029-2b0845dc783e/go.mod h1:emZCQorbCU4vsT4fOWvOPXz4eW1wZW4PmDk9uLelYpA= golang.org/x/tools v0.48.0 h1:3+hClM1aLL5mjMKm5ovokw9epgRXPuu2tILgismM6RE= golang.org/x/tools v0.48.0/go.mod h1:08xX0orndb/F7jJxGDicx061tyd5pcMto75YMAXr6lk= +golang.org/x/xerrors v0.0.0-20190717185122-a985d3407aa7/go.mod h1:I/5z698sn9Ka8TeJc9MKroUUfqBBauWjQqLJ2OPfmY0= +golang.org/x/xerrors v0.0.0-20191011141410-1b5146add898/go.mod h1:I/5z698sn9Ka8TeJc9MKroUUfqBBauWjQqLJ2OPfmY0= +golang.org/x/xerrors v0.0.0-20191204190536-9bdfabe68543/go.mod h1:I/5z698sn9Ka8TeJc9MKroUUfqBBauWjQqLJ2OPfmY0= +golang.org/x/xerrors v0.0.0-20200804184101-5ec99f83aff1/go.mod h1:I/5z698sn9Ka8TeJc9MKroUUfqBBauWjQqLJ2OPfmY0= gonum.org/v1/gonum v0.17.0 h1:VbpOemQlsSMrYmn7T2OUvQ4dqxQXU+ouZFQsZOx50z4= gonum.org/v1/gonum v0.17.0/go.mod h1:El3tOrEuMpv2UdMrbNlKEh9vd86bmQ6vqIcDwxEOc1E= google.golang.org/genproto/googleapis/api v0.0.0-20260526163538-3dc84a4a5aaa h1:Kjn0N0tCrDgiAFW+lGO4JZ3ck44CehvJQMAwj9QF0G8= @@ -225,17 +357,34 @@ google.golang.org/genproto/googleapis/rpc v0.0.0-20260526163538-3dc84a4a5aaa h1: google.golang.org/genproto/googleapis/rpc v0.0.0-20260526163538-3dc84a4a5aaa/go.mod h1:4Hqkh8ycfw05ld/3BWL7rJOSfebL2Q+DVDeRgYgxUU8= google.golang.org/grpc v1.83.2 h1:EManeRomTObA0BU7I8vXgg/78uE5MJ9M8B39EX2WscU= google.golang.org/grpc v1.83.2/go.mod h1:YPI1hK3kDked6iHvgX3tR0y+nX/qpMFKhPgFsokw1S8= +google.golang.org/protobuf v0.0.0-20200109180630-ec00e32a8dfd/go.mod h1:DFci5gLYBciE7Vtevhsrf46CRTquxDuWsQurQQe4oz8= +google.golang.org/protobuf v0.0.0-20200221191635-4d8936d0db64/go.mod h1:kwYJMbMJ01Woi6D6+Kah6886xMZcty6N08ah7+eCXa0= +google.golang.org/protobuf v0.0.0-20200228230310-ab0ca4ff8a60/go.mod h1:cfTl7dwQJ+fmap5saPgwCLgHXTUD7jkjRqWcaiX5VyM= +google.golang.org/protobuf v1.20.1-0.20200309200217-e05f789c0967/go.mod h1:A+miEFZTKqfCUM6K7xSMQL9OKL/b6hQv+e19PK+JZNE= +google.golang.org/protobuf v1.21.0/go.mod h1:47Nbq4nVaFHyn7ilMalzfO3qCViNmqZ2kzikPIcrTAo= +google.golang.org/protobuf v1.23.0/go.mod h1:EGpADcykh3NcUnDUJcl1+ZksZNG86OlYog2l/sGQquU= +google.golang.org/protobuf v1.26.0-rc.1/go.mod h1:jlhhOSvTdKEhbULTjvd4ARK9grFBp09yW+WbY/TyQbw= +google.golang.org/protobuf v1.26.0/go.mod h1:9q0QmTI4eRPtz6boOQmLYwt+qCgq0jsYwAQnmE0givc= google.golang.org/protobuf v1.36.12-0.20260120151049-f2248ac996af h1:+5/Sw3GsDNlEmu7TfklWKPdQ0Ykja5VEmq2i817+jbI= google.golang.org/protobuf v1.36.12-0.20260120151049-f2248ac996af/go.mod h1:HTf+CrKn2C3g5S8VImy6tdcUvCska2kB7j23XfzDpco= gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0= +gopkg.in/check.v1 v1.0.0-20190902080502-41f04d3bba15/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0= gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c h1:Hei/4ADfdWqJk1ZMxUNpqntNwaWcugrBjAiHlqqRiVk= gopkg.in/check.v1 v1.0.0-20201130134442-10cb98267c6c/go.mod h1:JHkPIbrfpd72SG/EVd6muEfDQjcINNoR0C8j2r3qZ4Q= gopkg.in/evanphx/json-patch.v4 v4.13.0 h1:czT3CmqEaQ1aanPc5SdlgQrrEIb8w/wwCvWWnfEbYzo= gopkg.in/evanphx/json-patch.v4 v4.13.0/go.mod h1:p8EYWUEYMpynmqDbY58zCKCFZw8pRWMG4EsWvDvM72M= +gopkg.in/fsnotify.v1 v1.4.7/go.mod h1:Tz8NjZHkW78fSQdbUxIjBTcgA1z1m8ZHf0WmKUhAMys= gopkg.in/inf.v0 v0.9.1 h1:73M5CoZyi3ZLMOyDlQh031Cx6N9NDJ2Vvfl76EDAgDc= gopkg.in/inf.v0 v0.9.1/go.mod h1:cWUDdTG/fYaXco+Dcufb5Vnc6Gp2YChqWtbxRZE0mXw= +gopkg.in/tomb.v1 v1.0.0-20141024135613-dd632973f1e7 h1:uRGJdciOHaEIrze2W8Q3AKkepLTh2hOroT7a+7czfdQ= +gopkg.in/tomb.v1 v1.0.0-20141024135613-dd632973f1e7/go.mod h1:dt/ZhP58zS4L8KSrWDmTeBkI65Dw0HsyUHuEVlX15mw= +gopkg.in/yaml.v2 v2.2.1/go.mod h1:hI93XBmqTisBFMUTm0b8Fm+jr3Dg1NNxqwp+5A1VGuI= +gopkg.in/yaml.v2 v2.2.2/go.mod h1:hI93XBmqTisBFMUTm0b8Fm+jr3Dg1NNxqwp+5A1VGuI= +gopkg.in/yaml.v2 v2.2.4/go.mod h1:hI93XBmqTisBFMUTm0b8Fm+jr3Dg1NNxqwp+5A1VGuI= +gopkg.in/yaml.v2 v2.3.0/go.mod h1:hI93XBmqTisBFMUTm0b8Fm+jr3Dg1NNxqwp+5A1VGuI= gopkg.in/yaml.v2 v2.4.0 h1:D8xgwECY7CYvx+Y2n4sBz93Jn9JRvxdiyyo8CTfuKaY= gopkg.in/yaml.v2 v2.4.0/go.mod h1:RDklbk79AGWmwhnvt/jBztapEOGDOx6ZbXqjP6csGnQ= +gopkg.in/yaml.v3 v3.0.0-20191026110619-0b21df46bc1d/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM= gopkg.in/yaml.v3 v3.0.0-20200313102051-9f266ea9e77c/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM= gopkg.in/yaml.v3 v3.0.1 h1:fxVm/GzAzEWqLHuvctI91KS9hhNmmWOoWu0XTYJS7CA= gopkg.in/yaml.v3 v3.0.1/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM= diff --git a/csi-driver/internal/clusters/clusters.go b/csi-driver/internal/clusters/clusters.go index 49276c17b..a71616866 100644 --- a/csi-driver/internal/clusters/clusters.go +++ b/csi-driver/internal/clusters/clusters.go @@ -15,6 +15,7 @@ import ( "os" "strings" + atlascp "github.com/simplyblock/atlas/controlplane" "github.com/simplyblock/atlas/errs/deferrers" "github.com/simplyblock/atlas/lvol" "k8s.io/klog" @@ -102,35 +103,11 @@ func List() ([]string, error) { // pool. poolIDOrName may be a pool UUID (used as-is), a pool name (resolved via // the API), or empty (no pool context, so only cluster-level operations work). func Client(ctx context.Context, clusterID, poolIDOrName string) (*controlplane.ClusterClient, error) { - clusters, err := Load() + clusterConfig, credential, err := resolve(clusterID) if err != nil { return nil, err } - var clusterConfig *Config - for _, cluster := range clusters.Clusters { - if cluster.ClusterID == clusterID { - clusterConfig = &cluster - break - } - } - - if clusterConfig == nil { - return nil, fmt.Errorf("failed to find secret for clusterID %s: %w", clusterID, controlplane.ErrClusterNotFound) - } - - if clusterConfig.ClusterEndpoint == "" { - return nil, fmt.Errorf("invalid cluster configuration for clusterID %s: missing endpoint", clusterID) - } - - credential := credentialFor(clusterConfig) - if credential == "" { - return nil, fmt.Errorf( - "invalid cluster configuration for clusterID %s: no cluster_secret and no API token available", - clusterID, - ) - } - klog.Infof("Simplyblock client created for ClusterID:%s, Endpoint:%s", clusterConfig.ClusterID, clusterConfig.ClusterEndpoint, @@ -152,6 +129,54 @@ func Client(ctx context.Context, clusterID, poolIDOrName string) (*controlplane. return client, nil } +// ReplicationClient creates the generated, atlas-lib control-plane client for +// a cluster, sharing this package's secret.json resolution and credential +// precedence with Client. It is a sibling rather than a Client return-type +// change: Client's hand-rolled controlplane.ClusterClient is what every +// existing Volume/Snapshot/Clone RPC already depends on, and only the +// Replication service (csi-addons) is built against the generated client. +func ReplicationClient(_ context.Context, clusterID string) (*atlascp.Client, error) { + clusterConfig, credential, err := resolve(clusterID) + if err != nil { + return nil, err + } + return atlascp.New(atlascp.Config{Endpoint: clusterConfig.ClusterEndpoint, Token: credential}) +} + +// resolve looks up the named cluster's endpoint and credential in the driver's +// secret.json. It is the lookup Client and ReplicationClient share, so a +// cluster missing from the secret is invisible to neither or both, never one. +func resolve(clusterID string) (*Config, string, error) { + clusters, err := Load() + if err != nil { + return nil, "", err + } + + var clusterConfig *Config + for _, cluster := range clusters.Clusters { + if cluster.ClusterID == clusterID { + clusterConfig = &cluster + break + } + } + + if clusterConfig == nil { + return nil, "", fmt.Errorf("failed to find secret for clusterID %s: %w", clusterID, controlplane.ErrClusterNotFound) + } + if clusterConfig.ClusterEndpoint == "" { + return nil, "", fmt.Errorf("invalid cluster configuration for clusterID %s: missing endpoint", clusterID) + } + + credential := credentialFor(clusterConfig) + if credential == "" { + return nil, "", fmt.Errorf( + "invalid cluster configuration for clusterID %s: no cluster_secret and no API token available", + clusterID, + ) + } + return clusterConfig, credential, nil +} + // credentialFor returns the bearer credential for a cluster: the API token when // SPDKCSI_API_TOKEN_PATH is set and names a readable, non-empty file, otherwise // the cluster_secret from the secret entry. diff --git a/csi-driver/internal/csi/common/server.go b/csi-driver/internal/csi/common/server.go index 8ab1cde6b..411cc969c 100644 --- a/csi-driver/internal/csi/common/server.go +++ b/csi-driver/internal/csi/common/server.go @@ -13,7 +13,11 @@ import ( ) type NonBlockingGRPCServer interface { - Start(endpoint string, ids csi.IdentityServer, cs csi.ControllerServer, ns csi.NodeServer) + // register is called once per extra service (csi-addons Identity and + // Replication) with the same *grpc.Server the CSI services register on, + // so csicommon never has to import csi-addons: the caller builds the + // closures and this package only invokes them. + Start(endpoint string, ids csi.IdentityServer, cs csi.ControllerServer, ns csi.NodeServer, register ...func(*grpc.Server)) Wait() Stop() ForceStop() @@ -33,10 +37,11 @@ func (s *nonBlockingGRPCServer) Start( ids csi.IdentityServer, cs csi.ControllerServer, ns csi.NodeServer, + register ...func(*grpc.Server), ) { s.wg.Add(1) - go s.serve(endpoint, ids, cs, ns) + go s.serve(endpoint, ids, cs, ns, register) } func (s *nonBlockingGRPCServer) Wait() { @@ -56,6 +61,7 @@ func (s *nonBlockingGRPCServer) serve( ids csi.IdentityServer, cs csi.ControllerServer, ns csi.NodeServer, + register []func(*grpc.Server), ) { var err error @@ -112,6 +118,9 @@ func (s *nonBlockingGRPCServer) serve( if ns != nil { csi.RegisterNodeServer(server, ns) } + for _, r := range register { + r(server) + } klog.Infof("Listening for connections on address: %#v", listener.Addr()) diff --git a/csi-driver/internal/csi/controller/errorclass.go b/csi-driver/internal/csi/controller/errorclass.go index a5b9f0a8c..1e29731f6 100644 --- a/csi-driver/internal/csi/controller/errorclass.go +++ b/csi-driver/internal/csi/controller/errorclass.go @@ -8,6 +8,7 @@ import ( "google.golang.org/grpc/codes" + atlascp "github.com/simplyblock/atlas/controlplane" "github.com/simplyblock/csi-driver/internal/controlplane" ) @@ -100,6 +101,13 @@ func httpStatusOf(err error) int { if errors.As(err, &httpErr) { return httpErr.StatusCode } + // The Replication service talks to the control plane through the + // generated atlas-lib client instead, whose error type reports its + // status through a method rather than a public field. + var statusErr *atlascp.StatusError + if errors.As(err, &statusErr) { + return statusErr.HTTPStatus() + } return 0 } diff --git a/csi-driver/internal/csi/controller/errorclass_rpc.go b/csi-driver/internal/csi/controller/errorclass_rpc.go index 16dd58c2c..7cd9317e5 100644 --- a/csi-driver/internal/csi/controller/errorclass_rpc.go +++ b/csi-driver/internal/csi/controller/errorclass_rpc.go @@ -108,6 +108,15 @@ func classifyValidateVolumeCapabilitiesError(err error) classifiedError { func classifyListSnapshotsError(err error) classifiedError { return classifiedError{ListSnapshotsErrorClassifier.Classify(err), err} } +func classifyEnableVolumeReplicationError(err error) classifiedError { + return classifiedError{EnableVolumeReplicationErrorClassifier.Classify(err), err} +} +func classifyDisableVolumeReplicationError(err error) classifiedError { + return classifiedError{DisableVolumeReplicationErrorClassifier.Classify(err), err} +} +func classifyGetVolumeReplicationInfoError(err error) classifiedError { + return classifiedError{GetVolumeReplicationInfoErrorClassifier.Classify(err), err} +} // Dispositions reused across RPCs. var ( @@ -120,6 +129,10 @@ var ( // resolveConflict: a 409 must be resolved by looking up the existing object // (same source and params → return it as success, otherwise AlreadyExists). resolveConflict = controlPlaneErrorClass{Idempotent: true} + // cutoverInFlight: a 409 on a replication verb means a cutover is + // currently running, not a conflicting object to resolve → ABORTED, + // retryable, since Ramen re-drives every reconcile until it clears. + cutoverInFlight = controlPlaneErrorClass{Code: codes.Aborted, Retryable: true} ) // Preconfigured per-RPC classifiers. Every RPC that talks to the control plane @@ -154,4 +167,20 @@ var ( // ListSnapshotsErrorClassifier has no operation-specific statuses: every // status is handled generically. ListSnapshotsErrorClassifier = errorClassifier{} + + // EnableVolumeReplicationErrorClassifier: a different-policy attach is a + // 412 (design §10), already generic (FailedPrecondition); a 404 means the + // volume itself does not exist. + EnableVolumeReplicationErrorClassifier = errorClassifier{overrides: map[int]controlPlaneErrorClass{ + http.StatusNotFound: sourceNotFound, + }} + // DisableVolumeReplicationErrorClassifier: a 409 means a cutover is in + // flight (design §10), not a conflicting object to resolve. + DisableVolumeReplicationErrorClassifier = errorClassifier{overrides: map[int]controlPlaneErrorClass{ + http.StatusNotFound: sourceNotFound, + http.StatusConflict: cutoverInFlight, + }} + GetVolumeReplicationInfoErrorClassifier = errorClassifier{overrides: map[int]controlPlaneErrorClass{ + http.StatusNotFound: sourceNotFound, + }} ) diff --git a/csi-driver/internal/csi/controller/mock_controlplane_test.go b/csi-driver/internal/csi/controller/mock_controlplane_test.go index a33407973..fa011abb7 100644 --- a/csi-driver/internal/csi/controller/mock_controlplane_test.go +++ b/csi-driver/internal/csi/controller/mock_controlplane_test.go @@ -27,6 +27,12 @@ type mockVolume struct { Size int64 Status string // defaults to "online" when empty GroupID string // consistency group id, "" for a non-member + + // ReplicationPolicyID is the policy this volume currently follows, "" + // when none. Set by a PUT carrying replication_policy_id (a string + // attaches, an explicit JSON null detaches; the key's absence, as an + // ordinary resize PUT sends, leaves it untouched). + ReplicationPolicyID string } // status returns the volume's reported status, defaulting to `online`. @@ -103,13 +109,27 @@ type mockSBCLI struct { // It lets a test drive an RPC through every control-plane response and assert // the resulting gRPC code. injectStatus func(r *http.Request) int + + // replicationStatus, keyed by volume id, is the raw JSON body GET + // .../replication/status serves for that volume. A test sets it directly + // rather than the mock deriving it, since the derivation itself + // (get_replication_info) is sbcli's, already covered there; this mock + // only has to prove the driver maps whatever shape the endpoint returns. + replicationStatus map[string]map[string]any + + // replicationPUTStatus, when set, makes every PUT carrying + // replication_policy_id respond with this HTTP status instead of the + // normal idempotent update, modeling a backend refusal (e.g. a policy + // that is not active) or a transient failure. + replicationPUTStatus int } func newMockSBCLI() *mockSBCLI { m := &mockSBCLI{ - volumes: make(map[string]*mockVolume), - snapshots: make(map[string]*mockSnapshot), - groups: make(map[string]*mockGroup), + volumes: make(map[string]*mockVolume), + snapshots: make(map[string]*mockSnapshot), + groups: make(map[string]*mockGroup), + replicationStatus: make(map[string]map[string]any), } mux := http.NewServeMux() @@ -128,6 +148,10 @@ func newMockSBCLI() *mockSBCLI { "PUT /api/v2/clusters/{clusterID}/storage-pools/{poolID}/volumes/{volumeID}/", m.locked(m.handleResizeVolume), ) + mux.HandleFunc( + "GET /api/v2/clusters/{clusterID}/storage-pools/{poolID}/volumes/{volumeID}/replication/status", + m.locked(m.handleReplicationStatus), + ) mux.HandleFunc( "GET /api/v2/clusters/{clusterID}/storage-pools/{poolID}/volumes/{volumeID}/connect", m.locked(m.handleVolumeConnect), @@ -262,16 +286,57 @@ func (m *mockSBCLI) handleResizeVolume(w http.ResponseWriter, r *http.Request) { if volume == nil { return } - var body struct { - Size int64 `json:"size"` + if m.replicationPUTStatus != 0 { + writeJSON(w, m.replicationPUTStatus, map[string]string{"detail": "injected status"}) + return } - _ = json.NewDecoder(r.Body).Decode(&body) - if body.Size > 0 { - volume.Size = body.Size + raw, _ := io.ReadAll(r.Body) + var fields map[string]json.RawMessage + _ = json.Unmarshal(raw, &fields) + + if sizeRaw, ok := fields["size"]; ok { + var size int64 + if err := json.Unmarshal(sizeRaw, &size); err == nil && size > 0 { + volume.Size = size + } + } + // A key present with a JSON null attaches nothing (detach); a key present + // with a string attaches that policy; the key's absence (an ordinary + // resize PUT) leaves the volume's policy untouched -- omitted and null + // are different requests, which is exactly the distinction this mock + // exists to exercise. + if policyRaw, ok := fields["replication_policy_id"]; ok { + if string(policyRaw) == "null" { + volume.ReplicationPolicyID = "" + } else { + var policyID string + _ = json.Unmarshal(policyRaw, &policyID) + volume.ReplicationPolicyID = policyID + } } w.WriteHeader(http.StatusNoContent) } +// handleReplicationStatus serves the typed steady-state status a test +// configured via replicationStatus, or a default "not_replicating" body for +// a volume nothing has configured -- the same "never a 404" contract P0-1 +// promises for a volume that exists but never replicated. +func (m *mockSBCLI) handleReplicationStatus(w http.ResponseWriter, r *http.Request) { + volumeID := r.PathValue("volumeID") + if m.lookupVolume(w, volumeID) == nil { + return + } + body, ok := m.replicationStatus[volumeID] + if !ok { + body = map[string]any{ + "role": "none", "state": "not_replicating", + "outstanding_count": 0, "outstanding_bytes": 0, + "failing_count": 0, "max_retry_reached": false, "resyncing": false, + } + } + writeJSON(w, http.StatusOK, body) +} + func (m *mockSBCLI) handleVolumeConnect(w http.ResponseWriter, r *http.Request) { volumeID := r.PathValue("volumeID") if m.lookupVolume(w, volumeID) == nil { diff --git a/csi-driver/internal/csi/controller/replication.go b/csi-driver/internal/csi/controller/replication.go new file mode 100644 index 000000000..fdcfda386 --- /dev/null +++ b/csi-driver/internal/csi/controller/replication.go @@ -0,0 +1,113 @@ +// The csi-addons Replication service: EnableVolumeReplication, +// DisableVolumeReplication, and GetVolumeReplicationInfo (design §5.1). Each +// verb is a thin adapter onto the atlas-lib control-plane client's +// replication calls, resolved through the same {clusterID}:{poolID}:{lvolID} +// handle every other RPC uses. The remaining Replication verbs +// (PromoteVolume, DemoteVolume, ResyncVolume) fall through to the embedded +// UnimplementedControllerServer until Phase 2. +package controller + +import ( + "context" + + "github.com/csi-addons/spec/lib/go/replication" + "google.golang.org/grpc/codes" + "google.golang.org/grpc/status" + "google.golang.org/protobuf/types/known/timestamppb" + + "github.com/simplyblock/csi-driver/internal/clusters" + csicommon "github.com/simplyblock/csi-driver/internal/csi/common" +) + +// replicationPolicyParam is the VolumeReplicationClass parameter key naming +// the backend policy to attach. It is spelled as an id rather than the +// design's own `replicationPolicy` (a name, §7.1), because resolving a name +// to an id needs a policy-list-and-match call this phase does not yet wrap +// in atlas-lib. A future change adds that resolution and accepts either. +const replicationPolicyParam = "replicationPolicyID" + +// EnableVolumeReplication attaches the volume to the policy named by the +// VolumeReplicationClass. Attaching to the policy the volume already follows +// is success (the backend's own idempotency, P0-2). +// +// The design (§5.1, §10) wants a different-policy attach refused with +// FAILED_PRECONDITION, because a silent re-attach forces a full re-sync. That +// refusal is NOT implemented here: it needs the volume's CURRENTLY attached +// policy id to compare against, and neither the P0-1 status read nor any +// other backend endpoint exposes it today (attach_policy in sbcli's +// replication_policy_controller.py detaches and re-attaches on a policy +// change without refusing). Until the backend adds that field, a +// different-policy attach silently re-syncs, exactly as it does through +// every other existing caller of this same endpoint. +func (cs *Server) EnableVolumeReplication( + ctx context.Context, + req *replication.EnableVolumeReplicationRequest, +) (*replication.EnableVolumeReplicationResponse, error) { + policyID := req.GetParameters()[replicationPolicyParam] + if policyID == "" { + return nil, status.Errorf(codes.InvalidArgument, "VolumeReplicationClass parameter %q is required", replicationPolicyParam) + } + h, err := csicommon.ParseVolumeHandle(req.GetVolumeId()) + if err != nil { + return nil, status.Error(codes.InvalidArgument, err.Error()) + } + client, err := clusters.ReplicationClient(ctx, h.ClusterID) + if err != nil { + return nil, status.Error(codes.Unavailable, err.Error()) + } + if err := client.EnableVolumeReplication(ctx, h.Handle(), policyID); err != nil { + return nil, classifyEnableVolumeReplicationError(err) + } + return &replication.EnableVolumeReplicationResponse{}, nil +} + +// DisableVolumeReplication detaches the volume from whatever policy it +// follows. Detaching an already-detached volume is success; a cutover in +// flight (409) is a retryable ABORTED, since Ramen re-drives every reconcile. +func (cs *Server) DisableVolumeReplication( + ctx context.Context, + req *replication.DisableVolumeReplicationRequest, +) (*replication.DisableVolumeReplicationResponse, error) { + h, err := csicommon.ParseVolumeHandle(req.GetVolumeId()) + if err != nil { + return nil, status.Error(codes.InvalidArgument, err.Error()) + } + client, err := clusters.ReplicationClient(ctx, h.ClusterID) + if err != nil { + return nil, status.Error(codes.Unavailable, err.Error()) + } + if err := client.DisableVolumeReplication(ctx, h.Handle()); err != nil { + return nil, classifyDisableVolumeReplicationError(err) + } + return &replication.DisableVolumeReplicationResponse{}, nil +} + +// GetVolumeReplicationInfo returns the volume's replicated-life status. +// +// The spec's response carries only LastSyncTime in this version +// (github.com/csi-addons/spec v0.2.0); lastSyncBytes/lastSyncDuration are not +// yet part of the wire contract this driver builds against, so they cannot +// be set even though the backend status read already computes them (design +// §6.1). Ramen's condition derivation (§6.2) does not depend on them. +func (cs *Server) GetVolumeReplicationInfo( + ctx context.Context, + req *replication.GetVolumeReplicationInfoRequest, +) (*replication.GetVolumeReplicationInfoResponse, error) { + h, err := csicommon.ParseVolumeHandle(req.GetVolumeId()) + if err != nil { + return nil, status.Error(codes.InvalidArgument, err.Error()) + } + client, err := clusters.ReplicationClient(ctx, h.ClusterID) + if err != nil { + return nil, status.Error(codes.Unavailable, err.Error()) + } + info, err := client.GetVolumeReplicationInfo(ctx, h.Handle()) + if err != nil { + return nil, classifyGetVolumeReplicationInfoError(err) + } + resp := &replication.GetVolumeReplicationInfoResponse{} + if info.LastReplicatedAt != nil { + resp.LastSyncTime = timestamppb.New(*info.LastReplicatedAt) + } + return resp, nil +} diff --git a/csi-driver/internal/csi/controller/replication_test.go b/csi-driver/internal/csi/controller/replication_test.go new file mode 100644 index 000000000..b8f86be49 --- /dev/null +++ b/csi-driver/internal/csi/controller/replication_test.go @@ -0,0 +1,241 @@ +package controller + +import ( + "context" + "testing" + + "github.com/csi-addons/spec/lib/go/replication" + "google.golang.org/grpc/codes" + "google.golang.org/grpc/status" +) + +const ( + testReplVolumeID = "88888888-8888-8888-8888-888888888888" + testReplVolID = sanityClusterID + ":" + sanityPoolUUID + ":" + testReplVolumeID + testReplPolicyID = "77777777-7777-7777-7777-777777777777" +) + +func newReplicationTestServer(t *testing.T, mock *mockSBCLI) *Server { + t.Helper() + mock.volumes[testReplVolumeID] = &mockVolume{UUID: testReplVolumeID, Name: "repl-vol", Size: 1 << 30} + return newTestControllerServer(t, mock) +} + +func TestEnableVolumeReplication(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + + _, err := cs.EnableVolumeReplication(context.Background(), &replication.EnableVolumeReplicationRequest{ + VolumeId: testReplVolID, + Parameters: map[string]string{replicationPolicyParam: testReplPolicyID}, + }) + if err != nil { + t.Fatal(err) + } + if got := mock.volumes[testReplVolumeID].ReplicationPolicyID; got != testReplPolicyID { + t.Errorf("ReplicationPolicyID = %q, want %q", got, testReplPolicyID) + } +} + +// Repeating an enable already in effect is success without a second +// meaningful change (P0-2's own idempotency; the mock does not distinguish, +// it just re-applies the same value). +func TestEnableVolumeReplicationRepeatedIsIdempotent(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + req := &replication.EnableVolumeReplicationRequest{ + VolumeId: testReplVolID, + Parameters: map[string]string{replicationPolicyParam: testReplPolicyID}, + } + if _, err := cs.EnableVolumeReplication(context.Background(), req); err != nil { + t.Fatal(err) + } + if _, err := cs.EnableVolumeReplication(context.Background(), req); err != nil { + t.Errorf("second EnableVolumeReplication = %v, want nil (idempotent)", err) + } +} + +func TestEnableVolumeReplicationMissingPolicyParam(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + + _, err := cs.EnableVolumeReplication(context.Background(), &replication.EnableVolumeReplicationRequest{ + VolumeId: testReplVolID, + }) + st, _ := status.FromError(err) + if st.Code() != codes.InvalidArgument { + t.Errorf("code = %v, want InvalidArgument", st.Code()) + } +} + +func TestEnableVolumeReplicationBackendRefusal(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + mock.replicationPUTStatus = 412 + + _, err := cs.EnableVolumeReplication(context.Background(), &replication.EnableVolumeReplicationRequest{ + VolumeId: testReplVolID, + Parameters: map[string]string{replicationPolicyParam: testReplPolicyID}, + }) + st, _ := status.FromError(err) + if st.Code() != codes.FailedPrecondition { + t.Errorf("code = %v, want FailedPrecondition", st.Code()) + } +} + +func TestEnableVolumeReplicationMalformedVolumeHandle(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + + _, err := cs.EnableVolumeReplication(context.Background(), &replication.EnableVolumeReplicationRequest{ + VolumeId: "not-a-valid-handle", + Parameters: map[string]string{replicationPolicyParam: testReplPolicyID}, + }) + st, _ := status.FromError(err) + if st.Code() != codes.InvalidArgument { + t.Errorf("code = %v, want InvalidArgument", st.Code()) + } +} + +// A cluster this deployment's secret.json has no entry for is unreachable in +// the same sense a network partition would be: there is no client to make the +// call with, so the RPC must not be confused with a backend-side refusal. +func TestEnableVolumeReplicationUnknownCluster(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + + unregisteredClusterID := "99999999-9999-9999-9999-999999999999" + unknownClusterVolID := unregisteredClusterID + ":" + sanityPoolUUID + ":" + testReplVolumeID + + _, err := cs.EnableVolumeReplication(context.Background(), &replication.EnableVolumeReplicationRequest{ + VolumeId: unknownClusterVolID, + Parameters: map[string]string{replicationPolicyParam: testReplPolicyID}, + }) + st, _ := status.FromError(err) + if st.Code() != codes.Unavailable { + t.Errorf("code = %v, want Unavailable", st.Code()) + } +} + +func TestDisableVolumeReplication(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + mock.volumes[testReplVolumeID].ReplicationPolicyID = testReplPolicyID + + _, err := cs.DisableVolumeReplication(context.Background(), &replication.DisableVolumeReplicationRequest{ + VolumeId: testReplVolID, + }) + if err != nil { + t.Fatal(err) + } + if got := mock.volumes[testReplVolumeID].ReplicationPolicyID; got != "" { + t.Errorf("ReplicationPolicyID = %q, want cleared", got) + } +} + +// Disabling a volume that follows no policy is success, matching the +// backend's own idempotency (P0-2): there is nothing to detach. +func TestDisableVolumeReplicationNotAttachedIsSuccess(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + + _, err := cs.DisableVolumeReplication(context.Background(), &replication.DisableVolumeReplicationRequest{ + VolumeId: testReplVolID, + }) + if err != nil { + t.Errorf("DisableVolumeReplication = %v, want nil", err) + } +} + +func TestDisableVolumeReplicationDuringCutoverIsAborted(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + mock.replicationPUTStatus = 409 + + _, err := cs.DisableVolumeReplication(context.Background(), &replication.DisableVolumeReplicationRequest{ + VolumeId: testReplVolID, + }) + st, _ := status.FromError(err) + if st.Code() != codes.Aborted { + t.Errorf("code = %v, want Aborted (retryable)", st.Code()) + } +} + +func TestGetVolumeReplicationInfo(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + mock.replicationStatus[testReplVolumeID] = map[string]any{ + "role": "source", "state": "in_sync", + "last_replicated_at": "2026-09-17T12:00:00Z", + "lag_seconds": 42, + "outstanding_count": 0, "outstanding_bytes": 0, + "failing_count": 0, "max_retry_reached": false, "resyncing": false, + } + + resp, err := cs.GetVolumeReplicationInfo(context.Background(), &replication.GetVolumeReplicationInfoRequest{ + VolumeId: testReplVolID, + }) + if err != nil { + t.Fatal(err) + } + if resp.LastSyncTime == nil || resp.LastSyncTime.AsTime().Unix() != 1789646400 { + t.Errorf("LastSyncTime = %v, want 2026-09-17T12:00:00Z", resp.LastSyncTime) + } +} + +// A volume that never replicated is a valid answer, never a 404: the mock's +// default GET .../replication/status body for a volume nothing configured. +func TestGetVolumeReplicationInfoNeverReplicated(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + + resp, err := cs.GetVolumeReplicationInfo(context.Background(), &replication.GetVolumeReplicationInfoRequest{ + VolumeId: testReplVolID, + }) + if err != nil { + t.Fatal(err) + } + if resp.LastSyncTime != nil { + t.Errorf("LastSyncTime = %v, want nil", resp.LastSyncTime) + } +} + +func TestGetVolumeReplicationInfoUnknownVolume(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + delete(mock.volumes, testReplVolumeID) + + _, err := cs.GetVolumeReplicationInfo(context.Background(), &replication.GetVolumeReplicationInfoRequest{ + VolumeId: testReplVolID, + }) + st, _ := status.FromError(err) + if st.Code() != codes.NotFound { + t.Errorf("code = %v, want NotFound", st.Code()) + } +} + +// The remaining Replication verbs (Phase 2) fall through to the embedded +// UnimplementedControllerServer until they are implemented. +func TestUnimplementedReplicationVerbsAreUnimplemented(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + + _, err := cs.PromoteVolume(context.Background(), &replication.PromoteVolumeRequest{VolumeId: testReplVolID}) + st, _ := status.FromError(err) + if st.Code() != codes.Unimplemented { + t.Errorf("PromoteVolume code = %v, want Unimplemented", st.Code()) + } +} diff --git a/csi-driver/internal/csi/controller/server.go b/csi-driver/internal/csi/controller/server.go index ca9b630d8..46e3f801b 100644 --- a/csi-driver/internal/csi/controller/server.go +++ b/csi-driver/internal/csi/controller/server.go @@ -4,6 +4,7 @@ package controller import ( "github.com/container-storage-interface/spec/lib/go/csi" + "github.com/csi-addons/spec/lib/go/replication" "k8s.io/client-go/kubernetes" csicommon "github.com/simplyblock/csi-driver/internal/csi/common" @@ -15,6 +16,11 @@ type Server struct { // groupsnapshot.go; embedding the unimplemented server satisfies the // interface's forward-compat guard for any method not overridden. csi.UnimplementedGroupControllerServer + // The csi-addons Replication service (design §5) is implemented in + // replication.go, for the three Phase 1 verbs; the rest + // (PromoteVolume, DemoteVolume, ResyncVolume) fall through to this + // embedded default until Phase 2. + replication.UnimplementedControllerServer volumeLocks *csicommon.VolumeLocks // kubeClient reads/patches PVC annotations (host_id resolution, placement-hint // cleanup). Built once at construction and reused, and nil when no in-cluster diff --git a/csi-driver/internal/csi/csiaddons/identity/identity.go b/csi-driver/internal/csi/csiaddons/identity/identity.go new file mode 100644 index 000000000..8def40c96 --- /dev/null +++ b/csi-driver/internal/csi/csiaddons/identity/identity.go @@ -0,0 +1,61 @@ +// Package identity serves the csi-addons Identity service: GetIdentity, +// GetCapabilities (advertising VOLUME_REPLICATION), and Probe. It is a +// distinct service from the CSI spec's own Identity (served by +// internal/csi/identity), not an extension of it, because the +// kubernetes-csi-addons controller-manager and the CSI sidecars probe two +// separate protocols on the same socket. +package identity + +import ( + "context" + + "github.com/csi-addons/spec/lib/go/identity" +) + +// Server implements the csi-addons Identity service. +type Server struct { + identity.UnimplementedIdentityServer + name string + version string +} + +// New returns an identity.Server reporting name and version as this driver's +// own (the same values the CSI spec's Identity service reports), since both +// protocols identify the one plugin process serving them. +func New(name, version string) *Server { + return &Server{name: name, version: version} +} + +func (s *Server) GetIdentity(context.Context, *identity.GetIdentityRequest) (*identity.GetIdentityResponse, error) { + return &identity.GetIdentityResponse{Name: s.name, VendorVersion: s.version}, nil +} + +func (s *Server) GetCapabilities( + context.Context, *identity.GetCapabilitiesRequest, +) (*identity.GetCapabilitiesResponse, error) { + return &identity.GetCapabilitiesResponse{ + Capabilities: []*identity.Capability{ + { + Type: &identity.Capability_Service_{ + Service: &identity.Capability_Service{ + Type: identity.Capability_Service_CONTROLLER_SERVICE, + }, + }, + }, + { + Type: &identity.Capability_VolumeReplication_{ + VolumeReplication: &identity.Capability_VolumeReplication{ + Type: identity.Capability_VolumeReplication_VOLUME_REPLICATION, + }, + }, + }, + }, + }, nil +} + +func (s *Server) Probe(context.Context, *identity.ProbeRequest) (*identity.ProbeResponse, error) { + // Ready left nil: per the spec, absent means "assume ready." The + // Replication service's RPCs are stateless and idempotent by contract + // (design §5), so there is no initialization phase to report against. + return &identity.ProbeResponse{}, nil +} diff --git a/csi-driver/internal/csi/csiaddons/identity/identity_test.go b/csi-driver/internal/csi/csiaddons/identity/identity_test.go new file mode 100644 index 000000000..ad47630d1 --- /dev/null +++ b/csi-driver/internal/csi/csiaddons/identity/identity_test.go @@ -0,0 +1,55 @@ +package identity + +import ( + "context" + "testing" + + "github.com/csi-addons/spec/lib/go/identity" +) + +func TestGetIdentity(t *testing.T) { + s := New("test.csi.simplyblock.io", "v1.2.3") + resp, err := s.GetIdentity(context.Background(), &identity.GetIdentityRequest{}) + if err != nil { + t.Fatal(err) + } + if resp.Name != "test.csi.simplyblock.io" || resp.VendorVersion != "v1.2.3" { + t.Errorf("GetIdentity = %+v", resp) + } +} + +func TestGetCapabilitiesAdvertisesVolumeReplication(t *testing.T) { + s := New("test.csi.simplyblock.io", "v1.2.3") + resp, err := s.GetCapabilities(context.Background(), &identity.GetCapabilitiesRequest{}) + if err != nil { + t.Fatal(err) + } + var sawVolumeReplication, sawControllerService bool + for _, c := range resp.Capabilities { + if vr := c.GetVolumeReplication(); vr != nil && vr.Type == identity.Capability_VolumeReplication_VOLUME_REPLICATION { + sawVolumeReplication = true + } + if svc := c.GetService(); svc != nil && svc.Type == identity.Capability_Service_CONTROLLER_SERVICE { + sawControllerService = true + } + } + if !sawVolumeReplication { + t.Error("capabilities do not advertise VOLUME_REPLICATION") + } + if !sawControllerService { + t.Error("capabilities do not advertise CONTROLLER_SERVICE") + } +} + +func TestProbeReportsReady(t *testing.T) { + s := New("test.csi.simplyblock.io", "v1.2.3") + resp, err := s.Probe(context.Background(), &identity.ProbeRequest{}) + if err != nil { + t.Fatal(err) + } + // Ready left nil means "assume ready" per the spec; asserting nil pins + // that choice rather than a stray true/false creeping in later. + if resp.Ready != nil { + t.Errorf("Ready = %v, want nil (assume ready)", resp.Ready) + } +} diff --git a/csi-driver/internal/driver/driver.go b/csi-driver/internal/driver/driver.go index 4493cd311..b80731943 100644 --- a/csi-driver/internal/driver/driver.go +++ b/csi-driver/internal/driver/driver.go @@ -28,6 +28,9 @@ import ( "fmt" "github.com/container-storage-interface/spec/lib/go/csi" + csiaddonsidentity "github.com/csi-addons/spec/lib/go/identity" + csiaddonsreplication "github.com/csi-addons/spec/lib/go/replication" + "google.golang.org/grpc" "k8s.io/client-go/kubernetes" "k8s.io/client-go/rest" "k8s.io/klog" @@ -41,6 +44,7 @@ import ( csicommon "github.com/simplyblock/csi-driver/internal/csi/common" "github.com/simplyblock/csi-driver/internal/csi/controller" "github.com/simplyblock/csi-driver/internal/csi/identity" + csiaddonsidentityserver "github.com/simplyblock/csi-driver/internal/csi/csiaddons/identity" "github.com/simplyblock/csi-driver/internal/csi/node" "github.com/simplyblock/csi-driver/internal/csilink" "github.com/simplyblock/csi-driver/internal/guardian" @@ -131,8 +135,23 @@ func Run(conf *config.Config) { } } + // The csi-addons Identity and Replication services register alongside the + // CSI services on the same socket. Identity is always registered (it just + // answers capability probes); Replication only when this process serves + // the controller (cs is nil on a node-only process). + register := []func(*grpc.Server){ + func(gs *grpc.Server) { + csiaddonsidentity.RegisterIdentityServer(gs, csiaddonsidentityserver.New(conf.DriverName, conf.DriverVersion)) + }, + } + if cs != nil { + register = append(register, func(gs *grpc.Server) { + csiaddonsreplication.RegisterControllerServer(gs, cs) + }) + } + s := csicommon.NewNonBlockingGRPCServer() - s.Start(conf.Endpoint, ids, cs, ns) + s.Start(conf.Endpoint, ids, cs, ns, register...) s.Wait() } diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_simplyblockdrivers.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_simplyblockdrivers.yaml index 82ab4ff4c..e5d075f08 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_simplyblockdrivers.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_simplyblockdrivers.yaml @@ -299,13 +299,23 @@ spec: type: object sidecarImages: description: |- - SidecarImages overrides the six CSI sidecars, one field each. Unset takes - the version this operator release ships. + SidecarImages overrides the seven CSI sidecars, one field each. Unset + takes the version this operator release ships. properties: attacher: description: Attacher is csi-attacher, on the controller plugin. pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ type: string + csiAddons: + description: |- + CSIAddons is the kubernetes-csi-addons sidecar, on the controller + plugin. It connects to the plugin's socket, probes the csi-addons + Identity service for capabilities, and publishes a CSIAddonsNode so the + kubernetes-csi-addons controller-manager (design + design-csi-addons-replication.md §4.1) can reach the Replication + service this driver serves. + pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ + type: string healthMonitor: description: |- HealthMonitor is csi-external-health-monitor-controller, on the diff --git a/operator/api/v1alpha2/simplyblockdriver_types.go b/operator/api/v1alpha2/simplyblockdriver_types.go index 49de327fa..873c046b1 100644 --- a/operator/api/v1alpha2/simplyblockdriver_types.go +++ b/operator/api/v1alpha2/simplyblockdriver_types.go @@ -84,6 +84,16 @@ type SidecarImages struct { // +kubebuilder:validation:Pattern=`^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$` // +optional NodeDriverRegistrar string `json:"nodeDriverRegistrar,omitempty"` + + // CSIAddons is the kubernetes-csi-addons sidecar, on the controller + // plugin. It connects to the plugin's socket, probes the csi-addons + // Identity service for capabilities, and publishes a CSIAddonsNode so the + // kubernetes-csi-addons controller-manager (design + // design-csi-addons-replication.md §4.1) can reach the Replication + // service this driver serves. + // +kubebuilder:validation:Pattern=`^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$` + // +optional + CSIAddons string `json:"csiAddons,omitempty"` } // DriverTLSProvider is where the TLS certificate on this connection comes @@ -208,8 +218,8 @@ type SimplyblockDriverSpec struct { // +optional NodeResources corev1.ResourceRequirements `json:"nodeResources,omitempty"` - // SidecarImages overrides the six CSI sidecars, one field each. Unset takes - // the version this operator release ships. + // SidecarImages overrides the seven CSI sidecars, one field each. Unset + // takes the version this operator release ships. // +optional SidecarImages SidecarImages `json:"sidecarImages,omitempty"` diff --git a/operator/config/crd/bases/storage.simplyblock.io_simplyblockdrivers.yaml b/operator/config/crd/bases/storage.simplyblock.io_simplyblockdrivers.yaml index 82ab4ff4c..e5d075f08 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_simplyblockdrivers.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_simplyblockdrivers.yaml @@ -299,13 +299,23 @@ spec: type: object sidecarImages: description: |- - SidecarImages overrides the six CSI sidecars, one field each. Unset takes - the version this operator release ships. + SidecarImages overrides the seven CSI sidecars, one field each. Unset + takes the version this operator release ships. properties: attacher: description: Attacher is csi-attacher, on the controller plugin. pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ type: string + csiAddons: + description: |- + CSIAddons is the kubernetes-csi-addons sidecar, on the controller + plugin. It connects to the plugin's socket, probes the csi-addons + Identity service for capabilities, and publishes a CSIAddonsNode so the + kubernetes-csi-addons controller-manager (design + design-csi-addons-replication.md §4.1) can reach the Replication + service this driver serves. + pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ + type: string healthMonitor: description: |- HealthMonitor is csi-external-health-monitor-controller, on the diff --git a/operator/docs/designs/design-csi-addons-replication.md b/operator/docs/designs/design-csi-addons-replication.md index af2dbf3f2..c3e5be50a 100644 --- a/operator/docs/designs/design-csi-addons-replication.md +++ b/operator/docs/designs/design-csi-addons-replication.md @@ -1,6 +1,6 @@ # Design Document: csi-addons Volume Replication -**Status:** Draft +**Status:** Phase 1 Implemented **Author:** Israel Geoffrey (geoffrey1330) **Date:** 2026-09-16 (last updated 2026-09-17) **Test Plan:** [`tests/test-plan-csi-addons-replication.md`](../tests/test-plan-csi-addons-replication.md) @@ -9,12 +9,12 @@ ## Phasing Overview -| Phase | Status | Scope | Sections | -|-------------|---------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|--------------| -| **Phase 1** | Planned | The csi-addons machinery and the steady-state contract: CRDs, controller-manager, sidecar, the Replication and csi-addons Identity gRPC services with `EnableVolumeReplication`, `DisableVolumeReplication`, and `GetVolumeReplicationInfo`, backed by a typed backend status endpoint | §4, §5.1, §6 | -| **Phase 2** | Planned | The lifecycle verbs: `PromoteVolume` (planned and forced), `DemoteVolume`, and `ResyncVolume`, validated end to end against a Ramen `VolumeReplicationGroup` in async mode | §5.2, §9 | -| **Phase 3** | Planned | peerClasses convention and preflight, and the replication observability surface (lag, backlog, RPO compliance) | §7, §11 | -| **Phase 4** | Planned | Test failover: the latest-replicated-snapshot read, the `drtest-*` conventions, and the two drill modes (bubble and test cluster) composed from clone, replication, and the real failover | §14 | +| Phase | Status | Scope | Sections | +|-------------|-------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|--------------| +| **Phase 1** | Implemented | The csi-addons machinery and the steady-state contract: CRDs, controller-manager, sidecar, the Replication and csi-addons Identity gRPC services with `EnableVolumeReplication`, `DisableVolumeReplication`, and `GetVolumeReplicationInfo`, backed by a typed backend status endpoint | §4, §5.1, §6 | +| **Phase 2** | Planned | The lifecycle verbs: `PromoteVolume` (planned and forced), `DemoteVolume`, and `ResyncVolume`, validated end to end against a Ramen `VolumeReplicationGroup` in async mode | §5.2, §9 | +| **Phase 3** | Planned | peerClasses convention and preflight, and the replication observability surface (lag, backlog, RPO compliance) | §7, §11 | +| **Phase 4** | Planned | Test failover: the latest-replicated-snapshot read, the `drtest-*` conventions, and the two drill modes (bubble and test cluster) composed from clone, replication, and the real failover | §14 | Phase 1 is independently useful: a `VolumeReplication` object per PVC whose status truthfully reports the relationship, which no surface provides today. Phase 2 makes the object drivable, which is what Ramen actually needs. Phase 3 makes the whole thing operable at fleet scale. Phase 4 turns the same primitives into a rehearsal: a failover that can be drilled, in a bubble or against a test cluster, without touching production replication. @@ -24,14 +24,14 @@ The phase numbers above are this document's own, not the DR storage foundation g ## Phase 0 — External Prerequisites -| # | Prerequisite | Kind | Blocks | Status | -|------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------|---------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| P0-1 | A typed, steady-state per-volume replication status read: `GET .../volumes/{id}/replication/status` serving what `lvol_controller.get_replication_info` computes today (state, lag, outstanding bytes, failure counters), available for the volume's whole replicated life | Control plane (`sbcli`) | Phase 1 | Not shipped | -| P0-2 | Idempotent attach and detach: attaching a volume to the policy it already follows returns success, and detaching a non-attached volume returns success | Control plane (`sbcli`) | Phase 1 | Not shipped | -| P0-3 | A standalone demote verb: `POST .../volumes/{id}/replication/demote` that converges the peer while still serving (repeated snapshot-and-ship until the remaining delta is small), then quiesces, ships the final delta, confirms it landed on the peer, and fences the data path | Control plane (`sbcli`) | Phase 2 | Not shipped | -| P0-4 | An `rpo_target_seconds` field on `ReplicationPolicy`, so RPO compliance is computable against a declared target rather than the derived lag budget | Control plane (`sbcli`) | Phase 3 | Not shipped | -| P0-5 | csi-addons upstream: the `VolumeReplication` and `VolumeReplicationClass` CRDs (`replication.storage.openshift.io/v1alpha1`), the kubernetes-csi-addons controller-manager image, and the csi-addons sidecar image | Ecosystem | Phase 1 | Vendored in the chart at v0.15.0 behind `csiaddons.create` (all twelve upstream CRDs, since the stock manager starts a controller per kind); sidecar wiring is Phase 1 | -| P0-6 | A latest-replicated-snapshot read: per volume, and per consistency group as one complete generation, the newest fully replicated snapshot on the secondary addressed as a cloneable object | Control plane (`sbcli`) | Phase 4 | Not shipped | +| # | Prerequisite | Kind | Blocks | Status | +|------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------|---------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| P0-1 | A typed, steady-state per-volume replication status read: `GET .../volumes/{id}/replication/status` serving what `lvol_controller.get_replication_info` computes today (state, lag, outstanding bytes, failure counters), available for the volume's whole replicated life | Control plane (`sbcli`) | Phase 1 | Shipped | +| P0-2 | Idempotent attach and detach: attaching a volume to the policy it already follows returns success, and detaching a non-attached volume returns success | Control plane (`sbcli`) | Phase 1 | Shipped | +| P0-3 | A standalone demote verb: `POST .../volumes/{id}/replication/demote` that converges the peer while still serving (repeated snapshot-and-ship until the remaining delta is small), then quiesces, ships the final delta, confirms it landed on the peer, and fences the data path | Control plane (`sbcli`) | Phase 2 | Not shipped | +| P0-4 | An `rpo_target_seconds` field on `ReplicationPolicy`, so RPO compliance is computable against a declared target rather than the derived lag budget | Control plane (`sbcli`) | Phase 3 | Shipped | +| P0-5 | csi-addons upstream: the `VolumeReplication` and `VolumeReplicationClass` CRDs (`replication.storage.openshift.io/v1alpha1`), the kubernetes-csi-addons controller-manager image, and the csi-addons sidecar image | Ecosystem | Phase 1 | Vendored in the chart at v0.15.0 behind `csiaddons.create` (all twelve upstream CRDs, since the stock manager starts a controller per kind); sidecar wiring shipped in Phase 1 | +| P0-6 | A latest-replicated-snapshot read: per volume, and per consistency group as one complete generation, the newest fully replicated snapshot on the secondary addressed as a cloneable object | Control plane (`sbcli`) | Phase 4 | Shipped | Everything else the adapter needs already exists: the attach and detach calls, failover, the failback and commit pair, the relationship read, and the backlog arithmetic inside `get_replication_info`. The adapter is thin precisely because the engine is complete. What is missing is the shape Ramen can drive. @@ -181,7 +181,7 @@ A reader who stops here has the model: the engine is unchanged, the csi-addons s The plugin registers two additional gRPC services on the existing socket, beside the CSI services, following the GroupController precedent in `csicommon.NonBlockingGRPCServer`: - **csi-addons Identity:** `GetIdentity`, `GetCapabilities` (advertising `VOLUME_REPLICATION`), and `Probe`. This is a distinct service from CSI Identity, so it is a new small server type, not an extension of the existing one. -- **Replication:** the six verbs of §5, implemented on the controller `Server` through an embedded `UnimplementedReplicationServer` from `github.com/csi-addons/spec`, mirroring how the GroupController embeds its unimplemented base. +- **Replication:** the six verbs of §5, implemented on the controller `Server` through an embedded `replication.UnimplementedControllerServer` from `github.com/csi-addons/spec` (pinned at v0.2.0), mirroring how the GroupController embeds its unimplemented base. `NonBlockingGRPCServer.Start` today takes exactly the three CSI servers. It gains a registration hook so the driver package can register additional services without `csicommon` importing csi-addons. @@ -197,11 +197,11 @@ The generated control-plane client already declares every replication endpoint, ### 5.1 Phase 1 verbs -| Verb | Backend mapping | Semantics | -|----------------------------|------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| `EnableVolumeReplication` | `PUT .../volumes/{v}` body `{"replication_policy_id": }` | The policy comes from the `VolumeReplicationClass` parameters (§7). Attach is synchronous on the backend; already-attached to the same policy is success (P0-2). Attached to a *different* policy is `FAILED_PRECONDITION`, because silently re-attaching forces a full re-sync. | -| `DisableVolumeReplication` | `PUT .../volumes/{v}` body `{"replication_policy_id": null}` | Detach. A 409 (cutover in flight) maps to `ABORTED`, retryable. Not-attached is success (P0-2). | -| `GetVolumeReplicationInfo` | `GET .../volumes/{v}/replication/status` (P0-1) | Returns `lastSyncTime` (newest fully replicated snapshot's creation time), `lastSyncDuration` (last cycle duration), and `lastSyncBytes` (last shipped snapshot's used size). | +| Verb | Backend mapping | Semantics | +|----------------------------|------------------------------------------------------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| `EnableVolumeReplication` | `PUT .../volumes/{v}` body `{"replication_policy_id": }` | The policy comes from the `VolumeReplicationClass` parameters (§7). Attach is synchronous on the backend; already-attached to the same policy is success (P0-2). Attaching to a *different* policy is **not** refused: sbcli's `attach_policy` silently re-attaches onto the new policy rather than returning `FAILED_PRECONDITION`, and no endpoint yet exposes the policy id a volume is currently attached to for the driver to compare against before attaching (§15, Open Question 3). | +| `DisableVolumeReplication` | `PUT .../volumes/{v}` body `{"replication_policy_id": null}` | Detach. A 409 (cutover in flight) maps to `ABORTED`, retryable. Not-attached is success (P0-2). | +| `GetVolumeReplicationInfo` | `GET .../volumes/{v}/replication/status` (P0-1) | Returns `lastSyncTime` (newest fully replicated snapshot's creation time). `csi-addons/spec` v0.2.0's `GetVolumeReplicationInfoResponse` carries no `lastSyncDuration` or `lastSyncBytes` field, so the status read's cycle-duration and shipped-size data has no response field to land in until a newer spec version adds one. | ### 5.2 Phase 2 verbs @@ -437,7 +437,9 @@ A drill that silently perturbed replication would be worse than no drill. Before ## 15. Open Questions -| # | Question | Owner | -|-----|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|---------------| -| 1 | **Demote semantics for the application.** The P0-3 demote fences the volume (ANA inaccessible) after the final flush, and with convergence folded into the verb it is now the only place a planned swap can stall. Ramen relocation unmounts the workload first, so the fence is ordinarily unopposed. Confirm the verb's behavior when writes are still in flight at quiesce (block versus fail), whether the converge phase has its own budget separate from the quiesced flush, and whether a timeout in either phase must abort back to serving primary. | Backend team | -| 2 | **Where the preflight lives.** §7.2 attaches peerClasses validation to the `ReplicationPair` reconciler. If the redesign retires the pair kind, the preflight needs a new home (the `SimplyblockDriver`, or a standalone check job). | Operator team | +| # | Question | Owner | +|-----|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-----------------------| +| 1 | **Demote semantics for the application.** The P0-3 demote fences the volume (ANA inaccessible) after the final flush, and with convergence folded into the verb it is now the only place a planned swap can stall. Ramen relocation unmounts the workload first, so the fence is ordinarily unopposed. Confirm the verb's behavior when writes are still in flight at quiesce (block versus fail), whether the converge phase has its own budget separate from the quiesced flush, and whether a timeout in either phase must abort back to serving primary. | Backend team | +| 2 | **Where the preflight lives.** §7.2 attaches peerClasses validation to the `ReplicationPair` reconciler. If the redesign retires the pair kind, the preflight needs a new home (the `SimplyblockDriver`, or a standalone check job). | Operator team | +| 3 | **Different-policy enable refusal has no data source.** §5.1's `EnableVolumeReplication` was designed to refuse an attach to a different policy with `FAILED_PRECONDITION`. Phase 1 found that sbcli's `attach_policy` silently re-attaches instead of refusing, and no endpoint returns the policy id a volume is currently attached to, for the driver to compare against. Either the backend adds that read, or this design accepts the silent re-sync as the behavior. | Backend team | +| 4 | **`lastSyncDuration` and `lastSyncBytes` have no response field.** §5.1's `GetVolumeReplicationInfo` was designed to return all three fields; `csi-addons/spec` v0.2.0's `GetVolumeReplicationInfoResponse` carries only `lastSyncTime`. Confirm whether a newer spec version adds the other two, or whether they surface some other way (a `VolumeReplication` annotation, a metric). | Ecosystem/driver team | diff --git a/operator/docs/tests/test-plan-csi-addons-replication.md b/operator/docs/tests/test-plan-csi-addons-replication.md index 19bbec5fd..692b47c95 100644 --- a/operator/docs/tests/test-plan-csi-addons-replication.md +++ b/operator/docs/tests/test-plan-csi-addons-replication.md @@ -4,9 +4,9 @@ Related design: [`designs/design-csi-addons-replication.md`](../designs/design-c Scope is the CSI driver's Replication service, the operator's preflight and coexistence rules, and the deployment of the csi-addons machinery. The replication engine itself (snapshot shipping, failover cloning, the cutover task runner) is the control plane's to prove and is exercised here only through the adapter's boundary. The kubernetes-csi-addons controller-manager is stock upstream and is not re-tested; what is tested is this driver's conformance to the contract it drives. -Scenario IDs are permanent and are never reused or renumbered. `U-` is unit (no cluster: mock control plane, fake `client.Client`), `I-` is integration (the sidecar and controller-manager against the driver with a mock backend), `E-` is end-to-end (two live simplyblock clusters), and `M-` is manual. Types are `Positive`, `Negative`, `Boundary`, and `Regression`. A `—` in the `Test` column means nothing implements the scenario yet, and every such row reappears in §6 with its reason. +Scenario IDs are permanent and are never reused or renumbered. `U-` is unit (no cluster: mock control plane, fake `client.Client`), `I-` is integration (the sidecar and controller-manager against the driver with a mock backend), `E-` is end-to-end (two live simplyblock clusters), and `M-` is manual. Types are `Positive`, `Negative`, `Boundary`, and `Regression`. A `—` in the `Test` column means nothing implements the scenario yet, and every such row reappears in §7 with its reason. -The plan is the target coverage for a Draft design: every row is `—` until the work lands. +Phase 1 (the csi-addons machinery, §4, §5.1's three verbs, and §6's steady-state contract) has landed; its unit rows below are filled in. Phase 2 (promote, demote, resync, and the operator's preflight and coexistence controllers) has not started, and the integration and E2E tiers wait on a test bed neither phase has built yet. --- @@ -18,17 +18,22 @@ The Replication service against a mock control plane, and the operator pieces ag File: `csi-driver/internal/csi/controller/replication_test.go` (planned) -| # | Scenario | Type | Test | -|------|---------------------------------------------------------------------------------------------------------------|----------|------| -| U-01 | Enable on an unattached volume: the attach call carries the class's policy, and the RPC succeeds | Positive | — | -| U-02 | Enable on a volume already attached to the same policy: success, no second attach call (idempotency) | Boundary | — | -| U-03 | Enable on a volume attached to a different policy: `FAILED_PRECONDITION` naming both policies, no attach call | Negative | — | -| U-04 | Disable on an attached volume: the detach call is made, success | Positive | — | -| U-05 | Disable on a non-attached volume: success without a backend call (idempotency) | Boundary | — | -| U-06 | Disable while a cutover is in flight (backend 409): `ABORTED`, retryable | Negative | — | -| U-07 | Info returns `lastSyncTime`, `lastSyncDuration`, and `lastSyncBytes` from the status read | Positive | — | -| U-08 | A malformed volume handle: `INVALID_ARGUMENT` before any backend call | Negative | — | -| U-09 | Backend unreachable: `UNAVAILABLE`, and no condition flap is implied by the error | Negative | — | +| # | Scenario | Type | Test | +|------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|------------|------------------------------------------------------| +| U-01 | Enable on an unattached volume: the attach call carries the class's policy, and the RPC succeeds | Positive | `TestEnableVolumeReplication` | +| U-02 | Enable on a volume already attached to the same policy: success, no second attach call (idempotency) | Boundary | `TestEnableVolumeReplicationRepeatedIsIdempotent` | +| U-03 | Enable on a volume attached to a different policy: `FAILED_PRECONDITION` naming both policies, no attach call | Negative | — | +| U-04 | Disable on an attached volume: the detach call is made, success | Positive | `TestDisableVolumeReplication` | +| U-05 | Disable on a non-attached volume: success without a backend call (idempotency) | Boundary | `TestDisableVolumeReplicationNotAttachedIsSuccess` | +| U-06 | Disable while a cutover is in flight (backend 409): `ABORTED`, retryable | Negative | `TestDisableVolumeReplicationDuringCutoverIsAborted` | +| U-07 | Info returns `lastSyncTime` from the status read (`lastSyncDuration` and `lastSyncBytes` are not part of `GetVolumeReplicationInfoResponse` in csi-addons/spec v0.2.0, the version this driver builds against) | Positive | `TestGetVolumeReplicationInfo` | +| U-08 | A malformed volume handle: `INVALID_ARGUMENT` before any backend call | Negative | `TestEnableVolumeReplicationMalformedVolumeHandle` | +| U-09 | A cluster ID with no entry in this deployment's secret: `UNAVAILABLE`, distinct from a backend-side refusal | Negative | `TestEnableVolumeReplicationUnknownCluster` | +| U-28 | Enable without the `VolumeReplicationClass` policy parameter: `INVALID_ARGUMENT` before any backend call | Negative | `TestEnableVolumeReplicationMissingPolicyParam` | +| U-29 | Enable the backend refuses (412): `FAILED_PRECONDITION` carrying the backend's own reason | Negative | `TestEnableVolumeReplicationBackendRefusal` | +| U-30 | Info on a volume that never replicated: a nil `lastSyncTime`, never a `NOT_FOUND` | Boundary | `TestGetVolumeReplicationInfoNeverReplicated` | +| U-31 | Info on a volume id the backend does not recognize: `NOT_FOUND` | Negative | `TestGetVolumeReplicationInfoUnknownVolume` | +| U-32 | The remaining Replication verbs (`PromoteVolume`, `DemoteVolume`, `ResyncVolume`) fall through to `UNIMPLEMENTED` until Phase 2 lands | Regression | `TestUnimplementedReplicationVerbsAreUnimplemented` | ### Replication Verbs: Promote, Demote, Resync (design §5.2) @@ -69,6 +74,27 @@ Files: `operator/internal/controller/peerclasses_preflight_test.go`, `operator/i | U-26 | `PVCAnnotationWatcher` skips a PVC whose volume has a `VolumeReplication`: no slot is created, a skip is recorded | Negative | — | | U-27 | Annotation added and later a `VolumeReplication` appears: the existing slot is not deleted by the adapter, and the enable is refused per the one-owner rule | Boundary | — | +### csi-addons Identity Service (design §4) + +File: `csi-driver/internal/csi/csiaddons/identity/identity_test.go` + +| # | Scenario | Type | Test | +|------|--------------------------------------------------------------------------------|----------|--------------------------------------------------| +| U-33 | `GetIdentity` returns this driver's name and version | Positive | `TestGetIdentity` | +| U-34 | `GetCapabilities` advertises `VOLUME_REPLICATION` and `CONTROLLER_SERVICE` | Positive | `TestGetCapabilitiesAdvertisesVolumeReplication` | +| U-35 | `Probe` reports ready (a nil `Ready` per the spec, not a stray `true`/`false`) | Positive | `TestProbeReportsReady` | + +### Operator: Sidecar Deployment and RBAC (design §4.1) + +File: `operator/internal/controllers/driver/workloads_test.go`, `operator/internal/controllers/driver/rbac_test.go` + +| # | Scenario | Type | Test | +|------|-----------------------------------------------------------------------------------------------------------------------------------------|----------|---------------------------------------------------------------| +| U-36 | The csi-addons sidecar is appended after the plugin container, never inserted, and is addressed at the plugin's own socket | Positive | `TestCSIAddonsSidecarIsAppliedAfterThePlugin` | +| U-37 | The sidecar advertises its own pod (IP, name, namespace, UID) through the downward API, since the StatefulSet runs on the host network | Positive | `TestCSIAddonsSidecarAdvertisesItsOwnPod` | +| U-38 | The sidecar's grant is a namespaced Role bound to the controller plugin's account, not a ClusterRole, since CSIAddonsNode is namespaced | Positive | `TestCSIAddonsRoleIsNamespacedAndBoundToTheControllerAccount` | +| U-39 | The namespaced Role's rules are scoped to the sidecar's own job: its CSIAddonsNode and its own leader-election Lease | Positive | `TestCSIAddonsRoleRulesAreScopedToItsOwnJob` | + --- ## 2. Integration Tests @@ -153,23 +179,26 @@ Two live simplyblock clusters with the chart-deployed csi-addons machinery. The ## 6. Coverage Summary -| Class | Scenarios | Covered | Not covered | -|-------------|-----------|---------|-------------| -| Unit | 27 | 0 | U-01 … U-27 | -| Integration | 7 | 0 | I-01 … I-07 | -| E2E | 7 | 0 | E-01 … E-07 | -| Manual | 2 | 0 | M-01, M-02 | +| Class | Scenarios | Covered | Not covered | +|-------------|-----------|---------|-------------------| +| Unit | 39 | 20 | U-03, U-10 … U-27 | +| Integration | 7 | 0 | I-01 … I-07 | +| E2E | 7 | 0 | E-01 … E-07 | +| Manual | 2 | 0 | M-01, M-02 | -Every scenario is uncovered because the design is Draft. The counts are the target, and each `Test` column fills in as the work lands. +Phase 1 landed the driver's Replication and Identity services, the error classifier, and the operator's sidecar and RBAC wiring, covering every Phase 1 unit scenario except U-03 (§7). Phase 2 (promote, demote, resync, and the operator's preflight and coexistence controllers) has not started, and neither has a sidecar-and-controller-manager integration suite or a live two-cluster E2E bed, so those tiers remain fully uncovered. --- ## 7. What Is Not Yet Covered -| # | Gap | Reason | -|-------------|-------------------------------------------------------------------------------------------------------------------|---------------------------------------------------------------------------------------------------------------------------------------------| -| U-01 … U-27 | The verb adapter, condition derivation, preflight, and coexistence units | The adapter does not exist; blocked on P0-1 and P0-2 for the Phase 1 verbs and P0-3 for demote | -| I-01 … I-07 | The sidecar and controller-manager loop | CRDs and controller-manager vendored in the chart (P0-5, `csiaddons.create`); blocked on the Phase 1 sidecar and driver Replication service | -| E-01 … E-07 | The live lifecycle and the Ramen gate | Blocked on Phase 1 and 2 landing, plus a two-cluster test bed with Ramen dr-cluster installed for E-06 and E-07 | -| — | Repeated resync, class drift after verification, annotated-volume migration onto the adapter, cascaded topologies | Beyond the first coverage pass, recorded so the gaps are explicit rather than assumed covered | -| M-01, M-02 | Demote under writes; concurrent ownership race | Need failure injection and precise timing a live two-cluster run does not automate yet | +| # | Gap | Reason | +|-------------|-------------------------------------------------------------------------------------------------------------------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| U-03 | Different-policy enable refused with `FAILED_PRECONDITION` | Not implemented: sbcli's `attach_policy` silently re-attaches onto the new policy rather than refusing, and no endpoint exposes the policy id a volume is currently attached to, for the driver to compare against before attaching (design §5.1 assumed this refusal exists; it does not against the current backend) | +| U-10 … U-17 | Promote, demote, and resync unit coverage | Phase 2: the three verbs fall through to `UnimplementedControllerServer`; P0-3 (demote's backend support) is also still blocked | +| U-18 … U-22 | Condition derivation (`Completed`/`Degraded`/`Resyncing`) from the status read | Not implemented: `csi-addons/spec` v0.2.0's `GetVolumeReplicationInfoResponse` carries only `lastSyncTime`, with no per-condition field at all; deriving these needs either a newer spec version or belongs in the controller-manager's own reconcile, neither examined yet | +| U-23 … U-27 | Preflight (`peerClasses` verification) and coexistence (`PVCReplicationController`, the one-owner rule) | Out of Phase 1's scope: the auto-adapter and preflight webhook are a separate, unbuilt subsystem | +| I-01 … I-07 | The sidecar and controller-manager loop | The driver's Replication and Identity services and the sidecar container now exist (Phase 1); no envtest/kind suite exercises them against the real kubernetes-csi-addons controller-manager yet | +| E-01 … E-07 | The live lifecycle and the Ramen gate | Blocked on Phase 2 landing, plus a two-cluster test bed with Ramen dr-cluster installed for E-06 and E-07 | +| — | Repeated resync, class drift after verification, annotated-volume migration onto the adapter, cascaded topologies | Beyond the first coverage pass, recorded so the gaps are explicit rather than assumed covered | +| M-01, M-02 | Demote under writes; concurrent ownership race | Need failure injection and precise timing a live two-cluster run does not automate yet | diff --git a/operator/internal/controllers/driver/names.go b/operator/internal/controllers/driver/names.go index 10c08fdf1..3794b4c57 100644 --- a/operator/internal/controllers/driver/names.go +++ b/operator/internal/controllers/driver/names.go @@ -26,6 +26,13 @@ var clusterRoleComponents = []string{ // rather than the controller plugin's. const nodeComponent = "node" +// csiAddonsComponent is not a clusterRoleComponents entry: CSIAddonsNode is a +// namespaced kind (helm-charts' csiaddons.openshift.io_csiaddonsnodes.yaml +// sets scope: Namespaced), so its sidecar gets a namespaced Role instead of a +// sixth ClusterRole, per rbac-hardening's preference for the narrowest scope +// that works. See rbac.go's csiAddonsRoleRules. +const csiAddonsComponent = "csi-addons" + // objectNames is the whole naming surface of one deployment. type objectNames struct { prefix string @@ -81,6 +88,17 @@ func (n objectNames) clusterRoleBinding(component string) string { return n.prefix + component + "-binding" } +// role and roleBinding name the namespaced counterparts. Only csiAddonsComponent +// uses these today; the suffixes match clusterRole/clusterRoleBinding because +// the two are different Kinds and do not share a namespace with each other. +func (n objectNames) role(component string) string { + return n.prefix + component + "-role" +} + +func (n objectNames) roleBinding(component string) string { + return n.prefix + component + "-binding" +} + // driverName is spec.driverName with the CRD's default applied, so that code // reading it does not have to care whether admission had a chance to default it. func driverName(d *simplyblockv1alpha2.SimplyblockDriver) string { diff --git a/operator/internal/controllers/driver/rbac.go b/operator/internal/controllers/driver/rbac.go index 778099065..1b482f92f 100644 --- a/operator/internal/controllers/driver/rbac.go +++ b/operator/internal/controllers/driver/rbac.go @@ -1,5 +1,7 @@ -// The RBAC the two plugins need: a ServiceAccount each, and the five ClusterRole -// and ClusterRoleBinding pairs behind them. +// The RBAC the two plugins need: a ServiceAccount each, the five ClusterRole and +// ClusterRoleBinding pairs behind the sidecars that watch cluster-scoped kinds, +// and one namespaced Role and RoleBinding pair for the csi-addons sidecar, whose +// CSIAddonsNode is namespaced. // // The rules are the ones the chart applies, because adoption reconciles toward // the state that is running and a rule this file widens or narrows is a @@ -8,7 +10,7 @@ // runs a second set of its own and each set says which sidecar needs what. // // Specified by operator/docs/designs/crd-redesign/design-simplyblockdriver.md -// §4.1 and §4.3. +// §4.1 and §4.3, and operator/docs/designs/design-csi-addons-replication.md §4.1. package driver @@ -30,6 +32,8 @@ var ( storage = []string{"storage.k8s.io"} snapshot = []string{"snapshot.storage.k8s.io"} groupsnapshot = []string{"groupsnapshot.storage.k8s.io"} + csiaddons = []string{"csiaddons.openshift.io"} + coordination = []string{"coordination.k8s.io"} ) // clusterRoleRules is the rule set of each of the five roles, keyed by the @@ -87,6 +91,49 @@ var clusterRoleRules = map[string][]rbacv1.PolicyRule{ }, } +// csiAddonsRoleRules is the csi-addons sidecar's rule set, granted as a +// namespaced Role rather than added to clusterRoleRules: the sidecar only ever +// touches its own CSIAddonsNode and its own leader-election Lease, both in +// this deployment's namespace, and neither kind justifies a cluster-wide grant. +var csiAddonsRoleRules = []rbacv1.PolicyRule{ + // rbac-justified: the sidecar publishes and maintains exactly one + // CSIAddonsNode, naming itself, so the kubernetes-csi-addons + // controller-manager (design-csi-addons-replication.md §4.1) can find its + // endpoint. It does not read any other driver's CSIAddonsNode. + rule(csiaddons, []string{"csiaddonsnodes"}, "get", "list", "watch", "create", "update", "delete"), + rule(csiaddons, []string{"csiaddonsnodes/status"}, "get", "update", "patch"), + // rbac-justified: only one replica of the controller StatefulSet serves + // CONTROLLER_SERVICE requests at a time; the Lease is how the sidecar + // replicas elect that one, in this namespace only. + rule(coordination, []string{"leases"}, "get", "list", "watch", "create", "update", "delete"), + rule(core, []string{"events"}, "create", "patch"), +} + +func csiAddonsRole(d *simplyblockv1alpha2.SimplyblockDriver) *rbacv1.Role { + n := names(d) + return &rbacv1.Role{ + ObjectMeta: metav1.ObjectMeta{Name: n.role(csiAddonsComponent), Namespace: d.Namespace}, + Rules: csiAddonsRoleRules, + } +} + +func csiAddonsRoleBinding(d *simplyblockv1alpha2.SimplyblockDriver) *rbacv1.RoleBinding { + n := names(d) + return &rbacv1.RoleBinding{ + ObjectMeta: metav1.ObjectMeta{Name: n.roleBinding(csiAddonsComponent), Namespace: d.Namespace}, + Subjects: []rbacv1.Subject{{ + Kind: rbacv1.ServiceAccountKind, + Name: n.controllerServiceAccount, + Namespace: d.Namespace, + }}, + RoleRef: rbacv1.RoleRef{ + APIGroup: rbacv1.GroupName, + Kind: "Role", + Name: n.role(csiAddonsComponent), + }, + } +} + // serviceAccountFor names the account each role is bound to. The node plugin has // its own, and the controller plugin's sidecars share one. func serviceAccountFor(n objectNames, component string) string { diff --git a/operator/internal/controllers/driver/rbac_test.go b/operator/internal/controllers/driver/rbac_test.go index dbfb94cc4..b5dd58f10 100644 --- a/operator/internal/controllers/driver/rbac_test.go +++ b/operator/internal/controllers/driver/rbac_test.go @@ -154,6 +154,70 @@ func TestNoRuleIsAWildcard(t *testing.T) { } } +// The csi-addons sidecar is the one component whose grant is a namespaced Role: +// CSIAddonsNode is namespaced (helm-charts' +// csiaddons.openshift.io_csiaddonsnodes.yaml sets scope: Namespaced), so a +// ClusterRole would be wider than the sidecar's own job. +func TestCSIAddonsRoleIsNamespacedAndBoundToTheControllerAccount(t *testing.T) { + d := testDriver("simplyblock") + n := names(d) + + role := csiAddonsRole(d) + if role.Namespace != d.Namespace { + t.Errorf("role namespace = %q, want %q", role.Namespace, d.Namespace) + } + if len(role.Rules) == 0 { + t.Error("csi-addons role has no rules, so its sidecar can do nothing") + } + + binding := csiAddonsRoleBinding(d) + if binding.Namespace != d.Namespace { + t.Errorf("binding namespace = %q, want %q", binding.Namespace, d.Namespace) + } + if len(binding.Subjects) != 1 || binding.Subjects[0].Name != n.controllerServiceAccount || + binding.Subjects[0].Namespace != d.Namespace { + t.Errorf("binding subject = %+v, want the controller account in %q", + binding.Subjects, d.Namespace) + } + if binding.RoleRef.Kind != "Role" || binding.RoleRef.Name != role.Name { + t.Errorf("roleRef = %+v, want Role %q", binding.RoleRef, role.Name) + } +} + +// The rule set is exactly what the sidecar's own job needs: its CSIAddonsNode +// and its leader-election Lease, both scoped to this namespace. +func TestCSIAddonsRoleRulesAreScopedToItsOwnJob(t *testing.T) { + tests := []struct { + group string + resource string + verbs []string + }{ + {"csiaddons.openshift.io", "csiaddonsnodes", []string{"get", "list", "watch", "create", "update", "delete"}}, + {"csiaddons.openshift.io", "csiaddonsnodes/status", []string{"get", "update", "patch"}}, + {"coordination.k8s.io", "leases", []string{"get", "list", "watch", "create", "update", "delete"}}, + {"", "events", []string{"create", "patch"}}, + } + + for _, tc := range tests { + t.Run(tc.resource, func(t *testing.T) { + var found *rbacv1.PolicyRule + for i, r := range csiAddonsRoleRules { + if len(r.APIGroups) == 1 && r.APIGroups[0] == tc.group && + len(r.Resources) == 1 && r.Resources[0] == tc.resource { + found = &csiAddonsRoleRules[i] + break + } + } + if found == nil { + t.Fatalf("no rule for %s in the csi-addons role", tc.resource) + } + if !slices.Equal(found.Verbs, tc.verbs) { + t.Errorf("verbs = %v, want %v", found.Verbs, tc.verbs) + } + }) + } +} + // The csi-snapshotter sidecar watches VolumeGroupSnapshotContent and drives // the GroupController when the CSIVolumeGroupSnapshot gate is on // (design-consistency-groups.md §9, P0-4). The chart granted these on the diff --git a/operator/internal/controllers/driver/registration_test.go b/operator/internal/controllers/driver/registration_test.go index 500f79931..303606984 100644 --- a/operator/internal/controllers/driver/registration_test.go +++ b/operator/internal/controllers/driver/registration_test.go @@ -131,6 +131,7 @@ func TestSidecarsDefaultToTheOperatorsRelease(t *testing.T) { snapshotter: defaultSnapshotterImage, healthMonitor: defaultHealthMonitorImage, nodeDriverRegistrar: defaultNodeDriverRegistrarImage, + csiAddons: defaultCSIAddonsImage, } if got != want { t.Errorf("sidecars = %+v, want %+v", got, want) diff --git a/operator/internal/controllers/driver/sidecars.go b/operator/internal/controllers/driver/sidecars.go index 0c65dc244..3e4491abb 100644 --- a/operator/internal/controllers/driver/sidecars.go +++ b/operator/internal/controllers/driver/sidecars.go @@ -1,4 +1,4 @@ -// The six CSI sidecar images, and which of them a deployment gets. +// The seven CSI sidecar images, and which of them a deployment gets. // // The versions below are this operator release's, meaning the combination it was // tested against, and spec.sidecarImages overrides one at a time. The field @@ -25,6 +25,12 @@ const ( defaultSnapshotterImage = "quay.io/simplyblock-io/csi-snapshotter:v8.2.0" defaultHealthMonitorImage = "quay.io/simplyblock-io/csi-external-health-monitor-controller:v0.14.0" defaultNodeDriverRegistrarImage = "quay.io/simplyblock-io/csi-node-driver-registrar:v2.12.0" + // defaultCSIAddonsImage is the kubernetes-csi-addons sidecar (upstream + // quay.io/csiaddons/k8s-sidecar), pinned at the same v0.15.0 the chart's + // controller-manager runs (design P0-5). Named for the eventual + // quay.io/simplyblock-io mirror this field's validation pattern requires, + // which does not exist yet; mirroring it is a release task. + defaultCSIAddonsImage = "quay.io/simplyblock-io/csi-addons-sidecar:v0.15.0" ) // resolvedSidecars is the image each sidecar runs, after the overrides. @@ -35,6 +41,7 @@ type resolvedSidecars struct { snapshotter string healthMonitor string nodeDriverRegistrar string + csiAddons string } func sidecars(d *simplyblockv1alpha2.SimplyblockDriver) resolvedSidecars { @@ -46,6 +53,7 @@ func sidecars(d *simplyblockv1alpha2.SimplyblockDriver) resolvedSidecars { snapshotter: orDefault(o.Snapshotter, defaultSnapshotterImage), healthMonitor: orDefault(o.HealthMonitor, defaultHealthMonitorImage), nodeDriverRegistrar: orDefault(o.NodeDriverRegistrar, defaultNodeDriverRegistrarImage), + csiAddons: orDefault(o.CSIAddons, defaultCSIAddonsImage), } } diff --git a/operator/internal/controllers/driver/simplyblockdriver_controller.go b/operator/internal/controllers/driver/simplyblockdriver_controller.go index 1e12ec8cb..b0e269689 100644 --- a/operator/internal/controllers/driver/simplyblockdriver_controller.go +++ b/operator/internal/controllers/driver/simplyblockdriver_controller.go @@ -257,7 +257,7 @@ func (r *SimplyblockDriverReconciler) event( func (r *SimplyblockDriverReconciler) desired( d *simplyblockv1alpha2.SimplyblockDriver, image string, ) []client.Object { - objects := make([]client.Object, 0, 18) + objects := make([]client.Object, 0, 20) for _, sa := range serviceAccounts(d) { objects = append(objects, sa) @@ -271,6 +271,7 @@ func (r *SimplyblockDriverReconciler) desired( for _, crb := range clusterRoleBindings(d) { objects = append(objects, crb) } + objects = append(objects, csiAddonsRole(d), csiAddonsRoleBinding(d)) objects = append(objects, nodeDaemonSet(d, image), controllerStatefulSet(d, image), csiDriver(d)) if snapshotsEnabled(d) { objects = append(objects, volumeSnapshotClass(d)) diff --git a/operator/internal/controllers/driver/simplyblockdriver_controller_test.go b/operator/internal/controllers/driver/simplyblockdriver_controller_test.go index 14b9187f3..88d443567 100644 --- a/operator/internal/controllers/driver/simplyblockdriver_controller_test.go +++ b/operator/internal/controllers/driver/simplyblockdriver_controller_test.go @@ -59,6 +59,10 @@ func TestDesiredCoversTheWholeObjectSet(t *testing.T) { counts["role"]++ case *rbacv1.ClusterRoleBinding: counts["binding"]++ + case *rbacv1.Role: + counts["namespacedRole"]++ + case *rbacv1.RoleBinding: + counts["namespacedBinding"]++ case *appsv1.DaemonSet: counts["ds"]++ case *appsv1.StatefulSet: @@ -72,6 +76,7 @@ func TestDesiredCoversTheWholeObjectSet(t *testing.T) { want := map[string]int{ "sa": 2, "cm": 2, "role": 5, "binding": 5, + "namespacedRole": 1, "namespacedBinding": 1, "ds": 1, "sts": 1, "csidriver": 1, // the VolumeSnapshotClass, which is unstructured "other": 1, diff --git a/operator/internal/controllers/driver/workloads.go b/operator/internal/controllers/driver/workloads.go index 357675fbc..79f2cdbea 100644 --- a/operator/internal/controllers/driver/workloads.go +++ b/operator/internal/controllers/driver/workloads.go @@ -14,6 +14,8 @@ package driver import ( + "strconv" + appsv1 "k8s.io/api/apps/v1" corev1 "k8s.io/api/core/v1" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" @@ -247,6 +249,7 @@ func controllerStatefulSet(d *simplyblockv1alpha2.SimplyblockDriver, image strin "--leader-election=false", }), controllerPluginContainer(d, image), + csiAddonsSidecarContainer(d, s.csiAddons, sidecarMount), } // The snapshotter runs privileged in the chart, and the health monitor // exposes a port. Both are properties of the container rather than of the @@ -306,6 +309,56 @@ func controllerSidecar( } } +// csiAddonsControllerPort is where this sidecar's own gRPC server listens for +// the kubernetes-csi-addons controller-manager (design +// design-csi-addons-replication.md §4.1). The manager never hardcodes it: it +// reads the endpoint the sidecar published on its own CSIAddonsNode, so this +// number only has to be free and to agree with the container port below it. +const csiAddonsControllerPort int32 = 9070 + +// csiAddonsSidecarContainer runs the kubernetes-csi-addons sidecar (upstream +// quay.io/csiaddons/k8s-sidecar), which is a separate binary from the +// controller-manager P0-5 already vendored: it connects to this pod's own CSI +// socket to probe the csi-addons Identity/Replication services this driver +// now serves alongside its CSI ones (driver.go), publishes a CSIAddonsNode +// naming this pod so the manager can find it, and leader-elects across +// replicas of this StatefulSet before serving controller requests. +// +// The pod runs on the host network (controllerStatefulSet), so the sidecar +// advertises its own pod IP rather than a Service DNS name, the same way the +// existing csi-provisioner/-snapshotter/etc. sidecars address the plugin's +// socket by path instead of by name. +func csiAddonsSidecarContainer( + d *simplyblockv1alpha2.SimplyblockDriver, image string, mounts []corev1.VolumeMount, +) corev1.Container { + return corev1.Container{ + Name: "csi-addons", + Image: image, + ImagePullPolicy: pullPolicy(d), + Args: []string{ + verbosity, + "--csi-addons-address=" + controllerSocketPath, + "--controller-ip=$(POD_IP)", + "--controller-port=" + strconv.Itoa(int(csiAddonsControllerPort)), + "--pod=$(POD_NAME)", + "--namespace=$(POD_NAMESPACE)", + "--pod-uid=$(POD_UID)", + "--leader-election-namespace=$(POD_NAMESPACE)", + }, + Env: []corev1.EnvVar{ + fieldRefEnv("POD_IP", "status.podIP"), + fieldRefEnv("POD_NAME", "metadata.name"), + fieldRefEnv("POD_NAMESPACE", "metadata.namespace"), + fieldRefEnv("POD_UID", "metadata.uid"), + }, + Ports: []corev1.ContainerPort{ + {ContainerPort: csiAddonsControllerPort, Name: "csi-addons", Protocol: corev1.ProtocolTCP}, + }, + Resources: d.Spec.ControllerResources, + VolumeMounts: mounts, + } +} + func controllerPluginContainer(d *simplyblockv1alpha2.SimplyblockDriver, image string) corev1.Container { return corev1.Container{ Name: "csi-controller", diff --git a/operator/internal/controllers/driver/workloads_test.go b/operator/internal/controllers/driver/workloads_test.go index 2c95c801e..a2c8b53d5 100644 --- a/operator/internal/controllers/driver/workloads_test.go +++ b/operator/internal/controllers/driver/workloads_test.go @@ -95,6 +95,64 @@ func TestSnapshotterSidecarIsAppliedRegardlessOfTheToggle(t *testing.T) { } } +// The csi-addons sidecar is appended after the plugin container (index 6), +// never inserted: containerStatefulSet's positional tweaks (containers[1]'s +// SecurityContext, containers[4]'s Ports) only stay pointed at the snapshotter +// and health-monitor if nothing ahead of them shifts. +func TestCSIAddonsSidecarIsAppliedAfterThePlugin(t *testing.T) { + d := testDriver("simplyblock") + containers := controllerStatefulSet(d, testImage).Spec.Template.Spec.Containers + + if len(containers) != 7 { + t.Fatalf("got %d containers, want 7", len(containers)) + } + if containers[5].Name != "csi-controller" || containers[6].Name != "csi-addons" { + t.Errorf("containers[5:7] = %q, %q, want csi-controller, csi-addons", + containers[5].Name, containers[6].Name) + } + + addr, ok := argValue(&containers[6], "--csi-addons-address") + if !ok || addr != controllerSocketPath { + t.Errorf("csi-addons --csi-addons-address = %q, want %q", addr, controllerSocketPath) + } +} + +// The sidecar advertises this pod, not a Service DNS name: the StatefulSet runs +// on the host network, so its own pod IP is what the kubernetes-csi-addons +// controller-manager can actually reach. +func TestCSIAddonsSidecarAdvertisesItsOwnPod(t *testing.T) { + d := testDriver("simplyblock") + c := containerNamed(controllerStatefulSet(d, testImage).Spec.Template.Spec.Containers, "csi-addons") + if c == nil { + t.Fatal("csi-addons is not applied") + } + + wantFieldPaths := map[string]string{ + "POD_IP": "status.podIP", "POD_NAME": "metadata.name", + "POD_NAMESPACE": "metadata.namespace", "POD_UID": "metadata.uid", + } + for _, e := range c.Env { + want, known := wantFieldPaths[e.Name] + if !known { + continue + } + delete(wantFieldPaths, e.Name) + if e.ValueFrom == nil || e.ValueFrom.FieldRef == nil || e.ValueFrom.FieldRef.FieldPath != want { + t.Errorf("%s field path = %+v, want %q", e.Name, e.ValueFrom, want) + } + } + for name := range wantFieldPaths { + t.Errorf("no %s env var", name) + } + + if ip, ok := argValue(c, "--controller-ip"); !ok || ip != "$(POD_IP)" { + t.Errorf("--controller-ip = %q, want $(POD_IP)", ip) + } + if ns, ok := argValue(c, "--leader-election-namespace"); !ok || ns != "$(POD_NAMESPACE)" { + t.Errorf("--leader-election-namespace = %q, want $(POD_NAMESPACE)", ns) + } +} + // U-33 and U-34: the driver name reaches the kubelet registration path and the // hostPath the node plugin mounts, which is what design §9 Q3 says the chart // writes literally. diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_simplyblockdrivers.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_simplyblockdrivers.yaml index 82ab4ff4c..e5d075f08 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_simplyblockdrivers.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_simplyblockdrivers.yaml @@ -299,13 +299,23 @@ spec: type: object sidecarImages: description: |- - SidecarImages overrides the six CSI sidecars, one field each. Unset takes - the version this operator release ships. + SidecarImages overrides the seven CSI sidecars, one field each. Unset + takes the version this operator release ships. properties: attacher: description: Attacher is csi-attacher, on the controller plugin. pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ type: string + csiAddons: + description: |- + CSIAddons is the kubernetes-csi-addons sidecar, on the controller + plugin. It connects to the plugin's socket, probes the csi-addons + Identity service for capabilities, and publishes a CSIAddonsNode so the + kubernetes-csi-addons controller-manager (design + design-csi-addons-replication.md §4.1) can reach the Replication + service this driver serves. + pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ + type: string healthMonitor: description: |- HealthMonitor is csi-external-health-monitor-controller, on the diff --git a/shared/openapi.json b/shared/openapi.json index d3d3cfa81..11021dcee 100644 --- a/shared/openapi.json +++ b/shared/openapi.json @@ -775,6 +775,131 @@ } } }, + "/api/v2/clusters/{cluster_id}/alerts/": { + "get": { + "summary": "Clusters:Alerts:List", + "description": "The conditions in this cluster that currently need an operator.\n\nThis is not the event log. An alert appears only while it is still true\nand disappears on its own once it is not: the node comes back ONLINE, the\ndevice comes back, the cluster leaves degraded. Conditions an operator\ncaused on purpose -- a node they shut down, a device they removed -- are\nnot alerts and are not listed.\n\nBy default only what is wrong NOW is returned -- every entry has\n``status: firing``. Pass ``history=true`` to also get the ones that have\nsince resolved, each with its ``resolved_at``, or ``history_seconds=N``\nfor just the recent past. Either way both transitions are written to the\ncluster event log as ALERT_RAISED / ALERT_RESOLVED, so a resolution\nreaches an operator whether or not anyone asks for history here.\n\nCritical sorts before warning, and firing before resolved.", + "operationId": "clusters_alerts_list_api_v2_clusters__cluster_id__alerts__get", + "security": [ + { + "HTTPBearer": [] + } + ], + "parameters": [ + { + "name": "cluster_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Cluster Id" + } + }, + { + "name": "severity", + "in": "query", + "required": false, + "schema": { + "anyOf": [ + { + "enum": [ + "critical", + "warning" + ], + "type": "string" + }, + { + "type": "null" + } + ], + "description": "Only return alerts of this severity", + "title": "Severity" + }, + "description": "Only return alerts of this severity" + }, + { + "name": "history", + "in": "query", + "required": false, + "schema": { + "type": "boolean", + "description": "Also return alerts that have already resolved", + "default": false, + "title": "History" + }, + "description": "Also return alerts that have already resolved" + }, + { + "name": "history_seconds", + "in": "query", + "required": false, + "schema": { + "anyOf": [ + { + "type": "integer", + "minimum": 1 + }, + { + "type": "null" + } + ], + "description": "Limit the history to alerts resolved within this many seconds. Implies history=true.", + "title": "History Seconds" + }, + "description": "Limit the history to alerts resolved within this many seconds. Implies history=true." + }, + { + "name": "status", + "in": "query", + "required": false, + "schema": { + "anyOf": [ + { + "enum": [ + "firing", + "resolved" + ], + "type": "string" + }, + { + "type": "null" + } + ], + "description": "Only return alerts in this state", + "title": "Status" + }, + "description": "Only return alerts in this state" + } + ], + "responses": { + "200": { + "description": "Successful Response", + "content": { + "application/json": { + "schema": { + "type": "array", + "items": { + "$ref": "#/components/schemas/AlertDTO" + }, + "title": "Response Clusters Alerts List Api V2 Clusters Cluster Id Alerts Get" + } + } + } + }, + "422": { + "description": "Validation Error", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/HTTPValidationError" + } + } + } + } + } + } + }, "/api/v2/clusters/{cluster_id}/storage-nodes/": { "get": { "summary": "Clusters:Storage-Nodes:List", @@ -4178,6 +4303,75 @@ } } }, + "/api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/status": { + "get": { + "tags": [ + "replication" + ], + "summary": "Clusters:Storage-Pools:Volumes:Replication:Status", + "description": "The typed steady-state replication status.\n\nUnlike the relationship read above, which serves cutover records and 404s\nfor a volume's whole healthy replicated life, this endpoint always answers\nfor a volume that exists: ``state: not_replicating, role: none`` is the\nvalid answer for an unreplicated volume. The csi-addons adapter derives\nits conditions and ``lastSyncTime`` from this read on every reconcile.", + "operationId": "clusters_storage_pools_volumes_replication_status_api_v2_clusters__cluster_id__storage_pools__pool_id__volumes__volume_id__replication_status_get", + "security": [ + { + "HTTPBearer": [] + } + ], + "parameters": [ + { + "name": "cluster_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Cluster Id" + } + }, + { + "name": "pool_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Pool Id" + } + }, + { + "name": "volume_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Volume Id" + } + } + ], + "responses": { + "200": { + "description": "Successful Response", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/ReplicationStatusDTO" + } + } + } + }, + "422": { + "description": "Validation Error", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/HTTPValidationError" + } + } + } + } + } + } + }, "/api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/start": { "post": { "tags": [ @@ -4777,6 +4971,22 @@ "title": "Watch" }, "description": "Stream state changes as Server-Sent Events instead of returning a plain response: a `snapshot` event with the current state first, then `created`/`updated`/`deleted` events carrying the full resource representation. A `deleted` event carries the resource's final state when it is still retrievable (e.g. a volume whose status became `deleted`), or an empty object once it is gone entirely. Streams do not support resume; reconnecting clients receive a fresh snapshot. Changes written by pre-upgrade components may take up to 30 seconds to appear." + }, + { + "name": "consistency_group", + "in": "query", + "required": false, + "schema": { + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "title": "Consistency Group" + } } ], "responses": { @@ -6171,6 +6381,65 @@ } } }, + "/api/v2/clusters/{cluster_id}/replication/relationships/{lvol_id}/latest-snapshot": { + "get": { + "tags": [ + "replication" + ], + "summary": "Clusters:Replication:Relationships:Latest-Snapshot", + "description": "The volume's newest fully replicated snapshot, on the secondary, as a\ncloneable object. Exists for the volume's whole replicated life: a\ntest-failover drill (design \u00a714) resolves its test point through this\nread, without touching the real replication state to find out what it is.", + "operationId": "clusters_replication_relationships_latest_snapshot_api_v2_clusters__cluster_id__replication_relationships__lvol_id__latest_snapshot_get", + "security": [ + { + "HTTPBearer": [] + } + ], + "parameters": [ + { + "name": "lvol_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Lvol Id" + } + }, + { + "name": "cluster_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Cluster Id" + } + } + ], + "responses": { + "200": { + "description": "Successful Response", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/ReplicatedSnapshotDTO" + } + } + } + }, + "422": { + "description": "Validation Error", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/HTTPValidationError" + } + } + } + } + } + } + }, "/api/v2/clusters/{cluster_id}/replication/targets/": { "get": { "tags": [ @@ -6742,10 +7011,14 @@ } } }, - "/api/v2/management-nodes/": { + "/api/v2/clusters/{cluster_id}/replication/policies/{policy_id}/latest-generation": { "get": { - "summary": "Management Nodes:List", - "operationId": "management_nodes_list_api_v2_management_nodes__get", + "tags": [ + "replication" + ], + "summary": "Clusters:Replication:Policies:Latest-Generation", + "description": "The consistency group's newest fully replicated generation, every\nmember as a cloneable object on the secondary. Refused as a 400 when the\npolicy has no consistency group, when no generation is complete for\nevery current member yet, or when members are already split across\ngenerations: the same refusal a real group fail-over applies, so a drill\nnever addresses a mixed-generation cut.", + "operationId": "clusters_replication_policies_latest_generation_api_v2_clusters__cluster_id__replication_policies__policy_id__latest_generation_get", "security": [ { "HTTPBearer": [] @@ -6754,13 +7027,23 @@ "parameters": [ { "name": "cluster_id", - "in": "query", + "in": "path", "required": true, "schema": { "type": "string", "format": "uuid", "title": "Cluster Id" } + }, + { + "name": "policy_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Policy Id" + } } ], "responses": { @@ -6769,11 +7052,7 @@ "content": { "application/json": { "schema": { - "type": "array", - "items": { - "$ref": "#/components/schemas/ManagementNodeDTO" - }, - "title": "Response Management Nodes List Api V2 Management Nodes Get" + "$ref": "#/components/schemas/ReplicatedGenerationDTO" } } } @@ -6791,11 +7070,622 @@ } } }, - "/api/v2/management-nodes/{management_node_id}/": { + "/api/v2/clusters/{cluster_id}/consistency-groups/": { "get": { - "summary": "Management Node:Detail", - "operationId": "management_node_detail_api_v2_management_nodes__management_node_id___get", - "security": [ + "tags": [ + "consistency-groups" + ], + "summary": "Clusters:Consistency-Groups:List", + "description": "List the cluster's consistency groups, or resolve one by name (\u00a710).\n\nReturns an empty list when ``name`` matches no group, so a caller can probe\nexistence without a 404.", + "operationId": "clusters_consistency_groups_list_api_v2_clusters__cluster_id__consistency_groups__get", + "security": [ + { + "HTTPBearer": [] + } + ], + "parameters": [ + { + "name": "cluster_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Cluster Id" + } + }, + { + "name": "name", + "in": "query", + "required": false, + "schema": { + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "title": "Name" + } + } + ], + "responses": { + "200": { + "description": "Successful Response", + "content": { + "application/json": { + "schema": { + "type": "array", + "items": { + "$ref": "#/components/schemas/ConsistencyGroupDTO" + }, + "title": "Response Clusters Consistency Groups List Api V2 Clusters Cluster Id Consistency Groups Get" + } + } + } + }, + "422": { + "description": "Validation Error", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/HTTPValidationError" + } + } + } + } + } + } + }, + "/api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/": { + "get": { + "tags": [ + "consistency-groups" + ], + "summary": "Clusters:Consistency-Groups:Detail", + "operationId": "clusters_consistency_groups_detail_api_v2_clusters__cluster_id__consistency_groups__group_id___get", + "security": [ + { + "HTTPBearer": [] + } + ], + "parameters": [ + { + "name": "cluster_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Cluster Id" + } + }, + { + "name": "group_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Group Id" + } + } + ], + "responses": { + "200": { + "description": "Successful Response", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/ConsistencyGroupDTO" + } + } + } + }, + "422": { + "description": "Validation Error", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/HTTPValidationError" + } + } + } + } + } + } + }, + "/api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/members": { + "get": { + "tags": [ + "consistency-groups" + ], + "summary": "Clusters:Consistency-Groups:Members", + "operationId": "clusters_consistency_groups_members_api_v2_clusters__cluster_id__consistency_groups__group_id__members_get", + "security": [ + { + "HTTPBearer": [] + } + ], + "parameters": [ + { + "name": "cluster_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Cluster Id" + } + }, + { + "name": "group_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Group Id" + } + } + ], + "responses": { + "200": { + "description": "Successful Response", + "content": { + "application/json": { + "schema": { + "type": "array", + "items": { + "$ref": "#/components/schemas/ConsistencyGroupMemberDTO" + }, + "title": "Response Clusters Consistency Groups Members Api V2 Clusters Cluster Id Consistency Groups Group Id Members Get" + } + } + } + }, + "422": { + "description": "Validation Error", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/HTTPValidationError" + } + } + } + } + } + }, + "post": { + "tags": [ + "consistency-groups" + ], + "summary": "Clusters:Consistency-Groups:Members:Join", + "description": "Join an EXISTING volume to the group (design \u00a74.5, Phase 4 late join).\n\nValidates the pinned placement, the pool, the member cap, and the one-way\nrule; a refusal is a 409 naming the precondition. Idempotent: joining a\ncurrent member returns its membership row unchanged.", + "operationId": "clusters_consistency_groups_members_join_api_v2_clusters__cluster_id__consistency_groups__group_id__members_post", + "security": [ + { + "HTTPBearer": [] + } + ], + "parameters": [ + { + "name": "cluster_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Cluster Id" + } + }, + { + "name": "group_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Group Id" + } + } + ], + "requestBody": { + "required": true, + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/ConsistencyGroupMemberJoinDTO" + } + } + } + }, + "responses": { + "200": { + "description": "Successful Response", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/ConsistencyGroupMemberDTO" + } + } + } + }, + "422": { + "description": "Validation Error", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/HTTPValidationError" + } + } + } + } + } + } + }, + "/api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/members/{lvol_id}": { + "delete": { + "tags": [ + "consistency-groups" + ], + "summary": "Clusters:Consistency-Groups:Members:Detach", + "description": "Detach a member: close its epoch one-way, preserving prior generations (\u00a78.2).", + "operationId": "clusters_consistency_groups_members_detach_api_v2_clusters__cluster_id__consistency_groups__group_id__members__lvol_id__delete", + "security": [ + { + "HTTPBearer": [] + } + ], + "parameters": [ + { + "name": "lvol_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "title": "Lvol Id" + } + }, + { + "name": "cluster_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Cluster Id" + } + }, + { + "name": "group_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Group Id" + } + } + ], + "responses": { + "204": { + "description": "Successful Response" + }, + "422": { + "description": "Validation Error", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/HTTPValidationError" + } + } + } + } + } + } + }, + "/api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/snapshots": { + "get": { + "tags": [ + "consistency-groups" + ], + "summary": "Clusters:Consistency-Groups:Snapshots:List", + "operationId": "clusters_consistency_groups_snapshots_list_api_v2_clusters__cluster_id__consistency_groups__group_id__snapshots_get", + "security": [ + { + "HTTPBearer": [] + } + ], + "parameters": [ + { + "name": "cluster_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Cluster Id" + } + }, + { + "name": "group_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Group Id" + } + } + ], + "responses": { + "200": { + "description": "Successful Response", + "content": { + "application/json": { + "schema": { + "type": "array", + "items": { + "$ref": "#/components/schemas/ConsistencyGroupGenerationDTO" + }, + "title": "Response Clusters Consistency Groups Snapshots List Api V2 Clusters Cluster Id Consistency Groups Group Id Snapshots Get" + } + } + } + }, + "422": { + "description": "Validation Error", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/HTTPValidationError" + } + } + } + } + } + }, + "post": { + "tags": [ + "consistency-groups" + ], + "summary": "Clusters:Consistency-Groups:Snapshots:Take", + "description": "Take one crash-consistent generation across every current member (\u00a75).", + "operationId": "clusters_consistency_groups_snapshots_take_api_v2_clusters__cluster_id__consistency_groups__group_id__snapshots_post", + "security": [ + { + "HTTPBearer": [] + } + ], + "parameters": [ + { + "name": "cluster_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Cluster Id" + } + }, + { + "name": "group_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Group Id" + } + } + ], + "responses": { + "200": { + "description": "Successful Response", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/ConsistencyGroupGenerationDTO" + } + } + } + }, + "422": { + "description": "Validation Error", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/HTTPValidationError" + } + } + } + } + } + } + }, + "/api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/snapshots/{seq}": { + "get": { + "tags": [ + "consistency-groups" + ], + "summary": "Clusters:Consistency-Groups:Snapshots:Detail", + "operationId": "clusters_consistency_groups_snapshots_detail_api_v2_clusters__cluster_id__consistency_groups__group_id__snapshots__seq__get", + "security": [ + { + "HTTPBearer": [] + } + ], + "parameters": [ + { + "name": "seq", + "in": "path", + "required": true, + "schema": { + "type": "integer", + "title": "Seq" + } + }, + { + "name": "cluster_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Cluster Id" + } + }, + { + "name": "group_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Group Id" + } + } + ], + "responses": { + "200": { + "description": "Successful Response", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/ConsistencyGroupGenerationDTO" + } + } + } + }, + "422": { + "description": "Validation Error", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/HTTPValidationError" + } + } + } + } + } + }, + "delete": { + "tags": [ + "consistency-groups" + ], + "summary": "Clusters:Consistency-Groups:Snapshots:Delete", + "description": "Delete one generation and all its member snapshots; never the group (\u00a710).", + "operationId": "clusters_consistency_groups_snapshots_delete_api_v2_clusters__cluster_id__consistency_groups__group_id__snapshots__seq__delete", + "security": [ + { + "HTTPBearer": [] + } + ], + "parameters": [ + { + "name": "seq", + "in": "path", + "required": true, + "schema": { + "type": "integer", + "title": "Seq" + } + }, + { + "name": "cluster_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Cluster Id" + } + }, + { + "name": "group_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Group Id" + } + } + ], + "responses": { + "204": { + "description": "Successful Response" + }, + "422": { + "description": "Validation Error", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/HTTPValidationError" + } + } + } + } + } + } + }, + "/api/v2/management-nodes/": { + "get": { + "summary": "Management Nodes:List", + "operationId": "management_nodes_list_api_v2_management_nodes__get", + "security": [ + { + "HTTPBearer": [] + } + ], + "parameters": [ + { + "name": "cluster_id", + "in": "query", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Cluster Id" + } + } + ], + "responses": { + "200": { + "description": "Successful Response", + "content": { + "application/json": { + "schema": { + "type": "array", + "items": { + "$ref": "#/components/schemas/ManagementNodeDTO" + }, + "title": "Response Management Nodes List Api V2 Management Nodes Get" + } + } + } + }, + "422": { + "description": "Validation Error", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/HTTPValidationError" + } + } + } + } + } + } + }, + "/api/v2/management-nodes/{management_node_id}/": { + "get": { + "summary": "Management Node:Detail", + "operationId": "management_node_detail_api_v2_management_nodes__management_node_id___get", + "security": [ { "HTTPBearer": [] } @@ -6873,6 +7763,122 @@ }, "components": { "schemas": { + "AlertDTO": { + "properties": { + "id": { + "type": "string", + "title": "Id" + }, + "kind": { + "type": "string", + "title": "Kind" + }, + "severity": { + "type": "string", + "enum": [ + "critical", + "warning" + ], + "title": "Severity" + }, + "status": { + "type": "string", + "enum": [ + "firing", + "resolved" + ], + "title": "Status" + }, + "message": { + "type": "string", + "title": "Message" + }, + "cluster_id": { + "type": "string", + "format": "uuid", + "title": "Cluster Id" + }, + "node_id": { + "anyOf": [ + { + "type": "string", + "format": "uuid" + }, + { + "type": "null" + } + ], + "title": "Node Id" + }, + "device_id": { + "anyOf": [ + { + "type": "string", + "format": "uuid" + }, + { + "type": "null" + } + ], + "title": "Device Id" + }, + "since": { + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "title": "Since" + }, + "first_seen": { + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "title": "First Seen" + }, + "resolved_at": { + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "title": "Resolved At" + }, + "details": { + "additionalProperties": true, + "type": "object", + "title": "Details" + } + }, + "type": "object", + "required": [ + "id", + "kind", + "severity", + "status", + "message", + "cluster_id", + "node_id", + "device_id", + "since", + "first_seen", + "resolved_at", + "details" + ], + "title": "AlertDTO", + "description": "One condition that currently needs an operator.\n\nDeliberately NOT an EventObj. An event is a journal entry -- it happened,\nit is kept forever, and nothing ever retracts it. An alert is a claim\nabout the present that goes away by itself when it stops being true, so\nit carries the object it is about and the time the condition started\nrather than the time something was logged. ``id`` is derived from the\nkind and the object, so it is stable across polls and a consumer can\ndedupe on it without keeping state." + }, "BackupConfigParams": { "properties": { "access_key_id": { @@ -7614,52 +8620,223 @@ }, "hugepages_mem": { "type": "integer", - "minimum": 0.0, - "title": "Hugepages Mem", - "default": 0 + "minimum": 0.0, + "title": "Hugepages Mem", + "default": 0 + }, + "spdk_vcpu_count": { + "type": "integer", + "minimum": 0.0, + "title": "Spdk Vcpu Count" + }, + "device_mode": { + "type": "string", + "enum": [ + "nvme", + "lblk" + ], + "title": "Device Mode", + "default": "nvme" + }, + "inline_checksum": { + "type": "boolean", + "title": "Inline Checksum", + "default": false + }, + "atomic_4k": { + "type": "boolean", + "title": "Atomic 4K", + "default": false + } + }, + "type": "object", + "required": [ + "max_subsys", + "spdk_vcpu_count" + ], + "title": "ClusterParams" + }, + "CommitParams": { + "properties": { + "delete_source": { + "type": "boolean", + "title": "Delete Source", + "default": false + } + }, + "type": "object", + "title": "CommitParams" + }, + "ConsistencyGroupDTO": { + "properties": { + "id": { + "type": "string", + "format": "uuid", + "title": "Id" + }, + "cluster_id": { + "type": "string", + "format": "uuid", + "title": "Cluster Id" + }, + "name": { + "type": "string", + "title": "Name" + }, + "node_id": { + "anyOf": [ + { + "type": "string", + "format": "uuid" + }, + { + "type": "null" + } + ], + "title": "Node Id" + }, + "lvs_name": { + "type": "string", + "title": "Lvs Name", + "default": "" + }, + "member_count": { + "type": "integer", + "title": "Member Count" + }, + "last_group_seq": { + "type": "integer", + "title": "Last Group Seq" + } + }, + "type": "object", + "required": [ + "id", + "cluster_id", + "name", + "member_count", + "last_group_seq" + ], + "title": "ConsistencyGroupDTO", + "description": "A standalone consistency group summary (design \u00a710)." + }, + "ConsistencyGroupGenerationDTO": { + "properties": { + "group_seq": { + "type": "integer", + "title": "Group Seq" + }, + "created_at": { + "type": "integer", + "title": "Created At" + }, + "expected": { + "type": "integer", + "title": "Expected" + }, + "present": { + "type": "integer", + "title": "Present" + }, + "complete": { + "type": "boolean", + "title": "Complete" + }, + "members": { + "items": { + "$ref": "#/components/schemas/ConsistencyGroupGenerationMemberDTO" + }, + "type": "array", + "title": "Members" + } + }, + "type": "object", + "required": [ + "group_seq", + "created_at", + "expected", + "present", + "complete", + "members" + ], + "title": "ConsistencyGroupGenerationDTO", + "description": "One generation of a consistency group (design \u00a76.3)." + }, + "ConsistencyGroupGenerationMemberDTO": { + "properties": { + "lvol_id": { + "type": "string", + "title": "Lvol Id" + }, + "snapshot_id": { + "type": "string", + "title": "Snapshot Id" + }, + "ready": { + "type": "boolean", + "title": "Ready" + } + }, + "type": "object", + "required": [ + "lvol_id", + "snapshot_id", + "ready" + ], + "title": "ConsistencyGroupGenerationMemberDTO" + }, + "ConsistencyGroupMemberDTO": { + "properties": { + "lvol_id": { + "type": "string", + "title": "Lvol Id" + }, + "joined_seq": { + "type": "integer", + "title": "Joined Seq" }, - "spdk_vcpu_count": { + "removed_seq": { "type": "integer", - "minimum": 0.0, - "title": "Spdk Vcpu Count" + "title": "Removed Seq" }, - "device_mode": { + "node_id": { "type": "string", - "enum": [ - "nvme", - "lblk" - ], - "title": "Device Mode", - "default": "nvme" + "title": "Node Id" }, - "inline_checksum": { - "type": "boolean", - "title": "Inline Checksum", - "default": false + "lvs_name": { + "type": "string", + "title": "Lvs Name" }, - "atomic_4k": { + "online": { "type": "boolean", - "title": "Atomic 4K", - "default": false + "title": "Online" } }, "type": "object", "required": [ - "max_subsys", - "spdk_vcpu_count" + "lvol_id", + "joined_seq", + "removed_seq", + "node_id", + "lvs_name", + "online" ], - "title": "ClusterParams" + "title": "ConsistencyGroupMemberDTO", + "description": "One current member of a consistency group (design \u00a710 /members)." }, - "CommitParams": { + "ConsistencyGroupMemberJoinDTO": { "properties": { - "delete_source": { - "type": "boolean", - "title": "Delete Source", - "default": false + "lvol_id": { + "type": "string", + "title": "Lvol Id" } }, "type": "object", - "title": "CommitParams" + "required": [ + "lvol_id" + ], + "title": "ConsistencyGroupMemberJoinDTO", + "description": "Request body for the late join of an existing volume (design \u00a74.5)." }, "DeviceDTO": { "properties": { @@ -8286,6 +9463,18 @@ ], "title": "Keep Replicated" }, + "rpo_target_seconds": { + "anyOf": [ + { + "type": "integer", + "minimum": 0.0 + }, + { + "type": "null" + } + ], + "title": "Rpo Target Seconds" + }, "consistency_group": { "type": "boolean", "title": "Consistency Group", @@ -8327,6 +9516,103 @@ ], "title": "ReplicateLVolParams" }, + "ReplicatedGenerationDTO": { + "properties": { + "group_seq": { + "type": "integer", + "minimum": 0.0, + "title": "Group Seq" + }, + "members": { + "items": { + "$ref": "#/components/schemas/ReplicatedSnapshotDTO" + }, + "type": "array", + "title": "Members" + } + }, + "type": "object", + "required": [ + "group_seq", + "members" + ], + "title": "ReplicatedGenerationDTO", + "description": "One complete, fully replicated consistency-group generation, every\nmember addressed as a cloneable object on the secondary." + }, + "ReplicatedSnapshotDTO": { + "properties": { + "snapshot_id": { + "type": "string", + "format": "uuid", + "title": "Snapshot Id" + }, + "cluster_id": { + "type": "string", + "format": "uuid", + "title": "Cluster Id" + }, + "pool_id": { + "anyOf": [ + { + "type": "string", + "format": "uuid" + }, + { + "type": "null" + } + ], + "title": "Pool Id" + }, + "lvol_id": { + "anyOf": [ + { + "type": "string", + "format": "uuid" + }, + { + "type": "null" + } + ], + "title": "Lvol Id" + }, + "size": { + "type": "integer", + "minimum": 0.0, + "title": "Size" + }, + "used_size": { + "type": "integer", + "minimum": 0.0, + "title": "Used Size" + }, + "created_at": { + "type": "string", + "format": "date-time", + "title": "Created At" + }, + "group_id": { + "type": "string", + "title": "Group Id", + "default": "" + }, + "group_seq": { + "type": "integer", + "minimum": 0.0, + "title": "Group Seq", + "default": 0 + } + }, + "type": "object", + "required": [ + "snapshot_id", + "cluster_id", + "size", + "used_size", + "created_at" + ], + "title": "ReplicatedSnapshotDTO", + "description": "A fully replicated snapshot on the secondary, addressed as a cloneable\nobject. ``lvol_id`` is the volume the snapshot belongs to on the\nSECONDARY cluster, not the source volume the caller asked about, because\nthat is the identity the ordinary CSI clone path resolves a\n``dataSource`` against." + }, "ReplicationPolicyDTO": { "properties": { "id": { @@ -8373,6 +9659,18 @@ ], "title": "Status" }, + "rpo_target_seconds": { + "anyOf": [ + { + "type": "integer", + "minimum": 0.0 + }, + { + "type": "null" + } + ], + "title": "Rpo Target Seconds" + }, "consistency_group": { "type": "boolean", "title": "Consistency Group", @@ -8604,6 +9902,127 @@ "type": "object", "title": "ReplicationStartParams" }, + "ReplicationStatusDTO": { + "properties": { + "role": { + "type": "string", + "enum": [ + "source", + "secondary", + "failed_over", + "none" + ], + "title": "Role" + }, + "state": { + "type": "string", + "enum": [ + "in_sync", + "replicating", + "lagging", + "degraded", + "error", + "not_replicating" + ], + "title": "State" + }, + "last_replicated_at": { + "anyOf": [ + { + "type": "string", + "format": "date-time" + }, + { + "type": "null" + } + ], + "title": "Last Replicated At" + }, + "lag_seconds": { + "anyOf": [ + { + "type": "integer", + "minimum": 0.0 + }, + { + "type": "null" + } + ], + "title": "Lag Seconds" + }, + "lag_budget_seconds": { + "anyOf": [ + { + "type": "integer", + "minimum": 0.0 + }, + { + "type": "null" + } + ], + "title": "Lag Budget Seconds" + }, + "outstanding_count": { + "type": "integer", + "minimum": 0.0, + "title": "Outstanding Count", + "default": 0 + }, + "outstanding_bytes": { + "type": "integer", + "minimum": 0.0, + "title": "Outstanding Bytes", + "default": 0 + }, + "failing_count": { + "type": "integer", + "minimum": 0.0, + "title": "Failing Count", + "default": 0 + }, + "max_retry_reached": { + "type": "boolean", + "title": "Max Retry Reached", + "default": false + }, + "last_cycle_bytes": { + "anyOf": [ + { + "type": "integer", + "minimum": 0.0 + }, + { + "type": "null" + } + ], + "title": "Last Cycle Bytes" + }, + "last_cycle_seconds": { + "anyOf": [ + { + "type": "integer", + "minimum": 0.0 + }, + { + "type": "null" + } + ], + "title": "Last Cycle Seconds" + }, + "resyncing": { + "type": "boolean", + "title": "Resyncing", + "default": false + } + }, + "type": "object", + "required": [ + "role", + "state" + ], + "title": "ReplicationStatusDTO", + "description": "The typed steady-state replication status of one volume.\n\nServes what ``lvol_controller.get_replication_info`` computes, for the\nvolume's WHOLE replicated life \u2014 unlike ``ReplicationRelationshipDTO``,\nwhich only exists once a cutover or fail-over has created a relationship\nrecord. ``state: not_replicating, role: none`` is a valid answer, never a\n404, because the csi-addons adapter polls this on every reconcile." + }, "ReplicationTargetDTO": { "properties": { "id": { @@ -8722,6 +10141,14 @@ "type": "string", "format": "date-time", "title": "Created At" + }, + "group_id": { + "type": "string", + "title": "Group Id" + }, + "group_seq": { + "type": "integer", + "title": "Group Seq" } }, "type": "object", @@ -8734,7 +10161,9 @@ "used_size", "migrating", "lvol", - "created_at" + "created_at", + "group_id", + "group_seq" ], "title": "SnapshotDTO" }, @@ -9872,6 +11301,16 @@ "type": "boolean", "title": "From Source", "default": true + }, + "group_id": { + "type": "string", + "title": "Group Id", + "default": "" + }, + "group_seq": { + "type": "integer", + "title": "Group Seq", + "default": 0 } }, "type": "object", @@ -10018,6 +11457,17 @@ "type": "boolean", "title": "Delete Snap On Lvol Delete", "default": false + }, + "consistency_group": { + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "title": "Consistency Group" } }, "type": "object", @@ -10205,6 +11655,17 @@ ], "title": "Replication Policy" }, + "consistency_group": { + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "title": "Consistency Group" + }, "encrypt": { "type": "boolean", "title": "Encrypt", diff --git a/test/integration/controlplane/cpsim.gen.go b/test/integration/controlplane/cpsim.gen.go index a3e8650fa..198045f1a 100644 --- a/test/integration/controlplane/cpsim.gen.go +++ b/test/integration/controlplane/cpsim.gen.go @@ -16,6 +16,42 @@ import ( openapi_types "github.com/oapi-codegen/runtime/types" ) +// Defines values for AlertDTOSeverity. +const ( + AlertDTOSeverityCritical AlertDTOSeverity = "critical" + AlertDTOSeverityWarning AlertDTOSeverity = "warning" +) + +// Valid indicates whether the value is a known member of the AlertDTOSeverity enum. +func (e AlertDTOSeverity) Valid() bool { + switch e { + case AlertDTOSeverityCritical: + return true + case AlertDTOSeverityWarning: + return true + default: + return false + } +} + +// Defines values for AlertDTOStatus. +const ( + AlertDTOStatusFiring AlertDTOStatus = "firing" + AlertDTOStatusResolved AlertDTOStatus = "resolved" +) + +// Valid indicates whether the value is a known member of the AlertDTOStatus enum. +func (e AlertDTOStatus) Valid() bool { + switch e { + case AlertDTOStatusFiring: + return true + case AlertDTOStatusResolved: + return true + default: + return false + } +} + // Defines values for ClusterDTOStatus. const ( ClusterDTOStatusActive ClusterDTOStatus = "active" @@ -259,6 +295,60 @@ func (e ReplicationStartParamsMode) Valid() bool { } } +// Defines values for ReplicationStatusDTORole. +const ( + ReplicationStatusDTORoleFailedOver ReplicationStatusDTORole = "failed_over" + ReplicationStatusDTORoleNone ReplicationStatusDTORole = "none" + ReplicationStatusDTORoleSecondary ReplicationStatusDTORole = "secondary" + ReplicationStatusDTORoleSource ReplicationStatusDTORole = "source" +) + +// Valid indicates whether the value is a known member of the ReplicationStatusDTORole enum. +func (e ReplicationStatusDTORole) Valid() bool { + switch e { + case ReplicationStatusDTORoleFailedOver: + return true + case ReplicationStatusDTORoleNone: + return true + case ReplicationStatusDTORoleSecondary: + return true + case ReplicationStatusDTORoleSource: + return true + default: + return false + } +} + +// Defines values for ReplicationStatusDTOState. +const ( + ReplicationStatusDTOStateDegraded ReplicationStatusDTOState = "degraded" + ReplicationStatusDTOStateError ReplicationStatusDTOState = "error" + ReplicationStatusDTOStateInSync ReplicationStatusDTOState = "in_sync" + ReplicationStatusDTOStateLagging ReplicationStatusDTOState = "lagging" + ReplicationStatusDTOStateNotReplicating ReplicationStatusDTOState = "not_replicating" + ReplicationStatusDTOStateReplicating ReplicationStatusDTOState = "replicating" +) + +// Valid indicates whether the value is a known member of the ReplicationStatusDTOState enum. +func (e ReplicationStatusDTOState) Valid() bool { + switch e { + case ReplicationStatusDTOStateDegraded: + return true + case ReplicationStatusDTOStateError: + return true + case ReplicationStatusDTOStateInSync: + return true + case ReplicationStatusDTOStateLagging: + return true + case ReplicationStatusDTOStateNotReplicating: + return true + case ReplicationStatusDTOStateReplicating: + return true + default: + return false + } +} + // Defines values for ReplicationTargetDTOStatus. const ( ReplicationTargetDTOStatusActive ReplicationTargetDTOStatus = "active" @@ -487,6 +577,42 @@ func (e ClustersCreateApiV2ClustersPostParamsResponseFormat) Valid() bool { } } +// Defines values for ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsSeverity. +const ( + ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsSeverityCritical ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsSeverity = "critical" + ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsSeverityWarning ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsSeverity = "warning" +) + +// Valid indicates whether the value is a known member of the ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsSeverity enum. +func (e ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsSeverity) Valid() bool { + switch e { + case ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsSeverityCritical: + return true + case ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsSeverityWarning: + return true + default: + return false + } +} + +// Defines values for ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsStatus. +const ( + ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsStatusFiring ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsStatus = "firing" + ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsStatusResolved ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsStatus = "resolved" +) + +// Valid indicates whether the value is a known member of the ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsStatus enum. +func (e ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsStatus) Valid() bool { + switch e { + case ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsStatusFiring: + return true + case ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsStatusResolved: + return true + default: + return false + } +} + // Defines values for ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostParamsResponseFormat. const ( ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostParamsResponseFormatEmpty ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostParamsResponseFormat = "empty" @@ -634,6 +760,36 @@ func (e ClustersSubsystemsMigrationsCreateApiV2ClustersClusterIdSubsystemsNqnMig } } +// AlertDTO One condition that currently needs an operator. +// +// Deliberately NOT an EventObj. An event is a journal entry -- it happened, +// it is kept forever, and nothing ever retracts it. An alert is a claim +// about the present that goes away by itself when it stops being true, so +// it carries the object it is about and the time the condition started +// rather than the time something was logged. “id“ is derived from the +// kind and the object, so it is stable across polls and a consumer can +// dedupe on it without keeping state. +type AlertDTO struct { + ClusterId openapi_types.UUID `json:"cluster_id"` + Details map[string]interface{} `json:"details"` + DeviceId *openapi_types.UUID `json:"device_id"` + FirstSeen *string `json:"first_seen"` + Id string `json:"id"` + Kind string `json:"kind"` + Message string `json:"message"` + NodeId *openapi_types.UUID `json:"node_id"` + ResolvedAt *string `json:"resolved_at"` + Severity AlertDTOSeverity `json:"severity"` + Since *string `json:"since"` + Status AlertDTOStatus `json:"status"` +} + +// AlertDTOSeverity defines model for AlertDTO.Severity. +type AlertDTOSeverity string + +// AlertDTOStatus defines model for AlertDTO.Status. +type AlertDTOStatus string + // BackupConfigParams defines model for BackupConfigParams. type BackupConfigParams struct { AccessKeyId *string `json:"access_key_id,omitempty"` @@ -797,6 +953,49 @@ type CommitParams struct { DeleteSource *bool `json:"delete_source,omitempty"` } +// ConsistencyGroupDTO A standalone consistency group summary (design §10). +type ConsistencyGroupDTO struct { + ClusterId openapi_types.UUID `json:"cluster_id"` + Id openapi_types.UUID `json:"id"` + LastGroupSeq int `json:"last_group_seq"` + LvsName *string `json:"lvs_name,omitempty"` + MemberCount int `json:"member_count"` + Name string `json:"name"` + NodeId *openapi_types.UUID `json:"node_id,omitempty"` +} + +// ConsistencyGroupGenerationDTO One generation of a consistency group (design §6.3). +type ConsistencyGroupGenerationDTO struct { + Complete bool `json:"complete"` + CreatedAt int `json:"created_at"` + Expected int `json:"expected"` + GroupSeq int `json:"group_seq"` + Members []ConsistencyGroupGenerationMemberDTO `json:"members"` + Present int `json:"present"` +} + +// ConsistencyGroupGenerationMemberDTO defines model for ConsistencyGroupGenerationMemberDTO. +type ConsistencyGroupGenerationMemberDTO struct { + LvolId string `json:"lvol_id"` + Ready bool `json:"ready"` + SnapshotId string `json:"snapshot_id"` +} + +// ConsistencyGroupMemberDTO One current member of a consistency group (design §10 /members). +type ConsistencyGroupMemberDTO struct { + JoinedSeq int `json:"joined_seq"` + LvolId string `json:"lvol_id"` + LvsName string `json:"lvs_name"` + NodeId string `json:"node_id"` + Online bool `json:"online"` + RemovedSeq int `json:"removed_seq"` +} + +// ConsistencyGroupMemberJoinDTO Request body for the late join of an existing volume (design §4.5). +type ConsistencyGroupMemberJoinDTO struct { + LvolId string `json:"lvol_id"` +} + // DeviceDTO defines model for DeviceDTO. type DeviceDTO struct { BdevType *string `json:"bdev_type,omitempty"` @@ -931,6 +1130,7 @@ type PolicyParams struct { KeepReplicated *int `json:"keep_replicated,omitempty"` Mode *PolicyParamsMode `json:"mode,omitempty"` PolicyName string `json:"policy_name"` + RpoTargetSeconds *int `json:"rpo_target_seconds,omitempty"` TargetId openapi_types.UUID `json:"target_id"` } @@ -947,6 +1147,30 @@ type ReplicateLVolParams struct { LvolId openapi_types.UUID `json:"lvol_id"` } +// ReplicatedGenerationDTO One complete, fully replicated consistency-group generation, every +// member addressed as a cloneable object on the secondary. +type ReplicatedGenerationDTO struct { + GroupSeq int `json:"group_seq"` + Members []ReplicatedSnapshotDTO `json:"members"` +} + +// ReplicatedSnapshotDTO A fully replicated snapshot on the secondary, addressed as a cloneable +// object. “lvol_id“ is the volume the snapshot belongs to on the +// SECONDARY cluster, not the source volume the caller asked about, because +// that is the identity the ordinary CSI clone path resolves a +// “dataSource“ against. +type ReplicatedSnapshotDTO struct { + ClusterId openapi_types.UUID `json:"cluster_id"` + CreatedAt time.Time `json:"created_at"` + GroupId *string `json:"group_id,omitempty"` + GroupSeq *int `json:"group_seq,omitempty"` + LvolId *openapi_types.UUID `json:"lvol_id,omitempty"` + PoolId *openapi_types.UUID `json:"pool_id,omitempty"` + Size int `json:"size"` + SnapshotId openapi_types.UUID `json:"snapshot_id"` + UsedSize int `json:"used_size"` +} + // ReplicationPolicyDTO defines model for ReplicationPolicyDTO. type ReplicationPolicyDTO struct { ClusterId openapi_types.UUID `json:"cluster_id"` @@ -959,6 +1183,7 @@ type ReplicationPolicyDTO struct { KeepReplicated int `json:"keep_replicated"` Mode ReplicationPolicyDTOMode `json:"mode"` PolicyName string `json:"policy_name"` + RpoTargetSeconds *int `json:"rpo_target_seconds,omitempty"` Status ReplicationPolicyDTOStatus `json:"status"` TargetId openapi_types.UUID `json:"target_id"` } @@ -1006,6 +1231,34 @@ type ReplicationStartParams struct { // ReplicationStartParamsMode defines model for ReplicationStartParams.Mode. type ReplicationStartParamsMode string +// ReplicationStatusDTO The typed steady-state replication status of one volume. +// +// Serves what “lvol_controller.get_replication_info“ computes, for the +// volume's WHOLE replicated life — unlike “ReplicationRelationshipDTO“, +// which only exists once a cutover or fail-over has created a relationship +// record. “state: not_replicating, role: none“ is a valid answer, never a +// 404, because the csi-addons adapter polls this on every reconcile. +type ReplicationStatusDTO struct { + FailingCount *int `json:"failing_count,omitempty"` + LagBudgetSeconds *int `json:"lag_budget_seconds,omitempty"` + LagSeconds *int `json:"lag_seconds,omitempty"` + LastCycleBytes *int `json:"last_cycle_bytes,omitempty"` + LastCycleSeconds *int `json:"last_cycle_seconds,omitempty"` + LastReplicatedAt *time.Time `json:"last_replicated_at,omitempty"` + MaxRetryReached *bool `json:"max_retry_reached,omitempty"` + OutstandingBytes *int `json:"outstanding_bytes,omitempty"` + OutstandingCount *int `json:"outstanding_count,omitempty"` + Resyncing *bool `json:"resyncing,omitempty"` + Role ReplicationStatusDTORole `json:"role"` + State ReplicationStatusDTOState `json:"state"` +} + +// ReplicationStatusDTORole defines model for ReplicationStatusDTO.Role. +type ReplicationStatusDTORole string + +// ReplicationStatusDTOState defines model for ReplicationStatusDTO.State. +type ReplicationStatusDTOState string + // ReplicationTargetDTO defines model for ReplicationTargetDTO. type ReplicationTargetDTO struct { ClusterId openapi_types.UUID `json:"cluster_id"` @@ -1028,6 +1281,8 @@ type RootModelUnionCreateParamsCloneParams struct { // SnapshotDTO defines model for SnapshotDTO. type SnapshotDTO struct { CreatedAt time.Time `json:"created_at"` + GroupId string `json:"group_id"` + GroupSeq int `json:"group_seq"` HealthCheck bool `json:"health_check"` Id openapi_types.UUID `json:"id"` Lvol *string `json:"lvol"` @@ -1223,6 +1478,8 @@ type VolumeDTO struct { DoReplicate *bool `json:"do_replicate,omitempty"` Fabric string `json:"fabric"` FromSource *bool `json:"from_source,omitempty"` + GroupId *string `json:"group_id,omitempty"` + GroupSeq *int `json:"group_seq,omitempty"` HealthCheck bool `json:"health_check"` HighAvailability bool `json:"high_availability"` Hostname string `json:"hostname"` @@ -1277,6 +1534,7 @@ type UnderscoreBackupSourceSwitchParams struct { // UnderscoreCloneParams defines model for _CloneParams. type UnderscoreCloneParams struct { + ConsistencyGroup *string `json:"consistency_group,omitempty"` DeleteSnapOnLvolDelete *bool `json:"delete_snap_on_lvol_delete,omitempty"` Name string `json:"name"` PvcName *string `json:"pvc_name,omitempty"` @@ -1294,6 +1552,7 @@ type UnderscoreContinueParams struct { // UnderscoreCreateParams defines model for _CreateParams. type UnderscoreCreateParams struct { AllowedHosts *[]string `json:"allowed_hosts,omitempty"` + ConsistencyGroup *string `json:"consistency_group,omitempty"` DoReplicate *bool `json:"do_replicate,omitempty"` Encrypt *bool `json:"encrypt,omitempty"` Fabric *string `json:"fabric,omitempty"` @@ -1399,6 +1658,27 @@ type ClustersDetailApiV2ClustersClusterIdGetParams struct { Watch *bool `form:"watch,omitempty" json:"watch,omitempty"` } +// ClustersAlertsListApiV2ClustersClusterIdAlertsGetParams defines parameters for ClustersAlertsListApiV2ClustersClusterIdAlertsGet. +type ClustersAlertsListApiV2ClustersClusterIdAlertsGetParams struct { + // Severity Only return alerts of this severity + Severity *ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsSeverity `form:"severity,omitempty" json:"severity,omitempty"` + + // History Also return alerts that have already resolved + History *bool `form:"history,omitempty" json:"history,omitempty"` + + // HistorySeconds Limit the history to alerts resolved within this many seconds. Implies history=true. + HistorySeconds *int `form:"history_seconds,omitempty" json:"history_seconds,omitempty"` + + // Status Only return alerts in this state + Status *ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsStatus `form:"status,omitempty" json:"status,omitempty"` +} + +// ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsSeverity defines parameters for ClustersAlertsListApiV2ClustersClusterIdAlertsGet. +type ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsSeverity string + +// ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsStatus defines parameters for ClustersAlertsListApiV2ClustersClusterIdAlertsGet. +type ClustersAlertsListApiV2ClustersClusterIdAlertsGetParamsStatus string + // ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostParams defines parameters for ClustersBackupsCreateApiV2ClustersClusterIdBackupsPost. type ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostParams struct { ResponseFormat *ClustersBackupsCreateApiV2ClustersClusterIdBackupsPostParamsResponseFormat `form:"response-format,omitempty" json:"response-format,omitempty"` @@ -1421,6 +1701,11 @@ type ClustersCapacityApiV2ClustersClusterIdCapacityGetParams struct { History *string `form:"history,omitempty" json:"history,omitempty"` } +// ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetParams defines parameters for ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGet. +type ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetParams struct { + Name *string `form:"name,omitempty" json:"name,omitempty"` +} + // ClustersIostatsApiV2ClustersClusterIdIostatsGetParams defines parameters for ClustersIostatsApiV2ClustersClusterIdIostatsGet. type ClustersIostatsApiV2ClustersClusterIdIostatsGetParams struct { History *string `form:"history,omitempty" json:"history,omitempty"` @@ -1557,7 +1842,8 @@ type ClustersStoragePoolsIostatsApiV2ClustersClusterIdStoragePoolsPoolIdIostatsG // ClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsGetParams defines parameters for ClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsGet. type ClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsGetParams struct { // Watch Stream state changes as Server-Sent Events instead of returning a plain response: a `snapshot` event with the current state first, then `created`/`updated`/`deleted` events carrying the full resource representation. A `deleted` event carries the resource's final state when it is still retrievable (e.g. a volume whose status became `deleted`), or an empty object once it is gone entirely. Streams do not support resume; reconnecting clients receive a fresh snapshot. Changes written by pre-upgrade components may take up to 30 seconds to appear. - Watch *bool `form:"watch,omitempty" json:"watch,omitempty"` + Watch *bool `form:"watch,omitempty" json:"watch,omitempty"` + ConsistencyGroup *string `form:"consistency_group,omitempty" json:"consistency_group,omitempty"` } // ClustersStoragePoolsSnapshotsDetailApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdGetParams defines parameters for ClustersStoragePoolsSnapshotsDetailApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsSnapshotIdGet. @@ -1689,6 +1975,9 @@ type ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePostJSONRequestBo // ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostJSONRequestBody defines body for ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPost for application/json ContentType. type ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostJSONRequestBody = UnderscoreBackupSourceSwitchParams +// ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostJSONRequestBody defines body for ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPost for application/json ContentType. +type ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostJSONRequestBody = ConsistencyGroupMemberJoinDTO + // ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostJSONRequestBody defines body for ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPost for application/json ContentType. type ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostJSONRequestBody = PolicyParams @@ -2026,6 +2315,9 @@ type ServerInterface interface { // ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPost Clusters:Addreplication // (POST /api/v2/clusters/{cluster_id}/addreplication) ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPost(w http.ResponseWriter, r *http.Request, clusterId openapi_types.UUID) + // ClustersAlertsListApiV2ClustersClusterIdAlertsGet Clusters:Alerts:List + // (GET /api/v2/clusters/{cluster_id}/alerts/) + ClustersAlertsListApiV2ClustersClusterIdAlertsGet(w http.ResponseWriter, r *http.Request, clusterId openapi_types.UUID, params ClustersAlertsListApiV2ClustersClusterIdAlertsGetParams) // ClustersBackupsListApiV2ClustersClusterIdBackupsGet Clusters:Backups:List // (GET /api/v2/clusters/{cluster_id}/backups/) ClustersBackupsListApiV2ClustersClusterIdBackupsGet(w http.ResponseWriter, r *http.Request, clusterId openapi_types.UUID) @@ -2073,6 +2365,33 @@ type ServerInterface interface { // ClustersCapacityApiV2ClustersClusterIdCapacityGet Clusters:Capacity // (GET /api/v2/clusters/{cluster_id}/capacity) ClustersCapacityApiV2ClustersClusterIdCapacityGet(w http.ResponseWriter, r *http.Request, clusterId openapi_types.UUID, params ClustersCapacityApiV2ClustersClusterIdCapacityGetParams) + // ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGet Clusters:Consistency-Groups:List + // (GET /api/v2/clusters/{cluster_id}/consistency-groups/) + ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGet(w http.ResponseWriter, r *http.Request, clusterId openapi_types.UUID, params ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetParams) + // ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGet Clusters:Consistency-Groups:Detail + // (GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/) + ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGet(w http.ResponseWriter, r *http.Request, clusterId openapi_types.UUID, groupId openapi_types.UUID) + // ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGet Clusters:Consistency-Groups:Members + // (GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/members) + ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGet(w http.ResponseWriter, r *http.Request, clusterId openapi_types.UUID, groupId openapi_types.UUID) + // ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPost Clusters:Consistency-Groups:Members:Join + // (POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/members) + ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPost(w http.ResponseWriter, r *http.Request, clusterId openapi_types.UUID, groupId openapi_types.UUID) + // ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDelete Clusters:Consistency-Groups:Members:Detach + // (DELETE /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/members/{lvol_id}) + ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDelete(w http.ResponseWriter, r *http.Request, clusterId openapi_types.UUID, groupId openapi_types.UUID, lvolId string) + // ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGet Clusters:Consistency-Groups:Snapshots:List + // (GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/snapshots) + ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGet(w http.ResponseWriter, r *http.Request, clusterId openapi_types.UUID, groupId openapi_types.UUID) + // ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPost Clusters:Consistency-Groups:Snapshots:Take + // (POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/snapshots) + ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPost(w http.ResponseWriter, r *http.Request, clusterId openapi_types.UUID, groupId openapi_types.UUID) + // ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDelete Clusters:Consistency-Groups:Snapshots:Delete + // (DELETE /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/snapshots/{seq}) + ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDelete(w http.ResponseWriter, r *http.Request, clusterId openapi_types.UUID, groupId openapi_types.UUID, seq int) + // ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGet Clusters:Consistency-Groups:Snapshots:Detail + // (GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/snapshots/{seq}) + ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGet(w http.ResponseWriter, r *http.Request, clusterId openapi_types.UUID, groupId openapi_types.UUID, seq int) // ClustersExpandApiV2ClustersClusterIdExpandPost Clusters:Expand // (POST /api/v2/clusters/{cluster_id}/expand) ClustersExpandApiV2ClustersClusterIdExpandPost(w http.ResponseWriter, r *http.Request, clusterId openapi_types.UUID) @@ -2100,9 +2419,15 @@ type ServerInterface interface { // ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPost Clusters:Replication:Policies:Failover // (POST /api/v2/clusters/{cluster_id}/replication/policies/{policy_id}/failover) ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPost(w http.ResponseWriter, r *http.Request, clusterId openapi_types.UUID, policyId openapi_types.UUID) + // ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGet Clusters:Replication:Policies:Latest-Generation + // (GET /api/v2/clusters/{cluster_id}/replication/policies/{policy_id}/latest-generation) + ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGet(w http.ResponseWriter, r *http.Request, clusterId openapi_types.UUID, policyId openapi_types.UUID) // ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGet Clusters:Replication:Relationships:Detail // (GET /api/v2/clusters/{cluster_id}/replication/relationships/{lvol_id}) ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGet(w http.ResponseWriter, r *http.Request, clusterId openapi_types.UUID, lvolId openapi_types.UUID) + // ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGet Clusters:Replication:Relationships:Latest-Snapshot + // (GET /api/v2/clusters/{cluster_id}/replication/relationships/{lvol_id}/latest-snapshot) + ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGet(w http.ResponseWriter, r *http.Request, clusterId openapi_types.UUID, lvolId openapi_types.UUID) // ClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGet Clusters:Replication:Targets:List // (GET /api/v2/clusters/{cluster_id}/replication/targets/) ClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGet(w http.ResponseWriter, r *http.Request, clusterId openapi_types.UUID) @@ -2289,6 +2614,9 @@ type ServerInterface interface { // ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPost Clusters:Storage-Pools:Volumes:Replication:Start // (POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/start) ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPost(w http.ResponseWriter, r *http.Request, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID) + // ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGet Clusters:Storage-Pools:Volumes:Replication:Status + // (GET /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/status) + ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGet(w http.ResponseWriter, r *http.Request, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID) // ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPost Clusters:Storage-Pools:Volumes:Replication:Stop // (POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/stop) ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPost(w http.ResponseWriter, r *http.Request, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID) @@ -2614,6 +2942,87 @@ func (siw *ServerInterfaceWrapper) ClustersAddreplicationApiV2ClustersClusterIdA handler.ServeHTTP(w, r) } +// ClustersAlertsListApiV2ClustersClusterIdAlertsGet operation middleware +func (siw *ServerInterfaceWrapper) ClustersAlertsListApiV2ClustersClusterIdAlertsGet(w http.ResponseWriter, r *http.Request) { + + var err error + _ = err + + // ------------- Path parameter "cluster_id" ------------- + var clusterId openapi_types.UUID + + err = runtime.BindStyledParameterWithOptions("simple", "cluster_id", r.PathValue("cluster_id"), &clusterId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + if err != nil { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "cluster_id", Err: err}) + return + } + + // Parameter object where we will unmarshal all parameters from the context + var params ClustersAlertsListApiV2ClustersClusterIdAlertsGetParams + + // ------------- Optional query parameter "severity" ------------- + + err = runtime.BindQueryParameterWithOptions("form", true, false, "severity", r.URL.Query(), ¶ms.Severity, runtime.BindQueryParameterOptions{Type: "", Format: ""}) + if err != nil { + var requiredError *runtime.RequiredParameterError + if errors.As(err, &requiredError) { + siw.ErrorHandlerFunc(w, r, &RequiredParamError{ParamName: "severity"}) + } else { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "severity", Err: err}) + } + return + } + + // ------------- Optional query parameter "history" ------------- + + err = runtime.BindQueryParameterWithOptions("form", true, false, "history", r.URL.Query(), ¶ms.History, runtime.BindQueryParameterOptions{Type: "boolean", Format: ""}) + if err != nil { + var requiredError *runtime.RequiredParameterError + if errors.As(err, &requiredError) { + siw.ErrorHandlerFunc(w, r, &RequiredParamError{ParamName: "history"}) + } else { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "history", Err: err}) + } + return + } + + // ------------- Optional query parameter "history_seconds" ------------- + + err = runtime.BindQueryParameterWithOptions("form", true, false, "history_seconds", r.URL.Query(), ¶ms.HistorySeconds, runtime.BindQueryParameterOptions{Type: "", Format: ""}) + if err != nil { + var requiredError *runtime.RequiredParameterError + if errors.As(err, &requiredError) { + siw.ErrorHandlerFunc(w, r, &RequiredParamError{ParamName: "history_seconds"}) + } else { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "history_seconds", Err: err}) + } + return + } + + // ------------- Optional query parameter "status" ------------- + + err = runtime.BindQueryParameterWithOptions("form", true, false, "status", r.URL.Query(), ¶ms.Status, runtime.BindQueryParameterOptions{Type: "", Format: ""}) + if err != nil { + var requiredError *runtime.RequiredParameterError + if errors.As(err, &requiredError) { + siw.ErrorHandlerFunc(w, r, &RequiredParamError{ParamName: "status"}) + } else { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "status", Err: err}) + } + return + } + + handler := http.Handler(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + siw.Handler.ClustersAlertsListApiV2ClustersClusterIdAlertsGet(w, r, clusterId, params) + })) + + for _, middleware := range siw.HandlerMiddlewares { + handler = middleware(handler) + } + + handler.ServeHTTP(w, r) +} + // ClustersBackupsListApiV2ClustersClusterIdBackupsGet operation middleware func (siw *ServerInterfaceWrapper) ClustersBackupsListApiV2ClustersClusterIdBackupsGet(w http.ResponseWriter, r *http.Request) { @@ -2752,14 +3161,313 @@ func (siw *ServerInterfaceWrapper) ClustersBackupPoliciesDeleteApiV2ClustersClus // ------------- Path parameter "policy_id" ------------- var policyId openapi_types.UUID - err = runtime.BindStyledParameterWithOptions("simple", "policy_id", r.PathValue("policy_id"), &policyId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + err = runtime.BindStyledParameterWithOptions("simple", "policy_id", r.PathValue("policy_id"), &policyId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + if err != nil { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "policy_id", Err: err}) + return + } + + handler := http.Handler(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + siw.Handler.ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDelete(w, r, clusterId, policyId) + })) + + for _, middleware := range siw.HandlerMiddlewares { + handler = middleware(handler) + } + + handler.ServeHTTP(w, r) +} + +// ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPost operation middleware +func (siw *ServerInterfaceWrapper) ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPost(w http.ResponseWriter, r *http.Request) { + + var err error + _ = err + + // ------------- Path parameter "cluster_id" ------------- + var clusterId openapi_types.UUID + + err = runtime.BindStyledParameterWithOptions("simple", "cluster_id", r.PathValue("cluster_id"), &clusterId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + if err != nil { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "cluster_id", Err: err}) + return + } + + // ------------- Path parameter "policy_id" ------------- + var policyId openapi_types.UUID + + err = runtime.BindStyledParameterWithOptions("simple", "policy_id", r.PathValue("policy_id"), &policyId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + if err != nil { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "policy_id", Err: err}) + return + } + + handler := http.Handler(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + siw.Handler.ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPost(w, r, clusterId, policyId) + })) + + for _, middleware := range siw.HandlerMiddlewares { + handler = middleware(handler) + } + + handler.ServeHTTP(w, r) +} + +// ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPost operation middleware +func (siw *ServerInterfaceWrapper) ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPost(w http.ResponseWriter, r *http.Request) { + + var err error + _ = err + + // ------------- Path parameter "cluster_id" ------------- + var clusterId openapi_types.UUID + + err = runtime.BindStyledParameterWithOptions("simple", "cluster_id", r.PathValue("cluster_id"), &clusterId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + if err != nil { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "cluster_id", Err: err}) + return + } + + // ------------- Path parameter "policy_id" ------------- + var policyId openapi_types.UUID + + err = runtime.BindStyledParameterWithOptions("simple", "policy_id", r.PathValue("policy_id"), &policyId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + if err != nil { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "policy_id", Err: err}) + return + } + + handler := http.Handler(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + siw.Handler.ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPost(w, r, clusterId, policyId) + })) + + for _, middleware := range siw.HandlerMiddlewares { + handler = middleware(handler) + } + + handler.ServeHTTP(w, r) +} + +// ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGet operation middleware +func (siw *ServerInterfaceWrapper) ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGet(w http.ResponseWriter, r *http.Request) { + + var err error + _ = err + + // ------------- Path parameter "cluster_id" ------------- + var clusterId openapi_types.UUID + + err = runtime.BindStyledParameterWithOptions("simple", "cluster_id", r.PathValue("cluster_id"), &clusterId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + if err != nil { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "cluster_id", Err: err}) + return + } + + // Parameter object where we will unmarshal all parameters from the context + var params ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetParams + + // ------------- Optional query parameter "backup_id" ------------- + + err = runtime.BindQueryParameterWithOptions("form", true, false, "backup_id", r.URL.Query(), ¶ms.BackupId, runtime.BindQueryParameterOptions{Type: "", Format: ""}) + if err != nil { + var requiredError *runtime.RequiredParameterError + if errors.As(err, &requiredError) { + siw.ErrorHandlerFunc(w, r, &RequiredParamError{ParamName: "backup_id"}) + } else { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "backup_id", Err: err}) + } + return + } + + // ------------- Optional query parameter "lvol_name" ------------- + + err = runtime.BindQueryParameterWithOptions("form", true, false, "lvol_name", r.URL.Query(), ¶ms.LvolName, runtime.BindQueryParameterOptions{Type: "", Format: ""}) + if err != nil { + var requiredError *runtime.RequiredParameterError + if errors.As(err, &requiredError) { + siw.ErrorHandlerFunc(w, r, &RequiredParamError{ParamName: "lvol_name"}) + } else { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "lvol_name", Err: err}) + } + return + } + + handler := http.Handler(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + siw.Handler.ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGet(w, r, clusterId, params) + })) + + for _, middleware := range siw.HandlerMiddlewares { + handler = middleware(handler) + } + + handler.ServeHTTP(w, r) +} + +// ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPost operation middleware +func (siw *ServerInterfaceWrapper) ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPost(w http.ResponseWriter, r *http.Request) { + + var err error + _ = err + + // ------------- Path parameter "cluster_id" ------------- + var clusterId openapi_types.UUID + + err = runtime.BindStyledParameterWithOptions("simple", "cluster_id", r.PathValue("cluster_id"), &clusterId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + if err != nil { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "cluster_id", Err: err}) + return + } + + handler := http.Handler(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + siw.Handler.ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPost(w, r, clusterId) + })) + + for _, middleware := range siw.HandlerMiddlewares { + handler = middleware(handler) + } + + handler.ServeHTTP(w, r) +} + +// ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePost operation middleware +func (siw *ServerInterfaceWrapper) ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePost(w http.ResponseWriter, r *http.Request) { + + var err error + _ = err + + // ------------- Path parameter "cluster_id" ------------- + var clusterId openapi_types.UUID + + err = runtime.BindStyledParameterWithOptions("simple", "cluster_id", r.PathValue("cluster_id"), &clusterId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + if err != nil { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "cluster_id", Err: err}) + return + } + + handler := http.Handler(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + siw.Handler.ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePost(w, r, clusterId) + })) + + for _, middleware := range siw.HandlerMiddlewares { + handler = middleware(handler) + } + + handler.ServeHTTP(w, r) +} + +// ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPost operation middleware +func (siw *ServerInterfaceWrapper) ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPost(w http.ResponseWriter, r *http.Request) { + + var err error + _ = err + + // ------------- Path parameter "cluster_id" ------------- + var clusterId openapi_types.UUID + + err = runtime.BindStyledParameterWithOptions("simple", "cluster_id", r.PathValue("cluster_id"), &clusterId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + if err != nil { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "cluster_id", Err: err}) + return + } + + handler := http.Handler(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + siw.Handler.ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPost(w, r, clusterId) + })) + + for _, middleware := range siw.HandlerMiddlewares { + handler = middleware(handler) + } + + handler.ServeHTTP(w, r) +} + +// ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGet operation middleware +func (siw *ServerInterfaceWrapper) ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGet(w http.ResponseWriter, r *http.Request) { + + var err error + _ = err + + // ------------- Path parameter "cluster_id" ------------- + var clusterId openapi_types.UUID + + err = runtime.BindStyledParameterWithOptions("simple", "cluster_id", r.PathValue("cluster_id"), &clusterId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + if err != nil { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "cluster_id", Err: err}) + return + } + + handler := http.Handler(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + siw.Handler.ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGet(w, r, clusterId) + })) + + for _, middleware := range siw.HandlerMiddlewares { + handler = middleware(handler) + } + + handler.ServeHTTP(w, r) +} + +// ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGet operation middleware +func (siw *ServerInterfaceWrapper) ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGet(w http.ResponseWriter, r *http.Request) { + + var err error + _ = err + + // ------------- Path parameter "cluster_id" ------------- + var clusterId openapi_types.UUID + + err = runtime.BindStyledParameterWithOptions("simple", "cluster_id", r.PathValue("cluster_id"), &clusterId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + if err != nil { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "cluster_id", Err: err}) + return + } + + // ------------- Path parameter "backup_id" ------------- + var backupId openapi_types.UUID + + err = runtime.BindStyledParameterWithOptions("simple", "backup_id", r.PathValue("backup_id"), &backupId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + if err != nil { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "backup_id", Err: err}) + return + } + + handler := http.Handler(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + siw.Handler.ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGet(w, r, clusterId, backupId) + })) + + for _, middleware := range siw.HandlerMiddlewares { + handler = middleware(handler) + } + + handler.ServeHTTP(w, r) +} + +// ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDelete operation middleware +func (siw *ServerInterfaceWrapper) ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDelete(w http.ResponseWriter, r *http.Request) { + + var err error + _ = err + + // ------------- Path parameter "cluster_id" ------------- + var clusterId openapi_types.UUID + + err = runtime.BindStyledParameterWithOptions("simple", "cluster_id", r.PathValue("cluster_id"), &clusterId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + if err != nil { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "cluster_id", Err: err}) + return + } + + // ------------- Path parameter "volume_id" ------------- + var volumeId openapi_types.UUID + + err = runtime.BindStyledParameterWithOptions("simple", "volume_id", r.PathValue("volume_id"), &volumeId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) if err != nil { - siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "policy_id", Err: err}) + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "volume_id", Err: err}) return } handler := http.Handler(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { - siw.Handler.ClustersBackupPoliciesDeleteApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDelete(w, r, clusterId, policyId) + siw.Handler.ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDelete(w, r, clusterId, volumeId) })) for _, middleware := range siw.HandlerMiddlewares { @@ -2769,8 +3477,8 @@ func (siw *ServerInterfaceWrapper) ClustersBackupPoliciesDeleteApiV2ClustersClus handler.ServeHTTP(w, r) } -// ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPost operation middleware -func (siw *ServerInterfaceWrapper) ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPost(w http.ResponseWriter, r *http.Request) { +// ClustersCapacityApiV2ClustersClusterIdCapacityGet operation middleware +func (siw *ServerInterfaceWrapper) ClustersCapacityApiV2ClustersClusterIdCapacityGet(w http.ResponseWriter, r *http.Request) { var err error _ = err @@ -2784,17 +3492,24 @@ func (siw *ServerInterfaceWrapper) ClustersBackupPoliciesAttachApiV2ClustersClus return } - // ------------- Path parameter "policy_id" ------------- - var policyId openapi_types.UUID + // Parameter object where we will unmarshal all parameters from the context + var params ClustersCapacityApiV2ClustersClusterIdCapacityGetParams - err = runtime.BindStyledParameterWithOptions("simple", "policy_id", r.PathValue("policy_id"), &policyId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + // ------------- Optional query parameter "history" ------------- + + err = runtime.BindQueryParameterWithOptions("form", true, false, "history", r.URL.Query(), ¶ms.History, runtime.BindQueryParameterOptions{Type: "", Format: ""}) if err != nil { - siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "policy_id", Err: err}) + var requiredError *runtime.RequiredParameterError + if errors.As(err, &requiredError) { + siw.ErrorHandlerFunc(w, r, &RequiredParamError{ParamName: "history"}) + } else { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "history", Err: err}) + } return } handler := http.Handler(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { - siw.Handler.ClustersBackupPoliciesAttachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdAttachPost(w, r, clusterId, policyId) + siw.Handler.ClustersCapacityApiV2ClustersClusterIdCapacityGet(w, r, clusterId, params) })) for _, middleware := range siw.HandlerMiddlewares { @@ -2804,8 +3519,8 @@ func (siw *ServerInterfaceWrapper) ClustersBackupPoliciesAttachApiV2ClustersClus handler.ServeHTTP(w, r) } -// ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPost operation middleware -func (siw *ServerInterfaceWrapper) ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPost(w http.ResponseWriter, r *http.Request) { +// ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGet operation middleware +func (siw *ServerInterfaceWrapper) ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGet(w http.ResponseWriter, r *http.Request) { var err error _ = err @@ -2819,17 +3534,24 @@ func (siw *ServerInterfaceWrapper) ClustersBackupPoliciesDetachApiV2ClustersClus return } - // ------------- Path parameter "policy_id" ------------- - var policyId openapi_types.UUID + // Parameter object where we will unmarshal all parameters from the context + var params ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetParams - err = runtime.BindStyledParameterWithOptions("simple", "policy_id", r.PathValue("policy_id"), &policyId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + // ------------- Optional query parameter "name" ------------- + + err = runtime.BindQueryParameterWithOptions("form", true, false, "name", r.URL.Query(), ¶ms.Name, runtime.BindQueryParameterOptions{Type: "", Format: ""}) if err != nil { - siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "policy_id", Err: err}) + var requiredError *runtime.RequiredParameterError + if errors.As(err, &requiredError) { + siw.ErrorHandlerFunc(w, r, &RequiredParamError{ParamName: "name"}) + } else { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "name", Err: err}) + } return } handler := http.Handler(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { - siw.Handler.ClustersBackupPoliciesDetachApiV2ClustersClusterIdBackupsBackupPoliciesPolicyIdDetachPost(w, r, clusterId, policyId) + siw.Handler.ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGet(w, r, clusterId, params) })) for _, middleware := range siw.HandlerMiddlewares { @@ -2839,8 +3561,8 @@ func (siw *ServerInterfaceWrapper) ClustersBackupPoliciesDetachApiV2ClustersClus handler.ServeHTTP(w, r) } -// ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGet operation middleware -func (siw *ServerInterfaceWrapper) ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGet(w http.ResponseWriter, r *http.Request) { +// ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGet operation middleware +func (siw *ServerInterfaceWrapper) ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGet(w http.ResponseWriter, r *http.Request) { var err error _ = err @@ -2854,37 +3576,17 @@ func (siw *ServerInterfaceWrapper) ClustersBackupsExportApiV2ClustersClusterIdBa return } - // Parameter object where we will unmarshal all parameters from the context - var params ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGetParams - - // ------------- Optional query parameter "backup_id" ------------- - - err = runtime.BindQueryParameterWithOptions("form", true, false, "backup_id", r.URL.Query(), ¶ms.BackupId, runtime.BindQueryParameterOptions{Type: "", Format: ""}) - if err != nil { - var requiredError *runtime.RequiredParameterError - if errors.As(err, &requiredError) { - siw.ErrorHandlerFunc(w, r, &RequiredParamError{ParamName: "backup_id"}) - } else { - siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "backup_id", Err: err}) - } - return - } - - // ------------- Optional query parameter "lvol_name" ------------- + // ------------- Path parameter "group_id" ------------- + var groupId openapi_types.UUID - err = runtime.BindQueryParameterWithOptions("form", true, false, "lvol_name", r.URL.Query(), ¶ms.LvolName, runtime.BindQueryParameterOptions{Type: "", Format: ""}) + err = runtime.BindStyledParameterWithOptions("simple", "group_id", r.PathValue("group_id"), &groupId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) if err != nil { - var requiredError *runtime.RequiredParameterError - if errors.As(err, &requiredError) { - siw.ErrorHandlerFunc(w, r, &RequiredParamError{ParamName: "lvol_name"}) - } else { - siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "lvol_name", Err: err}) - } + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "group_id", Err: err}) return } handler := http.Handler(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { - siw.Handler.ClustersBackupsExportApiV2ClustersClusterIdBackupsExportGet(w, r, clusterId, params) + siw.Handler.ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGet(w, r, clusterId, groupId) })) for _, middleware := range siw.HandlerMiddlewares { @@ -2894,8 +3596,8 @@ func (siw *ServerInterfaceWrapper) ClustersBackupsExportApiV2ClustersClusterIdBa handler.ServeHTTP(w, r) } -// ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPost operation middleware -func (siw *ServerInterfaceWrapper) ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPost(w http.ResponseWriter, r *http.Request) { +// ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGet operation middleware +func (siw *ServerInterfaceWrapper) ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGet(w http.ResponseWriter, r *http.Request) { var err error _ = err @@ -2909,8 +3611,17 @@ func (siw *ServerInterfaceWrapper) ClustersBackupsImportApiV2ClustersClusterIdBa return } + // ------------- Path parameter "group_id" ------------- + var groupId openapi_types.UUID + + err = runtime.BindStyledParameterWithOptions("simple", "group_id", r.PathValue("group_id"), &groupId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + if err != nil { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "group_id", Err: err}) + return + } + handler := http.Handler(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { - siw.Handler.ClustersBackupsImportApiV2ClustersClusterIdBackupsImportPost(w, r, clusterId) + siw.Handler.ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGet(w, r, clusterId, groupId) })) for _, middleware := range siw.HandlerMiddlewares { @@ -2920,8 +3631,8 @@ func (siw *ServerInterfaceWrapper) ClustersBackupsImportApiV2ClustersClusterIdBa handler.ServeHTTP(w, r) } -// ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePost operation middleware -func (siw *ServerInterfaceWrapper) ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePost(w http.ResponseWriter, r *http.Request) { +// ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPost operation middleware +func (siw *ServerInterfaceWrapper) ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPost(w http.ResponseWriter, r *http.Request) { var err error _ = err @@ -2935,8 +3646,17 @@ func (siw *ServerInterfaceWrapper) ClustersBackupsRestoreApiV2ClustersClusterIdB return } + // ------------- Path parameter "group_id" ------------- + var groupId openapi_types.UUID + + err = runtime.BindStyledParameterWithOptions("simple", "group_id", r.PathValue("group_id"), &groupId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + if err != nil { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "group_id", Err: err}) + return + } + handler := http.Handler(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { - siw.Handler.ClustersBackupsRestoreApiV2ClustersClusterIdBackupsRestorePost(w, r, clusterId) + siw.Handler.ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPost(w, r, clusterId, groupId) })) for _, middleware := range siw.HandlerMiddlewares { @@ -2946,8 +3666,8 @@ func (siw *ServerInterfaceWrapper) ClustersBackupsRestoreApiV2ClustersClusterIdB handler.ServeHTTP(w, r) } -// ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPost operation middleware -func (siw *ServerInterfaceWrapper) ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPost(w http.ResponseWriter, r *http.Request) { +// ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDelete operation middleware +func (siw *ServerInterfaceWrapper) ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDelete(w http.ResponseWriter, r *http.Request) { var err error _ = err @@ -2961,8 +3681,26 @@ func (siw *ServerInterfaceWrapper) ClustersBackupsSourceSwitchApiV2ClustersClust return } + // ------------- Path parameter "group_id" ------------- + var groupId openapi_types.UUID + + err = runtime.BindStyledParameterWithOptions("simple", "group_id", r.PathValue("group_id"), &groupId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + if err != nil { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "group_id", Err: err}) + return + } + + // ------------- Path parameter "lvol_id" ------------- + var lvolId string + + err = runtime.BindStyledParameterWithOptions("simple", "lvol_id", r.PathValue("lvol_id"), &lvolId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "", ValueIsUnescaped: true}) + if err != nil { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "lvol_id", Err: err}) + return + } + handler := http.Handler(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { - siw.Handler.ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPost(w, r, clusterId) + siw.Handler.ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDelete(w, r, clusterId, groupId, lvolId) })) for _, middleware := range siw.HandlerMiddlewares { @@ -2972,8 +3710,8 @@ func (siw *ServerInterfaceWrapper) ClustersBackupsSourceSwitchApiV2ClustersClust handler.ServeHTTP(w, r) } -// ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGet operation middleware -func (siw *ServerInterfaceWrapper) ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGet(w http.ResponseWriter, r *http.Request) { +// ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGet operation middleware +func (siw *ServerInterfaceWrapper) ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGet(w http.ResponseWriter, r *http.Request) { var err error _ = err @@ -2987,8 +3725,17 @@ func (siw *ServerInterfaceWrapper) ClustersBackupsSourcesApiV2ClustersClusterIdB return } + // ------------- Path parameter "group_id" ------------- + var groupId openapi_types.UUID + + err = runtime.BindStyledParameterWithOptions("simple", "group_id", r.PathValue("group_id"), &groupId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + if err != nil { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "group_id", Err: err}) + return + } + handler := http.Handler(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { - siw.Handler.ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGet(w, r, clusterId) + siw.Handler.ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGet(w, r, clusterId, groupId) })) for _, middleware := range siw.HandlerMiddlewares { @@ -2998,8 +3745,8 @@ func (siw *ServerInterfaceWrapper) ClustersBackupsSourcesApiV2ClustersClusterIdB handler.ServeHTTP(w, r) } -// ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGet operation middleware -func (siw *ServerInterfaceWrapper) ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGet(w http.ResponseWriter, r *http.Request) { +// ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPost operation middleware +func (siw *ServerInterfaceWrapper) ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPost(w http.ResponseWriter, r *http.Request) { var err error _ = err @@ -3013,17 +3760,17 @@ func (siw *ServerInterfaceWrapper) ClustersBackupsDetailApiV2ClustersClusterIdBa return } - // ------------- Path parameter "backup_id" ------------- - var backupId openapi_types.UUID + // ------------- Path parameter "group_id" ------------- + var groupId openapi_types.UUID - err = runtime.BindStyledParameterWithOptions("simple", "backup_id", r.PathValue("backup_id"), &backupId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + err = runtime.BindStyledParameterWithOptions("simple", "group_id", r.PathValue("group_id"), &groupId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) if err != nil { - siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "backup_id", Err: err}) + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "group_id", Err: err}) return } handler := http.Handler(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { - siw.Handler.ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGet(w, r, clusterId, backupId) + siw.Handler.ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPost(w, r, clusterId, groupId) })) for _, middleware := range siw.HandlerMiddlewares { @@ -3033,8 +3780,8 @@ func (siw *ServerInterfaceWrapper) ClustersBackupsDetailApiV2ClustersClusterIdBa handler.ServeHTTP(w, r) } -// ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDelete operation middleware -func (siw *ServerInterfaceWrapper) ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDelete(w http.ResponseWriter, r *http.Request) { +// ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDelete operation middleware +func (siw *ServerInterfaceWrapper) ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDelete(w http.ResponseWriter, r *http.Request) { var err error _ = err @@ -3048,17 +3795,26 @@ func (siw *ServerInterfaceWrapper) ClustersBackupsDeleteApiV2ClustersClusterIdBa return } - // ------------- Path parameter "volume_id" ------------- - var volumeId openapi_types.UUID + // ------------- Path parameter "group_id" ------------- + var groupId openapi_types.UUID - err = runtime.BindStyledParameterWithOptions("simple", "volume_id", r.PathValue("volume_id"), &volumeId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + err = runtime.BindStyledParameterWithOptions("simple", "group_id", r.PathValue("group_id"), &groupId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) if err != nil { - siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "volume_id", Err: err}) + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "group_id", Err: err}) + return + } + + // ------------- Path parameter "seq" ------------- + var seq int + + err = runtime.BindStyledParameterWithOptions("simple", "seq", r.PathValue("seq"), &seq, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "integer", Format: "", ValueIsUnescaped: true}) + if err != nil { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "seq", Err: err}) return } handler := http.Handler(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { - siw.Handler.ClustersBackupsDeleteApiV2ClustersClusterIdBackupsVolumeIdDelete(w, r, clusterId, volumeId) + siw.Handler.ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDelete(w, r, clusterId, groupId, seq) })) for _, middleware := range siw.HandlerMiddlewares { @@ -3068,8 +3824,8 @@ func (siw *ServerInterfaceWrapper) ClustersBackupsDeleteApiV2ClustersClusterIdBa handler.ServeHTTP(w, r) } -// ClustersCapacityApiV2ClustersClusterIdCapacityGet operation middleware -func (siw *ServerInterfaceWrapper) ClustersCapacityApiV2ClustersClusterIdCapacityGet(w http.ResponseWriter, r *http.Request) { +// ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGet operation middleware +func (siw *ServerInterfaceWrapper) ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGet(w http.ResponseWriter, r *http.Request) { var err error _ = err @@ -3083,24 +3839,26 @@ func (siw *ServerInterfaceWrapper) ClustersCapacityApiV2ClustersClusterIdCapacit return } - // Parameter object where we will unmarshal all parameters from the context - var params ClustersCapacityApiV2ClustersClusterIdCapacityGetParams + // ------------- Path parameter "group_id" ------------- + var groupId openapi_types.UUID - // ------------- Optional query parameter "history" ------------- + err = runtime.BindStyledParameterWithOptions("simple", "group_id", r.PathValue("group_id"), &groupId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + if err != nil { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "group_id", Err: err}) + return + } - err = runtime.BindQueryParameterWithOptions("form", true, false, "history", r.URL.Query(), ¶ms.History, runtime.BindQueryParameterOptions{Type: "", Format: ""}) + // ------------- Path parameter "seq" ------------- + var seq int + + err = runtime.BindStyledParameterWithOptions("simple", "seq", r.PathValue("seq"), &seq, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "integer", Format: "", ValueIsUnescaped: true}) if err != nil { - var requiredError *runtime.RequiredParameterError - if errors.As(err, &requiredError) { - siw.ErrorHandlerFunc(w, r, &RequiredParamError{ParamName: "history"}) - } else { - siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "history", Err: err}) - } + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "seq", Err: err}) return } handler := http.Handler(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { - siw.Handler.ClustersCapacityApiV2ClustersClusterIdCapacityGet(w, r, clusterId, params) + siw.Handler.ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGet(w, r, clusterId, groupId, seq) })) for _, middleware := range siw.HandlerMiddlewares { @@ -3432,6 +4190,41 @@ func (siw *ServerInterfaceWrapper) ClustersReplicationPoliciesFailoverApiV2Clust handler.ServeHTTP(w, r) } +// ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGet operation middleware +func (siw *ServerInterfaceWrapper) ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGet(w http.ResponseWriter, r *http.Request) { + + var err error + _ = err + + // ------------- Path parameter "cluster_id" ------------- + var clusterId openapi_types.UUID + + err = runtime.BindStyledParameterWithOptions("simple", "cluster_id", r.PathValue("cluster_id"), &clusterId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + if err != nil { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "cluster_id", Err: err}) + return + } + + // ------------- Path parameter "policy_id" ------------- + var policyId openapi_types.UUID + + err = runtime.BindStyledParameterWithOptions("simple", "policy_id", r.PathValue("policy_id"), &policyId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + if err != nil { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "policy_id", Err: err}) + return + } + + handler := http.Handler(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + siw.Handler.ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGet(w, r, clusterId, policyId) + })) + + for _, middleware := range siw.HandlerMiddlewares { + handler = middleware(handler) + } + + handler.ServeHTTP(w, r) +} + // ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGet operation middleware func (siw *ServerInterfaceWrapper) ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGet(w http.ResponseWriter, r *http.Request) { @@ -3467,6 +4260,41 @@ func (siw *ServerInterfaceWrapper) ClustersReplicationRelationshipsDetailApiV2Cl handler.ServeHTTP(w, r) } +// ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGet operation middleware +func (siw *ServerInterfaceWrapper) ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGet(w http.ResponseWriter, r *http.Request) { + + var err error + _ = err + + // ------------- Path parameter "cluster_id" ------------- + var clusterId openapi_types.UUID + + err = runtime.BindStyledParameterWithOptions("simple", "cluster_id", r.PathValue("cluster_id"), &clusterId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + if err != nil { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "cluster_id", Err: err}) + return + } + + // ------------- Path parameter "lvol_id" ------------- + var lvolId openapi_types.UUID + + err = runtime.BindStyledParameterWithOptions("simple", "lvol_id", r.PathValue("lvol_id"), &lvolId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + if err != nil { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "lvol_id", Err: err}) + return + } + + handler := http.Handler(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + siw.Handler.ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGet(w, r, clusterId, lvolId) + })) + + for _, middleware := range siw.HandlerMiddlewares { + handler = middleware(handler) + } + + handler.ServeHTTP(w, r) +} + // ClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGet operation middleware func (siw *ServerInterfaceWrapper) ClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGet(w http.ResponseWriter, r *http.Request) { @@ -5132,6 +5960,19 @@ func (siw *ServerInterfaceWrapper) ClustersStoragePoolsSnapshotsListApiV2Cluster return } + // ------------- Optional query parameter "consistency_group" ------------- + + err = runtime.BindQueryParameterWithOptions("form", true, false, "consistency_group", r.URL.Query(), ¶ms.ConsistencyGroup, runtime.BindQueryParameterOptions{Type: "", Format: ""}) + if err != nil { + var requiredError *runtime.RequiredParameterError + if errors.As(err, &requiredError) { + siw.ErrorHandlerFunc(w, r, &RequiredParamError{ParamName: "consistency_group"}) + } else { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "consistency_group", Err: err}) + } + return + } + handler := http.Handler(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { siw.Handler.ClustersStoragePoolsSnapshotsListApiV2ClustersClusterIdStoragePoolsPoolIdSnapshotsGet(w, r, clusterId, poolId, params) })) @@ -6360,6 +7201,50 @@ func (siw *ServerInterfaceWrapper) ClustersStoragePoolsVolumesReplicationStartAp handler.ServeHTTP(w, r) } +// ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGet operation middleware +func (siw *ServerInterfaceWrapper) ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGet(w http.ResponseWriter, r *http.Request) { + + var err error + _ = err + + // ------------- Path parameter "cluster_id" ------------- + var clusterId openapi_types.UUID + + err = runtime.BindStyledParameterWithOptions("simple", "cluster_id", r.PathValue("cluster_id"), &clusterId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + if err != nil { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "cluster_id", Err: err}) + return + } + + // ------------- Path parameter "pool_id" ------------- + var poolId openapi_types.UUID + + err = runtime.BindStyledParameterWithOptions("simple", "pool_id", r.PathValue("pool_id"), &poolId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + if err != nil { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "pool_id", Err: err}) + return + } + + // ------------- Path parameter "volume_id" ------------- + var volumeId openapi_types.UUID + + err = runtime.BindStyledParameterWithOptions("simple", "volume_id", r.PathValue("volume_id"), &volumeId, runtime.BindStyledParameterOptions{ParamLocation: runtime.ParamLocationPath, Explode: false, Required: true, Type: "string", Format: "uuid", ValueIsUnescaped: true}) + if err != nil { + siw.ErrorHandlerFunc(w, r, &InvalidParamFormatError{ParamName: "volume_id", Err: err}) + return + } + + handler := http.Handler(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + siw.Handler.ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGet(w, r, clusterId, poolId, volumeId) + })) + + for _, middleware := range siw.HandlerMiddlewares { + handler = middleware(handler) + } + + handler.ServeHTTP(w, r) +} + // ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPost operation middleware func (siw *ServerInterfaceWrapper) ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPost(w http.ResponseWriter, r *http.Request) { @@ -7165,6 +8050,7 @@ func HandlerWithOptions(si ServerInterface, options StdHTTPServerOptions) http.H m.HandleFunc(http.MethodPut+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/{$}", wrapper.ClustersUpdateApiV2ClustersClusterIdPut) m.HandleFunc(http.MethodPost+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/activate", wrapper.ClustersActivateApiV2ClustersClusterIdActivatePost) m.HandleFunc(http.MethodPost+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/addreplication", wrapper.ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPost) + m.HandleFunc(http.MethodGet+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/alerts/{$}", wrapper.ClustersAlertsListApiV2ClustersClusterIdAlertsGet) m.HandleFunc(http.MethodGet+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/backups/{$}", wrapper.ClustersBackupsListApiV2ClustersClusterIdBackupsGet) m.HandleFunc(http.MethodPost+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/backups/{$}", wrapper.ClustersBackupsCreateApiV2ClustersClusterIdBackupsPost) m.HandleFunc(http.MethodGet+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/backups/{backup_id}/{$}", wrapper.ClustersBackupsDetailApiV2ClustersClusterIdBackupsBackupIdGet) @@ -7180,6 +8066,15 @@ func HandlerWithOptions(si ServerInterface, options StdHTTPServerOptions) http.H m.HandleFunc(http.MethodPost+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/backups/source-switch", wrapper.ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPost) m.HandleFunc(http.MethodGet+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/backups/sources", wrapper.ClustersBackupsSourcesApiV2ClustersClusterIdBackupsSourcesGet) m.HandleFunc(http.MethodGet+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/capacity", wrapper.ClustersCapacityApiV2ClustersClusterIdCapacityGet) + m.HandleFunc(http.MethodGet+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/consistency-groups/{$}", wrapper.ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGet) + m.HandleFunc(http.MethodGet+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/{$}", wrapper.ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGet) + m.HandleFunc(http.MethodGet+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/members", wrapper.ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGet) + m.HandleFunc(http.MethodPost+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/members", wrapper.ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPost) + m.HandleFunc(http.MethodDelete+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/members/{lvol_id}", wrapper.ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDelete) + m.HandleFunc(http.MethodGet+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/snapshots", wrapper.ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGet) + m.HandleFunc(http.MethodPost+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/snapshots", wrapper.ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPost) + m.HandleFunc(http.MethodDelete+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/snapshots/{seq}", wrapper.ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDelete) + m.HandleFunc(http.MethodGet+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/snapshots/{seq}", wrapper.ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGet) m.HandleFunc(http.MethodPost+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/expand", wrapper.ClustersExpandApiV2ClustersClusterIdExpandPost) m.HandleFunc(http.MethodGet+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/iostats", wrapper.ClustersIostatsApiV2ClustersClusterIdIostatsGet) m.HandleFunc(http.MethodGet+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/logs", wrapper.ClustersLogsApiV2ClustersClusterIdLogsGet) @@ -7189,7 +8084,9 @@ func HandlerWithOptions(si ServerInterface, options StdHTTPServerOptions) http.H m.HandleFunc(http.MethodDelete+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/replication/policies/{policy_id}/{$}", wrapper.ClustersReplicationPoliciesDeleteApiV2ClustersClusterIdReplicationPoliciesPolicyIdDelete) m.HandleFunc(http.MethodGet+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/replication/policies/{policy_id}/{$}", wrapper.ClustersReplicationPoliciesDetailApiV2ClustersClusterIdReplicationPoliciesPolicyIdGet) m.HandleFunc(http.MethodPost+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/replication/policies/{policy_id}/failover", wrapper.ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplicationPoliciesPolicyIdFailoverPost) + m.HandleFunc(http.MethodGet+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/replication/policies/{policy_id}/latest-generation", wrapper.ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGet) m.HandleFunc(http.MethodGet+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/replication/relationships/{lvol_id}", wrapper.ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGet) + m.HandleFunc(http.MethodGet+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/replication/relationships/{lvol_id}/latest-snapshot", wrapper.ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGet) m.HandleFunc(http.MethodGet+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/replication/targets/{$}", wrapper.ClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGet) m.HandleFunc(http.MethodPost+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/replication/targets/{$}", wrapper.ClustersReplicationTargetsCreateApiV2ClustersClusterIdReplicationTargetsPost) m.HandleFunc(http.MethodDelete+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/replication/targets/{target_id}/{$}", wrapper.ClustersReplicationTargetsDeleteApiV2ClustersClusterIdReplicationTargetsTargetIdDelete) @@ -7251,6 +8148,7 @@ func HandlerWithOptions(si ServerInterface, options StdHTTPServerOptions) http.H m.HandleFunc(http.MethodPost+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/failback", wrapper.ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPost) m.HandleFunc(http.MethodPost+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/failover", wrapper.ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPost) m.HandleFunc(http.MethodPost+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/start", wrapper.ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPost) + m.HandleFunc(http.MethodGet+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/status", wrapper.ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGet) m.HandleFunc(http.MethodPost+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/stop", wrapper.ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPost) m.HandleFunc(http.MethodGet+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/tasks", wrapper.ClustersStoragePoolsVolumesReplicationTasksApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTasksGet) m.HandleFunc(http.MethodPost+" "+options.BaseURL+"/api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/trigger", wrapper.ClustersStoragePoolsVolumesReplicationTriggerApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationTriggerPost) diff --git a/test/integration/controlplane/unimplemented.gen.go b/test/integration/controlplane/unimplemented.gen.go index e536851d1..4ae1cd0ce 100644 --- a/test/integration/controlplane/unimplemented.gen.go +++ b/test/integration/controlplane/unimplemented.gen.go @@ -1,6 +1,6 @@ // Code generated by ./gen; DO NOT EDIT. // -// A 501 stub for each of cpsim.gen.go's 112 endpoints that handlers.go does not implement. +// A 501 stub for each of cpsim.gen.go's 125 endpoints that handlers.go does not implement. // 9 are implemented; the rest answer "not implemented," which is a truthful // answer and a distinguishable one — a 404 from here would look like a missing // volume rather than a missing simulator. @@ -57,6 +57,10 @@ func (s *Server) ClustersAddreplicationApiV2ClustersClusterIdAddreplicationPost( notImplemented(w, r) } +func (s *Server) ClustersAlertsListApiV2ClustersClusterIdAlertsGet(w http.ResponseWriter, r *http.Request, _ openapi_types.UUID, _ ClustersAlertsListApiV2ClustersClusterIdAlertsGetParams) { + notImplemented(w, r) +} + func (s *Server) ClustersBackupsListApiV2ClustersClusterIdBackupsGet(w http.ResponseWriter, r *http.Request, _ openapi_types.UUID) { notImplemented(w, r) } @@ -117,6 +121,42 @@ func (s *Server) ClustersCapacityApiV2ClustersClusterIdCapacityGet(w http.Respon notImplemented(w, r) } +func (s *Server) ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGet(w http.ResponseWriter, r *http.Request, _ openapi_types.UUID, _ ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetParams) { + notImplemented(w, r) +} + +func (s *Server) ClustersConsistencyGroupsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdGet(w http.ResponseWriter, r *http.Request, _ openapi_types.UUID, _ openapi_types.UUID) { + notImplemented(w, r) +} + +func (s *Server) ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGet(w http.ResponseWriter, r *http.Request, _ openapi_types.UUID, _ openapi_types.UUID) { + notImplemented(w, r) +} + +func (s *Server) ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPost(w http.ResponseWriter, r *http.Request, _ openapi_types.UUID, _ openapi_types.UUID) { + notImplemented(w, r) +} + +func (s *Server) ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDelete(w http.ResponseWriter, r *http.Request, _ openapi_types.UUID, _ openapi_types.UUID, _ string) { + notImplemented(w, r) +} + +func (s *Server) ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGet(w http.ResponseWriter, r *http.Request, _ openapi_types.UUID, _ openapi_types.UUID) { + notImplemented(w, r) +} + +func (s *Server) ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPost(w http.ResponseWriter, r *http.Request, _ openapi_types.UUID, _ openapi_types.UUID) { + notImplemented(w, r) +} + +func (s *Server) ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDelete(w http.ResponseWriter, r *http.Request, _ openapi_types.UUID, _ openapi_types.UUID, _ int) { + notImplemented(w, r) +} + +func (s *Server) ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGet(w http.ResponseWriter, r *http.Request, _ openapi_types.UUID, _ openapi_types.UUID, _ int) { + notImplemented(w, r) +} + func (s *Server) ClustersExpandApiV2ClustersClusterIdExpandPost(w http.ResponseWriter, r *http.Request, _ openapi_types.UUID) { notImplemented(w, r) } @@ -153,10 +193,18 @@ func (s *Server) ClustersReplicationPoliciesFailoverApiV2ClustersClusterIdReplic notImplemented(w, r) } +func (s *Server) ClustersReplicationPoliciesLatestGenerationApiV2ClustersClusterIdReplicationPoliciesPolicyIdLatestGenerationGet(w http.ResponseWriter, r *http.Request, _ openapi_types.UUID, _ openapi_types.UUID) { + notImplemented(w, r) +} + func (s *Server) ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGet(w http.ResponseWriter, r *http.Request, _ openapi_types.UUID, _ openapi_types.UUID) { notImplemented(w, r) } +func (s *Server) ClustersReplicationRelationshipsLatestSnapshotApiV2ClustersClusterIdReplicationRelationshipsLvolIdLatestSnapshotGet(w http.ResponseWriter, r *http.Request, _ openapi_types.UUID, _ openapi_types.UUID) { + notImplemented(w, r) +} + func (s *Server) ClustersReplicationTargetsListApiV2ClustersClusterIdReplicationTargetsGet(w http.ResponseWriter, r *http.Request, _ openapi_types.UUID) { notImplemented(w, r) } @@ -377,6 +425,10 @@ func (s *Server) ClustersStoragePoolsVolumesReplicationStartApiV2ClustersCluster notImplemented(w, r) } +func (s *Server) ClustersStoragePoolsVolumesReplicationStatusApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStatusGet(w http.ResponseWriter, r *http.Request, _ openapi_types.UUID, _ openapi_types.UUID, _ openapi_types.UUID) { + notImplemented(w, r) +} + func (s *Server) ClustersStoragePoolsVolumesReplicationStopApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStopPost(w http.ResponseWriter, r *http.Request, _ openapi_types.UUID, _ openapi_types.UUID, _ openapi_types.UUID) { notImplemented(w, r) } From b0098f31a883961dfa1dfd06f719ad75e22914ae Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Fri, 18 Sep 2026 10:12:15 +0100 Subject: [PATCH 063/206] Implement Phase 2 lifecycle verbs (Promote/Demote/Resync) with P0-3 demote endpoint and planned-gate force-escalation fix --- atlas-lib/controlplane/replication.go | 85 ++++++ atlas-lib/controlplane/replication_test.go | 129 +++++++++ atlas-lib/internal/cpapi/cpapi.gen.go | 256 +++++++++++++++++- .../internal/csi/controller/errorclass_rpc.go | 29 ++ .../csi/controller/mock_controlplane_test.go | 69 +++++ .../internal/csi/controller/replication.go | 102 ++++++- .../controller/replication_lifecycle_test.go | 217 +++++++++++++++ .../csi/controller/replication_test.go | 13 - .../designs/design-csi-addons-replication.md | 46 ++-- .../tests/test-plan-csi-addons-replication.md | 102 +++---- shared/openapi.json | 77 +++++- 11 files changed, 1034 insertions(+), 91 deletions(-) create mode 100644 csi-driver/internal/csi/controller/replication_lifecycle_test.go diff --git a/atlas-lib/controlplane/replication.go b/atlas-lib/controlplane/replication.go index ce5b942a5..4b88df585 100644 --- a/atlas-lib/controlplane/replication.go +++ b/atlas-lib/controlplane/replication.go @@ -138,3 +138,88 @@ func (c *Client) GetVolumeReplicationInfo(ctx context.Context, h lvol.VolumeHand } return replicationStatusFromDTO(*d), nil } + +// PromoteVolume brings the volume up as primary on this cluster. +// +// force=true is the unplanned path: it ignores demote state entirely, +// because its whole premise is that the peer may never have been reachable +// to demote. force=false is the planned path, gated on a completed demote +// (P0-3): a 409 (demote still converging, retryable) or 412 (no demote was +// ever requested) surfaces as a *StatusError the caller classifies -- the +// 409/412 split matters because the vendored csi-addons controller +// auto-escalates ANY FAILED_PRECONDITION from a force=false promote to +// force=true inline, with no wait-and-retry grace period of its own. +func (c *Client) PromoteVolume(ctx context.Context, h lvol.VolumeHandle, force bool) error { + cluster, pool, volume, err := h.Split() + if err != nil { + return err + } + planned := !force + params := &cpapi.ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPostParams{ + Planned: &planned, + } + resp, err := c.api.ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPostWithResponse( + ctx, cluster, pool, volume, params) + if err != nil { + return fmt.Errorf("promote volume %s: %w", h, err) + } + if code := resp.StatusCode(); code != http.StatusOK && code != http.StatusNoContent { + return respError("promote volume "+string(h), code, resp.Body) + } + return nil +} + +// DemoteVolume fences the source and confirms the last write replicated +// (P0-3) -- the lossless half of a planned swap. Synchronous and +// re-drivable, not queued: it returns done=false while the final snapshot is +// still converging, and the caller (the driver's DemoteVolume RPC) is +// expected to call this again rather than block, matching the backend's own +// call-repeatedly contract. +func (c *Client) DemoteVolume(ctx context.Context, h lvol.VolumeHandle) (bool, error) { + cluster, pool, volume, err := h.Split() + if err != nil { + return false, err + } + resp, err := c.api.ClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePostWithResponse( + ctx, cluster, pool, volume) + if err != nil { + return false, fmt.Errorf("demote volume %s: %w", h, err) + } + switch code := resp.StatusCode(); code { + case http.StatusNoContent: + return true, nil + case http.StatusAccepted: + return false, nil + default: + return false, respError("demote volume "+string(h), code, resp.Body) + } +} + +// ResyncVolume reconciles a diverged copy back onto the current primary's +// history. It configures the reverse direction only -- it never cuts over, +// matching the design's own "it never merges": cutover is PromoteVolume's job +// on a separate, later call. sourceClusterID selects the source explicitly +// when it isn't the cluster's configured default; "" leaves it unset. +func (c *Client) ResyncVolume(ctx context.Context, h lvol.VolumeHandle, sourceClusterID string) error { + cluster, pool, volume, err := h.Split() + if err != nil { + return err + } + var body cpapi.FailbackParams + if sourceClusterID != "" { + id, err := parseUUID("source cluster id", sourceClusterID) + if err != nil { + return err + } + body.SourceClusterId = &id + } + resp, err := c.api.ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostWithResponse( + ctx, cluster, pool, volume, body) + if err != nil { + return fmt.Errorf("resync volume %s: %w", h, err) + } + if code := resp.StatusCode(); code != http.StatusOK && code != http.StatusNoContent { + return respError("resync volume "+string(h), code, resp.Body) + } + return nil +} diff --git a/atlas-lib/controlplane/replication_test.go b/atlas-lib/controlplane/replication_test.go index 1057e7801..9567f0fee 100644 --- a/atlas-lib/controlplane/replication_test.go +++ b/atlas-lib/controlplane/replication_test.go @@ -171,3 +171,132 @@ func TestClientGetVolumeReplicationInfoNotFound(t *testing.T) { t.Errorf("err = %v, want ErrNotFound", err) } } + +func TestClientPromoteVolumeForcedIgnoresDemoteState(t *testing.T) { + var gotQuery string + c := newTestClient(t, func(w http.ResponseWriter, r *http.Request) { + if !strings.HasSuffix(r.URL.Path, "/replication/failover") { + t.Errorf("unexpected path %q", r.URL.Path) + } + gotQuery = r.URL.RawQuery + w.WriteHeader(http.StatusNoContent) + }) + if err := c.PromoteVolume(context.Background(), testHandle, true); err != nil { + t.Fatal(err) + } + if strings.Contains(gotQuery, "planned=true") { + t.Errorf("query = %q, forced promote must not ask for the planned gate", gotQuery) + } +} + +func TestClientPromoteVolumePlannedSendsThePlannedFlag(t *testing.T) { + var gotQuery string + c := newTestClient(t, func(w http.ResponseWriter, r *http.Request) { + gotQuery = r.URL.RawQuery + w.WriteHeader(http.StatusNoContent) + }) + if err := c.PromoteVolume(context.Background(), testHandle, false); err != nil { + t.Fatal(err) + } + if !strings.Contains(gotQuery, "planned=true") { + t.Errorf("query = %q, want planned=true", gotQuery) + } +} + +// The whole point of the planned gate: a demote still converging must surface +// as something the driver can map to codes.Aborted (retryable), never +// codes.FailedPrecondition -- the vendored csi-addons controller escalates +// ANY FAILED_PRECONDITION from a force=false promote to force=true inline, +// with no wait-and-retry grace period of its own. +func TestClientPromoteVolumeWhileDemoteConvergingIsA409(t *testing.T) { + c := newTestClient(t, func(w http.ResponseWriter, r *http.Request) { + w.WriteHeader(http.StatusConflict) + }) + err := c.PromoteVolume(context.Background(), testHandle, false) + var se *StatusError + if !errors.As(err, &se) || se.StatusCode != http.StatusConflict { + t.Fatalf("err = %v, want a *StatusError carrying 409", err) + } +} + +func TestClientPromoteVolumeWithNoDemoteIsA412(t *testing.T) { + c := newTestClient(t, func(w http.ResponseWriter, r *http.Request) { + w.WriteHeader(http.StatusPreconditionFailed) + }) + err := c.PromoteVolume(context.Background(), testHandle, false) + var se *StatusError + if !errors.As(err, &se) || se.StatusCode != http.StatusPreconditionFailed { + t.Fatalf("err = %v, want a *StatusError carrying 412", err) + } +} + +func TestClientDemoteVolumeNotYetDone(t *testing.T) { + c := newTestClient(t, func(w http.ResponseWriter, r *http.Request) { + if !strings.HasSuffix(r.URL.Path, "/replication/demote") { + t.Errorf("unexpected path %q", r.URL.Path) + } + w.WriteHeader(http.StatusAccepted) + }) + done, err := c.DemoteVolume(context.Background(), testHandle) + if err != nil { + t.Fatal(err) + } + if done { + t.Error("done = true, want false while still converging") + } +} + +func TestClientDemoteVolumeDone(t *testing.T) { + c := newTestClient(t, func(w http.ResponseWriter, r *http.Request) { + w.WriteHeader(http.StatusNoContent) + }) + done, err := c.DemoteVolume(context.Background(), testHandle) + if err != nil { + t.Fatal(err) + } + if !done { + t.Error("done = false, want true") + } +} + +func TestClientDemoteVolumeFailureIsAnError(t *testing.T) { + c := newTestClient(t, func(w http.ResponseWriter, r *http.Request) { + w.WriteHeader(http.StatusInternalServerError) + }) + if _, err := c.DemoteVolume(context.Background(), testHandle); err == nil { + t.Error("want an error on a genuine backend failure") + } +} + +func TestClientResyncVolume(t *testing.T) { + var gotBody string + c := newTestClient(t, func(w http.ResponseWriter, r *http.Request) { + if !strings.HasSuffix(r.URL.Path, "/replication/failback") { + t.Errorf("unexpected path %q", r.URL.Path) + } + b, _ := io.ReadAll(r.Body) + gotBody = string(b) + w.WriteHeader(http.StatusNoContent) + }) + if err := c.ResyncVolume(context.Background(), testHandle, testCluster); err != nil { + t.Fatal(err) + } + if !strings.Contains(gotBody, `"source_cluster_id":"`+testCluster+`"`) { + t.Errorf("request body = %q, want source_cluster_id %s", gotBody, testCluster) + } +} + +func TestClientResyncVolumeWithoutSourceCluster(t *testing.T) { + var gotBody string + c := newTestClient(t, func(w http.ResponseWriter, r *http.Request) { + b, _ := io.ReadAll(r.Body) + gotBody = string(b) + w.WriteHeader(http.StatusNoContent) + }) + if err := c.ResyncVolume(context.Background(), testHandle, ""); err != nil { + t.Fatal(err) + } + if strings.Contains(gotBody, "source_cluster_id") { + t.Errorf("request body = %q, want no source_cluster_id when none is given", gotBody) + } +} diff --git a/atlas-lib/internal/cpapi/cpapi.gen.go b/atlas-lib/internal/cpapi/cpapi.gen.go index 2f8e51a0d..7bbeb5c91 100644 --- a/atlas-lib/internal/cpapi/cpapi.gen.go +++ b/atlas-lib/internal/cpapi/cpapi.gen.go @@ -1901,7 +1901,8 @@ type ClustersStoragePoolsVolumesReplicationCommitApiV2ClustersClusterIdStoragePo // ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPostParams defines parameters for ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPost. type ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPostParams struct { - Generation *int `form:"generation,omitempty" json:"generation,omitempty"` + Generation *int `form:"generation,omitempty" json:"generation,omitempty"` + Planned *bool `form:"planned,omitempty" json:"planned,omitempty"` } // ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPostJSONBody defines parameters for ClustersStoragePoolsVolumesReplicationStartApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationStartPost. @@ -3216,6 +3217,19 @@ type ClientInterface interface { // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/cutover-proceed (the `ClustersStoragePoolsVolumesReplicationCutoverProceedApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCutoverProceedPost` operationId). ClustersStoragePoolsVolumesReplicationCutoverProceedApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCutoverProceedPost(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) + // ClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePost Clusters:Storage-Pools:Volumes:Replication:Demote + // + // Fence the source and confirm the last write replicated (P0-3). + // + // Synchronous and re-drivable, not queued: each call does only the work its + // current state calls for (fence + trigger the final snapshot once, then + // just check whether it has landed), so the caller re-invokes this route + // until it reports 204. A 202 means still waiting -- call again, the same + // way `GET .../status` is re-read rather than pushed. + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/demote (the `ClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePost` operationId). + ClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePost(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) + // ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostWithBody Clusters:Storage-Pools:Volumes:Replication:Failback // // Point replication back at a source cluster. The cutover itself is @@ -3249,6 +3263,18 @@ type ClientInterface interface { // is the recovery path for a logical corruption, which the newest copy has // faithfully replicated. // + // ``planned=True`` gates on a completed demote (P0-3) so a planned swap + // loses nothing: 412 when no demote was ever requested for this volume (the + // caller's premise that the source is reachable to demote was wrong, and a + // 412 is what lets the csi-addons controller's own force-escalation take + // over), 409 while demote is still converging (retryable -- 409 must never + // become a code the controller reads as permission to force, since that + // controller escalates on ANY FAILED_PRECONDITION from a force=false + // promote with no wait-and-retry grace period of its own). Unplanned + // failover (the default) ignores demote state entirely, unchanged from + // today: its whole premise is that the source may never have been + // reachable to demote. + // // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/failover (the `ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPost` operationId). ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPost(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, params *ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPostParams, reqEditors ...RequestEditorFn) (*http.Response, error) @@ -5572,6 +5598,29 @@ func (c *Client) ClustersStoragePoolsVolumesReplicationCutoverProceedApiV2Cluste return c.Client.Do(req) } +// ClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePost Clusters:Storage-Pools:Volumes:Replication:Demote +// +// Fence the source and confirm the last write replicated (P0-3). +// +// Synchronous and re-drivable, not queued: each call does only the work its +// current state calls for (fence + trigger the final snapshot once, then +// just check whether it has landed), so the caller re-invokes this route +// until it reports 204. A 202 means still waiting -- call again, the same +// way `GET .../status` is re-read rather than pushed. +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/demote (the `ClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePost` operationId). +func (c *Client) ClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePost(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePostRequest(c.Server, clusterId, poolId, volumeId) + if err != nil { + return nil, err + } + req = req.WithContext(ctx) + if err := c.applyEditors(ctx, req, reqEditors); err != nil { + return nil, err + } + return c.Client.Do(req) +} + // ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostWithBody Clusters:Storage-Pools:Volumes:Replication:Failback // // Point replication back at a source cluster. The cutover itself is @@ -5625,6 +5674,18 @@ func (c *Client) ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClus // is the recovery path for a logical corruption, which the newest copy has // faithfully replicated. // +// “planned=True“ gates on a completed demote (P0-3) so a planned swap +// loses nothing: 412 when no demote was ever requested for this volume (the +// caller's premise that the source is reachable to demote was wrong, and a +// 412 is what lets the csi-addons controller's own force-escalation take +// over), 409 while demote is still converging (retryable -- 409 must never +// become a code the controller reads as permission to force, since that +// controller escalates on ANY FAILED_PRECONDITION from a force=false +// promote with no wait-and-retry grace period of its own). Unplanned +// failover (the default) ignores demote state entirely, unchanged from +// today: its whole premise is that the source may never have been +// reachable to demote. +// // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/failover (the `ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPost` operationId). func (c *Client) ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPost(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, params *ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPostParams, reqEditors ...RequestEditorFn) (*http.Response, error) { req, err := NewClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPostRequest(c.Server, clusterId, poolId, volumeId, params) @@ -11801,6 +11862,54 @@ func NewClustersStoragePoolsVolumesReplicationCutoverProceedApiV2ClustersCluster return req, nil } +// NewClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePostRequest constructs an http.Request for the ClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePost method +func NewClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID) (*http.Request, error) { + var err error + + var pathParam0 string + + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + var pathParam1 string + + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "pool_id", poolId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + var pathParam2 string + + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "volume_id", volumeId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + serverURL, err := url.Parse(server) + if err != nil { + return nil, err + } + + operationPath := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s/replication/demote", pathParam0, pathParam1, pathParam2) + if operationPath[0] == '/' { + operationPath = "." + operationPath + } + + queryURL, err := serverURL.Parse(operationPath) + if err != nil { + return nil, err + } + + req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) + if err != nil { + return nil, err + } + + return req, nil +} + // NewClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostRequest calls the generic ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPost builder with application/json body func NewClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostRequest(server string, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, body ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostJSONRequestBody) (*http.Request, error) { var bodyReader io.Reader @@ -11923,6 +12032,18 @@ func NewClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStor } + if params.Planned != nil { + + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "planned", *params.Planned, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "boolean", Format: ""}); err != nil { + return nil, err + } else { + for _, qp := range strings.Split(queryFrag, "&") { + rawQueryFragments = append(rawQueryFragments, qp) + } + } + + } + if encoded := queryValues.Encode(); encoded != "" { rawQueryFragments = append(rawQueryFragments, encoded) } @@ -13967,6 +14088,21 @@ type ClientWithResponsesInterface interface { // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/cutover-proceed (the `ClustersStoragePoolsVolumesReplicationCutoverProceedApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCutoverProceedPost` operationId). ClustersStoragePoolsVolumesReplicationCutoverProceedApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCutoverProceedPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationCutoverProceedApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCutoverProceedPostResponse, error) + // ClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePostWithResponse Clusters:Storage-Pools:Volumes:Replication:Demote + // + // Fence the source and confirm the last write replicated (P0-3). + // + // Synchronous and re-drivable, not queued: each call does only the work its + // current state calls for (fence + trigger the final snapshot once, then + // just check whether it has landed), so the caller re-invokes this route + // until it reports 204. A 202 means still waiting -- call again, the same + // way `GET .../status` is re-read rather than pushed. + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/demote (the `ClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePost` operationId). + ClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePostResponse, error) + // ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostWithBodyWithResponse Clusters:Storage-Pools:Volumes:Replication:Failback // // Point replication back at a source cluster. The cutover itself is @@ -14000,6 +14136,18 @@ type ClientWithResponsesInterface interface { // is the recovery path for a logical corruption, which the newest copy has // faithfully replicated. // + // ``planned=True`` gates on a completed demote (P0-3) so a planned swap + // loses nothing: 412 when no demote was ever requested for this volume (the + // caller's premise that the source is reachable to demote was wrong, and a + // 412 is what lets the csi-addons controller's own force-escalation take + // over), 409 while demote is still converging (retryable -- 409 must never + // become a code the controller reads as permission to force, since that + // controller escalates on ANY FAILED_PRECONDITION from a force=false + // promote with no wait-and-retry grace period of its own). Unplanned + // failover (the default) ignores demote state entirely, unchanged from + // today: its whole premise is that the source may never have been + // reachable to demote. + // // Returns a wrapper object for the known response body format(s). // // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/failover (the `ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPost` operationId). @@ -18895,6 +19043,47 @@ func (r ClustersStoragePoolsVolumesReplicationCutoverProceedApiV2ClustersCluster return "" } +type ClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePostResponse struct { + Body []byte + HTTPResponse *http.Response + // JSON422 the response for an HTTP 422 `application/json` response + JSON422 *HTTPValidationError +} + +// GetJSON422 returns the response for an HTTP 422 `application/json` response +func (r ClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePostResponse) GetJSON422() *HTTPValidationError { + return r.JSON422 +} + +// GetBody returns the raw response body bytes +func (r ClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePostResponse) GetBody() []byte { + return r.Body +} + +// Status returns HTTPResponse.Status +func (r ClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePostResponse) Status() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Status + } + return http.StatusText(0) +} + +// StatusCode returns HTTPResponse.StatusCode +func (r ClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePostResponse) StatusCode() int { + if r.HTTPResponse != nil { + return r.HTTPResponse.StatusCode + } + return 0 +} + +// ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers +func (r ClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePostResponse) ContentType() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Header.Get("Content-Type") + } + return "" +} + type ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostResponse struct { Body []byte HTTPResponse *http.Response @@ -21593,6 +21782,27 @@ func (c *ClientWithResponses) ClustersStoragePoolsVolumesReplicationCutoverProce return ParseClustersStoragePoolsVolumesReplicationCutoverProceedApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationCutoverProceedPostResponse(rsp) } +// ClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePostWithResponse Clusters:Storage-Pools:Volumes:Replication:Demote +// +// Fence the source and confirm the last write replicated (P0-3). +// +// Synchronous and re-drivable, not queued: each call does only the work its +// current state calls for (fence + trigger the final snapshot once, then +// just check whether it has landed), so the caller re-invokes this route +// until it reports 204. A 202 means still waiting -- call again, the same +// way `GET .../status` is re-read rather than pushed. +// +// Returns a wrapper object for the known response body format(s). +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/demote (the `ClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePost` operationId). +func (c *ClientWithResponses) ClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePostResponse, error) { + rsp, err := c.ClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePost(ctx, clusterId, poolId, volumeId, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePostResponse(rsp) +} + // ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostWithBodyWithResponse Clusters:Storage-Pools:Volumes:Replication:Failback // // Point replication back at a source cluster. The cutover itself is @@ -21638,6 +21848,18 @@ func (c *ClientWithResponses) ClustersStoragePoolsVolumesReplicationFailbackApiV // is the recovery path for a logical corruption, which the newest copy has // faithfully replicated. // +// “planned=True“ gates on a completed demote (P0-3) so a planned swap +// loses nothing: 412 when no demote was ever requested for this volume (the +// caller's premise that the source is reachable to demote was wrong, and a +// 412 is what lets the csi-addons controller's own force-escalation take +// over), 409 while demote is still converging (retryable -- 409 must never +// become a code the controller reads as permission to force, since that +// controller escalates on ANY FAILED_PRECONDITION from a force=false +// promote with no wait-and-retry grace period of its own). Unplanned +// failover (the default) ignores demote state entirely, unchanged from +// today: its whole premise is that the source may never have been +// reachable to demote. +// // Returns a wrapper object for the known response body format(s). // // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/failover (the `ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPost` operationId). @@ -25272,6 +25494,38 @@ func ParseClustersStoragePoolsVolumesReplicationCutoverProceedApiV2ClustersClust return response, nil } +// ParseClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePostResponse parses an HTTP response from a ClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePostWithResponse call +func ParseClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePostResponse(rsp *http.Response) (*ClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePostResponse, error) { + bodyBytes, err := io.ReadAll(rsp.Body) + defer func() { _ = rsp.Body.Close() }() + if err != nil { + return nil, err + } + + response := &ClustersStoragePoolsVolumesReplicationDemoteApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationDemotePostResponse{ + Body: bodyBytes, + HTTPResponse: rsp, + } + + switch { + case rsp.StatusCode == 202: + break // No content-type + + case rsp.StatusCode == 204: + break // No content-type + + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: + var dest HTTPValidationError + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON422 = &dest + + } + + return response, nil +} + // ParseClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostResponse parses an HTTP response from a ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostWithResponse call func ParseClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostResponse(rsp *http.Response) (*ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailbackPostResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) diff --git a/csi-driver/internal/csi/controller/errorclass_rpc.go b/csi-driver/internal/csi/controller/errorclass_rpc.go index 7cd9317e5..c1355b92d 100644 --- a/csi-driver/internal/csi/controller/errorclass_rpc.go +++ b/csi-driver/internal/csi/controller/errorclass_rpc.go @@ -117,6 +117,15 @@ func classifyDisableVolumeReplicationError(err error) classifiedError { func classifyGetVolumeReplicationInfoError(err error) classifiedError { return classifiedError{GetVolumeReplicationInfoErrorClassifier.Classify(err), err} } +func classifyPromoteVolumeError(err error) classifiedError { + return classifiedError{PromoteVolumeErrorClassifier.Classify(err), err} +} +func classifyDemoteVolumeError(err error) classifiedError { + return classifiedError{DemoteVolumeErrorClassifier.Classify(err), err} +} +func classifyResyncVolumeError(err error) classifiedError { + return classifiedError{ResyncVolumeErrorClassifier.Classify(err), err} +} // Dispositions reused across RPCs. var ( @@ -183,4 +192,24 @@ var ( GetVolumeReplicationInfoErrorClassifier = errorClassifier{overrides: map[int]controlPlaneErrorClass{ http.StatusNotFound: sourceNotFound, }} + // PromoteVolumeErrorClassifier: a 409 on a planned promote means demote is + // still converging (design §5.2) -- reuses cutoverInFlight's ABORTED, + // retryable disposition, since the vendored csi-addons controller + // auto-escalates ANY FAILED_PRECONDITION to force=true inline with no + // wait-and-retry grace period of its own. A 412 falls through to the + // generic classifier's FailedPrecondition unchanged: that is the one case + // meant to let the controller's own escalation take over. + PromoteVolumeErrorClassifier = errorClassifier{overrides: map[int]controlPlaneErrorClass{ + http.StatusNotFound: sourceNotFound, + http.StatusConflict: cutoverInFlight, + }} + // DemoteVolumeErrorClassifier classifies only genuine backend failures. + // "Not yet done" (202) is not an error at atlas-lib's DemoteVolume, so it + // never reaches this classifier -- the RPC handler checks it directly. + DemoteVolumeErrorClassifier = errorClassifier{overrides: map[int]controlPlaneErrorClass{ + http.StatusNotFound: sourceNotFound, + }} + ResyncVolumeErrorClassifier = errorClassifier{overrides: map[int]controlPlaneErrorClass{ + http.StatusNotFound: sourceNotFound, + }} ) diff --git a/csi-driver/internal/csi/controller/mock_controlplane_test.go b/csi-driver/internal/csi/controller/mock_controlplane_test.go index fa011abb7..75942a6b5 100644 --- a/csi-driver/internal/csi/controller/mock_controlplane_test.go +++ b/csi-driver/internal/csi/controller/mock_controlplane_test.go @@ -122,6 +122,25 @@ type mockSBCLI struct { // normal idempotent update, modeling a backend refusal (e.g. a policy // that is not active) or a transient failure. replicationPUTStatus int + + // failoverStatus, when set, is the HTTP status POST .../failover answers + // with instead of its default success (204) -- modeling the planned + // gate's 409 (demote still converging) and 412 (no demote requested). + failoverStatus int + // lastFailoverQuery captures the raw query string of the last failover + // call, so a test can assert the driver actually sent planned=true/false + // rather than only checking the resulting gRPC code. + lastFailoverQuery string + + // demoteStatus, when set, is the HTTP status POST .../demote answers with + // instead of its default success (204). 202 models "still converging." + demoteStatus int + + // failbackStatus, when set, is the HTTP status POST .../failback answers + // with instead of its default success (204). + failbackStatus int + // lastFailbackBody captures the raw JSON body of the last failback call. + lastFailbackBody []byte } func newMockSBCLI() *mockSBCLI { @@ -152,6 +171,18 @@ func newMockSBCLI() *mockSBCLI { "GET /api/v2/clusters/{clusterID}/storage-pools/{poolID}/volumes/{volumeID}/replication/status", m.locked(m.handleReplicationStatus), ) + mux.HandleFunc( + "POST /api/v2/clusters/{clusterID}/storage-pools/{poolID}/volumes/{volumeID}/replication/failover", + m.locked(m.handleFailover), + ) + mux.HandleFunc( + "POST /api/v2/clusters/{clusterID}/storage-pools/{poolID}/volumes/{volumeID}/replication/demote", + m.locked(m.handleDemote), + ) + mux.HandleFunc( + "POST /api/v2/clusters/{clusterID}/storage-pools/{poolID}/volumes/{volumeID}/replication/failback", + m.locked(m.handleFailback), + ) mux.HandleFunc( "GET /api/v2/clusters/{clusterID}/storage-pools/{poolID}/volumes/{volumeID}/connect", m.locked(m.handleVolumeConnect), @@ -337,6 +368,44 @@ func (m *mockSBCLI) handleReplicationStatus(w http.ResponseWriter, r *http.Reque writeJSON(w, http.StatusOK, body) } +func (m *mockSBCLI) handleFailover(w http.ResponseWriter, r *http.Request) { + volumeID := r.PathValue("volumeID") + if m.lookupVolume(w, volumeID) == nil { + return + } + m.lastFailoverQuery = r.URL.RawQuery + if m.failoverStatus != 0 { + writeJSON(w, m.failoverStatus, map[string]string{"detail": "injected status"}) + return + } + w.WriteHeader(http.StatusNoContent) +} + +func (m *mockSBCLI) handleDemote(w http.ResponseWriter, r *http.Request) { + volumeID := r.PathValue("volumeID") + if m.lookupVolume(w, volumeID) == nil { + return + } + if m.demoteStatus != 0 { + writeJSON(w, m.demoteStatus, map[string]bool{"demoted": m.demoteStatus == http.StatusNoContent}) + return + } + w.WriteHeader(http.StatusNoContent) +} + +func (m *mockSBCLI) handleFailback(w http.ResponseWriter, r *http.Request) { + volumeID := r.PathValue("volumeID") + if m.lookupVolume(w, volumeID) == nil { + return + } + m.lastFailbackBody, _ = io.ReadAll(r.Body) + if m.failbackStatus != 0 { + writeJSON(w, m.failbackStatus, map[string]string{"detail": "injected status"}) + return + } + w.WriteHeader(http.StatusNoContent) +} + func (m *mockSBCLI) handleVolumeConnect(w http.ResponseWriter, r *http.Request) { volumeID := r.PathValue("volumeID") if m.lookupVolume(w, volumeID) == nil { diff --git a/csi-driver/internal/csi/controller/replication.go b/csi-driver/internal/csi/controller/replication.go index fdcfda386..dc7a02fab 100644 --- a/csi-driver/internal/csi/controller/replication.go +++ b/csi-driver/internal/csi/controller/replication.go @@ -1,10 +1,9 @@ // The csi-addons Replication service: EnableVolumeReplication, -// DisableVolumeReplication, and GetVolumeReplicationInfo (design §5.1). Each -// verb is a thin adapter onto the atlas-lib control-plane client's -// replication calls, resolved through the same {clusterID}:{poolID}:{lvolID} -// handle every other RPC uses. The remaining Replication verbs -// (PromoteVolume, DemoteVolume, ResyncVolume) fall through to the embedded -// UnimplementedControllerServer until Phase 2. +// DisableVolumeReplication, GetVolumeReplicationInfo (design §5.1), and the +// Phase 2 lifecycle verbs PromoteVolume, DemoteVolume, and ResyncVolume +// (design §5.2). Each verb is a thin adapter onto the atlas-lib control-plane +// client's replication calls, resolved through the same +// {clusterID}:{poolID}:{lvolID} handle every other RPC uses. package controller import ( @@ -26,6 +25,13 @@ import ( // in atlas-lib. A future change adds that resolution and accepts either. const replicationPolicyParam = "replicationPolicyID" +// sourceClusterIDParam is the VolumeReplicationClass parameter naming the +// cluster to resync from, when it isn't the one the backend already has on +// record for this volume's relationship. Optional: sbcli's replication_failback +// resolves it from the existing relationship when omitted (the common case, +// design §5.2's "Recovered source"). +const sourceClusterIDParam = "sourceClusterID" + // EnableVolumeReplication attaches the volume to the policy named by the // VolumeReplicationClass. Attaching to the policy the volume already follows // is success (the backend's own idempotency, P0-2). @@ -111,3 +117,87 @@ func (cs *Server) GetVolumeReplicationInfo( } return resp, nil } + +// PromoteVolume brings the volume up as primary on this cluster (design +// §5.2). Force=true is the unplanned path: it clones the last fully +// replicated generation and ignores demote state entirely, because its whole +// premise is that the peer may never have been reachable to demote. +// Force=false is the planned path, refused with ABORTED (retryable) while a +// demote is still converging and FAILED_PRECONDITION when no demote was ever +// requested -- the split matters because the vendored csi-addons controller +// auto-escalates ANY FAILED_PRECONDITION from a force=false promote to +// force=true inline, with no wait-and-retry grace period of its own. +func (cs *Server) PromoteVolume( + ctx context.Context, + req *replication.PromoteVolumeRequest, +) (*replication.PromoteVolumeResponse, error) { + h, err := csicommon.ParseVolumeHandle(req.GetVolumeId()) + if err != nil { + return nil, status.Error(codes.InvalidArgument, err.Error()) + } + client, err := clusters.ReplicationClient(ctx, h.ClusterID) + if err != nil { + return nil, status.Error(codes.Unavailable, err.Error()) + } + if err := client.PromoteVolume(ctx, h.Handle(), req.GetForce()); err != nil { + return nil, classifyPromoteVolumeError(err) + } + return &replication.PromoteVolumeResponse{}, nil +} + +// DemoteVolume fences the source and confirms the last write replicated +// (P0-3) -- the lossless half of a planned swap. Synchronous and +// non-blocking: it never waits out the backend's own convergence loop. +// While still converging it returns ABORTED (retryable), matching the actual +// upstream reconciler's requeue-until-ready behavior for a Secondary +// transition that has not yet settled, rather than holding the RPC open. +func (cs *Server) DemoteVolume( + ctx context.Context, + req *replication.DemoteVolumeRequest, +) (*replication.DemoteVolumeResponse, error) { + h, err := csicommon.ParseVolumeHandle(req.GetVolumeId()) + if err != nil { + return nil, status.Error(codes.InvalidArgument, err.Error()) + } + client, err := clusters.ReplicationClient(ctx, h.ClusterID) + if err != nil { + return nil, status.Error(codes.Unavailable, err.Error()) + } + done, err := client.DemoteVolume(ctx, h.Handle()) + if err != nil { + return nil, classifyDemoteVolumeError(err) + } + if !done { + return nil, status.Error(codes.Aborted, "demote is still converging") + } + return &replication.DemoteVolumeResponse{}, nil +} + +// ResyncVolume reconciles a diverged copy back onto the current primary's +// history. It configures the reverse direction only and reports readiness +// off the ordinary lag read -- it never cuts over, matching the design's own +// "it never merges" (§5.2): cutover is PromoteVolume's job, on a separate, +// later call. +func (cs *Server) ResyncVolume( + ctx context.Context, + req *replication.ResyncVolumeRequest, +) (*replication.ResyncVolumeResponse, error) { + h, err := csicommon.ParseVolumeHandle(req.GetVolumeId()) + if err != nil { + return nil, status.Error(codes.InvalidArgument, err.Error()) + } + client, err := clusters.ReplicationClient(ctx, h.ClusterID) + if err != nil { + return nil, status.Error(codes.Unavailable, err.Error()) + } + sourceClusterID := req.GetParameters()[sourceClusterIDParam] + if err := client.ResyncVolume(ctx, h.Handle(), sourceClusterID); err != nil { + return nil, classifyResyncVolumeError(err) + } + info, err := client.GetVolumeReplicationInfo(ctx, h.Handle()) + if err != nil { + return nil, classifyGetVolumeReplicationInfoError(err) + } + ready := info.LagSeconds == nil || info.LagBudgetSeconds == nil || *info.LagSeconds <= *info.LagBudgetSeconds + return &replication.ResyncVolumeResponse{Ready: ready}, nil +} diff --git a/csi-driver/internal/csi/controller/replication_lifecycle_test.go b/csi-driver/internal/csi/controller/replication_lifecycle_test.go new file mode 100644 index 000000000..b2af814f0 --- /dev/null +++ b/csi-driver/internal/csi/controller/replication_lifecycle_test.go @@ -0,0 +1,217 @@ +package controller + +import ( + "context" + "net/http" + "strings" + "testing" + + "github.com/csi-addons/spec/lib/go/replication" + "google.golang.org/grpc/codes" + "google.golang.org/grpc/status" +) + +func TestPromoteVolumeForced(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + + _, err := cs.PromoteVolume(context.Background(), &replication.PromoteVolumeRequest{ + VolumeId: testReplVolID, Force: true, + }) + if err != nil { + t.Fatal(err) + } + if strings.Contains(mock.lastFailoverQuery, "planned=true") { + t.Errorf("query = %q, forced promote must not ask for the planned gate", mock.lastFailoverQuery) + } +} + +func TestPromoteVolumePlannedSendsThePlannedFlag(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + + _, err := cs.PromoteVolume(context.Background(), &replication.PromoteVolumeRequest{ + VolumeId: testReplVolID, Force: false, + }) + if err != nil { + t.Fatal(err) + } + if !strings.Contains(mock.lastFailoverQuery, "planned=true") { + t.Errorf("query = %q, want planned=true", mock.lastFailoverQuery) + } +} + +// The whole point of the planned gate: a demote still converging must map to +// ABORTED (retryable), never FAILED_PRECONDITION -- the vendored csi-addons +// controller auto-escalates ANY FAILED_PRECONDITION from a force=false +// promote to force=true inline, in the same reconcile, with no +// wait-and-retry grace period of its own. Mapping this to FailedPrecondition +// would silently force through a promote while demote is still converging. +func TestPromoteVolumePlannedWhileDemoteConvergingIsAborted(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + mock.failoverStatus = http.StatusConflict + + _, err := cs.PromoteVolume(context.Background(), &replication.PromoteVolumeRequest{ + VolumeId: testReplVolID, Force: false, + }) + st, _ := status.FromError(err) + if st.Code() != codes.Aborted { + t.Errorf("code = %v, want Aborted (retryable, never auto-forced by the vendored controller)", st.Code()) + } +} + +// The case that SHOULD let the vendored controller's own force-escalation +// take over: no demote was ever requested, so this planned attempt only +// makes sense as a genuinely unplanned failover the caller mislabeled. +func TestPromoteVolumePlannedWithNoDemoteIsFailedPrecondition(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + mock.failoverStatus = http.StatusPreconditionFailed + + _, err := cs.PromoteVolume(context.Background(), &replication.PromoteVolumeRequest{ + VolumeId: testReplVolID, Force: false, + }) + st, _ := status.FromError(err) + if st.Code() != codes.FailedPrecondition { + t.Errorf("code = %v, want FailedPrecondition", st.Code()) + } +} + +func TestDemoteVolumeDone(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + + _, err := cs.DemoteVolume(context.Background(), &replication.DemoteVolumeRequest{ + VolumeId: testReplVolID, + }) + if err != nil { + t.Fatal(err) + } +} + +// Non-blocking and re-driven: a still-converging demote must not hang the +// RPC or report success, it must fail with a retryable code so the +// controller-manager's own reconcile loop re-invokes DemoteVolume later -- +// matching the actual upstream reconciler's requeue-until-ready behavior. +func TestDemoteVolumeNotYetDoneIsAborted(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + mock.demoteStatus = http.StatusAccepted + + _, err := cs.DemoteVolume(context.Background(), &replication.DemoteVolumeRequest{ + VolumeId: testReplVolID, + }) + st, _ := status.FromError(err) + if st.Code() != codes.Aborted { + t.Errorf("code = %v, want Aborted (retryable)", st.Code()) + } +} + +func TestDemoteVolumeBackendFailureIsUnavailable(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + mock.demoteStatus = http.StatusInternalServerError + + _, err := cs.DemoteVolume(context.Background(), &replication.DemoteVolumeRequest{ + VolumeId: testReplVolID, + }) + if err == nil { + t.Fatal("want an error on a genuine backend failure") + } +} + +func TestResyncVolume(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + + _, err := cs.ResyncVolume(context.Background(), &replication.ResyncVolumeRequest{ + VolumeId: testReplVolID, + }) + if err != nil { + t.Fatal(err) + } +} + +func TestResyncVolumeForwardsTheSourceClusterParameter(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + + _, err := cs.ResyncVolume(context.Background(), &replication.ResyncVolumeRequest{ + VolumeId: testReplVolID, + Parameters: map[string]string{sourceClusterIDParam: sanityClusterID}, + }) + if err != nil { + t.Fatal(err) + } + if !strings.Contains(string(mock.lastFailbackBody), `"source_cluster_id":"`+sanityClusterID+`"`) { + t.Errorf("failback body = %q, want source_cluster_id %s", mock.lastFailbackBody, sanityClusterID) + } +} + +func TestResyncVolumeReadyReflectsLag(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + mock.replicationStatus[testReplVolumeID] = map[string]any{ + "role": "source", "state": "in_sync", + "lag_seconds": 120, "lag_budget_seconds": 900, + "outstanding_count": 0, "outstanding_bytes": 0, + "failing_count": 0, "max_retry_reached": false, "resyncing": true, + } + + resp, err := cs.ResyncVolume(context.Background(), &replication.ResyncVolumeRequest{ + VolumeId: testReplVolID, + }) + if err != nil { + t.Fatal(err) + } + if !resp.Ready { + t.Error("Ready = false, want true: lag (120s) is inside the budget (900s)") + } +} + +func TestResyncVolumeNotReadyWhileLagExceedsBudget(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + mock.replicationStatus[testReplVolumeID] = map[string]any{ + "role": "source", "state": "in_sync", + "lag_seconds": 1800, "lag_budget_seconds": 900, + "outstanding_count": 0, "outstanding_bytes": 0, + "failing_count": 0, "max_retry_reached": false, "resyncing": true, + } + + resp, err := cs.ResyncVolume(context.Background(), &replication.ResyncVolumeRequest{ + VolumeId: testReplVolID, + }) + if err != nil { + t.Fatal(err) + } + if resp.Ready { + t.Error("Ready = true, want false: lag (1800s) exceeds the budget (900s)") + } +} + +func TestResyncVolumeBackendFailureIsUnavailable(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + mock.failbackStatus = http.StatusInternalServerError + + _, err := cs.ResyncVolume(context.Background(), &replication.ResyncVolumeRequest{ + VolumeId: testReplVolID, + }) + if err == nil { + t.Fatal("want an error on a genuine backend failure") + } +} diff --git a/csi-driver/internal/csi/controller/replication_test.go b/csi-driver/internal/csi/controller/replication_test.go index b8f86be49..d27ca3764 100644 --- a/csi-driver/internal/csi/controller/replication_test.go +++ b/csi-driver/internal/csi/controller/replication_test.go @@ -226,16 +226,3 @@ func TestGetVolumeReplicationInfoUnknownVolume(t *testing.T) { } } -// The remaining Replication verbs (Phase 2) fall through to the embedded -// UnimplementedControllerServer until they are implemented. -func TestUnimplementedReplicationVerbsAreUnimplemented(t *testing.T) { - mock := newMockSBCLI() - defer mock.Close() - cs := newReplicationTestServer(t, mock) - - _, err := cs.PromoteVolume(context.Background(), &replication.PromoteVolumeRequest{VolumeId: testReplVolID}) - st, _ := status.FromError(err) - if st.Code() != codes.Unimplemented { - t.Errorf("PromoteVolume code = %v, want Unimplemented", st.Code()) - } -} diff --git a/operator/docs/designs/design-csi-addons-replication.md b/operator/docs/designs/design-csi-addons-replication.md index c3e5be50a..1871b0e55 100644 --- a/operator/docs/designs/design-csi-addons-replication.md +++ b/operator/docs/designs/design-csi-addons-replication.md @@ -12,7 +12,7 @@ | Phase | Status | Scope | Sections | |-------------|-------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|--------------| | **Phase 1** | Implemented | The csi-addons machinery and the steady-state contract: CRDs, controller-manager, sidecar, the Replication and csi-addons Identity gRPC services with `EnableVolumeReplication`, `DisableVolumeReplication`, and `GetVolumeReplicationInfo`, backed by a typed backend status endpoint | §4, §5.1, §6 | -| **Phase 2** | Planned | The lifecycle verbs: `PromoteVolume` (planned and forced), `DemoteVolume`, and `ResyncVolume`, validated end to end against a Ramen `VolumeReplicationGroup` in async mode | §5.2, §9 | +| **Phase 2** | Implemented | The lifecycle verbs: `PromoteVolume` (planned and forced), `DemoteVolume`, and `ResyncVolume`. Validation end to end against a Ramen `VolumeReplicationGroup` in async mode is still outstanding (§12, E-06/E-07) | §5.2, §9 | | **Phase 3** | Planned | peerClasses convention and preflight, and the replication observability surface (lag, backlog, RPO compliance) | §7, §11 | | **Phase 4** | Planned | Test failover: the latest-replicated-snapshot read, the `drtest-*` conventions, and the two drill modes (bubble and test cluster) composed from clone, replication, and the real failover | §14 | @@ -24,14 +24,14 @@ The phase numbers above are this document's own, not the DR storage foundation g ## Phase 0 — External Prerequisites -| # | Prerequisite | Kind | Blocks | Status | -|------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------|---------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| P0-1 | A typed, steady-state per-volume replication status read: `GET .../volumes/{id}/replication/status` serving what `lvol_controller.get_replication_info` computes today (state, lag, outstanding bytes, failure counters), available for the volume's whole replicated life | Control plane (`sbcli`) | Phase 1 | Shipped | -| P0-2 | Idempotent attach and detach: attaching a volume to the policy it already follows returns success, and detaching a non-attached volume returns success | Control plane (`sbcli`) | Phase 1 | Shipped | -| P0-3 | A standalone demote verb: `POST .../volumes/{id}/replication/demote` that converges the peer while still serving (repeated snapshot-and-ship until the remaining delta is small), then quiesces, ships the final delta, confirms it landed on the peer, and fences the data path | Control plane (`sbcli`) | Phase 2 | Not shipped | -| P0-4 | An `rpo_target_seconds` field on `ReplicationPolicy`, so RPO compliance is computable against a declared target rather than the derived lag budget | Control plane (`sbcli`) | Phase 3 | Shipped | -| P0-5 | csi-addons upstream: the `VolumeReplication` and `VolumeReplicationClass` CRDs (`replication.storage.openshift.io/v1alpha1`), the kubernetes-csi-addons controller-manager image, and the csi-addons sidecar image | Ecosystem | Phase 1 | Vendored in the chart at v0.15.0 behind `csiaddons.create` (all twelve upstream CRDs, since the stock manager starts a controller per kind); sidecar wiring shipped in Phase 1 | -| P0-6 | A latest-replicated-snapshot read: per volume, and per consistency group as one complete generation, the newest fully replicated snapshot on the secondary addressed as a cloneable object | Control plane (`sbcli`) | Phase 4 | Shipped | +| # | Prerequisite | Kind | Blocks | Status | +|------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------|---------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| P0-1 | A typed, steady-state per-volume replication status read: `GET .../volumes/{id}/replication/status` serving what `lvol_controller.get_replication_info` computes today (state, lag, outstanding bytes, failure counters), available for the volume's whole replicated life | Control plane (`sbcli`) | Phase 1 | Shipped | +| P0-2 | Idempotent attach and detach: attaching a volume to the policy it already follows returns success, and detaching a non-attached volume returns success | Control plane (`sbcli`) | Phase 1 | Shipped | +| P0-3 | A standalone demote verb: `POST .../volumes/{id}/replication/demote` that fences the source, ships a final internal snapshot, and confirms it landed on the peer before recording the volume demoted | Control plane (`sbcli`) | Phase 2 | Shipped (narrower than originally scoped here -- see §5.2's note on why it does not reuse the live-cutover engine) | +| P0-4 | An `rpo_target_seconds` field on `ReplicationPolicy`, so RPO compliance is computable against a declared target rather than the derived lag budget | Control plane (`sbcli`) | Phase 3 | Shipped | +| P0-5 | csi-addons upstream: the `VolumeReplication` and `VolumeReplicationClass` CRDs (`replication.storage.openshift.io/v1alpha1`), the kubernetes-csi-addons controller-manager image, and the csi-addons sidecar image | Ecosystem | Phase 1 | Vendored in the chart at v0.15.0 behind `csiaddons.create` (all twelve upstream CRDs, since the stock manager starts a controller per kind); sidecar wiring shipped in Phase 1 | +| P0-6 | A latest-replicated-snapshot read: per volume, and per consistency group as one complete generation, the newest fully replicated snapshot on the secondary addressed as a cloneable object | Control plane (`sbcli`) | Phase 4 | Shipped | Everything else the adapter needs already exists: the attach and detach calls, failover, the failback and commit pair, the relationship read, and the backlog arithmetic inside `get_replication_info`. The adapter is thin precisely because the engine is complete. What is missing is the shape Ramen can drive. @@ -160,7 +160,7 @@ A reader who stops here has the model: the engine is unchanged, the csi-addons s **One relationship, addressed from either cluster.** The `VolumeReplication.spec.dataSource` resolves to a PV whose handle is `{clusterID}:{poolID}:{volumeID}`. The driver resolves the backend from the handle through `clusters.Client`, exactly as every other RPC does, so the DR cluster's driver can drive the same backend relationship as the source cluster's without either side holding special state. This is the same stateless addressing the node plugin already uses to redirect a staged volume after failover. -**The adapter holds no state.** csi-addons RPCs are stateless and idempotent by contract. Every answer the driver gives is derived on the spot from the backend status read and the relationship record. There is no driver-side cache, no persisted step, and no state machine. A verb whose backend work outlives the call (a demote converging a busy peer) reports `Completed=False` until the backend reflects the target state, and Ramen's re-drive is the retry loop. +**The adapter holds no state.** csi-addons RPCs are stateless and idempotent by contract. Every answer the driver gives is derived on the spot from the backend status read and the relationship record. There is no driver-side cache, no persisted step, and no state machine. `PromoteVolumeResponse` and `DemoteVolumeResponse` carry no fields at all in `csi-addons/spec` v0.2.0 -- there is no response field to report partial progress in. A verb whose backend work outlives the call (a demote converging a busy peer) instead returns a retryable `ABORTED` error, and the vendored controller-manager's own reconcile requeue is the retry loop; only once the backend reports the target state reached does the call return success, which the controller then reflects as the CR's `Completed` condition. **The operator's part is small and off the data path.** The kinds that author the backend state (`ReplicationPair` for the target, `ReplicationPolicy` for cadence and retention) keep working unchanged, and the classes name what they author. On top of them the operator runs the peerClasses preflight (§7.2), surfacing convention drift as events on the pair, and teaches the `PVCAnnotationWatcher` the one-owner rule (§8) so the legacy annotation path and a `VolumeReplication` never fight over one volume. Everything imperative it used to own (`ReplicationOps`, the commit cutover, the cutover-proceed handshake) is off this contract and confined to the legacy path. @@ -205,14 +205,14 @@ The generated control-plane client already declares every replication endpoint, ### 5.2 Phase 2 verbs -| Verb | Backend mapping | Semantics | -|------------------------------------|-------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| `PromoteVolume` with `force=true` | `POST .../volumes/{v}/replication/failover` | Unplanned promote: clone the last fully replicated generation on the target, retire the (assumed dead) source. Idempotent by the backend's own NQN probe. `Completed=True` once the relationship reads `failed_over`. | -| `PromoteVolume` with `force=false` | `POST .../volumes/{v}/replication/failover` with the planned gate | The same promote as the forced form, refused unless the peer already holds every acknowledged write, which is exactly the state a completed demote leaves behind: the source fenced and the final flush confirmed. A planned promote against an undemoted or lagging source surfaces as `FAILED_PRECONDITION`. `Completed=True` once the relationship reports the target active. | -| `DemoteVolume` | `POST .../volumes/{v}/replication/demote` (P0-3) | Quiesce (ANA inaccessible), ship a final internal snapshot, wait until it carries the replicated marker, fence the data path, and record the role. This is the lossless half of a planned swap: after demote, the peer's planned promote loses nothing. `Completed=True` only when the final flush has landed on the peer. | -| `ResyncVolume` | `POST .../volumes/{v}/replication/failback` | Reverse the shipping direction, seeded by `data_uuid` matching so a recovered source resyncs by delta rather than full copy. `Resyncing=True` while the catch-up runs; it clears when the reverse lag is inside the budget. Resync reconciles the diverged old primary from the current primary; it never merges. | +| Verb | Backend mapping | Semantics | +|------------------------------------|--------------------------------------------------------------------------------------------------------------------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| `PromoteVolume` with `force=true` | `POST .../volumes/{v}/replication/failover` (already shipped, unchanged) | Unplanned promote: clone the last fully replicated generation on the target, retire the (assumed dead) source. Idempotent by the backend's own NQN probe. Ignores demote state entirely -- its whole premise is that the peer may never have been reachable to demote. | +| `PromoteVolume` with `force=false` | `POST .../volumes/{v}/replication/failover?planned=true` (P0-3 adds the `planned` parameter to the existing route) | The same promote as the forced form, refused unless the peer already holds every acknowledged write, which is exactly the state a completed demote leaves behind: the source fenced and the final flush confirmed. Demote still converging surfaces as `ABORTED` (retryable); no demote ever requested surfaces as `FAILED_PRECONDITION` -- the split matters because the vendored csi-addons controller auto-escalates ANY `FAILED_PRECONDITION` from a force=false promote to force=true inline, with no wait-and-retry grace period of its own (§15, resolves the risk this design's original `FAILED_PRECONDITION`-only wording would have created). | +| `DemoteVolume` | `POST .../volumes/{v}/replication/demote` (P0-3, new route) | Fence the source (ANA inaccessible) FIRST, then trigger and confirm one final internal snapshot. This is the lossless half of a planned swap: after demote, the peer's planned promote loses nothing. Synchronous and re-drivable, not queued: returns `ABORTED` while the snapshot is still converging, success only once confirmed. Deliberately does not reuse `replication/commit`'s live-cutover engine (shrink rounds, the synchronous hub transfer) -- that machinery exists to bound a freeze window against a *live* writer, and by the time Kubernetes calls `DemoteVolume` the workload has already unmounted, so there is no moving target to protect against, only the ordinary fence-before-snapshot discipline. It also never touches a target volume; that is `PromoteVolume`'s job, on a separate, later call, possibly on a different cluster. | +| `ResyncVolume` | `POST .../volumes/{v}/replication/failback` | Reverse the shipping direction, seeded by `data_uuid` matching so a recovered source resyncs by delta rather than full copy. Response's `ready` field is false while lag exceeds the budget and true once caught up. Resync reconciles the diverged old primary from the current primary; it never merges, and never calls `commit` -- that would be a cutover, not a resync. | -**Promote is one operation; planned and unplanned differ only in what precedes it.** Both forms clone the last fully replicated generation on the target and serve it under the preserved NVMe identity, exactly as failover does today. The planned form is lossless not because it runs different machinery but because a completed demote guarantees the last replicated generation contains every acknowledged write, and the planned gate refuses the promote until that holds. The forced form skips the gate and accepts the RPO loss, because its premise is that the source is gone. The engine's commit cutover (`replication/commit`, the `FN_REPLICATION_FINAL` runner with its shrink rounds and cutover-proceed handshake) is deliberately NOT part of this contract: its convergence job moves into the demote verb (P0-3), and it remains only behind the legacy `ReplicationOps` migration path (§8, §13). +**Promote is one operation; planned and unplanned differ only in what precedes it.** Both forms clone the last fully replicated generation on the target and serve it under the preserved NVMe identity, exactly as failover does today. The planned form is lossless not because it runs different machinery but because a completed demote guarantees the last replicated generation contains every acknowledged write, and the planned gate refuses the promote until that holds. The forced form skips the gate and accepts the RPO loss, because its premise is that the source is gone. The engine's commit cutover (`replication/commit`, the `FN_REPLICATION_FINAL` runner with its shrink rounds and cutover-proceed handshake) is deliberately NOT part of this contract and remains behind only the legacy `ReplicationOps` migration path (§8, §13) -- its convergence job does NOT move into the demote verb as this design originally assumed; see the `DemoteVolume` row above for why that engine is the wrong shape for a post-unmount demote. --- @@ -437,9 +437,9 @@ A drill that silently perturbed replication would be worse than no drill. Before ## 15. Open Questions -| # | Question | Owner | -|-----|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-----------------------| -| 1 | **Demote semantics for the application.** The P0-3 demote fences the volume (ANA inaccessible) after the final flush, and with convergence folded into the verb it is now the only place a planned swap can stall. Ramen relocation unmounts the workload first, so the fence is ordinarily unopposed. Confirm the verb's behavior when writes are still in flight at quiesce (block versus fail), whether the converge phase has its own budget separate from the quiesced flush, and whether a timeout in either phase must abort back to serving primary. | Backend team | -| 2 | **Where the preflight lives.** §7.2 attaches peerClasses validation to the `ReplicationPair` reconciler. If the redesign retires the pair kind, the preflight needs a new home (the `SimplyblockDriver`, or a standalone check job). | Operator team | -| 3 | **Different-policy enable refusal has no data source.** §5.1's `EnableVolumeReplication` was designed to refuse an attach to a different policy with `FAILED_PRECONDITION`. Phase 1 found that sbcli's `attach_policy` silently re-attaches instead of refusing, and no endpoint returns the policy id a volume is currently attached to, for the driver to compare against. Either the backend adds that read, or this design accepts the silent re-sync as the behavior. | Backend team | -| 4 | **`lastSyncDuration` and `lastSyncBytes` have no response field.** §5.1's `GetVolumeReplicationInfo` was designed to return all three fields; `csi-addons/spec` v0.2.0's `GetVolumeReplicationInfoResponse` carries only `lastSyncTime`. Confirm whether a newer spec version adds the other two, or whether they surface some other way (a `VolumeReplication` annotation, a metric). | Ecosystem/driver team | +| # | Question | Owner | +|-------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------| +| ~~1~~ | **Resolved.** Demote fences (ANA inaccessible) BEFORE taking the final snapshot, mirroring `fence_source_paths`'s existing invariant, and does not block on in-flight writes: a write that lands after the fence queues against the now-dark path (NVMe multipath semantics) and fails via the client's own `ctrl_loss_tmo`, since nothing in the demote-only flow ever lights a target path to drain it onto -- that is a separate, later `PromoteVolume` call. Demote's own fence-and-snapshot sequence proceeds independently of whatever is queued client-side; it never waits on it. This is a fail choice, not a block, and it trusts Ramen's unmount-first contract rather than adding a defensive "no live session" check before fencing -- a genuine simplification, flagged in case a caller outside Ramen's contract ever drives this verb directly. | Backend team (resolved) | +| 2 | **Where the preflight lives.** §7.2 attaches peerClasses validation to the `ReplicationPair` reconciler. If the redesign retires the pair kind, the preflight needs a new home (the `SimplyblockDriver`, or a standalone check job). | Operator team | +| 3 | **Different-policy enable refusal has no data source.** §5.1's `EnableVolumeReplication` was designed to refuse an attach to a different policy with `FAILED_PRECONDITION`. Phase 1 found that sbcli's `attach_policy` silently re-attaches instead of refusing, and no endpoint returns the policy id a volume is currently attached to, for the driver to compare against. Either the backend adds that read, or this design accepts the silent re-sync as the behavior. | Backend team | +| 4 | **`lastSyncDuration` and `lastSyncBytes` have no response field.** §5.1's `GetVolumeReplicationInfo` was designed to return all three fields; `csi-addons/spec` v0.2.0's `GetVolumeReplicationInfoResponse` carries only `lastSyncTime`. Confirm whether a newer spec version adds the other two, or whether they surface some other way (a `VolumeReplication` annotation, a metric). | Ecosystem/driver team | diff --git a/operator/docs/tests/test-plan-csi-addons-replication.md b/operator/docs/tests/test-plan-csi-addons-replication.md index 692b47c95..1db5bdc8b 100644 --- a/operator/docs/tests/test-plan-csi-addons-replication.md +++ b/operator/docs/tests/test-plan-csi-addons-replication.md @@ -6,7 +6,7 @@ Scope is the CSI driver's Replication service, the operator's preflight and coex Scenario IDs are permanent and are never reused or renumbered. `U-` is unit (no cluster: mock control plane, fake `client.Client`), `I-` is integration (the sidecar and controller-manager against the driver with a mock backend), `E-` is end-to-end (two live simplyblock clusters), and `M-` is manual. Types are `Positive`, `Negative`, `Boundary`, and `Regression`. A `—` in the `Test` column means nothing implements the scenario yet, and every such row reappears in §7 with its reason. -Phase 1 (the csi-addons machinery, §4, §5.1's three verbs, and §6's steady-state contract) has landed; its unit rows below are filled in. Phase 2 (promote, demote, resync, and the operator's preflight and coexistence controllers) has not started, and the integration and E2E tiers wait on a test bed neither phase has built yet. +Phase 1 (the csi-addons machinery, §4, §5.1's three verbs, and §6's steady-state contract) and Phase 2 (§5.2's lifecycle verbs and P0-3) have both landed; their unit rows below are filled in. The operator's preflight and coexistence controllers (peerClasses, `PVCReplicationController`) remain a separate, unbuilt subsystem, and the integration and E2E tiers wait on a test bed neither phase has built yet. --- @@ -18,37 +18,45 @@ The Replication service against a mock control plane, and the operator pieces ag File: `csi-driver/internal/csi/controller/replication_test.go` (planned) -| # | Scenario | Type | Test | -|------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|------------|------------------------------------------------------| -| U-01 | Enable on an unattached volume: the attach call carries the class's policy, and the RPC succeeds | Positive | `TestEnableVolumeReplication` | -| U-02 | Enable on a volume already attached to the same policy: success, no second attach call (idempotency) | Boundary | `TestEnableVolumeReplicationRepeatedIsIdempotent` | -| U-03 | Enable on a volume attached to a different policy: `FAILED_PRECONDITION` naming both policies, no attach call | Negative | — | -| U-04 | Disable on an attached volume: the detach call is made, success | Positive | `TestDisableVolumeReplication` | -| U-05 | Disable on a non-attached volume: success without a backend call (idempotency) | Boundary | `TestDisableVolumeReplicationNotAttachedIsSuccess` | -| U-06 | Disable while a cutover is in flight (backend 409): `ABORTED`, retryable | Negative | `TestDisableVolumeReplicationDuringCutoverIsAborted` | -| U-07 | Info returns `lastSyncTime` from the status read (`lastSyncDuration` and `lastSyncBytes` are not part of `GetVolumeReplicationInfoResponse` in csi-addons/spec v0.2.0, the version this driver builds against) | Positive | `TestGetVolumeReplicationInfo` | -| U-08 | A malformed volume handle: `INVALID_ARGUMENT` before any backend call | Negative | `TestEnableVolumeReplicationMalformedVolumeHandle` | -| U-09 | A cluster ID with no entry in this deployment's secret: `UNAVAILABLE`, distinct from a backend-side refusal | Negative | `TestEnableVolumeReplicationUnknownCluster` | -| U-28 | Enable without the `VolumeReplicationClass` policy parameter: `INVALID_ARGUMENT` before any backend call | Negative | `TestEnableVolumeReplicationMissingPolicyParam` | -| U-29 | Enable the backend refuses (412): `FAILED_PRECONDITION` carrying the backend's own reason | Negative | `TestEnableVolumeReplicationBackendRefusal` | -| U-30 | Info on a volume that never replicated: a nil `lastSyncTime`, never a `NOT_FOUND` | Boundary | `TestGetVolumeReplicationInfoNeverReplicated` | -| U-31 | Info on a volume id the backend does not recognize: `NOT_FOUND` | Negative | `TestGetVolumeReplicationInfoUnknownVolume` | -| U-32 | The remaining Replication verbs (`PromoteVolume`, `DemoteVolume`, `ResyncVolume`) fall through to `UNIMPLEMENTED` until Phase 2 lands | Regression | `TestUnimplementedReplicationVerbsAreUnimplemented` | +| # | Scenario | Type | Test | +|----------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|----------|------------------------------------------------------| +| U-01 | Enable on an unattached volume: the attach call carries the class's policy, and the RPC succeeds | Positive | `TestEnableVolumeReplication` | +| U-02 | Enable on a volume already attached to the same policy: success, no second attach call (idempotency) | Boundary | `TestEnableVolumeReplicationRepeatedIsIdempotent` | +| U-03 | Enable on a volume attached to a different policy: `FAILED_PRECONDITION` naming both policies, no attach call | Negative | — | +| U-04 | Disable on an attached volume: the detach call is made, success | Positive | `TestDisableVolumeReplication` | +| U-05 | Disable on a non-attached volume: success without a backend call (idempotency) | Boundary | `TestDisableVolumeReplicationNotAttachedIsSuccess` | +| U-06 | Disable while a cutover is in flight (backend 409): `ABORTED`, retryable | Negative | `TestDisableVolumeReplicationDuringCutoverIsAborted` | +| U-07 | Info returns `lastSyncTime` from the status read (`lastSyncDuration` and `lastSyncBytes` are not part of `GetVolumeReplicationInfoResponse` in csi-addons/spec v0.2.0, the version this driver builds against) | Positive | `TestGetVolumeReplicationInfo` | +| U-08 | A malformed volume handle: `INVALID_ARGUMENT` before any backend call | Negative | `TestEnableVolumeReplicationMalformedVolumeHandle` | +| U-09 | A cluster ID with no entry in this deployment's secret: `UNAVAILABLE`, distinct from a backend-side refusal | Negative | `TestEnableVolumeReplicationUnknownCluster` | +| U-28 | Enable without the `VolumeReplicationClass` policy parameter: `INVALID_ARGUMENT` before any backend call | Negative | `TestEnableVolumeReplicationMissingPolicyParam` | +| U-29 | Enable the backend refuses (412): `FAILED_PRECONDITION` carrying the backend's own reason | Negative | `TestEnableVolumeReplicationBackendRefusal` | +| U-30 | Info on a volume that never replicated: a nil `lastSyncTime`, never a `NOT_FOUND` | Boundary | `TestGetVolumeReplicationInfoNeverReplicated` | +| U-31 | Info on a volume id the backend does not recognize: `NOT_FOUND` | Negative | `TestGetVolumeReplicationInfoUnknownVolume` | +| ~~U-32~~ | **Retired.** Was "the remaining verbs fall through to `UNIMPLEMENTED`"; moot now that Phase 2 implements all six `csi-addons/spec` v0.2.0 methods, so `TestUnimplementedReplicationVerbsAreUnimplemented` was deleted rather than kept red. | — | — | ### Replication Verbs: Promote, Demote, Resync (design §5.2) -File: `csi-driver/internal/csi/controller/replication_lifecycle_test.go` (planned) - -| # | Scenario | Type | Test | -|------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------|----------|------| -| U-10 | Forced promote: the failover endpoint is called, and `Completed` follows the relationship reaching `failed_over` | Positive | — | -| U-11 | Forced promote repeated after completion: success without a second failover (backend idempotency honored) | Boundary | — | -| U-12 | Planned promote after a completed demote: the failover endpoint is called with the planned gate, `Completed=True` once the relationship reports the target active | Positive | — | -| U-13 | Planned promote without a completed demote (source live or peer lagging): the planned gate refuses, `FAILED_PRECONDITION` naming the un-flushed tail | Negative | — | -| U-14 | Demote: the demote endpoint is called, and `Completed` is `True` only after the flush-confirmed response | Positive | — | -| U-15 | Demote repeated on a demoted volume: success (idempotency) | Boundary | — | -| U-16 | Resync: the failback endpoint is called and `Resyncing=True` while the status read reports the catch-up | Positive | — | -| U-17 | Resync completion: `Resyncing` clears when the reverse lag is inside the budget | Boundary | — | +File: `csi-driver/internal/csi/controller/replication_lifecycle_test.go` + +| # | Scenario | Type | Test | +|------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| U-10 | Forced promote: the failover endpoint is called without `planned=true`, ignoring demote state entirely | Positive | `TestPromoteVolumeForced` | +| U-11 | Forced promote repeated after completion: success without a second failover (backend idempotency honored) | Boundary | — | +| U-12 | Planned promote sends `planned=true` on the failover call | Positive | `TestPromoteVolumePlannedSendsThePlannedFlag` | +| U-13 | Planned promote with no demote ever requested: `FAILED_PRECONDITION`, the one case meant to let the vendored controller's own force-escalation take over | Negative | `TestPromoteVolumePlannedWithNoDemoteIsFailedPrecondition` | +| U-14 | Demote: the demote endpoint is called, and the RPC succeeds only once the backend confirms the final snapshot landed | Positive | `TestDemoteVolumeDone` | +| U-15 | Demote repeated on a demoted volume: success (idempotency) | Boundary | `sbcli: test_demote_is_idempotent_once_done` (backend tier; the driver's `TestDemoteVolumeDone` exercises the same success path) | +| U-16 | Resync: the failback endpoint is called, forwarding the `sourceClusterID` class parameter when given | Positive | `TestResyncVolume`, `TestResyncVolumeForwardsTheSourceClusterParameter` | +| U-17 | Resync's `ready` field reflects lag against budget: false while lag exceeds it, true once caught up | Boundary | `TestResyncVolumeNotReadyWhileLagExceedsBudget`, `TestResyncVolumeReadyReflectsLag` | +| U-40 | Planned promote while demote is still converging: `ABORTED` (retryable) — never `FAILED_PRECONDITION`, which the vendored controller auto-escalates to a forced, lossy promote inline with no wait-and-retry grace period of its own | Negative | `TestPromoteVolumePlannedWhileDemoteConvergingIsAborted` | +| U-41 | Demote still converging: `ABORTED` (retryable), non-blocking — the RPC never waits out the backend's own convergence loop | Boundary | `TestDemoteVolumeNotYetDoneIsAborted` | +| U-42 | `demote_lvol` fences the source strictly before triggering the final snapshot, never after (a write landing in the gap would be silently lost) | Positive | `sbcli: test_demote_fences_before_triggering_the_final_snapshot` | +| U-43 | `demote_lvol` re-invoked while pending: checks the marker only, never re-fences or re-triggers | Boundary | `sbcli: test_demote_does_not_refence_or_retrigger_once_pending` | +| U-44 | `demote_lvol` completes once the triggered snapshot carries the replicated marker | Positive | `sbcli: test_demote_completes_once_the_snapshot_carries_the_replicated_marker` | +| U-45 | `demote_lvol` surfaces a snapshot-creation failure without recording pending state | Negative | `sbcli: test_demote_surfaces_a_snapshot_creation_failure` | +| U-46 | The `failover` route's three-way planned-gate branch: demoted proceeds, converging is 409, no relationship is 412 | Positive/Negative | `sbcli: test_planned_failover_proceeds_once_demoted`, `test_planned_failover_while_demote_is_converging_is_409`, `test_planned_failover_without_any_demote_is_412` | +| U-47 | Unplanned failover ignores demote state, unchanged from before P0-3 | Regression | `sbcli: test_unplanned_failover_ignores_demote_state` | ### Condition Derivation (design §6.2) @@ -165,28 +173,28 @@ Two live simplyblock clusters with the chart-deployed csi-addons machinery. The ## 5. Axis Coverage -| Axis | Values covered | IDs | Not covered | -|------------------|------------------------------------------------------------------------|---------------------------------------|-----------------------------------------------------------| -| Verb lifecycle | enable, disable, info, forced promote, planned promote, demote, resync | U-01 … U-17, I-02 … I-04, E-01 … E-04 | — | -| Idempotency | repeat enable, disable, promote, demote; re-drive after restart | U-02, U-05, U-11, U-15, I-07 | repeated resync | -| Conditions | healthy, degraded, error, staleness, resyncing, disabled | U-18 … U-22, E-05 | condition behavior across backend upgrade | -| Coexistence | slot skip, one-owner refusal, concurrent claim | U-26, U-27, M-02 | migration of an annotated volume onto a VolumeReplication | -| peerClasses | verified, missing class, mispaired policies | U-23 … U-25 | drift after verification | -| Orchestrator | direct kubectl lifecycle, Ramen VRG async | I-02 … I-06, E-06, E-07 | Ramen hub failover of multiple apps | -| Cluster topology | two clusters, one relationship addressed from both sides | E-01 … E-07 | three-cluster (cascaded) topologies | +| Axis | Values covered | IDs | Not covered | +|------------------|------------------------------------------------------------------------|-----------------------------------------------------------------------------|-----------------------------------------------------------| +| Verb lifecycle | enable, disable, info, forced promote, planned promote, demote, resync | U-01, U-02, U-04 … U-10, U-12 … U-17, U-40 … U-47, I-02 … I-04, E-01 … E-04 | — | +| Idempotency | repeat enable, disable, demote; re-drive after restart | U-02, U-05, U-15, I-07 | repeated promote (U-11), repeated resync | +| Conditions | healthy, degraded, error, staleness, resyncing, disabled | U-18 … U-22, E-05 | condition behavior across backend upgrade | +| Coexistence | slot skip, one-owner refusal, concurrent claim | U-26, U-27, M-02 | migration of an annotated volume onto a VolumeReplication | +| peerClasses | verified, missing class, mispaired policies | U-23 … U-25 | drift after verification | +| Orchestrator | direct kubectl lifecycle, Ramen VRG async | I-02 … I-06, E-06, E-07 | Ramen hub failover of multiple apps | +| Cluster topology | two clusters, one relationship addressed from both sides | E-01 … E-07 | three-cluster (cascaded) topologies | --- ## 6. Coverage Summary -| Class | Scenarios | Covered | Not covered | -|-------------|-----------|---------|-------------------| -| Unit | 39 | 20 | U-03, U-10 … U-27 | -| Integration | 7 | 0 | I-01 … I-07 | -| E2E | 7 | 0 | E-01 … E-07 | -| Manual | 2 | 0 | M-01, M-02 | +| Class | Scenarios | Covered | Not covered | +|-------------|-----------|---------|-------------------------| +| Unit | 46 | 34 | U-03, U-11, U-18 … U-27 | +| Integration | 7 | 0 | I-01 … I-07 | +| E2E | 7 | 0 | E-01 … E-07 | +| Manual | 2 | 0 | M-01, M-02 | -Phase 1 landed the driver's Replication and Identity services, the error classifier, and the operator's sidecar and RBAC wiring, covering every Phase 1 unit scenario except U-03 (§7). Phase 2 (promote, demote, resync, and the operator's preflight and coexistence controllers) has not started, and neither has a sidecar-and-controller-manager integration suite or a live two-cluster E2E bed, so those tiers remain fully uncovered. +Phase 1 landed the driver's Replication and Identity services, the error classifier, and the operator's sidecar and RBAC wiring, covering every Phase 1 unit scenario except U-03 (§7). Phase 2 landed P0-3 (the demote endpoint and the planned gate on `failover`, in sbcli) and the driver's `PromoteVolume`/`DemoteVolume`/`ResyncVolume`, covering every Phase 2 unit scenario except U-11 (§7). The operator's preflight and coexistence controllers (peerClasses, `PVCReplicationController`) remain a separate, unbuilt subsystem, and neither a sidecar-and-controller-manager integration suite nor a live two-cluster E2E bed exists yet, so those tiers remain fully uncovered. --- @@ -195,9 +203,9 @@ Phase 1 landed the driver's Replication and Identity services, the error classif | # | Gap | Reason | |-------------|-------------------------------------------------------------------------------------------------------------------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | U-03 | Different-policy enable refused with `FAILED_PRECONDITION` | Not implemented: sbcli's `attach_policy` silently re-attaches onto the new policy rather than refusing, and no endpoint exposes the policy id a volume is currently attached to, for the driver to compare against before attaching (design §5.1 assumed this refusal exists; it does not against the current backend) | -| U-10 … U-17 | Promote, demote, and resync unit coverage | Phase 2: the three verbs fall through to `UnimplementedControllerServer`; P0-3 (demote's backend support) is also still blocked | +| U-11 | Forced promote repeated after completion: success without a second failover | Not implemented: the backend's own idempotency (`_failover_volumes`' `already failed_over` skip) is real and exercised transitively, but no driver-level test asserts it directly for `PromoteVolume` | | U-18 … U-22 | Condition derivation (`Completed`/`Degraded`/`Resyncing`) from the status read | Not implemented: `csi-addons/spec` v0.2.0's `GetVolumeReplicationInfoResponse` carries only `lastSyncTime`, with no per-condition field at all; deriving these needs either a newer spec version or belongs in the controller-manager's own reconcile, neither examined yet | -| U-23 … U-27 | Preflight (`peerClasses` verification) and coexistence (`PVCReplicationController`, the one-owner rule) | Out of Phase 1's scope: the auto-adapter and preflight webhook are a separate, unbuilt subsystem | +| U-23 … U-27 | Preflight (`peerClasses` verification) and coexistence (`PVCReplicationController`, the one-owner rule) | Out of Phase 1 and Phase 2's scope: the auto-adapter and preflight webhook are a separate, unbuilt subsystem | | I-01 … I-07 | The sidecar and controller-manager loop | The driver's Replication and Identity services and the sidecar container now exist (Phase 1); no envtest/kind suite exercises them against the real kubernetes-csi-addons controller-manager yet | | E-01 … E-07 | The live lifecycle and the Ramen gate | Blocked on Phase 2 landing, plus a two-cluster test bed with Ramen dr-cluster installed for E-06 and E-07 | | — | Repeated resync, class drift after verification, annotated-volume migration onto the adapter, cascaded topologies | Beyond the first coverage pass, recorded so the gaps are explicit rather than assumed covered | diff --git a/shared/openapi.json b/shared/openapi.json index 11021dcee..12eff74ef 100644 --- a/shared/openapi.json +++ b/shared/openapi.json @@ -4579,7 +4579,7 @@ "replication" ], "summary": "Clusters:Storage-Pools:Volumes:Replication:Failover", - "description": "Bring the volume up on the target cluster.\n\nThe counterpart's id is read back from this volume's replication\nrelationship, its connection paths from the target volume's `connect`.\n\n``generation`` selects WHICH retained point-in-time to come up on: 0 (the\ndefault) is the newest, 1 the one before it, and so on through the\nhistory a retention schedule keeps. Failing over to an older generation\nis the recovery path for a logical corruption, which the newest copy has\nfaithfully replicated.", + "description": "Bring the volume up on the target cluster.\n\nThe counterpart's id is read back from this volume's replication\nrelationship, its connection paths from the target volume's `connect`.\n\n``generation`` selects WHICH retained point-in-time to come up on: 0 (the\ndefault) is the newest, 1 the one before it, and so on through the\nhistory a retention schedule keeps. Failing over to an older generation\nis the recovery path for a logical corruption, which the newest copy has\nfaithfully replicated.\n\n``planned=True`` gates on a completed demote (P0-3) so a planned swap\nloses nothing: 412 when no demote was ever requested for this volume (the\ncaller's premise that the source is reachable to demote was wrong, and a\n412 is what lets the csi-addons controller's own force-escalation take\nover), 409 while demote is still converging (retryable -- 409 must never\nbecome a code the controller reads as permission to force, since that\ncontroller escalates on ANY FAILED_PRECONDITION from a force=false\npromote with no wait-and-retry grace period of its own). Unplanned\nfailover (the default) ignores demote state entirely, unchanged from\ntoday: its whole premise is that the source may never have been\nreachable to demote.", "operationId": "clusters_storage_pools_volumes_replication_failover_api_v2_clusters__cluster_id__storage_pools__pool_id__volumes__volume_id__replication_failover_post", "security": [ { @@ -4626,6 +4626,16 @@ "default": 0, "title": "Generation" } + }, + { + "name": "planned", + "in": "query", + "required": false, + "schema": { + "type": "boolean", + "default": false, + "title": "Planned" + } } ], "responses": { @@ -4724,6 +4734,71 @@ } } }, + "/api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/demote": { + "post": { + "tags": [ + "replication" + ], + "summary": "Clusters:Storage-Pools:Volumes:Replication:Demote", + "description": "Fence the source and confirm the last write replicated (P0-3).\n\nSynchronous and re-drivable, not queued: each call does only the work its\ncurrent state calls for (fence + trigger the final snapshot once, then\njust check whether it has landed), so the caller re-invokes this route\nuntil it reports 204. A 202 means still waiting -- call again, the same\nway `GET .../status` is re-read rather than pushed.", + "operationId": "clusters_storage_pools_volumes_replication_demote_api_v2_clusters__cluster_id__storage_pools__pool_id__volumes__volume_id__replication_demote_post", + "security": [ + { + "HTTPBearer": [] + } + ], + "parameters": [ + { + "name": "cluster_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Cluster Id" + } + }, + { + "name": "pool_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Pool Id" + } + }, + { + "name": "volume_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Volume Id" + } + } + ], + "responses": { + "204": { + "description": "Successful Response" + }, + "202": { + "description": "Accepted" + }, + "422": { + "description": "Validation Error", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/HTTPValidationError" + } + } + } + } + } + } + }, "/api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/failback": { "post": { "tags": [ From bdab53f32eedbd3026789c097262453eb5895b49 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Fri, 18 Sep 2026 12:28:21 +0100 Subject: [PATCH 064/206] support pattern quay.io/csiaddons --- .../csi/controller/replication_test.go | 5 ++--- csi-driver/internal/driver/driver.go | 2 +- ...age.simplyblock.io_simplyblockdrivers.yaml | 8 +++++++- .../api/v1alpha2/simplyblockdriver_types.go | 8 +++++++- ...age.simplyblock.io_simplyblockdrivers.yaml | 8 +++++++- operator/dist/install.yaml | 20 +++++++++++++++++-- ...age.simplyblock.io_simplyblockdrivers.yaml | 8 +++++++- 7 files changed, 49 insertions(+), 10 deletions(-) diff --git a/csi-driver/internal/csi/controller/replication_test.go b/csi-driver/internal/csi/controller/replication_test.go index d27ca3764..c358fcb21 100644 --- a/csi-driver/internal/csi/controller/replication_test.go +++ b/csi-driver/internal/csi/controller/replication_test.go @@ -177,8 +177,8 @@ func TestGetVolumeReplicationInfo(t *testing.T) { mock.replicationStatus[testReplVolumeID] = map[string]any{ "role": "source", "state": "in_sync", "last_replicated_at": "2026-09-17T12:00:00Z", - "lag_seconds": 42, - "outstanding_count": 0, "outstanding_bytes": 0, + "lag_seconds": 42, + "outstanding_count": 0, "outstanding_bytes": 0, "failing_count": 0, "max_retry_reached": false, "resyncing": false, } @@ -225,4 +225,3 @@ func TestGetVolumeReplicationInfoUnknownVolume(t *testing.T) { t.Errorf("code = %v, want NotFound", st.Code()) } } - diff --git a/csi-driver/internal/driver/driver.go b/csi-driver/internal/driver/driver.go index b80731943..ca2111e0f 100644 --- a/csi-driver/internal/driver/driver.go +++ b/csi-driver/internal/driver/driver.go @@ -43,8 +43,8 @@ import ( "github.com/simplyblock/csi-driver/internal/config" csicommon "github.com/simplyblock/csi-driver/internal/csi/common" "github.com/simplyblock/csi-driver/internal/csi/controller" - "github.com/simplyblock/csi-driver/internal/csi/identity" csiaddonsidentityserver "github.com/simplyblock/csi-driver/internal/csi/csiaddons/identity" + "github.com/simplyblock/csi-driver/internal/csi/identity" "github.com/simplyblock/csi-driver/internal/csi/node" "github.com/simplyblock/csi-driver/internal/csilink" "github.com/simplyblock/csi-driver/internal/guardian" diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_simplyblockdrivers.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_simplyblockdrivers.yaml index e5d075f08..62e6cd5d8 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_simplyblockdrivers.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_simplyblockdrivers.yaml @@ -314,7 +314,13 @@ spec: kubernetes-csi-addons controller-manager (design design-csi-addons-replication.md §4.1) can reach the Replication service this driver serves. - pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ + + Unlike the other sidecars above, this one's upstream home is the + csi-addons project's own registry, not simplyblock's: the allowlist + carries quay.io/csiaddons alongside the simplyblock registries so a + deployment can run the stock kubernetes-csi-addons sidecar image + directly, ahead of (or instead of) a quay.io/simplyblock-io mirror. + pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block|quay\.io/csiaddons)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ type: string healthMonitor: description: |- diff --git a/operator/api/v1alpha2/simplyblockdriver_types.go b/operator/api/v1alpha2/simplyblockdriver_types.go index 873c046b1..1b8b55989 100644 --- a/operator/api/v1alpha2/simplyblockdriver_types.go +++ b/operator/api/v1alpha2/simplyblockdriver_types.go @@ -91,7 +91,13 @@ type SidecarImages struct { // kubernetes-csi-addons controller-manager (design // design-csi-addons-replication.md §4.1) can reach the Replication // service this driver serves. - // +kubebuilder:validation:Pattern=`^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$` + // + // Unlike the other sidecars above, this one's upstream home is the + // csi-addons project's own registry, not simplyblock's: the allowlist + // carries quay.io/csiaddons alongside the simplyblock registries so a + // deployment can run the stock kubernetes-csi-addons sidecar image + // directly, ahead of (or instead of) a quay.io/simplyblock-io mirror. + // +kubebuilder:validation:Pattern=`^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block|quay\.io/csiaddons)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$` // +optional CSIAddons string `json:"csiAddons,omitempty"` } diff --git a/operator/config/crd/bases/storage.simplyblock.io_simplyblockdrivers.yaml b/operator/config/crd/bases/storage.simplyblock.io_simplyblockdrivers.yaml index e5d075f08..62e6cd5d8 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_simplyblockdrivers.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_simplyblockdrivers.yaml @@ -314,7 +314,13 @@ spec: kubernetes-csi-addons controller-manager (design design-csi-addons-replication.md §4.1) can reach the Replication service this driver serves. - pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ + + Unlike the other sidecars above, this one's upstream home is the + csi-addons project's own registry, not simplyblock's: the allowlist + carries quay.io/csiaddons alongside the simplyblock registries so a + deployment can run the stock kubernetes-csi-addons sidecar image + directly, ahead of (or instead of) a quay.io/simplyblock-io mirror. + pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block|quay\.io/csiaddons)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ type: string healthMonitor: description: |- diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index 27b6a50b7..2d521f443 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -2547,13 +2547,29 @@ spec: type: object sidecarImages: description: |- - SidecarImages overrides the six CSI sidecars, one field each. Unset takes - the version this operator release ships. + SidecarImages overrides the seven CSI sidecars, one field each. Unset + takes the version this operator release ships. properties: attacher: description: Attacher is csi-attacher, on the controller plugin. pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ type: string + csiAddons: + description: |- + CSIAddons is the kubernetes-csi-addons sidecar, on the controller + plugin. It connects to the plugin's socket, probes the csi-addons + Identity service for capabilities, and publishes a CSIAddonsNode so the + kubernetes-csi-addons controller-manager (design + design-csi-addons-replication.md §4.1) can reach the Replication + service this driver serves. + + Unlike the other sidecars above, this one's upstream home is the + csi-addons project's own registry, not simplyblock's: the allowlist + carries quay.io/csiaddons alongside the simplyblock registries so a + deployment can run the stock kubernetes-csi-addons sidecar image + directly, ahead of (or instead of) a quay.io/simplyblock-io mirror. + pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block|quay\.io/csiaddons)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ + type: string healthMonitor: description: |- HealthMonitor is csi-external-health-monitor-controller, on the diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_simplyblockdrivers.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_simplyblockdrivers.yaml index e5d075f08..62e6cd5d8 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_simplyblockdrivers.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_simplyblockdrivers.yaml @@ -314,7 +314,13 @@ spec: kubernetes-csi-addons controller-manager (design design-csi-addons-replication.md §4.1) can reach the Replication service this driver serves. - pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ + + Unlike the other sidecars above, this one's upstream home is the + csi-addons project's own registry, not simplyblock's: the allowlist + carries quay.io/csiaddons alongside the simplyblock registries so a + deployment can run the stock kubernetes-csi-addons sidecar image + directly, ahead of (or instead of) a quay.io/simplyblock-io mirror. + pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block|quay\.io/csiaddons)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ type: string healthMonitor: description: |- From 8c7984089a8973f6a250bb7831804f8dfe1cbce3 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Fri, 18 Sep 2026 12:41:07 +0100 Subject: [PATCH 065/206] updated simplyblockdriver_controller roles --- .../templates/roles/manager_role.yaml | 32 +++++++++++++++++++ operator/config/rbac/role.yaml | 32 +++++++++++++++++++ .../driver/simplyblockdriver_controller.go | 16 ++++++++++ 3 files changed, 80 insertions(+) diff --git a/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml b/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml index 3ce61450d..029e6a7e5 100644 --- a/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml @@ -161,6 +161,36 @@ rules: - patch - update - watch +- apiGroups: + - coordination.k8s.io + resources: + - leases + verbs: + - create + - delete + - get + - list + - update + - watch +- apiGroups: + - csiaddons.openshift.io + resources: + - csiaddonsnodes + verbs: + - create + - delete + - get + - list + - update + - watch +- apiGroups: + - csiaddons.openshift.io + resources: + - csiaddonsnodes/status + verbs: + - get + - patch + - update - apiGroups: - discovery.k8s.io resources: @@ -207,6 +237,8 @@ rules: resources: - clusterrolebindings - clusterroles + - rolebindings + - roles verbs: - bind - create diff --git a/operator/config/rbac/role.yaml b/operator/config/rbac/role.yaml index f6dadf0e4..bdd175326 100644 --- a/operator/config/rbac/role.yaml +++ b/operator/config/rbac/role.yaml @@ -161,6 +161,36 @@ rules: - patch - update - watch +- apiGroups: + - coordination.k8s.io + resources: + - leases + verbs: + - create + - delete + - get + - list + - update + - watch +- apiGroups: + - csiaddons.openshift.io + resources: + - csiaddonsnodes + verbs: + - create + - delete + - get + - list + - update + - watch +- apiGroups: + - csiaddons.openshift.io + resources: + - csiaddonsnodes/status + verbs: + - get + - patch + - update - apiGroups: - discovery.k8s.io resources: @@ -207,6 +237,8 @@ rules: resources: - clusterrolebindings - clusterroles + - rolebindings + - roles verbs: - bind - create diff --git a/operator/internal/controllers/driver/simplyblockdriver_controller.go b/operator/internal/controllers/driver/simplyblockdriver_controller.go index b0e269689..fe1f166fc 100644 --- a/operator/internal/controllers/driver/simplyblockdriver_controller.go +++ b/operator/internal/controllers/driver/simplyblockdriver_controller.go @@ -66,8 +66,24 @@ const ( // +kubebuilder:rbac:groups=apps,resources=daemonsets;statefulsets,verbs=get;list;watch;create;update;patch;delete // +kubebuilder:rbac:groups="",resources=serviceaccounts;configmaps,verbs=get;list;watch;create;update;patch;delete // +kubebuilder:rbac:groups=rbac.authorization.k8s.io,resources=clusterroles;clusterrolebindings,verbs=get;list;watch;create;update;patch;delete;escalate;bind +// rbac-justified: rbac.go's csiAddonsRole/csiAddonsRoleBinding are a +// namespaced Role, not a ClusterRole (CSIAddonsNode is namespaced) -- this +// mirrors the ClusterRole marker above for the same reason: the manager +// creates RBAC on behalf of the sidecars it deploys, and escalate/bind is +// what Kubernetes' escalation prevention requires to do that, capped by the +// manager's own role, which is why the resources this grants are named +// individually below rather than left open-ended. +// +kubebuilder:rbac:groups=rbac.authorization.k8s.io,resources=roles;rolebindings,verbs=get;list;watch;create;update;patch;delete;escalate;bind // +kubebuilder:rbac:groups=storage.k8s.io,resources=csidrivers,verbs=get;list;watch;create;update;patch;delete // +kubebuilder:rbac:groups=snapshot.storage.k8s.io,resources=volumesnapshotclasses,verbs=get;list;watch;create;update;patch;delete +// rbac-justified: the csi-addons sidecar (rbac.go's csiAddonsRole) needs to +// publish and update its own CSIAddonsNode and hold its own leader-election +// Lease, and the manager creates that namespaced Role on the sidecar's +// behalf -- RBAC escalation prevention requires the manager to already hold +// what it grants, capped by holding only these same resources itself. +// +kubebuilder:rbac:groups=csiaddons.openshift.io,resources=csiaddonsnodes,verbs=get;list;watch;create;update;delete +// +kubebuilder:rbac:groups=csiaddons.openshift.io,resources=csiaddonsnodes/status,verbs=get;update;patch +// +kubebuilder:rbac:groups=coordination.k8s.io,resources=leases,verbs=get;list;watch;create;update;delete // SimplyblockDriverReconciler applies the CSI driver deployment. type SimplyblockDriverReconciler struct { From 0e62ef6143e57dafddda067db0fb100135b3bb28 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Fri, 18 Sep 2026 12:49:50 +0100 Subject: [PATCH 066/206] added node-id env field to csi-addons container --- operator/internal/controllers/driver/workloads.go | 6 ++++++ operator/internal/controllers/driver/workloads_test.go | 8 +++++++- 2 files changed, 13 insertions(+), 1 deletion(-) diff --git a/operator/internal/controllers/driver/workloads.go b/operator/internal/controllers/driver/workloads.go index 79f2cdbea..73d51bca4 100644 --- a/operator/internal/controllers/driver/workloads.go +++ b/operator/internal/controllers/driver/workloads.go @@ -338,6 +338,7 @@ func csiAddonsSidecarContainer( Args: []string{ verbosity, "--csi-addons-address=" + controllerSocketPath, + "--node-id=$(NODE_ID)", "--controller-ip=$(POD_IP)", "--controller-port=" + strconv.Itoa(int(csiAddonsControllerPort)), "--pod=$(POD_NAME)", @@ -346,6 +347,11 @@ func csiAddonsSidecarContainer( "--leader-election-namespace=$(POD_NAMESPACE)", }, Env: []corev1.EnvVar{ + // Required, not cosmetic: csiaddonsnode.Manager.Node ("the + // hostname of the system where the sidecar is running") rejects + // an empty value with "invalid configuration: missing node" + // before it ever creates the CSIAddonsNode object. + fieldRefEnv("NODE_ID", "spec.nodeName"), fieldRefEnv("POD_IP", "status.podIP"), fieldRefEnv("POD_NAME", "metadata.name"), fieldRefEnv("POD_NAMESPACE", "metadata.namespace"), diff --git a/operator/internal/controllers/driver/workloads_test.go b/operator/internal/controllers/driver/workloads_test.go index a2c8b53d5..ab15c45d7 100644 --- a/operator/internal/controllers/driver/workloads_test.go +++ b/operator/internal/controllers/driver/workloads_test.go @@ -128,7 +128,7 @@ func TestCSIAddonsSidecarAdvertisesItsOwnPod(t *testing.T) { } wantFieldPaths := map[string]string{ - "POD_IP": "status.podIP", "POD_NAME": "metadata.name", + "NODE_ID": "spec.nodeName", "POD_IP": "status.podIP", "POD_NAME": "metadata.name", "POD_NAMESPACE": "metadata.namespace", "POD_UID": "metadata.uid", } for _, e := range c.Env { @@ -145,6 +145,12 @@ func TestCSIAddonsSidecarAdvertisesItsOwnPod(t *testing.T) { t.Errorf("no %s env var", name) } + // Required, not cosmetic: csiaddonsnode.Manager.Node rejects an empty + // value with "invalid configuration: missing node" before it ever + // creates the CSIAddonsNode object (confirmed against a live cluster). + if nodeID, ok := argValue(c, "--node-id"); !ok || nodeID != "$(NODE_ID)" { + t.Errorf("--node-id = %q, want $(NODE_ID)", nodeID) + } if ip, ok := argValue(c, "--controller-ip"); !ok || ip != "$(POD_IP)" { t.Errorf("--controller-ip = %q, want $(POD_IP)", ip) } From 3a8dbbe7ad0c7c151d4e91b9f50e838b1f07933a Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Fri, 18 Sep 2026 14:51:13 +0100 Subject: [PATCH 067/206] fixed csiaddon endpoint --- .../designs/design-csi-addons-replication.md | 43 ++++++++++++++----- .../internal/controllers/driver/workloads.go | 20 ++++++--- .../controllers/driver/workloads_test.go | 18 +++++--- 3 files changed, 58 insertions(+), 23 deletions(-) diff --git a/operator/docs/designs/design-csi-addons-replication.md b/operator/docs/designs/design-csi-addons-replication.md index 1871b0e55..31599284f 100644 --- a/operator/docs/designs/design-csi-addons-replication.md +++ b/operator/docs/designs/design-csi-addons-replication.md @@ -24,6 +24,7 @@ The phase numbers above are this document's own, not the DR storage foundation g ## Phase 0 — External Prerequisites +<<<<<<< Updated upstream | # | Prerequisite | Kind | Blocks | Status | |------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------|---------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | P0-1 | A typed, steady-state per-volume replication status read: `GET .../volumes/{id}/replication/status` serving what `lvol_controller.get_replication_info` computes today (state, lag, outstanding bytes, failure counters), available for the volume's whole replicated life | Control plane (`sbcli`) | Phase 1 | Shipped | @@ -32,6 +33,16 @@ The phase numbers above are this document's own, not the DR storage foundation g | P0-4 | An `rpo_target_seconds` field on `ReplicationPolicy`, so RPO compliance is computable against a declared target rather than the derived lag budget | Control plane (`sbcli`) | Phase 3 | Shipped | | P0-5 | csi-addons upstream: the `VolumeReplication` and `VolumeReplicationClass` CRDs (`replication.storage.openshift.io/v1alpha1`), the kubernetes-csi-addons controller-manager image, and the csi-addons sidecar image | Ecosystem | Phase 1 | Vendored in the chart at v0.15.0 behind `csiaddons.create` (all twelve upstream CRDs, since the stock manager starts a controller per kind); sidecar wiring shipped in Phase 1 | | P0-6 | A latest-replicated-snapshot read: per volume, and per consistency group as one complete generation, the newest fully replicated snapshot on the secondary addressed as a cloneable object | Control plane (`sbcli`) | Phase 4 | Shipped | +======= +| # | Prerequisite | Kind | Blocks | Status | +|------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------|---------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| P0-1 | A typed, steady-state per-volume replication status read: `GET .../volumes/{id}/replication/status` serving what `lvol_controller.get_replication_info` computes today (state, lag, outstanding bytes, failure counters), available for the volume's whole replicated life | Control plane (`sbcli`) | Phase 1 | Not shipped | +| P0-2 | Idempotent attach and detach: attaching a volume to the policy it already follows returns success, and detaching a non-attached volume returns success | Control plane (`sbcli`) | Phase 1 | Not shipped | +| P0-3 | A standalone demote verb: `POST .../volumes/{id}/replication/demote` that converges the peer while still serving (repeated snapshot-and-ship until the remaining delta is small), then quiesces, ships the final delta, confirms it landed on the peer, and fences the data path | Control plane (`sbcli`) | Phase 2 | Not shipped | +| P0-4 | An `rpo_target_seconds` field on `ReplicationPolicy`, so RPO compliance is computable against a declared target rather than the derived lag budget | Control plane (`sbcli`) | Phase 3 | Not shipped | +| P0-5 | csi-addons upstream: the `VolumeReplication` and `VolumeReplicationClass` CRDs (`replication.storage.openshift.io/v1alpha1`), the kubernetes-csi-addons controller-manager image, and the csi-addons sidecar image | Ecosystem | Phase 1 | Vendored in the chart at v0.15.0 behind `csiaddons.create` (all twelve upstream CRDs, since the stock manager starts a controller per kind); sidecar wiring is Phase 1 | +| P0-6 | A latest-replicated-snapshot read: per volume, and per consistency group as one complete generation, the newest fully replicated snapshot on the secondary addressed as a cloneable object | Control plane (`sbcli`) | Phase 4 | Shipped: `GET .../replication/relationships/{lvol}/latest-snapshot` and `GET .../replication/policies/{policy}/latest-generation` | +>>>>>>> Stashed changes Everything else the adapter needs already exists: the attach and detach calls, failover, the failback and commit pair, the relationship read, and the backlog arithmetic inside `get_replication_info`. The adapter is thin precisely because the engine is complete. What is missing is the shape Ramen can drive. @@ -309,17 +320,18 @@ The consolidation direction (§13) is that the annotation path becomes a compati Every endpoint is scoped as today: volume-scoped under `/api/v2/clusters/{c}/storage-pools/{p}/volumes/{v}`, cluster-scoped under `/api/v2/clusters/{c}/replication`. -| Method | Endpoint | Notes | -|--------|---------------------------------------------------------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| `PUT` | `.../volumes/{v}` (`replication_policy_id`) | Existing attach and detach. P0-2 makes both idempotent: same-policy attach and non-attached detach return success. | -| `GET` | `.../volumes/{v}/replication/status` | **New (P0-1).** The typed steady-state status of §6.1. Never 404s for a volume that exists; `state: not_replicating, role: none` is a valid answer. | -| `POST` | `.../volumes/{v}/replication/failover` (+ planned gate) | Existing; the one promote, both forms. Gains a planned form that is refused unless the peer holds every acknowledged write (a completed demote). Idempotent by NQN probe. | -| `POST` | `.../volumes/{v}/replication/demote` | **New (P0-3).** Quiesce, final ship, confirm on peer, fence. Idempotent: demoting a demoted volume returns success. | -| `POST` | `.../volumes/{v}/replication/failback` | Existing. Resync (direction reversal, delta-seeded). | -| `POST` | `.../volumes/{v}/replication/commit` | Existing, unchanged, and NOT part of this contract: it stays behind the legacy `ReplicationOps` migration path only (§13). | -| `GET` | `.../replication/relationships/{lvol}` | Existing, unchanged. Cutover records only; the node redirect depends on its survive-deletion semantics. | -| `GET` | `.../replication/relationships/{lvol}/latest-snapshot` | **New (P0-6).** The newest fully replicated snapshot for the volume, as a cloneable snapshot handle. A consistency-group form returns one complete generation's member snapshots. Exposes what the failover path already computes internally. | -| `POST` | `.../volumes/{v}/replication/cutover-proceed` | Existing, unchanged, legacy path only: the adapter never reaches it, because the commit cutover is off this contract. | +| Method | Endpoint | Notes | +|--------|---------------------------------------------------------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| `PUT` | `.../volumes/{v}` (`replication_policy_id`) | Existing attach and detach. P0-2 makes both idempotent: same-policy attach and non-attached detach return success. | +| `GET` | `.../volumes/{v}/replication/status` | **New (P0-1).** The typed steady-state status of §6.1. Never 404s for a volume that exists; `state: not_replicating, role: none` is a valid answer. | +| `POST` | `.../volumes/{v}/replication/failover` (+ planned gate) | Existing; the one promote, both forms. Gains a planned form that is refused unless the peer holds every acknowledged write (a completed demote). Idempotent by NQN probe. | +| `POST` | `.../volumes/{v}/replication/demote` | **New (P0-3).** Quiesce, final ship, confirm on peer, fence. Idempotent: demoting a demoted volume returns success. | +| `POST` | `.../volumes/{v}/replication/failback` | Existing. Resync (direction reversal, delta-seeded). | +| `POST` | `.../volumes/{v}/replication/commit` | Existing, unchanged, and NOT part of this contract: it stays behind the legacy `ReplicationOps` migration path only (§13). | +| `GET` | `.../replication/relationships/{lvol}` | Existing, unchanged. Cutover records only; the node redirect depends on its survive-deletion semantics. | +| `GET` | `.../replication/relationships/{lvol}/latest-snapshot` | **Shipped (P0-6).** The newest fully replicated snapshot for the volume, as a cloneable snapshot handle. Exposes what the failover path already computes internally. | +| `GET` | `.../replication/policies/{policy}/latest-generation` | **Shipped (P0-6), group form.** One complete, fully replicated consistency-group generation, every member as a cloneable snapshot handle. Reuses the group fail-over's mixed-generation refusal instead of its side effect. | +| `POST` | `.../volumes/{v}/replication/cutover-proceed` | Existing, unchanged, legacy path only: the adapter never reaches it, because the commit cutover is off this contract. | The unused backend verbs the operator never calls (`start`, `stop`, `trigger`, `tasks`) are unaffected, and `start` and `stop` remain the policy-less legacy path. @@ -437,9 +449,18 @@ A drill that silently perturbed replication would be worse than no drill. Before ## 15. Open Questions +<<<<<<< Updated upstream | # | Question | Owner | |-------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------| | ~~1~~ | **Resolved.** Demote fences (ANA inaccessible) BEFORE taking the final snapshot, mirroring `fence_source_paths`'s existing invariant, and does not block on in-flight writes: a write that lands after the fence queues against the now-dark path (NVMe multipath semantics) and fails via the client's own `ctrl_loss_tmo`, since nothing in the demote-only flow ever lights a target path to drain it onto -- that is a separate, later `PromoteVolume` call. Demote's own fence-and-snapshot sequence proceeds independently of whatever is queued client-side; it never waits on it. This is a fail choice, not a block, and it trusts Ramen's unmount-first contract rather than adding a defensive "no live session" check before fencing -- a genuine simplification, flagged in case a caller outside Ramen's contract ever drives this verb directly. | Backend team (resolved) | | 2 | **Where the preflight lives.** §7.2 attaches peerClasses validation to the `ReplicationPair` reconciler. If the redesign retires the pair kind, the preflight needs a new home (the `SimplyblockDriver`, or a standalone check job). | Operator team | | 3 | **Different-policy enable refusal has no data source.** §5.1's `EnableVolumeReplication` was designed to refuse an attach to a different policy with `FAILED_PRECONDITION`. Phase 1 found that sbcli's `attach_policy` silently re-attaches instead of refusing, and no endpoint returns the policy id a volume is currently attached to, for the driver to compare against. Either the backend adds that read, or this design accepts the silent re-sync as the behavior. | Backend team | | 4 | **`lastSyncDuration` and `lastSyncBytes` have no response field.** §5.1's `GetVolumeReplicationInfo` was designed to return all three fields; `csi-addons/spec` v0.2.0's `GetVolumeReplicationInfoResponse` carries only `lastSyncTime`. Confirm whether a newer spec version adds the other two, or whether they surface some other way (a `VolumeReplication` annotation, a metric). | Ecosystem/driver team | +======= +| # | Question | Owner | +|---|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------| +| 1 | **Demote semantics for the application.** The P0-3 demote fences the volume (ANA inaccessible) after the final flush, and with convergence folded into the verb it is now the only place a planned swap can stall. This is also the one verb the planned promote's lossless guarantee entirely depends on (§5.2): a planned promote is refused unless a completed demote already fenced the source and confirmed the final delta landed, so an unresolved failure mode here is an unresolved gap in the whole "zero loss" claim. Ramen relocation unmounts the workload first, so the fence is ordinarily unopposed, but that is Ramen's choreography, not a guarantee the driver can rely on: a stuck termination, a stale mount that never released, or a demote invoked outside Ramen's normal flow can all leave writes still arriving when quiesce fires. Confirm the verb's behavior when writes are still in flight at quiesce (block versus fail), whether the converge phase has its own budget separate from the quiesced flush, and whether a timeout in either phase must abort back to serving primary or leave the volume fenced with no automatic recovery. | Backend team | +| 2 | **Where the preflight lives.** §7.2 attaches peerClasses validation to the `ReplicationPair` reconciler. If the redesign retires the pair kind, the preflight needs a new home (the `SimplyblockDriver`, or a standalone check job). | Operator team | +| 3 | **Per-volume policy granularity.** A `VolumeReplicationClass` names one policy, and today one policy implies one target and cadence for all its volumes. Confirm one class per (policy, cadence) is an acceptable authoring model for Ramen's `replicationClassSelector`, or whether per-volume interval overrides are needed. | Operator / Backend team | +| 4 | **Visibility of `drtest-` clones.** The test-cluster mode's clones on the secondary are replication sources only, never served. Decide whether the backend creates them as internal volumes (hidden from listings, exempt from the per-node subsystem cap, like the shipping path's landing volumes) or as ordinary volumes under a naming convention. | Backend team | +>>>>>>> Stashed changes diff --git a/operator/internal/controllers/driver/workloads.go b/operator/internal/controllers/driver/workloads.go index 73d51bca4..70ca95027 100644 --- a/operator/internal/controllers/driver/workloads.go +++ b/operator/internal/controllers/driver/workloads.go @@ -324,10 +324,20 @@ const csiAddonsControllerPort int32 = 9070 // naming this pod so the manager can find it, and leader-elects across // replicas of this StatefulSet before serving controller requests. // -// The pod runs on the host network (controllerStatefulSet), so the sidecar -// advertises its own pod IP rather than a Service DNS name, the same way the -// existing csi-provisioner/-snapshotter/etc. sidecars address the plugin's -// socket by path instead of by name. +// The endpoint the sidecar advertises MUST be the pod://. +// form, never a bare ip:port: the controller-manager's own resolveEndpoint +// (internal/controller/csiaddons/csiaddonsnode_controller.go, v0.15.0) +// hard-requires url.Parse's Scheme to equal "pod" and rejects everything +// else with "endpoint scheme %q not supported" -- confirmed against a live +// cluster, where passing --controller-ip produced the OTHER branch of the +// sidecar's own BuildEndpointURL (a bare ":", no scheme at all) +// and the controller-manager failed every connection attempt with "first +// path segment in URL cannot contain colon" until it gave up and deleted +// the CSIAddonsNode. Omitting --controller-ip is the fix, not a missing +// feature: the manager resolves the pod's CURRENT IP itself via a live API +// read at connection time (see resolveEndpoint), which is more robust than +// baking a static IP into the object anyway -- it survives the pod +// restarting with a new IP without anyone having to update anything. func csiAddonsSidecarContainer( d *simplyblockv1alpha2.SimplyblockDriver, image string, mounts []corev1.VolumeMount, ) corev1.Container { @@ -339,7 +349,6 @@ func csiAddonsSidecarContainer( verbosity, "--csi-addons-address=" + controllerSocketPath, "--node-id=$(NODE_ID)", - "--controller-ip=$(POD_IP)", "--controller-port=" + strconv.Itoa(int(csiAddonsControllerPort)), "--pod=$(POD_NAME)", "--namespace=$(POD_NAMESPACE)", @@ -352,7 +361,6 @@ func csiAddonsSidecarContainer( // an empty value with "invalid configuration: missing node" // before it ever creates the CSIAddonsNode object. fieldRefEnv("NODE_ID", "spec.nodeName"), - fieldRefEnv("POD_IP", "status.podIP"), fieldRefEnv("POD_NAME", "metadata.name"), fieldRefEnv("POD_NAMESPACE", "metadata.namespace"), fieldRefEnv("POD_UID", "metadata.uid"), diff --git a/operator/internal/controllers/driver/workloads_test.go b/operator/internal/controllers/driver/workloads_test.go index ab15c45d7..bc29fecfa 100644 --- a/operator/internal/controllers/driver/workloads_test.go +++ b/operator/internal/controllers/driver/workloads_test.go @@ -117,9 +117,13 @@ func TestCSIAddonsSidecarIsAppliedAfterThePlugin(t *testing.T) { } } -// The sidecar advertises this pod, not a Service DNS name: the StatefulSet runs -// on the host network, so its own pod IP is what the kubernetes-csi-addons -// controller-manager can actually reach. +// The sidecar advertises itself as pod://., never a bare +// ip:port: the controller-manager's own resolveEndpoint (v0.15.0) hard-requires +// url.Parse's Scheme to equal "pod" and rejects everything else with +// "endpoint scheme %q not supported" -- confirmed against a live cluster, +// where passing --controller-ip produced a bare ":" (no scheme) +// that the controller-manager could never parse, so it kept failing every +// connection attempt and deleting the CSIAddonsNode in a tight loop. func TestCSIAddonsSidecarAdvertisesItsOwnPod(t *testing.T) { d := testDriver("simplyblock") c := containerNamed(controllerStatefulSet(d, testImage).Spec.Template.Spec.Containers, "csi-addons") @@ -128,7 +132,7 @@ func TestCSIAddonsSidecarAdvertisesItsOwnPod(t *testing.T) { } wantFieldPaths := map[string]string{ - "NODE_ID": "spec.nodeName", "POD_IP": "status.podIP", "POD_NAME": "metadata.name", + "NODE_ID": "spec.nodeName", "POD_NAME": "metadata.name", "POD_NAMESPACE": "metadata.namespace", "POD_UID": "metadata.uid", } for _, e := range c.Env { @@ -151,8 +155,10 @@ func TestCSIAddonsSidecarAdvertisesItsOwnPod(t *testing.T) { if nodeID, ok := argValue(c, "--node-id"); !ok || nodeID != "$(NODE_ID)" { t.Errorf("--node-id = %q, want $(NODE_ID)", nodeID) } - if ip, ok := argValue(c, "--controller-ip"); !ok || ip != "$(POD_IP)" { - t.Errorf("--controller-ip = %q, want $(POD_IP)", ip) + // --controller-ip must NEVER be set: it takes BuildEndpointURL's bare + // ip:port branch, which this controller-manager version cannot parse. + if _, ok := argValue(c, "--controller-ip"); ok { + t.Error("--controller-ip is set; the pod:// addressing scheme requires omitting it") } if ns, ok := argValue(c, "--leader-election-namespace"); !ok || ns != "$(POD_NAMESPACE)" { t.Errorf("--leader-election-namespace = %q, want $(POD_NAMESPACE)", ns) From 989a49b3780555774202c36326276ad6964efda3 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Fri, 18 Sep 2026 14:53:25 +0100 Subject: [PATCH 068/206] fixed merge conflict in design doc --- .../designs/design-csi-addons-replication.md | 20 ------------------- 1 file changed, 20 deletions(-) diff --git a/operator/docs/designs/design-csi-addons-replication.md b/operator/docs/designs/design-csi-addons-replication.md index 31599284f..a061ac5d8 100644 --- a/operator/docs/designs/design-csi-addons-replication.md +++ b/operator/docs/designs/design-csi-addons-replication.md @@ -24,16 +24,6 @@ The phase numbers above are this document's own, not the DR storage foundation g ## Phase 0 — External Prerequisites -<<<<<<< Updated upstream -| # | Prerequisite | Kind | Blocks | Status | -|------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------|---------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| P0-1 | A typed, steady-state per-volume replication status read: `GET .../volumes/{id}/replication/status` serving what `lvol_controller.get_replication_info` computes today (state, lag, outstanding bytes, failure counters), available for the volume's whole replicated life | Control plane (`sbcli`) | Phase 1 | Shipped | -| P0-2 | Idempotent attach and detach: attaching a volume to the policy it already follows returns success, and detaching a non-attached volume returns success | Control plane (`sbcli`) | Phase 1 | Shipped | -| P0-3 | A standalone demote verb: `POST .../volumes/{id}/replication/demote` that fences the source, ships a final internal snapshot, and confirms it landed on the peer before recording the volume demoted | Control plane (`sbcli`) | Phase 2 | Shipped (narrower than originally scoped here -- see §5.2's note on why it does not reuse the live-cutover engine) | -| P0-4 | An `rpo_target_seconds` field on `ReplicationPolicy`, so RPO compliance is computable against a declared target rather than the derived lag budget | Control plane (`sbcli`) | Phase 3 | Shipped | -| P0-5 | csi-addons upstream: the `VolumeReplication` and `VolumeReplicationClass` CRDs (`replication.storage.openshift.io/v1alpha1`), the kubernetes-csi-addons controller-manager image, and the csi-addons sidecar image | Ecosystem | Phase 1 | Vendored in the chart at v0.15.0 behind `csiaddons.create` (all twelve upstream CRDs, since the stock manager starts a controller per kind); sidecar wiring shipped in Phase 1 | -| P0-6 | A latest-replicated-snapshot read: per volume, and per consistency group as one complete generation, the newest fully replicated snapshot on the secondary addressed as a cloneable object | Control plane (`sbcli`) | Phase 4 | Shipped | -======= | # | Prerequisite | Kind | Blocks | Status | |------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------|---------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------| | P0-1 | A typed, steady-state per-volume replication status read: `GET .../volumes/{id}/replication/status` serving what `lvol_controller.get_replication_info` computes today (state, lag, outstanding bytes, failure counters), available for the volume's whole replicated life | Control plane (`sbcli`) | Phase 1 | Not shipped | @@ -42,7 +32,6 @@ The phase numbers above are this document's own, not the DR storage foundation g | P0-4 | An `rpo_target_seconds` field on `ReplicationPolicy`, so RPO compliance is computable against a declared target rather than the derived lag budget | Control plane (`sbcli`) | Phase 3 | Not shipped | | P0-5 | csi-addons upstream: the `VolumeReplication` and `VolumeReplicationClass` CRDs (`replication.storage.openshift.io/v1alpha1`), the kubernetes-csi-addons controller-manager image, and the csi-addons sidecar image | Ecosystem | Phase 1 | Vendored in the chart at v0.15.0 behind `csiaddons.create` (all twelve upstream CRDs, since the stock manager starts a controller per kind); sidecar wiring is Phase 1 | | P0-6 | A latest-replicated-snapshot read: per volume, and per consistency group as one complete generation, the newest fully replicated snapshot on the secondary addressed as a cloneable object | Control plane (`sbcli`) | Phase 4 | Shipped: `GET .../replication/relationships/{lvol}/latest-snapshot` and `GET .../replication/policies/{policy}/latest-generation` | ->>>>>>> Stashed changes Everything else the adapter needs already exists: the attach and detach calls, failover, the failback and commit pair, the relationship read, and the backlog arithmetic inside `get_replication_info`. The adapter is thin precisely because the engine is complete. What is missing is the shape Ramen can drive. @@ -449,18 +438,9 @@ A drill that silently perturbed replication would be worse than no drill. Before ## 15. Open Questions -<<<<<<< Updated upstream -| # | Question | Owner | -|-------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------| -| ~~1~~ | **Resolved.** Demote fences (ANA inaccessible) BEFORE taking the final snapshot, mirroring `fence_source_paths`'s existing invariant, and does not block on in-flight writes: a write that lands after the fence queues against the now-dark path (NVMe multipath semantics) and fails via the client's own `ctrl_loss_tmo`, since nothing in the demote-only flow ever lights a target path to drain it onto -- that is a separate, later `PromoteVolume` call. Demote's own fence-and-snapshot sequence proceeds independently of whatever is queued client-side; it never waits on it. This is a fail choice, not a block, and it trusts Ramen's unmount-first contract rather than adding a defensive "no live session" check before fencing -- a genuine simplification, flagged in case a caller outside Ramen's contract ever drives this verb directly. | Backend team (resolved) | -| 2 | **Where the preflight lives.** §7.2 attaches peerClasses validation to the `ReplicationPair` reconciler. If the redesign retires the pair kind, the preflight needs a new home (the `SimplyblockDriver`, or a standalone check job). | Operator team | -| 3 | **Different-policy enable refusal has no data source.** §5.1's `EnableVolumeReplication` was designed to refuse an attach to a different policy with `FAILED_PRECONDITION`. Phase 1 found that sbcli's `attach_policy` silently re-attaches instead of refusing, and no endpoint returns the policy id a volume is currently attached to, for the driver to compare against. Either the backend adds that read, or this design accepts the silent re-sync as the behavior. | Backend team | -| 4 | **`lastSyncDuration` and `lastSyncBytes` have no response field.** §5.1's `GetVolumeReplicationInfo` was designed to return all three fields; `csi-addons/spec` v0.2.0's `GetVolumeReplicationInfoResponse` carries only `lastSyncTime`. Confirm whether a newer spec version adds the other two, or whether they surface some other way (a `VolumeReplication` annotation, a metric). | Ecosystem/driver team | -======= | # | Question | Owner | |---|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------| | 1 | **Demote semantics for the application.** The P0-3 demote fences the volume (ANA inaccessible) after the final flush, and with convergence folded into the verb it is now the only place a planned swap can stall. This is also the one verb the planned promote's lossless guarantee entirely depends on (§5.2): a planned promote is refused unless a completed demote already fenced the source and confirmed the final delta landed, so an unresolved failure mode here is an unresolved gap in the whole "zero loss" claim. Ramen relocation unmounts the workload first, so the fence is ordinarily unopposed, but that is Ramen's choreography, not a guarantee the driver can rely on: a stuck termination, a stale mount that never released, or a demote invoked outside Ramen's normal flow can all leave writes still arriving when quiesce fires. Confirm the verb's behavior when writes are still in flight at quiesce (block versus fail), whether the converge phase has its own budget separate from the quiesced flush, and whether a timeout in either phase must abort back to serving primary or leave the volume fenced with no automatic recovery. | Backend team | | 2 | **Where the preflight lives.** §7.2 attaches peerClasses validation to the `ReplicationPair` reconciler. If the redesign retires the pair kind, the preflight needs a new home (the `SimplyblockDriver`, or a standalone check job). | Operator team | | 3 | **Per-volume policy granularity.** A `VolumeReplicationClass` names one policy, and today one policy implies one target and cadence for all its volumes. Confirm one class per (policy, cadence) is an acceptable authoring model for Ramen's `replicationClassSelector`, or whether per-volume interval overrides are needed. | Operator / Backend team | | 4 | **Visibility of `drtest-` clones.** The test-cluster mode's clones on the secondary are replication sources only, never served. Decide whether the backend creates them as internal volumes (hidden from listings, exempt from the per-node subsystem cap, like the shipping path's landing volumes) or as ordinary volumes under a naming convention. | Backend team | ->>>>>>> Stashed changes From 7d5d6638eef7d89ff6e51678667f771e97650e80 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Fri, 18 Sep 2026 15:13:29 +0100 Subject: [PATCH 069/206] added rbac verb TokenReview --- .../designs/design-csi-addons-replication.md | 6 ++-- operator/internal/controllers/driver/rbac.go | 35 +++++++++++++++++++ .../internal/controllers/driver/rbac_test.go | 24 +++++++++++++ .../driver/simplyblockdriver_controller.go | 4 +-- .../simplyblockdriver_controller_test.go | 2 +- 5 files changed, 65 insertions(+), 6 deletions(-) diff --git a/operator/docs/designs/design-csi-addons-replication.md b/operator/docs/designs/design-csi-addons-replication.md index a061ac5d8..65b825445 100644 --- a/operator/docs/designs/design-csi-addons-replication.md +++ b/operator/docs/designs/design-csi-addons-replication.md @@ -441,6 +441,6 @@ A drill that silently perturbed replication would be worse than no drill. Before | # | Question | Owner | |---|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------| | 1 | **Demote semantics for the application.** The P0-3 demote fences the volume (ANA inaccessible) after the final flush, and with convergence folded into the verb it is now the only place a planned swap can stall. This is also the one verb the planned promote's lossless guarantee entirely depends on (§5.2): a planned promote is refused unless a completed demote already fenced the source and confirmed the final delta landed, so an unresolved failure mode here is an unresolved gap in the whole "zero loss" claim. Ramen relocation unmounts the workload first, so the fence is ordinarily unopposed, but that is Ramen's choreography, not a guarantee the driver can rely on: a stuck termination, a stale mount that never released, or a demote invoked outside Ramen's normal flow can all leave writes still arriving when quiesce fires. Confirm the verb's behavior when writes are still in flight at quiesce (block versus fail), whether the converge phase has its own budget separate from the quiesced flush, and whether a timeout in either phase must abort back to serving primary or leave the volume fenced with no automatic recovery. | Backend team | -| 2 | **Where the preflight lives.** §7.2 attaches peerClasses validation to the `ReplicationPair` reconciler. If the redesign retires the pair kind, the preflight needs a new home (the `SimplyblockDriver`, or a standalone check job). | Operator team | -| 3 | **Per-volume policy granularity.** A `VolumeReplicationClass` names one policy, and today one policy implies one target and cadence for all its volumes. Confirm one class per (policy, cadence) is an acceptable authoring model for Ramen's `replicationClassSelector`, or whether per-volume interval overrides are needed. | Operator / Backend team | -| 4 | **Visibility of `drtest-` clones.** The test-cluster mode's clones on the secondary are replication sources only, never served. Decide whether the backend creates them as internal volumes (hidden from listings, exempt from the per-node subsystem cap, like the shipping path's landing volumes) or as ordinary volumes under a naming convention. | Backend team | +| 2 | **Where the preflight lives.** §7.2 attaches peerClasses validation to the `ReplicationPair` reconciler. If the redesign retires the pair kind, the preflight needs a new home (the `SimplyblockDriver`, or a standalone check job). | Operator team | +| 3 | **Per-volume policy granularity.** A `VolumeReplicationClass` names one policy, and today one policy implies one target and cadence for all its volumes. Confirm one class per (policy, cadence) is an acceptable authoring model for Ramen's `replicationClassSelector`, or whether per-volume interval overrides are needed. | Operator / Backend team | +| 4 | **Visibility of `drtest-` clones.** The test-cluster mode's clones on the secondary are replication sources only, never served. Decide whether the backend creates them as internal volumes (hidden from listings, exempt from the per-node subsystem cap, like the shipping path's landing volumes) or as ordinary volumes under a naming convention. | Backend team | diff --git a/operator/internal/controllers/driver/rbac.go b/operator/internal/controllers/driver/rbac.go index 1b482f92f..358d5eedf 100644 --- a/operator/internal/controllers/driver/rbac.go +++ b/operator/internal/controllers/driver/rbac.go @@ -134,6 +134,41 @@ func csiAddonsRoleBinding(d *simplyblockv1alpha2.SimplyblockDriver) *rbacv1.Role } } +// authDelegatorClusterRole is the well-known, built-in ClusterRole every +// component that validates bearer tokens via TokenReview binds to, rather +// than each defining its own copy of the same two-verb rule. +const authDelegatorClusterRole = "system:auth-delegator" + +// csiAddonsAuthDelegatorBinding grants the controller plugin's account +// tokenreviews.authentication.k8s.io:create, cluster-scoped since TokenReview +// has no namespaced form. Required, not optional: the csi-addons sidecar's +// gRPC server authenticates every incoming call from the controller-manager +// by reviewing its bearer token (internal/kubernetes/token/grpc.go, +// --enable-auth defaults to true) — confirmed against a live cluster, where +// omitting this left every connection attempt failing with "failed to +// review token ... is forbidden ... at the cluster scope". Binding to the +// built-in role rather than a hand-rolled ClusterRole needs no new marker on +// the operator's own ClusterRole: the operator already holds `bind` on +// every ClusterRole unconditionally (rbac.go's clusterroles;clusterrolebindings +// marker), which is what Kubernetes' escalation prevention checks for +// referencing an existing role instead of granting its permissions directly. +func csiAddonsAuthDelegatorBinding(d *simplyblockv1alpha2.SimplyblockDriver) *rbacv1.ClusterRoleBinding { + n := names(d) + return &rbacv1.ClusterRoleBinding{ + ObjectMeta: metav1.ObjectMeta{Name: n.clusterRoleBinding("csi-addons-auth-delegator")}, + Subjects: []rbacv1.Subject{{ + Kind: rbacv1.ServiceAccountKind, + Name: n.controllerServiceAccount, + Namespace: d.Namespace, + }}, + RoleRef: rbacv1.RoleRef{ + APIGroup: rbacv1.GroupName, + Kind: "ClusterRole", + Name: authDelegatorClusterRole, + }, + } +} + // serviceAccountFor names the account each role is bound to. The node plugin has // its own, and the controller plugin's sidecars share one. func serviceAccountFor(n objectNames, component string) string { diff --git a/operator/internal/controllers/driver/rbac_test.go b/operator/internal/controllers/driver/rbac_test.go index b5dd58f10..cf414cb8c 100644 --- a/operator/internal/controllers/driver/rbac_test.go +++ b/operator/internal/controllers/driver/rbac_test.go @@ -184,6 +184,30 @@ func TestCSIAddonsRoleIsNamespacedAndBoundToTheControllerAccount(t *testing.T) { } } +// The csi-addons sidecar's gRPC server authenticates every incoming call +// from the controller-manager via TokenReview (internal/kubernetes/token/grpc.go, +// --enable-auth defaults to true), which is cluster-scoped and so cannot be +// granted by the namespaced Role above -- confirmed against a live cluster, +// where omitting this left every connection failing with "failed to review +// token ... is forbidden ... at the cluster scope". +func TestCSIAddonsSidecarCanAuthenticateIncomingCalls(t *testing.T) { + d := testDriver("simplyblock") + n := names(d) + + binding := csiAddonsAuthDelegatorBinding(d) + if binding.Namespace != "" { + t.Errorf("namespace = %q, want \"\" (ClusterRoleBinding is cluster-scoped)", binding.Namespace) + } + if len(binding.Subjects) != 1 || binding.Subjects[0].Name != n.controllerServiceAccount || + binding.Subjects[0].Namespace != d.Namespace { + t.Errorf("binding subject = %+v, want the controller account in %q", + binding.Subjects, d.Namespace) + } + if binding.RoleRef.Kind != "ClusterRole" || binding.RoleRef.Name != "system:auth-delegator" { + t.Errorf("roleRef = %+v, want the built-in ClusterRole system:auth-delegator", binding.RoleRef) + } +} + // The rule set is exactly what the sidecar's own job needs: its CSIAddonsNode // and its leader-election Lease, both scoped to this namespace. func TestCSIAddonsRoleRulesAreScopedToItsOwnJob(t *testing.T) { diff --git a/operator/internal/controllers/driver/simplyblockdriver_controller.go b/operator/internal/controllers/driver/simplyblockdriver_controller.go index fe1f166fc..843f4064c 100644 --- a/operator/internal/controllers/driver/simplyblockdriver_controller.go +++ b/operator/internal/controllers/driver/simplyblockdriver_controller.go @@ -273,7 +273,7 @@ func (r *SimplyblockDriverReconciler) event( func (r *SimplyblockDriverReconciler) desired( d *simplyblockv1alpha2.SimplyblockDriver, image string, ) []client.Object { - objects := make([]client.Object, 0, 20) + objects := make([]client.Object, 0, 21) for _, sa := range serviceAccounts(d) { objects = append(objects, sa) @@ -287,7 +287,7 @@ func (r *SimplyblockDriverReconciler) desired( for _, crb := range clusterRoleBindings(d) { objects = append(objects, crb) } - objects = append(objects, csiAddonsRole(d), csiAddonsRoleBinding(d)) + objects = append(objects, csiAddonsRole(d), csiAddonsRoleBinding(d), csiAddonsAuthDelegatorBinding(d)) objects = append(objects, nodeDaemonSet(d, image), controllerStatefulSet(d, image), csiDriver(d)) if snapshotsEnabled(d) { objects = append(objects, volumeSnapshotClass(d)) diff --git a/operator/internal/controllers/driver/simplyblockdriver_controller_test.go b/operator/internal/controllers/driver/simplyblockdriver_controller_test.go index 88d443567..84ae9a1e1 100644 --- a/operator/internal/controllers/driver/simplyblockdriver_controller_test.go +++ b/operator/internal/controllers/driver/simplyblockdriver_controller_test.go @@ -75,7 +75,7 @@ func TestDesiredCoversTheWholeObjectSet(t *testing.T) { } want := map[string]int{ - "sa": 2, "cm": 2, "role": 5, "binding": 5, + "sa": 2, "cm": 2, "role": 5, "binding": 6, "namespacedRole": 1, "namespacedBinding": 1, "ds": 1, "sts": 1, "csidriver": 1, // the VolumeSnapshotClass, which is unstructured From 0057b4711ac450f10a574e09663709c1c9002505 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Fri, 18 Sep 2026 15:48:43 +0100 Subject: [PATCH 070/206] modified helm-charts/charts/simplyblock-operator/templates/rbac-csi-addons-controller.yaml --- .../templates/rbac-csi-addons-controller.yaml | 18 ++++++++++++++++++ 1 file changed, 18 insertions(+) diff --git a/helm-charts/charts/simplyblock-operator/templates/rbac-csi-addons-controller.yaml b/helm-charts/charts/simplyblock-operator/templates/rbac-csi-addons-controller.yaml index 7f333720c..1a4879189 100644 --- a/helm-charts/charts/simplyblock-operator/templates/rbac-csi-addons-controller.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/rbac-csi-addons-controller.yaml @@ -93,6 +93,24 @@ rules: - patch - update - watch + # rbac-justified: an addition beyond the stock upstream manifest, which + # omits this too (it's an opt-in feature, not something every deployment + # uses). The manager's sidecar-facing ReplicationServer proxy + # (internal/sidecar/service/volumereplication.go) hard-requires resolving + # a Secret named by VolumeReplicationClass's + # replication.storage.openshift.io/replication-secret-name(space) + # parameters before it will proxy ANY Replication RPC, even to a driver + # (like this one) that never reads the secret's own contents -- confirmed + # against a live cluster, where every EnableVolumeReplication attempt + # failed with "Failed to get secret ... resource name may not be empty" + # before ever reaching the sidecar. get only, no list/watch: the manager + # reads one named Secret per VolumeReplicationClass, never enumerates. + - apiGroups: + - "" + resources: + - secrets + verbs: + - get - apiGroups: - "" resources: From 3c44d9dddac97e210f9a46b0cf6a327871cd0561 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Fri, 18 Sep 2026 16:51:34 +0200 Subject: [PATCH 071/206] feat(discovery): the probe reads what the rules need, and the rules read it A draft is a list of disks somebody approves, and five of the decisions behind it were being made on readings the probe never took. Each of the five was a way for a run to propose the wrong thing quietly. The interface kind. Every software interface a worker has sits under devices/virtual, so a rule that refused virtual devices refused a bonded or tagged management network along with the cluster's own plumbing -- which is to say it refused the ordinary enterprise host, and every draft written for one named no management interface at all. inventory now reads the device type the kernel publishes and the lower_*/upper_* links around it, so a bond, a VLAN, a VXLAN and a veth are four answers rather than one. The management rule turns on that: it admits the kinds an address can be bound to, resolves an aggregate's speed and memory node through its members, and keeps a bridge or an overlay only where the cluster's own address is on it, since nothing readable separates a hypervisor's bridge from a CNI one except which address it carries. The device class, on both sides. The class rule checked the transport for an NVMe run and only the path for a block run, and an NVMe disk has a path like every other block device, so a block draft could name both classes. A cluster is built out of one, so the check is symmetric now. The subsystem NQN. A volume this product exported and a worker attached is a namespace like any other, and handing one back to a cluster as backend storage would give a volume's own bytes away as free space. blockdev reads subsysnqn and a rule parses it, so such a disk is refused as this fleet's own rather than as a device on a bus the run does not scan -- which is what the class rule said, and is both true and useless to somebody asking why a machine full of disks proposed none. The iSCSI session. A LUN reached the SCSI stack looking exactly like a local disk, so a cluster could be built on storage across a network without anybody deciding that. The transport is read from the session the SCSI transport class creates, and a LUN is taken only where the allow list names it. Whether a controller is held. InUse was a bool, so a controller nothing holds and a controller whose process table could not be read were one value, and only the first is safe to reclaim. It is a pointer now, with Held and Free as the two questions a caller has, and claimableControllers offers only what was checked. Two smaller things the same pass turned up. The explanation counted refusals by the sentence they rendered as rather than by the rule the refusal carries, so a hundred disks declined by one allow list read as a hundred findings; it counts by rule now, and the allow-and-deny reason names the list rather than the device. And a size was rendered to four significant digits, so a byte under a tebibyte printed as 1024G -- above the value, and a decimal its own parser refuses -- while a refusal re-rendered the bound instead of quoting what the filter said. Whole units now render exactly, nothing rounds up, and the bound is the administrator's own text. Report version goes to 6 for the interface kind and stack, the subsystem NQN, and the in-use pointer. Co-Authored-By: Claude Opus 5 (1M context) --- atlas-lib/blockdev/iscsi_test.go | 48 +++ atlas-lib/blockdev/scan.go | 40 +++ atlas-lib/blockdev/subsysnqn_test.go | 59 ++++ atlas-lib/inventory/netiface.go | 47 ++- atlas-lib/inventory/netiface_test.go | 1 + atlas-lib/inventory/netkind.go | 188 +++++++++++ atlas-lib/inventory/netkind_test.go | 295 ++++++++++++++++++ atlas-lib/pci/inuse_test.go | 42 +++ atlas-lib/pci/pci_test.go | 19 +- atlas-lib/pci/scan.go | 31 +- atlas-lib/pci/userspace.go | 15 +- .../deployment/operatorops_controller.go | 45 ++- .../deployment/operatorops_unit_test.go | 3 +- operator/internal/discovery/grouping.go | 20 +- operator/internal/discovery/iscsi.go | 69 ++++ operator/internal/discovery/iscsi_test.go | 97 ++++++ operator/internal/discovery/mgmtiface.go | 184 +++++++++-- operator/internal/discovery/netstack.go | 141 +++++++++ operator/internal/discovery/netstack_test.go | 262 ++++++++++++++++ operator/internal/discovery/plan.go | 152 ++++++--- operator/internal/discovery/plan_test.go | 102 +++++- operator/internal/discovery/rules.go | 118 +++++-- operator/internal/discovery/rules_test.go | 128 +++++++- operator/internal/discovery/simplyblock.go | 50 +++ .../internal/discovery/simplyblock_test.go | 120 +++++++ operator/internal/nodeprobe/collect.go | 30 +- operator/internal/nodeprobe/netstack_test.go | 118 +++++++ operator/internal/nodeprobe/report.go | 62 +++- 28 files changed, 2315 insertions(+), 171 deletions(-) create mode 100644 atlas-lib/blockdev/iscsi_test.go create mode 100644 atlas-lib/blockdev/subsysnqn_test.go create mode 100644 atlas-lib/inventory/netkind.go create mode 100644 atlas-lib/inventory/netkind_test.go create mode 100644 atlas-lib/pci/inuse_test.go create mode 100644 operator/internal/discovery/iscsi.go create mode 100644 operator/internal/discovery/iscsi_test.go create mode 100644 operator/internal/discovery/netstack.go create mode 100644 operator/internal/discovery/netstack_test.go create mode 100644 operator/internal/discovery/simplyblock.go create mode 100644 operator/internal/discovery/simplyblock_test.go create mode 100644 operator/internal/nodeprobe/netstack_test.go diff --git a/atlas-lib/blockdev/iscsi_test.go b/atlas-lib/blockdev/iscsi_test.go new file mode 100644 index 000000000..93420a04f --- /dev/null +++ b/atlas-lib/blockdev/iscsi_test.go @@ -0,0 +1,48 @@ +// That the scan names an iSCSI LUN for what it is. +// +// An iSCSI disk reaches the kernel through the SCSI stack, so every marker the +// transport reading looks for says SCSI: it hangs off a hostN, it is addressed +// as a SCSI target, and it presents an sdX. What separates it from a disk in +// the machine is one segment of its device path, the session the SCSI transport +// class creates for it, and without reading that a LUN on the other side of a +// network is indistinguishable from a disk on a cable inside the chassis. + +package blockdev + +import "testing" + +// iscsiHost is a worker with one local SAS disk and one iSCSI LUN, which is the +// pair the reading has to tell apart. +func iscsiHost() tree { + const ( + local = "devices/pci0000:00/0000:18:00.0/host0/end_device-0:0/target0:0:0/0:0:0:0/block/sda" + lun = "devices/platform/host6/session1/target6:0:0/6:0:0:0/block/sdb" + ) + t2 := tree{files: map[string]string{}, links: map[string]string{}} + + t2.blockAttrs(local, "8:0", 3125627568, true, false, false) + t2.links["class/block/sda"] = "../../" + local + + t2.blockAttrs(lun, "8:16", 2097152000, false, false, false) + t2.links["class/block/sdb"] = "../../" + lun + + t2.files["self/mountinfo"] = "25 1 0:24 / / rw - overlay overlay rw" + t2.files["swaps"] = "Filename\t\t\t\tType\t\tSize\t\tUsed\t\tPriority\n" + return t2 +} + +func TestScanNamesAnISCSILUNForWhatItIs(t *testing.T) { + root := iscsiHost().write(t) + + disks, err := Scan(ScanConfig{SysfsRoot: root}) + if err != nil { + t.Fatalf("scan the block devices: %v", err) + } + + if got := scanned(t, disks, "sdb").Transport; got != TransportISCSI { + t.Errorf("the LUN reads as %q, want %q", got, TransportISCSI) + } + if got := scanned(t, disks, "sda").Transport; got != TransportSAS { + t.Errorf("the local disk reads as %q, want %q", got, TransportSAS) + } +} diff --git a/atlas-lib/blockdev/scan.go b/atlas-lib/blockdev/scan.go index 9ba40b7f1..2cad9b38d 100644 --- a/atlas-lib/blockdev/scan.go +++ b/atlas-lib/blockdev/scan.go @@ -191,6 +191,17 @@ const ( // TransportSCSI is a SCSI disk whose bus the tree does not narrow further. TransportSCSI Transport = "SCSI" + // TransportISCSI is a LUN reached over iSCSI, which is a disk on the other + // side of a network presented through the SCSI stack. + // + // It is a separate transport from TransportSCSI, and the distinction is the + // whole reason it is read. Every marker an iSCSI LUN carries says SCSI: it + // hangs off a host, it is addressed as a target, and it presents an sdX. The + // one thing that says where its bytes are is the session the SCSI transport + // class creates for it, and a caller deciding whether to hand a disk to a + // storage cluster is deciding about bytes that are somewhere else. + TransportISCSI Transport = "iSCSI" + // TransportVirtio is a paravirtualized disk, which is what a virtual worker // has. TransportVirtio Transport = "Virtio" @@ -246,6 +257,16 @@ type Disk struct { // NUMANodeUnknown. NUMANode int + // SubsystemNQN is the NVMe Qualified Name of the subsystem the namespace + // belongs to, and is empty for a device on any other bus. + // + // It is what identifies a namespace rather than placing it. The transport + // says a fabric namespace came from somewhere else, which is an inference + // from where its controllers are; the NQN says what it is, and a caller + // that needs to recognize its own product's volumes among a machine's + // disks has nothing else to read. + SubsystemNQN string + // Partitions is the kernel names of the partitions on this device, // ascending. A disk with any is a disk something has already divided up, // whether or not those partitions carry anything. @@ -331,6 +352,7 @@ func scanOne(cfg ScanConfig, dir, name string) (Disk, error) { disk.Virtual = sysfs.IsVirtual(resolved) disk.PCIAddress = sysfs.PCIAddressOf(resolved) disk.NUMANode = numaNodeOf(resolved) + disk.SubsystemNQN = subsystemNQNOf(resolved) if transport, ok := nvmeSubsystemTransport(resolved); ok { disk.Transport = transport } else { @@ -425,6 +447,12 @@ func transportOf(resolved string) Transport { case strings.HasPrefix(segment, "end_device-"), strings.HasPrefix(segment, "sas_"), strings.HasPrefix(segment, "expander-"): return TransportSAS + case numbered(segment, "session"): + // The SCSI transport class names an iSCSI session sessionN and + // nothing else in the tree is named that way. The host segment + // above it has already set the SCSI fallback, and this overrides + // it, because where a LUN's bytes are is the more specific answer. + return TransportISCSI case numbered(segment, "host"): scsi = true } @@ -504,6 +532,18 @@ func nvmeSubsystemTransport(resolved string) (Transport, bool) { return TransportNVMeFabric, true } +// subsystemNQNOf reads the NQN of the NVMe subsystem a namespace belongs to, +// and returns the empty string for a device that is on no NVMe subsystem. +// +// One read covers both shapes the kernel presents. A namespace reached through +// a multipath head sits under its nvme-subsysN directory, and one reached +// through a single controller sits under that controller: in both, the +// subsysnqn attribute is in the parent, which is why this reads the parent +// rather than deciding which shape it is looking at first. +func subsystemNQNOf(resolved string) string { + return sysfs.String(filepath.Dir(filepath.Clean(resolved)), "subsysnqn") +} + // controllerName matches the controller entries of a subsystem directory, // nvme0 and nvme12, and not the namespaces (nvme0n1), the per-controller legs // (nvme0c0n1), or the generic character devices (ng0n1) that sit beside them. diff --git a/atlas-lib/blockdev/subsysnqn_test.go b/atlas-lib/blockdev/subsysnqn_test.go new file mode 100644 index 000000000..beb64374c --- /dev/null +++ b/atlas-lib/blockdev/subsysnqn_test.go @@ -0,0 +1,59 @@ +// That the scan reports the subsystem NQN of an NVMe device. +// +// The NQN is what identifies a namespace as a volume this product exported, +// which the transport can only infer: a fabric namespace is somebody else's +// bytes, and a fabric namespace whose NQN names a simplyblock logical volume is +// this fleet's own. The two answers need different words, and only the NQN +// distinguishes them. + +package blockdev + +import "testing" + +func TestScanReportsTheSubsystemNQNOfAMultipathNamespace(t *testing.T) { + host := multipathHost() + const nqn = "nqn.2023-02.io.simplyblock:c30a691a-1d2e-4f3a-9b8c-5d6e7f809a1b:lvol:792e184c-0a1b-2c3d-4e5f-60718293a4b5" + host.files["devices/virtual/nvme-subsystem/nvme-subsys0/subsysnqn"] = nqn + + root := host.write(t) + disks, err := Scan(ScanConfig{SysfsRoot: root}) + if err != nil { + t.Fatalf("scan the block devices: %v", err) + } + + if got := scanned(t, disks, "nvme3n1").SubsystemNQN; got != nqn { + t.Errorf("the fabric namespace reports NQN %q, want the subsystem's", got) + } +} + +func TestScanReportsTheSubsystemNQNOfAControllerNamespace(t *testing.T) { + // A namespace the kernel presents through a controller rather than a + // subsystem keeps its NQN one directory up, in the same place. + host := storageHost() + const nqn = "nqn.2019-08.org.qemu:local" + host.files["devices/pci0000:00/0000:5e:00.0/nvme/nvme0/subsysnqn"] = nqn + + root := host.write(t) + disks, err := Scan(ScanConfig{SysfsRoot: root}) + if err != nil { + t.Fatalf("scan the block devices: %v", err) + } + + if got := scanned(t, disks, "nvme0n1").SubsystemNQN; got != nqn { + t.Errorf("the local namespace reports NQN %q, want the controller's", got) + } +} + +func TestScanReportsNoNQNForADeviceThatIsNotNVMe(t *testing.T) { + root := storageHost().write(t) + disks, err := Scan(ScanConfig{SysfsRoot: root}) + if err != nil { + t.Fatalf("scan the block devices: %v", err) + } + + for _, name := range []string{"sda", "sdb", "vda"} { + if got := scanned(t, disks, name).SubsystemNQN; got != "" { + t.Errorf("%s reports NQN %q, and is on no NVMe subsystem", name, got) + } + } +} diff --git a/atlas-lib/inventory/netiface.go b/atlas-lib/inventory/netiface.go index 6e35cf412..70060bb24 100644 --- a/atlas-lib/inventory/netiface.go +++ b/atlas-lib/inventory/netiface.go @@ -92,6 +92,28 @@ type Interface struct { // socket reaches its NIC across the interconnect. NUMANode int + // Kind is what sort of device the interface is, from the device type its + // driver registered. It is what separates a bond or a tagged VLAN, which a + // management address sits on in most fleets, from a veth or a CNI bridge, + // which are the cluster's own plumbing: all of them are virtual, and only + // the kind tells them apart. + Kind LinkKind + + // Lower is what this interface is built on, ascending by name: the members + // of a bond or a bridge, or the single parent of a VLAN or a macvlan. It is + // empty for an interface built on nothing. + // + // It is the only route from an aggregate to the hardware under it. A bond + // carries no slot, no driver, and no memory node of its own, so a caller + // that has to know where a bonded management network physically lands reads + // the members and looks them up in the same reading. + Lower []string + + // Upper is what is built on this interface, ascending by name. It is the + // direction that matters for a NIC holding no address of its own: on a host + // whose management network is tagged, the address is on a VLAN above it. + Upper []string + // Bridge reports whether the interface is a software bridge. // // It is separate from Virtual, which a bridge also is, because the two @@ -207,20 +229,27 @@ func readInterface(dir, name string) Interface { iface.OperState = LinkUnknown } + // What the interface is stacked on is read before anything else, because it + // is the one reading that answers for a device the rest of this function + // returns early on: a bond has no slot and no driver, and its members are + // where both of those are. + iface.Lower, iface.Upper = stackAt(dir) + // The class entry is a symlink into the device tree, and where that tree // says the interface sits is what decides whether it is backed by // hardware. A device under devices/virtual has none. resolved, err := filepath.EvalSymlinks(dir) - if err != nil { - return iface - } - iface.Virtual = sysfs.IsVirtual(resolved) - // A bridge exports a bridge/ directory whatever it is named, which is what - // makes this a reading rather than a guess at br0 and cni0 and docker0. - if entries, err := os.Stat(filepath.Join(dir, "bridge")); err == nil && entries.IsDir() { - iface.Bridge = true + if err == nil { + iface.Virtual = sysfs.IsVirtual(resolved) + // A bridge exports a bridge/ directory whatever it is named, which is + // what makes this a reading rather than a guess at br0 and cni0 and + // docker0. + if entries, err := os.Stat(filepath.Join(dir, "bridge")); err == nil && entries.IsDir() { + iface.Bridge = true + } } - if iface.Virtual { + iface.Kind = kindOf(dir, iface.Virtual, iface.Loopback, iface.Bridge) + if err != nil || iface.Virtual { return iface } diff --git a/atlas-lib/inventory/netiface_test.go b/atlas-lib/inventory/netiface_test.go index 40f5779b8..dbe9cf465 100644 --- a/atlas-lib/inventory/netiface_test.go +++ b/atlas-lib/inventory/netiface_test.go @@ -105,6 +105,7 @@ func TestReadInterfacesReportsAPhysicalNICWhole(t *testing.T) { Driver: "mlx5_core", PCIAddress: "0000:3b:00.0", NUMANode: 0, + Kind: LinkPhysical, } if got := byName(t, ifaces, "eth0"); !reflect.DeepEqual(got, want) { t.Errorf("read %+v, want %+v", got, want) diff --git a/atlas-lib/inventory/netkind.go b/atlas-lib/inventory/netkind.go new file mode 100644 index 000000000..764cd9298 --- /dev/null +++ b/atlas-lib/inventory/netkind.go @@ -0,0 +1,188 @@ +// What kind of device a network interface is, and what it is stacked on. +// +// Every software interface a host has sits under devices/virtual: a bridge, a +// bond, a VLAN, a VXLAN, a veth to a pod. The physical-or-not reading in +// netiface.go therefore collapses all of them into one answer. That answer is not enough +// to decide anything: a management address on a bond or on a tagged VLAN is an +// ordinary configuration, and the same address on a CNI bridge or a veth is the +// cluster's own plumbing. The two cases need opposite decisions and read +// identically without the kind. +// +// The kind comes from the device type the kernel publishes in the interface's +// uevent, which is the driver's own word for what it registered: `bond`, +// `bridge`, `vlan`, `vxlan`, `macvlan`, `ipvlan`, `team`. A driver that +// registers no device type gets no DEVTYPE line, and such an interface is +// reported as virtual with no kind rather than guessed at from its name. A host +// is free to call its veth pairs and its dummy interfaces anything at all. +// +// The stack comes from the lower_* and upper_* links every stacked interface +// exports, one per relation and in both directions. One reading answers both +// questions that matter: the members of a bond or a bridge, which is the only +// route from an aggregate to the slot, the memory node, and the link speed of +// the hardware under it, and what is stacked above a NIC, which is where the +// address lives on a host whose management network is tagged. + +package inventory + +import ( + "os" + "path/filepath" + "slices" + "strings" + + "github.com/simplyblock/atlas/internal/sysfs" +) + +// LinkKind is what sort of device an interface is, in the kernel's own +// spelling where the kernel has one. +type LinkKind string + +const ( + // LinkPhysical is an interface backed by hardware. + LinkPhysical LinkKind = "physical" + + // LinkLoopback is the loopback interface, decided by its ARPHRD type. + LinkLoopback LinkKind = "loopback" + + // LinkBridge is a software bridge. It has members and is the one aggregate + // whose address says more about who configured it than about the host: a + // cluster's CNI bridge and a hypervisor host's guest bridge are the same + // kind of device put to opposite purposes. + LinkBridge LinkKind = "bridge" + + // LinkBond is a link aggregation. It has members, and it rather than any of + // them is what an address is bound to. + LinkBond LinkKind = "bond" + + // LinkTeam is the other link aggregation, teamd's. + LinkTeam LinkKind = "team" + + // LinkVLAN is a tagged interface over a single parent, which is how most + // fleets separate a management network from the data one. + LinkVLAN LinkKind = "vlan" + + // LinkVXLAN is an overlay. It is bindable, and on a Kubernetes worker it is + // usually the cluster's own pod network rather than anything a storage node + // should serve over. + LinkVXLAN LinkKind = "vxlan" + + // LinkMACVLAN and LinkIPVLAN are the two ways of giving one NIC several + // identities. + LinkMACVLAN LinkKind = "macvlan" + LinkIPVLAN LinkKind = "ipvlan" + + // LinkVirtual is an interface with no hardware behind it whose driver + // registered no device type: a veth, a dummy, a tunnel. It is the honest + // answer for a device the kernel does not name, and it is never bindable, + // because nothing that reaches this value has been identified. + LinkVirtual LinkKind = "virtual" +) + +// Bindable reports whether an address held by an interface of this kind is one +// a service could bind and something outside the host could reach it on. +// +// Loopback is out because it reaches nothing, and LinkVirtual is out because it +// is the value for a device nothing identified: a veth to a pod and a dummy +// interface both land there, and admitting an unidentified device would be +// admitting those. A bridge is in, which is a change from refusing every +// bridge: a hypervisor host's management address lives on one, and whether a +// given bridge is that or a CNI bridge is answered by the addresses it holds +// rather than by its kind. +func (k LinkKind) Bindable() bool { + switch k { + case LinkPhysical, LinkBond, LinkTeam, LinkVLAN, LinkVXLAN, LinkMACVLAN, LinkIPVLAN, LinkBridge: + return true + default: + return false + } +} + +// Aggregate reports whether an interface of this kind is built out of members, +// as opposed to derived from a single parent. +// +// The distinction is what a caller needs it for: an aggregate's members are +// several and interchangeable, so the hardware facts under it have to be +// combined, where a derived interface has exactly one parent and inherits its. +func (k LinkKind) Aggregate() bool { + return k == LinkBridge || k == LinkBond || k == LinkTeam +} + +// deviceTypeKey is the line the kernel writes the driver's device type on, in +// the generic uevent every device directory carries. +const deviceTypeKey = "DEVTYPE=" + +// kindOf is the kind of the interface at dir, given what the physical-or-not, +// loopback, and bridge readings already concluded. +// +// The device type is asked first and the directories a driver exports second, +// because the type is the driver's own word and the directories are evidence of +// it: a kernel that publishes no DEVTYPE still has bonding/ under every bond and +// bridge/ under every bridge, and a tree captured from one is the case the +// fallback exists for. +func kindOf(dir string, virtual, loopback, bridge bool) LinkKind { + if loopback { + return LinkLoopback + } + + switch declared := LinkKind(deviceTypeAt(dir)); declared { + case LinkBridge, LinkBond, LinkTeam, LinkVLAN, LinkVXLAN, LinkMACVLAN, LinkIPVLAN: + return declared + } + if bridge { + return LinkBridge + } + if isDir(filepath.Join(dir, "bonding")) { + return LinkBond + } + + // No device type and no driver directory, so nothing identified it. What + // the tree says about hardware is then the whole answer. + if virtual { + return LinkVirtual + } + return LinkPhysical +} + +// isDir reports whether the path is a directory, which is how a driver's own +// export is recognized. +func isDir(path string) bool { + info, err := os.Stat(path) + return err == nil && info.IsDir() +} + +// deviceTypeAt reads the DEVTYPE the kernel publishes for the device at dir, and +// the empty string for a device whose driver registered no type. +func deviceTypeAt(dir string) string { + for _, line := range strings.Split(sysfs.String(dir, "uevent"), "\n") { + if value, found := strings.CutPrefix(strings.TrimSpace(line), deviceTypeKey); found { + return value + } + } + return "" +} + +// stackAt is what the interface at dir is built on and what is built on it, +// each ascending by name so that two readings of one host are comparable. +// +// The links are read by name rather than followed, because the name is what +// identifies the other interface in the same reading and following the link +// would answer a question nobody asked. +func stackAt(dir string) (lower, upper []string) { + entries, err := os.ReadDir(dir) + if err != nil { + return nil, nil + } + + for _, entry := range entries { + if name, found := strings.CutPrefix(entry.Name(), "lower_"); found { + lower = append(lower, name) + continue + } + if name, found := strings.CutPrefix(entry.Name(), "upper_"); found { + upper = append(upper, name) + } + } + slices.Sort(lower) + slices.Sort(upper) + return lower, upper +} diff --git a/atlas-lib/inventory/netkind_test.go b/atlas-lib/inventory/netkind_test.go new file mode 100644 index 000000000..e5ec6af8f --- /dev/null +++ b/atlas-lib/inventory/netkind_test.go @@ -0,0 +1,295 @@ +// What the interface reader reports for a stacked host: a bond over two NICs, a +// VLAN over the bond, a bridge over a third NIC, an overlay, a veth, and +// loopback. +// +// The tree is the one this reading exists for. Every one of those devices sits +// under devices/virtual, so the physical-or-not reading alone collapses them +// into one answer, and which of them a management address can be bound to +// differs for each. + +package inventory + +import ( + "reflect" + "testing" +) + +// stackedNetHost is a bonded host with a tagged management network, a bridge for +// guests, and a cluster's own plumbing beside it. +// +// bond0 carries no address of its own: the address is on the VLAN above it, +// which is how a tagged management network is configured and the case the +// physical-or-not reading cannot describe. +func stackedNetHost() fixture { + const ( + eth0 = "devices/pci0000:00/0000:3b:00.0/net/eth0/" + eth1 = "devices/pci0000:00/0000:3b:00.1/net/eth1/" + eth2 = "devices/pci0000:00/0000:af:00.0/net/eth2/" + bond = "devices/virtual/net/bond0/" + vlan = "devices/virtual/net/bond0.100/" + br = "devices/virtual/net/br0/" + vx = "devices/virtual/net/vxlan.calico/" + veth = "devices/virtual/net/veth7a1c/" + lo = "devices/virtual/net/lo/" + ) + f := fixture{files: map[string]string{ + eth0 + "operstate": "up", + eth0 + "type": "1", + eth0 + "speed": "25000", + eth0 + "uevent": "INTERFACE=eth0\nIFINDEX=2", + + eth1 + "operstate": "up", + eth1 + "type": "1", + eth1 + "speed": "25000", + eth1 + "uevent": "INTERFACE=eth1\nIFINDEX=3", + + eth2 + "operstate": "up", + eth2 + "type": "1", + eth2 + "speed": "10000", + eth2 + "uevent": "INTERFACE=eth2\nIFINDEX=4", + + bond + "operstate": "up", + bond + "type": "1", + bond + "speed": "50000", + bond + "uevent": "INTERFACE=bond0\nIFINDEX=5\nDEVTYPE=bond", + bond + "bonding/slaves": "eth0 eth1", + + vlan + "operstate": "up", + vlan + "type": "1", + vlan + "uevent": "INTERFACE=bond0.100\nIFINDEX=6\nDEVTYPE=vlan", + + br + "operstate": "up", + br + "type": "1", + br + "uevent": "INTERFACE=br0\nIFINDEX=7\nDEVTYPE=bridge", + br + "bridge/root_id": "8000.0c42a15bc312", + + vx + "operstate": "unknown", + vx + "type": "1", + vx + "uevent": "INTERFACE=vxlan.calico\nIFINDEX=8\nDEVTYPE=vxlan", + + // A veth names no DEVTYPE, which is the kernel's answer for a device + // whose driver registers no device type. It stays unnamed rather than + // guessed at. + veth + "operstate": "up", + veth + "type": "1", + veth + "uevent": "INTERFACE=veth7a1c\nIFINDEX=9", + + lo + "operstate": "unknown", + lo + "type": "772", + lo + "uevent": "INTERFACE=lo\nIFINDEX=1", + + "devices/pci0000:00/0000:3b:00.0/numa_node": "0", + "devices/pci0000:00/0000:3b:00.1/numa_node": "0", + "devices/pci0000:00/0000:af:00.0/numa_node": "1", + }} + f.links = map[string]string{ + "class/net/eth0": "../../devices/pci0000:00/0000:3b:00.0/net/eth0", + "class/net/eth1": "../../devices/pci0000:00/0000:3b:00.1/net/eth1", + "class/net/eth2": "../../devices/pci0000:00/0000:af:00.0/net/eth2", + "class/net/bond0": "../../devices/virtual/net/bond0", + "class/net/bond0.100": "../../devices/virtual/net/bond0.100", + "class/net/br0": "../../devices/virtual/net/br0", + "class/net/vxlan.calico": "../../devices/virtual/net/vxlan.calico", + "class/net/veth7a1c": "../../devices/virtual/net/veth7a1c", + "class/net/lo": "../../devices/virtual/net/lo", + + eth0 + "device": "../..", + eth1 + "device": "../..", + eth2 + "device": "../..", + + // The stack, as the kernel exports it: one link per relation, in both + // directions. + bond + "lower_eth0": "../../../pci0000:00/0000:3b:00.0/net/eth0", + bond + "lower_eth1": "../../../pci0000:00/0000:3b:00.1/net/eth1", + bond + "upper_bond0.100": "../bond0.100", + eth0 + "upper_bond0": "../../../../virtual/net/bond0", + eth1 + "upper_bond0": "../../../../virtual/net/bond0", + vlan + "lower_bond0": "../bond0", + br + "lower_eth2": "../../../pci0000:00/0000:af:00.0/net/eth2", + eth2 + "upper_br0": "../../../../virtual/net/br0", + } + return f +} + +func TestReadInterfacesNamesTheKindOfEachDevice(t *testing.T) { + root := stackedNetHost().write(t) + + ifaces, err := ReadInterfaces(Config{SysfsRoot: root, ProcRoot: root}) + if err != nil { + t.Fatalf("read the interfaces: %v", err) + } + + want := map[string]LinkKind{ + "eth0": LinkPhysical, + "eth1": LinkPhysical, + "eth2": LinkPhysical, + "bond0": LinkBond, + "bond0.100": LinkVLAN, + "br0": LinkBridge, + "vxlan.calico": LinkVXLAN, + "veth7a1c": LinkVirtual, + "lo": LinkLoopback, + } + for name, kind := range want { + if got := byName(t, ifaces, name).Kind; got != kind { + t.Errorf("%s reads as kind %q, want %q", name, got, kind) + } + } +} + +func TestReadInterfacesReportsTheMembersOfABondAndABridge(t *testing.T) { + // A bond and a bridge carry no slot, no driver, and no memory node of their + // own, so the members are the only route to the hardware underneath one. + root := stackedNetHost().write(t) + + ifaces, err := ReadInterfaces(Config{SysfsRoot: root, ProcRoot: root}) + if err != nil { + t.Fatalf("read the interfaces: %v", err) + } + + if got := byName(t, ifaces, "bond0").Lower; !reflect.DeepEqual(got, []string{"eth0", "eth1"}) { + t.Errorf("bond0 reports members %v, want both NICs ascending", got) + } + if got := byName(t, ifaces, "br0").Lower; !reflect.DeepEqual(got, []string{"eth2"}) { + t.Errorf("br0 reports members %v, want eth2", got) + } + if got := byName(t, ifaces, "eth0").Lower; len(got) != 0 { + t.Errorf("a physical NIC reports members %v, and is built on nothing", got) + } +} + +func TestReadInterfacesReportsTheParentOfADerivedDevice(t *testing.T) { + root := stackedNetHost().write(t) + + ifaces, err := ReadInterfaces(Config{SysfsRoot: root, ProcRoot: root}) + if err != nil { + t.Fatalf("read the interfaces: %v", err) + } + + if got := byName(t, ifaces, "bond0.100").Lower; !reflect.DeepEqual(got, []string{"bond0"}) { + t.Errorf("the VLAN reports %v as what it is built on, want bond0", got) + } +} + +func TestReadInterfacesReportsWhatIsStackedOnAnInterface(t *testing.T) { + // The direction that matters for a NIC holding no address: what above it + // might hold one. + root := stackedNetHost().write(t) + + ifaces, err := ReadInterfaces(Config{SysfsRoot: root, ProcRoot: root}) + if err != nil { + t.Fatalf("read the interfaces: %v", err) + } + + for _, name := range []string{"eth0", "eth1"} { + if got := byName(t, ifaces, name).Upper; !reflect.DeepEqual(got, []string{"bond0"}) { + t.Errorf("%s reports %v stacked on it, want bond0", name, got) + } + } + if got := byName(t, ifaces, "bond0").Upper; !reflect.DeepEqual(got, []string{"bond0.100"}) { + t.Errorf("bond0 reports %v stacked on it, want the VLAN", got) + } + if got := byName(t, ifaces, "eth2").Upper; !reflect.DeepEqual(got, []string{"br0"}) { + t.Errorf("eth2 reports %v stacked on it, want br0", got) + } +} + +func TestReadInterfacesStillMarksEveryStackedDeviceVirtual(t *testing.T) { + // The kind is an addition and not a replacement: a caller reading Virtual + // keeps the answer it had. + root := stackedNetHost().write(t) + + ifaces, err := ReadInterfaces(Config{SysfsRoot: root, ProcRoot: root}) + if err != nil { + t.Fatalf("read the interfaces: %v", err) + } + + for _, name := range []string{"bond0", "bond0.100", "br0", "vxlan.calico", "veth7a1c", "lo"} { + if !byName(t, ifaces, name).Virtual { + t.Errorf("%s is under devices/virtual and was not marked virtual", name) + } + } + if !byName(t, ifaces, "br0").Bridge { + t.Error("br0 exports a bridge directory and was not marked a bridge") + } + for _, name := range []string{"bond0", "bond0.100"} { + if byName(t, ifaces, name).Bridge { + t.Errorf("%s was marked a bridge", name) + } + } +} + +func TestLinkKindBindableSeparatesWhatAnAddressCanBeBoundTo(t *testing.T) { + // The question the management-interface rule asks of a kind, kept beside the + // kinds so that a kind added later has to answer it. + bindable := map[LinkKind]bool{ + LinkPhysical: true, + LinkBond: true, + LinkVLAN: true, + LinkVXLAN: true, + LinkMACVLAN: true, + LinkIPVLAN: true, + LinkTeam: true, + LinkBridge: true, + LinkLoopback: false, + LinkVirtual: false, + } + for kind, want := range bindable { + if got := kind.Bindable(); got != want { + t.Errorf("%s.Bindable() is %v, want %v", kind, got, want) + } + } +} + +func TestLinkKindAggregatesAreTheOnesWithMembers(t *testing.T) { + for _, kind := range []LinkKind{LinkBond, LinkBridge, LinkTeam} { + if !kind.Aggregate() { + t.Errorf("%s holds members and does not report itself as an aggregate", kind) + } + } + for _, kind := range []LinkKind{LinkPhysical, LinkVLAN, LinkVXLAN, LinkLoopback, LinkVirtual} { + if kind.Aggregate() { + t.Errorf("%s reports itself as an aggregate", kind) + } + } +} + +// kernelWithoutDeviceTypes is a bridge and a bond on a kernel that publishes no +// DEVTYPE for either, which is what a tree captured from an older kernel looks +// like. The directories the two drivers export are the fallback. +func kernelWithoutDeviceTypes() fixture { + const ( + br = "devices/virtual/net/br-mgmt/" + bond = "devices/virtual/net/bond1/" + ) + f := fixture{files: map[string]string{ + br + "operstate": "up", + br + "type": "1", + br + "bridge/stp_state": "0", + + bond + "operstate": "up", + bond + "type": "1", + bond + "bonding/slaves": "", + }} + f.links = map[string]string{ + "class/net/br-mgmt": "../../devices/virtual/net/br-mgmt", + "class/net/bond1": "../../devices/virtual/net/bond1", + } + return f +} + +func TestReadInterfacesFallsBackToWhatTheDriverExports(t *testing.T) { + root := kernelWithoutDeviceTypes().write(t) + + ifaces, err := ReadInterfaces(Config{SysfsRoot: root, ProcRoot: root}) + if err != nil { + t.Fatalf("read the interfaces: %v", err) + } + + if got := byName(t, ifaces, "br-mgmt").Kind; got != LinkBridge { + t.Errorf("a bridge with no DEVTYPE reads as %q, want %q", got, LinkBridge) + } + if got := byName(t, ifaces, "bond1").Kind; got != LinkBond { + t.Errorf("a bond with no DEVTYPE reads as %q, want %q", got, LinkBond) + } +} diff --git a/atlas-lib/pci/inuse_test.go b/atlas-lib/pci/inuse_test.go new file mode 100644 index 000000000..0e15b64d1 --- /dev/null +++ b/atlas-lib/pci/inuse_test.go @@ -0,0 +1,42 @@ +// That a controller nobody checked is not reported as one nobody is using. +// +// The two answers used to be one value. A device nothing holds and a device +// whose process table could not be read both left InUse false, and only the +// first is safe to reclaim, so a caller reading the field alone could not tell +// a free disk from a disk it never asked about. + +package pci + +import "testing" + +func TestAnUncheckedControllerReportsNoAnswer(t *testing.T) { + // A device the check never looked at, which is every device before + // CheckHolders runs. + device := Device{Address: "0000:5e:00.0", Driver: DriverVFIO} + + if device.InUse != nil { + t.Errorf("a device nothing checked reports %v, want no answer", *device.InUse) + } + if device.Held() { + t.Error("a device nothing checked reports itself held") + } + if device.Free() { + t.Error("a device nothing checked reports itself free, which is the whole defect") + } +} + +func TestACheckedControllerReportsWhatWasFound(t *testing.T) { + held := Device{Address: "0000:5e:00.0", Driver: DriverVFIO} + held.InUse = new(bool) + *held.InUse = true + + free := Device{Address: "0000:5f:00.0", Driver: DriverVFIO} + free.InUse = new(bool) + + if !held.Held() || held.Free() { + t.Error("a held device does not report itself held") + } + if free.Held() || !free.Free() { + t.Error("a device nothing holds does not report itself free") + } +} diff --git a/atlas-lib/pci/pci_test.go b/atlas-lib/pci/pci_test.go index 406c02573..78d2452a6 100644 --- a/atlas-lib/pci/pci_test.go +++ b/atlas-lib/pci/pci_test.go @@ -353,19 +353,19 @@ func TestCheckHoldersMarksOnlyTheDeviceSomethingHolds(t *testing.T) { // The holder is a hypervisor rather than this product, which changes // nothing: the question is whether anything is driving the disk. held := nvmeSlot(t, checked, "0000:00:02.0") - if !held.InUse { + if !held.Held() { t.Error("a controller a process holds open is not marked in use") } idle := nvmeSlot(t, checked, "0000:00:03.0") - if idle.InUse { - t.Error("a controller nothing holds is marked in use") + if !idle.Free() { + t.Error("a controller nothing holds is not marked free") } } -// The failure that matters: a process table that cannot be read leaves InUse -// false, and the only thing standing between that and a reclaimed disk is the -// error. So it is returned rather than swallowed, and the devices come back -// regardless, because a machine whose procfs is unreadable still has +// The failure that matters: a process table that cannot be read leaves the +// answer unset rather than false, so a device nothing established as free is +// not one a caller can take. The error is returned as well, and the devices +// come back regardless, because a machine whose procfs is unreadable still has // controllers worth reporting. func TestCheckHoldersReportsWhatItCouldNotDetermine(t *testing.T) { h := takenWorker() @@ -387,8 +387,11 @@ func TestCheckHoldersReportsWhatItCouldNotDetermine(t *testing.T) { t.Fatal("the controllers were dropped along with the answer") } for _, device := range checked { - if device.InUse { + if device.Held() { t.Errorf("%s was marked in use by a check that never ran", device.Address) } + if device.Free() { + t.Errorf("%s was marked free by a check that never ran", device.Address) + } } } diff --git a/atlas-lib/pci/scan.go b/atlas-lib/pci/scan.go index 2f74c6810..4c68b03f0 100644 --- a/atlas-lib/pci/scan.go +++ b/atlas-lib/pci/scan.go @@ -119,17 +119,32 @@ type Device struct { // was the one in slot 02.0. UIODevices []string - // InUse reports whether anything holds one of UIODevices open. + // InUse reports whether anything holds one of UIODevices open, and is nil + // for a device nothing asked about. // - // [Scan] does not set it, because answering needs the process table and a - // scan reads sysfs. [CheckHolders] fills it in, and a false on a device - // neither of them looked at is the zero value rather than an answer: a - // caller that needs to tell the two apart has to know which produced the - // device, which is why the check records its failures rather than leaving - // this field to carry them. - InUse bool + // The three states are the point. [Scan] does not set it, because answering + // needs the process table and a scan reads sysfs; [CheckHolders] sets it on + // every device it could check and leaves it nil on every device it could + // not. A device nothing holds and a device whose process table could not be + // read are then different values rather than one, which is what keeps a + // caller from reclaiming a disk it never established was free. + // + // Read it through [Device.Held] and [Device.Free], which are the two + // questions a caller actually has and neither of which is the negation of + // the other. + InUse *bool } +// Held reports whether something is known to hold the device open. +// +// A device nothing checked is not held, because nothing established that it +// was. It is not free either: see [Device.Free]. +func (d Device) Held() bool { return d.InUse != nil && *d.InUse } + +// Free reports whether the device was checked and found to be held by nothing, +// which is the only state in which it is safe to take. +func (d Device) Free() bool { return d.InUse != nil && !*d.InUse } + // IsNVMe reports whether the device is an NVMe controller. func (d Device) IsNVMe() bool { return strings.HasPrefix(d.Class, classNVMePrefix) } diff --git a/atlas-lib/pci/userspace.go b/atlas-lib/pci/userspace.go index f8bbaf07c..f51f49854 100644 --- a/atlas-lib/pci/userspace.go +++ b/atlas-lib/pci/userspace.go @@ -71,12 +71,12 @@ func HeldBy(cfg Config, device Device) ([]Holder, error) { // CheckHolders fills in InUse for every device given, and returns what it could // not determine alongside the devices it could. // -// The failures are returned rather than folded into InUse because the two are -// not the same answer. A device nothing holds and a device that could not be -// checked both leave InUse false, and only one of them is safe to reclaim, so a -// caller that drops the error has quietly turned unknown into free. The -// devices come back either way: a machine whose process table could not be read -// still has controllers worth reporting. +// The failures are returned as well as left on the device, because the two +// answers are not the same. A device nothing holds and a device that could not +// be checked are a set InUse and a nil one, and only the first is safe to +// reclaim: a caller reading the field gets the distinction whether or not it +// reads the error. The devices come back either way, since a machine whose +// process table could not be read still has controllers worth reporting. func CheckHolders(cfg Config, devices []Device) ([]Device, error) { out := make([]Device, 0, len(devices)) var errs []error @@ -97,7 +97,8 @@ func CheckHolders(cfg Config, devices []Device) ([]Device, error) { out = append(out, device) continue } - device.InUse = len(holders) > 0 + held := len(holders) > 0 + device.InUse = &held out = append(out, device) } return out, errors.Join(errs...) diff --git a/operator/internal/controllers/deployment/operatorops_controller.go b/operator/internal/controllers/deployment/operatorops_controller.go index 4d0b03f11..d94e69d63 100644 --- a/operator/internal/controllers/deployment/operatorops_controller.go +++ b/operator/internal/controllers/deployment/operatorops_controller.go @@ -330,6 +330,11 @@ func (r *OperatorOpsReconciler) inspect( spec = &simplyblockv1alpha2.DiscoverSpec{} } + // Refused before a single worker is probed. See refuseUnreadableFilter. + if err := refuseUnreadableFilter(spec); err != nil { + return false, err + } + var nodes corev1.NodeList options := []client.ListOption{} if len(spec.NodeSelector) > 0 { @@ -511,6 +516,12 @@ func (r *OperatorOpsReconciler) write( spec = &simplyblockv1alpha2.DiscoverSpec{} } + // Asked again at the step that applies the filter, so that the guard sits + // where the value is used and not only where the run was settled. + if err := refuseUnreadableFilter(spec); err != nil { + return false, err + } + reports, err := r.reportsFor(ctx, ops) if err != nil { return false, err @@ -536,10 +547,7 @@ func (r *OperatorOpsReconciler) write( } filter := spec.DeviceFilter - planner := discoverypkg.Planner{ - Class: discoverypkg.ClassOf(filter), - KubeNodes: kubeNodes, - } + planner := discoverypkg.Planner{KubeNodes: kubeNodes} plan := planner.Plan(collected, filter) if len(plan.NodeSets) == 0 { @@ -589,6 +597,35 @@ func (r *OperatorOpsReconciler) write( return true, r.status(ctx, ops) } +// refuseUnreadableFilter refuses a run whose device filter names a size range +// nothing can read. +// +// An unreadable range builds no size rule, which widens the filter to every +// disk on every worker rather than narrowing it to none, and the refusal has no +// device to attach itself to so it reaches neither the refusal list nor the +// run's explanation. The draft that results is indistinguishable from one a run +// meant to write. +// +// The admission webhook refuses such a run at the request, so reaching this +// means the webhook is not installed or was bypassed. Both steps that read the +// filter ask, because the one that settles the run should not probe a fleet for +// a draft that cannot be right, and the one that applies it should not depend on +// the other having asked. +func refuseUnreadableFilter(spec *simplyblockv1alpha2.DiscoverSpec) error { + filter := spec.DeviceFilter + if filter == nil || filter.DriveSizeRange == "" { + return nil + } + if _, _, err := discoverypkg.ParseSizeRange(filter.DriveSizeRange); err != nil { + return refusef(OperationFailed, + "spec.discover.deviceFilter.driveSizeRange is %q, which cannot be read: %v. "+ + "A run whose range cannot be read applies no size filter at all, so the draft "+ + "would name every disk on every worker rather than the ones asked for", + filter.DriveSizeRange, err) + } + return nil +} + // draftFor builds the document, and the notes explaining the numbers in it that // were not read off the hardware. func (r *OperatorOpsReconciler) draftFor( diff --git a/operator/internal/controllers/deployment/operatorops_unit_test.go b/operator/internal/controllers/deployment/operatorops_unit_test.go index 9d84a4f47..b3456f34e 100644 --- a/operator/internal/controllers/deployment/operatorops_unit_test.go +++ b/operator/internal/controllers/deployment/operatorops_unit_test.go @@ -33,6 +33,7 @@ import ( "sigs.k8s.io/controller-runtime/pkg/client/interceptor" "github.com/simplyblock/atlas/blockdev" + "github.com/simplyblock/atlas/ptr" simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" @@ -640,7 +641,7 @@ func heldReportConfigMap(t *testing.T, node string, addresses ...string) *corev1 NUMANode: -1, // Held rather than idle, which is what makes the machine unusable: // an idle binding is a disk the draft now claims. - InUse: true, + InUse: ptr.To(true), }) } diff --git a/operator/internal/discovery/grouping.go b/operator/internal/discovery/grouping.go index e0a3b76cb..17db9f49c 100644 --- a/operator/internal/discovery/grouping.go +++ b/operator/internal/discovery/grouping.go @@ -55,12 +55,16 @@ type Worker struct { // was given no node objects. Kube KubeNode - // MgmtInterface is the interface the draft names for management, or empty - // when the machine presents none that would serve. It is part of what makes - // two workers groupable: a NodeGroup names one interface for every worker in - // it, so machines that call theirs different things describe different - // groups however identical their disks are. - MgmtInterface string + // Mgmt is the interface the draft names for management, with what the stack + // says about it: the physical interfaces underneath a bond or a bridge, the + // speed they add up to, and the memory node they sit on. Its name is empty + // when the machine presents no interface that would serve. + // + // The name is part of what makes two workers groupable: a NodeGroup names + // one interface for every worker in it, so machines that call theirs + // different things describe different groups however identical their disks + // are. + Mgmt Management } // Addresses is how the draft names this worker's devices, ascending and without @@ -132,14 +136,14 @@ func (GroupByHardware) Group(workers []Worker) []Group { for _, worker := range workers { addresses := worker.Addresses() - signature := worker.Class.signature(addresses, worker.MgmtInterface) + signature := worker.Class.signature(addresses, worker.Mgmt.Name) group, seen := bySignature[signature] if !seen { group = &Group{ Class: worker.Class, Addresses: addresses, - MgmtInterface: worker.MgmtInterface, + MgmtInterface: worker.Mgmt.Name, } bySignature[signature] = group order = append(order, signature) diff --git a/operator/internal/discovery/iscsi.go b/operator/internal/discovery/iscsi.go new file mode 100644 index 000000000..511678d1a --- /dev/null +++ b/operator/internal/discovery/iscsi.go @@ -0,0 +1,69 @@ +// The rule that makes an iSCSI LUN an instruction rather than a default. +// +// Every other bus a run scans is a cable inside the chassis. iSCSI is not: a +// LUN is a disk on the other side of a network, and a storage cluster built on +// one runs every write of its data path over that network, on top of whatever +// the target is doing with the bytes at the far end. +// +// Whether that is wanted is a question about the deployment and not about the +// hardware, and it is exactly the kind of question the approval gate exists for +// — except that a draft proposing a LUN is a draft a reviewer has to notice and +// strike, and a fifty-worker document is not one anybody reads closely enough. +// So the default is the other way around: a LUN reaches a draft only where +// somebody named it, and naming it is the decision. + +package discovery + +import ( + "fmt" + "strings" + + "github.com/simplyblock/atlas/blockdev" + "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" +) + +// ISCSIRule admits an iSCSI LUN only when the run's allow list names it. +type ISCSIRule struct { + // Class is how a device of this run is named, which is what the allow list + // is written in. + Class DeviceClass + + // Allow is the run's allow list for that class. Empty names nothing, so + // every LUN is refused. + Allow []string +} + +func (ISCSIRule) Name() string { return "iSCSI" } + +// Admit refuses an iSCSI device the allow list does not name. +// +// It reads the allow list itself rather than leaving it to the allow-and-deny +// rule, because that rule is only built when a list was given: with no filter +// at all there is nothing to refuse a LUN, and with a list given for another +// reason a LUN would be admitted by being named alongside everything else. What +// this rule states is that a LUN is never taken by default, and that holds +// whether or not a run carries a filter. +func (r ISCSIRule) Admit(_ nodeprobe.Report, device nodeprobe.Device) (bool, string) { + if device.Transport != string(blockdev.TransportISCSI) { + return true, "" + } + + address := r.Class.Address(device) + for _, allowed := range r.Allow { + if strings.EqualFold(address, allowed) { + return true, "" + } + } + return false, fmt.Sprintf( + "it is an iSCSI LUN, which is storage on the other side of a network, so a run "+ + "takes one only where the allow list names it and this one names %s", + describeAllowList(r.Allow)) +} + +// describeAllowList renders what the run was told to take, for the refusal. +func describeAllowList(allow []string) string { + if len(allow) == 0 { + return "nothing" + } + return strings.Join(allow, ", ") +} diff --git a/operator/internal/discovery/iscsi_test.go b/operator/internal/discovery/iscsi_test.go new file mode 100644 index 000000000..09c9860ae --- /dev/null +++ b/operator/internal/discovery/iscsi_test.go @@ -0,0 +1,97 @@ +// That an iSCSI LUN reaches a draft only where somebody named it. +// +// A LUN is a disk on the other side of a network, and a fleet that wants one in +// a storage cluster has decided something a discovery run cannot decide for it: +// that the network between the worker and the target is one a data path should +// run over. So the run does not propose one, and taking it is an instruction +// rather than a default. + +package discovery + +import ( + "strings" + "testing" + + "github.com/simplyblock/atlas/blockdev" + "github.com/simplyblock/atlas/ptr" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" +) + +// attachedLUN is an iSCSI disk as the probe reports one. +func attachedLUN(name string, size uint64) nodeprobe.Device { + device := blockDisk(name, 0, size) + device.Transport = string(blockdev.TransportISCSI) + return device +} + +func TestAnISCSILUNIsNotProposedUnlessItIsNamed(t *testing.T) { + device := attachedLUN("sdb", 2*tb) + + ok, why := admit(ISCSIRule{Class: ClassBlock}, device) + if ok { + t.Fatal("a run proposed an iSCSI LUN nobody asked for") + } + if !strings.Contains(why, "iSCSI") || !strings.Contains(why, "allow list") { + t.Errorf("the reason %q does not say it is iSCSI and has to be named", why) + } +} + +func TestAnNamedISCSILUNIsAdmitted(t *testing.T) { + device := attachedLUN("sdb", 2*tb) + rule := ISCSIRule{Class: ClassBlock, Allow: []string{"/dev/sdb"}} + + if ok, why := admit(rule, device); !ok { + t.Errorf("a named iSCSI LUN was refused: %s", why) + } + + // Naming a different disk does not name this one. + other := ISCSIRule{Class: ClassBlock, Allow: []string{"/dev/sdc"}} + if ok, _ := admit(other, device); ok { + t.Error("an allow list naming another disk admitted this LUN") + } +} + +func TestTheRuleLeavesEveryOtherBusAlone(t *testing.T) { + for _, transport := range []blockdev.Transport{ + blockdev.TransportSATA, blockdev.TransportSAS, + blockdev.TransportSCSI, blockdev.TransportVirtio, + } { + device := blockDisk("sda", 0, tb) + device.Transport = string(transport) + if ok, why := admit(ISCSIRule{Class: ClassBlock}, device); !ok { + t.Errorf("a %s disk was refused by the iSCSI rule: %s", transport, why) + } + } +} + +func TestAnISCSILUNIsRefusedByAWholeRunUnlessNamed(t *testing.T) { + report := report("worker-1") + report.Devices = []nodeprobe.Device{attachedLUN("sdb", 2*tb), blockDisk("vdb", 0, 2*tb)} + + block := &simplyblockv1alpha2.DeviceFilter{EnableLogicalBlockDevices: ptr.To(true)} + plan := Planner{}.Plan([]nodeprobe.Report{report}, + &simplyblockv1alpha2.DeviceFilter{EnableLogicalBlockDevices: block.EnableLogicalBlockDevices}) + + if len(plan.NodeSets) != 1 { + t.Fatalf("built %+v", plan.NodeSets) + } + group := plan.NodeSets[0].Groups[0] + if len(group.Devices.Block) != 1 || group.Devices.Block[0] != "/dev/vdb" { + t.Errorf("the group names %v, want the virtio disk alone", group.Devices.Block) + } + + // Naming it is what takes it, and the allow list then bounds the draft to + // what it names, which is the allow list's own rule. + named := &simplyblockv1alpha2.DeviceFilter{ + EnableLogicalBlockDevices: ptr.To(true), + BlockAllowList: []string{"/dev/sdb", "/dev/vdb"}, + } + plan = Planner{}.Plan([]nodeprobe.Report{report}, named) + if len(plan.NodeSets) != 1 { + t.Fatalf("built %+v", plan.NodeSets) + } + if got := plan.NodeSets[0].Groups[0].Devices.Block; len(got) != 2 { + t.Errorf("the group names %v, want both disks once the LUN is named", got) + } +} diff --git a/operator/internal/discovery/mgmtiface.go b/operator/internal/discovery/mgmtiface.go index e35224571..f26b555ce 100644 --- a/operator/internal/discovery/mgmtiface.go +++ b/operator/internal/discovery/mgmtiface.go @@ -6,81 +6,205 @@ // complete. So the draft names one, and names it from what the probe read rather // than leaving it to be defaulted somewhere further down. // -// The rule is a ranking rather than a match, because a fleet's machines do not +// The rule is a ladder rather than a match, because a fleet's machines do not // agree on what their NICs are called and a draft that named eth0 everywhere // would be wrong on the machines that call it ens5f0. What every candidate has -// in common is the shape of the answer: hardware, up, and holding an address -// something can reach it on. +// in common is the shape of the answer: a device something can be bound to, up, +// and holding an address something can reach it on. +// +// The kind of device is the rung that carries the most weight, and it is the one +// this rule could not read until the probe reported it. Every software interface +// a worker has sits under devices/virtual: a bond, a tagged VLAN, a CNI bridge, +// a veth to a pod. A rule that refused virtual devices therefore refused a +// bonded or tagged management network along with the cluster's own plumbing, +// which is to say it refused the ordinary enterprise host. The kinds are told apart now, and +// the two that need a judgment rather than a rule get one: a bridge and an +// overlay are named only where the cluster's own address is on them, because the +// same bridge is a hypervisor host's management network or a pod network +// depending on nothing readable but that. package discovery import ( "cmp" + "fmt" "net" "slices" + "strings" + "github.com/simplyblock/atlas/inventory" "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" ) +// Management is the interface a draft names for management, with the evidence +// behind the choice. +type Management struct { + // Name is the interface, or empty when the machine presents none that would + // serve. + Name string + + // Kind is what sort of device it is, in inventory.LinkKind's spelling. + Kind string + + // Members are the physical interfaces underneath it, ascending by name: the + // slaves of a bond, the ports of a bridge, or the hardware a VLAN's parent + // resolves to. It is empty for an interface that is itself physical. + Members []string + + // SpeedMbps is what the interface can carry, resolved through the stack + // where the interface reports nothing of its own. + SpeedMbps int + + // NUMANode is the memory node the hardware underneath it sits on, or + // inventory.NUMANodeUnknown when there is none or the members disagree. + NUMANode int + + // Reason says what was chosen and why, for the record a reviewer reads. + Reason string +} + // ManagementInterface is the interface a draft names for this worker, or empty // when the machine presents none that would serve. +func ManagementInterface(report nodeprobe.Report, nodeAddress string) string { + return ManagementOf(report, nodeAddress).Name +} + +// ManagementOf applies the ladder and returns what it chose. // // nodeAddress is the address the cluster already reaches the machine on, and // when it is known the interface holding it wins outright. That is not a // preference among equals: the operator addresses the worker by that address // everywhere else it talks to it, so any other choice would have the two halves -// of one deployment describing different networks. +// of one deployment describing different networks. It is also what settles the +// two kinds no rule can settle, since a bridge or an overlay carrying the +// cluster's own address is the management network by definition. // -// Returning empty is a real answer rather than a failure. A machine whose only -// addressed interfaces are its cluster's own bridges has no management interface, -// and a draft that named one anyway would produce a storage node the rest of the -// fleet cannot reach. -func ManagementInterface(report nodeprobe.Report, nodeAddress string) string { - candidates := make([]nodeprobe.Interface, 0, len(report.Interfaces)) +// An empty name is a real answer rather than a failure. A machine whose only +// addressed interfaces are its cluster's pod network has no management +// interface, and a draft that named one anyway would produce a storage node the +// rest of the fleet cannot reach. +func ManagementOf(report nodeprobe.Report, nodeAddress string) Management { + index := interfacesByName(report) + + var candidates []nodeprobe.Interface for _, iface := range report.Interfaces { - if !servesManagement(iface) { + holdsNodeAddress := nodeAddress != "" && slices.Contains(iface.Addresses, nodeAddress) + if !servesManagement(iface, holdsNodeAddress) { continue } - if nodeAddress != "" && slices.Contains(iface.Addresses, nodeAddress) { - return iface.Name + if holdsNodeAddress { + return resolve(iface, index, fmt.Sprintf( + "it holds %s, the address the cluster reaches the machine on", nodeAddress)) } candidates = append(candidates, iface) } if len(candidates) == 0 { - return "" + return Management{NUMANode: inventory.NUMANodeUnknown} } - // Fastest first, then by name. The name is what makes the choice stable: the - // probe's reading order is the kernel's, and a draft that changed between two - // runs of the same fleet is one a reviewer cannot diff. + // Fastest first, then the simpler kind, then by name. The name is what makes + // the choice stable: the probe's reading order is the kernel's, and a draft + // that changed between two runs of the same fleet is one a reviewer cannot + // diff. + speeds := make(map[string]int, len(candidates)) + for _, iface := range candidates { + speeds[iface.Name] = effectiveSpeed(iface.Name, index, map[string]bool{}) + } slices.SortFunc(candidates, func(a, b nodeprobe.Interface) int { - if a.SpeedMbps != b.SpeedMbps { - return cmp.Compare(b.SpeedMbps, a.SpeedMbps) - } - return cmp.Compare(a.Name, b.Name) + return cmp.Or( + cmp.Compare(speeds[b.Name], speeds[a.Name]), + cmp.Compare(simplicity(interfaceKind(a)), simplicity(interfaceKind(b))), + cmp.Compare(a.Name, b.Name), + ) }) - return candidates[0].Name + + best := candidates[0] + why := fmt.Sprintf("it is the fastest interface holding a reachable address, at %d Mbps", + speeds[best.Name]) + if len(candidates) == 1 { + why = "it is the only interface holding a reachable address" + } + return resolve(best, index, why) +} + +// resolve fills in what the stack says about the interface that was chosen. +func resolve(iface nodeprobe.Interface, index map[string]nodeprobe.Interface, why string) Management { + out := Management{ + Name: iface.Name, + Kind: string(interfaceKind(iface)), + Members: physicalUnder(iface.Lower, index, map[string]bool{}), + SpeedMbps: effectiveSpeed(iface.Name, index, map[string]bool{}), + } + + out.NUMANode = iface.NUMANode + if len(out.Members) > 0 { + out.NUMANode = numaNodeOf(out.Members, index) + } + + out.Reason = fmt.Sprintf("%s is %s and %s", out.Name, describeKind(out.Kind, out.Members), why) + return out +} + +// describeKind names what sort of device was chosen, and what is under it where +// that is the only place the hardware appears. +func describeKind(kind string, members []string) string { + if inventory.LinkKind(kind) == inventory.LinkPhysical { + return "a physical interface" + } + if len(members) == 0 { + return "a " + kind + " interface" + } + return fmt.Sprintf("a %s interface over %s", kind, strings.Join(members, " and ")) } // servesManagement reports whether an interface could carry a storage node's // management traffic at all. // -// Each condition rules out a machine this product has actually been deployed -// onto. The virtual ones are the cluster's own plumbing — a CNI bridge, a veth -// to a pod, a flannel overlay, loopback — and every worker has several holding -// addresses that reach nothing outside the node. A link that is down keeps the -// address it was configured with and carries nothing. An interface with no -// address is what the control plane refuses by name. -func servesManagement(iface nodeprobe.Interface) bool { - if iface.Virtual || iface.Bridge || iface.Loopback { +// Each condition rules out a machine this product has been deployed onto. An +// unidentified virtual device is a veth to a pod or a dummy, and there are +// dozens of the first on every worker. A link that is down keeps the address it +// was configured with and carries nothing. An interface with no address is what +// the control plane refuses by name. +// +// The bridge and overlay condition is the one that is a judgment. Both kinds can +// be bound and both are ordinarily somebody else's network: a CNI bridge and a +// flannel overlay hold the pod network, and a hypervisor host's bridge holds the +// address the cluster reaches the machine on. Nothing readable separates the two +// except which address is on them, so that is what separates them here. +func servesManagement(iface nodeprobe.Interface, holdsNodeAddress bool) bool { + kind := interfaceKind(iface) + if !kind.Bindable() { return false } if iface.State != "" && iface.State != "up" && iface.State != "unknown" { return false } + if (kind == inventory.LinkBridge || kind == inventory.LinkVXLAN) && !holdsNodeAddress { + return false + } return slices.ContainsFunc(iface.Addresses, reachable) } +// simplicity orders the kinds for a tie, lowest first: the fewest layers +// between the address and the wire. +// +// It breaks a tie and never more than that, because the layering is not what +// decides whether an interface works. Two interfaces carrying the same traffic +// over the same hardware are equally usable, and the untagged one is the +// simpler answer for a reviewer to read. +func simplicity(kind inventory.LinkKind) int { + switch kind { + case inventory.LinkPhysical: + return 0 + case inventory.LinkBond, inventory.LinkTeam: + return 1 + case inventory.LinkVLAN, inventory.LinkMACVLAN, inventory.LinkIPVLAN: + return 2 + default: + return 3 + } +} + // reachable reports whether an address is one something could contact the // machine on. // diff --git a/operator/internal/discovery/netstack.go b/operator/internal/discovery/netstack.go new file mode 100644 index 000000000..865c66c1a --- /dev/null +++ b/operator/internal/discovery/netstack.go @@ -0,0 +1,141 @@ +// Resolving an interface through the stack it sits in. +// +// An aggregate reports nothing about the hardware under it. A bond has no slot, +// no driver, no memory node, and usually no link speed of its own, because the +// kernel puts it under devices/virtual along with everything else that has no +// hardware. What it can carry, and where in the machine it lands, are facts +// about its members, and a VLAN over that bond inherits both from the bond. +// +// So the two questions the management rule asks of a candidate, how fast it is +// and which memory node it is on, are answered by walking down the stack the +// report carries, rather than by reading the interface and getting zero. + +package discovery + +import ( + "slices" + + "github.com/simplyblock/atlas/inventory" + "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" +) + +// interfaceKind is what sort of device an interface is, taking the probe's own +// answer where there is one. +// +// The fallback is for an interface whose kind the probe left empty, which is a +// fixture or a probe that could not read the device type rather than an older +// schema: a report from before the kind existed is refused by version. What was +// readable before the kind is then the whole answer, and it is read the way it +// was: virtual and unidentified is not something to bind to. +func interfaceKind(iface nodeprobe.Interface) inventory.LinkKind { + if kind := inventory.LinkKind(iface.Kind); kind != "" { + return kind + } + switch { + case iface.Loopback: + return inventory.LinkLoopback + case iface.Bridge: + return inventory.LinkBridge + case iface.Virtual: + return inventory.LinkVirtual + default: + return inventory.LinkPhysical + } +} + +// interfacesByName indexes a report's interfaces so that a stack walk can +// resolve a name to the interface it belongs to. +func interfacesByName(report nodeprobe.Report) map[string]nodeprobe.Interface { + out := make(map[string]nodeprobe.Interface, len(report.Interfaces)) + for _, iface := range report.Interfaces { + out[iface.Name] = iface + } + return out +} + +// effectiveSpeed is what an interface can carry, in megabits per second. +// +// An interface that reports a speed is taken at its word. One that does not is +// resolved through what it is built on: an aggregate carries the sum of its +// members, because that is what aggregation is for, and a derived interface +// carries what its single parent carries, because it is the same wire with a +// tag on it. +// +// The walk is guarded against revisiting an interface. Nothing sysfs exports +// can produce a cycle, and a resolution that would hang on one is a resolution +// that trusts its input. +func effectiveSpeed(name string, index map[string]nodeprobe.Interface, seen map[string]bool) int { + iface, known := index[name] + if !known || seen[name] { + return 0 + } + seen[name] = true + + if iface.SpeedMbps > 0 { + return iface.SpeedMbps + } + if interfaceKind(iface).Aggregate() { + total := 0 + for _, member := range iface.Lower { + total += effectiveSpeed(member, index, seen) + } + return total + } + for _, parent := range iface.Lower { + if speed := effectiveSpeed(parent, index, seen); speed > 0 { + return speed + } + } + return 0 +} + +// physicalUnder is the physical interfaces reachable downward from the names +// given, ascending and without repeats. +// +// It is the only route from an aggregate to the hardware under it, and it is +// what lets a draft say which slots a bonded management network lands in. +func physicalUnder(names []string, index map[string]nodeprobe.Interface, seen map[string]bool) []string { + var out []string + for _, name := range names { + iface, known := index[name] + if !known || seen[name] { + continue + } + seen[name] = true + + if interfaceKind(iface) == inventory.LinkPhysical { + out = append(out, iface.Name) + continue + } + out = append(out, physicalUnder(iface.Lower, index, seen)...) + } + slices.Sort(out) + return slices.Compact(out) +} + +// numaNodeOf is the memory node a set of interfaces agrees on, and +// NUMANodeUnknown when they do not. +// +// Disagreement is a real answer and not a gap: a bond whose members are in two +// sockets has no affinity, and reporting one of the two would claim an affinity +// the interface does not have. +func numaNodeOf(names []string, index map[string]nodeprobe.Interface) int { + node := inventory.NUMANodeUnknown + for _, name := range names { + iface, known := index[name] + if !known { + continue + } + if iface.NUMANode == inventory.NUMANodeUnknown { + return inventory.NUMANodeUnknown + } + if node == inventory.NUMANodeUnknown { + node = iface.NUMANode + continue + } + if node != iface.NUMANode { + return inventory.NUMANodeUnknown + } + } + return node +} diff --git a/operator/internal/discovery/netstack_test.go b/operator/internal/discovery/netstack_test.go new file mode 100644 index 000000000..cdfdca5f1 --- /dev/null +++ b/operator/internal/discovery/netstack_test.go @@ -0,0 +1,262 @@ +// The management-interface rule over a stacked host, and the stack resolution it +// rests on. +// +// Every case here is a host this product has been deployed onto and the earlier +// rule named no interface on: a bonded management network, a tagged one, a +// hypervisor host whose address is on a bridge. All three are virtual devices, +// which is why the kind reading had to exist before the rule could tell them +// from the veth pairs beside them. + +package discovery + +import ( + "slices" + "testing" + + "github.com/simplyblock/atlas/inventory" + "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" +) + +// stacked is the interface list of a bonded host with a tagged management +// network: two NICs in a bond, a VLAN over the bond, a bridge for guests over a +// third NIC, the cluster's own overlay, and a veth to a pod. +// +// Which interface holds the management address is the caller's to set, because +// that is the only thing that differs between the cases below. +func stacked(addressed map[string][]string) []nodeprobe.Interface { + kinds := []struct { + name string + kind inventory.LinkKind + speed int + lower []string + upper []string + }{ + {name: "bond0", kind: inventory.LinkBond, lower: []string{"eth0", "eth1"}, upper: []string{"bond0.100"}}, + {name: "bond0.100", kind: inventory.LinkVLAN, lower: []string{"bond0"}}, + {name: "br0", kind: inventory.LinkBridge, lower: []string{"eth2"}}, + {name: "eth0", kind: inventory.LinkPhysical, speed: 25000, upper: []string{"bond0"}}, + {name: "eth1", kind: inventory.LinkPhysical, speed: 25000, upper: []string{"bond0"}}, + {name: "eth2", kind: inventory.LinkPhysical, speed: 10000, upper: []string{"br0"}}, + {name: "flannel.1", kind: inventory.LinkVXLAN}, + {name: "lo", kind: inventory.LinkLoopback}, + {name: "veth7a1c", kind: inventory.LinkVirtual}, + } + + out := make([]nodeprobe.Interface, 0, len(kinds)) + for _, entry := range kinds { + out = append(out, nodeprobe.Interface{ + Name: entry.name, + Kind: string(entry.kind), + SpeedMbps: entry.speed, + State: "up", + Lower: entry.lower, + Upper: entry.upper, + Virtual: entry.kind != inventory.LinkPhysical, + Bridge: entry.kind == inventory.LinkBridge, + Loopback: entry.kind == inventory.LinkLoopback, + NUMANode: 0, + Addresses: addressed[entry.name], + }) + } + return out +} + +// stackedReport is that host as a report, with the management address where the +// case wants it. +func stackedReport(addressed map[string][]string) nodeprobe.Report { + out := report("worker-1") + out.Interfaces = stacked(addressed) + return out +} + +func TestABondHoldingTheNodeAddressIsNamed(t *testing.T) { + // The bond is what an address is bound to on a bonded host. Its members hold + // none, so the earlier rule found no candidate at all and named nothing. + r := stackedReport(map[string][]string{"bond0": {"10.10.10.113"}}) + + if got := ManagementInterface(r, "10.10.10.113"); got != "bond0" { + t.Errorf("named %q, want the bond holding the node's address", got) + } +} + +func TestATaggedVLANHoldingTheNodeAddressIsNamed(t *testing.T) { + r := stackedReport(map[string][]string{"bond0.100": {"10.10.10.113"}}) + + if got := ManagementInterface(r, "10.10.10.113"); got != "bond0.100" { + t.Errorf("named %q, want the VLAN holding the node's address", got) + } +} + +func TestAHostBridgeHoldingTheNodeAddressIsNamed(t *testing.T) { + // A hypervisor host keeps its own address on the bridge its guests are on, + // and that address is the one the cluster reaches the machine by. + r := stackedReport(map[string][]string{"br0": {"192.168.1.10"}}) + + if got := ManagementInterface(r, "192.168.1.10"); got != "br0" { + t.Errorf("named %q, want the bridge holding the node's address", got) + } +} + +func TestABridgeHoldingSomebodyElsesAddressIsNotNamed(t *testing.T) { + // The same kind of device put to the opposite purpose: a CNI bridge holds + // the pod network and nothing outside the node reaches the machine on it. + r := stackedReport(map[string][]string{"br0": {"10.42.2.1"}}) + + if got := ManagementInterface(r, ""); got != "" { + t.Errorf("named %q, want nothing: the bridge carries somebody else's network", got) + } +} + +func TestAnOverlayIsNamedOnlyWhenTheClusterItselfUsesIt(t *testing.T) { + pod := stackedReport(map[string][]string{"flannel.1": {"10.42.2.0"}}) + if got := ManagementInterface(pod, ""); got != "" { + t.Errorf("named %q, want nothing: the overlay is the pod network", got) + } + + // A cluster that genuinely addresses its nodes over an overlay says so by + // the node address being on it, and then it is the management interface. + own := stackedReport(map[string][]string{"flannel.1": {"10.42.2.0"}}) + if got := ManagementInterface(own, "10.42.2.0"); got != "flannel.1" { + t.Errorf("named %q, want the overlay the node address is on", got) + } +} + +func TestAnUnidentifiedVirtualDeviceIsNeverNamed(t *testing.T) { + // A veth reports no kind of its own, and admitting a device nothing + // identified would admit every pod link on the machine. + r := stackedReport(map[string][]string{"veth7a1c": {"10.42.2.7"}}) + + if got := ManagementInterface(r, "10.42.2.7"); got != "" { + t.Errorf("named %q, want nothing: nothing identified the device", got) + } + if got := ManagementInterface(stackedReport(map[string][]string{"lo": {"127.0.0.1"}}), ""); got != "" { + t.Errorf("named %q, want nothing for loopback", got) + } +} + +func TestTheEffectiveSpeedOfAnAggregateIsItsMembers(t *testing.T) { + // A bond reports no speed of its own, so ranking it on what it reports puts + // it behind every physical NIC. What it can carry is what its members can. + r := stackedReport(map[string][]string{ + "bond0": {"10.10.10.113"}, + "eth2": {"192.168.1.10"}, + }) + + if got := ManagementInterface(r, ""); got != "bond0" { + t.Errorf("named %q, want the bond: 2x25G carries more than one 10G NIC", got) + } +} + +func TestADerivedInterfaceInheritsTheSpeedOfWhatItIsBuiltOn(t *testing.T) { + r := stackedReport(map[string][]string{ + "bond0.100": {"10.10.10.113"}, + "eth2": {"192.168.1.10"}, + }) + + if got := ManagementInterface(r, ""); got != "bond0.100" { + t.Errorf("named %q, want the VLAN over the bond", got) + } +} + +func TestAPhysicalInterfaceWinsATieWithADerivedOne(t *testing.T) { + // Both carry the same traffic over the same hardware, and the untagged one + // is the simpler answer for a reviewer to read. + r := report("worker-1") + r.Interfaces = []nodeprobe.Interface{ + { + Name: "eth0", Kind: string(inventory.LinkPhysical), SpeedMbps: 25000, + State: "up", Addresses: []string{"192.168.1.10"}, Upper: []string{"eth0.100"}, + }, + { + Name: "eth0.100", Kind: string(inventory.LinkVLAN), SpeedMbps: 25000, Virtual: true, + State: "up", Addresses: []string{"10.10.10.113"}, Lower: []string{"eth0"}, + }, + } + + if got := ManagementInterface(r, ""); got != "eth0" { + t.Errorf("named %q, want the physical interface", got) + } +} + +func TestTheHardwareUnderTheChosenInterfaceIsReported(t *testing.T) { + // What a bond amounts to is not readable from the bond: it carries no slot, + // no driver, and no memory node. The members are the only route to all three. + r := stackedReport(map[string][]string{"bond0": {"10.10.10.113"}}) + + mgmt := ManagementOf(r, "10.10.10.113") + if mgmt.Name != "bond0" || mgmt.Kind != string(inventory.LinkBond) { + t.Fatalf("chose %+v", mgmt) + } + if !slices.Equal(mgmt.Members, []string{"eth0", "eth1"}) { + t.Errorf("reported members %v, want both NICs", mgmt.Members) + } + if mgmt.SpeedMbps != 50000 { + t.Errorf("reported %d Mbps, want the sum of the members", mgmt.SpeedMbps) + } + if mgmt.NUMANode != 0 { + t.Errorf("reported memory node %d, want 0", mgmt.NUMANode) + } + if mgmt.Reason == "" { + t.Error("the choice carries no reason") + } +} + +func TestTheHardwareUnderADerivedInterfaceResolvesThroughItsParent(t *testing.T) { + r := stackedReport(map[string][]string{"bond0.100": {"10.10.10.113"}}) + + mgmt := ManagementOf(r, "10.10.10.113") + if !slices.Equal(mgmt.Members, []string{"eth0", "eth1"}) { + t.Errorf("the VLAN reports members %v, want the bond's NICs", mgmt.Members) + } +} + +func TestMembersOnDifferentMemoryNodesReportNone(t *testing.T) { + // A bond across two sockets has no memory node, and reporting one of them + // would claim an affinity the interface does not have. + r := stackedReport(map[string][]string{"bond0": {"10.10.10.113"}}) + for i := range r.Interfaces { + if r.Interfaces[i].Name == "eth1" { + r.Interfaces[i].NUMANode = 1 + } + } + + if got := ManagementOf(r, "10.10.10.113").NUMANode; got != inventory.NUMANodeUnknown { + t.Errorf("reported memory node %d, want %d", got, inventory.NUMANodeUnknown) + } +} + +func TestAnInterfaceThatNamesNoKindFallsBackToWhatElseWasReported(t *testing.T) { + // A report written before the kind existed is refused outright, but a + // fixture or a probe that left the field empty should still be read the way + // it was before: physical unless something said otherwise. + r := report("worker-1") + r.Interfaces = []nodeprobe.Interface{ + {Name: "eth0", State: "up", Addresses: []string{"192.168.1.10"}, SpeedMbps: 10000}, + {Name: "cni0", State: "up", Addresses: []string{"10.42.2.1"}, Virtual: true, Bridge: true}, + {Name: "flannel.1", State: "up", Addresses: []string{"10.42.2.0"}, Virtual: true}, + } + + if got := ManagementInterface(r, ""); got != "eth0" { + t.Errorf("named %q, want the one interface nothing marked virtual", got) + } +} + +func TestAStackThatPointsAtItselfTerminates(t *testing.T) { + // Nothing in sysfs produces a cycle, and a resolution that would hang on one + // is a resolution that trusts its input. + r := report("worker-1") + r.Interfaces = []nodeprobe.Interface{ + { + Name: "bond0", Kind: string(inventory.LinkBond), State: "up", Virtual: true, + Addresses: []string{"10.10.10.113"}, Lower: []string{"bond1"}, + }, + { + Name: "bond1", Kind: string(inventory.LinkBond), State: "up", Virtual: true, + Lower: []string{"bond0"}, + }, + } + + if got := ManagementOf(r, "10.10.10.113").Name; got != "bond0" { + t.Errorf("chose %q", got) + } +} diff --git a/operator/internal/discovery/plan.go b/operator/internal/discovery/plan.go index 5d89d7c70..5198748a9 100644 --- a/operator/internal/discovery/plan.go +++ b/operator/internal/discovery/plan.go @@ -26,12 +26,14 @@ import ( // this package documents, so a caller replaces the one decision it wants to // change and leaves the rest. type Planner struct { - // Class is which of the two classes of backend storage the run scans. - // Empty is ClassNVMe. - Class DeviceClass - // DeviceRules are applied to every reported device, in order. Nil is the // default set, which is what BasicDeviceRules returns. + // + // The class a run scans is not a field here. It is the filter's to state, + // through EnableLogicalBlockDevices, and a second statement of it could + // only ever agree or be wrong: a planner scanning NVMe with a filter + // carrying block lists read those lists on a branch that never ran and + // dropped them without a word. DeviceRules []DeviceRule // WorkerRules are applied to every worker whose devices survived. Nil is @@ -98,23 +100,36 @@ func (p Plan) DeviceCount() int { // Explain says why the plan holds nothing, one line per worker. // -// A worker is refused either as a whole — its controllers are on a userspace -// driver, so the kernel presents no disk — or one device at a time, and the two -// need different answers. The first is on the worker's own refusal. The second -// leaves a worker-level reason that is the arithmetic ("no device survived the -// rules") and puts the reason on the devices, so the device refusals are folded -// in behind it, deduplicated and counted. +// A worker is refused either as a whole, because its controllers are on a +// userspace driver and the kernel presents no disk, or one device at a time, and +// the two need different answers. The first is on the worker's own refusal. The +// second leaves a worker-level reason that is the arithmetic ("no device +// survived the rules") and puts the reason on the devices, so the device +// refusals are folded in behind it and counted. +// +// They are counted by the rule that made them, which is the fact the refusal +// carries. Counting by the sentence instead meant a rule whose sentence names +// the device never counted at all: forty disks declined by one allow list +// produced forty clauses, each a sentence long, where the reader wanted a rule +// and a number. Where a rule's sentences agree the sentence is still quoted, +// because for most rules it is the substance. // // Refusals that were pre-filters are left out of both. A machine presents // sixteen network block devices and four disks, and a line saying the sixteen // were not whole disks is true, longer than the rest of the message, and not // the answer to anything. func (p Plan) Explain() []string { + type byRule struct { + rule string + count int + order []string + reasons map[string]int + } type perWorker struct { - worker string - line string - reasons []string - counts map[string]int + worker string + line string + order []string + rules map[string]*byRule } order := make([]string, 0, len(p.Refusals)) @@ -123,7 +138,7 @@ func (p Plan) Explain() []string { if found, ok := byWorker[worker]; ok { return found } - fresh := &perWorker{worker: worker, counts: map[string]int{}} + fresh := &perWorker{worker: worker, rules: map[string]*byRule{}} byWorker[worker] = fresh order = append(order, worker) return fresh @@ -137,28 +152,49 @@ func (p Plan) Explain() []string { if refusal.PreFilter { continue } + entry := at(refusal.Worker) - if _, seen := entry.counts[refusal.Reason]; !seen { - entry.reasons = append(entry.reasons, refusal.Reason) + group, seen := entry.rules[refusal.Rule] + if !seen { + group = &byRule{rule: refusal.Rule, reasons: map[string]int{}} + entry.rules[refusal.Rule] = group + entry.order = append(entry.order, refusal.Rule) } - entry.counts[refusal.Reason]++ + group.count++ + if _, counted := group.reasons[refusal.Reason]; !counted { + group.order = append(group.order, refusal.Reason) + } + group.reasons[refusal.Reason]++ } lines := make([]string, 0, len(order)) for _, worker := range order { entry := byWorker[worker] - if entry.line == "" && len(entry.reasons) == 0 { + if entry.line == "" && len(entry.order) == 0 { // Every refusal on it was a pre-filter and the worker itself was // admitted, so there is nothing about it to explain. continue } - detail := make([]string, 0, len(entry.reasons)) - for _, reason := range entry.reasons { - if count := entry.counts[reason]; count > 1 { - detail = append(detail, fmt.Sprintf("%d devices: %s", count, reason)) + + detail := make([]string, 0, len(entry.order)) + for _, rule := range entry.order { + group := entry.rules[rule] + counted := fmt.Sprintf("%d devices declined by %s", group.count, group.rule) + if group.count == 1 { + counted = "1 device declined by " + group.rule + } + if len(group.order) == 1 { + detail = append(detail, counted+": "+group.order[0]) continue } - detail = append(detail, "1 device: "+reason) + // The rule refused them for different reasons, and the reasons are + // the substance: a disk refused for being mounted and one refused + // for carrying a partition table are two findings. + parts := make([]string, 0, len(group.order)) + for _, reason := range group.order { + parts = append(parts, fmt.Sprintf("%d %s", group.reasons[reason], reason)) + } + detail = append(detail, counted+": "+strings.Join(parts, ", ")) } switch { @@ -221,7 +257,10 @@ func claimableControllers(report nodeprobe.Report, class DeviceClass) []nodeprob var out []nodeprobe.Device for _, controller := range report.NVMeControllers { - if !controller.BoundToUserspace() || controller.InUse { + if !controller.BoundToUserspace() || !controller.Free() { + // Free rather than "not held": a controller the probe could not + // check is not one this may offer. Reading the unchecked state as + // free is how a disk something is driving reaches a draft. continue } if _, already := presented[controller.Address]; already { @@ -241,15 +280,26 @@ func claimableControllers(report nodeprobe.Report, class DeviceClass) []nodeprob return out } -// BasicDeviceRules is the default device pipeline for a class and a filter. +// BasicDeviceRules is the default device pipeline for a filter. // // The order is deliberate and is the order a reader wants the refusal in: what // the device is, then whether it is free, then whether this run wants it. A // partition refused as "not in the allow list" would be a true statement and // the wrong one. -func BasicDeviceRules(class DeviceClass, filter *simplyblockv1alpha2.DeviceFilter) []DeviceRule { +// +// The volume rule sits second, ahead of the class rule, for the same reason +// read the other way. A disk that is one of this fleet's own volumes is refused +// by the class rule too, as a device on a bus this run does not scan, and that +// answer is both true and useless to somebody asking why a machine full of +// disks proposed none. Putting the specific answer first is what puts it in +// front of the reviewer, since the class rule's is a pre-filter and never +// reaches the explanation. +func BasicDeviceRules(filter *simplyblockv1alpha2.DeviceFilter) []DeviceRule { + class := ClassOf(filter) + rules := []DeviceRule{ WholeDiskRule{}, + SimplyblockVolumeRule{}, ClassRule{Class: class}, } @@ -258,9 +308,21 @@ func BasicDeviceRules(class DeviceClass, filter *simplyblockv1alpha2.DeviceFilte rules = append(rules, AvailableRule{AllowPartitioned: allowPartitioned}) if filter == nil { - return rules + // A run with no filter names nothing, so every LUN is refused and every + // other device is taken on its own terms. + return append(rules, ISCSIRule{Class: class}) } + // An iSCSI LUN is refused unless the run's allow list names it, and that + // holds for a run carrying no filter too, which is why the rule is added + // before the filter's own lists and not beside them. + rules = append(rules, ISCSIRule{Class: class, Allow: allowListFor(filter)}) + + // Each class reads its own lists, and the class is the filter's own answer, + // so the branch cannot be the one the filter was not written for. A PCI + // filter beside EnableLogicalBlockDevices describes devices the run will + // never look at, which is why the API refuses that combination rather than + // leaving it to be ignored here. if class == ClassNVMe { if len(filter.PcieAllowList) > 0 || len(filter.PcieDenyList) > 0 { rules = append(rules, AllowDenyRule{ @@ -277,31 +339,43 @@ func BasicDeviceRules(class DeviceClass, filter *simplyblockv1alpha2.DeviceFilte } if filter.DriveSizeRange != "" { - // A range that cannot be parsed is not a range that admits everything. - // ParseSizeRange is called by the caller that validates the spec, and - // this one skips what it cannot read rather than silently widening the - // filter; the run reports the parse failure separately. + // A range that cannot be read builds no rule, which widens the filter + // to every disk rather than narrowing it to none. That is why a run + // carrying one never reaches here: the admission webhook refuses it at + // the request, and the Inspecting step refuses it again for a cluster + // whose webhook is not installed. Skipping is what is left for a caller + // that built a Planner directly, and it is the same answer as omitting + // the field. if min, max, err := ParseSizeRange(filter.DriveSizeRange); err == nil { - rules = append(rules, SizeRule{Min: min, Max: max}) + rules = append(rules, SizeRule{Min: min, Max: max, Spec: filter.DriveSizeRange}) } } return rules } +// allowListFor is the run's allow list in the vocabulary of the class it scans, +// which is the list an iSCSI LUN has to be named in. +func allowListFor(filter *simplyblockv1alpha2.DeviceFilter) []string { + if filter == nil { + return nil + } + if ClassOf(filter) == ClassBlock { + return filter.BlockAllowList + } + return filter.PcieAllowList +} + // Plan runs the pipeline over a fleet's reports. // // The reports are taken in whatever order they arrived and the output does not // depend on it: workers are sorted by name before grouping, so two runs over // one fleet produce the same draft. func (p Planner) Plan(reports []nodeprobe.Report, filter *simplyblockv1alpha2.DeviceFilter) Plan { - class := p.Class - if class == "" { - class = ClassNVMe - } + class := ClassOf(filter) deviceRules := p.DeviceRules if deviceRules == nil { - deviceRules = BasicDeviceRules(class, filter) + deviceRules = BasicDeviceRules(filter) } workerRules := p.WorkerRules if workerRules == nil { @@ -366,7 +440,7 @@ func (p Planner) Plan(reports []nodeprobe.Report, filter *simplyblockv1alpha2.De Class: class, PlacementReason: why, Kube: kube, - MgmtInterface: ManagementInterface(report, kube.InternalIP), + Mgmt: ManagementOf(report, kube.InternalIP), }) } diff --git a/operator/internal/discovery/plan_test.go b/operator/internal/discovery/plan_test.go index 9e8dcef6c..2ec4a0278 100644 --- a/operator/internal/discovery/plan_test.go +++ b/operator/internal/discovery/plan_test.go @@ -237,7 +237,7 @@ func TestPlanScansTheBlockClassWhenAskedTo(t *testing.T) { )} filter := &simplyblockv1alpha2.DeviceFilter{EnableLogicalBlockDevices: ptr.To(true)} - plan := Planner{Class: ClassOf(filter)}.Plan(fleet, filter) + plan := Planner{}.Plan(fleet, filter) if plan.Class != ClassBlock { t.Fatalf("the plan is for class %q, want %q", plan.Class, ClassBlock) @@ -336,11 +336,11 @@ func TestPlanHonorsTheSeamsItWasGiven(t *testing.T) { func TestAnIdleUserspaceControllerIsPlannedOn(t *testing.T) { worker := report("worker-1") worker.NVMeControllers = []nodeprobe.Controller{ - {Address: "0000:00:02.0", Driver: "uio_pci_generic", NUMANode: 0}, - {Address: "0000:00:03.0", Driver: "uio_pci_generic", NUMANode: 0}, + {Address: "0000:00:02.0", Driver: "uio_pci_generic", NUMANode: 0, InUse: ptr.To(false)}, + {Address: "0000:00:03.0", Driver: "uio_pci_generic", NUMANode: 0, InUse: ptr.To(false)}, } - plan := Planner{Class: ClassNVMe}.Plan([]nodeprobe.Report{worker}, nil) + plan := Planner{}.Plan([]nodeprobe.Report{worker}, nil) if len(plan.Workers) != 1 { t.Fatalf("planned %d workers, want the one: %s", len(plan.Workers), plan.Summary()) @@ -365,10 +365,10 @@ func TestAnIdleUserspaceControllerIsPlannedOn(t *testing.T) { func TestAHeldUserspaceControllerIsNotPlannedOn(t *testing.T) { worker := report("worker-1") worker.NVMeControllers = []nodeprobe.Controller{ - {Address: "0000:00:02.0", Driver: "uio_pci_generic", NUMANode: 0, InUse: true}, + {Address: "0000:00:02.0", Driver: "uio_pci_generic", NUMANode: 0, InUse: ptr.To(true)}, } - plan := Planner{Class: ClassNVMe}.Plan([]nodeprobe.Report{worker}, nil) + plan := Planner{}.Plan([]nodeprobe.Report{worker}, nil) if len(plan.Workers) != 0 { t.Errorf("a worker whose only controller is in use was planned on: %s", plan.Summary()) @@ -384,7 +384,7 @@ func TestAKernelBoundControllerIsNotCountedTwice(t *testing.T) { {Address: "0000:00:02.0", Driver: "nvme", NUMANode: 0}, } - plan := Planner{Class: ClassNVMe}.Plan([]nodeprobe.Report{worker}, nil) + plan := Planner{}.Plan([]nodeprobe.Report{worker}, nil) var addresses []string for _, set := range plan.NodeSets { @@ -396,3 +396,91 @@ func TestAKernelBoundControllerIsNotCountedTwice(t *testing.T) { t.Errorf("the draft names %v, want the one disk once", addresses) } } + +// The class a run scans is the filter's to state, and the planner has no second +// opinion about it. +// +// Two statements of one fact can disagree, and this one disagreed silently: a +// planner told nothing scanned NVMe, so a filter asking for logical block +// devices had its allow and deny lists read on the branch that never runs and +// dropped without a word. There is nothing for a caller to keep in step now, +// because there is only one place the class is written down. +func TestTheFilterDecidesTheClassWithoutBeingToldTwice(t *testing.T) { + fleet := []nodeprobe.Report{report("worker-1", + blockDisk("vdb", 0, 2*tb), + blockDisk("vdc", 0, 2*tb), + )} + filter := &simplyblockv1alpha2.DeviceFilter{ + EnableLogicalBlockDevices: ptr.To(true), + BlockDenyList: []string{"/dev/vdc"}, + } + + plan := Planner{}.Plan(fleet, filter) + + if plan.Class != ClassBlock { + t.Fatalf("the plan is for class %q, want %q", plan.Class, ClassBlock) + } + if len(plan.NodeSets) != 1 || len(plan.NodeSets[0].Groups) != 1 { + t.Fatalf("built %+v", plan.NodeSets) + } + if got := plan.NodeSets[0].Groups[0].Devices.Block; !slices.Equal(got, []string{"/dev/vdb"}) { + t.Errorf("the group names %v, want the deny list to have been applied", got) + } +} + +// The explanation groups refusals by the rule that made them, not by the words +// they came out as. +// +// Every refusal carries the rule that produced it, so the count is a fact the +// data already holds. Grouping on the rendered sentence instead meant a rule +// whose sentence names the device never grouped at all: forty disks declined by +// one allow list produced forty clauses, each a sentence long, where the reader +// wanted one line and a number. +func TestTheExplanationCountsByRuleRatherThanByWording(t *testing.T) { + filter := &simplyblockv1alpha2.DeviceFilter{PcieAllowList: []string{"0000:ff:00.0"}} + fleet := []nodeprobe.Report{report("worker-1", + disk("nvme0n1", "0000:5e:00.0", 0, tb), + disk("nvme1n1", "0000:5f:00.0", 0, tb), + disk("nvme2n1", "0000:af:00.0", 0, tb), + disk("nvme3n1", "0000:b0:00.0", 0, tb), + )} + + plan := Planner{}.Plan(fleet, filter) + + explained := plan.Explain() + if len(explained) != 1 { + t.Fatalf("explained %d workers, want 1: %v", len(explained), explained) + } + line := explained[0] + + if !strings.Contains(line, "4 devices") { + t.Errorf("the line %q does not count the four disks together", line) + } + // One clause, not four: the addresses are what the refusals differ in, and + // they are not what the reader is counting. + if strings.Count(line, "allow list") > 1 { + t.Errorf("the line %q repeats the rule once per device", line) + } +} + +// Two rules refusing a worker's disks are two clauses, because the reason a +// disk was left out is the rule that left it out. +func TestTheExplanationKeepsTheRulesApart(t *testing.T) { + filter := &simplyblockv1alpha2.DeviceFilter{ + PcieDenyList: []string{"0000:5e:00.0"}, + DriveSizeRange: "2T-4T", + } + fleet := []nodeprobe.Report{report("worker-1", + disk("nvme0n1", "0000:5e:00.0", 0, 3*tb), + disk("nvme1n1", "0000:5f:00.0", 0, tb), + )} + + plan := Planner{}.Plan(fleet, filter) + + line := strings.Join(plan.Explain(), " ") + for _, rule := range []string{"allow and deny lists", "size range"} { + if !strings.Contains(line, rule) { + t.Errorf("the explanation %q does not name the %s rule", line, rule) + } + } +} diff --git a/operator/internal/discovery/rules.go b/operator/internal/discovery/rules.go index 0e25c3b32..9bbd692b2 100644 --- a/operator/internal/discovery/rules.go +++ b/operator/internal/discovery/rules.go @@ -158,12 +158,20 @@ func (r AvailableRule) Admit(_ nodeprobe.Report, device nodeprobe.Device) (bool, return false, "the probe refused it: " + strings.Join(reasons, ", ") } -// ClassRule admits a device that can be named in the class the run is scanning. +// ClassRule admits a device that can be named in the class the run is scanning, +// and refuses one of the other class. // // It is a rule rather than a precondition because the failure is worth // reporting: an NVMe run against a worker whose disks are virtio finds devices // it cannot name, and "no NVMe devices" is a more useful answer than an empty // draft. +// +// The check is symmetric, and it has to be. A cluster is built out of one class +// of backend storage, so the two classes cannot both reach one draft — and +// while an NVMe run recognizes its own class by the bus, a block run cannot +// recognize its own by the address, because an NVMe device has a path like +// every other block device. Naming devices by path is what the block class does +// rather than what makes a device one, so the bus is what both sides read. type ClassRule struct { Class DeviceClass } @@ -179,12 +187,27 @@ func (r ClassRule) Admit(_ nodeprobe.Report, device nodeprobe.Device) (bool, str return false, fmt.Sprintf("this run scans NVMe devices and the device is on %s", transportOrNone(device.Transport)) } + if r.Class == ClassBlock && onNVMe(device.Transport) { + return false, fmt.Sprintf("this run scans logical block devices and the device is on %s", + transportOrNone(device.Transport)) + } if address := r.Class.Address(device); address == "" { return false, fmt.Sprintf("it has no %s to name it by", addressKind(r.Class)) } return true, "" } +// onNVMe reports whether a bus is one of the two the NVMe class covers. +// +// A fabric namespace is in, and it is the one worth naming: it is a volume +// something else exported, so it is neither class's candidate, and refusing it +// here rather than leaving it to the availability rule keeps a block run's +// refusal about what the device is rather than about what is holding it. +func onNVMe(transport string) bool { + return transport == string(blockdev.TransportNVMe) || + transport == string(blockdev.TransportNVMeFabric) +} + // transportOrNone names a transport for a message, including the empty one. func transportOrNone(transport string) string { if transport == "" { @@ -234,12 +257,17 @@ type AllowDenyRule struct { func (AllowDenyRule) Name() string { return "allow and deny lists" } +// The reason names the list and not the address. Which device was refused is +// already on the refusal, and putting it in the sentence too made every such +// refusal a different sentence, so a hundred disks declined by one list read as +// a hundred separate findings rather than one list and a number. + func (r AllowDenyRule) Admit(_ nodeprobe.Report, device nodeprobe.Device) (bool, string) { address := r.Class.Address(device) for _, denied := range r.Deny { if strings.EqualFold(address, denied) { - return false, fmt.Sprintf("%s is in the deny list", address) + return false, "it is in the deny list" } } if len(r.Allow) == 0 { @@ -250,7 +278,7 @@ func (r AllowDenyRule) Admit(_ nodeprobe.Report, device nodeprobe.Device) (bool, return true, "" } } - return false, fmt.Sprintf("%s is not in the allow list", address) + return false, "it is not in the allow list" } // ModelRule admits a device whose model string contains the wanted text. @@ -278,22 +306,40 @@ func (r ModelRule) Admit(_ nodeprobe.Report, device nodeprobe.Device) (bool, str type SizeRule struct { // Min and Max are inclusive bounds in bytes. A zero Max is no upper bound. Min, Max uint64 + + // Spec is the range as the filter wrote it, which is what a refusal quotes. + // + // The bound is not re-rendered from the parsed number, because that is a + // different string: a filter naming 1920G would be quoted back as 1.875T, + // and a reviewer comparing the refusal against what they wrote would be + // comparing two spellings of one number. What they can act on is what they + // typed. + Spec string } func (SizeRule) Name() string { return "size range" } func (r SizeRule) Admit(_ nodeprobe.Report, device nodeprobe.Device) (bool, string) { if device.SizeBytes < r.Min { - return false, fmt.Sprintf("it is %s and the range starts at %s", - humanBytes(device.SizeBytes), humanBytes(r.Min)) + return false, fmt.Sprintf("it is %s and the range %s starts above it", + humanBytes(device.SizeBytes), r.describeRange(r.Min)) } if r.Max > 0 && device.SizeBytes > r.Max { - return false, fmt.Sprintf("it is %s and the range ends at %s", - humanBytes(device.SizeBytes), humanBytes(r.Max)) + return false, fmt.Sprintf("it is %s and the range %s ends below it", + humanBytes(device.SizeBytes), r.describeRange(r.Max)) } return true, "" } +// describeRange is the range as the filter wrote it, and the bound rendered +// when the rule was built by hand rather than from a filter. +func (r SizeRule) describeRange(bound uint64) string { + if r.Spec != "" { + return r.Spec + } + return humanBytes(bound) +} + // WorkerHasDevices admits a worker that has at least one device left. // // It is the only worker rule, and it is the one that cannot be omitted: a @@ -320,12 +366,21 @@ func (WorkerHasDevices) Admit(report nodeprobe.Report, admitted []nodeprobe.Devi // Something holding them means the machine is serving whatever that is, // which need not be this product: a userspace binding is also how a // disk is passed through to a guest. - var busy []nodeprobe.Controller + var busy, unchecked []nodeprobe.Controller for _, controller := range bound { - if controller.InUse { + switch { + case controller.Held(): busy = append(busy, controller) + case !controller.Free(): + unchecked = append(unchecked, controller) } } + if len(busy) == 0 && len(unchecked) > 0 { + return false, fmt.Sprintf( + "it presents no usable block device, and whether anything is driving %d of its "+ + "NVMe controllers (%s) could not be established, so they are not offered", + len(unchecked), describeControllers(unchecked)) + } if len(busy) > 0 { return false, fmt.Sprintf( "it presents no usable block device, and %d of its NVMe controllers (%s) are "+ @@ -368,17 +423,40 @@ func (WorkerWasReadable) Admit(report nodeprobe.Report, _ []nodeprobe.Device) (b len(report.Unreadable), strings.Join(report.Unreadable, "; ")) } -// humanBytes renders a size the way an administrator writes one, so that a -// refusal quotes the same units the filter was written in. +// humanBytes renders a size the way an administrator writes one. +// +// Two properties, and the old rendering had neither. A whole number of units is +// written as that whole number, so the sizes a fleet is actually built out of +// read as 3T rather than as 3.001T and parse back to the byte they came from. +// Anything else is truncated rather than rounded, because a size that reads as +// more than the device holds is one that says the disk is bigger than it is: +// the old rendering printed a byte under a tebibyte as 1024G, which is not a +// approximation of the value, it is above it. func humanBytes(bytes uint64) string { - switch { - case bytes >= 1<<40: - return fmt.Sprintf("%.4gT", float64(bytes)/float64(uint64(1)<<40)) - case bytes >= 1<<30: - return fmt.Sprintf("%.4gG", float64(bytes)/float64(uint64(1)<<30)) - case bytes >= 1<<20: - return fmt.Sprintf("%.4gM", float64(bytes)/float64(uint64(1)<<20)) - default: - return fmt.Sprintf("%dB", bytes) + for _, unit := range []struct { + size uint64 + suffix string + }{ + {uint64(1) << 40, "T"}, + {uint64(1) << 30, "G"}, + {uint64(1) << 20, "M"}, + } { + if bytes < unit.size { + continue + } + if bytes%unit.size == 0 { + return fmt.Sprintf("%d%s", bytes/unit.size, unit.suffix) + } + // Truncated to two decimals, which is the precision a disk is sold in + // and never more than the device holds. A trailing zero is dropped, so + // a size and a half reads as 1.5T rather than 1.50T. + hundredths := (bytes % unit.size) * 100 / unit.size + decimals := fmt.Sprintf("%02d", hundredths) + decimals = strings.TrimRight(decimals, "0") + if decimals == "" { + return fmt.Sprintf("%d%s", bytes/unit.size, unit.suffix) + } + return fmt.Sprintf("%d.%s%s", bytes/unit.size, decimals, unit.suffix) } + return fmt.Sprintf("%dB", bytes) } diff --git a/operator/internal/discovery/rules_test.go b/operator/internal/discovery/rules_test.go index 9e992d2af..453a44667 100644 --- a/operator/internal/discovery/rules_test.go +++ b/operator/internal/discovery/rules_test.go @@ -17,7 +17,12 @@ import ( "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" ) -const tb = uint64(1) << 40 +const ( + mib = uint64(1) << 20 + gib = uint64(1) << 30 + tib = uint64(1) << 40 + tb = tib +) // admit runs one rule and returns its verdict and reason. func admit(rule DeviceRule, device nodeprobe.Device) (bool, string) { @@ -154,12 +159,12 @@ func TestSizeRuleBoundsBothEnds(t *testing.T) { if ok, why := admit(rule, small); ok { t.Error("admitted a disk below the range") - } else if !strings.Contains(why, "range starts at") { + } else if !strings.Contains(why, "starts above it") { t.Errorf("the reason %q does not say which end it missed", why) } if ok, why := admit(rule, large); ok { t.Error("admitted a disk above the range") - } else if !strings.Contains(why, "range ends at") { + } else if !strings.Contains(why, "ends below it") { t.Errorf("the reason %q does not say which end it missed", why) } if ok, _ := admit(rule, disk("nvme2n1", "0000:af:00.0", 0, 2*tb)); !ok { @@ -214,8 +219,8 @@ func TestWorkerHasDevicesExplainsAMachineWhoseDisksAreAlreadyDriven(t *testing.T // full of disks. worker := report("worker-1") worker.NVMeControllers = []nodeprobe.Controller{ - {Address: "0000:00:02.0", Driver: "uio_pci_generic", InUse: true}, - {Address: "0000:00:03.0", Driver: "uio_pci_generic", InUse: true}, + {Address: "0000:00:02.0", Driver: "uio_pci_generic", InUse: ptr.To(true)}, + {Address: "0000:00:03.0", Driver: "uio_pci_generic", InUse: ptr.To(true)}, } ok, why := (WorkerHasDevices{}).Admit(worker, nil) @@ -324,8 +329,8 @@ func TestParseSizeRange(t *testing.T) { func TestWorkerHasDevicesSeparatesReclaimableFromInUse(t *testing.T) { idle := report("worker-1") idle.NVMeControllers = []nodeprobe.Controller{ - {Address: "0000:00:02.0", Driver: "uio_pci_generic"}, - {Address: "0000:00:03.0", Driver: "uio_pci_generic"}, + {Address: "0000:00:02.0", Driver: "uio_pci_generic", InUse: ptr.To(false)}, + {Address: "0000:00:03.0", Driver: "uio_pci_generic", InUse: ptr.To(false)}, } _, why := (WorkerHasDevices{}).Admit(idle, nil) @@ -335,7 +340,7 @@ func TestWorkerHasDevicesSeparatesReclaimableFromInUse(t *testing.T) { busy := report("worker-2") busy.NVMeControllers = []nodeprobe.Controller{ - {Address: "0000:00:04.0", Driver: "vfio-pci", InUse: true}, + {Address: "0000:00:04.0", Driver: "vfio-pci", InUse: ptr.To(true)}, } _, why = (WorkerHasDevices{}).Admit(busy, nil) @@ -346,3 +351,110 @@ func TestWorkerHasDevicesSeparatesReclaimableFromInUse(t *testing.T) { t.Errorf("a held controller was offered for reclaiming: %q", why) } } + +// A cluster is built out of one class of backend storage, so a run that scans +// logical block devices refuses an NVMe device rather than naming it by its +// path. +// +// The rule used to check the transport on one side only: an NVMe run refused a +// virtio disk, and a block run admitted an NVMe disk, because an NVMe disk has +// a path like any other. That put both classes in one draft, and a draft is one +// cluster. +func TestClassRuleRefusesTheOtherClassOnABlockRun(t *testing.T) { + nvme := disk("nvme0n1", "0000:5e:00.0", 0, tb) + + if ok, why := admit(ClassRule{Class: ClassBlock}, nvme); ok { + t.Error("a block run admitted an NVMe disk, so a draft could hold both classes") + } else if !strings.Contains(why, "NVMe") { + t.Errorf("the reason %q does not say what bus it is on", why) + } + + // A fabric namespace is a volume something else exported, and it is the + // other class read the other way: the probe refuses it first, and this is + // the rule that keeps it out of a block draft on its own terms. + fabric := blockDisk("nvme1n1", 0, tb) + fabric.Transport = string(blockdev.TransportNVMeFabric) + if ok, why := admit(ClassRule{Class: ClassBlock}, fabric); ok { + t.Error("a block run admitted a fabric namespace") + } else if !strings.Contains(why, "NVMeFabric") { + t.Errorf("the reason %q does not say what bus it is on", why) + } + + // Every other bus is what a block run is for. + for _, transport := range []blockdev.Transport{ + blockdev.TransportVirtio, blockdev.TransportSATA, + blockdev.TransportSAS, blockdev.TransportSCSI, + } { + device := blockDisk("sda", 0, tb) + device.Transport = string(transport) + if ok, why := admit(ClassRule{Class: ClassBlock}, device); !ok { + t.Errorf("a block run declined a %s disk: %s", transport, why) + } + } +} + +// A refusal quotes the bound the way the filter wrote it. +// +// It used to re-render the parsed number, which is a different string: the +// renderer prints four significant digits, so a bound of 1920G came back as +// 1.875T and a reviewer comparing the refusal against their own filter was +// comparing two spellings of one number. What they wrote is what they can act +// on, so that is what the refusal says. +func TestTheRefusalQuotesTheBoundAsItWasWritten(t *testing.T) { + rules := BasicDeviceRules(&simplyblockv1alpha2.DeviceFilter{DriveSizeRange: "1920G-4T"}) + + var size SizeRule + for _, rule := range rules { + if found, is := rule.(SizeRule); is { + size = found + } + } + if size.Spec == "" { + t.Fatal("the size rule does not carry the range as it was written") + } + + _, why := size.Admit(nodeprobe.Report{}, disk("nvme0n1", "0000:5e:00.0", 0, 512*gib)) + if !strings.Contains(why, "1920G-4T") { + t.Errorf("the reason %q does not quote the range the filter carried", why) + } +} + +// A size a reviewer reads is never rounded up, because a rendering that reads +// as more than the device holds is one that says the disk is bigger than it is. +func TestASizeIsNeverRoundedUp(t *testing.T) { + for _, c := range []struct { + bytes uint64 + want string + }{ + {tib, "1T"}, + {3 * tib, "3T"}, + {1536 * gib, "1.5T"}, + {tib - 1, "1023.99G"}, + {512 * mib, "512M"}, + {1023, "1023B"}, + {0, "0B"}, + } { + if got := humanBytes(c.bytes); got != c.want { + t.Errorf("%d bytes renders as %q, want %q", c.bytes, got, c.want) + } + } +} + +// An exact multiple of a unit is written as that unit, with no decimals and no +// loss, which is what makes the common sizes readable. +func TestAnExactSizeIsWrittenExactly(t *testing.T) { + for _, size := range []uint64{mib, gib, tib, 3 * tib, 512 * gib, 100 * gib} { + rendered := humanBytes(size) + if strings.Contains(rendered, ".") { + t.Errorf("%d bytes is a whole number of units and renders as %q", size, rendered) + } + parsed, err := ParseSize(rendered) + if err != nil { + t.Errorf("%d bytes renders as %q, which the parser refuses: %v", size, rendered, err) + continue + } + if parsed != size { + t.Errorf("%d bytes renders as %q, which reads back as %d", size, rendered, parsed) + } + } +} diff --git a/operator/internal/discovery/simplyblock.go b/operator/internal/discovery/simplyblock.go new file mode 100644 index 000000000..fca911573 --- /dev/null +++ b/operator/internal/discovery/simplyblock.go @@ -0,0 +1,50 @@ +// The rule that keeps a fleet's own volumes out of its own draft. +// +// A simplyblock volume attached to a worker is an NVMe-oF namespace, and the +// kernel presents it as a disk like any other. Nothing about its shape says +// whose it is: a fabric namespace is a volume somebody exported, and a fabric +// namespace whose subsystem NQN names a simplyblock cluster and a logical +// volume is one this product exported. Only the second is a disk a draft would +// be giving back to the cluster it came from, and only the NQN tells them apart. +// +// The class rule already refuses every fabric namespace, so this rule changes +// no draft today. What it changes is the answer: a reviewer asking why a +// machine full of disks proposed none is told that the disks are the fleet's +// own volumes rather than that they are on a bus the run does not scan. It is +// also the check that survives the class rule being relaxed, which is the one +// way a cluster could be told to take its own bytes as free space. + +package discovery + +import ( + "fmt" + + "github.com/simplyblock/atlas/nqn" + "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" +) + +// SimplyblockVolumeRule refuses a device that is one of this product's own +// logical volumes. +type SimplyblockVolumeRule struct{} + +func (SimplyblockVolumeRule) Name() string { return "simplyblock volume" } + +// Admit refuses a device whose subsystem NQN parses as a simplyblock +// logical-volume subsystem. +// +// Parsing rather than matching a prefix is what makes the refusal specific. The +// NQN carries the cluster and the volume it belongs to, so the reason can name +// the cluster, and a subsystem of this product that is not a logical volume +// does not read as one. +func (SimplyblockVolumeRule) Admit(_ nodeprobe.Report, device nodeprobe.Device) (bool, string) { + if device.SubsystemNQN == "" { + return true, "" + } + subsystem, isVolume := nqn.Parse(device.SubsystemNQN) + if !isVolume { + return true, "" + } + return false, fmt.Sprintf( + "it is a simplyblock logical volume of cluster %s, which this fleet already serves", + subsystem.ClusterID) +} diff --git a/operator/internal/discovery/simplyblock_test.go b/operator/internal/discovery/simplyblock_test.go new file mode 100644 index 000000000..3722a9734 --- /dev/null +++ b/operator/internal/discovery/simplyblock_test.go @@ -0,0 +1,120 @@ +// That a run never proposes a disk that is already a simplyblock volume. +// +// A volume this product exported and some node attached is an NVMe-oF namespace +// like any other, and the machine it is attached to presents it as a disk. What +// separates it from a disk the fleet owns is its subsystem NQN, which names the +// cluster and the logical volume it belongs to. Handing one back to a cluster as +// backend storage would give a volume's own bytes away as free space, and would +// do it to the product's own data. + +package discovery + +import ( + "strings" + "testing" + + "github.com/simplyblock/atlas/blockdev" + "github.com/simplyblock/atlas/nqn" + "github.com/simplyblock/atlas/ptr" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" +) + +// attached is a simplyblock volume as a worker that has connected it sees: a +// namespace on a fabric, carrying the NQN the product builds. +func attached(name, clusterID, volumeID string) nodeprobe.Device { + device := blockDisk(name, 0, tb) + device.Transport = string(blockdev.TransportNVMeFabric) + device.SubsystemNQN = nqn.Make(clusterID, volumeID) + return device +} + +func TestASimplyblockVolumeIsNeverProposed(t *testing.T) { + const cluster = "c30a691a-1d2e-4f3a-9b8c-5d6e7f809a1b" + volume := attached("nvme3n1", cluster, "792e184c-0a1b-2c3d-4e5f-60718293a4b5") + + ok, why := admit(SimplyblockVolumeRule{}, volume) + if ok { + t.Fatal("a run admitted a disk that is one of this product's own volumes") + } + if !strings.Contains(why, cluster) { + t.Errorf("the reason %q does not name the cluster the volume belongs to", why) + } +} + +func TestADiskThatIsNotAVolumeIsUntouchedByTheRule(t *testing.T) { + for _, device := range []nodeprobe.Device{ + disk("nvme0n1", "0000:5e:00.0", 0, tb), + blockDisk("vdb", 0, tb), + } { + if ok, why := admit(SimplyblockVolumeRule{}, device); !ok { + t.Errorf("%s was refused as a simplyblock volume: %s", device.Name, why) + } + } + + // A fabric namespace some other product exported is refused for being on a + // fabric, by the class rule, and not by this one: what this rule says is + // that the disk is ours, and it is not. + foreign := blockDisk("nvme4n1", 0, tb) + foreign.Transport = string(blockdev.TransportNVMeFabric) + foreign.SubsystemNQN = "nqn.2019-08.org.ceph:rbd.pool.image" + if ok, _ := admit(SimplyblockVolumeRule{}, foreign); !ok { + t.Error("a namespace another product exported was called a simplyblock volume") + } +} + +func TestTheRuleIsInEveryRunsPipeline(t *testing.T) { + // It has to hold for both classes and for a run with no filter at all: a + // worker with volumes attached is the ordinary case on a fleet this product + // already serves, and the rule that keeps them out cannot be one a filter + // switches on. + volume := attached("nvme3n1", "c30a691a", "792e184c") + + runs := []*simplyblockv1alpha2.DeviceFilter{ + nil, + {EnableLogicalBlockDevices: ptr.To(true)}, + } + + for _, filter := range runs { + var found bool + for _, rule := range BasicDeviceRules(filter) { + if _, isRule := rule.(SimplyblockVolumeRule); isRule { + found = true + } + } + if !found { + t.Errorf("a %s run has no rule refusing this product's own volumes", ClassOf(filter)) + } + } + + // And the whole pipeline refuses it, whichever class is scanned. + for _, filter := range runs { + report := report("worker-1") + report.Devices = []nodeprobe.Device{volume} + plan := Planner{}.Plan([]nodeprobe.Report{report}, filter) + if len(plan.NodeSets) != 0 { + t.Errorf("a %s run drafted a simplyblock volume", ClassOf(filter)) + } + } +} + +func TestTheRefusalIsReadableRatherThanFilteredAway(t *testing.T) { + // A worker whose disks are all attached volumes is a worker somebody will + // ask about, so the answer has to survive into the explanation rather than + // being folded away as a device of the wrong class. + report := report("worker-1") + report.Devices = []nodeprobe.Device{ + attached("nvme3n1", "c30a691a", "792e184c"), + attached("nvme4n1", "c30a691a", "8a3f0b12"), + } + + plan := Planner{}.Plan([]nodeprobe.Report{report}, nil) + + explained := strings.Join(plan.Explain(), " ") + if !strings.Contains(explained, "simplyblock") { + t.Errorf("the explanation %q does not say the disks are already volumes", explained) + } + if !strings.Contains(explained, "2 devices") { + t.Errorf("the explanation %q does not count the two disks together", explained) + } +} diff --git a/operator/internal/nodeprobe/collect.go b/operator/internal/nodeprobe/collect.go index 19702b9fb..173265578 100644 --- a/operator/internal/nodeprobe/collect.go +++ b/operator/internal/nodeprobe/collect.go @@ -106,6 +106,9 @@ func interfacesOf(ifaces []inventory.Interface) []Interface { Virtual: iface.Virtual, Loopback: iface.Loopback, Bridge: iface.Bridge, + Kind: string(iface.Kind), + Lower: iface.Lower, + Upper: iface.Upper, Addresses: iface.Addresses, }) } @@ -116,19 +119,20 @@ func devicesOf(candidates []blockdev.Candidate) []Device { out := make([]Device, 0, len(candidates)) for _, c := range candidates { device := Device{ - Name: c.Name, - Path: c.Path, - PCIAddress: c.PCIAddress, - SizeBytes: c.SizeBytes, - Kind: string(c.Kind), - Transport: string(c.Transport), - Vendor: c.Vendor, - Model: c.Model, - Serial: c.Serial, - Rotational: c.Rotational, - NUMANode: c.NUMANode, - Available: c.Available(), - ContentType: c.Reading.Type, + Name: c.Name, + Path: c.Path, + PCIAddress: c.PCIAddress, + SizeBytes: c.SizeBytes, + Kind: string(c.Kind), + Transport: string(c.Transport), + Vendor: c.Vendor, + Model: c.Model, + Serial: c.Serial, + Rotational: c.Rotational, + NUMANode: c.NUMANode, + SubsystemNQN: c.SubsystemNQN, + Available: c.Available(), + ContentType: c.Reading.Type, } // A device refused before anything was opened carries ContentUnknown, // and the report leaves the field empty rather than writing "Unknown": diff --git a/operator/internal/nodeprobe/netstack_test.go b/operator/internal/nodeprobe/netstack_test.go new file mode 100644 index 000000000..beb6db01d --- /dev/null +++ b/operator/internal/nodeprobe/netstack_test.go @@ -0,0 +1,118 @@ +// That the interface kind and the stack around it survive the trip into the +// report, and that a reader refuses the schema that predates them. +// +// The three fields are what make the management-interface rule possible: every +// software interface is virtual, and only the kind separates a bond or a tagged +// VLAN from a veth. A report that carried the readings and dropped these would +// leave the rule with the same collapsed answer it had. + +package nodeprobe + +import ( + "encoding/json" + "slices" + "testing" + "time" + + "github.com/simplyblock/atlas/inventory" +) + +func TestFromInventoryCarriesTheKindAndTheStack(t *testing.T) { + report := FromInventory("worker-1", time.Now(), inventory.Inventory{ + Interfaces: []inventory.Interface{ + { + Name: "bond0", Kind: inventory.LinkBond, Virtual: true, + Lower: []string{"eth0", "eth1"}, Upper: []string{"bond0.100"}, + SpeedMbps: 50000, OperState: inventory.LinkUp, + }, + { + Name: "bond0.100", Kind: inventory.LinkVLAN, Virtual: true, + Lower: []string{"bond0"}, Addresses: []string{"10.0.0.11"}, + OperState: inventory.LinkUp, + }, + { + Name: "eth0", Kind: inventory.LinkPhysical, + Upper: []string{"bond0"}, PCIAddress: "0000:3b:00.0", NUMANode: 0, + }, + }, + }, nil) + + byName := map[string]Interface{} + for _, iface := range report.Interfaces { + byName[iface.Name] = iface + } + + bond := byName["bond0"] + if bond.Kind != string(inventory.LinkBond) { + t.Errorf("the bond reports kind %q, want %q", bond.Kind, inventory.LinkBond) + } + if !slices.Equal(bond.Lower, []string{"eth0", "eth1"}) { + t.Errorf("the bond reports members %v, want both NICs", bond.Lower) + } + if !slices.Equal(bond.Upper, []string{"bond0.100"}) { + t.Errorf("the bond reports %v stacked on it, want the VLAN", bond.Upper) + } + if vlan := byName["bond0.100"]; vlan.Kind != string(inventory.LinkVLAN) || + !slices.Equal(vlan.Lower, []string{"bond0"}) { + t.Errorf("the VLAN reads as %+v", vlan) + } + if nic := byName["eth0"]; nic.Kind != string(inventory.LinkPhysical) || + !slices.Equal(nic.Upper, []string{"bond0"}) { + t.Errorf("the NIC reads as %+v", nic) + } +} + +func TestTheKindAndTheStackSurviveTheJSON(t *testing.T) { + // The report is a wire format between a probe pod and the operator, so a + // field that is not marshaled is a field the discovery step never sees. + encoded, err := Encode(Report{ + Node: "worker-1", + Interfaces: []Interface{{ + Name: "bond0", Kind: string(inventory.LinkBond), + Lower: []string{"eth0", "eth1"}, Upper: []string{"bond0.100"}, + }}, + }) + if err != nil { + t.Fatalf("encode: %v", err) + } + + decoded, err := Decode(encoded) + if err != nil { + t.Fatalf("decode: %v", err) + } + if len(decoded.Interfaces) != 1 { + t.Fatalf("decoded %d interfaces", len(decoded.Interfaces)) + } + iface := decoded.Interfaces[0] + if iface.Kind != string(inventory.LinkBond) || + !slices.Equal(iface.Lower, []string{"eth0", "eth1"}) || + !slices.Equal(iface.Upper, []string{"bond0.100"}) { + t.Errorf("the round trip produced %+v", iface) + } + + var raw map[string]any + if err := json.Unmarshal(encoded, &raw); err != nil { + t.Fatalf("unmarshal: %v", err) + } + first := raw["interfaces"].([]any)[0].(map[string]any) + for _, key := range []string{"kind", "lower", "upper"} { + if _, carried := first[key]; !carried { + t.Errorf("the encoded interface carries no %q", key) + } + } +} + +func TestAReportFromTheSchemaBeforeTheStackIsRefused(t *testing.T) { + // A Job keeps the image it started with, so an operator upgraded mid-run + // reads reports from the previous probe. One written before the kind existed + // describes its bonds as virtual and nothing else, and reading it as current + // would put the rule back where it was. + before := []byte(`{"version":3,"node":"worker-1"}`) + + if _, err := Decode(before); err == nil { + t.Fatal("a report from the previous schema was accepted") + } + if ReportVersion < 4 { + t.Errorf("the report is version %d, and the kind and the stack arrived at version 4", ReportVersion) + } +} diff --git a/operator/internal/nodeprobe/report.go b/operator/internal/nodeprobe/report.go index 3c4263e79..a38c75352 100644 --- a/operator/internal/nodeprobe/report.go +++ b/operator/internal/nodeprobe/report.go @@ -38,7 +38,7 @@ import ( // because a probe pod outlives the operator that created it across an upgrade: // the image is pinned in the Job, and a Job already running keeps the image it // started with. -const ReportVersion = 3 +const ReportVersion = 6 // Report is one worker's inventory as the probe found it. type Report struct { @@ -214,10 +214,34 @@ type Interface struct { // cluster's own CNI leaves on every worker. It is separate from Virtual // because the two answer different questions: Virtual says the interface is // backed by no hardware, and Bridge says it is carrying somebody else's - // traffic, which is what disqualifies it from being a management address - // even where it holds one. + // traffic. Bridge bool `json:"bridge,omitempty"` + // Kind is what sort of device the interface is, in the spelling + // inventory.LinkKind uses: `physical`, `loopback`, `bridge`, `bond`, `team`, + // `vlan`, `vxlan`, `macvlan`, `ipvlan`, or `virtual` for a device the kernel + // does not identify. + // + // It is the field the management-interface rule turns on. Every software + // interface a worker has is virtual, so without the kind a bond and a tagged + // VLAN, which is where most fleets put their management address, read the + // same as a veth to a pod. + Kind string `json:"kind,omitempty"` + + // Lower is what this interface is built on, ascending by name: the members + // of a bond or a bridge, or the single parent of a VLAN or a macvlan. + // + // It is how a draft reaches the hardware under an aggregate. A bond reports + // no slot, no driver, and no memory node of its own, so a reader that has + // to know where a bonded management network lands looks its members up in + // the same report. + Lower []string `json:"lower,omitempty"` + + // Upper is what is built on this interface, ascending by name. It is the + // direction that answers for a NIC holding no address: on a host whose + // management network is tagged, the address is on a VLAN above it. + Upper []string `json:"upper,omitempty"` + // Addresses are the IP addresses the interface holds, without a prefix // length. They are what makes a management interface identifiable: the one // a draft names is the one carrying the address the cluster already reaches @@ -250,6 +274,16 @@ type Device struct { // NUMANode is the memory node the device hangs off, or NUMANodeUnknown. NUMANode int `json:"numaNode"` + // SubsystemNQN is the NVMe Qualified Name of the subsystem the namespace + // belongs to, and is empty for a device on any other bus. + // + // It is what tells a volume this product exported from a disk the fleet + // owns. Both are namespaces, both are presented as disks, and the transport + // separates them only by inference: what an NQN says is which cluster and + // which logical volume the bytes belong to, which is the answer a draft + // needs before it proposes handing them to a cluster. + SubsystemNQN string `json:"subsystemNQN,omitempty"` + // Available reports whether the device may be handed over: a whole disk, on // a bus the scan recognized, that nothing is using and that positively // reads as holding nothing. @@ -291,7 +325,8 @@ type Controller struct { // NUMANode is the memory node it hangs off, or NUMANodeUnknown. NUMANode int `json:"numaNode"` - // InUse reports whether anything holds the controller open. + // InUse reports whether anything holds the controller open, and is absent + // for a controller the probe could not check. // // It is the question Driver cannot answer and the one that decides whether // a controller can be reclaimed: a userspace binding nothing is driving is @@ -299,13 +334,22 @@ type Controller struct { // service, which may belong to a hypervisor guest or another product rather // than to this one. // - // False on a controller the probe could not check is the zero value and not - // an answer. A probe that failed to read the process table says so in - // Unreadable, so a reader deciding whether to reclaim has to find this - // report free of such an entry first. - InUse bool `json:"inUse,omitempty"` + // The third state is why this is a pointer. A controller nothing holds and + // a controller whose process table could not be read are different answers, + // and only the first is one a draft may act on: reading them as one value + // is how a disk something is driving gets proposed as free. Read it through + // Held and Free, which are the two questions a reader has and neither of + // which is the negation of the other. + InUse *bool `json:"inUse,omitempty"` } +// Held reports whether something is known to hold the controller open. +func (c Controller) Held() bool { return c.InUse != nil && *c.InUse } + +// Free reports whether the probe checked the controller and found nothing +// holding it, which is the only state a draft may claim it in. +func (c Controller) Free() bool { return c.InUse != nil && !*c.InUse } + // BoundToUserspace reports whether a userspace-IO driver owns the controller. // // It reads the driver and says nothing about who bound it or whether anything From 79140e157053d6b366e8ee1195966ef625b162ff Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Fri, 18 Sep 2026 16:51:49 +0200 Subject: [PATCH 072/206] feat(webhook): a discovery run whose size range cannot be read is refused An unreadable driveSizeRange built no size rule at all, so the filter widened to every disk on every worker rather than narrowing to none, and the draft a reviewer approved was the fleet's whole storage instead of the slice they asked for. Nothing downstream could catch it: the refusal has no device to attach itself to, so it reached neither the refusal list nor the run's explanation, and the document was indistinguishable from one a run meant to write. BasicDeviceRules already assumed this webhook existed -- "ParseSizeRange is called by the caller that validates the spec, and the run reports the parse failure separately" -- and no such caller did. OperatorOps was also the one Ops kind in the group carrying no validator at all. The refusal names the field, quotes the parse failure, says what would otherwise happen, and shows the forms a range takes, because a reviewer whose filter was refused needs the form and not only the verdict. Update is guarded as well as create: the filter is not immutable, and a run edited before Probing still decides what the draft holds. failurePolicy=Fail, on the same grounds as the other Ops validators: the webhook server runs inside the operator pod, so its availability tracks the operator's, and a run admitted while the operator is down is a run nothing would perform. The controller asks the same question twice more, for a cluster whose webhook is not installed or was bypassed -- at Inspecting, so a fleet is not probed for a draft that cannot be right, and at Writing, so the step applying the filter does not depend on the other having asked. Co-Authored-By: Claude Opus 5 (1M context) --- .../simplyblock-operator-webhook.yaml | 20 +++ operator/cmd/main.go | 4 + operator/config/webhook/manifests.yaml | 20 +++ .../internal/webhook/operatorops_validator.go | 90 +++++++++++++ .../webhook/operatorops_validator_test.go | 123 ++++++++++++++++++ 5 files changed, 257 insertions(+) create mode 100644 operator/internal/webhook/operatorops_validator.go create mode 100644 operator/internal/webhook/operatorops_validator_test.go diff --git a/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml b/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml index 7e11e8b20..bf194bd4c 100644 --- a/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml @@ -126,6 +126,26 @@ webhooks: resources: - controlplaneops sideEffects: None +- admissionReviewVersions: + - v1 + clientConfig: + service: + name: simplyblock-operator-webhook-service + namespace: {{ .Release.Namespace }} + path: /validate-storage-simplyblock-io-v1alpha2-operatorops + failurePolicy: Fail + name: voperatorops.simplyblock.io + rules: + - apiGroups: + - storage.simplyblock.io + apiVersions: + - v1alpha2 + operations: + - CREATE + - UPDATE + resources: + - operatorops + sideEffects: None - admissionReviewVersions: - v1 clientConfig: diff --git a/operator/cmd/main.go b/operator/cmd/main.go index e7c27cb23..afc9285ad 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -966,6 +966,10 @@ func main() { &webhook.Admission{Handler: &internalwebhook.StorageClusterOpsValidator{}}) setupLog.Info("registered storageclusterops validating webhook") + mgr.GetWebhookServer().Register("/validate-storage-simplyblock-io-v1alpha2-operatorops", + &webhook.Admission{Handler: &internalwebhook.OperatorOpsValidator{}}) + setupLog.Info("registered operatorops validating webhook") + mgr.GetWebhookServer().Register("/validate-storage-simplyblock-io-v1alpha2-storagenodeops", &webhook.Admission{Handler: &internalwebhook.StorageNodeOpsValidator{}}) setupLog.Info("registered storagenodeops validating webhook") diff --git a/operator/config/webhook/manifests.yaml b/operator/config/webhook/manifests.yaml index 087fcd824..7f94881fc 100644 --- a/operator/config/webhook/manifests.yaml +++ b/operator/config/webhook/manifests.yaml @@ -108,6 +108,26 @@ webhooks: resources: - controlplaneops sideEffects: None +- admissionReviewVersions: + - v1 + clientConfig: + service: + name: webhook-service + namespace: system + path: /validate-storage-simplyblock-io-v1alpha2-operatorops + failurePolicy: Fail + name: voperatorops.simplyblock.io + rules: + - apiGroups: + - storage.simplyblock.io + apiVersions: + - v1alpha2 + operations: + - CREATE + - UPDATE + resources: + - operatorops + sideEffects: None - admissionReviewVersions: - v1 clientConfig: diff --git a/operator/internal/webhook/operatorops_validator.go b/operator/internal/webhook/operatorops_validator.go new file mode 100644 index 000000000..54d339fae --- /dev/null +++ b/operator/internal/webhook/operatorops_validator.go @@ -0,0 +1,90 @@ +// The OperatorOps guard: a validating webhook that refuses a discovery run +// whose device filter names a size range nothing can read. +// +// The filter is the one part of a run that is a rule rather than a value, and +// the size range is the one part of the filter that has to be parsed. A range +// the parser cannot read is not a range that admits nothing: the pipeline +// builds no size rule at all, so the filter silently widens to every disk on +// every worker, and the draft a reviewer approves is the fleet's whole storage +// rather than the slice they asked for. +// +// Nothing downstream can catch it. The refusal has no device to attach itself +// to, so it reaches neither the refusal list nor the run's explanation, and the +// draft that results is indistinguishable from one written by a run that meant +// to take everything. +// +// Admission is therefore where it belongs, and it is where the parser's own +// caller already assumed it was: the comment in BasicDeviceRules says the spec +// is validated by the caller and the failure is reported separately, which was +// true of nothing until this existed. + +package webhook + +import ( + "context" + "encoding/json" + "fmt" + "net/http" + + admissionv1 "k8s.io/api/admission/v1" + "sigs.k8s.io/controller-runtime/pkg/webhook/admission" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/discovery" +) + +// +kubebuilder:webhook:path=/validate-storage-simplyblock-io-v1alpha2-operatorops,mutating=false,failurePolicy=fail,sideEffects=None,groups=storage.simplyblock.io,resources=operatorops,verbs=create;update,versions=v1alpha2,name=voperatorops.simplyblock.io,admissionReviewVersions=v1 + +// OperatorOpsValidator refuses a run whose filter cannot be read. +// +// failurePolicy=Fail, because the webhook server runs inside the operator pod: +// its availability tracks the operator's, and a run admitted while the operator +// is down is a run nothing would perform anyway. +// +// It holds no client. The only thing the decision rests on is the object being +// admitted, which is what makes the answer the same on every replica. +type OperatorOpsValidator struct{} + +// Handle refuses a create or an update whose device filter does not parse. +// +// Update is guarded as well as create, because the filter is not immutable and +// a run edited before it reaches Probing is a run whose filter still decides +// what the draft holds. +func (v *OperatorOpsValidator) Handle(_ context.Context, req admission.Request) admission.Response { + if req.Operation != admissionv1.Create && req.Operation != admissionv1.Update { + return admission.Allowed("") + } + if len(req.Object.Raw) == 0 { + return admission.Allowed("") + } + + var ops simplyblockv1alpha2.OperatorOps + if err := json.Unmarshal(req.Object.Raw, &ops); err != nil { + return admission.Errored(http.StatusBadRequest, err) + } + + if refusal := unreadableSizeRange(ops.Spec.Discover); refusal != "" { + return admission.Denied(refusal) + } + return admission.Allowed("") +} + +// unreadableSizeRange is the refusal for a size range that does not parse, and +// the empty string for a run that names none or names one that does. +func unreadableSizeRange(spec *simplyblockv1alpha2.DiscoverSpec) string { + if spec == nil || spec.DeviceFilter == nil || spec.DeviceFilter.DriveSizeRange == "" { + return "" + } + + rangeSpec := spec.DeviceFilter.DriveSizeRange + if _, _, err := discovery.ParseSizeRange(rangeSpec); err != nil { + return fmt.Sprintf( + "spec.discover.deviceFilter.driveSizeRange is %q, which cannot be read: %v. "+ + "A run whose range cannot be read applies no size filter at all, so the draft "+ + "would name every disk on every worker rather than the ones asked for. "+ + "Write a range as 100G-2T, as 500G- for no upper bound, as -2T for no lower "+ + "one, or as a single size such as 1T, which means exactly that size.", + rangeSpec, err) + } + return "" +} diff --git a/operator/internal/webhook/operatorops_validator_test.go b/operator/internal/webhook/operatorops_validator_test.go new file mode 100644 index 000000000..c0e967fb7 --- /dev/null +++ b/operator/internal/webhook/operatorops_validator_test.go @@ -0,0 +1,123 @@ +// That a discovery run naming a size range nothing can read is refused at the +// request rather than accepted and quietly widened. + +package webhook + +import ( + "context" + "encoding/json" + "strings" + "testing" + + admissionv1 "k8s.io/api/admission/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/runtime" + "sigs.k8s.io/controller-runtime/pkg/webhook/admission" + + "github.com/simplyblock/atlas/ptr" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// discoverWith is a run carrying the device filter given. +func discoverWith(filter *simplyblockv1alpha2.DeviceFilter) *simplyblockv1alpha2.OperatorOps { + return &simplyblockv1alpha2.OperatorOps{ + ObjectMeta: metav1.ObjectMeta{Name: "oops-1", Namespace: "simplyblock"}, + Spec: simplyblockv1alpha2.OperatorOpsSpec{ + Action: simplyblockv1alpha2.OperatorOpsActionDiscover, + Discover: &simplyblockv1alpha2.DiscoverSpec{DeviceFilter: filter}, + }, + } +} + +// reviewOf is the admission request for creating the run. +func reviewOf(t *testing.T, ops *simplyblockv1alpha2.OperatorOps) admission.Request { + t.Helper() + raw, err := json.Marshal(ops) + if err != nil { + t.Fatalf("render the run: %v", err) + } + return admission.Request{AdmissionRequest: admissionv1.AdmissionRequest{ + Operation: admissionv1.Create, + Object: runtime.RawExtension{Raw: raw}, + }} +} + +func TestADriveSizeRangeNothingCanReadIsRefused(t *testing.T) { + validator := &OperatorOpsValidator{} + + for _, spec := range []string{"2T-1T", "abc", "1X", "1T-2T-3T", "0100G", ""} { + if spec == "" { + continue + } + run := discoverWith(&simplyblockv1alpha2.DeviceFilter{DriveSizeRange: spec}) + response := validator.Handle(context.Background(), reviewOf(t, run)) + + if response.Allowed { + t.Errorf("a run whose driveSizeRange is %q was admitted", spec) + continue + } + if !strings.Contains(response.Result.Message, "driveSizeRange") { + t.Errorf("the refusal of %q does not name the field: %s", spec, response.Result.Message) + } + } +} + +func TestADriveSizeRangeThatReadsIsAdmitted(t *testing.T) { + validator := &OperatorOpsValidator{} + + for _, spec := range []string{"100G-2T", "500G-", "-2T", "1T", "2048", "1TiB-4TiB"} { + run := discoverWith(&simplyblockv1alpha2.DeviceFilter{DriveSizeRange: spec}) + response := validator.Handle(context.Background(), reviewOf(t, run)) + + if !response.Allowed { + t.Errorf("a run whose driveSizeRange is %q was refused: %s", + spec, response.Result.Message) + } + } +} + +func TestARunWithNoRangeToReadIsAdmitted(t *testing.T) { + validator := &OperatorOpsValidator{} + + for _, run := range []*simplyblockv1alpha2.OperatorOps{ + discoverWith(nil), + discoverWith(&simplyblockv1alpha2.DeviceFilter{}), + discoverWith(&simplyblockv1alpha2.DeviceFilter{EnableLogicalBlockDevices: ptr.To(true)}), + { + ObjectMeta: metav1.ObjectMeta{Name: "oops-2", Namespace: "simplyblock"}, + Spec: simplyblockv1alpha2.OperatorOpsSpec{ + Action: simplyblockv1alpha2.OperatorOpsActionDiscover, + }, + }, + } { + if response := validator.Handle(context.Background(), reviewOf(t, run)); !response.Allowed { + t.Errorf("a run naming no size range was refused: %s", response.Result.Message) + } + } +} + +func TestTheRefusalSaysWhatARangeLooksLike(t *testing.T) { + // A reviewer whose filter was refused needs the form, not only the verdict. + validator := &OperatorOpsValidator{} + run := discoverWith(&simplyblockv1alpha2.DeviceFilter{DriveSizeRange: "2T-1T"}) + + response := validator.Handle(context.Background(), reviewOf(t, run)) + + for _, fragment := range []string{"100G-2T", "500G-", "-2T"} { + if !strings.Contains(response.Result.Message, fragment) { + t.Errorf("the refusal does not show the %s form: %s", fragment, response.Result.Message) + } + } +} + +func TestAnObjectThatDoesNotDecodeIsAnError(t *testing.T) { + validator := &OperatorOpsValidator{} + request := admission.Request{AdmissionRequest: admissionv1.AdmissionRequest{ + Operation: admissionv1.Create, + Object: runtime.RawExtension{Raw: []byte("{ not an object")}, + }} + + if response := validator.Handle(context.Background(), request); response.Allowed { + t.Error("an object that does not decode was admitted") + } +} From f17b2fb516fc46f3718431b2bc275bdca7ab0afc Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Fri, 18 Sep 2026 16:52:12 +0200 Subject: [PATCH 073/206] test(discovery): 179 fleets the generator is run over, and two harnesses The generator had unit tests over hand-built reports and nothing that drove a fleet through the whole path. What a reviewer approves is a document, and the parts of it the planner has no hand in -- the name, the environment, the cluster block, whether the run refused outright -- were the parts nothing covered. Each case is a directory: the OperatorOps, the node objects, and one ConfigMap per worker holding the report a probe would have written. They are generated rather than hand-typed, because the objects are built through the same types the probe and the operator use, so a fixture that would not decode never gets written. The case list is read from the document that enumerates them, so a row with no fixture and a fixture with no row are both errors that stop the run -- which is what keeps 179 of them honest. Two harnesses, because they answer different questions. The first drives every case through the writing step and compares against a recorded document, refusal list, note list, and failure message. A recording is change detection and nothing more: it was produced by running the code, so it asserts what the code already did. The file says so, because a golden nobody read is a record of behavior rather than of intent, and reading each one against its row is a step no harness can take. The second is the part that is not circular. It checks properties against things written for other purposes: every drafted document is applied to a real apiserver with this repository's own CRDs, run through the ClusterDeploymentConfig controller's own draft validation, and checked against the reports it was built from, so a worker or a device the draft names that no probe reported is a failure. Two findings are allow-listed by name rather than skipped, so the list is the specification and anything else is a generator that has started producing something new. The generator keeps recordings across a regeneration. It used to clear the tree, which meant regenerating deleted every expectation and the next recording silently agreed with whatever the new inputs produced; it now prunes only directories no case claims and only the files it owns, so a changed input shows up as a failing diff. Co-Authored-By: Claude Opus 5 (1M context) --- operator/hack/discoveryfixtures/cases_dev.go | 255 +++ operator/hack/discoveryfixtures/cases_held.go | 327 ++++ operator/hack/discoveryfixtures/cases_net.go | 323 ++++ operator/hack/discoveryfixtures/cases_numa.go | 127 ++ operator/hack/discoveryfixtures/cases_pci.go | 265 +++ operator/hack/discoveryfixtures/cases_role.go | 241 +++ operator/hack/discoveryfixtures/cases_size.go | 128 ++ operator/hack/discoveryfixtures/cases_tmpl.go | 101 + operator/hack/discoveryfixtures/document.go | 95 + operator/hack/discoveryfixtures/fixtures.go | 647 +++++++ operator/hack/discoveryfixtures/main.go | 243 +++ .../deployment/discovery_cases_test.go | 379 ++++ .../deployment/discovery_invariants_test.go | 319 ++++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 33 + .../nodes.yaml | 89 + .../ops.yaml | 14 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 33 + .../nodes.yaml | 89 + .../ops.yaml | 14 + .../reports/worker-01.yaml | 9 + .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 33 + .../nodes.yaml | 89 + .../ops.yaml | 14 + .../reports/worker-01.yaml | 11 + .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ .../cm/cm-04-a-report-naming-no-node/case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 33 + .../cm-04-a-report-naming-no-node/nodes.yaml | 89 + .../cm/cm-04-a-report-naming-no-node/ops.yaml | 14 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 33 + .../nodes.yaml | 89 + .../ops.yaml | 14 + .../reports/worker-01.yaml | 174 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../ops.yaml | 12 + ...l.a-very-long-suffix-nobody-shortened.yaml | 175 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 156 ++ .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 1663 +++++++++++++++++ .../dev/dev-01-four-nvme-disks/case.md | 7 + .../dev-01-four-nvme-disks/expected-notes.txt | 5 + .../dev/dev-01-four-nvme-disks/expected.yaml | 32 + .../dev/dev-01-four-nvme-disks/nodes.yaml | 29 + .../dev/dev-01-four-nvme-disks/ops.yaml | 12 + .../reports/worker-01.yaml | 175 ++ .../dev/dev-02-four-block-disks/case.md | 7 + .../expected-notes.txt | 5 + .../dev/dev-02-four-block-disks/expected.yaml | 32 + .../dev/dev-02-four-block-disks/nodes.yaml | 29 + .../dev/dev-02-four-block-disks/ops.yaml | 15 + .../reports/worker-01.yaml | 167 ++ .../dev-03-block-disks-on-an-nvme-run/case.md | 7 + .../expected-error.txt | 1 + .../expected-refusals.txt | 5 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 167 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 2 + .../expected.yaml | 30 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 171 ++ .../case.md | 9 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 2 + .../expected.yaml | 30 + .../nodes.yaml | 29 + .../ops.yaml | 15 + .../reports/worker-01.yaml | 171 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 30 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 163 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 1 + .../expected.yaml | 29 + .../nodes.yaml | 29 + .../ops.yaml | 15 + .../reports/worker-01.yaml | 149 ++ .../case.md | 9 + .../expected-notes.txt | 5 + .../expected.yaml | 30 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 152 ++ .../dev/dev-09-ten-nvme-disks/case.md | 7 + .../dev-09-ten-nvme-disks/expected-notes.txt | 5 + .../dev/dev-09-ten-nvme-disks/expected.yaml | 38 + .../dev/dev-09-ten-nvme-disks/nodes.yaml | 29 + .../dev/dev-09-ten-nvme-disks/ops.yaml | 12 + .../reports/worker-01.yaml | 247 +++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 3 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 205 ++ .../case.md | 9 + .../expected-notes.txt | 5 + .../expected.yaml | 156 ++ .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 1663 +++++++++++++++++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 30 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 151 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 1 + .../expected.yaml | 29 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 155 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 1 + .../expected.yaml | 29 + .../nodes.yaml | 29 + .../ops.yaml | 15 + .../reports/worker-01.yaml | 153 ++ .../case.md | 7 + .../expected-error.txt | 1 + .../expected-refusals.txt | 4 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 175 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 1 + .../expected.yaml | 29 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 155 ++ .../dev-17-an-iscsi-lun-nobody-named/case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 1 + .../expected.yaml | 29 + .../nodes.yaml | 29 + .../dev-17-an-iscsi-lun-nobody-named/ops.yaml | 15 + .../reports/worker-01.yaml | 147 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 30 + .../nodes.yaml | 29 + .../ops.yaml | 18 + .../reports/worker-01.yaml | 147 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 1 + .../expected.yaml | 29 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 149 ++ .../case.md | 7 + .../expected-error.txt | 1 + .../expected-refusals.txt | 3 + .../nodes.yaml | 89 + .../ops.yaml | 14 + .../reports/worker-01.yaml | 125 ++ .../reports/worker-02.yaml | 125 ++ .../reports/worker-03.yaml | 125 ++ .../case.md | 7 + .../expected-error.txt | 1 + .../expected-refusals.txt | 15 + .../nodes.yaml | 89 + .../ops.yaml | 14 + .../reports/worker-01.yaml | 195 ++ .../reports/worker-02.yaml | 195 ++ .../reports/worker-03.yaml | 195 ++ .../fail-03-every-report-unreadable/case.md | 7 + .../expected-error.txt | 1 + .../nodes.yaml | 89 + .../fail-03-every-report-unreadable/ops.yaml | 14 + .../reports/worker-01.yaml | 11 + .../reports/worker-02.yaml | 11 + .../reports/worker-03.yaml | 11 + .../case.md | 7 + .../expected-error.txt | 1 + .../expected-refusals.txt | 15 + .../nodes.yaml | 89 + .../ops.yaml | 17 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ .../case.md | 9 + .../expected-notes.txt | 5 + .../expected.yaml | 33 + .../nodes.yaml | 80 + .../ops.yaml | 14 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ .../case.md | 9 + .../expected-notes.txt | 5 + .../expected.yaml | 34 + .../nodes.yaml | 89 + .../ops.yaml | 14 + .../reports/worker-01.yaml | 140 ++ .../reports/worker-02.yaml | 140 ++ .../reports/worker-03.yaml | 140 ++ .../case.md | 9 + .../expected-notes.txt | 5 + .../expected.yaml | 34 + .../nodes.yaml | 89 + .../ops.yaml | 14 + .../reports/worker-01.yaml | 146 ++ .../reports/worker-02.yaml | 146 ++ .../reports/worker-03.yaml | 146 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 138 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 5 + .../expected.yaml | 62 + .../nodes.yaml | 959 ++++++++++ .../ops.yaml | 43 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ .../reports/worker-04.yaml | 175 ++ .../reports/worker-05.yaml | 175 ++ .../reports/worker-06.yaml | 175 ++ .../reports/worker-07.yaml | 175 ++ .../reports/worker-08.yaml | 175 ++ .../reports/worker-09.yaml | 175 ++ .../reports/worker-10.yaml | 175 ++ .../reports/worker-11.yaml | 175 ++ .../reports/worker-12.yaml | 175 ++ .../reports/worker-13.yaml | 175 ++ .../reports/worker-14.yaml | 175 ++ .../reports/worker-15.yaml | 175 ++ .../reports/worker-16.yaml | 175 ++ .../reports/worker-17.yaml | 195 ++ .../reports/worker-18.yaml | 175 ++ .../reports/worker-19.yaml | 175 ++ .../reports/worker-20.yaml | 175 ++ .../reports/worker-21.yaml | 175 ++ .../reports/worker-22.yaml | 175 ++ .../reports/worker-23.yaml | 175 ++ .../reports/worker-24.yaml | 175 ++ .../reports/worker-25.yaml | 175 ++ .../reports/worker-26.yaml | 175 ++ .../reports/worker-27.yaml | 175 ++ .../reports/worker-28.yaml | 175 ++ .../reports/worker-29.yaml | 175 ++ .../reports/worker-30.yaml | 175 ++ .../reports/worker-31.yaml | 175 ++ .../reports/worker-32.yaml | 175 ++ .../case.md | 7 + .../expected-error.txt | 1 + .../expected-refusals.txt | 3 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 151 ++ .../filt/filt-01-a-pci-deny-list/case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 3 + .../filt-01-a-pci-deny-list/expected.yaml | 30 + .../filt/filt-01-a-pci-deny-list/nodes.yaml | 29 + .../filt/filt-01-a-pci-deny-list/ops.yaml | 16 + .../reports/worker-01.yaml | 201 ++ .../filt/filt-02-a-pci-allow-list/case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 3 + .../filt-02-a-pci-allow-list/expected.yaml | 30 + .../filt/filt-02-a-pci-allow-list/nodes.yaml | 29 + .../filt/filt-02-a-pci-allow-list/ops.yaml | 17 + .../reports/worker-01.yaml | 201 ++ .../filt/filt-03-a-model-substring/case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 3 + .../filt-03-a-model-substring/expected.yaml | 30 + .../filt/filt-03-a-model-substring/nodes.yaml | 29 + .../filt/filt-03-a-model-substring/ops.yaml | 15 + .../reports/worker-01.yaml | 201 ++ .../filt/filt-04-a-size-range/case.md | 7 + .../filt-04-a-size-range/expected-notes.txt | 5 + .../expected-refusals.txt | 3 + .../filt/filt-04-a-size-range/expected.yaml | 30 + .../filt/filt-04-a-size-range/nodes.yaml | 29 + .../filt/filt-04-a-size-range/ops.yaml | 15 + .../reports/worker-01.yaml | 201 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 1 + .../expected.yaml | 30 + .../nodes.yaml | 29 + .../ops.yaml | 15 + .../reports/worker-01.yaml | 163 ++ .../case.md | 9 + .../expected-error.txt | 1 + .../nodes.yaml | 29 + .../ops.yaml | 15 + .../reports/worker-01.yaml | 201 ++ .../filt/filt-07-a-block-deny-list/case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 1 + .../filt-07-a-block-deny-list/expected.yaml | 33 + .../filt/filt-07-a-block-deny-list/nodes.yaml | 29 + .../filt/filt-07-a-block-deny-list/ops.yaml | 17 + .../reports/worker-01.yaml | 187 ++ .../filt/filt-08-a-block-allow-list/case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 4 + .../filt-08-a-block-allow-list/expected.yaml | 30 + .../filt-08-a-block-allow-list/nodes.yaml | 29 + .../filt/filt-08-a-block-allow-list/ops.yaml | 18 + .../reports/worker-01.yaml | 187 ++ .../filt-09-a-partition-table-waived/case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 1 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../filt-09-a-partition-table-waived/ops.yaml | 15 + .../reports/worker-01.yaml | 201 ++ .../case.md | 7 + .../expected-error.txt | 1 + .../expected-refusals.txt | 6 + .../nodes.yaml | 29 + .../ops.yaml | 16 + .../reports/worker-01.yaml | 201 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 3 + .../expected.yaml | 30 + .../nodes.yaml | 29 + .../ops.yaml | 17 + .../reports/worker-01.yaml | 201 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 4 + .../expected.yaml | 29 + .../nodes.yaml | 29 + .../ops.yaml | 19 + .../reports/worker-01.yaml | 201 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 3 + .../expected.yaml | 31 + .../nodes.yaml | 29 + .../ops.yaml | 16 + .../reports/worker-01.yaml | 187 ++ .../filt/filt-14-no-filter-at-all/case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 2 + .../filt-14-no-filter-at-all/expected.yaml | 31 + .../filt/filt-14-no-filter-at-all/nodes.yaml | 29 + .../filt/filt-14-no-filter-at-all/ops.yaml | 12 + .../reports/worker-01.yaml | 201 ++ .../fleet/fleet-01-one-worker/case.md | 7 + .../fleet-01-one-worker/expected-notes.txt | 5 + .../fleet/fleet-01-one-worker/expected.yaml | 32 + .../fleet/fleet-01-one-worker/nodes.yaml | 29 + .../fleet/fleet-01-one-worker/ops.yaml | 12 + .../reports/worker-01.yaml | 175 ++ .../fleet-02-three-uniform-workers/case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 34 + .../fleet-02-three-uniform-workers/nodes.yaml | 89 + .../fleet-02-three-uniform-workers/ops.yaml | 14 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 63 + .../nodes.yaml | 959 ++++++++++ .../ops.yaml | 43 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ .../reports/worker-04.yaml | 175 ++ .../reports/worker-05.yaml | 175 ++ .../reports/worker-06.yaml | 175 ++ .../reports/worker-07.yaml | 175 ++ .../reports/worker-08.yaml | 175 ++ .../reports/worker-09.yaml | 175 ++ .../reports/worker-10.yaml | 175 ++ .../reports/worker-11.yaml | 175 ++ .../reports/worker-12.yaml | 175 ++ .../reports/worker-13.yaml | 175 ++ .../reports/worker-14.yaml | 175 ++ .../reports/worker-15.yaml | 175 ++ .../reports/worker-16.yaml | 175 ++ .../reports/worker-17.yaml | 175 ++ .../reports/worker-18.yaml | 175 ++ .../reports/worker-19.yaml | 175 ++ .../reports/worker-20.yaml | 175 ++ .../reports/worker-21.yaml | 175 ++ .../reports/worker-22.yaml | 175 ++ .../reports/worker-23.yaml | 175 ++ .../reports/worker-24.yaml | 175 ++ .../reports/worker-25.yaml | 175 ++ .../reports/worker-26.yaml | 175 ++ .../reports/worker-27.yaml | 175 ++ .../reports/worker-28.yaml | 175 ++ .../reports/worker-29.yaml | 175 ++ .../reports/worker-30.yaml | 175 ++ .../reports/worker-31.yaml | 175 ++ .../reports/worker-32.yaml | 175 ++ .../fleet/fleet-04-no-reports-at-all/case.md | 7 + .../expected-error.txt | 1 + .../fleet-04-no-reports-at-all/nodes.yaml | 89 + .../fleet/fleet-04-no-reports-at-all/ops.yaml | 14 + .../case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 6 + .../expected.yaml | 61 + .../nodes.yaml | 959 ++++++++++ .../ops.yaml | 43 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ .../reports/worker-04.yaml | 175 ++ .../reports/worker-05.yaml | 175 ++ .../reports/worker-06.yaml | 175 ++ .../reports/worker-07.yaml | 175 ++ .../reports/worker-08.yaml | 175 ++ .../reports/worker-09.yaml | 175 ++ .../reports/worker-10.yaml | 175 ++ .../reports/worker-11.yaml | 161 ++ .../reports/worker-12.yaml | 175 ++ .../reports/worker-13.yaml | 175 ++ .../reports/worker-14.yaml | 175 ++ .../reports/worker-15.yaml | 175 ++ .../reports/worker-16.yaml | 175 ++ .../reports/worker-17.yaml | 175 ++ .../reports/worker-18.yaml | 175 ++ .../reports/worker-19.yaml | 175 ++ .../reports/worker-20.yaml | 175 ++ .../reports/worker-21.yaml | 175 ++ .../reports/worker-22.yaml | 161 ++ .../reports/worker-23.yaml | 175 ++ .../reports/worker-24.yaml | 175 ++ .../reports/worker-25.yaml | 175 ++ .../reports/worker-26.yaml | 175 ++ .../reports/worker-27.yaml | 175 ++ .../reports/worker-28.yaml | 175 ++ .../reports/worker-29.yaml | 175 ++ .../reports/worker-30.yaml | 175 ++ .../reports/worker-31.yaml | 175 ++ .../reports/worker-32.yaml | 175 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 63 + .../nodes.yaml | 959 ++++++++++ .../ops.yaml | 43 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ .../reports/worker-04.yaml | 175 ++ .../reports/worker-05.yaml | 175 ++ .../reports/worker-06.yaml | 175 ++ .../reports/worker-07.yaml | 175 ++ .../reports/worker-08.yaml | 175 ++ .../reports/worker-09.yaml | 175 ++ .../reports/worker-10.yaml | 175 ++ .../reports/worker-11.yaml | 175 ++ .../reports/worker-12.yaml | 175 ++ .../reports/worker-13.yaml | 175 ++ .../reports/worker-14.yaml | 175 ++ .../reports/worker-15.yaml | 175 ++ .../reports/worker-16.yaml | 175 ++ .../reports/worker-17.yaml | 175 ++ .../reports/worker-18.yaml | 175 ++ .../reports/worker-19.yaml | 175 ++ .../reports/worker-20.yaml | 175 ++ .../reports/worker-21.yaml | 175 ++ .../reports/worker-22.yaml | 175 ++ .../reports/worker-23.yaml | 175 ++ .../reports/worker-24.yaml | 175 ++ .../reports/worker-25.yaml | 175 ++ .../reports/worker-26.yaml | 175 ++ .../reports/worker-27.yaml | 175 ++ .../reports/worker-28.yaml | 175 ++ .../reports/worker-29.yaml | 175 ++ .../reports/worker-30.yaml | 175 ++ .../reports/worker-31.yaml | 175 ++ .../reports/worker-32.yaml | 175 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 33 + .../nodes.yaml | 59 + .../ops.yaml | 13 + .../reports/report-2.yaml | 175 ++ .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 33 + .../nodes.yaml | 89 + .../ops.yaml | 13 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ .../held/held-01-every-disk-mounted/case.md | 7 + .../expected-error.txt | 1 + .../expected-refusals.txt | 5 + .../held-01-every-disk-mounted/nodes.yaml | 29 + .../held/held-01-every-disk-mounted/ops.yaml | 12 + .../reports/worker-01.yaml | 195 ++ .../case.md | 7 + .../expected-error.txt | 1 + .../expected-refusals.txt | 5 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 195 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 159 ++ .../case.md | 7 + .../expected-error.txt | 1 + .../expected-refusals.txt | 1 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 159 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 185 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 159 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 209 +++ .../case.md | 7 + .../expected-error.txt | 1 + .../expected-refusals.txt | 1 + .../nodes.yaml | 29 + .../ops.yaml | 15 + .../reports/worker-01.yaml | 159 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 16 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 319 ++++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 30 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 159 ++ .../case.md | 7 + .../expected-error.txt | 1 + .../expected-refusals.txt | 1 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 144 ++ .../net-01-one-addressed-physical-nic/case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 163 ++ .../net/net-02-the-faster-of-two-nics/case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../net-02-the-faster-of-two-nics/nodes.yaml | 26 + .../net-02-the-faster-of-two-nics/ops.yaml | 12 + .../reports/worker-01.yaml | 175 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 175 ++ .../net-04-every-interface-unusable/case.md | 9 + .../expected-notes.txt | 5 + .../expected.yaml | 31 + .../nodes.yaml | 26 + .../net-04-every-interface-unusable/ops.yaml | 12 + .../reports/worker-01.yaml | 201 ++ .../net/net-05-no-interfaces-reported/case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 31 + .../net-05-no-interfaces-reported/nodes.yaml | 26 + .../net-05-no-interfaces-reported/ops.yaml | 12 + .../reports/worker-01.yaml | 149 ++ .../case.md | 9 + .../expected-notes.txt | 5 + .../expected.yaml | 42 + .../nodes.yaml | 53 + .../ops.yaml | 13 + .../reports/worker-01.yaml | 163 ++ .../reports/worker-02.yaml | 163 ++ .../net/net-07-two-equal-nics/case.md | 7 + .../net-07-two-equal-nics/expected-notes.txt | 5 + .../net/net-07-two-equal-nics/expected.yaml | 32 + .../net/net-07-two-equal-nics/nodes.yaml | 26 + .../net/net-07-two-equal-nics/ops.yaml | 12 + .../reports/worker-01.yaml | 175 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 26 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 174 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 179 ++ .../net-10-a-global-ipv6-address-only/case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 26 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 163 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 26 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 164 ++ .../net-12-loopback-and-nothing-else/case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 31 + .../nodes.yaml | 26 + .../net-12-loopback-and-nothing-else/ops.yaml | 12 + .../reports/worker-01.yaml | 163 ++ .../net-13-the-node-address-on-a-bond/case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 218 +++ .../net-14-the-node-address-on-a-vlan/case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 178 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 174 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 218 +++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 195 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 295 +++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 31 + .../nodes.yaml | 26 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 196 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 31 + .../nodes.yaml | 26 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 175 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 26 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 163 ++ .../case.md | 9 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 26 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 175 ++ .../case.md | 9 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 187 ++ .../net-24-a-fast-link-and-a-slow-one/case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 26 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 175 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 26 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 175 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 26 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 192 ++ .../net-27-a-link-up-with-no-partner/case.md | 9 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 26 + .../net-27-a-link-up-with-no-partner/ops.yaml | 12 + .../reports/worker-01.yaml | 163 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 26 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 231 +++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 26 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 231 +++ .../net-30-a-bond-across-two-sockets/case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../net-30-a-bond-across-two-sockets/ops.yaml | 12 + .../reports/worker-01.yaml | 218 +++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 31 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 174 ++ .../case.md | 9 + .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 31 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 163 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 221 +++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 26 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 195 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 165 ++ .../net/net-37-a-team-interface/case.md | 7 + .../expected-notes.txt | 5 + .../net/net-37-a-team-interface/expected.yaml | 32 + .../net/net-37-a-team-interface/nodes.yaml | 29 + .../net/net-37-a-team-interface/ops.yaml | 12 + .../reports/worker-01.yaml | 192 ++ .../net/net-38-a-bridge-over-a-bond/case.md | 7 + .../expected-notes.txt | 5 + .../net-38-a-bridge-over-a-bond/expected.yaml | 32 + .../net-38-a-bridge-over-a-bond/nodes.yaml | 29 + .../net/net-38-a-bridge-over-a-bond/ops.yaml | 12 + .../reports/worker-01.yaml | 207 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 178 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 26 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 183 ++ .../case.md | 9 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 218 +++ .../case.md | 9 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 218 +++ .../numa/numa-01-one-memory-node/case.md | 7 + .../expected-notes.txt | 5 + .../numa-01-one-memory-node/expected.yaml | 32 + .../numa/numa-01-one-memory-node/nodes.yaml | 29 + .../numa/numa-01-one-memory-node/ops.yaml | 12 + .../reports/worker-01.yaml | 175 ++ .../numa-02-two-nodes-evenly-split/case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 2 + .../expected.yaml | 30 + .../numa-02-two-nodes-evenly-split/nodes.yaml | 29 + .../numa-02-two-nodes-evenly-split/ops.yaml | 12 + .../reports/worker-01.yaml | 181 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 1 + .../expected.yaml | 31 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 181 ++ .../numa/numa-04-count-beats-capacity/case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 31 + .../numa-04-count-beats-capacity/nodes.yaml | 29 + .../numa-04-count-beats-capacity/ops.yaml | 12 + .../reports/worker-01.yaml | 193 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 30 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 181 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 2 + .../expected.yaml | 30 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 213 +++ .../numa/numa-07-four-memory-nodes/case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 6 + .../numa-07-four-memory-nodes/expected.yaml | 30 + .../numa/numa-07-four-memory-nodes/nodes.yaml | 29 + .../numa/numa-07-four-memory-nodes/ops.yaml | 12 + .../reports/worker-01.yaml | 337 ++++ .../numa/numa-08-eight-memory-nodes/case.md | 7 + .../expected-notes.txt | 4 + .../expected-refusals.txt | 7 + .../numa-08-eight-memory-nodes/expected.yaml | 30 + .../numa-08-eight-memory-nodes/nodes.yaml | 29 + .../numa/numa-08-eight-memory-nodes/ops.yaml | 12 + .../reports/worker-01.yaml | 385 ++++ .../numa-09-every-device-on-no-node/case.md | 7 + .../expected-notes.txt | 4 + .../expected.yaml | 31 + .../nodes.yaml | 29 + .../numa-09-every-device-on-no-node/ops.yaml | 12 + .../reports/worker-01.yaml | 181 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 2 + .../expected.yaml | 30 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 181 ++ .../numa-11-all-devices-placement/case.md | 9 + .../case.md | 9 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 1 + .../expected.yaml | 37 + .../nodes.yaml | 59 + .../ops.yaml | 13 + .../reports/worker-01.yaml | 151 ++ .../reports/worker-02.yaml | 157 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 38 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 253 +++ .../numa-14-a-large-two-socket-worker/case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 5 + .../expected.yaml | 33 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 417 +++++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 2 + .../expected.yaml | 30 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 135 ++ .../pci/pci-01-a-fleet-that-agrees/case.md | 7 + .../expected-notes.txt | 5 + .../pci-01-a-fleet-that-agrees/expected.yaml | 63 + .../pci/pci-01-a-fleet-that-agrees/nodes.yaml | 959 ++++++++++ .../pci/pci-01-a-fleet-that-agrees/ops.yaml | 43 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ .../reports/worker-04.yaml | 175 ++ .../reports/worker-05.yaml | 175 ++ .../reports/worker-06.yaml | 175 ++ .../reports/worker-07.yaml | 175 ++ .../reports/worker-08.yaml | 175 ++ .../reports/worker-09.yaml | 175 ++ .../reports/worker-10.yaml | 175 ++ .../reports/worker-11.yaml | 175 ++ .../reports/worker-12.yaml | 175 ++ .../reports/worker-13.yaml | 175 ++ .../reports/worker-14.yaml | 175 ++ .../reports/worker-15.yaml | 175 ++ .../reports/worker-16.yaml | 175 ++ .../reports/worker-17.yaml | 175 ++ .../reports/worker-18.yaml | 175 ++ .../reports/worker-19.yaml | 175 ++ .../reports/worker-20.yaml | 175 ++ .../reports/worker-21.yaml | 175 ++ .../reports/worker-22.yaml | 175 ++ .../reports/worker-23.yaml | 175 ++ .../reports/worker-24.yaml | 175 ++ .../reports/worker-25.yaml | 175 ++ .../reports/worker-26.yaml | 175 ++ .../reports/worker-27.yaml | 175 ++ .../reports/worker-28.yaml | 175 ++ .../reports/worker-29.yaml | 175 ++ .../reports/worker-30.yaml | 175 ++ .../reports/worker-31.yaml | 175 ++ .../reports/worker-32.yaml | 175 ++ .../pci-02-two-workers-that-do-not/case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 42 + .../pci-02-two-workers-that-do-not/nodes.yaml | 59 + .../pci-02-two-workers-that-do-not/ops.yaml | 13 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../pci/pci-03-sixteen-and-sixteen/case.md | 7 + .../expected-notes.txt | 5 + .../pci-03-sixteen-and-sixteen/expected.yaml | 72 + .../pci/pci-03-sixteen-and-sixteen/nodes.yaml | 959 ++++++++++ .../pci/pci-03-sixteen-and-sixteen/ops.yaml | 43 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ .../reports/worker-04.yaml | 175 ++ .../reports/worker-05.yaml | 175 ++ .../reports/worker-06.yaml | 175 ++ .../reports/worker-07.yaml | 175 ++ .../reports/worker-08.yaml | 175 ++ .../reports/worker-09.yaml | 175 ++ .../reports/worker-10.yaml | 175 ++ .../reports/worker-11.yaml | 175 ++ .../reports/worker-12.yaml | 175 ++ .../reports/worker-13.yaml | 175 ++ .../reports/worker-14.yaml | 175 ++ .../reports/worker-15.yaml | 175 ++ .../reports/worker-16.yaml | 175 ++ .../reports/worker-17.yaml | 175 ++ .../reports/worker-18.yaml | 175 ++ .../reports/worker-19.yaml | 175 ++ .../reports/worker-20.yaml | 175 ++ .../reports/worker-21.yaml | 175 ++ .../reports/worker-22.yaml | 175 ++ .../reports/worker-23.yaml | 175 ++ .../reports/worker-24.yaml | 175 ++ .../reports/worker-25.yaml | 175 ++ .../reports/worker-26.yaml | 175 ++ .../reports/worker-27.yaml | 175 ++ .../reports/worker-28.yaml | 175 ++ .../reports/worker-29.yaml | 175 ++ .../reports/worker-30.yaml | 175 ++ .../reports/worker-31.yaml | 175 ++ .../reports/worker-32.yaml | 175 ++ .../pci/pci-04-every-worker-distinct/case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 278 +++ .../pci-04-every-worker-distinct/nodes.yaml | 959 ++++++++++ .../pci/pci-04-every-worker-distinct/ops.yaml | 43 + .../reports/worker-01.yaml | 151 ++ .../reports/worker-02.yaml | 151 ++ .../reports/worker-03.yaml | 151 ++ .../reports/worker-04.yaml | 151 ++ .../reports/worker-05.yaml | 151 ++ .../reports/worker-06.yaml | 151 ++ .../reports/worker-07.yaml | 151 ++ .../reports/worker-08.yaml | 151 ++ .../reports/worker-09.yaml | 151 ++ .../reports/worker-10.yaml | 151 ++ .../reports/worker-11.yaml | 151 ++ .../reports/worker-12.yaml | 151 ++ .../reports/worker-13.yaml | 151 ++ .../reports/worker-14.yaml | 151 ++ .../reports/worker-15.yaml | 151 ++ .../reports/worker-16.yaml | 151 ++ .../reports/worker-17.yaml | 151 ++ .../reports/worker-18.yaml | 151 ++ .../reports/worker-19.yaml | 151 ++ .../reports/worker-20.yaml | 151 ++ .../reports/worker-21.yaml | 151 ++ .../reports/worker-22.yaml | 151 ++ .../reports/worker-23.yaml | 151 ++ .../reports/worker-24.yaml | 151 ++ .../reports/worker-25.yaml | 151 ++ .../reports/worker-26.yaml | 151 ++ .../reports/worker-27.yaml | 151 ++ .../reports/worker-28.yaml | 151 ++ .../reports/worker-29.yaml | 151 ++ .../reports/worker-30.yaml | 151 ++ .../reports/worker-31.yaml | 151 ++ .../reports/worker-32.yaml | 151 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 175 ++ .../pci-06-ten-disks-across-two-buses/case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 38 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 247 +++ .../pci-07-a-five-digit-pci-domain/case.md | 9 + .../expected-notes.txt | 5 + .../expected.yaml | 30 + .../pci-07-a-five-digit-pci-domain/nodes.yaml | 29 + .../pci-07-a-five-digit-pci-domain/ops.yaml | 12 + .../reports/worker-01.yaml | 151 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 38 + .../nodes.yaml | 59 + .../ops.yaml | 13 + .../reports/worker-01.yaml | 151 ++ .../reports/worker-02.yaml | 151 ++ .../pci/pci-09-unpadded-worker-names/case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 72 + .../pci-09-unpadded-worker-names/nodes.yaml | 959 ++++++++++ .../pci/pci-09-unpadded-worker-names/ops.yaml | 43 + .../reports/worker-1.yaml | 175 ++ .../reports/worker-10.yaml | 175 ++ .../reports/worker-11.yaml | 175 ++ .../reports/worker-12.yaml | 175 ++ .../reports/worker-13.yaml | 175 ++ .../reports/worker-14.yaml | 175 ++ .../reports/worker-15.yaml | 175 ++ .../reports/worker-16.yaml | 175 ++ .../reports/worker-17.yaml | 175 ++ .../reports/worker-18.yaml | 175 ++ .../reports/worker-19.yaml | 175 ++ .../reports/worker-2.yaml | 175 ++ .../reports/worker-20.yaml | 175 ++ .../reports/worker-21.yaml | 175 ++ .../reports/worker-22.yaml | 175 ++ .../reports/worker-23.yaml | 175 ++ .../reports/worker-24.yaml | 175 ++ .../reports/worker-25.yaml | 175 ++ .../reports/worker-26.yaml | 175 ++ .../reports/worker-27.yaml | 175 ++ .../reports/worker-28.yaml | 175 ++ .../reports/worker-29.yaml | 175 ++ .../reports/worker-3.yaml | 175 ++ .../reports/worker-30.yaml | 175 ++ .../reports/worker-31.yaml | 175 ++ .../reports/worker-32.yaml | 175 ++ .../reports/worker-4.yaml | 175 ++ .../reports/worker-5.yaml | 175 ++ .../reports/worker-6.yaml | 175 ++ .../reports/worker-7.yaml | 175 ++ .../reports/worker-8.yaml | 175 ++ .../reports/worker-9.yaml | 175 ++ .../pci-10-three-agree-and-two-do-not/case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 50 + .../nodes.yaml | 149 ++ .../ops.yaml | 16 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ .../reports/worker-04.yaml | 151 ++ .../reports/worker-05.yaml | 151 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 114 ++ .../nodes.yaml | 959 ++++++++++ .../ops.yaml | 43 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ .../reports/worker-04.yaml | 175 ++ .../reports/worker-05.yaml | 175 ++ .../reports/worker-06.yaml | 175 ++ .../reports/worker-07.yaml | 175 ++ .../reports/worker-08.yaml | 175 ++ .../reports/worker-09.yaml | 175 ++ .../reports/worker-10.yaml | 175 ++ .../reports/worker-11.yaml | 175 ++ .../reports/worker-12.yaml | 175 ++ .../reports/worker-13.yaml | 175 ++ .../reports/worker-14.yaml | 175 ++ .../reports/worker-15.yaml | 175 ++ .../reports/worker-16.yaml | 175 ++ .../reports/worker-17.yaml | 175 ++ .../reports/worker-18.yaml | 175 ++ .../reports/worker-19.yaml | 175 ++ .../reports/worker-20.yaml | 175 ++ .../reports/worker-21.yaml | 175 ++ .../reports/worker-22.yaml | 175 ++ .../reports/worker-23.yaml | 175 ++ .../reports/worker-24.yaml | 175 ++ .../reports/worker-25.yaml | 175 ++ .../reports/worker-26.yaml | 175 ++ .../reports/worker-27.yaml | 151 ++ .../reports/worker-28.yaml | 151 ++ .../reports/worker-29.yaml | 151 ++ .../reports/worker-30.yaml | 151 ++ .../reports/worker-31.yaml | 151 ++ .../reports/worker-32.yaml | 151 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 48 + .../nodes.yaml | 239 +++ .../ops.yaml | 19 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ .../reports/worker-04.yaml | 175 ++ .../reports/worker-05.yaml | 175 ++ .../reports/worker-06.yaml | 175 ++ .../reports/worker-07.yaml | 175 ++ .../reports/worker-08.yaml | 175 ++ .../pci-13-one-layout-two-capacities/case.md | 9 + .../expected-notes.txt | 5 + .../expected.yaml | 31 + .../nodes.yaml | 59 + .../pci-13-one-layout-two-capacities/ops.yaml | 13 + .../reports/worker-01.yaml | 151 ++ .../reports/worker-02.yaml | 151 ++ .../pci-14-a-worker-holding-a-subset/case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 42 + .../nodes.yaml | 89 + .../pci-14-a-worker-holding-a-subset/ops.yaml | 14 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 163 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 34 + .../nodes.yaml | 89 + .../ops.yaml | 14 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 48 + .../nodes.yaml | 185 ++ .../ops.yaml | 17 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ .../reports/worker-04.yaml | 175 ++ .../reports/worker-05.yaml | 175 ++ .../reports/worker-06.yaml | 175 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 45 + .../nodes.yaml | 94 + .../ops.yaml | 16 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 32 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 175 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 32 + .../ops.yaml | 14 + .../reports/worker-01.yaml | 175 ++ .../role/role-06-the-master-spelling/case.md | 7 + .../expected-notes.txt | 5 + .../role-06-the-master-spelling/expected.yaml | 32 + .../role-06-the-master-spelling/nodes.yaml | 31 + .../role/role-06-the-master-spelling/ops.yaml | 14 + .../reports/worker-01.yaml | 175 ++ .../case.md | 9 + .../case.md | 9 + .../expected-notes.txt | 5 + .../expected.yaml | 33 + .../nodes.yaml | 64 + .../ops.yaml | 13 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 31 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 175 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 85 + .../nodes.yaml | 1023 ++++++++++ .../ops.yaml | 45 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ .../reports/worker-04.yaml | 175 ++ .../reports/worker-05.yaml | 175 ++ .../reports/worker-06.yaml | 175 ++ .../reports/worker-07.yaml | 175 ++ .../reports/worker-08.yaml | 175 ++ .../reports/worker-09.yaml | 175 ++ .../reports/worker-10.yaml | 175 ++ .../reports/worker-11.yaml | 175 ++ .../reports/worker-12.yaml | 175 ++ .../reports/worker-13.yaml | 175 ++ .../reports/worker-14.yaml | 175 ++ .../reports/worker-15.yaml | 175 ++ .../reports/worker-16.yaml | 175 ++ .../reports/worker-17.yaml | 175 ++ .../reports/worker-18.yaml | 175 ++ .../reports/worker-19.yaml | 175 ++ .../reports/worker-20.yaml | 175 ++ .../reports/worker-21.yaml | 175 ++ .../reports/worker-22.yaml | 175 ++ .../reports/worker-23.yaml | 175 ++ .../reports/worker-24.yaml | 175 ++ .../reports/worker-25.yaml | 175 ++ .../reports/worker-26.yaml | 175 ++ .../reports/worker-27.yaml | 175 ++ .../reports/worker-28.yaml | 175 ++ .../reports/worker-29.yaml | 175 ++ .../reports/worker-30.yaml | 175 ++ .../reports/worker-31.yaml | 175 ++ .../reports/worker-32.yaml | 175 ++ .../size/size-01-eight-logical-cpus/case.md | 7 + .../expected-notes.txt | 5 + .../size-01-eight-logical-cpus/expected.yaml | 32 + .../size-01-eight-logical-cpus/nodes.yaml | 29 + .../size/size-01-eight-logical-cpus/ops.yaml | 12 + .../reports/worker-01.yaml | 151 ++ .../size/size-02-two-physical-cores/case.md | 7 + .../expected-notes.txt | 5 + .../size-02-two-physical-cores/expected.yaml | 32 + .../size-02-two-physical-cores/nodes.yaml | 29 + .../size/size-02-two-physical-cores/ops.yaml | 12 + .../reports/worker-01.yaml | 140 ++ .../size/size-03-one-socket-196-vcpus/case.md | 9 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../size-03-one-socket-196-vcpus/nodes.yaml | 29 + .../size-03-one-socket-196-vcpus/ops.yaml | 12 + .../reports/worker-01.yaml | 329 ++++ .../size-04-two-sockets-196-vcpus/case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 2 + .../expected.yaml | 30 + .../size-04-two-sockets-196-vcpus/nodes.yaml | 29 + .../size-04-two-sockets-196-vcpus/ops.yaml | 12 + .../reports/worker-01.yaml | 345 ++++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 34 + .../nodes.yaml | 89 + .../ops.yaml | 14 + .../reports/worker-01.yaml | 151 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 241 +++ .../size-06-four-gibibytes-of-memory/case.md | 9 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../size-06-four-gibibytes-of-memory/ops.yaml | 12 + .../reports/worker-01.yaml | 146 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 166 ++ .../case.md | 9 + .../size-09-pages-already-set-aside/case.md | 9 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 2 + .../expected.yaml | 30 + .../nodes.yaml | 29 + .../size-09-pages-already-set-aside/ops.yaml | 12 + .../reports/worker-01.yaml | 182 ++ .../case.md | 9 + .../expected-notes.txt | 4 + .../expected-refusals.txt | 6 + .../expected.yaml | 31 + .../nodes.yaml | 89 + .../ops.yaml | 14 + .../reports/worker-01.yaml | 181 ++ .../reports/worker-02.yaml | 162 ++ .../reports/worker-03.yaml | 181 ++ .../case.md | 7 + .../expected-notes.txt | 4 + .../expected.yaml | 31 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 170 ++ .../case.md | 7 + .../expected-notes.txt | 4 + .../expected.yaml | 31 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 163 ++ .../case.md | 7 + .../expected-notes.txt | 4 + .../expected-refusals.txt | 2 + .../expected.yaml | 29 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 181 ++ .../size/size-14-swap-in-use/case.md | 9 + .../size-14-swap-in-use/expected-notes.txt | 5 + .../size/size-14-swap-in-use/expected.yaml | 32 + .../size/size-14-swap-in-use/nodes.yaml | 29 + .../size/size-14-swap-in-use/ops.yaml | 12 + .../reports/worker-01.yaml | 176 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 31 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 170 ++ .../case.md | 9 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 166 ++ .../case.md | 9 + .../expected-notes.txt | 4 + .../expected.yaml | 31 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 151 ++ .../case.md | 9 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 171 ++ .../case.md | 9 + .../expected-notes.txt | 4 + .../expected.yaml | 31 + .../nodes.yaml | 29 + .../ops.yaml | 12 + .../reports/worker-01.yaml | 171 ++ .../size-20-two-page-sizes-at-once/case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 32 + .../size-20-two-page-sizes-at-once/nodes.yaml | 29 + .../size-20-two-page-sizes-at-once/ops.yaml | 12 + .../reports/worker-01.yaml | 182 ++ .../size/size-21-no-hugetlbfs-at-all/case.md | 9 + .../expected-notes.txt | 4 + .../size-21-no-hugetlbfs-at-all/expected.yaml | 31 + .../size-21-no-hugetlbfs-at-all/nodes.yaml | 29 + .../size/size-21-no-hugetlbfs-at-all/ops.yaml | 12 + .../reports/worker-01.yaml | 156 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 34 + .../nodes.yaml | 89 + .../ops.yaml | 14 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 34 + .../nodes.yaml | 89 + .../ops.yaml | 14 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ .../tmpl-03-a-generated-config-name/case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 34 + .../nodes.yaml | 89 + .../tmpl-03-a-generated-config-name/ops.yaml | 14 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ .../case.md | 9 + .../expected-notes.txt | 5 + .../expected.yaml | 34 + .../nodes.yaml | 89 + .../ops.yaml | 16 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ .../tmpl/tmpl-05-a-growth-document/case.md | 7 + .../expected-notes.txt | 2 + .../tmpl-05-a-growth-document/expected.yaml | 29 + .../tmpl/tmpl-05-a-growth-document/nodes.yaml | 89 + .../tmpl/tmpl-05-a-growth-document/ops.yaml | 16 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ .../case.md | 9 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 6 + .../expected.yaml | 32 + .../nodes.yaml | 89 + .../ops.yaml | 14 + .../reports/worker-01.yaml | 181 ++ .../reports/worker-02.yaml | 181 ++ .../reports/worker-03.yaml | 181 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected.yaml | 34 + .../nodes.yaml | 89 + .../ops.yaml | 14 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ .../case.md | 7 + .../expected-notes.txt | 5 + .../expected-refusals.txt | 3 + .../expected.yaml | 33 + .../nodes.yaml | 89 + .../ops.yaml | 19 + .../reports/worker-01.yaml | 175 ++ .../reports/worker-02.yaml | 175 ++ .../reports/worker-03.yaml | 175 ++ 1490 files changed, 131981 insertions(+) create mode 100644 operator/hack/discoveryfixtures/cases_dev.go create mode 100644 operator/hack/discoveryfixtures/cases_held.go create mode 100644 operator/hack/discoveryfixtures/cases_net.go create mode 100644 operator/hack/discoveryfixtures/cases_numa.go create mode 100644 operator/hack/discoveryfixtures/cases_pci.go create mode 100644 operator/hack/discoveryfixtures/cases_role.go create mode 100644 operator/hack/discoveryfixtures/cases_size.go create mode 100644 operator/hack/discoveryfixtures/cases_tmpl.go create mode 100644 operator/hack/discoveryfixtures/document.go create mode 100644 operator/hack/discoveryfixtures/fixtures.go create mode 100644 operator/hack/discoveryfixtures/main.go create mode 100644 operator/internal/controllers/deployment/discovery_cases_test.go create mode 100644 operator/internal/controllers/deployment/discovery_invariants_test.go create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/reports/ip-10-0-1-23.eu-central-1.compute.internal.a-very-long-suffix-nobody-shortened.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-03-block-disks-on-an-nvme-run/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-03-block-disks-on-an-nvme-run/expected-error.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-03-block-disks-on-an-nvme-run/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-03-block-disks-on-an-nvme-run/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-03-block-disks-on-an-nvme-run/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-03-block-disks-on-an-nvme-run/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-15-every-disk-is-an-attached-volume/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-15-every-disk-is-an-attached-volume/expected-error.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-15-every-disk-is-an-attached-volume/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-15-every-disk-is-an-attached-volume/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-15-every-disk-is-an-attached-volume/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-15-every-disk-is-an-attached-volume/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/expected-error.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/expected-error.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-03-every-report-unreadable/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-03-every-report-unreadable/expected-error.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-03-every-report-unreadable/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-03-every-report-unreadable/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-03-every-report-unreadable/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-03-every-report-unreadable/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-03-every-report-unreadable/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/expected-error.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-04.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-05.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-06.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-07.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-08.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-09.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-10.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-11.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-12.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-13.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-14.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-15.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-16.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-17.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-18.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-19.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-20.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-21.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-22.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-23.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-24.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-25.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-26.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-27.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-28.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-29.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-30.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-31.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-32.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-10-a-worker-whose-disks-are-partitions/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-10-a-worker-whose-disks-are-partitions/expected-error.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-10-a-worker-whose-disks-are-partitions/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-10-a-worker-whose-disks-are-partitions/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-10-a-worker-whose-disks-are-partitions/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fail/fail-10-a-worker-whose-disks-are-partitions/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-06-a-size-range-that-counts-backward/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-06-a-size-range-that-counts-backward/expected-error.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-06-a-size-range-that-counts-backward/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-06-a-size-range-that-counts-backward/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-06-a-size-range-that-counts-backward/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-10-a-filter-that-matches-nothing/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-10-a-filter-that-matches-nothing/expected-error.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-10-a-filter-that-matches-nothing/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-10-a-filter-that-matches-nothing/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-10-a-filter-that-matches-nothing/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-10-a-filter-that-matches-nothing/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-04.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-05.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-06.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-07.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-08.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-09.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-10.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-11.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-12.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-13.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-14.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-15.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-16.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-17.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-18.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-19.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-20.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-21.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-22.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-23.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-24.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-25.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-26.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-27.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-28.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-29.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-30.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-31.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-32.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-04-no-reports-at-all/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-04-no-reports-at-all/expected-error.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-04-no-reports-at-all/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-04-no-reports-at-all/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-04.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-05.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-06.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-07.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-08.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-09.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-10.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-11.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-12.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-13.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-14.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-15.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-16.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-17.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-18.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-19.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-20.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-21.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-22.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-23.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-24.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-25.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-26.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-27.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-28.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-29.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-30.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-31.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-32.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-04.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-05.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-06.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-07.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-08.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-09.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-10.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-11.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-12.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-13.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-14.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-15.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-16.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-17.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-18.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-19.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-20.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-21.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-22.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-23.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-24.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-25.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-26.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-27.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-28.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-29.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-30.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-31.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-32.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/reports/report-2.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-01-every-disk-mounted/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-01-every-disk-mounted/expected-error.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-01-every-disk-mounted/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-01-every-disk-mounted/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-01-every-disk-mounted/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-01-every-disk-mounted/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-02-every-disk-in-a-device-mapper-stack/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-02-every-disk-in-a-device-mapper-stack/expected-error.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-02-every-disk-in-a-device-mapper-stack/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-02-every-disk-in-a-device-mapper-stack/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-02-every-disk-in-a-device-mapper-stack/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-02-every-disk-in-a-device-mapper-stack/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-04-four-userspace-controllers-in-use/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-04-four-userspace-controllers-in-use/expected-error.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-04-four-userspace-controllers-in-use/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-04-four-userspace-controllers-in-use/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-04-four-userspace-controllers-in-use/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-04-four-userspace-controllers-in-use/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-08-idle-controllers-on-a-block-run/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-08-idle-controllers-on-a-block-run/expected-error.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-08-idle-controllers-on-a-block-run/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-08-idle-controllers-on-a-block-run/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-08-idle-controllers-on-a-block-run/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-08-idle-controllers-on-a-block-run/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-11-controllers-nothing-could-check/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-11-controllers-nothing-could-check/expected-error.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-11-controllers-nothing-could-check/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-11-controllers-nothing-could-check/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-11-controllers-nothing-could-check/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/held/held-11-controllers-nothing-could-check/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-32-a-stack-that-points-at-itself/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-11-all-devices-placement/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-04.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-05.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-06.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-07.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-08.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-09.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-10.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-11.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-12.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-13.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-14.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-15.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-16.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-17.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-18.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-19.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-20.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-21.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-22.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-23.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-24.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-25.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-26.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-27.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-28.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-29.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-30.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-31.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-32.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-04.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-05.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-06.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-07.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-08.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-09.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-10.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-11.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-12.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-13.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-14.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-15.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-16.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-17.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-18.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-19.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-20.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-21.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-22.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-23.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-24.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-25.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-26.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-27.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-28.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-29.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-30.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-31.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-32.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-04.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-05.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-06.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-07.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-08.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-09.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-10.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-11.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-12.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-13.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-14.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-15.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-16.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-17.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-18.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-19.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-20.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-21.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-22.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-23.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-24.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-25.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-26.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-27.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-28.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-29.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-30.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-31.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-32.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-1.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-10.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-11.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-12.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-13.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-14.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-15.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-16.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-17.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-18.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-19.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-2.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-20.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-21.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-22.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-23.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-24.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-25.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-26.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-27.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-28.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-29.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-3.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-30.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-31.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-32.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-4.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-5.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-6.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-7.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-8.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-9.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/reports/worker-04.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/reports/worker-05.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-04.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-05.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-06.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-07.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-08.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-09.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-10.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-11.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-12.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-13.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-14.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-15.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-16.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-17.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-18.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-19.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-20.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-21.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-22.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-23.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-24.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-25.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-26.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-27.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-28.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-29.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-30.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-31.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-32.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-04.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-05.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-06.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-07.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-08.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/reports/worker-04.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/reports/worker-05.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/reports/worker-06.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-07-a-plan-with-no-node-objects/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-04.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-05.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-06.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-07.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-08.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-09.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-10.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-11.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-12.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-13.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-14.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-15.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-16.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-17.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-18.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-19.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-20.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-21.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-22.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-23.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-24.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-25.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-26.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-27.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-28.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-29.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-30.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-31.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-32.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-08-the-fully-readable-worker-rule/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/reports/worker-03.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/case.md create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/expected-notes.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/expected-refusals.txt create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/expected.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/nodes.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/ops.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/reports/worker-01.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/reports/worker-02.yaml create mode 100644 operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/reports/worker-03.yaml diff --git a/operator/hack/discoveryfixtures/cases_dev.go b/operator/hack/discoveryfixtures/cases_dev.go new file mode 100644 index 000000000..0fc770694 --- /dev/null +++ b/operator/hack/discoveryfixtures/cases_dev.go @@ -0,0 +1,255 @@ +// The device class and inventory shape cases, §2 of the document. +// +// Each is a worker whose disks differ in one way from the four free NVMe disks +// a draft is ordinarily built from: another class, another bus, another kind of +// entry in the block layer, or a count at the edge of what the schema holds. + +package main + +import ( + "fmt" + + "github.com/simplyblock/atlas/blockdev" + "github.com/simplyblock/atlas/nqn" + "github.com/simplyblock/atlas/ptr" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" +) + +// blockClass is the filter that makes a run scan logical block devices. +func blockClass() *simplyblockv1alpha2.DiscoverSpec { + return &simplyblockv1alpha2.DiscoverSpec{ + DeviceFilter: &simplyblockv1alpha2.DeviceFilter{ + EnableLogicalBlockDevices: ptr.To(true), + }, + } +} + +// fourNVMe is the disk set a case inherits when it is about something else: two +// disks on each memory node, all the same size. +func fourNVMe() []nodeprobe.Device { + return []nodeprobe.Device{ + nvme("nvme0n1", "0000:5e:00.0", 0, 3*tb), + nvme("nvme1n1", "0000:5f:00.0", 0, 3*tb), + nvme("nvme2n1", "0000:af:00.0", 1, 3*tb), + nvme("nvme3n1", "0000:b0:00.0", 1, 3*tb), + } +} + +// oneNodeFourNVMe is the same four disks with every one of them on memory node +// 0, for a case that is about the disks rather than about the placement. +func oneNodeFourNVMe() []nodeprobe.Device { + return []nodeprobe.Device{ + nvme("nvme0n1", "0000:5e:00.0", 0, 3*tb), + nvme("nvme1n1", "0000:5f:00.0", 0, 3*tb), + nvme("nvme2n1", "0000:60:00.0", 0, 3*tb), + nvme("nvme3n1", "0000:61:00.0", 0, 3*tb), + } +} + +// virtioDisks is n virtio disks, which have a path and no PCI address. +func virtioDisks(n int) []nodeprobe.Device { + out := make([]nodeprobe.Device, 0, n) + for i := 0; i < n; i++ { + out = append(out, blk(fmt.Sprintf("vd%c", 'b'+rune(i)), 0, 2*tb)) + } + return out +} + +// manyNVMe is n disks in consecutive slots on one memory node. +func manyNVMe(n int, size uint64) []nodeprobe.Device { + out := make([]nodeprobe.Device, 0, n) + for i := 0; i < n; i++ { + out = append(out, nvme( + fmt.Sprintf("nvme%dn1", i), fmt.Sprintf("0000:%02x:00.0", 0x5e+i), 0, size)) + } + return out +} + +func devCases() map[string]Case { + single := cpu(1, 16, 2) + + cases := map[string]Case{ + "DEV-01": { + Family: "dev", Slug: "four-nvme-disks", + Reports: []nodeprobe.Report{host("worker-01", single, disks(oneNodeFourNVMe()...))}, + }, + "DEV-02": { + Family: "dev", Slug: "four-block-disks", + Discover: blockClass(), + Reports: []nodeprobe.Report{host("worker-01", single, disks(virtioDisks(4)...))}, + }, + "DEV-03": { + Family: "dev", Slug: "block-disks-on-an-nvme-run", + Reports: []nodeprobe.Report{host("worker-01", single, disks(virtioDisks(4)...))}, + }, + "DEV-04": { + Family: "dev", Slug: "mixed-classes-on-an-nvme-run", + Reports: []nodeprobe.Report{host("worker-01", single, disks( + nvme("nvme0n1", "0000:5e:00.0", 0, 3*tb), + nvme("nvme1n1", "0000:5f:00.0", 0, 3*tb), + blk("vdb", 0, 2*tb), + blk("vdc", 0, 2*tb), + ))}, + }, + "DEV-05": { + Family: "dev", Slug: "mixed-classes-on-a-block-run", Gap: "G-1", + Discover: blockClass(), + Reports: []nodeprobe.Report{host("worker-01", single, disks( + nvme("nvme0n1", "0000:5e:00.0", 0, 3*tb), + nvme("nvme1n1", "0000:5f:00.0", 0, 3*tb), + blk("vdb", 0, 2*tb), + blk("vdc", 0, 2*tb), + ))}, + }, + "DEV-06": { + Family: "dev", Slug: "one-controller-two-namespaces", + Reports: []nodeprobe.Report{host("worker-01", single, disks( + nvme("nvme0n1", "0000:5e:00.0", 0, tb), + nvme("nvme0n2", "0000:5e:00.0", 0, tb), + nvme("nvme1n1", "0000:5f:00.0", 0, tb), + ))}, + }, + "DEV-07": { + Family: "dev", Slug: "a-sata-disk-beside-an-nvme-one", + Discover: blockClass(), + Reports: []nodeprobe.Report{host("worker-01", single, disks( + nvme("nvme0n1", "0000:5e:00.0", 0, 3*tb), + blk("sda", 0, 4*tb, transported(blockdev.TransportSATA)), + ))}, + }, + "DEV-08": { + Family: "dev", Slug: "a-spinning-disk-beside-an-ssd", Gap: "G-2", + Reports: []nodeprobe.Report{host("worker-01", single, disks( + nvme("nvme0n1", "0000:5e:00.0", 0, 3*tb), + nvme("nvme1n1", "0000:5f:00.0", 0, 16*tb, spinning()), + ))}, + }, + "DEV-09": { + Family: "dev", Slug: "ten-nvme-disks", + Reports: []nodeprobe.Report{host("worker-01", single, disks(manyNVMe(10, 3*tb)...))}, + }, + "DEV-10": { + Family: "dev", Slug: "partitions-and-loopbacks-beside-disks", + Reports: []nodeprobe.Report{host("worker-01", single, disks(append( + oneNodeFourNVMe(), + nvme("nvme0n1p1", "0000:5e:00.0", 0, 512*gb, partOf()), + blk("loop0", 0, 64*gb, looped()), + blk("loop1", 0, 64*gb, looped()), + )...))}, + }, + "DEV-11": { + Family: "dev", Slug: "a-worker-at-the-selection-ceiling", + Note: "128 is the MaxItems of a group's device selection, so this is the largest worker the schema can describe.", + Reports: []nodeprobe.Report{host("worker-01", single, disks(manyNVMe(128, 512*gb)...))}, + }, + "DEV-12": { + Family: "dev", Slug: "a-disk-that-reports-no-size", + Reports: []nodeprobe.Report{host("worker-01", single, disks( + nvme("nvme0n1", "0000:5e:00.0", 0, 0), + nvme("nvme1n1", "0000:5f:00.0", 0, 0), + ))}, + }, + "DEV-13": { + Family: "dev", Slug: "an-attached-simplyblock-volume", + Reports: []nodeprobe.Report{host("worker-01", single, disks( + attachedVolume("nvme3n1", ourCluster, "792e184c-0a1b-2c3d-4e5f-60718293a4b5"), + nvme("nvme0n1", "0000:5e:00.0", 0, 3*tb), + ))}, + }, + "DEV-14": { + Family: "dev", Slug: "an-attached-volume-on-a-block-run", + Discover: blockClass(), + Reports: []nodeprobe.Report{host("worker-01", single, disks( + attachedVolume("nvme3n1", ourCluster, "792e184c-0a1b-2c3d-4e5f-60718293a4b5"), + blk("vdb", 0, 2*tb), + ))}, + }, + "DEV-15": { + Family: "dev", Slug: "every-disk-is-an-attached-volume", + Reports: []nodeprobe.Report{host("worker-01", single, disks( + attachedVolume("nvme3n1", ourCluster, "792e184c-0a1b-2c3d-4e5f-60718293a4b5"), + attachedVolume("nvme4n1", ourCluster, "8a3f0b12-3c4d-5e6f-7081-92a3b4c5d6e7"), + attachedVolume("nvme5n1", ourCluster, "b1c2d3e4-f506-1728-394a-5b6c7d8e9f01"), + ))}, + }, + "DEV-17": { + Family: "dev", Slug: "an-iscsi-lun-nobody-named", + Discover: blockClass(), + Reports: []nodeprobe.Report{host("worker-01", single, disks( + iscsiLUN("sdb", 2*tb), + blk("vdb", 0, 2*tb), + ))}, + }, + "DEV-18": { + Family: "dev", Slug: "an-iscsi-lun-the-allow-list-names", + Discover: &simplyblockv1alpha2.DiscoverSpec{ + DeviceFilter: &simplyblockv1alpha2.DeviceFilter{ + EnableLogicalBlockDevices: ptr.To(true), + BlockAllowList: []string{"/dev/sdb", "/dev/vdb"}, + }, + }, + Reports: []nodeprobe.Report{host("worker-01", single, disks( + iscsiLUN("sdb", 2*tb), + blk("vdb", 0, 2*tb), + ))}, + }, + "DEV-19": { + Family: "dev", Slug: "an-iscsi-lun-on-an-nvme-run", + Reports: []nodeprobe.Report{host("worker-01", single, disks( + iscsiLUN("sdb", 2*tb), + nvme("nvme0n1", "0000:5e:00.0", 0, 3*tb), + ))}, + }, + "DEV-16": { + Family: "dev", Slug: "a-fabric-namespace-of-another-product", + Reports: []nodeprobe.Report{host("worker-01", single, disks( + foreignVolume("nvme3n1"), + nvme("nvme0n1", "0000:5e:00.0", 0, 3*tb), + ))}, + }, + } + + // Every case here is about disks, so none of them says anything about the + // node objects, and all of them need one: a document naming a worker the + // cluster does not have is a document its own validation rejects. + for id, c := range cases { + c.Nodes = kubeFleet(c.Reports) + cases[id] = c + } + return cases +} + +// ourCluster is the cluster the attached volumes below belong to, which is what +// a refusal names. +const ourCluster = "c30a691a-1d2e-4f3a-9b8c-5d6e7f809a1b" + +// attachedVolume is a simplyblock volume as the worker that connected it sees +// one: a namespace on a fabric, carrying the NQN this product builds. +// +// The probe refuses it for being on a fabric, which is what a real report +// carries, and the rule that names it as this fleet's own reads the NQN rather +// than the refusal. +func attachedVolume(name, clusterID, volumeID string) nodeprobe.Device { + device := blk(name, 0, tb, refused(blockdev.ReasonFabricNamespace)) + device.Transport = string(blockdev.TransportNVMeFabric) + device.SubsystemNQN = nqn.Make(clusterID, volumeID) + return device +} + +// iscsiLUN is a disk on the other side of a network, which the kernel presents +// through the SCSI stack like any local disk. +func iscsiLUN(name string, size uint64) nodeprobe.Device { + device := blk(name, 0, size) + device.Transport = string(blockdev.TransportISCSI) + return device +} + +// foreignVolume is a namespace something else exported, which is on a fabric +// and is nobody's simplyblock volume. +func foreignVolume(name string) nodeprobe.Device { + device := blk(name, 0, tb, refused(blockdev.ReasonFabricNamespace)) + device.Transport = string(blockdev.TransportNVMeFabric) + device.SubsystemNQN = "nqn.2019-08.org.ceph:rbd.pool.image" + return device +} diff --git a/operator/hack/discoveryfixtures/cases_held.go b/operator/hack/discoveryfixtures/cases_held.go new file mode 100644 index 000000000..676bb0c4d --- /dev/null +++ b/operator/hack/discoveryfixtures/cases_held.go @@ -0,0 +1,327 @@ +// The cases for devices the probe declined and controllers a userspace driver +// holds, §10 of the document, the refusal cases of §11, the report-validity +// cases of §12, and the document-shape cases of §13. +// +// The four sections share a shape: each is a fleet the run either cannot build +// a draft from or can only partly, and what is being checked is the account the +// run gives of itself rather than the document it writes. + +package main + +import ( + "encoding/json" + "fmt" + "strings" + + corev1 "k8s.io/api/core/v1" + + "github.com/simplyblock/atlas/blockdev" + "github.com/simplyblock/atlas/inventory" + "github.com/simplyblock/atlas/pci" + "github.com/simplyblock/atlas/ptr" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" +) + +// idleControllers is n NVMe controllers a userspace driver holds and nothing is +// driving, which is what a machine that has run this product before presents. +func idleControllers(n int, driver string) []nodeprobe.Controller { + out := make([]nodeprobe.Controller, 0, n) + for i := 0; i < n; i++ { + out = append(out, controller(fmt.Sprintf("0000:%02x:00.0", 0x5e+i), driver, 0)) + } + return out +} + +func heldCases() map[string]Case { + one := func(report nodeprobe.Report) Case { + return Case{Family: "held", Reports: []nodeprobe.Report{report}, + Nodes: kubeFleet([]nodeprobe.Report{report})} + } + allMounted := func() []nodeprobe.Device { + out := layoutA.disks(3 * tb) + for i := range out { + refused(blockdev.ReasonMounted)(&out[i]) + } + return out + } + allHeld := func() []nodeprobe.Device { + out := layoutA.disks(3 * tb) + for i := range out { + refused(blockdev.ReasonStacked)(&out[i]) + } + return out + } + + busy := idleControllers(4, pci.DriverUIOGeneric) + for i := range busy { + busy[i].InUse = ptr.To(true) + } + halfBusy := idleControllers(4, pci.DriverUIOGeneric) + halfBusy[2].InUse, halfBusy[3].InUse = ptr.To(true), ptr.To(true) + + // A controller the probe could not check: the answer is absent rather than + // false, so nothing may claim it. + unchecked := idleControllers(2, pci.DriverUIOGeneric) + for i := range unchecked { + unchecked[i].InUse = nil + } + + loopbacks := make([]nodeprobe.Device, 0, 16) + for i := 0; i < 16; i++ { + loopbacks = append(loopbacks, blk(fmt.Sprintf("loop%d", i), 0, 64*gb, looped())) + } + + cases := map[string]Case{ + "HELD-01": one(host("worker-01", cpu(1, 16, 2), disks(allMounted()...))), + "HELD-02": one(host("worker-01", cpu(1, 16, 2), disks(allHeld()...))), + "HELD-03": one(host("worker-01", cpu(1, 16, 2), + controllers(idleControllers(4, pci.DriverUIOGeneric)...))), + "HELD-04": one(host("worker-01", cpu(1, 16, 2), controllers(busy...))), + "HELD-05": one(host("worker-01", cpu(1, 16, 2), + disks( + nvme("nvme0n1", "0000:af:00.0", 0, 3*tb), + nvme("nvme1n1", "0000:b0:00.0", 0, 3*tb), + ), + controllers( + controller("0000:5e:00.0", pci.DriverUIOGeneric, 0), + controller("0000:5f:00.0", pci.DriverUIOGeneric, 0), + controller("0000:af:00.0", "nvme", 0), + controller("0000:b0:00.0", "nvme", 0), + ))), + "HELD-06": one(host("worker-01", cpu(1, 16, 2), + controllers(idleControllers(4, pci.DriverVFIO)...))), + "HELD-07": one(host("worker-01", cpu(1, 16, 2), + disks(layoutA.disks(3*tb)...), + controllers( + controller("0000:5e:00.0", "nvme", 0), + controller("0000:5f:00.0", "nvme", 0), + controller("0000:af:00.0", "nvme", 0), + controller("0000:b0:00.0", "nvme", 0), + ))), + "HELD-09": one(host("worker-01", cpu(1, 16, 2), + disks(append(loopbacks, layoutA.disks(3*tb)...)...))), + "HELD-10": one(host("worker-01", cpu(1, 16, 2), controllers(halfBusy...))), + "HELD-11": one(host("worker-01", cpu(1, 16, 2), controllers(unchecked...), + unreadable("read the process table: permission denied"))), + } + + // The same idle controllers on a run that scans logical block devices, + // which names a device by a path a controller has none of. + cases["HELD-08"] = Case{ + Family: "held", + Discover: blockClass(), + Reports: []nodeprobe.Report{host("worker-01", cpu(1, 16, 2), + controllers(idleControllers(4, pci.DriverUIOGeneric)...))}, + Nodes: []corev1.Node{kubeNode("worker-01", reachableAt(managementAddress("worker-01")))}, + } + + slugs := map[string]string{ + "HELD-01": "every-disk-mounted", + "HELD-02": "every-disk-in-a-device-mapper-stack", + "HELD-03": "four-idle-userspace-controllers", + "HELD-04": "four-userspace-controllers-in-use", + "HELD-05": "idle-controllers-beside-kernel-disks", + "HELD-06": "an-idle-controller-on-vfio-pci", + "HELD-07": "kernel-controllers-already-presented", + "HELD-08": "idle-controllers-on-a-block-run", + "HELD-09": "sixteen-loopbacks-beside-four-disks", + "HELD-10": "half-the-controllers-in-use", + "HELD-11": "controllers-nothing-could-check", + } + for id, slug := range slugs { + entry := cases[id] + entry.Slug = slug + cases[id] = entry + } + return cases +} + +func failCases() map[string]Case { + bare := func(n int, build func(name string) nodeprobe.Report) Case { + reports := fleet(n, func(_ int, name string) nodeprobe.Report { return build(name) }) + return Case{Family: "fail", Reports: reports, Nodes: kubeFleet(reports)} + } + + mounted := func(name string) nodeprobe.Report { + out := layoutA.disks(3 * tb) + for i := range out { + refused(blockdev.ReasonMounted)(&out[i]) + } + return host(name, cpu(1, 16, 2), disks(out...)) + } + partitions := func(name string) nodeprobe.Report { + return host(name, cpu(1, 16, 2), disks( + nvme("nvme0n1p1", "0000:5e:00.0", 0, 512*gb, partOf()), + nvme("nvme0n1p2", "0000:5e:00.0", 0, 512*gb, partOf()), + )) + } + + cases := map[string]Case{ + "FAIL-01": bare(3, func(name string) nodeprobe.Report { + return host(name, cpu(1, 16, 2)) + }), + "FAIL-02": bare(3, mounted), + "FAIL-06": bare(3, func(name string) nodeprobe.Report { + return host(name, cpu(1, 2, 1), mem(16*gb, 14*gb, 16*gb), + disks(layoutA.disks(3*tb)...)) + }), + "FAIL-07": bare(3, func(name string) nodeprobe.Report { + return host(name, cpu(1, 4, 2), mem(4*gb, 900*mb, 4*gb), + disks(layoutA.disks(3*tb)...)) + }), + "FAIL-10": bare(1, partitions), + } + + // Every report unreadable, which is a probe that ran and wrote nothing a + // reader can use. + broken := fleet(3, func(_ int, name string) nodeprobe.Report { + return host(name, cpu(1, 16, 2), disks(layoutA.disks(3*tb)...)) + }) + cases["FAIL-03"] = Case{ + Family: "fail", Reports: broken, Nodes: kubeFleet(broken), + Amend: func(_ string, maps []*corev1.ConfigMap) []*corev1.ConfigMap { + for _, cm := range maps { + cm.Data[nodeprobe.ReportKey] = "{ this is not a report" + } + return maps + }, + } + + // A filter that excludes every disk in the fleet. + excluded := fleet(3, func(_ int, name string) nodeprobe.Report { + return host(name, cpu(1, 16, 2), disks(layoutA.disks(3*tb)...)) + }) + cases["FAIL-04"] = Case{ + Family: "fail", Reports: excluded, Nodes: kubeFleet(excluded), + Discover: filtered(func(f *simplyblockv1alpha2.DeviceFilter) { + f.DriveSizeRange = "100T-200T" + }), + } + + // Disks everywhere and no interface anything reaches the machines on. + unreachable := fleet(3, func(_ int, name string) nodeprobe.Report { + return host(name, cpu(1, 16, 2), disks(layoutA.disks(3*tb)...), ifaces( + nic("cni0", inventory.LinkBridge, holding("10.42.2.1")), + nic("lo", inventory.LinkLoopback, holding("127.0.0.1")), + )) + }) + unreachableNodes := make([]corev1.Node, 0, len(unreachable)) + for _, report := range unreachable { + unreachableNodes = append(unreachableNodes, kubeNode(report.Node)) + } + cases["FAIL-05"] = Case{Family: "fail", Reports: unreachable, Nodes: unreachableNodes} + + // A worker whose processor tree the probe could not read, with everything + // else in order. + blind := host("worker-01", noCPUTopology(), disks(layoutA.disks(3*tb)...), + unreadable("read /sys/devices/system/cpu: permission denied")) + cases["FAIL-08"] = Case{Family: "fail", Reports: []nodeprobe.Report{blind}, + Nodes: kubeFleet([]nodeprobe.Report{blind})} + + // One worker of thirty-two refused, which is an event rather than a + // failure. + nearlyAll := fleet(32, func(index int, name string) nodeprobe.Report { + if index == 17 { + return mounted(name) + } + return host(name, cpu(1, 16, 2), disks(layoutA.disks(3*tb)...)) + }) + cases["FAIL-09"] = Case{Family: "fail", Reports: nearlyAll, Nodes: kubeFleet(nearlyAll)} + + slugs := map[string]string{ + "FAIL-01": "no-devices-and-no-controllers", + "FAIL-02": "every-disk-in-the-fleet-mounted", + "FAIL-03": "every-report-unreadable", + "FAIL-04": "a-filter-that-excludes-the-fleet", + "FAIL-05": "no-usable-interface-anywhere", + "FAIL-06": "every-worker-under-the-vcpu-floor", + "FAIL-07": "every-worker-at-four-gibibytes", + "FAIL-08": "an-unreadable-processor-tree", + "FAIL-09": "one-worker-of-thirty-two-refused", + "FAIL-10": "a-worker-whose-disks-are-partitions", + } + gaps := map[string]string{"FAIL-05": "G-7", "FAIL-06": "G-10", "FAIL-07": "G-4"} + for id, slug := range slugs { + entry := cases[id] + entry.Slug = slug + entry.Gap = gaps[id] + cases[id] = entry + } + return cases +} + +func configMapCases() map[string]Case { + three := fleet(3, func(_ int, name string) nodeprobe.Report { + return host(name, cpu(1, 16, 2), disks(layoutA.disks(3*tb)...)) + }) + withNodes := func(reports []nodeprobe.Report, amend func(string, []*corev1.ConfigMap) []*corev1.ConfigMap) Case { + return Case{Family: "cm", Reports: reports, Nodes: kubeFleet(reports), Amend: amend} + } + + // A report written by the probe of the previous schema, which named no + // interface kind and therefore described its bonds as virtual and nothing + // else. + previousVersion := func(_ string, maps []*corev1.ConfigMap) []*corev1.ConfigMap { + var report map[string]any + _ = json.Unmarshal([]byte(maps[0].Data[nodeprobe.ReportKey]), &report) + report["version"] = nodeprobe.ReportVersion - 1 + rewritten, _ := json.MarshalIndent(report, "", " ") + maps[0].Data[nodeprobe.ReportKey] = string(rewritten) + return maps + } + + cases := map[string]Case{ + "CM-01": withNodes(three, previousVersion), + "CM-02": withNodes(three, func(_ string, maps []*corev1.ConfigMap) []*corev1.ConfigMap { + delete(maps[0].Data, nodeprobe.ReportKey) + return maps + }), + "CM-03": withNodes(three, func(_ string, maps []*corev1.ConfigMap) []*corev1.ConfigMap { + maps[0].Data[nodeprobe.ReportKey] = `{"version": 4, "node": "worker-01",` + return maps + }), + "CM-04": withNodes(three, func(_ string, maps []*corev1.ConfigMap) []*corev1.ConfigMap { + maps[0].Data[nodeprobe.ReportKey] = strings.Replace( + maps[0].Data[nodeprobe.ReportKey], `"node": "worker-01"`, `"node": ""`, 1) + return maps + }), + "CM-05": withNodes(three, func(_ string, maps []*corev1.ConfigMap) []*corev1.ConfigMap { + delete(maps[0].Labels, nodeprobe.LabelRun) + return maps + }), + } + + // A worker whose name is longer than a label value may be, so that the + // label is truncated and the report inside it is not. + const long = "ip-10-0-1-23.eu-central-1.compute.internal.a-very-long-suffix-nobody-shortened" + longNamed := host(long, cpu(1, 16, 2), disks(layoutA.disks(3*tb)...)) + cases["CM-06"] = Case{ + Family: "cm", Reports: []nodeprobe.Report{longNamed}, + Nodes: []corev1.Node{kubeNode(long, reachableAt("10.10.10.1"))}, + } + + // A report of a worker with as many devices as the schema will hold, which + // is the largest object a probe writes. + large := host("worker-01", cpu(1, 16, 2), disks(manyNVMe(128, 512*gb)...)) + cases["CM-07"] = Case{ + Family: "cm", Reports: []nodeprobe.Report{large}, + Nodes: kubeFleet([]nodeprobe.Report{large}), + } + + slugs := map[string]string{ + "CM-01": "a-report-from-the-previous-schema", + "CM-02": "a-configmap-with-no-report-key", + "CM-03": "a-report-that-does-not-parse", + "CM-04": "a-report-naming-no-node", + "CM-05": "a-configmap-without-the-run-label", + "CM-06": "a-worker-name-longer-than-a-label", + "CM-07": "a-report-of-a-hundred-and-twenty-eight-devices", + } + for id, slug := range slugs { + entry := cases[id] + entry.Slug = slug + cases[id] = entry + } + return cases +} diff --git a/operator/hack/discoveryfixtures/cases_net.go b/operator/hack/discoveryfixtures/cases_net.go new file mode 100644 index 000000000..1378049fb --- /dev/null +++ b/operator/hack/discoveryfixtures/cases_net.go @@ -0,0 +1,323 @@ +// The network-interface cases, §7 of the document. +// +// The ladder that picks a management interface is five rungs, and each case +// here is a host where exactly one of them decides. The stacked kinds are the +// reason the section is the longest: a bond, a VLAN, a bridge, and a veth are +// all virtual devices, and which of them an address can be bound to differs for +// every one. +// +// Every worker carries the same four disks, because a worker with none is +// refused before its interfaces are read and the case would be about nothing. + +package main + +import ( + "fmt" + + corev1 "k8s.io/api/core/v1" + + "github.com/simplyblock/atlas/inventory" + "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" +) + +// reachedAt is the address the cluster reaches worker-01 on, which is what +// makes an interface the management one outright. +const reachedAt = "10.10.10.1" + +// netHost is a worker whose disks are ordinary and whose interfaces are the +// case. +func netHost(name string, list ...nodeprobe.Interface) nodeprobe.Report { + return host(name, cpu(1, 16, 2), disks(spread(3*tb, 4)...), ifaces(list...)) +} + +// known is the case where Kubernetes knows the address it reaches each worker +// on, which is the ordinary run. +func known(reports ...nodeprobe.Report) Case { + return Case{Family: "net", Reports: reports, Nodes: kubeFleet(reports)} +} + +// unknown is the case where no node object names an address, so the ranking +// decides rather than the address. +func unknown(reports ...nodeprobe.Report) Case { + nodes := make([]corev1.Node, 0, len(reports)) + for _, report := range reports { + nodes = append(nodes, kubeNode(report.Node)) + } + return Case{Family: "net", Reports: reports, Nodes: nodes} +} + +// bonded is the interface list of a host whose management network is a bond +// over two NICs, with the address wherever the case puts it. +func bonded(addresses map[string][]string) []nodeprobe.Interface { + holds := func(name string) ifaceOpt { return holding(addresses[name]...) } + return []nodeprobe.Interface{ + nic("bond0", inventory.LinkBond, over("eth0", "eth1"), under("bond0.100"), holds("bond0")), + nic("bond0.100", inventory.LinkVLAN, over("bond0"), holds("bond0.100")), + nic("eth0", inventory.LinkPhysical, at(25000), on(0), under("bond0"), + slotted("0000:3b:00.0"), holds("eth0")), + nic("eth1", inventory.LinkPhysical, at(25000), on(0), under("bond0"), + slotted("0000:3b:00.1"), holds("eth1")), + nic("lo", inventory.LinkLoopback, holding("127.0.0.1")), + } +} + +func netCases() map[string]Case { + // The cluster's own plumbing, which every worker of every Kubernetes fleet + // carries and none of it is a management interface. + plumbing := []nodeprobe.Interface{ + nic("cni0", inventory.LinkBridge, holding("10.42.2.1")), + nic("flannel.1", inventory.LinkVXLAN, holding("10.42.2.0")), + nic("lo", inventory.LinkLoopback, holding("127.0.0.1")), + } + + veths := func(n int) []nodeprobe.Interface { + out := make([]nodeprobe.Interface, 0, n) + for i := 0; i < n; i++ { + out = append(out, nic(fmt.Sprintf("veth%04x", 0x7a10+i), inventory.LinkVirtual, + holding(fmt.Sprintf("10.42.2.%d", 10+i)))) + } + return out + } + + cases := map[string]Case{ + "NET-01": known(netHost("worker-01", + nic("eth0", inventory.LinkPhysical, at(10000), holding(reachedAt)))), + + "NET-02": unknown(netHost("worker-01", + nic("eth0", inventory.LinkPhysical, at(1000), holding("192.168.10.1")), + nic("eth1", inventory.LinkPhysical, at(10000), holding("10.10.11.1")))), + + "NET-03": known(netHost("worker-01", + nic("eth0", inventory.LinkPhysical, at(1000), holding(reachedAt)), + nic("eth1", inventory.LinkPhysical, at(10000), holding("10.10.11.1")))), + + "NET-04": unknown(netHost("worker-01", + nic("eth0", inventory.LinkPhysical, at(1000), linkState("down"), holding("192.168.10.1")), + nic("eth1", inventory.LinkPhysical, at(10000), holding("169.254.1.1")), + nic("cni0", inventory.LinkBridge, at(1000), holding("10.42.2.1")), + nic("br0", inventory.LinkBridge, at(10000), holding("192.168.1.1")))), + + "NET-05": unknown(host("worker-01", cpu(1, 16, 2), disks(spread(3*tb, 4)...), ifaces())), + + "NET-06": unknown( + netHost("worker-01", nic("eth0", inventory.LinkPhysical, at(25000), holding("192.168.10.1"))), + netHost("worker-02", nic("ens5f0", inventory.LinkPhysical, at(25000), holding("192.168.10.2")))), + + "NET-07": unknown(netHost("worker-01", + nic("eth1", inventory.LinkPhysical, at(10000), holding("10.10.11.1")), + nic("eth0", inventory.LinkPhysical, at(10000), holding("192.168.10.1")))), + + "NET-08": unknown(netHost("worker-01", + nic("eth0", inventory.LinkPhysical, holding("192.168.10.1")), + nic("eth1", inventory.LinkPhysical, at(10000), linkState("down"), holding("10.10.11.1")))), + + "NET-09": known(netHost("worker-01", + nic("br0", inventory.LinkBridge, over("eth0"), holding(reachedAt)), + nic("eth0", inventory.LinkPhysical, at(10000), under("br0"), slotted("0000:3b:00.0")))), + + "NET-10": unknown(netHost("worker-01", + nic("eth0", inventory.LinkPhysical, at(10000), holding("2001:db8::1")))), + + "NET-11": unknown(netHost("worker-01", + nic("eth0", inventory.LinkPhysical, at(10000), holding("169.254.1.1", "192.168.10.1")))), + + "NET-12": unknown(netHost("worker-01", + nic("lo", inventory.LinkLoopback, holding("127.0.0.1")))), + + "NET-13": known(netHost("worker-01", bonded(map[string][]string{"bond0": {reachedAt}})...)), + + "NET-14": known(netHost("worker-01", + nic("eth0", inventory.LinkPhysical, at(25000), under("eth0.100"), slotted("0000:3b:00.0")), + nic("eth0.100", inventory.LinkVLAN, over("eth0"), holding(reachedAt)))), + + "NET-15": known(netHost("worker-01", + nic("eth0", inventory.LinkPhysical, at(10000), holding(reachedAt)), + nic("vxlan.calico", inventory.LinkVXLAN, holding("10.42.2.0")))), + + "NET-16": known(netHost("worker-01", bonded(map[string][]string{"bond0.100": {reachedAt}})...)), + + "NET-17": known(netHost("worker-01", + nic("eth0", inventory.LinkPhysical, at(25000), under("macvlan0", "ipvlan0"), + holding(reachedAt)), + nic("macvlan0", inventory.LinkMACVLAN, over("eth0"), holding("192.168.10.5")), + nic("ipvlan0", inventory.LinkIPVLAN, over("eth0"), holding("192.168.10.6")))), + + "NET-18": known(netHost("worker-01", append(veths(12), + nic("eth0", inventory.LinkPhysical, at(10000), holding(reachedAt)))...)), + + "NET-19": unknown(netHost("worker-01", + nic("cni0", inventory.LinkBridge, over("veth7a10"), holding("10.42.2.1")), + nic("docker0", inventory.LinkBridge, holding("172.17.0.1")), + nic("eth0", inventory.LinkPhysical, at(10000), slotted("0000:3b:00.0")))), + + "NET-20": unknown(netHost("worker-01", + nic("eth0", inventory.LinkPhysical, at(10000), linkState("dormant"), holding("192.168.10.1")), + nic("eth1", inventory.LinkPhysical, at(10000), linkState("lowerlayerdown"), + holding("10.10.11.1")))), + + "NET-21": unknown(netHost("worker-01", + nic("eth0", inventory.LinkPhysical, at(10000), linkState("unknown"), + holding("192.168.10.1")))), + + "NET-22": unknown(netHost("worker-01", + nic("ens5f0", inventory.LinkPhysical, at(10000), holding("192.168.10.1")), + nic("eth0", inventory.LinkPhysical, at(25000), holding("10.10.11.1")))), + + "NET-23": known(netHost("worker-01", + nic("eth0", inventory.LinkPhysical, at(10000), holding(reachedAt)), + nic("ens5f1", inventory.LinkPhysical, at(100000), on(0), holding("192.168.20.1")), + nic("ens5f2", inventory.LinkPhysical, at(100000), on(0), holding("192.168.21.1")))), + + "NET-24": unknown(netHost("worker-01", + nic("eth0", inventory.LinkPhysical, at(10000), holding("192.168.10.1")), + nic("ens5f0", inventory.LinkPhysical, at(100000), holding("192.168.20.1")))), + + "NET-25": unknown(netHost("worker-01", + nic("eth0", inventory.LinkPhysical, at(10000), frames(9000), holding("192.168.10.1")), + nic("eth1", inventory.LinkPhysical, at(10000), holding("10.10.11.1")))), + + "NET-26": unknown(netHost("worker-01", + nic("br0", inventory.LinkBridge, at(100000), over("eth1"), holding("192.168.1.1")), + nic("eth0", inventory.LinkPhysical, at(10000), holding("192.168.10.1")), + nic("eth1", inventory.LinkPhysical, at(100000), under("br0"), slotted("0000:3b:00.0")))), + + "NET-27": unknown(netHost("worker-01", + nic("eth0", inventory.LinkPhysical, at(10000), holding("192.168.10.1")))), + + "NET-28": unknown(netHost("worker-01", append( + bonded(map[string][]string{"bond0": {"192.168.10.1"}}), + nic("eth2", inventory.LinkPhysical, at(10000), slotted("0000:af:00.0"), + holding("192.168.20.1")))...)), + + "NET-29": unknown(netHost("worker-01", append( + bonded(map[string][]string{"bond0.100": {"192.168.10.1"}}), + nic("eth2", inventory.LinkPhysical, at(10000), slotted("0000:af:00.0"), + holding("192.168.20.1")))...)), + + "NET-31": known(netHost("worker-01", + nic("veth7a1c", inventory.LinkVirtual, holding(reachedAt)), + nic("lo", inventory.LinkLoopback, holding("127.0.0.1")))), + + "NET-33": known(netHost("worker-01", + nic("eth0", inventory.LinkPhysical, at(10000), linkState("down"), holding(reachedAt)))), + + "NET-34": known(netHost("worker-01", + bonded(map[string][]string{"bond0": {reachedAt}, "bond0.100": {reachedAt}})...)), + + "NET-35": unknown(netHost("worker-01", append( + bonded(map[string][]string{"bond0": {"192.168.10.1"}})[:1], + nic("eth0", inventory.LinkPhysical, at(25000), on(0), under("bond0"), + slotted("0000:3b:00.0")), + nic("eth1", inventory.LinkPhysical, at(25000), on(0), under("bond0"), + slotted("0000:3b:00.1")))...)), + + "NET-36": known(netHost("worker-01", + nic("bond0", inventory.LinkBond, over("eth9"), holding(reachedAt)))), + + "NET-37": known(netHost("worker-01", + nic("team0", inventory.LinkTeam, over("eth0", "eth1"), holding(reachedAt)), + nic("eth0", inventory.LinkPhysical, at(25000), on(0), under("team0"), + slotted("0000:3b:00.0")), + nic("eth1", inventory.LinkPhysical, at(25000), on(0), under("team0"), + slotted("0000:3b:00.1")))), + + "NET-38": known(netHost("worker-01", + nic("br0", inventory.LinkBridge, over("bond0"), holding(reachedAt)), + nic("bond0", inventory.LinkBond, over("eth0", "eth1"), under("br0")), + nic("eth0", inventory.LinkPhysical, at(25000), on(0), under("bond0"), + slotted("0000:3b:00.0")), + nic("eth1", inventory.LinkPhysical, at(25000), on(0), under("bond0"), + slotted("0000:3b:00.1")))), + + "NET-39": known(netHost("worker-01", + nic("eth0", inventory.LinkPhysical, at(25000), on(0), under("macvlan0"), + slotted("0000:3b:00.0")), + nic("macvlan0", inventory.LinkMACVLAN, over("eth0"), holding(reachedAt)))), + + "NET-40": unknown(netHost("worker-01", + nic("eth0", inventory.LinkPhysical, at(10000), unkinded(), holding("192.168.10.1")), + nic("cni0", inventory.LinkBridge, unkinded(), holding("10.42.2.1")), + nic("flannel.1", inventory.LinkVXLAN, unkinded(), holding("10.42.2.0")))), + + "NET-41": known(netHost("worker-01", bonded(map[string][]string{"bond0": {reachedAt}})...)), + + "NET-42": known(netHost("worker-01", bonded(map[string][]string{"bond0": {reachedAt}})...)), + } + + // A bond whose members sit in two sockets, which has no memory node at all. + split := bonded(map[string][]string{"bond0": {reachedAt}}) + for i := range split { + if split[i].Name == "eth1" { + split[i].NUMANode = 1 + } + } + cases["NET-30"] = known(netHost("worker-01", split...)) + + // The cluster's own plumbing beside nothing else, which is the worker a + // draft has to leave without an interface. + cases["NET-19"] = unknown(netHost("worker-01", append(plumbing, + nic("eth0", inventory.LinkPhysical, at(10000), slotted("0000:3b:00.0")))...)) + + // The seam case: a stack that points at itself cannot be written by the + // kernel, so it is built directly in the discovery package. + cases["NET-32"] = Case{ + Family: "net", Slug: "a-stack-that-points-at-itself", + Note: "Driven in the discovery package, because no probe can write a report whose " + + "lower_* links form a cycle and the case is about the resolution terminating anyway.", + } + + slugs := map[string]string{ + "NET-01": "one-addressed-physical-nic", + "NET-02": "the-faster-of-two-nics", + "NET-03": "the-node-address-beats-the-faster-link", + "NET-04": "every-interface-unusable", + "NET-05": "no-interfaces-reported", + "NET-06": "two-workers-naming-their-nics-differently", + "NET-07": "two-equal-nics", + "NET-08": "an-addressed-nic-reporting-no-speed", + "NET-09": "the-node-address-on-a-host-bridge", + "NET-10": "a-global-ipv6-address-only", + "NET-11": "a-link-local-and-a-routable-address", + "NET-12": "loopback-and-nothing-else", + "NET-13": "the-node-address-on-a-bond", + "NET-14": "the-node-address-on-a-vlan", + "NET-15": "an-overlay-beside-an-addressed-nic", + "NET-16": "the-node-address-on-a-vlan-over-a-bond", + "NET-17": "a-macvlan-and-an-ipvlan-over-one-nic", + "NET-18": "twelve-veths-beside-one-nic", + "NET-19": "the-clusters-own-plumbing-alone", + "NET-20": "a-dormant-link-and-a-down-lower-layer", + "NET-21": "a-link-whose-state-is-unknown", + "NET-22": "a-fleet-that-knows-which-nic-it-means", + "NET-23": "a-fleet-that-knows-its-data-nics", + "NET-24": "a-fast-link-and-a-slow-one", + "NET-25": "jumbo-frames-against-standard-ones", + "NET-26": "a-fast-bridge-against-a-slower-nic", + "NET-27": "a-link-up-with-no-partner", + "NET-28": "an-aggregate-with-no-speed-of-its-own", + "NET-29": "a-vlan-inheriting-the-bonds-speed", + "NET-30": "a-bond-across-two-sockets", + "NET-31": "a-veth-holding-the-node-address", + "NET-33": "the-node-address-on-a-down-link", + "NET-34": "the-node-address-on-two-interfaces", + "NET-35": "an-aggregate-reporting-its-own-speed", + "NET-36": "a-member-the-report-does-not-carry", + "NET-37": "a-team-interface", + "NET-38": "a-bridge-over-a-bond", + "NET-39": "a-macvlan-holding-the-node-address", + "NET-40": "an-interface-naming-no-kind", + "NET-41": "a-bonded-data-path-ranked-for-placement", + "NET-42": "what-the-draft-says-about-a-bonded-host", + } + gaps := map[string]string{ + "NET-04": "G-7", "NET-22": "G-23", "NET-23": "G-24", "NET-27": "G-26", + "NET-41": "G-27", "NET-42": "G-28", "NET-06": "G-25", + } + for id, slug := range slugs { + entry := cases[id] + entry.Slug = slug + entry.Gap = gaps[id] + cases[id] = entry + } + return cases +} diff --git a/operator/hack/discoveryfixtures/cases_numa.go b/operator/hack/discoveryfixtures/cases_numa.go new file mode 100644 index 000000000..1561dab9c --- /dev/null +++ b/operator/hack/discoveryfixtures/cases_numa.go @@ -0,0 +1,127 @@ +// The NUMA topology cases, §3 of the document. +// +// The placement ranks a worker's memory nodes by unclaimed device count, then +// capacity, then physical cores, then a real node ahead of the bucket that is +// not one, then the node identifier. Each case here leaves exactly one of those +// comparisons deciding, so that a change to the ranking fails the case that +// names the rung it changed. + +package main + +import ( + "fmt" + + "github.com/simplyblock/atlas/inventory" + "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" +) + +// spread lays n disks of one size across the memory nodes given, one slot each +// in a bus per node, so that two workers built the same way name the same +// addresses. +func spread(size uint64, perNode ...int) []nodeprobe.Device { + var out []nodeprobe.Device + index := 0 + for node, count := range perNode { + for i := 0; i < count; i++ { + out = append(out, nvme( + fmt.Sprintf("nvme%dn1", index), + fmt.Sprintf("0000:%02x:00.%d", 0x5e+node*0x20, i), + node, size)) + index++ + } + } + return out +} + +func numaCases() map[string]Case { + twoNode := func(reports ...nodeprobe.Report) Case { + return Case{Family: "numa", Reports: reports, Nodes: kubeFleet(reports)} + } + + oneNode := host("worker-01", cpu(1, 16, 2), disks(spread(3*tb, 4)...)) + + unknownNode := func(devices ...nodeprobe.Device) []nodeprobe.Device { + for i := range devices { + devices[i].NUMANode = inventory.NUMANodeUnknown + } + return devices + } + + cases := map[string]Case{ + "NUMA-01": twoNode(oneNode), + "NUMA-02": twoNode(host("worker-01", disks(spread(3*tb, 2, 2)...))), + "NUMA-03": twoNode(host("worker-01", disks(spread(3*tb, 1, 3)...))), + "NUMA-04": twoNode(host("worker-01", disks(append( + spread(8*tb, 2), spread(tb, 0, 3)...)...))), + "NUMA-05": twoNode(host("worker-01", disks(append( + spread(tb, 2), spread(2*tb, 0, 2)...)...))), + "NUMA-06": twoNode(host("worker-01", cpuNodes(2, 8, 24), disks(spread(3*tb, 2, 2)...))), + "NUMA-07": twoNode(host("worker-01", cpu(4, 16, 2), disks(spread(3*tb, 2, 2, 2, 2)...))), + "NUMA-08": twoNode(host("worker-01", cpu(8, 8, 2), + disks(spread(3*tb, 1, 1, 1, 1, 1, 3, 1, 1)...))), + "NUMA-09": twoNode(host("worker-01", disks(unknownNode(spread(3*tb, 4)...)...))), + "NUMA-10": twoNode(host("worker-01", disks(append( + spread(3*tb, 2), unknownNode( + nvme("nvme8n1", "0000:c0:00.0", 0, 3*tb), + nvme("nvme9n1", "0000:c1:00.0", 0, 3*tb), + )...)...))), + "NUMA-12": { + Family: "numa", + Note: "Both workers hand over the same two addresses, and the grouper reads the " + + "addresses rather than the topology, so they share a group.", + Reports: []nodeprobe.Report{ + host("worker-01", cpu(1, 16, 2), disks( + nvme("nvme0n1", "0000:5e:00.0", 0, 3*tb), + nvme("nvme1n1", "0000:5e:00.1", 0, 3*tb), + )), + host("worker-02", disks( + nvme("nvme0n1", "0000:5e:00.0", 0, 3*tb), + nvme("nvme1n1", "0000:5e:00.1", 1, 3*tb), + )), + }, + }, + "NUMA-13": twoNode(host("worker-01", disks(spread(3*tb, 0, 10)...))), + "NUMA-14": twoNode(host("worker-01", cpu(2, 49, 2), + mem(1024*gb, 1000*gb, 512*gb, 512*gb), + pages(pagePool(gb, 256, 256)), + disks(spread(3*tb, 5, 5)...))), + "NUMA-15": twoNode(host("worker-01", noNUMATopology(2, 8, 2), disks(spread(3*tb, 2, 2)...))), + + // The seam case: the controller never substitutes a placement, so there + // is nothing to drive it from and the row carries its own reasoning. + "NUMA-11": { + Family: "numa", Slug: "all-devices-placement", + Note: "Driven in the discovery package with Planner{Placement: AllDevices{}} over the " + + "NUMA-02 fleet, because the controller always builds the default Planner.", + }, + } + + // NUMA-12 states its workers directly rather than through the helper, so it + // is the one case here that would otherwise carry no node objects. + twelve := cases["NUMA-12"] + twelve.Nodes = kubeFleet(twelve.Reports) + cases["NUMA-12"] = twelve + + slugs := map[string]string{ + "NUMA-01": "one-memory-node", + "NUMA-02": "two-nodes-evenly-split", + "NUMA-03": "two-nodes-one-against-three", + "NUMA-04": "count-beats-capacity", + "NUMA-05": "capacity-breaks-the-count-tie", + "NUMA-06": "cores-break-the-capacity-tie", + "NUMA-07": "four-memory-nodes", + "NUMA-08": "eight-memory-nodes", + "NUMA-09": "every-device-on-no-node", + "NUMA-10": "a-real-node-against-the-unknown-bucket", + "NUMA-12": "one-node-and-two-node-workers-agree-on-addresses", + "NUMA-13": "every-disk-on-the-second-node", + "NUMA-14": "a-large-two-socket-worker", + "NUMA-15": "no-memory-node-carries-cores", + } + for id, slug := range slugs { + entry := cases[id] + entry.Slug = slug + cases[id] = entry + } + return cases +} diff --git a/operator/hack/discoveryfixtures/cases_pci.go b/operator/hack/discoveryfixtures/cases_pci.go new file mode 100644 index 000000000..581895204 --- /dev/null +++ b/operator/hack/discoveryfixtures/cases_pci.go @@ -0,0 +1,265 @@ +// The PCI addressing and fleet-grouping cases, §5 and §6 of the document. +// +// A group's device selection is shared by every worker in it, so grouping is a +// statement about what the machines have rather than a presentation choice. +// These cases vary how much of a fleet agrees: all of it, none of it, and the +// partial agreements in between that are what a real fleet looks like after a +// few years of replacements. + +package main + +import ( + "fmt" + + corev1 "k8s.io/api/core/v1" + + "github.com/simplyblock/atlas/blockdev" + "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" +) + +// layout is one arrangement of slots, which is what two workers have to agree +// on to share a group. +type layout []string + +var ( + layoutA = layout{"0000:5e:00.0", "0000:5f:00.0", "0000:af:00.0", "0000:b0:00.0"} + layoutB = layout{"0000:3b:00.0", "0000:3c:00.0", "0000:d8:00.0", "0000:d9:00.0"} +) + +// distinct is a layout nothing else in the fleet shares, keyed off the worker's +// number so that every such worker differs from every other. +func distinct(index int) layout { + return layout{ + fmt.Sprintf("0000:%02x:00.0", 0x10+index), + fmt.Sprintf("0000:%02x:00.1", 0x10+index), + } +} + +// at builds the disks of a layout, all on memory node 0 and all the same size. +func (l layout) disks(size uint64) []nodeprobe.Device { + out := make([]nodeprobe.Device, 0, len(l)) + for i, address := range l { + out = append(out, nvme(fmt.Sprintf("nvme%dn1", i), address, 0, size)) + } + return out +} + +// workersOn is n workers all handing over the same layout. +func workersOn(n int, l layout) []nodeprobe.Report { + return fleet(n, func(_ int, name string) nodeprobe.Report { + return host(name, cpu(1, 16, 2), disks(l.disks(3*tb)...)) + }) +} + +func pciCases() map[string]Case { + withNodes := func(reports []nodeprobe.Report) Case { + return Case{Family: "pci", Reports: reports, Nodes: kubeFleet(reports)} + } + + // Sixteen on one layout and sixteen on another. + twoLayouts := fleet(32, func(index int, name string) nodeprobe.Report { + chosen := layoutA + if index > 16 { + chosen = layoutB + } + return host(name, cpu(1, 16, 2), disks(chosen.disks(3*tb)...)) + }) + + // Twenty on one, six on another, and six that agree with nobody: the shape + // of a fleet nobody bought all at once. + stragglers := fleet(32, func(index int, name string) nodeprobe.Report { + switch { + case index <= 20: + return host(name, cpu(1, 16, 2), disks(layoutA.disks(3*tb)...)) + case index <= 26: + return host(name, cpu(1, 16, 2), disks(layoutB.disks(3*tb)...)) + default: + return host(name, cpu(1, 16, 2), disks(distinct(index).disks(3*tb)...)) + } + }) + + // Eight identical machines but for one disk of one of them in another slot. + oddOneOut := fleet(8, func(index int, name string) nodeprobe.Report { + chosen := layoutA + if index == 5 { + chosen = layout{"0000:5e:00.0", "0000:5f:00.0", "0000:af:00.0", "0000:c8:00.0"} + } + return host(name, cpu(1, 16, 2), disks(chosen.disks(3*tb)...)) + }) + + // Three on one layout, two agreeing with nobody. + partial := fleet(5, func(index int, name string) nodeprobe.Report { + chosen := layoutA + if index > 3 { + chosen = distinct(index) + } + return host(name, cpu(1, 16, 2), disks(chosen.disks(3*tb)...)) + }) + + cases := map[string]Case{ + "PCI-01": withNodes(workersOn(32, layoutA)), + "PCI-02": withNodes([]nodeprobe.Report{ + host("worker-01", cpu(1, 16, 2), disks(layoutA.disks(3*tb)...)), + host("worker-02", cpu(1, 16, 2), disks(layoutB.disks(3*tb)...)), + }), + "PCI-03": withNodes(twoLayouts), + "PCI-04": withNodes(fleet(32, func(index int, name string) nodeprobe.Report { + return host(name, cpu(1, 16, 2), disks(distinct(index).disks(3*tb)...)) + })), + "PCI-05": withNodes([]nodeprobe.Report{host("worker-01", cpu(1, 16, 2), disks( + nvme("nvme3n1", "0000:b0:00.0", 0, 3*tb), + nvme("nvme2n1", "0000:af:00.0", 0, 3*tb), + nvme("nvme1n1", "0000:5f:00.0", 0, 3*tb), + nvme("nvme0n1", "0000:5e:00.0", 0, 3*tb), + ))}), + "PCI-06": withNodes([]nodeprobe.Report{host("worker-01", cpu(1, 16, 2), disks( + layout{ + "0000:5e:00.0", "0000:5e:00.1", "0000:5f:00.0", "0000:5f:00.1", "0000:60:00.0", + "0000:af:00.0", "0000:af:00.1", "0000:b0:00.0", "0000:b0:00.1", "0000:b1:00.0", + }.disks(3*tb)...))}), + "PCI-07": withNodes([]nodeprobe.Report{host("worker-01", cpu(1, 16, 2), disks( + nvme("nvme0n1", "10000:01:00.0", 0, 3*tb), + nvme("nvme1n1", "10000:02:00.0", 0, 3*tb), + ))}), + "PCI-08": withNodes([]nodeprobe.Report{ + host("worker-01", cpu(1, 16, 2), disks( + nvme("nvme0n1", "0000:5E:00.0", 0, 3*tb), + nvme("nvme1n1", "0000:5F:00.0", 0, 3*tb), + )), + host("worker-02", cpu(1, 16, 2), disks( + nvme("nvme0n1", "0000:5e:00.0", 0, 3*tb), + nvme("nvme1n1", "0000:5f:00.0", 0, 3*tb), + )), + }), + "PCI-10": withNodes(partial), + "PCI-11": withNodes(stragglers), + "PCI-12": withNodes(oddOneOut), + "PCI-13": withNodes([]nodeprobe.Report{ + host("worker-01", cpu(1, 16, 2), disks(layoutA[:2].disks(2*tb)...)), + host("worker-02", cpu(1, 16, 2), disks(layoutA[:2].disks(4*tb)...)), + }), + "PCI-14": withNodes([]nodeprobe.Report{ + host("worker-01", cpu(1, 16, 2), disks(layoutA.disks(3*tb)...)), + host("worker-02", cpu(1, 16, 2), disks(layoutA.disks(3*tb)...)), + host("worker-03", cpu(1, 16, 2), disks(layoutA[:3].disks(3*tb)...)), + }), + } + + // Unpadded names, which sort worker-1, worker-10, worker-11 and decide the + // order the groups are numbered in. + unpadded := make([]nodeprobe.Report, 0, 32) + for i := 1; i <= 32; i++ { + chosen := layoutA + if i > 16 { + chosen = layoutB + } + unpadded = append(unpadded, + host(fmt.Sprintf("worker-%d", i), cpu(1, 16, 2), disks(chosen.disks(3*tb)...))) + } + cases["PCI-09"] = withNodes(unpadded) + + slugs := map[string]string{ + "PCI-01": "a-fleet-that-agrees", + "PCI-02": "two-workers-that-do-not", + "PCI-03": "sixteen-and-sixteen", + "PCI-04": "every-worker-distinct", + "PCI-05": "addresses-reported-descending", + "PCI-06": "ten-disks-across-two-buses", + "PCI-07": "a-five-digit-pci-domain", + "PCI-08": "uppercase-hex-against-lowercase", + "PCI-09": "unpadded-worker-names", + "PCI-10": "three-agree-and-two-do-not", + "PCI-11": "a-majority-layout-with-stragglers", + "PCI-12": "one-slot-moved-on-one-worker", + "PCI-13": "one-layout-two-capacities", + "PCI-14": "a-worker-holding-a-subset", + } + gaps := map[string]string{"PCI-07": "G-6", "PCI-13": "G-13"} + for id, slug := range slugs { + entry := cases[id] + entry.Slug = slug + entry.Gap = gaps[id] + cases[id] = entry + } + return cases +} + +func fleetCases() map[string]Case { + withNodes := func(reports []nodeprobe.Report) Case { + return Case{Family: "fleet", Reports: reports, Nodes: kubeFleet(reports)} + } + + uniform32 := workersOn(32, layoutA) + + // Three workers of the thirty-two have nothing the rules would take. + someRefused := fleet(32, func(index int, name string) nodeprobe.Report { + if index%11 == 0 { + return host(name, cpu(1, 16, 2), disks( + nvme("nvme0n1", "0000:5e:00.0", 0, 3*tb, refused(blockdev.ReasonMounted)), + nvme("nvme1n1", "0000:5f:00.0", 0, 3*tb, refused(blockdev.ReasonMounted)), + )) + } + return host(name, cpu(1, 16, 2), disks(layoutA.disks(3*tb)...)) + }) + + // The same fleet with the reports reversed, which is what a different + // listing order looks like from the operator's side. + reversed := make([]nodeprobe.Report, 0, len(uniform32)) + for i := len(uniform32) - 1; i >= 0; i-- { + reversed = append(reversed, uniform32[i]) + } + + cases := map[string]Case{ + "FLEET-01": withNodes(workersOn(1, layoutA)), + "FLEET-02": withNodes(workersOn(3, layoutA)), + "FLEET-03": withNodes(uniform32), + "FLEET-05": withNodes(someRefused), + "FLEET-06": withNodes(reversed), + } + + // A run whose probes wrote nothing at all. + cases["FLEET-04"] = Case{ + Family: "fleet", Workers: []string{"worker-01", "worker-02", "worker-03"}, + Nodes: []corev1.Node{ + kubeNode("worker-01", reachableAt("10.10.10.1")), + kubeNode("worker-02", reachableAt("10.10.10.2")), + kubeNode("worker-03", reachableAt("10.10.10.3")), + }, + } + + // Two ConfigMaps carrying a report for one node, which is what a probe + // restarted under a second object name leaves behind. + doubled := workersOn(2, layoutA) + cases["FLEET-07"] = Case{ + Family: "fleet", Reports: doubled, Nodes: kubeFleet(doubled), + Amend: func(run string, maps []*corev1.ConfigMap) []*corev1.ConfigMap { + copied := maps[0].DeepCopy() + copied.Name += "-again" + return append(maps, copied) + }, + } + + // A report for a machine this run is not about. + foreign := workersOn(3, layoutA) + cases["FLEET-08"] = Case{ + Family: "fleet", Reports: foreign, Nodes: kubeFleet(foreign), + Workers: []string{"worker-01", "worker-02"}, + } + + slugs := map[string]string{ + "FLEET-01": "one-worker", + "FLEET-02": "three-uniform-workers", + "FLEET-03": "thirty-two-uniform-workers", + "FLEET-04": "no-reports-at-all", + "FLEET-05": "three-of-thirty-two-refused", + "FLEET-06": "the-same-fleet-in-another-order", + "FLEET-07": "two-reports-for-one-worker", + "FLEET-08": "a-report-this-run-is-not-about", + } + for id, slug := range slugs { + entry := cases[id] + entry.Slug = slug + cases[id] = entry + } + return cases +} diff --git a/operator/hack/discoveryfixtures/cases_role.go b/operator/hack/discoveryfixtures/cases_role.go new file mode 100644 index 000000000..d2cfcd475 --- /dev/null +++ b/operator/hack/discoveryfixtures/cases_role.go @@ -0,0 +1,241 @@ +// The node-role cases, §8 of the document, and the filter cases, §9. +// +// A role is a label rather than a taint, and the two do not coincide: +// Kubernetes taints its control-plane nodes and OpenShift usually does not taint +// its infrastructure ones. The role cases are therefore about node objects +// rather than about reports, and every worker in them carries the same disks. +// +// The filter cases are the opposite shape: one worker with a deliberately mixed +// set of disks, and a different filter on each run. + +package main + +import ( + corev1 "k8s.io/api/core/v1" + + "github.com/simplyblock/atlas/blockdev" + "github.com/simplyblock/atlas/ptr" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" +) + +// sameDisks is n workers that agree on everything, so that a role case's groups +// differ only by what the node objects say. +func sameDisks(n int) []nodeprobe.Report { + return fleet(n, func(_ int, name string) nodeprobe.Report { + return host(name, cpu(1, 16, 2), disks(layoutA.disks(3*tb)...)) + }) +} + +func roleCases() map[string]Case { + three := sameDisks(3) + six := sameDisks(6) + + // Three infrastructure nodes and three plain workers, identical hardware. + tiered := make([]corev1.Node, 0, len(six)) + for i, report := range six { + opts := []nodeOpt{reachableAt(managementAddress(report.Node))} + if i < 3 { + opts = append(opts, role("infra")) + } + tiered = append(tiered, kubeNode(report.Node, opts...)) + } + + // A control-plane node with free disks beside two workers. + mixedRoles := []corev1.Node{ + kubeNode(three[0].Node, reachableAt(managementAddress(three[0].Node)), + role("control-plane"), tainted("node-role.kubernetes.io/control-plane", "", + corev1.TaintEffectNoSchedule)), + kubeNode(three[1].Node, reachableAt(managementAddress(three[1].Node))), + kubeNode(three[2].Node, reachableAt(managementAddress(three[2].Node))), + } + + // Thirty-two machines across the three tiers. + large := sameDisks(32) + spreadRoles := make([]corev1.Node, 0, len(large)) + for i, report := range large { + opts := []nodeOpt{reachableAt(managementAddress(report.Node))} + switch { + case i < 4: + opts = append(opts, role("infra")) + case i < 7: + opts = append(opts, role("control-plane")) + default: + opts = append(opts, role("worker")) + } + spreadRoles = append(spreadRoles, kubeNode(report.Node, opts...)) + } + + withControlPlane := &simplyblockv1alpha2.DiscoverSpec{ + EnableControlPlaneNodes: ptr.To(true), + } + + cases := map[string]Case{ + "ROLE-01": {Family: "role", Reports: three, Nodes: kubeFleet(three)}, + "ROLE-02": {Family: "role", Reports: six, Nodes: tiered}, + "ROLE-03": {Family: "role", Reports: three, Nodes: mixedRoles, Discover: withControlPlane}, + "ROLE-04": {Family: "role", Reports: three[:1], + Nodes: []corev1.Node{kubeNode(three[0].Node, + reachableAt(managementAddress(three[0].Node)), role("infra"), role("worker"))}}, + "ROLE-05": {Family: "role", Reports: three[:1], Discover: withControlPlane, + Nodes: []corev1.Node{kubeNode(three[0].Node, + reachableAt(managementAddress(three[0].Node)), role("control-plane"), role("worker"))}}, + "ROLE-06": {Family: "role", Reports: three[:1], Discover: withControlPlane, + Nodes: []corev1.Node{kubeNode(three[0].Node, + reachableAt(managementAddress(three[0].Node)), role("master"))}}, + "ROLE-08": {Family: "role", Reports: three[:2], + Nodes: []corev1.Node{ + kubeNode(three[0].Node, reachableAt(managementAddress(three[0].Node)), cordoned()), + kubeNode(three[1].Node, reachableAt(managementAddress(three[1].Node)), + tainted("storage", "drained", corev1.TaintEffectNoExecute)), + }}, + "ROLE-09": {Family: "role", Reports: three[:1], + Nodes: []corev1.Node{kubeNode(three[0].Node, + reachableAt(managementAddress(three[0].Node)), role("database"))}}, + "ROLE-10": {Family: "role", Reports: large, Nodes: spreadRoles, Discover: withControlPlane}, + + // The seam case: a plan built with no node objects at all, which the + // controller never does because it always lists them. + "ROLE-07": {Family: "role", Slug: "a-plan-with-no-node-objects", + Note: "Driven in the discovery package with a Planner carrying no KubeNodes, because the " + + "controller always reads the node objects for the workers its run settled on."}, + } + + slugs := map[string]string{ + "ROLE-01": "workers-with-no-role-label", + "ROLE-02": "an-infrastructure-tier-beside-the-workers", + "ROLE-03": "a-control-plane-node-with-disks", + "ROLE-04": "a-node-labeled-infra-and-worker", + "ROLE-05": "a-node-labeled-control-plane-and-worker", + "ROLE-06": "the-master-spelling", + "ROLE-08": "a-cordoned-node-and-a-tainted-one", + "ROLE-09": "a-role-this-product-does-not-know", + "ROLE-10": "thirty-two-machines-across-three-tiers", + } + gaps := map[string]string{"ROLE-08": "G-8"} + for id, slug := range slugs { + entry := cases[id] + entry.Slug = slug + entry.Gap = gaps[id] + cases[id] = entry + } + return cases +} + +// mixedDisks is the worker every filter case is run against: four disks of two +// models and three sizes, one of them partitioned and one mounted, so that each +// filter has something to take and something to leave. +func mixedDisks() []nodeprobe.Device { + return []nodeprobe.Device{ + nvme("nvme0n1", "0000:5e:00.0", 0, 2*tb), + nvme("nvme1n1", "0000:5f:00.0", 0, 2*tb, modeled("INTEL SSDPF2KX038TZ")), + nvme("nvme2n1", "0000:af:00.0", 0, 512*gb), + nvme("nvme3n1", "0000:b0:00.0", 0, 2*tb, refused(blockdev.ReasonPartitioned)), + nvme("nvme4n1", "0000:b1:00.0", 0, 2*tb, + refused(blockdev.ReasonPartitioned, blockdev.ReasonMounted)), + } +} + +// filtered is a run carrying one device filter. +func filtered(edit func(*simplyblockv1alpha2.DeviceFilter)) *simplyblockv1alpha2.DiscoverSpec { + filter := &simplyblockv1alpha2.DeviceFilter{} + edit(filter) + return &simplyblockv1alpha2.DiscoverSpec{DeviceFilter: filter} +} + +func filterCases() map[string]Case { + nvmeWorker := func() []nodeprobe.Report { + return []nodeprobe.Report{host("worker-01", cpu(1, 16, 2), disks(mixedDisks()...))} + } + blockWorker := func() []nodeprobe.Report { + return []nodeprobe.Report{host("worker-01", cpu(1, 16, 2), disks( + blk("sda", 0, 512*gb), + blk("sdb", 0, 2*tb), + blk("sdc", 0, 2*tb), + blk("vdb", 0, 4*tb), + blk("vdc", 0, 4*tb), + blk("vdd", 0, 4*tb), + ))} + } + + with := func(reports []nodeprobe.Report, spec *simplyblockv1alpha2.DiscoverSpec) Case { + return Case{Family: "filt", Reports: reports, Nodes: kubeFleet(reports), Discover: spec} + } + block := func(edit func(*simplyblockv1alpha2.DeviceFilter)) *simplyblockv1alpha2.DiscoverSpec { + return filtered(func(f *simplyblockv1alpha2.DeviceFilter) { + f.EnableLogicalBlockDevices = ptr.To(true) + edit(f) + }) + } + + cases := map[string]Case{ + "FILT-01": with(nvmeWorker(), filtered(func(f *simplyblockv1alpha2.DeviceFilter) { + f.PcieDenyList = []string{"0000:5e:00.0"} + })), + "FILT-02": with(nvmeWorker(), filtered(func(f *simplyblockv1alpha2.DeviceFilter) { + f.PcieAllowList = []string{"0000:5e:00.0", "0000:5f:00.0"} + })), + "FILT-03": with(nvmeWorker(), filtered(func(f *simplyblockv1alpha2.DeviceFilter) { + f.PcieModel = "MZQL2" + })), + "FILT-04": with(nvmeWorker(), filtered(func(f *simplyblockv1alpha2.DeviceFilter) { + f.DriveSizeRange = "1T-4T" + })), + "FILT-05": with([]nodeprobe.Report{host("worker-01", cpu(1, 16, 2), disks( + nvme("nvme0n1", "0000:5e:00.0", 0, 2*tb), + nvme("nvme1n1", "0000:5f:00.0", 0, 2*tb), + nvme("nvme2n1", "0000:af:00.0", 0, 1920*gb), + ))}, filtered(func(f *simplyblockv1alpha2.DeviceFilter) { f.DriveSizeRange = "2T" })), + "FILT-06": with(nvmeWorker(), filtered(func(f *simplyblockv1alpha2.DeviceFilter) { + f.DriveSizeRange = "2T-1T" + })), + "FILT-07": with(blockWorker(), block(func(f *simplyblockv1alpha2.DeviceFilter) { + f.BlockDenyList = []string{"/dev/sda"} + })), + "FILT-08": with(blockWorker(), block(func(f *simplyblockv1alpha2.DeviceFilter) { + f.BlockAllowList = []string{"/dev/vdb", "/dev/vdc"} + })), + "FILT-09": with(nvmeWorker(), filtered(func(f *simplyblockv1alpha2.DeviceFilter) { + f.EnablePartitionedDevices = ptr.To(true) + })), + "FILT-10": with(nvmeWorker(), filtered(func(f *simplyblockv1alpha2.DeviceFilter) { + f.PcieAllowList = []string{"0000:ff:00.0"} + })), + "FILT-11": with(nvmeWorker(), filtered(func(f *simplyblockv1alpha2.DeviceFilter) { + f.PcieAllowList = []string{"0000:5E:00.0", "0000:5F:00.0"} + })), + "FILT-12": with(nvmeWorker(), filtered(func(f *simplyblockv1alpha2.DeviceFilter) { + f.PcieAllowList = []string{"0000:5e:00.0", "0000:5f:00.0"} + f.PcieDenyList = []string{"0000:5e:00.0"} + })), + "FILT-13": with(blockWorker(), block(func(f *simplyblockv1alpha2.DeviceFilter) { + f.DriveSizeRange = "3T-8T" + })), + "FILT-14": with(nvmeWorker(), nil), + } + + slugs := map[string]string{ + "FILT-01": "a-pci-deny-list", + "FILT-02": "a-pci-allow-list", + "FILT-03": "a-model-substring", + "FILT-04": "a-size-range", + "FILT-05": "a-bare-size-is-both-bounds", + "FILT-06": "a-size-range-that-counts-backward", + "FILT-07": "a-block-deny-list", + "FILT-08": "a-block-allow-list", + "FILT-09": "a-partition-table-waived", + "FILT-10": "a-filter-that-matches-nothing", + "FILT-11": "an-allow-list-in-uppercase", + "FILT-12": "one-address-allowed-and-denied", + "FILT-13": "a-size-range-on-a-block-run", + "FILT-14": "no-filter-at-all", + } + gaps := map[string]string{"FILT-06": "G-9"} + for id, slug := range slugs { + entry := cases[id] + entry.Slug = slug + entry.Gap = gaps[id] + cases[id] = entry + } + return cases +} diff --git a/operator/hack/discoveryfixtures/cases_size.go b/operator/hack/discoveryfixtures/cases_size.go new file mode 100644 index 000000000..9cabb5575 --- /dev/null +++ b/operator/hack/discoveryfixtures/cases_size.go @@ -0,0 +1,128 @@ +// The node-size cases, §4 of the document: vCPU, memory, and the huge pages a +// host already has. +// +// simplyblock allocates its own huge pages, so what these cases vary is the +// baseline the new allocation is added to, not a prerequisite. A host with +// nothing set aside is the ordinary starting state and several of the cases are +// there to record that the generator treats it as a reason to propose nothing. + +package main + +import ( + corev1 "k8s.io/api/core/v1" + + "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" +) + +func sizeCases() map[string]Case { + one := func(report nodeprobe.Report) Case { + return Case{Family: "size", Reports: []nodeprobe.Report{report}, + Nodes: kubeFleet([]nodeprobe.Report{report})} + } + fourDisks := disks(spread(3*tb, 4)...) + + cases := map[string]Case{ + "SIZE-01": one(host("worker-01", cpu(1, 4, 2), fourDisks)), + "SIZE-02": one(host("worker-01", cpu(1, 2, 1), mem(16*gb, 14*gb, 16*gb), fourDisks)), + "SIZE-03": one(host("worker-01", cpu(1, 98, 2), mem(1024*gb, 1000*gb, 1024*gb), + pages(pagePool(gb, 256)), fourDisks)), + "SIZE-04": one(host("worker-01", cpu(2, 49, 2), mem(1024*gb, 1000*gb, 512*gb, 512*gb), + pages(pagePool(gb, 128, 128)), disks(spread(3*tb, 2, 2)...))), + "SIZE-06": one(host("worker-01", cpu(1, 4, 2), mem(4*gb, 900*mb, 4*gb), fourDisks)), + "SIZE-07": one(host("worker-01", cpu(1, 16, 2), mem(0, 0), + unreadable("read /proc/meminfo: permission denied"), fourDisks)), + "SIZE-09": one(host("worker-01", pages(pagePool(gb, 256, 256)), + reserved(512*gb), disks(spread(3*tb, 2, 2)...))), + "SIZE-11": one(host("worker-01", cpu(1, 16, 2), pages(pagePool(2*mb, 256)), fourDisks)), + "SIZE-12": one(host("worker-01", cpu(1, 16, 2), + pages(nodeprobe.HugePagePool{SizeBytes: gb, Total: 64, Free: 64}), fourDisks)), + "SIZE-13": one(host("worker-01", pages(pagePool(gb, 0, 128)), + disks(spread(3*tb, 2, 2)...))), + "SIZE-14": one(host("worker-01", cpu(1, 16, 2), swap(8*gb, 0), fourDisks)), + "SIZE-16": one(host("worker-01", cpu(1, 16, 2), mem(56*gb, 40*gb, 56*gb), + reserved(200*gb), pages(pagePool(gb, 200)), fourDisks)), + "SIZE-17": one(host("worker-01", cpu(1, 16, 2), mem(1024*gb, 1000*gb, 1024*gb), + noPages(), fourDisks)), + "SIZE-18": one(host("worker-01", cpu(1, 16, 2), pages(pagePool(gb, 256)), + reserved(256*gb), fourDisks)), + "SIZE-19": one(host("worker-01", cpu(1, 16, 2), + pages(nodeprobe.HugePagePool{SizeBytes: gb, Total: 256, Free: 0, + NUMANodes: []nodeprobe.NUMAHugePages{{Node: 0, Total: 256, Free: 0}}}), + reserved(256*gb), fourDisks)), + "SIZE-20": one(host("worker-01", cpu(1, 16, 2), + pages(pagePool(2*mb, 1024), pagePool(gb, 32)), fourDisks)), + "SIZE-21": one(host("worker-01", cpu(1, 16, 2), noPages(), fourDisks)), + } + + // A fleet whose chosen memory nodes carry different core counts, which is + // what the cluster's one vCPU count has to fit. + uneven := []nodeprobe.Report{ + host("worker-01", cpu(1, 4, 2), disks(spread(3*tb, 4)...)), + host("worker-02", cpu(1, 16, 2), disks(spread(3*tb, 4)...)), + host("worker-03", cpu(1, 49, 2), disks(spread(3*tb, 4)...)), + } + cases["SIZE-05"] = Case{Family: "size", Reports: uneven, Nodes: kubeFleet(uneven)} + + // One worker of three has nothing set aside on the node it is placed on, + // and the whole cluster's proposal follows from it. + mixed := []nodeprobe.Report{ + host("worker-01", pages(pagePool(gb, 128, 128)), disks(spread(3*tb, 2, 2)...)), + host("worker-02", noPages(), disks(spread(3*tb, 2, 2)...)), + host("worker-03", pages(pagePool(gb, 128, 128)), disks(spread(3*tb, 2, 2)...)), + } + cases["SIZE-10"] = Case{Family: "size", Reports: mixed, Nodes: kubeFleet(mixed)} + + // What Kubernetes says about the machine, which is not what the machine + // says about itself. + held := host("worker-01", cpu(1, 16, 2), mem(256*gb, 240*gb, 256*gb), disks(spread(3*tb, 4)...)) + cases["SIZE-15"] = Case{ + Family: "size", Reports: []nodeprobe.Report{held}, + Nodes: []corev1.Node{kubeNode("worker-01", + reachableAt(managementAddress("worker-01")), + sized("32", "256Gi", "28", "180Gi"), + schedulableHugePages("1Gi", "64Gi"))}, + } + + // The seam case: WorkerWasReadable is off by default and the controller + // never turns it on. + cases["SIZE-08"] = Case{ + Family: "size", Slug: "the-fully-readable-worker-rule", + Note: "Driven in the discovery package with Planner{WorkerRules: []WorkerRule{WorkerHasDevices{}, " + + "WorkerWasReadable{}}} over the SIZE-07 report, because the controller never substitutes a worker rule.", + } + + slugs := map[string]string{ + "SIZE-01": "eight-logical-cpus", + "SIZE-02": "two-physical-cores", + "SIZE-03": "one-socket-196-vcpus", + "SIZE-04": "two-sockets-196-vcpus", + "SIZE-05": "a-fleet-of-uneven-core-counts", + "SIZE-06": "four-gibibytes-of-memory", + "SIZE-07": "an-unreadable-memory-reading", + "SIZE-09": "pages-already-set-aside", + "SIZE-10": "one-worker-with-nothing-set-aside", + "SIZE-11": "a-reservation-under-a-gigabyte", + "SIZE-12": "a-reservation-with-no-node-breakdown", + "SIZE-13": "pages-on-the-node-not-chosen", + "SIZE-14": "swap-in-use", + "SIZE-15": "allocatable-far-under-capacity", + "SIZE-16": "no-room-for-an-allocation-on-top", + "SIZE-17": "room-and-nothing-set-aside", + "SIZE-18": "a-reuse-instruction-with-nowhere-to-put-it", + "SIZE-19": "every-page-promised-to-a-mapping", + "SIZE-20": "two-page-sizes-at-once", + "SIZE-21": "no-hugetlbfs-at-all", + } + gaps := map[string]string{ + "SIZE-03": "G-3", "SIZE-06": "G-4", "SIZE-09": "G-14", "SIZE-10": "G-15", + "SIZE-14": "G-5", "SIZE-16": "G-16", "SIZE-17": "G-15 and G-16", + "SIZE-18": "G-17", "SIZE-19": "G-18", "SIZE-21": "G-19", + } + for id, slug := range slugs { + entry := cases[id] + entry.Slug = slug + entry.Gap = gaps[id] + cases[id] = entry + } + return cases +} diff --git a/operator/hack/discoveryfixtures/cases_tmpl.go b/operator/hack/discoveryfixtures/cases_tmpl.go new file mode 100644 index 000000000..d218b0e40 --- /dev/null +++ b/operator/hack/discoveryfixtures/cases_tmpl.go @@ -0,0 +1,101 @@ +// The document-shape cases, §13 of the document, and the table every family is +// collected into. +// +// A template case is about the half of a draft no probe reports: the cluster +// block, its name, and the fields a reviewer is being asked to approve without +// any reading behind them. The fleet under each is therefore the plainest one +// available, so that nothing in the document distracts from the block itself. + +package main + +import ( + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" +) + +func templateCases() map[string]Case { + plain := func() []nodeprobe.Report { + return fleet(3, func(_ int, name string) nodeprobe.Report { + return host(name, cpu(1, 16, 2), disks(layoutA.disks(3*tb)...)) + }) + } + with := func(spec *simplyblockv1alpha2.DiscoverSpec) Case { + reports := plain() + return Case{Family: "tmpl", Reports: reports, Nodes: kubeFleet(reports), Discover: spec} + } + + // A two-socket fleet, which is where the absent socket layout costs a + // cluster the second half of every worker. + twoSocket := fleet(3, func(_ int, name string) nodeprobe.Report { + return host(name, disks(spread(3*tb, 2, 2)...)) + }) + + cases := map[string]Case{ + "TMPL-01": with(nil), + "TMPL-02": with(nil), + "TMPL-03": with(nil), + "TMPL-05": with(&simplyblockv1alpha2.DiscoverSpec{ClusterRef: "existing-cluster"}), + "TMPL-06": {Family: "tmpl", Reports: twoSocket, Nodes: kubeFleet(twoSocket)}, + "TMPL-07": with(nil), + "TMPL-08": with(&simplyblockv1alpha2.DiscoverSpec{ + DeviceFilter: &simplyblockv1alpha2.DeviceFilter{ + PcieDenyList: []string{"0000:b0:00.0"}, + DriveSizeRange: "1T-8T", + }, + }), + } + + // A run whose name pushes the cluster it proposes past what a + // StorageCluster name may be. + cases["TMPL-04"] = with(&simplyblockv1alpha2.DiscoverSpec{ + ConfigName: "discovered-the-quarterly-storage-expansion-for-the-frankfurt-racks", + }) + + slugs := map[string]string{ + "TMPL-01": "every-draft-formats-its-drives", + "TMPL-02": "the-subsystem-count-is-not-a-reading", + "TMPL-03": "a-generated-config-name", + "TMPL-04": "a-cluster-name-past-its-limit", + "TMPL-05": "a-growth-document", + "TMPL-06": "a-two-socket-fleet-and-no-socket-layout", + "TMPL-07": "the-environment-is-copied-through", + "TMPL-08": "the-filter-is-resolved-and-not-carried", + } + gaps := map[string]string{"TMPL-04": "G-11", "TMPL-06": "G-12"} + for id, slug := range slugs { + entry := cases[id] + entry.Slug = slug + entry.Gap = gaps[id] + cases[id] = entry + } + + // The environment case says what it varies, since the value is on the run + // rather than in the fleet. + openshift := cases["TMPL-07"] + openshift.Environment = simplyblockv1alpha2.KubernetesEnvironmentOpenShift + cases["TMPL-07"] = openshift + + return cases +} + +// allCases is every family's table, merged. +// +// A case identifier that two families claim is a mistake worth stopping for: +// the second would silently overwrite the first and one of the two directories +// would never be written. +func allCases() map[string]Case { + out := map[string]Case{} + for _, family := range []map[string]Case{ + devCases(), numaCases(), sizeCases(), pciCases(), fleetCases(), + netCases(), roleCases(), filterCases(), heldCases(), failCases(), + configMapCases(), templateCases(), + } { + for id, c := range family { + if _, repeated := out[id]; repeated { + panic("two families claim " + id) + } + out[id] = c + } + } + return out +} diff --git a/operator/hack/discoveryfixtures/document.go b/operator/hack/discoveryfixtures/document.go new file mode 100644 index 000000000..b12f7417f --- /dev/null +++ b/operator/hack/discoveryfixtures/document.go @@ -0,0 +1,95 @@ +// Reading the case rows out of the document. +// +// The document is the specification and this generator is its reader, which is +// why the rows are parsed rather than restated here: a mutation described one +// way in the document and another way in a fixture is a case nobody can check, +// and the two drift the first time a row is edited. +// +// The parse is deliberately shallow. It takes the identifier, the two prose +// columns, and the harness out of any table row that opens with a case +// identifier, and it knows nothing about which section the row was in. + +package main + +import ( + "fmt" + "os" + "regexp" + "strings" +) + +// Row is one case as the document states it. +type Row struct { + // ID is the case identifier, such as NET-13. + ID string + + // Mutation and Expected are the row's two prose columns, in the document's + // own words. + Mutation string + Expected string + + // Harness is CM for a case driven from this directory, GO for one that + // substitutes a Planner seam and has no directory to be driven from. + Harness string +} + +// caseRow matches a table row whose first cell is a case identifier. The +// families are listed rather than matched loosely, so that a table of something +// else that happens to start with a dashed word is not read as a case. +var caseRow = regexp.MustCompile( + `^\|\s*((?:DEV|NUMA|SIZE|PCI|FLEET|NET|ROLE|FILT|HELD|FAIL|CM|TMPL)-\d+)\s*\|(.*)$`) + +// readRows reads every case the document states, keyed by identifier. +func readRows(path string) (map[string]Row, error) { + content, err := os.ReadFile(path) + if err != nil { + return nil, fmt.Errorf("read the case document: %w", err) + } + + rows := map[string]Row{} + for _, line := range strings.Split(string(content), "\n") { + match := caseRow.FindStringSubmatch(line) + if match == nil { + continue + } + + cells := strings.Split(strings.TrimSuffix(strings.TrimSpace(match[2]), "|"), "|") + if len(cells) < 3 { + return nil, fmt.Errorf("the row for %s has %d cells, want the mutation, the expectation, and the harness", + match[1], len(cells)+1) + } + row := Row{ + ID: match[1], + Mutation: strings.TrimSpace(cells[0]), + Expected: strings.TrimSpace(cells[1]), + Harness: strings.Trim(strings.TrimSpace(cells[2]), "`"), + } + if seen, repeated := rows[row.ID]; repeated { + return nil, fmt.Errorf("%s appears twice: %q and %q", row.ID, seen.Mutation, row.Mutation) + } + rows[row.ID] = row + } + if len(rows) == 0 { + return nil, fmt.Errorf("%s holds no case rows", path) + } + return rows, nil +} + +// describe renders the case.md a directory carries: what the case is, in the +// document's own words, so that a failure is readable without the document +// open beside it. +func describe(row Row, c Case) string { + var out strings.Builder + fmt.Fprintf(&out, "# %s\n\n", row.ID) + fmt.Fprintf(&out, "**Mutation.** %s\n\n", row.Mutation) + fmt.Fprintf(&out, "**Expected.** %s\n\n", row.Expected) + fmt.Fprintf(&out, "**Harness.** `%s`\n", row.Harness) + if c.Gap != "" { + fmt.Fprintf(&out, "\n**Gap.** %s. This case records what the generator does today, "+ + "so that the day it changes the diff is the finding.\n", c.Gap) + } + if c.Note != "" { + fmt.Fprintf(&out, "\n**Note.** %s\n", c.Note) + } + return out.String() +} diff --git a/operator/hack/discoveryfixtures/fixtures.go b/operator/hack/discoveryfixtures/fixtures.go new file mode 100644 index 000000000..2066e21fe --- /dev/null +++ b/operator/hack/discoveryfixtures/fixtures.go @@ -0,0 +1,647 @@ +// The builders every case is written with, and the objects a case directory +// holds. +// +// A case says one thing about a fleet and inherits the rest, so the builders +// default to a plausible machine and take options for the part that matters: +// two memory nodes, sixteen cores, a data NIC on each, which is the host this +// product was developed against. A case that is about a memory layout says only +// the layout, and one about a NIC says only the NIC. +// +// The objects are built through the same types the probe and the operator use, +// so a fixture that does not decode is a fixture that never gets written. + +package main + +import ( + "fmt" + "os" + "strconv" + "strings" + "time" + + corev1 "k8s.io/api/core/v1" + "k8s.io/apimachinery/pkg/api/resource" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "sigs.k8s.io/yaml" + + "github.com/simplyblock/atlas/blockdev" + "github.com/simplyblock/atlas/inventory" + "github.com/simplyblock/atlas/ptr" + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" +) + +// Sizes a case names a disk or a machine by. +const ( + mb = uint64(1) << 20 + gb = uint64(1) << 30 + tb = uint64(1) << 40 +) + +// namespace is where every case's objects live, matching the chart's default. +const namespace = "simplyblock" + +// probedAt is when every report says it was collected. +// +// It is fixed rather than taken from the clock so that regenerating the tree +// produces no diff: a timestamp that moved on every run would put 171 changed +// files in front of a reviewer looking for the one case that actually changed. +var probedAt = metav1.NewTime(time.Date(2026, time.September, 18, 9, 0, 0, 0, time.UTC)) + +// Case is one case's inputs. +type Case struct { + // Family is the directory the case is filed under, and Slug names it + // within that directory. + Family string + Slug string + + // Gap is the finding this case records today's behavior for, when it + // records one. + Gap string + + // Note is anything a reader of the directory needs that the document's row + // does not say. + Note string + + // Reports are the workers the probes reported, and Nodes is what + // Kubernetes says about the same machines. Nodes may be empty, which is a + // run whose workers carry no node object. + Reports []nodeprobe.Report + Nodes []corev1.Node + + // Workers overrides which workers the run settled on. Empty is every + // worker a report was written for, which is what an ordinary run produces. + Workers []string + + // Discover is the run's own configuration, and Environment is what the run + // concluded the distribution is. + Discover *simplyblockv1alpha2.DiscoverSpec + Environment simplyblockv1alpha2.KubernetesEnvironment + + // Amend rewrites the ConfigMaps after they are built, for a case that is + // about the object rather than about the fleet in it. + Amend func(run string, maps []*corev1.ConfigMap) []*corev1.ConfigMap +} + +// Directory is the case's directory name: the identifier and a slug of the +// mutation, so that a failure names the case without anybody looking it up. +func (c Case) Directory(id string) string { + return strings.ToLower(id) + "-" + c.Slug +} + +// runName is what the run is called, and is both the OperatorOps name and the +// label its reports are selected by. +func runName(id string) string { return "discover-" + strings.ToLower(id) } + +// operatorOpsFor is the run a case is driven from. +func operatorOpsFor(id string, c Case) *simplyblockv1alpha2.OperatorOps { + workers := c.Workers + if workers == nil { + for _, report := range c.Reports { + workers = append(workers, report.Node) + } + } + + environment := c.Environment + if environment == "" { + environment = simplyblockv1alpha2.KubernetesEnvironmentK3s + } + + return &simplyblockv1alpha2.OperatorOps{ + TypeMeta: metav1.TypeMeta{ + APIVersion: simplyblockv1alpha2.GroupVersion.String(), + Kind: "OperatorOps", + }, + ObjectMeta: metav1.ObjectMeta{Name: runName(id), Namespace: namespace}, + Spec: simplyblockv1alpha2.OperatorOpsSpec{ + Action: simplyblockv1alpha2.OperatorOpsActionDiscover, + Discover: c.Discover, + }, + Status: simplyblockv1alpha2.OperatorOpsStatus{ + Workers: workers, + Environment: environment, + }, + } +} + +// reportConfigMap is the object a probe would have written for one report. +func reportConfigMap(id string, report nodeprobe.Report) (*corev1.ConfigMap, error) { + cm, err := nodeprobe.ConfigMap(namespace, runName(id), nil, report) + if err != nil { + return nil, err + } + cm.TypeMeta = metav1.TypeMeta{APIVersion: "v1", Kind: "ConfigMap"} + return cm, nil +} + +// --- reports --------------------------------------------------------------- + +// hostOpt changes one thing about the default machine. +type hostOpt func(*nodeprobe.Report) + +// host is a worker: two memory nodes, eight cores each with two threads, 256 +// GiB of memory, sixteen 1 GiB huge pages free per node, and a 25G data NIC on +// each node holding a routable address. +// +// Everything a case does not name, it inherits from here. +func host(name string, opts ...hostOpt) nodeprobe.Report { + report := nodeprobe.Report{ + Version: nodeprobe.ReportVersion, + Node: name, + ProbedAt: probedAt, + CPU: cpuTopology(2, 8, 2), + Memory: memoryOf(256*gb, 240*gb, 128*gb, 128*gb), + HugePages: []nodeprobe.HugePagePool{pagePool(gb, 16, 16)}, + Interfaces: []nodeprobe.Interface{ + nic("eth0", inventory.LinkPhysical, at(25000), on(0), holding(managementAddress(name))), + nic("eth1", inventory.LinkPhysical, at(25000), on(1), holding(dataAddress(name))), + }, + } + for _, opt := range opts { + opt(&report) + } + return report +} + +// managementAddress and dataAddress are the two addresses a worker holds, keyed +// off the digits in its name so that a fleet's machines do not all claim one +// address and the node object can name the same one the report does. +func managementAddress(node string) string { return fmt.Sprintf("10.10.10.%d", hostNumber(node)) } +func dataAddress(node string) string { return fmt.Sprintf("10.10.11.%d", hostNumber(node)) } + +// hostNumber is the trailing number of a worker's name, and 1 for a name that +// ends in none. +func hostNumber(node string) int { + digits := "" + for i := len(node) - 1; i >= 0 && node[i] >= '0' && node[i] <= '9'; i-- { + digits = string(node[i]) + digits + } + number, err := strconv.Atoi(digits) + if err != nil || number == 0 { + return 1 + } + return number +} + +// cpuTopology is a processor inventory of the shape given: one memory node per +// socket, with the CPU identifiers handed out node by node the way a kernel +// numbers them. +func cpuTopology(sockets, coresPerNode, threadsPerCore int) nodeprobe.CPU { + cpu := nodeprobe.CPU{ + Sockets: sockets, + PhysicalCores: sockets * coresPerNode, + ThreadsPerCore: threadsPerCore, + HyperThreading: threadsPerCore > 1, + OnlineCPUs: sockets * coresPerNode * threadsPerCore, + } + next := 0 + for node := 0; node < sockets; node++ { + ids := make([]int, 0, coresPerNode*threadsPerCore) + for i := 0; i < coresPerNode*threadsPerCore; i++ { + ids = append(ids, next) + next++ + } + cpu.NUMANodes = append(cpu.NUMANodes, nodeprobe.NUMACPUs{ + Node: node, OnlineCPUs: ids, PhysicalCores: coresPerNode, + }) + } + return cpu +} + +// `cpu` sets the processor inventory. +func cpu(sockets, coresPerNode, threadsPerCore int) hostOpt { + return func(r *nodeprobe.Report) { r.CPU = cpuTopology(sockets, coresPerNode, threadsPerCore) } +} + +// cpuNodes sets an asymmetric processor tree, for a machine whose memory nodes +// do not carry the same number of cores. +func cpuNodes(threadsPerCore int, coresPerNode ...int) hostOpt { + return func(r *nodeprobe.Report) { + out := nodeprobe.CPU{Sockets: len(coresPerNode), ThreadsPerCore: threadsPerCore, + HyperThreading: threadsPerCore > 1} + next := 0 + for node, cores := range coresPerNode { + ids := make([]int, 0, cores*threadsPerCore) + for i := 0; i < cores*threadsPerCore; i++ { + ids = append(ids, next) + next++ + } + out.PhysicalCores += cores + out.OnlineCPUs += len(ids) + out.NUMANodes = append(out.NUMANodes, nodeprobe.NUMACPUs{ + Node: node, OnlineCPUs: ids, PhysicalCores: cores, + }) + } + r.CPU = out + } +} + +// noNUMATopology is a machine whose cores the probe read and could not attribute +// to a memory node. +func noNUMATopology(sockets, coresPerNode, threadsPerCore int) hostOpt { + return func(r *nodeprobe.Report) { + out := cpuTopology(sockets, coresPerNode, threadsPerCore) + out.NUMANodes = nil + r.CPU = out + } +} + +// noCPUTopology is a machine whose processor tree the probe could not read. +func noCPUTopology() hostOpt { + return func(r *nodeprobe.Report) { r.CPU = nodeprobe.CPU{} } +} + +// memoryOf is a memory reading, with the per-node split given. +func memoryOf(total, available uint64, perNode ...uint64) nodeprobe.Memory { + out := nodeprobe.Memory{TotalBytes: total, AvailableBytes: available, FreeBytes: available / 4} + for node, bytes := range perNode { + out.NUMANodes = append(out.NUMANodes, nodeprobe.NUMAMemory{ + Node: node, TotalBytes: bytes, FreeBytes: bytes / 2, + }) + } + return out +} + +// mem sets the machine's memory. +func mem(total, available uint64, perNode ...uint64) hostOpt { + return func(r *nodeprobe.Report) { r.Memory = memoryOf(total, available, perNode...) } +} + +// reserved records memory already held as huge pages, which meminfo reports +// separately because it is out of the general pool. +func reserved(bytes uint64) hostOpt { + return func(r *nodeprobe.Report) { r.Memory.HugePagesBytes = bytes } +} + +// swap records the swap a host has and how much of it is gone. +func swap(total, free uint64) hostOpt { + return func(r *nodeprobe.Report) { r.Memory.SwapTotalBytes, r.Memory.SwapFreeBytes = total, free } +} + +// pagePool is one huge-page size's allocation, with the per-node split given as +// the pages free on each. Every page is counted as free, which is a host whose +// reservation nothing has taken yet. +func pagePool(size uint64, freePerNode ...uint64) nodeprobe.HugePagePool { + pool := nodeprobe.HugePagePool{SizeBytes: size} + for node, free := range freePerNode { + pool.Total += free + pool.Free += free + pool.NUMANodes = append(pool.NUMANodes, nodeprobe.NUMAHugePages{ + Node: node, Total: free, Free: free, + }) + } + return pool +} + +// pages sets the huge-page pools. +func pages(pools ...nodeprobe.HugePagePool) hostOpt { + return func(r *nodeprobe.Report) { r.HugePages = pools } +} + +// noPages is a host with nothing set aside, which is every host before anybody +// prepared it. +func noPages() hostOpt { + return func(r *nodeprobe.Report) { r.HugePages = nil } +} + +// ifaces replaces the machine's network interfaces. +func ifaces(list ...nodeprobe.Interface) hostOpt { + return func(r *nodeprobe.Report) { r.Interfaces = list } +} + +// disks sets the machine's block devices. +func disks(list ...nodeprobe.Device) hostOpt { + return func(r *nodeprobe.Report) { r.Devices = list } +} + +// controllers sets the NVMe controllers on the machine's PCI bus, which is what +// tells a reader about disks the kernel presents no block device for. +func controllers(list ...nodeprobe.Controller) hostOpt { + return func(r *nodeprobe.Report) { r.NVMeControllers = list } +} + +// unreadable records what the probe could not read. +func unreadable(sentences ...string) hostOpt { + return func(r *nodeprobe.Report) { r.Unreadable = sentences } +} + +// --- interfaces ------------------------------------------------------------ + +// ifaceOpt changes one thing about an interface. +type ifaceOpt func(*nodeprobe.Interface) + +// `nic` is one network interface of the kind given, up and on memory node 0 +// unless a case says otherwise. +func nic(name string, kind inventory.LinkKind, opts ...ifaceOpt) nodeprobe.Interface { + iface := nodeprobe.Interface{ + Name: name, + Kind: string(kind), + State: "up", + MTU: 1500, + NUMANode: inventory.NUMANodeUnknown, + Virtual: kind != inventory.LinkPhysical, + Bridge: kind == inventory.LinkBridge, + Loopback: kind == inventory.LinkLoopback, + } + if kind == inventory.LinkPhysical { + iface.NUMANode = 0 + iface.Driver = "mlx5_core" + } + for _, opt := range opts { + opt(&iface) + } + return iface +} + +// at is the interface's negotiated link speed in megabits per second. +func at(speed int) ifaceOpt { + return func(i *nodeprobe.Interface) { i.SpeedMbps = speed } +} + +// holding is the addresses the interface carries. +func holding(addresses ...string) ifaceOpt { + return func(i *nodeprobe.Interface) { i.Addresses = addresses } +} + +// on is the memory node the interface's hardware hangs off. +func on(node int) ifaceOpt { + return func(i *nodeprobe.Interface) { i.NUMANode = node } +} + +// linkState overrides the kernel's operstate for the interface. +func linkState(state string) ifaceOpt { + return func(i *nodeprobe.Interface) { i.State = state } +} + +// over is what the interface is built on: the members of a bond or a bridge, or +// the parent of a VLAN. +func over(lower ...string) ifaceOpt { + return func(i *nodeprobe.Interface) { i.Lower = lower } +} + +// under is what is built on the interface. +func under(upper ...string) ifaceOpt { + return func(i *nodeprobe.Interface) { i.Upper = upper } +} + +// frames is the interface's MTU, for a case about jumbo frames. +func frames(mtu int) ifaceOpt { + return func(i *nodeprobe.Interface) { i.MTU = mtu } +} + +// slotted is the PCI address of the interface's hardware. +func slotted(address string) ifaceOpt { + return func(i *nodeprobe.Interface) { i.PCIAddress = address } +} + +// unkinded strips the kind, which is a probe that could not read the device +// type rather than an older schema. +func unkinded() ifaceOpt { + return func(i *nodeprobe.Interface) { i.Kind = "" } +} + +// --- devices --------------------------------------------------------------- + +// devOpt changes one thing about a block device. +type devOpt func(*nodeprobe.Device) + +// `nvme` is one free NVMe disk in a slot, on a memory node. +func nvme(name, address string, node int, size uint64, opts ...devOpt) nodeprobe.Device { + return device(nodeprobe.Device{ + Name: name, + Path: "/dev/" + name, + PCIAddress: address, + SizeBytes: size, + Kind: string(blockdev.KindDisk), + Transport: string(blockdev.TransportNVMe), + Model: "SAMSUNG MZQL23T8HCLS-00A07", + NUMANode: node, + Available: true, + Content: "Blank", + }, opts...) +} + +// blk is one free disk of the other class: a virtio disk with a path and no PCI +// address a draft could name it by. +func blk(name string, node int, size uint64, opts ...devOpt) nodeprobe.Device { + return device(nodeprobe.Device{ + Name: name, + Path: "/dev/" + name, + SizeBytes: size, + Kind: string(blockdev.KindDisk), + Transport: string(blockdev.TransportVirtio), + NUMANode: node, + Available: true, + Content: "Blank", + }, opts...) +} + +// device applies a device's options. +func device(d nodeprobe.Device, opts ...devOpt) nodeprobe.Device { + for _, opt := range opts { + opt(&d) + } + return d +} + +// refused is a device the probe declined, on the grounds given. +func refused(reasons ...blockdev.Reason) devOpt { + return func(d *nodeprobe.Device) { + d.Available = false + d.Content = "" + for _, reason := range reasons { + d.Rejections = append(d.Rejections, nodeprobe.Rejection{ + Reason: string(reason), Detail: "as the probe found it", + }) + } + } +} + +// modeled overrides the device's model string. +func modeled(model string) devOpt { + return func(d *nodeprobe.Device) { d.Model = model } +} + +// spinning marks the device a rotational one. +func spinning() devOpt { + return func(d *nodeprobe.Device) { d.Rotational = true; d.Model = "SEAGATE ST16000NM" } +} + +// partOf makes the device a partition rather than a whole disk. +func partOf() devOpt { + return func(d *nodeprobe.Device) { d.Kind = string(blockdev.KindPartition) } +} + +// looped makes the device a loopback device, which is what a machine presents +// dozens of and a draft can use none of. +func looped() devOpt { + return func(d *nodeprobe.Device) { + d.Kind = string(blockdev.KindLoop) + d.Transport = "" + d.PCIAddress = "" + } +} + +// transported overrides the bus the device sits on. +func transported(transport blockdev.Transport) devOpt { + return func(d *nodeprobe.Device) { d.Transport = string(transport) } +} + +// controller is one NVMe controller on the PCI bus, bound to the driver given. +func controller(address, driver string, node int, opts ...func(*nodeprobe.Controller)) nodeprobe.Controller { + c := nodeprobe.Controller{ + Address: address, Driver: driver, NUMANode: node, + Vendor: "0x144d", Product: "0xa80a", + // Checked and found free, which is the state a draft may claim. + InUse: ptr.To(false), + } + for _, opt := range opts { + opt(&c) + } + return c +} + +// held marks a controller something is driving, which is a disk in service +// rather than one to reclaim. +func held() func(*nodeprobe.Controller) { + return func(c *nodeprobe.Controller) { c.InUse = ptr.To(true) } +} + +// unchecked is a controller the probe could not ask about, which is neither +// held nor free. +func unchecked() func(*nodeprobe.Controller) { + return func(c *nodeprobe.Controller) { c.InUse = nil } +} + +// --- Kubernetes nodes ------------------------------------------------------ + +// nodeOpt changes one thing about a node object. +type nodeOpt func(*corev1.Node) + +// kubeNode is what Kubernetes says about a worker: 32 cores and 256 GiB, of +// which the kubelet holds some back. +func kubeNode(name string, opts ...nodeOpt) corev1.Node { + node := corev1.Node{ + TypeMeta: metav1.TypeMeta{APIVersion: "v1", Kind: "Node"}, + ObjectMeta: metav1.ObjectMeta{Name: name, Labels: map[string]string{}}, + Status: corev1.NodeStatus{ + Capacity: corev1.ResourceList{ + corev1.ResourceCPU: resource.MustParse("32"), + corev1.ResourceMemory: resource.MustParse("256Gi"), + }, + Allocatable: corev1.ResourceList{ + corev1.ResourceCPU: resource.MustParse("31500m"), + corev1.ResourceMemory: resource.MustParse("250Gi"), + }, + }, + } + for _, opt := range opts { + opt(&node) + } + return node +} + +// role labels the node with what it is for, in Kubernetes' own convention. +func role(name string) nodeOpt { + return func(n *corev1.Node) { n.Labels["node-role.kubernetes.io/"+name] = "" } +} + +// labeled puts one label on the node. +func labeled(key, value string) nodeOpt { + return func(n *corev1.Node) { n.Labels[key] = value } +} + +// reachableAt is the address the cluster reaches the machine on. +func reachableAt(address string) nodeOpt { + return func(n *corev1.Node) { + n.Status.Addresses = append(n.Status.Addresses, + corev1.NodeAddress{Type: corev1.NodeInternalIP, Address: address}) + } +} + +// sized overrides what the kubelet found and what it will schedule against. +func sized(capacityCPU, capacityMemory, allocatableCPU, allocatableMemory string) nodeOpt { + return func(n *corev1.Node) { + n.Status.Capacity[corev1.ResourceCPU] = resource.MustParse(capacityCPU) + n.Status.Capacity[corev1.ResourceMemory] = resource.MustParse(capacityMemory) + n.Status.Allocatable[corev1.ResourceCPU] = resource.MustParse(allocatableCPU) + n.Status.Allocatable[corev1.ResourceMemory] = resource.MustParse(allocatableMemory) + } +} + +// schedulableHugePages is a huge-page size the kubelet will schedule against. +func schedulableHugePages(size, quantity string) nodeOpt { + return func(n *corev1.Node) { + name := corev1.ResourceName("hugepages-" + size) + n.Status.Capacity[name] = resource.MustParse(quantity) + n.Status.Allocatable[name] = resource.MustParse(quantity) + } +} + +// cordoned marks the node unschedulable. +func cordoned() nodeOpt { + return func(n *corev1.Node) { n.Spec.Unschedulable = true } +} + +// tainted puts a taint on the node. +func tainted(key, value string, effect corev1.TaintEffect) nodeOpt { + return func(n *corev1.Node) { + n.Spec.Taints = append(n.Spec.Taints, corev1.Taint{Key: key, Value: value, Effect: effect}) + } +} + +// --- fleets ---------------------------------------------------------------- + +// fleet is n workers named worker-01 upward, each built by the function given. +// +// The names are padded because the draft's groups are numbered by their first +// worker's name and the ordering is lexicographic: worker-10 sorts before +// worker-2, and a fixture that made a reviewer work that out would be a fixture +// about nothing. +func fleet(n int, build func(index int, name string) nodeprobe.Report) []nodeprobe.Report { + reports := make([]nodeprobe.Report, 0, n) + for i := 1; i <= n; i++ { + name := fmt.Sprintf("worker-%02d", i) + reports = append(reports, build(i, name)) + } + return reports +} + +// kubeFleet is the node objects for a fleet, with each worker reachable on its +// own address. +func kubeFleet(reports []nodeprobe.Report, opts ...nodeOpt) []corev1.Node { + nodes := make([]corev1.Node, 0, len(reports)) + for _, report := range reports { + all := append([]nodeOpt{reachableAt(managementAddress(report.Node))}, opts...) + nodes = append(nodes, kubeNode(report.Node, all...)) + } + return nodes +} + +// --- writing --------------------------------------------------------------- + +// writeYAML writes one object. +func writeYAML(path string, object any) error { + out, err := yaml.Marshal(object) + if err != nil { + return fmt.Errorf("render %s: %w", path, err) + } + return os.WriteFile(path, out, 0o644) +} + +// writeDocuments writes a list of objects as a multi-document file, which is +// what a reviewer expects of a file called nodes.yaml. +func writeDocuments[T any](path string, objects []T) error { + var out strings.Builder + for i := range objects { + rendered, err := yaml.Marshal(objects[i]) + if err != nil { + return fmt.Errorf("render %s: %w", path, err) + } + if i > 0 { + out.WriteString("---\n") + } + out.Write(rendered) + } + return os.WriteFile(path, []byte(out.String()), 0o644) +} diff --git a/operator/hack/discoveryfixtures/main.go b/operator/hack/discoveryfixtures/main.go new file mode 100644 index 000000000..35d244345 --- /dev/null +++ b/operator/hack/discoveryfixtures/main.go @@ -0,0 +1,243 @@ +// Writes the discovery generator's test fixtures: one directory per case in +// discovery-generator-test-cases.md, holding the probe reports, the Kubernetes +// nodes, and the OperatorOps that case is run from. +// +// It is a program rather than a table inside a test because the fixtures are +// the thing under review. A reviewer reading a directory of ConfigMaps is +// reading the fleet a case describes. The same cases expressed as Go literals +// inside a test file are readable only to somebody already holding the +// harness. The generated tree is committed, and this exists to write it again +// when a case changes rather than to be run by the tests. +// +// The case list is read from the document, and the inputs from the tables in +// this package, so neither can gain a case the other does not have: a row with +// no builder and a builder with no row are both errors that stop the run. +// +// Only the inputs are written. What each case should produce is settled after +// the tree is reviewed, which is the order §15 of the document sets out. +package main + +import ( + "flag" + "fmt" + "os" + "path/filepath" + "sort" + "strings" + + corev1 "k8s.io/api/core/v1" +) + +func main() { + document := flag.String("document", "../discovery-generator-test-cases.md", + "the case document to read the rows from") + out := flag.String("out", "internal/controllers/deployment/testdata/discovery", + "the directory to write the case tree into") + flag.Parse() + + if err := run(*document, *out); err != nil { + fmt.Fprintln(os.Stderr, "discoveryfixtures:", err) + os.Exit(1) + } +} + +func run(document, out string) error { + rows, err := readRows(document) + if err != nil { + return err + } + + builders := allCases() + if err := reconcile(rows, builders); err != nil { + return err + } + + // Directories no case claims any more are removed, so a renamed or + // withdrawn case leaves nothing behind for a harness to keep loading. What + // is deliberately kept is the recorded expectation of a case that still + // exists: regenerating changes the inputs, and the point of the recording + // is that the test then fails with the diff rather than quietly agreeing + // with whatever the new inputs produce. + if err := pruneWithdrawn(out, rows, builders); err != nil { + return err + } + + ids := make([]string, 0, len(builders)) + for id := range builders { + ids = append(ids, id) + } + sort.Strings(ids) + + written := 0 + for _, id := range ids { + if err := writeCase(out, rows[id], builders[id]); err != nil { + return fmt.Errorf("%s: %w", id, err) + } + written++ + } + fmt.Printf("wrote %d cases into %s\n", written, out) + return nil +} + +// reconcile refuses a document and a builder table that do not describe the +// same set of cases, which is the one way this generator can go quietly wrong. +func reconcile(rows map[string]Row, builders map[string]Case) error { + var missing, extra []string + for id := range rows { + if _, built := builders[id]; !built { + missing = append(missing, id) + } + } + for id := range builders { + if _, documented := rows[id]; !documented { + extra = append(extra, id) + } + } + sort.Strings(missing) + sort.Strings(extra) + + var problems []string + if len(missing) > 0 { + problems = append(problems, fmt.Sprintf( + "%d case(s) in the document have no fixture: %s", + len(missing), strings.Join(missing, ", "))) + } + if len(extra) > 0 { + problems = append(problems, fmt.Sprintf( + "%d fixture(s) are in no row of the document: %s", + len(extra), strings.Join(extra, ", "))) + } + if len(problems) > 0 { + return fmt.Errorf("%s", strings.Join(problems, "; ")) + } + return nil +} + +// pruneWithdrawn removes every case directory no builder claims, and the input +// files of the ones that remain. +// +// The inputs are cleared rather than overwritten because a case whose fleet +// shrank would otherwise keep the report of a worker it no longer has. The +// expectation files are left alone. +func pruneWithdrawn(out string, rows map[string]Row, builders map[string]Case) error { + wanted := map[string]struct{}{} + for id, c := range builders { + wanted[filepath.Join(out, c.Family, c.Directory(rows[id].ID))] = struct{}{} + } + + entries, err := os.ReadDir(out) + if err != nil { + if os.IsNotExist(err) { + return nil + } + return fmt.Errorf("read %s: %w", out, err) + } + + for _, family := range entries { + if !family.IsDir() { + continue + } + familyDir := filepath.Join(out, family.Name()) + cases, err := os.ReadDir(familyDir) + if err != nil { + return fmt.Errorf("read %s: %w", familyDir, err) + } + for _, entry := range cases { + dir := filepath.Join(familyDir, entry.Name()) + if _, keep := wanted[dir]; !keep { + if err := os.RemoveAll(dir); err != nil { + return fmt.Errorf("remove %s: %w", dir, err) + } + continue + } + if err := clearInputs(dir); err != nil { + return err + } + } + } + return nil +} + +// clearInputs removes the files this generator owns, leaving the recorded +// expectations in place. +func clearInputs(dir string) error { + for _, name := range []string{caseFile, opsFile, nodesFile} { + if err := os.Remove(filepath.Join(dir, name)); err != nil && !os.IsNotExist(err) { + return fmt.Errorf("remove %s: %w", filepath.Join(dir, name), err) + } + } + reports := filepath.Join(dir, reportsDir) + if err := os.RemoveAll(reports); err != nil { + return fmt.Errorf("remove %s: %w", reports, err) + } + return nil +} + +// The files this generator owns. Everything else in a case directory is the +// harness's recording and is never touched here. +const ( + caseFile = "case.md" + opsFile = "ops.yaml" + nodesFile = "nodes.yaml" + reportsDir = "reports" +) + +// writeCase writes one case directory: what it is, and the objects it is run +// from. +func writeCase(out string, row Row, c Case) error { + dir := filepath.Join(out, c.Family, c.Directory(row.ID)) + if err := os.MkdirAll(dir, 0o755); err != nil { + return err + } + + if err := os.WriteFile(filepath.Join(dir, caseFile), []byte(describe(row, c)), 0o644); err != nil { + return err + } + + // A GO case substitutes a Planner seam the controller never does, so it has + // no objects to be run from and its directory carries the row alone. + if row.Harness == "GO" { + return nil + } + + if err := writeYAML(filepath.Join(dir, opsFile), operatorOpsFor(row.ID, c)); err != nil { + return err + } + if len(c.Nodes) > 0 { + if err := writeDocuments(filepath.Join(dir, nodesFile), c.Nodes); err != nil { + return err + } + } + + maps := make([]*corev1.ConfigMap, 0, len(c.Reports)) + for _, report := range c.Reports { + cm, err := reportConfigMap(row.ID, report) + if err != nil { + return err + } + maps = append(maps, cm) + } + if c.Amend != nil { + maps = c.Amend(runName(row.ID), maps) + } + if len(maps) == 0 { + return nil + } + + if err := os.MkdirAll(filepath.Join(dir, reportsDir), 0o755); err != nil { + return err + } + for i, cm := range maps { + // The file is named for the worker rather than for the object, whose + // own name is a digest: a directory whose every entry is + // `sb-nodeprobe-` is a directory nobody can read. + name := fmt.Sprintf("report-%d", i) + if i < len(c.Reports) && c.Reports[i].Node != "" { + name = c.Reports[i].Node + } + if err := writeYAML(filepath.Join(dir, reportsDir, name+".yaml"), cm); err != nil { + return err + } + } + return nil +} diff --git a/operator/internal/controllers/deployment/discovery_cases_test.go b/operator/internal/controllers/deployment/discovery_cases_test.go new file mode 100644 index 000000000..51869865d --- /dev/null +++ b/operator/internal/controllers/deployment/discovery_cases_test.go @@ -0,0 +1,379 @@ +// The case tree under testdata/discovery, driven through the discovery run. +// +// Each directory there is a fleet: the probe reports, the node objects, and the +// OperatorOps a run is started from. This drives every one of them through the +// step that turns reports into a document, and compares what came out against +// what the directory says should. +// +// It goes through the reconciler rather than through the planner because the +// document is the product. What a reviewer approves is a +// ClusterDeploymentConfig, and the parts of it the planner has no hand in are +// exactly the parts a fixture is worth having for: the name, the environment, +// the cluster block, and whether the run refused outright. +// +// The expectations are files rather than assertions in Go. A case that produces +// a different document produces a diff a reviewer reads, which is the only form +// in which 171 outcomes can be reviewed at all. The same run with -update +// rewrites them, so recording a new case is not transcription. +// +// What that buys is a record of today's behavior and nothing more. A recorded +// file is evidence that the output has not changed since somebody read it; it +// is not evidence that the output is right. Reading each one against the row in +// its case.md is the step that supplies the second, and it is a step no harness +// can do. + +package deployment + +import ( + "context" + "flag" + "fmt" + "os" + "path/filepath" + "sort" + "strings" + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/client-go/tools/events" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + "sigs.k8s.io/yaml" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// update rewrites every expectation from what the run produced, for recording a +// case rather than checking one. +var update = flag.Bool("update", false, + "rewrite the expected files of every discovery case from what the run produced") + +// caseRoot is the tree the cases live in. +const caseRoot = "testdata/discovery" + +// The files a case directory holds. The inputs are written by +// hack/discoveryfixtures and the expectations by this test. +const ( + caseFile = "case.md" + opsFile = "ops.yaml" + nodesFile = "nodes.yaml" + reportsDir = "reports" + expectedFile = "expected.yaml" + refusalsFile = "expected-refusals.txt" + notesFile = "expected-notes.txt" + errorFile = "expected-error.txt" +) + +// discoveryCase is one directory, loaded. +type discoveryCase struct { + // Name is the path under the case root, which is what a failure names. + Name string + + // Dir is where the case's files are. + Dir string + + // Ops is the run, and Objects is everything the run reads: the node + // objects and the ConfigMaps the probes wrote. + Ops *simplyblockv1alpha2.OperatorOps + Objects []client.Object +} + +func TestDiscoveryCases(t *testing.T) { + cases := loadDiscoveryCases(t) + if len(cases) == 0 { + t.Fatalf("no cases under %s: the tree is written by hack/discoveryfixtures", caseRoot) + } + + for _, c := range cases { + t.Run(c.Name, func(t *testing.T) { + outcome := runDiscoveryCase(t, c) + compareDiscoveryOutcome(t, c, outcome) + }) + } +} + +// outcome is everything one case produced. +type outcome struct { + // Config is the document the run wrote, or nil when it wrote none. + Config *simplyblockv1alpha2.ClusterDeploymentConfig + + // Refusals are the device and worker refusals the run reported, and Notes + // are the sentences accounting for the numbers in the cluster block. Both + // reach a reviewer as events, which is where they are read from. + Refusals []string + Notes []string + + // Failure is why the run wrote no document. + Failure string +} + +// runDiscoveryCase drives one case through the writing step. +func runDiscoveryCase(t *testing.T, c discoveryCase) outcome { + t.Helper() + + scheme := opsScheme(t) + ops := c.Ops.DeepCopy() + objects := append([]client.Object{ops}, c.Objects...) + + recorder := events.NewFakeRecorder(1024) + fakeClient := fake.NewClientBuilder(). + WithScheme(scheme). + WithObjects(objects...). + WithStatusSubresource(&simplyblockv1alpha2.OperatorOps{}). + Build() + + reconciler := &OperatorOpsReconciler{ + Client: fakeClient, Scheme: scheme, Recorder: recorder, ProbeImage: opsImage, + } + + _, err := reconciler.write(context.Background(), ops) + + out := outcome{} + if err != nil { + out.Failure = err.Error() + } + + // The document is read back from the client rather than taken from the + // reconciler, so that what is compared is what a reviewer would get. + var written simplyblockv1alpha2.ClusterDeploymentConfigList + if listErr := fakeClient.List(context.Background(), &written); listErr != nil { + t.Fatalf("list the documents the run wrote: %v", listErr) + } + switch len(written.Items) { + case 0: + case 1: + out.Config = &written.Items[0] + default: + t.Fatalf("the run wrote %d documents, and a run writes one", len(written.Items)) + } + + out.Refusals, out.Notes = drainEvents(recorder) + return out +} + +// drainEvents separates what the run said into the refusals and the notes, +// which are the two things a reviewer is owed and are read from different +// events. +func drainEvents(recorder *events.FakeRecorder) (refusals, notes []string) { + for { + select { + case line := <-recorder.Events: + message := eventMessage(line) + switch { + case strings.Contains(line, DeviceDeclined): + refusals = append(refusals, message) + case strings.Contains(line, ConfigWritten): + notes = append(notes, message) + } + default: + return refusals, notes + } + } +} + +// eventMessage is the message out of a recorded event, which the fake renders +// as a single line with the type and the reason ahead of it. +func eventMessage(line string) string { + parts := strings.SplitN(line, " ", 3) + if len(parts) < 3 { + return line + } + return parts[2] +} + +// compareDiscoveryOutcome checks one case against its recorded expectations, or +// records them. +func compareDiscoveryOutcome(t *testing.T, c discoveryCase, got outcome) { + t.Helper() + + document := "" + if got.Config != nil { + rendered, err := yaml.Marshal(comparableConfig(got.Config)) + if err != nil { + t.Fatalf("render the document: %v", err) + } + document = string(rendered) + } + + files := map[string]string{ + expectedFile: document, + refusalsFile: lines(got.Refusals), + notesFile: lines(got.Notes), + errorFile: trailing(got.Failure), + } + + if *update { + for name, content := range files { + path := filepath.Join(c.Dir, name) + if content == "" { + if err := os.Remove(path); err != nil && !os.IsNotExist(err) { + t.Fatalf("remove %s: %v", path, err) + } + continue + } + if err := os.WriteFile(path, []byte(content), 0o644); err != nil { + t.Fatalf("write %s: %v", path, err) + } + } + return + } + + for name, content := range files { + path := filepath.Join(c.Dir, name) + want, err := os.ReadFile(path) + if err != nil && !os.IsNotExist(err) { + t.Fatalf("read %s: %v", path, err) + } + if string(want) == content { + continue + } + t.Errorf("%s does not match.\n--- recorded ---\n%s\n--- produced ---\n%s\n"+ + "Run `go test ./internal/controllers/deployment/ -run TestDiscoveryCases -update` "+ + "to record this, and read the result against the row in case.md before committing it.", + name, want, content) + } +} + +// comparableConfig is the document with the fields a fake client invents +// stripped, so that a diff is about the draft rather than about the harness. +func comparableConfig( + config *simplyblockv1alpha2.ClusterDeploymentConfig, +) *simplyblockv1alpha2.ClusterDeploymentConfig { + out := config.DeepCopy() + out.ResourceVersion = "" + out.CreationTimestamp = metav1.Time{} + out.ManagedFields = nil + out.UID = "" + out.Generation = 0 + out.TypeMeta = metav1.TypeMeta{ + APIVersion: simplyblockv1alpha2.GroupVersion.String(), + Kind: "ClusterDeploymentConfig", + } + return out +} + +// lines renders a list as a file, one entry per line. +func lines(entries []string) string { + if len(entries) == 0 { + return "" + } + return strings.Join(entries, "\n") + "\n" +} + +// trailing renders a single value as a file, and nothing for an empty one. +func trailing(value string) string { + if value == "" { + return "" + } + return value + "\n" +} + +// loadDiscoveryCases reads every case directory under the root. +// +// A directory carrying a case.md is a case, which is what lets the tree be +// organized into families without the loader knowing what a family is. +func loadDiscoveryCases(t *testing.T) []discoveryCase { + t.Helper() + + var cases []discoveryCase + err := filepath.WalkDir(caseRoot, func(path string, entry os.DirEntry, err error) error { + if err != nil { + return err + } + if entry.IsDir() || entry.Name() != caseFile { + return nil + } + + dir := filepath.Dir(path) + body, err := os.ReadFile(path) + if err != nil { + return err + } + // A case whose harness is the discovery package has no objects to be + // driven from, and its directory carries the row alone. + if strings.Contains(string(body), "**Harness.** `GO`") { + return nil + } + + loaded, err := loadDiscoveryCase(dir) + if err != nil { + return fmt.Errorf("%s: %w", dir, err) + } + loaded.Name = strings.TrimPrefix(dir, caseRoot+string(filepath.Separator)) + cases = append(cases, loaded) + return nil + }) + if err != nil { + t.Fatalf("walk %s: %v", caseRoot, err) + } + + sort.Slice(cases, func(i, j int) bool { return cases[i].Name < cases[j].Name }) + return cases +} + +// loadDiscoveryCase reads one directory's objects. +func loadDiscoveryCase(dir string) (discoveryCase, error) { + out := discoveryCase{Dir: dir} + + opsRaw, err := os.ReadFile(filepath.Join(dir, opsFile)) + if err != nil { + return out, fmt.Errorf("read the run: %w", err) + } + var ops simplyblockv1alpha2.OperatorOps + if err := yaml.Unmarshal(opsRaw, &ops); err != nil { + return out, fmt.Errorf("parse the run: %w", err) + } + out.Ops = &ops + + nodes, err := readDocuments[corev1.Node](filepath.Join(dir, nodesFile)) + if err != nil { + return out, fmt.Errorf("read the nodes: %w", err) + } + for i := range nodes { + out.Objects = append(out.Objects, &nodes[i]) + } + + entries, err := os.ReadDir(filepath.Join(dir, reportsDir)) + if err != nil && !os.IsNotExist(err) { + return out, fmt.Errorf("list the reports: %w", err) + } + for _, entry := range entries { + raw, err := os.ReadFile(filepath.Join(dir, reportsDir, entry.Name())) + if err != nil { + return out, err + } + var cm corev1.ConfigMap + if err := yaml.Unmarshal(raw, &cm); err != nil { + return out, fmt.Errorf("parse %s: %w", entry.Name(), err) + } + out.Objects = append(out.Objects, &cm) + } + return out, nil +} + +// readDocuments reads a multi-document file, and nothing for a file that is not +// there: a case with no node objects is a run whose workers Kubernetes says +// nothing about. +func readDocuments[T any](path string) ([]T, error) { + raw, err := os.ReadFile(path) + if err != nil { + if os.IsNotExist(err) { + return nil, nil + } + return nil, err + } + + var out []T + for _, document := range strings.Split(string(raw), "\n---\n") { + if strings.TrimSpace(document) == "" { + continue + } + var object T + if err := yaml.Unmarshal([]byte(document), &object); err != nil { + return nil, err + } + out = append(out, object) + } + return out, nil +} diff --git a/operator/internal/controllers/deployment/discovery_invariants_test.go b/operator/internal/controllers/deployment/discovery_invariants_test.go new file mode 100644 index 000000000..7b6ca38d5 --- /dev/null +++ b/operator/internal/controllers/deployment/discovery_invariants_test.go @@ -0,0 +1,319 @@ +// What must hold of every document the discovery run writes, whatever fleet it +// was written from. +// +// The recorded expectation beside each case says what the run produced. It +// cannot say whether that was right: it was written by running the code that +// produced it, so it asserts what the code already did. This file is the other +// half, and every check in it is made against something written for another +// purpose by somebody else: +// +// - A real apiserver with this repository's own CRDs installed. A draft it +// refuses is a document nobody can apply, however plausible the YAML looks, +// and the schema and its CEL rules are generated from the API types rather +// than restated here. +// - The ClusterDeploymentConfig controller's own draft validation, which is +// what decides whether a reviewer can approve the thing. It was written for +// hand-written documents and knows nothing about discovery. +// - The probe reports themselves. Every worker and every device a draft names +// has to trace back to a report, which is the one property a generator can +// violate silently and catastrophically. +// +// A finding here is a finding about the generator rather than about a fixture, +// which is the whole reason for keeping the two files apart. + +package deployment + +import ( + "context" + "fmt" + "os" + "path/filepath" + "slices" + "sort" + "strings" + "testing" + + corev1 "k8s.io/api/core/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" +) + +// TestEveryDraftedDocumentIsOneTheAPIServerAccepts applies what every case +// produced to a real apiserver. +// +// This is the check the recorded expectations cannot make. A golden file says +// the run wrote this YAML; only the apiserver says whether the YAML is a +// document that can exist, and the schema it decides with is generated from the +// API types rather than written for this test. +func TestEveryDraftedDocumentIsOneTheAPIServerAccepts(t *testing.T) { + api := apiServer(t) + ctx := context.Background() + + namespace := &corev1.Namespace{ObjectMeta: metav1.ObjectMeta{Name: "discovery-cases"}} + if err := api.Create(ctx, namespace); err != nil && !apierrors.IsAlreadyExists(err) { + t.Fatalf("create the namespace to apply into: %v", err) + } + + var refused []string + for _, c := range withoutRecordedGaps(t, loadDiscoveryCases(t)) { + out := runDiscoveryCase(t, c) + if out.Config == nil { + continue + } + + // The name is the case's rather than the run's, so that a hundred and + // fifty documents can be applied into one namespace without colliding, + // and the length a run would have produced is checked separately below. + applied := out.Config.DeepCopy() + applied.Namespace = namespace.Name + applied.Name = applyableName(c.Name) + applied.ResourceVersion = "" + + if err := api.Create(ctx, applied); err != nil { + refused = append(refused, fmt.Sprintf("%s: %s", c.Name, err.Error())) + continue + } + _ = api.Delete(ctx, applied) + } + + sort.Strings(refused) + for _, problem := range refused { + t.Errorf("the run wrote a document the apiserver refuses: %s", problem) + } +} + +// TestEveryDraftedClusterNameIsOneAClusterCanHave checks the half of the name +// the apiserver cannot: the cluster the document proposes does not exist yet, so +// nothing refuses a name too long for it until the run tries to create it. +func TestEveryDraftedClusterNameIsOneAClusterCanHave(t *testing.T) { + const clusterNameLimit = 63 + + var findings []string + for _, c := range withoutRecordedGaps(t, loadDiscoveryCases(t)) { + out := runDiscoveryCase(t, c) + if out.Config == nil || out.Config.Spec.Cluster == nil { + continue + } + if name := out.Config.Spec.Cluster.Name; len(name) > clusterNameLimit { + findings = append(findings, fmt.Sprintf("%s: %q is %d characters", + c.Name, name, len(name))) + } + } + + sort.Strings(findings) + for _, problem := range findings { + t.Errorf("the run proposed a cluster whose name no StorageCluster can carry: %s", problem) + } +} + +// knownUnbuildable are the two findings a generated document is currently +// allowed to carry, each recorded as a gap in discovery-generator-rules.md. +// +// They are listed rather than skipped so that the list is the specification: a +// document failing its own validation for any other reason is a generator that +// has started producing something new, and that is what this test is for. +var knownUnbuildable = map[string]string{ + // A worker presenting no interface that would serve is drafted anyway, with + // the group's management interface empty. The rules document records it as + // the generator writing a group it already knows cannot deploy. + NoManagementInterface: "the ladder found no interface and the group was drafted regardless", + + // The grouper splits workers on their management interface while the + // expansion binds one interface per cluster, so a fleet whose machines name + // their NICs differently produces exactly the document this refuses. + WorkerNotFound: "the groups name different management interfaces", +} + +// TestEveryDraftedDocumentPassesItsOwnValidation runs each document through the +// check a reviewer's document gets before it can be approved. +// +// A draft that fails it is one the reviewer is told is wrong, which for a +// hand-written document is the point and for a generated one is the generator +// handing somebody a document it already knows cannot deploy. +func TestEveryDraftedDocumentPassesItsOwnValidation(t *testing.T) { + var findings []string + for _, c := range loadDiscoveryCases(t) { + out := runDiscoveryCase(t, c) + if out.Config == nil { + continue + } + for _, found := range validateDraft(t, c, out.Config) { + if _, known := knownUnbuildable[found.reason]; known { + continue + } + findings = append(findings, fmt.Sprintf("%s: %s: %s", c.Name, found.reason, found.message)) + } + } + + sort.Strings(findings) + for _, problem := range findings { + t.Errorf("the run wrote a document its own validation rejects: %s", problem) + } +} + +// TestEveryDraftedDocumentNamesOnlyWhatWasReported checks the draft against the +// reports it was built from. +// +// It is the property that cannot be read off any single document: a device +// named for a worker that never reported it is a cluster told to take a disk +// that may belong to something else, and nothing downstream would catch it, +// because the address is well formed and the worker exists. +func TestEveryDraftedDocumentNamesOnlyWhatWasReported(t *testing.T) { + var findings []string + for _, c := range loadDiscoveryCases(t) { + out := runDiscoveryCase(t, c) + if out.Config == nil { + continue + } + for _, problem := range unreportedContent(c, out.Config) { + findings = append(findings, fmt.Sprintf("%s: %s", c.Name, problem)) + } + } + + sort.Strings(findings) + for _, problem := range findings { + t.Errorf("the run named something no probe reported: %s", problem) + } +} + +// TestNoDraftedDocumentPlacesAWorkerTwice is the invariant the expansion +// depends on. +// +// A StorageNode is identified by its cluster, its worker, and its slot, so a +// worker in two groups expands into one node carrying whichever group the +// expansion reached first. The document's own validation reports it; a +// generator should be incapable of producing it. +func TestNoDraftedDocumentPlacesAWorkerTwice(t *testing.T) { + var findings []string + for _, c := range loadDiscoveryCases(t) { + out := runDiscoveryCase(t, c) + if out.Config == nil { + continue + } + if repeated := duplicateWorkers(out.Config); len(repeated) > 0 { + findings = append(findings, fmt.Sprintf("%s: %s", c.Name, strings.Join(repeated, ", "))) + } + } + + sort.Strings(findings) + for _, problem := range findings { + t.Errorf("the run placed a worker in more than one group: %s", problem) + } +} + +// withoutRecordedGaps drops the cases that exist to hold today's output against +// a finding. +// +// A case carrying a gap number is one the document already says is wrong, and +// the expectation beside it is the record that it has not silently changed. +// Holding those to a property they are known to break would turn every finding +// into a failing test and leave no signal for the cases that should hold. +func withoutRecordedGaps(t *testing.T, cases []discoveryCase) []discoveryCase { + t.Helper() + + var out []discoveryCase + for _, c := range cases { + body, err := os.ReadFile(filepath.Join(c.Dir, caseFile)) + if err != nil { + t.Fatalf("read %s: %v", c.Dir, err) + } + if strings.Contains(string(body), "**Gap.**") { + continue + } + out = append(out, c) + } + return out +} + +// validateDraft runs the document through the ClusterDeploymentConfig +// controller's own draft validation, against the same nodes the case carries. +func validateDraft( + t *testing.T, c discoveryCase, config *simplyblockv1alpha2.ClusterDeploymentConfig, +) []finding { + t.Helper() + + scheme := opsScheme(t) + objects := append([]client.Object{config.DeepCopy()}, c.Objects...) + fakeClient := fake.NewClientBuilder().WithScheme(scheme).WithObjects(objects...).Build() + + reconciler := &ClusterDeploymentConfigReconciler{Client: fakeClient, Scheme: scheme} + findings, err := reconciler.validate(context.Background(), config) + if err != nil { + t.Fatalf("%s: validate the document: %v", c.Name, err) + } + return findings +} + +// unreportedContent is every worker and every device address the document names +// that no probe report accounts for. +func unreportedContent( + c discoveryCase, config *simplyblockv1alpha2.ClusterDeploymentConfig, +) []string { + reported := map[string]map[string]struct{}{} + for _, object := range c.Objects { + cm, isConfigMap := object.(*corev1.ConfigMap) + if !isConfigMap { + continue + } + report, err := nodeprobe.ReportFromConfigMap(cm) + if err != nil { + continue + } + + names := map[string]struct{}{} + for _, device := range report.Devices { + if device.PCIAddress != "" { + names[device.PCIAddress] = struct{}{} + } + if device.Path != "" { + names[device.Path] = struct{}{} + } + } + for _, controller := range report.NVMeControllers { + names[controller.Address] = struct{}{} + } + reported[report.Node] = names + } + + var problems []string + for _, set := range config.Spec.NodeSets { + for _, group := range set.Groups { + var addresses []string + if group.Devices != nil { + addresses = append(addresses, group.Devices.NVMe...) + addresses = append(addresses, group.Devices.Block...) + } + for _, worker := range group.Workers { + names, known := reported[worker] + if !known { + problems = append(problems, "worker "+worker+" has no report") + continue + } + for _, address := range addresses { + if _, found := names[address]; !found { + problems = append(problems, fmt.Sprintf( + "%s is named for worker %s, which reported no such device", + address, worker)) + } + } + } + } + } + slices.Sort(problems) + return slices.Compact(problems) +} + +// applyableName is a name for the copy of a document applied to the apiserver, +// derived from the case's path so that a refusal names the case. +func applyableName(caseName string) string { + name := strings.ReplaceAll(caseName, "/", "-") + if len(name) > 200 { + name = name[:200] + } + return strings.Trim(name, "-") +} diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/case.md b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/case.md new file mode 100644 index 000000000..edf89d15f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/case.md @@ -0,0 +1,7 @@ +# CM-01 + +**Mutation.** `version: 4` in one report of three, the schema before the subsystem NQN + +**Expected.** `ReportUnreadable` naming both versions. The other two are drafted + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/expected-notes.txt new file mode 100644 index 000000000..a749d795e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-02, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-02, which is the smallest of the fleet +wrote discovered-discover-cm-01 in Draft: 2 workers with 8 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/expected.yaml new file mode 100644 index 000000000..902a7e015 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/expected.yaml @@ -0,0 +1,33 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-cm-01 + name: discovered-discover-cm-01 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-cm-01-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-02 + - worker-03 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/nodes.yaml new file mode 100644 index 000000000..98de08ee9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/nodes.yaml @@ -0,0 +1,89 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/ops.yaml new file mode 100644 index 000000000..1e4d42caf --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/ops.yaml @@ -0,0 +1,14 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-cm-01 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/reports/worker-01.yaml new file mode 100644 index 000000000..e165f80de --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "cpu": { + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ], + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2 + }, + "devices": [ + { + "available": true, + "content": "Blank", + "kind": "Disk", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "name": "nvme0n1", + "numaNode": 0, + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "transport": "NVMe" + }, + { + "available": true, + "content": "Blank", + "kind": "Disk", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "name": "nvme1n1", + "numaNode": 0, + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "transport": "NVMe" + }, + { + "available": true, + "content": "Blank", + "kind": "Disk", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "name": "nvme2n1", + "numaNode": 0, + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "transport": "NVMe" + }, + { + "available": true, + "content": "Blank", + "kind": "Disk", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "name": "nvme3n1", + "numaNode": 0, + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "transport": "NVMe" + } + ], + "hugePages": [ + { + "free": 32, + "numaNodes": [ + { + "free": 16, + "node": 0, + "total": 16 + }, + { + "free": 16, + "node": 1, + "total": 16 + } + ], + "sizeBytes": 1073741824, + "total": 32 + } + ], + "interfaces": [ + { + "addresses": [ + "10.10.10.1" + ], + "driver": "mlx5_core", + "kind": "physical", + "mtu": 1500, + "name": "eth0", + "numaNode": 0, + "speedMbps": 25000, + "state": "up" + }, + { + "addresses": [ + "10.10.11.1" + ], + "driver": "mlx5_core", + "kind": "physical", + "mtu": 1500, + "name": "eth1", + "numaNode": 1, + "speedMbps": 25000, + "state": "up" + } + ], + "memory": { + "availableBytes": 257698037760, + "freeBytes": 64424509440, + "numaNodes": [ + { + "freeBytes": 68719476736, + "node": 0, + "totalBytes": 137438953472 + }, + { + "freeBytes": 68719476736, + "node": 1, + "totalBytes": 137438953472 + } + ], + "totalBytes": 274877906944 + }, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "version": 5 + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-cm-01 + name: sb-nodeprobe-discover-cm-01-worker-01-4c81e50d + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/reports/worker-02.yaml new file mode 100644 index 000000000..26e972675 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-cm-01 + name: sb-nodeprobe-discover-cm-01-worker-02-6b221646 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/reports/worker-03.yaml new file mode 100644 index 000000000..1329e4189 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-cm-01 + name: sb-nodeprobe-discover-cm-01-worker-03-52a3d756 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/case.md b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/case.md new file mode 100644 index 000000000..a70435aad --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/case.md @@ -0,0 +1,7 @@ +# CM-02 + +**Mutation.** A ConfigMap with no `report.json` key + +**Expected.** `ReportUnreadable` saying the probe did not finish writing it + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/expected-notes.txt new file mode 100644 index 000000000..3d2c5161f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-02, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-02, which is the smallest of the fleet +wrote discovered-discover-cm-02 in Draft: 2 workers with 8 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/expected.yaml new file mode 100644 index 000000000..1cd89073b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/expected.yaml @@ -0,0 +1,33 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-cm-02 + name: discovered-discover-cm-02 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-cm-02-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-02 + - worker-03 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/nodes.yaml new file mode 100644 index 000000000..98de08ee9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/nodes.yaml @@ -0,0 +1,89 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/ops.yaml new file mode 100644 index 000000000..1c30c8a7a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/ops.yaml @@ -0,0 +1,14 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-cm-02 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/reports/worker-01.yaml new file mode 100644 index 000000000..8cc6b6680 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/reports/worker-01.yaml @@ -0,0 +1,9 @@ +apiVersion: v1 +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-cm-02 + name: sb-nodeprobe-discover-cm-02-worker-01-54a2b7e2 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/reports/worker-02.yaml new file mode 100644 index 000000000..887ce54a8 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-cm-02 + name: sb-nodeprobe-discover-cm-02-worker-02-b52001ce + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/reports/worker-03.yaml new file mode 100644 index 000000000..7889a3c38 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-cm-02 + name: sb-nodeprobe-discover-cm-02-worker-03-b1c85bfb + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/case.md b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/case.md new file mode 100644 index 000000000..ae2076f54 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/case.md @@ -0,0 +1,7 @@ +# CM-03 + +**Mutation.** `report.json` holding malformed JSON + +**Expected.** `ReportUnreadable` quoting the parse failure + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/expected-notes.txt new file mode 100644 index 000000000..b67a69df2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-02, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-02, which is the smallest of the fleet +wrote discovered-discover-cm-03 in Draft: 2 workers with 8 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/expected.yaml new file mode 100644 index 000000000..c48adba96 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/expected.yaml @@ -0,0 +1,33 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-cm-03 + name: discovered-discover-cm-03 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-cm-03-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-02 + - worker-03 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/nodes.yaml new file mode 100644 index 000000000..98de08ee9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/nodes.yaml @@ -0,0 +1,89 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/ops.yaml new file mode 100644 index 000000000..bba54418e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/ops.yaml @@ -0,0 +1,14 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-cm-03 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/reports/worker-01.yaml new file mode 100644 index 000000000..d064b6036 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/reports/worker-01.yaml @@ -0,0 +1,11 @@ +apiVersion: v1 +data: + report.json: '{"version": 4, "node": "worker-01",' +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-cm-03 + name: sb-nodeprobe-discover-cm-03-worker-01-120c4748 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/reports/worker-02.yaml new file mode 100644 index 000000000..3c3df6f6e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-cm-03 + name: sb-nodeprobe-discover-cm-03-worker-02-aca8e9d7 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/reports/worker-03.yaml new file mode 100644 index 000000000..5d8bf107d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-cm-03 + name: sb-nodeprobe-discover-cm-03-worker-03-308a2575 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/case.md b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/case.md new file mode 100644 index 000000000..119073026 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/case.md @@ -0,0 +1,7 @@ +# CM-04 + +**Mutation.** A report with an empty `node` + +**Expected.** Refused: nothing can be attributed to it + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/expected-notes.txt new file mode 100644 index 000000000..f9a1e897d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-02, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-02, which is the smallest of the fleet +wrote discovered-discover-cm-04 in Draft: 2 workers with 8 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/expected.yaml new file mode 100644 index 000000000..ea24c4e4a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/expected.yaml @@ -0,0 +1,33 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-cm-04 + name: discovered-discover-cm-04 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-cm-04-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-02 + - worker-03 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/nodes.yaml new file mode 100644 index 000000000..98de08ee9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/nodes.yaml @@ -0,0 +1,89 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/ops.yaml new file mode 100644 index 000000000..2abdee482 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/ops.yaml @@ -0,0 +1,14 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-cm-04 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/reports/worker-01.yaml new file mode 100644 index 000000000..a086c3e1f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-cm-04 + name: sb-nodeprobe-discover-cm-04-worker-01-4a1c9b79 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/reports/worker-02.yaml new file mode 100644 index 000000000..5a4a98584 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-cm-04 + name: sb-nodeprobe-discover-cm-04-worker-02-a95bda55 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/reports/worker-03.yaml new file mode 100644 index 000000000..32f637a3b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-cm-04 + name: sb-nodeprobe-discover-cm-04-worker-03-268308f1 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/case.md b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/case.md new file mode 100644 index 000000000..5d819c315 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/case.md @@ -0,0 +1,7 @@ +# CM-05 + +**Mutation.** A ConfigMap without the run label + +**Expected.** Not listed, so not read + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/expected-notes.txt new file mode 100644 index 000000000..a5301ac33 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-02, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-02, which is the smallest of the fleet +wrote discovered-discover-cm-05 in Draft: 2 workers with 8 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/expected.yaml new file mode 100644 index 000000000..ff1e69f27 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/expected.yaml @@ -0,0 +1,33 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-cm-05 + name: discovered-discover-cm-05 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-cm-05-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-02 + - worker-03 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/nodes.yaml new file mode 100644 index 000000000..98de08ee9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/nodes.yaml @@ -0,0 +1,89 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/ops.yaml new file mode 100644 index 000000000..0f1a5cfef --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/ops.yaml @@ -0,0 +1,14 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-cm-05 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/reports/worker-01.yaml new file mode 100644 index 000000000..6f8773cc7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/reports/worker-01.yaml @@ -0,0 +1,174 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + name: sb-nodeprobe-discover-cm-05-worker-01-65978d83 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/reports/worker-02.yaml new file mode 100644 index 000000000..37c34b26b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-cm-05 + name: sb-nodeprobe-discover-cm-05-worker-02-37832d23 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/reports/worker-03.yaml new file mode 100644 index 000000000..34b8a3cb7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-cm-05 + name: sb-nodeprobe-discover-cm-05-worker-03-ba214f0c + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/case.md b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/case.md new file mode 100644 index 000000000..5e07255db --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/case.md @@ -0,0 +1,7 @@ +# CM-06 + +**Mutation.** A report from a node whose name exceeds 63 characters + +**Expected.** Drafted under its full name: the label is truncated and the report is not + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/expected-notes.txt new file mode 100644 index 000000000..27357dbde --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on ip-10-0-1-23.eu-central-1.compute.internal.a-very-long-suffix-nobody-shortened, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on ip-10-0-1-23.eu-central-1.compute.internal.a-very-long-suffix-nobody-shortened, which is the smallest of the fleet +wrote discovered-discover-cm-06 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/expected.yaml new file mode 100644 index 000000000..a3f21f11e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-cm-06 + name: discovered-discover-cm-06 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-cm-06-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - ip-10-0-1-23.eu-central-1.compute.internal.a-very-long-suffix-nobody-shortened + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/nodes.yaml new file mode 100644 index 000000000..cbb71d196 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: ip-10-0-1-23.eu-central-1.compute.internal.a-very-long-suffix-nobody-shortened +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/ops.yaml new file mode 100644 index 000000000..250db31ea --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-cm-06 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - ip-10-0-1-23.eu-central-1.compute.internal.a-very-long-suffix-nobody-shortened diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/reports/ip-10-0-1-23.eu-central-1.compute.internal.a-very-long-suffix-nobody-shortened.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/reports/ip-10-0-1-23.eu-central-1.compute.internal.a-very-long-suffix-nobody-shortened.yaml new file mode 100644 index 000000000..49e632226 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/reports/ip-10-0-1-23.eu-central-1.compute.internal.a-very-long-suffix-nobody-shortened.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "ip-10-0-1-23.eu-central-1.compute.internal.a-very-long-suffix-nobody-shortened", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: ip-10-0-1-23.eu-central-1.compute.internal.a-very-long-suffix-n + storage.simplyblock.io/nodeprobe-run: discover-cm-06 + name: sb-nodeprobe-discover-cm-06-ip-10-0-1-23.eu-central-1-de65fe1c + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/case.md b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/case.md new file mode 100644 index 000000000..01e3d3191 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/case.md @@ -0,0 +1,7 @@ +# CM-07 + +**Mutation.** A report of a 128-device worker near the ConfigMap ceiling + +**Expected.** Decoded and drafted + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/expected-notes.txt new file mode 100644 index 000000000..7f274d953 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-cm-07 in Draft: 1 workers with 128 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/expected.yaml new file mode 100644 index 000000000..6d097b9dc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/expected.yaml @@ -0,0 +1,156 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-cm-07 + name: discovered-discover-cm-07 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-cm-07-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:60:00.0 + - 0000:61:00.0 + - 0000:62:00.0 + - 0000:63:00.0 + - 0000:64:00.0 + - 0000:65:00.0 + - 0000:66:00.0 + - 0000:67:00.0 + - 0000:68:00.0 + - 0000:69:00.0 + - 0000:6a:00.0 + - 0000:6b:00.0 + - 0000:6c:00.0 + - 0000:6d:00.0 + - 0000:6e:00.0 + - 0000:6f:00.0 + - 0000:70:00.0 + - 0000:71:00.0 + - 0000:72:00.0 + - 0000:73:00.0 + - 0000:74:00.0 + - 0000:75:00.0 + - 0000:76:00.0 + - 0000:77:00.0 + - 0000:78:00.0 + - 0000:79:00.0 + - 0000:7a:00.0 + - 0000:7b:00.0 + - 0000:7c:00.0 + - 0000:7d:00.0 + - 0000:7e:00.0 + - 0000:7f:00.0 + - 0000:80:00.0 + - 0000:81:00.0 + - 0000:82:00.0 + - 0000:83:00.0 + - 0000:84:00.0 + - 0000:85:00.0 + - 0000:86:00.0 + - 0000:87:00.0 + - 0000:88:00.0 + - 0000:89:00.0 + - 0000:8a:00.0 + - 0000:8b:00.0 + - 0000:8c:00.0 + - 0000:8d:00.0 + - 0000:8e:00.0 + - 0000:8f:00.0 + - 0000:90:00.0 + - 0000:91:00.0 + - 0000:92:00.0 + - 0000:93:00.0 + - 0000:94:00.0 + - 0000:95:00.0 + - 0000:96:00.0 + - 0000:97:00.0 + - 0000:98:00.0 + - 0000:99:00.0 + - 0000:9a:00.0 + - 0000:9b:00.0 + - 0000:9c:00.0 + - 0000:9d:00.0 + - 0000:9e:00.0 + - 0000:9f:00.0 + - 0000:a0:00.0 + - 0000:a1:00.0 + - 0000:a2:00.0 + - 0000:a3:00.0 + - 0000:a4:00.0 + - 0000:a5:00.0 + - 0000:a6:00.0 + - 0000:a7:00.0 + - 0000:a8:00.0 + - 0000:a9:00.0 + - 0000:aa:00.0 + - 0000:ab:00.0 + - 0000:ac:00.0 + - 0000:ad:00.0 + - 0000:ae:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + - 0000:b1:00.0 + - 0000:b2:00.0 + - 0000:b3:00.0 + - 0000:b4:00.0 + - 0000:b5:00.0 + - 0000:b6:00.0 + - 0000:b7:00.0 + - 0000:b8:00.0 + - 0000:b9:00.0 + - 0000:ba:00.0 + - 0000:bb:00.0 + - 0000:bc:00.0 + - 0000:bd:00.0 + - 0000:be:00.0 + - 0000:bf:00.0 + - 0000:c0:00.0 + - 0000:c1:00.0 + - 0000:c2:00.0 + - 0000:c3:00.0 + - 0000:c4:00.0 + - 0000:c5:00.0 + - 0000:c6:00.0 + - 0000:c7:00.0 + - 0000:c8:00.0 + - 0000:c9:00.0 + - 0000:ca:00.0 + - 0000:cb:00.0 + - 0000:cc:00.0 + - 0000:cd:00.0 + - 0000:ce:00.0 + - 0000:cf:00.0 + - 0000:d0:00.0 + - 0000:d1:00.0 + - 0000:d2:00.0 + - 0000:d3:00.0 + - 0000:d4:00.0 + - 0000:d5:00.0 + - 0000:d6:00.0 + - 0000:d7:00.0 + - 0000:d8:00.0 + - 0000:d9:00.0 + - 0000:da:00.0 + - 0000:db:00.0 + - 0000:dc:00.0 + - 0000:dd:00.0 + mgmtInterface: eth0 + name: group-1-nvme-128x512G + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/ops.yaml new file mode 100644 index 000000000..8db0c9fb1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-cm-07 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/reports/worker-01.yaml new file mode 100644 index 000000000..4cba535d4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/reports/worker-01.yaml @@ -0,0 +1,1663 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:60:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:61:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme4n1", + "path": "/dev/nvme4n1", + "pciAddress": "0000:62:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme5n1", + "path": "/dev/nvme5n1", + "pciAddress": "0000:63:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme6n1", + "path": "/dev/nvme6n1", + "pciAddress": "0000:64:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme7n1", + "path": "/dev/nvme7n1", + "pciAddress": "0000:65:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme8n1", + "path": "/dev/nvme8n1", + "pciAddress": "0000:66:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme9n1", + "path": "/dev/nvme9n1", + "pciAddress": "0000:67:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme10n1", + "path": "/dev/nvme10n1", + "pciAddress": "0000:68:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme11n1", + "path": "/dev/nvme11n1", + "pciAddress": "0000:69:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme12n1", + "path": "/dev/nvme12n1", + "pciAddress": "0000:6a:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme13n1", + "path": "/dev/nvme13n1", + "pciAddress": "0000:6b:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme14n1", + "path": "/dev/nvme14n1", + "pciAddress": "0000:6c:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme15n1", + "path": "/dev/nvme15n1", + "pciAddress": "0000:6d:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme16n1", + "path": "/dev/nvme16n1", + "pciAddress": "0000:6e:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme17n1", + "path": "/dev/nvme17n1", + "pciAddress": "0000:6f:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme18n1", + "path": "/dev/nvme18n1", + "pciAddress": "0000:70:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme19n1", + "path": "/dev/nvme19n1", + "pciAddress": "0000:71:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme20n1", + "path": "/dev/nvme20n1", + "pciAddress": "0000:72:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme21n1", + "path": "/dev/nvme21n1", + "pciAddress": "0000:73:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme22n1", + "path": "/dev/nvme22n1", + "pciAddress": "0000:74:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme23n1", + "path": "/dev/nvme23n1", + "pciAddress": "0000:75:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme24n1", + "path": "/dev/nvme24n1", + "pciAddress": "0000:76:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme25n1", + "path": "/dev/nvme25n1", + "pciAddress": "0000:77:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme26n1", + "path": "/dev/nvme26n1", + "pciAddress": "0000:78:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme27n1", + "path": "/dev/nvme27n1", + "pciAddress": "0000:79:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme28n1", + "path": "/dev/nvme28n1", + "pciAddress": "0000:7a:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme29n1", + "path": "/dev/nvme29n1", + "pciAddress": "0000:7b:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme30n1", + "path": "/dev/nvme30n1", + "pciAddress": "0000:7c:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme31n1", + "path": "/dev/nvme31n1", + "pciAddress": "0000:7d:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme32n1", + "path": "/dev/nvme32n1", + "pciAddress": "0000:7e:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme33n1", + "path": "/dev/nvme33n1", + "pciAddress": "0000:7f:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme34n1", + "path": "/dev/nvme34n1", + "pciAddress": "0000:80:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme35n1", + "path": "/dev/nvme35n1", + "pciAddress": "0000:81:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme36n1", + "path": "/dev/nvme36n1", + "pciAddress": "0000:82:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme37n1", + "path": "/dev/nvme37n1", + "pciAddress": "0000:83:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme38n1", + "path": "/dev/nvme38n1", + "pciAddress": "0000:84:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme39n1", + "path": "/dev/nvme39n1", + "pciAddress": "0000:85:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme40n1", + "path": "/dev/nvme40n1", + "pciAddress": "0000:86:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme41n1", + "path": "/dev/nvme41n1", + "pciAddress": "0000:87:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme42n1", + "path": "/dev/nvme42n1", + "pciAddress": "0000:88:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme43n1", + "path": "/dev/nvme43n1", + "pciAddress": "0000:89:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme44n1", + "path": "/dev/nvme44n1", + "pciAddress": "0000:8a:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme45n1", + "path": "/dev/nvme45n1", + "pciAddress": "0000:8b:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme46n1", + "path": "/dev/nvme46n1", + "pciAddress": "0000:8c:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme47n1", + "path": "/dev/nvme47n1", + "pciAddress": "0000:8d:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme48n1", + "path": "/dev/nvme48n1", + "pciAddress": "0000:8e:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme49n1", + "path": "/dev/nvme49n1", + "pciAddress": "0000:8f:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme50n1", + "path": "/dev/nvme50n1", + "pciAddress": "0000:90:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme51n1", + "path": "/dev/nvme51n1", + "pciAddress": "0000:91:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme52n1", + "path": "/dev/nvme52n1", + "pciAddress": "0000:92:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme53n1", + "path": "/dev/nvme53n1", + "pciAddress": "0000:93:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme54n1", + "path": "/dev/nvme54n1", + "pciAddress": "0000:94:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme55n1", + "path": "/dev/nvme55n1", + "pciAddress": "0000:95:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme56n1", + "path": "/dev/nvme56n1", + "pciAddress": "0000:96:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme57n1", + "path": "/dev/nvme57n1", + "pciAddress": "0000:97:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme58n1", + "path": "/dev/nvme58n1", + "pciAddress": "0000:98:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme59n1", + "path": "/dev/nvme59n1", + "pciAddress": "0000:99:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme60n1", + "path": "/dev/nvme60n1", + "pciAddress": "0000:9a:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme61n1", + "path": "/dev/nvme61n1", + "pciAddress": "0000:9b:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme62n1", + "path": "/dev/nvme62n1", + "pciAddress": "0000:9c:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme63n1", + "path": "/dev/nvme63n1", + "pciAddress": "0000:9d:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme64n1", + "path": "/dev/nvme64n1", + "pciAddress": "0000:9e:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme65n1", + "path": "/dev/nvme65n1", + "pciAddress": "0000:9f:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme66n1", + "path": "/dev/nvme66n1", + "pciAddress": "0000:a0:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme67n1", + "path": "/dev/nvme67n1", + "pciAddress": "0000:a1:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme68n1", + "path": "/dev/nvme68n1", + "pciAddress": "0000:a2:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme69n1", + "path": "/dev/nvme69n1", + "pciAddress": "0000:a3:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme70n1", + "path": "/dev/nvme70n1", + "pciAddress": "0000:a4:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme71n1", + "path": "/dev/nvme71n1", + "pciAddress": "0000:a5:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme72n1", + "path": "/dev/nvme72n1", + "pciAddress": "0000:a6:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme73n1", + "path": "/dev/nvme73n1", + "pciAddress": "0000:a7:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme74n1", + "path": "/dev/nvme74n1", + "pciAddress": "0000:a8:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme75n1", + "path": "/dev/nvme75n1", + "pciAddress": "0000:a9:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme76n1", + "path": "/dev/nvme76n1", + "pciAddress": "0000:aa:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme77n1", + "path": "/dev/nvme77n1", + "pciAddress": "0000:ab:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme78n1", + "path": "/dev/nvme78n1", + "pciAddress": "0000:ac:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme79n1", + "path": "/dev/nvme79n1", + "pciAddress": "0000:ad:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme80n1", + "path": "/dev/nvme80n1", + "pciAddress": "0000:ae:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme81n1", + "path": "/dev/nvme81n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme82n1", + "path": "/dev/nvme82n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme83n1", + "path": "/dev/nvme83n1", + "pciAddress": "0000:b1:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme84n1", + "path": "/dev/nvme84n1", + "pciAddress": "0000:b2:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme85n1", + "path": "/dev/nvme85n1", + "pciAddress": "0000:b3:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme86n1", + "path": "/dev/nvme86n1", + "pciAddress": "0000:b4:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme87n1", + "path": "/dev/nvme87n1", + "pciAddress": "0000:b5:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme88n1", + "path": "/dev/nvme88n1", + "pciAddress": "0000:b6:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme89n1", + "path": "/dev/nvme89n1", + "pciAddress": "0000:b7:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme90n1", + "path": "/dev/nvme90n1", + "pciAddress": "0000:b8:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme91n1", + "path": "/dev/nvme91n1", + "pciAddress": "0000:b9:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme92n1", + "path": "/dev/nvme92n1", + "pciAddress": "0000:ba:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme93n1", + "path": "/dev/nvme93n1", + "pciAddress": "0000:bb:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme94n1", + "path": "/dev/nvme94n1", + "pciAddress": "0000:bc:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme95n1", + "path": "/dev/nvme95n1", + "pciAddress": "0000:bd:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme96n1", + "path": "/dev/nvme96n1", + "pciAddress": "0000:be:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme97n1", + "path": "/dev/nvme97n1", + "pciAddress": "0000:bf:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme98n1", + "path": "/dev/nvme98n1", + "pciAddress": "0000:c0:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme99n1", + "path": "/dev/nvme99n1", + "pciAddress": "0000:c1:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme100n1", + "path": "/dev/nvme100n1", + "pciAddress": "0000:c2:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme101n1", + "path": "/dev/nvme101n1", + "pciAddress": "0000:c3:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme102n1", + "path": "/dev/nvme102n1", + "pciAddress": "0000:c4:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme103n1", + "path": "/dev/nvme103n1", + "pciAddress": "0000:c5:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme104n1", + "path": "/dev/nvme104n1", + "pciAddress": "0000:c6:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme105n1", + "path": "/dev/nvme105n1", + "pciAddress": "0000:c7:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme106n1", + "path": "/dev/nvme106n1", + "pciAddress": "0000:c8:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme107n1", + "path": "/dev/nvme107n1", + "pciAddress": "0000:c9:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme108n1", + "path": "/dev/nvme108n1", + "pciAddress": "0000:ca:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme109n1", + "path": "/dev/nvme109n1", + "pciAddress": "0000:cb:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme110n1", + "path": "/dev/nvme110n1", + "pciAddress": "0000:cc:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme111n1", + "path": "/dev/nvme111n1", + "pciAddress": "0000:cd:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme112n1", + "path": "/dev/nvme112n1", + "pciAddress": "0000:ce:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme113n1", + "path": "/dev/nvme113n1", + "pciAddress": "0000:cf:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme114n1", + "path": "/dev/nvme114n1", + "pciAddress": "0000:d0:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme115n1", + "path": "/dev/nvme115n1", + "pciAddress": "0000:d1:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme116n1", + "path": "/dev/nvme116n1", + "pciAddress": "0000:d2:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme117n1", + "path": "/dev/nvme117n1", + "pciAddress": "0000:d3:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme118n1", + "path": "/dev/nvme118n1", + "pciAddress": "0000:d4:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme119n1", + "path": "/dev/nvme119n1", + "pciAddress": "0000:d5:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme120n1", + "path": "/dev/nvme120n1", + "pciAddress": "0000:d6:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme121n1", + "path": "/dev/nvme121n1", + "pciAddress": "0000:d7:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme122n1", + "path": "/dev/nvme122n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme123n1", + "path": "/dev/nvme123n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme124n1", + "path": "/dev/nvme124n1", + "pciAddress": "0000:da:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme125n1", + "path": "/dev/nvme125n1", + "pciAddress": "0000:db:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme126n1", + "path": "/dev/nvme126n1", + "pciAddress": "0000:dc:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme127n1", + "path": "/dev/nvme127n1", + "pciAddress": "0000:dd:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-cm-07 + name: sb-nodeprobe-discover-cm-07-worker-01-ffd0a260 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/case.md b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/case.md new file mode 100644 index 000000000..87da7f9bc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/case.md @@ -0,0 +1,7 @@ +# DEV-01 + +**Mutation.** 4 NVMe disks, one memory node, NVMe run + +**Expected.** 1 group, `devices.nvme` holds the 4 addresses ascending, name `group-1-nvme-4x3T` + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/expected-notes.txt new file mode 100644 index 000000000..19b879293 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-dev-01 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/expected.yaml new file mode 100644 index 000000000..c50e61d3e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-dev-01 + name: discovered-discover-dev-01 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-dev-01-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:60:00.0 + - 0000:61:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/ops.yaml new file mode 100644 index 000000000..125484235 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-dev-01 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/reports/worker-01.yaml new file mode 100644 index 000000000..bfbda426c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:60:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:61:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-dev-01 + name: sb-nodeprobe-discover-dev-01-worker-01-37459fbb + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/case.md b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/case.md new file mode 100644 index 000000000..b8a1f5bc1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/case.md @@ -0,0 +1,7 @@ +# DEV-02 + +**Mutation.** 4 virtio disks, block run + +**Expected.** 1 group, `devices.block` holds the 4 paths, class `block` + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/expected-notes.txt new file mode 100644 index 000000000..3c1d9d9ed --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-dev-02 in Draft: 1 workers with 4 block devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/expected.yaml new file mode 100644 index 000000000..a16ce0258 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-dev-02 + name: discovered-discover-dev-02 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-dev-02-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + block: + - /dev/vdb + - /dev/vdc + - /dev/vdd + - /dev/vde + mgmtInterface: eth0 + name: group-1-block-4x2T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/ops.yaml new file mode 100644 index 000000000..f689592cb --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/ops.yaml @@ -0,0 +1,15 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-dev-02 + namespace: simplyblock +spec: + action: Discover + discover: + deviceFilter: + enableLogicalBlockDevices: true +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/reports/worker-01.yaml new file mode 100644 index 000000000..1aadccbbe --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/reports/worker-01.yaml @@ -0,0 +1,167 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "vdb", + "path": "/dev/vdb", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "vdc", + "path": "/dev/vdc", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "vdd", + "path": "/dev/vdd", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "vde", + "path": "/dev/vde", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-dev-02 + name: sb-nodeprobe-discover-dev-02-worker-01-5354bc65 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-03-block-disks-on-an-nvme-run/case.md b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-03-block-disks-on-an-nvme-run/case.md new file mode 100644 index 000000000..8ed38525f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-03-block-disks-on-an-nvme-run/case.md @@ -0,0 +1,7 @@ +# DEV-03 + +**Mutation.** 4 virtio disks, NVMe run + +**Expected.** No node sets. Every device refused by `device class`, the worker by `has devices` + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-03-block-disks-on-an-nvme-run/expected-error.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-03-block-disks-on-an-nvme-run/expected-error.txt new file mode 100644 index 000000000..27b1e7e6e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-03-block-disks-on-an-nvme-run/expected-error.txt @@ -0,0 +1 @@ +no worker has a device this run would use: worker-01: no device of it survived the device rules diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-03-block-disks-on-an-nvme-run/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-03-block-disks-on-an-nvme-run/expected-refusals.txt new file mode 100644 index 000000000..8187480d2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-03-block-disks-on-an-nvme-run/expected-refusals.txt @@ -0,0 +1,5 @@ +worker-01/vdb: declined by device class because this run scans NVMe devices and the device is on Virtio +worker-01/vdc: declined by device class because this run scans NVMe devices and the device is on Virtio +worker-01/vdd: declined by device class because this run scans NVMe devices and the device is on Virtio +worker-01/vde: declined by device class because this run scans NVMe devices and the device is on Virtio +worker-01: declined by has devices because no device of it survived the device rules diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-03-block-disks-on-an-nvme-run/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-03-block-disks-on-an-nvme-run/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-03-block-disks-on-an-nvme-run/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-03-block-disks-on-an-nvme-run/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-03-block-disks-on-an-nvme-run/ops.yaml new file mode 100644 index 000000000..46acbd32e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-03-block-disks-on-an-nvme-run/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-dev-03 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-03-block-disks-on-an-nvme-run/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-03-block-disks-on-an-nvme-run/reports/worker-01.yaml new file mode 100644 index 000000000..2c8887d95 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-03-block-disks-on-an-nvme-run/reports/worker-01.yaml @@ -0,0 +1,167 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "vdb", + "path": "/dev/vdb", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "vdc", + "path": "/dev/vdc", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "vdd", + "path": "/dev/vdd", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "vde", + "path": "/dev/vde", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-dev-03 + name: sb-nodeprobe-discover-dev-03-worker-01-706c5260 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/case.md b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/case.md new file mode 100644 index 000000000..4deb5f71c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/case.md @@ -0,0 +1,7 @@ +# DEV-04 + +**Mutation.** 2 NVMe and 2 virtio disks on one worker, NVMe run + +**Expected.** Only the 2 PCI addresses reach the group + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/expected-notes.txt new file mode 100644 index 000000000..3fafbe5b6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-dev-04 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 2 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/expected-refusals.txt new file mode 100644 index 000000000..7c9654fec --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/expected-refusals.txt @@ -0,0 +1,2 @@ +worker-01/vdb: declined by device class because this run scans NVMe devices and the device is on Virtio +worker-01/vdc: declined by device class because this run scans NVMe devices and the device is on Virtio diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/expected.yaml new file mode 100644 index 000000000..97fcbaa2e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/expected.yaml @@ -0,0 +1,30 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-dev-04 + name: discovered-discover-dev-04 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-dev-04-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + mgmtInterface: eth0 + name: group-1-nvme-2x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/ops.yaml new file mode 100644 index 000000000..6b359f444 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-dev-04 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/reports/worker-01.yaml new file mode 100644 index 000000000..807ae0df9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/reports/worker-01.yaml @@ -0,0 +1,171 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "vdb", + "path": "/dev/vdb", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "vdc", + "path": "/dev/vdc", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-dev-04 + name: sb-nodeprobe-discover-dev-04-worker-01-7b2a73f9 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/case.md b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/case.md new file mode 100644 index 000000000..d53262124 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/case.md @@ -0,0 +1,9 @@ +# DEV-05 + +**Mutation.** The same worker, block run + +**Expected.** Only the two virtio disks. The NVMe pair is refused for being the other class + +**Harness.** `CM` + +**Gap.** G-1. This case records what the generator does today, so that the day it changes the diff is the finding. diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/expected-notes.txt new file mode 100644 index 000000000..f8cad067a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-dev-05 in Draft: 1 workers with 2 block devices, in 1 group(s) across 1 node set(s); 2 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/expected-refusals.txt new file mode 100644 index 000000000..2914dfdba --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/expected-refusals.txt @@ -0,0 +1,2 @@ +worker-01/nvme0n1: declined by device class because this run scans logical block devices and the device is on NVMe +worker-01/nvme1n1: declined by device class because this run scans logical block devices and the device is on NVMe diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/expected.yaml new file mode 100644 index 000000000..f1ef93ef5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/expected.yaml @@ -0,0 +1,30 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-dev-05 + name: discovered-discover-dev-05 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-dev-05-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + block: + - /dev/vdb + - /dev/vdc + mgmtInterface: eth0 + name: group-1-block-2x2T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/ops.yaml new file mode 100644 index 000000000..e8e53a94e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/ops.yaml @@ -0,0 +1,15 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-dev-05 + namespace: simplyblock +spec: + action: Discover + discover: + deviceFilter: + enableLogicalBlockDevices: true +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/reports/worker-01.yaml new file mode 100644 index 000000000..f05a6ca42 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/reports/worker-01.yaml @@ -0,0 +1,171 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "vdb", + "path": "/dev/vdb", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "vdc", + "path": "/dev/vdc", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-dev-05 + name: sb-nodeprobe-discover-dev-05-worker-01-21ce09de + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/case.md b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/case.md new file mode 100644 index 000000000..433e89a8f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/case.md @@ -0,0 +1,7 @@ +# DEV-06 + +**Mutation.** One controller presenting two namespaces at one PCI address + +**Expected.** The address is named once. `DeviceCount` is 1 for that controller + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/expected-notes.txt new file mode 100644 index 000000000..c16005f59 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-dev-06 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/expected.yaml new file mode 100644 index 000000000..3798c6fbb --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/expected.yaml @@ -0,0 +1,30 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-dev-06 + name: discovered-discover-dev-06 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-dev-06-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + mgmtInterface: eth0 + name: group-1-nvme-2x1.5T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/ops.yaml new file mode 100644 index 000000000..5f345f2f9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-dev-06 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/reports/worker-01.yaml new file mode 100644 index 000000000..64a2231a5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/reports/worker-01.yaml @@ -0,0 +1,163 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 1099511627776, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme0n2", + "path": "/dev/nvme0n2", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 1099511627776, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 1099511627776, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-dev-06 + name: sb-nodeprobe-discover-dev-06-worker-01-b7fb1bb9 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/case.md b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/case.md new file mode 100644 index 000000000..0626d3699 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/case.md @@ -0,0 +1,7 @@ +# DEV-07 + +**Mutation.** A SATA disk beside an NVMe one, block run + +**Expected.** Only the SATA disk. The NVMe run takes only the NVMe disk + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/expected-notes.txt new file mode 100644 index 000000000..542045b0e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-dev-07 in Draft: 1 workers with 1 block devices, in 1 group(s) across 1 node set(s); 1 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/expected-refusals.txt new file mode 100644 index 000000000..4617d6253 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/expected-refusals.txt @@ -0,0 +1 @@ +worker-01/nvme0n1: declined by device class because this run scans logical block devices and the device is on NVMe diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/expected.yaml new file mode 100644 index 000000000..1622e6e6d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/expected.yaml @@ -0,0 +1,29 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-dev-07 + name: discovered-discover-dev-07 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-dev-07-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + block: + - /dev/sda + mgmtInterface: eth0 + name: group-1-block-1x4T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/ops.yaml new file mode 100644 index 000000000..8f35912b8 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/ops.yaml @@ -0,0 +1,15 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-dev-07 + namespace: simplyblock +spec: + action: Discover + discover: + deviceFilter: + enableLogicalBlockDevices: true +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/reports/worker-01.yaml new file mode 100644 index 000000000..36bff9a69 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/reports/worker-01.yaml @@ -0,0 +1,149 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "sda", + "path": "/dev/sda", + "sizeBytes": 4398046511104, + "kind": "Disk", + "transport": "SATA", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-dev-07 + name: sb-nodeprobe-discover-dev-07-worker-01-f250cba7 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/case.md b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/case.md new file mode 100644 index 000000000..98290e2dc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/case.md @@ -0,0 +1,9 @@ +# DEV-08 + +**Mutation.** A rotational HDD beside an SSD, both free + +**Expected.** Both admitted: no rule reads `Rotational`. See §14, gap G-2 + +**Harness.** `CM` + +**Gap.** G-2. This case records what the generator does today, so that the day it changes the diff is the finding. diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/expected-notes.txt new file mode 100644 index 000000000..3f32dc8cc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-dev-08 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/expected.yaml new file mode 100644 index 000000000..90f1b8733 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/expected.yaml @@ -0,0 +1,30 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-dev-08 + name: discovered-discover-dev-08 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-dev-08-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + mgmtInterface: eth0 + name: group-1-nvme-2x9.5T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/ops.yaml new file mode 100644 index 000000000..97b2d07ba --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-dev-08 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/reports/worker-01.yaml new file mode 100644 index 000000000..6084c6787 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/reports/worker-01.yaml @@ -0,0 +1,152 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 17592186044416, + "kind": "Disk", + "transport": "NVMe", + "model": "SEAGATE ST16000NM", + "rotational": true, + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-dev-08 + name: sb-nodeprobe-discover-dev-08-worker-01-a42fd76b + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/case.md b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/case.md new file mode 100644 index 000000000..24d17786c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/case.md @@ -0,0 +1,7 @@ +# DEV-09 + +**Mutation.** 10 NVMe disks, one memory node + +**Expected.** All 10 in one group, name `group-1-nvme-10x3T` + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/expected-notes.txt new file mode 100644 index 000000000..d90b28001 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-dev-09 in Draft: 1 workers with 10 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/expected.yaml new file mode 100644 index 000000000..05b1993e2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/expected.yaml @@ -0,0 +1,38 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-dev-09 + name: discovered-discover-dev-09 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-dev-09-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:60:00.0 + - 0000:61:00.0 + - 0000:62:00.0 + - 0000:63:00.0 + - 0000:64:00.0 + - 0000:65:00.0 + - 0000:66:00.0 + - 0000:67:00.0 + mgmtInterface: eth0 + name: group-1-nvme-10x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/ops.yaml new file mode 100644 index 000000000..bf1fea7fd --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-dev-09 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/reports/worker-01.yaml new file mode 100644 index 000000000..0bb2431f7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/reports/worker-01.yaml @@ -0,0 +1,247 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:60:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:61:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme4n1", + "path": "/dev/nvme4n1", + "pciAddress": "0000:62:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme5n1", + "path": "/dev/nvme5n1", + "pciAddress": "0000:63:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme6n1", + "path": "/dev/nvme6n1", + "pciAddress": "0000:64:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme7n1", + "path": "/dev/nvme7n1", + "pciAddress": "0000:65:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme8n1", + "path": "/dev/nvme8n1", + "pciAddress": "0000:66:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme9n1", + "path": "/dev/nvme9n1", + "pciAddress": "0000:67:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-dev-09 + name: sb-nodeprobe-discover-dev-09-worker-01-fa115c9e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/case.md b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/case.md new file mode 100644 index 000000000..70fbc5528 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/case.md @@ -0,0 +1,7 @@ +# DEV-10 + +**Mutation.** A partition and a loopback device beside 4 disks + +**Expected.** Both pre-filtered: absent from `Plan.Explain`, present in `RefusalLines` + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/expected-notes.txt new file mode 100644 index 000000000..2c3cec7ca --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-dev-10 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 3 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/expected-refusals.txt new file mode 100644 index 000000000..5eb9b8c3a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/expected-refusals.txt @@ -0,0 +1,3 @@ +worker-01/nvme0n1p1: declined by whole disk because it is a Partition rather than a whole disk +worker-01/loop0: declined by whole disk because it is a Loop rather than a whole disk +worker-01/loop1: declined by whole disk because it is a Loop rather than a whole disk diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/expected.yaml new file mode 100644 index 000000000..e06b18c79 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-dev-10 + name: discovered-discover-dev-10 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-dev-10-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:60:00.0 + - 0000:61:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/ops.yaml new file mode 100644 index 000000000..7930ac403 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-dev-10 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/reports/worker-01.yaml new file mode 100644 index 000000000..7282d8371 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/reports/worker-01.yaml @@ -0,0 +1,205 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:60:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:61:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme0n1p1", + "path": "/dev/nvme0n1p1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 549755813888, + "kind": "Partition", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "loop0", + "path": "/dev/loop0", + "sizeBytes": 68719476736, + "kind": "Loop", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "loop1", + "path": "/dev/loop1", + "sizeBytes": 68719476736, + "kind": "Loop", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-dev-10 + name: sb-nodeprobe-discover-dev-10-worker-01-c4813e68 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/case.md b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/case.md new file mode 100644 index 000000000..6f5ea53c6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/case.md @@ -0,0 +1,9 @@ +# DEV-11 + +**Mutation.** 128 NVMe disks on one worker + +**Expected.** One group at the selection's `MaxItems`. The document still applies + +**Harness.** `CM` + +**Note.** 128 is the MaxItems of a group's device selection, so this is the largest worker the schema can describe. diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/expected-notes.txt new file mode 100644 index 000000000..d5cf092f5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-dev-11 in Draft: 1 workers with 128 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/expected.yaml new file mode 100644 index 000000000..20232b9e0 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/expected.yaml @@ -0,0 +1,156 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-dev-11 + name: discovered-discover-dev-11 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-dev-11-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:60:00.0 + - 0000:61:00.0 + - 0000:62:00.0 + - 0000:63:00.0 + - 0000:64:00.0 + - 0000:65:00.0 + - 0000:66:00.0 + - 0000:67:00.0 + - 0000:68:00.0 + - 0000:69:00.0 + - 0000:6a:00.0 + - 0000:6b:00.0 + - 0000:6c:00.0 + - 0000:6d:00.0 + - 0000:6e:00.0 + - 0000:6f:00.0 + - 0000:70:00.0 + - 0000:71:00.0 + - 0000:72:00.0 + - 0000:73:00.0 + - 0000:74:00.0 + - 0000:75:00.0 + - 0000:76:00.0 + - 0000:77:00.0 + - 0000:78:00.0 + - 0000:79:00.0 + - 0000:7a:00.0 + - 0000:7b:00.0 + - 0000:7c:00.0 + - 0000:7d:00.0 + - 0000:7e:00.0 + - 0000:7f:00.0 + - 0000:80:00.0 + - 0000:81:00.0 + - 0000:82:00.0 + - 0000:83:00.0 + - 0000:84:00.0 + - 0000:85:00.0 + - 0000:86:00.0 + - 0000:87:00.0 + - 0000:88:00.0 + - 0000:89:00.0 + - 0000:8a:00.0 + - 0000:8b:00.0 + - 0000:8c:00.0 + - 0000:8d:00.0 + - 0000:8e:00.0 + - 0000:8f:00.0 + - 0000:90:00.0 + - 0000:91:00.0 + - 0000:92:00.0 + - 0000:93:00.0 + - 0000:94:00.0 + - 0000:95:00.0 + - 0000:96:00.0 + - 0000:97:00.0 + - 0000:98:00.0 + - 0000:99:00.0 + - 0000:9a:00.0 + - 0000:9b:00.0 + - 0000:9c:00.0 + - 0000:9d:00.0 + - 0000:9e:00.0 + - 0000:9f:00.0 + - 0000:a0:00.0 + - 0000:a1:00.0 + - 0000:a2:00.0 + - 0000:a3:00.0 + - 0000:a4:00.0 + - 0000:a5:00.0 + - 0000:a6:00.0 + - 0000:a7:00.0 + - 0000:a8:00.0 + - 0000:a9:00.0 + - 0000:aa:00.0 + - 0000:ab:00.0 + - 0000:ac:00.0 + - 0000:ad:00.0 + - 0000:ae:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + - 0000:b1:00.0 + - 0000:b2:00.0 + - 0000:b3:00.0 + - 0000:b4:00.0 + - 0000:b5:00.0 + - 0000:b6:00.0 + - 0000:b7:00.0 + - 0000:b8:00.0 + - 0000:b9:00.0 + - 0000:ba:00.0 + - 0000:bb:00.0 + - 0000:bc:00.0 + - 0000:bd:00.0 + - 0000:be:00.0 + - 0000:bf:00.0 + - 0000:c0:00.0 + - 0000:c1:00.0 + - 0000:c2:00.0 + - 0000:c3:00.0 + - 0000:c4:00.0 + - 0000:c5:00.0 + - 0000:c6:00.0 + - 0000:c7:00.0 + - 0000:c8:00.0 + - 0000:c9:00.0 + - 0000:ca:00.0 + - 0000:cb:00.0 + - 0000:cc:00.0 + - 0000:cd:00.0 + - 0000:ce:00.0 + - 0000:cf:00.0 + - 0000:d0:00.0 + - 0000:d1:00.0 + - 0000:d2:00.0 + - 0000:d3:00.0 + - 0000:d4:00.0 + - 0000:d5:00.0 + - 0000:d6:00.0 + - 0000:d7:00.0 + - 0000:d8:00.0 + - 0000:d9:00.0 + - 0000:da:00.0 + - 0000:db:00.0 + - 0000:dc:00.0 + - 0000:dd:00.0 + mgmtInterface: eth0 + name: group-1-nvme-128x512G + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/ops.yaml new file mode 100644 index 000000000..2d6425893 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-dev-11 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/reports/worker-01.yaml new file mode 100644 index 000000000..c076ad6ae --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/reports/worker-01.yaml @@ -0,0 +1,1663 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:60:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:61:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme4n1", + "path": "/dev/nvme4n1", + "pciAddress": "0000:62:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme5n1", + "path": "/dev/nvme5n1", + "pciAddress": "0000:63:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme6n1", + "path": "/dev/nvme6n1", + "pciAddress": "0000:64:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme7n1", + "path": "/dev/nvme7n1", + "pciAddress": "0000:65:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme8n1", + "path": "/dev/nvme8n1", + "pciAddress": "0000:66:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme9n1", + "path": "/dev/nvme9n1", + "pciAddress": "0000:67:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme10n1", + "path": "/dev/nvme10n1", + "pciAddress": "0000:68:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme11n1", + "path": "/dev/nvme11n1", + "pciAddress": "0000:69:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme12n1", + "path": "/dev/nvme12n1", + "pciAddress": "0000:6a:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme13n1", + "path": "/dev/nvme13n1", + "pciAddress": "0000:6b:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme14n1", + "path": "/dev/nvme14n1", + "pciAddress": "0000:6c:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme15n1", + "path": "/dev/nvme15n1", + "pciAddress": "0000:6d:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme16n1", + "path": "/dev/nvme16n1", + "pciAddress": "0000:6e:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme17n1", + "path": "/dev/nvme17n1", + "pciAddress": "0000:6f:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme18n1", + "path": "/dev/nvme18n1", + "pciAddress": "0000:70:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme19n1", + "path": "/dev/nvme19n1", + "pciAddress": "0000:71:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme20n1", + "path": "/dev/nvme20n1", + "pciAddress": "0000:72:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme21n1", + "path": "/dev/nvme21n1", + "pciAddress": "0000:73:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme22n1", + "path": "/dev/nvme22n1", + "pciAddress": "0000:74:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme23n1", + "path": "/dev/nvme23n1", + "pciAddress": "0000:75:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme24n1", + "path": "/dev/nvme24n1", + "pciAddress": "0000:76:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme25n1", + "path": "/dev/nvme25n1", + "pciAddress": "0000:77:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme26n1", + "path": "/dev/nvme26n1", + "pciAddress": "0000:78:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme27n1", + "path": "/dev/nvme27n1", + "pciAddress": "0000:79:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme28n1", + "path": "/dev/nvme28n1", + "pciAddress": "0000:7a:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme29n1", + "path": "/dev/nvme29n1", + "pciAddress": "0000:7b:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme30n1", + "path": "/dev/nvme30n1", + "pciAddress": "0000:7c:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme31n1", + "path": "/dev/nvme31n1", + "pciAddress": "0000:7d:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme32n1", + "path": "/dev/nvme32n1", + "pciAddress": "0000:7e:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme33n1", + "path": "/dev/nvme33n1", + "pciAddress": "0000:7f:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme34n1", + "path": "/dev/nvme34n1", + "pciAddress": "0000:80:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme35n1", + "path": "/dev/nvme35n1", + "pciAddress": "0000:81:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme36n1", + "path": "/dev/nvme36n1", + "pciAddress": "0000:82:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme37n1", + "path": "/dev/nvme37n1", + "pciAddress": "0000:83:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme38n1", + "path": "/dev/nvme38n1", + "pciAddress": "0000:84:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme39n1", + "path": "/dev/nvme39n1", + "pciAddress": "0000:85:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme40n1", + "path": "/dev/nvme40n1", + "pciAddress": "0000:86:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme41n1", + "path": "/dev/nvme41n1", + "pciAddress": "0000:87:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme42n1", + "path": "/dev/nvme42n1", + "pciAddress": "0000:88:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme43n1", + "path": "/dev/nvme43n1", + "pciAddress": "0000:89:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme44n1", + "path": "/dev/nvme44n1", + "pciAddress": "0000:8a:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme45n1", + "path": "/dev/nvme45n1", + "pciAddress": "0000:8b:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme46n1", + "path": "/dev/nvme46n1", + "pciAddress": "0000:8c:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme47n1", + "path": "/dev/nvme47n1", + "pciAddress": "0000:8d:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme48n1", + "path": "/dev/nvme48n1", + "pciAddress": "0000:8e:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme49n1", + "path": "/dev/nvme49n1", + "pciAddress": "0000:8f:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme50n1", + "path": "/dev/nvme50n1", + "pciAddress": "0000:90:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme51n1", + "path": "/dev/nvme51n1", + "pciAddress": "0000:91:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme52n1", + "path": "/dev/nvme52n1", + "pciAddress": "0000:92:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme53n1", + "path": "/dev/nvme53n1", + "pciAddress": "0000:93:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme54n1", + "path": "/dev/nvme54n1", + "pciAddress": "0000:94:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme55n1", + "path": "/dev/nvme55n1", + "pciAddress": "0000:95:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme56n1", + "path": "/dev/nvme56n1", + "pciAddress": "0000:96:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme57n1", + "path": "/dev/nvme57n1", + "pciAddress": "0000:97:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme58n1", + "path": "/dev/nvme58n1", + "pciAddress": "0000:98:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme59n1", + "path": "/dev/nvme59n1", + "pciAddress": "0000:99:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme60n1", + "path": "/dev/nvme60n1", + "pciAddress": "0000:9a:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme61n1", + "path": "/dev/nvme61n1", + "pciAddress": "0000:9b:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme62n1", + "path": "/dev/nvme62n1", + "pciAddress": "0000:9c:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme63n1", + "path": "/dev/nvme63n1", + "pciAddress": "0000:9d:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme64n1", + "path": "/dev/nvme64n1", + "pciAddress": "0000:9e:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme65n1", + "path": "/dev/nvme65n1", + "pciAddress": "0000:9f:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme66n1", + "path": "/dev/nvme66n1", + "pciAddress": "0000:a0:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme67n1", + "path": "/dev/nvme67n1", + "pciAddress": "0000:a1:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme68n1", + "path": "/dev/nvme68n1", + "pciAddress": "0000:a2:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme69n1", + "path": "/dev/nvme69n1", + "pciAddress": "0000:a3:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme70n1", + "path": "/dev/nvme70n1", + "pciAddress": "0000:a4:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme71n1", + "path": "/dev/nvme71n1", + "pciAddress": "0000:a5:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme72n1", + "path": "/dev/nvme72n1", + "pciAddress": "0000:a6:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme73n1", + "path": "/dev/nvme73n1", + "pciAddress": "0000:a7:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme74n1", + "path": "/dev/nvme74n1", + "pciAddress": "0000:a8:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme75n1", + "path": "/dev/nvme75n1", + "pciAddress": "0000:a9:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme76n1", + "path": "/dev/nvme76n1", + "pciAddress": "0000:aa:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme77n1", + "path": "/dev/nvme77n1", + "pciAddress": "0000:ab:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme78n1", + "path": "/dev/nvme78n1", + "pciAddress": "0000:ac:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme79n1", + "path": "/dev/nvme79n1", + "pciAddress": "0000:ad:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme80n1", + "path": "/dev/nvme80n1", + "pciAddress": "0000:ae:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme81n1", + "path": "/dev/nvme81n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme82n1", + "path": "/dev/nvme82n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme83n1", + "path": "/dev/nvme83n1", + "pciAddress": "0000:b1:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme84n1", + "path": "/dev/nvme84n1", + "pciAddress": "0000:b2:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme85n1", + "path": "/dev/nvme85n1", + "pciAddress": "0000:b3:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme86n1", + "path": "/dev/nvme86n1", + "pciAddress": "0000:b4:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme87n1", + "path": "/dev/nvme87n1", + "pciAddress": "0000:b5:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme88n1", + "path": "/dev/nvme88n1", + "pciAddress": "0000:b6:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme89n1", + "path": "/dev/nvme89n1", + "pciAddress": "0000:b7:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme90n1", + "path": "/dev/nvme90n1", + "pciAddress": "0000:b8:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme91n1", + "path": "/dev/nvme91n1", + "pciAddress": "0000:b9:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme92n1", + "path": "/dev/nvme92n1", + "pciAddress": "0000:ba:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme93n1", + "path": "/dev/nvme93n1", + "pciAddress": "0000:bb:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme94n1", + "path": "/dev/nvme94n1", + "pciAddress": "0000:bc:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme95n1", + "path": "/dev/nvme95n1", + "pciAddress": "0000:bd:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme96n1", + "path": "/dev/nvme96n1", + "pciAddress": "0000:be:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme97n1", + "path": "/dev/nvme97n1", + "pciAddress": "0000:bf:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme98n1", + "path": "/dev/nvme98n1", + "pciAddress": "0000:c0:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme99n1", + "path": "/dev/nvme99n1", + "pciAddress": "0000:c1:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme100n1", + "path": "/dev/nvme100n1", + "pciAddress": "0000:c2:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme101n1", + "path": "/dev/nvme101n1", + "pciAddress": "0000:c3:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme102n1", + "path": "/dev/nvme102n1", + "pciAddress": "0000:c4:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme103n1", + "path": "/dev/nvme103n1", + "pciAddress": "0000:c5:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme104n1", + "path": "/dev/nvme104n1", + "pciAddress": "0000:c6:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme105n1", + "path": "/dev/nvme105n1", + "pciAddress": "0000:c7:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme106n1", + "path": "/dev/nvme106n1", + "pciAddress": "0000:c8:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme107n1", + "path": "/dev/nvme107n1", + "pciAddress": "0000:c9:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme108n1", + "path": "/dev/nvme108n1", + "pciAddress": "0000:ca:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme109n1", + "path": "/dev/nvme109n1", + "pciAddress": "0000:cb:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme110n1", + "path": "/dev/nvme110n1", + "pciAddress": "0000:cc:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme111n1", + "path": "/dev/nvme111n1", + "pciAddress": "0000:cd:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme112n1", + "path": "/dev/nvme112n1", + "pciAddress": "0000:ce:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme113n1", + "path": "/dev/nvme113n1", + "pciAddress": "0000:cf:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme114n1", + "path": "/dev/nvme114n1", + "pciAddress": "0000:d0:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme115n1", + "path": "/dev/nvme115n1", + "pciAddress": "0000:d1:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme116n1", + "path": "/dev/nvme116n1", + "pciAddress": "0000:d2:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme117n1", + "path": "/dev/nvme117n1", + "pciAddress": "0000:d3:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme118n1", + "path": "/dev/nvme118n1", + "pciAddress": "0000:d4:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme119n1", + "path": "/dev/nvme119n1", + "pciAddress": "0000:d5:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme120n1", + "path": "/dev/nvme120n1", + "pciAddress": "0000:d6:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme121n1", + "path": "/dev/nvme121n1", + "pciAddress": "0000:d7:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme122n1", + "path": "/dev/nvme122n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme123n1", + "path": "/dev/nvme123n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme124n1", + "path": "/dev/nvme124n1", + "pciAddress": "0000:da:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme125n1", + "path": "/dev/nvme125n1", + "pciAddress": "0000:db:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme126n1", + "path": "/dev/nvme126n1", + "pciAddress": "0000:dc:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme127n1", + "path": "/dev/nvme127n1", + "pciAddress": "0000:dd:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-dev-11 + name: sb-nodeprobe-discover-dev-11-worker-01-cf2f6228 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/case.md b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/case.md new file mode 100644 index 000000000..212419c7a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/case.md @@ -0,0 +1,7 @@ +# DEV-12 + +**Mutation.** A disk whose `Kind` is `disk` and whose `SizeBytes` is 0 + +**Expected.** Admitted. The group is named `unsized` + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/expected-notes.txt new file mode 100644 index 000000000..6f1896307 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-dev-12 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/expected.yaml new file mode 100644 index 000000000..cd4d32dfa --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/expected.yaml @@ -0,0 +1,30 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-dev-12 + name: discovered-discover-dev-12 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-dev-12-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + mgmtInterface: eth0 + name: group-1-nvme-2xunsized + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/ops.yaml new file mode 100644 index 000000000..03ebda963 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-dev-12 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/reports/worker-01.yaml new file mode 100644 index 000000000..3de4170fe --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/reports/worker-01.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 0, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 0, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-dev-12 + name: sb-nodeprobe-discover-dev-12-worker-01-b9b1da6b + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/case.md b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/case.md new file mode 100644 index 000000000..301e219b8 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/case.md @@ -0,0 +1,7 @@ +# DEV-13 + +**Mutation.** A simplyblock volume attached to the worker, NVMe run + +**Expected.** Refused as this fleet's own volume, naming the cluster it belongs to + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/expected-notes.txt new file mode 100644 index 000000000..ee54f0b0f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-dev-13 in Draft: 1 workers with 1 nvme devices, in 1 group(s) across 1 node set(s); 1 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/expected-refusals.txt new file mode 100644 index 000000000..81fd7794a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/expected-refusals.txt @@ -0,0 +1 @@ +worker-01/nvme3n1: declined by simplyblock volume because it is a simplyblock logical volume of cluster c30a691a-1d2e-4f3a-9b8c-5d6e7f809a1b, which this fleet already serves diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/expected.yaml new file mode 100644 index 000000000..aa815d4d8 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/expected.yaml @@ -0,0 +1,29 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-dev-13 + name: discovered-discover-dev-13 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-dev-13-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + mgmtInterface: eth0 + name: group-1-nvme-1x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/ops.yaml new file mode 100644 index 000000000..885931e1a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-dev-13 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/reports/worker-01.yaml new file mode 100644 index 000000000..1072ab96b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/reports/worker-01.yaml @@ -0,0 +1,155 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "sizeBytes": 1099511627776, + "kind": "Disk", + "transport": "NVMeFabric", + "numaNode": 0, + "subsystemNQN": "nqn.2023-02.io.simplyblock:c30a691a-1d2e-4f3a-9b8c-5d6e7f809a1b:lvol:792e184c-0a1b-2c3d-4e5f-60718293a4b5", + "available": false, + "rejections": [ + { + "reason": "FabricNamespace", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-dev-13 + name: sb-nodeprobe-discover-dev-13-worker-01-7d0f45fe + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/case.md b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/case.md new file mode 100644 index 000000000..5ad02f414 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/case.md @@ -0,0 +1,7 @@ +# DEV-14 + +**Mutation.** The same volume on a block run + +**Expected.** Refused the same way, since the rule is in both pipelines and reads no filter + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/expected-notes.txt new file mode 100644 index 000000000..ae524ac76 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-dev-14 in Draft: 1 workers with 1 block devices, in 1 group(s) across 1 node set(s); 1 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/expected-refusals.txt new file mode 100644 index 000000000..81fd7794a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/expected-refusals.txt @@ -0,0 +1 @@ +worker-01/nvme3n1: declined by simplyblock volume because it is a simplyblock logical volume of cluster c30a691a-1d2e-4f3a-9b8c-5d6e7f809a1b, which this fleet already serves diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/expected.yaml new file mode 100644 index 000000000..270d13716 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/expected.yaml @@ -0,0 +1,29 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-dev-14 + name: discovered-discover-dev-14 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-dev-14-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + block: + - /dev/vdb + mgmtInterface: eth0 + name: group-1-block-1x2T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/ops.yaml new file mode 100644 index 000000000..454cb70ef --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/ops.yaml @@ -0,0 +1,15 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-dev-14 + namespace: simplyblock +spec: + action: Discover + discover: + deviceFilter: + enableLogicalBlockDevices: true +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/reports/worker-01.yaml new file mode 100644 index 000000000..6ccf9aa59 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/reports/worker-01.yaml @@ -0,0 +1,153 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "sizeBytes": 1099511627776, + "kind": "Disk", + "transport": "NVMeFabric", + "numaNode": 0, + "subsystemNQN": "nqn.2023-02.io.simplyblock:c30a691a-1d2e-4f3a-9b8c-5d6e7f809a1b:lvol:792e184c-0a1b-2c3d-4e5f-60718293a4b5", + "available": false, + "rejections": [ + { + "reason": "FabricNamespace", + "detail": "as the probe found it" + } + ] + }, + { + "name": "vdb", + "path": "/dev/vdb", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-dev-14 + name: sb-nodeprobe-discover-dev-14-worker-01-17707fee + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-15-every-disk-is-an-attached-volume/case.md b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-15-every-disk-is-an-attached-volume/case.md new file mode 100644 index 000000000..6e45cc771 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-15-every-disk-is-an-attached-volume/case.md @@ -0,0 +1,7 @@ +# DEV-15 + +**Mutation.** A worker whose every disk is an attached volume + +**Expected.** No draft, and the explanation counts them together rather than listing each + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-15-every-disk-is-an-attached-volume/expected-error.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-15-every-disk-is-an-attached-volume/expected-error.txt new file mode 100644 index 000000000..c73155346 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-15-every-disk-is-an-attached-volume/expected-error.txt @@ -0,0 +1 @@ +no worker has a device this run would use: worker-01: no device of it survived the device rules (3 devices declined by simplyblock volume: it is a simplyblock logical volume of cluster c30a691a-1d2e-4f3a-9b8c-5d6e7f809a1b, which this fleet already serves) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-15-every-disk-is-an-attached-volume/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-15-every-disk-is-an-attached-volume/expected-refusals.txt new file mode 100644 index 000000000..fe214fe3a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-15-every-disk-is-an-attached-volume/expected-refusals.txt @@ -0,0 +1,4 @@ +worker-01/nvme3n1: declined by simplyblock volume because it is a simplyblock logical volume of cluster c30a691a-1d2e-4f3a-9b8c-5d6e7f809a1b, which this fleet already serves +worker-01/nvme4n1: declined by simplyblock volume because it is a simplyblock logical volume of cluster c30a691a-1d2e-4f3a-9b8c-5d6e7f809a1b, which this fleet already serves +worker-01/nvme5n1: declined by simplyblock volume because it is a simplyblock logical volume of cluster c30a691a-1d2e-4f3a-9b8c-5d6e7f809a1b, which this fleet already serves +worker-01: declined by has devices because no device of it survived the device rules diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-15-every-disk-is-an-attached-volume/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-15-every-disk-is-an-attached-volume/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-15-every-disk-is-an-attached-volume/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-15-every-disk-is-an-attached-volume/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-15-every-disk-is-an-attached-volume/ops.yaml new file mode 100644 index 000000000..de950c5b5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-15-every-disk-is-an-attached-volume/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-dev-15 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-15-every-disk-is-an-attached-volume/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-15-every-disk-is-an-attached-volume/reports/worker-01.yaml new file mode 100644 index 000000000..38fe0e1e0 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-15-every-disk-is-an-attached-volume/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "sizeBytes": 1099511627776, + "kind": "Disk", + "transport": "NVMeFabric", + "numaNode": 0, + "subsystemNQN": "nqn.2023-02.io.simplyblock:c30a691a-1d2e-4f3a-9b8c-5d6e7f809a1b:lvol:792e184c-0a1b-2c3d-4e5f-60718293a4b5", + "available": false, + "rejections": [ + { + "reason": "FabricNamespace", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme4n1", + "path": "/dev/nvme4n1", + "sizeBytes": 1099511627776, + "kind": "Disk", + "transport": "NVMeFabric", + "numaNode": 0, + "subsystemNQN": "nqn.2023-02.io.simplyblock:c30a691a-1d2e-4f3a-9b8c-5d6e7f809a1b:lvol:8a3f0b12-3c4d-5e6f-7081-92a3b4c5d6e7", + "available": false, + "rejections": [ + { + "reason": "FabricNamespace", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme5n1", + "path": "/dev/nvme5n1", + "sizeBytes": 1099511627776, + "kind": "Disk", + "transport": "NVMeFabric", + "numaNode": 0, + "subsystemNQN": "nqn.2023-02.io.simplyblock:c30a691a-1d2e-4f3a-9b8c-5d6e7f809a1b:lvol:b1c2d3e4-f506-1728-394a-5b6c7d8e9f01", + "available": false, + "rejections": [ + { + "reason": "FabricNamespace", + "detail": "as the probe found it" + } + ] + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-dev-15 + name: sb-nodeprobe-discover-dev-15-worker-01-f280cf1e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/case.md b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/case.md new file mode 100644 index 000000000..2ca97f907 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/case.md @@ -0,0 +1,7 @@ +# DEV-16 + +**Mutation.** A fabric namespace another product exported + +**Expected.** Refused for being on a fabric, not as a simplyblock volume: the NQN does not parse as one + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/expected-notes.txt new file mode 100644 index 000000000..3f2b3a7c1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-dev-16 in Draft: 1 workers with 1 nvme devices, in 1 group(s) across 1 node set(s); 1 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/expected-refusals.txt new file mode 100644 index 000000000..047ac5d21 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/expected-refusals.txt @@ -0,0 +1 @@ +worker-01/nvme3n1: declined by device class because this run scans NVMe devices and the device is on NVMeFabric diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/expected.yaml new file mode 100644 index 000000000..ef64f524d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/expected.yaml @@ -0,0 +1,29 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-dev-16 + name: discovered-discover-dev-16 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-dev-16-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + mgmtInterface: eth0 + name: group-1-nvme-1x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/ops.yaml new file mode 100644 index 000000000..f104e8811 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-dev-16 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/reports/worker-01.yaml new file mode 100644 index 000000000..f504de91a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/reports/worker-01.yaml @@ -0,0 +1,155 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "sizeBytes": 1099511627776, + "kind": "Disk", + "transport": "NVMeFabric", + "numaNode": 0, + "subsystemNQN": "nqn.2019-08.org.ceph:rbd.pool.image", + "available": false, + "rejections": [ + { + "reason": "FabricNamespace", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-dev-16 + name: sb-nodeprobe-discover-dev-16-worker-01-3e7ff2bf + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/case.md b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/case.md new file mode 100644 index 000000000..be136313a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/case.md @@ -0,0 +1,7 @@ +# DEV-17 + +**Mutation.** An iSCSI LUN beside a virtio disk, block run, no allow list + +**Expected.** Only the virtio disk. A LUN is storage across a network and is never taken by default + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/expected-notes.txt new file mode 100644 index 000000000..d9d261829 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-dev-17 in Draft: 1 workers with 1 block devices, in 1 group(s) across 1 node set(s); 1 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/expected-refusals.txt new file mode 100644 index 000000000..987914354 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/expected-refusals.txt @@ -0,0 +1 @@ +worker-01/sdb: declined by iSCSI because it is an iSCSI LUN, which is storage on the other side of a network, so a run takes one only where the allow list names it and this one names nothing diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/expected.yaml new file mode 100644 index 000000000..45307ac1c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/expected.yaml @@ -0,0 +1,29 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-dev-17 + name: discovered-discover-dev-17 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-dev-17-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + block: + - /dev/vdb + mgmtInterface: eth0 + name: group-1-block-1x2T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/ops.yaml new file mode 100644 index 000000000..e9bf06e2f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/ops.yaml @@ -0,0 +1,15 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-dev-17 + namespace: simplyblock +spec: + action: Discover + discover: + deviceFilter: + enableLogicalBlockDevices: true +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/reports/worker-01.yaml new file mode 100644 index 000000000..8ff873afd --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/reports/worker-01.yaml @@ -0,0 +1,147 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "sdb", + "path": "/dev/sdb", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "iSCSI", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "vdb", + "path": "/dev/vdb", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-dev-17 + name: sb-nodeprobe-discover-dev-17-worker-01-4b81a475 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/case.md b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/case.md new file mode 100644 index 000000000..bfd585f73 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/case.md @@ -0,0 +1,7 @@ +# DEV-18 + +**Mutation.** The same worker with the allow list naming the LUN + +**Expected.** Both disks. Naming it is the decision a run cannot make for a fleet + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/expected-notes.txt new file mode 100644 index 000000000..3b6e63f4f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-dev-18 in Draft: 1 workers with 2 block devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/expected.yaml new file mode 100644 index 000000000..6c538117b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/expected.yaml @@ -0,0 +1,30 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-dev-18 + name: discovered-discover-dev-18 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-dev-18-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + block: + - /dev/sdb + - /dev/vdb + mgmtInterface: eth0 + name: group-1-block-2x2T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/ops.yaml new file mode 100644 index 000000000..c618bc26e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/ops.yaml @@ -0,0 +1,18 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-dev-18 + namespace: simplyblock +spec: + action: Discover + discover: + deviceFilter: + blockAllowList: + - /dev/sdb + - /dev/vdb + enableLogicalBlockDevices: true +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/reports/worker-01.yaml new file mode 100644 index 000000000..d43b82b6c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/reports/worker-01.yaml @@ -0,0 +1,147 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "sdb", + "path": "/dev/sdb", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "iSCSI", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "vdb", + "path": "/dev/vdb", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-dev-18 + name: sb-nodeprobe-discover-dev-18-worker-01-36d529c3 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/case.md b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/case.md new file mode 100644 index 000000000..b3354837f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/case.md @@ -0,0 +1,7 @@ +# DEV-19 + +**Mutation.** An iSCSI LUN on an NVMe run + +**Expected.** Refused for being the other class, before the iSCSI rule is reached + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/expected-notes.txt new file mode 100644 index 000000000..7ad132066 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-dev-19 in Draft: 1 workers with 1 nvme devices, in 1 group(s) across 1 node set(s); 1 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/expected-refusals.txt new file mode 100644 index 000000000..b0716fcfc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/expected-refusals.txt @@ -0,0 +1 @@ +worker-01/sdb: declined by device class because this run scans NVMe devices and the device is on iSCSI diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/expected.yaml new file mode 100644 index 000000000..7f61d3c99 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/expected.yaml @@ -0,0 +1,29 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-dev-19 + name: discovered-discover-dev-19 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-dev-19-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + mgmtInterface: eth0 + name: group-1-nvme-1x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/ops.yaml new file mode 100644 index 000000000..289420eca --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-dev-19 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/reports/worker-01.yaml new file mode 100644 index 000000000..90d9fd1c8 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/reports/worker-01.yaml @@ -0,0 +1,149 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "sdb", + "path": "/dev/sdb", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "iSCSI", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-dev-19 + name: sb-nodeprobe-discover-dev-19-worker-01-7a631456 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/case.md b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/case.md new file mode 100644 index 000000000..86a2601e8 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/case.md @@ -0,0 +1,7 @@ +# FAIL-01 + +**Mutation.** Every worker reports no devices and no controllers + +**Expected.** The run fails: no worker has a device this run would use, one line per worker + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/expected-error.txt b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/expected-error.txt new file mode 100644 index 000000000..1ac92bcf3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/expected-error.txt @@ -0,0 +1 @@ +no worker has a device this run would use: worker-01: no device of it survived the device rules; worker-02: no device of it survived the device rules; worker-03: no device of it survived the device rules diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/expected-refusals.txt new file mode 100644 index 000000000..17045595c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/expected-refusals.txt @@ -0,0 +1,3 @@ +worker-01: declined by has devices because no device of it survived the device rules +worker-02: declined by has devices because no device of it survived the device rules +worker-03: declined by has devices because no device of it survived the device rules diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/nodes.yaml new file mode 100644 index 000000000..98de08ee9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/nodes.yaml @@ -0,0 +1,89 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/ops.yaml new file mode 100644 index 000000000..9ce26a46f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/ops.yaml @@ -0,0 +1,14 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-fail-01 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/reports/worker-01.yaml new file mode 100644 index 000000000..2c1824fa3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/reports/worker-01.yaml @@ -0,0 +1,125 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-fail-01 + name: sb-nodeprobe-discover-fail-01-worker-01-d59552a0 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/reports/worker-02.yaml new file mode 100644 index 000000000..212e3678b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/reports/worker-02.yaml @@ -0,0 +1,125 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-fail-01 + name: sb-nodeprobe-discover-fail-01-worker-02-ec92782f + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/reports/worker-03.yaml new file mode 100644 index 000000000..7a8707688 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-01-no-devices-and-no-controllers/reports/worker-03.yaml @@ -0,0 +1,125 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-fail-01 + name: sb-nodeprobe-discover-fail-01-worker-03-8568380e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/case.md b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/case.md new file mode 100644 index 000000000..b7e29af4e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/case.md @@ -0,0 +1,7 @@ +# FAIL-02 + +**Mutation.** Every disk in the fleet mounted + +**Expected.** The same, each line carrying the device reasons and their counts + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/expected-error.txt b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/expected-error.txt new file mode 100644 index 000000000..ee8f2c3b3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/expected-error.txt @@ -0,0 +1 @@ +no worker has a device this run would use: worker-01: no device of it survived the device rules (4 devices declined by available: the probe refused it: Mounted); worker-02: no device of it survived the device rules (4 devices declined by available: the probe refused it: Mounted); worker-03: no device of it survived the device rules (4 devices declined by available: the probe refused it: Mounted) diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/expected-refusals.txt new file mode 100644 index 000000000..775557fc8 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/expected-refusals.txt @@ -0,0 +1,15 @@ +worker-01/nvme0n1: declined by available because the probe refused it: Mounted +worker-01/nvme1n1: declined by available because the probe refused it: Mounted +worker-01/nvme2n1: declined by available because the probe refused it: Mounted +worker-01/nvme3n1: declined by available because the probe refused it: Mounted +worker-01: declined by has devices because no device of it survived the device rules +worker-02/nvme0n1: declined by available because the probe refused it: Mounted +worker-02/nvme1n1: declined by available because the probe refused it: Mounted +worker-02/nvme2n1: declined by available because the probe refused it: Mounted +worker-02/nvme3n1: declined by available because the probe refused it: Mounted +worker-02: declined by has devices because no device of it survived the device rules +worker-03/nvme0n1: declined by available because the probe refused it: Mounted +worker-03/nvme1n1: declined by available because the probe refused it: Mounted +worker-03/nvme2n1: declined by available because the probe refused it: Mounted +worker-03/nvme3n1: declined by available because the probe refused it: Mounted +worker-03: declined by has devices because no device of it survived the device rules diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/nodes.yaml new file mode 100644 index 000000000..98de08ee9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/nodes.yaml @@ -0,0 +1,89 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/ops.yaml new file mode 100644 index 000000000..5afd0ded2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/ops.yaml @@ -0,0 +1,14 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-fail-02 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/reports/worker-01.yaml new file mode 100644 index 000000000..1f1514300 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/reports/worker-01.yaml @@ -0,0 +1,195 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-fail-02 + name: sb-nodeprobe-discover-fail-02-worker-01-359dd640 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/reports/worker-02.yaml new file mode 100644 index 000000000..e4943d8af --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/reports/worker-02.yaml @@ -0,0 +1,195 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-fail-02 + name: sb-nodeprobe-discover-fail-02-worker-02-9bb52b28 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/reports/worker-03.yaml new file mode 100644 index 000000000..5fb4e1642 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/reports/worker-03.yaml @@ -0,0 +1,195 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-fail-02 + name: sb-nodeprobe-discover-fail-02-worker-03-552265ab + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-03-every-report-unreadable/case.md b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-03-every-report-unreadable/case.md new file mode 100644 index 000000000..545c560fc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-03-every-report-unreadable/case.md @@ -0,0 +1,7 @@ +# FAIL-03 + +**Mutation.** Every report unreadable + +**Expected.** `ReportUnreadable` per ConfigMap, then the run fails for want of reports + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-03-every-report-unreadable/expected-error.txt b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-03-every-report-unreadable/expected-error.txt new file mode 100644 index 000000000..0ff97c845 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-03-every-report-unreadable/expected-error.txt @@ -0,0 +1 @@ +the probe reports are gone, so there is nothing to write diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-03-every-report-unreadable/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-03-every-report-unreadable/nodes.yaml new file mode 100644 index 000000000..98de08ee9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-03-every-report-unreadable/nodes.yaml @@ -0,0 +1,89 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-03-every-report-unreadable/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-03-every-report-unreadable/ops.yaml new file mode 100644 index 000000000..3b0052f1c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-03-every-report-unreadable/ops.yaml @@ -0,0 +1,14 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-fail-03 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-03-every-report-unreadable/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-03-every-report-unreadable/reports/worker-01.yaml new file mode 100644 index 000000000..9ce842d34 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-03-every-report-unreadable/reports/worker-01.yaml @@ -0,0 +1,11 @@ +apiVersion: v1 +data: + report.json: '{ this is not a report' +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-fail-03 + name: sb-nodeprobe-discover-fail-03-worker-01-fdd789f1 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-03-every-report-unreadable/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-03-every-report-unreadable/reports/worker-02.yaml new file mode 100644 index 000000000..c13338cdc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-03-every-report-unreadable/reports/worker-02.yaml @@ -0,0 +1,11 @@ +apiVersion: v1 +data: + report.json: '{ this is not a report' +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-fail-03 + name: sb-nodeprobe-discover-fail-03-worker-02-0c47ceec + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-03-every-report-unreadable/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-03-every-report-unreadable/reports/worker-03.yaml new file mode 100644 index 000000000..cc45afb78 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-03-every-report-unreadable/reports/worker-03.yaml @@ -0,0 +1,11 @@ +apiVersion: v1 +data: + report.json: '{ this is not a report' +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-fail-03 + name: sb-nodeprobe-discover-fail-03-worker-03-bec46d84 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/case.md b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/case.md new file mode 100644 index 000000000..3c4f77fb5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/case.md @@ -0,0 +1,7 @@ +# FAIL-04 + +**Mutation.** A filter that excludes every disk + +**Expected.** The run fails. The refusal lines name the filter rule + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/expected-error.txt b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/expected-error.txt new file mode 100644 index 000000000..b2a177666 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/expected-error.txt @@ -0,0 +1 @@ +no worker has a device this run would use: worker-01: no device of it survived the device rules (4 devices declined by size range: it is 3T and the range 100T-200T starts above it); worker-02: no device of it survived the device rules (4 devices declined by size range: it is 3T and the range 100T-200T starts above it); worker-03: no device of it survived the device rules (4 devices declined by size range: it is 3T and the range 100T-200T starts above it) diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/expected-refusals.txt new file mode 100644 index 000000000..056b10a0c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/expected-refusals.txt @@ -0,0 +1,15 @@ +worker-01/nvme0n1: declined by size range because it is 3T and the range 100T-200T starts above it +worker-01/nvme1n1: declined by size range because it is 3T and the range 100T-200T starts above it +worker-01/nvme2n1: declined by size range because it is 3T and the range 100T-200T starts above it +worker-01/nvme3n1: declined by size range because it is 3T and the range 100T-200T starts above it +worker-01: declined by has devices because no device of it survived the device rules +worker-02/nvme0n1: declined by size range because it is 3T and the range 100T-200T starts above it +worker-02/nvme1n1: declined by size range because it is 3T and the range 100T-200T starts above it +worker-02/nvme2n1: declined by size range because it is 3T and the range 100T-200T starts above it +worker-02/nvme3n1: declined by size range because it is 3T and the range 100T-200T starts above it +worker-02: declined by has devices because no device of it survived the device rules +worker-03/nvme0n1: declined by size range because it is 3T and the range 100T-200T starts above it +worker-03/nvme1n1: declined by size range because it is 3T and the range 100T-200T starts above it +worker-03/nvme2n1: declined by size range because it is 3T and the range 100T-200T starts above it +worker-03/nvme3n1: declined by size range because it is 3T and the range 100T-200T starts above it +worker-03: declined by has devices because no device of it survived the device rules diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/nodes.yaml new file mode 100644 index 000000000..98de08ee9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/nodes.yaml @@ -0,0 +1,89 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/ops.yaml new file mode 100644 index 000000000..934a7726c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/ops.yaml @@ -0,0 +1,17 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-fail-04 + namespace: simplyblock +spec: + action: Discover + discover: + deviceFilter: + driveSizeRange: 100T-200T +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/reports/worker-01.yaml new file mode 100644 index 000000000..b31ea954f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-fail-04 + name: sb-nodeprobe-discover-fail-04-worker-01-c8f720f7 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/reports/worker-02.yaml new file mode 100644 index 000000000..d7a44f93a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-fail-04 + name: sb-nodeprobe-discover-fail-04-worker-02-4982a317 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/reports/worker-03.yaml new file mode 100644 index 000000000..3554e90ed --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-fail-04 + name: sb-nodeprobe-discover-fail-04-worker-03-fb70a92b + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/case.md b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/case.md new file mode 100644 index 000000000..cc0cf600e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/case.md @@ -0,0 +1,9 @@ +# FAIL-05 + +**Mutation.** Disks fine, no usable interface anywhere + +**Expected.** **Contested.** A draft is written with an empty `mgmtInterface`. See §14, gap G-7 + +**Harness.** `CM` + +**Gap.** G-7. This case records what the generator does today, so that the day it changes the diff is the finding. diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/expected-notes.txt new file mode 100644 index 000000000..4bf46af0a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-fail-05 in Draft: 3 workers with 12 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/expected.yaml new file mode 100644 index 000000000..946e71c31 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/expected.yaml @@ -0,0 +1,33 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-fail-05 + name: discovered-discover-fail-05 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-fail-05-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - worker-02 + - worker-03 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/nodes.yaml new file mode 100644 index 000000000..c8d4bdfaf --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/nodes.yaml @@ -0,0 +1,80 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/ops.yaml new file mode 100644 index 000000000..bcc974e0f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/ops.yaml @@ -0,0 +1,14 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-fail-05 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/reports/worker-01.yaml new file mode 100644 index 000000000..768c8a253 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "cni0", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "bridge": true, + "kind": "bridge", + "addresses": [ + "10.42.2.1" + ] + }, + { + "name": "lo", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "loopback": true, + "kind": "loopback", + "addresses": [ + "127.0.0.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-fail-05 + name: sb-nodeprobe-discover-fail-05-worker-01-a45ee4ce + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/reports/worker-02.yaml new file mode 100644 index 000000000..693fa122f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "cni0", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "bridge": true, + "kind": "bridge", + "addresses": [ + "10.42.2.1" + ] + }, + { + "name": "lo", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "loopback": true, + "kind": "loopback", + "addresses": [ + "127.0.0.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-fail-05 + name: sb-nodeprobe-discover-fail-05-worker-02-d795fa99 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/reports/worker-03.yaml new file mode 100644 index 000000000..61a714bcb --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "cni0", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "bridge": true, + "kind": "bridge", + "addresses": [ + "10.42.2.1" + ] + }, + { + "name": "lo", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "loopback": true, + "kind": "loopback", + "addresses": [ + "127.0.0.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-fail-05 + name: sb-nodeprobe-discover-fail-05-worker-03-a922ee29 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/case.md b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/case.md new file mode 100644 index 000000000..9426dfda3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/case.md @@ -0,0 +1,9 @@ +# FAIL-06 + +**Mutation.** Every worker under the vCPU floor + +**Expected.** **Contested.** A draft is written with `vcpuCount: 4`. See §14, gap G-10 + +**Harness.** `CM` + +**Gap.** G-10. This case records what the generator does today, so that the day it changes the diff is the finding. diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/expected-notes.txt new file mode 100644 index 000000000..f227a61d4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is the API's minimum of 4, and the smallest placement (worker-01) has only 2 cores, so this cluster asks for more than that worker has +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-fail-06 in Draft: 3 workers with 12 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/expected.yaml new file mode 100644 index 000000000..ed5bf480c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/expected.yaml @@ -0,0 +1,34 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-fail-06 + name: discovered-discover-fail-06 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-fail-06-cluster + vcpuCount: 4 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - worker-02 + - worker-03 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/nodes.yaml new file mode 100644 index 000000000..98de08ee9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/nodes.yaml @@ -0,0 +1,89 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/ops.yaml new file mode 100644 index 000000000..ef8a7ce60 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/ops.yaml @@ -0,0 +1,14 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-fail-06 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/reports/worker-01.yaml new file mode 100644 index 000000000..643d1fb6b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/reports/worker-01.yaml @@ -0,0 +1,140 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 2, + "physicalCores": 2, + "sockets": 1, + "threadsPerCore": 1, + "hyperThreading": false, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1 + ], + "physicalCores": 2 + } + ] + }, + "memory": { + "totalBytes": 17179869184, + "freeBytes": 3758096384, + "availableBytes": 15032385536, + "numaNodes": [ + { + "node": 0, + "totalBytes": 17179869184, + "freeBytes": 8589934592 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-fail-06 + name: sb-nodeprobe-discover-fail-06-worker-01-17449dcf + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/reports/worker-02.yaml new file mode 100644 index 000000000..061a08d3e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/reports/worker-02.yaml @@ -0,0 +1,140 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 2, + "physicalCores": 2, + "sockets": 1, + "threadsPerCore": 1, + "hyperThreading": false, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1 + ], + "physicalCores": 2 + } + ] + }, + "memory": { + "totalBytes": 17179869184, + "freeBytes": 3758096384, + "availableBytes": 15032385536, + "numaNodes": [ + { + "node": 0, + "totalBytes": 17179869184, + "freeBytes": 8589934592 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-fail-06 + name: sb-nodeprobe-discover-fail-06-worker-02-e7325f7e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/reports/worker-03.yaml new file mode 100644 index 000000000..1c571115c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/reports/worker-03.yaml @@ -0,0 +1,140 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 2, + "physicalCores": 2, + "sockets": 1, + "threadsPerCore": 1, + "hyperThreading": false, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1 + ], + "physicalCores": 2 + } + ] + }, + "memory": { + "totalBytes": 17179869184, + "freeBytes": 3758096384, + "availableBytes": 15032385536, + "numaNodes": [ + { + "node": 0, + "totalBytes": 17179869184, + "freeBytes": 8589934592 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-fail-06 + name: sb-nodeprobe-discover-fail-06-worker-03-13750814 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/case.md b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/case.md new file mode 100644 index 000000000..bb4cec36c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/case.md @@ -0,0 +1,9 @@ +# FAIL-07 + +**Mutation.** Every worker at 4 GiB of RAM + +**Expected.** **Contested.** A draft is written unchanged. See §14, gap G-4 + +**Harness.** `CM` + +**Gap.** G-4. This case records what the generator does today, so that the day it changes the diff is the finding. diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/expected-notes.txt new file mode 100644 index 000000000..3f41158c5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 4, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-fail-07 in Draft: 3 workers with 12 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/expected.yaml new file mode 100644 index 000000000..eb45e0ab1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/expected.yaml @@ -0,0 +1,34 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-fail-07 + name: discovered-discover-fail-07 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-fail-07-cluster + vcpuCount: 4 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - worker-02 + - worker-03 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/nodes.yaml new file mode 100644 index 000000000..98de08ee9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/nodes.yaml @@ -0,0 +1,89 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/ops.yaml new file mode 100644 index 000000000..cfefbd345 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/ops.yaml @@ -0,0 +1,14 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-fail-07 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/reports/worker-01.yaml new file mode 100644 index 000000000..a901584b2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/reports/worker-01.yaml @@ -0,0 +1,146 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 8, + "physicalCores": 4, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7 + ], + "physicalCores": 4 + } + ] + }, + "memory": { + "totalBytes": 4294967296, + "freeBytes": 235929600, + "availableBytes": 943718400, + "numaNodes": [ + { + "node": 0, + "totalBytes": 4294967296, + "freeBytes": 2147483648 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-fail-07 + name: sb-nodeprobe-discover-fail-07-worker-01-43813dc5 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/reports/worker-02.yaml new file mode 100644 index 000000000..7a4b2079c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/reports/worker-02.yaml @@ -0,0 +1,146 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 8, + "physicalCores": 4, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7 + ], + "physicalCores": 4 + } + ] + }, + "memory": { + "totalBytes": 4294967296, + "freeBytes": 235929600, + "availableBytes": 943718400, + "numaNodes": [ + { + "node": 0, + "totalBytes": 4294967296, + "freeBytes": 2147483648 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-fail-07 + name: sb-nodeprobe-discover-fail-07-worker-02-d0cfc1e2 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/reports/worker-03.yaml new file mode 100644 index 000000000..dffee50cf --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/reports/worker-03.yaml @@ -0,0 +1,146 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 8, + "physicalCores": 4, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7 + ], + "physicalCores": 4 + } + ] + }, + "memory": { + "totalBytes": 4294967296, + "freeBytes": 235929600, + "availableBytes": 943718400, + "numaNodes": [ + { + "node": 0, + "totalBytes": 4294967296, + "freeBytes": 2147483648 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-fail-07 + name: sb-nodeprobe-discover-fail-07-worker-03-cc978276 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/case.md b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/case.md new file mode 100644 index 000000000..79fecda1b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/case.md @@ -0,0 +1,7 @@ +# FAIL-08 + +**Mutation.** A single worker whose CPU tree is unreadable, disks fine + +**Expected.** A draft with `vcpuCount: 4` and the note saying no worker reported its cores + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/expected-notes.txt new file mode 100644 index 000000000..6af1eca4d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is the API's minimum of 4, because no worker reported the cores of the memory node it was placed on +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-fail-08 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/expected.yaml new file mode 100644 index 000000000..f0ec1216f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-fail-08 + name: discovered-discover-fail-08 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-fail-08-cluster + vcpuCount: 4 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/ops.yaml new file mode 100644 index 000000000..9f4bf4178 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-fail-08 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/reports/worker-01.yaml new file mode 100644 index 000000000..219c21ac5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/reports/worker-01.yaml @@ -0,0 +1,138 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 0, + "physicalCores": 0, + "sockets": 0, + "threadsPerCore": 0, + "hyperThreading": false + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ], + "unreadable": [ + "read /sys/devices/system/cpu: permission denied" + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-fail-08 + name: sb-nodeprobe-discover-fail-08-worker-01-ba59f814 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/case.md b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/case.md new file mode 100644 index 000000000..ce6266041 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/case.md @@ -0,0 +1,7 @@ +# FAIL-09 + +**Mutation.** One worker of 32 refused, the rest fine + +**Expected.** A draft of 31. The refusal is an event, never a failure + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/expected-notes.txt new file mode 100644 index 000000000..a9f48baf6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-fail-09 in Draft: 31 workers with 124 nvme devices, in 1 group(s) across 1 node set(s); 5 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/expected-refusals.txt new file mode 100644 index 000000000..47f0b0708 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/expected-refusals.txt @@ -0,0 +1,5 @@ +worker-17/nvme0n1: declined by available because the probe refused it: Mounted +worker-17/nvme1n1: declined by available because the probe refused it: Mounted +worker-17/nvme2n1: declined by available because the probe refused it: Mounted +worker-17/nvme3n1: declined by available because the probe refused it: Mounted +worker-17: declined by has devices because no device of it survived the device rules diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/expected.yaml new file mode 100644 index 000000000..37d662263 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/expected.yaml @@ -0,0 +1,62 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: discovered-discover-fail-09 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-fail-09-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - worker-02 + - worker-03 + - worker-04 + - worker-05 + - worker-06 + - worker-07 + - worker-08 + - worker-09 + - worker-10 + - worker-11 + - worker-12 + - worker-13 + - worker-14 + - worker-15 + - worker-16 + - worker-18 + - worker-19 + - worker-20 + - worker-21 + - worker-22 + - worker-23 + - worker-24 + - worker-25 + - worker-26 + - worker-27 + - worker-28 + - worker-29 + - worker-30 + - worker-31 + - worker-32 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/nodes.yaml new file mode 100644 index 000000000..c25e3f497 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/nodes.yaml @@ -0,0 +1,959 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-04 +spec: {} +status: + addresses: + - address: 10.10.10.4 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-05 +spec: {} +status: + addresses: + - address: 10.10.10.5 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-06 +spec: {} +status: + addresses: + - address: 10.10.10.6 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-07 +spec: {} +status: + addresses: + - address: 10.10.10.7 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-08 +spec: {} +status: + addresses: + - address: 10.10.10.8 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-09 +spec: {} +status: + addresses: + - address: 10.10.10.9 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-10 +spec: {} +status: + addresses: + - address: 10.10.10.10 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-11 +spec: {} +status: + addresses: + - address: 10.10.10.11 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-12 +spec: {} +status: + addresses: + - address: 10.10.10.12 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-13 +spec: {} +status: + addresses: + - address: 10.10.10.13 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-14 +spec: {} +status: + addresses: + - address: 10.10.10.14 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-15 +spec: {} +status: + addresses: + - address: 10.10.10.15 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-16 +spec: {} +status: + addresses: + - address: 10.10.10.16 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-17 +spec: {} +status: + addresses: + - address: 10.10.10.17 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-18 +spec: {} +status: + addresses: + - address: 10.10.10.18 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-19 +spec: {} +status: + addresses: + - address: 10.10.10.19 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-20 +spec: {} +status: + addresses: + - address: 10.10.10.20 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-21 +spec: {} +status: + addresses: + - address: 10.10.10.21 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-22 +spec: {} +status: + addresses: + - address: 10.10.10.22 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-23 +spec: {} +status: + addresses: + - address: 10.10.10.23 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-24 +spec: {} +status: + addresses: + - address: 10.10.10.24 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-25 +spec: {} +status: + addresses: + - address: 10.10.10.25 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-26 +spec: {} +status: + addresses: + - address: 10.10.10.26 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-27 +spec: {} +status: + addresses: + - address: 10.10.10.27 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-28 +spec: {} +status: + addresses: + - address: 10.10.10.28 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-29 +spec: {} +status: + addresses: + - address: 10.10.10.29 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-30 +spec: {} +status: + addresses: + - address: 10.10.10.30 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-31 +spec: {} +status: + addresses: + - address: 10.10.10.31 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-32 +spec: {} +status: + addresses: + - address: 10.10.10.32 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/ops.yaml new file mode 100644 index 000000000..b1512ef94 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/ops.yaml @@ -0,0 +1,43 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-fail-09 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 + - worker-04 + - worker-05 + - worker-06 + - worker-07 + - worker-08 + - worker-09 + - worker-10 + - worker-11 + - worker-12 + - worker-13 + - worker-14 + - worker-15 + - worker-16 + - worker-17 + - worker-18 + - worker-19 + - worker-20 + - worker-21 + - worker-22 + - worker-23 + - worker-24 + - worker-25 + - worker-26 + - worker-27 + - worker-28 + - worker-29 + - worker-30 + - worker-31 + - worker-32 diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-01.yaml new file mode 100644 index 000000000..cf91202ab --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-01-bb6ca5cd + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-02.yaml new file mode 100644 index 000000000..e797fe1d5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-02-2de5b1f0 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-03.yaml new file mode 100644 index 000000000..6f09e1763 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-03-d8171272 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-04.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-04.yaml new file mode 100644 index 000000000..8b32391fc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-04.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-04", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.4" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.4" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-04 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-04-27c84795 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-05.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-05.yaml new file mode 100644 index 000000000..6785305ca --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-05.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-05", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.5" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.5" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-05 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-05-961bfe00 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-06.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-06.yaml new file mode 100644 index 000000000..44945ceb8 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-06.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-06", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.6" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.6" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-06 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-06-4e5b0ce2 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-07.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-07.yaml new file mode 100644 index 000000000..17720e32c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-07.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-07", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.7" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.7" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-07 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-07-2f0ff4d0 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-08.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-08.yaml new file mode 100644 index 000000000..a0abddae8 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-08.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-08", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.8" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.8" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-08 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-08-3a93ced3 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-09.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-09.yaml new file mode 100644 index 000000000..c462e823c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-09.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-09", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.9" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.9" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-09 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-09-72d2784f + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-10.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-10.yaml new file mode 100644 index 000000000..9214403b1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-10.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-10", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.10" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.10" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-10 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-10-76fc532e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-11.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-11.yaml new file mode 100644 index 000000000..904f0ed59 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-11.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-11", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.11" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.11" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-11 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-11-5c0dd0b2 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-12.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-12.yaml new file mode 100644 index 000000000..3fe4c3a14 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-12.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-12", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.12" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.12" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-12 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-12-5f8049b0 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-13.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-13.yaml new file mode 100644 index 000000000..c5e7746bc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-13.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-13", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.13" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.13" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-13 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-13-3ebb6dc2 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-14.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-14.yaml new file mode 100644 index 000000000..e3f7ee4f2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-14.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-14", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.14" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.14" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-14 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-14-3c20c71e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-15.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-15.yaml new file mode 100644 index 000000000..84bba5495 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-15.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-15", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.15" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.15" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-15 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-15-32de2353 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-16.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-16.yaml new file mode 100644 index 000000000..c04a0a423 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-16.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-16", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.16" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.16" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-16 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-16-978ec102 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-17.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-17.yaml new file mode 100644 index 000000000..bce04ec83 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-17.yaml @@ -0,0 +1,195 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-17", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.17" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.17" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-17 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-17-ae948e5e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-18.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-18.yaml new file mode 100644 index 000000000..a29eccedd --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-18.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-18", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.18" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.18" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-18 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-18-e8ad36a1 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-19.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-19.yaml new file mode 100644 index 000000000..fd25b0dd9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-19.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-19", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.19" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.19" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-19 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-19-22f92df7 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-20.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-20.yaml new file mode 100644 index 000000000..20bf57ef2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-20.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-20", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.20" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.20" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-20 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-20-ced5ee46 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-21.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-21.yaml new file mode 100644 index 000000000..e40c19d10 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-21.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-21", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.21" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.21" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-21 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-21-e2789752 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-22.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-22.yaml new file mode 100644 index 000000000..abdc5070c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-22.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-22", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.22" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.22" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-22 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-22-92eada9a + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-23.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-23.yaml new file mode 100644 index 000000000..01a62b47e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-23.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-23", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.23" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.23" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-23 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-23-83775843 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-24.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-24.yaml new file mode 100644 index 000000000..1a7d674bb --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-24.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-24", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.24" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.24" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-24 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-24-bbb08a99 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-25.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-25.yaml new file mode 100644 index 000000000..91c883226 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-25.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-25", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.25" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.25" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-25 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-25-fdba7cf8 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-26.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-26.yaml new file mode 100644 index 000000000..96cecd194 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-26.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-26", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.26" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.26" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-26 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-26-c9bf700f + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-27.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-27.yaml new file mode 100644 index 000000000..66ca9d922 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-27.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-27", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.27" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.27" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-27 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-27-50b1b22c + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-28.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-28.yaml new file mode 100644 index 000000000..56c3914f1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-28.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-28", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.28" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.28" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-28 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-28-1b04bb93 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-29.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-29.yaml new file mode 100644 index 000000000..2ae40695b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-29.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-29", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.29" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.29" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-29 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-29-7919aca8 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-30.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-30.yaml new file mode 100644 index 000000000..c2f89262c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-30.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-30", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.30" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.30" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-30 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-30-3b56bd99 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-31.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-31.yaml new file mode 100644 index 000000000..236c3d5ad --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-31.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-31", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.31" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.31" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-31 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-31-cfff9eb2 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-32.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-32.yaml new file mode 100644 index 000000000..57f5cad9a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/reports/worker-32.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-32", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.32" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.32" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-32 + storage.simplyblock.io/nodeprobe-run: discover-fail-09 + name: sb-nodeprobe-discover-fail-09-worker-32-7d2f16a4 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-10-a-worker-whose-disks-are-partitions/case.md b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-10-a-worker-whose-disks-are-partitions/case.md new file mode 100644 index 000000000..8d215da48 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-10-a-worker-whose-disks-are-partitions/case.md @@ -0,0 +1,7 @@ +# FAIL-10 + +**Mutation.** A worker whose only disks are partitions, `enablePartitionedDevices` unset + +**Expected.** Refused by `whole disk`, pre-filtered, so `Explain` gives the worker line alone + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-10-a-worker-whose-disks-are-partitions/expected-error.txt b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-10-a-worker-whose-disks-are-partitions/expected-error.txt new file mode 100644 index 000000000..27b1e7e6e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-10-a-worker-whose-disks-are-partitions/expected-error.txt @@ -0,0 +1 @@ +no worker has a device this run would use: worker-01: no device of it survived the device rules diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-10-a-worker-whose-disks-are-partitions/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-10-a-worker-whose-disks-are-partitions/expected-refusals.txt new file mode 100644 index 000000000..11b0ad0a9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-10-a-worker-whose-disks-are-partitions/expected-refusals.txt @@ -0,0 +1,3 @@ +worker-01/nvme0n1p1: declined by whole disk because it is a Partition rather than a whole disk +worker-01/nvme0n1p2: declined by whole disk because it is a Partition rather than a whole disk +worker-01: declined by has devices because no device of it survived the device rules diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-10-a-worker-whose-disks-are-partitions/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-10-a-worker-whose-disks-are-partitions/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-10-a-worker-whose-disks-are-partitions/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-10-a-worker-whose-disks-are-partitions/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-10-a-worker-whose-disks-are-partitions/ops.yaml new file mode 100644 index 000000000..275c344b9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-10-a-worker-whose-disks-are-partitions/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-fail-10 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-10-a-worker-whose-disks-are-partitions/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-10-a-worker-whose-disks-are-partitions/reports/worker-01.yaml new file mode 100644 index 000000000..8f8ecba07 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-10-a-worker-whose-disks-are-partitions/reports/worker-01.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1p1", + "path": "/dev/nvme0n1p1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 549755813888, + "kind": "Partition", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme0n1p2", + "path": "/dev/nvme0n1p2", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 549755813888, + "kind": "Partition", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-fail-10 + name: sb-nodeprobe-discover-fail-10-worker-01-1d361fec + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/case.md b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/case.md new file mode 100644 index 000000000..70aefcadb --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/case.md @@ -0,0 +1,7 @@ +# FILT-01 + +**Mutation.** `pcieDenyList` naming the boot slot on a uniform fleet + +**Expected.** That address in no group. The refusal appears in `Explain`, not pre-filtered + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/expected-notes.txt new file mode 100644 index 000000000..5692a767f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-filt-01 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 3 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/expected-refusals.txt new file mode 100644 index 000000000..cddb1a732 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/expected-refusals.txt @@ -0,0 +1,3 @@ +worker-01/nvme0n1: declined by allow and deny lists because it is in the deny list +worker-01/nvme3n1: declined by available because the probe refused it: Partitioned +worker-01/nvme4n1: declined by available because the probe refused it: Partitioned, Mounted diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/expected.yaml new file mode 100644 index 000000000..cf8d5e3ce --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/expected.yaml @@ -0,0 +1,30 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-filt-01 + name: discovered-discover-filt-01 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-filt-01-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5f:00.0 + - 0000:af:00.0 + mgmtInterface: eth0 + name: group-1-nvme-2x1.25T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/ops.yaml new file mode 100644 index 000000000..eb346f0aa --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/ops.yaml @@ -0,0 +1,16 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-filt-01 + namespace: simplyblock +spec: + action: Discover + discover: + deviceFilter: + pcieDenyList: + - 0000:5e:00.0 +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/reports/worker-01.yaml new file mode 100644 index 000000000..f877b30d9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/reports/worker-01.yaml @@ -0,0 +1,201 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "INTEL SSDPF2KX038TZ", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Partitioned", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme4n1", + "path": "/dev/nvme4n1", + "pciAddress": "0000:b1:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Partitioned", + "detail": "as the probe found it" + }, + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-filt-01 + name: sb-nodeprobe-discover-filt-01-worker-01-b753d021 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/case.md b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/case.md new file mode 100644 index 000000000..7498c0b9f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/case.md @@ -0,0 +1,7 @@ +# FILT-02 + +**Mutation.** `pcieAllowList` of 2 addresses against 10 disks + +**Expected.** Only those 2 reach the group + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/expected-notes.txt new file mode 100644 index 000000000..543b17b5a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-filt-02 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 3 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/expected-refusals.txt new file mode 100644 index 000000000..b52555788 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/expected-refusals.txt @@ -0,0 +1,3 @@ +worker-01/nvme2n1: declined by allow and deny lists because it is not in the allow list +worker-01/nvme3n1: declined by available because the probe refused it: Partitioned +worker-01/nvme4n1: declined by available because the probe refused it: Partitioned, Mounted diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/expected.yaml new file mode 100644 index 000000000..442342a93 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/expected.yaml @@ -0,0 +1,30 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-filt-02 + name: discovered-discover-filt-02 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-filt-02-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + mgmtInterface: eth0 + name: group-1-nvme-2x2T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/ops.yaml new file mode 100644 index 000000000..6908dc174 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/ops.yaml @@ -0,0 +1,17 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-filt-02 + namespace: simplyblock +spec: + action: Discover + discover: + deviceFilter: + pcieAllowList: + - 0000:5e:00.0 + - 0000:5f:00.0 +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/reports/worker-01.yaml new file mode 100644 index 000000000..1667c2ac7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/reports/worker-01.yaml @@ -0,0 +1,201 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "INTEL SSDPF2KX038TZ", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Partitioned", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme4n1", + "path": "/dev/nvme4n1", + "pciAddress": "0000:b1:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Partitioned", + "detail": "as the probe found it" + }, + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-filt-02 + name: sb-nodeprobe-discover-filt-02-worker-01-b75b6691 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/case.md b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/case.md new file mode 100644 index 000000000..7edaadd0a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/case.md @@ -0,0 +1,7 @@ +# FILT-03 + +**Mutation.** `pcieModel: MZQL2` against a mixed-model worker + +**Expected.** Only the matching disks. The refusal quotes both strings + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/expected-notes.txt new file mode 100644 index 000000000..7a68f5848 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-filt-03 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 3 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/expected-refusals.txt new file mode 100644 index 000000000..c010f223d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/expected-refusals.txt @@ -0,0 +1,3 @@ +worker-01/nvme1n1: declined by model because its model "INTEL SSDPF2KX038TZ" does not contain "MZQL2" +worker-01/nvme3n1: declined by available because the probe refused it: Partitioned +worker-01/nvme4n1: declined by available because the probe refused it: Partitioned, Mounted diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/expected.yaml new file mode 100644 index 000000000..6dd8866e3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/expected.yaml @@ -0,0 +1,30 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-filt-03 + name: discovered-discover-filt-03 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-filt-03-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:af:00.0 + mgmtInterface: eth0 + name: group-1-nvme-2x1.25T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/ops.yaml new file mode 100644 index 000000000..f6bbfb158 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/ops.yaml @@ -0,0 +1,15 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-filt-03 + namespace: simplyblock +spec: + action: Discover + discover: + deviceFilter: + pcieModel: MZQL2 +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/reports/worker-01.yaml new file mode 100644 index 000000000..12b8d0f06 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/reports/worker-01.yaml @@ -0,0 +1,201 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "INTEL SSDPF2KX038TZ", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Partitioned", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme4n1", + "path": "/dev/nvme4n1", + "pciAddress": "0000:b1:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Partitioned", + "detail": "as the probe found it" + }, + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-filt-03 + name: sb-nodeprobe-discover-filt-03-worker-01-f16e0e1b + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/case.md b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/case.md new file mode 100644 index 000000000..a38da4f38 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/case.md @@ -0,0 +1,7 @@ +# FILT-04 + +**Mutation.** `driveSizeRange: 1T-4T` with a 512 GiB disk present + +**Expected.** The small disk refused, the range quoted in the reason + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/expected-notes.txt new file mode 100644 index 000000000..8a4f770b7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-filt-04 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 3 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/expected-refusals.txt new file mode 100644 index 000000000..0bd43ba52 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/expected-refusals.txt @@ -0,0 +1,3 @@ +worker-01/nvme2n1: declined by size range because it is 512G and the range 1T-4T starts above it +worker-01/nvme3n1: declined by available because the probe refused it: Partitioned +worker-01/nvme4n1: declined by available because the probe refused it: Partitioned, Mounted diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/expected.yaml new file mode 100644 index 000000000..feb200d8a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/expected.yaml @@ -0,0 +1,30 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-filt-04 + name: discovered-discover-filt-04 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-filt-04-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + mgmtInterface: eth0 + name: group-1-nvme-2x2T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/ops.yaml new file mode 100644 index 000000000..ed008c363 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/ops.yaml @@ -0,0 +1,15 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-filt-04 + namespace: simplyblock +spec: + action: Discover + discover: + deviceFilter: + driveSizeRange: 1T-4T +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/reports/worker-01.yaml new file mode 100644 index 000000000..6ae89f93f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/reports/worker-01.yaml @@ -0,0 +1,201 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "INTEL SSDPF2KX038TZ", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Partitioned", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme4n1", + "path": "/dev/nvme4n1", + "pciAddress": "0000:b1:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Partitioned", + "detail": "as the probe found it" + }, + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-filt-04 + name: sb-nodeprobe-discover-filt-04-worker-01-a2fb2ce6 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/case.md b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/case.md new file mode 100644 index 000000000..8d7892dcd --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/case.md @@ -0,0 +1,7 @@ +# FILT-05 + +**Mutation.** `driveSizeRange: 2T` against 2 TiB and 1.92 TB disks + +**Expected.** Only the exact 2 TiB disks: a bare size is both bounds + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/expected-notes.txt new file mode 100644 index 000000000..59d196d92 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-filt-05 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 1 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/expected-refusals.txt new file mode 100644 index 000000000..5379b5b38 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/expected-refusals.txt @@ -0,0 +1 @@ +worker-01/nvme2n1: declined by size range because it is 1.87T and the range 2T starts above it diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/expected.yaml new file mode 100644 index 000000000..fe050fc7f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/expected.yaml @@ -0,0 +1,30 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-filt-05 + name: discovered-discover-filt-05 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-filt-05-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + mgmtInterface: eth0 + name: group-1-nvme-2x2T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/ops.yaml new file mode 100644 index 000000000..ca90209f9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/ops.yaml @@ -0,0 +1,15 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-filt-05 + namespace: simplyblock +spec: + action: Discover + discover: + deviceFilter: + driveSizeRange: 2T +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/reports/worker-01.yaml new file mode 100644 index 000000000..8a152bfa6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/reports/worker-01.yaml @@ -0,0 +1,163 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 2061584302080, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-filt-05 + name: sb-nodeprobe-discover-filt-05-worker-01-e85e825a + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-06-a-size-range-that-counts-backward/case.md b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-06-a-size-range-that-counts-backward/case.md new file mode 100644 index 000000000..e1f1d20a0 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-06-a-size-range-that-counts-backward/case.md @@ -0,0 +1,9 @@ +# FILT-06 + +**Mutation.** `driveSizeRange: 2T-1T` + +**Expected.** The run is refused, naming the field and what a range looks like. Admission refuses it at the request + +**Harness.** `CM` + +**Gap.** G-9. This case records what the generator does today, so that the day it changes the diff is the finding. diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-06-a-size-range-that-counts-backward/expected-error.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-06-a-size-range-that-counts-backward/expected-error.txt new file mode 100644 index 000000000..7a86cf07c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-06-a-size-range-that-counts-backward/expected-error.txt @@ -0,0 +1 @@ +spec.discover.deviceFilter.driveSizeRange is "2T-1T", which cannot be read: the size range "2T-1T" counts backward. A run whose range cannot be read applies no size filter at all, so the draft would name every disk on every worker rather than the ones asked for diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-06-a-size-range-that-counts-backward/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-06-a-size-range-that-counts-backward/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-06-a-size-range-that-counts-backward/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-06-a-size-range-that-counts-backward/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-06-a-size-range-that-counts-backward/ops.yaml new file mode 100644 index 000000000..76aaf43ba --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-06-a-size-range-that-counts-backward/ops.yaml @@ -0,0 +1,15 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-filt-06 + namespace: simplyblock +spec: + action: Discover + discover: + deviceFilter: + driveSizeRange: 2T-1T +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-06-a-size-range-that-counts-backward/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-06-a-size-range-that-counts-backward/reports/worker-01.yaml new file mode 100644 index 000000000..18fc15d7d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-06-a-size-range-that-counts-backward/reports/worker-01.yaml @@ -0,0 +1,201 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "INTEL SSDPF2KX038TZ", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Partitioned", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme4n1", + "path": "/dev/nvme4n1", + "pciAddress": "0000:b1:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Partitioned", + "detail": "as the probe found it" + }, + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-filt-06 + name: sb-nodeprobe-discover-filt-06-worker-01-87895835 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/case.md b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/case.md new file mode 100644 index 000000000..04bb6a0c3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/case.md @@ -0,0 +1,7 @@ +# FILT-07 + +**Mutation.** `blockDenyList: /dev/sda` on a block run + +**Expected.** The root disk in no group + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/expected-notes.txt new file mode 100644 index 000000000..2e1733ecc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-filt-07 in Draft: 1 workers with 5 block devices, in 1 group(s) across 1 node set(s); 1 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/expected-refusals.txt new file mode 100644 index 000000000..41053130d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/expected-refusals.txt @@ -0,0 +1 @@ +worker-01/sda: declined by allow and deny lists because it is in the deny list diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/expected.yaml new file mode 100644 index 000000000..18228c567 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/expected.yaml @@ -0,0 +1,33 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-filt-07 + name: discovered-discover-filt-07 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-filt-07-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + block: + - /dev/sdb + - /dev/sdc + - /dev/vdb + - /dev/vdc + - /dev/vdd + mgmtInterface: eth0 + name: group-1-block-5x3.19T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/ops.yaml new file mode 100644 index 000000000..94e16739d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/ops.yaml @@ -0,0 +1,17 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-filt-07 + namespace: simplyblock +spec: + action: Discover + discover: + deviceFilter: + blockDenyList: + - /dev/sda + enableLogicalBlockDevices: true +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/reports/worker-01.yaml new file mode 100644 index 000000000..abcf07446 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/reports/worker-01.yaml @@ -0,0 +1,187 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "sda", + "path": "/dev/sda", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "sdb", + "path": "/dev/sdb", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "sdc", + "path": "/dev/sdc", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "vdb", + "path": "/dev/vdb", + "sizeBytes": 4398046511104, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "vdc", + "path": "/dev/vdc", + "sizeBytes": 4398046511104, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "vdd", + "path": "/dev/vdd", + "sizeBytes": 4398046511104, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-filt-07 + name: sb-nodeprobe-discover-filt-07-worker-01-1399faf4 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/case.md b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/case.md new file mode 100644 index 000000000..5430e4a1d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/case.md @@ -0,0 +1,7 @@ +# FILT-08 + +**Mutation.** `blockAllowList` of 2 paths against 6 block devices + +**Expected.** Only those 2 + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/expected-notes.txt new file mode 100644 index 000000000..d2254588f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-filt-08 in Draft: 1 workers with 2 block devices, in 1 group(s) across 1 node set(s); 4 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/expected-refusals.txt new file mode 100644 index 000000000..b282d4cb0 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/expected-refusals.txt @@ -0,0 +1,4 @@ +worker-01/sda: declined by allow and deny lists because it is not in the allow list +worker-01/sdb: declined by allow and deny lists because it is not in the allow list +worker-01/sdc: declined by allow and deny lists because it is not in the allow list +worker-01/vdd: declined by allow and deny lists because it is not in the allow list diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/expected.yaml new file mode 100644 index 000000000..14589cbf3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/expected.yaml @@ -0,0 +1,30 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-filt-08 + name: discovered-discover-filt-08 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-filt-08-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + block: + - /dev/vdb + - /dev/vdc + mgmtInterface: eth0 + name: group-1-block-2x4T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/ops.yaml new file mode 100644 index 000000000..2e017f74d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/ops.yaml @@ -0,0 +1,18 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-filt-08 + namespace: simplyblock +spec: + action: Discover + discover: + deviceFilter: + blockAllowList: + - /dev/vdb + - /dev/vdc + enableLogicalBlockDevices: true +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/reports/worker-01.yaml new file mode 100644 index 000000000..42b0d027a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/reports/worker-01.yaml @@ -0,0 +1,187 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "sda", + "path": "/dev/sda", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "sdb", + "path": "/dev/sdb", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "sdc", + "path": "/dev/sdc", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "vdb", + "path": "/dev/vdb", + "sizeBytes": 4398046511104, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "vdc", + "path": "/dev/vdc", + "sizeBytes": 4398046511104, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "vdd", + "path": "/dev/vdd", + "sizeBytes": 4398046511104, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-filt-08 + name: sb-nodeprobe-discover-filt-08-worker-01-7f4e984b + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/case.md b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/case.md new file mode 100644 index 000000000..30f656b52 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/case.md @@ -0,0 +1,7 @@ +# FILT-09 + +**Mutation.** `enablePartitionedDevices` with a GPT-only refusal + +**Expected.** Admitted. A disk also mounted stays refused + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/expected-notes.txt new file mode 100644 index 000000000..7fdff9f56 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-filt-09 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 1 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/expected-refusals.txt new file mode 100644 index 000000000..e2d1dfd43 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/expected-refusals.txt @@ -0,0 +1 @@ +worker-01/nvme4n1: declined by available because the probe refused it: Partitioned, Mounted diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/expected.yaml new file mode 100644 index 000000000..341109a81 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-filt-09 + name: discovered-discover-filt-09 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-filt-09-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x1.62T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/ops.yaml new file mode 100644 index 000000000..d80211c3f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/ops.yaml @@ -0,0 +1,15 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-filt-09 + namespace: simplyblock +spec: + action: Discover + discover: + deviceFilter: + enablePartitionedDevices: true +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/reports/worker-01.yaml new file mode 100644 index 000000000..ec9373495 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/reports/worker-01.yaml @@ -0,0 +1,201 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "INTEL SSDPF2KX038TZ", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Partitioned", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme4n1", + "path": "/dev/nvme4n1", + "pciAddress": "0000:b1:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Partitioned", + "detail": "as the probe found it" + }, + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-filt-09 + name: sb-nodeprobe-discover-filt-09-worker-01-0e3bbb8e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-10-a-filter-that-matches-nothing/case.md b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-10-a-filter-that-matches-nothing/case.md new file mode 100644 index 000000000..89b39b76b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-10-a-filter-that-matches-nothing/case.md @@ -0,0 +1,7 @@ +# FILT-10 + +**Mutation.** A filter that matches nothing on any worker + +**Expected.** No node sets. The run fails naming the filter rule per worker + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-10-a-filter-that-matches-nothing/expected-error.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-10-a-filter-that-matches-nothing/expected-error.txt new file mode 100644 index 000000000..beb26ec8e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-10-a-filter-that-matches-nothing/expected-error.txt @@ -0,0 +1 @@ +no worker has a device this run would use: worker-01: no device of it survived the device rules (3 devices declined by allow and deny lists: it is not in the allow list, 2 devices declined by available: 1 the probe refused it: Partitioned, 1 the probe refused it: Partitioned, Mounted) diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-10-a-filter-that-matches-nothing/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-10-a-filter-that-matches-nothing/expected-refusals.txt new file mode 100644 index 000000000..59a7a170c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-10-a-filter-that-matches-nothing/expected-refusals.txt @@ -0,0 +1,6 @@ +worker-01/nvme0n1: declined by allow and deny lists because it is not in the allow list +worker-01/nvme1n1: declined by allow and deny lists because it is not in the allow list +worker-01/nvme2n1: declined by allow and deny lists because it is not in the allow list +worker-01/nvme3n1: declined by available because the probe refused it: Partitioned +worker-01/nvme4n1: declined by available because the probe refused it: Partitioned, Mounted +worker-01: declined by has devices because no device of it survived the device rules diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-10-a-filter-that-matches-nothing/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-10-a-filter-that-matches-nothing/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-10-a-filter-that-matches-nothing/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-10-a-filter-that-matches-nothing/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-10-a-filter-that-matches-nothing/ops.yaml new file mode 100644 index 000000000..6c9317fe6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-10-a-filter-that-matches-nothing/ops.yaml @@ -0,0 +1,16 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-filt-10 + namespace: simplyblock +spec: + action: Discover + discover: + deviceFilter: + pcieAllowList: + - 0000:ff:00.0 +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-10-a-filter-that-matches-nothing/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-10-a-filter-that-matches-nothing/reports/worker-01.yaml new file mode 100644 index 000000000..9ef4a9887 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-10-a-filter-that-matches-nothing/reports/worker-01.yaml @@ -0,0 +1,201 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "INTEL SSDPF2KX038TZ", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Partitioned", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme4n1", + "path": "/dev/nvme4n1", + "pciAddress": "0000:b1:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Partitioned", + "detail": "as the probe found it" + }, + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-filt-10 + name: sb-nodeprobe-discover-filt-10-worker-01-931c7280 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/case.md b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/case.md new file mode 100644 index 000000000..18d81b20e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/case.md @@ -0,0 +1,7 @@ +# FILT-11 + +**Mutation.** An allow list written in uppercase against lowercase addresses + +**Expected.** Matched: the comparison folds case + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/expected-notes.txt new file mode 100644 index 000000000..6aee0237c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-filt-11 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 3 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/expected-refusals.txt new file mode 100644 index 000000000..b52555788 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/expected-refusals.txt @@ -0,0 +1,3 @@ +worker-01/nvme2n1: declined by allow and deny lists because it is not in the allow list +worker-01/nvme3n1: declined by available because the probe refused it: Partitioned +worker-01/nvme4n1: declined by available because the probe refused it: Partitioned, Mounted diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/expected.yaml new file mode 100644 index 000000000..a5d9e245e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/expected.yaml @@ -0,0 +1,30 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-filt-11 + name: discovered-discover-filt-11 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-filt-11-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + mgmtInterface: eth0 + name: group-1-nvme-2x2T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/ops.yaml new file mode 100644 index 000000000..ccc9595af --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/ops.yaml @@ -0,0 +1,17 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-filt-11 + namespace: simplyblock +spec: + action: Discover + discover: + deviceFilter: + pcieAllowList: + - 0000:5E:00.0 + - 0000:5F:00.0 +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/reports/worker-01.yaml new file mode 100644 index 000000000..634a0eb88 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/reports/worker-01.yaml @@ -0,0 +1,201 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "INTEL SSDPF2KX038TZ", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Partitioned", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme4n1", + "path": "/dev/nvme4n1", + "pciAddress": "0000:b1:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Partitioned", + "detail": "as the probe found it" + }, + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-filt-11 + name: sb-nodeprobe-discover-filt-11-worker-01-975dc296 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/case.md b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/case.md new file mode 100644 index 000000000..a941b2e6a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/case.md @@ -0,0 +1,7 @@ +# FILT-12 + +**Mutation.** One address in both the allow and the deny list + +**Expected.** Refused: deny is evaluated first + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/expected-notes.txt new file mode 100644 index 000000000..0354faecc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-filt-12 in Draft: 1 workers with 1 nvme devices, in 1 group(s) across 1 node set(s); 4 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/expected-refusals.txt new file mode 100644 index 000000000..ba7cbb326 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/expected-refusals.txt @@ -0,0 +1,4 @@ +worker-01/nvme0n1: declined by allow and deny lists because it is in the deny list +worker-01/nvme2n1: declined by allow and deny lists because it is not in the allow list +worker-01/nvme3n1: declined by available because the probe refused it: Partitioned +worker-01/nvme4n1: declined by available because the probe refused it: Partitioned, Mounted diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/expected.yaml new file mode 100644 index 000000000..f77d3be04 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/expected.yaml @@ -0,0 +1,29 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-filt-12 + name: discovered-discover-filt-12 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-filt-12-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5f:00.0 + mgmtInterface: eth0 + name: group-1-nvme-1x2T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/ops.yaml new file mode 100644 index 000000000..a840889ea --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/ops.yaml @@ -0,0 +1,19 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-filt-12 + namespace: simplyblock +spec: + action: Discover + discover: + deviceFilter: + pcieAllowList: + - 0000:5e:00.0 + - 0000:5f:00.0 + pcieDenyList: + - 0000:5e:00.0 +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/reports/worker-01.yaml new file mode 100644 index 000000000..c6a87f9dd --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/reports/worker-01.yaml @@ -0,0 +1,201 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "INTEL SSDPF2KX038TZ", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Partitioned", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme4n1", + "path": "/dev/nvme4n1", + "pciAddress": "0000:b1:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Partitioned", + "detail": "as the probe found it" + }, + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-filt-12 + name: sb-nodeprobe-discover-filt-12-worker-01-a4cd7012 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/case.md b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/case.md new file mode 100644 index 000000000..00a06b073 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/case.md @@ -0,0 +1,7 @@ +# FILT-13 + +**Mutation.** `driveSizeRange` on a block run + +**Expected.** It narrows the block class, which is the class being scanned + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/expected-notes.txt new file mode 100644 index 000000000..8cca2d5f0 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-filt-13 in Draft: 1 workers with 3 block devices, in 1 group(s) across 1 node set(s); 3 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/expected-refusals.txt new file mode 100644 index 000000000..ee4b52908 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/expected-refusals.txt @@ -0,0 +1,3 @@ +worker-01/sda: declined by size range because it is 512G and the range 3T-8T starts above it +worker-01/sdb: declined by size range because it is 2T and the range 3T-8T starts above it +worker-01/sdc: declined by size range because it is 2T and the range 3T-8T starts above it diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/expected.yaml new file mode 100644 index 000000000..1dbee7037 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/expected.yaml @@ -0,0 +1,31 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-filt-13 + name: discovered-discover-filt-13 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-filt-13-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + block: + - /dev/vdb + - /dev/vdc + - /dev/vdd + mgmtInterface: eth0 + name: group-1-block-3x4T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/ops.yaml new file mode 100644 index 000000000..6a7732e7d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/ops.yaml @@ -0,0 +1,16 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-filt-13 + namespace: simplyblock +spec: + action: Discover + discover: + deviceFilter: + driveSizeRange: 3T-8T + enableLogicalBlockDevices: true +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/reports/worker-01.yaml new file mode 100644 index 000000000..f0165a493 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/reports/worker-01.yaml @@ -0,0 +1,187 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "sda", + "path": "/dev/sda", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "sdb", + "path": "/dev/sdb", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "sdc", + "path": "/dev/sdc", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "vdb", + "path": "/dev/vdb", + "sizeBytes": 4398046511104, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "vdc", + "path": "/dev/vdc", + "sizeBytes": 4398046511104, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "vdd", + "path": "/dev/vdd", + "sizeBytes": 4398046511104, + "kind": "Disk", + "transport": "Virtio", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-filt-13 + name: sb-nodeprobe-discover-filt-13-worker-01-f8b6b9c9 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/case.md b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/case.md new file mode 100644 index 000000000..ab80d2618 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/case.md @@ -0,0 +1,7 @@ +# FILT-14 + +**Mutation.** No `deviceFilter` at all + +**Expected.** Every free whole NVMe disk reaches the draft + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/expected-notes.txt new file mode 100644 index 000000000..a73059461 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-filt-14 in Draft: 1 workers with 3 nvme devices, in 1 group(s) across 1 node set(s); 2 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/expected-refusals.txt new file mode 100644 index 000000000..8c74e3c6e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/expected-refusals.txt @@ -0,0 +1,2 @@ +worker-01/nvme3n1: declined by available because the probe refused it: Partitioned +worker-01/nvme4n1: declined by available because the probe refused it: Partitioned, Mounted diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/expected.yaml new file mode 100644 index 000000000..cd8917c22 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/expected.yaml @@ -0,0 +1,31 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-filt-14 + name: discovered-discover-filt-14 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-filt-14-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + mgmtInterface: eth0 + name: group-1-nvme-3x1.5T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/ops.yaml new file mode 100644 index 000000000..5b407c29c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-filt-14 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/reports/worker-01.yaml new file mode 100644 index 000000000..3d0195b7a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/reports/worker-01.yaml @@ -0,0 +1,201 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "INTEL SSDPF2KX038TZ", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 549755813888, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Partitioned", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme4n1", + "path": "/dev/nvme4n1", + "pciAddress": "0000:b1:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Partitioned", + "detail": "as the probe found it" + }, + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-filt-14 + name: sb-nodeprobe-discover-filt-14-worker-01-8bcef7b2 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/case.md b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/case.md new file mode 100644 index 000000000..cb6eaae9d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/case.md @@ -0,0 +1,7 @@ +# FLEET-01 + +**Mutation.** 1 worker, 4 disks + +**Expected.** 1 group of 1 + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/expected-notes.txt new file mode 100644 index 000000000..b4c9266da --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-fleet-01 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/expected.yaml new file mode 100644 index 000000000..ceff29a87 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-fleet-01 + name: discovered-discover-fleet-01 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-fleet-01-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/ops.yaml new file mode 100644 index 000000000..88e38890d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-fleet-01 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/reports/worker-01.yaml new file mode 100644 index 000000000..fb2927283 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-fleet-01 + name: sb-nodeprobe-discover-fleet-01-worker-01-c8c3c7b0 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/case.md b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/case.md new file mode 100644 index 000000000..9efadfe81 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/case.md @@ -0,0 +1,7 @@ +# FLEET-02 + +**Mutation.** 3 uniform workers + +**Expected.** 1 group of 3. The summary reads 3 workers, 12 devices, 1 group, 1 node set + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/expected-notes.txt new file mode 100644 index 000000000..0d641fc59 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-fleet-02 in Draft: 3 workers with 12 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/expected.yaml new file mode 100644 index 000000000..55db0db40 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/expected.yaml @@ -0,0 +1,34 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-fleet-02 + name: discovered-discover-fleet-02 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-fleet-02-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - worker-02 + - worker-03 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/nodes.yaml new file mode 100644 index 000000000..98de08ee9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/nodes.yaml @@ -0,0 +1,89 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/ops.yaml new file mode 100644 index 000000000..db0def1b9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/ops.yaml @@ -0,0 +1,14 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-fleet-02 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/reports/worker-01.yaml new file mode 100644 index 000000000..70b1a4209 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-fleet-02 + name: sb-nodeprobe-discover-fleet-02-worker-01-dcb5e737 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/reports/worker-02.yaml new file mode 100644 index 000000000..13d22e695 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-fleet-02 + name: sb-nodeprobe-discover-fleet-02-worker-02-cb207a98 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/reports/worker-03.yaml new file mode 100644 index 000000000..c3dc5b2d7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-fleet-02 + name: sb-nodeprobe-discover-fleet-02-worker-03-fe3614df + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/case.md b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/case.md new file mode 100644 index 000000000..9f404effc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/case.md @@ -0,0 +1,7 @@ +# FLEET-03 + +**Mutation.** 32 uniform workers + +**Expected.** 1 group of 32, under the 200-worker ceiling + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/expected-notes.txt new file mode 100644 index 000000000..549147508 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-fleet-03 in Draft: 32 workers with 128 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/expected.yaml new file mode 100644 index 000000000..f4f6d3e97 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/expected.yaml @@ -0,0 +1,63 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: discovered-discover-fleet-03 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-fleet-03-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - worker-02 + - worker-03 + - worker-04 + - worker-05 + - worker-06 + - worker-07 + - worker-08 + - worker-09 + - worker-10 + - worker-11 + - worker-12 + - worker-13 + - worker-14 + - worker-15 + - worker-16 + - worker-17 + - worker-18 + - worker-19 + - worker-20 + - worker-21 + - worker-22 + - worker-23 + - worker-24 + - worker-25 + - worker-26 + - worker-27 + - worker-28 + - worker-29 + - worker-30 + - worker-31 + - worker-32 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/nodes.yaml new file mode 100644 index 000000000..c25e3f497 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/nodes.yaml @@ -0,0 +1,959 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-04 +spec: {} +status: + addresses: + - address: 10.10.10.4 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-05 +spec: {} +status: + addresses: + - address: 10.10.10.5 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-06 +spec: {} +status: + addresses: + - address: 10.10.10.6 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-07 +spec: {} +status: + addresses: + - address: 10.10.10.7 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-08 +spec: {} +status: + addresses: + - address: 10.10.10.8 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-09 +spec: {} +status: + addresses: + - address: 10.10.10.9 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-10 +spec: {} +status: + addresses: + - address: 10.10.10.10 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-11 +spec: {} +status: + addresses: + - address: 10.10.10.11 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-12 +spec: {} +status: + addresses: + - address: 10.10.10.12 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-13 +spec: {} +status: + addresses: + - address: 10.10.10.13 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-14 +spec: {} +status: + addresses: + - address: 10.10.10.14 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-15 +spec: {} +status: + addresses: + - address: 10.10.10.15 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-16 +spec: {} +status: + addresses: + - address: 10.10.10.16 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-17 +spec: {} +status: + addresses: + - address: 10.10.10.17 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-18 +spec: {} +status: + addresses: + - address: 10.10.10.18 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-19 +spec: {} +status: + addresses: + - address: 10.10.10.19 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-20 +spec: {} +status: + addresses: + - address: 10.10.10.20 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-21 +spec: {} +status: + addresses: + - address: 10.10.10.21 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-22 +spec: {} +status: + addresses: + - address: 10.10.10.22 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-23 +spec: {} +status: + addresses: + - address: 10.10.10.23 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-24 +spec: {} +status: + addresses: + - address: 10.10.10.24 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-25 +spec: {} +status: + addresses: + - address: 10.10.10.25 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-26 +spec: {} +status: + addresses: + - address: 10.10.10.26 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-27 +spec: {} +status: + addresses: + - address: 10.10.10.27 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-28 +spec: {} +status: + addresses: + - address: 10.10.10.28 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-29 +spec: {} +status: + addresses: + - address: 10.10.10.29 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-30 +spec: {} +status: + addresses: + - address: 10.10.10.30 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-31 +spec: {} +status: + addresses: + - address: 10.10.10.31 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-32 +spec: {} +status: + addresses: + - address: 10.10.10.32 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/ops.yaml new file mode 100644 index 000000000..d6069c83a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/ops.yaml @@ -0,0 +1,43 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-fleet-03 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 + - worker-04 + - worker-05 + - worker-06 + - worker-07 + - worker-08 + - worker-09 + - worker-10 + - worker-11 + - worker-12 + - worker-13 + - worker-14 + - worker-15 + - worker-16 + - worker-17 + - worker-18 + - worker-19 + - worker-20 + - worker-21 + - worker-22 + - worker-23 + - worker-24 + - worker-25 + - worker-26 + - worker-27 + - worker-28 + - worker-29 + - worker-30 + - worker-31 + - worker-32 diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-01.yaml new file mode 100644 index 000000000..1a0205d92 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-01-bdb6f719 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-02.yaml new file mode 100644 index 000000000..2695afe75 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-02-bc0c962f + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-03.yaml new file mode 100644 index 000000000..a06b8b42a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-03-c576d1bf + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-04.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-04.yaml new file mode 100644 index 000000000..148f76832 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-04.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-04", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.4" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.4" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-04 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-04-c7d6255e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-05.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-05.yaml new file mode 100644 index 000000000..3ef92d381 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-05.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-05", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.5" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.5" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-05 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-05-8af85b52 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-06.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-06.yaml new file mode 100644 index 000000000..b2c8938cc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-06.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-06", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.6" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.6" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-06 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-06-e0fd66a9 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-07.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-07.yaml new file mode 100644 index 000000000..c180bc54e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-07.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-07", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.7" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.7" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-07 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-07-7349eedb + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-08.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-08.yaml new file mode 100644 index 000000000..ef7d73300 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-08.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-08", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.8" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.8" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-08 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-08-264dafce + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-09.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-09.yaml new file mode 100644 index 000000000..9d54522d5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-09.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-09", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.9" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.9" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-09 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-09-c093dade + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-10.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-10.yaml new file mode 100644 index 000000000..207657017 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-10.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-10", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.10" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.10" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-10 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-10-738fbe03 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-11.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-11.yaml new file mode 100644 index 000000000..fc1071d2e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-11.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-11", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.11" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.11" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-11 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-11-7153980b + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-12.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-12.yaml new file mode 100644 index 000000000..06346a1ef --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-12.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-12", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.12" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.12" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-12 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-12-e6c5a8d9 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-13.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-13.yaml new file mode 100644 index 000000000..ce7bd93cb --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-13.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-13", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.13" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.13" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-13 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-13-371685ad + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-14.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-14.yaml new file mode 100644 index 000000000..768d35068 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-14.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-14", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.14" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.14" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-14 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-14-c03bd461 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-15.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-15.yaml new file mode 100644 index 000000000..26e862d2c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-15.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-15", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.15" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.15" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-15 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-15-300d3876 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-16.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-16.yaml new file mode 100644 index 000000000..8ad080113 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-16.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-16", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.16" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.16" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-16 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-16-ad4954ee + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-17.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-17.yaml new file mode 100644 index 000000000..8521c77e1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-17.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-17", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.17" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.17" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-17 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-17-2f221ed0 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-18.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-18.yaml new file mode 100644 index 000000000..03ff788a7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-18.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-18", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.18" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.18" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-18 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-18-e86a0a9b + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-19.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-19.yaml new file mode 100644 index 000000000..e9b7351ec --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-19.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-19", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.19" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.19" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-19 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-19-1ccf010f + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-20.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-20.yaml new file mode 100644 index 000000000..0ff860919 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-20.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-20", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.20" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.20" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-20 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-20-79c76254 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-21.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-21.yaml new file mode 100644 index 000000000..5fa42cb9b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-21.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-21", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.21" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.21" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-21 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-21-adc493aa + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-22.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-22.yaml new file mode 100644 index 000000000..4b4b8ad97 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-22.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-22", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.22" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.22" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-22 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-22-409b6dfd + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-23.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-23.yaml new file mode 100644 index 000000000..755f38a64 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-23.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-23", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.23" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.23" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-23 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-23-683fda49 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-24.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-24.yaml new file mode 100644 index 000000000..d829e84e1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-24.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-24", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.24" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.24" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-24 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-24-486d01c5 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-25.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-25.yaml new file mode 100644 index 000000000..f650ff6bf --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-25.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-25", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.25" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.25" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-25 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-25-ce6bf93f + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-26.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-26.yaml new file mode 100644 index 000000000..d0f240dd5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-26.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-26", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.26" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.26" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-26 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-26-4d1d908a + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-27.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-27.yaml new file mode 100644 index 000000000..1c7f04e07 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-27.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-27", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.27" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.27" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-27 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-27-a87307d0 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-28.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-28.yaml new file mode 100644 index 000000000..13809f4eb --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-28.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-28", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.28" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.28" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-28 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-28-d32d828b + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-29.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-29.yaml new file mode 100644 index 000000000..0ff33d64f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-29.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-29", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.29" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.29" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-29 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-29-30825a9d + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-30.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-30.yaml new file mode 100644 index 000000000..6ff346ec5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-30.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-30", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.30" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.30" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-30 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-30-ac58633c + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-31.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-31.yaml new file mode 100644 index 000000000..045e728bc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-31.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-31", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.31" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.31" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-31 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-31-af969486 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-32.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-32.yaml new file mode 100644 index 000000000..2be44f20e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/reports/worker-32.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-32", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.32" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.32" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-32 + storage.simplyblock.io/nodeprobe-run: discover-fleet-03 + name: sb-nodeprobe-discover-fleet-03-worker-32-564b71fc + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-04-no-reports-at-all/case.md b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-04-no-reports-at-all/case.md new file mode 100644 index 000000000..5ba139418 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-04-no-reports-at-all/case.md @@ -0,0 +1,7 @@ +# FLEET-04 + +**Mutation.** No reports at all + +**Expected.** The run fails: the probe reports are gone + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-04-no-reports-at-all/expected-error.txt b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-04-no-reports-at-all/expected-error.txt new file mode 100644 index 000000000..0ff97c845 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-04-no-reports-at-all/expected-error.txt @@ -0,0 +1 @@ +the probe reports are gone, so there is nothing to write diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-04-no-reports-at-all/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-04-no-reports-at-all/nodes.yaml new file mode 100644 index 000000000..98de08ee9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-04-no-reports-at-all/nodes.yaml @@ -0,0 +1,89 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-04-no-reports-at-all/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-04-no-reports-at-all/ops.yaml new file mode 100644 index 000000000..f7a6fa0e2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-04-no-reports-at-all/ops.yaml @@ -0,0 +1,14 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-fleet-04 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/case.md b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/case.md new file mode 100644 index 000000000..6d0ed6c4e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/case.md @@ -0,0 +1,7 @@ +# FLEET-05 + +**Mutation.** 32 workers, 3 with every disk mounted + +**Expected.** 29 workers drafted. 3 worker refusals with their device reasons folded in + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/expected-notes.txt new file mode 100644 index 000000000..6572794ab --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-fleet-05 in Draft: 30 workers with 120 nvme devices, in 1 group(s) across 1 node set(s); 6 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/expected-refusals.txt new file mode 100644 index 000000000..c35adeab8 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/expected-refusals.txt @@ -0,0 +1,6 @@ +worker-11/nvme0n1: declined by available because the probe refused it: Mounted +worker-11/nvme1n1: declined by available because the probe refused it: Mounted +worker-11: declined by has devices because no device of it survived the device rules +worker-22/nvme0n1: declined by available because the probe refused it: Mounted +worker-22/nvme1n1: declined by available because the probe refused it: Mounted +worker-22: declined by has devices because no device of it survived the device rules diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/expected.yaml new file mode 100644 index 000000000..5ddf35823 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/expected.yaml @@ -0,0 +1,61 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: discovered-discover-fleet-05 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-fleet-05-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - worker-02 + - worker-03 + - worker-04 + - worker-05 + - worker-06 + - worker-07 + - worker-08 + - worker-09 + - worker-10 + - worker-12 + - worker-13 + - worker-14 + - worker-15 + - worker-16 + - worker-17 + - worker-18 + - worker-19 + - worker-20 + - worker-21 + - worker-23 + - worker-24 + - worker-25 + - worker-26 + - worker-27 + - worker-28 + - worker-29 + - worker-30 + - worker-31 + - worker-32 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/nodes.yaml new file mode 100644 index 000000000..c25e3f497 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/nodes.yaml @@ -0,0 +1,959 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-04 +spec: {} +status: + addresses: + - address: 10.10.10.4 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-05 +spec: {} +status: + addresses: + - address: 10.10.10.5 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-06 +spec: {} +status: + addresses: + - address: 10.10.10.6 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-07 +spec: {} +status: + addresses: + - address: 10.10.10.7 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-08 +spec: {} +status: + addresses: + - address: 10.10.10.8 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-09 +spec: {} +status: + addresses: + - address: 10.10.10.9 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-10 +spec: {} +status: + addresses: + - address: 10.10.10.10 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-11 +spec: {} +status: + addresses: + - address: 10.10.10.11 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-12 +spec: {} +status: + addresses: + - address: 10.10.10.12 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-13 +spec: {} +status: + addresses: + - address: 10.10.10.13 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-14 +spec: {} +status: + addresses: + - address: 10.10.10.14 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-15 +spec: {} +status: + addresses: + - address: 10.10.10.15 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-16 +spec: {} +status: + addresses: + - address: 10.10.10.16 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-17 +spec: {} +status: + addresses: + - address: 10.10.10.17 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-18 +spec: {} +status: + addresses: + - address: 10.10.10.18 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-19 +spec: {} +status: + addresses: + - address: 10.10.10.19 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-20 +spec: {} +status: + addresses: + - address: 10.10.10.20 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-21 +spec: {} +status: + addresses: + - address: 10.10.10.21 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-22 +spec: {} +status: + addresses: + - address: 10.10.10.22 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-23 +spec: {} +status: + addresses: + - address: 10.10.10.23 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-24 +spec: {} +status: + addresses: + - address: 10.10.10.24 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-25 +spec: {} +status: + addresses: + - address: 10.10.10.25 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-26 +spec: {} +status: + addresses: + - address: 10.10.10.26 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-27 +spec: {} +status: + addresses: + - address: 10.10.10.27 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-28 +spec: {} +status: + addresses: + - address: 10.10.10.28 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-29 +spec: {} +status: + addresses: + - address: 10.10.10.29 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-30 +spec: {} +status: + addresses: + - address: 10.10.10.30 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-31 +spec: {} +status: + addresses: + - address: 10.10.10.31 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-32 +spec: {} +status: + addresses: + - address: 10.10.10.32 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/ops.yaml new file mode 100644 index 000000000..5045f7262 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/ops.yaml @@ -0,0 +1,43 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-fleet-05 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 + - worker-04 + - worker-05 + - worker-06 + - worker-07 + - worker-08 + - worker-09 + - worker-10 + - worker-11 + - worker-12 + - worker-13 + - worker-14 + - worker-15 + - worker-16 + - worker-17 + - worker-18 + - worker-19 + - worker-20 + - worker-21 + - worker-22 + - worker-23 + - worker-24 + - worker-25 + - worker-26 + - worker-27 + - worker-28 + - worker-29 + - worker-30 + - worker-31 + - worker-32 diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-01.yaml new file mode 100644 index 000000000..ad18c05d0 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-01-ce747d90 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-02.yaml new file mode 100644 index 000000000..0e8bd35e6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-02-d002bb4e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-03.yaml new file mode 100644 index 000000000..08b20627d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-03-99315af9 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-04.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-04.yaml new file mode 100644 index 000000000..ebacbaa24 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-04.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-04", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.4" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.4" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-04 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-04-4532477b + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-05.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-05.yaml new file mode 100644 index 000000000..0044e6d4d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-05.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-05", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.5" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.5" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-05 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-05-39dc449a + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-06.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-06.yaml new file mode 100644 index 000000000..cffbf4177 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-06.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-06", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.6" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.6" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-06 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-06-f0832b3d + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-07.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-07.yaml new file mode 100644 index 000000000..d97eb9e31 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-07.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-07", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.7" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.7" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-07 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-07-c0db9644 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-08.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-08.yaml new file mode 100644 index 000000000..2cca64559 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-08.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-08", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.8" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.8" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-08 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-08-8a4e2fbc + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-09.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-09.yaml new file mode 100644 index 000000000..310432f6a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-09.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-09", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.9" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.9" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-09 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-09-0a07822d + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-10.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-10.yaml new file mode 100644 index 000000000..eaffd8142 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-10.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-10", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.10" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.10" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-10 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-10-49f05c7a + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-11.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-11.yaml new file mode 100644 index 000000000..2c4ddb181 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-11.yaml @@ -0,0 +1,161 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-11", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.11" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.11" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-11 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-11-193e8fce + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-12.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-12.yaml new file mode 100644 index 000000000..deb254ad9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-12.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-12", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.12" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.12" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-12 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-12-8a1abfc6 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-13.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-13.yaml new file mode 100644 index 000000000..2268d8a1b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-13.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-13", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.13" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.13" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-13 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-13-4ee86784 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-14.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-14.yaml new file mode 100644 index 000000000..671db3e01 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-14.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-14", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.14" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.14" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-14 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-14-67f5c7a1 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-15.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-15.yaml new file mode 100644 index 000000000..b80099388 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-15.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-15", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.15" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.15" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-15 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-15-90315e57 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-16.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-16.yaml new file mode 100644 index 000000000..8e0b7e6eb --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-16.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-16", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.16" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.16" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-16 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-16-375fc11f + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-17.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-17.yaml new file mode 100644 index 000000000..385ccdfd6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-17.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-17", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.17" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.17" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-17 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-17-330af242 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-18.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-18.yaml new file mode 100644 index 000000000..34a5a8d36 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-18.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-18", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.18" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.18" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-18 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-18-1557f640 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-19.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-19.yaml new file mode 100644 index 000000000..c3abd5e23 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-19.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-19", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.19" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.19" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-19 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-19-1466d300 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-20.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-20.yaml new file mode 100644 index 000000000..1748957b8 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-20.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-20", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.20" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.20" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-20 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-20-208754ea + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-21.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-21.yaml new file mode 100644 index 000000000..5116e1478 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-21.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-21", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.21" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.21" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-21 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-21-8a286efc + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-22.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-22.yaml new file mode 100644 index 000000000..401f632b1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-22.yaml @@ -0,0 +1,161 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-22", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.22" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.22" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-22 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-22-5d20f55f + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-23.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-23.yaml new file mode 100644 index 000000000..423dcfbc1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-23.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-23", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.23" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.23" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-23 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-23-9fa3c516 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-24.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-24.yaml new file mode 100644 index 000000000..30856dff5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-24.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-24", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.24" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.24" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-24 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-24-f857439d + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-25.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-25.yaml new file mode 100644 index 000000000..dc8d47fa2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-25.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-25", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.25" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.25" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-25 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-25-6a057cef + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-26.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-26.yaml new file mode 100644 index 000000000..c45192308 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-26.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-26", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.26" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.26" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-26 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-26-f7e8a534 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-27.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-27.yaml new file mode 100644 index 000000000..6fdd4d6aa --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-27.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-27", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.27" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.27" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-27 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-27-993c2b53 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-28.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-28.yaml new file mode 100644 index 000000000..27bd1b9c4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-28.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-28", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.28" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.28" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-28 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-28-05d8ee3b + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-29.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-29.yaml new file mode 100644 index 000000000..b5bdbd649 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-29.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-29", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.29" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.29" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-29 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-29-f4e3c309 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-30.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-30.yaml new file mode 100644 index 000000000..bcc5eba37 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-30.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-30", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.30" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.30" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-30 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-30-cc0a78c9 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-31.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-31.yaml new file mode 100644 index 000000000..eb4d488fd --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-31.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-31", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.31" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.31" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-31 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-31-8b2b0abf + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-32.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-32.yaml new file mode 100644 index 000000000..575ff5bd4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/reports/worker-32.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-32", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.32" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.32" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-32 + storage.simplyblock.io/nodeprobe-run: discover-fleet-05 + name: sb-nodeprobe-discover-fleet-05-worker-32-641fe473 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/case.md b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/case.md new file mode 100644 index 000000000..ebcef901c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/case.md @@ -0,0 +1,7 @@ +# FLEET-06 + +**Mutation.** The same 32 reports, ConfigMaps listed in a different order + +**Expected.** A byte-identical document + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/expected-notes.txt new file mode 100644 index 000000000..aef485fda --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-fleet-06 in Draft: 32 workers with 128 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/expected.yaml new file mode 100644 index 000000000..203d4afce --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/expected.yaml @@ -0,0 +1,63 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: discovered-discover-fleet-06 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-fleet-06-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - worker-02 + - worker-03 + - worker-04 + - worker-05 + - worker-06 + - worker-07 + - worker-08 + - worker-09 + - worker-10 + - worker-11 + - worker-12 + - worker-13 + - worker-14 + - worker-15 + - worker-16 + - worker-17 + - worker-18 + - worker-19 + - worker-20 + - worker-21 + - worker-22 + - worker-23 + - worker-24 + - worker-25 + - worker-26 + - worker-27 + - worker-28 + - worker-29 + - worker-30 + - worker-31 + - worker-32 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/nodes.yaml new file mode 100644 index 000000000..423ecc4c0 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/nodes.yaml @@ -0,0 +1,959 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-32 +spec: {} +status: + addresses: + - address: 10.10.10.32 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-31 +spec: {} +status: + addresses: + - address: 10.10.10.31 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-30 +spec: {} +status: + addresses: + - address: 10.10.10.30 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-29 +spec: {} +status: + addresses: + - address: 10.10.10.29 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-28 +spec: {} +status: + addresses: + - address: 10.10.10.28 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-27 +spec: {} +status: + addresses: + - address: 10.10.10.27 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-26 +spec: {} +status: + addresses: + - address: 10.10.10.26 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-25 +spec: {} +status: + addresses: + - address: 10.10.10.25 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-24 +spec: {} +status: + addresses: + - address: 10.10.10.24 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-23 +spec: {} +status: + addresses: + - address: 10.10.10.23 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-22 +spec: {} +status: + addresses: + - address: 10.10.10.22 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-21 +spec: {} +status: + addresses: + - address: 10.10.10.21 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-20 +spec: {} +status: + addresses: + - address: 10.10.10.20 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-19 +spec: {} +status: + addresses: + - address: 10.10.10.19 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-18 +spec: {} +status: + addresses: + - address: 10.10.10.18 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-17 +spec: {} +status: + addresses: + - address: 10.10.10.17 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-16 +spec: {} +status: + addresses: + - address: 10.10.10.16 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-15 +spec: {} +status: + addresses: + - address: 10.10.10.15 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-14 +spec: {} +status: + addresses: + - address: 10.10.10.14 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-13 +spec: {} +status: + addresses: + - address: 10.10.10.13 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-12 +spec: {} +status: + addresses: + - address: 10.10.10.12 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-11 +spec: {} +status: + addresses: + - address: 10.10.10.11 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-10 +spec: {} +status: + addresses: + - address: 10.10.10.10 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-09 +spec: {} +status: + addresses: + - address: 10.10.10.9 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-08 +spec: {} +status: + addresses: + - address: 10.10.10.8 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-07 +spec: {} +status: + addresses: + - address: 10.10.10.7 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-06 +spec: {} +status: + addresses: + - address: 10.10.10.6 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-05 +spec: {} +status: + addresses: + - address: 10.10.10.5 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-04 +spec: {} +status: + addresses: + - address: 10.10.10.4 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/ops.yaml new file mode 100644 index 000000000..105d37e9d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/ops.yaml @@ -0,0 +1,43 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-fleet-06 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-32 + - worker-31 + - worker-30 + - worker-29 + - worker-28 + - worker-27 + - worker-26 + - worker-25 + - worker-24 + - worker-23 + - worker-22 + - worker-21 + - worker-20 + - worker-19 + - worker-18 + - worker-17 + - worker-16 + - worker-15 + - worker-14 + - worker-13 + - worker-12 + - worker-11 + - worker-10 + - worker-09 + - worker-08 + - worker-07 + - worker-06 + - worker-05 + - worker-04 + - worker-03 + - worker-02 + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-01.yaml new file mode 100644 index 000000000..6c028938d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-01-ffe744b6 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-02.yaml new file mode 100644 index 000000000..1b2eda311 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-02-26dc3e6b + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-03.yaml new file mode 100644 index 000000000..76e067834 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-03-625c9b15 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-04.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-04.yaml new file mode 100644 index 000000000..7bff43f98 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-04.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-04", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.4" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.4" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-04 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-04-df682948 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-05.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-05.yaml new file mode 100644 index 000000000..460b6a92c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-05.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-05", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.5" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.5" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-05 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-05-84cb6c65 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-06.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-06.yaml new file mode 100644 index 000000000..56efb8d8f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-06.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-06", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.6" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.6" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-06 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-06-6c4f128d + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-07.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-07.yaml new file mode 100644 index 000000000..5a0c288a3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-07.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-07", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.7" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.7" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-07 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-07-1969a1be + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-08.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-08.yaml new file mode 100644 index 000000000..e4b70d23f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-08.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-08", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.8" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.8" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-08 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-08-e8c9177c + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-09.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-09.yaml new file mode 100644 index 000000000..8c47b4bb8 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-09.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-09", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.9" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.9" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-09 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-09-b34dbbd0 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-10.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-10.yaml new file mode 100644 index 000000000..abe22c79b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-10.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-10", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.10" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.10" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-10 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-10-b733b1f9 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-11.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-11.yaml new file mode 100644 index 000000000..5ae6cd67a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-11.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-11", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.11" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.11" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-11 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-11-1bc47d14 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-12.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-12.yaml new file mode 100644 index 000000000..7fb18e9f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-12.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-12", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.12" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.12" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-12 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-12-b0f25f46 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-13.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-13.yaml new file mode 100644 index 000000000..f5424e318 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-13.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-13", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.13" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.13" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-13 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-13-7fbbc189 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-14.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-14.yaml new file mode 100644 index 000000000..af133abe1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-14.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-14", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.14" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.14" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-14 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-14-579bb178 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-15.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-15.yaml new file mode 100644 index 000000000..fa7173cfc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-15.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-15", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.15" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.15" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-15 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-15-692323df + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-16.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-16.yaml new file mode 100644 index 000000000..c36a2b4d4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-16.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-16", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.16" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.16" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-16 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-16-9f967c4e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-17.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-17.yaml new file mode 100644 index 000000000..73a4b543a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-17.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-17", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.17" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.17" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-17 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-17-26afd598 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-18.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-18.yaml new file mode 100644 index 000000000..3373bf2a6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-18.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-18", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.18" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.18" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-18 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-18-ef135668 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-19.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-19.yaml new file mode 100644 index 000000000..a33b5844e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-19.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-19", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.19" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.19" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-19 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-19-ec74aff5 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-20.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-20.yaml new file mode 100644 index 000000000..b88531c3e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-20.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-20", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.20" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.20" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-20 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-20-4df14367 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-21.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-21.yaml new file mode 100644 index 000000000..f0a961b51 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-21.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-21", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.21" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.21" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-21 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-21-2ecf2959 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-22.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-22.yaml new file mode 100644 index 000000000..b31fe4586 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-22.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-22", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.22" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.22" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-22 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-22-70dbcabb + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-23.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-23.yaml new file mode 100644 index 000000000..9b72fbfd8 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-23.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-23", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.23" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.23" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-23 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-23-ddf51465 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-24.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-24.yaml new file mode 100644 index 000000000..da1fc78bc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-24.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-24", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.24" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.24" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-24 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-24-a1846eb1 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-25.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-25.yaml new file mode 100644 index 000000000..f83515d65 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-25.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-25", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.25" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.25" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-25 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-25-91238e5f + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-26.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-26.yaml new file mode 100644 index 000000000..ff612da9e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-26.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-26", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.26" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.26" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-26 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-26-5cc7b9a6 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-27.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-27.yaml new file mode 100644 index 000000000..dd7bac4c6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-27.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-27", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.27" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.27" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-27 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-27-60441017 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-28.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-28.yaml new file mode 100644 index 000000000..c52b346cc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-28.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-28", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.28" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.28" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-28 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-28-4f95993c + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-29.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-29.yaml new file mode 100644 index 000000000..c4e5d7883 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-29.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-29", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.29" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.29" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-29 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-29-c2d6e962 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-30.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-30.yaml new file mode 100644 index 000000000..c2eaddff4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-30.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-30", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.30" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.30" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-30 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-30-7f6f42d6 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-31.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-31.yaml new file mode 100644 index 000000000..dda91a1c5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-31.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-31", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.31" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.31" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-31 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-31-afad5007 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-32.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-32.yaml new file mode 100644 index 000000000..434e02956 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/reports/worker-32.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-32", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.32" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.32" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-32 + storage.simplyblock.io/nodeprobe-run: discover-fleet-06 + name: sb-nodeprobe-discover-fleet-06-worker-32-b730392c + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/case.md b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/case.md new file mode 100644 index 000000000..87870ca9d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/case.md @@ -0,0 +1,7 @@ +# FLEET-07 + +**Mutation.** Two ConfigMaps carrying a report for one node + +**Expected.** One report per node. The draft names the worker once + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/expected-notes.txt new file mode 100644 index 000000000..9ac93bb70 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-fleet-07 in Draft: 2 workers with 8 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/expected.yaml new file mode 100644 index 000000000..25be74039 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/expected.yaml @@ -0,0 +1,33 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-fleet-07 + name: discovered-discover-fleet-07 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-fleet-07-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - worker-02 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/nodes.yaml new file mode 100644 index 000000000..1fba89979 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/nodes.yaml @@ -0,0 +1,59 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/ops.yaml new file mode 100644 index 000000000..a633c320d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/ops.yaml @@ -0,0 +1,13 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-fleet-07 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/reports/report-2.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/reports/report-2.yaml new file mode 100644 index 000000000..6728a1168 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/reports/report-2.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-fleet-07 + name: sb-nodeprobe-discover-fleet-07-worker-01-c893b6f1-again + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/reports/worker-01.yaml new file mode 100644 index 000000000..bf290eca7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-fleet-07 + name: sb-nodeprobe-discover-fleet-07-worker-01-c893b6f1 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/reports/worker-02.yaml new file mode 100644 index 000000000..7e3644ff7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-fleet-07 + name: sb-nodeprobe-discover-fleet-07-worker-02-dc8a4407 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/case.md b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/case.md new file mode 100644 index 000000000..50c11855d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/case.md @@ -0,0 +1,7 @@ +# FLEET-08 + +**Mutation.** A report for a node absent from `status.workers` + +**Expected.** Skipped without an event: it is not this run's evidence + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/expected-notes.txt new file mode 100644 index 000000000..969946ccd --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-fleet-08 in Draft: 2 workers with 8 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/expected.yaml new file mode 100644 index 000000000..d645ea6ef --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/expected.yaml @@ -0,0 +1,33 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-fleet-08 + name: discovered-discover-fleet-08 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-fleet-08-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - worker-02 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/nodes.yaml new file mode 100644 index 000000000..98de08ee9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/nodes.yaml @@ -0,0 +1,89 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/ops.yaml new file mode 100644 index 000000000..aed8990f1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/ops.yaml @@ -0,0 +1,13 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-fleet-08 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/reports/worker-01.yaml new file mode 100644 index 000000000..db3e9e4dc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-fleet-08 + name: sb-nodeprobe-discover-fleet-08-worker-01-73b253ae + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/reports/worker-02.yaml new file mode 100644 index 000000000..964248a21 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-fleet-08 + name: sb-nodeprobe-discover-fleet-08-worker-02-930ebac5 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/reports/worker-03.yaml new file mode 100644 index 000000000..38f8a9d67 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-fleet-08 + name: sb-nodeprobe-discover-fleet-08-worker-03-b3438334 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-01-every-disk-mounted/case.md b/operator/internal/controllers/deployment/testdata/discovery/held/held-01-every-disk-mounted/case.md new file mode 100644 index 000000000..bc3393a7e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-01-every-disk-mounted/case.md @@ -0,0 +1,7 @@ +# HELD-01 + +**Mutation.** Every disk mounted + +**Expected.** The worker refused. Its device reasons folded into one line, counted + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-01-every-disk-mounted/expected-error.txt b/operator/internal/controllers/deployment/testdata/discovery/held/held-01-every-disk-mounted/expected-error.txt new file mode 100644 index 000000000..ca5fadb41 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-01-every-disk-mounted/expected-error.txt @@ -0,0 +1 @@ +no worker has a device this run would use: worker-01: no device of it survived the device rules (4 devices declined by available: the probe refused it: Mounted) diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-01-every-disk-mounted/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/held/held-01-every-disk-mounted/expected-refusals.txt new file mode 100644 index 000000000..b1c56b8ce --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-01-every-disk-mounted/expected-refusals.txt @@ -0,0 +1,5 @@ +worker-01/nvme0n1: declined by available because the probe refused it: Mounted +worker-01/nvme1n1: declined by available because the probe refused it: Mounted +worker-01/nvme2n1: declined by available because the probe refused it: Mounted +worker-01/nvme3n1: declined by available because the probe refused it: Mounted +worker-01: declined by has devices because no device of it survived the device rules diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-01-every-disk-mounted/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-01-every-disk-mounted/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-01-every-disk-mounted/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-01-every-disk-mounted/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-01-every-disk-mounted/ops.yaml new file mode 100644 index 000000000..3a7b3bb73 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-01-every-disk-mounted/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-held-01 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-01-every-disk-mounted/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-01-every-disk-mounted/reports/worker-01.yaml new file mode 100644 index 000000000..ed9850a1c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-01-every-disk-mounted/reports/worker-01.yaml @@ -0,0 +1,195 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Mounted", + "detail": "as the probe found it" + } + ] + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-held-01 + name: sb-nodeprobe-discover-held-01-worker-01-28c6b8cd + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-02-every-disk-in-a-device-mapper-stack/case.md b/operator/internal/controllers/deployment/testdata/discovery/held/held-02-every-disk-in-a-device-mapper-stack/case.md new file mode 100644 index 000000000..46fdad2f0 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-02-every-disk-in-a-device-mapper-stack/case.md @@ -0,0 +1,7 @@ +# HELD-02 + +**Mutation.** Every disk held in a device-mapper stack + +**Expected.** The same shape with the other reason + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-02-every-disk-in-a-device-mapper-stack/expected-error.txt b/operator/internal/controllers/deployment/testdata/discovery/held/held-02-every-disk-in-a-device-mapper-stack/expected-error.txt new file mode 100644 index 000000000..0b9b4c155 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-02-every-disk-in-a-device-mapper-stack/expected-error.txt @@ -0,0 +1 @@ +no worker has a device this run would use: worker-01: no device of it survived the device rules (4 devices declined by available: the probe refused it: Stacked) diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-02-every-disk-in-a-device-mapper-stack/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/held/held-02-every-disk-in-a-device-mapper-stack/expected-refusals.txt new file mode 100644 index 000000000..3e392751f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-02-every-disk-in-a-device-mapper-stack/expected-refusals.txt @@ -0,0 +1,5 @@ +worker-01/nvme0n1: declined by available because the probe refused it: Stacked +worker-01/nvme1n1: declined by available because the probe refused it: Stacked +worker-01/nvme2n1: declined by available because the probe refused it: Stacked +worker-01/nvme3n1: declined by available because the probe refused it: Stacked +worker-01: declined by has devices because no device of it survived the device rules diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-02-every-disk-in-a-device-mapper-stack/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-02-every-disk-in-a-device-mapper-stack/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-02-every-disk-in-a-device-mapper-stack/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-02-every-disk-in-a-device-mapper-stack/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-02-every-disk-in-a-device-mapper-stack/ops.yaml new file mode 100644 index 000000000..438d49092 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-02-every-disk-in-a-device-mapper-stack/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-held-02 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-02-every-disk-in-a-device-mapper-stack/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-02-every-disk-in-a-device-mapper-stack/reports/worker-01.yaml new file mode 100644 index 000000000..e81826979 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-02-every-disk-in-a-device-mapper-stack/reports/worker-01.yaml @@ -0,0 +1,195 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Stacked", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Stacked", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Stacked", + "detail": "as the probe found it" + } + ] + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": false, + "rejections": [ + { + "reason": "Stacked", + "detail": "as the probe found it" + } + ] + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-held-02 + name: sb-nodeprobe-discover-held-02-worker-01-7b6c33ef + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/case.md b/operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/case.md new file mode 100644 index 000000000..2c55a74fc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/case.md @@ -0,0 +1,7 @@ +# HELD-03 + +**Mutation.** 4 NVMe controllers on `uio_pci_generic`, idle, no block devices + +**Expected.** 4 synthesized devices, group `group-1-nvme-4xunsized` + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/expected-notes.txt new file mode 100644 index 000000000..0980eabcb --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-held-03 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/expected.yaml new file mode 100644 index 000000000..39902da23 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-held-03 + name: discovered-discover-held-03 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-held-03-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:60:00.0 + - 0000:61:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4xunsized + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/ops.yaml new file mode 100644 index 000000000..7da26f8f5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-held-03 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/reports/worker-01.yaml new file mode 100644 index 000000000..6d6d7731a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/reports/worker-01.yaml @@ -0,0 +1,159 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "nvmeControllers": [ + { + "address": "0000:5e:00.0", + "driver": "uio_pci_generic", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0, + "inUse": false + }, + { + "address": "0000:5f:00.0", + "driver": "uio_pci_generic", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0, + "inUse": false + }, + { + "address": "0000:60:00.0", + "driver": "uio_pci_generic", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0, + "inUse": false + }, + { + "address": "0000:61:00.0", + "driver": "uio_pci_generic", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0, + "inUse": false + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-held-03 + name: sb-nodeprobe-discover-held-03-worker-01-55fbb016 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-04-four-userspace-controllers-in-use/case.md b/operator/internal/controllers/deployment/testdata/discovery/held/held-04-four-userspace-controllers-in-use/case.md new file mode 100644 index 000000000..d47246461 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-04-four-userspace-controllers-in-use/case.md @@ -0,0 +1,7 @@ +# HELD-04 + +**Mutation.** The same controllers with `inUse: true` + +**Expected.** No devices. The worker refused with the message saying something is driving its disks + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-04-four-userspace-controllers-in-use/expected-error.txt b/operator/internal/controllers/deployment/testdata/discovery/held/held-04-four-userspace-controllers-in-use/expected-error.txt new file mode 100644 index 000000000..b780f9ec4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-04-four-userspace-controllers-in-use/expected-error.txt @@ -0,0 +1 @@ +no worker has a device this run would use: worker-01: it presents no usable block device, and 4 of its NVMe controllers (0000:5e:00.0 on uio_pci_generic, 0000:5f:00.0 on uio_pci_generic, 0000:60:00.0 on uio_pci_generic, 0000:61:00.0 on uio_pci_generic) are bound to a userspace driver and in use, so something is driving its disks diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-04-four-userspace-controllers-in-use/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/held/held-04-four-userspace-controllers-in-use/expected-refusals.txt new file mode 100644 index 000000000..b9a29cff5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-04-four-userspace-controllers-in-use/expected-refusals.txt @@ -0,0 +1 @@ +worker-01: declined by has devices because it presents no usable block device, and 4 of its NVMe controllers (0000:5e:00.0 on uio_pci_generic, 0000:5f:00.0 on uio_pci_generic, 0000:60:00.0 on uio_pci_generic, 0000:61:00.0 on uio_pci_generic) are bound to a userspace driver and in use, so something is driving its disks diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-04-four-userspace-controllers-in-use/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-04-four-userspace-controllers-in-use/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-04-four-userspace-controllers-in-use/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-04-four-userspace-controllers-in-use/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-04-four-userspace-controllers-in-use/ops.yaml new file mode 100644 index 000000000..3a50c2f09 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-04-four-userspace-controllers-in-use/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-held-04 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-04-four-userspace-controllers-in-use/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-04-four-userspace-controllers-in-use/reports/worker-01.yaml new file mode 100644 index 000000000..29ee86dfc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-04-four-userspace-controllers-in-use/reports/worker-01.yaml @@ -0,0 +1,159 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "nvmeControllers": [ + { + "address": "0000:5e:00.0", + "driver": "uio_pci_generic", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0, + "inUse": true + }, + { + "address": "0000:5f:00.0", + "driver": "uio_pci_generic", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0, + "inUse": true + }, + { + "address": "0000:60:00.0", + "driver": "uio_pci_generic", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0, + "inUse": true + }, + { + "address": "0000:61:00.0", + "driver": "uio_pci_generic", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0, + "inUse": true + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-held-04 + name: sb-nodeprobe-discover-held-04-worker-01-d9950e34 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/case.md b/operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/case.md new file mode 100644 index 000000000..21037f26b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/case.md @@ -0,0 +1,7 @@ +# HELD-05 + +**Mutation.** 2 idle userspace controllers beside 2 kernel-presented disks + +**Expected.** 4 addresses in the group, the kernel-bound pair counted once + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/expected-notes.txt new file mode 100644 index 000000000..3cccaae78 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-held-05 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/expected.yaml new file mode 100644 index 000000000..69be29fa0 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-held-05 + name: discovered-discover-held-05 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-held-05-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x1.5T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/ops.yaml new file mode 100644 index 000000000..4059573d7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-held-05 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/reports/worker-01.yaml new file mode 100644 index 000000000..89df61de6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/reports/worker-01.yaml @@ -0,0 +1,185 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ], + "nvmeControllers": [ + { + "address": "0000:5e:00.0", + "driver": "uio_pci_generic", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0, + "inUse": false + }, + { + "address": "0000:5f:00.0", + "driver": "uio_pci_generic", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0, + "inUse": false + }, + { + "address": "0000:af:00.0", + "driver": "nvme", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0, + "inUse": false + }, + { + "address": "0000:b0:00.0", + "driver": "nvme", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0, + "inUse": false + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-held-05 + name: sb-nodeprobe-discover-held-05-worker-01-d55c608b + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/case.md b/operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/case.md new file mode 100644 index 000000000..99cf74ad2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/case.md @@ -0,0 +1,7 @@ +# HELD-06 + +**Mutation.** An idle controller on `vfio-pci` + +**Expected.** Claimable on the same terms as `uio_pci_generic` + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/expected-notes.txt new file mode 100644 index 000000000..d85aac69c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-held-06 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/expected.yaml new file mode 100644 index 000000000..3d3a966e8 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-held-06 + name: discovered-discover-held-06 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-held-06-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:60:00.0 + - 0000:61:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4xunsized + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/ops.yaml new file mode 100644 index 000000000..aba93d67a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-held-06 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/reports/worker-01.yaml new file mode 100644 index 000000000..844f2b23f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/reports/worker-01.yaml @@ -0,0 +1,159 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "nvmeControllers": [ + { + "address": "0000:5e:00.0", + "driver": "vfio-pci", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0, + "inUse": false + }, + { + "address": "0000:5f:00.0", + "driver": "vfio-pci", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0, + "inUse": false + }, + { + "address": "0000:60:00.0", + "driver": "vfio-pci", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0, + "inUse": false + }, + { + "address": "0000:61:00.0", + "driver": "vfio-pci", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0, + "inUse": false + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-held-06 + name: sb-nodeprobe-discover-held-06-worker-01-077efd70 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/case.md b/operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/case.md new file mode 100644 index 000000000..5c291278d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/case.md @@ -0,0 +1,7 @@ +# HELD-07 + +**Mutation.** A kernel-bound controller whose block device is already reported + +**Expected.** Named once, not twice + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/expected-notes.txt new file mode 100644 index 000000000..bb8c25be4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-held-07 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/expected.yaml new file mode 100644 index 000000000..617767bbf --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-held-07 + name: discovered-discover-held-07 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-held-07-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/ops.yaml new file mode 100644 index 000000000..cff49173c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-held-07 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/reports/worker-01.yaml new file mode 100644 index 000000000..079bde6ad --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/reports/worker-01.yaml @@ -0,0 +1,209 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ], + "nvmeControllers": [ + { + "address": "0000:5e:00.0", + "driver": "nvme", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0, + "inUse": false + }, + { + "address": "0000:5f:00.0", + "driver": "nvme", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0, + "inUse": false + }, + { + "address": "0000:af:00.0", + "driver": "nvme", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0, + "inUse": false + }, + { + "address": "0000:b0:00.0", + "driver": "nvme", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0, + "inUse": false + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-held-07 + name: sb-nodeprobe-discover-held-07-worker-01-a0d22df9 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-08-idle-controllers-on-a-block-run/case.md b/operator/internal/controllers/deployment/testdata/discovery/held/held-08-idle-controllers-on-a-block-run/case.md new file mode 100644 index 000000000..0e0c0bba9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-08-idle-controllers-on-a-block-run/case.md @@ -0,0 +1,7 @@ +# HELD-08 + +**Mutation.** Idle userspace controllers on a block run + +**Expected.** No devices at all: a controller with no block device has no path. Worker refused + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-08-idle-controllers-on-a-block-run/expected-error.txt b/operator/internal/controllers/deployment/testdata/discovery/held/held-08-idle-controllers-on-a-block-run/expected-error.txt new file mode 100644 index 000000000..54cda083c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-08-idle-controllers-on-a-block-run/expected-error.txt @@ -0,0 +1 @@ +no worker has a device this run would use: worker-01: it presents no usable block device, and 4 of its NVMe controllers (0000:5e:00.0 on uio_pci_generic, 0000:5f:00.0 on uio_pci_generic, 0000:60:00.0 on uio_pci_generic, 0000:61:00.0 on uio_pci_generic) are bound to a userspace driver and nothing is using them, so the disks are there to be reclaimed diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-08-idle-controllers-on-a-block-run/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/held/held-08-idle-controllers-on-a-block-run/expected-refusals.txt new file mode 100644 index 000000000..53b6b8186 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-08-idle-controllers-on-a-block-run/expected-refusals.txt @@ -0,0 +1 @@ +worker-01: declined by has devices because it presents no usable block device, and 4 of its NVMe controllers (0000:5e:00.0 on uio_pci_generic, 0000:5f:00.0 on uio_pci_generic, 0000:60:00.0 on uio_pci_generic, 0000:61:00.0 on uio_pci_generic) are bound to a userspace driver and nothing is using them, so the disks are there to be reclaimed diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-08-idle-controllers-on-a-block-run/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-08-idle-controllers-on-a-block-run/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-08-idle-controllers-on-a-block-run/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-08-idle-controllers-on-a-block-run/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-08-idle-controllers-on-a-block-run/ops.yaml new file mode 100644 index 000000000..c5964919f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-08-idle-controllers-on-a-block-run/ops.yaml @@ -0,0 +1,15 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-held-08 + namespace: simplyblock +spec: + action: Discover + discover: + deviceFilter: + enableLogicalBlockDevices: true +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-08-idle-controllers-on-a-block-run/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-08-idle-controllers-on-a-block-run/reports/worker-01.yaml new file mode 100644 index 000000000..5df995cec --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-08-idle-controllers-on-a-block-run/reports/worker-01.yaml @@ -0,0 +1,159 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "nvmeControllers": [ + { + "address": "0000:5e:00.0", + "driver": "uio_pci_generic", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0, + "inUse": false + }, + { + "address": "0000:5f:00.0", + "driver": "uio_pci_generic", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0, + "inUse": false + }, + { + "address": "0000:60:00.0", + "driver": "uio_pci_generic", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0, + "inUse": false + }, + { + "address": "0000:61:00.0", + "driver": "uio_pci_generic", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0, + "inUse": false + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-held-08 + name: sb-nodeprobe-discover-held-08-worker-01-d0a61743 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/case.md b/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/case.md new file mode 100644 index 000000000..84983c1e1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/case.md @@ -0,0 +1,7 @@ +# HELD-09 + +**Mutation.** 16 loopback devices and 4 free disks + +**Expected.** The 4 disks drafted. The loopbacks absent from `Explain` and present in `RefusalLines` + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/expected-notes.txt new file mode 100644 index 000000000..6bdb42422 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-held-09 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 16 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/expected-refusals.txt new file mode 100644 index 000000000..f0b1f6803 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/expected-refusals.txt @@ -0,0 +1,16 @@ +worker-01/loop0: declined by whole disk because it is a Loop rather than a whole disk +worker-01/loop1: declined by whole disk because it is a Loop rather than a whole disk +worker-01/loop2: declined by whole disk because it is a Loop rather than a whole disk +worker-01/loop3: declined by whole disk because it is a Loop rather than a whole disk +worker-01/loop4: declined by whole disk because it is a Loop rather than a whole disk +worker-01/loop5: declined by whole disk because it is a Loop rather than a whole disk +worker-01/loop6: declined by whole disk because it is a Loop rather than a whole disk +worker-01/loop7: declined by whole disk because it is a Loop rather than a whole disk +worker-01/loop8: declined by whole disk because it is a Loop rather than a whole disk +worker-01/loop9: declined by whole disk because it is a Loop rather than a whole disk +worker-01/loop10: declined by whole disk because it is a Loop rather than a whole disk +worker-01/loop11: declined by whole disk because it is a Loop rather than a whole disk +worker-01/loop12: declined by whole disk because it is a Loop rather than a whole disk +worker-01/loop13: declined by whole disk because it is a Loop rather than a whole disk +worker-01/loop14: declined by whole disk because it is a Loop rather than a whole disk +worker-01/loop15: declined by whole disk because it is a Loop rather than a whole disk diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/expected.yaml new file mode 100644 index 000000000..e37328793 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-held-09 + name: discovered-discover-held-09 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-held-09-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/ops.yaml new file mode 100644 index 000000000..5bce1dd6d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-held-09 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/reports/worker-01.yaml new file mode 100644 index 000000000..0507eebd6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/reports/worker-01.yaml @@ -0,0 +1,319 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "loop0", + "path": "/dev/loop0", + "sizeBytes": 68719476736, + "kind": "Loop", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "loop1", + "path": "/dev/loop1", + "sizeBytes": 68719476736, + "kind": "Loop", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "loop2", + "path": "/dev/loop2", + "sizeBytes": 68719476736, + "kind": "Loop", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "loop3", + "path": "/dev/loop3", + "sizeBytes": 68719476736, + "kind": "Loop", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "loop4", + "path": "/dev/loop4", + "sizeBytes": 68719476736, + "kind": "Loop", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "loop5", + "path": "/dev/loop5", + "sizeBytes": 68719476736, + "kind": "Loop", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "loop6", + "path": "/dev/loop6", + "sizeBytes": 68719476736, + "kind": "Loop", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "loop7", + "path": "/dev/loop7", + "sizeBytes": 68719476736, + "kind": "Loop", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "loop8", + "path": "/dev/loop8", + "sizeBytes": 68719476736, + "kind": "Loop", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "loop9", + "path": "/dev/loop9", + "sizeBytes": 68719476736, + "kind": "Loop", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "loop10", + "path": "/dev/loop10", + "sizeBytes": 68719476736, + "kind": "Loop", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "loop11", + "path": "/dev/loop11", + "sizeBytes": 68719476736, + "kind": "Loop", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "loop12", + "path": "/dev/loop12", + "sizeBytes": 68719476736, + "kind": "Loop", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "loop13", + "path": "/dev/loop13", + "sizeBytes": 68719476736, + "kind": "Loop", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "loop14", + "path": "/dev/loop14", + "sizeBytes": 68719476736, + "kind": "Loop", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "loop15", + "path": "/dev/loop15", + "sizeBytes": 68719476736, + "kind": "Loop", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-held-09 + name: sb-nodeprobe-discover-held-09-worker-01-6d447c0f + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/case.md b/operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/case.md new file mode 100644 index 000000000..aff2e83e0 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/case.md @@ -0,0 +1,7 @@ +# HELD-10 + +**Mutation.** Half the controllers idle and half in use, no block devices + +**Expected.** Only the idle half is offered + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/expected-notes.txt new file mode 100644 index 000000000..b2cda69b3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-held-10 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/expected.yaml new file mode 100644 index 000000000..2f0575538 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/expected.yaml @@ -0,0 +1,30 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-held-10 + name: discovered-discover-held-10 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-held-10-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + mgmtInterface: eth0 + name: group-1-nvme-2xunsized + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/ops.yaml new file mode 100644 index 000000000..abc2e6bc3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-held-10 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/reports/worker-01.yaml new file mode 100644 index 000000000..4c3716250 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/reports/worker-01.yaml @@ -0,0 +1,159 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "nvmeControllers": [ + { + "address": "0000:5e:00.0", + "driver": "uio_pci_generic", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0, + "inUse": false + }, + { + "address": "0000:5f:00.0", + "driver": "uio_pci_generic", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0, + "inUse": false + }, + { + "address": "0000:60:00.0", + "driver": "uio_pci_generic", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0, + "inUse": true + }, + { + "address": "0000:61:00.0", + "driver": "uio_pci_generic", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0, + "inUse": true + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-held-10 + name: sb-nodeprobe-discover-held-10-worker-01-9d484ab1 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-11-controllers-nothing-could-check/case.md b/operator/internal/controllers/deployment/testdata/discovery/held/held-11-controllers-nothing-could-check/case.md new file mode 100644 index 000000000..37bc88a83 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-11-controllers-nothing-could-check/case.md @@ -0,0 +1,7 @@ +# HELD-11 + +**Mutation.** Userspace controllers whose process table the probe could not read + +**Expected.** Refused, and the message says whether anything is driving them could not be established. An unchecked controller is not a free one + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-11-controllers-nothing-could-check/expected-error.txt b/operator/internal/controllers/deployment/testdata/discovery/held/held-11-controllers-nothing-could-check/expected-error.txt new file mode 100644 index 000000000..538779616 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-11-controllers-nothing-could-check/expected-error.txt @@ -0,0 +1 @@ +no worker has a device this run would use: worker-01: it presents no usable block device, and whether anything is driving 2 of its NVMe controllers (0000:5e:00.0 on uio_pci_generic, 0000:5f:00.0 on uio_pci_generic) could not be established, so they are not offered diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-11-controllers-nothing-could-check/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/held/held-11-controllers-nothing-could-check/expected-refusals.txt new file mode 100644 index 000000000..858e788d9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-11-controllers-nothing-could-check/expected-refusals.txt @@ -0,0 +1 @@ +worker-01: declined by has devices because it presents no usable block device, and whether anything is driving 2 of its NVMe controllers (0000:5e:00.0 on uio_pci_generic, 0000:5f:00.0 on uio_pci_generic) could not be established, so they are not offered diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-11-controllers-nothing-could-check/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-11-controllers-nothing-could-check/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-11-controllers-nothing-could-check/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-11-controllers-nothing-could-check/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-11-controllers-nothing-could-check/ops.yaml new file mode 100644 index 000000000..f3c58e504 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-11-controllers-nothing-could-check/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-held-11 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-11-controllers-nothing-could-check/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-11-controllers-nothing-could-check/reports/worker-01.yaml new file mode 100644 index 000000000..fd81860f5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-11-controllers-nothing-could-check/reports/worker-01.yaml @@ -0,0 +1,144 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "nvmeControllers": [ + { + "address": "0000:5e:00.0", + "driver": "uio_pci_generic", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0 + }, + { + "address": "0000:5f:00.0", + "driver": "uio_pci_generic", + "vendor": "0x144d", + "product": "0xa80a", + "numaNode": 0 + } + ], + "unreadable": [ + "read the process table: permission denied" + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-held-11 + name: sb-nodeprobe-discover-held-11-worker-01-349735ae + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/case.md new file mode 100644 index 000000000..23d13d90e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/case.md @@ -0,0 +1,7 @@ +# NET-01 + +**Mutation.** One 10G interface, up, holding a routable address + +**Expected.** `mgmtInterface` names it + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/expected-notes.txt new file mode 100644 index 000000000..376a82db0 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-01 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/expected.yaml new file mode 100644 index 000000000..b6ba36b2d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-01 + name: discovered-discover-net-01 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-01-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/ops.yaml new file mode 100644 index 000000000..53dc19c7f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-01 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/reports/worker-01.yaml new file mode 100644 index 000000000..0bff5c714 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/reports/worker-01.yaml @@ -0,0 +1,163 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 10000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-01 + name: sb-nodeprobe-discover-net-01-worker-01-23a3436e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/case.md new file mode 100644 index 000000000..7e89fc0d2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/case.md @@ -0,0 +1,7 @@ +# NET-02 + +**Mutation.** A 1G and a 10G interface, both up and addressed, no node address on either + +**Expected.** The 10G is named + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/expected-notes.txt new file mode 100644 index 000000000..9b19a4c5b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-02 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/expected.yaml new file mode 100644 index 000000000..b362cd5d3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-02 + name: discovered-discover-net-02 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-02-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth1 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/nodes.yaml new file mode 100644 index 000000000..94be916b6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/nodes.yaml @@ -0,0 +1,26 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/ops.yaml new file mode 100644 index 000000000..fac7014bb --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-02 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/reports/worker-01.yaml new file mode 100644 index 000000000..791a67eb7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 1000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "192.168.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 10000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-02 + name: sb-nodeprobe-discover-net-02-worker-01-ae8e1fa5 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/case.md new file mode 100644 index 000000000..0c2578403 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/case.md @@ -0,0 +1,7 @@ +# NET-03 + +**Mutation.** The node's `InternalIP` on the 1G while the 10G is faster + +**Expected.** The 1G is named: the node address wins outright + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/expected-notes.txt new file mode 100644 index 000000000..979d19332 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-03 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/expected.yaml new file mode 100644 index 000000000..59a300a5f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-03 + name: discovered-discover-net-03 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-03-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/ops.yaml new file mode 100644 index 000000000..ca6bb919b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-03 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/reports/worker-01.yaml new file mode 100644 index 000000000..ebaf51534 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 1000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 10000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-03 + name: sb-nodeprobe-discover-net-03-worker-01-8520b11b + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/case.md new file mode 100644 index 000000000..ae5ed06f8 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/case.md @@ -0,0 +1,9 @@ +# NET-04 + +**Mutation.** Four interfaces, 1G and 10G, all unusable: down, bridged, or link-local only + +**Expected.** `mgmtInterface` empty and the draft still written. See §14, gap G-7 + +**Harness.** `CM` + +**Gap.** G-7. This case records what the generator does today, so that the day it changes the diff is the finding. diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/expected-notes.txt new file mode 100644 index 000000000..727152069 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-04 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/expected.yaml new file mode 100644 index 000000000..d47e16295 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/expected.yaml @@ -0,0 +1,31 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-04 + name: discovered-discover-net-04 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-04-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/nodes.yaml new file mode 100644 index 000000000..94be916b6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/nodes.yaml @@ -0,0 +1,26 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/ops.yaml new file mode 100644 index 000000000..5bd747b26 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-04 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/reports/worker-01.yaml new file mode 100644 index 000000000..91cde6dbd --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/reports/worker-01.yaml @@ -0,0 +1,201 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 1000, + "mtu": 1500, + "state": "down", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "192.168.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 10000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "169.254.1.1" + ] + }, + { + "name": "cni0", + "speedMbps": 1000, + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "bridge": true, + "kind": "bridge", + "addresses": [ + "10.42.2.1" + ] + }, + { + "name": "br0", + "speedMbps": 10000, + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "bridge": true, + "kind": "bridge", + "addresses": [ + "192.168.1.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-04 + name: sb-nodeprobe-discover-net-04-worker-01-24a3e4af + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/case.md new file mode 100644 index 000000000..006285ca9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/case.md @@ -0,0 +1,7 @@ +# NET-05 + +**Mutation.** `interfaces` absent from the report + +**Expected.** The same: an empty interface and a written draft + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/expected-notes.txt new file mode 100644 index 000000000..96e405a0f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-05 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/expected.yaml new file mode 100644 index 000000000..fce2f50d2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/expected.yaml @@ -0,0 +1,31 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-05 + name: discovered-discover-net-05 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-05-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/nodes.yaml new file mode 100644 index 000000000..94be916b6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/nodes.yaml @@ -0,0 +1,26 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/ops.yaml new file mode 100644 index 000000000..cb4b71111 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-05 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/reports/worker-01.yaml new file mode 100644 index 000000000..a404ed7c2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/reports/worker-01.yaml @@ -0,0 +1,149 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-05 + name: sb-nodeprobe-discover-net-05-worker-01-e9fdaf8e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/case.md new file mode 100644 index 000000000..0a934dbfc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/case.md @@ -0,0 +1,9 @@ +# NET-06 + +**Mutation.** Two workers with identical disks calling their NIC `eth0` and `ens5f0` + +**Expected.** 2 groups, which the draft validation then reports as unbuildable. See §14, gap G-25 + +**Harness.** `CM` + +**Gap.** G-25. This case records what the generator does today, so that the day it changes the diff is the finding. diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/expected-notes.txt new file mode 100644 index 000000000..015f3a906 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-06 in Draft: 2 workers with 8 nvme devices, in 2 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/expected.yaml new file mode 100644 index 000000000..ce8e871d3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/expected.yaml @@ -0,0 +1,42 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-06 + name: discovered-discover-net-06 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-06-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: ens5f0 + name: group-2-nvme-4x3T + workers: + - worker-02 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/nodes.yaml new file mode 100644 index 000000000..234ddc29e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/nodes.yaml @@ -0,0 +1,53 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/ops.yaml new file mode 100644 index 000000000..66d150d8e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/ops.yaml @@ -0,0 +1,13 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-06 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/reports/worker-01.yaml new file mode 100644 index 000000000..6b65ba85f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/reports/worker-01.yaml @@ -0,0 +1,163 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "192.168.10.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-06 + name: sb-nodeprobe-discover-net-06-worker-01-12ed112b + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/reports/worker-02.yaml new file mode 100644 index 000000000..c5809863c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/reports/worker-02.yaml @@ -0,0 +1,163 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "ens5f0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "192.168.10.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-net-06 + name: sb-nodeprobe-discover-net-06-worker-02-3facc229 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/case.md new file mode 100644 index 000000000..af28e781d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/case.md @@ -0,0 +1,7 @@ +# NET-07 + +**Mutation.** Two 10G interfaces, `eth1` and `eth0`, both addressed + +**Expected.** `eth0` is named. A second run names it again + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/expected-notes.txt new file mode 100644 index 000000000..fe609fd1e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-07 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/expected.yaml new file mode 100644 index 000000000..5f0092277 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-07 + name: discovered-discover-net-07 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-07-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/nodes.yaml new file mode 100644 index 000000000..94be916b6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/nodes.yaml @@ -0,0 +1,26 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/ops.yaml new file mode 100644 index 000000000..6c4ad69f2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-07 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/reports/worker-01.yaml new file mode 100644 index 000000000..b3bf6a9f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth1", + "speedMbps": 10000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + }, + { + "name": "eth0", + "speedMbps": 10000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "192.168.10.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-07 + name: sb-nodeprobe-discover-net-07-worker-01-30691e51 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/case.md new file mode 100644 index 000000000..e72700b42 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/case.md @@ -0,0 +1,7 @@ +# NET-08 + +**Mutation.** An addressed physical interface reporting speed 0 beside a 10G interface that is down + +**Expected.** The addressed one is named + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/expected-notes.txt new file mode 100644 index 000000000..c5835aa2b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-08 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/expected.yaml new file mode 100644 index 000000000..f8652bea3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-08 + name: discovered-discover-net-08 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-08-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/nodes.yaml new file mode 100644 index 000000000..94be916b6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/nodes.yaml @@ -0,0 +1,26 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/ops.yaml new file mode 100644 index 000000000..1ceec2367 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-08 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/reports/worker-01.yaml new file mode 100644 index 000000000..487d6b4fe --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/reports/worker-01.yaml @@ -0,0 +1,174 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "192.168.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 10000, + "mtu": 1500, + "state": "down", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-08 + name: sb-nodeprobe-discover-net-08-worker-01-0e668ece + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/case.md new file mode 100644 index 000000000..6e48acdc4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/case.md @@ -0,0 +1,7 @@ +# NET-09 + +**Mutation.** A worker whose `InternalIP` sits on a bridge, as a hypervisor host's does + +**Expected.** `br0` named. A bridge carrying the cluster's own address is the management network + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/expected-notes.txt new file mode 100644 index 000000000..e1ede782a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-09 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/expected.yaml new file mode 100644 index 000000000..d867c0c2c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-09 + name: discovered-discover-net-09 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-09-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: br0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/ops.yaml new file mode 100644 index 000000000..36e42796f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-09 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/reports/worker-01.yaml new file mode 100644 index 000000000..c5965041a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/reports/worker-01.yaml @@ -0,0 +1,179 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "br0", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "bridge": true, + "kind": "bridge", + "lower": [ + "eth0" + ], + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth0", + "speedMbps": 10000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:3b:00.0", + "numaNode": 0, + "kind": "physical", + "upper": [ + "br0" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-09 + name: sb-nodeprobe-discover-net-09-worker-01-6ccc8073 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/case.md new file mode 100644 index 000000000..7ccef989f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/case.md @@ -0,0 +1,7 @@ +# NET-10 + +**Mutation.** One interface holding only a global IPv6 address + +**Expected.** Named: reachability is not address-family dependent + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/expected-notes.txt new file mode 100644 index 000000000..869d2e1fb --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-10 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/expected.yaml new file mode 100644 index 000000000..092036692 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-10 + name: discovered-discover-net-10 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-10-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/nodes.yaml new file mode 100644 index 000000000..94be916b6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/nodes.yaml @@ -0,0 +1,26 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/ops.yaml new file mode 100644 index 000000000..6c6ba403e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-10 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/reports/worker-01.yaml new file mode 100644 index 000000000..de5b4d719 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/reports/worker-01.yaml @@ -0,0 +1,163 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 10000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "2001:db8::1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-10 + name: sb-nodeprobe-discover-net-10-worker-01-d27b1648 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/case.md new file mode 100644 index 000000000..6faf34265 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/case.md @@ -0,0 +1,7 @@ +# NET-11 + +**Mutation.** An interface holding both a link-local and a routable address + +**Expected.** Named on the routable one + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/expected-notes.txt new file mode 100644 index 000000000..11e2257cb --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-11 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/expected.yaml new file mode 100644 index 000000000..c3850c6d1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-11 + name: discovered-discover-net-11 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-11-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/nodes.yaml new file mode 100644 index 000000000..94be916b6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/nodes.yaml @@ -0,0 +1,26 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/ops.yaml new file mode 100644 index 000000000..23798231b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-11 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/reports/worker-01.yaml new file mode 100644 index 000000000..ec5e22d47 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/reports/worker-01.yaml @@ -0,0 +1,164 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 10000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "169.254.1.1", + "192.168.10.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-11 + name: sb-nodeprobe-discover-net-11-worker-01-96feff1b + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/case.md new file mode 100644 index 000000000..15eb3cbda --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/case.md @@ -0,0 +1,7 @@ +# NET-12 + +**Mutation.** One interface and it is the loopback + +**Expected.** `mgmtInterface` empty: loopback is refused by its ARPHRD type, not its name + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/expected-notes.txt new file mode 100644 index 000000000..fedc2414b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-12 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/expected.yaml new file mode 100644 index 000000000..34c3e4983 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/expected.yaml @@ -0,0 +1,31 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-12 + name: discovered-discover-net-12 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-12-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/nodes.yaml new file mode 100644 index 000000000..94be916b6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/nodes.yaml @@ -0,0 +1,26 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/ops.yaml new file mode 100644 index 000000000..8c1957633 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-12 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/reports/worker-01.yaml new file mode 100644 index 000000000..1681eb493 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/reports/worker-01.yaml @@ -0,0 +1,163 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "lo", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "loopback": true, + "kind": "loopback", + "addresses": [ + "127.0.0.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-12 + name: sb-nodeprobe-discover-net-12-worker-01-170ee41e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/case.md new file mode 100644 index 000000000..3f2a543e7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/case.md @@ -0,0 +1,7 @@ +# NET-13 + +**Mutation.** `bond0` holding the node address over two unaddressed physical slaves + +**Expected.** `bond0` named, with `eth0` and `eth1` reported as the hardware under it + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/expected-notes.txt new file mode 100644 index 000000000..a4b5dfa06 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-13 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/expected.yaml new file mode 100644 index 000000000..577d2e9b4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-13 + name: discovered-discover-net-13 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-13-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: bond0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/ops.yaml new file mode 100644 index 000000000..562d66a05 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-13 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/reports/worker-01.yaml new file mode 100644 index 000000000..f0c63d743 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/reports/worker-01.yaml @@ -0,0 +1,218 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "bond0", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "bond", + "lower": [ + "eth0", + "eth1" + ], + "upper": [ + "bond0.100" + ], + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "bond0.100", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "vlan", + "lower": [ + "bond0" + ] + }, + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:3b:00.0", + "numaNode": 0, + "kind": "physical", + "upper": [ + "bond0" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:3b:00.1", + "numaNode": 0, + "kind": "physical", + "upper": [ + "bond0" + ] + }, + { + "name": "lo", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "loopback": true, + "kind": "loopback", + "addresses": [ + "127.0.0.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-13 + name: sb-nodeprobe-discover-net-13-worker-01-24517756 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/case.md new file mode 100644 index 000000000..27ecb218b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/case.md @@ -0,0 +1,7 @@ +# NET-14 + +**Mutation.** `eth0.100` holding the node address, the parent `eth0` unaddressed + +**Expected.** `eth0.100` named, resolving to `eth0` for its slot and memory node + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/expected-notes.txt new file mode 100644 index 000000000..cba6728c9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-14 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/expected.yaml new file mode 100644 index 000000000..61bdee7f2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-14 + name: discovered-discover-net-14 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-14-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0.100 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/ops.yaml new file mode 100644 index 000000000..364fce826 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-14 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/reports/worker-01.yaml new file mode 100644 index 000000000..97f520dab --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/reports/worker-01.yaml @@ -0,0 +1,178 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:3b:00.0", + "numaNode": 0, + "kind": "physical", + "upper": [ + "eth0.100" + ] + }, + { + "name": "eth0.100", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "vlan", + "lower": [ + "eth0" + ], + "addresses": [ + "10.10.10.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-14 + name: sb-nodeprobe-discover-net-14-worker-01-b78b5ab4 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/case.md new file mode 100644 index 000000000..a2c7dedb2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/case.md @@ -0,0 +1,7 @@ +# NET-15 + +**Mutation.** `vxlan.calico` holding an overlay address beside `eth0` holding the node address + +**Expected.** `eth0` named. The overlay is refused for holding an address the cluster does not reach the machine on + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/expected-notes.txt new file mode 100644 index 000000000..49b1d8e28 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-15 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/expected.yaml new file mode 100644 index 000000000..3ad553fec --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-15 + name: discovered-discover-net-15 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-15-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/ops.yaml new file mode 100644 index 000000000..664e0e55a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-15 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/reports/worker-01.yaml new file mode 100644 index 000000000..a7c8281cc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/reports/worker-01.yaml @@ -0,0 +1,174 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 10000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "vxlan.calico", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "vxlan", + "addresses": [ + "10.42.2.0" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-15 + name: sb-nodeprobe-discover-net-15-worker-01-441be5c7 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/case.md new file mode 100644 index 000000000..677e7a2a7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/case.md @@ -0,0 +1,7 @@ +# NET-16 + +**Mutation.** `bond0.100` over `bond0` over two slaves, the node address on the VLAN + +**Expected.** `bond0.100` named, resolving through the bond to both NICs + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/expected-notes.txt new file mode 100644 index 000000000..e30b7e9ea --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-16 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/expected.yaml new file mode 100644 index 000000000..3aaa0aaef --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-16 + name: discovered-discover-net-16 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-16-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: bond0.100 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/ops.yaml new file mode 100644 index 000000000..9d2b96f22 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-16 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/reports/worker-01.yaml new file mode 100644 index 000000000..a9e9b822b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/reports/worker-01.yaml @@ -0,0 +1,218 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "bond0", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "bond", + "lower": [ + "eth0", + "eth1" + ], + "upper": [ + "bond0.100" + ] + }, + { + "name": "bond0.100", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "vlan", + "lower": [ + "bond0" + ], + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:3b:00.0", + "numaNode": 0, + "kind": "physical", + "upper": [ + "bond0" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:3b:00.1", + "numaNode": 0, + "kind": "physical", + "upper": [ + "bond0" + ] + }, + { + "name": "lo", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "loopback": true, + "kind": "loopback", + "addresses": [ + "127.0.0.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-16 + name: sb-nodeprobe-discover-net-16-worker-01-62190b65 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/case.md new file mode 100644 index 000000000..c65a29921 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/case.md @@ -0,0 +1,7 @@ +# NET-17 + +**Mutation.** A `macvlan` and an `ipvlan` over `eth0`, all three addressed + +**Expected.** `eth0` named: the three carry the same traffic and the physical one is the simplest + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/expected-notes.txt new file mode 100644 index 000000000..b1f94f0ff --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-17 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/expected.yaml new file mode 100644 index 000000000..8bee3a73e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-17 + name: discovered-discover-net-17 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-17-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/ops.yaml new file mode 100644 index 000000000..6d7dc47ed --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-17 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/reports/worker-01.yaml new file mode 100644 index 000000000..68c8169c5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/reports/worker-01.yaml @@ -0,0 +1,195 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "upper": [ + "macvlan0", + "ipvlan0" + ], + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "macvlan0", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "macvlan", + "lower": [ + "eth0" + ], + "addresses": [ + "192.168.10.5" + ] + }, + { + "name": "ipvlan0", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "ipvlan", + "lower": [ + "eth0" + ], + "addresses": [ + "192.168.10.6" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-17 + name: sb-nodeprobe-discover-net-17-worker-01-967d7e3e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/case.md new file mode 100644 index 000000000..66170b791 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/case.md @@ -0,0 +1,7 @@ +# NET-18 + +**Mutation.** Twelve `veth` interfaces holding pod-CIDR addresses beside one physical NIC + +**Expected.** The physical NIC named, whatever the veth count + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/expected-notes.txt new file mode 100644 index 000000000..b66e10811 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-18 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/expected.yaml new file mode 100644 index 000000000..8ed81fcbf --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-18 + name: discovered-discover-net-18 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-18-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/ops.yaml new file mode 100644 index 000000000..3c5aace94 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-18 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/reports/worker-01.yaml new file mode 100644 index 000000000..ddd2aeb50 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/reports/worker-01.yaml @@ -0,0 +1,295 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "veth7a10", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "virtual", + "addresses": [ + "10.42.2.10" + ] + }, + { + "name": "veth7a11", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "virtual", + "addresses": [ + "10.42.2.11" + ] + }, + { + "name": "veth7a12", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "virtual", + "addresses": [ + "10.42.2.12" + ] + }, + { + "name": "veth7a13", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "virtual", + "addresses": [ + "10.42.2.13" + ] + }, + { + "name": "veth7a14", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "virtual", + "addresses": [ + "10.42.2.14" + ] + }, + { + "name": "veth7a15", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "virtual", + "addresses": [ + "10.42.2.15" + ] + }, + { + "name": "veth7a16", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "virtual", + "addresses": [ + "10.42.2.16" + ] + }, + { + "name": "veth7a17", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "virtual", + "addresses": [ + "10.42.2.17" + ] + }, + { + "name": "veth7a18", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "virtual", + "addresses": [ + "10.42.2.18" + ] + }, + { + "name": "veth7a19", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "virtual", + "addresses": [ + "10.42.2.19" + ] + }, + { + "name": "veth7a1a", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "virtual", + "addresses": [ + "10.42.2.20" + ] + }, + { + "name": "veth7a1b", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "virtual", + "addresses": [ + "10.42.2.21" + ] + }, + { + "name": "eth0", + "speedMbps": 10000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-18 + name: sb-nodeprobe-discover-net-18-worker-01-9f20b64e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/case.md new file mode 100644 index 000000000..b9be7a60a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/case.md @@ -0,0 +1,7 @@ +# NET-19 + +**Mutation.** `cni0` and `docker0` both addressed, one physical NIC with no address + +**Expected.** `mgmtInterface` empty: a bridge is refused and an unaddressed NIC is not a candidate + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/expected-notes.txt new file mode 100644 index 000000000..a6aea2be3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-19 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/expected.yaml new file mode 100644 index 000000000..fdc602045 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/expected.yaml @@ -0,0 +1,31 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-19 + name: discovered-discover-net-19 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-19-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/nodes.yaml new file mode 100644 index 000000000..94be916b6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/nodes.yaml @@ -0,0 +1,26 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/ops.yaml new file mode 100644 index 000000000..111a1f013 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-19 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/reports/worker-01.yaml new file mode 100644 index 000000000..184d53c5e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/reports/worker-01.yaml @@ -0,0 +1,196 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "cni0", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "bridge": true, + "kind": "bridge", + "addresses": [ + "10.42.2.1" + ] + }, + { + "name": "flannel.1", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "vxlan", + "addresses": [ + "10.42.2.0" + ] + }, + { + "name": "lo", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "loopback": true, + "kind": "loopback", + "addresses": [ + "127.0.0.1" + ] + }, + { + "name": "eth0", + "speedMbps": 10000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:3b:00.0", + "numaNode": 0, + "kind": "physical" + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-19 + name: sb-nodeprobe-discover-net-19-worker-01-4da347a1 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/case.md new file mode 100644 index 000000000..ba30616a6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/case.md @@ -0,0 +1,7 @@ +# NET-20 + +**Mutation.** An interface whose state is `dormant`, and one whose state is `lowerlayerdown` + +**Expected.** Both passed over: only `up` and `unknown` are admitted + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/expected-notes.txt new file mode 100644 index 000000000..7742515ef --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-20 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/expected.yaml new file mode 100644 index 000000000..74d52cfc4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/expected.yaml @@ -0,0 +1,31 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-20 + name: discovered-discover-net-20 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-20-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/nodes.yaml new file mode 100644 index 000000000..94be916b6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/nodes.yaml @@ -0,0 +1,26 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/ops.yaml new file mode 100644 index 000000000..4b5d13ed9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-20 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/reports/worker-01.yaml new file mode 100644 index 000000000..f2ae7a132 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 10000, + "mtu": 1500, + "state": "dormant", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "192.168.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 10000, + "mtu": 1500, + "state": "lowerlayerdown", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-20 + name: sb-nodeprobe-discover-net-20-worker-01-be5715db + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/case.md new file mode 100644 index 000000000..0f5c02af9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/case.md @@ -0,0 +1,7 @@ +# NET-21 + +**Mutation.** An interface whose state is `unknown` holding a routable address + +**Expected.** Named. A driver that does not track carrier is not a reason to refuse it + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/expected-notes.txt new file mode 100644 index 000000000..7b4133b82 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-21 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/expected.yaml new file mode 100644 index 000000000..124eca09c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-21 + name: discovered-discover-net-21 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-21-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/nodes.yaml new file mode 100644 index 000000000..94be916b6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/nodes.yaml @@ -0,0 +1,26 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/ops.yaml new file mode 100644 index 000000000..9fd6f1716 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-21 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/reports/worker-01.yaml new file mode 100644 index 000000000..94554b960 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/reports/worker-01.yaml @@ -0,0 +1,163 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 10000, + "mtu": 1500, + "state": "unknown", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "192.168.10.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-21 + name: sb-nodeprobe-discover-net-21-worker-01-41de615b + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/case.md new file mode 100644 index 000000000..cdb010d02 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/case.md @@ -0,0 +1,9 @@ +# NET-22 + +**Mutation.** A run told that `ens5f0` is the management interface + +**Expected.** **Contested.** No input expresses it, so the ranking decides regardless. See §14, gap G-23 + +**Harness.** `CM` + +**Gap.** G-23. This case records what the generator does today, so that the day it changes the diff is the finding. diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/expected-notes.txt new file mode 100644 index 000000000..f5be9fb9b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-22 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/expected.yaml new file mode 100644 index 000000000..bb48b4964 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-22 + name: discovered-discover-net-22 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-22-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/nodes.yaml new file mode 100644 index 000000000..94be916b6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/nodes.yaml @@ -0,0 +1,26 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/ops.yaml new file mode 100644 index 000000000..4981c3645 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-22 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/reports/worker-01.yaml new file mode 100644 index 000000000..7bb5aee78 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "ens5f0", + "speedMbps": 10000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "192.168.10.1" + ] + }, + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-22 + name: sb-nodeprobe-discover-net-22-worker-01-e125d0a5 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/case.md new file mode 100644 index 000000000..8f29d3827 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/case.md @@ -0,0 +1,9 @@ +# NET-23 + +**Mutation.** A run told that `ens5f1` and `ens5f2` are the data NICs + +**Expected.** **Contested.** `dataInterfaces` is never written, so the draft leaves the data plane unnamed. See §14, gap G-24 + +**Harness.** `CM` + +**Gap.** G-24. This case records what the generator does today, so that the day it changes the diff is the finding. diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/expected-notes.txt new file mode 100644 index 000000000..51f26af2f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-23 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/expected.yaml new file mode 100644 index 000000000..015ada7e2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-23 + name: discovered-discover-net-23 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-23-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/ops.yaml new file mode 100644 index 000000000..0506e9a99 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-23 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/reports/worker-01.yaml new file mode 100644 index 000000000..b846adb58 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/reports/worker-01.yaml @@ -0,0 +1,187 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 10000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "ens5f1", + "speedMbps": 100000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "192.168.20.1" + ] + }, + { + "name": "ens5f2", + "speedMbps": 100000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "192.168.21.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-23 + name: sb-nodeprobe-discover-net-23-worker-01-c9d44a49 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/case.md new file mode 100644 index 000000000..669b11826 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/case.md @@ -0,0 +1,7 @@ +# NET-24 + +**Mutation.** Two physical NICs, a 10G and a 100G, both up and addressed + +**Expected.** The 100G named for management and no data NIC named at all, so the fast link is proposed for the wrong plane + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/expected-notes.txt new file mode 100644 index 000000000..13332c1fa --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-24 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/expected.yaml new file mode 100644 index 000000000..55ebd38ba --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-24 + name: discovered-discover-net-24 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-24-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: ens5f0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/nodes.yaml new file mode 100644 index 000000000..94be916b6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/nodes.yaml @@ -0,0 +1,26 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/ops.yaml new file mode 100644 index 000000000..68cb77db8 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-24 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/reports/worker-01.yaml new file mode 100644 index 000000000..b9ce6fe78 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 10000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "192.168.10.1" + ] + }, + { + "name": "ens5f0", + "speedMbps": 100000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "192.168.20.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-24 + name: sb-nodeprobe-discover-net-24-worker-01-de28a2f9 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/case.md new file mode 100644 index 000000000..bbbb5689f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/case.md @@ -0,0 +1,7 @@ +# NET-25 + +**Mutation.** Two otherwise equal 10G NICs at MTU 9000 and MTU 1500 + +**Expected.** The lower name wins. MTU is reported and not ranked on + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/expected-notes.txt new file mode 100644 index 000000000..a304c26c2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-25 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/expected.yaml new file mode 100644 index 000000000..62ace41db --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-25 + name: discovered-discover-net-25 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-25-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/nodes.yaml new file mode 100644 index 000000000..94be916b6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/nodes.yaml @@ -0,0 +1,26 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/ops.yaml new file mode 100644 index 000000000..9f0b0f5af --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-25 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/reports/worker-01.yaml new file mode 100644 index 000000000..6d0264396 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 10000, + "mtu": 9000, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "192.168.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 10000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-25 + name: sb-nodeprobe-discover-net-25-worker-01-62d5c539 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/case.md new file mode 100644 index 000000000..c18580d81 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/case.md @@ -0,0 +1,7 @@ +# NET-26 + +**Mutation.** A 100G bridge beside a 10G physical NIC, both addressed + +**Expected.** The 10G named: the kind ladder is decided before the speed + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/expected-notes.txt new file mode 100644 index 000000000..059400ffc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-26 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/expected.yaml new file mode 100644 index 000000000..aba65d5b7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-26 + name: discovered-discover-net-26 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-26-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/nodes.yaml new file mode 100644 index 000000000..94be916b6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/nodes.yaml @@ -0,0 +1,26 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/ops.yaml new file mode 100644 index 000000000..02753b160 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-26 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/reports/worker-01.yaml new file mode 100644 index 000000000..3399d4e9c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/reports/worker-01.yaml @@ -0,0 +1,192 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "br0", + "speedMbps": 100000, + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "bridge": true, + "kind": "bridge", + "lower": [ + "eth1" + ], + "addresses": [ + "192.168.1.1" + ] + }, + { + "name": "eth0", + "speedMbps": 10000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "192.168.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 100000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:3b:00.0", + "numaNode": 0, + "kind": "physical", + "upper": [ + "br0" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-26 + name: sb-nodeprobe-discover-net-26-worker-01-edb87fa8 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/case.md new file mode 100644 index 000000000..2f877a478 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/case.md @@ -0,0 +1,9 @@ +# NET-27 + +**Mutation.** An interface reporting state `up` with no link partner + +**Expected.** Named: the report drops `carrier`, so a link with no partner reads as a working one. See §14, gap G-26 + +**Harness.** `CM` + +**Gap.** G-26. This case records what the generator does today, so that the day it changes the diff is the finding. diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/expected-notes.txt new file mode 100644 index 000000000..cecf62d8d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-27 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/expected.yaml new file mode 100644 index 000000000..e445991cf --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-27 + name: discovered-discover-net-27 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-27-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/nodes.yaml new file mode 100644 index 000000000..94be916b6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/nodes.yaml @@ -0,0 +1,26 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/ops.yaml new file mode 100644 index 000000000..6d833ca62 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-27 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/reports/worker-01.yaml new file mode 100644 index 000000000..fb28ddc7b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/reports/worker-01.yaml @@ -0,0 +1,163 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 10000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "192.168.10.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-27 + name: sb-nodeprobe-discover-net-27-worker-01-7a41e52a + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/case.md new file mode 100644 index 000000000..6d6017569 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/case.md @@ -0,0 +1,7 @@ +# NET-28 + +**Mutation.** A bond reporting no speed over two 25G NICs, beside an addressed 10G NIC + +**Expected.** The bond named: an aggregate carries the sum of its members + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/expected-notes.txt new file mode 100644 index 000000000..341fca074 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-28 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/expected.yaml new file mode 100644 index 000000000..852d5b206 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-28 + name: discovered-discover-net-28 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-28-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: bond0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/nodes.yaml new file mode 100644 index 000000000..94be916b6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/nodes.yaml @@ -0,0 +1,26 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/ops.yaml new file mode 100644 index 000000000..416ee375b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-28 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/reports/worker-01.yaml new file mode 100644 index 000000000..3549a8ee0 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/reports/worker-01.yaml @@ -0,0 +1,231 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "bond0", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "bond", + "lower": [ + "eth0", + "eth1" + ], + "upper": [ + "bond0.100" + ], + "addresses": [ + "192.168.10.1" + ] + }, + { + "name": "bond0.100", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "vlan", + "lower": [ + "bond0" + ] + }, + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:3b:00.0", + "numaNode": 0, + "kind": "physical", + "upper": [ + "bond0" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:3b:00.1", + "numaNode": 0, + "kind": "physical", + "upper": [ + "bond0" + ] + }, + { + "name": "lo", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "loopback": true, + "kind": "loopback", + "addresses": [ + "127.0.0.1" + ] + }, + { + "name": "eth2", + "speedMbps": 10000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:af:00.0", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "192.168.20.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-28 + name: sb-nodeprobe-discover-net-28-worker-01-65ec0106 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/case.md new file mode 100644 index 000000000..de04f4878 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/case.md @@ -0,0 +1,7 @@ +# NET-29 + +**Mutation.** A VLAN reporting no speed over that bond, beside the same 10G NIC + +**Expected.** The VLAN named: a derived interface carries what its parent carries + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/expected-notes.txt new file mode 100644 index 000000000..8e4589e2c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-29 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/expected.yaml new file mode 100644 index 000000000..49018dc9b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-29 + name: discovered-discover-net-29 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-29-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: bond0.100 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/nodes.yaml new file mode 100644 index 000000000..94be916b6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/nodes.yaml @@ -0,0 +1,26 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/ops.yaml new file mode 100644 index 000000000..411196ac4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-29 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/reports/worker-01.yaml new file mode 100644 index 000000000..e9c296461 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/reports/worker-01.yaml @@ -0,0 +1,231 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "bond0", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "bond", + "lower": [ + "eth0", + "eth1" + ], + "upper": [ + "bond0.100" + ] + }, + { + "name": "bond0.100", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "vlan", + "lower": [ + "bond0" + ], + "addresses": [ + "192.168.10.1" + ] + }, + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:3b:00.0", + "numaNode": 0, + "kind": "physical", + "upper": [ + "bond0" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:3b:00.1", + "numaNode": 0, + "kind": "physical", + "upper": [ + "bond0" + ] + }, + { + "name": "lo", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "loopback": true, + "kind": "loopback", + "addresses": [ + "127.0.0.1" + ] + }, + { + "name": "eth2", + "speedMbps": 10000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:af:00.0", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "192.168.20.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-29 + name: sb-nodeprobe-discover-net-29-worker-01-99f6137a + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/case.md new file mode 100644 index 000000000..a671728e4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/case.md @@ -0,0 +1,7 @@ +# NET-30 + +**Mutation.** The chosen bond's members in sockets 0 and 1 + +**Expected.** The memory node is reported as unknown rather than as one of the two + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/expected-notes.txt new file mode 100644 index 000000000..f22772fcc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-30 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/expected.yaml new file mode 100644 index 000000000..6ee8feb01 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-30 + name: discovered-discover-net-30 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-30-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: bond0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/ops.yaml new file mode 100644 index 000000000..b2a10135b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-30 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/reports/worker-01.yaml new file mode 100644 index 000000000..6d1fc7868 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/reports/worker-01.yaml @@ -0,0 +1,218 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "bond0", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "bond", + "lower": [ + "eth0", + "eth1" + ], + "upper": [ + "bond0.100" + ], + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "bond0.100", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "vlan", + "lower": [ + "bond0" + ] + }, + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:3b:00.0", + "numaNode": 0, + "kind": "physical", + "upper": [ + "bond0" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:3b:00.1", + "numaNode": 1, + "kind": "physical", + "upper": [ + "bond0" + ] + }, + { + "name": "lo", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "loopback": true, + "kind": "loopback", + "addresses": [ + "127.0.0.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-30 + name: sb-nodeprobe-discover-net-30-worker-01-2ab1668e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/case.md new file mode 100644 index 000000000..3d49e1cd6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/case.md @@ -0,0 +1,7 @@ +# NET-31 + +**Mutation.** A veth holding the node's own address + +**Expected.** Nothing named: an unidentified virtual device is refused at the first rung, node address or not + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/expected-notes.txt new file mode 100644 index 000000000..ddf501103 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-31 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/expected.yaml new file mode 100644 index 000000000..4260ac8ab --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/expected.yaml @@ -0,0 +1,31 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-31 + name: discovered-discover-net-31 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-31-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/ops.yaml new file mode 100644 index 000000000..785c3cb1a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-31 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/reports/worker-01.yaml new file mode 100644 index 000000000..dea217ab1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/reports/worker-01.yaml @@ -0,0 +1,174 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "veth7a1c", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "virtual", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "lo", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "loopback": true, + "kind": "loopback", + "addresses": [ + "127.0.0.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-31 + name: sb-nodeprobe-discover-net-31-worker-01-78210588 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-32-a-stack-that-points-at-itself/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-32-a-stack-that-points-at-itself/case.md new file mode 100644 index 000000000..563400043 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-32-a-stack-that-points-at-itself/case.md @@ -0,0 +1,9 @@ +# NET-32 + +**Mutation.** A bond whose `lower` names a second bond whose `lower` names the first + +**Expected.** The bond is named and the resolution terminates + +**Harness.** `GO` + +**Note.** Driven in the discovery package, because no probe can write a report whose lower_* links form a cycle and the case is about the resolution terminating anyway. diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/case.md new file mode 100644 index 000000000..a5f0f65af --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/case.md @@ -0,0 +1,7 @@ +# NET-33 + +**Mutation.** The node's `InternalIP` on a link whose state is `down` + +**Expected.** Nothing named: the state rung is applied before the address wins, so a link carrying nothing is refused however right its address is + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/expected-notes.txt new file mode 100644 index 000000000..299cd946e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-33 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/expected.yaml new file mode 100644 index 000000000..ab38ab56d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/expected.yaml @@ -0,0 +1,31 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-33 + name: discovered-discover-net-33 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-33-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/ops.yaml new file mode 100644 index 000000000..0391e6e4f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-33 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/reports/worker-01.yaml new file mode 100644 index 000000000..d001ded5e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/reports/worker-01.yaml @@ -0,0 +1,163 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 10000, + "mtu": 1500, + "state": "down", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-33 + name: sb-nodeprobe-discover-net-33-worker-01-5e5c17fd + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/case.md new file mode 100644 index 000000000..19d49dd1f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/case.md @@ -0,0 +1,7 @@ +# NET-34 + +**Mutation.** The node's `InternalIP` on both `bond0` and `bond0.100` + +**Expected.** The first by the report's own order, which is the kernel's and is ascending by name, so two runs agree + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/expected-notes.txt new file mode 100644 index 000000000..dc02c5a77 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-34 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/expected.yaml new file mode 100644 index 000000000..92f31b367 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-34 + name: discovered-discover-net-34 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-34-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: bond0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/ops.yaml new file mode 100644 index 000000000..796003e63 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-34 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/reports/worker-01.yaml new file mode 100644 index 000000000..a0cc34195 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/reports/worker-01.yaml @@ -0,0 +1,221 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "bond0", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "bond", + "lower": [ + "eth0", + "eth1" + ], + "upper": [ + "bond0.100" + ], + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "bond0.100", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "vlan", + "lower": [ + "bond0" + ], + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:3b:00.0", + "numaNode": 0, + "kind": "physical", + "upper": [ + "bond0" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:3b:00.1", + "numaNode": 0, + "kind": "physical", + "upper": [ + "bond0" + ] + }, + { + "name": "lo", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "loopback": true, + "kind": "loopback", + "addresses": [ + "127.0.0.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-34 + name: sb-nodeprobe-discover-net-34-worker-01-01c6ea18 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/case.md new file mode 100644 index 000000000..71e6543cd --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/case.md @@ -0,0 +1,7 @@ +# NET-35 + +**Mutation.** A bond reporting 50000 Mbps of its own over two 25G members + +**Expected.** Its own reading is taken, and the members are not summed on top of it + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/expected-notes.txt new file mode 100644 index 000000000..7ecb77b6b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-35 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/expected.yaml new file mode 100644 index 000000000..71a3786f8 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-35 + name: discovered-discover-net-35 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-35-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: bond0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/nodes.yaml new file mode 100644 index 000000000..94be916b6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/nodes.yaml @@ -0,0 +1,26 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/ops.yaml new file mode 100644 index 000000000..3f3834bbe --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-35 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/reports/worker-01.yaml new file mode 100644 index 000000000..91d24b568 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/reports/worker-01.yaml @@ -0,0 +1,195 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "bond0", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "bond", + "lower": [ + "eth0", + "eth1" + ], + "upper": [ + "bond0.100" + ], + "addresses": [ + "192.168.10.1" + ] + }, + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:3b:00.0", + "numaNode": 0, + "kind": "physical", + "upper": [ + "bond0" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:3b:00.1", + "numaNode": 0, + "kind": "physical", + "upper": [ + "bond0" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-35 + name: sb-nodeprobe-discover-net-35-worker-01-7af30d58 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/case.md new file mode 100644 index 000000000..3158f20f1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/case.md @@ -0,0 +1,7 @@ +# NET-36 + +**Mutation.** A bond whose `lower` names an interface the report does not carry + +**Expected.** Named, with no members and no resolved speed: a member nothing describes contributes nothing + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/expected-notes.txt new file mode 100644 index 000000000..3c47feb21 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-36 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/expected.yaml new file mode 100644 index 000000000..282311f75 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-36 + name: discovered-discover-net-36 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-36-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: bond0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/ops.yaml new file mode 100644 index 000000000..f11792d47 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-36 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/reports/worker-01.yaml new file mode 100644 index 000000000..149efdd50 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/reports/worker-01.yaml @@ -0,0 +1,165 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "bond0", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "bond", + "lower": [ + "eth9" + ], + "addresses": [ + "10.10.10.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-36 + name: sb-nodeprobe-discover-net-36-worker-01-061df04d + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/case.md new file mode 100644 index 000000000..de6ffa768 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/case.md @@ -0,0 +1,7 @@ +# NET-37 + +**Mutation.** A `team` interface over two 25G NICs, holding the node address + +**Expected.** Named and resolved as a bond is, both being aggregates + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/expected-notes.txt new file mode 100644 index 000000000..1dfb54136 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-37 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/expected.yaml new file mode 100644 index 000000000..632a8f388 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-37 + name: discovered-discover-net-37 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-37-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: team0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/ops.yaml new file mode 100644 index 000000000..dab431c43 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-37 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/reports/worker-01.yaml new file mode 100644 index 000000000..0e39386e5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/reports/worker-01.yaml @@ -0,0 +1,192 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "team0", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "team", + "lower": [ + "eth0", + "eth1" + ], + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:3b:00.0", + "numaNode": 0, + "kind": "physical", + "upper": [ + "team0" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:3b:00.1", + "numaNode": 0, + "kind": "physical", + "upper": [ + "team0" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-37 + name: sb-nodeprobe-discover-net-37-worker-01-e285f225 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/case.md new file mode 100644 index 000000000..24bfec324 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/case.md @@ -0,0 +1,7 @@ +# NET-38 + +**Mutation.** A bridge over a bond over two NICs, the node address on the bridge + +**Expected.** Named, and the members resolve two levels down to the two NICs + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/expected-notes.txt new file mode 100644 index 000000000..7fbcdc81a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-38 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/expected.yaml new file mode 100644 index 000000000..c14638172 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-38 + name: discovered-discover-net-38 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-38-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: br0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/ops.yaml new file mode 100644 index 000000000..a5abe08ec --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-38 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/reports/worker-01.yaml new file mode 100644 index 000000000..ae22a5d6d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/reports/worker-01.yaml @@ -0,0 +1,207 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "br0", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "bridge": true, + "kind": "bridge", + "lower": [ + "bond0" + ], + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "bond0", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "bond", + "lower": [ + "eth0", + "eth1" + ], + "upper": [ + "br0" + ] + }, + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:3b:00.0", + "numaNode": 0, + "kind": "physical", + "upper": [ + "bond0" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:3b:00.1", + "numaNode": 0, + "kind": "physical", + "upper": [ + "bond0" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-38 + name: sb-nodeprobe-discover-net-38-worker-01-2888964d + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/case.md new file mode 100644 index 000000000..38ad82f73 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/case.md @@ -0,0 +1,7 @@ +# NET-39 + +**Mutation.** A `macvlan` over `eth0` holding the node address + +**Expected.** Named, with `eth0` as its member and `eth0`'s speed as its own + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/expected-notes.txt new file mode 100644 index 000000000..f471247f2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-39 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/expected.yaml new file mode 100644 index 000000000..7a2286487 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-39 + name: discovered-discover-net-39 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-39-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: macvlan0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/ops.yaml new file mode 100644 index 000000000..1331f7ba7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-39 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/reports/worker-01.yaml new file mode 100644 index 000000000..2c6117f9c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/reports/worker-01.yaml @@ -0,0 +1,178 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:3b:00.0", + "numaNode": 0, + "kind": "physical", + "upper": [ + "macvlan0" + ] + }, + { + "name": "macvlan0", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "macvlan", + "lower": [ + "eth0" + ], + "addresses": [ + "10.10.10.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-39 + name: sb-nodeprobe-discover-net-39-worker-01-921c32d7 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/case.md new file mode 100644 index 000000000..283e660d3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/case.md @@ -0,0 +1,7 @@ +# NET-40 + +**Mutation.** An interface carrying no `kind`, and nothing marking it virtual + +**Expected.** Read as physical, which is what the fields before the kind said about it + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/expected-notes.txt new file mode 100644 index 000000000..d55cd4b6b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-40 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/expected.yaml new file mode 100644 index 000000000..9e6d0a805 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-40 + name: discovered-discover-net-40 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-40-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/nodes.yaml new file mode 100644 index 000000000..94be916b6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/nodes.yaml @@ -0,0 +1,26 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/ops.yaml new file mode 100644 index 000000000..80015dd89 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-40 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/reports/worker-01.yaml new file mode 100644 index 000000000..358ac0c2e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/reports/worker-01.yaml @@ -0,0 +1,183 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 10000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "addresses": [ + "192.168.10.1" + ] + }, + { + "name": "cni0", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "bridge": true, + "addresses": [ + "10.42.2.1" + ] + }, + { + "name": "flannel.1", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "addresses": [ + "10.42.2.0" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-40 + name: sb-nodeprobe-discover-net-40-worker-01-b433623f + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/case.md new file mode 100644 index 000000000..6128b54bd --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/case.md @@ -0,0 +1,9 @@ +# NET-41 + +**Mutation.** A worker whose only fast NIC is a 2x25G bond, ranked for placement + +**Expected.** **Contested.** The memory node's fastest NIC reads as 0 Mbps: the placement reads the raw speed and the raw memory node, neither of which a bond has. See §14, gap G-27 + +**Harness.** `CM` + +**Gap.** G-27. This case records what the generator does today, so that the day it changes the diff is the finding. diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/expected-notes.txt new file mode 100644 index 000000000..ca74cac69 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-41 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/expected.yaml new file mode 100644 index 000000000..fa215ca0c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-41 + name: discovered-discover-net-41 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-41-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: bond0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/ops.yaml new file mode 100644 index 000000000..eb00a7281 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-41 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/reports/worker-01.yaml new file mode 100644 index 000000000..21e70d2b3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/reports/worker-01.yaml @@ -0,0 +1,218 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "bond0", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "bond", + "lower": [ + "eth0", + "eth1" + ], + "upper": [ + "bond0.100" + ], + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "bond0.100", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "vlan", + "lower": [ + "bond0" + ] + }, + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:3b:00.0", + "numaNode": 0, + "kind": "physical", + "upper": [ + "bond0" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:3b:00.1", + "numaNode": 0, + "kind": "physical", + "upper": [ + "bond0" + ] + }, + { + "name": "lo", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "loopback": true, + "kind": "loopback", + "addresses": [ + "127.0.0.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-41 + name: sb-nodeprobe-discover-net-41-worker-01-cc01f736 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/case.md new file mode 100644 index 000000000..7ad683aa8 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/case.md @@ -0,0 +1,9 @@ +# NET-42 + +**Mutation.** Any successful run on a bonded host + +**Expected.** **Contested.** The draft names the interface and says nothing about the slots, the speed, or the memory node resolved under it. See §14, gap G-28 + +**Harness.** `CM` + +**Gap.** G-28. This case records what the generator does today, so that the day it changes the diff is the finding. diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/expected-notes.txt new file mode 100644 index 000000000..ba78ccf0c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-net-42 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/expected.yaml new file mode 100644 index 000000000..3c59303fa --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-net-42 + name: discovered-discover-net-42 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-net-42-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: bond0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/ops.yaml new file mode 100644 index 000000000..06d818e8f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-net-42 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/reports/worker-01.yaml new file mode 100644 index 000000000..19c8a8d45 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/reports/worker-01.yaml @@ -0,0 +1,218 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "bond0", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "bond", + "lower": [ + "eth0", + "eth1" + ], + "upper": [ + "bond0.100" + ], + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "bond0.100", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "kind": "vlan", + "lower": [ + "bond0" + ] + }, + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:3b:00.0", + "numaNode": 0, + "kind": "physical", + "upper": [ + "bond0" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "pciAddress": "0000:3b:00.1", + "numaNode": 0, + "kind": "physical", + "upper": [ + "bond0" + ] + }, + { + "name": "lo", + "mtu": 1500, + "state": "up", + "numaNode": -1, + "virtual": true, + "loopback": true, + "kind": "loopback", + "addresses": [ + "127.0.0.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-net-42 + name: sb-nodeprobe-discover-net-42-worker-01-2f9b039b + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/case.md b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/case.md new file mode 100644 index 000000000..49aa20880 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/case.md @@ -0,0 +1,7 @@ +# NUMA-01 + +**Mutation.** 1 memory node, 4 disks on it + +**Expected.** All 4 used, reason "every unclaimed device is on NUMA node 0" + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/expected-notes.txt new file mode 100644 index 000000000..45d7ba75e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-numa-01 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/expected.yaml new file mode 100644 index 000000000..de45da834 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-numa-01 + name: discovered-discover-numa-01 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-numa-01-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/ops.yaml new file mode 100644 index 000000000..bbc24f53f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-numa-01 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/reports/worker-01.yaml new file mode 100644 index 000000000..72542bc71 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-numa-01 + name: sb-nodeprobe-discover-numa-01-worker-01-381d5b88 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/case.md b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/case.md new file mode 100644 index 000000000..049beaf22 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/case.md @@ -0,0 +1,7 @@ +# NUMA-02 + +**Mutation.** 2 memory nodes, 2 disks each, equal size, equal cores + +**Expected.** Node 0 chosen on the id tie-break. 2 disks refused by the placement + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/expected-notes.txt new file mode 100644 index 000000000..79d7e6860 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 8, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-numa-02 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 2 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/expected-refusals.txt new file mode 100644 index 000000000..e5d0ed2ac --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/expected-refusals.txt @@ -0,0 +1,2 @@ +worker-01/nvme2n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 1, so it was chosen +worker-01/nvme3n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 1, so it was chosen diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/expected.yaml new file mode 100644 index 000000000..52e4e407d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/expected.yaml @@ -0,0 +1,30 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-numa-02 + name: discovered-discover-numa-02 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-numa-02-cluster + vcpuCount: 8 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + mgmtInterface: eth0 + name: group-1-nvme-2x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/ops.yaml new file mode 100644 index 000000000..5e32c3917 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-numa-02 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/reports/worker-01.yaml new file mode 100644 index 000000000..02801fa5f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/reports/worker-01.yaml @@ -0,0 +1,181 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 2, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15 + ], + "physicalCores": 8 + }, + { + "node": 1, + "onlineCPUs": [ + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 8 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:7e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:7e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-numa-02 + name: sb-nodeprobe-discover-numa-02-worker-01-1574d813 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/case.md b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/case.md new file mode 100644 index 000000000..0d068e151 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/case.md @@ -0,0 +1,7 @@ +# NUMA-03 + +**Mutation.** 2 memory nodes, 1 disk on node 0 and 3 on node 1 + +**Expected.** Node 1 chosen on count. The single disk refused + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/expected-notes.txt new file mode 100644 index 000000000..edffb2773 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 8, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-numa-03 in Draft: 1 workers with 3 nvme devices, in 1 group(s) across 1 node set(s); 1 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/expected-refusals.txt new file mode 100644 index 000000000..c0e86410f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/expected-refusals.txt @@ -0,0 +1 @@ +worker-01/nvme0n1: declined by most available NUMA node because NUMA node 1 carries 3 unclaimed devices (9T) against 1 (3T) on NUMA node 0, so it was chosen diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/expected.yaml new file mode 100644 index 000000000..39132b0e5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/expected.yaml @@ -0,0 +1,31 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-numa-03 + name: discovered-discover-numa-03 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-numa-03-cluster + vcpuCount: 8 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:7e:00.0 + - 0000:7e:00.1 + - 0000:7e:00.2 + mgmtInterface: eth0 + name: group-1-nvme-3x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/ops.yaml new file mode 100644 index 000000000..f31376c01 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-numa-03 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/reports/worker-01.yaml new file mode 100644 index 000000000..5977ffbdc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/reports/worker-01.yaml @@ -0,0 +1,181 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 2, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15 + ], + "physicalCores": 8 + }, + { + "node": 1, + "onlineCPUs": [ + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 8 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:7e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:7e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:7e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-numa-03 + name: sb-nodeprobe-discover-numa-03-worker-01-3bafa483 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/case.md b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/case.md new file mode 100644 index 000000000..29e34fffd --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/case.md @@ -0,0 +1,7 @@ +# NUMA-04 + +**Mutation.** 2 memory nodes, 2 × 8 TiB on node 0 and 3 × 1 TiB on node 1 + +**Expected.** Node 1 chosen: count beats capacity + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/expected-notes.txt new file mode 100644 index 000000000..e6abb4fc5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 8, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-numa-04 in Draft: 1 workers with 3 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/expected.yaml new file mode 100644 index 000000000..749e830c3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/expected.yaml @@ -0,0 +1,31 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-numa-04 + name: discovered-discover-numa-04 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-numa-04-cluster + vcpuCount: 8 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:7e:00.0 + - 0000:7e:00.1 + - 0000:7e:00.2 + mgmtInterface: eth0 + name: group-1-nvme-3x1T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/ops.yaml new file mode 100644 index 000000000..28b625110 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-numa-04 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/reports/worker-01.yaml new file mode 100644 index 000000000..e22e7006f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/reports/worker-01.yaml @@ -0,0 +1,193 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 2, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15 + ], + "physicalCores": 8 + }, + { + "node": 1, + "onlineCPUs": [ + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 8 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 8796093022208, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 8796093022208, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:7e:00.0", + "sizeBytes": 1099511627776, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:7e:00.1", + "sizeBytes": 1099511627776, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:7e:00.2", + "sizeBytes": 1099511627776, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-numa-04 + name: sb-nodeprobe-discover-numa-04-worker-01-f6b1662a + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/case.md b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/case.md new file mode 100644 index 000000000..71fc3c023 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/case.md @@ -0,0 +1,7 @@ +# NUMA-05 + +**Mutation.** 2 memory nodes, 2 disks each, 1 TiB on node 0 and 2 TiB on node 1 + +**Expected.** Node 1 chosen on the capacity tie-break + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/expected-notes.txt new file mode 100644 index 000000000..81a10fa1e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 8, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-numa-05 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/expected.yaml new file mode 100644 index 000000000..bcda41e81 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/expected.yaml @@ -0,0 +1,30 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-numa-05 + name: discovered-discover-numa-05 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-numa-05-cluster + vcpuCount: 8 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:7e:00.0 + - 0000:7e:00.1 + mgmtInterface: eth0 + name: group-1-nvme-2x2T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/ops.yaml new file mode 100644 index 000000000..1ca48510b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-numa-05 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/reports/worker-01.yaml new file mode 100644 index 000000000..a9aacf760 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/reports/worker-01.yaml @@ -0,0 +1,181 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 2, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15 + ], + "physicalCores": 8 + }, + { + "node": 1, + "onlineCPUs": [ + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 8 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 1099511627776, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 1099511627776, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:7e:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:7e:00.1", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-numa-05 + name: sb-nodeprobe-discover-numa-05-worker-01-7a53cdb4 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/case.md b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/case.md new file mode 100644 index 000000000..ba84c7a00 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/case.md @@ -0,0 +1,7 @@ +# NUMA-06 + +**Mutation.** 2 memory nodes, 2 disks each of one size, 8 cores on node 0 and 24 on node 1 + +**Expected.** Node 1 chosen on the core tie-break + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/expected-notes.txt new file mode 100644 index 000000000..4341f3e39 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 24, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-numa-06 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 2 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/expected-refusals.txt new file mode 100644 index 000000000..876f779a7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/expected-refusals.txt @@ -0,0 +1,2 @@ +worker-01/nvme0n1: declined by most available NUMA node because NUMA node 1 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 0, so it was chosen +worker-01/nvme1n1: declined by most available NUMA node because NUMA node 1 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 0, so it was chosen diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/expected.yaml new file mode 100644 index 000000000..57a6816ba --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/expected.yaml @@ -0,0 +1,30 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-numa-06 + name: discovered-discover-numa-06 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-numa-06-cluster + vcpuCount: 24 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:7e:00.0 + - 0000:7e:00.1 + mgmtInterface: eth0 + name: group-1-nvme-2x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/ops.yaml new file mode 100644 index 000000000..eeea4be58 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-numa-06 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/reports/worker-01.yaml new file mode 100644 index 000000000..4b1a10dcd --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/reports/worker-01.yaml @@ -0,0 +1,213 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 64, + "physicalCores": 32, + "sockets": 2, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15 + ], + "physicalCores": 8 + }, + { + "node": 1, + "onlineCPUs": [ + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31, + 32, + 33, + 34, + 35, + 36, + 37, + 38, + 39, + 40, + 41, + 42, + 43, + 44, + 45, + 46, + 47, + 48, + 49, + 50, + 51, + 52, + 53, + 54, + 55, + 56, + 57, + 58, + 59, + 60, + 61, + 62, + 63 + ], + "physicalCores": 24 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:7e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:7e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-numa-06 + name: sb-nodeprobe-discover-numa-06-worker-01-993de559 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/case.md b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/case.md new file mode 100644 index 000000000..27665e3c6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/case.md @@ -0,0 +1,7 @@ +# NUMA-07 + +**Mutation.** 4 memory nodes, 2 disks each (NPS4) + +**Expected.** Node 0 chosen. 6 disks refused. `vcpuCount` is node 0's physical cores + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/expected-notes.txt new file mode 100644 index 000000000..40d07e240 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-numa-07 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 6 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/expected-refusals.txt new file mode 100644 index 000000000..2dd22b9fc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/expected-refusals.txt @@ -0,0 +1,6 @@ +worker-01/nvme2n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 1, so it was chosen +worker-01/nvme3n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 1, so it was chosen +worker-01/nvme4n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 1, so it was chosen +worker-01/nvme5n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 1, so it was chosen +worker-01/nvme6n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 1, so it was chosen +worker-01/nvme7n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 1, so it was chosen diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/expected.yaml new file mode 100644 index 000000000..d1f1a8982 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/expected.yaml @@ -0,0 +1,30 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-numa-07 + name: discovered-discover-numa-07 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-numa-07-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + mgmtInterface: eth0 + name: group-1-nvme-2x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/ops.yaml new file mode 100644 index 000000000..2a863f52d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-numa-07 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/reports/worker-01.yaml new file mode 100644 index 000000000..ca09bd802 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/reports/worker-01.yaml @@ -0,0 +1,337 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 128, + "physicalCores": 64, + "sockets": 4, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + }, + { + "node": 1, + "onlineCPUs": [ + 32, + 33, + 34, + 35, + 36, + 37, + 38, + 39, + 40, + 41, + 42, + 43, + 44, + 45, + 46, + 47, + 48, + 49, + 50, + 51, + 52, + 53, + 54, + 55, + 56, + 57, + 58, + 59, + 60, + 61, + 62, + 63 + ], + "physicalCores": 16 + }, + { + "node": 2, + "onlineCPUs": [ + 64, + 65, + 66, + 67, + 68, + 69, + 70, + 71, + 72, + 73, + 74, + 75, + 76, + 77, + 78, + 79, + 80, + 81, + 82, + 83, + 84, + 85, + 86, + 87, + 88, + 89, + 90, + 91, + 92, + 93, + 94, + 95 + ], + "physicalCores": 16 + }, + { + "node": 3, + "onlineCPUs": [ + 96, + 97, + 98, + 99, + 100, + 101, + 102, + 103, + 104, + 105, + 106, + 107, + 108, + 109, + 110, + 111, + 112, + 113, + 114, + 115, + 116, + 117, + 118, + 119, + 120, + 121, + 122, + 123, + 124, + 125, + 126, + 127 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:7e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:7e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme4n1", + "path": "/dev/nvme4n1", + "pciAddress": "0000:9e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 2, + "available": true, + "content": "Blank" + }, + { + "name": "nvme5n1", + "path": "/dev/nvme5n1", + "pciAddress": "0000:9e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 2, + "available": true, + "content": "Blank" + }, + { + "name": "nvme6n1", + "path": "/dev/nvme6n1", + "pciAddress": "0000:be:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 3, + "available": true, + "content": "Blank" + }, + { + "name": "nvme7n1", + "path": "/dev/nvme7n1", + "pciAddress": "0000:be:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 3, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-numa-07 + name: sb-nodeprobe-discover-numa-07-worker-01-f1b2c482 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/case.md b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/case.md new file mode 100644 index 000000000..61cd3c244 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/case.md @@ -0,0 +1,7 @@ +# NUMA-08 + +**Mutation.** 8 memory nodes, 1 disk each but 3 on node 5 + +**Expected.** Node 5 chosen on count + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/expected-notes.txt new file mode 100644 index 000000000..b5429a99a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/expected-notes.txt @@ -0,0 +1,4 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 8, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +wrote discovered-discover-numa-08 in Draft: 1 workers with 3 nvme devices, in 1 group(s) across 1 node set(s); 7 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/expected-refusals.txt new file mode 100644 index 000000000..5a6fa0f99 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/expected-refusals.txt @@ -0,0 +1,7 @@ +worker-01/nvme0n1: declined by most available NUMA node because NUMA node 5 carries 3 unclaimed devices (9T) against 1 (3T) on NUMA node 0, so it was chosen +worker-01/nvme1n1: declined by most available NUMA node because NUMA node 5 carries 3 unclaimed devices (9T) against 1 (3T) on NUMA node 0, so it was chosen +worker-01/nvme2n1: declined by most available NUMA node because NUMA node 5 carries 3 unclaimed devices (9T) against 1 (3T) on NUMA node 0, so it was chosen +worker-01/nvme3n1: declined by most available NUMA node because NUMA node 5 carries 3 unclaimed devices (9T) against 1 (3T) on NUMA node 0, so it was chosen +worker-01/nvme4n1: declined by most available NUMA node because NUMA node 5 carries 3 unclaimed devices (9T) against 1 (3T) on NUMA node 0, so it was chosen +worker-01/nvme8n1: declined by most available NUMA node because NUMA node 5 carries 3 unclaimed devices (9T) against 1 (3T) on NUMA node 0, so it was chosen +worker-01/nvme9n1: declined by most available NUMA node because NUMA node 5 carries 3 unclaimed devices (9T) against 1 (3T) on NUMA node 0, so it was chosen diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/expected.yaml new file mode 100644 index 000000000..1ed6466bd --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/expected.yaml @@ -0,0 +1,30 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-numa-08 + name: discovered-discover-numa-08 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + name: discovered-discover-numa-08-cluster + vcpuCount: 8 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:fe:00.0 + - 0000:fe:00.1 + - 0000:fe:00.2 + mgmtInterface: eth0 + name: group-1-nvme-3x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/ops.yaml new file mode 100644 index 000000000..953005cff --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-numa-08 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/reports/worker-01.yaml new file mode 100644 index 000000000..7e19b748c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/reports/worker-01.yaml @@ -0,0 +1,385 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 128, + "physicalCores": 64, + "sockets": 8, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15 + ], + "physicalCores": 8 + }, + { + "node": 1, + "onlineCPUs": [ + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 8 + }, + { + "node": 2, + "onlineCPUs": [ + 32, + 33, + 34, + 35, + 36, + 37, + 38, + 39, + 40, + 41, + 42, + 43, + 44, + 45, + 46, + 47 + ], + "physicalCores": 8 + }, + { + "node": 3, + "onlineCPUs": [ + 48, + 49, + 50, + 51, + 52, + 53, + 54, + 55, + 56, + 57, + 58, + 59, + 60, + 61, + 62, + 63 + ], + "physicalCores": 8 + }, + { + "node": 4, + "onlineCPUs": [ + 64, + 65, + 66, + 67, + 68, + 69, + 70, + 71, + 72, + 73, + 74, + 75, + 76, + 77, + 78, + 79 + ], + "physicalCores": 8 + }, + { + "node": 5, + "onlineCPUs": [ + 80, + 81, + 82, + 83, + 84, + 85, + 86, + 87, + 88, + 89, + 90, + 91, + 92, + 93, + 94, + 95 + ], + "physicalCores": 8 + }, + { + "node": 6, + "onlineCPUs": [ + 96, + 97, + 98, + 99, + 100, + 101, + 102, + 103, + 104, + 105, + 106, + 107, + 108, + 109, + 110, + 111 + ], + "physicalCores": 8 + }, + { + "node": 7, + "onlineCPUs": [ + 112, + 113, + 114, + 115, + 116, + 117, + 118, + 119, + 120, + 121, + 122, + 123, + 124, + 125, + 126, + 127 + ], + "physicalCores": 8 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:7e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:9e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 2, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:be:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 3, + "available": true, + "content": "Blank" + }, + { + "name": "nvme4n1", + "path": "/dev/nvme4n1", + "pciAddress": "0000:de:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 4, + "available": true, + "content": "Blank" + }, + { + "name": "nvme5n1", + "path": "/dev/nvme5n1", + "pciAddress": "0000:fe:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 5, + "available": true, + "content": "Blank" + }, + { + "name": "nvme6n1", + "path": "/dev/nvme6n1", + "pciAddress": "0000:fe:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 5, + "available": true, + "content": "Blank" + }, + { + "name": "nvme7n1", + "path": "/dev/nvme7n1", + "pciAddress": "0000:fe:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 5, + "available": true, + "content": "Blank" + }, + { + "name": "nvme8n1", + "path": "/dev/nvme8n1", + "pciAddress": "0000:11e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 6, + "available": true, + "content": "Blank" + }, + { + "name": "nvme9n1", + "path": "/dev/nvme9n1", + "pciAddress": "0000:13e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 7, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-numa-08 + name: sb-nodeprobe-discover-numa-08-worker-01-d2ca2f09 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/case.md b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/case.md new file mode 100644 index 000000000..ee91ba36a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/case.md @@ -0,0 +1,7 @@ +# NUMA-09 + +**Mutation.** Every device reports `numaNode: -1` + +**Expected.** The unknown bucket is used. `vcpuCount` floors at 4, `minHugePagesSize` unset, both noted + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/expected-notes.txt new file mode 100644 index 000000000..b1bc9fad7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/expected-notes.txt @@ -0,0 +1,4 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is the API's minimum of 4, because no worker reported the cores of the memory node it was placed on +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +wrote discovered-discover-numa-09 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/expected.yaml new file mode 100644 index 000000000..4e1881b14 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/expected.yaml @@ -0,0 +1,31 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-numa-09 + name: discovered-discover-numa-09 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + name: discovered-discover-numa-09-cluster + vcpuCount: 4 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/ops.yaml new file mode 100644 index 000000000..99e8334dd --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-numa-09 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/reports/worker-01.yaml new file mode 100644 index 000000000..67c6ac496 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/reports/worker-01.yaml @@ -0,0 +1,181 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 2, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15 + ], + "physicalCores": 8 + }, + { + "node": 1, + "onlineCPUs": [ + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 8 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": -1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": -1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": -1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": -1, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-numa-09 + name: sb-nodeprobe-discover-numa-09-worker-01-450940ef + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/case.md b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/case.md new file mode 100644 index 000000000..e35e86fec --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/case.md @@ -0,0 +1,7 @@ +# NUMA-10 + +**Mutation.** 2 disks on node 0 and 2 on `-1`, node 0 carrying cores + +**Expected.** Node 0 chosen on cores. The unknown bucket is never preferred + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/expected-notes.txt new file mode 100644 index 000000000..cd01cfc5d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 8, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-numa-10 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 2 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/expected-refusals.txt new file mode 100644 index 000000000..02f9278c1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/expected-refusals.txt @@ -0,0 +1,2 @@ +worker-01/nvme8n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on no memory node in particular, so it was chosen +worker-01/nvme9n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on no memory node in particular, so it was chosen diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/expected.yaml new file mode 100644 index 000000000..89ebc347e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/expected.yaml @@ -0,0 +1,30 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-numa-10 + name: discovered-discover-numa-10 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-numa-10-cluster + vcpuCount: 8 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + mgmtInterface: eth0 + name: group-1-nvme-2x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/ops.yaml new file mode 100644 index 000000000..1604ce68d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-numa-10 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/reports/worker-01.yaml new file mode 100644 index 000000000..206f12698 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/reports/worker-01.yaml @@ -0,0 +1,181 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 2, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15 + ], + "physicalCores": 8 + }, + { + "node": 1, + "onlineCPUs": [ + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 8 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme8n1", + "path": "/dev/nvme8n1", + "pciAddress": "0000:c0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": -1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme9n1", + "path": "/dev/nvme9n1", + "pciAddress": "0000:c1:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": -1, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-numa-10 + name: sb-nodeprobe-discover-numa-10-worker-01-07082c6d + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-11-all-devices-placement/case.md b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-11-all-devices-placement/case.md new file mode 100644 index 000000000..2181006fe --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-11-all-devices-placement/case.md @@ -0,0 +1,9 @@ +# NUMA-11 + +**Mutation.** The same worker with `Placement: AllDevices` + +**Expected.** Every disk used. No single chosen node, so `vcpuCount` floors and huge pages go unset + +**Harness.** `GO` + +**Note.** Driven in the discovery package with Planner{Placement: AllDevices{}} over the NUMA-02 fleet, because the controller always builds the default Planner. diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/case.md b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/case.md new file mode 100644 index 000000000..0514ad407 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/case.md @@ -0,0 +1,9 @@ +# NUMA-12 + +**Mutation.** A 1-node worker and a 2-node worker whose chosen addresses coincide + +**Expected.** One group: grouping reads addresses and the interface, never the topology + +**Harness.** `CM` + +**Note.** Both workers hand over the same two addresses, and the grouper reads the addresses rather than the topology, so they share a group. diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/expected-notes.txt new file mode 100644 index 000000000..2bafa25f1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 8, the cores of the memory node chosen on worker-02, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-numa-12 in Draft: 2 workers with 3 nvme devices, in 2 group(s) across 1 node set(s); 1 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/expected-refusals.txt new file mode 100644 index 000000000..bbca70d31 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/expected-refusals.txt @@ -0,0 +1 @@ +worker-02/nvme1n1: declined by most available NUMA node because NUMA node 0 carries 1 unclaimed devices (3T) against 1 (3T) on NUMA node 1, so it was chosen diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/expected.yaml new file mode 100644 index 000000000..3545852eb --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/expected.yaml @@ -0,0 +1,37 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-numa-12 + name: discovered-discover-numa-12 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-numa-12-cluster + vcpuCount: 8 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + mgmtInterface: eth0 + name: group-1-nvme-2x3T + workers: + - worker-01 + - devices: + nvme: + - 0000:5e:00.0 + mgmtInterface: eth0 + name: group-2-nvme-1x3T + workers: + - worker-02 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/nodes.yaml new file mode 100644 index 000000000..1fba89979 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/nodes.yaml @@ -0,0 +1,59 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/ops.yaml new file mode 100644 index 000000000..2dc88fefc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/ops.yaml @@ -0,0 +1,13 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-numa-12 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/reports/worker-01.yaml new file mode 100644 index 000000000..893c8ab1e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/reports/worker-01.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-numa-12 + name: sb-nodeprobe-discover-numa-12-worker-01-19e26185 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/reports/worker-02.yaml new file mode 100644 index 000000000..da1d35141 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/reports/worker-02.yaml @@ -0,0 +1,157 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 2, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15 + ], + "physicalCores": 8 + }, + { + "node": 1, + "onlineCPUs": [ + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 8 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-numa-12 + name: sb-nodeprobe-discover-numa-12-worker-02-2561aa88 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/case.md b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/case.md new file mode 100644 index 000000000..f104e7d24 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/case.md @@ -0,0 +1,7 @@ +# NUMA-13 + +**Mutation.** 2 memory nodes, all 10 disks on node 1 and none on node 0 + +**Expected.** Node 1 chosen with nothing to choose. The CPU note names node 1's cores + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/expected-notes.txt new file mode 100644 index 000000000..83949a031 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 8, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-numa-13 in Draft: 1 workers with 10 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/expected.yaml new file mode 100644 index 000000000..934472d8f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/expected.yaml @@ -0,0 +1,38 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-numa-13 + name: discovered-discover-numa-13 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-numa-13-cluster + vcpuCount: 8 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:7e:00.0 + - 0000:7e:00.1 + - 0000:7e:00.2 + - 0000:7e:00.3 + - 0000:7e:00.4 + - 0000:7e:00.5 + - 0000:7e:00.6 + - 0000:7e:00.7 + - 0000:7e:00.8 + - 0000:7e:00.9 + mgmtInterface: eth0 + name: group-1-nvme-10x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/ops.yaml new file mode 100644 index 000000000..7ad4fdde4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-numa-13 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/reports/worker-01.yaml new file mode 100644 index 000000000..069af4746 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/reports/worker-01.yaml @@ -0,0 +1,253 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 2, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15 + ], + "physicalCores": 8 + }, + { + "node": 1, + "onlineCPUs": [ + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 8 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:7e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:7e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:7e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:7e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme4n1", + "path": "/dev/nvme4n1", + "pciAddress": "0000:7e:00.4", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme5n1", + "path": "/dev/nvme5n1", + "pciAddress": "0000:7e:00.5", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme6n1", + "path": "/dev/nvme6n1", + "pciAddress": "0000:7e:00.6", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme7n1", + "path": "/dev/nvme7n1", + "pciAddress": "0000:7e:00.7", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme8n1", + "path": "/dev/nvme8n1", + "pciAddress": "0000:7e:00.8", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme9n1", + "path": "/dev/nvme9n1", + "pciAddress": "0000:7e:00.9", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-numa-13 + name: sb-nodeprobe-discover-numa-13-worker-01-b92c52b4 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/case.md b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/case.md new file mode 100644 index 000000000..d24dbebf1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/case.md @@ -0,0 +1,7 @@ +# NUMA-14 + +**Mutation.** A large worker: 2 nodes × 49 cores, 5 disks each, huge pages on both + +**Expected.** 5 disks in the draft, `vcpuCount` 49, `minHugePagesSize` from node 0's free pages alone + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/expected-notes.txt new file mode 100644 index 000000000..b1758ea7e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 49, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 256G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-numa-14 in Draft: 1 workers with 5 nvme devices, in 1 group(s) across 1 node set(s); 5 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/expected-refusals.txt new file mode 100644 index 000000000..e88cd23a6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/expected-refusals.txt @@ -0,0 +1,5 @@ +worker-01/nvme5n1: declined by most available NUMA node because NUMA node 0 carries 5 unclaimed devices (15T) against 5 (15T) on NUMA node 1, so it was chosen +worker-01/nvme6n1: declined by most available NUMA node because NUMA node 0 carries 5 unclaimed devices (15T) against 5 (15T) on NUMA node 1, so it was chosen +worker-01/nvme7n1: declined by most available NUMA node because NUMA node 0 carries 5 unclaimed devices (15T) against 5 (15T) on NUMA node 1, so it was chosen +worker-01/nvme8n1: declined by most available NUMA node because NUMA node 0 carries 5 unclaimed devices (15T) against 5 (15T) on NUMA node 1, so it was chosen +worker-01/nvme9n1: declined by most available NUMA node because NUMA node 0 carries 5 unclaimed devices (15T) against 5 (15T) on NUMA node 1, so it was chosen diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/expected.yaml new file mode 100644 index 000000000..3e0a904a7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/expected.yaml @@ -0,0 +1,33 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-numa-14 + name: discovered-discover-numa-14 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 256G + name: discovered-discover-numa-14-cluster + vcpuCount: 49 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + - 0000:5e:00.4 + mgmtInterface: eth0 + name: group-1-nvme-5x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/ops.yaml new file mode 100644 index 000000000..f32c6cd9a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-numa-14 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/reports/worker-01.yaml new file mode 100644 index 000000000..dd3a61809 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/reports/worker-01.yaml @@ -0,0 +1,417 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 196, + "physicalCores": 98, + "sockets": 2, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31, + 32, + 33, + 34, + 35, + 36, + 37, + 38, + 39, + 40, + 41, + 42, + 43, + 44, + 45, + 46, + 47, + 48, + 49, + 50, + 51, + 52, + 53, + 54, + 55, + 56, + 57, + 58, + 59, + 60, + 61, + 62, + 63, + 64, + 65, + 66, + 67, + 68, + 69, + 70, + 71, + 72, + 73, + 74, + 75, + 76, + 77, + 78, + 79, + 80, + 81, + 82, + 83, + 84, + 85, + 86, + 87, + 88, + 89, + 90, + 91, + 92, + 93, + 94, + 95, + 96, + 97 + ], + "physicalCores": 49 + }, + { + "node": 1, + "onlineCPUs": [ + 98, + 99, + 100, + 101, + 102, + 103, + 104, + 105, + 106, + 107, + 108, + 109, + 110, + 111, + 112, + 113, + 114, + 115, + 116, + 117, + 118, + 119, + 120, + 121, + 122, + 123, + 124, + 125, + 126, + 127, + 128, + 129, + 130, + 131, + 132, + 133, + 134, + 135, + 136, + 137, + 138, + 139, + 140, + 141, + 142, + 143, + 144, + 145, + 146, + 147, + 148, + 149, + 150, + 151, + 152, + 153, + 154, + 155, + 156, + 157, + 158, + 159, + 160, + 161, + 162, + 163, + 164, + 165, + 166, + 167, + 168, + 169, + 170, + 171, + 172, + 173, + 174, + 175, + 176, + 177, + 178, + 179, + 180, + 181, + 182, + 183, + 184, + 185, + 186, + 187, + 188, + 189, + 190, + 191, + 192, + 193, + 194, + 195 + ], + "physicalCores": 49 + } + ] + }, + "memory": { + "totalBytes": 1099511627776, + "freeBytes": 268435456000, + "availableBytes": 1073741824000, + "numaNodes": [ + { + "node": 0, + "totalBytes": 549755813888, + "freeBytes": 274877906944 + }, + { + "node": 1, + "totalBytes": 549755813888, + "freeBytes": 274877906944 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 512, + "free": 512, + "numaNodes": [ + { + "node": 0, + "total": 256, + "free": 256 + }, + { + "node": 1, + "total": 256, + "free": 256 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme4n1", + "path": "/dev/nvme4n1", + "pciAddress": "0000:5e:00.4", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme5n1", + "path": "/dev/nvme5n1", + "pciAddress": "0000:7e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme6n1", + "path": "/dev/nvme6n1", + "pciAddress": "0000:7e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme7n1", + "path": "/dev/nvme7n1", + "pciAddress": "0000:7e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme8n1", + "path": "/dev/nvme8n1", + "pciAddress": "0000:7e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme9n1", + "path": "/dev/nvme9n1", + "pciAddress": "0000:7e:00.4", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-numa-14 + name: sb-nodeprobe-discover-numa-14-worker-01-d472dee0 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/case.md b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/case.md new file mode 100644 index 000000000..83d77fc65 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/case.md @@ -0,0 +1,7 @@ +# NUMA-15 + +**Mutation.** `cpu.numaNodes` empty while devices report nodes 0 and 1 + +**Expected.** Placement falls through to the id tie-break. `vcpuCount` floors at 4 with its note + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/expected-notes.txt new file mode 100644 index 000000000..2ca82468b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is the API's minimum of 4, because no worker reported the cores of the memory node it was placed on +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-numa-15 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 2 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/expected-refusals.txt new file mode 100644 index 000000000..e5d0ed2ac --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/expected-refusals.txt @@ -0,0 +1,2 @@ +worker-01/nvme2n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 1, so it was chosen +worker-01/nvme3n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 1, so it was chosen diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/expected.yaml new file mode 100644 index 000000000..19709d839 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/expected.yaml @@ -0,0 +1,30 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-numa-15 + name: discovered-discover-numa-15 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-numa-15-cluster + vcpuCount: 4 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + mgmtInterface: eth0 + name: group-1-nvme-2x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/ops.yaml new file mode 100644 index 000000000..d4144a59a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-numa-15 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/reports/worker-01.yaml new file mode 100644 index 000000000..979a50e93 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/reports/worker-01.yaml @@ -0,0 +1,135 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 2, + "threadsPerCore": 2, + "hyperThreading": true + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:7e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:7e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-numa-15 + name: sb-nodeprobe-discover-numa-15-worker-01-da726d44 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/case.md b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/case.md new file mode 100644 index 000000000..1e4c98518 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/case.md @@ -0,0 +1,7 @@ +# PCI-01 + +**Mutation.** 32 workers, identical addresses + +**Expected.** 1 group of 32 workers + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/expected-notes.txt new file mode 100644 index 000000000..975ea8ee5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-pci-01 in Draft: 32 workers with 128 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/expected.yaml new file mode 100644 index 000000000..0ec964637 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/expected.yaml @@ -0,0 +1,63 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: discovered-discover-pci-01 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-pci-01-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - worker-02 + - worker-03 + - worker-04 + - worker-05 + - worker-06 + - worker-07 + - worker-08 + - worker-09 + - worker-10 + - worker-11 + - worker-12 + - worker-13 + - worker-14 + - worker-15 + - worker-16 + - worker-17 + - worker-18 + - worker-19 + - worker-20 + - worker-21 + - worker-22 + - worker-23 + - worker-24 + - worker-25 + - worker-26 + - worker-27 + - worker-28 + - worker-29 + - worker-30 + - worker-31 + - worker-32 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/nodes.yaml new file mode 100644 index 000000000..c25e3f497 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/nodes.yaml @@ -0,0 +1,959 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-04 +spec: {} +status: + addresses: + - address: 10.10.10.4 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-05 +spec: {} +status: + addresses: + - address: 10.10.10.5 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-06 +spec: {} +status: + addresses: + - address: 10.10.10.6 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-07 +spec: {} +status: + addresses: + - address: 10.10.10.7 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-08 +spec: {} +status: + addresses: + - address: 10.10.10.8 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-09 +spec: {} +status: + addresses: + - address: 10.10.10.9 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-10 +spec: {} +status: + addresses: + - address: 10.10.10.10 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-11 +spec: {} +status: + addresses: + - address: 10.10.10.11 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-12 +spec: {} +status: + addresses: + - address: 10.10.10.12 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-13 +spec: {} +status: + addresses: + - address: 10.10.10.13 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-14 +spec: {} +status: + addresses: + - address: 10.10.10.14 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-15 +spec: {} +status: + addresses: + - address: 10.10.10.15 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-16 +spec: {} +status: + addresses: + - address: 10.10.10.16 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-17 +spec: {} +status: + addresses: + - address: 10.10.10.17 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-18 +spec: {} +status: + addresses: + - address: 10.10.10.18 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-19 +spec: {} +status: + addresses: + - address: 10.10.10.19 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-20 +spec: {} +status: + addresses: + - address: 10.10.10.20 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-21 +spec: {} +status: + addresses: + - address: 10.10.10.21 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-22 +spec: {} +status: + addresses: + - address: 10.10.10.22 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-23 +spec: {} +status: + addresses: + - address: 10.10.10.23 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-24 +spec: {} +status: + addresses: + - address: 10.10.10.24 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-25 +spec: {} +status: + addresses: + - address: 10.10.10.25 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-26 +spec: {} +status: + addresses: + - address: 10.10.10.26 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-27 +spec: {} +status: + addresses: + - address: 10.10.10.27 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-28 +spec: {} +status: + addresses: + - address: 10.10.10.28 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-29 +spec: {} +status: + addresses: + - address: 10.10.10.29 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-30 +spec: {} +status: + addresses: + - address: 10.10.10.30 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-31 +spec: {} +status: + addresses: + - address: 10.10.10.31 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-32 +spec: {} +status: + addresses: + - address: 10.10.10.32 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/ops.yaml new file mode 100644 index 000000000..0f85dca2b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/ops.yaml @@ -0,0 +1,43 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-pci-01 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 + - worker-04 + - worker-05 + - worker-06 + - worker-07 + - worker-08 + - worker-09 + - worker-10 + - worker-11 + - worker-12 + - worker-13 + - worker-14 + - worker-15 + - worker-16 + - worker-17 + - worker-18 + - worker-19 + - worker-20 + - worker-21 + - worker-22 + - worker-23 + - worker-24 + - worker-25 + - worker-26 + - worker-27 + - worker-28 + - worker-29 + - worker-30 + - worker-31 + - worker-32 diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-01.yaml new file mode 100644 index 000000000..e60b149b7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-01-7d92dc1d + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-02.yaml new file mode 100644 index 000000000..7bf065fe0 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-02-c289e1a6 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-03.yaml new file mode 100644 index 000000000..251a3186c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-03-fe9b2756 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-04.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-04.yaml new file mode 100644 index 000000000..67f314063 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-04.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-04", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.4" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.4" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-04 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-04-9cd39c46 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-05.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-05.yaml new file mode 100644 index 000000000..16a959b55 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-05.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-05", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.5" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.5" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-05 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-05-77ae00c7 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-06.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-06.yaml new file mode 100644 index 000000000..938e9aa29 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-06.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-06", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.6" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.6" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-06 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-06-4e5c6e75 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-07.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-07.yaml new file mode 100644 index 000000000..91e03adda --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-07.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-07", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.7" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.7" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-07 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-07-e1bd41d4 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-08.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-08.yaml new file mode 100644 index 000000000..e477db948 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-08.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-08", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.8" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.8" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-08 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-08-b8f97fc2 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-09.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-09.yaml new file mode 100644 index 000000000..bc7ee55d9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-09.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-09", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.9" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.9" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-09 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-09-89c78385 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-10.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-10.yaml new file mode 100644 index 000000000..1f98792dc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-10.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-10", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.10" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.10" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-10 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-10-bbb55e52 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-11.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-11.yaml new file mode 100644 index 000000000..e3d386dd8 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-11.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-11", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.11" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.11" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-11 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-11-bb25efa5 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-12.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-12.yaml new file mode 100644 index 000000000..b7e4bc5d7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-12.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-12", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.12" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.12" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-12 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-12-1815b141 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-13.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-13.yaml new file mode 100644 index 000000000..f3cc6571b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-13.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-13", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.13" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.13" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-13 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-13-12377513 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-14.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-14.yaml new file mode 100644 index 000000000..a96d6c8e3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-14.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-14", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.14" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.14" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-14 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-14-0d7f1d60 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-15.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-15.yaml new file mode 100644 index 000000000..2f2a01fd2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-15.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-15", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.15" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.15" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-15 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-15-654e0b00 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-16.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-16.yaml new file mode 100644 index 000000000..e60d69918 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-16.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-16", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.16" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.16" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-16 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-16-c5b53cd0 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-17.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-17.yaml new file mode 100644 index 000000000..c889d7c54 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-17.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-17", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.17" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.17" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-17 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-17-47d29168 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-18.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-18.yaml new file mode 100644 index 000000000..5fb824740 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-18.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-18", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.18" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.18" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-18 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-18-34f030fa + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-19.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-19.yaml new file mode 100644 index 000000000..480b08b6a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-19.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-19", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.19" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.19" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-19 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-19-b12224c3 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-20.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-20.yaml new file mode 100644 index 000000000..92fddb9ec --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-20.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-20", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.20" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.20" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-20 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-20-ebd70816 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-21.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-21.yaml new file mode 100644 index 000000000..8afb7ac6d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-21.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-21", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.21" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.21" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-21 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-21-f6a9aac3 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-22.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-22.yaml new file mode 100644 index 000000000..c5b0e97bd --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-22.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-22", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.22" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.22" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-22 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-22-2d62616c + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-23.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-23.yaml new file mode 100644 index 000000000..012e5a1fd --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-23.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-23", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.23" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.23" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-23 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-23-7bb9e25d + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-24.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-24.yaml new file mode 100644 index 000000000..97d8f4d46 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-24.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-24", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.24" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.24" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-24 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-24-edd02edb + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-25.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-25.yaml new file mode 100644 index 000000000..7a1bfe97f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-25.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-25", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.25" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.25" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-25 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-25-ce323435 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-26.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-26.yaml new file mode 100644 index 000000000..c380deb00 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-26.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-26", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.26" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.26" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-26 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-26-53b6466c + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-27.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-27.yaml new file mode 100644 index 000000000..f8ad60e03 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-27.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-27", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.27" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.27" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-27 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-27-c9dbb83b + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-28.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-28.yaml new file mode 100644 index 000000000..f93734c70 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-28.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-28", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.28" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.28" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-28 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-28-9487ee56 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-29.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-29.yaml new file mode 100644 index 000000000..e4e9f4bf5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-29.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-29", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.29" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.29" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-29 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-29-9e0a59c1 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-30.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-30.yaml new file mode 100644 index 000000000..dc10ecebc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-30.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-30", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.30" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.30" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-30 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-30-c1a761ff + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-31.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-31.yaml new file mode 100644 index 000000000..97f529f08 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-31.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-31", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.31" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.31" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-31 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-31-a8f4d33b + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-32.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-32.yaml new file mode 100644 index 000000000..4319933ad --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/reports/worker-32.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-32", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.32" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.32" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-32 + storage.simplyblock.io/nodeprobe-run: discover-pci-01 + name: sb-nodeprobe-discover-pci-01-worker-32-47ce7fc4 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/case.md b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/case.md new file mode 100644 index 000000000..8a70a9513 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/case.md @@ -0,0 +1,7 @@ +# PCI-02 + +**Mutation.** 2 workers, 4 disks each, different slots + +**Expected.** 2 groups of 1 worker each + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/expected-notes.txt new file mode 100644 index 000000000..36d92737d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-pci-02 in Draft: 2 workers with 8 nvme devices, in 2 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/expected.yaml new file mode 100644 index 000000000..501e594db --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/expected.yaml @@ -0,0 +1,42 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-pci-02 + name: discovered-discover-pci-02 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-pci-02-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - devices: + nvme: + - 0000:3b:00.0 + - 0000:3c:00.0 + - 0000:d8:00.0 + - 0000:d9:00.0 + mgmtInterface: eth0 + name: group-2-nvme-4x3T + workers: + - worker-02 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/nodes.yaml new file mode 100644 index 000000000..1fba89979 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/nodes.yaml @@ -0,0 +1,59 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/ops.yaml new file mode 100644 index 000000000..8d2e36ac8 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/ops.yaml @@ -0,0 +1,13 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-pci-02 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/reports/worker-01.yaml new file mode 100644 index 000000000..a91cdcbea --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-pci-02 + name: sb-nodeprobe-discover-pci-02-worker-01-f0f3067d + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/reports/worker-02.yaml new file mode 100644 index 000000000..6588b0250 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-pci-02 + name: sb-nodeprobe-discover-pci-02-worker-02-ab00718b + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/case.md b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/case.md new file mode 100644 index 000000000..5e3fc2481 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/case.md @@ -0,0 +1,7 @@ +# PCI-03 + +**Mutation.** 32 workers: 16 on layout A and 16 on layout B + +**Expected.** 2 groups of 16, ordered by their first worker's name + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/expected-notes.txt new file mode 100644 index 000000000..bacfdf776 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-pci-03 in Draft: 32 workers with 128 nvme devices, in 2 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/expected.yaml new file mode 100644 index 000000000..40b9e9a78 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/expected.yaml @@ -0,0 +1,72 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: discovered-discover-pci-03 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-pci-03-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - worker-02 + - worker-03 + - worker-04 + - worker-05 + - worker-06 + - worker-07 + - worker-08 + - worker-09 + - worker-10 + - worker-11 + - worker-12 + - worker-13 + - worker-14 + - worker-15 + - worker-16 + - devices: + nvme: + - 0000:3b:00.0 + - 0000:3c:00.0 + - 0000:d8:00.0 + - 0000:d9:00.0 + mgmtInterface: eth0 + name: group-2-nvme-4x3T + workers: + - worker-17 + - worker-18 + - worker-19 + - worker-20 + - worker-21 + - worker-22 + - worker-23 + - worker-24 + - worker-25 + - worker-26 + - worker-27 + - worker-28 + - worker-29 + - worker-30 + - worker-31 + - worker-32 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/nodes.yaml new file mode 100644 index 000000000..c25e3f497 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/nodes.yaml @@ -0,0 +1,959 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-04 +spec: {} +status: + addresses: + - address: 10.10.10.4 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-05 +spec: {} +status: + addresses: + - address: 10.10.10.5 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-06 +spec: {} +status: + addresses: + - address: 10.10.10.6 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-07 +spec: {} +status: + addresses: + - address: 10.10.10.7 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-08 +spec: {} +status: + addresses: + - address: 10.10.10.8 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-09 +spec: {} +status: + addresses: + - address: 10.10.10.9 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-10 +spec: {} +status: + addresses: + - address: 10.10.10.10 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-11 +spec: {} +status: + addresses: + - address: 10.10.10.11 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-12 +spec: {} +status: + addresses: + - address: 10.10.10.12 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-13 +spec: {} +status: + addresses: + - address: 10.10.10.13 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-14 +spec: {} +status: + addresses: + - address: 10.10.10.14 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-15 +spec: {} +status: + addresses: + - address: 10.10.10.15 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-16 +spec: {} +status: + addresses: + - address: 10.10.10.16 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-17 +spec: {} +status: + addresses: + - address: 10.10.10.17 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-18 +spec: {} +status: + addresses: + - address: 10.10.10.18 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-19 +spec: {} +status: + addresses: + - address: 10.10.10.19 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-20 +spec: {} +status: + addresses: + - address: 10.10.10.20 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-21 +spec: {} +status: + addresses: + - address: 10.10.10.21 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-22 +spec: {} +status: + addresses: + - address: 10.10.10.22 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-23 +spec: {} +status: + addresses: + - address: 10.10.10.23 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-24 +spec: {} +status: + addresses: + - address: 10.10.10.24 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-25 +spec: {} +status: + addresses: + - address: 10.10.10.25 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-26 +spec: {} +status: + addresses: + - address: 10.10.10.26 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-27 +spec: {} +status: + addresses: + - address: 10.10.10.27 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-28 +spec: {} +status: + addresses: + - address: 10.10.10.28 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-29 +spec: {} +status: + addresses: + - address: 10.10.10.29 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-30 +spec: {} +status: + addresses: + - address: 10.10.10.30 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-31 +spec: {} +status: + addresses: + - address: 10.10.10.31 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-32 +spec: {} +status: + addresses: + - address: 10.10.10.32 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/ops.yaml new file mode 100644 index 000000000..dd3773d00 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/ops.yaml @@ -0,0 +1,43 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-pci-03 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 + - worker-04 + - worker-05 + - worker-06 + - worker-07 + - worker-08 + - worker-09 + - worker-10 + - worker-11 + - worker-12 + - worker-13 + - worker-14 + - worker-15 + - worker-16 + - worker-17 + - worker-18 + - worker-19 + - worker-20 + - worker-21 + - worker-22 + - worker-23 + - worker-24 + - worker-25 + - worker-26 + - worker-27 + - worker-28 + - worker-29 + - worker-30 + - worker-31 + - worker-32 diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-01.yaml new file mode 100644 index 000000000..878ea776d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-01-fc5f0aaf + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-02.yaml new file mode 100644 index 000000000..1ae1b4691 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-02-3fd467a5 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-03.yaml new file mode 100644 index 000000000..5ae4c07cd --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-03-b4778c39 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-04.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-04.yaml new file mode 100644 index 000000000..d962153b6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-04.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-04", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.4" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.4" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-04 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-04-9733558a + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-05.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-05.yaml new file mode 100644 index 000000000..e0b7f8fa9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-05.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-05", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.5" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.5" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-05 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-05-99c3fcba + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-06.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-06.yaml new file mode 100644 index 000000000..d94fa3fef --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-06.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-06", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.6" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.6" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-06 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-06-0121789f + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-07.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-07.yaml new file mode 100644 index 000000000..6d9b2f831 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-07.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-07", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.7" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.7" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-07 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-07-6f2b20b5 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-08.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-08.yaml new file mode 100644 index 000000000..ec3b0e346 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-08.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-08", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.8" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.8" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-08 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-08-8c689e04 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-09.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-09.yaml new file mode 100644 index 000000000..503e8db64 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-09.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-09", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.9" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.9" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-09 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-09-3dbc4617 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-10.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-10.yaml new file mode 100644 index 000000000..718c7d4c0 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-10.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-10", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.10" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.10" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-10 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-10-cda25316 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-11.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-11.yaml new file mode 100644 index 000000000..6da3ba442 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-11.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-11", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.11" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.11" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-11 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-11-9f87676c + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-12.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-12.yaml new file mode 100644 index 000000000..66e97c88b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-12.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-12", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.12" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.12" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-12 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-12-8b6a9128 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-13.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-13.yaml new file mode 100644 index 000000000..fb6837328 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-13.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-13", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.13" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.13" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-13 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-13-7e1ba59e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-14.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-14.yaml new file mode 100644 index 000000000..b62fc9388 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-14.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-14", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.14" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.14" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-14 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-14-4a98d5a6 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-15.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-15.yaml new file mode 100644 index 000000000..06c7df200 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-15.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-15", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.15" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.15" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-15 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-15-92cabd34 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-16.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-16.yaml new file mode 100644 index 000000000..e9cbd01c0 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-16.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-16", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.16" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.16" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-16 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-16-69aec2a4 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-17.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-17.yaml new file mode 100644 index 000000000..ea5a77e01 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-17.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-17", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.17" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.17" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-17 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-17-f9d16d74 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-18.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-18.yaml new file mode 100644 index 000000000..f74678759 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-18.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-18", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.18" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.18" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-18 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-18-8082b52b + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-19.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-19.yaml new file mode 100644 index 000000000..6fbc6f25d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-19.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-19", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.19" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.19" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-19 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-19-fd27131d + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-20.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-20.yaml new file mode 100644 index 000000000..0dbb5eede --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-20.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-20", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.20" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.20" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-20 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-20-c7c77091 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-21.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-21.yaml new file mode 100644 index 000000000..15a8672d4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-21.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-21", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.21" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.21" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-21 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-21-06e0b3f6 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-22.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-22.yaml new file mode 100644 index 000000000..b7e1cc12a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-22.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-22", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.22" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.22" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-22 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-22-89a90764 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-23.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-23.yaml new file mode 100644 index 000000000..dd257c881 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-23.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-23", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.23" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.23" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-23 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-23-2b03ab51 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-24.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-24.yaml new file mode 100644 index 000000000..225e49949 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-24.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-24", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.24" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.24" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-24 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-24-d0f01afe + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-25.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-25.yaml new file mode 100644 index 000000000..529989f74 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-25.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-25", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.25" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.25" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-25 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-25-858d117d + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-26.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-26.yaml new file mode 100644 index 000000000..690cd2dc1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-26.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-26", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.26" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.26" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-26 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-26-edd2ae2b + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-27.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-27.yaml new file mode 100644 index 000000000..a18d5d059 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-27.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-27", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.27" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.27" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-27 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-27-c1a0a056 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-28.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-28.yaml new file mode 100644 index 000000000..e1662d014 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-28.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-28", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.28" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.28" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-28 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-28-fddee6d1 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-29.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-29.yaml new file mode 100644 index 000000000..6cc4c1ca4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-29.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-29", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.29" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.29" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-29 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-29-0b8f8ce1 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-30.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-30.yaml new file mode 100644 index 000000000..7c24b44e0 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-30.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-30", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.30" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.30" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-30 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-30-23a3d30e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-31.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-31.yaml new file mode 100644 index 000000000..26974f5d4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-31.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-31", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.31" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.31" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-31 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-31-b750ffed + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-32.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-32.yaml new file mode 100644 index 000000000..a60435ab5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/reports/worker-32.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-32", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.32" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.32" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-32 + storage.simplyblock.io/nodeprobe-run: discover-pci-03 + name: sb-nodeprobe-discover-pci-03-worker-32-45ee7f9e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/case.md b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/case.md new file mode 100644 index 000000000..adae1c041 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/case.md @@ -0,0 +1,7 @@ +# PCI-04 + +**Mutation.** 32 workers, every layout distinct + +**Expected.** 32 groups in one node set, under the 64-group ceiling + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/expected-notes.txt new file mode 100644 index 000000000..d732febc9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-pci-04 in Draft: 32 workers with 64 nvme devices, in 32 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/expected.yaml new file mode 100644 index 000000000..69ab3f467 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/expected.yaml @@ -0,0 +1,278 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: discovered-discover-pci-04 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-pci-04-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - "0000:11:00.0" + - "0000:11:00.1" + mgmtInterface: eth0 + name: group-1-nvme-2x3T + workers: + - worker-01 + - devices: + nvme: + - "0000:12:00.0" + - "0000:12:00.1" + mgmtInterface: eth0 + name: group-2-nvme-2x3T + workers: + - worker-02 + - devices: + nvme: + - "0000:13:00.0" + - "0000:13:00.1" + mgmtInterface: eth0 + name: group-3-nvme-2x3T + workers: + - worker-03 + - devices: + nvme: + - "0000:14:00.0" + - "0000:14:00.1" + mgmtInterface: eth0 + name: group-4-nvme-2x3T + workers: + - worker-04 + - devices: + nvme: + - "0000:15:00.0" + - "0000:15:00.1" + mgmtInterface: eth0 + name: group-5-nvme-2x3T + workers: + - worker-05 + - devices: + nvme: + - "0000:16:00.0" + - "0000:16:00.1" + mgmtInterface: eth0 + name: group-6-nvme-2x3T + workers: + - worker-06 + - devices: + nvme: + - "0000:17:00.0" + - "0000:17:00.1" + mgmtInterface: eth0 + name: group-7-nvme-2x3T + workers: + - worker-07 + - devices: + nvme: + - "0000:18:00.0" + - "0000:18:00.1" + mgmtInterface: eth0 + name: group-8-nvme-2x3T + workers: + - worker-08 + - devices: + nvme: + - "0000:19:00.0" + - "0000:19:00.1" + mgmtInterface: eth0 + name: group-9-nvme-2x3T + workers: + - worker-09 + - devices: + nvme: + - 0000:1a:00.0 + - 0000:1a:00.1 + mgmtInterface: eth0 + name: group-10-nvme-2x3T + workers: + - worker-10 + - devices: + nvme: + - 0000:1b:00.0 + - 0000:1b:00.1 + mgmtInterface: eth0 + name: group-11-nvme-2x3T + workers: + - worker-11 + - devices: + nvme: + - 0000:1c:00.0 + - 0000:1c:00.1 + mgmtInterface: eth0 + name: group-12-nvme-2x3T + workers: + - worker-12 + - devices: + nvme: + - 0000:1d:00.0 + - 0000:1d:00.1 + mgmtInterface: eth0 + name: group-13-nvme-2x3T + workers: + - worker-13 + - devices: + nvme: + - 0000:1e:00.0 + - 0000:1e:00.1 + mgmtInterface: eth0 + name: group-14-nvme-2x3T + workers: + - worker-14 + - devices: + nvme: + - 0000:1f:00.0 + - 0000:1f:00.1 + mgmtInterface: eth0 + name: group-15-nvme-2x3T + workers: + - worker-15 + - devices: + nvme: + - "0000:20:00.0" + - "0000:20:00.1" + mgmtInterface: eth0 + name: group-16-nvme-2x3T + workers: + - worker-16 + - devices: + nvme: + - "0000:21:00.0" + - "0000:21:00.1" + mgmtInterface: eth0 + name: group-17-nvme-2x3T + workers: + - worker-17 + - devices: + nvme: + - "0000:22:00.0" + - "0000:22:00.1" + mgmtInterface: eth0 + name: group-18-nvme-2x3T + workers: + - worker-18 + - devices: + nvme: + - "0000:23:00.0" + - "0000:23:00.1" + mgmtInterface: eth0 + name: group-19-nvme-2x3T + workers: + - worker-19 + - devices: + nvme: + - "0000:24:00.0" + - "0000:24:00.1" + mgmtInterface: eth0 + name: group-20-nvme-2x3T + workers: + - worker-20 + - devices: + nvme: + - "0000:25:00.0" + - "0000:25:00.1" + mgmtInterface: eth0 + name: group-21-nvme-2x3T + workers: + - worker-21 + - devices: + nvme: + - "0000:26:00.0" + - "0000:26:00.1" + mgmtInterface: eth0 + name: group-22-nvme-2x3T + workers: + - worker-22 + - devices: + nvme: + - "0000:27:00.0" + - "0000:27:00.1" + mgmtInterface: eth0 + name: group-23-nvme-2x3T + workers: + - worker-23 + - devices: + nvme: + - "0000:28:00.0" + - "0000:28:00.1" + mgmtInterface: eth0 + name: group-24-nvme-2x3T + workers: + - worker-24 + - devices: + nvme: + - "0000:29:00.0" + - "0000:29:00.1" + mgmtInterface: eth0 + name: group-25-nvme-2x3T + workers: + - worker-25 + - devices: + nvme: + - 0000:2a:00.0 + - 0000:2a:00.1 + mgmtInterface: eth0 + name: group-26-nvme-2x3T + workers: + - worker-26 + - devices: + nvme: + - 0000:2b:00.0 + - 0000:2b:00.1 + mgmtInterface: eth0 + name: group-27-nvme-2x3T + workers: + - worker-27 + - devices: + nvme: + - 0000:2c:00.0 + - 0000:2c:00.1 + mgmtInterface: eth0 + name: group-28-nvme-2x3T + workers: + - worker-28 + - devices: + nvme: + - 0000:2d:00.0 + - 0000:2d:00.1 + mgmtInterface: eth0 + name: group-29-nvme-2x3T + workers: + - worker-29 + - devices: + nvme: + - 0000:2e:00.0 + - 0000:2e:00.1 + mgmtInterface: eth0 + name: group-30-nvme-2x3T + workers: + - worker-30 + - devices: + nvme: + - 0000:2f:00.0 + - 0000:2f:00.1 + mgmtInterface: eth0 + name: group-31-nvme-2x3T + workers: + - worker-31 + - devices: + nvme: + - "0000:30:00.0" + - "0000:30:00.1" + mgmtInterface: eth0 + name: group-32-nvme-2x3T + workers: + - worker-32 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/nodes.yaml new file mode 100644 index 000000000..c25e3f497 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/nodes.yaml @@ -0,0 +1,959 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-04 +spec: {} +status: + addresses: + - address: 10.10.10.4 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-05 +spec: {} +status: + addresses: + - address: 10.10.10.5 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-06 +spec: {} +status: + addresses: + - address: 10.10.10.6 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-07 +spec: {} +status: + addresses: + - address: 10.10.10.7 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-08 +spec: {} +status: + addresses: + - address: 10.10.10.8 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-09 +spec: {} +status: + addresses: + - address: 10.10.10.9 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-10 +spec: {} +status: + addresses: + - address: 10.10.10.10 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-11 +spec: {} +status: + addresses: + - address: 10.10.10.11 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-12 +spec: {} +status: + addresses: + - address: 10.10.10.12 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-13 +spec: {} +status: + addresses: + - address: 10.10.10.13 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-14 +spec: {} +status: + addresses: + - address: 10.10.10.14 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-15 +spec: {} +status: + addresses: + - address: 10.10.10.15 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-16 +spec: {} +status: + addresses: + - address: 10.10.10.16 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-17 +spec: {} +status: + addresses: + - address: 10.10.10.17 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-18 +spec: {} +status: + addresses: + - address: 10.10.10.18 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-19 +spec: {} +status: + addresses: + - address: 10.10.10.19 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-20 +spec: {} +status: + addresses: + - address: 10.10.10.20 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-21 +spec: {} +status: + addresses: + - address: 10.10.10.21 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-22 +spec: {} +status: + addresses: + - address: 10.10.10.22 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-23 +spec: {} +status: + addresses: + - address: 10.10.10.23 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-24 +spec: {} +status: + addresses: + - address: 10.10.10.24 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-25 +spec: {} +status: + addresses: + - address: 10.10.10.25 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-26 +spec: {} +status: + addresses: + - address: 10.10.10.26 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-27 +spec: {} +status: + addresses: + - address: 10.10.10.27 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-28 +spec: {} +status: + addresses: + - address: 10.10.10.28 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-29 +spec: {} +status: + addresses: + - address: 10.10.10.29 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-30 +spec: {} +status: + addresses: + - address: 10.10.10.30 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-31 +spec: {} +status: + addresses: + - address: 10.10.10.31 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-32 +spec: {} +status: + addresses: + - address: 10.10.10.32 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/ops.yaml new file mode 100644 index 000000000..9d03ae795 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/ops.yaml @@ -0,0 +1,43 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-pci-04 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 + - worker-04 + - worker-05 + - worker-06 + - worker-07 + - worker-08 + - worker-09 + - worker-10 + - worker-11 + - worker-12 + - worker-13 + - worker-14 + - worker-15 + - worker-16 + - worker-17 + - worker-18 + - worker-19 + - worker-20 + - worker-21 + - worker-22 + - worker-23 + - worker-24 + - worker-25 + - worker-26 + - worker-27 + - worker-28 + - worker-29 + - worker-30 + - worker-31 + - worker-32 diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-01.yaml new file mode 100644 index 000000000..6c6f7178d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-01.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:11:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:11:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-01-e83c41a4 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-02.yaml new file mode 100644 index 000000000..9ede94d6c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-02.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:12:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:12:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-02-2dc2bc0c + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-03.yaml new file mode 100644 index 000000000..fe2f7bd5a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-03.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:13:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:13:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-03-b168db77 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-04.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-04.yaml new file mode 100644 index 000000000..d52b4460e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-04.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-04", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.4" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.4" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:14:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:14:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-04 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-04-ddbaf6ac + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-05.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-05.yaml new file mode 100644 index 000000000..f46695238 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-05.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-05", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.5" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.5" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:15:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:15:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-05 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-05-7cce4907 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-06.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-06.yaml new file mode 100644 index 000000000..9367cbc52 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-06.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-06", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.6" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.6" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:16:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:16:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-06 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-06-f23ab09a + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-07.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-07.yaml new file mode 100644 index 000000000..0d32057df --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-07.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-07", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.7" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.7" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:17:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:17:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-07 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-07-3b75c393 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-08.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-08.yaml new file mode 100644 index 000000000..134d658ac --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-08.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-08", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.8" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.8" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:18:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:18:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-08 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-08-8a6e0588 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-09.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-09.yaml new file mode 100644 index 000000000..d9b1526f9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-09.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-09", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.9" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.9" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:19:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:19:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-09 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-09-9b08de7a + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-10.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-10.yaml new file mode 100644 index 000000000..c7e94a2a9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-10.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-10", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.10" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.10" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:1a:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:1a:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-10 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-10-583e1bf2 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-11.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-11.yaml new file mode 100644 index 000000000..f7b6adc9e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-11.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-11", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.11" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.11" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:1b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:1b:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-11 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-11-eb0ad3ce + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-12.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-12.yaml new file mode 100644 index 000000000..9011d61f8 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-12.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-12", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.12" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.12" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:1c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:1c:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-12 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-12-27dbbb37 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-13.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-13.yaml new file mode 100644 index 000000000..f0b05472b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-13.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-13", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.13" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.13" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:1d:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:1d:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-13 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-13-522485f3 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-14.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-14.yaml new file mode 100644 index 000000000..9e56b5921 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-14.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-14", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.14" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.14" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:1e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:1e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-14 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-14-0fff409c + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-15.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-15.yaml new file mode 100644 index 000000000..20f083770 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-15.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-15", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.15" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.15" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:1f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:1f:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-15 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-15-10aadfc5 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-16.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-16.yaml new file mode 100644 index 000000000..3173533e4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-16.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-16", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.16" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.16" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:20:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:20:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-16 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-16-c4aff7e5 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-17.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-17.yaml new file mode 100644 index 000000000..2b6666d8a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-17.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-17", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.17" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.17" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:21:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:21:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-17 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-17-7f603a55 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-18.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-18.yaml new file mode 100644 index 000000000..cd3aed016 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-18.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-18", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.18" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.18" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:22:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:22:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-18 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-18-36fb1144 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-19.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-19.yaml new file mode 100644 index 000000000..2b0445014 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-19.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-19", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.19" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.19" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:23:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:23:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-19 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-19-178a0429 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-20.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-20.yaml new file mode 100644 index 000000000..95facc667 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-20.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-20", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.20" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.20" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:24:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:24:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-20 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-20-02e8ddbc + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-21.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-21.yaml new file mode 100644 index 000000000..457f130db --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-21.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-21", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.21" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.21" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:25:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:25:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-21 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-21-51ffd70e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-22.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-22.yaml new file mode 100644 index 000000000..ff7444cf8 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-22.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-22", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.22" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.22" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:26:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:26:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-22 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-22-37e512e4 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-23.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-23.yaml new file mode 100644 index 000000000..66fd14771 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-23.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-23", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.23" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.23" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:27:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:27:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-23 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-23-e051838c + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-24.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-24.yaml new file mode 100644 index 000000000..281e30123 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-24.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-24", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.24" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.24" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:28:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:28:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-24 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-24-84b5e9d7 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-25.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-25.yaml new file mode 100644 index 000000000..d38e2f41f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-25.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-25", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.25" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.25" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:29:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:29:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-25 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-25-c4e0abb5 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-26.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-26.yaml new file mode 100644 index 000000000..535ce5847 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-26.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-26", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.26" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.26" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:2a:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:2a:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-26 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-26-290a221b + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-27.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-27.yaml new file mode 100644 index 000000000..3ee6abf50 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-27.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-27", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.27" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.27" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:2b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:2b:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-27 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-27-23fc1d04 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-28.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-28.yaml new file mode 100644 index 000000000..7744a654e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-28.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-28", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.28" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.28" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:2c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:2c:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-28 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-28-bf064c64 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-29.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-29.yaml new file mode 100644 index 000000000..609fa8a0d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-29.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-29", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.29" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.29" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:2d:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:2d:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-29 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-29-bb2e7bec + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-30.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-30.yaml new file mode 100644 index 000000000..6a0807d1f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-30.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-30", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.30" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.30" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:2e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:2e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-30 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-30-413ac3c7 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-31.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-31.yaml new file mode 100644 index 000000000..3f54b5ce8 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-31.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-31", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.31" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.31" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:2f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:2f:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-31 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-31-7b331297 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-32.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-32.yaml new file mode 100644 index 000000000..b8b1bcce3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/reports/worker-32.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-32", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.32" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.32" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:30:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:30:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-32 + storage.simplyblock.io/nodeprobe-run: discover-pci-04 + name: sb-nodeprobe-discover-pci-04-worker-32-36901f5b + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/case.md b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/case.md new file mode 100644 index 000000000..a7e62bf93 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/case.md @@ -0,0 +1,7 @@ +# PCI-05 + +**Mutation.** Addresses reported in descending order + +**Expected.** `devices.nvme` ascending: the draft sorts + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/expected-notes.txt new file mode 100644 index 000000000..ff726a85f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-pci-05 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/expected.yaml new file mode 100644 index 000000000..760378848 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-pci-05 + name: discovered-discover-pci-05 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-pci-05-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/ops.yaml new file mode 100644 index 000000000..c2697a732 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-pci-05 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/reports/worker-01.yaml new file mode 100644 index 000000000..c5f0eddac --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-pci-05 + name: sb-nodeprobe-discover-pci-05-worker-01-e2bd502c + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/case.md b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/case.md new file mode 100644 index 000000000..1b1ac208d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/case.md @@ -0,0 +1,7 @@ +# PCI-06 + +**Mutation.** 10 disks across two buses, `0000:5e:*` and `0000:af:*` + +**Expected.** One group, addresses ascending across both buses + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/expected-notes.txt new file mode 100644 index 000000000..a0dfefe3b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-pci-06 in Draft: 1 workers with 10 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/expected.yaml new file mode 100644 index 000000000..6d8e4368c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/expected.yaml @@ -0,0 +1,38 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-pci-06 + name: discovered-discover-pci-06 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-pci-06-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5f:00.0 + - 0000:5f:00.1 + - 0000:60:00.0 + - 0000:af:00.0 + - 0000:af:00.1 + - 0000:b0:00.0 + - 0000:b0:00.1 + - 0000:b1:00.0 + mgmtInterface: eth0 + name: group-1-nvme-10x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/ops.yaml new file mode 100644 index 000000000..2714cfd63 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-pci-06 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/reports/worker-01.yaml new file mode 100644 index 000000000..2e585aefe --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/reports/worker-01.yaml @@ -0,0 +1,247 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5f:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme4n1", + "path": "/dev/nvme4n1", + "pciAddress": "0000:60:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme5n1", + "path": "/dev/nvme5n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme6n1", + "path": "/dev/nvme6n1", + "pciAddress": "0000:af:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme7n1", + "path": "/dev/nvme7n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme8n1", + "path": "/dev/nvme8n1", + "pciAddress": "0000:b0:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme9n1", + "path": "/dev/nvme9n1", + "pciAddress": "0000:b1:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-pci-06 + name: sb-nodeprobe-discover-pci-06-worker-01-df460ee5 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/case.md b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/case.md new file mode 100644 index 000000000..91d00b3ca --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/case.md @@ -0,0 +1,9 @@ +# PCI-07 + +**Mutation.** A five-digit PCI domain, `10000:01:00.0` + +**Expected.** **Contested.** The generator names it and the schema's item pattern refuses the document. See §14, gap G-6 + +**Harness.** `CM` + +**Gap.** G-6. This case records what the generator does today, so that the day it changes the diff is the finding. diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/expected-notes.txt new file mode 100644 index 000000000..401902a64 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-pci-07 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/expected.yaml new file mode 100644 index 000000000..0a2c0298c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/expected.yaml @@ -0,0 +1,30 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-pci-07 + name: discovered-discover-pci-07 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-pci-07-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - "10000:01:00.0" + - "10000:02:00.0" + mgmtInterface: eth0 + name: group-1-nvme-2x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/ops.yaml new file mode 100644 index 000000000..0ea32da33 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-pci-07 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/reports/worker-01.yaml new file mode 100644 index 000000000..6906212e9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/reports/worker-01.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "10000:01:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "10000:02:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-pci-07 + name: sb-nodeprobe-discover-pci-07-worker-01-7998a144 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/case.md b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/case.md new file mode 100644 index 000000000..31a48b767 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/case.md @@ -0,0 +1,7 @@ +# PCI-08 + +**Mutation.** Uppercase hex in the report, `0000:5E:00.0` + +**Expected.** Named as reported. Grouping is byte-exact, so a fleet mixing cases splits + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/expected-notes.txt new file mode 100644 index 000000000..9b87e4a35 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-pci-08 in Draft: 2 workers with 4 nvme devices, in 2 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/expected.yaml new file mode 100644 index 000000000..e0683793f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/expected.yaml @@ -0,0 +1,38 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-pci-08 + name: discovered-discover-pci-08 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-pci-08-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5E:00.0 + - 0000:5F:00.0 + mgmtInterface: eth0 + name: group-1-nvme-2x3T + workers: + - worker-01 + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + mgmtInterface: eth0 + name: group-2-nvme-2x3T + workers: + - worker-02 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/nodes.yaml new file mode 100644 index 000000000..1fba89979 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/nodes.yaml @@ -0,0 +1,59 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/ops.yaml new file mode 100644 index 000000000..04756bbd7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/ops.yaml @@ -0,0 +1,13 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-pci-08 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/reports/worker-01.yaml new file mode 100644 index 000000000..39ac1ae48 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/reports/worker-01.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5E:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5F:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-pci-08 + name: sb-nodeprobe-discover-pci-08-worker-01-58d6dda5 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/reports/worker-02.yaml new file mode 100644 index 000000000..4cee2ade7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/reports/worker-02.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-pci-08 + name: sb-nodeprobe-discover-pci-08-worker-02-2d445300 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/case.md b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/case.md new file mode 100644 index 000000000..c285aa75e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/case.md @@ -0,0 +1,7 @@ +# PCI-09 + +**Mutation.** Workers named `worker-1` … `worker-32` with two layouts + +**Expected.** Group numbering follows lexicographic worker order: `worker-1`, `worker-10`, `worker-11` + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/expected-notes.txt new file mode 100644 index 000000000..93254abff --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-1, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-1, which is the smallest of the fleet +wrote discovered-discover-pci-09 in Draft: 32 workers with 128 nvme devices, in 2 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/expected.yaml new file mode 100644 index 000000000..7996e2a59 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/expected.yaml @@ -0,0 +1,72 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: discovered-discover-pci-09 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-pci-09-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-1 + - worker-10 + - worker-11 + - worker-12 + - worker-13 + - worker-14 + - worker-15 + - worker-16 + - worker-2 + - worker-3 + - worker-4 + - worker-5 + - worker-6 + - worker-7 + - worker-8 + - worker-9 + - devices: + nvme: + - 0000:3b:00.0 + - 0000:3c:00.0 + - 0000:d8:00.0 + - 0000:d9:00.0 + mgmtInterface: eth0 + name: group-2-nvme-4x3T + workers: + - worker-17 + - worker-18 + - worker-19 + - worker-20 + - worker-21 + - worker-22 + - worker-23 + - worker-24 + - worker-25 + - worker-26 + - worker-27 + - worker-28 + - worker-29 + - worker-30 + - worker-31 + - worker-32 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/nodes.yaml new file mode 100644 index 000000000..c4c813a9b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/nodes.yaml @@ -0,0 +1,959 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-1 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-2 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-3 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-4 +spec: {} +status: + addresses: + - address: 10.10.10.4 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-5 +spec: {} +status: + addresses: + - address: 10.10.10.5 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-6 +spec: {} +status: + addresses: + - address: 10.10.10.6 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-7 +spec: {} +status: + addresses: + - address: 10.10.10.7 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-8 +spec: {} +status: + addresses: + - address: 10.10.10.8 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-9 +spec: {} +status: + addresses: + - address: 10.10.10.9 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-10 +spec: {} +status: + addresses: + - address: 10.10.10.10 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-11 +spec: {} +status: + addresses: + - address: 10.10.10.11 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-12 +spec: {} +status: + addresses: + - address: 10.10.10.12 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-13 +spec: {} +status: + addresses: + - address: 10.10.10.13 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-14 +spec: {} +status: + addresses: + - address: 10.10.10.14 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-15 +spec: {} +status: + addresses: + - address: 10.10.10.15 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-16 +spec: {} +status: + addresses: + - address: 10.10.10.16 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-17 +spec: {} +status: + addresses: + - address: 10.10.10.17 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-18 +spec: {} +status: + addresses: + - address: 10.10.10.18 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-19 +spec: {} +status: + addresses: + - address: 10.10.10.19 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-20 +spec: {} +status: + addresses: + - address: 10.10.10.20 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-21 +spec: {} +status: + addresses: + - address: 10.10.10.21 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-22 +spec: {} +status: + addresses: + - address: 10.10.10.22 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-23 +spec: {} +status: + addresses: + - address: 10.10.10.23 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-24 +spec: {} +status: + addresses: + - address: 10.10.10.24 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-25 +spec: {} +status: + addresses: + - address: 10.10.10.25 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-26 +spec: {} +status: + addresses: + - address: 10.10.10.26 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-27 +spec: {} +status: + addresses: + - address: 10.10.10.27 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-28 +spec: {} +status: + addresses: + - address: 10.10.10.28 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-29 +spec: {} +status: + addresses: + - address: 10.10.10.29 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-30 +spec: {} +status: + addresses: + - address: 10.10.10.30 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-31 +spec: {} +status: + addresses: + - address: 10.10.10.31 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-32 +spec: {} +status: + addresses: + - address: 10.10.10.32 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/ops.yaml new file mode 100644 index 000000000..596786566 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/ops.yaml @@ -0,0 +1,43 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-pci-09 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-1 + - worker-2 + - worker-3 + - worker-4 + - worker-5 + - worker-6 + - worker-7 + - worker-8 + - worker-9 + - worker-10 + - worker-11 + - worker-12 + - worker-13 + - worker-14 + - worker-15 + - worker-16 + - worker-17 + - worker-18 + - worker-19 + - worker-20 + - worker-21 + - worker-22 + - worker-23 + - worker-24 + - worker-25 + - worker-26 + - worker-27 + - worker-28 + - worker-29 + - worker-30 + - worker-31 + - worker-32 diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-1.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-1.yaml new file mode 100644 index 000000000..2c2e34fcb --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-1.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-1", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-1 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-1-af8b3008 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-10.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-10.yaml new file mode 100644 index 000000000..72c1489fe --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-10.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-10", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.10" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.10" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-10 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-10-2b289882 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-11.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-11.yaml new file mode 100644 index 000000000..341fe1c7a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-11.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-11", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.11" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.11" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-11 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-11-e5f5abdc + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-12.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-12.yaml new file mode 100644 index 000000000..70fd344a1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-12.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-12", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.12" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.12" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-12 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-12-e70c6afc + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-13.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-13.yaml new file mode 100644 index 000000000..0a43d2590 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-13.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-13", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.13" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.13" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-13 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-13-6f1cb330 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-14.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-14.yaml new file mode 100644 index 000000000..c032b1aa5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-14.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-14", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.14" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.14" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-14 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-14-a8f7c1ba + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-15.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-15.yaml new file mode 100644 index 000000000..f6c9f3c31 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-15.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-15", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.15" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.15" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-15 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-15-3c59240a + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-16.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-16.yaml new file mode 100644 index 000000000..0d7b40c55 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-16.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-16", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.16" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.16" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-16 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-16-29792a65 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-17.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-17.yaml new file mode 100644 index 000000000..17c0de652 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-17.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-17", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.17" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.17" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-17 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-17-fee26890 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-18.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-18.yaml new file mode 100644 index 000000000..cb6d5df33 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-18.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-18", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.18" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.18" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-18 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-18-5e67933e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-19.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-19.yaml new file mode 100644 index 000000000..231a03c8d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-19.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-19", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.19" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.19" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-19 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-19-414a05e1 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-2.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-2.yaml new file mode 100644 index 000000000..e8ee62ef3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-2.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-2", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-2 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-2-c9b63fb5 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-20.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-20.yaml new file mode 100644 index 000000000..e5a75f84e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-20.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-20", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.20" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.20" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-20 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-20-51b7e182 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-21.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-21.yaml new file mode 100644 index 000000000..0ab7d9ab5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-21.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-21", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.21" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.21" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-21 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-21-1cb84e83 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-22.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-22.yaml new file mode 100644 index 000000000..9f783eebe --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-22.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-22", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.22" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.22" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-22 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-22-dbeb4bd3 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-23.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-23.yaml new file mode 100644 index 000000000..56ac18a58 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-23.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-23", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.23" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.23" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-23 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-23-5ac5dbb3 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-24.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-24.yaml new file mode 100644 index 000000000..50278884b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-24.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-24", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.24" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.24" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-24 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-24-3d896f90 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-25.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-25.yaml new file mode 100644 index 000000000..e59e2a906 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-25.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-25", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.25" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.25" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-25 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-25-6ad09dff + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-26.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-26.yaml new file mode 100644 index 000000000..72c3a580a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-26.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-26", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.26" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.26" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-26 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-26-3ecd3a8f + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-27.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-27.yaml new file mode 100644 index 000000000..4e7992f39 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-27.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-27", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.27" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.27" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-27 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-27-4391721d + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-28.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-28.yaml new file mode 100644 index 000000000..f736ac8db --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-28.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-28", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.28" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.28" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-28 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-28-84c5aeb8 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-29.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-29.yaml new file mode 100644 index 000000000..afa9c2b25 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-29.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-29", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.29" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.29" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-29 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-29-12991303 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-3.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-3.yaml new file mode 100644 index 000000000..193baf854 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-3.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-3", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-3 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-3-c094af95 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-30.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-30.yaml new file mode 100644 index 000000000..849eb09c0 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-30.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-30", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.30" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.30" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-30 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-30-023f6fbe + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-31.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-31.yaml new file mode 100644 index 000000000..b54a8ff12 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-31.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-31", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.31" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.31" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-31 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-31-349e40a0 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-32.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-32.yaml new file mode 100644 index 000000000..0fbde381a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-32.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-32", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.32" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.32" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-32 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-32-32835ba6 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-4.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-4.yaml new file mode 100644 index 000000000..56f9e4ec6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-4.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-4", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.4" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.4" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-4 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-4-2d6a8786 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-5.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-5.yaml new file mode 100644 index 000000000..c77da2e88 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-5.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-5", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.5" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.5" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-5 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-5-575ca7e2 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-6.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-6.yaml new file mode 100644 index 000000000..4d7d5324d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-6.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-6", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.6" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.6" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-6 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-6-4ca79564 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-7.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-7.yaml new file mode 100644 index 000000000..20b3e0428 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-7.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-7", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.7" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.7" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-7 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-7-d92f4f96 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-8.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-8.yaml new file mode 100644 index 000000000..4b0b8c797 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-8.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-8", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.8" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.8" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-8 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-8-b405a43c + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-9.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-9.yaml new file mode 100644 index 000000000..7b86ee3be --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/reports/worker-9.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-9", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.9" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.9" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-9 + storage.simplyblock.io/nodeprobe-run: discover-pci-09 + name: sb-nodeprobe-discover-pci-09-worker-9-40965de6 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/case.md b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/case.md new file mode 100644 index 000000000..a39ac2597 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/case.md @@ -0,0 +1,7 @@ +# PCI-10 + +**Mutation.** 5 workers: 3 on layout A, the other 2 each distinct + +**Expected.** 3 groups, of 3, 1, and 1 workers: partial homogeneity still groups what it can + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/expected-notes.txt new file mode 100644 index 000000000..6198c7d5b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-pci-10 in Draft: 5 workers with 16 nvme devices, in 3 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/expected.yaml new file mode 100644 index 000000000..4cc80e6aa --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/expected.yaml @@ -0,0 +1,50 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-pci-10 + name: discovered-discover-pci-10 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-pci-10-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - worker-02 + - worker-03 + - devices: + nvme: + - "0000:14:00.0" + - "0000:14:00.1" + mgmtInterface: eth0 + name: group-2-nvme-2x3T + workers: + - worker-04 + - devices: + nvme: + - "0000:15:00.0" + - "0000:15:00.1" + mgmtInterface: eth0 + name: group-3-nvme-2x3T + workers: + - worker-05 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/nodes.yaml new file mode 100644 index 000000000..4f0192d65 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/nodes.yaml @@ -0,0 +1,149 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-04 +spec: {} +status: + addresses: + - address: 10.10.10.4 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-05 +spec: {} +status: + addresses: + - address: 10.10.10.5 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/ops.yaml new file mode 100644 index 000000000..94d2cc640 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/ops.yaml @@ -0,0 +1,16 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-pci-10 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 + - worker-04 + - worker-05 diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/reports/worker-01.yaml new file mode 100644 index 000000000..070976800 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-pci-10 + name: sb-nodeprobe-discover-pci-10-worker-01-3142c336 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/reports/worker-02.yaml new file mode 100644 index 000000000..8457c4952 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-pci-10 + name: sb-nodeprobe-discover-pci-10-worker-02-a2462f84 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/reports/worker-03.yaml new file mode 100644 index 000000000..e7553bfc2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-pci-10 + name: sb-nodeprobe-discover-pci-10-worker-03-159e2243 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/reports/worker-04.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/reports/worker-04.yaml new file mode 100644 index 000000000..4f2ed92b4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/reports/worker-04.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-04", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.4" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.4" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:14:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:14:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-04 + storage.simplyblock.io/nodeprobe-run: discover-pci-10 + name: sb-nodeprobe-discover-pci-10-worker-04-0151e243 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/reports/worker-05.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/reports/worker-05.yaml new file mode 100644 index 000000000..734af5fa9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/reports/worker-05.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-05", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.5" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.5" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:15:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:15:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-05 + storage.simplyblock.io/nodeprobe-run: discover-pci-10 + name: sb-nodeprobe-discover-pci-10-worker-05-7c6684e5 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/case.md b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/case.md new file mode 100644 index 000000000..8fb5dcf46 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/case.md @@ -0,0 +1,7 @@ +# PCI-11 + +**Mutation.** 32 workers: 20 on layout A, 6 on layout B, 6 each distinct + +**Expected.** 8 groups, of 20, 6, and six of 1, numbered by their first worker's name + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/expected-notes.txt new file mode 100644 index 000000000..392203dbb --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-pci-11 in Draft: 32 workers with 116 nvme devices, in 8 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/expected.yaml new file mode 100644 index 000000000..128d9d9f6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/expected.yaml @@ -0,0 +1,114 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: discovered-discover-pci-11 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-pci-11-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - worker-02 + - worker-03 + - worker-04 + - worker-05 + - worker-06 + - worker-07 + - worker-08 + - worker-09 + - worker-10 + - worker-11 + - worker-12 + - worker-13 + - worker-14 + - worker-15 + - worker-16 + - worker-17 + - worker-18 + - worker-19 + - worker-20 + - devices: + nvme: + - 0000:3b:00.0 + - 0000:3c:00.0 + - 0000:d8:00.0 + - 0000:d9:00.0 + mgmtInterface: eth0 + name: group-2-nvme-4x3T + workers: + - worker-21 + - worker-22 + - worker-23 + - worker-24 + - worker-25 + - worker-26 + - devices: + nvme: + - 0000:2b:00.0 + - 0000:2b:00.1 + mgmtInterface: eth0 + name: group-3-nvme-2x3T + workers: + - worker-27 + - devices: + nvme: + - 0000:2c:00.0 + - 0000:2c:00.1 + mgmtInterface: eth0 + name: group-4-nvme-2x3T + workers: + - worker-28 + - devices: + nvme: + - 0000:2d:00.0 + - 0000:2d:00.1 + mgmtInterface: eth0 + name: group-5-nvme-2x3T + workers: + - worker-29 + - devices: + nvme: + - 0000:2e:00.0 + - 0000:2e:00.1 + mgmtInterface: eth0 + name: group-6-nvme-2x3T + workers: + - worker-30 + - devices: + nvme: + - 0000:2f:00.0 + - 0000:2f:00.1 + mgmtInterface: eth0 + name: group-7-nvme-2x3T + workers: + - worker-31 + - devices: + nvme: + - "0000:30:00.0" + - "0000:30:00.1" + mgmtInterface: eth0 + name: group-8-nvme-2x3T + workers: + - worker-32 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/nodes.yaml new file mode 100644 index 000000000..c25e3f497 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/nodes.yaml @@ -0,0 +1,959 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-04 +spec: {} +status: + addresses: + - address: 10.10.10.4 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-05 +spec: {} +status: + addresses: + - address: 10.10.10.5 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-06 +spec: {} +status: + addresses: + - address: 10.10.10.6 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-07 +spec: {} +status: + addresses: + - address: 10.10.10.7 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-08 +spec: {} +status: + addresses: + - address: 10.10.10.8 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-09 +spec: {} +status: + addresses: + - address: 10.10.10.9 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-10 +spec: {} +status: + addresses: + - address: 10.10.10.10 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-11 +spec: {} +status: + addresses: + - address: 10.10.10.11 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-12 +spec: {} +status: + addresses: + - address: 10.10.10.12 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-13 +spec: {} +status: + addresses: + - address: 10.10.10.13 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-14 +spec: {} +status: + addresses: + - address: 10.10.10.14 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-15 +spec: {} +status: + addresses: + - address: 10.10.10.15 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-16 +spec: {} +status: + addresses: + - address: 10.10.10.16 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-17 +spec: {} +status: + addresses: + - address: 10.10.10.17 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-18 +spec: {} +status: + addresses: + - address: 10.10.10.18 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-19 +spec: {} +status: + addresses: + - address: 10.10.10.19 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-20 +spec: {} +status: + addresses: + - address: 10.10.10.20 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-21 +spec: {} +status: + addresses: + - address: 10.10.10.21 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-22 +spec: {} +status: + addresses: + - address: 10.10.10.22 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-23 +spec: {} +status: + addresses: + - address: 10.10.10.23 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-24 +spec: {} +status: + addresses: + - address: 10.10.10.24 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-25 +spec: {} +status: + addresses: + - address: 10.10.10.25 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-26 +spec: {} +status: + addresses: + - address: 10.10.10.26 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-27 +spec: {} +status: + addresses: + - address: 10.10.10.27 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-28 +spec: {} +status: + addresses: + - address: 10.10.10.28 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-29 +spec: {} +status: + addresses: + - address: 10.10.10.29 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-30 +spec: {} +status: + addresses: + - address: 10.10.10.30 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-31 +spec: {} +status: + addresses: + - address: 10.10.10.31 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-32 +spec: {} +status: + addresses: + - address: 10.10.10.32 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/ops.yaml new file mode 100644 index 000000000..cd8eba5ce --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/ops.yaml @@ -0,0 +1,43 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-pci-11 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 + - worker-04 + - worker-05 + - worker-06 + - worker-07 + - worker-08 + - worker-09 + - worker-10 + - worker-11 + - worker-12 + - worker-13 + - worker-14 + - worker-15 + - worker-16 + - worker-17 + - worker-18 + - worker-19 + - worker-20 + - worker-21 + - worker-22 + - worker-23 + - worker-24 + - worker-25 + - worker-26 + - worker-27 + - worker-28 + - worker-29 + - worker-30 + - worker-31 + - worker-32 diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-01.yaml new file mode 100644 index 000000000..3d11f5f4b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-01-6e83338d + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-02.yaml new file mode 100644 index 000000000..9834432df --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-02-d4b3ca64 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-03.yaml new file mode 100644 index 000000000..07b8ac66c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-03-96f166fb + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-04.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-04.yaml new file mode 100644 index 000000000..5c5a8d8f7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-04.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-04", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.4" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.4" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-04 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-04-062c2ea5 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-05.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-05.yaml new file mode 100644 index 000000000..e899c6dce --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-05.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-05", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.5" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.5" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-05 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-05-38200420 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-06.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-06.yaml new file mode 100644 index 000000000..c5aa86f8a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-06.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-06", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.6" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.6" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-06 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-06-14bcd4ba + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-07.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-07.yaml new file mode 100644 index 000000000..f23b40e99 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-07.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-07", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.7" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.7" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-07 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-07-5468f366 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-08.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-08.yaml new file mode 100644 index 000000000..9230599ad --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-08.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-08", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.8" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.8" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-08 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-08-616b65fe + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-09.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-09.yaml new file mode 100644 index 000000000..4df1d61c0 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-09.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-09", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.9" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.9" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-09 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-09-8d8888e6 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-10.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-10.yaml new file mode 100644 index 000000000..5a376ffa8 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-10.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-10", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.10" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.10" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-10 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-10-ba32bbf3 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-11.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-11.yaml new file mode 100644 index 000000000..858db190b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-11.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-11", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.11" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.11" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-11 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-11-fd4699cd + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-12.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-12.yaml new file mode 100644 index 000000000..d48826bf7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-12.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-12", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.12" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.12" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-12 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-12-0568dde0 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-13.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-13.yaml new file mode 100644 index 000000000..f6301823d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-13.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-13", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.13" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.13" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-13 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-13-dac429bd + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-14.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-14.yaml new file mode 100644 index 000000000..ec0108e19 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-14.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-14", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.14" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.14" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-14 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-14-8891755e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-15.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-15.yaml new file mode 100644 index 000000000..02325d640 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-15.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-15", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.15" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.15" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-15 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-15-c8643383 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-16.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-16.yaml new file mode 100644 index 000000000..8ae0330cf --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-16.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-16", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.16" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.16" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-16 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-16-0536f5de + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-17.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-17.yaml new file mode 100644 index 000000000..c095e4e94 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-17.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-17", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.17" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.17" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-17 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-17-67edd34e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-18.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-18.yaml new file mode 100644 index 000000000..660435557 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-18.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-18", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.18" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.18" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-18 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-18-8430d2f2 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-19.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-19.yaml new file mode 100644 index 000000000..014d79901 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-19.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-19", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.19" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.19" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-19 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-19-d316694a + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-20.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-20.yaml new file mode 100644 index 000000000..9a0b5d6c3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-20.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-20", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.20" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.20" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-20 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-20-17777297 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-21.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-21.yaml new file mode 100644 index 000000000..0fa6a33a2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-21.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-21", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.21" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.21" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-21 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-21-7915b992 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-22.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-22.yaml new file mode 100644 index 000000000..769a53d52 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-22.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-22", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.22" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.22" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-22 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-22-1ba82e89 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-23.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-23.yaml new file mode 100644 index 000000000..5f36f3417 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-23.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-23", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.23" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.23" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-23 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-23-cb5096b2 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-24.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-24.yaml new file mode 100644 index 000000000..940f9c6de --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-24.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-24", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.24" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.24" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-24 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-24-9670e8ff + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-25.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-25.yaml new file mode 100644 index 000000000..b1f9ba2c7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-25.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-25", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.25" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.25" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-25 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-25-329e71e1 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-26.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-26.yaml new file mode 100644 index 000000000..f4482b5b6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-26.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-26", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.26" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.26" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:3b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:3c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:d8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:d9:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-26 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-26-619fbf66 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-27.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-27.yaml new file mode 100644 index 000000000..dc3014e88 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-27.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-27", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.27" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.27" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:2b:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:2b:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-27 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-27-0a1c64a3 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-28.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-28.yaml new file mode 100644 index 000000000..fa56f6796 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-28.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-28", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.28" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.28" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:2c:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:2c:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-28 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-28-c12fea3d + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-29.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-29.yaml new file mode 100644 index 000000000..3c573d7ac --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-29.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-29", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.29" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.29" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:2d:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:2d:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-29 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-29-bb47b73b + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-30.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-30.yaml new file mode 100644 index 000000000..37d250c95 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-30.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-30", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.30" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.30" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:2e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:2e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-30 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-30-08e2b00e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-31.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-31.yaml new file mode 100644 index 000000000..def6ad48b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-31.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-31", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.31" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.31" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:2f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:2f:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-31 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-31-3579672e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-32.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-32.yaml new file mode 100644 index 000000000..3d401bf11 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/reports/worker-32.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-32", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.32" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.32" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:30:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:30:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-32 + storage.simplyblock.io/nodeprobe-run: discover-pci-11 + name: sb-nodeprobe-discover-pci-11-worker-32-ee38071a + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/case.md b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/case.md new file mode 100644 index 000000000..380c6b634 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/case.md @@ -0,0 +1,7 @@ +# PCI-12 + +**Mutation.** 8 workers uniform but for one whose fourth disk is in another slot + +**Expected.** 2 groups, of 7 and 1. One slot moved is a second group, which is the guess a reviewer regroups + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/expected-notes.txt new file mode 100644 index 000000000..7ef2eea37 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-pci-12 in Draft: 8 workers with 32 nvme devices, in 2 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/expected.yaml new file mode 100644 index 000000000..06b888542 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/expected.yaml @@ -0,0 +1,48 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-pci-12 + name: discovered-discover-pci-12 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-pci-12-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - worker-02 + - worker-03 + - worker-04 + - worker-06 + - worker-07 + - worker-08 + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:c8:00.0 + mgmtInterface: eth0 + name: group-2-nvme-4x3T + workers: + - worker-05 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/nodes.yaml new file mode 100644 index 000000000..3822338a4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/nodes.yaml @@ -0,0 +1,239 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-04 +spec: {} +status: + addresses: + - address: 10.10.10.4 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-05 +spec: {} +status: + addresses: + - address: 10.10.10.5 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-06 +spec: {} +status: + addresses: + - address: 10.10.10.6 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-07 +spec: {} +status: + addresses: + - address: 10.10.10.7 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-08 +spec: {} +status: + addresses: + - address: 10.10.10.8 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/ops.yaml new file mode 100644 index 000000000..f05056da9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/ops.yaml @@ -0,0 +1,19 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-pci-12 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 + - worker-04 + - worker-05 + - worker-06 + - worker-07 + - worker-08 diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-01.yaml new file mode 100644 index 000000000..209ce7027 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-pci-12 + name: sb-nodeprobe-discover-pci-12-worker-01-f5390938 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-02.yaml new file mode 100644 index 000000000..f606fda11 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-pci-12 + name: sb-nodeprobe-discover-pci-12-worker-02-82b6ff29 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-03.yaml new file mode 100644 index 000000000..40cc4953a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-pci-12 + name: sb-nodeprobe-discover-pci-12-worker-03-de421b46 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-04.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-04.yaml new file mode 100644 index 000000000..ed29c442e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-04.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-04", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.4" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.4" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-04 + storage.simplyblock.io/nodeprobe-run: discover-pci-12 + name: sb-nodeprobe-discover-pci-12-worker-04-8b4e4a82 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-05.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-05.yaml new file mode 100644 index 000000000..5b344ad96 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-05.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-05", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.5" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.5" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:c8:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-05 + storage.simplyblock.io/nodeprobe-run: discover-pci-12 + name: sb-nodeprobe-discover-pci-12-worker-05-64df5835 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-06.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-06.yaml new file mode 100644 index 000000000..d2febc7a1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-06.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-06", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.6" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.6" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-06 + storage.simplyblock.io/nodeprobe-run: discover-pci-12 + name: sb-nodeprobe-discover-pci-12-worker-06-ea30430f + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-07.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-07.yaml new file mode 100644 index 000000000..9b7d68ec7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-07.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-07", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.7" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.7" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-07 + storage.simplyblock.io/nodeprobe-run: discover-pci-12 + name: sb-nodeprobe-discover-pci-12-worker-07-3b913cad + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-08.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-08.yaml new file mode 100644 index 000000000..4e159d7d2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/reports/worker-08.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-08", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.8" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.8" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-08 + storage.simplyblock.io/nodeprobe-run: discover-pci-12 + name: sb-nodeprobe-discover-pci-12-worker-08-d87f60f9 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/case.md b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/case.md new file mode 100644 index 000000000..4c9cd0ef6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/case.md @@ -0,0 +1,9 @@ +# PCI-13 + +**Mutation.** 2 workers on one layout carrying 2 TiB and 4 TiB disks + +**Expected.** **Contested.** One group, named `group-1-nvme-2x2T` for the first worker's capacity. See §14, gap G-13 + +**Harness.** `CM` + +**Gap.** G-13. This case records what the generator does today, so that the day it changes the diff is the finding. diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/expected-notes.txt new file mode 100644 index 000000000..db7a0b5e9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-pci-13 in Draft: 2 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/expected.yaml new file mode 100644 index 000000000..d288fd32d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/expected.yaml @@ -0,0 +1,31 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-pci-13 + name: discovered-discover-pci-13 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-pci-13-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + mgmtInterface: eth0 + name: group-1-nvme-2x2T + workers: + - worker-01 + - worker-02 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/nodes.yaml new file mode 100644 index 000000000..1fba89979 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/nodes.yaml @@ -0,0 +1,59 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/ops.yaml new file mode 100644 index 000000000..0ad6333fb --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/ops.yaml @@ -0,0 +1,13 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-pci-13 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/reports/worker-01.yaml new file mode 100644 index 000000000..d5c2898a6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/reports/worker-01.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 2199023255552, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-pci-13 + name: sb-nodeprobe-discover-pci-13-worker-01-1f5e0512 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/reports/worker-02.yaml new file mode 100644 index 000000000..f4c03462c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/reports/worker-02.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 4398046511104, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 4398046511104, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-pci-13 + name: sb-nodeprobe-discover-pci-13-worker-02-c4a760f1 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/case.md b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/case.md new file mode 100644 index 000000000..23318eb00 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/case.md @@ -0,0 +1,7 @@ +# PCI-14 + +**Mutation.** 3 workers, the third holding 3 of the 4 slots the others hold + +**Expected.** 2 groups: the address set matches whole or not at all, never as a subset + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/expected-notes.txt new file mode 100644 index 000000000..5704abd9d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-pci-14 in Draft: 3 workers with 11 nvme devices, in 2 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/expected.yaml new file mode 100644 index 000000000..4cc7b03de --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/expected.yaml @@ -0,0 +1,42 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-pci-14 + name: discovered-discover-pci-14 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-pci-14-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - worker-02 + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + mgmtInterface: eth0 + name: group-2-nvme-3x3T + workers: + - worker-03 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/nodes.yaml new file mode 100644 index 000000000..98de08ee9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/nodes.yaml @@ -0,0 +1,89 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/ops.yaml new file mode 100644 index 000000000..f4bdc243b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/ops.yaml @@ -0,0 +1,14 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-pci-14 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/reports/worker-01.yaml new file mode 100644 index 000000000..a698d3946 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-pci-14 + name: sb-nodeprobe-discover-pci-14-worker-01-79d7cf06 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/reports/worker-02.yaml new file mode 100644 index 000000000..207933995 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-pci-14 + name: sb-nodeprobe-discover-pci-14-worker-02-ab8ba038 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/reports/worker-03.yaml new file mode 100644 index 000000000..463df2a93 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/reports/worker-03.yaml @@ -0,0 +1,163 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-pci-14 + name: sb-nodeprobe-discover-pci-14-worker-03-c3c43f8a + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/case.md b/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/case.md new file mode 100644 index 000000000..77d1e6ce5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/case.md @@ -0,0 +1,7 @@ +# ROLE-01 + +**Mutation.** 3 workers, no role labels + +**Expected.** One node set, `discovered` + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/expected-notes.txt new file mode 100644 index 000000000..a0ddbb388 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-role-01 in Draft: 3 workers with 12 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/expected.yaml new file mode 100644 index 000000000..0f99b2543 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/expected.yaml @@ -0,0 +1,34 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-role-01 + name: discovered-discover-role-01 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-role-01-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - worker-02 + - worker-03 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/nodes.yaml new file mode 100644 index 000000000..98de08ee9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/nodes.yaml @@ -0,0 +1,89 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/ops.yaml new file mode 100644 index 000000000..32aadc222 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/ops.yaml @@ -0,0 +1,14 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-role-01 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/reports/worker-01.yaml new file mode 100644 index 000000000..aa8664b90 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-role-01 + name: sb-nodeprobe-discover-role-01-worker-01-8b820cb9 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/reports/worker-02.yaml new file mode 100644 index 000000000..c5cc2270f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-role-01 + name: sb-nodeprobe-discover-role-01-worker-02-ed19770e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/reports/worker-03.yaml new file mode 100644 index 000000000..e7e0eb665 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-role-01 + name: sb-nodeprobe-discover-role-01-worker-03-3d6c9982 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/case.md b/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/case.md new file mode 100644 index 000000000..dca6d94b2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/case.md @@ -0,0 +1,7 @@ +# ROLE-02 + +**Mutation.** 3 `infra` and 3 `worker` nodes, identical hardware + +**Expected.** 2 node sets, `infra` first, one hardware group split across them + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/expected-notes.txt new file mode 100644 index 000000000..eae4e4cb3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-role-02 in Draft: 6 workers with 24 nvme devices, in 2 group(s) across 2 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/expected.yaml new file mode 100644 index 000000000..d2739d403 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/expected.yaml @@ -0,0 +1,48 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-role-02 + name: discovered-discover-role-02 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-role-02-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - worker-02 + - worker-03 + name: infra + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-04 + - worker-05 + - worker-06 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/nodes.yaml new file mode 100644 index 000000000..08079b60b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/nodes.yaml @@ -0,0 +1,185 @@ +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/infra: "" + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/infra: "" + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/infra: "" + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-04 +spec: {} +status: + addresses: + - address: 10.10.10.4 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-05 +spec: {} +status: + addresses: + - address: 10.10.10.5 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-06 +spec: {} +status: + addresses: + - address: 10.10.10.6 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/ops.yaml new file mode 100644 index 000000000..6c9ddad3a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/ops.yaml @@ -0,0 +1,17 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-role-02 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 + - worker-04 + - worker-05 + - worker-06 diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/reports/worker-01.yaml new file mode 100644 index 000000000..b9d658258 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-role-02 + name: sb-nodeprobe-discover-role-02-worker-01-632eb4d8 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/reports/worker-02.yaml new file mode 100644 index 000000000..987d2f4ca --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-role-02 + name: sb-nodeprobe-discover-role-02-worker-02-011a35be + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/reports/worker-03.yaml new file mode 100644 index 000000000..e63c0c323 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-role-02 + name: sb-nodeprobe-discover-role-02-worker-03-305ec053 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/reports/worker-04.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/reports/worker-04.yaml new file mode 100644 index 000000000..8ba875ea4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/reports/worker-04.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-04", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.4" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.4" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-04 + storage.simplyblock.io/nodeprobe-run: discover-role-02 + name: sb-nodeprobe-discover-role-02-worker-04-56e3e015 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/reports/worker-05.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/reports/worker-05.yaml new file mode 100644 index 000000000..070bba0a9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/reports/worker-05.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-05", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.5" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.5" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-05 + storage.simplyblock.io/nodeprobe-run: discover-role-02 + name: sb-nodeprobe-discover-role-02-worker-05-db79cddd + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/reports/worker-06.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/reports/worker-06.yaml new file mode 100644 index 000000000..8004ac10d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/reports/worker-06.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-06", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.6" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.6" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-06 + storage.simplyblock.io/nodeprobe-run: discover-role-02 + name: sb-nodeprobe-discover-role-02-worker-06-03c47c22 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/case.md b/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/case.md new file mode 100644 index 000000000..ba93a3c92 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/case.md @@ -0,0 +1,7 @@ +# ROLE-03 + +**Mutation.** A `control-plane` node with free disks beside 2 workers + +**Expected.** A third node set, `control-plane`, ordered last + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/expected-notes.txt new file mode 100644 index 000000000..7249e8bd5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-role-03 in Draft: 3 workers with 12 nvme devices, in 2 group(s) across 2 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/expected.yaml new file mode 100644 index 000000000..2ca88f1e1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/expected.yaml @@ -0,0 +1,45 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-role-03 + name: discovered-discover-role-03 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-role-03-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-02 + - worker-03 + name: discovered + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: control-plane +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/nodes.yaml new file mode 100644 index 000000000..aa5311869 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/nodes.yaml @@ -0,0 +1,94 @@ +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/control-plane: "" + name: worker-01 +spec: + taints: + - effect: NoSchedule + key: node-role.kubernetes.io/control-plane +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/ops.yaml new file mode 100644 index 000000000..b67851425 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/ops.yaml @@ -0,0 +1,16 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-role-03 + namespace: simplyblock +spec: + action: Discover + discover: + enableControlPlaneNodes: true +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/reports/worker-01.yaml new file mode 100644 index 000000000..74a82a8cb --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-role-03 + name: sb-nodeprobe-discover-role-03-worker-01-01eb86a8 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/reports/worker-02.yaml new file mode 100644 index 000000000..18cd1c7f8 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-role-03 + name: sb-nodeprobe-discover-role-03-worker-02-4ddb8fb7 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/reports/worker-03.yaml new file mode 100644 index 000000000..3dbf5edb6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-role-03 + name: sb-nodeprobe-discover-role-03-worker-03-dbecf7f8 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/case.md b/operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/case.md new file mode 100644 index 000000000..bd7069532 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/case.md @@ -0,0 +1,7 @@ +# ROLE-04 + +**Mutation.** A node labeled `infra` and `worker` + +**Expected.** `infra` + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/expected-notes.txt new file mode 100644 index 000000000..4c58f215c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-role-04 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/expected.yaml new file mode 100644 index 000000000..a3d2e8b09 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-role-04 + name: discovered-discover-role-04 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-role-04-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: infra +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/nodes.yaml new file mode 100644 index 000000000..9c7d69725 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/nodes.yaml @@ -0,0 +1,32 @@ +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/infra: "" + node-role.kubernetes.io/worker: "" + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/ops.yaml new file mode 100644 index 000000000..115e7df72 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-role-04 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/reports/worker-01.yaml new file mode 100644 index 000000000..310e4a254 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-role-04 + name: sb-nodeprobe-discover-role-04-worker-01-b3bbaca7 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/case.md b/operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/case.md new file mode 100644 index 000000000..cbf67b08d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/case.md @@ -0,0 +1,7 @@ +# ROLE-05 + +**Mutation.** A node labeled `control-plane` and `worker` + +**Expected.** `control-plane`: the most restrictive role wins + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/expected-notes.txt new file mode 100644 index 000000000..ecaf23632 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-role-05 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/expected.yaml new file mode 100644 index 000000000..25cdcaf08 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-role-05 + name: discovered-discover-role-05 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-role-05-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: control-plane +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/nodes.yaml new file mode 100644 index 000000000..55b2e7b32 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/nodes.yaml @@ -0,0 +1,32 @@ +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/control-plane: "" + node-role.kubernetes.io/worker: "" + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/ops.yaml new file mode 100644 index 000000000..86195b073 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/ops.yaml @@ -0,0 +1,14 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-role-05 + namespace: simplyblock +spec: + action: Discover + discover: + enableControlPlaneNodes: true +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/reports/worker-01.yaml new file mode 100644 index 000000000..009c15f0d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-role-05 + name: sb-nodeprobe-discover-role-05-worker-01-817ba5f1 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/case.md b/operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/case.md new file mode 100644 index 000000000..32e6be6db --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/case.md @@ -0,0 +1,7 @@ +# ROLE-06 + +**Mutation.** A node labeled `master` + +**Expected.** `control-plane`: both spellings are read + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/expected-notes.txt new file mode 100644 index 000000000..24a1efa18 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-role-06 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/expected.yaml new file mode 100644 index 000000000..83e1b597a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-role-06 + name: discovered-discover-role-06 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-role-06-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: control-plane +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/nodes.yaml new file mode 100644 index 000000000..a5c618731 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/nodes.yaml @@ -0,0 +1,31 @@ +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/master: "" + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/ops.yaml new file mode 100644 index 000000000..c9b509d5e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/ops.yaml @@ -0,0 +1,14 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-role-06 + namespace: simplyblock +spec: + action: Discover + discover: + enableControlPlaneNodes: true +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/reports/worker-01.yaml new file mode 100644 index 000000000..fc7c98ee5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-role-06 + name: sb-nodeprobe-discover-role-06-worker-01-87864f15 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-07-a-plan-with-no-node-objects/case.md b/operator/internal/controllers/deployment/testdata/discovery/role/role-07-a-plan-with-no-node-objects/case.md new file mode 100644 index 000000000..387379ba4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-07-a-plan-with-no-node-objects/case.md @@ -0,0 +1,9 @@ +# ROLE-07 + +**Mutation.** No `nodes.yaml` at all + +**Expected.** Every worker a `Worker`. The interface is chosen with no address hint + +**Harness.** `GO` + +**Note.** Driven in the discovery package with a Planner carrying no KubeNodes, because the controller always reads the node objects for the workers its run settled on. diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/case.md b/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/case.md new file mode 100644 index 000000000..ca09e681a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/case.md @@ -0,0 +1,9 @@ +# ROLE-08 + +**Mutation.** A cordoned node and a tainted node, both with free disks + +**Expected.** Both reach the draft. See §14, gap G-8 + +**Harness.** `CM` + +**Gap.** G-8. This case records what the generator does today, so that the day it changes the diff is the finding. diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/expected-notes.txt new file mode 100644 index 000000000..b1053274b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-role-08 in Draft: 2 workers with 8 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/expected.yaml new file mode 100644 index 000000000..348d54ef8 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/expected.yaml @@ -0,0 +1,33 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-role-08 + name: discovered-discover-role-08 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-role-08-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - worker-02 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/nodes.yaml new file mode 100644 index 000000000..55019f4b2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/nodes.yaml @@ -0,0 +1,64 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: + unschedulable: true +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: + taints: + - effect: NoExecute + key: storage + value: drained +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/ops.yaml new file mode 100644 index 000000000..cce05a650 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/ops.yaml @@ -0,0 +1,13 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-role-08 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/reports/worker-01.yaml new file mode 100644 index 000000000..0e8a1a9fb --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-role-08 + name: sb-nodeprobe-discover-role-08-worker-01-0acca693 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/reports/worker-02.yaml new file mode 100644 index 000000000..a94d60211 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-role-08 + name: sb-nodeprobe-discover-role-08-worker-02-557b67db + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/case.md b/operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/case.md new file mode 100644 index 000000000..0578f9986 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/case.md @@ -0,0 +1,7 @@ +# ROLE-09 + +**Mutation.** A node carrying an unrecognized role label + +**Expected.** Treated as a worker: an unknown role is not a refusal + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/expected-notes.txt new file mode 100644 index 000000000..48561a796 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-role-09 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/expected.yaml new file mode 100644 index 000000000..f1412054a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-role-09 + name: discovered-discover-role-09 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-role-09-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/nodes.yaml new file mode 100644 index 000000000..4010d31c2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/nodes.yaml @@ -0,0 +1,31 @@ +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/database: "" + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/ops.yaml new file mode 100644 index 000000000..545113731 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-role-09 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/reports/worker-01.yaml new file mode 100644 index 000000000..6afd80765 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-role-09 + name: sb-nodeprobe-discover-role-09-worker-01-aa562200 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/case.md b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/case.md new file mode 100644 index 000000000..d920acfd6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/case.md @@ -0,0 +1,7 @@ +# ROLE-10 + +**Mutation.** 32 workers across `infra`, `worker`, and `control-plane` + +**Expected.** 3 node sets, each holding its role's share of every hardware group + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/expected-notes.txt new file mode 100644 index 000000000..6e192dc9d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-role-10 in Draft: 32 workers with 128 nvme devices, in 3 group(s) across 3 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/expected.yaml new file mode 100644 index 000000000..fb959f1c5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/expected.yaml @@ -0,0 +1,85 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: discovered-discover-role-10 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-role-10-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - worker-02 + - worker-03 + - worker-04 + name: infra + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-08 + - worker-09 + - worker-10 + - worker-11 + - worker-12 + - worker-13 + - worker-14 + - worker-15 + - worker-16 + - worker-17 + - worker-18 + - worker-19 + - worker-20 + - worker-21 + - worker-22 + - worker-23 + - worker-24 + - worker-25 + - worker-26 + - worker-27 + - worker-28 + - worker-29 + - worker-30 + - worker-31 + - worker-32 + name: discovered + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-05 + - worker-06 + - worker-07 + name: control-plane +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/nodes.yaml new file mode 100644 index 000000000..c4ebe6637 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/nodes.yaml @@ -0,0 +1,1023 @@ +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/infra: "" + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/infra: "" + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/infra: "" + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/infra: "" + name: worker-04 +spec: {} +status: + addresses: + - address: 10.10.10.4 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/control-plane: "" + name: worker-05 +spec: {} +status: + addresses: + - address: 10.10.10.5 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/control-plane: "" + name: worker-06 +spec: {} +status: + addresses: + - address: 10.10.10.6 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/control-plane: "" + name: worker-07 +spec: {} +status: + addresses: + - address: 10.10.10.7 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/worker: "" + name: worker-08 +spec: {} +status: + addresses: + - address: 10.10.10.8 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/worker: "" + name: worker-09 +spec: {} +status: + addresses: + - address: 10.10.10.9 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/worker: "" + name: worker-10 +spec: {} +status: + addresses: + - address: 10.10.10.10 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/worker: "" + name: worker-11 +spec: {} +status: + addresses: + - address: 10.10.10.11 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/worker: "" + name: worker-12 +spec: {} +status: + addresses: + - address: 10.10.10.12 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/worker: "" + name: worker-13 +spec: {} +status: + addresses: + - address: 10.10.10.13 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/worker: "" + name: worker-14 +spec: {} +status: + addresses: + - address: 10.10.10.14 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/worker: "" + name: worker-15 +spec: {} +status: + addresses: + - address: 10.10.10.15 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/worker: "" + name: worker-16 +spec: {} +status: + addresses: + - address: 10.10.10.16 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/worker: "" + name: worker-17 +spec: {} +status: + addresses: + - address: 10.10.10.17 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/worker: "" + name: worker-18 +spec: {} +status: + addresses: + - address: 10.10.10.18 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/worker: "" + name: worker-19 +spec: {} +status: + addresses: + - address: 10.10.10.19 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/worker: "" + name: worker-20 +spec: {} +status: + addresses: + - address: 10.10.10.20 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/worker: "" + name: worker-21 +spec: {} +status: + addresses: + - address: 10.10.10.21 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/worker: "" + name: worker-22 +spec: {} +status: + addresses: + - address: 10.10.10.22 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/worker: "" + name: worker-23 +spec: {} +status: + addresses: + - address: 10.10.10.23 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/worker: "" + name: worker-24 +spec: {} +status: + addresses: + - address: 10.10.10.24 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/worker: "" + name: worker-25 +spec: {} +status: + addresses: + - address: 10.10.10.25 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/worker: "" + name: worker-26 +spec: {} +status: + addresses: + - address: 10.10.10.26 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/worker: "" + name: worker-27 +spec: {} +status: + addresses: + - address: 10.10.10.27 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/worker: "" + name: worker-28 +spec: {} +status: + addresses: + - address: 10.10.10.28 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/worker: "" + name: worker-29 +spec: {} +status: + addresses: + - address: 10.10.10.29 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/worker: "" + name: worker-30 +spec: {} +status: + addresses: + - address: 10.10.10.30 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/worker: "" + name: worker-31 +spec: {} +status: + addresses: + - address: 10.10.10.31 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + labels: + node-role.kubernetes.io/worker: "" + name: worker-32 +spec: {} +status: + addresses: + - address: 10.10.10.32 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/ops.yaml new file mode 100644 index 000000000..837259e21 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/ops.yaml @@ -0,0 +1,45 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-role-10 + namespace: simplyblock +spec: + action: Discover + discover: + enableControlPlaneNodes: true +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 + - worker-04 + - worker-05 + - worker-06 + - worker-07 + - worker-08 + - worker-09 + - worker-10 + - worker-11 + - worker-12 + - worker-13 + - worker-14 + - worker-15 + - worker-16 + - worker-17 + - worker-18 + - worker-19 + - worker-20 + - worker-21 + - worker-22 + - worker-23 + - worker-24 + - worker-25 + - worker-26 + - worker-27 + - worker-28 + - worker-29 + - worker-30 + - worker-31 + - worker-32 diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-01.yaml new file mode 100644 index 000000000..c355c8770 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-01-df572e68 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-02.yaml new file mode 100644 index 000000000..280282adf --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-02-c6e18385 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-03.yaml new file mode 100644 index 000000000..96ca89a87 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-03-f2c9ce43 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-04.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-04.yaml new file mode 100644 index 000000000..0cfee8326 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-04.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-04", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.4" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.4" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-04 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-04-9f5b7430 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-05.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-05.yaml new file mode 100644 index 000000000..47b59e7e0 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-05.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-05", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.5" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.5" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-05 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-05-a5c7bec9 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-06.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-06.yaml new file mode 100644 index 000000000..d6b47190c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-06.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-06", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.6" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.6" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-06 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-06-034ebbf6 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-07.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-07.yaml new file mode 100644 index 000000000..499163460 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-07.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-07", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.7" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.7" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-07 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-07-d6495312 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-08.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-08.yaml new file mode 100644 index 000000000..1380f4b77 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-08.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-08", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.8" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.8" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-08 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-08-a613e069 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-09.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-09.yaml new file mode 100644 index 000000000..587976ea4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-09.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-09", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.9" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.9" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-09 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-09-3119ba29 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-10.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-10.yaml new file mode 100644 index 000000000..639623bc5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-10.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-10", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.10" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.10" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-10 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-10-6e955126 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-11.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-11.yaml new file mode 100644 index 000000000..c98764bc9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-11.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-11", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.11" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.11" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-11 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-11-83a80602 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-12.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-12.yaml new file mode 100644 index 000000000..758dce96d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-12.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-12", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.12" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.12" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-12 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-12-e56df2a1 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-13.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-13.yaml new file mode 100644 index 000000000..41396c40b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-13.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-13", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.13" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.13" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-13 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-13-513010af + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-14.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-14.yaml new file mode 100644 index 000000000..631ee12fe --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-14.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-14", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.14" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.14" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-14 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-14-51233507 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-15.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-15.yaml new file mode 100644 index 000000000..2330232fe --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-15.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-15", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.15" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.15" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-15 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-15-c8ed93bd + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-16.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-16.yaml new file mode 100644 index 000000000..e362aa6f9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-16.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-16", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.16" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.16" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-16 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-16-b675852d + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-17.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-17.yaml new file mode 100644 index 000000000..9a4655d14 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-17.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-17", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.17" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.17" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-17 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-17-7a5d5c49 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-18.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-18.yaml new file mode 100644 index 000000000..c9afe982e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-18.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-18", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.18" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.18" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-18 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-18-1ab16bba + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-19.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-19.yaml new file mode 100644 index 000000000..92ad1fe7e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-19.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-19", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.19" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.19" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-19 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-19-9fbc0b2a + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-20.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-20.yaml new file mode 100644 index 000000000..4e846651f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-20.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-20", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.20" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.20" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-20 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-20-6f47e2a9 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-21.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-21.yaml new file mode 100644 index 000000000..5dd47e9ab --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-21.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-21", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.21" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.21" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-21 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-21-5d876b97 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-22.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-22.yaml new file mode 100644 index 000000000..e2fb2543e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-22.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-22", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.22" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.22" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-22 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-22-3ca79df9 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-23.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-23.yaml new file mode 100644 index 000000000..864bacb0d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-23.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-23", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.23" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.23" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-23 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-23-b62949b8 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-24.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-24.yaml new file mode 100644 index 000000000..483b1ebbe --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-24.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-24", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.24" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.24" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-24 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-24-0a56beff + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-25.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-25.yaml new file mode 100644 index 000000000..211f885a4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-25.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-25", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.25" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.25" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-25 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-25-c2d8d806 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-26.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-26.yaml new file mode 100644 index 000000000..e284a3505 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-26.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-26", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.26" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.26" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-26 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-26-0defbc59 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-27.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-27.yaml new file mode 100644 index 000000000..f03decc32 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-27.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-27", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.27" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.27" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-27 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-27-af1ec3ed + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-28.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-28.yaml new file mode 100644 index 000000000..ca72d0f63 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-28.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-28", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.28" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.28" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-28 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-28-3834c2ff + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-29.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-29.yaml new file mode 100644 index 000000000..8cc109f15 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-29.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-29", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.29" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.29" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-29 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-29-7cdbe886 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-30.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-30.yaml new file mode 100644 index 000000000..4c73e3918 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-30.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-30", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.30" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.30" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-30 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-30-07d6b73c + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-31.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-31.yaml new file mode 100644 index 000000000..2d7a7599b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-31.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-31", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.31" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.31" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-31 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-31-176e702a + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-32.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-32.yaml new file mode 100644 index 000000000..fc169cdec --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/reports/worker-32.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-32", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.32" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.32" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-32 + storage.simplyblock.io/nodeprobe-run: discover-role-10 + name: sb-nodeprobe-discover-role-10-worker-32-627d7353 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/case.md b/operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/case.md new file mode 100644 index 000000000..d4ca7e2a2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/case.md @@ -0,0 +1,7 @@ +# SIZE-01 + +**Mutation.** 8 logical CPUs, 4 physical cores on the chosen node + +**Expected.** `vcpuCount: 4`, with the ordinary derivation note, not the floor note + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/expected-notes.txt new file mode 100644 index 000000000..1ca15afb2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 4, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-size-01 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/expected.yaml new file mode 100644 index 000000000..8058d192d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-size-01 + name: discovered-discover-size-01 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-size-01-cluster + vcpuCount: 4 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/ops.yaml new file mode 100644 index 000000000..4ae296544 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-size-01 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/reports/worker-01.yaml new file mode 100644 index 000000000..c7f617203 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/reports/worker-01.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 8, + "physicalCores": 4, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7 + ], + "physicalCores": 4 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-size-01 + name: sb-nodeprobe-discover-size-01-worker-01-bb4ffce3 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/case.md b/operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/case.md new file mode 100644 index 000000000..d30ce164f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/case.md @@ -0,0 +1,7 @@ +# SIZE-02 + +**Mutation.** 2 logical CPUs, 2 physical cores + +**Expected.** `vcpuCount: 4` with the note saying the cluster asks for more than the worker has. Not a refusal + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/expected-notes.txt new file mode 100644 index 000000000..34b160f05 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is the API's minimum of 4, and the smallest placement (worker-01) has only 2 cores, so this cluster asks for more than that worker has +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-size-02 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/expected.yaml new file mode 100644 index 000000000..9d8686a2b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-size-02 + name: discovered-discover-size-02 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-size-02-cluster + vcpuCount: 4 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/ops.yaml new file mode 100644 index 000000000..d0fee5c1f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-size-02 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/reports/worker-01.yaml new file mode 100644 index 000000000..be1f9ff53 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/reports/worker-01.yaml @@ -0,0 +1,140 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 2, + "physicalCores": 2, + "sockets": 1, + "threadsPerCore": 1, + "hyperThreading": false, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1 + ], + "physicalCores": 2 + } + ] + }, + "memory": { + "totalBytes": 17179869184, + "freeBytes": 3758096384, + "availableBytes": 15032385536, + "numaNodes": [ + { + "node": 0, + "totalBytes": 17179869184, + "freeBytes": 8589934592 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-size-02 + name: sb-nodeprobe-discover-size-02-worker-01-f0b8ccb5 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/case.md b/operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/case.md new file mode 100644 index 000000000..02b43615a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/case.md @@ -0,0 +1,9 @@ +# SIZE-03 + +**Mutation.** 196 logical CPUs, 98 physical cores, 1 memory node, 1 TiB RAM + +**Expected.** `vcpuCount: 98`, the whole machine. See §14, gap G-3 + +**Harness.** `CM` + +**Gap.** G-3. This case records what the generator does today, so that the day it changes the diff is the finding. diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/expected-notes.txt new file mode 100644 index 000000000..f08a46bea --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 98, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 256G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-size-03 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/expected.yaml new file mode 100644 index 000000000..c5aa3541f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-size-03 + name: discovered-discover-size-03 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 256G + name: discovered-discover-size-03-cluster + vcpuCount: 98 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/ops.yaml new file mode 100644 index 000000000..284e49da1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-size-03 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/reports/worker-01.yaml new file mode 100644 index 000000000..8f880a5bf --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/reports/worker-01.yaml @@ -0,0 +1,329 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 196, + "physicalCores": 98, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31, + 32, + 33, + 34, + 35, + 36, + 37, + 38, + 39, + 40, + 41, + 42, + 43, + 44, + 45, + 46, + 47, + 48, + 49, + 50, + 51, + 52, + 53, + 54, + 55, + 56, + 57, + 58, + 59, + 60, + 61, + 62, + 63, + 64, + 65, + 66, + 67, + 68, + 69, + 70, + 71, + 72, + 73, + 74, + 75, + 76, + 77, + 78, + 79, + 80, + 81, + 82, + 83, + 84, + 85, + 86, + 87, + 88, + 89, + 90, + 91, + 92, + 93, + 94, + 95, + 96, + 97, + 98, + 99, + 100, + 101, + 102, + 103, + 104, + 105, + 106, + 107, + 108, + 109, + 110, + 111, + 112, + 113, + 114, + 115, + 116, + 117, + 118, + 119, + 120, + 121, + 122, + 123, + 124, + 125, + 126, + 127, + 128, + 129, + 130, + 131, + 132, + 133, + 134, + 135, + 136, + 137, + 138, + 139, + 140, + 141, + 142, + 143, + 144, + 145, + 146, + 147, + 148, + 149, + 150, + 151, + 152, + 153, + 154, + 155, + 156, + 157, + 158, + 159, + 160, + 161, + 162, + 163, + 164, + 165, + 166, + 167, + 168, + 169, + 170, + 171, + 172, + 173, + 174, + 175, + 176, + 177, + 178, + 179, + 180, + 181, + 182, + 183, + 184, + 185, + 186, + 187, + 188, + 189, + 190, + 191, + 192, + 193, + 194, + 195 + ], + "physicalCores": 98 + } + ] + }, + "memory": { + "totalBytes": 1099511627776, + "freeBytes": 268435456000, + "availableBytes": 1073741824000, + "numaNodes": [ + { + "node": 0, + "totalBytes": 1099511627776, + "freeBytes": 549755813888 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 256, + "free": 256, + "numaNodes": [ + { + "node": 0, + "total": 256, + "free": 256 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-size-03 + name: sb-nodeprobe-discover-size-03-worker-01-21661b72 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/case.md b/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/case.md new file mode 100644 index 000000000..714a71393 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/case.md @@ -0,0 +1,7 @@ +# SIZE-04 + +**Mutation.** 196 logical CPUs across 2 nodes, 1 TiB RAM, 512 GiB per node + +**Expected.** `vcpuCount: 49`, the chosen node's cores + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/expected-notes.txt new file mode 100644 index 000000000..73842c79a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 49, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 128G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-size-04 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 2 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/expected-refusals.txt new file mode 100644 index 000000000..e5d0ed2ac --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/expected-refusals.txt @@ -0,0 +1,2 @@ +worker-01/nvme2n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 1, so it was chosen +worker-01/nvme3n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 1, so it was chosen diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/expected.yaml new file mode 100644 index 000000000..a514f920d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/expected.yaml @@ -0,0 +1,30 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-size-04 + name: discovered-discover-size-04 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 128G + name: discovered-discover-size-04-cluster + vcpuCount: 49 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + mgmtInterface: eth0 + name: group-1-nvme-2x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/ops.yaml new file mode 100644 index 000000000..bb6e27532 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-size-04 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/reports/worker-01.yaml new file mode 100644 index 000000000..7dc4afbd1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/reports/worker-01.yaml @@ -0,0 +1,345 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 196, + "physicalCores": 98, + "sockets": 2, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31, + 32, + 33, + 34, + 35, + 36, + 37, + 38, + 39, + 40, + 41, + 42, + 43, + 44, + 45, + 46, + 47, + 48, + 49, + 50, + 51, + 52, + 53, + 54, + 55, + 56, + 57, + 58, + 59, + 60, + 61, + 62, + 63, + 64, + 65, + 66, + 67, + 68, + 69, + 70, + 71, + 72, + 73, + 74, + 75, + 76, + 77, + 78, + 79, + 80, + 81, + 82, + 83, + 84, + 85, + 86, + 87, + 88, + 89, + 90, + 91, + 92, + 93, + 94, + 95, + 96, + 97 + ], + "physicalCores": 49 + }, + { + "node": 1, + "onlineCPUs": [ + 98, + 99, + 100, + 101, + 102, + 103, + 104, + 105, + 106, + 107, + 108, + 109, + 110, + 111, + 112, + 113, + 114, + 115, + 116, + 117, + 118, + 119, + 120, + 121, + 122, + 123, + 124, + 125, + 126, + 127, + 128, + 129, + 130, + 131, + 132, + 133, + 134, + 135, + 136, + 137, + 138, + 139, + 140, + 141, + 142, + 143, + 144, + 145, + 146, + 147, + 148, + 149, + 150, + 151, + 152, + 153, + 154, + 155, + 156, + 157, + 158, + 159, + 160, + 161, + 162, + 163, + 164, + 165, + 166, + 167, + 168, + 169, + 170, + 171, + 172, + 173, + 174, + 175, + 176, + 177, + 178, + 179, + 180, + 181, + 182, + 183, + 184, + 185, + 186, + 187, + 188, + 189, + 190, + 191, + 192, + 193, + 194, + 195 + ], + "physicalCores": 49 + } + ] + }, + "memory": { + "totalBytes": 1099511627776, + "freeBytes": 268435456000, + "availableBytes": 1073741824000, + "numaNodes": [ + { + "node": 0, + "totalBytes": 549755813888, + "freeBytes": 274877906944 + }, + { + "node": 1, + "totalBytes": 549755813888, + "freeBytes": 274877906944 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 256, + "free": 256, + "numaNodes": [ + { + "node": 0, + "total": 128, + "free": 128 + }, + { + "node": 1, + "total": 128, + "free": 128 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:7e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:7e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-size-04 + name: sb-nodeprobe-discover-size-04-worker-01-0c753d3e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/case.md b/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/case.md new file mode 100644 index 000000000..8af5e9ec5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/case.md @@ -0,0 +1,7 @@ +# SIZE-05 + +**Mutation.** A fleet whose chosen nodes carry 4, 16, and 49 cores + +**Expected.** `vcpuCount: 4`, the note naming the smallest worker + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/expected-notes.txt new file mode 100644 index 000000000..9c109ad65 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 4, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-size-05 in Draft: 3 workers with 12 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/expected.yaml new file mode 100644 index 000000000..383e759a0 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/expected.yaml @@ -0,0 +1,34 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-size-05 + name: discovered-discover-size-05 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-size-05-cluster + vcpuCount: 4 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - worker-02 + - worker-03 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/nodes.yaml new file mode 100644 index 000000000..98de08ee9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/nodes.yaml @@ -0,0 +1,89 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/ops.yaml new file mode 100644 index 000000000..f4aeacdd0 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/ops.yaml @@ -0,0 +1,14 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-size-05 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/reports/worker-01.yaml new file mode 100644 index 000000000..6b1427852 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/reports/worker-01.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 8, + "physicalCores": 4, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7 + ], + "physicalCores": 4 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-size-05 + name: sb-nodeprobe-discover-size-05-worker-01-afe48f86 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/reports/worker-02.yaml new file mode 100644 index 000000000..3a3855317 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-size-05 + name: sb-nodeprobe-discover-size-05-worker-02-0fc3e79e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/reports/worker-03.yaml new file mode 100644 index 000000000..8c622c656 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/reports/worker-03.yaml @@ -0,0 +1,241 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 98, + "physicalCores": 49, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31, + 32, + 33, + 34, + 35, + 36, + 37, + 38, + 39, + 40, + 41, + 42, + 43, + 44, + 45, + 46, + 47, + 48, + 49, + 50, + 51, + 52, + 53, + 54, + 55, + 56, + 57, + 58, + 59, + 60, + 61, + 62, + 63, + 64, + 65, + 66, + 67, + 68, + 69, + 70, + 71, + 72, + 73, + 74, + 75, + 76, + 77, + 78, + 79, + 80, + 81, + 82, + 83, + 84, + 85, + 86, + 87, + 88, + 89, + 90, + 91, + 92, + 93, + 94, + 95, + 96, + 97 + ], + "physicalCores": 49 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-size-05 + name: sb-nodeprobe-discover-size-05-worker-03-465caf88 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/case.md b/operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/case.md new file mode 100644 index 000000000..0c0a94755 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/case.md @@ -0,0 +1,9 @@ +# SIZE-06 + +**Mutation.** 4 GiB total, 900 MiB available, 4 free disks + +**Expected.** Admitted unchanged: no rule reads memory. See §14, gap G-4 + +**Harness.** `CM` + +**Gap.** G-4. This case records what the generator does today, so that the day it changes the diff is the finding. diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/expected-notes.txt new file mode 100644 index 000000000..ac64fa13c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 4, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-size-06 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/expected.yaml new file mode 100644 index 000000000..d7663cbd6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-size-06 + name: discovered-discover-size-06 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-size-06-cluster + vcpuCount: 4 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/ops.yaml new file mode 100644 index 000000000..250451587 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-size-06 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/reports/worker-01.yaml new file mode 100644 index 000000000..bdba84f78 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/reports/worker-01.yaml @@ -0,0 +1,146 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 8, + "physicalCores": 4, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7 + ], + "physicalCores": 4 + } + ] + }, + "memory": { + "totalBytes": 4294967296, + "freeBytes": 235929600, + "availableBytes": 943718400, + "numaNodes": [ + { + "node": 0, + "totalBytes": 4294967296, + "freeBytes": 2147483648 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-size-06 + name: sb-nodeprobe-discover-size-06-worker-01-1b51480d + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/case.md b/operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/case.md new file mode 100644 index 000000000..210b3fec4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/case.md @@ -0,0 +1,7 @@ +# SIZE-07 + +**Mutation.** `memory` zero and an `unreadable` entry for it + +**Expected.** Admitted by default. The draft is written from the disks + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/expected-notes.txt new file mode 100644 index 000000000..ea67da423 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-size-07 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/expected.yaml new file mode 100644 index 000000000..7ee2ec660 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-size-07 + name: discovered-discover-size-07 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-size-07-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/ops.yaml new file mode 100644 index 000000000..c2d8cfa57 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-size-07 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/reports/worker-01.yaml new file mode 100644 index 000000000..92b86ecc4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/reports/worker-01.yaml @@ -0,0 +1,166 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 0, + "freeBytes": 0, + "availableBytes": 0 + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ], + "unreadable": [ + "read /proc/meminfo: permission denied" + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-size-07 + name: sb-nodeprobe-discover-size-07-worker-01-4c9abfa4 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-08-the-fully-readable-worker-rule/case.md b/operator/internal/controllers/deployment/testdata/discovery/size/size-08-the-fully-readable-worker-rule/case.md new file mode 100644 index 000000000..6c6d7f6b3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-08-the-fully-readable-worker-rule/case.md @@ -0,0 +1,9 @@ +# SIZE-08 + +**Mutation.** The same report with `WorkerRules: WorkerWasReadable` + +**Expected.** The worker is refused, the message quoting the unreadable readings + +**Harness.** `GO` + +**Note.** Driven in the discovery package with Planner{WorkerRules: []WorkerRule{WorkerHasDevices{}, WorkerWasReadable{}}} over the SIZE-07 report, because the controller never substitutes a worker rule. diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/case.md b/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/case.md new file mode 100644 index 000000000..541b4494c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/case.md @@ -0,0 +1,9 @@ +# SIZE-09 + +**Mutation.** 512 × 1 GiB pages pre-allocated, 256 free per node + +**Expected.** **Contested.** `minHugePagesSize: 256G`: another workload's reservation becomes simplyblock's own floor. See §14, gap G-14 + +**Harness.** `CM` + +**Gap.** G-14. This case records what the generator does today, so that the day it changes the diff is the finding. diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/expected-notes.txt new file mode 100644 index 000000000..8d77f9fd7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 8, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 256G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-size-09 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 2 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/expected-refusals.txt new file mode 100644 index 000000000..e5d0ed2ac --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/expected-refusals.txt @@ -0,0 +1,2 @@ +worker-01/nvme2n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 1, so it was chosen +worker-01/nvme3n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 1, so it was chosen diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/expected.yaml new file mode 100644 index 000000000..87571ad98 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/expected.yaml @@ -0,0 +1,30 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-size-09 + name: discovered-discover-size-09 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 256G + name: discovered-discover-size-09-cluster + vcpuCount: 8 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + mgmtInterface: eth0 + name: group-1-nvme-2x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/ops.yaml new file mode 100644 index 000000000..65f8071f0 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-size-09 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/reports/worker-01.yaml new file mode 100644 index 000000000..cd4671bc7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/reports/worker-01.yaml @@ -0,0 +1,182 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 2, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15 + ], + "physicalCores": 8 + }, + { + "node": 1, + "onlineCPUs": [ + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 8 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "hugePagesBytes": 549755813888, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 512, + "free": 512, + "numaNodes": [ + { + "node": 0, + "total": 256, + "free": 256 + }, + { + "node": 1, + "total": 256, + "free": 256 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:7e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:7e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-size-09 + name: sb-nodeprobe-discover-size-09-worker-01-5ce69e44 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/case.md b/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/case.md new file mode 100644 index 000000000..055bf5d3a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/case.md @@ -0,0 +1,9 @@ +# SIZE-10 + +**Mutation.** One worker of three with no pre-allocated pages on its chosen node + +**Expected.** **Contested.** Unset for the whole cluster, the note naming that worker. Nothing pre-allocated is the ordinary case, not a reason to propose nothing. See §14, gap G-15 + +**Harness.** `CM` + +**Gap.** G-15. This case records what the generator does today, so that the day it changes the diff is the finding. diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/expected-notes.txt new file mode 100644 index 000000000..861adb587 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/expected-notes.txt @@ -0,0 +1,4 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 8, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +wrote discovered-discover-size-10 in Draft: 3 workers with 6 nvme devices, in 1 group(s) across 1 node set(s); 6 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/expected-refusals.txt new file mode 100644 index 000000000..fc4f7afb6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/expected-refusals.txt @@ -0,0 +1,6 @@ +worker-01/nvme2n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 1, so it was chosen +worker-01/nvme3n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 1, so it was chosen +worker-02/nvme2n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 1, so it was chosen +worker-02/nvme3n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 1, so it was chosen +worker-03/nvme2n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 1, so it was chosen +worker-03/nvme3n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 1, so it was chosen diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/expected.yaml new file mode 100644 index 000000000..48ffa0828 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/expected.yaml @@ -0,0 +1,31 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-size-10 + name: discovered-discover-size-10 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + name: discovered-discover-size-10-cluster + vcpuCount: 8 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + mgmtInterface: eth0 + name: group-1-nvme-2x3T + workers: + - worker-01 + - worker-02 + - worker-03 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/nodes.yaml new file mode 100644 index 000000000..98de08ee9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/nodes.yaml @@ -0,0 +1,89 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/ops.yaml new file mode 100644 index 000000000..a438d409b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/ops.yaml @@ -0,0 +1,14 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-size-10 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/reports/worker-01.yaml new file mode 100644 index 000000000..d9e970e9b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/reports/worker-01.yaml @@ -0,0 +1,181 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 2, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15 + ], + "physicalCores": 8 + }, + { + "node": 1, + "onlineCPUs": [ + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 8 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 256, + "free": 256, + "numaNodes": [ + { + "node": 0, + "total": 128, + "free": 128 + }, + { + "node": 1, + "total": 128, + "free": 128 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:7e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:7e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-size-10 + name: sb-nodeprobe-discover-size-10-worker-01-35000808 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/reports/worker-02.yaml new file mode 100644 index 000000000..420370eb1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/reports/worker-02.yaml @@ -0,0 +1,162 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 2, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15 + ], + "physicalCores": 8 + }, + { + "node": 1, + "onlineCPUs": [ + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 8 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:7e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:7e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-size-10 + name: sb-nodeprobe-discover-size-10-worker-02-68302997 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/reports/worker-03.yaml new file mode 100644 index 000000000..e4fd279cf --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/reports/worker-03.yaml @@ -0,0 +1,181 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 2, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15 + ], + "physicalCores": 8 + }, + { + "node": 1, + "onlineCPUs": [ + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 8 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 256, + "free": 256, + "numaNodes": [ + { + "node": 0, + "total": 128, + "free": 128 + }, + { + "node": 1, + "total": 128, + "free": 128 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:7e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:7e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-size-10 + name: sb-nodeprobe-discover-size-10-worker-03-448162d8 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/case.md b/operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/case.md new file mode 100644 index 000000000..b1563b9bc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/case.md @@ -0,0 +1,7 @@ +# SIZE-11 + +**Mutation.** 256 × 2 MiB pages free on the chosen node, 512 MiB in all + +**Expected.** Unset, the note saying the smallest reservation is under a gigabyte. A 1.5 GiB reading truncates to `1G` + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/expected-notes.txt new file mode 100644 index 000000000..6f685dc6f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/expected-notes.txt @@ -0,0 +1,4 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +wrote discovered-discover-size-11 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/expected.yaml new file mode 100644 index 000000000..0d3f70fc4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/expected.yaml @@ -0,0 +1,31 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-size-11 + name: discovered-discover-size-11 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + name: discovered-discover-size-11-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/ops.yaml new file mode 100644 index 000000000..0382f028a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-size-11 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/reports/worker-01.yaml new file mode 100644 index 000000000..938b8e3f2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/reports/worker-01.yaml @@ -0,0 +1,170 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 2097152, + "total": 256, + "free": 256, + "numaNodes": [ + { + "node": 0, + "total": 256, + "free": 256 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-size-11 + name: sb-nodeprobe-discover-size-11-worker-01-4b21c543 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/case.md b/operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/case.md new file mode 100644 index 000000000..6063efc92 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/case.md @@ -0,0 +1,7 @@ +# SIZE-12 + +**Mutation.** A pool with a total and an empty `numaNodes` + +**Expected.** Unset: the reservation exists and its distribution is unknown + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/expected-notes.txt new file mode 100644 index 000000000..0e0ad126f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/expected-notes.txt @@ -0,0 +1,4 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +wrote discovered-discover-size-12 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/expected.yaml new file mode 100644 index 000000000..8c90d651d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/expected.yaml @@ -0,0 +1,31 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-size-12 + name: discovered-discover-size-12 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + name: discovered-discover-size-12-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/ops.yaml new file mode 100644 index 000000000..b7b3755d2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-size-12 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/reports/worker-01.yaml new file mode 100644 index 000000000..885ac6ed7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/reports/worker-01.yaml @@ -0,0 +1,163 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 64, + "free": 64 + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-size-12 + name: sb-nodeprobe-discover-size-12-worker-01-97e789f8 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/case.md b/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/case.md new file mode 100644 index 000000000..41a78d634 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/case.md @@ -0,0 +1,7 @@ +# SIZE-13 + +**Mutation.** Pages pre-allocated on node 1 while the placement chooses node 0 + +**Expected.** Unset, the note naming the worker. The placement does not read huge pages, and the pool on the other node is neither counted nor reported + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/expected-notes.txt new file mode 100644 index 000000000..64d3c3110 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/expected-notes.txt @@ -0,0 +1,4 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 8, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +wrote discovered-discover-size-13 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 2 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/expected-refusals.txt new file mode 100644 index 000000000..e5d0ed2ac --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/expected-refusals.txt @@ -0,0 +1,2 @@ +worker-01/nvme2n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 1, so it was chosen +worker-01/nvme3n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 1, so it was chosen diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/expected.yaml new file mode 100644 index 000000000..9105f88ec --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/expected.yaml @@ -0,0 +1,29 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-size-13 + name: discovered-discover-size-13 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + name: discovered-discover-size-13-cluster + vcpuCount: 8 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + mgmtInterface: eth0 + name: group-1-nvme-2x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/ops.yaml new file mode 100644 index 000000000..da62d5bdd --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-size-13 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/reports/worker-01.yaml new file mode 100644 index 000000000..c063bf0fe --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/reports/worker-01.yaml @@ -0,0 +1,181 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 2, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15 + ], + "physicalCores": 8 + }, + { + "node": 1, + "onlineCPUs": [ + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 8 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 128, + "free": 128, + "numaNodes": [ + { + "node": 0, + "total": 0, + "free": 0 + }, + { + "node": 1, + "total": 128, + "free": 128 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:7e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:7e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-size-13 + name: sb-nodeprobe-discover-size-13-worker-01-4919bb90 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/case.md b/operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/case.md new file mode 100644 index 000000000..74e676cf6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/case.md @@ -0,0 +1,9 @@ +# SIZE-14 + +**Mutation.** Swap total 8 GiB, swap free 0 + +**Expected.** No effect on the draft. See §14, gap G-5 + +**Harness.** `CM` + +**Gap.** G-5. This case records what the generator does today, so that the day it changes the diff is the finding. diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/expected-notes.txt new file mode 100644 index 000000000..723807fda --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-size-14 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/expected.yaml new file mode 100644 index 000000000..0753065d4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-size-14 + name: discovered-discover-size-14 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-size-14-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/ops.yaml new file mode 100644 index 000000000..66a8d22cb --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-size-14 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/reports/worker-01.yaml new file mode 100644 index 000000000..dfb3684ae --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/reports/worker-01.yaml @@ -0,0 +1,176 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "swapTotalBytes": 8589934592, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-size-14 + name: sb-nodeprobe-discover-size-14-worker-01-579216d1 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/case.md b/operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/case.md new file mode 100644 index 000000000..1a2914a85 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/case.md @@ -0,0 +1,7 @@ +# SIZE-15 + +**Mutation.** `nodes.yaml` allocatable memory far under capacity + +**Expected.** Recorded on the `KubeNode`, no effect on the draft + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/expected-notes.txt new file mode 100644 index 000000000..815891ccf --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-size-15 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/expected.yaml new file mode 100644 index 000000000..98b7d62f9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-size-15 + name: discovered-discover-size-15 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-size-15-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/nodes.yaml new file mode 100644 index 000000000..d951b922a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/nodes.yaml @@ -0,0 +1,31 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: "28" + hugepages-1Gi: 64Gi + memory: 180Gi + capacity: + cpu: "32" + hugepages-1Gi: 64Gi + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/ops.yaml new file mode 100644 index 000000000..92d5dc0a8 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-size-15 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/reports/worker-01.yaml new file mode 100644 index 000000000..54e089898 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/reports/worker-01.yaml @@ -0,0 +1,170 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 274877906944, + "freeBytes": 137438953472 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-size-15 + name: sb-nodeprobe-discover-size-15-worker-01-febefbb4 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/case.md b/operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/case.md new file mode 100644 index 000000000..df500e174 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/case.md @@ -0,0 +1,9 @@ +# SIZE-16 + +**Mutation.** 256 GiB RAM, 200 GiB already in huge pages, 40 GiB available + +**Expected.** **Contested.** `minHugePagesSize: 200G` and no arithmetic against the 40 GiB left. See §14, gap G-16 + +**Harness.** `CM` + +**Gap.** G-16. This case records what the generator does today, so that the day it changes the diff is the finding. diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/expected-notes.txt new file mode 100644 index 000000000..7db22e3f9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 200G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-size-16 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/expected.yaml new file mode 100644 index 000000000..05efb2bad --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-size-16 + name: discovered-discover-size-16 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 200G + name: discovered-discover-size-16-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/ops.yaml new file mode 100644 index 000000000..3667b97f8 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-size-16 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/reports/worker-01.yaml new file mode 100644 index 000000000..954e95e67 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/reports/worker-01.yaml @@ -0,0 +1,166 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 60129542144, + "freeBytes": 10737418240, + "availableBytes": 42949672960, + "hugePagesBytes": 214748364800, + "numaNodes": [ + { + "node": 0, + "totalBytes": 60129542144, + "freeBytes": 30064771072 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 200, + "free": 200, + "numaNodes": [ + { + "node": 0, + "total": 200, + "free": 200 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-size-16 + name: sb-nodeprobe-discover-size-16-worker-01-44d22c6a + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/case.md b/operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/case.md new file mode 100644 index 000000000..54c110a0a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/case.md @@ -0,0 +1,9 @@ +# SIZE-17 + +**Mutation.** 1 TiB RAM, nothing pre-allocated, the whole machine available + +**Expected.** **Contested.** Nothing proposed, though the room for an allocation is there and reported. See §14, gaps G-15 and G-16 + +**Harness.** `CM` + +**Gap.** G-15 and G-16. This case records what the generator does today, so that the day it changes the diff is the finding. diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/expected-notes.txt new file mode 100644 index 000000000..f86ee2302 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/expected-notes.txt @@ -0,0 +1,4 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +wrote discovered-discover-size-17 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/expected.yaml new file mode 100644 index 000000000..258d29e00 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/expected.yaml @@ -0,0 +1,31 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-size-17 + name: discovered-discover-size-17 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + name: discovered-discover-size-17-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/ops.yaml new file mode 100644 index 000000000..5ce6cf477 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-size-17 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/reports/worker-01.yaml new file mode 100644 index 000000000..b304b6fe4 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/reports/worker-01.yaml @@ -0,0 +1,151 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 1099511627776, + "freeBytes": 268435456000, + "availableBytes": 1073741824000, + "numaNodes": [ + { + "node": 0, + "totalBytes": 1099511627776, + "freeBytes": 549755813888 + } + ] + }, + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-size-17 + name: sb-nodeprobe-discover-size-17-worker-01-fbea9b24 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/case.md b/operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/case.md new file mode 100644 index 000000000..1e914362d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/case.md @@ -0,0 +1,9 @@ +# SIZE-18 + +**Mutation.** A run meant to take 128 of the host's 256 pre-allocated pages + +**Expected.** **Contested.** No field expresses it. `spec.discover` carries no huge-page input, and neither `minHugePagesSize` nor `spdkSystemMemory` is written by a run. See §14, gap G-17 + +**Harness.** `CM` + +**Gap.** G-17. This case records what the generator does today, so that the day it changes the diff is the finding. diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/expected-notes.txt new file mode 100644 index 000000000..cafc79400 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 256G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-size-18 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/expected.yaml new file mode 100644 index 000000000..1f742150a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-size-18 + name: discovered-discover-size-18 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 256G + name: discovered-discover-size-18-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/ops.yaml new file mode 100644 index 000000000..3a88e85a1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-size-18 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/reports/worker-01.yaml new file mode 100644 index 000000000..ee537a842 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/reports/worker-01.yaml @@ -0,0 +1,171 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "hugePagesBytes": 274877906944, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 256, + "free": 256, + "numaNodes": [ + { + "node": 0, + "total": 256, + "free": 256 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-size-18 + name: sb-nodeprobe-discover-size-18-worker-01-c6cc26c0 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/case.md b/operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/case.md new file mode 100644 index 000000000..a92471d41 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/case.md @@ -0,0 +1,9 @@ +# SIZE-19 + +**Mutation.** A pool of 256 pages, all promised to a mapping: `resv` 256, `free` 0 + +**Expected.** Indistinguishable from an untouched pool of 256, because the report drops `resv_hugepages`. See §14, gap G-18 + +**Harness.** `CM` + +**Gap.** G-18. This case records what the generator does today, so that the day it changes the diff is the finding. diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/expected-notes.txt new file mode 100644 index 000000000..7248e4ea1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/expected-notes.txt @@ -0,0 +1,4 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +wrote discovered-discover-size-19 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/expected.yaml new file mode 100644 index 000000000..56b88f5eb --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/expected.yaml @@ -0,0 +1,31 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-size-19 + name: discovered-discover-size-19 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + name: discovered-discover-size-19-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/ops.yaml new file mode 100644 index 000000000..04dd08f55 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-size-19 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/reports/worker-01.yaml new file mode 100644 index 000000000..39b70de44 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/reports/worker-01.yaml @@ -0,0 +1,171 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "hugePagesBytes": 274877906944, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 256, + "free": 0, + "numaNodes": [ + { + "node": 0, + "total": 256, + "free": 0 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-size-19 + name: sb-nodeprobe-discover-size-19-worker-01-42ec74b1 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/case.md b/operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/case.md new file mode 100644 index 000000000..4618176b1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/case.md @@ -0,0 +1,7 @@ +# SIZE-20 + +**Mutation.** A 1 GiB pool and a 2 MiB pool, both with free pages on the chosen node + +**Expected.** The two sizes are summed into one figure, and nothing says which size the node should take + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/expected-notes.txt new file mode 100644 index 000000000..5df61aaf5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 34G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-size-20 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/expected.yaml new file mode 100644 index 000000000..9172f72e1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-size-20 + name: discovered-discover-size-20 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 34G + name: discovered-discover-size-20-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/ops.yaml new file mode 100644 index 000000000..5085336fb --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-size-20 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/reports/worker-01.yaml new file mode 100644 index 000000000..e6b605d0f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/reports/worker-01.yaml @@ -0,0 +1,182 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 2097152, + "total": 1024, + "free": 1024, + "numaNodes": [ + { + "node": 0, + "total": 1024, + "free": 1024 + } + ] + }, + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 32, + "free": 32 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-size-20 + name: sb-nodeprobe-discover-size-20-worker-01-9ef17276 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/case.md b/operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/case.md new file mode 100644 index 000000000..760dfe980 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/case.md @@ -0,0 +1,9 @@ +# SIZE-21 + +**Mutation.** No hugetlbfs at all, `hugePages` absent from the report + +**Expected.** The same unset output as a machine whose pools are full, so the two are not told apart. See §14, gap G-19 + +**Harness.** `CM` + +**Gap.** G-19. This case records what the generator does today, so that the day it changes the diff is the finding. diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/expected-notes.txt new file mode 100644 index 000000000..e97928389 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/expected-notes.txt @@ -0,0 +1,4 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +wrote discovered-discover-size-21 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/expected.yaml new file mode 100644 index 000000000..2df28e9cc --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/expected.yaml @@ -0,0 +1,31 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-size-21 + name: discovered-discover-size-21 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + name: discovered-discover-size-21-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + - 0000:5e:00.2 + - 0000:5e:00.3 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/nodes.yaml new file mode 100644 index 000000000..07f9cf5f3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/nodes.yaml @@ -0,0 +1,29 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/ops.yaml new file mode 100644 index 000000000..ba3e02e97 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/ops.yaml @@ -0,0 +1,12 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-size-21 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/reports/worker-01.yaml new file mode 100644 index 000000000..353a37303 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/reports/worker-01.yaml @@ -0,0 +1,156 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:5e:00.2", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:5e:00.3", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-size-21 + name: sb-nodeprobe-discover-size-21-worker-01-26c78f43 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/case.md b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/case.md new file mode 100644 index 000000000..0c8dc2a93 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/case.md @@ -0,0 +1,7 @@ +# TMPL-01 + +**Mutation.** Any successful run + +**Expected.** `spec.approved` false, `enableDriveFormat` true, and the note about formatting + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/expected-notes.txt new file mode 100644 index 000000000..111703651 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-tmpl-01 in Draft: 3 workers with 12 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/expected.yaml new file mode 100644 index 000000000..e6897f3a2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/expected.yaml @@ -0,0 +1,34 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-tmpl-01 + name: discovered-discover-tmpl-01 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-tmpl-01-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - worker-02 + - worker-03 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/nodes.yaml new file mode 100644 index 000000000..98de08ee9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/nodes.yaml @@ -0,0 +1,89 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/ops.yaml new file mode 100644 index 000000000..6473a3f7c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/ops.yaml @@ -0,0 +1,14 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-tmpl-01 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/reports/worker-01.yaml new file mode 100644 index 000000000..55fc4127c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-tmpl-01 + name: sb-nodeprobe-discover-tmpl-01-worker-01-70c1a3b7 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/reports/worker-02.yaml new file mode 100644 index 000000000..aa365cbcd --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-tmpl-01 + name: sb-nodeprobe-discover-tmpl-01-worker-02-a36df8a1 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/reports/worker-03.yaml new file mode 100644 index 000000000..2d4c96208 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-tmpl-01 + name: sb-nodeprobe-discover-tmpl-01-worker-03-32e107cd + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/case.md b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/case.md new file mode 100644 index 000000000..f6f62d82f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/case.md @@ -0,0 +1,7 @@ +# TMPL-02 + +**Mutation.** Any successful run + +**Expected.** `maxSubsystemCount: 30` with the note saying it is not a reading + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/expected-notes.txt new file mode 100644 index 000000000..f21beef43 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-tmpl-02 in Draft: 3 workers with 12 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/expected.yaml new file mode 100644 index 000000000..3a3c09262 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/expected.yaml @@ -0,0 +1,34 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-tmpl-02 + name: discovered-discover-tmpl-02 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-tmpl-02-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - worker-02 + - worker-03 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/nodes.yaml new file mode 100644 index 000000000..98de08ee9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/nodes.yaml @@ -0,0 +1,89 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/ops.yaml new file mode 100644 index 000000000..80fd26958 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/ops.yaml @@ -0,0 +1,14 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-tmpl-02 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/reports/worker-01.yaml new file mode 100644 index 000000000..aa7e52bb2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-tmpl-02 + name: sb-nodeprobe-discover-tmpl-02-worker-01-fbc376d6 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/reports/worker-02.yaml new file mode 100644 index 000000000..510c12aa5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-tmpl-02 + name: sb-nodeprobe-discover-tmpl-02-worker-02-917ff35c + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/reports/worker-03.yaml new file mode 100644 index 000000000..c83aba638 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-tmpl-02 + name: sb-nodeprobe-discover-tmpl-02-worker-03-77a6ffd7 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/case.md b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/case.md new file mode 100644 index 000000000..cefb4945d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/case.md @@ -0,0 +1,7 @@ +# TMPL-03 + +**Mutation.** `spec.discover.configName` unset + +**Expected.** The name is `discovered-` and the cluster `discovered--cluster` + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/expected-notes.txt new file mode 100644 index 000000000..ba1c098ba --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-tmpl-03 in Draft: 3 workers with 12 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/expected.yaml new file mode 100644 index 000000000..2eb3831c7 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/expected.yaml @@ -0,0 +1,34 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-tmpl-03 + name: discovered-discover-tmpl-03 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-tmpl-03-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - worker-02 + - worker-03 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/nodes.yaml new file mode 100644 index 000000000..98de08ee9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/nodes.yaml @@ -0,0 +1,89 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/ops.yaml new file mode 100644 index 000000000..ddeeb6fbd --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/ops.yaml @@ -0,0 +1,14 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-tmpl-03 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/reports/worker-01.yaml new file mode 100644 index 000000000..783655204 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-tmpl-03 + name: sb-nodeprobe-discover-tmpl-03-worker-01-63aa463b + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/reports/worker-02.yaml new file mode 100644 index 000000000..ca433d06f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-tmpl-03 + name: sb-nodeprobe-discover-tmpl-03-worker-02-83ad10c2 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/reports/worker-03.yaml new file mode 100644 index 000000000..b803eb1f2 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-tmpl-03 + name: sb-nodeprobe-discover-tmpl-03-worker-03-6476d6ac + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/case.md b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/case.md new file mode 100644 index 000000000..eb6872200 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/case.md @@ -0,0 +1,9 @@ +# TMPL-04 + +**Mutation.** An `OperatorOps` name long enough to push the cluster past 63 + +**Expected.** **Contested.** The document is written and `CreatingCluster` can never succeed. See §14, gap G-11 + +**Harness.** `CM` + +**Gap.** G-11. This case records what the generator does today, so that the day it changes the diff is the finding. diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/expected-notes.txt new file mode 100644 index 000000000..4247c2f59 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-the-quarterly-storage-expansion-for-the-frankfurt-racks in Draft: 3 workers with 12 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/expected.yaml new file mode 100644 index 000000000..06ce8f3b5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/expected.yaml @@ -0,0 +1,34 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-tmpl-04 + name: discovered-the-quarterly-storage-expansion-for-the-frankfurt-racks + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-the-quarterly-storage-expansion-for-the-frankfurt-racks-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - worker-02 + - worker-03 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/nodes.yaml new file mode 100644 index 000000000..98de08ee9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/nodes.yaml @@ -0,0 +1,89 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/ops.yaml new file mode 100644 index 000000000..acc6e550d --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/ops.yaml @@ -0,0 +1,16 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-tmpl-04 + namespace: simplyblock +spec: + action: Discover + discover: + configName: discovered-the-quarterly-storage-expansion-for-the-frankfurt-racks +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/reports/worker-01.yaml new file mode 100644 index 000000000..c36584786 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-tmpl-04 + name: sb-nodeprobe-discover-tmpl-04-worker-01-9736fa02 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/reports/worker-02.yaml new file mode 100644 index 000000000..7e2ebc1b6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-tmpl-04 + name: sb-nodeprobe-discover-tmpl-04-worker-02-f833124e + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/reports/worker-03.yaml new file mode 100644 index 000000000..da3daf75a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-tmpl-04 + name: sb-nodeprobe-discover-tmpl-04-worker-03-808cedc1 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/case.md b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/case.md new file mode 100644 index 000000000..dc72dfcfd --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/case.md @@ -0,0 +1,7 @@ +# TMPL-05 + +**Mutation.** `spec.discover.clusterRef` set + +**Expected.** No `spec.cluster`, and `clusterRef` carried through with the growth note + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/expected-notes.txt new file mode 100644 index 000000000..5d2e0c28b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/expected-notes.txt @@ -0,0 +1,2 @@ +the draft grows the existing cluster existing-cluster, so it proposes no cluster layout +wrote discovered-discover-tmpl-05 in Draft: 3 workers with 12 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/expected.yaml new file mode 100644 index 000000000..98b7a0e9c --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/expected.yaml @@ -0,0 +1,29 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-tmpl-05 + name: discovered-discover-tmpl-05 + namespace: simplyblock +spec: + approved: false + clusterRef: existing-cluster + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - worker-02 + - worker-03 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/nodes.yaml new file mode 100644 index 000000000..98de08ee9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/nodes.yaml @@ -0,0 +1,89 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/ops.yaml new file mode 100644 index 000000000..57f71de03 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/ops.yaml @@ -0,0 +1,16 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-tmpl-05 + namespace: simplyblock +spec: + action: Discover + discover: + clusterRef: existing-cluster +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/reports/worker-01.yaml new file mode 100644 index 000000000..1294cba01 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-tmpl-05 + name: sb-nodeprobe-discover-tmpl-05-worker-01-b7723703 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/reports/worker-02.yaml new file mode 100644 index 000000000..c04c62fec --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-tmpl-05 + name: sb-nodeprobe-discover-tmpl-05-worker-02-815d6280 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/reports/worker-03.yaml new file mode 100644 index 000000000..f441dccdb --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-05-a-growth-document/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-tmpl-05 + name: sb-nodeprobe-discover-tmpl-05-worker-03-72ce050a + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/case.md b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/case.md new file mode 100644 index 000000000..97b0c3f26 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/case.md @@ -0,0 +1,9 @@ +# TMPL-06 + +**Mutation.** Any run on a two-socket fleet + +**Expected.** `socketsToUse` and `nodesPerSocket` unset, so one node per worker on socket 0. See §14, gap G-12 + +**Harness.** `CM` + +**Gap.** G-12. This case records what the generator does today, so that the day it changes the diff is the finding. diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/expected-notes.txt new file mode 100644 index 000000000..879f58581 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 8, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-tmpl-06 in Draft: 3 workers with 6 nvme devices, in 1 group(s) across 1 node set(s); 6 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/expected-refusals.txt new file mode 100644 index 000000000..fc4f7afb6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/expected-refusals.txt @@ -0,0 +1,6 @@ +worker-01/nvme2n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 1, so it was chosen +worker-01/nvme3n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 1, so it was chosen +worker-02/nvme2n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 1, so it was chosen +worker-02/nvme3n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 1, so it was chosen +worker-03/nvme2n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 1, so it was chosen +worker-03/nvme3n1: declined by most available NUMA node because NUMA node 0 carries 2 unclaimed devices (6T) against 2 (6T) on NUMA node 1, so it was chosen diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/expected.yaml new file mode 100644 index 000000000..f723190cf --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/expected.yaml @@ -0,0 +1,32 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-tmpl-06 + name: discovered-discover-tmpl-06 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-tmpl-06-cluster + vcpuCount: 8 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5e:00.1 + mgmtInterface: eth0 + name: group-1-nvme-2x3T + workers: + - worker-01 + - worker-02 + - worker-03 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/nodes.yaml new file mode 100644 index 000000000..98de08ee9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/nodes.yaml @@ -0,0 +1,89 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/ops.yaml new file mode 100644 index 000000000..584abfa8e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/ops.yaml @@ -0,0 +1,14 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-tmpl-06 + namespace: simplyblock +spec: + action: Discover +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/reports/worker-01.yaml new file mode 100644 index 000000000..e5e81195b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/reports/worker-01.yaml @@ -0,0 +1,181 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 2, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15 + ], + "physicalCores": 8 + }, + { + "node": 1, + "onlineCPUs": [ + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 8 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:7e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:7e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-tmpl-06 + name: sb-nodeprobe-discover-tmpl-06-worker-01-3ccf91bf + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/reports/worker-02.yaml new file mode 100644 index 000000000..217c16107 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/reports/worker-02.yaml @@ -0,0 +1,181 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 2, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15 + ], + "physicalCores": 8 + }, + { + "node": 1, + "onlineCPUs": [ + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 8 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:7e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:7e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-tmpl-06 + name: sb-nodeprobe-discover-tmpl-06-worker-02-abec8778 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/reports/worker-03.yaml new file mode 100644 index 000000000..cc5e870ab --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/reports/worker-03.yaml @@ -0,0 +1,181 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 2, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15 + ], + "physicalCores": 8 + }, + { + "node": 1, + "onlineCPUs": [ + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 8 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:7e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:7e:00.1", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 1, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-tmpl-06 + name: sb-nodeprobe-discover-tmpl-06-worker-03-efcee559 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/case.md b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/case.md new file mode 100644 index 000000000..dfdcdd253 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/case.md @@ -0,0 +1,7 @@ +# TMPL-07 + +**Mutation.** `spec.environment` from the run's status + +**Expected.** Copied verbatim into the document + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/expected-notes.txt new file mode 100644 index 000000000..fbd13e1a0 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-tmpl-07 in Draft: 3 workers with 12 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/expected.yaml new file mode 100644 index 000000000..7c3f4a333 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/expected.yaml @@ -0,0 +1,34 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-tmpl-07 + name: discovered-discover-tmpl-07 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-tmpl-07-cluster + vcpuCount: 16 + environment: OpenShift + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + - 0000:b0:00.0 + mgmtInterface: eth0 + name: group-1-nvme-4x3T + workers: + - worker-01 + - worker-02 + - worker-03 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/nodes.yaml new file mode 100644 index 000000000..98de08ee9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/nodes.yaml @@ -0,0 +1,89 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/ops.yaml new file mode 100644 index 000000000..5dabbfc1f --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/ops.yaml @@ -0,0 +1,14 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-tmpl-07 + namespace: simplyblock +spec: + action: Discover +status: + environment: OpenShift + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/reports/worker-01.yaml new file mode 100644 index 000000000..b9789b1a0 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-tmpl-07 + name: sb-nodeprobe-discover-tmpl-07-worker-01-652238fd + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/reports/worker-02.yaml new file mode 100644 index 000000000..12e2e902b --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-tmpl-07 + name: sb-nodeprobe-discover-tmpl-07-worker-02-5930546d + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/reports/worker-03.yaml new file mode 100644 index 000000000..ce783785a --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-tmpl-07 + name: sb-nodeprobe-discover-tmpl-07-worker-03-cfd73a9c + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/case.md b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/case.md new file mode 100644 index 000000000..31761a305 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/case.md @@ -0,0 +1,7 @@ +# TMPL-08 + +**Mutation.** Any run + +**Expected.** No `deviceFilter` anywhere in the document: the resolved list is the record + +**Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/expected-notes.txt new file mode 100644 index 000000000..0ab6483c1 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/expected-notes.txt @@ -0,0 +1,5 @@ +enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. +vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes +maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +wrote discovered-discover-tmpl-08 in Draft: 3 workers with 9 nvme devices, in 1 group(s) across 1 node set(s); 3 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/expected-refusals.txt b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/expected-refusals.txt new file mode 100644 index 000000000..6c5791410 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/expected-refusals.txt @@ -0,0 +1,3 @@ +worker-01/nvme3n1: declined by allow and deny lists because it is in the deny list +worker-02/nvme3n1: declined by allow and deny lists because it is in the deny list +worker-03/nvme3n1: declined by allow and deny lists because it is in the deny list diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/expected.yaml new file mode 100644 index 000000000..d122d03d6 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/expected.yaml @@ -0,0 +1,33 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: ClusterDeploymentConfig +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-run: discover-tmpl-08 + name: discovered-discover-tmpl-08 + namespace: simplyblock +spec: + approved: false + cluster: + enableDriveFormat: true + maxSubsystemCount: 30 + minHugePagesSize: 16G + name: discovered-discover-tmpl-08-cluster + vcpuCount: 16 + environment: K3s + nodeSets: + - groups: + - devices: + nvme: + - 0000:5e:00.0 + - 0000:5f:00.0 + - 0000:af:00.0 + mgmtInterface: eth0 + name: group-1-nvme-3x3T + workers: + - worker-01 + - worker-02 + - worker-03 + name: discovered +status: + step: {} diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/nodes.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/nodes.yaml new file mode 100644 index 000000000..98de08ee9 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/nodes.yaml @@ -0,0 +1,89 @@ +apiVersion: v1 +kind: Node +metadata: + name: worker-01 +spec: {} +status: + addresses: + - address: 10.10.10.1 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-02 +spec: {} +status: + addresses: + - address: 10.10.10.2 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" +--- +apiVersion: v1 +kind: Node +metadata: + name: worker-03 +spec: {} +status: + addresses: + - address: 10.10.10.3 + type: InternalIP + allocatable: + cpu: 31500m + memory: 250Gi + capacity: + cpu: "32" + memory: 256Gi + daemonEndpoints: + kubeletEndpoint: + Port: 0 + nodeInfo: + architecture: "" + bootID: "" + containerRuntimeVersion: "" + kernelVersion: "" + kubeProxyVersion: "" + kubeletVersion: "" + machineID: "" + operatingSystem: "" + osImage: "" + systemUUID: "" diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/ops.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/ops.yaml new file mode 100644 index 000000000..a0e2aab9e --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/ops.yaml @@ -0,0 +1,19 @@ +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: discover-tmpl-08 + namespace: simplyblock +spec: + action: Discover + discover: + deviceFilter: + driveSizeRange: 1T-8T + pcieDenyList: + - 0000:b0:00.0 +status: + environment: K3s + step: {} + workers: + - worker-01 + - worker-02 + - worker-03 diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/reports/worker-01.yaml new file mode 100644 index 000000000..80a5301f5 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/reports/worker-01.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-01", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.1" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.1" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-01 + storage.simplyblock.io/nodeprobe-run: discover-tmpl-08 + name: sb-nodeprobe-discover-tmpl-08-worker-01-e21002a7 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/reports/worker-02.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/reports/worker-02.yaml new file mode 100644 index 000000000..2886436cf --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/reports/worker-02.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-02", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.2" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.2" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-02 + storage.simplyblock.io/nodeprobe-run: discover-tmpl-08 + name: sb-nodeprobe-discover-tmpl-08-worker-02-f2510a35 + namespace: simplyblock diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/reports/worker-03.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/reports/worker-03.yaml new file mode 100644 index 000000000..8ab8e20a3 --- /dev/null +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/reports/worker-03.yaml @@ -0,0 +1,175 @@ +apiVersion: v1 +data: + report.json: |- + { + "version": 6, + "node": "worker-03", + "probedAt": "2026-09-18T09:00:00Z", + "cpu": { + "onlineCPUs": 32, + "physicalCores": 16, + "sockets": 1, + "threadsPerCore": 2, + "hyperThreading": true, + "numaNodes": [ + { + "node": 0, + "onlineCPUs": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31 + ], + "physicalCores": 16 + } + ] + }, + "memory": { + "totalBytes": 274877906944, + "freeBytes": 64424509440, + "availableBytes": 257698037760, + "numaNodes": [ + { + "node": 0, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + }, + { + "node": 1, + "totalBytes": 137438953472, + "freeBytes": 68719476736 + } + ] + }, + "hugePages": [ + { + "sizeBytes": 1073741824, + "total": 32, + "free": 32, + "numaNodes": [ + { + "node": 0, + "total": 16, + "free": 16 + }, + { + "node": 1, + "total": 16, + "free": 16 + } + ] + } + ], + "interfaces": [ + { + "name": "eth0", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 0, + "kind": "physical", + "addresses": [ + "10.10.10.3" + ] + }, + { + "name": "eth1", + "speedMbps": 25000, + "mtu": 1500, + "state": "up", + "driver": "mlx5_core", + "numaNode": 1, + "kind": "physical", + "addresses": [ + "10.10.11.3" + ] + } + ], + "devices": [ + { + "name": "nvme0n1", + "path": "/dev/nvme0n1", + "pciAddress": "0000:5e:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme1n1", + "path": "/dev/nvme1n1", + "pciAddress": "0000:5f:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme2n1", + "path": "/dev/nvme2n1", + "pciAddress": "0000:af:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + }, + { + "name": "nvme3n1", + "path": "/dev/nvme3n1", + "pciAddress": "0000:b0:00.0", + "sizeBytes": 3298534883328, + "kind": "Disk", + "transport": "NVMe", + "model": "SAMSUNG MZQL23T8HCLS-00A07", + "numaNode": 0, + "available": true, + "content": "Blank" + } + ] + } +kind: ConfigMap +metadata: + labels: + storage.simplyblock.io/component: nodeprobe + storage.simplyblock.io/nodeprobe-node: worker-03 + storage.simplyblock.io/nodeprobe-run: discover-tmpl-08 + name: sb-nodeprobe-discover-tmpl-08-worker-03-618e59b8 + namespace: simplyblock From 3c325979a6bde24d916e905be60ec556767c7299 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Fri, 18 Sep 2026 16:52:26 +0200 Subject: [PATCH 074/206] docs(discovery): the rules the generator applies, and the cases that hold them Two working documents at the repository root, both written to be argued with. discovery-generator-rules.md is every rule between a fleet's probe reports and the drafted ClusterDeploymentConfig, with the precedence between them and the exact condition each turns on. Every rule carries where it came from: specified by a design, grounded in a constraint outside the code, asserted as a preference, invented while the code was written, or stated by a design and never built. That marker is the point of the document. The recorded expectations prove the rules were applied consistently and say nothing about whether they are the right rules, and several were never written down anywhere -- they were settled in code and the recordings now carry them as though they were agreed. It ends with three lists: the rules that need a decision, the defects found while cataloguing, and the one rule a design states that the code does not implement. Some of those are already fixed in the two commits before this one and are marked so; the rest are open, and the ranking questions in particular are product decisions rather than engineering ones. discovery-generator-test-cases.md is the case matrix the fixtures are generated from: 179 rows across twelve families, each with the mutation, the expected outcome, and the harness it runs in. A row marked contested records what the generator does today against a finding, so the day the behavior changes the diff is the finding rather than a surprise. They are at the root rather than under operator/docs because neither is finished: the first is a list somebody has to rule on, and the second moves every time a rule does. Co-Authored-By: Claude Opus 5 (1M context) --- discovery-generator-rules.md | 757 ++++++++++++++++++++++++++++++ discovery-generator-test-cases.md | 491 +++++++++++++++++++ 2 files changed, 1248 insertions(+) create mode 100644 discovery-generator-rules.md create mode 100644 discovery-generator-test-cases.md diff --git a/discovery-generator-rules.md b/discovery-generator-rules.md new file mode 100644 index 000000000..00358edb2 --- /dev/null +++ b/discovery-generator-rules.md @@ -0,0 +1,757 @@ + + +# The rules the discovery generator applies + +**What this is.** Every rule between a fleet's probe reports and the drafted +`ClusterDeploymentConfig`, with its precedence, its exact condition, and its +provenance. + +**Why it exists.** The 155 recorded expectations under +`operator/internal/controllers/deployment/testdata/discovery/` were produced by +running the generator over the fixtures. That proves the rules were applied +consistently and proves nothing about whether they are the right rules. Several +were never specified anywhere: they were settled in code, and the expectations +now carry them as though they were agreed. This is the list to ratify, correct, +or overrule. + +**How to read it.** Every rule carries a provenance marker, and the markers are +the point. + +| Marker | Meaning | +|-------------|--------------------------------------------------------------------------------------------------------------------------------------------------------------| +| `SPECIFIED` | A design document states the rule, and the code implements what was written | +| `GROUNDED` | No design states it, and the code cites a constraint outside itself: a control-plane refusal, an API server limit, an addressing requirement, or determinism | +| `ASSERTED` | No design states it, and the reason the code gives is a preference rather than a constraint | +| `INVENTED` | No design states it and no reason is given. It was settled while the code was written | +| `UNBUILT` | A design states the rule and the code does not implement it | + +A `SPECIFIED` or `GROUNDED` rule needs checking for accuracy. An `ASSERTED` or +`INVENTED` rule needs a decision. §10 collects the second kind. + +## 1. The pipeline, in order + +A run is three steps, and only the third produces a document. + +1. **Inspecting** settles which workers the run is about and which distribution + the cluster runs. Nothing after it re-derives either. +2. **Probing** starts one Job per worker, each writing a report into a + ConfigMap. The probe decides nothing: it reports every block device, refused + ones included, with the grounds for each. +3. **Writing** turns the reports into a document. + +Writing is itself ordered, and the order is load-bearing. + +``` +reports, sorted by worker name + └─ per worker: + 1. synthesize a device for every idle userspace NVMe controller §4.4 + 2. apply the device rules, first refusal wins §4.1 + 3. apply the worker rules, first refusal wins §4.3 + 4. choose which memory node to place on §5 + 5. record every admitted device the placement did not take §5.5 + 6. choose the management interface §7 + └─ over the workers that survived: + 7. group them by what they hand over §6.1 + 8. split the groups into node sets by role §6.4 + └─ over the plan: + 9. derive the cluster block §8 + 10. assemble, name, and write the document §9 +``` + +Devices are filtered before workers because whether a worker is worth including +is mostly a question about what is left of it. `GROUNDED`. + +## 2. Which workers the run is about + +Decided once, in Inspecting, and persisted. `GROUNDED`: a worker that joins the +cluster while the probes run is not in this run's draft, and re-listing in +Probing would put it there with no report to describe it. + +A node is a candidate when all four hold: + +| # | Condition | Provenance | +|-----|------------------------------------------------------------------|-------------| +| 1 | It matches `spec.discover.nodeSelector`, when one is set | `SPECIFIED` | +| 2 | No `StorageNode` in the namespace already names it | `SPECIFIED` | +| 3 | It is schedulable | `GROUNDED` | +| 4 | Its role holds storage nodes, or the control-plane opt-in is set | `GROUNDED` | + +**Schedulable** means not cordoned and carrying no `NoSchedule` or `NoExecute` +taint. `PreferNoSchedule` does not exclude. `GROUNDED`: a probe Job is pinned +with `spec.nodeName` and would run on either, and a worker the cluster is not +scheduling to is not one to hand to a storage cluster. + +**Role** is read from `node-role.kubernetes.io/` labels, and the most +restrictive wins: control-plane, then infra, then worker. Both `master` and +`control-plane` are read, neither preferred. An unrecognized role is ignored +rather than refused. Infra nodes are used **without** an opt-in, control-plane +nodes only with one. `GROUNDED`, and the asymmetry is argued: an OpenShift infra +node is the tier a cluster's own infrastructure runs on, and simplyblock storage +is infrastructure. + +A declined worker gets an event naming the reason. A node already carrying a +`StorageNode` is the one exclusion that stays **silent**, against the adjacent +comment claiming both exclusions are otherwise explained. + +**Environment** is concluded once from the whole node list, not per probe. +`SPECIFIED`. Signals, strongest first: registered API groups, then the kubelet +version string, then the OS image prefix, then node labels and annotations. When +several distributions match, precedence is OpenShift, Talos, K3s, Rancher. +`GROUNDED`: a Rancher-managed K3s cluster carries both sets of markers, and what +decides the host's shape is the distribution that installed the kubelet. Nothing +matching yields `Vanilla`. A failure to list API groups is an event, not a +failure, since half the evidence still yields a conclusion. +## 3. What the probe decides, and what it refuses to decide + +The probe decides nothing about what a cluster should use. `SPECIFIED`, and the +rule is load-bearing for everything below it: every block device is reported, +refused ones included, each with the grounds it was refused on, because the run's +device filter belongs to the operator that holds the fleet's reports and an +administrator has to be able to be told why the disk they expected is not a +candidate. + +Three consequences the generator depends on: + +- **The probe accumulates grounds, the generator stops at the first.** A device + arrives carrying every reason it is unusable, and is refused for one of them. + That is what makes the partition waiver safe: it admits a device whose **only** + ground is a partition table, so a boot disk stays out on its mount. +- **The report is a wire format with a version, and a reader refuses a version it + does not know.** `GROUNDED`: a Job keeps the image it started with, so an + operator upgraded mid-run reads reports from the previous probe, and a field + that changed meaning would be read as the current one. +- **The report is the only place a node's name survives exactly.** Object names + and labels are sanitized and truncated to fit a 63-byte budget, so the node a + report is about is the field inside it. + +What the probe reads and the generator never consults is itself a list worth +reviewing: per-node memory, online CPUs, swap, MTU, MAC address, driver, PCI +address of an interface, huge-page totals as distinct from free pages, and every +huge-page reading the kubelet publishes. None of them enters any rule. + +## 4. Which devices reach the draft + +Four stages, and a device dropped at one never reaches the next. Two of them are +the probe's and two are the generator's. + +| Stage | Owner | Decides | Short-circuits | +|-------|-----------|---------------------------------------------------|-------------------------| +| A | probe | Which sysfs entries exist, and their kind and bus | no | +| B | probe | Every ground a device is refused on | accumulates all grounds | +| C | generator | Admit or refuse each device | **first refusal wins** | +| D | generator | Admit or refuse the whole worker | **first refusal wins** | + +The probe accumulates every ground, and the generator stops at the first. A device +therefore carries all the reasons it is unusable and is refused for one of them. + +### 4.1 The device rules, in evaluation order + +| # | Rule | Present when | Pre-filter | Provenance | +|-----|----------------------|-----------------------------------------------|------------|------------| +| 1 | Whole disk | always | **yes** | `GROUNDED` | +| 2 | Device class | always | **yes** | `GROUNDED` | +| 3 | Available | always | no | `GROUNDED` | +| 4 | Allow and deny lists | the scanned class's allow or deny list is set | no | `ASSERTED` | +| 5 | Model | NVMe run and `pcieModel` is set | no | `GROUNDED` | +| 6 | Size range | `driveSizeRange` is set **and parses** | no | `ASSERTED` | + +The order is stated and justified: what the device is, then whether it is free, +then whether this run wants it. "A partition refused as 'not in the allow list' +would be a true statement and the wrong one." + +**Pre-filter** marks a rule whose refusal says nothing about the fleet's +storage. Pre-filtered refusals are emitted as events and are **excluded** from +the per-worker explanation the run's status carries. Only rules 1 and 2 set it. + +### 4.2 Each rule's exact condition + +**1. Whole disk**: refuses unless `kind == "Disk"`, exact and case-sensitive. +An empty kind refuses. `GROUNDED`: it is redundant against the probe, and it is +the one rule whose absence would be silent, because a partition admitted by a +waiver would be handed to a cluster as a disk. + +**2. Device class**: symmetric, and it has to be, because a cluster is built out +of one class. + +- On an **NVMe** run: refuses unless `transport == "NVMe"` exactly, then refuses + unless the PCI address is non-empty. +- On a **block** run: refuses `NVMe` and `NVMeFabric` alike, then refuses unless + the path is non-empty. + +The block side used to check only the path, so an NVMe disk reached a block draft +by its path and one document could name both classes. Naming devices by path is +what the block class does rather than what makes a device one, so the bus is what +both sides read. + +**2a. Simplyblock volume**: refuses a device whose subsystem NQN parses as a +simplyblock logical-volume subsystem, naming the cluster it belongs to. +`GROUNDED`: a volume this product exported and a worker attached is a namespace +like any other, and handing one back to a cluster as backend storage would give +a volume's own bytes away as free space. + +It sits ahead of the class rule deliberately. The class rule refuses every +fabric namespace too, but as a pre-filter, so its answer never reaches the +explanation: a reviewer asking why a machine full of disks proposed none would be +told the disks are on a bus the run does not scan, rather than that they are the +fleet's own volumes. It is also the check that survives the class rule being +relaxed, which is the one way a cluster could be told to take its own bytes. + +Parsing rather than matching a prefix is what makes it specific. A subsystem of +this product that is not a logical volume does not read as one, and a namespace +another product exported falls through to the class rule. + +**3. Available**: admits a device the probe found free. One waiver, and it is +narrow: with `enablePartitionedDevices`, a device whose **only** ground is a +partition table is admitted. A device also mounted, held, or unreadable is not. +`GROUNDED`, and the narrowness is enforced on the probe's side too. + +The refusal quotes the ground *names* joined by commas, not the details. A +device marked unavailable with no grounds at all, which is inconsistent data the +probe does not write, reads "the probe did not report it as available and gave no +reason." + +**5. iSCSI**: refuses an iSCSI LUN the run's allow list does not name. +`GROUNDED`: every other bus a run scans is a cable inside the chassis, and a LUN +is a disk on the other side of a network, so a cluster built on one runs every +write of its data path over that network. Whether that is wanted is a question +about the deployment rather than about the hardware. + +The default is refusal rather than proposal-and-review, because a draft +proposing a LUN is one a reviewer has to notice and strike, and a fifty-worker +document is not one anybody reads that closely. Naming it in the allow list is +the decision. + +It reads the allow list itself rather than leaving it to the rule below, because +that rule is only built when a list was given: with no filter there would be +nothing to refuse a LUN, and with a list given for another reason a LUN would be +admitted by being named alongside everything else. + +It also sits ahead of that rule so that an unnamed LUN says it is a LUN rather +than that it is not in the allow list, which is the same true-but-wrong-answer +the whole-disk ordering exists to avoid. + +**6. Allow and deny lists**: deny is evaluated first and wins, and an **empty allow +list means allow everything** rather than allow nothing. Matching is case-folded. + +Two things asserted and not justified: deny-before-allow, and the case folding. +The folding is defensible for a PCI address and is **wrong for a block path**: +`/dev/SDB` matches `/dev/sdb`, and Linux paths are case-sensitive. + +**7. Model**: case-folded **substring** match. `GROUNDED`: a model string is +padded, versioned, and vendor-formatted, so `MZQL2` is what an administrator +writes. A device with an empty model is refused by any non-empty filter. + +**8. Size range**: both bounds are **inclusive**, and a zero maximum means **no upper +bound**. Applies to whichever class is scanned. + +### 4.3 The worker rules + +**Has devices**: refuses a worker with nothing admitted, and says one of three +things. `GROUNDED` throughout, with the distinction spelled out: a machine whose +disks are on a userspace driver has no block devices at all, so "no device +survived the rules" is true and useless. + +| Situation | What the reviewer is told | +|-----------------------------------------------------|-------------------------------------------| +| Userspace-bound controllers, something driving them | something is driving its disks | +| Userspace-bound controllers, nothing driving them | the disks are there to be reclaimed | +| Neither | no device of it survived the device rules | + +**Fully readable**: refuses a worker whose probe could not read everything. Off +by default, and **nothing in the operator ever turns it on**. `GROUNDED`: a +worker whose CPU tree could not be read still has disks worth reviewing. + +### 4.4 Devices the generator invents + +An NVMe controller bound to a userspace driver presents no block device, so it +cannot reach a draft through the device reading at all. The generator synthesizes +one for each such controller that **nothing is using** and whose address no +reported device already carries. `GROUNDED`, and the trade is stated: everything +the disk would have said about itself is lost, so it reaches the draft unsized +and uninspected, and the group it lands in is named for a count rather than a +capacity. + +**Two interactions worth deciding on**, neither documented: + +- A **model filter refuses every synthesized controller**, because its model is + empty and the match is a substring. +- A **size range with any lower bound refuses every synthesized controller**, + because its size is zero. + +So a run that filters by model or by size silently excludes exactly the disks the +synthesis exists to offer. + +### 4.5 Findings + +**F-4.1. An unparsable size range silently admits everything, and the code says +otherwise.** `plan.go:281-283` states: "ParseSizeRange is called by the caller +that validates the spec, and this one skips what it cannot read rather than +silently widening the filter, and the run reports the parse failure separately." + +No such caller exists. `ParseSizeRange` is called from exactly one place outside +its own tests. `driveSizeRange` carries no schema pattern and no CEL rule. So an +unparsable range drops the rule entirely, every size passes, and **nothing +reports it**: not a refusal, not an event, not the summary. This is `G-9`, and +it is worse than recorded: the comment asserts a safeguard that was never built. + +**F-4.2. Eleven places drop a device or a worker with no record.** The ones that +matter: + +- A userspace controller that is **in use** is skipped silently. The only trace + is the worker-level message, and only if the worker ends up with nothing. +- A worker whose report cannot be decoded raises an event but produces **no + refusal**, so it is absent from the explanation the status carries. +- A worker whose report names a node the run is not about is dropped in silence. +- ~~`blockAllowList` and `blockDenyList` are silently ignored when the + planner's class and the filter disagree.~~ **Fixed.** The planner carried a + class field beside the filter's own `enableLogicalBlockDevices`, so one fact + had two statements and they could disagree: a planner told nothing scanned + NVMe, read the PCI lists, and dropped the block lists on a branch that never + ran. The field is gone, and the class is read from the filter, where it was + always written. + +**F-4.3. `InUse: false` is not distinguished from "never read."** The report's +own comment says a reader deciding whether to reclaim a controller "has to find +this report free of such an entry [in `unreadable`] first." Nothing performs that +check, and the rule that would, Fully readable, is off by default. A probe that +could not read the process table therefore yields controllers that look idle. + +**F-4.4. Fixed: the explanation counts by rule.** It used to group on the +rendered sentence, and the allow-and-deny refusal embedded the address in its +sentence, so a hundred declined devices produced a hundred clauses. It now +groups on the rule the refusal already carries, and the allow-and-deny reason +names the list rather than the device, which the refusal holds separately. Where +one rule's reasons genuinely differ, a mounted disk against a partitioned one, +each is counted rather than dropped. + +**F-4.5. Fixed: the bound is quoted as written, and a size is never rounded +up.** The bound used to be re-rendered from the parsed number, so a filter +naming `1920G` was quoted back as `1.875T`. The rule now carries the range as +the filter wrote it. + +The renderer was worse than lossy. Four significant digits rounded, so a byte +under a tebibyte printed as `1024G`, which reads as more than the value, and its +own parser refuses every decimal it produced. A whole number of units is now +written exactly and parses back to the byte it came from, and anything else is +truncated to two decimals. +## 5. Placement: which part of a worker is used + +A worker's admitted devices are grouped by the memory node they hang off, and +one of those groups is taken. The rest are left behind and recorded as refused. + +**The buckets exist only where an admitted device is.** A memory node with 64 +cores, 128 GiB of huge pages, and a 100 GbE port but no admitted device does not +appear in the ranking at all, and cannot be chosen. The CPU, huge-page, and NIC +readings are attached to buckets that already exist and are otherwise discarded. +`placement.go:139-171`. `INVENTED`, asserted by construction with no comment. + +### 5.1 The ranking, in order + +Applied only when two or more buckets exist. Every key is compared in turn until +one of them separates the two nodes. + +| # | Key | Direction | Provenance | Stated reason | +|-----|-------------------------------------------|------------|------------|---------------------------------------------------------------------------------------------------------------------------| +| 1 | Admitted device **count** | descending | `GROUNDED` | Usable space is bounded by the erasure-coding stripe, which is a count of devices, so four small disks beat one large one | +| 2 | Combined **capacity** in bytes | descending | `ASSERTED` | "Capacity breaks a tie in the count" | +| 3 | **Physical cores** of the node | descending | `ASSERTED` | "Cores break a tie in capacity" | +| 4 | A real node before the **unknown bucket** | real first | `GROUNDED` | The bucket is numbered -1, so comparing ids alone would rank it above node 0 | +| 5 | **Node id** | ascending | `GROUNDED` | Two runs against one worker have to choose the same node | + +`placement.go:101-120`. + +Key 4 is the one worth reading twice. The unknown bucket's id is `-1`, and key 5 +is ascending, so removing key 4 would make "no memory node in particular" win +every full tie against node 0. It is not merely theoretical: a worker whose CPU +topology could not be read has zero cores against every bucket, so keys 1 to 3 +can all tie, and the comparison falls through to it. + +**Two questions this ranking does not ask**, both recorded in the code as known +and both with a cost: + +- Where the data NIC is. A node with four disks and no fast NIC is preferred + over one with three and a 100 GbE port. +- Whether huge pages are reserved on the node. A node with the disks and no + reserved memory cannot start SPDK at all, and this ranks it first. + +### 5.2 What is read and not ranked on + +| Reading | Carried as | In the key? | +|--------------------------|---------------------------------------|-------------| +| Online CPUs per node | `NUMANodeResources.OnlineCPUs` | No | +| Free huge pages per node | `NUMANodeResources.FreeHugePageBytes` | No | +| Fastest NIC per node | `NUMANodeResources.FastestNICMbps` | No | +| Memory per node | nothing, the struct has no field | No | + +The per-node memory reading is collected by the probe for a stated purpose, that a +storage node pinned to a socket draws its memory from that socket, and the +placement never consults it. + +### 5.3 The NIC reading, and what a bond does to it + +The fastest NIC per node is computed from the **raw** `speedMbps` against the +**raw** `numaNode` of every interface in the report, with no stack traversal and +no kind filter. `placement.go:167-171`. + +A bond, a team, a VLAN, a macvlan, a bridge, and a veth all sit under +`devices/virtual`, and the interface reader returns early for those before it +reads a memory node, so every one of them arrives carrying `numaNode: -1`. The +consequence is exact: an aggregate's speed is credited to the unknown bucket, +and only when that bucket happens to exist because some admitted device also +reported no memory node. Otherwise, the reading is dropped. + +No real memory node is ever credited with a bonded or tagged link as such. What +saves the reading in practice is that the report carries the bond's members too, +each with a real node and a real speed, so the members are counted individually. + +This is the one place the stack resolution added for the management interface +was not applied: `netstack.go` resolves an aggregate through its members and +reports a memory node only where they agree, and `NUMANodeBreakdown` does none +of it. + +### 5.4 The sentence the placement produces + +Every chosen worker carries a sentence saying what was chosen and why, and every +device left behind is recorded as refused with that same sentence as its reason. + +Three defects in it, all observable in the recorded expectations: + +1. It always says "devices" plural, so a one-device winner reads `carries 1 + unclaimed devices`. +2. It names only the runner-up. A three-node machine's third node is never + mentioned. +3. **It asserts the count and capacity comparison even when neither decided.** A + tie settled by key 3, 4, or 5 renders as `NUMA node 0 carries 2 unclaimed + devices (2T) against 2 (2T) on no memory node in particular, so it was + chosen`, which is equal numbers either side of a "so." + +`Worker.PlacementReason` holds the sentence and **nothing in the repository +reads it**. Its only path to a reviewer is as the reason on the per-device +refusals, which become `DeviceDeclined` events. + +### 5.5 Leftover devices + +A device that survived every rule and was not taken by the placement is recorded +as a refusal whose rule is the placement's name and whose reason is the whole +winner sentence, identical on every leftover device of that worker. + +Membership is tested by **device name equality**, not by identity or address. + +These refusals are not pre-filters, so they would reach the per-worker +explanation, but the explanation is only assembled when the run produced no +node sets at all, and a worker with leftovers is by construction in the plan. So +they surface as events, inflate the refusal count in the run's summary, and do +not count toward the device count. + +**A knock-on worth checking:** leftover devices are absent from the worker's +address list, which is what the grouping signature is computed over. Two +identical machines placed on different memory nodes therefore hand over +different address sets and land in different groups. + +### 5.6 The alternative placement + +`AllDevices` takes every admitted device regardless of memory node. It is not +used by the run, because the controller always builds the default, and it is a seam +for a fleet whose machines have one memory node. + +Its downstream effect is worth stating because it is not local: a worker whose +devices span nodes has no single chosen node, which makes the cluster block's +vCPU count fall to the API floor and its huge-page size go unset for the whole +fleet, and the note the reviewer reads then says the worker "has no huge pages +reserved on the memory node it was placed on," which is not what happened. +## 6. Grouping and node sets + +### 6.1 What makes two workers one group + +A group's device selection is shared by every worker in it, so grouping is a +claim about the machines rather than a presentation choice. Two workers share a +group when three things match exactly: + +1. The device class of the run. +2. The **name** of the management interface. +3. The sorted, deduplicated list of device addresses. + +Hashed together, in that order. `SPECIFIED` for grouping by identical hardware, +which the design calls a guess at intent that a reviewer regroups. The +management interface is in the key with a `GROUNDED` reason: a `NodeGroup` names +one interface for every worker it lists. + +**What is deliberately not in the key**, and what each omission costs: + +| Not in the key | Consequence | +|----------------------------------------------------|-------------------------------------------------------------------------------------------------------------| +| Device **sizes** | A 1.92 TB fleet and a 3.84 TB fleet at the same slots are one group, named for the first worker's capacity | +| Device **count**, as distinct from the address set | One controller with two namespaces and one with a single namespace group together | +| NUMA topology | No document field expresses it, so nothing is lost at the group level | +| Model, vendor, serial, rotational | A mixed-model fleet at identical slots is one group. Model can only exclude devices, never split groups | +| Kube role | Deliberate: the split by role happens after grouping, so identical infra machines stay one group | +| Memory, CPU, huge pages | Two machines with identical disks and very different RAM are one group, and `spdkSystemMemory` is never set | +| Interface kind, speed, members | Two workers whose `eth0` is a bare NIC on one and a bond on the other are one group | + +**Address handling.** Addresses are deduplicated (`GROUNDED`: a controller with +two namespaces reports two devices at one address) and sorted ascending. The sort +normalizes probe enumeration order, so two workers whose kernels enumerated the +same disks differently still group together. + +Case is **not** normalized. A probe reporting `0000:5E:00.0` and another +reporting `0000:5e:00.0` produce different signatures and therefore different +groups, while the allow-and-deny rule *does* fold case. `INVENTED`, and +inconsistent with the filter's treatment of the same string. + +### 6.2 Ordering and numbering + +Groups are ordered ascending by their first worker's name, and numbered from one +in that order. `GROUNDED`: two runs over one fleet produce the groups in the same +order and their generated names are stable. + +The comparison is lexicographic, so `worker-10` sorts before `worker-2`. A +stability limit the comment does not state: numbering is stable across re-runs +over an **unchanged** fleet only. Adding a machine that sorts earliest shifts +every later group's number. + +### 6.3 The group name + +The pattern is `group---x`, for example +`group-1-nvme-4x3.492T`. `GROUNDED`: it is positional rather than derived from +the hardware, because the name is what a reviewer edits and `group-1` invites +that where a hash does not. + +The size is the **first worker's** total device bytes divided by the number of +deduplicated addresses. Consequences: + +| Fleet | Rendered | Note | +|--------------------------------------|--------------|-----------------------------------------------------------------------------------| +| 4 × 3.84 TB at four slots | `4x3.492T` | Binary divisor under an SI-looking letter | +| One controller, two 1 TiB namespaces | `1x2T` | States 2T for something the draft names once | +| One 2 TiB disk and one 1 TiB disk | `2x1.5T` | An average no disk has | +| Claimed userspace controllers | `8xunsized` | `GROUNDED`: naming it `0 B` would state a capacity where there is only an absence | +| 10 240 TiB | `1.024e+04T` | Exponent notation inside a group name | + +**A false claim in the code.** The comment on the size computation says it is +"the same for every worker in it by construction." The signature hashes addresses +and not sizes, so it is not. This is the one place in that file where a stated +rationale asserts more than the code guarantees. + +### 6.4 Node sets + +One node set per role, ordered infra, then workers, then control-plane. +`GROUNDED`: infra nodes are almost always the ones somebody meant to be the +storage, and on OpenShift they do not count against a subscription's core limit, +so proposing them first in a block of their own lets a reviewer take that +placement by deleting the other block. + +Names are `infra`, `discovered`, and `control-plane`. Splitting happens **after** +grouping, `GROUNDED`, so two infra nodes with identical disks stay one group. + +**A consequence of that order:** a hardware group whose members span two roles +produces two `NodeGroup`s **with the identical name** in two different node sets, +each still stating the address count and size computed before the split. Names +are unique within a set and not across the document. + +## 7. The management interface + +### 7.1 The ladder + +Applied per interface, in report order, which is ascending by name. + +| # | Rung | Provenance | +|-----|-------------------------------------------------------------------------|------------------------------------| +| 1 | The kind must be bindable | `INVENTED` | +| 2 | The state must be `up`, `unknown`, or empty | `ASSERTED` | +| 3 | A bridge or an overlay is refused unless it holds the node's address | `ASSERTED` | +| 4 | It must hold at least one reachable address | `GROUNDED` | +| 5 | The interface holding the node's `InternalIP` wins, and returns at once | `GROUNDED` | +| 6 | Otherwise: resolved speed descending, kind simplicity, then name | `ASSERTED`, `INVENTED`, `GROUNDED` | + +**Bindable** is exactly physical, bond, team, VLAN, VXLAN, macvlan, ipvlan, and +bridge. Loopback, an unidentified virtual device, and any unrecognized kind are +refused. Refusing loopback and unidentified devices is argued. Admitting the rest is an +enumeration with no stated reason. + +**Reachable** refuses an unparsable address, link-local unicast and multicast, +loopback, and the unspecified address. `GROUNDED`: the control plane refuses a +node whose management interface it cannot find an IP on, and it refuses it inside +the `node_add` task rather than at the request. + +**Simplicity tiers** are physical 0, bond and team 1, VLAN and macvlan and ipvlan +2, everything else 3. The *role* of the key is argued, the tiers themselves are +not, and the stated metric ("the fewest layers between the address and the wire") +does not produce them: a VLAN over a bond is two layers and a bare VLAN is one, +and both are tier 2. Tier 3 is unreachable given the bindable set. + +### 7.2 Resolving through the stack + +An aggregate reports no speed, no slot, and no memory node of its own, so those +are resolved downward through its members. + +- **Speed:** an interface's own reading wins if it is above zero, and the walk + stops there. Otherwise, an aggregate sums its members and a derived interface + takes its first parent with a speed. `GROUNDED` for the split, `ASSERTED` for + the self-report short circuit. +- **Members:** the physical interfaces reachable downward, sorted and + deduplicated. Empty for an interface that is itself physical. +- **Memory node:** the node every member agrees on, and unknown when they do not. + `GROUNDED`: a bond whose members are in two sockets has no affinity, and + reporting one of them would claim an affinity the interface does not have. + +### 7.3 The interaction that matters most + +**The eligibility gate runs before the node-address win.** An interface holding +the cluster's own address is skipped entirely, not even kept as a candidate, when +it is down, of an unidentified kind, or holding no otherwise-reachable address. + +The adjacent comment states the opposite: "when it is known the interface holding +it wins outright. That is not a preference among equals." `INVENTED`, and no test +covers it. + +What happens instead: the draft names some other interface on a different +network, with a reason reading "it is the fastest interface holding a reachable +address," which does not mention that the cluster's own address was elsewhere. +That is precisely the outcome the rule was written to prevent. + +### 7.4 What the choice records, and what is read + +The chosen interface carries its kind, its members, its resolved speed, its +memory node, and a sentence explaining the choice. **Only the name is ever read** +downstream, for the group signature and the document's `mgmtInterface`. + +The sentence's stated purpose is "the record a reviewer reads," and it is +rendered in no template, event, or status. In the one case a reviewer most needs +explained, a worker with no serving interface, the sentence is **empty** and no +refusal is recorded. +## 8. The cluster block + +Two of the cluster's fields are required by the API and cannot be read off a +worker. The file that derives them states its own standard: nothing there has to +be right, it has to be plausible, stated, and easy to correct. + +| Field | Value | Provenance | +|------------------------|----------------------------------------------------------------|------------| +| `name` | the document's name plus `-cluster` | `ASSERTED` | +| `maxSubsystemCount` | always 30 | `GROUNDED` | +| `enableDriveFormat` | always true | `GROUNDED` | +| `vcpuCount` | the smallest chosen memory node's physical cores, floored at 4 | `GROUNDED` | +| `minHugePagesSize` | the smallest free huge pages on any chosen node, or unset | `ASSERTED` | +| `socketsToUse` | never set | `INVENTED` | +| `nodesPerSocket` | never set | `INVENTED` | +| `stripe` | never set | `INVENTED` | +| `fabricType` | never set | `INVENTED` | +| `enableFailureDomains` | never set | `INVENTED` | + +**`maxSubsystemCount`** is the middle of the API's range, chosen so that a +reviewer who has not thought about it gets a working cluster and one who has can +see the number was not derived. Nothing a probe reports bears on it. + +**`enableDriveFormat`** is on the document because it is destructive and the +document is what somebody approves. The cluster's own field is immutable once the +cluster exists, so a default nobody saw could not be undone. + +**`vcpuCount`** takes the smallest worker because the control plane assumes the +count uniform across a cluster's nodes. The floor of 4 is the best-grounded +number in the generator: below it, sbcli's core layout assigns no NVMe-oF poller +core at all. Where the smallest worker cannot meet the floor, the floor is used +anyway and the note says the cluster asks for more than that worker has. + +A worker contributes a reading only when **every** chosen device sits on one +memory node. A worker whose devices span nodes contributes nothing, and if no +worker contributes, the value is the floor. + +**`minHugePagesSize`** reads the **free** pages on the chosen node, summed across +page sizes, and proposes the smallest across the fleet. It returns unset as soon +as **any** worker has none, and it truncates to whole gibibytes. + +**The five fields never set** each have a cost. The most expensive is +`socketsToUse`: empty means socket 0 alone, so a two-socket fleet is drafted as a +single-socket deployment and the disks the placement chose on socket 1 are in the +document while the cluster is laid out for socket 0. The layout is immutable on +the cluster it lands on. + +## 9. Assembling and naming the document + +| Rule | Value | Provenance | +|---------------------------------------|-----------------------------------------------------|-------------| +| Document name | `spec.discover.configName`, else `discovered-` | `ASSERTED` | +| `spec.approved` | always false, including on a re-run | `GROUNDED` | +| `spec.environment` | copied from what Inspecting concluded | `SPECIFIED` | +| Growth branch | `clusterRef` set means no cluster block is proposed | `GROUNDED` | +| The filter | never written into the document | `SPECIFIED` | +| An existing document of the same name | adopted, left as it is, with an event | `GROUNDED` | +| Owner reference | none. The document outlives the run | `GROUNDED` | + +**The report objects** are named by a formula holding to 63 bytes rather than the +253 a ConfigMap may have. `GROUNDED` in an observed API server rejection: the Job +that writes a report is named the same way, and a Job's name reaches its pods as +a label, where 63 is the limit. The digest is unconditional because the run and +the node join on a separator both may contain. + +Neither the name nor the labels can be read back as the values that produced +them. The node a report is about is the field inside the report, which is the only +place it appears as the cluster spells it. + +## 10. Rules that need a decision + +Every entry is `ASSERTED` or `INVENTED`: no design states it, and the code gives +no constraint for it. Ordered by what a wrong answer costs. + +| # | Rule | Where | The question | +|-----|---------------------------------------------------------------------------------|-------|-----------------------------------------------------------------------------------| +| 1 | The eligibility gate runs before the node-address win | §7.3 | Should a down link holding the cluster's own address be passed over in silence? | +| 2 | Resolved speed descending is the first sort key for an interface | §7.1 | Is the fastest link the management interface, or the data one? | +| 3 | A bridge or overlay is admitted only with the node's address | §7.1 | Is a bridge holding that address a management interface, or a machine to look at? | +| 4 | An aggregate's speed is the sum of its members | §7.2 | Is a bridge an aggregate for this purpose? Its ports are not one flow's bandwidth | +| 5 | The simplicity tiers | §7.1 | Why does a bond outrank a VLAN, and what is tier 3 for? | +| 6 | Capacity, then cores, as the placement's tie-breaks | §5.1 | Is more capacity the right second key, given the stripe argument for the first? | +| 7 | The placement ignores huge pages and NIC locality | §5.1 | Both are recorded as known. Which should enter the ranking? | +| 8 | Buckets exist only where an admitted device is | §5 | A node with cores, pages, and a fast NIC but no free disk is invisible | +| 9 | `minHugePagesSize` is the smallest **free** pages, unset if any worker has none | §8 | simplyblock allocates its own pages, so this reads a baseline as a requirement | +| 10 | The five cluster fields never set | §8 | `socketsToUse` above all: the second socket of every worker is left out | +| 11 | Deny before allow, and an empty allow list means allow everything | §4.2 | Both are conventions worth stating rather than discovering | +| 12 | Allow and deny fold case, including for block paths | §4.2 | `/dev/SDB` matches `/dev/sdb`, and Linux paths are case-sensitive | +| 13 | Device addresses are grouped case-sensitively | §6.1 | The opposite convention to the filter, on the same strings | +| 14 | A model or size filter silently excludes every claimed controller | §4.4 | They are unsized and unmodeled by construction | +| 15 | The document name is the run's name, not a timestamp | §9 | The field's own documentation says timestamp | + +## 11. Defects found while cataloging + +Each is a contradiction between the code and its own stated intent, or a path +that cannot do what it claims. None of them is a matter of taste. + +| # | Finding | Where | +|-----|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|--------| +| 1 | ~~An unparsable `driveSizeRange` silently admits every device.~~ **Fixed.** An admission webhook on `OperatorOps` refuses the run at the request, and both steps that read the filter refuse it again | §4.5 | +| 2 | **Both "huge pages unset" notes are computed and thrown away.** A draft that omits the field never explains why, against the file's own standard that every derived number carries a sentence | §8 | +| 3 | **The derived cluster name is unbounded against a 63-character limit.** `configName` has no maximum and a run name may be 253, so the create fails and the step retries until its deadline | §9 | +| 4 | **The draft's run label is the raw run name**, where every other use of that label is sanitized. An uppercase letter or 64 characters produces a create that can never succeed | §9 | +| 5 | **The group size comment claims every worker in a group has the same disks.** The signature hashes addresses, not sizes | §6.3 | +| 6 | **The placement's sentence asserts a comparison that did not decide.** A tie settled by cores or node id still reads "carries 2 devices (2T) against 2 (2T), so it was chosen," and says "1 unclaimed devices" for a single disk | §5.4 | +| 7 | **The no-worker failure message says infra nodes are excluded.** They are admitted without an opt-in | §2 | +| 8 | **`InUse: false` is not distinguished from never read.** The report's own comment says a reader must check for an unreadable entry first. Nothing does, and the rule that would is off by default | §4.5 | +| 9 | **A worker whose report cannot be decoded produces no refusal**, only an event, so it is absent from the explanation the status carries | §4.5 | +| 10 | **`Reason` is empty in the one case a reviewer most needs explained**: a worker with no serving interface | §7.4 | +| 11 | **Five of six recorded interface facts have no consumer**, including the sentence whose stated purpose is a record a reviewer reads | §7.4 | +| 12 | **`Upper` is collected, shipped, and read by nothing.** The documented use, a NIC holding no address with the address on a VLAN above it, is never implemented | §7.2 | +| 13 | **A wireless interface is a tier-0 management candidate.** `DEVTYPE=wlan` falls through the kind switch to physical | §7.1 | +| 14 | **A team or a VLAN on a kernel publishing no `DEVTYPE` becomes unbindable.** Bonds and bridges have a directory fallback, these do not | §7.1 | +| 15 | **Carrier and duplex are dropped from the report**, against a header claiming no filtering and no judgment. The rule admits `unknown` state, which is exactly the case carrier would settle | §7.1 | +| 16 | **Five schema limits are unchecked before writing**: groups per node set, workers per group, devices per selection, and the two address patterns. Each surfaces as a create rejection rather than a finding | §6, §9 | +| 17 | **The unknown memory-node bucket sorts first here and last in atlas-lib.** Same sentinel, opposite convention | §5 | +| 18 | **Two constants share one event wire value** in a package whose own premise is that a reason is an API | §9 | + +## 12. Rules that are specified and not built + +| Rule | Stated in | State | +|--------------------------------------------------------------------------------------------|-----------------------------------------------------------------------|-----------| +| `failureDomain` is seeded from `topology.kubernetes.io/zone` and left unset where no label | `design-clusterdeploymentconfig.md` §8.2, and the API field's own doc | `UNBUILT` | + +Nothing in the operator reads that label, and `failureDomain` is never written. +The test plan's `U-52` and `U-53` are the rows for it, both unimplemented. + +It is not cosmetic. A **growth** document against a cluster with +`enableFailureDomains` set is generated already failing its own validation, +because every group is required to carry a domain and none does. A creating +document escapes only because the generator never sets `enableFailureDomains` +either. diff --git a/discovery-generator-test-cases.md b/discovery-generator-test-cases.md new file mode 100644 index 000000000..01f54b4cb --- /dev/null +++ b/discovery-generator-test-cases.md @@ -0,0 +1,491 @@ + + +# Discovery generator: test case mutations + +**Scope.** The generator is the path from probe reports to a drafted +`ClusterDeploymentConfig`: `nodeprobe.ReportFromConfigMap` → +`discovery.Planner.Plan` → `discovery.ClusterTemplateFor` → +`OperatorOpsReconciler.draftFor`. Everything after approval, meaning validation +and expansion into a `StorageCluster` and `StorageNode` objects, is the +`ClusterDeploymentConfig` controller's other half, covered by +`operator/docs/tests/test-plan-clusterdeploymentconfig.md` §1 (`U-01` onward). +Nothing here duplicates those IDs. + +**What one case is.** A directory of input objects and one expected output: + +```text +operator/internal/controllers/deployment/testdata/discovery/ + net/ + net-13-bond-holds-node-address/ + case.md # the row this directory is, copied from this document + ops.yaml # the OperatorOps, carrying spec.discover and its deviceFilter + nodes.yaml # corev1.Node objects: labels, taints, addresses, capacity + reports/.yaml # one ConfigMap per worker, the report JSON under report.json + expected.yaml # the ClusterDeploymentConfig the run should write + expected-refusals.txt # Plan.RefusalLines(), one per line, absent when there are none + net-14-.../ + numa/ + size/ +``` + +One directory per case, under its family. The family level is for reading rather +than for the harness, which walks the tree and takes every directory holding a +`case.md`: a hundred and seventy directories in one listing is a set nobody +reviews, and eleven families of a dozen is. + +The directory name carries the case identifier and a slug of the mutation, so +that a failure names the case without anybody looking it up. + +A case with no `expected.yaml` is a refusal case: `expected-error.txt` holds the +message the run fails with, which is what `kubectl get operatorops` shows. + +**Two harnesses.** The `Harness` column names which one a row belongs to. + +| Harness | Where | Covers | +|---------|------------------------------------------------------------------|-----------------------------------------------------------------| +| `CM` | `deployment` package, table-driven over the testdata directories | The whole path, report decoding and the draft's YAML included | +| `GO` | `discovery` package, built reports | Cases needing a `Planner` seam the controller never substitutes | + +A `GO` row exists because the controller always builds the default `Planner`: +`AllDevices`, `SingleNodeSet`, and `WorkerWasReadable` are reachable only by +constructing one directly. + +--- + +## 1. Mutation axes + +| Axis | Values exercised | +|------------------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| Device class | NVMe only, logical block only, both on one worker, neither | +| Device count per node | 0, 1, 4, 10, and one worker at 128 (the schema's `MaxItems`) | +| Device presentation | Kernel-presented disk, partition, loopback, multi-namespace controller, idle userspace controller, held userspace controller, attached simplyblock volume, iSCSI LUN | +| Device state | Available, mounted, held open, in a device-mapper stack, partitioned, unreadable | +| NUMA topology | 1, 2, 4, and 8 memory nodes. Devices on one node, split evenly, split unevenly, and on none | +| vCPU count | 2, 8, 32, and 196 logical CPUs, with and without simultaneous multithreading | +| Memory | 4 GiB, 64 GiB, 256 GiB, and 1 TiB total, with and without a per-node breakdown | +| Huge pages already allocated | None at all, no hugetlbfs, a pool with no free pages, 256 GiB free per node, a total without a per-node breakdown, a pool only on the node not chosen, 2 MiB and 1 GiB pools together | +| Huge-page headroom | Room for simplyblock's allocation on top of the existing pool, exactly enough, and short of it | +| Reuse instruction | A run told to take a set number of the pre-allocated pages, and a run told nothing | +| PCI addressing | Identical across the fleet, identical across a subset of it, one odd worker out, differing per worker, one address set a subset of another, unsorted, two buses, a five-digit domain | +| Interface count | One interface, one that is only loopback, two physical, and a dozen of mixed kinds | +| Interface kind | Physical, bridge, bond, VLAN, VLAN over bond, VXLAN, macvlan, veth, and loopback | +| Interface link | Up, down, dormant, unknown state, 1G, 10G, 100G, unknown speed, MTU 1500 and 9000 | +| Interface address | A routable v4 address, a global v6 address, link-local only, none, and the node's own `InternalIP` | +| Interface instruction | A run told which interface is management, told which are data NICs, and told neither | +| Stack resolution | An aggregate reporting no speed of its own, a derived interface over one, members on one memory node and on two, and a stack that points at itself | +| Fleet size | 0, 1, 3, and 32 workers | +| Fleet homogeneity | Uniform, two layouts, a majority layout with stragglers, one odd worker out, every worker distinct, a mix of NUMA topologies | +| Node role | Unlabeled, `worker`, `infra`, `control-plane`, `master`, and machines carrying two | +| Filter | Each of the seven `DeviceFilter` fields, singly and in the two combinations the CEL rule permits | +| Document shape | Creating (a cluster template) and growing (`spec.discover.clusterRef`) | +| Report validity | Current version, an older version, absent key, malformed JSON, no node name, a foreign node | + +--- + +## 2. Device class and inventory shape + +| ID | Mutation | Expected | Harness | +|--------|-------------------------------------------------------------|-------------------------------------------------------------------------------------------|---------| +| DEV-01 | 4 NVMe disks, one memory node, NVMe run | 1 group, `devices.nvme` holds the 4 addresses ascending, name `group-1-nvme-4x3T` | `CM` | +| DEV-02 | 4 virtio disks, block run | 1 group, `devices.block` holds the 4 paths, class `block` | `CM` | +| DEV-03 | 4 virtio disks, NVMe run | No node sets. Every device refused by `device class`, the worker by `has devices` | `CM` | +| DEV-04 | 2 NVMe and 2 virtio disks on one worker, NVMe run | Only the 2 PCI addresses reach the group | `CM` | +| DEV-05 | The same worker, block run | Only the two virtio disks. The NVMe pair is refused for being the other class | `CM` | +| DEV-06 | One controller presenting two namespaces at one PCI address | The address is named once. `DeviceCount` is 1 for that controller | `CM` | +| DEV-07 | A SATA disk beside an NVMe one, block run | Only the SATA disk. The NVMe run takes only the NVMe disk | `CM` | +| DEV-08 | A rotational HDD beside an SSD, both free | Both admitted: no rule reads `Rotational`. See §14, gap G-2 | `CM` | +| DEV-09 | 10 NVMe disks, one memory node | All 10 in one group, name `group-1-nvme-10x3T` | `CM` | +| DEV-10 | A partition and a loopback device beside 4 disks | Both pre-filtered: absent from `Plan.Explain`, present in `RefusalLines` | `CM` | +| DEV-11 | 128 NVMe disks on one worker | One group at the selection's `MaxItems`. The document still applies | `CM` | +| DEV-12 | A disk whose `Kind` is `disk` and whose `SizeBytes` is 0 | Admitted. The group is named `unsized` | `CM` | +| DEV-13 | A simplyblock volume attached to the worker, NVMe run | Refused as this fleet's own volume, naming the cluster it belongs to | `CM` | +| DEV-14 | The same volume on a block run | Refused the same way, since the rule is in both pipelines and reads no filter | `CM` | +| DEV-15 | A worker whose every disk is an attached volume | No draft, and the explanation counts them together rather than listing each | `CM` | +| DEV-16 | A fabric namespace another product exported | Refused for being on a fabric, not as a simplyblock volume: the NQN does not parse as one | `CM` | +| DEV-17 | An iSCSI LUN beside a virtio disk, block run, no allow list | Only the virtio disk. A LUN is storage across a network and is never taken by default | `CM` | +| DEV-18 | The same worker with the allow list naming the LUN | Both disks. Naming it is the decision a run cannot make for a fleet | `CM` | +| DEV-19 | An iSCSI LUN on an NVMe run | Refused for being the other class, before the iSCSI rule is reached | `CM` | + +## 3. NUMA topology + +The placement ranks a worker's memory nodes by unclaimed device count, then +capacity, then physical cores, then a real node ahead of the unknown bucket, +then the node id. Each row below pins one of those tie-breaks. + +| ID | Mutation | Expected | Harness | +|---------|------------------------------------------------------------------------------|-------------------------------------------------------------------------------------------|---------| +| NUMA-01 | 1 memory node, 4 disks on it | All 4 used, reason "every unclaimed device is on NUMA node 0" | `CM` | +| NUMA-02 | 2 memory nodes, 2 disks each, equal size, equal cores | Node 0 chosen on the id tie-break. 2 disks refused by the placement | `CM` | +| NUMA-03 | 2 memory nodes, 1 disk on node 0 and 3 on node 1 | Node 1 chosen on count. The single disk refused | `CM` | +| NUMA-04 | 2 memory nodes, 2 × 8 TiB on node 0 and 3 × 1 TiB on node 1 | Node 1 chosen: count beats capacity | `CM` | +| NUMA-05 | 2 memory nodes, 2 disks each, 1 TiB on node 0 and 2 TiB on node 1 | Node 1 chosen on the capacity tie-break | `CM` | +| NUMA-06 | 2 memory nodes, 2 disks each of one size, 8 cores on node 0 and 24 on node 1 | Node 1 chosen on the core tie-break | `CM` | +| NUMA-07 | 4 memory nodes, 2 disks each (NPS4) | Node 0 chosen. 6 disks refused. `vcpuCount` is node 0's physical cores | `CM` | +| NUMA-08 | 8 memory nodes, 1 disk each but 3 on node 5 | Node 5 chosen on count | `CM` | +| NUMA-09 | Every device reports `numaNode: -1` | The unknown bucket is used. `vcpuCount` floors at 4, `minHugePagesSize` unset, both noted | `CM` | +| NUMA-10 | 2 disks on node 0 and 2 on `-1`, node 0 carrying cores | Node 0 chosen on cores. The unknown bucket is never preferred | `CM` | +| NUMA-11 | The same worker with `Placement: AllDevices` | Every disk used. No single chosen node, so `vcpuCount` floors and huge pages go unset | `GO` | +| NUMA-12 | A 1-node worker and a 2-node worker whose chosen addresses coincide | One group: grouping reads addresses and the interface, never the topology | `CM` | +| NUMA-13 | 2 memory nodes, all 10 disks on node 1 and none on node 0 | Node 1 chosen with nothing to choose. The CPU note names node 1's cores | `CM` | +| NUMA-14 | A large worker: 2 nodes × 49 cores, 5 disks each, huge pages on both | 5 disks in the draft, `vcpuCount` 49, `minHugePagesSize` from node 0's free pages alone | `CM` | +| NUMA-15 | `cpu.numaNodes` empty while devices report nodes 0 and 1 | Placement falls through to the id tie-break. `vcpuCount` floors at 4 with its note | `CM` | + +## 4. Node size: vCPU, memory, and huge pages + +simplyblock allocates its own huge pages: the node writes the pool it needs, and +`cluster.minHugePagesSize` is the floor on what it writes, rendered into the +per-node configuration as `MAX_HUGE_PAGES_SIZE`. What discovery reads a host's +pool for is therefore not whether SPDK will find pages to consume. It is the +baseline the new total is added to, which is four questions: + +1. Are pages allocated already, and how many? +2. Are there none, on a kernel that has hugetlbfs and on one that has not? +3. Does the machine have memory for simplyblock's allocation on top of the + existing pool? +4. Is the run told to take a stated number of the pre-allocated pages instead of + adding its own? + +The generator answers none of them. It reads the free pages of the chosen memory +node and proposes them as the cluster's floor, and it proposes nothing at all +when a worker has none, on the stated grounds that SPDK consumes pages rather +than reserving them. Every row below records that output, and §14 carries the +seven findings it produces, `G-14` through `G-20`. + +| ID | Mutation | Expected | Harness | +|---------|------------------------------------------------------------------------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|---------| +| SIZE-01 | 8 logical CPUs, 4 physical cores on the chosen node | `vcpuCount: 4`, with the ordinary derivation note, not the floor note | `CM` | +| SIZE-02 | 2 logical CPUs, 2 physical cores | `vcpuCount: 4` with the note saying the cluster asks for more than the worker has. Not a refusal | `CM` | +| SIZE-03 | 196 logical CPUs, 98 physical cores, 1 memory node, 1 TiB RAM | `vcpuCount: 98`, the whole machine. See §14, gap G-3 | `CM` | +| SIZE-04 | 196 logical CPUs across 2 nodes, 1 TiB RAM, 512 GiB per node | `vcpuCount: 49`, the chosen node's cores | `CM` | +| SIZE-05 | A fleet whose chosen nodes carry 4, 16, and 49 cores | `vcpuCount: 4`, the note naming the smallest worker | `CM` | +| SIZE-06 | 4 GiB total, 900 MiB available, 4 free disks | Admitted unchanged: no rule reads memory. See §14, gap G-4 | `CM` | +| SIZE-07 | `memory` zero and an `unreadable` entry for it | Admitted by default. The draft is written from the disks | `CM` | +| SIZE-08 | The same report with `WorkerRules: WorkerWasReadable` | The worker is refused, the message quoting the unreadable readings | `GO` | +| SIZE-09 | 512 × 1 GiB pages pre-allocated, 256 free per node | **Contested.** `minHugePagesSize: 256G`: another workload's reservation becomes simplyblock's own floor. See §14, gap G-14 | `CM` | +| SIZE-10 | One worker of three with no pre-allocated pages on its chosen node | **Contested.** Unset for the whole cluster, the note naming that worker. Nothing pre-allocated is the ordinary case, not a reason to propose nothing. See §14, gap G-15 | `CM` | +| SIZE-11 | 256 × 2 MiB pages free on the chosen node, 512 MiB in all | Unset, the note saying the smallest reservation is under a gigabyte. A 1.5 GiB reading truncates to `1G` | `CM` | +| SIZE-12 | A pool with a total and an empty `numaNodes` | Unset: the reservation exists and its distribution is unknown | `CM` | +| SIZE-13 | Pages pre-allocated on node 1 while the placement chooses node 0 | Unset, the note naming the worker. The placement does not read huge pages, and the pool on the other node is neither counted nor reported | `CM` | +| SIZE-14 | Swap total 8 GiB, swap free 0 | No effect on the draft. See §14, gap G-5 | `CM` | +| SIZE-15 | `nodes.yaml` allocatable memory far under capacity | Recorded on the `KubeNode`, no effect on the draft | `CM` | +| SIZE-16 | 256 GiB RAM, 200 GiB already in huge pages, 40 GiB available | **Contested.** `minHugePagesSize: 200G` and no arithmetic against the 40 GiB left. See §14, gap G-16 | `CM` | +| SIZE-17 | 1 TiB RAM, nothing pre-allocated, the whole machine available | **Contested.** Nothing proposed, though the room for an allocation is there and reported. See §14, gaps G-15 and G-16 | `CM` | +| SIZE-18 | A run meant to take 128 of the host's 256 pre-allocated pages | **Contested.** No field expresses it. `spec.discover` carries no huge-page input, and neither `minHugePagesSize` nor `spdkSystemMemory` is written by a run. See §14, gap G-17 | `CM` | +| SIZE-19 | A pool of 256 pages, all promised to a mapping: `resv` 256, `free` 0 | Indistinguishable from an untouched pool of 256, because the report drops `resv_hugepages`. See §14, gap G-18 | `CM` | +| SIZE-20 | A 1 GiB pool and a 2 MiB pool, both with free pages on the chosen node | The two sizes are summed into one figure, and nothing says which size the node should take | `CM` | +| SIZE-21 | No hugetlbfs at all, `hugePages` absent from the report | The same unset output as a machine whose pools are full, so the two are not told apart. See §14, gap G-19 | `CM` | + +## 5. PCI addressing and disk count + +| ID | Mutation | Expected | Harness | +|--------|--------------------------------------------------------------------|------------------------------------------------------------------------------------------------------------|---------| +| PCI-01 | 32 workers, identical addresses | 1 group of 32 workers | `CM` | +| PCI-02 | 2 workers, 4 disks each, different slots | 2 groups of 1 worker each | `CM` | +| PCI-03 | 32 workers: 16 on layout A and 16 on layout B | 2 groups of 16, ordered by their first worker's name | `CM` | +| PCI-04 | 32 workers, every layout distinct | 32 groups in one node set, under the 64-group ceiling | `CM` | +| PCI-05 | Addresses reported in descending order | `devices.nvme` ascending: the draft sorts | `CM` | +| PCI-06 | 10 disks across two buses, `0000:5e:*` and `0000:af:*` | One group, addresses ascending across both buses | `CM` | +| PCI-07 | A five-digit PCI domain, `10000:01:00.0` | **Contested.** The generator names it and the schema's item pattern refuses the document. See §14, gap G-6 | `CM` | +| PCI-08 | Uppercase hex in the report, `0000:5E:00.0` | Named as reported. Grouping is byte-exact, so a fleet mixing cases splits | `CM` | +| PCI-09 | Workers named `worker-1` … `worker-32` with two layouts | Group numbering follows lexicographic worker order: `worker-1`, `worker-10`, `worker-11` | `CM` | +| PCI-10 | 5 workers: 3 on layout A, the other 2 each distinct | 3 groups, of 3, 1, and 1 workers: partial homogeneity still groups what it can | `CM` | +| PCI-11 | 32 workers: 20 on layout A, 6 on layout B, 6 each distinct | 8 groups, of 20, 6, and six of 1, numbered by their first worker's name | `CM` | +| PCI-12 | 8 workers uniform but for one whose fourth disk is in another slot | 2 groups, of 7 and 1. One slot moved is a second group, which is the guess a reviewer regroups | `CM` | +| PCI-13 | 2 workers on one layout carrying 2 TiB and 4 TiB disks | **Contested.** One group, named `group-1-nvme-2x2T` for the first worker's capacity. See §14, gap G-13 | `CM` | +| PCI-14 | 3 workers, the third holding 3 of the 4 slots the others hold | 2 groups: the address set matches whole or not at all, never as a subset | `CM` | + +## 6. Fleet size and grouping + +| ID | Mutation | Expected | Harness | +|----------|-------------------------------------------------------------|----------------------------------------------------------------------------|---------| +| FLEET-01 | 1 worker, 4 disks | 1 group of 1 | `CM` | +| FLEET-02 | 3 uniform workers | 1 group of 3. The summary reads 3 workers, 12 devices, 1 group, 1 node set | `CM` | +| FLEET-03 | 32 uniform workers | 1 group of 32, under the 200-worker ceiling | `CM` | +| FLEET-04 | No reports at all | The run fails: the probe reports are gone | `CM` | +| FLEET-05 | 32 workers, 3 with every disk mounted | 29 workers drafted. 3 worker refusals with their device reasons folded in | `CM` | +| FLEET-06 | The same 32 reports, ConfigMaps listed in a different order | A byte-identical document | `CM` | +| FLEET-07 | Two ConfigMaps carrying a report for one node | One report per node. The draft names the worker once | `CM` | +| FLEET-08 | A report for a node absent from `status.workers` | Skipped without an event: it is not this run's evidence | `CM` | + +## 7. Network interfaces + +The draft names one management interface per group, and the rule is a ladder +rather than a match, because a fleet's machines do not agree on what their NICs +are called. What `ManagementInterface` applies today, in order: + +1. Refuse a kind nothing can be bound to: loopback, and a virtual device the + kernel does not identify, which is every veth and dummy on the machine. +2. Refuse anything whose state is neither `up` nor `unknown`. +3. Refuse anything holding no reachable address, which is to say link-local, + loopback, and unspecified addresses do not count. +4. Refuse a bridge or an overlay that does not hold the node's own address. + Both kinds can be bound and both are ordinarily somebody else's network: a + CNI bridge and a flannel overlay carry the pod network, and a hypervisor + host's bridge carries the address the cluster reaches the machine on. Which + of the two a given device is can be read from nothing but the address on it. +5. Of what survives, the interface holding the node's `InternalIP` wins + outright, whatever kind it is. +6. Failing that, the fastest link, with the simpler kind breaking that tie and + the lowest name breaking the rest. A bond and a VLAN report no speed of + their own, so the speed is resolved down the stack: an aggregate carries the + sum of its members, and a derived interface carries what its parent carries. + +Rungs 1, 4, and 6 are new, and they are what the interface kind in report +version 4 bought. Before it, every software interface read as "virtual" and was +refused, so a bonded or tagged management network, which is the ordinary +enterprise host, yielded no interface on any worker. + +Two things the ladder does not cover, and both are inputs rather than readings: +which interface a fleet means for management, and which interfaces it means for +the data plane. Neither is expressible today (see `G-23` and `G-24`), so the +rows for them record what a run does in their absence. + +### Where each rung came from + +No design document specifies any of this. `design-clusterdeploymentconfig.md` +shows `mgmtInterface: eth1` in an example and states no rule for arriving at it, +so the ladder was assembled from a control-plane failure and from judgment. The +two are not the same thing, and a fixture whose outcome turns on the second is +recording a proposal rather than an expectation. + +| Rung | Where it came from | Status | +|-------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------|--------------------| +| A node with no management address is refused | The control plane's own refusal, "No management interface with IP found in provided interfaces" (commit 69e9f3f2) | Grounded | +| The interface holding the node's `InternalIP` wins outright | The operator addresses the worker by that address everywhere else, so any other choice splits one deployment across two networks | Grounded | +| A link that is down is passed over | A link that is down keeps its address and carries nothing | Grounded | +| A link-local address does not count | It is configured without anybody assigning it and routes nowhere | Grounded | +| The cluster's own plumbing never wins | Stated as intent in commit 69e9f3f2. It was implemented as "anything virtual," which is what `G-21` records as wrong | Grounded as intent | +| The lowest name breaks a tie | Two runs over one fleet have to produce the same document. Which tiebreak is arbitrary; that there is one is not | Grounded | +| **The fastest link wins when no node address matches** | Asserted in commit 69e9f3f2 with no reason given | **Unratified** | +| **Which kinds can be bound at all** | Added with the kind reading. Loopback and an unidentified virtual device are safe to refuse; admitting the rest is a choice | **Unratified** | +| **A bridge or an overlay only with the node's own address** | Added with the kind reading | **Unratified** | +| **An aggregate carries the sum of its members** | Added with the stack resolution | **Unratified** | +| **A simpler kind breaks a speed tie** | Added with the stack resolution | **Unratified** | + +Twenty of the forty-two cases in this section carry no node address on a +candidate interface, so an unratified rung decides them. Their recorded +documents are evidence that the rule was applied consistently and no evidence +that the rule is right. + +The question each unratified rung is really asking: + +1. **Is the fastest link the management interface?** Management traffic is + light, and the fastest link is what a data path wants. Naming it for + management may be taking the wrong NIC for the wrong plane, which is what + `NET-24` records. The alternatives are the slowest addressed link, the lowest + name outright, and refusing to choose at all where a fleet has not said. +2. **Should a run choose at all when nothing identifies the interface?** A draft + naming the wrong NIC is approved as readily as one naming the right one. The + alternative is to name none and make the reviewer say, which trades a wrong + document for one that cannot be approved unedited. +3. **Is a bridge holding the node's address the management interface, or a + machine a reviewer should look at?** It is an ordinary hypervisor host and it + is also what a misconfigured worker looks like. + +| ID | Mutation | Expected | Harness | +|--------|---------------------------------------------------------------------------------------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------------|---------| +| NET-01 | One 10G interface, up, holding a routable address | `mgmtInterface` names it | `CM` | +| NET-02 | A 1G and a 10G interface, both up and addressed, no node address on either | The 10G is named | `CM` | +| NET-03 | The node's `InternalIP` on the 1G while the 10G is faster | The 1G is named: the node address wins outright | `CM` | +| NET-04 | Four interfaces, 1G and 10G, all unusable: down, bridged, or link-local only | `mgmtInterface` empty and the draft still written. See §14, gap G-7 | `CM` | +| NET-05 | `interfaces` absent from the report | The same: an empty interface and a written draft | `CM` | +| NET-06 | Two workers with identical disks calling their NIC `eth0` and `ens5f0` | 2 groups, which the draft validation then reports as unbuildable. See §14, gap G-25 | `CM` | +| NET-07 | Two 10G interfaces, `eth1` and `eth0`, both addressed | `eth0` is named. A second run names it again | `CM` | +| NET-08 | An addressed physical interface reporting speed 0 beside a 10G interface that is down | The addressed one is named | `CM` | +| NET-09 | A worker whose `InternalIP` sits on a bridge, as a hypervisor host's does | `br0` named. A bridge carrying the cluster's own address is the management network | `CM` | +| NET-10 | One interface holding only a global IPv6 address | Named: reachability is not address-family dependent | `CM` | +| NET-11 | An interface holding both a link-local and a routable address | Named on the routable one | `CM` | +| NET-12 | One interface and it is the loopback | `mgmtInterface` empty: loopback is refused by its ARPHRD type, not its name | `CM` | +| NET-13 | `bond0` holding the node address over two unaddressed physical slaves | `bond0` named, with `eth0` and `eth1` reported as the hardware under it | `CM` | +| NET-14 | `eth0.100` holding the node address, the parent `eth0` unaddressed | `eth0.100` named, resolving to `eth0` for its slot and memory node | `CM` | +| NET-15 | `vxlan.calico` holding an overlay address beside `eth0` holding the node address | `eth0` named. The overlay is refused for holding an address the cluster does not reach the machine on | `CM` | +| NET-16 | `bond0.100` over `bond0` over two slaves, the node address on the VLAN | `bond0.100` named, resolving through the bond to both NICs | `CM` | +| NET-17 | A `macvlan` and an `ipvlan` over `eth0`, all three addressed | `eth0` named: the three carry the same traffic and the physical one is the simplest | `CM` | +| NET-18 | Twelve `veth` interfaces holding pod-CIDR addresses beside one physical NIC | The physical NIC named, whatever the veth count | `CM` | +| NET-19 | `cni0` and `docker0` both addressed, one physical NIC with no address | `mgmtInterface` empty: a bridge is refused and an unaddressed NIC is not a candidate | `CM` | +| NET-20 | An interface whose state is `dormant`, and one whose state is `lowerlayerdown` | Both passed over: only `up` and `unknown` are admitted | `CM` | +| NET-21 | An interface whose state is `unknown` holding a routable address | Named. A driver that does not track carrier is not a reason to refuse it | `CM` | +| NET-22 | A run told that `ens5f0` is the management interface | **Contested.** No input expresses it, so the ranking decides regardless. See §14, gap G-23 | `CM` | +| NET-23 | A run told that `ens5f1` and `ens5f2` are the data NICs | **Contested.** `dataInterfaces` is never written, so the draft leaves the data plane unnamed. See §14, gap G-24 | `CM` | +| NET-24 | Two physical NICs, a 10G and a 100G, both up and addressed | The 100G named for management and no data NIC named at all, so the fast link is proposed for the wrong plane | `CM` | +| NET-25 | Two otherwise equal 10G NICs at MTU 9000 and MTU 1500 | The lower name wins. MTU is reported and not ranked on | `CM` | +| NET-26 | A 100G bridge beside a 10G physical NIC, both addressed | The 10G named: the kind ladder is decided before the speed | `CM` | +| NET-27 | An interface reporting state `up` with no link partner | Named: the report drops `carrier`, so a link with no partner reads as a working one. See §14, gap G-26 | `CM` | +| NET-28 | A bond reporting no speed over two 25G NICs, beside an addressed 10G NIC | The bond named: an aggregate carries the sum of its members | `CM` | +| NET-29 | A VLAN reporting no speed over that bond, beside the same 10G NIC | The VLAN named: a derived interface carries what its parent carries | `CM` | +| NET-30 | The chosen bond's members in sockets 0 and 1 | The memory node is reported as unknown rather than as one of the two | `CM` | +| NET-31 | A veth holding the node's own address | Nothing named: an unidentified virtual device is refused at the first rung, node address or not | `CM` | +| NET-32 | A bond whose `lower` names a second bond whose `lower` names the first | The bond is named and the resolution terminates | `GO` | +| NET-33 | The node's `InternalIP` on a link whose state is `down` | Nothing named: the state rung is applied before the address wins, so a link carrying nothing is refused however right its address is | `CM` | +| NET-34 | The node's `InternalIP` on both `bond0` and `bond0.100` | The first by the report's own order, which is the kernel's and is ascending by name, so two runs agree | `CM` | +| NET-35 | A bond reporting 50000 Mbps of its own over two 25G members | Its own reading is taken, and the members are not summed on top of it | `CM` | +| NET-36 | A bond whose `lower` names an interface the report does not carry | Named, with no members and no resolved speed: a member nothing describes contributes nothing | `CM` | +| NET-37 | A `team` interface over two 25G NICs, holding the node address | Named and resolved as a bond is, both being aggregates | `CM` | +| NET-38 | A bridge over a bond over two NICs, the node address on the bridge | Named, and the members resolve two levels down to the two NICs | `CM` | +| NET-39 | A `macvlan` over `eth0` holding the node address | Named, with `eth0` as its member and `eth0`'s speed as its own | `CM` | +| NET-40 | An interface carrying no `kind`, and nothing marking it virtual | Read as physical, which is what the fields before the kind said about it | `CM` | +| NET-41 | A worker whose only fast NIC is a 2x25G bond, ranked for placement | **Contested.** The memory node's fastest NIC reads as 0 Mbps: the placement reads the raw speed and the raw memory node, neither of which a bond has. See §14, gap G-27 | `CM` | +| NET-42 | Any successful run on a bonded host | **Contested.** The draft names the interface and says nothing about the slots, the speed, or the memory node resolved under it. See §14, gap G-28 | `CM` | + +## 8. Node roles and node sets + +| ID | Mutation | Expected | Harness | +|---------|----------------------------------------------------------|-----------------------------------------------------------------------|---------| +| ROLE-01 | 3 workers, no role labels | One node set, `discovered` | `CM` | +| ROLE-02 | 3 `infra` and 3 `worker` nodes, identical hardware | 2 node sets, `infra` first, one hardware group split across them | `CM` | +| ROLE-03 | A `control-plane` node with free disks beside 2 workers | A third node set, `control-plane`, ordered last | `CM` | +| ROLE-04 | A node labeled `infra` and `worker` | `infra` | `CM` | +| ROLE-05 | A node labeled `control-plane` and `worker` | `control-plane`: the most restrictive role wins | `CM` | +| ROLE-06 | A node labeled `master` | `control-plane`: both spellings are read | `CM` | +| ROLE-07 | No `nodes.yaml` at all | Every worker a `Worker`. The interface is chosen with no address hint | `GO` | +| ROLE-08 | A cordoned node and a tainted node, both with free disks | Both reach the draft. See §14, gap G-8 | `CM` | +| ROLE-09 | A node carrying an unrecognized role label | Treated as a worker: an unknown role is not a refusal | `CM` | +| ROLE-10 | 32 workers across `infra`, `worker`, and `control-plane` | 3 node sets, each holding its role's share of every hardware group | `CM` | + +## 9. Filters + +| ID | Mutation | Expected | Harness | +|---------|----------------------------------------------------------------|-------------------------------------------------------------------------------------------------------|---------| +| FILT-01 | `pcieDenyList` naming the boot slot on a uniform fleet | That address in no group. The refusal appears in `Explain`, not pre-filtered | `CM` | +| FILT-02 | `pcieAllowList` of 2 addresses against 10 disks | Only those 2 reach the group | `CM` | +| FILT-03 | `pcieModel: MZQL2` against a mixed-model worker | Only the matching disks. The refusal quotes both strings | `CM` | +| FILT-04 | `driveSizeRange: 1T-4T` with a 512 GiB disk present | The small disk refused, the range quoted in the reason | `CM` | +| FILT-05 | `driveSizeRange: 2T` against 2 TiB and 1.92 TB disks | Only the exact 2 TiB disks: a bare size is both bounds | `CM` | +| FILT-06 | `driveSizeRange: 2T-1T` | The run is refused, naming the field and what a range looks like. Admission refuses it at the request | `CM` | +| FILT-07 | `blockDenyList: /dev/sda` on a block run | The root disk in no group | `CM` | +| FILT-08 | `blockAllowList` of 2 paths against 6 block devices | Only those 2 | `CM` | +| FILT-09 | `enablePartitionedDevices` with a GPT-only refusal | Admitted. A disk also mounted stays refused | `CM` | +| FILT-10 | A filter that matches nothing on any worker | No node sets. The run fails naming the filter rule per worker | `CM` | +| FILT-11 | An allow list written in uppercase against lowercase addresses | Matched: the comparison folds case | `CM` | +| FILT-12 | One address in both the allow and the deny list | Refused: deny is evaluated first | `CM` | +| FILT-13 | `driveSizeRange` on a block run | It narrows the block class, which is the class being scanned | `CM` | +| FILT-14 | No `deviceFilter` at all | Every free whole NVMe disk reaches the draft | `CM` | + +## 10. Probe-refused devices and userspace controllers + +| ID | Mutation | Expected | Harness | +|---------|--------------------------------------------------------------------|------------------------------------------------------------------------------------------------------------------------------------|---------| +| HELD-01 | Every disk mounted | The worker refused. Its device reasons folded into one line, counted | `CM` | +| HELD-02 | Every disk held in a device-mapper stack | The same shape with the other reason | `CM` | +| HELD-03 | 4 NVMe controllers on `uio_pci_generic`, idle, no block devices | 4 synthesized devices, group `group-1-nvme-4xunsized` | `CM` | +| HELD-04 | The same controllers with `inUse: true` | No devices. The worker refused with the message saying something is driving its disks | `CM` | +| HELD-05 | 2 idle userspace controllers beside 2 kernel-presented disks | 4 addresses in the group, the kernel-bound pair counted once | `CM` | +| HELD-06 | An idle controller on `vfio-pci` | Claimable on the same terms as `uio_pci_generic` | `CM` | +| HELD-07 | A kernel-bound controller whose block device is already reported | Named once, not twice | `CM` | +| HELD-08 | Idle userspace controllers on a block run | No devices at all: a controller with no block device has no path. Worker refused | `CM` | +| HELD-09 | 16 loopback devices and 4 free disks | The 4 disks drafted. The loopbacks absent from `Explain` and present in `RefusalLines` | `CM` | +| HELD-10 | Half the controllers idle and half in use, no block devices | Only the idle half is offered | `CM` | +| HELD-11 | Userspace controllers whose process table the probe could not read | Refused, and the message says whether anything is driving them could not be established. An unchecked controller is not a free one | `CM` | + +## 11. Refusal cases + +| ID | Mutation | Expected | Harness | +|---------|----------------------------------------------------------------------------|-----------------------------------------------------------------------------------|---------| +| FAIL-01 | Every worker reports no devices and no controllers | The run fails: no worker has a device this run would use, one line per worker | `CM` | +| FAIL-02 | Every disk in the fleet mounted | The same, each line carrying the device reasons and their counts | `CM` | +| FAIL-03 | Every report unreadable | `ReportUnreadable` per ConfigMap, then the run fails for want of reports | `CM` | +| FAIL-04 | A filter that excludes every disk | The run fails. The refusal lines name the filter rule | `CM` | +| FAIL-05 | Disks fine, no usable interface anywhere | **Contested.** A draft is written with an empty `mgmtInterface`. See §14, gap G-7 | `CM` | +| FAIL-06 | Every worker under the vCPU floor | **Contested.** A draft is written with `vcpuCount: 4`. See §14, gap G-10 | `CM` | +| FAIL-07 | Every worker at 4 GiB of RAM | **Contested.** A draft is written unchanged. See §14, gap G-4 | `CM` | +| FAIL-08 | A single worker whose CPU tree is unreadable, disks fine | A draft with `vcpuCount: 4` and the note saying no worker reported its cores | `CM` | +| FAIL-09 | One worker of 32 refused, the rest fine | A draft of 31. The refusal is an event, never a failure | `CM` | +| FAIL-10 | A worker whose only disks are partitions, `enablePartitionedDevices` unset | Refused by `whole disk`, pre-filtered, so `Explain` gives the worker line alone | `CM` | + +## 12. Report and ConfigMap validity + +| ID | Mutation | Expected | Harness | +|-------|--------------------------------------------------------------------------|---------------------------------------------------------------------------|---------| +| CM-01 | `version: 4` in one report of three, the schema before the subsystem NQN | `ReportUnreadable` naming both versions. The other two are drafted | `CM` | +| CM-02 | A ConfigMap with no `report.json` key | `ReportUnreadable` saying the probe did not finish writing it | `CM` | +| CM-03 | `report.json` holding malformed JSON | `ReportUnreadable` quoting the parse failure | `CM` | +| CM-04 | A report with an empty `node` | Refused: nothing can be attributed to it | `CM` | +| CM-05 | A ConfigMap without the run label | Not listed, so not read | `CM` | +| CM-06 | A report from a node whose name exceeds 63 characters | Drafted under its full name: the label is truncated and the report is not | `CM` | +| CM-07 | A report of a 128-device worker near the ConfigMap ceiling | Decoded and drafted | `CM` | + +## 13. Document shape and the cluster template + +| ID | Mutation | Expected | Harness | +|---------|---------------------------------------------------------------|---------------------------------------------------------------------------------------------------|---------| +| TMPL-01 | Any successful run | `spec.approved` false, `enableDriveFormat` true, and the note about formatting | `CM` | +| TMPL-02 | Any successful run | `maxSubsystemCount: 30` with the note saying it is not a reading | `CM` | +| TMPL-03 | `spec.discover.configName` unset | The name is `discovered-` and the cluster `discovered--cluster` | `CM` | +| TMPL-04 | An `OperatorOps` name long enough to push the cluster past 63 | **Contested.** The document is written and `CreatingCluster` can never succeed. See §14, gap G-11 | `CM` | +| TMPL-05 | `spec.discover.clusterRef` set | No `spec.cluster`, and `clusterRef` carried through with the growth note | `CM` | +| TMPL-06 | Any run on a two-socket fleet | `socketsToUse` and `nodesPerSocket` unset, so one node per worker on socket 0. See §14, gap G-12 | `CM` | +| TMPL-07 | `spec.environment` from the run's status | Copied verbatim into the document | `CM` | +| TMPL-08 | Any run | No `deviceFilter` anywhere in the document: the resolved list is the record | `CM` | + +--- + +## 14. Gaps and contested expectations + +Each is a case above whose expected value is what the code does rather than what +the design says, or a behavior no case can assert because nothing implements it. +A fixture is still written for every one. The row records today's output so that +a later change to it is visible as a diff rather than a surprise. + +| # | Finding | Cases | +|------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|------------------------| +| G-1 | **Fixed.** `ClassRule` checked the transport on the NVMe side only, so a block run admitted an NVMe disk by its path and a draft could name both classes. A cluster is built out of one class, so the check is symmetric now: a block run refuses `NVMe` and `NVMeFabric` alike | DEV-05, DEV-07 | +| G-2 | Nothing reads `Rotational`, so a spinning disk is offered to a cluster on the same terms as an SSD | DEV-08 | +| G-3 | `vcpuCount` is the chosen node's whole physical core count, leaving nothing for the system on a large worker | SIZE-03 | +| G-4 | No worker rule reads memory: a 4 GiB machine is drafted as a storage node | SIZE-06, FAIL-07 | +| G-5 | Swap is reported and unread, though a host swapping is already oversubscribed | SIZE-14 | +| G-6 | The generator does not hold an address to the schema's PCI pattern, so a five-digit domain produces a document the API server refuses | PCI-07 | +| G-7 | No worker rule requires a management interface, so the generator writes a group it already knows cannot deploy. The `ClusterDeploymentConfig` controller's draft validation does report it as `NoManagementInterface`, one step and one object later than the run that produced it | NET-04, FAIL-05 | +| G-8 | A cordoned or tainted node reaches the draft, because the role labels are read and `Unschedulable` is not | ROLE-08 | +| G-9 | **Fixed.** An unreadable `driveSizeRange` dropped the rule and admitted every disk, and the comment claimed a validating caller reported it. There is one now: an admission webhook on `OperatorOps` refuses the run at the request, and both steps that read the filter refuse it again for a cluster whose webhook is not installed | FILT-06 | +| G-10 | A fleet under the vCPU floor is drafted at the floor, so approval produces nodes the workers cannot host | SIZE-02, FAIL-06 | +| G-11 | The cluster name is not checked against its own 63-character limit when derived | TMPL-04 | +| G-12 | The generator chooses one memory node per worker and proposes no `socketsToUse`, so the second socket of every dual-socket worker is left out of the cluster entirely | TMPL-06, NUMA-14 | +| G-13 | A group's signature is its addresses and its interface, so workers of differing capacity share a group and the group is named for whichever of them sorts first. Verified against the code | PCI-13 | +| G-14 | `minHugePagesSize` is the floor on the pool simplyblock allocates for itself, and the generator sets it to the free pages already on the chosen node, so another workload's reservation becomes this cluster's allocation floor | SIZE-09, SIZE-16 | +| G-15 | A worker with nothing pre-allocated leaves the field unset, on the stated grounds that SPDK consumes pages and does not reserve them. simplyblock allocates its own, so an empty pool is the ordinary starting state. The same premise heads `atlas-lib/inventory/hugepages.go` | SIZE-10, SIZE-17 | +| G-16 | Nothing checks that the host has memory for the allocation on top of the existing pool. `memory.availableBytes`, `memory.hugePagesBytes`, and the per-node `freeBytes` are all reported and all unread | SIZE-16, SIZE-17 | +| G-17 | No input says to take a stated number of the pre-allocated pages. `spec.discover` carries no huge-page field, and the only two knobs that exist are never written by a run: the cluster's `minHugePagesSize` and a group's `spdkSystemMemory` | SIZE-18 | +| G-18 | `hugePagesOf` drops `resv_hugepages`, which `inventory.HugePagePool` reads, so a draft cannot tell pages promised to a mapping from pages genuinely free | SIZE-19 | +| G-19 | A kernel with no hugetlbfs and a host whose pools are fully taken produce the same unset field and the same note, so a reviewer cannot tell a machine that cannot hold huge pages from one that has none left | SIZE-21 | +| G-20 | The cluster's field is documented as a floor and rendered into the per-node configuration as `MAX_HUGE_PAGES_SIZE`. Which of the two the node honors decides what a run should propose | SIZE-09, SIZE-16 | +| G-21 | **Fixed.** `servesManagement` refused every virtual interface, and a bond, a VLAN, a VXLAN, and a macvlan are all virtual, so a fleet whose management address sat on a bond or a VLAN yielded no interface on any worker. The ladder now refuses a kind rather than the absence of hardware | NET-13, NET-14, NET-16 | +| G-22 | **Fixed.** The report carried `virtual`, `loopback`, and `bridge` and nothing else, so a bond, a VLAN, a VXLAN, and a veth were indistinguishable in it. `inventory.Interface` now reads the device type the kernel publishes and the `lower_*` and `upper_*` links around it, and the report carries `kind`, `lower`, and `upper` at version 4 | NET-13, NET-14, NET-17 | +| G-23 | No input predefines the management interface. `DiscoverSpec` carries no interface field, so the ranking cannot be overridden for a fleet that knows which NIC it means | NET-22 | +| G-24 | No run writes `dataInterfaces`. Every drafted document leaves the data plane unnamed, and the fastest link is proposed for management instead | NET-23, NET-24 | +| G-25 | The grouper splits workers on their management interface while the expansion binds one interface per cluster, so a fleet whose machines name their NICs differently produces exactly the document `conflictingInterfaces` reports as unbuildable | NET-06 | +| G-26 | `interfacesOf` drops `carrier` and `duplex`, which `inventory.Interface` reads, so a link that is up with no partner is not distinguishable from one carrying traffic | NET-27 | +| G-27 | `NUMANodeBreakdown` ranks a memory node's NICs by the raw `speedMbps` and the raw `numaNode`, and a bond has neither. A worker whose data path is bonded or tagged therefore counts as having no NIC at all on every memory node, though the resolution that would answer both now exists | NET-41 | +| G-28 | `Management` carries the members, the resolved speed, and the memory node onto `Worker.Mgmt`, and nothing reads them. No note, field, or event says which slots a bonded management network lands in, so the evidence is gathered and not reported | NET-42 | + +## 15. Generation order + +The fixtures come first and whole. Every case in this document is a directory of +input objects, and the inputs are a statement of what the fleet was, which is +settled by the row rather than by anything the harness does with it. Writing the +complete set before any code reads it keeps the two apart: what is being tested +is reviewable as a tree of hosts, and the harness is then written against a set +that is already fixed rather than growing to fit the cases it happens to load. + +1. **Every case directory, all of them, inputs only.** `ops.yaml`, `nodes.yaml`, + `reports/*.yaml`, and the `case.md` naming the row. This is the bulk of the + work and none of it depends on a harness existing. +2. **Review the tree.** A family at a time, against its section here. A wrong + fixture is a test that passes for the wrong reason, and it is far cheaper to + catch in a directory of YAML than in a golden file. +3. **The harness**, walking the tree and driving each case through the run. It + is written once, against the whole set, and a case it cannot load is a gap in + the harness rather than a case to drop. +4. **The expected outputs.** For a row stating a value, written by hand from the + row. For the rest, recorded from a run and then read against the row before + being committed, because a golden file nobody read is a record of what the + code did rather than of what it should do. +5. **The `GO` rows**, in the `discovery` package, since they substitute a + `Planner` seam the controller never does and have no directory. + +The contested rows are written like any other and keep their gap number in +`case.md`. They record today's output deliberately, so that the day one of them +changes, the diff is the finding rather than a surprise. From 2c348f8eb22d8700002b0af64a32b80594fd2862 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Fri, 18 Sep 2026 16:57:57 +0200 Subject: [PATCH 075/206] fix(storagenode): a teardown reads the drain's phase, not the node's lock The lock being clear was taken to mean the drain had finished. It also means the drain has not started: the pass before the teardown raises the removal, and the node holds no lock until the operation picks it up. A teardown reading the empty field on that pass takes "not started" for "finished", drops the finalizer, and the object goes -- and with it the operation it owns. What is left is a backend node still running, with data on it, and nothing in Kubernetes tracking it. The signal is the removal's own phase now, read by name from the operation the node raised for itself. An operation that is not there is a drain that has not started rather than one that is over, for the same reason: a record deleted out of band is raised again on the next pass. The suite this arrived with covers the operation machinery the fix sits in, one file per subject rather than one per type, because what breaks is a path and not a function: the four single-call actions and their re-entry, the step spine every action advances on, how a drain classifies the volumes it finds, where its movable volumes go, the node's lock and the three paths that release it, node migration and the per-node configuration it clones, the gates a node passes before it is added, certificate rotation, what a provisioned node's status says and where each field comes from, what wakes a reconcile, and surviving a Kubernetes worker being drained underneath all of it. The four existing suites that named their node inline now take the name from the shared fixture, so a test that raises an operation and a test that reads one agree about which object they are talking about. Co-Authored-By: Claude Opus 5 (1M context) --- operator/docs/tests/test-plan-storagenode.md | 1127 ++++++++++------- .../internal/controllers/node/actions_test.go | 219 ++++ .../internal/controllers/node/advance_test.go | 310 +++++ .../controllers/node/classify_test.go | 255 ++++ .../internal/controllers/node/drain_test.go | 498 ++++++++ .../controllers/node/fixtures_test.go | 370 ++++++ .../controllers/node/hostmaintenance_test.go | 362 ++++++ .../internal/controllers/node/migrate_test.go | 438 +++++++ .../internal/controllers/node/opslock_test.go | 284 +++++ .../controllers/node/peertargets_test.go | 214 ++++ .../controllers/node/pernodeconfig_test.go | 149 +++ .../controllers/node/provisioning_test.go | 269 ++++ .../controllers/node/remove_fanout_test.go | 2 +- .../controllers/node/remove_gone_test.go | 2 +- .../controllers/node/rotation_test.go | 115 ++ .../controllers/node/selfbudget_test.go | 4 +- .../node/storagenode_controller.go | 39 +- .../controllers/node/syncstatus_test.go | 481 +++++++ .../controllers/node/unregister_test.go | 2 +- .../internal/controllers/node/watches_test.go | 128 ++ 20 files changed, 4806 insertions(+), 462 deletions(-) create mode 100644 operator/internal/controllers/node/actions_test.go create mode 100644 operator/internal/controllers/node/advance_test.go create mode 100644 operator/internal/controllers/node/classify_test.go create mode 100644 operator/internal/controllers/node/drain_test.go create mode 100644 operator/internal/controllers/node/fixtures_test.go create mode 100644 operator/internal/controllers/node/hostmaintenance_test.go create mode 100644 operator/internal/controllers/node/migrate_test.go create mode 100644 operator/internal/controllers/node/opslock_test.go create mode 100644 operator/internal/controllers/node/peertargets_test.go create mode 100644 operator/internal/controllers/node/pernodeconfig_test.go create mode 100644 operator/internal/controllers/node/provisioning_test.go create mode 100644 operator/internal/controllers/node/rotation_test.go create mode 100644 operator/internal/controllers/node/syncstatus_test.go create mode 100644 operator/internal/controllers/node/watches_test.go diff --git a/operator/docs/tests/test-plan-storagenode.md b/operator/docs/tests/test-plan-storagenode.md index dcfb8d875..f89cca606 100644 --- a/operator/docs/tests/test-plan-storagenode.md +++ b/operator/docs/tests/test-plan-storagenode.md @@ -14,444 +14,634 @@ answer, never the control plane's own behavior. Scenario IDs are permanent and are never reused or renumbered. A `—` in the `Test` column means nothing implements the scenario yet, and every such row -reappears in §6 with its reason. - -Rows that are runnable against the shipped API name the shipped spelling, and -rows in a planned block name the target. The parent reference is -`spec.storageNodeSetRef` today and `spec.clusterRef` after the reparent in design -§3.1, the per-node block is `spec.overrides` today and `spec.config` after, the -operation's target is `spec.storageNodeRef` today and `spec.nodeRef` after, and -its position is `status.subPhase` today and `status.step` after, and the slot -ordinal is `spec.socketIndex` today and `spec.slot` after. Action values are -lowercase today (`remove`, `migrate`) and PascalCase after the rename in design -§6.3 (`Remove`, `Migrate`), so a row naming an action names the spelling its -harness would use. - -| Class | Prefix | Harness | -|-------------|--------|------------------------------------------------------------------------| -| Unit | `U-` | No cluster: pure functions, a fake `client.Client`, and a mock backend | -| Integration | `I-` | Full reconcile loop against `envtest` and a mock backend | -| E2E | `E-` | Live simplyblock cluster, real data path | -| Manual | `M-` | Needs failure injection or orchestration not automated yet | +reappears in §6 with its reason. A struck-through ID is a scenario that does not +describe the shipped system, kept because review history cites it, and its +`Test` column names the row that replaced it. + +Both kinds ship at `v1alpha2` and every row names that spelling: `spec.clusterRef` +for the parent, `spec.config` for the per-node block, `spec.slot` for the socket +ordinal, `spec.nodeRef` for the operation's target, `status.step` for its +position, and PascalCase action values (`Remove`, `Migrate`). The reconcilers live +in `operator/internal/controllers/node/`, and the retired `StorageNodeSet` that +several rows were first written against is gone: the workload objects belong to +the `StorageCluster`, and the node's own object name is derived by the +`ClusterDeploymentConfig` expansion that creates it. + +| Class | Prefix | Harness | +|-------------|--------|----------------------------------------------------------------------------------| +| Unit | `U-` | No cluster: pure functions, a fake `client.Client`, and a scripted control plane | +| Integration | `I-` | Full reconcile loop against `envtest` and a mock backend | +| E2E | `E-` | Live simplyblock cluster, real data path | +| Manual | `M-` | Needs failure injection or orchestration not automated yet | --- ## 1. Unit Tests Pure functions and single reconcile calls against a fake client, with the control -plane replaced by a mock HTTP server. No Kubernetes API server is involved. +plane replaced by the scripted surface in +`operator/internal/controllers/node/fixtures_test.go`, which answers what a case +states and records what was asked of it. No Kubernetes API server is involved. ### Entity: Object Naming (design §3.1) -File: `operator/internal/controllers/node/storagenodeset_storagenode_unit_test.go` - -| # | Scenario | Type | Test | -|------|---------------------------------------------------------------------|----------|-----------------------------------------------------| -| U-01 | A generated name is non-empty, lowercase, and at most 63 characters | Positive | `TestStorageNodeCRName_SimpleCase` | -| U-02 | A parent name long enough to overflow is truncated to a valid label | Boundary | `TestStorageNodeCRName_TruncatesLongNames` | -| U-03 | Two calls with the same parent name produce different names | Positive | `TestStorageNodeCRName_IsRandomPerCall` | -| U-04 | A generated name contains only DNS-label characters | Positive | `TestStorageNodeCRName_IsDNSLabelSafe` | -| U-05 | Uppercase and underscores in a worker hostname are replaced | Negative | `TestSanitiseDNSLabel_ReplacesInvalidChars` | -| U-06 | Leading and trailing hyphens are stripped from a sanitized label | Boundary | `TestSanitiseDNSLabel_StripsLeadingTrailingHyphens` | -| U-07 | A name collision on create is retried with a fresh identifier | Negative | — | -| U-08 | The generated name encodes neither the worker nor the socket | Negative | — | +File: `operator/internal/controllers/deployment/nodename_test.go` + +The name is derived from the cluster, the worker, and the slot through atlas-lib's +formula, so it is deterministic and the same three inputs always produce it again. +The rows that described a random identifier are struck: an expansion finds the +node it already created by the worker and slot its spec records, so the name is a +function of those three inputs and of nothing else. + +| # | Scenario | Type | Test | +|----------|----------------------------------------------------------------------|----------|-------------------------------------------| +| U-01 | A generated name is non-empty, lowercase, and at most 63 characters | Positive | `TestANodeNameFitsTheLabelItIsCopiedInto` | +| U-02 | A parent name long enough to overflow is truncated to a valid label | Boundary | `TestANodeNameFitsTheLabelItIsCopiedInto` | +| ~~U-03~~ | ~~Two calls with the same parent name produce different names~~ | Positive | Superseded by U-265 | +| U-04 | A generated name contains only DNS-label characters | Positive | `TestANodeNameFitsTheLabelItIsCopiedInto` | +| ~~U-05~~ | ~~Uppercase and underscores in a worker hostname are replaced~~ | Negative | Superseded by U-265 | +| ~~U-06~~ | ~~Leading and trailing hyphens are stripped from a sanitized label~~ | Boundary | Superseded by U-265 | +| ~~U-07~~ | ~~A name collision on create is retried with a fresh identifier~~ | Negative | Superseded by U-266 | +| ~~U-08~~ | ~~The generated name encodes neither the worker nor the socket~~ | Negative | Superseded by U-265 | +| U-265 | The same cluster, worker, and slot derive the same name every time | Positive | — | +| U-266 | Two workers whose names share a long prefix do not collide | Boundary | — | ### Entity: Provisioning Gates (design §4.2) -File: `operator/internal/controllers/node/storagenode_controller_unit_test.go` - -| # | Scenario | Type | Test | -|------|--------------------------------------------------------------------------------------|----------|--------------------------------------------------------------| -| U-09 | `enableFailureDomains` set and no fault group declared: provisioning is held | Negative | `TestCheckFailureDomain_BlocksWhenEnabledAndNotSet` | -| U-10 | `enableFailureDomains` set and a fault group present: provisioning proceeds | Positive | `TestCheckFailureDomain_AllowsWhenFailureDomainSet` | -| U-11 | `enableFailureDomains` unset: the fault group is not required | Negative | `TestCheckFailureDomain_SkipsWhenFeatureDisabled` | -| U-12 | A `failureDomain` label of `0`: a label like any other, not read as unset | Boundary | — | -| U-13 | Held provisioning emits `FailureDomainMissing` and issues no `POST` | Negative | — | -| U-14 | The worker's storage-node API answers: the host check passes | Positive | `TestCheckNodeInfoReachable` | -| U-15 | The worker's storage-node API is unreachable: held, no `POST` | Negative | `TestStorageNodeSetReconcileUnreachableNodeInfoRequeues` | -| U-16 | TLS is enabled and the CA is missing: the host check fails informatively | Negative | `TestCheckNodeInfoReachableTLSMissingCA` | -| U-17 | The host check retries until the endpoint answers | Positive | `TestWaitForNodeInfoReachable` | -| U-18 | No in-flight sibling: the in-flight count is zero | Boundary | `TestCountInFlightNodes_ZeroWhenNonePosted` | -| U-19 | Siblings past `Posting` without a UUID are counted as in flight | Positive | `TestCountInFlightNodes_CountsSiblingsWithPostedAtAndNoUUID` | -| U-20 | The node counts every sibling but itself | Boundary | `TestCountInFlightNodes_ExcludesSelf` | -| U-21 | Two nodes on one worker count as one in-flight worker, not two | Boundary | `TestCountInFlightNodes_DeduplicatesByWorker` | -| U-22 | A worker already in flight does not block another worker under the limit | Positive | `TestParallelNodeAddContinuesPastPendingWorker` | -| U-23 | `maxParallelNodeAdds` reached: the node holds at `AwaitingSlot` and issues no `POST` | Negative | — | -| U-24 | Workers hosting a FoundationDB pod are identified | Positive | `TestFDBWorkerSet` | -| U-25 | A FoundationDB worker holds while another FoundationDB worker is in flight | Negative | — | -| U-26 | A FoundationDB worker holds even when `maxParallelNodeAdds` allows more | Boundary | — | -| U-27 | A non-FoundationDB worker is not held by a FoundationDB worker in flight | Negative | — | +Files: `operator/internal/controllers/node/provisioning_test.go`, +`slot_race_test.go`, `resolve_retry_test.go` + +| # | Scenario | Type | Test | +|-------|--------------------------------------------------------------------------------------|----------|----------------------------------------------------| +| U-09 | `enableFailureDomains` set and no fault group declared: provisioning is held | Negative | `TestANodeWithNoFaultGroupIsHeldRatherThanRefused` | +| U-10 | `enableFailureDomains` set and a fault group present: provisioning proceeds | Positive | `TestANodeThatDeclaresItsFaultGroupPasses` | +| U-11 | `enableFailureDomains` unset: the fault group is not required | Negative | `TestANodeThatDeclaresItsFaultGroupPasses` | +| U-12 | A `failureDomain` label of `0`: a label like any other, not read as unset | Boundary | — | +| U-13 | Held provisioning emits `FailureDomainMissing` and issues no `POST` | Negative | `TestANodeWithNoFaultGroupIsHeldRatherThanRefused` | +| U-14 | The worker's storage-node API answers: the host check passes | Positive | — | +| U-15 | The worker's storage-node API is unreachable: held, no `POST` | Negative | — | +| U-16 | TLS is enabled and the CA is missing: the host check fails informatively | Negative | — | +| U-17 | The host check retries until the endpoint answers | Positive | — | +| U-18 | Nothing in flight and one slot free: exactly one of three waiting nodes takes it | Boundary | `TestOnlyOneNodeTakesAFreeSlot` | +| U-19 | Siblings past `Posting` without a UUID hold the slot | Positive | `TestANodeInFlightFillsTheCap` | +| U-20 | The node counts every sibling but itself | Boundary | — | +| U-21 | Two nodes on one worker count as one in-flight worker, not two | Boundary | — | +| U-22 | A worker already in flight does not block another worker under the limit | Positive | `TestACapOfTwoAdmitsTwo` | +| U-23 | `maxParallelNodeAdds` reached: the node holds at `AwaitingSlot` and issues no `POST` | Negative | `TestANodeInFlightFillsTheCap` | +| U-24 | Workers hosting a FoundationDB pod are identified | Positive | — | +| U-25 | A FoundationDB worker holds while another FoundationDB worker is in flight | Negative | — | +| U-26 | A FoundationDB worker holds even when `maxParallelNodeAdds` allows more | Boundary | — | +| U-27 | A non-FoundationDB worker is not held by a FoundationDB worker in flight | Negative | — | +| U-267 | The same waiting node wins the slot on every pass, so nobody overtakes it | Positive | `TestTheChoiceIsStable` | ### Entity: The Provisioning Claim (design §4.2) -File: `operator/internal/controllers/node/storagenode_controller_unit_test.go` - -| # | Scenario | Type | Test | -|------|---------------------------------------------------------------------------------------|----------|------| -| U-28 | The transition into `Posting` is persisted before the `POST` is issued | Positive | — | -| U-29 | A second reconciler at the same `resourceVersion`: 409, backs off, issues no `POST` | Negative | — | -| U-30 | A sibling socket already at `Posting`: this object enters `Resolving` without posting | Negative | — | -| U-31 | A sibling socket at `Resolving`: this object enters `Resolving` without posting | Negative | — | -| U-32 | The `POST` returns 5xx: the step stays `Posting` and is retried | Negative | — | -| U-33 | The `POST` returns 4xx: the step stays `Posting`, and the body is in the event | Negative | — | -| U-34 | The `POST` times out: the step is not advanced and no second `POST` is issued | Negative | — | -| U-35 | The `POST` succeeds: the step advances to `Resolving` | Positive | — | +File: `operator/internal/controllers/node/provisioning_test.go` + +The claim is the transition into `Posting`, written as an optimistic-lock patch so +that two objects for one worker cannot both conclude the add is theirs to make. + +| # | Scenario | Type | Test | +|-------|---------------------------------------------------------------------------------------|----------|-----------------------------------------------------------| +| U-28 | The transition into `Posting` is persisted before the `POST` is issued | Positive | — | +| U-29 | A second reconciler at the same `resourceVersion`: 409, backs off, issues no `POST` | Negative | — | +| U-30 | A sibling socket already at `Posting`: this object enters `Resolving` without posting | Negative | — | +| U-31 | A sibling socket at `Resolving`: this object enters `Resolving` without posting | Negative | — | +| U-32 | The `POST` returns 5xx: the step stays `Posting` and is retried | Negative | — | +| U-33 | The `POST` returns 4xx: the step stays `Posting`, and the body is in the event | Negative | — | +| U-34 | The `POST` times out: the step is not advanced and no second `POST` is issued | Negative | — | +| U-35 | The `POST` succeeds: the step advances to `Resolving` | Positive | `TestPostingIssuesTheAddOnce` | +| U-268 | The add carries the node's own images, memory, and journal settings | Positive | `TestTheAddCarriesWhatTheNodeSaysAboutItself` | +| U-269 | A node stating no journal settings is added with the documented defaults | Boundary | `TestAnUnstatedJournalIsTheDefaultRatherThanNothing` | +| U-270 | A fault group that is a number is sent as one, and a named group is not sent | Boundary | `TestOnlyAFaultGroupThatIsANumberIsSentToTheControlPlane` | +| U-271 | A step no provisioning path declares: refused rather than stalled | Negative | `TestAStepNoProvisioningPathDeclaresIsRefused` | ### Entity: UUID Resolution and Adoption (design §4.2, §4.3) -File: `operator/internal/controllers/node/storagenode_controller_unit_test.go` - -| # | Scenario | Type | Test | -|------|------------------------------------------------------------------------------------|----------|------------------------------------------| -| U-36 | The worker's internal IP is resolved from the Kubernetes `Node` | Positive | `TestGetNodeInternalIP` | -| U-37 | The Kubernetes `Node` carries no internal address: held, not failed | Negative | `TestGetNodeInternalIPNoAddress` | -| U-38 | One backend node at the worker's IP: matched to `slot` 0 | Positive | — | -| U-39 | Two backend nodes at one IP: sorted by RPC port and matched by `slot` | Positive | — | -| U-40 | `slot` beyond the number of backend nodes at the IP: held, not matched | Boundary | — | -| U-41 | No backend node at the worker's IP: `Resolving` holds | Negative | `TestPollNodeOnlinePaths` | -| U-42 | The control plane errors during resolution: held, and the step is not advanced | Negative | `TestPollNodeOnlineErrorAndTimeoutPaths` | -| U-43 | The `Resolving` deadline expires: the node is marked failed rather than polling on | Boundary | — | -| U-44 | An upgrade Secret is present: the node adopts without a `POST` | Positive | — | -| U-45 | An upgrade Secret is present but empty: falls through to the normal path | Negative | — | -| U-46 | A backend node already exists at the worker's IP: adopted rather than added | Positive | — | -| U-47 | Adoption records the backend UUID and leaves `status.phase` at `Online` | Positive | — | +Files: `operator/internal/controllers/node/provisioning_test.go`, +`resolve_retry_test.go` + +| # | Scenario | Type | Test | +|----------|------------------------------------------------------------------------------------|----------|-----------------------------------------------------| +| U-36 | The worker's internal IP is resolved from the Kubernetes `Node` | Positive | — | +| U-37 | The Kubernetes `Node` carries no internal address: held, not failed | Negative | — | +| U-38 | One backend node at the worker's IP: matched to `slot` 0 | Positive | `TestANodeAlreadyAtTheWorkersAddressIsAdopted` | +| U-39 | Two backend nodes at one IP: sorted by RPC port and matched by `slot` | Positive | — | +| U-40 | `slot` beyond the number of backend nodes at the IP: held, not matched | Boundary | — | +| U-41 | No backend node at the worker's IP: `Resolving` holds | Negative | `TestResolvingWaitsWhileTheAddIsStillRunning` | +| U-42 | The control plane errors during resolution: held, and the step is not advanced | Negative | — | +| U-43 | The `Resolving` deadline expires: the node is marked failed rather than polling on | Boundary | — | +| U-44 | An upgrade Secret is present: the node adopts without a `POST` | Positive | `TestAnUpgradeAdoptionDivertsBeforeTheHostIsProbed` | +| ~~U-45~~ | ~~An upgrade Secret is present but empty: falls through to the normal path~~ | Negative | Superseded by U-272 | +| U-46 | A backend node already exists at the worker's IP: adopted rather than added | Positive | `TestANodeAlreadyAtTheWorkersAddressIsAdopted` | +| U-47 | Adoption records the backend UUID and leaves `status.phase` at `Online` | Positive | — | +| U-272 | The upgrade Secret's presence is the whole signal, whatever it holds | Boundary | `TestAnUpgradeAdoptionDivertsBeforeTheHostIsProbed` | +| U-273 | The `node_add` left the task window without producing a node: the add is reissued | Negative | `TestResolvingAsksAgainWhenTheAddIsOver` | +| U-274 | No `node_add` in the window at all: the add is reissued rather than waited out | Boundary | `TestResolvingAsksAgainWhenNoAddIsInTheWindow` | +| U-275 | An unrelated task in the window does not hold `Resolving` open | Negative | `TestAnUnrelatedRunningTaskDoesNotHoldResolving` | ### Entity: Steady-State Sync (design §4.4) -File: `operator/internal/controllers/node/storagenode_controller_unit_test.go` - -| # | Scenario | Type | Test | -|-------|-------------------------------------------------------------------------------|----------|------| -| U-48 | A streamed node object writes `status.status`, `health`, and the resources | Positive | — | -| U-49 | Nothing changed since the last pass: the reconcile issues no status patch | Negative | — | -| U-50 | The control-plane assigned fault group differs from the requested one | Positive | — | -| U-51 | A malformed streamed object: an error rather than a nil dereference | Negative | — | -| U-52 | `status.observedGeneration` matches `metadata.generation` after a sync | Positive | — | -| U-236 | A node reporting 3 of 4 devices online: `devices.online` 3, `devices.total` 4 | Positive | — | -| U-237 | A node whose devices have not been reported: `devices` absent, not `{0, 0}` | Boundary | — | -| U-238 | A node with zero devices online: `devices.online` is 0 and present | Boundary | — | -| U-249 | The first capacity sample: `resources.capacity` written with its `sampledAt` | Positive | — | -| U-250 | A sample under the one-percent threshold: the reconcile issues no patch | Negative | — | -| U-251 | A sample over it: the new used size and the new sample time are written | Positive | — | -| U-252 | The total changed because a device joined: written whatever the used delta is | Boundary | — | -| U-253 | The capacity source is unreachable: the node is published without `capacity` | Negative | — | -| U-254 | A node the exporter has never measured: `capacity` absent rather than zeros | Boundary | — | +File: `operator/internal/controllers/node/syncstatus_test.go` + +| # | Scenario | Type | Test | +|-------|--------------------------------------------------------------------------------|----------|---------------------------------------------------------------| +| U-48 | A streamed node object writes `status.status`, `health`, and the resources | Positive | `TestAPushedNodeIsReadFromTheStreamRatherThanAskedFor` | +| U-49 | Nothing changed since the last pass: the reconcile issues no status patch | Negative | `TestAPassThatFoundNothingNewWritesNothing` | +| U-50 | The control-plane assigned fault group differs from the requested one | Positive | — | +| U-51 | A malformed streamed object: an error rather than a nil dereference | Negative | — | +| U-52 | `status.observedGeneration` matches `metadata.generation` after a sync | Positive | — | +| U-236 | A node reporting 3 of 4 devices online: `devices.online` 3, `devices.total` 4 | Positive | — | +| U-237 | A node whose devices have not been reported: `devices` absent, not `{0, 0}` | Boundary | — | +| U-238 | A node with zero devices online: `devices.online` is 0 and present | Boundary | — | +| U-249 | The first capacity sample: `resources.capacity` written with its `sampledAt` | Positive | `TestAFirstCapacityReadingIsAlwaysWritten` | +| U-250 | A sample under the one-percent threshold: the reconcile issues no patch | Negative | `TestASampleIsWrittenOnlyWhenItSaysSomethingNew` | +| U-251 | A sample over it: the new used size and the new sample time are written | Positive | `TestASampleIsWrittenOnlyWhenItSaysSomethingNew` | +| U-252 | The total changed because a device joined: written whatever the used delta is | Boundary | `TestASampleIsWrittenOnlyWhenItSaysSomethingNew` | +| U-253 | The capacity source is unreachable: the node is published without `capacity` | Negative | `TestAFailingCapacitySourceStillLeavesTheNodePublished` | +| U-254 | A node the exporter has never measured: `capacity` absent rather than zeros | Boundary | `TestANodeNobodySampledCarriesNoCapacity` | +| U-276 | A node the stream has delivered costs no control-plane request to read back | Positive | `TestAPushedNodeIsReadFromTheStreamRatherThanAskedFor` | +| U-277 | A node the stream has not delivered is read from the control plane instead | Boundary | `TestANodeTheStreamHasNotDeliveredIsAskedFor` | +| U-278 | Each control-plane status maps to the phase this operator publishes | Positive | `TestThePhaseIsThisOperatorsReadingOfWhatTheControlPlaneSays` | +| U-279 | A status this operator has never heard of: `Failed` rather than a guess | Negative | `TestThePhaseIsThisOperatorsReadingOfWhatTheControlPlaneSays` | +| U-280 | A stored UUID the control plane has forgotten: the object goes back to Pending | Boundary | `TestANodeTheControlPlaneHasForgottenGoesBackToProvisioning` | +| U-281 | A reading that did move is written rather than held back by the threshold | Positive | `TestAPassThatFoundSomethingNewWritesIt` | ### Entity: Deletion (design §4.5) -File: `operator/internal/controllers/node/storagenode_controller_unit_test.go` - -| # | Scenario | Type | Test | -|------|---------------------------------------------------------------------------------|----------|-----------------------------------------------------------| -| U-53 | A node that never got a UUID: the finalizer is removed with no operation raised | Boundary | `TestHandleDeletion_RemovesFinalizerWhenNeverProvisioned` | -| U-54 | An online node deleted: a `Remove` operation is raised and owned by the node | Positive | `TestEnsureRemoveOps_CreatesOpsWhenMissing` | -| U-55 | A second reconcile while the drain runs: no second operation is created | Negative | `TestEnsureRemoveOps_IdempotentWhenAlreadyExists` | -| U-56 | `status.activeOpsRef` still set: the finalizer is held | Negative | — | -| U-57 | The raised operation failed: the finalizer is held rather than force-removed | Negative | — | -| U-58 | A suspended node deleted: a `Remove` operation is still raised | Positive | — | -| U-59 | An offline node deleted: no operation is raised and the finalizer is removed | Boundary | — | - -### Entity: The workerNode Webhook (design §3.1) +File: `operator/internal/controllers/node/syncstatus_test.go` + +| # | Scenario | Type | Test | +|----------|----------------------------------------------------------------------------------|------------|-----------------------------------------------------| +| U-53 | A node that never got a UUID: the finalizer is removed with no operation raised | Boundary | `TestANodeThatWasNeverProvisionedIsDeletedOutright` | +| U-54 | An online node deleted: a `Remove` operation is raised and owned by the node | Positive | `TestANodeWithDataOnItIsDrainedBeforeItGoes` | +| U-55 | A second reconcile while the drain runs: no second operation is created | Negative | — | +| ~~U-56~~ | ~~`status.activeOpsRef` still set: the finalizer is held~~ | Negative | Superseded by U-282 | +| U-57 | The raised operation failed: the finalizer is held rather than force-removed | Negative | — | +| U-58 | A suspended node deleted: a `Remove` operation is still raised | Positive | — | +| ~~U-59~~ | ~~An offline node deleted: no operation is raised and the finalizer is removed~~ | Boundary | Superseded by U-282 | +| U-282 | A drain that has not reached a terminal phase holds the finalizer | Regression | `TestANodeWithDataOnItIsDrainedBeforeItGoes` | +| U-283 | The drain finished: the finalizer is released and the object goes | Positive | `TestANodeWithDataOnItIsDrainedBeforeItGoes` | +| U-284 | A cordoned worker raises a `HostMaintenance` operation the node owns | Positive | `TestACordonedWorkerRaisesItsMaintenanceWindow` | +| U-285 | A second pass finds the window it raised rather than raising another | Negative | `TestACordonedWorkerRaisesItsMaintenanceWindow` | + +`U-282` is the regression row for +`2026-09-17-node-teardown-reads-an-unheld-lock-as-a-finished-drain`: the teardown +held the finalizer while `status.activeOpsRef` was set, and a drain raised one +line earlier has not taken that lock yet, so the first pass read "not started" as +"finished" and deleted the object while its backend node was still running. +`U-56` and `U-59` are struck because the lock is no longer the signal and a node +with a UUID always gets its drain, whatever the control plane reports about it. + +### Entity: The Admission Guard (design §3.1, §3.2, §3.4) File: `operator/internal/webhook/storagenode_validator_test.go` -| # | Scenario | Type | Test | -|-------|-------------------------------------------------------------------------------------|----------|----------------------------| -| U-60 | A user changing `spec.workerNode`: denied with the migration hint | Negative | `TestStorageNodeValidator` | -| U-61 | The operator's service account changing `spec.workerNode`: allowed | Positive | `TestStorageNodeValidator` | -| U-62 | An update that does not touch `spec.workerNode`: allowed without inspection | Negative | `TestStorageNodeValidator` | -| U-63 | A create rather than an update: allowed, since there is no old value | Boundary | `TestStorageNodeValidator` | -| U-64 | A service account in another namespace named like the operator's: denied | Negative | — | -| U-239 | A user changing `spec.config.pcieAllowList`: denied | Negative | — | -| U-240 | The operator merging `newSsdPcie` into `spec.config.pcieAllowList`: allowed | Positive | — | -| U-241 | A user changing `spec.config.sizing.vcpuCount`: denied | Negative | — | -| U-242 | The operator re-sizing `spec.config.sizing`: allowed | Positive | — | -| U-243 | An update touching none of the guarded fields: admitted without inspection | Negative | — | -| U-255 | A `config.deviceNames` entry that is a path on an `NVMe` cluster: denied | Negative | — | -| U-256 | A `config.deviceNames` of PCI addresses on an `NVMe` cluster: admitted | Positive | — | -| U-257 | A `config.deviceNames` entry that is an address on a `LogicalBlock` cluster: denied | Negative | — | -| U-258 | A list holding an address and a path: denied whichever class the cluster is | Negative | — | -| U-259 | A bare device name: read as a path and classed as block, not as unknown | Boundary | — | -| U-260 | `config.pcieDenyList` set on a `LogicalBlock` cluster: denied | Negative | — | -| U-261 | `config.pcieDenyList` set on an `NVMe` cluster: admitted | Positive | — | +| # | Scenario | Type | Test | +|-------|-------------------------------------------------------------------------------------|----------|-----------------------------------------------------| +| U-60 | A user changing `spec.workerNode`: denied with the migration hint | Negative | `TestTheOperatorOnlyFieldsAreRefusedToEveryoneElse` | +| U-61 | The operator's service account changing `spec.workerNode`: allowed | Positive | `TestTheOperatorOnlyFieldsAreRefusedToEveryoneElse` | +| U-62 | An update that does not touch `spec.workerNode`: allowed without inspection | Negative | `TestAnUpdateTouchingNoGuardedFieldIsAdmitted` | +| U-63 | A create rather than an update: allowed, since there is no old value | Boundary | `TestTheOperatorMayCreateANodeSizedAgainstTheFleet` | +| U-64 | A service account in another namespace named like the operator's: denied | Negative | — | +| U-239 | A user changing `spec.config.pcieAllowList`: denied | Negative | `TestTheOperatorOnlyFieldsAreRefusedToEveryoneElse` | +| U-240 | The operator merging `newSsdPcie` into `spec.config.pcieAllowList`: allowed | Positive | `TestTheOperatorOnlyFieldsAreRefusedToEveryoneElse` | +| U-241 | A user changing `spec.config.sizing.vcpuCount`: denied | Negative | `TestTheOperatorOnlyFieldsAreRefusedToEveryoneElse` | +| U-242 | The operator re-sizing `spec.config.sizing`: allowed | Positive | `TestTheOperatorOnlyFieldsAreRefusedToEveryoneElse` | +| U-243 | An update touching none of the guarded fields: admitted without inspection | Negative | `TestAnUpdateTouchingNoGuardedFieldIsAdmitted` | +| U-255 | A `config.deviceNames` entry that is a path on an `NVMe` cluster: denied | Negative | `TestDeviceNamesMustBeOfTheClusterSClass` | +| U-256 | A `config.deviceNames` of PCI addresses on an `NVMe` cluster: admitted | Positive | `TestDeviceNamesOfTheClusterSClassAreAdmitted` | +| U-257 | A `config.deviceNames` entry that is an address on a `LogicalBlock` cluster: denied | Negative | `TestDeviceNamesMustBeOfTheClusterSClass` | +| U-258 | A list holding an address and a path: denied whichever class the cluster is | Negative | `TestDeviceNamesMustBeOfTheClusterSClass` | +| U-259 | A bare device name: read as a path and classed as block, not as unknown | Boundary | `TestDeviceNamesOfTheClusterSClassAreAdmitted` | +| U-260 | `config.pcieDenyList` set on a `LogicalBlock` cluster: denied | Negative | `TestThePCIFiltersAreRefusedOnALogicalBlockCluster` | +| U-261 | `config.pcieDenyList` set on an `NVMe` cluster: admitted | Positive | `TestThePCIFiltersAreRefusedOnALogicalBlockCluster` | +| U-286 | A cluster stating no device class is read as `NVMe` | Boundary | `TestAClusterWithNoStatedClassIsNVMe` | +| U-287 | A node naming no `StorageCluster`: refused at admission | Negative | `TestANodeNamingNoClusterIsRefused` | +| U-305 | A user's node whose sizing differs from the fleet's: refused at create | Negative | `TestAUserSNodeMustAgreeWithTheFleetSSizing` | +| U-306 | The operator's node sized against the fleet mid-roll: admitted | Positive | `TestTheOperatorMayCreateANodeSizedAgainstTheFleet` | + +### Operation: The Deletion Guard (design §7.4) + +File: `operator/internal/webhook/storagenodeops_validator_test.go` + +| # | Scenario | Type | Test | +|-------|---------------------------------------------------------------------------------|----------|--------------------------------------------------------------| +| U-288 | A delete at a step the graph declares no abort edge from: refused | Negative | `TestStorageNodeOpsDeleteIsRefusedWhereNoAbortEdgeExists` | +| U-289 | The refusal names what the record still owes, rather than only refusing | Positive | `TestTheRefusalAtPromotingSaysWhatTheRecordStillOwes` | +| U-290 | A delete at a step with an abort edge: admitted | Positive | `TestStorageNodeOpsDeleteIsAllowedWhereTheAbortEdgeExists` | +| U-291 | A delete of a terminal operation: admitted, since it is a record | Positive | `TestStorageNodeOpsDeleteIsAllowedOnceTerminal` | +| U-292 | A delete before the operation started: admitted | Boundary | `TestStorageNodeOpsDeleteIsAllowedBeforeTheOperationStarted` | +| U-293 | A delete with no object to read: admitted rather than failing closed on nothing | Boundary | `TestStorageNodeOpsDeleteIsAllowedWithNoObjectToRead` | +| U-294 | A create is not this guard's business | Negative | `TestStorageNodeOpsCreateIsNotThisGuardsBusiness` | +| U-295 | The refusal table and the graph's own unabortable steps agree | Positive | `TestTheNodeRefusalTableAndTheGraphAgree` | +| U-296 | Every undeletable step states what it is doing, so a refusal is actionable | Positive | `TestEveryUndeletableNodeStepStatesWhatItIsDoing` | + +### Entity: Conversion to and from `v1alpha1` (design §15.1) + +File: `operator/api/v1alpha1/storagenode_conversion_test.go` + +| # | Scenario | Type | Test | +|-------|---------------------------------------------------------------------------------|------------|------------------------------------------------------------| +| U-297 | `spec.clusterRef` is read from the controller owner the stored object carries | Positive | `TestStorageNodeReadsItsClusterFromTheControllerOwner` | +| U-298 | An object with no controller owner converts with no cluster rather than failing | Boundary | `TestStorageNodeWithNoControllerConvertsWithNoCluster` | +| U-299 | The device summary is read in the order it was written, not the documented one | Regression | `TestStorageNodeDeviceSummaryIsReadInTheOrderItWasWritten` | +| U-300 | An unparsable device summary produces no device block rather than zeros | Boundary | `TestStorageNodeUnparsableDeviceSummaryHasNoBlock` | +| U-301 | A failure-domain label survives the trip down to the retired version | Positive | `TestStorageNodeFailureDomainLabelSurvivesTheTripDown` | +| U-302 | A failure-domain index becomes its digits on the way up | Positive | `TestStorageNodeFailureDomainIndexBecomesItsDigits` | +| U-303 | The four per-node fields that reach nothing survive the round trip | Boundary | `TestStorageNodeDeadPerNodeFieldsSurviveTheRoundTrip` | +| U-304 | The sizing block survives the trip down | Positive | `TestStorageNodeSizingSurvivesTheTripDown` | ### Workload: DaemonSet, Services, and RBAC (design §5.1) -Files: `operator/internal/controllers/node/storagenodeset_controller_unit_test.go`, -`operator/internal/utils/storage_nodeset_ds_test.go` - -| # | Scenario | Type | Test | -|-------|-----------------------------------------------------------------------------------------------------|------------|--------------------------------------------------------------------------| -| U-65 | No DaemonSet present: one is created | Positive | `TestStorageNodeSetDaemonSetReconcileCreatesWhenMissing` | -| U-66 | A DaemonSet present: it is updated in place rather than recreated | Positive | `TestStorageNodeSetDaemonSetReconcileUpdatesExisting` | -| U-67 | TLS disabled: the pod template carries no serving-certificate mount | Negative | `TestStorageNodeSetDaemonSetReconcileTLSDisabled` | -| U-68 | TLS enabled: the pod template mounts the serving certificate | Positive | `TestStorageNodeSetDaemonSetReconcileTLSEnabled` | -| U-69 | The cert-manager provider: the Certificate is created alongside | Positive | `TestStorageNodeSetDaemonSetReconcileTLSCertManagerProvider` | -| U-70 | User-supplied container resources override the defaults | Positive | `TestBuildStorageNodeSetDaemonSetUserResourcesOverrideDefaults` | -| U-71 | The ServiceAccount carries an owner reference to its parent | Positive | `TestStorageNodeSetReconcileServiceAccountHasOwnerReference` | -| U-72 | ClusterRoleBinding names include the namespace, so two namespaces do not collide | Positive | `TestBuildStorageNodeSetClusterRoleBindingNameIncludesNamespace` | -| U-73 | Namespace-specific ClusterRoleBindings are created | Positive | `TestStorageNodeSetReconcileCreatesNamespaceSpecificClusterRoleBindings` | -| U-74 | The SPDK proxy Service is created | Positive | `TestReconcileSpdkProxyService` | -| U-75 | Every workload object carries an owner reference to the `StorageCluster` | Positive | — | -| U-76 | Deleting the `StorageCluster` cascades to every workload object | Positive | — | -| U-264 | No `clusterImage`: the image is taken from the singleton `ControlPlane`, read at the stored version | Regression | `TestStorageNodeSetDaemonSetImageFallsBackToControlPlane` | +Files: `operator/internal/controllers/node/rotation_test.go`, +`enrollment_test.go`, `operator/internal/utils/storage_node_workload_test.go` + +The objects belong to the `StorageCluster` now, and the reconcile that writes them +is `StorageNodeWorkloadReconciler`. Rows that named the retired `StorageNodeSet` +name the cluster instead. + +| # | Scenario | Type | Test | +|-------|-----------------------------------------------------------------------------------------------------------|------------|------------------------------------------------------------------| +| U-65 | No DaemonSet present: one is created | Positive | `TestARotatedCertificateRollsTheStoragePods` | +| U-66 | A DaemonSet present: it is updated in place rather than recreated | Positive | `TestARotatedCertificateRollsTheStoragePods` | +| U-67 | TLS disabled: the pod template carries no serving-certificate mount | Negative | — | +| U-68 | TLS enabled: the pod template mounts the serving certificate | Positive | — | +| U-69 | The cert-manager provider: the Certificate is created alongside | Positive | — | +| U-70 | User-supplied container resources override the defaults | Positive | `TestBuildStorageNodeDaemonSetUserResourcesOverrideDefaults` | +| U-71 | The ServiceAccount carries an owner reference to its parent | Positive | — | +| U-72 | ClusterRoleBinding names include the namespace, so two namespaces do not collide | Positive | `TestBuildStorageNodeSetClusterRoleBindingNameIncludesNamespace` | +| U-73 | Namespace-specific ClusterRoleBindings are created | Positive | — | +| U-74 | The SPDK proxy Service is created | Positive | — | +| U-75 | Every workload object carries an owner reference to the `StorageCluster` | Positive | — | +| U-76 | Deleting the `StorageCluster` cascades to every workload object | Positive | — | +| U-264 | No image on the cluster: the image is taken from the singleton `ControlPlane`, read at the stored version | Regression | — | +| U-307 | Neither the cluster nor the `ControlPlane` states an image: the write is refused | Negative | `TestAWorkloadWithNoImageAnywhereIsRefused` | +| U-308 | The config generator mounts `/dev` and `/sys`, which the init container reads | Positive | `TestBuildStorageNodeDaemonSetConfigGeneratorMountsDevAndSys` | +| U-309 | A conflict on the DaemonSet write is retried against a fresh read | Regression | `TestTheDaemonSetWriteRetriesAConflict` | ### Workload: Storage-Plane Labels (design §5.2) -File: `operator/internal/controllers/node/storagenodeset_controller_unit_test.go` +File: `operator/internal/controllers/node/enrollment_test.go` -| # | Scenario | Type | Test | -|------|---------------------------------------------------------------------------------|----------|-------------------------------------| -| U-77 | Workers gain the node-set label and carry no retired node-type label | Positive | `TestStorageNodeSetLabelingHelpers` | -| U-78 | A node that has come online gets its per-slot storage-node-uuid label | Positive | — | -| U-79 | A node with no UUID yet contributes no slot label | Negative | — | -| U-80 | The slot label key is stable across a UUID change, and only the value moves | Positive | — | -| U-81 | A worker hosting nodes of two clusters carries two non-colliding slot keys | Positive | — | -| U-82 | A stale slot label whose node is gone is removed | Positive | — | -| U-83 | The node `List` fails: the reconcile aborts and deletes no label | Negative | — | -| U-84 | A worker with no storage node: its slot labels are removed, the others are kept | Boundary | — | +| # | Scenario | Type | Test | +|------|---------------------------------------------------------------------------------|----------|-------------------------------------------------| +| U-77 | Workers gain the cluster's storage-plane label | Positive | `TestTheWorkloadEnrollsEveryWorkerItHasANodeOn` | +| U-78 | A node that has come online gets its per-slot storage-node-uuid label | Positive | — | +| U-79 | A node with no UUID yet contributes no slot label | Negative | — | +| U-80 | The slot label key is stable across a UUID change, and only the value moves | Positive | — | +| U-81 | A worker hosting nodes of two clusters carries two non-colliding slot keys | Positive | — | +| U-82 | A stale slot label whose node is gone is removed | Positive | — | +| U-83 | The node `List` fails: the reconcile aborts and deletes no label | Negative | — | +| U-84 | A worker with no storage node: its slot labels are removed, the others are kept | Boundary | — | ### Workload: Per-Node Configuration (design §5.3) -File: `operator/internal/controllers/node/storagenodeset_storagenode_unit_test.go` - -| # | Scenario | Type | Test | -|-------|--------------------------------------------------------------------------------------------|----------|---------------------------------------------------------------------| -| U-85 | Cluster sizing values appear in a worker's env file | Positive | `TestBuildPerNodeEnvFile_UsesClusterSizingValues` | -| U-86 | Cluster sizing is identical in every worker's entry | Positive | `TestBuildPerNodeEnvFile_ClusterSizingIdenticalAcrossWorkers` | -| U-87 | Every key the init container reads is present, including the empty ones | Boundary | `TestBuildPerNodeEnvFile_ContainsAllRequiredKeys` | -| U-88 | The cluster is missing its required sizing: the write is refused with a named error | Negative | `TestReconcilePerNodeConfigMap_RejectsClusterMissingRequiredSizing` | -| U-89 | Every worker gets an entry carrying the cluster sizing | Positive | `TestReconcilePerNodeConfigMap_WritesClusterSizingForEveryWorker` | -| U-90 | Two nodes with different device filters get different entries | Positive | — | -| U-91 | A node with no `spec.config`: its entry carries the empty values, not missing keys | Boundary | — | -| U-92 | A device list containing a shell metacharacter is quoted rather than interpolated | Negative | — | -| U-262 | `VCPU_COUNT` and `MAX_HUGE_PAGES_SIZE` come from the node's own sizing | Positive | — | -| U-263 | Two nodes mid-roll: their entries differ in those two keys and agree on `MAX_SUBSYS_COUNT` | Boundary | — | -| U-244 | A `deviceNames` entry that is a PCI address reaches the node as one | Positive | — | -| U-245 | A `deviceNames` entry that is a device path reaches the node as one | Positive | — | -| U-246 | A mixed `deviceNames` list: both forms reach the node, in the order given | Boundary | — | -| U-247 | `deviceNames` set alongside `pcieAllowList`: the explicit list wins (design §3.1) | Boundary | — | -| U-93 | The ConfigMap is written before the DaemonSet on a fresh reconcile | Positive | — | -| U-94 | A node deleted: its entry is removed from the ConfigMap | Positive | — | +File: `operator/internal/controllers/node/pernodeconfig_test.go` + +| # | Scenario | Type | Test | +|-------|--------------------------------------------------------------------------------------------|----------|--------------------------------------------------------| +| U-85 | Cluster sizing values appear in a worker's env file | Positive | — | +| U-86 | Cluster sizing is identical in every worker's entry | Positive | — | +| U-87 | Every key the init container reads is present, including the empty ones | Boundary | — | +| U-88 | The cluster is missing its required sizing: the write is refused with a named error | Negative | — | +| U-89 | Every worker gets an entry carrying the cluster sizing | Positive | — | +| U-90 | Two nodes with different device filters get different entries | Positive | — | +| U-91 | A node with no `spec.config`: its entry carries the empty values, not missing keys | Boundary | — | +| U-92 | A device list containing a shell metacharacter is quoted rather than interpolated | Negative | — | +| U-262 | `VCPU_COUNT` and `MAX_HUGE_PAGES_SIZE` come from the node's own sizing | Positive | — | +| U-263 | Two nodes mid-roll: their entries differ in those two keys and agree on `MAX_SUBSYS_COUNT` | Boundary | — | +| U-244 | A `deviceNames` entry that is a PCI address reaches the node as one | Positive | — | +| U-245 | A `deviceNames` entry that is a device path reaches the node as one | Positive | — | +| U-246 | A mixed `deviceNames` list: both forms reach the node, in the order given | Boundary | — | +| U-247 | `deviceNames` set alongside `pcieAllowList`: the explicit list wins (design §3.1) | Boundary | — | +| U-93 | The ConfigMap is written before the DaemonSet on a fresh reconcile | Positive | — | +| U-94 | A node deleted: its entry is removed from the ConfigMap | Positive | — | +| U-310 | A relocation clones the source worker's entry onto the target | Positive | `TestTheTargetInheritsTheSourcesEntryWithTheNewDrives` | +| U-311 | Cloning from a worker that has no entry is refused rather than writing an empty one | Negative | `TestCloningFromAWorkerWithNoEntryIsRefused` | +| U-312 | Cloning before the ConfigMap exists is not a failure: the ordinary pass writes it | Boundary | `TestNothingToCloneFromIsNotAFailure` | +| U-313 | The source worker's own entry is left alone by the clone | Negative | `TestTheTargetInheritsTheSourcesEntryWithTheNewDrives` | ### Workload: Endpoints and Certificate Rotation (design §5.4) -Files: `operator/internal/controllers/node/storagenodeset_controller_unit_test.go`, -`operator/internal/utils/storage_nodeset_ds_test.go` - -| # | Scenario | Type | Test | -|-------|------------------------------------------------------------------------------------|----------|----------------------------------------------------------------------| -| U-95 | The per-pod address is built from the worker's hostname label and the namespace | Positive | `TestStorageNodeSetAPIAddress` | -| U-96 | The EndpointSlice check matches the address builder's output exactly | Positive | `TestEndpointSliceHasWorker_MatchesBuilderOutput` | -| U-97 | A dotted worker hostname is truncated to a valid endpoint label | Boundary | `TestBuildSpdkProxyEndpointSlice_DottedNodeNameTruncates` | -| U-98 | Two workers whose first label segment collides: the build fails rather than merges | Negative | `TestBuildSpdkProxyEndpointSlice_CollidingFirstLabelFails` | -| U-99 | SPDK proxy EndpointSlices are built from the running pods | Positive | `TestReconcileSpdkProxyEndpointSlices` | -| U-100 | Two pods sharing a first label segment: reported rather than silently merged | Negative | `TestReconcileSpdkProxyEndpointSlices_DuplicateFirstSegment` | -| U-101 | The RPC port is read from the pod's environment, falling back to its name | Boundary | `TestExtractSpdkProxyRpcPort_FallbackToPodName` | -| U-102 | The serving certificate's revision is stamped onto the pod template | Positive | `TestStorageNodeSetDaemonSetTLSSecretRevisionAnnotation` | -| U-103 | The certificate Secret rotates: the DaemonSet rolls | Positive | `TestStorageNodeSetDaemonSetReconcileRollsOnTLSSecretRevisionChange` | -| U-104 | The TLS serving environment reaches the container | Positive | `TestStorageNodeSetDaemonSetSBTLSServeEnv` | -| U-105 | A Secret that is not the storage-node-api certificate: the predicate ignores it | Negative | `TestIsStorageNodeSetTLSSecretPredicate` | -| U-106 | The certificate Secret changes: every affected object in the namespace is enqueued | Positive | `TestTLSSecretToStorageNodeSetRequestsEnqueuesAllInNamespace` | -| U-107 | Certificates and Services are reconciled together for the cert-manager provider | Positive | `TestReconcileServicesAndServingCertificatesForCertManagerProvider` | +Files: `operator/internal/controllers/node/rotation_test.go`, `migrate_test.go`, +`operator/internal/utils/storage_node_workload_test.go` + +| # | Scenario | Type | Test | +|-------|------------------------------------------------------------------------------------|----------|------------------------------------------------------------| +| U-95 | The per-pod address is built from the worker's hostname label and the namespace | Positive | `TestStorageNodeSetAPIAddress` | +| U-96 | The EndpointSlice check matches the address builder's output exactly | Positive | `TestPreparingWaitsForTheNameToResolveAndNotOnlyForThePod` | +| U-97 | A dotted worker hostname is truncated to a valid endpoint label | Boundary | `TestBuildSpdkProxyEndpointSlice_DottedNodeNameTruncates` | +| U-98 | Two workers whose first label segment collides: the build fails rather than merges | Negative | `TestBuildSpdkProxyEndpointSlice_CollidingFirstLabelFails` | +| U-99 | SPDK proxy EndpointSlices are built from the running pods | Positive | — | +| U-100 | Two pods sharing a first label segment: reported rather than silently merged | Negative | — | +| U-101 | The RPC port is read from the pod's environment, falling back to its name | Boundary | — | +| U-102 | The serving certificate's revision is stamped onto the pod template | Positive | `TestARotatedCertificateRollsTheStoragePods` | +| U-103 | The certificate Secret rotates: the DaemonSet rolls | Positive | `TestARotatedCertificateRollsTheStoragePods` | +| U-104 | The TLS serving environment reaches the container | Positive | — | +| U-105 | A Secret that is not the storage-node-api certificate: the predicate ignores it | Negative | — | +| U-106 | The certificate Secret changes: every affected object in the namespace is enqueued | Positive | — | +| U-107 | Certificates and Services are reconciled together for the cert-manager provider | Positive | — | +| U-314 | TLS disabled: nothing is stamped, so no pod is rolled by the absence | Boundary | `TestWithoutTLSNothingIsStampedOnTheTemplate` | ### Operation: Lifecycle and Lock (design §7.1, §11) -File: `operator/internal/controllers/node/storagenodeops_controller_unit_test.go` - -| # | Scenario | Type | Test | -|-------|------------------------------------------------------------------------------------|----------|-----------------------------------------------------------| -| U-108 | The lock is free: it is acquired and the phase becomes `Running` | Positive | `TestAcquireLock_SetsActiveOpsRefAndTransitionsToRunning` | -| U-109 | Another operation holds the lock: this one stays `Pending` and requeues | Negative | `TestAcquireLock_RequeuesWhenAnotherOpsActive` | -| U-110 | A `Remove` acquiring the lock enters its graph's initial step | Positive | `TestAcquireLock_RemoveDrainSetsValidatingSubPhase` | -| U-111 | Success: the phase is `Succeeded` and the lock is cleared | Positive | `TestSucceedOps_SetsPhaseAndClearsLock` | -| U-112 | Failure: the phase is `Failed` with a message, and the lock is cleared | Positive | `TestFailOps_SetsPhaseAndClearsLock` | -| U-113 | A release by a non-owner: the lock is left alone | Negative | `TestReleaseLock_OnlyClearsIfOwner` | -| U-114 | Advancing the step persists it before the next side effect | Positive | `TestAdvanceSubPhase_UpdatesSubPhaseAndResetsTrigger` | -| U-115 | An unknown action: the operation fails terminally with the action in the message | Negative | `TestDispatch_UnknownActionFails` | -| U-116 | The target node does not exist: the operation fails with a not-found message | Negative | — | -| U-117 | A terminal operation re-reconciled: no side effect, and the lock is released again | Negative | — | -| U-118 | Two reconcilers acquiring one free lock: the loser gets 409 and requeues | Negative | — | -| U-119 | The operation is deleted while `Running`: the finalizer releases the lock | Positive | — | -| U-120 | Operations on two different nodes run without contending | Positive | — | -| U-121 | The cluster is not active: the operation holds and emits `ClusterNotReady` | Negative | — | -| U-122 | The cluster is rebalancing: the operation holds rather than proceeding | Negative | — | -| U-123 | The cluster becomes active: the held operation resumes with no further input | Positive | — | -| U-124 | A node event wakes a queued operation before its requeue interval elapses | Positive | — | +Files: `operator/internal/controllers/node/opslock_test.go`, `advance_test.go`, +`watches_test.go`, `remove_deadlock_test.go` + +| # | Scenario | Type | Test | +|-------|-------------------------------------------------------------------------------------|----------|------------------------------------------------------------| +| U-108 | The lock is free: it is acquired and the phase becomes `Running` | Positive | `TestTakingTheLockIsWhatStartsTheOperation` | +| U-109 | Another operation holds the lock: this one stays `Pending` and requeues | Negative | `TestAnOperationWaitsForTheOneHoldingTheNode` | +| U-110 | An operation acquiring the lock enters its graph's initial step with a deadline | Positive | `TestTheFirstPassArmsTheStepAMachineIsBornIn` | +| U-111 | Success: the phase is `Succeeded` and the lock is cleared | Positive | `TestFinishingWritesTheOutcomeAndLetsTheNodeGo` | +| U-112 | Failure: the phase is `Failed` with a message, and the lock is cleared | Positive | `TestAStepThatOutlivedItsDeadlineFailsTheOperation` | +| U-113 | A release by a non-owner: the lock is left alone | Negative | `TestALateReleaseDoesNotUnlockSomebodyElsesNode` | +| U-114 | Advancing persists the next step and its deadline before the side effect | Positive | `TestAFinishedStepEntersTheNextWithItsOwnDeadline` | +| U-115 | An unknown action: the operation fails terminally with the reason in the message | Negative | `TestAStepThatBelongsToNoActionEndsTheOperation` | +| U-116 | The target node does not exist: the operation fails with a not-found message | Negative | `TestAnOperationAgainstAMissingNodeFails` | +| U-117 | A terminal operation re-reconciled: no side effect, and the lock is released again | Negative | `TestATerminalOperationStillReleasesALockItLeftBehind` | +| U-118 | Two reconcilers acquiring one free lock: the loser gets 409 and requeues | Negative | — | +| U-119 | The operation is deleted while `Running`: the finalizer releases the lock | Positive | `TestDeletingAnOperationUnlocksTheNodeFirst` | +| U-120 | Operations on two different nodes run without contending | Positive | — | +| U-121 | The cluster is not active: the operation holds and emits `ClusterNotReady` | Negative | — | +| U-122 | The cluster is rebalancing: the operation holds rather than proceeding | Negative | `TestAnOperationHoldsWhileItsClusterIsRebalancing` | +| U-123 | The cluster becomes active: the held operation resumes with no further input | Positive | — | +| U-124 | A node event wakes a queued operation before its requeue interval elapses | Positive | `TestANodeEventWakesTheOperationsWaitingOnIt` | +| U-315 | The holder re-reading its own lock keeps it, since every pass goes through this | Boundary | `TestTheHolderKeepsItsOwnLock` | +| U-316 | The finalizer is taken on the pass before anything is locked | Positive | `TestAnOperationTakesItsFinalizerBeforeItTakesAnything` | +| U-317 | A removal runs whatever the cluster says, because removal is how a cluster recovers | Boundary | `TestRemovalDoesNotWaitOnTheCluster` | +| U-318 | Every other action waits on the cluster gate | Negative | `TestEveryOtherActionWaitsOnTheCluster` | +| U-319 | A removal proceeds against a rebalancing cluster rather than holding | Boundary | `TestARemovalRunsAgainstARebalancingCluster` | +| U-320 | `status.observedGeneration` advances, which is how an observed abort is visible | Positive | `TestTheObservedGenerationMovesWhenTheOperationIsLookedAt` | +| U-321 | A cluster event wakes its own nodes and nobody else's | Positive | `TestAClusterEventWakesItsOwnNodes` | +| U-322 | A worker event wakes the nodes that run on it, which is how a cordon arrives | Positive | `TestAWorkerEventWakesTheNodesOnIt` | ### Operation: The Single-Step Actions (design §7.3) -File: `operator/internal/controllers/node/storagenodeops_controller_unit_test.go` - -| # | Scenario | Type | Test | -|-------|----------------------------------------------------------------------------------|----------|------| -| U-125 | `Suspend`: the call is issued, and the step completes when the node is suspended | Positive | — | -| U-126 | `Resume`: the step completes when the node is online | Positive | — | -| U-127 | `Shutdown`: the step completes when the node is offline | Positive | — | -| U-128 | `Restart`: `reattachVolume` and `force` are passed through when set | Positive | — | -| U-129 | The node is already at the target state: the call is not issued at all | Negative | — | -| U-130 | The call returns 5xx: the step is retried and the phase does not advance | Negative | — | -| U-131 | The call returns 4xx: the step is retried, and the body reaches the event | Negative | — | -| U-132 | The call is retried after a timeout: the endpoint is called at most once more | Negative | — | -| U-133 | The node never reaches the target state: the step's deadline expires and fails | Boundary | — | -| U-134 | A late response after the deadline expired: ignored, no second commit | Negative | — | +File: `operator/internal/controllers/node/actions_test.go` + +| # | Scenario | Type | Test | +|-------|----------------------------------------------------------------------------------|----------|----------------------------------------------------------| +| U-125 | `Suspend`: the call is issued, and the step completes when the node is suspended | Positive | `TestACallIsIssuedWhenTheNodeIsNotThereYet` | +| U-126 | `Resume`: the step completes when the node is online | Positive | `TestTheWaitIsOverWhenTheNodeReportsWhatTheActionWasFor` | +| U-127 | `Shutdown`: the step completes when the node is offline | Positive | `TestTheWaitIsOverWhenTheNodeReportsWhatTheActionWasFor` | +| U-128 | `Restart`: `reattachVolume` and `force` are passed through when set | Positive | `TestOnlyTheFlagsTheOperationStatesAreSent` | +| U-129 | The node is already at the target state: the call is not issued at all | Negative | `TestACallIsSkippedWhenTheNodeIsAlreadyThere` | +| U-130 | The call returns 5xx: the step is retried and the phase does not advance | Negative | — | +| U-131 | The call returns 4xx: the step is retried, and the body reaches the event | Negative | — | +| U-132 | The call is retried after a timeout: the endpoint is called at most once more | Negative | — | +| U-133 | The node never reaches the target state: the step's deadline expires and fails | Boundary | `TestAStepThatOutlivedItsDeadlineFailsTheOperation` | +| U-134 | A late response after the deadline expired: ignored, no second commit | Negative | — | +| U-323 | A restart has no state to skip on, so it is issued against an online node | Boundary | `TestARestartIsIssuedAgainstAnOnlineNode` | +| U-324 | An unstated flag is not sent, since not sending is not the same as sending false | Boundary | `TestOnlyTheFlagsTheOperationStatesAreSent` | +| U-325 | An operation against a node with no backend UUID: terminal, not retried | Negative | `TestAnUnprovisionedNodeEndsTheOperation` | ### Operation: Volume Classification (design §8.1) -File: `operator/internal/controllers/node/drain_unit_test.go` - -| # | Scenario | Type | Test | -|-------|---------------------------------------------------------------------------------|----------|---------------------------------------------------------------| -| U-135 | A volume matching a `PersistentVolume`: classified as PV-managed | Positive | `TestMatchVolumesToPVs_PVManaged` | -| U-136 | A volume whose claim carries the pin annotation: classified as pinned | Negative | `TestMatchVolumesToPVs_Pinned` | -| U-137 | A volume matching no `PersistentVolume`: classified as unmanaged | Negative | `TestMatchVolumesToPVs_Unmanaged` | -| U-138 | A volume matching the system filter: excluded from every bucket | Negative | `TestMatchVolumesToPVs_SystemVolumeSkipped` | -| U-139 | A node with no volumes: no drain work is produced and no error is returned | Boundary | `TestMatchVolumesToPVs_EmptyNodeSkipsMigration` | -| U-140 | A node holding only system volumes: nothing is migrated | Boundary | `TestMatchVolumesToPVs_OnlySystemVolumes` | -| U-141 | The default system filter matches the rebalancer's benchmark volume names | Positive | `TestResolveOpsSystemVolumeFilter_UsesDefaultWhenNoDrain` | -| U-142 | A custom `systemVolumeFilterRegex` replaces the default | Positive | `TestResolveOpsSystemVolumeFilter_UsesCustomPattern` | -| U-143 | A malformed `systemVolumeFilterRegex`: the operation fails with the parse error | Negative | `TestResolveOpsSystemVolumeFilter_InvalidPatternReturnsError` | -| U-144 | A volume that is both pinned and unmanaged: both blockers are reported | Boundary | — | -| U-145 | A claim in another namespace with the same name: not read as this one's pin | Negative | — | -| U-146 | An empty `systemVolumeFilterRegex`: matches nothing rather than everything | Boundary | — | +File: `operator/internal/controllers/node/classify_test.go` + +| # | Scenario | Type | Test | +|-------|--------------------------------------------------------------------------------------------|----------|------------------------------------------------------| +| U-135 | A volume matching a `PersistentVolume`: classified as PV-managed | Positive | `TestAVolumeKubernetesAccountsForIsMovable` | +| U-136 | A volume whose claim carries the pin annotation: classified as pinned | Negative | `TestAPinnedVolumeBlocksRatherThanMoving` | +| U-137 | A volume matching no `PersistentVolume`: classified as unmanaged | Negative | `TestAVolumeNothingAccountsForBlocks` | +| U-138 | A volume matching the system filter: excluded from every bucket | Negative | `TestABenchmarkVolumeIsTheDrainsToDelete` | +| U-139 | A node with no volumes: no drain work is produced and no error is returned | Boundary | `TestADrainIsDoneWhenTheNodeHoldsNothingMovable` | +| U-140 | A node holding only system volumes: nothing is migrated | Boundary | `TestABenchmarkVolumeIsTheDrainsToDelete` | +| U-141 | The default system filter matches the rebalancer's benchmark volume names | Positive | `TestTheDefaultFilterMatchesTheBenchmarkVolumesOnly` | +| U-142 | A custom `systemVolumeFilterRegex` replaces the default | Positive | `TestAStatedFilterReplacesTheDefault` | +| U-143 | A malformed `systemVolumeFilterRegex`: the operation fails with the parse error | Negative | `TestAPatternThatDoesNotCompileEndsTheOperation` | +| U-144 | A volume that is both pinned and unmanaged: both blockers are reported | Boundary | — | +| U-145 | A claim in another namespace with the same name: not read as this one's pin | Negative | — | +| U-146 | An empty `systemVolumeFilterRegex`: falls back to the default rather than matching nothing | Boundary | — | +| U-326 | A claim that cannot be read: counted where it blocks, and the census says so | Negative | `TestAClaimThatCannotBeReadMakesTheCensusIncomplete` | +| U-327 | A volume whose delete the backend has accepted is not counted | Boundary | `TestAVolumeAlreadyBeingDeletedIsNotCounted` | +| U-328 | A peer's volumes in the same pools are not this node's to move | Negative | `TestAPeersVolumesAreNotThisNodesToMove` | +| U-329 | A `PersistentVolume` of another driver accounts for nothing | Negative | `TestAnotherDriversVolumeAccountsForNothing` | +| U-330 | A `PersistentVolume` with no claim cannot be pinned, so it is movable | Boundary | `TestAnUnclaimedVolumeIsMovable` | +| U-331 | A volume handle naming no volume indexes nothing | Boundary | `TestAnEmptyVolumeHandleIndexesNothing` | ### Operation: The Remove Graph (design §8.2, §8.3) -File: `operator/internal/controllers/node/storagenodeops_controller_unit_test.go` - -| # | Scenario | Type | Test | -|-------|----------------------------------------------------------------------------------|----------|---------------------------------------------| -| U-147 | Validation clear: the step advances to `Suspending` | Positive | — | -| U-148 | Pinned volumes present: the step holds and no suspend is issued | Negative | — | -| U-149 | Unmanaged volumes present: the step holds and no suspend is issued | Negative | — | -| U-150 | The pin is removed: the next reconcile advances without further input | Positive | — | -| U-151 | The node is already suspended: no suspend call is issued and the step advances | Negative | — | -| U-152 | Migration targets are spread evenly across the online peers | Positive | `TestRoundRobinDistributesEvenly` | -| U-153 | No online peer: the step holds and emits `NoMigrationTarget` | Negative | `TestRoundRobinErrorsWhenNoTargetAvailable` | -| U-154 | Offline peers are excluded from target selection | Negative | `TestRoundRobinSkipsOfflineNodes` | -| U-155 | Exactly one online peer: every volume goes to it | Boundary | — | -| U-156 | Every migration completed: the step advances to `Verifying` | Positive | — | -| U-157 | Verification finds non-system volumes: the step holds and retries | Negative | — | -| U-158 | Verification finds only system volumes: they are deleted, then the step advances | Positive | — | -| U-159 | A system-volume delete returns 404: treated as success | Boundary | — | -| U-160 | A system-volume delete is rejected: the node is resumed and the operation fails | Negative | — | -| U-161 | The node delete returns 200, 204, or 404: the operation succeeds | Boundary | — | -| U-162 | The node delete returns 5xx: retried, and the operation does not fail | Negative | — | -| U-163 | The node delete is rejected: the node is resumed and the operation fails | Negative | — | -| U-164 | A resume that itself fails: the operation still reaches `Failed`, with an event | Negative | — | -| U-165 | `spec.abort` set during `Validating`: `Aborted` with no resume call issued | Boundary | — | -| U-166 | `spec.abort` set during `MigratingVolumes`: migrations deleted, node resumed | Positive | — | -| U-167 | A step's deadline expires mid-drain: the node is resumed and the operation fails | Boundary | — | -| U-168 | `status.drain.volumesTotal` is written once and not recomputed on later passes | Positive | — | -| U-169 | A drain with zero volumes: `volumesTotal` is 0 and the step advances immediately | Boundary | — | - -### Operation: PersistentVolumeOps Lifecycle (design §8.4) - -File: `operator/internal/controllers/node/drain_unit_test.go` - -| # | Scenario | Type | Test | -|-------|--------------------------------------------------------------------------------------------------------------------------|----------|--------------------------------------------------| -| U-170 | Generated migration names are valid DNS labels | Positive | `TestDrainMigrationNameIsDNSValid` | -| U-171 | Two long volume names sharing a prefix produce distinct migration names | Boundary | `TestDrainMigrationNameNoCollisionOnLongPVNames` | -| U-172 | An operator restart mid-drain does not recreate existing migration objects | Negative | — | -| U-173 | A completed migration is deleted and the counter is written first | Positive | — | -| U-174 | A failed migration is deleted and replaced against a fresh target | Positive | — | -| U-175 | A migration deleted out of band is recreated rather than counted as complete | Negative | — | -| U-176 | Every migration carries `spec.creatorRef` with the operation's UID | Positive | — | -| U-177 | The same volume name in two namespaces produces two distinct objects | Boundary | — | -| U-248 | Every migration carries `storage.simplyblock.io/managed-by: storagenodeops`, and a `List` on the label finds the fan-out | Positive | — | +Files: `operator/internal/controllers/node/drain_test.go`, `peertargets_test.go`, +`advance_test.go`, `remove_gone_test.go` + +| # | Scenario | Type | Test | +|-------|------------------------------------------------------------------------------------------|------------|----------------------------------------------------------------------------------------| +| U-147 | Validation clear: the step advances to `Suspending` | Positive | `TestValidationWritesTheTotalTheDrainIsMeasuredAgainst` | +| U-148 | Pinned volumes present: the step holds and no suspend is issued | Negative | `TestAPinnedVolumeStopsTheDrainBeforeItSuspendsAnything` | +| U-149 | Unmanaged volumes present: the step holds and no suspend is issued | Negative | `TestAnUnmanagedVolumeStopsTheDrain` | +| U-150 | The pin is removed: the next reconcile advances without further input | Positive | — | +| U-151 | The node is already suspended: no suspend call is issued and the step advances | Negative | `TestTheSuspendIsSkippedWhenTheNodeIsAlreadyOutOfService` | +| U-152 | Migration targets are spread evenly across the online peers | Positive | `TestTheVolumesAreSpreadOverEveryOnlinePeer` | +| U-153 | No online peer: the step holds and emits `NoMigrationTarget` | Negative | `TestADrainWithNowhereToMoveToHolds` | +| U-154 | Offline peers are excluded from target selection | Negative | `TestAnOfflinePeerIsNotATarget` | +| U-155 | Exactly one online peer: every volume goes to it | Boundary | — | +| U-156 | Every migration completed: the step advances to `Verifying` | Positive | `TestADrainIsDoneWhenTheNodeHoldsNothingMovable` | +| U-157 | Verification finds non-system volumes: the step holds and retries | Negative | `TestVerificationHoldsWhileAUsersVolumeIsStillThere` | +| U-158 | Verification finds only system volumes: they are deleted, then the step advances | Positive | `TestVerificationDeletesTheBenchmarkVolumesAndRereads` | +| U-159 | A system-volume delete returns 404: treated as success | Boundary | — | +| U-160 | A system-volume delete is rejected: the node is resumed and the operation fails | Negative | `TestABenchmarkVolumeThatCannotBeDeletedEndsTheDrain` | +| U-161 | The node delete returns 200, 204, or 404: the operation succeeds | Boundary | `TestAnAcceptedRemovalFinishesTheDrain` | +| U-162 | The node delete returns 5xx: retried, and the operation does not fail | Negative | — | +| U-163 | The node delete is rejected: the node is resumed and the operation fails | Negative | `TestARefusedRemovalEndsTheDrain`, `TestAStepThatOutlivedItsDeadlineFailsTheOperation` | +| U-164 | A resume that itself fails: the operation still reaches `Failed`, with an event | Negative | — | +| U-165 | `spec.abort` set during `Validating`: `Aborted` with no resume call issued | Boundary | — | +| U-166 | `spec.abort` set during `MigratingVolumes`: migrations deleted, node resumed | Positive | `TestAnAbortAtAnAbortableStepStopsAndResumesTheNode` | +| U-167 | A step's deadline expires mid-drain: the node is resumed and the operation fails | Boundary | `TestAStepThatOutlivedItsDeadlineFailsTheOperation` | +| U-168 | `status.drain.volumesTotal` is written once and not recomputed on later passes | Positive | `TestValidationWritesTheTotalTheDrainIsMeasuredAgainst` | +| U-169 | A drain with zero volumes: `volumesTotal` is 0 and the step advances immediately | Boundary | `TestADrainIsDoneWhenTheNodeHoldsNothingMovable` | +| U-332 | A node the control plane has forgotten: every step of the removal is done | Regression | `TestARemovalOfAGoneNodeIsDoneAtEveryStep` | +| U-333 | An incomplete census is retried rather than reported as a blocker | Negative | `TestAnIncompleteCensusIsRetriedRatherThanReportedAsABlocker` | +| U-334 | A node already offline is past the state a suspend produces, so none is issued | Boundary | `TestTheSuspendIsSkippedWhenTheNodeIsAlreadyOutOfService` | +| U-335 | An online node is suspended, and the step waits for the backend to report it | Positive | `TestAnOnlineNodeIsSuspendedAndWaitedFor` | +| U-336 | An empty node passes verification | Boundary | `TestAnEmptyNodePassesVerification` | +| U-337 | A suspended peer is not a migration target either | Negative | `TestASuspendedPeerIsNotATarget` | +| U-338 | The drained node is never chosen as its own target | Negative | `TestTheDrainedNodeIsNeverItsOwnTarget` | +| U-339 | The peer assignment is the same on every pass, so a retry is deliberate | Positive | `TestTheAssignmentIsTheSameOnEveryPass` | +| U-340 | The peers come from the stream once it has delivered, and the control plane is not asked | Positive | `TestThePeersComeFromTheStreamOnceItHasDelivered` | +| U-341 | A stream that has not delivered falls back to the control plane | Boundary | `TestAnUndeliveredStreamFallsBackToTheControlPlane` | + +### Operation: The Fan-Out (design §8.4) + +Files: `operator/internal/controllers/node/drain_test.go`, `remove_fanout_test.go` + +| # | Scenario | Type | Test | +|-------|----------------------------------------------------------------------------------------------------------|------------|-----------------------------------------------------------| +| U-170 | Generated migration names are valid DNS labels | Positive | `TestEveryMovableVolumeIsGivenAMove` | +| U-171 | Two long volume names sharing a prefix produce distinct migration names | Boundary | — | +| U-172 | An operator restart mid-drain does not recreate existing migration objects | Negative | `TestAVolumeAlreadyMovingIsNotGivenASecondMove` | +| U-173 | A completed migration is deleted and the counter is written first | Positive | `TestFinishedMovesAreRecordedAndThenReaped` | +| U-174 | A failed migration is deleted and replaced against a fresh target | Positive | `TestAFailedMoveIsRetriedRatherThanFailingTheDrain` | +| U-175 | A migration deleted out of band is recreated rather than counted as complete | Negative | `TestEveryMovableVolumeIsGivenAMove` | +| U-176 | Every migration carries `spec.creatorRef` with the operation's UID | Positive | `TestTheFanOutRecordsItsCreatorWithoutOwningTheOperation` | +| U-177 | The same volume name in two namespaces produces two distinct objects | Boundary | — | +| U-248 | Every migration carries `storage.simplyblock.io/drain-node`, and a `List` on the label finds the fan-out | Positive | `TestTheFanOutIsFoundAgainByItsLabel` | +| U-342 | The cluster-scoped kind carries no owner reference, which garbage collection would follow | Regression | `TestTheFanOutRecordsItsCreatorWithoutOwningTheOperation` | +| U-343 | With the registered kind in use, the move is owned by the drain as it always was | Positive | `TestTheLegacyFanOutIsStillOwnedByItsDrain` | +| U-344 | Deleting a drain reaps the fan-out it raised | Positive | `TestDeletingADrainStopsWhatItFannedOut` | +| U-345 | A move still copying holds the deletion rather than being reaped mid-copy | Negative | `TestADrainBeingDeletedWaitsForAMoveStillRunning` | +| U-346 | A retried move is announced, so a drain that keeps retrying is not read as stalled | Positive | `TestAFailedMoveIsRetriedRatherThanFailingTheDrain` | ### Operation: The Migrate Graph (design §9) -File: `operator/internal/controllers/node/storagenodeops_migrate_config_unit_test.go` - -| # | Scenario | Type | Test | -|-------|----------------------------------------------------------------------------------|----------|---------------------------------------------------| -| U-178 | The target worker's configuration is cloned from the source | Positive | `TestEnsureMigratedWorkerConfig` | -| U-179 | The topology re-point moves the node's configuration to the target | Positive | `TestReconcileMigratedTopologyMigratesNodeConfig` | -| U-180 | `newSsdPcie` addresses are merged into the effective allow list | Positive | `TestMergePcieAllowedIntoEnvFile` | -| U-181 | Merging two PCI lists produces no duplicates | Positive | `TestMergePcieList` | -| U-182 | A quoted shell list is parsed back into its elements | Positive | `TestParseShellCSV` | -| U-183 | An empty PCI list merged with additions yields just the additions | Boundary | `TestMergePcieList` | -| U-184 | The target worker does not exist: the operation fails informatively | Negative | — | -| U-185 | The target worker is the node's current worker: the operation fails immediately | Negative | — | -| U-186 | `spec.migrate` absent for `action: Migrate`: rejected | Negative | — | -| U-187 | The target's pod is not `Ready`: `Preparing` holds and no restart is issued | Negative | — | -| U-188 | The pod is `Ready` but not yet in the EndpointSlice: `Preparing` still holds | Boundary | — | -| U-189 | The pod is `Ready` and in the EndpointSlice: the step advances | Positive | — | -| U-190 | The relocation restart is forced by default | Positive | — | -| U-191 | An explicit `spec.force: false` is honored rather than overridden | Negative | — | -| U-192 | The node is still online after the restart call: `Relocating` holds | Negative | — | -| U-193 | The node has left online: the step advances to `AwaitingNode` | Positive | — | -| U-194 | The node returns online: the step advances to `Promoting` | Positive | — | -| U-195 | The promote is issued exactly once across several reconciles | Negative | — | -| U-196 | `spec.abort` during `Preparing`: `Aborted`, and no control-plane call was made | Positive | — | -| U-197 | `spec.abort` during `Promoting`: refused by the graph, and the operation runs on | Negative | — | -| U-198 | The topology re-point happens after the promote, never before | Positive | — | -| U-199 | The source worker loses its storage-plane labels when no node remains on it | Positive | — | -| U-200 | The source worker keeps its labels when another node still runs there | Boundary | — | +Files: `operator/internal/controllers/node/migrate_test.go`, +`pernodeconfig_test.go` + +| # | Scenario | Type | Test | +|-------|----------------------------------------------------------------------------------|----------|---------------------------------------------------------------| +| U-178 | The target worker's configuration is cloned from the source | Positive | `TestPreparingWritesTheTargetsConfigurationAndLabelsIt` | +| U-179 | The topology re-point moves the node onto the target worker | Positive | `TestThePromoteIsFollowedByTheTopologyRepoint` | +| U-180 | `newSsdPcie` addresses are merged into the effective allow list | Positive | `TestTheTargetInheritsTheSourcesEntryWithTheNewDrives` | +| U-181 | Merging two PCI lists produces no duplicates | Positive | `TestTheBoundDrivesJoinTheListInTheOrderItWasWritten` | +| U-182 | A quoted shell list is parsed back into its elements | Positive | `TestAListIsReadBackTheWayItWasWritten` | +| U-183 | An empty PCI list merged with additions yields just the additions | Boundary | `TestAnEntryWithNoAllowListGetsOne` | +| U-184 | The target worker does not exist: the operation fails informatively | Negative | `TestATargetThatCannotHostTheNodeIsRefused` | +| U-185 | The target worker is the node's current worker: the step is already done | Boundary | `TestARelocationToTheHostTheNodeIsOnIsAlreadyDone` | +| U-186 | `spec.migrate` absent for `action: Migrate`: the operation is terminal | Negative | `TestARelocationWithNoTargetEndsTheOperation` | +| U-187 | The target's pod is not `Ready`: `Preparing` holds and no restart is issued | Negative | `TestPreparingWritesTheTargetsConfigurationAndLabelsIt` | +| U-188 | The pod is `Ready` but not yet in the EndpointSlice: `Preparing` still holds | Boundary | `TestPreparingWaitsForTheNameToResolveAndNotOnlyForThePod` | +| U-189 | The pod is `Ready` and in the EndpointSlice: the step advances | Positive | `TestPreparingWaitsForTheNameToResolveAndNotOnlyForThePod` | +| U-190 | The relocation restart is forced by default | Positive | `TestTheRelocationRestartIsAimedAtTheTargetAndForced` | +| U-191 | An explicit `spec.force: false` is honored rather than overridden | Negative | `TestAStatedForceOutranksTheRelocationsDefault` | +| U-192 | The node is still online after the restart call: `Relocating` holds | Negative | `TestTheRelocationRestartIsAimedAtTheTargetAndForced` | +| U-193 | The node has left online: the step advances to `AwaitingNode` | Positive | `TestTheRelocationIsOverWhenTheNodeHasLeftOnline` | +| U-194 | The node returns online: the step advances to `Promoting` | Positive | `TestTheWaitForTheRelocatedNodeEndsWhenItIsBack` | +| U-195 | The promote is issued exactly once across several reconciles | Negative | `TestAPromoteThatAlreadyLandedIsNotIssuedAgain` | +| U-196 | `spec.abort` during `Preparing`: `Aborted`, and no control-plane call was made | Positive | — | +| U-197 | `spec.abort` during `Promoting`: refused by the graph, and the operation runs on | Negative | `TestAnAbortThatArrivedTooLateIsRefusedAndTheOperationRunsOn` | +| U-198 | The topology re-point happens after the promote, never before | Positive | `TestThePromoteIsFollowedByTheTopologyRepoint` | +| U-199 | The source worker loses its storage-plane labels when no node remains on it | Positive | — | +| U-200 | The source worker keeps its labels when another node still runs there | Boundary | — | +| U-347 | A target worker that is not `Ready`: refused rather than waited on | Negative | `TestATargetThatCannotHostTheNodeIsRefused` | +| U-348 | The promote is refused while the node has not come back | Boundary | `TestAPromoteIsRefusedWhileTheNodeIsNotBack` | +| U-349 | The restart names the target's own per-pod address | Positive | `TestTheRelocationRestartIsAimedAtTheTargetAndForced` | +| U-350 | A worker that has reported no readiness condition is not read as `Ready` | Boundary | `TestAWorkerThatHasNotReportedIsNotReady` | +| U-351 | A migration binding no drive leaves the allow list byte for byte as it was | Boundary | `TestAMigrationThatBindsNoDriveLeavesTheListAlone` | ### Operation: The Host Maintenance Graph (design §10) -File: `operator/internal/controllers/node/storagenodeops_controller_unit_test.go` - -| # | Scenario | Type | Test | -|-------|-------------------------------------------------------------------------------------|----------|------| -| U-201 | A worker becomes unschedulable: a `HostMaintenance` operation is raised | Positive | — | -| U-202 | The worker is uncordoned before the operation starts: it completes as a no-op | Negative | — | -| U-203 | A second reconcile while one is running: no second operation is created | Negative | — | -| U-204 | The concurrency limit is reached: the operation holds at `Holding` | Negative | — | -| U-205 | Two nodes on one worker: the pair counts as one worker against the limit | Boundary | — | -| U-206 | The limit is the cluster's effective value, not `spec.maxConcurrentWorkerRestarts` | Boundary | — | -| U-207 | A blocking budget is created before the shutdown call | Positive | — | -| U-208 | The node reports offline: the budget is relaxed | Positive | — | -| U-209 | The budget is relaxed only after the node is offline, never before | Negative | — | -| U-210 | The storage pod is gone: the step advances to `AwaitingHost` | Positive | — | -| U-211 | The worker's API answers again: the restart is issued | Positive | — | -| U-212 | The node is already online when `Restarting` is entered: no restart call is issued | Negative | — | -| U-213 | The node comes back online: the budget and the drain label are removed | Positive | — | -| U-214 | The operation fails: the budget is still removed, so the worker stays drainable | Negative | — | -| U-215 | A stale operator self-budget from a previous crash is cleaned up | Negative | — | -| U-216 | The `AwaitingHost` deadline expires: the operation fails and the node stays offline | Boundary | — | - -### Ops Shape: The Step Machine (design §6.3, not yet built) - -File: `operator/internal/controllers/node/storagenodeops_machine_test.go`, planned. - -Seven graphs declared as one `statemachine.MultiConfig`, none of which exists -yet. - -| # | Scenario | Type | Test | -|-------|------------------------------------------------------------------------------------|----------|------| -| U-217 | Every declared graph builds, including the ones the action under test does not use | Positive | — | -| U-218 | An action with no declared graph: `ErrUnknownAction` rather than a stall | Negative | — | -| U-219 | `Remove` transitioning to `Promoting`: rejected as an illegal transition | Negative | — | -| U-220 | `Migrate` transitioning to `Removing`: rejected as an illegal transition | Negative | — | -| U-221 | `HostMaintenance` transitioning to `Suspending`: rejected | Negative | — | -| U-222 | An empty `status.step`: restores to the action's declared initial state | Boundary | — | -| U-223 | A step value that belongs to a different action: restoration fails informatively | Negative | — | -| U-224 | A step value outside the enum: restoration fails rather than stalling | Negative | — | -| U-225 | The snapshot round-trips through `Snapshot` and `FromSnapshot` unchanged | Positive | — | -| U-226 | A deadline persisted and restored is the same absolute instant | Positive | — | -| U-227 | A deadline that passed while the operator was down: restores as expired | Boundary | — | -| U-228 | A step with no deadline: restores with none rather than with a zero instant | Boundary | — | -| U-229 | A terminal step: `IsTerminal` is true and no transition is attempted | Boundary | — | -| U-230 | The outer phase machine is separate from the step machine | Positive | — | -| U-231 | Every state each graph declares appears in the step `Enum` marker | Boundary | — | -| U-232 | Every state each graph declares appears in the `status.step` CEL rule | Boundary | — | -| U-233 | The CEL rule names no value the graphs do not declare | Negative | — | -| U-234 | A stored step from another action: `ErrUnknownState`, naming the declared set | Negative | — | -| U-235 | A restore that fails: the operation is `Failed` with the error, not requeued | Negative | — | - ---- +Files: `operator/internal/controllers/node/hostmaintenance_test.go`, +`selfbudget_test.go`, `syncstatus_test.go` + +| # | Scenario | Type | Test | +|-------|-------------------------------------------------------------------------------------|------------|----------------------------------------------------------| +| U-201 | A worker becomes unschedulable: a `HostMaintenance` operation is raised | Positive | `TestACordonedWorkerRaisesItsMaintenanceWindow` | +| U-202 | The worker is uncordoned before the operation starts: it completes as a no-op | Negative | — | +| U-203 | A second reconcile while one is running: no second operation is created | Negative | `TestACordonedWorkerRaisesItsMaintenanceWindow` | +| U-204 | The concurrency limit is reached: the operation holds at `Holding` | Negative | `TestOneWorkerAtATimeIsWhatAnUnstatedConcurrencyMeans` | +| U-205 | Two nodes on one worker: the pair counts as one worker against the limit | Boundary | `TestASiblingSocketOfTheSameWorkerIsNotASecondWorker` | +| U-206 | The limit is the cluster's effective value rather than what it asked for | Boundary | `TestTheClusterSaysHowManyWorkersMayBeDownAtOnce` | +| U-207 | A blocking budget is created before the shutdown call | Positive | `TestTheEvictionIsBlockedBeforeTheNodeIsTakenDown` | +| U-208 | The node reports offline: the budget is relaxed | Positive | `TestReleasingRelaxesTheBudgetAndWaitsForThePodToGo` | +| U-209 | The budget is relaxed only after the node is offline, never before | Negative | — | +| U-210 | The storage pod is gone: the step advances to `AwaitingHost` | Positive | `TestReleasingRelaxesTheBudgetAndWaitsForThePodToGo` | +| U-211 | The worker's API answers again: the restart is issued | Positive | `TestTheNodeIsRestartedOnceTheHostIsBack` | +| U-212 | The node is already online when `Restarting` is entered: no restart call is issued | Negative | `TestTheNodeIsRestartedOnceTheHostIsBack` | +| U-213 | The node comes back online: the budget and the pod label are removed | Positive | `TestCleanupLeavesTheWorkerDrainableAgain` | +| U-214 | The operation fails: the budget is still removed, so the worker stays drainable | Negative | `TestCleanupLeavesTheWorkerDrainableAgain` | +| U-215 | A stale operator self-budget from a previous crash is cleaned up | Negative | `TestAStaleSelfBudgetIsCleanedUp` | +| U-216 | The `AwaitingHost` deadline expires: the operation fails and the node stays offline | Boundary | — | +| U-352 | A window still at `Holding` occupies no slot, so a deployment cannot deadlock | Boundary | `TestAWindowStillWaitingHoldsNoSlot` | +| U-353 | A finished window occupies no slot either | Boundary | `TestAFinishedWindowHoldsNoSlot` | +| U-354 | A window in another cluster does not hold this one back | Negative | `TestAWindowInAnotherClusterDoesNotHoldThisOneBack` | +| U-355 | An offline node needs no shutdown, and none is issued | Boundary | `TestAnOfflineNodeNeedsNoShutdown` | +| U-356 | A node mid-restart is waited for rather than shut down into an in-flight restart | Boundary | `TestANodeMidRestartIsWaitedForRatherThanShutDown` | +| U-357 | The manager holds its own eviction on the worker it is draining | Positive | `TestTheManagerHoldsItselfOnTheWorkerItIsDraining` | +| U-358 | The manager holds nothing on somebody else's worker | Negative | `TestTheManagerDoesNotHoldItselfOnSomebodyElsesWorker` | +| U-359 | A manager that does not know where it runs holds nothing | Boundary | `TestAManagerThatDoesNotKnowWhereItRunsHoldsNothing` | +| U-360 | Releasing the manager's own budget is what lets the drain finish | Positive | `TestReleasingTheManagerIsWhatLetsTheDrainFinish` | +| U-361 | The self-budget carries the labels its group is selected by | Positive | `TestTheSelfBudgetCarriesTheGroupsLabels` | +| U-362 | The window holds the manager before anything that can fail | Regression | `TestTheWindowHoldsTheManagerBeforeAnythingThatCanFail` | +| U-363 | The window lets the manager go when it lets the storage pod go | Regression | `TestTheWindowLetsTheManagerGoWhenItLetsTheStoragePodGo` | +| U-364 | A step of another action reached under this one: terminal rather than stalled | Negative | `TestAStepOfAnotherActionEndsTheWindow` | + +### Ops Shape: The Step Machine (design §6.3) + +File: `operator/internal/controllers/node/graphs_test.go` + +Seven graphs are declared as one `statemachine.MultiConfig`, and the provisioning +machine is declared beside them. The three lists that have to agree are the +graph's states, the `Enum` marker on `status.step.state`, and the CEL rule the +CRD carries. + +| # | Scenario | Type | Test | +|-------|------------------------------------------------------------------------------------|----------|-------------------------------------------------------| +| U-217 | Every declared graph builds, including the ones the action under test does not use | Positive | `TestEveryActionDeclaresAGraph` | +| U-218 | An action with no declared graph: refused rather than stalled | Negative | `TestAStepThatBelongsToNoActionEndsTheOperation` | +| U-219 | `Remove` transitioning to `Promoting`: rejected as an illegal transition | Negative | `TestAStepOfAnotherActionIsRejected` | +| U-220 | `Migrate` transitioning to `Removing`: rejected as an illegal transition | Negative | `TestAStepOfAnotherActionIsRejected` | +| U-221 | `HostMaintenance` transitioning to `Suspending`: rejected | Negative | `TestAStepOfAnotherActionIsRejected` | +| U-222 | An empty `status.step`: restores to the action's declared initial state | Boundary | `TestTheFirstPassArmsTheStepAMachineIsBornIn` | +| U-223 | A step value that belongs to a different action: restoration fails informatively | Negative | `TestAStepOfAnotherActionIsRejected` | +| U-224 | A step value outside the enum: restoration fails rather than stalling | Negative | `TestAStepThisOperatorCannotResumeEndsTheOperation` | +| U-225 | The snapshot round-trips through `Snapshot` and `FromSnapshot` unchanged | Positive | — | +| U-226 | A deadline persisted and restored is the same absolute instant | Positive | — | +| U-227 | A deadline that passed while the operator was down: restores as expired | Boundary | `TestAStepThatOutlivedItsDeadlineFailsTheOperation` | +| U-228 | A step with no deadline: restores with none rather than with a zero instant | Boundary | — | +| U-229 | A terminal step: `IsTerminal` is true and no transition is attempted | Boundary | `TestTheLastStepFinishingEndsTheOperation` | +| U-230 | The outer phase machine is separate from the step machine | Positive | — | +| U-231 | Every state each graph declares appears in the step `Enum` marker | Boundary | `TestTheStepEnumCoversEveryDeclaredState` | +| U-232 | Every state each graph declares appears in the `status.step` CEL rule | Boundary | `TestTheCELRuleCoversEveryDeclaredState` | +| U-233 | The CEL rule names no value the graphs do not declare | Negative | `TestTheCELRuleCoversEveryDeclaredState` | +| U-234 | A stored step from another action: refused, naming the declared set | Negative | `TestAStepOfAnotherActionIsRejected` | +| U-235 | A restore that fails: the operation is `Failed` with the error, not requeued | Negative | `TestAStepThisOperatorCannotResumeEndsTheOperation` | +| U-365 | The provisioning graph's states cover its own `Enum` and CEL rule | Boundary | `TestTheNodeStepEnumAndRuleCoverTheProvisioningGraph` | +| U-366 | No step past the point of no return declares an abort edge | Negative | `TestNoStepPastThePointOfNoReturnIsAbortable` | +| U-367 | The steps an abort stops from are the ones that unwind cleanly | Positive | `TestTheStepsAnAbortStopsCleanly` | +| U-368 | Every drain step past the suspend owes the resume | Boundary | `TestTheDrainStepsPastTheSuspendUnwind` | +| U-369 | Every step carries a budget, so none is the step that cannot time out | Boundary | `TestEveryStepHasABudget` | +| U-370 | The remove graph validates before it suspends | Positive | `TestTheRemoveGraphValidatesBeforeItSuspends` | +| U-371 | The migrate graph splits the restart from the wait | Positive | `TestTheMigrateGraphSplitsTheRestartFromTheWait` | +| U-372 | The host maintenance graph is the six-step window | Positive | `TestTheHostMaintenanceGraphIsTheSixStepWindow` | +| U-373 | The four single-step actions share one line | Positive | `TestTheSingleStepActionsShareOneLine` | +| U-374 | Adoption is reachable from both provisioning gates | Boundary | `TestAdoptionIsReachableFromBothGates` | +| U-375 | `AwaitingSlot` may go straight to `Resolving` when a sibling claimed the worker | Boundary | `TestAwaitingSlotMayGoStraightToResolving` | + +### Entity: Pod Placement (design §13.1) + +File: `operator/internal/controllers/node/podscheduling_test.go` + +A pod the scheduler refused is the one Kubernetes-side failure the control plane +cannot see, and the scheduler has already written down why. + +| # | Scenario | Type | Test | +|-------|-----------------------------------------------------------|----------|--------------------------------------------| +| U-376 | An unplaceable pod is announced on the node it belongs to | Positive | `TestAnUnplaceablePodIsAnnouncedOnItsNode` | +| U-377 | A placed pod is announced nowhere | Negative | `TestAPlacedPodIsAnnouncedNowhere` | +| U-378 | Another worker's scheduling failure is not this node's | Negative | `TestAnotherWorkersFailureIsNotThisNodes` | +| U-379 | A gated pod is a wait rather than a failure | Boundary | `TestAGatedPodIsAWaitRatherThanAFailure` | + +### Workload: Device Partitioning (design §5.3) + +File: `operator/internal/controllers/node/partitions_test.go` + +| # | Scenario | Type | Test | +|-------|----------------------------------------------------------------------|------------|-------------------------------------------------------| +| U-380 | The partition count matches what the deployed fleets were built with | Regression | `TestThePartitionCountMatchesWhatFleetsWereBuiltWith` | + +### Node Streams (design §4.4) + +File: `operator/internal/controllers/node/unregister_test.go` + +| # | Scenario | Type | Test | +|-------|---------------------------------------------------------------------------|----------|----------------------------------------------| +| U-381 | A node whose object goes closes its device stream, with the cluster known | Positive | `TestTheDeviceStreamClosesWithTheCluster` | +| U-382 | The same, with the cluster already gone | Boundary | `TestTheDeviceStreamClosesWithoutTheCluster` | +| U-383 | A node that never resolved a UUID closes nothing | Boundary | `TestANodeWithNoIdClosesNothing` | ## 2. Integration Tests -Full reconcile loop against a real Kubernetes API server via `envtest`, driven by -`TestControllers` in `operator/internal/controllers/node/suite_test.go`, with the +Full reconcile loop against a real Kubernetes API server via `envtest`, with the control plane still mocked. These cover what a fake client cannot: real admission, real `resourceVersion` semantics, and real watch delivery. +**This package has no `envtest` suite.** The one this class was written against +lived beside the retired reconcilers in `operator/internal/controller/` and went +with them, so every row below is uncovered and §6 carries the class as one gap +rather than as thirty. + ### Admission and Validation (design §3.1, §3.2, §3.4, §6.1) | # | Scenario | Type | Test | @@ -500,21 +690,21 @@ admission, real `resourceVersion` semantics, and real watch delivery. ### Controller Behavior Under a Real API Server (design §4, §7, §11) -| # | Scenario | Type | Test | -|------|---------------------------------------------------------------------------------|----------|-------------------| -| I-16 | Reconciling a not-found resource returns no requeue | Negative | `TestControllers` | -| I-17 | Two operations for one node: the second stays `Pending`, then runs | Positive | — | -| I-18 | The lock is released: the queued operation wakes from the node watch | Positive | — | -| I-19 | `kubectl delete` on a `Running` operation: `activeOpsRef` is cleared | Positive | — | -| I-20 | Operations on two nodes of one cluster run concurrently without interference | Positive | — | -| I-21 | Operations on nodes of two clusters run concurrently without interference | Positive | — | -| I-22 | An operation targeting a node with no `status.uuid`: fails informatively | Negative | — | -| I-23 | An operation in namespace A cannot lock a same-named node in namespace B | Negative | — | -| I-24 | Deleting the `StorageCluster` cascades to its nodes and to the workload objects | Positive | — | -| I-25 | Two reconcilers racing the `Posting` claim: exactly one `POST` is issued | Negative | — | -| I-26 | Two nodes with the same name in two namespaces: both workloads are independent | Positive | — | -| I-27 | The namespace is deleted mid-drain: the operation terminates and releases | Negative | — | -| I-28 | The controller's role covers every object the workload reconcile touches | Positive | — | +| # | Scenario | Type | Test | +|------|---------------------------------------------------------------------------------|----------|------| +| I-16 | Reconciling a not-found resource returns no requeue | Negative | — | +| I-17 | Two operations for one node: the second stays `Pending`, then runs | Positive | — | +| I-18 | The lock is released: the queued operation wakes from the node watch | Positive | — | +| I-19 | `kubectl delete` on a `Running` operation: `activeOpsRef` is cleared | Positive | — | +| I-20 | Operations on two nodes of one cluster run concurrently without interference | Positive | — | +| I-21 | Operations on nodes of two clusters run concurrently without interference | Positive | — | +| I-22 | An operation targeting a node with no `status.uuid`: fails informatively | Negative | — | +| I-23 | An operation in namespace A cannot lock a same-named node in namespace B | Negative | — | +| I-24 | Deleting the `StorageCluster` cascades to its nodes and to the workload objects | Positive | — | +| I-25 | Two reconcilers racing the `Posting` claim: exactly one `POST` is issued | Negative | — | +| I-26 | Two nodes with the same name in two namespaces: both workloads are independent | Positive | — | +| I-27 | The namespace is deleted mid-drain: the operation terminates and releases | Negative | — | +| I-28 | The controller's role covers every object the workload reconcile touches | Positive | — | --- @@ -666,68 +856,95 @@ eviction, the kubelet, and the reboot. 5. Run fio throughout and confirm no I/O errors and matching checksums. 6. Confirm no budget and no drain label survives the run. +--- + + --- ## 5. Coverage Summary | Class | Scenarios | Covered | Not covered | |-------------|-----------|---------|-------------| -| Unit | 263 | 87 | 176 | -| Integration | 54 | 1 | 53 | +| Unit | 375 | 270 | 105 | +| Integration | 54 | 0 | 54 | | E2E | 26 | 0 | 26 | | Manual | 5 | 0 | 5 | -| **Total** | **348** | **88** | **260** | +| **Total** | **460** | **270** | **190** | -Eighty-seven of the eighty-eight covered scenarios are unit tests, and they -concentrate in three places: the workload builders, the drain's volume -classification, and the operation lock. That is the right distribution, because a -defect in any of the three corrupts state rather than failing safely. Every -covered scenario outside a unit test is `I-16`. +Eight further scenarios are struck through: they describe a system this one no +longer is, and each names the row that replaced it. They are excluded from the +counts. -Eighty-four distinct test functions cover those eighty-eight scenarios, because a -table-driven test satisfies one identifier per subtest. +Two hundred and seven distinct test functions cover the two hundred and seventy +covered scenarios, because a table-driven test satisfies one identifier per +subtest and several rows are two halves of one assertion. -The ratio is what a target-model plan looks like against a registered API. Most -uncovered rows are not gaps in testing so much as scenarios for behavior that does -not exist yet, and §6 separates the two. +Every covered scenario is a unit test, and the concentration is the shape of the +package rather than an accident of effort: both reconcilers are written so that +every decision is a predicate over state a fake client and a scripted control +plane can present, which is what makes the operation machine testable without a +cluster. What that buys is also its limit. Admission, `resourceVersion` +semantics, watch delivery, and garbage collection are the API server's, and no +row that depends on one of them is covered. --- ## 6. What Is Not Yet Covered -| # | Gap | Reason | -|----------------------------|-----------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| U-07, U-08 | Name collision retry, and the name carrying no worker | The retry path exists and is unexercised. The negative assertion about the name has never been written down | -| U-12 | A failure-domain label of `0` read as set rather than unset | The label form retires the zero-versus-unset ambiguity the integer had, since an unset label is the empty string, and this row is what keeps a label spelled `0` from reintroducing it | -| U-23, U-25 … U-27 | The parallel-add hold and the FoundationDB serialization rules | `isFDBWorkerBlocked` has no test reference, and it is the rule that keeps a control plane from losing quorum during an expansion | -| U-28 … U-35 | The whole `Posting` claim | Planned, not built. The guard today is `status.postedAt` plus a `List`, which the optimistic-lock claim of design §4.2 replaces | -| U-38 … U-40, U-43 | Positional UUID resolution for multi-socket workers | `pollUUIDFromBackend` has no test reference, and the RPC-port ordering it depends on is an assumption nothing asserts | -| U-44 … U-47 | Every adoption route | No test reference at all, and adoption is the migration path off Helm and the path every pre-existing cluster takes | -| U-48 … U-52, U-236 … U-238 | Steady-state sync, and the device counts | `syncStatus` has no test reference. The early return is what keeps an idle cluster from writing, and the device rows pin an ordering the string form got backward (design §15.1) | -| U-249 … U-254 | The node's capacity sample and its write threshold | Planned, not built. Nothing samples a node's occupancy today, and the threshold is what keeps a node that is serving I/O from reconciling itself (design §3.3) | -| U-56 … U-59 | Deletion while an operation runs, and the non-online cases | The finalizer hold has no test, and it is what stops a delete from orphaning a backend node | -| U-64 | A look-alike service account in another namespace | The prefix match is string-based, and nothing asserts it cannot be spoofed by a namespace name | -| U-75, U-76 | Workload ownership by the `StorageCluster` | Planned, not built. The objects are owned by `StorageNodeSet` today, and the reparent is design §15.3 | -| U-78 … U-84 | The per-slot storage-node-uuid labels | The most consequential labels in the operator, and the only covered part is the per-set label. A wrong key here breaks CSI provisioning | -| U-90 … U-94, U-262, U-263 | Per-node divergence in the ConfigMap | Only the cluster-uniform half is covered. The shell-quoting row matters because device names reach a sourced file, and the two sizing rows keep the node's own values out of the cluster's key | -| U-244 … U-247, I-40 … I-45 | The widened `deviceNames` | Design §3.1 has the field take a PCI address and a device path in one list. Nothing implements either form yet, and `I-44` and `I-45` are what hold the item pattern | -| U-255 … U-261 | The device class the cluster fixes | Design §3.4 has the webhook compare every entry against `StorageCluster.spec.deviceClass` and refuse the PCI filters on a block cluster. Neither the field nor the comparison exists, and `U-258` is the row that keeps a mixed list out of a cluster of either class | -| U-116 … U-124 | Terminal re-reconcile, the acquire race, and the cluster gate | The lock's happy paths are covered and its concurrent ones are not. The cluster gate applies to one action today, design §7.1 widens it | -| U-125 … U-134 | Every single-step action | `runSimpleAction` has no test reference, and its skip-if-already-there behavior is what design §7.2 replaces `status.triggered` with | -| U-144 … U-146 | Classification boundaries | The buckets are covered individually and their overlaps are not | -| U-147 … U-169 | The remove graph past classification | Only the helpers are covered. Every step, every hold, and the whole resume path are untested | -| U-172 … U-177, U-248 | `PersistentVolumeOps` lifecycle, and the reference and label that replace the owner reference | The name generation is covered and the lifecycle is not | -| U-184 … U-200 | The migrate graph | Only the configuration merge is covered. The DNS gate, the two-step restart observation, and the abort refusal are untested | -| U-201 … U-216 | The whole host maintenance action | Planned, not built. It is a separate controller driving `StorageNodeSet.status.drainCoordination` today, which design §10 replaces | -| U-217 … U-235 | The step machine, and the three lists that have to agree | Planned, not built. `atlas-lib/statemachine` has no consumer in either component yet. `U-231` to `U-233` are what the shared `KubeSnapshot` makes necessary: the step values live in the graph, in the `Enum` marker, and in the CEL rule, and only a test keeps the three level | -| I-01 … I-15, I-46 … I-54 | Every admission rule | Needs `envtest`, because CEL, `Required`, and defaulting are enforced by the API server and a fake client applies none of them | -| I-17 … I-28 | Lock behavior, cascade, and isolation under a real API server | Needs `envtest` for real `resourceVersion` conflicts, real watch delivery, and real garbage collection | -| I-37 … I-39 | The deployment config's ephemerality | Nothing to test yet: `ClusterDeploymentConfig` does not exist, and the registered model rewrites `spec.overrides` from the set on every reconcile, which is the behavior design §3.1 replaces | -| E-01 … E-26 | All end-to-end scenarios | Needs a live cluster. The e2e harness under `test/` is not committed yet | -| M-01 … M-05 | Crash consistency, host death, degraded clusters, and the roll | Need process kills, host power control, and a real OS upgrade | -| Metrics | The eleven metrics of design §13.2 | Designed, not built. Nothing exports a metric for either kind today | -| Events | The twenty reasons of design §13.1 | Fourteen reasons exist under different names, and no test asserts any of them | -| Retention | Nothing deletes a terminal `StorageNodeOps` | Feature does not exist. Design §16, Q6 | +| # | Gap | Reason | +|------------------------------------------|--------------------------------------------------------------------------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| U-12 | A failure-domain label of `0` read as set rather than unset | The label form retires the zero-versus-unset ambiguity the integer had, and this row is what keeps a label spelled `0` from reintroducing it | +| U-14 … U-17 | The worker's storage-node API probe | `HostAnswers` makes a real HTTP request, so a unit test would either reach the network or assert a stub. It needs `envtest` with a served endpoint, or an e2e row | +| U-20, U-21 | Counting in-flight workers rather than objects | The cap is covered by worker. That two sockets of one host count once is the multi-socket case, and the slot suite uses one node per worker | +| U-24 … U-27 | The FoundationDB serialization rules | `hostsFoundationDB` has no test, and it is the rule that keeps the control plane's own store from losing quorum during an expansion | +| U-28 … U-34 | The optimistic-lock claim, and the `POST`'s failure modes | The claim is structural: it is the transition into `Posting`, and nothing asserts the ordering or the 409. The control-plane failure paths need a scripted client that fails | +| U-36, U-37 | The worker's internal address | `workerAddress` is exercised through adoption and asserted nowhere | +| U-39, U-40, U-42, U-43 | Positional resolution for multi-socket workers | The RPC-port ordering the slot match depends on is an assumption nothing asserts | +| U-47 | What adoption writes | The divert to `Adopting` is covered and `resolveUUID`'s write is not | +| U-50 … U-52 | The assigned fault group, a malformed reading, and `observedGeneration` | `applyReading` writes all three and no row asserts them | +| U-55, U-57, U-58 | Deletion's remaining cases | `U-57` is the open decision below rather than an unwritten test | +| U-64 | A look-alike service account in another namespace | The identity match is string-based, and nothing asserts it cannot be spoofed by a namespace name | +| U-67 … U-69, U-71, U-73 … U-76 | The workload's TLS mounts, its RBAC, and its ownership | The DaemonSet's shape is covered by the builder's own tests, and what the reconcile does with it is covered only for the image and the certificate revision | +| U-78 … U-84 | The per-slot storage-node-uuid labels | The most consequential labels in the operator, and only the cluster-wide one is covered. A wrong key here breaks CSI provisioning | +| U-85 … U-94, U-262, U-263 | The per-node configuration's contents | The clone is covered and the render is not: the assertions that named every key died with the retired set, and `ReconcileConfig` is exercised only through enrollment | +| U-99 … U-101 | The SPDK proxy EndpointSlices | The builders are covered in `internal/utils` and the reconcile that applies them is not | +| U-104 … U-107 | The TLS environment, the Secret predicate, and the cert-manager provider | `reconcileCertificates` is the least covered path in the workload reconcile | +| U-118, U-120, U-121, U-123 | The lock's concurrent cases, and the inactive cluster | Two reconcilers racing one free lock needs real `resourceVersion` semantics. The rebalancing half of the gate is covered and the inactive half is not | +| U-130 … U-132, U-134 | The single-step actions' control-plane failures | The scripted control plane can refuse a call, and no row yet distinguishes a 4xx from a 5xx or asserts what reaches the event | +| U-144 … U-146 | Classification boundaries | The buckets are covered individually and their overlaps are not. `U-146` records that an empty pattern falls back to the default, which is the shipped behavior and was written as its opposite | +| U-150, U-155, U-159, U-162, U-164, U-165 | The drain's remaining edges | `U-162` is written against a behavior the code does not have: a refused removal is terminal whatever its status, so either the row or `drainRemove` has to move | +| U-171, U-177 | Migration naming under collision and across namespaces | The formula is atlas-lib's and is tested there. What is missing is the assertion that this caller's inputs cannot collide | +| U-196, U-199, U-200 | The relocation's abort, and the source worker's labels | `ReleaseWorker` decides whether a worker keeps its labels and no row exercises the two-socket case | +| U-202, U-209, U-216 | The maintenance window's remaining edges | An uncordon mid-window, the ordering of the release against the node going offline, and the deadline that is a detection mechanism rather than a recovery one | +| U-225, U-226, U-228, U-230 | The snapshot's round trip | Covered by `atlas-lib/statemachine`'s own suite for the type, and not by this package for the values it stores | +| U-236 … U-238 | The device summary | `applyReading` writes the pair and the absent case, and the rows that pin the ordering the string form got backward are unwritten | +| U-244 … U-247 | The widened `deviceNames` | The field takes a PCI address and a device path in one list, and nothing asserts either form reaches the node | +| U-264 … U-266 | The image fallback, and the name's determinism | The fallback to the singleton `ControlPlane` is exercised only through its refusal, and the name formula's determinism is asserted nowhere | +| I-01 … I-54 | Every integration row | This package has no `envtest` suite. CEL, `Required`, defaulting, real conflicts, watch delivery, and garbage collection are all the API server's, and a fake client has none | +| E-01 … E-26 | All end-to-end scenarios | Needs a live cluster. The e2e harness under `test/` does not cover this kind yet | +| M-01 … M-05 | Crash consistency, host death, degraded clusters, and the roll | Need process kills, host power control, and a real OS upgrade | +| Metrics | The eleven metrics of design §13.2 | The series are declared in `metrics.go` and observed by the reconcilers, and no row asserts a label set or a value | +| Events | The twenty reasons of design §13.1 | Nine reasons are asserted through the rows above. The rest are raised and unasserted | +| Retention | Nothing deletes a terminal `StorageNodeOps` | Feature does not exist. Design §16, Q6 | + +### Two decisions this plan is holding + +**`U-57`, a failed removal and the finalizer.** The row states that a `Failed` +removal holds the node's finalizer, which is what the deleted +`TestHandleDeletion_FailedRemoveOpsBlocksFinalizerRemoval` asserted after the +2026-08-13 incident: `Failed` means the backend node was never removed, so +letting the object go orphans a live node. The shipped teardown releases on any +terminal phase and argues for it in prose, on the grounds that an object nobody +can delete is worse. Both positions are defensible and they contradict, so the +row stays uncovered until one of them is written down as the rule. + +**The failure-domain balance gate.** The removal's own pre-check, +`fdRemovalBalanceCheck`, went with the rework and has no equivalent: balance now +rests entirely on the control plane's `check_fd_admission_for_remove`, which +`drainRemove` reads as a terminal refusal. No row in this plan describes an +operator-side check, because the operator no longer makes one. `M-03` is the +scenario that would find out whether the control plane's refusal arrives early +enough to matter. ### Axis coverage @@ -737,8 +954,8 @@ combination nothing exercises. | Axis | Value | Scenarios | |---------------------------|----------------------------|-----------------------------------------------------------------| | Cluster node count | Single node | U-155, E-24 | -| | Two nodes | U-155 | -| | Three nodes | E-13, E-20, E-23, M-02, M-03 | +| | Two nodes | U-155, U-337 | +| | Three nodes | U-152, U-339, E-13, E-20, E-23, M-02, M-03 | | | Five or more | E-02, E-25, M-05 | | | Asymmetric node sizes | — | | Sockets per worker | One | Every scenario except those below | @@ -746,37 +963,42 @@ combination nothing exercises. | Namespace count | Single namespace | Every scenario except those below | | | Multiple namespaces | U-72, U-145, U-177, I-23, I-26 | | simplyblock cluster count | One cluster | Every scenario except those below | -| | Several in one Kubernetes | U-81, I-21 | +| | Several in one Kubernetes | U-81, U-328, U-354, I-21 | | | Cross-cluster | — (not applicable: no node operation spans Kubernetes clusters) | | Failure domains | Not enabled | U-11 | -| | Enabled and set | U-10 | +| | Enabled and set | U-10, U-270 | | | Enabled and unset | U-09, U-13 | | | Partially set | — | -| Object scale | Zero volumes | U-139, U-140, U-169, E-11 | -| | A handful | U-152, E-10, E-13 | +| Object scale | Zero volumes | U-139, U-169, U-336, E-11 | +| | A handful | U-152, U-335, E-10, E-13 | | | 100 or more | E-25 | -| Lifecycle and restart | Mid-step restart | U-172, M-01 | -| | Terminal re-reconcile | U-117 | -| | Deletion mid-operation | U-119, I-19, I-27 | -| | Host death mid-operation | M-02 | +| Lifecycle and restart | Mid-step restart | U-172, U-195, M-01 | +| | Terminal re-reconcile | U-117, U-315 | +| | Deletion mid-operation | U-119, U-344, U-345, I-19, I-27 | +| | Host death mid-operation | U-332, M-02 | +| Reading source | The stream, once delivered | U-276, U-340 | +| | The poll, until it is | U-277, U-341 | +| | A node neither one has | U-280 | | Device class | NVMe, matching | U-256, U-261 | | | Block, matching | U-259 | | | An entry of the other one | U-255, U-257 | | | A list holding both | U-258 | | | A filter for the other one | U-260 | +| | Unstated | U-286 | | Capacity sampling | A first reading | U-249 | | | Below the write threshold | U-250 | | | Above it, or a new total | U-251, U-252 | | | Unreachable or unmeasured | U-253, U-254 | -| Actor | Operator-raised operation | U-54, U-201 | -| | User-created operation | Every operation scenario except those two | -| | Webhook path | U-60 … U-64, I-08, I-09 | +| Actor | Operator-raised operation | U-54, U-201, U-284 | +| | User-created operation | Every operation scenario except those three | +| | Webhook path | U-60 … U-64, U-288 … U-296, I-08, I-09 | **The asymmetric-node row is the significant blank.** Every drain and every migration picks a target from the online peers by round-robin, which spreads by count rather than by capacity, so a cluster whose nodes differ in size -concentrates on the smallest one exactly as readily as on the largest. Nothing -here exercises that, and design §8.2 does not claim otherwise. +concentrates on the smallest one exactly as readily as on the largest. `U-339` +pins the spread as deliberate and nothing exercises what it costs, and design +§8.2 does not claim otherwise. **The partially set failure-domain row is the second.** A cluster with `enableFailureDomains` set where some nodes declare a group and others do not is @@ -786,5 +1008,6 @@ per-node, so the cluster runs in a mixed state nothing reports on. **The multi-namespace rows are covered thinly and matter more than the count suggests.** Two namespaces are two independent deployments, and the workload objects are namespaced while the storage-plane labels and the ClusterRoleBindings -are not. `U-72` covers the binding names and `U-81` covers the label keys, and -nothing covers what happens when two namespaces claim the same worker. +are not. `U-72` covers the binding names, `U-354` covers a maintenance window not +reaching across a cluster boundary, and nothing covers what happens when two +namespaces claim the same worker. diff --git a/operator/internal/controllers/node/actions_test.go b/operator/internal/controllers/node/actions_test.go new file mode 100644 index 000000000..2673e6a75 --- /dev/null +++ b/operator/internal/controllers/node/actions_test.go @@ -0,0 +1,219 @@ +// The four actions that are one call and one wait. +// +// Shutdown, Restart, Suspend, and Resume each issue a single request and then +// watch for the state it produces. Both halves are written against current state +// rather than against a transition: the call is skipped when the node is already +// at or past where that call would put it, and the wait is a predicate over what +// the control plane reports now. +// +// That is what makes a step recorded without its side effect having fired safe to +// re-enter, which is the whole reason this operator keeps no `triggered` flag +// beside the step. A suspend that was issued and whose response was lost is +// re-entered, reads a suspended node, and advances without a second suspend. +// +// design-storagenode.md §7.2 and §7.3. + +package node + +import ( + "context" + "errors" + "testing" + + "github.com/simplyblock/atlas/ptr" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// requested runs the Requesting step of one action against a node the control +// plane reports in the given state, and reports what was asked of it. +func requested( + t *testing.T, + action simplyblockv1alpha2.StorageNodeOpsAction, + status string, +) (*scriptedControlPlane, bool) { + t.Helper() + api := aControlPlane().reporting(status) + r, _ := anOpsWorld(t, api) + + done, err := r.perform(context.Background(), + anOperation("an-operation", action), stepRequesting) + if err != nil { + t.Fatalf("the %s request: %v", action, err) + } + return api, done +} + +// A node that is already where the call would put it receives no call, and the +// step is finished. One call per action, and the state each skips on is the state +// that call produces. +func TestACallIsSkippedWhenTheNodeIsAlreadyThere(t *testing.T) { + cases := []struct { + action simplyblockv1alpha2.StorageNodeOpsAction + status string + call string + }{ + {simplyblockv1alpha2.StorageNodeOpsActionShutdown, nodeStatusOffline, "ShutdownNode"}, + {simplyblockv1alpha2.StorageNodeOpsActionSuspend, nodeStatusSuspended, "Suspend"}, + {simplyblockv1alpha2.StorageNodeOpsActionResume, nodeStatusOnline, "Resume"}, + } + for _, c := range cases { + t.Run(string(c.action), func(t *testing.T) { + api, done := requested(t, c.action, c.status) + + if asked := api.asked(c.call); asked != 0 { + t.Errorf("%s was issued %d time(s) against a node already %s", + c.call, asked, c.status) + } + if !done { + t.Error("the step did not finish against a node already where it was going") + } + }) + } +} + +// A node that is not there yet receives the call. +func TestACallIsIssuedWhenTheNodeIsNotThereYet(t *testing.T) { + cases := []struct { + action simplyblockv1alpha2.StorageNodeOpsAction + status string + call string + }{ + {simplyblockv1alpha2.StorageNodeOpsActionShutdown, nodeStatusOnline, "ShutdownNode"}, + {simplyblockv1alpha2.StorageNodeOpsActionSuspend, nodeStatusOnline, "Suspend"}, + {simplyblockv1alpha2.StorageNodeOpsActionResume, nodeStatusSuspended, "Resume"}, + } + for _, c := range cases { + t.Run(string(c.action), func(t *testing.T) { + api, _ := requested(t, c.action, c.status) + + if asked := api.asked(c.call); asked != 1 { + t.Errorf("%s was issued %d time(s), want once", c.call, asked) + } + }) + } +} + +// A restart has no state of its own to skip on: a node is online before one and +// online after it. What guards the second call is the persisted step and the wait +// that follows, which does not finish until the node is back. +func TestARestartIsIssuedAgainstAnOnlineNode(t *testing.T) { + api, _ := requested(t, simplyblockv1alpha2.StorageNodeOpsActionRestart, nodeStatusOnline) + + if asked := api.asked("RestartNode"); asked != 1 { + t.Errorf("RestartNode was issued %d time(s), want once", asked) + } +} + +// The two modifiers travel only when the operation states them, because the +// control plane defaults them itself and not sending one is not the same as +// sending false. +func TestOnlyTheFlagsTheOperationStatesAreSent(t *testing.T) { + api := aControlPlane() + r, _ := anOpsWorld(t, api) + + ops := anOperation("a-restart", simplyblockv1alpha2.StorageNodeOpsActionRestart) + if _, err := r.perform(context.Background(), ops, stepRequesting); err != nil { + t.Fatalf("the restart request: %v", err) + } + if len(api.restarts) != 1 { + t.Fatalf("%d restarts were issued, want one", len(api.restarts)) + } + if api.restarts[0].Force || api.restarts[0].ReattachVolume { + t.Errorf("params = %+v, want both flags left to the control plane's own defaults", + api.restarts[0]) + } + + stated := aControlPlane() + r, _ = anOpsWorld(t, stated) + ops = anOperation("a-forced-restart", simplyblockv1alpha2.StorageNodeOpsActionRestart) + ops.Spec.Force = ptr.To(true) + ops.Spec.ReattachVolume = ptr.To(true) + if _, err := r.perform(context.Background(), ops, stepRequesting); err != nil { + t.Fatalf("the forced restart request: %v", err) + } + if !stated.restarts[0].Force || !stated.restarts[0].ReattachVolume { + t.Errorf("params = %+v, want what the operation stated", stated.restarts[0]) + } +} + +// The wait is over when the node reports the state the action was for, and not +// before. +func TestTheWaitIsOverWhenTheNodeReportsWhatTheActionWasFor(t *testing.T) { + cases := []struct { + action simplyblockv1alpha2.StorageNodeOpsAction + wanted string + other string + }{ + {simplyblockv1alpha2.StorageNodeOpsActionShutdown, nodeStatusOffline, nodeStatusOnline}, + {simplyblockv1alpha2.StorageNodeOpsActionRestart, nodeStatusOnline, nodeStatusInRestart}, + {simplyblockv1alpha2.StorageNodeOpsActionSuspend, nodeStatusSuspended, nodeStatusOnline}, + {simplyblockv1alpha2.StorageNodeOpsActionResume, nodeStatusOnline, nodeStatusSuspended}, + } + for _, c := range cases { + t.Run(string(c.action), func(t *testing.T) { + r, _ := anOpsWorld(t, aControlPlane().reporting(c.other)) + ops := anOperation("an-operation", c.action) + + done, err := r.perform(context.Background(), ops, stepAwaiting) + if err != nil { + t.Fatalf("the wait: %v", err) + } + if done { + t.Errorf("the wait finished against a node reporting %s", c.other) + } + + r, _ = anOpsWorld(t, aControlPlane().reporting(c.wanted)) + done, err = r.perform(context.Background(), ops, stepAwaiting) + if err != nil { + t.Fatalf("the wait: %v", err) + } + if !done { + t.Errorf("the wait did not finish against a node reporting %s", c.wanted) + } + }) + } +} + +// An action that runs a graph of its own does not issue a single request, and a +// step reached under one that does is a hand-edited object or a downgrade. +// Neither resolves by reconciling again, so both are terminal. +func TestAStepThatBelongsToNoActionEndsTheOperation(t *testing.T) { + r, _ := anOpsWorld(t, aControlPlane()) + + _, err := r.perform(context.Background(), + anOperation("a-drain", simplyblockv1alpha2.StorageNodeOpsActionRemove), stepRequesting) + + var fatal *terminalStepError + if !errors.As(err, &fatal) { + t.Errorf("err = %v, want the terminal kind for an action that issues no single request", err) + } + + _, err = r.perform(context.Background(), + anOperation("an-operation", simplyblockv1alpha2.StorageNodeOpsActionSuspend), step("Nowhere")) + if !errors.As(err, &fatal) { + t.Errorf("err = %v, want the terminal kind for a step no action declares", err) + } +} + +// A node the object has no UUID for has not been provisioned, which no number of +// passes will change. +func TestAnUnprovisionedNodeEndsTheOperation(t *testing.T) { + node := anOpsNode() + node.Status.UUID = "" + r, apiClient := anOpsWorld(t, aControlPlane()) + if err := apiClient.Delete(context.Background(), anOpsNode()); err != nil { + t.Fatalf("clearing the provisioned node: %v", err) + } + if err := apiClient.Create(context.Background(), node); err != nil { + t.Fatalf("seeding the unprovisioned node: %v", err) + } + + _, err := r.perform(context.Background(), + anOperation("an-operation", simplyblockv1alpha2.StorageNodeOpsActionSuspend), stepRequesting) + + var fatal *terminalStepError + if !errors.As(err, &fatal) { + t.Errorf("err = %v, want the terminal kind for a node with no backend behind it", err) + } +} diff --git a/operator/internal/controllers/node/advance_test.go b/operator/internal/controllers/node/advance_test.go new file mode 100644 index 000000000..f64481c79 --- /dev/null +++ b/operator/internal/controllers/node/advance_test.go @@ -0,0 +1,310 @@ +// The spine every action runs on: one step per pass, and the four ways out. +// +// Nothing here blocks. A pass asks whether the current step has finished and +// either requeues or writes the next step down, so a drain that takes an hour +// costs the controller nothing while it waits. The step is written before the +// side effect it performs, which is what makes a process dying between the two +// safe: every completion condition is a predicate over current state and every +// call is skipped when the node is already where it would put it. +// +// The four ways out are finishing, failing, being aborted where the graph allows +// it, and holding — and the last two are the ones a reader cannot tell from a +// stalled controller without an event, which is what the reasons in events.go +// exist for. +// +// design-storagenode.md §7. + +package node + +import ( + "context" + "testing" + "time" + + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + + "github.com/simplyblock/atlas/statemachine" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/cpinformer/subscriptions" + "github.com/simplyblock/simplyblock-operator/internal/utils" +) + +// deliveredCluster is the cluster stream's cache, holding one cluster. +type deliveredCluster struct { + reading subscriptions.ClusterDTO + synced bool +} + +func (d *deliveredCluster) Lookup(string) (subscriptions.ClusterDTO, bool) { + return d.reading, d.reading.ID != "" +} + +func (d *deliveredCluster) SyncedRoot() bool { return d.synced } + +// anAdvancingOperation is an operation at the given step, with a deadline that +// has not passed. +func anAdvancingOperation( + name string, action simplyblockv1alpha2.StorageNodeOpsAction, at step, +) *simplyblockv1alpha2.StorageNodeOps { + ops := anOperation(name, action) + ops.Finalizers = []string{OpsFinalizer} + ops.Status.Phase = simplyblockv1alpha2.StorageNodeOpsPhaseRunning + deadline := metav1.NewTime(time.Now().Add(time.Hour)) + ops.Status.Step = statemachine.KubeSnapshot{State: string(at), Deadline: &deadline} + return ops +} + +// pass runs one reconcile of the operation. +func pass(t *testing.T, r *StorageNodeOpsReconciler, name string) { + t.Helper() + _, err := r.Reconcile(context.Background(), ctrlRequest(name)) + if err != nil { + t.Fatalf("reconciling: %v", err) + } +} + +// The first pass of an admitted operation puts it in its initial step with a +// deadline, and says so. A step with no deadline is a step nothing can ever time +// out, and the initial one is the step a machine is born in — so it is the one +// whose deadline nothing else would set. +func TestTheFirstPassArmsTheStepAMachineIsBornIn(t *testing.T) { + ops := anOperation("a-suspend", simplyblockv1alpha2.StorageNodeOpsActionSuspend) + ops.Finalizers = []string{OpsFinalizer} + r, apiClient := anOpsWorld(t, aControlPlane(), ops) + + pass(t, r, "a-suspend") + + got := operationRead(t, apiClient, "a-suspend") + if got.Status.Step.State != string(stepRequesting) { + t.Errorf("step = %q, want the step a suspend starts in", got.Status.Step.State) + } + if got.Status.Step.Deadline == nil { + t.Error("the first step carries no deadline, so it is the one step that cannot time out") + } + if got.Status.Phase != simplyblockv1alpha2.StorageNodeOpsPhaseRunning { + t.Errorf("phase = %q, want Running", got.Status.Phase) + } + if !announcedReason(r, OperationStarted) { + t.Error("nothing announced that the operation had taken the node and started") + } +} + +// A step that has finished moves the operation to the next one, and the next +// step's deadline travels with it. +func TestAFinishedStepEntersTheNextWithItsOwnDeadline(t *testing.T) { + // A suspend whose node is already suspended finishes Requesting at once. + api := aControlPlane().reporting(nodeStatusSuspended) + ops := anAdvancingOperation("a-suspend", + simplyblockv1alpha2.StorageNodeOpsActionSuspend, stepRequesting) + r, apiClient := anOpsWorld(t, api, ops) + lockedBy(t, apiClient, "a-suspend") + + pass(t, r, "a-suspend") + + got := operationRead(t, apiClient, "a-suspend") + if got.Status.Step.State != string(stepAwaiting) { + t.Errorf("step = %q, want the wait that follows the request", got.Status.Step.State) + } + if got.Status.Step.Deadline == nil { + t.Error("the step was entered without a deadline") + } +} + +// The last step finishing is the operation finishing, and the node is released. +func TestTheLastStepFinishingEndsTheOperation(t *testing.T) { + api := aControlPlane().reporting(nodeStatusSuspended) + ops := anAdvancingOperation("a-suspend", + simplyblockv1alpha2.StorageNodeOpsActionSuspend, stepAwaiting) + r, apiClient := anOpsWorld(t, api, ops) + lockedBy(t, apiClient, "a-suspend") + + pass(t, r, "a-suspend") + + got := operationRead(t, apiClient, "a-suspend") + if got.Status.Phase != simplyblockv1alpha2.StorageNodeOpsPhaseSucceeded { + t.Errorf("phase = %q, want Succeeded", got.Status.Phase) + } + if holder := lockHolder(t, apiClient); holder != "" { + t.Errorf("the node is still held by %q after the operation finished", holder) + } +} + +// An abort at a step the graph declares abortable stops the operation there, and +// the node is put back into service on the way out: a node past the suspend is +// serving nothing, and leaving it that way takes capacity out of the cluster for +// as long as nobody notices. +func TestAnAbortAtAnAbortableStepStopsAndResumesTheNode(t *testing.T) { + api := aControlPlane().reporting(nodeStatusSuspended) + ops := anAdvancingOperation("a-drain", + simplyblockv1alpha2.StorageNodeOpsActionRemove, stepMigratingVolumes) + ops.Spec.Abort = true + r, apiClient := anOpsWorld(t, api, ops) + r.Mover = &scriptedMover{} + lockedBy(t, apiClient, "a-drain") + + pass(t, r, "a-drain") + + got := operationRead(t, apiClient, "a-drain") + if got.Status.Phase != simplyblockv1alpha2.StorageNodeOpsPhaseAborted { + t.Errorf("phase = %q, want Aborted", got.Status.Phase) + } + if asked := api.asked("Resume"); asked != 1 { + t.Errorf("Resume was issued %d time(s), want the node put back into service", asked) + } + if !announcedReason(r, OperationAborted) { + t.Error("nothing announced the abort") + } +} + +// An abort at a step the control plane is part-way through is refused, and +// refusing is the point: stopping there would leave nothing driving the node +// back to a state somebody can reason about. The operation carries on and says +// so, which is not a failure of it. +func TestAnAbortThatArrivedTooLateIsRefusedAndTheOperationRunsOn(t *testing.T) { + api := aControlPlane() + ops := anAdvancingOperation("a-relocation", + simplyblockv1alpha2.StorageNodeOpsActionMigrate, stepPromoting) + ops.Spec.Abort = true + ops.Spec.Migrate = &simplyblockv1alpha2.MigrateSpec{TargetWorkerNode: opsTarget} + r, apiClient := anOpsWorld(t, api, ops) + lockedBy(t, apiClient, "a-relocation") + + pass(t, r, "a-relocation") + + got := operationRead(t, apiClient, "a-relocation") + if got.Status.Phase != simplyblockv1alpha2.StorageNodeOpsPhaseRunning { + t.Errorf("phase = %q, want the operation still running", got.Status.Phase) + } + if got.Status.Message == "" { + t.Error("nothing says why the abort was not honored") + } + if asked := api.asked("Promote"); asked != 0 { + t.Errorf("Promote was issued %d time(s) on the pass that answered the abort", asked) + } +} + +// A cluster that is mid-rebalance is a cluster whose layout an operation would +// either be refused by or succeed into inconsistently, so the operation holds +// and says which. +func TestAnOperationHoldsWhileItsClusterIsRebalancing(t *testing.T) { + ops := anAdvancingOperation("a-suspend", + simplyblockv1alpha2.StorageNodeOpsActionSuspend, stepRequesting) + api := aControlPlane() + r, apiClient := anOpsWorld(t, api, ops) + r.Clusters = &deliveredCluster{synced: true, reading: subscriptions.ClusterDTO{ + ID: opsClusterID, Status: utils.ClusterStatusActive, Rebalancing: true, + }} + lockedBy(t, apiClient, "a-suspend") + + pass(t, r, "a-suspend") + + if asked := api.asked("Suspend"); asked != 0 { + t.Errorf("Suspend was issued %d time(s) into a rebalancing cluster", asked) + } + if !announcedReason(r, ClusterNotReady) { + t.Error("nothing announced the hold, which is what tells it from a stalled controller") + } + if operationRead(t, apiClient, "a-suspend").Status.Message == "" { + t.Error("the operation says nothing about what it is waiting for") + } +} + +// A removal runs whatever the cluster says about itself, because a node is +// removed from an unready cluster precisely to make the cluster ready. +func TestARemovalRunsAgainstARebalancingCluster(t *testing.T) { + api := aControlPlane() + ops := anAdvancingOperation("a-drain", + simplyblockv1alpha2.StorageNodeOpsActionRemove, stepValidating) + r, apiClient := anOpsWorld(t, api, ops) + r.Mover = &scriptedMover{} + r.Clusters = &deliveredCluster{synced: true, reading: subscriptions.ClusterDTO{ + ID: opsClusterID, Status: utils.ClusterStatusActive, Rebalancing: true, + }} + lockedBy(t, apiClient, "a-drain") + + pass(t, r, "a-drain") + + got := operationRead(t, apiClient, "a-drain") + if got.Status.Step.State != string(stepSuspending) { + t.Errorf("step = %q, want the drain past validation despite the rebalance", + got.Status.Step.State) + } +} + +// A step that outlived its deadline fails the operation rather than retrying +// forever, and the node is resumed where the step it failed on left it +// suspended. +func TestAStepThatOutlivedItsDeadlineFailsTheOperation(t *testing.T) { + api := aControlPlane().reporting(nodeStatusSuspended) + ops := anAdvancingOperation("a-drain", + simplyblockv1alpha2.StorageNodeOpsActionRemove, stepMigratingVolumes) + expired := metav1.NewTime(time.Now().Add(-time.Minute)) + ops.Status.Step.Deadline = &expired + r, apiClient := anOpsWorld(t, api, ops) + r.Mover = &scriptedMover{} + lockedBy(t, apiClient, "a-drain") + + pass(t, r, "a-drain") + + got := operationRead(t, apiClient, "a-drain") + if got.Status.Phase != simplyblockv1alpha2.StorageNodeOpsPhaseFailed { + t.Errorf("phase = %q, want Failed", got.Status.Phase) + } + if !announcedReason(r, StepDeadlineExceeded) { + t.Error("nothing announced the expiry, so a failed operation looks like a slow one") + } + if asked := api.asked("Resume"); asked != 1 { + t.Errorf("Resume was issued %d time(s); a drain that failed past the suspend owes it", + asked) + } + if holder := lockHolder(t, apiClient); holder != "" { + t.Errorf("the node is still held by %q after the operation failed", holder) + } +} + +// A step this operator does not recognize is a downgrade, a hand-edited object, +// or a rename that shipped without a conversion. None of them resolves by +// reconciling again, so the operation is terminal with the reason in its status. +func TestAStepThisOperatorCannotResumeEndsTheOperation(t *testing.T) { + ops := anAdvancingOperation("a-suspend", + simplyblockv1alpha2.StorageNodeOpsActionSuspend, step("SomethingFromAnotherVersion")) + r, apiClient := anOpsWorld(t, aControlPlane(), ops) + lockedBy(t, apiClient, "a-suspend") + + pass(t, r, "a-suspend") + + got := operationRead(t, apiClient, "a-suspend") + if got.Status.Phase != simplyblockv1alpha2.StorageNodeOpsPhaseFailed { + t.Errorf("phase = %q, want Failed", got.Status.Phase) + } + if got.Status.Message == "" { + t.Error("the record says nothing about why the operation could not be resumed") + } +} + +// Observing the spec is what tells a controller that has not looked from one +// that looked and declined, which on this kind is precisely what a user who set +// spec.abort needs to see. +func TestTheObservedGenerationMovesWhenTheOperationIsLookedAt(t *testing.T) { + ops := anAdvancingOperation("a-suspend", + simplyblockv1alpha2.StorageNodeOpsActionSuspend, stepRequesting) + ops.Generation = 3 + r, apiClient := anOpsWorld(t, aControlPlane(), ops) + lockedBy(t, apiClient, "a-suspend") + + pass(t, r, "a-suspend") + + if got := operationRead(t, apiClient, "a-suspend"); got.Status.ObservedGeneration == 0 { + t.Error("nothing records that the operation was looked at") + } +} + +// ctrlRequest names one operation of the fixtures' namespace. +func ctrlRequest(name string) ctrl.Request { + return ctrl.Request{NamespacedName: client.ObjectKey{ + Namespace: opsNamespace, Name: name, + }} +} diff --git a/operator/internal/controllers/node/classify_test.go b/operator/internal/controllers/node/classify_test.go new file mode 100644 index 000000000..565266b78 --- /dev/null +++ b/operator/internal/controllers/node/classify_test.go @@ -0,0 +1,255 @@ +// How a drain sorts the volumes it finds on a node. +// +// The four buckets are what the whole of the Remove action is decided from: +// managed volumes are moved, system volumes are deleted, and a pinned or +// unmanaged one blocks the drain before it has suspended anything. Every one of +// the drain's five steps asks this question and two of them ask nothing else, so +// a volume placed in the wrong bucket is either data moved that somebody pinned +// where it is or data destroyed that nothing in Kubernetes was tracking. +// +// design-storagenode.md §8.1. + +package node + +import ( + "context" + "errors" + "testing" + + "sigs.k8s.io/controller-runtime/pkg/client" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/webapi" +) + +// onNode is one volume the control plane reports as living on the node being +// drained. +func onNode(uuid, name string) webapi.VolumeInfo { + return webapi.VolumeInfo{UUID: uuid, Name: name, PrimaryNodeUUID: opsNodeID} +} + +// censusOf classifies what the given control plane reports, through a drain of +// the node the fixtures describe. +func censusOf( + t *testing.T, api *scriptedControlPlane, objects ...client.Object, +) volumeCensus { + t.Helper() + r, _ := anOpsWorld(t, api, objects...) + census, err := r.classify(context.Background(), + anOperation("a-drain", simplyblockv1alpha2.StorageNodeOpsActionRemove), + opsClusterID, opsNodeID) + if err != nil { + t.Fatalf("classifying the node's volumes: %v", err) + } + return census +} + +// A volume a PersistentVolume accounts for and nothing pins is movable, and it +// is carried with that PersistentVolume's name because the migration is +// addressed by the Kubernetes object rather than by the backend volume. +func TestAVolumeKubernetesAccountsForIsMovable(t *testing.T) { + census := censusOf(t, + aControlPlane().holding(onNode("volume-1", "pvc-abc")), + aPersistentVolume("pv-1", "volume-1"), aClaim("pv-1", false)) + + if len(census.Managed) != 1 { + t.Fatalf("managed = %v, want the one volume a PersistentVolume accounts for", census.Managed) + } + if census.Managed[0].PVName != "pv-1" { + t.Errorf("the movable volume names %q, and the move is addressed by the object", + census.Managed[0].PVName) + } + if len(census.Pinned)+len(census.Unmanaged)+len(census.System) != 0 { + t.Errorf("the volume landed in a second bucket as well: %+v", census) + } +} + +// A pinned claim blocks. Moving it would violate the pin, and the operator does +// not remove somebody's placement decision on their behalf. +func TestAPinnedVolumeBlocksRatherThanMoving(t *testing.T) { + census := censusOf(t, + aControlPlane().holding(onNode("volume-1", "pvc-abc")), + aPersistentVolume("pv-1", "volume-1"), aClaim("pv-1", true)) + + if len(census.Pinned) != 1 || census.Pinned[0] != "volume-1" { + t.Errorf("pinned = %v, want the volume whose claim carries the annotation", census.Pinned) + } + if len(census.Managed) != 0 { + t.Errorf("a pinned volume was also queued to move: %v", census.Managed) + } +} + +// A volume with no PersistentVolume behind it blocks, and blocking is the only +// safe answer: migrating it moves data nothing in Kubernetes tracks, and +// deleting it destroys data nothing in Kubernetes tracks. +func TestAVolumeNothingAccountsForBlocks(t *testing.T) { + census := censusOf(t, aControlPlane().holding(onNode("volume-orphan", "hand-made"))) + + if len(census.Unmanaged) != 1 || census.Unmanaged[0] != "volume-orphan" { + t.Errorf("unmanaged = %v, want the volume no PersistentVolume accounts for", census.Unmanaged) + } +} + +// A PersistentVolume of another driver accounts for nothing here, so the volume +// behind it is as unaccounted for as one with no object at all. +func TestAnotherDriversVolumeAccountsForNothing(t *testing.T) { + foreign := aPersistentVolume("pv-1", "volume-1") + foreign.Spec.CSI.Driver = "ebs.csi.aws.com" + + census := censusOf(t, + aControlPlane().holding(onNode("volume-1", "pvc-abc")), + foreign, aClaim("pv-1", false)) + + if len(census.Unmanaged) != 1 { + t.Errorf("unmanaged = %v, want the volume a foreign driver's object does not account for", + census.Unmanaged) + } +} + +// A PersistentVolume with no claim cannot be pinned, because the annotation +// lives on the claim. It is movable. +func TestAnUnclaimedVolumeIsMovable(t *testing.T) { + unclaimed := aPersistentVolume("pv-1", "volume-1") + unclaimed.Spec.ClaimRef = nil + + census := censusOf(t, + aControlPlane().holding(onNode("volume-1", "pvc-abc")), unclaimed) + + if len(census.Managed) != 1 { + t.Errorf("managed = %v, want the volume no claim can pin", census.Managed) + } +} + +// The rebalancer's benchmark volumes are per-node artifacts, so they are neither +// moved nor counted as blockers: verification deletes them, which is why the +// bucket carries the pool the delete is addressed by. +func TestABenchmarkVolumeIsTheDrainsToDelete(t *testing.T) { + census := censusOf(t, + aControlPlane().holding(onNode("volume-bench", "sb-fio-baseline-read"))) + + if len(census.System) != 1 { + t.Fatalf("system = %v, want the benchmark volume", census.System) + } + if census.System[0].PoolUUID != opsPool { + t.Errorf("the benchmark volume carries pool %q, and a delete is addressed by both", + census.System[0].PoolUUID) + } + if len(census.Managed)+len(census.Unmanaged) != 0 { + t.Errorf("the benchmark volume was also treated as a user's: %+v", census) + } +} + +// A volume whose delete has been accepted is on its way out, and the backend +// keeps reporting it briefly. Counting it would hold a drain open against a +// volume that is already leaving. +func TestAVolumeAlreadyBeingDeletedIsNotCounted(t *testing.T) { + leaving := onNode("volume-1", "pvc-abc") + leaving.Status = volumeStatusInDeletion + + census := censusOf(t, aControlPlane().holding(leaving)) + + if len(census.Managed)+len(census.Pinned)+len(census.Unmanaged)+len(census.System) != 0 { + t.Errorf("a volume in deletion was counted: %+v", census) + } +} + +// Only the node being drained is the drain's business. A peer's volumes are in +// the same pools and must not be. +func TestAPeersVolumesAreNotThisNodesToMove(t *testing.T) { + elsewhere := webapi.VolumeInfo{UUID: "volume-2", Name: "pvc-def", PrimaryNodeUUID: opsPeerID} + + census := censusOf(t, aControlPlane().holding(elsewhere)) + + if len(census.Managed)+len(census.Unmanaged) != 0 { + t.Errorf("another node's volume was counted: %+v", census) + } +} + +// A claim that cannot be read is counted as unmanaged for safety and recorded as +// incomplete, which is the flag that tells the caller this is a transient false +// positive rather than a volume nothing accounts for. An API server that is +// briefly away is not permission to move a volume nobody could classify. +func TestAClaimThatCannotBeReadMakesTheCensusIncomplete(t *testing.T) { + r, _ := anOpsWorldWith(t, aControlPlane().holding(onNode("volume-1", "pvc-abc")), + refusingClaims(), aPersistentVolume("pv-1", "volume-1"), aClaim("pv-1", false)) + + census, err := r.classify(context.Background(), + anOperation("a-drain", simplyblockv1alpha2.StorageNodeOpsActionRemove), + opsClusterID, opsNodeID) + if err != nil { + t.Fatalf("classifying the node's volumes: %v", err) + } + + if !census.Incomplete { + t.Error("the census does not say it is incomplete, so the drain would act on a guess") + } + if len(census.Unmanaged) != 1 { + t.Errorf("unmanaged = %v, want the unclassifiable volume counted where it blocks", + census.Unmanaged) + } +} + +// A PersistentVolume of this driver whose handle names no volume indexes +// nothing, rather than indexing every such object under one empty key. +func TestAnEmptyVolumeHandleIndexesNothing(t *testing.T) { + blank := aPersistentVolume("pv-1", "volume-1") + blank.Spec.CSI.VolumeHandle = "" + + census := censusOf(t, + aControlPlane().holding(onNode("volume-1", "pvc-abc")), + blank, aClaim("pv-1", false)) + + if len(census.Unmanaged) != 1 { + t.Errorf("unmanaged = %v, want the volume an empty handle accounts for nothing about", + census.Unmanaged) + } +} + +// The default pattern is the CRD's, and it matches what the rebalancer makes +// rather than what a user does. +func TestTheDefaultFilterMatchesTheBenchmarkVolumesOnly(t *testing.T) { + filter, err := systemVolumeFilter( + anOperation("a-drain", simplyblockv1alpha2.StorageNodeOpsActionRemove)) + if err != nil { + t.Fatalf("compiling the default pattern: %v", err) + } + if !filter.MatchString("sb-fio-baseline-read") { + t.Error("the default pattern does not match the rebalancer's own benchmark volume") + } + if filter.MatchString("pvc-a-users-volume") { + t.Error("the default pattern matches a user's volume, which verification would delete") + } +} + +// An operation that states a pattern is asking for that one instead. +func TestAStatedFilterReplacesTheDefault(t *testing.T) { + ops := anOperation("a-drain", simplyblockv1alpha2.StorageNodeOpsActionRemove) + custom := "^bench-.*" + ops.Spec.Remove = &simplyblockv1alpha2.RemoveSpec{SystemVolumeFilterRegex: &custom} + + filter, err := systemVolumeFilter(ops) + if err != nil { + t.Fatalf("compiling the stated pattern: %v", err) + } + if !filter.MatchString("bench-read") { + t.Error("the stated pattern does not match what it says it matches") + } + if filter.MatchString("sb-fio-baseline-read") { + t.Error("the default pattern is still being applied beside the stated one") + } +} + +// A pattern that does not compile is fatal rather than retried: the expression +// is in the spec, and no number of passes will make it parse. +func TestAPatternThatDoesNotCompileEndsTheOperation(t *testing.T) { + ops := anOperation("a-drain", simplyblockv1alpha2.StorageNodeOpsActionRemove) + broken := "[" + ops.Spec.Remove = &simplyblockv1alpha2.RemoveSpec{SystemVolumeFilterRegex: &broken} + + _, err := systemVolumeFilter(ops) + + var fatal *terminalStepError + if !errors.As(err, &fatal) { + t.Errorf("err = %v, want the terminal kind: retrying cannot make a pattern parse", err) + } +} diff --git a/operator/internal/controllers/node/drain_test.go b/operator/internal/controllers/node/drain_test.go new file mode 100644 index 000000000..3c9ebd363 --- /dev/null +++ b/operator/internal/controllers/node/drain_test.go @@ -0,0 +1,498 @@ +// Draining a node before it leaves. +// +// Removing a storage node destroys it, so every logical volume whose data lives +// on it has to be somewhere else first. The five steps are ordered around that +// one fact: validation runs before the suspend, because suspending a node whose +// drain cannot complete takes capacity out of the cluster and leaves it out for +// as long as the blocker goes unnoticed; verification runs after the migration, +// because the census is the authority on what is left rather than the counter of +// what moved; and the removal is the last step rather than the operation. +// +// design-storagenode.md §8. + +package node + +import ( + "context" + "errors" + "testing" + + "k8s.io/client-go/tools/events" + "sigs.k8s.io/controller-runtime/pkg/client" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + vmigration "github.com/simplyblock/simplyblock-operator/internal/volumemigration" +) + +// scriptedMover is the fan-out as a drain sees it: whatever moves the test says +// are outstanding, plus whatever the drain raised during the case. +// +// The kind that carries a move is the deployment's and not the drain's, which is +// exactly why this is an interface the drain talks through and why a fake here +// is a fair stand-in for either one. +type scriptedMover struct { + moves []vmigration.Move + started []vmigration.MoveRequest + deleted []string +} + +func (m *scriptedMover) Start(_ context.Context, request vmigration.MoveRequest) error { + m.started = append(m.started, request) + return nil +} + +func (m *scriptedMover) List( + _ context.Context, _ string, _ map[string]string, +) ([]vmigration.Move, error) { + return m.moves, nil +} + +func (m *scriptedMover) Get(_ context.Context, name, _ string) (vmigration.Move, error) { + for _, move := range m.moves { + if move.Name == name { + return move, nil + } + } + return vmigration.Move{}, errors.New("no such move") +} + +func (m *scriptedMover) Delete(_ context.Context, move vmigration.Move) error { + m.deleted = append(m.deleted, move.Name) + remaining := m.moves[:0] + for _, existing := range m.moves { + if existing.Name != move.Name { + remaining = append(remaining, existing) + } + } + m.moves = remaining + return nil +} + +// aDrain is the operation these cases run. +func aDrain() *simplyblockv1alpha2.StorageNodeOps { + return anOperation("a-drain", simplyblockv1alpha2.StorageNodeOpsActionRemove) +} + +// aDraining builds the world, with the drain object already in it so that its +// status can be written and read back. +func aDraining( + t *testing.T, api *scriptedControlPlane, mover *scriptedMover, objects ...client.Object, +) (*StorageNodeOpsReconciler, client.Client) { + t.Helper() + r, apiClient := anOpsWorld(t, api, append(objects, aDrain())...) + r.Mover = mover + return r, apiClient +} + +// A pinned claim stops the drain where nothing has been done yet, and says which +// annotation to remove from which volume. Blocking here rather than later is +// what leaves the node fully operational while somebody decides. +func TestAPinnedVolumeStopsTheDrainBeforeItSuspendsAnything(t *testing.T) { + api := aControlPlane().holding(onNode("volume-1", "pvc-abc")) + r, _ := aDraining(t, api, &scriptedMover{}, + aPersistentVolume("pv-1", "volume-1"), aClaim("pv-1", true)) + + _, err := r.perform(context.Background(), aDrain(), stepValidating) + + var blocked *blockedStepError + if !errors.As(err, &blocked) { + t.Fatalf("err = %v, want the drain held by the pin", err) + } + if blocked.reason != DrainBlocked { + t.Errorf("the hold is announced as %q, want %q", blocked.reason, DrainBlocked) + } + if asked := api.asked("Suspend"); asked != 0 { + t.Errorf("the node was suspended %d time(s) by a drain that cannot finish", asked) + } +} + +// A volume nothing in Kubernetes accounts for blocks the same way, and for the +// stronger reason: there is no object to move and deleting it would destroy data +// nobody is tracking. +func TestAnUnmanagedVolumeStopsTheDrain(t *testing.T) { + api := aControlPlane().holding(onNode("volume-orphan", "hand-made")) + r, _ := aDraining(t, api, &scriptedMover{}) + + _, err := r.perform(context.Background(), aDrain(), stepValidating) + + var blocked *blockedStepError + if !errors.As(err, &blocked) { + t.Fatalf("err = %v, want the drain held by the volume nothing accounts for", err) + } +} + +// Validation fixes the total every later step reports progress against, and +// writes it once. A node with nothing movable on it is a total of zero rather +// than an absent one. +func TestValidationWritesTheTotalTheDrainIsMeasuredAgainst(t *testing.T) { + api := aControlPlane().holding(onNode("volume-1", "pvc-abc"), onNode("volume-2", "pvc-def")) + r, apiClient := aDraining(t, api, &scriptedMover{}, + aPersistentVolume("pv-1", "volume-1"), aClaim("pv-1", false), + aPersistentVolume("pv-2", "volume-2"), aClaim("pv-2", false)) + + done, err := r.perform(context.Background(), aDrain(), stepValidating) + if err != nil { + t.Fatalf("validating: %v", err) + } + if !done { + t.Error("validation did not finish although nothing blocks the drain") + } + + got := operationRead(t, apiClient, "a-drain") + if got.Status.Drain == nil || got.Status.Drain.VolumesTotal != 2 { + t.Errorf("drain = %+v, want the two movable volumes counted", got.Status.Drain) + } + if got.Status.Drain.VolumesMigrated != 0 { + t.Errorf("%d volumes are already recorded as moved before any move was raised", + got.Status.Drain.VolumesMigrated) + } +} + +// A census that could not be completed is retried rather than acted on: a claim +// that could not be read put a volume where it blocks, and reporting that as a +// blocker would name a volume that is in fact accounted for. +func TestAnIncompleteCensusIsRetriedRatherThanReportedAsABlocker(t *testing.T) { + api := aControlPlane().holding(onNode("volume-1", "pvc-abc")) + r, _ := anOpsWorldWith(t, api, refusingClaims(), + aDrain(), aPersistentVolume("pv-1", "volume-1"), aClaim("pv-1", false)) + + _, err := r.perform(context.Background(), aDrain(), stepValidating) + if err == nil { + t.Fatal("the drain acted on a census it knows is incomplete") + } + var blocked *blockedStepError + if errors.As(err, &blocked) { + t.Errorf("err = %v, want an ordinary retry rather than a hold naming a volume", err) + } +} + +// The suspend is skipped against a node already at or past where it would put +// it, which is what makes re-entering the step after a lost response harmless. +func TestTheSuspendIsSkippedWhenTheNodeIsAlreadyOutOfService(t *testing.T) { + for _, status := range []string{nodeStatusSuspended, nodeStatusOffline} { + t.Run(status, func(t *testing.T) { + api := aControlPlane().reporting(status) + r, _ := aDraining(t, api, &scriptedMover{}) + + done, err := r.perform(context.Background(), aDrain(), stepSuspending) + if err != nil { + t.Fatalf("suspending: %v", err) + } + if !done { + t.Errorf("the step did not finish against a node already %s", status) + } + if asked := api.asked("Suspend"); asked != 0 { + t.Errorf("Suspend was issued %d time(s) against a node already %s", asked, status) + } + }) + } +} + +// An online node is suspended, and the step does not finish on the call: it +// finishes when the control plane reports the node suspended. +func TestAnOnlineNodeIsSuspendedAndWaitedFor(t *testing.T) { + api := aControlPlane() + r, _ := aDraining(t, api, &scriptedMover{}) + + done, err := r.perform(context.Background(), aDrain(), stepSuspending) + if err != nil { + t.Fatalf("suspending: %v", err) + } + if done { + t.Error("the step finished on the call rather than on the node reporting suspended") + } + if asked := api.asked("Suspend"); asked != 1 { + t.Errorf("Suspend was issued %d time(s), want once", asked) + } +} + +// Every movable volume gets a move, named so that a second pass finds the object +// it made rather than making another. +func TestEveryMovableVolumeIsGivenAMove(t *testing.T) { + api := aControlPlane(). + withPeer(opsPeerID, nodeStatusOnline). + holding(onNode("volume-1", "pvc-abc"), onNode("volume-2", "pvc-def")) + mover := &scriptedMover{} + r, _ := aDraining(t, api, mover, + aPersistentVolume("pv-1", "volume-1"), aClaim("pv-1", false), + aPersistentVolume("pv-2", "volume-2"), aClaim("pv-2", false)) + + done, err := r.perform(context.Background(), aDrain(), stepMigratingVolumes) + if err != nil { + t.Fatalf("migrating: %v", err) + } + if done { + t.Error("the step finished on the pass that raised the moves") + } + if len(mover.started) != 2 { + t.Fatalf("%d moves were raised for two movable volumes", len(mover.started)) + } + for _, request := range mover.started { + if request.TargetNodeUUID != opsPeerID { + t.Errorf("%s is being moved to %q, want the cluster's one online peer", + request.PVName, request.TargetNodeUUID) + } + if request.Name != migrationName(opsNodeID, request.PVName) { + t.Errorf("the move is named %q, and a second pass would raise it again", + request.Name) + } + if request.Labels[drainNodeLabel] != opsNodeID { + t.Errorf("the move carries %v, and the drain could not find its own fan-out", + request.Labels) + } + } +} + +// A move already raised is not raised again, which is what makes the step safe +// to re-enter on every pass while the copies run. +func TestAVolumeAlreadyMovingIsNotGivenASecondMove(t *testing.T) { + api := aControlPlane(). + withPeer(opsPeerID, nodeStatusOnline). + holding(onNode("volume-1", "pvc-abc")) + mover := &scriptedMover{moves: []vmigration.Move{{ + Name: migrationName(opsNodeID, "pv-1"), PVName: "pv-1", Phase: vmigration.MoveRunning, + }}} + r, apiClient := aDraining(t, api, mover, + aPersistentVolume("pv-1", "volume-1"), aClaim("pv-1", false)) + + done, err := r.perform(context.Background(), aDrain(), stepMigratingVolumes) + if err != nil { + t.Fatalf("migrating: %v", err) + } + if done { + t.Error("the step finished while a move was still running") + } + if len(mover.started) != 0 { + t.Errorf("%d further moves were raised for a volume already moving", len(mover.started)) + } + + got := operationRead(t, apiClient, "a-drain") + if got.Status.Drain == nil || got.Status.Drain.VolumesMigrated != 0 { + t.Errorf("drain = %+v, want nothing recorded as moved while the move runs", + got.Status.Drain) + } +} + +// A failed move is deleted and replaced against a fresh target rather than +// failing the drain: the volume is still on the node, and the peer that could +// not take it is not the only peer. +func TestAFailedMoveIsRetriedRatherThanFailingTheDrain(t *testing.T) { + api := aControlPlane(). + withPeer(opsPeerID, nodeStatusOnline). + holding(onNode("volume-1", "pvc-abc")) + mover := &scriptedMover{moves: []vmigration.Move{{ + Name: migrationName(opsNodeID, "pv-1"), + PVName: "pv-1", + Phase: vmigration.MoveFailed, + Message: "the target refused the copy", + }}} + r, _ := aDraining(t, api, mover, + aPersistentVolume("pv-1", "volume-1"), aClaim("pv-1", false)) + ops := aDrain() + + done, err := r.perform(context.Background(), ops, stepMigratingVolumes) + if err != nil { + t.Fatalf("migrating: %v", err) + } + if done { + t.Error("the step finished although a volume is still on the node") + } + if len(mover.deleted) != 1 { + t.Errorf("the failed move was deleted %d time(s), so nothing would replace it", + len(mover.deleted)) + } + if !announced(r.Recorder.(*events.FakeRecorder), MigrationRetried) { + t.Error("nothing announced the retry, so a drain that keeps retrying looks like one that stalled") + } + + // The next pass raises it again, against a target chosen afresh. + if _, err := r.perform(context.Background(), ops, stepMigratingVolumes); err != nil { + t.Fatalf("migrating: %v", err) + } + if len(mover.started) != 1 { + t.Errorf("%d moves were raised on the pass after the retry, want the replacement", + len(mover.started)) + } +} + +// Every move finished is progress recorded and the objects reaped, and the +// counter is written before the delete so that a crash between the two leaves +// the progress recorded rather than lost. +func TestFinishedMovesAreRecordedAndThenReaped(t *testing.T) { + api := aControlPlane().withPeer(opsPeerID, nodeStatusOnline) + mover := &scriptedMover{moves: []vmigration.Move{ + {Name: migrationName(opsNodeID, "pv-1"), PVName: "pv-1", Phase: vmigration.MoveSucceeded}, + {Name: migrationName(opsNodeID, "pv-2"), PVName: "pv-2", Phase: vmigration.MoveSucceeded}, + }} + r, apiClient := aDraining(t, api, mover) + + done, err := r.perform(context.Background(), aDrain(), stepMigratingVolumes) + if err != nil { + t.Fatalf("migrating: %v", err) + } + if done { + t.Error("the step finished on the pass that reaped the moves rather than on a fresh census") + } + + got := operationRead(t, apiClient, "a-drain") + if got.Status.Drain == nil || got.Status.Drain.VolumesMigrated != 2 { + t.Errorf("drain = %+v, want both moves recorded", got.Status.Drain) + } + if len(mover.deleted) != 2 { + t.Errorf("%d of two finished moves were reaped, and a hundred-volume drain leaves "+ + "a hundred objects behind", len(mover.deleted)) + } +} + +// Nothing movable left and nothing outstanding is a drain that is done, and the +// census is the authority on that rather than the counter: the counter records +// what this operation did, and the census is what is on the node. +func TestADrainIsDoneWhenTheNodeHoldsNothingMovable(t *testing.T) { + r, _ := aDraining(t, aControlPlane().withPeer(opsPeerID, nodeStatusOnline), &scriptedMover{}) + + done, err := r.perform(context.Background(), aDrain(), stepMigratingVolumes) + if err != nil { + t.Fatalf("migrating: %v", err) + } + if !done { + t.Error("the step did not finish against a node with nothing left to move") + } + if !announced(r.Recorder.(*events.FakeRecorder), DrainCompleted) { + t.Error("nothing announced that the node had been emptied") + } +} + +// Verification deletes the benchmark volumes the migration skipped, and does not +// call the node empty on the pass that deleted them: the control plane's +// deletion is asynchronous, so a later pass has to re-read. +func TestVerificationDeletesTheBenchmarkVolumesAndRereads(t *testing.T) { + api := aControlPlane().holding(onNode("volume-bench", "sb-fio-baseline-read")) + r, _ := aDraining(t, api, &scriptedMover{}) + + done, err := r.perform(context.Background(), aDrain(), stepVerifying) + if err != nil { + t.Fatalf("verifying: %v", err) + } + if done { + t.Error("the step called the node empty on the pass that asked for the deletions") + } + if asked := api.asked("DeleteVolume"); asked != 1 { + t.Errorf("DeleteVolume was issued %d time(s), want one per benchmark volume", asked) + } +} + +// A node that still holds a user's volume after the migration is not a node to +// remove, and the drain holds rather than destroying it. +func TestVerificationHoldsWhileAUsersVolumeIsStillThere(t *testing.T) { + api := aControlPlane().holding(onNode("volume-1", "pvc-abc")) + r, _ := aDraining(t, api, &scriptedMover{}, + aPersistentVolume("pv-1", "volume-1"), aClaim("pv-1", false)) + + _, err := r.perform(context.Background(), aDrain(), stepVerifying) + + var blocked *blockedStepError + if !errors.As(err, &blocked) { + t.Fatalf("err = %v, want the removal held while the node still holds a volume", err) + } +} + +// A benchmark volume that cannot be deleted is one the removal would destroy, so +// the operation fails rather than proceeding. +func TestABenchmarkVolumeThatCannotBeDeletedEndsTheDrain(t *testing.T) { + api := aControlPlane(). + holding(onNode("volume-bench", "sb-fio-baseline-read")). + refusing("DeleteVolume", errors.New("the control plane refused")) + r, _ := aDraining(t, api, &scriptedMover{}) + + _, err := r.perform(context.Background(), aDrain(), stepVerifying) + + var fatal *terminalStepError + if !errors.As(err, &fatal) { + t.Errorf("err = %v, want the terminal kind: the node still holds the volume", err) + } +} + +// An empty node passes verification. +func TestAnEmptyNodePassesVerification(t *testing.T) { + r, _ := aDraining(t, aControlPlane(), &scriptedMover{}) + + done, err := r.perform(context.Background(), aDrain(), stepVerifying) + if err != nil { + t.Fatalf("verifying: %v", err) + } + if !done { + t.Error("an empty node did not pass verification") + } +} + +// The removal is the last step, and a refusal is the control plane's answer +// about what the cluster can afford to lose. Retrying cannot change it, so the +// operation fails and the unwind puts the node back into service. +func TestARefusedRemovalEndsTheDrain(t *testing.T) { + api := aControlPlane().refusing("RemoveNode", errors.New("the cluster cannot lose this node")) + r, _ := aDraining(t, api, &scriptedMover{}) + + _, err := r.perform(context.Background(), aDrain(), stepRemoving) + + var fatal *terminalStepError + if !errors.As(err, &fatal) { + t.Errorf("err = %v, want the terminal kind for a removal the control plane refused", err) + } +} + +// An accepted removal is the end of the drain. +func TestAnAcceptedRemovalFinishesTheDrain(t *testing.T) { + api := aControlPlane() + r, _ := aDraining(t, api, &scriptedMover{}) + + done, err := r.perform(context.Background(), aDrain(), stepRemoving) + if err != nil { + t.Fatalf("removing: %v", err) + } + if !done { + t.Error("the step did not finish although the removal was accepted") + } + if asked := api.asked("RemoveNode"); asked != 1 { + t.Errorf("RemoveNode was issued %d time(s), want once", asked) + } +} + +// Deleting a drain stops what it fanned out, so nothing keeps copying for an +// operation that no longer exists. +func TestDeletingADrainStopsWhatItFannedOut(t *testing.T) { + mover := &scriptedMover{moves: []vmigration.Move{{ + Name: migrationName(opsNodeID, "pv-1"), PVName: "pv-1", Phase: vmigration.MoveSucceeded, + }}} + r, _ := aDraining(t, aControlPlane(), mover) + + pending, err := r.cascadeMigrations(context.Background(), aDrain()) + if err != nil { + t.Fatalf("cascading: %v", err) + } + if pending { + t.Error("a finished move was reported as still running") + } + if len(mover.deleted) != 1 { + t.Errorf("%d moves were reaped, want the drain's whole fan-out", len(mover.deleted)) + } +} + +// A move still running holds the deletion, because the cascade's own deletes +// have to be admissible before the operation goes. +func TestADrainBeingDeletedWaitsForAMoveStillRunning(t *testing.T) { + mover := &scriptedMover{moves: []vmigration.Move{{ + Name: migrationName(opsNodeID, "pv-1"), PVName: "pv-1", Phase: vmigration.MoveRunning, + }}} + r, _ := aDraining(t, aControlPlane(), mover) + + pending, err := r.cascadeMigrations(context.Background(), aDrain()) + if err != nil { + t.Fatalf("cascading: %v", err) + } + if !pending { + t.Error("a move still copying was not reported as pending") + } + if len(mover.deleted) != 0 { + t.Errorf("%d moves were reaped mid-copy", len(mover.deleted)) + } +} diff --git a/operator/internal/controllers/node/fixtures_test.go b/operator/internal/controllers/node/fixtures_test.go new file mode 100644 index 000000000..d38a3ec7d --- /dev/null +++ b/operator/internal/controllers/node/fixtures_test.go @@ -0,0 +1,370 @@ +// The world the operation suites in this package drive, and the control plane +// they drive it against. +// +// Every step of every action is a call against the control plane followed by a +// predicate over what it answers next, so a suite that exercises a step needs +// both halves scripted: what the backend reports now, and what it was asked to +// do. One fake rather than one per suite, because the seven actions ask for +// different things and assert the same way — a restart issued once, a suspend +// skipped because the node was already suspended, a system volume deleted. +// +// It lives in a file of its own for the reason internal/controllers/testsupport +// does: a fixture beside one suite becomes a second copy beside the next. + +package node + +import ( + "context" + "errors" + "fmt" + "testing" + + corev1 "k8s.io/api/core/v1" + discoveryv1 "k8s.io/api/discovery/v1" + policyv1 "k8s.io/api/policy/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/client-go/tools/events" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + "sigs.k8s.io/controller-runtime/pkg/client/interceptor" + + atlaskube "github.com/simplyblock/atlas/kube" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/testsupport" + "github.com/simplyblock/simplyblock-operator/internal/utils" + "github.com/simplyblock/simplyblock-operator/internal/webapi" +) + +const ( + // The one deployment every suite built on these fixtures runs in: a cluster + // of one object, a node of one object, and the worker it runs on. + opsNamespace = "simplyblock" + opsCluster = "a-cluster" + opsWorker = "worker-1" + opsTarget = "worker-2" + + // opsNodeID is the backend node behind the object, and opsPeerID a second + // node of the same cluster — the one a drain moves volumes to. + opsNodeID = "node-1111" + opsPeerID = "node-2222" + + // opsClusterID is what the control plane knows the cluster by. + opsClusterID = "cluster-1111" + + // opsNodeName is the node object every operation here targets, and opsPool + // the one pool its volumes live in. + opsNodeName = "a-node" + opsPool = "pool-1" +) + +// scriptedControlPlane answers what a test scripted and records what it was +// asked. +// +// The embedded interface is nil on purpose: a call no suite scripts is a test +// reaching past what it is about, and a panic says so where a zero value would +// quietly pass. +type scriptedControlPlane struct { + ControlPlane + + // nodes is what the control plane currently reports, by backend id. A test + // that wants a step's second pass to see a different state rewrites an entry + // between passes, which is what the backend doing the thing looks like from + // here. + nodes map[string]NodeReading + + // pools and volumes are what a drain's classification walks. The volumes are + // keyed by pool, because that is the shape the control plane serves them in + // and the shape the walk reads them in. + pools []webapi.StoragePoolInfo + volumes map[string][]webapi.VolumeInfo + + // refuse fails one call by method name, which is how a suite reaches the + // paths that only a refusing control plane produces. + refuse map[string]error + + // calls is every call in order, each spelled as its method, a colon, and its + // argument, which is what an assertion about a call being skipped or issued + // once reads. + calls []string + + // restarts carries the parameters of each restart, because three actions + // issue one and they differ precisely in what they fill in. + restarts []RestartParams +} + +// aControlPlane reports one online node and nothing else. +func aControlPlane() *scriptedControlPlane { + return &scriptedControlPlane{ + nodes: map[string]NodeReading{ + opsNodeID: {UUID: opsNodeID, Status: nodeStatusOnline, ManagementIP: "10.0.0.1"}, + }, + volumes: map[string][]webapi.VolumeInfo{}, + refuse: map[string]error{}, + } +} + +// reporting replaces what the control plane says about the node under +// operation, which is how a test moves the backend between two passes of a step. +func (c *scriptedControlPlane) reporting(status string) *scriptedControlPlane { + reading := c.nodes[opsNodeID] + reading.UUID, reading.Status = opsNodeID, status + c.nodes[opsNodeID] = reading + return c +} + +// withPeer adds a second node to the cluster, which is what a drain needs one of +// to have anywhere to move to. +func (c *scriptedControlPlane) withPeer(nodeID, status string) *scriptedControlPlane { + c.nodes[nodeID] = NodeReading{UUID: nodeID, Status: status} + return c +} + +// holding puts volumes in the cluster's pool. +func (c *scriptedControlPlane) holding(volumes ...webapi.VolumeInfo) *scriptedControlPlane { + if len(c.pools) == 0 { + c.pools = append(c.pools, webapi.StoragePoolInfo{UUID: opsPool, Name: opsPool}) + } + c.volumes[opsPool] = append(c.volumes[opsPool], volumes...) + return c +} + +// refusing makes one method answer with an error. +func (c *scriptedControlPlane) refusing(method string, err error) *scriptedControlPlane { + c.refuse[method] = err + return c +} + +// asked is how many times one method was called, which is what "the call was +// skipped" and "the call was issued once" are both assertions about. +func (c *scriptedControlPlane) asked(method string) int { + count := 0 + for _, call := range c.calls { + if call == method || len(call) > len(method) && call[:len(method)+1] == method+":" { + count++ + } + } + return count +} + +func (c *scriptedControlPlane) record(method, argument string) error { + c.calls = append(c.calls, method+":"+argument) + return c.refuse[method] +} + +func (c *scriptedControlPlane) StorageNode( + _ context.Context, _, nodeID string, +) (NodeReading, bool, error) { + if err := c.record("StorageNode", nodeID); err != nil { + return NodeReading{}, false, err + } + reading, found := c.nodes[nodeID] + return reading, found, nil +} + +func (c *scriptedControlPlane) StorageNodes( + _ context.Context, clusterID string, +) ([]NodeReading, error) { + if err := c.record("StorageNodes", clusterID); err != nil { + return nil, err + } + readings := make([]NodeReading, 0, len(c.nodes)) + for _, reading := range c.nodes { + readings = append(readings, reading) + } + return readings, nil +} + +func (c *scriptedControlPlane) AddNode( + _ context.Context, clusterID string, _ utils.StorageNodeSetAddParams, +) error { + return c.record("AddNode", clusterID) +} + +func (c *scriptedControlPlane) Suspend(_ context.Context, _, nodeID string) error { + return c.record("Suspend", nodeID) +} + +func (c *scriptedControlPlane) Resume(_ context.Context, _, nodeID string) error { + return c.record("Resume", nodeID) +} + +func (c *scriptedControlPlane) ShutdownNode(_ context.Context, _, nodeID string) error { + return c.record("ShutdownNode", nodeID) +} + +func (c *scriptedControlPlane) RestartNode( + _ context.Context, _, nodeID string, params RestartParams, +) error { + c.restarts = append(c.restarts, params) + return c.record("RestartNode", nodeID) +} + +func (c *scriptedControlPlane) Promote(_ context.Context, _, nodeID string) error { + return c.record("Promote", nodeID) +} + +func (c *scriptedControlPlane) RemoveNode(_ context.Context, _, nodeID string) error { + return c.record("RemoveNode", nodeID) +} + +func (c *scriptedControlPlane) StoragePools( + _ context.Context, clusterID string, +) ([]webapi.StoragePoolInfo, error) { + if err := c.record("StoragePools", clusterID); err != nil { + return nil, err + } + return c.pools, nil +} + +func (c *scriptedControlPlane) PoolVolumes( + _ context.Context, _, poolID string, +) ([]webapi.VolumeInfo, error) { + if err := c.record("PoolVolumes", poolID); err != nil { + return nil, err + } + return c.volumes[poolID], nil +} + +func (c *scriptedControlPlane) DeleteVolume(_ context.Context, _, _, volumeID string) error { + return c.record("DeleteVolume", volumeID) +} + +// anOpsNode is the node every operation in these suites targets: provisioned, +// on opsWorker, holding no lock. +func anOpsNode() *simplyblockv1alpha2.StorageNode { + node := &simplyblockv1alpha2.StorageNode{ + ObjectMeta: metav1.ObjectMeta{Name: opsNodeName, Namespace: opsNamespace}, + Spec: simplyblockv1alpha2.StorageNodeSpec{ + ClusterRef: opsCluster, + WorkerNode: opsWorker, + }, + } + node.Status.UUID = opsNodeID + return node +} + +// anOpsCluster is the node's cluster, already created in the control plane. +func anOpsCluster() *simplyblockv1alpha2.StorageCluster { + cluster := &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{Name: opsCluster, Namespace: opsNamespace}, + } + cluster.Status.UUID = opsClusterID + return cluster +} + +// anOperation is one operation of the given action against anOpsNode. +func anOperation( + name string, action simplyblockv1alpha2.StorageNodeOpsAction, +) *simplyblockv1alpha2.StorageNodeOps { + return &simplyblockv1alpha2.StorageNodeOps{ + ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: opsNamespace}, + Spec: simplyblockv1alpha2.StorageNodeOpsSpec{ + NodeRef: opsNodeName, + Action: action, + }, + } +} + +// aWorker is a Kubernetes node, Ready unless a test says otherwise. +func aWorker(name string, ready bool) *corev1.Node { + status := corev1.ConditionFalse + if ready { + status = corev1.ConditionTrue + } + return &corev1.Node{ + ObjectMeta: metav1.ObjectMeta{Name: name}, + Status: corev1.NodeStatus{ + Addresses: []corev1.NodeAddress{{Type: corev1.NodeInternalIP, Address: "10.0.0.1"}}, + Conditions: []corev1.NodeCondition{{Type: corev1.NodeReady, Status: status}}, + }, + } +} + +// anOpsWorld builds the reconciler over a fake client holding the node, its +// cluster, and whatever else the suite put in. +func anOpsWorld( + t *testing.T, api ControlPlane, objects ...client.Object, +) (*StorageNodeOpsReconciler, client.Client) { + t.Helper() + return anOpsWorldWith(t, api, interceptor.Funcs{}, objects...) +} + +// anOpsWorldWith scripts the client's answers too, which is how a suite reaches +// a read that fails — the census is written to treat one of those as a fact +// about the census rather than about the volume. +func anOpsWorldWith( + t *testing.T, api ControlPlane, funcs interceptor.Funcs, objects ...client.Object, +) (*StorageNodeOpsReconciler, client.Client) { + t.Helper() + scheme := testsupport.NewScheme(t, + corev1.AddToScheme, policyv1.AddToScheme, discoveryv1.AddToScheme) + + world := append([]client.Object{anOpsNode(), anOpsCluster()}, objects...) + apiClient := fake.NewClientBuilder().WithScheme(scheme). + WithObjects(world...). + WithStatusSubresource( + &simplyblockv1alpha2.StorageNodeOps{}, + &simplyblockv1alpha2.StorageNode{}, + &simplyblockv1alpha2.StorageCluster{}, + ). + WithInterceptorFuncs(funcs). + Build() + + reconciler := &StorageNodeOpsReconciler{ + Client: apiClient, + Scheme: scheme, + Recorder: events.NewFakeRecorder(64), + API: api, + Workload: &Workload{Client: apiClient}, + } + return reconciler, apiClient +} + +// aPersistentVolume is one simplyblock volume as Kubernetes accounts for it, +// claimed by a PersistentVolumeClaim named after it. +func aPersistentVolume(name, volumeUUID string) *corev1.PersistentVolume { + return &corev1.PersistentVolume{ + ObjectMeta: metav1.ObjectMeta{Name: name}, + Spec: corev1.PersistentVolumeSpec{ + PersistentVolumeSource: corev1.PersistentVolumeSource{ + CSI: &corev1.CSIPersistentVolumeSource{ + Driver: utils.CSIProvisioner, + VolumeHandle: fmt.Sprintf("%s:%s:%s", opsClusterID, opsPool, volumeUUID), + }, + }, + ClaimRef: &corev1.ObjectReference{ + Namespace: opsNamespace, + Name: name + "-claim", + }, + }, + } +} + +// aClaim is the claim behind such a volume, pinned or not. +func aClaim(name string, pinned bool) *corev1.PersistentVolumeClaim { + claim := &corev1.PersistentVolumeClaim{ + ObjectMeta: metav1.ObjectMeta{Name: name + "-claim", Namespace: opsNamespace}, + } + if pinned { + claim.Annotations = map[string]string{atlaskube.AnnoSelectedStorageNode: opsNodeID} + } + return claim +} + +// refusingClaims fails every claim read, which is the transient failure the +// census is written to treat as a fact about the census rather than as a fact +// about the volume. +func refusingClaims() interceptor.Funcs { + return interceptor.Funcs{ + Get: func( + ctx context.Context, c client.WithWatch, key client.ObjectKey, + object client.Object, options ...client.GetOption, + ) error { + if _, claim := object.(*corev1.PersistentVolumeClaim); claim { + return errors.New("the API server is briefly away") + } + return c.Get(ctx, key, object, options...) + }, + } +} diff --git a/operator/internal/controllers/node/hostmaintenance_test.go b/operator/internal/controllers/node/hostmaintenance_test.go new file mode 100644 index 000000000..13537fed2 --- /dev/null +++ b/operator/internal/controllers/node/hostmaintenance_test.go @@ -0,0 +1,362 @@ +// Surviving a Kubernetes worker being drained. +// +// The window runs backward from an ordinary disruption budget: one that allows no +// disruption at all is created *before* the backend node is shut down, so +// `kubectl drain` blocks on it while the node is taken down gracefully, and +// relaxing it is what lets the drain proceed. The ordering is the whole +// mechanism — a drain that reaches the pod before the budget exists kills the +// SPDK process under a running backend node, which is exactly what the action +// exists to prevent. +// +// The gate in front of it is cluster-wide and counted by worker rather than by +// operation, because two nodes on one host go into maintenance together and the +// pair is one worker's worth of unavailability. +// +// design-storagenode.md §10. + +package node + +import ( + "context" + "errors" + "testing" + + corev1 "k8s.io/api/core/v1" + policyv1 "k8s.io/api/policy/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" + "sigs.k8s.io/controller-runtime/pkg/client" + + "github.com/simplyblock/atlas/ptr" + "github.com/simplyblock/atlas/statemachine" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// aWindow is a maintenance window on the fixture's node. +func aWindow(name string) *simplyblockv1alpha2.StorageNodeOps { + return anOperation(name, simplyblockv1alpha2.StorageNodeOpsActionHostMaintenance) +} + +// aRunningWindow is somebody else's window, past Holding and therefore holding a +// slot, against a node of its own on the given worker. +func aRunningWindow( + name, nodeName, worker string, +) (*simplyblockv1alpha2.StorageNodeOps, *simplyblockv1alpha2.StorageNode) { + ops := aWindow(name) + ops.Spec.NodeRef = nodeName + ops.Status.Phase = simplyblockv1alpha2.StorageNodeOpsPhaseRunning + ops.Status.Step = kubeStep(stepShuttingDown) + + node := anOpsNode() + node.Name = nodeName + node.Spec.WorkerNode = worker + return ops, node +} + +// allowing states how many workers the cluster can afford to have down at once, +// which is the smaller of what it asks for and the fault tolerance the control +// plane reports. +func allowing(t *testing.T, apiClient client.Client, windows int32) { + t.Helper() + var cluster simplyblockv1alpha2.StorageCluster + key := client.ObjectKey{Namespace: opsNamespace, Name: opsCluster} + if err := apiClient.Get(context.Background(), key, &cluster); err != nil { + t.Fatalf("reading the cluster: %v", err) + } + cluster.Status.MaxConcurrentWorkerRestarts = ptr.To(windows) + if err := apiClient.Status().Update(context.Background(), &cluster); err != nil { + t.Fatalf("stating the cluster's concurrency: %v", err) + } +} + +// held runs the gate and reports whether it held this window back. +func held(t *testing.T, r *StorageNodeOpsReconciler, ops *simplyblockv1alpha2.StorageNodeOps) bool { + t.Helper() + done, err := r.perform(context.Background(), ops, stepHolding) + if err == nil { + return !done + } + var blocked *blockedStepError + if !errors.As(err, &blocked) { + t.Fatalf("holding: %v", err) + } + if blocked.reason != MaintenanceQueued { + t.Errorf("the hold is announced as %q, want %q", blocked.reason, MaintenanceQueued) + } + return true +} + +// A cluster that says nothing about concurrency takes one worker at a time, +// which is the safe reading of a cluster whose fault tolerance is unknown. +func TestOneWorkerAtATimeIsWhatAnUnstatedConcurrencyMeans(t *testing.T) { + other, itsNode := aRunningWindow("another-window", "another-node", "worker-9") + r, _ := anOpsWorld(t, aControlPlane(), other, itsNode) + + if !held(t, r, aWindow("a-window")) { + t.Error("a second worker entered maintenance although the cluster stated no concurrency") + } +} + +// What the cluster says it can afford is what the gate admits. +func TestTheClusterSaysHowManyWorkersMayBeDownAtOnce(t *testing.T) { + other, itsNode := aRunningWindow("another-window", "another-node", "worker-9") + r, apiClient := anOpsWorld(t, aControlPlane(), other, itsNode) + allowing(t, apiClient, 2) + + if held(t, r, aWindow("a-window")) { + t.Error("the window waited although the cluster allows two workers at once") + } +} + +// A window still in Holding holds no slot. Counting one would let a deployment +// deadlock with every window waiting for every other. +func TestAWindowStillWaitingHoldsNoSlot(t *testing.T) { + queued, itsNode := aRunningWindow("another-window", "another-node", "worker-9") + queued.Status.Step = kubeStep(stepHolding) + r, _ := anOpsWorld(t, aControlPlane(), queued, itsNode) + + if held(t, r, aWindow("a-window")) { + t.Error("a window queued behind the same gate was counted as occupying it") + } +} + +// A finished window holds nothing either, whatever its last step was. +func TestAFinishedWindowHoldsNoSlot(t *testing.T) { + over, itsNode := aRunningWindow("another-window", "another-node", "worker-9") + over.Status.Phase = simplyblockv1alpha2.StorageNodeOpsPhaseSucceeded + r, _ := anOpsWorld(t, aControlPlane(), over, itsNode) + + if held(t, r, aWindow("a-window")) { + t.Error("a window that has finished was still counted as occupying a slot") + } +} + +// The count is by worker, so a sibling socket of the same host is not a second +// worker's worth of unavailability. This is the multi-socket case the gate +// exists to get right. +func TestASiblingSocketOfTheSameWorkerIsNotASecondWorker(t *testing.T) { + sibling, itsNode := aRunningWindow("the-siblings-window", "the-sibling", opsWorker) + r, _ := anOpsWorld(t, aControlPlane(), sibling, itsNode) + + if held(t, r, aWindow("a-window")) { + t.Error("the second socket of the worker already in maintenance waited for itself") + } +} + +// Another cluster's maintenance is another cluster's business. +func TestAWindowInAnotherClusterDoesNotHoldThisOneBack(t *testing.T) { + elsewhere, itsNode := aRunningWindow("another-clusters-window", "another-node", "worker-9") + itsNode.Spec.ClusterRef = "another-cluster" + r, _ := anOpsWorld(t, aControlPlane(), elsewhere, itsNode) + + if held(t, r, aWindow("a-window")) { + t.Error("a window in another cluster held this one back") + } +} + +// The budget exists before the shutdown is issued, which is the ordering the +// whole action rests on. +func TestTheEvictionIsBlockedBeforeTheNodeIsTakenDown(t *testing.T) { + api := aControlPlane() + r, apiClient := anOpsWorld(t, api, aReadyStoragePod(opsWorker)) + + done, err := r.perform(context.Background(), aWindow("a-window"), stepShuttingDown) + if err != nil { + t.Fatalf("shutting down: %v", err) + } + if done { + t.Error("the step finished while the node was still online") + } + if asked := api.asked("ShutdownNode"); asked != 1 { + t.Errorf("ShutdownNode was issued %d time(s), want once", asked) + } + + budget := budgetFor(t, apiClient, opsWorker) + if budget == nil { + t.Fatal("no budget holds the eviction, so a drain would evict the pod under a live node") + } + if budget.Spec.MaxUnavailable == nil || budget.Spec.MaxUnavailable.IntVal != 0 { + t.Errorf("maxUnavailable = %v, want 0 while the node is being taken down", + budget.Spec.MaxUnavailable) + } +} + +// A node that is already offline is where the shutdown was taking it. +func TestAnOfflineNodeNeedsNoShutdown(t *testing.T) { + api := aControlPlane().reporting(nodeStatusOffline) + r, _ := anOpsWorld(t, api, aReadyStoragePod(opsWorker)) + + done, err := r.perform(context.Background(), aWindow("a-window"), stepShuttingDown) + if err != nil { + t.Fatalf("shutting down: %v", err) + } + if !done { + t.Error("the step did not finish against a node that is already offline") + } + if asked := api.asked("ShutdownNode"); asked != 0 { + t.Errorf("ShutdownNode was issued %d time(s) against an offline node", asked) + } +} + +// Mid-restart is not a state to shut down from: the call would be refused, and +// the node is on its way somewhere anyway. +func TestANodeMidRestartIsWaitedForRatherThanShutDown(t *testing.T) { + api := aControlPlane().reporting(nodeStatusInRestart) + r, _ := anOpsWorld(t, api, aReadyStoragePod(opsWorker)) + + done, err := r.perform(context.Background(), aWindow("a-window"), stepShuttingDown) + if err != nil { + t.Fatalf("shutting down: %v", err) + } + if done { + t.Error("the step finished against a node whose restart is still running") + } + if asked := api.asked("ShutdownNode"); asked != 0 { + t.Errorf("ShutdownNode was issued %d time(s) into an in-flight restart", asked) + } +} + +// Releasing relaxes the budget the window created and finishes only once the pod +// has actually gone, which is what the drain was waiting to do. +func TestReleasingRelaxesTheBudgetAndWaitsForThePodToGo(t *testing.T) { + r, apiClient := anOpsWorld(t, aControlPlane(), aReadyStoragePod(opsWorker)) + ops := aWindow("a-window") + + if _, err := r.perform(context.Background(), ops, stepShuttingDown); err != nil { + t.Fatalf("shutting down: %v", err) + } + + done, err := r.perform(context.Background(), ops, stepReleasing) + if err != nil { + t.Fatalf("releasing: %v", err) + } + if done { + t.Error("the step finished while the storage pod was still on the worker") + } + budget := budgetFor(t, apiClient, opsWorker) + if budget == nil || budget.Spec.MaxUnavailable == nil || + budget.Spec.MaxUnavailable.IntVal != 1 { + t.Errorf("the budget is %v, want the one eviction the drain is waiting on", budget) + } + + if err := apiClient.Delete(context.Background(), aReadyStoragePod(opsWorker)); err != nil { + t.Fatalf("evicting the storage pod: %v", err) + } + done, err = r.perform(context.Background(), ops, stepReleasing) + if err != nil { + t.Fatalf("releasing: %v", err) + } + if !done { + t.Error("the step did not finish although the pod has left the worker") + } +} + +// The node comes back on the host it left, and a restart is not issued against +// one that is already back or already on its way. +func TestTheNodeIsRestartedOnceTheHostIsBack(t *testing.T) { + api := aControlPlane().reporting(nodeStatusOffline) + r, _ := anOpsWorld(t, api) + + done, err := r.perform(context.Background(), aWindow("a-window"), stepRestarting) + if err != nil { + t.Fatalf("restarting: %v", err) + } + if done { + t.Error("the step finished before the node had come back") + } + if asked := api.asked("RestartNode"); asked != 1 { + t.Errorf("RestartNode was issued %d time(s), want once", asked) + } + if api.restarts[0].NodeAddress != "" { + t.Errorf("the restart names address %q; the node is coming back on the host it left", + api.restarts[0].NodeAddress) + } + + restarting, _ := anOpsWorld(t, aControlPlane().reporting(nodeStatusInRestart)) + if done, err := restarting.perform(context.Background(), + aWindow("a-window"), stepRestarting); err != nil || done { + t.Errorf("done, err = %v, %v; a restart in flight is waited for", done, err) + } + + back := aControlPlane() + online, _ := anOpsWorld(t, back) + done, err = online.perform(context.Background(), aWindow("a-window"), stepRestarting) + if err != nil { + t.Fatalf("restarting: %v", err) + } + if !done { + t.Error("the step did not finish against a node that is online again") + } + if asked := back.asked("RestartNode"); asked != 0 { + t.Errorf("RestartNode was issued %d time(s) against a node already back", asked) + } +} + +// Cleanup takes the budget away, so the worker is drainable by the ordinary +// rules again. A budget left behind is what would make it undrainable forever. +func TestCleanupLeavesTheWorkerDrainableAgain(t *testing.T) { + r, apiClient := anOpsWorld(t, aControlPlane(), aReadyStoragePod(opsWorker)) + ops := aWindow("a-window") + + if _, err := r.perform(context.Background(), ops, stepShuttingDown); err != nil { + t.Fatalf("shutting down: %v", err) + } + done, err := r.perform(context.Background(), ops, stepCleanup) + if err != nil { + t.Fatalf("cleaning up: %v", err) + } + if !done { + t.Error("the step did not finish") + } + + if budget := budgetFor(t, apiClient, opsWorker); budget != nil { + t.Errorf("the budget %s outlived the window it belonged to", budget.Name) + } + var pod corev1.Pod + key := client.ObjectKey{Namespace: opsNamespace, Name: "storage-node-" + opsWorker} + if err := apiClient.Get(context.Background(), key, &pod); err != nil { + t.Fatalf("reading the storage pod: %v", err) + } + if _, labeled := pod.Labels[maintenanceLabel]; labeled { + t.Error("the pod still carries the window's label, which a later budget would select") + } +} + +// A step that belongs to another action is a hand-edited object or a downgrade, +// and neither resolves by reconciling again. +func TestAStepOfAnotherActionEndsTheWindow(t *testing.T) { + r, _ := anOpsWorld(t, aControlPlane()) + + _, err := r.performMaintenanceStep(context.Background(), aWindow("a-window"), stepPromoting) + + var fatal *terminalStepError + if !errors.As(err, &fatal) { + t.Errorf("err = %v, want the terminal kind for a step of another action", err) + } +} + +// budgetFor reads one window's budget, or reports that there is none. +func budgetFor( + t *testing.T, apiClient client.Client, worker string, +) *policyv1.PodDisruptionBudget { + t.Helper() + var budget policyv1.PodDisruptionBudget + key := client.ObjectKey{ + Namespace: opsNamespace, + Name: maintenanceBudgetName(opsCluster, worker), + } + err := apiClient.Get(context.Background(), key, &budget) + if apierrors.IsNotFound(err) { + return nil + } + if err != nil { + t.Fatalf("reading the maintenance budget: %v", err) + } + return &budget +} + +// kubeStep is a persisted step, which is what an operation somebody else is +// running carries in its status. +func kubeStep(s step) statemachine.KubeSnapshot { + return statemachine.KubeSnapshot{State: string(s)} +} diff --git a/operator/internal/controllers/node/migrate_test.go b/operator/internal/controllers/node/migrate_test.go new file mode 100644 index 000000000..833d2f3c6 --- /dev/null +++ b/operator/internal/controllers/node/migrate_test.go @@ -0,0 +1,438 @@ +// Relocating a node onto another worker. +// +// A migration moves the SPDK process and nothing else: the node keeps its backend +// UUID, its partitions, and its logical volumes, and what changes is the machine +// it runs on. The four steps are ordered the way they are because each of them +// guards the next — the target is put into the storage plane and its name made +// resolvable before the restart is aimed at it, the departure is observed before +// the promote, and the Kubernetes view is re-pointed only once the promote has +// landed. +// +// The promote is the point of no return, so the things that can be refused are +// refused before it: a target that is not a node of this cluster, one that is not +// Ready, and a node that is not back online. +// +// design-storagenode.md §9. + +package node + +import ( + "context" + "errors" + "testing" + + corev1 "k8s.io/api/core/v1" + discoveryv1 "k8s.io/api/discovery/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + + atlaskube "github.com/simplyblock/atlas/kube" + "github.com/simplyblock/atlas/ptr" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/utils" +) + +// aRelocation is the operation these cases run, aimed at opsTarget. +func aRelocation() *simplyblockv1alpha2.StorageNodeOps { + ops := anOperation("a-relocation", simplyblockv1alpha2.StorageNodeOpsActionMigrate) + ops.Spec.Migrate = &simplyblockv1alpha2.MigrateSpec{TargetWorkerNode: opsTarget} + return ops +} + +// aReadyStoragePod is the storage-node pod on one worker, running and ready, +// which is what Preparing waits for before it asks about DNS. +func aReadyStoragePod(worker string) *corev1.Pod { + pod := &corev1.Pod{ + ObjectMeta: metav1.ObjectMeta{ + Name: "storage-node-" + worker, + Namespace: opsNamespace, + Labels: map[string]string{ + atlaskube.LabelApp: atlaskube.AppStorageNode, + atlaskube.LabelStorageNodeSet: opsCluster, + }, + }, + Spec: corev1.PodSpec{NodeName: worker}, + Status: corev1.PodStatus{ + Phase: corev1.PodRunning, + Conditions: []corev1.PodCondition{{Type: corev1.PodReady, Status: corev1.ConditionTrue}}, + }, + } + return pod +} + +// aPublishedName is the EndpointSlice entry the control plane resolves the +// worker's node_address through. +func aPublishedName(workers ...string) *discoveryv1.EndpointSlice { + slice := &discoveryv1.EndpointSlice{ + ObjectMeta: metav1.ObjectMeta{ + Name: atlaskube.StorageNodeSetAPIEndpointSliceName(opsCluster), + Namespace: opsNamespace, + }, + AddressType: discoveryv1.AddressTypeIPv4, + } + for _, worker := range workers { + hostname := utils.NodeHostnameLabel(worker) + slice.Endpoints = append(slice.Endpoints, discoveryv1.Endpoint{ + Hostname: &hostname, + Addresses: []string{"10.0.0.2"}, + }) + } + return slice +} + +// aPerNodeConfig is the cluster's per-node configuration, holding one entry for +// the worker the node is leaving. +func aPerNodeConfig(entry string) *corev1.ConfigMap { + return &corev1.ConfigMap{ + ObjectMeta: metav1.ObjectMeta{ + Name: PerNodeConfigMapName(opsCluster), + Namespace: opsNamespace, + }, + Data: map[string]string{opsWorker: entry}, + } +} + +// An operation naming nowhere to go cannot be run, and no number of passes will +// give it a target. +func TestARelocationWithNoTargetEndsTheOperation(t *testing.T) { + r, _ := anOpsWorld(t, aControlPlane()) + ops := anOperation("a-relocation", simplyblockv1alpha2.StorageNodeOpsActionMigrate) + + _, err := r.perform(context.Background(), ops, stepPreparing) + + var fatal *terminalStepError + if !errors.As(err, &fatal) { + t.Errorf("err = %v, want the terminal kind for an operation with nowhere to relocate to", err) + } +} + +// A node already on the target is what a re-run of a finished migration looks +// like. Saying so beats issuing a restart that moves a node onto the host it is +// already on. +func TestARelocationToTheHostTheNodeIsOnIsAlreadyDone(t *testing.T) { + node := anOpsNode() + node.Spec.WorkerNode = opsTarget + r, apiClient := anOpsWorld(t, aControlPlane()) + replaceNode(t, apiClient, node) + + done, err := r.perform(context.Background(), aRelocation(), stepPreparing) + if err != nil { + t.Fatalf("preparing: %v", err) + } + if !done { + t.Error("the step did not finish against a node already on the target host") + } +} + +// A target that is not a worker of this Kubernetes cluster, or one that is not +// Ready, is refused rather than waited on: both are the operation naming +// somewhere it cannot go. +func TestATargetThatCannotHostTheNodeIsRefused(t *testing.T) { + t.Run("not a worker of this cluster", func(t *testing.T) { + r, _ := anOpsWorld(t, aControlPlane()) + + _, err := r.perform(context.Background(), aRelocation(), stepPreparing) + + var fatal *terminalStepError + if !errors.As(err, &fatal) { + t.Errorf("err = %v, want the terminal kind for a target that does not exist", err) + } + }) + + t.Run("not Ready", func(t *testing.T) { + r, _ := anOpsWorld(t, aControlPlane(), aWorker(opsTarget, false)) + + _, err := r.perform(context.Background(), aRelocation(), stepPreparing) + + var fatal *terminalStepError + if !errors.As(err, &fatal) { + t.Errorf("err = %v, want the terminal kind for a target that is not Ready", err) + } + }) +} + +// Preparing writes the target's configuration and labels it into the storage +// plane, in that order: the entry has to exist by the time the pod's init +// container sources it. +func TestPreparingWritesTheTargetsConfigurationAndLabelsIt(t *testing.T) { + config := aPerNodeConfig("MAX_SUBSYS_COUNT=10\nPCI_ALLOWED='0000:02:00.0'\n") + r, apiClient := anOpsWorld(t, aControlPlane(), aWorker(opsTarget, true), config) + + ops := aRelocation() + ops.Spec.Migrate.NewSsdPcie = []string{"0000:04:00.0"} + + done, err := r.perform(context.Background(), ops, stepPreparing) + if err != nil { + t.Fatalf("preparing: %v", err) + } + if done { + t.Error("the step finished before the target's pod was ready") + } + + var written corev1.ConfigMap + key := client.ObjectKey{Namespace: opsNamespace, Name: PerNodeConfigMapName(opsCluster)} + if err := apiClient.Get(context.Background(), key, &written); err != nil { + t.Fatalf("reading the per-node configuration: %v", err) + } + entry, cloned := written.Data[opsTarget] + if !cloned { + t.Fatal("the target has no per-node configuration, so its init container would find none") + } + if entry != "MAX_SUBSYS_COUNT=10\nPCI_ALLOWED='0000:02:00.0,0000:04:00.0'\n" { + t.Errorf("the cloned entry is %q, want the source's with the migration's drives merged in", + entry) + } + + var worker corev1.Node + if err := apiClient.Get(context.Background(), client.ObjectKey{Name: opsTarget}, &worker); err != nil { + t.Fatalf("reading the target worker: %v", err) + } + if worker.Labels[atlaskube.LabelStorageNodeSet] != opsCluster { + t.Errorf("the target carries %v, want the cluster's storage-plane label", worker.Labels) + } +} + +// Preparing blocks on the published name rather than on pod readiness. The +// control plane resolves node_address itself, and readiness happens before the +// EndpointSlice is published, so a restart issued on readiness alone is aimed at +// a name that does not yet exist. +func TestPreparingWaitsForTheNameToResolveAndNotOnlyForThePod(t *testing.T) { + config := aPerNodeConfig("PCI_ALLOWED=''\n") + + unpublished, _ := anOpsWorld(t, aControlPlane(), + aWorker(opsTarget, true), config, aReadyStoragePod(opsTarget)) + done, err := unpublished.perform(context.Background(), aRelocation(), stepPreparing) + if err != nil { + t.Fatalf("preparing: %v", err) + } + if done { + t.Error("the step finished against a name the control plane could not resolve") + } + + published, _ := anOpsWorld(t, aControlPlane(), aWorker(opsTarget, true), config, + aReadyStoragePod(opsTarget), aPublishedName(opsTarget)) + done, err = published.perform(context.Background(), aRelocation(), stepPreparing) + if err != nil { + t.Fatalf("preparing: %v", err) + } + if !done { + t.Error("the step did not finish although the target's name is published") + } +} + +// The restart is aimed at the target's own name, and it is forced by default: a +// migration relocates a node that is still online, and the control plane refuses +// a non-forced restart of one that is not already offline. +func TestTheRelocationRestartIsAimedAtTheTargetAndForced(t *testing.T) { + api := aControlPlane() + r, _ := anOpsWorld(t, api, aWorker(opsTarget, true)) + + done, err := r.perform(context.Background(), aRelocation(), stepRelocating) + if err != nil { + t.Fatalf("relocating: %v", err) + } + if done { + t.Error("the step finished while the node was still reporting online") + } + if len(api.restarts) != 1 { + t.Fatalf("%d restarts were issued, want one", len(api.restarts)) + } + if api.restarts[0].NodeAddress != utils.StorageNodeSetAPIAddress(opsTarget, opsNamespace) { + t.Errorf("the restart names %q, want the target's per-pod address", + api.restarts[0].NodeAddress) + } + if !api.restarts[0].Force { + t.Error("the restart is not forced, and the control plane refuses it against an online node") + } +} + +// An operation that states the flag means it, including to turn the default off. +func TestAStatedForceOutranksTheRelocationsDefault(t *testing.T) { + api := aControlPlane() + r, _ := anOpsWorld(t, api, aWorker(opsTarget, true)) + ops := aRelocation() + ops.Spec.Force = ptr.To(false) + + if _, err := r.perform(context.Background(), ops, stepRelocating); err != nil { + t.Fatalf("relocating: %v", err) + } + if api.restarts[0].Force { + t.Error("the operation asked for an unforced restart and was given a forced one") + } +} + +// A node that has left online has begun its restart, which is the observation +// this step exists to make. There is nothing left to issue. +func TestTheRelocationIsOverWhenTheNodeHasLeftOnline(t *testing.T) { + api := aControlPlane().reporting(nodeStatusInRestart) + r, _ := anOpsWorld(t, api, aWorker(opsTarget, true)) + + done, err := r.perform(context.Background(), aRelocation(), stepRelocating) + if err != nil { + t.Fatalf("relocating: %v", err) + } + if !done { + t.Error("the step did not finish although the node had left online") + } + if asked := api.asked("RestartNode"); asked != 0 { + t.Errorf("a second restart was issued %d time(s) against a node already restarting", asked) + } +} + +// The wait for the node is over when it is online again, and not while it is on +// its way there. +func TestTheWaitForTheRelocatedNodeEndsWhenItIsBack(t *testing.T) { + restarting, _ := anOpsWorld(t, aControlPlane().reporting(nodeStatusInRestart), + aWorker(opsTarget, true)) + done, err := restarting.perform(context.Background(), aRelocation(), stepAwaitingNode) + if err != nil { + t.Fatalf("awaiting the node: %v", err) + } + if done { + t.Error("the wait finished against a node still restarting") + } + + back, _ := anOpsWorld(t, aControlPlane(), aWorker(opsTarget, true)) + done, err = back.perform(context.Background(), aRelocation(), stepAwaitingNode) + if err != nil { + t.Fatalf("awaiting the node: %v", err) + } + if !done { + t.Error("the wait did not finish against a node that is online again") + } +} + +// The promote is separately guarded, which is what makes the negative predicate +// in Relocating tolerable: promoting into an in-flight restart leaves the +// relocated devices stuck. +func TestAPromoteIsRefusedWhileTheNodeIsNotBack(t *testing.T) { + api := aControlPlane().reporting(nodeStatusInRestart) + r, _ := anOpsWorld(t, api, aWorker(opsTarget, true)) + + done, err := r.perform(context.Background(), aRelocation(), stepPromoting) + if err != nil { + t.Fatalf("promoting: %v", err) + } + if done { + t.Error("the step finished against a node that has not come back") + } + if asked := api.asked("Promote"); asked != 0 { + t.Errorf("Promote was issued %d time(s) into an in-flight restart", asked) + } +} + +// The promote lands and the Kubernetes view follows it, never the other way +// round: re-pointing first would leave the object describing a relocation the +// control plane had not performed. +func TestThePromoteIsFollowedByTheTopologyRepoint(t *testing.T) { + api := aControlPlane() + r, apiClient := anOpsWorld(t, api, aWorker(opsTarget, true)) + ops := aRelocation() + ops.Spec.Migrate.NewSsdPcie = []string{"0000:04:00.0"} + + done, err := r.perform(context.Background(), ops, stepPromoting) + if err != nil { + t.Fatalf("promoting: %v", err) + } + if !done { + t.Error("the step did not finish although the promote landed") + } + if asked := api.asked("Promote"); asked != 1 { + t.Errorf("Promote was issued %d time(s), want once", asked) + } + + var node simplyblockv1alpha2.StorageNode + key := client.ObjectKey{Namespace: opsNamespace, Name: opsNodeName} + if err := apiClient.Get(context.Background(), key, &node); err != nil { + t.Fatalf("reading the node: %v", err) + } + if node.Spec.WorkerNode != opsTarget { + t.Errorf("the node still names worker %q, so Kubernetes and the control plane disagree", + node.Spec.WorkerNode) + } + if len(node.Spec.Config.PcieAllowList) != 1 || + node.Spec.Config.PcieAllowList[0] != "0000:04:00.0" { + t.Errorf("the allow list is %v, want the drives the migration bound on the target", + node.Spec.Config.PcieAllowList) + } +} + +// A node whose object already names the target was promoted by an earlier pass. +// The re-point is the last thing the step does, so its presence is the record +// that the promote landed, and a second one must not be issued. +func TestAPromoteThatAlreadyLandedIsNotIssuedAgain(t *testing.T) { + node := anOpsNode() + node.Spec.WorkerNode = opsTarget + api := aControlPlane() + r, apiClient := anOpsWorld(t, api, aWorker(opsTarget, true)) + replaceNode(t, apiClient, node) + + done, err := r.migratePromote(context.Background(), aRelocation(), node, + opsTarget, opsClusterID, opsNodeID) + if err != nil { + t.Fatalf("promoting: %v", err) + } + if !done { + t.Error("the step did not finish against a node that has already been promoted") + } + if asked := api.asked("Promote"); asked != 0 { + t.Errorf("Promote was issued %d time(s) against a node already on its target", asked) + } +} + +// The drives a migration bound are added to what the node already had, in the +// order somebody wrote them: the field is a user's, and one the operator +// rewrites should come back recognizable. +func TestTheBoundDrivesJoinTheListInTheOrderItWasWritten(t *testing.T) { + merged := mergePCIAddresses( + []string{"0000:02:00.0", "0000:03:00.0"}, + []string{"0000:03:00.0", "", "0000:04:00.0"}) + + want := []string{"0000:02:00.0", "0000:03:00.0", "0000:04:00.0"} + if len(merged) != len(want) { + t.Fatalf("merged = %v, want %v", merged, want) + } + for i := range want { + if merged[i] != want[i] { + t.Fatalf("merged = %v, want %v", merged, want) + } + } +} + +// Nothing to add leaves the list exactly as it was. +func TestAMigrationThatBindsNoDriveLeavesTheListAlone(t *testing.T) { + existing := []string{"0000:02:00.0"} + if merged := mergePCIAddresses(existing, nil); len(merged) != 1 { + t.Errorf("merged = %v, want the list untouched", merged) + } +} + +// A worker with no Ready condition at all has not reported one, which is not the +// same as having reported that it is Ready. +func TestAWorkerThatHasNotReportedIsNotReady(t *testing.T) { + if workerReady(&corev1.Node{}) { + t.Error("a worker with no conditions was read as Ready") + } + if !workerReady(aWorker(opsWorker, true)) { + t.Error("a Ready worker was read as not Ready") + } +} + +// replaceNode swaps the fixture's node for one a case has rewritten. +func replaceNode(t *testing.T, apiClient client.Client, node *simplyblockv1alpha2.StorageNode) { + t.Helper() + existing := anOpsNode() + if err := apiClient.Delete(context.Background(), existing); err != nil { + t.Fatalf("clearing the fixture's node: %v", err) + } + status := *node.Status.DeepCopy() + node.ResourceVersion = "" + if err := apiClient.Create(context.Background(), node); err != nil { + t.Fatalf("seeding the case's node: %v", err) + } + node.Status = status + if err := apiClient.Status().Update(context.Background(), node); err != nil { + t.Fatalf("seeding the case's node status: %v", err) + } +} diff --git a/operator/internal/controllers/node/opslock_test.go b/operator/internal/controllers/node/opslock_test.go new file mode 100644 index 000000000..b45be956f --- /dev/null +++ b/operator/internal/controllers/node/opslock_test.go @@ -0,0 +1,284 @@ +// The node's lock, and the three paths that let go of it. +// +// One operation touches one node at a time, and the lock that guarantees it is a +// field on the node rather than anything this controller holds in memory: two +// operator replicas, or one replica across a restart, have to reach the same +// answer. Acquisition is an optimistic-lock patch for the same reason — two +// operations can both read an empty field and both conclude the lock is free, and +// only one of them can have written it. +// +// Releasing is the half that goes wrong quietly. A release that does not check +// ownership clears a lock somebody else now holds, which is worse than never +// releasing: two operations then run against one node with neither knowing. And a +// lock nothing releases leaves a node held by an operation that has finished or +// has been deleted, which is why the release happens on the terminal transition, +// again on every pass of a terminal operation, and again from the finalizer. +// +// design-storagenode.md §7.1 and §11. + +package node + +import ( + "context" + "testing" + + "k8s.io/client-go/tools/events" + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// lockHolder is what the node currently says holds it. +func lockHolder(t *testing.T, apiClient client.Client) string { + t.Helper() + var node simplyblockv1alpha2.StorageNode + key := client.ObjectKey{Namespace: opsNamespace, Name: opsNodeName} + if err := apiClient.Get(context.Background(), key, &node); err != nil { + t.Fatalf("reading the node: %v", err) + } + return node.Status.ActiveOpsRef +} + +// operationRead is the operation as the API server now holds it. +func operationRead( + t *testing.T, apiClient client.Client, name string, +) *simplyblockv1alpha2.StorageNodeOps { + t.Helper() + var ops simplyblockv1alpha2.StorageNodeOps + key := client.ObjectKey{Namespace: opsNamespace, Name: name} + if err := apiClient.Get(context.Background(), key, &ops); err != nil { + t.Fatalf("reading the operation: %v", err) + } + return &ops +} + +// held marks the node as locked by the named operation. +func lockedBy(t *testing.T, apiClient client.Client, name string) { + t.Helper() + var node simplyblockv1alpha2.StorageNode + key := client.ObjectKey{Namespace: opsNamespace, Name: opsNodeName} + if err := apiClient.Get(context.Background(), key, &node); err != nil { + t.Fatalf("reading the node: %v", err) + } + node.Status.ActiveOpsRef = name + if err := apiClient.Status().Update(context.Background(), &node); err != nil { + t.Fatalf("seeding the lock: %v", err) + } +} + +// Taking the lock is what moves the operation out of Pending, and both halves +// have to be visible: the node says who holds it, and the operation says it is +// running and since when. +func TestTakingTheLockIsWhatStartsTheOperation(t *testing.T) { + ops := anOperation("a-suspend", simplyblockv1alpha2.StorageNodeOpsActionSuspend) + r, apiClient := anOpsWorld(t, aControlPlane(), ops) + + acquired, err := r.acquireLock(context.Background(), ops) + if err != nil { + t.Fatalf("acquiring the lock: %v", err) + } + if !acquired { + t.Fatal("the operation did not take a lock nothing else holds") + } + if holder := lockHolder(t, apiClient); holder != "a-suspend" { + t.Errorf("the node is held by %q, want the operation that took it", holder) + } + + got := operationRead(t, apiClient, "a-suspend") + if got.Status.Phase != simplyblockv1alpha2.StorageNodeOpsPhaseRunning { + t.Errorf("phase = %q, want Running", got.Status.Phase) + } + if got.Status.StartedAt == nil { + t.Error("nothing records when the operation started, which its duration is measured from") + } +} + +// An operation whose node is held waits, and says so. It must not take the lock, +// and it must not fail: the holder finishing is what wakes it. +func TestAnOperationWaitsForTheOneHoldingTheNode(t *testing.T) { + ops := anOperation("a-suspend", simplyblockv1alpha2.StorageNodeOpsActionSuspend) + r, apiClient := anOpsWorld(t, aControlPlane(), ops) + lockedBy(t, apiClient, "somebody-elses-operation") + + acquired, err := r.acquireLock(context.Background(), ops) + if err != nil { + t.Fatalf("acquiring the lock: %v", err) + } + if acquired { + t.Fatal("the operation took a lock another operation holds") + } + if holder := lockHolder(t, apiClient); holder != "somebody-elses-operation" { + t.Errorf("the lock is now held by %q, and the holder was overwritten", holder) + } + + got := operationRead(t, apiClient, "a-suspend") + if got.Status.Phase != simplyblockv1alpha2.StorageNodeOpsPhasePending { + t.Errorf("phase = %q, want Pending: a queued operation has not started", got.Status.Phase) + } + if !announcedReason(r, OperationQueued) { + t.Error("nothing announced the wait, which is what tells it from a stalled controller") + } +} + +// The holder re-reading its own lock still holds it. Every pass of a running +// operation goes through this, so it has to be a no-op rather than a second +// acquisition. +func TestTheHolderKeepsItsOwnLock(t *testing.T) { + ops := anOperation("a-suspend", simplyblockv1alpha2.StorageNodeOpsActionSuspend) + r, apiClient := anOpsWorld(t, aControlPlane(), ops) + lockedBy(t, apiClient, "a-suspend") + + acquired, err := r.acquireLock(context.Background(), ops) + if err != nil { + t.Fatalf("acquiring the lock: %v", err) + } + if !acquired { + t.Error("the operation lost a lock it already held") + } +} + +// An operation against a node that is not there cannot be run, and holding it +// open would leave a record nobody can resolve. +func TestAnOperationAgainstAMissingNodeFails(t *testing.T) { + ops := anOperation("a-suspend", simplyblockv1alpha2.StorageNodeOpsActionSuspend) + ops.Spec.NodeRef = "no-such-node" + r, apiClient := anOpsWorld(t, aControlPlane(), ops) + + acquired, err := r.acquireLock(context.Background(), ops) + if err != nil { + t.Fatalf("acquiring the lock: %v", err) + } + if acquired { + t.Fatal("the operation claims to hold a node that does not exist") + } + + got := operationRead(t, apiClient, "a-suspend") + if got.Status.Phase != simplyblockv1alpha2.StorageNodeOpsPhaseFailed { + t.Errorf("phase = %q, want Failed", got.Status.Phase) + } + if got.Status.Message == "" { + t.Error("the failure says nothing about which node is missing") + } +} + +// A release only clears a lock that still names this operation. A pass that +// started before the lock changed hands would otherwise unlock a node somebody +// else is working on. +func TestALateReleaseDoesNotUnlockSomebodyElsesNode(t *testing.T) { + ops := anOperation("a-suspend", simplyblockv1alpha2.StorageNodeOpsActionSuspend) + r, apiClient := anOpsWorld(t, aControlPlane(), ops) + lockedBy(t, apiClient, "the-operation-that-took-it-next") + + if err := r.releaseLock(context.Background(), ops); err != nil { + t.Fatalf("releasing the lock: %v", err) + } + if holder := lockHolder(t, apiClient); holder != "the-operation-that-took-it-next" { + t.Errorf("the lock is now %q, and an operation cleared one it did not hold", holder) + } +} + +// Finishing writes the outcome and lets the node go, in that order: the release +// is what admits the next operation, and it must not admit one before the record +// of this operation is durable. +func TestFinishingWritesTheOutcomeAndLetsTheNodeGo(t *testing.T) { + ops := anOperation("a-suspend", simplyblockv1alpha2.StorageNodeOpsActionSuspend) + r, apiClient := anOpsWorld(t, aControlPlane(), ops) + lockedBy(t, apiClient, "a-suspend") + + _, err := r.finish(context.Background(), ops, + simplyblockv1alpha2.StorageNodeOpsPhaseSucceeded, "the node was suspended") + if err != nil { + t.Fatalf("finishing: %v", err) + } + + got := operationRead(t, apiClient, "a-suspend") + if got.Status.Phase != simplyblockv1alpha2.StorageNodeOpsPhaseSucceeded { + t.Errorf("phase = %q, want Succeeded", got.Status.Phase) + } + if got.Status.CompletedAt == nil { + t.Error("nothing records when the operation ended") + } + if holder := lockHolder(t, apiClient); holder != "" { + t.Errorf("the node is still held by %q after the operation finished", holder) + } +} + +// A terminal operation is a record, and a record does nothing — except let go of +// a lock it still holds. That covers the pass that crashed between persisting +// the phase and clearing the field, which would otherwise leave the node held by +// a finished operation for good. +func TestATerminalOperationStillReleasesALockItLeftBehind(t *testing.T) { + ops := anOperation("a-suspend", simplyblockv1alpha2.StorageNodeOpsActionSuspend) + ops.Finalizers = []string{OpsFinalizer} + ops.Status.Phase = simplyblockv1alpha2.StorageNodeOpsPhaseSucceeded + r, apiClient := anOpsWorld(t, aControlPlane(), ops) + lockedBy(t, apiClient, "a-suspend") + + if _, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: client.ObjectKeyFromObject(ops), + }); err != nil { + t.Fatalf("reconciling the finished operation: %v", err) + } + + if holder := lockHolder(t, apiClient); holder != "" { + t.Errorf("the node is held by %q, which has already finished", holder) + } +} + +// Deleting a running operation releases the node before the object goes. +// Without it, a `kubectl delete` mid-drain would leave the node locked by an +// object that no longer exists, and nothing would ever unlock it. +func TestDeletingAnOperationUnlocksTheNodeFirst(t *testing.T) { + ops := anOperation("a-suspend", simplyblockv1alpha2.StorageNodeOpsActionSuspend) + ops.Finalizers = []string{OpsFinalizer} + ops.Status.Phase = simplyblockv1alpha2.StorageNodeOpsPhaseRunning + r, apiClient := anOpsWorld(t, aControlPlane(), ops) + lockedBy(t, apiClient, "a-suspend") + + if err := apiClient.Delete(context.Background(), ops); err != nil { + t.Fatalf("deleting the operation: %v", err) + } + if _, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: client.ObjectKeyFromObject(ops), + }); err != nil { + t.Fatalf("reconciling the deleted operation: %v", err) + } + + if holder := lockHolder(t, apiClient); holder != "" { + t.Errorf("the node is held by %q, an operation that has been deleted", holder) + } + var gone simplyblockv1alpha2.StorageNodeOps + err := apiClient.Get(context.Background(), client.ObjectKeyFromObject(ops), &gone) + if err == nil && controllerutil.ContainsFinalizer(&gone, OpsFinalizer) { + t.Error("the finalizer is still there, so the object is held by its own teardown") + } +} + +// The first pass of a new operation puts the finalizer on, because the lock it +// is about to take has to be released even if the object is deleted mid-flight. +func TestAnOperationTakesItsFinalizerBeforeItTakesAnything(t *testing.T) { + ops := anOperation("a-suspend", simplyblockv1alpha2.StorageNodeOpsActionSuspend) + r, apiClient := anOpsWorld(t, aControlPlane(), ops) + + if _, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: client.ObjectKeyFromObject(ops), + }); err != nil { + t.Fatalf("reconciling: %v", err) + } + + got := operationRead(t, apiClient, "a-suspend") + if !controllerutil.ContainsFinalizer(got, OpsFinalizer) { + t.Error("the operation carries no finalizer, so a delete mid-flight would strand the lock") + } + if holder := lockHolder(t, apiClient); holder != "" { + t.Errorf("the node was locked by %q on the pass that only added the finalizer", holder) + } +} + +// announcedReason reports whether the operation raised one reason, which is what +// a queued operation owes whoever is wondering why nothing is happening. +func announcedReason(r *StorageNodeOpsReconciler, reason string) bool { + return announced(r.Recorder.(*events.FakeRecorder), reason) +} diff --git a/operator/internal/controllers/node/peertargets_test.go b/operator/internal/controllers/node/peertargets_test.go new file mode 100644 index 000000000..548f9dd81 --- /dev/null +++ b/operator/internal/controllers/node/peertargets_test.go @@ -0,0 +1,214 @@ +// Where a drain's movable volumes go. +// +// The choice is round-robin over the cluster's online peers, and the two +// properties that matter are spread and stability: concentrating a drained +// node's volumes on whichever peer sorts first refills one node with what +// another was holding, and an assignment that reshuffles between passes means a +// volume whose migration failed comes back against a different target for a +// reason nobody chose. +// +// A drain with no online peer holds rather than fails, because the condition is +// resolved by another node coming back and failing the operation would only mean +// starting it again afterward. +// +// design-storagenode.md §8.2. + +package node + +import ( + "context" + "errors" + "testing" + + "github.com/simplyblock/simplyblock-operator/internal/cpinformer" + "github.com/simplyblock/simplyblock-operator/internal/cpinformer/subscriptions" +) + +// movable is the census's movable half, one entry per PersistentVolume. +func movable(names ...string) []managedVolume { + volumes := make([]managedVolume, 0, len(names)) + for _, name := range names { + volumes = append(volumes, managedVolume{PVName: name, VolumeUUID: name + "-uuid"}) + } + return volumes +} + +// targetsOf assigns the volumes against the given control plane. +func targetsOf( + t *testing.T, api *scriptedControlPlane, volumes []managedVolume, +) (map[string]string, error) { + t.Helper() + r, _ := anOpsWorld(t, api) + return r.peerTargets(context.Background(), opsClusterID, opsNodeID, volumes) +} + +// Six volumes over two peers is three each, rather than six on the peer that +// happened to sort first. +func TestTheVolumesAreSpreadOverEveryOnlinePeer(t *testing.T) { + api := aControlPlane(). + withPeer(opsPeerID, nodeStatusOnline). + withPeer("node-3333", nodeStatusOnline) + + targets, err := targetsOf(t, api, movable("pv-a", "pv-b", "pv-c", "pv-d", "pv-e", "pv-f")) + if err != nil { + t.Fatalf("choosing targets: %v", err) + } + if len(targets) != 6 { + t.Fatalf("%d of 6 volumes were assigned a target", len(targets)) + } + + share := map[string]int{} + for _, target := range targets { + share[target]++ + } + for _, peer := range []string{opsPeerID, "node-3333"} { + if share[peer] != 3 { + t.Errorf("peer %s took %d of six volumes, want an even three", peer, share[peer]) + } + } +} + +// The node being drained is not somewhere to drain to. +func TestTheDrainedNodeIsNeverItsOwnTarget(t *testing.T) { + api := aControlPlane().withPeer(opsPeerID, nodeStatusOnline) + + targets, err := targetsOf(t, api, movable("pv-a", "pv-b")) + if err != nil { + t.Fatalf("choosing targets: %v", err) + } + for volume, target := range targets { + if target == opsNodeID { + t.Errorf("%s was assigned back to the node being drained", volume) + } + } +} + +// A peer that is not online cannot take a volume, so it is not offered one. +func TestAnOfflinePeerIsNotATarget(t *testing.T) { + api := aControlPlane(). + withPeer(opsPeerID, nodeStatusOffline). + withPeer("node-3333", nodeStatusOnline) + + targets, err := targetsOf(t, api, movable("pv-a", "pv-b")) + if err != nil { + t.Fatalf("choosing targets: %v", err) + } + for volume, target := range targets { + if target != "node-3333" { + t.Errorf("%s was assigned to %s, which is not online", volume, target) + } + } +} + +// A suspended peer is equally not a target: it accepts no new placement, which +// is the whole of what suspending it did. +func TestASuspendedPeerIsNotATarget(t *testing.T) { + api := aControlPlane().withPeer(opsPeerID, nodeStatusSuspended) + + _, err := targetsOf(t, api, movable("pv-a")) + + var blocked *blockedStepError + if !errors.As(err, &blocked) { + t.Fatalf("err = %v, want the drain held for want of a peer", err) + } +} + +// No online peer is a stall rather than a failure: another node coming back +// resolves it, and failing the drain would only mean starting it again. +func TestADrainWithNowhereToMoveToHolds(t *testing.T) { + _, err := targetsOf(t, aControlPlane(), movable("pv-a")) + + var blocked *blockedStepError + if !errors.As(err, &blocked) { + t.Fatalf("err = %v, want the drain held rather than failed", err) + } + if blocked.reason != NoMigrationTarget { + t.Errorf("the hold is announced as %q, want %q", blocked.reason, NoMigrationTarget) + } +} + +// The same volumes reach the same peers on every pass, so a migration that +// failed and is recreated is reassigned deliberately by the caller rather than +// by the peer list having reshuffled. +func TestTheAssignmentIsTheSameOnEveryPass(t *testing.T) { + api := aControlPlane(). + withPeer(opsPeerID, nodeStatusOnline). + withPeer("node-3333", nodeStatusOnline). + withPeer("node-4444", nodeStatusOnline) + volumes := movable("pv-a", "pv-b", "pv-c", "pv-d") + + first, err := targetsOf(t, api, volumes) + if err != nil { + t.Fatalf("choosing targets: %v", err) + } + second, err := targetsOf(t, api, volumes) + if err != nil { + t.Fatalf("choosing targets again: %v", err) + } + + for volume, target := range first { + if second[volume] != target { + t.Errorf("%s was assigned to %s and then to %s", volume, target, second[volume]) + } + } +} + +// Once the stream has delivered the cluster's snapshot, the peers are read from +// it and the control plane is not asked at all. +func TestThePeersComeFromTheStreamOnceItHasDelivered(t *testing.T) { + api := aControlPlane() + r, _ := anOpsWorld(t, api) + r.Nodes = &deliveredNodes{synced: true, nodes: []subscriptions.NodeDTO{ + {ID: opsNodeID, Status: nodeStatusOnline}, + {ID: opsPeerID, Status: nodeStatusOnline}, + }} + + targets, err := r.peerTargets(context.Background(), opsClusterID, opsNodeID, movable("pv-a")) + if err != nil { + t.Fatalf("choosing targets: %v", err) + } + if targets["pv-a"] != opsPeerID { + t.Errorf("pv-a was assigned to %q, want the peer the stream reported", targets["pv-a"]) + } + if asked := api.asked("StorageNodes"); asked != 0 { + t.Errorf("the control plane was asked %d time(s) for what the stream already holds", asked) + } +} + +// A cache that has not delivered its snapshot is not evidence of anything. An +// empty unsynced cache and a cluster with no peers look identical, and reading +// the first as the second would hold a drain that has every peer it needs. +func TestAnUndeliveredStreamFallsBackToTheControlPlane(t *testing.T) { + api := aControlPlane().withPeer(opsPeerID, nodeStatusOnline) + r, _ := anOpsWorld(t, api) + r.Nodes = &deliveredNodes{synced: false} + + targets, err := r.peerTargets(context.Background(), opsClusterID, opsNodeID, movable("pv-a")) + if err != nil { + t.Fatalf("choosing targets: %v", err) + } + if targets["pv-a"] != opsPeerID { + t.Errorf("pv-a was assigned to %q, want the peer the control plane reported", + targets["pv-a"]) + } +} + +// deliveredNodes is a storage-node cache, holding whatever the stream is said to +// have delivered and reporting synced only when it has. +type deliveredNodes struct { + nodes []subscriptions.NodeDTO + synced bool +} + +func (d *deliveredNodes) Lookup(nodeID string) (cpinformer.Scope, subscriptions.NodeDTO, bool) { + for _, node := range d.nodes { + if node.ID == nodeID { + return cpinformer.Scope{opsClusterID}, node, true + } + } + return nil, subscriptions.NodeDTO{}, false +} + +func (d *deliveredNodes) List(cpinformer.Scope) []subscriptions.NodeDTO { return d.nodes } + +func (d *deliveredNodes) Synced(cpinformer.Scope) bool { return d.synced } diff --git a/operator/internal/controllers/node/pernodeconfig_test.go b/operator/internal/controllers/node/pernodeconfig_test.go new file mode 100644 index 000000000..346c19023 --- /dev/null +++ b/operator/internal/controllers/node/pernodeconfig_test.go @@ -0,0 +1,149 @@ +// Cloning one worker's per-node configuration onto another. +// +// The entry is a shell-sourceable env file the storage-node pod's init container +// reads, and a relocation needs the target's to exist before the pod starts +// there: a pod that reaches the node configuration script with no entry fails +// with --max-subsys-count=0, a long way from the cause. +// +// The clone edits the rendered text rather than re-rendering from the node, +// because the node's own spec.config.pcieAllowList has not been rewritten yet — +// that happens after the promote, and this runs before the relocation. +// +// design-storagenode.md §5.3 and §9. + +package node + +import ( + "context" + "strings" + "testing" + + corev1 "k8s.io/api/core/v1" + "sigs.k8s.io/controller-runtime/pkg/client" +) + +// aRenderedEntry is one worker's entry as the ordinary pass writes it. +const aRenderedEntry = "MAX_SUBSYS_COUNT=10\n" + + "MAX_HUGE_PAGES_SIZE=''\n" + + "VCPU_COUNT=8\n" + + "PCI_ALLOWED='0000:02:00.0,0000:03:00.0'\n" + + "PCI_BLOCKED=''\n" + +// entryFor reads one worker's entry out of the cluster's per-node configuration. +func entryFor(t *testing.T, apiClient client.Client, worker string) (string, bool) { + t.Helper() + var configMap corev1.ConfigMap + key := client.ObjectKey{Namespace: opsNamespace, Name: PerNodeConfigMapName(opsCluster)} + if err := apiClient.Get(context.Background(), key, &configMap); err != nil { + t.Fatalf("reading the per-node configuration: %v", err) + } + entry, ok := configMap.Data[worker] + return entry, ok +} + +// The target's entry is the source's, with the drives the migration is binding +// merged into the allow list so the host binds them on start and they survive a +// later rebuild. +func TestTheTargetInheritsTheSourcesEntryWithTheNewDrives(t *testing.T) { + r, apiClient := anOpsWorld(t, aControlPlane(), aPerNodeConfig(aRenderedEntry)) + + err := r.Workload.CloneWorkerConfig(context.Background(), opsNamespace, opsCluster, + opsWorker, opsTarget, []string{"0000:03:00.0", "0000:04:00.0"}) + if err != nil { + t.Fatalf("cloning the configuration: %v", err) + } + + entry, cloned := entryFor(t, apiClient, opsTarget) + if !cloned { + t.Fatal("the target has no entry, so its init container would find none") + } + if !strings.Contains(entry, "PCI_ALLOWED='0000:02:00.0,0000:03:00.0,0000:04:00.0'") { + t.Errorf("the target's allow list is not the merge of both:\n%s", entry) + } + for _, line := range strings.Split(aRenderedEntry, "\n") { + if line == "" || strings.HasPrefix(line, "PCI_ALLOWED=") { + continue + } + if !strings.Contains(entry, line) { + t.Errorf("%q did not survive the clone:\n%s", line, entry) + } + } + + if source, _ := entryFor(t, apiClient, opsWorker); source != aRenderedEntry { + t.Errorf("the source's own entry was rewritten by the clone:\n%s", source) + } +} + +// A worker with no entry to clone from is an error rather than an empty entry +// written out: an empty one fails inside the pod, where the cause is hard to +// see. +func TestCloningFromAWorkerWithNoEntryIsRefused(t *testing.T) { + r, _ := anOpsWorld(t, aControlPlane(), aPerNodeConfig(aRenderedEntry)) + + err := r.Workload.CloneWorkerConfig(context.Background(), opsNamespace, opsCluster, + "worker-9", opsTarget, nil) + + if err == nil { + t.Error("cloning from a worker that has no configuration was accepted") + } +} + +// A cluster whose configuration has not been written yet has nothing to clone, +// which is not a failure: the ordinary pass writes the map, and the relocation +// holds until the target's pod is ready, which cannot happen before it exists. +func TestNothingToCloneFromIsNotAFailure(t *testing.T) { + r, _ := anOpsWorld(t, aControlPlane()) + + err := r.Workload.CloneWorkerConfig(context.Background(), opsNamespace, opsCluster, + opsWorker, opsTarget, nil) + + if err != nil { + t.Errorf("cloning before the configuration exists: %v", err) + } +} + +// An entry that predates the allow-list field gets one appended, which is what a +// node whose list was empty needs; the init container reads the last assignment. +func TestAnEntryWithNoAllowListGetsOne(t *testing.T) { + merged := mergeAllowedIntoEntry("MAX_SUBSYS_COUNT=10\n", []string{"0000:05:00.0"}) + + if !strings.Contains(merged, "MAX_SUBSYS_COUNT=10") { + t.Errorf("the entry lost what it already said:\n%s", merged) + } + if !strings.Contains(merged, "PCI_ALLOWED='0000:05:00.0'") { + t.Errorf("the drives the migration bound are not in the entry:\n%s", merged) + } +} + +// Nothing to add leaves the entry byte for byte as it was, so a clone of a +// migration binding no drive is the source's own text. +func TestAnEntryIsUntouchedWhenThereIsNothingToAdd(t *testing.T) { + if merged := mergeAllowedIntoEntry(aRenderedEntry, nil); merged != aRenderedEntry { + t.Errorf("the entry changed although nothing was added:\n%s", merged) + } +} + +// Reading a list back undoes the quoting the render applied, because the merge +// is a round trip through the text this package wrote. +func TestAListIsReadBackTheWayItWasWritten(t *testing.T) { + cases := map[string][]string{ + "'0000:02:00.0,0000:03:00.0'": {"0000:02:00.0", "0000:03:00.0"}, + "0000:02:00.0": {"0000:02:00.0"}, + "'a, b ,c'": {"a", "b", "c"}, + "''": nil, + "": nil, + } + for written, want := range cases { + got := parseShellList(written) + if len(got) != len(want) { + t.Errorf("parseShellList(%q) = %v, want %v", written, got, want) + continue + } + for i := range want { + if got[i] != want[i] { + t.Errorf("parseShellList(%q) = %v, want %v", written, got, want) + break + } + } + } +} diff --git a/operator/internal/controllers/node/provisioning_test.go b/operator/internal/controllers/node/provisioning_test.go new file mode 100644 index 000000000..8fb4ff24d --- /dev/null +++ b/operator/internal/controllers/node/provisioning_test.go @@ -0,0 +1,269 @@ +// The gates a node passes before it is added, and what the add carries. +// +// Adding a backend node is not idempotent and the call adds every socket of a +// worker at once, so the machine's job is to be sure exactly once that the add is +// both possible and necessary. Two of its steps divert to adoption for that +// reason: a deployment being taken over wholesale, and a backend node already at +// the worker's address, which is what a POST whose response was lost looks like +// from here. +// +// The parameters are the node describing itself. Every value but the subsystem +// cap comes from its own spec.config, which is what stopped a fleet default being +// the source of truth and the node a cache of it. +// +// design-storagenode.md §4.2 and §4.3. + +package node + +import ( + "context" + "errors" + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + + "github.com/simplyblock/atlas/ptr" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// anUnprovisionedNode is a node with no backend behind it yet, at the step being +// exercised. +func anUnprovisionedNode(at nodeStep) *simplyblockv1alpha2.StorageNode { + node := anOpsNode() + node.Status.UUID = "" + node.Status.Step.State = string(at) + return node +} + +// A cluster that requires a fault group holds a node that declares none, rather +// than rejecting it: the value can arrive later, and holding is what makes +// filling it in sufficient. +func TestANodeWithNoFaultGroupIsHeldRatherThanRefused(t *testing.T) { + cluster := anOpsCluster() + cluster.Spec.EnableFailureDomains = ptr.To(true) + r, _ := aSteadyNode(t, aControlPlane()) + + next, done, err := r.checkConfig(context.Background(), + anUnprovisionedNode(stepCheckingConfig), cluster) + + var blocked *blockedStepError + if !errors.As(err, &blocked) { + t.Fatalf("err = %v, want the node held for want of a fault group", err) + } + if blocked.reason != FailureDomainMissing { + t.Errorf("the hold is announced as %q, want %q", blocked.reason, FailureDomainMissing) + } + if done || next != stepCheckingConfig { + t.Errorf("next, done = %s, %v; a held node stays where it is", next, done) + } +} + +// A node that declares one passes, and so does every node of a cluster that does +// not ask for fault groups at all. +func TestANodeThatDeclaresItsFaultGroupPasses(t *testing.T) { + cluster := anOpsCluster() + cluster.Spec.EnableFailureDomains = ptr.To(true) + node := anUnprovisionedNode(stepCheckingConfig) + node.Spec.Config.FailureDomain = "rack-1" + r, _ := aSteadyNode(t, aControlPlane()) + + next, done, err := r.checkConfig(context.Background(), node, cluster) + if err != nil { + t.Fatalf("checking the configuration: %v", err) + } + if !done || next != stepAwaitingSlot { + t.Errorf("next, done = %s, %v; want the node moving on to wait for a slot", next, done) + } + + indifferent, done, err := r.checkConfig(context.Background(), + anUnprovisionedNode(stepCheckingConfig), anOpsCluster()) + if err != nil { + t.Fatalf("checking the configuration: %v", err) + } + if !done || indifferent != stepAwaitingSlot { + t.Errorf("a cluster that asks for no fault groups held a node that declares none") + } +} + +// A backend node already at the worker's address is adopted rather than added. +// That covers a POST whose response was lost after the control plane committed, +// as well as a node this operator never added. +func TestANodeAlreadyAtTheWorkersAddressIsAdopted(t *testing.T) { + api := aControlPlane() + r, _ := aSteadyNode(t, api, aWorker(opsWorker, true)) + + next, done, err := r.checkConfig(context.Background(), + anUnprovisionedNode(stepCheckingConfig), anOpsCluster()) + if err != nil { + t.Fatalf("checking the configuration: %v", err) + } + if !done || next != stepAdopting { + t.Errorf("next, done = %s, %v; want the node diverted to adoption", next, done) + } +} + +// A deployment being taken over wholesale diverts before the host check, because +// an adopted node is already running and its API answering is not this +// operator's precondition to establish. +func TestAnUpgradeAdoptionDivertsBeforeTheHostIsProbed(t *testing.T) { + secret := &corev1.Secret{ObjectMeta: metav1.ObjectMeta{ + Name: "simplyblock-" + opsCluster + "-upgrade", + Namespace: opsNamespace, + }} + r, _ := aSteadyNode(t, aControlPlane(), secret) + + next, done, err := r.checkHost(context.Background(), + anUnprovisionedNode(stepCheckingHost), anOpsCluster()) + if err != nil { + t.Fatalf("checking the host: %v", err) + } + if !done || next != stepAdopting { + t.Errorf("next, done = %s, %v; want the node diverted to adoption", next, done) + } +} + +// The add carries what the node says about itself, and the cluster contributes +// only what belongs to the fleet. +func TestTheAddCarriesWhatTheNodeSaysAboutItself(t *testing.T) { + node := anUnprovisionedNode(stepPosting) + node.Spec.Config.SpdkImage = "example.test/spdk:v1" + node.Spec.Config.SpdkProxyImage = "example.test/spdk-proxy:v1" + node.Spec.Config.SpdkSystemMemory = "4g" + node.Spec.Config.JournalManager = &simplyblockv1alpha2.JournalManagerSpec{ + PercentPerDevice: ptr.To(int32(5)), Count: ptr.To(int32(2)), + } + + cluster := anOpsCluster() + cluster.Spec.StorageNodes = &simplyblockv1alpha2.StorageNodesSpec{ + MgmtInterface: "eth0", + DataInterfaces: []string{"eth1"}, + EnableFormat4K: ptr.To(true), + } + + r, _ := aSteadyNode(t, aControlPlane()) + params := r.addParams(node, cluster) + + if params.SPDKImage != "example.test/spdk:v1" || + params.SPDKProxyImage != "example.test/spdk-proxy:v1" || + params.SpdkSystemMemory != "4g" { + t.Errorf("params = %+v, want the images and memory the node declares", params) + } + if params.JMPercent != 5 || params.HaJMCount != 2 { + t.Errorf("journal = %d%% over %d, want what the node declares", + params.JMPercent, params.HaJMCount) + } + if params.InterfaceName != "eth0" || len(params.DataNics) != 1 { + t.Errorf("params = %+v, want the interfaces the fleet declares", params) + } + if !params.Format4K { + t.Error("the fleet asked for 4K formatting and the add does not carry it") + } + if params.CRName != opsCluster || params.CRNameSpace != opsNamespace || + params.CRPlural != "storageclusters" { + t.Errorf("params = %+v, want the cluster object the node belongs to", params) + } + if params.NodeAddress == "" { + t.Error("the add names no address, which is what the control plane resolves the node by") + } +} + +// A node that declares no journal settings is added with the defaults the +// control plane was always given, rather than with zeros. +func TestAnUnstatedJournalIsTheDefaultRatherThanNothing(t *testing.T) { + r, _ := aSteadyNode(t, aControlPlane()) + + params := r.addParams(anUnprovisionedNode(stepPosting), anOpsCluster()) + + if params.JMPercent != 3 || params.HaJMCount != 3 { + t.Errorf("journal = %d%% over %d, want the defaults 3 and 3", + params.JMPercent, params.HaJMCount) + } +} + +// The two vocabularies differ: this API names a fault group after the rack or +// the power feed somebody would say out loud, and the control plane indexes one. +// A label seeded from an index sends its number, and a name is left to the +// control plane to assign. +func TestOnlyAFaultGroupThatIsANumberIsSentToTheControlPlane(t *testing.T) { + r, _ := aSteadyNode(t, aControlPlane()) + + numbered := anUnprovisionedNode(stepPosting) + numbered.Spec.Config.FailureDomain = "2" + if index := r.addParams(numbered, anOpsCluster()).FailureDomain; index == nil || *index != 2 { + t.Errorf("failureDomain = %v, want the index the label spells", index) + } + + named := anUnprovisionedNode(stepPosting) + named.Spec.Config.FailureDomain = "rack-1" + if index := r.addParams(named, anOpsCluster()).FailureDomain; index != nil { + t.Errorf("failureDomain = %v, want none: the control plane has no field for a name", + index) + } + + if index := r.addParams(anUnprovisionedNode(stepPosting), anOpsCluster()).FailureDomain; index != nil { + t.Errorf("failureDomain = %v, want none for a node that declares no group", index) + } +} + +// Posting is the one call the machine makes, and the claim that guards it was +// made by the transition into the step rather than by anything here. +func TestPostingIssuesTheAddOnce(t *testing.T) { + api := aControlPlane() + r, _ := aSteadyNode(t, api) + + next, done, err := r.performNodeStep(context.Background(), + anUnprovisionedNode(stepPosting), anOpsCluster(), stepPosting) + if err != nil { + t.Fatalf("posting: %v", err) + } + if !done || next != stepResolving { + t.Errorf("next, done = %s, %v; want the node waiting for the UUID the add produces", + next, done) + } + if asked := api.asked("AddNode"); asked != 1 { + t.Errorf("AddNode was issued %d time(s), want once", asked) + } +} + +// A step no provisioning path declares is a downgrade or a hand-edited object, +// and reconciling again does not resolve either. +func TestAStepNoProvisioningPathDeclaresIsRefused(t *testing.T) { + r, _ := aSteadyNode(t, aControlPlane()) + + _, _, err := r.performNodeStep(context.Background(), + anUnprovisionedNode(stepPosting), anOpsCluster(), nodeStep("Nowhere")) + + if err == nil { + t.Error("a step that belongs to no provisioning path was accepted") + } +} + +// A cluster the control plane has not created yet is nothing to add a node to, +// so the node holds and says which of the two it is waiting for. +func TestANodeWaitsForItsClusterToExistInTheControlPlane(t *testing.T) { + r, apiClient := aSteadyNode(t, aControlPlane()) + + var cluster simplyblockv1alpha2.StorageCluster + key := client.ObjectKey{Namespace: opsNamespace, Name: opsCluster} + if err := apiClient.Get(context.Background(), key, &cluster); err != nil { + t.Fatalf("reading the cluster: %v", err) + } + cluster.Status.UUID = "" + if err := apiClient.Status().Update(context.Background(), &cluster); err != nil { + t.Fatalf("clearing the cluster's UUID: %v", err) + } + node := nodeRead(t, apiClient) + node.Status.UUID = "" + if err := apiClient.Status().Update(context.Background(), node); err != nil { + t.Fatalf("clearing the node's UUID: %v", err) + } + + settle(t, r) + + if message := nodeRead(t, apiClient).Status.Message; message == "" { + t.Error("the node says nothing about what it is waiting for") + } +} diff --git a/operator/internal/controllers/node/remove_fanout_test.go b/operator/internal/controllers/node/remove_fanout_test.go index 41d64af50..d3a41aeb5 100644 --- a/operator/internal/controllers/node/remove_fanout_test.go +++ b/operator/internal/controllers/node/remove_fanout_test.go @@ -47,7 +47,7 @@ func aFanOut(t *testing.T, mover func(client.Client, *runtime.Scheme) vmigration scheme := testsupport.NewScheme(t) node := &simplyblockv1alpha2.StorageNode{ - ObjectMeta: metav1.ObjectMeta{Name: "a-node", Namespace: "simplyblock"}, + ObjectMeta: metav1.ObjectMeta{Name: opsNodeName, Namespace: "simplyblock"}, Spec: simplyblockv1alpha2.StorageNodeSpec{ClusterRef: "a-cluster"}, } node.Status.UUID = aDrainedNodeID diff --git a/operator/internal/controllers/node/remove_gone_test.go b/operator/internal/controllers/node/remove_gone_test.go index 4f154a7a8..b31017630 100644 --- a/operator/internal/controllers/node/remove_gone_test.go +++ b/operator/internal/controllers/node/remove_gone_test.go @@ -60,7 +60,7 @@ func aRemover(t *testing.T) *StorageNodeOpsReconciler { scheme := testsupport.NewScheme(t) node := &simplyblockv1alpha2.StorageNode{ - ObjectMeta: metav1.ObjectMeta{Name: "a-node", Namespace: "simplyblock"}, + ObjectMeta: metav1.ObjectMeta{Name: opsNodeName, Namespace: "simplyblock"}, Spec: simplyblockv1alpha2.StorageNodeSpec{ClusterRef: "a-cluster"}, } node.Status.UUID = "node-uuid" diff --git a/operator/internal/controllers/node/rotation_test.go b/operator/internal/controllers/node/rotation_test.go new file mode 100644 index 000000000..664987af2 --- /dev/null +++ b/operator/internal/controllers/node/rotation_test.go @@ -0,0 +1,115 @@ +// What makes the storage-node pods restart when their certificate changes. +// +// The serving certificate is mounted from a Secret, and a mounted Secret is +// updated in place: the file under the pod changes and the process holding it +// keeps serving the old one. Nothing about that is visible from the DaemonSet, +// whose template is identical before and after the rotation, so the pods are +// never rolled and the API keeps presenting an expired certificate. +// +// Stamping the Secret's resourceVersion into the pod template is what turns a +// rotation into a template change, which is the one thing a DaemonSet does roll +// on. It is stamped only where TLS is served, because a deployment that serves +// plaintext has no certificate to rotate. + +package node + +import ( + "context" + "testing" + + appsv1 "k8s.io/api/apps/v1" + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + + "github.com/simplyblock/simplyblock-operator/internal/utils" +) + +// aServingSecret is the storage-node API's certificate, as cert-manager or the +// operator's own rotator writes it. +func aServingSecret() *corev1.Secret { + return &corev1.Secret{ObjectMeta: metav1.ObjectMeta{ + Name: utils.SecretNameStorageNodeSetAPITLS, + Namespace: enrollNamespace, + }} +} + +// templateRevision is what the pod template says the certificate was at. +func templateRevision(t *testing.T, apiClient client.Client) string { + t.Helper() + var daemonSet appsv1.DaemonSet + var sets appsv1.DaemonSetList + if err := apiClient.List(context.Background(), &sets, + client.InNamespace(enrollNamespace)); err != nil { + t.Fatalf("listing the storage-node workload: %v", err) + } + if len(sets.Items) != 1 { + t.Fatalf("%d DaemonSets were written, want the cluster's one", len(sets.Items)) + } + daemonSet = sets.Items[0] + return daemonSet.Spec.Template.Annotations[utils.AnnotationTLSSecretRevision] +} + +// A rotation changes the pod template, which is what a DaemonSet rolls on. +func TestARotatedCertificateRollsTheStoragePods(t *testing.T) { + secret := aServingSecret() + r := aWorkloadReconciler(t, aSizedCluster(), secret) + r.TLSEnabled = true + + if err := r.reconcileDaemonSet(context.Background(), aSizedCluster()); err != nil { + t.Fatalf("writing the workload: %v", err) + } + before := templateRevision(t, r.Client) + if before == "" { + t.Fatal("the template records no certificate revision, so a rotation rolls nothing") + } + + // Rewriting the Secret is what a rotation is, and it is the only change. + var rotated corev1.Secret + key := client.ObjectKeyFromObject(secret) + if err := r.Get(context.Background(), key, &rotated); err != nil { + t.Fatalf("reading the certificate: %v", err) + } + rotated.Data = map[string][]byte{"tls.crt": []byte("a fresh certificate")} + if err := r.Update(context.Background(), &rotated); err != nil { + t.Fatalf("rotating the certificate: %v", err) + } + + if err := r.reconcileDaemonSet(context.Background(), aSizedCluster()); err != nil { + t.Fatalf("writing the workload: %v", err) + } + if after := templateRevision(t, r.Client); after == before { + t.Errorf("the template still records %q, so the pods keep serving the old certificate", + after) + } +} + +// A deployment that serves plaintext has no certificate to rotate, so nothing is +// stamped and no pod is rolled by the absence. +func TestWithoutTLSNothingIsStampedOnTheTemplate(t *testing.T) { + r := aWorkloadReconciler(t, aSizedCluster(), aServingSecret()) + + if err := r.reconcileDaemonSet(context.Background(), aSizedCluster()); err != nil { + t.Fatalf("writing the workload: %v", err) + } + + if revision := templateRevision(t, r.Client); revision != "" { + t.Errorf("the template records certificate revision %q on a deployment that serves "+ + "plaintext", revision) + } +} + +// A cluster that states no image takes the ControlPlane's, so a deployment +// states its version once. A cluster that has neither is an error rather than a +// workload written with an empty image, which schedules pods that cannot start. +func TestAWorkloadWithNoImageAnywhereIsRefused(t *testing.T) { + r := aWorkloadReconciler(t, aSizedCluster()) + cluster := aSizedCluster() + cluster.Spec.StorageNodes.Image = "" + + err := r.reconcileDaemonSet(context.Background(), cluster) + + if err == nil { + t.Error("a workload was written with no image for its containers") + } +} diff --git a/operator/internal/controllers/node/selfbudget_test.go b/operator/internal/controllers/node/selfbudget_test.go index 412b19bb2..74049f401 100644 --- a/operator/internal/controllers/node/selfbudget_test.go +++ b/operator/internal/controllers/node/selfbudget_test.go @@ -220,7 +220,7 @@ var errUnreachableControlPlane = errors.New("the control plane is unreachable") // that would have created either. func TestTheWindowHoldsTheManagerBeforeAnythingThatCanFail(t *testing.T) { node := &simplyblockv1alpha2.StorageNode{ - ObjectMeta: metav1.ObjectMeta{Name: "a-node", Namespace: budgetNamespace}, + ObjectMeta: metav1.ObjectMeta{Name: opsNodeName, Namespace: budgetNamespace}, Spec: simplyblockv1alpha2.StorageNodeSpec{ ClusterRef: budgetCluster, WorkerNode: managerWorker, @@ -264,7 +264,7 @@ func TestTheWindowHoldsTheManagerBeforeAnythingThatCanFail(t *testing.T) { // same deadlock, one object over. func TestTheWindowLetsTheManagerGoWhenItLetsTheStoragePodGo(t *testing.T) { node := &simplyblockv1alpha2.StorageNode{ - ObjectMeta: metav1.ObjectMeta{Name: "a-node", Namespace: budgetNamespace}, + ObjectMeta: metav1.ObjectMeta{Name: opsNodeName, Namespace: budgetNamespace}, Spec: simplyblockv1alpha2.StorageNodeSpec{ ClusterRef: budgetCluster, WorkerNode: managerWorker, diff --git a/operator/internal/controllers/node/storagenode_controller.go b/operator/internal/controllers/node/storagenode_controller.go index b856f11d1..c74b6e9d6 100644 --- a/operator/internal/controllers/node/storagenode_controller.go +++ b/operator/internal/controllers/node/storagenode_controller.go @@ -945,11 +945,19 @@ func (r *StorageNodeReconciler) teardown( return ctrl.Result{}, err } - // The lock being clear is what says the drain has finished, whatever its - // outcome. A failed removal leaves the lock released and the operation as the - // record of why, so the object is not held forever by a drain nobody is going - // to retry. - if node.Status.ActiveOpsRef != "" { + // The drain reaching a terminal phase is what says it has finished, whatever + // its outcome. A failed removal leaves the operation as the record of why, so + // the object is not held forever by a drain nobody is going to retry. + // + // The node's lock is not the signal, and the difference is not cosmetic: a + // drain that has just been raised holds no lock yet, so a teardown reading the + // empty field would take "not started" for "finished" and drop the finalizer + // on the same pass that asked for the drain. The object then goes, and with it + // the operation it owns, and the backend node is left running with its data on + // it and nothing in Kubernetes tracking it. + if finished, err := r.drainFinished(ctx, node); err != nil { + return ctrl.Result{}, err + } else if !finished { return ctrl.Result{RequeueAfter: nodeRetry}, nil } @@ -958,6 +966,27 @@ func (r *StorageNodeReconciler) teardown( return ctrl.Result{}, r.Update(ctx, node) } +// drainFinished reports whether the removal this node raised for itself has +// reached a terminal phase. +// +// An operation that is not there is a drain that has not started rather than one +// that is over: the pass before this one raises it, and one deleted out of band +// is raised again. Treating a missing record as a finished drain is the same +// mistake as treating an unheld lock as one (§4.5). +func (r *StorageNodeReconciler) drainFinished( + ctx context.Context, node *simplyblockv1alpha2.StorageNode, +) (bool, error) { + var ops simplyblockv1alpha2.StorageNodeOps + key := types.NamespacedName{Name: node.Name + "-remove", Namespace: node.Namespace} + if err := r.Get(ctx, key, &ops); err != nil { + if apierrors.IsNotFound(err) { + return false, nil + } + return false, fmt.Errorf("read the removal of node %s: %w", node.Name, err) + } + return terminalOps(ops.Status.Phase), nil +} + // ensureOps raises one operation the entity created for itself, idempotently by // name, and reports whether it created one. // diff --git a/operator/internal/controllers/node/syncstatus_test.go b/operator/internal/controllers/node/syncstatus_test.go new file mode 100644 index 000000000..98852b90a --- /dev/null +++ b/operator/internal/controllers/node/syncstatus_test.go @@ -0,0 +1,481 @@ +// What a provisioned node's object says about it, and where that comes from. +// +// Steady state is the stream's cache: the control plane pushes a node's changes, +// so reading one back costs no request, and the poll that remains is the +// correctness floor for a stream nobody has noticed is dead. The fallback matters +// as much as the cache — a node the stream has not delivered still has to be +// readable, or a cold cache would leave the object's status frozen until the +// first snapshot lands. +// +// Occupancy is the one figure in neither the node list nor the stream. It exists +// only in the metrics the control plane exports, and it is written under +// hysteresis for a reason that is not economy: the reconciler watches its own +// objects, so writing a freshly sampled number every pass would make a node +// reconcile itself in a loop for as long as any I/O was happening. +// +// design-storagenode.md §3.3, §4.4, and §12. + +package node + +import ( + "context" + "errors" + "testing" + "time" + + corev1 "k8s.io/api/core/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" + "k8s.io/client-go/tools/events" + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + "github.com/simplyblock/atlas/prometheus" + "github.com/simplyblock/atlas/ptr" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/testsupport" + "github.com/simplyblock/simplyblock-operator/internal/cpinformer/subscriptions" +) + +// sampledCapacity is a metrics endpoint that answers with what a test scripted, +// or fails the way one that is momentarily away does. +type sampledCapacity struct { + samples map[string]prometheus.Capacity + err error +} + +func (s *sampledCapacity) NodeCapacity( + _ context.Context, _ string, +) (map[string]prometheus.Capacity, error) { + if s.err != nil { + return nil, s.err + } + return s.samples, nil +} + +// aSteadyNode builds the entity reconciler over a provisioned node and its +// cluster. +func aSteadyNode( + t *testing.T, api ControlPlane, objects ...client.Object, +) (*StorageNodeReconciler, client.Client) { + t.Helper() + scheme := testsupport.NewScheme(t, corev1.AddToScheme) + + node := anOpsNode() + node.Finalizers = []string{NodeFinalizer} + world := append([]client.Object{node, anOpsCluster()}, objects...) + + apiClient := fake.NewClientBuilder().WithScheme(scheme). + WithObjects(world...). + WithStatusSubresource( + &simplyblockv1alpha2.StorageNode{}, + &simplyblockv1alpha2.StorageNodeOps{}, + &simplyblockv1alpha2.StorageCluster{}, + ). + WithIndex(&simplyblockv1alpha2.StorageNode{}, clusterRefField, + func(o client.Object) []string { + return []string{o.(*simplyblockv1alpha2.StorageNode).Spec.ClusterRef} + }). + Build() + + return &StorageNodeReconciler{ + Client: apiClient, + Scheme: scheme, + Recorder: events.NewFakeRecorder(64), + API: api, + Workload: &Workload{Client: apiClient}, + }, apiClient +} + +// nodeRead is the node as the API server now holds it. +func nodeRead(t *testing.T, apiClient client.Client) *simplyblockv1alpha2.StorageNode { + t.Helper() + var node simplyblockv1alpha2.StorageNode + key := client.ObjectKey{Namespace: opsNamespace, Name: opsNodeName} + if err := apiClient.Get(context.Background(), key, &node); err != nil { + t.Fatalf("reading the node: %v", err) + } + return &node +} + +// settle runs one reconcile of the node. +func settle(t *testing.T, r *StorageNodeReconciler) ctrl.Result { + t.Helper() + result, err := r.Reconcile(context.Background(), ctrl.Request{ + NamespacedName: client.ObjectKey{Namespace: opsNamespace, Name: opsNodeName}, + }) + if err != nil { + t.Fatalf("reconciling: %v", err) + } + return result +} + +// A node the stream has delivered is read from the cache, and the control plane +// is not asked for what it has already pushed. +func TestAPushedNodeIsReadFromTheStreamRatherThanAskedFor(t *testing.T) { + api := aControlPlane() + r, apiClient := aSteadyNode(t, api) + r.Nodes = &deliveredNodes{synced: true, nodes: []subscriptions.NodeDTO{{ + ID: opsNodeID, Status: nodeStatusOnline, ManagementIP: "192.168.10.112", + HealthCheck: true, Hostname: "vm02_4420", CPUCount: 6, Volumes: 3, + RPCPort: 4420, LvolPort: 4426, NVMeOFPort: 4421, + }}} + + settle(t, r) + + if asked := api.asked("StorageNode"); asked != 0 { + t.Errorf("the control plane was asked %d time(s) for a node the stream delivered", asked) + } + node := nodeRead(t, apiClient) + if node.Status.Status != nodeStatusOnline || !node.Status.Health { + t.Errorf("status/health = %q/%v, want what the stream reported", + node.Status.Status, node.Status.Health) + } + if node.Status.Hostname != "vm02_4420" { + t.Errorf("hostname = %q; the control plane's name for a node is not the worker's", + node.Status.Hostname) + } + if node.Status.Ports == nil || node.Status.Ports.Management != "192.168.10.112" { + t.Errorf("ports = %+v, want the addresses the stream carried", node.Status.Ports) + } + if node.Status.Resources == nil || ptr.From(node.Status.Resources.Volumes, 0) != 3 { + t.Errorf("resources = %+v, want the volume count the stream carried", node.Status.Resources) + } + if node.Status.Phase != simplyblockv1alpha2.StorageNodePhaseOnline { + t.Errorf("phase = %q, want Online", node.Status.Phase) + } +} + +// A node the stream has not delivered is read from the control plane, because a +// cold cache must not freeze a node's status until the first snapshot lands. +func TestANodeTheStreamHasNotDeliveredIsAskedFor(t *testing.T) { + api := aControlPlane().reporting(nodeStatusInCreation) + r, apiClient := aSteadyNode(t, api) + r.Nodes = &deliveredNodes{synced: false} + + settle(t, r) + + if asked := api.asked("StorageNode"); asked != 1 { + t.Errorf("the control plane was asked %d time(s), want once for an undelivered node", + asked) + } + if phase := nodeRead(t, apiClient).Status.Phase; phase != + simplyblockv1alpha2.StorageNodePhaseProvisioning { + t.Errorf("phase = %q, want Provisioning for a node the control plane is still creating", + phase) + } +} + +// The phase is the operator's reading of the lifecycle the control plane +// reports, and the two vocabularies are deliberately separate. +func TestThePhaseIsThisOperatorsReadingOfWhatTheControlPlaneSays(t *testing.T) { + cases := []struct { + reading NodeReading + want simplyblockv1alpha2.StorageNodePhase + }{ + {NodeReading{Status: nodeStatusOnline}, simplyblockv1alpha2.StorageNodePhaseOnline}, + {NodeReading{Status: nodeStatusActive}, simplyblockv1alpha2.StorageNodePhaseOnline}, + { + NodeReading{Status: nodeStatusOnline, DevicesCount: 4, OnlineDevicesCount: 3}, + simplyblockv1alpha2.StorageNodePhaseDegraded, + }, + {NodeReading{Status: nodeStatusSuspended}, simplyblockv1alpha2.StorageNodePhaseOffline}, + {NodeReading{Status: nodeStatusOffline}, simplyblockv1alpha2.StorageNodePhaseOffline}, + { + NodeReading{Status: nodeStatusInCreation}, + simplyblockv1alpha2.StorageNodePhaseProvisioning, + }, + { + NodeReading{Status: nodeStatusInRestart}, + simplyblockv1alpha2.StorageNodePhaseProvisioning, + }, + {NodeReading{Status: "unreachable"}, simplyblockv1alpha2.StorageNodePhaseFailed}, + {NodeReading{Status: "a status this operator has never heard of"}, + simplyblockv1alpha2.StorageNodePhaseFailed}, + } + for _, c := range cases { + if got := phaseOf(c.reading); got != c.want { + t.Errorf("a node reporting %q with %d of %d devices online is %q, want %q", + c.reading.Status, c.reading.OnlineDevicesCount, c.reading.DevicesCount, + got, c.want) + } + } +} + +// A stored UUID the control plane no longer reports is a cluster that was reset +// and its nodes recreated. The object goes back to provisioning rather than +// reporting a node that is not there. +func TestANodeTheControlPlaneHasForgottenGoesBackToProvisioning(t *testing.T) { + api := aControlPlane() + delete(api.nodes, opsNodeID) + r, apiClient := aSteadyNode(t, api) + + settle(t, r) + + node := nodeRead(t, apiClient) + if node.Status.UUID != "" { + t.Errorf("the object still names backend node %q, which the control plane has forgotten", + node.Status.UUID) + } + if node.Status.Phase != simplyblockv1alpha2.StorageNodePhasePending { + t.Errorf("phase = %q, want Pending", node.Status.Phase) + } +} + +// The first sample is always worth writing: an object that says nothing about +// its occupancy says nothing a reader can use. +func TestAFirstCapacityReadingIsAlwaysWritten(t *testing.T) { + sampled := prometheus.Capacity{ + Total: 112303538176, Used: 422576128, SampledAt: time.Unix(1788423117, 0).UTC(), + } + r, apiClient := aSteadyNode(t, aControlPlane()) + r.Capacity = &sampledCapacity{samples: map[string]prometheus.Capacity{opsNodeID: sampled}} + + settle(t, r) + + capacity := nodeRead(t, apiClient).Status.Resources.Capacity + if capacity == nil || capacity.UsedBytes == nil || capacity.TotalBytes == nil { + t.Fatalf("capacity = %+v, want the sample that was taken", capacity) + } + if *capacity.UsedBytes != sampled.Used || *capacity.TotalBytes != sampled.Total { + t.Errorf("capacity = %d of %d, want %d of %d", + *capacity.UsedBytes, *capacity.TotalBytes, sampled.Used, sampled.Total) + } + if capacity.SampledAt == nil { + t.Error("nothing records when the sample was taken, so its age cannot be judged") + } +} + +// Whether a sample is worth writing is decided against what the object already +// says, because every write schedules another reconcile. +func TestASampleIsWrittenOnlyWhenItSaysSomethingNew(t *testing.T) { + const total = int64(100_000_000_000) + recorded := &simplyblockv1alpha2.StorageNodeCapacity{ + TotalBytes: ptr.To(total), UsedBytes: ptr.To(int64(50_000_000_000)), + } + + cases := []struct { + name string + sample prometheus.Capacity + want bool + }{ + { + "a use that barely moved", + prometheus.Capacity{Total: total, Used: 50_100_000_000, SampledAt: time.Now()}, + false, + }, + { + "a use that moved by a percent of the node", + prometheus.Capacity{Total: total, Used: 51_000_000_000, SampledAt: time.Now()}, + true, + }, + { + "a total that changed, which is a device joining or leaving", + prometheus.Capacity{Total: total + 1, Used: 50_000_000_000, SampledAt: time.Now()}, + true, + }, + { + "nothing measured this node at all", + prometheus.Capacity{}, + false, + }, + } + for _, c := range cases { + t.Run(c.name, func(t *testing.T) { + if got := worthWriting(recorded, c.sample); got != c.want { + t.Errorf("worthWriting = %v, want %v", got, c.want) + } + }) + } + + if !worthWriting(nil, prometheus.Capacity{Total: total, Used: 1, SampledAt: time.Now()}) { + t.Error("a first reading was not written, so the object would say nothing about occupancy") + } +} + +// A capacity source that is momentarily away is not a reason to publish nothing: +// the rest of the status is correct without it. +func TestAFailingCapacitySourceStillLeavesTheNodePublished(t *testing.T) { + r, apiClient := aSteadyNode(t, aControlPlane()) + r.Capacity = &sampledCapacity{err: errors.New("the metrics endpoint is away")} + + settle(t, r) + + node := nodeRead(t, apiClient) + if node.Status.Phase != simplyblockv1alpha2.StorageNodePhaseOnline { + t.Errorf("phase = %q, want the node published without its occupancy", node.Status.Phase) + } + if node.Status.Resources != nil && node.Status.Resources.Capacity != nil { + t.Errorf("capacity = %+v, want nothing said about a figure nobody measured", + node.Status.Resources.Capacity) + } +} + +// A node nothing sampled carries no capacity, rather than a zero one that reads +// as an empty node. +func TestANodeNobodySampledCarriesNoCapacity(t *testing.T) { + r, apiClient := aSteadyNode(t, aControlPlane()) + r.Capacity = &sampledCapacity{samples: map[string]prometheus.Capacity{ + "another-node": {Total: 1, Used: 1, SampledAt: time.Now()}, + }} + + settle(t, r) + + node := nodeRead(t, apiClient) + if node.Status.Resources != nil && node.Status.Resources.Capacity != nil { + t.Errorf("capacity = %+v, want none for a node nothing sampled", + node.Status.Resources.Capacity) + } +} + +// A cordoned worker takes its storage-node pod with it, so the node is taken +// down deliberately rather than killed underneath a running SPDK process. The +// operator raises the window; a user does not have to. +func TestACordonedWorkerRaisesItsMaintenanceWindow(t *testing.T) { + cordoned := aWorker(opsWorker, true) + cordoned.Spec.Unschedulable = true + r, apiClient := aSteadyNode(t, aControlPlane(), cordoned) + + settle(t, r) + + var window simplyblockv1alpha2.StorageNodeOps + key := client.ObjectKey{Namespace: opsNamespace, Name: "a-node-maintenance"} + if err := apiClient.Get(context.Background(), key, &window); err != nil { + t.Fatalf("reading the maintenance window: %v", err) + } + if window.Spec.Action != simplyblockv1alpha2.StorageNodeOpsActionHostMaintenance { + t.Errorf("the operation raised is a %s", window.Spec.Action) + } + if len(window.OwnerReferences) != 1 || window.OwnerReferences[0].Name != opsNodeName { + t.Errorf("owner references = %v, want the node that raised it for itself", + window.OwnerReferences) + } + + // The second pass finds the window it raised rather than raising another. + settle(t, r) + var windows simplyblockv1alpha2.StorageNodeOpsList + if err := apiClient.List(context.Background(), &windows); err != nil { + t.Fatalf("listing the operations: %v", err) + } + if len(windows.Items) != 1 { + t.Errorf("%d windows were raised for one cordon", len(windows.Items)) + } +} + +// A node with no backend behind it has nothing to drain, so the object goes +// straight away. +func TestANodeThatWasNeverProvisionedIsDeletedOutright(t *testing.T) { + r, apiClient := aSteadyNode(t, aControlPlane()) + node := nodeRead(t, apiClient) + node.Status.UUID = "" + if err := apiClient.Status().Update(context.Background(), node); err != nil { + t.Fatalf("clearing the node's UUID: %v", err) + } + if err := apiClient.Delete(context.Background(), node); err != nil { + t.Fatalf("deleting the node: %v", err) + } + + settle(t, r) + + var gone simplyblockv1alpha2.StorageNode + err := apiClient.Get(context.Background(), + client.ObjectKey{Namespace: opsNamespace, Name: opsNodeName}, &gone) + if err == nil { + t.Errorf("the object is still there with finalizers %v", gone.Finalizers) + } +} + +// A node with data on it is drained before the object goes, by a Remove +// operation the node owns, and the finalizer is held until that operation has +// finished. +// +// Regression: 2026-09-17-node-teardown-reads-an-unheld-lock-as-a-finished-drain. +// The teardown held the finalizer while status.activeOpsRef was set, and a drain +// raised one line earlier has not taken the lock yet — so the first teardown pass +// read "not started" as "finished" and dropped the finalizer immediately. The +// object went, the operation it owns went with it, and the backend node was left +// running with its data on it and nothing in Kubernetes tracking it. +func TestANodeWithDataOnItIsDrainedBeforeItGoes(t *testing.T) { + r, apiClient := aSteadyNode(t, aControlPlane()) + node := nodeRead(t, apiClient) + if err := apiClient.Delete(context.Background(), node); err != nil { + t.Fatalf("deleting the node: %v", err) + } + + settle(t, r) + + drain := drainRaisedFor(t, apiClient) + if drain.Spec.Action != simplyblockv1alpha2.StorageNodeOpsActionRemove { + t.Errorf("the operation raised is a %s", drain.Spec.Action) + } + if len(drain.OwnerReferences) != 1 || drain.OwnerReferences[0].Name != opsNodeName { + t.Errorf("owner references = %v, want the node that raised it for itself", + drain.OwnerReferences) + } + if len(finalizersOn(t, apiClient)) == 0 { + t.Fatal("the finalizer came off on the pass that raised the drain, so the object " + + "goes while the backend node is still there with its data on it") + } + + // A drain that is running holds it too, which is the long middle of a real + // removal. + finishDrain(t, apiClient, simplyblockv1alpha2.StorageNodeOpsPhaseRunning) + settle(t, r) + if len(finalizersOn(t, apiClient)) == 0 { + t.Fatal("the finalizer came off while the drain was still running") + } + + // A drain that is over releases it, whatever its outcome: the operation stays + // as the record, and an object nobody can delete would be worse. + finishDrain(t, apiClient, simplyblockv1alpha2.StorageNodeOpsPhaseSucceeded) + settle(t, r) + + var gone simplyblockv1alpha2.StorageNode + err := apiClient.Get(context.Background(), + client.ObjectKey{Namespace: opsNamespace, Name: opsNodeName}, &gone) + if err == nil { + t.Errorf("the object is still held by %v after its drain finished", gone.Finalizers) + } +} + +// finalizersOn is what still holds the node, and nothing at all when the object +// has already gone. +func finalizersOn(t *testing.T, apiClient client.Client) []string { + t.Helper() + var node simplyblockv1alpha2.StorageNode + key := client.ObjectKey{Namespace: opsNamespace, Name: opsNodeName} + err := apiClient.Get(context.Background(), key, &node) + if apierrors.IsNotFound(err) { + return nil + } + if err != nil { + t.Fatalf("reading the node: %v", err) + } + return node.Finalizers +} + +// drainRaisedFor reads the removal the node raised for itself. +func drainRaisedFor( + t *testing.T, apiClient client.Client, +) *simplyblockv1alpha2.StorageNodeOps { + t.Helper() + var drain simplyblockv1alpha2.StorageNodeOps + key := client.ObjectKey{Namespace: opsNamespace, Name: "a-node-remove"} + if err := apiClient.Get(context.Background(), key, &drain); err != nil { + t.Fatalf("reading the drain: %v", err) + } + return &drain +} + +// finishDrain moves the removal to the phase a case is about. +func finishDrain( + t *testing.T, apiClient client.Client, phase simplyblockv1alpha2.StorageNodeOpsPhase, +) { + t.Helper() + drain := drainRaisedFor(t, apiClient) + drain.Status.Phase = phase + if err := apiClient.Status().Update(context.Background(), drain); err != nil { + t.Fatalf("moving the drain to %s: %v", phase, err) + } +} diff --git a/operator/internal/controllers/node/unregister_test.go b/operator/internal/controllers/node/unregister_test.go index c4d8f073e..223c92f00 100644 --- a/operator/internal/controllers/node/unregister_test.go +++ b/operator/internal/controllers/node/unregister_test.go @@ -26,7 +26,7 @@ import ( func aResolvedNode(uuid string) *simplyblockv1alpha2.StorageNode { node := &simplyblockv1alpha2.StorageNode{ - ObjectMeta: metav1.ObjectMeta{Name: "a-node", Namespace: "simplyblock"}, + ObjectMeta: metav1.ObjectMeta{Name: opsNodeName, Namespace: "simplyblock"}, } node.Status.UUID = uuid return node diff --git a/operator/internal/controllers/node/watches_test.go b/operator/internal/controllers/node/watches_test.go new file mode 100644 index 000000000..c00a83a62 --- /dev/null +++ b/operator/internal/controllers/node/watches_test.go @@ -0,0 +1,128 @@ +// What wakes a reconcile, and what a pass that changed nothing writes. +// +// The mappings are what make the queue move. An operation waiting for a lock has +// nothing of its own to react to, so a controller watching only its own kind +// would leave every queued operation waiting out a requeue interval after the +// lock frees; a node waiting for its cluster's UUID would wait out the same +// interval after the cluster gets one; and a cordon would reach the node it is +// about only by the slow backstop. +// +// The other half is the opposite discipline. Both reconcilers watch their own +// objects, so a status write schedules another pass — which means a pass that +// found nothing new has to write nothing at all, or a node serving I/O +// reconciles itself in a loop for as long as the I/O lasts. +// +// design-crd-model.md §3.2 and design-storagenode.md §12. + +package node + +import ( + "context" + "testing" + + corev1 "k8s.io/api/core/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/testsupport" +) + +// A cluster event wakes its own nodes and nobody else's, which is what the index +// on spec.clusterRef is for. +func TestAClusterEventWakesItsOwnNodes(t *testing.T) { + elsewhere := anOpsNode() + elsewhere.Name = "another-clusters-node" + elsewhere.Spec.ClusterRef = "another-cluster" + r, _ := aSteadyNode(t, aControlPlane(), elsewhere) + + requests := r.nodesOf(context.Background(), anOpsCluster()) + + if len(requests) != 1 || requests[0].Name != opsNodeName { + t.Errorf("the cluster woke %v, want its own node alone", requests) + } +} + +// A worker event wakes the nodes running on that worker, which is how a cordon +// reaches the node whose maintenance window it is about. +func TestAWorkerEventWakesTheNodesOnIt(t *testing.T) { + elsewhere := anOpsNode() + elsewhere.Name = "a-node-on-another-worker" + elsewhere.Spec.WorkerNode = "worker-9" + r, _ := aSteadyNode(t, aControlPlane(), elsewhere) + + requests := r.nodesOn(context.Background(), aWorker(opsWorker, true)) + + if len(requests) != 1 || requests[0].Name != opsNodeName { + t.Errorf("the worker woke %v, want the node that runs on it", requests) + } +} + +// A node event wakes every unfinished operation targeting it, which is what lets +// a released lock start the next operation immediately rather than after a +// requeue interval. +func TestANodeEventWakesTheOperationsWaitingOnIt(t *testing.T) { + queued := anOperation("a-suspend", simplyblockv1alpha2.StorageNodeOpsActionSuspend) + over := anOperation("a-finished-restart", simplyblockv1alpha2.StorageNodeOpsActionRestart) + over.Status.Phase = simplyblockv1alpha2.StorageNodeOpsPhaseSucceeded + + scheme := testsupport.NewScheme(t, corev1.AddToScheme) + apiClient := fake.NewClientBuilder().WithScheme(scheme). + WithObjects(anOpsNode(), anOpsCluster(), queued, over). + WithIndex(&simplyblockv1alpha2.StorageNodeOps{}, nodeRefField, + func(o client.Object) []string { + return []string{o.(*simplyblockv1alpha2.StorageNodeOps).Spec.NodeRef} + }). + Build() + r := &StorageNodeOpsReconciler{Client: apiClient, Scheme: scheme} + + requests := r.operationsOn(context.Background(), anOpsNode()) + + if len(requests) != 1 || requests[0].Name != "a-suspend" { + t.Errorf("the node woke %v, want the operation that is still waiting for it", requests) + } +} + +// A pass that found what the object already says writes nothing, because every +// write schedules another pass and a node under load would never settle. +func TestAPassThatFoundNothingNewWritesNothing(t *testing.T) { + api := aControlPlane() + api.nodes[opsNodeID] = NodeReading{ + UUID: opsNodeID, Status: nodeStatusOnline, ManagementIP: "10.0.0.1", + Health: true, Hostname: "vm02_4420", CPUCount: 6, Volumes: 3, + RPCPort: 4420, LvolPort: 4426, NVMeOFPort: 4421, + } + r, apiClient := aSteadyNode(t, api) + + settle(t, r) + first := nodeRead(t, apiClient).ResourceVersion + + settle(t, r) + second := nodeRead(t, apiClient).ResourceVersion + + if first != second { + t.Errorf("the object moved from %s to %s on a pass that learned nothing, "+ + "and every write schedules another pass", first, second) + } +} + +// A reading that did move is written, or the object would report a node it has +// stopped describing. +func TestAPassThatFoundSomethingNewWritesIt(t *testing.T) { + api := aControlPlane() + r, apiClient := aSteadyNode(t, api) + + settle(t, r) + before := nodeRead(t, apiClient).ResourceVersion + + api.reporting(nodeStatusSuspended) + settle(t, r) + + after := nodeRead(t, apiClient) + if after.ResourceVersion == before { + t.Error("the node still reports online after the control plane suspended it") + } + if after.Status.Phase != simplyblockv1alpha2.StorageNodePhaseOffline { + t.Errorf("phase = %q, want Offline for a suspended node", after.Status.Phase) + } +} From 4ed1ca34f4a7e246143a354381b5e3fefd5a0de0 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Fri, 18 Sep 2026 16:25:58 +0100 Subject: [PATCH 076/206] resolve volume ID via a shared volumeIDFrom(req) helper that checks ReplicationSource.Volume.VolumeId --- .../internal/csi/controller/replication.go | 31 ++++++-- .../controller/replication_lifecycle_test.go | 43 +++++++++++ .../csi/controller/replication_test.go | 74 +++++++++++++++++++ 3 files changed, 142 insertions(+), 6 deletions(-) diff --git a/csi-driver/internal/csi/controller/replication.go b/csi-driver/internal/csi/controller/replication.go index dc7a02fab..f472ba2d6 100644 --- a/csi-driver/internal/csi/controller/replication.go +++ b/csi-driver/internal/csi/controller/replication.go @@ -32,6 +32,25 @@ const replicationPolicyParam = "replicationPolicyID" // design §5.2's "Recovered source"). const sourceClusterIDParam = "sourceClusterID" +// volumeIDCarrier is every Replication request type: each exposes the legacy +// flat VolumeId field and the ReplicationSource oneof. +type volumeIDCarrier interface { + GetVolumeId() string + GetReplicationSource() *replication.ReplicationSource +} + +// volumeIDFrom resolves a request's volume ID. The real upstream sidecar +// (kubernetes-csi-addons v0.15.0, internal/sidecar/service.ReplicationServer) +// proxies every Replication RPC through ReplicationSource and never sets the +// legacy flat VolumeId, so that field is checked first; the flat field is +// kept as a fallback for any caller that still sends it. +func volumeIDFrom(req volumeIDCarrier) string { + if v := req.GetReplicationSource().GetVolume().GetVolumeId(); v != "" { + return v + } + return req.GetVolumeId() +} + // EnableVolumeReplication attaches the volume to the policy named by the // VolumeReplicationClass. Attaching to the policy the volume already follows // is success (the backend's own idempotency, P0-2). @@ -53,7 +72,7 @@ func (cs *Server) EnableVolumeReplication( if policyID == "" { return nil, status.Errorf(codes.InvalidArgument, "VolumeReplicationClass parameter %q is required", replicationPolicyParam) } - h, err := csicommon.ParseVolumeHandle(req.GetVolumeId()) + h, err := csicommon.ParseVolumeHandle(volumeIDFrom(req)) if err != nil { return nil, status.Error(codes.InvalidArgument, err.Error()) } @@ -74,7 +93,7 @@ func (cs *Server) DisableVolumeReplication( ctx context.Context, req *replication.DisableVolumeReplicationRequest, ) (*replication.DisableVolumeReplicationResponse, error) { - h, err := csicommon.ParseVolumeHandle(req.GetVolumeId()) + h, err := csicommon.ParseVolumeHandle(volumeIDFrom(req)) if err != nil { return nil, status.Error(codes.InvalidArgument, err.Error()) } @@ -99,7 +118,7 @@ func (cs *Server) GetVolumeReplicationInfo( ctx context.Context, req *replication.GetVolumeReplicationInfoRequest, ) (*replication.GetVolumeReplicationInfoResponse, error) { - h, err := csicommon.ParseVolumeHandle(req.GetVolumeId()) + h, err := csicommon.ParseVolumeHandle(volumeIDFrom(req)) if err != nil { return nil, status.Error(codes.InvalidArgument, err.Error()) } @@ -131,7 +150,7 @@ func (cs *Server) PromoteVolume( ctx context.Context, req *replication.PromoteVolumeRequest, ) (*replication.PromoteVolumeResponse, error) { - h, err := csicommon.ParseVolumeHandle(req.GetVolumeId()) + h, err := csicommon.ParseVolumeHandle(volumeIDFrom(req)) if err != nil { return nil, status.Error(codes.InvalidArgument, err.Error()) } @@ -155,7 +174,7 @@ func (cs *Server) DemoteVolume( ctx context.Context, req *replication.DemoteVolumeRequest, ) (*replication.DemoteVolumeResponse, error) { - h, err := csicommon.ParseVolumeHandle(req.GetVolumeId()) + h, err := csicommon.ParseVolumeHandle(volumeIDFrom(req)) if err != nil { return nil, status.Error(codes.InvalidArgument, err.Error()) } @@ -182,7 +201,7 @@ func (cs *Server) ResyncVolume( ctx context.Context, req *replication.ResyncVolumeRequest, ) (*replication.ResyncVolumeResponse, error) { - h, err := csicommon.ParseVolumeHandle(req.GetVolumeId()) + h, err := csicommon.ParseVolumeHandle(volumeIDFrom(req)) if err != nil { return nil, status.Error(codes.InvalidArgument, err.Error()) } diff --git a/csi-driver/internal/csi/controller/replication_lifecycle_test.go b/csi-driver/internal/csi/controller/replication_lifecycle_test.go index b2af814f0..6818eaff7 100644 --- a/csi-driver/internal/csi/controller/replication_lifecycle_test.go +++ b/csi-driver/internal/csi/controller/replication_lifecycle_test.go @@ -82,6 +82,23 @@ func TestPromoteVolumePlannedWithNoDemoteIsFailedPrecondition(t *testing.T) { } } +// Same real-world shape as TestEnableVolumeReplicationUsesReplicationSourceWhenVolumeIdIsEmpty: +// the vendored controller-manager/sidecar chain sends every Replication RPC, +// Promote included, via ReplicationSource with the legacy flat VolumeId left +// empty. +func TestPromoteVolumeUsesReplicationSourceWhenVolumeIdIsEmpty(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + + _, err := cs.PromoteVolume(context.Background(), &replication.PromoteVolumeRequest{ + ReplicationSource: replicationSourceFor(testReplVolID), Force: true, + }) + if err != nil { + t.Fatal(err) + } +} + func TestDemoteVolumeDone(t *testing.T) { mock := newMockSBCLI() defer mock.Close() @@ -128,6 +145,19 @@ func TestDemoteVolumeBackendFailureIsUnavailable(t *testing.T) { } } +func TestDemoteVolumeUsesReplicationSourceWhenVolumeIdIsEmpty(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + + _, err := cs.DemoteVolume(context.Background(), &replication.DemoteVolumeRequest{ + ReplicationSource: replicationSourceFor(testReplVolID), + }) + if err != nil { + t.Fatal(err) + } +} + func TestResyncVolume(t *testing.T) { mock := newMockSBCLI() defer mock.Close() @@ -141,6 +171,19 @@ func TestResyncVolume(t *testing.T) { } } +func TestResyncVolumeUsesReplicationSourceWhenVolumeIdIsEmpty(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + + _, err := cs.ResyncVolume(context.Background(), &replication.ResyncVolumeRequest{ + ReplicationSource: replicationSourceFor(testReplVolID), + }) + if err != nil { + t.Fatal(err) + } +} + func TestResyncVolumeForwardsTheSourceClusterParameter(t *testing.T) { mock := newMockSBCLI() defer mock.Close() diff --git a/csi-driver/internal/csi/controller/replication_test.go b/csi-driver/internal/csi/controller/replication_test.go index c358fcb21..578356891 100644 --- a/csi-driver/internal/csi/controller/replication_test.go +++ b/csi-driver/internal/csi/controller/replication_test.go @@ -21,6 +21,18 @@ func newReplicationTestServer(t *testing.T, mock *mockSBCLI) *Server { return newTestControllerServer(t, mock) } +// replicationSourceFor builds the ReplicationSource the real +// kubernetes-csi-addons v0.15.0 sidecar sends on every Replication RPC +// instead of the legacy flat VolumeId field (internal/sidecar/service's +// ReplicationServer proxy never sets it). +func replicationSourceFor(volumeID string) *replication.ReplicationSource { + return &replication.ReplicationSource{ + Type: &replication.ReplicationSource_Volume{ + Volume: &replication.ReplicationSource_VolumeSource{VolumeId: volumeID}, + }, + } +} + func TestEnableVolumeReplication(t *testing.T) { mock := newMockSBCLI() defer mock.Close() @@ -57,6 +69,28 @@ func TestEnableVolumeReplicationRepeatedIsIdempotent(t *testing.T) { } } +// The real controller-manager/sidecar chain (v0.15.0) sends the volume +// identity via ReplicationSource, leaving the legacy flat VolumeId field +// empty -- confirmed against a live cluster, where every Replication RPC +// failed with "invalid volume handle \"\"" despite the controller-manager's +// own log showing it resolved a correct handle. +func TestEnableVolumeReplicationUsesReplicationSourceWhenVolumeIdIsEmpty(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + + _, err := cs.EnableVolumeReplication(context.Background(), &replication.EnableVolumeReplicationRequest{ + ReplicationSource: replicationSourceFor(testReplVolID), + Parameters: map[string]string{replicationPolicyParam: testReplPolicyID}, + }) + if err != nil { + t.Fatal(err) + } + if got := mock.volumes[testReplVolumeID].ReplicationPolicyID; got != testReplPolicyID { + t.Errorf("ReplicationPolicyID = %q, want %q", got, testReplPolicyID) + } +} + func TestEnableVolumeReplicationMissingPolicyParam(t *testing.T) { mock := newMockSBCLI() defer mock.Close() @@ -155,6 +189,23 @@ func TestDisableVolumeReplicationNotAttachedIsSuccess(t *testing.T) { } } +func TestDisableVolumeReplicationUsesReplicationSourceWhenVolumeIdIsEmpty(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + mock.volumes[testReplVolumeID].ReplicationPolicyID = testReplPolicyID + + _, err := cs.DisableVolumeReplication(context.Background(), &replication.DisableVolumeReplicationRequest{ + ReplicationSource: replicationSourceFor(testReplVolID), + }) + if err != nil { + t.Fatal(err) + } + if got := mock.volumes[testReplVolumeID].ReplicationPolicyID; got != "" { + t.Errorf("ReplicationPolicyID = %q, want cleared", got) + } +} + func TestDisableVolumeReplicationDuringCutoverIsAborted(t *testing.T) { mock := newMockSBCLI() defer mock.Close() @@ -211,6 +262,29 @@ func TestGetVolumeReplicationInfoNeverReplicated(t *testing.T) { } } +func TestGetVolumeReplicationInfoUsesReplicationSourceWhenVolumeIdIsEmpty(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + mock.replicationStatus[testReplVolumeID] = map[string]any{ + "role": "source", "state": "in_sync", + "last_replicated_at": "2026-09-17T12:00:00Z", + "lag_seconds": 42, + "outstanding_count": 0, "outstanding_bytes": 0, + "failing_count": 0, "max_retry_reached": false, "resyncing": false, + } + + resp, err := cs.GetVolumeReplicationInfo(context.Background(), &replication.GetVolumeReplicationInfoRequest{ + ReplicationSource: replicationSourceFor(testReplVolID), + }) + if err != nil { + t.Fatal(err) + } + if resp.LastSyncTime == nil || resp.LastSyncTime.AsTime().Unix() != 1789646400 { + t.Errorf("LastSyncTime = %v, want 2026-09-17T12:00:00Z", resp.LastSyncTime) + } +} + func TestGetVolumeReplicationInfoUnknownVolume(t *testing.T) { mock := newMockSBCLI() defer mock.Close() From 04e271d4926a37026ee10a7a13670ce890e51dc9 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Fri, 18 Sep 2026 18:53:10 +0100 Subject: [PATCH 077/206] read role via GetVolumeReplicationInfo before calling the backend's failover --- .../internal/csi/controller/replication.go | 27 +++++++ .../controller/replication_lifecycle_test.go | 78 +++++++++++++++++++ .../designs/design-csi-addons-replication.md | 4 +- 3 files changed, 108 insertions(+), 1 deletion(-) diff --git a/csi-driver/internal/csi/controller/replication.go b/csi-driver/internal/csi/controller/replication.go index f472ba2d6..540983794 100644 --- a/csi-driver/internal/csi/controller/replication.go +++ b/csi-driver/internal/csi/controller/replication.go @@ -137,6 +137,16 @@ func (cs *Server) GetVolumeReplicationInfo( return resp, nil } +// alreadyPromotedRoles are the ReplicationStatus.Role values for which +// PromoteVolume has nothing to do: the volume is already the live source +// (never failed over) or already sits on this side from an earlier promote. +// Checked before any backend failover call -- see PromoteVolume's own +// comment for why this check exists at all. +var alreadyPromotedRoles = map[string]bool{ + "source": true, + "failed_over": true, +} + // PromoteVolume brings the volume up as primary on this cluster (design // §5.2). Force=true is the unplanned path: it clones the last fully // replicated generation and ignores demote state entirely, because its whole @@ -146,6 +156,16 @@ func (cs *Server) GetVolumeReplicationInfo( // requested -- the split matters because the vendored csi-addons controller // auto-escalates ANY FAILED_PRECONDITION from a force=false promote to // force=true inline, with no wait-and-retry grace period of its own. +// +// The vendored controller-manager has no "already primary" awareness of its +// own: markVolumeAsPrimary calls Promote unconditionally, every time a +// VolumeReplication first declares primary intent -- including day-one +// protection of a volume that has always lived at this cluster and has never +// failed over. Left unguarded, that call would reach the backend's failover +// and wrongly clone-and-retire against a healthy source (confirmed against a +// live cluster). So this checks the volume's current role first and treats +// an already-source or already-failed-over volume as a no-op, before either +// the force value or the planned gate ever comes into play. func (cs *Server) PromoteVolume( ctx context.Context, req *replication.PromoteVolumeRequest, @@ -158,6 +178,13 @@ func (cs *Server) PromoteVolume( if err != nil { return nil, status.Error(codes.Unavailable, err.Error()) } + info, err := client.GetVolumeReplicationInfo(ctx, h.Handle()) + if err != nil { + return nil, classifyGetVolumeReplicationInfoError(err) + } + if alreadyPromotedRoles[info.Role] { + return &replication.PromoteVolumeResponse{}, nil + } if err := client.PromoteVolume(ctx, h.Handle(), req.GetForce()); err != nil { return nil, classifyPromoteVolumeError(err) } diff --git a/csi-driver/internal/csi/controller/replication_lifecycle_test.go b/csi-driver/internal/csi/controller/replication_lifecycle_test.go index 6818eaff7..d6f0382ee 100644 --- a/csi-driver/internal/csi/controller/replication_lifecycle_test.go +++ b/csi-driver/internal/csi/controller/replication_lifecycle_test.go @@ -82,6 +82,84 @@ func TestPromoteVolumePlannedWithNoDemoteIsFailedPrecondition(t *testing.T) { } } +// The vendored csi-addons controller has no "already primary" awareness of +// its own (markVolumeAsPrimary calls Promote unconditionally, every time a +// VolumeReplication first declares primary intent -- including day-one +// protection of a volume that has always lived at this cluster, never +// failed over). So the driver must supply that check itself: a volume +// already reporting role=source has nothing to fail over, and calling the +// backend's failover would wrongly clone+retire against a healthy source -- +// confirmed against a live cluster, where protecting a brand-new PVC +// (VolumeReplication created directly as primary) triggered a real target- +// cluster clone with no failover ever intended. +func TestPromoteVolumeAlreadySourceIsNoOp(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + mock.replicationStatus[testReplVolumeID] = map[string]any{ + "role": "source", "state": "in_sync", + "outstanding_count": 0, "outstanding_bytes": 0, + "failing_count": 0, "max_retry_reached": false, "resyncing": false, + } + + _, err := cs.PromoteVolume(context.Background(), &replication.PromoteVolumeRequest{ + VolumeId: testReplVolID, Force: true, + }) + if err != nil { + t.Fatal(err) + } + if mock.lastFailoverQuery != "" { + t.Errorf("failover query = %q, want no failover call: the volume is already the live source", mock.lastFailoverQuery) + } +} + +// A repeat promote after an earlier one already succeeded is the same +// no-op: nothing to fail over a second time. +func TestPromoteVolumeAlreadyFailedOverIsNoOp(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + mock.replicationStatus[testReplVolumeID] = map[string]any{ + "role": "failed_over", "state": "in_sync", + "outstanding_count": 0, "outstanding_bytes": 0, + "failing_count": 0, "max_retry_reached": false, "resyncing": false, + } + + _, err := cs.PromoteVolume(context.Background(), &replication.PromoteVolumeRequest{ + VolumeId: testReplVolID, Force: false, + }) + if err != nil { + t.Fatal(err) + } + if mock.lastFailoverQuery != "" { + t.Errorf("failover query = %q, want no failover call: already promoted", mock.lastFailoverQuery) + } +} + +// A volume genuinely holding the replica side (role=secondary) is the real +// promotion case, and must still reach the backend's failover exactly as +// before. +func TestPromoteVolumeSecondaryProceedsToFailover(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + mock.replicationStatus[testReplVolumeID] = map[string]any{ + "role": "secondary", "state": "in_sync", + "outstanding_count": 0, "outstanding_bytes": 0, + "failing_count": 0, "max_retry_reached": false, "resyncing": false, + } + + _, err := cs.PromoteVolume(context.Background(), &replication.PromoteVolumeRequest{ + VolumeId: testReplVolID, Force: true, + }) + if err != nil { + t.Fatal(err) + } + if mock.lastFailoverQuery == "" { + t.Error("want a failover call: the volume genuinely holds the replica side") + } +} + // Same real-world shape as TestEnableVolumeReplicationUsesReplicationSourceWhenVolumeIdIsEmpty: // the vendored controller-manager/sidecar chain sends every Replication RPC, // Promote included, via ReplicationSource with the legacy flat VolumeId left diff --git a/operator/docs/designs/design-csi-addons-replication.md b/operator/docs/designs/design-csi-addons-replication.md index 65b825445..3ac9ce863 100644 --- a/operator/docs/designs/design-csi-addons-replication.md +++ b/operator/docs/designs/design-csi-addons-replication.md @@ -1,6 +1,6 @@ # Design Document: csi-addons Volume Replication -**Status:** Phase 1 Implemented +**Status:** Phase 2 Implemented **Author:** Israel Geoffrey (geoffrey1330) **Date:** 2026-09-16 (last updated 2026-09-17) **Test Plan:** [`tests/test-plan-csi-addons-replication.md`](../tests/test-plan-csi-addons-replication.md) @@ -214,6 +214,8 @@ The generated control-plane client already declares every replication endpoint, **Promote is one operation; planned and unplanned differ only in what precedes it.** Both forms clone the last fully replicated generation on the target and serve it under the preserved NVMe identity, exactly as failover does today. The planned form is lossless not because it runs different machinery but because a completed demote guarantees the last replicated generation contains every acknowledged write, and the planned gate refuses the promote until that holds. The forced form skips the gate and accepts the RPO loss, because its premise is that the source is gone. The engine's commit cutover (`replication/commit`, the `FN_REPLICATION_FINAL` runner with its shrink rounds and cutover-proceed handshake) is deliberately NOT part of this contract and remains behind only the legacy `ReplicationOps` migration path (§8, §13) -- its convergence job does NOT move into the demote verb as this design originally assumed; see the `DemoteVolume` row above for why that engine is the wrong shape for a post-unmount demote. +**A role check precedes both forms, before either the force value or the planned gate comes into play.** The vendored controller-manager's `markVolumeAsPrimary` calls Promote unconditionally, every time a `VolumeReplication` first declares primary intent (`internal/controller/replication.storage/volumereplication_controller.go`) -- including day-one protection of a volume that has always lived at this cluster and has never failed over, which is the common case, not the exceptional one. Left unguarded, that call reaches the backend's failover and clones-and-retires against a healthy, never-demoted source. So the driver reads the volume's current `role` (§6.1's status read) first: `source` or `failed_over` means there is nothing to fail over -- already the live primary, or already promoted by an earlier call -- and the RPC returns success without ever calling `replication/failover`. Only `role=secondary` (or the still-unattached `none`) falls through to the promote described above. + --- ## 6. Steady-State Status and Conditions From 3a831fd95f001b8f75397512695a2e6ee93f0866 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Fri, 18 Sep 2026 19:15:01 +0100 Subject: [PATCH 078/206] reverted read role via GetVolumeReplicationInfo before calling the backend's failover --- .../internal/csi/controller/replication.go | 27 ------- .../controller/replication_lifecycle_test.go | 78 ------------------- .../designs/design-csi-addons-replication.md | 3 +- 3 files changed, 2 insertions(+), 106 deletions(-) diff --git a/csi-driver/internal/csi/controller/replication.go b/csi-driver/internal/csi/controller/replication.go index 540983794..f472ba2d6 100644 --- a/csi-driver/internal/csi/controller/replication.go +++ b/csi-driver/internal/csi/controller/replication.go @@ -137,16 +137,6 @@ func (cs *Server) GetVolumeReplicationInfo( return resp, nil } -// alreadyPromotedRoles are the ReplicationStatus.Role values for which -// PromoteVolume has nothing to do: the volume is already the live source -// (never failed over) or already sits on this side from an earlier promote. -// Checked before any backend failover call -- see PromoteVolume's own -// comment for why this check exists at all. -var alreadyPromotedRoles = map[string]bool{ - "source": true, - "failed_over": true, -} - // PromoteVolume brings the volume up as primary on this cluster (design // §5.2). Force=true is the unplanned path: it clones the last fully // replicated generation and ignores demote state entirely, because its whole @@ -156,16 +146,6 @@ var alreadyPromotedRoles = map[string]bool{ // requested -- the split matters because the vendored csi-addons controller // auto-escalates ANY FAILED_PRECONDITION from a force=false promote to // force=true inline, with no wait-and-retry grace period of its own. -// -// The vendored controller-manager has no "already primary" awareness of its -// own: markVolumeAsPrimary calls Promote unconditionally, every time a -// VolumeReplication first declares primary intent -- including day-one -// protection of a volume that has always lived at this cluster and has never -// failed over. Left unguarded, that call would reach the backend's failover -// and wrongly clone-and-retire against a healthy source (confirmed against a -// live cluster). So this checks the volume's current role first and treats -// an already-source or already-failed-over volume as a no-op, before either -// the force value or the planned gate ever comes into play. func (cs *Server) PromoteVolume( ctx context.Context, req *replication.PromoteVolumeRequest, @@ -178,13 +158,6 @@ func (cs *Server) PromoteVolume( if err != nil { return nil, status.Error(codes.Unavailable, err.Error()) } - info, err := client.GetVolumeReplicationInfo(ctx, h.Handle()) - if err != nil { - return nil, classifyGetVolumeReplicationInfoError(err) - } - if alreadyPromotedRoles[info.Role] { - return &replication.PromoteVolumeResponse{}, nil - } if err := client.PromoteVolume(ctx, h.Handle(), req.GetForce()); err != nil { return nil, classifyPromoteVolumeError(err) } diff --git a/csi-driver/internal/csi/controller/replication_lifecycle_test.go b/csi-driver/internal/csi/controller/replication_lifecycle_test.go index d6f0382ee..6818eaff7 100644 --- a/csi-driver/internal/csi/controller/replication_lifecycle_test.go +++ b/csi-driver/internal/csi/controller/replication_lifecycle_test.go @@ -82,84 +82,6 @@ func TestPromoteVolumePlannedWithNoDemoteIsFailedPrecondition(t *testing.T) { } } -// The vendored csi-addons controller has no "already primary" awareness of -// its own (markVolumeAsPrimary calls Promote unconditionally, every time a -// VolumeReplication first declares primary intent -- including day-one -// protection of a volume that has always lived at this cluster, never -// failed over). So the driver must supply that check itself: a volume -// already reporting role=source has nothing to fail over, and calling the -// backend's failover would wrongly clone+retire against a healthy source -- -// confirmed against a live cluster, where protecting a brand-new PVC -// (VolumeReplication created directly as primary) triggered a real target- -// cluster clone with no failover ever intended. -func TestPromoteVolumeAlreadySourceIsNoOp(t *testing.T) { - mock := newMockSBCLI() - defer mock.Close() - cs := newReplicationTestServer(t, mock) - mock.replicationStatus[testReplVolumeID] = map[string]any{ - "role": "source", "state": "in_sync", - "outstanding_count": 0, "outstanding_bytes": 0, - "failing_count": 0, "max_retry_reached": false, "resyncing": false, - } - - _, err := cs.PromoteVolume(context.Background(), &replication.PromoteVolumeRequest{ - VolumeId: testReplVolID, Force: true, - }) - if err != nil { - t.Fatal(err) - } - if mock.lastFailoverQuery != "" { - t.Errorf("failover query = %q, want no failover call: the volume is already the live source", mock.lastFailoverQuery) - } -} - -// A repeat promote after an earlier one already succeeded is the same -// no-op: nothing to fail over a second time. -func TestPromoteVolumeAlreadyFailedOverIsNoOp(t *testing.T) { - mock := newMockSBCLI() - defer mock.Close() - cs := newReplicationTestServer(t, mock) - mock.replicationStatus[testReplVolumeID] = map[string]any{ - "role": "failed_over", "state": "in_sync", - "outstanding_count": 0, "outstanding_bytes": 0, - "failing_count": 0, "max_retry_reached": false, "resyncing": false, - } - - _, err := cs.PromoteVolume(context.Background(), &replication.PromoteVolumeRequest{ - VolumeId: testReplVolID, Force: false, - }) - if err != nil { - t.Fatal(err) - } - if mock.lastFailoverQuery != "" { - t.Errorf("failover query = %q, want no failover call: already promoted", mock.lastFailoverQuery) - } -} - -// A volume genuinely holding the replica side (role=secondary) is the real -// promotion case, and must still reach the backend's failover exactly as -// before. -func TestPromoteVolumeSecondaryProceedsToFailover(t *testing.T) { - mock := newMockSBCLI() - defer mock.Close() - cs := newReplicationTestServer(t, mock) - mock.replicationStatus[testReplVolumeID] = map[string]any{ - "role": "secondary", "state": "in_sync", - "outstanding_count": 0, "outstanding_bytes": 0, - "failing_count": 0, "max_retry_reached": false, "resyncing": false, - } - - _, err := cs.PromoteVolume(context.Background(), &replication.PromoteVolumeRequest{ - VolumeId: testReplVolID, Force: true, - }) - if err != nil { - t.Fatal(err) - } - if mock.lastFailoverQuery == "" { - t.Error("want a failover call: the volume genuinely holds the replica side") - } -} - // Same real-world shape as TestEnableVolumeReplicationUsesReplicationSourceWhenVolumeIdIsEmpty: // the vendored controller-manager/sidecar chain sends every Replication RPC, // Promote included, via ReplicationSource with the legacy flat VolumeId left diff --git a/operator/docs/designs/design-csi-addons-replication.md b/operator/docs/designs/design-csi-addons-replication.md index 3ac9ce863..4698b2fc6 100644 --- a/operator/docs/designs/design-csi-addons-replication.md +++ b/operator/docs/designs/design-csi-addons-replication.md @@ -214,7 +214,7 @@ The generated control-plane client already declares every replication endpoint, **Promote is one operation; planned and unplanned differ only in what precedes it.** Both forms clone the last fully replicated generation on the target and serve it under the preserved NVMe identity, exactly as failover does today. The planned form is lossless not because it runs different machinery but because a completed demote guarantees the last replicated generation contains every acknowledged write, and the planned gate refuses the promote until that holds. The forced form skips the gate and accepts the RPO loss, because its premise is that the source is gone. The engine's commit cutover (`replication/commit`, the `FN_REPLICATION_FINAL` runner with its shrink rounds and cutover-proceed handshake) is deliberately NOT part of this contract and remains behind only the legacy `ReplicationOps` migration path (§8, §13) -- its convergence job does NOT move into the demote verb as this design originally assumed; see the `DemoteVolume` row above for why that engine is the wrong shape for a post-unmount demote. -**A role check precedes both forms, before either the force value or the planned gate comes into play.** The vendored controller-manager's `markVolumeAsPrimary` calls Promote unconditionally, every time a `VolumeReplication` first declares primary intent (`internal/controller/replication.storage/volumereplication_controller.go`) -- including day-one protection of a volume that has always lived at this cluster and has never failed over, which is the common case, not the exceptional one. Left unguarded, that call reaches the backend's failover and clones-and-retires against a healthy, never-demoted source. So the driver reads the volume's current `role` (§6.1's status read) first: `source` or `failed_over` means there is nothing to fail over -- already the live primary, or already promoted by an earlier call -- and the RPC returns success without ever calling `replication/failover`. Only `role=secondary` (or the still-unattached `none`) falls through to the promote described above. +**Every Promote clones, including the first one a volume's own home cluster ever sees.** The vendored controller-manager has no "already primary" awareness of its own: `markVolumeAsPrimary` calls Promote unconditionally, every time a `VolumeReplication` first declares primary intent (`internal/controller/replication.storage/volumereplication_controller.go`), including day-one protection of a volume that has always lived here and has never failed over. `get_replication_info`'s `role` field cannot distinguish that case from a volume genuinely awaiting a planned re-promotion after a completed demote -- `demote_lvol` writes only `LVol.replication_demote_state`, a field `role`'s computation never reads, so a fenced, demoted volume still reports `role: source`, identically to one that was never touched at all. A role-based short-circuit before the backend call was tried and reverted for exactly this reason: it silently skipped the real `failover?planned=true` call a completed demote is waiting on, leaving the volume fenced while csi-addons reported `Completed=True`. So the clone-on-first-promote for a never-demoted volume is accepted as designed, not guarded against: `force=true`'s existing "ignores demote state entirely" premise already covers it, matching Test 6 of `regression_test/21/test_csi_addons_replication.sh`, which asserts a never-demoted volume auto-escalates through to a clone. Resolving the resulting cost -- an unwanted clone at day-one protection time -- needs a signal this design does not yet have (§15, Open Question 5). --- @@ -446,3 +446,4 @@ A drill that silently perturbed replication would be worse than no drill. Before | 2 | **Where the preflight lives.** §7.2 attaches peerClasses validation to the `ReplicationPair` reconciler. If the redesign retires the pair kind, the preflight needs a new home (the `SimplyblockDriver`, or a standalone check job). | Operator team | | 3 | **Per-volume policy granularity.** A `VolumeReplicationClass` names one policy, and today one policy implies one target and cadence for all its volumes. Confirm one class per (policy, cadence) is an acceptable authoring model for Ramen's `replicationClassSelector`, or whether per-volume interval overrides are needed. | Operator / Backend team | | 4 | **Visibility of `drtest-` clones.** The test-cluster mode's clones on the secondary are replication sources only, never served. Decide whether the backend creates them as internal volumes (hidden from listings, exempt from the per-node subsystem cap, like the shipping path's landing volumes) or as ordinary volumes under a naming convention. | Backend team | +| 5 | **Avoiding the clone on day-one protection.** §5.2's promote always clones, including the first `PromoteVolume` a healthy, never-demoted volume ever receives -- the vendored controller-manager calls Promote unconditionally the first time any `VolumeReplication` declares primary intent, and `role` cannot tell that case apart from a volume genuinely awaiting a planned re-promotion (a demoted volume still reports `role: source`). A correct fix needs a signal this design does not expose today, most likely `replication_demote_state` (or an equivalent marker) surfaced through `get_replication_info`, so the driver can distinguish "never touched by the swap machinery" from "fenced, waiting on the real promote." Until that signal exists, every Ramen-protected volume's first reconcile materializes an unused clone on the peer, which also has no cleanup path today (`cleanup()` in the regression suite only removes Kubernetes objects). | Backend team | From c13021bc3cda80c408719b5f83e54b1c09cb0cd6 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Fri, 18 Sep 2026 22:09:43 +0200 Subject: [PATCH 079/206] feat(erasurecoding): the schemes, and how many nodes each one needs Two facts that live nowhere they are enforced. The scheme set is simplyblock_core's SUPPORTED_ERASURE_CODING_SCHEMES, and the control plane checks it on the cluster create -- by which point the operator has written a StorageCluster and an administrator has approved the document that produced it. A refusal arriving there is a refusal arriving after the only moment anybody could have acted on it. The minimum-node table is the product documentation's, and nothing checks it at all. The control plane's activation gate counts devices, ndcs+npcs+1, and never nodes, so a cluster with too few nodes for its stripe activates and then loses data on the first failure it was configured to survive. The count is ndcs+npcs nodes to place a stripe across, plus one spare per tolerated failure for the rebuild to land on: 1 for 1+0, 3 for 1+1, 4 for 2+1, 6 for 4+1, 5 for 1+2, 6 for 2+2, and 8 for 4+2. A package of its own because three unrelated callers ask the same questions: the deployment config's validation and its approval webhook, the cluster operation's activation gate, and the discovery run that proposes a scheme for a fleet it has just inspected. A leaf with no Kubernetes dependencies beyond the API types is what lets all three import it without any of them importing each other. Co-Authored-By: Claude Opus 5 (1M context) --- operator/internal/erasurecoding/scheme.go | 121 ++++++++++++++++ .../internal/erasurecoding/scheme_test.go | 131 ++++++++++++++++++ 2 files changed, 252 insertions(+) create mode 100644 operator/internal/erasurecoding/scheme.go create mode 100644 operator/internal/erasurecoding/scheme_test.go diff --git a/operator/internal/erasurecoding/scheme.go b/operator/internal/erasurecoding/scheme.go new file mode 100644 index 000000000..5197dbaec --- /dev/null +++ b/operator/internal/erasurecoding/scheme.go @@ -0,0 +1,121 @@ +// The erasure-coding schemes simplyblock supports, and how many storage nodes +// each one needs before a cluster may be built with it. +// +// Both facts are mirrored from elsewhere, and neither is enforced where it is +// first written down. The scheme set is simplyblock_core's +// SUPPORTED_ERASURE_CODING_SCHEMES, which the control plane checks on the +// create call — by which point the operator has already written a +// StorageCluster and an administrator has already approved the document that +// produced it. The minimum-node table is the product documentation's, and the +// control plane checks nothing resembling it: its activation gate counts +// devices (ndcs+npcs+1) and never nodes, so a cluster with too few nodes for +// its stripe activates and loses data on the first failure it was configured to +// survive. +// +// The rules live in a package of their own because three unrelated callers ask +// them: the deployment config's validation and its approval webhook, the +// cluster operation's activation gate, and the discovery run that proposes a +// scheme for a fleet it has just inspected. A leaf package with no Kubernetes +// dependencies beyond the API types is what lets all three import it. + +package erasurecoding + +import ( + "fmt" + "strings" + + "github.com/simplyblock/atlas/ptr" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// Scheme is one erasure-coding layout: the data chunks a stripe carries and the +// parity chunks protecting them, which the documentation writes as k+m and the +// control plane as distr_ndcs and distr_npcs. +type Scheme struct { + // DataChunks is ndcs, the k of the notation. + DataChunks int + + // ParityChunks is npcs, the m of the notation, and therefore how many + // simultaneous failures the cluster survives. + ParityChunks int +} + +// supported is the control plane's set, in the order the documentation lists +// it: the failure tolerances in turn, and the data widths ascending within +// each. The order is the one a reviewer reading a refusal sees, so it is the +// documentation's rather than sorted. +var supported = []Scheme{ + {DataChunks: 1, ParityChunks: 0}, + {DataChunks: 1, ParityChunks: 1}, + {DataChunks: 2, ParityChunks: 1}, + {DataChunks: 4, ParityChunks: 1}, + {DataChunks: 1, ParityChunks: 2}, + {DataChunks: 2, ParityChunks: 2}, + {DataChunks: 4, ParityChunks: 2}, +} + +// Supported is every scheme a cluster may be built with. +func Supported() []Scheme { + out := make([]Scheme, len(supported)) + copy(out, supported) + return out +} + +// IsSupported reports whether the control plane would accept this scheme. +func (s Scheme) IsSupported() bool { + for _, candidate := range supported { + if candidate == s { + return true + } + } + return false +} + +// MinimumNodes is how many storage nodes a cluster needs before it may be built +// with this scheme. +// +// It is ndcs+npcs nodes to place a stripe across, plus one spare node per +// tolerated failure for the rebuild to land on: a cluster with exactly +// ndcs+npcs nodes survives a failure with no node left to reconstruct the lost +// chunks onto, so it is one failure from unprotected and a second from data +// loss. That reproduces the documented table exactly, including the 1+0 that +// protects nothing and therefore needs no spare. +func (s Scheme) MinimumNodes() int { + return s.DataChunks + 2*s.ParityChunks +} + +// String is the k+m notation the documentation, the control plane's Mod column, +// and StorageCluster.status.erasureCodingScheme all use. +func (s Scheme) String() string { + return fmt.Sprintf("%d+%d", s.DataChunks, s.ParityChunks) +} + +// SchemeOf reads a cluster's stripe, filling in what an unstated one means. +// +// An absent stripe, or one stating only half of itself, is 1+1: that is what +// the control plane's API defaults each field to, and what the operator sends +// for a cluster whose spec says nothing. A document that never mentions erasure +// coding therefore describes a cluster needing three nodes, which is the whole +// reason the default has to be resolved here rather than left unknown. +func SchemeOf(stripe *simplyblockv1alpha2.StripeSpec) Scheme { + if stripe == nil { + return Scheme{DataChunks: 1, ParityChunks: 1} + } + return Scheme{ + DataChunks: ptr.IntFrom(stripe.DataChunks, 1), + ParityChunks: ptr.IntFrom(stripe.ParityChunks, 1), + } +} + +// SupportedNotation renders the supported set as one phrase for a message, +// because a refusal that names the scheme it rejected without naming the +// alternatives leaves the reviewer to go and find them. +func SupportedNotation() string { + rendered := make([]string, 0, len(supported)) + for _, scheme := range supported { + rendered = append(rendered, scheme.String()) + } + last := len(rendered) - 1 + return strings.Join(rendered[:last], ", ") + ", and " + rendered[last] +} diff --git a/operator/internal/erasurecoding/scheme_test.go b/operator/internal/erasurecoding/scheme_test.go new file mode 100644 index 000000000..69489b368 --- /dev/null +++ b/operator/internal/erasurecoding/scheme_test.go @@ -0,0 +1,131 @@ +// The schemes and the minimum-node table, checked against the two authorities +// they are mirrored from: simplyblock_core's SUPPORTED_ERASURE_CODING_SCHEMES +// and the product documentation's table. +// +// Both are written out here rather than derived, because a test that computes +// its expectation the way the code does would pass whatever the code said. + +package erasurecoding + +import ( + "testing" + + "github.com/simplyblock/atlas/ptr" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// The seven pairs simplyblock_core's SUPPORTED_ERASURE_CODING_SCHEMES holds, and +// nothing else. A scheme the control plane refuses is one a cluster is never +// built with, so admitting it here only moves the refusal to a place where the +// document is already immutable. +func TestTheSupportedSchemesAreTheControlPlanes(t *testing.T) { + want := []Scheme{ + {DataChunks: 1, ParityChunks: 0}, + {DataChunks: 1, ParityChunks: 1}, + {DataChunks: 2, ParityChunks: 1}, + {DataChunks: 4, ParityChunks: 1}, + {DataChunks: 1, ParityChunks: 2}, + {DataChunks: 2, ParityChunks: 2}, + {DataChunks: 4, ParityChunks: 2}, + } + + got := Supported() + if len(got) != len(want) { + t.Fatalf("Supported() has %d schemes, want %d: %v", len(got), len(want), got) + } + for i := range want { + if got[i] != want[i] { + t.Errorf("Supported()[%d] = %s, want %s", i, got[i], want[i]) + } + if !want[i].IsSupported() { + t.Errorf("%s reports itself unsupported", want[i]) + } + } +} + +// Everything else is refused, including the schemes that merely look plausible: +// a 3+1 is neither in the control plane's set nor buildable by it. +func TestAnUnlistedSchemeIsNotSupported(t *testing.T) { + for _, scheme := range []Scheme{ + {DataChunks: 3, ParityChunks: 1}, + {DataChunks: 8, ParityChunks: 2}, + {DataChunks: 2, ParityChunks: 0}, + {DataChunks: 4, ParityChunks: 0}, + {DataChunks: 1, ParityChunks: 3}, + {DataChunks: 0, ParityChunks: 1}, + } { + if scheme.IsSupported() { + t.Errorf("%s is reported supported", scheme) + } + } +} + +// The minimum-node table of the product documentation, verbatim. The formula +// behind it is ndcs+npcs nodes to place a stripe on plus one spare per tolerated +// failure, and the table is what the formula is checked against rather than the +// other way around. +func TestTheMinimumNodeCountIsTheDocumentedOne(t *testing.T) { + for _, testCase := range []struct { + scheme Scheme + nodes int + }{ + {Scheme{DataChunks: 1, ParityChunks: 0}, 1}, + {Scheme{DataChunks: 1, ParityChunks: 1}, 3}, + {Scheme{DataChunks: 2, ParityChunks: 1}, 4}, + {Scheme{DataChunks: 4, ParityChunks: 1}, 6}, + {Scheme{DataChunks: 1, ParityChunks: 2}, 5}, + {Scheme{DataChunks: 2, ParityChunks: 2}, 6}, + {Scheme{DataChunks: 4, ParityChunks: 2}, 8}, + } { + if got := testCase.scheme.MinimumNodes(); got != testCase.nodes { + t.Errorf("%s needs %d nodes, want %d", testCase.scheme, got, testCase.nodes) + } + } +} + +// A scheme renders as the notation the documentation, the control plane, and +// the cluster's own status all use, so a message names what a reviewer can look +// up. +func TestASchemeRendersAsTheDocumentedNotation(t *testing.T) { + if got := (Scheme{DataChunks: 2, ParityChunks: 1}).String(); got != "2+1" { + t.Errorf("String() = %q, want %q", got, "2+1") + } +} + +// An unstated stripe is 1+1, which is what the control plane defaults to and +// what the operator has always sent for a cluster that states none. A document +// that says nothing about erasure coding therefore describes a cluster needing +// three nodes, and that is the thing the minimum has to be checked against. +func TestAnUnstatedStripeIsOnePlusOne(t *testing.T) { + for name, stripe := range map[string]*simplyblockv1alpha2.StripeSpec{ + "absent": nil, + "empty": {}, + "parity alone": {ParityChunks: ptr.To(int32(1))}, + "data chunks alone": {DataChunks: ptr.To(int32(1))}, + } { + if got := SchemeOf(stripe); got != (Scheme{DataChunks: 1, ParityChunks: 1}) { + t.Errorf("SchemeOf(%s) = %s, want 1+1", name, got) + } + } +} + +// A stated stripe is read as it is written, including the 1+0 that is the only +// scheme a fleet of one or two workers can carry. +func TestAStatedStripeIsReadAsItIsWritten(t *testing.T) { + stripe := &simplyblockv1alpha2.StripeSpec{ + DataChunks: ptr.To(int32(1)), ParityChunks: ptr.To(int32(0)), + } + if got := SchemeOf(stripe); got != (Scheme{DataChunks: 1, ParityChunks: 0}) { + t.Errorf("SchemeOf(1+0) = %s, want 1+0", got) + } +} + +// The supported set renders as one phrase, because a refusal that names the +// scheme it rejected without naming the alternatives leaves the reviewer to +// find them. +func TestTheSupportedSetRendersAsAPhrase(t *testing.T) { + if got := SupportedNotation(); got != "1+0, 1+1, 2+1, 4+1, 1+2, 2+2, and 4+2" { + t.Errorf("SupportedNotation() = %q", got) + } +} From 9c52668d0164b67ad67798d9ef933a5479bb8ce1 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Fri, 18 Sep 2026 22:09:59 +0200 Subject: [PATCH 080/206] feat(api): the erasure-coding scheme is checked where it is written A stripe naming a scheme simplyblock does not support was storable, and the refusal came from the control plane on the cluster create. For a cluster a deployment config produced, that is several steps and one irreversible approval after the apply that stated it: the document is immutable once approved, so the reviewer who could have corrected the scheme is told about it after the point they could act. The CEL rule holds the pair to the same seven the control plane does, and it reads an unstated half as 1, because the fields are optional and a scheme with one of them omitted is still a scheme. The node minimum the same schemes carry is not expressible here and the comment says so: the nodes are objects of their own, and a cluster is created before any of them exists. It is answered by the deployment config's validation, by its approval webhook, and by the cluster's activation gate, which is where the count is knowable. Regenerates the four committed copies of each CRD. Co-Authored-By: Claude Opus 5 (1M context) --- ...mplyblock.io_clusterdeploymentconfigs.yaml | 7 ++ ...torage.simplyblock.io_storageclusters.yaml | 5 ++ operator/api/v1alpha1/storagecluster_types.go | 5 ++ operator/api/v1alpha2/storagecluster_types.go | 15 +++++ ...mplyblock.io_clusterdeploymentconfigs.yaml | 7 ++ ...torage.simplyblock.io_storageclusters.yaml | 5 ++ operator/dist/install.yaml | 66 +++++++++++++++---- ...mplyblock.io_clusterdeploymentconfigs.yaml | 7 ++ ...torage.simplyblock.io_storageclusters.yaml | 5 ++ 9 files changed, 111 insertions(+), 11 deletions(-) diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml index 7f68f56a5..b2a648ceb 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_clusterdeploymentconfigs.yaml @@ -189,6 +189,13 @@ spec: minimum: 0 type: integer type: object + x-kubernetes-validations: + - message: the erasure-coding scheme must be one of 1+0, 1+1, + 2+1, 4+1, 1+2, 2+2, or 4+2, written as dataChunks+parityChunks, + and an unstated half is 1 + rule: '[has(self.dataChunks) ? self.dataChunks : 1, has(self.parityChunks) + ? self.parityChunks : 1] in [[1, 0], [1, 1], [2, 1], [4, 1], + [1, 2], [2, 2], [4, 2]]' vcpuCount: description: |- VCPUCount is the number of vCPUs allocated to SPDK on each storage node of diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusters.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusters.yaml index 5abed659d..49335795d 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusters.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusters.yaml @@ -200,6 +200,9 @@ spec: format: int32 type: integer type: object + x-kubernetes-validations: + - message: the erasure-coding scheme must be one of 1+0, 1+1, 2+1, 4+1, 1+2, 2+2, or 4+2, written as dataChunks+parityChunks, and an unstated half is 1 + rule: '[has(self.dataChunks) ? self.dataChunks : 1, has(self.parityChunks) ? self.parityChunks : 1] in [[1, 0], [1, 1], [2, 1], [4, 1], [1, 2], [2, 2], [4, 2]]' vcpuCount: description: |- VCPUCount is the number of vCPUs allocated to SPDK on each storage node. @@ -1153,6 +1156,8 @@ spec: x-kubernetes-validations: - message: field is immutable rule: self == oldSelf + - message: the erasure-coding scheme must be one of 1+0, 1+1, 2+1, 4+1, 1+2, 2+2, or 4+2, written as dataChunks+parityChunks, and an unstated half is 1 + rule: '[has(self.dataChunks) ? self.dataChunks : 1, has(self.parityChunks) ? self.parityChunks : 1] in [[1, 0], [1, 1], [2, 1], [4, 1], [1, 2], [2, 2], [4, 2]]' vcpuCount: description: |- VCPUCount is the number of vCPUs allocated to SPDK on each storage node, diff --git a/operator/api/v1alpha1/storagecluster_types.go b/operator/api/v1alpha1/storagecluster_types.go index c95f7b535..d2fc85495 100644 --- a/operator/api/v1alpha1/storagecluster_types.go +++ b/operator/api/v1alpha1/storagecluster_types.go @@ -27,6 +27,11 @@ type CapacityThresholdSpec struct { ProvisionedCapacity *int32 `json:"provisionedCapacity,omitempty"` } +// StripeSpec is the erasure-coding layout. The rule is the hub type's, declared +// here as well because the apiserver validates the version an apply was written +// in: a scheme the control plane refuses would otherwise reach the cluster +// through this version and be refused by the cluster create instead. +// +kubebuilder:validation:XValidation:rule="[has(self.dataChunks) ? self.dataChunks : 1, has(self.parityChunks) ? self.parityChunks : 1] in [[1, 0], [1, 1], [2, 1], [4, 1], [1, 2], [2, 2], [4, 2]]",message="the erasure-coding scheme must be one of 1+0, 1+1, 2+1, 4+1, 1+2, 2+2, or 4+2, written as dataChunks+parityChunks, and an unstated half is 1" type StripeSpec struct { // +operator-sdk:csv:customresourcedefinitions:type=spec,displayName="Data Chunks" // DataChunks defines the number of data chunks in the erasure-coding layout. diff --git a/operator/api/v1alpha2/storagecluster_types.go b/operator/api/v1alpha2/storagecluster_types.go index db9907fbb..4e0414e68 100644 --- a/operator/api/v1alpha2/storagecluster_types.go +++ b/operator/api/v1alpha2/storagecluster_types.go @@ -93,6 +93,21 @@ const ( // StripeSpec is the erasure-coding layout: how many data chunks a stripe // carries and how many parity chunks protect them. +// +// The pair is one of the seven schemes simplyblock supports, and the rule below +// is the same set the control plane holds in SUPPORTED_ERASURE_CODING_SCHEMES. +// It is stated here as well because the control plane's refusal arrives at the +// cluster create, which is several steps and — for a cluster a deployment config +// produced — one irreversible approval after the apply that stated the scheme. +// +// Each scheme also has a storage-node count below which it must not be used, +// which is ndcs+npcs nodes to place a stripe across plus one spare per tolerated +// failure to rebuild onto: 1 for 1+0, 3 for 1+1, 4 for 2+1, 6 for 4+1, 5 for +// 1+2, 6 for 2+2, and 8 for 4+2. That is not expressible here, because the nodes +// are objects of their own and a cluster is created before any of them exists. +// It is answered by the deployment config's validation, by its approval webhook, +// and by the cluster's activation gate. +// +kubebuilder:validation:XValidation:rule="[has(self.dataChunks) ? self.dataChunks : 1, has(self.parityChunks) ? self.parityChunks : 1] in [[1, 0], [1, 1], [2, 1], [4, 1], [1, 2], [2, 2], [4, 2]]",message="the erasure-coding scheme must be one of 1+0, 1+1, 2+1, 4+1, 1+2, 2+2, or 4+2, written as dataChunks+parityChunks, and an unstated half is 1" type StripeSpec struct { // DataChunks is the number of data chunks per stripe (ndcs). // +kubebuilder:validation:Minimum=1 diff --git a/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml b/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml index 7f68f56a5..b2a648ceb 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_clusterdeploymentconfigs.yaml @@ -189,6 +189,13 @@ spec: minimum: 0 type: integer type: object + x-kubernetes-validations: + - message: the erasure-coding scheme must be one of 1+0, 1+1, + 2+1, 4+1, 1+2, 2+2, or 4+2, written as dataChunks+parityChunks, + and an unstated half is 1 + rule: '[has(self.dataChunks) ? self.dataChunks : 1, has(self.parityChunks) + ? self.parityChunks : 1] in [[1, 0], [1, 1], [2, 1], [4, 1], + [1, 2], [2, 2], [4, 2]]' vcpuCount: description: |- VCPUCount is the number of vCPUs allocated to SPDK on each storage node of diff --git a/operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml b/operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml index 5abed659d..49335795d 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml @@ -200,6 +200,9 @@ spec: format: int32 type: integer type: object + x-kubernetes-validations: + - message: the erasure-coding scheme must be one of 1+0, 1+1, 2+1, 4+1, 1+2, 2+2, or 4+2, written as dataChunks+parityChunks, and an unstated half is 1 + rule: '[has(self.dataChunks) ? self.dataChunks : 1, has(self.parityChunks) ? self.parityChunks : 1] in [[1, 0], [1, 1], [2, 1], [4, 1], [1, 2], [2, 2], [4, 2]]' vcpuCount: description: |- VCPUCount is the number of vCPUs allocated to SPDK on each storage node. @@ -1153,6 +1156,8 @@ spec: x-kubernetes-validations: - message: field is immutable rule: self == oldSelf + - message: the erasure-coding scheme must be one of 1+0, 1+1, 2+1, 4+1, 1+2, 2+2, or 4+2, written as dataChunks+parityChunks, and an unstated half is 1 + rule: '[has(self.dataChunks) ? self.dataChunks : 1, has(self.parityChunks) ? self.parityChunks : 1] in [[1, 0], [1, 1], [2, 1], [4, 1], [1, 2], [2, 2], [4, 2]]' vcpuCount: description: |- VCPUCount is the number of vCPUs allocated to SPDK on each storage node, diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index 243cc6ee0..cfe93c684 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -862,6 +862,13 @@ spec: minimum: 0 type: integer type: object + x-kubernetes-validations: + - message: the erasure-coding scheme must be one of 1+0, 1+1, + 2+1, 4+1, 1+2, 2+2, or 4+2, written as dataChunks+parityChunks, + and an unstated half is 1 + rule: '[has(self.dataChunks) ? self.dataChunks : 1, has(self.parityChunks) + ? self.parityChunks : 1] in [[1, 0], [1, 1], [2, 1], [4, 1], + [1, 2], [2, 2], [4, 2]]' vcpuCount: description: |- VCPUCount is the number of vCPUs allocated to SPDK on each storage node of @@ -5153,6 +5160,13 @@ spec: format: int32 type: integer type: object + x-kubernetes-validations: + - message: the erasure-coding scheme must be one of 1+0, 1+1, 2+1, + 4+1, 1+2, 2+2, or 4+2, written as dataChunks+parityChunks, and + an unstated half is 1 + rule: '[has(self.dataChunks) ? self.dataChunks : 1, has(self.parityChunks) + ? self.parityChunks : 1] in [[1, 0], [1, 1], [2, 1], [4, 1], [1, + 2], [2, 2], [4, 2]]' vcpuCount: description: |- VCPUCount is the number of vCPUs allocated to SPDK on each storage node. @@ -6139,6 +6153,12 @@ spec: x-kubernetes-validations: - message: field is immutable rule: self == oldSelf + - message: the erasure-coding scheme must be one of 1+0, 1+1, 2+1, + 4+1, 1+2, 2+2, or 4+2, written as dataChunks+parityChunks, and + an unstated half is 1 + rule: '[has(self.dataChunks) ? self.dataChunks : 1, has(self.parityChunks) + ? self.parityChunks : 1] in [[1, 0], [1, 1], [2, 1], [4, 1], [1, + 2], [2, 2], [4, 2]]' vcpuCount: description: |- VCPUCount is the number of vCPUs allocated to SPDK on each storage node, @@ -6519,10 +6539,12 @@ spec: rule: '!has(self.state) || self.state in [''Claiming'',''CheckingControlPlane'',''ResolvingConfig'',''Creating'',''Adopting'',''Persisting'']' tasks: description: |- - Tasks are the control plane's running and pending jobs, newest first and - capped at twenty. Completed and canceled tasks are not here: they leave - the list and become events, so the length tracks concurrency rather than - history. + Tasks are the control plane's running and pending jobs, capped at twenty + and in the order the control plane reports them: its TaskDTO carries no + creation date, so newest-first is not orderable from what is on the wire + (design-storagecluster.md §12.1). Completed and canceled tasks are not + here: they leave the list and become events, so the length tracks + concurrency rather than history. items: description: |- ClusterTask is one asynchronous job the control plane is running, as of the @@ -9475,9 +9497,10 @@ spec: properties: allowedNodes: description: |- - AllowedNodes restricts which storage nodes may host this pool's volumes. - Empty means every node in the cluster. Narrowing it stops new volumes - landing on the removed nodes and leaves the existing ones where they are. + AllowedNodes restricts which hosts may carry this pool's volumes, by + Kubernetes Node name. Empty means every node in the cluster. Narrowing it + stops new volumes landing on the removed nodes and leaves the existing + ones where they are. The list is left exactly as authored: a name that no longer resolves is dropped from Status.AllowedNodes rather than pruned from here, so a node @@ -9661,10 +9684,11 @@ spec: type: string allowedNodes: description: |- - AllowedNodes is Spec.AllowedNodes resolved against the StorageNodes that - exist, which is what the control plane is sent. An empty list here is not - the same as an absent Spec.AllowedNodes: absent means every node, and - empty after resolution means the pool can place nothing. + AllowedNodes is Spec.AllowedNodes resolved against the Node objects that + exist, which is what the control plane's host list and the per-pool node + labels are derived from. An empty list here is not the same as an absent + Spec.AllowedNodes: absent means every node, and empty after resolution + means the pool can place nothing. items: type: string type: array @@ -11424,6 +11448,26 @@ webhooks: resources: - controlplaneops sideEffects: None +- admissionReviewVersions: + - v1 + clientConfig: + service: + name: simplyblock-operator-webhook-service + namespace: simplyblock-operator-system + path: /validate-storage-simplyblock-io-v1alpha2-operatorops + failurePolicy: Fail + name: voperatorops.simplyblock.io + rules: + - apiGroups: + - storage.simplyblock.io + apiVersions: + - v1alpha2 + operations: + - CREATE + - UPDATE + resources: + - operatorops + sideEffects: None - admissionReviewVersions: - v1 clientConfig: diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml index 7f68f56a5..b2a648ceb 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_clusterdeploymentconfigs.yaml @@ -189,6 +189,13 @@ spec: minimum: 0 type: integer type: object + x-kubernetes-validations: + - message: the erasure-coding scheme must be one of 1+0, 1+1, + 2+1, 4+1, 1+2, 2+2, or 4+2, written as dataChunks+parityChunks, + and an unstated half is 1 + rule: '[has(self.dataChunks) ? self.dataChunks : 1, has(self.parityChunks) + ? self.parityChunks : 1] in [[1, 0], [1, 1], [2, 1], [4, 1], + [1, 2], [2, 2], [4, 2]]' vcpuCount: description: |- VCPUCount is the number of vCPUs allocated to SPDK on each storage node of diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusters.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusters.yaml index 5abed659d..49335795d 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusters.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusters.yaml @@ -200,6 +200,9 @@ spec: format: int32 type: integer type: object + x-kubernetes-validations: + - message: the erasure-coding scheme must be one of 1+0, 1+1, 2+1, 4+1, 1+2, 2+2, or 4+2, written as dataChunks+parityChunks, and an unstated half is 1 + rule: '[has(self.dataChunks) ? self.dataChunks : 1, has(self.parityChunks) ? self.parityChunks : 1] in [[1, 0], [1, 1], [2, 1], [4, 1], [1, 2], [2, 2], [4, 2]]' vcpuCount: description: |- VCPUCount is the number of vCPUs allocated to SPDK on each storage node. @@ -1153,6 +1156,8 @@ spec: x-kubernetes-validations: - message: field is immutable rule: self == oldSelf + - message: the erasure-coding scheme must be one of 1+0, 1+1, 2+1, 4+1, 1+2, 2+2, or 4+2, written as dataChunks+parityChunks, and an unstated half is 1 + rule: '[has(self.dataChunks) ? self.dataChunks : 1, has(self.parityChunks) ? self.parityChunks : 1] in [[1, 0], [1, 1], [2, 1], [4, 1], [1, 2], [2, 2], [4, 2]]' vcpuCount: description: |- VCPUCount is the number of vCPUs allocated to SPDK on each storage node, From 606936cb1daf40cd4f189f6c4fa7a7f8272e0b3b Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Fri, 18 Sep 2026 22:10:14 +0200 Subject: [PATCH 081/206] feat(cluster): an activation waits until the fleet can carry its stripe The control plane's activation gate counts devices, ndcs+npcs+1 of them, and never nodes. A fleet of four configured 4+2 passes it -- six devices are easy on four machines -- and what activates is a cluster whose stripe has nowhere to place its chunks and nothing to rebuild onto. It then loses data on the second failure it was configured to survive, which is the failure somebody chose 4+2 to survive. The rule is the product documentation's, and nothing in the system enforced it until now. It gates the activation rather than the create, because that is where the layout stops being a proposal. A StorageCluster is created before any of its nodes exists, so a node count at create time is a number about nothing; by activation the fleet is whatever it is going to be. Waiting rather than failing, and StripeNodesNotReady says which count is short of which. A fleet that is still coming up is the ordinary case, and an activation that failed on it would have to be raised again by hand for a condition that resolves itself. Co-Authored-By: Claude Opus 5 (1M context) --- .../internal/controllers/cluster/actions.go | 7 + .../cluster/cel_validation_test.go | 91 +++++++++++ .../controllers/cluster/erasurecoding.go | 108 +++++++++++++ .../controllers/cluster/erasurecoding_test.go | 146 ++++++++++++++++++ .../internal/controllers/cluster/events.go | 6 + 5 files changed, 358 insertions(+) create mode 100644 operator/internal/controllers/cluster/erasurecoding.go create mode 100644 operator/internal/controllers/cluster/erasurecoding_test.go diff --git a/operator/internal/controllers/cluster/actions.go b/operator/internal/controllers/cluster/actions.go index 3eaa3f51e..a9f2d22b6 100644 --- a/operator/internal/controllers/cluster/actions.go +++ b/operator/internal/controllers/cluster/actions.go @@ -81,6 +81,13 @@ func (r *StorageClusterOpsReconciler) request( if active { return true, nil } + // After the active check rather than before it, so that a re-activation + // of a cluster that is already serving is never held: the gate exists to + // stop a layout being brought up wrong, and a live cluster's layout is + // already a fact. + if ready, err := r.stripeNodesReady(ctx, ops); err != nil || !ready { + return false, err + } if err := r.API.Activate(ctx, clusterID); err != nil { return false, fmt.Errorf("activate cluster %s: %w", ops.Spec.ClusterRef, err) } diff --git a/operator/internal/controllers/cluster/cel_validation_test.go b/operator/internal/controllers/cluster/cel_validation_test.go index fa616df6c..ce8540435 100644 --- a/operator/internal/controllers/cluster/cel_validation_test.go +++ b/operator/internal/controllers/cluster/cel_validation_test.go @@ -387,3 +387,94 @@ func TestStorageClusterNameIsBoundedAtALabelsLimit(t *testing.T) { _ = apiClient.Delete(ctx, pool) } } + +// The erasure-coding scheme is one of the seven the control plane accepts, and +// the schema is what says so. +// +// It is CEL rather than a webhook because it is a statement about the document's +// structure: the pair is wrong or right on its own, without reference to any +// cluster, node, or fleet. What it buys is that the refusal arrives at the apply +// rather than from the control plane's cluster create, which is several steps +// and one approval later and leaves a StorageCluster nothing can use behind. +func TestStorageClusterCELAcceptsOnlyTheSupportedErasureCodingSchemes(t *testing.T) { + apiClient := apiServer(t) + + for _, testCase := range []struct { + data, parity int32 + denied bool + }{ + {data: 1, parity: 0}, + {data: 1, parity: 1}, + {data: 2, parity: 1}, + {data: 4, parity: 1}, + {data: 1, parity: 2}, + {data: 2, parity: 2}, + {data: 4, parity: 2}, + {data: 3, parity: 1, denied: true}, + {data: 8, parity: 2, denied: true}, + {data: 2, parity: 0, denied: true}, + {data: 4, parity: 0, denied: true}, + {data: 1, parity: 3, denied: true}, + {data: 16, parity: 4, denied: true}, + } { + cluster := &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{GenerateName: "stripe-", Namespace: "default"}, + Spec: simplyblockv1alpha2.StorageClusterSpec{ + MaxSubsystemCount: ptr.To(int32(10)), + VCPUCount: ptr.To(int32(6)), + Stripe: &simplyblockv1alpha2.StripeSpec{ + DataChunks: ptr.To(testCase.data), ParityChunks: ptr.To(testCase.parity), + }, + }, + } + + err := apiClient.Create(context.Background(), cluster) + switch { + case testCase.denied && err == nil: + t.Errorf("%d+%d was accepted, and the control plane refuses it", + testCase.data, testCase.parity) + case testCase.denied && !strings.Contains(err.Error(), "erasure-coding scheme"): + t.Errorf("%d+%d was refused for the wrong reason: %v", + testCase.data, testCase.parity, err) + case !testCase.denied && err != nil: + t.Errorf("%d+%d was refused: %v", testCase.data, testCase.parity, err) + } + if err == nil { + if err := apiClient.Delete(context.Background(), cluster); err != nil { + t.Fatalf("clean up: %v", err) + } + } + } +} + +// A cluster stating half a stripe is stating the control plane's default for the +// other half, so the pair the rule sees is the pair the cluster is created with. +func TestStorageClusterCELReadsAnUnstatedHalfAsTheDefault(t *testing.T) { + apiClient := apiServer(t) + + for _, stripe := range []*simplyblockv1alpha2.StripeSpec{ + {}, + {DataChunks: ptr.To(int32(2))}, + {ParityChunks: ptr.To(int32(0))}, + } { + cluster := &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{GenerateName: "stripe-half-", Namespace: "default"}, + Spec: simplyblockv1alpha2.StorageClusterSpec{ + MaxSubsystemCount: ptr.To(int32(10)), + VCPUCount: ptr.To(int32(6)), + Stripe: stripe, + }, + } + + // {} is 1+1, {dataChunks: 2} is 2+1, and {parityChunks: 0} is 1+0: all + // three are supported, and a rule reading an absent field as absent + // rather than as its default would refuse or admit the wrong ones. + if err := apiClient.Create(context.Background(), cluster); err != nil { + t.Errorf("a half-stated stripe %+v was refused: %v", stripe, err) + continue + } + if err := apiClient.Delete(context.Background(), cluster); err != nil { + t.Fatalf("clean up: %v", err) + } + } +} diff --git a/operator/internal/controllers/cluster/erasurecoding.go b/operator/internal/controllers/cluster/erasurecoding.go new file mode 100644 index 000000000..b2130f9f0 --- /dev/null +++ b/operator/internal/controllers/cluster/erasurecoding.go @@ -0,0 +1,108 @@ +// The activation gate for a cluster whose storage nodes are too few for its +// erasure-coding scheme. +// +// Unlike the failure-domain rules next door, this one is not mirrored from the +// control plane, because the control plane does not have it. Its own activation +// gate counts devices — ndcs+npcs+1 of them — and never nodes, so a fleet of +// four configured 4+2 passes it: six devices are easy on four nodes, and the +// stripe still has nowhere to place its chunks and nothing to rebuild onto. The +// rule enforced here is the one the product documentation states, which nothing +// in the system enforced until this file. +// +// It gates an activation rather than a cluster create, because that is where the +// layout stops being a proposal: a StorageCluster is created before its nodes +// exist, so a node count at create time is a number about nothing. The +// deployment config's validation and its approval webhook answer the same +// question earlier, against the document; this answers it for the cluster and +// the nodes themselves, however they were written. + +package cluster + +import ( + "context" + "fmt" + + corev1 "k8s.io/api/core/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/erasurecoding" +) + +// ActivationNodeCountViolation reports why a cluster with this many storage +// nodes must not be activated with this stripe, and the empty string when it +// may be. +// +// An unstated stripe is 1+1, which is what the control plane defaults to, so a +// cluster that says nothing about erasure coding is held to the three nodes 1+1 +// needs rather than to nothing. +func ActivationNodeCountViolation(stripe *simplyblockv1alpha2.StripeSpec, nodes int) string { + scheme := erasurecoding.SchemeOf(stripe) + minimum := scheme.MinimumNodes() + if nodes >= minimum { + return "" + } + + if scheme.ParityChunks == 0 { + return fmt.Sprintf( + "the cluster's stripe is %s, which needs at least %d storage node, and the "+ + "cluster has %d", scheme, minimum, nodes) + } + return fmt.Sprintf( + "the cluster's stripe is %s, which needs at least %d storage nodes (%d to place "+ + "a stripe across and %d to rebuild onto after a failure), and the cluster "+ + "has %d", scheme, minimum, scheme.DataChunks+scheme.ParityChunks, + scheme.ParityChunks, nodes) +} + +// stripeNodesReady gates an activation on the cluster having the storage nodes +// its scheme requires, and reports the hold rather than failing. +// +// It holds rather than fails for the reason the failure-domain gate beside it +// does: a cluster whose nodes are still being created is a cluster that will +// meet the rule shortly, and an operation that failed on a count taken too early +// is one somebody has to notice and re-issue. What it counts is the StorageNode +// objects of the cluster rather than what the control plane reports, because the +// objects are what the deployment states it will have; whether each of them is +// online is the expansion's own gate, one step earlier. +func (r *StorageClusterOpsReconciler) stripeNodesReady( + ctx context.Context, ops *simplyblockv1alpha2.StorageClusterOps, +) (bool, error) { + var cluster simplyblockv1alpha2.StorageCluster + key := client.ObjectKey{Name: ops.Spec.ClusterRef, Namespace: ops.Namespace} + if err := r.Get(ctx, key, &cluster); err != nil { + return false, err + } + + nodes, err := storageNodeCount(ctx, r.Client, cluster.Namespace, cluster.Name) + if err != nil { + return false, err + } + + reason := ActivationNodeCountViolation(cluster.Spec.Stripe, nodes) + if reason == "" { + return true, nil + } + + r.Recorder.Eventf(ops, nil, corev1.EventTypeWarning, + StripeNodesNotReady, StripeNodesNotReady, + "The activation is waiting on the nodes the erasure coding requires: %s", reason) + return false, nil +} + +// storageNodeCount is how many storage nodes belong to one cluster. +func storageNodeCount( + ctx context.Context, c client.Client, namespace, clusterName string, +) (int, error) { + var nodes simplyblockv1alpha2.StorageNodeList + if err := c.List(ctx, &nodes, client.InNamespace(namespace)); err != nil { + return 0, fmt.Errorf("read the nodes of cluster %s: %w", clusterName, err) + } + count := 0 + for i := range nodes.Items { + if nodes.Items[i].Spec.ClusterRef == clusterName { + count++ + } + } + return count, nil +} diff --git a/operator/internal/controllers/cluster/erasurecoding_test.go b/operator/internal/controllers/cluster/erasurecoding_test.go new file mode 100644 index 000000000..584ec5bcd --- /dev/null +++ b/operator/internal/controllers/cluster/erasurecoding_test.go @@ -0,0 +1,146 @@ +// The activation gate for a cluster whose fleet is too small for its stripe. +// +// It is the backstop the deployment config's checks cannot be: a cluster and its +// nodes can be written by hand, by a chart, or by a document from before the +// checks existed, and an activation is the moment the layout stops being a +// proposal. The control plane refuses none of it — its own gate counts devices +// rather than nodes — so a four-node cluster configured 4+2 activates, and what +// it bought is a cluster that loses data on the second failure it was configured +// to survive. + +package cluster + +import ( + "testing" + + "github.com/simplyblock/atlas/ptr" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/webapi" +) + +// suspendedCluster is a control-plane reading of a cluster that is not active, +// which is the only state an activation is asked for from. +func suspendedCluster() webapi.ClusterResponse { + reading := activeCluster() + reading.Status = "suspended" + return reading +} + +// nodeOfTestCluster is one storage node of the cluster under test. +func nodeOfTestCluster(name, worker string) *simplyblockv1alpha2.StorageNode { + return &simplyblockv1alpha2.StorageNode{ + ObjectMeta: objectMeta(name), + Spec: simplyblockv1alpha2.StorageNodeSpec{ + ClusterRef: testClusterName, + WorkerNode: worker, + Slot: ptr.To(int32(0)), + }, + } +} + +// withStripe states a cluster's erasure coding. +func withStripe(data, parity int32) func(*simplyblockv1alpha2.StorageCluster) { + return func(c *simplyblockv1alpha2.StorageCluster) { + c.Spec.Stripe = &simplyblockv1alpha2.StripeSpec{ + DataChunks: ptr.To(data), ParityChunks: ptr.To(parity), + } + c.Status.Status = "suspended" + c.Status.Phase = simplyblockv1alpha2.StorageClusterPhasePending + } +} + +// A 1+1 cluster needs three storage nodes: two to place a stripe across and one +// spare to rebuild onto. Two nodes is a cluster that survives no failure it was +// configured to survive, so the activation waits rather than firing. +func TestAnActivationBelowTheStripesMinimumIsHeld(t *testing.T) { + api := &fakeControlPlane{cluster: func(string) (webapi.ClusterResponse, error) { + return suspendedCluster(), nil + }} + rec := &recorder{} + r := newOpsReconciler(t, api, rec, + newTestCluster(withStripe(1, 1)), + newTestOps(simplyblockv1alpha2.StorageClusterOpsActionActivate), + nodeOfTestCluster("node-1", "worker-1"), + nodeOfTestCluster("node-2", "worker-2")) + + ops, _ := reconcileOps(t, r, 6) + if api.activateCalls != 0 { + t.Errorf("the control plane was asked to activate %d times, want 0 below the minimum", + api.activateCalls) + } + if ops.Status.Phase == simplyblockv1alpha2.StorageClusterOpsPhaseSucceeded { + t.Errorf("the operation reported success without activating anything") + } + if !rec.has(StripeNodesNotReady) { + t.Errorf("nothing said why the activation is waiting; events: %+v", rec.events) + } +} + +// The same cluster with the third node is activated, because the hold is on the +// node count and not on the stripe. +func TestAnActivationWithTheNodesTheStripeNeedsIsRequested(t *testing.T) { + api := &fakeControlPlane{cluster: func(string) (webapi.ClusterResponse, error) { + return suspendedCluster(), nil + }} + r := newOpsReconciler(t, api, &recorder{}, + newTestCluster(withStripe(1, 1)), + newTestOps(simplyblockv1alpha2.StorageClusterOpsActionActivate), + nodeOfTestCluster("node-1", "worker-1"), + nodeOfTestCluster("node-2", "worker-2"), + nodeOfTestCluster("node-3", "worker-3")) + + reconcileOps(t, r, 6) + if api.activateCalls != 1 { + t.Errorf("the control plane was asked to activate %d times, want 1", api.activateCalls) + } +} + +// A cluster that is already active is not held, whatever its node count says. +// The gate exists to stop a layout being brought up wrong, and a live cluster is +// one whose layout is already the fact; holding a re-activation on it would wedge +// a cluster that is serving. +func TestAReactivationOfALiveClusterIsNotHeld(t *testing.T) { + api := &fakeControlPlane{cluster: func(string) (webapi.ClusterResponse, error) { + return activeCluster(), nil + }} + r := newOpsReconciler(t, api, &recorder{}, + newTestCluster(func(c *simplyblockv1alpha2.StorageCluster) { + c.Spec.Stripe = &simplyblockv1alpha2.StripeSpec{ + DataChunks: ptr.To(int32(4)), ParityChunks: ptr.To(int32(2)), + } + }), + newTestOps(simplyblockv1alpha2.StorageClusterOpsActionActivate)) + + ops, _ := reconcileOps(t, r, 6) + if ops.Status.Phase != simplyblockv1alpha2.StorageClusterOpsPhaseSucceeded { + t.Errorf("phase = %q, want Succeeded (message: %s)", ops.Status.Phase, ops.Status.Message) + } +} + +// The rule itself, stated against the cases the callers cannot reach: a cluster +// with no nodes reported yet, and the 1+0 that needs one. +func TestTheActivationNodeCountRule(t *testing.T) { + for _, testCase := range []struct { + data, parity int32 + nodes int + held bool + }{ + {1, 0, 1, false}, + {1, 0, 0, true}, + {1, 1, 2, true}, + {1, 1, 3, false}, + {2, 1, 4, false}, + {4, 2, 7, true}, + {4, 2, 8, false}, + } { + scheme := simplyblockv1alpha2.StripeSpec{ + DataChunks: ptr.To(testCase.data), ParityChunks: ptr.To(testCase.parity), + } + reason := ActivationNodeCountViolation(&scheme, testCase.nodes) + if held := reason != ""; held != testCase.held { + t.Errorf("%d+%d with %d nodes: held = %v, want %v (%s)", + testCase.data, testCase.parity, testCase.nodes, held, testCase.held, reason) + } + } +} diff --git a/operator/internal/controllers/cluster/events.go b/operator/internal/controllers/cluster/events.go index 580e6184d..308c0a647 100644 --- a/operator/internal/controllers/cluster/events.go +++ b/operator/internal/controllers/cluster/events.go @@ -72,4 +72,10 @@ const ( // FailureDomainNotReady says an activation is waiting because the cluster's // failure domains do not yet hold an equal number of hosts. FailureDomainNotReady = "FailureDomainNotReady" + + // StripeNodesNotReady says an activation is waiting because the cluster has + // fewer storage nodes than its erasure-coding scheme needs. Nothing below the + // operator reports it: the control plane's own activation gate counts devices + // rather than nodes. + StripeNodesNotReady = "StripeNodesNotReady" ) From 032d58ccc6516da1b1d06fbf6232109659b61d4f Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Fri, 18 Sep 2026 22:10:31 +0200 Subject: [PATCH 082/206] feat(deployment): a draft is told its fleet is too small for its stripe, twice The scheme and the node count it needs are the one part of a deployment nothing below this operator checks. The control plane validates the scheme on the cluster create, which lands after the approval, and its activation gate counts devices rather than nodes, so a document naming 4+2 over four workers is admitted by everything and produces a cluster that loses data on the second failure it was configured to survive. Both moments, for the same reason the rest of this document's validation has two. A draft is told on every reconcile, while it is still editable and while a reviewer can act on it, which is the whole point of the gate. The approving edit is refused, because an approved document is immutable and telling somebody afterwards is telling them about something they can no longer change. The count is the nodes the document itself describes, which is what makes the answer available before any StorageNode exists: a draft states its workers and its slots, so the fleet it would produce is knowable from the document alone. StripeUnsupported and StripeBelowMinimumNodes are separate reasons because they ask for different corrections. The first is a scheme nothing supports and the edit is the scheme. The second is a scheme this fleet cannot carry, and the edit is either the scheme or the fleet. Co-Authored-By: Claude Opus 5 (1M context) --- .../deployment/cel_validation_test.go | 36 +++ ...clusterdeploymentconfig_controller_test.go | 16 +- .../controllers/deployment/erasurecoding.go | 238 ++++++++++++++++++ .../deployment/erasurecoding_test.go | 218 ++++++++++++++++ .../internal/controllers/deployment/events.go | 8 + .../controllers/deployment/expansion.go | 14 +- .../controllers/deployment/validation.go | 8 + .../clusterdeploymentconfig_validator.go | 19 +- .../clusterdeploymentconfig_validator_test.go | 73 +++++- 9 files changed, 621 insertions(+), 9 deletions(-) create mode 100644 operator/internal/controllers/deployment/erasurecoding.go create mode 100644 operator/internal/controllers/deployment/erasurecoding_test.go diff --git a/operator/internal/controllers/deployment/cel_validation_test.go b/operator/internal/controllers/deployment/cel_validation_test.go index 3bed1a3a3..59562fc49 100644 --- a/operator/internal/controllers/deployment/cel_validation_test.go +++ b/operator/internal/controllers/deployment/cel_validation_test.go @@ -169,3 +169,39 @@ func TestAnApprovedDocumentStillTakesLabelsAndStatus(t *testing.T) { t.Fatalf("the controller cannot report on an approved document: %v", err) } } + +// The scheme rule reaches this kind too, because the document states the +// cluster's stripe with the cluster's own type. A rule declared on the shared +// type and reaching only one of the two schemas would leave the document — the +// thing a reviewer approves — the unguarded one. +func TestTheDocumentsSchemaRefusesAnUnsupportedScheme(t *testing.T) { + apiClient := apiServer(t) + + config := &simplyblockv1alpha2.ClusterDeploymentConfig{ + ObjectMeta: metav1.ObjectMeta{GenerateName: "stripe-", Namespace: "default"}, + Spec: simplyblockv1alpha2.ClusterDeploymentConfigSpec{ + Cluster: &simplyblockv1alpha2.ClusterTemplate{ + Name: "a-cluster", + VCPUCount: ptr.To(int32(4)), + MaxSubsystemCount: ptr.To(int32(30)), + Stripe: &simplyblockv1alpha2.StripeSpec{ + DataChunks: ptr.To(int32(3)), ParityChunks: ptr.To(int32(1)), + }, + }, + NodeSets: []simplyblockv1alpha2.NodeSet{{ + Name: "discovered", + Groups: []simplyblockv1alpha2.NodeGroup{{ + Name: "group-1", Workers: []string{"worker-1"}, + }}, + }}, + }, + } + + err := apiClient.Create(context.Background(), config) + if err == nil { + t.Fatal("a document stating 3+1 was stored") + } + if !strings.Contains(err.Error(), "erasure-coding scheme") { + t.Fatalf("refused for the wrong reason: %v", err) + } +} diff --git a/operator/internal/controllers/deployment/clusterdeploymentconfig_controller_test.go b/operator/internal/controllers/deployment/clusterdeploymentconfig_controller_test.go index 0cf291c7e..78df80a6d 100644 --- a/operator/internal/controllers/deployment/clusterdeploymentconfig_controller_test.go +++ b/operator/internal/controllers/deployment/clusterdeploymentconfig_controller_test.go @@ -33,6 +33,11 @@ const ( // aDocument is a valid, approved two-worker NVMe deployment, so a case states only // what it is about. +// +// It states 1+0 because two workers carry no other scheme: every redundant one +// needs at least three storage nodes, and a document that says nothing about +// erasure coding means the control plane's 1+1. A case about the stripe says so +// by overwriting the field. func aDocument( mutate func(*simplyblockv1alpha2.ClusterDeploymentConfig), ) *simplyblockv1alpha2.ClusterDeploymentConfig { @@ -46,6 +51,9 @@ func aDocument( MaxSubsystemCount: ptr.To(int32(20)), VCPUCount: ptr.To(int32(8)), MinHugePagesSize: "100G", + Stripe: &simplyblockv1alpha2.StripeSpec{ + DataChunks: ptr.To(int32(1)), ParityChunks: ptr.To(int32(0)), + }, }, NodeSets: []simplyblockv1alpha2.NodeSet{{ Name: "rack-a", @@ -67,7 +75,10 @@ func aDocument( return config } -// aCluster is a StorageCluster the control plane has already created. +// aCluster is a StorageCluster the control plane has already created. It states +// the 1+0 the document fixture states, and for the same reason: a growth +// document is answered against its cluster's stripe, and two workers carry no +// redundant scheme. func aCluster( mutate func(*simplyblockv1alpha2.StorageCluster), ) *simplyblockv1alpha2.StorageCluster { @@ -76,6 +87,9 @@ func aCluster( Spec: simplyblockv1alpha2.StorageClusterSpec{ VCPUCount: ptr.To(int32(8)), MinHugePagesSize: "100G", + Stripe: &simplyblockv1alpha2.StripeSpec{ + DataChunks: ptr.To(int32(1)), ParityChunks: ptr.To(int32(0)), + }, }, Status: simplyblockv1alpha2.StorageClusterStatus{UUID: "cluster-uuid"}, } diff --git a/operator/internal/controllers/deployment/erasurecoding.go b/operator/internal/controllers/deployment/erasurecoding.go new file mode 100644 index 000000000..c3f58b935 --- /dev/null +++ b/operator/internal/controllers/deployment/erasurecoding.go @@ -0,0 +1,238 @@ +// What a document's erasure coding is answered against, and the counting the +// answer needs. +// +// The scheme decides how many storage nodes the deployment must have: ndcs+npcs +// of them to place a stripe across, plus one spare per tolerated failure for the +// rebuild to land on. Nothing below this operator checks it. The control plane +// validates the scheme itself on the cluster create and counts devices at +// activation, never nodes, so a four-worker fleet configured 4+2 activates, and +// what it has bought is a cluster that loses data on the second failure it was +// configured to survive. +// +// Both callers of this file are here for the same reason the rest of the +// document's validation has two: a draft is told on every reconcile so it can be +// edited, and the approving edit is refused because §3.2 makes an approved +// document immutable. StripeChecks is therefore exported and takes a reader +// rather than a reconciler, and the two callers differ only in what they wrap the +// result in. +// +// The counting is by slot rather than by worker because the expansion creates +// nodes by slot: a growth document naming a worker the cluster already has a node +// on produces nothing there, and counting that worker again would report a +// minimum as met by a node nobody is going to create. + +package deployment + +import ( + "context" + "fmt" + + apierrors "k8s.io/apimachinery/pkg/api/errors" + "sigs.k8s.io/controller-runtime/pkg/client" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/erasurecoding" +) + +// StripeCheck is one thing wrong with what a document says about erasure +// coding, carrying the reason it is counted and reported under as well as the +// sentence somebody reads. +type StripeCheck struct { + Reason string + Message string +} + +// footprint is a set of storage nodes, identified the way the expansion +// identifies them, and the workers they sit on. Both numbers are needed: the +// minimum is a node count, and a node count reached by running several nodes on +// one worker has not bought the independent spare the count was asking for. +type footprint struct { + slots map[slotKey]struct{} + workers map[string]struct{} +} + +func newFootprint() footprint { + return footprint{slots: map[slotKey]struct{}{}, workers: map[string]struct{}{}} +} + +func (f footprint) add(worker string, slot int32) { + f.slots[slotKey{worker: worker, slot: slot}] = struct{}{} + f.workers[worker] = struct{}{} +} + +func (f footprint) merge(other footprint) { + for key := range other.slots { + f.slots[key] = struct{}{} + } + for worker := range other.workers { + f.workers[worker] = struct{}{} + } +} + +// nodes is how many storage nodes the deployment ends up with. +func (f footprint) nodes() int { return len(f.slots) } + +// hosts is how many workers those nodes sit on. +func (f footprint) hosts() int { return len(f.workers) } + +// StripeChecks answers a document's erasure coding against the deployment it +// describes. +// +// It returns at most one check, unlike the validation it feeds, because the three +// things that can be wrong here are one thing each: a scheme the control plane +// refuses has no minimum to be short of, and a deployment short of the node count +// is not also separately short of the worker count. Reporting two of them would +// put two findings on one fact. +// +// A document that names no cluster at all, or one that grows a cluster which is +// not there, is answered by the checks that own those refusals; this one has +// nothing to say about a deployment whose shape is not yet known and says +// nothing. +func StripeChecks( + ctx context.Context, + reader client.Reader, + namespace string, + config *simplyblockv1alpha2.ClusterDeploymentConfig, +) ([]StripeCheck, error) { + scheme, subject, slots, total, known, err := stripeSubject(ctx, reader, namespace, config) + if err != nil || !known { + return nil, err + } + total.merge(plannedFootprint(config, slots)) + + if !scheme.IsSupported() { + // The minimum of a scheme that does not exist is not a fact about + // anything, so it is not reported beside it. + return []StripeCheck{{ + Reason: StripeUnsupported, + Message: fmt.Sprintf( + "%s is %s, which is not a scheme simplyblock supports (%s); the control "+ + "plane refuses the cluster it would be created with, and that refusal "+ + "lands after approval has made this document immutable", + subject, scheme, erasurecoding.SupportedNotation()), + }}, nil + } + + minimum := scheme.MinimumNodes() + if total.nodes() < minimum { + return []StripeCheck{{ + Reason: StripeBelowMinimumNodes, + Message: fmt.Sprintf("%s is %s, which needs at least %d %s%s, and this document %s %d", + subject, scheme, minimum, plural(minimum, "storage node", "storage nodes"), + placementPhrase(scheme), leavesPhrase(config), total.nodes()), + }}, nil + } + + if total.hosts() < minimum { + return []StripeCheck{{ + Reason: StripeBelowMinimumWorkers, + Message: fmt.Sprintf( + "%s is %s, which needs at least %d storage nodes, and the %d this "+ + "deployment has sit on %d %s; every node of a worker fails with the "+ + "worker, so the spare the scheme rebuilds onto is not a spare", + subject, scheme, minimum, total.nodes(), total.hosts(), + plural(total.hosts(), "worker", "workers")), + }}, nil + } + return nil, nil +} + +// placementPhrase accounts for the minimum, because a number a reviewer cannot +// account for is a number they cannot act on: it is the stripe's own width plus +// a spare for each failure the stripe survives. +func placementPhrase(scheme erasurecoding.Scheme) string { + if scheme.ParityChunks == 0 { + return "" + } + return fmt.Sprintf(" (%d to place a stripe across and %d %s to rebuild onto)", + scheme.DataChunks+scheme.ParityChunks, scheme.ParityChunks, + plural(scheme.ParityChunks, "spare", "spares")) +} + +// leavesPhrase says whether the count that follows is what the document builds +// or what it leaves behind, which are different things for a growth document. +func leavesPhrase(config *simplyblockv1alpha2.ClusterDeploymentConfig) string { + if config.Spec.ClusterRef != "" { + return "leaves the cluster with" + } + return "produces" +} + +// stripeSubject resolves what the document's erasure coding is, what to call it +// in a message, how many nodes each of its workers runs, and what the cluster +// already has. known is false for a document whose cluster cannot be resolved, +// which is a refusal of its own elsewhere. +func stripeSubject( + ctx context.Context, + reader client.Reader, + namespace string, + config *simplyblockv1alpha2.ClusterDeploymentConfig, +) (scheme erasurecoding.Scheme, subject string, slots int32, existing footprint, known bool, err error) { + if config.Spec.ClusterRef == "" { + template := config.Spec.Cluster + if template == nil { + return scheme, "", 0, existing, false, nil + } + return erasurecoding.SchemeOf(template.Stripe), "spec.cluster.stripe", + slotsOf(template.SocketsToUse, template.NodesPerSocket), newFootprint(), true, nil + } + + var cluster simplyblockv1alpha2.StorageCluster + key := client.ObjectKey{Namespace: namespace, Name: config.Spec.ClusterRef} + switch err := reader.Get(ctx, key, &cluster); { + case apierrors.IsNotFound(err): + return scheme, "", 0, existing, false, nil + case err != nil: + return scheme, "", 0, existing, false, + fmt.Errorf("reading StorageCluster %s: %w", config.Spec.ClusterRef, err) + } + + held, err := existingFootprint(ctx, reader, namespace, cluster.Name) + if err != nil { + return scheme, "", 0, existing, false, err + } + return erasurecoding.SchemeOf(cluster.Spec.Stripe), + fmt.Sprintf("the stripe of cluster %s", cluster.Name), + slotsPerWorker(&cluster), held, true, nil +} + +// plannedFootprint is every storage node the document would produce, one per +// slot per worker. +func plannedFootprint( + config *simplyblockv1alpha2.ClusterDeploymentConfig, slots int32, +) footprint { + planned := newFootprint() + for _, set := range config.Spec.NodeSets { + for _, group := range set.Groups { + for _, worker := range group.Workers { + for slot := int32(0); slot < slots; slot++ { + planned.add(worker, slot) + } + } + } + } + return planned +} + +// existingFootprint is every storage node a cluster already has. +func existingFootprint( + ctx context.Context, reader client.Reader, namespace, cluster string, +) (footprint, error) { + var nodes simplyblockv1alpha2.StorageNodeList + if err := reader.List(ctx, &nodes, client.InNamespace(namespace)); err != nil { + return footprint{}, fmt.Errorf("listing the nodes of cluster %s: %w", cluster, err) + } + held := newFootprint() + for i := range nodes.Items { + node := &nodes.Items[i] + if node.Spec.ClusterRef != cluster { + continue + } + slot := int32(0) + if node.Spec.Slot != nil { + slot = *node.Spec.Slot + } + held.add(node.Spec.WorkerNode, slot) + } + return held, nil +} diff --git a/operator/internal/controllers/deployment/erasurecoding_test.go b/operator/internal/controllers/deployment/erasurecoding_test.go new file mode 100644 index 000000000..4baec82d0 --- /dev/null +++ b/operator/internal/controllers/deployment/erasurecoding_test.go @@ -0,0 +1,218 @@ +// What a document's erasure coding is answered against: the scheme the control +// plane would accept, and the storage nodes the deployment actually produces. +// +// The cases are written against validate() rather than against the counting +// helpers, because the thing being guaranteed is that a reviewer is told before +// approval, and a helper that counts correctly while nothing calls it +// guarantees nothing. + +package deployment + +import ( + "context" + "strings" + "testing" + + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + + "github.com/simplyblock/atlas/ptr" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// aNode is a storage node the cluster already has, which is what a growth +// document adds to. +func aNode(name, worker string, slot int32) client.Object { + return &simplyblockv1alpha2.StorageNode{ + ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: theNamespace}, + Spec: simplyblockv1alpha2.StorageNodeSpec{ + ClusterRef: theCluster, + WorkerNode: worker, + Slot: ptr.To(slot), + }, + } +} + +// findingsOf validates a draft against the objects the cluster carries. +func findingsOf( + t *testing.T, + config *simplyblockv1alpha2.ClusterDeploymentConfig, + objects ...client.Object, +) []finding { + t.Helper() + r := reconcilerFor(t, append(objects, config)...) + findings, err := r.validate(context.Background(), config) + if err != nil { + t.Fatalf("validate: %v", err) + } + return findings +} + +// only asserts one finding under one reason and returns it, because a document +// with two problems reported as one is a document whose second problem is +// discovered after the first is fixed. +func only(t *testing.T, findings []finding, reason string) finding { + t.Helper() + if len(findings) != 1 || findings[0].reason != reason { + t.Fatalf("findings = %+v, want one %s", findings, reason) + } + return findings[0] +} + +// A document that says nothing about erasure coding describes a 1+1 cluster, +// because that is what the control plane defaults to, and 1+1 needs three +// storage nodes. Two workers cannot carry it, and saying so while the document +// is a draft is the only moment it can be corrected. +func TestADraftTooSmallForTheDefaultStripeIsReported(t *testing.T) { + config := aDocument(func(c *simplyblockv1alpha2.ClusterDeploymentConfig) { + c.Spec.Approved = false + c.Spec.Cluster.Stripe = nil + }) + + found := only(t, findingsOf(t, config, workers("worker-1", "worker-2")...), + StripeBelowMinimumNodes) + for _, want := range []string{"1+1", "3", "2"} { + if !strings.Contains(found.message, want) { + t.Errorf("the finding does not mention %q: %s", want, found.message) + } + } +} + +// A scheme the control plane's supported set does not hold is refused here, +// where the document can still be edited, rather than by the cluster create +// that runs after approval has made it immutable. +func TestADraftWhoseSchemeTheControlPlaneRefusesIsReported(t *testing.T) { + config := aDocument(func(c *simplyblockv1alpha2.ClusterDeploymentConfig) { + c.Spec.Approved = false + c.Spec.Cluster.Stripe = &simplyblockv1alpha2.StripeSpec{ + DataChunks: ptr.To(int32(3)), ParityChunks: ptr.To(int32(1)), + } + }) + + found := only(t, findingsOf(t, config, workers("worker-1", "worker-2")...), + StripeUnsupported) + for _, want := range []string{"3+1", "2+1"} { + if !strings.Contains(found.message, want) { + t.Errorf("the finding does not mention %q: %s", want, found.message) + } + } +} + +// The minimum of a scheme that does not exist is not a fact, so an unsupported +// scheme is reported once rather than followed by a node count derived from it. +func TestAnUnsupportedSchemeIsReportedWithoutItsNodeCount(t *testing.T) { + config := aDocument(func(c *simplyblockv1alpha2.ClusterDeploymentConfig) { + c.Spec.Approved = false + c.Spec.Cluster.Stripe = &simplyblockv1alpha2.StripeSpec{ + DataChunks: ptr.To(int32(4)), ParityChunks: ptr.To(int32(3)), + } + }) + + only(t, findingsOf(t, config, workers("worker-1", "worker-2")...), StripeUnsupported) +} + +// 1+0 is the one scheme a fleet of two carries, and a document stating it is +// accepted: the guarantee is that a deployment cannot fall below its stripe's +// minimum, not that every deployment is redundant. +func TestADraftStatingNoRedundancyIsAccepted(t *testing.T) { + config := aDocument(func(c *simplyblockv1alpha2.ClusterDeploymentConfig) { + c.Spec.Approved = false + }) + + if findings := findingsOf(t, config, workers("worker-1", "worker-2")...); len(findings) != 0 { + t.Errorf("a 1+0 document on two workers reported %+v", findings) + } +} + +// Two storage nodes on one worker die with the worker, so a fleet that reaches +// the node count only by running several nodes per socket has not reached the +// spare the scheme requires. +func TestNodesSharingAWorkerAreNotSpares(t *testing.T) { + config := aDocument(func(c *simplyblockv1alpha2.ClusterDeploymentConfig) { + c.Spec.Approved = false + c.Spec.Cluster.Stripe = &simplyblockv1alpha2.StripeSpec{ + DataChunks: ptr.To(int32(1)), ParityChunks: ptr.To(int32(1)), + } + c.Spec.Cluster.NodesPerSocket = ptr.To(int32(2)) + }) + + found := only(t, findingsOf(t, config, workers("worker-1", "worker-2")...), + StripeBelowMinimumWorkers) + if !strings.Contains(found.message, "worker") { + t.Errorf("the finding does not say what is short: %s", found.message) + } +} + +// A growth document is answered against the cluster it grows, so the nodes that +// are already there count toward the minimum. +func TestAGrowthDocumentCountsTheNodesTheClusterHas(t *testing.T) { + config := aDocument(func(c *simplyblockv1alpha2.ClusterDeploymentConfig) { + c.Spec.Approved = false + c.Spec.ClusterRef = theCluster + c.Spec.Cluster = nil + c.Spec.NodeSets[0].Groups[0].Workers = []string{"worker-3", "worker-4"} + }) + cluster := aCluster(func(c *simplyblockv1alpha2.StorageCluster) { + c.Spec.Stripe = &simplyblockv1alpha2.StripeSpec{ + DataChunks: ptr.To(int32(2)), ParityChunks: ptr.To(int32(1)), + } + }) + objects := append(workers("worker-1", "worker-2", "worker-3", "worker-4"), + cluster, aNode("node-1", "worker-1", 0), aNode("node-2", "worker-2", 0)) + + if findings := findingsOf(t, config, objects...); len(findings) != 0 { + t.Errorf("a growth document reaching the minimum reported %+v", findings) + } +} + +// The same document against a cluster that has one node is still short, and the +// message counts what the cluster would end up with rather than what the +// document adds. +func TestAGrowthDocumentThatStaysBelowTheMinimumIsReported(t *testing.T) { + config := aDocument(func(c *simplyblockv1alpha2.ClusterDeploymentConfig) { + c.Spec.Approved = false + c.Spec.ClusterRef = theCluster + c.Spec.Cluster = nil + c.Spec.NodeSets[0].Groups[0].Workers = []string{"worker-3"} + }) + cluster := aCluster(func(c *simplyblockv1alpha2.StorageCluster) { + c.Spec.Stripe = &simplyblockv1alpha2.StripeSpec{ + DataChunks: ptr.To(int32(2)), ParityChunks: ptr.To(int32(1)), + } + }) + objects := append(workers("worker-1", "worker-3"), + cluster, aNode("node-1", "worker-1", 0)) + + found := only(t, findingsOf(t, config, objects...), StripeBelowMinimumNodes) + for _, want := range []string{"2+1", "4", "2"} { + if !strings.Contains(found.message, want) { + t.Errorf("the finding does not mention %q: %s", want, found.message) + } + } +} + +// A growth document naming a worker the cluster already has a node on adds +// nothing there, because a node is identified by its worker and slot and the +// expansion creates only the slots that are empty. Counting it twice would +// report a minimum as met by a node that does not exist. +func TestAGrowthDocumentDoesNotCountASlotTwice(t *testing.T) { + config := aDocument(func(c *simplyblockv1alpha2.ClusterDeploymentConfig) { + c.Spec.Approved = false + c.Spec.ClusterRef = theCluster + c.Spec.Cluster = nil + c.Spec.NodeSets[0].Groups[0].Workers = []string{"worker-1", "worker-2"} + }) + cluster := aCluster(func(c *simplyblockv1alpha2.StorageCluster) { + c.Spec.Stripe = &simplyblockv1alpha2.StripeSpec{ + DataChunks: ptr.To(int32(1)), ParityChunks: ptr.To(int32(1)), + } + }) + objects := append(workers("worker-1", "worker-2"), + cluster, aNode("node-1", "worker-1", 0), aNode("node-2", "worker-2", 0)) + + found := only(t, findingsOf(t, config, objects...), StripeBelowMinimumNodes) + if !strings.Contains(found.message, "2") { + t.Errorf("the finding counts the re-named slots twice: %s", found.message) + } +} diff --git a/operator/internal/controllers/deployment/events.go b/operator/internal/controllers/deployment/events.go index 05413df0f..b4799c762 100644 --- a/operator/internal/controllers/deployment/events.go +++ b/operator/internal/controllers/deployment/events.go @@ -23,6 +23,14 @@ const ( // nodes to bind their management address to. NoManagementInterface = "NoManagementInterface" + // What the document says about erasure coding, which is the one part of a + // deployment nothing below the operator checks: the control plane validates + // the scheme on the cluster create and counts devices at activation, never + // nodes, so a fleet too small for its stripe is admitted everywhere else. + StripeUnsupported = "StripeUnsupported" + StripeBelowMinimumNodes = "StripeBelowMinimumNodes" + StripeBelowMinimumWorkers = "StripeBelowMinimumWorkers" + // AwaitingNodes is the expansion waiting for a node it created to come // online before it asks for the cluster to be activated. AwaitingNodes = "AwaitingNodes" diff --git a/operator/internal/controllers/deployment/expansion.go b/operator/internal/controllers/deployment/expansion.go index dcede39ac..695b9197a 100644 --- a/operator/internal/controllers/deployment/expansion.go +++ b/operator/internal/controllers/deployment/expansion.go @@ -448,14 +448,22 @@ func slotsPerWorker(cluster *simplyblockv1alpha2.StorageCluster) int32 { if workload == nil { return 1 } - sockets := int32(len(workload.SocketsToUse)) + return slotsOf(workload.SocketsToUse, workload.NodesPerSocket) +} + +// slotsOf is the same arithmetic against the fields themselves, because the +// document states the layout on its cluster template and the cluster states it +// on its node workload, and a deployment counted by one rule and built by +// another is a deployment whose validation means nothing. +func slotsOf(socketsToUse []string, nodesPerSocket *int32) int32 { + sockets := int32(len(socketsToUse)) if sockets == 0 { // An empty list means socket 0 alone (design-storagenode.md §5.1). sockets = 1 } perSocket := int32(1) - if workload.NodesPerSocket != nil && *workload.NodesPerSocket > 0 { - perSocket = *workload.NodesPerSocket + if nodesPerSocket != nil && *nodesPerSocket > 0 { + perSocket = *nodesPerSocket } return sockets * perSocket } diff --git a/operator/internal/controllers/deployment/validation.go b/operator/internal/controllers/deployment/validation.go index 026b1d23e..8c67cf843 100644 --- a/operator/internal/controllers/deployment/validation.go +++ b/operator/internal/controllers/deployment/validation.go @@ -91,6 +91,14 @@ func (r *ClusterDeploymentConfigReconciler) validate( }) } + stripe, err := StripeChecks(ctx, r.Client, config.Namespace, config) + if err != nil { + return nil, err + } + for _, check := range stripe { + findings = append(findings, finding{reason: check.Reason, message: check.Message}) + } + if found := conflictingInterfaces(config); found != "" { findings = append(findings, finding{reason: WorkerNotFound, message: found}) } diff --git a/operator/internal/webhook/clusterdeploymentconfig_validator.go b/operator/internal/webhook/clusterdeploymentconfig_validator.go index bfc92ba7a..06b936813 100644 --- a/operator/internal/webhook/clusterdeploymentconfig_validator.go +++ b/operator/internal/webhook/clusterdeploymentconfig_validator.go @@ -31,7 +31,7 @@ // that was edited and what to do instead where CEL names the rule that failed. // // Specified by operator/docs/designs/crd-redesign/design-clusterdeploymentconfig.md -// §5, whose four checks are checkApproval below. +// §5, whose five checks are checkApproval below. package webhook @@ -57,6 +57,7 @@ import ( // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=clusterdeploymentconfigs,verbs=get;list;watch // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storageclusters,verbs=get;list;watch +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodes,verbs=get;list;watch // +kubebuilder:rbac:groups="",resources=nodes,verbs=get;list;watch // ClusterDeploymentConfigValidator answers the approving edit against the @@ -134,7 +135,7 @@ func (v *ClusterDeploymentConfigValidator) Handle( // under as well as the sentence the reviewer reads. // // The reason is the vocabulary the controller's own validation events use, for -// the four checks both perform, so the two counters of §9.2 can be read against +// the five checks both perform, so the two counters of §9.2 can be read against // each other: a rejection under a reason the draft never reported is a gap in // the draft's validation rather than a reviewer's slip. type problem struct { @@ -165,7 +166,7 @@ func afterApproval( return admission.Allowed("") } -// checkApproval is §5.1's four checks, and returns everything wrong rather than +// checkApproval is §5.1's five checks, and returns everything wrong rather than // the first thing wrong. // // A reviewer fixing a document that is about to become immutable wants the whole @@ -190,6 +191,18 @@ func (v *ClusterDeploymentConfigValidator) checkApproval( }) } + // What the document says about erasure coding, against the deployment it + // describes. It is the one check here that nothing downstream repeats: the + // control plane validates the scheme on the cluster create, which lands after + // this edit, and its activation gate counts devices rather than nodes. + stripe, err := deployment.StripeChecks(ctx, v.Client, namespace, config) + if err != nil { + return nil, err + } + for _, check := range stripe { + problems = append(problems, problem{reason: check.Reason, message: check.Message}) + } + cluster, refused, err := v.resolveCluster(ctx, namespace, config) if err != nil { return nil, err diff --git a/operator/internal/webhook/clusterdeploymentconfig_validator_test.go b/operator/internal/webhook/clusterdeploymentconfig_validator_test.go index e6720de9f..fc8a1cb57 100644 --- a/operator/internal/webhook/clusterdeploymentconfig_validator_test.go +++ b/operator/internal/webhook/clusterdeploymentconfig_validator_test.go @@ -41,7 +41,9 @@ func testWorker(name string) *corev1.Node { // testDeploymentStorageCluster is the cluster every test in this file names, // because the checks under test are about whether a cluster of that name exists -// and what class it is, never about which name it carries. +// and what class it is, never about which name it carries. It carries the 1+0 a +// cluster of one node can carry, so that a case about something else is not also +// a case about erasure coding; a case about the stripe states its own. func testDeploymentStorageCluster( class simplyblockv1alpha2.StorageClusterDeviceClass, ) *simplyblockv1alpha2.StorageCluster { @@ -50,7 +52,12 @@ func testDeploymentStorageCluster( Name: testDeploymentCluster, Namespace: testDeploymentNamespace, }, - Spec: simplyblockv1alpha2.StorageClusterSpec{DeviceClass: class}, + Spec: simplyblockv1alpha2.StorageClusterSpec{ + DeviceClass: class, + Stripe: &simplyblockv1alpha2.StripeSpec{ + DataChunks: ptr.To(int32(1)), ParityChunks: ptr.To(int32(0)), + }, + }, } } @@ -60,6 +67,10 @@ const testSecondCluster = "rack-two" // testConfig is a document that passes every check: one group, one worker that // exists, NVMe devices, and a cluster of its own to create. +// +// It states 1+0 because one worker carries no other scheme: every redundant one +// needs at least three storage nodes, and a document saying nothing about +// erasure coding means the control plane's 1+1. func testConfig() *simplyblockv1alpha2.ClusterDeploymentConfig { return &simplyblockv1alpha2.ClusterDeploymentConfig{ ObjectMeta: metav1.ObjectMeta{ @@ -71,6 +82,9 @@ func testConfig() *simplyblockv1alpha2.ClusterDeploymentConfig { Name: testDeploymentCluster, MaxSubsystemCount: ptr.To(int32(10)), VCPUCount: ptr.To(int32(4)), + Stripe: &simplyblockv1alpha2.StripeSpec{ + DataChunks: ptr.To(int32(1)), ParityChunks: ptr.To(int32(0)), + }, }, NodeSets: []simplyblockv1alpha2.NodeSet{{ Name: "rack-b", @@ -116,6 +130,18 @@ func withBlockDevices( return config } +// testStorageNode is a node the cluster a growth document names already has. +func testStorageNode(name, worker string) *simplyblockv1alpha2.StorageNode { + return &simplyblockv1alpha2.StorageNode{ + ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: testDeploymentNamespace}, + Spec: simplyblockv1alpha2.StorageNodeSpec{ + ClusterRef: testDeploymentCluster, + WorkerNode: worker, + Slot: ptr.To(int32(0)), + }, + } +} + func configRaw(t *testing.T, config *simplyblockv1alpha2.ClusterDeploymentConfig) runtime.RawExtension { t.Helper() raw, err := json.Marshal(config) @@ -388,3 +414,46 @@ func TestTheRequestNamespaceIsUsedWhenTheObjectCarriesNone(t *testing.T) { mustDeny(t, review(t, nil, admissionv1.Create, nil, config), "worker-1") } + +// A document whose fleet is too small for its erasure coding is refused at the +// approving edit, which is the last moment it can be corrected: the control +// plane validates the scheme on the cluster create and counts devices at +// activation, never nodes, so nothing after this edit refuses it. +func TestApprovingADocumentTooSmallForItsStripeIsRefused(t *testing.T) { + config := testConfig() + config.Spec.Cluster.Stripe = &simplyblockv1alpha2.StripeSpec{ + DataChunks: ptr.To(int32(2)), ParityChunks: ptr.To(int32(1)), + } + + mustDeny(t, approve(t, []client.Object{testWorker("worker-1")}, config), + "2+1", "4") +} + +// A scheme the control plane's supported set does not hold is refused here too, +// because the create that refuses it runs after approval has made the document +// immutable. +func TestApprovingAnUnsupportedSchemeIsRefused(t *testing.T) { + config := testConfig() + config.Spec.Cluster.Stripe = &simplyblockv1alpha2.StripeSpec{ + DataChunks: ptr.To(int32(3)), ParityChunks: ptr.To(int32(1)), + } + + mustDeny(t, approve(t, []client.Object{testWorker("worker-1")}, config), "3+1") +} + +// A growth document is answered against the cluster it grows, so the nodes that +// cluster already has count toward the minimum and the approval stands. +func TestApprovingAGrowthDocumentCountsTheClustersNodes(t *testing.T) { + config := grows(testConfig(), testDeploymentCluster) + config.Spec.NodeSets[0].Groups[0].Workers = []string{"worker-3"} + + cluster := testDeploymentStorageCluster(simplyblockv1alpha2.StorageClusterDeviceClassNVMe) + cluster.Spec.Stripe = &simplyblockv1alpha2.StripeSpec{ + DataChunks: ptr.To(int32(1)), ParityChunks: ptr.To(int32(1)), + } + + mustAllow(t, approve(t, []client.Object{ + testWorker("worker-3"), cluster, + testStorageNode("node-1", "worker-1"), testStorageNode("node-2", "worker-2"), + }, config)) +} From aa997538add0243bd6aaf055f863e48f939b9aea Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Fri, 18 Sep 2026 22:10:48 +0200 Subject: [PATCH 083/206] feat(discovery): a draft proposes a scheme its fleet can carry The cluster block left stripe unset, so a reviewer approved a document that said nothing about erasure coding and got whatever the StorageCluster defaulted to -- a layout they never saw, on the one field that decides how many failures the deployment survives. The proposal is bounded by the fleet the run just counted, because a scheme needs ndcs+npcs nodes to place a stripe across and one spare per tolerated failure to rebuild onto. Under three nodes there is no redundancy to be had and the draft says 1+0 rather than proposing a scheme the fleet cannot carry; at three it is 1+1; above that 2+1, which is the smallest scheme that survives a failure and rebuilds without a second copy of everything. 1+0 on a small fleet is a real answer and the note says what it costs, because a two-node draft that silently proposed 2+1 would be a document whose activation then waits forever on a third node nobody is bringing. The note is the point of the whole file: nothing a probe reports decides a stripe, so the number is a starting point, and the sentence beside it is what tells a reviewer that and what to weigh when changing it. Every recorded expectation moves with it, since a drafted document now carries a cluster block it did not before. Co-Authored-By: Claude Opus 5 (1M context) --- .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../dev-01-four-nvme-disks/expected-notes.txt | 1 + .../dev/dev-01-four-nvme-disks/expected.yaml | 3 + .../expected-notes.txt | 1 + .../dev/dev-02-four-block-disks/expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../dev-09-ten-nvme-disks/expected-notes.txt | 1 + .../dev/dev-09-ten-nvme-disks/expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../filt-01-a-pci-deny-list/expected.yaml | 3 + .../expected-notes.txt | 1 + .../filt-02-a-pci-allow-list/expected.yaml | 3 + .../expected-notes.txt | 1 + .../filt-03-a-model-substring/expected.yaml | 3 + .../filt-04-a-size-range/expected-notes.txt | 1 + .../filt/filt-04-a-size-range/expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../filt-07-a-block-deny-list/expected.yaml | 3 + .../expected-notes.txt | 1 + .../filt-08-a-block-allow-list/expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../filt-14-no-filter-at-all/expected.yaml | 3 + .../fleet-01-one-worker/expected-notes.txt | 1 + .../fleet/fleet-01-one-worker/expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../net-07-two-equal-nics/expected-notes.txt | 1 + .../net/net-07-two-equal-nics/expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../net/net-37-a-team-interface/expected.yaml | 3 + .../expected-notes.txt | 1 + .../net-38-a-bridge-over-a-bond/expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../numa-01-one-memory-node/expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../numa-07-four-memory-nodes/expected.yaml | 3 + .../expected-notes.txt | 1 + .../numa-08-eight-memory-nodes/expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../pci-01-a-fleet-that-agrees/expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../pci-03-sixteen-and-sixteen/expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../role-06-the-master-spelling/expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../size-01-eight-logical-cpus/expected.yaml | 3 + .../expected-notes.txt | 1 + .../size-02-two-physical-cores/expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../size-14-swap-in-use/expected-notes.txt | 1 + .../size/size-14-swap-in-use/expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../size-21-no-hugetlbfs-at-all/expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + .../expected-notes.txt | 1 + .../expected.yaml | 3 + operator/internal/discovery/template.go | 96 +++++++++++++++ .../discovery/template_stripe_test.go | 109 ++++++++++++++++++ 320 files changed, 841 insertions(+) create mode 100644 operator/internal/discovery/template_stripe_test.go diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/expected-notes.txt index a749d795e..50cd37238 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-02, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-02, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 2, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-cm-01 in Draft: 2 workers with 8 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/expected.yaml index 902a7e015..93bfcd82f 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-01-a-report-from-the-previous-schema/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-cm-01-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/expected-notes.txt index 3d2c5161f..73e22013c 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-02, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-02, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 2, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-cm-02 in Draft: 2 workers with 8 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/expected.yaml index 1cd89073b..edafe3a5f 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-02-a-configmap-with-no-report-key/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-cm-02-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/expected-notes.txt index b67a69df2..fb448dd3b 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-02, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-02, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 2, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-cm-03 in Draft: 2 workers with 8 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/expected.yaml index c48adba96..823d6d0a0 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-03-a-report-that-does-not-parse/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-cm-03-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/expected-notes.txt index f9a1e897d..e394603fb 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-02, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-02, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 2, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-cm-04 in Draft: 2 workers with 8 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/expected.yaml index ea24c4e4a..b1a4440d3 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-04-a-report-naming-no-node/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-cm-04-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/expected-notes.txt index a5301ac33..e201c9d79 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-02, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-02, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 2, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-cm-05 in Draft: 2 workers with 8 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/expected.yaml index ff1e69f27..e4870b799 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-05-a-configmap-without-the-run-label/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-cm-05-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/expected-notes.txt index 27357dbde..544b2d350 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on ip-10-0-1-23.eu-central-1.compute.internal.a-very-long-suffix-nobody-shortened, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on ip-10-0-1-23.eu-central-1.compute.internal.a-very-long-suffix-nobody-shortened, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-cm-06 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/expected.yaml index a3f21f11e..496bfb0ac 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-06-a-worker-name-longer-than-a-label/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-cm-06-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/expected-notes.txt index 7f274d953..f097df39c 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-cm-07 in Draft: 1 workers with 128 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/expected.yaml index 6d097b9dc..679afb70b 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/cm/cm-07-a-report-of-a-hundred-and-twenty-eight-devices/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-cm-07-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/expected-notes.txt index 19b879293..f63de5bc3 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-dev-01 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/expected.yaml index c50e61d3e..ca97edf77 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-01-four-nvme-disks/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-dev-01-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/expected-notes.txt index 3c1d9d9ed..da017e4e1 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-dev-02 in Draft: 1 workers with 4 block devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/expected.yaml index a16ce0258..6d9e6d0b5 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-02-four-block-disks/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-dev-02-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/expected-notes.txt index 3fafbe5b6..edcc75cd4 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-dev-04 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 2 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/expected.yaml index 97fcbaa2e..0db9237e6 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-04-mixed-classes-on-an-nvme-run/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-dev-04-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/expected-notes.txt index f8cad067a..6dc74fedd 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-dev-05 in Draft: 1 workers with 2 block devices, in 1 group(s) across 1 node set(s); 2 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/expected.yaml index f1ef93ef5..05f3dc4ca 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-05-mixed-classes-on-a-block-run/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-dev-05-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/expected-notes.txt index c16005f59..6b8d584f2 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-dev-06 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/expected.yaml index 3798c6fbb..cb01cf74d 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-06-one-controller-two-namespaces/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-dev-06-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/expected-notes.txt index 542045b0e..26845742f 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-dev-07 in Draft: 1 workers with 1 block devices, in 1 group(s) across 1 node set(s); 1 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/expected.yaml index 1622e6e6d..0926f0c1f 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-07-a-sata-disk-beside-an-nvme-one/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-dev-07-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/expected-notes.txt index 3f32dc8cc..c62790ece 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-dev-08 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/expected.yaml index 90f1b8733..782b67c09 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-08-a-spinning-disk-beside-an-ssd/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-dev-08-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/expected-notes.txt index d90b28001..d7360ae18 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-dev-09 in Draft: 1 workers with 10 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/expected.yaml index 05b1993e2..707f4c466 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-09-ten-nvme-disks/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-dev-09-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/expected-notes.txt index 2c3cec7ca..4dec13acc 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-dev-10 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 3 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/expected.yaml index e06b18c79..f10c5e3b8 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-10-partitions-and-loopbacks-beside-disks/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-dev-10-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/expected-notes.txt index d5cf092f5..54e4a731f 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-dev-11 in Draft: 1 workers with 128 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/expected.yaml index 20232b9e0..b46b937a6 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-11-a-worker-at-the-selection-ceiling/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-dev-11-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/expected-notes.txt index 6f1896307..744d1e80b 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-dev-12 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/expected.yaml index cd4d32dfa..8ae99d9ab 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-12-a-disk-that-reports-no-size/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-dev-12-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/expected-notes.txt index ee54f0b0f..237b12654 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-dev-13 in Draft: 1 workers with 1 nvme devices, in 1 group(s) across 1 node set(s); 1 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/expected.yaml index aa815d4d8..ef4fbd577 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-13-an-attached-simplyblock-volume/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-dev-13-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/expected-notes.txt index ae524ac76..c98fff4a7 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-dev-14 in Draft: 1 workers with 1 block devices, in 1 group(s) across 1 node set(s); 1 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/expected.yaml index 270d13716..35dab5239 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-14-an-attached-volume-on-a-block-run/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-dev-14-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/expected-notes.txt index 3f2b3a7c1..7ee634341 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-dev-16 in Draft: 1 workers with 1 nvme devices, in 1 group(s) across 1 node set(s); 1 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/expected.yaml index ef64f524d..73b9b7bec 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-16-a-fabric-namespace-of-another-product/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-dev-16-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/expected-notes.txt index d9d261829..52846f49a 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-dev-17 in Draft: 1 workers with 1 block devices, in 1 group(s) across 1 node set(s); 1 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/expected.yaml index 45307ac1c..86baa215d 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-17-an-iscsi-lun-nobody-named/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-dev-17-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/expected-notes.txt index 3b6e63f4f..49346381e 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-dev-18 in Draft: 1 workers with 2 block devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/expected.yaml index 6c538117b..76e5fa980 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-18-an-iscsi-lun-the-allow-list-names/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-dev-18-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/expected-notes.txt index 7ad132066..6bcb33584 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-dev-19 in Draft: 1 workers with 1 nvme devices, in 1 group(s) across 1 node set(s); 1 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/expected.yaml index 7f61d3c99..41c3e2ced 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-19-an-iscsi-lun-on-an-nvme-run/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-dev-19-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/expected-notes.txt index 4bf46af0a..8ea243dd0 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+1, the only redundant scheme a fleet of 3 carries: it needs 3 storage nodes and the next scheme up needs more. A cluster's stripe cannot be changed afterward wrote discovered-discover-fail-05 in Draft: 3 workers with 12 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/expected.yaml index 946e71c31..69667e244 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-05-no-usable-interface-anywhere/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-fail-05-cluster + stripe: + dataChunks: 1 + parityChunks: 1 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/expected-notes.txt index f227a61d4..e86a738aa 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is the API's minimum of 4, and the smallest placement (worker-01) has only 2 cores, so this cluster asks for more than that worker has maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+1, the only redundant scheme a fleet of 3 carries: it needs 3 storage nodes and the next scheme up needs more. A cluster's stripe cannot be changed afterward wrote discovered-discover-fail-06 in Draft: 3 workers with 12 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/expected.yaml index ed5bf480c..a127a5b8f 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-06-every-worker-under-the-vcpu-floor/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-fail-06-cluster + stripe: + dataChunks: 1 + parityChunks: 1 vcpuCount: 4 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/expected-notes.txt index 3f41158c5..54e72215e 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 4, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+1, the only redundant scheme a fleet of 3 carries: it needs 3 storage nodes and the next scheme up needs more. A cluster's stripe cannot be changed afterward wrote discovered-discover-fail-07 in Draft: 3 workers with 12 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/expected.yaml index eb45e0ab1..ee77abba2 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-07-every-worker-at-four-gibibytes/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-fail-07-cluster + stripe: + dataChunks: 1 + parityChunks: 1 vcpuCount: 4 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/expected-notes.txt index 6af1eca4d..a76dbdc54 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is the API's minimum of 4, because no worker reported the cores of the memory node it was placed on maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-fail-08 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/expected.yaml index f0ec1216f..0289b3e01 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-08-an-unreadable-processor-tree/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-fail-08-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 4 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/expected-notes.txt index a9f48baf6..638576e76 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 2+1, which needs 4 storage nodes and this fleet of 31 has them: it is the widest stripe to propose without knowing the workload. This fleet could also carry 1+1, 4+1, 1+2, 2+2, 4+2, which trade capacity for a second tolerated failure or the other way about, and a cluster's stripe cannot be changed afterward wrote discovered-discover-fail-09 in Draft: 31 workers with 124 nvme devices, in 1 group(s) across 1 node set(s); 5 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/expected.yaml index 37d662263..e09e0e50e 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-09-one-worker-of-thirty-two-refused/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-fail-09-cluster + stripe: + dataChunks: 2 + parityChunks: 1 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/expected-notes.txt index 5692a767f..65d24b6a9 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-filt-01 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 3 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/expected.yaml index cf8d5e3ce..b0b59de0d 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-01-a-pci-deny-list/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-filt-01-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/expected-notes.txt index 543b17b5a..33c6aaf6f 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-filt-02 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 3 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/expected.yaml index 442342a93..6da5d21d0 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-02-a-pci-allow-list/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-filt-02-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/expected-notes.txt index 7a68f5848..ee0c43c13 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-filt-03 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 3 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/expected.yaml index 6dd8866e3..a53c01709 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-03-a-model-substring/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-filt-03-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/expected-notes.txt index 8a4f770b7..f3cd20d42 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-filt-04 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 3 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/expected.yaml index feb200d8a..3b9513ee0 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-04-a-size-range/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-filt-04-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/expected-notes.txt index 59d196d92..b2bef00df 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-filt-05 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 1 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/expected.yaml index fe050fc7f..2fb2965d2 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-05-a-bare-size-is-both-bounds/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-filt-05-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/expected-notes.txt index 2e1733ecc..3ca523b15 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-filt-07 in Draft: 1 workers with 5 block devices, in 1 group(s) across 1 node set(s); 1 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/expected.yaml index 18228c567..a5b32e585 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-07-a-block-deny-list/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-filt-07-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/expected-notes.txt index d2254588f..b6ac192f3 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-filt-08 in Draft: 1 workers with 2 block devices, in 1 group(s) across 1 node set(s); 4 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/expected.yaml index 14589cbf3..03449edca 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-08-a-block-allow-list/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-filt-08-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/expected-notes.txt index 7fdff9f56..509a9fccd 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-filt-09 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 1 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/expected.yaml index 341109a81..59d66a48f 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-09-a-partition-table-waived/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-filt-09-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/expected-notes.txt index 6aee0237c..d5bf6df0b 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-filt-11 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 3 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/expected.yaml index a5d9e245e..257f51f5f 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-11-an-allow-list-in-uppercase/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-filt-11-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/expected-notes.txt index 0354faecc..66918b406 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-filt-12 in Draft: 1 workers with 1 nvme devices, in 1 group(s) across 1 node set(s); 4 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/expected.yaml index f77d3be04..8cb548311 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-12-one-address-allowed-and-denied/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-filt-12-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/expected-notes.txt index 8cca2d5f0..a21f02d60 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-filt-13 in Draft: 1 workers with 3 block devices, in 1 group(s) across 1 node set(s); 3 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/expected.yaml index 1dbee7037..a64bf548f 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-13-a-size-range-on-a-block-run/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-filt-13-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/expected-notes.txt index a73059461..65e8b9f42 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-filt-14 in Draft: 1 workers with 3 nvme devices, in 1 group(s) across 1 node set(s); 2 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/expected.yaml index cd8917c22..eb76f3033 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-14-no-filter-at-all/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-filt-14-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/expected-notes.txt index b4c9266da..0a17b70f5 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-fleet-01 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/expected.yaml index ceff29a87..0dba8570c 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-01-one-worker/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-fleet-01-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/expected-notes.txt index 0d641fc59..d9865143b 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+1, the only redundant scheme a fleet of 3 carries: it needs 3 storage nodes and the next scheme up needs more. A cluster's stripe cannot be changed afterward wrote discovered-discover-fleet-02 in Draft: 3 workers with 12 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/expected.yaml index 55db0db40..4c4944931 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-02-three-uniform-workers/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-fleet-02-cluster + stripe: + dataChunks: 1 + parityChunks: 1 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/expected-notes.txt index 549147508..12ff1875f 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 2+1, which needs 4 storage nodes and this fleet of 32 has them: it is the widest stripe to propose without knowing the workload. This fleet could also carry 1+1, 4+1, 1+2, 2+2, 4+2, which trade capacity for a second tolerated failure or the other way about, and a cluster's stripe cannot be changed afterward wrote discovered-discover-fleet-03 in Draft: 32 workers with 128 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/expected.yaml index f4f6d3e97..05e5da841 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-03-thirty-two-uniform-workers/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-fleet-03-cluster + stripe: + dataChunks: 2 + parityChunks: 1 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/expected-notes.txt index 6572794ab..9b3d948c4 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 2+1, which needs 4 storage nodes and this fleet of 30 has them: it is the widest stripe to propose without knowing the workload. This fleet could also carry 1+1, 4+1, 1+2, 2+2, 4+2, which trade capacity for a second tolerated failure or the other way about, and a cluster's stripe cannot be changed afterward wrote discovered-discover-fleet-05 in Draft: 30 workers with 120 nvme devices, in 1 group(s) across 1 node set(s); 6 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/expected.yaml index 5ddf35823..a51c3043d 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-05-three-of-thirty-two-refused/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-fleet-05-cluster + stripe: + dataChunks: 2 + parityChunks: 1 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/expected-notes.txt index aef485fda..bdacca976 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 2+1, which needs 4 storage nodes and this fleet of 32 has them: it is the widest stripe to propose without knowing the workload. This fleet could also carry 1+1, 4+1, 1+2, 2+2, 4+2, which trade capacity for a second tolerated failure or the other way about, and a cluster's stripe cannot be changed afterward wrote discovered-discover-fleet-06 in Draft: 32 workers with 128 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/expected.yaml index 203d4afce..ca536d998 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-06-the-same-fleet-in-another-order/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-fleet-06-cluster + stripe: + dataChunks: 2 + parityChunks: 1 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/expected-notes.txt index 9ac93bb70..370298249 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 2, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-fleet-07 in Draft: 2 workers with 8 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/expected.yaml index 25be74039..a79dd8f9d 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-07-two-reports-for-one-worker/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-fleet-07-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/expected-notes.txt index 969946ccd..556b5ce76 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 2, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-fleet-08 in Draft: 2 workers with 8 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/expected.yaml index d645ea6ef..fc2ba2f19 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/fleet/fleet-08-a-report-this-run-is-not-about/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-fleet-08-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/expected-notes.txt index 0980eabcb..7c88986a1 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-held-03 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/expected.yaml index 39902da23..92bfdbc83 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-03-four-idle-userspace-controllers/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-held-03-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/expected-notes.txt index 3cccaae78..f352148d5 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-held-05 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/expected.yaml index 69be29fa0..b2ff314bf 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-05-idle-controllers-beside-kernel-disks/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-held-05-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/expected-notes.txt index d85aac69c..6d2d45a15 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-held-06 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/expected.yaml index 3d3a966e8..95aef2e1f 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-06-an-idle-controller-on-vfio-pci/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-held-06-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/expected-notes.txt index bb8c25be4..af8de9f0c 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-held-07 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/expected.yaml index 617767bbf..0cb8f8e57 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-07-kernel-controllers-already-presented/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-held-07-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/expected-notes.txt index 6bdb42422..c401001f3 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-held-09 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 16 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/expected.yaml index e37328793..563c69551 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-09-sixteen-loopbacks-beside-four-disks/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-held-09-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/expected-notes.txt index b2cda69b3..c914069b8 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-held-10 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/expected.yaml index 2f0575538..8a926a95e 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-10-half-the-controllers-in-use/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-held-10-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/expected-notes.txt index 376a82db0..8bafe1a17 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-01 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/expected.yaml index b6ba36b2d..27a3d3a46 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-01-one-addressed-physical-nic/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-01-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/expected-notes.txt index 9b19a4c5b..5e4241774 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-02 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/expected.yaml index b362cd5d3..4d40e7cd1 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-02-the-faster-of-two-nics/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-02-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/expected-notes.txt index 979d19332..2173cfe96 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-03 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/expected.yaml index 59a300a5f..e5f5c590c 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-03-the-node-address-beats-the-faster-link/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-03-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/expected-notes.txt index 727152069..6fd3459b2 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-04 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/expected.yaml index d47e16295..8a6922e36 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-04-every-interface-unusable/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-04-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/expected-notes.txt index 96e405a0f..64ee4a0d9 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-05 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/expected.yaml index fce2f50d2..3deda9ef2 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-05-no-interfaces-reported/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-05-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/expected-notes.txt index 015f3a906..98670aa3c 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 2, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-06 in Draft: 2 workers with 8 nvme devices, in 2 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/expected.yaml index ce8e871d3..e35944e16 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-06-two-workers-naming-their-nics-differently/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-06-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/expected-notes.txt index fe609fd1e..700480028 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-07 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/expected.yaml index 5f0092277..47c9369db 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-07-two-equal-nics/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-07-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/expected-notes.txt index c5835aa2b..8552c9f4d 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-08 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/expected.yaml index f8652bea3..92734b433 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-08-an-addressed-nic-reporting-no-speed/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-08-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/expected-notes.txt index e1ede782a..084aa2b37 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-09 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/expected.yaml index d867c0c2c..822d157f0 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-09-the-node-address-on-a-host-bridge/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-09-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/expected-notes.txt index 869d2e1fb..52382a3d1 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-10 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/expected.yaml index 092036692..59d6333f4 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-10-a-global-ipv6-address-only/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-10-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/expected-notes.txt index 11e2257cb..d62e0aa2f 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-11 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/expected.yaml index c3850c6d1..8585c4136 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-11-a-link-local-and-a-routable-address/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-11-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/expected-notes.txt index fedc2414b..c79e9ef01 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-12 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/expected.yaml index 34c3e4983..94833faff 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-12-loopback-and-nothing-else/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-12-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/expected-notes.txt index a4b5dfa06..fef55ffd5 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-13 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/expected.yaml index 577d2e9b4..d4730f12d 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-13-the-node-address-on-a-bond/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-13-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/expected-notes.txt index cba6728c9..0067d0271 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-14 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/expected.yaml index 61bdee7f2..5761a548e 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-14-the-node-address-on-a-vlan/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-14-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/expected-notes.txt index 49b1d8e28..56bd83a77 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-15 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/expected.yaml index 3ad553fec..f9c8a3415 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-15-an-overlay-beside-an-addressed-nic/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-15-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/expected-notes.txt index e30b7e9ea..0f428da46 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-16 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/expected.yaml index 3aaa0aaef..267d077ab 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-16-the-node-address-on-a-vlan-over-a-bond/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-16-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/expected-notes.txt index b1f94f0ff..4a828b99b 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-17 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/expected.yaml index 8bee3a73e..8075805b3 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-17-a-macvlan-and-an-ipvlan-over-one-nic/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-17-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/expected-notes.txt index b66e10811..4315370c1 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-18 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/expected.yaml index 8ed81fcbf..08bb0ab86 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-18-twelve-veths-beside-one-nic/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-18-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/expected-notes.txt index a6aea2be3..257a1b7fa 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-19 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/expected.yaml index fdc602045..9892637b4 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-19-the-clusters-own-plumbing-alone/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-19-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/expected-notes.txt index 7742515ef..ac519d606 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-20 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/expected.yaml index 74d52cfc4..a92f880a8 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-20-a-dormant-link-and-a-down-lower-layer/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-20-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/expected-notes.txt index 7b4133b82..83863defd 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-21 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/expected.yaml index 124eca09c..8a5edcb0f 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-21-a-link-whose-state-is-unknown/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-21-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/expected-notes.txt index f5be9fb9b..e2298b687 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-22 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/expected.yaml index bb48b4964..565e2380c 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-22-a-fleet-that-knows-which-nic-it-means/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-22-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/expected-notes.txt index 51f26af2f..645071399 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-23 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/expected.yaml index 015ada7e2..a31121bfa 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-23-a-fleet-that-knows-its-data-nics/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-23-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/expected-notes.txt index 13332c1fa..783f550f6 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-24 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/expected.yaml index 55ebd38ba..67b674d4d 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-24-a-fast-link-and-a-slow-one/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-24-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/expected-notes.txt index a304c26c2..18245a2dc 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-25 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/expected.yaml index 62ace41db..1870c948d 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-25-jumbo-frames-against-standard-ones/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-25-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/expected-notes.txt index 059400ffc..fa1a0f5c6 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-26 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/expected.yaml index aba65d5b7..c6f189a7b 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-26-a-fast-bridge-against-a-slower-nic/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-26-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/expected-notes.txt index cecf62d8d..de5c71a23 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-27 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/expected.yaml index e445991cf..c2173b0f9 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-27-a-link-up-with-no-partner/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-27-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/expected-notes.txt index 341fca074..aed2061f1 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-28 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/expected.yaml index 852d5b206..8214df6eb 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-28-an-aggregate-with-no-speed-of-its-own/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-28-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/expected-notes.txt index 8e4589e2c..8ab190c9e 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-29 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/expected.yaml index 49018dc9b..4dbde4be7 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-29-a-vlan-inheriting-the-bonds-speed/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-29-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/expected-notes.txt index f22772fcc..7cf0c901a 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-30 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/expected.yaml index 6ee8feb01..2c1d47aeb 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-30-a-bond-across-two-sockets/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-30-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/expected-notes.txt index ddf501103..2e26dc1fc 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-31 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/expected.yaml index 4260ac8ab..a69356eb8 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-31-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/expected-notes.txt index 299cd946e..a59a1c5a1 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-33 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/expected.yaml index ab38ab56d..216279d6a 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-33-the-node-address-on-a-down-link/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-33-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/expected-notes.txt index dc02c5a77..e6e9f77f1 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-34 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/expected.yaml index 92f31b367..12aec7c5a 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-34-the-node-address-on-two-interfaces/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-34-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/expected-notes.txt index 7ecb77b6b..e00822e73 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-35 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/expected.yaml index 71a3786f8..dd590925d 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-35-an-aggregate-reporting-its-own-speed/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-35-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/expected-notes.txt index 3c47feb21..359864473 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-36 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/expected.yaml index 282311f75..62a709f5a 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-36-a-member-the-report-does-not-carry/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-36-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/expected-notes.txt index 1dfb54136..e2a3022f8 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-37 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/expected.yaml index 632a8f388..14d69feed 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-37-a-team-interface/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-37-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/expected-notes.txt index 7fbcdc81a..4a0fc96cb 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-38 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/expected.yaml index c14638172..f7d9f9606 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-38-a-bridge-over-a-bond/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-38-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/expected-notes.txt index f471247f2..750390a89 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-39 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/expected.yaml index 7a2286487..734f3e801 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-39-a-macvlan-holding-the-node-address/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-39-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/expected-notes.txt index d55cd4b6b..fb5d81b69 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-40 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/expected.yaml index 9e6d0a805..99b7aa62a 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-40-an-interface-naming-no-kind/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-40-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/expected-notes.txt index ca74cac69..6caffac66 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-41 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/expected.yaml index fa215ca0c..55e4847ec 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-41-a-bonded-data-path-ranked-for-placement/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-41-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/expected-notes.txt index ba78ccf0c..dec055e14 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-net-42 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/expected.yaml index 3c59303fa..94a345417 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-42-what-the-draft-says-about-a-bonded-host/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-net-42-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/expected-notes.txt index 45d7ba75e..644e06e76 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-numa-01 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/expected.yaml index de45da834..cda63e163 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-01-one-memory-node/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-numa-01-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/expected-notes.txt index 79d7e6860..805f9e82d 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 8, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-numa-02 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 2 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/expected.yaml index 52e4e407d..ca53db612 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-02-two-nodes-evenly-split/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-numa-02-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 8 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/expected-notes.txt index edffb2773..442a2bcbb 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 8, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-numa-03 in Draft: 1 workers with 3 nvme devices, in 1 group(s) across 1 node set(s); 1 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/expected.yaml index 39132b0e5..bb34b2228 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-03-two-nodes-one-against-three/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-numa-03-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 8 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/expected-notes.txt index e6abb4fc5..da73ebd30 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 8, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-numa-04 in Draft: 1 workers with 3 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/expected.yaml index 749e830c3..9e1c0cb83 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-04-count-beats-capacity/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-numa-04-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 8 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/expected-notes.txt index 81a10fa1e..ae4a1c837 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 8, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-numa-05 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/expected.yaml index bcda41e81..d7ab1d1ae 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-05-capacity-breaks-the-count-tie/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-numa-05-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 8 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/expected-notes.txt index 4341f3e39..40a5f4d46 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 24, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-numa-06 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 2 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/expected.yaml index 57a6816ba..30fb3d81a 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-06-cores-break-the-capacity-tie/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-numa-06-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 24 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/expected-notes.txt index 40d07e240..b91053b9c 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-numa-07 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 6 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/expected.yaml index d1f1a8982..d48b8a354 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-07-four-memory-nodes/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-numa-07-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/expected-notes.txt index b5429a99a..6232f2083 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/expected-notes.txt @@ -1,4 +1,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. vcpuCount is 8, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-numa-08 in Draft: 1 workers with 3 nvme devices, in 1 group(s) across 1 node set(s); 7 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/expected.yaml index 1ed6466bd..0fa1e979c 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-08-eight-memory-nodes/expected.yaml @@ -12,6 +12,9 @@ spec: enableDriveFormat: true maxSubsystemCount: 30 name: discovered-discover-numa-08-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 8 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/expected-notes.txt index b1bc9fad7..9e440778e 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/expected-notes.txt @@ -1,4 +1,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. vcpuCount is the API's minimum of 4, because no worker reported the cores of the memory node it was placed on maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-numa-09 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/expected.yaml index 4e1881b14..4425fb39a 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-09-every-device-on-no-node/expected.yaml @@ -12,6 +12,9 @@ spec: enableDriveFormat: true maxSubsystemCount: 30 name: discovered-discover-numa-09-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 4 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/expected-notes.txt index cd01cfc5d..1937927a2 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 8, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-numa-10 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 2 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/expected.yaml index 89ebc347e..4acac0d42 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-10-a-real-node-against-the-unknown-bucket/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-numa-10-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 8 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/expected-notes.txt index 2bafa25f1..893bcae85 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 8, the cores of the memory node chosen on worker-02, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 2, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-numa-12 in Draft: 2 workers with 3 nvme devices, in 2 group(s) across 1 node set(s); 1 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/expected.yaml index 3545852eb..5b67465ba 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-12-one-node-and-two-node-workers-agree-on-addresses/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-numa-12-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 8 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/expected-notes.txt index 83949a031..5bef1e6aa 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 8, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-numa-13 in Draft: 1 workers with 10 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/expected.yaml index 934472d8f..f9335cd74 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-13-every-disk-on-the-second-node/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-numa-13-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 8 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/expected-notes.txt index b1758ea7e..168d791f1 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 49, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 256G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-numa-14 in Draft: 1 workers with 5 nvme devices, in 1 group(s) across 1 node set(s); 5 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/expected.yaml index 3e0a904a7..6d189cd6c 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-14-a-large-two-socket-worker/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 256G name: discovered-discover-numa-14-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 49 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/expected-notes.txt index 2ca82468b..064adb72a 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is the API's minimum of 4, because no worker reported the cores of the memory node it was placed on maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-numa-15 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 2 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/expected.yaml index 19709d839..24d49f206 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/numa/numa-15-no-memory-node-carries-cores/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-numa-15-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 4 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/expected-notes.txt index 975ea8ee5..2bbb0ffb1 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 2+1, which needs 4 storage nodes and this fleet of 32 has them: it is the widest stripe to propose without knowing the workload. This fleet could also carry 1+1, 4+1, 1+2, 2+2, 4+2, which trade capacity for a second tolerated failure or the other way about, and a cluster's stripe cannot be changed afterward wrote discovered-discover-pci-01 in Draft: 32 workers with 128 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/expected.yaml index 0ec964637..4b3d2a293 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-01-a-fleet-that-agrees/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-pci-01-cluster + stripe: + dataChunks: 2 + parityChunks: 1 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/expected-notes.txt index 36d92737d..1493e9121 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 2, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-pci-02 in Draft: 2 workers with 8 nvme devices, in 2 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/expected.yaml index 501e594db..50ca5b627 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-02-two-workers-that-do-not/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-pci-02-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/expected-notes.txt index bacfdf776..5bd09f2d3 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 2+1, which needs 4 storage nodes and this fleet of 32 has them: it is the widest stripe to propose without knowing the workload. This fleet could also carry 1+1, 4+1, 1+2, 2+2, 4+2, which trade capacity for a second tolerated failure or the other way about, and a cluster's stripe cannot be changed afterward wrote discovered-discover-pci-03 in Draft: 32 workers with 128 nvme devices, in 2 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/expected.yaml index 40b9e9a78..b731169e5 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-03-sixteen-and-sixteen/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-pci-03-cluster + stripe: + dataChunks: 2 + parityChunks: 1 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/expected-notes.txt index d732febc9..659bdaa8a 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 2+1, which needs 4 storage nodes and this fleet of 32 has them: it is the widest stripe to propose without knowing the workload. This fleet could also carry 1+1, 4+1, 1+2, 2+2, 4+2, which trade capacity for a second tolerated failure or the other way about, and a cluster's stripe cannot be changed afterward wrote discovered-discover-pci-04 in Draft: 32 workers with 64 nvme devices, in 32 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/expected.yaml index 69ab3f467..aa166e750 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-04-every-worker-distinct/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-pci-04-cluster + stripe: + dataChunks: 2 + parityChunks: 1 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/expected-notes.txt index ff726a85f..c17a7b466 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-pci-05 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/expected.yaml index 760378848..46f4b1469 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-05-addresses-reported-descending/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-pci-05-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/expected-notes.txt index a0dfefe3b..bf5aaf437 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-pci-06 in Draft: 1 workers with 10 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/expected.yaml index 6d8e4368c..e8cadd2ae 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-06-ten-disks-across-two-buses/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-pci-06-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/expected-notes.txt index 401902a64..a9c818f34 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-pci-07 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/expected.yaml index 0a2c0298c..4d9127ef0 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-07-a-five-digit-pci-domain/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-pci-07-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/expected-notes.txt index 9b87e4a35..a5e61476d 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 2, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-pci-08 in Draft: 2 workers with 4 nvme devices, in 2 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/expected.yaml index e0683793f..1b460a1af 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-08-uppercase-hex-against-lowercase/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-pci-08-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/expected-notes.txt index 93254abff..d43f5aaaa 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-1, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-1, which is the smallest of the fleet +stripe is 2+1, which needs 4 storage nodes and this fleet of 32 has them: it is the widest stripe to propose without knowing the workload. This fleet could also carry 1+1, 4+1, 1+2, 2+2, 4+2, which trade capacity for a second tolerated failure or the other way about, and a cluster's stripe cannot be changed afterward wrote discovered-discover-pci-09 in Draft: 32 workers with 128 nvme devices, in 2 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/expected.yaml index 7996e2a59..272d84e80 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-09-unpadded-worker-names/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-pci-09-cluster + stripe: + dataChunks: 2 + parityChunks: 1 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/expected-notes.txt index 6198c7d5b..14d276dd2 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 2+1, which needs 4 storage nodes and this fleet of 5 has them: it is the widest stripe to propose without knowing the workload. This fleet could also carry 1+1, 1+2, which trade capacity for a second tolerated failure or the other way about, and a cluster's stripe cannot be changed afterward wrote discovered-discover-pci-10 in Draft: 5 workers with 16 nvme devices, in 3 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/expected.yaml index 4cc80e6aa..74f74c108 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-10-three-agree-and-two-do-not/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-pci-10-cluster + stripe: + dataChunks: 2 + parityChunks: 1 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/expected-notes.txt index 392203dbb..d0233b467 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 2+1, which needs 4 storage nodes and this fleet of 32 has them: it is the widest stripe to propose without knowing the workload. This fleet could also carry 1+1, 4+1, 1+2, 2+2, 4+2, which trade capacity for a second tolerated failure or the other way about, and a cluster's stripe cannot be changed afterward wrote discovered-discover-pci-11 in Draft: 32 workers with 116 nvme devices, in 8 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/expected.yaml index 128d9d9f6..a42fd5daa 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-11-a-majority-layout-with-stragglers/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-pci-11-cluster + stripe: + dataChunks: 2 + parityChunks: 1 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/expected-notes.txt index 7ef2eea37..3f0d9476e 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 2+1, which needs 4 storage nodes and this fleet of 8 has them: it is the widest stripe to propose without knowing the workload. This fleet could also carry 1+1, 4+1, 1+2, 2+2, 4+2, which trade capacity for a second tolerated failure or the other way about, and a cluster's stripe cannot be changed afterward wrote discovered-discover-pci-12 in Draft: 8 workers with 32 nvme devices, in 2 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/expected.yaml index 06b888542..7e17e9e5a 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-12-one-slot-moved-on-one-worker/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-pci-12-cluster + stripe: + dataChunks: 2 + parityChunks: 1 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/expected-notes.txt index db7a0b5e9..15c149eef 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 2, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-pci-13 in Draft: 2 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/expected.yaml index d288fd32d..82fa9ab62 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-13-one-layout-two-capacities/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-pci-13-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/expected-notes.txt index 5704abd9d..a9d0eabdc 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+1, the only redundant scheme a fleet of 3 carries: it needs 3 storage nodes and the next scheme up needs more. A cluster's stripe cannot be changed afterward wrote discovered-discover-pci-14 in Draft: 3 workers with 11 nvme devices, in 2 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/expected.yaml index 4cc7b03de..b2c7cdf23 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/pci/pci-14-a-worker-holding-a-subset/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-pci-14-cluster + stripe: + dataChunks: 1 + parityChunks: 1 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/expected-notes.txt index a0ddbb388..e5d8eeb35 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+1, the only redundant scheme a fleet of 3 carries: it needs 3 storage nodes and the next scheme up needs more. A cluster's stripe cannot be changed afterward wrote discovered-discover-role-01 in Draft: 3 workers with 12 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/expected.yaml index 0f99b2543..b04318d16 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-01-workers-with-no-role-label/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-role-01-cluster + stripe: + dataChunks: 1 + parityChunks: 1 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/expected-notes.txt index eae4e4cb3..16d7208db 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 2+1, which needs 4 storage nodes and this fleet of 6 has them: it is the widest stripe to propose without knowing the workload. This fleet could also carry 1+1, 4+1, 1+2, 2+2, which trade capacity for a second tolerated failure or the other way about, and a cluster's stripe cannot be changed afterward wrote discovered-discover-role-02 in Draft: 6 workers with 24 nvme devices, in 2 group(s) across 2 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/expected.yaml index d2739d403..e0784276e 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-02-an-infrastructure-tier-beside-the-workers/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-role-02-cluster + stripe: + dataChunks: 2 + parityChunks: 1 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/expected-notes.txt index 7249e8bd5..8992d89b7 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+1, the only redundant scheme a fleet of 3 carries: it needs 3 storage nodes and the next scheme up needs more. A cluster's stripe cannot be changed afterward wrote discovered-discover-role-03 in Draft: 3 workers with 12 nvme devices, in 2 group(s) across 2 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/expected.yaml index 2ca88f1e1..e858a897d 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-03-a-control-plane-node-with-disks/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-role-03-cluster + stripe: + dataChunks: 1 + parityChunks: 1 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/expected-notes.txt index 4c58f215c..344c6b925 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-role-04 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/expected.yaml index a3d2e8b09..e5808d6f2 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-04-a-node-labeled-infra-and-worker/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-role-04-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/expected-notes.txt index ecaf23632..9596ff557 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-role-05 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/expected.yaml index 25cdcaf08..8135bdd07 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-05-a-node-labeled-control-plane-and-worker/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-role-05-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/expected-notes.txt index 24a1efa18..4578c603c 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-role-06 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/expected.yaml index 83e1b597a..aeebf5ad3 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-06-the-master-spelling/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-role-06-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/expected-notes.txt index b1053274b..d510d9758 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 2, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-role-08 in Draft: 2 workers with 8 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/expected.yaml index 348d54ef8..a65dca438 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-08-a-cordoned-node-and-a-tainted-one/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-role-08-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/expected-notes.txt index 48561a796..13625378a 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-role-09 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/expected.yaml index f1412054a..e6492cbbe 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-09-a-role-this-product-does-not-know/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-role-09-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/expected-notes.txt index 6e192dc9d..63b255645 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 2+1, which needs 4 storage nodes and this fleet of 32 has them: it is the widest stripe to propose without knowing the workload. This fleet could also carry 1+1, 4+1, 1+2, 2+2, 4+2, which trade capacity for a second tolerated failure or the other way about, and a cluster's stripe cannot be changed afterward wrote discovered-discover-role-10 in Draft: 32 workers with 128 nvme devices, in 3 group(s) across 3 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/expected.yaml index fb959f1c5..0d316910c 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/role/role-10-thirty-two-machines-across-three-tiers/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-role-10-cluster + stripe: + dataChunks: 2 + parityChunks: 1 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/expected-notes.txt index 1ca15afb2..7163399b1 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 4, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-size-01 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/expected.yaml index 8058d192d..accde56bf 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-01-eight-logical-cpus/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-size-01-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 4 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/expected-notes.txt index 34b160f05..386026d2e 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is the API's minimum of 4, and the smallest placement (worker-01) has only 2 cores, so this cluster asks for more than that worker has maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-size-02 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/expected.yaml index 9d8686a2b..43181bcbd 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-02-two-physical-cores/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-size-02-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 4 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/expected-notes.txt index f08a46bea..b3f3b3442 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 98, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 256G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-size-03 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/expected.yaml index c5aa3541f..42b99860a 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-03-one-socket-196-vcpus/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 256G name: discovered-discover-size-03-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 98 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/expected-notes.txt index 73842c79a..4c8e9678d 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 49, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 128G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-size-04 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 2 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/expected.yaml index a514f920d..9da1b774c 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-04-two-sockets-196-vcpus/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 128G name: discovered-discover-size-04-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 49 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/expected-notes.txt index 9c109ad65..c9f77a3c2 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 4, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+1, the only redundant scheme a fleet of 3 carries: it needs 3 storage nodes and the next scheme up needs more. A cluster's stripe cannot be changed afterward wrote discovered-discover-size-05 in Draft: 3 workers with 12 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/expected.yaml index 383e759a0..701d3156b 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-05-a-fleet-of-uneven-core-counts/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-size-05-cluster + stripe: + dataChunks: 1 + parityChunks: 1 vcpuCount: 4 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/expected-notes.txt index ac64fa13c..581c52415 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 4, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-size-06 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/expected.yaml index d7663cbd6..0b18624c3 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-06-four-gibibytes-of-memory/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-size-06-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 4 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/expected-notes.txt index ea67da423..d9d317c7f 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-size-07 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/expected.yaml index 7ee2ec660..be37dcf04 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-07-an-unreadable-memory-reading/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-size-07-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/expected-notes.txt index 8d77f9fd7..d064cbbf3 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 8, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 256G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-size-09 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 2 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/expected.yaml index 87571ad98..f5e18a5fa 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-09-pages-already-set-aside/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 256G name: discovered-discover-size-09-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 8 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/expected-notes.txt index 861adb587..a3361c2f6 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/expected-notes.txt @@ -1,4 +1,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. vcpuCount is 8, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +stripe is 1+1, the only redundant scheme a fleet of 3 carries: it needs 3 storage nodes and the next scheme up needs more. A cluster's stripe cannot be changed afterward wrote discovered-discover-size-10 in Draft: 3 workers with 6 nvme devices, in 1 group(s) across 1 node set(s); 6 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/expected.yaml index 48ffa0828..cf31d97a1 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-10-one-worker-with-nothing-set-aside/expected.yaml @@ -12,6 +12,9 @@ spec: enableDriveFormat: true maxSubsystemCount: 30 name: discovered-discover-size-10-cluster + stripe: + dataChunks: 1 + parityChunks: 1 vcpuCount: 8 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/expected-notes.txt index 6f685dc6f..0939f22e1 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/expected-notes.txt @@ -1,4 +1,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-size-11 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/expected.yaml index 0d3f70fc4..10dc986dc 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-11-a-reservation-under-a-gigabyte/expected.yaml @@ -12,6 +12,9 @@ spec: enableDriveFormat: true maxSubsystemCount: 30 name: discovered-discover-size-11-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/expected-notes.txt index 0e0ad126f..773127464 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/expected-notes.txt @@ -1,4 +1,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-size-12 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/expected.yaml index 8c90d651d..4d0e3829f 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-12-a-reservation-with-no-node-breakdown/expected.yaml @@ -12,6 +12,9 @@ spec: enableDriveFormat: true maxSubsystemCount: 30 name: discovered-discover-size-12-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/expected-notes.txt index 64d3c3110..2a4099ef4 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/expected-notes.txt @@ -1,4 +1,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. vcpuCount is 8, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-size-13 in Draft: 1 workers with 2 nvme devices, in 1 group(s) across 1 node set(s); 2 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/expected.yaml index 9105f88ec..09f2f5f9d 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-13-pages-on-the-node-not-chosen/expected.yaml @@ -12,6 +12,9 @@ spec: enableDriveFormat: true maxSubsystemCount: 30 name: discovered-discover-size-13-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 8 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/expected-notes.txt index 723807fda..f265ba09e 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-size-14 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/expected.yaml index 0753065d4..f57ce9fda 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-14-swap-in-use/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-size-14-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/expected-notes.txt index 815891ccf..68f726867 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-size-15 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/expected.yaml index 98b7d62f9..0b6711d16 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-15-allocatable-far-under-capacity/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-size-15-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/expected-notes.txt index 7db22e3f9..49031e00c 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 200G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-size-16 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/expected.yaml index 05efb2bad..aa55d009f 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-16-no-room-for-an-allocation-on-top/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 200G name: discovered-discover-size-16-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/expected-notes.txt index f86ee2302..9c3acd5d6 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/expected-notes.txt @@ -1,4 +1,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-size-17 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/expected.yaml index 258d29e00..64dc15813 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-17-room-and-nothing-set-aside/expected.yaml @@ -12,6 +12,9 @@ spec: enableDriveFormat: true maxSubsystemCount: 30 name: discovered-discover-size-17-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/expected-notes.txt index cafc79400..797791b99 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 256G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-size-18 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/expected.yaml index 1f742150a..c8421a68b 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-18-a-reuse-instruction-with-nowhere-to-put-it/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 256G name: discovered-discover-size-18-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/expected-notes.txt index 7248e4ea1..98963d380 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/expected-notes.txt @@ -1,4 +1,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-size-19 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/expected.yaml index 56b88f5eb..2e0bc0bcf 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-19-every-page-promised-to-a-mapping/expected.yaml @@ -12,6 +12,9 @@ spec: enableDriveFormat: true maxSubsystemCount: 30 name: discovered-discover-size-19-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/expected-notes.txt index 5df61aaf5..dae06fe27 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 34G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-size-20 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/expected.yaml index 9172f72e1..0fb28dec8 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-20-two-page-sizes-at-once/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 34G name: discovered-discover-size-20-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/expected-notes.txt index e97928389..1b95a381e 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/expected-notes.txt @@ -1,4 +1,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a storage node takes it: a drive that carries anything is not usable otherwise. This is the line to remove if any of them should be left alone. vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware +stripe is 1+0, which protects nothing: every redundant scheme needs at least three storage nodes and this fleet has 1, so a third worker is what makes 1+1 possible. A cluster's stripe cannot be changed afterward wrote discovered-discover-size-21 in Draft: 1 workers with 4 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/expected.yaml index 2df28e9cc..0041411be 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/size/size-21-no-hugetlbfs-at-all/expected.yaml @@ -12,6 +12,9 @@ spec: enableDriveFormat: true maxSubsystemCount: 30 name: discovered-discover-size-21-cluster + stripe: + dataChunks: 1 + parityChunks: 0 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/expected-notes.txt index 111703651..90f8618bf 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+1, the only redundant scheme a fleet of 3 carries: it needs 3 storage nodes and the next scheme up needs more. A cluster's stripe cannot be changed afterward wrote discovered-discover-tmpl-01 in Draft: 3 workers with 12 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/expected.yaml index e6897f3a2..f1a312e41 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-01-every-draft-formats-its-drives/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-tmpl-01-cluster + stripe: + dataChunks: 1 + parityChunks: 1 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/expected-notes.txt index f21beef43..c63962d39 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+1, the only redundant scheme a fleet of 3 carries: it needs 3 storage nodes and the next scheme up needs more. A cluster's stripe cannot be changed afterward wrote discovered-discover-tmpl-02 in Draft: 3 workers with 12 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/expected.yaml index 3a3c09262..499b5b103 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-02-the-subsystem-count-is-not-a-reading/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-tmpl-02-cluster + stripe: + dataChunks: 1 + parityChunks: 1 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/expected-notes.txt index ba1c098ba..c9e98ceb9 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+1, the only redundant scheme a fleet of 3 carries: it needs 3 storage nodes and the next scheme up needs more. A cluster's stripe cannot be changed afterward wrote discovered-discover-tmpl-03 in Draft: 3 workers with 12 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/expected.yaml index 2eb3831c7..7e7e8781d 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-03-a-generated-config-name/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-tmpl-03-cluster + stripe: + dataChunks: 1 + parityChunks: 1 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/expected-notes.txt index 4247c2f59..b170b0161 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+1, the only redundant scheme a fleet of 3 carries: it needs 3 storage nodes and the next scheme up needs more. A cluster's stripe cannot be changed afterward wrote discovered-the-quarterly-storage-expansion-for-the-frankfurt-racks in Draft: 3 workers with 12 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/expected.yaml index 06ce8f3b5..cf516b083 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-04-a-cluster-name-past-its-limit/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-the-quarterly-storage-expansion-for-the-frankfurt-racks-cluster + stripe: + dataChunks: 1 + parityChunks: 1 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/expected-notes.txt index 879f58581..d3d7ce922 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 8, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+1, the only redundant scheme a fleet of 3 carries: it needs 3 storage nodes and the next scheme up needs more. A cluster's stripe cannot be changed afterward wrote discovered-discover-tmpl-06 in Draft: 3 workers with 6 nvme devices, in 1 group(s) across 1 node set(s); 6 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/expected.yaml index f723190cf..2a1983e39 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-06-a-two-socket-fleet-and-no-socket-layout/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-tmpl-06-cluster + stripe: + dataChunks: 1 + parityChunks: 1 vcpuCount: 8 environment: K3s nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/expected-notes.txt index fbd13e1a0..d4271bb1a 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+1, the only redundant scheme a fleet of 3 carries: it needs 3 storage nodes and the next scheme up needs more. A cluster's stripe cannot be changed afterward wrote discovered-discover-tmpl-07 in Draft: 3 workers with 12 nvme devices, in 1 group(s) across 1 node set(s); 0 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/expected.yaml index 7c3f4a333..93ea75fce 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-07-the-environment-is-copied-through/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-tmpl-07-cluster + stripe: + dataChunks: 1 + parityChunks: 1 vcpuCount: 16 environment: OpenShift nodeSets: diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/expected-notes.txt b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/expected-notes.txt index 0ab6483c1..a12696cdb 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/expected-notes.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/expected-notes.txt @@ -2,4 +2,5 @@ enableDriveFormat is set, so every drive listed here is formatted before a stora vcpuCount is 16, the cores of the memory node chosen on worker-01, which is the smallest of the fleet: the control plane assumes this uniform across a cluster's nodes maxSubsystemCount is 30, which is the middle of the range this API accepts and not a reading: how many subsystems a node should serve follows from the workload rather than from the hardware minHugePagesSize is 16G, the huge pages free on the memory node chosen on worker-01, which is the smallest of the fleet +stripe is 1+1, the only redundant scheme a fleet of 3 carries: it needs 3 storage nodes and the next scheme up needs more. A cluster's stripe cannot be changed afterward wrote discovered-discover-tmpl-08 in Draft: 3 workers with 9 nvme devices, in 1 group(s) across 1 node set(s); 3 refusal(s) diff --git a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/expected.yaml b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/expected.yaml index d122d03d6..26dbbc110 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/expected.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/tmpl/tmpl-08-the-filter-is-resolved-and-not-carried/expected.yaml @@ -13,6 +13,9 @@ spec: maxSubsystemCount: 30 minHugePagesSize: 16G name: discovered-discover-tmpl-08-cluster + stripe: + dataChunks: 1 + parityChunks: 1 vcpuCount: 16 environment: K3s nodeSets: diff --git a/operator/internal/discovery/template.go b/operator/internal/discovery/template.go index dfa68fd02..57d7b8663 100644 --- a/operator/internal/discovery/template.go +++ b/operator/internal/discovery/template.go @@ -14,9 +14,12 @@ package discovery import ( "fmt" + "strings" "github.com/simplyblock/atlas/ptr" simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + + "github.com/simplyblock/simplyblock-operator/internal/erasurecoding" ) const ( @@ -84,9 +87,102 @@ func ClusterTemplateFor(name string, plan Plan) ClusterTemplate { out.Template.MinHugePagesSize = size out.Notes = append(out.Notes, note) } + + stripe, note := stripeFor(plan) + out.Template.Stripe = stripe + out.Notes = append(out.Notes, note) return out } +// stripeFor proposes the erasure-coding scheme for the fleet the run found. +// +// It is stated rather than left out, and that is the whole point of it. An +// unstated stripe is the control plane's 1+1, which needs three storage nodes, +// so a draft that says nothing about erasure coding proposes a scheme a fleet of +// one or two cannot carry — and says it in the one field a reviewer cannot +// correct after approval, since the cluster's stripe is immutable. +// +// The ladder is deliberately short. 2+1 is the widest stripe the product +// documentation recommends without knowing the workload, and it is proposed +// wherever the fleet meets its minimum; 1+1 is what a fleet of exactly three +// carries; 1+0 is what is left, and it protects nothing, which the note says in +// those words. Everything beyond that — the tolerance of two failures that 1+2, +// 2+2, and 4+2 buy, and the capacity 4+1 saves — is a trade the reviewer makes +// against a workload this run knows nothing about, so the note names the +// alternatives rather than picking one. +func stripeFor(plan Plan) (*simplyblockv1alpha2.StripeSpec, string) { + nodes := fleetSize(plan) + + scheme := erasurecoding.Scheme{DataChunks: 2, ParityChunks: 1} + switch { + case nodes < 3: + scheme = erasurecoding.Scheme{DataChunks: 1, ParityChunks: 0} + case nodes < scheme.MinimumNodes(): + scheme = erasurecoding.Scheme{DataChunks: 1, ParityChunks: 1} + } + + stripe := &simplyblockv1alpha2.StripeSpec{ + DataChunks: ptr.To(int32(scheme.DataChunks)), + ParityChunks: ptr.To(int32(scheme.ParityChunks)), + } + return stripe, stripeNote(scheme, nodes) +} + +// stripeNote accounts for the proposal, which for this field means saying what +// the fleet rules out as well as what it allows. +func stripeNote(scheme erasurecoding.Scheme, nodes int) string { + if scheme.ParityChunks == 0 { + return fmt.Sprintf( + "stripe is %s, which protects nothing: every redundant scheme needs at least "+ + "three storage nodes and this fleet has %d, so a third worker is what makes "+ + "1+1 possible. A cluster's stripe cannot be changed afterward", + scheme, nodes) + } + + alternatives := alternativesFor(scheme, nodes) + if len(alternatives) == 0 { + return fmt.Sprintf( + "stripe is %s, the only redundant scheme a fleet of %d carries: it needs %d "+ + "storage nodes and the next scheme up needs more. A cluster's stripe cannot "+ + "be changed afterward", + scheme, nodes, scheme.MinimumNodes()) + } + return fmt.Sprintf( + "stripe is %s, which needs %d storage nodes and this fleet of %d has them: it is "+ + "the widest stripe to propose without knowing the workload. This fleet could "+ + "also carry %s, which trade capacity for a second tolerated failure or the "+ + "other way about, and a cluster's stripe cannot be changed afterward", + scheme, scheme.MinimumNodes(), nodes, strings.Join(alternatives, ", ")) +} + +// alternativesFor is every other supported scheme the fleet is large enough for, +// in the documentation's order. +func alternativesFor(chosen erasurecoding.Scheme, nodes int) []string { + var out []string + for _, scheme := range erasurecoding.Supported() { + if scheme == chosen || scheme.ParityChunks == 0 || scheme.MinimumNodes() > nodes { + continue + } + out = append(out, scheme.String()) + } + return out +} + +// fleetSize is how many storage nodes the draft produces, which is one per +// worker: the template proposes neither socketsToUse nor nodesPerSocket, so +// every worker runs a single node. +func fleetSize(plan Plan) int { + workers := map[string]struct{}{} + for _, set := range plan.NodeSets { + for _, group := range set.Groups { + for _, worker := range group.Workers { + workers[worker] = struct{}{} + } + } + } + return len(workers) +} + // vcpuCountFor is the smallest chosen memory node's core count across the // fleet, floored at the API's minimum. func vcpuCountFor(plan Plan) (int32, string) { diff --git a/operator/internal/discovery/template_stripe_test.go b/operator/internal/discovery/template_stripe_test.go new file mode 100644 index 000000000..a42160f6c --- /dev/null +++ b/operator/internal/discovery/template_stripe_test.go @@ -0,0 +1,109 @@ +// What the draft proposes for erasure coding, against the fleet the run found. +// +// The proposal is the one number in the template that a reviewer cannot correct +// after approval, because the cluster's stripe is immutable, and it is also the +// one the document's validation refuses outright: a scheme needing more storage +// nodes than the fleet has is a draft nobody can approve. So the rule this file +// holds the generator to is that whatever it proposes, the fleet it was derived +// from can carry it. + +package discovery + +import ( + "strings" + "testing" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + + "github.com/simplyblock/simplyblock-operator/internal/erasurecoding" +) + +// aFleetOf is a plan whose node sets name count workers, one group, which is +// what a discovery run produces for a uniform fleet. +func aFleetOf(count int) Plan { + workers := make([]string, 0, count) + for i := 1; i <= count; i++ { + workers = append(workers, "worker-"+string(rune('0'+i))) + } + return Plan{NodeSets: []simplyblockv1alpha2.NodeSet{{ + Name: "discovered", + Groups: []simplyblockv1alpha2.NodeGroup{{Name: "group-1", Workers: workers}}, + }}} +} + +func proposedScheme(t *testing.T, workers int) erasurecoding.Scheme { + t.Helper() + template := ClusterTemplateFor("a-cluster", aFleetOf(workers)) + if template.Template.Stripe == nil { + t.Fatalf("the draft for %d workers proposes no stripe at all", workers) + } + return erasurecoding.SchemeOf(template.Template.Stripe) +} + +// The ladder the generator walks. A fleet that cannot carry a redundant scheme +// is proposed 1+0 rather than a scheme it fails validation against, and a fleet +// that can is proposed the widest stripe that leaves a spare per failure. +func TestTheDraftProposesASchemeTheFleetCanCarry(t *testing.T) { + for _, testCase := range []struct { + workers int + scheme string + }{ + {1, "1+0"}, + {2, "1+0"}, + {3, "1+1"}, + {4, "2+1"}, + {6, "2+1"}, + {9, "2+1"}, + } { + if got := proposedScheme(t, testCase.workers).String(); got != testCase.scheme { + t.Errorf("a fleet of %d is proposed %s, want %s", + testCase.workers, got, testCase.scheme) + } + } +} + +// The invariant behind the ladder, held against every fleet size rather than +// the tabulated ones: the proposal is a scheme the control plane accepts, and +// the fleet meets its minimum. A draft failing either is one the reviewer can +// only rewrite. +func TestTheProposedSchemeIsOneTheFleetAndTheControlPlaneBothAccept(t *testing.T) { + for workers := 0; workers <= 9; workers++ { + scheme := proposedScheme(t, workers) + if !scheme.IsSupported() { + t.Errorf("a fleet of %d is proposed %s, which the control plane refuses", + workers, scheme) + } + if workers > 0 && scheme.MinimumNodes() > workers { + t.Errorf("a fleet of %d is proposed %s, which needs %d storage nodes", + workers, scheme, scheme.MinimumNodes()) + } + } +} + +// A derived number a reviewer cannot account for is one they cannot correct +// with confidence, so the note says what was proposed and what the fleet it was +// derived from was. +func TestTheDraftAccountsForTheSchemeItProposed(t *testing.T) { + template := ClusterTemplateFor("a-cluster", aFleetOf(4)) + + joined := strings.Join(template.Notes, "\n") + for _, want := range []string{"2+1", "4"} { + if !strings.Contains(joined, want) { + t.Errorf("the notes do not mention %q:\n%s", want, joined) + } + } +} + +// A fleet of two carries nothing redundant, and the note says so rather than +// leaving a reviewer to discover that the draft they approved protects nothing. +func TestAFleetTooSmallForRedundancyIsSaidSo(t *testing.T) { + template := ClusterTemplateFor("a-cluster", aFleetOf(2)) + + joined := strings.ToLower(strings.Join(template.Notes, "\n")) + if !strings.Contains(joined, "1+0") { + t.Fatalf("the notes do not name the scheme proposed:\n%s", joined) + } + if !strings.Contains(joined, "1+1") || !strings.Contains(joined, "third") { + t.Errorf("the notes do not say what a redundant scheme would need:\n%s", joined) + } +} From 0b801a9b3656af2abdc7b0a57713c3071e6cb01c Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Fri, 18 Sep 2026 22:10:56 +0200 Subject: [PATCH 084/206] docs(crd-redesign): the designs carry the erasure-coding rules The scheme set, the node minimum each one implies, and the three places they are answered: the deployment config's validation and its approval webhook, and the cluster's activation gate. The test plans gain the rows for each, including the ones that record what nothing below this operator checks -- a scheme validated only at the cluster create, and an activation gate that counts devices and never nodes. Co-Authored-By: Claude Opus 5 (1M context) --- .../design-clusterdeploymentconfig.md | 80 +++++--- .../crd-redesign/design-storagecluster.md | 24 +++ .../test-plan-clusterdeploymentconfig.md | 185 +++++++++++------- .../docs/tests/test-plan-storagecluster.md | 13 ++ 4 files changed, 206 insertions(+), 96 deletions(-) diff --git a/operator/docs/designs/crd-redesign/design-clusterdeploymentconfig.md b/operator/docs/designs/crd-redesign/design-clusterdeploymentconfig.md index fb2a4bb59..6134de156 100644 --- a/operator/docs/designs/crd-redesign/design-clusterdeploymentconfig.md +++ b/operator/docs/designs/crd-redesign/design-clusterdeploymentconfig.md @@ -356,6 +356,27 @@ can, and `DeviceNotFound` in `status.message` is what says so while the document is still editable. Approving without reading it is how a device mistake becomes an immutable document, and no mechanism below the reviewer prevents that. +**The erasure coding is the other thing only this operator checks.** A scheme +decides how many storage nodes the deployment must have: `ndcs+npcs` of them to +place a stripe across, plus one spare per tolerated failure to rebuild onto, +which is 3 for 1+1, 4 for 2+1, 6 for 4+1, 5 for 1+2, 6 for 2+2, and 8 for 4+2. +The control plane validates the scheme itself on the cluster create and counts +devices at activation — `ndcs+npcs+1` of them — and never nodes, so a fleet of +four configured 4+2 is admitted everywhere below this document and produces a +cluster that loses data on the second failure it was configured to survive. +Validation counts the nodes the document produces, one per slot per worker, adds +the ones the cluster already has for a growth document, and reports +`StripeBelowMinimumNodes` when the total is short. A document that says nothing +about erasure coding is counted as 1+1, because that is what the control plane +defaults to. + +**Nodes that share a worker are counted, and then counted again.** The minimum is +a node count, and a fleet that reaches it by running several nodes per socket on +two workers has not bought the independent spare the count was asking for, since +every node of a worker fails with the worker. `StripeBelowMinimumWorkers` is that +case reported separately, so that the remedy — another machine rather than +another slot — is the one stated. + ### 4.2 The expansion machine ``` @@ -476,10 +497,20 @@ it found in `status.message` (§4.1) so that a reviewer fixes it in place. API.** Every worker named by every group has to exist as a Node, `spec.clusterRef` has to resolve to a `StorageCluster` when it is set and to nothing when it is not (§6), the class the groups name has to match that cluster's -`spec.deviceClass` where one is named, and no other approved config may already -own the cluster this one would create. All four are answerable from objects the -operator already caches, which is what makes them cheap enough to answer inside an -admission request. +`spec.deviceClass` where one is named, no other approved config may already own +the cluster this one would create, and the deployment has to have the storage +nodes its erasure-coding scheme requires (§4.1). All five are answerable from +objects the operator already caches, which is what makes them cheap enough to +answer inside an admission request. + +**The erasure-coding check is the one with nothing behind it.** The other four +are refused again later by something — a node create, the expansion's own +refusals — where this one is refused by nothing: the control plane accepts a +cluster whose fleet is too small for its stripe, activates it, and serves from +it. The schema refuses an unsupported scheme at the apply, which is CEL's half +(`StripeSpec` in +[`design-storagecluster.md`](design-storagecluster.md) §3.1), and the node count +is this webhook's, because a schema cannot count objects that do not exist yet. **The class check is the one of the four that has a schema half.** That every group agrees is CEL's (§3.1), and it holds from the first draft. What admission @@ -816,25 +847,28 @@ Both kinds are new, so both tables are new infrastructure. ### 9.1 Kubernetes events -| Event | Type | Reason | On | -|----------------------------------------------------------|-----------|--------------------------|---------------------------| -| A draft names a worker that does not exist | `Warning` | `WorkerNotFound` | `ClusterDeploymentConfig` | -| A draft names a device no node advertises | `Warning` | `DeviceNotFound` | `ClusterDeploymentConfig` | -| A draft's devices are not the class its cluster uses | `Warning` | `DeviceClassMismatch` | `ClusterDeploymentConfig` | -| A draft is valid and awaiting approval | `Normal` | `AwaitingApproval` | `ClusterDeploymentConfig` | -| Expansion is held because the control plane is not ready | `Warning` | `ControlPlaneNotReady` | `ClusterDeploymentConfig` | -| Expansion refused: the cluster already exists | `Warning` | `ClusterExists` | `ClusterDeploymentConfig` | -| Expansion refused: `clusterRef` names no cluster | `Warning` | `ClusterNotFound` | `ClusterDeploymentConfig` | -| Expansion created the cluster | `Normal` | `ClusterCreated` | `ClusterDeploymentConfig` | -| Expansion created the nodes | `Normal` | `NodesCreated` | `ClusterDeploymentConfig` | -| A step's deadline expired | `Warning` | `StepDeadlineExceeded` | `ClusterDeploymentConfig` | -| Discovery could not read a node's devices | `Warning` | `DeviceInspectionFailed` | `OperatorOps` | -| Discovery wrote a config | `Normal` | `ConfigWritten` | `OperatorOps` | -| The run is waiting for another to finish (§7) | `Normal` | `OperationQueued` | `OperatorOps` | -| The run started | `Normal` | `OperationStarted` | `OperatorOps` | -| The operation finished successfully | `Normal` | `OperationSucceeded` | `OperatorOps` | -| The operation failed | `Warning` | `OperationFailed` | `OperatorOps` | -| The run was aborted and its unwind finished | `Normal` | `OperationAborted` | `OperatorOps` | +| Event | Type | Reason | On | +|----------------------------------------------------------|-----------|-----------------------------|---------------------------| +| A draft names a worker that does not exist | `Warning` | `WorkerNotFound` | `ClusterDeploymentConfig` | +| A draft names a device no node advertises | `Warning` | `DeviceNotFound` | `ClusterDeploymentConfig` | +| A draft's devices are not the class its cluster uses | `Warning` | `DeviceClassMismatch` | `ClusterDeploymentConfig` | +| A draft's scheme is one the control plane refuses | `Warning` | `StripeUnsupported` | `ClusterDeploymentConfig` | +| A draft has fewer nodes than its scheme requires | `Warning` | `StripeBelowMinimumNodes` | `ClusterDeploymentConfig` | +| A draft's nodes sit on too few workers for its scheme | `Warning` | `StripeBelowMinimumWorkers` | `ClusterDeploymentConfig` | +| A draft is valid and awaiting approval | `Normal` | `AwaitingApproval` | `ClusterDeploymentConfig` | +| Expansion is held because the control plane is not ready | `Warning` | `ControlPlaneNotReady` | `ClusterDeploymentConfig` | +| Expansion refused: the cluster already exists | `Warning` | `ClusterExists` | `ClusterDeploymentConfig` | +| Expansion refused: `clusterRef` names no cluster | `Warning` | `ClusterNotFound` | `ClusterDeploymentConfig` | +| Expansion created the cluster | `Normal` | `ClusterCreated` | `ClusterDeploymentConfig` | +| Expansion created the nodes | `Normal` | `NodesCreated` | `ClusterDeploymentConfig` | +| A step's deadline expired | `Warning` | `StepDeadlineExceeded` | `ClusterDeploymentConfig` | +| Discovery could not read a node's devices | `Warning` | `DeviceInspectionFailed` | `OperatorOps` | +| Discovery wrote a config | `Normal` | `ConfigWritten` | `OperatorOps` | +| The run is waiting for another to finish (§7) | `Normal` | `OperationQueued` | `OperatorOps` | +| The run started | `Normal` | `OperationStarted` | `OperatorOps` | +| The operation finished successfully | `Normal` | `OperationSucceeded` | `OperatorOps` | +| The operation failed | `Warning` | `OperationFailed` | `OperatorOps` | +| The run was aborted and its unwind finished | `Normal` | `OperationAborted` | `OperatorOps` | **No event reports a rejected approval.** An admission rejection fails the request, so what the administrator gets is the webhook's message on their own diff --git a/operator/docs/designs/crd-redesign/design-storagecluster.md b/operator/docs/designs/crd-redesign/design-storagecluster.md index d395b9bfb..39eeeec25 100644 --- a/operator/docs/designs/crd-redesign/design-storagecluster.md +++ b/operator/docs/designs/crd-redesign/design-storagecluster.md @@ -145,6 +145,30 @@ affinity-based placement for storage components. `deviceClass` names the one cla of backend storage the cluster is built out of. All nine are enforced immutable, in two spellings that mean the same thing (§3.2). +**`stripe` is one of seven schemes, and the schema is what says so.** simplyblock +supports 1+0, 1+1, 2+1, 4+1, 1+2, 2+2, and 4+2, which is the set the control +plane holds in `SUPPORTED_ERASURE_CODING_SCHEMES` and refuses a cluster create +outside of. A type-level CEL rule on `StripeSpec` states the same set, so the +refusal arrives at the apply rather than from a cluster create that runs several +steps — and, for a cluster a deployment config produced, one irreversible +approval — later +([`design-clusterdeploymentconfig.md`](design-clusterdeploymentconfig.md) §5.1). +An unstated half is 1, which is what the control plane defaults each field to. + +**Each scheme also has a storage-node count below which it must not be used, and +the activation gate is where that is enforced.** The minimum is `ndcs+npcs` nodes +to place a stripe across plus one spare per tolerated failure to rebuild onto: 1 +for 1+0, 3 for 1+1, 4 for 2+1, 6 for 4+1, 5 for 1+2, 6 for 2+2, and 8 for 4+2. +Nothing below the operator enforces it — the control plane's own activation gate +counts devices, `ndcs+npcs+1` of them, and never nodes — so a cluster of four +nodes configured 4+2 activates today and loses data on the second failure it was +configured to survive. It cannot be a rule on this type, because the nodes are +objects of their own and a cluster is created before any of them exists, so +`Activate` holds on it instead and says why under `StripeNodesNotReady` (§6.3). +The hold is placed after the "already active" check, so a re-activation of a +cluster that is already serving is never held: the gate exists to stop a layout +being brought up wrong, and a live cluster's layout is already a fact. + **A cluster is built out of one class of device, and `deviceClass` is which.** NVMe devices are the class simplyblock has always accepted and logical block devices are the class 26.4 adds, and a cluster uses one of them. The two differ in diff --git a/operator/docs/tests/test-plan-clusterdeploymentconfig.md b/operator/docs/tests/test-plan-clusterdeploymentconfig.md index 18971cf3c..6068b1623 100644 --- a/operator/docs/tests/test-plan-clusterdeploymentconfig.md +++ b/operator/docs/tests/test-plan-clusterdeploymentconfig.md @@ -172,23 +172,30 @@ The handler is a pure function of an admission request and a lister, so every decision it makes is a unit test. Only its wiring needs `envtest`, which is `I-24` onward. -| # | Scenario | Type | Test | -|-------|-------------------------------------------------------------------------------------------------------|----------|------| -| U-86 | Creating a draft naming a worker that does not exist: admitted | Positive | — | -| U-87 | Creating a draft naming a device no node advertises: admitted | Positive | — | -| U-88 | Editing an unapproved draft into a worse state: admitted | Positive | — | -| U-89 | Approving a document whose workers all exist: admitted | Positive | — | -| U-90 | Approving a document naming a worker that does not exist: denied, naming the worker | Negative | — | -| U-91 | Approving a document naming an unavailable device: admitted, since devices are not an admission check | Boundary | — | -| U-92 | Approving with `clusterRef` set to a cluster that does not exist: denied | Negative | — | -| U-93 | Approving with no `clusterRef` while the named cluster exists: denied | Negative | — | -| U-94 | Approving while another approved config already owns the cluster: denied | Negative | — | -| U-95 | Editing an approved document: denied, naming the field and pointing at `clusterRef` | Negative | — | -| U-96 | Withdrawing approval: denied | Negative | — | -| U-97 | Editing only `metadata` on an approved document: admitted | Boundary | — | -| U-98 | A draft created by the operator's own service account: admitted on the same path | Boundary | — | -| U-99 | An approving edit by the operator's own service account: validated, not exempted | Negative | — | -| U-100 | A request whose object does not decode: `Errored`, not `Allowed` | Negative | — | +| # | Scenario | Type | Test | +|-------|---------------------------------------------------------------------------------------------------------|----------|-------------------------------------------------------| +| U-86 | Creating a draft naming a worker that does not exist: admitted | Positive | — | +| U-87 | Creating a draft naming a device no node advertises: admitted | Positive | — | +| U-88 | Editing an unapproved draft into a worse state: admitted | Positive | — | +| U-89 | Approving a document whose workers all exist: admitted | Positive | — | +| U-90 | Approving a document naming a worker that does not exist: denied, naming the worker | Negative | — | +| U-91 | Approving a document naming an unavailable device: admitted, since devices are not an admission check | Boundary | — | +| U-92 | Approving with `clusterRef` set to a cluster that does not exist: denied | Negative | — | +| U-93 | Approving with no `clusterRef` while the named cluster exists: denied | Negative | — | +| U-94 | Approving while another approved config already owns the cluster: denied | Negative | — | +| U-95 | Editing an approved document: denied, naming the field and pointing at `clusterRef` | Negative | — | +| U-96 | Withdrawing approval: denied | Negative | — | +| U-97 | Editing only `metadata` on an approved document: admitted | Boundary | — | +| U-98 | A draft created by the operator's own service account: admitted on the same path | Boundary | — | +| U-99 | An approving edit by the operator's own service account: validated, not exempted | Negative | — | +| U-100 | A request whose object does not decode: `Errored`, not `Allowed` | Negative | — | +| U-164 | Approving a document whose fleet is below its scheme's minimum: denied, naming the scheme and the count | Negative | `TestApprovingADocumentTooSmallForItsStripeIsRefused` | +| U-165 | Approving a document stating a scheme the control plane refuses: denied | Negative | `TestApprovingAnUnsupportedSchemeIsRefused` | +| U-166 | Approving a growth document: the cluster's existing nodes count toward the minimum | Positive | `TestApprovingAGrowthDocumentCountsTheClustersNodes` | + +`U-164` is the last refusal there is. The control plane accepts a cluster whose +fleet is too small for its stripe, activates it, and serves from it, so an +approval admitted here is a deployment nothing else refuses. `U-98` and `U-99` are the pair that keeps design §5.2's rule true: discovery writes drafts and needs no exemption, and an exemption for approvals would be a @@ -322,6 +329,37 @@ File: `operator/internal/controllers/deployment/clusterdeploymentconfig_expand_t `U-149` is the row that keeps the cap where it belongs. A per-node copy could only repeat the cluster's value, and design §3.1 leaves it out for that reason. +### Erasure Coding and the Node Minimum (design §4.1, §5.1) + +Files: `operator/internal/controllers/deployment/erasurecoding_test.go`, +`operator/internal/erasurecoding/scheme_test.go`, +`operator/internal/discovery/template_stripe_test.go` + +| # | Scenario | Type | Test | +|-------|----------------------------------------------------------------------------------------------|----------|------------------------------------------------------------------| +| U-151 | The supported set is the control plane's seven schemes and nothing else | Positive | `TestTheSupportedSchemesAreTheControlPlanes` | +| U-152 | The minimum node count is the documented table, scheme by scheme | Positive | `TestTheMinimumNodeCountIsTheDocumentedOne` | +| U-153 | An unstated stripe, and a half-stated one, are read as 1+1 | Boundary | `TestAnUnstatedStripeIsOnePlusOne` | +| U-154 | A two-worker draft that states no stripe: `StripeBelowMinimumNodes`, still `Draft` | Negative | `TestADraftTooSmallForTheDefaultStripeIsReported` | +| U-155 | A draft stating 3+1: `StripeUnsupported`, and no node count derived from it | Negative | `TestADraftWhoseSchemeTheControlPlaneRefusesIsReported` | +| U-156 | A two-worker draft stating 1+0: no finding, because the minimum is met | Positive | `TestADraftStatingNoRedundancyIsAccepted` | +| U-157 | 1+1 reached by two nodes per socket on two workers: `StripeBelowMinimumWorkers` | Negative | `TestNodesSharingAWorkerAreNotSpares` | +| U-158 | A growth document: the cluster's existing nodes count toward the minimum | Positive | `TestAGrowthDocumentCountsTheNodesTheClusterHas` | +| U-159 | A growth document still short: the message counts what the cluster ends up with | Negative | `TestAGrowthDocumentThatStaysBelowTheMinimumIsReported` | +| U-160 | A growth document re-naming a worker the cluster already has a node on: the slot counts once | Boundary | `TestAGrowthDocumentDoesNotCountASlotTwice` | +| U-161 | A discovery draft proposes a scheme its own fleet satisfies, at every fleet size | Positive | `TestTheProposedSchemeIsOneTheFleetAndTheControlPlaneBothAccept` | +| U-162 | A fleet of one or two is proposed 1+0, and the note says what a third worker would allow | Boundary | `TestAFleetTooSmallForRedundancyIsSaidSo` | +| U-163 | A fleet of four or more is proposed 2+1, with the alternatives named | Positive | `TestTheDraftAccountsForTheSchemeItProposed` | + +`U-157` is the row that makes the minimum mean what it says. The count is of +storage nodes, and a fleet that reaches it by running several nodes on one worker +has bought no independent spare, because every node of a worker fails with it. + +`U-161` is the invariant behind the generator rather than a case of it. A run +writes a draft and the draft's own validation refuses a scheme its fleet cannot +carry, so a generator that proposed one would produce documents nobody can +approve. + --- ## 2. Integration Tests @@ -329,62 +367,63 @@ repeat the cluster's value, and design §3.1 leaves it out for that reason. Full reconcile loop against a real Kubernetes API server via `envtest`. The immutability rules are CEL and cannot be exercised any other way. -| # | Scenario | Type | Test | -|----------|------------------------------------------------------------------------------------------------------------------------|----------|------| -| I-01 | An approved config edited: rejected by the immutability rule | Negative | — | -| I-02 | An unapproved config edited: accepted | Positive | — | -| I-03 | `spec.approved` set true, then false: the withdrawal is rejected | Negative | — | -| I-04 | `spec.approved` false, then true: accepted | Positive | — | -| I-05 | An approved config edited only in `metadata`: accepted, since the rule is on spec | Boundary | — | -| I-06 | `spec.nodeSets` omitted: rejected as `Required` | Negative | — | -| I-07 | `spec.nodeSets` empty: rejected by `MinItems` | Boundary | — | -| I-08 | `spec.environment` outside the enum: rejected | Negative | — | -| I-09 | A group with 201 workers: rejected by `MaxItems` | Boundary | — | -| I-10 | A group with duplicate workers: rejected by `listType=set` | Negative | — | -| ~~I-11~~ | `spec.nodeSets[].sizing` omitted: rejected as `Required`. Withdrawn: the field is gone, and `I-52` is what replaces it | — | — | -| I-12 | `OperatorOps.spec.action` outside the enum: rejected | Negative | — | -| I-13 | `OperatorOps.spec.action` changed after creation: rejected as immutable | Negative | — | -| I-14 | Short names `cdc` and `oops` resolve to the same lists as the full kinds | Positive | — | -| I-15 | A full expansion against a real API server: cluster and nodes exist afterward | Positive | — | -| I-16 | Deleting the config afterward: the cluster and nodes survive | Positive | — | -| I-17 | Two configs in two namespaces with the same name: neither reads the other | Negative | — | -| I-18 | Two configs in one namespace naming one cluster: the second is refused | Negative | — | -| I-19 | The controller's role covers every object the expansion creates | Positive | — | -| I-20 | A `devices` block with neither `nvme` nor `block`: rejected by the CEL rule | Negative | — | -| I-21 | A `devices` block with only `nvme`: accepted | Boundary | — | -| I-22 | A `devices` block with duplicate `nvme` entries: rejected by `listType=set` | Negative | — | -| I-23 | A `devices` block carrying a filter field such as `pcieDenyList`: rejected | Negative | — | -| I-34 | `devices.nvme` holding a well-formed PCI address: accepted | Positive | — | -| I-35 | `devices.nvme` holding a device path: rejected by the item pattern | Negative | — | -| I-36 | `devices.nvme` holding a truncated PCI address: rejected by the item pattern | Negative | — | -| I-37 | `devices.block` holding a path under `/dev`: accepted | Positive | — | -| I-38 | `devices.block` holding a bare device name: rejected by the item pattern | Negative | — | -| I-39 | `devices.block` holding a path outside `/dev`: rejected by the item pattern | Negative | — | -| I-24 | The webhook is registered for `create` and `update` on the kind | Positive | — | -| I-25 | An approving apply naming a missing worker: rejected by the API server | Negative | — | -| I-26 | The same document with the worker created first: accepted | Positive | — | -| I-27 | An invalid draft applied unapproved: accepted, and `status.message` reports it | Positive | — | -| I-28 | An edit to an approved document: rejected, and the stored object is unchanged | Negative | — | -| I-29 | `enableLogicalBlockDevices` outside a boolean: rejected by the schema | Negative | — | -| I-30 | A group whose `devices.block` names a path: accepted | Positive | — | -| I-31 | `enablePartitionedDevices` outside a boolean: rejected by the schema | Negative | — | -| I-32 | Approving a config naming a mounted device: accepted by the API server, then `Failed` at `Validating` | Negative | — | -| I-33 | Approving a config naming a partitioned device: accepted and expanded | Positive | — | -| I-40 | A group naming both `nvme` and `block`: rejected by the selection's CEL rule | Negative | — | -| I-41 | Two groups of one node set naming different classes: rejected by the spec's CEL rule | Negative | — | -| I-42 | Two node sets naming different classes: rejected by the same rule | Negative | — | -| I-43 | Every group of every node set naming `block`: accepted | Positive | — | -| I-44 | `enableLogicalBlockDevices` with a `pcieDenyList`: rejected by the filter's CEL rule | Negative | — | -| I-45 | `blockDenyList` with `enableLogicalBlockDevices` unset: rejected by the same rule | Negative | — | -| I-46 | `enableLogicalBlockDevices` with a `blockAllowList`: accepted | Positive | — | -| I-47 | The PCI filters with `enableLogicalBlockDevices` unset: accepted | Positive | — | -| I-48 | A group's `failureDomain` of `rack-b`: accepted | Positive | — | -| I-49 | A `failureDomain` holding a slash, and one of 64 characters: both rejected by the schema | Boundary | — | -| I-50 | `spec.cluster.maxSubsystemCount` omitted on a creating document: rejected as `Required` | Negative | — | -| I-51 | A node set's `sizing` carrying `maxSubsystemCount`: pruned rather than stored | Boundary | — | -| I-52 | A node set carrying a `sizing` block at all: pruned rather than stored | Boundary | — | -| I-53 | `spec.cluster.vcpuCount` omitted on a creating document: rejected as `Required` | Negative | — | -| I-54 | `spec.cluster.minHugePagesSize` omitted: accepted, and each node uses the computed minimum | Boundary | — | +| # | Scenario | Type | Test | +|----------|------------------------------------------------------------------------------------------------------------------------|----------|----------------------------------------------------| +| I-01 | An approved config edited: rejected by the immutability rule | Negative | — | +| I-02 | An unapproved config edited: accepted | Positive | — | +| I-03 | `spec.approved` set true, then false: the withdrawal is rejected | Negative | — | +| I-04 | `spec.approved` false, then true: accepted | Positive | — | +| I-05 | An approved config edited only in `metadata`: accepted, since the rule is on spec | Boundary | — | +| I-06 | `spec.nodeSets` omitted: rejected as `Required` | Negative | — | +| I-07 | `spec.nodeSets` empty: rejected by `MinItems` | Boundary | — | +| I-08 | `spec.environment` outside the enum: rejected | Negative | — | +| I-09 | A group with 201 workers: rejected by `MaxItems` | Boundary | — | +| I-10 | A group with duplicate workers: rejected by `listType=set` | Negative | — | +| ~~I-11~~ | `spec.nodeSets[].sizing` omitted: rejected as `Required`. Withdrawn: the field is gone, and `I-52` is what replaces it | — | — | +| I-12 | `OperatorOps.spec.action` outside the enum: rejected | Negative | — | +| I-13 | `OperatorOps.spec.action` changed after creation: rejected as immutable | Negative | — | +| I-14 | Short names `cdc` and `oops` resolve to the same lists as the full kinds | Positive | — | +| I-15 | A full expansion against a real API server: cluster and nodes exist afterward | Positive | — | +| I-16 | Deleting the config afterward: the cluster and nodes survive | Positive | — | +| I-17 | Two configs in two namespaces with the same name: neither reads the other | Negative | — | +| I-18 | Two configs in one namespace naming one cluster: the second is refused | Negative | — | +| I-19 | The controller's role covers every object the expansion creates | Positive | — | +| I-20 | A `devices` block with neither `nvme` nor `block`: rejected by the CEL rule | Negative | — | +| I-21 | A `devices` block with only `nvme`: accepted | Boundary | — | +| I-22 | A `devices` block with duplicate `nvme` entries: rejected by `listType=set` | Negative | — | +| I-23 | A `devices` block carrying a filter field such as `pcieDenyList`: rejected | Negative | — | +| I-34 | `devices.nvme` holding a well-formed PCI address: accepted | Positive | — | +| I-35 | `devices.nvme` holding a device path: rejected by the item pattern | Negative | — | +| I-36 | `devices.nvme` holding a truncated PCI address: rejected by the item pattern | Negative | — | +| I-37 | `devices.block` holding a path under `/dev`: accepted | Positive | — | +| I-38 | `devices.block` holding a bare device name: rejected by the item pattern | Negative | — | +| I-39 | `devices.block` holding a path outside `/dev`: rejected by the item pattern | Negative | — | +| I-24 | The webhook is registered for `create` and `update` on the kind | Positive | — | +| I-25 | An approving apply naming a missing worker: rejected by the API server | Negative | — | +| I-26 | The same document with the worker created first: accepted | Positive | — | +| I-27 | An invalid draft applied unapproved: accepted, and `status.message` reports it | Positive | — | +| I-28 | An edit to an approved document: rejected, and the stored object is unchanged | Negative | — | +| I-29 | `enableLogicalBlockDevices` outside a boolean: rejected by the schema | Negative | — | +| I-30 | A group whose `devices.block` names a path: accepted | Positive | — | +| I-31 | `enablePartitionedDevices` outside a boolean: rejected by the schema | Negative | — | +| I-32 | Approving a config naming a mounted device: accepted by the API server, then `Failed` at `Validating` | Negative | — | +| I-33 | Approving a config naming a partitioned device: accepted and expanded | Positive | — | +| I-40 | A group naming both `nvme` and `block`: rejected by the selection's CEL rule | Negative | — | +| I-41 | Two groups of one node set naming different classes: rejected by the spec's CEL rule | Negative | — | +| I-42 | Two node sets naming different classes: rejected by the same rule | Negative | — | +| I-43 | Every group of every node set naming `block`: accepted | Positive | — | +| I-44 | `enableLogicalBlockDevices` with a `pcieDenyList`: rejected by the filter's CEL rule | Negative | — | +| I-45 | `blockDenyList` with `enableLogicalBlockDevices` unset: rejected by the same rule | Negative | — | +| I-46 | `enableLogicalBlockDevices` with a `blockAllowList`: accepted | Positive | — | +| I-47 | The PCI filters with `enableLogicalBlockDevices` unset: accepted | Positive | — | +| I-48 | A group's `failureDomain` of `rack-b`: accepted | Positive | — | +| I-49 | A `failureDomain` holding a slash, and one of 64 characters: both rejected by the schema | Boundary | — | +| I-50 | `spec.cluster.maxSubsystemCount` omitted on a creating document: rejected as `Required` | Negative | — | +| I-51 | A node set's `sizing` carrying `maxSubsystemCount`: pruned rather than stored | Boundary | — | +| I-52 | A node set carrying a `sizing` block at all: pruned rather than stored | Boundary | — | +| I-53 | `spec.cluster.vcpuCount` omitted on a creating document: rejected as `Required` | Negative | — | +| I-54 | `spec.cluster.minHugePagesSize` omitted: accepted, and each node uses the computed minimum | Boundary | — | +| I-55 | A document whose cluster template states a scheme outside the supported seven: rejected by the schema | Negative | `TestTheDocumentsSchemaRefusesAnUnsupportedScheme` | --- diff --git a/operator/docs/tests/test-plan-storagecluster.md b/operator/docs/tests/test-plan-storagecluster.md index 972991164..ffd18f0e4 100644 --- a/operator/docs/tests/test-plan-storagecluster.md +++ b/operator/docs/tests/test-plan-storagecluster.md @@ -173,6 +173,17 @@ File: `operator/internal/controllers/cluster/storageclusterops_controller_test.g | U-77 | `CancelTask`: the cancel is issued once and the operation waits for the task to leave `status.tasks` | Positive | `TestCancelTaskWaitsForTheTaskToLeaveTheList` | | U-78 | `CancelTask` naming a task already gone: succeeds with no call | Boundary | `TestCancelingATaskThatIsAlreadyGoneSucceedsWithoutCalling` | | U-79 | `CancelTask` with no `spec.cancelTask.taskID`: terminal rather than requeued | Negative | `TestCancelTaskWithNoTaskIDFails` | +| U-95 | `Activate` on a cluster with fewer nodes than its scheme needs: held, `StripeNodesNotReady`, no call | Negative | `TestAnActivationBelowTheStripesMinimumIsHeld` | +| U-96 | The same cluster once the missing node exists: the activation is requested | Positive | `TestAnActivationWithTheNodesTheStripeNeedsIsRequested` | +| U-97 | `Activate` on a cluster already `active` and below the minimum: not held, because its layout is already a fact | Boundary | `TestAReactivationOfALiveClusterIsNotHeld` | +| U-98 | The node minimum itself, scheme by scheme, including the 1+0 that needs one node | Boundary | `TestTheActivationNodeCountRule` | + +`U-95` through `U-98` are in +`operator/internal/controllers/cluster/erasurecoding_test.go`, beside the rule +they exercise, and `U-95` is the only place this is refused. The control plane's +own activation gate counts devices, `ndcs+npcs+1` of them, and never nodes, so a +cluster whose fleet is too small for its stripe activates and serves from it +([`design-storagecluster.md`](../designs/crd-redesign/design-storagecluster.md) §3.1). ### Operation Reconciler: Rolling Restart (design §7) @@ -353,6 +364,8 @@ File: `operator/internal/controllers/cluster/cel_validation_test.go` | I-27 | `enableAtomic4kWrites` without `enableChecksumValidation`: rejected | Negative | `TestStorageClusterCELRejectsAtomic4kWritesWithoutChecksumValidation` | | I-28 | Either checksum field changed or cleared after creation: rejected as immutable | Negative | `TestStorageClusterChecksumValidationFieldsAreImmutable` | | I-29 | The default pool a cluster is created with picks up the CRD's declared defaults | Positive | `TestTheDefaultPoolIsFormattedXFS` | +| I-30 | Each of the seven supported schemes at creation: accepted; 3+1, 8+2, 2+0, 4+0, 1+3, and 16+4: rejected | Negative | `TestStorageClusterCELAcceptsOnlyTheSupportedErasureCodingSchemes` | +| I-31 | A stripe stating one half only: read as the control plane's default for the other, and accepted | Boundary | `TestStorageClusterCELReadsAnUnstatedHalfAsTheDefault` | `I-01` is answered by a unit test rather than an integration one: a not-found read needs no API server to be a not-found read, and the row is kept because the ID is From 20a8bf89eb88d19d954248a196b3c074a53f4d4a Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Fri, 18 Sep 2026 22:16:08 +0200 Subject: [PATCH 085/206] chore(repo): the loose files at the repository root leave it Fifteen manifests and notes from working against a live cluster, and two working documents of my own. A storage class, a fio job, a local-path provisioner, several cluster and migration manifests, two config dumps, and a hub analysis: none of them is repository content. They are what somebody had open while debugging, and carrying them means every checkout gets one person's scratch directory and every reader has to work out which of the twenty-two files at the root are the project. The files stay where they are. This removes them from the index and nothing else, so anybody still using them keeps using them. The two of mine are the same judgment applied to my own work: a rules catalogue that exists to be argued with and discarded once its questions are answered, and a case matrix that was put at the root because it was unfinished, which is a reason to keep it out rather than a reason to keep it there. Co-Authored-By: Claude Opus 5 (1M context) --- analysis-management-hub.md | 1472 ----------------------------- discovery-generator-rules.md | 757 --------------- discovery-generator-test-cases.md | 491 ---------- fio-quick.yaml | 104 -- job.md | 96 -- local-path-storage.yaml | 167 ---- sc.yaml | 27 - test-cluster-111.yaml | 97 -- test-cluster-okd.yaml | 93 -- test-cluster.yaml | 91 -- test-collection.yaml | 614 ------------ test-deployment-config.yaml | 32 - test-discovery-config.yaml | 47 - test-migration.yaml | 11 - test-new-crd.yaml | 15 - test-node-migration.yaml | 32 - test-pod.yaml | 37 - 17 files changed, 4183 deletions(-) delete mode 100644 analysis-management-hub.md delete mode 100644 discovery-generator-rules.md delete mode 100644 discovery-generator-test-cases.md delete mode 100644 fio-quick.yaml delete mode 100644 job.md delete mode 100644 local-path-storage.yaml delete mode 100644 sc.yaml delete mode 100644 test-cluster-111.yaml delete mode 100644 test-cluster-okd.yaml delete mode 100644 test-cluster.yaml delete mode 100644 test-collection.yaml delete mode 100644 test-deployment-config.yaml delete mode 100644 test-discovery-config.yaml delete mode 100644 test-migration.yaml delete mode 100644 test-new-crd.yaml delete mode 100644 test-node-migration.yaml delete mode 100644 test-pod.yaml diff --git a/analysis-management-hub.md b/analysis-management-hub.md deleted file mode 100644 index 318aa4148..000000000 --- a/analysis-management-hub.md +++ /dev/null @@ -1,1472 +0,0 @@ -# Fleet Management on Open Cluster Management - -**Status:** Design sketch -**Author:** Christoph Engelbert (noctarius) -**Date:** 2026-09-11 -**Design:** none yet. This is the input to `design-management-hub.md`. - -A single simplyblock control plane runs outside Kubernetes and manages several Kubernetes clusters, each of which runs -the operator. This document specifies that arrangement on Open Cluster Management: the kinds the hub owns, the kinds the -managed cluster keeps, what crosses between them, and what the hub serves to a management console. The link protocol, -the enrollment handshake, and the controllers belong to the design document that follows it. - ---- - -## Table of Contents - -1. [The Shape](#1-the-shape) -2. [Separate Kinds on Each Side](#2-separate-kinds-on-each-side) -3. [The Standalone Constraint](#3-the-standalone-constraint) -4. [What the Control Plane Already Knows](#4-what-the-control-plane-already-knows) -5. [The Two Bands](#5-the-two-bands) -6. [What Crosses, and What Does Not](#6-what-crosses-and-what-does-not) -7. [Band B: The Hub-Side Kinds](#7-band-b-the-hub-side-kinds) -8. [The Hub as a Gateway for `storage.simplyblock.io`](#8-the-hub-as-a-gateway-for-storagesimplyblockio) -9. [The Three Transports](#9-the-three-transports) -10. [The Add-Ons](#10-the-add-ons) -11. [Open Questions](#11-open-questions) - -Appendix: - -- [Appendix A: `fleet.simplyblock.io` types](#appendix-a-fleetsimplyblockio-types) - ---- - -## 1. The Shape - -There is exactly one control plane, and it is either local to a Kubernetes cluster or remote. A remote one fronts -several Kubernetes clusters, each running its own operator, which is the deployment -[`design-controlplane.md`](operator/docs/designs/crd-redesign/design-controlplane.md) §5.2 anticipates when it says an -external control plane may be shared and the operator must assume it is. - -A hub runs beside that control plane, and each managed cluster runs an add-on. The hub holds the desired shape of every -managed cluster and ships it down, and the managed cluster reconciles it exactly as a standalone deployment does. - -The hub holds no copy of a managed object. Desired state lives in hub-owned kinds distinct from the managed kinds, and -current state is served from the control plane the hub sits beside, joined with what the add-on reports. §2 states the -reasons. - ---- - -## 2. Separate Kinds on Each Side - -The hub and the managed cluster share no kind. A hub-side object carries intent and a managed-side object is what the -operator reconciles, and no object exists on both sides. - -The alternative is a digital twin, where the hub owns each object's spec, the add-on replays it into a local copy no -user may write, the operator writes the local status, and the add-on mirrors that back. Six properties of this API rule -it out. - -- **`metadata.generation` is issued by the local API server**, so a mirrored status carries a generation the hub never - saw. [`design-crd-model.md`](operator/docs/designs/crd-redesign/design-crd-model.md) §7.9 makes `observedGeneration` - mandatory because it is the only field that says a status is current, and across a twin it reports a - definite-looking wrong answer. One field also cannot say whether the generation it reports is the hub's or the - member's, so each side reads the other's write as an unobserved change and answers it, without end. -- **Defaulting and pruning mean the spec written is never the spec read back**, so an agent that writes, reads, diffs, - and resynchronizes flaps forever. -- **UIDs do not travel.** `spec.creatorRef` carries one and - [`design-persistentvolumeops.md`](operator/docs/designs/crd-redesign/design-persistentvolumeops.md) §11.1 calls it the - load-bearing part, and owner references carry them too. The ownership spine cannot be reconstructed on the far side. -- **A refused write has nowhere to be reported.** Three designs put `failurePolicy: Fail` webhooks in front of creates, - and a refused create writes no object, so there is no status to mirror. -- **`Ops` kinds are garbage-collected**, so an agent that reads local absence as work to do re-creates them, and a - re-created `StorageNodeOps` with the `Remove` action drains a node a second time. -- **Cluster-scoped kinds have no hub-side layout**, and object names collide across members. - -Each of these is a property of the twin, and each stops applying once the two sides share no kind. - -[RamenDR](https://github.com/RamenDR/ramen), the disaster-recovery orchestrator behind OpenShift Data Foundation, is -built the same way. Its hub holds `DRPolicy`, `DRCluster`, and `DRPlacementControl`, and its managed clusters hold -`VolumeReplicationGroup`, `DRClusterConfig`, and `MaintenanceMode`. The two sets share no kind. A -`VolumeReplicationGroup` is composed by a hub controller and shipped as an opaque payload, and no hub-side copy of it -exists. One binary runs in one of two modes, `DRHubType` or `DRClusterType`, chosen in `cmd/main.go`, and each mode -registers a different set of controllers. - ---- - -## 3. The Standalone Constraint - -A standalone deployment against a local control plane keeps working exactly as it does today. The constraint is stated -as a test, so that a fleet concern cannot reach the product: - -> A standalone installation installs no hub kind, runs no add-on, and its `CustomResourceDefinition` set, its RBAC, and -> its webhook configurations are byte-identical to what it installs today. - -Two consequences follow. Hub kinds live in their own API group, installed only on the hub. And the agent is an add-on -rather than a mode of the operator, so the operator binary and its RBAC are unchanged. RamenDR selects its mode in one -binary, and this arrangement does not. - -The operator therefore behaves identically in both deployments. It reconciles objects somebody wrote, and it draws no -distinction between a person and a work agent as the writer. - ---- - -## 4. What the Control Plane Already Knows - -The one control plane is the fleet's state store. Every design in this group reads its backend state from that control -plane's stream rather than from Kubernetes -([`design-crd-model.md`](operator/docs/designs/crd-redesign/design-crd-model.md) §7.7), and a hub sitting beside it -reads the same stream. - -| Fact | Where the hub reads it | -|--------------------------------------------------------|--------------------------------------------------------------------------------------------------------------------------------------------------------------| -| Cluster, node, pool, and volume state | The control plane, directly | -| Device inventory and health | The control plane. `StorageDevice` is a pure projection of it ([`design-storagedevice.md`](operator/docs/designs/crd-redesign/design-storagedevice.md) §5.1) | -| Capacity and occupancy measurements | The control plane's own metrics | -| Which worker a node runs on, and whether its pod is up | The add-on. Kubernetes knows this and the control plane does not | -| What a worker has before a cluster exists on it | The add-on, as the discovery inventory | -| Driver version and rollout state | The add-on | -| Claim-to-volume binding | The add-on, which is why a measurement kind is named after the claim | - -The add-on therefore carries the Kubernetes-shaped half of the picture, which is small and specific, and it carries -intent downward. Most of what a console displays about storage is on the hub's side of the wire before the add-on -reports anything. - -A disconnect degrades both halves. The control-plane half stays current only while the member's storage nodes still -reach the control plane, and an add-on that cannot reach the hub and nodes that cannot reach the control plane are -usually the same network event. The hub can always report when each half was last current, which is what separates a -partial answer from a wrong one. - ---- - -## 5. The Two Bands - -| Band | Group | Installed on | Contents | -|------|-------------------------------------|--------------------------------------|-------------------------------------------------------------------------------| -| A | `storage.simplyblock.io`, unchanged | Every standalone and managed cluster | Every kind in the target model, exactly as designed | -| B | `fleet.simplyblock.io`, new | The hub only | Intent kinds naming a member and composing a Band A object | -| — | Open Cluster Management | Hub and managed | `ManagedCluster`, `ManifestWork`, `ManagedClusterAddOn`, `ManagedClusterView` | - -Band B stacks on Band A. Nothing is removed from Band A and nothing in it references Band B, so §3's test passes by -construction. In a managed cluster the Band A objects are the same objects, applied by the work agent instead of by a -person. - -References run one way. A hub kind names a member, and no managed kind names a hub, which is what lets a managed -cluster lose its hub and keep reconciling. - ---- - -## 6. What Crosses, and What Does Not - -| Kind | Standalone | Managed cluster | Crosses as | -|---------------------------|-------------------------|--------------------------------------|-----------------------------------------------------------------------------| -| `ControlPlane` | `source.managed` | `source.external`, same kind | Nothing | -| `ControlPlaneOps` | Applies | No action applies | Nothing | -| `SimplyblockDriver` | Written by a person | `ManifestWork` payload | Down, from `DriverDeployment` | -| `ClusterDeploymentConfig` | Written or discovered | `ManifestWork` payload | Down, from `ClusterDeployment` | -| `OperatorOps` discovery | Written by a person | Runs locally, triggered by a payload | Down as a trigger via `FleetOperation`, up as inventory | -| `StorageCluster` | Written by a person | Expanded from the config, as today | Nothing. State comes from the control plane | -| `StorageNode` | Expanded by its cluster | Expanded by its cluster | Nothing. The hub reads the control plane | -| `StorageDevice` | Projected | Projected | Nothing. The hub reads the control plane | -| `StoragePool` | Written by a person | `ManifestWork` payload | Down, from a `StoragePool` payload. The backend half is the control plane's | -| `StorageClass` | Authored by a person | `ManifestWork` payload | Down, from `StorageClassDeployment` | -| Every `Ops` kind | Written by a person | Payload, or raised locally | Down via `FleetOperation` | -| `PersistentVolumeOps` | Written or fanned out | Almost always fanned out locally | Nothing. Cluster-scoped, high cardinality, local origin | -| `XyzMetrics` | Served | Served | Nothing. Both read the control plane | - -Six kinds never cross in either direction. - -A local object the control plane does not know is read straight from the member. A `StorageClass` is one, so a console -listing the classes that draw on a pool asks for them through a `ManagedClusterView` rather than finding them in the -projection. The same holds for any local object a Band A kind produced rather than mirrored. - ---- - -## 7. Band B: The Hub-Side Kinds - -`fleet.simplyblock.io/v1alpha1`, installed on the hub and nowhere else. Every convention below is -[`design-crd-model.md`](operator/docs/designs/crd-redesign/design-crd-model.md) §3 and §7 applied unchanged, so that the -two groups are one API to learn. The prefix for every label and annotation these kinds define is -`fleet.simplyblock.io`, one prefix per group, and metrics keep the `simplyblock___` form. - -The types are Appendix A, whole and as they are to be written. Each section below quotes the field its argument turns on -and no more, which is the shape -[`design-storagenode.md`](operator/docs/designs/crd-redesign/design-storagenode.md) §3 follows. - -| Kind | Category | Short | Holds | -|--------------------------|----------|-------|-------------------------------------------------------------------------| -| `StorageFleet` | Entity | `sf` | The one control plane, and the fleet's defaults. Singleton | -| `FleetMember` | Entity | `fm` | One enrolled Kubernetes cluster, its link, its inventory, and a roll-up | -| `ClusterDeployment` | Entity | `cd` | The composed `ClusterDeploymentConfig` for one member, and its approval | -| `DriverDeployment` | Entity | `dd` | The `SimplyblockDriver` payload for one member | -| `StorageClassDeployment` | Entity | `scd` | One `StorageClass` drawing on a pool in one member | -| `FleetOperation` | Action | `fop` | One Band A `Ops` object, shipped into one member | - -Five of the six are entities, because building out a cluster is desired state: a document, a driver version, the classes -that consume the capacity, and a member to put them in. Issuing an operation is the one imperative act, and -`FleetOperation` carries it. - -Ownership is a two-level tree. `FleetMember` owns every object that names it, so detaching a member collects its -deployments, its classes, and its operations. `StorageFleet` owns nothing, because a member outlives an edit to the -fleet's defaults and a singleton at the root of the tree makes one deletion a fleet-wide cascade. - -### 7.1 Three Namespaces - -Namespace means four different things across the boundary, and the distinction between any two of them is a privilege -boundary. - -| Namespace | Side | Holds | Is the boundary for | -|------------------------------------|---------|-------------------------------------------------------------------------------|----------------------------------------| -| A tenant namespace | Hub | Every Band B object: the member, its deployments, its classes, its operations | Who may operate on this tenant's fleet | -| The member's OCM namespace | Hub | `ManifestWork`, `ManagedClusterView`, and the add-on's registration | Open Cluster Management's own layout | -| The operator's namespace | Managed | The operator workload, its webhooks, and its own configuration | Who may operate the installation | -| Any namespace a deployment chooses | Managed | Every Band A object a payload creates or a cluster expands | What a team may consume in the member | - -The first two are both on the hub and stay separate. Open Cluster Management requires a `ManifestWork` to live in its -`ManagedCluster`'s namespace, which is OCM's addressing rather than a tenancy decision, and a tenant needs no rights -there: write access to a member's OCM namespace is write access to any payload into that member, which is every -permission the fleet holds over it. The Band B object is therefore a tenant's to write, the `ManifestWork` is the -controller's, and the controller is the only thing that crosses between them. - -The last two are both in the member, and the operator's own namespace is not where its objects live. The operator -workload runs in `simplyblock-system` and watches cluster-wide: `WATCH_NAMESPACE` is set by the chart and read nowhere, -and `operator/internal/upgrade/scope.go` states that a `StorageCluster` in any namespace belongs to the installation. A -`StorageCluster`, its nodes, and its pools therefore live wherever a deployment puts them, which is a choice a -standalone deployment already makes and which the hub makes too. - -A payload states its namespace rather than inheriting one. There is no default to fall back on, so the document a -`ClusterDeployment` carries names the namespace its objects are created in, and `status.result.namespace` records what -was used. It bears no relation to the hub-side namespace the intent was written in. - -The same fact fixes what the admission guard keys on. The set of namespaces a deployment may use is not knowable in -advance, so §10's guard keys on the work agent's field ownership, which travels with the object wherever it is written. - -The `ManifestWork` objects these controllers create live in the member's OCM namespace, which is a different namespace -from the Band B object's, so an owner reference is unavailable for the reason -[`design-crd-model.md`](operator/docs/designs/crd-redesign/design-crd-model.md) §5 gives for the `StorageClass`. They -carry `fleet.simplyblock.io/managed-by`, and a finalizer performs the cascade, which is the group's existing pattern. - -### 7.2 StorageFleet and FleetMember - -`StorageFleet` is a singleton named `simplyblock`, matching `ControlPlane`'s convention: there is exactly one control -plane, and no kind carries a reference by which a controller could select among several. Its spec is where that control -plane is, and it is immutable as a block, because re-pointing a live fleet at another control plane produces a different -fleet. - -`FleetMember` is one enrolled Kubernetes cluster, and its spec holds two fields: - -```go -// ClusterRef names the Open Cluster Management ManagedCluster this member is. -// That kind is cluster-scoped and its names are unique, so a bare name is -// unambiguous. -// +kubebuilder:validation:Required -// +k8s:immutable -ClusterRef string `json:"clusterRef"` - -// DetachPolicy is what happens to what the fleet applied when this member is -// deleted. Retain leaves the member a working standalone deployment. -// +kubebuilder:validation:Enum=Retain;Delete -// +kubebuilder:default=Retain -// +optional -DetachPolicy FleetDetachPolicy `json:"detachPolicy,omitempty"` -``` - -Everything else about a member is observed rather than declared. The status splits into three blocks that come from -different places and go stale independently: `link` is the add-on's own report, `inventory` is a summary of the -discovery reports, and `storage` is the roll-up read from the control plane. Each carries the time it was last current, -for the reason §4 gives. - -Detaching is expressed as a deletion. Deleting a `FleetMember` is the instruction to unenroll, and a finalizer performs -it: patch every `ManifestWork` for the member to `propagationPolicy: Orphan` when the policy is `Retain`, delete them, -then withdraw the add-on. Orphaning before withdrawing is what leaves a working standalone deployment behind, and it is -why `deleteOption.propagationPolicy` is load-bearing. - -`status.inventory` is a summary rather than the reports. A discovery run produces one `ConfigMap` per worker, each -admitted up to a mebibyte (`operator/internal/nodeprobe/configmap.go`), so a fifty-worker member is fifty objects and -tens of megabytes. A fleet-wide list needs the counts, the device classes, and the raw capacity. The full reports stay -in the member and are read through a `ManagedClusterView` when a deployment is composed. - -### 7.3 ClusterDeployment and DriverDeployment - -Both carry a Band A spec as their payload, and both are entities rather than operations, because a deployment is applied -rather than operated, which is the reading `ClusterDeploymentConfig` already has. - -```go -// Config is the document that will be applied in the member, composed from that -// member's inventory. It is a typed Band A spec rather than an opaque blob -// because it is the thing a reviewer reads. -// +kubebuilder:validation:Required -Config storagev1alpha2.ClusterDeploymentConfigSpec `json:"config"` - -// Approved is the instruction to build, and setting it is the deployment rather -// than a review stage in front of one. -// +optional -Approved bool `json:"approved,omitempty"` -``` - -Two CEL rules carry what markers cannot: - -```go -// +kubebuilder:validation:XValidation:rule="!oldSelf.approved || self.config == oldSelf.config",message="config is immutable once approved" -// +kubebuilder:validation:XValidation:rule="!oldSelf.approved || self.approved",message="approval cannot be withdrawn" -``` - -The document is editable while it is a draft and frozen once approved, which is the one place this kind departs from -`ClusterDeploymentConfig`'s immutability. A draft composed at the hub is edited at the hub, and the Band A object is -created only once the hub-side document is settled, so what ships down is always an already-approved document and the -member never holds a draft nobody approved. - -`spec.approvedBy` records the identity that approved. A webhook sets it from the request's user info, because the work -agent replaying the edit into the member is otherwise the only approver the audit trail carries, and -[`design-clusterdeploymentconfig.md`](operator/docs/designs/crd-redesign/design-clusterdeploymentconfig.md) §5 made -approval a spec field for its audit surface. - -`DriverDeployment` is the same shape around `SimplyblockDriverSpec`, without the approval, because a driver version is -an edit rather than a deployment gate. It is one object per member, so a fleet-wide rollout is a set of them, and §11 -carries the kind that would generate the set. - -### 7.4 StorageClassDeployment - -A class is how a claim asks a pool for capacity, and -[`design-storagepool.md`](operator/docs/designs/crd-redesign/design-storagepool.md) §5 settles the two facts that decide -the hub's shape for it. - -A class is authored rather than generated, and a pool may have zero or more. One pool can back a class with compression -on, another with it off, one formatted `ext4`, one formatted `xfs`, a permissive ceiling for a batch tenant, and a tight -one for a latency-sensitive one. Configuring a class for a pool that already exists is therefore the ordinary case, and -it changes nothing about the pool. - -The assignment is three labels on the class. `storage.simplyblock.io/namespace`, `.../cluster`, and `.../pool` are what -say which pool a class draws from, in either direction, and `StoragePool.status.storageClassNames` is the pool's side of -the same selector. Neither side owns the other. - -One class is expanded locally and the rest come from the hub. A cluster creation writes one class for its default pool, -carrying `storage.simplyblock.io/managed-by: storagecluster`, and that happens in a managed cluster exactly as it does -in a standalone one, because it is Band A behavior the hub does not participate in. Every class beyond that first one is -authored, and on a managed cluster the hub authors it. - -A `StorageClass` is therefore an independent object, shipped as its own payload. - -```go -// Pool is the pool this class draws on, in the member. The three fields become -// the three assignment labels on the class, so an author states the pool and -// never writes a label by hand. -// +kubebuilder:validation:Required -// +k8s:immutable -Pool StorageClassPoolRef `json:"pool"` - -// Template is what the class carries beyond its assignment. Parameters are -// written in the superseding QoS spelling only, which is what the operator's own -// generated class does. -// +kubebuilder:validation:Required -Template StorageClassTemplate `json:"template"` -``` - -A namespaced hub kind is what makes a cluster-scoped object delegable. `StorageClass` is cluster-scoped, and -`resourceNames` covers `get`, `update`, and `delete` but never `create` or `list`, so RBAC cannot express "may author -classes for this pool" against the class itself. A single cluster offers no namespaced object to hang that permission -on. The hub offers one: a pool's owner holds verbs on `storageclassdeployments` in their own namespace and needs no -rights on `StorageClass` anywhere, and the work agent writes the cluster-scoped object in the member under its own -identity. The delegation boundary sits at the hub, and the cardinality stays whatever the pool needs. - -Deleting the Band B object deletes the class, without colliding with the operator. A `ManifestWork` with the default -`propagationPolicy: Foreground` removes what it applied, and §6 of the pool design lets the operator delete only a class -carrying `storage.simplyblock.io/managed-by: storagecluster`, which is the one class it generates for a default pool. A -hub-shipped class carries the fleet's marker, so each side deletes what it created and refuses on the other's. - -`spec.className` is immutable, because it is the object's name in the member and a rename produces a different class. -Everything under `template` is editable, so re-tuning a ceiling is an edit rather than a delete and a re-create. - -### 7.5 FleetOperation - -One kind ships any Band A `Ops` object into one member. - -```go -// Operation is the operation to run in the member. Exactly one member is set, -// and each is the spec of the Band A kind it names. -// +kubebuilder:validation:XValidation:rule="[has(self.operatorOps),has(self.storageClusterOps),has(self.storageNodeOps),has(self.storageDeviceOps),has(self.storagePoolOps),has(self.storageBackupOps)].filter(x, x).size() == 1",message="set exactly one operation" -// +kubebuilder:validation:Required -// +k8s:immutable -Operation FleetOperationTarget `json:"operation"` -``` - -The single kind is a deliberate break with the `Ops` convention. That convention holds that a kind ending in -`Ops` is one-shot and names one target, and this kind is both. What it does not name is a Band B entity: its target is -an object in a member, and the parameters belong to the Band A kind whose spec it carries. One hub kind per managed -`Ops` kind reintroduces the twin for the one category where re-creating an object is dangerous. - -The discriminated block follows the `ControlPlane.spec.source` pattern: one member per variant, exactly one set, -enforced in CEL. The block is typed rather than an opaque `RawExtension`, because this is the payload that performs a -side effect and it is the one that most needs validation. - -Status comes back in two pieces, which is how a managed operation is observed from the hub: - -```go -// Remote is what the member's own object reports, filled from the status -// feedback rules on the ManifestWork. Scalars only, which is what a feedback -// rule carries and what a list view needs. -// +optional -Remote *RemoteOpsStatus `json:"remote,omitempty"` -``` - -`RemoteOpsStatus` holds the phase, the step's state, the message, and the observed generation. Anything beyond that is a -`ManagedClusterView` against the object the payload created, on demand, which returns a whole and current object rather -than a copy of one. - -The block enumerates the target set rather than what compiles today. Three of its members, `StorageDeviceOpsSpec`, -`StoragePoolOpsSpec`, and `StorageBackupOpsSpec`, are specified in their designs and declared nowhere in `operator/api` -yet, and two more are still at `v1alpha1`. That is the state of the tree rather than a constraint on this design: Band -A's redesign and Band B are one body of work, and the members exist by the time either ships. - ---- - -## 8. The Hub as a Gateway for `storage.simplyblock.io` - -A console reads the same kinds on the hub that it reads in a single cluster: one set of types, one client, one -authorization model. §5 keeps Band A off the hub as stored objects, and a kind is served by whatever is registered for -its group, which a `CustomResourceDefinition` is one of two ways to be. - -The operator already runs an aggregated API server. `operator/internal/metricsapi/` is the implementation, and -`helm-charts/charts/simplyblock-operator/templates/metrics-apiserver.yaml` ships the `APIService` and the two delegation -bindings. [`design-crd-model.md`](operator/docs/designs/crd-redesign/design-crd-model.md) §7.13 states the trade: a kind -served that way is computed when a client asks for it and is never persisted, which is what `metrics.k8s.io` does for -`PodMetrics`. Its only consumer today is one measurement kind, which is the thinnest use of the mechanism rather than -its purpose. A serving path, a certificate rotator, an `APIService`, and the authorization delegation are all built, and -none of them knows what group it is serving. - -The hub therefore registers `storage.simplyblock.io` as an `APIService` rather than as `CustomResourceDefinition` -objects, and serves it by calling the control plane it sits beside. `kubectl -n cluster-a get storageclusters` answers -on the hub, and the console's read path is the code it already has. - -Two properties carry that. A client cannot tell the difference, because the group, the versions, the kinds, and the -verbs are identical over the wire whichever way they are registered. And authorization is unchanged, because the -delegation bindings the chart already ships are what make an ordinary `RoleBinding` on an aggregated group work. - -A cluster serves a group one way or the other, never both, and each cluster chooses independently. A standalone or -managed cluster is CRD-backed and the hub is aggregated, so §3's test is untouched: the hub is neither. - -### 8.1 Namespacing - -The projection serves a member's objects in a namespace on the hub, and which one follows from the tenancy model rather -than from Open Cluster Management. Serving them in the tenant namespace of §7.1 lets one `RoleBinding` cover a tenant's -intent and the state it produced, at the cost of a disambiguator when one tenant holds several members. Serving them in -the member's own namespace inverts both. Either is a hub-side namespace and not the member's operator namespace. - -Under either choice, names that would collide across members do not, and cluster-scoped Band A kinds become namespaced -on the hub with no translation the console has to know about. - -### 8.2 Versions, and What the Projection Omits - -Registration is per group-version. An `APIService` is named `.` and carries the two as separate fields, -which `operator/config/apiservice/apiservice.yaml` shows for `v1alpha1.metrics.simplyblock.io`. A gateway for -`storage.simplyblock.io/v1alpha2` is one object, and serving `v1alpha1` beside it is a second. - -The gateway serves one version. A managed cluster cannot predate the arrangement, so it is a new installation, it starts -at `v1alpha2`, and the gateway registers `v1alpha2.storage.simplyblock.io` and nothing else. No conversion is needed, -because no second version is in play. - -The cost is deferred rather than absent. An aggregated server gets no conversion webhook the way a -`CustomResourceDefinition` does, so it is responsible for every version it claims. The first hub that serves two -versions of the group at once owns a conversion the operator has already written for its own upgrade path. - -A status in this group is two things, and only one is the control plane's. -[`design-storagedevice.md`](operator/docs/designs/crd-redesign/design-storagedevice.md) §4.2 draws the line in one kind -and it holds across the group: `status.phase` is the operator's own view, and `status.deviceStatus` is the control -plane's string kept in its own spelling. Everything on the operator's side is absent from a control-plane projection, -which means `status.phase`, `status.conditions`, `status.message`, `status.observedGeneration`, `status.activeOpsRef`, -`status.step`, and every `Ops` object. - -The projection therefore serves an inventory console. It answers what the fleet has and how full it is. What an operator -is doing about a node right now comes from §9. - ---- - -## 9. The Three Transports - -Open Cluster Management defines three, and a console uses a different one per view. - -| Transport | Direction | Carries | -|-----------------------|----------------|------------------------------------------------------------------------------------------------------------------------------| -| `ManifestWork` | Hub to managed | A complete resource as an opaque payload, applied locally by the work agent, with its own conditions from pending to applied | -| Status feedback rules | Managed to hub | Named scalars selected by JSON path out of the applied resource, on a poll or a watch | -| `ManagedClusterView` | Managed to hub | One named object, fetched on demand, returned as raw JSON | - -Alongside the projection of §8 they make four tiers: - -| Tier | Mechanism | Serves | Costs | -|------|----------------------|------------------------------------------------------|----------------------------------------------| -| 1 | The projection | Fleet-wide lists and search, from the control plane | Nothing per member. It never leaves the hub | -| 2 | Status feedback | An operation's phase, step, and message in a list | Scalars only, so a phase and not a condition | -| 3 | `ManagedClusterView` | One real object, whole, from the member's API server | A round trip, and the link has to be up | -| 4 | `cluster-proxy` | The member's API server directly, including events | The same, plus an add-on and a tunnel | - -Tier 2 is what makes a fleet list of operations useful. A phase, a step state, and an observed generation are single -values, which is what a feedback rule carries, so a list of what is running across the fleet needs no per-object round -trip. - -Tier 4 is optional and answers what the other three cannot. The `cluster-proxy` add-on runs a proxy server on the hub -and an agent in each member, and the agent dials the tunnel outbound, so a hub component reaches a member's -`kube-apiserver` by naming the member. Through it a console reads the actual objects and the actual `Events`, which live -in no object at all and which -[`design-crd-model.md`](operator/docs/designs/crd-redesign/design-crd-model.md) §3.3 identifies as the only thing -separating a queued operation from a new one. It uses the same outbound direction the work agent already does, so it -adds no egress path and no credential. - -Tiers 3 and 4 answer about one member at a time. Which members have a node in `Degraded` is a fan-out over N clusters -with N failure modes, which is what tier 1 answers. - ---- - -## 10. The Add-Ons - -Three, of which the first two are required. - -| Add-on | Does | -|-------------------------|-------------------------------------------------------------------------------------------------------------| -| The inventory publisher | Publishes the merged discovery inventory after a discovery operation, and the Kubernetes-shaped facts of §4 | -| The admission guard | Refuses a spec write to an agent-applied or agent-expanded object from any identity but the work agent | -| `cluster-proxy` | Optional. Tier 4 of §9, where single-object access is not enough | - -The guard ships with the add-on, which is what keeps §3's test true: a standalone cluster carries no guard because it -carries no add-on, and the operator's webhook configuration is unchanged. It keys on the work agent's field ownership -rather than on a list of kinds, so it needs no per-kind configuration. - -Field-level exceptions are declared in the payload. `ManifestWork`'s `updateStrategy.type: ServerSideApply` takes a -`fieldManager` and an `ignoreFields` list, whose `condition: OnSpokePresent` stops reconciling a path once the resource -exists in the member. That is how the three `StorageNode` fields a migration writes stay the operator's while the rest -of the spec stays the hub's, and it is declarative and travels with the payload. - -Two other `ManifestWork` settings are load-bearing. `updateStrategy.type: CreateOnly` ensures creation and never -re-applies, which is what makes a one-shot `Ops` payload safe. And `deleteOption.propagationPolicy` carries the detach -semantics of §7.2. - ---- - -## 11. Open Questions - -- **Both halves of the picture can be stale at once.** §4 states the reason: an add-on that cannot reach the hub and - nodes that cannot reach the control plane are usually the same event. Every block of `FleetMember.status` carries when - it was last current, and a console that renders them as one panel loses the distinction. -- **The gateway owes a conversion implementation the first time it serves two versions of the group** (§8.2). Not now, - because a member is a new installation and starts at `v1alpha2`. -- **`status.step.deadline` is an absolute `metav1.Time`.** Clock skew makes a step's remaining time wrong at the hub. A - start instant and a duration fix it, and the blast radius is `atlas-lib/statemachine`, - [`design-crd-model.md`](operator/docs/designs/crd-redesign/design-crd-model.md) §3.1, and every `Ops` kind's - `status.step`. -- **A discovery report that cannot be parsed is skipped with an event**, and its worker is left out of the draft. - `FleetMember.status.inventory.unreadableReports` is where that becomes visible at the hub, as a count rather than a - reason. -- **A fleet-wide driver rollout is a set of `DriverDeployment` objects.** A kind generating the set from a selection has - a partial-success outcome to represent, which the `Ops` shape does not model, so it is an entity with a per-member - status list and it needs its own working out. -- **Tenancy is mostly RBAC, and the quantities are an allocation envelope beside it.** Band B kinds are namespaced and - the projection serves a member in its own namespace, so a tenant is a set of namespaces and ordinary `RoleBinding` - objects separate them, including on the aggregated group. A quantity is what RBAC cannot express, and a separate - draft, `simplyblock-multicluster-rbac-design.md`, is where that lives: an admin-owned, cluster-scoped allocation - naming the nodes and devices a namespace may claim and the cores and hugepages it may take, enforced by a - `ValidatingAdmissionPolicy` with the allocation as its `paramKind`, and a `ResourceQuota` per consumer namespace for - capacity. That draft also settles whether a tenant owns whole Kubernetes clusters or pools inside a shared one, which - is the structural choice underneath all of it. `resourceNames` does not apply to `list` or `watch`, so a pool admin in - a shared namespace either lists every pool or is refused, with no filtered middle. -- **Neither of that draft's two arguments against the kinds here holds.** Its treatment of `StorageClass` assumes one - class per pool, and §7.4 states why the count is whatever the pool needs and why a namespaced hub kind delegates the - cluster-scoped object. Its rule that an action is a resource rather than a field would split `FleetOperation` into one - kind per operation family, and §7.5 settles that the other way. - ---- - -## Appendix A: `fleet.simplyblock.io` types - -The six kinds of §7, whole. Package `fleet/v1alpha1`, importing `storagev1alpha2` -for the Band A specs §7 explains the coupling to, and `statemachine` for the step snapshot of -[`design-crd-model.md`](operator/docs/designs/crd-redesign/design-crd-model.md) §3.1. - -```go -// +kubebuilder:object:root=true -// +kubebuilder:subresource:status -// +kubebuilder:resource:scope=Namespaced,shortName=sf -// +kubebuilder:printcolumn:name="Phase",type=string,JSONPath=".status.phase" -// +kubebuilder:printcolumn:name="Endpoint",type=string,JSONPath=".status.endpoint" -// +kubebuilder:printcolumn:name="Version",type=string,JSONPath=".status.version" -// +kubebuilder:printcolumn:name="Members",type=integer,JSONPath=".status.memberCount" -// +kubebuilder:printcolumn:name="Age",type=date,JSONPath=".metadata.creationTimestamp" - -// StorageFleet is the control plane a fleet is built on, and the fleet's -// defaults. One object per hub, named simplyblock. -type StorageFleet struct { - metav1.TypeMeta `json:",inline"` - metav1.ObjectMeta `json:"metadata,omitempty"` - - Spec StorageFleetSpec `json:"spec,omitempty"` - Status StorageFleetStatus `json:"status,omitempty"` -} - -// +kubebuilder:object:root=true - -// StorageFleetList is a list of StorageFleet. -type StorageFleetList struct { - metav1.TypeMeta `json:",inline"` - metav1.ListMeta `json:"metadata,omitempty"` - Items []StorageFleet `json:"items"` -} - -// StorageFleetSpec is where the fleet's control plane is, and what its -// deployments inherit. -type StorageFleetSpec struct { - // ControlPlane is the control plane every member of this fleet is joined to. - // +kubebuilder:validation:Required - // +k8s:immutable - ControlPlane FleetControlPlane `json:"controlPlane"` - - // Defaults are the values a member inherits when it omits them. - // +optional - Defaults *FleetDefaults `json:"defaults,omitempty"` -} - -// FleetControlPlane addresses the management API. It is the shape -// ControlPlane.spec.source.external takes, because it describes the same -// deployment from the other side. -type FleetControlPlane struct { - // Endpoint is the management API's base URL. - // +kubebuilder:validation:Pattern=`^https?://[a-zA-Z0-9.-]+(:[0-9]{1,5})?(/.*)?$` - // +kubebuilder:validation:Required - Endpoint string `json:"endpoint"` - - // CredentialsSecretRef names a Secret in this namespace holding the token. - // +kubebuilder:validation:Required - CredentialsSecretRef corev1.LocalObjectReference `json:"credentialsSecretRef"` -} - -// FleetDefaults are fleet-wide values a member inherits. -type FleetDefaults struct { - // OpsRetention is how long a terminal FleetOperation is kept. - // +optional - OpsRetention *metav1.Duration `json:"opsRetention,omitempty"` - - // InventoryRefreshInterval is how often a member is asked to re-probe. - // +optional - InventoryRefreshInterval *metav1.Duration `json:"inventoryRefreshInterval,omitempty"` -} - -// StorageFleetStatus is what the control plane reports about itself. -type StorageFleetStatus struct { - // Phase is whether the control plane answers. - // +optional - Phase StorageFleetPhase `json:"phase,omitempty"` - - // Endpoint is the resolved base URL, so that a reader asks status not spec. - // +optional - Endpoint string `json:"endpoint,omitempty"` - - // Version is the management API's reported version. - // +optional - Version string `json:"version,omitempty"` - - // MemberCount is how many FleetMember objects this fleet has. - // +optional - MemberCount int32 `json:"memberCount,omitempty"` - - // LastChecked is when the readiness probe last ran. - // +optional - LastChecked *metav1.Time `json:"lastChecked,omitempty"` - - // Message is why the phase is what it is. - // +optional - Message string `json:"message,omitempty"` - - // ObservedGeneration is the generation this status was computed from. - // +optional - ObservedGeneration int64 `json:"observedGeneration,omitempty"` - - // +listType=map - // +listMapKey=type - // +optional - Conditions []metav1.Condition `json:"conditions,omitempty"` -} - -// StorageFleetPhase is whether the fleet's control plane answers. There is no -// Installing value: the control plane exists before the hub does. -// +kubebuilder:validation:Enum=Available;Degraded;Unavailable -type StorageFleetPhase string - -const ( - StorageFleetPhaseAvailable StorageFleetPhase = "Available" - StorageFleetPhaseDegraded StorageFleetPhase = "Degraded" - StorageFleetPhaseUnavailable StorageFleetPhase = "Unavailable" -) -``` - -```go -// +kubebuilder:object:root=true -// +kubebuilder:subresource:status -// +kubebuilder:resource:scope=Namespaced,shortName=fm -// +kubebuilder:printcolumn:name="Cluster",type=string,JSONPath=".spec.clusterRef" -// +kubebuilder:printcolumn:name="Phase",type=string,JSONPath=".status.phase" -// +kubebuilder:printcolumn:name="Workers",type=integer,JSONPath=".status.inventory.eligibleWorkers" -// +kubebuilder:printcolumn:name="Nodes",type=integer,JSONPath=".status.storage.nodesOnline" -// +kubebuilder:printcolumn:name="Contact",type=date,JSONPath=".status.link.lastContact" -// +kubebuilder:printcolumn:name="Age",type=date,JSONPath=".metadata.creationTimestamp" - -// FleetMember is one enrolled Kubernetes cluster. Its spec says which cluster -// and what a detachment does, because everything else about a member is -// observed rather than declared. -type FleetMember struct { - metav1.TypeMeta `json:",inline"` - metav1.ObjectMeta `json:"metadata,omitempty"` - - Spec FleetMemberSpec `json:"spec,omitempty"` - Status FleetMemberStatus `json:"status,omitempty"` -} - -// +kubebuilder:object:root=true - -// FleetMemberList is a list of FleetMember. -type FleetMemberList struct { - metav1.TypeMeta `json:",inline"` - metav1.ListMeta `json:"metadata,omitempty"` - Items []FleetMember `json:"items"` -} - -// FleetMemberSpec identifies the Kubernetes cluster a member is. -type FleetMemberSpec struct { - // ClusterRef names the Open Cluster Management ManagedCluster this member is. - // That kind is cluster-scoped and its names are unique, so a bare name is - // unambiguous. - // +kubebuilder:validation:Required - // +k8s:immutable - ClusterRef string `json:"clusterRef"` - - // DetachPolicy is what happens to what the fleet applied when this member is - // deleted. Retain orphans it, leaving a working standalone deployment. - // +kubebuilder:default=Retain - // +optional - DetachPolicy FleetDetachPolicy `json:"detachPolicy,omitempty"` - - // InventoryRefreshInterval is how often the hub asks the member to re-probe - // its workers. Absent inherits the fleet's default, and zero means never. - // +optional - InventoryRefreshInterval *metav1.Duration `json:"inventoryRefreshInterval,omitempty"` -} - -// FleetDetachPolicy is what a member's deletion does to what the fleet applied. -// +kubebuilder:validation:Enum=Retain;Delete -type FleetDetachPolicy string - -const ( - FleetDetachPolicyRetain FleetDetachPolicy = "Retain" - FleetDetachPolicyDelete FleetDetachPolicy = "Delete" -) - -// FleetMemberStatus is the member's readiness and three blocks that come from -// different places and go stale independently. -type FleetMemberStatus struct { - // Phase is the member's readiness to be deployed into. - // +optional - Phase FleetMemberPhase `json:"phase,omitempty"` - - // Link is what the add-on reports about itself, and it goes stale with it. - // +optional - Link *FleetMemberLink `json:"link,omitempty"` - - // Inventory is a summary and never the reports, which stay in the member. - // +optional - Inventory *FleetMemberInventory `json:"inventory,omitempty"` - - // Storage is read from the control plane rather than from the member. - // +optional - Storage *FleetMemberStorage `json:"storage,omitempty"` - - // Message is why the phase is what it is. - // +optional - Message string `json:"message,omitempty"` - - // ObservedGeneration is the generation this status was computed from. - // +optional - ObservedGeneration int64 `json:"observedGeneration,omitempty"` - - // +listType=map - // +listMapKey=type - // +optional - Conditions []metav1.Condition `json:"conditions,omitempty"` -} - -// FleetMemberPhase is a member's readiness to be deployed into. -// +kubebuilder:validation:Enum=Pending;Enrolling;Ready;Degraded;Unreachable;Detaching -type FleetMemberPhase string - -const ( - FleetMemberPhasePending FleetMemberPhase = "Pending" - FleetMemberPhaseEnrolling FleetMemberPhase = "Enrolling" - FleetMemberPhaseReady FleetMemberPhase = "Ready" - FleetMemberPhaseDegraded FleetMemberPhase = "Degraded" - FleetMemberPhaseUnreachable FleetMemberPhase = "Unreachable" - FleetMemberPhaseDetaching FleetMemberPhase = "Detaching" -) - -// FleetMemberLink is what the add-on reports about the member it runs in. -type FleetMemberLink struct { - // AddOnVersion is the add-on's own version. - // +optional - AddOnVersion string `json:"addOnVersion,omitempty"` - - // OperatorVersion is the simplyblock operator the member runs. - // +optional - OperatorVersion string `json:"operatorVersion,omitempty"` - - // DriverVersion is the CSI driver the member runs. - // +optional - DriverVersion string `json:"driverVersion,omitempty"` - - // StorageAPIVersions are the versions of storage.simplyblock.io the member - // serves, which is what makes schema skew visible before a payload is pruned - // by it rather than after. - // +optional - StorageAPIVersions []string `json:"storageAPIVersions,omitempty"` - - // LastContact is when the add-on last reported. - // +optional - LastContact *metav1.Time `json:"lastContact,omitempty"` -} - -// FleetMemberInventory summarizes what a member's workers have. The reports -// themselves stay in the member, because a fifty-worker set is tens of -// megabytes and one member's copy is all a composition needs. -type FleetMemberInventory struct { - // Workers is how many the member has. - // +optional - Workers int32 `json:"workers,omitempty"` - - // EligibleWorkers is how many a cluster could be built on. - // +optional - EligibleWorkers int32 `json:"eligibleWorkers,omitempty"` - - // DeviceClasses are the classes present, so a fleet-wide list can answer - // which members could take a cluster of a class without fetching a report. - // +optional - DeviceClasses []string `json:"deviceClasses,omitempty"` - - // RawCapacity is the sum of the candidate devices. - // +optional - RawCapacity *resource.Quantity `json:"rawCapacity,omitempty"` - - // Revision identifies the report set this summary was computed from. - // +optional - Revision string `json:"revision,omitempty"` - - // ObservedAt is when the summary was computed, which is what stops a stale - // one reading as current. - // +optional - ObservedAt *metav1.Time `json:"observedAt,omitempty"` - - // UnreadableReports counts workers whose report could not be parsed, which - // is otherwise reported only by an event inside the member. - // +optional - UnreadableReports int32 `json:"unreadableReports,omitempty"` -} - -// FleetMemberStorage is the control plane's count of what the member holds. -type FleetMemberStorage struct { - // Clusters is how many storage clusters the member holds. - // +optional - Clusters int32 `json:"clusters,omitempty"` - - // Nodes and NodesOnline are the member's storage nodes and how many answer. - // +optional - Nodes int32 `json:"nodes,omitempty"` - // +optional - NodesOnline int32 `json:"nodesOnline,omitempty"` - - // Pools is how many pools the member's clusters carry. - // +optional - Pools int32 `json:"pools,omitempty"` - - // DevicesDegraded is how many devices are serving and should not be. - // +optional - DevicesDegraded int32 `json:"devicesDegraded,omitempty"` - - // ObservedAt is when the control plane was last read, so that a stale - // projection is distinguishable from a current one. - // +optional - ObservedAt *metav1.Time `json:"observedAt,omitempty"` -} -``` - -```go -// +kubebuilder:object:root=true -// +kubebuilder:subresource:status -// +kubebuilder:resource:scope=Namespaced,shortName=cd -// +kubebuilder:printcolumn:name="Member",type=string,JSONPath=".spec.memberRef" -// +kubebuilder:printcolumn:name="Approved",type=boolean,JSONPath=".spec.approved" -// +kubebuilder:printcolumn:name="Phase",type=string,JSONPath=".status.phase" -// +kubebuilder:printcolumn:name="Step",type=string,JSONPath=".status.step.state" -// +kubebuilder:printcolumn:name="Cluster",type=string,JSONPath=".status.result.storageClusterName" -// +kubebuilder:printcolumn:name="Age",type=date,JSONPath=".metadata.creationTimestamp" - -// ClusterDeployment is the intent to build one storage cluster in one member, -// the document that will be applied there, and the record of it landing. -type ClusterDeployment struct { - metav1.TypeMeta `json:",inline"` - metav1.ObjectMeta `json:"metadata,omitempty"` - - Spec ClusterDeploymentSpec `json:"spec,omitempty"` - Status ClusterDeploymentStatus `json:"status,omitempty"` -} - -// +kubebuilder:object:root=true - -// ClusterDeploymentList is a list of ClusterDeployment. -type ClusterDeploymentList struct { - metav1.TypeMeta `json:",inline"` - metav1.ListMeta `json:"metadata,omitempty"` - Items []ClusterDeployment `json:"items"` -} - -// ClusterDeploymentSpec is the document and the approval. The document is -// editable while it is a draft and frozen once approved, which is the one place -// this kind departs from ClusterDeploymentConfig's immutability. -// +kubebuilder:validation:XValidation:rule="!oldSelf.approved || self.config == oldSelf.config",message="config is immutable once approved" -// +kubebuilder:validation:XValidation:rule="!oldSelf.approved || self.approved",message="approval cannot be withdrawn" -type ClusterDeploymentSpec struct { - // MemberRef names the FleetMember this cluster is built in. - // +kubebuilder:validation:Required - // +k8s:immutable - MemberRef string `json:"memberRef"` - - // Config is the document that will be applied in the member, composed from - // that member's inventory. It is a typed Band A spec rather than an opaque - // blob because it is the thing a reviewer reads. - // +kubebuilder:validation:Required - Config storagev1alpha2.ClusterDeploymentConfigSpec `json:"config"` - - // Approved is the instruction to build, and setting it is the deployment - // rather than a review stage in front of one. - // +optional - Approved bool `json:"approved,omitempty"` - - // ApprovedBy records the identity that approved, which the service account - // replaying the edit into the member would otherwise erase. A webhook sets - // it from the request's user info. - // +optional - ApprovedBy string `json:"approvedBy,omitempty"` -} - -// ClusterDeploymentStatus is how far the build-out got, and what it produced. -type ClusterDeploymentStatus struct { - // Phase is where the build-out is. - // +optional - Phase ClusterDeploymentPhase `json:"phase,omitempty"` - - // Step is the delivery machine's position. - // +kubebuilder:validation:XValidation:rule="!has(self.state) || self.state in ['Composing','Delivering','Applying','Expanding','Verifying']",message="unknown step" - // +optional - Step statemachine.KubeSnapshot `json:"step,omitempty"` - - // ConfigFingerprint is a hash of spec.config, and Delivery.AppliedFingerprint - // is the one that reached the member. The pair is what makes an edit's - // arrival answerable without a generation the hub did not issue. - // +optional - ConfigFingerprint string `json:"configFingerprint,omitempty"` - - // Delivery is what the ManifestWork carrying this document reports. - // +optional - Delivery *DeliveryStatus `json:"delivery,omitempty"` - - // Result names what the member built. - // +optional - Result *ClusterDeploymentResult `json:"result,omitempty"` - - // Message is why the phase is what it is. - // +optional - Message string `json:"message,omitempty"` - - // StartedAt is when delivery began. - // +optional - StartedAt *metav1.Time `json:"startedAt,omitempty"` - - // CompletedAt is when the build-out reached a terminal phase. - // +optional - CompletedAt *metav1.Time `json:"completedAt,omitempty"` - - // ObservedGeneration is the generation this status was computed from. - // +optional - ObservedGeneration int64 `json:"observedGeneration,omitempty"` - - // +listType=map - // +listMapKey=type - // +optional - Conditions []metav1.Condition `json:"conditions,omitempty"` -} - -// ClusterDeploymentPhase is where a build-out is. Draft is a phase and not a -// step, because a document waiting for approval is not a delivery that stalled. -// +kubebuilder:validation:Enum=Draft;Delivering;Deploying;Ready;Degraded;Failed -type ClusterDeploymentPhase string - -const ( - ClusterDeploymentPhaseDraft ClusterDeploymentPhase = "Draft" - ClusterDeploymentPhaseDelivering ClusterDeploymentPhase = "Delivering" - ClusterDeploymentPhaseDeploying ClusterDeploymentPhase = "Deploying" - ClusterDeploymentPhaseReady ClusterDeploymentPhase = "Ready" - ClusterDeploymentPhaseDegraded ClusterDeploymentPhase = "Degraded" - ClusterDeploymentPhaseFailed ClusterDeploymentPhase = "Failed" -) - -// ClusterDeploymentStep is one step of the delivery machine. -// +kubebuilder:validation:Enum=Composing;Delivering;Applying;Expanding;Verifying -type ClusterDeploymentStep string - -const ( - ClusterDeploymentStepComposing ClusterDeploymentStep = "Composing" - ClusterDeploymentStepDelivering ClusterDeploymentStep = "Delivering" - ClusterDeploymentStepApplying ClusterDeploymentStep = "Applying" - ClusterDeploymentStepExpanding ClusterDeploymentStep = "Expanding" - ClusterDeploymentStepVerifying ClusterDeploymentStep = "Verifying" -) - -// ClusterDeploymentResult identifies what the member built, so that a console -// reaches it through the projection without re-deriving the name. -type ClusterDeploymentResult struct { - // StorageClusterName is the object's name in the member. - // +optional - StorageClusterName string `json:"storageClusterName,omitempty"` - - // StorageClusterUUID is the control plane's identifier, which is what the - // roll-up and the projection address the same object by. - // +optional - StorageClusterUUID string `json:"storageClusterUUID,omitempty"` - - // Namespace is where the objects were created in the member. - // +optional - Namespace string `json:"namespace,omitempty"` - - // NodesExpected and NodesReady are what the document asked for and what - // answered. - // +optional - NodesExpected int32 `json:"nodesExpected,omitempty"` - // +optional - NodesReady int32 `json:"nodesReady,omitempty"` -} - -// DeliveryStatus is what the ManifestWork carrying a payload reports, and the -// only place a refused write can be seen from the hub. It is shared by every -// kind in this group that delivers one. -type DeliveryStatus struct { - // ManifestWorkName is the object in the member's namespace on the hub. - // +optional - ManifestWorkName string `json:"manifestWorkName,omitempty"` - - // AppliedFingerprint is the fingerprint of what actually reached the member. - // +optional - AppliedFingerprint string `json:"appliedFingerprint,omitempty"` - - // Conditions mirror the ManifestWork's Applied, Available, and Degraded, - // each written against this object's own generation, which the hub issued - // and can therefore compare. - // +listType=map - // +listMapKey=type - // +optional - Conditions []metav1.Condition `json:"conditions,omitempty"` -} -``` - -```go -// +kubebuilder:object:root=true -// +kubebuilder:subresource:status -// +kubebuilder:resource:scope=Namespaced,shortName=dd -// +kubebuilder:printcolumn:name="Member",type=string,JSONPath=".spec.memberRef" -// +kubebuilder:printcolumn:name="Phase",type=string,JSONPath=".status.phase" -// +kubebuilder:printcolumn:name="Desired",type=string,JSONPath=".spec.driver.image" -// +kubebuilder:printcolumn:name="Running",type=string,JSONPath=".status.runningVersion" -// +kubebuilder:printcolumn:name="Age",type=date,JSONPath=".metadata.creationTimestamp" - -// DriverDeployment is the CSI driver one member should run. It is one object -// per member, so a fleet-wide rollout is a set of them. -type DriverDeployment struct { - metav1.TypeMeta `json:",inline"` - metav1.ObjectMeta `json:"metadata,omitempty"` - - Spec DriverDeploymentSpec `json:"spec,omitempty"` - Status DriverDeploymentStatus `json:"status,omitempty"` -} - -// +kubebuilder:object:root=true - -// DriverDeploymentList is a list of DriverDeployment. -type DriverDeploymentList struct { - metav1.TypeMeta `json:",inline"` - metav1.ListMeta `json:"metadata,omitempty"` - Items []DriverDeployment `json:"items"` -} - -// DriverDeploymentSpec is the driver payload for one member. There is no -// approval field: a driver version is an edit rather than a deployment gate. -type DriverDeploymentSpec struct { - // MemberRef names the FleetMember this driver is installed in. - // +kubebuilder:validation:Required - // +k8s:immutable - MemberRef string `json:"memberRef"` - - // Driver is the SimplyblockDriver spec that will be applied in the member. - // +kubebuilder:validation:Required - Driver storagev1alpha2.SimplyblockDriverSpec `json:"driver"` -} - -// DriverDeploymentStatus is what the member's driver reports back. -type DriverDeploymentStatus struct { - // Phase is where the installation is. - // +optional - Phase DriverDeploymentPhase `json:"phase,omitempty"` - - // RunningVersion is what the member's driver reports, fed back from the - // applied object rather than assumed from the image in the spec. - // +optional - RunningVersion string `json:"runningVersion,omitempty"` - - // Delivery is what the ManifestWork carrying this driver reports. - // +optional - Delivery *DeliveryStatus `json:"delivery,omitempty"` - - // Message is why the phase is what it is. - // +optional - Message string `json:"message,omitempty"` - - // ObservedGeneration is the generation this status was computed from. - // +optional - ObservedGeneration int64 `json:"observedGeneration,omitempty"` - - // +listType=map - // +listMapKey=type - // +optional - Conditions []metav1.Condition `json:"conditions,omitempty"` -} - -// DriverDeploymentPhase is where a driver installation is. -// +kubebuilder:validation:Enum=Delivering;Installing;Ready;Degraded;Failed -type DriverDeploymentPhase string - -const ( - DriverDeploymentPhaseDelivering DriverDeploymentPhase = "Delivering" - DriverDeploymentPhaseInstalling DriverDeploymentPhase = "Installing" - DriverDeploymentPhaseReady DriverDeploymentPhase = "Ready" - DriverDeploymentPhaseDegraded DriverDeploymentPhase = "Degraded" - DriverDeploymentPhaseFailed DriverDeploymentPhase = "Failed" -) -``` - -```go -// +kubebuilder:object:root=true -// +kubebuilder:subresource:status -// +kubebuilder:resource:scope=Namespaced,shortName=scd -// +kubebuilder:printcolumn:name="Member",type=string,JSONPath=".spec.memberRef" -// +kubebuilder:printcolumn:name="Class",type=string,JSONPath=".spec.className" -// +kubebuilder:printcolumn:name="Pool",type=string,JSONPath=".spec.pool.pool" -// +kubebuilder:printcolumn:name="Phase",type=string,JSONPath=".status.phase" -// +kubebuilder:printcolumn:name="Age",type=date,JSONPath=".metadata.creationTimestamp" - -// StorageClassDeployment is one StorageClass drawing on one pool in one member. -// A pool may have zero or more, because nothing about a pool implies a single -// way to consume it. -type StorageClassDeployment struct { - metav1.TypeMeta `json:",inline"` - metav1.ObjectMeta `json:"metadata,omitempty"` - - Spec StorageClassDeploymentSpec `json:"spec,omitempty"` - Status StorageClassDeploymentStatus `json:"status,omitempty"` -} - -// +kubebuilder:object:root=true - -// StorageClassDeploymentList is a list of StorageClassDeployment. -type StorageClassDeploymentList struct { - metav1.TypeMeta `json:",inline"` - metav1.ListMeta `json:"metadata,omitempty"` - Items []StorageClassDeployment `json:"items"` -} - -// StorageClassDeploymentSpec is the class to write in the member, and the pool -// it draws on. -type StorageClassDeploymentSpec struct { - // MemberRef names the FleetMember the class is written in. - // +kubebuilder:validation:Required - // +k8s:immutable - MemberRef string `json:"memberRef"` - - // ClassName is the StorageClass name in the member. It is immutable because - // it is the object's name there, and a rename is a different class. - // +kubebuilder:validation:Pattern=`^[a-z0-9]([-a-z0-9]*[a-z0-9])?$` - // +kubebuilder:validation:MaxLength=253 - // +kubebuilder:validation:Required - // +k8s:immutable - ClassName string `json:"className"` - - // Pool is the pool this class draws on. The three fields become the three - // assignment labels on the class, so an author states the pool and never - // writes a label by hand. - // +kubebuilder:validation:Required - // +k8s:immutable - Pool StorageClassPoolRef `json:"pool"` - - // Template is what the class carries beyond its assignment. - // +kubebuilder:validation:Required - Template StorageClassTemplate `json:"template"` -} - -// StorageClassPoolRef locates a pool in the member. A class may draw on a pool -// in any namespace, which is why all three parts are stated rather than -// inherited from wherever the class happens to be written. -type StorageClassPoolRef struct { - // Namespace is the pool's namespace in the member. - // +kubebuilder:validation:Required - Namespace string `json:"namespace"` - - // Cluster is the StorageCluster the pool belongs to. - // +kubebuilder:validation:Required - Cluster string `json:"cluster"` - - // Pool is the StoragePool's name. - // +kubebuilder:validation:Required - Pool string `json:"pool"` -} - -// StorageClassTemplate is the class body. Everything here is editable, which is -// what makes re-tuning a ceiling an edit rather than a delete and a re-create. -type StorageClassTemplate struct { - // Parameters are the driver parameters, written in the superseding QoS - // spelling only, which is what the operator's own generated class does. The - // three assignment labels are not parameters and are not written here. - // +optional - Parameters map[string]string `json:"parameters,omitempty"` - - // ReclaimPolicy is what happens to a volume when its claim goes. - // +kubebuilder:validation:Enum=Delete;Retain - // +optional - ReclaimPolicy *corev1.PersistentVolumeReclaimPolicy `json:"reclaimPolicy,omitempty"` - - // VolumeBindingMode is when a volume is provisioned. - // +kubebuilder:validation:Enum=Immediate;WaitForFirstConsumer - // +optional - VolumeBindingMode *storagev1.VolumeBindingMode `json:"volumeBindingMode,omitempty"` - - // DisableVolumeExpansion turns off online expansion. The negative spelling - // is what makes the zero value the default, and every class this product - // writes today allows expansion. - // +optional - DisableVolumeExpansion bool `json:"disableVolumeExpansion,omitempty"` - - // MountOptions are passed to the mount of a volume of this class. - // +optional - MountOptions []string `json:"mountOptions,omitempty"` -} - -// StorageClassDeploymentStatus is whether the class reached the member. -type StorageClassDeploymentStatus struct { - // Phase is where the class is. - // +optional - Phase StorageClassDeploymentPhase `json:"phase,omitempty"` - - // Delivery is what the ManifestWork carrying this class reports. - // +optional - Delivery *DeliveryStatus `json:"delivery,omitempty"` - - // Message is why the phase is what it is. - // +optional - Message string `json:"message,omitempty"` - - // ObservedGeneration is the generation this status was computed from. - // +optional - ObservedGeneration int64 `json:"observedGeneration,omitempty"` - - // +listType=map - // +listMapKey=type - // +optional - Conditions []metav1.Condition `json:"conditions,omitempty"` -} - -// StorageClassDeploymentPhase is where a class is. PoolMissing is its own value -// because a class naming a pool that is not there is a mistake somebody fixes, -// and not a delivery that failed. -// +kubebuilder:validation:Enum=Delivering;Ready;PoolMissing;Degraded;Failed -type StorageClassDeploymentPhase string - -const ( - StorageClassDeploymentPhaseDelivering StorageClassDeploymentPhase = "Delivering" - StorageClassDeploymentPhaseReady StorageClassDeploymentPhase = "Ready" - StorageClassDeploymentPhasePoolMissing StorageClassDeploymentPhase = "PoolMissing" - StorageClassDeploymentPhaseDegraded StorageClassDeploymentPhase = "Degraded" - StorageClassDeploymentPhaseFailed StorageClassDeploymentPhase = "Failed" -) -``` - -```go -// +kubebuilder:object:root=true -// +kubebuilder:subresource:status -// +kubebuilder:resource:scope=Namespaced,shortName=fop -// +kubebuilder:printcolumn:name="Member",type=string,JSONPath=".spec.memberRef" -// +kubebuilder:printcolumn:name="Kind",type=string,JSONPath=".status.target.kind" -// +kubebuilder:printcolumn:name="Phase",type=string,JSONPath=".status.phase" -// +kubebuilder:printcolumn:name="Remote",type=string,JSONPath=".status.remote.phase" -// +kubebuilder:printcolumn:name="Step",type=string,JSONPath=".status.remote.step" -// +kubebuilder:printcolumn:name="Age",type=date,JSONPath=".metadata.creationTimestamp" - -// FleetOperation ships one Band A Ops object into one member and reports what it -// does there. One kind rather than one per managed Ops kind, because the hub's -// question is which member and which operation, and the parameters belong to the -// kind whose spec it carries. -type FleetOperation struct { - metav1.TypeMeta `json:",inline"` - metav1.ObjectMeta `json:"metadata,omitempty"` - - Spec FleetOperationSpec `json:"spec,omitempty"` - Status FleetOperationStatus `json:"status,omitempty"` -} - -// +kubebuilder:object:root=true - -// FleetOperationList is a list of FleetOperation. -type FleetOperationList struct { - metav1.TypeMeta `json:",inline"` - metav1.ListMeta `json:"metadata,omitempty"` - Items []FleetOperation `json:"items"` -} - -// FleetOperationSpec is a request, and every field but Abort is immutable. -type FleetOperationSpec struct { - // MemberRef names the FleetMember the operation runs in. - // +kubebuilder:validation:Required - // +k8s:immutable - MemberRef string `json:"memberRef"` - - // Operation is what to run there. Exactly one member is set. - // +kubebuilder:validation:Required - // +k8s:immutable - Operation FleetOperationTarget `json:"operation"` - - // Abort asks the member's own object to stop, by setting spec.abort on it - // through the payload. It is the one mutable field on this spec, and the - // member's graph decides whether the step it is on accepts it. - // +optional - Abort bool `json:"abort,omitempty"` -} - -// FleetOperationTarget carries the spec of exactly one Band A Ops kind. The -// discriminated block is the shape ControlPlane.spec.source uses, and it is -// typed rather than opaque because this is the payload that performs a side -// effect. -// +kubebuilder:validation:XValidation:rule="[has(self.operatorOps),has(self.storageClusterOps),has(self.storageNodeOps),has(self.storageDeviceOps),has(self.storagePoolOps),has(self.storageBackupOps)].filter(x, x).size() == 1",message="set exactly one operation" -type FleetOperationTarget struct { - // OperatorOps runs an operator-level operation, today a discovery run. - // +optional - OperatorOps *storagev1alpha2.OperatorOpsSpec `json:"operatorOps,omitempty"` - - // StorageClusterOps runs a cluster-level operation. - // +optional - StorageClusterOps *storagev1alpha2.StorageClusterOpsSpec `json:"storageClusterOps,omitempty"` - - // StorageNodeOps runs a node-level operation. - // +optional - StorageNodeOps *storagev1alpha2.StorageNodeOpsSpec `json:"storageNodeOps,omitempty"` - - // StorageDeviceOps runs a device-level operation. - // +optional - StorageDeviceOps *storagev1alpha2.StorageDeviceOpsSpec `json:"storageDeviceOps,omitempty"` - - // StoragePoolOps runs a pool-level operation. - // +optional - StoragePoolOps *storagev1alpha2.StoragePoolOpsSpec `json:"storagePoolOps,omitempty"` - - // StorageBackupOps runs a backup or restore. - // +optional - StorageBackupOps *storagev1alpha2.StorageBackupOpsSpec `json:"storageBackupOps,omitempty"` -} - -// FleetOperationStatus is the hub's view of an operation running in a member. -type FleetOperationStatus struct { - // Phase is the hub's own progress, which is about delivery rather than about - // the operation. What the operation is doing is in Remote. - // +optional - Phase FleetOperationPhase `json:"phase,omitempty"` - - // Target names the object the payload created in the member, so that a - // ManagedClusterView for the whole object needs no re-derivation. - // +optional - Target *FleetOperationTargetRef `json:"target,omitempty"` - - // Remote is what the member's own object reports, filled from the status - // feedback rules on the ManifestWork. Scalars only, which is what a feedback - // rule carries and what a list view needs. Anything beyond it is a - // ManagedClusterView against Target. - // +optional - Remote *RemoteOpsStatus `json:"remote,omitempty"` - - // Delivery is what the ManifestWork carrying the payload reports. - // +optional - Delivery *DeliveryStatus `json:"delivery,omitempty"` - - // Message is why the phase is what it is. - // +optional - Message string `json:"message,omitempty"` - - // StartedAt is when the payload was delivered. - // +optional - StartedAt *metav1.Time `json:"startedAt,omitempty"` - - // CompletedAt is when the operation reached a terminal remote phase, and is - // what retention measures against. - // +optional - CompletedAt *metav1.Time `json:"completedAt,omitempty"` - - // ObservedGeneration is the generation this status was computed from. It - // advances at most twice, and the second advance is the signal that Abort - // was observed. - // +optional - ObservedGeneration int64 `json:"observedGeneration,omitempty"` - - // +listType=map - // +listMapKey=type - // +optional - Conditions []metav1.Condition `json:"conditions,omitempty"` -} - -// FleetOperationPhase is the hub's progress at getting the operation to run. -// +kubebuilder:validation:Enum=Pending;Delivering;Running;Succeeded;Failed;Aborted -type FleetOperationPhase string - -const ( - FleetOperationPhasePending FleetOperationPhase = "Pending" - FleetOperationPhaseDelivering FleetOperationPhase = "Delivering" - FleetOperationPhaseRunning FleetOperationPhase = "Running" - FleetOperationPhaseSucceeded FleetOperationPhase = "Succeeded" - FleetOperationPhaseFailed FleetOperationPhase = "Failed" - FleetOperationPhaseAborted FleetOperationPhase = "Aborted" -) - -// FleetOperationTargetRef locates the object the payload created. -type FleetOperationTargetRef struct { - // Kind is the Band A kind that was created. - // +optional - Kind string `json:"kind,omitempty"` - - // Namespace and Name locate it in the member. - // +optional - Namespace string `json:"namespace,omitempty"` - // +optional - Name string `json:"name,omitempty"` -} - -// RemoteOpsStatus is what a feedback rule can carry off a Band A Ops object. -// Every field is a scalar, because a JSON path feedback rule selects one value. -type RemoteOpsStatus struct { - // Phase is the member object's own phase, in its own spelling. - // +optional - Phase string `json:"phase,omitempty"` - - // Step is the state name of the member object's step machine. The deadline - // beside it in the member is an absolute instant and is deliberately not - // carried, because it is meaningless against the hub's clock. - // +optional - Step string `json:"step,omitempty"` - - // Message is the member object's status message. - // +optional - Message string `json:"message,omitempty"` - - // ObservedGeneration is the member object's, and is comparable only against - // generations in the member. - // +optional - ObservedGeneration int64 `json:"observedGeneration,omitempty"` - - // ObservedAt is when the feedback last arrived. - // +optional - ObservedAt *metav1.Time `json:"observedAt,omitempty"` -} -``` diff --git a/discovery-generator-rules.md b/discovery-generator-rules.md deleted file mode 100644 index 00358edb2..000000000 --- a/discovery-generator-rules.md +++ /dev/null @@ -1,757 +0,0 @@ - - -# The rules the discovery generator applies - -**What this is.** Every rule between a fleet's probe reports and the drafted -`ClusterDeploymentConfig`, with its precedence, its exact condition, and its -provenance. - -**Why it exists.** The 155 recorded expectations under -`operator/internal/controllers/deployment/testdata/discovery/` were produced by -running the generator over the fixtures. That proves the rules were applied -consistently and proves nothing about whether they are the right rules. Several -were never specified anywhere: they were settled in code, and the expectations -now carry them as though they were agreed. This is the list to ratify, correct, -or overrule. - -**How to read it.** Every rule carries a provenance marker, and the markers are -the point. - -| Marker | Meaning | -|-------------|--------------------------------------------------------------------------------------------------------------------------------------------------------------| -| `SPECIFIED` | A design document states the rule, and the code implements what was written | -| `GROUNDED` | No design states it, and the code cites a constraint outside itself: a control-plane refusal, an API server limit, an addressing requirement, or determinism | -| `ASSERTED` | No design states it, and the reason the code gives is a preference rather than a constraint | -| `INVENTED` | No design states it and no reason is given. It was settled while the code was written | -| `UNBUILT` | A design states the rule and the code does not implement it | - -A `SPECIFIED` or `GROUNDED` rule needs checking for accuracy. An `ASSERTED` or -`INVENTED` rule needs a decision. §10 collects the second kind. - -## 1. The pipeline, in order - -A run is three steps, and only the third produces a document. - -1. **Inspecting** settles which workers the run is about and which distribution - the cluster runs. Nothing after it re-derives either. -2. **Probing** starts one Job per worker, each writing a report into a - ConfigMap. The probe decides nothing: it reports every block device, refused - ones included, with the grounds for each. -3. **Writing** turns the reports into a document. - -Writing is itself ordered, and the order is load-bearing. - -``` -reports, sorted by worker name - └─ per worker: - 1. synthesize a device for every idle userspace NVMe controller §4.4 - 2. apply the device rules, first refusal wins §4.1 - 3. apply the worker rules, first refusal wins §4.3 - 4. choose which memory node to place on §5 - 5. record every admitted device the placement did not take §5.5 - 6. choose the management interface §7 - └─ over the workers that survived: - 7. group them by what they hand over §6.1 - 8. split the groups into node sets by role §6.4 - └─ over the plan: - 9. derive the cluster block §8 - 10. assemble, name, and write the document §9 -``` - -Devices are filtered before workers because whether a worker is worth including -is mostly a question about what is left of it. `GROUNDED`. - -## 2. Which workers the run is about - -Decided once, in Inspecting, and persisted. `GROUNDED`: a worker that joins the -cluster while the probes run is not in this run's draft, and re-listing in -Probing would put it there with no report to describe it. - -A node is a candidate when all four hold: - -| # | Condition | Provenance | -|-----|------------------------------------------------------------------|-------------| -| 1 | It matches `spec.discover.nodeSelector`, when one is set | `SPECIFIED` | -| 2 | No `StorageNode` in the namespace already names it | `SPECIFIED` | -| 3 | It is schedulable | `GROUNDED` | -| 4 | Its role holds storage nodes, or the control-plane opt-in is set | `GROUNDED` | - -**Schedulable** means not cordoned and carrying no `NoSchedule` or `NoExecute` -taint. `PreferNoSchedule` does not exclude. `GROUNDED`: a probe Job is pinned -with `spec.nodeName` and would run on either, and a worker the cluster is not -scheduling to is not one to hand to a storage cluster. - -**Role** is read from `node-role.kubernetes.io/` labels, and the most -restrictive wins: control-plane, then infra, then worker. Both `master` and -`control-plane` are read, neither preferred. An unrecognized role is ignored -rather than refused. Infra nodes are used **without** an opt-in, control-plane -nodes only with one. `GROUNDED`, and the asymmetry is argued: an OpenShift infra -node is the tier a cluster's own infrastructure runs on, and simplyblock storage -is infrastructure. - -A declined worker gets an event naming the reason. A node already carrying a -`StorageNode` is the one exclusion that stays **silent**, against the adjacent -comment claiming both exclusions are otherwise explained. - -**Environment** is concluded once from the whole node list, not per probe. -`SPECIFIED`. Signals, strongest first: registered API groups, then the kubelet -version string, then the OS image prefix, then node labels and annotations. When -several distributions match, precedence is OpenShift, Talos, K3s, Rancher. -`GROUNDED`: a Rancher-managed K3s cluster carries both sets of markers, and what -decides the host's shape is the distribution that installed the kubelet. Nothing -matching yields `Vanilla`. A failure to list API groups is an event, not a -failure, since half the evidence still yields a conclusion. -## 3. What the probe decides, and what it refuses to decide - -The probe decides nothing about what a cluster should use. `SPECIFIED`, and the -rule is load-bearing for everything below it: every block device is reported, -refused ones included, each with the grounds it was refused on, because the run's -device filter belongs to the operator that holds the fleet's reports and an -administrator has to be able to be told why the disk they expected is not a -candidate. - -Three consequences the generator depends on: - -- **The probe accumulates grounds, the generator stops at the first.** A device - arrives carrying every reason it is unusable, and is refused for one of them. - That is what makes the partition waiver safe: it admits a device whose **only** - ground is a partition table, so a boot disk stays out on its mount. -- **The report is a wire format with a version, and a reader refuses a version it - does not know.** `GROUNDED`: a Job keeps the image it started with, so an - operator upgraded mid-run reads reports from the previous probe, and a field - that changed meaning would be read as the current one. -- **The report is the only place a node's name survives exactly.** Object names - and labels are sanitized and truncated to fit a 63-byte budget, so the node a - report is about is the field inside it. - -What the probe reads and the generator never consults is itself a list worth -reviewing: per-node memory, online CPUs, swap, MTU, MAC address, driver, PCI -address of an interface, huge-page totals as distinct from free pages, and every -huge-page reading the kubelet publishes. None of them enters any rule. - -## 4. Which devices reach the draft - -Four stages, and a device dropped at one never reaches the next. Two of them are -the probe's and two are the generator's. - -| Stage | Owner | Decides | Short-circuits | -|-------|-----------|---------------------------------------------------|-------------------------| -| A | probe | Which sysfs entries exist, and their kind and bus | no | -| B | probe | Every ground a device is refused on | accumulates all grounds | -| C | generator | Admit or refuse each device | **first refusal wins** | -| D | generator | Admit or refuse the whole worker | **first refusal wins** | - -The probe accumulates every ground, and the generator stops at the first. A device -therefore carries all the reasons it is unusable and is refused for one of them. - -### 4.1 The device rules, in evaluation order - -| # | Rule | Present when | Pre-filter | Provenance | -|-----|----------------------|-----------------------------------------------|------------|------------| -| 1 | Whole disk | always | **yes** | `GROUNDED` | -| 2 | Device class | always | **yes** | `GROUNDED` | -| 3 | Available | always | no | `GROUNDED` | -| 4 | Allow and deny lists | the scanned class's allow or deny list is set | no | `ASSERTED` | -| 5 | Model | NVMe run and `pcieModel` is set | no | `GROUNDED` | -| 6 | Size range | `driveSizeRange` is set **and parses** | no | `ASSERTED` | - -The order is stated and justified: what the device is, then whether it is free, -then whether this run wants it. "A partition refused as 'not in the allow list' -would be a true statement and the wrong one." - -**Pre-filter** marks a rule whose refusal says nothing about the fleet's -storage. Pre-filtered refusals are emitted as events and are **excluded** from -the per-worker explanation the run's status carries. Only rules 1 and 2 set it. - -### 4.2 Each rule's exact condition - -**1. Whole disk**: refuses unless `kind == "Disk"`, exact and case-sensitive. -An empty kind refuses. `GROUNDED`: it is redundant against the probe, and it is -the one rule whose absence would be silent, because a partition admitted by a -waiver would be handed to a cluster as a disk. - -**2. Device class**: symmetric, and it has to be, because a cluster is built out -of one class. - -- On an **NVMe** run: refuses unless `transport == "NVMe"` exactly, then refuses - unless the PCI address is non-empty. -- On a **block** run: refuses `NVMe` and `NVMeFabric` alike, then refuses unless - the path is non-empty. - -The block side used to check only the path, so an NVMe disk reached a block draft -by its path and one document could name both classes. Naming devices by path is -what the block class does rather than what makes a device one, so the bus is what -both sides read. - -**2a. Simplyblock volume**: refuses a device whose subsystem NQN parses as a -simplyblock logical-volume subsystem, naming the cluster it belongs to. -`GROUNDED`: a volume this product exported and a worker attached is a namespace -like any other, and handing one back to a cluster as backend storage would give -a volume's own bytes away as free space. - -It sits ahead of the class rule deliberately. The class rule refuses every -fabric namespace too, but as a pre-filter, so its answer never reaches the -explanation: a reviewer asking why a machine full of disks proposed none would be -told the disks are on a bus the run does not scan, rather than that they are the -fleet's own volumes. It is also the check that survives the class rule being -relaxed, which is the one way a cluster could be told to take its own bytes. - -Parsing rather than matching a prefix is what makes it specific. A subsystem of -this product that is not a logical volume does not read as one, and a namespace -another product exported falls through to the class rule. - -**3. Available**: admits a device the probe found free. One waiver, and it is -narrow: with `enablePartitionedDevices`, a device whose **only** ground is a -partition table is admitted. A device also mounted, held, or unreadable is not. -`GROUNDED`, and the narrowness is enforced on the probe's side too. - -The refusal quotes the ground *names* joined by commas, not the details. A -device marked unavailable with no grounds at all, which is inconsistent data the -probe does not write, reads "the probe did not report it as available and gave no -reason." - -**5. iSCSI**: refuses an iSCSI LUN the run's allow list does not name. -`GROUNDED`: every other bus a run scans is a cable inside the chassis, and a LUN -is a disk on the other side of a network, so a cluster built on one runs every -write of its data path over that network. Whether that is wanted is a question -about the deployment rather than about the hardware. - -The default is refusal rather than proposal-and-review, because a draft -proposing a LUN is one a reviewer has to notice and strike, and a fifty-worker -document is not one anybody reads that closely. Naming it in the allow list is -the decision. - -It reads the allow list itself rather than leaving it to the rule below, because -that rule is only built when a list was given: with no filter there would be -nothing to refuse a LUN, and with a list given for another reason a LUN would be -admitted by being named alongside everything else. - -It also sits ahead of that rule so that an unnamed LUN says it is a LUN rather -than that it is not in the allow list, which is the same true-but-wrong-answer -the whole-disk ordering exists to avoid. - -**6. Allow and deny lists**: deny is evaluated first and wins, and an **empty allow -list means allow everything** rather than allow nothing. Matching is case-folded. - -Two things asserted and not justified: deny-before-allow, and the case folding. -The folding is defensible for a PCI address and is **wrong for a block path**: -`/dev/SDB` matches `/dev/sdb`, and Linux paths are case-sensitive. - -**7. Model**: case-folded **substring** match. `GROUNDED`: a model string is -padded, versioned, and vendor-formatted, so `MZQL2` is what an administrator -writes. A device with an empty model is refused by any non-empty filter. - -**8. Size range**: both bounds are **inclusive**, and a zero maximum means **no upper -bound**. Applies to whichever class is scanned. - -### 4.3 The worker rules - -**Has devices**: refuses a worker with nothing admitted, and says one of three -things. `GROUNDED` throughout, with the distinction spelled out: a machine whose -disks are on a userspace driver has no block devices at all, so "no device -survived the rules" is true and useless. - -| Situation | What the reviewer is told | -|-----------------------------------------------------|-------------------------------------------| -| Userspace-bound controllers, something driving them | something is driving its disks | -| Userspace-bound controllers, nothing driving them | the disks are there to be reclaimed | -| Neither | no device of it survived the device rules | - -**Fully readable**: refuses a worker whose probe could not read everything. Off -by default, and **nothing in the operator ever turns it on**. `GROUNDED`: a -worker whose CPU tree could not be read still has disks worth reviewing. - -### 4.4 Devices the generator invents - -An NVMe controller bound to a userspace driver presents no block device, so it -cannot reach a draft through the device reading at all. The generator synthesizes -one for each such controller that **nothing is using** and whose address no -reported device already carries. `GROUNDED`, and the trade is stated: everything -the disk would have said about itself is lost, so it reaches the draft unsized -and uninspected, and the group it lands in is named for a count rather than a -capacity. - -**Two interactions worth deciding on**, neither documented: - -- A **model filter refuses every synthesized controller**, because its model is - empty and the match is a substring. -- A **size range with any lower bound refuses every synthesized controller**, - because its size is zero. - -So a run that filters by model or by size silently excludes exactly the disks the -synthesis exists to offer. - -### 4.5 Findings - -**F-4.1. An unparsable size range silently admits everything, and the code says -otherwise.** `plan.go:281-283` states: "ParseSizeRange is called by the caller -that validates the spec, and this one skips what it cannot read rather than -silently widening the filter, and the run reports the parse failure separately." - -No such caller exists. `ParseSizeRange` is called from exactly one place outside -its own tests. `driveSizeRange` carries no schema pattern and no CEL rule. So an -unparsable range drops the rule entirely, every size passes, and **nothing -reports it**: not a refusal, not an event, not the summary. This is `G-9`, and -it is worse than recorded: the comment asserts a safeguard that was never built. - -**F-4.2. Eleven places drop a device or a worker with no record.** The ones that -matter: - -- A userspace controller that is **in use** is skipped silently. The only trace - is the worker-level message, and only if the worker ends up with nothing. -- A worker whose report cannot be decoded raises an event but produces **no - refusal**, so it is absent from the explanation the status carries. -- A worker whose report names a node the run is not about is dropped in silence. -- ~~`blockAllowList` and `blockDenyList` are silently ignored when the - planner's class and the filter disagree.~~ **Fixed.** The planner carried a - class field beside the filter's own `enableLogicalBlockDevices`, so one fact - had two statements and they could disagree: a planner told nothing scanned - NVMe, read the PCI lists, and dropped the block lists on a branch that never - ran. The field is gone, and the class is read from the filter, where it was - always written. - -**F-4.3. `InUse: false` is not distinguished from "never read."** The report's -own comment says a reader deciding whether to reclaim a controller "has to find -this report free of such an entry [in `unreadable`] first." Nothing performs that -check, and the rule that would, Fully readable, is off by default. A probe that -could not read the process table therefore yields controllers that look idle. - -**F-4.4. Fixed: the explanation counts by rule.** It used to group on the -rendered sentence, and the allow-and-deny refusal embedded the address in its -sentence, so a hundred declined devices produced a hundred clauses. It now -groups on the rule the refusal already carries, and the allow-and-deny reason -names the list rather than the device, which the refusal holds separately. Where -one rule's reasons genuinely differ, a mounted disk against a partitioned one, -each is counted rather than dropped. - -**F-4.5. Fixed: the bound is quoted as written, and a size is never rounded -up.** The bound used to be re-rendered from the parsed number, so a filter -naming `1920G` was quoted back as `1.875T`. The rule now carries the range as -the filter wrote it. - -The renderer was worse than lossy. Four significant digits rounded, so a byte -under a tebibyte printed as `1024G`, which reads as more than the value, and its -own parser refuses every decimal it produced. A whole number of units is now -written exactly and parses back to the byte it came from, and anything else is -truncated to two decimals. -## 5. Placement: which part of a worker is used - -A worker's admitted devices are grouped by the memory node they hang off, and -one of those groups is taken. The rest are left behind and recorded as refused. - -**The buckets exist only where an admitted device is.** A memory node with 64 -cores, 128 GiB of huge pages, and a 100 GbE port but no admitted device does not -appear in the ranking at all, and cannot be chosen. The CPU, huge-page, and NIC -readings are attached to buckets that already exist and are otherwise discarded. -`placement.go:139-171`. `INVENTED`, asserted by construction with no comment. - -### 5.1 The ranking, in order - -Applied only when two or more buckets exist. Every key is compared in turn until -one of them separates the two nodes. - -| # | Key | Direction | Provenance | Stated reason | -|-----|-------------------------------------------|------------|------------|---------------------------------------------------------------------------------------------------------------------------| -| 1 | Admitted device **count** | descending | `GROUNDED` | Usable space is bounded by the erasure-coding stripe, which is a count of devices, so four small disks beat one large one | -| 2 | Combined **capacity** in bytes | descending | `ASSERTED` | "Capacity breaks a tie in the count" | -| 3 | **Physical cores** of the node | descending | `ASSERTED` | "Cores break a tie in capacity" | -| 4 | A real node before the **unknown bucket** | real first | `GROUNDED` | The bucket is numbered -1, so comparing ids alone would rank it above node 0 | -| 5 | **Node id** | ascending | `GROUNDED` | Two runs against one worker have to choose the same node | - -`placement.go:101-120`. - -Key 4 is the one worth reading twice. The unknown bucket's id is `-1`, and key 5 -is ascending, so removing key 4 would make "no memory node in particular" win -every full tie against node 0. It is not merely theoretical: a worker whose CPU -topology could not be read has zero cores against every bucket, so keys 1 to 3 -can all tie, and the comparison falls through to it. - -**Two questions this ranking does not ask**, both recorded in the code as known -and both with a cost: - -- Where the data NIC is. A node with four disks and no fast NIC is preferred - over one with three and a 100 GbE port. -- Whether huge pages are reserved on the node. A node with the disks and no - reserved memory cannot start SPDK at all, and this ranks it first. - -### 5.2 What is read and not ranked on - -| Reading | Carried as | In the key? | -|--------------------------|---------------------------------------|-------------| -| Online CPUs per node | `NUMANodeResources.OnlineCPUs` | No | -| Free huge pages per node | `NUMANodeResources.FreeHugePageBytes` | No | -| Fastest NIC per node | `NUMANodeResources.FastestNICMbps` | No | -| Memory per node | nothing, the struct has no field | No | - -The per-node memory reading is collected by the probe for a stated purpose, that a -storage node pinned to a socket draws its memory from that socket, and the -placement never consults it. - -### 5.3 The NIC reading, and what a bond does to it - -The fastest NIC per node is computed from the **raw** `speedMbps` against the -**raw** `numaNode` of every interface in the report, with no stack traversal and -no kind filter. `placement.go:167-171`. - -A bond, a team, a VLAN, a macvlan, a bridge, and a veth all sit under -`devices/virtual`, and the interface reader returns early for those before it -reads a memory node, so every one of them arrives carrying `numaNode: -1`. The -consequence is exact: an aggregate's speed is credited to the unknown bucket, -and only when that bucket happens to exist because some admitted device also -reported no memory node. Otherwise, the reading is dropped. - -No real memory node is ever credited with a bonded or tagged link as such. What -saves the reading in practice is that the report carries the bond's members too, -each with a real node and a real speed, so the members are counted individually. - -This is the one place the stack resolution added for the management interface -was not applied: `netstack.go` resolves an aggregate through its members and -reports a memory node only where they agree, and `NUMANodeBreakdown` does none -of it. - -### 5.4 The sentence the placement produces - -Every chosen worker carries a sentence saying what was chosen and why, and every -device left behind is recorded as refused with that same sentence as its reason. - -Three defects in it, all observable in the recorded expectations: - -1. It always says "devices" plural, so a one-device winner reads `carries 1 - unclaimed devices`. -2. It names only the runner-up. A three-node machine's third node is never - mentioned. -3. **It asserts the count and capacity comparison even when neither decided.** A - tie settled by key 3, 4, or 5 renders as `NUMA node 0 carries 2 unclaimed - devices (2T) against 2 (2T) on no memory node in particular, so it was - chosen`, which is equal numbers either side of a "so." - -`Worker.PlacementReason` holds the sentence and **nothing in the repository -reads it**. Its only path to a reviewer is as the reason on the per-device -refusals, which become `DeviceDeclined` events. - -### 5.5 Leftover devices - -A device that survived every rule and was not taken by the placement is recorded -as a refusal whose rule is the placement's name and whose reason is the whole -winner sentence, identical on every leftover device of that worker. - -Membership is tested by **device name equality**, not by identity or address. - -These refusals are not pre-filters, so they would reach the per-worker -explanation, but the explanation is only assembled when the run produced no -node sets at all, and a worker with leftovers is by construction in the plan. So -they surface as events, inflate the refusal count in the run's summary, and do -not count toward the device count. - -**A knock-on worth checking:** leftover devices are absent from the worker's -address list, which is what the grouping signature is computed over. Two -identical machines placed on different memory nodes therefore hand over -different address sets and land in different groups. - -### 5.6 The alternative placement - -`AllDevices` takes every admitted device regardless of memory node. It is not -used by the run, because the controller always builds the default, and it is a seam -for a fleet whose machines have one memory node. - -Its downstream effect is worth stating because it is not local: a worker whose -devices span nodes has no single chosen node, which makes the cluster block's -vCPU count fall to the API floor and its huge-page size go unset for the whole -fleet, and the note the reviewer reads then says the worker "has no huge pages -reserved on the memory node it was placed on," which is not what happened. -## 6. Grouping and node sets - -### 6.1 What makes two workers one group - -A group's device selection is shared by every worker in it, so grouping is a -claim about the machines rather than a presentation choice. Two workers share a -group when three things match exactly: - -1. The device class of the run. -2. The **name** of the management interface. -3. The sorted, deduplicated list of device addresses. - -Hashed together, in that order. `SPECIFIED` for grouping by identical hardware, -which the design calls a guess at intent that a reviewer regroups. The -management interface is in the key with a `GROUNDED` reason: a `NodeGroup` names -one interface for every worker it lists. - -**What is deliberately not in the key**, and what each omission costs: - -| Not in the key | Consequence | -|----------------------------------------------------|-------------------------------------------------------------------------------------------------------------| -| Device **sizes** | A 1.92 TB fleet and a 3.84 TB fleet at the same slots are one group, named for the first worker's capacity | -| Device **count**, as distinct from the address set | One controller with two namespaces and one with a single namespace group together | -| NUMA topology | No document field expresses it, so nothing is lost at the group level | -| Model, vendor, serial, rotational | A mixed-model fleet at identical slots is one group. Model can only exclude devices, never split groups | -| Kube role | Deliberate: the split by role happens after grouping, so identical infra machines stay one group | -| Memory, CPU, huge pages | Two machines with identical disks and very different RAM are one group, and `spdkSystemMemory` is never set | -| Interface kind, speed, members | Two workers whose `eth0` is a bare NIC on one and a bond on the other are one group | - -**Address handling.** Addresses are deduplicated (`GROUNDED`: a controller with -two namespaces reports two devices at one address) and sorted ascending. The sort -normalizes probe enumeration order, so two workers whose kernels enumerated the -same disks differently still group together. - -Case is **not** normalized. A probe reporting `0000:5E:00.0` and another -reporting `0000:5e:00.0` produce different signatures and therefore different -groups, while the allow-and-deny rule *does* fold case. `INVENTED`, and -inconsistent with the filter's treatment of the same string. - -### 6.2 Ordering and numbering - -Groups are ordered ascending by their first worker's name, and numbered from one -in that order. `GROUNDED`: two runs over one fleet produce the groups in the same -order and their generated names are stable. - -The comparison is lexicographic, so `worker-10` sorts before `worker-2`. A -stability limit the comment does not state: numbering is stable across re-runs -over an **unchanged** fleet only. Adding a machine that sorts earliest shifts -every later group's number. - -### 6.3 The group name - -The pattern is `group---x`, for example -`group-1-nvme-4x3.492T`. `GROUNDED`: it is positional rather than derived from -the hardware, because the name is what a reviewer edits and `group-1` invites -that where a hash does not. - -The size is the **first worker's** total device bytes divided by the number of -deduplicated addresses. Consequences: - -| Fleet | Rendered | Note | -|--------------------------------------|--------------|-----------------------------------------------------------------------------------| -| 4 × 3.84 TB at four slots | `4x3.492T` | Binary divisor under an SI-looking letter | -| One controller, two 1 TiB namespaces | `1x2T` | States 2T for something the draft names once | -| One 2 TiB disk and one 1 TiB disk | `2x1.5T` | An average no disk has | -| Claimed userspace controllers | `8xunsized` | `GROUNDED`: naming it `0 B` would state a capacity where there is only an absence | -| 10 240 TiB | `1.024e+04T` | Exponent notation inside a group name | - -**A false claim in the code.** The comment on the size computation says it is -"the same for every worker in it by construction." The signature hashes addresses -and not sizes, so it is not. This is the one place in that file where a stated -rationale asserts more than the code guarantees. - -### 6.4 Node sets - -One node set per role, ordered infra, then workers, then control-plane. -`GROUNDED`: infra nodes are almost always the ones somebody meant to be the -storage, and on OpenShift they do not count against a subscription's core limit, -so proposing them first in a block of their own lets a reviewer take that -placement by deleting the other block. - -Names are `infra`, `discovered`, and `control-plane`. Splitting happens **after** -grouping, `GROUNDED`, so two infra nodes with identical disks stay one group. - -**A consequence of that order:** a hardware group whose members span two roles -produces two `NodeGroup`s **with the identical name** in two different node sets, -each still stating the address count and size computed before the split. Names -are unique within a set and not across the document. - -## 7. The management interface - -### 7.1 The ladder - -Applied per interface, in report order, which is ascending by name. - -| # | Rung | Provenance | -|-----|-------------------------------------------------------------------------|------------------------------------| -| 1 | The kind must be bindable | `INVENTED` | -| 2 | The state must be `up`, `unknown`, or empty | `ASSERTED` | -| 3 | A bridge or an overlay is refused unless it holds the node's address | `ASSERTED` | -| 4 | It must hold at least one reachable address | `GROUNDED` | -| 5 | The interface holding the node's `InternalIP` wins, and returns at once | `GROUNDED` | -| 6 | Otherwise: resolved speed descending, kind simplicity, then name | `ASSERTED`, `INVENTED`, `GROUNDED` | - -**Bindable** is exactly physical, bond, team, VLAN, VXLAN, macvlan, ipvlan, and -bridge. Loopback, an unidentified virtual device, and any unrecognized kind are -refused. Refusing loopback and unidentified devices is argued. Admitting the rest is an -enumeration with no stated reason. - -**Reachable** refuses an unparsable address, link-local unicast and multicast, -loopback, and the unspecified address. `GROUNDED`: the control plane refuses a -node whose management interface it cannot find an IP on, and it refuses it inside -the `node_add` task rather than at the request. - -**Simplicity tiers** are physical 0, bond and team 1, VLAN and macvlan and ipvlan -2, everything else 3. The *role* of the key is argued, the tiers themselves are -not, and the stated metric ("the fewest layers between the address and the wire") -does not produce them: a VLAN over a bond is two layers and a bare VLAN is one, -and both are tier 2. Tier 3 is unreachable given the bindable set. - -### 7.2 Resolving through the stack - -An aggregate reports no speed, no slot, and no memory node of its own, so those -are resolved downward through its members. - -- **Speed:** an interface's own reading wins if it is above zero, and the walk - stops there. Otherwise, an aggregate sums its members and a derived interface - takes its first parent with a speed. `GROUNDED` for the split, `ASSERTED` for - the self-report short circuit. -- **Members:** the physical interfaces reachable downward, sorted and - deduplicated. Empty for an interface that is itself physical. -- **Memory node:** the node every member agrees on, and unknown when they do not. - `GROUNDED`: a bond whose members are in two sockets has no affinity, and - reporting one of them would claim an affinity the interface does not have. - -### 7.3 The interaction that matters most - -**The eligibility gate runs before the node-address win.** An interface holding -the cluster's own address is skipped entirely, not even kept as a candidate, when -it is down, of an unidentified kind, or holding no otherwise-reachable address. - -The adjacent comment states the opposite: "when it is known the interface holding -it wins outright. That is not a preference among equals." `INVENTED`, and no test -covers it. - -What happens instead: the draft names some other interface on a different -network, with a reason reading "it is the fastest interface holding a reachable -address," which does not mention that the cluster's own address was elsewhere. -That is precisely the outcome the rule was written to prevent. - -### 7.4 What the choice records, and what is read - -The chosen interface carries its kind, its members, its resolved speed, its -memory node, and a sentence explaining the choice. **Only the name is ever read** -downstream, for the group signature and the document's `mgmtInterface`. - -The sentence's stated purpose is "the record a reviewer reads," and it is -rendered in no template, event, or status. In the one case a reviewer most needs -explained, a worker with no serving interface, the sentence is **empty** and no -refusal is recorded. -## 8. The cluster block - -Two of the cluster's fields are required by the API and cannot be read off a -worker. The file that derives them states its own standard: nothing there has to -be right, it has to be plausible, stated, and easy to correct. - -| Field | Value | Provenance | -|------------------------|----------------------------------------------------------------|------------| -| `name` | the document's name plus `-cluster` | `ASSERTED` | -| `maxSubsystemCount` | always 30 | `GROUNDED` | -| `enableDriveFormat` | always true | `GROUNDED` | -| `vcpuCount` | the smallest chosen memory node's physical cores, floored at 4 | `GROUNDED` | -| `minHugePagesSize` | the smallest free huge pages on any chosen node, or unset | `ASSERTED` | -| `socketsToUse` | never set | `INVENTED` | -| `nodesPerSocket` | never set | `INVENTED` | -| `stripe` | never set | `INVENTED` | -| `fabricType` | never set | `INVENTED` | -| `enableFailureDomains` | never set | `INVENTED` | - -**`maxSubsystemCount`** is the middle of the API's range, chosen so that a -reviewer who has not thought about it gets a working cluster and one who has can -see the number was not derived. Nothing a probe reports bears on it. - -**`enableDriveFormat`** is on the document because it is destructive and the -document is what somebody approves. The cluster's own field is immutable once the -cluster exists, so a default nobody saw could not be undone. - -**`vcpuCount`** takes the smallest worker because the control plane assumes the -count uniform across a cluster's nodes. The floor of 4 is the best-grounded -number in the generator: below it, sbcli's core layout assigns no NVMe-oF poller -core at all. Where the smallest worker cannot meet the floor, the floor is used -anyway and the note says the cluster asks for more than that worker has. - -A worker contributes a reading only when **every** chosen device sits on one -memory node. A worker whose devices span nodes contributes nothing, and if no -worker contributes, the value is the floor. - -**`minHugePagesSize`** reads the **free** pages on the chosen node, summed across -page sizes, and proposes the smallest across the fleet. It returns unset as soon -as **any** worker has none, and it truncates to whole gibibytes. - -**The five fields never set** each have a cost. The most expensive is -`socketsToUse`: empty means socket 0 alone, so a two-socket fleet is drafted as a -single-socket deployment and the disks the placement chose on socket 1 are in the -document while the cluster is laid out for socket 0. The layout is immutable on -the cluster it lands on. - -## 9. Assembling and naming the document - -| Rule | Value | Provenance | -|---------------------------------------|-----------------------------------------------------|-------------| -| Document name | `spec.discover.configName`, else `discovered-` | `ASSERTED` | -| `spec.approved` | always false, including on a re-run | `GROUNDED` | -| `spec.environment` | copied from what Inspecting concluded | `SPECIFIED` | -| Growth branch | `clusterRef` set means no cluster block is proposed | `GROUNDED` | -| The filter | never written into the document | `SPECIFIED` | -| An existing document of the same name | adopted, left as it is, with an event | `GROUNDED` | -| Owner reference | none. The document outlives the run | `GROUNDED` | - -**The report objects** are named by a formula holding to 63 bytes rather than the -253 a ConfigMap may have. `GROUNDED` in an observed API server rejection: the Job -that writes a report is named the same way, and a Job's name reaches its pods as -a label, where 63 is the limit. The digest is unconditional because the run and -the node join on a separator both may contain. - -Neither the name nor the labels can be read back as the values that produced -them. The node a report is about is the field inside the report, which is the only -place it appears as the cluster spells it. - -## 10. Rules that need a decision - -Every entry is `ASSERTED` or `INVENTED`: no design states it, and the code gives -no constraint for it. Ordered by what a wrong answer costs. - -| # | Rule | Where | The question | -|-----|---------------------------------------------------------------------------------|-------|-----------------------------------------------------------------------------------| -| 1 | The eligibility gate runs before the node-address win | §7.3 | Should a down link holding the cluster's own address be passed over in silence? | -| 2 | Resolved speed descending is the first sort key for an interface | §7.1 | Is the fastest link the management interface, or the data one? | -| 3 | A bridge or overlay is admitted only with the node's address | §7.1 | Is a bridge holding that address a management interface, or a machine to look at? | -| 4 | An aggregate's speed is the sum of its members | §7.2 | Is a bridge an aggregate for this purpose? Its ports are not one flow's bandwidth | -| 5 | The simplicity tiers | §7.1 | Why does a bond outrank a VLAN, and what is tier 3 for? | -| 6 | Capacity, then cores, as the placement's tie-breaks | §5.1 | Is more capacity the right second key, given the stripe argument for the first? | -| 7 | The placement ignores huge pages and NIC locality | §5.1 | Both are recorded as known. Which should enter the ranking? | -| 8 | Buckets exist only where an admitted device is | §5 | A node with cores, pages, and a fast NIC but no free disk is invisible | -| 9 | `minHugePagesSize` is the smallest **free** pages, unset if any worker has none | §8 | simplyblock allocates its own pages, so this reads a baseline as a requirement | -| 10 | The five cluster fields never set | §8 | `socketsToUse` above all: the second socket of every worker is left out | -| 11 | Deny before allow, and an empty allow list means allow everything | §4.2 | Both are conventions worth stating rather than discovering | -| 12 | Allow and deny fold case, including for block paths | §4.2 | `/dev/SDB` matches `/dev/sdb`, and Linux paths are case-sensitive | -| 13 | Device addresses are grouped case-sensitively | §6.1 | The opposite convention to the filter, on the same strings | -| 14 | A model or size filter silently excludes every claimed controller | §4.4 | They are unsized and unmodeled by construction | -| 15 | The document name is the run's name, not a timestamp | §9 | The field's own documentation says timestamp | - -## 11. Defects found while cataloging - -Each is a contradiction between the code and its own stated intent, or a path -that cannot do what it claims. None of them is a matter of taste. - -| # | Finding | Where | -|-----|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|--------| -| 1 | ~~An unparsable `driveSizeRange` silently admits every device.~~ **Fixed.** An admission webhook on `OperatorOps` refuses the run at the request, and both steps that read the filter refuse it again | §4.5 | -| 2 | **Both "huge pages unset" notes are computed and thrown away.** A draft that omits the field never explains why, against the file's own standard that every derived number carries a sentence | §8 | -| 3 | **The derived cluster name is unbounded against a 63-character limit.** `configName` has no maximum and a run name may be 253, so the create fails and the step retries until its deadline | §9 | -| 4 | **The draft's run label is the raw run name**, where every other use of that label is sanitized. An uppercase letter or 64 characters produces a create that can never succeed | §9 | -| 5 | **The group size comment claims every worker in a group has the same disks.** The signature hashes addresses, not sizes | §6.3 | -| 6 | **The placement's sentence asserts a comparison that did not decide.** A tie settled by cores or node id still reads "carries 2 devices (2T) against 2 (2T), so it was chosen," and says "1 unclaimed devices" for a single disk | §5.4 | -| 7 | **The no-worker failure message says infra nodes are excluded.** They are admitted without an opt-in | §2 | -| 8 | **`InUse: false` is not distinguished from never read.** The report's own comment says a reader must check for an unreadable entry first. Nothing does, and the rule that would is off by default | §4.5 | -| 9 | **A worker whose report cannot be decoded produces no refusal**, only an event, so it is absent from the explanation the status carries | §4.5 | -| 10 | **`Reason` is empty in the one case a reviewer most needs explained**: a worker with no serving interface | §7.4 | -| 11 | **Five of six recorded interface facts have no consumer**, including the sentence whose stated purpose is a record a reviewer reads | §7.4 | -| 12 | **`Upper` is collected, shipped, and read by nothing.** The documented use, a NIC holding no address with the address on a VLAN above it, is never implemented | §7.2 | -| 13 | **A wireless interface is a tier-0 management candidate.** `DEVTYPE=wlan` falls through the kind switch to physical | §7.1 | -| 14 | **A team or a VLAN on a kernel publishing no `DEVTYPE` becomes unbindable.** Bonds and bridges have a directory fallback, these do not | §7.1 | -| 15 | **Carrier and duplex are dropped from the report**, against a header claiming no filtering and no judgment. The rule admits `unknown` state, which is exactly the case carrier would settle | §7.1 | -| 16 | **Five schema limits are unchecked before writing**: groups per node set, workers per group, devices per selection, and the two address patterns. Each surfaces as a create rejection rather than a finding | §6, §9 | -| 17 | **The unknown memory-node bucket sorts first here and last in atlas-lib.** Same sentinel, opposite convention | §5 | -| 18 | **Two constants share one event wire value** in a package whose own premise is that a reason is an API | §9 | - -## 12. Rules that are specified and not built - -| Rule | Stated in | State | -|--------------------------------------------------------------------------------------------|-----------------------------------------------------------------------|-----------| -| `failureDomain` is seeded from `topology.kubernetes.io/zone` and left unset where no label | `design-clusterdeploymentconfig.md` §8.2, and the API field's own doc | `UNBUILT` | - -Nothing in the operator reads that label, and `failureDomain` is never written. -The test plan's `U-52` and `U-53` are the rows for it, both unimplemented. - -It is not cosmetic. A **growth** document against a cluster with -`enableFailureDomains` set is generated already failing its own validation, -because every group is required to carry a domain and none does. A creating -document escapes only because the generator never sets `enableFailureDomains` -either. diff --git a/discovery-generator-test-cases.md b/discovery-generator-test-cases.md deleted file mode 100644 index 01f54b4cb..000000000 --- a/discovery-generator-test-cases.md +++ /dev/null @@ -1,491 +0,0 @@ - - -# Discovery generator: test case mutations - -**Scope.** The generator is the path from probe reports to a drafted -`ClusterDeploymentConfig`: `nodeprobe.ReportFromConfigMap` → -`discovery.Planner.Plan` → `discovery.ClusterTemplateFor` → -`OperatorOpsReconciler.draftFor`. Everything after approval, meaning validation -and expansion into a `StorageCluster` and `StorageNode` objects, is the -`ClusterDeploymentConfig` controller's other half, covered by -`operator/docs/tests/test-plan-clusterdeploymentconfig.md` §1 (`U-01` onward). -Nothing here duplicates those IDs. - -**What one case is.** A directory of input objects and one expected output: - -```text -operator/internal/controllers/deployment/testdata/discovery/ - net/ - net-13-bond-holds-node-address/ - case.md # the row this directory is, copied from this document - ops.yaml # the OperatorOps, carrying spec.discover and its deviceFilter - nodes.yaml # corev1.Node objects: labels, taints, addresses, capacity - reports/.yaml # one ConfigMap per worker, the report JSON under report.json - expected.yaml # the ClusterDeploymentConfig the run should write - expected-refusals.txt # Plan.RefusalLines(), one per line, absent when there are none - net-14-.../ - numa/ - size/ -``` - -One directory per case, under its family. The family level is for reading rather -than for the harness, which walks the tree and takes every directory holding a -`case.md`: a hundred and seventy directories in one listing is a set nobody -reviews, and eleven families of a dozen is. - -The directory name carries the case identifier and a slug of the mutation, so -that a failure names the case without anybody looking it up. - -A case with no `expected.yaml` is a refusal case: `expected-error.txt` holds the -message the run fails with, which is what `kubectl get operatorops` shows. - -**Two harnesses.** The `Harness` column names which one a row belongs to. - -| Harness | Where | Covers | -|---------|------------------------------------------------------------------|-----------------------------------------------------------------| -| `CM` | `deployment` package, table-driven over the testdata directories | The whole path, report decoding and the draft's YAML included | -| `GO` | `discovery` package, built reports | Cases needing a `Planner` seam the controller never substitutes | - -A `GO` row exists because the controller always builds the default `Planner`: -`AllDevices`, `SingleNodeSet`, and `WorkerWasReadable` are reachable only by -constructing one directly. - ---- - -## 1. Mutation axes - -| Axis | Values exercised | -|------------------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| Device class | NVMe only, logical block only, both on one worker, neither | -| Device count per node | 0, 1, 4, 10, and one worker at 128 (the schema's `MaxItems`) | -| Device presentation | Kernel-presented disk, partition, loopback, multi-namespace controller, idle userspace controller, held userspace controller, attached simplyblock volume, iSCSI LUN | -| Device state | Available, mounted, held open, in a device-mapper stack, partitioned, unreadable | -| NUMA topology | 1, 2, 4, and 8 memory nodes. Devices on one node, split evenly, split unevenly, and on none | -| vCPU count | 2, 8, 32, and 196 logical CPUs, with and without simultaneous multithreading | -| Memory | 4 GiB, 64 GiB, 256 GiB, and 1 TiB total, with and without a per-node breakdown | -| Huge pages already allocated | None at all, no hugetlbfs, a pool with no free pages, 256 GiB free per node, a total without a per-node breakdown, a pool only on the node not chosen, 2 MiB and 1 GiB pools together | -| Huge-page headroom | Room for simplyblock's allocation on top of the existing pool, exactly enough, and short of it | -| Reuse instruction | A run told to take a set number of the pre-allocated pages, and a run told nothing | -| PCI addressing | Identical across the fleet, identical across a subset of it, one odd worker out, differing per worker, one address set a subset of another, unsorted, two buses, a five-digit domain | -| Interface count | One interface, one that is only loopback, two physical, and a dozen of mixed kinds | -| Interface kind | Physical, bridge, bond, VLAN, VLAN over bond, VXLAN, macvlan, veth, and loopback | -| Interface link | Up, down, dormant, unknown state, 1G, 10G, 100G, unknown speed, MTU 1500 and 9000 | -| Interface address | A routable v4 address, a global v6 address, link-local only, none, and the node's own `InternalIP` | -| Interface instruction | A run told which interface is management, told which are data NICs, and told neither | -| Stack resolution | An aggregate reporting no speed of its own, a derived interface over one, members on one memory node and on two, and a stack that points at itself | -| Fleet size | 0, 1, 3, and 32 workers | -| Fleet homogeneity | Uniform, two layouts, a majority layout with stragglers, one odd worker out, every worker distinct, a mix of NUMA topologies | -| Node role | Unlabeled, `worker`, `infra`, `control-plane`, `master`, and machines carrying two | -| Filter | Each of the seven `DeviceFilter` fields, singly and in the two combinations the CEL rule permits | -| Document shape | Creating (a cluster template) and growing (`spec.discover.clusterRef`) | -| Report validity | Current version, an older version, absent key, malformed JSON, no node name, a foreign node | - ---- - -## 2. Device class and inventory shape - -| ID | Mutation | Expected | Harness | -|--------|-------------------------------------------------------------|-------------------------------------------------------------------------------------------|---------| -| DEV-01 | 4 NVMe disks, one memory node, NVMe run | 1 group, `devices.nvme` holds the 4 addresses ascending, name `group-1-nvme-4x3T` | `CM` | -| DEV-02 | 4 virtio disks, block run | 1 group, `devices.block` holds the 4 paths, class `block` | `CM` | -| DEV-03 | 4 virtio disks, NVMe run | No node sets. Every device refused by `device class`, the worker by `has devices` | `CM` | -| DEV-04 | 2 NVMe and 2 virtio disks on one worker, NVMe run | Only the 2 PCI addresses reach the group | `CM` | -| DEV-05 | The same worker, block run | Only the two virtio disks. The NVMe pair is refused for being the other class | `CM` | -| DEV-06 | One controller presenting two namespaces at one PCI address | The address is named once. `DeviceCount` is 1 for that controller | `CM` | -| DEV-07 | A SATA disk beside an NVMe one, block run | Only the SATA disk. The NVMe run takes only the NVMe disk | `CM` | -| DEV-08 | A rotational HDD beside an SSD, both free | Both admitted: no rule reads `Rotational`. See §14, gap G-2 | `CM` | -| DEV-09 | 10 NVMe disks, one memory node | All 10 in one group, name `group-1-nvme-10x3T` | `CM` | -| DEV-10 | A partition and a loopback device beside 4 disks | Both pre-filtered: absent from `Plan.Explain`, present in `RefusalLines` | `CM` | -| DEV-11 | 128 NVMe disks on one worker | One group at the selection's `MaxItems`. The document still applies | `CM` | -| DEV-12 | A disk whose `Kind` is `disk` and whose `SizeBytes` is 0 | Admitted. The group is named `unsized` | `CM` | -| DEV-13 | A simplyblock volume attached to the worker, NVMe run | Refused as this fleet's own volume, naming the cluster it belongs to | `CM` | -| DEV-14 | The same volume on a block run | Refused the same way, since the rule is in both pipelines and reads no filter | `CM` | -| DEV-15 | A worker whose every disk is an attached volume | No draft, and the explanation counts them together rather than listing each | `CM` | -| DEV-16 | A fabric namespace another product exported | Refused for being on a fabric, not as a simplyblock volume: the NQN does not parse as one | `CM` | -| DEV-17 | An iSCSI LUN beside a virtio disk, block run, no allow list | Only the virtio disk. A LUN is storage across a network and is never taken by default | `CM` | -| DEV-18 | The same worker with the allow list naming the LUN | Both disks. Naming it is the decision a run cannot make for a fleet | `CM` | -| DEV-19 | An iSCSI LUN on an NVMe run | Refused for being the other class, before the iSCSI rule is reached | `CM` | - -## 3. NUMA topology - -The placement ranks a worker's memory nodes by unclaimed device count, then -capacity, then physical cores, then a real node ahead of the unknown bucket, -then the node id. Each row below pins one of those tie-breaks. - -| ID | Mutation | Expected | Harness | -|---------|------------------------------------------------------------------------------|-------------------------------------------------------------------------------------------|---------| -| NUMA-01 | 1 memory node, 4 disks on it | All 4 used, reason "every unclaimed device is on NUMA node 0" | `CM` | -| NUMA-02 | 2 memory nodes, 2 disks each, equal size, equal cores | Node 0 chosen on the id tie-break. 2 disks refused by the placement | `CM` | -| NUMA-03 | 2 memory nodes, 1 disk on node 0 and 3 on node 1 | Node 1 chosen on count. The single disk refused | `CM` | -| NUMA-04 | 2 memory nodes, 2 × 8 TiB on node 0 and 3 × 1 TiB on node 1 | Node 1 chosen: count beats capacity | `CM` | -| NUMA-05 | 2 memory nodes, 2 disks each, 1 TiB on node 0 and 2 TiB on node 1 | Node 1 chosen on the capacity tie-break | `CM` | -| NUMA-06 | 2 memory nodes, 2 disks each of one size, 8 cores on node 0 and 24 on node 1 | Node 1 chosen on the core tie-break | `CM` | -| NUMA-07 | 4 memory nodes, 2 disks each (NPS4) | Node 0 chosen. 6 disks refused. `vcpuCount` is node 0's physical cores | `CM` | -| NUMA-08 | 8 memory nodes, 1 disk each but 3 on node 5 | Node 5 chosen on count | `CM` | -| NUMA-09 | Every device reports `numaNode: -1` | The unknown bucket is used. `vcpuCount` floors at 4, `minHugePagesSize` unset, both noted | `CM` | -| NUMA-10 | 2 disks on node 0 and 2 on `-1`, node 0 carrying cores | Node 0 chosen on cores. The unknown bucket is never preferred | `CM` | -| NUMA-11 | The same worker with `Placement: AllDevices` | Every disk used. No single chosen node, so `vcpuCount` floors and huge pages go unset | `GO` | -| NUMA-12 | A 1-node worker and a 2-node worker whose chosen addresses coincide | One group: grouping reads addresses and the interface, never the topology | `CM` | -| NUMA-13 | 2 memory nodes, all 10 disks on node 1 and none on node 0 | Node 1 chosen with nothing to choose. The CPU note names node 1's cores | `CM` | -| NUMA-14 | A large worker: 2 nodes × 49 cores, 5 disks each, huge pages on both | 5 disks in the draft, `vcpuCount` 49, `minHugePagesSize` from node 0's free pages alone | `CM` | -| NUMA-15 | `cpu.numaNodes` empty while devices report nodes 0 and 1 | Placement falls through to the id tie-break. `vcpuCount` floors at 4 with its note | `CM` | - -## 4. Node size: vCPU, memory, and huge pages - -simplyblock allocates its own huge pages: the node writes the pool it needs, and -`cluster.minHugePagesSize` is the floor on what it writes, rendered into the -per-node configuration as `MAX_HUGE_PAGES_SIZE`. What discovery reads a host's -pool for is therefore not whether SPDK will find pages to consume. It is the -baseline the new total is added to, which is four questions: - -1. Are pages allocated already, and how many? -2. Are there none, on a kernel that has hugetlbfs and on one that has not? -3. Does the machine have memory for simplyblock's allocation on top of the - existing pool? -4. Is the run told to take a stated number of the pre-allocated pages instead of - adding its own? - -The generator answers none of them. It reads the free pages of the chosen memory -node and proposes them as the cluster's floor, and it proposes nothing at all -when a worker has none, on the stated grounds that SPDK consumes pages rather -than reserving them. Every row below records that output, and §14 carries the -seven findings it produces, `G-14` through `G-20`. - -| ID | Mutation | Expected | Harness | -|---------|------------------------------------------------------------------------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|---------| -| SIZE-01 | 8 logical CPUs, 4 physical cores on the chosen node | `vcpuCount: 4`, with the ordinary derivation note, not the floor note | `CM` | -| SIZE-02 | 2 logical CPUs, 2 physical cores | `vcpuCount: 4` with the note saying the cluster asks for more than the worker has. Not a refusal | `CM` | -| SIZE-03 | 196 logical CPUs, 98 physical cores, 1 memory node, 1 TiB RAM | `vcpuCount: 98`, the whole machine. See §14, gap G-3 | `CM` | -| SIZE-04 | 196 logical CPUs across 2 nodes, 1 TiB RAM, 512 GiB per node | `vcpuCount: 49`, the chosen node's cores | `CM` | -| SIZE-05 | A fleet whose chosen nodes carry 4, 16, and 49 cores | `vcpuCount: 4`, the note naming the smallest worker | `CM` | -| SIZE-06 | 4 GiB total, 900 MiB available, 4 free disks | Admitted unchanged: no rule reads memory. See §14, gap G-4 | `CM` | -| SIZE-07 | `memory` zero and an `unreadable` entry for it | Admitted by default. The draft is written from the disks | `CM` | -| SIZE-08 | The same report with `WorkerRules: WorkerWasReadable` | The worker is refused, the message quoting the unreadable readings | `GO` | -| SIZE-09 | 512 × 1 GiB pages pre-allocated, 256 free per node | **Contested.** `minHugePagesSize: 256G`: another workload's reservation becomes simplyblock's own floor. See §14, gap G-14 | `CM` | -| SIZE-10 | One worker of three with no pre-allocated pages on its chosen node | **Contested.** Unset for the whole cluster, the note naming that worker. Nothing pre-allocated is the ordinary case, not a reason to propose nothing. See §14, gap G-15 | `CM` | -| SIZE-11 | 256 × 2 MiB pages free on the chosen node, 512 MiB in all | Unset, the note saying the smallest reservation is under a gigabyte. A 1.5 GiB reading truncates to `1G` | `CM` | -| SIZE-12 | A pool with a total and an empty `numaNodes` | Unset: the reservation exists and its distribution is unknown | `CM` | -| SIZE-13 | Pages pre-allocated on node 1 while the placement chooses node 0 | Unset, the note naming the worker. The placement does not read huge pages, and the pool on the other node is neither counted nor reported | `CM` | -| SIZE-14 | Swap total 8 GiB, swap free 0 | No effect on the draft. See §14, gap G-5 | `CM` | -| SIZE-15 | `nodes.yaml` allocatable memory far under capacity | Recorded on the `KubeNode`, no effect on the draft | `CM` | -| SIZE-16 | 256 GiB RAM, 200 GiB already in huge pages, 40 GiB available | **Contested.** `minHugePagesSize: 200G` and no arithmetic against the 40 GiB left. See §14, gap G-16 | `CM` | -| SIZE-17 | 1 TiB RAM, nothing pre-allocated, the whole machine available | **Contested.** Nothing proposed, though the room for an allocation is there and reported. See §14, gaps G-15 and G-16 | `CM` | -| SIZE-18 | A run meant to take 128 of the host's 256 pre-allocated pages | **Contested.** No field expresses it. `spec.discover` carries no huge-page input, and neither `minHugePagesSize` nor `spdkSystemMemory` is written by a run. See §14, gap G-17 | `CM` | -| SIZE-19 | A pool of 256 pages, all promised to a mapping: `resv` 256, `free` 0 | Indistinguishable from an untouched pool of 256, because the report drops `resv_hugepages`. See §14, gap G-18 | `CM` | -| SIZE-20 | A 1 GiB pool and a 2 MiB pool, both with free pages on the chosen node | The two sizes are summed into one figure, and nothing says which size the node should take | `CM` | -| SIZE-21 | No hugetlbfs at all, `hugePages` absent from the report | The same unset output as a machine whose pools are full, so the two are not told apart. See §14, gap G-19 | `CM` | - -## 5. PCI addressing and disk count - -| ID | Mutation | Expected | Harness | -|--------|--------------------------------------------------------------------|------------------------------------------------------------------------------------------------------------|---------| -| PCI-01 | 32 workers, identical addresses | 1 group of 32 workers | `CM` | -| PCI-02 | 2 workers, 4 disks each, different slots | 2 groups of 1 worker each | `CM` | -| PCI-03 | 32 workers: 16 on layout A and 16 on layout B | 2 groups of 16, ordered by their first worker's name | `CM` | -| PCI-04 | 32 workers, every layout distinct | 32 groups in one node set, under the 64-group ceiling | `CM` | -| PCI-05 | Addresses reported in descending order | `devices.nvme` ascending: the draft sorts | `CM` | -| PCI-06 | 10 disks across two buses, `0000:5e:*` and `0000:af:*` | One group, addresses ascending across both buses | `CM` | -| PCI-07 | A five-digit PCI domain, `10000:01:00.0` | **Contested.** The generator names it and the schema's item pattern refuses the document. See §14, gap G-6 | `CM` | -| PCI-08 | Uppercase hex in the report, `0000:5E:00.0` | Named as reported. Grouping is byte-exact, so a fleet mixing cases splits | `CM` | -| PCI-09 | Workers named `worker-1` … `worker-32` with two layouts | Group numbering follows lexicographic worker order: `worker-1`, `worker-10`, `worker-11` | `CM` | -| PCI-10 | 5 workers: 3 on layout A, the other 2 each distinct | 3 groups, of 3, 1, and 1 workers: partial homogeneity still groups what it can | `CM` | -| PCI-11 | 32 workers: 20 on layout A, 6 on layout B, 6 each distinct | 8 groups, of 20, 6, and six of 1, numbered by their first worker's name | `CM` | -| PCI-12 | 8 workers uniform but for one whose fourth disk is in another slot | 2 groups, of 7 and 1. One slot moved is a second group, which is the guess a reviewer regroups | `CM` | -| PCI-13 | 2 workers on one layout carrying 2 TiB and 4 TiB disks | **Contested.** One group, named `group-1-nvme-2x2T` for the first worker's capacity. See §14, gap G-13 | `CM` | -| PCI-14 | 3 workers, the third holding 3 of the 4 slots the others hold | 2 groups: the address set matches whole or not at all, never as a subset | `CM` | - -## 6. Fleet size and grouping - -| ID | Mutation | Expected | Harness | -|----------|-------------------------------------------------------------|----------------------------------------------------------------------------|---------| -| FLEET-01 | 1 worker, 4 disks | 1 group of 1 | `CM` | -| FLEET-02 | 3 uniform workers | 1 group of 3. The summary reads 3 workers, 12 devices, 1 group, 1 node set | `CM` | -| FLEET-03 | 32 uniform workers | 1 group of 32, under the 200-worker ceiling | `CM` | -| FLEET-04 | No reports at all | The run fails: the probe reports are gone | `CM` | -| FLEET-05 | 32 workers, 3 with every disk mounted | 29 workers drafted. 3 worker refusals with their device reasons folded in | `CM` | -| FLEET-06 | The same 32 reports, ConfigMaps listed in a different order | A byte-identical document | `CM` | -| FLEET-07 | Two ConfigMaps carrying a report for one node | One report per node. The draft names the worker once | `CM` | -| FLEET-08 | A report for a node absent from `status.workers` | Skipped without an event: it is not this run's evidence | `CM` | - -## 7. Network interfaces - -The draft names one management interface per group, and the rule is a ladder -rather than a match, because a fleet's machines do not agree on what their NICs -are called. What `ManagementInterface` applies today, in order: - -1. Refuse a kind nothing can be bound to: loopback, and a virtual device the - kernel does not identify, which is every veth and dummy on the machine. -2. Refuse anything whose state is neither `up` nor `unknown`. -3. Refuse anything holding no reachable address, which is to say link-local, - loopback, and unspecified addresses do not count. -4. Refuse a bridge or an overlay that does not hold the node's own address. - Both kinds can be bound and both are ordinarily somebody else's network: a - CNI bridge and a flannel overlay carry the pod network, and a hypervisor - host's bridge carries the address the cluster reaches the machine on. Which - of the two a given device is can be read from nothing but the address on it. -5. Of what survives, the interface holding the node's `InternalIP` wins - outright, whatever kind it is. -6. Failing that, the fastest link, with the simpler kind breaking that tie and - the lowest name breaking the rest. A bond and a VLAN report no speed of - their own, so the speed is resolved down the stack: an aggregate carries the - sum of its members, and a derived interface carries what its parent carries. - -Rungs 1, 4, and 6 are new, and they are what the interface kind in report -version 4 bought. Before it, every software interface read as "virtual" and was -refused, so a bonded or tagged management network, which is the ordinary -enterprise host, yielded no interface on any worker. - -Two things the ladder does not cover, and both are inputs rather than readings: -which interface a fleet means for management, and which interfaces it means for -the data plane. Neither is expressible today (see `G-23` and `G-24`), so the -rows for them record what a run does in their absence. - -### Where each rung came from - -No design document specifies any of this. `design-clusterdeploymentconfig.md` -shows `mgmtInterface: eth1` in an example and states no rule for arriving at it, -so the ladder was assembled from a control-plane failure and from judgment. The -two are not the same thing, and a fixture whose outcome turns on the second is -recording a proposal rather than an expectation. - -| Rung | Where it came from | Status | -|-------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------|--------------------| -| A node with no management address is refused | The control plane's own refusal, "No management interface with IP found in provided interfaces" (commit 69e9f3f2) | Grounded | -| The interface holding the node's `InternalIP` wins outright | The operator addresses the worker by that address everywhere else, so any other choice splits one deployment across two networks | Grounded | -| A link that is down is passed over | A link that is down keeps its address and carries nothing | Grounded | -| A link-local address does not count | It is configured without anybody assigning it and routes nowhere | Grounded | -| The cluster's own plumbing never wins | Stated as intent in commit 69e9f3f2. It was implemented as "anything virtual," which is what `G-21` records as wrong | Grounded as intent | -| The lowest name breaks a tie | Two runs over one fleet have to produce the same document. Which tiebreak is arbitrary; that there is one is not | Grounded | -| **The fastest link wins when no node address matches** | Asserted in commit 69e9f3f2 with no reason given | **Unratified** | -| **Which kinds can be bound at all** | Added with the kind reading. Loopback and an unidentified virtual device are safe to refuse; admitting the rest is a choice | **Unratified** | -| **A bridge or an overlay only with the node's own address** | Added with the kind reading | **Unratified** | -| **An aggregate carries the sum of its members** | Added with the stack resolution | **Unratified** | -| **A simpler kind breaks a speed tie** | Added with the stack resolution | **Unratified** | - -Twenty of the forty-two cases in this section carry no node address on a -candidate interface, so an unratified rung decides them. Their recorded -documents are evidence that the rule was applied consistently and no evidence -that the rule is right. - -The question each unratified rung is really asking: - -1. **Is the fastest link the management interface?** Management traffic is - light, and the fastest link is what a data path wants. Naming it for - management may be taking the wrong NIC for the wrong plane, which is what - `NET-24` records. The alternatives are the slowest addressed link, the lowest - name outright, and refusing to choose at all where a fleet has not said. -2. **Should a run choose at all when nothing identifies the interface?** A draft - naming the wrong NIC is approved as readily as one naming the right one. The - alternative is to name none and make the reviewer say, which trades a wrong - document for one that cannot be approved unedited. -3. **Is a bridge holding the node's address the management interface, or a - machine a reviewer should look at?** It is an ordinary hypervisor host and it - is also what a misconfigured worker looks like. - -| ID | Mutation | Expected | Harness | -|--------|---------------------------------------------------------------------------------------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------------|---------| -| NET-01 | One 10G interface, up, holding a routable address | `mgmtInterface` names it | `CM` | -| NET-02 | A 1G and a 10G interface, both up and addressed, no node address on either | The 10G is named | `CM` | -| NET-03 | The node's `InternalIP` on the 1G while the 10G is faster | The 1G is named: the node address wins outright | `CM` | -| NET-04 | Four interfaces, 1G and 10G, all unusable: down, bridged, or link-local only | `mgmtInterface` empty and the draft still written. See §14, gap G-7 | `CM` | -| NET-05 | `interfaces` absent from the report | The same: an empty interface and a written draft | `CM` | -| NET-06 | Two workers with identical disks calling their NIC `eth0` and `ens5f0` | 2 groups, which the draft validation then reports as unbuildable. See §14, gap G-25 | `CM` | -| NET-07 | Two 10G interfaces, `eth1` and `eth0`, both addressed | `eth0` is named. A second run names it again | `CM` | -| NET-08 | An addressed physical interface reporting speed 0 beside a 10G interface that is down | The addressed one is named | `CM` | -| NET-09 | A worker whose `InternalIP` sits on a bridge, as a hypervisor host's does | `br0` named. A bridge carrying the cluster's own address is the management network | `CM` | -| NET-10 | One interface holding only a global IPv6 address | Named: reachability is not address-family dependent | `CM` | -| NET-11 | An interface holding both a link-local and a routable address | Named on the routable one | `CM` | -| NET-12 | One interface and it is the loopback | `mgmtInterface` empty: loopback is refused by its ARPHRD type, not its name | `CM` | -| NET-13 | `bond0` holding the node address over two unaddressed physical slaves | `bond0` named, with `eth0` and `eth1` reported as the hardware under it | `CM` | -| NET-14 | `eth0.100` holding the node address, the parent `eth0` unaddressed | `eth0.100` named, resolving to `eth0` for its slot and memory node | `CM` | -| NET-15 | `vxlan.calico` holding an overlay address beside `eth0` holding the node address | `eth0` named. The overlay is refused for holding an address the cluster does not reach the machine on | `CM` | -| NET-16 | `bond0.100` over `bond0` over two slaves, the node address on the VLAN | `bond0.100` named, resolving through the bond to both NICs | `CM` | -| NET-17 | A `macvlan` and an `ipvlan` over `eth0`, all three addressed | `eth0` named: the three carry the same traffic and the physical one is the simplest | `CM` | -| NET-18 | Twelve `veth` interfaces holding pod-CIDR addresses beside one physical NIC | The physical NIC named, whatever the veth count | `CM` | -| NET-19 | `cni0` and `docker0` both addressed, one physical NIC with no address | `mgmtInterface` empty: a bridge is refused and an unaddressed NIC is not a candidate | `CM` | -| NET-20 | An interface whose state is `dormant`, and one whose state is `lowerlayerdown` | Both passed over: only `up` and `unknown` are admitted | `CM` | -| NET-21 | An interface whose state is `unknown` holding a routable address | Named. A driver that does not track carrier is not a reason to refuse it | `CM` | -| NET-22 | A run told that `ens5f0` is the management interface | **Contested.** No input expresses it, so the ranking decides regardless. See §14, gap G-23 | `CM` | -| NET-23 | A run told that `ens5f1` and `ens5f2` are the data NICs | **Contested.** `dataInterfaces` is never written, so the draft leaves the data plane unnamed. See §14, gap G-24 | `CM` | -| NET-24 | Two physical NICs, a 10G and a 100G, both up and addressed | The 100G named for management and no data NIC named at all, so the fast link is proposed for the wrong plane | `CM` | -| NET-25 | Two otherwise equal 10G NICs at MTU 9000 and MTU 1500 | The lower name wins. MTU is reported and not ranked on | `CM` | -| NET-26 | A 100G bridge beside a 10G physical NIC, both addressed | The 10G named: the kind ladder is decided before the speed | `CM` | -| NET-27 | An interface reporting state `up` with no link partner | Named: the report drops `carrier`, so a link with no partner reads as a working one. See §14, gap G-26 | `CM` | -| NET-28 | A bond reporting no speed over two 25G NICs, beside an addressed 10G NIC | The bond named: an aggregate carries the sum of its members | `CM` | -| NET-29 | A VLAN reporting no speed over that bond, beside the same 10G NIC | The VLAN named: a derived interface carries what its parent carries | `CM` | -| NET-30 | The chosen bond's members in sockets 0 and 1 | The memory node is reported as unknown rather than as one of the two | `CM` | -| NET-31 | A veth holding the node's own address | Nothing named: an unidentified virtual device is refused at the first rung, node address or not | `CM` | -| NET-32 | A bond whose `lower` names a second bond whose `lower` names the first | The bond is named and the resolution terminates | `GO` | -| NET-33 | The node's `InternalIP` on a link whose state is `down` | Nothing named: the state rung is applied before the address wins, so a link carrying nothing is refused however right its address is | `CM` | -| NET-34 | The node's `InternalIP` on both `bond0` and `bond0.100` | The first by the report's own order, which is the kernel's and is ascending by name, so two runs agree | `CM` | -| NET-35 | A bond reporting 50000 Mbps of its own over two 25G members | Its own reading is taken, and the members are not summed on top of it | `CM` | -| NET-36 | A bond whose `lower` names an interface the report does not carry | Named, with no members and no resolved speed: a member nothing describes contributes nothing | `CM` | -| NET-37 | A `team` interface over two 25G NICs, holding the node address | Named and resolved as a bond is, both being aggregates | `CM` | -| NET-38 | A bridge over a bond over two NICs, the node address on the bridge | Named, and the members resolve two levels down to the two NICs | `CM` | -| NET-39 | A `macvlan` over `eth0` holding the node address | Named, with `eth0` as its member and `eth0`'s speed as its own | `CM` | -| NET-40 | An interface carrying no `kind`, and nothing marking it virtual | Read as physical, which is what the fields before the kind said about it | `CM` | -| NET-41 | A worker whose only fast NIC is a 2x25G bond, ranked for placement | **Contested.** The memory node's fastest NIC reads as 0 Mbps: the placement reads the raw speed and the raw memory node, neither of which a bond has. See §14, gap G-27 | `CM` | -| NET-42 | Any successful run on a bonded host | **Contested.** The draft names the interface and says nothing about the slots, the speed, or the memory node resolved under it. See §14, gap G-28 | `CM` | - -## 8. Node roles and node sets - -| ID | Mutation | Expected | Harness | -|---------|----------------------------------------------------------|-----------------------------------------------------------------------|---------| -| ROLE-01 | 3 workers, no role labels | One node set, `discovered` | `CM` | -| ROLE-02 | 3 `infra` and 3 `worker` nodes, identical hardware | 2 node sets, `infra` first, one hardware group split across them | `CM` | -| ROLE-03 | A `control-plane` node with free disks beside 2 workers | A third node set, `control-plane`, ordered last | `CM` | -| ROLE-04 | A node labeled `infra` and `worker` | `infra` | `CM` | -| ROLE-05 | A node labeled `control-plane` and `worker` | `control-plane`: the most restrictive role wins | `CM` | -| ROLE-06 | A node labeled `master` | `control-plane`: both spellings are read | `CM` | -| ROLE-07 | No `nodes.yaml` at all | Every worker a `Worker`. The interface is chosen with no address hint | `GO` | -| ROLE-08 | A cordoned node and a tainted node, both with free disks | Both reach the draft. See §14, gap G-8 | `CM` | -| ROLE-09 | A node carrying an unrecognized role label | Treated as a worker: an unknown role is not a refusal | `CM` | -| ROLE-10 | 32 workers across `infra`, `worker`, and `control-plane` | 3 node sets, each holding its role's share of every hardware group | `CM` | - -## 9. Filters - -| ID | Mutation | Expected | Harness | -|---------|----------------------------------------------------------------|-------------------------------------------------------------------------------------------------------|---------| -| FILT-01 | `pcieDenyList` naming the boot slot on a uniform fleet | That address in no group. The refusal appears in `Explain`, not pre-filtered | `CM` | -| FILT-02 | `pcieAllowList` of 2 addresses against 10 disks | Only those 2 reach the group | `CM` | -| FILT-03 | `pcieModel: MZQL2` against a mixed-model worker | Only the matching disks. The refusal quotes both strings | `CM` | -| FILT-04 | `driveSizeRange: 1T-4T` with a 512 GiB disk present | The small disk refused, the range quoted in the reason | `CM` | -| FILT-05 | `driveSizeRange: 2T` against 2 TiB and 1.92 TB disks | Only the exact 2 TiB disks: a bare size is both bounds | `CM` | -| FILT-06 | `driveSizeRange: 2T-1T` | The run is refused, naming the field and what a range looks like. Admission refuses it at the request | `CM` | -| FILT-07 | `blockDenyList: /dev/sda` on a block run | The root disk in no group | `CM` | -| FILT-08 | `blockAllowList` of 2 paths against 6 block devices | Only those 2 | `CM` | -| FILT-09 | `enablePartitionedDevices` with a GPT-only refusal | Admitted. A disk also mounted stays refused | `CM` | -| FILT-10 | A filter that matches nothing on any worker | No node sets. The run fails naming the filter rule per worker | `CM` | -| FILT-11 | An allow list written in uppercase against lowercase addresses | Matched: the comparison folds case | `CM` | -| FILT-12 | One address in both the allow and the deny list | Refused: deny is evaluated first | `CM` | -| FILT-13 | `driveSizeRange` on a block run | It narrows the block class, which is the class being scanned | `CM` | -| FILT-14 | No `deviceFilter` at all | Every free whole NVMe disk reaches the draft | `CM` | - -## 10. Probe-refused devices and userspace controllers - -| ID | Mutation | Expected | Harness | -|---------|--------------------------------------------------------------------|------------------------------------------------------------------------------------------------------------------------------------|---------| -| HELD-01 | Every disk mounted | The worker refused. Its device reasons folded into one line, counted | `CM` | -| HELD-02 | Every disk held in a device-mapper stack | The same shape with the other reason | `CM` | -| HELD-03 | 4 NVMe controllers on `uio_pci_generic`, idle, no block devices | 4 synthesized devices, group `group-1-nvme-4xunsized` | `CM` | -| HELD-04 | The same controllers with `inUse: true` | No devices. The worker refused with the message saying something is driving its disks | `CM` | -| HELD-05 | 2 idle userspace controllers beside 2 kernel-presented disks | 4 addresses in the group, the kernel-bound pair counted once | `CM` | -| HELD-06 | An idle controller on `vfio-pci` | Claimable on the same terms as `uio_pci_generic` | `CM` | -| HELD-07 | A kernel-bound controller whose block device is already reported | Named once, not twice | `CM` | -| HELD-08 | Idle userspace controllers on a block run | No devices at all: a controller with no block device has no path. Worker refused | `CM` | -| HELD-09 | 16 loopback devices and 4 free disks | The 4 disks drafted. The loopbacks absent from `Explain` and present in `RefusalLines` | `CM` | -| HELD-10 | Half the controllers idle and half in use, no block devices | Only the idle half is offered | `CM` | -| HELD-11 | Userspace controllers whose process table the probe could not read | Refused, and the message says whether anything is driving them could not be established. An unchecked controller is not a free one | `CM` | - -## 11. Refusal cases - -| ID | Mutation | Expected | Harness | -|---------|----------------------------------------------------------------------------|-----------------------------------------------------------------------------------|---------| -| FAIL-01 | Every worker reports no devices and no controllers | The run fails: no worker has a device this run would use, one line per worker | `CM` | -| FAIL-02 | Every disk in the fleet mounted | The same, each line carrying the device reasons and their counts | `CM` | -| FAIL-03 | Every report unreadable | `ReportUnreadable` per ConfigMap, then the run fails for want of reports | `CM` | -| FAIL-04 | A filter that excludes every disk | The run fails. The refusal lines name the filter rule | `CM` | -| FAIL-05 | Disks fine, no usable interface anywhere | **Contested.** A draft is written with an empty `mgmtInterface`. See §14, gap G-7 | `CM` | -| FAIL-06 | Every worker under the vCPU floor | **Contested.** A draft is written with `vcpuCount: 4`. See §14, gap G-10 | `CM` | -| FAIL-07 | Every worker at 4 GiB of RAM | **Contested.** A draft is written unchanged. See §14, gap G-4 | `CM` | -| FAIL-08 | A single worker whose CPU tree is unreadable, disks fine | A draft with `vcpuCount: 4` and the note saying no worker reported its cores | `CM` | -| FAIL-09 | One worker of 32 refused, the rest fine | A draft of 31. The refusal is an event, never a failure | `CM` | -| FAIL-10 | A worker whose only disks are partitions, `enablePartitionedDevices` unset | Refused by `whole disk`, pre-filtered, so `Explain` gives the worker line alone | `CM` | - -## 12. Report and ConfigMap validity - -| ID | Mutation | Expected | Harness | -|-------|--------------------------------------------------------------------------|---------------------------------------------------------------------------|---------| -| CM-01 | `version: 4` in one report of three, the schema before the subsystem NQN | `ReportUnreadable` naming both versions. The other two are drafted | `CM` | -| CM-02 | A ConfigMap with no `report.json` key | `ReportUnreadable` saying the probe did not finish writing it | `CM` | -| CM-03 | `report.json` holding malformed JSON | `ReportUnreadable` quoting the parse failure | `CM` | -| CM-04 | A report with an empty `node` | Refused: nothing can be attributed to it | `CM` | -| CM-05 | A ConfigMap without the run label | Not listed, so not read | `CM` | -| CM-06 | A report from a node whose name exceeds 63 characters | Drafted under its full name: the label is truncated and the report is not | `CM` | -| CM-07 | A report of a 128-device worker near the ConfigMap ceiling | Decoded and drafted | `CM` | - -## 13. Document shape and the cluster template - -| ID | Mutation | Expected | Harness | -|---------|---------------------------------------------------------------|---------------------------------------------------------------------------------------------------|---------| -| TMPL-01 | Any successful run | `spec.approved` false, `enableDriveFormat` true, and the note about formatting | `CM` | -| TMPL-02 | Any successful run | `maxSubsystemCount: 30` with the note saying it is not a reading | `CM` | -| TMPL-03 | `spec.discover.configName` unset | The name is `discovered-` and the cluster `discovered--cluster` | `CM` | -| TMPL-04 | An `OperatorOps` name long enough to push the cluster past 63 | **Contested.** The document is written and `CreatingCluster` can never succeed. See §14, gap G-11 | `CM` | -| TMPL-05 | `spec.discover.clusterRef` set | No `spec.cluster`, and `clusterRef` carried through with the growth note | `CM` | -| TMPL-06 | Any run on a two-socket fleet | `socketsToUse` and `nodesPerSocket` unset, so one node per worker on socket 0. See §14, gap G-12 | `CM` | -| TMPL-07 | `spec.environment` from the run's status | Copied verbatim into the document | `CM` | -| TMPL-08 | Any run | No `deviceFilter` anywhere in the document: the resolved list is the record | `CM` | - ---- - -## 14. Gaps and contested expectations - -Each is a case above whose expected value is what the code does rather than what -the design says, or a behavior no case can assert because nothing implements it. -A fixture is still written for every one. The row records today's output so that -a later change to it is visible as a diff rather than a surprise. - -| # | Finding | Cases | -|------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|------------------------| -| G-1 | **Fixed.** `ClassRule` checked the transport on the NVMe side only, so a block run admitted an NVMe disk by its path and a draft could name both classes. A cluster is built out of one class, so the check is symmetric now: a block run refuses `NVMe` and `NVMeFabric` alike | DEV-05, DEV-07 | -| G-2 | Nothing reads `Rotational`, so a spinning disk is offered to a cluster on the same terms as an SSD | DEV-08 | -| G-3 | `vcpuCount` is the chosen node's whole physical core count, leaving nothing for the system on a large worker | SIZE-03 | -| G-4 | No worker rule reads memory: a 4 GiB machine is drafted as a storage node | SIZE-06, FAIL-07 | -| G-5 | Swap is reported and unread, though a host swapping is already oversubscribed | SIZE-14 | -| G-6 | The generator does not hold an address to the schema's PCI pattern, so a five-digit domain produces a document the API server refuses | PCI-07 | -| G-7 | No worker rule requires a management interface, so the generator writes a group it already knows cannot deploy. The `ClusterDeploymentConfig` controller's draft validation does report it as `NoManagementInterface`, one step and one object later than the run that produced it | NET-04, FAIL-05 | -| G-8 | A cordoned or tainted node reaches the draft, because the role labels are read and `Unschedulable` is not | ROLE-08 | -| G-9 | **Fixed.** An unreadable `driveSizeRange` dropped the rule and admitted every disk, and the comment claimed a validating caller reported it. There is one now: an admission webhook on `OperatorOps` refuses the run at the request, and both steps that read the filter refuse it again for a cluster whose webhook is not installed | FILT-06 | -| G-10 | A fleet under the vCPU floor is drafted at the floor, so approval produces nodes the workers cannot host | SIZE-02, FAIL-06 | -| G-11 | The cluster name is not checked against its own 63-character limit when derived | TMPL-04 | -| G-12 | The generator chooses one memory node per worker and proposes no `socketsToUse`, so the second socket of every dual-socket worker is left out of the cluster entirely | TMPL-06, NUMA-14 | -| G-13 | A group's signature is its addresses and its interface, so workers of differing capacity share a group and the group is named for whichever of them sorts first. Verified against the code | PCI-13 | -| G-14 | `minHugePagesSize` is the floor on the pool simplyblock allocates for itself, and the generator sets it to the free pages already on the chosen node, so another workload's reservation becomes this cluster's allocation floor | SIZE-09, SIZE-16 | -| G-15 | A worker with nothing pre-allocated leaves the field unset, on the stated grounds that SPDK consumes pages and does not reserve them. simplyblock allocates its own, so an empty pool is the ordinary starting state. The same premise heads `atlas-lib/inventory/hugepages.go` | SIZE-10, SIZE-17 | -| G-16 | Nothing checks that the host has memory for the allocation on top of the existing pool. `memory.availableBytes`, `memory.hugePagesBytes`, and the per-node `freeBytes` are all reported and all unread | SIZE-16, SIZE-17 | -| G-17 | No input says to take a stated number of the pre-allocated pages. `spec.discover` carries no huge-page field, and the only two knobs that exist are never written by a run: the cluster's `minHugePagesSize` and a group's `spdkSystemMemory` | SIZE-18 | -| G-18 | `hugePagesOf` drops `resv_hugepages`, which `inventory.HugePagePool` reads, so a draft cannot tell pages promised to a mapping from pages genuinely free | SIZE-19 | -| G-19 | A kernel with no hugetlbfs and a host whose pools are fully taken produce the same unset field and the same note, so a reviewer cannot tell a machine that cannot hold huge pages from one that has none left | SIZE-21 | -| G-20 | The cluster's field is documented as a floor and rendered into the per-node configuration as `MAX_HUGE_PAGES_SIZE`. Which of the two the node honors decides what a run should propose | SIZE-09, SIZE-16 | -| G-21 | **Fixed.** `servesManagement` refused every virtual interface, and a bond, a VLAN, a VXLAN, and a macvlan are all virtual, so a fleet whose management address sat on a bond or a VLAN yielded no interface on any worker. The ladder now refuses a kind rather than the absence of hardware | NET-13, NET-14, NET-16 | -| G-22 | **Fixed.** The report carried `virtual`, `loopback`, and `bridge` and nothing else, so a bond, a VLAN, a VXLAN, and a veth were indistinguishable in it. `inventory.Interface` now reads the device type the kernel publishes and the `lower_*` and `upper_*` links around it, and the report carries `kind`, `lower`, and `upper` at version 4 | NET-13, NET-14, NET-17 | -| G-23 | No input predefines the management interface. `DiscoverSpec` carries no interface field, so the ranking cannot be overridden for a fleet that knows which NIC it means | NET-22 | -| G-24 | No run writes `dataInterfaces`. Every drafted document leaves the data plane unnamed, and the fastest link is proposed for management instead | NET-23, NET-24 | -| G-25 | The grouper splits workers on their management interface while the expansion binds one interface per cluster, so a fleet whose machines name their NICs differently produces exactly the document `conflictingInterfaces` reports as unbuildable | NET-06 | -| G-26 | `interfacesOf` drops `carrier` and `duplex`, which `inventory.Interface` reads, so a link that is up with no partner is not distinguishable from one carrying traffic | NET-27 | -| G-27 | `NUMANodeBreakdown` ranks a memory node's NICs by the raw `speedMbps` and the raw `numaNode`, and a bond has neither. A worker whose data path is bonded or tagged therefore counts as having no NIC at all on every memory node, though the resolution that would answer both now exists | NET-41 | -| G-28 | `Management` carries the members, the resolved speed, and the memory node onto `Worker.Mgmt`, and nothing reads them. No note, field, or event says which slots a bonded management network lands in, so the evidence is gathered and not reported | NET-42 | - -## 15. Generation order - -The fixtures come first and whole. Every case in this document is a directory of -input objects, and the inputs are a statement of what the fleet was, which is -settled by the row rather than by anything the harness does with it. Writing the -complete set before any code reads it keeps the two apart: what is being tested -is reviewable as a tree of hosts, and the harness is then written against a set -that is already fixed rather than growing to fit the cases it happens to load. - -1. **Every case directory, all of them, inputs only.** `ops.yaml`, `nodes.yaml`, - `reports/*.yaml`, and the `case.md` naming the row. This is the bulk of the - work and none of it depends on a harness existing. -2. **Review the tree.** A family at a time, against its section here. A wrong - fixture is a test that passes for the wrong reason, and it is far cheaper to - catch in a directory of YAML than in a golden file. -3. **The harness**, walking the tree and driving each case through the run. It - is written once, against the whole set, and a case it cannot load is a gap in - the harness rather than a case to drop. -4. **The expected outputs.** For a row stating a value, written by hand from the - row. For the rest, recorded from a run and then read against the row before - being committed, because a golden file nobody read is a record of what the - code did rather than of what it should do. -5. **The `GO` rows**, in the `discovery` package, since they substitute a - `Planner` seam the controller never does and have no directory. - -The contested rows are written like any other and keep their gap number in -`case.md`. They record today's output deliberately, so that the day one of them -changes, the diff is the finding rather than a surprise. diff --git a/fio-quick.yaml b/fio-quick.yaml deleted file mode 100644 index 500eda7c6..000000000 --- a/fio-quick.yaml +++ /dev/null @@ -1,104 +0,0 @@ -# Quick standalone fio Deployment — mirrors a single fio_migration_test.py workload pod. -# Uses the existing XFS StorageClass cloned from the live pool by the test harness. -# -# kubectl -n default apply -f test/fio-quick.yaml -# kubectl -n default logs -f deploy/fio-quick # live eta status -# kubectl -n default exec deploy/fio-quick -- cat /logs/result.json # final JSON after run -# kubectl -n default delete -f test/fio-quick.yaml # cleanup (Deployment + PVC) -# -# NOTE: the simplyblock volume is ReadWriteOnce, so this Deployment is pinned to -# replicas: 1 — a single PVC cannot be mounted by multiple pods. For N independent -# fio volumes, use a StatefulSet with volumeClaimTemplates instead. -# -# Knobs match fio_migration_test.py defaults: 4k randrw 70/30, iodepth 16, numjobs 4, -# direct=1/libaio, time_based 600s, 1 GiB file on the simplyblock XFS volume. -apiVersion: v1 -kind: PersistentVolumeClaim -metadata: - name: fio-quick-pvc - labels: - app: fio-quick - -spec: - accessModes: ["ReadWriteOnce"] - storageClassName: simplyblock-default-simplyblock-cluster-pool1-xfs - resources: - requests: - storage: 10Gi ---- -apiVersion: apps/v1 -kind: Deployment -metadata: - name: fio-quick - labels: { app: fio-quick } -spec: - replicas: 1 - selector: - matchLabels: { app: fio-quick } - strategy: - type: Recreate # RWO PVC: old pod must release the volume before the new one mounts - template: - metadata: - labels: { app: fio-quick } - spec: - terminationGracePeriodSeconds: 5 - # keep off control-plane nodes (same intent as the test's worker pinning) - affinity: - nodeAffinity: - requiredDuringSchedulingIgnoredDuringExecution: - nodeSelectorTerms: - - matchExpressions: - - key: node-role.kubernetes.io/control-plane - operator: DoesNotExist - containers: - - name: fio - image: alpine:3 - imagePullPolicy: IfNotPresent - command: - - sh - - -c - - | - set -u - echo "[pod] $(date -u +%FT%TZ) installing fio" - apk add --no-cache fio >/dev/null 2>&1 || { echo "[pod] apk add fio FAILED"; exit 90; } - mkdir -p /logs - echo "[pod] $(date -u +%FT%TZ) starting fio" - fio \ - --name=fiotest \ - --filename=/data/fiotest \ - --size=1G \ - --ioengine=libaio \ - --direct=1 \ - --rw=randrw \ - --rwmixread=70 \ - --bs=4k \ - --iodepth=16 \ - --numjobs=4 \ - --group_reporting \ - --time_based \ - --runtime=1200 \ - --continue_on_error=all \ - --percentile_list=50:95:99:99.9 \ - --write_iops_log=/logs/iops \ - --write_lat_log=/logs/lat \ - --write_bw_log=/logs/bw \ - --log_avg_msec=1000 \ - --eta=always \ - --eta-newline=5 \ - --output=/logs/result.json \ - --output-format=json - rc=$? - echo "$rc" > /logs/fio.rc - echo "[pod] $(date -u +%FT%TZ) fio exited rc=$rc" - # keep the container alive so logs can be collected after the run - sleep 100000 - volumeMounts: - - { name: data, mountPath: /data } - - { name: logs, mountPath: /logs } - resources: - requests: { cpu: 250m, memory: 256Mi } - volumes: - - name: data - persistentVolumeClaim: { claimName: fio-quick-pvc } - - name: logs - emptyDir: {} diff --git a/job.md b/job.md deleted file mode 100644 index e8a1438e6..000000000 --- a/job.md +++ /dev/null @@ -1,96 +0,0 @@ -Kubernetes Operator & CSI Engineer (m/f/d) -simplyblock · Remote (EU time zones) or Berlin · Full-time · Junior to Intermediate - -## Why this role exists - -At simplyblock, we build Kubernetes-native, software-defined NVMe-over-Fabrics block storage. Customers run it -underneath databases and stateful workloads where downtime is measured in money and latency is measured in microseconds. - -Our operator and CSI driver are the part of that product customers actually talk to, and user experience is our key -metric. The operator turns roughly twenty custom resources into a running, self-healing storage cluster, managing -everything nodes, devices, pools, snapshots, backups, replication, migrations, and upgrades. The CSI driver turns -customer PVC requests into NVMe-oF volumes and keeps them attached throughout failovers, reboots, and path changes. - -Both are growing fast. We are looking for someone early in their career but already well-versed in Go, who wants to -spend the next few years becoming genuinely excellent at Kubernetes controllers and storage. - -We are a small company, which means the same person designs a CRD in the morning and explains to a customer's platform -team in the afternoon why their cluster did what it did. Both halves belong to the role, and the second one is not a -tax on the first: the questions customers ask are where the next fix usually comes from. - -## What you'll do - -- Write and extend our controllers: reconcilers that never block, complex processes modeled as multistep operations in - persisted phases. You will start on well-scoped controllers and grow into the ones that move data. -- Work on the CSI driver, from provisioning and attachment down to the NVMe-oF paths, multipath states, and mounts. - That includes the unglamorous half that matters most: self-healing after something crashed halfway through. -- Design and evolve the API. CRDs are user experience: validation and immutability where they belong, typed phases and - conditions, and conversions that keep an already-shipped field from breaking a customer. -- Write the failing test first. Every fix and every behavior change starts with a test describing the wanted behavior, - at the level that proves it: fake clients, a fake simplyblock, or a live cluster with real NVMe devices. -- Reproduce and fix real cluster behavior. A drain that stalls, a migration that loses writes, a node that never gets - re-probed. You read logs, build the reproducer, and write the regression test that proves the problem. This part of - the job is hard, and it is also the most fun. -- Work directly with customers. You will join a support conversation with a customer's engineers, read their logs with - them, explain what their cluster is doing in terms a DBA or a platform team can act on, agree on the next step, and - carry the outcome back into an issue, a regression test, or a fix. - -## What you bring - -- Go, in production. One to three years is the shape we have in mind: interfaces, contexts, goroutines, and errors - hold no surprises, and you have shipped code that survived contact with users. -- Kubernetes fluency as a user, and curiosity as a developer. PVCs, StorageClasses, DaemonSets, RBAC, and lifecycle - should be familiar. Having written a Kubernetes controller before is a strong plus. Wanting to is the minimum. -- Linux fundamentals. Block devices, filesystems, mounts, `/sys` and `/dev`, and enough of a systems instinct to be - suspicious of the right things. -- Composure in front of a customer. Under pressure, you hold a professional technical conversation with their engineers: - listen before diagnosing, separate what is known from what is still a guess, say when you have no answer yet rather - than promising a date. Underpromise, overdeliver. -- Testing as a habit, not a chore. You would rather spend an hour making a failure reproducible than an afternoon - guessing. -- Research and prototyping as part of your workflow. You read code to find out how things really work, and you throw - away the branch that led nowhere without mourning it. -- AI-assisted engineering, with judgment. We use Claude daily for code and log analysis, reproductions, refactors, and - documentation. Use it to make yourself faster, and learn where it (quietly) does not help. -- Communication that does not wait to be asked. Design documents, commit messages, issue comments, and standups are - how a distributed team thinks together, and asking early beats guessing quietly. -- The temperament for distributed systems. Comfortable with not knowing yet, the patience to keep digging, and the - honesty to say "this test passes, but I don't understand why." - -## Preferred - -- The CSI specification, or any prior work on a storage or networking plugin -- NVMe, NVMe-oF, iSCSI, SPDK, or other userspace / kernel-bypass storage stacks -- kubebuilder, envtest, Ginkgo, or client-go beyond the basics -- Helm chart authoring, OLM bundles, or OpenShift -- gRPC, and reading an OpenAPI-generated client without flinching -- One or more of Python, JavaScript/TypeScript, C/C++, Bash -- Fluent English required with German or other European languages a plus - -## What you don't need - -- Storage expertise. We will teach you NVMe-oF, and nobody here was born knowing what an ANA group is. -- A CV full of infrastructure companies. Two good years and clear momentum beat five unremarkable ones. -- Certainty. You will be reviewed, mentored, and occasionally told to throw a branch away. That is the job working - correctly. - -## Your first six months - -- Month 1: ship small, real changes across the operator and the CSI driver. Validation markers, a webhook check, a - status condition. Analyzing the first logs to make yourself familiar with the system components. -- Month 2–3: own features end to end. From design document to merged tests. Take your first on-cluster bug from a - vague symptom to a regression test and successful fix. Sit in on customer calls, first listening, then answering the - parts you know. -- Month 4–6: you carry weight. The whole system is your territory: you follow a symptom across components to the - mechanism, and you know where the fix belongs. Your features are loved because the CRD, the events, and the failure - messages carry the user experience. You take an escalation alone, and we hear about it from the satisfied customer. - -## What we offer - -- A technically challenging product where your commits reach production clusters, not a sandbox -- Review and mentoring from the engineers who wrote the operator, the CSI driver, and the storage engine underneath -- A codebase with strong and explicit conventions, being built with user experience in mind -- [Compensation range] -- [Equity / VSOP] -- [Remote setup, hardware budget, learning budget, conference attendance] -- [Holiday allowance, other benefits] diff --git a/local-path-storage.yaml b/local-path-storage.yaml deleted file mode 100644 index c0342ec66..000000000 --- a/local-path-storage.yaml +++ /dev/null @@ -1,167 +0,0 @@ -apiVersion: v1 -kind: Namespace -metadata: - name: local-path-storage - ---- -apiVersion: v1 -kind: ServiceAccount -metadata: - name: local-path-provisioner-service-account - namespace: local-path-storage - ---- -apiVersion: rbac.authorization.k8s.io/v1 -kind: Role -metadata: - name: local-path-provisioner-role - namespace: local-path-storage -rules: - - apiGroups: [""] - resources: ["pods"] - verbs: ["get", "list", "watch", "create", "patch", "update", "delete"] - ---- -apiVersion: rbac.authorization.k8s.io/v1 -kind: ClusterRole -metadata: - name: local-path-provisioner-role -rules: - - apiGroups: [""] - resources: - ["nodes", "persistentvolumeclaims", "configmaps", "pods", "pods/log"] - verbs: ["get", "list", "watch"] - - apiGroups: [""] - resources: ["persistentvolumes"] - verbs: ["get", "list", "watch", "create", "patch", "update", "delete"] - - apiGroups: [""] - resources: ["events"] - verbs: ["create", "patch"] - - apiGroups: ["storage.k8s.io"] - resources: ["storageclasses"] - verbs: ["get", "list", "watch"] - ---- -apiVersion: rbac.authorization.k8s.io/v1 -kind: RoleBinding -metadata: - name: local-path-provisioner-bind - namespace: local-path-storage -roleRef: - apiGroup: rbac.authorization.k8s.io - kind: Role - name: local-path-provisioner-role -subjects: - - kind: ServiceAccount - name: local-path-provisioner-service-account - namespace: local-path-storage - ---- -apiVersion: rbac.authorization.k8s.io/v1 -kind: ClusterRoleBinding -metadata: - name: local-path-provisioner-bind -roleRef: - apiGroup: rbac.authorization.k8s.io - kind: ClusterRole - name: local-path-provisioner-role -subjects: - - kind: ServiceAccount - name: local-path-provisioner-service-account - namespace: local-path-storage - ---- -apiVersion: apps/v1 -kind: Deployment -metadata: - name: local-path-provisioner - namespace: local-path-storage -spec: - replicas: 1 - selector: - matchLabels: - app: local-path-provisioner - template: - metadata: - labels: - app: local-path-provisioner - spec: - serviceAccountName: local-path-provisioner-service-account - containers: - - name: local-path-provisioner - image: docker.io/rancher/local-path-provisioner:v0.0.36 - imagePullPolicy: IfNotPresent - command: - - local-path-provisioner - - --debug - - start - - --config - - /etc/config/config.json - volumeMounts: - - name: config-volume - mountPath: /etc/config/ - env: - - name: ALLOW_UNSAFE_HELPER_POD_TEMPLATE - value: "true" - - name: POD_NAMESPACE - valueFrom: - fieldRef: - fieldPath: metadata.namespace - - name: CONFIG_MOUNT_PATH - value: /etc/config/ - volumes: - - name: config-volume - configMap: - name: local-path-config - ---- -apiVersion: storage.k8s.io/v1 -kind: StorageClass -metadata: - name: local-path -provisioner: rancher.io/local-path -volumeBindingMode: WaitForFirstConsumer -reclaimPolicy: Delete - ---- -kind: ConfigMap -apiVersion: v1 -metadata: - name: local-path-config - namespace: local-path-storage -data: - config.json: |- - { - "nodePathMap":[ - { - "node":"DEFAULT_PATH_FOR_NON_LISTED_NODES", - "paths":["/opt/local-path-provisioner"] - } - ] - } - setup: |- - #!/bin/sh - set -eu - mkdir -m 0777 -p "$VOL_DIR" - teardown: |- - #!/bin/sh - set -eu - rm -rf "$VOL_DIR" - helperPod.yaml: |- - apiVersion: v1 - kind: Pod - metadata: - name: helper-pod - spec: - priorityClassName: system-node-critical - tolerations: - - key: node.kubernetes.io/disk-pressure - operator: Exists - effect: NoSchedule - containers: - - name: helper-pod - image: docker.io/library/busybox - imagePullPolicy: IfNotPresent - securityContext: - privileged: True - runAsUser: 0 diff --git a/sc.yaml b/sc.yaml deleted file mode 100644 index 4090888e0..000000000 --- a/sc.yaml +++ /dev/null @@ -1,27 +0,0 @@ -allowVolumeExpansion: true -apiVersion: storage.k8s.io/v1 -kind: StorageClass -metadata: - labels: - simplyblock.io/auto-restart-on-pathloss: "true" - name: simplyblock-default-simplyblock-cluster-pool1-xfs -parameters: - cluster_id: d5dcb59d-dec8-4fb8-a9e2-e61ade9b7275 - compression: "False" - csi.storage.k8s.io/fstype: xfs - distr_ndcs: "1" - distr_npcs: "1" - encryption: "False" - fabric: tcp - lvol_priority_class: "0" - max_namespace_per_subsys: "1" - pool_name: pool1 - qos_r_mbytes: "0" - qos_rw_iops: "0" - qos_rw_mbytes: "0" - qos_w_mbytes: "0" - replicate: "False" - tune2fs_reserved_blocks: "0" -provisioner: csi.simplyblock.io -reclaimPolicy: Delete -volumeBindingMode: WaitForFirstConsumer diff --git a/test-cluster-111.yaml b/test-cluster-111.yaml deleted file mode 100644 index 98261eb5e..000000000 --- a/test-cluster-111.yaml +++ /dev/null @@ -1,97 +0,0 @@ ---- -apiVersion: storage.simplyblock.io/v1alpha2 -kind: StorageCluster -metadata: - name: simplyblock-cluster - namespace: default -spec: - enableNodeAffinity: false - maxSubsystemCount: 10 - vcpuCount: 6 - volumeMigrationSettings: - # There is no enabled switch at v1alpha2. Migration cannot be turned off, - # because a drain, a rebalance, and a device replacement are all performed - # by moving volumes (design-storagecluster.md §3.1). - # - # rebalancerImage is used for the migration path-validation Job (must include nvme-cli). - # - # This tag must track the branch the operator itself is deployed from. The Job runs a - # *different* binary than the operator, so a stale tag here silently pairs a current - # operator with an old validator — the operator still fans the Jobs out per node while - # the Job inside them runs whatever that tag was built from. Nothing in the logs states - # the mismatch; it only shows up as the Job's log format not matching the code. - rebalancerImage: docker.io/simplyblock/simplyblock-rebalancer:multi-namespace-volume-migration - # Automatic, latency-driven rebalancing. Whether it runs at all is the - # top-level switch; the block below only tunes it, and migrationEnabled is - # now spelled as the disable it inverts, so leaving it out is the dry run - # turned off. - #enableVolumeAutoPlacement: true - #volumeAutoPlacement: - # enableLatencyBenchmark: true - # disableMigration: false - # metricsBackend: Prometheus - # prometheusURL: http://simplyblock-prometheus.simplyblock:9090 - warningThreshold: - capacity: 80 - provisionedCapacity: 100 - criticalThreshold: - capacity: 90 - provisionedCapacity: 100 ---- -apiVersion: storage.simplyblock.io/v1alpha1 -kind: StorageNodeSet -metadata: - name: simplyblock-node - namespace: default -spec: - clusterName: simplyblock-cluster - # Stated explicitly because the ControlPlane fallback cannot be read: the - # operator resolves it through the v1alpha1 ControlPlane, which now needs the - # conversion webhook that is not deployed here. - # clusterImage: "public.ecr.aws/simply-block/simplyblock:main" - # spdkImage: "public.ecr.aws/simply-block/ultra:debuge-migration-latest" # the SPDK container - skipKubeletConfiguration: false - enableCpuTopology: false - socketsToUse: - - "0" - mgmtIfname: eth0 - workerNodes: - - vm02.simplyblock4.localdomain - - vm03.simplyblock4.localdomain - - vm04.simplyblock4.localdomain ---- -apiVersion: storage.simplyblock.io/v1alpha2 -kind: StoragePool -metadata: - name: pool1 - namespace: default -spec: - clusterRef: simplyblock-cluster ---- -# The CSI driver. The chart used to render the node DaemonSet, the controller -# StatefulSet, their RBAC, the node ConfigMaps, and the CSIDriver registration -# directly. It renders none of them now, and this object is what the operator -# turns into all eighteen. -# -# It lives in the operator's namespace rather than beside the StorageCluster -# above. The plugins mount simplyblock-csi-secret-v2, which the StorageCluster -# reconciler writes into the operator's namespace, so a driver deployed anywhere -# else comes up with no credentials to reach the control plane with. -# -# The name is "simplyblock" because the operator derives every object it applies -# as -csi-. That reproduces the names the chart wrote literally, -# which is what lets an already-running deployment be taken over in place rather -# than deleted and rebuilt. -# -# It names no image on purpose. Unset takes the operator's own registry and tag -# with the CSI driver's repository, so the driver follows the operator -# reconciling it and an operator upgrade carries it forward without a second -# edit here. That needs the chart to have set SB_OPERATOR_IMAGE on the -# operator's deployment; an operator deployed before that was added reports -# NoImage and wants spec.image stated. -apiVersion: storage.simplyblock.io/v1alpha2 -kind: SimplyblockDriver -metadata: - name: simplyblock - namespace: simplyblock -spec: {} diff --git a/test-cluster-okd.yaml b/test-cluster-okd.yaml deleted file mode 100644 index be8b38a19..000000000 --- a/test-cluster-okd.yaml +++ /dev/null @@ -1,93 +0,0 @@ ---- -apiVersion: storage.simplyblock.io/v1alpha1 -kind: StorageCluster -metadata: - name: simplyblock-cluster - namespace: default -spec: - isSingleNode: false - enableNodeAffinity: false - strictNodeAntiAffinity: false - volumeMigrationSettings: - # enabled turns on volume migration for this cluster (defaults to true). - enabled: true - # rebalancerImage is used for the migration path-validation Job (must include nvme-cli). - #rebalancerImage: docker.io/simplyblock/simplyblock-rebalancer:volume-autorebalancing - #autoRebalancing: - #enabled: true - #latencyBenchmarkEnabled: true - #migrationEnabled: true - #prometheusURL: http://simplyblock-prometheus.simplyblock:9090 - warningThreshold: - capacity: 80 - provisionedCapacity: 100 - criticalThreshold: - capacity: 90 - provisionedCapacity: 100 ---- -apiVersion: storage.simplyblock.io/v1alpha1 -kind: StorageNodeSet -metadata: - name: simplyblock-node - namespace: default -spec: - pcieAllowList: - - 0000:02:00.0 - - 0000:03:00.0 - maxParallelNodeAdds: 1 - #spdkProxyImage: "public.ecr.aws/simply-block/simplyblock:main-2774-bd790547" - clusterName: simplyblock-cluster - maxLogicalVolumeCount: 10 - partitions: 1 - corePercentage: 50 - skipKubeletConfiguration: false - enableCpuTopology: true - openShiftCluster: true - forceFormat4K: true - nodesPerSocket: 1 - socketsToUse: - - "0" - mgmtIfname: br-ex - dataIfname: - - ens16np0 - workerNodes: - - worker-1 - - worker-2 - - worker-3 - - worker-4 ---- -apiVersion: storage.simplyblock.io/v1alpha1 -kind: Pool -metadata: - name: pool1 - namespace: default -spec: - clusterName: simplyblock-cluster ---- -# The CSI driver. The chart used to render the node DaemonSet, the controller -# StatefulSet, their RBAC, the node ConfigMaps, and the CSIDriver registration -# directly. It renders none of them now, and this object is what the operator -# turns into all eighteen. -# -# It lives in the operator's namespace rather than beside the StorageCluster -# above. The plugins mount simplyblock-csi-secret-v2, which the StorageCluster -# reconciler writes into the operator's namespace, so a driver deployed anywhere -# else comes up with no credentials to reach the control plane with. -# -# The name is "simplyblock" because the operator derives every object it applies -# as -csi-. That reproduces the names the chart wrote literally, -# which is what lets an already-running deployment be taken over in place rather -# than deleted and rebuilt. -# -# It names no image on purpose. Unset takes the operator's own registry and tag -# with the CSI driver's repository, so the driver follows the operator -# reconciling it and an operator upgrade carries it forward without a second -# edit here. That needs the chart to have set SB_OPERATOR_IMAGE on the -# operator's deployment; an operator deployed before that was added reports -# NoImage and wants spec.image stated. -apiVersion: storage.simplyblock.io/v1alpha2 -kind: SimplyblockDriver -metadata: - name: simplyblock - namespace: simplyblock -spec: {} diff --git a/test-cluster.yaml b/test-cluster.yaml deleted file mode 100644 index e09d563e5..000000000 --- a/test-cluster.yaml +++ /dev/null @@ -1,91 +0,0 @@ ---- -apiVersion: storage.simplyblock.io/v1alpha1 -kind: StorageCluster -metadata: - name: simplyblock-cluster - namespace: default -spec: - isSingleNode: false - enableNodeAffinity: false - strictNodeAntiAffinity: false - volumeMigrationSettings: - # enabled turns on volume migration for this cluster (defaults to true). - enabled: true - # rebalancerImage is used for the migration path-validation Job (must include nvme-cli). - rebalancerImage: docker.io/simplyblock/simplyblock-rebalancer:volume-autorebalancing - #autoRebalancing: - # enabled: true - # latencyBenchmarkEnabled: true - # migrationEnabled: true - # prometheusURL: http://simplyblock-prometheus.simplyblock:9090 - warningThreshold: - capacity: 80 - provisionedCapacity: 100 - criticalThreshold: - capacity: 90 - provisionedCapacity: 100 ---- -apiVersion: storage.simplyblock.io/v1alpha1 -kind: StorageNodeSet -metadata: - name: simplyblock-node - namespace: default -spec: - clusterName: simplyblock-cluster - maxLogicalVolumeCount: 10 - partitions: 1 - corePercentage: 50 - skipKubeletConfiguration: false - enableCpuTopology: true - openShiftCluster: true - forceFormat4K: true - nodesPerSocket: 1 - socketsToUse: - - "0" - mgmtIfname: br-ex - dataIfname: - - enp2s0f0 - # - enp2s0f1 - workerNodes: - - worker-0.ocp.simplyblock.ai - - worker-1.ocp.simplyblock.ai - - worker-2.ocp.simplyblock.ai -# - worker-3.ocp.simplyblock.ai -# - worker-4.ocp.simplyblock.ai -# - worker-5.ocp.simplyblock.ai ---- -apiVersion: storage.simplyblock.io/v1alpha1 -kind: Pool -metadata: - name: pool1 - namespace: default -spec: - clusterName: simplyblock-cluster ---- -# The CSI driver. The chart used to render the node DaemonSet, the controller -# StatefulSet, their RBAC, the node ConfigMaps, and the CSIDriver registration -# directly. It renders none of them now, and this object is what the operator -# turns into all eighteen. -# -# It lives in the operator's namespace rather than beside the StorageCluster -# above. The plugins mount simplyblock-csi-secret-v2, which the StorageCluster -# reconciler writes into the operator's namespace, so a driver deployed anywhere -# else comes up with no credentials to reach the control plane with. -# -# The name is "simplyblock" because the operator derives every object it applies -# as -csi-. That reproduces the names the chart wrote literally, -# which is what lets an already-running deployment be taken over in place rather -# than deleted and rebuilt. -# -# It names no image on purpose. Unset takes the operator's own registry and tag -# with the CSI driver's repository, so the driver follows the operator -# reconciling it and an operator upgrade carries it forward without a second -# edit here. That needs the chart to have set SB_OPERATOR_IMAGE on the -# operator's deployment; an operator deployed before that was added reports -# NoImage and wants spec.image stated. -apiVersion: storage.simplyblock.io/v1alpha2 -kind: SimplyblockDriver -metadata: - name: simplyblock - namespace: simplyblock -spec: {} diff --git a/test-collection.yaml b/test-collection.yaml deleted file mode 100644 index 2da521ae3..000000000 --- a/test-collection.yaml +++ /dev/null @@ -1,614 +0,0 @@ -apiVersion: v1 -data: - report.json: |- - { - "version": 3, - "node": "vm03.simplyblock4.localdomain", - "probedAt": "2026-09-16T15:25:42Z", - "cpu": { - "onlineCPUs": 12, - "physicalCores": 12, - "sockets": 1, - "threadsPerCore": 1, - "hyperThreading": false, - "numaNodes": [ - { - "node": 0, - "onlineCPUs": [ - 0, - 1, - 2, - 3, - 4, - 5, - 6, - 7, - 8, - 9, - 10, - 11 - ], - "physicalCores": 12 - } - ] - }, - "memory": { - "totalBytes": 20700946432, - "freeBytes": 3447087104, - "availableBytes": 11525099520, - "hugePagesBytes": 7516192768, - "swapTotalBytes": 3221221376, - "swapFreeBytes": 2874433536, - "numaNodes": [ - { - "node": 0, - "totalBytes": 20700946432, - "freeBytes": 3447087104 - } - ] - }, - "hugePages": [ - { - "sizeBytes": 2097152, - "total": 3584, - "free": 3584, - "numaNodes": [ - { - "node": 0, - "total": 3584, - "free": 3584 - } - ] - }, - { - "sizeBytes": 1073741824, - "total": 0, - "free": 0, - "numaNodes": [ - { - "node": 0, - "total": 0, - "free": 0 - } - ] - } - ], - "interfaces": [ - { - "name": "cni0", - "macAddress": "22:75:52:54:80:bf", - "speedMbps": 10000, - "mtu": 1450, - "state": "up", - "numaNode": -1, - "virtual": true, - "bridge": true, - "addresses": [ - "10.42.2.1" - ] - }, - { - "name": "eth0", - "macAddress": "bc:24:11:15:e6:15", - "mtu": 1500, - "state": "up", - "driver": "virtio_net", - "pciAddress": "0000:06:12.0", - "numaNode": -1, - "addresses": [ - "192.168.10.113" - ] - }, - { - "name": "eth1", - "macAddress": "c2:39:fd:5f:2e:46", - "speedMbps": 40000, - "mtu": 9000, - "state": "up", - "driver": "mlx4_core", - "pciAddress": "0000:06:10.0", - "numaNode": -1, - "addresses": [ - "10.10.10.113" - ] - }, - { - "name": "flannel.1", - "macAddress": "e6:62:2e:90:75:b3", - "mtu": 1450, - "state": "unknown", - "numaNode": -1, - "virtual": true, - "addresses": [ - "10.42.2.0" - ] - }, - { - "name": "lo", - "macAddress": "00:00:00:00:00:00", - "mtu": 65536, - "state": "unknown", - "numaNode": -1, - "virtual": true, - "loopback": true, - "addresses": [ - "127.0.0.1" - ] - }, - { - "name": "veth0aadfdd7", - "macAddress": "e6:47:4e:18:29:3e", - "speedMbps": 10000, - "mtu": 1450, - "state": "up", - "numaNode": -1, - "virtual": true - }, - { - "name": "veth9c4bb771", - "macAddress": "6e:0d:66:c9:1c:e9", - "speedMbps": 10000, - "mtu": 1450, - "state": "up", - "numaNode": -1, - "virtual": true - } - ], - "devices": [ - { - "name": "dm-0", - "path": "/dev/dm-0", - "sizeBytes": 27913093120, - "kind": "DeviceMapper", - "rotational": true, - "numaNode": -1, - "available": false, - "rejections": [ - { - "reason": "NotAWholeDisk", - "detail": "the device is a DeviceMapper" - }, - { - "reason": "Mounted", - "detail": "mounted at [/]" - }, - { - "reason": "Busy", - "detail": "the kernel refused an exclusive open, and gives no reason" - } - ] - }, - { - "name": "dm-1", - "path": "/dev/dm-1", - "sizeBytes": 3221225472, - "kind": "DeviceMapper", - "rotational": true, - "numaNode": -1, - "available": false, - "rejections": [ - { - "reason": "NotAWholeDisk", - "detail": "the device is a DeviceMapper" - }, - { - "reason": "SwapArea", - "detail": "the device is an active swap area" - }, - { - "reason": "Busy", - "detail": "the kernel refused an exclusive open, and gives no reason" - } - ] - }, - { - "name": "nbd0", - "path": "/dev/nbd0", - "sizeBytes": 0, - "kind": "Network", - "numaNode": -1, - "available": false, - "rejections": [ - { - "reason": "NotAWholeDisk", - "detail": "the device is a Network" - }, - { - "reason": "NoCapacity", - "detail": "the device reports a size of zero" - } - ] - }, - { - "name": "nbd1", - "path": "/dev/nbd1", - "sizeBytes": 0, - "kind": "Network", - "numaNode": -1, - "available": false, - "rejections": [ - { - "reason": "NotAWholeDisk", - "detail": "the device is a Network" - }, - { - "reason": "NoCapacity", - "detail": "the device reports a size of zero" - } - ] - }, - { - "name": "nbd10", - "path": "/dev/nbd10", - "sizeBytes": 0, - "kind": "Network", - "numaNode": -1, - "available": false, - "rejections": [ - { - "reason": "NotAWholeDisk", - "detail": "the device is a Network" - }, - { - "reason": "NoCapacity", - "detail": "the device reports a size of zero" - } - ] - }, - { - "name": "nbd11", - "path": "/dev/nbd11", - "sizeBytes": 0, - "kind": "Network", - "numaNode": -1, - "available": false, - "rejections": [ - { - "reason": "NotAWholeDisk", - "detail": "the device is a Network" - }, - { - "reason": "NoCapacity", - "detail": "the device reports a size of zero" - } - ] - }, - { - "name": "nbd12", - "path": "/dev/nbd12", - "sizeBytes": 0, - "kind": "Network", - "numaNode": -1, - "available": false, - "rejections": [ - { - "reason": "NotAWholeDisk", - "detail": "the device is a Network" - }, - { - "reason": "NoCapacity", - "detail": "the device reports a size of zero" - } - ] - }, - { - "name": "nbd13", - "path": "/dev/nbd13", - "sizeBytes": 0, - "kind": "Network", - "numaNode": -1, - "available": false, - "rejections": [ - { - "reason": "NotAWholeDisk", - "detail": "the device is a Network" - }, - { - "reason": "NoCapacity", - "detail": "the device reports a size of zero" - } - ] - }, - { - "name": "nbd14", - "path": "/dev/nbd14", - "sizeBytes": 0, - "kind": "Network", - "numaNode": -1, - "available": false, - "rejections": [ - { - "reason": "NotAWholeDisk", - "detail": "the device is a Network" - }, - { - "reason": "NoCapacity", - "detail": "the device reports a size of zero" - } - ] - }, - { - "name": "nbd15", - "path": "/dev/nbd15", - "sizeBytes": 0, - "kind": "Network", - "numaNode": -1, - "available": false, - "rejections": [ - { - "reason": "NotAWholeDisk", - "detail": "the device is a Network" - }, - { - "reason": "NoCapacity", - "detail": "the device reports a size of zero" - } - ] - }, - { - "name": "nbd2", - "path": "/dev/nbd2", - "sizeBytes": 0, - "kind": "Network", - "numaNode": -1, - "available": false, - "rejections": [ - { - "reason": "NotAWholeDisk", - "detail": "the device is a Network" - }, - { - "reason": "NoCapacity", - "detail": "the device reports a size of zero" - } - ] - }, - { - "name": "nbd3", - "path": "/dev/nbd3", - "sizeBytes": 0, - "kind": "Network", - "numaNode": -1, - "available": false, - "rejections": [ - { - "reason": "NotAWholeDisk", - "detail": "the device is a Network" - }, - { - "reason": "NoCapacity", - "detail": "the device reports a size of zero" - } - ] - }, - { - "name": "nbd4", - "path": "/dev/nbd4", - "sizeBytes": 0, - "kind": "Network", - "numaNode": -1, - "available": false, - "rejections": [ - { - "reason": "NotAWholeDisk", - "detail": "the device is a Network" - }, - { - "reason": "NoCapacity", - "detail": "the device reports a size of zero" - } - ] - }, - { - "name": "nbd5", - "path": "/dev/nbd5", - "sizeBytes": 0, - "kind": "Network", - "numaNode": -1, - "available": false, - "rejections": [ - { - "reason": "NotAWholeDisk", - "detail": "the device is a Network" - }, - { - "reason": "NoCapacity", - "detail": "the device reports a size of zero" - } - ] - }, - { - "name": "nbd6", - "path": "/dev/nbd6", - "sizeBytes": 0, - "kind": "Network", - "numaNode": -1, - "available": false, - "rejections": [ - { - "reason": "NotAWholeDisk", - "detail": "the device is a Network" - }, - { - "reason": "NoCapacity", - "detail": "the device reports a size of zero" - } - ] - }, - { - "name": "nbd7", - "path": "/dev/nbd7", - "sizeBytes": 0, - "kind": "Network", - "numaNode": -1, - "available": false, - "rejections": [ - { - "reason": "NotAWholeDisk", - "detail": "the device is a Network" - }, - { - "reason": "NoCapacity", - "detail": "the device reports a size of zero" - } - ] - }, - { - "name": "nbd8", - "path": "/dev/nbd8", - "sizeBytes": 0, - "kind": "Network", - "numaNode": -1, - "available": false, - "rejections": [ - { - "reason": "NotAWholeDisk", - "detail": "the device is a Network" - }, - { - "reason": "NoCapacity", - "detail": "the device reports a size of zero" - } - ] - }, - { - "name": "nbd9", - "path": "/dev/nbd9", - "sizeBytes": 0, - "kind": "Network", - "numaNode": -1, - "available": false, - "rejections": [ - { - "reason": "NotAWholeDisk", - "detail": "the device is a Network" - }, - { - "reason": "NoCapacity", - "detail": "the device reports a size of zero" - } - ] - }, - { - "name": "sda", - "path": "/dev/sda", - "pciAddress": "0000:09:01.0", - "sizeBytes": 32212254720, - "kind": "Disk", - "transport": "Virtio", - "vendor": "QEMU", - "model": "QEMU HARDDISK", - "rotational": true, - "numaNode": -1, - "available": false, - "rejections": [ - { - "reason": "Mounted", - "detail": "mounted at [/boot]" - }, - { - "reason": "Busy", - "detail": "the kernel refused an exclusive open, and gives no reason" - }, - { - "reason": "Partitioned", - "detail": "the device carries the partitions [sda1 sda2]" - } - ] - }, - { - "name": "sda1", - "path": "/dev/sda1", - "pciAddress": "0000:09:01.0", - "sizeBytes": 1073741824, - "kind": "Partition", - "transport": "Virtio", - "numaNode": -1, - "available": false, - "rejections": [ - { - "reason": "NotAWholeDisk", - "detail": "the device is a Partition" - }, - { - "reason": "Mounted", - "detail": "mounted at [/boot]" - }, - { - "reason": "Busy", - "detail": "the kernel refused an exclusive open, and gives no reason" - } - ] - }, - { - "name": "sda2", - "path": "/dev/sda2", - "pciAddress": "0000:09:01.0", - "sizeBytes": 31137464320, - "kind": "Partition", - "transport": "Virtio", - "numaNode": -1, - "available": false, - "rejections": [ - { - "reason": "NotAWholeDisk", - "detail": "the device is a Partition" - }, - { - "reason": "Stacked", - "detail": "held by [dm-0 dm-1]" - }, - { - "reason": "Busy", - "detail": "the kernel refused an exclusive open, and gives no reason" - } - ] - } - ], - "nvmeControllers": [ - { - "address": "0000:00:02.0", - "driver": "uio_pci_generic", - "vendor": "0x1b36", - "product": "0x0010", - "numaNode": -1 - }, - { - "address": "0000:00:03.0", - "driver": "uio_pci_generic", - "vendor": "0x1b36", - "product": "0x0010", - "numaNode": -1 - }, - { - "address": "0000:00:04.0", - "driver": "uio_pci_generic", - "vendor": "0x1b36", - "product": "0x0010", - "numaNode": -1 - }, - { - "address": "0000:00:05.0", - "driver": "uio_pci_generic", - "vendor": "0x1b36", - "product": "0x0010", - "numaNode": -1 - } - ] - } -kind: ConfigMap -metadata: - creationTimestamp: "2026-09-16T15:25:42Z" - labels: - storage.simplyblock.io/component: nodeprobe - storage.simplyblock.io/nodeprobe-node: vm03.simplyblock4.localdomain - storage.simplyblock.io/nodeprobe-run: initial-discovery - name: sb-nodeprobe-initial-discovery-vm03.simplyblock4.local-74c6caeb - namespace: simplyblock - ownerReferences: - - apiVersion: storage.simplyblock.io/v1alpha2 - kind: OperatorOps - name: initial-discovery - uid: 47f9af07-a7f5-4f34-8c0c-fb94a77afefa - resourceVersion: "15100415" - uid: b81f1d34-81a6-4bae-8a7c-e3243530528d diff --git a/test-deployment-config.yaml b/test-deployment-config.yaml deleted file mode 100644 index c5343e558..000000000 --- a/test-deployment-config.yaml +++ /dev/null @@ -1,32 +0,0 @@ -apiVersion: storage.simplyblock.io/v1alpha2 -kind: ClusterDeploymentConfig -metadata: - creationTimestamp: "2026-09-15T13:55:35Z" - generation: 1 - labels: - storage.simplyblock.io/component: nodeprobe - storage.simplyblock.io/nodeprobe-run: initial-discovery - name: discovered-initial-discovery - namespace: simplyblock - resourceVersion: "14949006" - uid: 4c99c053-c94e-44c2-88bb-e76bf37c3357 -spec: - cluster: - maxSubsystemCount: 30 - name: discovered-initial-discovery-cluster - vcpuCount: 4 - environment: K3s - nodeSets: - - groups: - - devices: - nvme: - - "0000:00:02.0" - - "0000:00:03.0" - - "0000:00:04.0" - - "0000:00:05.0" - name: group-1-nvme-4xunsized - workers: - - vm02.simplyblock4.localdomain - - vm03.simplyblock4.localdomain - - vm04.simplyblock4.localdomain - name: discovered diff --git a/test-discovery-config.yaml b/test-discovery-config.yaml deleted file mode 100644 index 8911f06c3..000000000 --- a/test-discovery-config.yaml +++ /dev/null @@ -1,47 +0,0 @@ -apiVersion: storage.simplyblock.io/v1alpha2 -kind: ClusterDeploymentConfig -metadata: - creationTimestamp: "2026-09-15T15:23:38Z" - generation: 2 - labels: - storage.simplyblock.io/component: nodeprobe - storage.simplyblock.io/nodeprobe-run: initial-discovery - storage.simplyblock.io/ready-to-deploy: "true" - name: discovered-initial-discovery - namespace: simplyblock - resourceVersion: "14960753" - uid: 69cdac9f-af1e-4ded-907a-fdc36b8485ae -spec: - approved: true - cluster: - maxSubsystemCount: 30 - name: discovered-initial-discovery-cluster - vcpuCount: 4 - environment: K3s - nodeSets: - - groups: - - devices: - nvme: - - "0000:00:02.0" - - "0000:00:03.0" - - "0000:00:04.0" - - "0000:00:05.0" - mgmtInterface: eth0 - name: group-1-nvme-4x54.12G - workers: - - vm02.simplyblock4.localdomain - - vm03.simplyblock4.localdomain - - vm04.simplyblock4.localdomain - name: discovered -status: - clusterRef: discovered-initial-discovery-cluster - message: expanded into cluster discovered-initial-discovery-cluster and 3 node(s) - nodeRefs: - - discovered-initial-discovery-cluster-vm02.simplyblock4.localdomain-0 - - discovered-initial-discovery-cluster-vm03.simplyblock4.localdomain-0 - - discovered-initial-discovery-cluster-vm04.simplyblock4.localdomain-0 - observedGeneration: 2 - phase: Expanded - step: - deadline: "2026-09-15T15:37:47Z" - state: CreatingNodes diff --git a/test-migration.yaml b/test-migration.yaml deleted file mode 100644 index d14a00c92..000000000 --- a/test-migration.yaml +++ /dev/null @@ -1,11 +0,0 @@ ---- -apiVersion: storage.simplyblock.io/v1alpha1 -kind: VolumeMigration -metadata: - name: test-migration - namespace: default -spec: - # PV name whose backing logical volume should be migrated. - pvName: pvc-968cff4f-a199-4964-88f0-7cfccb5251d9 - # UUID of the destination storage node. - targetNodeUUID: 4e53efdd-86c9-424f-940c-e437eb6a2e95 diff --git a/test-new-crd.yaml b/test-new-crd.yaml deleted file mode 100644 index aa4c5de49..000000000 --- a/test-new-crd.yaml +++ /dev/null @@ -1,15 +0,0 @@ ---- -apiVersion: storage.simplyblock.io/v1alpha2 -kind: StoragePool -metadata: - name: pool1 - namespace: default -spec: - clusterRef: simplyblock-cluster ---- -apiVersion: storage.simplyblock.io/v1alpha2 -kind: SimplyblockDriver -metadata: - name: simplyblock - namespace: simplyblock -spec: {} diff --git a/test-node-migration.yaml b/test-node-migration.yaml deleted file mode 100644 index 1f944bab3..000000000 --- a/test-node-migration.yaml +++ /dev/null @@ -1,32 +0,0 @@ -# Migrate storage node 314c90cb-1da5-42f0-9f6c-cab4b9d220cc from worker-2 to worker-4. -# -# A migration is NOT a drain: the node keeps its UUID and its partitions / logical -# volume assignments follow it. It is a storage-node restart pointed at the target -# host's storage-node-api pod (node_address). -# -# The StorageNodeOpsReconciler (action=migrate): -# 1. validates targetWorkerNode exists and is Ready -# 2. labels the target worker (io.simplyblock.node-type=...) so the storage-node -# DaemonSet schedules a pod there, then waits for that pod to be Ready -# 3. POST /clusters/{c}/storage-nodes/{uuid}/restart -# { force, reattach_volume, node_address: } -# 4. polls until the node reports online on worker-4 -# 5. re-points the StorageNode CR's spec.workerNode to worker-4 and swaps the -# owning StorageNodeSet.spec.workerNodes (worker-2 -> worker-4), then Succeeds -# -# spec.workerNode is only mutable by the operator: the StorageNode validating -# webhook rejects user-driven changes and steers them to this ops instead. -# -# kubectl apply -f test-node-migration.yaml -# kubectl get snops,storagenodes -n default -w -apiVersion: storage.simplyblock.io/v1alpha1 -kind: StorageNodeOps -metadata: - name: migrate-314c90cb-to-worker-4 - namespace: default -spec: - storageNodeRef: simplyblock-node-jqameh - action: migrate - targetWorkerNode: worker-4 - # reattachVolume: true # optional — reattach volumes during the restart - # force: true # optional — force the restart where the backend allows it diff --git a/test-pod.yaml b/test-pod.yaml deleted file mode 100644 index 7bb2c76a7..000000000 --- a/test-pod.yaml +++ /dev/null @@ -1,37 +0,0 @@ ---- -kind: PersistentVolumeClaim -apiVersion: v1 -metadata: - name: spdkcsi-pvc2 - annotations: - # Pin this volume's backing logical volume to a specific storage node. - # The value must be a valid storage-node UUID in the volume's cluster; the - # validating webhook rejects an unknown UUID at write time. Changing the - # value triggers a VolumeMigration onto the new node. - # Replace with a real storage-node UUID from your cluster. - simplyblock.io/pinned-volume: "5cbe989e-ff32-4d65-a0c2-9250c42a3112" -spec: - accessModes: - - ReadWriteOnce - resources: - requests: - storage: 10Gi - storageClassName: simplyblock-default-simplyblock-cluster-pool1-xfs ---- -kind: Pod -apiVersion: v1 -metadata: - name: spdkcsi-test2 -spec: - containers: - - name: alpine - image: alpine:3 - imagePullPolicy: "IfNotPresent" - command: ["sleep", "365d"] - volumeMounts: - - mountPath: "/spdkvol" - name: spdk-volume - volumes: - - name: spdk-volume - persistentVolumeClaim: - claimName: spdkcsi-pvc2 From 04b2f177d2e798e65b752638319174dccc3592a9 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Fri, 18 Sep 2026 22:10:31 +0200 Subject: [PATCH 086/206] feat(deployment): a draft is told its fleet is too small for its stripe, twice The scheme and the node count it needs are the one part of a deployment nothing below this operator checks. The control plane validates the scheme on the cluster create, which lands after the approval, and its activation gate counts devices rather than nodes, so a document naming 4+2 over four workers is admitted by everything and produces a cluster that loses data on the second failure it was configured to survive. Both moments, for the same reason the rest of this document's validation has two. A draft is told on every reconcile, while it is still editable and while a reviewer can act on it, which is the whole point of the gate. The approving edit is refused, because an approved document is immutable and telling somebody afterwards is telling them about something they can no longer change. The count is the nodes the document itself describes, which is what makes the answer available before any StorageNode exists: a draft states its workers and its slots, so the fleet it would produce is knowable from the document alone. StripeUnsupported and StripeBelowMinimumNodes are separate reasons because they ask for different corrections. The first is a scheme nothing supports and the edit is the scheme. The second is a scheme this fleet cannot carry, and the edit is either the scheme or the fleet. Co-Authored-By: Claude Opus 5 (1M context) From 4337197a221da69d5bf584051fd0b5aa4069aa0d Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Sat, 19 Sep 2026 18:20:52 +0200 Subject: [PATCH 087/206] fix(operator): the two gates develop has been failing unseen Neither is this branch's. A push to develop runs only the image builds, which take `branches: ["**"]`, and every gate below takes `branches: [main]` plus `pull_request`, so nothing has run them on develop since it diverged. They surface on the first pull request opened from it. k8s.io/utils became a direct dependency when the webhook tests took k8s.io/utils/ptr in 032d58cc, and go.mod still lists it as indirect. PersistentVolumeOps and StorageDeviceOps have CRD bases and no entry in the CSV's owned list, which leaves them out of the Provided APIs panel on the install page. The file is hand-maintained rather than generated, so nothing added them with the types. Co-Authored-By: Claude Opus 5 (1M context) --- ...implyblock-operator.clusterserviceversion.yaml | 15 +++++++++++++++ operator/go.mod | 2 +- 2 files changed, 16 insertions(+), 1 deletion(-) diff --git a/operator/config/manifests/bases/simplyblock-operator.clusterserviceversion.yaml b/operator/config/manifests/bases/simplyblock-operator.clusterserviceversion.yaml index 7134d8db0..802b65438 100644 --- a/operator/config/manifests/bases/simplyblock-operator.clusterserviceversion.yaml +++ b/operator/config/manifests/bases/simplyblock-operator.clusterserviceversion.yaml @@ -220,6 +220,16 @@ spec: kind: SimplyblockDriver name: simplyblockdrivers.storage.simplyblock.io version: v1alpha2 + - description: PersistentVolumeOps is a single operation performed against one PersistentVolume. + It is the one Ops kind in this group whose target is a core Kubernetes type + rather than a kind this group defines, so it locks its target with an annotation + rather than a status field, is cluster-scoped because its target is, cannot + be owned by the namespaced operation that created it, and derives its cluster, + pool, and volume from the volume's CSI handle rather than being told. + displayName: Persistent Volume Ops + kind: PersistentVolumeOps + name: persistentvolumeops.storage.simplyblock.io + version: v1alpha2 - description: StorageBackupOps is a single operation performed against one StorageBackup. It runs to a terminal phase and stays afterward as the audit record of what was restored, into which pool, and how it ended. The claim a restore produces @@ -246,6 +256,11 @@ spec: kind: StorageDevice name: storagedevices.storage.simplyblock.io version: v1alpha2 + - description: StorageDeviceOps is a single operation performed against one StorageDevice. + displayName: Storage Device Ops + kind: StorageDeviceOps + name: storagedeviceops.storage.simplyblock.io + version: v1alpha2 - description: StoragePoolOps is a single operation performed against one StoragePool. Analogous to a Kubernetes Job, it drives an action to completion and records the result, and only one may be active per pool at a time, which the pool's diff --git a/operator/go.mod b/operator/go.mod index 8e8dd0dab..df4e93a08 100644 --- a/operator/go.mod +++ b/operator/go.mod @@ -29,6 +29,7 @@ require ( k8s.io/client-go v0.36.2 k8s.io/component-base v0.36.2 k8s.io/kube-openapi v0.0.0-20260317180543-43fb72c5454a + k8s.io/utils v0.0.0-20260210185600-b8788abfbbc2 sigs.k8s.io/controller-runtime v0.24.1 sigs.k8s.io/yaml v1.6.0 ) @@ -196,7 +197,6 @@ require ( k8s.io/kms v0.36.2 // indirect k8s.io/kubectl v0.36.2 // indirect k8s.io/streaming v0.36.2 // indirect - k8s.io/utils v0.0.0-20260210185600-b8788abfbbc2 // indirect oras.land/oras-go/v2 v2.6.1 // indirect sigs.k8s.io/apiserver-network-proxy/konnectivity-client v0.34.0 // indirect sigs.k8s.io/json v0.0.0-20250730193827-2d320260d730 // indirect From 80ed261a9a2dfdc3a72c7c06047e98e0c1e6c577 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Sat, 19 Sep 2026 18:58:51 +0200 Subject: [PATCH 088/206] fix(atlas-lib): a captured tree is read without the host it is read on sysfs carries no IP addresses, so the interface reader takes them from a seam whose default answers from this process's own network namespace. The tests named a root and left the seam alone, which meant every reading they asserted against was the fixture plus whatever the machine running it had configured. It is silent where it is wrong. The fixture's interfaces carry the ordinary names, so a host with an `eth0` or a `lo` of its own fills them in: the suite passed on a developer's machine, where neither exists, and failed on a runner that had both, on an address nothing in the fixture declares. Every test that reads a fixture now names the seam, through one helper, and a new case asserts the property directly so that a reading which picks the host up again fails on any machine that has an interface the fixture also names. Co-Authored-By: Claude Opus 5 (1M context) --- atlas-lib/inventory/fixture_test.go | 16 ++++++++++++ atlas-lib/inventory/netiface_test.go | 37 +++++++++++++++++++++++----- atlas-lib/inventory/netkind_test.go | 12 ++++----- 3 files changed, 53 insertions(+), 12 deletions(-) diff --git a/atlas-lib/inventory/fixture_test.go b/atlas-lib/inventory/fixture_test.go index cddfaa795..fb8e7c009 100644 --- a/atlas-lib/inventory/fixture_test.go +++ b/atlas-lib/inventory/fixture_test.go @@ -24,6 +24,22 @@ type fixture struct { dirs []string } +// syntheticHost reads the tree at root and nothing else. +// +// The address reader has to be named, because sysfs does not carry addresses +// and the default answers from this process's own network namespace. A test +// that left it out would assert a fixture against whatever the machine running +// it happens to have configured, and it would do so silently: the fixture's +// interface names are the ordinary ones, so any host with an `eth0` or a `lo` +// of its own fills them in. +func syntheticHost(root string) Config { + return Config{ + SysfsRoot: root, + ProcRoot: root, + InterfaceAddresses: func() (map[string][]string, error) { return nil, nil }, + } +} + // write materializes the fixture and returns its root. func (f fixture) write(t *testing.T) string { t.Helper() diff --git a/atlas-lib/inventory/netiface_test.go b/atlas-lib/inventory/netiface_test.go index dbe9cf465..50677135b 100644 --- a/atlas-lib/inventory/netiface_test.go +++ b/atlas-lib/inventory/netiface_test.go @@ -89,7 +89,7 @@ func byName(t *testing.T, ifaces []Interface, name string) Interface { func TestReadInterfacesReportsAPhysicalNICWhole(t *testing.T) { root := netHost().write(t) - ifaces, err := ReadInterfaces(Config{SysfsRoot: root, ProcRoot: root}) + ifaces, err := ReadInterfaces(syntheticHost(root)) if err != nil { t.Fatalf("read the interfaces: %v", err) } @@ -119,7 +119,7 @@ func TestReadInterfacesReportsNoSpeedForALinkThatIsDown(t *testing.T) { // nobody was going to use. root := netHost().write(t) - ifaces, err := ReadInterfaces(Config{SysfsRoot: root, ProcRoot: root}) + ifaces, err := ReadInterfaces(syntheticHost(root)) if err != nil { t.Fatalf("read the interfaces: %v", err) } @@ -142,7 +142,7 @@ func TestReadInterfacesReportsNoSpeedForALinkThatIsDown(t *testing.T) { func TestReadInterfacesMarksTheVirtualOnesAsVirtual(t *testing.T) { root := netHost().write(t) - ifaces, err := ReadInterfaces(Config{SysfsRoot: root, ProcRoot: root}) + ifaces, err := ReadInterfaces(syntheticHost(root)) if err != nil { t.Fatalf("read the interfaces: %v", err) } @@ -172,7 +172,7 @@ func TestReadInterfacesMarksTheVirtualOnesAsVirtual(t *testing.T) { func TestReadInterfacesIsOrderedByName(t *testing.T) { root := netHost().write(t) - ifaces, err := ReadInterfaces(Config{SysfsRoot: root, ProcRoot: root}) + ifaces, err := ReadInterfaces(syntheticHost(root)) if err != nil { t.Fatalf("read the interfaces: %v", err) } @@ -192,7 +192,7 @@ func TestReadInterfacesIsOrderedByName(t *testing.T) { func TestReadInterfacesReportsNoneRatherThanFailingWithoutTheClassDirectory(t *testing.T) { root := fixture{files: map[string]string{"meminfo": "MemTotal: 1024 kB"}}.write(t) - ifaces, err := ReadInterfaces(Config{SysfsRoot: root, ProcRoot: root}) + ifaces, err := ReadInterfaces(syntheticHost(root)) if err != nil { t.Fatalf("read the interfaces of a tree without class/net: %v", err) } @@ -216,7 +216,7 @@ func TestABridgeIsReadFromItsOwnDirectory(t *testing.T) { t.Fatal(err) } - ifaces, err := ReadInterfaces(Config{SysfsRoot: filepath.Join(root, "sys")}) + ifaces, err := ReadInterfaces(syntheticHost(filepath.Join(root, "sys"))) if err != nil { t.Fatalf("read the interfaces: %v", err) } @@ -289,3 +289,28 @@ func TestAFailedAddressReadStillReportsTheInterfaces(t *testing.T) { t.Errorf("addresses were invented: %v", ifaces[0].Addresses) } } + +// A reading of a captured tree carries only what the tree declares. +// +// sysfs holds no addresses, so the reader supplies them, and its default +// answers from this process's own network namespace. A test that reads a +// fixture without naming a reader therefore asserts against whatever the +// machine running it has configured, and it does so silently: the fixture's +// interfaces carry the ordinary names, so any host with an `eth0` or a `lo` of +// its own fills them in. That is how this suite passed on a developer's machine +// and failed on a runner, which had both. +func TestAFixtureTakesNoAddressesFromTheRunningHost(t *testing.T) { + ifaces, err := ReadInterfaces(syntheticHost(netHost().write(t))) + if err != nil { + t.Fatalf("read the interfaces: %v", err) + } + if len(ifaces) == 0 { + t.Fatal("the fixture produced no interfaces to check") + } + for _, iface := range ifaces { + if len(iface.Addresses) != 0 { + t.Errorf("%s carries %v, which the fixture does not declare and the host does", + iface.Name, iface.Addresses) + } + } +} diff --git a/atlas-lib/inventory/netkind_test.go b/atlas-lib/inventory/netkind_test.go index e5ec6af8f..c0eb1ae93 100644 --- a/atlas-lib/inventory/netkind_test.go +++ b/atlas-lib/inventory/netkind_test.go @@ -114,7 +114,7 @@ func stackedNetHost() fixture { func TestReadInterfacesNamesTheKindOfEachDevice(t *testing.T) { root := stackedNetHost().write(t) - ifaces, err := ReadInterfaces(Config{SysfsRoot: root, ProcRoot: root}) + ifaces, err := ReadInterfaces(syntheticHost(root)) if err != nil { t.Fatalf("read the interfaces: %v", err) } @@ -142,7 +142,7 @@ func TestReadInterfacesReportsTheMembersOfABondAndABridge(t *testing.T) { // own, so the members are the only route to the hardware underneath one. root := stackedNetHost().write(t) - ifaces, err := ReadInterfaces(Config{SysfsRoot: root, ProcRoot: root}) + ifaces, err := ReadInterfaces(syntheticHost(root)) if err != nil { t.Fatalf("read the interfaces: %v", err) } @@ -161,7 +161,7 @@ func TestReadInterfacesReportsTheMembersOfABondAndABridge(t *testing.T) { func TestReadInterfacesReportsTheParentOfADerivedDevice(t *testing.T) { root := stackedNetHost().write(t) - ifaces, err := ReadInterfaces(Config{SysfsRoot: root, ProcRoot: root}) + ifaces, err := ReadInterfaces(syntheticHost(root)) if err != nil { t.Fatalf("read the interfaces: %v", err) } @@ -176,7 +176,7 @@ func TestReadInterfacesReportsWhatIsStackedOnAnInterface(t *testing.T) { // might hold one. root := stackedNetHost().write(t) - ifaces, err := ReadInterfaces(Config{SysfsRoot: root, ProcRoot: root}) + ifaces, err := ReadInterfaces(syntheticHost(root)) if err != nil { t.Fatalf("read the interfaces: %v", err) } @@ -199,7 +199,7 @@ func TestReadInterfacesStillMarksEveryStackedDeviceVirtual(t *testing.T) { // keeps the answer it had. root := stackedNetHost().write(t) - ifaces, err := ReadInterfaces(Config{SysfsRoot: root, ProcRoot: root}) + ifaces, err := ReadInterfaces(syntheticHost(root)) if err != nil { t.Fatalf("read the interfaces: %v", err) } @@ -281,7 +281,7 @@ func kernelWithoutDeviceTypes() fixture { func TestReadInterfacesFallsBackToWhatTheDriverExports(t *testing.T) { root := kernelWithoutDeviceTypes().write(t) - ifaces, err := ReadInterfaces(Config{SysfsRoot: root, ProcRoot: root}) + ifaces, err := ReadInterfaces(syntheticHost(root)) if err != nil { t.Fatalf("read the interfaces: %v", err) } From f2f2dd34a223e23f642a1f1aa2b18493913f5fc2 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Sat, 19 Sep 2026 18:59:10 +0200 Subject: [PATCH 089/206] chore(operator): the fifteen lint findings the discovery work left All of them are in fixtures, and none changes what a fixture produces: the generator writes the same 179 cases byte for byte. Four parameters and two variadics were inert. Every caller passed the same memory node, the same cluster, or nothing at all, so the signatures said a case could vary something no case varies. They now say what is true, and a case that needs the axis back adds it with its first caller. Four helpers had no caller. `fourNVMe` was superseded by the one-node set beside it, and `held`, `unchecked`, and `labeled` were vocabulary written for a case that ended up expressed another way. Git holds them for the case that wants one. The repeated literals become constants. A fixture's interface name and the assertion against it have to be the same string, and a case that disagreed with its own fixture by a character would pass while asserting nothing. The cluster's own plumbing becomes a function rather than a slice. One case put a NIC beside it with `append`, which writes into the backing array every other case shares, and building a list per call removes both that and the preallocation finding. Co-Authored-By: Claude Opus 5 (1M context) --- operator/hack/discoveryfixtures/cases_dev.go | 53 ++++++++----------- operator/hack/discoveryfixtures/cases_held.go | 20 +++---- operator/hack/discoveryfixtures/cases_net.go | 17 ++++-- operator/hack/discoveryfixtures/cases_role.go | 12 ++--- operator/hack/discoveryfixtures/fixtures.go | 47 ++++++---------- .../controllers/cluster/erasurecoding_test.go | 8 ++- .../deployment/erasurecoding_test.go | 13 ++--- operator/internal/discovery/fixture_test.go | 7 +-- operator/internal/discovery/iscsi_test.go | 6 +-- operator/internal/discovery/mgmtiface_test.go | 38 ++++++++----- operator/internal/discovery/netstack_test.go | 46 ++++++++-------- operator/internal/discovery/plan_test.go | 8 +-- operator/internal/discovery/rules_test.go | 6 +-- .../internal/discovery/simplyblock_test.go | 6 +-- 14 files changed, 141 insertions(+), 146 deletions(-) diff --git a/operator/hack/discoveryfixtures/cases_dev.go b/operator/hack/discoveryfixtures/cases_dev.go index 0fc770694..99d273ece 100644 --- a/operator/hack/discoveryfixtures/cases_dev.go +++ b/operator/hack/discoveryfixtures/cases_dev.go @@ -25,17 +25,6 @@ func blockClass() *simplyblockv1alpha2.DiscoverSpec { } } -// fourNVMe is the disk set a case inherits when it is about something else: two -// disks on each memory node, all the same size. -func fourNVMe() []nodeprobe.Device { - return []nodeprobe.Device{ - nvme("nvme0n1", "0000:5e:00.0", 0, 3*tb), - nvme("nvme1n1", "0000:5f:00.0", 0, 3*tb), - nvme("nvme2n1", "0000:af:00.0", 1, 3*tb), - nvme("nvme3n1", "0000:b0:00.0", 1, 3*tb), - } -} - // oneNodeFourNVMe is the same four disks with every one of them on memory node // 0, for a case that is about the disks rather than about the placement. func oneNodeFourNVMe() []nodeprobe.Device { @@ -51,7 +40,7 @@ func oneNodeFourNVMe() []nodeprobe.Device { func virtioDisks(n int) []nodeprobe.Device { out := make([]nodeprobe.Device, 0, n) for i := 0; i < n; i++ { - out = append(out, blk(fmt.Sprintf("vd%c", 'b'+rune(i)), 0, 2*tb)) + out = append(out, blk(fmt.Sprintf("vd%c", 'b'+rune(i)), 2*tb)) } return out } @@ -88,8 +77,8 @@ func devCases() map[string]Case { Reports: []nodeprobe.Report{host("worker-01", single, disks( nvme("nvme0n1", "0000:5e:00.0", 0, 3*tb), nvme("nvme1n1", "0000:5f:00.0", 0, 3*tb), - blk("vdb", 0, 2*tb), - blk("vdc", 0, 2*tb), + blk("vdb", 2*tb), + blk("vdc", 2*tb), ))}, }, "DEV-05": { @@ -98,8 +87,8 @@ func devCases() map[string]Case { Reports: []nodeprobe.Report{host("worker-01", single, disks( nvme("nvme0n1", "0000:5e:00.0", 0, 3*tb), nvme("nvme1n1", "0000:5f:00.0", 0, 3*tb), - blk("vdb", 0, 2*tb), - blk("vdc", 0, 2*tb), + blk("vdb", 2*tb), + blk("vdc", 2*tb), ))}, }, "DEV-06": { @@ -115,7 +104,7 @@ func devCases() map[string]Case { Discover: blockClass(), Reports: []nodeprobe.Report{host("worker-01", single, disks( nvme("nvme0n1", "0000:5e:00.0", 0, 3*tb), - blk("sda", 0, 4*tb, transported(blockdev.TransportSATA)), + blk("sda", 4*tb, transported(blockdev.TransportSATA)), ))}, }, "DEV-08": { @@ -134,8 +123,8 @@ func devCases() map[string]Case { Reports: []nodeprobe.Report{host("worker-01", single, disks(append( oneNodeFourNVMe(), nvme("nvme0n1p1", "0000:5e:00.0", 0, 512*gb, partOf()), - blk("loop0", 0, 64*gb, looped()), - blk("loop1", 0, 64*gb, looped()), + blk("loop0", 64*gb, looped()), + blk("loop1", 64*gb, looped()), )...))}, }, "DEV-11": { @@ -153,7 +142,7 @@ func devCases() map[string]Case { "DEV-13": { Family: "dev", Slug: "an-attached-simplyblock-volume", Reports: []nodeprobe.Report{host("worker-01", single, disks( - attachedVolume("nvme3n1", ourCluster, "792e184c-0a1b-2c3d-4e5f-60718293a4b5"), + attachedVolume("nvme3n1", "792e184c-0a1b-2c3d-4e5f-60718293a4b5"), nvme("nvme0n1", "0000:5e:00.0", 0, 3*tb), ))}, }, @@ -161,16 +150,16 @@ func devCases() map[string]Case { Family: "dev", Slug: "an-attached-volume-on-a-block-run", Discover: blockClass(), Reports: []nodeprobe.Report{host("worker-01", single, disks( - attachedVolume("nvme3n1", ourCluster, "792e184c-0a1b-2c3d-4e5f-60718293a4b5"), - blk("vdb", 0, 2*tb), + attachedVolume("nvme3n1", "792e184c-0a1b-2c3d-4e5f-60718293a4b5"), + blk("vdb", 2*tb), ))}, }, "DEV-15": { Family: "dev", Slug: "every-disk-is-an-attached-volume", Reports: []nodeprobe.Report{host("worker-01", single, disks( - attachedVolume("nvme3n1", ourCluster, "792e184c-0a1b-2c3d-4e5f-60718293a4b5"), - attachedVolume("nvme4n1", ourCluster, "8a3f0b12-3c4d-5e6f-7081-92a3b4c5d6e7"), - attachedVolume("nvme5n1", ourCluster, "b1c2d3e4-f506-1728-394a-5b6c7d8e9f01"), + attachedVolume("nvme3n1", "792e184c-0a1b-2c3d-4e5f-60718293a4b5"), + attachedVolume("nvme4n1", "8a3f0b12-3c4d-5e6f-7081-92a3b4c5d6e7"), + attachedVolume("nvme5n1", "b1c2d3e4-f506-1728-394a-5b6c7d8e9f01"), ))}, }, "DEV-17": { @@ -178,7 +167,7 @@ func devCases() map[string]Case { Discover: blockClass(), Reports: []nodeprobe.Report{host("worker-01", single, disks( iscsiLUN("sdb", 2*tb), - blk("vdb", 0, 2*tb), + blk("vdb", 2*tb), ))}, }, "DEV-18": { @@ -191,7 +180,7 @@ func devCases() map[string]Case { }, Reports: []nodeprobe.Report{host("worker-01", single, disks( iscsiLUN("sdb", 2*tb), - blk("vdb", 0, 2*tb), + blk("vdb", 2*tb), ))}, }, "DEV-19": { @@ -230,17 +219,17 @@ const ourCluster = "c30a691a-1d2e-4f3a-9b8c-5d6e7f809a1b" // The probe refuses it for being on a fabric, which is what a real report // carries, and the rule that names it as this fleet's own reads the NQN rather // than the refusal. -func attachedVolume(name, clusterID, volumeID string) nodeprobe.Device { - device := blk(name, 0, tb, refused(blockdev.ReasonFabricNamespace)) +func attachedVolume(name, volumeID string) nodeprobe.Device { + device := blk(name, tb, refused(blockdev.ReasonFabricNamespace)) device.Transport = string(blockdev.TransportNVMeFabric) - device.SubsystemNQN = nqn.Make(clusterID, volumeID) + device.SubsystemNQN = nqn.Make(ourCluster, volumeID) return device } // iscsiLUN is a disk on the other side of a network, which the kernel presents // through the SCSI stack like any local disk. func iscsiLUN(name string, size uint64) nodeprobe.Device { - device := blk(name, 0, size) + device := blk(name, size) device.Transport = string(blockdev.TransportISCSI) return device } @@ -248,7 +237,7 @@ func iscsiLUN(name string, size uint64) nodeprobe.Device { // foreignVolume is a namespace something else exported, which is on a fabric // and is nobody's simplyblock volume. func foreignVolume(name string) nodeprobe.Device { - device := blk(name, 0, tb, refused(blockdev.ReasonFabricNamespace)) + device := blk(name, tb, refused(blockdev.ReasonFabricNamespace)) device.Transport = string(blockdev.TransportNVMeFabric) device.SubsystemNQN = "nqn.2019-08.org.ceph:rbd.pool.image" return device diff --git a/operator/hack/discoveryfixtures/cases_held.go b/operator/hack/discoveryfixtures/cases_held.go index 676bb0c4d..b717ca1a1 100644 --- a/operator/hack/discoveryfixtures/cases_held.go +++ b/operator/hack/discoveryfixtures/cases_held.go @@ -28,7 +28,7 @@ import ( func idleControllers(n int, driver string) []nodeprobe.Controller { out := make([]nodeprobe.Controller, 0, n) for i := 0; i < n; i++ { - out = append(out, controller(fmt.Sprintf("0000:%02x:00.0", 0x5e+i), driver, 0)) + out = append(out, controller(fmt.Sprintf("0000:%02x:00.0", 0x5e+i), driver)) } return out } @@ -69,7 +69,7 @@ func heldCases() map[string]Case { loopbacks := make([]nodeprobe.Device, 0, 16) for i := 0; i < 16; i++ { - loopbacks = append(loopbacks, blk(fmt.Sprintf("loop%d", i), 0, 64*gb, looped())) + loopbacks = append(loopbacks, blk(fmt.Sprintf("loop%d", i), 64*gb, looped())) } cases := map[string]Case{ @@ -84,20 +84,20 @@ func heldCases() map[string]Case { nvme("nvme1n1", "0000:b0:00.0", 0, 3*tb), ), controllers( - controller("0000:5e:00.0", pci.DriverUIOGeneric, 0), - controller("0000:5f:00.0", pci.DriverUIOGeneric, 0), - controller("0000:af:00.0", "nvme", 0), - controller("0000:b0:00.0", "nvme", 0), + controller("0000:5e:00.0", pci.DriverUIOGeneric), + controller("0000:5f:00.0", pci.DriverUIOGeneric), + controller("0000:af:00.0", "nvme"), + controller("0000:b0:00.0", "nvme"), ))), "HELD-06": one(host("worker-01", cpu(1, 16, 2), controllers(idleControllers(4, pci.DriverVFIO)...))), "HELD-07": one(host("worker-01", cpu(1, 16, 2), disks(layoutA.disks(3*tb)...), controllers( - controller("0000:5e:00.0", "nvme", 0), - controller("0000:5f:00.0", "nvme", 0), - controller("0000:af:00.0", "nvme", 0), - controller("0000:b0:00.0", "nvme", 0), + controller("0000:5e:00.0", "nvme"), + controller("0000:5f:00.0", "nvme"), + controller("0000:af:00.0", "nvme"), + controller("0000:b0:00.0", "nvme"), ))), "HELD-09": one(host("worker-01", cpu(1, 16, 2), disks(append(loopbacks, layoutA.disks(3*tb)...)...))), diff --git a/operator/hack/discoveryfixtures/cases_net.go b/operator/hack/discoveryfixtures/cases_net.go index 1378049fb..2e1b105cf 100644 --- a/operator/hack/discoveryfixtures/cases_net.go +++ b/operator/hack/discoveryfixtures/cases_net.go @@ -64,10 +64,17 @@ func bonded(addresses map[string][]string) []nodeprobe.Interface { func netCases() map[string]Case { // The cluster's own plumbing, which every worker of every Kubernetes fleet // carries and none of it is a management interface. - plumbing := []nodeprobe.Interface{ - nic("cni0", inventory.LinkBridge, holding("10.42.2.1")), - nic("flannel.1", inventory.LinkVXLAN, holding("10.42.2.0")), - nic("lo", inventory.LinkLoopback, holding("127.0.0.1")), + // + // It builds a list rather than being one, because the case that puts a NIC + // beside it would otherwise append into the slice every other case shares. + plumbing := func(beside ...nodeprobe.Interface) []nodeprobe.Interface { + out := make([]nodeprobe.Interface, 0, 3+len(beside)) + out = append(out, + nic("cni0", inventory.LinkBridge, holding("10.42.2.1")), + nic("flannel.1", inventory.LinkVXLAN, holding("10.42.2.0")), + nic("lo", inventory.LinkLoopback, holding("127.0.0.1")), + ) + return append(out, beside...) } veths := func(n int) []nodeprobe.Interface { @@ -255,7 +262,7 @@ func netCases() map[string]Case { // The cluster's own plumbing beside nothing else, which is the worker a // draft has to leave without an interface. - cases["NET-19"] = unknown(netHost("worker-01", append(plumbing, + cases["NET-19"] = unknown(netHost("worker-01", plumbing( nic("eth0", inventory.LinkPhysical, at(10000), slotted("0000:3b:00.0")))...)) // The seam case: a stack that points at itself cannot be written by the diff --git a/operator/hack/discoveryfixtures/cases_role.go b/operator/hack/discoveryfixtures/cases_role.go index d2cfcd475..c4d885e6e 100644 --- a/operator/hack/discoveryfixtures/cases_role.go +++ b/operator/hack/discoveryfixtures/cases_role.go @@ -149,12 +149,12 @@ func filterCases() map[string]Case { } blockWorker := func() []nodeprobe.Report { return []nodeprobe.Report{host("worker-01", cpu(1, 16, 2), disks( - blk("sda", 0, 512*gb), - blk("sdb", 0, 2*tb), - blk("sdc", 0, 2*tb), - blk("vdb", 0, 4*tb), - blk("vdc", 0, 4*tb), - blk("vdd", 0, 4*tb), + blk("sda", 512*gb), + blk("sdb", 2*tb), + blk("sdc", 2*tb), + blk("vdb", 4*tb), + blk("vdc", 4*tb), + blk("vdd", 4*tb), ))} } diff --git a/operator/hack/discoveryfixtures/fixtures.go b/operator/hack/discoveryfixtures/fixtures.go index 2066e21fe..a71b0b659 100644 --- a/operator/hack/discoveryfixtures/fixtures.go +++ b/operator/hack/discoveryfixtures/fixtures.go @@ -422,15 +422,16 @@ func nvme(name, address string, node int, size uint64, opts ...devOpt) nodeprobe } // blk is one free disk of the other class: a virtio disk with a path and no PCI -// address a draft could name it by. -func blk(name string, node int, size uint64, opts ...devOpt) nodeprobe.Device { +// address a draft could name it by. It sits on the first memory node, which is +// where every case that reaches for this class of disk puts it. +func blk(name string, size uint64, opts ...devOpt) nodeprobe.Device { return device(nodeprobe.Device{ Name: name, Path: "/dev/" + name, SizeBytes: size, Kind: string(blockdev.KindDisk), Transport: string(blockdev.TransportVirtio), - NUMANode: node, + NUMANode: 0, Available: true, Content: "Blank", }, opts...) @@ -487,30 +488,18 @@ func transported(transport blockdev.Transport) devOpt { return func(d *nodeprobe.Device) { d.Transport = string(transport) } } -// controller is one NVMe controller on the PCI bus, bound to the driver given. -func controller(address, driver string, node int, opts ...func(*nodeprobe.Controller)) nodeprobe.Controller { - c := nodeprobe.Controller{ - Address: address, Driver: driver, NUMANode: node, +// controller is one NVMe controller on the PCI bus, bound to the driver given, +// on the first memory node. Which node a controller sits on is the subject of +// the NUMA cases, and those build their controllers rather than calling this. +func controller(address, driver string) nodeprobe.Controller { + return nodeprobe.Controller{ + Address: address, Driver: driver, NUMANode: 0, Vendor: "0x144d", Product: "0xa80a", - // Checked and found free, which is the state a draft may claim. + // Checked and found free, which is the state a draft may claim. A + // controller in any other state is the subject of the cases that build + // one directly. InUse: ptr.To(false), } - for _, opt := range opts { - opt(&c) - } - return c -} - -// held marks a controller something is driving, which is a disk in service -// rather than one to reclaim. -func held() func(*nodeprobe.Controller) { - return func(c *nodeprobe.Controller) { c.InUse = ptr.To(true) } -} - -// unchecked is a controller the probe could not ask about, which is neither -// held nor free. -func unchecked() func(*nodeprobe.Controller) { - return func(c *nodeprobe.Controller) { c.InUse = nil } } // --- Kubernetes nodes ------------------------------------------------------ @@ -546,11 +535,6 @@ func role(name string) nodeOpt { return func(n *corev1.Node) { n.Labels["node-role.kubernetes.io/"+name] = "" } } -// labeled puts one label on the node. -func labeled(key, value string) nodeOpt { - return func(n *corev1.Node) { n.Labels[key] = value } -} - // reachableAt is the address the cluster reaches the machine on. func reachableAt(address string) nodeOpt { return func(n *corev1.Node) { @@ -609,11 +593,10 @@ func fleet(n int, build func(index int, name string) nodeprobe.Report) []nodepro // kubeFleet is the node objects for a fleet, with each worker reachable on its // own address. -func kubeFleet(reports []nodeprobe.Report, opts ...nodeOpt) []corev1.Node { +func kubeFleet(reports []nodeprobe.Report) []corev1.Node { nodes := make([]corev1.Node, 0, len(reports)) for _, report := range reports { - all := append([]nodeOpt{reachableAt(managementAddress(report.Node))}, opts...) - nodes = append(nodes, kubeNode(report.Node, all...)) + nodes = append(nodes, kubeNode(report.Node, reachableAt(managementAddress(report.Node)))) } return nodes } diff --git a/operator/internal/controllers/cluster/erasurecoding_test.go b/operator/internal/controllers/cluster/erasurecoding_test.go index 584ec5bcd..7341f2cde 100644 --- a/operator/internal/controllers/cluster/erasurecoding_test.go +++ b/operator/internal/controllers/cluster/erasurecoding_test.go @@ -19,11 +19,15 @@ import ( "github.com/simplyblock/simplyblock-operator/internal/webapi" ) +// statusSuspended is the control plane's own spelling of a cluster that is not +// serving, which is the reading every case here is written against. +const statusSuspended = "suspended" + // suspendedCluster is a control-plane reading of a cluster that is not active, // which is the only state an activation is asked for from. func suspendedCluster() webapi.ClusterResponse { reading := activeCluster() - reading.Status = "suspended" + reading.Status = statusSuspended return reading } @@ -45,7 +49,7 @@ func withStripe(data, parity int32) func(*simplyblockv1alpha2.StorageCluster) { c.Spec.Stripe = &simplyblockv1alpha2.StripeSpec{ DataChunks: ptr.To(data), ParityChunks: ptr.To(parity), } - c.Status.Status = "suspended" + c.Status.Status = statusSuspended c.Status.Phase = simplyblockv1alpha2.StorageClusterPhasePending } } diff --git a/operator/internal/controllers/deployment/erasurecoding_test.go b/operator/internal/controllers/deployment/erasurecoding_test.go index 4baec82d0..9b994b935 100644 --- a/operator/internal/controllers/deployment/erasurecoding_test.go +++ b/operator/internal/controllers/deployment/erasurecoding_test.go @@ -22,14 +22,15 @@ import ( ) // aNode is a storage node the cluster already has, which is what a growth -// document adds to. -func aNode(name, worker string, slot int32) client.Object { +// document adds to. Every case puts its node in the first slot, because what +// the cases differ in is how many nodes there are rather than where they sit. +func aNode(name, worker string) client.Object { return &simplyblockv1alpha2.StorageNode{ ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: theNamespace}, Spec: simplyblockv1alpha2.StorageNodeSpec{ ClusterRef: theCluster, WorkerNode: worker, - Slot: ptr.To(slot), + Slot: ptr.To(int32(0)), }, } } @@ -159,7 +160,7 @@ func TestAGrowthDocumentCountsTheNodesTheClusterHas(t *testing.T) { } }) objects := append(workers("worker-1", "worker-2", "worker-3", "worker-4"), - cluster, aNode("node-1", "worker-1", 0), aNode("node-2", "worker-2", 0)) + cluster, aNode("node-1", "worker-1"), aNode("node-2", "worker-2")) if findings := findingsOf(t, config, objects...); len(findings) != 0 { t.Errorf("a growth document reaching the minimum reported %+v", findings) @@ -182,7 +183,7 @@ func TestAGrowthDocumentThatStaysBelowTheMinimumIsReported(t *testing.T) { } }) objects := append(workers("worker-1", "worker-3"), - cluster, aNode("node-1", "worker-1", 0)) + cluster, aNode("node-1", "worker-1")) found := only(t, findingsOf(t, config, objects...), StripeBelowMinimumNodes) for _, want := range []string{"2+1", "4", "2"} { @@ -209,7 +210,7 @@ func TestAGrowthDocumentDoesNotCountASlotTwice(t *testing.T) { } }) objects := append(workers("worker-1", "worker-2"), - cluster, aNode("node-1", "worker-1", 0), aNode("node-2", "worker-2", 0)) + cluster, aNode("node-1", "worker-1"), aNode("node-2", "worker-2")) found := only(t, findingsOf(t, config, objects...), StripeBelowMinimumNodes) if !strings.Contains(found.message, "2") { diff --git a/operator/internal/discovery/fixture_test.go b/operator/internal/discovery/fixture_test.go index bf678b8d1..2544d79ab 100644 --- a/operator/internal/discovery/fixture_test.go +++ b/operator/internal/discovery/fixture_test.go @@ -45,15 +45,16 @@ func refused(d nodeprobe.Device, reasons ...blockdev.Reason) nodeprobe.Device { } // blockDisk is a free disk of the other class: a virtio disk with a path and no -// PCI address the draft would name it by. -func blockDisk(name string, numaNode int, sizeBytes uint64) nodeprobe.Device { +// PCI address the draft would name it by. It sits on the first memory node, +// which is where every case that uses it puts its disks. +func blockDisk(name string, sizeBytes uint64) nodeprobe.Device { return nodeprobe.Device{ Name: name, Path: "/dev/" + name, SizeBytes: sizeBytes, Kind: string(blockdev.KindDisk), Transport: string(blockdev.TransportVirtio), - NUMANode: numaNode, + NUMANode: 0, Available: true, Content: "Blank", } diff --git a/operator/internal/discovery/iscsi_test.go b/operator/internal/discovery/iscsi_test.go index 09c9860ae..af269f9e6 100644 --- a/operator/internal/discovery/iscsi_test.go +++ b/operator/internal/discovery/iscsi_test.go @@ -20,7 +20,7 @@ import ( // attachedLUN is an iSCSI disk as the probe reports one. func attachedLUN(name string, size uint64) nodeprobe.Device { - device := blockDisk(name, 0, size) + device := blockDisk(name, size) device.Transport = string(blockdev.TransportISCSI) return device } @@ -57,7 +57,7 @@ func TestTheRuleLeavesEveryOtherBusAlone(t *testing.T) { blockdev.TransportSATA, blockdev.TransportSAS, blockdev.TransportSCSI, blockdev.TransportVirtio, } { - device := blockDisk("sda", 0, tb) + device := blockDisk("sda", tb) device.Transport = string(transport) if ok, why := admit(ISCSIRule{Class: ClassBlock}, device); !ok { t.Errorf("a %s disk was refused by the iSCSI rule: %s", transport, why) @@ -67,7 +67,7 @@ func TestTheRuleLeavesEveryOtherBusAlone(t *testing.T) { func TestAnISCSILUNIsRefusedByAWholeRunUnlessNamed(t *testing.T) { report := report("worker-1") - report.Devices = []nodeprobe.Device{attachedLUN("sdb", 2*tb), blockDisk("vdb", 0, 2*tb)} + report.Devices = []nodeprobe.Device{attachedLUN("sdb", 2*tb), blockDisk("vdb", 2*tb)} block := &simplyblockv1alpha2.DeviceFilter{EnableLogicalBlockDevices: ptr.To(true)} plan := Planner{}.Plan([]nodeprobe.Report{report}, diff --git a/operator/internal/discovery/mgmtiface_test.go b/operator/internal/discovery/mgmtiface_test.go index 70206e8fb..1c5552db4 100644 --- a/operator/internal/discovery/mgmtiface_test.go +++ b/operator/internal/discovery/mgmtiface_test.go @@ -13,6 +13,16 @@ import ( "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" ) +// The interface names the cases in this package are written in terms of. They +// are constants because one fixture's name and the assertion against it have to +// be the same string, and a case that disagreed with its own fixture by a +// character would pass while asserting nothing. +const ( + eth0 = "eth0" + eth1 = "eth1" + bond0 = "bond0" +) + func iface(name string, edit func(*nodeprobe.Interface)) nodeprobe.Interface { out := nodeprobe.Interface{Name: name, State: "up", NUMANode: 0} if edit != nil { @@ -30,16 +40,16 @@ func iface(name string, edit func(*nodeprobe.Interface)) nodeprobe.Interface { func TestTheInterfaceHoldingTheNodeAddressWins(t *testing.T) { report := report("worker-1") report.Interfaces = []nodeprobe.Interface{ - iface("eth1", func(i *nodeprobe.Interface) { + iface(eth1, func(i *nodeprobe.Interface) { i.Addresses = []string{"10.10.10.113"} i.SpeedMbps = 40000 }), - iface("eth0", func(i *nodeprobe.Interface) { + iface(eth0, func(i *nodeprobe.Interface) { i.Addresses = []string{"192.168.10.113"} }), } - if got := ManagementInterface(report, "192.168.10.113"); got != "eth0" { + if got := ManagementInterface(report, "192.168.10.113"); got != eth0 { t.Errorf("named %q, want the interface holding the node's address", got) } } @@ -49,17 +59,17 @@ func TestTheInterfaceHoldingTheNodeAddressWins(t *testing.T) { func TestTheFastestAddressedPhysicalInterfaceIsNamed(t *testing.T) { report := report("worker-1") report.Interfaces = []nodeprobe.Interface{ - iface("eth0", func(i *nodeprobe.Interface) { + iface(eth0, func(i *nodeprobe.Interface) { i.Addresses = []string{"192.168.10.113"} i.SpeedMbps = 1000 }), - iface("eth1", func(i *nodeprobe.Interface) { + iface(eth1, func(i *nodeprobe.Interface) { i.Addresses = []string{"10.10.10.113"} i.SpeedMbps = 40000 }), } - if got := ManagementInterface(report, ""); got != "eth1" { + if got := ManagementInterface(report, ""); got != eth1 { t.Errorf("named %q, want the fastest addressed interface", got) } } @@ -100,7 +110,7 @@ func TestTheClustersOwnInterfacesAreNeverNamed(t *testing.T) { func TestAnInterfaceWithNoAddressIsNotNamed(t *testing.T) { report := report("worker-1") report.Interfaces = []nodeprobe.Interface{ - iface("eth0", func(i *nodeprobe.Interface) { i.SpeedMbps = 40000 }), + iface(eth0, func(i *nodeprobe.Interface) { i.SpeedMbps = 40000 }), } if got := ManagementInterface(report, ""); got != "" { @@ -113,18 +123,18 @@ func TestAnInterfaceWithNoAddressIsNotNamed(t *testing.T) { func TestALinkThatIsDownIsPassedOver(t *testing.T) { report := report("worker-1") report.Interfaces = []nodeprobe.Interface{ - iface("eth0", func(i *nodeprobe.Interface) { + iface(eth0, func(i *nodeprobe.Interface) { i.Addresses = []string{"192.168.10.113"} i.SpeedMbps = 40000 i.State = "down" }), - iface("eth1", func(i *nodeprobe.Interface) { + iface(eth1, func(i *nodeprobe.Interface) { i.Addresses = []string{"10.10.10.113"} i.SpeedMbps = 1000 }), } - if got := ManagementInterface(report, ""); got != "eth1" { + if got := ManagementInterface(report, ""); got != eth1 { t.Errorf("named %q, want the interface that is up", got) } } @@ -133,7 +143,7 @@ func TestALinkThatIsDownIsPassedOver(t *testing.T) { func TestALinkLocalAddressDoesNotCount(t *testing.T) { report := report("worker-1") report.Interfaces = []nodeprobe.Interface{ - iface("eth0", func(i *nodeprobe.Interface) { + iface(eth0, func(i *nodeprobe.Interface) { i.Addresses = []string{"169.254.1.1", "fe80::1"} }), } @@ -148,8 +158,8 @@ func TestALinkLocalAddressDoesNotCount(t *testing.T) { func TestTheChoiceIsStableAcrossRuns(t *testing.T) { report := report("worker-1") report.Interfaces = []nodeprobe.Interface{ - iface("eth1", func(i *nodeprobe.Interface) { i.Addresses = []string{"10.10.10.113"} }), - iface("eth0", func(i *nodeprobe.Interface) { i.Addresses = []string{"192.168.10.113"} }), + iface(eth1, func(i *nodeprobe.Interface) { i.Addresses = []string{"10.10.10.113"} }), + iface(eth0, func(i *nodeprobe.Interface) { i.Addresses = []string{"192.168.10.113"} }), } first := ManagementInterface(report, "") @@ -157,7 +167,7 @@ func TestTheChoiceIsStableAcrossRuns(t *testing.T) { if second := ManagementInterface(report, ""); second != first { t.Errorf("the reading order changed the answer: %q then %q", first, second) } - if first != "eth0" { + if first != eth0 { t.Errorf("named %q, want the first by name", first) } } diff --git a/operator/internal/discovery/netstack_test.go b/operator/internal/discovery/netstack_test.go index cdfdca5f1..bc73cc1a7 100644 --- a/operator/internal/discovery/netstack_test.go +++ b/operator/internal/discovery/netstack_test.go @@ -31,11 +31,11 @@ func stacked(addressed map[string][]string) []nodeprobe.Interface { lower []string upper []string }{ - {name: "bond0", kind: inventory.LinkBond, lower: []string{"eth0", "eth1"}, upper: []string{"bond0.100"}}, - {name: "bond0.100", kind: inventory.LinkVLAN, lower: []string{"bond0"}}, + {name: bond0, kind: inventory.LinkBond, lower: []string{eth0, eth1}, upper: []string{"bond0.100"}}, + {name: "bond0.100", kind: inventory.LinkVLAN, lower: []string{bond0}}, {name: "br0", kind: inventory.LinkBridge, lower: []string{"eth2"}}, - {name: "eth0", kind: inventory.LinkPhysical, speed: 25000, upper: []string{"bond0"}}, - {name: "eth1", kind: inventory.LinkPhysical, speed: 25000, upper: []string{"bond0"}}, + {name: eth0, kind: inventory.LinkPhysical, speed: 25000, upper: []string{bond0}}, + {name: eth1, kind: inventory.LinkPhysical, speed: 25000, upper: []string{bond0}}, {name: "eth2", kind: inventory.LinkPhysical, speed: 10000, upper: []string{"br0"}}, {name: "flannel.1", kind: inventory.LinkVXLAN}, {name: "lo", kind: inventory.LinkLoopback}, @@ -72,9 +72,9 @@ func stackedReport(addressed map[string][]string) nodeprobe.Report { func TestABondHoldingTheNodeAddressIsNamed(t *testing.T) { // The bond is what an address is bound to on a bonded host. Its members hold // none, so the earlier rule found no candidate at all and named nothing. - r := stackedReport(map[string][]string{"bond0": {"10.10.10.113"}}) + r := stackedReport(map[string][]string{bond0: {"10.10.10.113"}}) - if got := ManagementInterface(r, "10.10.10.113"); got != "bond0" { + if got := ManagementInterface(r, "10.10.10.113"); got != bond0 { t.Errorf("named %q, want the bond holding the node's address", got) } } @@ -138,11 +138,11 @@ func TestTheEffectiveSpeedOfAnAggregateIsItsMembers(t *testing.T) { // A bond reports no speed of its own, so ranking it on what it reports puts // it behind every physical NIC. What it can carry is what its members can. r := stackedReport(map[string][]string{ - "bond0": {"10.10.10.113"}, - "eth2": {"192.168.1.10"}, + bond0: {"10.10.10.113"}, + "eth2": {"192.168.1.10"}, }) - if got := ManagementInterface(r, ""); got != "bond0" { + if got := ManagementInterface(r, ""); got != bond0 { t.Errorf("named %q, want the bond: 2x25G carries more than one 10G NIC", got) } } @@ -164,16 +164,16 @@ func TestAPhysicalInterfaceWinsATieWithADerivedOne(t *testing.T) { r := report("worker-1") r.Interfaces = []nodeprobe.Interface{ { - Name: "eth0", Kind: string(inventory.LinkPhysical), SpeedMbps: 25000, + Name: eth0, Kind: string(inventory.LinkPhysical), SpeedMbps: 25000, State: "up", Addresses: []string{"192.168.1.10"}, Upper: []string{"eth0.100"}, }, { Name: "eth0.100", Kind: string(inventory.LinkVLAN), SpeedMbps: 25000, Virtual: true, - State: "up", Addresses: []string{"10.10.10.113"}, Lower: []string{"eth0"}, + State: "up", Addresses: []string{"10.10.10.113"}, Lower: []string{eth0}, }, } - if got := ManagementInterface(r, ""); got != "eth0" { + if got := ManagementInterface(r, ""); got != eth0 { t.Errorf("named %q, want the physical interface", got) } } @@ -181,13 +181,13 @@ func TestAPhysicalInterfaceWinsATieWithADerivedOne(t *testing.T) { func TestTheHardwareUnderTheChosenInterfaceIsReported(t *testing.T) { // What a bond amounts to is not readable from the bond: it carries no slot, // no driver, and no memory node. The members are the only route to all three. - r := stackedReport(map[string][]string{"bond0": {"10.10.10.113"}}) + r := stackedReport(map[string][]string{bond0: {"10.10.10.113"}}) mgmt := ManagementOf(r, "10.10.10.113") - if mgmt.Name != "bond0" || mgmt.Kind != string(inventory.LinkBond) { + if mgmt.Name != bond0 || mgmt.Kind != string(inventory.LinkBond) { t.Fatalf("chose %+v", mgmt) } - if !slices.Equal(mgmt.Members, []string{"eth0", "eth1"}) { + if !slices.Equal(mgmt.Members, []string{eth0, eth1}) { t.Errorf("reported members %v, want both NICs", mgmt.Members) } if mgmt.SpeedMbps != 50000 { @@ -205,7 +205,7 @@ func TestTheHardwareUnderADerivedInterfaceResolvesThroughItsParent(t *testing.T) r := stackedReport(map[string][]string{"bond0.100": {"10.10.10.113"}}) mgmt := ManagementOf(r, "10.10.10.113") - if !slices.Equal(mgmt.Members, []string{"eth0", "eth1"}) { + if !slices.Equal(mgmt.Members, []string{eth0, eth1}) { t.Errorf("the VLAN reports members %v, want the bond's NICs", mgmt.Members) } } @@ -213,9 +213,9 @@ func TestTheHardwareUnderADerivedInterfaceResolvesThroughItsParent(t *testing.T) func TestMembersOnDifferentMemoryNodesReportNone(t *testing.T) { // A bond across two sockets has no memory node, and reporting one of them // would claim an affinity the interface does not have. - r := stackedReport(map[string][]string{"bond0": {"10.10.10.113"}}) + r := stackedReport(map[string][]string{bond0: {"10.10.10.113"}}) for i := range r.Interfaces { - if r.Interfaces[i].Name == "eth1" { + if r.Interfaces[i].Name == eth1 { r.Interfaces[i].NUMANode = 1 } } @@ -231,12 +231,12 @@ func TestAnInterfaceThatNamesNoKindFallsBackToWhatElseWasReported(t *testing.T) // it was before: physical unless something said otherwise. r := report("worker-1") r.Interfaces = []nodeprobe.Interface{ - {Name: "eth0", State: "up", Addresses: []string{"192.168.1.10"}, SpeedMbps: 10000}, + {Name: eth0, State: "up", Addresses: []string{"192.168.1.10"}, SpeedMbps: 10000}, {Name: "cni0", State: "up", Addresses: []string{"10.42.2.1"}, Virtual: true, Bridge: true}, {Name: "flannel.1", State: "up", Addresses: []string{"10.42.2.0"}, Virtual: true}, } - if got := ManagementInterface(r, ""); got != "eth0" { + if got := ManagementInterface(r, ""); got != eth0 { t.Errorf("named %q, want the one interface nothing marked virtual", got) } } @@ -247,16 +247,16 @@ func TestAStackThatPointsAtItselfTerminates(t *testing.T) { r := report("worker-1") r.Interfaces = []nodeprobe.Interface{ { - Name: "bond0", Kind: string(inventory.LinkBond), State: "up", Virtual: true, + Name: bond0, Kind: string(inventory.LinkBond), State: "up", Virtual: true, Addresses: []string{"10.10.10.113"}, Lower: []string{"bond1"}, }, { Name: "bond1", Kind: string(inventory.LinkBond), State: "up", Virtual: true, - Lower: []string{"bond0"}, + Lower: []string{bond0}, }, } - if got := ManagementOf(r, "10.10.10.113").Name; got != "bond0" { + if got := ManagementOf(r, "10.10.10.113").Name; got != bond0 { t.Errorf("chose %q", got) } } diff --git a/operator/internal/discovery/plan_test.go b/operator/internal/discovery/plan_test.go index 2ec4a0278..11d187ba4 100644 --- a/operator/internal/discovery/plan_test.go +++ b/operator/internal/discovery/plan_test.go @@ -232,8 +232,8 @@ func TestPlanScansTheBlockClassWhenAskedTo(t *testing.T) { // The block class names devices by path, and a virtio disk has no PCI // address at all, so the same fleet yields a block draft or nothing. fleet := []nodeprobe.Report{report("worker-1", - blockDisk("vdb", 0, 2*tb), - blockDisk("vdc", 0, 2*tb), + blockDisk("vdb", 2*tb), + blockDisk("vdc", 2*tb), )} filter := &simplyblockv1alpha2.DeviceFilter{EnableLogicalBlockDevices: ptr.To(true)} @@ -407,8 +407,8 @@ func TestAKernelBoundControllerIsNotCountedTwice(t *testing.T) { // because there is only one place the class is written down. func TestTheFilterDecidesTheClassWithoutBeingToldTwice(t *testing.T) { fleet := []nodeprobe.Report{report("worker-1", - blockDisk("vdb", 0, 2*tb), - blockDisk("vdc", 0, 2*tb), + blockDisk("vdb", 2*tb), + blockDisk("vdc", 2*tb), )} filter := &simplyblockv1alpha2.DeviceFilter{ EnableLogicalBlockDevices: ptr.To(true), diff --git a/operator/internal/discovery/rules_test.go b/operator/internal/discovery/rules_test.go index 453a44667..4707624fd 100644 --- a/operator/internal/discovery/rules_test.go +++ b/operator/internal/discovery/rules_test.go @@ -62,7 +62,7 @@ func TestAvailableRuleWaivesOnlyAPartitionTable(t *testing.T) { func TestClassRuleRefusesADeviceItCannotName(t *testing.T) { nvme := disk("nvme0n1", "0000:5e:00.0", 0, tb) - virtio := blockDisk("vda", 0, tb) + virtio := blockDisk("vda", tb) if ok, _ := admit(ClassRule{Class: ClassNVMe}, nvme); !ok { t.Error("an NVMe run declined an NVMe disk") @@ -372,7 +372,7 @@ func TestClassRuleRefusesTheOtherClassOnABlockRun(t *testing.T) { // A fabric namespace is a volume something else exported, and it is the // other class read the other way: the probe refuses it first, and this is // the rule that keeps it out of a block draft on its own terms. - fabric := blockDisk("nvme1n1", 0, tb) + fabric := blockDisk("nvme1n1", tb) fabric.Transport = string(blockdev.TransportNVMeFabric) if ok, why := admit(ClassRule{Class: ClassBlock}, fabric); ok { t.Error("a block run admitted a fabric namespace") @@ -385,7 +385,7 @@ func TestClassRuleRefusesTheOtherClassOnABlockRun(t *testing.T) { blockdev.TransportVirtio, blockdev.TransportSATA, blockdev.TransportSAS, blockdev.TransportSCSI, } { - device := blockDisk("sda", 0, tb) + device := blockDisk("sda", tb) device.Transport = string(transport) if ok, why := admit(ClassRule{Class: ClassBlock}, device); !ok { t.Errorf("a block run declined a %s disk: %s", transport, why) diff --git a/operator/internal/discovery/simplyblock_test.go b/operator/internal/discovery/simplyblock_test.go index 3722a9734..8775cf45d 100644 --- a/operator/internal/discovery/simplyblock_test.go +++ b/operator/internal/discovery/simplyblock_test.go @@ -23,7 +23,7 @@ import ( // attached is a simplyblock volume as a worker that has connected it sees: a // namespace on a fabric, carrying the NQN the product builds. func attached(name, clusterID, volumeID string) nodeprobe.Device { - device := blockDisk(name, 0, tb) + device := blockDisk(name, tb) device.Transport = string(blockdev.TransportNVMeFabric) device.SubsystemNQN = nqn.Make(clusterID, volumeID) return device @@ -45,7 +45,7 @@ func TestASimplyblockVolumeIsNeverProposed(t *testing.T) { func TestADiskThatIsNotAVolumeIsUntouchedByTheRule(t *testing.T) { for _, device := range []nodeprobe.Device{ disk("nvme0n1", "0000:5e:00.0", 0, tb), - blockDisk("vdb", 0, tb), + blockDisk("vdb", tb), } { if ok, why := admit(SimplyblockVolumeRule{}, device); !ok { t.Errorf("%s was refused as a simplyblock volume: %s", device.Name, why) @@ -55,7 +55,7 @@ func TestADiskThatIsNotAVolumeIsUntouchedByTheRule(t *testing.T) { // A fabric namespace some other product exported is refused for being on a // fabric, by the class rule, and not by this one: what this rule says is // that the disk is ours, and it is not. - foreign := blockDisk("nvme4n1", 0, tb) + foreign := blockDisk("nvme4n1", tb) foreign.Transport = string(blockdev.TransportNVMeFabric) foreign.SubsystemNQN = "nqn.2019-08.org.ceph:rbd.pool.image" if ok, _ := admit(SimplyblockVolumeRule{}, foreign); !ok { From 92625ac11bf8017d7f6fcb95c1f8f6c1f1e184b2 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Sat, 19 Sep 2026 18:59:27 +0200 Subject: [PATCH 090/206] fix(helm): the hostpath class is rendered only where its provisioner is A default install created `StorageClass/local-hostpath`, which names the hostpath provisioner, while `controlplane.csiHostpathDriver.enabled` defaults to false and installs none. The class is then advertised by `kubectl get storageclass` and binds nothing: a PVC pointed at it waits for a provisioner that is never coming, with nothing in the cluster saying why. It predates the guard that was dropped in #544. That guard read `operator.enabled`, which was true in any deployment anybody wants, so the class was rendered unconditionally before it went as well. check-rendered-objects.sh takes the assertion, because the script exists for exactly this class of silence: it already covers objects that must be present, and this is the first that must be absent. It fails on the unchanged chart, naming the class rendered without its provisioner, and it covers the other direction too, so enabling the driver has to produce both. Co-Authored-By: Claude Opus 5 (1M context) --- .../templates/controlplane_storageclass.yaml | 9 ++++- helm-charts/scripts/check-rendered-objects.sh | 39 +++++++++++++++++++ 2 files changed, 47 insertions(+), 1 deletion(-) diff --git a/helm-charts/charts/simplyblock-operator/templates/controlplane_storageclass.yaml b/helm-charts/charts/simplyblock-operator/templates/controlplane_storageclass.yaml index 6486b445a..67607fcfe 100644 --- a/helm-charts/charts/simplyblock-operator/templates/controlplane_storageclass.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/controlplane_storageclass.yaml @@ -1,3 +1,10 @@ +{{- /* +The class names the hostpath provisioner, so it is rendered only where this +release installs one. A class whose provisioner is absent is advertised by +`kubectl get storageclass` and binds nothing: a PVC pointed at it waits for a +provisioner that is never coming, with nothing in the cluster saying why. +*/}} +{{- if .Values.controlplane.csiHostpathDriver.enabled }} --- apiVersion: storage.k8s.io/v1 kind: StorageClass @@ -21,4 +28,4 @@ allowedTopologies: - {{ . }} {{- end }} {{- end }} - +{{- end }} diff --git a/helm-charts/scripts/check-rendered-objects.sh b/helm-charts/scripts/check-rendered-objects.sh index 9001ef0fc..d3fda624e 100755 --- a/helm-charts/scripts/check-rendered-objects.sh +++ b/helm-charts/scripts/check-rendered-objects.sh @@ -64,7 +64,46 @@ check() { fi } +# render prints the objects a profile produces, with the extra settings given. +render() { + local profile="$1" + shift + helm template sb "$CHART" --namespace simplyblock \ + --set deployment.profile="$profile" \ + --set controlplane.managed.endpoint=https://cp.example.com \ + "$@" 2>/dev/null | objects +} + +# checkPair asserts that a StorageClass and the provisioner it names are +# rendered together. +# +# A class naming a provisioner nothing installed is worse than no class: it is +# advertised by `kubectl get storageclass`, a PVC can be pointed at it, and that +# PVC then waits for a provisioner that is never coming, with nothing in the +# cluster saying why. +checkPair() { + local present + + present="$(render standalone)" + if printf '%s\n' "$present" | grep -qxF "StorageClass/local-hostpath"; then + echo " hostpath: StorageClass/local-hostpath is rendered while its provisioner is not installed" + fail=1 + else + echo " hostpath: no StorageClass without its provisioner" + fi + + present="$(render standalone --set controlplane.csiHostpathDriver.enabled=true)" + local want + for want in "StorageClass/local-hostpath" "CSIDriver/hostpath.csi.k8s.io"; do + if ! printf '%s\n' "$present" | grep -qxF "$want"; then + echo " hostpath: MISSING ${want} with the driver enabled" + fail=1 + fi + done +} + check standalone "${COMMON[@]}" check managed "${COMMON[@]}" +checkPair exit "$fail" From ce9436f2173fa3890558f7167b8d631efbff5b3c Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Sat, 19 Sep 2026 18:59:44 +0200 Subject: [PATCH 091/206] test(operator): the metrics group is covered at every kind it serves The scheme publishes five kinds and the tests named two of them. Removing StorageNodeMetrics from the scheme entirely left the whole package green, and the failure would have arrived the first time the kube-apiserver proxied a request for it, as a 500 in somebody's cluster. That is the discovery this package's tests exist to move to the build. One kind has to appear in three places: the scheme, the OpenAPI definitions, and the resource map the installer takes. A roster names them once and the three checks read from it, so a sixth kind fails until all three carry it. The roster is written out rather than read back from the scheme, because one derived from the thing under test agrees with it by construction. The install exercises all five storages, as the server does. With the same kind removed, the three now fail: the scheme does not know it, the roster names what the scheme does not serve, and the installer cannot resolve its OpenAPI model. Co-Authored-By: Claude Opus 5 (1M context) --- operator/internal/metricsapi/scheme_test.go | 106 +++++++++++++++----- 1 file changed, 83 insertions(+), 23 deletions(-) diff --git a/operator/internal/metricsapi/scheme_test.go b/operator/internal/metricsapi/scheme_test.go index 355335d9b..b04b5d2f5 100644 --- a/operator/internal/metricsapi/scheme_test.go +++ b/operator/internal/metricsapi/scheme_test.go @@ -28,30 +28,76 @@ import ( metricsv1alpha2 "github.com/simplyblock/simplyblock-operator/api/metrics/v1alpha2" ) +// servedKinds is every kind this group publishes, against the resource it is +// served at. +// +// It is written out rather than read back from the scheme, because a roster +// derived from the thing under test agrees with it by construction. What this +// one is for is the opposite: to fail when a kind is added to the scheme and +// forgotten in the OpenAPI definitions or in the resource map, which are three +// places one kind has to appear in and which nothing else holds together. +var servedKinds = map[string]string{ + "LogicalVolumeMetrics": ResourceName, + "StorageDeviceMetrics": DeviceResourceName, + "StoragePoolMetrics": PoolResourceName, + "StorageClusterMetrics": ClusterResourceName, + "StorageNodeMetrics": NodeResourceName, +} + // Every kind the group serves, at the version it is served at. A kind registered // under the wrong version is a 404 on the route a client was told to use, which // nothing else in this package would catch. func TestSchemeKnowsTheServedKinds(t *testing.T) { - for _, object := range []runtime.Object{ - &metricsv1alpha2.LogicalVolumeMetrics{}, - &metricsv1alpha2.LogicalVolumeMetricsList{}, - &metricsv1alpha2.StorageDeviceMetrics{}, - &metricsv1alpha2.StorageDeviceMetricsList{}, - } { - kinds, _, err := Scheme.ObjectKinds(object) - if err != nil { - t.Errorf("ObjectKinds(%T): %v", object, err) - continue - } - found := false - for _, kind := range kinds { - if kind.GroupVersion() == metricsv1alpha2.GroupVersion { - found = true + for kind := range servedKinds { + for _, name := range []string{kind, kind + "List"} { + object, err := Scheme.New(metricsv1alpha2.GroupVersion.WithKind(name)) + if err != nil { + t.Errorf("the scheme does not know %s: %v", name, err) + continue + } + kinds, _, err := Scheme.ObjectKinds(object) + if err != nil { + t.Errorf("ObjectKinds(%T): %v", object, err) + continue + } + found := false + for _, registered := range kinds { + if registered.GroupVersion() == metricsv1alpha2.GroupVersion { + found = true + } + } + if !found { + t.Errorf("%T registered as %v, want %s", object, kinds, metricsv1alpha2.GroupVersion) } } - if !found { - t.Errorf("%T registered as %v, want %s", object, kinds, metricsv1alpha2.GroupVersion) + } +} + +// The scheme publishes the roster and nothing besides. +// +// Adding a type to this scheme is what publishes it, so a kind here that the +// roster does not name is served with no storage behind it and appears in the +// discovery document as a resource that answers nothing. +func TestTheSchemeServesExactlyTheRoster(t *testing.T) { + want := map[string]bool{} + for kind := range servedKinds { + want[kind] = true + want[kind+"List"] = true + } + + for kind := range Scheme.KnownTypes(metricsv1alpha2.GroupVersion) { + if !strings.Contains(kind, "Metrics") { + // The meta kinds every group registers: the list and get options, + // and the watch event. + continue } + if !want[kind] { + t.Errorf("the scheme serves %s, which the roster does not name", kind) + } + delete(want, kind) + } + for kind := range want { + t.Errorf("the roster names %s, which the scheme does not serve", kind) } } @@ -124,10 +170,22 @@ func TestAPIGroupInstalls(t *testing.T) { } group := genericapiserver.NewDefaultAPIGroupInfo(metricsv1alpha2.GroupName, Scheme, ParameterCodec, Codecs) - group.VersionedResourcesStorageMap[metricsv1alpha2.GroupVersion.Version] = map[string]rest.Storage{ - ResourceName: NewStorage(fakeVolumes{}, nil, nil), - DeviceResourceName: NewDeviceStorage(nil, nil), + // Every resource the real server installs. The installer refuses a storage + // whose kind the scheme does not know, and it refuses it at startup, so a + // resource left out here is one this test would have passed without. + storages := map[string]rest.Storage{ + ResourceName: NewStorage(fakeVolumes{}, nil, nil), + DeviceResourceName: NewDeviceStorage(nil, nil), + PoolResourceName: NewPoolStorage(nil, nil), + ClusterResourceName: NewClusterStorage(nil, nil), + NodeResourceName: NewNodeStorage(nil, nil), } + for kind, name := range servedKinds { + if _, ok := storages[name]; !ok { + t.Errorf("%s is served at %q, which this install does not exercise", kind, name) + } + } + group.VersionedResourcesStorageMap[metricsv1alpha2.GroupVersion.Version] = storages if err := server.InstallAPIGroup(&group); err != nil { t.Fatalf("InstallAPIGroup: %v", err) } @@ -188,9 +246,11 @@ func TestCodecRoundTripsADeviceReading(t *testing.T) { func TestOpenAPIDefinitionsCoverEveryServedKind(t *testing.T) { definitions := openAPIDefinitions(func(string) spec.Ref { return spec.Ref{} }) const pkg = "github.com/simplyblock/simplyblock-operator/api/metrics/v1alpha2." - for _, kind := range []string{"LogicalVolumeMetrics", "StorageDeviceMetrics"} { - if _, ok := definitions[pkg+kind]; !ok { - t.Errorf("no OpenAPI definition for %s", pkg+kind) + for kind := range servedKinds { + for _, name := range []string{kind, kind + "List"} { + if _, ok := definitions[pkg+name]; !ok { + t.Errorf("no OpenAPI definition for %s", pkg+name) + } } } } From 4df8e883cd5ffea70e24f71fd97258770dacbcb3 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Sun, 20 Sep 2026 13:51:46 +0200 Subject: [PATCH 092/206] fix(operator): the driver the chart writes is admitted before the operator serves MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A first `helm install` of the chart failed at the `SimplyblockDriver`: Internal error occurred: failed calling webhook "vsimplyblockdriver.simplyblock.io": no endpoints available for service "simplyblock-operator-webhook-service" The chart renders that object beside the Deployment that validates it, so the create arrives seconds after the webhook configuration is registered and about a minute before the operator answers on the service, which the cert rotator lengthens by exiting twice to pick up its refreshed serving material. Under `failurePolicy: Fail` the release stops there, on a cluster holding an operator, a control plane, and no CSI driver, and the only way out is a second run of the same command. It is every first install, not a race that sometimes lands. The validator was never the enforcement of record. design-simplyblockdriver.md §3.4 already specifies the controller half for exactly this window: a driver admitted while the webhook was unavailable is not the oldest in the cluster, so it holds at `Installing`, applies nothing, and emits `DuplicateDriver`. What `Fail` bought was a clearer message at the moment of typing, and it cost the install. The test reads the generated webhook configuration rather than a copy of it and registers it against an apiserver with nothing serving behind it, which is the condition a first install is in. On the unchanged tree the create is refused by name, so a marker edit that puts the policy back is what makes it fail again. The chart is the only writer of the field: the operator patches `caBundle` into these configurations at runtime and nothing else, so this reaches a cluster through `helm upgrade` without a new image. Co-Authored-By: Claude Opus 5 (1M context) --- .../simplyblock-operator-webhook.yaml | 2 +- operator/config/webhook/manifests.yaml | 2 +- operator/dist/install.yaml | 2 +- .../crd-redesign/design-simplyblockdriver.md | 10 +- .../docs/tests/test-plan-simplyblockdriver.md | 60 ++++--- .../webhook/install_admission_test.go | 160 ++++++++++++++++++ .../webhook/simplyblockdriver_validator.go | 19 ++- 7 files changed, 219 insertions(+), 36 deletions(-) create mode 100644 operator/internal/webhook/install_admission_test.go diff --git a/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml b/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml index bf194bd4c..b372c5e4b 100644 --- a/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml @@ -192,7 +192,7 @@ webhooks: name: simplyblock-operator-webhook-service namespace: {{ .Release.Namespace }} path: /validate-storage-simplyblock-io-v1alpha2-simplyblockdriver - failurePolicy: Fail + failurePolicy: Ignore name: vsimplyblockdriver.simplyblock.io rules: - apiGroups: diff --git a/operator/config/webhook/manifests.yaml b/operator/config/webhook/manifests.yaml index 7f94881fc..c51b858bf 100644 --- a/operator/config/webhook/manifests.yaml +++ b/operator/config/webhook/manifests.yaml @@ -174,7 +174,7 @@ webhooks: name: webhook-service namespace: system path: /validate-storage-simplyblock-io-v1alpha2-simplyblockdriver - failurePolicy: Fail + failurePolicy: Ignore name: vsimplyblockdriver.simplyblock.io rules: - apiGroups: diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index cfe93c684..d85e89fa8 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -11514,7 +11514,7 @@ webhooks: name: simplyblock-operator-webhook-service namespace: simplyblock-operator-system path: /validate-storage-simplyblock-io-v1alpha2-simplyblockdriver - failurePolicy: Fail + failurePolicy: Ignore name: vsimplyblockdriver.simplyblock.io rules: - apiGroups: diff --git a/operator/docs/designs/crd-redesign/design-simplyblockdriver.md b/operator/docs/designs/crd-redesign/design-simplyblockdriver.md index 0da745ecb..42b5599ae 100644 --- a/operator/docs/designs/crd-redesign/design-simplyblockdriver.md +++ b/operator/docs/designs/crd-redesign/design-simplyblockdriver.md @@ -290,8 +290,14 @@ no object name reaches. **Enforcement is a validating webhook**, `SimplyblockDriverValidator` in `operator/internal/webhook/simplyblockdriver_validator.go`, which denies a -`CREATE` where a `SimplyblockDriver` already exists in any namespace and carries -`failurePolicy=fail` like every other validator the operator serves. The +`CREATE` where a `SimplyblockDriver` already exists in any namespace. It carries +`failurePolicy=ignore` rather than the `fail` most of the operator's validators +carry, and the install is what decides that: the chart renders this object beside +the `Deployment` that serves the webhook (§8), so on a first install the `CREATE` +arrives while the operator is still coming up, and under `fail` the release stops +on it and leaves a cluster holding an operator, a control plane, and no driver. +The singleton survives the difference because the webhook is the message rather +than the enforcement, and the controller below refuses the same object. The `ControlPlane` singleton is enforced by convention instead ([`design-controlplane.md`](design-controlplane.md) §3.1), and what separates the two is what a second object does: a `ControlPlane` under another name is ignored diff --git a/operator/docs/tests/test-plan-simplyblockdriver.md b/operator/docs/tests/test-plan-simplyblockdriver.md index e16649e95..41e2c1980 100644 --- a/operator/docs/tests/test-plan-simplyblockdriver.md +++ b/operator/docs/tests/test-plan-simplyblockdriver.md @@ -270,37 +270,45 @@ control-plane upgrade is the surprise the design declines to build. Full reconcile loop against a real API server via `envtest`. The immutability and defaulting rules are admission and cannot be exercised any other way. -| # | Scenario | Type | Test | -|----------|------------------------------------------------------------------------------|--------------|--------------------| -| I-01 | `spec.driverName` changed after creation: rejected as immutable | Negative | — | -| I-02 | `spec.driverName` unset: defaulted to `csi.simplyblock.io` | Boundary | — | -| I-03 | `spec.image` outside the trusted registries: rejected by the pattern | Negative | — | -| I-04 | `spec.image` omitted: rejected as `Required` | Negative | — | -| I-05 | `spec.controllerReplicas` of 0: rejected by the minimum | Boundary | — | -| I-06 | `spec.controllerReplicas` unset: defaulted to 1 | Boundary | — | -| I-07 | `spec.imagePullPolicy` outside the enum: rejected | Negative | — | -| I-08 | The short name `sbd` resolves to the same list as the full kind | Positive | — | -| I-09 | A full apply against a real API server: every object exists afterward | Positive | — | -| I-10 | Deleting the object: garbage collection removes every applied child | Positive | — | -| I-11 | The controller's role covers every object the apply creates | Positive | — | -| ~~I-12~~ | ~~Two drivers with two `driverName` values in one namespace: both accepted~~ | ~~Boundary~~ | Superseded by I-18 | -| I-13 | `spec.enableVolumeSnapshots` unset: defaulted to true | Boundary | — | -| I-14 | Field ownership is taken from a manager that already holds the fields | Positive | — | -| I-15 | An adopted object keeps its UID across the reconcile that adopts it | Positive | — | -| I-16 | The finalizer deletes the cluster-scoped objects the label marks | Positive | — | -| I-17 | A cluster-scoped object carrying another controller's label is left | Negative | — | -| I-18 | A second `SimplyblockDriver` in the same namespace: denied at admission | Negative | — | -| I-19 | A second one in another namespace: denied, whatever its `driverName` | Negative | — | -| I-20 | The first `SimplyblockDriver` in an empty cluster: admitted | Positive | — | -| I-21 | The only object updated, not created: admitted, since the rule is CREATE | Boundary | — | -| I-22 | The only object deleted and another created: admitted | Boundary | — | -| I-23 | A `spec.sidecarImages` entry outside the trusted registries: rejected | Negative | — | -| I-24 | `spec.image` of the empty string: rejected, since Required admits it | Negative | — | +| # | Scenario | Type | Test | +|----------|------------------------------------------------------------------------------|--------------|-------------------------------------------------------------| +| I-01 | `spec.driverName` changed after creation: rejected as immutable | Negative | — | +| I-02 | `spec.driverName` unset: defaulted to `csi.simplyblock.io` | Boundary | — | +| I-03 | `spec.image` outside the trusted registries: rejected by the pattern | Negative | — | +| I-04 | `spec.image` omitted: rejected as `Required` | Negative | — | +| I-05 | `spec.controllerReplicas` of 0: rejected by the minimum | Boundary | — | +| I-06 | `spec.controllerReplicas` unset: defaulted to 1 | Boundary | — | +| I-07 | `spec.imagePullPolicy` outside the enum: rejected | Negative | — | +| I-08 | The short name `sbd` resolves to the same list as the full kind | Positive | — | +| I-09 | A full apply against a real API server: every object exists afterward | Positive | — | +| I-10 | Deleting the object: garbage collection removes every applied child | Positive | — | +| I-11 | The controller's role covers every object the apply creates | Positive | — | +| ~~I-12~~ | ~~Two drivers with two `driverName` values in one namespace: both accepted~~ | ~~Boundary~~ | Superseded by I-18 | +| I-13 | `spec.enableVolumeSnapshots` unset: defaulted to true | Boundary | — | +| I-14 | Field ownership is taken from a manager that already holds the fields | Positive | — | +| I-15 | An adopted object keeps its UID across the reconcile that adopts it | Positive | — | +| I-16 | The finalizer deletes the cluster-scoped objects the label marks | Positive | — | +| I-17 | A cluster-scoped object carrying another controller's label is left | Negative | — | +| I-18 | A second `SimplyblockDriver` in the same namespace: denied at admission | Negative | — | +| I-19 | A second one in another namespace: denied, whatever its `driverName` | Negative | — | +| I-20 | The first `SimplyblockDriver` in an empty cluster: admitted | Positive | — | +| I-21 | The only object updated, not created: admitted, since the rule is CREATE | Boundary | — | +| I-22 | The only object deleted and another created: admitted | Boundary | — | +| I-23 | A `spec.sidecarImages` entry outside the trusted registries: rejected | Negative | — | +| I-24 | `spec.image` of the empty string: rejected, since Required admits it | Negative | — | +| I-25 | The chart's object admitted while the operator is not yet serving | Regression | `TestBootstrapDriverIsAdmittedWhileTheOperatorIsNotServing` | `I-12` asserted that two drivers in one namespace were accepted, which design §3.4 now rejects. The row keeps its ID struck through, and `I-18` is what replaced it. +`I-25` is the row a first install turns on. The chart writes its +`SimplyblockDriver` beside the `Deployment` that serves this webhook, so the +object arrives while nothing answers on the service, and under `failurePolicy: +Fail` the release stopped there and left a cluster with no CSI driver +(2026-09-20). It reads the generated webhook configuration rather than a copy, so +a marker edit that puts the policy back is what makes it fail. + `I-18` to `I-22` are design §3.4's webhook, and the last three are the boundary the rule is written against. It denies a `CREATE` where an object already exists, so editing the one that does must not be denied and replacing it must not be diff --git a/operator/internal/webhook/install_admission_test.go b/operator/internal/webhook/install_admission_test.go new file mode 100644 index 000000000..9dd33f398 --- /dev/null +++ b/operator/internal/webhook/install_admission_test.go @@ -0,0 +1,160 @@ +// What a fresh install's own objects must survive: admission, on a cluster where +// the operator serving the webhooks is part of the same install and is not up +// yet. +// +// The chart renders a SimplyblockDriver beside the Deployment that validates it +// (design-simplyblockdriver.md §8), so on a first install the API server is asked +// to admit that object seconds after the webhook configuration is registered and +// roughly a minute before the operator answers on the service. Whether that +// succeeds is decided by one field of the generated configuration, which is why +// this test reads the generated file rather than a copy of it. +// +// It lives beside the validators rather than under a chart test because the field +// it pins is written as a kubebuilder marker in this package, and `helm-sync` +// copies it into the chart from here. + +package webhook + +import ( + "bytes" + "context" + "errors" + "io" + "os" + "path/filepath" + "sync" + "testing" + + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" + utilyaml "k8s.io/apimachinery/pkg/util/yaml" + "k8s.io/client-go/kubernetes/scheme" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/envtest" + logf "sigs.k8s.io/controller-runtime/pkg/log" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +var ( + sharedEnvOnce sync.Once + sharedEnv *envtest.Environment + sharedEnvClient client.Client + sharedEnvErr error +) + +// TestMain stops the shared apiserver once, after every test that used one. +func TestMain(m *testing.M) { + code := m.Run() + if sharedEnv != nil { + _ = sharedEnv.Stop() + } + os.Exit(code) +} + +// apiServer returns a client against a real apiserver with this repository's +// CRDs installed, starting one on first use. +// +// A real one is the point here: admission is the apiserver's, and a fake client +// admits everything, so the behavior this file is about is invisible below this +// level. +func apiServer(t *testing.T) client.Client { + t.Helper() + sharedEnvOnce.Do(func() { + if err := simplyblockv1alpha2.AddToScheme(scheme.Scheme); err != nil { + sharedEnvErr = err + return + } + sharedEnv = &envtest.Environment{ + CRDDirectoryPaths: []string{filepath.Join("..", "..", "config", "crd", "bases")}, + ErrorIfCRDPathMissing: true, + BinaryAssetsDirectory: envTestBinaryDir(), + } + cfg, err := sharedEnv.Start() + if err != nil { + sharedEnvErr = err + return + } + sharedEnvClient, sharedEnvErr = client.New(cfg, client.Options{Scheme: scheme.Scheme}) + }) + if sharedEnvErr != nil { + t.Fatalf("starting the test apiserver: %v", sharedEnvErr) + } + return sharedEnvClient +} + +// envTestBinaryDir locates the envtest asset binaries in the repository-root +// .bin that `make setup-envtest` populates, so that the test runs from an editor +// as well as from the Makefile. +func envTestBinaryDir() string { + basePath := filepath.Join("..", "..", "..", ".bin", "k8s") + entries, err := os.ReadDir(basePath) + if err != nil { + logf.Log.Error(err, "the envtest assets could not be read", "path", basePath) + return "" + } + for _, entry := range entries { + if entry.IsDir() { + return filepath.Join(basePath, entry.Name()) + } + } + return "" +} + +// registerGeneratedWebhooks applies the webhook configurations exactly as +// `make manifests` generates them, and points nothing at a running server. +// +// The generated clientConfig names a service in a namespace neither of which +// exists here, which is the same condition a first install is in: the +// configuration is registered, and the endpoint behind it is not there yet. What +// the apiserver does about that is the failurePolicy's answer, and that is the +// field under test. +func registerGeneratedWebhooks(t *testing.T, c client.Client) { + t.Helper() + + raw, err := os.ReadFile(filepath.Join("..", "..", "config", "webhook", "manifests.yaml")) + if err != nil { + t.Fatalf("reading the generated webhook manifests: %v", err) + } + + decoder := utilyaml.NewYAMLOrJSONDecoder(bytes.NewReader(raw), 4096) + for { + object := &unstructured.Unstructured{} + if err := decoder.Decode(object); err != nil { + if errors.Is(err, io.EOF) { + break + } + t.Fatalf("decoding the generated webhook manifests: %v", err) + } + if len(object.Object) == 0 { + continue + } + if err := c.Create(context.Background(), object); err != nil { + t.Fatalf("registering %s/%s: %v", object.GetKind(), object.GetName(), err) + } + t.Cleanup(func() { _ = c.Delete(context.Background(), object) }) + } +} + +// TestBootstrapDriverIsAdmittedWhileTheOperatorIsNotServing covers the object the +// chart writes on a first install. +// +// Regression: 2026-09-20-fresh-install-blocked-by-driver-webhook — a first +// `helm install` of the chart failed with `Internal error occurred: failed +// calling webhook "vsimplyblockdriver.simplyblock.io": no endpoints available for +// service "simplyblock-operator-webhook-service"`. The release stopped at the +// SimplyblockDriver, which left a cluster carrying an operator, a control plane, +// and no CSI driver, and the install had to be run a second time to complete. +func TestBootstrapDriverIsAdmittedWhileTheOperatorIsNotServing(t *testing.T) { + c := apiServer(t) + registerGeneratedWebhooks(t, c) + + driver := &simplyblockv1alpha2.SimplyblockDriver{ + ObjectMeta: metav1.ObjectMeta{Name: "simplyblock", Namespace: "default"}, + } + if err := c.Create(context.Background(), driver); err != nil { + t.Fatalf("the chart's SimplyblockDriver was refused while the operator was not serving, "+ + "which is every first install of the chart: %v", err) + } + t.Cleanup(func() { _ = c.Delete(context.Background(), driver) }) +} diff --git a/operator/internal/webhook/simplyblockdriver_validator.go b/operator/internal/webhook/simplyblockdriver_validator.go index 5dd54a212..376d8766c 100644 --- a/operator/internal/webhook/simplyblockdriver_validator.go +++ b/operator/internal/webhook/simplyblockdriver_validator.go @@ -21,7 +21,7 @@ import ( "github.com/simplyblock/simplyblock-operator/internal/controllers/driver" ) -// +kubebuilder:webhook:path=/validate-storage-simplyblock-io-v1alpha2-simplyblockdriver,mutating=false,failurePolicy=fail,sideEffects=None,groups=storage.simplyblock.io,resources=simplyblockdrivers,verbs=create,versions=v1alpha2,name=vsimplyblockdriver.simplyblock.io,admissionReviewVersions=v1 +// +kubebuilder:webhook:path=/validate-storage-simplyblock-io-v1alpha2-simplyblockdriver,mutating=false,failurePolicy=ignore,sideEffects=None,groups=storage.simplyblock.io,resources=simplyblockdrivers,verbs=create,versions=v1alpha2,name=vsimplyblockdriver.simplyblock.io,admissionReviewVersions=v1 // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=simplyblockdrivers,verbs=get;list;watch @@ -38,10 +38,19 @@ import ( // plugin per driver name, so two node plugins on one worker contend for a path // no object name reaches. // -// failurePolicy=Fail, like every other validator here. What it blocks while -// unavailable is the creation of a driver, which is a deployment-time action -// rather than a data-path one, and the object written in that window is caught -// by the controller instead. +// failurePolicy=Ignore, unlike most validators here, and the install is what +// decides it. The chart renders a SimplyblockDriver beside the Deployment that +// serves this webhook, so the one CREATE a first install performs arrives +// seconds after the webhook configuration is registered and about a minute +// before the operator answers on the service. Under Fail the release stops +// there, on a cluster left holding an operator, a control plane, and no CSI +// driver, and the only way out is to run the install a second time. +// +// Nothing is given up by ignoring it, because this webhook was never the +// enforcement of record: a driver admitted while it was unavailable is refused +// by the controller, which holds the object at Installing, applies nothing, and +// emits DuplicateDriver (§3.4). What Fail bought was a clearer message at the +// moment of typing, and it cost every first install. // // The rule is CREATE and not UPDATE, because an edit to the object that already // exists is not a second one, and a rule over every operation would lock the From 07d62f754788ac5db53dc7c8dd334ea34d51b082 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Sun, 20 Sep 2026 13:51:52 +0200 Subject: [PATCH 093/206] fix(operator): the initial discovery run outlasts the webhook it depends on MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A fresh install raised no discovery run, and the operator log carried it once: ERROR initial-discovery the initial discovery run could not be created {"error": "Internal error occurred: failed calling webhook \"voperatorops.simplyblock.io\": ... dial tcp 10.130.2.137:9443: connect: connection refused"} The address is the operator's own pod, on its own webhook port. An OperatorOps is validated by a webhook this same process serves, which makes it the one write in the operator whose admission depends on the operator, and the create races the webhook server the manager is still starting. Start is not a loop. It attempted once, logged, and returned, so a few seconds of unavailability cost the installation its draft for good: nothing re-raises the run, and an administrator finds an empty namespace where the fleet's disks should have been listed. The guard that declines the run once anything exists is what makes it unrecoverable rather than merely late. Weakening the webhook is the wrong end of it here, because the operator can wait and the chart cannot. The create now waits the window out, and only for an unreachable webhook: Internal, ServiceUnavailable, and Timeout are waited out, while Forbidden and Invalid are answers that repeating the write cannot change. The second test is what holds that line, and it passes on the unchanged tree too, because its value is in what it stops the fix from becoming. The budget is two minutes at half-second intervals, which covers a cold start plus the cert rotator's one restart. Past it the run is not raised, which is the behavior that was there before and which an administrator already has a remedy for. design-clusterdeploymentconfig.md gains §8.4, which had no prose at all for the automatic first run, and the test plan gains the section its nine existing cases were missing. Co-Authored-By: Claude Opus 5 (1M context) --- .../design-clusterdeploymentconfig.md | 27 +++++++ .../test-plan-clusterdeploymentconfig.md | 30 +++++++ .../controllers/deployment/bootstrap.go | 66 +++++++++++++++- .../controllers/deployment/bootstrap_test.go | 78 +++++++++++++++++++ 4 files changed, 197 insertions(+), 4 deletions(-) diff --git a/operator/docs/designs/crd-redesign/design-clusterdeploymentconfig.md b/operator/docs/designs/crd-redesign/design-clusterdeploymentconfig.md index 6134de156..8c1662dff 100644 --- a/operator/docs/designs/crd-redesign/design-clusterdeploymentconfig.md +++ b/operator/docs/designs/crd-redesign/design-clusterdeploymentconfig.md @@ -839,6 +839,33 @@ re-run. A second discovery writes a second document rather than editing the first, because the first may have been reviewed and edited, and overwriting a reviewer's corrections with a fresh guess is the worst behavior available. +### 8.4 The first run raises itself + +**An install that has nothing raises one `Discover` run by itself**, so that an +administrator finds a draft of what the fleet has rather than an empty namespace. +It is raised by a `Runnable` in the operator rather than by a reconciler, because +there is no object whose desired state it converges toward and the question has +one answer per installation. It is declined the moment anything already exists: +an `OperatorOps` of any kind, a `ClusterDeploymentConfig`, a `StorageCluster`, or +a fleet with no worker that holds storage without being asked to. The guard is +most of the behavior, because probing puts a Job on every worker, and against a +deployed fleet that cost falls on every worker of it once per operator upgrade. + +**The run is named rather than generated**, so the create is idempotent on top of +that guard, and so an administrator who does not want it can say so by writing an +object under that name. + +**It is the one write in the operator whose admission depends on the operator.** +An `OperatorOps` is validated by `voperatorops.simplyblock.io`, which this same +process serves, so the create races the webhook server the manager is still +starting and the API server answers `failed calling webhook ... connection +refused` until it is listening. The create therefore waits that window out rather +than attempting once: a `Runnable` that is not a loop turns a few seconds of +unavailability into an installation that never gets its draft, with nothing left +to raise it. Only an unreachable webhook is waited out, and a webhook that +answered and refused the spec has given an answer that repeating the write cannot +change. + --- ## 9. Observability diff --git a/operator/docs/tests/test-plan-clusterdeploymentconfig.md b/operator/docs/tests/test-plan-clusterdeploymentconfig.md index 6068b1623..c1dc9a7a1 100644 --- a/operator/docs/tests/test-plan-clusterdeploymentconfig.md +++ b/operator/docs/tests/test-plan-clusterdeploymentconfig.md @@ -127,6 +127,36 @@ File: `operator/internal/controllers/deployment/operatorops_discover_test.go` | U-61 | `configName` unset: a generated name that cannot collide with the first run | Boundary | — | | U-62 | Discovery changes nothing: no cluster, no node, no control-plane write | Negative | — | +### The Automatic First Run (design §8.4) + +File: `operator/internal/controllers/deployment/bootstrap_test.go` + +The run these rows cover is the one nobody asks for. Probing puts a Job on every +worker, so the guard is most of the behavior: `U-168` to `U-172` are the states +that decline, and each of them is enough on its own. + +| # | Scenario | Type | Test | +|-------|---------------------------------------------------------------------------------------|------------|-------------------------------------------------| +| U-167 | An install with nothing in it: one `Discover` run, narrowing nothing | Positive | `TestAFreshInstallRaisesOneDiscoveryRun` | +| U-168 | A previous result of any of the three kinds: declined, and it says which | Negative | `TestAPreviousResultDeclinesTheRun` | +| U-169 | A restart after the first run: nothing further is raised | Negative | `TestARestartRaisesNothingFurther` | +| U-170 | An object already carrying that name: left as it is, not replaced | Boundary | `TestAnObjectByThatNameIsNotReplaced` | +| U-171 | A cluster with no node to inspect: no run | Negative | `TestAClusterWithNothingToInspectRaisesNoRun` | +| U-172 | A fleet whose every worker is cordoned: no run | Negative | `TestACordonedFleetRaisesNoRun` | +| U-173 | One usable worker: enough to raise it | Boundary | `TestAWorkerIsEnoughToRaiseTheRun` | +| U-174 | The run waives a partition table, so a fleet that has held data is not reported empty | Positive | `TestTheInitialRunWaivesAPartitionTable` | +| U-175 | One replica asks, because the three reads are not idempotent | Positive | `TestTheCheckIsLeaderElected` | +| U-176 | The operator's own webhook not serving yet: the create is waited out, not lost | Regression | `TestTheRunOutlastsAWebhookThatIsNotServingYet` | +| U-177 | A webhook that answered and refused: not retried, because waiting changes nothing | Negative | `TestARejectedRunIsNotRetried` | + +`U-176` is the row a fresh install turns on. An `OperatorOps` is validated by a +webhook this same operator serves, so the create races the webhook server the +manager is still starting, and `Start` is not a loop: the single attempt lost the +draft for the lifetime of the installation, and an administrator found an empty +namespace where the fleet's disks should have been (2026-09-20). `U-177` is what +keeps the remedy from becoming a write repeated until the deadline against a +webhook that already gave its answer. + ### OperatorOps Lifecycle (design §7) File: `operator/internal/controllers/deployment/operatorops_controller_unit_test.go` diff --git a/operator/internal/controllers/deployment/bootstrap.go b/operator/internal/controllers/deployment/bootstrap.go index aac8b8ac7..222b72612 100644 --- a/operator/internal/controllers/deployment/bootstrap.go +++ b/operator/internal/controllers/deployment/bootstrap.go @@ -45,10 +45,12 @@ package deployment import ( "context" "fmt" + "time" corev1 "k8s.io/api/core/v1" apierrors "k8s.io/apimachinery/pkg/api/errors" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/util/wait" ctrl "sigs.k8s.io/controller-runtime" "sigs.k8s.io/controller-runtime/pkg/client" logf "sigs.k8s.io/controller-runtime/pkg/log" @@ -63,6 +65,19 @@ import ( // object with that name and deleting nothing. const InitialDiscoveryName = "initial-discovery" +// How long the create waits for the operator's own webhook to begin serving, and +// how often it asks. +// +// The budget covers a cold start rather than an outage: the webhook server is +// listening within seconds of the manager starting, and the operator restarts +// once more when its cert rotator refreshes the serving material, so the window +// this crosses is a cold start plus one restart. Past the budget the run is not +// raised, which is the behavior an administrator already has a remedy for. +const ( + createRetryInterval = 500 * time.Millisecond + createRetryDeadline = 2 * time.Minute +) + // InitialDiscovery raises one Discover run on an install that has nothing. // // It is a Runnable rather than a reconciler because it is not reconciling @@ -129,10 +144,7 @@ func (d *InitialDiscovery) Start(ctx context.Context) error { }, }, } - if err := d.Create(ctx, run); err != nil { - if apierrors.IsAlreadyExists(err) { - return nil - } + if err := d.createRun(ctx, run); err != nil { // Failing here would crash the manager over a convenience. The operator // works without the run; what is lost is the draft an administrator would // otherwise have found waiting. @@ -145,6 +157,52 @@ func (d *InitialDiscovery) Start(ctx context.Context) error { return nil } +// createRun writes the run, waiting out a webhook that is not serving yet. +// +// This is the one write in the operator whose admission depends on the operator: +// an OperatorOps is validated by voperatorops.simplyblock.io, which this same +// process serves, so the create races the webhook server the manager is still +// starting and the API server answers `failed calling webhook ... connection +// refused` until it is listening. Start is not a loop, so a single attempt lost +// the run for the lifetime of the installation rather than for a few seconds. +// +// Only an unreachable webhook is waited out. A webhook that answered and refused +// the spec has given an answer, and repeating the same write until the deadline +// would not change it. +func (d *InitialDiscovery) createRun(ctx context.Context, run *simplyblockv1alpha2.OperatorOps) error { + lastErr := error(nil) + err := wait.PollUntilContextTimeout(ctx, createRetryInterval, createRetryDeadline, true, + func(ctx context.Context) (bool, error) { + switch err := d.Create(ctx, run); { + case err == nil, apierrors.IsAlreadyExists(err): + return true, nil + case webhookUnreachable(err): + lastErr = err + return false, nil + default: + return false, err + } + }) + if wait.Interrupted(err) && lastErr != nil { + // The deadline says how long it waited, and the refusal says what it + // waited for. The second is the one worth reading in a log. + return lastErr + } + return err +} + +// webhookUnreachable reports whether the API server refused the write because it +// could not call an admission webhook, rather than because a webhook rejected +// what was written. +// +// An unreachable webhook surfaces as an Internal error naming it, and a refusal +// surfaces as Forbidden or Invalid, which this deliberately does not cover. +func webhookUnreachable(err error) bool { + return apierrors.IsInternalError(err) || + apierrors.IsServiceUnavailable(err) || + apierrors.IsTimeout(err) +} + // declineReason answers whether anything already exists, and says which thing. An // empty string means the install has nothing and the run is worth raising. func (d *InitialDiscovery) declineReason(ctx context.Context) (string, error) { diff --git a/operator/internal/controllers/deployment/bootstrap_test.go b/operator/internal/controllers/deployment/bootstrap_test.go index 46c589e53..46389ee55 100644 --- a/operator/internal/controllers/deployment/bootstrap_test.go +++ b/operator/internal/controllers/deployment/bootstrap_test.go @@ -9,12 +9,16 @@ package deployment import ( "context" + "fmt" "testing" corev1 "k8s.io/api/core/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/runtime/schema" "sigs.k8s.io/controller-runtime/pkg/client" "sigs.k8s.io/controller-runtime/pkg/client/fake" + "sigs.k8s.io/controller-runtime/pkg/client/interceptor" simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" "github.com/simplyblock/simplyblock-operator/internal/controllers/testsupport" @@ -29,6 +33,32 @@ func discoveryFor(t *testing.T, objects ...client.Object) *InitialDiscovery { } } +// discoveryRefusingCreates builds the check against a client that answers the +// first refusals writes get while an admission webhook is unreachable, and +// admits the write after that. It is how the startup window is reproduced +// without an apiserver: what the API server returns in that window is an +// Internal error naming the webhook it could not call. +func discoveryRefusingCreates(t *testing.T, refusals int, objects ...client.Object) (*InitialDiscovery, *int) { + t.Helper() + scheme := testsupport.NewScheme(t, corev1.AddToScheme) + attempts := 0 + inner := fake.NewClientBuilder().WithScheme(scheme).WithObjects(objects...).Build() + refusing := interceptor.NewClient(inner, interceptor.Funcs{ + Create: func(ctx context.Context, c client.WithWatch, obj client.Object, opts ...client.CreateOption) error { + attempts++ + if attempts <= refusals { + return apierrors.NewInternalError(fmt.Errorf( + `failed calling webhook "voperatorops.simplyblock.io": failed to call webhook: ` + + `Post "https://simplyblock-operator-webhook-service.simplyblock.svc:443/` + + `validate-storage-simplyblock-io-v1alpha2-operatorops?timeout=10s": ` + + `dial tcp 10.130.2.137:9443: connect: connection refused`)) + } + return c.Create(ctx, obj, opts...) + }, + }) + return &InitialDiscovery{Client: refusing, Namespace: theNamespace}, &attempts +} + // raised reports whether the run exists, which is what every case here asserts. func raised(t *testing.T, d *InitialDiscovery) bool { t.Helper() @@ -265,3 +295,51 @@ func TestTheInitialRunWaivesAPartitionTable(t *testing.T) { t.Error("the initial run refuses a disk for carrying a partition table") } } + +// TestTheRunOutlastsAWebhookThatIsNotServingYet covers the one write in the +// operator whose admission depends on the operator. +// +// Regression: 2026-09-20-initial-discovery-lost-to-its-own-webhook — on a fresh +// install the run was never raised, and the operator log carried +// `the initial discovery run could not be created ... failed calling webhook +// "voperatorops.simplyblock.io": ... connect: connection refused`. The create +// raced the webhook server this same process was still starting, Start is not a +// loop, and the single attempt lost the draft for the lifetime of the +// installation: an administrator found an empty namespace where the fleet's +// disks should have been listed, with nothing to re-raise it. +func TestTheRunOutlastsAWebhookThatIsNotServingYet(t *testing.T) { + d, attempts := discoveryRefusingCreates(t, 3, &corev1.Node{ObjectMeta: metav1.ObjectMeta{Name: "worker-1"}}) + + if err := d.Start(context.Background()); err != nil { + t.Fatalf("Start: %v", err) + } + if !raised(t, d) { + t.Fatalf("the run was lost to a webhook that was not serving yet, after %d attempt(s)", *attempts) + } +} + +// A webhook that answers is an answer, and a rejected run is not retried until +// the deadline: the spec is what it is, and waiting cannot change it. +func TestARejectedRunIsNotRetried(t *testing.T) { + scheme := testsupport.NewScheme(t, corev1.AddToScheme) + attempts := 0 + inner := fake.NewClientBuilder().WithScheme(scheme). + WithObjects(&corev1.Node{ObjectMeta: metav1.ObjectMeta{Name: "worker-1"}}).Build() + denying := interceptor.NewClient(inner, interceptor.Funcs{ + Create: func(ctx context.Context, c client.WithWatch, obj client.Object, opts ...client.CreateOption) error { + attempts++ + return apierrors.NewForbidden( + schema.GroupResource{Group: "storage.simplyblock.io", Resource: "operatorops"}, + InitialDiscoveryName, + fmt.Errorf(`admission webhook "voperatorops.simplyblock.io" denied the request`)) + }, + }) + d := &InitialDiscovery{Client: denying, Namespace: theNamespace} + + if err := d.Start(context.Background()); err != nil { + t.Fatalf("Start: %v", err) + } + if attempts != 1 { + t.Errorf("a denial was retried %d times; a rejected spec is an answer", attempts) + } +} From d39bc1e4496a63be62b58bd437f27247161d41c1 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Sun, 20 Sep 2026 15:09:00 +0200 Subject: [PATCH 094/206] fix(operator): a refusal nobody can read is a refusal that was not delivered A discovery run that finds nothing writes why, and on a six-worker fleet the API server threw it away: Server rejected event (will not retry!) ... "initial-discovery..." is invalid: message: Invalid value: "": can have at most 1024 characters The message carried one clause per worker, so it was 1193 characters at six workers and 5609 at thirty-two: the larger the fleet, the more certain the loss. `kubectl describe operatorops` then showed a run that had failed for no stated reason, and the only surviving copy was the operator's own log, which is where this was found. The two halves of the answer do not compress the same way, and that is the whole of the fix. A device refusal repeats, because every worker of a uniform fleet declines its disks for the same reason, and the counts are the finding: a fleet with no disks and a fleet whose disks are all held are different answers. A worker refusal does not repeat, because it names the controllers it found and where they are, and one worker's addresses are not another's. So devices are counted, workers are listed, and the list is what gives way when the budget runs out. Worker attribution is not lost either way: every device refusal already went out as its own DeviceDeclined event naming the worker and the device. boundEventMessage is the second half, at the one funnel every event of this reconciler leaves through. It is not the design, it is what keeps the message that outgrew the bound from being dropped in silence, and the cut says it happened so that a truncated list is not read as a whole one. There are 127 other emission sites in the operator and any of them can be built from a list. Six recorded expectations changed, each read against its case.md before being re-recorded. filt-10 is the clearest gain: the old wording nested "1 the probe refused it: Partitioned, 1 the probe refused it: Partitioned, Mounted" inside a parenthetical. Co-Authored-By: Claude Opus 5 (1M context) --- .../deployment/operatorops_controller.go | 37 +++- .../deployment/operatorops_events_test.go | 159 ++++++++++++++++++ .../expected-error.txt | 2 +- .../expected-error.txt | 2 +- .../expected-error.txt | 2 +- .../expected-error.txt | 2 +- .../expected-error.txt | 2 +- .../expected-error.txt | 2 +- operator/internal/discovery/plan.go | 65 +++++++ 9 files changed, 262 insertions(+), 11 deletions(-) create mode 100644 operator/internal/controllers/deployment/operatorops_events_test.go diff --git a/operator/internal/controllers/deployment/operatorops_controller.go b/operator/internal/controllers/deployment/operatorops_controller.go index d94e69d63..f6eb0eecd 100644 --- a/operator/internal/controllers/deployment/operatorops_controller.go +++ b/operator/internal/controllers/deployment/operatorops_controller.go @@ -29,7 +29,6 @@ import ( "errors" "fmt" "slices" - "strings" "time" batchv1 "k8s.io/api/batch/v1" @@ -560,14 +559,18 @@ func (r *OperatorOpsReconciler) write( for _, refusal := range plan.RefusalLines() { r.event(ops, corev1.EventTypeNormal, DeviceDeclined, refusal) } - why := plan.Explain() - if len(why) == 0 { + if len(plan.Explain()) == 0 { // No machine was refused by name, so the run had no worker to refuse. return false, refusef(OperationFailed, "no worker has a device this run would use: %s", plan.Summary()) } + // Counted rather than listed per worker, because this message is carried + // by an event and an event is refused past 1024 characters. The + // per-worker detail is not lost: every device refusal above went out as + // its own DeviceDeclined event. return false, refusef(OperationFailed, - "no worker has a device this run would use: %s", strings.Join(why, "; ")) + "no worker has a device this run would use: %s", + plan.ExplainWithin(maxEventMessage-len(refusalPreamble))) } config, notes := r.draftFor(ops, spec, plan) @@ -865,7 +868,7 @@ func (r *OperatorOpsReconciler) event( if action == "" { action = string(ops.Spec.Action) } - r.Recorder.Eventf(ops, nil, eventType, reason, action, "%s", message) + r.Recorder.Eventf(ops, nil, eventType, reason, action, "%s", boundEventMessage(message)) } // schedulable reports whether a node is one work can be placed on. @@ -958,3 +961,27 @@ func (r *OperatorOpsReconciler) SetupWithManager(mgr ctrl.Manager) error { Named("operatorops"). Complete(r) } + +// refusalPreamble is what every refusal of this kind opens with, and what the +// explanation has to fit inside the event alongside. +const refusalPreamble = "no worker has a device this run would use: " + +// maxEventMessage is what the Kubernetes event API takes. A longer message is +// not truncated by the API server, it is refused, and the recorder does not +// retry, so the bound is this side's to keep. +const maxEventMessage = 1024 + +// boundEventMessage cuts a message to what an event carries. +// +// It is the last guard rather than the design: a message built from a list is +// written to stay within the bound, and this is what keeps the one that was not +// from being dropped in silence. The cut says it happened, because a reader who +// cannot see that the text ends early will read a truncated list as the whole +// list. +func boundEventMessage(message string) string { + if len(message) <= maxEventMessage { + return message + } + const ellipsis = " […]" + return message[:maxEventMessage-len(ellipsis)] + ellipsis +} diff --git a/operator/internal/controllers/deployment/operatorops_events_test.go b/operator/internal/controllers/deployment/operatorops_events_test.go new file mode 100644 index 000000000..5d827fbeb --- /dev/null +++ b/operator/internal/controllers/deployment/operatorops_events_test.go @@ -0,0 +1,159 @@ +// What a refusal has to fit into, which is an event. +// +// The Kubernetes event API caps a message at 1024 characters and rejects a +// longer one outright, so a refusal whose length grows with the fleet is a +// refusal that stops being delivered on exactly the fleets where it matters +// most. These cases pin the bound and the accounting it has to keep. + +package deployment + +import ( + "fmt" + "strings" + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/client-go/tools/events" + + "github.com/simplyblock/atlas/blockdev" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" + "sigs.k8s.io/controller-runtime/pkg/client" +) + +// declinedFleet builds a run over workers whose every device is refused, in the +// shape a re-install finds: two disks carrying data and one holding the root +// filesystem, on each worker. +func declinedFleet(t *testing.T, workers int) discoveryCase { + t.Helper() + + ops := &simplyblockv1alpha2.OperatorOps{ + ObjectMeta: metav1.ObjectMeta{Name: opsName, Namespace: opsNamespace}, + Spec: simplyblockv1alpha2.OperatorOpsSpec{Action: simplyblockv1alpha2.OperatorOpsActionDiscover}, + } + + var objects []client.Object + for i := 1; i <= workers; i++ { + node := fmt.Sprintf("worker-%02d", i) + ops.Status.Workers = append(ops.Status.Workers, node) + objects = append(objects, &corev1.Node{ + ObjectMeta: metav1.ObjectMeta{Name: node}, + Status: corev1.NodeStatus{ + Addresses: []corev1.NodeAddress{{Type: corev1.NodeInternalIP, Address: fmt.Sprintf("10.10.10.%d", i)}}, + }, + }) + objects = append(objects, declinedReport(t, node)) + } + return discoveryCase{Ops: ops, Objects: objects} +} + +// declinedReport is one worker's probe result with nothing usable on it. +func declinedReport(t *testing.T, node string) *corev1.ConfigMap { + t.Helper() + + const tb = uint64(1) << 40 + report := nodeprobe.Report{ + Version: nodeprobe.ReportVersion, + Node: node, + CPU: nodeprobe.CPU{ + OnlineCPUs: 32, PhysicalCores: 16, Sockets: 2, ThreadsPerCore: 2, HyperThreading: true, + NUMANodes: []nodeprobe.NUMACPUs{{Node: 0, OnlineCPUs: []int{0, 1, 2, 3}, PhysicalCores: 16}}, + }, + HugePages: []nodeprobe.HugePagePool{{ + SizeBytes: 1 << 30, Total: 32, Free: 32, + NUMANodes: []nodeprobe.NUMAHugePages{{Node: 0, Total: 32, Free: 32}}, + }}, + } + for i := range 2 { + report.Devices = append(report.Devices, nodeprobe.Device{ + Name: fmt.Sprintf("nvme%dn1", i), Path: fmt.Sprintf("/dev/nvme%dn1", i), + PCIAddress: fmt.Sprintf("0000:0%d:00.0", i), SizeBytes: 2 * tb, + Kind: string(blockdev.KindDisk), Transport: string(blockdev.TransportNVMe), + Available: false, Content: "Foreign", + Rejections: []nodeprobe.Rejection{{ + Reason: string(blockdev.ReasonNotBlank), + Detail: "no known signature, and the probed regions are not empty: first non-zero byte at 0", + }}, + }) + } + report.Devices = append(report.Devices, nodeprobe.Device{ + Name: "nvme2n1", Path: "/dev/nvme2n1", PCIAddress: "0000:02:00.0", SizeBytes: tb, + Kind: string(blockdev.KindDisk), Transport: string(blockdev.TransportNVMe), + Available: false, + Rejections: []nodeprobe.Rejection{ + {Reason: string(blockdev.ReasonMounted), Detail: "mounted at [/boot /sysroot /var]"}, + {Reason: string(blockdev.ReasonBusy), Detail: "the kernel refused an exclusive open, and gives no reason"}, + {Reason: string(blockdev.ReasonPartitioned), Detail: "the device carries the partitions [nvme2n1p1 nvme2n1p2]"}, + }, + }) + + cm, err := nodeprobe.ConfigMap(opsNamespace, opsName, nil, report) + if err != nil { + t.Fatalf("render the report ConfigMap: %v", err) + } + return cm +} + +// TestTheRefusalFitsAnEventWhateverTheFleetSize covers the message a run writes +// when no worker has a usable device. +// +// Regression: 2026-09-20-discovery-refusal-too-long-for-an-event — the message +// carried one clause per worker, so at six workers it reached 1193 characters +// and the API server refused the event with `is invalid: message: Invalid +// value: "": can have at most 1024 characters`, twice, and the recorder does +// not retry. `kubectl describe operatorops` then showed a run that had failed +// for no stated reason, and the only surviving copy of why was the operator's +// own log. It grows with the fleet, so the larger the cluster the more certain +// the loss. +func TestTheRefusalFitsAnEventWhateverTheFleetSize(t *testing.T) { + for _, workers := range []int{6, 32} { + t.Run(fmt.Sprintf("%d-workers", workers), func(t *testing.T) { + got := runDiscoveryCase(t, declinedFleet(t, workers)) + + if got.Failure == "" { + t.Fatal("a fleet with nothing usable produced no refusal") + } + if len(got.Failure) > maxEventMessage { + t.Errorf("the refusal is %d characters and an event carries %d, so it is dropped:\n%s", + len(got.Failure), maxEventMessage, got.Failure) + } + }) + } +} + +// The bound is not an excuse to stop saying what happened: the counts are what +// distinguishes a fleet with no disks from one whose disks are all held. +func TestTheRefusalStillAccountsForEveryDevice(t *testing.T) { + const workers = 6 + got := runDiscoveryCase(t, declinedFleet(t, workers)) + + for _, want := range []string{"18", string(blockdev.ReasonNotBlank), string(blockdev.ReasonMounted)} { + if !strings.Contains(got.Failure, want) { + t.Errorf("the refusal does not mention %q, so a reader cannot tell what was found:\n%s", + want, got.Failure) + } + } +} + +// TestALongEventIsDeliveredRatherThanDropped covers every other message this +// reconciler emits, since any of them may be built from a list. +func TestALongEventIsDeliveredRatherThanDropped(t *testing.T) { + recorder := events.NewFakeRecorder(4) + r := &OperatorOpsReconciler{Recorder: recorder} + ops := &simplyblockv1alpha2.OperatorOps{ + ObjectMeta: metav1.ObjectMeta{Name: opsName, Namespace: opsNamespace}, + } + + r.event(ops, corev1.EventTypeWarning, OperationFailed, strings.Repeat("x", 4000)) + + select { + case line := <-recorder.Events: + if got := len(eventMessage(line)); got > maxEventMessage { + t.Errorf("the event carries %d characters and the API server takes %d", got, maxEventMessage) + } + default: + t.Fatal("no event was recorded") + } +} diff --git a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-15-every-disk-is-an-attached-volume/expected-error.txt b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-15-every-disk-is-an-attached-volume/expected-error.txt index c73155346..31254c0f9 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/dev/dev-15-every-disk-is-an-attached-volume/expected-error.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/dev/dev-15-every-disk-is-an-attached-volume/expected-error.txt @@ -1 +1 @@ -no worker has a device this run would use: worker-01: no device of it survived the device rules (3 devices declined by simplyblock volume: it is a simplyblock logical volume of cluster c30a691a-1d2e-4f3a-9b8c-5d6e7f809a1b, which this fleet already serves) +no worker has a device this run would use: 3 device(s) across 1 worker(s), none usable: 3 declined by simplyblock volume: it is a simplyblock logical volume of cluster c30a691a-1d2e-4f3a-9b8c-5d6e7f809a1b, which this fleet already serves diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/expected-error.txt b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/expected-error.txt index ee8f2c3b3..3e9f0db23 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/expected-error.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-02-every-disk-in-the-fleet-mounted/expected-error.txt @@ -1 +1 @@ -no worker has a device this run would use: worker-01: no device of it survived the device rules (4 devices declined by available: the probe refused it: Mounted); worker-02: no device of it survived the device rules (4 devices declined by available: the probe refused it: Mounted); worker-03: no device of it survived the device rules (4 devices declined by available: the probe refused it: Mounted) +no worker has a device this run would use: 12 device(s) across 3 worker(s), none usable: 12 declined by available: the probe refused it: Mounted diff --git a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/expected-error.txt b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/expected-error.txt index b2a177666..96e449e59 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/expected-error.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/fail/fail-04-a-filter-that-excludes-the-fleet/expected-error.txt @@ -1 +1 @@ -no worker has a device this run would use: worker-01: no device of it survived the device rules (4 devices declined by size range: it is 3T and the range 100T-200T starts above it); worker-02: no device of it survived the device rules (4 devices declined by size range: it is 3T and the range 100T-200T starts above it); worker-03: no device of it survived the device rules (4 devices declined by size range: it is 3T and the range 100T-200T starts above it) +no worker has a device this run would use: 12 device(s) across 3 worker(s), none usable: 12 declined by size range: it is 3T and the range 100T-200T starts above it diff --git a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-10-a-filter-that-matches-nothing/expected-error.txt b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-10-a-filter-that-matches-nothing/expected-error.txt index beb26ec8e..d9820102a 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/filt/filt-10-a-filter-that-matches-nothing/expected-error.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/filt/filt-10-a-filter-that-matches-nothing/expected-error.txt @@ -1 +1 @@ -no worker has a device this run would use: worker-01: no device of it survived the device rules (3 devices declined by allow and deny lists: it is not in the allow list, 2 devices declined by available: 1 the probe refused it: Partitioned, 1 the probe refused it: Partitioned, Mounted) +no worker has a device this run would use: 5 device(s) across 1 worker(s), none usable: 3 declined by allow and deny lists: it is not in the allow list; 1 declined by available: the probe refused it: Partitioned; 1 declined by available: the probe refused it: Partitioned, Mounted diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-01-every-disk-mounted/expected-error.txt b/operator/internal/controllers/deployment/testdata/discovery/held/held-01-every-disk-mounted/expected-error.txt index ca5fadb41..88766220c 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/held/held-01-every-disk-mounted/expected-error.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-01-every-disk-mounted/expected-error.txt @@ -1 +1 @@ -no worker has a device this run would use: worker-01: no device of it survived the device rules (4 devices declined by available: the probe refused it: Mounted) +no worker has a device this run would use: 4 device(s) across 1 worker(s), none usable: 4 declined by available: the probe refused it: Mounted diff --git a/operator/internal/controllers/deployment/testdata/discovery/held/held-02-every-disk-in-a-device-mapper-stack/expected-error.txt b/operator/internal/controllers/deployment/testdata/discovery/held/held-02-every-disk-in-a-device-mapper-stack/expected-error.txt index 0b9b4c155..1e43c5010 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/held/held-02-every-disk-in-a-device-mapper-stack/expected-error.txt +++ b/operator/internal/controllers/deployment/testdata/discovery/held/held-02-every-disk-in-a-device-mapper-stack/expected-error.txt @@ -1 +1 @@ -no worker has a device this run would use: worker-01: no device of it survived the device rules (4 devices declined by available: the probe refused it: Stacked) +no worker has a device this run would use: 4 device(s) across 1 worker(s), none usable: 4 declined by available: the probe refused it: Stacked diff --git a/operator/internal/discovery/plan.go b/operator/internal/discovery/plan.go index 5198748a9..eef253495 100644 --- a/operator/internal/discovery/plan.go +++ b/operator/internal/discovery/plan.go @@ -118,6 +118,71 @@ func (p Plan) DeviceCount() int { // sixteen network block devices and four disks, and a line saying the sixteen // were not whole disks is true, longer than the rest of the message, and not // the answer to anything. +// ExplainWithin is what Explain says, written to fit a budget. +// +// Explain writes a clause per worker, which is the detail a reviewer wants and a +// length that grows with the fleet. A Kubernetes event carries 1024 characters +// and is refused rather than truncated past it, so the run that fails on a large +// fleet is exactly the run whose reason never reaches anybody. +// +// The two halves of the answer do not compress the same way. A device refusal +// repeats: every worker of a uniform fleet declines its disks for the same +// reason, and the counts are the finding, since a fleet with no disks and a +// fleet whose disks are all held are different answers. A worker refusal does +// not: it names the controllers it found and where they are, and one worker's +// addresses are not another's. So the devices are counted and the workers are +// listed, and it is the list that gives way when the budget runs out. +func (p Plan) ExplainWithin(budget int) string { + declined := 0 + workers := map[string]bool{} + var deviceOrder []string + byDevice := map[string]int{} + workerLines := p.Explain() + + for _, refusal := range p.Refusals { + if refusal.Device == "" || refusal.PreFilter { + continue + } + workers[refusal.Worker] = true + declined++ + key := refusal.Rule + ": " + refusal.Reason + if _, counted := byDevice[key]; !counted { + deviceOrder = append(deviceOrder, key) + } + byDevice[key]++ + } + + // A fleet whose every refusal is about a device is the case the per-worker + // list says nothing extra about, so the counts replace it outright. + if declined > 0 && len(deviceOrder) > 0 { + counts := make([]string, 0, len(deviceOrder)) + for _, key := range deviceOrder { + counts = append(counts, fmt.Sprintf("%d declined by %s", byDevice[key], key)) + } + return fmt.Sprintf("%d device(s) across %d worker(s), none usable: %s", + declined, len(workers), strings.Join(counts, "; ")) + } + + // Otherwise, the refusals are about the machines, and what they name cannot + // be counted away. As many as the budget holds are written, and the rest are + // reported as a number so that the reader knows the list is not the whole of + // it. + kept, used := 0, 0 + for _, line := range workerLines { + cost := len(line) + len("; ") + if kept > 0 && used+cost > budget-len(fmt.Sprintf("; and %d more worker(s)", len(workerLines))) { + break + } + used += cost + kept++ + } + if kept >= len(workerLines) { + return strings.Join(workerLines, "; ") + } + return fmt.Sprintf("%s; and %d more worker(s)", + strings.Join(workerLines[:kept], "; "), len(workerLines)-kept) +} + func (p Plan) Explain() []string { type byRule struct { rule string From 1afe517f1b0f03e29ee71b3af31bcc40614d2a80 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Sun, 20 Sep 2026 15:09:08 +0200 Subject: [PATCH 095/206] feat(atlas-lib): the one on-disk signature this product writes and could not read A fleet that had held a simplyblock cluster reported that it had no disks at all. Every device was refused: NotBlank: no known signature, and the probed regions are not empty: first non-zero byte at 0 The bytes at 0 were `ALCEML_STORAGE`, the superblock a storage node writes. It was the only format the catalog did not know that this product itself produces, and neither blkid nor wipefs knows it either, which is also why `wipefs -a` does not remove one. So a re-install over its own disks read them as somebody else's data, and the run meant to show a fleet what it has said it had nothing. It is its own Content rather than another ContentForeign row, because the two answer different questions. Foreign content belongs to somebody else and the answer is always no. This content belongs to a simplyblock deployment, so the question is which one, and the device itself settles it: a disk a storage node is driving is bound to a userspace driver, which takes the block device away entirely, so a disk whose superblock can be read at all is one no node currently holds. The candidate is therefore offered rather than refused, and the case is written out in the switch rather than left to fall through it, because a content added later must not become available by default the way this one would have. A disk something is still holding is refused for the holding, which is the reason that matters: the superblock says whose it was and the usage says whether it is free now. Offering it costs nothing that was not already guarded. A device an existing StorageNode names is excluded before this is reached, and what a run produces is a draft nobody has approved. The image is captured from a device this product wrote, on a Micron 7450 in an OKD worker, because no local tool can produce that superblock: it is the one fixture in the tree that cannot be regenerated by formatting a loop device, and its manifest says so. Its head carries 112 non-zero bytes and its tail is entirely zero, so no data of anybody's is in the capture. The magic is matched at an absolute offset rather than against LBA 0, which the capture is the evidence for: the device is 4Kn, and an offset counted in blocks would have been resolved against the wrong one. Co-Authored-By: Claude Opus 5 (1M context) --- atlas-lib/README.md | 8 +- atlas-lib/blockdev/alceml_candidate_test.go | 72 ++++++++++++++++++ atlas-lib/blockdev/candidate.go | 11 +++ atlas-lib/blockdev/content.go | 14 ++++ atlas-lib/blockdev/content_test.go | 4 + atlas-lib/blockdev/signatures.go | 27 +++++++ .../testdata/images/alceml/head.bin.gz | Bin 0 -> 1225 bytes .../testdata/images/alceml/manifest.json | 15 ++++ .../testdata/images/alceml/tail.bin.gz | Bin 0 -> 1060 bytes .../design-device-content-detection.md | 49 ++++++++---- .../test-plan-device-content-detection.md | 25 +++--- 11 files changed, 197 insertions(+), 28 deletions(-) create mode 100644 atlas-lib/blockdev/alceml_candidate_test.go create mode 100644 atlas-lib/blockdev/testdata/images/alceml/head.bin.gz create mode 100644 atlas-lib/blockdev/testdata/images/alceml/manifest.json create mode 100644 atlas-lib/blockdev/testdata/images/alceml/tail.bin.gz diff --git a/atlas-lib/README.md b/atlas-lib/README.md index c7b9d6518..33ff23bcd 100644 --- a/atlas-lib/README.md +++ b/atlas-lib/README.md @@ -760,6 +760,10 @@ case reading.Content == blockdev.ContentFilesystem: // helper that formats on its own probe. case reading.Content == blockdev.ContentStackLayer: // An LVM physical-volume label. Activate the stack; do not pvcreate. +case reading.Content == blockdev.ContentSimplyblock: + // This product's own storage superblock, from a deployment that is gone: a + // device a storage node is driving is bound to a userspace driver and is + // not a block device at all, so one readable here is held by nobody. default: // ContentForeign: somebody else's data. Refuse, and say what was found // through reading.Detail. @@ -771,7 +775,9 @@ which is the path a node service takes rather than resolving one from `/dev`. _Today:_ the reading is built and covered by images captured from devices real tools formatted (`atlas-lib/blockdev/testdata/images`, regenerated by -`hack/blockdev/capture-image.sh`). Its consumers are still on the blkid probe: +`hack/blockdev/capture-image.sh`). The `alceml` image is the exception that has +to be captured off a device a storage node wrote, since no formatting tool +produces that superblock and neither blkid nor wipefs knows it. Its consumers are still on the blkid probe: the CSI driver's `NodeStageVolume` calls `BlkidProber` through `probeDiskFormat` in `csi-driver/internal/csi/node`, and moving it onto `Read` is Phase 1 of [`design-device-content-detection.md`](../operator/docs/designs/design-device-content-detection.md). diff --git a/atlas-lib/blockdev/alceml_candidate_test.go b/atlas-lib/blockdev/alceml_candidate_test.go new file mode 100644 index 000000000..1cf745ef3 --- /dev/null +++ b/atlas-lib/blockdev/alceml_candidate_test.go @@ -0,0 +1,72 @@ +// What a caller is told about a device this product already took. +// +// The reading is half the answer and the candidate is the other half, because a +// discovery run acts on whether the device was offered rather than on what the +// bytes were called. + +package blockdev + +import ( + "context" + "testing" +) + +// TestADeviceThisProductWroteIsOffered covers the disks a re-install finds. +// +// Regression: 2026-09-20-alceml-superblock-read-as-unknown-bytes — a fleet that +// had held a simplyblock cluster reported having no disks at all. Every device +// carried an alceml superblock, which blkid and wipefs both read as nothing, so +// the reading called it bytes matching no known signature and the candidate was +// refused NotBlank. The run meant to show a fleet what it has said it had +// nothing, and nothing distinguished that from a fleet whose disks genuinely +// belong to something else. +func TestADeviceThisProductWroteIsOffered(t *testing.T) { + const size = 6251233968 * 512 + taken := blank(size) + taken.bytes[0] = []byte("ALCEML_STORAGE\x00\x00") + + in := inspectorOver(t, map[string]sparse{"nvme0n1": taken}, nil) + + cands, err := in.Candidates(context.Background()) + if err != nil { + t.Fatalf("collect the candidates: %v", err) + } + + got := found(t, cands, "nvme0n1") + if got.RejectedFor(ReasonNotBlank) { + t.Error("a device this product wrote is refused as carrying an unknown format") + } + if !got.Available() { + t.Errorf("a device no storage node holds was not offered: %v", got.Rejections) + } + if got.Reading.Content != ContentSimplyblock { + t.Errorf("it reads as %v, want %v", got.Reading.Content, ContentSimplyblock) + } +} + +// A disk this product wrote and something is still using is refused for the +// use, which is the reason that matters: the superblock says whose it was and +// the usage says whether it is free now. +func TestADeviceThisProductWroteIsStillRefusedWhileHeld(t *testing.T) { + const size = 6251233968 * 512 + taken := blank(size) + taken.bytes[0] = []byte("ALCEML_STORAGE\x00\x00") + + held := func(path string) error { + if path == "/dev/nvme0n1" { + return ErrDeviceBusy + } + return nil + } + in := inspectorOver(t, map[string]sparse{"nvme0n1": taken}, held) + + cands, err := in.Candidates(context.Background()) + if err != nil { + t.Fatalf("collect the candidates: %v", err) + } + + got := found(t, cands, "nvme0n1") + if got.Available() { + t.Error("a device something is still using was offered") + } +} diff --git a/atlas-lib/blockdev/candidate.go b/atlas-lib/blockdev/candidate.go index 6b0748a04..f8ca26391 100644 --- a/atlas-lib/blockdev/candidate.go +++ b/atlas-lib/blockdev/candidate.go @@ -273,6 +273,17 @@ func judge(ctx context.Context, prober *Prober, disk Disk, usage Usage) Candidat c.reject(ReasonNotBlank, reading.Detail) case ContentFilesystem, ContentStackLayer: c.reject(ReasonNotBlank, reading.Detail) + case ContentSimplyblock: + // Not a rejection, and the reading is what says so: a device a storage + // node is driving is bound to a userspace driver, which takes the block + // device away, so a device whose superblock can be read here is one no + // node currently holds. What is left on it is a previous deployment's, + // and the reading carries that to the caller, which decides whether to + // take it back. + // + // It is written out rather than left to fall through the switch, because + // a content this package adds later must not become available by + // default the way this one would have. case ContentUnknown: // Read never returns it, and a reading that carries it anyway is one // nothing established. Refusing is the only safe reading of that. diff --git a/atlas-lib/blockdev/content.go b/atlas-lib/blockdev/content.go index e9da85514..64cf4f375 100644 --- a/atlas-lib/blockdev/content.go +++ b/atlas-lib/blockdev/content.go @@ -46,6 +46,18 @@ const ( // ContentForeign means the device carries something else: a recognized // format this driver does not create, or bytes that match nothing known. ContentForeign + + // ContentSimplyblock means the device carries this product's own storage + // superblock, written by a storage node rather than by any formatting tool. + // + // It is separate from ContentForeign because the two answer different + // questions for a caller deciding what to offer. Foreign content belongs to + // somebody else and the answer is always no. This content belongs to a + // simplyblock deployment, so the question becomes which one: a device a + // storage node is driving is bound to a userspace driver and is not a block + // device at all, so one that is readable here is a device no node currently + // holds. + ContentSimplyblock ) // String names the content for a log line, an event, and a test failure. @@ -59,6 +71,8 @@ func (c Content) String() string { return "StackLayer" case ContentForeign: return "Foreign" + case ContentSimplyblock: + return "Simplyblock" case ContentUnknown: return "Unknown" default: diff --git a/atlas-lib/blockdev/content_test.go b/atlas-lib/blockdev/content_test.go index a0fda5320..5c8a5de4d 100644 --- a/atlas-lib/blockdev/content_test.go +++ b/atlas-lib/blockdev/content_test.go @@ -52,6 +52,10 @@ var catalog = []want{ {"mdraid-11", ContentForeign, "linux_raid_member", "U-12: md metadata 1.1, at offset 0"}, {"mdraid-12", ContentForeign, "linux_raid_member", "U-12: md metadata 1.2, at offset 4096"}, {"zfs", ContentForeign, "zfs_member", "U-12: ZFS vdev labels"}, + // The one capture no local tool can reproduce: the superblock is written by + // a storage node, so the image comes off a device this product wrote. + {"alceml", ContentSimplyblock, "simplyblock_alceml", + "U-65: an alceml superblock, which blkid and wipefs both read as nothing"}, // U-15: the only reading that permits a format. {"blank", ContentBlank, "", "U-15: a device that has never been written to"}, } diff --git a/atlas-lib/blockdev/signatures.go b/atlas-lib/blockdev/signatures.go index a41389304..2444fbc0d 100644 --- a/atlas-lib/blockdev/signatures.go +++ b/atlas-lib/blockdev/signatures.go @@ -94,8 +94,15 @@ func detect(r regions) []find { var detectors = []func(regions) (find, bool){ detectExt, detectXFS, detectLVM2, detectLUKS, detectGPT, detectMBR, detectExFAT, detectFAT, detectBtrfs, detectSwap, detectMDRaid, detectZFS, + detectAlceml, } +// The storage superblock a storage node writes, at the very start of a device +// it has taken. The magic is a fixed sixteen bytes and the fields after it are +// not read here: what this answers is whose the device is, and the layout +// behind the magic is the storage node's to change. +var alcemlMagic = []byte("ALCEML_STORAGE\x00\x00") + // The ext superblock sits at 1024, so its magic is at 1080 and its three // feature words follow at 1116, 1120, and 1124. const ( @@ -112,6 +119,26 @@ const ( ext4RoCompat2 = 0x0100 | 0x0200 | 0x0400 // quota, bigalloc, metadata_csum ) +// detectAlceml names a device this product already took. +// +// It is the one signature here that no external tool writes, and the one whose +// absence was a real cost rather than a wording problem: blkid does not know it +// and neither does wipefs, so a device carrying it reads as bytes matching +// nothing, and a discovery run over a fleet that had held a simplyblock cluster +// reported that the fleet had no disks at all. Naming it is what lets a caller +// tell its own deployment's leftovers from somebody else's data. +// +// The magic is at offset 0 whatever the logical block size, so it is read as an +// absolute offset rather than against LBA 0. The captured image is 4Kn, and an +// offset counted in blocks would have been read against the wrong one. +func detectAlceml(r regions) (find, bool) { + if !r.eq(0, alcemlMagic) { + return find{}, false + } + return find{ContentSimplyblock, "simplyblock_alceml", 0, + "a simplyblock storage superblock at 0"}, true +} + // detectExt names the exact member of the ext family. The distinction is worth // making because a reading that rounded every ext filesystem up to ext4 would // disagree with the claim annotation for a volume this driver did not create, diff --git a/atlas-lib/blockdev/testdata/images/alceml/head.bin.gz b/atlas-lib/blockdev/testdata/images/alceml/head.bin.gz new file mode 100644 index 0000000000000000000000000000000000000000..7d159836c5e05acd1ecefc46fdcaeb870779c9e7 GIT binary patch literal 1225 zcmb2|=HOU-V|^AAb4F@nie6G?9>d#<`rgikB5V(mjdY}R%}$^4N!ooUR_FHlhw~Ra zRmzySaQec8i+WmSjsk1jYMJeuLXSQY7wA5jFyoNGmS~ytlh2>#et$Bm@?zdp@6+7f zH){k9h3_w8Vqjp9;Cd#%hKvjh42Kqc_0Q#qUjSr|g3%Bd4S^9B0(Q(Y MuPl}`FbFUJ0C$!Sl>h($ literal 0 HcmV?d00001 diff --git a/operator/docs/designs/design-device-content-detection.md b/operator/docs/designs/design-device-content-detection.md index 1b7b2543c..0a871078a 100644 --- a/operator/docs/designs/design-device-content-detection.md +++ b/operator/docs/designs/design-device-content-detection.md @@ -301,22 +301,23 @@ hangs and then exits 2. ## 5. Signature Catalog -| Format | `Content` | Where | Match | -|------------------|---------------------|-----------------------------------------------|------------------------------------------------------------------------------| -| ext2, ext3, ext4 | `ContentFilesystem` | 1080, little-endian `uint16` | `0xEF53`, with the feature words at 1116, 1120, and 1124 deciding the family | -| XFS | `ContentFilesystem` | 0, big-endian `uint32` | `XFSB` | -| LVM2 | `ContentStackLayer` | start of one of the first four 512-byte units | `LABELONE`, with `LVM2 001` at offset 24 of the label | -| LUKS | `ContentForeign` | 0 | `LUKS\xba\xbe` | -| GPT | `ContentForeign` | LBA 1, and the last LBA for the backup header | `EFI PART`, the GPT header signature | -| MBR | `ContentForeign` | 510 | `0x55AA` with a non-empty partition entry | -| FAT12, FAT16 | `ContentForeign` | 0x36, with the BPB at 0x0B validated | `FAT12 `, `FAT16 `, or `FAT ` | -| FAT32 | `ContentForeign` | 0x52, with the BPB at 0x0B validated | `FAT32 ` | -| exFAT | `ContentForeign` | 3 | `EXFAT ` | -| Btrfs | `ContentForeign` | 65600 | `_BHRfS_M` | -| swap | `ContentForeign` | page size minus 10 | `SWAPSPACE2` or `SWAP-SPACE` | -| md-raid | `ContentForeign` | 0, 4096, or 8 KiB from the end | `0xa92b4efc` | -| ZFS | `ContentForeign` | vdev label offsets | `0x00bab10c` | -| anything else | `ContentForeign` | anywhere in the probed regions | a non-zero byte | +| Format | `Content` | Where | Match | +|------------------|----------------------|-----------------------------------------------|------------------------------------------------------------------------------| +| ext2, ext3, ext4 | `ContentFilesystem` | 1080, little-endian `uint16` | `0xEF53`, with the feature words at 1116, 1120, and 1124 deciding the family | +| XFS | `ContentFilesystem` | 0, big-endian `uint32` | `XFSB` | +| LVM2 | `ContentStackLayer` | start of one of the first four 512-byte units | `LABELONE`, with `LVM2 001` at offset 24 of the label | +| LUKS | `ContentForeign` | 0 | `LUKS\xba\xbe` | +| GPT | `ContentForeign` | LBA 1, and the last LBA for the backup header | `EFI PART`, the GPT header signature | +| MBR | `ContentForeign` | 510 | `0x55AA` with a non-empty partition entry | +| FAT12, FAT16 | `ContentForeign` | 0x36, with the BPB at 0x0B validated | `FAT12 `, `FAT16 `, or `FAT ` | +| FAT32 | `ContentForeign` | 0x52, with the BPB at 0x0B validated | `FAT32 ` | +| exFAT | `ContentForeign` | 3 | `EXFAT ` | +| Btrfs | `ContentForeign` | 65600 | `_BHRfS_M` | +| swap | `ContentForeign` | page size minus 10 | `SWAPSPACE2` or `SWAP-SPACE` | +| md-raid | `ContentForeign` | 0, 4096, or 8 KiB from the end | `0xa92b4efc` | +| ZFS | `ContentForeign` | vdev label offsets | `0x00bab10c` | +| alceml | `ContentSimplyblock` | 0 | `ALCEML_STORAGE\0\0`, the superblock a storage node writes | +| anything else | `ContentForeign` | anywhere in the probed regions | a non-zero byte | **An offset counted in logical blocks is resolved against the device, not against 512.** The GPT header is at LBA 1 and its backup is at the last LBA, which is @@ -359,6 +360,22 @@ therefore match the MBR row, and that ambiguity is cosmetic rather than consequential: both readings are `ContentForeign`, both refuse, and only the wording of the refusal differs. +**The one signature this product writes is the one no external tool can.** An +alceml superblock is written by a storage node rather than by any formatting +tool, and neither `blkid` nor `wipefs` knows it, so before it was cataloged a +device carrying one read as bytes matching nothing and was refused as foreign. +A fleet that had held a simplyblock cluster therefore reported that it had no +disks at all, which is the opposite of what a discovery run is for. It is its own +`Content` rather than another `ContentForeign` row because the two answer +different questions: foreign content is somebody else's and the answer is always +no, while this content is a previous deployment's and the question is whether the +deployment is still there. A device a storage node is driving is bound to a +userspace driver, which takes the block device away entirely, so a device whose +superblock can be read at all is one no node currently holds. That is why the +reading offers it rather than refusing it, and why the offer costs nothing: the +run still excludes a device an existing `StorageNode` names, and what it produces +is a draft nobody has approved. + **The order of evaluation is head signatures, then tail signatures, then the zero test.** A device carrying both an LVM label and a stale filesystem signature is reported as the LVM label, because the label is at the lower offset and is what diff --git a/operator/docs/tests/test-plan-device-content-detection.md b/operator/docs/tests/test-plan-device-content-detection.md index 51ffe5a0b..060c602b6 100644 --- a/operator/docs/tests/test-plan-device-content-detection.md +++ b/operator/docs/tests/test-plan-device-content-detection.md @@ -87,6 +87,9 @@ against its own manifest. | U-62 | Every fixture's bytes match the checksum its manifest records | Positive | `TestFixturesMatchTheirManifests` | | U-63 | ext4 images captured from two `e2fsprogs` generations both decode as ext4 | Regression | — | | U-64 | An md metadata-1.1 member, which `blkid` reports nothing for and exits 2 on, is not `ContentBlank` | Regression | `TestMdMetadata11IsNotBlankEvenThoughBlkidSaysNothing` | +| U-65 | An alceml superblock, which blkid and wipefs both read as nothing, is `ContentSimplyblock` | Regression | `TestReadingOfCapturedImages` | +| U-66 | A device carrying only that superblock is offered, since a driven one is not a block device at all | Regression | `TestADeviceThisProductWroteIsOffered` | +| U-67 | The same device while the kernel will not hand it over is refused for the use | Negative | `TestADeviceThisProductWroteIsStillRefusedWhileHeld` | ### The Blank Rule (design §4.1) @@ -288,17 +291,17 @@ the old one accepted, before the shadow is removed. Which topologies the matrix exercises. An axis with no bearing on the reading is argued rather than listed as a gap. -| Axis | Values covered | IDs | Not covered | -|-----------------|------------------------------------------------------------------------------------------------------------------|---------------------------------------|-----------------------------------------------------------------------| -| Device content | blank, ext2, ext3, ext4, XFS, LVM label, LUKS, GPT, MBR, FAT16, FAT32, exFAT, Btrfs, swap, md-raid, ZFS, garbage | U-01…U-15, U-54…U-61, I-01…I-04, I-13 | A second stack layer, which does not exist yet (design §14 Q4) | -| Device state | healthy, all paths down, warm cache, slow | U-24…U-28, I-05…I-08 | **Serving zeros successfully: M-01, and it is design §7.2** | -| Device size | smaller than two regions, exactly two regions, normal, zero-length | U-21, U-22, U-23 | Multi-terabyte, where only the tail seek's cost differs | -| Stack layer | raw namespace, LVM physical volume | U-45…U-50, I-04, E-07 | VDO logical volume and a striped logical volume, both Phase 2 | -| Block size | 512-byte logical blocks, 4Kn | U-60, I-11 | A 512e device, whose logical size is what these offsets follow anyway | -| Cache state | cold, stale after a content change | I-08, I-09 | — | -| Namespace scope | single namespace | every `E-` row | Multi-namespace, argued below | -| Cluster size | single node | every `I-` row | Three-node and larger, argued below | -| Cluster count | single cluster | every `E-` row | Multi-cluster, argued below | +| Axis | Values covered | IDs | Not covered | +|-----------------|--------------------------------------------------------------------------------------------------------------------------|---------------------------------------|-----------------------------------------------------------------------| +| Device content | blank, ext2, ext3, ext4, XFS, LVM label, LUKS, GPT, MBR, FAT16, FAT32, exFAT, Btrfs, swap, md-raid, ZFS, alceml, garbage | U-01…U-15, U-54…U-67, I-01…I-04, I-13 | A second stack layer, which does not exist yet (design §14 Q4) | +| Device state | healthy, all paths down, warm cache, slow | U-24…U-28, I-05…I-08 | **Serving zeros successfully: M-01, and it is design §7.2** | +| Device size | smaller than two regions, exactly two regions, normal, zero-length | U-21, U-22, U-23 | Multi-terabyte, where only the tail seek's cost differs | +| Stack layer | raw namespace, LVM physical volume | U-45…U-50, I-04, E-07 | VDO logical volume and a striped logical volume, both Phase 2 | +| Block size | 512-byte logical blocks, 4Kn | U-60, I-11 | A 512e device, whose logical size is what these offsets follow anyway | +| Cache state | cold, stale after a content change | I-08, I-09 | — | +| Namespace scope | single namespace | every `E-` row | Multi-namespace, argued below | +| Cluster size | single node | every `I-` row | Three-node and larger, argued below | +| Cluster count | single cluster | every `E-` row | Multi-cluster, argued below | The last three axes do not discriminate here. The reading is a node-local read of one device's bytes, taken with no reference to a namespace, a peer node, or a From 346fdfce7f82f87ff64c0b6af70494258827c486 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Sun, 20 Sep 2026 16:22:14 +0200 Subject: [PATCH 096/206] fix(operator): a PCI address reaches the allow list, not the namespace-name list Every entry of spec.config.deviceNames went into NVME_DEVICES, which the storage node's init container passes as --nvme-devices. That flag's argparse destination is nvme_names, and the backend resolves it with query_nvme_ssd_by_namespace_names, comparing each entry against the NameSpace values of `nvme list`: Address: 0000:01:00.0 | NameSpaces: ['nvme0n1'] Address: 0000:0b:00.0 | NameSpaces: ['nvme2n1'] A PCI address never matches one. So a discovery draft that named the disks by PCI address, which is what a discovery run produces for an NVMe cluster, selected no device at all: ERROR: There are no enough SSD devices on system, you may run 'sbctl sn clean-devices' ... PCI addresses are what --pci-allowed takes, and the API already says the list carries both spellings: deviceNames admits an address, a path, or a bare name, and the class every entry belongs to is the cluster's to declare. So the spelling decides the variable. atlas-lib's pci.ParseAddress is the predicate, rather than a second copy of the same regular expression, and the addresses are merged into the allow list with mergePCIAddresses so an explicit pcieAllowList and the device list become one list instead of one silently winning. Three of the four cases are red on the unchanged tree, reproducing a live worker's env.sh exactly. The fourth, a bare name staying in NVME_DEVICES, passes before and after: it is what stops the fix from moving everything to the allow list and breaking the spelling that was already right. The hostPath the generated configuration lands on carries a TODO rather than a change. It is one path per worker for every cluster, so two clusters on one worker replace each other's document, and which way that should be resolved -- key the path by cluster, or refuse the second node in admission -- is a question about whether co-tenancy is supported at all. Co-Authored-By: Claude Opus 5 (1M context) --- .../controllers/node/pernodeconfig.go | 35 ++++++- .../controllers/node/pernodeconfig_test.go | 93 +++++++++++++++++++ .../internal/utils/storage_node_workload.go | 16 ++++ 3 files changed, 142 insertions(+), 2 deletions(-) diff --git a/operator/internal/controllers/node/pernodeconfig.go b/operator/internal/controllers/node/pernodeconfig.go index f86062491..db67058eb 100644 --- a/operator/internal/controllers/node/pernodeconfig.go +++ b/operator/internal/controllers/node/pernodeconfig.go @@ -31,6 +31,7 @@ import ( "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil" atlaskube "github.com/simplyblock/atlas/kube" + "github.com/simplyblock/atlas/pci" "github.com/simplyblock/atlas/ptr" simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" @@ -200,9 +201,11 @@ func renderNodeConfig( fmt.Fprintf(&entry, "MAX_SUBSYS_COUNT=%s\n", ptr.StringOrDefault(cluster.Spec.MaxSubsystemCount, "")) fmt.Fprintf(&entry, "MAX_HUGE_PAGES_SIZE=%s\n", utils.ShellQuote(config.Sizing.MinHugePagesSize)) fmt.Fprintf(&entry, "VCPU_COUNT=%s\n", ptr.StringOrDefault(config.Sizing.VCPUCount, "")) - fmt.Fprintf(&entry, "PCI_ALLOWED=%s\n", utils.ShellQuote(strings.Join(config.PcieAllowList, ","))) + addresses, names := splitDeviceNames(config.DeviceNames) + fmt.Fprintf(&entry, "PCI_ALLOWED=%s\n", + utils.ShellQuote(strings.Join(mergePCIAddresses(config.PcieAllowList, addresses), ","))) fmt.Fprintf(&entry, "PCI_BLOCKED=%s\n", utils.ShellQuote(strings.Join(config.PcieDenyList, ","))) - fmt.Fprintf(&entry, "NVME_DEVICES=%s\n", utils.ShellQuote(strings.Join(config.DeviceNames, ","))) + fmt.Fprintf(&entry, "NVME_DEVICES=%s\n", utils.ShellQuote(strings.Join(names, ","))) fmt.Fprintf(&entry, "DEVICE_MODEL=%s\n", utils.ShellQuote(config.PcieModel)) fmt.Fprintf(&entry, "SIZE_RANGE=%s\n", utils.ShellQuote(config.DriveSizeRange)) if jm := config.JournalManager; jm != nil { @@ -215,6 +218,34 @@ func renderNodeConfig( return entry.String() } +// splitDeviceNames sorts one device list into the two variables the storage +// node's configuration script reads it through. +// +// spec.config.deviceNames carries both spellings a device is named by, and the +// backend resolves them by different means: --pci-allowed takes PCI addresses +// and matches them against the controllers sysfs enumerates, while +// --nvme-devices is the namespace-name channel, whose argparse destination is +// nvme_names and whose lookup compares each entry against the NameSpace values +// of `nvme list` — nvme0n1 and its siblings. A PCI address sent through the +// second matches nothing, and matching nothing is not an error there: the +// configure selects no device, fails with `There are no enough SSD devices on +// system`, and its return value is discarded, so the pod starts on whatever +// configuration the host already had. +// +// So the spelling decides the variable, which is what the API says it does: +// design-storagenode.md's deviceNames admits an address, a path, or a bare +// name, and the class it belongs to is the cluster's to declare. +func splitDeviceNames(devices []string) (addresses, names []string) { + for _, device := range devices { + if address, err := pci.ParseAddress(device); err == nil { + addresses = append(addresses, address) + continue + } + names = append(names, device) + } + return addresses, names +} + // mergeAllowedIntoEntry rewrites one entry's PCI_ALLOWED to include the addresses // a migration is binding on the target host. // diff --git a/operator/internal/controllers/node/pernodeconfig_test.go b/operator/internal/controllers/node/pernodeconfig_test.go index 346c19023..7add6076e 100644 --- a/operator/internal/controllers/node/pernodeconfig_test.go +++ b/operator/internal/controllers/node/pernodeconfig_test.go @@ -20,6 +20,8 @@ import ( corev1 "k8s.io/api/core/v1" "sigs.k8s.io/controller-runtime/pkg/client" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" ) // aRenderedEntry is one worker's entry as the ordinary pass writes it. @@ -147,3 +149,94 @@ func TestAListIsReadBackTheWayItWasWritten(t *testing.T) { } } } + +// renderedFor is the entry the ordinary pass writes for one node's device list. +func renderedFor(t *testing.T, class simplyblockv1alpha2.StorageClusterDeviceClass, names ...string) map[string]string { + t.Helper() + cluster := &simplyblockv1alpha2.StorageCluster{ + Spec: simplyblockv1alpha2.StorageClusterSpec{DeviceClass: class}, + } + node := &simplyblockv1alpha2.StorageNode{ + Spec: simplyblockv1alpha2.StorageNodeSpec{ + Config: simplyblockv1alpha2.StorageNodeConfig{DeviceNames: names}, + }, + } + out := map[string]string{} + for _, line := range strings.Split(renderNodeConfig(cluster, node), "\n") { + key, value, found := strings.Cut(line, "=") + if found { + out[key] = strings.Trim(value, "'") + } + } + return out +} + +// TestAPCIAddressReachesTheAllowList covers the device list of an NVMe cluster. +// +// Regression: 2026-09-20-pci-addresses-written-to-the-namespace-name-variable — +// every entry of spec.config.deviceNames went to NVME_DEVICES, which the storage +// node's init container passes as --nvme-devices, whose argparse destination is +// nvme_names and which the backend resolves with +// query_nvme_ssd_by_namespace_names against `nvme list` namespace names such as +// nvme0n1. A PCI address never matches one, so the configure selected no device +// and failed with `There are no enough SSD devices on system`. Its return value +// is discarded by node_configure.py, so the init container still exited 0, the +// pod started on whatever configuration the host already had, and the node_add +// that read it was refused for carrying the wrong device class. PCI addresses +// are what --pci-allowed takes. +func TestAPCIAddressReachesTheAllowList(t *testing.T) { + got := renderedFor(t, simplyblockv1alpha2.StorageClusterDeviceClassNVMe, "0000:01:00.0", "0000:0b:00.0") + + if allowed := got["PCI_ALLOWED"]; allowed != "0000:01:00.0,0000:0b:00.0" { + t.Errorf("PCI_ALLOWED = %q, and a PCI address is what the allow list takes", allowed) + } + if devices := got["NVME_DEVICES"]; devices != "" { + t.Errorf("NVME_DEVICES = %q, which the backend matches against namespace names", devices) + } +} + +// A bare name is a namespace name, which is exactly what NVME_DEVICES is +// resolved against, so it stays there. +func TestABareNameStaysANamespaceName(t *testing.T) { + got := renderedFor(t, simplyblockv1alpha2.StorageClusterDeviceClassNVMe, "nvme0n1", "nvme2n1") + + if devices := got["NVME_DEVICES"]; devices != "nvme0n1,nvme2n1" { + t.Errorf("NVME_DEVICES = %q, and a bare name is a namespace name", devices) + } + if allowed := got["PCI_ALLOWED"]; allowed != "" { + t.Errorf("PCI_ALLOWED = %q, and a namespace name is not a PCI address", allowed) + } +} + +// One list carries both spellings, and each goes where the backend reads it. +func TestBothSpellingsGoWhereTheyAreRead(t *testing.T) { + got := renderedFor(t, simplyblockv1alpha2.StorageClusterDeviceClassNVMe, "0000:01:00.0", "nvme2n1") + + if allowed := got["PCI_ALLOWED"]; allowed != "0000:01:00.0" { + t.Errorf("PCI_ALLOWED = %q", allowed) + } + if devices := got["NVME_DEVICES"]; devices != "nvme2n1" { + t.Errorf("NVME_DEVICES = %q", devices) + } +} + +// An explicit allow list and PCI addresses in the device list are one allow +// list, not two, and neither silently drops the other. +func TestTheAllowListAndTheDeviceListAreMerged(t *testing.T) { + cluster := &simplyblockv1alpha2.StorageCluster{ + Spec: simplyblockv1alpha2.StorageClusterSpec{DeviceClass: simplyblockv1alpha2.StorageClusterDeviceClassNVMe}, + } + node := &simplyblockv1alpha2.StorageNode{ + Spec: simplyblockv1alpha2.StorageNodeSpec{ + Config: simplyblockv1alpha2.StorageNodeConfig{ + PcieAllowList: []string{"0000:02:00.0"}, + DeviceNames: []string{"0000:01:00.0"}, + }, + }, + } + entry := renderNodeConfig(cluster, node) + if !strings.Contains(entry, "PCI_ALLOWED='0000:01:00.0,0000:02:00.0'") && + !strings.Contains(entry, "PCI_ALLOWED='0000:02:00.0,0000:01:00.0'") { + t.Errorf("the allow list and the device list did not merge:\n%s", entry) + } +} diff --git a/operator/internal/utils/storage_node_workload.go b/operator/internal/utils/storage_node_workload.go index 43d7782c0..cdb3652e0 100644 --- a/operator/internal/utils/storage_node_workload.go +++ b/operator/internal/utils/storage_node_workload.go @@ -170,6 +170,22 @@ fi` }, }, { + // The node configuration the init container generates and the + // storage node reads, at one path per worker rather than per + // cluster: the backend's own constant is + // /etc/simplyblock/sn_config_file, which this directory backs. + // + // TODO: it is therefore shared by every storage node on the worker, + // and a second cluster's node regenerates the whole document rather + // than taking a slot in it, so two clusters on one worker replace + // each other's configuration. Nothing here or in the backend keys a + // slot by cluster, and the operator does not see across clusters + // either: awaitSlot reads this cluster's nodes alone. Whether one + // worker may carry nodes of two clusters at all is the question to + // settle first — if it may, this path and the backend's constant + // have to carry the cluster; if it may not, a StorageNode whose + // worker already holds another cluster's node belongs in the + // validating webhook rather than in a file both of them overwrite. Name: "etc-simplyblock", VolumeSource: corev1.VolumeSource{ HostPath: &corev1.HostPathVolumeSource{ From bce0f52221820936ad2aeff3da83da058d67e553 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Sun, 20 Sep 2026 17:35:32 +0200 Subject: [PATCH 097/206] fix(operator): a node survives its worker's reboot, and is recognized when it returns MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two faults on one path, in the same region of one file, so they land together: the add of the first node of a fresh cluster could neither survive the reboot its own MachineConfig caused nor recognize the node it had produced. --- The node the add produced was invisible Resolving matched a backend node by comparing the control plane's mgmt_ip with the worker's Kubernetes InternalIP. Those are the same value only where the storage plane and Kubernetes share a network. Where they do not, they never match: control plane : worker-5_4422 mgmt_ip 192.168.10.15 online health_check true Kubernetes : worker-5 InternalIP 10.0.0.15 so every reading was dropped, the node the add had just produced was invisible, and the step reported that the add had finished without producing a node. It then re-posted once a second and would have failed the node on its deadline with a healthy backend node running underneath it. The operator was matching on a value the control plane was never going to report. node_address is sent as a per-pod DNS name; the backend resolves it and records the address the node holds on its own management interface, which is the storage plane's. The firmware UUID is the identity both sides already have. Kubernetes reads it from the host and reports it as Node.status.nodeInfo.systemUUID, the control plane reads it from the same place and reports it as system_uuid, and on the cluster this was found on the two are byte-identical and unique across all six workers. Neither side invents it and no network carries it. The address is kept for the control plane that reports no UUID, which is what every deployment matched this way so far has been. A UUID that disagrees is a node of another machine and does not fall back to the address, because falling back there would undo the whole point. StorageNodeReconciler.workerAddress is superseded by workerIdentity and removed; the workload reconciler keeps its own, which is still used. --- The worker went away in the middle of it The storage pool's MachineConfig is applied by rebooting the machine, so the first node of a fresh cluster is cordoned, drained, rebooted and uncordoned in the middle of being added: Cordon → Drain → Reboot into rendered-storage-6398f8-… → Uncordon → NodeDone Nothing on the provisioning path represented that. Every step of it waits on something the worker does, so none of them could progress meanwhile, and the step kept the deadline it entered with: the node would have been failed for a reboot it was correctly waiting out. AwaitingWorker is that wait written down. Posting and Resolving enter it when the worker stops being Ready or is cordoned, it carries its own hour-long budget -- the cordon-to-NodeDone cycle took eleven minutes on the cluster this was found on -- and it leaves for CheckingHost when the machine is back. The check runs before the deadline test, which is what stops a step spending its budget on a machine that is coming back. It returns to the first step rather than to the step it left, because what a reboot interrupts is not resumable in the middle: a Posting resumed after the fact would ask for a second node, and the add it was waiting on may well have landed while the machine was away. CheckingHost is where an existing backend node is found, and its Adopting edge takes that node over instead of adding it twice. It holds the claim while it waits, so maxParallelNodeAdds stays closed and no second worker is handed a configuration change on top of a reboot already running. That is also why only the two steps that have claimed the worker divert into it: awaitSlot reads a claimed sibling on the same worker as an add that already happened and a reason to skip Posting, and a node that never posted must not be read that way. The snapshot carries state and deadline and no origin, so the two cases cannot be told apart after the fact. The steps before the claim wait where they are instead, and AwaitingSlot now declines to take a slot at all while its own worker is away. Terminality was derived from the graph: IsTerminal means "no outgoing edges", so giving Resolving an edge would have ended the success path. The condition was already the step -- every non-terminal step returns a different successor while Resolving and Adopting return themselves -- and that is now what is tested. One rule was written and withdrawn: a worker the API server does not have is not reported away. slot_race_test.go caught it, arbitrating slots without Node fixtures, and the rule would have made every step of the path depend on a Node object that a node being torn down no longer has. It is a different condition with its own handling. Co-Authored-By: Claude Opus 5 (1M context) --- .../storage.simplyblock.io_storagenodes.yaml | 2 +- operator/api/v1alpha2/storagenode_types.go | 16 +- .../storage.simplyblock.io_storagenodes.yaml | 2 +- operator/dist/install.yaml | 2 +- .../crd-redesign/design-storagenode.md | 44 ++- operator/docs/tests/test-plan-storagenode.md | 80 +++--- .../controllers/node/awaitworker_test.go | 253 ++++++++++++++++++ .../internal/controllers/node/controlplane.go | 14 +- operator/internal/controllers/node/events.go | 8 + operator/internal/controllers/node/graphs.go | 21 +- .../internal/controllers/node/graphs_test.go | 6 +- .../controllers/node/resolve_identity_test.go | 160 +++++++++++ .../node/storagenode_controller.go | 170 +++++++++++- .../storage.simplyblock.io_storagenodes.yaml | 2 +- 14 files changed, 713 insertions(+), 67 deletions(-) create mode 100644 operator/internal/controllers/node/awaitworker_test.go create mode 100644 operator/internal/controllers/node/resolve_identity_test.go diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagenodes.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagenodes.yaml index 53e5f76a6..73682b427 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagenodes.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagenodes.yaml @@ -869,7 +869,7 @@ spec: type: object x-kubernetes-validations: - message: unknown step - rule: '!has(self.state) || self.state in [''CheckingHost'',''CheckingConfig'',''AwaitingSlot'',''Posting'',''Resolving'',''Adopting'']' + rule: '!has(self.state) || self.state in [''CheckingHost'',''CheckingConfig'',''AwaitingSlot'',''Posting'',''Resolving'',''Adopting'',''AwaitingWorker'']' uptime: description: Uptime is the node uptime as the control plane reports it. type: string diff --git a/operator/api/v1alpha2/storagenode_types.go b/operator/api/v1alpha2/storagenode_types.go index e83174683..011d66c17 100644 --- a/operator/api/v1alpha2/storagenode_types.go +++ b/operator/api/v1alpha2/storagenode_types.go @@ -61,7 +61,7 @@ const ( // StorageNodeStep is one step of the provisioning path. There is one graph rather // than a MultiConfig, because an entity has no spec.action to key one on. -// +kubebuilder:validation:Enum=CheckingHost;CheckingConfig;AwaitingSlot;Posting;Resolving;Adopting +// +kubebuilder:validation:Enum=CheckingHost;CheckingConfig;AwaitingSlot;Posting;Resolving;Adopting;AwaitingWorker type StorageNodeStep string const ( @@ -89,6 +89,18 @@ const ( // StorageNodeStepAdopting takes over a backend node the operator did not add. StorageNodeStepAdopting StorageNodeStep = "Adopting" + + // StorageNodeStepAwaitingWorker holds while the machine the node is being + // added to is not there: not Ready, or cordoned ahead of a drain. + // + // A worker is rebooted whenever a MachineConfig reaches it, and the storage + // pool's own config is one, so the first node of a fresh cluster is rebooted + // in the middle of being added. Every other step is waiting on something the + // worker does, so none of them can make progress meanwhile, and the step's + // deadline would be spent on a machine that is coming back. This is the wait + // written down: it carries its own budget, and it leaves for CheckingHost so + // the path is walked again rather than resumed in the middle of a claim. + StorageNodeStepAwaitingWorker StorageNodeStep = "AwaitingWorker" ) // JournalManagerSpec tunes the journal managers on one storage node. @@ -428,7 +440,7 @@ type StorageNodeStatus struct { // Step is the position of the provisioning machine, as the shared // statemachine.KubeSnapshot. The rule is what an Enum marker would do if a // marker could reach a field of a shared type. - // +kubebuilder:validation:XValidation:rule="!has(self.state) || self.state in ['CheckingHost','CheckingConfig','AwaitingSlot','Posting','Resolving','Adopting']",message="unknown step" + // +kubebuilder:validation:XValidation:rule="!has(self.state) || self.state in ['CheckingHost','CheckingConfig','AwaitingSlot','Posting','Resolving','Adopting','AwaitingWorker']",message="unknown step" // +optional Step statemachine.KubeSnapshot `json:"step,omitempty"` diff --git a/operator/config/crd/bases/storage.simplyblock.io_storagenodes.yaml b/operator/config/crd/bases/storage.simplyblock.io_storagenodes.yaml index 53e5f76a6..73682b427 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_storagenodes.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_storagenodes.yaml @@ -869,7 +869,7 @@ spec: type: object x-kubernetes-validations: - message: unknown step - rule: '!has(self.state) || self.state in [''CheckingHost'',''CheckingConfig'',''AwaitingSlot'',''Posting'',''Resolving'',''Adopting'']' + rule: '!has(self.state) || self.state in [''CheckingHost'',''CheckingConfig'',''AwaitingSlot'',''Posting'',''Resolving'',''Adopting'',''AwaitingWorker'']' uptime: description: Uptime is the node uptime as the control plane reports it. type: string diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index d85e89fa8..6660e6eb0 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -8138,7 +8138,7 @@ spec: type: object x-kubernetes-validations: - message: unknown step - rule: '!has(self.state) || self.state in [''CheckingHost'',''CheckingConfig'',''AwaitingSlot'',''Posting'',''Resolving'',''Adopting'']' + rule: '!has(self.state) || self.state in [''CheckingHost'',''CheckingConfig'',''AwaitingSlot'',''Posting'',''Resolving'',''Adopting'',''AwaitingWorker'']' uptime: description: Uptime is the node uptime as the control plane reports it. diff --git a/operator/docs/designs/crd-redesign/design-storagenode.md b/operator/docs/designs/crd-redesign/design-storagenode.md index ff23fba61..34f4f2ec9 100644 --- a/operator/docs/designs/crd-redesign/design-storagenode.md +++ b/operator/docs/designs/crd-redesign/design-storagenode.md @@ -760,11 +760,12 @@ transition of the node's own machine: │ a slot is free ▼ Posting ← Status().Patch with MergeFromWithOptimisticLock, then - │ POST /storage-nodes - ▼ - Resolving ← match this slot against the cluster's node list - │ UUID found - ▼ + │ POST /storage-nodes │ worker away → AwaitingWorker + ▼ │ + Resolving ← match this slot against │ worker away → AwaitingWorker + │ the cluster's node list │ + │ UUID found ▼ + ▼ AwaitingWorker ← back? → CheckingHost phase: Online ``` @@ -782,6 +783,33 @@ its sibling's step rather than its `postedAt`: a sibling at `Posting` or beyond means the worker has been claimed, and the second object enters `Resolving` directly. +**A worker that goes away mid-add is waited for, not counted against.** The +storage pool's MachineConfig is applied by rebooting the machine, so the first +node of a fresh cluster is cordoned, drained, rebooted, and uncordoned in the +middle of being added. Every step of this path waits on something the worker +does, so none of them can progress meanwhile, and a step that kept its deadline +through the reboot would fail the node for the reboot. `AwaitingWorker` is that +wait written down: `Posting` and `Resolving` enter it when the worker stops being +`Ready` or is cordoned, it carries its own hour-long budget, and it leaves for +`CheckingHost` when the machine is back. + +**It returns to the first step rather than to the step it left**, because what a +reboot interrupts is not resumable in the middle: a `Posting` resumed after the +fact would ask for a second node, and the add it was waiting on may well have +landed while the machine was away. `CheckingHost` is where a backend node that +already exists is found, and its `Adopting` edge takes that node over instead of +adding it twice. + +**It holds the claim while it waits.** A node in `AwaitingWorker` still counts as +having claimed its worker, so `maxParallelNodeAdds` stays closed and no second +worker is handed a configuration change on top of a reboot already running — the +same reason the cap exists at all. That is also why only the two steps that have +claimed the worker divert into it: a claimed sibling on the same worker is read as +an add that already happened and a reason to skip `Posting`, and a node that never +posted must not be read that way. The steps before the claim wait where they are, +and `AwaitingSlot` declines to take a slot at all while its worker is away, which +is what stops an add being posted against a machine that cannot answer it. + **`AwaitingSlot` is where two independent serialization rules live.** `maxParallelNodeAdds` caps how many workers may be in flight at once, counted by distinct worker rather than by object so that a two-socket host consumes one slot. @@ -808,7 +836,7 @@ type StorageNodePhase string // StorageNodeStep is one step of the provisioning path. There is one graph // rather than a MultiConfig, because an entity has no spec.action to key one on. -// +kubebuilder:validation:Enum=CheckingHost;CheckingConfig;AwaitingSlot;Posting;Resolving;Adopting +// +kubebuilder:validation:Enum=CheckingHost;CheckingConfig;AwaitingSlot;Posting;Resolving;Adopting;AwaitingWorker type StorageNodeStep string ``` @@ -1788,6 +1816,8 @@ starts and the operation's name is not something they know yet. | Provisioning is held because no fault group is declared | `Warning` | `FailureDomainMissing` | `StorageNode` | | Provisioning is held because the worker's API does not answer | `Warning` | `HostUnreachable` | `StorageNode` | | Provisioning is held because no node-add slot is free | `Normal` | `AwaitingSlot` | `StorageNode` | +| The worker is not Ready or is cordoned, so the step is held | `Warning` | `WorkerAway` | `StorageNode` | +| The worker came back and provisioning starts again | `Normal` | `WorkerReturned` | `StorageNode` | | The node was adopted rather than added | `Normal` | `NodeAdopted` | `StorageNode` | | The node came online | `Normal` | `NodeOnline` | `StorageNode` | | The node's pod reported a scheduling failure | `Warning` | `PodSchedulingFailed` | `StorageNode` | @@ -2093,7 +2123,7 @@ const ( // StorageNodeStep is one step of the provisioning path. There is one graph // rather than a MultiConfig, because an entity has no spec.action to key one on. -// +kubebuilder:validation:Enum=CheckingHost;CheckingConfig;AwaitingSlot;Posting;Resolving;Adopting +// +kubebuilder:validation:Enum=CheckingHost;CheckingConfig;AwaitingSlot;Posting;Resolving;Adopting;AwaitingWorker type StorageNodeStep string const ( diff --git a/operator/docs/tests/test-plan-storagenode.md b/operator/docs/tests/test-plan-storagenode.md index f89cca606..1c5ed0bc7 100644 --- a/operator/docs/tests/test-plan-storagenode.md +++ b/operator/docs/tests/test-plan-storagenode.md @@ -566,38 +566,54 @@ machine is declared beside them. The three lists that have to agree are the graph's states, the `Enum` marker on `status.step.state`, and the CEL rule the CRD carries. -| # | Scenario | Type | Test | -|-------|------------------------------------------------------------------------------------|----------|-------------------------------------------------------| -| U-217 | Every declared graph builds, including the ones the action under test does not use | Positive | `TestEveryActionDeclaresAGraph` | -| U-218 | An action with no declared graph: refused rather than stalled | Negative | `TestAStepThatBelongsToNoActionEndsTheOperation` | -| U-219 | `Remove` transitioning to `Promoting`: rejected as an illegal transition | Negative | `TestAStepOfAnotherActionIsRejected` | -| U-220 | `Migrate` transitioning to `Removing`: rejected as an illegal transition | Negative | `TestAStepOfAnotherActionIsRejected` | -| U-221 | `HostMaintenance` transitioning to `Suspending`: rejected | Negative | `TestAStepOfAnotherActionIsRejected` | -| U-222 | An empty `status.step`: restores to the action's declared initial state | Boundary | `TestTheFirstPassArmsTheStepAMachineIsBornIn` | -| U-223 | A step value that belongs to a different action: restoration fails informatively | Negative | `TestAStepOfAnotherActionIsRejected` | -| U-224 | A step value outside the enum: restoration fails rather than stalling | Negative | `TestAStepThisOperatorCannotResumeEndsTheOperation` | -| U-225 | The snapshot round-trips through `Snapshot` and `FromSnapshot` unchanged | Positive | — | -| U-226 | A deadline persisted and restored is the same absolute instant | Positive | — | -| U-227 | A deadline that passed while the operator was down: restores as expired | Boundary | `TestAStepThatOutlivedItsDeadlineFailsTheOperation` | -| U-228 | A step with no deadline: restores with none rather than with a zero instant | Boundary | — | -| U-229 | A terminal step: `IsTerminal` is true and no transition is attempted | Boundary | `TestTheLastStepFinishingEndsTheOperation` | -| U-230 | The outer phase machine is separate from the step machine | Positive | — | -| U-231 | Every state each graph declares appears in the step `Enum` marker | Boundary | `TestTheStepEnumCoversEveryDeclaredState` | -| U-232 | Every state each graph declares appears in the `status.step` CEL rule | Boundary | `TestTheCELRuleCoversEveryDeclaredState` | -| U-233 | The CEL rule names no value the graphs do not declare | Negative | `TestTheCELRuleCoversEveryDeclaredState` | -| U-234 | A stored step from another action: refused, naming the declared set | Negative | `TestAStepOfAnotherActionIsRejected` | -| U-235 | A restore that fails: the operation is `Failed` with the error, not requeued | Negative | `TestAStepThisOperatorCannotResumeEndsTheOperation` | -| U-365 | The provisioning graph's states cover its own `Enum` and CEL rule | Boundary | `TestTheNodeStepEnumAndRuleCoverTheProvisioningGraph` | -| U-366 | No step past the point of no return declares an abort edge | Negative | `TestNoStepPastThePointOfNoReturnIsAbortable` | -| U-367 | The steps an abort stops from are the ones that unwind cleanly | Positive | `TestTheStepsAnAbortStopsCleanly` | -| U-368 | Every drain step past the suspend owes the resume | Boundary | `TestTheDrainStepsPastTheSuspendUnwind` | -| U-369 | Every step carries a budget, so none is the step that cannot time out | Boundary | `TestEveryStepHasABudget` | -| U-370 | The remove graph validates before it suspends | Positive | `TestTheRemoveGraphValidatesBeforeItSuspends` | -| U-371 | The migrate graph splits the restart from the wait | Positive | `TestTheMigrateGraphSplitsTheRestartFromTheWait` | -| U-372 | The host maintenance graph is the six-step window | Positive | `TestTheHostMaintenanceGraphIsTheSixStepWindow` | -| U-373 | The four single-step actions share one line | Positive | `TestTheSingleStepActionsShareOneLine` | -| U-374 | Adoption is reachable from both provisioning gates | Boundary | `TestAdoptionIsReachableFromBothGates` | -| U-375 | `AwaitingSlot` may go straight to `Resolving` when a sibling claimed the worker | Boundary | `TestAwaitingSlotMayGoStraightToResolving` | +| # | Scenario | Type | Test | +|-------|------------------------------------------------------------------------------------|------------|----------------------------------------------------------| +| U-217 | Every declared graph builds, including the ones the action under test does not use | Positive | `TestEveryActionDeclaresAGraph` | +| U-218 | An action with no declared graph: refused rather than stalled | Negative | `TestAStepThatBelongsToNoActionEndsTheOperation` | +| U-219 | `Remove` transitioning to `Promoting`: rejected as an illegal transition | Negative | `TestAStepOfAnotherActionIsRejected` | +| U-220 | `Migrate` transitioning to `Removing`: rejected as an illegal transition | Negative | `TestAStepOfAnotherActionIsRejected` | +| U-221 | `HostMaintenance` transitioning to `Suspending`: rejected | Negative | `TestAStepOfAnotherActionIsRejected` | +| U-222 | An empty `status.step`: restores to the action's declared initial state | Boundary | `TestTheFirstPassArmsTheStepAMachineIsBornIn` | +| U-223 | A step value that belongs to a different action: restoration fails informatively | Negative | `TestAStepOfAnotherActionIsRejected` | +| U-224 | A step value outside the enum: restoration fails rather than stalling | Negative | `TestAStepThisOperatorCannotResumeEndsTheOperation` | +| U-225 | The snapshot round-trips through `Snapshot` and `FromSnapshot` unchanged | Positive | — | +| U-226 | A deadline persisted and restored is the same absolute instant | Positive | — | +| U-227 | A deadline that passed while the operator was down: restores as expired | Boundary | `TestAStepThatOutlivedItsDeadlineFailsTheOperation` | +| U-228 | A step with no deadline: restores with none rather than with a zero instant | Boundary | — | +| U-229 | A terminal step: `IsTerminal` is true and no transition is attempted | Boundary | `TestTheLastStepFinishingEndsTheOperation` | +| U-230 | The outer phase machine is separate from the step machine | Positive | — | +| U-231 | Every state each graph declares appears in the step `Enum` marker | Boundary | `TestTheStepEnumCoversEveryDeclaredState` | +| U-232 | Every state each graph declares appears in the `status.step` CEL rule | Boundary | `TestTheCELRuleCoversEveryDeclaredState` | +| U-233 | The CEL rule names no value the graphs do not declare | Negative | `TestTheCELRuleCoversEveryDeclaredState` | +| U-234 | A stored step from another action: refused, naming the declared set | Negative | `TestAStepOfAnotherActionIsRejected` | +| U-235 | A restore that fails: the operation is `Failed` with the error, not requeued | Negative | `TestAStepThisOperatorCannotResumeEndsTheOperation` | +| U-365 | The provisioning graph's states cover its own `Enum` and CEL rule | Boundary | `TestTheNodeStepEnumAndRuleCoverTheProvisioningGraph` | +| U-366 | No step past the point of no return declares an abort edge | Negative | `TestNoStepPastThePointOfNoReturnIsAbortable` | +| U-367 | The steps an abort stops from are the ones that unwind cleanly | Positive | `TestTheStepsAnAbortStopsCleanly` | +| U-368 | Every drain step past the suspend owes the resume | Boundary | `TestTheDrainStepsPastTheSuspendUnwind` | +| U-369 | Every step carries a budget, so none is the step that cannot time out | Boundary | `TestEveryStepHasABudget` | +| U-370 | The remove graph validates before it suspends | Positive | `TestTheRemoveGraphValidatesBeforeItSuspends` | +| U-371 | The migrate graph splits the restart from the wait | Positive | `TestTheMigrateGraphSplitsTheRestartFromTheWait` | +| U-372 | The host maintenance graph is the six-step window | Positive | `TestTheHostMaintenanceGraphIsTheSixStepWindow` | +| U-373 | The four single-step actions share one line | Positive | `TestTheSingleStepActionsShareOneLine` | +| U-374 | Adoption is reachable from both provisioning gates | Boundary | `TestAdoptionIsReachableFromBothGates` | +| U-375 | `AwaitingSlot` may go straight to `Resolving` when a sibling claimed the worker | Boundary | `TestAwaitingSlotMayGoStraightToResolving` | +| U-384 | A worker that is not Ready or is cordoned holds `Posting` at `AwaitingWorker` | Regression | `TestAWorkerThatWentAwayHoldsTheNodeRatherThanFailingIt` | +| U-385 | The held step carries its own budget rather than the one it was diverted from | Regression | `TestTheHeldNodeGetsAFreshBudget` | +| U-386 | The worker coming back starts the path again at `CheckingHost` | Positive | `TestAWorkerThatCameBackRestartsThePath` | +| U-387 | A held node emits `WorkerAway` rather than holding silently | Positive | `TestTheHeldNodeStaysAndSaysSo` | +| U-388 | A node that never claimed its worker takes no slot while the worker is away | Negative | `TestAnUnclaimedNodeTakesNoSlotWhileItsWorkerIsAway` | +| U-389 | A held node keeps its claim, so the node-add cap stays closed | Regression | `TestAHeldNodeKeepsItsClaim` | +| U-390 | A sibling posts no add while another worker is held at `AwaitingWorker` | Negative | `TestASiblingWaitsWhileAnotherWorkerIsHeld` | + +`U-384` to `U-390` are the reboot. The storage pool's MachineConfig is applied by +rebooting the machine, so the first node of a fresh cluster is cordoned, drained +and rebooted in the middle of being added (2026-09-20). `U-389` and `U-390` are +the pair that matter most: a held node keeps its claim, so the node-add cap stays +closed and no second worker is given a configuration change while the first is +still coming back. `U-388` is the other direction, and it is why only the two +steps that have claimed the worker divert — a node that never posted must not be +read by its sibling as an add that already happened. ### Entity: Pod Placement (design §13.1) diff --git a/operator/internal/controllers/node/awaitworker_test.go b/operator/internal/controllers/node/awaitworker_test.go new file mode 100644 index 000000000..4f31848f8 --- /dev/null +++ b/operator/internal/controllers/node/awaitworker_test.go @@ -0,0 +1,253 @@ +// What provisioning does while the worker it is adding is not there. +// +// A worker is cordoned, drained, rebooted and uncordoned whenever a +// MachineConfig reaches it, and the storage pool's own config is one: CPU +// isolation and huge pages are applied to a node by rebooting it, so the first +// node of a fresh cluster is rebooted in the middle of being added. The step +// that was running keeps its deadline through all of it and the control plane +// keeps being asked for a node the worker cannot produce. + +package node + +import ( + "context" + "strings" + "testing" + "time" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/client-go/tools/events" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/testsupport" +) + +// aWorkerIn is a Kubernetes node in the state the arguments describe. It is +// separate from aWorker because these cases turn on the cordon and the system +// UUID, which that fixture does not carry. +func aWorkerIn(ready bool, cordoned bool) *corev1.Node { + status := corev1.ConditionTrue + if !ready { + status = corev1.ConditionFalse + } + return &corev1.Node{ + ObjectMeta: metav1.ObjectMeta{Name: "worker-1"}, + Spec: corev1.NodeSpec{Unschedulable: cordoned}, + Status: corev1.NodeStatus{ + Addresses: []corev1.NodeAddress{{Type: corev1.NodeInternalIP, Address: workerKubernetesIP}}, + NodeInfo: corev1.NodeSystemInfo{SystemUUID: workerSystemUUID}, + Conditions: []corev1.NodeCondition{{Type: corev1.NodeReady, Status: status}}, + }, + } +} + +// aProvisioningNode is a node partway through the path, at the step given. +func aProvisioningNode(step nodeStep, deadline time.Duration) *simplyblockv1alpha2.StorageNode { + node := aResolvingNode() + node.Status.Step.State = string(step) + at := metav1.NewTime(time.Now().Add(deadline)) + node.Status.Step.Deadline = &at + return node +} + +// aProvisioner drives the provisioning path over one worker. +func aProvisioner(t *testing.T, worker *corev1.Node, node *simplyblockv1alpha2.StorageNode) ( + *StorageNodeReconciler, *simplyblockv1alpha2.StorageCluster, client.Client, *int, +) { + t.Helper() + scheme := testsupport.NewScheme(t, corev1.AddToScheme) + cluster := aClusterWithTasks() + apiClient := fake.NewClientBuilder().WithScheme(scheme). + WithObjects(worker, cluster, node). + WithStatusSubresource(&simplyblockv1alpha2.StorageNode{}). + Build() + + adds := 0 + return &StorageNodeReconciler{ + Client: apiClient, + Scheme: scheme, + Recorder: events.NewFakeRecorder(64), + API: countingBackend{adds: &adds}, + }, cluster, apiClient, &adds +} + +// stepOf reads back the step the provisioning pass recorded. +func stepOf(t *testing.T, apiClient client.Client, node *simplyblockv1alpha2.StorageNode) nodeStep { + t.Helper() + var read simplyblockv1alpha2.StorageNode + key := client.ObjectKeyFromObject(node) + if err := apiClient.Get(context.Background(), key, &read); err != nil { + t.Fatalf("reading the node back: %v", err) + } + return nodeStep(read.Status.Step.State) +} + +// TestAWorkerThatWentAwayHoldsTheNodeRatherThanFailingIt covers the reboot a +// MachineConfig performs in the middle of an add. +// +// Regression: 2026-09-20-a-rebooting-worker-burned-the-step-deadline — the +// storage pool's MachineConfig cordoned, drained and rebooted worker-5 while its +// node was being added. Nothing in the path represented that, so the step kept +// the deadline it entered with, the control plane kept being asked for a node, +// and the step would have failed the node for outliving a budget it spent +// waiting for a machine to come back. +func TestAWorkerThatWentAwayHoldsTheNodeRatherThanFailingIt(t *testing.T) { + for _, away := range []struct { + name string + worker *corev1.Node + }{ + {"not ready", aWorkerIn(false, false)}, + {"cordoned", aWorkerIn(true, true)}, + {"drained and rebooting", aWorkerIn(false, true)}, + } { + t.Run(away.name, func(t *testing.T) { + node := aProvisioningNode(stepPosting, time.Minute) + r, cluster, apiClient, adds := aProvisioner(t, away.worker, node) + + if _, err := r.provision(context.Background(), node, cluster); err != nil { + t.Fatalf("provision: %v", err) + } + if got := stepOf(t, apiClient, node); got != stepAwaitingWorker { + t.Errorf("the step is %q, want AwaitingWorker while the machine is not there", got) + } + if *adds != 0 { + t.Errorf("the control plane was asked for a node %d time(s) on a machine that is not there", *adds) + } + }) + } +} + +// The deadline of the step it left does not keep running: entering the new step +// sets that step's own budget, which is what makes the wait a wait rather than a +// countdown to failure. +func TestTheHeldNodeGetsAFreshBudget(t *testing.T) { + node := aProvisioningNode(stepPosting, 5*time.Second) + r, cluster, apiClient, _ := aProvisioner(t, aWorkerIn(false, true), node) + + if _, err := r.provision(context.Background(), node, cluster); err != nil { + t.Fatalf("provision: %v", err) + } + + var read simplyblockv1alpha2.StorageNode + if err := apiClient.Get(context.Background(), client.ObjectKeyFromObject(node), &read); err != nil { + t.Fatalf("reading the node back: %v", err) + } + if got := nodeStep(read.Status.Step.State); got != stepAwaitingWorker { + t.Fatalf("the step is %q, so this is not measuring the held step's budget", got) + } + if read.Status.Step.Deadline == nil { + t.Fatal("the held step carries no deadline, so it can never be given up on") + } + if left := time.Until(read.Status.Step.Deadline.Time); left <= 5*time.Second { + t.Errorf("the held step inherited %v, so it is still counting down the step it left", left) + } +} + +// A worker that came back starts the path again from its first step, which is +// where a backend node that was created before the reboot is adopted rather than +// added a second time. +func TestAWorkerThatCameBackRestartsThePath(t *testing.T) { + node := aProvisioningNode(stepAwaitingWorker, time.Hour) + r, cluster, apiClient, _ := aProvisioner(t, aWorkerIn(true, false), node) + + if _, err := r.provision(context.Background(), node, cluster); err != nil { + t.Fatalf("provision: %v", err) + } + if got := stepOf(t, apiClient, node); got != stepCheckingHost { + t.Errorf("the step is %q, want CheckingHost so the path is walked again", got) + } +} + +// While it is still away the node stays put, and says so: a step that holds +// without an event is indistinguishable from a reconcile that never ran. +func TestTheHeldNodeStaysAndSaysSo(t *testing.T) { + node := aProvisioningNode(stepAwaitingWorker, time.Hour) + r, cluster, apiClient, _ := aProvisioner(t, aWorkerIn(false, true), node) + + if _, err := r.provision(context.Background(), node, cluster); err != nil { + t.Fatalf("provision: %v", err) + } + if got := stepOf(t, apiClient, node); got != stepAwaitingWorker { + t.Errorf("the step is %q, want it to stay at AwaitingWorker", got) + } + + recorder, ok := r.Recorder.(*events.FakeRecorder) + if !ok { + t.Fatal("the recorder is not the fake one") + } + select { + case line := <-recorder.Events: + if !strings.Contains(line, WorkerAway) { + t.Errorf("the event was %q, want one naming %s", line, WorkerAway) + } + default: + t.Error("the node was held with no event, so nothing says why it is not progressing") + } +} + +// A node that has not claimed its worker is held where it is rather than +// diverted, because a claimed sibling on the same worker is read as an add that +// already happened. It still refuses the slot, which is what stops a second +// worker being given a configuration change while the first is rebooting. +func TestAnUnclaimedNodeTakesNoSlotWhileItsWorkerIsAway(t *testing.T) { + node := aProvisioningNode(stepAwaitingSlot, time.Hour) + r, cluster, apiClient, adds := aProvisioner(t, aWorkerIn(false, true), node) + + if _, err := r.provision(context.Background(), node, cluster); err != nil { + t.Fatalf("provision: %v", err) + } + if got := stepOf(t, apiClient, node); got != stepAwaitingSlot { + t.Errorf("the step is %q, want it to wait at AwaitingSlot", got) + } + if *adds != 0 { + t.Errorf("a node-add was posted %d time(s) for a machine that is not there", *adds) + } +} + +// The claim survives the wait, so the cluster's node-add cap stays closed and no +// sibling starts an add of its own while a worker reboots. +func TestAHeldNodeKeepsItsClaim(t *testing.T) { + held := aProvisioningNode(stepAwaitingWorker, time.Hour) + if !claimedWorker(held) { + t.Error("a node held for its worker released its claim, so a sibling may start an add beside it") + } +} + +// A sibling of a held node stays at AwaitingSlot: the cap counts the held node, +// so the sibling's worker is not handed a configuration change alongside the +// reboot already running. +func TestASiblingWaitsWhileAnotherWorkerIsHeld(t *testing.T) { + held := aProvisioningNode(stepAwaitingWorker, time.Hour) + held.Name = "a-cluster-worker-9-0" + held.Spec.WorkerNode = "worker-9" + + waiting := aProvisioningNode(stepAwaitingSlot, time.Hour) + scheme := testsupport.NewScheme(t, corev1.AddToScheme) + cluster := aClusterWithTasks() + apiClient := fake.NewClientBuilder().WithScheme(scheme). + WithObjects(aWorkerIn(true, false), &corev1.Node{ + ObjectMeta: metav1.ObjectMeta{Name: "worker-9"}, + }, cluster, held, waiting). + WithStatusSubresource(&simplyblockv1alpha2.StorageNode{}). + Build() + adds := 0 + r := &StorageNodeReconciler{ + Client: apiClient, + Scheme: scheme, + Recorder: events.NewFakeRecorder(64), + API: countingBackend{adds: &adds}, + } + + if _, err := r.provision(context.Background(), waiting, cluster); err != nil { + t.Fatalf("provision: %v", err) + } + if got := stepOf(t, apiClient, waiting); got != stepAwaitingSlot { + t.Errorf("the sibling moved to %q while another worker was held", got) + } + if adds != 0 { + t.Errorf("the sibling posted %d add(s) while another worker was rebooting", adds) + } +} diff --git a/operator/internal/controllers/node/controlplane.go b/operator/internal/controllers/node/controlplane.go index f14a61b60..f14f92404 100644 --- a/operator/internal/controllers/node/controlplane.go +++ b/operator/internal/controllers/node/controlplane.go @@ -37,9 +37,17 @@ import ( // condition in the package is a predicate over it and the fields it needs are the // ones named here. type NodeReading struct { - UUID string `json:"id"` - Status string `json:"status"` - ManagementIP string `json:"mgmt_ip"` + UUID string `json:"id"` + Status string `json:"status"` + ManagementIP string `json:"mgmt_ip"` + + // SystemUUID is the host's firmware identity, which Kubernetes reports for + // the same machine as Node.status.nodeInfo.systemUUID. It is what says which + // worker a reading is on: mgmt_ip is an address on the storage plane, and a + // deployment whose storage traffic has its own network reports one the + // Kubernetes node object never carries. + SystemUUID string `json:"system_uuid"` + Health bool `json:"health_check"` Hostname string `json:"hostname"` Uptime string `json:"uptime"` diff --git a/operator/internal/controllers/node/events.go b/operator/internal/controllers/node/events.go index ea5e79836..91f99056f 100644 --- a/operator/internal/controllers/node/events.go +++ b/operator/internal/controllers/node/events.go @@ -29,6 +29,14 @@ const ( HostUnreachable = "HostUnreachable" AwaitingSlot = "AwaitingSlot" + // WorkerAway is the machine this node is being added to being not Ready or + // cordoned, which is every MachineConfig reboot and every drain. + WorkerAway = "WorkerAway" + + // WorkerReturned is that machine coming back, and the path being walked + // again from its first step. + WorkerReturned = "WorkerReturned" + // NodeAddGaveUp is the add this node was waiting on leaving the control // plane's task window without having produced a node. NodeAddGaveUp = "NodeAddGaveUp" diff --git a/operator/internal/controllers/node/graphs.go b/operator/internal/controllers/node/graphs.go index a3c54ce85..9cee809f4 100644 --- a/operator/internal/controllers/node/graphs.go +++ b/operator/internal/controllers/node/graphs.go @@ -63,6 +63,8 @@ const ( stepPosting = simplyblockv1alpha2.StorageNodeStepPosting stepResolving = simplyblockv1alpha2.StorageNodeStepResolving stepAdopting = simplyblockv1alpha2.StorageNodeStepAdopting + + stepAwaitingWorker = simplyblockv1alpha2.StorageNodeStepAwaitingWorker ) // How long each operation step may take before it is reported as stuck. @@ -113,6 +115,12 @@ const ( postingDeadline = 10 * time.Minute resolvingDeadline = 45 * time.Minute adoptingDeadline = 10 * time.Minute + + // A worker comes back from a MachineConfig reboot in minutes: the cordon, + // the drain, the reboot itself and the uncordon took eleven of them on the + // cluster this was written for. The budget is the one that says a machine is + // not coming back rather than the one that says it is slow, so it is an hour. + awaitingWorkerDeadline = time.Hour ) // deadline is the entry hook every state here carries: it sets the step's budget @@ -277,11 +285,18 @@ func provisioningGraph() statemachine.Config[nodeStep] { OnEnter: deadline[nodeStep](awaitingSlotDeadline), }, stepPosting: { - To: []nodeStep{stepResolving}, + To: []nodeStep{stepResolving, stepAwaitingWorker}, OnEnter: deadline[nodeStep](postingDeadline), }, - stepResolving: {OnEnter: deadline[nodeStep](resolvingDeadline)}, - stepAdopting: {OnEnter: deadline[nodeStep](adoptingDeadline)}, + stepResolving: { + To: []nodeStep{stepAwaitingWorker}, + OnEnter: deadline[nodeStep](resolvingDeadline), + }, + stepAwaitingWorker: { + To: []nodeStep{stepCheckingHost}, + OnEnter: deadline[nodeStep](awaitingWorkerDeadline), + }, + stepAdopting: {OnEnter: deadline[nodeStep](adoptingDeadline)}, }, } } diff --git a/operator/internal/controllers/node/graphs_test.go b/operator/internal/controllers/node/graphs_test.go index 06f4da17e..b8c8bfebb 100644 --- a/operator/internal/controllers/node/graphs_test.go +++ b/operator/internal/controllers/node/graphs_test.go @@ -70,7 +70,8 @@ func TestTheCELRuleCoversEveryDeclaredState(t *testing.T) { func TestTheNodeStepEnumAndRuleCoverTheProvisioningGraph(t *testing.T) { declared := statemachine.DeclaredStates(provisioningGraph()) want := []string{ - "Adopting", "AwaitingSlot", "CheckingConfig", "CheckingHost", "Posting", "Resolving", + "Adopting", "AwaitingSlot", "AwaitingWorker", "CheckingConfig", "CheckingHost", + "Posting", "Resolving", } if diff := cmp.Diff(want, declared); diff != "" { t.Errorf("the provisioning graph and the Enum marker disagree (-marker +graph):\n%s", diff) @@ -107,7 +108,8 @@ const opsStepCELRule = "!has(self.state) || self.state in " + "'ShuttingDown','Releasing','AwaitingHost','Restarting','Cleanup']" const nodeStepCELRule = "!has(self.state) || self.state in " + - "['CheckingHost','CheckingConfig','AwaitingSlot','Posting','Resolving','Adopting']" + "['CheckingHost','CheckingConfig','AwaitingSlot','Posting','Resolving','Adopting'," + + "'AwaitingWorker']" // Every action the API accepts needs a graph, or an operation of that action fails // at its first pass with ErrUnknownAction rather than doing anything. diff --git a/operator/internal/controllers/node/resolve_identity_test.go b/operator/internal/controllers/node/resolve_identity_test.go new file mode 100644 index 000000000..484dda97a --- /dev/null +++ b/operator/internal/controllers/node/resolve_identity_test.go @@ -0,0 +1,160 @@ +// Which backend node belongs to which worker. +// +// Resolving has to recognize the node the add just produced, and the only thing +// that makes that possible is an identity both sides report. The control plane +// names a node by an address on the storage plane, which is the network its +// nodes talk to each other over and is not required to be the one Kubernetes +// runs on. + +package node + +import ( + "context" + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/client-go/tools/events" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/testsupport" + "github.com/simplyblock/simplyblock-operator/internal/utils" +) + +// The two networks of the cluster this was found on, and the host identity both +// sides carry: Kubernetes reports it as Node.status.nodeInfo.systemUUID and the +// control plane as the node's system_uuid, both read from the same firmware. +const ( + workerKubernetesIP = "10.0.0.15" + workerStorageIP = "192.168.10.15" + workerSystemUUID = "7aa0847c-6308-11ef-b0d9-02310e2a5e00" +) + +// backendNodes answers with the readings it was built from. +type backendNodes struct { + ControlPlane + readings []NodeReading +} + +func (b backendNodes) StorageNodes(context.Context, string) ([]NodeReading, error) { + return b.readings, nil +} + +// adds counts the node_add calls a case provoked, which is what separates +// "the step held" from "the step held after asking for a node anyway". +type countingBackend struct { + ControlPlane + adds *int +} + +func (countingBackend) StorageNodes(context.Context, string) ([]NodeReading, error) { + return nil, nil +} + +func (c countingBackend) AddNode(context.Context, string, utils.StorageNodeSetAddParams) error { + *c.adds++ + return nil +} + +// aMatcher builds a reconciler over one worker whose Kubernetes address and +// storage-plane address differ, which is the ordinary shape of a deployment +// whose storage traffic is on its own network. +func aMatcher(t *testing.T, readings []NodeReading, systemUUID string) ( + *StorageNodeReconciler, *simplyblockv1alpha2.StorageNode, *simplyblockv1alpha2.StorageCluster, +) { + t.Helper() + scheme := testsupport.NewScheme(t, corev1.AddToScheme) + worker := &corev1.Node{ + ObjectMeta: metav1.ObjectMeta{Name: "worker-1"}, + Status: corev1.NodeStatus{ + Addresses: []corev1.NodeAddress{{Type: corev1.NodeInternalIP, Address: workerKubernetesIP}}, + NodeInfo: corev1.NodeSystemInfo{SystemUUID: systemUUID}, + }, + } + node := aResolvingNode() + cluster := aClusterWithTasks() + + apiClient := fake.NewClientBuilder().WithScheme(scheme). + WithObjects(worker, cluster, node). + WithStatusSubresource(&simplyblockv1alpha2.StorageNode{}). + Build() + + return &StorageNodeReconciler{ + Client: apiClient, + Scheme: scheme, + Recorder: events.NewFakeRecorder(64), + API: backendNodes{readings: readings}, + }, node, cluster +} + +// TestANodeIsFoundWhenTheStoragePlaneIsItsOwnNetwork covers the node the add +// produced on a cluster whose storage network is not the Kubernetes one. +// +// Regression: 2026-09-20-backend-node-matched-by-the-kubernetes-address — the +// match kept a reading only where its mgmt_ip was the worker's Kubernetes +// InternalIP. The control plane reports the address the node holds on the +// storage plane, which on a two-network cluster is a different subnet +// (192.168.10.15 against 10.0.0.15), so every reading was dropped. The node was +// online and healthy and the operator reported that the add had produced +// nothing, re-posted once a second, and would have failed the node on its +// deadline with a backend node running underneath it. +func TestANodeIsFoundWhenTheStoragePlaneIsItsOwnNetwork(t *testing.T) { + r, node, cluster := aMatcher(t, []NodeReading{{ + UUID: "4a304439-2909-4199-ad2f-b8624d66a13d", + Status: nodeStatusOnline, + ManagementIP: workerStorageIP, + SystemUUID: workerSystemUUID, + Hostname: "worker-1_4422", + RPCPort: 4422, + }}, workerSystemUUID) + + reading, found, err := r.matchBackendNode(context.Background(), node, cluster) + if err != nil { + t.Fatalf("matchBackendNode: %v", err) + } + if !found { + t.Fatal("the node the add produced was not recognized, so the add reads as having produced nothing") + } + if reading.UUID != "4a304439-2909-4199-ad2f-b8624d66a13d" { + t.Errorf("matched %q", reading.UUID) + } +} + +// A node of another worker is not this worker's, whatever network it is on. +func TestANodeOfAnotherWorkerIsNotMatched(t *testing.T) { + r, node, cluster := aMatcher(t, []NodeReading{{ + UUID: "someone-else", + Status: nodeStatusOnline, + ManagementIP: "192.168.10.99", + SystemUUID: "ef7f45c7-ac2b-9a4f-ee58-08bfb8a41207", + RPCPort: 4422, + }}, workerSystemUUID) + + _, found, err := r.matchBackendNode(context.Background(), node, cluster) + if err != nil { + t.Fatalf("matchBackendNode: %v", err) + } + if found { + t.Error("another worker's node was taken for this one") + } +} + +// A control plane that reports no system UUID is still matched on the address, +// which is what every single-network deployment has always been matched on. +func TestTheAddressStillMatchesWhereNoSystemUUIDIsReported(t *testing.T) { + r, node, cluster := aMatcher(t, []NodeReading{{ + UUID: "older-control-plane", + Status: nodeStatusOnline, + ManagementIP: workerKubernetesIP, + RPCPort: 4422, + }}, workerSystemUUID) + + _, found, err := r.matchBackendNode(context.Background(), node, cluster) + if err != nil { + t.Fatalf("matchBackendNode: %v", err) + } + if !found { + t.Error("a reading with no system UUID stopped matching on the address it always did") + } +} diff --git a/operator/internal/controllers/node/storagenode_controller.go b/operator/internal/controllers/node/storagenode_controller.go index c74b6e9d6..3969d3a67 100644 --- a/operator/internal/controllers/node/storagenode_controller.go +++ b/operator/internal/controllers/node/storagenode_controller.go @@ -27,6 +27,7 @@ import ( "fmt" "slices" "strconv" + "strings" "time" corev1 "k8s.io/api/core/v1" @@ -317,6 +318,38 @@ func (r *StorageNodeReconciler) provision( } current := machine.CurrentState() + + // A worker that is not there is checked before the deadline, because a step + // waiting on it cannot make progress and spending its budget on a machine + // that is rebooting would fail the node for the reboot. + // + // Only the two steps that have claimed the worker divert. The claim is what + // AwaitingWorker has to carry: it counts as claimed (claimedWorker), so the + // cluster's node-add cap keeps holding while the machine is away and no + // second worker is handed a configuration change on top of a reboot already + // running. A step that had not claimed anything cannot be held the same way, + // because a claimed sibling on the same worker is read as an add that already + // happened and a reason to skip Posting, and a node diverted out of + // AwaitingSlot never posted. Those steps wait where they are instead: AwaitingSlot declines to + // take a slot while the worker is away, and the host check simply does not + // answer. + if current == stepPosting || current == stepResolving { + away, reason, err := r.workerAway(ctx, node.Spec.WorkerNode) + if err != nil { + return ctrl.Result{RequeueAfter: nodeRetry}, err + } + if away { + r.emit(node, corev1.EventTypeWarning, WorkerAway, fmt.Sprintf( + "Worker %s %s, so %s is held until it is back", node.Spec.WorkerNode, reason, current)) + if err := machine.TransitionTo(ctx, stepAwaitingWorker); err != nil { + return ctrl.Result{}, fmt.Errorf("enter step %s: %w", stepAwaitingWorker, err) + } + snapshot := statemachine.ToKube(machine.Snapshot()) + return ctrl.Result{RequeueAfter: nodeRetry}, + r.recordStep(ctx, node, stepAwaitingWorker, snapshot.Deadline) + } + } + if machine.TimeoutReached() { return ctrl.Result{}, r.fail(ctx, node, fmt.Sprintf("step %s outlived its deadline", current)) @@ -338,10 +371,15 @@ func (r *StorageNodeReconciler) provision( fmt.Sprintf("waiting on %s", current)) } - if machine.IsTerminal() { + if next == current { // Resolving and Adopting both end with a UUID on the object, which the - // step that reached them has already written. The next pass is steady - // state. + // step that reached them has already written, and both report themselves + // as their own successor. The next pass is steady state. + // + // The test is the step rather than machine.IsTerminal, which answers + // whether the graph gives the state an exit: Resolving has one now, to + // the step that holds while the worker is away, and a finished Resolving + // is still the end of the path. return ctrl.Result{RequeueAfter: nodeAdvance}, nil } @@ -377,6 +415,8 @@ func (r *StorageNodeReconciler) performNodeStep( return stepResolving, true, r.postNode(ctx, node, cluster) case stepResolving: return r.resolve(ctx, node, cluster) + case stepAwaitingWorker: + return r.awaitWorker(ctx, node) case stepAdopting: done, err := r.resolveUUID(ctx, node, cluster) return stepAdopting, done, err @@ -385,6 +425,60 @@ func (r *StorageNodeReconciler) performNodeStep( } } +// awaitWorker holds while the worker is away and starts the path again when it +// is back. +// +// It returns to CheckingHost rather than to the step it was diverted from, +// because what a reboot interrupts is not resumable in the middle: a Posting +// resumed after the fact would ask for a second node, and the add it was waiting +// on may well have landed. CheckingHost is where a backend node that already +// exists is found, and its Adopting edge is what takes the node over instead of +// adding it twice. +func (r *StorageNodeReconciler) awaitWorker( + ctx context.Context, node *simplyblockv1alpha2.StorageNode, +) (nodeStep, bool, error) { + away, reason, err := r.workerAway(ctx, node.Spec.WorkerNode) + if err != nil { + return stepAwaitingWorker, false, err + } + if away { + return stepAwaitingWorker, false, blockedf(WorkerAway, + "worker %s %s", node.Spec.WorkerNode, reason) + } + r.emit(node, corev1.EventTypeNormal, WorkerReturned, fmt.Sprintf( + "Worker %s is back, and provisioning starts again from the host check", + node.Spec.WorkerNode)) + return stepCheckingHost, true, nil +} + +// workerAway reports whether the machine is unavailable, and says which of the +// two conditions it is so the event names the cause rather than the outcome. +// +// The cordon counts as well as the readiness, because a drain begins with one: +// catching it there is what stops an add being asked for against a machine that +// is about to go down. +// +// A worker the API server does not have is not reported away. It is a different +// condition with its own handling, and answering it here would make every step +// of this path depend on a Node object that a node being torn down no longer +// has. +func (r *StorageNodeReconciler) workerAway( + ctx context.Context, worker string, +) (bool, string, error) { + var object corev1.Node + if err := r.Get(ctx, types.NamespacedName{Name: worker}, &object); err != nil { + return false, "", client.IgnoreNotFound(err) + } + switch { + case !workerReady(&object): + return true, "is not Ready", nil + case object.Spec.Unschedulable: + return true, "is cordoned", nil + default: + return false, "", nil + } +} + // checkHost holds until the worker's storage-node API answers, and diverts to // adoption when this deployment is being taken over wholesale or when a backend // node is already at the worker's address. @@ -463,6 +557,19 @@ func (r *StorageNodeReconciler) awaitSlot( node *simplyblockv1alpha2.StorageNode, cluster *simplyblockv1alpha2.StorageCluster, ) (nodeStep, bool, error) { + // A slot taken for a machine that is not there is an add posted against a + // worker that cannot answer it. The wait costs nothing, because the slot is + // still free when the machine comes back, and taking it would put a second + // worker's configuration change alongside a reboot already under way. + away, reason, err := r.workerAway(ctx, node.Spec.WorkerNode) + if err != nil { + return stepAwaitingSlot, false, err + } + if away { + return stepAwaitingSlot, false, blockedf(WorkerAway, + "worker %s %s, so no node-add slot is taken for it", node.Spec.WorkerNode, reason) + } + siblings, err := r.clusterNodeObjects(ctx, node) if err != nil { return stepAwaitingSlot, false, err @@ -554,7 +661,11 @@ func (r *StorageNodeReconciler) awaitSlot( // transition into Posting or anything past it. func claimedWorker(node *simplyblockv1alpha2.StorageNode) bool { switch nodeStep(node.Status.Step.State) { - case stepPosting, stepResolving, stepAdopting: + case stepPosting, stepResolving, stepAdopting, stepAwaitingWorker: + // AwaitingWorker is only reachable from Posting and Resolving, so a node + // in it has claimed its worker and its add is still outstanding. Holding + // the claim across the wait is what keeps the cluster's node-add cap + // closed while a machine reboots. return true default: return node.Status.UUID != "" @@ -684,8 +795,8 @@ func (r *StorageNodeReconciler) matchBackendNode( node *simplyblockv1alpha2.StorageNode, cluster *simplyblockv1alpha2.StorageCluster, ) (NodeReading, bool, error) { - address, err := r.workerAddress(ctx, node.Spec.WorkerNode) - if err != nil || address == "" { + worker, err := r.workerIdentity(ctx, node.Spec.WorkerNode) + if err != nil || (worker.address == "" && worker.systemUUID == "") { return NodeReading{}, false, err } @@ -696,7 +807,7 @@ func (r *StorageNodeReconciler) matchBackendNode( var onWorker []NodeReading for _, reading := range readings { - if reading.ManagementIP == address && reading.UUID != "" { + if reading.UUID != "" && worker.owns(reading) { onWorker = append(onWorker, reading) } } @@ -1118,21 +1229,52 @@ func (r *StorageNodeReconciler) clusterNodeObjects( return nodes.Items, nil } -// workerAddress is the worker's internal IP, which is what the control plane -// reports as a backend node's management address. -func (r *StorageNodeReconciler) workerAddress( +// workerIdentity is what the two sides of a node both know about one machine. +// +// The firmware UUID is the identity: Kubernetes reads it from the host and +// reports it as Node.status.nodeInfo.systemUUID, and the control plane reads it +// from the same place and reports it as a node's system_uuid. Neither invents +// it and no network carries it. +// +// The address is kept for the control plane that reports no UUID. It was the +// only match there was, and it works wherever the storage plane and Kubernetes +// share a network, which is every deployment that has been matched this way so +// far. +type workerIdentity struct { + address string + systemUUID string +} + +// owns reports whether a reading is a node of this worker. +// +// The UUID decides it whenever both sides have one, because it is the answer +// that does not depend on which network the reading's address is on. Falling +// back to the address where the reading carries no UUID keeps the older control +// plane working; falling back where it carries a different one would undo the +// whole point, so a UUID that disagrees is a node of another machine. +func (w workerIdentity) owns(reading NodeReading) bool { + if reading.SystemUUID != "" && w.systemUUID != "" { + return strings.EqualFold(reading.SystemUUID, w.systemUUID) + } + return w.address != "" && reading.ManagementIP == w.address +} + +// workerIdentity reads both from the worker's Node object. +func (r *StorageNodeReconciler) workerIdentity( ctx context.Context, worker string, -) (string, error) { +) (workerIdentity, error) { var object corev1.Node if err := r.Get(ctx, types.NamespacedName{Name: worker}, &object); err != nil { - return "", client.IgnoreNotFound(err) + return workerIdentity{}, client.IgnoreNotFound(err) } + identity := workerIdentity{systemUUID: object.Status.NodeInfo.SystemUUID} for _, address := range object.Status.Addresses { if address.Type == corev1.NodeInternalIP { - return address.Address, nil + identity.address = address.Address + break } } - return "", nil + return identity, nil } // hostsFoundationDB reports whether a worker runs a FoundationDB pod, which is diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagenodes.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagenodes.yaml index 53e5f76a6..73682b427 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagenodes.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagenodes.yaml @@ -869,7 +869,7 @@ spec: type: object x-kubernetes-validations: - message: unknown step - rule: '!has(self.state) || self.state in [''CheckingHost'',''CheckingConfig'',''AwaitingSlot'',''Posting'',''Resolving'',''Adopting'']' + rule: '!has(self.state) || self.state in [''CheckingHost'',''CheckingConfig'',''AwaitingSlot'',''Posting'',''Resolving'',''Adopting'',''AwaitingWorker'']' uptime: description: Uptime is the node uptime as the control plane reports it. type: string From 01a57c8c7bb3da4a09e29dabbd9887e6fcf07a2a Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Sun, 20 Sep 2026 18:09:50 +0200 Subject: [PATCH 098/206] fix(operator): the step that asks for the add again can reach the one that posts enter step Posting: statemachine: illegal transition Resolving -> Posting Resolving returns Posting when the add it was waiting on left the control plane's task window without producing a node, which is the whole point of that branch: an add that is over without a node is one to ask for again. The graph gave Resolving no exit, so the transition was refused, the reconcile errored before re-posting, and controller-runtime backed off and ran it again. What that looked like on a cluster is a node emitting NodeAddGaveUp once a second for its entire deadline while exactly one add was ever sent. The retry the step exists for never happened, and the event said it had. TestResolvingAsksAgainWhenTheAddIsOver covers the same decision one level down and passed throughout. It reads the step resolve returns and never performs the transition, and the transition is the only place the graph is consulted. The new case goes through provision, so the edge is exercised rather than the intent. Co-Authored-By: Claude Opus 5 (1M context) --- .../controllers/node/awaitworker_test.go | 31 +++++++++++++++++++ operator/internal/controllers/node/graphs.go | 7 ++++- 2 files changed, 37 insertions(+), 1 deletion(-) diff --git a/operator/internal/controllers/node/awaitworker_test.go b/operator/internal/controllers/node/awaitworker_test.go index 4f31848f8..334ae2391 100644 --- a/operator/internal/controllers/node/awaitworker_test.go +++ b/operator/internal/controllers/node/awaitworker_test.go @@ -251,3 +251,34 @@ func TestASiblingWaitsWhileAnotherWorkerIsHeld(t *testing.T) { t.Errorf("the sibling posted %d add(s) while another worker was rebooting", adds) } } + +// TestAnAddThatGaveUpIsActuallyReposted covers the retry the resolve step asks +// for, through the transition that performs it. +// +// Regression: 2026-09-20-resolving-could-not-reach-posting — resolve returns +// Posting when the add it was waiting on left the task window without producing +// a node, and the graph gave Resolving no exit, so the transition was refused: +// +// enter step Posting: statemachine: illegal transition Resolving -> Posting +// +// The reconcile errored before re-posting and backed off, so the node emitted +// NodeAddGaveUp once a second for its whole deadline while exactly one add was +// ever sent. The step that exists to ask again could not. +// +// TestResolvingAsksAgainWhenTheAddIsOver covers the same decision one level +// down, and passed throughout: it reads the step resolve returns and never +// performs the transition, which is the only place the graph is consulted. +func TestAnAddThatGaveUpIsActuallyReposted(t *testing.T) { + node := aProvisioningNode(stepResolving, time.Hour) + r, cluster, apiClient, _ := aProvisioner(t, aWorkerIn(true, false), node) + cluster.Status.Tasks = []simplyblockv1alpha2.ClusterTask{ + {ID: "task-1", Type: "node_add", Status: "done"}, + } + + if _, err := r.provision(context.Background(), node, cluster); err != nil { + t.Fatalf("provision: %v", err) + } + if got := stepOf(t, apiClient, node); got != stepPosting { + t.Errorf("the step is %q, want Posting so the add is actually asked for again", got) + } +} diff --git a/operator/internal/controllers/node/graphs.go b/operator/internal/controllers/node/graphs.go index 9cee809f4..4618ebeca 100644 --- a/operator/internal/controllers/node/graphs.go +++ b/operator/internal/controllers/node/graphs.go @@ -289,7 +289,12 @@ func provisioningGraph() statemachine.Config[nodeStep] { OnEnter: deadline[nodeStep](postingDeadline), }, stepResolving: { - To: []nodeStep{stepAwaitingWorker}, + // Posting is an exit because an add that left the control + // plane's task window without producing a node is one to ask for + // again, and asking is re-entering the step that posts. Without + // the edge the transition is refused and the retry the step + // exists for never happens. + To: []nodeStep{stepPosting, stepAwaitingWorker}, OnEnter: deadline[nodeStep](resolvingDeadline), }, stepAwaitingWorker: { From 2bf5a0376c48904fa71fee0c3c84636972b880bc Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Sun, 20 Sep 2026 20:02:54 +0200 Subject: [PATCH 099/206] feat(atlas-lib): a machine's own sysfs, and the reading that tells a pod link from a bridge MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two devices sysfs describes identically: the bridge a node's address lives on and a pod's link into the host. On OpenShift with OVN-Kubernetes both resolve under devices/virtual, export no bridge directory and carry no driver link, so nothing about what they are separates them. Checked on the machine: br-ex ifindex=10 iflink=10 ip -d link: openvswitch 3070a31ea98cce0 ifindex=14 iflink=2 ip -d link: veth iflink is what separates them, and every interface carries it: one that points at itself stands alone, and one naming another index is derived from or paired with that device. Interface.Peered is that reading. It does not say which of the two. A VLAN points at the parent it tags and reads the same way — enp8s0.4000 on the same host, iflink=2 against ifindex=9 — so a caller keeping pod links out wants this together with the kind rather than instead of it: a VLAN is declared (DEVTYPE=vlan, a lower link) and a veth's peer is in another namespace with nothing declaring it at all. The field's documentation says so, because the first rule written on it refused the VLAN. The transcript is the other half. testdata/hosts/okd-worker.json.gz is every path under the trees the readers walk, taken off worker-5 of the lab fleet: 3,244 directories, 16,335 files, 1,635 symlinks, 128 KiB compressed, plus the two procfs files the readers need. It is one file rather than a checked-in tree because sysfs is mostly empty files and symlinks, which no review and no checkout survives intact; hostFixture materializes it once per process. It earned itself immediately: the VLAN above is why the first rule was wrong, and no synthetic fixture had one. TestTheCapturedHostAnswersEveryReader runs ReadCPU, ReadHugePages and ReadInterfaces over it, so a reader added later gets a real machine for free. hack/inventory/capture-host.py takes the next host. Co-Authored-By: Claude Opus 5 (1M context) --- atlas-lib/inventory/hostfixture_test.go | 142 ++++++++++++++++++ atlas-lib/inventory/netiface.go | 16 ++ atlas-lib/inventory/netiface_test.go | 98 ++++++++++++ .../testdata/hosts/okd-worker.json.gz | Bin 0 -> 129487 bytes hack/inventory/capture-host.py | 100 ++++++++++++ 5 files changed, 356 insertions(+) create mode 100644 atlas-lib/inventory/hostfixture_test.go create mode 100644 atlas-lib/inventory/testdata/hosts/okd-worker.json.gz create mode 100755 hack/inventory/capture-host.py diff --git a/atlas-lib/inventory/hostfixture_test.go b/atlas-lib/inventory/hostfixture_test.go new file mode 100644 index 000000000..3e976ca35 --- /dev/null +++ b/atlas-lib/inventory/hostfixture_test.go @@ -0,0 +1,142 @@ +// Whole-host transcripts: a real machine's sysfs, captured and replayed. +// +// The fixtures beside this one are written by hand, one attribute at a time, +// which is what a test of a single reading wants. This is the other kind: every +// path under the trees the readers walk, taken off a machine a storage node runs +// on, so a reading can be checked against what a host actually exports rather +// than against what the fixture's author remembered to include. +// +// It is one file per host rather than a directory of thousands, because sysfs is +// mostly empty files and symlinks and a checked-in tree of those is unreviewable +// and does not survive a checkout on a filesystem without symlinks. The capture +// is a JSON document of three maps -- directories, file contents, symlink +// targets -- materialized into a temporary tree per test. +// +// Capturing another host: run the walker in hack/inventory/capture-host.py on it +// and drop the result beside this file. The trees are small; the OKD worker below +// is 16k files and 128 KiB compressed. + +package inventory + +import ( + "compress/gzip" + "encoding/json" + "os" + "path/filepath" + "sync" + "testing" +) + +// hostTranscript is the captured shape of one machine's sysfs. +type hostTranscript struct { + Dirs []string `json:"dirs"` + Files map[string]string `json:"files"` + Links map[string]string `json:"links"` +} + +// materialized caches one transcript per process. A host is sixteen thousand +// files and writing them takes seconds, which is worth paying once and not once +// per test: nothing reading a transcript writes to it. +var materialized sync.Map + +// hostFixture materializes a captured host under a temporary root and returns +// it, ready to be handed to a Config as SysfsRoot. The tree is written once per +// process and shared by every test that asks for that host. +// +// A symlink whose target cannot be created is written anyway: sysfs is full of +// links that point outside the captured subtree, and a reader that follows one +// gets the same "not found" it would get on a host where the target is a device +// that has since gone. Failing the fixture instead would make the capture depend +// on how much of the machine was walked. +func hostFixture(t *testing.T, name string) string { + t.Helper() + + if root, ok := materialized.Load(name); ok { + return root.(string) + } + + handle, err := os.Open(filepath.Join("testdata", "hosts", name+".json.gz")) + if err != nil { + t.Fatalf("open the host transcript: %v", err) + } + defer func() { _ = handle.Close() }() + + reader, err := gzip.NewReader(handle) + if err != nil { + t.Fatalf("read the host transcript: %v", err) + } + defer func() { _ = reader.Close() }() + + var host hostTranscript + if err := json.NewDecoder(reader).Decode(&host); err != nil { + t.Fatalf("decode the host transcript: %v", err) + } + + // Not t.TempDir: the tree outlives the test that first asked for it, and Go + // removes the process's temporary directories when it exits. + root, err := os.MkdirTemp("", "inventory-host-") + if err != nil { + t.Fatalf("create the host root: %v", err) + } + for _, dir := range host.Dirs { + if err := os.MkdirAll(filepath.Join(root, dir), 0o755); err != nil { + t.Fatalf("create %s: %v", dir, err) + } + } + for path, content := range host.Files { + full := filepath.Join(root, path) + if err := os.MkdirAll(filepath.Dir(full), 0o755); err != nil { + t.Fatalf("create the parent of %s: %v", path, err) + } + if err := os.WriteFile(full, []byte(content), 0o644); err != nil { + t.Fatalf("write %s: %v", path, err) + } + } + for path, target := range host.Links { + full := filepath.Join(root, path) + if err := os.MkdirAll(filepath.Dir(full), 0o755); err != nil { + t.Fatalf("create the parent of %s: %v", path, err) + } + _ = os.Symlink(target, full) + } + materialized.Store(name, root) + return root +} + +// TestTheCapturedHostAnswersEveryReader is what makes the transcript worth +// keeping: one machine, read by all of them. +// +// A reader added later gets a real host to run against for free, and a reader +// that starts disagreeing with a machine says so here rather than on a cluster. +func TestTheCapturedHostAnswersEveryReader(t *testing.T) { + root := hostFixture(t, "okd-worker") + host := syntheticHost(root) + + cpu, err := ReadCPU(host) + if err != nil { + t.Fatalf("ReadCPU: %v", err) + } + if cpu.OnlineCount == 0 || cpu.PhysicalCores == 0 { + t.Errorf("the captured host reads as having no CPUs: online=%d cores=%d", + cpu.OnlineCount, cpu.PhysicalCores) + } + if cpu.Sockets == 0 { + t.Error("the captured host reads as having no sockets") + } + + pages, err := ReadHugePages(host) + if err != nil { + t.Fatalf("ReadHugePages: %v", err) + } + if len(pages.Pools) == 0 { + t.Error("the captured host reads as having no huge-page pools") + } + + ifaces, err := ReadInterfaces(host) + if err != nil { + t.Fatalf("ReadInterfaces: %v", err) + } + if len(ifaces) == 0 { + t.Error("the captured host reads as having no interfaces") + } +} diff --git a/atlas-lib/inventory/netiface.go b/atlas-lib/inventory/netiface.go index 70060bb24..c2cc03d0e 100644 --- a/atlas-lib/inventory/netiface.go +++ b/atlas-lib/inventory/netiface.go @@ -74,6 +74,19 @@ type Interface struct { // a veth, a bond, a VLAN, or loopback. Virtual bool + // Peered reports whether the interface names another device as its link, + // which is read from iflink: an interface that points at itself stands + // alone, and one whose iflink is another index is derived from or paired + // with that device. + // + // It does not say which of those. A VLAN points at the parent it tags and a + // veth points at the peer on the other side, and both read the same here. + // What separates them is that a VLAN is identified — the kernel declares + // DEVTYPE=vlan and exports a lower link — while a veth's peer is in another + // namespace and nothing declares it at all. So a caller keeping pod links + // out wants this together with the kind, not instead of it. + Peered bool + // Loopback reports whether this is the loopback interface, decided by its // ARPHRD type rather than by its name. Loopback bool @@ -225,6 +238,9 @@ func readInterface(dir, name string) Interface { Loopback: sysfs.String(dir, "type") == loopbackARPHRD, NUMANode: NUMANodeUnknown, } + if index, link := sysfs.Int(0, dir, "ifindex"), sysfs.Int(0, dir, "iflink"); index != 0 && link != 0 { + iface.Peered = index != link + } if iface.OperState == "" { iface.OperState = LinkUnknown } diff --git a/atlas-lib/inventory/netiface_test.go b/atlas-lib/inventory/netiface_test.go index 50677135b..35c0ab168 100644 --- a/atlas-lib/inventory/netiface_test.go +++ b/atlas-lib/inventory/netiface_test.go @@ -314,3 +314,101 @@ func TestAFixtureTakesNoAddressesFromTheRunningHost(t *testing.T) { } } } + +// TestAnOpenShiftWorkerReadsItsNodeBridgeAsVirtual is the reading a real +// OVN-Kubernetes host produces, and the reason a caller cannot decide what an +// interface is for from its kind. +// +// br-ex carries the address the cluster reaches the machine on. It resolves +// under devices/virtual and exports no bridge directory, so it reads as virtual, +// exactly as an unnamed veth does. A caller that refused a virtual device named +// the fastest physical NIC instead and handed the control plane an address +// nothing else on the machine answers to (2026-09-20). +func TestAnOpenShiftWorkerReadsItsNodeBridgeAsVirtual(t *testing.T) { + root := hostFixture(t, "okd-worker") + + ifaces, err := ReadInterfaces(syntheticHost(root)) + if err != nil { + t.Fatalf("ReadInterfaces: %v", err) + } + + byName := make(map[string]Interface, len(ifaces)) + for _, iface := range ifaces { + byName[iface.Name] = iface + } + + brex, ok := byName["br-ex"] + if !ok { + t.Fatal("br-ex is missing from the reading") + } + if !brex.Virtual { + t.Error("br-ex reads as backed by hardware, and it resolves under devices/virtual") + } + if brex.Bridge { + t.Error("br-ex reads as a bridge, and it exports no bridge directory") + } + if brex.Kind == LinkBridge { + t.Errorf("br-ex is kind %q; nothing in sysfs says bridge, so nothing may conclude it", brex.Kind) + } + + // The physical NIC the old rule preferred, for the contrast that makes the + // point: it is the one with a device link and a speed, and it is not the one + // the node is reached on. + nic, ok := byName["enp2s0f0"] + if !ok { + t.Fatal("enp2s0f0 is missing from the reading") + } + if nic.Virtual { + t.Error("enp2s0f0 reads as virtual, and it resolves under a PCI device") + } + if nic.SpeedMbps == 0 { + t.Error("enp2s0f0 reports no speed, and the capture has one") + } +} + +// TestAPeerDeviceIsToldFromAnOtherwiseIdenticalVirtualOne is the reading that +// separates the two devices sysfs otherwise describes the same way. +// +// br-ex and a pod's veth both resolve under devices/virtual, export no bridge +// directory and carry no driver link, so nothing about what they are separates +// them. What does is iflink: a veth is one half of a pair and names its peer, so +// its iflink is the peer's index and not its own. Everything else points at +// itself. +// +// The distinction is load-bearing for anything choosing an interface to bind: a +// caller that refused every virtual device to keep pod links out also refused +// the bridge the node's own address lives on (2026-09-20). +func TestAPeerDeviceIsToldFromAnOtherwiseIdenticalVirtualOne(t *testing.T) { + root := hostFixture(t, "okd-worker") + + ifaces, err := ReadInterfaces(syntheticHost(root)) + if err != nil { + t.Fatalf("ReadInterfaces: %v", err) + } + byName := make(map[string]Interface, len(ifaces)) + for _, iface := range ifaces { + byName[iface.Name] = iface + } + + // The captured host's OVN pod links: ifindex 14, iflink 2. + veth, ok := byName["3070a31ea98cce0"] + if !ok { + t.Fatal("the captured host's veth is missing from the reading") + } + if !veth.Peered { + t.Error("a veth does not read as peered, so nothing tells it from the node's own bridge") + } + + brex, ok := byName["br-ex"] + if !ok { + t.Fatal("br-ex is missing from the reading") + } + if brex.Peered { + t.Error("br-ex reads as peered, and its iflink is its own index") + } + for _, name := range []string{"enp2s0f0", "lo", "genev_sys_6081"} { + if iface, ok := byName[name]; ok && iface.Peered { + t.Errorf("%s reads as peered, and its iflink is its own index", name) + } + } +} diff --git a/atlas-lib/inventory/testdata/hosts/okd-worker.json.gz b/atlas-lib/inventory/testdata/hosts/okd-worker.json.gz new file mode 100644 index 0000000000000000000000000000000000000000..09f127cccc9377080292167a53ce23b92eb98c69 GIT binary patch literal 129487 zcmcG#1z45c_Ag3zcjKZ<8bm<4yIV>cl#uRPgmiaFr*ue2gQy@O(%sUz==(1HzPNf+ z7~GW$*490+r~C@Bo;f^D24wkhcidxr*Y19CQugso^o95NqEj=8-_7Ce+PeFqVQfjK z^Np9F!q0XN*#?Ww!*7$yyJKTlyN5@9vrE681^7H(U46Q{n>usP%e%i3dGtNIJ51R| zJvyI_svna{@%b{4=4uHCie||=LoOWIx+_+GJ=$aYGmoVh;-gPpKDR%KkG_BTwDmD_!9#l75L=lq ziX+b_a7nZk#(WNc{l+9wb=Tb%qxMlYz~Ul&(iH=|OA>KP638Df{%xI6KFCbqy{XZ% zlT5xat>5c>sZS-VBy=UO;**EOc35ZB0PzWiGP0Q-DS`>HJlzf7Hze_Nh2Sy;iToJe z@ej5FDEO1&=~_sjDQP&>Ph0vJlA-4ORrGA)Arj;7myH*@Jne;LkdxP2g303QKXmT; zwk@ieA24x>UCMmsCRMSX{C4gnBNvtDud*%tbQv_Mx&Rv>xFuHXz)qil zsch(iFg}n3dzLi$B;rRc7CaG2Ae9I#V%UijrHDKl#>eUOv6qBq*g=={ z>&oInGMZ#0o=WN!Cz+WcAZ!28#A$z zNFwKIux9u7;zw(AF0=iTm_}ViqJ^}Vcz+zeBlz8~Osb>LdndWgPZYm@(D%)_x2|P2 zfa9W=qD;=WiOaXgR%=AwnZws?WGt~C<9T-}TxhRa7eg;E7BPpP6*q>}-}<8D0yDo&svo=2M`vNl(mpyF8q`n90YZZW(QPWS3Q)B26V zL_SIE9t6T%-2KG&}fDo>QJ|)dF90;Hp6pqZs0O*csvvv9d4yzkv+ikkpU_(-Z#nNkCFxY4>{X=4-x| z2iJO%=WFMIZ}Va-N8#%YM&}q!iiXf zeNG-y6(MGKtZW0S-{Nl!ASM6L zlv?!ybk*wld><>g8Ji**g4Lx1b$WbQRlLk7~i^l zeSWJnpKg+>|1@tUsYZZCJj(sjoJ_{T`#LK2c?qu5>S=SPe!xe4b$g9S@=g`LZfdq- zx~)F0?R}otul)6`njDyIr>|UjhIt)*8>9tfH-~;IH-2>$k+Tl+-j9*(@GTV|NP8+f z*{7^}8_@6=uvUO}VKsTpX!JRZ)QIQFYkjKu0 zOHDL$8pJZ=?JyhDwJL1+Yp=YxUkZ=fk!U>?@f?Ipwlg4=2wIC}<+xxr%4k=Vm+2Po z2`}#a*egDxuj-7n)_dvU`sGQ0!{qCaCE~l)T_Tq|`3N8a%Lb--#8ACNoG*<1Xef0O z5>*j6tPK5olV-T@h;b@#fT4Dopau^7w}B!pfmabvrv;Lbrzg;ocur5CDZyNfH-Y(b zDn;@lY?*&IA7aj#wWRvYUeHM#WnfMv4lk;yI9YN;s6;GsU`|6GI^{vvVCjkM6(tqe zrc-5-^{Sr6FxTlB&XxABIWQMIw`?j|Q8%w^)=~qeWS3JfYZDECn=0)$e{BoJnTUr% z;6fo@K_M=g>jVrHY>M<=C!aF`KrBqZWwW(IK# zlXnJjwOTOFyL$cP12_pb`G`$?bD)5q=76qp5hAnmYDqRM4IvLMLX%JdaCyCrC& zLI69_P#h}r?uU&0?~MxnEnDUN*uy9D)5%h`In{qEQqW`FbQQXl`(qBD%m>NPG2W{n zV4^0;3xu_cZ>pwUJOG(OZE`>b|LhjqjA75$mM@+}lRo(nRNeu?v_w%)2xq%5I3-Kp z%7h#~oBn+@IM1Jc&m`$*32zf`WCjPkTbl;WnxN@21wO)y z3ay-gE>*rtLWF3Rt1W6%c9bWs+hqvJn(Nd1V&2rm4bC48CIgkXv(EeMK3>bfdpsf& zs*Dv%?!}}U3JQpaCqIkEiLH$rNLPo$bAS8HAv917MY}NuM-5Mc@HH*MIcmYXW?VR= zP&Mg!6j(`#viGk;193rG_^H5T)pELtL=9G7;uS(=ruxYI; zvU~y6gOTO|lO{;B`vb+z?VmdpK_0VO8w>th2)wEjIKBY^QjYvJDsbhl=DAv1#)s+> ziL=bT5PxN!(bmrnHRYqv#Tb~i?Drp0m@CpAm_EV@&fM9RV#T4%IP7+_htmMV}Q)KXEqgm!yH|lf#W^!$|S4~y08ap*KiDelTGflV*0xZ+R%QE}uF+ee`0>@7#3lvXWq+)<0WmMm#;HNUo#Lj~FlO%GCDROItJUULT@nG7``a}hODg4XbPw0SHTnH z{m7vuO>tq#?e0$3=MjiJ!jm@C8JvCF0m8=YI^<>{vX0j$^WK6ARw1+I$qpq^ZIm_5 zgWnr^Hv|i9)CsDG*HS{q-}43&GECpRDim@eEU0}JHh~mZ8C}rGAJ&CQrAsQgJV%1W z?`dR-0dN0`WNyVdg$tuP#>^sQNPCd8ol03SKuTXwt_so$fkPFvvhfU0p z-27hg1ES6Rj#d?8X7fAR2AO24C%(YCGQ&?NRna)ojR)!;lyCeLw*i5;nFR-1jE_&_ zZv)_7Nh&)psKjG)SI|2}N3Pmfr6GjMu#8+kkr9%5)!?=RfymlrV`~iT(Pv4;!*5x{ zw{weFY=s>?a1#w#F;El#iWx%qGPkwl-I$AY8n=(pip2bEI2g4#o=P(+u4nr|I*cjT zvtvhE*L~${bo#m3PzY)^BGEAH>{sFBeNN6k9!w9j9t3gA)zA>;rpKR`-D%Cnz zb-f##kHWy+`Q}TSGyYYVglo4i78zdvy^0Rmg6FK>BAg!BJ||)#jC51|`LEVn_Rz?> zs-F?QSZ!58Iy#4O=sPnw3VAZG(P$a`W=>4utb;B-17SN%>ib0g_$qPt)Hi30ef>$| zgsD-JD}v9KmGNQi?2<mY_TY_%1VNQsrcni1AsYYJHZOU|U6J@pyaEFz8^ z-2H(|U(?1(zxjf3lSKACIfELy!!`mPtVGymSK~=@ZClaNi2&<=2@LvYV2uAY@C&v! zEsYMZVkCtL$9FY&7M(Hnwsy?Gv3E)cOtSCzMk&3{4t@uIP` z=B~zg;~j>mxE+#=Ez%cBDjRJLQb+K6mIn2%4fMLRNWphKpD?snLA_5YMdf6%W zEQ29w#{(?ch!twP*mna~ulWcwf2x^y-T;pu)p+KCA7q^;d&VDh9FwaCbLAH){d|3{ z^l*C4OE|qL5`uSCq#y(Ij$M0;zzN>m2 z@KX8NE+?rUjqMUCnK-&Kj?S)ccFh~a{y~_zgLGd~Ulm>ZsgT5`5$aD1zkbfE1FtCU z>uu}`i{3Iue$WXXI!P{i-x#HZoVdO)u+X^wR_pC3tc?1^n+jD<6{jb@X(;M*;P9Lj zpPXwUf|h{A)6;UNVFQBiu(-{U3hNPcfl4slrM)0~b)0zEgeGP@3o)(NiBg;v2PePn z)K=A+B5WdtWri91K=ujwfF-WlYJ3wDlv9@jOHRK}M+>izmP(T0x0ceB;eQ3vjv&-P z)CZ-+VU}YA$k9tD$bu&eB_}ch_Ukl@$sSxAb+jFMBwo^qe{jrSezy41!|iE;w_nnY z-{fLR(lvQ)kyTS}f7ftxgW#oG6ADhDx`0GN;6{L{31^h_gx4lDRaWCqvGD#Gd(K~c zr^R5(ly4AntxLYj3t#*kQJrVw*Ju&rOcVujVyB`?u3e~+3O8AG@;;cZa?zbg4aVd< z9ga0;UGkegi|WFi%$j8IB<`kzWb1ze*9JSwuAXbg;$u%_~5Jq+jdYC3%lI_Hk_+VOfPM^td|4s-CNJ!#; zFCkx>nzMpG34tAhxB|_PFCRmpN8o$3uzlK~T-KmAHKLvRpyoAK0(^;hRK9Rlo#*|jwZ$fUKFPmH zX46xnrI~Q5p#80Dli}G!)pXSHiIuFIZ-*D|Kgxo5;)fRzdzsZKd&k6f7_erYJS(HH z*ZDQTFcDDohUT8|_nwdjkqnDFup`$2nZJB3h**>GC-dKjw_BDih7cQ9pdQ)tS>={K z;hL;I7Gm9rd~xE4=lVOeukagQ1#3a%DeuuRnOIQ3?<;2V)wB@F>Gj|q)%c2J?0T{>%F4 zb6rOO5Qy<6L_43ocCjPj$^3JQt^;AtfuqsgBbdj8A z;g@m6kZZlA}dRB4~i0-u*|olw8jYOwr%r26*`r#-{$+GW zw5Z`Y7Cc#!5bpqg^h8HIA=+De%DMKJX?LW@)y9*97BLd1pq3Rl$6myxF!T;wVkTC0 zt*@{?1~-)lGFEWaX7WTJ;YCzkVBdU32570b1Y#{aUXj8)GKu_3nM zmqOUf&-JVs?v(R277K1wLZcn84ul4~X`lLD>A`8Nz=Dngi9Nf~g;x-%j^VJGf{+RS z9+Q)TxR_yqfk=(6F7?j2W$Dxg2Y=_dJh3i=kI>w_v#$!Il z0~p3*|EmvjL07nNN9%}7L+BlG#Gcd~$07< z(6NZhO30x}0vz=d1d(oZ-#}sr49797@d|u$H!=$xbs36pKFx7Z*l)oG#WxTS0>gX^ zE3pEf-5tL8i3U+b3dJ{q7=pub3^HDk?C1_>3c}V#5YfVgs7bF#V*VDUa3DDJ#~_Ur z$To721w&SW6yGA6C5kW5nOhdP<)1hZpkcRi zK}WO*U2%Aj(E_Q%URwUu6Q5(qQ_njrVSjk4rpKE4`USx~rP{i<-K#*f~$x z!*<%{Md$Xz?=PMYOTmfJ;lYR`z#K;+1gg@(_|m}U)G{8xW18rU*Ayg-FNMCKm^0EPL>o*qx^O3z%9Lq-+TGbegV1ty(_@x*!RUk3`U*PY|o1!;=HLTn^~9@lcCo zXo95Z@#HK30w|1Vg7&FA{}Oy@g8CWpE6=92rJyZPNr(@Zi{PPK zz}47B#XE&Tb@mSx73s~RTE&h#ww5g`q)kbpO$nh*c}pAfiZ(@)Hb#&DRNONa#1L9mMe0RDRPzy6$D98;ebZN zorWZHN+pN;K>f2g4>l+P+@Yc*y?N*bv0yC~uof0rYc!NMHdIR`lou;hZCiYJ{Yk$q z&LL;n85@chL8t&~FEzsISkUp17)0~Q6exa#G=2m(egrjs5H^0~I&Sa;d+!|C3lbu5 z+e3M;q_az%=0E27t)$SAsy@3rRTBO^F07%NIr_~^Fl@7F#8VP8^ud{6KMV1FPB?cY zSfP+WtC$?|?$IZv);P61C^Mu8`xtCFpooi#MmfDkd`E|ldq`YJ$ zwvcAn5)3kB^O5{t2r%e# zC0N#s1ei0ZB^2OkKVT^h=o?8ikz%HYl6&5R5D=bHPxJxnsL)WRKm^XdFzRR2fB_Ly zl7Pkw1yG`(D1J($Y-HvN<2-vle}{nw3^Jj>W(=^R=>IJY1MCarpl|>NAL`c)n9D=` z;-GmdG5f1X{)aF#-wTPn|1HdQI=3Eg=+1k3VoqMw5ic~e#|p!w6&Lzwkzjb z&`SntgJz^Gv|N!FM7h++>X6+G70Y23lbSa2&>p&s(dJ6UoK8-yTM3i(+DS6I_({ALDBtW5rQB6+gXFyK&m}KtR)K(b@xq3Y8BJ zLP+$n9QUze2~Mll1YsV#9vBKRY5qg`{SU>*o9_y1%|4mYyl!nDH(8wj1%5rYSBy1~E^b zITe$=Nj+j=A$V8OU9l1@pA z?1{JZ^W;fZ1$UnYzl7Xy7T58#6a_S2c|X!tf zRe^cB-D9jwyfC64{;(_CgvV!S5HG~Ljv%NPyQq$4$li zDQR6dmLFN{(}i)n5cAzH{?$xj&h}{~dQTQJ#oGCtXRA@84e@PS7{0$zURliSZs!Z~ zu!lREMp&vw?J&gmY+*S5hr<28DgNQn__s&?Z}cF*<6s(Lt{N3kFtsqKy-{Xr;|qFg z54Szd{tn)~li?VhKS-|*bt+ML<);JOCmVdUO8$ac2Yj?&0CB=cV|`Xo>xz$N4-ij4 z2M}L;$VHZ#UN_o)cKl`p6J>0anuh02pOH@LQ%c7es6l$_oPtcK>kr-uYO9%x^L zQZ8~v;8ZQarS<)q`s?Lvr&>B*EMJ6E+;T@GRLzP@>!&pJf0eUcYU%i}e0h;_kvkHs zYB^9^zoDsr#>{r3rQ^r)MFi}Y*W0UVwpm(xuBms+%=QppDY|j0>UG%Ph1f&VCY zA~O+EyY8B7LB3%RVwqZC%3*y`Y zMV4T1P>^>DDe~hyehX+=-*ue=yMp1*fnXn2VEhn~cUT3)OnP8%{-~fH{|81;b5G=* zb;`d3Km{P$+rX|&`11s;??cbKqww)fDNO@C0%3b$k(iK4R(=aX!SEOa6sEv_z#}4) zFbQ;v^}r&bVQ~WO{!0aT{GS*>11!NJdHx*$DhL1q5Mv7lz6c&e4eK~D(Oebg1UWF` zOmymm`r#Zq08;97S?Y9N>U3J_bX@9mSn9N2>a4~JC_`mwxB!-YPy^E%H=3f`Gfu2o^VX@!fH_fbzVXHPOy zOfqLrFjGu0XOA;ej5B9{XQuehoIS=&F}A9@O@OscFuFw$yG5X~MS!(MFuF+)yGfw3 zNr1IUFuHLp>w?DUf`;q@ITN^Zrg-4R3Lps=jq0NwSXEhbLUy*>b{c*qu)w@fyxX1I zvl_R8!{}57FK~BF{s2*4Jy%{mR$kp#Ufoe%-B4a#QC?k8UY${1ol;)?jv263;Aw2KDR2h%se#!#KA5GtQwoDKG zQ|12ORsZw?aR2g}tG*Zq00FN?!(Tv3(2z_(thT1EL?FU~!KW|Pz+dgDdMBjGB{1B3 z@+!P=?Wu`#~1g_dZqcZh_HtY8aLc1UD8tn!Qnr;M} z8cWa{5zv5fZcs+Amj7?V0G^`>LBS3+d78prxo#T6Ub%Mqu0r)lX2^as#k*n zD_!}hlpHY*`lGO2ZyrjwsmQzjydpp5J9WWR3D+LQr7nZJ#1=>PfS&ml-T&Y6Dfhl* z!5Ey!F7)3)^h5CzO(_>6?$|h}9N_S}-8v=i@bjTB(3X!&q2l;nLEmI|ZO{btq(G-y zu2QE;yYhpkqWwoP#eC@a05*U_{I^lecdVh9Uz93l`>k%8!2Vn9G=lxN4>#v;Tsu-L zT0ZpVciite-(2*L5x?+J-+3=8%IfwxDvx-vj8L1b_C2>tddb;0;EfeNB~2oFs(fa% zK#;Xvovl(p&+GM9o8oekI8aIwLa!6IIB^HB|Btmar24iyFz)LQbd4;#brb7|22Dje zPfaRy^0ir~DhO?Vylb+%aI>>jCbmH(0LD&}+Mn$I-*Y39F|d16cudvLr!&o~NvAVS ztFN{5mSSRmYPg+hlO(s~ajZ98^#GSl)7xlz4)djsWrNJWtxsoeLuJ-~-pOwAinBRH{#L;uaSrzXL^i-uT z+a&9fkI?oOCoK3X>0Zr#DjYE9&j}#e_rFR;%ZocGb>&mF?!3KS8afDf&F~e7w&=qy zRh|9V7Z<{GcaEV+Qr{b%_{)|5`cU-6O+bTaSC*1BX-jF`! zYHnE8U#Ay7Jlw9_H}oX6|NPY9dn$Z8PVy}8tHZnPtRLn&ygu#v-veu!&%T|DP|_8C zJAW%UEV_*$Y;bP0(J~O|xc%$ErA_bk1Q%D-fp^B}&~ADS_B>#;yrCxgsk77TesAj@ zyYaMs5w+P&60tfMg&z_3H5yH{^Dd!OElb3TE>2JNlVC&fesOV(O?VqZgc$;aIRXTz z9zl-~0YZq-`g>rZnKMx*>WI^21rPHhCJeToqtWOSN}aMsZ0O*8u6}Zw|0L{mb&@4w ztmKJ*Rxku=u+{*GAD)P`1Va|k)~E^k)kNS75K(6c=xj9OyGhcF#o=@iQ6YqMD?0Ik zhT_2?!6YNHa+Oa&S;3*N(eTLDCs^f%E1rNp1&1!6;kmHNd8*+^UPssv(wS?$ay`d}mWo;(Oy_`>*vLE_!^PeiOjm`Lc+4u- zTm1yJC75m~M4~Sh%a+e#xCGS4kLhIgdu#w~0Yv8sO)p(82d zK*nu18NUrQgpAw#M?CJQu;^Aki~ed*VPO~vRY@EG^R)!D&X3u~H(2~Lo~tq}`WjfO z2EBTNNlMMvf6V)$_(y!Qvg8C1$qX8axiYNijx5o|UT&lsw5u+O5jV-$NXBh8884p| z*1-!!L9Gl^g78`NR)g^0U}jM8_3sP6{4PRBGn0#k`&t!D@39q!bzBjMvodD_bv{MayTS`aPH>!rAG&4#{%_o7v3<-nRKjZY=SVXJ@a7#IflG3#1Vc-riN_N*4 zH^u>WvVdJ7sKk(v0+6O{_>I*{QBj_+^00C5fW^aYz!X@7 z`2L0iRvpFwxNXUPo>xZRm}Y(=QpcjyKtHf?7XZcj8OW4r;$P9)fI|xuK}`esV**9& zKe_%>1T|0uHn3~q5r+Q_tyQfUh*A53nr03#{Ee0dJ>$JMfJwj~6KH-z+mduQO81A~ay-@tZa*9?JzNa5HxOS^b#Y5o?hc*E zO0HCR3kPhzxc;*E@_geEu*fpcDwAU@DG$ryIQ=?h@ z1>N9@YM<%5sgibv`VLla$Y~q0^;N?Drvi!U-Idj6EJ!gQAzQNtN!CY8ONm?{jS(y- zi32I80u4kUQ`{xlTC*tW(Yn${pS*7$qrcB>EMX)Yjgo@14BvYVKpJyAWaCP{t7eLYj>=f7M&u92nb*N@KKb7cgWlBV5m zu+FECCFRgjqpJmT=(aO2vSJxB@<%CizmnWmn_T?FyT+vOMw;Yf%9hif&T2)jm6~}= zFA0w`vO>prLVmoidaCXdH`doHQa6SnjIOM14iYUI@TG37e4YT)SM!lj9CdU<$Ie%7 zQ2{eYQki)XZEE42W2ifXuCg-!b0|uSrK-x-5^Kqt`2Nt|3Z` zKe=G7-jO0R8e%+W!(0iiDUMb7Ww-|k9WBN(zFr7= z0a6#+HJcl^b!~;s8|~u_n)FIXa@D||3{>4VBkN|RlA$Zw5vD2$!YvNB4FnJsuz8rh z(r3Q+Da{Eh*&r$Gn=$xOS@i;>ipoj?RGaj!JFJMRgfK@g{1+sYssuLsd*9T!?F%K! zJb>w`1dQH|ZkX*w%K=Nd(q^bHFtSBq)^UaLr=t~bT=jUOwmRT&UtG_tj&(A6EPZe- zl-fTg8{f?_VI((~o%t3u!K@>|Cu}T(>`z7~s5B@$s*sJ4^akC3ktpZmQ z!tQS&q(VuJPKXrceJ4uB&x1gP_8mg~?5Rl_^q#qABo`o=e z$Xmt|NOvjoWu<1{{4^XtfMK+y;D|%!nHfGyxRcbt)e{vN#+|^s&>{n)v|-mO zQ=FlX?_j5~K|z$5BfhrVmb>mpsp@Xadre!btT1F&NEt*FeZhKnMo0GOC#_+X&hA&D z%8P|Ln3_+VE-d?M2|Sbci0ULH`)cFtH%cRXP9e8s^JbBkcyG|d?+~l?uAJ!^M`;4$ewmS{Bcv5`Kj-2=WIJzFs zdv71WkO%PHLowum-uIFFb|v7E8{(g`FgBgUA^cQ$U*(aXvNTF|tl6H5|4ZER^tQA4 z(c4kSE7r*w!$!zY&I{d!6J74Zj7YMPp+sl*ALg5hx&N&=C9A=%v*xc zI&E29uV_i5Mk(c$L^I1?viip3C#(IcM99 z+Hf4d|H9DPAHm+R0>?PP7J}Zf!;QtxDcJ8}Z8_t4vO!~Hs*sW96hwHvWt#%~x~uo( zWssWCB@Byc#nszXb<(9VeXH+e6blQTpJ-o6P?1k{>q~^_u%9Q03$d;cUZao zAZcXA#EAFsdHnfhB{Ny{ffOP?3h`MMbT}R-559fS5+2U+CCnNaYT+vIIH}H#cySkr z9jQSn@?!~e8PCoGWmpDLk_lE{_JQZ@2G>drk5)c9X@sQlZG66=jhqdh*_rZGU?%} z{@;FuNwpprUp=;YDG!8v`*p+LQR4`Z)Q%dHvyWP>MRS|Fjg?~?E{50fS%rd(L}~Ng zVFR1f*G_ES=8S z&qve|hm&C699*T~%_`EB^T)&d5G&){dwZmNNlSC(S&_OiC+*MQJ=~xt(|abqkaOHPokw_uIJ@Q z$P$aP6+F9>`8dZ>8(4|_q}q4*8S}4fJ^nTFnOcYbpVv;W`+q*K1GrhOXKU~bBKzbg zZmEEwo6J|4AZuFx;YMLHD$$s9{DusQR5Lc=a)NJ6vm7nCDj^qn0!x*1% zYQ(=k^GrvQNd}%jz%z#V1qOkM)nm?LdDCMUTGL5!#kHdrMrOy{t@Y8v>BixYGgwl5 zbB>=dqbtXPt!lSj>^ennYWvH6A(qi`GuTytP_!gn%fg;GxL4ZUodQ=LDL&B;933sm z8BI(xMZSK^A5Rjv%F2;&%8*>kkZ}0hwfWuzbhXjH>U~dUD$UoOYZv(RN8pzVahg}o z%6#T8g|9r^LT*$rfTEpu5A)j$4p@ov3 zRw2S@30aJhr&eJGVTj#6>B*v`#wrh`W|L5jXTd%>%hX%#LUD&5)p`Cm*Qph=Mjr_c zorGeHT3uP2*WE9pRXU%9IH>7z@8rm2klyW;m-0&3Yf|y3xmcP_;k;>D`GFWLnLL8Y z_e4(hh2ePQXbSL7Q<@9B;gsgOxz1^h35u+iTT^y;@1Yj$C;@9SMZg4o#}U53ID_+W ze7sp(oPC)X{$bcDi#==UUT*L`y=7ybXg;j!VTVajEMZ-ry@a)s|4WB5N;5_djZ(d> zMuneI2X@I=C1*YEIhlhdmd6eDte-jOQ=@p%G&f?PXztsMXh>7k_&fJ)6X4cq#=+c0f(L!Ck-;sM|@?350$9UN9JY{IoR+f0bW_k-Qf( zhh3!6Q0Qeg$}_KzGg>V^Gz=n2}dl+%%W?+P~ipQ8yJ|g|5ozyWDB{%%56FgwF5x9iB4!DK9nX zZi(+t057?t(3jjS=u7S~@RDmYf%>MOE~k)v_uP+#%3aPA$#{QC zMm>?=5jI*-k_@_ru%(INSsf)i>(_gXuo>KCc)DzNUdP0u*M&g$&mFpdCcyqN>TOA7 zw`wAl$}cd#a$z`(wr73KDIw&mApetlyS4fztGW@Mg=%BYaKLJOk)Po_bABXJ&r#JE zeBPU=<3x6(5nhV!O;)=boGU0``&|<%?aid*4>2Nc zqk8Z8GKJ&{3WjHDAiZ#Mmm&C0IjrOK9Ls_g=)C^<#!sO!AD6429!I-+xIS$Ru1Ik| z2g2rV-A-vPLM>GqOX~@p5;bSh7}m?HOuR{TA6zWPsusgl6J5pM#?Dvtx6FD+ze8@j z)Tz^VZHYnc+-i2AnV!d|hGX;!J{vn{rMs~CM!~d1GK;o!@Z*iVks-P@KcDSdUj_F| zjO(e=B-SLB6fb@ONi2Rkjy4BrM!#iC^tALvrG6uj6Hh~w73~%mc^!QN?mR#FiS}0+ zU&IF7$)WGx9LnDBTDdN@evGv4eItc8sz=*q!;7F7Wuj`!DzsSPM4D`GDnZr zEQSw{wQewCQ7KyBTmMddhUvSfGO1tthnq7Dyct0yLx)hTKB&-5_Sgqh@z9BeeqPE&F(5do*x)U>C=B?mr%o$eDtl;+3 zQBC>qO$U@9`&PZ%y!a26JUs-;kkqfCj`Kc@cwwN=_G_W98>wsIc zwO4th@l)$rLJz9YoKt`gxC+gTu~MSeN*je$pD_4Er9^+W)tU1;Vc~YQw9Y4rO?{MXCMGj%8o0=&}g&} zCF4E1q8CB0I~BN90r*Hzf#s@;T%wYBQVfeOHQRz-RKf{I$sn*>>xT;0A?uCQ!_4aw zIi;|nw@RcUdg?@@cxKBW2;0kq9>;XkFZ7kJunn-8b8CL{$b@VUv!)PuB zMWLz#Q61MM~$d^2_HDuWY{^fA7Mt zALSA>ls(V5q8#a{IuJ^pr7X-8DlT2+3ly&!Eo zevdSroEMiLrzGaheNsD^Nd3SD(bfD=@Qq} zNiPwl)xP{@YVy^c>r=4vy2*+$R`rg)L1ZzTtlH0rtbJLrnG_eibp0gS65>dZ>-5} zCAVU(xH;^}k1U%beyTlQV2%y?6oZI@8yI3!;hNZ|`GcEP9`<34GTOnRpSUHvEW){| z=gE)m=TbD!?}6L>%(JC2FVHtUYxJs_LKac`>9UwQO2es<#bA2DlPz9!Ms}q-b?O4z zjn4W)f>iBQGBi_@{pa8u<^C_L0eODQ1u3-ab`_&ZUdshRptCN?9NI@fmw+4;i(&rn2d?B##oQ^Km%E22Y+b&ZY{jzmS4D_ z>;o{VLb<-ZuD^Z%VsW@qh387M`~M^uYo>ZnjiOG1s5tvvW4<0 zQ#=Jfds+25R50gmwlA041N|1uGR&7tHzV%AFI?W1u+RDT+Oed4B8B9r5Km8tycTN%={7*r2YTv`H=00ciu9hcjPo)wPYpvP}^BF zKCzBASc$p%^Gqhz-L)&J3Gx1DYwR^T%ffZY37RCoq}J-I=OkmTdC0>czrbwZTPfV3 zX-EJ6+_5Mv_FpJoANHBRmfEdha5Ka{DrVMr&R$za@LQ%GU%wb|8e93O7tXo^k?Zoe zfx35nfxE)Z^5?uoqU1J1l)*0;*BI(R*``LRa2pXIvD6{e{hytwp8w-D;gB$5`EMJg zMC0$C_rbL5vAy?CT=_!3uACE~;X7Wy?MQHWU@q*zm&8Gi6of>&J{dPPQ}LluV#=N@}W6=!QYdJil^Ps#O3`pUkDNssTN z2l)5Vo{#E5Uj~^1w-AX;;1<#g{4F29L=JEZF;fL@A;8&}>`JfvI&|43dHrxC(r4U# zQZW*E|8Ru&e{)FW`^OF;cwar>bUgLZ2fsLQ*t&Lmpy=5~Aii*1vytC!*PH+T2*hsA z|GFpv?BuX@#E7p0^7^I_EPfcwwP%}tw?Ct4QUER@CaNQuf>6)fpZi;h67yASc6}PEip;qHr8WAVQ;1N z>5b(Y`e)%-FUv;8GWqOYISdgsxwM!@s&qX6)*Fd#SmdHFJM3dA%gk2U%SX`NHS-ot zC@CUu5eby;7*~E=PZXUtYdo>@-RHC=&olTx9~8`6?$vyXfkn@pM@#rq z_h06$c9|3g2hD2J7QYVKKQcsx7kysqv90$s&19=7{CuIS7yRQ*NIiq0m(+<2 zw}AE@x#aVM8}IWnTNj6>`JlgipL@z>QV0tuV^O6uApacSxIwS>tz5)P;hXiot*QCE zDmk%n++*tk>E?aj5z8RUr+jF5;9m3>bS#jSEC5{yK##lewmsB`ciXkRvQ%zf>Z-JN zxr^t0?(dXIA{&=IB%y(V<8y=b-BLL3`=oxj#H_9&m zAKu zsW~dz)1(YrE<$0xFYL}$lebJuz9LD9Mv0hvJB~ig7d*Ds1}a4|```&N>q6Yl^$Uj` zOp-R1M`L;2uNap?W(CRi*@Cw%RukwU^kZ46xLs)JFP{jL%&c*}`4l2>e zZoaI(N()MW)5>SVodV#q0XQY0nvI=uVa7 zPWoxXd5m2k!ys?5mGMk(+m&HpV8RN#*s?O*TI&+qMd9qPury)i!eS+m0q=z~78jNX z?MudrKJ`?}68k=->wN0x@BQi)eFT^__)+K2I?(^Y@e2Xv&zBUAtcjPaTf%)`mwsx{ z^lKcH9))i@?Dcm4||A>?1l7NCg>2vBV~G*aYQot-{xg1IGDBt$LY08`PF z4@nnQwWZAgQ0qzN~PhW67*nt}Tv{yx3F6yNEQpMLIWj5hi&!n`K9U;L$-HX3cBmWY zcoLyo@>AzBOytaO*p6RzZ`dU@F6H2@PPEE?K=s&xa|fseg|mz9!_}4W1%<;1*ruX6 zU8RT=b?*^!y1L2axjJmY&Mf@EH(p)M$p~=oQWI_F<|2KChRsoFZOQw<$+v}_>@x|o z;ONZn-86QM#n0b3XF;)QOqmo&DklyryBYYnxuXaDU6^N}U#)L+@++t33*SLWejNY* zt)FK;y8AHL;l5uI^og}Y1^H(9e%6E+f}MsU+hs(^nwsD5ASi29*N@i^#_;MqJ)y^X z=9OOLusOE7=QDhZi?63Nh;#}7Ia}8OdJ&FJJX8;MJ+30u@8H)dj3MXmc%_ajt6kJ! zkEBEgXB@$L7dnh%l9o9?9avdr2TcHQNV=MDe?bQhw{#A-yn7SyDrK^H`i;3=JF_vu zAN>8QJ$F|f*7u?zl}{5YTe^^?D?@$yoN;{H;?nb?hb)Br+_6WO)}u@;5vfn*w*RE8`=iDT83{{a=zcV zS~W&?b_xnQ>*WiD?)}F$q4FT2u*Hzsde)=R#7OMonU=L3_D23j zr2p=%>hGwuQviQ{-qUXJ2A}|(6ExTQSR52Q_8eh$BnfMqQg?n0xFpfb3O3QOzM$NF z&XlRpVL5(I#B-{92a~zLir$`fxmjF=>R5-4FbRqVyZ@Ht^nJK(gC6FI%Fq9*F! zk&&CJusyEPkZ05D*cONK@r9P}vTr? z*R|ND^Hg>D(?QcQ_LYEF)UO%kcPZ{O$DPy42~%yKKBuo^GkZ9r;++)zL6OK995Ez2 z`oB4p{xxPtl&5bHKN>00T(W$`yFwfI@K1k#R-~DOm!t4v(2~#w!lse75Zt?O$4~g< zA3GtQBzJ%)yiU|uc*{D{mN+cLHMq@0aOTx)3!A*EAuDqX^GCds5Lca9L&C7BcUSiL zl7~oJ^jW~FSaq>3W7Z&zc5b_Pce)i9$K>XKZr+rf=@xGJbR$;g>~AxqDs|$v3l(*0 z?SVH3O?ap09mSzLfp%X|J?E$l_g|qlta7T5nR#>h4+y5McH7kG7v_(X>-*hY_1j(1 z5q4M)-=-~kKd3ZSd6Pn!YX7iB51&tVT4X!aK(Bk);Dg2L&(Z|At5!2vh#*^awrqQRS2kB$j%3`yD;$n1N)u_bPIn9FwukJGOxU z`yGaPl1*w4CBvJ8@n^L|gqC}yz0Jy7ne0F=sss7z$2y}6?;!kjxs!9}4wfKX*i9!$ zr^Qeh8n~n~;#iD6)9B$Qoc8V9HX6R2Oi^q_S#1jpYxvQVk3#u&#_^@s%u0C_mWFo~ z?RtJ%)OFeO+vbdBD)3;-8`-nK_7>-~AE*_2^zRb%3xeKuRz-NNp2hP$RW(>0-0i)& zxKXmpe!5}V<9Mo9lHD@ANrv%x*Ym&%%(1>M;%jpwWrFe^Cc)wDP$^LxQ&`ekGsC94 z?IUfN>=yXR_yTaeJ9B<9+_-H6+rci{4h~aiWwA>bAStn_T8fgR(d{ZyMf{CQ80M; z7_pZ});bf_MWrGuWr=vrO^dvF3n2BvPc@J6pruQeB6zdfuNxKJq%+UwaQS&@)-CaVG#YbW@ z|Df`6TAZL0LugDTE}^=dygBg;$KBRZ1bg5}n|#<>dB4*fd^fh5%zys56yb*C=&Zv! z=yT|_mGPbwg<&!@5t?nLMUjZWdwB_3^sP3m>^p-FOOQrI{4j zR>?nHRct_hoQ)0^Zf#X(Q{(dQ|7Eamvy(34vVoR{zwtXKhP}v2r1g(%FKNSvuwA9b^?L!|5d=k{0|kDmH8Hvr+Cp1Hs@92 z@zOsm;3q6>e?I9?!PbMnw-zTM;U4GJw0nrXaY<)3U>OwB%n87z@54(Cge*JiU6_!)v9g zN&6j)6&uiNXp=)b+rL#_i<^BuZA(-Q)_&ukxMD#SAk89eL!fqfm!kEjf-P<0W^pzvFm&B_YUqoN2M?Xk&{g0@D>=Qs-tcVI-$BhSsApmw4CX=N&yf_=zB5wx5ZWr^ zopgqx3z3onvzXx_HeVg@*cn(RdMp|9NLrU9b*t<6``dlL&Za~Eg}{T14yG-FUogvtvcr0G^2 z;>C!b!1=QjI8llAZ(vJ>bCv3b7$dd+&h>sCb}H)A;4%9@_RH;ndC>iJzq%qxXu7y> zM$yJo2C|M-ws9@x;~AEr3g8#PNzdcrofr254`bKQyu& z_3zgy-H?>=MJ>M<{Wuamoq>(QK@IBwyrGRDv_;CvoXskosg@t?RY~J}t+bX#EDs+H zG%ol@j~?0*BK%$amiuUuhCSeK{%i&mx%u*XI`!O;3pxUa1#f)Pfr zwoftlbj;2_U#64e=SMbpsH}7VZ0xD_yQ0?+N6bUTDP$|-3vI$lZQL#b1kx7>YcI+7nNQKj0?EGA*HR#b>B?0$~$(%#>mhDBKGA! zjTkv|IP+#taDc5}OuMAmcc`n%&>6~QR`6H9Za_1r3$Z0J%QjwW?Rlc&SXnoOKG)OtfbY)F`_Q@nfo3yW?_%9BX zit88#Sv7d+X3~AuNud|g07ebnIr-I(sdqP{{{zX)C)5$?_>teBDw$8T^5Uhd3bijO(zmU>+)$}zQ3~Pf0-VQxZ%u&NnLyR?b zd`p!}VQR406B>p2G;L-nysQ75a`g6j|7QH;bPuo!Q9L?^`8qi_Cp%**vp^q)-6(4w zWm@qXf(G4-$ZZ|ZxodPVyPod835xkdIM*C2oPp;)bfev_;2Z6OXjK`}{ zWxRW7@<{LC_UvujM&SSs=MG^U!`00q@eO6g${p2>mD%_5w3~ZJ)@HGmw{lzjJE)pZ zPL}SFsSfLws9BGFx1;`zlgw(xp2fi`FP#WWm{e~9%3;FZk1C8ZH?5A~mY2EBae9|e zPl02`O2G*9x!+Og@(`j*zua?*es5u>AQMAE)D%rb9V$jZE%_o&pZg8YIYx z_p)Nt-$UfdkWxfZ>4UK~5k8e5s(*h&l~RgCKaBjp|4Y?OWGbaHP?u{gNOX{+UQU#+ zEzGc(oK{`IT1sR|8U+;Nwa|{B<5-|^ZakjHLQ68)H z(P`?A@6BC9FU=I@&-S;Q>&!5?K=Iz8be{-PPl?_I;{a7}qS;c8+o2B8hLV$%FGj1F zs_s#85WM(^V`k;7J_^&6xo#|dYt4B|uy?0PTBUxsZI)4a4hBT8JE!e_`Mb3=tbKWv z6{^rlCMNFn|hjEV?1j4uFqUs})-JJAwb(1Zba0lUnIkA|<)!C2?sMFKPL*S4Nu`~B*EmZ>05nFV`B9CE&1t&#gL!1$njgNS7zO0kj>FA7Yb2QehO z_u!CJu*fQjG*hn;si%ku%dk;Y(j*M8iVaIUHrRw4Ib`d)7c0S!5hK9;xG~IzmVxtB2Hsb+>)T6Y1f&E$W{|ldxXFF73P)B21FVS zM({7_BBcK_fcmyu{K&Q1gbU3C!uUn2?ghw7Oy6w40k^AC?}by{i*l=%y5k6QV+f(h zH9#a=2%~J7m6%rsk_e)}So(NKy*(~fliv4oE3w*4Bs~r^(JHYctL`Qz+<-Ex?$dRH zZ^x|W+-Mz!b-&z!NICP2@a-7gRT)S*MxjgN;&6_L+T-F=zKGND|G)(s)?v8-!@kTj zB9>c`mt=&c{{v?OVf+e*Te8J%)73W{boeL^4kzE{Mkg;beXaLaLYE6YM9hBqt%L=H zRZ<(CxZ}Sw9A=dC8W-<##HG-#!>Y*)E3w1W48G z8e8au8xZNx7T&wKU1rR)DQ>t&MT{=ijMX%Z)w>L$8Y6Wolfm;Jbya+k%RWv1(lg-g zN_n5%?bO5JLM!;nzK5flMo@xlL(j#A5v*37{Rdt9L$;Xin#ETJYQa#*SLC6H9ZZf7^#aGN6{$j^RG_Nf3ybu;p0q&(CMmpBbV(@ zkR9WPF;EMHLM-4^ypR_x;8eWA?P}e}e$*UEcl*_H=)%hr7S^Okf6N+Dz1%^;0*m;% z@Pum-pT1=xM9OlxBIN3`b#>+;I;dr$zI4K+P}8~^c!0iEI?lYp(^KzjGJQF zgu0*s>bCpI@%7qh2w7*u9U1I9(FNK=FNY_Yd*~^G);a$oCkQ9Ehn^r+=2FNBr?rc| zW($YdMK8C7({hi~(@cfX>#DdTm)TC3u#DqZChwwq&1MA)jN|LV6RvUm5!;Cn8Tiu) zxq{=v5d(prVs+%<4B>P@|IqnMZimzQ*R-BfoypVZnS+{w`dj@l24aO3p_FamWDB7r z3!%hhCEE!$Arx(M6QOZEeI@tQY$bn}LP={|o@qTdT7HEEN$cUv-;AS?6T$W$pXlg0aQ_^DZnaXeXrw_|szQ$Q#xITlzrFeVI*tnT^{#QMj&KGOdXVSLEy*;z5`OMA*s##9jx9E=Q zlt0l>My~UCHoZtl>V4u&lT*gH=pVEY&MsuYz3kuf=wH6KWWP^#LUq8uqSSvJ6Qb_z zx9jD73FMf?^?%as3$%MwtJPP+w>Isy7vhQ8b-YM%*VG+12#2vLX#-6&9e6Xc7}62( z)wTfe`Hqj>UztzD9cK#(WCa8j3?f+U7Q(8so?iY+U3JeM$_FZyF`C6f~sm z)#K+zUcDP1V{!0KlzgEZLrjfpdHDLrOx5UqRK07RgPQj1kzYYdp*}eQg-C0R+`}=* zwNp*vs!m)fW_4pQTEF~dvrAi66)ytrtM)l`ma+zwSYP_F@8@SK12CvpdkYkJ@Vp`7 z{%caI^CsRoqk9}>m{AWxczeZBa(PH&IDTW6sC536i0qwFSrmQd!Kv%OGkP5U+N0t2 zdxh|+yF0(G@*Zc2wTAgOE%!Y5vP13{hqGc5>g(Zrr@u#l^%$pR@m22@DIeX0WQA#G z^_VdceTqLG?*ReuHOvSK$jP3Lb8UQvZA;+UDiD}s{P-Pjelw|JXXf8l6iiPo8vQue zDm`m(SLbaO539gR@O!jt0Va>rcRJHlu+v*D*DqVIeDa}tdr&PquOKNk4rl$(ZC9IW zcV#HcE9I`H1UgZc#`A*<7ZSOf`*z%mbL#1n{UlFDrkBda3mGa8f*-5c3EM^#Rjk%H z;yF2WDNbJ0U9Ij#2E&$rFw!KAbznQbd26x&ocC~8E2Oww^QOOi`$bp<8LTfAjR3F+ zZ$OJYU7W2#?)@A8sNmUu(i=tn6pOyWe#&8VI7GLA@F(8Fl%AHd!k`_1k3 z_3l*}@J7TU*8ko&jSf3vtH;T@v$YJfi6CD0@h(cO|WP9 zh7M>Or1Lb0;TZ>I`&9?+%}(D5v}y$Op`FU0+7hRB=fo~0jyvMdZ^$BE*?hjkR1Wo) zu@UrG1k$k{7?9@4DWW?h0;?4enGosF8#AWKm^msFo+L9K=XvN18oBwS?QD$ex&G{S z;rC|~0`9(OIwjdxJTBcx$}(TyHm>t=Wd{{+5&iL6oVsaZS;s^726O}%JzhKp@vT0d zU*DctHPf%a4j=8$yr=uCz(vnlV>AydEZLp@9So`cZKpFL=aW2>P#Cx2n3r+9JaJGe z3(idMZ-@i#&Fu8*&hu(u{x#s(evPTrD>-DsVjEa8GA6DbM0B&Z>;6*cxZhr&t;_&xY)&m z9x17cZ5&-J^Axko^lm>1e*);ReI|)gH>p=gUnw1nKbR7THoVhXd)j+C?Ufm5f1L_y zBVonKvsAZKVAA_A0I;E-D11#YS4PKiTMXp^cjmy3mrr`1bM~&o6%?`U3%{`w^|awt zz&Z#AC5A4)P>ao>SHw38J~+3@SR+3nOv*r1VJ}O{F*;rcIn-+KL87Nc@}&&3s)&A< zOCpJ=_21Q8ssPbDju*xk!QFnH7xHwD*;we^Uzcm88a9X=ThY?rB4qRa zv@pNc^@59WlYkublccNSZkN{S&JVPqYn9r3>kj-dPPmzYa3P<4IGr{Q+OHsm2mw+u zw}LX_cu3tpw{yqs*su{|CN4Xi^>1f}U5t8q;m!U2g7xqu!0qZ zHG)IL0X9WtB2kB*{_FQ?I+OIU1FOE1bw&(c-K8!811_LqP_WF zL010A4#r^)Fyp7KR(LoSoPe|PBBQOspDdqH~jo304q;KR#15}fb32nh6Ka!A`6 zFOf7w$#Jk!EFLryk2U1HM5&iwrdXh8CK9Xuw}l9Ji@KbbEcG&Zip76dkfdIImtr9U z2UC;t`Y#0wn&k_RqzcWn>T8<3w3xv%hyLd5pma=bXlm(mScK=2p8n@iqM5Ej@F8~K zsqK4VVri9~-}@Hl*a+^I8+HjXGQ&D*B9bG}Yx;#eY z%&s7OR7HbxCuKMAg;YDb+oI7BK;+QAJY7O*Zwo?pkNtMkJMLRolv2q< zfBMMvx^rVcbKc&&FXKrM^#`N@s7pH@tje?*rnl+k17YTr2SL{|mE#$TPHB?gqb*xU zPG2$mbsA>AVipN9z`M3`?hRmh6OllqM1%e8&~lqtd2Sx{n0P@XyzGPvKQk~EuHv5Y z(Xl9JR(vB;x+MnC3h)IQ%YiVp`Dv_c4OwQ#m}BG~YV;j&!zif{9?gMHbx~zi@$Kt& zEy7Nvq}Pf2e2P}lETQi*UI}!gU1Pl7#rUHV{44tvT2hw~S{D?_(W0$Vp$SrZOPrE;UM-ksieQu9aUMtmYWU@cC1MGz)Qo9h~FCwN(FL#Ss9zExn<}` zt+lGTO&?G_3V(>wVj%Hb(oKDm2sXk)O%K2<_#w*2G?j9WU`Wh1WnhNfs&AL8;rV?D_2k_|2m_VdR<&9#G{o>NDor=O0tvR?NT|4qt?KVm zG}QLg5OEbp)!%9;VI=UH1mrLyF1jT|Yer91)ec3meR?F%Xl#A`{9La{gez2}86?z# z!w;!{IjJaLpN7csv|Y9I1kIC8S&>#N2d5JoQ{Is)BB0r2JEH z#5O3H^V6YnuJBGk+KC2c%=C{dLN=`yW;U&gyj(f*jn>9w;RN8-<*m)jMZD(2p&{pR z7&vjOpJ2YkhI+VwaONdCaG?3|7!Oa5vT~;dIc2@`lK)8oZ61!qzk<;jL&lQf${lF) z)TP2J8{kS_{Y%LNcy-GBIL+|io=vWDvXY1M@hk)CDU<;-R^a^nJUJ-g>}7R$>P8N% zfZKl`(+C`<5Z>N8KVDuqKdyKTr`6t?@t0Ns9M#vePTBn*ec>>0nHAw3zA$qD*VhxS z??1Hus#4z8nlYVKGbbi?ms~J|FX#Cok!f7!Cn&BUJC+I*Lj{WFgbTcSF7!PdVW$lE z7S8c4T(VfW284DxJcRH^^%hPB8UDEP+gg_k`2Oj-`SEagWjI_kX}Ah+;i|qp_!m&9 zhi9Fd5yWBkUpy133mH=9gkoz7cjWDK7}Qv_DGbXbTQ!Hv@i~7ZlF&&6pL@#7r+q3l zD4d(2pl+~FE=i+Hs8ppe%$+NJgA~*C%92jCgI{iclgiURNnby+z{~ReRS(sSr(D5k zVtJ;j$$Ot%bfl3*s!xWA1x`}x3TTOQl`?roagwbBk?^2Q6Akv#zxbrc%HvU#8br*^ zuv0aZeJ{ZyPpA~5FbuJ%3P&)c{3D8wnp`EK-ufsDIml1NB`r1dmu$cM=CSi`zbAj) zj?RQ280GEQVD?On?6MuHX3mqQiZD1S;k*j)P z%1KRjazl*&31H@BX2s*&%@R4h*Pu_PMzlctP-%9qi`YB1=#GIIB z^|~vTBInNws`_9K@)(EVSC(bprMaD8>+bmlr!8*^lX_Y#hZ~!?VDN2t*??XRI_q=OU|h9v0b_)t!gci^J%o*^9*smeY{Y=+PVN zJpiD6EB~AIhcsiGa9cy;Bkg6|TSzo-@Y_T>A~XTh>Ijex?ziBGEsXatpw37ARpK&p zZ)#oKCm}+A3t@}&tg|0tT@TxN9V;Z-onYg(WldfD`J3UO+_<4T-uY+&!KYZ&E|YGvlmRqbk+jd*u$ESlW{0p zE`R#ptR10WD2=a9lCY%A2+RicU)NEsK#8AE^|BPtRUtswvGc7#EA&oD{~-g7$ooh- zNvQG=Nn>4NQ%Y=I=~gUTac2qSmq3c~6#QW#ai1~4f)iEbLTpd+tA_Ofid#7@#&~w4 zHgnzy;~UF#hvsZMc1M<~k~*?%rQ-c5#h%3up8<^bp@{wpan~o?&N_=# zzKShq>xLb8&Zi-D7b>SJ>!40t5y&tm>7Z{U$64Y7Yhio7Yke8W1@~>OP|(m!UVb+n zK)z4Sb(ZeQz(JnYL2;nI+fa3_C`F+XLX}TdmD!E(6dsWwChYaVy|T5Er1wUK zlEmO^ijy-YhwJ;h-gfU#g`hiSymQRxS zc&$-uF<;OQez;!ZZTG?vgk5|TZE5*xr}*M<5x@P}#cw~WL#U<27ZhC(^)}7ADlA9G zuXwbUiH8)r_PAwr)_-2i<6Pwgx?It@9AJsnv8HuqQE@W57-_#xtw5`-FePYenfIyr zL^3f6eHJx86PRFE z6|Zeue&5b{rdRixo|n?!J-FVka1RPR&N4D2bqwI#^CW&lo0m$;M@e)S%!AE6QQHd1 z0znIeYvb{X_eb|F_?4ZiaGh1|)%17byCv@VAY~EdUZ-nkbS!~Eeu0@!F?B2Ur)53> zH@}Nv%)#YdFIaKmlMF~9~DfeebKgUwDYlX5$pp?zr zwWxMio>Kecs2S@zk&cj)&K~CKk>Yaw#I;lAxEk!?@%`1@^LRWA*p>j$XWD#JfAo09 zvUaY0zp~HCa26rnN#}d#>3iz|)*Vz@RXN=ufi)j&ifBI!+nVl4qpr5S5deCMeWHzb zh_O3JiHHRXnq8t96QHyb+EzS!OffN48w&|sPK;#aIeVSFz{HA{iG8w7GSfSntIv;C z3VNChWOh`qiKfIrL8m&E2y80+# zJYBFT8P;n!i+?3KZ%;aV@c*OyUt9lQB_^jw#ha!04=l1-8-1QMPLKWj4;gX5nwpRa z6I`jDz9q~K-m|$;5?klaxwvOHeV@eRVT+B#za5)2_CHm)d%+-GXLZP5@%yz)LmFREiZ zrL(OFpe!N;!F`~=e5 z;VtIDrh%;A7E(#89*sX++C2ykV~@+Gn%57ER$^4r*9}u@?0i0nG??DiQ{)|A$n#1# z9|`Z8HQetO4+RZg$n_sM>8y?~RZ`5K$Eiw1@h_$<7y4)g=&z0siM#maoIfm>O>Szj z?L*2nkqW%))ZfLfj}JwuUV2%1y^E-ff*%VtNdBB4V)!h^{DE4{4UcbCtJLIqkiNEE zALQU1JH$%co~g`op|9ozY5siVTFdH@FVk)O71m!eqj+Z{nI0$`Lv%hU`@Ht^vS1m)-}79fX6L) zE^n}K{I{k>(C%XumDT0X(xhEKA%jkAmU`OGRcf-|EBg9+GMtNE5U4c{c=+Sg00KI?Jpgs=<_sDUK@zFDnfY_Y-uT)cpgV&t#@I-4agRDX^ww9QR6Vd-aN5HgpRg7( zxv8q1&scP@zcCwI$e(LO6cKaogA0ju3n5;kzVo4uw$=UtP7nJKx})`6+I$08J9@9z z9dwqqAg{xHvsfk$(!HcDI`*45pvEuJ!_wN(ZulU%z&pX@9 z`enm5!VP3nzhFc-?3fTGD-D3-5T4fwcgDdYuohtaGxC?be!g+eUZoz%vZG6@9KS_< zpc^lcwZlg(lfO}0=bPT|c2P3bM%7b_u6@mG&+@1j!W>keQv)yij?s;EcIycjYY%|1 z`h*-Rfr{!{3q;UQJn0V+D}5sNelm+i=0iEJ{+Cd>6hz0IV@UE~ZxP}(ZyAwAC~nKs zQzbDcR2IS8S)T;!UY7t4Uk- zTF##K{$Vqo;Y}%w6-z66SWR}tWJcR7nPA{MIVR(L;Dh6K1dGoG=!nVE!&`%5h|AN_ zzj(&~^%1w;ikM@QnVpuOj`-lZUIZLbJ9&|FQk|*vgMmOK9*{k-H0i@}fymUQPww(_ zocz00hD&YqBdS(|nwkh?yNW_&u{o#3xvOlZ?my}~=HxE)y$4acKYtt&k8<8rx6i2% z-=E%Edz;plafu9g`DY9=-200?F1F$Dt|OBx7nOly;V z{Pa7tAtF9|-CR~dAhejkH8<;DGXjjkDU{PN&wTSV=_r35KBv8^Z7>hxWI}e-TKVhn zfxl&I$IO+*x1Piyc8EpIae9L_D{Wv*dXMU9oN3pYhly_TH}b4)IEFKq(MkPy@N~?) zg837dbMlO)?dYHP`-Fz*gX%c?y#a-z)wJ>23j|Ik8JPi?3axw}`2EGV7YN@s>uuzy zUt9W$w+Y;>&R^LF1^A}52TVfGdjq_lKOIkRHVwM-*mf%giqQHh(CIrZ2xHX*{7lJX zRg@$f+Hb&<+C*GBC~*8wKFCOB7Gc_|L3mwBTgGU(v#twyV}_7PX^C&T%9baOYB*b3 zyztb{x`f$WS^B(42tdegCMWlgsm^d{i|BXy;Ev8ogR@d|c9Cc2)I^DS3d{M?iwNF_ z{m)nDKgcFDZ)~c2%qLR$)qzBrg`x(Zrz9CBxnP5KB7FAWhaJ!aRD`xSzmNvMo1bIA zB;>j4{+%`n>CKwMYq+277I7$myB2QmBDKQa6LWl@)Z)8Gh5mZp4}n1TOm-OG-Bj3@e-hcdF>?YDjcfmc}S`CpZ$ve zHGcl*%C7cpp~X*r1cVF?L)d3)8cpt)r4xD0@lF^I@cwS!u;b(f{F_Ih8|CEG25gxq znqIE&GMdSJD7CFbwIOgwQ*MZ`vT9c>8qlS8EC5%!5oI(?gh#+pcL_2Ik5hwRyvi{08%DMj?t< z1J^jF`p%v#8-=urr_7C-Ty5C8E7bw|@20PPsk46Yc#*n$N(WwOSafdPebro) zoR?T?d)KQM57_CKH-TCTU)WdL)hV3)-2T}JOb8o05Z<||=+Evid;b%xS2|L3VpT|y z(h5w#@tr#!!=_T>Vb)C%1t~Yu=~4Rd!$L)AQ)0&?T=gs z{+fF+Ga}`H2cw?`)(#FP16AP?BEMXnF+XsS zGE9$$nYG;Mn8nmlU;_ime;U(b1AWVJqNqRe9j}lrBe4N$;=PLque`ikdbAk}Q5;~X z$BL%(Uh3~Tn!Q6OA5NAkC5y~xTrLt6OJ|QQ>}f=lvXr!j{(?0aC6@Rg`4}8&cEiKQRlC@ zA8CP`dwj^}ogh30G$n|43!Rm!&kVRhTRCk9A@d&D4UAHO!w zxn~5oD7zS^>K)E~^Q})TlgH@=zYwpSlwH(dSR-psk-u(q*c1BlG2bhrnl1*z;nsPv0LX{;>MA}y{`lQx7P&ix8EegUE zW6K{_X@boU?67I~poKLrQ`d$(ZL(CPR63lD>jCH1q?<`z1y3SY987h-*S)k0F~JWY zjncV%u8OGAZoFPG8YrLp`6!%1`=&0kqWo84e4svp*r0RqjgNFHDt<{Hntr7N6@i1Sw*NyLk+<-Dxl+ zKRkCl3rd&$hg~JZx-~Z!HsEWC9=V82X_2}prT&Bf+bj9*``d{w(tNcfHL@(tX$zq|T7%Y7@s z{OiL=2L7GtNQx?t4|*YlQcg_nO?&+c8f}KwJ_ZdlQ5B#DmohKZM5)S(7BGIo+_}H} zu6DU&MUKIzoap@(2uS|)NBOFSElLI?s#I&-@D0Sz`sMmhkUp2utf0bGzOutHm65{; zXQ8rky%8XH|4g**iFls8F>SYk?fb#y0q$s+AL%vxQf+pN5%#$R%b`Te5{A+5OrxqQ zi9af7M!z$UelO34q|DcWYyG@*v?otS4wen|guo1C*4~y^4hGtjjsM@tQ8pYuX}-=M zj<2FVxn5Thice~ZXv)p#Q$whO)}hpfy;LJWByCj5tbIG&VTB^beyw5UD94Hun>1+K z+8fkd)v{sY#PIr+v$Op=k7*SCHA7=nMFOc(2RLC*GGn?k3*@EDc@2G7dD^7eIN5YZ zk_fQaEPNlrLwo8Z+yNV%2E$G}p|#nCPgj@Di!8Uil7}QI&)1$G1FV8O9JWrL9kcYt z`ZiI>PxVY&CH}5^WyV>F(MQ2=6HIYKJP2T%N3s@oQ7a+#mv}Fe;Z)9(+%BR(KI=^j z0JjE>$6p%$IXQ3l@yE*o0pu;yW}zp?s!o@R38ybfs{t4^6s0(E(lu#AwpnYpU7*%L zz!7b{+xS3-rvJrk{cMs;=7i(@Sy!xN5R0>M3)l4)OU;cZTFOPSW04`>x>|D6boSOj ze79)pye%(@K++OnpPLHGDY+}fj~^h<*9xf+e*o_#LTO+ZrgWP#q?nkY)6Vu{G)f9>)vc%BkaRPzO|nHsw3r&2hfA*dl^S}Krk1Ws6+5K4MND(GG!6s9>nAn& zW0@DZeOJF<6KXpV()`V`jx@53I8rXdvmcWk%~MwglI?+Qj*26V`^l?Clz%}y=vD{l zn_{W|=5>=DXPKImbSfRs8~4>4V-1}M4gTh9KCKq%I03UiuWmceQaS+H+fCJ(Gw62gs>(rI_C=-v2y;?2X1X9$%AO&|1QHw{baITMyn|upK->yXMui^6P|pBPLY4A z=IGq9`0&+;W);IcyC+!nyp4UluJE7sIGrLNp}(CQ76T8$S}*ha_OdI^ve)Agug)2L zEVpl;;#}}GT`DTJ1k8xLVLeEc+G~aAWmQ_#I&`45+RA}TcB?ctsT&n|!nDefgFR=i z$`wT9wYiufN)WoU{`aCSAUoR}i}s*U!rF!hmD_y1L0Gaz2-gz>bN=_CKLIH;vx>kj zzaB7G>0Qs=aC!cDk82A*A`>fkXJ=>*|^QA zq=OSKds!0<9W~zKu|6SBeedD1n6Q&yQL{clLH+A{!;fNR z?@3u(J?ej#*IX3l2!)TOo?6~vO>JWApA&ytN(;cm9fVp4)meM$RY~kJI5J7N{n^6q z&KaY@5Fy92MUz-;P4suiQf~WQCnNS@21tIp%9{`b?&<07?DbQ@$(KFMn`@Eu@6VZ}=z&4C|O0?8K_iIh>ilsyGo%2LuH z9#!fhcS(h#y}SrA63Z$dj0G8t^eri!-v=dS%cTm6sHfxXfAo@>EvTuTQIx9UG2SOz zaPg|BoImuu$cXxXSbGb&I+8ANbfP4<6WlEXcMtCF9^Bo12myk-yE_CfTpSYI-R0u$ z?(lDhneUs~-M9bU_uj7)Zk?)Ar>ahORdpTfGb(z8OzeroAPZ!D$Gny9J=G?Uyw)8s zrve`@FCPmP9}fl}$E(aQ5ZwWbR~m_ys)Nk$ysD0}`iMzh#*`b3G zFZffXPoGPFQ4W?rb8!xd~+z z&3(po8Orxo1$DnY+Zr*LCc#|ru*ta?TGL6PqzMM%?dE(iuWiiO(7=)Izswe0yJ3Cz z1r@J}ZB+Y&jVCluWCasRq(q@np&i*l&;l9rw!=E_z_?tLriz%p5fb~Q@5HNZEIi|4u#*d{H}oK7g_uD>OU%;S*#+nJxJ34CPG(=(D8(oWrr3 zBE%p>IJ=chYD_BBbOTV&#ogmtR&9%X>cN)Ru+*?rwA2`=nhu}g-w{7>X!ST<^;ied z)kQ~!ydi_M_XHVsYO3t+n}&^*>wYlYa2T84=cqHVHJ*x3A1qgBb;9sqYjTiYOq-u? z+?W^8pYiVlP0bI@6a%8E9`L+}hy7`I@;U~&Uj8U8M0vW|EFCi}RzIG_Sj}LKG{-3;JJ+*~mgT8tij%iJT#5GgHh4l`KQ-G1 z)L8Sb9S%Z_tjpVu+3q%Sd20vqHag}mb#0bUg2lDZVv-wSqwdCZxSh9zR4_~vi|3+i z2CvVn?cE%k*hy2T97$5Xt9z|^Sgb)oWT)BscAV^;yrOTqd9+jVX!E1D+3K3q8@abd zOk;W9B5ZFul;c&-HraP&H=bQ@pPODV@>DimGzVRh=d-5Y*`!$F#atYRCNP|Cz&S;g zRn#YKXq<|#r=K9>^Yg;PB=?ot+qN!RF7ohwxTBLgd)T>$FC~c^3zE~5{wio>9FflL zo}|Wf*X$nzWoxeysUs<4sab~ar_=Z|v3@2fxMRyy|J5sLm({zMdZrs27)U&MIM?Y4 zH`ikeAtFNiR-LYGbe`7Go_|ezL-qEa6GQ~m zSg%#Gnzx^S$q+t^50lFDIN@D#5e;e`3`GUIqBtpY#maVezwlq3XgFkbT5}fOswCKZRG!j&zWL_d3!fIgJAva84pV8d z@)>RY^KhZ!2BiFvw=+4OU{*wKwSDz|-!1$kJF*Ki#)F3KC7maEmBiJizv337s3~u; z9QEJSVE-F&6+RxT7rf_QS0Dz-An=*PtU~FH?fQotFV7Gx+@<1VXPbwH8l#dOM<-R^q)bnVc63Y%0qAj+RWF@^IJ-T<$2q+`0ybjo_Ooa}7&++lze{+vB`1 zNj$UcJzB-Y$Bv`u;?tfgp2(`w(3P{N6I%~wX zvc{5HYGy?nIpd~(8`M`;Qy*Wjhmz-vVON5s-79XX%2WE?-BYqYXiG5PY5<7ED#I%{T%Xn5IbzWL9<>lhpj6EiyB*@m(C(gij0F&G<7A#BPyv?EMlJGz8ROGRj{y&;L(051H6-mU z+rgOUk4?iHztXMcZ<_|Bro5?q<-TWofSeb{I2kIp$-7N#9m5(DL;m#YtTN_du3dPi zpuHln)upRzU$;ObeZ`qbjqzkC;Q|;+^qIG$R_#x!%`2Q+rEXW^r4{zj&1A$lC&6bo z0Bb)jdYgIE5kxa+jrqYyqY4_*K&5BrmfAgEG_GmbK*Rd$NS7S9u?m?(A>sYDWU}E^ zlKR1C*`b{aYrl_JcJS+s@t5}Yyrtl;r+r%T6km*5u2g)vo5t^nTm84KXJ6K{cJ)$z z5$6!w{b4lZ9j~1I2=d7x8DN+9Do0?Uyfweb!qq~p!7F_@N;fg)4^a+F+wZ0?n`>}{ zC7|1XpPg(MNVyl#SElU_R$$1U%icS=n{LN=yawI$K4hv6EwU;cw3)5&qMP5J z+HZfrzh1O2+)Za84|AFn>7`}nb!+A^t1zD)N;GJPy0f+Rr}{ z&vp&}z#xW#7lg<8{*QXkbLgK{*=-2ymOlSU=b7LidOgG|Nz|$|m&z;1th)LBP?qQi z{M%m`M6Pu2 z7k&m0p2;Um^4o6|8RmWK+B+mP0Y}c@<_u2Gq}C$Do{HCKn=#0g_(3I-Pq^)~jMZ)# zByrqe=NTt)`2X6|ROp&lRWPWQb*96!V+HrrvEoMjlZIHrfvhCX^~G+B zEHvwd)P}JbBE03bjAxw(*^E8XE|@RA6_uAx*j5dsSzz=XRIzoAYLz(P5sJ=?Fqv|m zF5IsSwq5?)uB}TAk(V^!)TTpCT zxix|Kf^8GKek41i3?4t{oUv}A?DZIle`%HNe? z-53TSb&ztkiCZ+aZ$*&uZt})nr#k6F*3(+$LlrSAJyy~sM|8%?h!(otgx}FwY9l1% zilj%I`RcJ>b5-pSqb7-K9q|qx(tfX zq<_~X(1jb`c38(yXMykrld(zvL*+m;{*CUM_bHJ`!Gxd}+3CGK9d@{PnUwXqt2E1tl@9Q;_1*>gqDHAq%Ey@Ky^pK|L+MwwUEf?Ej(z%l z`3`=OTn-anenE15*Z8e#sSww^>%iv2aRK^m`YUqR6gSjgr8`!8xjXsjrP#D6Iy-RM zYM#2iCxY}%GMm--eg`&7MlU>=+>{=|2M>R)$L~tUDUn}B?egbY_H8x6wH^MFr_RbGi^RtIr39EV%9n+dUm(l&R?>)AuPbHQV(@!|XRW zeuE}l!G{WbLpH(-zdxA#GThR<`bEbn8e%8%dY9NK3Rz0H_%gs=;j~skkiabGq4nxYF&3X+Ce+zbSPk3i7xqo_N-t;2 z!tg1$n8en?gdcvYndwrU+!D-k_fjTtagLVVwO2ut`3nGhdgTE6DV zuKsCE0pk%&YG0)*)n9&PyMnN?3Y-IOEAZcrYk`;S0b}ebk62hkyRs!lHygo(ZuC-B zZiOj}Z1X}p?sw!06*!cfwMh&$BR`L`C_?mdtpeoOQa?f_(4Fkp!dsQ^g!RZ;-gN2H zUZu>}EW54*+vmd#@wptnZLR@Er)bo|$l7P)PBo>uyOh(EL* zDyNl}9?~(^S$GqV5q~%|`1qiH^If0CmR+tpF0eVG$)1Ypx3=Hg1=_NBd1QAS_N2qd zH|>7uY@|+NUcYSVUv$0#2cI&wsJ86CsI9{#y+`fn|ywc&s3u*g>AyU#mWTgnYNiuRxBE83} z_7kz9mGVivgZYU6^-jE{%|KV{jLG?`#n_gir~C|^(S0S{%==rH^>$|OXn+A62HU}O z==Ynm(ZI_H7rBUZ{>FHjJd(>3G%o3;ZiIk`pYM%(9O5Dxf?J0oITODZmoc7$=D}i{ z<-_dV@X6DmtH~e)^_}(a^WkOP2nuUm85h}0!l^-wLN(h=E=so1($VH_Z(D4*LUAvX zcf$+QR4R|EaO#@>G-e5GelpSCw7!O^pQm9@tIX5Pj(m5_PcdG!*wVTv6 z<}%8t#d$8HqGWQdw_`lLKK2Y&c7nX_K9ED&3B6J#Mv=lj+UR<5XYwq_T3}0`O z>DRyQwr%Iy-(0M1Y%ur1M@`$fUZ``;Z7P8q^!f(6SXm!yk#xQ%p|!ZUJMQN2)oq1z z?G=NfYvRdS%Mq4M`8wjzCL%0&up>8V(K+)5XPB;t``;*X=V^&8z{?wl6g`aQ$X)!7 zAp|F$!d_ILOs>(p`KhqaG{!M01za?*aiXX;vCJ?YRZrMGTfd?v0AaOad^!Vc-aefH zNc=;s=a={A5Uv|QH8XR-$2`!3nV;;W`Ep!td0=TYAt(Pn0HC3`|0g?XvH9Rg8`_=6 z$na?3wF=Jf98@yNYP`u>c)7RQVR~c7kr);|SKKid?GSX?Llkp~2ELHRt}@iSu!~bPc@2B`z6>@@rtSio-KdTV`KH zr@O9)MP6&c)5*dYlQ-X&j`L7kw19+`(8-l%#J;~F8;}e6q;vJ|$1;btZ!0>GZB`X= z5Xsw@It%2l)X=qQOVa?)5{&J3` zWMM~AjcEQwUhxEH;RL5EzOWW1Um940aRI$T zKa!4PJ!D$=4cxPDP;0qYV>_2XxG8(!Iyc>WJa#T$6aEf6uf-H*iH2cb2)wRkAX4iC z%Np^*2ic0CjT)!{TNAqPG&kLazY|73&-V6N_7tux5QYQ1JaOrviAv=L$am#)|yY*yn$S#pAAPCH$#>K=GYD<3}J#*gud1W6uZ;B+2^+ z44~~nU+HSgE7bga?T{R7Gl5$;T_^KfBn zF!odIM2HerP6EXJ_f@?vjT1631}N8Z_k$o=Xp+oJaMW>V3+RO^Fd z2Z>QTIg=>&=|tl;jjB%OLnPCZk@|oxxdOc#8u-3mhgr(Snhg{(pUii9ChS=&f%#c| z`B?!gJp@_1M+!oP%;ZwJtphQ-g@R$~F1A2^R$|31qh7RnwK?bo$$Ea)VzW!ES^5{G z7B+;1nIuYTB0gjW%fkr!OGkv6b_*H*VN*;HGDZ&mJw85*r-6}ZP22>Vuy7$kR>K}$6^3vQ@!k?30-ma zCB6SaM}K{Lu@08~_6WFyqSd^D#E5+zAv#op($`;gw~H z+}$MUI9TmOBgzqZz}i>#bWXPH+On$j+ER4tjzxL=_Zj*ZlutK01Gv!>cTYFk54h1& zcTYDu6u8kQUf8tn8;4E3rd+cLW|Z(x?C_NOqC`xL)_n~ItWQ}puPymcHTe7ZPTjjM zqKmFAM@_wwPncTu(8B;^Y4*F)vcEl!8wkp60R$v z?NO6tp*hvpIXR9s#(hqNTBEb`s-yMfPZPe1#T1~zDBjP2Na~xUn7MswPBrY@K-8pNNwlbX^0%nTJUcOSkR{NqdZ1fG zPUd(h5v^}A?B0t{1Zs-ysw14JnN0XOi0yKapUjv{v;_g(iTVkX3iOZ_I!f!Dkakn9 z`jhK7BNZN?PhfALJrQI=e78OEQ(_j=)r#O>&{N`h=@v}G zcC}ioi`*1`PWS)?{{sFZF{#jg`!QW$+tXes%OsHn5Af`Nr8Z`lS z0J8WiXY-;GAnbY|Bw(ghL>Aq(0rG+cD{s6ZY6*FFX zoK2f$C47aw^rm=S6Prv=! zT9)9bj@uQE(j{jp@~a3m-c!Oy^w;0beBTK^X9MN4e@x?Oz(L zoDpZQc8_f7dc$6?H12AX8@KRE)Nz$Y(Rc-&^I`mVAOYUEjMwVM3)?zx12ID9eJ%;1 zqirI3TFf~7ihmthP)?IFY@K*s0jpfw3<2T!u_>W>E}%V;RjY)yKir0){IA2c%eYYR~S8B3_Zv}r*q-Wd6~m;7Aoycs}<;f z$Qwm7LfT#&Ce)$V>yL%;1|&e zV*eEy>pKw4GR%ze920Dpn&R@mEKuxd?`KTbglH3tW<^eUseLVy!B9qHgQMs#5_sjjR84V=6Sn*JAz(E*Ay2#GFKTS_py;ah3B``7V zy^MRCTNR()1IBl)h-X8gtwR{&L&?87;r<3PNlPC-X^EWZLobbGb-e!h<3}KZnPjLA z>*3cGmQr=Q0Wv_Dz#<{e6We?Zh~2IWN!$$(pZ&y9@)P7xP2VtxETX?xJt!D5I2@7h zzT3Mw8M{8;4Y>6=oJU}Ly#AN3xMvdO;~&oALP6`XqU8n}fQ541b(hpl2k^Sxi$#3D zvOiIJ`mnZey_?n5u(U8S(=*o>ojer~QI*)?d9Ui7*)-~BavB14-TjC$04vg?^k|Uc zBKMP4omD&Ph(7wrMX=tv?>cX_ohm2sC=BW{x1dieAU z8{9oXvdyq4XECdz#eFUBr3(X&U2oJ)wbE5nzj44QE-r=5)g%m$F}rzb@y~@d5;39{ zUWvv%Rjf0f01tyU5sMkWQG_kr8$bPSj!BX{=enrif7nW3i zIU??wV+YdiXf|dgQM~-ZVkS31kSYJ6omgH0x2N|ekB1GEPUD*YSnS@{?8F#%)fCHB zoyMEIbT#X|6`ID($&rEO&%1E!LSz04QkU0IIvz*Y2zS3*p=hPg7+E-t3& z=%l0=VF$rpXJDy&xDY0(WAHIAO(7V2g)Id1IB?yTbh|eTf9DEA)pWCB@>=`FB&pEO zSI}LmiV^Y2ALPWM%z(;{5iFH*eZo-r$t3!6`(A%9SPAKrg$>W&msVWEyBd0XD*^?& z*%GVs;chsjNp4ghvx?3C;-)^Py;A<=3v8XM=uaW73ARnZmP~HpHdOb~)8OVN4xoqUOo?)hRS)#5&{3tp$YoOtR}uN?s1{ z*t2~lv2Z_6ZhamV4oF}%W45i5+vpyiE0pBC+xlWBvQHYWK7dy!jbEL{t4DQbg6{FY z_0Z?z<<2#dc_4lTG_Sn7dW91JobtD^)DKVZE9Rxbz;VXFz?tv!x6L*V(^X?&P43Xu z`cW*(M;M>$e^H1$Wf9;q_kP_EZ}VHfI{ss&?}vxVHtc(Fj~Gi9iBWlMeO0R{ykA6B zr^sv$aqo<1L$-Fd5{p5Bv4AkAMK|z+D>OJ~ARv5e(arXm#h_}?Kg9QoK<9D=VYU#9 zfkmr7hccC;sUI9BRI7gocsgQ)2%BNNk}%tl#Q@yvFL`d!Z4m&M2>i&eHA(?_L;n4* z3Bm0hyYTDe?R(6L}46UA~h&ey4a#1I}Ofp zR#za^jb&WviO#rbI-DbX1e~L~u|O)&pVT&efmHMvSz{EI@rgeO71$G46pW!`?^JdB zbC}J@<^vrntt4c{^137j6oQ4D0?M2o-9?)MV4_XS6T)Io9s^tg{7Usrfl0n2!rh~* zhzTP6N|qG?r>Ef|s6n*`lz^Wz#{qOho^eZN&T*w7|N8wOnO5A`JCt@1GVC2-eDzS> z{`IAx%mcn3nFq@inO4CT1r6`9D>+NYbJ1AFxe-~$Era14fnjS^TNKPOU{`7w3V_{M z3~Ji^_uN?4K0XnuoHG`&{i}aSJC(M0jbf9&Dt_0?x=4ChSmL&~#xdu1|E3Y>~*=gpn;>as+N9-|^2e&F{6PRX0wDXksw78Zp%E zmQo9vUh%EXtd1_n9L)NmqdH<838P5WZng(`Me35Z5~S5DUJ3JKbsu+7&6z%SBoZ6y zS7vd>7PUKi#TFt>ai}nk_i7HEhUd?`$g>w@&+6eCH-$?kk=jIHH^d!$E5`HDkV%^7 zSQ|0#wG72c)K!84jmYwM(-cV{hHxr*hUusk3TwhKu7r`CaOC36uf64hQ+ZY$Gg=O< z{oKLYo_I*@rYbGIMPEgKHn3|wu6WVq6b`j5LGX;rOZ3rJEA#+L`8u_k}HJoBfGYoc+ZyCp!>mNzzGI(Q!Rko2@#&>A%N~)u&jK-!_ z_ccG&AM;!4dBbf-a<0Ak^%^8%Hh%SjomPF&eB@T*Hk@At`cf*D{7|Vj#0tsq# zAf5|3ELZ)R;hxX~mwx44So_#`--eWAa=RD*f?}WTQbTl$mR8H;KIF-9hugc)d28=( z>^9nYv-}pMdY2N9F?oI#Jm>Af<~_j5q2_o6-tCjwoYEW9I$FNj*<0P-H>a29FzPZj zVx=B_+Dxul@vR}=9+z%b=ErGYpt2vhb_*~`WcF@1glwgS!>1*JzM*p(U6T-A8!U_h z`*puawokGPt2R*{@fa@2Lc@dJppQdd@!<09Oiua`gkrX1lqs1geVoziO?o5WX5mSK z8JX?&U@`|-J2|CB>57_k8wwg{lx$aJUMc$GZPSqa2?W@Q(G{*r^yNs7i)ql-F?JS; zBzDR#%PHy}1xXc2u8Wlt$ZyBCkOuPJH%Zi2c~2;4-K6tNnQuEI!!J}E8fj4;v^;OR zcblEEpQt~;S3&6TOlED0FnJ!Af)RKgeej31VNSJiIb+X0ZO59r<^*ec(zpG{*u4aO z)MEYUkP3C%cUiMJg+ZP_^A@6dST-mP(N3Fh;h;PHfn_p^ev$`MA4;@Sb=M#S_cf+|fJmomuz>~28~ujtr`4qU11BO- zfk?kZh@o>FK&Qdfr;eu2R$?$JAlDb>G$`TwJfjG#OAJO;GMyo>Nc4qw4f@&MF&N?C zVKS9n2Zxq^dKz&L9&yh zG^!%ZjbYL>AcVPFL?bArMy!)#Ke}S@MO20Ku0e@CBJkhEpk#~)b54+_-q`GE{brbC zM+B+_hd2Vg#jOTPX)t%IBs)u#dlGWfLe%yn=!*y5+$k3u^qZxJr~%z70lMV~bc+M% zR`pZ2$_^Z>m3Ky=F+uR(LewlH>5GLJj0#Ehl^YEDGhZDufM&`B);j29)I`ptotrv=LMak+i#MQc)uxRLj8={dSOGW+pn*|_5IEiS$M{N z%jixe6;<|Zq*yPRl6X^02&n)SYw_N?{x<&JNoW`yN!%}u(g(G4gH33Qan=*^22 z>R&v^*)q6q+@k71Vp8Sz{bf&Vtfrt3nh>gkk;2~V!jPB5adKaR3{ew4V!p`}7WScx zk^8b^@C+&o`&5%;|C|M~_IoJklo+5xGyx2t0u=-F+fWoU%r}$%AukmOvwy-%3j4$w zpo*f~FcncaaA@&+uQ5PNHif)&R~OEl+K;@+ZJz!ZDd)0sVA^*ua;r#`{gWj|?gyts zvr;z&T?c@h3J{{s{uu)_Gy#Cs0Qd<2`zYwnfp*n_cBACJi~}W!Aulz0DYg{H;@>;! zmm=y0U2?i7eAtd{?5EacAXl3##V?lK#H7|onr zg23?hR$j&NptBXX*;rqR!TF{v7vX3)k>`g2d(7~XS#yN&ek0maLH=` zc^COT#X{n9D5}MP63E{}QRxi)ClgCXc;Iy6s*pyq+>Wm)bz170dm`f+h>(dN;5V@D z)WEh3gz(lxe9J4j+SKAkTYmpXz^TJe+aY@*%NVpadGmHBs!O5Bk_hM$?PQSKIM` z)j6o?H) zs70SrJJ=rD!sf{EDr|NI0(#5Tf11DxLlv7s{TG=e_V) z)Qs=8ls*X`b=+Dv9Pwi^20Q;X(Mf=*=W{HG+vGu#*P~%STIhlJQ0h?`KUv|S4yG|B zjrG06bD*~TG9{Rj@si~pYFe~1iv--Xsd2iw^8eI-qBouuTd!ykYA~Z(ORi?27$-yM zxTR|_--2$R<3;pE*d_omqSAXtn_-Ra&R|;iP+bXs5zrDAv8r^jZ*X_v0#Bw4-nvr; zmPpnn24GeBswnJx`|-}g@cWh0OA(T2t);n%qIY6tpZTe|!BBh_s$6{Gs z$aUsoJs7!cDmZ?fAoQiF3Rm9-s()Qgj+^AJHo=N|1&)3VhwiB!2~jjL?nhbo;}!gVM;Z zoj@sXKdQh!UCiW|)SS1W`Od-XJ}LRAyR;lIR2b4fq#qqW?qWYg8UIWA@diJo zrPqo559vqrM@kjdc{kAnWb| zU{%FV<|iF)-gpD|Or?@Cx9l%%fKp@>8BmG@$hUPs(cW%c@5-|tAPW}dG7+N|8#3i7 zUhgQK1{TBEye{nJxo)f2Tw_#(o_`a`m?}BOo~#-vaU6Cj{obB3*)=+?e_@__K(O`V z;XN6;Q^B`A#bYw12z$TsO9%YUz`QbFw1vVwRY$LQTr*q6*gEZPVZDQ#Y-h`QtajgN z=k~cfd&pk22lOSw5z3-etEevA0o05!@O2Wk`mg%KQI}QNF2rqEp{n_z95%w7LrmLA zncG3T5!ypASC~^ zEgf)MW;We>*`DZPyXS>2N6``#Bwi)&yJovXX?1&Mu{*RP(&06RCNv{C*$#I#C?kla zhZT0uryfh2#_A~S#dLijd^hYO287MjcRXJuxp{DRBI-1#0z=N3prHo208{Qr`6<&o z=i7$_oySI1+j7x+@3gv+AEEYy4|_Riq6)MXkLBoh@tV7F{12c6nT?!JOPfjjL#MTK zWY<~viX|?v9(dX~-UV!p{dz;&;Vo>@_*Z$75J(fNAheuUp!L|-n@0GYFR(R)gD`jdKx7AsrH1KvsnecoKr?@8OmE!&XcvEwJ@vQv#|9tIY@VQQW*b~kSgWC zK!I^+(E(~dmXa2_QD_3RwYZ*hlZGR`N0YZu#a*Cy0h>~>QU;f8M})A-)UTieMW+&u z7IC9cEKSyaEIG}^4Q2|2JV!r}a454ZgLze?^NLz_{_ba9^C~brv{7i%UGxwZGY45Z zv)M_B-ekc@O{H2koZ06DO6wD8Zu9ZAOIZMEC;!W2K~MEpwQT0H*nX_ia~U3>NNb{G z1ncn(AU)V)9E!zVoWg8Y^!h$SuS4HMPsxZdB+xiCt?bo($;d1RMFPN&r$KHdcmN{@1*v5VFZLRRvJ||~ z&1>x8JTH-pUKB_K#+wc>g=`AoHdT%f*K-u~z{Z+=UhG13E(5zM6)zNacnQw<#PnlfVh>nM@Swl)L_V~_t%rvO_v3%@ z*f32DU`@Cwv47K9)mAez(ZU@!@E2#!IMff~RMBWEnm8EOj^ck{PCwT9JLhJu1QxXa^1(qh4n@7- zJHgBr?pk9MmP%9V|G<2s%%We#2lo!NCIu-p4oSMh>kW}pxJLzw>Xhm9t2(QrFte*4 zdVr$F;v{sQSb#`u2EJ^G%csE)lq9ME(bOmA=#Pt)Q>2~<3W)${ z0?eobAGnvB2GBFOS{M-K^+Xmz}a^g`(i)-FvE_ zbs0tejzaa)m+OI4{!F-zxFuWChFwLc#Q>>aj$;5V72lID_1QL^1r=|_jY3#Xq`CCZ z9B{XcLnbGETsjMSSf7Ta4={*q=%Y~xmM2v*{cOhDZ)zuGhvUFly$}pL3ktoUjY5(% zsS*HQsP}}7s6RgO5~4Klq4ZdWSATpf#RFgvs24LkPkb1GLYWKQIOLO`yEwPL!P9z! zNlew0rMeynE%9IQCM(yfvtY8gclj3UvUik!cPE34O0npLC%(db-RzE1sZ&1sEs`OW zJcGW0C6m6v2xne8y;)&4qGRM!nAVrZJHlQkswUhsw{#yGY>lq?#+W7C-RrvN z2m8y{^`)AgLU+KxJO5qP0(3*(8y>*tp?eyRyo}Lyl~#IDW;B0aBX4hW4{#lF z5jQ+rDU1(9uZ&O4ZkfmH62X_I-6yJ~yIJx9KcC(oGp`m1S( zfw}JRpY9OkYlzp0+OtDXexgE^__T=X4yXYH@qEVWGL_4c*5bjqtZC8|YmA-CKXEyI zw=MdI&c=Vi2;`qcI;G1!kk9nAd&+*`M2B>*bUueoO%6$RJD2h(HV=X8z#gmpv0^cX zo!8a2ClP!Om4t`D)E8go%Mj?9StMRrJkP}IyQ1&)@b2~wWPtE11ge5yO3#C^;o*#5 zRK|=z%?(iZMXNrDR-edi96pH(rTp$Q@q^9VB8Ed(FBI;|T?!Ibk&Z>F9oX zgD5PTo3}nOz1Y}C>)6OdP9nDk7v*}fa|tN4aM7rd7x$^wwHIEl;@il4mwNx-2Te9mgpI~C$IL<%ZHt&#*1JN za)}!_QaN^hrMaw#(hhci*l6==^M-S(42eF0fprEit>`L2wg$?dR2krzuV-X2Ee5XP zxp%>D)+}k)dGFrzWNkc>c~iFhn9>0SJvtW`UmXg!K!iAHmLDCjeC|LY?7&JLh_P4NB4^<>-do&B#1vHfW+;p_kB1)%&80*=FH!eR;nU-f*E(G(b# z<&jHAgSZM`Cvrp)i%{@?RB~pD$6_W;AR2whiMl0GxgtS2_mJ-bDJT`v!x~*eC7~fO zjl~nbB2=kYsE;FiI(=e{{Xd;Pm0$BedFIra%D$ffeEK?MjriG(0G~cs?g8=?c?!Vj zj7F&R0-U0(jdS0)^6Zzbm_@T{kiTOt^KqCVw8x29#{mfES}5AL@lwzZPNl;DaNomt zS;-W)Z6*zT%L<-q&PxeV8GIk~uEgS$sB}nn^6|DV-%blr#`IvJaRvl0sBeQz#>C-fkNYA zZMa`MxL7b)6UnTJuxe&?;^TJh`3MTu^`OcSc?~thYiza8raaZ|?lk_l`=2smOxz%S z<$b{vkKFdVcXsgxIw^*8i!qaYX_n1)@-M9#J>)MpRJTw3bat}w-L7*?5_&2<&wPF! z6u2F1FEnhmj51uSc=u>O3?=9NW}l314?;jVjVpWD87yR%gIE#8a~RF z{5;&+KdxO+Ak+{kN^3D8if1(%9_#0^O$g z6bI>>c_p}bD8m=}FC$1qqHx6QmiS0}e2}BohB6PB*-wpS&EK=z6NusAK5v&ttm%Mj zyW2bKnv_#(aZIbIfT<%y2sxh49oNaee~N&b7h<{gEdfM+RErg)b65OJn_VBRI>LO) zNlk&qGsIVu({BIs+N@7^wLTvQdd^>eo_EquqFrAd{`~VtUkIYsK9ltLs75?^&$wgL z)Yvjg-Ikj$ByT%J*&!_jP#n6g+9^5KBRA|Qh>&>t-Wj#K!vX zyvYp@ji4Jg1OzmL$205rlKW%oiljC}=Sn2anmo8*IBM2Wt`XJNlg(S(q#2=^Jy2g86=T4m@@ST=<9dRjG)l&2ojPwzdV9rw|1ZO6WGh`NkTVC~ee0 zn2bxmF#-_pQ4#-o^t6`5Xt3&xm+Gu-?SO$x&Bw6GX7L!O9T7T`7IivrSKhB92cEGM z7c~pCl`^T(3)kLG4%<1S4Z9uzcOnRT+dk~pgNrPWEczI~4<)r#HQDBXvX&-{C ziAVK|?Ir|zpC#G|DBg|cEzoZ}6`Q-h_!@nV-)1{$*hemiaoKY5*;~x8FHKZ0P6;iu zR67;jP~JU6UES&zvvTh(@6tY;L4PH9_rLj9V6Z8cU0W1&c>Lbfb$71FMkk*P z2+v=A>N*Rqb^kxSy#-WU%hoQs6Os@?5(sVy?(Q0b1=q$kXyfj(1Hs*;k>Jo+f;0}n z-L-+n8+VrmUMG8>bN_SiJLBGS|2N*C=c-jzb1iy}nzj0?`OT^VeOgnKa|@b8gh(@n z0Q7=^5kWZ3fxx}5K;o~`asi><{`qBJc5QzOsk9^T=wej@uI;)GYn>Xn!wssc{`gX2 zR~1{~F0)nf&gv6`)eJ+XouzB@HEPb4ylb;Zl+?cGi>x#g%-%0Xv=eZ751p?jwv_-n ztzB4Jl(k3!eV*ZWoy55P5j7VK8)r8cMTPKozFIKd5*$o_?z00fNZ!HTjpn4omXB5F z6IXpsoMz$Lq-XL%5oF;?CS5}sAwQ$h^Cdd*rv53ilmGWWi7>@l;YVQyg&*^6`R1TS z-#+v0xGEGXY2yqizgN z3`h>6nNdV6H&;EM;C8IiXSrLrELZIg%Ckwxwj?s`os%T^9_h*Y)#mMgWSIX(d6{L2 zyZ_yfNcMG`Tr}xU6MNq`|A>;On)S~4pMm4PT<_&_D0 zSjejMzck&dkx$ZTVx4PEkNir=M+~xgtJ2O5?m)H|cfUF$vNF*ocP|?$w?86(e|496 zTGyiuhlcoM-?DM#W>25fiF2!Cn2>q%$vQymTKD{oR^DO$ei=Ra1gl(n^dXG2_7v=& zrlT%VW}T7mR&fDHY?qQf1A83Qo{3%1)OHb?b3!ELCU6I6e9R{k?CAEZEZ8F~e?;z7 zRlF0h{ZMnTtT=TPN#uH`-;W(mCj_Ug=o}MPSK?kGxxKAq6!^Zn+A|V=VI$w+jwdv< zBrG&bKRJ5@XHgg&Iaf(&&7X0km)$V7&Q`*$5`+2v6y()_&zVPogQYoY=BSmT;M0Hl zs9G7}i-J&ozu;zT#CYj0#e7*uCH$1mkFcaCK=5i9?n zy+wZ+YrE0TmE%pBtE!%!GdK=Ji{&XRa<$xP#d67NI89-5wpqJ*yfp=nEM!LXWzp2u ze^1bay_X%PSoz63z`ou>FqN&rISwu^8#cG)oN>(5Yvwo3X@9cF_=2KmBs#5ot8;S~ zaB*(n%GI+&zB6y;)Ut?6JJjqKwBdVn-R=slKk^baH9Hkuw6oaecgRLJMnPW@7Fa?maQ|-n9j^)Q0 zKF@zt@=c7B7&%w8d$`}Zk!}E0*@YFcucDXDrzm*;IUR{zYwO)R9cjKF>C@!0Hd(J+ zYn$HSvep5yMPzF8_-2(6wYE-2VlFfs-&(@XJa@Z!Gp_z911T&tY>Po`J*OAtqH1mV zU2G?jWw3^Azll!tbPXEnA+yS71?4lJvp{bbyoueGm7IASIq7Xr(!qki3+>2W;w#q2=52$ ziI$gTh!5E};;SHWjR@qhBB3Q;DX&b1%jo20G7ZHh-*P1AyICSTpGtEjwqCiR1Upze zX=x0;DU6qu9lq+lZ&xDIen?20suW+476#rkD3b5Xr{ z!<@iatG%TPo?&NfKi6q}&Ux;(wCeGGS#rN5Sm1%oBR+3)a5cxX&s2vVW@v}RQzObd zistCI)Nmo;PWXuMe!5y0NCZu9r*$7=`?O9}hakQ& z@1SZcYj&qdQ`ED-d#(|8t5F}Tot0MDan*5UOcW|9n&4qf&1Kfo$r)2Sge*!=Z~jxN(=!msm?bJgL`G9O? zWjt~@h4XIlTaHTtc<&6&V6`u?4BT03yMRdZ=h#+J2RubUGu|l2>q^6wvVAY|5L5)*;shI z&6iJ!BKy=0v;d8R2lT3r9%mITwgoCbObcanZU%F|;bz`p^K7{tM#Q^jqh}n^O$cT1 z4=>IFh8i;-{g?UlS?IF(w=){vwHd(fdlkbG*IS}Agb<|WPZT};`iuMVQEP4Z9h{AN z8m33aUuJuX^LfmieUK}zFS!DXe|zILu)D6N^%Ha0t$#P&;VYCye2hc12({Wep3x+? zKl_(_-M)<&r+^KbMMm|in(8;A_g61iTbM0d6rg9PE0)3{Mc3(SZU(0=+EBeN^Snwi z4R2wKuFb?SMK|imf-t+;{Danl`LBsE{PId7)I`|3DVT<9xO&q4o1YeCd+CDFM)Zwr znh)KRX=EXhIS{A#N{BfDh8oOi-^Yh1H1U$l_hV^e$q9AY!KjOGRv&Aka6DP*bXAbH z9qHtaY$R{Tilq<{vfv9cDVnajEjlk%-z6wRyhDZ*H)W5k++3r=ycLT1e~!|Sy31%ftEnovyw3Af^UYc6NL0>|kw9awq;w=2Ts+_rApgyW6MBjkKjyxLQ0KE651{df~IrCXuqSSeh-c=GAn>76wm!^r8} zSFxNM`{*Y(Z4_S=yy`A$GAH^EEQM@|5oL zD+e!7Yo7$Dxcq55%4#g$ee~WgU=hR84TaydjA~K7aaQB{Rr`vPaUIlf&5(We8?v2U z;?-R>rud5Ye(L!ek9ru25$)lQa<89~u+R&bf`k_=q8wOLj7Ldx9?=UB6;xJv73~E? zt>If2rJx4*)qgTGuJh1eGt8M;&TBwX^A?utY7KT*FdpsG55UCfwXPOgsoY(4W6lr% zuy%1z?yi255$>zi1fgyG&>{4c2vxByNE$5r@YWiQ11eY$Ad1PYKS%;phv{oO~Ffkonr4P#bSAtaenFY~IB!j$GcU+v+ zql-4POh`<}RPVcfhfGxeN(}lu%Hra~FxC3byJd!D6BoXfuHP6Y>XPNf?RKa+`ozeT zV;c0GYg5q((W3wybBp6d7*L@x7k!a!^!03?sCa^k*Sp^(F5tJs5!J{jXyc)R*UgNn zHuGf^=4g0OOV&4#mBa;cpa3;Trz;uTY5IC;I^j=LWLB77E^Wv5U(DA6epa!SG2ej& zjBR4xr~cqZlMR^C6jr4!b2rGST#h^BeqTn4?uTv`0iE5`;2&y023l7 zt-kI;+DJ)9c$ZEhT`yLD7J#`a%jCM?-W>N}R*$9%cwv4po?6Bu>*qz6xh*G8Cr}x_ zFG5$G=(QPAvH1&;p_%>*BBQ?fjmV19V`#t#a_2>w(D<*Rje`sw?m4>)5K5FbOf{b4 z3o9bItS#*Fg1RU{yBcOSoL4FOal<{mdN@n4NhYsSd3+kIML(YbPN=wspnOL{VtM+? zg5Hg->r~zgQolhzjO=5Ni+1+%wYviMRvZ3_Gxq{u%euS=&TV_UOiLf`W1s+jmXi)P z{k9h$M{+fvYsR~D^V(>R`u02l4~6xzOQ zK)Sq~_Fyv4Uu#W~Vta6g!TgORhhoY@K8t$J<}PD%bY>d5YBqEERw;cFyHd=;fS$B| z7w~RA!6XmR=>{5U>axBTy(iu?aYWxg{xnmKSNlNR{AzaX%|u;BYix1>?IM^qGh~jj z-+a@~_knpC^>Dok%T0J{wOhg=QUzG%XvDYVlM|!JYc(9|;7@UUTn8z2_ANWKYAua$ z-x?IJyrA(r*TfeF=GNER^9+cq^Q8-H^SB8I4RrcsB_6{QPd84k4af>tqowBkuHW}X zeY&^v`ndGeCMfe4cE8#+@;vnRIqt}+m;jl(Vlu%59oGKWk@?$%=awN&v$E-l$qPYR z?)%QzvQMS*_*bas?k&u3LCIZ%^Y<8LYID`->SQ>rac^3qR&+r)tx0tEk^DCFD$Hi# zo|&&A(NoRuM@kMifQf=!D{Mde>F}oO(nK48x(sDWOkScI+Ie|$bS`ZPzP3(0@vju4 zcqXo#(%l(o6qCZ)29Dmsg+1H6Z#jwu-k`BiB`Ilr)NCyh9GE(Uw{jAI#eG z^~jfF zF(Ow!%9Obh5v5w&w2iq^c*~ru>UNR+HoKy)SfZt>% zJ>x-mS(S6HHH)JnlV8Q>9{Asn9pw|b-nQ65pjLX@&nYVw<>z1ZVN|7H%(I!oE4naS zEzXZDm!>%*+{Pvl`x60WA5&pcJ!woQGa_!YQzaQao^JpzZ6G9d6#=L;5s58<(0iYG zZA>M9xNOE?|!C1~4L zzW1}h109eYw4gy;qp zgTHdPZ4XX^Oy{1Mu$%L%s`i6n)0bB7scIY#w?I+NMV3`RUTz+YyH!M; zx2dCWK#|R{r*m%qAWAd$B((K}*exk9GK)fglUoJ?bdaS{V)pj_^^W7&mcSaE*l^I$ zmlr*czipc@4eVct<}ju-W$Qss8C)G~mp>qtc$Y#VAj8x@{zk^b%}RuvX<peevjmWPS7$Z5TuaAeOC*i-ew zo4K_hK?P3^x~Za0Dd)94Mjt`UqWN_jUU+1UTc*KbP7-)3KQ|iur2%F@jbFI$`00dp zY-gKKIYFRWXWh{?Jakhzjgb0Sl7oLU39s|<>C5jqf4%m;;!yn?pW6S4 zNgJd)THV%aD{BX(>-}U^PvFnT$yYFuNZpY$58Kd74dQ5ifyJU-$8iXFOTEJArZJJy zv~2AQz)ufJyf#flT$d~i2Y%QtY@T=q(oqFwyVU9OxA1-@gTbe3WGWj->wc!F_RWbq z?<=MKxkH7mRIdtgfpPYz6ie&!xM0o5XAHWi7CONugY99H#kxG_s^y>!juNx3cw5kL z^?-Z9c1s>KB4@Wizs@ub;;&<-W&SbO7%0vk@hKR4drO@ES!`5sUhY90IIE(=KGlF{ z`<001*W9Yu3)-nPTijPH{A+1O(XJ2G#Mj(5tGPSR@8`9b3TQ$;LN;<4y#3@mSaCXE3Bby5K$YrK@){IM3-jA?p|MHPDkVEbvhTEN zCaGdd-FL8Pr{!F5IC1n`Tvoule{(pXdi}nB&XiMTMvhxALm{i_NXOLgnOue_;mEBt zWQt!&Len+uI7$;q!0)(9r*=2KyVY_m)n^P1VEjm~zufvFYJF2WYyDcsd(`B02^mia z$yTv5w&J;oNeab#3M6{ymW``+qjHvdLdEn9hvib? zJGKmzjJu>bS>yfqDrv-*IQmwoT)44lb~>s5u3-zt-8-`DiRbiRYhA5fg<;$JgPYrB zJgT+ZlXj4fLal6blgg3iWgp>#yvROr4LoVA>k0W)c#meGN@iik#9*Or@6DscQUiX- zcJ2{<0ak&Ry)}-VN77_aU%q?Y!jogOW`(R8O0qcxbAZ<}<*h>BrVDk7p6ZxchSCeX zS`ETxMV5s|jf1^_eI7lr9(4> zrxQFa?rVhtSC1qZn#%ZoP%sDYdr&vClF9Fnx>7gCn;sIEB z``@ZcdTNa-^=%j9H-2Q0Q#BmOv}Y1Ai4>wsr$pX9d0$+Yav*)7=k9e&E1MRWvhU;? zKOcrdW#dt5C(_fBQa9(5b?U);%lh*eY@TqQI69hFV@L(VkS4M5UW4znXr8gS8 z7FyuHR;pDHUN4L1S>opd;&v?YU;DuLukjjPPP`hB(&Tc&;<~#Yu*C0?tW&toXjw~) ziY)VLb~!P!%kVzZp&?tH?hCGSiHl6b`#8LyQwK85p}%4~+f)#iN&}#F+3`1{1q43$ zVw(;PD;_wm(IwfZNu9)@t=iQ*&xsgtRO{H82soybf1dz}!QPivFP{qZAIAhuwRuDW zMu7BnS&6-!lJu6SV~fzb_-ZxN?=V%q3=r?EiBaQAii@ffBMsoz_*PoVoV3ZlEse}Nps6PQV3JccuegMC@OAOP*dUrs7ntK?f0sbYi?pD=GO(cs#wXR( zNB!^+uS8Sj8-grqm(|qFZbn*xaF1t=JvRf0pC$dpm(z|J1s=p=D+B7jYq=*W3MUj< zxxXSslxf+WRHX`#S zk!ofCd1jGLKFW8Wh|JeT6V7amr-;MeIp$}% zOusG3URML(+B>RlPFO-fea;+2W7oye!OQ|kR~#~$_i2A}%i)8yLJnisQ+ zYZHEMRJACju*30PfPLp_3x0>ldtHG8%ctQ#Qp;L3W;t=g`~=4b0GnYfqUbs+N>rQ% z$l*xK-_BRmO}Dmi5ZQP#^c3mcQ`B;|y-8rf%6oa^flI6$mC#p(oQt0aQV8x7ygeSm z6lrnK3FQ@umA=r1hl@|_qX?M4bPJBwuMXcJJ(fC&PVcT4C6dv7juniR{FP2D6l6m8 zF%;wsL>>}=-GZ|H&e@%TTd%V_xm?neR?`@}{#2GvU`(h3%e|G+O{987RsJ$wIFmD6 z6_b%c?iT1kDPGt{oHQ9}*Yq$U9bh)Be!k zrr6F93q9QY7RuA{6^my>^bO_|EpmIKDOv?0nT3?64uWdNq+*?7E`t_-x)(+9G4vo$ zgDFYD#`-9oVqBYoa#FcT9b@iAFNj#B!0tuNTIge^v~Wp0FalZRWW$p_a`5QJ5t5dZ=+)|>w5XpT)6Wtb*Trb;owJ2^gx3V!xj>d zgMRYOB7l{4&H^oLIuTw({hZK0hbGMO3ES8+7rAOtLTsDSE^(r-Y-7dh_=Pjb#e@+q zSd!Oy&6+K+!MZCkpu}bM9hRTgy~2_$DCxp*$f~Uh%B>e;cRQ0rwH9ezi}EJ-tb3tZ zpAq-;#`SbtL1fAJLH%ZQucHfR7)QCTP8kDl#wPIbo3G(Slq;*n`9k>n;=ph5KE+EgxZd z!(uzy8BVl+yc(2GSAz94f=CQ0g(%(i+1SXg0{!}EtvmEjhj3Lo>}L@~Bq-fB@Ry=^ zthQ)athP$15kwtNB8WPXI;?-oJ7oF+DN2vlp2`2L6D4G|9sNs1@(t_oWVb(dT7?^> ziJ(_dd2j>~*84T|6{Xu4hX8Vu1=_#}j>JV9J`-KP&OF6YdNhdNg^2NH(6MIWD8;_( zWMDBF-n`R09}~eibeOR*U;IvfeCEE6 zUS9R+2uq{OvKyPbpwL~9Ha08UBfSEB;Ljl0Gn~9t|vMjX!{@w9l zaoDqieu`~kK;VqBZ@(+bW8_{(?1T}$-U{Q!IZDpkdWQt=_yQ&4*tmOp-BFFQY(%KB6BLe#*S!1FSqa( zY;Vo&N?R#$f3u0v=t-iGE0aA0LAi7}TK*b|+1hn~o&)#S{8^))ihBIIwE`d|IePN< z2yZ7Mgp};R!S@n>MFQp?-}L9NLazH=)NK)MsW;MQgoXr;u8;mf4!#xE{TDyfFJ^F) zxpy!Q3$$Bh&*!!PXGc^sWv3527g*S*0C^{x+ipu}<=VAuaq`$a+E74lk-%{o91R^R zDpS>YT?-QCAziQBfnf|bVfMh zi!uii0G(a{uUm2p$Y+SNZ7q6Lb;XKnL_cMu``A~tY2HlUe9lyG!r5Tl+6?CwNi_}d z^nQF-R(D*_g-gCYFCakz`h>*i9Md7r_^}u?<71)DXjB=7@esIZ3Qvo;3xk}T0j9>=MtNOMO-dhX2^^xnAn(~o^c7G=8-oFw4l{QmWEt%(S z@~4V2{ig`s=YM;4C$n|Yz}oKiPAT6TTl7$;#A0w%?SOX zu{&bmRu?fl6jEQ?JZj;UVJVdy*>isJgwU8lw^ z@S*6DMZ}%_HHy`3+i9|0#$aN7`dZ_1*FZ`5CJ&xn=E9QdHg#s&=0@s8iqQD*M9Z8F z{`G?z0k`1mbbUr_mY3t)_mRnzz(bj>sOK`;X!QCA8HO|ol_J*asw@>B^SXraQK$gE zRjIf={vY{-p3Jm)liY|3pA^1Pdv3QsyXW07(VQ|ep<`EIJ5ju39jsXOairHI%d73y zdqJ8gKD3-Uj`E}s)Hx9SLm^*KQJ@fr%5Rr*_ovJ>3(5SMvOA9IiwPlKNLa0)tYyc# zilDBozk8K_;fQt|rcGfLSjaxkr*Qq0AA7L8*rH0I=qrx*Fwwbiv;CodeHEtu33c-q z8zwmV`FIj)_yi`YpNYnFl}m${rvY~=TNW_h!xLqC!z%K6o5-N)c!C7+XAaKdbo%U@ z6P9x(S(HG;cqtL0fc~&hJo`HC?$&%3J293=y=gJgpF@usKhC8L!fKQqRqW(_VOJ3G9+D zE4T_1g}A~~Qp!gH5R;25!r4p0Ydm*u|1EE z1yLUCCFMk`_kSq|sbbx&->e$VX?=~Q;D0XbcUvT*>FkS%#PyS9@t7XF>qE5)@tol+ zb{U8k+YCfF65k8`lEuGpew$rJQ%9NVOz9&h;-l%l!*e8_N5Xq#v3JOCcrVRussEhM z$!^6qDHOciGnyrEO$<~gdE2^u#*T)2e=DEBB_L8i5Lu|%4kYc9We;hcA`f9_9#BRv zQN-*n0S&qgAht_fiIFov52!>OzIZ%f9O5*hIxeB;sD?YYO8kiEXZmaA(_9 z!}!2``m* zrWXirmpfJ5Opq`<{7JEY>b*G>_4ZS$;Ji5=;t)fq%v(lB6H^uB&z3^ zr*)7|`uwwqS_W4}7eM`zs=Plg5KALanoKdRs~9yYYV#xiPiJz5hA#z{k|x=m0Hdkw zLNKWftwZ)Ap-kya&S#g(qFHjU6XYGWg^^hQEcR?TyYtd%v=e0XF_h@^TS&0w+k!CU zp%9(QcF>=f6wDU=el?rljPrw_T2Ff7S~mZN7#8LfTg9W*>@+9Rf?VZ+6$`qK?nnKTuk}EFZ_ePZH49#ipo_>kYL-y zc8lvcvyI^z$vJ~@y0^IqN67s$zA6g?Tl+lI{KG{O)SVUQS&oz8hiCf`AP{r930Dwg zGx78q;%B49zS!$3@S07=ZV{RH&E+2D9Sa%^#q(yvXRh>_Ff&Hxl zuV`JJ1QxR`I#R}h!gUf>R&hO$MI?V^#lD8?1R#q$g2Q#hkW!7TMh*pOjWwF(4sH9i^S;wXJv(AOza6u{c5shut3zY6M z>_=-!()n`uthO#__%k+B#9NQM%S0b{GwuB8V3mRd{=9|@gw%fvB{HQ&3ccY%&I|n7 zPbl3?X@9BEKkiPQ)3pD1dLH#P!hB@(JjOwm+@$F}u_YP1t@76ho%|Swg7HSC;E3Jm zbw7E?%Ygl?yOG&?BSCJuvrmmAE=D^&T^Gck$OxiM#4cwwhx0tQ^ETZctio5@p$BD| z*?@8>+`z>T=e>kYq*|A%(WSRVQqw;?O#r*G3m!CV7`?X?^CBBp>D@gA5&&p$J{Uw| zywg7J4p#$I9c=GSqc2X7KE@hJ2U{$c3&ssA1KdgAh1NXpKIl7=cVr<0w4WOSBHEYw z1d&VeLL035!sT+00`Pcu)flgJEgj&KX1Bm;fuZkk{`SVepvvH!rM(T9D~ADN)+u4~ z#w-NeY+{ZwRr$V-+XJ~s_&lmXip^4!(bC~<$YCp^2EG1?rN{$RU4?r|KKbA@ofO0f zps_;uUMKBHPq)~l1V(0V*kzGsfk;Xb;fV7Q64!G}nreOV#E+mJ7NXaJ{Q+AcOEX{! zA_&zVe?{X8u#{6Q67Uy1(pDHG@ z&25d?jjWB83=k>mUB!!x2*svSQE|_huQ(pm4VYd5bw2WCv}@i)rDiU?&*eWm?S0si zP)H#^+0-ib#j{Qtu*;uxw0d@T) z$!*A)TgpH~^{pSIrC+QOoRnCzfIS1t+Upz7V0%^tNaFcz*30BPB=Ov|@e7+Tac2M} z1*^D_OBS2_1rT?+W)+YL5YMdSFhX&FscF2iaV zbfGq7()*fLSZPyf>%ukE84_jE2MShkL=$O~nbp+IB>U>@jVG01j`>SLp$SUsT}T$y z#rhYv<-*NWncFg)P+RyZ0~+@W>XsmK`z~LSCs_G{F7`yWAfdrf_(mW=XMZ?ux3vcZ zMDY0#>hMuU;{+0aQ8tApyX;C47m|Q1nw`fdw@FNGV9l?_Z3~^Q5vxPU9XT%feiNF# zXFs-<;5v(ltG%J6B~G6s1YC#a6wYXV$LxN5va@*_$Fcb#)3rq ziDd(IF}dj*&D=#8%~309gW|U&0&Zf4+;qin3J2Jlbighed;ae|q=JuE-aK!o_&ff0 zPsaVpKWM%GoK#t68K4-=xqMz7bn?8KR_HMkdGv>poi{}Tp|43iWsv9VZ)nXaew;q% zlrLVVB>jRLD)mibBb(dM=M(bi@`3AMP9>=Q*zG zMP@CL*~QKgaN3oPKc?uFbh(2?%`qsnHJIDCip?4kbCR5(vg1!TeplJaAdvFne{rp3|ugsw=!MIL{uLyZ8DIadxEgp@Od_ zUvskhMjH$5*`Y#wNf@s)aU{RM(yyk`6}&~<=nUzeRy8(iS;xB^HCa?HHiAJ{KFVpZ z9eBXmMK~d(rskbN*K!f+iDmYg=gQBc({0du>$oZIn3<{N z_Tvae{gLYj?k<0%WDTn-^hx^!RYcrDw>?cBe18|* z9!R6ubEYNQ`C|#UJ=IJ0@^soI!lhMwa(I(&x9vw)U3yCC{5u!{*!`ugjgLv3hTan{ zzSw4`Z!(@!DzwxU4UhYa)iVBotI;3$Ig)erwED6K4lDsSzsBNug;t%!&9gRssn8Ho z;jVjhaNj0`LFXq#*L-8EKlGq-CaT2;Z)XP+FqS$@ic4eZV9W_ zFU5x$wr24ru6{s0ztTxkiE=wi4nVs*Fo^VfHJ`Y}?~kgvJa(Jc5H9j^)ZV2PS#9ij z8+OunSr3-+A!`J0Qwen4w1IkZV4ql~H=cEARS0?j3eFKxI&Y{KMaC)7F8-+U ztLDoy)44iUowrWBscxfp*;Eif^a&pcBvva5l}?9|f;R$c=LCJd`+MyFW`Z@k9#;Wb zKmHZTAX#MNT4ZpzYb;rWSH$p}{q?15$HQqUI+FO$=dT>$?S^X`c(+bbPO+~~pVBzV z2~jW1?^Tjh@8@ezr4Ea)){|Yb=9gIUPJrGW=L8Cz>Pr}54LiR25ZU)?cta!beY7T* zfqGw7p+uUU*^4jZie6$w+_eSB>Hqfi&fX54!VLpu=OHF8LiSATg^={;V>P*3sXEFG zXwGm2=(K$D01383Y9s{HWG2TCFL&1<3nx-#9=B`8Q z{oen>JB41*s%#(dVS`YRySBW|EsP&Gb~xoz5A}HJ=+-P@EC$8+iVVNX;uub(3qJGd zF161JYPh{K_xM+*ng7He<+fz77%^**6}d`zvM^nPHb-#qbZf%amffGf>c zp($4yxwW@6GR#?exI;efp5R4*@gXza;Sd1>4$VoxW$vt3tIVw@RNoUr43w94(;*O| zf+Ie+Td!Q)bp@T~`(j~)4XKHeReT+2J@iKVN(}_*e%ncxBHlh;!&zEz7J7fvcw3vf7ZPc!qw9CR;4PihEC z;vW0{NwU9!BY*NK#oB@H8z7q&cVSXadU~6^MEORNr<=fO#yqsiy45?C=qu?gOW{b& z(8a)ZfJyG%33t!h#Q=&>s_p9bmA7}2?Q7doP&HHJFSeNEZ?^aa%>i-+>O3LFYyomw zAlrBM`dX_DrT+WxJm@o@!o65AWWWeB4yLH5fsUTf%LA_Y&jP&96e)^cFV6jdUON|h zv1Jw3F3Q0QR9m4K*=B*6v3xw*~$z;YEG9H<(N)rmsi{iI` zbH!_aamDVxxZ=#;T=CDpx#C(%vV~t<@z(#u6*EACc(!~zR~DNc&B>b9F>9UTnncA z;4iyRarEExZ=*B7AQ;AeZmh|R5lr{Dd{c+MYz8d1u)ft$q@RXG&#W#{5M^U7U7D3tN)`-v%+vtzU}8h^9m=P&$~` zzV)eA!ElD9$8FKf;9*V-lSdGbiE7@=O98HP;)QPhXL@%Q^l#|hV9MACaqF7r?fLzV zYQ`uN0Sb_U*Vt;kp%;`4Gdmwew}XdvmYcf4o18ccm}MrPDR626zSoqVGBN=O*pAc` zHmNbVwGgrGe|DF8&jJV=I3--i_xjd;nEqEEkd)UtU|ed#xv*e`SzhnlFhy1`%FNx0 zL7ik;g?X{fewtIx7VvZcMJuhJRh?OCp@cDf98-s@7#mZA%S;riHOMh5G;Oefst-_T zu1(tIM3*FZX#l}0E&b&KvVi}$=v}j$;mTCuzvx|+8jcJDk@;7C`VBpfVp?bRg31(= zvJ7qGccpLNt6(igEr&J$`2V;zK1=zSQP4FEPXeNRq@Er>0@25d)A%o|r6Cl)*kcc< z4P}P}sowa=$p0Y3mdbC44yVAOWTC0i=aP*21RB@QHD{-5;b^+#B0Zw6(a7#MqBI)kYmilRUM?2y>d{|6B#r6+!aeTUm{pAr z@t65{3J)9J2@h{=^1h?=vG>Jms=I#glPV-)p!bl_g`cl@O)d&%Rd}e{3%GMO$lUiW zXMs6;mLEc+?R>9}yRhX2X~N_xXXp(m(+)^y(F3JMkPe_%(nP#Jyt?;yT?jMZr8o&AKWe}7|HEo0urghgTeoX+g;5D|IO{X z7a_S_Qkjati!GCEAtbkpyyID)dr^RIto}CMc82k{gUt`IBX8!NVeeOuKZ4w&W~s9% zATKsU_PhB$>uE)STPIni1eMh&Qph32g&{WMbe}=CVYAffqpTI&_@to=Q# zot1)S-4dbKjzZgSUjL#;M|zj!F)zMpuF|k9_}kZPh5~GXw|pSIsoE@16f#6e$3lp5y(zpThbB6KNe4*3jF2kuqSQ#YFx5M2+u#*XW67bst~y-IgU zLv~S14S`pK7f#7;Y54>&0i|n%;)VMu2sSzO`A3R*MWnUAW%?Mcy!BM-A&}&tYC_rJ zr=fPZnJGEbl5hIZgdM&26iu7a`|-t-J7~!F)_rQqzb3RpU}qptto`|*39_=&TEzh! zzXO$@Dlg5*vgm~FNd(lWD}0&DWyn}a|6B{#Dyh=;lJ_%=;!^H_K-A$S_RDWf?>9ur z9NPk}{Ov69P>6EiD8Jn#*4y@q{dX}-1D_)B40se>6QHVv&ObMbcz?sj4TMd7hy9rF zAHF7D0XwaUfQoz>$dyWmqnP9DiQk0J=v;a7b5NA^3j%ziu(^ZmKefXpaR+ z-Al1Jw!e)P89BNC&4%vN$LR2zP=C;izZx6SdM^NKW$8$5bT-=^VlShIuJBE09d<%v zwX7dpavX?*AiR=U zw*rvwg$=p9=f^cCAM){hyT$@EZeR_1d((J^%P=d9`DMKm9iQSzX`|-Vmn`ojzBY8n zoTNmm+ervGJuVwil>!c(asW`Rfw9L=%m9FumL%?;3~#MVS@C^uQUSjO2s&T(Y63hi zjy|#*((T@UyWKZGtCBKJuRR%uKMsbD#H9?yX%491_l?hXo6&0tjSO!9HV3#)KC?Ke z>pK!5S@(0WR6ml5^fcCnFBMG`xII!KhgWdH+tbvM0kNhTRhBaM`7-l0OjlcnjP8Ll zGtNJfUEgllo@t%wrB}BKoagxJJ-Ej&KHKz-;r6B{Jl;He;|1>(Ho$#Z3CWP~v)S!Q zKiCRwfj7SM33Fu5LN1JZYGO4^!@Zy@^yVbGD?sj_+~fr6h(U`ByU4g>)S*8fO9gZ0 zl5Ud+`xZrf=Od$T3i)nyHrc|}pp+L0u-E2k#V(4o%Rwo=!gozoBeKAVN)f~+$^=B> z5|B(rOVq+DdG+Bn;^W(A{9;^r8HcEr6ql(oKGjX9=*0^l0>x_Jrt`~u06x>mld!Rq zVS|ayDh?Yv%e#%PDtSRq7pGGmQKHi}l|+^9AFnd}@pD?^ZI&$K%^4K*9@d>%QMv;K1&V{_0lauVu}w+Sm+P(tfs zf?h@J$qH?#E+=a|-6X1+_rtLLDqzkcxyOw8G~54VHlmJ#GioS7xF~Wc zL@qjN==Gx4dQ_?Ik*B(Fz3(dOVTMG?qWwfgpU?aQx2x&Dir zl_+zrYDY^}8~W2MmL{W($+4o`w^$InchB$Vpx2q-1lfxQqS1DyVnr*Q_yfnmu5Z1N zS)m?~bK)_z9gAV*v?OMp#k_%No^8EukhPA2oNbeI;qK>> z2%h(QFX`A8oo@)~qIps?kQ4omsViFj_hw~|C-tIkjnIu{XKaun)tz*vYAmaENdP^p z;1oBMeTD%Ae-U$~vhUr|*4t0|nGK5<=I|7L>~}};%pgXtB=RSIRwbtUw$e_1Hi#RN z3S+VV*7}+t3gtg{?!8UWoDS}zPnf9_tSqecA(6g+>UTX!1d$T5G|6i4_i01bYK1;_ z!9BlCo;c+l7a6@xAMuNro7qWLC9bTFi!kxh!E9b={b=*ItMN%Iy{rB=gRzRfBsa%e z28DycqlQ9N+OjZ%c-T@O>|h(yw&W^cUgm!JLfX&XOxLPDI_N6-)*{4DhYYl-LHQb1{SS(+8lJeL}1Jz6CAzyhPN81alo!rsA znLNN^N&P??bbboGn#w6?LF-_A*?-ds`IwcQ=rCOu&xzbe@%~clpTVR;Q8)ZBX zpPW4JG4P{&bdd7Tpz=%Hz99^QwO)iuZ&Q*3&~<0#%dfgC2I5WRk|@a7bi~t>+H?8D zM_+)lDk!L!g2N9jTa_~wb&qGEN6C>! zl690F#b73{d!;my=ii}+gxpZ{I|M}5RPEKWD7>zaW?|)QhPU(G6@c9WKb(>4kXO>c zl-TsHhAlz(pXP9dBOqa5&I!M~eX$ zy{GLyAlQIy={Dgc&l#X25&m;eqxzKYpN8`L1l*H;`lws!xC-?j_Wv)?fxf--`TshL z>%Wv7X{@YwN~fpSHN|6B@bkHhsAN}LDkwCy#$z9DuXh5FHF5FSBKt@sV(z-H@z}{u zj;kj~h4^@EVLA6W+N3%#^=TxC~KKqSOtBV&(- z`=pbd9mfF~>CRd-G$aZNm_0mFRyyV!55-C6n8Y-%pB%=%a>-D>mR}4a;B8pOGgg{6 zQ@f?Nt`*|-?8)>ArK)$p{pZuQO$?$#YYO*`F9VNNArcSm`3fyC+=40g9lBxq=fe1( zoiATGinuk>+Ie`6meM@O&3LaS(~M96r(DDxv+qZ z{O!iFt4SI!szlR0UieLdO zlmJ0&G${coQ4s~D2?C*okc>15EfY{cDI!=vQ3z21>7ewILWx9?B1J$TfV7072vVg< zks>Pk2K4*x_wK##t^a?mz4khpv)egya?0L&X7uVzFHO!dU9M|lM zZEaiXJo9C0xz((v@Y#%JWNyp5Yhqv3)}O42v4hXI{cL%xuC0b{eaDCUs>U_$w#Qxn z;Qsr{;@-F`{OQ)K1MidJTbCQs%fFbfg|43Ge9ehEKQ(rB*GIY3 z#pVw`dg#SBs>(G#5FJ7*-efC&_te(>t@~>B=E{s%+fS}M2b!M!)#G?(y>b7wintz4 z16i%PFBN_n17AuS2A`QLK--?AxWFy{9$o)Gg&Kz)l~1ek>B>i|HGlHN$KIU(`1)4K zfWLQ!Nz=;Ld~_egdvIJ?Lgej(kw0%<*%2~S*d*dw%5;3FSsK#+&5j|FocvTY*$SH& z`It09U6pgsP`bV6+4UP5t#8?VKfjg!Sovy~G;_)|Mk#5ma;4CJGP%Ut51N?g>L*(c zd7n}nO>K*$bR~9Zr0NjZDQ%IzP;LfhnbNxT?oH;ZfGSEHAtn_j5_Y#E`mkGxb0hRK zH8Sqn#;heB+Wd#cvE=TPW5B8FByo2Q=QK6s4F z=4jMU-|{qGANhTBZe=k0NyDww?}2uC5)0%Xt5tEU>xJ45!&|v(PMv>0?F6vQm3ib11h(koF*Pmq7Tg(!#{r@+4t3MezMl#`qXt^s%&KT2}mOm>;lJm+5Q`H+u zGKumzMmf=h=v^Nd`yBF^Z!RqPe@I{&Sk2ufZy8OeL%J7Uer* z1J;Gk@Aw(GHa7?RRIZQO&;OsK|9knTK+CP1tx#n4SAW3h7e~&V(C;HZGM9s0$LC)b z7mjj%i2qu_zmCfaSV)%0FgvJilv-eK%xC!d>-r<6pf7T&1(R`7R@{N`m6qbSfNWejx(1!`Anm)sLft9sKE z)S1ibrVKPxE!`zGT|$V~cXGXClpuriftp@Hh-NpbH$@A^sF^bI{?PE|7#Z&D%8KcE z0QFY32jrwmmcgBY>ck-aCj^0w6-S+*NJvu0AWh+{BpLFhq>Z?vUrO@NDZuynX|?{v z$_xId!qj&EABCoeLhankUS}_ya2T0V&i$OXK5_KcLhb;``%OXDx>Nk(AV&;Y)urzG z!flN;a_hkU`@7GtoB6Dr9x&4Ut#Rrb`QxGUD+k9+8c07H7w7-M47$9%7I-jMg+9#3h_kBWQigdx?-|Gn%Uyx$NL~M>H;@wH1vhQch#&u8bhHCdb_GE=lw$k@6 zEIPXD|DCJmc}2y^Mv>%(#Rh{BN^wW&=IKF9ocu#voO}Z=PPT^XEIWc;#l<&2_25R=V7`H4oti)Vzw1P5F`o3ku+RO%`|_O{GM$AQ8kFu#f=0KIR^jvYTOO0oj_1n5A9Dss z<2NZ13N4f4Pgk|BbWU8D{h5~%b9z>GHS{Hp6#TN_mW0*GFY1UCHQp0zxQG0O(4(?5MVi!7o7&hLvaw1#1^xSwV;vXE59au<6qjZ`>Qf;*%lslj zkNuJ{M@BwJtr$Wd-Dq$;)g%##Qyg&KXFz>hf7vG^xbSm~rGUr{x7OFQXZ=6fo7Nfi zY0=N0FV^BWc56a;qO_6y0?`r6gt{_tg)j&q_wQvFR*1sy^@;Bk#>R64+GB|+1w6P*Gp+!(JvSsdSz(b*n_%uKcMVd(05FCSAFXN45PZ zDMt`AkMps+ai7LizjVQDx?DZKY-CQcgv!oO1Eyu=Ic9dv>}ybT*z&ysbSkzd2Oz!g zZ?6ZVzpOe|jwED!{B+oAuJ4~I#_ZFA?hli@=qdKoGHvVM0<RXbGT*IuMb7%W_AJoGd98M_&~1YEU9#Cg-*&e+y~MM? z-M1zaNuGT#8ax`#U%j%HkO)6`u?+b_1~yl4bGoW(xf9M<$BMy8nok zdhMIJR1kM_Xy(neUyX`5a%_29xj9_$_m3tIv*fnl{R#t@-PfB6NVU05|cM&sF+cRR1FJQ+f^1%JtcQF-)Qi~r3bIu zDc1eIsrtSvULv|`<$}D!9pA)Vw3jwR1@Sj0Y2Gx=q8sL|Yuv|4sBAblv@fXB9k=N>G^7mf<-IFZ+(Kcr&O614YzAKft7|GZ$ z-|w$pLGfDP>}I{WM`M@$Cm|P zh-C!vk=?5yKA?T~HKK?2)lRkL*OW*6ZxTe$SaqR&Z(`Y3>|Y120o_+amz5rQm`-d* z4tz=Ub6{V8cQ@KPRbi84`869c-dE1q`S^0vs`cwD-|b!tYIT&dFJ--qzWGdLA}5>t z!%ivCyI95UIqm6oGv$KgbT^$OUeha;?JciQF+`H=U$s55MNpZhwNniNHB&2|@Z#oKRcss{}f4m&EWf zADHoKaK-w9)#}p0%DiuVaYRBSQTFXyP1Cj$vOrCyf6L%1T_5r5Uw#fPtrnCBlkKfv!X6c0VH)pyKKb-feZt6y z=KJsXvMTJ$%byB5RZ*Bvwq1%V0iUwTn9(!8GQUN|{v$ZFwe0p`$6eU1_$?P5kIb74 zSq@zo0*1=9BSsD#%1SW5bVN+BDZ#KoX`7y^(#nXFwi091$qhnBj5iQMgz(Ccr5|_N zSmDp5?t4#^vaTIYBbAgHLf?SfKtaA6KjV7e2|NRIToku8=r}4itHPc=)K_WH0Toxo zd1H&s#Yh5?_aYF{eMb@z7p{=(JQng1TB#2_1rH)7Pkt#r+Gln8BN6FjwGH;H%T#Y@7*e* zGo-V__5s88L8@|Hn@EP}1{~?T*M39B@8$;h+>o(6E$DJnX6b3T&^x^dwrLI{rQxMx-hsXT zN9ZTdo1{p}wCxj|ctGM!{x2XVk^Tilk*)8xF`NlC8iGH8$t0ImG?^ zG5qUJs_71qGE1bWyElt-+Lbv9Wc2@7VK7IQNKKHYq;_Qkf$#N6I#t+0bvxavgg{8& zu5~9_-Ks#29m>l8#=V;=_u7?Z1i#n2>r{yveHMw?vF_vpX*dMg3ROae7(kVy5U&nY zxpUn~53Ro?mx72PtD%O;9Q;G<^I60Sk|YliBbSJ$4_p-dsOG!x zd;Kfqmw11iNq+nU6i!_U(9seb=pyv^%v(YB~?Krcqf&e68MD5;AyH z#46&a6By-SqOwS0>ffA;Gg@4a@#hGKpZ>_pUHZ|d>`Sz458wjn|KNxS}!#68hi1Bh~QOo506KP@?4#cTWYV)sF<}) zsB_;Wn{B_T@~T>yqd8A9+ICYVv06Dt^jrOttyR6(wUXKSEZKdM4c+*<1j|TKt4I6r z6cwi#_egB6Ma77UlUfl9cj zN)BAI;Rz%Jj1<)rmgE$6zNmlTn_u=|pKhe6Q{7(4FwtVWm;j%9MJ0^{vPzu-kM6GOPBBKfg8{A>L@MML1 zsr&Co+bTrLDZ?5XF}(QPcwXZldp<;C`b@ca?@Q=TU)L*tJ7lOV){w>F!|Rd4RFbUU zU`H62)~A$XXjLY!eyK9Pvka${B`H-VFMc&|i!wqSYIuOzi&r;(P~3p5`))mQAl-Qz z>wyZxx2|{cUR|J#GYTE`gZoLoqD(^;Mi8&Z3Rg*5uEUOmIMEz3+Ag`s#tU_=^-itT z=Tp4#&WXVdE!sXKOFKH6T_ugRr!m7JnZ(HMI$(x?#GfQE_VFHp(FoUsPJ ziz6+ZrPq_IzV!V&{- z;xeF;^I{&l2iGI>mCiH(JqJemu|0gO1FOQlZl1j#o|=4@*2V^ToPO#a9OJ)RCld&c zv_NP#2yqVLFYrEAnG|kJSQ2nHap70ZDcvDvWqS{jquWrj5}@9dB6@mOak)V0M++kL znY}}rx3sF(?B2&RuY|;=t6*X=xX4H&xrC%1SG%{{%lX}8v<4})iT9d=nxCl~4^z-s_knv2Twwov!k zDDaHI65NR$C>70zv2y5u1`?qGMI&&^cfI7H`BEBp;B*_|I-I?Xu*Xch07_+6Orc!FALUQWBo8+goC|br-4S_?E$oHz+OElHgb z?mosDQS@-mRSWes>yG|-eKjQhgC=w4wES$DyF71ZJ}okbp$@O_o)vQA-hTf4{peKiUrc6hBVtz5x0d=gSpl=aeP@&)D6gaaNgFGgI$#$OvI}cRX72MGnwoQ6p$U zOf@(voWTKltfw7*1$G1PYU49F*muqX7DOQ`RV6zQX|Rvk;GDtfj>k%3@utC1OhGDb zxuBK~vC6AvaP~68@{oG;!=Pu6%wh~L9T!*Ahdub{t>5eRe!UXSuvq)~m;=Mh_c*iz z`8T%Hy`$5R1$#>x*dd%!%XYX8A22d{j+i3jPCmgyU)NswLVp=0X)Ltt&@tkXc3OB^ z;KaL{Egmn)XWU!TI?j;4`uuZ8+QK@-muTej@<-1cpQV6Ry@@N0cU7=*HP2Jt#yP@v zUrWof%eLl>vW`7T9q_Vrq^$w*+a>RuEw(zudB&--qSVVvp*K~-VLpfdo^mIjvI6a*U&^=KS#&S1JYW)19 zQMDut_8)Sq<`8EjBGtDLs)X}%rDsDwVQs9{~Uasvqe5uG`j;OQSHcS-vo0{ z4g~E^oOWl9R~t>gT2BK*_^C93*B6jkZ=|Vw9saI*Qo93aVn)14W!PWYWOLf!`OD_Y zX?N$egB&lv3?v1)&%m2Y^YEFV_eEYsaI+Ga@jp+({F9MOs>IJOo2<_N!3yzhvd&lQ z5zl6;UU;`?;ELEY=r56-xszegeU4*p#4$&4%yl{D`keL)oOT0_`Pt3#9a=h>md=IK zRl~ix&fo-_dEKG;4m?wTj7&L9j2Ib|CDw)>j6w>9$u+`HJw~cWA(svj^9~c;Wr*&w z8#v%Gt;s4Nj^-=97*0!V?Ei}@ABLCcal#%|aMZJeq(7ocFej}(J@zp;H*&+-v-q(t zjA5YUOy+YYdpMI#oXL^S7@U4~rtUiy|qDA|-3o!^Es3#8Ev?<9%9(6iyO{K7o_m z&M>&_Opa?xQjhA8i*j)$EB=R^&c{g0XyoFd4V{gTXia5b??SvAm?Z+M?RX4#h1_=QJDn}ucQe6@FJPlXE&&y{%BoJX=ZdM$3$J!U&9&dsR z^LUq9ai*K#VvTTIBRs!p1E4}*BU}tBG{IY=kd1O%@9TB)7r9{im34>K!H2Wj|621VeBa^Wh&G7{iKXGO@>H2LY#qE?$Z{eaDoQ^b3S>D#6%(eAR=8(Ks0SbBOB{`h&k_#LKg{e z!UlVLa%BZvyaJwF0jE}Oz(we1lWp*^Hh8KHKG_C;XS2M$u;F-N!)|+Wt37$55?JN0;V$I(UhN!W2Q*OS9^M+H z?%Z>8l#Lb5nFnWv*uA6d-Z#|Z25NC1wRm6yFPtlsG5droY*nFjsB zU8`5QagwhQp;Yg;{5mp6I62Qz`1Us6j`UeYPNPC2Q zC9pIPcC=FUwF_aXkWm3l_XW^mjGYm0LsBh-vF#CUH<)`V>e!3=XDF z8rd1K)Os#kMDG5J*D?Kd?gRegPp#hXEbwF3urs2m`WiMQcmd4X76Cv!-pr#zphZ19 z1Cn&S044#czREoLj$JmlUCE1Ca3hAN0OnSjH2iJ#Rnq*>d0W%K-d%slyqTXNd0|w2 zH5-y&0j$*)abOf&6nZ)a+Jy%VgR%vFe^DXc2^a#w!|u$$@E$yrFMd5= zp^Db7bSB~QOiVWxkzWM()Phahso8QgnDe=^8SbkqAs1Id&aZ@6t%MxEfkY$n^8lX= zu<1S}TaE&e3|JyG44I=ckVJG$j>@U*Vm^vAFWOi>FD@kwSiA!gc2T%-<|oj`#^!X` zu3E{J=3%2e*lU*Vfm`>pE>9w2U!Lf{J6jPHk^FM(9RsVY6eAyxKp~n^0834#`gHfP zdnt9F`nQ5Uv0&37cJw)lbST=`FF#Jv5|Nk%B*lRHFYrmkFI`3R3r2K#oqh#8Ow8`M z5}d}bHte3Okv~j=n>c7PK)23*x$M5xC8J;5-U8HM2WU+mM8f!Fm^h;1x~fH_~?R z)@HSA@KZyeQ4>2lg(}T&V=P|~N3%url>!1@%&cLs36dplW4s%>nXiDP32?xld(Ty7 zf%p?1dYo%X((wFhNz#0;KNhhgQwW35z%tc6*=1KUS}!txi3bXzP3OPJ?Gc`xNL*+w-hsMKw8tg-^ z(fW^4fmS)=On>M5KH=x!U%I*Z%n{;g!11$8S{R7qXEV4dE`n&*Rdc#igMejJ!yHUf zs`?P$!E;F0DUH^|P%!1tgNg5E2KYgZ=3z6q&@9)y zxIH?|*U2HgUn$EUo~fT@p5DXGx?iPDPnzjH|D3WjQ^*{ld6o%1WzYxp5&tH~tT`e% z9k34pYk#JuTG3@mX6Cco*r>nA*F=$I?rBKs{dEu>kPXV!ALrKGBl7bew`T7%4F9$c zCmP1K2vBikI=unw@3JHIQHA|!n4lscu@*eaJ2kFVAIXjo$0(4n{Mw2iy%7>v&*!S; z&^hUgRX3KNR|I(1g7w>}v2wI9#v!|v5RDr(h#NJMH)`Z=oJ+2R(=cc({p2O)a45T3 z2%`(lFc^j1SbFItCgTA+ix7z-C_;Ke*b%#^!m^usI~OjQzF^eD8qVRsFK?1&nQaHzEc4;3?mp<`))3XZAj#+Hu_=^T>i z7xOWPoelwQ@o_wJ<3cUMG3xztKln}Pq6vV?F=mDN2&lCeP0AM1)>=L#pADQta)L2J=Z zek*bMN~o{N6@G4}uss+U46c5oMW2q#s>zML{!He~&oj?tPOqez>0WwLn=8Ei8OCIj z2a;4j(-kqW$!4r&noxCBi>^R2)0!RaLj6U)CNv2x{xB1&+*ndA)okPw_k%aT5sL+Z zjk)0&^;-0@ZzbS7fj!WDxh}zT(4r>=0jXaJJKtmoP`AhZl)AL*VK5vY#q%gy z+ijsBl`;5la_||3twjfd0OVT9sktso7H;8SWpbSsfwdA7;jEf+qi#^%0zUg> zF%bH1Z4z7m*sy>-+(l{K>@UnSjN62{@twm43-4d+rc29ya@1ZS+>Xlb7|@~zLM<5E zG%v=C`d|9Kmw-zI_EgtYg&EOUcl#0V@kiF~v!d5z0{X4Dn^7-3#f+V%g8_i`LJ2rV zU^jG8Zs%GHn^AY4ruPQ`zC#3dOV`x7ryPIr`D13&Ws-y;g74E@qdm7CgD|7S96^GN;;cyRlV7q!ZOI^K2y`)sVq*#3kL--uhX|Cc1pE864`tZOz>emvK z`Ga$eb+Zk05eB+secgi*VM+r0!h+f9ZRzQ;>FM!0y7!Gb+ua_jqXpE}cd2*(qrSA@ z_PO?*xr!!yYF?~=RxGeX?OK8uf3S>!E>lyNOH-G8O84NkFr{%m;ZZ){VLpj!vB2wU zOT(_8Gh59cwYEF9n(MZ<8zrUhNJ{4gM*0W-nN5eRHvwK-ip2NcfiLGqdnS)fQyTv4EF{u4WwNUopDIJgR%SpQGa{%@l2Wzn`_m7~W}Uc^e7M!!5`H5S4uAAE^V$AfC!9V$XB z6?C;qcd3dxLq$DbW#&9SPX+HTi+2~myKl$4FFU!Xw(_;cB&Kg6r4!`UN?M)6lEPcs z`Rd!h)V0^vwHMTx->GW{>)K7~+U4uoFa1(oN_fp@c6z4n4eBrQG~^>`DXPXFw%Wys zf5mEUAH0uBwQ_+DMr!fU!HD)Igm(`w&q0{SS;d6A*kc04LznA9_6YD=T`&>t_PiX@ z7tVd+C=8Y@dJ9|8yoZJ6>+g#!Nxd8(*o{@pE5UjVJJq};-IF^e!`JDeobX)a14HCA zLnMVEQqo8D#-HPU8xpghSL>RhNuSCk^AiVncLyq#E)*Aie*0&I5p=H6X$)z8%q^=2 z3lhS$A7i9lbFSfv=00Hr`*0NJ#c-~<-4vo9!!0b!YyVWS@hr}q#5iW$oOYx~p1%j9 z7fO}q<-x%C3KT^haYZbKW4cR?Gpkln$v6QoFwxOAfQ~rhFsfF~r|b0C;L-Q9ZF-$v zxmu3qt~J#yAvKBINOfM208{K%OrYTs-nb&i>-N|p7_56Y0(%HJcA>;E%e@BdsNHWs6&92o zs_umO!4a;GQ;jczAHnk9WO7Rz^7=%<)TI!4CO}a}iK7+aTylt6??9&de95utuDbhG+q>_+eJuzoMq2J=4moy*?9Auho()qsK_sLC zMY>GKIPlUPiqTj9Hj1yy!cJ;xQU6~Ax#f!PvJbhmMe5xytpRnzMnARtw2}P%hxTGv zMtczr;7|y=`Zl{dj9q<~UHyPvEkV7r2NNnsllkC(Vs(oehU)~4hlEAkJR2-|rQXFu zR`pyj3aUV2xP&oWvNRqoEPqZB@PvDkA=1*vd%2JIKp!uvkGFq@?T9qLg!HNfheFwo zyQumbyArPLG@i>?epqpm`sQX9%dd##M`QU@is0WIk?jzN6LP*99D=Ojr`{37aB*R{ zcsDgE(|A;AJRptdIhO2n)B;XNq=rzfxkCR1<kY5P3h$E!)e#KefsrQEi?AX zS$L)KkH3t$&5btzDij$Zk6l0(T|g$J${+-E)1-d!e)Ma-Llje>hyf+QLW3znnEVAev((5+p7zTRnp!j|iI?cx}`(g%9 zgSY4)7e39MLw+irOG{cx6E^4Hp9W7zgF`IsXOYKrkwv;27<(J+{FxdG+TnsiK-oOM zzo-yDJWNSG?d%k4cq|eW&ug%?q#K%xh?Ji!9UUJg^K>O8YmqJkm;rUAa>7ObuCn=@ z-_6M=<146w%0%eo+x_qZd!8DblA_bc&i{u$Wv-H12-bj6#e z=>Q_%nGo<2bM<0hmnXrD)&KNG{h4Lzjj8xmb^08^Zm6rQHuris#I8kj{a|wOP+Q!9 zjeV9(Mf`|t3Ghz!S--EAFmEsg^Sv9fOt7EsD(%g6W||H7JUcKQ3>W#igk9)*Jvvc? z)YhEiJHP?$PRQ{Cp3u$>8aafGky8Q)7|4|*z3qJVf~kE5_8U$AMR54lb!pw~$JR7s zh1(at`lG1O9AiXHE$RP@HeCNF>F~FrD%*2P?6|HqDn;0Gak)Y3QWs@oU1B%tQ=`sN zqfDt$M${+^YE(H!@*v&73P7+z$}0CHGU2ZXNT3^Bdovt<6hUNmqO3!paWkh2(hcP4 z1}K0WR0208Fco?K=epDZW_5z2S?tF1)Q%?@7POnC8|((iA4=c|5J`Y%TEST*a9jyI zzXYy9KrTY$5Ojm1>Fh?xoMeooEZu;IZUEcVWDk&?0dhA$c4fkSaJ0T?5sah_&4A}$ z;72on(F|m023N6U%X;~TA3N@dDl}!hwM3v{=jXEE-r0)${!?3c@(um~m~Q|9EZ;yT z-{5N(>y=^FqQ`qJQH zP$3=O+D2*c1dzLj(S6qwKhYRFppPBWKr~(l8Y)(j%DD#iC^8?4Y=RG zbvMp*e|~H=T_GFLo9{dDRQ)6}$rl`cGk6uO#SdyQp{3{~a)a1sgV=CngTC?026nH> z-&l$Ap;6M5pPJJD7eS+>VzfIir)8g&*s0NJ8pzyzKk2$R+f{3*+iL-{^I{6T{3P=D zN#qL+4MvoJ2nL0T1y``GJthJFW`+`ectgJ=xvtzxTeL{cG(yV{^|N zk4)oO>hS3MtDBBG)fo9iggM#8oXj>SPe6x|z31R`dMZN!deh3kKogDUMdJn0BYn%- z8Y-?(fN|Z7qDTr`7_~%76P8dj&d!ewG=2XGwZi`arWGCnk`+G2YD6O!hBr8(*!ivF zbO<-;*yMQ=@-CfNgGbj~!;!`to|d`p-l6W^;co2g20(@0@$TNw-Mv%Y0i1jn_-l_{ z;a6{!%qYdm@27p#U>+CFP_KBOUmS0RSI%wV$!(CyZ2&g`mfK*k37Fi5?J51^H?`fo zAur3p?}#Pa#5ql-o&<^ z-do?DV-~1uons23HUFX>6%{2^UfkZW=Y)drERDbGXu-|S z5JZ!-%n_(_O!@z1vbpo$Nv6WUjxB#B{vAf<2p|nWpJRT8)bBx);CZlEb41`-=4T#u z#z8a*WsZQS0VQXdc_Cm2FGXJrO#<^^Zs~CIaVatfW|padCh_FK`2UdUGTrZNG-Mb` z>1cd4Lh;#6IhUnB4x4a8E9}Ir^e8bGtQ@0@_xHcIf9~&ioL%w?m=f+%~oQQYIA(L13|A%uPPFPL}+~R~vKAW@ka|;%8^TDf+x<((ybP8EQj1ka7=H zm%4mM{<7nC_U)a9Fs(4~!99|Cg*ZiDIS*F(r<+6mgVW&OocF=-05xCF*yAd|3>=(n zd;z)Ul*=wZ$V&sB=rZ&EbfYjuUmQ)^ng_EsN1Xa|$DmfIp-J1*fVnd@Z_Me_D_L8h zMLthP`{$WF$kif&rq&+iMw51%Ba+g9rQ4td7klG|YNAQ5d9eTQ5rfb{M(8q*IqH1S zLDHnJg0~Ltf)9p*q?gPXZi=)>UL54x>`Mbw&oQ0vf=v<>>4SN3IcfdF-8YTo*wHwO z^ojg9Sa!eHuiHig&D7{#rFx~H^FNBsEUlPXkn=wrVxYw3K<~u4buO{WvCYAY<>1xi z;MMHVK3>Q}l@`Vrqp*lIN2dBcwlhHy^1(w6_!IOv7&GbQW@viTFK!8W%j-5{6Yzq# zCy@8ti#giBjy`B(jIu?Dl>&@E9&i#>T5Qvs?)nNCZ|QXOy5}VzTnZGSv4_c%;r|Ezymg!s!8hNs#`zg{SYvbe8 zxJV2FM+cIa;K`JIPT|959&YTdox>p$gpU+v`3kfVhKS>>jqNcAF&dzXV`jbBQ&~o1Zucdz(eY?^kNe=wW9wof|jc5-rVha*;0+VR|ZhDhIFBj z_au`s|Fpsi!Z;{GxUv9V@1)*Qyb%B2M@$3NCDw+8vPb*~V4j|2 zpM^mYNJgj)3yMH0|5pU^MF5kw2=+BmT@Dt;K@muWN?>Ik>}#dEXm85jFmD0Gjjd4QV^CYXIe3ES` zbxGi1&lc`o!=EKS8s(4N42!|CE;}K_>wx1Am^2s>hu|>uX)e&!=T2mSZ25O79g1ez zMP92TzGn{$LP5w(tvEdk_?IWkfs8nQXT-%u029N6V)T7F90r-@B8g*R7;$b*fR6l| z6WJ$M{UsiSi!=dRMZ4Xlp&jrE9LwGraU+W9afsNb$6+MVT#z^xj1fn6Mu;~7FQS=2 zjuP#4$nK~kDOG|fc=$AUBBtRbxsX%tiO5D}-(VQ}GC#~E=F zjll63X3%Zg@<&12Lqxk(S;)F~QOLE0xftZ9nn#WB9dCr35t@&g9!H3Mp)?nM9Ltpv zH|vZ@hRieJ)XAMT=9CGn{Z~SR9>@M|9phv@T=s5@%Qxd@zX+tva1t-4e)^%0R|He) z0W5|&tiy@8LlYLqDabG|k5m-1()Ys)iMAL-#Y^q(5+?3A+ zy?%g!K|9j}qL@b{iTb9T=AS}D^XD5eYeaz)U9=8RdBAkqORQJoMCj9mW3UR_bFi6u zou3DUh_5F2oGwQ&wy}>jQu^TLrFT{-6)7>ioi|ZS7_2cY_YN~LojlynwdT#~b z4cRioZmzJIeto!fhz)(P4GP^SE`U&?)G7!iN?qSbl=?$9%8s~iqW~|&v_PDdfXWS~ zPdgRLl;TjKOsW4iyZClcFMSfgkN%uu@4c{FWxOole(r%qkYt?#E`aQMs&%Lh^>`u1 z)SeEn1bkOP_FF33sL+6scX`gL)t(NeHr#U+42v~~SB@lUJU`hl_xs=U8%>yHQ=lGz zTCtH5aj(__UX_e4@;dwJ#_>haTpr5HYoi|CdLc74+J25rsi9hvY^c!PX4}(&3V?Z? z2`+)`DfSf2CRDFZ%&5P*ky^j4_U4I25NRoN-v26ds-8^|v7z!5U`%c4P-=tqRVMfe zWH+)YQB-RM8|wB_03K*6->H_ZViS6}?-b9|odt24rGT{$Q+@)pe#53jY|vRmx@~Rf-CwO4V$V$rNDDlmcuo zCO8hV+t`#uDm&D6BUj3HBUj3HwvLVZi(E{y>bcL|w;f9Gk!Oa@MyEZo%m(gLOI`xVCRnl}hU|eRS7gIqokb@3fKm~Y)>_W`;HXQtdsJS>_PqY% z_VG`7-{UU3v@U%=(7CAI`B{CjhA`^XI(K`>(VU)V z9Rk|adm%Gksh0i%&+g4FZZ{k}Ob`!b4bE1o;nb854vr(%o*zX&tIjfVLuFRuKP-ts(wxqv6{|skeCtA4WBhqf*IHM&u|7 zauhqVdyF%sX}BBB%xmFgRJ*Sq(JFMpg(a&sD3 z9gum9y_GW+DVi8Ain}K&VSmdispP5_P1e`!)J)bJ)L-NaMGj&T`Ne9gx#eG*xWZJ~ zs#0^%+y4kGC=#2H(qbaY{&(mgPF8JX z4{5coHd=VZ?>l41%4+{)+=+Aa`j7|`x2!TOsQM4b|4ggWeH+sJkX!aNENJPkQ2xe{ z&_leJ?O69DFzgW%r>MWb@}gH<}5E50_)~M zVD|#YhRUK@E<6ROZm5CF8D7gOr|oKkm%u)mZ#6;XT!~rlxH6*&Oy+I@{5ndg(1U8Dg-Z^YHu&{OCL|I*$yUXL|`C z4Q)$G;A*Upw*W%$La0ALePlH~cvz%p_hmkEv&m(Gf~W?TZg=i0^jWtb}%fkxRVG z3rzAwb!}!FF%1UeR@wT}&(TY~%s-jze=x~MiqnUP5BP~^&Az~RTADT9+}?b#($)zs z?F7elf}J~uv^s}4I){ileSW>4rF94X%w$SeuJ(u{itTR2)JZcp`NV8s)JHOlN}Y-2 zOs3`cOv_13_6bb(ZNnVz` zPw(~AquoSeBJCnEq0yEF`Vek3%=WaQeOEC1tZ`DRM zKPVIozb<<@*I%DW{-cjX&w$6M75! zu6trvsiELs$UqIUO7b@F{15QYAK;Haz+pdzUj7(T`%wdlN0P-O$K#NZaY)lRBwgI* zM;@6Yr$4NpEfu|K@h&koDUQH24=0Asyc*c~}p^Z!AzE5GFJG`_jrgUy$&Mf-6OkFD;?n>5=LTy!830sSaHYbG%0H!8=f>YeHlTE6x|ZOLRB)dNash+QxkeLxm6sm)oR=KXZ^7m;&_pX~2`;?{ zONb!9SU|fYBh{&CqK$azZQg^Y%L7cObFn%3*Q5P-=<$yPmy*B|0?3z7pxrsx>5S_L zmPZyA8`{6H&E=Cw>T7wA!f=UfQ2yK3asz?yiC}Xk;OY(3(MDYKU+)XvhyhnJB4zBc z90OnhUW;Sn&*CkQ${XTh82NtmOqv#oKRM5~LPR9Ej|TaY1FEZu%}Iu<6H`a?-lTU_ z6TA@#&ftPx=5!~2626II64ERb?AW8GQvD_vyG9D_z5-Y0RJulPPHjR5vLI4ERBkwFfsBAF;tfvn?nRw2VEFGRuS9< zMhT(a6xbX9RGuo@fP}s&(ej#<=ipU<%2GvpanRpU z5gZ8t>k%Md0=@xAI(Q2y2d~zMEXTmxIzA3BBIrMdT`}&@E$YISP__~dSaf`fy6}w4 zVm9n^~cha^IoSyupzi@Cxik2(`Tm=h2Pfw{A)+pG+#Bj4z*j zTmC`oIq82j(GsT8^x{aae+4!of&No&kX8iHhs4-K$_u}@Ux9lFk$(s-yqc*99uOcu zU4pg&eKI%?A4js9ilA)>IE@tQ>f?Qvu}IK?SBE2c=OTn2Vx$$>g()dKPnN?#j4qiC zUy=+76Y(LzzOuh-v2hvBdGjJ=D6xrba2`b}O(TxvN)lmF^HDKM_M8l4BEz;5;BI&HEh5S}KBgU=&y;3T)yPIFCG)rU6GX#s8Hu zn*TXvlCAE#aU`#(2tEe_1ZEOI&zF-4&Lcpjc}+ztwbEVX{Ki*hg#@ep&%s*bXNO9R zD(E&g)#nKBUxRv@Exgrd#GW1sA_wt zPJ5`BHQoPX(x4(Z8Uo%Wg8EX!uedV=x*3vS%L(DuAgT@}6+sjMvWNis{1UdDn5qLv z6@NnH3Nf~v4Q>tefz-@=1*RuN7Lh`IgFd#US<;dTTuUFOfmpSbD1ns@)%ES0;l}y9qWgm|c(tHt@MK(hFxmf`R}QUV*Cs$w{CG1XMGR zXvH*o0|)z4Zwa3lmlZoExHT(P2b3fEuYyuUNCzTlE&GK|z;1hdp<)#gGT<^CKi(`) zj`dsPS-N1+V~Hb%mJ{QE<-3*$u)PEwCc>7p!mUZEI)K#bR1q`|0hrU88<_~2EMsnBL1H{87gEBC$uF3PRDT5;4}RE0;g6q6batE3AIl6CtK98Q6X?T5Ve4kT2T!Z zQcNAILffl}ZNE;<_4DEG1Npo)ZfXJCU6yk1O(&h_V$jia>l~8L-IMb5e8KeLOX8^# zOb#MYbt`Pt06fW;3x+9>y^;%j)p$A*8^sn$r+8h!kK5pliilSqt6}W%PVD1ArZ>+tvz{-Lcg@225nVwB}!5+nVM-`b92303-i6V-k16^2Z zJt^7&Mnj-26fIEzsywZL!7T#{7342w--wUkE>?9#S^@VbS^+EmPtGq`w+w*lbIV}m zNs*Q+azY-m{iNuOckDIoa$NUABs3Y?!r#(vahtbuME3T;TNK}IuSDoS2jIFjz;pXP zQ;iBm0*%uLM%AfMf+xxReE1JEKoKgSL9l5TEoFiwWs!ES(Ed4WVKogxNV`Z|7RevA zs4m^Vge^>=K}2g81%3PCI727oz;H-W{=%>Oalah05SY;#sn5V?ke>mj`)_WLW&W$l zjvUg?8`|Fy*{{yW`_pu%%`|}qLB?kwsa zp8-g_$Swoy=?2yRgDo70>_2GXwWv{}ArsI{?|&PquLv0N8sN2x(%yq5MUZwDP){%9 zL+#rkB0ka#>O_S(lhoHSv{x^OdD94zewjY6LtMzh2T=VuY$1K5{+;UxGH!!A8b!Z2 zk(Wa_krDxVTJ!N5w2JqjqM8q02DHA)%OuYfGj+rkPQdlUs1Zh724C+N0XIjIOh`L% zX#Y#>QU*11Be7+38t}o+UzTpuH%E_r_>qs01{NNrR0!>RMRHe=g$&UCo7h4FxPCDe zf{fF^@m|r$Yp~>1q?~dRBh7r3k=6C8jjH7|@$lJSVw~75^{!6xH&h5Hrvcr)BD)Jq zyX(-e5V$@t>cwgB7?>Xp{(2SZrvmq!XqFees72x6Od9C4C64C27F|K|UrW4(wEOEh zj0$m;)1dfXQQF^WVEDrdFKsyf&)P$H=Ifz<4h`gyAfdoxk6<+Gy^oeN7PO6x`&NptN#(34Gfa)Nh zCldI>2TQn}!DI*1$znQ91|P>y?_WpYcX@};#OvUv!Kl&DlO`7KCZ#Z08HzcVcTFty zWbWV_KN*-dQQv99WNXOhCPh=jM$IW6u>YKr36os`%_Zt=eF>A*qL?#z#-c4xSDgC5 z-G#Zc$y(x4TsMotgY8In_scmh`Gn)!J1nv8K@?7W3uQjA*vp9n61jCqw-JO2j??xPrr9xD5=69bRmYGc%ONF{On>7J9g#FefEENu8HW$Rc=i(USTG?kd zH7pg@;x=;tQlf08+B4m*@CRs4wm_JV2h()7J3K##ve<1bRTm%> zXO?1wkJ2xyE+?g2&pJPrXHo)je zwf>7Zn>Rd8NHxHNWjYr}#U2RTUl7AS=;Bm*OuqT>TyQ?1&1wVJaZ(M~Qd?KGPvcZN zAm0RH(+N@ykUOQyj+MaX1DxKGwX#n4#-v(NBuI1)weJ6s{?3;zs#F`Z+2$V8MPnI2 zA1?o9a~Np1GR;(sGG?3J2`~T#Y+|~EECY6#XSO$6_X9qODy{#cxDqPkyG<*n5)k09 z4ZI8BxXUu#fu(|+OCGQZv|VBQ(^A-!*F#M0^ z0j|K=87jyza8?}4kiyyML$T=zg_oX--snfMiI-6g@MNA|o(9m*RWDTn`b^UZ92F!^ z89T@Y+Vl*DGkAWV0j8}D8>msFy;!jfz%$Pvx)>p*-}jefzUzGRc}H?YzLCX$DsyY1 z6Pr7KfO>>Idga4Fx!yVzS2`8T|q!bwi_0ji(jp-#Z8d04K~=q4xN@y8uV3bw9|Q zhV+Z$wV#5xw1fulG|q{|xtOkTrq)sMGlsP;mSwU|Eb9$j8-+EhbunW?zll$-^_nvt z>EGv--YmaKiyXchf4+|O^9}&BVSrf`(WlvpzY7`oLO>-27;~;7QaA<9e;U5(5?n>} z{9j;lzYJal0i3!3PUqM_m!^V_Bf2>g=F>HNbwdU?H|jw8BWwFnj$IYe&ii5lQvebT zoZ7FTn@o5i@a<&ys@3^!tUr(IzX|iF!7JVK?Dq$-TY>rGz?|9PtCzViWCOTa1GwqM zfNaKqY*m14#eZc(0W)I(|2P0w+?cQq=w<;vu@0o6vi99Tlsg%~tn)K~ux;*yRC@nO zU}*ph2vniRfwV^1IHc%&bLxm!VDjT>UZx@AHviwC%`bMgn8mfc=j?so$oJ={P%5F2_BKC3(QfE zcWE3_J!Ifqyuxh@EhoBqaXoS`CzAVRMfykTIU#A+Y=(y9bHW2H0~NqiJkE;L0oX+u z9hq9iNoR*I9T zRp|b6pwf%o*3!McBA)86^wT55=KUC(byQYGQ~{gSmTzp^tSB!!KQ;D@3V^W~34VP6 zlZsM6|JmP9#Pw<_Q3}^8p#B^z^kCm=KKu()<&lw}JjQ19C~G^SfbIE~FSKnogaKJ_tRI$jnXerL^tiI= zxcST7_Wwou*-xr%zMDIR@9i<0mfWfP!v}tXwr2(056|S6%-6Ow5@%>022pY!D8&8Q zl1ZHJ)u5r|qWcTO4Zx_m`BN15DA^YJU2U7^mQ8?11Y6MrnJ4 z+FUT6_WZU4Z4(#3rs(Fo^<%p!H*!9<5YSm-FJNIori-Ok(M?Y=|Lm!v%b)G~gpQf} zK|0K=XBN3!iuRRfUlPs}N+R3$*JbGAguhs=F#%a(ZoWT$9LIFyXFCsPw1~z^)*G`t zUsJMg#J6%af4mUI*4cp&P*qL zFc{?+akdL(3tcbGDrG<0JiJb^3Oe&u;_mMpwa07*)zzZ+@iW!T+sndArZ{v1BU;%$ z+t0tZ5fyN}x{5!p+gvRP0eowVELUy*0IYxOT8QdE0^F)GN?`^LXD2 zEK!e>9?VvvSRXD!Y)7N2(a|nP>8<^;BgfrN#dIZcW#rGf42^QR{pw-bcsQ z2c=Y;HNc*|9uc>Bx?PhbQGX4S@~7o+1`_NbT*_7@?A$vC(Ww@8?i4bYdDjs33&RF$ z+8SIBMk%siUz#4=j_z+s>2!vvj(8w;H-l~0ogGUR0!~f`mn|y#=jg^7w`}+0m4nCY z#egrox0X6Puu|CIH%>~N`zt$-JDsE+iU7ZmBEGIv4XLVj@j2NVkG9^KFgZo~Qgdi| zBTZQdV&vxcJ&$>J4tsm|_f}^D4_F|3-(7bi(GB3L;F0lN;jmp~mGq3SOmThv*@=_S zZ3@&OzPj^E#i&rvGW)@i-NI-!Wmj;eVYNT2HZDuKF;)>3U%A+^+IM`oc6boi?fgA( zbG3hGYc0l%3C)Buv^_K(bEJB}tWd({sX8)_IjY+Yvfxu)*pyLmXhc{+XSb}TSy^LXR zlx1xIE7t3XJA`diRaZ(MRd;G)y6aDab~R5A7lYjC#~UXo9d?&p5|NFze1Yjlyb&n^ z=ms>4_D_TKXRyaQXEPQjCmSV!`;LLq8Bu`$^P|ZA9=E0aI*jGXw&^@vtxtU=-KbI(1EL!A}=j|H%~_F z=7w90QjB^Hc||%}I-sysn!tfR`VPIZV3p}KZ`JDcQ6c1NqS3Ryw44EguV6c5AGDz5 z>8S8mvM(*4t4!_CkBgJWZu}??8H?qq#?5&#)=DOgdeweSG#033z3Wu`j^4lsZ2JH* z!kc$eRf!weHFA91Kk&%$lHMrq*RM}Uk$qLJ;|^mKGx-A(I>imyBNoSfH}(r8H=fBS zXB!D9e>DiQyR`Gnekp&TPp3Eu7z(yS2HTCE6gckXSB+@<2ZR&|ue+9X_@1b9YQa_#3i?rF-0LPmswhaCoEPg4dKgjM~2EBU-+zf~DW zJ;HZ+6%)2Qb@yn4Ao!(nke%|#_f@X2-2_65qweOBh)aFY(TGc{7-Bo>RW4#X`5gL= zJITg^)nF%NK|}J0J#dgJ{rY)i{#rr!S4!HEOIKGbUXKOoj@Vo8TTVs{e4{K5VK*XH zuCzS(=yX&6 z?s8;Tt}e@#J{xWx$tO^*#=K-FJsvyV+SUEm1qfm+pj>Sbu;-G}21EIuo+H zsdD5qnLA!4XIghHX{1>8 zMYt)HM;=5MmqL#Mke<&R*u~tAs)wYU0tt;F-+H|IKQ1L+tEqNTVo-{QI)c4JxZNs8QC7iI+**?A`b1l;u?eS|A zEj%9V$3CPjQdRAIpJDB!8xnEt4Ap_R#$`n~uTG`a&UE_J1)kQNyiFEL98J!ZUOip= zRs|VEH>?DCy~Q<7j1{vRW|ugZWS2%j0?ibIpDSloKP%bJ?_b4qc60(iyA`bD-#{2> z9*2|i4)pdtks-;X%-_r&?3&kEPRGxU{p^|B?#$mFo$A{gnl;KC zb+MlvR-Z8jozAvS9nJ^8wCqj5H>)0GYf zaTRECSLGn!rPjFYH_G@knWN3%v-zVG?gwZ4ZGF92*#@}NVde26!#Eq|li78f(Ico% z&fKYQqu=4o+H?Tkc7IKTWyl37Z(>l4R;=0mo*ife`PyfEwtaZ8auD>ETL+`RIu_o) zdboSK8JBW)CAZXm;hN9>R%`Ep*IQMb?K54>W%PteZg@(toFysl^zFxHnvgI@jbB$N zr27K>HXe(|Va`sg{A691j(zWUon#5*?+N&gLN!*YzL3y9GzQn<9n-Q zpO`KelD}RgUm4GpB5vX&ZW17F;v>GB5b~i-h02;EAU!1bu$28WYL|$5xHYPO6+TZ^Az?k>8cP5$YiGx?ytzZ7g~8+1_xAQ&*Pv7kN~+Us(U*f|%gp zn`krQTsx~i(3@r?HTfcZ+C$A$ZrVf6Y=;cdORd;LJ>El|*+VV42ru+ddoGv5$@(xO z_LJs=H0{=o33E@m#0PST#&U_z%Q@tK5B4x$lb}_2--|cQjfTl4HAXv*z-$=JK=V3bN*Mvc8pJ zeJjM;YbAHjyG)Mt?E^X9Y7Z^##Z-yKRK3MitHo6E#nhI?RL{lK{9dGg_L9d>w#C%( z9#Q!`%Yjdp26>jrd6rV2EUEJ?yYnpHYd=^y_q3_eI-MNba+@`pmo=K7HCm80nv)eF z!-^1MUG|WB>I`^7JdiV}_9)U`1WPW0^%ud`i(vfXt;W?asAU=VUTS&EuLTRP9o?(3 zxxJtVohSGESMLvx-^X;`r!^$gHDtmrwsW`uIFyR~VG9OKYWeLPcKl&q3dwVd#Ne=Y zjwpe!F9`B!4{Ivuxwroxt9Eow60W-!(o<0pw?c~O?VhzCS@$1VjyA$SEB}v zENV%*+ow8_qWnJ(a6)tc%*tn&U>R$u>TtIuK= zxt9CDnra+~QjbaPu{D)`(%pCZWZh4g)LgBpJd^Hb6p#mHNv*n_LkP3V^%0Uk}n^XhoR)TOUM&T#5BZ{nmqqmoIeIk+`!`4 z@`rsYBu^|7(*PDHQXuRTg1kq|dZX|k72FWyiO;O#C6eww)hBxhOi{It$0ptF(I>O5 z6N8U8EqUyGAx|{6j?Vya3dj=+#Wc#>IhgpvxRovSvc$Br#i4sj0CBg01@Dms&xtQ` zu6>!ZS9MtL%*Rf*b1d?QNtiOJo#!T?Q~Vfz*dIM6?#I^gdP#Q^^~u_vFmbzD$9p8* zO)DY4i3j#1ppyomlN&(!vzP{;QyYJn6~LfIOasvAu_0N)KRR&(I>`c*pD}UkTgO`^ z-4)X#duYPMZE79=Ss;w@e`S?=Q6_Ym{ykRezrku@7^qubdSBpEBW{o$#m3O@G}xz| z(w8P+*H8L&#c#w#wi8fcnLkXylnL9zT=YvduAKv=Pc~d2_9LyG1Hm6Ah#)WjDE8w& ztVWoNv>*I?t7l9m|7i6C(8{Zjy!^A+4?wF9{xF9x-*NiOT*r{C zK>qKB|2Gcn#5Bf_^#1=G(n8u6gsff|sKYanBYHbEQLZSsYaOiHr1?vUxfO#5Y&qpt z3Hg0fYY}ZjvLgf47+AYGkm>iU+s%PJdjGMtNFje1P+`u;fx?$tBF3065T@4fcSHOS zX4wCo8SOub72}lD1QeN=w03i#tmPJoF@lns{)64^e-vW`bnPw>W0Xv40?O_E=hh-O z0E$71#q}Su8go96o##!8y&XE*t`uR{Z~b+x&dA5Yv9_H4|IpX_C*S`EhXuQTH{yTq z(1a>k4awe?h^4;Z4>SBUOf(v9x$RJ+=H1ee8(pT1Me#@GbFo2fGFl^GB`pUui29%S(3Gr5=njcL%+W z!ys`c%8nD+8{;csChNpbhLny+h?L-DOQ8tGqO~ehkQ~KQ2k`hQqJMtM+xcvll|)E1 z9};}*BYPa!rMhH#X5Y&}0X%SeV4rIkw~Rj+6_&%*V;)Z^%1Q6<;5Vn{mfvb3M!g-U zyFP=5`ujecFB-;)xRQ)q%QFXyd1L3- zy;m)LorsT@W1)q!$r>t$5W4!{o__tb?FK4?qSFoZU(=7*b6*8+8b>c zolBW|a@ONEzxtBnUDSFU@N{V7)a5$yCdx60aiY$9HG^({H}G^cg>!p?_1bHe$#MVv zdHl@y;hb|6YSe_+atJ_!RhERv!J(J{<2k-rH|RxB;9u(bhD1}6nzlWSTG8#j6d!%zHZ0P3=5x|T5(FZ>N?zp|XJ`X9$j!3xF6TgbF#?0l7F6Za( zPB@>=*zH%qluJ(b@Z;Nuz^~oi(8e3nK4nuvTVPqVH;iU-U#w-Li`c7$y;$D zw4=@2AHx1wemXW2zlOct<{oDRon^tC^R`>ZtA7o<+<)7#W(ZlIIJy4x&1i%2IF39e zFJ*=}(BH9QdFm@}ojsUUDIp)f{W)7HlP6$rdgjm$_^`}<$OI7H}*#@5za~Q9g|AP;MU9(=>JklA1Bsr`8VY^{k_9QCYFw;cDH7qG_T!i z9QW{liF2quKxMHTdaJVgvcTRne6fhB0gpRUZk}ikR#z`~Ebp$NgApS*KgAUe-Oo~x z?GE<+`bPO^lWPv*YrwBMo@{S#ob2y*IPR3D_kJIk5}KG@>cF`KyStyd1WE^No$dRb z3{4ynsOCyGoZ)Bu7UFbpr-yl~`Ci&}icI76W2a-=k8OMp{=g{RkC2MR!n?u18~aAP z4lf=nR|_kFQq&mIHFLavP-e5Utyd3jK_V<($ir96SSG%AEO+#8HxG`iP2ie3FP?!w z)}-$^+&T{J3zocrrhN>FaEH7`gRUHUdTwv_&f*`$;Qfz%P>1se`M`sV!Euv_j;Ti& zI^0CV-r|qe{%!op?spw3HFDOiS=idHko|*&_+jUW;o6hYriWF*@|Ce4O03eZEk7>E z+b_WyIatXy?56vel){9+3$M{onx7R+H9N&&P>@G zw&Owic^#vtbL(AulW9HauGi~5!o;AmH!Evo&gpyj(7p6M&(Oy)HIy+q5~k)iO~!~n zkHmYvdVFpA7^k+_bH<|=f07RW_2-fx%_2J?jqvJ>q~34BiSWy)Z@dOQSImg(P)(>; zYq7t9=@tz6E96I@mvtC5foY2Fk>4JX^!aPFqvp@6Uy3%ca*=wrYD_!Byg7Sf6JFye zN(oO`xojIbITjnaIGE5p9KHcMD&@h?RC)ppRm1{}9m2v{x$wc3dV1b5QPZ)yx;Rn( ztMd2NEn2^PS!AU1WX;V%TI%gAFw(68%@RvJ36Geltyo=DDHqlAQBgBhQq~z$c$uXG zl!tXj@&n4U{@TJC7KW2JDR}AST zTrEV&+d6cY8dbIsSDR0>sCG;_e;nNt@qlrxg6@v6h88N5SVvqQV52svQ*&9MsT+NA z=hyE$p*au_pvZhVi-m6qN50%ohWC_4-a5&G7YavxIsObED2-a!GNop16wFzZqE+wM zQ`ZZyz@9qCG%v;KVyiz+gadL=Pr&>@fD-ENz*{E@Y+HLT^YZ?bF7stsg2-sA&gFzCQI3@SAN zgYMRVS{%@zmLe4JZw=@xh~A5Qtr671^9A4v2|!;<@QS%vzaR9M^X7Pt9&0Y6aN{2H z6VDhZdf&!keBDJw3dnIvwPx-U*#Lo!Jw;DZE<0iMHN&znY|(mR_8Fq#gR* zmpXw+yPEhfTg&hfZ>EH#I9;o$XSC0csMP5dRTI?Xo>(2w&_3TLP?u6vPFRi8v0Cn> zdA_}&W}qyekQb-@CyS47#Bz|`d2OA&xwXF&KaU#Sy}WqR-{0$JaVhs(bU$XP7vG8B zW*_IotdCS-IK3PLPdt!*EE5ZZTr0RPd~d1Fa&YCUW!G+cwf$T5^aqw*!i`01N~K3! zU4}ZrW2QR|IALna)r7&(e01r4B>$4RFA6CU79R8-UmwGd_is1_ z57+lQSNLaiYGFjjzil6G&&1QRR~u(Slrzx{UX5+<)4xu-m>Hw&#Nm$f3N>is4k%DX zvj;IE?!nc8b7$54kGnYJh6g?`pH&6bHy&EL1=V1>yAWN@6=7Gr>oZj>#C1nIOx#V6 zoRpQTbPuLqH2A#>|N3CGbY00gOT{^1_4at{(Gl?1QVhJ(siH_>Y>d*mEhYLWfACBr z04anoEb8H_;lKUlaSGD_0)JIW_f^KhQWQVLg!?Ch0L$r_^#>vXk8t}?MDTpvS)=1A zPPinijuO8$i>Q`Ykr{G?vKhp6ty23v*E_10hpp)jttu-yy{M26PCr}2Fa^k9C~bD8 zVw5vlC~X!o&re+plrs9K3|9v+>lQ-=|Cu{_30#v$VW zK0Ls9k&IvR+Z%de&FH6pL}nTolz8&q@li+$hs?(0?k%n3+ds6Q?h>`|@mJjWp=~gi zPMuK}bK{_O5I=|q{!VIrL{VFhP9jsn;mjcQKWBR8Y->#vPdQoc^BMVhe=m0YvpkNk zu+a?U8}`3r8gSsZXP+K0eQ>IK!@u@0Ut)T8XJLE8HK^gaVq*rJRe>(k@y|&(r9sK< zS>pcPFi_yZcz$tkG5dOi`)tQq@5HE7tt0nh4sOrz)~|d;>IZgx0}LDx0>J2ItZG5j~Jo4`IYUHlUM(*$;D_%77_LxIoT21=PqGd7&U_kfr9VlM z^20C2jY7AVybj_FFaK1eEH`#WW`a#+ot4&Xbh6_K&^RaOwHKr5^KJucni71Bi{oq- zhP;K>Cm03``Jsxz`*3yNwDDeVWq&y4@P`>{&+~*0@t9{g$0z}9t{O`6-2B%*pA?D2 zXXd2wyDuCk^V=;op70NHtXP6t12c!!&~_Vlq;yX=LOe^2^nPw!vj#&o-O+`VgiF4SW%jy?V!K&B;dMO&o&saoI|_*eTUc)|KJx#sl=}VYPBzFyw4! zrfW93yv-SXwmR<|ba<+RlJ|BFKIrnsofS-rba*Q_o>lMR&bpLAzf+e-Nh90F+djg_ zGO;18#f3X6YrGpZ51P2%q!T(RSJZqnwy1P8C1@R4VvMxBpE5?B!6+@_0+S65_&Ua8`0K%3r#^l zG&2|;%5&_+NYYp;)S7d2_(}xTp+qh*%y8+ROQ0l&V80K-x|lgGgNY~nzR%dZ#?-(sjb8;D6-uP$$8JdUIW4=8 zxst#GGi=SN?q(I{eKNB54h(A|NUWFq=oHH!P6#J=e)apN&l~MPSSSZ|Jt0{#&M{JY zaa{54+odsG^Q>5M(R+2Bo6y7SO&bGwX(YnDX&n3GX?m9$5HV4+Ww6VNw8GyFNkXfI z-b{HfyWR4u<{dKkMm!rqoSAZYpHZfc7pvZE8&A6xIa=H{Q!rj9@_yJc$)vSSgld-Es7^VfUA>x7AhKHTXDetZ(T|`0TOnqOpFy^BpznRJ1<$ z)F9!&IDFojE()!$b!zZ%t2*izj$iiFz;)kPea@NgEm|LPB^V)dTzqr0`f3}F|7I|v z=D7I!lyee6FamyDoH%{a`m}G%(}Lrd3r3V37fn+pp zJP$f9^$OnDX%D)vUj59Bl1sX&QC*AtLivcHjFQ9|54!AL{p%Se*C##b!h7|9){oyH z$?Ve)cxrehzAoy)5wTU$~zX^2*j77 zl0@K4X1~4`a0YBjXX%N^zmag6_L1S0Cv{Oei~NL-3=;r4IUaP$J^JR43{5RGUdiRz z>m?69)K3zap@^)TOEpI%u58>Kj0hgT{azX*^Sqei>31fQx684$-%2^MtNXTmPLJq@`^~zfTX+(t+21_>Rfx^HJeM0OT^?qhI?xpN^&6f-=djP}NS|A1Q6; zByw?L%--+#ah&(*^rn08a2vXO7@e^LX*-6P5s7dSr)Clz_x~Z-PJ9*l{M&+rxf`2^ zGBx-X4VarQHRGCK3b|k+NHB#$Fop8x!^_`oNCXI-z@oJlE%UI0+u!P0%mc2MO}k<$ z1?_i$bVGg`nY1$n_bfxJ!Q`61cO=B761)t1c z7?eTXFxX0SZ8H&3Q80%%_=*cAsulG1ZP?;)ljWyN?ptBQSDNd|EqIzi*FIf|{&Pi> z`)ah+Rhnnv3AN4pRu*biBK)S{6B-vx{cDS--$ewhQ!~VWYTbx^0YY&&)<`=uHb>g= zOBZd9m6f<-;#-v^-!l(*I)TGTCdbO|yJOORf$WK-i(JRb=G-tw@0bV3oxs=JFlj9y zdpYT%+Oe`zqZ*6OX_(|2<^ix1_*tz*=N#-SxpWavjm6_x*jL7j)}8Mldr>E_c#Xvy z;4Fc3(R%<70Ph6Z(@Ph1107dP8Wv7s+b2yx5VIE|@fc8p{O-m^y*$Fi;$`6t~lm3vgx%_J%;(VQ8#u%KgF<(O4NZFh(Nnpm8w<>jE(o zIt^*KU|?+^W=^M}jlXbjn5D=8Qb1?^0>Njj>?1(vCn%j77@NG{53q&<^8kb^PD7W0 zNJ5#VB!J-ma>!toGQ9PDav2CbnpukWf{yC!1vhFZpVq>GT8rKe5Ha*?o*Ra~3p6Sx z?I0dckic98assQK=a;%WCixf0m`K_I$gf#9jKMqRDuB~fH%xL1$QUrM9V=H6hp~n* zd&xKrxz<|nwS(lzq*q94En25xteMPSCQd_vcgD&rT`{rWLGpsqD_?+)J4PGE>~(#t zEY}UA9dps5nSo_8O0O`Em8k+}Uom@WISp+BD048@XlAdMK&Qsy0RS1n?8V|V1eBcE zPLRCPg(tpmAbB9fhrk&2Vr+k`jAjy+Ng}4F&b%zW0wBBv5D27KxB<5_FxJ=$4o$Td-r-R4#vrF5Brra8K?hJA z;KncmtC*}_FcFqk_!;@M12lsG4Mv0fuhnMWzb4HOGKOFJ_DQ7fmq^1;k&o>nB|k)} zzn8s?w74H`(L{tPI&9wAY93o^hG@E47&uxO*;+VQTi97xc$r$fFK@oCA6}py-ePxE z^BpZc;+CMnlT=+DMpJE{4{0I`LyX0|^)}1_Y0s}~XA-_p`S{-4z}3RTaq6_E`3r?g z_qUAp=JD`wMff`H7lEHb0zXUn1;KjY51-349$G|3HGc|i-t=q!!lW`EMB08uQ;+5a zJq>*!&rbm!mICgWB`ucaj$7PX2(F*PTnbOr8E1>=XQFLolGk}ssMtP+isZf$sSXlB z1h`cCZLB(+UWrt(@a4IB=~^Es`4+Pl;zo?XjhGT9a$%;J0>&{xMu$o_Pe1+H?{=tQ z=OP^XN!t`n(h&x&&7RluDzkrSqgf_y(xXdSRiSOt>-KD@Y=3A$OQ`SHP|Li~z>23X z&L7PQ@*Wn`#0b-vlwKzX5B-n~W#$Pb%$HOPH}?oJ_X{+?=WA}_Y3}4|-aYayrsrE| zQpof0kVVgsTI&#Wa^bh4oNtwB--Z&u^+kPK3H>I9f5p>I_zXc9O-G#kg;dbsD!Ayn zNG_E~9YvW6llcoe^B`(-FOa#x6>|?#b2CEorza5Y1Bf;r5;Bcm89^WSphJ^_Nw1zo z^_?W$I?^mW><~WeC_NB(vSpYvEnG2K`lJD3RttIY6=H;h+%JY06+nXWAR+hA^>XMw zG4vmPG#BlGb>IGM^1ikHKHcq&j?#4j-3h}Fb){;q5F`QE{k7`7QHu_d%IXlZjzC8~&>STcYtTyXlXl{(xJdol)AWGI%H3Z@vEYKKHH$7`GZA~{#SAgjd#t7(!8-BQ4 zss=JMx?*NRYUV*`=5u1Ibub+EMk*C}bT4N zLezc^3ZI5*4_nK3t$7oLYA0B`J~!l>iz0Ixz%OZ)n!2Ih`xFZ=Xo3?d5~UPyxP}IcRN==S3N~f zEd`>Q0=Z}SA+$t|(DcQL$;YkQ`lZ^R)3uc&wJ|-lmF>01&9y|H?!MOU1^Vu7>h2G| zyzjX6zN0v`gFjW^sc;HJ$S}vKLvDz~deYfZ8Znq`f zWO>{Q=-hls-7N90-sCUeU{>$MUE>f@Ou*n@QVb5tj+v;QYtz^?W%TXDoMO;Y)iYZUb(CVM2<*XMkIPO*#f2Krwg?ux_35W1<+W0ETB=6eBTs4~oIJe{Hfc zctMIm8i4++OM}9NV;jIR;Oh!70YLmV<);*}Oe$ldGwq@Xa0F8fhRK=)hs$mTAewwki?4Jj7XP~F}YL&e3SGum_nMz zz2Rzg-Q9yja0@QM-QC^Y-Q6v?1qmA5J-7zFlk9!Jdry1k zwf6qLAFyf->7$RTF{_HT<}CVU6>+V{c5RiV5IKq*TFY=lGVD2569EeuM2;6}pxM7w zOp3@MX5UH`*Du?KYfas^RSb%2Ee-fkA#$K?{(VW>wpHd1$uQtrO`O^)=0)T%2EbGL zWx0U(d$!8xAsI>lRPRj5 z@F5xXfG1!GkN^l$h5e5O0Ez;{8`~iAZ% zJ|ag~U*8i?JUr9u3eL%xYoh@whi1NaO!C+)FByc zoZ6wpj3F6cxmF{8Y?-(sa){WqVkPuXkOG3qEfafO>!B@M7Uqx)Ca%@rG(ZzEBqNS< z6@GNfL==%@$G$ZUIwS+ewiPR`f1;3cRmS(9m+=0HLm=vBpql!tGIVH71BU66YgNX-e*zqk_idScL*xMDSULX)hPOpD+eI2efwchHfoeHBEG zlmlA_AVn6gzuZy&u>flL9UI`6z`2SGU;q~xFl+xo+aGwFcWC7b>7RfBEdR*_z~UX+ zGMU(59rxf`)dh9X`BoznPtH;hu8}=)Af|-3%)!x6QgYSm(D1UuBu97)x_Zhb8iPig z?Y!0Z=S0iL>R>TqSF_Brf{C*wBV@^de`QvEe9!3`39@We)LQk4Od5IDdY_=;`&HoG zfTH*1#nYpSF8ycV@j-~~KXK3s>kbVBC!;CgH0p7(Ke?8u)a#lp{a3Gdbn154lBZl; zFh}>+;&`9)Rs8LzLMHkruFn1(8NS?{7=(1uC(heDx1@`8a~z^;R$)9r-ls?%q0`>0 z=H!>O($`5J?!SIj%=Os!e(KzS-|4Kcc4iWSX6Ez-zHP}@Pu_VmbVAg&dc6!}?|!bt z9y68@5XcWAz>kn^5>EBqYm{g@{cvXD1U_3@674V(QyN@qwVe{m|N0#co#T~D%OS}q%NsfNZ>4e=RT z5|iX4*75MPJ<aX_HITc5nPOQ$Uu2PuF~7_8|DuY&6r$zD94#6SrQ;K~KVkx4jA+r| zisIu`Xk(>^6zdqNvBU=bOO4>s;PT>;Rmgpw+A0KR8iWgBQ zhVT#Fk6Q?TpXD$yR-6~O=#I?r)@kgvV2;A7S|)F;$RXk3x&-QRJ zHZgFFFzOlYkHJ`8?-3YPmJA!@VC7E7fARk@>&)iUo-VMuR+X)aWa-H;uW*)gKIVg@ zr>rEz-h>Nh)!D#ZL^45eMo?j+e~)jF%DVBxCVaw#lFtp(tj^8V(f$ZvYc$^HWa62qT6-(JgpQGoB$JBa2o6Hi*Z zYl@kFjp3A z!6WinQj7YV$li|lH55PYB5NpnecY@p`v{WBHctjUmTmz9+E5Ifv|d@JfD*@s7Cuv7 zER3qKprUn{$+l;^EU2NJhskS6od|9s8)R@;S;h`RWLXa3Kbj631YqS+a$A<;_{OuL zNeG+DB7EPfFV@t*0I(7h*w_LgihCaY5_1)$e^5%PQE};oaetoF(ww2CGf9qP9S==2 zCYz`?FyWGF$V}DHouOqgQKdFlr8ZF|H&-P$QKdImr8iDC`h^mq`ES|&Tb5!^3e}(f zzn1O4u1208;gJ3<+c<6a$X_!gSf2-l%Zy0T;0ojIRA?im#TDxasj)-`g-eaVlVEaD zG?Y2bY!^@w(UaN4z7H4!&umsM(DvJDC@Ye>F3f?zvszL^UW=Jj#{xY@>$;%wIX;mM zO~}twR!s0%Ls=E;6af!9j)n;YTs}PZp` zLUz#ZT~w%P5iYAN2Hpn6ua+!=Ti2%3e7Czhqj$mfd-jOJ4OND zlW79?<5H&M0z{NyBWbW==3L4|u#~3dbdHm0SDaS4nPm#_0E9HCL$NTEdvywY|V8Q-Zq2l_J^4?I%8jhT->tR(~uvfo@%1hQq^8|Sr3*@JhT%d zA@F-47#`(|8@ia;?*k4@{MCQ&wiGB1Xz5~v#gDnJ`F1z#Q( zcjd>ud#NF)LpZXU!v{{hV@mOT?D|O5cb3)34%r?O1iHisL|>2?GCZ(QZ?m?7tydCW zXof5@qN59ZsFpI$3ECK0TZD4C)YxkF>UUx}>9DXqBC5_dv~ZpXt|#nKtZ}JW7H9je zqNorQNId?bpkO>G9yBy{t_Zdmh~voOq$;<(PCOPW=vM2-2fK6uYP+Gm1@oYw@7ZS^ zgqX^<h=)s6B7CSN#F4Ua8(2{L-kUjNPKYxCVb4CTI9ch0z<}?5AxKWl@ z!lpZ4MSpC$Vj_5QLkj1`4C}AXryG!;`K}U0MGnl7*pq$cCxPlK6|pDO5+K_Vv(vgp zbep?urHlEs(h|fK9@p`N2(pXZ7Uj!N(n2&Hvda#BoVbcEfuFDf()0;YDb-Qb+EBq1 z>FC?4UvHY9)+GG0_gAdR450xqdz>IP@9B^oSWt)cwY3<5i?ab=iuL35Y!Vx{BhUS) zY{+n7I`nqN@w7Ydwt>a%LvUYQ{|<1ZlC8FGxhu5l{|U@AtCpo%5Ht%0B>6GaCW;{cOZrwAOpoM2!+ zF1Z0ToQ%{B7?n^#Vya?%wUT`|ic@lCOJkvi<_}sD^JHX}ap`p9!5{YVCR|JnorM~@ zKPpvbDph_|%FI;C{HWBPsni~$9jzrpZ~D)vDdH|qn?m}Z`Rl*UPybGB|DhUvAVkNH z+?_ETNXGwX>XKki4i8cqm83?OB-<%f$4rXMHxN=`4h|1e7=@<(BuJ*DX*<2SijGT5 zI~wF0HttMBrCy~GyklV}8n>Co3dT5{n2z}aBdtxid=S~7*V3+$R0>M~8P)Yz9=!1K zNlYpq6%Q&dnh^wi3LaS2KH>gd)|BvpWo=0WSk|x=fn`k|6mv1p?GAq4QbEDO;zb0IQm*x2FB{ zx6)+7Vn~V4sqqPONn~iN${wPW{1}TdOUhTp+TRi;wW*>|AorvogcHkgo#4g2RS+OD zapsfD4#e_|k&*D6!WOUdzPE3+DfoS)MJlHUpYSUm|9+igp_IB>l8*V>R?^N&yHQ4D zJLWXDC9Q8R$A&B#t}F|yPEttkFK=Z5tvr){3{sj{oK8rdRYo2xZ)^dRFq7)8ctIue zqA+q(#Ql-jsf>I`S}T0>ds_*6DD*jaaFBzXz9ivBV84jHtop@#`Zva4tmBZ}K<}AZ zRVQCaCouHNggSA9;jjV$^_T_iW#r6b^dE!~ z`za&G0TMJ3Mv{k>^??(&Viq)jg(r_PWDTnil!f>I$J>WpLSXwK1PulE83ourpalK9 zeF&M!m%NxO_<(8ls0hCw@k5)4FA?3Kmq6N5Dt;z|nDQcI2G~E)QsiU|6gv`%au0xf z@(sY%(GpRx1-gD1ooNP5NN}_M#6(}7Fbm4ma^b|B_i)z zOpfaz8WAYrVo?O&umQiSk()K|IP-+*o1#Lg!L4I378fc!=oA_>1EK38;A#XmnuPnk zjmeB;3}q+1jhora0p>rm9~3Ghra0t%c-J0Cl^+tz=-niTR6LYIW3VzVLdqtpLGN`2sFa-hEC_09@$pccewo6Eemrat>`v)LEN6JQe~+fa z5AkVfJJjv>dnD+-=l8sE|e@JC)#BxSLqT9rku;-{?|)ya;}v)z>if!MPaW}M@+wGJkgK)N^~0G5wn07(N+W;og650jb>?eAD7MmWQqxFp_8S<3T$Mrgck{e?|hnqE}+e7_|> zAyP~dkbvGg5QQJ}*p~R)jT_0Lq5?Bz-H8QRyBYo$vmVOTan5<)(#;AKW*lyRp)F z4{nujP6^!L3Mpr?M!;CU=A5A8^CEU4<8CPyEZ$URgJp_NijV0dciF2T)1k%(C)7i# zMiEF)+|zim%+DPIy-;)Qj1Ps_uzt~%2t=DpP4oc|cRlpM0@DjcLKG{--;Y91wxpv? zb!wrXr?pZ*lUd3la(fPJU_V7qBY^zSXr~Y5Q&}|rS#r|MxR<&XpchHb?lf?}xbYKe?5sts8pXlqb6Dj`LYBZ!?)hHy1+!Zpk^fN`)=?O@qh&K}G zbK?e>g1{Y3Nt3+&2`d7$oqCw>_cCAakB2`?2>h%-_Zd-zdZ!^?7&+H-37{khZ+5J+I#r-$c9blBc4E;&Niv3#;{16s`)=)P^zh)o)7X2JGt(pe*X zR)uh3-poM1PLb2sPp6yTsiP~wR6s?zGXtYD13u@XfRE*oe>T9KyyEQJB3<0XK@TwM zmx}(*242rvo(C29o%M$3YgjM~Vz4Qv)TLRlE++W+M}v`ActI8LVGz+kC3s9u+J4`D z-Kw(v{D&ZZaN$IXdN>L<>b7YJ-!1g^-IF;`SzM=w!~X9VFnpL7*PoF<7C}~zijcsE z$!>D>%L4H+X%nWw0*FxD32k|Wkt9!9B|vgNLi2HEVo7WvJleLhxQu5xtuXu+zc9s% zVm=fyh+(2tv-KCkeVSrERDRPyb^CF~e@(!U6Z70vl72dD;Gp-)fpqYaQ_E|&|C43oBQ3ZKC~|^Zvn7hs;c)46s(5bOp+=_P00slulFy{{)7eApHD?2 zGSfFuhSmnCgG?ai3o`K>f{31F1$cH^j)^A^FNGX8sY$N|Zy%N;aKbo43hH88^6gjIVlSnLP9B z$(Rs8k630aK{5SS-|_Hzzhrn4tX6Y}(X#}xwj88H%>k7qP;ELfzRaD)t&{M)xb*n- zW#gX1$lM^IkVvBk_wadf@3{Jx-pr(tTkf@ETlMv#6B=ZT1k$PI@%8V`o{kWEs2Q`} zKL&e&;!|@7<@D^_#om{1-qbaaYUvmRU{m-E-E@Y32=WTY?=apB8q_Oi>yb8Vz+ zx+;d2jtr7_-Rv>N=%Uukwf}f2^CuV7eJDf8DS3)`yK6D<=FnQpc~L`@&mtfl_gicn zOu~NY6;T}3Tg-UbE9ji;yz;Pov*PGkx3Q%A!q3B*XR*V8w%?XF_A8+2Ji_gyy~V`m z@<5=c=a19N940Kqz1P!4XEw&-nvS#mGsArh-RQY<-u?!Sw~5O{h5en$utsx-I$l7x zTTf5A(E@JeOOP!y)XH5^@&)D!sScqg5y49!i~_hqd8@+mdA2^D*d$hgGwv zFOSn30zcRzHcT_NJ3=P*|JZqe3m_E`jK$95N)HR^MJ3+NFAuuhN9b3cGQ75BrghcM z4>oW9?7DA=mU8X>#RP+d;X~Ko*U8CIYjg~6km*8 z6wMQ`lL)HMFmg+_t=*i~FJ5^&7^>g8y`Bx9F2CxW{r2g~;q#&vpiKFs<80M*z_9Uh zI^UCwv*Og&_IwW9gM9V1iBe#^OK#TQ{&XXW&Rgy6VeOBOo9Z89mJEyxmth_szD~{; zjkluP(VUI$8s7W+Dh8fShIut~`FXq=>8SG-S-5opNSlCh`P-JBwH_LiFK{lFz_T$a z8)rz%h%~2{vp&EvYkWw#DEdv-J`epq&yEj_0+THpXK#0}UJqLv+1ce~&JV4b504La zY&g*brXiu~-kmWsbY3>5D@!n>0>Q!eD+yg=xB}0wpY4)chLfwTW44!ZPrdHP?@M@I zMp@?5(GLaSH(x3dtCO~hT%SDO)T5KVT;8twJ2DJ(#5T$a#tLlR#M8Z6xuP2x#y5ME z3|yprmN#C!Z?}9@YpS*5X-r^@YQxcR+iBtER;*W9PSen#7J^T21#Nwfk7@?cNE?@v%m2HK55c<#{$)}UHhgyrIOO7L`S(OKxmtz>NQbS%*ZXjUHI zV&tw|C=VeLcOj%Zf1zI{kXG;F&}5=%W{`B@+%h5706`o~8`v!lw!@YyMBq(5!cRDh zS`5^Dj5&RVL>DvgbbL`3=*F)-8sO=8q7Klk9HINjX5Js^#Bz(6ScA0W$G;R)Vfrno zo2{YZ`m3rO9ujh>SI8U!8d6+O$WVA(OLJ$CxT&E8^l`c*;NYq#r5SrRO)E7bzB20 z-E#uDaKhi#=Hpd&M;eN~$Jv7Qlye4oa1EK}=PBcyqOK01P2O|zqxA)rgin*?q8yFH zLKWnqINTnf%kX-n$J1G`8^5m3KDaDe!!!DCG=%R2azLq1_>b9@ZmxqIoe|ON=pvC? zAmt5-op3ZWMoU$Rd@Ol7hNi$qv?m^X;f?#17fzb$Dsa z5~LNtvlR)+?|>$}67`xX1qQ%dl4xWR&6;d2wmd z-tna9X+|yZ!;@5kfjs4z&82j&Z99!7-x1n6f&QS>Rnh=5;Tt6-cbgDlN+93*i(mpE zpt6+V{w1N50FY>vDscTvA|sIfd~NDH-o_cLX@2Z9ZCPj7?n42ljnmwHk z1K&P2y;2#)#i@2_c&)#Q%lDM`%g!f6Mr;d7URu07bgsRvwIS84@dOX8*+?~_o!3N{ z+O`Iqu$WQs1hWyc>{?+KGJpSq^Q%~fryqxCZrL2tvlR8`(eG)m?Sw|r5=*SlL+97& z^~Lu*ElLL+!-`zSAq0A|TYEb+Buvejp+w5;?Mo9HMk`PgJz~_I1FiKaUZ-O4Y1K7W zn~#_*MW>Xo3_lmoZpLzSivnJd@FDh({~Rr7r!9^;ePJ$gzixkwV`WP}BBX$g&Y-7Z z{tC5BJ&N}J2>#{Ujj+AR6bW~CM5O368e+vqf zgGFF$ecyW^QHU{lgtN_+8Y1v>XPsp@5=^T)^`4G0Pd8Y0&irV&T`G2?DErk<~1uZZ{ zg2JQ&>I|ZH(Rr|dVmWnoQIDLjYfP@jRWF3y$FzidyGZRQ?Ml+Ea3%LBK>?pnHI^y{ zG3TF5e@`)4&GF(i{LQqwUh$cOfA6P)OaX7ZcW{KIWaQPfPGiBlqpdZw7ABX|4TzWg9XSw~x@Nt(VyQ%|m` zEAx1K!qTbDt%{zbW`1wKseQF`Az|udhej>J@)KYAlHu@WQcRBiZ1s*dcNSgmk-dQf z{gTi?&DjlA(Clf?`cls&?NGtgz^{=@9i6@5W|k-C(}CauF>5R$Nq<3yOSOJU-$uV)|_9NDNOC@ZoleoNHSzi zz*!L^abO1)<;eJd=FifcU{5XmLI+Q%3tiZ|cVjJ~*mES2h0hQlZ#N=&6VXFVeCD9`zE`5%D6jxoRh zRU?4HM)qdFg{S^FLSn%^edP)ONbEO^cz9NKMwr-hYzAjB{@YvIar3Ye3u2zTs5&}d zBL5GAk$vX6Nc|2>joU$zllnWns`kwR9elkuQf>hiH6o;<;)WtFK8y>k$MpZt}6#7>b z9hAWpAG9f-#YdR7mQ^P*C&`hBs--#5GCIQ1RjRn+6UhS802qlak5cVa37#LF! zN;;z)$M0iT$MimSxn?bnd#$DKqn^bShK!M3+BApa)H{dbGL~a|IfpWzG@ge(yMz)M z6*kdk1SQr{L}M&G0Tkh3B*1H0w91&e6wZXYb$0QxnzM`BD{eQhhgz6xc~ag6y(B*1 zQ99?dt%WX(oRR~wJhgT%&)i-|zx?uQ${OC|)mk2f)!FzNB8#0%DUH(^BAZ(+GealO z?>*Cyw8l1u4xUw@mHoH%479TTwr0ks5LwI}^qgIcQgZODXZE{)pN%jG@cb^E@LF19 z>hN@FmCAUk@7TG!THE@IAN>c86Oufk`QG?`c>C%lPVR5`!MzTWjK7GIh8H7YO^AnchhR-&fX2FKS9(OZ~45ISC& zRkXbiicVvbRv!07@7rnN^P>r;gb{q${KPRePHd~)4di^ixC-AMQZ?I2nfSQDk|L_y z9Zjn0j?bw=&z8)^;q)nWeXP2aRj)6a8s~8g7vu^Q8qb>+U1{*OL$r=a=`sziTE-Lb z2?a5xi!DzAcyZ~-+w3M%juG;EDc<`Gc5Wi@& zAtR*4LZbc56=h8gh4lNAOZ<{6`j{umgp7b<6?)T%a%7HnWKK3SOBNn&HcgBzBN`wh z@GBIOCcjsVl@l3(KfgD^Op*j!sT#6L;7QyWbOI;s_$HQnty57r%fH-EUWR3kd$u#K zY!b1o;a#V#t&=I=TXD_l6t{s!cBD$WCN0`RezZW`jv&>7{L@ORy(;u9vXJM~Pd$1{PNMNvqNUN;qD*dtFC}7gSf$(~ z`IEx)8puwF<&$Bh*9uOn2zAoI;>G6cp{f`pD)%bKXV#dE37aa!MwR|l%N&+9xl=sQ zg&9a0AqT_a9Rg?>PIhX*M*-CE@u!mWt`0iuDx_iUn%`7&B4mHfL#<|PkM9&ydU*OQ zo;7UmUv}$L{>IbwldFe~`ZinHnN)50d%x5O7450v+p01Jl9z5ZzAj0~l8}&>98<{` z4Etf+lQfet6i0ZsY>>8VW2!n2pQw{65r^ zbf)#9N{P7!gY+?$ zi$gYH)}^?SN}sp6MyEa$wn`Ux|% z4EbuWN&UmmS|AP^34Lb98@azHuG(Y_Wx(960GwT7vZ%$EZSD`TJGlCuPE8iRLpOfO z6|}uP7Z#^|r4kAmsQ)C$u3j_}-RIkqYX`C2m=vcgj^KmUBq$+8R{+@*iH&LGFvtQ1 z2aBxYh{g77Jja^SDN{WEp?H_s0>_42UEBjznjRL2qABj-7`2AGW!#vC0HrOym}AFh zJ(x00Z#00Cj_jxswf=&}g`s|-xro&Ln!9+q=jb8}iuq`(LU#QqQb*5T!KG|->r(c| zx$b!3k6Y6c|I}GIPIcpWO;cnyleVib|MOER;}rk;I^Pe0(c?vo1snRx4P8kAM--US zKt=OLUs^lr+L;bBn%H%)%zdQ}vxsFeEr5GJ&kY+Tz#YTtPsicrmVMF)D-U?miiBmH zVI3d0n3&4DfaL>b|R?LU`wzgB;5bk^<3 zb}itBRW@9_dQ0^qN>R_=jY?7FO9=j9g2vI$Q&ur%2@{e(zEa?57{OjomvU*h*v+yY zW3x`rs?I-kT(awFyZw=m!!k&4nAZ*iXrHsHEv)z?T+B+X+q0?>*R}OaDoPMl-FNkq z9HKJF)y*|)XSA~>4&x*alD!ioD>4|>k1hj}5DP-2hMeD>^o{-BlQTKiHx@@!ZPDH- zhEWYhuY^}KWpV$J?=apeRX8t;S^49kcdRi#1b-zjul>kB31MBEP<^Q^APJ94ovw9A z_ERM-ow`Q#q{&Td)Fj;#d7LW8*V_6WD4LaUeCKZ=!N zotQOi4O*(P`1XIcC3XE1sQxnG%R)H5_rFe6dOm}mY>#_}OL+tbaEf34ypMj z@^>$0Ehabj1-f&=mp!rJZ^CZh=CWle#pDd{cdk!e_u-?y9>saSa_wC%gTGRJSF+%x znG_I1a=eoC=I_2TiErjd?>(%3le|BhI1}d!Vj6|JeD(|<>W=-IO#X>gwI?>(DWMUK zv#2?N18hM6_E=%x&M5v@Zo@6Cm#0j zS8XRMHDOwA6~kU+F@?G1Vsz6#gBi^X(q{%01Fazd)R33f!H8zDPN8jNhr)5R16pB+_G_x7jW| zCr25H@LB>tCT9-!s0F`$JX!{82KlIk+~<0O^Nn;3`3N()=}duX5vNb&8WY3v_3173 z+4XgJ*yS1YqtwrT<|~$*i_3eg%VTSO7Joh*U)Rn16Yr0rNocrL@bmLpjSF%3Xs)R+ zAf5AM)^{)_d6&A#A##6+G$BfKub_?N*tTwZZ6pw$ujpTi)?L;~mE>zrh`yS+I~rP@ z;ZxR(sk$?kG4H~urA}2Uw~?+WyXncKs$xQqVk^3}MmgHgobgiDlrMZ8lrf)8)hafX z&ZwL3j-;}322+JnmcExY=SzA`VJSiw%C9SKT#_Z8%~US7kq(`u!f_bClQHLM zUQS^uI<@|ut}mU|O=VR)^JZV%IH=CLkl5TIWt8DCb|-7j6ZeS0Qlw*O-Ih&tI;8r< zT->;VppGL$wZc(!s`98endpoeF_b=I%@L;UAlKGSbxN-EW>DN{^Mh`KrRe3Ch2Bu& zUB`zLf2zI5;>HZ=xAC+Ye8ZpW%FE~EaW<1w8yrP1qQ`*I5U(ol6KBe4h8+Qw-273} z?>LuwiSQ1#0^qHf73N#q_TZe}CjHP0D2H)E>8z(;N3NIOIKJ+LjY`DfeN+TTSYta( z&v%{3H;6OhXYqM*LmpNYjz^5vT0LJQ(co8o!3c2hPXpee^*C6UB8iD1gc%OBN+pAe z2{FP5xf{6!@?T%4rke;eAmUD;`C-|7JpzSH7f?evbr2@-EjW$43HocZ7VXvw-EMK_A=SUHJCG$rx(8WY>Rkz_l zaEy5P#HrCNwExK_UPI^sP9J>g(WOY-_kQkj$9BAL!S2mK zaE)bfxkq6=B{rylkh#9k71;=$PEJ|N_JI#L4TbDrf_Emt{yoh2E5^8X?w)FG8yHT+n(o(lNPjGJ7(XFm!}>tpeR^O|2T zoH>%KKCZ_2ynHpyMo-5PcyY|cQSv=^7MQNeO6XJWIExjY9i zXbk+RB!P+VDSW{-AU94(jiqMVX?PuM&bte&5pj>m>7Oh+5<{ApAmPj$Ci_+O9y^>wyg3% zGuI$^8+t zhAm)knI1Ci7`PvZJ;a;My?J?K-qpIqyS}C}l`^p(^9m=FaKSbIQ3mrQ^LY9|MRuCV zTkdwo%K(Nva-ITzjUqT#q*z`Qj>QNi2#a-kODP6z2bP7B(4-->k6AOH(4)XC*FOo1 z+1og{@{1@@g1Iv^wOSgHhat8NxvOV~MBZ-0ybg6B!3KY+0?huU75pqNWvn{gk0UtF z1QeN%+H~!)$Q40khky)p#GWKWw`fO^ipj}dMcPXxQ`cDgX}2^nHTie~Q$@N(UarQz zd~GJ=mL5+y#A)d(k)yheh^g9AL^iVN(@-Yz5;zHsM;HkWj*U^4TaPA?de+WnCAhe%2N zn#a_Ft!7pCuZ$9pPz~wAAGm}BMkS^0iCjl+{O&$*kQD6`A-qUleM-XYR6-ezv9g+D z1Q1X$?Xl+r4IH2jd1e}S!{qH zBVnc=Zn61Thte(g*u$fYWI=1v%}r9Rcqo+ay$QJ!$YklbEITKgzEXG zHY{LE{?~@*{n;?-($D6QMXBZV-Mn;IKCXT>8>Aeu!Thhb_^QUhexo$zi?+)2u=I~% z?&{!O!@+AiPr?c#LG_sl14FJ#jbAe&4y+EWNbsh>DPl+Fb;woEzWC`SX@;7^d{t0^ zWPDYFyjci`d2X85yjI)U>|lNI+3W_{+3%v|3tp?8?B;h-Gay2G7hUn5Zf85ci<}Lz z=^<~Xc~4!l@9gBCK75su_o(B`#-0|rZai?~j`HBf` z3?y5l9R`l_b2p3L&+2!2#;uS4GfRFBIJEv7cJ|#cBCCHscK&ZzZXn3gJC(u7txJ|F^6Y-LbKL5k1KTX5_b9({ ztF7O21Hue*xRKLSADXn`WykzYt%n6jP3muI_g)VMf=ClTI~`iXa{8mNfzE4Oz26%D zH@8cYh^PqIc8%Yi@kVxRq!ib3G?)G15fQMD>c0!)e{5Sw^6zIUZx5pqV4&U9t-1tV zw_?Y~xKSiNd>{iHAJxya`iS2vWPOR$n{{+M-=9C3CxsM%BDj_#yA+Ya|LYefIy@ii zDZ?AHigAyWA%OU%?-Ju@JLmf%+s_zo7=Tjmw+BtO2a`=x=$WUXYDPo@KMqS`6lARN zHK@{-mL@Y6E!#Khi1YC|PXY1(5B)m~2?|6qT57I2eo9m9-(W!%m;#%h#!gvcV~ z+$@n69Lc0Q!p`1PF2h=ku6&ImSVtkZ_z z!gs$)4=oTY;eVP>koAy=Mr$0gY_^0a-F}sRU>f%bMhLhg?KXs#n}Q2oU>aah_IeDU zJU?#`jKFY5ifssSHU-(ez?MLu>W`d0%t8Iz#2JC)jFsIGWdH8>?Gdv61ESF^M=XLZ zVMe##H(#)}noqu7`am=YM57gsSUOw66+rF>wsePJ1cy7)ycnWD@o&?HP-|1L89)Vv zsz1dULFbIM+7N91?zeg%k_)G^A?W7(t4AE$jBNoq=u%p}Ww=dYzErJmg?nrgZ zr`)dT|01^`1NhG!mxC4Qnh z!{5)TKEubyPgs+sI0Cz)1+?zy(ISv}NtMA51qTH_&~tivM#k1jL++OIJ&P0cd;2 z6W{&cS00_CZwhjPZZ0V_Iq+jJl=o|64t$gpxxA+;)lGPbpA@p4Xk5JCE!L}e$07F) z#})_g#z@VXS;|xQmOOQ{229B2>U^jA4@wqwdf2hSFd0iw{S5g8_|Y>m($Q=i{_lsB zPhmfBp2F5)p0>+6I(k;9{_A^8-rIZcf&G7Xuhqdzfq3cStM=^w>`Vr!G2hAs|bT;Fg&t4^z3DkNE z+t*}0OHTM4`iPZps)dgW430|GPltDbAernz>XFV~lna9g8CTUfIYxmGC?gJyZFKj1noe;O?+T@MvcY*e*n{cpu%McHHYfF zPzGOjqbI1d$oX#P|Ht7wEMUkcarS}!;iJDe-nW&Q!Ts)Mm)8s17EPblB|Z6e?~UwR gPDdT#`a(~9z$1vDW43$D9fT>t<8 literal 0 HcmV?d00001 diff --git a/hack/inventory/capture-host.py b/hack/inventory/capture-host.py new file mode 100755 index 000000000..8e48e8681 --- /dev/null +++ b/hack/inventory/capture-host.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 +# Captures a machine's sysfs into the transcript atlas-lib/inventory's tests +# replay, so a reading can be checked against what a host actually exports. +# +# It is run by hand on a machine worth keeping a transcript of, not in CI: +# +# oc debug node/ --quiet -- chroot /host python3 - \ +# < hack/inventory/capture-host.py \ +# | gzip -9 > atlas-lib/inventory/testdata/hosts/.json.gz +# +# Only the trees the readers walk are taken, and a symlink into /sys/devices is +# followed a little way so the class directories have the attributes behind them. +# The output is one JSON document of three maps -- directories, file contents, +# and symlink targets -- with every path relative to /sys. + +import json +import os + +SYS = "/sys" + +# The seeds are the readers' own entry points: interfaces, NVMe controllers, +# block devices, the PCI bus, CPU topology, NUMA, and huge pages. +SEEDS = [ + "class/net", + "class/nvme", + "block", + "bus/pci/devices", + "devices/system/cpu", + "devices/system/node", + "kernel/mm/hugepages", +] + +MAX_FILE_BYTES = 64 * 1024 +MAX_DEPTH = 8 +# A device a class entry points at is walked, but only a little way: the whole +# device tree reached from one link is most of the machine. +MAX_DEVICE_DEPTH = 4 + +dirs = set() +files = {} +links = {} +seen = set() + + +def relative(path): + return os.path.relpath(path, SYS) + + +def walk(start, depth=0): + if depth > MAX_DEPTH or start in seen: + return + seen.add(start) + try: + entries = sorted(os.listdir(start)) + except OSError: + return + dirs.add(relative(start)) + for name in entries: + path = os.path.join(start, name) + if os.path.islink(path): + try: + links[relative(path)] = os.readlink(path) + except OSError: + continue + target = os.path.realpath(path) + if target.startswith(SYS + "/devices") and depth < MAX_DEVICE_DEPTH: + walk(target, depth + 1) + elif os.path.isdir(path): + walk(path, depth + 1) + elif os.path.isfile(path): + # Most of sysfs refuses a read for a reason that is not an error: an + # attribute that does not apply to this device, a link with no + # carrier, a file the kernel answers only for root. What was + # readable is what the transcript carries. + try: + if os.stat(path).st_size > MAX_FILE_BYTES: + continue + with open(path, "rb") as handle: + files[relative(path)] = handle.read(MAX_FILE_BYTES).decode("utf-8", "replace") + except OSError: + continue + + +for seed in SEEDS: + walk(os.path.join(SYS, seed)) + +# The readers take one root for both trees, so the few procfs files they read +# are carried in the same transcript: the memory reading and the affinity mask +# come from there rather than from sysfs. +for rel, path in (("meminfo", "/proc/meminfo"), ("self/status", "/proc/self/status")): + try: + with open(path) as handle: + files[rel] = handle.read() + except OSError: + continue +dirs.add("self") + +print(json.dumps( + {"dirs": sorted(dirs), "files": dict(sorted(files.items())), "links": dict(sorted(links.items()))}, + indent=1, sort_keys=True)) From 1cd0693cba9eeb0cb702077062c25010fd274d95 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Sun, 20 Sep 2026 20:03:15 +0200 Subject: [PATCH 100/206] fix(operator): the management interface is the one the node is reached on MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A discovery run named enp2s0f0 on every worker of an OKD fleet — the fastest physical NIC, on the storage plane. The node is reached on br-ex: br-ex 10.0.0.15/24 the Kubernetes InternalIP enp2s0f0 192.168.10.15/24 named instead ManagementOf returns immediately for the interface holding the node's address, and never got the chance: servesManagement refused br-ex first, because the probe reports it as virtual and LinkVirtual is not Bindable. The bridge exemption below that rung was never reached. What followed is the whole of it. The add carries the named interface, so the control plane recorded 192.168.10.15 as the node's mgmt_ip, and matchBackendNode compares mgmt_ip against the worker's Kubernetes InternalIP — in main and now. The node came up online and healthy and no StorageNode could ever be matched to it, so the add read as having produced nothing and was asked for again, forever. In main a human set MgmtIfname to the right NIC and none of this arose; the rework moved the choice here and the choice was wrong. So the rung decides on what an interface holds rather than on what it is. Two are still refused, both for what they are: loopback reaches nothing off the machine, and a device that nothing identified which also names another as its link is one end of a veth pair whose other end is in a pod. Both halves of that second test carry weight — a VLAN names its parent the same way and is declared, and refusing on the link alone refuses a tagged interface a fleet is entitled to be reached on. NET-31 is restored. "A veth holding the node's own address names nothing" was right, and was overturned earlier in the day because nothing could tell that veth from br-ex. Its report now marks the device peered, its expectation is the recorded one again, and its case.md says why both halves are needed. The new case is the machine itself: fifteen interfaces as the inventory read them off worker-5, asserting br-ex is named, that no pod link ever is, and that with no address to match the honest answer is the physical NIC. Co-Authored-By: Claude Opus 5 (1M context) --- .../case.md | 11 +- .../reports/worker-01.yaml | 1 + operator/internal/discovery/mgmtiface.go | 36 +++++- .../internal/discovery/mgmtiface_okd_test.go | 120 ++++++++++++++++++ operator/internal/discovery/mgmtiface_test.go | 39 ++++++ operator/internal/discovery/netstack_test.go | 78 ++++++++++-- operator/internal/nodeprobe/collect.go | 1 + operator/internal/nodeprobe/report.go | 8 +- 8 files changed, 277 insertions(+), 17 deletions(-) create mode 100644 operator/internal/discovery/mgmtiface_okd_test.go diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/case.md b/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/case.md index 3d49e1cd6..465b6ce36 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/case.md +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/case.md @@ -2,6 +2,15 @@ **Mutation.** A veth holding the node's own address -**Expected.** Nothing named: an unidentified virtual device is refused at the first rung, node address or not +**Expected.** Nothing named: the device names another as its link and nothing +declared what it is, which together are a pod's link into the host — and +admitting one would admit every link the cluster's CNI leaves behind. + +Both halves carry weight. The kind alone refused too much: on OpenShift the +node's address sits on `br-ex`, which is equally undeclared and equally virtual, +and refusing it named the fastest storage NIC instead and handed the control +plane an address nothing else answers to (2026-09-20). The link alone refuses too +much the other way, since a VLAN names its parent exactly as a veth names its +peer. **Harness.** `CM` diff --git a/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/reports/worker-01.yaml b/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/reports/worker-01.yaml index dea217ab1..b8c69c3ce 100644 --- a/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/reports/worker-01.yaml +++ b/operator/internal/controllers/deployment/testdata/discovery/net/net-31-a-veth-holding-the-node-address/reports/worker-01.yaml @@ -95,6 +95,7 @@ data: "state": "up", "numaNode": -1, "virtual": true, + "peered": true, "kind": "virtual", "addresses": [ "10.10.10.1" diff --git a/operator/internal/discovery/mgmtiface.go b/operator/internal/discovery/mgmtiface.go index f26b555ce..4af764453 100644 --- a/operator/internal/discovery/mgmtiface.go +++ b/operator/internal/discovery/mgmtiface.go @@ -173,13 +173,43 @@ func describeKind(kind string, members []string) string { // except which address is on them, so that is what separates them here. func servesManagement(iface nodeprobe.Interface, holdsNodeAddress bool) bool { kind := interfaceKind(iface) - if !kind.Bindable() { + if iface.State != "" && iface.State != "up" && iface.State != "unknown" { return false } - if iface.State != "" && iface.State != "up" && iface.State != "unknown" { + + // The interface holding the address the cluster reaches this machine on is + // the management interface, whatever kind the probe called it. It is not a + // preference among candidates: the operator addresses the worker by that + // address everywhere else, and it matches the backend node the control plane + // reports against it, so naming any other interface hands the control plane + // an address the operator cannot recognize the node by. + // + // The kind cannot be trusted to decide this. On OpenShift with + // OVN-Kubernetes the node's own address lives on br-ex, which the probe + // reports as virtual rather than as a bridge, so the bridge exemption below + // never reached it and Bindable discarded it first. + // + // Two are still refused. Loopback reaches nothing off the machine. And an + // interface that nothing identified, which also names another device as its + // link, is one end of a veth pair whose other end is in a pod: admitting it + // would admit every link the cluster's CNI leaves on the host, which is what + // the kind test was reaching for and missing. + // + // Both halves of that are needed. A VLAN names its parent the same way, and + // refusing on the link alone would refuse a tagged interface a fleet is + // perfectly entitled to be reached on — the kernel declares that one, so it + // is not unidentified. + if holdsNodeAddress { + unidentifiedPeer := kind == inventory.LinkVirtual && iface.Peered + return kind != inventory.LinkLoopback && !unidentifiedPeer + } + + if !kind.Bindable() { return false } - if (kind == inventory.LinkBridge || kind == inventory.LinkVXLAN) && !holdsNodeAddress { + // A bridge or an overlay that does not hold that address is somebody else's + // network — the cluster's own fabric, most often — and is never named. + if kind == inventory.LinkBridge || kind == inventory.LinkVXLAN { return false } return slices.ContainsFunc(iface.Addresses, reachable) diff --git a/operator/internal/discovery/mgmtiface_okd_test.go b/operator/internal/discovery/mgmtiface_okd_test.go new file mode 100644 index 000000000..866018e7e --- /dev/null +++ b/operator/internal/discovery/mgmtiface_okd_test.go @@ -0,0 +1,120 @@ +// The management interface of a real OVN-Kubernetes worker. +// +// Every other case in this package is a shape reduced to the two or three +// interfaces it turns on. This one is the whole machine: fifteen interfaces as +// atlas-lib/inventory read them off worker-5 of the lab fleet on 2026-09-20, +// which is the fleet the choice was wrong on. + +package discovery + +import ( + "testing" + + "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" +) + +// okdWorker is that machine's interface list. The names are the machine's own. +func okdWorker() nodeprobe.Report { + iface := func(name string, edit func(*nodeprobe.Interface)) nodeprobe.Interface { + out := nodeprobe.Interface{Name: name, State: "up", NUMANode: -1, Virtual: true, Kind: "virtual"} + edit(&out) + return out + } + pod := func(name string) nodeprobe.Interface { + return iface(name, func(i *nodeprobe.Interface) { + i.Peered = true + i.SpeedMbps = 10000 + }) + } + + report := report("worker-5.ocp.simplyblock.ai") + report.Interfaces = []nodeprobe.Interface{ + pod("3070a31ea98cce0"), + pod("74a4fe9314cec06"), + pod("b7960592644e3e1"), + pod("ee1dd2e3fa266ff"), + pod("feb8ae72b6d20d4"), + // The node's own bridge: undeclared, standing alone, and holding the + // address the cluster reaches the machine on. + iface("br-ex", func(i *nodeprobe.Interface) { + i.State = "unknown" + i.Addresses = []string{"10.0.0.15", "169.254.0.2"} + }), + iface("br-int", func(i *nodeprobe.Interface) { i.State = "down" }), + iface("ovs-system", func(i *nodeprobe.Interface) { i.State = "down" }), + iface("genev_sys_6081", func(i *nodeprobe.Interface) { + i.State = "unknown" + i.Kind = "geneve" + }), + iface("ovn-k8s-mp0", func(i *nodeprobe.Interface) { + i.State = "unknown" + i.Addresses = []string{"10.130.2.2"} + }), + iface("lo", func(i *nodeprobe.Interface) { + i.State = "unknown" + i.Kind = "loopback" + i.Loopback = true + i.Addresses = []string{"127.0.0.1"} + }), + // The storage plane: the fastest physical NICs, and the ones the old + // rule named. + iface("enp2s0f0", func(i *nodeprobe.Interface) { + i.Virtual, i.Kind, i.SpeedMbps = false, "physical", 10000 + i.Addresses = []string{"192.168.10.15"} + }), + iface("enp2s0f1", func(i *nodeprobe.Interface) { + i.Virtual, i.Kind, i.SpeedMbps = false, "physical", 10000 + i.Addresses = []string{"192.168.20.15"} + }), + iface("enp8s0", func(i *nodeprobe.Interface) { + i.Virtual, i.Kind, i.SpeedMbps = false, "physical", 1000 + }), + iface("enp8s0.4000", func(i *nodeprobe.Interface) { + i.Kind, i.SpeedMbps, i.Peered = "vlan", 1000, true + i.Lower = []string{"enp8s0"} + }), + } + return report +} + +// TestTheOKDWorkerNamesTheBridgeItsNodeAddressIsOn is the whole thread in one +// assertion. +// +// Regression: 2026-09-20-management-interface-chosen-off-the-cluster-network — +// the run named enp2s0f0, the fastest physical NIC, because br-ex reads as +// virtual and undeclared and was refused before its address was ever consulted. +// The control plane then recorded 192.168.10.15 as the node's mgmt_ip, and the +// operator matches a backend node by the worker's InternalIP: 10.0.0.15. The +// node came up online and healthy and no StorageNode could ever be matched to +// it. +func TestTheOKDWorkerNamesTheBridgeItsNodeAddressIsOn(t *testing.T) { + if got := ManagementInterface(okdWorker(), "10.0.0.15"); got != "br-ex" { + t.Errorf("named %q, want br-ex: it holds the address the operator matches by", got) + } +} + +// Without an address to match, the machine's own bridge is not a candidate and +// the storage NIC is the honest answer: nothing then says which network the +// cluster is on, which is why the node address is passed at all. +func TestTheOKDWorkerFallsBackToAPhysicalNIC(t *testing.T) { + got := ManagementInterface(okdWorker(), "") + if got == "br-ex" { + t.Error("named br-ex with no address to match, which nothing in the reading justifies") + } + if got != "enp2s0f0" { + t.Errorf("named %q, want the fastest physical NIC holding a reachable address", got) + } +} + +// The pod links the CNI leaves on the host are never named, whatever address +// they are asked about. +func TestTheOKDWorkerNeverNamesAPodLink(t *testing.T) { + for _, address := range []string{"10.0.0.15", "10.130.2.2", ""} { + got := ManagementInterface(okdWorker(), address) + for _, link := range []string{"3070a31ea98cce0", "74a4fe9314cec06", "b7960592644e3e1"} { + if got == link { + t.Errorf("named the pod link %q for address %q", got, address) + } + } + } +} diff --git a/operator/internal/discovery/mgmtiface_test.go b/operator/internal/discovery/mgmtiface_test.go index 1c5552db4..46b02e3f3 100644 --- a/operator/internal/discovery/mgmtiface_test.go +++ b/operator/internal/discovery/mgmtiface_test.go @@ -21,6 +21,9 @@ const ( eth0 = "eth0" eth1 = "eth1" bond0 = "bond0" + // vlan100 is the tagged interface on bond0, which names its parent as its + // link exactly as a veth names its peer. + vlan100 = "bond0.100" ) func iface(name string, edit func(*nodeprobe.Interface)) nodeprobe.Interface { @@ -54,6 +57,42 @@ func TestTheInterfaceHoldingTheNodeAddressWins(t *testing.T) { } } +// TestTheInterfaceHoldingTheNodeAddressWinsWhateverItsKind is the same rule on +// the machines it was failing on. +// +// Regression: 2026-09-20-management-interface-chosen-off-the-cluster-network — +// on OpenShift with OVN-Kubernetes a node's InternalIP lives on br-ex, which the +// probe reports as kind `virtual`. LinkVirtual is not Bindable, so +// servesManagement discarded it before the bridge exemption could apply, and the +// fastest physical NIC was named instead. The control plane then recorded that +// NIC's address as the node's mgmt_ip, and the operator — which matches a +// backend node by the worker's InternalIP, in main and now — could never match +// it: 192.168.10.15 against 10.0.0.15. The node came up online and healthy and +// was invisible to the operator that asked for it. +func TestTheInterfaceHoldingTheNodeAddressWinsWhateverItsKind(t *testing.T) { + for _, kind := range []string{"virtual", "bridge", "", "physical"} { + t.Run("kind="+kind, func(t *testing.T) { + report := report("worker-1") + report.Interfaces = []nodeprobe.Interface{ + iface("enp2s0f0", func(i *nodeprobe.Interface) { + i.Addresses = []string{"192.168.10.15"} + i.SpeedMbps = 10000 + i.Kind = "physical" + }), + iface("br-ex", func(i *nodeprobe.Interface) { + i.Addresses = []string{"10.0.0.15", "169.254.0.2"} + i.Kind = kind + }), + } + + if got := ManagementInterface(report, "10.0.0.15"); got != "br-ex" { + t.Errorf("named %q, want br-ex: it holds the address the operator "+ + "matches the backend node by", got) + } + }) + } +} + // With no address to match, a physical interface holding an address is named, // and the fastest such one wins. func TestTheFastestAddressedPhysicalInterfaceIsNamed(t *testing.T) { diff --git a/operator/internal/discovery/netstack_test.go b/operator/internal/discovery/netstack_test.go index bc73cc1a7..701f0b6f4 100644 --- a/operator/internal/discovery/netstack_test.go +++ b/operator/internal/discovery/netstack_test.go @@ -31,8 +31,8 @@ func stacked(addressed map[string][]string) []nodeprobe.Interface { lower []string upper []string }{ - {name: bond0, kind: inventory.LinkBond, lower: []string{eth0, eth1}, upper: []string{"bond0.100"}}, - {name: "bond0.100", kind: inventory.LinkVLAN, lower: []string{bond0}}, + {name: bond0, kind: inventory.LinkBond, lower: []string{eth0, eth1}, upper: []string{vlan100}}, + {name: vlan100, kind: inventory.LinkVLAN, lower: []string{bond0}}, {name: "br0", kind: inventory.LinkBridge, lower: []string{"eth2"}}, {name: eth0, kind: inventory.LinkPhysical, speed: 25000, upper: []string{bond0}}, {name: eth1, kind: inventory.LinkPhysical, speed: 25000, upper: []string{bond0}}, @@ -63,6 +63,19 @@ func stacked(addressed map[string][]string) []nodeprobe.Interface { // stackedReport is that host as a report, with the management address where the // case wants it. +// peeredReport is stackedReport with the named interfaces marked as naming +// another device as their link, which is what iflink reports for a veth's peer +// and for a VLAN's parent alike. +func peeredReport(addressed map[string][]string, peered ...string) nodeprobe.Report { + out := stackedReport(addressed) + for i := range out.Interfaces { + if slices.Contains(peered, out.Interfaces[i].Name) { + out.Interfaces[i].Peered = true + } + } + return out +} + func stackedReport(addressed map[string][]string) nodeprobe.Report { out := report("worker-1") out.Interfaces = stacked(addressed) @@ -80,9 +93,9 @@ func TestABondHoldingTheNodeAddressIsNamed(t *testing.T) { } func TestATaggedVLANHoldingTheNodeAddressIsNamed(t *testing.T) { - r := stackedReport(map[string][]string{"bond0.100": {"10.10.10.113"}}) + r := stackedReport(map[string][]string{vlan100: {"10.10.10.113"}}) - if got := ManagementInterface(r, "10.10.10.113"); got != "bond0.100" { + if got := ManagementInterface(r, "10.10.10.113"); got != vlan100 { t.Errorf("named %q, want the VLAN holding the node's address", got) } } @@ -122,18 +135,59 @@ func TestAnOverlayIsNamedOnlyWhenTheClusterItselfUsesIt(t *testing.T) { } func TestAnUnidentifiedVirtualDeviceIsNeverNamed(t *testing.T) { - // A veth reports no kind of its own, and admitting a device nothing - // identified would admit every pod link on the machine. - r := stackedReport(map[string][]string{"veth7a1c": {"10.42.2.7"}}) + // A veth reports no kind of its own and names its peer as its link, and + // admitting a device nothing identified would admit every pod link on the + // machine. What it holds is a pod's address, and the node is reached on a + // different one. + r := peeredReport(map[string][]string{"veth7a1c": {"10.42.2.7"}}, "veth7a1c") - if got := ManagementInterface(r, "10.42.2.7"); got != "" { + if got := ManagementInterface(r, "10.0.0.15"); got != "" { t.Errorf("named %q, want nothing: nothing identified the device", got) } + if got := ManagementInterface(r, ""); got != "" { + t.Errorf("named %q with no address to match, want nothing", got) + } if got := ManagementInterface(stackedReport(map[string][]string{"lo": {"127.0.0.1"}}), ""); got != "" { t.Errorf("named %q, want nothing for loopback", got) } } +// The exception, and the reason the rule above is about what a device holds +// rather than about what it is. +// +// An interface holding the address the cluster reaches the machine on is the +// management interface whatever the probe called it, which is the same answer +// TestABondHoldingTheNodeAddressIsNamed and the overlay case give. Only loopback +// is still refused, since an address on it reaches nothing off the machine. +// On OpenShift that address is on br-ex, which sysfs describes no better than it +// describes a veth: same type, no bridge directory, no device link. +func TestADeviceHoldingTheNodeAddressIsNamedEvenUnidentified(t *testing.T) { + // The machine's own virtual device: nothing identifies it either, and it + // stands alone rather than naming a peer. This is br-ex on an OVN host. + r := stackedReport(map[string][]string{"veth7a1c": {"10.0.0.15"}}) + + if got := ManagementInterface(r, "10.0.0.15"); got != "veth7a1c" { + t.Errorf("named %q, want the device holding the address the node is reached on", got) + } + + // The same device, now one end of a pair: that is a pod's link, and it is + // refused however the address got onto it. + peered := peeredReport(map[string][]string{"veth7a1c": {"10.0.0.15"}}, "veth7a1c") + if got := ManagementInterface(peered, "10.0.0.15"); got != "" { + t.Errorf("named %q, want nothing: a device naming a peer is a pod's link", got) + } + + // A VLAN names its parent the same way and is identified, so it stays + // eligible: refusing on the link alone would refuse a tagged interface. + tagged := peeredReport(map[string][]string{vlan100: {"10.0.0.15"}}, vlan100) + if got := ManagementInterface(tagged, "10.0.0.15"); got != vlan100 { + t.Errorf("named %q, want the VLAN: its link is its parent, and it is declared", got) + } + if got := ManagementInterface(stackedReport(map[string][]string{"lo": {"127.0.0.1"}}), "127.0.0.1"); got != "" { + t.Errorf("named %q, want nothing: loopback reaches nothing off the machine", got) + } +} + func TestTheEffectiveSpeedOfAnAggregateIsItsMembers(t *testing.T) { // A bond reports no speed of its own, so ranking it on what it reports puts // it behind every physical NIC. What it can carry is what its members can. @@ -149,11 +203,11 @@ func TestTheEffectiveSpeedOfAnAggregateIsItsMembers(t *testing.T) { func TestADerivedInterfaceInheritsTheSpeedOfWhatItIsBuiltOn(t *testing.T) { r := stackedReport(map[string][]string{ - "bond0.100": {"10.10.10.113"}, - "eth2": {"192.168.1.10"}, + vlan100: {"10.10.10.113"}, + "eth2": {"192.168.1.10"}, }) - if got := ManagementInterface(r, ""); got != "bond0.100" { + if got := ManagementInterface(r, ""); got != vlan100 { t.Errorf("named %q, want the VLAN over the bond", got) } } @@ -202,7 +256,7 @@ func TestTheHardwareUnderTheChosenInterfaceIsReported(t *testing.T) { } func TestTheHardwareUnderADerivedInterfaceResolvesThroughItsParent(t *testing.T) { - r := stackedReport(map[string][]string{"bond0.100": {"10.10.10.113"}}) + r := stackedReport(map[string][]string{vlan100: {"10.10.10.113"}}) mgmt := ManagementOf(r, "10.10.10.113") if !slices.Equal(mgmt.Members, []string{eth0, eth1}) { diff --git a/operator/internal/nodeprobe/collect.go b/operator/internal/nodeprobe/collect.go index 173265578..e9a2e3e52 100644 --- a/operator/internal/nodeprobe/collect.go +++ b/operator/internal/nodeprobe/collect.go @@ -104,6 +104,7 @@ func interfacesOf(ifaces []inventory.Interface) []Interface { PCIAddress: iface.PCIAddress, NUMANode: iface.NUMANode, Virtual: iface.Virtual, + Peered: iface.Peered, Loopback: iface.Loopback, Bridge: iface.Bridge, Kind: string(iface.Kind), diff --git a/operator/internal/nodeprobe/report.go b/operator/internal/nodeprobe/report.go index a38c75352..eeedf56fe 100644 --- a/operator/internal/nodeprobe/report.go +++ b/operator/internal/nodeprobe/report.go @@ -207,7 +207,13 @@ type Interface struct { // NUMANode is the memory node the interface hangs off, or NUMANodeUnknown. NUMANode int `json:"numaNode"` - Virtual bool `json:"virtual,omitempty"` + Virtual bool `json:"virtual,omitempty"` + + // Peered reports whether the interface is one end of a pair, which is what + // a pod's link into the host is. It is read from iflink and is the one + // reading that tells such a link from the machine's own virtual devices: + // sysfs describes them identically otherwise. + Peered bool `json:"peered,omitempty"` Loopback bool `json:"loopback,omitempty"` // Bridge reports whether the interface is a software bridge, which a From cffa4019a586a41a80a12ad97f27130ed8252730 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Sun, 20 Sep 2026 20:40:37 +0200 Subject: [PATCH 101/206] fix(rbac): the document that owns the activate can be named as its owner MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit storageclusterops "...-activate" is forbidden: cannot set blockOwnerDeletion if an ownerReference refers to a resource you can't set finalizers on: , The expansion raises a StorageClusterOps and names the ClusterDeploymentConfig as its owner. SetControllerReference always sets blockOwnerDeletion, and OwnerReferencesPermissionEnforcement requires update on the owner's finalizers subresource before it will accept that flag. Twenty-three kinds carried the grant and the one being named did not, so the activate was refused on every pass and the expansion sat at Activating with every node online. It is an admission plugin OpenShift enables by default and vanilla Kubernetes does not — confirmed in this cluster's own apiserver config, which is why the same code walks this step on a K3s lab and stops here. Nothing about the operator changed; the platform asks a question the other one does not. The test is over the generated role rather than over a reconcile, because a reconcile proves nothing where the plugin is off, and an envtest apiserver runs its client as an administrator who holds every permission there is. It names the kinds the operator sets a controller reference to, which is the list every call site resolves to, and it fails for any of them whose finalizers are ungranted — so the next owner added is caught here rather than on a customer's cluster. goconst findings in mgmtiface_okd_test.go are folded in. They are mine from the commit before this one, where I did not re-run the linter after adding the file. Co-Authored-By: Claude Opus 5 (1M context) --- .../templates/roles/manager_role.yaml | 1 + operator/config/rbac/role.yaml | 1 + operator/dist/install.yaml | 1 + .../clusterdeploymentconfig_controller.go | 7 ++ .../deployment/ownerfinalizers_test.go | 78 +++++++++++++++++++ .../internal/discovery/mgmtiface_okd_test.go | 33 +++++--- 6 files changed, 109 insertions(+), 12 deletions(-) create mode 100644 operator/internal/controllers/deployment/ownerfinalizers_test.go diff --git a/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml b/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml index a01259b53..8aa237daa 100644 --- a/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml @@ -308,6 +308,7 @@ rules: - backupimports/finalizers - backuppolicies/finalizers - backuprestores/finalizers + - clusterdeploymentconfigs/finalizers - controlplaneops/finalizers - controlplanes/finalizers - operatorops/finalizers diff --git a/operator/config/rbac/role.yaml b/operator/config/rbac/role.yaml index 0090a347e..269aee4e1 100644 --- a/operator/config/rbac/role.yaml +++ b/operator/config/rbac/role.yaml @@ -308,6 +308,7 @@ rules: - backupimports/finalizers - backuppolicies/finalizers - backuprestores/finalizers + - clusterdeploymentconfigs/finalizers - controlplaneops/finalizers - controlplanes/finalizers - operatorops/finalizers diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index 6660e6eb0..7a3f6443a 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -10728,6 +10728,7 @@ rules: - backupimports/finalizers - backuppolicies/finalizers - backuprestores/finalizers + - clusterdeploymentconfigs/finalizers - controlplaneops/finalizers - controlplanes/finalizers - operatorops/finalizers diff --git a/operator/internal/controllers/deployment/clusterdeploymentconfig_controller.go b/operator/internal/controllers/deployment/clusterdeploymentconfig_controller.go index 7164f6a9e..71ec2179c 100644 --- a/operator/internal/controllers/deployment/clusterdeploymentconfig_controller.go +++ b/operator/internal/controllers/deployment/clusterdeploymentconfig_controller.go @@ -85,6 +85,13 @@ type ClusterDeploymentConfigReconciler struct { // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=clusterdeploymentconfigs,verbs=get;list;watch;update;patch // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=clusterdeploymentconfigs/status,verbs=get;update;patch + +// The document owns what the expansion raises from it, and an owner reference +// carrying blockOwnerDeletion needs update on the owner's finalizers wherever +// OwnerReferencesPermissionEnforcement runs — which is every OpenShift cluster, +// and no vanilla one. Without it the activate this controller creates is +// refused outright. +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=clusterdeploymentconfigs/finalizers,verbs=update // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storageclusters,verbs=get;list;watch;create // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodes,verbs=get;list;watch;create // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storageclusterops,verbs=get;list;watch;create diff --git a/operator/internal/controllers/deployment/ownerfinalizers_test.go b/operator/internal/controllers/deployment/ownerfinalizers_test.go new file mode 100644 index 000000000..d75541738 --- /dev/null +++ b/operator/internal/controllers/deployment/ownerfinalizers_test.go @@ -0,0 +1,78 @@ +// Every kind the operator owns children of needs its finalizers granted. +// +// OwnerReferencesPermissionEnforcement is an admission plugin OpenShift enables +// by default and vanilla Kubernetes does not. Where it runs, setting +// blockOwnerDeletion on an ownerReference requires update on the owner's +// finalizers subresource — and controllerutil.SetControllerReference always sets +// that flag, so every owner the operator names needs the grant. +// +// Without it the create is refused, and the message names neither the missing +// permission nor the kind clearly: +// +// storageclusterops ... is forbidden: cannot set blockOwnerDeletion if an +// ownerReference refers to a resource you can't set finalizers on: , +// +// It passes on a cluster where the plugin is off, which is why this is a test +// against the generated role rather than something a K3s run would have caught. + +package deployment + +import ( + "os" + "path/filepath" + "slices" + "strings" + "testing" + + "sigs.k8s.io/yaml" +) + +// ownerKinds are the kinds the operator sets a controller reference to, which +// is the list every SetControllerReference call site resolves to. A new owner +// belongs here and in a marker, and this test is what says so. +var ownerKinds = []string{ + "clusterdeploymentconfigs", // the expansion's activate and its children + "controlplanes", + "simplyblockdrivers", + "storageclusters", + "storagenodes", +} + +// managerRole is the ClusterRole controller-gen writes from the markers. +type managerRole struct { + Rules []struct { + APIGroups []string `json:"apiGroups"` + Resources []string `json:"resources"` + Verbs []string `json:"verbs"` + } `json:"rules"` +} + +func TestEveryOwnerKindGrantsItsFinalizers(t *testing.T) { + raw, err := os.ReadFile(filepath.Join("..", "..", "..", "config", "rbac", "role.yaml")) + if err != nil { + t.Fatalf("read the generated manager role: %v", err) + } + var role managerRole + if err := yaml.Unmarshal(raw, &role); err != nil { + t.Fatalf("parse the generated manager role: %v", err) + } + + granted := map[string]bool{} + for _, rule := range role.Rules { + if !slices.Contains(rule.Verbs, "update") { + continue + } + for _, resource := range rule.Resources { + if strings.HasSuffix(resource, "/finalizers") { + granted[resource] = true + } + } + } + + for _, kind := range ownerKinds { + if !granted[kind+"/finalizers"] { + t.Errorf("%s is owned by something the operator creates and its finalizers are not granted, "+ + "so every child of it is refused where OwnerReferencesPermissionEnforcement runs", kind) + } + } +} diff --git a/operator/internal/discovery/mgmtiface_okd_test.go b/operator/internal/discovery/mgmtiface_okd_test.go index 866018e7e..9a7564c67 100644 --- a/operator/internal/discovery/mgmtiface_okd_test.go +++ b/operator/internal/discovery/mgmtiface_okd_test.go @@ -13,6 +13,15 @@ import ( "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" ) +// The machine's own vocabulary: the states the kernel reported, the kinds the +// probe read, and the bridge the node is reached on. +const ( + stateUnknown = "unknown" + stateDown = "down" + kindPhysical = "physical" + nodeBridge = "br-ex" +) + // okdWorker is that machine's interface list. The names are the machine's own. func okdWorker() nodeprobe.Report { iface := func(name string, edit func(*nodeprobe.Interface)) nodeprobe.Interface { @@ -36,22 +45,22 @@ func okdWorker() nodeprobe.Report { pod("feb8ae72b6d20d4"), // The node's own bridge: undeclared, standing alone, and holding the // address the cluster reaches the machine on. - iface("br-ex", func(i *nodeprobe.Interface) { - i.State = "unknown" + iface(nodeBridge, func(i *nodeprobe.Interface) { + i.State = stateUnknown i.Addresses = []string{"10.0.0.15", "169.254.0.2"} }), - iface("br-int", func(i *nodeprobe.Interface) { i.State = "down" }), - iface("ovs-system", func(i *nodeprobe.Interface) { i.State = "down" }), + iface("br-int", func(i *nodeprobe.Interface) { i.State = stateDown }), + iface("ovs-system", func(i *nodeprobe.Interface) { i.State = stateDown }), iface("genev_sys_6081", func(i *nodeprobe.Interface) { - i.State = "unknown" + i.State = stateUnknown i.Kind = "geneve" }), iface("ovn-k8s-mp0", func(i *nodeprobe.Interface) { - i.State = "unknown" + i.State = stateUnknown i.Addresses = []string{"10.130.2.2"} }), iface("lo", func(i *nodeprobe.Interface) { - i.State = "unknown" + i.State = stateUnknown i.Kind = "loopback" i.Loopback = true i.Addresses = []string{"127.0.0.1"} @@ -59,15 +68,15 @@ func okdWorker() nodeprobe.Report { // The storage plane: the fastest physical NICs, and the ones the old // rule named. iface("enp2s0f0", func(i *nodeprobe.Interface) { - i.Virtual, i.Kind, i.SpeedMbps = false, "physical", 10000 + i.Virtual, i.Kind, i.SpeedMbps = false, kindPhysical, 10000 i.Addresses = []string{"192.168.10.15"} }), iface("enp2s0f1", func(i *nodeprobe.Interface) { - i.Virtual, i.Kind, i.SpeedMbps = false, "physical", 10000 + i.Virtual, i.Kind, i.SpeedMbps = false, kindPhysical, 10000 i.Addresses = []string{"192.168.20.15"} }), iface("enp8s0", func(i *nodeprobe.Interface) { - i.Virtual, i.Kind, i.SpeedMbps = false, "physical", 1000 + i.Virtual, i.Kind, i.SpeedMbps = false, kindPhysical, 1000 }), iface("enp8s0.4000", func(i *nodeprobe.Interface) { i.Kind, i.SpeedMbps, i.Peered = "vlan", 1000, true @@ -88,7 +97,7 @@ func okdWorker() nodeprobe.Report { // node came up online and healthy and no StorageNode could ever be matched to // it. func TestTheOKDWorkerNamesTheBridgeItsNodeAddressIsOn(t *testing.T) { - if got := ManagementInterface(okdWorker(), "10.0.0.15"); got != "br-ex" { + if got := ManagementInterface(okdWorker(), "10.0.0.15"); got != nodeBridge { t.Errorf("named %q, want br-ex: it holds the address the operator matches by", got) } } @@ -98,7 +107,7 @@ func TestTheOKDWorkerNamesTheBridgeItsNodeAddressIsOn(t *testing.T) { // cluster is on, which is why the node address is passed at all. func TestTheOKDWorkerFallsBackToAPhysicalNIC(t *testing.T) { got := ManagementInterface(okdWorker(), "") - if got == "br-ex" { + if got == nodeBridge { t.Error("named br-ex with no address to match, which nothing in the reading justifies") } if got != "enp2s0f0" { From 0cc90a62946fadd5ed9d089e678334f300ab35fd Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Sun, 20 Sep 2026 20:55:42 +0200 Subject: [PATCH 102/206] test(atlas-lib): a second host, captured with its NVMe controllers unbound worker-0 of the lab fleet after a day of failed adds: two of its three NVMe controllers bound to nothing, the third the kernel's boot disk. It is the state the scan has to keep finding, because a controller with no driver has no block device and no character device and is invisible to everything that looks for either. The transcript answers it the way the first one answers the bridge question: off the machine rather than from a fixture's author. The capture script now takes the mount table and swaps as well, so a host's block-device reading has something to read. Co-Authored-By: Claude Opus 5 (1M context) --- atlas-lib/inventory/netiface_test.go | 53 ++++++++++++++++++ .../hosts/okd-worker-unbound-nvme.json.gz | Bin 0 -> 130197 bytes hack/inventory/capture-host.py | 4 +- 3 files changed, 56 insertions(+), 1 deletion(-) create mode 100644 atlas-lib/inventory/testdata/hosts/okd-worker-unbound-nvme.json.gz diff --git a/atlas-lib/inventory/netiface_test.go b/atlas-lib/inventory/netiface_test.go index 35c0ab168..0cd574f3b 100644 --- a/atlas-lib/inventory/netiface_test.go +++ b/atlas-lib/inventory/netiface_test.go @@ -11,6 +11,7 @@ package inventory import ( + "context" "errors" "os" "path/filepath" @@ -412,3 +413,55 @@ func TestAPeerDeviceIsToldFromAnOtherwiseIdenticalVirtualOne(t *testing.T) { } } } + +// TestAControllerNothingOwnsIsStillFound is the machine a failed deployment +// walks away from. +// +// Binding an NVMe controller to a userspace driver takes its namespaces from the +// kernel, and rebinding it to the kernel's own driver is how they come back. A run that unbinds and +// then fails leaves the controller owned by nothing at all: no kernel driver, no +// userspace driver, no block device, nothing under /dev. The disk is gone from +// everything that looks for disks, and stays gone until somebody binds it by +// hand. +// +// The PCI class is what keeps it findable, because the bus says what a device is +// whoever is driving it. Captured from worker-0 of the lab fleet on 2026-09-20, +// where two of three controllers had been left that way by the day's failed +// adds. +func TestAControllerNothingOwnsIsStillFound(t *testing.T) { + root := hostFixture(t, "okd-worker-unbound-nvme") + + // Collect reports what it could not read rather than failing: this host's + // process table is not in the transcript, so the holder check has nothing to + // answer from. What the controllers are is read from sysfs regardless, and + // that is what this is about. + inv, err := Collect(context.Background(), syntheticHost(root)) + if err != nil { + t.Logf("the reading reported: %v", err) + } + + byAddress := make(map[string]string, len(inv.NVMeControllers)) + for _, controller := range inv.NVMeControllers { + byAddress[controller.Address] = controller.Driver + } + if len(byAddress) != 3 { + t.Fatalf("found %d NVMe controllers, want the machine's 3: %v", len(byAddress), byAddress) + } + + if driver := byAddress["0000:02:00.0"]; driver != "nvme" { + t.Errorf("the boot controller reads as driver %q, want nvme", driver) + } + // The two a failed add unbound. Nothing owns them, which is a different + // state from a userspace driver owning them and is the one that was + // invisible to everything downstream. + for _, address := range []string{"0000:01:00.0", "0000:0b:00.0"} { + driver, found := byAddress[address] + if !found { + t.Errorf("the controller at %s is missing, so the disk is gone from the inventory", address) + continue + } + if driver != "" { + t.Errorf("%s reads as driven by %q, and the capture has no driver link for it", address, driver) + } + } +} diff --git a/atlas-lib/inventory/testdata/hosts/okd-worker-unbound-nvme.json.gz b/atlas-lib/inventory/testdata/hosts/okd-worker-unbound-nvme.json.gz new file mode 100644 index 0000000000000000000000000000000000000000..78ed6df0c87310ffc6ed513955bb59b5a501b878 GIT binary patch literal 130197 zcmce;1yq!6*EY-mQqqWYHwe<*ASE&g3W{`hcL@wKbR*K;Euge?cZYO$mk9pX;O!Il z`#kUZ|9`FTW1Z`~&VBBE?0xLLuNa0IC_+$B;D8_aUp5wotS&b87Do20PS*N1PS%F3 z*3MQ&Y~~I&)^K}`+aE~;qUnEpOQRM%9G~RNv4v~eNo`2UKq*j}*T|IW-77R3Smu|{ zYo|f#S3fkQz~p~ zf=#}#DFZe|!KMn>Bn2~;g~1}KId^xbsE4F=eo3=be(~NgHJY3nssc7QVd{r4LoL`O z0-L;G69R0y_do+%8Nep!y)PJy!UAxR2J<2|()%Rq*Ts2>!@B>BQPiV`E1TYg_ip>P zM`%sR(qHLkufMgYMenPS=w8Y;8cigre#J_@z22f>eF&g*rR} zLJ2n;5|W0QFnVZPTttX-4O+hK@oWW1B)gT4)k*rpbB&!(Xz|(OuAf}xWs8JU8*#-r3S87 zStnuZ#nWyXci7?a8*ap6WxmWvIMIA%d!Bq9zbzJ)q(p~3Aj6tYDC0(`?Zz+T#;@%r z4Bd{;gOgwa2abqRY2OHUFb=6UoUDCiMMOhUag)5gU=ly-ZBvF#FM){$wpz{fS&HkD z0*$amSbGo1KU?2mRBghc8_%PGCgD&wH}h+Ls4;~W1)`~s5hjitvmZxthke1U)>&o8 zP@LZ}Q!m^L+F4aCPCd2dX>T639DO77%{cxQrcgfDj-1HmvU#X)#!M#7{?9}aPUg$N z&k^~H=ynvPgh&C}A>>qcuo^Z$6Vj}GRpYup9aw7eRs8CfAM!np_J!0a z`3REi79w7R;7@OQ`p9!-ZVY)ytciX;o=U`WCRr&zKAZQ~xZz`=a*(&p2-|1!;9M%R z$j{GKc_DwXL!76VV@FlYsOy*3YIY-Mn_ceBYo*)b^)t#PSOPpqGHV9&QZ|Re_;LJ> zjS1sH!~7lSxpvz!Ge<#4X1oxOlHv}hme3+cL7(gXCwZ+-i>?4NBs?j|tHY~QtM}0Tvj_g{uXiqT9 zQPA}n(JdL#tr_mF%9L|&76BDvhRkgVgFsUX;b-4-*V;Ctl z5x(nv9^bx%EUU2{H{mKI7&5a)RRvkpjU7-N`>%m!#GZOZ9?xtfAMNmMs6XYJ8bdz? zq?1H}X_G#IX7blP^sCx&#pjg%OP003ee~%j#WX2`WOamsc)~c_3y}S_9R~?R5fA3H zW0SmSac+_gzwPBW{bfx-SGoAC^~YG@%Tz`Wr2M6d4Hk?)`2)@Vl0tkG zsO&a_4bqa%Z?FvXA*X(&%_}WVfy4Zigj&}JB`zKVhh5sgw=V#NkP5JGLue-nI~l>i z8IL17Jh9Sf(+&3-jHGilr^6s zD&=hJzda8?hkAINfF2*R?cr=0<453XDu<)Xj-@7NiU>UOKr`e|ym<6zpNVmH>{&pT z88|*uM63To`pIJ1oMFhT!t*D0$&}|0Bn=99qoqZ!nsUFZ<+jU%v}evXVNJqN?4=Uw zAZ-xnUlwVj+Vs9UmVeCVaFECPw`NzurBdZ@OncdKe2y1gd%YZrs9Noo7%9@8=)u{= zNej1}m5jJ0Jfpj+MSn`CBZLPv-p20cD3;1b9(7J3P1S_jI$p|mY+SdXC~F}& z-~PHUPsYOhXWzt>Y&a@ie;_W`YYtTcYv~a7%nqK+4zA1&zDzpyOgipNI?jxaz+V?&>186-mMzaKdCyIZ^AXhPB1Wjsd%pA#R;(O{ajQ*6)M8R$<$Gu4uM z&(T7DhO}6^i_I6m$TyOSP0qItCt_tNOi3sedcA1JrJ?jnNKd9O$b2(##6EM@&A2K# z-?<59p`{DkCb-{z25jd#%WNG;TXvD%V~k-+lTLQ9Q9^6zN4$wao~(80>eGHxftUN9 zFpK(T*ftwCY&!(lez6~l-5PWy&VL`FKlJ0+_#jGLT3}BY_H~f!ca*TSz>cuUt04aG zC?C=STf(qwUf$6Q`Ukx~u&;fqziRZ016TR%hr#0*y|f+UM|pKbCk58dmY1REHS1NC zZXd4bQBKcy9|{ky^Kt1_lPrJq>)w}ac4Wf*@- zqg(qohn?_^P$~X&h_A`3o4JGIWdL;)h;=So6rWBmF~U=R@+zWgTjmx*>yE!t{j9Z; zzIV4Z&vtsgKyaWgs@t&LgrWKknLN^KpD#zfqH6v#Qs8nHkMe*Cb>flAb8RuyT+ijEHNQjv}FIkqx(GUy#Fq|ls z(u^VbhHGg3Q$aB->^cngtB-Sc{3~~gmscZ>5{m}u8}`JYh-dZ!DP=*CBk$DS`;Noc z-mKQvO%tAa78i)Lp3Ie%#CsX?ld?#FS}ICUPTpBuVVphAkXT3PJgNUp8exCk1%H=w zp&gx!9S^duU!5UE&*plJ+PI509T9TgN##e!db;i&8{RxaR|{V#(`y=t)3Py{)#rfi zIjfN!xK|++lB%8q$Dk&?hV=oH+7|asIHnwL|LG`yf6#JJq7X$3AT;g z3|>n##_8*OFzNeGPqXkaPw}5g-zjfV!goQReN_m>cfoj+vq1Q|qmZoi)<4}Dfl@r6 z&CK}`_U#R-#J;!RbD*Do&e*tf@Uqt%1t=A!{#X=U( zbm&=wiQ^Tv+2UfMz8Q%-26o$yMf$AsOR4XrP;|1RXc7BV_BC=rq$b6WkVu?1yy^=H zR1VLJ<(FSvjp!XU*~Glxfn|>^+24h(3^;P|Wlg*y5d>}-x%e1~nK=2S#Jsd(LD13j3fqE0szBdpQ{;glW$*Qkwc&~%L0Z*7#ansjs$`(W zABi@9ipYUp(-0b#&vp4xjuzaoR^M4x;+8c%%WDTD$)dhJ6{-l&xb)h0d8g7m>%s% z(gP_Gx@`Y;vW?OB(UEKvT~$@f;fVayYYJ83?^l0>b1jttK}klOltrc`*|;6$-@|Le z<8cx7bCWL5Hp(U0ut3WvW^w?hp?tLbRJJlUN&W2X7<$bqxXiBe^3Q2L4M!Gf`3Xy5 z)m9uGJpTO}WdoLVg7v2|rgZb{eh~!fH=p&Ig2C@X9&t2NZ1ipUw|PFpi4uDx7}zYY zE9EHuSx@vySQ{0G@JokHnm6MAmNfc5lWxcUThhNT5`w02@s~ScFJJqs+N{oUR1*?m z=G7{TpMe|0it-j>;X9cx8r@yu>)R(}l@G8cJbSg&cUX~xvrrBxYe6!F$bB!^VmW)N zzDlbkEHFC3nexI3j&?q$I3d^*z*HSVE%P?;KIe2087PmZ8EQwSV z&36tw^R#WQXQ(;b;ltPp`H*sjja5UZHrITNbj|Rd_?At>Y9tYtw|zg4`x@PwB<%{UB)T?}BFl@v5s_3+l41pwPimG;5{n-0+%y!$ zTx(b3Q)oYAsseGdl;5nHBQVAzxGxpct=?(H`EvVtrf5~WkX)MF#Jjm2T&S)a@7fd*O-MFnci*Jfxs-}C%Fpi|m zhV0on`pHi0eW)&qZDO2(TF`Y?zT$%};1%~Q$;Vg;a>ah!^@xye+Bin>R9|`8xQ-26 zd82ZPfQ}_Bi(L5!j3ev&yYCqK#rksN8PRIIbuA#mp}oymS)5mG`RFI?%m`0;i9A1i zq+d^$)*T=%ed~h@qImBiry(aR3YG3@8KBbhWMjm6Q~7Xe&llI1$HR6&6eeO#k@^54 z-;n!8-;^%amHzxy$&hn+|vOUk$;BWO9kttRmk^d9z9vKmSTzK#V_rqCIbu zd)c#7amv47{TS`-kB74#E9_*ZA}ilXU@4egY*tqXncg(XU;mhMkNK_>{&DWh;nUTd zW8K-OU*mX9Jt>3wM9$FBLWcYIViZ51$;nA}y^zK$I~ z0nyd;1`NIym_!}QHLqVLCa3nHFb}1g=YU_ ziYBL=9ZMS|E}Kc41$4`Q8Tt!~tqlqOZd|CR#wQyCwEamvNqmq`e2@gbT|NMCcrXYG zUqom3;krT@T$194L@c9=267EE4jKm7(Q}1!CHN*Yg{~8aRVG- zgAd6+2OLpDB!?@N!I5O=(g(kmheQwpG^D5ENKbu`p28tL)kQjV{)NWvN}zkh_I*jK zrlsB%;_E8clf49t zkjo{c!xoGM1fm88gQL{Y77tFOLngv|!LP-G5_0K-bO=IUfDkn--m(lYdHJ;503aYz z8^!{ZqydZ5Ri|^wiwJUQgubB8elI+LTuLAvB48{a6Ez+xm#Adr682f@Op*8q0B|+;T7Aql+REd7^M%Ve1}I^)F=|==dU0fY3$#NIVhw zu28IxaK10mM7GER!FmWiX$_8|{O$!NLQh1_D-_csoMxEbg_rJ#O?Zi3_$RsW&&R?) zF@=9332)ui_wTzT)E{#4E#eAL`JRYmw;z;i1cA0{VlcvMuKw9zJSpZB?t-_&0}B*%Qb_3 zqo5SDOGgkyQtwC~2;edSf~dJc35pkxYsUXeT>kC488EVq^l8JyJv#W}ucSxmKo}$+ z+WtgFpO}Ut!uN$?wlSxQDp5ggXs@A+Id2wcitm6)Iv#O%C=!3pb}91_;0B=dSFARy z-}C`yM*j?}&i`Bd0K{?xas-T;0s2C6&GWDtV*%>FlKn#)uxe zK=m9VZ{Ts^Z#$ zalN18dduQ^OW}Hp;W`SbY*=Ua56C5WM4^49Had`I{vL=p4@F!EK>Yr}f8PJmdnwD&kK{UmpruY@q${-jU)yn7AmwsV2EW=^H5^NBQV4ahcK=j*J$&88EVq6mQ1F z4Lk?{nurbL4MR0!;i3TNJV;EAQXCo8A4rtSAyT}XlolIkI84j{P%zg(-XMNCN?l}B zL6-!dBt;()`m4uyo>0`L7Q{szc&n#9du%+CE1u&q7>O_B&)~`5LoSFP}a@o6ty`P(quKLi-_%Ho=SbgB9&C8?G%LYB>~f zPIBkF1iUpR=ptfiPWEV?=4u1S^AM%ViMG;#w$hHa(k8pyinh`syW9+MmJ8mR2;>~9 zOCpnU^Q?KFy4f1Xvl69A9dVJ_@9>c?Z?NG-kl~jgL+7AJyg`N+fseifiVI0MV^BBW zVOx6SJGh|K2p~?A_}O5zxCAVbNF8}m4S$L3FOBTajqFd2Z2cJ7A0F8nfNd&u4z^bA7u41DwqNM#x>dgj;5={wN=9UR|# zI1jP%lC&P?>*~YHS5uaPKQ`dTI7|iEDG8&b@X$;p2@;1P^k*5U2&S#riR|$7{=78M z3rb)h7X{oOe@p+EtDF$4kRMbbXAEFiS(zG0%@NOzr?Acoj9-SG+#+~hPB%3zR^45= z2_-QUx$>vH;2SfX|8yW)t5e*xf5-1bOSaBLXNXCQK^P`6>NCO*EMABM3r#s<_C#cP z1esGRvu0Beu+8e94a#~J$*s)oQ|nt2CKSo7%Z&sWKwdr4Rs{hI2bd)}U^xH+ zD(q;a(8m&^;v;ZK@L#Z&VU7u5p`?HX=Q!(e|9iJKEOgk5VY{{zlBMGJ3tun1*|w=I2aF9815mArz2AX*b9*U9>#J> z?eu?!slCs1+mDq%w??=A!6&QuiNQJRGb6FJ*6r)-lhq5U&!J*f-v=!eu58Z(1+2Fo z+~wslC%7fnikXa{K;(r}tkk{Y_RR{lGRra+Q?o(Kr;gP#ot1sduI@dJTd0TEwODQV z;N$}x5K3`fWa%$ViQ|-}MJh&|!y7W6)}pVkcD_j|7iO#fBJ;xy@cM?JjCgoWCntY$JOemtNZlak>jtO z%#d1Xg6G@`kz4vyBOSORR3kXJx0E9oxQ%dIh=d4QcXcnm$iM##d&3*f)U+~Kw$Zln zwky4%YpJ0NF0@lTv~%fx&=Q8pvlGYYa?`h%D*BClLd@bol`+5#?fk8PaQ%1PZxH(6 z9~^%c$KM8F76&jWVe@$%8I+S4%0-OKPK?X}TL2?omz%A{RMT&y1~JPHRb~|kOGg;? zoc%8a>EM6Y{RQIxgIf;m{M*nA?F5D|Y@W0rfbTVhA2fw;Hid7$UksW``i=M_W}%|W z;NgbS(-OXrV^56K<)&)+qu~EvbiYAMe{eoriT^%`Z5N+66TW~WdU3`z>&fr@!y9_^ zHWgFa#I4twcaE;sQ9Q`Kwm3GrMODE^A>H0p1q9w{jd4H978N#w0~5E#$Sh?a8LuCC z!v^jHpGy|#{}>G=hz&Xq8g!P);B4Nzp0>TCSh%{iTI_r6R;XR*F6ShmRV+&pV`sDD zt=$(RVOmlaLKHbk_-XQuSP2J82?zWA;4k6$Uj|_Fzi0u~KMa0*`_~z}gaa5%8V2mJ zvCD(J=RCKO6`{f>*p|3j>o*+?Rj!xsOG`hCz5cU+V2N8?qH}~ zi+Ou0|Kx|EniprEXpmD{dn;IfouT3^=FN@#lRKXxv85BR{VlQ<9K5K;AOzV1 zLjZnF9Da=*evgL+L?;en5C>6;gQ)Kp4B}K0K9YWL9}q!ksPK<)5gBPwl;yC~BPst< zJV^fUy1ziSe{jvH%6}gy%FQUs-8kuRxRiVC#=m^FB~gbz;H5vL6ep1OVM~V7;s@2? zB5o;Qr?W`-lp_8|;UR;#x*wbtD!d~d%5WUze-xJg`1s4GExaSJ)%rUajW~gfPbn#= zP8{A*4{>^qgFF^agdgM%2Y)7xxYdF>48cohknl-G1i1rTI_INoy}nYtKzp zWnldxuzns`KMSm%hSfB{P|?Rw(Sy9W3b8qEak*L|@jSW}xw;^?biw*k%kcFp^7mQD z{@yterlSs4Qwj3)1cKF0%jMb9J9k_mcU%&8+#5ICIXB#TH{2vQTn{(gcW$^sZnz|F zM_$4~R>DDA!a*;DgG7XbScQYggo7}IgKmX__JxAxg|0SI+80vVC&@fFF`Mj=7r#Pm z=2~3#my(>zMTDk5rhE@4vu(rtqBAvRH#{$RWlW}9hFSF*d9eUuGcLb3CT}?=zc(sx zIV!(5B5ye&zc(yzIV`_7ByTy?;<8E0yGeVoLHlKc)_H@LcZ2p~o%YK*t@Ao9?>g

~hBILyoh)qX}^NKEso&Ak%vlag|o2HJ7_~p{lJ6j(z z{$V6i&tDi#xAYY^^cC0i6_@lC=kyh)^cBbS6^HZ{`}7sN^cCAkiw-f<_A%3TG1Im& z(>5{F)-lsoG1HbY(-tw)<}uS|5SuFby-In@$`+TIr6kKE5#FoEB%V-=rWT|*9f(y~ z&weS0Z%Wt0!YgbK!E69|EO8NRA?iQr$zF1;t8uzqwb5oa03A!j%HyF>vJg$4^yDwO zw*8|3%>T3EpNLO?5G_?3koy2wxjYn#7NURxyX2Zu<21)~06LP0m08`C`_QNAv5y(+ zR$eunqy^~HBs_b`16z%UOI44X%~*ZLG~fV*k_9Mv5}yB$0xJ{`z)5)Qk_Re408>`D5e*%PzqrQ3ma4}WV;U$K)!2#_RG%g#`Aa?kES9St zx0StoR4Q z_a6vF3!G2)0dzG&fN}r|_z?}n%xbydNovB9XIfRfVBYv}!icN(7Xi!s$B4pu>K#eZ zn@;cga=&SMd9CFe>MJCSzbe>XQK0!7u{I-7tTHEKj;Pv>eh#_Xj&9Cf+~GnPnuSpU8+4Jk zhw_&`d(r}!4>uI{p5(SW7n#$V-eC8O&L?u3Vgm;dyo(Ai9yf|PzPv-dn+runrkF%# z9xY9#c;()8=Jk_U`)h?-_#Dj)l~Osncoi4B-!vu^^6%r6t9ZOkfvkA^+sSm=r?vg1;=wXmwB%iFQCLdTT^AE;UY3i=P0;;0 zE{+^U^TKgQyBtkPrud@RKGoiJmUXmP`wt4W)Nz^$^M8_O$@urE_XRRkR={`~i>?5a zLxqCHsX_&d#i?B7SFg*V%!g{jnI<#@y_;sRWST6FN$YU{o^7k>mooZ8gPV;r5zEEY zASNRf)tuZAmUUc(TFo5IR24QICiV1KmRG&_1#2ohY^A@`)j<7Q1AZ4k#oRT+P+1b_ z$B#D!A#v_+3LtSm)SEo`n2YK9E?yI2&ta_&hcJL#w8-gj`1PL7sYhRa_A>Y(`o(6T zJ(jbeKxM2W&K@;Kvs8r*%oJZ1%c9(i&%7S|%AU!I0M>pWdWrx4Y(Eg4iKhW*KM-BP zt9ZOaL8N%RMd6lZv5^FSGTQhC4c&jgjS5V!=vqNQ_awX-^=7&$u_0nVY3#Dx%Y!~Q zR{lj?B3o(fbH;W1ID0pRTF)F!UzJj2CiU#-=LPkDwbz{P)o%_x>Hperg@WY|(GsY1 z5S1N+95%Hby&O!n9i8ig`TX*aPv+G1C@m=>-Y!8`rH_u=ry6)Yt5R|+3Wvyx68=sbb#L+jV^>=K24*@*)cd?%_gbMR>Ama zg2X%K*tC5^Hmv-XS^T*O?XFE>RuxH4RdiU~dO(#Lr%riS=_>Gsf|u1rWO><4A(Mwj z=|@kRG|k*wayZHtS5yfMX~?@ayS|71=!vCe94rMl8dV%_;7Z=j8ODt$vn4dN>^PH< z(#0PdTrWnZ{GhZVB^Qy_G*L>qPNT_@kKcNxiv4(0N-_Tw_@G=9ub|TG)SL4vk&?+s zy{u@7T(Sr%U5q;K*x5DuT4?sR&7f58-DENCzE!{Zk1$F1+FW7Ev+A`CgU@x~U*V5K zfDhW~DgQi>qZ75B1NWGx_Ol%xJ{7YTqwTwDYoa6ZcCB%5C*?bX|23v6r=Lo{kf+H| z+2iud@%Y;)*5ZrY%2#tSA-^0nFXGF7U0fX6EH=7FpGbM%o^5`eIiEP+f7qor*j9~a zI7N}S`e>#8R{d~uJvZ`0y-w)r4EJF0xfbxJl{8ZC@n_ndNfAUlH0qYP(}H{}@4(JQr-GeahrniK_(( z?tQ5x!9Yp}?2!#hNtgq~0tkvbGR`o2)&@ z89RPf;(H^+bQTKTr6f;)xLRSZn60h|1{NWhyTqvn&* zP74L*dimF2;zkR|O>a~VA^sl7BUKS7Qxepi$U(EDw4$9}InwlyGVuauRZvcGgVHVuq6yi?pQ=5`a%J|962))bN<-5d2Y+DS$&;L3F2K(H znlWfg@lek&;zRkDE=g{i*X+Svl7!m^RH1hj*fO_oP_n$A_=^ZX&NLgX8T!Efc~kPu zoC8(jS7yQv2ddDsT5OqLnt%s$nT{vVXGVk zy51R90Q(NI6yZ69YVSxFWo%B>qy`+o{+U7Q&0NfHmob3L9>Ap`;F1$?Nz{r74zL}F zUJ)g3?Iy?2vx*4VS}06+N2coXSc~FOOiC)+%>+@%l%JkhIRGVvIDt zmN>RVp#gSA#7J86(*WC{fbzCVDouxs;2T?jAwR&7w6*HF;u2Lh7L@%56yVlI75ZZt z+u4=h--7LxT^gpIAsAZzB~QxH=C!ggS*skda|VV2?-;@Su8!nAt2dJsp=luEKPy9NM%-UM29O0{=?@*3ffejIeosf4vEQ-Qc_ zW~a>^)$t9(X8s^0JfJ_fF2gs)lFxw6Pw|52A)8&1RYnT(vm!h&RE$8exCQ zP+gwsCL3bqTddw{eAK{kFtv5$=75baa$T2Pnm=pwxNOY+eds2I-0ZsOox z6ujcmdg0b+%4+0m{mSWfE>1@@KGwUHPc$qeXk*!N8{6Yy(%&PEl?^qhlDS@jV@*wD z(c84&&S(oTmsI)ET-#ufN2%>h2Ue*dbmhbm3saYjmlsvFvbO}8=Hw#BshMv*vJRGK zu?teGp)YB-syflcdX%81FQ^d|WJX7y6Gzylz52QFM}E~PQux|89#?6N0t1>+1q{-7 zHS;6DYg_54LMW<|Gz*VNewB(;o~^2)2GK{;XQrdQ5$tkc% zp?i>h+>cvccwq5N_GC#=l)94JheqW&8Cm?hovXk#738wdec1RZf!aI4%GWuV;Fa<~ z^tRrI1nLobZ{}J*aET~u2Ix6HB#>#a)ysZJ5Kv*O7e|C+@Ckh^Jbbkca#q7=Dru(vy4spm#pC@LDx7v`ur_vbX9f#UUnzI z+cQJLREDedj)6+P?ygSe?i-u8&yWos2@+!!&wpMze)_(VSlHXmiz{6@xDO)Yel&xV zkDb7kp}~<*^*XPqUM<}PIptC(1K&+h$C~fMT6~r(5msL1wsE3lXB&_8ZdH+j=w@Ck zhH0Xr;>5P<32WJ_=ML+_2`yS3Vap!?IQH^A#uexS8pJN*!&T zM)wP;sLF0h)#U9f`x-5qCe{$a=M4(-8bkKggf_ArnlpBuoFZlxVFX1f#8Bk{bKqm* zLFfHqO3*P3CMq(q^0@cshM#6yHjCBzit{)hJCoB5Uhdt1hmsG|Dh|_v4$~?R(@qc5 z@(2zthZ+ynk8dqSt|e7lZ<#OSxB1^{3eT6CjXwcCGf6Ro_f*W}N3c%N?7~vu`{0Z_ zLxZ0|NSBU*@92vx^(FJyd2|-udGTzM3|E*n z8xopf*RiIoZ++%^w4PK?@J+b>hRS7-MuVRS$dNB2OqnS1&B?*7P!uPh^ZC^pbc zS-c(VXAx`-DIw&~Zm#EF5_4*Cdk~8#a9r}azsgt?E@5t-zvLHo7_yinCw)k`D?(uO z%I{J$7Qadfv($QVA(KfWjbyx8J{PG(`&2>qOT;)6py z#QFn21EZ3xP}$_Mmb)^;W3u4#O?*Ji!zXjUs0@bEzDn%gctj=`mG_7&QQ7zr*;Nf* z@Xkw-+z(&A6%Hwt+rE2ymzPwOR^gZ)IUyV~l6*%TGcYmg-Kdo?#e4A&{DyU^JxlR08 ziC{{wG0_PvR^J`ULUAo_5)*LyTAUkfHaW5L{{7ANtefZM*btavaiNih6_nuOyiIfd zCQx7SDUun($)53i_ky-M4lyB!+_A+opbAmOi(WWx$e-an;mBVUZ}djd-}7tJJA!@_ zx#zm$CYh?&ijFF2K}uU(IE_Q?JZW}4g?I>Fzj2~*L+t?W3c!tBD}FoAoazm&CFUJf z{gmqlk>+hvT9slJ&Z#D<;Bfe8)SDVvlHV|ZPh|tS+>w|gvYhg-LtV%~a>8oR!n%0u zT2DVtcI3uyA>2E>|^5NM0%Ei;ix?OLNRWg3nGd)fg!zae8H@5`{fKWS*?~G;Q7H}hEKx3jxFJQQ9j+L>AN~H{P5cdXbFb4MZg6GOv2Sy= z3PcT|Y{rVxB#&K(+HQH~F1n9c7f-un3bXyS-Y&-Wez$po5IJz zLaVu~nhSi6^(dtP$}1kn(nb&L+M>zL+KS%)9k8M zlP|4ioS$JnH)7{m1EPq~OKd(S%mtCf=@B2n77eZJei@l#?1~U-#Q-xO2D>+k0n5O5 z<-TPtfnZN-a%$X0 z?dW_nz(eR~QI0vK8`xp(6y{QPfMlv5rwJjl!!RD?`NTA(dt2(OZ_uzoOxnWZ3bn9p z*Z*4AIZ#wY!c9k8+7ecS#J5TOJ=WsLdrG{5V@MBw>c}zt2?Fn?Iw=<3 zu&1%jJ)RjaABH{UE}}mAkdg`2NSENPAs{|;3{@R3m(k9E1kB25Q*}1uUPMmK&AtC7 zr7F2no`G|MHV~z58r;z8eIi*=k;gOT8oWdd7Sdd0(yw@@wR$kmqrqybxmtiTVFWgWiMY5^jJ`3oGFgv_r%ctJEfEVnb1aiC z+iFQn*CNpQsL9$c`Ak!CAN!rxZeE)9ye@uH?^j8Ls%w}Qq{lKe^x@}Nyf!+ zK6$U4aBPML?@ZacU8+_9Ydu#ucGZJ1?^L6ASe7P-NSSW?S$8{Jov>{Mr$Ax&4HX* z(JjCEQdwd%1L~9?0@;jq%?~-qu!cfK-#bHITS{U%g}LTi3`%LKlJBf|}_o1LWzhBy4o zf+K+ZEFWAxzAR7r_!RZ=x5vn@O*8h(;QKF|KH6TYXMiwnicn4mJ~!5BEEr-Nl`Sn3 z+I-x93gXOZJHx?lO4i@o_Q`f0?~zQi_Kzxf`|SVCed=c>bwHgY{FX#wn6sIrCXZ90 zgr=C&c8iw06CkQO6%>?~!c&V#?-sMg<+GHs-QT5{i^FQp#77c>Bh&KV27L@s4lQXI zzKnFsdq6gTS9Fz6s)`_M^pRPRH|^E&%eUER+UuAU_$i?rg~5?-nP|YPuTRO!;qZOj zhXNhYB8~pM9z@iy32oMu@!8uA``?NY5=M8X5!WV55tB9qxifGVD##JR!SAeOi$bqD zB~7Wv10*ak8D*3)-JycFW(cw?@x;O_YkA>e{Fep3t4dzLD@d+1VrQonJOe+x0S|L` z!^P40$Q=;3Cdaau2tA7Z_p`Lyq+vnKNRd0+p~t3zhe)lc}Upt$_$i zIJbE%Ctmo&X+nVvB-AtP6tMnr$E=||AO^cIz`ylP{QV};N7JL6)-je63B6tIyu5B|Mq-HZBC-~)+ zQ#6KO;d zxMrWP6X!%H_0Api!sSvzwNJ@+=}&vTcb}ET5jpT!X6LcLI7N~-42T)^*~delj=7E7 zyoZZk;$l7!h*i*Sp4<6>eyvrnk4HY8k@B4&!YRgW9TI$!p)BFeiEuMY`RHAU zJ=jb#apxppq!3<5deiufW_wy?LVj?G=fhfq)`Y2FNs(ZWKwV+E`-+!j=?Ei*dK`)F za-o6&0v5-r`?|UU;#J7qtG2t~pASHET8&(&!SzGyxWnKl#AgW!`i_je>@!=VN+SD= z+qp!@3e3F@w(;)F@FOqy#+}XcQc@j@a#Qjhcj(hTI|5>ZV`1rn?h=v{@Ev!@#}Ie( zv4Wq75RzNRy!RSDWux7~&FnX6ey0OmA_)&;H?Qgc7rMRup_*AQMN66anpKxzRkzgtJd%;(Icry(WJ-z+_qX;K z2N(&S^lY}25TnjCmd3b}e@~BeCD)vi6T>RqrAlQGUY+AR42>UZ z&PQZ{WN*6U-4zmBf1Yik+3)G>b1?^Bq;Q=WE(-7gi&NMljA2SL(Ew_{J@IrGW?>ID-dyY1 z+Y?SmNnofPbG5(Sw$`hf)%?qS|4>iN^^6;*7#DaW@7VC==~}mG^@l13lBbJ3B5#vp z7(yyf4eg9mpD3;NoFivyW=%Nfh|IWND;@M42)KRra8BJzO<<@8pk*696(i;zqp6-o zb)F2DhKI?*GeSb(3xIbcEbKD+kZHA-8JJhQDaPaTd!^L9v=oLLlZ~D`Ghu;zx07Mh zY6LT|w>b>Fwv;R^z%K;G1i%m&IDv4{v|8#fu)V8NDrss80~7!+1K>ip6M!_08Muj8 z3-E342&hw17-RwUh?#Kh9b@#2bzqLjb*CeY_OXDQy{B_3X<7n9SKjfO-argq)&#kN z#|hWix=<|x7ElnV$!1R**XbJIy!H-py4|!|b~N?p_j`Dj!^zsnjN8dpzp$GlF}PIZ zbx*&V$CLe@_dKxRCs)bBZh(7XQ7$lfT4#*K5Cd1%(61jSLkzzR?^APd7o8ES-;BBC z@&FbCr#;^L*M`Oy1rpbrF*={vfU;jFr-Cnr4z3M<^#cnGv9u?vVluPKj+hQe*7j_J zcI0e&%{V?A;5S-h7v7vKVkK5-qnQ}fb^^>9vgEPwz`ZW#Mr$?*;k9ESJ?7hyoFq}# zXTZIW{jpYI5r|a=$%zErYUDNg2_=D6FPOdNC32I`Uen6o9&GD@Pc#142is}nzu!Fc zJ284~x=2>Paq#hb60q-lJ^ko-K1kczWm9JQTwJWeblAi=cqC)G*)Z$5(0Nr>c9JnE zV*%q(OHY=zq~jPyb;_(+9>aX083S`lT~FQ z!|w_c1LevJb0wYF7)A|zmI|4BOiKy_U6q@2%6V5t5uTYaOL}J-(`m3ta(syndeB`p zh915F#dm$>SD>HuS>XX`k^=Kl{?C^!rO`;t{6wm+&M3#E9;CFszzL#cgfcF>>$Uko z@!zoOwi^2R&q-=7uacB>{p=OAbt?vuq>Tv1QaZegr&s*QvOVIGo$gdkZXSoT_y*;L z|17)d0H^~+AYE-sFP&GyK=NV`GoLNa*2ei&w)PH=z=rTLWB6MihLfe$}Ti zK&&b z27KLkMi=-n@aRI1gD9?|pu+-t%)FTC!A%E$>Wfujo5YQt%{}!EThYO%-;-<`Z+h`} zf7Y|3`z?1LPtkNGw5~q?UX};Zdnsc(u`yS*xEYc;MebWUp-;oQVgL4>A0>y_S!=h0 zCr=gBl>9*{wZePHET!}YLU|mo?J>{e2j97i3>&9QFBW!R-D&wWxr==juo_&{SIUp@ zOfAg|xn7V>zn+|1{_#$s{}N&l&@ZFm=0CK(RX+VpqF=^}u~q74_S`N3WH(;jX7R2( zgfiCacx3|``Z(1vR@w9PVlxl;S4v;5TJXWaqZ?uko~PDFKXxW6DZ=p^(!RMnIGmq! zOw3Wp1?6hK$4q7P49ZK|{@JGcS+1*l-(Dbzo+W5Gj-}IcrH&)bm1#ru(=#u4lf8&* zN7}a@Re8&YGz()93U=>zDh!;()=D}vSEaw2hh1hl=#}p`zR((oh#cl2D0vwl49@Ic z7$(~LWo51TSO+}>Y_m`no)|N$%a1hJgxj(j+4W68ow3I*!|5qF=X-TQBi$=-MT&TPO}-?COt8MQYUqqxqr zOfr#qsMNE?!SjqUHN*H#=jlg0lQD(2&uQ{bWENWf33zdhPRC3qM|BPCxL= z2CMa%Ek&W?suI>z-B{br;h8im-Vu9>&yDI;vC5qiIEull|D5i!(sNtZR4K2Uc|pc$^0v{W?`Qymk>93I{?w7kg6pC1J#>1x~9Z(TKE6n_4JRc*B&(-FP$hM zv;c~4a$df|!Jj5$=0qLCab3Qm$XU!khRQNn)f&tt_vW~WI$8Fj3}rLwHYW6<_YtsJ z4cLej^Fspgu_lIJNy=~(0~;*G{FQngR!Ss%X#Wp$Zvhoo(ZAz$XDo#o0+FBw@6Ud9HJh;>DHwrzZYVvGJ6o}+JwM@9=3Gn=|a=`x> z`*XwSa@#F^XQ-(-!Ozsi#c!X}1qYI*OnXByk%)D>6RS==;l?uj!UpWy+TwP_p{DNX z>7Dl6(5gGkalA~(SRPp1{2(yWww-Yf@MZwt0n%lh8#1(4v~l_rF`A^jhsw&!R-Rb2 z2f^iHxd;L>)GiqiuYGMR?CXiKl&zx$?M&wkXKrhpiOcn0`}RtFZIb0|1QxA{&(`z% zmnk!vGI^L1Wv{aERLgvw2@NB-Gdri`kzejupK%#Axt=^nGl_O}cEl1fDTbabdE$tt zPk9Rxu^4(_EoVMQC~xC3!TMCj zS=%R&FiHPArSO$=@^hDK%)J949Jt!1lIk-~mV(y|e&CIVVk=U*(u0fi0%NBlXU^=c z$T(F^P;MM`T+cq0vyv9o2ak!wFcT@6=cmU`gP4ZefRev&Km9&jZq0~&-4#|?bUtvb z-Y=UfOi=dKp?Uu*;9o!buSp==_b79$FXt|Mtb+w++U|q~kdWRqWsW&j{*rs~Yw0oe zdqmDd95mzG1fSQbPL?r-Q?HAp_P%QSMQr`s+8bE137y6#R?(6+_7+!D!a ziMiL5KM)6b+O{BkQOs#U2(+JV5hYB;^-LQze|g(Ts*-)qp5+}PK)GNX$dX=m&df5e z=jU3lZuKU+s}}y+|3zXQ5f!${E@~#Ed|(%$#q$xT(hKy zt4Y(`7r>eeJY^$RdXn?0lE_M6n?epbGh}C4<$+1UTb2}E+IIGU_C;99g`jrn%nT#e z26NoOhoCHC8GU9%*l>O3!kwTZhI1|Ofteq7$9zw>&4ESkltG1tN!6yoSxngmD_m6a zVBHTwd>x!$^%!gF;5C{Wa(QC7v$Tz%nRU5rB9J~7!6rbSF8Svah8u36)*V0-Lu|tr2`B{8KMBKyPln7lc>?ay-{Yz?PD0pY-5jBn%ZL4B zWBvl9H)61r$Ay*9L6^_D4e`U+&Cq+g(g(QM+!-4k{T8?L$jhiGb({0fJ7VTFK#Rad z?co#5cCnBTimnD|O91ka5U&Mk`>HjtiuE5Z4d_L(Wg{wx49vK#5dELugGfQdsuaB!ETWj!2- z(-2s~uRR0@UsWT)>agDMBd-TDa40P9wMpM>7-9P?z*ZUB;*9OfeH544F3Kzvt@yd0 zC0Ww19gG)uF!d`o0>77(ELYZhNzu%tTr@>ls8BQ!`XL;r3{xlqiS=^}CU%1vRu9K& zXin&zL&dXy9ubAEcCwO<6VHfo&ynll*2t0jr7n-){&u)dc6_dbOUG+~irLR`oL5>^ zk@et`!?^R*F+IWq9~G&Bqe{gWC`cP@)VThI3@ANkeWdencEA|wkz_P6KqhU1X5fH> zZB1gps=Wd1o93)RO(|Ta7CB83w8X|I5VAxsN$yALXs0{&MRoC6l-k-AWsN}SRE;`A zo?wbP!_4!Ikuwk_*|>#U0vE~}0hrHSIYk)p($FXERoJe)oqXkaE<;pd#hpu4`pURa zwsTSBD%(5kG)1G^&-lkaZaGbIX%xb+h*b_Et}xLEyEMt23%FQ-Zg2H1 z`$nHmm}UB?uoYumRZ~C*Um!v6b1g4c&FH;x3BY?e8}bTS2GYh|=EA_Autq)mk$u)c zt(J^~&t~beM2KPt5z}nT(X_$jhK51Y19Z1yxFWT;R`KKINtMNXrdFY&92K6Hxp8f1 zyPV-LaqR(dHt4|T5t#}9D*Mjuy73py4vt9p$);x;ASqypT!s050oMUv>#u#w~d>w@}MWSVm z>HaOowd6kivTAw(zjK>l5sCBE4?8z%ffVDmsiQ0X^js`=$*|2^e6ViX0{Sy$Ahyf* z&Nd7oW6lpv^{Ja;t?wW-of^-()XJoe*wpwWr}R7~e4Lby#Bj9Ge2r~mly<9)YtV4m z-o!9N>z{`J_z;X-}MXSeJ+GR=yKN9<_7Ezt!GYD3S8`m0ZhsIigkf#D4 zUU9&pd$9DI1l(xp=8DKfK(~F-<}`fiLSPQ6Y2B*;497Yo<%yh>&eO&t8Dil*fMeNo zP&xRAqMDj48zHwPpM6K6K`{L7Y%|gew{JSz71}w=<&7T+5ykahHmH9j#*T`F#ngKl zp#BjJJ?dERrM^aOr12boyz#wy{Gs7L2b{kc&%KN{hEtCx($o0Q0zuFe0}bb3qq$@m zww1i%u-hzkY@lLF@sPj@b%S2QscW;OGii4OX-eDU1&0!#UcVi@~dbi-| zAJO=E(8TZFiA8gxQRvKxRo(b&HKN~1g{^>W&xTQiIn2Z@MNMgBgZ`oVz}y5*7DVQZ zG!(9>X#|*91{ZYvrom?febfAqT^76e-qnwNrEr5uEmQ_FzG#rH2D;==*p&p*b-WYh{$JQ{Zt@nytt>YH8(cxBx6p zo$|w0u5K+qh(24MvKGuQRcTkw>wJ z&K5B;M~wU%PGP0S^azEKqQScT2MVQCe;`Af-}D7^AJTR5S`_mL*$RYLj^pPgDDGjY z=AQC(VQMKL=hvt+1;i&3!~1mwq`XZl*^%3XJ{gL>tdHz*#jt26Pq9=ki%~DUVv^Zu zAEJHipR*1M7T;d4oP;K$z%VB2IQt&Wm#iz;C*7$$#fOy~?Ybrsc@Zn^y1QNn2<009 z6OjxtYBaD*WJ;;(Jnnd=((g{Q9+d>ve-&nU-8~e$__TRlb$+!PFV|^)ysjI0A|J5z zaJoLaM6q;?X$QaXvXC)vzV6Duj`ocpDH6Ww z98DoT68@DZ#yOh4D?gd+2NJaH9dciDelq0q==T@e?Qc&|>cl(~MA#6BJ(YtWo4;+j}`r8 zN_22GUY~8ntRBt+J1dS|wBZfG7S4?8@jGOUFIz*P0awQLbOXlqOVe#cBYU91+=WF<+HX7nnm`xc;DtyDklPIniGOBgMW5LW z`+1lgeRGKyf|wFrFUAVNS`VX`NP=)3CHWgS>)7y@)%qFL;6t}}dK1vFoz^VxN#GTO zFG*Yu!bms^x6lEshieALcFXgd0F8Gr zN{zMWyW6i3^3UM{h?kp;5xlY^7wC(GS?1|gEne;8_UO{eyM(5p6sp*mo|<3hNjx0Q z(g7}cOcge5+xbT9hTo`q!jk55fNs%nsrc? zUgn2B!I}Ki@Y8BWyXK)cx z=D^Cid_Ti;Yun~C0azyxCBY;LbK<*UqhA)nUbk>_Mf_zKfZIw2<1{z5I{M5Qx!8;1 zx*5IOuO;;!#d{Z1-@UDcBnR28-i9<0z6;B=rNRN;)@1rjqxIE%zdi7f8A>J&qe*}rM6 zM5Zhpwf5Fg6#g)5)VcXy$Hav~+#+s&u0^x4IgDJMt=l{8X3hB2EI^y`I&|ch3nZFT zJfctuL_>Sucp}rk< z(ZY)Mm@h@RLqe#udV4kmKh6LTLx4iGfYe;Obn<;JD&UglkoA4s_G^?q?_0%R(zO#k z9xX}p(iVA`O4O&*Z~Efj9I6I>Wq5a18k$-Sm+eT^*seiW9y3*~Zq|Icv{SCj>q9q) zM!i0BHv>aWMz2}LIM8lVsBPwF6$E`%>9+hf(BN+8Lnzrbnqj0^=zY&_iBn6Ok1K@I z%l;nFh%R?ycs%CFe|O`bl{P-+(mE_{i$kf|SbSu)&tG z^0H>{bvB2%G4FIg7sCenxe!y>xB=Nmfp?Mt=JCT(fK)u_vrri)TOHp zaa66XClw^bb6M*|HC5~MIpdjMtM{`kWza29`xsfkG)nS)7n^O&A_~|NR-KS;{1)aj z&oe%gDlDOcHha2fV)xQw%gxrDHou{VKwFFVhk9?f%d10_!M8*@p1&s`lt+eQibop& zt#0?{6LrvqttXu-s)Y|h2{y|}XLx$t;ursT!D-0ks~kU0C)k5^n2}}ja;@0uR2zZo z?-6jebG}MV(z@e?>_pCDohD`W(3AD>Z`!NEq2-Amxzf_Sq`lNk&v=YoHMi~np>|cX z95ziT$2s4zlLR!C%+PLz-kMzR=2H+;7jurB18tVITL^(oO$49Y8vI5iu8q@-Br6r4 zQipq}ev@8iYj9=DZ_+!X@E=KU?;p~u>BS&<{clL`c8HG-zOR zKiL}LP%;ScJHAGY9DuhHz&>PqI_(vADSUcvOat;3E&kL?sgo!-iIr`l&1kMRZYWCF zZ~UIj)TA3I))vBADxy;)kme+0b6y$)c3>{wrU^4pt{NZw<~eBZ<7>2lR59G;Ainp! zBqAVctmm=#7yGKC&~O}7eI`S+n>a7Gh`;^DflG*GX35(PO?^N@&g(lZxXM1&`lzaT zi$e2kd!V9r=n}}=YV+}NAIutt zvZl6%&c&H`2i*pwou9uoRDIMm&PM;-Re;QO_Qt6Vx2Fo!yevYy7!Yj7HPYG6e7(KP zC>*u{CKHy%eba8cd9xqAA-gt)bu=^4{@sjxoruvnH zjP2$54s74Nw_0#JS5nuCbLIi=qR7QvmB6>$9zk}tDcpEh?Hs&9baJj-6QE=)YOK2g zwZbczte^2e;i*F~J}GjW8L`W5LN{?W6;kef6d({8-LB?V$?}xlQ{X4sF;d{qH2kXh zW55X_i_v|$r^Vrbms7Q*59tYmVH%oVh=uK$aN{`wePsJNHy96h2A`W4y{iZVmJiKK zg5@3Aj_W~VcTrfqLSvuk8j7*hR|+ji+VfJxY@^LG$*lO0Mb&O!xcHuMl6H4qf%ZTr zUXJcS=Z!pa6P)8$prOw!-;EdPv*oMD2ho-MPl4qtb^pzC=%4Y$@5zw=K>_AH_KRI_ zN5i%fGbdqad||MZU(s~0dgA9!z5SN>68KvFH;w&>ma4<;Z9rS)`0$h7#pa}=m<%5} z(-+(w{kR~qVGG?|%w7c2j#vdmAce28ta)!6+VFY$2-P;VFtgDAi^WF;#Fs8F+~Fr{ z(tMJcTEC=pcQ00Q`?PUq@{^H-0tRLX;Uq;J90+Of{DZjUd-OBxzevvJ7n{%K6Akl& zgl6-VKk3d=LuUe!?k_0Qt@F-N{QAg<$m<5SEHz)BzXF*2W{>YRK{d(Q7D(~eA#cou z2Fw1?M}XcwOt83s!1D9@w7}*`>WgJM{+ya&LK9}q64gGls*L<4FRxln?VQIeHsUr9 zx2C2>H;vEB4C5gdJOqPfE4i4Bv(}XZGoiQVt{5ZE8s$p8IOju06LJzjSyh|#9k=D? zCU4pAA>zl@gtDA z*<=y;J$>WEop3nJ_92MIvkBg+r!sgkJ`crym)J{5Tvco8+&)>-u_1fbhW;Re?|KB1 zxw}Zhr0h3Gq~bQ7YEJUwNG!M4{ZiGS?s-J6s|WG$o@TjHX$nuO8w`2gSAd<*f@#EQ zOCT`pX-7wN0koos^?;r(M~*p>t~`ujc(uGV9W)yua#5pPcAzesMOpPrO@;z0!vLkY ztbC4}Igx?-tNR8u26Ln%FT!S$=~;WuZ$muo5d6dyi9&>pAjbK~Tme1shBheiymqx% z*1ter*|6{R zY{Sid)4*ld;bo}5;zq)*>-mR?VC{1EoR5K}Yh2&HvEmtY#<;YYxC(wx526X2)&7p? z+*_d~N=}p2j2tirPvJklMR|^uhI9x7~D^SA!N<@0`4rB za=0Gr{&LNI7y>zyyq_7?PiYvB_?O=10k0W{)AJToqr#JZ;W;k^` zlTbuMtuUU_!c-#j&MGH7o;+@BkE{*-`X+ zj-h3OdjT|)u&4ozFS5q@6blgFBRIVj&Sd=Zk4Sl(kC=O}Fx_*IpejO{j2}s5@}shp z#7AgzwlT-;C21O1t|12aMHPZ(v|;$x-%RZsyM8K^x+@oxWF-Ov4bBq5eTKADk)?@1 zp{`9k%owrKvPl)`N7SxH#4L4}Ps&ZYU?&{miX&f|Y3+YRV(v2r8!QDFL^frryXZp` zX!4M-lmfxKf?dj)2FdboMQ#$sjMWRsL-2?5Rk5`xHDZf!bIk(_D8e!!%z2U>>VUU~ z#T0|`l=_5m{t5+=jw*_*!nw;|3ML6qbFE|k^x}KK<(q#*{^{ax<1toH6`yt@wOn z&9-RF)IAF<{ zj;%40x~Rs%R=xDKq@}ZYg9a(kKp%ZAW$A3~puxg8m@ii2AhX(lmnHO2obF28q(ZC` z%uuxuKW%}Llt&wFl8pYhQ)u;Bhdcq4j1GcE|5u&D!6zxrYvW!v86;3%A?XAgdrndw zixn0q4=gPulAK94AscOy*ITWagGb9^g$bepOCyRzd{wNx-Rx7sYrK5R8DXNJLAaQI zdVbNun1@^&j7~Yh)!3AvO@5it$nj3A7WZhdn4EKO*fu9Af}d<07L1;Goj=&J%XX$M za0M2mKTnd?_Io3!#$;z1Sn=#8^B-5Zh$|__Sc4ir7hF zJDai3dy*43Dvv7ub1g)<&tJJKP`M{qc?7K79TsaQ*3Uv=tN|gAmJq{`V$_mB516D- z5SBqVhB~q`=qi&G(fL#%6zmcdg%Xu{Q~>~3tjV0rsAM)wg_!Om(mK#PqB)4TQcNJg zd^{?j3P}R$dg6t0f_p^YLb2=+}pbH6JgIITtVP zY$N4pX&{v#ob0P?Y)BSe_aRz*K2e;)N}A#^Jfu=PNz| zDLRm$IJ<(D>VXlG(-M+X$2o>plA$~bK{%9wmMlZjbWMd%J`~RgJ*^7+mxak}Y{;Uk zIiYn<5m&vY+RMU@V5;?TuDjYQUXHbFsEyV6vu>T{;n zA=sfPoWDfHpoJ|Jj7mZa1M?O&b1F6%uNRe|C)A059>=$nkouYCv{hgs*g#1q7BC$| zR|)jo**h5%uaoYx?;Q^>qDM-oWBEMZVJorUA-3NbZ6?@|`>UG_u~V->7e2P|qQ8KK zyp8K*f-3?{BeR;~5aguVby2PsOD1vNFmNHRJxYqFK*T|YJl0ilze_*RnPqlV`NQD8 z;W*GOeLFyJ{k>e3w5we%rW3x55_Fj@C*l_L(av6Nd0;JF2t8PcIa772_=7F-c;#$cwM%WFeY?!&{;(^keN-75gW@lA71_~ zGl;VI7Z+b9%3Ng>KKx4dHDw7)ybmhYG6gcs$K>vgtVgdY;EQj{mwf1*WO>4Z;vH_- zTfPXG9wHi?HZH0k9fq14Xo^@ET5yv~rbV63`Pa&D`@aFy9dj zm2>=**^v)g*DL*j_|ydG*I0nVK;5PdH4N@eu!c5laB&fAZM`@>8zwC$8RZ{B9!iS={O!B<5HH9`cUe|Yjj&czN*pHbQA&FSY zAg;);kAII6`rXk~-z{T$V+fRrQaPumi1iRM4!BM+Pi3N7&ui5=HXtfVO|JAwQ%cTD z!PnPcYT8Nm@oSO2EPn@b_Isi=j8kG|*PghwmGx3=3w8hQZVs;@s~J7-zJxrpr^;$o zsUx(aTj#oU7ne$s8{}SA=N0Q(?C+ZlgJK;uVNIET7kOn=J&^u|%NoBZCz((B6VETI zXGfk~$+9bXdl!AScOpq*;rl^fHg!?l=6NcSP{Bck3!5`suNtnmrQ%I)IOwvtdx^@a zx|9?Gt7JiYaN}^~S$r;097@3sVh*myN$wm-vSq^p89GEF=S>o>DV~HGO&g&b`}^eM zzi^$`A2x$SYzr(*kCRzvGM2TUy;~l9pX?XC?d;q(rcFc=A~cB-sC#gBdS=9wk2-@) zYU&K(hQ_{XfT!$=H_6y`YED0#?xUteN32Bdf3DUVtjJ)K+^jYMFMTQK z;i$*7c~(=dtshEuZFV1BMUdy^Cm8P8x#eMgP7kTY`8H22CvPvzZ21NU-{7Z;%_K7G zclf`>ppyRoz*8HA`2>)}tgerjslMT6mN@M*TaX%v zFds$#(J#deJ5Qw#Ssi*nEj{)ye?^`C_Tj9gb^C1lrh~7g(Y_}^4W_l- zJw24oxbl04r=+hfKBfMT0@-A;&l#yYGg)jTBi)x7r}Uu%DFa(;rxfKmZNI-5c}8hmwNI;uAVveqvppi1B*v!*b;=UhKDQpC|Y2l7`8iSu93s1#2?q z>G<;$qMc>SRoCY}@VzSTrjU7JW4~QHgONAhjwYaTT6KC_>%Q=lNcVn=!I{mrr9=$? z*bqZUpjqL@qV++TlN!I~L5?%w0aGH_x5W_=>?LCdsUZ1dLQZj2pL0UpeQr2+pvJuk z(IuaEcIgh^P4|F%)->1BW3g@<--5_cN&Lx%u8$VNW9jL%VeQWJ;IuVzbu}}=ZKTFo zU@;@(Q&Yul^J_*od^@A0hV1{1!@OB{S65gBk1Cg75a_avrMoDhJgxBBGREcM3t$9K ziI7rvkqj}#3~^e^>f}1b&)~$)lUR_2U-XmP2(DWK9rFVl@@J-qiJm-Z9-Ba(6sl^! zG8yXTD!8`GPsDuw4FetHn+oy?n>#t+4PCBtHU;2&>rb&h#&h`O!?)dq$6aL^Dj&R% z_jw<>M>DzhQh2SN?WmI69KOn3L9A|-d7JUZe}s~km$4*It-;(+j}h|f!8JNxm8qu# z_8cx&T-PmfPS_dNiM@qdbTwn+`G%%Ce)k|TJYOU+bqPn(b>pMltoc{ z_A|O_0Y-KZ8q79@^jsddY%( z+ox}EApUdomp~Qlpz}NZ;&ogI2;ak{tD~yqD@RZh&?(cA^lS%n5H)b+t3v^d*bWn2z^<_6rqY( z*~&wxasD5a(Yuu;aHS~qE#v(0iYUUeAK^S19kgw&4rO5OTgc5fYEvWrM;XmUl-$u!;pp*p{3S1CCJOP+#g6^`U z?Dp&FnpoebW{=MQ`wqOY_j|=Whr;jsIUrC>c4|_tqOX#q6DrMrG6mb?FN*y2zuodLy}4V?oAq} zlz*jD|BkD~rCSQ+F8`2+Vk?QQ!kklOMOVNsZ`@=W&`K=;MgJUui|dG|x}6?rb!CE$ zj;{N|IGcR(Ps8IqO)tv)w^E5B)~K?bi!Ik@-@Kn`T#d(34pjVI+%%0>jX>Dj2Mq=L zmHqDv{oC$kmU)*c+6H9yco!;x3=?80n0R<#r|sXzjCaTsGS#Hw0KnN%=B&>Qf;+P3Y3I|uua7N?TA}#oZ&Y}PsEC|%7pxv z9pfwA@kS=0Jdimw8&VRKs|>Nlfr`ahRg5$N=wiVP$IOL90qJmZwf2;t4}(H`r3CY2{m zup6D?Qa=T?p9YgE%y>>) zmZI8N*$N4!?0b+wQEkXlfekaEj0}B)B~l?|w*Y)t;D1_>epwK9OEcyjYy7dl@;zr_ zRK#!85=emn63+?_qyU(SXQc*G6q$*W$qAuL;6ud6ql{#P5T@XuDMF-3Aet;vBrLX^ z%+x?O<*}$hAQA|w5lDdqDglX;DUuX?Bb5nY%P9;FWK$W9%CpCZ@)M!tGd3tc5lV)z zLHP-(O!VDHSR^PvA!QI@%gOZ(td5C>Zt&!9lUh)dbXlZU zP?M#pfz>KwQMFK$QsYszP?Oi}IgHAjh2Kbn5J0_|Wcc<4zJY{*eA1vIv))V~5_fD4 zX%K{M^LsgIi2PEHZ=fUQcvJ)#{);GFz~cNjoDg~IK6iyo$~ABa zyqf-}di(u$6%AxY8}|B~u2smi?*hR^7HzCnrH-;0X+7Eg4;A${Ws3k?73{AAS}=^D zxAi$#tDQ*hV`%*h8ccxd2d(I-6#)H{s`X(VK&ox4^dY|bL~Fd!##EbWFZquxkri;s zt*U-sojqNvipB|pO-o(Q)-puZBXatZY|P(YYAjR)JFwjz29-u>H}0?T!fL)9A(le& z56&Ir{Sxr5ZaIwzWZNNOH`*jJ4wRf+O0Zt==or7HufSHPu9x(<=+X(4I~QcOi^|ua zYOYrV^MvOWQfx1#1lSjBFgMU0Z`qEJT~YL{$&Pz4yO+O_a3o$WKXR7sQ z9v7f}1#Eth^u16}4k_%W$=*ax-{D7T`S`<7(&9UjYfuP$g7j@Hdl`v11mFn|{YxLa zr?u^iy#rTI_v)K74sHS7XIuL_U9m+15c_$-CF32o$L5&e(Qh8*$(*JWFM$jq;d9e#ws z@NG4%)jS9BZ(QnhC%(^RVEgH-_Mg6|Wo7dGlk9V?{I9_@b}h35rteKmf+$iG^wX=6 zQaL&_-V7xMxW&GSKF(7~QlaG1PG8G3R=|hL z+sFNF* zTRh5cIcN8kpZjeeO*8Y#?alM@EO}nqtJWohVlMITo*KL8 zX394oX!s_T8V|F7E@;vh(4MTm;XYKfFS<8l5^dUl8$6MhDOohviJ9YIeI}Q$RJ6J` zht%_0pQayZzR3TY;@Q1^VmjLDi#$+ziK?ypZEp*;ZY7&god(*_rz2N(+H^2Ws*|LW zKombN+F^}fbrO9lZ?d5Sl{*LG|Cne{{U#cO|3owxeiMzou74sLqsvt5WqV)tuSkoL z(igy2A?;ROtts^UC+#k7d$1Df$g){Yjz2h$mu@sBR>tp7!GzVBd0JN%_M_}W|34ra zQ<~3xQ$iyT$NEGZT;sa!?og)%>O)NQHRBFW zgpyO;a&hVsB0%vw5_8^n;QY5zRsTN!*d35!#0mg@Sr`&mY^TdD_j1n;d zGFx8t4?oJE&U%qzc)sxDkLt$&AJM6-MI9J3cwWkI8y!N&f7N3>mP4Cpcmw%|RE@Kd@QNoLhnptR0c@IP8#rwR7jZu>%koF$ zihX+#D$g2Szif5o(p1f()@}i3=hYow0-CxaV!;_C)zmhKt7i zh9|*@O}V*K41-#Vg~(f+S&Ln&)ulQ4Uu}~H(%u}gmXZQIdyj@oD?9VG73O|bHj_%~ z8z(5T4kSbgF_I3gK zl%+mb_MORB7~a&CO$t2IAZ=hgUT62-9p6W8LMK~2&q*dRgZ~_>!N}H*w%L9JM2o*x zMd3KfctEa7*e}DV( zHnoVgzinECLnlKe^-kyDP6KN~EUWdC&AT+i+RK)t+Gm}NuuF-6U(@B;%b1>oe1lxr{F&r^#7Y zV`0NV@5?zPx*4X5W2QV3DLY3In076=MXE7&QsJA;Fn*sUj+E=RsPd9_!YfU+=LHb0N8IcyCs7u%?(L zP~Svtm2|+(2UQK~ra3S`cPB1gwMOk;j^VW+b|pr5T~vB#uGM&`KHN*@&pG0dVvgIe zgEH=E4i-C>Qa7d_;GjTR3DdR_Bw>cKkUV4a=Bo020!CA=zIGPm9gSER1Y#)8uj^DL zIvLJLlKwK60Ftnhc;TywI4N`LFBI>l?w`wxIB;LhAzd~1WzU`QHczD$mk-qPn|?1* zxl-!fZs_8bJ=5o4dp=8QPc*5?KC=53%9$^;`A|ly49{4E1Qi62=S-m!cH@hPTR~r9 z%1G?n8*iVZ7_<<>4l1EDH+-!iSi>Z~{hM90*6lkPY=F|;lsfsu#@1i2=!&-GR} z_a4r7r_hGTjQr9THaBh-iF65dbxCSoBF%-I$B=J-;xbzH+aA2kcbRF^ozZP3ztdOR zybBE-Ah_;4y`RWt+<@L4&o*z?A8zi%A<+M+LuV&e^)_Uw_6hcWp@(6exoVMdA+cBc zQuEyl@@hGSh0LdxgvFQBm_CNpBT=r0!PZ+E>qC&OAwJ!(9sPmV`i9rsH*UOl+z!Ow z=2}pKq%knGHw&Q>osM!_6+56>r9NpAK&`JVpl z@O?kL>AT~;TlQ=qgYkZfSns?I$NM1IqP8>b2reD4Rn7fQ?eXwUI)hPhiV#0zt+>n^ z&+E`#fk}Id7V?HU?&Vs_KHLX~QX37Q)%eJzdGaYjUT+t<#TDxwfUj*=i|^+r;?=l6BEG!%&e2 z{W_u2WBJsl9}5lXBOaY&*4iBN^9`4D6qwat3DEwp9}0@8_cAt@EBF`79Bx6Vzux5Q zitg3k!0s&%)rT#XI+(nDClC%Z6S{Fx5zarHl8I3%U6<$YaApXz6n`kR(E7Xz=T^oj7cfwY~jwz9#O+v|^Cnroi;>$#qO zq=fV9!&MyZm8luA2EU*ldB;Ysj_{owt`FiyY1WDR& zoU~AJipxf^A)q&yq1v8FCqqEM@IujV>ZF0 z#8D zR>&}8mkQ$?t~0Y}&uS2=wP(H?mbc58{8oe?5bHF3q>JXxX0Iy^Zr4H+TnA9Pyx15L zGMej~2Qtb5=k@n0$5~U6Ed#3j>?^Per&}^hU$V{$Z)1aWMhW@074VSHb`xkKR$Tig z)}yfGq;p@vdN23p4UFnG(m&=7Oz1Ygx|329@!2*1#YA+8Bb*DK)D;eVBCEW>iOz2e zesvdGkkuB9e;4}g$+tk+>`_IIXK?&QZRTKq(X>?j8_{T2X|xjeFEU(}^!9$X_}fP< z32-$D1l6y0U!H@Pwm@-T(Q`EtSyfbLbmKma!5{K$I%1Kq~EZ~_e`_j|?X?aoA zlriFEPG^EwD6Rra_}oWH1DQ!5^9>Oa9u{NU-TKN_g1^=mA#)x!Con6J9pwE3h!LNB z^bT+uHV4|sFe${6ZJ#y9I-I&&_oa-qfg95r@p)wb7i(_;R>zjKjRp$@f;$^0A-FpU z65K=ZV8PwpU4j$b-7Uf0-QC^YU2l_Q&Ybzang73Yp9}Qv-LLP%C^ zaFI*a{NLaS@$a$m!Q<(5Pg!(I;^}Qi(@4L><{13M(}n%nI*~}N3wzZjkw~NqThn^p zQW-V#rB)gK-vingn0s*+KDjssB`Q}-AZ>KE}9fkeSES#YDl>=B=_&gbfV;8v2rz5()3!msN6fF?F? zqOn%(^>mUDeOnQM^w@e(9g#R@YwXHO94aKF7GxYsu7_kgja{ zMHLxKy0lE|!kRxb5&Qd&>xWP8+Xi+fW=^ip{aq{)5>t}g*w`ORNDjGt3^%34Qk5KU zeC=8$<)^b=!L&u|*wh{jCl3+7B-=U5ul2=`e(j8>+IF?=@+v_td7!`*oF}?d)Nzrm z*u!$-N*F0PmEoJlE8t=tNbz|eU1y9J(-^F^akL6KA3}Kjg)>vh9p`(PS3*{7L`H(q%V|DWh z1BxYky3%X;EGySQ$$r1`rzWmSiu1T(G#FbbjlUnvDt)R0>d57*Dn{MN2=g!Xe1-*&eVQ!OYd zF6LCEXy9yliO3!6i4-`W=uUSq?T@qWHhD$KTsRSO=4^Azko1%Pv@CYp2$69vPQ2Z$ zEzj1?C8t))8~k0Df7uAsL(^>`)7&#WG{c~CdR6k+!*S#Gaq z;$~~CAY*&5NrEe6XyK&eXm`IaVlD7&ZI1+9ZD_VgiS*^xegd+?T^9plstb$q^$Qm;m(E)9Zs z`fv_P80q_8ah~GuATNC_(;(6ttnVQ+{PuSJyjItYD>-!>TBXx((%M^crvy?V|sd3r7uTa7(Xq*GohgTWZDwhWTjbF#$xxS!Xq)w>}BT0jX~0?SyQ2^ zHdE8BkW0Y*?2WJXfuA#I4~X!*enohCtx%V<4X^A=SkP{1Z7P5WkIo;f-|MGi;oY6e zT&1Fk`DO4II3wK07K6i#HOtTDgICPR{Oz?BLA3{Z$3;s0fes$%HA2jOl%WP%c2b~$ zyYd0=`!$jJA5XJUXX-3UBoww7Pm@u%TLf7e&bQ|zEL(0%Zp~XvMLHLV6p(-7wk%}4 z^=(2SXm*b75#A64%R&q5Fd%@#{k!x}K!#vx9hUL87G%J{|6hfFYUO{h{;AFn%})HU z!atZ#_1!Bbn*y?tRK@~yXY+9i&KB>aQENt)k!nZ&CB{4XH!+?JWK}c3+%j)qlOzE3 zIB3s)tJKDk3FHewvZMKZcQ+pR-q17o*WF#;UYThtw0Zs@+0z%yIKB^mRl(ovg3ddK z*>#nd5)AH+E^bQwISPiN+OBH8pd34A>l!R~gey#SFjz(xBYzg^VF+d{2u{2R*_K$R zY9kzMpW1Jo5~m3-(-l~kjIT>O>j`FQU2XoM7;?V0cP378EBwyjY`S5E@J)j!x&yMw zmV2e(KVmyFAhyeSj_vYxQNf^e7sMejzuT}9e4?Ag6eeo@aY=py6=~?LXXl3OPYI)S zlm{P_b@qwylZlDK+^9*HrC_wi@uq{R3_QuzJ#V~}#E3UEFy{&Kw`<*fN7P$)i(szG zco)?rbM`h7glKvm-1eyZy=p(_we_O)A@1H4(Mu$YN5j7&wTDk*GeKg>AaF6-M-4*F zdXy4p^1>Tm!u?HPiLd{$Su%w}Nngeu#264Z>);W?O>Z@k1yLrV*rl|Iq*;@j>w?Vx zYNDPn&J5;gpmX&7ncBDQ4@9`TPSpOaK->k9WcucaC0YyAli+}e5V!S}`ui1UIsuEo zB3e(K#ynb&av`@SGZn|ma=r4m46)b#_6y2n*R!N9D38)LQ&`uaR2SH=={aE2S3gCA zx%F}t2&lk8vUNi*HPoMjDURo0DiPj^03nW)w6-%3a%cuI!R%TPLv5(>-1;TWo_M^J zG?Jewrw5@sP0xaH^aRbtN2Ir$mJyL&1y0;7hps;#3^bfemj)4+Gae9#>^T=Mk-M0^n>|XN)?$03dWq*G)A0HF(WwWOTy#nh%KFxFpXpV2e2tGc4P3iw zXn*+PBiF6v)4_^71kMU)#afKgB0*bk=QRwM+wvKc!{q?-gwtxbqw6fG+{v5RDz)R0{#X6=-WaZ(zUoRiLax-kF?9X&29ZPfA2=jwT1$hz z*L_TM9`|o_GsbTR?=M^sPuj>kqVa7I$9t7ke9^@EO>_3a$7fky`fN!~4ZqtSpj=jA zq-uM$5#RbjKB5~MLcV&V-rp~sL6tTX&$H>X7-itpf=Drm72%(c$qEZu5R(tHAadW4 z$P0a;{Gf3rW5w4KU$34Y@z#sr`@?M728v3i{x!C2-w*_({CXy&~$l4HJ zj5rLyrNd${?KrY^S{ct&4{It~~LQ0-3yhnZh#LNu@cfPg(WDAmhWE7uHc`T7JoZu%Nem zgyja|opDwr6u_z`U2ra#lY9>DViPx%03*yxA4s zdDg#BF3;V7cSaXkPqRZ}m}ixbL^<;?jUUl*9?YC>8_U3`KYlD)kW>B4 zL>M&d_2(k-->j-bI+fhc?w!`o7CBcZvliWk-uVsK zDo*|Wt>1N|b)I|pYOC`!bjTW^UI(&#+TMpYJ$Wj6@(#RWpei+>P)aaf%DE33pm912 z?MNV)u30fDHBdg=eEq0!wsq~MHK76gKI{~PY&_&WoEdD8v?s`2wrDM5+FK9B;*Fg6 z;5uQyTe_IV(?BtG^;NqY#8W-mbo-g>pUg$qm<8(yF7ar0wsNBGhIigH{^3cpE?dzh zYVc)*e*^gzI%_)a(@$M3*gA*V-M+$0Q*#ajBl`Qr>;mv~Q^AFxS|yu>d&=ES?le=u zdzsbD-G;~SL@k)bM8uEVxqC%cP6sA6x2$ttnrt6FC{{f8B?aM8OgW~sbg92lOTU$# z+DS(ezs`f2{n+w@;Po^84DY0g(6H4U)JVSusG0!7G!e{ zxCh@bn$dicqMm8C9;bq@XFJ`c$WX%L8Yl6!<42qZ8zUF65@ifd=h{*n>}j7mZ=Krr zcLE+ehWztu2BAQf(wJX9TWTm;UyAEyf0(SS2C0dED{K9RUnt_arG9F-j$vwMv$O|4 zG^F7*egmyW)X#gb+^=}jFND_|BRE)WzA*Fp^W^P1z3~+g|K=U$eMZ#_g-&)EP-?n( zSm?f~xe!GNaG4GC5XCoX_!g6o&jM;}sA9gWS3_vk{*mYlu= zA`Eu$oqISYF;HO=PH{3ro)WK9&*M=41*L>m31#pFW$*)A&ux@8x1`B~sLO3@l#H9= z6W+Px&R7NIE4-yEwhEg;>YQWIZc1MKw+-^TQQ7c0doZo~Fxt-k0F(WTO^Z{zifqJc z9$8!5+(qE00_GPaj;Whk@vQvD{_SzOsDyXFlA_^uM`OHjcm}n<=)4&$JS4y3kWLf7 zU;_W7-(kq!EgZ!x5vaOCT}5TF`WNrnX2vw)>K_#_iXcdZeA|y~+D9Y!Noo=Y#qSXx zELV_g<5MQZKgALMi~Jh-sG4}I-YOu6AG1SsN?UemOeuSp z*RZOEgKn^OZrG57H%5}o9Fy6Tl6}gh;<~Kh=CIBZGVUJH0gE+O3F@58wnYo2iU~I| zbK8Asjdv9ErJ$tfq5L^e7}@`UXeZ=loH4Ge9Nl1nuzwvOTv^-g^HaX^>buAU>BB5f zh1&-0YPyNkm*kX5^2f`*S%Cxa2{$G&5&4fZy8kgksD*cZU@S)R2?5RNqq(IR&0!S4 z56Y?y-7Eu9zItAm%=qtyyK?42F zUj?Ak3~&0Uf0Td5Lg%udT%Jx+#+b07;SqjE2JHl|M+`zFt*qlS#&DB1x0$+dgmQqp z_<_Gm|8ZvCW&+y&YFfj>NZR!OR~gBl&>;Oc*x%Jj{)g-jC=4>UtwOh=7hf(Sk+fzD zc$;Ut=YcOfeQ;g}k&r(pG}d~Qk;)XyYCgTQi#*SXWZK8XVj*UmYc2&s4{_EmnhBV^ zjIr3xI5{e>X(nW&*IsO>>Sb{1j3oF(MAyEQs)A(9zVqObmZ2h|C0_;9Aoti)i$)S8 zGu}|7T(k7T^Y0t%em`-@zvVabkZ;HrPt&jDSDr|{N%ks|rDBkX-h3PSBa=4lBc#fK z3>Wl(qs?H&@%~|zXB{t=laxujNRl78hoR9O|0&+-<_bZ*0{_{L&$bB(?#VzkL8QW3 z(gOn$t@xn%cQ2iJ#L_;_C)cm4i#J!EJEw!L)b2F2;ElXiQ*G0hJyv4lvyK*Eg6tiX zPfQA?0qXA$goq|D1JzaL9v&#OI%r)Rq7`4s_h}?bdyfyo)!%=ME1SE@QwpR*Ix#nu1b0UeNK}p;f@|E;;q!eU&(2<-dFVZ7O6_ zKKbn;mD1w;S%)Hir8v^1O^R1aPDTpjR!iL2oq0DC{Rp$x*d}+9KTPUU&7Y>ZmmcZ8hKlCRo)l+rldP+Pi}bQe z%CB*MNRr-RG%5uZA1w-GuhB5f3sk=6Y=+ma_IQmPKJ~JcR6}PKgMOa4(kKw?;YzCM zo$(WQE;-WgD#VjYi71~QD@WbPOz@y7zqM^f87OJNe|ul>pTa*u@}IW9Hx=6c!{V3p zAA#?0uz%|F>j5>>e-ZxiZnx6q2?X&6!ut}wkHNVwu<1ZW&IcmA+#)>PzdrZ2?H#hOGge{BTK=34;)YMb>4(Ks;n z*D#sS+1Rt%!0ViCp__cJ=)J2xz4+p(hh?NEgk^+~?02A1cmq7{ymJ=pwblU)qSiq- zg4TgKur7qE#>0OR>kUMD@$Xzv8Z`=88IG%X@6hJk6wQcN%mOd5CtDRc&A=3wTNTYT zz36`4;ohs@KP0k;4&86~NuLdb5?DJLh0}pE#~8MokH4pgFf!h)*!KXJD><>wkiQ$( zUrl}*M4G&;Y*7DFN1|7I8C6+HEq=#0aN%-*5M?*S>A-e$f${}+4?jKnLj6i@rEmyy z4l1Pst59KRS-Hyn#>6+`V6G;*Yuy2~=KDd_2&HQ+%7h+NY0bM=_4t^i=A=(!49|T?iVAH5_QyAVuU11ZH<`~HGp_I&j<#KdjyN5c zkVo!9{&6DxvtG3OqvHQZ=cnW!ul-X&15EzX`K3Y(;QY(aQ(XAf-g;Fhf0e}IQL!IF zSYtxCLisr}#$npn6gDYWAB#UM0zEhD#-Ze31VevLGJ$LV7{f)vr;>OC-VNw?LoGG2 z&`I1tzFreDO5M}aS zX;At&aIsOamDumwUQD#{SC`R0cX)GH%Vy(<$2ohXao4le$IK+F_*J^{<3z)}0itEu zh@ULJ8d5pkS|D_$dNCrIzl2C^fzfb4$YV3H@;*2#g7=Bl>vsLj%zCx(DVyu&9+VL=aLNji z(WCbW+@GdCmDqWu_w$*t9QWNc_gNj~ZQsWm@}J#IwUU_vOB)ivK&DKrAx`9Lz|E--DaecxFugC*s2AB+J%+TawJS^ljkmb(p9M|u1NU$CU9)4i>bnvEz}&kJ?(tKP25J^?Q)mz|li@!xaz-SvVht(1w(h zgmy~As{a}%cB6<8=a$88^{}&WaWr*-dF1apEi3rY(XkPHw{rckBOP|>5&5qZoD7<+URrDO2^mWWVZS~0mfpfg z=}SgGjx-v?%3qX`E?0if6tDlNrza7GCKsLM(6dl=Qi`-JnFvn|48W9-GYYNm6}9ix z(}IoPIcZ>h+P+b5J&lj2l<>=l3wygSSv5k7KyZ$Fr`lP$0u`T?;879`ZkTSso4tel zfbu(fe<+jN0@(90F(xs{(=q2fPzenaDluxJ*ywLfoHf*v#^B{x# zQwKSagAL~MQ{eEdI%XO^-e)h*?~r0Fe4FSdvL;>6yRL`~BFjRE^{ z@95c)2K*C36%1!J-_^E2zt^VZ*u2iaG^?1H`NGIV4Z6pH_HSq#k2hGHr zQM4c5n9zfXB^u>2$u%$xU}pv5#oP2tEaA4_RgQ|D*-b40SWOE7qRHH5I^soe|LFk zx9ep%$uk17bO$h-17I@7MeGS^&2W7;!Wa1w@&;Qd!GvV`ViM0XE;}4NIy}HP+__Hp zRhckvfpE5Td5na%ekGblW-NccBLC=X;<(GOuU5j9%-X^gzOPRv!-%?Y7+93~JP0l0 zfQ=;_2DEbD*E^G8X`am5xq9%}C`CkCv_`_iumXHgM%|ycL@2UPirvrmTA%;3O6h4+>Fir-9z<)t!fivn`=LL!+DM%|J6@NrH`eBQ#yf&t9h z6q@jHr9eYEU|ZoU-ZL5Yb^@ogU70T$*Qk5e7+&HR3>5&A1;B_@`BDKeVgSr{eR!oi zB|Z}JR^gw%R04o78fY)~rTPg#%;1z}ec`LLH|mbof}cs4k`1IN69+cT;Ada~xwTQZ zd<#zLM47Mgx5==E7-sF)Sui9WIHl-6eT|hS!>C7aN;?2OZ6p0^UHG3kC9u2|AJEJY z;5TCV+6FemX#7>O*_k2%0XmbP@902U8tmrqrw>;8WFxpZhdKyx(a+t?$0yLAmPYkk z8o=!5`+a~&sF;1-pVbeSeKV|V2A9~gTPL0`#M!D_dKlL-Gl4Gzb~7v^nSEWD$uF^? zpD6ITkJ#7AG!fWK`-z0o0b6E2+@8&0;a!y(;d*AK?^^J|z#;Q}K4infjm%7+hjDO# zP7qg_ISQt?_>{`Lqj0*QA~m$f5w z_mXN6fU{)-0V6?b2Yx`F!4 zbSTP#dirqM^vrm~!m8REc_%ISiYksYrAGL*aHFOq&Pt;MubT#fhMFq~EQifU7c^Er zwXn-kvT%8=mw0Uh9+}F)U?s0K(y5s)ljoJd#AgOaeh8(@Bn2}LzxmNV++f0oCX+?w zt}A9|3h#HEIHY=0s*cW6JLxl3)QQc3fqU~j9DaAE!q$=2ebchxRYg#;Ugh%5?#nk} zqF-m)>3z3Ax~8V?VJJ-;7CI2gB&SSq>j!a6F2vG)eroY1I@QC0|4wn#zItJF$CXCLW$Z3PK&?;99egvwXGt35X?=~Tp;o@R z?OyL=ZV(PJDv9bTHT6M8B1DS=U!tC&9zl$SNC^|E+&W^T5Mo=WMk_MNW z2A=)g=WD)pMGnP|#WEIPwRg0MxxVPe3xjO>*a!n>=T`LLl!9`FdS9{&-A6H9U<@H+dIzuo$@e&pZERGg`Exn*DU8lceCsw$A&;fee^tc3Y zily)LCk7-sLCD`av5<{|kZXV~;~?zM!SsCf^t#>jo$WrKfG^<~7UMcd%h$6rmS_Qs z5deC95k$|IPG6}2qcg_ze3HdtM5FI~tpq5HU@^AiX!)vW^@kY;#Q{h%08%f2Bw+p< zX#=E{4Iq&NNOE+Y;sBD&GZLhpmHzVyjb4{RGbpa~msA(1gNWAvx-(d1nK)X}_FDbl zwSvCE5jKRYWYg>Ruy#^G1=EYw((C4Z4N}_0){5rWGA?;B8m;>qiCi;C2|yA{r`P?) z+(`u>y#4O)`e(g4$wvufKixW3xZ|pK_O=F)M}tv zukxUQGlhYJA!PDELVqSq2Pa3O_dxS+=7Hr`A%r5P2m?n333~{`X7Xsi{7iTXG)sG+ z!Po;$Y+>L^aG775!*~EY2A}~Mu+#KF>lP)10(pKl!iDOyesR+bL5_3+2=P78sI(wO zUoXtTzREmqeKc@l^V~v15%$=}7H;|JYraklAhZAo)xe&u0eC84&v?Qf{(nAL>u(&G z=O(f+a1TK24fUCD$OFwf_f#`B2B?(rVfZ=1nDeO%&nYvosT9@v6sAV zi^kkE7UwnXtm>-}_q%oRZ^QSL|%&1iomaoML9h zkBR+Ef8d`ZCjMOt++B-2J1Dv5>D#6C7}+a`=sOXTi;JEen}J2=_n;}a&N3<*GnyE` zu84Y?{v2&Jx8godvvt>_jcI@16OLhx)??F#kivQrs9Sd(RlH#;PvU)@hy=t7Z4B zI!$t-QKjzbaR!rEw&B*2RG@&NG4gb^8aT^-p}<&Wp@NZo{0yhtS&{x?G?R&1#pEi% z0HzyNl@P&JQZ8PPgm8ixCx=go8OM>P zauWo5eO|qPz;~Za7~N#fvjEz|HxzqUX&_wU_G`K1vi>!)n^l^&iSVk)aOb z)s*k*UvmWQZLcuFetObfCeOZ*;|TS``!z7Z#``Kd&?aOw8>F?bRSFFme-CU-?m;{SYp2rY9V6 ze0u4WR4Z{TT7Np^xG`RJ$DNU52xQ_0-cxQ!` zaUn-Q|Y90-L{mx)xY3ZM(EmJ}R??y?Xzxx0`ZIB^C0@^=#&%A!wYE%~@pM zh7%^o_d^njnD6WHOjWjc>w#hG7KU-Wc{g&*<7>>Iz?V^m{T zFs)I5$-j5mJWh_UQsbeHt{vVnQebvQz3NZ$5m;HEXCAlB5Qia00n#mG)rWgICZExz zJRSDX1qS#h`^EHE^!gEjtmXPty7Q~H{v^Sh{gJ+%qE~&uoLR6Ir_z+Sk3Sud;OOdX z8Ev9-8taaq=xp9Z)=oYxz6mgR!o=`Y78r0-aBv`8QzAQ@|KdJ-^_l-EthwoVbt4mCDiWv;W~z3+H2_!94!>-i&a4zJ1rP01LWF>>p;#s*up6eD~Ug8n{qTL z&&Uh10mX2<3W7WY+UY!Exa}P!*}yaI{B^PKT%=Dju~=}}>!p2#VBK2<8ma6!bd(R5 zy6Eh_&wLoxXFg2v|A7yq@P?DG&r@%+PSAl?ke(WFeV&K~5bcIi5f)wKGePAVp8?PT5m;)8Gl=sX5UR;- z23Pg~JJ$zQxI1Lvjepz>wmtP40-0!GG6S4;0$0GC9Wa;I4L=i{^L8hYE2OvG3KsWx z>yymNFqJRcTpzA&jejKz<`(D!TN9_PYM6iMR7{WZ$M~~*{_cDu{~oJ6YR7_M+-+r z1!8aXo3SDgsrG|R(iS-YsBb?w5?P=IK_*c=Kxe7#H=ha8_PyXXE0^(&av%u<7(eIB zV}i=xT5^Ib=>uZ{#t*getUlta7pfa`(`kV4B%i^gT!IU1OmHq2N*r{MiUDj#5IE23 zgX(e(GVtz?0rQ?`bA~{Gk3QOX(!gHRa)QS7ejB#y%Aev)!%* z@Q++Quo|CX;+Y*iR5hV^xK!aG21Za=4FJ9M7p~9^bX~(V*zMOo@^9SDv^+sEh4ODJ z*hP$m!r+CwGs7#UGQ*477~7<9es_`DiRIjI#B&z`FTI1A;pG3Zi%4+9{@F$1I6xP% z7(kA+Q1dI|i7SJ@3$7F;)E?ERwOtm;h|zIRVVWg}xX&D6aXc2uC~9Ah5S)t@0IW7J zSPaVjAl*YwP4uxQk?%jDwQ~+kS!rwtEn{49bI5$1_I|T(1H1b--LX zPCydJ1Qp!%)6dT7BqkL7ukKPfGe&Y0Uq`!(ozR?YDqAr=7?caZ@NC#oWZ2ed~qQG#mhVY$9&H_I)BhljYsF1DH~Bl-)hfblN?W8g6)Q9H22h1%2^P+xYy`aihVBG17TKj`o@KDic zj-C3|@>$@zLB2Mj3@SxK3mEi}64(vRG-WRdQ%Zb2Hrdcv^`4#P+68c$re~+A z{#!Dp9Y|K&1_fXf^wj>vp>6x)w9d=V?Vo$-m}ZN zKD!Jc1YCyk*=5+DU1l@r*<~1?T}JQOWfGrV1~6X-%mJ4f#C~=e{%4nAM6u6k*$PUm zpg^b7TOa=rP{tv*0+Yh{Ou%FRI{^>R!b%XNBoMM>Hwr$ z?sGy!yDj95yLUe}yVC$RJTt(+WB-o~ydL}Fo%n%;0>NhnUcm=4?ge?Ajq0N8F#|eX z0;W1as}rcYQo#nV%d?YSGqlzx!`||Py{7R}DbhWG>uHti6J=mJgxX`zFEf{(EL(PR zUb5q$nS1&6S1HDIb~5h{PNeoGj;qusi^tbCk%wImk|W1^uM9-BJsMu-x@mFwb7Qz} zJ(wg^u&~_;R`@di&srpm*LJed)%)75kq@E%hYr?5|DbaGqpq3mguih1M+f;?>3{q@ z+hqLF;lFzS53|44X6p~4cf=H;7%4P=8|RyDmQcTxqQ1edmA|%zt~AY0^2V&FWyuU7 zprz}?0jB5r`IeYpCHF4|cmjikUT&GeM^J2Mq#csswKBo_2w!X+xzBq;@U(|N1N`VT zcsHM?LW5kveX~YtS$lX~d^VlnZNn4|WWA2QaaD6Qvg6}Ev2YOtB?wbnbaTR}qN0gW zs+4r=A`x+|j$dg`4durg7R4_Ir}0^ghqjhKbj4|J4nXa;yamM~BxZSF3aP2;Fcq{x zUltW3qx^svV=D{X9Oq2ONqmA`KGyp2Wt@t_cy(#}aEWsN&{m_i$(P0EIV+%YrWL z_B?r{>+?($<2rQ`i|Wa^>WKvE1bKA)U-G_Cbnca1=_?48WWN-7o}9qELuEkf%$X0d{z^wYM7nCWvT%D?5j4tgCV%QWVGz;NV^RZpInQdz`Rug+0WIJ`bJ zOg*9Oh#Xr4x+YTFOB6zy(VU>}B2NAqzY`LeaS6%>N>WR=CVo)X9?5ciX z26f2un_pk*zqbCodHxRnP3M>Nm*T&e>IXibGSYwJ{MM;t`}OtzceZ~s1+agi4fD#{ zbx1I?@qb1_p+KkNpb+L6S9>k{YorGP-6;9ga<<8|wAT4;p}VK$$Bq?9#D-TWCTATl z66CoZgEWxzR*h^|jC9*c-a{Bn zmnv0`Q8QN1ix!|Ob}b1Xf5a{<(m`CSZJ>7QZ*>Ds&JQ`nwPE!mPhm?1Zs^dJvKr+5 z^J?6eZrtJhJ}@{VbofVQXM&#_lOpm!4(3@!@auNgHP=XmV7tq{?qa%MqcJSF4JnWA z@4~-0Y%K*2_Hy1;_!rj0TnlP+CWh0-RN45andOWxW;Yg7MgV7iFq zf_qw=OOM3=N-lz~#>^kTPcZM9+(lReD4GH6n=dFX-)l|WF(8^%_GQK~GTe^d+PO1#EE(@JzER`6 z{NfO8U#9I~HH1lGHqF4R*ye)Z>iW%f(AoDhx3Qu@(3=OmD4&*$s}ctMlk^*3|4mf% zF9SToot8W6y%fnNMh??A`(0l?7#u2;9u*I`B~L`)mZ>ov7EGW=K8*}N9>mO*-mGS_ zKbh(vulGFmk+Lr@8oj$P_h=wpy2;{tKrKnS-;>EP4BEVIFXEP%HOpS;`VkwP;5D?c zbkFVZu)fz>H+Qejl858_wA8Zc6H@yCA81n1qI9+Y0+YG7Sq+}GYQ%zlhe(gwsL86Z z2#fIATLQoRoRzK_%hyqDh<{!v+B z_o$#E&pyxX;bN_XTa@JS>cdO)BKR#Pll6qB)NGz#nIissMw2U{A!f(ow~oc=9aBza zL|JXDVBM~$FaC^u;=jOM_s7Lu=l0wO-EVNakK^8N_!Dwj|2F4ucXL#)BVU=Q9qNlH ze0Vs3a2`5re7yXidjC}ac-eDCXM3Pp_pq`0q^A zyZu=xhJ(16y^DFPdJ}glaV_5J4>Z4uRfy4`#Q=LT`#;Be@&6^QhsjbJEoLtAzbKqN?|44atwu!l7Yc;3>L3WhA zjLCQdg)eDgdPh|fzypc;I!Xt(^eRB3|SWsYoZpF*5gWGjUd%an9u!;H1s@V zn(o_e_Yeqt?O23R54=9A||iVx@`+#VH~87ue+PRlsK z<1z2Ins%Z^49Dv`!B6vE@KzzT{p^{?l<6=DOIQNc8qVYabloqa#Fm$%$kK29WeS(6 zYZ_-^wzvA;5+8rbR(W0FT3o;SczL#=`Qw67HTbYje+`#m6Mwi$s<~%&U+*rT*NVIB z2w&x1xW$1U|H$>ii^yI*KkL1O*rY=hNzYSCL)D7qoM4&FQ$Y6CW#$-a9n7(eO|ev+ zirh~1T*!yYh#jsqtBFYd8hfg6SdkgIvccUYSFq0nXd)tSj!RV}E6YwJ`H?|LuTTh& zVL63S9Kfs_&2h9Vq16;}ol)kD$T?t*5@0CrUJ{hT4yr2)DR?qSK7jkav7No2xCig7 zzLHCwUg~|!HF3o_-QOM+&a#{8KwpB`&qj(~l99CqZF>a$o>1z)bbu#Ao7GT*(q zAG$3GEp}uYlvANyv$knUN-KeFa7}>gbxAs6#`k|jdy8E-s*>VgYXv($uE9=ua z6j-MEgH-G_Fn(5zquuR&>=I)Q+ZrK^r}~l9^5eAJYX0ew*yR^H^)~8<#>f@uMnV|Z z3-0eOjz=Wogx8eoi)PidwS^9j`!|k9tdoX#X)`B2EEQD*JIw|Ul{Xh@7NNdTTm++} z0a0R6tRBlH-`w6;T8_2hH$JXPl>oc>ul~N!3^H{NM5+zMU-sjK^y^r?)s^Tf8p|VK zGbzx4vfBtwJN;n-PK~JZVQ#%vrnEZbJjA21X*h8VtKu~gmPf56Q zW_oYYekRTW8(In&Fx#{COLQ5S1c%gZ9C&OvrBADCZA2`ZxOK{{J@181j2A2EzN#m4 zjSw>nUz3heYwaSH(^~wH@=0nTph`(7>st|7-t2+Aj_e&BIv%{(Lcn)JschmhvxK@l zYtheWf;H$O7s>JhxqEH$Z)-G#VvicwV_4xViddEe1p5S1(|Y&vUBhzW7&6xA=lp}< zAXZAcuC91mG{EOqzP2D&KM={34j1|B-|ond_ionqAHOcnF148*PzcvGQtyAtE$$lH zEm%LRNFy~mU0x{0G5g@WH+;a#x|_m^b7gxm?E}3Dab7i_jCW?{H2A~USFCWMs>^Dq z@k612+!Eud!L;>sZ%rjIpy;NK_-Lf+sO&@6X5R#`G(E3&-C) zmLilJa$pVdlojmsV!+}X_c-grW3&59Bi2+4s?zNQjM2Rlw36b?`d5%Gm-kf|Cl{*k zAP>Jub#N#0AU)&%+dl%&U~?Dt2e0JTkgUp&+;J_(^ee@iKhS{*WLEMed*1ZwqRL2k zZ+9DUpec;o+IASaB7%-ID~{4Nw4H@TS|+NA<*L|7b0>RUyX2At&$Px$ExTwERO5GS zWmZ;o0-m~1y|Pf-;Lsn~YB`A%AS3&ODRG!fX4poD*)=$gkcA5zhVf%$nKMmf7iz{N z+seH@93cP9SD^M-o~()%4J@_x{w$}?=F*WrRMe>1K`yEOW3xA#rsk|wRe&Jlx@>kK zZ8wB7;T?pdabJ=|cJzJS78r+(5q@evL$yXrT!omZ1!3^*Pv*VH{OC6L>cz|W z&pVzoj8SrX@0b$m;`4PlfX9a2Du`NgzH3kF+nXf=W~Qac9MX|0-aMt?CTkRm=raoE z&FC{)df=NNIb&w$Z0-NKaN^;%mX{IPZb2HHJe5Rnf=f&5E}3)#!7v9=LMhAj?Vtkf zfUE+KtS>vlsB1s^RTWpn=tU77%*-R&EJe+OqL9gA;i%_8xv%!xLsU)ZljAKL{HegY z8s&_nY5o4;i2R9}i3uv}*3-ORRIJCP9aa3rG0} zCt7Xg=Pe!rtqJ>2#g4X*d!l!d1#4{;x-Kc>eKtztDMNcwoj9&ZSAt6UX87kG#WyIC zujw=-j|(gs78*B~hUQil3KyAh|J^B=un>2S=Cj*y!-(<06tBT_?#@WAMxUaswZ^3NJV(;4 z$JnLvyb@LAW4w`qKv_k&M|604vK?`hYPv6Ui-W0oTVa8o(HN+W0d5R5WVLY8 zZa(@=!{^|ESk@Uy2gibIjBzY%r+u!H99;sn!&a#~m~SYK?dPL0hM2UE1)O6X-;h&nB`qmJ$#hRCfX4aoMLJMgamRL9h0 zg5P{xn2ijrXyQK~)PpmU9QrIHW3!gYDG;BWfX0##KN$Z;H{S6v?oD*u8=5$@&R8b3 zSoFKTf=RhjL)#h^YvGx%CN($El>C|eqUZ&qN(=epbd9lc26Ij=$wqp6)&_*NuVlDu zB0j`~Zy$ts%|nG&9ExXCj=Yv^6P~}eX2RM3xB>Ce&gTPx_{onEd5tOMI>)p7tv%M! zTE?&seXToX=C4#a=S);ScB}z&Ae86Tcn3HS1+KEzcw?}$)OueYHIKlkev(A4MewrJwI6zI^ z*(8(O#Rw=0hNZv=aNV;8p2+$0QzSSiT*sukl}XyPqni$NM?-rO9-7W>#HA0W56JIc zq1vL`&4?8tBMT&yEG=~RF;1|>`x9xwOml~HRpvYKWn5-vbkypX=*+0qMZ<&nM(8d< zt3ySI@G4bdr&S512h5`j!J^qlmma&{t;!H(F@2H{WoItQF>|N6RAo(4COHyRLB=Wk zFLR7ja&d6v=5mHyV|_;aMyN2Zx+g5z5#d-%~i5!T+Z zi{MjPmyGwVoMtuk% z3IRkLppAtXV7(05$h-NasV26ROC%#Ba}s=->zwf(_#Y7)rG;ND z!Kb&ta@fV%D{~%vdhU`jy_jJ2Q--inekWJ0!Qt`j0(7(BHDZaJTTpshvrEi?d=&Li zL_xg9HYk|Fb&WTBmhYLYDlbs9pQKY+aw!*;6dt`Vf7WWg({(Sk61cHhJ}D`_29Yea z*_oF&{drbzavjU2nKi$!@%w%x#DoacS}ED})|%5ps&+s}Ilfit)TurJ8A*aE!kx=Z z)WMAT$r|+Zb!E0jgH$qxS<0bFlwsl1=@G&_?ZxNFCLQx-s_3AuY4Hpy+XT@Y^BTM2 z-ijf7S;PnpP=!P)PS-}lB5Oqg97*+p`!8RrJU~L|dkEvshpQcsx+Axacq>!4VJi{JwKQ#h(5D5ck$$aV_t<;NCGJNJ4NA z8VK%Af(C8eC1_)fyEHo?KyZgZBf;IJae`~(P6NTUn?MI?bUMl2zk7b?o;x#V?#wfP ze6^}p)vDD`KecMTZ@u+>%QUq;|CN0|Q*>!fKL#`UDVIl6*J7U8DgmN>VZL z{!Kb0NnSn&Nt^M2Iqw~edDz5~Yspz~kJ0K!;^>HiC80bWL~!6`34Sc49;DP`McVM$ zKKZ>(9p)62;q4Xju%aUK(~|6>$TNA_+mw%boTp6G(d*7dX3uz@12=!dWGvw>XR?U6=g$8HYM znLhBcBFlfa&uGcZ`zTq*G2_$ZxwsGdO72MgZTLmp<#s68vnrs?#vq8uG*`_ZAFUks zs<+jJ65o-Tz@|i9`E!dnKtayZ;+pekp2oA`+3AdXd^S1fz&!;HTm`I)iRa+ga7z47XZfvZY`!8d2d-2b z=OZLV#M0h(eBwSI^L$|ANSf>NW%sWJk5+kN<&fq2(s|8GI5<_>9TYww60rcg>2hHJ zjU37j(WI0*b_7>X(DuhU#X%*Zd4KA)8E7_t&q&zHH4soJxJ~S8E`Mn}4+KFj$}@M*B8e%qdz zNzhnG@p1Htv)tgosFnQpC*(onE{Q9WX(upEg!aT2(X+lc%Mt)4RP+S~*d9C{ipY9H z*T}Y6u=}IyAW5>;O*JG*IY*b>ig#uf9x-_tUSxg$$H@-WAen_c^+43lq*x5)(8nF; z(hQlR<&@!XxC3=9+RHQTvlT15;Z_}HdMsti8dH+)1{3ctUW^e>r<%ppXj*lpeV!{9 zVA6iLp9m5fED)9M`+>A4djRZBnOnC%C|Hw(sKkTtAcR1PVCUC-`fRd`u2i`;Tp^|! zz-O{+TV%?2j*pon%~9`UbZooqS_l-h!~+WJFnym}iSU#n^1fz(PNCR}jgIb>e~B59 z1f8BvaEj`0A~sj?_Z`ad;{+dB_JfD zak{`mU*EcWg6@L6eG9Kb46rUlFN!XvPv0X`H(`7$Cmygs-_{Z|<7QdrnGOnV4BO6a zT1<7E7<6C}r%OaMuc3@ziC1X)HSyT+Y55Jka=*oyq$@qp4oF;DC^;Z1Le^J{)K>q9 z{$r+NaAW|rQ&4Ttdg5rGgML4^Fzw->F>@-~l~L@Q)mnDQUv*sHix=uHptYPojc6Wm zBl*gkHZ_L))LN3`k(*gOT!nUn#4p3o=*icR}?P}gNl+lu8R9_G!ULIi%yn@9vy zKt`Fhd`T`gT7UR@-AZM)dK`J1NM*up+wcd)FpF)bMkeY}w#PRi)YHq z0LM&&Yl($pW8ce`g+(Si!v>JVSAH4}O+d1AR%&4&PT2ksVy?gjVT?B#&AVhVVWxS&zfY}uLl8|{RcIvkVFNQnLa)| zQT)$-B8xqPs1~Z#n8NuLy7!G|Z{JN&m4nK*YpWsp0E2}L>PN%I7+a=#q;&Zjy6vKK zq+FwU#O!!-?S2Jo1-IM8@13e1A;2tc{&P* z{1L23ex3(9L`!w!CXmWmdfkL#F}J`AQo8C3Jn#KFp~AWcZiBt=y#&_H3^q9)hDep& zSOu1~4!jH3_B5CV>CY=`ZRb{td7{gVx8JknKpDS{#f+Z$>C7`9?p{Z3>G}o=5g}ge zWmVUJSwxyc2AB5*qk1mFyerR>!!7DBAlr4Pj6Q5C#fwrs+t->`zMr;hebNu=yiDE? z&tx$fdbD1qEAG;AHb7e+W+Z~Q&5V)8A_+!cZZbZUwHCY^Vw}zkcwB#xg04RB-9esR z5ZIrb5gbr_I$lw^GA7&XTLMD*%srv$&O=|YWBZ4uChxBv7(K-;LY{u8w6@&kcvx_Y zE(cLv8$MNWB<{{KkvWFB9oDoSWe&k)zG5$ge`%+d(D!ZkuNB3N8o?Js`<64a!dK05 z;iA!C_7BV86E7Nt{yK&AjE|!|*Y1CvUP$og+2Wa}U;h9f+)#)5A};6rmC?djjDN5z zCZ;(hLEBYe#+kiYnux5pv&kIii;9ibQ()Q+i;!LEMjkb?T5(We#HVDVG+ANfqUiUb zVnghvwx{T-GN>{R)r(Z{=AHafy3qnv81XIIXvzG2xK$XjOH9*jDNu0W*LE>3-XOAM z5b!P8u;NIB9~8tvS~Xp2ZWSD6ScG=fc8lT+loPM?g!+SP0S+hH17*f6jne=Is87*` z_H81(Sj%OiX-F(}}BkJmF5YP=W25-DDq9|<-4?J;mt^~8=&W!eS zAXnFRoH_dtpp>|l{vlcqc{N@%gCqTPK22W0(LDUE9JtA_FrY>R{tpLP3Y$r``H?JCWdk-|AQ%k`^!pil+^ zCnQ*krj-Z&k~=GgzIyhhro<(6;I#UpKm#lPai4P>LWS@BJMh8Fx|fP_FA3r7T_0=ggSqqJKAliFLrZ&<5<#A z9N||xEOQJ|I#M7&`deVlTd*DRA``e3mX9gr_a*~nA0i*y|b3mXr z#Df_auO~K|(s42Dny%~*O;N*qkw)cCX_^Z^sSeLm2wdML($D!x(*!Mpr9$45{2w`~ zi=nCeE=^q1+or^6brz;({X1TjbYRAt4y*DTt>)RE7yJYC_6Zf*U& z4gk37`O)QaBJSgl_H#q8tsxeftejmogST}68_|D`Y~uYH#^xB{b9#E{7>>t^Y7yK* zV{sU_l3#lr9bd=`wV#a->F5x1WSF%0y97vZob%$YYo-lZY}0gL9H_jmyLwK-<$~7$ z7UIgZXh@IG?>KyQ^b}?LkfwgYs2cAB zKwz%l!KU1L--r5aTl@g*>~VDhuI9;`?D)pae2COx(NcHA)%#x7;fk%Z_IRVGR-&V~ zg5Ck$FdIT@-V;4(I)NqsLEA<5OX;AfwgsE?p|vM|;%&2dOFV=e%4MCvij`=*q*%l@ zClVSsIJgoMu44)Gs-Ff0pHu`VT87bsEpMtCnE{J* z+b9?c@_F%?G`|ENaD}$6HiY%C(Ik6YaEs-&ER#Ou7KE0ov?SvAdeG=ps38g~dMi;XRbOT#<~Puw>CU4^qb5AjqXn zl1{=nZJt`fnU;D7>qu@o=3{l>)=k1q1w^(rcMJTw8Bo(X-(7)Zu*=iIEa;c{Dm03 zlW6$PHfpyru2`SkQSj_yUE04}6gnj;#uRTkKjymb*5LkQqcLY&wP!qDK*Z-1KD#2} zZrH;;nakP8E01^ytESB3;!@I3%d?WxJN{g&s2zfnkRI?!RFC)OBw&M$m=Pd zGNanXRFi8?E-f7KHsty4XK7)g*eK=Pb?e=o%bZl;kYEP!Y^2vS*&46!*6&1#^w+2D ziRhynOuYkP5;?D{U~@=w9`QmXPrm^2)D&KF81D#2m?v4(Ut$kSZl8t>ZQnhe1qSje z-U;CPYyrJ1A%E3eTMKR2DqC8HdpuZgW))1OU$KMGMzdA>sb@as^Ez~pBP1x-pdxbtSZ}P3$let zKsm1nVWGhnBXXXIj{GvMx^cdT1W@rN3$xtKoT*~{UHrG&t!|I#Us*NX3*!?o!hyx5 zmwP`4DoD!5JeaqBW+C9bJZZ?TX-O904U9Lv@f`K=&=nZ4obZ_WaGzBjoX(ucP69GF zj-M%=$>l5I+wZM}3dB3P@uVaznQ}Z?85eWaJ$hJVhHFP95Q5n4a7&DW@t|UzN2thw z+07o$nN|b#BPfuMqByqnVx-sjZ~6PV$L}UJ-Te@U>?P!YvR6te%v0-xJF9BGm{>#N z1jezn8%28deUQ|n_^@x?6wy^ka>etIMY@v0`@p_sy!Eq+%p`;;-G}>Rd;i%~%h+@7 zdEz%evPgWm4Y|jAWRe(Gy%O*DaPb!<6l4da`X(sOb7sfkDE4WK;&7<5m^K0X0@tvp zm?FV0^I|6D*=}`)VSQ{p+x3%r(^SF(Z5kfX;rl2%7Ckl|mQ^E^Le3dFFW_<^(+|5E zn*HZ1hYw++aNoVnBeSCpG?T}ss6p70OtU?YoiJCh#`quPlX)GIM~n%Ml2uY9YVe?0 zJ|@b6Ayfa1P_k`xtnG_!&wv2`YiU12C5eI44gXr?8N9SNd+~T{dK!>?n9+|e;{h|a zp)X$octm+dq)jt!&yg3fGvef)2RT#8=sQ@-ZQk^FoNRURCy7c|g!@;tKFy5|KN@rMNg z@UFzr)c}qRaaqD$BaQ<4i+Q}wnjh9gOmy=SpF{SUv^ujqgNJ2Y^K>42Jz zY5p@>#hMq#>c%!>_c1uEYq#+83s_0umH?$m?>OsKrNZ&z$moq9p>>jz;m6}R?A408 zhn*2aE}DUR>#gI=QglfkOiYJN-i<`WWC{}HM84a6(_#gZP9jXP;W%znr=6`RH73iV zK~>)`uEIJtwS(&(e7*Zjn_cCuInSAo0TPFetGQ&h!;K*1(*{h zr1#JU!ZfZeoxizn{_wR6Qki6p$~&@rC2>6Nt-yp%1Z36PP|0ohw|Uo-C|_+_smg3~ zd9KR6QtP)*Bv^_dHmz-L}U`!}w18PwiaPL|=T0oxq@e5DNB~}mT4gL0mQfBR(fg(-lv(4NO8=JW^C9=uS z2ej#iLoHsZZAll`y>-P|G9sqAWgfwEPQ{s_%xoNaX=4+E&|ddQzZZT!nLl9sOWBoe z(RO*I;zqC|-dFGj-+?N2kgIsm2SOdFpV6B&Xx#=cZ@xHQ!J9Jp!INtKE1FPIjJx8UOH=~yi90dQ6y3@BxLj#w;?K^(ihHnvio#0v~(vpwU{(E z2wx}29V@fgrxw$}NnlpD6&uS8eR$ZCC4fqyC0-kkf&Z01eU1CeqiYIYbbhAlTCeEK zJK1DGMVO?=Mk4y$)lDk!CNCAZ?2Cf47|**0i&@Q^ zqlORbytl6TM_bOeQqIOP7g%R_Q^R=XcO;zuijbio0L5a-7<_y)u5GSbTf`>ECpPw! zKx|5r?o()ep?Nmj#AoYpcE$zTh)^PZXO{ADx44f(tSVv5zK!7Kbl200Jcgy7YsDMO z9ue18kl{igAIAfZHoI7M;noVZgXXg$Egy47<58;`8~^5G)4AYayENkLF{cs=G6x^B z`?L>npl>C+%8Vtu*cQf|!nt|jw0JnsLaATHl4wlBS&ZFc>=0AUkEwo)1NC`?1GUXs z&5t5bwr-dGP=}4XEgT&>)v>D0(XmS8?$j_1aB7gDpl~>O{Eb1icW5_mO{N%fn;tSE17e8#ssT|Zkh8X01cgmk^XA22| zM5i1-((fL>7(JEShgXj%KNgz+%Nc#Xcz@)Cn~N7ISDoeS#0pezb?Zhi84EW*0k%TH z!P<>fdIUaV({v?gJ{C&=v1()vm0sUotX7$v?Y934{cEm>;+^QyTU{h+BnbAW7HBcC z*Z~|-^hD7-QbP;YTo6w#aq-&s;agUh;_TS=hi`@Dv{o(E)2d-He|QgsUilYS>zn_j z#r0T(X@3~dIliEzXk?MrO&w@mOJFu{v>R2D&p)qHyi69~sVgnRkSQ`q9R`M`SS~6# zb2am>TP1FNn}I9!!%iiNhhy|E7hh$HKQtyX`Gz{0xE_=7c-37SmfX6mdL1wc19 z{|gUyq#HwgP4NOswd7=-M*sZH=1x-n_t@OrlWC$8Ko@WPGTvm6g#P(%-?-qg3-pgZ zb<`tWhR`VthwmF7`D@7A9{c7B$x7MM=0~;jBmlP{=ud9fgzo$;`V4^FjUQEjak8p1 zP;}glf0$2py)622pMGRbe)X9Jr+bH(fM+1-paP+k0UHBIVo}nXbaXo>^3{(8HR-A* z6;G_)qpjo5;eHhh>N3ov*9^VR*0l@c?P)l-q|+MpclNQ zmq-lOp1caW9#)=(u=2p5n7waNOud%ET*Ip4mR7OBIY%s7?JpGnRCAu>S&XM$LzLx@ z#K`t>eR3m-#buh{3Uy$+|$O!(HUmY+}eo1?)B7j;+1ck<(DU6)0x+68SKt{@Aewoq2_owtxDZF zj`M=^Z6My0@!6p>?kBpvs@AD^hcql&Psw=tpCh9)R`EYgG(m(MGY1~qn6`eovsqFSPCUx7EZ*3THK`?BKYAe7E97gR z7K&U@PWyOU#^ch!B-r*D_y#|8@oJG~NE%Of5BH%$6S+g#&$IDox*dq^vxJW27WA@3aSAzZQ`#We3Zh)imH2Ucb%v6N*fU3FeA{5NoKusC3F zk;fKI9F8|4QT1wjUj2Z5!4jf38OoIw_F(QaI}TBpK$hS0Fv`wPYwRuMLc4E`aWNlp zm=a7D7Mzu6zr5VvXIfh~yO6JEI?Qoz-?ri-T1pbAC4rXTSf!B*tFQ&&@* z86Tsr`aEYH+uAX!uOxYG!^DEt$1pnYxz8&GUXE^q zW1sSt1_mzWjmvw2f+wO&dX~MF-A9X_!bjvc%X5il^r)2 zMgAF&*;TO*-;oZx%JuGYh1iJdg%RvK_S8xz-T0(R8!_lcb*m_&q=Vb`mak&kC}1sP z`lZ6uxkfFm-d~*4mUbQ|0sq2IM`ub#H5lj-YgVSH|$Mt@xu z2ugI@@auZ|Z3*Y(=Dn#mh`hAHXI%q87=3^3x*|K$A1D(XSNn2Fk@-Nr!CRW8uaIRQ z1je@YCJSN7LAx{~_lCmTTYON`qQdRb#^SeXEyI~#D#`;!x?Fo)+mvVs?2jadB<5uQ z3v|Eno`^gCs&0f>zIR{FV}vi5R{Z_|W2*00>JkYzVIVH(<&jE^U_t}E5~%{5Uh$a% zJXi5KCQ;(q(^JoX!mu2A`b@!4qu>!k`G7h$>9q8KWN68tbXK>PUbX64G3PEyzDDq5 zY~sl9ZOCU83mJhaP1q+^4E8S-IoF)d!ElEnYHS96JGClCJnReHO0 zDP>ZjhudEU1g<~g3v9i4_H_EC5+Np0AyK3r(@Iw!fRGLgV;-X}-bu2hx07uV;st!}HBu{|(xw z)oyI4>EJYe_r<_+j8t-1}-jl)On|^4PO}Y2kI)s;TAuBx82l?sgAVr&Id4Z>6tkRgBI zlOd-hyOb%JQ^%YN+A&5(1}3&P%rv3phi8T2c;fa28$`9Ss5v!E>R^mk(63+Vj&~}A z-_)tW^=jF15jT+=lbz%jgx9;(e0hTF>GKHJvoNqL9gIQ#Yu}s-GXJ(@%&`Dt72ljf z7zN=kaBw}Xe>VXVQ;htO1@oj+!_#EKztO!) zBbCrEbYF1hvjGQ)xnR(}5C+|A{Dba~JUe6pEUx>GN38Aba0_i0%Qf9m=1NAoWOixA zw##Da@bhu31*z9&u^NPtS2C|Bgo+R$rg z6$Dh)qTlf9hyB2S`?dxBch%`^t$Y++wkut1Tknst(pErL!~Bz4J%*oGleU^7%y~?+ z0gX_ATA6#u;j&TG--BC!$8f{#4d-x^@&_7!HKtYl`v z{iO-s4=_P;8uH@o97Hoh$-mw}tM}^E2dqqNb;|*#1~-lic(%SL_@# zb3Sa>yh=PanVd%FIYF6VipVZjljIe9#0{4oLukV1UmT%+@IP{dmp1+wj-bJ52Lf{Q zF&+5{A8(`1QB4$O@ks1$#}qyg;z^G4M!x(x^IO@f*Z3-^mfNRgTPVlH#{@Gt>BACT z8_6v(Wzj>T8Rw{Ri^kOO;yCIIxyPX>Iv#^tJBYq;tm;&8S(Id?ez!&d)rB{@M|?YT zqhNvpnA%=hj#0ZAhm=sDjP0iN&{fRm4IeL(8?a6|@fy3MY~pRYA1RM2G5+#%k$t5G z%4sNci`ECetRIqefWIB-PB{VAa6YaPNqi2cVY1WHt5yV+4BC+D@P+TDYEkM;a~BG~ zD05KIB1GC6xx%5zzxGdWSD~0F%|jjt*z?sbdB1Wa)Z$f9>nZm@?B>L^>?zw7hT4}| zyHD&VN%AySJI<<+>+un4>6EWCRxMc^`($KJmCQVfr@^Tnr|oxGYq@X2E}R>oGmx)n zrFny+5AvaY!JrPlW(J-tSxA(m#>oTWOZD^JC03k}2%s{vx6tT3eIXHuKZUe3L*3)! z8S3CdY;pOCco2EJQPBjy5!t8Z&M5%W0?UeCxOJW83~Eu9nLLqbjay(!@8Bm)wwvxX zVjMAumvSuh(*sAYV0xj`Yu3h~nIoL??tF(Ib*Nm?_-1%G)G@O({e0t>06QR-3=r%| zf=Y<_UYJ-A1Ki_4dT`us&+)&le{YliMiE)zN7X zPJE{*Qz-JVE*8EV$%Zr)Q8GPx@sjhiF;~IzVZlzwF+z>x4mFr+#(;eZhSbU zl`?)L;vk(rK{T!E#)I+D(T89>bi@muckvQxMWtjckPj*4jQ^O@4G^_k4Y~MoO}^;O zck>15R&3aEaB_Ln!Cjv2I+&E756(EC4#8GKv0?bQy@#y-!ncSft{lYey~BA3OH&?F z$5MGn^Rna8j!Hs&(|iXR5RTZNf4^@U%*@-6`CC6do0Os|G&$TiM-~tlzWViKftfq8 z39Sk_rT*4j{~LRgzw6iq>zg>;e&H!LFDhFFSjGD&HPm!j$|%jht}Eur=a@DWl&4-{!27{Uc}T z3o8VV6GmE6jE3*awmi(y6ZPWd%r%(x;)C2DpmWe&euo-;scO$te>oR{SA2>6lO~g# z?Okq=dtQEKNI{6-tr=^)UOh(tb$v2V1B?4=g3UOO_zJ+kWoJXMm3!Mk1)%92=j>0S z7zfT47$4Mf5 zVQ#8;%^xY^j73}HUP@}yI92jB78swIpVWCN!-u9xuUH)iJY6$Df7`s>A!6e+QHSlE z9L&T0w{}ISu^slQ9D4jlh%aX|%n(L@a%zPi<-j;uy?<=E`lM;QiIs{@q=}tah|1Fl zxx0!Hu$FkvFalQYqkRRqV*K}ZcUio%dws+x*0;HTp8e#w9Dj3M0}RKl{mF6RK0)NU zJxk<|8MMI7WQS4rgQ^AZ&5pZ~hNyAtN#xp}V~U#|49d;y)aovF+UIW?3d zpx?se-gt=F`GC`aG%>MW(#LfzB7(lTT=k4F@!}mhV(894;?=WX;??#q@rwDEcxC=u zypm(@0YZO@SEBzeUYW8%zdvXl<+xs)z1c92X(hy&ThLmb-~f3K>-ss}gc=ErCYgDT zvo8*(szj`6@_y}gv#C^Vd9(>pJ$Nch;JNA!lb1Vo(ai*A_cl&-G~S?>R!mGJo>3EP zs8)-4Q<;H6^_0eQXNwR&VkWGxBeQ) z5=|2Dw0X%?$-lB~PX-V_b!yoo-AuO}Hwy&~3E~J~xl#>RXSimBO!}gO6(t@Y&**A) zWYpkmF$_Ty@bVfi?>_?dUjN{+6{6U&C!p(;8bWXl8I9)jw~Urv?06@*Tv`$`QqM;7 ziorAMb7HFXzSj(A!k-G@xGkl{?Eobb@FsKg(jqLdOsUGLwWh;WPIub>6CZux`@iC& zYV!1S&oF%SQxS%b&VBW&hzD7JypPz}xTfcrU6YjB?VC0`Y{_rh)_mv?qU}=kK5D^u zLx0&@JE=H=mP2~j=~edSSonHu-%%v)@e4+2#a=co+4yj)=D)t9k#*ngOKa8-37Rp7ZbE-Q_=7fP3{(3$W|O}Ux?lQi*Sl0n{_ z3d+@^nN$ZFX^IQy@S~Yys23MaQR?HAHapxcg@Z)lAW!7UVM{J}vAMrzC=7=Y=2d zfpH1YtW|Qel(bY_n5ZM6V#j)B(N0Noc4-X!b(@73ia0ghQp$WE`-$;Kn5ba(@%6Xx7BOuDe>*phln>-BLKAq1a^JF z4Nj<4W$Q6^y@1=)6KT#sMC|Z8-&Ty(^2||fV%N)w2O`u)e-(|XdFG5OFm{Wv6-+he zdO6-e#0hsGV%nSp$MaS_rvBY1>WP!1PYD=@$=?_XstmcfOmC>?3EgL<2y! z!zLVS*4IFCtx>jpc-|Nt!gVUBNCf*QX*2vu+US0gwnPkRKGdlKamM&Vd zqQGb4#5lU;vPf>(e~Ru8glc~Wr z)#5u$Q*pT3xr=x~qZd`lLWU0iE7wkw(3r7@c@Y~cCOOTTt*5wjX2(QJ$$M8x;*s2b zoq|XnInz-e4H=uv#X-5735bnze#Cg0EXuU*@HHXNYxFwokUw6hX5K!vz5bY)MW8?{ zb)mRsZgPGFn3H4wG4|7Kv3??Rt=XkBKcHpzyL=f3p=@-u*6E2=Sy^hqe-L+a+$YlO z&e&@5%P7rwqu{)cB<>HqBzoIvlW|6!J?(b`GI-zF{Q`Pz1EE&}?`rpZiym#~{1j%i z7$iiam1cyT!4_8g0U!-U0DhQ$-LzVvTDaHVsFVap``q?b{FUayh7yW+W?HLh*&@ZJ z4TduG6y-{DGt)Ko4G)(gi|Q?L2nKOS3TEx53|I;u)q@7D8Hzy!#ED3?ee7R8u^a;_ zf;Ip6#D2i|#IpVLi3MRI7d$D!L@rn?#zZdA*1|+CKok{s8Y$aHHJ??VCVyH+i+bEN zfN-WidK8pmKCftVTJ4^mwWzGR)R@#+pgozG)E%fj>YUVVlGLr7)G3nG%gnDc@X80w zU!iYO@jtdI`EU?Zs`r*yOj7c&+&(DdZ;i&NAps?r9k%lB0#(e1=~^1c99j$YXCDBo zz3w{2waUp4vI{s*;zj1YNkefD1^bQIQ|9$e8B#lJNJ7aoqhx)UGOCDkj_>CC$?lX_ z9!}ZRv3Hes)z`lx5_wSViWbL@R?8V8CU#sA*SkC#z!|N`=x0iEMZ0O(dtXl0Km!5{ z{NAN(%N(Kbj9YvJ2HD;Y3AIHVw>33+9bK&<;*B#?2uc+(8!%M4UJp<1lF)43vAo+I zYD$iH&QC%UHrnS&zAgI3F6MT%{QzADTi|HE;zZZ^)MC7>_s!d%R6`wbf5;+T6ECd2 z_fHGqDb69i4er8`6Sfga+jpABk7rDHHsa-aXL>IXL%TwY;@JgmWy~>d~L$KBWD^5RqPI!kiuoL1>)j) zH0Aj~tfs!?Y+w5j#Hd{#CPJF>S?GIFJ&-kjR1Hz5z8IDC9TDf81z;JN|Z~Ri~nIF<7apePG4%`0-tyMoIs7jM$E4YE4f)5yx}GmRy>uRFGJh z^mY|#%Is7to>G^5xU}U+3Uhp<#l%Q5~`vt_EH$NfzN%06{ zrm7YQ`;6Chi#0OxvG4<)1txnc zLTuz9k@db>$8BMwuMi38!2!7_8o%C==HboEe3Seb6)Equ+ymHN|GP2abXgns%g}U8 zF=|~of5a=kh;P{Z|i8;N1IW4|yQBUURg3ZTQcf=T99$vUqurpCuxY_6y;buYv0TTB7D z;D*+nkhJ>gX_2Rbwch$x+I_%p72xh~igBcM&q(%NZ~o+@v z{rTVCu-Oq(bt8Wn^j<%8_ul?&=$gs*SwM}c??jfYN71V}V!xOZ2uu~I*52-5*P3*F zw`LKYKo;2b>U6*VRLg>avDm0Q#2C5qh5w&D@+O1^6%!S^3mZk3jlP5uo>un}n21^0 zeRTu-$Ze#%+oLl#!%i`_WT(SiMZuc#$WA4xn^(iU8|`%|!Q&H3T^df0&ZaOe+>@Ou z@)R4orlvFvA)+kFPJkf>c%Y)7dS_ix51+aR$TqWlzOG1d1XiwWM38C(i72*I|0csT=SeCQwfPT)vC1!DnQ~X>nZzCOBV z&GoO9f<-*Jmux6wl^A$8!3|$P&F&BDd+e+f{@Q zz++SLY=~9Ur2i67Sb5V_*;2B0of-HbaY9A8HTyPEns3oUUcn zdOdv?FTrbnOm>0pn^m$nC+)eGSUoLh?|y-aE)u8KrE9T!8fE@Nt5BT#&Kq<5cAIxoF7Gtq@Khw4Tj zTw$okH&_PZqAtFfg>@Jc81NR}uRGdvjMNhwc5P)YbaW{kL@TTPwP{@R$}T^V|GLfV zEIQv4bFcPt15Yp%;dFVEe9?Mgask05oth1z&^&liGL8qKG!|&R`{wBtn;;<PGYCNOkih1`~MMeo$^r=XMu*yS4#!y)q zQ&1Bk7q|vMLSs|k_4{W;Pn_8e?6^%0*H6vp7ZVI4fs#Pa`F98u5|#VT~Fz%pw}(cH{kG@ zc!4OZ(ewy1>TAnv?Q@oY{iVM$dsyixs$SKjW~~3x+nuGe=U4oee~qsHocS(by4fvX zn%2;i_`E`2sBTm}r`AVXrKLIX`Jwe@w>zdLDe<`k2-Apyzu`mTbDbKOwLMIO^u+k0 z{Cum;?xNpKLK2^=cMk4u@MW8AcGs1qHPmBzc(>WDRho9r{EDeKSQa)opjqp}?Ov9q zr1gsFxwZ;kpJ1P0?U>b(0oyjUpbE{yWjf%6aNsdyjmm0u=YvUv8VI6=D5->(hAy!}iI^aQKP0eI0%hiRF|Y+J;plY`95Z5!VTVBkcImUSivJMF z=HSK0d6rg>z13C~n|DXEl-@|>OxinRkZNy{zmN50RtHc?@_+cUe-rNmfnAj@3}|L* z)Aq(+YhK|}-)lFkVTkXxAJ7yq;4ZOJwE7ijg2#Zn8N!wV6v!lY)*KTIzv_IxO~+IH zu1a&Fa2^X?!aAcW{EgA#>02b6QCqD^pM9PFNj^bSF{2crmN)L)OUw>$xNbL9Q+?DJ zoEi&F>Mby^H7IOV^txQ-{jke>MvZ({3xvW3XD))I3M+;~ruB;|}KZ;|ciU zELPyzjT7eE>GcopU)t%pE#B-LtPGtYfuk9#!xDd8^eAlb+3{!_v~I+xIDI<_9ta&> zG^j{#aD?h_90CdnF#)}l?OOkDdB}XL`1q!8Wom|);2Rq_a7aPIOQI#9dfC$>TWS6R zRi!?UdCsV}I4FUApVoiDeawLU<~Tr--(L4F%Anyx@TIsxG%OMj0kbJplI@>An!e=a zRU`fyGP(v9V(*=v|8d@Md4W<+y%M(S;!oXdt*;Q+h15EGU-HplHzB1&(7Z7bVr;l5&$m6r0l@Mkhl7U2tR;#7&}?lU(MQzHbv{jFH=^3)Hn(koC;`8*L{FKOAS ze@&VlE3A|h|MmR&Urwq|@3oh5#3`?ZTa;YV&TE;_H~)HhplRoq{9_NI+TPYjam1Bv zhFfHw%2dlAu-M->p*MgL4``psR4<}w8*OcE+urgg{%+Gn{D^rq-0hf5?MSA&&Og>3 zgy}C3Q-|q3^H?U^)_Te>mO{paUT&6nKLxZ|Ef{* za+vCI7iAZfW=0`;E25|mc5yjdsT+luBEm3U6f=Hj=|cv8xA~Te8NbRu$8Q6p&_cO< zR6VA@4!!|3O!q2{j2Jj;TP|OoZ5LOvpSn@p#MtPh{1bZ_sHlM5J>8+Ad_LcN6ZM6I zd2s*!ZNJ{B|I7Th!1dd*+p+yys&l@?X}OzT39O!Q36OCEPIDC0&aim17JZ#6LQXMn zK{70!hOUUdg^Q4|Mcrmj9eX#$QeC@NLZd(&a;MarO7MAqi^hOIN)C_Bh1;ShU$T5Z zz04>c$k(?DjJ(kczsQb%_79$=sS^nGvSJZS=L3m+kC4i;JI%F|__+QDCS=RKM5%Na* z;OFHZYGO;a>^F)rMMPR;FlKO$$g(G5PHx(7bk09yUbX)@I5GXzrAHcqfOOD6Ox>FO zMurPNW^k$qlglWktij&_O_zPWF*M2CjAD_)*lI0z3yKW9oFbzPNbHg0jjrdF-T&Au zYe|?_vdOJd{cb_rJ-pW<*<Z_zvmD&dEaB# zi=fT5 znE4~k!M2uxo^apgxn0Tp=FNoxTKeGQZ)%GLcGNV+xHTaUR1=xvZX z%S!CYVY{@aL7cB0nQZ(!0@n2bWpKHnqaZJi15`M8@yPDba4Oux#>4$`v&RGz#fa86 zp>Xo2mqSnZPbY^X%0@eKKJV4d>09f7CJ!-C@lY&ln4-^hACZnt$R+Z9Et@0u^pyAN zG&jAl?M>E zK0+7~UW!#8`IjP}6`nc=&-`Aa5&NTW9U*A1HK%%SlL(H~8oFXOIfC#qVEpFZu!)k~ zkZ6FOJTPAHQPdB8$n%WUjaPa2w?n=o+!V@5Lj&`y5A)3fj-7c2;tXFCtfmP?*+Gy_ zJGw~&Sd;I-BN&HweRiJUq@J?V3xU-T$ZPnkq&M6{27y_V;#FyT#5MwDZCcbKLMi-V zIo>4q!PRW_q(^bbSX@Q%e`Ff&z*l(jkB%f*?f#Bo$kZB27g~C?QBq2%^#j!~$0Is_(${FYCSkt@W+Fc1~u0d(X*U zGn1K}oG(_Tz|Zz(Jt*2n$6UXTwdNZhlCIFTSbORRs$5gNaseX! zi|S`AqW1f}8trtT+HY({XkPHRrtwEZRiiO)CA&|&CpHU7A4d%eZ$EC@pTI-jTnO*Y zOS|A!?tspPXQ5kMh>k9977l1RyDW6F3o)=}^`kmFecW6r(SBJfK&P~-Isd_8NIZ!Z z4#^I;hSfJuHMHFDEoouxTG6SfYQ2sK&wWaM5@y|;*Sp(yX4u~66xo!$U;vEjRn6t5 zY)CkSyP8?+aV4?PKU7o4aU24wd8cym!pn107o;XRuXwh}u)l0TSoj&UZmwP!76=C{PIo|uT^eZqeIJq8@`hC$9oO(g8v?CGEF zyxI(VpA{C@_|%hY!Ka$8e_p)y;rn=s4F3kB#x+L?f9-TysE+zO&=IR+X4cd(#f-(d zz1^cPhU8D`?4MMWN}v9{BSR~vTGD9O$mQzqRyvNi^stmM?M?Zc4i=ljdBap zX-Inbi((W)(fO}8kNvMf!(Dsk4X)k~eDkY&;mzx~_fLL%eONa>TKFU*CLbAp;ms*c z?vFa)?{MB9ZB?7%e6Ndlsb9rCp?*wFdYJWM-RK!)&&$u6@2d@c6BfeS+#{+UN4>oF zJg8?Lcu=1~U{7d^T1C}|>rXn2Z+-qg4sv_0&H3!7&JBNj@ztYL|A$hA;kK2oGtCz9 zuYX9r3_V$HANlLu^qGj(`R4(zAV0sv&M}sgcF*!K=i&o7YI9R4*u#v#cXh`xA;ac5 z$kD%!o+w3`l2X`45fmQs>V>0F+I3297`N+r4SO}mVJ$!(QU|^AJKzeTu8*#B%=u%8 zet32sJ12rO%kT|5wbCj*Xh0}_T~Iimd+6ph{p8u&98(m{-!n(z5e1H9@~~zP!>;wlq^m%PdiISB zIAvNdk_CSF)7<}fA@KJGrN+uZ6}l(CfEMG#M;oq>O%g7Ti-+&WKF?UXHck9BU)tn+ z|L^6in%n7}48?{i%d0;{d3NI8Mjws7Y94_ZoOF2szsLMihfFS#j4EV#^?pP=+pu3f zDx;sbA?(PUIany4&^hQE)8w!EbR&HI+%a+6SeqPi;%)62%O>%{E{ECNuClJIE_&D3 zSABB_l*@81c3p3FG1j5Yhxh_Q&t_9-XZttXvd~1M^P2x$GAPT` zD6`e@>~Ac$#r`9%+iku~oqCpb|D;oQ7GIs@Wwe|7MwMpDjJ3wacdxCbE><^Dchp^2 z^U*u6_R(kgRCCH<WeHAbbzy_d);(qY4n(wixYD8_%iG=hkxf;&MoO^8 zPMxy$GA&lwh33p7+0--HPAg8Os=ZKO zwoaegfWHgZ{`$j%I2~2B{%C3;$}RT&!p&DK@#{xvf1iGJp7<+DTNH%5G1&73!&-&y9GTcCT-~4^?ifx%GCgUc3={AX6?^kE}Wh z|07f$o(RbU#0FGD_1V!YK6;y!#kPeui~W0zn%NSU5|EZL!Wx&Ack1}K>C6wxguE7Z z!*3#Y)pX@lhN1B$nonhD6!z&EXIy*clkBV4t)FXJ=n$s;v5DAM_wk`#u95I6T-vN~ zcbN8*W+SOpc(uDt7Pcxxo0eIT*f-iazwEXCNjLIZj?t;rs9ac~irTY%8L#j3UFlkg zXa5Hf_QJnF5Ki(HkJNL#JUD-IeJmuB91?l6xy`oOsBbF09@k83lWqcVSU+86NA|hw zD5v%exkx>g$h6DnF^BdVW?$UBzu(6!*VaSI9YEn;ox8`MG~K-Ay-q4f7NR8%i0_rc z`9(3(1aF^6jc9y0P~44-n(%ley*KQLLW+M^95^{bw_Is2vvKxmj=hoWXo?_eY|RsY zNC^+=j-%N7PYQ5#Uqt9yNbP0XGvg@lb84z7!Vc3LA>;y|&G*IubMd`QzuBjCny2k4CH!XeSwLu# zdzmM@lMUav2OT3-SpQ<_Wg3blPq{!r4gm#zFhC%?Qy6X@Y*Ad-E1nHK=UfPeZ2 z2MCV>gupE|Tk}kPD?c?;{^GOyTb<)v16%GA@v)@nqHd==9GUTuAb>I-_Qb2pSc=@?VAhj3LZN;{|)@iaUbRf{Nj7mTts4#F>Y5x_qIYq4X@&^cyK*gw`4 zv&2FN8I^of^B#$N zswt?y=iTvx-rjg02)LY*FFT4jkN4MBXBu3TM9teIgGrEX*SQk!fhA(~Rom1rE#7#X z7dY%18a$jhol5P;giNI0kS(3t7q~ySw5>(f8Sa7KM?T8*la@shH1St!p4d`;>An5V z>RmaiQuPF1`Acv7dyP1}ypbcT0>|NpYoSc}Y|tAR!Ei^;zS`r=UZMi3f~)YxJ?(_! z*>}>Ia@sq9Gy)(sV)_x4Pz3BCD9Ft5P<#=v~t=vs=i%N#R8|NTDNlR3_^V@ZbQ zJ*iUMixsL{i$==aIycB4a)efOA@^#@V2idHGn1=~N|ZTxB@lB36J}R`(#*O=Qf|~W z1@EtA!8CxWpuV%xzy+Usb2za+tJBFo`Og=J3`8E*w1mxq4k%4;UwBvH4@B+M|dm&yQBYJk>>q+h- zWCZ>-L&5rl*gdGE%(J1D{F@IC-PhJBoEds^jj44_N!!$GeBTL@S7d+{F$7i6$5`V-_efWDl7JabU>srB1cu%FH zwi~mOqlJ=BA%e*+tekZ!+BWh95Yc6Try1|BgJgzsb$~l@6vgn33aQR(Mnb}nPl|g) zXzWuDl^U2m+I$?z?4zg`+@q)?!a3kFsezn#;uzb%MD)|zZcLFzVtw(R-TLJZ=jBvf zKjQt5USp<`uW`cKKA3-l*HCN-+Zb`0_pctkb?P=5HDhCfLwRXN5TGsB$1)qUO;Cq9CU4y1_X;%9t9ym~%`8EtlVT$Y_ZaL`Yx8iw zDaouchR^2^j&k6XcEUUNwX0e$mE+)+&-(-{v8N13JBLGI=A-W7c7NF!yf{0&bn&bT zT$&t-G~j5vyt#MaXW8%)#{x?^F8?F7wf{(-W%70RlJ3uCvLoe>&yXf}`(Lnmu#cR^ zX|t4nL$oIUkGvfzXSmCcU-LH?E{D|TFk8w36h|y5e)2D>{k4g9Z{YtV?MOL2lr%Z% ze<7t}AGwfu!F2Kh_dH5cIWSbS65MhRY9TqoG%%fPBsw}H&Jl3R^i<8Ns-tCL5!w)) z!b8o=SAo2|Z{12F4Tc1qd{ecoRcX;sLgJ(cT*Fo?x?f;7o*D%sj9m98`$=ToldDC2 zU#A9q8B&r*W@K}Kzx*FEXe~e7B5H7OGL#4#bwk+Q2!?zfA8#aH97SAu+slV#ycP}Y zT&;E3T`32jqMh{kQg#O#bRywRE^;)xQXA z%bPi*ciz~^8nU8$P@v!X%n;4C<_eNWH%fSYqJoPsY?~(eCLNrp#eU}u<9uxD(#Xx_ zv?bEZYU~QDuByJZoc@9It{>ZCQ}18@D9mRK67;+z z#Sc3EnJX#n=?sst3)}^L3z;yG@p?nuoj0k6lYG-(M$PnyH;=o zH6Nb(FSaE>) z#IV`tps~~dlY@Dqcoic=VDjj2A;dp+r}l|IHy0@Hb5 zAOV~s;tK|?l;Adm-~Aom@OS*ZzvGMk-iiGM*9DPuK>}V-{8QZV6rO_&1kcmmM)Q)z zbY2}OxAWO!?W2e>+xuTU+BqVG91#^=Py%4oiIcs}8;ao##qox2^MIoNVo_SL9Bt+x zhLu!qgB)TaSXHvgTP=V_0Pb!zjptq7MWYjA$|~pX*N{%oiPd`hzJl{~kBud~$5i&` zKTr+4?qGb`6B@b&0~8d%CqVatx_9A>uDW*QTRrc18B+X+&mpU|A(2!h`uQ1$FV7*MlLnw`1|Sk-6M*vF0CW;iARrxosF)ui0C9i+ZkZFHwNLeu3)#QC z!=Ptw2s%YT@#{F*1>VpMZwQd9TbSbwEpOos@A5Eh0j|o0Dn1?CSZ}+^j4ew&G*#@Y zBL4yB_X$_|73a5x^IOL)ea0<)#8mkA z1hn(l{%g>M!S|8S3kh;YAb|*IXe1O57_Q-#zTlR=Zpv@fzM~Z%N_;`q^BKN|k&Hhg zN?yYM>eJz6;IRa`F?rL{jp!YNmK}UGHL*iAMRtW+qzb_PC3kryzT7YLLTnDkR@90uVQVM^F%-avF1hk~GZnghf#f=?z9%a7A_{Me6 zi=d~~H{KQ!6bcDOg@l1E0Lt}3LVh8kt&l)8LKIlMJq>|7KWs2I%sW%l|Ni=|FvbN| zZ-IXerc|A6;YcDVkO=N1!oU`cNQC??w2=rzUtg5T_Snp^4eFcgQcg2^L+yBsCO?UQ zClNl82wEh9Er|fMyG|lVZs~jmpbZ1iPCsY}77FL!Iv?)7Xz+X}&?XU6x|5BW_-)26 zpV~OZw&>T5!<24gTj;PY_HFS@kqCd22!Qk$$9m%1R!B5D)0L4f}T+=&6|!tpKo zFjvMZ5F)^miDsQ-5!poTqi|@EN3Tl{qNeJ^7!{p}W{I+hibU-n+M8^vR9s{qRoSt4 z(gE~@|5Pqq=>+MhE0maNGn(+;_Au+yEjQ5qvAqgJZ9^>tz>Z_!csL%lXX^&AGzrqP zBfBfFygu<6>(%~X>i4OBfqY^8fVpsqP~z!_XE;p3g~tU{v>=+L)K6q*YUgPnzBn*e zOL06a9yOb~pxUqIyn?Bl8q;1CHG?s`-zV-wU0V85=gWO{U%bj+D$m^|bSuZ_Ro3e}up7JfxopviZI4 zoTuXo%_L~z)UKQ#ML$s|85iMduvxT3YOoL*C5T4tN26rWsKe;F%l%AnCc#%DFJ!J$ z^3ID)LZ>>Y-Wht+>7wVh^)tOP30#dlXc|~O?P@0B)PM0X9H3Aq zXlfBo!Gd=@l^QIz71>@iN&-En-p@SQ&&2dIQ!>f+GPSPs8IJdnL~Cf6pfqnsW zE2u|0?#x63WQ_U#_}#|IGcu3&^8CIZp8TNb!!khjV}RM}*yS_Ouq{B%pa8xAx(U=H z0XOzw;mFDq4}G)4NqBkot1>7=JHK%)O4II`I@y0NF65OcZ(=)df}c0BXA6KL$eY;3 zn-Jzj?4nYlFIT-`Ie*minLFf~PK`O$%N(V{1xA>A&7i5VIE9^C#(?~{F<|+R@jj~6 z(H&sa$SrJCa%Pc_oMn?-@0Z|hRLVXoMTtt;Nu>bvGE~Z5Dur)T58md3+V3;Bt(O_p z%VgaP{C0X~w-2brcR3o@wvD&v$1NdVrr{GoUSn%uG)`d~Z(=9!IA93yfT)!Hn{ut@ z-e$_&+f12zZ+GzQ{|JBEOh_-M{3ki7;xDC-VfN4^|Nc8pRpiv>^<Qp*SjL0$%oDYOYeN5Q1CI z(_2_OX1Pc#bdHRwx8n`YJ#5-={Z*cFbzUzwgl9Hker}@6cETsE!Eqn_KFhphPu^{L zXt3-btbw3wfG!7h)5GQKRea7miym%pwA0?VGW&HOLbF*q@S7;bsyaRodb)D|ak0FA zkh3ojK!jLclUSbCL8xWC)BQBN=N(s8j@3D@_#*#4?X5C!&;|6w43{f!IkGtpDN>U7 zDM^BqB)KgB%3ey61SLs^(l?944TabpbbHb9(x8YerO_Yv?zFef8B;ruF@ECEZiu6Fy$dyhf!`RG(DG++-sZTHb2x7k zoVO`%_AGAJ3|9hNG{uQWPjn?t9NI=n0#W+(yg`p0zn%2n%qCqN)pKG`>n+cT%pZ?c za8`0?F`W2;2_KIJkyjTH(zv;=-!R+qW=Hk;plkflV18(%RKxPe@Ldfe=(RrJfsm7b z#3@1&`i9+bdd1(d7V#rtcEkCahwBKzpCjL_%U?p!QQe(fYoSY;qK{$Muald;!GqW7 z;)7i4J(o0T5|riGgzXN&AL-&uZvS^9++RgGAYNY*+xmTf81M%8Bi(xzt9qwCOZG7g zQ{gqblSh@A;t6~4g!l1;O?bkbRr-hjZgLoCA_B0rbFF1BX}-DApfM^NH{#mEO-`(r zlDMR4a-G}-7|dY1d%4y-E@^sQC*NM9i+6F8d(PleO%hsu>_9IVFTw8nFiSb{Cl^J>1*!*|j3C zz=Ib~90n$&&-{X0U01kH&e?zi{h0-PcXRvK6>vX2gbu@g^P;L=?Y`guZFtW;;yJ4J z`|%}yd%7Ahj;%lc2@n28pIyUtuX2-l^-^~)X`X%p%fCsE_z8FWgssY4mKSp<&}@Bo zc4(pZ77;6#AN4wZAyEZ&n3z|Yks#(jW`bYjP2xYhjiF8Urdm^Rp{2y)8@II;`-GkB0$ zHMxYGJpXjt`v_6|Nlkt--a%(VJ|ML|+lZj;nI2-D+KEt~E*& zSb;bBWH06G`t@cGIMT^!=vEq?&?1Vl-z$pb4Iqskv(b zzLr)4iS;*X=QttByBORi_`b7^q#C+d9Y?{PywL=|5=>v5575r9q082B7MtXOap&pz zD_7{!%<;)17fX)~4G>lzbYWD(f%VUHMQ$@(KbS7vH4fi%5qrg*5Ng6$-5LJH)b)$0 z`4>}%FR_p)ZFh2R6I^&UKs%v^exlBzrO2y}!v|cM3kb}r@&9Mo87C#yKqM|=P29=d zm+6?90Cw*x}yz zZB`3Obj6BNI0N3~+-~^CJM=|yY@lPUvXoWq%9HOW-m036m;e0bjwRQghmK44$(z~b zG|rer{&g73HsO#GYa#k>SPYe8D%a>>jC7QmGngz_N@yY z8ur(Rt8{2+-bgWLk_sPv_bUeO{O*T6IZdy>hy#y?V|C~pzA-p$%Ya?aAtA_g2)%sZ zl`0*Y*?pKE=7jwtz8qJZ-3G_1a*0jebH0tx0&3L2+I|u~|W}V8P6+-um7REGt}KP zRNpg{(lcZ@C)aY=$ShI8j2r5^l=pUzc1A%lf>6AlP~3wrz7ITL2wyCNFMgg1(xK=qwvrjO#wE*lWN$6QG$RDY4z?To3*TJe39M(dt+knN&HVan5?Xa0?E_t@n7 za8P|+Jx=%b+km6B*t_)5_q3y)7CaIoJ`ziNB$oNelkFjV$VSbI7_07i$EtwBvb$B2>H>d0@O0D4}_mfv6U6*8< zI<5}`dwtgqc9f^QoWrE=mJa+BbdxsXq#79EAanT+ZT^67pVwfzT!AX>q>WM-(fys4 zcBj6TdBVVRIhH}FM4(k6J^u){^#17S!$;g6tpT&obJ~2nZ(r%N^xXTZG-Dg3XySD~ z?GwA;>++3_w})hfFaGowewT6HulCf?J8+|W(ZU0f{b7C@yGzQ{lZXnE+Ra*UJyfF) z$FHw*IDNEK^{}OlM;eisJ%5{RyWPtCc~{+QmvWAPy`-q4{OC4o8;dOBCoOF^ceviu z#`!G&zH_7L9p$PL_BPv3YHviqcYpBj6B!%90@?p~`i*(ul5|;TW!c#WDab?Sh@1sX$k38yDxUDi7pKqg`@pm+$)+EO@&V|DKYv^U~m3X-y%eGDoeca79>6TEBO2@132 z0OCumVU)5~l=8tSW%a11q({;e0X2$%iMNP%m;68ygtJowa3+QoQ{#BL)c!O39OksE z{D+4>#g)I*HBN{k;N>mS(Ivkl33B-knV+GQ0Ls!I}p%GR=psk zik&|2FC4~BD`Te*VjIjkdZ?P(!O|D2^7)Q=2|utQd{Dl?Z$V!0{!xd~#x;{i?Zuow zsO7W+DCM-5aXhWaNe+X}_fB3dl(?_)wH4R$?GytRv2x7{^qBRxmXsA>nfq5Fb~+k6 z9fS2u*aV$W2fV>4Li7EPWQESFMhtrcyrxNUTWc7!(p!Ved)&Pxx{68^I^!H z66|ylcDe-XS-y$sV(j#TEtCO4bFKZg$d;VnB8DwcV*+ydIV1x3?bA!;>o4j=Uetz8KwKNKD=&;2fMuZ~urwrn&n=Y?+mm#*np1 zWFxk!_%h4+Jh^7H@JS2Qcl+7K$-3^^!DY9?(dN??f5q6p2magNCRBhte4^E|VktJ^ z^W9sygTSv7xiTx6$L;_l4VIm^>v9ILSNs!osb-Yn>N^%tV8XgOSjW})Us=Dd}GW5-UCy5+tK&H61=?8rt{JdaD_s5&~*V3?cQbvI`5f2;plU--D}lnnPvmy@tGaIRpTJEHE3{Pq}W)BZs8mKXOp1o*ad^zsDc*}?mT zE~nqJu1|6|LbX3FBbKiOS3Mg^KM?E{$ZvQ$&OKQ@-1AwTW8jG-==JQ4dwPo{guVUO zTah1Yzb(VpM;|X%+!#L3t-e2qHh%QxW{b=(clQkmS8Fx@j!Eglho=^JV|PM&Cuvxr zbZii?Xla#Z`~tR1iwdRA+^~tB$a!XSruATCwH~iPIPTYA_@XYMxg}L zCHWIe@*7L?&498jA3q97F2)+1<0b(Y|y9L12OvMmM+pf+~rL~ZuMN-z7^IGZONJx2)kA;?%?#U{WOx5j~GvQuV?(ZuV!p>psDy* zG(gwOH^Z7Q#m+y#&X-{uExFWr;`i;&&pVWmH6xbIf>EI&@H)FQ!2U2-J-N!QswX#jv`CZEwJ6=%MCqepLYSDk3dO`*Uf~w8wy2Op9k0n67$#?(1( z+jLvHZR_3!wO9bR36Pqhc6m>1LIc%P&WFA==i2`f{Ubp(IV^#eBFXe%bzt3k_7}|7pg5dOa~Tj>k3GI;Hx~6 zg$*S%+*#jyw|~~H1@>fnp?Ks^9Y+ru_PJqTi+lMa<;d^uo11-y^mEgBC@yud4oyR{ zQo;yLW%$X@2PSIAPVwkq!DjfQpj$>dqxmk4!s;x}==jW7>Mv+zVC&8xbz{J%` z)O7L+{HF6U-j#aoBaluHx4iXi1%hcMo(j|cu?i0Z!k}X#gZqXqRCqmkP%@ZRfK;iI z%6jy=yAk^+y*zb41AC&sZb5y^ZR>x#9j!xGAXy%*#O&-nJ1t#8h}D6!k#l87<@$lg zZWbdY?h#SOc822*ER1E&9OdrO9%`KV5kV-bJhp#KSRsO-xCM3u0l=vU0y2Ui zvj%*nz95bTO`N~3UZXF)e$Zp?f)ji419Fk8WmYi+^q};Eyx|A=s1NdaYnuR+AHWqr zS(86H3*obfPWv&G5@EL^Ei>u^U;RDl=t z2!a}hQaZ6^YD#i}z8~O`#-UMNx4Eu&xf8%=bDKMnvib420v{4^qqfyk_SOe)^FST* zn7jC4a!Ij+yEZw*wifi5Q*A4M^^?DCbBXMM1@>n==C-vlb=sKwA_&?MgclKnpIh?( zn)nABzs)@)P~WHEFTzyhveNs}rmE~S;}oz?){xHS6>%G-U=_AafnpeHH=qYZIZ ziu*HvgmtQkZk>~0I@JIG2PUUG)$BXfJUZ2E610z5)S|LugcC&6eI*^re=B*TdXsK9 zu65qAD-p{xSB3^3fCei=!77_r1N4KNSUU)HGsPXK0>_A{UVX(w9DVFqjh>!3);aM~ zL>CIygAVIMDS!fouIWM7z);66eU~S07KIaUpXg#w%u6w$*D5|-WkSQaBQp~|eU#;| zzuPJP;Wq_&XIhbQJ{jA#PL%eg+i9iSO{LrYy(K-eaVwI*nI0e$N%*3hgph*d0G*1t}+;hms!H8WwI8^~! zmfL^;kUM)WOy!Vodoie~ylM7|b1+(p7I-o2JaDK0b^xOo-W@nqAzGH~01N{LE&Z=z z90b^lA%2@@y$7R7%Ml56?93tW@MMT5@PMBekE-bjs8U5~QhPC)Y6GxXT}0ATTz4!l zc^@rHXaLp$yzboO-waP_0A`y*HuqwbbmF@4y#9k2tnIh=uBG|(C_>oCq{vS2cWLv>)S2UHCM)e zOz)#L?OCukB3a7~szW&RkTZ6PSBSQ8O%GdoMZ@h(a+V1Ey8*4~NF^T%z?0AMZs(<_@#59 zX1svKUlRuBf7?Cq`*ns&4yay9!$HicE{Q%hr~6>Ybn;16gh)0HVnH2u>q`JKXxD`o z>Er|}MvD&bg9z-3=H0Wed(ba8ziI8dz)mNV4k1L+|J__2@hjHwma^6@OtQhhtyp>uXGvrugJOsb0@x<<4grkF|o zu7UXMzyK=Ct|n2f0?09ybKz(024k%@D&9LXCjYX%5cnRO@GRkP^G`QjM&^@y+m}4;0iG#XmG^kVqtvhoUhx z7HrsO5*coQfcS5=9IxP2o6=k|Fg4priy!e$oJe9NMg} z7~9OCQC#I7eLseL(Wd8dAvR2#L~el~K4Tdx9NcFYcKP|4v+T_b`ge8CVK$IK3>UWA z&=J;{{yuaK4;oPgNu!=rj*4=jB^`v5dGI?JtSFmvokVUnKwJr6NZ0Y;cnp?CA^^jM zd}T8PAv}r;MAH5#>Jol-_s!$Y>I(VFTnIwnpCSF=Kg}6`?+pKC>iWyn@|UT@uUN<} zAii7!gzyMX+k!@SV_N_T$>#6@SGE}aCFJ77-70vbb2LZ>25Ukhcl$9g!#H*&4}Kei z^#WdR@srGGkh>VH8kU^uwI$TY`xCvof?d3;Q|nzz{^ajP4DlgN=kH~bk4 zU@VmZr#k{u&Z(gQ&HlIV&kv}LIROPa(s0|KyhodhWkI0d3MBFg1B5b=aJ-S5Pb|* zX@HZTLp}jUeDr28@I3kQ-5tBkpEFa_t+6k>>yPQi@%p!6uoeTHPdQ{aU4%Yh-i>1? z@<@C%hy(_^J%_xZix37F3A{-?Abl=2#RBCyh?BYqMEhp?bZw?jC+@)>eNTpTG>;@q zqYDjiv~tMJfJ+o(k%@cjGO(FG1DokH5Y~bFBY-1@5fU0z+tt&i}T? z0quwM#G+-XK>_STc5g9HosX9Gf|`~>O*=w9M3ss$Ev_UJE`Un~?zG(|l{9i95oI z44iTesObi9N&iuZT%6v7>)tX;f<+kKA}qxsEZHI`$wKswg;Jsg>rI?m4E^<}-@an0 z*$2-*6`o-c96)V(E)7zaXXO+#zJy)32#c}^yJi84w17oeh=yAzg;}szF(sGdOJ01t zQqDCX&|IV*?Hw$VwCrCx6_lV}gir0eN`t5vLR1nBR5T4#YV=hE^;NRek>EeCmpYK51*fkcgN$IQgUu$ zavLIY%|db;JLMv`%RQntzMc+j<8Dm!{xmmKDlJ2w8vk+gOi!rc<2cj<;Sm~my#`D= zhblexn8=>^pxq<~h89*qbuEI9T6pMK)E}L%vYS-Zf~L0|J`z^q{YQK;I8+YNB&Ip% z>Fut-Pgzr@c1u=I$?Iv5jB~zf0ppQ&x08^Rji@H^x~o=mqiaYAMe<#@(;itH+Q&km zTyVR6NGV0KoH_En&38jY6`0Smrl-z(I*l;&2QH`QrmAKr+L%OWbuUp24BzMK^%I{= z4_*>tT@vcMB-DSYVX?kDO;zHMjY+yzw*z&nNq}oc**1G(hmjln{#V<73Qe7fMwME* z7u5tKYP4wd=(_*&}OWJ%9rq7F=zCTa(F4RURh$#Jt7kVJ*P!Zu)-Nj^G zsfER_19I8R)JCA{bj3D6+Uc*;m7wF(ZXvyZrEeeZS-RYP)w`!{bZ!!fzNc(N#Iw`s z+N(R@YSmV`tv>_;Pwvtw^PdaX|7BH7bOdd#t?v*O#I+;U;d+QhAEsYl>2UhpQq_uc zHXiqg(r4k~J_iO@R$t{B9JzpnrNX0*kQ*{#HPD+AWo7C)#1bEP$WY@v)xQrvn(kJi zy7#<|{e5Eo8STh;I6VKO?E(67ttRBcYgXMym)z7FdP7mZi$+n$+Q-s|t5i>2u#v7J z+FsC(B*WLe=nsYh`ntxQqA%K9I3g!Hd!sRni zT@&C>&TwM_bMH^#a<72WfAP2;0LdzY%+ElD;&5VuyaPg8 zkqJ^`c2P|vefSl9`1O6_&h`>*_-lckAh49FGU)h$Klr8$iZ}pjIsno#VOZxUS%QAK zkN7;Wwv&1Me&V*?Az*n;p;-18ZEXbg2#9V0(<1;tsYgJkM}YOQ(xr0G4mY6K-8Mj| zOyjHPD@xfxQ=ruS1hmj;KW%GmohqV`Dq@r>GOz_ext=PLpDNOpDne91)J0m~^mOq3 zU5JYFIIr0v@WuBcR$a;wWV||Ypr2xriQ-R0@n@s>X^0x0W{oGK&` z6eR(&uV}jk$Y7B2Nwc z2AxSpfm+bZmvIf#6F>ps(gaXI_<8~;AiRtdo1SQ2m^c6!0zA(tF|Rh|)0))H0>aeI z0>ad{^Aq-eL@==AK zQD0RqsFx5g#*_5I!D`_yM%sp8P_FGSM~uSnXk9I@exIlz@0 zZwFV!{CRvMA@|O)j-(RF^4dqe4b<=0fdaMeELtWL-->HLSB5hT8&@%Yyy79MVDW-# zoP}jBr(QH$NIug^ue!wR{)|sr3I3&q?ExQ0i=IP|k@wMtr&(c9MDiys#DZH`J8d=C zH!`gt>q2?vJ}2gTmzP&C!+|%S<$HBkZAX6!$~^bf-BL~|<1OYd;@c@L!$voTHw{O} z@J!fL8xnfI0n6T(n4KeMxm`=^!O!0N4{f`u(%(h5_zuY1KR*i0WM0I$XgASyG;fea zJqLUV)M3F8M8!s}A3knho>Ess(UHMBpP9QmvvTNe>ut;KJ4f2HZ&>z6bFADLNz_^a zG;g|(8lQo#P-E%X5_`1W7!K6cY#-jvKI*G9^idU7J$d>>nEv~#=FpIf!}h)Fv%3o? zZy0a$X;FN&MyVAZnWL;eMDb{`zF+F6b44^GFh@bGAPDi1pH_YCbd~L-s@~>IrW?kG z`Xjy^KeR3{n>RSWGrBn>Z{3Wce}DTO+gtsi=om6_P5ZhCv#2AYK^h;tLl$Kyi;|H= z?UhCCm!0#|Wh#5(U$^H3u}Z%uJo3bsw#(MXf^WtxLJKw9C(d6=2Np&*I$^p@X|Y4dWs2%Dm2{alp0>y7M)&4~J+;nr<7&d+y5a4cPtLfhy#ZK z{fgZLpoGPMzr=vUV!slwXp;H|_9*7wjk`o-hS#OaXgHqJXp!Ljh@%1YS^j(=Gd> zTQ;g|Qvy6u015E2TekiV7(JqsL`fSvl|;#qHT)xblpO`SWEsonuvG_*Q^2^L757!w zCgYRZ)R?>Rg4pgH^E+Tl3OG0kymkk?mIO|cMu{I}`q<->j$krh-MNUj%9PD;0jvEb zMh-N}j(&(Gz9Gpi9P^e!dGAAst1u6l#u0q=LfNdVPBX4e5`mm2fXnJt&< zU^TF}?W(t}g14=fw{0vy(3RubouicuegNN|In3m%Fi+l${2v4ZT>#?019lTfjmFZvK|soD6{u75fT z#?FP?v`Ek1wv&2%j`}b} z2bw=hQ~hh4vt#`GH1NvTaS8R#bA35G(!je8qP!2IW|dLis+*WSfSOg=!XXsZl4)Ht z_5JNeEzavMo^lys%wXT6gun5z#rfEd`PlC9vDNai&G)f&_pxR6<+%3deCx|;+L94{ zIc*k9X3*3GlS}*iLOpSZ|J= ze|GPLkC?mX%IR~hN)jrwM0K6hBkmqQ*t~dzjQ?vU8lMCtOrLAg0aUj= zGdTq>wO3B_v>vRa&vlnF3Z06Vav7+6qTvBt>2-BeMJ=!=lYHF{qn5UTVNlAHm5d+&4hch0-t>-zplX4b5^*S+ppGY=13JP%ZW$F{IDzU*V9 zShk?A0n^93OcZySZ0<5e@i_tL|I-BuCh0ZbVt@)1*y_BiWqDJ}FK*0kTWH#9`jQ8J z`FRxM38;XyxG@FjUxZ(dj7ne;H(s_alxa0Biz_>Sqd77%ictV6&{1&L_{qpo**tuC z>6rhAC`NgxfXgwOZ>y=Fozbn=i3z`I6AoA_w*?NGZZ_Xn%Wq!@x?%9LoGrF35IB1u zLz}jlmc1?m%h#mV+o6suHW`+|ApiS5zg;}hmh8Cdov z)#k{AVOQfO+mpN z2_1es&G?uon6@5YHgySZ=a*Opp`p&b($16y6RHn8i8>~6e$Uu7%E ztnp4fHgl%pUh|@WEo9K;so^PT!8zN)kv3DW__9SFc)w<|qa?HdltOoWS$-sR8`iJg z?8pEuD6lPT`0o68I*+Q(oK*q${%GV@bN-8{ez3A?t)?G%;Wa`j|DRj-xH3l`xKvnF zKfSo|eCyGb&bIKN%{2NV95CbaV1~P!s5j?_Mj4%j7KqsvhQ*c1X*PpL$KPsRWQFny ze#B>m-t|e0>UVvTyKG)RZ7|q5v=jD@|LG&wKOdCd9Nn}-brK=o$z}QC@P5bU#g9;@ z)cCL9nl)J%Z(lglWm=I^<_LrLKW@%PM;XOJ3$*Q>Ytl-h+4#`uPSHg-J8BjPe$B`M z!?(ZWH9}3d*}cj5PP zmdQGQv{QMXVs?yZ|HFJC`ukQ1IEVU8Y3{dk=fe9#re0sl7H!~O zsnJGf*$RTCjm@144~9**zmy$Nv$54>@5-^URaZGYgnPY=&I5%N4GOEJIbT?-r#U}8 zy5CP)+Mmw(U34@yCdCatEhY^!9&;)5A2SUai+=RgyDd6;X0tmb2mNZ7l;9F zo+Cy|p54h?W&ndIl$|i${8m%jMfAd|ii1Fvf8=s@bli)vu1kiAI ztZ>VF^BAdGb|+)mfkO8}*%{N#*0RcLEvCM=2HV^BKCtszUBCNctwYQ~|J610?u=;{ z@55#u_d?3HGW#TW7f;K3qZp|Ib|+fd0WyR=(bJCkIzELD{oUgabnw z%8U*2K0&w=;6hq472L7|%kG7KGp1kL%j^@5QS@S@(%GGE%MMt?!%fe-t3FO#zCFvr z_NhEr7{=Suj*lN@W0D>C?Oq54vs%mM-om>sw~Qd!oq+Ywy|5ba?kJl}fUhvhO8XkD zdC5wjgxxKPV`lbDLZ*<@*_ zq?sYD)wG$Rn=v2$`IwFy;_Dmrsj3>~VW1jUi2h-+M$&N=Q3}SGeOS4zEctzqt6SOi z_DoOzSrcLDt(N`y+KyPlmen4AP+rlAM>{tfO8bNBzEt=Jt(VM)ArTU=tu3E4tU2aT zn&V`Zz;g#BSf9ywiI}U>xEv{F-niIh2bxh7AC9|##Qrw!X(><_xnEOPU5B6T(LER& zH@Eb(@ma@v-WmD&a8uQFZ0t(Y?)c2;k7DzHuRcixk;-ip9(T#5J*6JCl{AX8*`C{3 zT|#X-dlY$vtmCsY_jfHJ1pX2%R?6Qg!XH` zc?hHl|M{(DLY*#BRnyiDwa0T;@ki5A#|G0J}Ng( zEN!aUSB}V{r^z~vLqwT)HBZN^Gq|(rJyla^qFHB(YQB{xEmoq@V~D$VDiP^>G+C$i5fMk6XxG`r z8v#%cNWKMf+?k;lkx;Aqe5BP&_NiGJwYujt>*w>va|Z;~>KY;I=fxYfG4vuWTvM}J zYIW?i>-my0tMno~YIP@R*U!TmwL6bHH)z-G>?1_-|77qqfHuGhJi|7n{rQDR4o?Gs zxojID;-(fTOQg4qk!(yt^H}tQI%rIa<7ogmmt7)6#J~*bl=ck0<-@SOd}=z@DJRed zi4B(WG!&;dh2}_g&|AX8_E^Om2a!AtVDU;oW=oUwC5url98*rM^p+>W_gG<#gULJ% zVUTqx_Xt-BwLpUxr)buNK)VGx^czp_x> zHh$vvhS93imXbyM9=QhSO1xpbmX_IA+D0wv{)IxJ8;nTbBd6^Y&pBWM`TgVBtZV$F zCB2X=R`MP>=WXMRr6aM_t_Xbl2IjeZp@@1{NOsh1ZKFo4ra_K&&EC&Zam^Z$3qEjr zLbndNu2Uk9(!#h-XV<%LwGt#|mmc;dVrM4}s}Uea&*c$&W#fwOg!%i7QAx4gN7r3{ zcXk?Oe12^;WJbv%YeLTXu$TgpQb1-h s4WR5n|6Kvu^c?@U=^mpen(qt9=;rx< z9@i(22cy*lm_A-4znkDlv)!yBCE=RV#yZ~zon4U#Da(wPD2 zs)2OH$I{WKC)v|>XG(0lmb4(5*mD7#Ilw z()QX?`pVpUS`3OKiVo;?;;O(mB#;k4KtPulblr57Rj=h>3L&oEL-$f5W(Nbc*f80b&TSt{@#a zSl6R2D34UYv*76DiA@Eya5Nvxk#Z-{{r{ks_uuFZT}=9)=`o2np6WJw;s#c>6D+j2 z3AskMk!a|Zo8YR{pc&aV{iID)Gs9^QndX-L7iIoWa^4QMsgeVHPDR&UHdBnKfx<*ff4Ax z0QA=f`k#z{B^&<;IJ_%-avcnT8d%r^Sn;5ZZlm@*KC40DFvx(Iy2dJ?j`WYZviH7s z8r;5joT1W`XX@3|ADoExc#r6JoUW06HGDIZImm}8F(ik+_%vs^z|xyXC&(v}!f*F^ z`a8pnF}(g530<*R$rlR081_~AIyZF+&%{8xc)9H5y1(jJ$BMhnl-uN-dS{acr}m5D zTjk4{R&MnC^e~y0%er?9l@&9`(<4Hu)WER2b+;3(&C?^Su2dtrIrG{9oMxM?Er~Ik zzwZvIX`h#Unv9FjPQIFWozpZjzyh}yE5Lz!l#LEgkQ|%)9J>GG&8_U$bI*DI+{{Lv zO0Y>Yp!U<@+u+N77ydT#ij2KvY~XdA4BfRCis^h>f=f9v;NaTaWLDh6*o$0> z@BKDDKY9k&0O`em^zIC)X$rYZ!W7*ZHXHA>p3rpbVdEgVW$`LL_S`FtKc?h*Hp{R3 z=6a{o#2XDaI@TtIhYTCZ@p!(h0=jfztnj68+p}FsU!VTLolX0)XH63DC+#oRk#3)> zClzf8v0d+3bJ6^`CcGM=NZQ_++{_h06-Ecgx(3v?RdwaMdnYz1))My0xGoTTx`v^F z0Xu(|#%7zs@vgbFg#9U(P5$730Rtz}<&lQ19ZRDf%kJ7F`~q>OLznCKF7{W0;u5MZ zIAnh%{E>X0wy0DH3D;3OPfBC0ZV1M=jP;{uYhUf66j8z{gCVHZTGFA-yJxP-N@in( z-9BxZ3K(I2|E*zqnq|t)?pW^VE~=hz(6x_aUMqG3n)bZimdy*A}FN&wt1*AKqW;SlR{FTJ^qp0&a0^6ep1)?L{Qc z?XO+L5a(8>(EX@1%TKzz=_0}JG`95mVgm8|gn8n@mL|qbsqRC=#(dRnu7UoQ?d_y0 zlPZ*@0%~M;X$swh-W2c~3B^6xoZRtW({Ar;CL!zbD=3s%@Gxm(qx^&A8fS5F?xY8J zK?;XiNHA#~O?=fi-`ldOyld%1NNMn`+C8r~O2X5T0`Ul+YSiU-O2t0RyGgEGF1aP^ zJ4pm%+XIH2&>4J3RS2m86}aw)-;jPx1kBd6hADMw-AnuIMt(?|5_eYC2Imc~*OOKb z2~BuQ80ww(qrRepL6Q?G@IlDzW|C`2?fY@>GC#_ylDd?J8OsJ|i+4NuyVJRJ;7PF_ z>-(eTSjR72hh6s4H$@+NABegJ<50_7lUoPi!E^)lJ#*;v6lU_?K<3dV@^2RNm=F#W zb-lk2ii9U_D+R~uJ?N<0jaEjPR)vgo?TCaDD|YCnt>uai5%`v_t@C;*2P?Z>pT7oh z_4RS>w>15(Dp6KG9JU;D@sbM)@Wo5+EMK?nOTy>v_jmQ9(*sb$Yp;4r+NX2B?zR)d z%K{SSl)>v@eb;DwnA4S4g8NaphDnk19%VBn9n0C3v2l^j;KAYg&Ym&n-S^8u^|dA) z@Hy9wPwD4Y2`&c{%4U|BpxXNLG9=&i**^4si9v08C057DlIDY6W*?MDoTv_frk#2| z=A@`QZuz*o8yi%O#W#8H`Go|4LJv|_9m5-~qPt)rhr6R=+J`c|SnNvuTHb}N(DCt( zSu1u7p*UEIWi9w?S8Xs=SwPutZ0?E9RdwIM^+m$i3VxIhkNUAvvMRh5kejNsllwSi zG(@}S#J=!FWgr=G{SgU$GWeirh0wLqPaqn9$|Vx|M}1(V9|iS;xypeaatFkz%??Sl z(rMQKZBl6nwoY=Tt8a~&Zo_1^D&^}!uy4rhW0dYyXj1yYP3&Al+tN~kB{n$ucBim( zE+i$5I6vRvnya_cwUz|ygK`GfA#a}<*{Xb-L%X}aF*VuoJ7s*U4@Eq2K)0;?nbb-+ zwae8T@|Kvj!&OAuYas1qlJ=ZP>u}O~-_{9#gP$qS#;LbN7p#AxCu)gV;npEsJG<7{ zOFJ+)_|UZhAEwJ?L+~=!^RTj@H$$+OL2rJid>m)kruy7+H#VD7TQTv#I}fI!XncVG zL^Adl$cE|rpU7U*#-W6QV7FZJw{{c8Z{dpGHk`_?x*K4+0^XPywBnjasHz}UZV%pf z&EGb@kIfvw_WBDHt!Xd0x}?B%$M!O-cGXx2S>_Xbykw&W4Y-)72+%9ncbf zrW|C#w$)sZOqK!5mUujI1iL#(Y(VW)9af-bOH!(~Xb3|oq_V0*c*tffAxL>=m+*Uc zB6qx9k07MHt+CzqH{#1_ zALRz|nK2CUKzW`E6?jlyuXF!k*A>?HfH)|yqL4;b9MwE}C&(jfKE;G zA3iR}VUm-z-Gv@onJYo=Z6In7h_xaxg6CRbz)#sV5^02xoAi`$5*~oJU;7cHaM0m7 z8L~OKfgVcC&D!zoQut$yu~e=bBi?jkZXreH-p=~fYABX0#mWoo6ig*{+ppx7qY&rqI8$#L_>FXgK1lI&t ze~?3jtTpALQVEB#xF($Ad>hRK z&w9@Z;!fvyo-~(x#>;W%LyUvAW~yGxZr1lX&&CMP#!}D549^?dav%9P?j*%5XldTf zWk&DL-xw%r$ zPE}m85H494m+Xa0=D@WsdfGma+X$2kW~Y(P<<X_kG;O76YmiMtl#DY)nqGHI$BJqgL1l4OtD;4z_Vq=7(DzQTG2-O7Dazmcr;z^#x7dG!{=7h_BsdsbB7hJ;!i;vSN<*rzdE&`l3E!7L zL+rimL`V5U{Tjrz(4&$^1tuD9MCGAxC;EUtbQt7X({AN$NY!s5UU~N23!`(s$7^IF zX3Z?7Iw+c`&CL45U@7x;G_YkJE$yBr?%jv=}%CtG@|kYCymSh+e<%a(LvrLRoP8P zoW)7EDWb2r^lNj*8;uUxEZo~WKFtd<)(l>D`rvrD_ZiZ_^>=CX!69H$VGfeqIZ2)4i__dh!jDvxymG0fc=ruzscu^jU+o=3f93z*R-ZoZg@~N<48(@W$(Ie~J}wTfLR1P(Kej|MmtWD!5O1O+;A% zb;fD7dA}(jm^eMPNyoevG8XWFv^%AUeN44kkA*!tx!n{%B+O0C);hvl)oDlft6rsy zrjRQ9xK#2kdKG`D_Yqp|K?T(vcvc5#evoc@@^wX_-pM|M=+}oz3|IQ5{8%dvb&#Pv z)<9aB4fptyhO!t_Jlvm}UP`|cP*y}*TV0&o7;C~~EgII;<&|}74=fdZII$&8{VLEJAq#>$AIepC)r8B)b$hFn6 zs}RynLK5qKcVmgc`^3Jc{%w~BQJ94$;^84~YOF&@Dkst6a78(`e6uH|cBMA$euwEu zN9r1Okd)RO&5Rv3!Po6K&|*FJ(nvMLKSnX^GeaYl(tf@_J3Dp{Jr*s2(#mzxE()&m zhmTZ7X}_LA1+%s@1rs(UE0tXAoNs*S76dm|Wgu4q?F|?F;x~&&#W!ew zyfxu@wftqGxPWR@;Lz3}eq6J5HM+8iAtE{c^`5>?hsLclKsH@1XyqZPc8 za9h>Rs835>Zmzd(IesCW_kR_;)*;*0ca@NtzN>h3jZ5z#XKb`GX?OoOTv@rn7fbi+ zq2z}zzgAH`%ry8lIZ_~D(#@pTMWi8k$$PNDe|~PwA6bihSJOEr8N+@$+w?-hisr{-%kES92i<{PN7SDs_5wYtT*=uN11|0 zQ+0QQ_lW9?N<@yTbfpg=4K?q?C&2H7bx$?mp4Wf>?O{ROS#@8`*tI0=l+EIUKh`L! zjn!dD6Pvm{!$Cf==?AlTH1@jFD6(L5X-UKq<*RIdwH|ABb(@2D*fC!_x;2m2Hn<#L_AN@9D+xLRwY;-$3hv0xyxLp$yr`^(w&1E3Cuj5u~ z=}cU6B?(Me3(Lc$!?AR_*rYwF4$1eQaQu5q1M|FV!kon}RX(AwFWM}=j`i`Ys=Q~p{G`wUh=JuwsjU?1(Why(}L*2+(GLLK1s}dVjF|R%_<*lp9os? z>aNDyN2MdQcA+VCA;^@Nt@-rg;u20OHI16fwf5R2s zd)Ys=+Q2!~YwTl7lsDw&Tnx`5FZvbz;}-haUz2=Rb?}JU)8sbt8c$?m&vTz^IYUp1 zr)cypBy(dIs{Gy_!!xZJ6*d?1OMgs^bwpl6&924a^D+yK5Jy&K5nR&KJ1XoPXrkbf zP;tu#5sM;YTc4j#z2{mSTV<9I)ib zI%3J>(c%l~_lzO{nhRz5G3zX%<6=7}n&q1pY-L`#4a8WdoP1c%7}PSP6AWKq8Qph{ z;(%W$TxKcJ`LMbDpkRLr&u7>tmQR=pxKOajV&u0@xCk;e9BtYa?j8|yY;4-q5jcps z^OzZ{*gY_8=kLh z7nsqznGq|e&J1u}9&$^zT5)?;af(`+gKx<9v(;MJGc{T!c0{?=^yD+emb23AqP~X_ zcfMA`4Yz|nGxLnZ+)3EYOrvO)*QmgRFW$+Cx$|K+la@`?D{2u%*FSpnjijrfG4_s< z3*_PtHI;8R`1}!GclwTZlx|@TT!9ke6U9u9=E%~#C)X_8JtRiUR_Yn`&iFS+R+Tl5 zMX?(rR{6aq<`-N;UZZ#}Ccb|+)8e^`nqRR^*todEnYqAxD;mSzu53 z$2FFI@Aa|iWsa;xk%oXys}_2Xya&HcSZq+==x4WIToVy#bob(^e0A2%;(=CU_KUp| z!A5Uuo|q--UzWzLe~i>6n5$9_-hUhx;d{Rk`H@fL+WT4mkHf-O-)}m8R5#O+Bvst_ zYJN+yCE-S&rB->_QiM33yXk;eU45Tjy-Z2tj-lQihJG~$brM9qOyT++Gd+C<(g}tp z{Nar<1*JQkdU^~i=nHm#hF@+7vGj4g#nBFUo-(Ql*J`y(T3w_ zk?$`KIG=i4iMFa7jeM-WU+%C^8pZd{wMJj+Slbk6-{~uSmml#*hU{md^1@=gx5na# zaNfg*Rf7(GW&B00I{fbh6TBZ2@f3bjki8d|3LcOCK7aFDdVm>cL=s&Z#z)ZJFk2*`e|Hn%M4IjT7vp z_X{N#d#)=5yJzcWyJps7`yqqaQ?J6m__q+IZ7MzRBN?ed<4j9mBb-eX`%Xo5h0g{y zm}UmJRM!ak}c&}tYIovXFGgsIuJIiRPlJH@mqTxi?zyi~6;< z%EN7Hu3=|UCGDRSY9HuGOMd;L-{Za=Y@(N2E+`P=Jg~IVU(>4f>Pr*`Jyx^un(piN z^m>x0&pm{JvUXqK6!=F|`-!VxFZD2aR3W5q>Lx>{t)<%nTYLRna)oEw$_39#dhA5k z`hN`%k*E#25iX~^T<37Y=F8}?f{yg)7jF8le0J9xi^a)fB<~7=!2l2bKlpxu&3=JB z@c(3s3HZeX(xTDcqR~Nr!E%9#pIcs?nHj8!CI^XTs7jvVG{D$K*6O&W@U4bUDP41m z#AhJs{?Csdrrmp?w4N1w1v-oc`4|=OBfz6B!B@N-Q8ob%O;HNe<7uZig_Xc&!&Iq&r{z-B6wWn~SYp%y((@YCh^?6W&^RWwqS;T8!E337%g|-xuvo zx*16w?xxG> zmxyNmely?n|zD~WT``OQCbF^z_r}A(kQaNI2xlC|%&$DWX zgPmyNw{d_iSq%wL6nU?o_$Y9DJr-xP>FAFi6`+_)J2kvg^;=c2B$Z2P^=d5FOEQ#R z70l<&N|!;I<1SV;Ucb<%=XCAsr0n$uLx+2#n>fNL{`TpaG}6}BaGB-UsrmrRioMLh z&$gO|c~%%-Zk}hId=C#VuswU$L%4(t6Q^w0T&ptI=?U3e%^Sh@sv)IH85jx{b^X^DOPbV z2iidHMdQ0ld2UaeR13a0Kh^eSPaB|1mexY5peHG!pP!X96aLwXRO#!Xd@slB;9srM z=s|{PQ^|Pq6mjJQxfkpW?p2e2?XRAuw>!%sFsGdE6oKzR5 zzhTS$r|rn}Cp+#b_r()WHK`~PASZ>Pr)|XPJzG_Na-VprP4$cwdYaXi`*Zx$42|QN z01xyVTUFq?uPNw0-8$*+ws46%LQ94F%@b~3@~7$)&z_$=$#mwl`+54AOWdtD0T8*W z7R9sFlPAxdIeq;+z3(ON(^@JM@lUuflB+&FdGg$u(`M)Chc0nPsZcz7a`GhKnbQ~l zV=l`j^1Oh;=VT{oDCm7pa>t!k;lA+X$yKt+$hX>yCnpLX=HM;LSSZerpJu0^zjTrt zaYiNX+>@VrWY<5CJy)WjM+e;#CU?M{%uu}Wgb4^kRxU*$+QI_dw6Bl}rIVFID1P>y zRuQ;NR>V*4KzTC5=n`2`1X+0?#m`%3RPfpKPmFJn-RC5C$UB)ad;W>B4%vMUil2`3 zPZG7s?n@oFjHAiQ9ZzO30vi~m`D=SgR&GP_GxGFNn4&1MazTK1TBYM6Sy3ukxjV(r z_A@F4s$}j)PeVfC&ssr zgYyG(C|S7z#n0ErfrQe?%1u6*ZOj4>z?zoa;qqVB8e}8E6h8;fsElfp(IPAhaH&Q^ zZVxp-lDwHt?eq*g{T_qttU35zB2i53NNk7KF}7?Q(Kr}l$6q&Yt?2x`fn9YW@9`-e z%-m(4|90cjHu)V+A-rg7p#^({%A(Pewo44OD2ni^>R_f7j# zC%=v!BrC{KM57BR;`r(!#bUk}u!x?nis=xtz#s!rj+W(|Bg(Rzm~f z&mI*ok*LA%N*J^%o|;?NqP|cF1>ULO7LAJ*KM7w>p_dtpG&+_2YVsTZX|wxP6yHDZ znpWhL))>j3&wC(N=vFh!x6)6_*RLV5b>WP8UcPl{dfT*Myo`J9(w}U;P)2#(-t)m} z^XI~xq1v<3JTf{iC*Zebi0LxtU)^(M;Pi1VZ63cVyB;!pX|?!_8EutKR1U&P`UX{% zjrK)^fH#Eb{c!A)>xY%kbEf0NnpoO3^02mXR`CEQ{$(MXJg9KcgMyjli!PddRbjWV ztkOYF3tL4zzp6B}u&lZV1AlN2vNW-*XM&uDwu;X6R?VHwbd7%AASHM(K#zNX{^T0{ zwn2)2tLRRKE6tk*sr$#RXr>^i@`VTJTvwWe1}XQ0f!2iwygA3N=MM(-aSwPuy3&L< zNKqUNoCYJg$0Ife0}CMRZ7eIz!N9ejRbgsaRSXRJHYpW=; zw~9&^`(B{^CFvFL1^GaZsT}1K(I9p1VBmT*YDMVV%M6j~jR$Ij*!Q0_CZ{0Dmt1t8 zzmh#qZ;K2oL`wzo6kA_&qtT6Y`gGqci@#ZlLaO-A+iBK#&cQ#F)f)M&QkIdI1|uY^ z!>+bV85$K&w!250COLUNuhvj|=Y;Gv%gX7JVx0nC-BOMlX1qV=S-*gXzFzZ;)(4Ms zWi{tz!m-BR8jQ26(a7nO63i@)zvdNvBj1UkyhcNCK+3%WL<6rMuP?rHZkd&)so46a zceL48reP7lPLEO2<1{n1Lus8pXxu%A-m z#rN}eq6Vym88id*r_T0&clHyt{phJK5&F>$i)9%*qERrv?F!6;AzqZb_m*HpKq3K7IIJ2N%QY5ia$-y8-E;1h1 zPF#UOq+DdsEECn%+#cv}^I=cF)0~KAqvC@>@-d>Pk=FV%VXBQZC(@5imeerFMHd!d zjHrIK^@KZ`doD~hoQ;acg$0Qb{p^O$_)Zhb3WGQv&jffG-)KVP*r;S-5E~bk8yL~J zYU?;R^tY)n)gLsW&)KLbU0511qOH}(QJ%6d;R<{NIgb$>Q}RSwYBzq*y9&$-pRTC9dluiquIK} zsC-?9URPUp%!dhuvQE@-EDKn`3vBI1wF{PKm=u@`J>n=m>0A(($D1xoh7Icu- z9aCXN&)F)^yA1Io(8irK4`h!c6?~(4!00mMh(ISc9ouFxrA8BBMN}~77oY_uMIFO! zAgu)^!h~YkD&7AgG}Al~hB*&oN;yD0ievtObrChp86fln1PYk56bL&VCIm)KX+}I|=dC*eZcw&n|cYsa+yqr z)jwg|0dfg|(N>xOAXgKnl-UiP)=3ipk?oCa(zV} z+fvzffLw!sM;zM@kShs<0B}GqNl{qMRh{ z5fbby5)c#?C^RrK&H5}_l~a2Erge1W4@b%u{flSLwx4-Za7KXP0@J|x#H8~~Cg+(L z&tDllmzZ$Q-soH+!@0!Y^v2hC1dF)^MYs!1l}$4POK)vkMNfS%9QnS{`+cMRJN1iD zsYR-@9aI8FG>MEfOaqX)1PGHMWR4M%IC$3h8na*lQ>PG9CwR>M^*k%pY^z6^R(C#F zC5V6D;Q!95{*2M_q|@8eb4I6g7*AgrJk^|ZD#!2?@zgb1!F-xdJ{rM1NT(1)u;47a zz*+W1qo)C{DCm?c5ZV-d0u&}^DR{lFddp7SmzrP)!k{iEQOuc8*fUV%^phthk{cIM z2wtNoG(BNzNRezwZhD^#2tyL!Cv_AMwS|ZsMFat1LV>})N!A~cs@0PF! zen%qpE#hkdB0s|+;{ZM(3D0PPXJo`*9y~}$IIuH1NMJZf_`Pp*ZBw9lLqKGsz;xOq zv$5pX3(IH_nB^TI%Y?fP8@C!*Uj#BP?m50)pEX*~W?a8K_@^o9PqyJ7 z!l`Q$0{PeuKCD3Ac!$uqK*1QBz*x$l!0ryyDo2h;wOs(tD4;++Ky2o0t-WiJZOL~` ztUaQQGGKq(A3E#*sNX-k$=}FnoBGx^zwsg?1%UxhA z^6b{)@A>{~Jsm~e1;(i+S<)rJ{jy&hqn{n6 z-@R?$%d@_h`+XTB*CU&A;U zs?A?G*0MHv`&D>9%Jyzc@Lqr6&FAG!B5z-g%wKkoT)ySF?5@4s&cEDDy}a7`iu%Rq z150LAEBB72-hw6LS4;1o)TFme4%}M0*zNSZy=LY|&0$SV)R&s_{F>g6HThCrl$X62 z1r{Tz7dMu0w>ogaeEpv4O`d)go_X1x)d`+WFFd^#7aZUInl=7qB=E7|j+#l)^WyuI z=Bni8?gX{1$Ts#jD$=q1WhaiOfn6NHL}gDGtIK)s=tU=|KP8#yIB2Ze%py*tplLj5VR14T?ioZV@Y`c|`8kaUeU7GariGoB#r%1JK&t+jq+$wZi(q%L za!0DRPf{_`!mc5ZPunJ`gpOOPuh2C2*R?vFn*|w$GsBhJ1>i zq#_SzJIU@6_jg1roNWt)okxUHg|jipV7^TvLX*)nHngx-Sxm;aNh)YKnbhPfa494)sWDd}a(H)6n2F>%dFgrIS zUp!hDfbq{F41i!p69|JVV2&FykK*qaQ+UdTEaq+dWDdwD0oaq#y4k>g5@Em}o)Rg8 zd5fORp#Z1|WM1mup*e(sBJlM<=23;G@FS3X;b`3^Fa_9Uj}3WIXk9PBr*Sez{4WVW z@ZNuV9D%k?=9tmK`sWb_O5rKgvd7$6|9VjV#RmMo07rmfNDEWOjrPue%nASA@YQ;cNI8oN{7WHE)!lXAz(Ww|4vDd?^Rb{CB) zgalCT5iq2qyTJPE{Ft-{%3W|r4g=-vfpTy+Bs3n~mCWvPWfma;lrxw>NC4#?xgp;H zW=CH@xpi4gRQseoQ0@&dB%`}N0sl#a1W@j!3?>RaX%CdUfIz-W{W~;=kO0cP@j$)< z%JCtP&~S8DBbWm0KsjJ|7lrNu<)+g(X%Cb;A^`}R{-?(gXxpScP_AztApw*-BYVsp zD0hSblsm!({K|kMzyQjrc_4@XQSRMybQcTo{F@2D5^SD4I52r{2TD%#={pHoOls%k zoDnUoX@82$$r6tOmjfAds>^I8bs&b|4%uR0H8a$(c1y&H>?$NC1MMeFRjY=dv)|Y=W%qHZ3{hx2^U2~#tztOK=df2T z<~7Me#nJT*x+QAJv>GSntO*_1Qd^+(aGz;a)!=r>HLWJK8T)GNM+kb=ekW5$Fm@(Q zW~65jRpMgbR9UY4z>PbNGnuSO5r(f$3%va~?D8;!`I`@6mz`>Dh7=VsLf&=FD}7@e zO&dF_TAabsS9rX7@LdPsK@qrhwe77XczC_=+u4;F7ulDAE@&{acMJw==9MIGl9}hpTv?efF*-}j#onz+pDX_0^-;nYPWeq7 zQvaI$R*;rE`#lp6`*)xrYiEU-F6VGrs=BHcGi@oh$`~Te@4z^-ptyp-IMblG{J=Q# zpg8lu+@zYR3TzWA-praSpY0iF;4MPGBeIrB%>Hl5zp4L2!ekq^`cj9rmZ-4iN_Dun zx4@D&xn7wr&df-$TDTxdY|{=5xr8t1t1}!GokL-F;4F8o^L2R+3cRhel59t{$-; zJn6r?$9hZYP!n5~>GB;E#bL~BUszid^@6&!@O1j(a|O0-8&i3;hG(3e2Ya{NM2K_Y z2c9N1&sI8nV>O~=MD*7*%rvbKPs@!C9zf^Je>{od{E{t=*h7bn#4~viSyqaOiMrN) z2UNr)?vq#!$ki!l9%C2YETC|y-VEApI+m#%pu|i+W}cw076D&va-s1`8xiBcO6Z_} zNw%4lh{mRvTmoss`&!Z<;6i;GX1j)S@vQx+7tv(QUCv|LtVYiC7YeT0b%L2TG(8!aWA2)Gr5>uTQrZ z3@a~JquSmNPZ{_(kDgZ>8Qza~jVCX?8SA77F++=bRuh2;Z@*+7fVm(U@AK(9{k`B< z-QZ{0>XG-Z7PULW{5oglSfk7C^NF~tyb4CsGc)%61rET;5GGH`2_Gc+r15{AX~dql99^sDliqDk*kKpz0b#fC&E9?LcR-5 z&NFyYW`_{t#}Xqw0K=@B06JVPKo_?Q4k{ZHR^^XUKpgn}Hh1o;Xlbc3w^V7Z05h#T z!k}4@;dF;xvMy(JS(A~f7F|&$o{BsI%`_s-R|J|#L>eUonps4eS$HXC%X~w+u|+ol z%LB;aV%Vt~B6&igX2GfdTk>z}|Bx6TLNVKG2L`-~eSOiIcCAi z%#j7T3d6Jo@xrY3zt!QVhcda~l!f?3{6!DICS@E?n9wnK8G4+mCBnXwmr4qsn!L;) zUhIC@|#MwIR$*(ZYi+jn3}v;u-pvV)*8tG%*z{cpd=ku^OPO_Jx54WY?(uNA~Fe^N8MlV>>s$}UVZo71gUq7w>nDli-xr*Jd{ zDk-?Gjd!MW{Ghg{CB^mu!bIBQ5M`)WWCEHIGPG%4hd4O&W$Vw^L z$?P41jNcVBUUz1gx_q*x;pM$#b>*cS0!t@Uby<;-stO$~7%yH4VbDBiYq|qIS@&&q znc{GnVs#nQa2Zo|ncZ-iU2z%9WW4$e)Ia0mUmZXFdFs9Jze)bL#PDNgJWh=WoA%8uKX}!L5u)vl@g0dl~V{z=%nmt zM$>t!c>gU)Ob!`tERha#FpN&wd_L}waawi9ge+_M>3n`XnG`pcV7ob(l;rD?vXSMn z#)uec9~q3#=|nyUs){q3m;zj2b|bDpwm5iQo@Gy&Ae1~UM1L(;$(A6*fjUB20et(% z*#((oJWk2BY;yeyFK32(5JXB1hG|n)Y*#USYt4U9C1Lv3K32)LZeY>DRgZu;m^{@? zPd8ABj=RXFz)&1bwmOV+J%@fxc_Tf5hlw;b^1yy#5 zzQlOWF;Uxi()^%gEC>gl7u{hvvS+cTTvMJzWJonrCk<4RVXm+#Kx#vktrnmJOg{^9 zOtX+7fEyq=I`TQt)i_uo#4u#5+fD&pzAy)CK%i0Ey>FPFyeJ*za0!!}XYRb3Tl3pw zU0-!o6=vYcM*Jk=V1D2hYv;L{?%UxqKXui7Bh`F$)g&X;Bz4t#Bh`8p)nrpCI&Kj#|0NM=u);sXJ_*dHy6o~=hYy|MSPF!~iR}6gAkA{3m#yIJ+uWD@z zqgTLu|$^<2BtQU*t(IyS~L>R=eq&$gQm|`}+=BR`u1?Y4Kres*<)DP5|7tEdE8_@N<^asN4OZ01Vt{k|JHRDhoSoGtkfP1)|IvX# z-9UxTqQx%aC$gWnn<%`5+52~3{7wn^faaiNi(6$y4RxVrTzm?&1FfG`focDw%+)Vc z2c;*HuS?35g>f!}V%fjR)&<@yDPz8)Yy6S@JK)Pt&8}$ihI~p)1$L}%(*3+_)mJSj zCnbdRYZq01WqhL~E};HndY0T}6vYLJ`IQ3;R!IUt7Oa1e1#1XhR%ER}dwQqsedp`M zg0E__r96pToN9*uGZvW5Be?RFIKQ@nvr|I`Xey{^t-`M`G8l#evQhzLp#+d6QhOS9 z*2%6pF>Yn1_H_0F?=7e_)r?d(PKnnE$SQ+ZkaHT=S4EZ;a%>=c_^CXT>=CXSBRiNgVyIAUHW4ghqpwYY$Z13(+`udX37 zMWc!G4oAuahmVGy%XI|ju_z)_vPeGpn|{Ru=*wO+rZE%f+a}<3o9TWy7^YrOwka@a zsRHBx8K56{id9N<`VoBjQry#Z;Ow*z`riJ z_pTi5(m>|ixvQrD75G%-RM3lDoVH&{VFJD<8Jo{1hfwF9HvBOmTa6eEKqn!Ps`_LC zP)#0y@4U@TGkc#_trQP1(=>pYzVebBfM|%$$O6oC{*{?>1I+Z*9AF~=Gvx-DDFDf< zCV&nXU@9zeCa=u&)Bn57lpA2C=C923=#`nK{fn7ay)sjPW5E=r{fn8tvcu9efSFdk zGE=8lW{OqP58;t~p7qWx7?K1KAqQY-btx!+$$G~C6Gb_n&w&Y0m|mcF*HvAB;Q#_U z^W)bos$dxc#;jdH{m1og9-=)LV5Yy*{{PH$T(*17EQWcsctdC$pntSLzZ!z3flj0X zI`Oku#uU(Bi%R3NB95?=uIP`na(S!R_BcSUvJhcT1%s!VX_7dzwR@}q=kFh|=Nx4_ z*37}l;=mw_d>sMCa$~Yrf&eL%fv5tkQGv|xNCDc*4WvvZ#$;J#T!E1TFUdX4R$u}q z7$%hDnP%w>1t-UgpD9m<(WaY0Gj=J_p@z5=m`bC{vC8zX=UYKUxKKw3W6L46jRQKy zCO}7xCWoY(rB>!e@(VXs0YS2|FicB3-zi?p+Sf=m6FYtyaj+n8i>1@fOxJa|tVvx} zkhv%mTcrR`r2tzc4^JfzTcrd~r35S9XkKwpJ+uDGXMP4c8CY^#@@8@cx%vP8H_88! z_?1LFHR1mc-_ZYO@oloLSgl-;Br%PiZ$$J}>Q}OwGF`CQwtO{aL6V>&`tKl@5SVam zxp5uabt)Kg8J=#t*F`lr7@RHcayA8lsKSHge^*qgqPh;izZTDQg+VZwl+xu)`ArUW zI@AHwX3~xj+8YdL0YF6XuCVFHF92N>v=B)Yx|D1wlY#aU!I0}NKV8q~Ko;S_5-Y$; zKj)K3l6L^i4>j14QwM2}^i85&%pj!6a4f#YpI%qWro_fGit%893Z-SmrF%x!49`28}lly8H;52#GuH-gud%&5VhTxh2nGSeD?S%MJju6S>7e}=b{qH*H# z9NKli``Pi+0dMl$iUBw_ejtY^x#mqrmZ`eC*x)@a!bl!DoHFV_R`y`30F6W9FF8`e z8mQVFC9wqFVESDMlAmt~-wB&#Li-B5b5M;j@j-qVxGUh$Go~Lzg7dT8J@Mh7%Nok9 z`SG##L$03u=&dHIpGX~6QTCCs#8mdzF0OKCALVU@WT?chm=8a8E(oFSm0Mchdt$Gw zEOYp75Ph)`J6S4DgFQZ%+)Cu#w5^{9UROxLy~+9D^0Jo63p#LcOB*T;Z%yBtouul} zaG0Qn#0BjX;2m9RROb+6NO2F`-MR2zLuf8*^J74F_&HM2Bb9^KpPIq?wn2V;M4_LZ zV5?AerAVkb^tbCw zbO1Ixs6-ApD9u;8f_V1=7HcSY7Hh-FB7!4t2qC(Z?B4r&NLvX-jrLoWLO!qsA|hVv{Yz*XmNz z2f3!LW?r%7^@+Qo_q<57D&`kGSiX+wZoWDpoVPPF9K=i+Xt%v9%rMuT+wme-L()~B z2vQb$rel~V817Fs*ag!vXvq6evS_e(?eLA?cYtP%Ki;c7S%h+xe6!|H`(s}Ot(+fL z3HiL}RU_z?gbeOmQ z{gVl2s^o4ABqa-IBK((NPE^#m1w0&U-PJ6NE{78Q@3bt(BNblI@eY_1jMdid+3?p`uzTAts1yhe$LZm-(cW!E;lO6w z&?xO zojbD!wx?u9Vr4IRTxySVY|+Dzf*Lp%-DAO@fXD1G(&m^*)3{QDsn#M&pRq8Yo)pW! zDA5NZB^z-rhGwe_xJNrk+L{M`>4&w-wI>fm3!#`>a;CRpeCJAS!DvH~4ozjncSf{~ zJnWABO<`W)HhA5OnnEEl%JoMK63E>IFXk?Y4Jurp^60GTi*0Ry)Gnjg1+4LRq!Ard zq+tw1(zQBB0>p5XHa+@pQ>X8Co@>ZHJd}^z1QMP|YqNJ{Y#caES2jw8f?K1=T}3M> zcnWGQ&nPbA>8_QRnKrp{cv@=1sVI))S&w3un80a8-?wik9R_3V{H*Z4A(PM4uX~uC zZeRM_4Qf>}?(VGB@H*>S)`#jpBoWej+uDp&(!E(z&LaGdV1KAk-Gt@NLw7<(Bq0^k zslXrus+ldp`FK1N_XXnf!xeb?AaZU4c%&Slg_vm=#%aeE9$`G%Up-n9>TSHu<$G{j zxc3#YMK&|Uk2l=uPiADUW=+6_X#(ecraG%=hH8rF?_cJ_?_7h5=TR>qzB1pp*O-NK zt<50GiY@s#+wojv%bbE%3a~1?1g8TY`@NSEZDUIyqoKl<>ByEZlbqZz-a|^H_@42X6ni8S}a|1^=nd~?--P&~{C)#933fP?U zAB1VYT_wDBM?D2U6Z!0)@n0pq%}_VzK?%0e3iesW^YubK+yfH}5Tp|2Df(8d*EaKgVB3OH1Wei`|bv6zVp8~rj0y~}Af8|!vZF3F&Z z^u3tgIOydkstt9&r$LwbPOzp>$$3f$Igmga)+l=jtQ-jKJ_Q6EbNG%Wg~n3_*O{nc zxP->g`PTfNBT);DLCO-=>kjGUPYsK{6xw>mcvu~$kJyBs@$2AC4U3ramn*|Y+|JS) zBqbWTW?5b{5TnjjE!USvM*2>)F$Ud1LU14T+oKZ?^o2dO$iPB~0nyjOL~qvyNn9Cf z+DEvH2?>*fdR(AWtE6u2r%4ds8_{^!oHxu$I0*dW}U4Wh9hI!kDG-NV)&ON!8^EdG#u@y5l_bN&+}g_BzcUSGwvc2M#Y|&m4>f7_3X7D zoolpkD3Wj39}fDb4^8du-GbzP&vZj~&CM$oOYHb?yt@+c5&D>h7!`H<@eA>)?Qzpm zQRo8k?0f1qSEi#QK64) zS&9YhZn_;F=fvK0s#IM}rJv%0Kq#N5gj!r=I3^BxPI_t&q;{UwW2PkTHPppw?5Ft9 zwK2K9cDxrISV24*0cj^eB#t@^i4Hjh`vaPcAQKwZ-2wUfEQ<1?{cB0h`hG?z8kd%bA$kn?jDTI!{S~`FLy9TUJbZcAayG{4(3gwyy7@ZA*4we&xsC3V5RGTCNv6>o znY)`|yT(N9QE5#`PtD8u1-taLZH1qB2`BEm;P_sitGTR=6;DbdeWZ!ouKA~K@23Mj zCw8^vjE7Mn8~TOs*chcmQpah`4<;)jbl#el*9no=$P+w=49^Evp3Ms?>hRq!*BC_> z8w_OIu39)sH;fp|bVkJore29~{4aM)!=;72^U55KE9?j0X$#vWUITCha3W}y;oTAa z^yByb(ZiC!>t;RZhZM< zWZwcjeAK!-*Z@l+6nR+SVBa;&gTW}wDpLM zgElaz-lO3i{}BUr68G6~=(hH?8n#92x2m+IPbVTN zD>DpD#m`zA&xXuNbZpj<&M$7aYnP+@+UnkS7tSXG&kql5b?bI_?cMsWy@eeIjh4(2 z&M3wFO{HR|7mdC5#LrGIPM@h3zF(_|S8kwlyI4M6tr&s6k51+TJGtk*AJ zIg?+mmKwK-U+!;J>6@Sh;u|+_ckC<)o)0T*8r(A~zdyI6SWtjkx1Y<*F^m7Cud{-x zo@*aY_!&m7!SQrpF=LSij+(=S6FKKRiXl6Y`7OWn&J(`12F=gm^+xamNbmR%^prrn zyiBWNNWUE3weqby*!28)o8yFUh3)y9>fy*?2kT?LYt8$!rvkVB!{<5e`|GC~%ePzl)18rpD7{j7_};dAN82qY6F0TB zo(V0Suq#cMS5H+J6Qc(iIIXQqwEVos?u~zNrJxFd<9pBQs^~IMZoEBIhtiM97q;9S zFMz{2W`Ofgd8CO33*9R=r@`RHX+5O~)uUtLZx@;^cSg)=!vpa+6pwcmO-%Lfj<{>2 zGoW>~y2Wy@m2l@=v^WHI;0t2S+Ljul8=LQ}oKEN8US2>E#Q0B+`~9t@e8ndNxz%?Y zMm0}O&90B0^>jps&Vzk{dOg^#;1}$a_}FRw=KE2-$yek(CifcaVwZ+VnXu^B5lTM7wGgeu@((wu`GANQ5G2TD5XFXA z@xi}HHwp!Blht4D2BFeSwhJ^@F~mD7Vdr>z;;6UWwuVFLPBy!yuJ z3uPFx*Y!(RQ``-579=DR%oySaHfG!uSN(S;MyN>dOoZs(ev#5G%M~!8LaYHECI`01 ztACL?(vc#B97f7Ed{c$!YxHV5j6P;TFpPRXjClRs7DSeb)GdnpmYD-Sh!Nh%1~Y`) z#|OsWB;JQ@(K^Cs8)lpbau_B1EtER0uoRq$DdO;$Q=6Xl@AsmazVMnW1Nvj}Fi7Cv zM^(<3}G& zs##LK;QZ-is#Ej8+7p7{`uNGYy#$Z3J>kK<$IC&3vAx)R%>5v=T|6W%?rP&8e;9fe z#4q3Js|W*}G$ij- z5p3y$EyHJLFZ=Ju;|36guJ+L8iVlX9LERmXY-)dAJz!=W(%1jsFX>?p{s9P@k-Gud?y?FUg2SD9z>LiF=VVNGHn38-C_`QnIp)nd?;CUduFZ=G;|J@q*bbMPCV;h$QdR2w#d8 zor$j4hRmlhOz$}sM|H~SswKP9r?zKx5>wX|3rXIg0tz7MFeIwJ=t)d#ha zj|s<$EgRqgynC<3^#vDw6e&*mX_OOPkKgza+<`}BwhTB5zB;D%b16Mt7iOg{LSNYGRzX?M!k>QW!KYIp%c8&R>t~kSk z&yY0uno!hIoZ>-eFiri(iQI7bHS!vtc`4A-p!9v2vTfNTP)aMm?_>&KXSmfRJe3>i|1~3oq&6J*xXp( zTRY9z2#jF=R1zaRI@m^n^>P^3`tUI4h2&Wjrc*Rjr*F95{I0y5$xVx%MFrzt)mNv8 zd?R8yXF2o(n+gUm0}g&6m1j|rPEj5pty&HxPksXqv7%>I4=S%9@U9xFqZwq`Zu%TO z&V(Vtj)PzP{ocgTZ@hufxytJO67E9PP;qw3U>NqH)$Kwg@LTAFeeKUNBJ38d#WPE%+@>dVIdbFt zC$!h&9(uFKW53o|LSrU__dQjrmn7^gnj7m}A-FfcxuP`2{DOS$>Ad05jgVP=_jfRd z9jws+;cm%dJ@nFV!o=nu$V0suzKrU-3|Y_=8>0fabeB7a0>ZUf6ibO&c)Aw7Z$BG_ zk%W$ih=zyZ~N6O)rV02(|Qm%4R&c0cu zBQCT<`|g5#ZJJ1_L8$eTd>pI$OFeYwR}N<#+-hhXhHPl-e-)ewZOQbE7%=|DdjxsP zb!0)V;j<8P$DarUy&+sdY<*m2Hkz=bI-*$prd|7GJ54>tbM)8IoYG>p zF8w<^Jror`#X>*xM+f%w!u9k7(JMIEC&?MmX#u`umf~FNe4=M|z*AV{+sbXAAhI#k4QF_CP zm`~kfqbQ4WuBH)j}q0f^pt}_2Sk^(8;PeMPEB6D~9OC$vt zWkgmn;~)lVbAEtNyx6;1*p*1Obb^^bb^v#veT<<4QF*Fv&ShL|0f%mab zZ>T417+$||Ky@F)=pNN0zYM8l0pl5lZpk?1dMmooOQh`7#jJU$L6GhFpQK9y7 z?2oyp#jam=3WDWc%%t2ej6A0f*O#$(i{V_%_@!N4KZLl)FYK~hZJ!S<`0kq$ZMWaP z>KKf!&_c<>lF}0fhY*~fO)sH6yKz#zD!>`3t961;T*GL7t$n`wj3I7z!Oaodi`{;0 zEz&vs;KH-da)v*wVN?)oA1Q8#das<<^@*LixGT)(Q;a_j#m3}J1R85hJ=c?W?tR*; zFt+fz*BBm^DdNeD|J`fo8iGU(5ZUE?iGpp zcoDA28gg$he<-jQ6Z-ZHov6p^inzz`-omYhRn?m`gs*cY;_$wnxTlHp>CG>LoXbn8 zg7j#!cK+|81tJP^F!b@!t9pV>K4F{TUZRrI#*-HYpH^s-u+!L#3}#2uqh*Cb-l8MD zkpjjB0~c5IyND~Bn0ncY)?!XihZk4E#<#{f%#Zftc+dai<>Zd`vqR6#pKpwua}G}T zWc+Y^tMDUGlu>Y*s>5VFro5dti&`!mLvfrap!!>z#g@ahcl{To$aN<6+SfLTJKBUc zB;)XngA{1eJ{V~7HBf1X~WpI=PiC3wF325S845&#%J|22%}d)Yt5 zu{<6mia5OcCiYK`o1?8<9U>_Y6(Ifd$IfW2s(J=@j=WKRg1LW$*MLy_uW*~eXIzGU zemQ}cC=`d0XgmAJb8h1do$t&iwC2Vz0@e)^GiY^U`18lw9gbZ3C3@UiO`E;zKhX>J zJX>0sntw$2?%|1rKF=R6j{AA@)r+!*G^jFPpBH9m(*8)u#7*0$IW^j^+?z9Ccf`}c zPGOsj*5t2+V&HwEV+DHZZZ9F&O|;*t8gE{ruER+q904 zK%wF>P6kZ)N-ku)6hFof!t>gaa7b*}x0}S@)6ia%NmK2*Mh}IS5Ye;GFKlf;-}D=e z<6?Ux4&P6BMl|>{fU?Wl1bj$llvp?4?3mWKw0bwFd)+ghEany7*g+B(=7N82|ClK; zbzww95$%Q+<}tZM7j}bFc7r2renZ{-CRz(BPOBnXtAe_Dlawem1Vp?3LFrtFlqf6& z_Wha87Ll|8LD≫lAFD`_{C8_hkw#uN%+ajs59v z5rf*G@7X`PYg$_3(D^WEq>L-*Z(n6Svn;56f*NJDY&|1~QtjleVWAS*xWA>RY#X&6 z!BVYV5}JZ7p;xjEpk#L*=M?N>O7{+EI2-ruNwcw)&Wh+52MW|YS-2K+1ogu*8MF4> zGkl34CTv~L+n`X!b*<>+%dsjHZ(*BiP-Ip_yP(rhyuYXAe8PkdK_-VTYpe9H(P)(p zrknsno2acW+J`enz41so`B-n9-BT^cvbJ`bPjmOR=zCQSzMP9Dhp<-0vW72%Z=C{( zK*^jXlilKtUm{D(Q+_qkOeIyj$fdcNTm`eOXSQrPe88w7=si_nbGs=O|I&FL^Z9hz zz0$1lkLDyqkC&bHe3MtJDsbXZ(wOz#Nymjl%i;B%Y5g2<9MQo0lD3`M%nj=>dL9Ad z=LdoN{rbHnw`+x2ex6G2hoR>s=DS{c(zglFB_pA6YoZ%e_*EI%f@@UxN##}P*}j*U z;nLPuuv)k-=>~oUd?KNo#PVw%lf2vhB~J|8J38oLw<^4Eh`HSyPACRPPrwLYPk2ptE+d!Op$dYQI20vNRFp@T z#7g?wX@2*Vsaup{7s9$oTYX$tJp>JC^fvb2-g~uZw3f*2iVBGN4~l(p@6D^~wY*#8 zfI~S7Ofw@uSpc2CY-dkyZN`qBu6s@6+9YX?=eo{>S*6plCgVim%Q!Y84LdTYZyX00 z+$(M{3!fk6m!s`e=8D=6DcX$h+1cG&D&{k4ZIC6`^FR`-yFLx6=X*BlXu*C>>7T&3;YMah5D%G253M8%iO z+uNLxe@}F$;k9s+TiH>@aD6&=l=OI-RM1oAovkJoZ;Hw==>_h_^hgW%$pDdRq^ls>gA2Z%q=eBEZf5LU4 z5InP2crwAxhD{*J*p+d{;eiV}lk>}RWU175Hkn-7lJzs)NHFGrV_SJ{OtzrNdz_U? zio0%gy!5`iDPNi(%8IDDSMXHjw+j7`#7l2IK#bDr>@aZLk^2}!!ZAp_}t-3`s z!%HY~ zlV=@x$({4ecaM&i8`W^b@t7(bBDT*3QU(-}Q={MvD~5>~kchsMv-fm!y)a|{;O-q` ztFXRy9}7aL6brVNxIRO9o|feIVUhWeAz~X}okN5@7x(BIQtoH;_Bx!2vu9FE5vQg` z0mkP*2M@P@6qO`RQ4ymI6DkJ1AcQwDEm08}8kxm&2&M;Q$eMS+hjeQ(G~$;SKHROS z(A`A_f95t%hQoKBE+ftierQ~vi^$G*X$zOBVe&Y4XJ2r$x<%&as~4Qp_83ryq&Ouh za7<|`HrAlpn__=m8qS9ty14u5XbmXLYvvU9MYIYtGHg*2(!ZAefbix*F=Ap4sQlqA zno;`&vCt~<7B1ANGh)4VDe_BjENWl3br5nrc@u5}J01hBON*C}@iFNBn@&Kelknt? zh&EQbrg|X9d$mBy5sU?Zb$AMcX)c)!wlk zb=S7DVtgsq-28%L%}y>Kcg@0J>(nga*Si6H7OevRmW2CN#pjHj+!m971h@QB?jPrx z9N)bQN3NTvj2Skc)wK2xLPn=&;t$j&tSo97^d%aX0;+!G@;22J}bst~V_}7E?ug3}vx!dt{wMDT_<%lrmF& ziCj-9u74t7mK4^f+q0)>78$s!osnk&kB7}{*V0?c0{ z%~l7!-aB4idnh(L8(uysAr$R~d~InLd~H~ArS}vCF}pkA7k}>_FCT8=7wART3yiPk zJ;!5tjpO`n;a>LTqDH%NFip~J@3P7_v1Z*3f}P`o*Q2DaSh4TYp$4&tS$Pqqx%SYB z?ww3w3@qmoJp%a5+yt)B48J`SR2zafQ(2O1?hjy9GG(H~R#G}))s%aohznyaQxj|(hZ;Gw@797Kq zEMZx^RC0Ko#!!t!kE|~ADIVcJ`uJly)!~U(p^WPjRhyxj8slZlMn{ z;Vb0S=s2wY3TvLW5E zh#mx`+@0Ij&^w*+11sF_uXa~gwboM3Jyh*a=bfc1u|6{g5MB?~FZYhYI1Z~V58K*a z&39)EWZ6m~r;Yq4V^vU(d$Z`siog7CcC5Ys1SSBE_s0+Da``9ETMJ8@N@fQrD@WTE zN`4P#4E%>3(7DxTlXAV)S#NKhZe)&Do3Yd9kU-e@huM>Y4|gfV&Id!EB$@9RPL{D2 zuB=) zKUjf!_jouRmweuAj=QnB>*&z%#>RU;viTmo!TlU|W$uBa^vY`H*z~1Ai~kSNCL7jZ=s1SsBf}xOQrnBnSh_SAt8}y$z|_M<9r;M z@j>MnqWeg{P@*TTzWn=!sFh4Jydl>Y`(z^Jw=+oAR)-t;lYHs2cpt(Xc^o%4Om%MM z`O_Qht{-T;vvx{@8Wqx8WknM4GPYUQ-|yWOuH`-&Ztj#yW~6#0)ATG|$)lBY-1W41 zJn!4X;T#ow@2-p1e5;Z0i_e?TW0&8Xigx~UdaH)#f2ux0cU|L*$2X#{_O9cXCK&(n z+yM8UqEG7MGVrF4e)n2^v$=`f!!yR38d&}v1E~DdX=3M*e<^R#MMHtBD&2`U=SAE-iT3%#U zZH7I&tVfu^@$S@f7=*wXGmf#$>cilpTB?t4MAQork_60cvu87_@uRv+lC*L+(iZfN zFGnn=)okynVtT;NR-vv`-gp~}GfNcU3_GjEc?y4eVo3__6@jn?od2kSOc&ZWo-p`tZM@j2(iV=HRs zq+UspcGK!vT=~)-%(tzt)xx8Tmzx--C!L9jVbYxwjf7Y&86r4Bdp1zfXqtBPQep=M zf+ywlyG6UUs%X~ys71e$HSmpnmB4fumAeBGZAYv6o-Igy1?J2xDH^Blvl(q=R>e>q z48AM0V6iXb7ihs{fvL%Du^3ExA{&>uFGv!HXoEULYg9K3# zX@F-0*uOxhHTh~4`+5byxVb|6t3!9bLqa@<$umQe7BkV2G+W)zVrh}R=V#U8O|{#H ztVR*k!l*dkuiA6OFkRly>V+>cx;f}gnFbGUdaah*U+n=~TzvL&hn<|1+zz_sU6hxe z&QeN1$CW9h5$7-c)0U z#j#QC(TS1g(dnwJ<@?C9_M7Y9DyGfzeCDdR^xd<9c?b<>|A8wB0kO#P`sq zy*y#Lk<+bGP^-vrwyDhwt+ZM#+c+Patx`81o0%t~Eqf6kTLoW)oIwmCuDMg?yoMFF zxcS|hW0egXinKAaPSMe8YOr4-l*1omtKKlzj!$+=gjYa!mQT~#GqIsCYB zz-62{+{?TUGlqFQb>r%D;|jlV;Cjt2UtMnDwNP0@-n4FauJ^BY(#1d?#9DI373m`z zylE&vO&m&QCih-tJK5A4s<&Yk-Pb1^Jkqia!r10w>*qTB*u9RlwI zx93QN)E`ct(T~r2tuY-ps8e9w=dWmx@ZxW%Mv=Fi1c$E`E==o;`CF6tw9XG+$gu3M zwQ~c0spVfB8mtcgtWksvR}`GywWrHa<;d9fU>9{`FA5 zI!2NyxKa12{4f=rKQEvzUyt~}Q-ve+uP6YWKQrfbpv*kP+Modqsy{dP$&00bJuk$w z$eCVZ-AdJymOHKsh0GOb_Q)&QPpGd(+pe2lwyn#Lwznu(d*!+WNIKq=jA3YRb*r)& zvnj(AF-bz8VzbM#DT@>_WibPlLOMdTPbC%JC}0Xwg(zc+6kv*whmb0Tkm86$9|(Vt z#F3nX625yMd(QsjoSkyZItuB>%#VLAQMat+2UbW1Rx)E-1%W>*Z+}W}f5m4Qg=d)D zN2KgWB+0w)Rq}B&`B+=lg|V&0l-}QSAAy6Z7pS+aX;QG&FTMam#pjyrN6WbL+W%xU zu%bfc9UA{*0^TCK({cqDwcK(%E}k&>!jsRRfWpPl|kqw!7U z2ne)HcLq^1uJDxMtK5ZLcv->Kc`0KcAZxIj(&v-fH?{y}fpJgH zL_)#zZ({anNY3aiYf@-CQlS-Lm`&lJmSDub;6?V3KEWzaj5$(@ildZ1Y|d!Q72%q` zU<>y~hy~dHAb%jQ5Ro-lV2`G=BVA|-wulDmfEuO0dQV;;B3?fteTC7RkH-)bO`A3S1laxLp z&S=jS;p@KOLw8t6ZutKo2dX3{yV*i7a4i6|0T3n-AaY+Y!9BU1fb8ZNdo&O;s3n-- zwYjiAP$;Yre38v+!DHd@UDvKvZgEwEzO7oc$RKzAHhFgr@8LywH)Zk?$le|P^ln6p zLJl9|C@eO=J&VJ-rx7$`pmyO!{} zG@}RINU@AUF6ZAJ37S~nT{|Dl^%gc>-G7){a;V^Q7%H2N+Bu*3fH}FffcV|?Hi&+O z(IR(k!RfB1Q$=y5aIZ&>1!E-wZ ziFI*7_NL@6S|}ka$bI&xMFQRd&#dld>A2G1k9@o5%@{1ZN#1HB@D3uzzn^1B2L4X) zBaNWN0E02vcGq2}^v=vQzl*9#PM`NxN;c_^+&8+fisd7oL71c4k#C_#U-fpF>l{N( z*|c36>l!jJ<$FUU4&)+B8(WrHVHi*O~DFY$Vn5^pO^v{L{Mz*ub@Bd}`=_py2iN9W^owE;z z^Oq(PghC@eU}>nNx`=;3!Ka-F_uuX^A%^%Nm-@P1DYUfB|6it9;IdU?SEfKx3{KHy z6DaZd>8_tuz&?jFHSD47E4yke!M{7|$s`n0tGpT16`!^IT@&d-3mdt7KrdPpr#%E- zRI+3JDH~YMX#tZAy53v6&=`=eI-_}9nk2YGH=)z(lQOomGU7}^){>oC6X`IqnF`M4 z3M2<|g|oz!y|Xtc!Kgd9`?^2bh-#3dhFjOq@KHEno3^#jT*5)2&{`LzE=@goUTx;j zaX#Nxq$6FzUrkc}K$hNp Date: Sun, 20 Sep 2026 20:55:43 +0200 Subject: [PATCH 103/206] fix(operator): a controller nothing owns is a disk, not an absence MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Five workers of a six-worker fleet were discovered as having no disks at all, and the cluster activated on the one machine whose add had finished. The disks were there: 0000:01:00.0 class=0x010802 driver= Micron 1.7T 0000:02:00.0 class=0x010802 driver=nvme boot disk 0000:0b:00.0 class=0x010802 driver= Micron 1.7T Claiming an NVMe controller unbinds it from the kernel before binding it to a userspace driver, so a run that fails in between leaves it owned by nothing: no namespaces, no block device, nothing under /dev. The day's failed adds left two per worker that way. claimableControllers offered a controller only when it was bound to a userspace driver and idle, so that state fell through every path at once. It presents no block device, so the device reading does not have it; it is not bound to userspace, so the reclaim path skips it. The disk is absent from the draft entirely and stays absent until somebody binds it by hand, which is not a fix — the next failed add puts it back. So the question becomes whether anything owns the controller rather than whether something owns it in a particular way. A controller with no driver qualifies outright: asking who holds it is asking about a character device no driver has created. A userspace binding still has to be idle, and unchecked still is not idle, because vfio-pci is equally how a disk is passed through to a guest. The enumeration was never the problem and is unchanged: atlas-lib's PCI scan finds controllers by class, whoever is driving them, and the probe has been reporting all three of worker-0's throughout. What was missing was a consumer willing to act on the third state the Controller type already documents. Co-Authored-By: Claude Opus 5 (1M context) --- operator/internal/discovery/plan.go | 31 +++++- .../discovery/unbound_controller_test.go | 98 +++++++++++++++++++ operator/internal/nodeprobe/report.go | 9 ++ 3 files changed, 134 insertions(+), 4 deletions(-) create mode 100644 operator/internal/discovery/unbound_controller_test.go diff --git a/operator/internal/discovery/plan.go b/operator/internal/discovery/plan.go index eef253495..50fbb702a 100644 --- a/operator/internal/discovery/plan.go +++ b/operator/internal/discovery/plan.go @@ -322,10 +322,7 @@ func claimableControllers(report nodeprobe.Report, class DeviceClass) []nodeprob var out []nodeprobe.Device for _, controller := range report.NVMeControllers { - if !controller.BoundToUserspace() || !controller.Free() { - // Free rather than "not held": a controller the probe could not - // check is not one this may offer. Reading the unchecked state as - // free is how a disk something is driving reaches a draft. + if !claimable(controller) { continue } if _, already := presented[controller.Address]; already { @@ -345,6 +342,32 @@ func claimableControllers(report nodeprobe.Report, class DeviceClass) []nodeprob return out } +// claimable reports whether a controller the kernel presents no block device for +// is one a run may offer anyway. +// +// Two of the four states qualify, for different reasons. +// +// A controller bound to a userspace driver qualifies only when nothing is +// driving it. Free rather than "not held": a controller the probe could not +// check is not one this may offer, and reading the unchecked state as free is +// how a disk something is driving reaches a draft. The holder need not be this +// product — vfio-pci is also how a disk is passed through to a guest. +// +// A controller bound to nothing at all qualifies outright, and asking who holds +// it would be asking about a character device that does not exist: a driver is +// what exposes one. This is the state a failed run leaves behind, because +// claiming an NVMe controller unbinds it from the kernel first, and an add that +// fails after that leaves it owned by nobody — no namespaces, no block device, +// and nothing in the device reading either. Five workers of a six-worker fleet +// were reported as having no disks at all that way, and the cluster came up on +// the one machine whose add had finished (2026-09-20). +func claimable(controller nodeprobe.Controller) bool { + if !controller.HasDriver() { + return true + } + return controller.BoundToUserspace() && controller.Free() +} + // BasicDeviceRules is the default device pipeline for a filter. // // The order is deliberate and is the order a reader wants the refusal in: what diff --git a/operator/internal/discovery/unbound_controller_test.go b/operator/internal/discovery/unbound_controller_test.go new file mode 100644 index 000000000..6ba8cb1af --- /dev/null +++ b/operator/internal/discovery/unbound_controller_test.go @@ -0,0 +1,98 @@ +// A controller nothing owns is still a disk. +// +// A controller has four states and only two of them present a block device: +// the kernel's driver has it, a userspace driver has it, nothing has it, or the +// probe could not say. The third is what a failed run leaves behind — binding an +// NVMe controller to a userspace driver takes its namespaces from the kernel, +// and a run that unbinds and then fails leaves it owned by nothing at all. + +package discovery + +import ( + "testing" + + "github.com/simplyblock/atlas/ptr" + + "github.com/simplyblock/simplyblock-operator/internal/nodeprobe" +) + +// TestAControllerNothingOwnsIsOffered covers the disks a re-run has to find +// again. +// +// Regression: 2026-09-20-unbound-controllers-vanish-from-discovery — +// claimableControllers offered a controller only when it was bound to a +// userspace driver and idle, so one bound to nothing was dropped: it presents no +// block device, so the device reading does not have it either, and the disk was +// absent from the draft entirely. Five of six workers of the lab fleet were +// reported as having nothing after the day's failed adds unbound their Microns, +// and the cluster activated on the one machine whose add had completed. +// +// Nothing can be driving a controller with no driver, which is what makes it +// offerable without asking who holds it: there is no character device to hold. +func TestAControllerNothingOwnsIsOffered(t *testing.T) { + worker := report("worker-0") + worker.NVMeControllers = []nodeprobe.Controller{ + // The state a failed add leaves: the capture of worker-0 has two of + // these and one kernel-driven boot disk. + {Address: "0000:01:00.0", NUMANode: -1}, + {Address: "0000:0b:00.0", NUMANode: -1}, + {Address: "0000:02:00.0", Driver: "nvme", NUMANode: -1}, + } + + offered := claimableControllers(worker, ClassNVMe) + + named := map[string]bool{} + for _, device := range offered { + named[device.PCIAddress] = true + } + for _, address := range []string{"0000:01:00.0", "0000:0b:00.0"} { + if !named[address] { + t.Errorf("the controller at %s is not offered, so a disk nothing owns is invisible", address) + } + } + if named["0000:02:00.0"] { + t.Error("the kernel-driven controller was offered a second time; the device reading already has it") + } +} + +// A userspace binding is the other reclaimable state, and it still turns on +// whether anything is driving it. +func TestAUserspaceBindingIsOfferedOnlyWhenIdle(t *testing.T) { + idle := report("worker-1") + idle.NVMeControllers = []nodeprobe.Controller{ + {Address: "0000:00:02.0", Driver: "uio_pci_generic", InUse: ptr.To(false), NUMANode: -1}, + } + if len(claimableControllers(idle, ClassNVMe)) != 1 { + t.Error("an idle userspace binding is not offered") + } + + busy := report("worker-2") + busy.NVMeControllers = []nodeprobe.Controller{ + {Address: "0000:00:04.0", Driver: "vfio-pci", InUse: ptr.To(true), NUMANode: -1}, + } + if len(claimableControllers(busy, ClassNVMe)) != 0 { + t.Error("a controller something is driving was offered; it may be a guest's disk") + } + + unchecked := report("worker-3") + unchecked.NVMeControllers = []nodeprobe.Controller{ + {Address: "0000:00:05.0", Driver: "vfio-pci", NUMANode: -1}, + } + if len(claimableControllers(unchecked, ClassNVMe)) != 0 { + t.Error("a controller the probe could not check was offered; unchecked is not idle") + } +} + +// A controller the kernel is presenting is left to the device reading, which +// knows its size, its content and whether anything is mounted on it. +func TestAPresentedControllerIsNotOfferedTwice(t *testing.T) { + worker := report("worker-4") + worker.NVMeControllers = []nodeprobe.Controller{{Address: "0000:01:00.0", NUMANode: -1}} + worker.Devices = []nodeprobe.Device{{ + Name: "nvme0n1", PCIAddress: "0000:01:00.0", Kind: "Disk", Transport: "NVMe", Available: true, + }} + + if got := claimableControllers(worker, ClassNVMe); len(got) != 0 { + t.Errorf("offered %d controller(s) the kernel already presents as block devices", len(got)) + } +} diff --git a/operator/internal/nodeprobe/report.go b/operator/internal/nodeprobe/report.go index eeedf56fe..73a26478a 100644 --- a/operator/internal/nodeprobe/report.go +++ b/operator/internal/nodeprobe/report.go @@ -361,6 +361,15 @@ func (c Controller) Free() bool { return c.InUse != nil && !*c.InUse } // It reads the driver and says nothing about who bound it or whether anything // is still driving it. InUse answers the second, and nothing answers the first, // because a binding carries no record of what made it. +// HasDriver reports whether anything at all is bound to the controller. +// +// The state it distinguishes is the one a failed claim leaves: taking an NVMe +// controller for a userspace driver unbinds it from the kernel first, so a run +// that stops in between leaves a controller with no driver, no namespaces and no +// block device. It is neither the kernel's nor a userspace driver's, and asking +// who holds it is asking about a character device no driver has created. +func (c Controller) HasDriver() bool { return c.Driver != "" } + func (c Controller) BoundToUserspace() bool { return c.Driver == pci.DriverUIOGeneric || c.Driver == pci.DriverVFIO } From 0a749ca64c7354565abb5d598d036da31fd95d68 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Sun, 20 Sep 2026 22:40:14 +0200 Subject: [PATCH 104/206] fix(operator): the node-add cap holds when the nodes cannot see each other MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The cap was an election: every node waiting for a slot ordered the waiting workers and advanced if its own place was within the number free. Two nodes reading one set reach one answer, and that is the whole of the assumption. Reconciles are served from an informer cache, and a cache filled moments ago — after a restart, a lease change, a resync — does not hold every sibling yet. Nodes reading different sets reach different answers, and each of them is alone in the set it can see. A node add reboots its host, so six workers electing themselves is six hosts rebooting at once against a cap of one. A slot is now taken rather than deduced. StorageCluster.status.provisioningSlots is the Provisioning phase's metadata: one entry per worker whose add is outstanding, naming the worker, the object that took it, and when. Taking one is an optimistic-locked patch of that single list, so exactly one node wins a given resourceVersion and the rest are refused and count again. The list is on the cluster rather than a field per node because six objects carry six resourceVersions, and six claims are six separate agreements. What is in flight is read from that list rather than counted off the siblings' step names. The inference is what wedged the lab cluster: a node stuck in Resolving held the cap and nothing could say so. A holder releases its own entry and no other, on each of the three ends of an add — the UUID arriving, the deadline expiring, the object being deleted — and an entry whose holder is gone, already has a UUID, or has failed is reaped by the next node to ask, so a cap cannot be left closed by an object that will never reconcile again. The name ordering stays as what it can honestly be: a tie-break so a node told to wait is not overtaken by one told to wait beside it, computed net of the workers already holding a slot. Adoption is now checked at the queue too. It was checked at CheckingHost and at CheckingConfig and nowhere after, and an object passes each of them once, so one that reached the queue before its backend node existed waited for a slot it had no use for and held the cap while it did. Two workers on the lab cluster had a running SPDK pod and a backend node the operator had added, and their objects sat at AwaitingSlot with no UUID. The graph gains the AwaitingSlot to Adopting edge, and the check runs before the worker's availability for the reason CheckingHost checks it first: an adopted node is already running. Every case was red first. The one that names the defect ran three reconcilers over one store, each able to list only itself, and reported three nodes taking a single free slot. Two further defects came out of the same suite rather than out of review: the tie-break ranked against workers that already held a slot, so a cap of two admitted one; and keying "do I hold one" on the object rather than the worker let a two-socket host take two slots and post two adds whenever more than one was free. Test plan: U-374 restated, U-391 through U-401 added. Co-Authored-By: Claude Opus 5 (1M context) --- ...torage.simplyblock.io_storageclusters.yaml | 46 ++- operator/api/v1alpha2/storagecluster_types.go | 36 ++ .../api/v1alpha2/zz_generated.deepcopy.go | 23 ++ ...torage.simplyblock.io_storageclusters.yaml | 46 ++- operator/dist/install.yaml | 47 ++- .../crd-redesign/design-storagecluster.md | 9 + .../crd-redesign/design-storagenode.md | 40 ++- operator/docs/tests/test-plan-storagenode.md | 61 ++-- .../controllers/node/adopt_at_slot_test.go | 116 +++++++ operator/internal/controllers/node/graphs.go | 5 +- .../internal/controllers/node/graphs_test.go | 14 +- .../controllers/node/provisioningslots.go | 182 ++++++++++ .../node/provisioningslots_test.go | 320 ++++++++++++++++++ .../controllers/node/slot_race_test.go | 72 ++-- .../node/storagenode_controller.go | 132 ++++++-- ...torage.simplyblock.io_storageclusters.yaml | 46 ++- 16 files changed, 1102 insertions(+), 93 deletions(-) create mode 100644 operator/internal/controllers/node/adopt_at_slot_test.go create mode 100644 operator/internal/controllers/node/provisioningslots.go create mode 100644 operator/internal/controllers/node/provisioningslots_test.go diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusters.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusters.yaml index 49335795d..dc8ff3afb 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusters.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storageclusters.yaml @@ -1368,7 +1368,6 @@ spec: - message: enableAtomic4kWrites requires enableChecksumValidation to be true rule: '!(has(self.enableAtomic4kWrites) && self.enableAtomic4kWrites) || (has(self.enableChecksumValidation) && self.enableChecksumValidation)' status: - description: StorageClusterStatus is the observed state of one backend cluster. properties: activeOpsRef: description: |- @@ -1430,6 +1429,51 @@ spec: - Unavailable - Suspended type: string + provisioningSlots: + description: |- + ProvisioningSlots are the workers whose node add is outstanding. The list + is the metadata of the Provisioning phase, and it is also the mutex that + caps concurrent adds at spec.storageNodes.maxParallelNodeAdds. + + It is one list on one object because that is what makes taking a slot + atomic. A node takes one with an optimistic-locked patch of this field, so + exactly one node wins a given resourceVersion and every other is told to + count again. A slot recorded per node could not do that: six objects carry + six resourceVersions, and two nodes reading a cold cache would both see the + same one free and both take it. + + A worker rather than an object is what holds a slot, because one POST adds + every socket of a worker and a two-socket host must consume one slot. + items: + description: |- + StorageClusterStatus is the observed state of one backend cluster. + ProvisioningSlot is one worker's hold on the cluster's node-add concurrency, + held from the moment its add is posted until the node has a backend UUID or + has given up. + properties: + node: + description: |- + Node is the StorageNode object that took the slot. A release removes only + the entry naming its own object, which is what keeps a node from freeing + somebody else's slot, and a slot whose object is gone is reaped. + type: string + takenAt: + description: |- + TakenAt is when the slot was taken, so a hold that outlives its node's + deadlines is visible in the object rather than only in the events. + format: date-time + type: string + worker: + description: |- + Worker is the Kubernetes node the add was posted for, and it is what the + cap counts. + type: string + required: + - node + - worker + type: object + maxItems: 64 + type: array realignedGeneration: description: |- RealignedGeneration is the generation the last successfully requested diff --git a/operator/api/v1alpha2/storagecluster_types.go b/operator/api/v1alpha2/storagecluster_types.go index 4e0414e68..398baa09d 100644 --- a/operator/api/v1alpha2/storagecluster_types.go +++ b/operator/api/v1alpha2/storagecluster_types.go @@ -773,6 +773,25 @@ type StorageClusterSpec struct { } // StorageClusterStatus is the observed state of one backend cluster. +// ProvisioningSlot is one worker's hold on the cluster's node-add concurrency, +// held from the moment its add is posted until the node has a backend UUID or +// has given up. +type ProvisioningSlot struct { + // Worker is the Kubernetes node the add was posted for, and it is what the + // cap counts. + Worker string `json:"worker"` + + // Node is the StorageNode object that took the slot. A release removes only + // the entry naming its own object, which is what keeps a node from freeing + // somebody else's slot, and a slot whose object is gone is reaped. + Node string `json:"node"` + + // TakenAt is when the slot was taken, so a hold that outlives its node's + // deadlines is visible in the object rather than only in the events. + // +optional + TakenAt metav1.Time `json:"takenAt,omitempty"` +} + type StorageClusterStatus struct { // Phase is the operator's own view of this cluster. // +optional @@ -858,6 +877,23 @@ type StorageClusterStatus struct { // +optional Tasks []ClusterTask `json:"tasks,omitempty"` + // ProvisioningSlots are the workers whose node add is outstanding. The list + // is the metadata of the Provisioning phase, and it is also the mutex that + // caps concurrent adds at spec.storageNodes.maxParallelNodeAdds. + // + // It is one list on one object because that is what makes taking a slot + // atomic. A node takes one with an optimistic-locked patch of this field, so + // exactly one node wins a given resourceVersion and every other is told to + // count again. A slot recorded per node could not do that: six objects carry + // six resourceVersions, and two nodes reading a cold cache would both see the + // same one free and both take it. + // + // A worker rather than an object is what holds a slot, because one POST adds + // every socket of a worker and a two-socket host must consume one slot. + // +kubebuilder:validation:MaxItems=64 + // +optional + ProvisioningSlots []ProvisioningSlot `json:"provisioningSlots,omitempty"` + // ActiveOpsRef names the StorageClusterOps currently allowed to operate on // this cluster. Empty when none is running. // +optional diff --git a/operator/api/v1alpha2/zz_generated.deepcopy.go b/operator/api/v1alpha2/zz_generated.deepcopy.go index b27c07325..de0cf0e1a 100644 --- a/operator/api/v1alpha2/zz_generated.deepcopy.go +++ b/operator/api/v1alpha2/zz_generated.deepcopy.go @@ -1425,6 +1425,22 @@ func (in *PoolLimitsStatus) DeepCopy() *PoolLimitsStatus { return out } +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *ProvisioningSlot) DeepCopyInto(out *ProvisioningSlot) { + *out = *in + in.TakenAt.DeepCopyInto(&out.TakenAt) +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new ProvisioningSlot. +func (in *ProvisioningSlot) DeepCopy() *ProvisioningSlot { + if in == nil { + return nil + } + out := new(ProvisioningSlot) + in.DeepCopyInto(out) + return out +} + // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *RebalancingMetrics) DeepCopyInto(out *RebalancingMetrics) { *out = *in @@ -2377,6 +2393,13 @@ func (in *StorageClusterStatus) DeepCopyInto(out *StorageClusterStatus) { *out = make([]ClusterTask, len(*in)) copy(*out, *in) } + if in.ProvisioningSlots != nil { + in, out := &in.ProvisioningSlots, &out.ProvisioningSlots + *out = make([]ProvisioningSlot, len(*in)) + for i := range *in { + (*in)[i].DeepCopyInto(&(*out)[i]) + } + } if in.RebalancingMetrics != nil { in, out := &in.RebalancingMetrics, &out.RebalancingMetrics *out = new(RebalancingMetrics) diff --git a/operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml b/operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml index 49335795d..dc8ff3afb 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_storageclusters.yaml @@ -1368,7 +1368,6 @@ spec: - message: enableAtomic4kWrites requires enableChecksumValidation to be true rule: '!(has(self.enableAtomic4kWrites) && self.enableAtomic4kWrites) || (has(self.enableChecksumValidation) && self.enableChecksumValidation)' status: - description: StorageClusterStatus is the observed state of one backend cluster. properties: activeOpsRef: description: |- @@ -1430,6 +1429,51 @@ spec: - Unavailable - Suspended type: string + provisioningSlots: + description: |- + ProvisioningSlots are the workers whose node add is outstanding. The list + is the metadata of the Provisioning phase, and it is also the mutex that + caps concurrent adds at spec.storageNodes.maxParallelNodeAdds. + + It is one list on one object because that is what makes taking a slot + atomic. A node takes one with an optimistic-locked patch of this field, so + exactly one node wins a given resourceVersion and every other is told to + count again. A slot recorded per node could not do that: six objects carry + six resourceVersions, and two nodes reading a cold cache would both see the + same one free and both take it. + + A worker rather than an object is what holds a slot, because one POST adds + every socket of a worker and a two-socket host must consume one slot. + items: + description: |- + StorageClusterStatus is the observed state of one backend cluster. + ProvisioningSlot is one worker's hold on the cluster's node-add concurrency, + held from the moment its add is posted until the node has a backend UUID or + has given up. + properties: + node: + description: |- + Node is the StorageNode object that took the slot. A release removes only + the entry naming its own object, which is what keeps a node from freeing + somebody else's slot, and a slot whose object is gone is reaped. + type: string + takenAt: + description: |- + TakenAt is when the slot was taken, so a hold that outlives its node's + deadlines is visible in the object rather than only in the events. + format: date-time + type: string + worker: + description: |- + Worker is the Kubernetes node the add was posted for, and it is what the + cap counts. + type: string + required: + - node + - worker + type: object + maxItems: 64 + type: array realignedGeneration: description: |- RealignedGeneration is the generation the last successfully requested diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index 7a3f6443a..13cdcceba 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -6375,8 +6375,6 @@ spec: rule: '!(has(self.enableAtomic4kWrites) && self.enableAtomic4kWrites) || (has(self.enableChecksumValidation) && self.enableChecksumValidation)' status: - description: StorageClusterStatus is the observed state of one backend - cluster. properties: activeOpsRef: description: |- @@ -6438,6 +6436,51 @@ spec: - Unavailable - Suspended type: string + provisioningSlots: + description: |- + ProvisioningSlots are the workers whose node add is outstanding. The list + is the metadata of the Provisioning phase, and it is also the mutex that + caps concurrent adds at spec.storageNodes.maxParallelNodeAdds. + + It is one list on one object because that is what makes taking a slot + atomic. A node takes one with an optimistic-locked patch of this field, so + exactly one node wins a given resourceVersion and every other is told to + count again. A slot recorded per node could not do that: six objects carry + six resourceVersions, and two nodes reading a cold cache would both see the + same one free and both take it. + + A worker rather than an object is what holds a slot, because one POST adds + every socket of a worker and a two-socket host must consume one slot. + items: + description: |- + StorageClusterStatus is the observed state of one backend cluster. + ProvisioningSlot is one worker's hold on the cluster's node-add concurrency, + held from the moment its add is posted until the node has a backend UUID or + has given up. + properties: + node: + description: |- + Node is the StorageNode object that took the slot. A release removes only + the entry naming its own object, which is what keeps a node from freeing + somebody else's slot, and a slot whose object is gone is reaped. + type: string + takenAt: + description: |- + TakenAt is when the slot was taken, so a hold that outlives its node's + deadlines is visible in the object rather than only in the events. + format: date-time + type: string + worker: + description: |- + Worker is the Kubernetes node the add was posted for, and it is what the + cap counts. + type: string + required: + - node + - worker + type: object + maxItems: 64 + type: array realignedGeneration: description: |- RealignedGeneration is the generation the last successfully requested diff --git a/operator/docs/designs/crd-redesign/design-storagecluster.md b/operator/docs/designs/crd-redesign/design-storagecluster.md index 39eeeec25..c24877a11 100644 --- a/operator/docs/designs/crd-redesign/design-storagecluster.md +++ b/operator/docs/designs/crd-redesign/design-storagecluster.md @@ -2150,6 +2150,15 @@ type StorageClusterStatus struct { // +optional Tasks []ClusterTask `json:"tasks,omitempty"` + // ProvisioningSlots are the workers whose node add is outstanding. The list + // is the metadata of the Provisioning phase, and it is also the mutex that + // caps concurrent adds at spec.storageNodes.maxParallelNodeAdds: taking a + // slot is an optimistic-locked patch of this one field, so exactly one node + // wins a given resourceVersion (design-storagenode.md §4.2). + // +kubebuilder:validation:MaxItems=64 + // +optional + ProvisioningSlots []ProvisioningSlot `json:"provisioningSlots,omitempty"` + // ActiveOpsRef names the StorageClusterOps currently allowed to operate on // this cluster. Empty when none is running. // +optional diff --git a/operator/docs/designs/crd-redesign/design-storagenode.md b/operator/docs/designs/crd-redesign/design-storagenode.md index 34f4f2ec9..9b4a5f48f 100644 --- a/operator/docs/designs/crd-redesign/design-storagenode.md +++ b/operator/docs/designs/crd-redesign/design-storagenode.md @@ -810,14 +810,48 @@ posted must not be read that way. The steps before the claim wait where they are and `AwaitingSlot` declines to take a slot at all while its worker is away, which is what stops an add being posted against a machine that cannot answer it. +**Adoption is checked at every gate before the add, the queue included.** A +backend node the worker already has is what `CheckingHost` and `CheckingConfig` +divert on, and an object passes each of them once. One that reached the queue +before its backend node existed — a previous operator's add, a `POST` whose +response was lost after the control plane committed, a rebuilt object — has no +add to ask for, and checking again where it waits is what recognizes that. It is +checked before the worker is, for the reason `CheckingHost` checks it first: an +adopted node is already running, and the machine being reachable is not this +operator's precondition to establish. + **`AwaitingSlot` is where two independent serialization rules live.** `maxParallelNodeAdds` caps how many workers may be in flight at once, counted by distinct worker rather than by object so that a two-socket host consumes one slot. Workers hosting a FoundationDB pod are always sequential regardless of that cap, because a node add reboots the host and two simultaneous FDB reboots reduce the -control plane's own fault tolerance. Both are predicates over the current state of -the cluster's other nodes, so the step re-evaluates them on every pass and holds -rather than failing. +control plane's own fault tolerance. The step re-evaluates both on every pass and +holds rather than failing. + +**A slot is taken, and the holders are the Provisioning phase's metadata on the +cluster.** `StorageCluster.status.provisioningSlots` is the list of workers whose +add is outstanding, each entry naming the worker, the `StorageNode` that took it, +and when. Taking one is an optimistic-locked patch of that single list, which is +what makes the cap hold: exactly one node wins a given `resourceVersion` and every +other is refused and counts again. A cap derived instead from the siblings' steps +holds only while every node reads the same set, and reconciles are served from an +informer cache — a cache filled moments ago, after a restart or a lease change, +gives each node a different set and every one of them is alone in its own. + +The list is on the cluster rather than a field per node for the same reason. Six +node objects carry six `resourceVersion`s, so a claim written to each is six +separate agreements and no mutual exclusion between them. + +A holder releases its own entry and no other, on each of the three ends of an add: +the UUID arriving, the step outliving its deadline, and the object being deleted. +A slot whose holder is gone, has a UUID already, or has failed is reaped by the +next node to ask for one, so a cap cannot be left closed by an object that will +never reconcile again. + +The ordering of the waiting workers stays, as a tie-break rather than as the cap. +It settles who tries first among nodes that can see each other, so a node told to +wait is not overtaken by one told to wait beside it, and it is computed net of the +workers already holding a slot. **`CheckingConfig` is a gate rather than a validation.** A cluster with `enableFailureDomains` set requires every node to declare a fault group, and a diff --git a/operator/docs/tests/test-plan-storagenode.md b/operator/docs/tests/test-plan-storagenode.md index 1c5ed0bc7..775d6fc4c 100644 --- a/operator/docs/tests/test-plan-storagenode.md +++ b/operator/docs/tests/test-plan-storagenode.md @@ -69,30 +69,41 @@ function of those three inputs and of nothing else. ### Entity: Provisioning Gates (design §4.2) Files: `operator/internal/controllers/node/provisioning_test.go`, -`slot_race_test.go`, `resolve_retry_test.go` - -| # | Scenario | Type | Test | -|-------|--------------------------------------------------------------------------------------|----------|----------------------------------------------------| -| U-09 | `enableFailureDomains` set and no fault group declared: provisioning is held | Negative | `TestANodeWithNoFaultGroupIsHeldRatherThanRefused` | -| U-10 | `enableFailureDomains` set and a fault group present: provisioning proceeds | Positive | `TestANodeThatDeclaresItsFaultGroupPasses` | -| U-11 | `enableFailureDomains` unset: the fault group is not required | Negative | `TestANodeThatDeclaresItsFaultGroupPasses` | -| U-12 | A `failureDomain` label of `0`: a label like any other, not read as unset | Boundary | — | -| U-13 | Held provisioning emits `FailureDomainMissing` and issues no `POST` | Negative | `TestANodeWithNoFaultGroupIsHeldRatherThanRefused` | -| U-14 | The worker's storage-node API answers: the host check passes | Positive | — | -| U-15 | The worker's storage-node API is unreachable: held, no `POST` | Negative | — | -| U-16 | TLS is enabled and the CA is missing: the host check fails informatively | Negative | — | -| U-17 | The host check retries until the endpoint answers | Positive | — | -| U-18 | Nothing in flight and one slot free: exactly one of three waiting nodes takes it | Boundary | `TestOnlyOneNodeTakesAFreeSlot` | -| U-19 | Siblings past `Posting` without a UUID hold the slot | Positive | `TestANodeInFlightFillsTheCap` | -| U-20 | The node counts every sibling but itself | Boundary | — | -| U-21 | Two nodes on one worker count as one in-flight worker, not two | Boundary | — | -| U-22 | A worker already in flight does not block another worker under the limit | Positive | `TestACapOfTwoAdmitsTwo` | -| U-23 | `maxParallelNodeAdds` reached: the node holds at `AwaitingSlot` and issues no `POST` | Negative | `TestANodeInFlightFillsTheCap` | -| U-24 | Workers hosting a FoundationDB pod are identified | Positive | — | -| U-25 | A FoundationDB worker holds while another FoundationDB worker is in flight | Negative | — | -| U-26 | A FoundationDB worker holds even when `maxParallelNodeAdds` allows more | Boundary | — | -| U-27 | A non-FoundationDB worker is not held by a FoundationDB worker in flight | Negative | — | -| U-267 | The same waiting node wins the slot on every pass, so nobody overtakes it | Positive | `TestTheChoiceIsStable` | +`slot_race_test.go`, `provisioningslots_test.go`, `resolve_retry_test.go` + +| # | Scenario | Type | Test | +|-------|--------------------------------------------------------------------------------------|------------|-----------------------------------------------------| +| U-09 | `enableFailureDomains` set and no fault group declared: provisioning is held | Negative | `TestANodeWithNoFaultGroupIsHeldRatherThanRefused` | +| U-10 | `enableFailureDomains` set and a fault group present: provisioning proceeds | Positive | `TestANodeThatDeclaresItsFaultGroupPasses` | +| U-11 | `enableFailureDomains` unset: the fault group is not required | Negative | `TestANodeThatDeclaresItsFaultGroupPasses` | +| U-12 | A `failureDomain` label of `0`: a label like any other, not read as unset | Boundary | — | +| U-13 | Held provisioning emits `FailureDomainMissing` and issues no `POST` | Negative | `TestANodeWithNoFaultGroupIsHeldRatherThanRefused` | +| U-14 | The worker's storage-node API answers: the host check passes | Positive | — | +| U-15 | The worker's storage-node API is unreachable: held, no `POST` | Negative | — | +| U-16 | TLS is enabled and the CA is missing: the host check fails informatively | Negative | — | +| U-17 | The host check retries until the endpoint answers | Positive | — | +| U-18 | Nothing in flight and one slot free: exactly one of three waiting nodes takes it | Boundary | `TestOnlyOneNodeTakesAFreeSlot` | +| U-19 | Siblings past `Posting` without a UUID hold the slot | Positive | `TestANodeInFlightFillsTheCap` | +| U-20 | The node counts every sibling but itself | Boundary | — | +| U-21 | Two nodes on one worker count as one in-flight worker, not two | Boundary | — | +| U-22 | A worker already in flight does not block another worker under the limit | Positive | `TestACapOfTwoAdmitsTwo` | +| U-23 | `maxParallelNodeAdds` reached: the node holds at `AwaitingSlot` and issues no `POST` | Negative | `TestANodeInFlightFillsTheCap` | +| U-24 | Workers hosting a FoundationDB pod are identified | Positive | — | +| U-25 | A FoundationDB worker holds while another FoundationDB worker is in flight | Negative | — | +| U-26 | A FoundationDB worker holds even when `maxParallelNodeAdds` allows more | Boundary | — | +| U-27 | A non-FoundationDB worker is not held by a FoundationDB worker in flight | Negative | — | +| U-267 | The same waiting node wins the slot on every pass, so nobody overtakes it | Positive | `TestTheChoiceIsStable` | +| U-391 | Three nodes reconciling against a cache holding only themselves: one slot admits one | Regression | `TestTheCapHoldsWhenNodesCannotSeeEachOther` | +| U-392 | The same, with two slots free: two are admitted and the third is not | Regression | `TestACapOfTwoHoldsWhenNodesCannotSeeEachOther` | +| U-393 | A slot taken against a cluster read the holder has not seen is refused | Regression | `TestASlotTakenFromAStaleReadIsRefused` | +| U-394 | Taking a slot records the worker and the object that took it on the cluster | Positive | `TestTakingASlotIsRecordedOnTheCluster` | +| U-395 | Re-entering `AwaitingSlot` with a slot held keeps the one entry | Boundary | `TestASlotIsNotTakenTwice` | +| U-396 | A node releases the slot it holds and leaves another node's alone | Boundary | `TestASlotIsReleasedOnlyByItsHolder` | +| U-397 | A slot whose `StorageNode` no longer exists is reaped, so the cap reopens | Regression | `TestASlotWhoseNodeIsGoneIsReaped` | +| U-398 | A slot whose holder already has its UUID is reaped | Regression | `TestASlotWhoseNodeIsFinishedIsReaped` | +| U-399 | A worker's second socket takes no second slot and resolves against the first | Regression | `TestASecondSocketDoesNotTakeASecondSlot` | +| U-400 | A backend node appearing while a node queues for a slot is adopted, not re-added | Regression | `TestANodeWaitingForASlotAdoptsTheNodeThatAppeared` | +| U-401 | No backend node for the worker: the queue takes its slot as before | Negative | `TestANodeWithNoBackendNodeStillTakesItsSlot` | ### Entity: The Provisioning Claim (design §4.2) @@ -596,7 +607,7 @@ CRD carries. | U-371 | The migrate graph splits the restart from the wait | Positive | `TestTheMigrateGraphSplitsTheRestartFromTheWait` | | U-372 | The host maintenance graph is the six-step window | Positive | `TestTheHostMaintenanceGraphIsTheSixStepWindow` | | U-373 | The four single-step actions share one line | Positive | `TestTheSingleStepActionsShareOneLine` | -| U-374 | Adoption is reachable from both provisioning gates | Boundary | `TestAdoptionIsReachableFromBothGates` | +| U-374 | Adoption is reachable from every gate before the add, the slot queue included | Boundary | `TestAdoptionIsReachableFromEveryGateBeforeTheAdd` | | U-375 | `AwaitingSlot` may go straight to `Resolving` when a sibling claimed the worker | Boundary | `TestAwaitingSlotMayGoStraightToResolving` | | U-384 | A worker that is not Ready or is cordoned holds `Posting` at `AwaitingWorker` | Regression | `TestAWorkerThatWentAwayHoldsTheNodeRatherThanFailingIt` | | U-385 | The held step carries its own budget rather than the one it was diverted from | Regression | `TestTheHeldNodeGetsAFreshBudget` | diff --git a/operator/internal/controllers/node/adopt_at_slot_test.go b/operator/internal/controllers/node/adopt_at_slot_test.go new file mode 100644 index 000000000..a67cb23bc --- /dev/null +++ b/operator/internal/controllers/node/adopt_at_slot_test.go @@ -0,0 +1,116 @@ +// A node whose backend node appears while it is queuing for a node-add slot. +// +// Adoption is what stops the operator adding a node the control plane already +// has, and the two gates before the queue both check for one. The queue itself +// did not: a node that reached AwaitingSlot before its backend node existed — +// because a previous operator posted the add, because the response was lost after +// the control plane committed, or because the object was rebuilt — waited for a +// slot it had no use for, and the check that would have found the node was behind +// it rather than in front. + +package node + +import ( + "context" + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/client-go/tools/events" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/testsupport" +) + +// aQueueWithABackendNode is a node waiting for a slot on a worker the control +// plane already reports a node for. +func aQueueWithABackendNode(t *testing.T, readings []NodeReading) ( + *StorageNodeReconciler, *simplyblockv1alpha2.StorageNode, *simplyblockv1alpha2.StorageCluster, +) { + t.Helper() + scheme := testsupport.NewScheme(t, corev1.AddToScheme) + + node := waitingNode("w1") + cluster := slotCluster(1) + worker := &corev1.Node{ + ObjectMeta: metav1.ObjectMeta{Name: "w1"}, + Status: corev1.NodeStatus{ + Addresses: []corev1.NodeAddress{ + {Type: corev1.NodeInternalIP, Address: workerKubernetesIP}, + }, + NodeInfo: corev1.NodeSystemInfo{SystemUUID: workerSystemUUID}, + Conditions: []corev1.NodeCondition{ + {Type: corev1.NodeReady, Status: corev1.ConditionTrue}, + }, + }, + } + + apiClient := fake.NewClientBuilder().WithScheme(scheme). + WithObjects(worker, cluster, node). + WithStatusSubresource(&simplyblockv1alpha2.StorageCluster{}). + WithIndex(&simplyblockv1alpha2.StorageNode{}, clusterRefField, + func(o client.Object) []string { + return []string{o.(*simplyblockv1alpha2.StorageNode).Spec.ClusterRef} + }). + Build() + + return &StorageNodeReconciler{ + Client: apiClient, + Scheme: scheme, + Recorder: events.NewFakeRecorder(64), + API: backendNodes{readings: readings}, + }, node, cluster +} + +// TestANodeWaitingForASlotAdoptsTheNodeThatAppeared covers the gate the queue +// was missing. +// +// Regression: 2026-09-20-adoption-is-a-one-shot-gate-before-the-queue — adoption +// was checked at CheckingHost and CheckingConfig and nowhere after, so a node +// that passed both before its backend node existed never looked again. On the +// cluster this was found on, two workers had a running SPDK pod and a backend +// node the operator had added, and their objects sat at AwaitingSlot with no +// UUID: each held the queue for an add that had already happened, and the cap +// meant the rest of the fleet queued behind them. +func TestANodeWaitingForASlotAdoptsTheNodeThatAppeared(t *testing.T) { + r, node, cluster := aQueueWithABackendNode(t, []NodeReading{{ + UUID: "4a304439-2909-4199-ad2f-b8624d66a13d", + Status: nodeStatusOnline, + ManagementIP: workerKubernetesIP, + SystemUUID: workerSystemUUID, + RPCPort: 4422, + }}) + + next, done, err := r.awaitSlot(context.Background(), node, cluster) + if err != nil { + t.Fatalf("awaitSlot: %v", err) + } + if !done || next != stepAdopting { + t.Fatalf("a node whose backend node already exists was sent to %s", next) + } + + var stored simplyblockv1alpha2.StorageCluster + key := client.ObjectKeyFromObject(cluster) + if err := r.Get(context.Background(), key, &stored); err != nil { + t.Fatalf("read the cluster back: %v", err) + } + if slots := stored.Status.ProvisioningSlots; len(slots) != 0 { + t.Errorf("a node that had nothing to add took %d node-add slot(s)", len(slots)) + } +} + +// With no backend node for the worker the queue behaves as it always has, so the +// check is a divert rather than a new way to decline a slot. +func TestANodeWithNoBackendNodeStillTakesItsSlot(t *testing.T) { + r, node, cluster := aQueueWithABackendNode(t, nil) + + next, done, err := r.awaitSlot(context.Background(), node, cluster) + if err != nil { + t.Fatalf("awaitSlot: %v", err) + } + if !done || next != stepPosting { + t.Fatalf("a node with nothing to adopt was sent to %s", next) + } +} diff --git a/operator/internal/controllers/node/graphs.go b/operator/internal/controllers/node/graphs.go index 4618ebeca..db8a02bf7 100644 --- a/operator/internal/controllers/node/graphs.go +++ b/operator/internal/controllers/node/graphs.go @@ -281,7 +281,10 @@ func provisioningGraph() statemachine.Config[nodeStep] { OnEnter: deadline[nodeStep](checkingConfigDeadline), }, stepAwaitingSlot: { - To: []nodeStep{stepPosting, stepResolving}, + // Adopting is an exit because the backend node a queuing node + // would have added can appear while it queues, and a node that + // has one to take over must not ask for a second. + To: []nodeStep{stepPosting, stepResolving, stepAdopting}, OnEnter: deadline[nodeStep](awaitingSlotDeadline), }, stepPosting: { diff --git a/operator/internal/controllers/node/graphs_test.go b/operator/internal/controllers/node/graphs_test.go index b8c8bfebb..faf0e232e 100644 --- a/operator/internal/controllers/node/graphs_test.go +++ b/operator/internal/controllers/node/graphs_test.go @@ -287,12 +287,16 @@ func assertLine(t *testing.T, a simplyblockv1alpha2.StorageNodeOpsAction, want [ } // The provisioning machine branches rather than running in a line: adoption -// diverts from the host check and from the configuration gate alike, which is what -// lets an upgrade Secret and a backend node found at the worker's address reach the -// same step. -func TestAdoptionIsReachableFromBothGates(t *testing.T) { +// diverts from the host check, from the configuration gate, and from the queue for +// a node-add slot alike, which is what lets an upgrade Secret and a backend node +// found at the worker's address reach the same step from anywhere before the add. +// +// The queue is the one that matters after a restart. A backend node can appear +// while an object waits there, and without the edge the only answer to one that +// has is a second add. +func TestAdoptionIsReachableFromEveryGateBeforeTheAdd(t *testing.T) { ctx := context.Background() - for _, from := range []nodeStep{stepCheckingHost, stepCheckingConfig} { + for _, from := range []nodeStep{stepCheckingHost, stepCheckingConfig, stepAwaitingSlot} { config := provisioningGraph() machine, err := statemachine.NewFromSnapshot(ctx, config, statemachine.Snapshot[nodeStep]{State: from}) diff --git a/operator/internal/controllers/node/provisioningslots.go b/operator/internal/controllers/node/provisioningslots.go new file mode 100644 index 000000000..5403070c5 --- /dev/null +++ b/operator/internal/controllers/node/provisioningslots.go @@ -0,0 +1,182 @@ +// The node-add slot: taking one, releasing it, and reaping the ones nobody will. +// +// A node add reboots its host, so the cluster caps how many may be outstanding at +// once. The cap used to be an election — every waiting node ordered the waiting +// workers and advanced if its own place was within the number free — which needs +// every node to be reading the same set. Reconciles are served from an informer +// cache, and a cache filled moments ago is not that set: a node whose siblings +// have not arrived yet elects itself, and so does every other one. +// +// A slot is therefore taken rather than deduced, and the record of who holds one +// lives in status.provisioningSlots on the StorageCluster. One list on one object +// is what makes the taking atomic: the append is an optimistic-locked patch, so +// exactly one node wins a given resourceVersion and the rest are told to count +// again. The same list on the six node objects could not do it, because six +// resourceVersions are six separate agreements. +// +// The ordering that used to be the cap stays, in awaitSlot, as what it can +// honestly be: a tie-break that keeps the same worker winning across passes so a +// node told to wait is not overtaken. Correctness is the patch. + +package node + +import ( + "context" + "fmt" + "slices" + "strings" + "time" + + apierrors "k8s.io/apimachinery/pkg/api/errors" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/types" + "k8s.io/client-go/util/retry" + "sigs.k8s.io/controller-runtime/pkg/client" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// takeSlot records this node's hold on the cluster's node-add concurrency and +// reports whether it got one. +// +// The cluster passed in is the one this reconcile read, and its resourceVersion +// is what the patch is conditioned on. A cluster that moved since the read means +// somebody else took a slot in between, so the patch is refused and the caller +// counts again on the next pass rather than appending to a list it has not seen. +func (r *StorageNodeReconciler) takeSlot( + ctx context.Context, + node *simplyblockv1alpha2.StorageNode, + cluster *simplyblockv1alpha2.StorageCluster, + held []simplyblockv1alpha2.ProvisioningSlot, + limit int32, +) (bool, error) { + if int32(len(held)) >= limit { + return false, nil + } + + taking := append(slices.Clone(held), simplyblockv1alpha2.ProvisioningSlot{ + Worker: node.Spec.WorkerNode, + Node: node.Name, + TakenAt: metav1.Now(), + }) + err := r.writeSlots(ctx, cluster, taking) + if apierrors.IsConflict(err) { + // Somebody else patched the list between this reconcile reading the + // cluster and writing to it, which is the race the lock exists for and + // not a failure. The count this node made is stale, so it makes it again + // on the next pass rather than appending to a list it never saw. + return false, nil + } + return err == nil, err +} + +// releaseSlot gives up the slot this node holds, and only that one. +// +// A release that cleared the list would free whichever add happened to be +// running alongside this one, which is the same discipline the cluster operation +// lock is released under: an owner clears its own entry or nothing. +// +// It is called on every path that ends a node's add — the UUID arriving, the +// deadline running out, the object going away — and it is safe to call when no +// slot is held, because the list is then already what it should be. +func (r *StorageNodeReconciler) releaseSlot( + ctx context.Context, node *simplyblockv1alpha2.StorageNode, +) error { + return retry.RetryOnConflict(retry.DefaultRetry, func() error { + var cluster simplyblockv1alpha2.StorageCluster + key := types.NamespacedName{Name: node.Spec.ClusterRef, Namespace: node.Namespace} + if err := r.Get(ctx, key, &cluster); err != nil { + // A cluster that is gone holds no slots. + return client.IgnoreNotFound(err) + } + + kept := slices.DeleteFunc( + slices.Clone(cluster.Status.ProvisioningSlots), + func(slot simplyblockv1alpha2.ProvisioningSlot) bool { + return slot.Node == node.Name + }) + if len(kept) == len(cluster.Status.ProvisioningSlots) { + return nil + } + return r.writeSlots(ctx, &cluster, kept) + }) +} + +// heldSlots is the recorded list with the entries nobody will ever release +// removed. +// +// Three things end a hold without the holder getting to say so: the object is +// deleted, the add it was taken for has already produced a UUID, and the node has +// failed. None of them are reachable from the holder's own reconcile in the case +// that matters — a deleted object has no reconcile left — so the next node to ask +// for a slot is what collects them. A slot nothing reaps is a cap that never +// reopens. +func (r *StorageNodeReconciler) heldSlots( + ctx context.Context, cluster *simplyblockv1alpha2.StorageCluster, +) ([]simplyblockv1alpha2.ProvisioningSlot, error) { + held := make([]simplyblockv1alpha2.ProvisioningSlot, 0, len(cluster.Status.ProvisioningSlots)) + for _, slot := range cluster.Status.ProvisioningSlots { + var holder simplyblockv1alpha2.StorageNode + key := types.NamespacedName{Name: slot.Node, Namespace: cluster.Namespace} + if err := r.Get(ctx, key, &holder); err != nil { + if apierrors.IsNotFound(err) { + continue + } + return nil, err + } + if holder.Status.UUID != "" || + holder.Status.Phase == simplyblockv1alpha2.StorageNodePhaseFailed || + !holder.DeletionTimestamp.IsZero() { + continue + } + held = append(held, slot) + } + return held, nil +} + +// writeSlots patches the list under the resourceVersion the cluster was read at, +// which is the whole of the mutual exclusion. +func (r *StorageNodeReconciler) writeSlots( + ctx context.Context, + cluster *simplyblockv1alpha2.StorageCluster, + slots []simplyblockv1alpha2.ProvisioningSlot, +) error { + patch := client.MergeFromWithOptions( + cluster.DeepCopy(), client.MergeFromWithOptimisticLock{}) + cluster.Status.ProvisioningSlots = slots + if err := r.Status().Patch(ctx, cluster, patch); err != nil { + return fmt.Errorf("record the node-add slots of cluster %s: %w", cluster.Name, err) + } + return nil +} + +// slotHolders is the workers currently holding a slot, which is what the +// FoundationDB rule and the blocked message are phrased over. +func slotHolders(slots []simplyblockv1alpha2.ProvisioningSlot) []string { + workers := make([]string, 0, len(slots)) + for _, slot := range slots { + workers = append(workers, slot.Worker) + } + slices.Sort(workers) + return workers +} + +// describeSlots renders the holders for the event that reports a node waiting, +// with how long each has been holding. +// +// The age is the part worth reading: a cap doing its job shows a slot seconds +// old, and a cap that has stuck shows one held for hours, and the two are the +// same event without it. +func describeSlots(slots []simplyblockv1alpha2.ProvisioningSlot) string { + parts := make([]string, 0, len(slots)) + for _, slot := range slots { + if slot.TakenAt.IsZero() { + parts = append(parts, slot.Worker) + continue + } + parts = append(parts, fmt.Sprintf("%s for %s", + slot.Worker, time.Since(slot.TakenAt.Time).Truncate(time.Second))) + } + slices.Sort(parts) + return strings.Join(parts, ", ") +} diff --git a/operator/internal/controllers/node/provisioningslots_test.go b/operator/internal/controllers/node/provisioningslots_test.go new file mode 100644 index 000000000..50ea4e7cc --- /dev/null +++ b/operator/internal/controllers/node/provisioningslots_test.go @@ -0,0 +1,320 @@ +// The node-add cap under reconcilers that do not see each other. +// +// The cap used to be an election: every waiting node ordered the waiting workers +// and advanced only if its own place was within the number free. Two nodes +// reading one set reach one answer, which is true and is the whole of the +// assumption. Reconciles are not served from the API server but from an informer +// cache, and a cache that has just been filled — an operator restart, a new lease +// holder, a resync — does not hold every sibling yet. Nodes reading different +// sets reach different answers, and each of them is alone in the set it can see. +// +// These cases starve the sibling listing to model that, and assert on the cap +// rather than on how it is enforced. + +package node + +import ( + "context" + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + "sigs.k8s.io/controller-runtime/pkg/client/interceptor" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/testsupport" +) + +// blindReconcilers builds one reconciler per node over a single shared backing +// store, each of which can list only its own node. +// +// The share is the point: the objects are one set of objects, as they are on a +// cluster, and only the listing is starved. Anything the reconcilers do to reach +// an agreement has to go through the store they have in common. +func blindReconcilers( + t *testing.T, + cluster *simplyblockv1alpha2.StorageCluster, + nodes ...*simplyblockv1alpha2.StorageNode, +) []*StorageNodeReconciler { + t.Helper() + scheme := testsupport.NewScheme(t, corev1.AddToScheme) + + objects := make([]client.Object, 0, 1+len(nodes)) + objects = append(objects, cluster) + for _, n := range nodes { + objects = append(objects, n) + } + shared := fake.NewClientBuilder().WithScheme(scheme). + WithObjects(objects...). + WithStatusSubresource(&simplyblockv1alpha2.StorageCluster{}). + WithIndex(&simplyblockv1alpha2.StorageNode{}, clusterRefField, + func(o client.Object) []string { + return []string{o.(*simplyblockv1alpha2.StorageNode).Spec.ClusterRef} + }). + Build() + + reconcilers := make([]*StorageNodeReconciler, 0, len(nodes)) + for _, node := range nodes { + only := node + blind := interceptor.NewClient(shared, interceptor.Funcs{ + List: func( + _ context.Context, _ client.WithWatch, list client.ObjectList, _ ...client.ListOption, + ) error { + nodeList, ok := list.(*simplyblockv1alpha2.StorageNodeList) + if !ok { + return nil + } + nodeList.Items = []simplyblockv1alpha2.StorageNode{*only} + return nil + }, + }) + reconcilers = append(reconcilers, &StorageNodeReconciler{Client: blind, Scheme: scheme}) + } + return reconcilers +} + +// admittedBlind reconciles every node once, each reading the cluster for itself +// first, and reports how many were sent on to post. +// +// The read per node is what a reconcile does: provision fetches the cluster at +// the top of every pass. Modeling it as one read shared by all three would be +// modeling three reconciles that began in the same instant, which is the case +// the optimistic lock is for and not the case a cap is measured in. +func admittedBlind( + t *testing.T, + reconcilers []*StorageNodeReconciler, + nodes ...*simplyblockv1alpha2.StorageNode, +) int { + t.Helper() + count := 0 + for i, node := range nodes { + next, _, err := reconcilers[i].awaitSlot( + context.Background(), node, storedCluster(t, reconcilers[i])) + if err != nil { + continue + } + if next == stepPosting { + count++ + } + } + return count +} + +// storedCluster reads the cluster back from a reconciler's store. +func storedCluster( + t *testing.T, r *StorageNodeReconciler, +) *simplyblockv1alpha2.StorageCluster { + t.Helper() + var cluster simplyblockv1alpha2.StorageCluster + key := client.ObjectKey{Name: "c", Namespace: "simplyblock"} + if err := r.Get(context.Background(), key, &cluster); err != nil { + t.Fatalf("read the cluster back: %v", err) + } + return &cluster +} + +// TestTheCapHoldsWhenNodesCannotSeeEachOther is the cap under a cold cache. +// +// Regression: 2026-09-20-node-add-cap-lost-under-a-cold-cache — the cap was an +// election over the set of waiting nodes, so a reconciler whose cache did not yet +// hold the siblings elected itself and posted. On a six-worker cluster that is +// six simultaneous adds against a cap of one, and a node add reboots its host: +// the cap exists so that two hosts do not reboot at once and cost the control +// plane its own fault tolerance. +func TestTheCapHoldsWhenNodesCannotSeeEachOther(t *testing.T) { + a, b, c := waitingNode("w1"), waitingNode("w2"), waitingNode("w3") + cluster := slotCluster(1) + reconcilers := blindReconcilers(t, cluster, a, b, c) + + if got := admittedBlind(t, reconcilers, a, b, c); got != 1 { + t.Errorf("%d nodes took a single free slot while none could see the others", got) + } +} + +// The cap holds above one as well, which is what separates a cap from a mutex. +func TestACapOfTwoHoldsWhenNodesCannotSeeEachOther(t *testing.T) { + a, b, c := waitingNode("w1"), waitingNode("w2"), waitingNode("w3") + cluster := slotCluster(2) + reconcilers := blindReconcilers(t, cluster, a, b, c) + + if got := admittedBlind(t, reconcilers, a, b, c); got != 2 { + t.Errorf("%d nodes took two free slots while none could see the others", got) + } +} + +// A node that took a slot is readable from the cluster, by the worker that holds +// it and the object that took it. +func TestTakingASlotIsRecordedOnTheCluster(t *testing.T) { + a := waitingNode("w1") + cluster := slotCluster(1) + reconcilers := blindReconcilers(t, cluster, a) + + next, _, err := reconcilers[0].awaitSlot(context.Background(), a, cluster.DeepCopy()) + if err != nil { + t.Fatalf("awaitSlot: %v", err) + } + if next != stepPosting { + t.Fatalf("the only waiting node was not admitted, it was sent to %s", next) + } + + slots := storedCluster(t, reconcilers[0]).Status.ProvisioningSlots + if len(slots) != 1 { + t.Fatalf("the cluster records %d slots after one was taken", len(slots)) + } + if slots[0].Worker != "w1" || slots[0].Node != a.Name { + t.Errorf("the slot reads worker %q node %q", slots[0].Worker, slots[0].Node) + } +} + +// Re-entering the step with a slot already held keeps the one entry, because a +// held slot is not a reason to take a second. +func TestASlotIsNotTakenTwice(t *testing.T) { + a := waitingNode("w1") + cluster := slotCluster(1) + reconcilers := blindReconcilers(t, cluster, a) + + for range 3 { + if _, _, err := reconcilers[0].awaitSlot( + context.Background(), a, storedCluster(t, reconcilers[0]), + ); err != nil { + t.Fatalf("awaitSlot: %v", err) + } + } + + if slots := storedCluster(t, reconcilers[0]).Status.ProvisioningSlots; len(slots) != 1 { + t.Errorf("three passes of one node left %d slots taken", len(slots)) + } +} + +// A slot is released by the node that took it and by nothing else, which is the +// same discipline the cluster operation lock is released under. +func TestASlotIsReleasedOnlyByItsHolder(t *testing.T) { + a, b := waitingNode("w1"), waitingNode("w2") + cluster := slotCluster(1) + cluster.Status.ProvisioningSlots = []simplyblockv1alpha2.ProvisioningSlot{ + {Worker: "w1", Node: a.Name, TakenAt: metav1.Now()}, + } + reconcilers := blindReconcilers(t, cluster, a, b) + + if err := reconcilers[1].releaseSlot(context.Background(), b); err != nil { + t.Fatalf("releaseSlot: %v", err) + } + if slots := storedCluster(t, reconcilers[1]).Status.ProvisioningSlots; len(slots) != 1 { + t.Fatalf("a node released a slot it did not hold, leaving %d", len(slots)) + } + + if err := reconcilers[0].releaseSlot(context.Background(), a); err != nil { + t.Fatalf("releaseSlot: %v", err) + } + if slots := storedCluster(t, reconcilers[0]).Status.ProvisioningSlots; len(slots) != 0 { + t.Errorf("the holder released its slot and %d remain", len(slots)) + } +} + +// A slot whose node object is gone is reaped, because nothing will ever release +// it and a cap held by a deleted object never reopens. +func TestASlotWhoseNodeIsGoneIsReaped(t *testing.T) { + b := waitingNode("w2") + cluster := slotCluster(1) + cluster.Status.ProvisioningSlots = []simplyblockv1alpha2.ProvisioningSlot{ + {Worker: "w1", Node: "c-w1-0", TakenAt: metav1.Now()}, + } + reconcilers := blindReconcilers(t, cluster, b) + + next, _, err := reconcilers[0].awaitSlot(context.Background(), b, cluster.DeepCopy()) + if err != nil { + t.Fatalf("awaitSlot: %v", err) + } + if next != stepPosting { + t.Fatalf("a slot held by an object that no longer exists still blocked the cap") + } +} + +// A slot whose node has already resolved its UUID is reaped too: the add it was +// taken for is finished, whatever the object has got round to writing. +func TestASlotWhoseNodeIsFinishedIsReaped(t *testing.T) { + done, b := waitingNode("w1"), waitingNode("w2") + done.Status.UUID = "4a304439-2909-4199-ad2f-b8624d66a13d" + cluster := slotCluster(1) + cluster.Status.ProvisioningSlots = []simplyblockv1alpha2.ProvisioningSlot{ + {Worker: "w1", Node: done.Name, TakenAt: metav1.Now()}, + } + reconcilers := blindReconcilers(t, cluster, done, b) + + next, _, err := reconcilers[1].awaitSlot(context.Background(), b, cluster.DeepCopy()) + if err != nil { + t.Fatalf("awaitSlot: %v", err) + } + if next != stepPosting { + t.Fatalf("a slot held by a node that already has its UUID still blocked the cap") + } +} + +// TestASlotTakenFromAStaleReadIsRefused is the mutual exclusion itself. +// +// Two nodes beginning a reconcile in the same instant both read a cluster with a +// free slot. The second one's patch is conditioned on the resourceVersion it +// read, which the first one's patch has moved, so it is refused and it counts +// again rather than appending to a list it never saw. Without the condition both +// appends land and the cap is the number of nodes. +func TestASlotTakenFromAStaleReadIsRefused(t *testing.T) { + a, b := waitingNode("w1"), waitingNode("w2") + cluster := slotCluster(2) + reconcilers := blindReconcilers(t, cluster, a, b) + + // Both hold the cluster as it was before either of them wrote. + stale := storedCluster(t, reconcilers[0]) + + first, _, err := reconcilers[0].awaitSlot(context.Background(), a, stale.DeepCopy()) + if err != nil { + t.Fatalf("awaitSlot: %v", err) + } + if first != stepPosting { + t.Fatalf("the first node was not admitted to a cluster with two free slots") + } + + second, _, err := reconcilers[1].awaitSlot(context.Background(), b, stale.DeepCopy()) + if err == nil && second == stepPosting { + t.Fatal("a node took a slot against a cluster it had not seen the last write to") + } + + if slots := storedCluster(t, reconcilers[0]).Status.ProvisioningSlots; len(slots) != 1 { + t.Errorf("two nodes patching one read left %d slots", len(slots)) + } +} + +// TestASecondSocketDoesNotTakeASecondSlot covers the two-socket worker. +// +// One POST adds every socket of a worker, so a host with two of them must +// consume one slot and the second object must resolve against the add the first +// one asked for. The step-based sibling check answers this a pass later than the +// slot does — a node that has taken a slot is still recorded at AwaitingSlot +// until its transition is written — so with more than one slot free the second +// socket reaches the cap with room in it. +func TestASecondSocketDoesNotTakeASecondSlot(t *testing.T) { + first := waitingNode("w1") + second := waitingNode("w1") + second.Name = "c-w1-1" + cluster := slotCluster(2) + reconcilers := blindReconcilers(t, cluster, first, second) + + if next, _, err := reconcilers[0].awaitSlot( + context.Background(), first, storedCluster(t, reconcilers[0]), + ); err != nil || next != stepPosting { + t.Fatalf("the first socket was sent to %s: %v", next, err) + } + + next, _, err := reconcilers[1].awaitSlot( + context.Background(), second, storedCluster(t, reconcilers[1])) + if err != nil { + t.Fatalf("awaitSlot: %v", err) + } + if next == stepPosting { + t.Error("the second socket of one worker posted an add of its own") + } + if slots := storedCluster(t, reconcilers[1]).Status.ProvisioningSlots; len(slots) != 1 { + t.Errorf("one worker holds %d slots", len(slots)) + } +} diff --git a/operator/internal/controllers/node/slot_race_test.go b/operator/internal/controllers/node/slot_race_test.go index a0747224d..fffb5f3c9 100644 --- a/operator/internal/controllers/node/slot_race_test.go +++ b/operator/internal/controllers/node/slot_race_test.go @@ -12,10 +12,10 @@ // holds once and then means nothing, which is the shape a cap fails in when it // is a count of other people's writes. // -// The answer is to decide from what is observed rather than from what has been -// written: the waiting nodes are ordered, and a node takes a slot only if its -// own place in that order is within the number free. Two nodes reading the same -// set reach the same answer without either having written anything. +// A slot is now taken rather than counted: the holders are a list on the cluster +// and the taking is an optimistic-locked patch of it. These cases are the cap's +// arithmetic — one free slot admits one, two admit two, a slot already held +// admits nobody — and provisioningslots_test.go is the mutual exclusion itself. package node @@ -63,16 +63,22 @@ func slotCluster(limit int32) *simplyblockv1alpha2.StorageCluster { return cluster } -func slotReconciler(t *testing.T, nodes ...*simplyblockv1alpha2.StorageNode) *StorageNodeReconciler { +func slotReconciler( + t *testing.T, + cluster *simplyblockv1alpha2.StorageCluster, + nodes ...*simplyblockv1alpha2.StorageNode, +) *StorageNodeReconciler { t.Helper() scheme := testsupport.NewScheme(t, corev1.AddToScheme) - objects := make([]client.Object, 0, len(nodes)) + objects := make([]client.Object, 0, 1+len(nodes)) + objects = append(objects, cluster) for _, n := range nodes { objects = append(objects, n) } apiClient := fake.NewClientBuilder().WithScheme(scheme). WithObjects(objects...). + WithStatusSubresource(&simplyblockv1alpha2.StorageCluster{}). WithIndex(&simplyblockv1alpha2.StorageNode{}, clusterRefField, func(o client.Object) []string { return []string{o.(*simplyblockv1alpha2.StorageNode).Spec.ClusterRef} @@ -82,17 +88,34 @@ func slotReconciler(t *testing.T, nodes ...*simplyblockv1alpha2.StorageNode) *St return &StorageNodeReconciler{Client: apiClient, Scheme: scheme} } +// heldBy is a cluster whose slots are already taken by the workers given, as a +// node that has posted its add and is waiting for the UUID leaves it. +func heldBy( + cluster *simplyblockv1alpha2.StorageCluster, nodes ...*simplyblockv1alpha2.StorageNode, +) *simplyblockv1alpha2.StorageCluster { + for _, node := range nodes { + cluster.Status.ProvisioningSlots = append( + cluster.Status.ProvisioningSlots, simplyblockv1alpha2.ProvisioningSlot{ + Worker: node.Spec.WorkerNode, Node: node.Name, TakenAt: metav1.Now(), + }) + } + return cluster +} + // admitted is how many of the waiting nodes would advance on one pass, each -// deciding for itself as it does in a real reconcile. +// reading the cluster and deciding for itself as it does in a real reconcile. func admitted( - t *testing.T, r *StorageNodeReconciler, - cluster *simplyblockv1alpha2.StorageCluster, - nodes ...*simplyblockv1alpha2.StorageNode, + t *testing.T, r *StorageNodeReconciler, nodes ...*simplyblockv1alpha2.StorageNode, ) int { t.Helper() count := 0 for _, node := range nodes { - next, _, err := r.awaitSlot(context.Background(), node, cluster) + var cluster simplyblockv1alpha2.StorageCluster + key := client.ObjectKey{Name: "c", Namespace: "simplyblock"} + if err := r.Get(context.Background(), key, &cluster); err != nil { + t.Fatalf("read the cluster: %v", err) + } + next, _, err := r.awaitSlot(context.Background(), node, &cluster) if err != nil { continue } @@ -107,9 +130,9 @@ func admitted( // takes the slot. func TestOnlyOneNodeTakesAFreeSlot(t *testing.T) { a, b, c := waitingNode("w1"), waitingNode("w2"), waitingNode("w3") - r := slotReconciler(t, a, b, c) + r := slotReconciler(t, slotCluster(1), a, b, c) - if got := admitted(t, r, slotCluster(1), a, b, c); got != 1 { + if got := admitted(t, r, a, b, c); got != 1 { t.Errorf("%d nodes took a single free slot", got) } } @@ -117,9 +140,9 @@ func TestOnlyOneNodeTakesAFreeSlot(t *testing.T) { // The cap is honored above one too: two free slots admit two of three. func TestACapOfTwoAdmitsTwo(t *testing.T) { a, b, c := waitingNode("w1"), waitingNode("w2"), waitingNode("w3") - r := slotReconciler(t, a, b, c) + r := slotReconciler(t, slotCluster(2), a, b, c) - if got := admitted(t, r, slotCluster(2), a, b, c); got != 2 { + if got := admitted(t, r, a, b, c); got != 2 { t.Errorf("%d nodes took two free slots", got) } } @@ -128,14 +151,21 @@ func TestACapOfTwoAdmitsTwo(t *testing.T) { // does not overtake one that was told to go on the next pass. func TestTheChoiceIsStable(t *testing.T) { a, b, c := waitingNode("w1"), waitingNode("w2"), waitingNode("w3") - r := slotReconciler(t, a, b, c) - cluster := slotCluster(1) + r := slotReconciler(t, slotCluster(1), a, b, c) - first, _, err := r.awaitSlot(context.Background(), a, cluster) + var cluster simplyblockv1alpha2.StorageCluster + key := client.ObjectKey{Name: "c", Namespace: "simplyblock"} + if err := r.Get(context.Background(), key, &cluster); err != nil { + t.Fatal(err) + } + first, _, err := r.awaitSlot(context.Background(), a, &cluster) if err != nil { t.Fatal(err) } - second, _, err := r.awaitSlot(context.Background(), a, cluster) + if err := r.Get(context.Background(), key, &cluster); err != nil { + t.Fatal(err) + } + second, _, err := r.awaitSlot(context.Background(), a, &cluster) if err != nil { t.Fatal(err) } @@ -149,9 +179,9 @@ func TestANodeInFlightFillsTheCap(t *testing.T) { inflight := waitingNode("w1") inflight.Status.Step.State = string(stepPosting) b, c := waitingNode("w2"), waitingNode("w3") - r := slotReconciler(t, inflight, b, c) + r := slotReconciler(t, heldBy(slotCluster(1), inflight), inflight, b, c) - if got := admitted(t, r, slotCluster(1), b, c); got != 0 { + if got := admitted(t, r, b, c); got != 0 { t.Errorf("%d nodes advanced past a full cap", got) } } diff --git a/operator/internal/controllers/node/storagenode_controller.go b/operator/internal/controllers/node/storagenode_controller.go index 3969d3a67..edea8af61 100644 --- a/operator/internal/controllers/node/storagenode_controller.go +++ b/operator/internal/controllers/node/storagenode_controller.go @@ -126,6 +126,8 @@ type NodeCapacitySource interface { // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodes/status,verbs=get;update;patch // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodes/finalizers,verbs=update // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagenodeops,verbs=get;list;watch;create;delete +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storageclusters,verbs=get;list;watch +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storageclusters/status,verbs=get;update;patch // +kubebuilder:rbac:groups="",resources=secrets,verbs=get;list;watch // SetupWithManager registers the controller. @@ -380,7 +382,11 @@ func (r *StorageNodeReconciler) provision( // whether the graph gives the state an exit: Resolving has one now, to // the step that holds while the worker is away, and a finished Resolving // is still the end of the path. - return ctrl.Result{RequeueAfter: nodeAdvance}, nil + // + // The add is over, so the slot it was taken for goes back. This is the + // successful one of the three ends, and the other two are the deadline in + // fail and the object going away in teardown. + return ctrl.Result{RequeueAfter: nodeAdvance}, r.releaseSlot(ctx, node) } if err := machine.TransitionTo(ctx, next); err != nil { @@ -539,7 +545,8 @@ func (r *StorageNodeReconciler) checkConfig( return stepAwaitingSlot, true, nil } -// awaitSlot is where two independent serialization rules live. +// awaitSlot is where two independent serialization rules live, and the last place +// a node that needs no add at all is recognized. // // maxParallelNodeAdds caps how many workers may be in flight at once, counted by // distinct worker rather than by object so that a two-socket host consumes one @@ -557,6 +564,24 @@ func (r *StorageNodeReconciler) awaitSlot( node *simplyblockv1alpha2.StorageNode, cluster *simplyblockv1alpha2.StorageCluster, ) (nodeStep, bool, error) { + // A backend node the worker already has is the answer to the question this + // queue exists to ask, so it is asked before the queue rather than only in + // front of it. The two gates before this one check the same thing and a node + // passes them once: one that reached the queue before its backend node + // existed — a previous operator's add, a POST whose response was lost after + // the control plane committed, a rebuilt object — would otherwise wait for a + // slot it has no use for, and hold the cap against the rest of the fleet + // while it did. + // + // It is checked before the worker is, for the reason CheckingHost checks it + // first: an adopted node is already running, and the machine being reachable + // is not this operator's precondition to establish. + if _, found, err := r.matchBackendNode(ctx, node, cluster); err != nil { + return stepAwaitingSlot, false, err + } else if found { + return stepAdopting, true, nil + } + // A slot taken for a machine that is not there is an add posted against a // worker that cannot answer it. The wait costs nothing, because the slot is // still free when the machine comes back, and taking it would put a second @@ -592,36 +617,50 @@ func (r *StorageNodeReconciler) awaitSlot( limit = *spec.MaxParallelNodeAdds } - inFlight := map[string]struct{}{} - for i := range siblings { - sibling := &siblings[i] - if sibling.Name == node.Name || sibling.Spec.WorkerNode == node.Spec.WorkerNode { - continue - } - if claimedWorker(sibling) && sibling.Status.UUID == "" { - inFlight[sibling.Spec.WorkerNode] = struct{}{} + // Who is in flight is read from the cluster's own record rather than counted + // off the siblings, because the siblings are read from a cache and the record + // is read from the object this node is about to patch. A cache that has just + // been filled holds some of the nodes and a cap counted from it is a cap each + // node applies to a different cluster. + held, err := r.heldSlots(ctx, cluster) + if err != nil { + return stepAwaitingSlot, false, err + } + // A slot held for this worker settles this node's answer, whichever object + // took it. Its own slot means it may post, and a slot its other socket took + // means the add has been asked for and this object resolves against it + // instead of asking again. + // + // The sibling check above answers the second case a pass later than this + // does, because a node that has taken a slot is still recorded at + // AwaitingSlot until its transition is written. With more than one slot free + // that pass is long enough for the second socket to post an add of its own. + if i := slices.IndexFunc(held, func(slot simplyblockv1alpha2.ProvisioningSlot) bool { + return slot.Worker == node.Spec.WorkerNode + }); i >= 0 { + if held[i].Node == node.Name { + return stepPosting, true, nil } + return stepResolving, true, nil } - available := limit - int32(len(inFlight)) - if available <= 0 { + + if int32(len(held)) >= limit { return stepAwaitingSlot, false, blockedf(AwaitingSlot, - "waiting for a node-add slot, %d of %d in flight", len(inFlight), limit) + "waiting for a node-add slot, %d of %d in flight: %s", + len(held), limit, describeSlots(held)) } + available := limit - int32(len(held)) - // Which of the waiting workers may take the free slots is decided from the - // set that is waiting, not from who has already written a claim. + // Which of the waiting workers should try for the free slots is decided from + // the set that is waiting, and it is a tie-break rather than the cap. // - // The difference is the whole of the cap. Reconciles are serialized per - // object and not across objects, so every node waiting for a slot reads the - // in-flight count before any of them has recorded taking one: the first add - // is correctly alone, and the instant it finishes every remaining node sees - // the same free slot and takes it. A cap that counts other people's writes - // holds exactly once. - // - // Ordering the contenders and admitting the first few needs nobody to have - // written anything. Two nodes reading one set reach one answer, and the - // answer does not change between passes, so a node told to wait is not - // overtaken by one told to wait beside it. + // It settles who goes first among nodes that can see each other, so a node + // told to wait is not overtaken by one told to wait beside it, and it needs + // nobody to have written anything: two nodes reading one set reach one + // answer. What it cannot do is hold the cap, because nodes reading a cold + // cache read different sets and each is alone in its own. The patch below is + // what holds the cap. + holders := slotHolders(held) contenders := []string{node.Spec.WorkerNode} for i := range siblings { sibling := &siblings[i] @@ -632,6 +671,13 @@ func (r *StorageNodeReconciler) awaitSlot( // Not waiting for a slot yet, so not competing for this one. continue } + if slices.Contains(holders, sibling.Spec.WorkerNode) { + // Already holding one, and the free slots are counted net of it, so + // ranking against it would charge this node for the same add twice. + // A worker is still at this step for the pass between taking its slot + // and the transition out being written. + continue + } contenders = append(contenders, sibling.Spec.WorkerNode) } slices.Sort(contenders) @@ -640,13 +686,13 @@ func (r *StorageNodeReconciler) awaitSlot( if rank := slices.Index(contenders, node.Spec.WorkerNode); int32(rank) >= available { return stepAwaitingSlot, false, blockedf(AwaitingSlot, "waiting for a node-add slot, %d of %d in flight and %d worker(s) ahead", - len(inFlight), limit, rank) + len(held), limit, rank) } // A FoundationDB worker waits for every other FoundationDB worker, whatever // the cap says. if r.hostsFoundationDB(ctx, node.Namespace, node.Spec.WorkerNode) { - for worker := range inFlight { + for _, worker := range holders { if r.hostsFoundationDB(ctx, node.Namespace, worker) { return stepAwaitingSlot, false, blockedf(AwaitingSlot, "worker %s hosts FoundationDB and worker %s is already being added", @@ -654,10 +700,19 @@ func (r *StorageNodeReconciler) awaitSlot( } } } + + took, err := r.takeSlot(ctx, node, cluster, held, limit) + if err != nil { + return stepAwaitingSlot, false, err + } + if !took { + return stepAwaitingSlot, false, blockedf(AwaitingSlot, + "another node took the last free node-add slot first") + } return stepPosting, true, nil } -// claimedWorker reports whether an object has claimed its worker, which is the +// claimedWorker reports// claimedWorker reports whether an object has claimed its worker, which is the // transition into Posting or anything past it. func claimedWorker(node *simplyblockv1alpha2.StorageNode) bool { switch nodeStep(node.Status.Step.State) { @@ -1047,6 +1102,12 @@ func (r *StorageNodeReconciler) teardown( if node.Status.UUID == "" { r.unregister(node) + // A node deleted while its add was outstanding is the one case where + // nothing is left to release the slot afterward, and a slot no reconcile + // will ever give back is a cap that never reopens. + if err := r.releaseSlot(ctx, node); err != nil { + return ctrl.Result{}, err + } controllerutil.RemoveFinalizer(node, NodeFinalizer) return ctrl.Result{}, r.Update(ctx, node) } @@ -1442,10 +1503,15 @@ func (r *StorageNodeReconciler) fail( ctx context.Context, node *simplyblockv1alpha2.StorageNode, message string, ) error { r.emit(node, corev1.EventTypeWarning, HostUnreachable, message) - return r.writeStatus(ctx, node, func(status *simplyblockv1alpha2.StorageNodeStatus) { - status.Phase = simplyblockv1alpha2.StorageNodePhaseFailed - status.Message = message - }) + // A node that has given up holds nothing. Releasing here rather than leaving + // it to the next node to reap is what keeps a cap from reading as full for as + // long as nobody happens to ask for a slot. + return errors.Join( + r.releaseSlot(ctx, node), + r.writeStatus(ctx, node, func(status *simplyblockv1alpha2.StorageNodeStatus) { + status.Phase = simplyblockv1alpha2.StorageNodePhaseFailed + status.Message = message + })) } // writeStatus applies the mutation and patches only when something changed, which diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusters.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusters.yaml index 49335795d..dc8ff3afb 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusters.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storageclusters.yaml @@ -1368,7 +1368,6 @@ spec: - message: enableAtomic4kWrites requires enableChecksumValidation to be true rule: '!(has(self.enableAtomic4kWrites) && self.enableAtomic4kWrites) || (has(self.enableChecksumValidation) && self.enableChecksumValidation)' status: - description: StorageClusterStatus is the observed state of one backend cluster. properties: activeOpsRef: description: |- @@ -1430,6 +1429,51 @@ spec: - Unavailable - Suspended type: string + provisioningSlots: + description: |- + ProvisioningSlots are the workers whose node add is outstanding. The list + is the metadata of the Provisioning phase, and it is also the mutex that + caps concurrent adds at spec.storageNodes.maxParallelNodeAdds. + + It is one list on one object because that is what makes taking a slot + atomic. A node takes one with an optimistic-locked patch of this field, so + exactly one node wins a given resourceVersion and every other is told to + count again. A slot recorded per node could not do that: six objects carry + six resourceVersions, and two nodes reading a cold cache would both see the + same one free and both take it. + + A worker rather than an object is what holds a slot, because one POST adds + every socket of a worker and a two-socket host must consume one slot. + items: + description: |- + StorageClusterStatus is the observed state of one backend cluster. + ProvisioningSlot is one worker's hold on the cluster's node-add concurrency, + held from the moment its add is posted until the node has a backend UUID or + has given up. + properties: + node: + description: |- + Node is the StorageNode object that took the slot. A release removes only + the entry naming its own object, which is what keeps a node from freeing + somebody else's slot, and a slot whose object is gone is reaped. + type: string + takenAt: + description: |- + TakenAt is when the slot was taken, so a hold that outlives its node's + deadlines is visible in the object rather than only in the events. + format: date-time + type: string + worker: + description: |- + Worker is the Kubernetes node the add was posted for, and it is what the + cap counts. + type: string + required: + - node + - worker + type: object + maxItems: 64 + type: array realignedGeneration: description: |- RealignedGeneration is the generation the last successfully requested From 5b1cf4ccad9acbd006d82a752ead51839c03a4aa Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Mon, 21 Sep 2026 09:32:40 +0200 Subject: [PATCH 105/206] feat(operator): the installed control plane serves TLS, and does by default #544 moved the control-plane install off the chart and into the operator. The chart had served it over TLS behind tls.enabled; the ControlPlane spec had no field to carry that decision, so the install was the plaintext one whatever a deployment asked for. The chart refused tls.enabled together with the standalone profile rather than installing something quieter than what was asked for, which made a TLS deployment unrenderable: the lab e2e has been failing on that refusal. spec.source.local.tls carries it now, and it defaults to on. The field names are DriverTLS's -- enableTLS, enableMutualTLS, provider -- because they are the same three decisions about the same connection seen from its two ends, and a reader who knows one should not have to learn the other. What differs is the default. The CSI driver's fields were added to describe deployments that already existed and default off; these default on, because an installed control plane holds every cluster definition, node registration, and volume record and is reached over the pod network by the operator, the CSI driver, and the metrics scrape alike. The helpers answer from the unset value rather than from a defaulted field. A reconciler builds these objects in tests and reads them through caches that never went past admission, and a TLS decision that depends on the API server having run reads as plaintext exactly where it is not checked. What the install carries: The environment the control-plane image has always spoken -- SB_TLS_SERVE, SB_TLS_PROVIDER, SB_TLS_CONNECT, SB_TLS_CLIENT_AUTH -- on every pod rather than only the one that serves. The task and monitoring pools are clients of the management API, and a plaintext pool cannot run the control plane's own tasks against an API that requires a certificate. The serving certificate. It was rendered beside the Service it certifies, and the install that replaced the chart took the Service and left the certificate, so every pod mounted a Secret nothing produced. An invariant test already in the package caught it. That test learned about Secrets a Certificate issues rather than being given an exemption by name, so it still catches one nothing produces. FoundationDB's peer TLS, which had been dropped whole: the certificate, the volume repeated per process class -- the FoundationDB operator replaces general.podTemplate wholesale for a class that overrides it, so stating it once reaches neither storage nor log -- and mainContainer.enableTls. The last is the half that matters and is a field rather than an environment variable: processes carrying a current certificate, key, and CA and no enableTls talk to each other in the clear while every mount looks right. Peer TLS follows the client-certificate decision rather than the serving one, because there is no anonymous mode between database processes. The operator reached it over the connection it already had for the management API. localEndpoint returned the plaintext scheme whatever the install did, and it is published as status.endpoint, which every control-plane call in the operator resolves through, so the scheme now reaches all of them at once. The atlas-lib client the data-protection band writes through was given the endpoint alone, so it built its own transport: the system trust store and no client certificate. controlplane.Config takes a Transport, and webapi.ControlPlaneConfig hands over the one the startup client already resolved from this pod's environment and mounted certificates. A first attempt built a second transport here, reading the same material through the Kubernetes API instead of the files; it is not in this change, because a second implementation of a primitive is worse than a dependency on the first. The metrics scrape verified nothing. It was https with insecure_skip_verify while the CA bundle sat unused beside it; it now mounts simplyblock-ca-bundle-tls and verifies against it. tls.enabled and tls.mutual_enabled default to on, so this is what a plain install produces. cert-manager is the prerequisite, and the chart still refuses to render without it -- the message now names both ways out, because a deployment reaching it did not ask for the thing it is being refused. Two notes for the release. An existing deployment that upgrades without pinning tls.enabled=false gets enableTLS: true in its SimplyblockDriver while its plugins run plaintext, and adoption then refuses the mismatch rather than reconfiguring a live data path, so the driver stops reconciling until the value is pinned or the deployment is migrated. The same flip restarts an existing control plane's pods. The SimplyblockDriver defaults are deliberately not touched. Flipping a shipped field's default would make the operator refuse to adopt every running plaintext deployment whose CR omits it, and the chart renders both ends from one value, which is what actually keeps them agreeing. Test plan: U-149 through U-161. Co-Authored-By: Claude Opus 5 (1M context) --- atlas-lib/controlplane/client.go | 14 +- atlas-lib/controlplane/transport_test.go | 60 ++++ .../storage.simplyblock.io_controlplanes.yaml | 29 ++ .../templates/controlplane_configmap.yaml | 6 +- .../templates/controlplane_cr.yaml | 12 + .../templates/validate-controlplane.yaml | 27 +- .../templates/validate-tls.yaml | 13 +- .../charts/simplyblock-operator/values.yaml | 29 +- .../api/v1alpha2/controlplane_tls_test.go | 73 +++++ operator/api/v1alpha2/controlplane_types.go | 88 ++++++ .../api/v1alpha2/zz_generated.deepcopy.go | 26 ++ .../storage.simplyblock.io_controlplanes.yaml | 29 ++ operator/dist/install.yaml | 30 ++ .../crd-redesign/design-controlplane.md | 48 ++- operator/docs/tests/test-plan-controlplane.md | 22 ++ .../controlplane/controlplane_controller.go | 4 +- .../controlplaneops_controller.go | 2 +- .../internal/controllers/controlplane/doc.go | 21 +- .../controllers/controlplane/endpoint.go | 16 +- .../controllers/controlplane/foundationdb.go | 95 +++++- .../controllers/controlplane/install_test.go | 2 +- .../controllers/controlplane/managementapi.go | 31 +- .../controllers/controlplane/podspec.go | 20 +- .../internal/controllers/controlplane/tls.go | 247 +++++++++++++++ .../controllers/controlplane/tls_test.go | 295 ++++++++++++++++++ .../controlplane/workloads_test.go | 46 ++- .../storage.simplyblock.io_controlplanes.yaml | 29 ++ operator/internal/webapi/controlplane.go | 24 +- .../webapi/controlplane_transport_test.go | 112 +++++++ 29 files changed, 1374 insertions(+), 76 deletions(-) create mode 100644 atlas-lib/controlplane/transport_test.go create mode 100644 operator/api/v1alpha2/controlplane_tls_test.go create mode 100644 operator/internal/controllers/controlplane/tls.go create mode 100644 operator/internal/controllers/controlplane/tls_test.go create mode 100644 operator/internal/webapi/controlplane_transport_test.go diff --git a/atlas-lib/controlplane/client.go b/atlas-lib/controlplane/client.go index e499a2fc9..8125704cc 100644 --- a/atlas-lib/controlplane/client.go +++ b/atlas-lib/controlplane/client.go @@ -16,6 +16,18 @@ type Config struct { Endpoint string // base URL of the control-plane API Token string // cluster secret / bearer token Timeout time.Duration // per-request timeout, where zero means a sane default + + // Transport carries the connection, and is what a caller reaching a control + // plane over TLS supplies: the certificate pool that verifies the endpoint, + // and the client certificate it presents where the control plane requires + // one. + // + // It is a RoundTripper rather than a *tls.Config because the decision is the + // caller's and not every caller reaches the endpoint the same way. Nil is + // http.DefaultTransport, which is the system trust store and no client + // certificate -- correct for a plaintext endpoint and for one signed by a CA + // the host already trusts. + Transport http.RoundTripper } // Client talks to the simplyblock control-plane v2 API. It wraps the @@ -35,7 +47,7 @@ func New(cfg Config) (*Client, error) { } api, err := cpapi.NewClientWithResponses( cfg.Endpoint, - cpapi.WithHTTPClient(&http.Client{Timeout: cfg.Timeout}), + cpapi.WithHTTPClient(&http.Client{Timeout: cfg.Timeout, Transport: cfg.Transport}), cpapi.WithRequestEditorFn(bearerAuth(cfg.Token)), ) if err != nil { diff --git a/atlas-lib/controlplane/transport_test.go b/atlas-lib/controlplane/transport_test.go new file mode 100644 index 000000000..e37d1e0a1 --- /dev/null +++ b/atlas-lib/controlplane/transport_test.go @@ -0,0 +1,60 @@ +// The connection a Client is given, rather than the one it would have built. +// +// A control plane behind TLS is reached with a pool that verifies it and, where +// it requires one, a certificate to present. Neither is expressible by an +// endpoint and a token, so a Client that always built its own transport could +// only ever reach a plaintext control plane -- which is what it did. + +package controlplane + +import ( + "net/http" + "testing" + "time" +) + +// countingTransport records that it carried a request and answers a bare 200. +type countingTransport struct{ calls int } + +func (t *countingTransport) RoundTrip(req *http.Request) (*http.Response, error) { + t.calls++ + return &http.Response{ + StatusCode: http.StatusOK, + Body: http.NoBody, + Request: req, + }, nil +} + +// A configured transport is the one the requests go over. +func TestTheConfiguredTransportCarriesTheRequests(t *testing.T) { + carrier := &countingTransport{} + client, err := New(Config{ + Endpoint: "https://control-plane.simplyblock.svc.cluster.local:5000", + Token: "a-token", + Timeout: time.Second, + Transport: carrier, + }) + if err != nil { + t.Fatalf("New: %v", err) + } + + // Any call will do: what is asserted is which connection it went over, not + // what came back, and a bare 200 with no body fails to decode by design. + _, _ = client.ListBackups(t.Context(), "3f2b1c4d-5e6f-4a7b-8c9d-0e1f2a3b4c5d") + + if carrier.calls == 0 { + t.Error("the client built its own transport, so a TLS endpoint is unreachable") + } +} + +// A Config with no transport is still valid, and is the plaintext case every +// existing caller is. +func TestNoTransportIsStillAClient(t *testing.T) { + client, err := New(Config{Endpoint: "http://control-plane:5000"}) + if err != nil { + t.Fatalf("New: %v", err) + } + if client == nil { + t.Error("a client with no transport was not built") + } +} diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_controlplanes.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_controlplanes.yaml index 140c3da8f..95afa9396 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_controlplanes.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_controlplanes.yaml @@ -333,6 +333,35 @@ spec: More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ type: object type: object + tls: + default: {} + description: |- + TLS is how this control plane serves and what it asks of its callers. + + The block is defaulted to its own zero value rather than left absent, so + that the field defaults inside it are applied to an object that does not + mention TLS at all. + properties: + enableMutualTLS: + default: true + description: |- + EnableMutualTLS additionally requires a caller to present a certificate + of its own, rather than reaching the control plane anonymously over the + encrypted connection EnableTLS alone provides. Ignored when EnableTLS is + false, the same as on DriverTLS. Unset is on. + type: boolean + enableTLS: + default: true + description: EnableTLS serves the management API over TLS. Unset is on. + type: boolean + provider: + default: cert-manager + description: Provider issues the serving certificate. + enum: + - cert-manager + - OpenShift + type: string + type: object tolerations: description: |- Tolerations are applied to every pod the operator installs for the control diff --git a/helm-charts/charts/simplyblock-operator/templates/controlplane_configmap.yaml b/helm-charts/charts/simplyblock-operator/templates/controlplane_configmap.yaml index 7f467e736..479d9f7df 100644 --- a/helm-charts/charts/simplyblock-operator/templates/controlplane_configmap.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/controlplane_configmap.yaml @@ -132,7 +132,11 @@ data: {{- if .Values.tls.enabled }} scheme: https tls_config: - insecure_skip_verify: true + # The bundle the chart's ClusterIssuer writes, rather than + # insecure_skip_verify: an encrypted scrape that verifies nothing is + # answered by anything that can reach the Service name. + ca_file: /etc/prometheus/ca/ca.crt + server_name: simplyblock-webappapi {{- if .Values.tls.mutual_enabled }} cert_file: /etc/prometheus/certs/tls.crt key_file: /etc/prometheus/certs/tls.key diff --git a/helm-charts/charts/simplyblock-operator/templates/controlplane_cr.yaml b/helm-charts/charts/simplyblock-operator/templates/controlplane_cr.yaml index d0b827c6c..ea8aa8c6a 100644 --- a/helm-charts/charts/simplyblock-operator/templates/controlplane_cr.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/controlplane_cr.yaml @@ -44,6 +44,18 @@ spec: {{- if .Values.controlplane.storageclass.name }} storageClassName: {{ .Values.controlplane.storageclass.name }} {{- end }} + # The chart states both halves explicitly rather than leaving the CR's + # defaults to decide. The CR defaults to serving TLS, and the chart's + # tls.enabled defaults to off, and a release that said nothing would get + # the CR's answer instead of the one its values state -- which is a + # deployment that comes up over TLS because a field was omitted. + tls: + enableTLS: {{ .Values.tls.enabled }} + enableMutualTLS: {{ and .Values.tls.enabled .Values.tls.mutual_enabled }} + # The chart spells the providers in lower case and this API spells + # OpenShift as the product does, the same translation + # simplyblockdriver_cr.yaml performs. + provider: {{ if eq .Values.tls.provider "openshift" }}OpenShift{{ else }}cert-manager{{ end }} {{- if .Values.controlplane.nodeSelector.create }} nodeSelector: {{ .Values.controlplane.nodeSelector.key }}: {{ .Values.controlplane.nodeSelector.value | quote }} diff --git a/helm-charts/charts/simplyblock-operator/templates/validate-controlplane.yaml b/helm-charts/charts/simplyblock-operator/templates/validate-controlplane.yaml index e6dbfb9ba..baaaae183 100644 --- a/helm-charts/charts/simplyblock-operator/templates/validate-controlplane.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/validate-controlplane.yaml @@ -1,9 +1,14 @@ {{/* -The two ways a control-plane profile can be asked for something it cannot do. +The ways a control-plane profile can be asked for something it cannot do. -Both fail at render time rather than at apply time, because each produces a -deployment that comes up and is quietly wrong: one serving plaintext where TLS -was asked for, and one pointed at no control plane at all. +They fail at render time rather than at apply time, because each produces a +deployment that comes up and is quietly wrong: one pointed at no control plane +at all, and one whose upgrade would prune the database. + +A TLS standalone deployment used to be refused here. The ControlPlane spec now +carries spec.source.local.tls, the operator's install honors it, and +controlplane_cr.yaml renders it from tls.enabled, so the deployment that could +only come up in plaintext is expressible and the refusal is gone. */}} @@ -11,20 +16,6 @@ was asked for, and one pointed at no control plane at all. {{- fail (printf "deployment.profile is %q: it is one of standalone (this cluster hosts its own control plane) or managed (a control plane elsewhere manages this cluster's storage)." .Values.deployment.profile) -}} {{- end -}} -{{/* -TLS has no field on the ControlPlane spec. design-controlplane.md §5.1 settles -which certificate issuer is detected, not whether a deployment wants one, so an -operator-installed control plane comes up in plaintext whatever tls.enabled says. - -That was survivable while the chart could still install the control plane -itself. It cannot any more, so a TLS deployment has nowhere to go and the honest -answer is to refuse rather than to install something quieter than what was asked -for. -*/}} -{{- if and (eq .Values.deployment.profile "standalone") .Values.tls.enabled -}} -{{- fail "deployment.profile=standalone cannot serve TLS: the ControlPlane spec has no field for it, so the operator's install would bring the control plane up in plaintext. Until the spec can express it, a TLS deployment needs a control plane this cluster does not host (deployment.profile=managed), or an operator release that carries the field." -}} -{{- end -}} - {{/* The managed profile needs somewhere to point. The CR template requires the endpoint too, and this exists so the message names the value rather than the diff --git a/helm-charts/charts/simplyblock-operator/templates/validate-tls.yaml b/helm-charts/charts/simplyblock-operator/templates/validate-tls.yaml index bf6d356c9..0dbfa4750 100644 --- a/helm-charts/charts/simplyblock-operator/templates/validate-tls.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/validate-tls.yaml @@ -1,11 +1,20 @@ +{{/* +TLS needs an issuer, and which issuers exist is a property of the cluster rather +than of the values. + +Both messages name the way out as well as the problem, because TLS is on by +default: a deployment reaching either of these did not ask for the thing it is +being refused, so a message that only states what is missing leaves a reader to +work out that there was a choice at all. +*/}} {{- if .Values.tls.enabled -}} {{- if eq .Values.tls.provider "openshift" -}} {{- if not (.Capabilities.APIVersions.Has "config.openshift.io/v1/ClusterVersion") -}} -{{- fail "tls.provider=openshift requires installing on an OpenShift cluster" -}} +{{- fail "tls.provider=openshift signs through the OpenShift service CA, and this is not an OpenShift cluster. Set tls.provider=cert-manager where cert-manager is installed, or tls.enabled=false to deploy without TLS." -}} {{- end -}} {{- else if eq .Values.tls.provider "cert-manager" -}} {{- if not (.Capabilities.APIVersions.Has "cert-manager.io/v1") -}} -{{- fail "tls.provider=cert-manager requires cert-manager.io/v1 to be installed in the cluster" -}} +{{- fail "TLS is on by default and cert-manager issues its certificates, and cert-manager.io/v1 is not served by this cluster. Install cert-manager (helm repo add jetstack https://charts.jetstack.io && helm install cert-manager jetstack/cert-manager -n cert-manager --create-namespace --set installCRDs=true), or set tls.enabled=false to deploy without TLS." -}} {{- end -}} {{- end -}} {{- end -}} diff --git a/helm-charts/charts/simplyblock-operator/values.yaml b/helm-charts/charts/simplyblock-operator/values.yaml index cf4631dce..0f0076e8f 100644 --- a/helm-charts/charts/simplyblock-operator/values.yaml +++ b/helm-charts/charts/simplyblock-operator/values.yaml @@ -683,12 +683,23 @@ prometheus: secret: secretName: simplyblock-prometheus-client-tls optional: true + # The CA the control plane's certificate is verified against. It is its + # own volume because the client certificate above is issued only for + # mutual TLS, and a scrape that encrypts without verifying is one an + # in-cluster impersonator answers. + - name: prometheus-ca-tls + secret: + secretName: simplyblock-ca-bundle-tls + optional: true extraVolumeMounts: - name: simplyblock-prometheus-config mountPath: /etc/simplyblock-config - name: prometheus-client-tls mountPath: /etc/prometheus/certs readOnly: true + - name: prometheus-ca-tls + mountPath: /etc/prometheus/ca + readOnly: true alertmanager: enabled: false @@ -714,7 +725,15 @@ prometheus: image: defaultbackend-amd64 tag: "1.5" tls: - enabled: false + # On by default. Everything this chart deploys reaches the control plane over + # the pod network, and the control plane holds every cluster definition, node + # registration, and volume record, so plaintext is a deployment's decision to + # state rather than the one it reaches by leaving a value alone. + # + # It requires an issuer: with provider=cert-manager the chart refuses to render + # where cert-manager.io/v1 is not served, and with provider=openshift it + # refuses off OpenShift. A cluster that has neither sets this to false. + enabled: true # Provider responsible for issuing serving certificates. One of: openshift, cert-manager. Required if enabled. provider: cert-manager @@ -731,5 +750,9 @@ tls: # secrets from the cert-manager controller namespace). namespace: cert-manager - # Mutual TLS (clients present and verify certificates). Only allowed when provider == cert-manager. - mutual_enabled: false + # Mutual TLS (clients present and verify certificates). Only allowed when + # provider == cert-manager. On by default, with tls.enabled: the chart issues a + # client certificate for each of the operator, the two CSI plugins, Prometheus, + # and FoundationDB's peers, so there is nothing left for a deployment to + # provision before it can be required. + mutual_enabled: true diff --git a/operator/api/v1alpha2/controlplane_tls_test.go b/operator/api/v1alpha2/controlplane_tls_test.go new file mode 100644 index 000000000..dcac49631 --- /dev/null +++ b/operator/api/v1alpha2/controlplane_tls_test.go @@ -0,0 +1,73 @@ +// What the TLS block means when it is not there. +// +// The decision these answer is reached from objects that never went past +// admission: a reconciler builds a LocalControlPlane in a test, a client reads +// one through a cache, an older object predates the field. Defaulting is the API +// server's and applies to none of those, so the zero value has to be the closed +// one on its own rather than because something filled it in. + +package v1alpha2 + +import ( + "testing" + + "github.com/simplyblock/atlas/ptr" +) + +// An object that says nothing about TLS serves it, which is the whole point of +// spelling the toggles as the disable. +func TestAControlPlaneThatSaysNothingServesTLS(t *testing.T) { + var local LocalControlPlane + if !local.ServesTLS() { + t.Error("a control plane with no TLS block came up in plaintext") + } + if !local.RequiresClientCertificate() { + t.Error("a control plane with no TLS block admitted anonymous callers") + } + if got := local.TLSProvider(); got != ControlPlaneTLSCertManager { + t.Errorf("the issuer defaulted to %q", got) + } +} + +// A nil spec is the same answer. It is reachable: source.local is a pointer, and +// a managed deployment leaves it unset. +func TestANilLocalControlPlaneStillReadsAsClosed(t *testing.T) { + var local *LocalControlPlane + if !local.ServesTLS() || !local.RequiresClientCertificate() { + t.Error("a nil local control plane read as plaintext") + } + if got := local.TLSProvider(); got != ControlPlaneTLSCertManager { + t.Errorf("the issuer of a nil spec is %q", got) + } +} + +// Each toggle is honored on its own, and the narrower one leaves the serving up. +func TestEachToggleIsHonored(t *testing.T) { + anonymous := LocalControlPlane{TLS: ControlPlaneTLS{EnableMutualTLS: ptr.To(false)}} + if !anonymous.ServesTLS() { + t.Error("dropping the client certificate also dropped the serving") + } + if anonymous.RequiresClientCertificate() { + t.Error("a client certificate is still required after it was disabled") + } +} + +// Mutual TLS without serving TLS is not a state to represent, so disabling the +// serving disables it rather than leaving the two to contradict each other. +func TestPlaintextNeverAsksForAClientCertificate(t *testing.T) { + plaintext := LocalControlPlane{TLS: ControlPlaneTLS{EnableTLS: ptr.To(false)}} + if plaintext.ServesTLS() { + t.Fatal("the serving was disabled and the control plane still serves TLS") + } + if plaintext.RequiresClientCertificate() { + t.Error("a plaintext control plane asks its callers for a certificate") + } +} + +// An explicitly named issuer is not overwritten by the default. +func TestANamedIssuerIsKept(t *testing.T) { + openshift := LocalControlPlane{TLS: ControlPlaneTLS{Provider: ControlPlaneTLSOpenShift}} + if got := openshift.TLSProvider(); got != ControlPlaneTLSOpenShift { + t.Errorf("the named issuer became %q", got) + } +} diff --git a/operator/api/v1alpha2/controlplane_types.go b/operator/api/v1alpha2/controlplane_types.go index 1001db524..5818dd96d 100644 --- a/operator/api/v1alpha2/controlplane_types.go +++ b/operator/api/v1alpha2/controlplane_types.go @@ -89,6 +89,85 @@ type FoundationDBSpec struct { Resources corev1.ResourceRequirements `json:"resources,omitempty"` } +// ControlPlaneTLSProvider is who issues the control plane's serving certificate. +// +// The two values are each product's own spelling rather than this group's +// PascalCase, because they are not this group's words: they are what +// internal/utils/tls.go already matches on and what the control-plane image +// reads out of SB_TLS_PROVIDER. +// +kubebuilder:validation:Enum=cert-manager;OpenShift +type ControlPlaneTLSProvider string + +const ( + // ControlPlaneTLSCertManager issues through cert-manager, from the CA the + // deployment's ClusterIssuer mints. + ControlPlaneTLSCertManager ControlPlaneTLSProvider = "cert-manager" + + // ControlPlaneTLSOpenShift issues through the OpenShift service CA, which + // signs from a Service annotation rather than from an object of its own. + ControlPlaneTLSOpenShift ControlPlaneTLSProvider = "OpenShift" +) + +// ControlPlaneTLS is how a control plane this operator installs serves, and what +// it requires of the clients that reach it. +// +// The field names are DriverTLS's, because they are the same three decisions +// about the same connection seen from its two ends, and a reader who knows one +// should not have to learn the other. What differs is the default: the CSI +// driver's fields were added to describe deployments that already existed and +// default off, and these default on. An installed control plane holds every +// cluster definition, node registration, and volume record, and is reached over +// the pod network by the operator, the CSI driver, and the metrics scrape alike, +// so plaintext is a decision to state rather than the state a deployment lands +// in by leaving the block out. +type ControlPlaneTLS struct { + // EnableTLS serves the management API over TLS. Unset is on. + // +kubebuilder:default=true + // +optional + EnableTLS *bool `json:"enableTLS,omitempty"` + + // EnableMutualTLS additionally requires a caller to present a certificate + // of its own, rather than reaching the control plane anonymously over the + // encrypted connection EnableTLS alone provides. Ignored when EnableTLS is + // false, the same as on DriverTLS. Unset is on. + // +kubebuilder:default=true + // +optional + EnableMutualTLS *bool `json:"enableMutualTLS,omitempty"` + + // Provider issues the serving certificate. + // +kubebuilder:default=cert-manager + // +optional + Provider ControlPlaneTLSProvider `json:"provider,omitempty"` +} + +// ServesTLS reports whether the installed control plane serves over TLS. +// +// It answers from the unset value rather than from a defaulted field, because an +// object built in a test or read through a cache never went past admission, and +// a TLS decision that depends on the API server having run is one that reads as +// plaintext exactly where it is not checked. +func (l *LocalControlPlane) ServesTLS() bool { + return l == nil || l.TLS.EnableTLS == nil || *l.TLS.EnableTLS +} + +// RequiresClientCertificate reports whether a caller has to present one. Mutual +// TLS is meaningless without serving TLS, so disabling the serving disables this +// too rather than leaving the two to contradict each other. +func (l *LocalControlPlane) RequiresClientCertificate() bool { + if !l.ServesTLS() { + return false + } + return l == nil || l.TLS.EnableMutualTLS == nil || *l.TLS.EnableMutualTLS +} + +// TLSProvider is the issuer to use, with the default applied. +func (l *LocalControlPlane) TLSProvider() ControlPlaneTLSProvider { + if l == nil || l.TLS.Provider == "" { + return ControlPlaneTLSCertManager + } + return l.TLS.Provider +} + // LocalControlPlane is a control plane this cluster hosts, installed and owned // by the operator. Its objects carry a controller reference to the ControlPlane, // so the ownership spine starts at a real edge rather than at a Helm release. @@ -134,6 +213,15 @@ type LocalControlPlane struct { // value already written down. // +optional NodeSelector map[string]string `json:"nodeSelector,omitempty"` + + // TLS is how this control plane serves and what it asks of its callers. + // + // The block is defaulted to its own zero value rather than left absent, so + // that the field defaults inside it are applied to an object that does not + // mention TLS at all. + // +kubebuilder:default={} + // +optional + TLS ControlPlaneTLS `json:"tls,omitempty"` } // ManagedControlPlane is a control plane somewhere else, which this cluster's diff --git a/operator/api/v1alpha2/zz_generated.deepcopy.go b/operator/api/v1alpha2/zz_generated.deepcopy.go index de0cf0e1a..2a0d1c58a 100644 --- a/operator/api/v1alpha2/zz_generated.deepcopy.go +++ b/operator/api/v1alpha2/zz_generated.deepcopy.go @@ -593,6 +593,31 @@ func (in *ControlPlaneStatus) DeepCopy() *ControlPlaneStatus { return out } +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *ControlPlaneTLS) DeepCopyInto(out *ControlPlaneTLS) { + *out = *in + if in.EnableTLS != nil { + in, out := &in.EnableTLS, &out.EnableTLS + *out = new(bool) + **out = **in + } + if in.EnableMutualTLS != nil { + in, out := &in.EnableMutualTLS, &out.EnableMutualTLS + *out = new(bool) + **out = **in + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new ControlPlaneTLS. +func (in *ControlPlaneTLS) DeepCopy() *ControlPlaneTLS { + if in == nil { + return nil + } + out := new(ControlPlaneTLS) + in.DeepCopyInto(out) + return out +} + // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *CreatorReference) DeepCopyInto(out *CreatorReference) { *out = *in @@ -904,6 +929,7 @@ func (in *LocalControlPlane) DeepCopyInto(out *LocalControlPlane) { (*out)[key] = val } } + in.TLS.DeepCopyInto(&out.TLS) } // DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new LocalControlPlane. diff --git a/operator/config/crd/bases/storage.simplyblock.io_controlplanes.yaml b/operator/config/crd/bases/storage.simplyblock.io_controlplanes.yaml index 140c3da8f..95afa9396 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_controlplanes.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_controlplanes.yaml @@ -333,6 +333,35 @@ spec: More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ type: object type: object + tls: + default: {} + description: |- + TLS is how this control plane serves and what it asks of its callers. + + The block is defaulted to its own zero value rather than left absent, so + that the field defaults inside it are applied to an object that does not + mention TLS at all. + properties: + enableMutualTLS: + default: true + description: |- + EnableMutualTLS additionally requires a caller to present a certificate + of its own, rather than reaching the control plane anonymously over the + encrypted connection EnableTLS alone provides. Ignored when EnableTLS is + false, the same as on DriverTLS. Unset is on. + type: boolean + enableTLS: + default: true + description: EnableTLS serves the management API over TLS. Unset is on. + type: boolean + provider: + default: cert-manager + description: Provider issues the serving certificate. + enum: + - cert-manager + - OpenShift + type: string + type: object tolerations: description: |- Tolerations are applied to every pod the operator installs for the control diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index 13cdcceba..a9cf29751 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -1497,6 +1497,36 @@ spec: More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ type: object type: object + tls: + default: {} + description: |- + TLS is how this control plane serves and what it asks of its callers. + + The block is defaulted to its own zero value rather than left absent, so + that the field defaults inside it are applied to an object that does not + mention TLS at all. + properties: + enableMutualTLS: + default: true + description: |- + EnableMutualTLS additionally requires a caller to present a certificate + of its own, rather than reaching the control plane anonymously over the + encrypted connection EnableTLS alone provides. Ignored when EnableTLS is + false, the same as on DriverTLS. Unset is on. + type: boolean + enableTLS: + default: true + description: EnableTLS serves the management API over + TLS. Unset is on. + type: boolean + provider: + default: cert-manager + description: Provider issues the serving certificate. + enum: + - cert-manager + - OpenShift + type: string + type: object tolerations: description: |- Tolerations are applied to every pod the operator installs for the control diff --git a/operator/docs/designs/crd-redesign/design-controlplane.md b/operator/docs/designs/crd-redesign/design-controlplane.md index b20b5fed9..b36d79e81 100644 --- a/operator/docs/designs/crd-redesign/design-controlplane.md +++ b/operator/docs/designs/crd-redesign/design-controlplane.md @@ -531,7 +531,7 @@ than by one detection rule. | `FoundationDBCluster` CRDs | A prerequisite. `Installing` holds and names it where `apps.foundationdb.org/v1beta2` is not served | | FoundationDB controller | Always applied, under the name the chart gave it, in the `ControlPlane`'s namespace | | Document store | None. A base deployment runs no MongoDB, and the object store takes the step's place | -| Certificate issuer | Not expressible. §12 Q2 records that TLS stays with the chart | +| Certificate issuer | `spec.source.local.tls.provider`, declared rather than detected, and defaulting to cert-manager | | `StorageClass` | Not detected. An unset class falls to the cluster's default (§12 Q7) | **The CRDs and the controller are split because the API server cannot tell them @@ -991,11 +991,47 @@ and the log collector stay with the chart: none appears in a step of §4.2's machine, every one is non-essential in §4.3's table, and they are gated behind one chart value this kind has no field for. -**TLS is the one configuration the install cannot express.** The chart serves the -control plane over TLS behind `tls.enabled`, and `LocalControlPlane` has no field -for it. The chart therefore refuses `tls.enabled` together with the `standalone` -profile rather than installing a control plane in plaintext, and closing that gap -is a field on this kind together with the issuer detection §5.1 does not perform. +**TLS is `spec.source.local.tls`, and it is on by default.** The block is +`DriverTLS`'s three fields under `DriverTLS`'s names — `enableTLS`, +`enableMutualTLS`, and `provider` — because they are the same three decisions +about the same connection seen from its two ends. What differs is the default: +the CSI driver's fields describe deployments that already existed and default +off, and these default on. An installed control plane holds every cluster +definition, node registration, and volume record, and is reached over the pod +network by the operator, the CSI driver, and the metrics scrape alike, so +plaintext is a decision to state rather than the one a deployment reaches by +leaving the block out. + +The install turns it into the environment the control-plane image has always +spoken — `SB_TLS_SERVE`, `SB_TLS_PROVIDER`, `SB_TLS_CONNECT`, +`SB_TLS_CLIENT_AUTH`, and FoundationDB's `FDB_TLS_*` — on every pod of the +install rather than only the one that serves. The task and monitoring pools are +clients of the management API, and a plaintext pool cannot reach an API that +requires a certificate. + +The serving certificate is the install's too. It was rendered beside the Service +it certifies, and the install that replaced the chart took the Service and left +the certificate, so every pod mounted a Secret nothing produced. The two issuers +produce it differently and neither is a choice this makes: cert-manager takes a +`Certificate` naming the Service's DNS names, and the OpenShift service CA takes +an annotation on the Service. + +The chart renders the block from `tls.enabled` and `tls.mutual_enabled` rather +than leaving it out, so a release gets the answer its values state rather than +this kind's default. Both of those values are on as well, which makes TLS what a +plain `helm install` produces. It has a prerequisite: cert-manager issues the +certificates, and the chart refuses to render where `cert-manager.io/v1` is not +served rather than installing something quieter than what the values say. The +refusal names both ways out, because a deployment reaching it did not ask for the +thing it is being refused. + +FoundationDB's own connections are the second listener, and they follow the +client-certificate decision rather than the serving one: peer TLS between +database processes has no anonymous mode, so every process authenticates to every +other or none of them do. `mainContainer.enableTls` is what turns those listeners +over, and it is a field on the cluster rather than an environment variable -- +processes carrying a current certificate, key, and CA and no `enableTls` talk to +each other in the clear while every mount looks right. **Q3: Whether backup belongs to the action or to the spec.** `FoundationDBBackup` describes a continuous backup, carrying a `backupState` and a diff --git a/operator/docs/tests/test-plan-controlplane.md b/operator/docs/tests/test-plan-controlplane.md index e313915c7..b90e4b833 100644 --- a/operator/docs/tests/test-plan-controlplane.md +++ b/operator/docs/tests/test-plan-controlplane.md @@ -155,6 +155,28 @@ would be the old `Degraded` under a new name. | U-145 | A default `StorageClass` exists: it is used, and no provisioner is applied | Positive | — | | U-146 | No default `StorageClass`: `Installing` holds rather than applying hostpath | Boundary | — | +### TLS (design §5.1, §12 Q2) + +Files: `operator/api/v1alpha2/controlplane_tls_test.go`, +`operator/internal/controllers/controlplane/tls_test.go`, +`operator/internal/webapi/controlplane_transport_test.go` + +| # | Scenario | Type | Test | +|-------|----------------------------------------------------------------------------------|------------|--------------------------------------------------------| +| U-149 | A spec with no `tls` block serves TLS and requires a client certificate | Boundary | `TestAControlPlaneThatSaysNothingServesTLS` | +| U-150 | A nil `source.local` reads as closed rather than as plaintext | Boundary | `TestANilLocalControlPlaneStillReadsAsClosed` | +| U-151 | `enableMutualTLS: false` leaves the listener serving and the caller anonymous | Positive | `TestEachToggleIsHonored` | +| U-152 | `enableTLS: false` never asks a caller for a certificate | Negative | `TestPlaintextNeverAsksForAClientCertificate` | +| U-153 | A named issuer is not overwritten by the default | Positive | `TestANamedIssuerIsKept` | +| U-154 | The installed management API carries `SB_TLS_SERVE` and mounts its certificate | Positive | `TestTheInstalledControlPlaneServesTLSByDefault` | +| U-155 | Mutual TLS carries the `FDB_TLS_*` trio, and anonymous TLS carries none of it | Boundary | `TestMutualTLSCarriesTheFoundationDBPeerFiles` | +| U-156 | A plaintext install carries no TLS environment, mount, or volume at all | Negative | `TestDisablingServingInstallsThePlaintextControlPlane` | +| U-157 | Every workload of the install carries the decision, not only the one that serves | Regression | `TestEveryControlPlaneWorkloadCarriesTheDecision` | +| U-158 | The OpenShift issuer projects its bundle beside the Secret | Positive | `TestTheOpenShiftIssuerProjectsItsBundle` | +| U-159 | The atlas-lib client is handed the startup client's verified connection | Regression | `TestTheAtlasClientIsGivenTheSameConnection` | +| U-160 | A plaintext deployment hands the atlas-lib client no transport | Negative | `TestAPlaintextDeploymentHandsOverNoTransport` | +| U-161 | Nothing the install mounts is a Secret no step of it issues | Regression | `TestNoWorkloadWaitsOnASecretTheInstallDoesNotCreate` | + ### Deletion (design §4.4) | # | Scenario | Type | Test | diff --git a/operator/internal/controllers/controlplane/controlplane_controller.go b/operator/internal/controllers/controlplane/controlplane_controller.go index af28868bf..8497e627c 100644 --- a/operator/internal/controllers/controlplane/controlplane_controller.go +++ b/operator/internal/controllers/controlplane/controlplane_controller.go @@ -299,7 +299,7 @@ func (r *ControlPlaneReconciler) performInstallStep( return true, "", applyAll(ctx, r.Client, cp, r.Scheme, managementAPIObjects(cp)) case stepAwaitingAPI: - ok, message := r.probe(ctx, cp.Namespace, managedAccess{endpoint: localEndpoint(cp.Namespace)}) + ok, message := r.probe(ctx, cp.Namespace, managedAccess{endpoint: localEndpoint(cp)}) if !ok { return false, fmt.Sprintf("the management API is not answering yet: %s", message), nil } @@ -324,7 +324,7 @@ func (r *ControlPlaneReconciler) steadyState( // A managed control plane is reached on the Service this install created, so // there is no token to present and no CA beyond the cluster's own. - access := managedAccess{endpoint: localEndpoint(cp.Namespace)} + access := managedAccess{endpoint: localEndpoint(cp)} ok, message := r.probe(ctx, cp.Namespace, access) components, err := observe(ctx, r.Client, cp.Namespace) diff --git a/operator/internal/controllers/controlplane/controlplaneops_controller.go b/operator/internal/controllers/controlplane/controlplaneops_controller.go index d70f09c1c..a799d16d6 100644 --- a/operator/internal/controllers/controlplane/controlplaneops_controller.go +++ b/operator/internal/controllers/controlplane/controlplaneops_controller.go @@ -541,7 +541,7 @@ func (r *ControlPlaneOpsReconciler) verify( ) (bool, string, error) { endpoint := target.Status.Endpoint if endpoint == "" { - endpoint = localEndpoint(target.Namespace) + endpoint = localEndpoint(target) } prober := r.Prober diff --git a/operator/internal/controllers/controlplane/doc.go b/operator/internal/controllers/controlplane/doc.go index e1e1fe07d..a4fb78957 100644 --- a/operator/internal/controllers/controlplane/doc.go +++ b/operator/internal/controllers/controlplane/doc.go @@ -44,14 +44,21 @@ // operator where a cluster has none, and the answer this package reaches is that // a base control plane does not need one. // -// # What is not here yet +// # TLS +// +// spec.source.local.tls is what the install reads, and it defaults to serving +// TLS and requiring a client certificate. tls.go turns it into the environment +// the control-plane image has always spoken -- SB_TLS_SERVE, SB_TLS_PROVIDER, +// SB_TLS_CONNECT, SB_TLS_CLIENT_AUTH, and FoundationDB's FDB_TLS_* -- onto every +// pod of the install rather than only the one that serves, because the task and +// monitoring pools are clients of the management API and a plaintext pool cannot +// reach an API that requires a certificate. // -// TLS. The chart serves the control plane over TLS behind tls.enabled, which has -// no field in the ControlPlane spec design-controlplane.md settles, and §5.1 -// says only that the issuer is detected rather than declared. The install this -// package performs is the plaintext one, which is the chart's default and what -// the reference deployment runs. The chart refuses to hand over a TLS-enabled -// deployment rather than quietly installing it without TLS. +// The serving certificate is applied here as well. It was the chart's, rendered +// beside the Service it certifies, and the install that replaced the chart took +// the Service without it, so every pod mounted a Secret nothing produced. +// +// # What is not here yet // // Adoption. An install that meets objects a Helm release already created takes // them over by server-side apply under a stable field manager, which is the same diff --git a/operator/internal/controllers/controlplane/endpoint.go b/operator/internal/controllers/controlplane/endpoint.go index a1a306cdf..f50344f3c 100644 --- a/operator/internal/controllers/controlplane/endpoint.go +++ b/operator/internal/controllers/controlplane/endpoint.go @@ -42,8 +42,20 @@ var credentialKeys = []string{"token", "secret"} var caBundleKeys = []string{"ca.crt", "tls.crt"} // localEndpoint is where the management API this install created answers. -func localEndpoint(namespace string) string { - return fmt.Sprintf("http://%s.%s.svc.cluster.local:%d", ComponentWebAPI, namespace, webAPIPort) +// +// The scheme follows the install rather than being fixed, which is the whole of +// the defect this replaces: the address was the plaintext scheme whatever the +// deployment asked for, so an install that served TLS was one this operator could +// no longer reach. Every control-plane call in the operator resolves through +// status.endpoint, which this is published as, so the scheme reaches all of them +// at once. +func localEndpoint(cp *simplyblockv1alpha2.ControlPlane) string { + scheme := "http" + if cp.Spec.Source.Local.ServesTLS() { + scheme = "https" + } + return fmt.Sprintf("%s://%s.%s.svc.cluster.local:%d", + scheme, ComponentWebAPI, cp.Namespace, webAPIPort) } // credentialsError is what a Secret that is missing or unusable produces. It is diff --git a/operator/internal/controllers/controlplane/foundationdb.go b/operator/internal/controllers/controlplane/foundationdb.go index 6df98244f..7dbab3929 100644 --- a/operator/internal/controllers/controlplane/foundationdb.go +++ b/operator/internal/controllers/controlplane/foundationdb.go @@ -91,7 +91,8 @@ const fdbCoordinatorDisk = "10G" // the controller, then the cluster the controller reconciles. func foundationDBObjects(cp *simplyblockv1alpha2.ControlPlane) []client.Object { ns := cp.Namespace - return []client.Object{ + //nolint:prealloc // the literal is the declaration; the append below is the conditional set + objects := []client.Object{ serviceAccount(ns, fdbOperatorServiceAccount), serviceAccount(ns, fdbPodServiceAccount), fdbManagerRoleObject(), @@ -103,6 +104,7 @@ func foundationDBObjects(cp *simplyblockv1alpha2.ControlPlane) []client.Object { fdbOperatorDeployment(cp), foundationDBCluster(cp), } + return append(objects, fdbPeerCertificate(cp)...) } // foundationDBClusterScoped is the half of that set the garbage collector will @@ -241,6 +243,7 @@ func fdbOperatorDeployment(cp *simplyblockv1alpha2.ControlPlane) *appsv1.Deploym logsVolume = "logs" ) labels := map[string]string{appLabel: ComponentFDBOperator} + peerTLS := fdbPeerTLS(cp) spec := corev1.PodSpec{ ServiceAccountName: fdbOperatorServiceAccount, @@ -249,11 +252,11 @@ func fdbOperatorDeployment(cp *simplyblockv1alpha2.ControlPlane) *appsv1.Deploym RunAsGroup: ptr.To(int64(4059)), FSGroup: ptr.To(int64(4059)), }, - Volumes: []corev1.Volume{ + Volumes: append([]corev1.Volume{ {Name: tmpVolume, VolumeSource: corev1.VolumeSource{EmptyDir: &corev1.EmptyDirVolumeSource{}}}, {Name: logsVolume, VolumeSource: corev1.VolumeSource{EmptyDir: &corev1.EmptyDirVolumeSource{}}}, {Name: binariesVolume, VolumeSource: corev1.VolumeSource{EmptyDir: &corev1.EmptyDirVolumeSource{}}}, - }, + }, peerVolumeIf(peerTLS)...), InitContainers: []corev1.Container{{ Name: "foundationdb-kubernetes-init-7-3", Image: fdbMonitorImage, @@ -274,12 +277,14 @@ func fdbOperatorDeployment(cp *simplyblockv1alpha2.ControlPlane) *appsv1.Deploym Image: fdbOperatorImage, Command: []string{"/manager"}, Args: []string{"--health-probe-bind-address=:9443"}, - Env: []corev1.EnvVar{{ + // The operator reconciles the database, so it reaches it the way its + // processes reach each other and needs the same material. + Env: append([]corev1.EnvVar{{ Name: "WATCH_NAMESPACE", ValueFrom: &corev1.EnvVarSource{ FieldRef: &corev1.ObjectFieldSelector{FieldPath: "metadata.namespace"}, }, - }}, + }}, peerEnvIf(peerTLS)...), Ports: []corev1.ContainerPort{{Name: "metrics", ContainerPort: 8080}}, Resources: corev1.ResourceRequirements{ Requests: corev1.ResourceList{ @@ -296,11 +301,11 @@ func fdbOperatorDeployment(cp *simplyblockv1alpha2.ControlPlane) *appsv1.Deploym AllowPrivilegeEscalation: ptr.To(false), Privileged: ptr.To(false), }, - VolumeMounts: []corev1.VolumeMount{ + VolumeMounts: append([]corev1.VolumeMount{ {Name: tmpVolume, MountPath: "/tmp"}, {Name: logsVolume, MountPath: "/var/log/fdb"}, {Name: binariesVolume, MountPath: "/usr/bin/fdb"}, - }, + }, peerMountIf(peerTLS)...), }}, TerminationGracePeriodSeconds: ptr.To(int64(10)), } @@ -337,12 +342,13 @@ func foundationDBCluster(cp *simplyblockv1alpha2.ControlPlane) *unstructured.Uns // The class templates are identical apart from their anti-affinity, and the // FoundationDB operator replaces general.podTemplate wholesale when a class // override exists, so each one repeats what general already said. + peerTLS := fdbPeerTLS(cp) classTemplate := func(class string) map[string]any { template := map[string]any{ "spec": map[string]any{ "serviceAccountName": fdbPodServiceAccount, "containers": []any{ - fdbContainer(fdb), + fdbContainer(fdb, peerTLS), }, "affinity": map[string]any{ "podAntiAffinity": map[string]any{ @@ -364,6 +370,12 @@ func foundationDBCluster(cp *simplyblockv1alpha2.ControlPlane) *unstructured.Uns if local := cp.Spec.Source.Local; local != nil && len(local.NodeSelector) > 0 { template["spec"].(map[string]any)["nodeSelector"] = toAnyMap(local.NodeSelector) } + if peerTLS { + // The operator replaces general.podTemplate wholesale for a class + // that overrides it, so the volume is repeated here rather than + // inherited. + template["spec"].(map[string]any)["volumes"] = fdbPeerVolume() + } return map[string]any{"podTemplate": template} } @@ -372,7 +384,7 @@ func foundationDBCluster(cp *simplyblockv1alpha2.ControlPlane) *unstructured.Uns "podTemplate": map[string]any{ "spec": map[string]any{ "serviceAccountName": fdbPodServiceAccount, - "containers": []any{fdbContainer(fdb)}, + "containers": []any{fdbContainer(fdb, peerTLS)}, "initContainers": []any{map[string]any{ "name": "foundationdb-kubernetes-init", "resources": map[string]any{ @@ -391,6 +403,9 @@ func foundationDBCluster(cp *simplyblockv1alpha2.ControlPlane) *unstructured.Uns general["podTemplate"].(map[string]any)["spec"].(map[string]any)["nodeSelector"] = toAnyMap(local.NodeSelector) } + if peerTLS { + general["podTemplate"].(map[string]any)["spec"].(map[string]any)["volumes"] = fdbPeerVolume() + } obj := &unstructured.Unstructured{Object: map[string]any{ "spec": map[string]any{ @@ -425,8 +440,8 @@ func foundationDBCluster(cp *simplyblockv1alpha2.ControlPlane) *unstructured.Uns "log": classTemplate("log"), }, "routing": map[string]any{"defineDNSLocalityFields": true}, - "mainContainer": map[string]any{"imageConfigs": []any{map[string]any{"baseImage": fdbMonitorImageRepository()}}}, - "sidecarContainer": map[string]any{"enableLivenessProbe": true, "enableReadinessProbe": false}, + "mainContainer": mainContainerSpec(peerTLS), + "sidecarContainer": sidecarContainerSpec(peerTLS), }, }} obj.SetGroupVersionKind(fdbClusterGVK) @@ -439,7 +454,7 @@ func foundationDBCluster(cp *simplyblockv1alpha2.ControlPlane) *unstructured.Uns // process. It runs as root because the FoundationDB image's data directory is // owned by it, which is the upstream image's arrangement rather than a choice // this install makes. -func fdbContainer(fdb *simplyblockv1alpha2.FoundationDBSpec) map[string]any { +func fdbContainer(fdb *simplyblockv1alpha2.FoundationDBSpec, peerTLS bool) map[string]any { requests := map[string]any{"cpu": "100m", "memory": "1Gi"} limits := map[string]any{"cpu": "500m", "memory": "4Gi"} if fdb != nil { @@ -456,11 +471,16 @@ func fdbContainer(fdb *simplyblockv1alpha2.FoundationDBSpec) map[string]any { limits["memory"] = q.String() } } - return map[string]any{ + container := map[string]any{ "name": "foundationdb", "resources": map[string]any{"requests": requests, "limits": limits}, "securityContext": map[string]any{"runAsUser": int64(0)}, } + if peerTLS { + container["env"] = fdbPeerEnv() + container["volumeMounts"] = fdbPeerMount() + } + return container } // volumeClaimSpec is what each FoundationDB process claims. An unset storage @@ -590,3 +610,52 @@ func (h fdbHealth) waitingOn() string { return "" } } + +// mainContainerSpec and sidecarContainerSpec carry the switch that turns +// FoundationDB's own listeners to TLS. +// +// It is a field on the cluster rather than an environment variable, and it is +// the half that matters: the FDB_TLS_* the pod templates carry is only the +// material, and processes with the material and no enableTls talk to each other +// in the clear while every certificate is mounted and current. +func mainContainerSpec(peerTLS bool) map[string]any { + spec := map[string]any{ + "imageConfigs": []any{map[string]any{"baseImage": fdbMonitorImageRepository()}}, + } + if peerTLS { + spec["enableTls"] = true + } + return spec +} + +func sidecarContainerSpec(peerTLS bool) map[string]any { + spec := map[string]any{"enableLivenessProbe": true, "enableReadinessProbe": false} + if peerTLS { + spec["enableTls"] = true + } + return spec +} + +// The three conditional halves of the FoundationDB operator's own peer TLS, +// written as helpers so the Deployment reads as one literal rather than as four +// branches around it. +func peerEnvIf(peerTLS bool) []corev1.EnvVar { + if !peerTLS { + return nil + } + return fdbOperatorPeerEnv() +} + +func peerVolumeIf(peerTLS bool) []corev1.Volume { + if !peerTLS { + return nil + } + return fdbOperatorPeerVolume() +} + +func peerMountIf(peerTLS bool) []corev1.VolumeMount { + if !peerTLS { + return nil + } + return fdbOperatorPeerMount() +} diff --git a/operator/internal/controllers/controlplane/install_test.go b/operator/internal/controllers/controlplane/install_test.go index d254e0cce..536c3e1bd 100644 --- a/operator/internal/controllers/controlplane/install_test.go +++ b/operator/internal/controllers/controlplane/install_test.go @@ -125,7 +125,7 @@ func TestReApplyingAStepCorrectsWhatWasChangedUnderIt(t *testing.T) { if err := c.Update(ctx, &api); err != nil { t.Fatalf("scale the management API down: %v", err) } - if err := c.Delete(ctx, webAPIService(testNamespace)); err != nil { + if err := c.Delete(ctx, webAPIService(cp)); err != nil { t.Fatalf("delete the Service: %v", err) } diff --git a/operator/internal/controllers/controlplane/managementapi.go b/operator/internal/controllers/controlplane/managementapi.go index 859679bde..d03ef131f 100644 --- a/operator/internal/controllers/controlplane/managementapi.go +++ b/operator/internal/controllers/controlplane/managementapi.go @@ -52,7 +52,8 @@ var restartOnClusterFileChange = map[string]string{ // of the one that serves. func managementAPIObjects(cp *simplyblockv1alpha2.ControlPlane) []client.Object { ns := cp.Namespace - return []client.Object{ + //nolint:prealloc // the literal is the declaration; the append below is the conditional set + objects := []client.Object{ serviceAccount(ns, serviceAccountName), sharedConfigMap(ns), controlPlaneClusterRole(), @@ -60,13 +61,14 @@ func managementAPIObjects(cp *simplyblockv1alpha2.ControlPlane) []client.Object serviceReaderClusterRole(), serviceReaderClusterRoleBinding(ns), webAPIDeployment(cp), - webAPIService(ns), + webAPIService(cp), tasksDeployment(cp), monitoringDeployment(cp), adminControlDeployment(cp), fdbExporterDeployment(cp), fdbExporterService(ns), } + return append(objects, servingCertificateObjects(cp)...) } // managementAPIClusterScoped is the half of that set the garbage collector will @@ -231,6 +233,7 @@ func webAPIDeployment(cp *simplyblockv1alpha2.ControlPlane) *appsv1.Deployment { Value: "system:serviceaccount:" + cp.Namespace + ":simplyblock-prometheus"}, } env = append(env, prometheusEnv()...) + env = append(env, tlsEnv(managed)...) spec := corev1.PodSpec{ ServiceAccountName: serviceAccountName, @@ -242,10 +245,11 @@ func webAPIDeployment(cp *simplyblockv1alpha2.ControlPlane) *appsv1.Deployment { Command: []string{"python3", "simplyblock_web/app.py"}, Ports: []corev1.ContainerPort{{ContainerPort: webAPIPort}}, Env: env, - VolumeMounts: []corev1.VolumeMount{clusterFileMount()}, + VolumeMounts: append([]corev1.VolumeMount{clusterFileMount()}, tlsMount(managed)...), Resources: webAPIResources(managed), }}, - Volumes: []corev1.Volume{clusterFileVolumeSource()}, + Volumes: append([]corev1.Volume{clusterFileVolumeSource()}, + tlsVolume(managed, ServingCertSecret)...), } scheduling(managed, &spec) @@ -298,9 +302,13 @@ func webAPIResources(managed *simplyblockv1alpha2.LocalControlPlane) corev1.Reso // webAPIService is what status.endpoint resolves to, and the name the CSI // driver's configuration and the metrics scrape both carry. -func webAPIService(namespace string) *corev1.Service { +func webAPIService(cp *simplyblockv1alpha2.ControlPlane) *corev1.Service { return &corev1.Service{ - ObjectMeta: metav1.ObjectMeta{Name: ComponentWebAPI, Namespace: namespace}, + ObjectMeta: metav1.ObjectMeta{ + Name: ComponentWebAPI, + Namespace: cp.Namespace, + Annotations: servingCertAnnotations(cp), + }, Spec: corev1.ServiceSpec{ Selector: map[string]string{appLabel: ComponentWebAPI}, Ports: []corev1.ServicePort{ @@ -387,8 +395,9 @@ func servicePoolDeployment( // directly, which are host addresses rather than Service names. HostNetwork: true, DNSPolicy: corev1.DNSClusterFirstWithHostNet, - Containers: containers(services, localImage(cp), pullPolicyOf(managed)), - Volumes: []corev1.Volume{clusterFileVolumeSource()}, + Containers: containers(services, managed, localImage(cp), pullPolicyOf(managed)), + Volumes: append([]corev1.Volume{clusterFileVolumeSource()}, + tlsVolume(managed, ServingCertSecret)...), } scheduling(managed, &spec) @@ -427,6 +436,7 @@ func adminControlDeployment(cp *simplyblockv1alpha2.ControlPlane) *appsv1.Deploy logLevelEnv(), } env = append(env, prometheusEnv()...) + env = append(env, tlsEnv(managed)...) spec := corev1.PodSpec{ ServiceAccountName: serviceAccountName, @@ -442,7 +452,7 @@ func adminControlDeployment(cp *simplyblockv1alpha2.ControlPlane) *appsv1.Deploy // act on a signal while a foreground sleep is running. Command: []string{"/bin/bash", "-c", "trap : TERM INT; sleep infinity & wait"}, Env: env, - VolumeMounts: []corev1.VolumeMount{clusterFileMount()}, + VolumeMounts: append([]corev1.VolumeMount{clusterFileMount()}, tlsMount(managed)...), Resources: corev1.ResourceRequirements{ Requests: corev1.ResourceList{ corev1.ResourceCPU: resource.MustParse("200m"), @@ -454,7 +464,8 @@ func adminControlDeployment(cp *simplyblockv1alpha2.ControlPlane) *appsv1.Deploy }, }, }}, - Volumes: []corev1.Volume{clusterFileVolumeSource()}, + Volumes: append([]corev1.Volume{clusterFileVolumeSource()}, + tlsVolume(managed, ServingCertSecret)...), } scheduling(managed, &spec) diff --git a/operator/internal/controllers/controlplane/podspec.go b/operator/internal/controllers/controlplane/podspec.go index 3bddeee42..5ea79f8db 100644 --- a/operator/internal/controllers/controlplane/podspec.go +++ b/operator/internal/controllers/controlplane/podspec.go @@ -145,10 +145,17 @@ type service struct { } // container builds one service container from the shared shape. -func (s service) container(image string, pullPolicy corev1.PullPolicy) corev1.Container { +func (s service) container( + local *simplyblockv1alpha2.LocalControlPlane, image string, pullPolicy corev1.PullPolicy, +) corev1.Container { env := append([]corev1.EnvVar{}, s.extraEnv...) env = append(env, prometheusEnv()...) env = append(env, logLevelEnv()) + // These processes are clients of the management API rather than servers, and + // SB_TLS_CONNECT is what tells them so: a pool left plaintext while the API + // it calls requires a certificate is a control plane that cannot run its own + // tasks. + env = append(env, tlsEnv(local)...) return corev1.Container{ Name: s.name, @@ -156,7 +163,7 @@ func (s service) container(image string, pullPolicy corev1.PullPolicy) corev1.Co ImagePullPolicy: pullPolicy, Command: []string{"python3", s.module}, Env: env, - VolumeMounts: []corev1.VolumeMount{clusterFileMount()}, + VolumeMounts: append([]corev1.VolumeMount{clusterFileMount()}, tlsMount(local)...), Resources: serviceResources(), } } @@ -164,10 +171,15 @@ func (s service) container(image string, pullPolicy corev1.PullPolicy) corev1.Co // containers builds every service of a pool, in the order they are declared, so // that the apply produces a stable list rather than one that reorders between // passes and rolls the Deployment for nothing. -func containers(services []service, image string, pullPolicy corev1.PullPolicy) []corev1.Container { +func containers( + services []service, + local *simplyblockv1alpha2.LocalControlPlane, + image string, + pullPolicy corev1.PullPolicy, +) []corev1.Container { out := make([]corev1.Container, 0, len(services)) for _, s := range services { - out = append(out, s.container(image, pullPolicy)) + out = append(out, s.container(local, image, pullPolicy)) } return out } diff --git a/operator/internal/controllers/controlplane/tls.go b/operator/internal/controllers/controlplane/tls.go new file mode 100644 index 000000000..abdfce29c --- /dev/null +++ b/operator/internal/controllers/controlplane/tls.go @@ -0,0 +1,247 @@ +// What an installed control plane needs in order to serve TLS. +// +// The control-plane image already speaks it, and has for as long as the chart +// installed the control plane itself: SB_TLS_SERVE turns the listener on, +// SB_TLS_PROVIDER says who signed the certificate, SB_TLS_CONNECT says what the +// control plane's own outbound calls do, SB_TLS_CLIENT_AUTH says what it demands +// of a caller, and the FDB_TLS_* trio is FoundationDB's peer TLS. None of it was +// reachable once the install moved off the chart, because the ControlPlane spec +// had no field to carry the decision, so the install was the plaintext one +// whatever a deployment asked for. +// +// The certificates themselves are not minted here. They are cert-manager +// Certificates and an OpenShift service-CA annotation, and the names below are +// the ones the chart has always issued them under, for the same reason names.go +// gives: a running deployment refers to them from places this operator does not +// control. + +package controlplane + +import ( + corev1 "k8s.io/api/core/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/utils" +) + +const ( + // tlsVolumeName is the volume every control-plane pod mounts its serving + // material from, and tlsMountPath is where the image looks for it. + tlsVolumeName = "tls" + tlsMountPath = "/etc/simplyblock/tls" + + // ServingCertSecret holds the management API's serving certificate. It is + // the secret the chart's Certificate issues into and the one a deployment + // migrating off the chart already has. + ServingCertSecret = "simplyblock-webappapi-tls" + + // openShiftCAConfigMap is where the OpenShift service CA publishes the + // bundle its own certificates verify against. It arrives as a ConfigMap + // under a key of its own, so the projection below renames it to the ca.crt + // the image reads. + openShiftCAConfigMap = "simplyblock-certificate-authority" + openShiftCAKey = "service-ca.crt" +) + +// tlsEnv is the TLS half of a control-plane container's environment. +// +// SB_TLS_CONNECT is the pair's subtlety: it is what the control plane does when +// it is the client, and it follows the client-certificate decision rather than +// the serving one. A deployment that serves TLS without mutual TLS connects +// anonymously, because there is no certificate for it to present. +func tlsEnv(local *simplyblockv1alpha2.LocalControlPlane) []corev1.EnvVar { + if !local.ServesTLS() { + return nil + } + + env := []corev1.EnvVar{ + {Name: "SB_TLS_SERVE", Value: "1"}, + {Name: "SB_TLS_PROVIDER", Value: string(local.TLSProvider())}, + } + + if !local.RequiresClientCertificate() { + return append(env, corev1.EnvVar{Name: "SB_TLS_CONNECT", Value: "anonymous"}) + } + + return append(env, + corev1.EnvVar{Name: "SB_TLS_CLIENT_AUTH", Value: "required"}, + corev1.EnvVar{Name: "SB_TLS_CONNECT", Value: "authenticated"}, + corev1.EnvVar{Name: "FDB_TLS_CERTIFICATE_FILE", Value: tlsMountPath + "/tls.crt"}, + corev1.EnvVar{Name: "FDB_TLS_KEY_FILE", Value: tlsMountPath + "/tls.key"}, + corev1.EnvVar{Name: "FDB_TLS_CA_FILE", Value: tlsMountPath + "/ca.crt"}, + ) +} + +// tlsMount is the mount that goes with it, and nothing when the deployment is +// plaintext. +func tlsMount(local *simplyblockv1alpha2.LocalControlPlane) []corev1.VolumeMount { + if !local.ServesTLS() { + return nil + } + return []corev1.VolumeMount{{ + Name: tlsVolumeName, + MountPath: tlsMountPath, + ReadOnly: true, + }} +} + +// tlsVolume is the material itself, which differs by issuer. +// +// cert-manager writes the certificate, its key, and the issuing CA into one +// Secret, so the Secret is the volume. The OpenShift service CA writes only the +// certificate and key there and publishes its bundle as a ConfigMap, so the two +// are projected together and the bundle's key is renamed to the ca.crt the image +// reads from either provider. +func tlsVolume(local *simplyblockv1alpha2.LocalControlPlane, secret string) []corev1.Volume { + if !local.ServesTLS() { + return nil + } + + if local.TLSProvider() == simplyblockv1alpha2.ControlPlaneTLSOpenShift { + return []corev1.Volume{{ + Name: tlsVolumeName, + VolumeSource: corev1.VolumeSource{ + Projected: &corev1.ProjectedVolumeSource{ + Sources: []corev1.VolumeProjection{ + {Secret: &corev1.SecretProjection{ + LocalObjectReference: corev1.LocalObjectReference{Name: secret}, + }}, + {ConfigMap: &corev1.ConfigMapProjection{ + LocalObjectReference: corev1.LocalObjectReference{Name: openShiftCAConfigMap}, + Items: []corev1.KeyToPath{{ + Key: openShiftCAKey, Path: "ca.crt", + }}, + }}, + }, + }, + }, + }} + } + + return []corev1.Volume{{ + Name: tlsVolumeName, + VolumeSource: corev1.VolumeSource{ + Secret: &corev1.SecretVolumeSource{SecretName: secret}, + }, + }} +} + +// servingCertificateObjects is what the install has to create so that the +// certificate its pods mount exists. +// +// It is the operator's job because the Service is. On the chart this lived next +// to the Service it certifies, and the install that replaced the chart took the +// Service without taking the certificate, so every pod mounted a Secret nothing +// produced. +// +// The two issuers do it differently and neither is a choice made here. +// cert-manager takes a Certificate naming the Service's DNS names, and the +// OpenShift service CA takes an annotation on the Service itself, which is why +// this returns objects for one and annotations for the other. +func servingCertificateObjects(cp *simplyblockv1alpha2.ControlPlane) []client.Object { + local := cp.Spec.Source.Local + if !local.ServesTLS() || local.TLSProvider() != simplyblockv1alpha2.ControlPlaneTLSCertManager { + return nil + } + return []client.Object{ + utils.BuildServiceServingCertificate(cp.Namespace, ComponentWebAPI, ServingCertSecret), + } +} + +// servingCertAnnotations is the OpenShift half: the service CA signs from an +// annotation on the Service and writes the result into the Secret it names. +func servingCertAnnotations(cp *simplyblockv1alpha2.ControlPlane) map[string]string { + local := cp.Spec.Source.Local + return utils.ServingCertServiceAnnotations( + local.ServesTLS(), string(local.TLSProvider()), ServingCertSecret) +} + +// FoundationDB's peer TLS. The database's own connections are a separate +// listener from the management API's, with its own certificate, its own mount +// path, and a switch on the FoundationDBCluster rather than an environment +// variable: FDB_TLS_* supplies the material and mainContainer.enableTls is what +// makes the processes listen for it. +const ( + fdbTLSVolumeName = "tls-fdb" + fdbTLSMountPath = "/var/fdb/tls" + + // FDBPeerCertSecret is the certificate the database's processes present to + // each other. It carries both usages, because every process is a server to + // its peers and a client of them. + FDBPeerCertSecret = "simplyblock-foundationdb-tls" +) + +// fdbPeerTLS reports whether the database's own connections are encrypted. +// +// It follows the client-certificate decision rather than the serving one. Peer +// TLS between database processes has no anonymous mode to fall back to: every +// process authenticates to every other or none of them do. +func fdbPeerTLS(cp *simplyblockv1alpha2.ControlPlane) bool { + return cp.Spec.Source.Local.RequiresClientCertificate() +} + +// fdbPeerEnv is the material FoundationDB reads, as the unstructured shape the +// FoundationDBCluster's pod templates take. +func fdbPeerEnv() []any { + return []any{ + map[string]any{"name": "FDB_TLS_CERTIFICATE_FILE", "value": fdbTLSMountPath + "/tls.crt"}, + map[string]any{"name": "FDB_TLS_KEY_FILE", "value": fdbTLSMountPath + "/tls.key"}, + map[string]any{"name": "FDB_TLS_CA_FILE", "value": fdbTLSMountPath + "/ca.crt"}, + } +} + +// fdbPeerVolume is the Secret the processes read it from. +func fdbPeerVolume() []any { + return []any{map[string]any{ + "name": fdbTLSVolumeName, + "secret": map[string]any{"secretName": FDBPeerCertSecret}, + }} +} + +// fdbPeerMount is where each process finds it. +func fdbPeerMount() []any { + return []any{map[string]any{ + "name": fdbTLSVolumeName, "mountPath": fdbTLSMountPath, "readOnly": true, + }} +} + +// fdbOperatorPeerEnv is the same material for the FoundationDB operator itself, +// which reconciles the cluster and has to reach it the way its processes do. +func fdbOperatorPeerEnv() []corev1.EnvVar { + return []corev1.EnvVar{ + {Name: "FDB_TLS_CERTIFICATE_FILE", Value: fdbTLSMountPath + "/tls.crt"}, + {Name: "FDB_TLS_KEY_FILE", Value: fdbTLSMountPath + "/tls.key"}, + {Name: "FDB_TLS_CA_FILE", Value: fdbTLSMountPath + "/ca.crt"}, + } +} + +// fdbOperatorPeerVolume and fdbOperatorPeerMount are the typed halves of the +// same, for the operator's own Deployment. +func fdbOperatorPeerVolume() []corev1.Volume { + return []corev1.Volume{{ + Name: fdbTLSVolumeName, + VolumeSource: corev1.VolumeSource{ + Secret: &corev1.SecretVolumeSource{SecretName: FDBPeerCertSecret}, + }, + }} +} + +func fdbOperatorPeerMount() []corev1.VolumeMount { + return []corev1.VolumeMount{{ + Name: fdbTLSVolumeName, MountPath: fdbTLSMountPath, ReadOnly: true, + }} +} + +// fdbPeerCertificate is the Certificate that issues it, and nothing where the +// deployment does not use peer TLS or signs through the OpenShift service CA -- +// that CA signs from a Service annotation, and these processes have no Service. +func fdbPeerCertificate(cp *simplyblockv1alpha2.ControlPlane) []client.Object { + if !fdbPeerTLS(cp) || + cp.Spec.Source.Local.TLSProvider() != simplyblockv1alpha2.ControlPlaneTLSCertManager { + return nil + } + return []client.Object{ + utils.BuildServiceServingCertificate(cp.Namespace, ComponentFDBCluster, FDBPeerCertSecret), + } +} diff --git a/operator/internal/controllers/controlplane/tls_test.go b/operator/internal/controllers/controlplane/tls_test.go new file mode 100644 index 000000000..aeb523fc3 --- /dev/null +++ b/operator/internal/controllers/controlplane/tls_test.go @@ -0,0 +1,295 @@ +// What the install carries when a deployment asks for TLS. +// +// The assertions are on the pod spec rather than on the helpers, because the +// defect this covers was not a wrong value: it was that no value reached the +// workload at all. A helper that returns the right environment and is wired into +// three of four pod specs is the same outage as one that returns nothing. + +package controlplane + +import ( + "slices" + "testing" + + appsv1 "k8s.io/api/apps/v1" + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" + + "github.com/simplyblock/atlas/ptr" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// aLocalControlPlane is the singleton with the TLS block given. +func aLocalControlPlane(tls simplyblockv1alpha2.ControlPlaneTLS) *simplyblockv1alpha2.ControlPlane { + return &simplyblockv1alpha2.ControlPlane{ + ObjectMeta: metav1.ObjectMeta{Name: SingletonName, Namespace: "simplyblock"}, + Spec: simplyblockv1alpha2.ControlPlaneSpec{ + Source: simplyblockv1alpha2.ControlPlaneSource{ + Local: &simplyblockv1alpha2.LocalControlPlane{ + Image: "docker.io/simplyblock/simplyblock:main", + TLS: tls, + }, + }, + }, + } +} + +// envOf reads one variable out of a container, and says whether it was set. +func envOf(container corev1.Container, name string) (string, bool) { + for _, e := range container.Env { + if e.Name == name { + return e.Value, true + } + } + return "", false +} + +// mountsTLS reports whether the container mounts the serving material. +func mountsTLS(container corev1.Container) bool { + return slices.ContainsFunc(container.VolumeMounts, func(m corev1.VolumeMount) bool { + return m.Name == tlsVolumeName && m.MountPath == tlsMountPath + }) +} + +// A control plane that says nothing about TLS serves it, because the block's +// zero value is the closed one. +func TestTheInstalledControlPlaneServesTLSByDefault(t *testing.T) { + deployment := webAPIDeployment(aLocalControlPlane(simplyblockv1alpha2.ControlPlaneTLS{})) + container := deployment.Spec.Template.Spec.Containers[0] + + if got, ok := envOf(container, "SB_TLS_SERVE"); !ok || got != "1" { + t.Errorf("the management API was installed without a TLS listener (SB_TLS_SERVE=%q, set=%v)", got, ok) + } + if got, _ := envOf(container, "SB_TLS_PROVIDER"); got != string(simplyblockv1alpha2.ControlPlaneTLSCertManager) { + t.Errorf("the issuer reached the pod as %q", got) + } + if got, _ := envOf(container, "SB_TLS_CLIENT_AUTH"); got != "required" { + t.Errorf("mutual TLS is on by default and the pod asks for %q", got) + } + if got, _ := envOf(container, "SB_TLS_CONNECT"); got != "authenticated" { + t.Errorf("the control plane's own calls connect as %q", got) + } + if !mountsTLS(container) { + t.Error("the serving certificate is not mounted, so the listener has nothing to present") + } +} + +// Mutual TLS is the FoundationDB peer decision too, and it is the only thing +// that turns the FDB_TLS_ trio on. +func TestMutualTLSCarriesTheFoundationDBPeerFiles(t *testing.T) { + mutual := webAPIDeployment(aLocalControlPlane(simplyblockv1alpha2.ControlPlaneTLS{})) + for _, name := range []string{"FDB_TLS_CERTIFICATE_FILE", "FDB_TLS_KEY_FILE", "FDB_TLS_CA_FILE"} { + if _, ok := envOf(mutual.Spec.Template.Spec.Containers[0], name); !ok { + t.Errorf("%s is not set, so FoundationDB's peers talk in the clear", name) + } + } + + anonymous := webAPIDeployment(aLocalControlPlane( + simplyblockv1alpha2.ControlPlaneTLS{EnableMutualTLS: ptr.To(false)})) + if _, ok := envOf(anonymous.Spec.Template.Spec.Containers[0], "FDB_TLS_CERTIFICATE_FILE"); ok { + t.Error("a deployment with no client certificates still configured FoundationDB peer TLS") + } +} + +// Dropping the client certificate leaves the listener up and the caller +// anonymous, which is the narrower of the two retreats. +func TestDisablingMutualLeavesTheListenerServing(t *testing.T) { + deployment := webAPIDeployment(aLocalControlPlane( + simplyblockv1alpha2.ControlPlaneTLS{EnableMutualTLS: ptr.To(false)})) + container := deployment.Spec.Template.Spec.Containers[0] + + if got, _ := envOf(container, "SB_TLS_SERVE"); got != "1" { + t.Error("dropping the client certificate took the listener down with it") + } + if got, _ := envOf(container, "SB_TLS_CONNECT"); got != "anonymous" { + t.Errorf("a deployment with no certificate to present connects as %q", got) + } + if _, ok := envOf(container, "SB_TLS_CLIENT_AUTH"); ok { + t.Error("callers are still required to present a certificate") + } + if !mountsTLS(container) { + t.Error("the serving certificate is no longer mounted") + } +} + +// The plaintext install is still reachable, and it carries nothing at all. +func TestDisablingServingInstallsThePlaintextControlPlane(t *testing.T) { + deployment := webAPIDeployment(aLocalControlPlane( + simplyblockv1alpha2.ControlPlaneTLS{EnableTLS: ptr.To(false)})) + container := deployment.Spec.Template.Spec.Containers[0] + + for _, name := range []string{"SB_TLS_SERVE", "SB_TLS_PROVIDER", "SB_TLS_CONNECT", "SB_TLS_CLIENT_AUTH"} { + if _, ok := envOf(container, name); ok { + t.Errorf("a plaintext install still set %s", name) + } + } + if mountsTLS(container) { + t.Error("a plaintext install mounts a serving certificate") + } + if slices.ContainsFunc(deployment.Spec.Template.Spec.Volumes, func(v corev1.Volume) bool { + return v.Name == tlsVolumeName + }) { + t.Error("a plaintext install carries the TLS volume") + } +} + +// TestEveryControlPlaneWorkloadCarriesTheDecision is the one that matters. +// +// The management API is the server and the rest are its clients, and a client +// pool left plaintext against an API that requires a certificate cannot run the +// control plane's own tasks. Each of the four pod specs is built separately, so +// each is asserted separately. +func TestEveryControlPlaneWorkloadCarriesTheDecision(t *testing.T) { + cp := aLocalControlPlane(simplyblockv1alpha2.ControlPlaneTLS{}) + + for _, workload := range []struct { + name string + deployment *appsv1.Deployment + }{ + {"webappapi", webAPIDeployment(cp)}, + {"tasks", tasksDeployment(cp)}, + {"monitoring", monitoringDeployment(cp)}, + {"admin-control", adminControlDeployment(cp)}, + } { + pod := workload.deployment.Spec.Template.Spec + if !slices.ContainsFunc(pod.Volumes, func(v corev1.Volume) bool { + return v.Name == tlsVolumeName + }) { + t.Errorf("%s carries no TLS volume", workload.name) + } + for _, container := range pod.Containers { + if got, _ := envOf(container, "SB_TLS_CONNECT"); got != "authenticated" { + t.Errorf("%s/%s connects as %q", workload.name, container.Name, got) + } + if !mountsTLS(container) { + t.Errorf("%s/%s mounts no certificate to present", workload.name, container.Name) + } + } + } +} + +// The OpenShift service CA publishes its bundle separately, so the volume is a +// projection rather than the Secret alone. +func TestTheOpenShiftIssuerProjectsItsBundle(t *testing.T) { + deployment := webAPIDeployment(aLocalControlPlane(simplyblockv1alpha2.ControlPlaneTLS{ + Provider: simplyblockv1alpha2.ControlPlaneTLSOpenShift, + })) + + var volume *corev1.Volume + for i := range deployment.Spec.Template.Spec.Volumes { + if deployment.Spec.Template.Spec.Volumes[i].Name == tlsVolumeName { + volume = &deployment.Spec.Template.Spec.Volumes[i] + } + } + if volume == nil { + t.Fatal("no TLS volume") + } + if volume.Projected == nil { + t.Fatal("the OpenShift issuer's volume is the Secret alone, so ca.crt is missing") + } + if len(volume.Projected.Sources) != 2 { + t.Fatalf("the projection carries %d sources", len(volume.Projected.Sources)) + } +} + +// nestedAny reads a path out of the FoundationDBCluster's unstructured spec. +func nestedAny(t *testing.T, obj map[string]any, path ...string) any { + t.Helper() + value, found, err := unstructured.NestedFieldNoCopy(obj, path...) + if err != nil { + t.Fatalf("read %v: %v", path, err) + } + if !found { + return nil + } + return value +} + +// TestTheDatabaseListensForTLSWhenItsPeersAreAuthenticated is the switch, not +// the material. +// +// Every process can carry a current certificate, a key, and a CA and still talk +// to its peers in the clear: what turns the listeners over is enableTls on the +// cluster, and it is a field rather than an environment variable. A deployment +// with the mounts and without the field is the shape that looks configured and +// is not. +func TestTheDatabaseListensForTLSWhenItsPeersAreAuthenticated(t *testing.T) { + cluster := foundationDBCluster(aLocalControlPlane(simplyblockv1alpha2.ControlPlaneTLS{})) + + if got := nestedAny(t, cluster.Object, "spec", "mainContainer", "enableTls"); got != true { + t.Errorf("the database's own listeners carry enableTls = %v", got) + } + if got := nestedAny(t, cluster.Object, "spec", "sidecarContainer", "enableTls"); got != true { + t.Errorf("the sidecar carries enableTls = %v", got) + } +} + +// Every process class carries the certificate, because the FoundationDB operator +// replaces general.podTemplate wholesale for a class that overrides it: a volume +// stated once on general reaches neither storage nor log. +func TestEveryProcessClassCarriesThePeerCertificate(t *testing.T) { + cluster := foundationDBCluster(aLocalControlPlane(simplyblockv1alpha2.ControlPlaneTLS{})) + + for _, class := range []string{"general", "storage", "log"} { + volumes := nestedAny(t, cluster.Object, + "spec", "processes", class, "podTemplate", "spec", "volumes") + if volumes == nil { + t.Errorf("process class %s carries no peer certificate volume", class) + continue + } + if len(volumes.([]any)) != 1 { + t.Errorf("process class %s carries %d volumes", class, len(volumes.([]any))) + } + + containers := nestedAny(t, cluster.Object, + "spec", "processes", class, "podTemplate", "spec", "containers") + first := containers.([]any)[0].(map[string]any) + if first["env"] == nil { + t.Errorf("process class %s has no FDB_TLS_ environment", class) + } + if first["volumeMounts"] == nil { + t.Errorf("process class %s mounts nothing to read the certificate from", class) + } + } +} + +// A deployment whose callers are anonymous has no peer TLS either. There is no +// anonymous mode between database processes: each authenticates to the others or +// none of them do. +func TestAnAnonymousDeploymentLeavesTheDatabaseInTheClear(t *testing.T) { + cluster := foundationDBCluster(aLocalControlPlane( + simplyblockv1alpha2.ControlPlaneTLS{EnableMutualTLS: ptr.To(false)})) + + if got := nestedAny(t, cluster.Object, "spec", "mainContainer", "enableTls"); got != nil { + t.Errorf("the database listens for TLS with no certificates issued for it (%v)", got) + } + for _, class := range []string{"general", "storage", "log"} { + if got := nestedAny(t, cluster.Object, + "spec", "processes", class, "podTemplate", "spec", "volumes"); got != nil { + t.Errorf("process class %s carries a certificate volume it has no use for", class) + } + } +} + +// The FoundationDB operator reconciles the database, so it reaches it the way +// its processes reach each other. +func TestTheDatabaseOperatorCarriesThePeerCertificate(t *testing.T) { + deployment := fdbOperatorDeployment(aLocalControlPlane(simplyblockv1alpha2.ControlPlaneTLS{})) + manager := deployment.Spec.Template.Spec.Containers[0] + + if _, ok := envOf(manager, "FDB_TLS_CERTIFICATE_FILE"); !ok { + t.Error("the database operator has no certificate, so it cannot reach a TLS cluster") + } + if !slices.ContainsFunc(manager.VolumeMounts, func(m corev1.VolumeMount) bool { + return m.MountPath == fdbTLSMountPath + }) { + t.Error("the database operator mounts nothing at the path its environment names") + } + if !slices.ContainsFunc(deployment.Spec.Template.Spec.Volumes, func(v corev1.Volume) bool { + return v.Secret != nil && v.Secret.SecretName == FDBPeerCertSecret + }) { + t.Error("the database operator's pod carries no peer certificate volume") + } +} diff --git a/operator/internal/controllers/controlplane/workloads_test.go b/operator/internal/controllers/controlplane/workloads_test.go index b4b2cdc51..340b5dd36 100644 --- a/operator/internal/controllers/controlplane/workloads_test.go +++ b/operator/internal/controllers/controlplane/workloads_test.go @@ -15,6 +15,7 @@ import ( appsv1 "k8s.io/api/apps/v1" corev1 "k8s.io/api/core/v1" + "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" "sigs.k8s.io/controller-runtime/pkg/client" "github.com/simplyblock/atlas/ptr" @@ -241,24 +242,45 @@ func TestTheObjectStoreTakesTheSameStorageClassAsTheDatabase(t *testing.T) { // Nothing the install applies names a Secret that the install does not also // create, because a workload waiting on a Secret nobody writes never starts and // reports the wait as a pod event rather than on the ControlPlane. +// +// "Creates" includes producing it indirectly. A cert-manager Certificate is not +// a Secret and writes one, so a workload mounting what a Certificate in the same +// install issues is waiting on something this install does produce. The set is +// read off the applied objects rather than listed here, so a certificate that +// stops being applied takes its exemption with it. func TestNoWorkloadWaitsOnASecretTheInstallDoesNotCreate(t *testing.T) { cp := localControlPlane() - for _, obj := range append(foundationDBObjects(cp), - append(datastoreObjects(cp), managementAPIObjects(cp)...)...) { + objects := append(foundationDBObjects(cp), + append(datastoreObjects(cp), managementAPIObjects(cp)...)...) + issued := secretsTheInstallIssues(objects) + + for _, obj := range objects { spec := podSpecOf(obj) if spec == nil { continue } for _, volume := range spec.Volumes { - if volume.Secret != nil { + if volume.Secret != nil && !issued[volume.Secret.SecretName] { t.Errorf("%s mounts Secret %q, which no step of the install creates", obj.GetName(), volume.Secret.SecretName) } + if volume.Projected == nil { + continue + } + for _, source := range volume.Projected.Sources { + if source.Secret != nil && !issued[source.Secret.Name] { + t.Errorf("%s projects Secret %q, which no step of the install creates", + obj.GetName(), source.Secret.Name) + } + } } for _, container := range spec.Containers { for _, env := range container.Env { - if env.ValueFrom != nil && env.ValueFrom.SecretKeyRef != nil { + if env.ValueFrom == nil || env.ValueFrom.SecretKeyRef == nil { + continue + } + if !issued[env.ValueFrom.SecretKeyRef.Name] { t.Errorf("%s/%s reads Secret %q, which no step of the install creates", obj.GetName(), container.Name, env.ValueFrom.SecretKeyRef.Name) } @@ -267,6 +289,22 @@ func TestNoWorkloadWaitsOnASecretTheInstallDoesNotCreate(t *testing.T) { } } +// secretsTheInstallIssues is the Secrets the applied objects produce without +// being one: today, what each cert-manager Certificate writes into. +func secretsTheInstallIssues(objects []client.Object) map[string]bool { + issued := map[string]bool{} + for _, obj := range objects { + u, ok := obj.(*unstructured.Unstructured) + if !ok || u.GetKind() != "Certificate" { + continue + } + if name, found, _ := unstructured.NestedString(u.Object, "spec", "secretName"); found { + issued[name] = true + } + } + return issued +} + // podSpecOf is the pod template of whatever workload kind an object is, or nil // for the objects that carry none. func podSpecOf(obj client.Object) *corev1.PodSpec { diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_controlplanes.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_controlplanes.yaml index 140c3da8f..95afa9396 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_controlplanes.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_controlplanes.yaml @@ -333,6 +333,35 @@ spec: More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ type: object type: object + tls: + default: {} + description: |- + TLS is how this control plane serves and what it asks of its callers. + + The block is defaulted to its own zero value rather than left absent, so + that the field defaults inside it are applied to an object that does not + mention TLS at all. + properties: + enableMutualTLS: + default: true + description: |- + EnableMutualTLS additionally requires a caller to present a certificate + of its own, rather than reaching the control plane anonymously over the + encrypted connection EnableTLS alone provides. Ignored when EnableTLS is + false, the same as on DriverTLS. Unset is on. + type: boolean + enableTLS: + default: true + description: EnableTLS serves the management API over TLS. Unset is on. + type: boolean + provider: + default: cert-manager + description: Provider issues the serving certificate. + enum: + - cert-manager + - OpenShift + type: string + type: object tolerations: description: |- Tolerations are applied to every pod the operator installs for the control diff --git a/operator/internal/webapi/controlplane.go b/operator/internal/webapi/controlplane.go index dc0683407..992a0b4df 100644 --- a/operator/internal/webapi/controlplane.go +++ b/operator/internal/webapi/controlplane.go @@ -12,6 +12,7 @@ package webapi import ( "fmt" + "net/http" "github.com/simplyblock/atlas/controlplane" atlaskube "github.com/simplyblock/atlas/kube" @@ -29,5 +30,26 @@ func ControlPlaneConfig() (controlplane.Config, error) { if err != nil { return controlplane.Config{}, fmt.Errorf("read the service-account token: %w", err) } - return controlplane.Config{Endpoint: NewClient().BaseURL, Token: token}, nil + // The startup client already resolved both halves of how the control plane + // is reached: the base URL carries the scheme, and its transport carries the + // CA bundle and the client certificate this pod mounts. Carrying only the URL + // left the atlas-lib client to build its own connection, which is the system + // trust store and no certificate -- a client that cannot verify this + // deployment's CA and cannot authenticate to a control plane that asks. + startup := NewClient() + return controlplane.Config{ + Endpoint: startup.BaseURL, + Token: token, + Transport: transportOf(startup), + }, nil +} + +// transportOf is the connection a client was built with, and nil where it was +// built with the default one. Nil is what controlplane.Config takes to mean the +// same default, so the plaintext case stays the zero value. +func transportOf(c *Client) http.RoundTripper { + if c == nil || c.HttpClient == nil { + return nil + } + return c.HttpClient.Transport } diff --git a/operator/internal/webapi/controlplane_transport_test.go b/operator/internal/webapi/controlplane_transport_test.go new file mode 100644 index 000000000..bec470067 --- /dev/null +++ b/operator/internal/webapi/controlplane_transport_test.go @@ -0,0 +1,112 @@ +// The connection the atlas-lib client is handed. +// +// Two clients reach the same control plane from this process: the one in this +// package, which resolves TLS from the pod's environment and its mounted +// certificates, and the atlas-lib one the data-protection band writes through. +// Only the first of them was ever told how. The second was given the endpoint +// alone, so it verified this deployment's CA against the system trust store and +// presented nothing where a certificate was required. + +package webapi + +import ( + "net/http" + "os" + "path/filepath" + "strings" + "testing" + + atlaskube "github.com/simplyblock/atlas/kube" + + "github.com/simplyblock/simplyblock-operator/internal/tlsutil" +) + +// TestTheAtlasClientIsGivenTheSameConnection covers the handover. +func TestTheAtlasClientIsGivenTheSameConnection(t *testing.T) { + t.Setenv("SB_TLS_SERVE", "1") + t.Setenv("SB_TLS_CONNECT", "authenticated") + t.Setenv("SIMPLYBLOCK_WEBAPI_BASE_URL", "") + resetTLSClientCacheForTest(t) + + origNamespacePath := tlsutil.OperatorNamespacePath + origCAPath := tlsutil.ServiceCABundlePath + origCertPath := tlsutil.ServiceClientCertificatePath + origKeyPath := tlsutil.ServiceClientKeyPath + t.Cleanup(func() { + tlsutil.OperatorNamespacePath = origNamespacePath + tlsutil.ServiceCABundlePath = origCAPath + tlsutil.ServiceClientCertificatePath = origCertPath + tlsutil.ServiceClientKeyPath = origKeyPath + }) + + nsPath, caPath, certPath, keyPath := writeNamespaceAndCertPair(t) + tlsutil.OperatorNamespacePath = nsPath + tlsutil.ServiceCABundlePath = caPath + tlsutil.ServiceClientCertificatePath = certPath + tlsutil.ServiceClientKeyPath = keyPath + + startup := NewClient() + if startup.initErr != nil { + t.Fatalf("the startup client: %v", startup.initErr) + } + + // The assertion is on the Config the band is built from, not on the helper + // that fills it in: a helper that returns the right transport and is not + // wired into the Config is the same unverified connection. + tokenPath := filepath.Join(t.TempDir(), "token") + if err := os.WriteFile(tokenPath, []byte("a-token"), 0o600); err != nil { + t.Fatalf("write the token: %v", err) + } + origTokenPath := atlaskube.ServiceAccountTokenPath + t.Cleanup(func() { atlaskube.ServiceAccountTokenPath = origTokenPath }) + atlaskube.ServiceAccountTokenPath = tokenPath + + cfg, err := ControlPlaneConfig() + if err != nil { + t.Fatalf("ControlPlaneConfig: %v", err) + } + if !strings.HasPrefix(cfg.Endpoint, "https://") { + t.Errorf("the atlas-lib client is pointed at %q", cfg.Endpoint) + } + + carried := cfg.Transport + if carried == nil { + t.Fatal("the atlas-lib client is handed no transport, so it builds the default one") + } + if got := transportOf(startup); got != carried { + t.Error("the Config carries a different connection than the startup client's") + } + + transport, ok := carried.(*http.Transport) + if !ok { + t.Fatalf("the transport is a %T", carried) + } + if transport.TLSClientConfig == nil || transport.TLSClientConfig.RootCAs == nil { + t.Error("the transport carries no CA, so this deployment's certificate cannot be verified") + } + if len(transport.TLSClientConfig.Certificates) != 1 { + t.Errorf("the transport presents %d client certificates, want the pod's one", + len(transport.TLSClientConfig.Certificates)) + } +} + +// A plaintext deployment hands over nothing, which is what the atlas-lib client +// takes to mean its own default and is what every existing caller is. +func TestAPlaintextDeploymentHandsOverNoTransport(t *testing.T) { + t.Setenv("SB_TLS_SERVE", "") + t.Setenv("SIMPLYBLOCK_WEBAPI_BASE_URL", "") + resetTLSClientCacheForTest(t) + + if carried := transportOf(NewClient()); carried != nil { + t.Errorf("a plaintext deployment handed over a %T", carried) + } +} + +// A nil client is not a panic. ControlPlaneConfig builds one and reads it back +// in the same breath, so this is defense rather than a reachable path, but it is +// a nil dereference in the operator's startup if it ever becomes one. +func TestTransportOfNilIsNil(t *testing.T) { + if transportOf(nil) != nil { + t.Error("a nil client reported a transport") + } +} From 3554e6a36d0b4a6c54d1b0c805c04bb6f6f0278e Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Mon, 21 Sep 2026 10:05:09 +0200 Subject: [PATCH 106/206] fix(ci): the chart's renders state the capability its TLS guard reads TLS became on by default in 5b1cf4cc, and the chart refuses a cluster that does not serve cert-manager's API. That refusal is right at install time and impossible to satisfy in a render: `helm template` asks no cluster anything, so `.Capabilities.APIVersions` never holds the entry and the default values cannot be rendered at all. Helm's lint workflow has failed on every push to develop since, and so has every pull request into it. The Template step was only the first casualty. check-rendered-objects.sh renders each profile the same way and reads an empty render as a missing set, so its output was "standalone: RENDER FAILED", "managed: RENDER FAILED", and then every required object reported missing -- a chart-content failure for a reason that had nothing to do with the chart's content. Both now pass --api-versions cert-manager.io/v1, which is what a renderer with a cluster behind it supplies and what Argo CD and Flux already pass from theirs. The guard keeps its meaning where it has one, which is the install. Co-Authored-By: Claude Opus 5 (1M context) --- .github/workflows/helm_lint.yaml | 8 +++++++- helm-charts/scripts/check-rendered-objects.sh | 12 ++++++++++++ 2 files changed, 19 insertions(+), 1 deletion(-) diff --git a/.github/workflows/helm_lint.yaml b/.github/workflows/helm_lint.yaml index e77f74467..2dc4f8190 100644 --- a/.github/workflows/helm_lint.yaml +++ b/.github/workflows/helm_lint.yaml @@ -31,8 +31,14 @@ jobs: - name: Lint run: helm lint helm-charts/charts/simplyblock-operator + # --api-versions states a capability the render has no cluster to read. + # TLS is on by default and the chart refuses a cluster that does not serve + # cert-manager's API; a cluster-less render serves nothing, so the default + # values cannot be rendered at all without saying so here. - name: Template - run: helm template simplyblock-operator helm-charts/charts/simplyblock-operator + run: > + helm template simplyblock-operator helm-charts/charts/simplyblock-operator + --api-versions cert-manager.io/v1 - name: Required objects run: helm-charts/scripts/check-rendered-objects.sh diff --git a/helm-charts/scripts/check-rendered-objects.sh b/helm-charts/scripts/check-rendered-objects.sh index d3fda624e..c15acf790 100755 --- a/helm-charts/scripts/check-rendered-objects.sh +++ b/helm-charts/scripts/check-rendered-objects.sh @@ -9,6 +9,16 @@ set -uo pipefail CHART="$(cd "$(dirname "${BASH_SOURCE[0]}")/../charts/simplyblock-operator" && pwd)" fail=0 +# The capabilities a cluster-less render has to be told about. +# +# TLS is on by default and the chart refuses a cluster that does not serve +# cert-manager's API, which is the right refusal at install time and an +# impossible one here: `helm template` asks no cluster anything, so the guard +# fires on every render. Stating the capability is what a renderer with a +# cluster behind it does, and without it every profile below reads as a render +# failure rather than as the objects it is meant to check. +CAPABILITIES=(--api-versions cert-manager.io/v1) + # Objects every profile renders, as `Kind/name`. COMMON=( "Deployment/simplyblock-operator" @@ -43,6 +53,7 @@ check() { local out present missing=0 out="$(helm template sb "$CHART" --namespace simplyblock \ + "${CAPABILITIES[@]}" \ --set deployment.profile="$profile" \ --set controlplane.managed.endpoint=https://cp.example.com 2>/dev/null)" if [ -z "$out" ]; then @@ -69,6 +80,7 @@ render() { local profile="$1" shift helm template sb "$CHART" --namespace simplyblock \ + "${CAPABILITIES[@]}" \ --set deployment.profile="$profile" \ --set controlplane.managed.endpoint=https://cp.example.com \ "$@" 2>/dev/null | objects From e73262bb2024f655596543431b13ab1c2784dfc4 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Mon, 21 Sep 2026 09:18:41 +0100 Subject: [PATCH 107/206] updated doc design-csi-addons-replication.md --- operator/docs/designs/design-csi-addons-replication.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/operator/docs/designs/design-csi-addons-replication.md b/operator/docs/designs/design-csi-addons-replication.md index 4698b2fc6..c822c5789 100644 --- a/operator/docs/designs/design-csi-addons-replication.md +++ b/operator/docs/designs/design-csi-addons-replication.md @@ -214,7 +214,7 @@ The generated control-plane client already declares every replication endpoint, **Promote is one operation; planned and unplanned differ only in what precedes it.** Both forms clone the last fully replicated generation on the target and serve it under the preserved NVMe identity, exactly as failover does today. The planned form is lossless not because it runs different machinery but because a completed demote guarantees the last replicated generation contains every acknowledged write, and the planned gate refuses the promote until that holds. The forced form skips the gate and accepts the RPO loss, because its premise is that the source is gone. The engine's commit cutover (`replication/commit`, the `FN_REPLICATION_FINAL` runner with its shrink rounds and cutover-proceed handshake) is deliberately NOT part of this contract and remains behind only the legacy `ReplicationOps` migration path (§8, §13) -- its convergence job does NOT move into the demote verb as this design originally assumed; see the `DemoteVolume` row above for why that engine is the wrong shape for a post-unmount demote. -**Every Promote clones, including the first one a volume's own home cluster ever sees.** The vendored controller-manager has no "already primary" awareness of its own: `markVolumeAsPrimary` calls Promote unconditionally, every time a `VolumeReplication` first declares primary intent (`internal/controller/replication.storage/volumereplication_controller.go`), including day-one protection of a volume that has always lived here and has never failed over. `get_replication_info`'s `role` field cannot distinguish that case from a volume genuinely awaiting a planned re-promotion after a completed demote -- `demote_lvol` writes only `LVol.replication_demote_state`, a field `role`'s computation never reads, so a fenced, demoted volume still reports `role: source`, identically to one that was never touched at all. A role-based short-circuit before the backend call was tried and reverted for exactly this reason: it silently skipped the real `failover?planned=true` call a completed demote is waiting on, leaving the volume fenced while csi-addons reported `Completed=True`. So the clone-on-first-promote for a never-demoted volume is accepted as designed, not guarded against: `force=true`'s existing "ignores demote state entirely" premise already covers it, matching Test 6 of `regression_test/21/test_csi_addons_replication.sh`, which asserts a never-demoted volume auto-escalates through to a clone. Resolving the resulting cost -- an unwanted clone at day-one protection time -- needs a signal this design does not yet have (§15, Open Question 5). +**The planned form's no-demote branch splits on the source's own health, not on `role`.** The vendored controller-manager has no "already primary" awareness of its own: `markVolumeAsPrimary` calls Promote unconditionally, every time a `VolumeReplication` first declares primary intent (`internal/controller/replication.storage/volumereplication_controller.go`), including day-one protection of a volume that has always lived here and has never failed over. `get_replication_info`'s `role` field cannot distinguish that case from a volume genuinely awaiting a planned re-promotion after a completed demote -- `demote_lvol` writes only `LVol.replication_demote_state`, a field `role`'s computation never reads, so a fenced, demoted volume still reports `role: source`, identically to one that was never touched at all. A role-based short-circuit before the backend call was tried and reverted for exactly this reason: it silently skipped the real `failover?planned=true` call a completed demote is waiting on, leaving the volume fenced while csi-addons reported `Completed=True`. The fix instead reads the source's *own storage node* status (`lvol_controller.replication_source_online`, mirroring the target-node health check `replicate_lvol_on_target_cluster` already makes for the destination side): when `planned=true` and no demote was ever requested, a genuinely online source has nothing to fail over and the call succeeds as the no-op it is, with no clone; a source that is not online falls through to `FAILED_PRECONDITION` exactly as before, letting the vendored controller's force-escalation run for a real disaster. This narrows, but does not close, the race a staleness-tolerant health field always carries: a source that died within the last health-check interval still reads online and is treated as a no-op for one reconcile, correcting itself once the node's status catches up and the controller retries. Test 6 of `regression_test/21/test_csi_addons_replication.sh` needs updating to match -- its `$FRESH_PVC_NAME` scenario now reaches Primary via this no-op path, not via force-escalation, since nothing in that scenario's setup makes the source anything but healthy. --- @@ -446,4 +446,4 @@ A drill that silently perturbed replication would be worse than no drill. Before | 2 | **Where the preflight lives.** §7.2 attaches peerClasses validation to the `ReplicationPair` reconciler. If the redesign retires the pair kind, the preflight needs a new home (the `SimplyblockDriver`, or a standalone check job). | Operator team | | 3 | **Per-volume policy granularity.** A `VolumeReplicationClass` names one policy, and today one policy implies one target and cadence for all its volumes. Confirm one class per (policy, cadence) is an acceptable authoring model for Ramen's `replicationClassSelector`, or whether per-volume interval overrides are needed. | Operator / Backend team | | 4 | **Visibility of `drtest-` clones.** The test-cluster mode's clones on the secondary are replication sources only, never served. Decide whether the backend creates them as internal volumes (hidden from listings, exempt from the per-node subsystem cap, like the shipping path's landing volumes) or as ordinary volumes under a naming convention. | Backend team | -| 5 | **Avoiding the clone on day-one protection.** §5.2's promote always clones, including the first `PromoteVolume` a healthy, never-demoted volume ever receives -- the vendored controller-manager calls Promote unconditionally the first time any `VolumeReplication` declares primary intent, and `role` cannot tell that case apart from a volume genuinely awaiting a planned re-promotion (a demoted volume still reports `role: source`). A correct fix needs a signal this design does not expose today, most likely `replication_demote_state` (or an equivalent marker) surfaced through `get_replication_info`, so the driver can distinguish "never touched by the swap machinery" from "fenced, waiting on the real promote." Until that signal exists, every Ramen-protected volume's first reconcile materializes an unused clone on the peer, which also has no cleanup path today (`cleanup()` in the regression suite only removes Kubernetes objects). | Backend team | +| 5 | ~~**Avoiding the clone on day-one protection.**~~ **Resolved:** `POST .../replication/failover?planned=true`'s no-demote branch now checks `lvol_controller.replication_source_online` (the source's own storage-node status) before falling through to `FAILED_PRECONDITION` -- an online source is a no-op (§5.2), so a healthy volume's first-ever `PromoteVolume` no longer materializes a clone. The remaining residual: a source that dies within the last health-check interval still briefly reads online, so one reconcile can treat a genuine disaster as a no-op before the node's status catches up and the controller retries -- bounded by the health-check detection window, not open-ended. | Backend team | From 6591a5f09580c5250a5d349936639cb4d14212b9 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Mon, 21 Sep 2026 10:23:47 +0200 Subject: [PATCH 108/206] fix(operator): the probe of a TLS control plane verifies it The install learned to serve TLS and localEndpoint learned to say so, and the readiness probe kept the client it had, which was none. A control plane serving TLS presents a certificate signed by the CA this deployment minted, which is in no system trust store, so every probe failed the handshake: the management API is not answering yet: Get "https://simplyblock-webappapi .simplyblock.svc.cluster.local:5000/api/v2/_meta/ready": tls: failed to verify certificate: x509: certificate signed by unknown authority The ControlPlane sat in AwaitingDependency while the management API beside it was up and answering every other caller in the operator -- the informer's watch and the metrics scrape both verified it, because both go through clients built from the material this pod mounts. It is a report about the wrong component: the control plane was ready, and the thing that could not reach it was the reader. localAccess resolves both halves, from the files the pod already mounts and webapi.NewClient already reads. The two callers that built a bare endpoint take it, and so does the ControlPlaneOps prober, which had the same gap one function along. The second fault the deployment showed: two Certificates were issuing simplyblock-foundationdb-tls. The chart's carries server auth and client auth, because every database process is a server to its peers and a client of them; the one added with the install carried the serving default alone, and whichever wrote last decided what the Secret held. The install issues it now, under the chart's name and with both usages, and the chart no longer does -- the database is the operator's install, so its certificate goes with it, the same handover the management API's serving certificate already made. That one is why the chart's template is edited here: leaving it would keep both issuers and the race with them. Test plan: U-162 through U-170. Co-Authored-By: Claude Opus 5 (1M context) --- .../templates/controlplane_certificates.yaml | 30 +--- operator/docs/tests/test-plan-controlplane.md | 39 +++-- .../controlplane/controlplane_controller.go | 17 ++- .../controlplaneops_controller.go | 11 +- .../controlplane/controlplaneops_test.go | 3 + .../controllers/controlplane/endpoint.go | 43 ++++++ .../controllers/controlplane/install_test.go | 3 + .../controlplane/localaccess_test.go | 137 ++++++++++++++++++ .../internal/controllers/controlplane/tls.go | 53 ++++++- .../controllers/controlplane/tls_test.go | 54 +++++++ 10 files changed, 339 insertions(+), 51 deletions(-) create mode 100644 operator/internal/controllers/controlplane/localaccess_test.go diff --git a/helm-charts/charts/simplyblock-operator/templates/controlplane_certificates.yaml b/helm-charts/charts/simplyblock-operator/templates/controlplane_certificates.yaml index 1e2b84855..78205a832 100644 --- a/helm-charts/charts/simplyblock-operator/templates/controlplane_certificates.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/controlplane_certificates.yaml @@ -55,31 +55,11 @@ spec: - digital signature - key encipherment {{- if .Values.tls.mutual_enabled }} ---- -# Peer cert for the FoundationDB cluster pods (server + client auth, since FDB -# peers communicate symmetrically). Also used by the FDB operator's fdbcli. -apiVersion: cert-manager.io/v1 -kind: Certificate -metadata: - name: simplyblock-foundationdb - namespace: {{ .Release.Namespace }} -spec: - commonName: simplyblock-foundationdb - secretName: simplyblock-foundationdb-tls - issuerRef: - kind: ClusterIssuer - name: simplyblock-certificate-authority-issuer - usages: - - digital signature - - key encipherment - - server auth - - client auth - dnsNames: - - simplyblock-fdb-cluster - - simplyblock-fdb-cluster.{{ .Release.Namespace }} - - simplyblock-fdb-cluster.{{ .Release.Namespace }}.svc - - simplyblock-fdb-cluster.{{ .Release.Namespace }}.svc.cluster.local - - '*.simplyblock-fdb-cluster.{{ .Release.Namespace }}.svc.cluster.local' +# The FoundationDB peer certificate is not here. The database is the operator's +# install, and its certificate went with it: the operator issues +# simplyblock-foundationdb, under this name and with these usages, as part of +# applying the cluster it certifies. Two Certificates naming one secretName are +# two controllers writing into one Secret. --- apiVersion: cert-manager.io/v1 kind: Certificate diff --git a/operator/docs/tests/test-plan-controlplane.md b/operator/docs/tests/test-plan-controlplane.md index b90e4b833..0d4d4c234 100644 --- a/operator/docs/tests/test-plan-controlplane.md +++ b/operator/docs/tests/test-plan-controlplane.md @@ -161,21 +161,30 @@ Files: `operator/api/v1alpha2/controlplane_tls_test.go`, `operator/internal/controllers/controlplane/tls_test.go`, `operator/internal/webapi/controlplane_transport_test.go` -| # | Scenario | Type | Test | -|-------|----------------------------------------------------------------------------------|------------|--------------------------------------------------------| -| U-149 | A spec with no `tls` block serves TLS and requires a client certificate | Boundary | `TestAControlPlaneThatSaysNothingServesTLS` | -| U-150 | A nil `source.local` reads as closed rather than as plaintext | Boundary | `TestANilLocalControlPlaneStillReadsAsClosed` | -| U-151 | `enableMutualTLS: false` leaves the listener serving and the caller anonymous | Positive | `TestEachToggleIsHonored` | -| U-152 | `enableTLS: false` never asks a caller for a certificate | Negative | `TestPlaintextNeverAsksForAClientCertificate` | -| U-153 | A named issuer is not overwritten by the default | Positive | `TestANamedIssuerIsKept` | -| U-154 | The installed management API carries `SB_TLS_SERVE` and mounts its certificate | Positive | `TestTheInstalledControlPlaneServesTLSByDefault` | -| U-155 | Mutual TLS carries the `FDB_TLS_*` trio, and anonymous TLS carries none of it | Boundary | `TestMutualTLSCarriesTheFoundationDBPeerFiles` | -| U-156 | A plaintext install carries no TLS environment, mount, or volume at all | Negative | `TestDisablingServingInstallsThePlaintextControlPlane` | -| U-157 | Every workload of the install carries the decision, not only the one that serves | Regression | `TestEveryControlPlaneWorkloadCarriesTheDecision` | -| U-158 | The OpenShift issuer projects its bundle beside the Secret | Positive | `TestTheOpenShiftIssuerProjectsItsBundle` | -| U-159 | The atlas-lib client is handed the startup client's verified connection | Regression | `TestTheAtlasClientIsGivenTheSameConnection` | -| U-160 | A plaintext deployment hands the atlas-lib client no transport | Negative | `TestAPlaintextDeploymentHandsOverNoTransport` | -| U-161 | Nothing the install mounts is a Secret no step of it issues | Regression | `TestNoWorkloadWaitsOnASecretTheInstallDoesNotCreate` | +| # | Scenario | Type | Test | +|-------|----------------------------------------------------------------------------------|------------|------------------------------------------------------------| +| U-149 | A spec with no `tls` block serves TLS and requires a client certificate | Boundary | `TestAControlPlaneThatSaysNothingServesTLS` | +| U-150 | A nil `source.local` reads as closed rather than as plaintext | Boundary | `TestANilLocalControlPlaneStillReadsAsClosed` | +| U-151 | `enableMutualTLS: false` leaves the listener serving and the caller anonymous | Positive | `TestEachToggleIsHonored` | +| U-152 | `enableTLS: false` never asks a caller for a certificate | Negative | `TestPlaintextNeverAsksForAClientCertificate` | +| U-153 | A named issuer is not overwritten by the default | Positive | `TestANamedIssuerIsKept` | +| U-154 | The installed management API carries `SB_TLS_SERVE` and mounts its certificate | Positive | `TestTheInstalledControlPlaneServesTLSByDefault` | +| U-155 | Mutual TLS carries the `FDB_TLS_*` trio, and anonymous TLS carries none of it | Boundary | `TestMutualTLSCarriesTheFoundationDBPeerFiles` | +| U-156 | A plaintext install carries no TLS environment, mount, or volume at all | Negative | `TestDisablingServingInstallsThePlaintextControlPlane` | +| U-157 | Every workload of the install carries the decision, not only the one that serves | Regression | `TestEveryControlPlaneWorkloadCarriesTheDecision` | +| U-158 | The OpenShift issuer projects its bundle beside the Secret | Positive | `TestTheOpenShiftIssuerProjectsItsBundle` | +| U-159 | The atlas-lib client is handed the startup client's verified connection | Regression | `TestTheAtlasClientIsGivenTheSameConnection` | +| U-160 | A plaintext deployment hands the atlas-lib client no transport | Negative | `TestAPlaintextDeploymentHandsOverNoTransport` | +| U-161 | Nothing the install mounts is a Secret no step of it issues | Regression | `TestNoWorkloadWaitsOnASecretTheInstallDoesNotCreate` | +| U-162 | The probe of a TLS control plane carries the CA this deployment minted | Regression | `TestTheProbeOfATLSControlPlaneCarriesTheCA` | +| U-163 | A plaintext probe needs no CA and does not fail for want of one | Negative | `TestThePlaintextProbeNeedsNoCA` | +| U-164 | The local address follows the install, either scheme | Positive | `TestTheLocalAddressFollowsTheInstall` | +| U-165 | The database listens for TLS only when its peers are authenticated | Boundary | `TestTheDatabaseListensForTLSWhenItsPeersAreAuthenticated` | +| U-166 | Every process class carries the peer certificate, not only general | Regression | `TestEveryProcessClassCarriesThePeerCertificate` | +| U-167 | An anonymous deployment leaves the database in the clear | Negative | `TestAnAnonymousDeploymentLeavesTheDatabaseInTheClear` | +| U-168 | The database operator carries the peer certificate it reconciles with | Positive | `TestTheDatabaseOperatorCarriesThePeerCertificate` | +| U-169 | One certificate issues the peer Secret, and carries both roles | Regression | `TestThePeerCertificateIsIssuedOnceAndForBothRoles` | +| U-170 | No peer certificate is issued without peer TLS | Negative | `TestNoPeerCertificateWithoutPeerTLS` | ### Deletion (design §4.4) diff --git a/operator/internal/controllers/controlplane/controlplane_controller.go b/operator/internal/controllers/controlplane/controlplane_controller.go index 8497e627c..5e0e78f2a 100644 --- a/operator/internal/controllers/controlplane/controlplane_controller.go +++ b/operator/internal/controllers/controlplane/controlplane_controller.go @@ -299,7 +299,11 @@ func (r *ControlPlaneReconciler) performInstallStep( return true, "", applyAll(ctx, r.Client, cp, r.Scheme, managementAPIObjects(cp)) case stepAwaitingAPI: - ok, message := r.probe(ctx, cp.Namespace, managedAccess{endpoint: localEndpoint(cp)}) + access, err := localAccess(cp) + if err != nil { + return false, err.Error(), nil + } + ok, message := r.probe(ctx, cp.Namespace, access) if !ok { return false, fmt.Sprintf("the management API is not answering yet: %s", message), nil } @@ -322,9 +326,14 @@ func (r *ControlPlaneReconciler) steadyState( return ctrl.Result{}, err } - // A managed control plane is reached on the Service this install created, so - // there is no token to present and no CA beyond the cluster's own. - access := managedAccess{endpoint: localEndpoint(cp)} + // An installed control plane is reached on the Service this install created, + // so there is no token to present -- and, where it serves TLS, the CA this + // deployment minted, which no system trust store holds. + access, err := localAccess(cp) + if err != nil { + r.announce(cp, simplyblockv1alpha2.ControlPlanePhaseDegraded, err.Error()) + return ctrl.Result{RequeueAfter: steadyStateInterval}, nil + } ok, message := r.probe(ctx, cp.Namespace, access) components, err := observe(ctx, r.Client, cp.Namespace) diff --git a/operator/internal/controllers/controlplane/controlplaneops_controller.go b/operator/internal/controllers/controlplane/controlplaneops_controller.go index a799d16d6..4b09a3277 100644 --- a/operator/internal/controllers/controlplane/controlplaneops_controller.go +++ b/operator/internal/controllers/controlplane/controlplaneops_controller.go @@ -539,14 +539,21 @@ func (r *ControlPlaneOpsReconciler) verify( ops *simplyblockv1alpha2.ControlPlaneOps, target *simplyblockv1alpha2.ControlPlane, ) (bool, string, error) { + // The published endpoint is the address alone, so the material to verify it + // with is resolved either way: a TLS control plane reached over a client that + // trusts only the system store fails the handshake, not the request. + access, err := localAccess(target) + if err != nil { + return false, err.Error(), nil + } endpoint := target.Status.Endpoint if endpoint == "" { - endpoint = localEndpoint(target) + endpoint = access.endpoint } prober := r.Prober if prober == nil { - prober = &HTTPProber{} + prober = &HTTPProber{Client: access.client} } if ok, reason := prober.Ready(ctx, endpoint); !ok { return false, fmt.Sprintf("the control plane is not answering yet: %s", reason), nil diff --git a/operator/internal/controllers/controlplane/controlplaneops_test.go b/operator/internal/controllers/controlplane/controlplaneops_test.go index f04ea94c6..d7eb4afff 100644 --- a/operator/internal/controllers/controlplane/controlplaneops_test.go +++ b/operator/internal/controllers/controlplane/controlplaneops_test.go @@ -420,6 +420,7 @@ func TestAnUpgradeWritesTheImageOntoTheEntity(t *testing.T) { // was asked for, which is what separates an upgrade that completed from a // rollout that failed back. func TestVerifyingFailsOnAVersionThatDisagrees(t *testing.T) { + mountedCA(t) cp := localControlPlane() cp.Status.Endpoint = "http://simplyblock-webappapi.simplyblock.svc.cluster.local:5000" @@ -452,6 +453,7 @@ func TestVerifyingFailsOnAVersionThatDisagrees(t *testing.T) { // Verifying passes when the reported version is the one asked for. func TestVerifyingPassesOnTheVersionThatWasAskedFor(t *testing.T) { + mountedCA(t) cp := localControlPlane() ops := opsFor(simplyblockv1alpha2.ControlPlaneOpsActionUpgrade) ops.Spec.Upgrade = &simplyblockv1alpha2.UpgradeSpec{ @@ -478,6 +480,7 @@ func TestVerifyingPassesOnTheVersionThatWasAskedFor(t *testing.T) { // cannot answer would make the action unusable, and the record of the operation // has to carry what was and was not verified. func TestVerifyingPassesAndSaysSoWhenNoVersionIsServed(t *testing.T) { + mountedCA(t) ctx := context.Background() cp := localControlPlane() ops := opsFor(simplyblockv1alpha2.ControlPlaneOpsActionUpgrade) diff --git a/operator/internal/controllers/controlplane/endpoint.go b/operator/internal/controllers/controlplane/endpoint.go index f50344f3c..bd760684c 100644 --- a/operator/internal/controllers/controlplane/endpoint.go +++ b/operator/internal/controllers/controlplane/endpoint.go @@ -29,6 +29,7 @@ import ( "sigs.k8s.io/controller-runtime/pkg/client" simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/tlsutil" ) // The keys a credentials Secret may carry. Two are accepted because the two @@ -58,6 +59,48 @@ func localEndpoint(cp *simplyblockv1alpha2.ControlPlane) string { scheme, ComponentWebAPI, cp.Namespace, webAPIPort) } +// localAccess is where an installed control plane answers and what to reach it +// with. +// +// The second half is the part that is easy to leave out. A control plane serving +// TLS presents a certificate signed by the deployment's own CA, which is in no +// system trust store, so a probe given the address alone fails the handshake and +// reports it as the control plane not being ready -- a sentence about the wrong +// component, on a control plane that is up and answering every other caller. +// +// The material is the one this pod already mounts, and the same files +// webapi.NewClient reads. Reading the Secrets through the API server instead +// would be a second way to answer a question the deployment has already +// answered, and the two would drift. +func localAccess(cp *simplyblockv1alpha2.ControlPlane) (managedAccess, error) { + access := managedAccess{endpoint: localEndpoint(cp)} + + local := cp.Spec.Source.Local + if !local.ServesTLS() { + return access, nil + } + + // The client certificate only where the control plane asks for one: a + // deployment serving TLS anonymously mounts no certificate to present, and + // naming the paths anyway fails on the files not being there. + certPath, keyPath := "", "" + if local.RequiresClientCertificate() { + certPath = tlsutil.ServiceClientCertificatePath + keyPath = tlsutil.ServiceClientKeyPath + } + + verified, err := tlsutil.BuildWebAPIClient( + cp.Namespace, tlsutil.ServiceCABundlePath, certPath, keyPath) + if err != nil { + return managedAccess{}, &credentialsError{message: fmt.Sprintf( + "this control plane serves TLS and the material to verify it with could not be "+ + "read from this pod: %v; a deployment whose operator mounts none sets "+ + "spec.source.local.tls.enableTLS to false", err)} + } + access.client = verified + return access, nil +} + // credentialsError is what a Secret that is missing or unusable produces. It is // its own type so the reconciler can emit the CredentialsError event for it and // EndpointUnreachable for everything else, which are different problems with diff --git a/operator/internal/controllers/controlplane/install_test.go b/operator/internal/controllers/controlplane/install_test.go index 536c3e1bd..84b501be3 100644 --- a/operator/internal/controllers/controlplane/install_test.go +++ b/operator/internal/controllers/controlplane/install_test.go @@ -273,6 +273,9 @@ func TestReadingAnAbsentFoundationDBClusterIsNotAnError(t *testing.T) { // AwaitingAPI holds on the probe rather than on the pod counts, because the // question it answers is whether the control plane can be reached at all. func TestAwaitingAPIHoldsUntilTheProbePasses(t *testing.T) { + // The fixture serves TLS, as an install does, so the process has to hold what + // the pod mounts before the probe is reached at all. + mountedCA(t) cp := localControlPlane() prober := &stubProber{ready: false, readyMessage: "connection refused"} r := &ControlPlaneReconciler{ diff --git a/operator/internal/controllers/controlplane/localaccess_test.go b/operator/internal/controllers/controlplane/localaccess_test.go new file mode 100644 index 000000000..e09f8e142 --- /dev/null +++ b/operator/internal/controllers/controlplane/localaccess_test.go @@ -0,0 +1,137 @@ +// What the operator reaches its own control plane with. +// +// The endpoint is half of it. A control plane serving TLS presents a certificate +// signed by the deployment's own CA, which is in no system trust store, so a +// probe given the address and nothing else fails the handshake rather than the +// request -- and reports it as the control plane not being ready, which is a +// sentence about the wrong component. + +package controlplane + +import ( + "crypto/rand" + "crypto/rsa" + "crypto/x509" + "crypto/x509/pkix" + "encoding/pem" + "math/big" + "os" + "path/filepath" + "testing" + "time" + + "github.com/simplyblock/atlas/ptr" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/tlsutil" +) + +// mountedCA writes what a mutual-TLS operator pod mounts -- the CA bundle, the +// client certificate, and its key, all three out of one Secret -- and points +// tlsutil at them for the duration of the test. +func mountedCA(t *testing.T) { + t.Helper() + + key, err := rsa.GenerateKey(rand.Reader, 2048) + if err != nil { + t.Fatalf("generate a key: %v", err) + } + template := &x509.Certificate{ + SerialNumber: big.NewInt(1), + Subject: pkix.Name{CommonName: "simplyblock-certificate-authority"}, + NotBefore: time.Now().Add(-time.Hour), + NotAfter: time.Now().Add(time.Hour), + IsCA: true, + BasicConstraintsValid: true, + } + der, err := x509.CreateCertificate(rand.Reader, template, template, &key.PublicKey, key) + if err != nil { + t.Fatalf("sign the CA: %v", err) + } + + dir := t.TempDir() + certPEM := pem.EncodeToMemory(&pem.Block{Type: "CERTIFICATE", Bytes: der}) + keyPEM := pem.EncodeToMemory(&pem.Block{ + Type: "RSA PRIVATE KEY", Bytes: x509.MarshalPKCS1PrivateKey(key), + }) + for name, content := range map[string][]byte{ + "ca.crt": certPEM, "tls.crt": certPEM, "tls.key": keyPEM, + } { + if err := os.WriteFile(filepath.Join(dir, name), content, 0o600); err != nil { + t.Fatalf("write %s: %v", name, err) + } + } + + originalCA := tlsutil.ServiceCABundlePath + originalCert := tlsutil.ServiceClientCertificatePath + originalKey := tlsutil.ServiceClientKeyPath + t.Cleanup(func() { + tlsutil.ServiceCABundlePath = originalCA + tlsutil.ServiceClientCertificatePath = originalCert + tlsutil.ServiceClientKeyPath = originalKey + }) + tlsutil.ServiceCABundlePath = filepath.Join(dir, "ca.crt") + tlsutil.ServiceClientCertificatePath = filepath.Join(dir, "tls.crt") + tlsutil.ServiceClientKeyPath = filepath.Join(dir, "tls.key") +} + +// TestTheProbeOfATLSControlPlaneCarriesTheCA is the defect. +// +// Regression: 2026-09-21-the-local-probe-verified-against-the-system-store — the +// install learned to serve TLS and localEndpoint learned to say so, and the +// probe kept the client it had, which was none. Every readiness read failed with +// "x509: certificate signed by unknown authority" and the ControlPlane reported +// AwaitingDependency forever, while the management API beside it was serving and +// answering every other caller in the operator. +func TestTheProbeOfATLSControlPlaneCarriesTheCA(t *testing.T) { + mountedCA(t) + + access, err := localAccess(aLocalControlPlane(simplyblockv1alpha2.ControlPlaneTLS{})) + if err != nil { + t.Fatalf("localAccess: %v", err) + } + if access.client == nil { + t.Fatal("the probe of a TLS control plane verifies against the system trust store, " + + "which does not hold this deployment's CA") + } +} + +// A plaintext install needs none of it, and must not fail for want of a CA file +// that its pod does not mount. +func TestThePlaintextProbeNeedsNoCA(t *testing.T) { + original := tlsutil.ServiceCABundlePath + t.Cleanup(func() { tlsutil.ServiceCABundlePath = original }) + tlsutil.ServiceCABundlePath = filepath.Join(t.TempDir(), "absent") + + access, err := localAccess(aLocalControlPlane( + simplyblockv1alpha2.ControlPlaneTLS{EnableTLS: ptr.To(false)})) + if err != nil { + t.Fatalf("a plaintext control plane could not be reached: %v", err) + } + if access.client != nil { + t.Error("a plaintext probe carries a TLS client") + } +} + +// The address follows the install either way, because it is published as +// status.endpoint and every control-plane call in the operator resolves it. +func TestTheLocalAddressFollowsTheInstall(t *testing.T) { + mountedCA(t) + + secure, err := localAccess(aLocalControlPlane(simplyblockv1alpha2.ControlPlaneTLS{})) + if err != nil { + t.Fatalf("localAccess: %v", err) + } + if got := secure.endpoint[:8]; got != "https://" { + t.Errorf("a TLS control plane is reached over %q", secure.endpoint) + } + + plain, err := localAccess(aLocalControlPlane( + simplyblockv1alpha2.ControlPlaneTLS{EnableTLS: ptr.To(false)})) + if err != nil { + t.Fatalf("localAccess: %v", err) + } + if got := plain.endpoint[:7]; got != "http://" { + t.Errorf("a plaintext control plane is reached at %q", plain.endpoint) + } +} diff --git a/operator/internal/controllers/controlplane/tls.go b/operator/internal/controllers/controlplane/tls.go index abdfce29c..848c599c1 100644 --- a/operator/internal/controllers/controlplane/tls.go +++ b/operator/internal/controllers/controlplane/tls.go @@ -18,7 +18,10 @@ package controlplane import ( + "fmt" + corev1 "k8s.io/api/core/v1" + "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" "sigs.k8s.io/controller-runtime/pkg/client" simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" @@ -170,6 +173,13 @@ const ( // each other. It carries both usages, because every process is a server to // its peers and a client of them. FDBPeerCertSecret = "simplyblock-foundationdb-tls" + + // fdbPeerCertificateName is the Certificate that issues it, under the name + // the chart used. Keeping the name is what makes this a handover rather than + // a second issuer for one Secret: two Certificates naming one secretName are + // two controllers writing to one place, and the narrower of them wins + // whenever it happens to write last. + fdbPeerCertificateName = "simplyblock-foundationdb" ) // fdbPeerTLS reports whether the database's own connections are encrypted. @@ -233,15 +243,48 @@ func fdbOperatorPeerMount() []corev1.VolumeMount { }} } -// fdbPeerCertificate is the Certificate that issues it, and nothing where the -// deployment does not use peer TLS or signs through the OpenShift service CA -- -// that CA signs from a Service annotation, and these processes have no Service. +// fdbPeerCertificate issues the material, and carries both usages. +// +// Every database process is a server to its peers and a client of them, so a +// certificate with only server auth fails the half of the handshake where it is +// the client. That is why this is its own builder rather than the serving one +// above: BuildServiceServingCertificate names a server and nothing else. +// +// Nothing is issued under the OpenShift service CA, which signs from an +// annotation on a Service, and these processes have none. func fdbPeerCertificate(cp *simplyblockv1alpha2.ControlPlane) []client.Object { if !fdbPeerTLS(cp) || cp.Spec.Source.Local.TLSProvider() != simplyblockv1alpha2.ControlPlaneTLSCertManager { return nil } - return []client.Object{ - utils.BuildServiceServingCertificate(cp.Namespace, ComponentFDBCluster, FDBPeerCertSecret), + + names := []any{ + ComponentFDBCluster, + fmt.Sprintf("%s.%s", ComponentFDBCluster, cp.Namespace), + fmt.Sprintf("%s.%s.svc", ComponentFDBCluster, cp.Namespace), + fmt.Sprintf("%s.%s.svc.cluster.local", ComponentFDBCluster, cp.Namespace), + fmt.Sprintf("*.%s.%s.svc.cluster.local", ComponentFDBCluster, cp.Namespace), } + + certificate := &unstructured.Unstructured{Object: map[string]any{ + "apiVersion": "cert-manager.io/v1", + "kind": "Certificate", + "metadata": map[string]any{ + "name": fdbPeerCertificateName, + "namespace": cp.Namespace, + }, + "spec": map[string]any{ + "commonName": fdbPeerCertificateName, + "secretName": FDBPeerCertSecret, + "issuerRef": map[string]any{ + "kind": "ClusterIssuer", + "name": utils.CertManagerClusterIssuerName, + }, + "usages": []any{ + "digital signature", "key encipherment", "server auth", "client auth", + }, + "dnsNames": names, + }, + }} + return []client.Object{certificate} } diff --git a/operator/internal/controllers/controlplane/tls_test.go b/operator/internal/controllers/controlplane/tls_test.go index aeb523fc3..a4f8bf3f3 100644 --- a/operator/internal/controllers/controlplane/tls_test.go +++ b/operator/internal/controllers/controlplane/tls_test.go @@ -194,6 +194,9 @@ func TestTheOpenShiftIssuerProjectsItsBundle(t *testing.T) { } } +// certificateKind is what an applied cert-manager object reports itself as. +const certificateKind = "Certificate" + // nestedAny reads a path out of the FoundationDBCluster's unstructured spec. func nestedAny(t *testing.T, obj map[string]any, path ...string) any { t.Helper() @@ -293,3 +296,54 @@ func TestTheDatabaseOperatorCarriesThePeerCertificate(t *testing.T) { t.Error("the database operator's pod carries no peer certificate volume") } } + +// The peer certificate is issued once, by whoever installs the database. +// +// Regression: 2026-09-21-two-certificates-one-secret — the install applied a +// second Certificate for simplyblock-foundationdb-tls beside the chart's, under +// a different name and with only the usages a server needs. Two controllers +// issuing into one Secret is a certificate that changes whenever either of them +// writes, and the narrower one takes client auth away from a database whose +// processes are each other's clients. +func TestThePeerCertificateIsIssuedOnceAndForBothRoles(t *testing.T) { + objects := foundationDBObjects(aLocalControlPlane(simplyblockv1alpha2.ControlPlaneTLS{})) + + var issued []*unstructured.Unstructured + for _, obj := range objects { + u, ok := obj.(*unstructured.Unstructured) + if !ok || u.GetKind() != certificateKind { + continue + } + name, _, _ := unstructured.NestedString(u.Object, "spec", "secretName") + if name == FDBPeerCertSecret { + issued = append(issued, u) + } + } + + if len(issued) != 1 { + t.Fatalf("%d certificates issue %s", len(issued), FDBPeerCertSecret) + } + if got := issued[0].GetName(); got != fdbPeerCertificateName { + t.Errorf("the peer certificate is named %q, and the chart's Secret is claimed by %q", + got, fdbPeerCertificateName) + } + + usages, _, _ := unstructured.NestedStringSlice(issued[0].Object, "spec", "usages") + for _, want := range []string{"server auth", "client auth"} { + if !slices.Contains(usages, want) { + t.Errorf("the peer certificate is missing %q, and every process is both", want) + } + } +} + +// A deployment with no peer TLS issues nothing for it. +func TestNoPeerCertificateWithoutPeerTLS(t *testing.T) { + objects := foundationDBObjects(aLocalControlPlane( + simplyblockv1alpha2.ControlPlaneTLS{EnableMutualTLS: ptr.To(false)})) + + for _, obj := range objects { + if u, ok := obj.(*unstructured.Unstructured); ok && u.GetKind() == certificateKind { + t.Errorf("a deployment with no peer TLS issues %s", u.GetName()) + } + } +} From 9322afcff0d826702b138a91f1efaf0bed95e8cf Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Mon, 21 Sep 2026 10:52:27 +0200 Subject: [PATCH 109/206] fix(operator): the exporter was given the cluster file and no keys Peer TLS rewrites the coordinators in the cluster file as :4500:tls, and every process holding that file has to present a certificate to open the database. The install gave the FoundationDB exporter the file and nothing else, so its client hung in the handshake, it never served /metrics, and the liveness probe killed and restarted it for as long as the deployment stayed up. The ControlPlane stayed Degraded on "simplyblock-fdb-exporter has 0 of 1 replicas ready" with every other component of the install ready. The three process classes, the FoundationDB operator, and the control plane's own pods were wired when peer TLS landed. The exporter is a client like all of them and was missed, which is the shape this kind of change fails in: the material is added to the workloads a reader thinks of as the database, and the one that only reads from it is somewhere else in the file. So the test is over the cluster file rather than over a list of workloads. Being given that file is what makes a process a client of the database, and a test naming the exporter would be a test of this fix rather than of the contract it broke. Test plan: U-171. Co-Authored-By: Claude Opus 5 (1M context) --- operator/docs/tests/test-plan-controlplane.md | 1 + .../controllers/controlplane/managementapi.go | 16 +++++---- .../controllers/controlplane/tls_test.go | 36 +++++++++++++++++++ 3 files changed, 47 insertions(+), 6 deletions(-) diff --git a/operator/docs/tests/test-plan-controlplane.md b/operator/docs/tests/test-plan-controlplane.md index 0d4d4c234..3b88601af 100644 --- a/operator/docs/tests/test-plan-controlplane.md +++ b/operator/docs/tests/test-plan-controlplane.md @@ -185,6 +185,7 @@ Files: `operator/api/v1alpha2/controlplane_tls_test.go`, | U-168 | The database operator carries the peer certificate it reconciles with | Positive | `TestTheDatabaseOperatorCarriesThePeerCertificate` | | U-169 | One certificate issues the peer Secret, and carries both roles | Regression | `TestThePeerCertificateIsIssuedOnceAndForBothRoles` | | U-170 | No peer certificate is issued without peer TLS | Negative | `TestNoPeerCertificateWithoutPeerTLS` | +| U-171 | Every workload holding the cluster file can open a :tls database | Regression | `TestEveryHolderOfTheClusterFileCanReachATLSDatabase` | ### Deletion (design §4.4) diff --git a/operator/internal/controllers/controlplane/managementapi.go b/operator/internal/controllers/controlplane/managementapi.go index d03ef131f..2c8037869 100644 --- a/operator/internal/controllers/controlplane/managementapi.go +++ b/operator/internal/controllers/controlplane/managementapi.go @@ -502,6 +502,10 @@ func adminControlDeployment(cp *simplyblockv1alpha2.ControlPlane) *appsv1.Deploy func fdbExporterDeployment(cp *simplyblockv1alpha2.ControlPlane) *appsv1.Deployment { const tmpVolume = "tmp" labels := map[string]string{appLabel: ComponentFDBExporter} + // The exporter opens the database like any other client, so peer TLS is its + // decision too: the cluster file it is given names :tls coordinators, and a + // client with no certificate hangs in the handshake rather than failing. + peerTLS := fdbPeerTLS(cp) spec := corev1.PodSpec{ SecurityContext: &corev1.PodSecurityContext{ @@ -510,16 +514,16 @@ func fdbExporterDeployment(cp *simplyblockv1alpha2.ControlPlane) *appsv1.Deploym RunAsGroup: ptr.To(int64(4059)), FSGroup: ptr.To(int64(4059)), }, - Volumes: []corev1.Volume{ + Volumes: append([]corev1.Volume{ {Name: tmpVolume, VolumeSource: corev1.VolumeSource{EmptyDir: &corev1.EmptyDirVolumeSource{}}}, clusterFileVolumeSource(), - }, + }, peerVolumeIf(peerTLS)...), Containers: []corev1.Container{{ Name: "exporter", Image: fdbExporterImage, - Env: []corev1.EnvVar{ + Env: append([]corev1.EnvVar{ {Name: "FDB_CLUSTER_FILE", Value: clusterFilePath}, - }, + }, peerEnvIf(peerTLS)...), Ports: []corev1.ContainerPort{{Name: "metrics", ContainerPort: fdbExporterPort}}, LivenessProbe: &corev1.Probe{ ProbeHandler: corev1.ProbeHandler{ @@ -540,7 +544,7 @@ func fdbExporterDeployment(cp *simplyblockv1alpha2.ControlPlane) *appsv1.Deploym AllowPrivilegeEscalation: ptr.To(false), Privileged: ptr.To(false), }, - VolumeMounts: []corev1.VolumeMount{ + VolumeMounts: append([]corev1.VolumeMount{ {Name: tmpVolume, MountPath: "/tmp"}, { Name: clusterFileVolume, @@ -548,7 +552,7 @@ func fdbExporterDeployment(cp *simplyblockv1alpha2.ControlPlane) *appsv1.Deploym SubPath: clusterFileSubURL, ReadOnly: true, }, - }, + }, peerMountIf(peerTLS)...), Resources: corev1.ResourceRequirements{ Requests: corev1.ResourceList{ corev1.ResourceCPU: resource.MustParse("50m"), diff --git a/operator/internal/controllers/controlplane/tls_test.go b/operator/internal/controllers/controlplane/tls_test.go index a4f8bf3f3..446376f11 100644 --- a/operator/internal/controllers/controlplane/tls_test.go +++ b/operator/internal/controllers/controlplane/tls_test.go @@ -347,3 +347,39 @@ func TestNoPeerCertificateWithoutPeerTLS(t *testing.T) { } } } + +// TestEveryHolderOfTheClusterFileCanReachATLSDatabase is the rule the exporter +// broke. +// +// Regression: 2026-09-21-the-exporter-was-given-the-cluster-file-and-no-keys — +// peer TLS rewrites the coordinators in the cluster file as :4500:tls, so every +// process holding that file has to present a certificate to open the database. +// The install gave the FoundationDB exporter the file and nothing else: its +// client hung in the handshake, it never served /metrics, and the liveness probe +// killed and restarted it for as long as the deployment stayed up. +// +// The rule is the cluster file rather than a list of workloads, because that is +// what makes a process a client of the database. A test naming the exporter +// would be a test of this fix rather than of the contract it broke. +func TestEveryHolderOfTheClusterFileCanReachATLSDatabase(t *testing.T) { + cp := aLocalControlPlane(simplyblockv1alpha2.ControlPlaneTLS{}) + + objects := append(foundationDBObjects(cp), managementAPIObjects(cp)...) + for _, obj := range objects { + spec := podSpecOf(obj) + if spec == nil { + continue + } + for _, container := range spec.Containers { + if !slices.ContainsFunc(container.VolumeMounts, func(m corev1.VolumeMount) bool { + return m.MountPath == clusterFilePath + }) { + continue + } + if _, ok := envOf(container, "FDB_TLS_CERTIFICATE_FILE"); !ok { + t.Errorf("%s/%s holds the cluster file and no certificate, so it cannot open "+ + "a database whose coordinators are :tls", obj.GetName(), container.Name) + } + } + } +} From 657afc00c53744c8b8834635ae79060c6553ad41 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Mon, 21 Sep 2026 10:53:10 +0200 Subject: [PATCH 110/206] style(operator): the brand spelling the gate asked for Two lines of the previous commit wrote the wire spelling of the cluster file's coordinator suffix in prose, where the house style is the acronym. The gate said so and the commit went out before its output was read. Co-Authored-By: Claude Opus 5 (1M context) --- operator/docs/tests/test-plan-controlplane.md | 2 +- operator/internal/controllers/controlplane/managementapi.go | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/operator/docs/tests/test-plan-controlplane.md b/operator/docs/tests/test-plan-controlplane.md index 3b88601af..a1a283deb 100644 --- a/operator/docs/tests/test-plan-controlplane.md +++ b/operator/docs/tests/test-plan-controlplane.md @@ -185,7 +185,7 @@ Files: `operator/api/v1alpha2/controlplane_tls_test.go`, | U-168 | The database operator carries the peer certificate it reconciles with | Positive | `TestTheDatabaseOperatorCarriesThePeerCertificate` | | U-169 | One certificate issues the peer Secret, and carries both roles | Regression | `TestThePeerCertificateIsIssuedOnceAndForBothRoles` | | U-170 | No peer certificate is issued without peer TLS | Negative | `TestNoPeerCertificateWithoutPeerTLS` | -| U-171 | Every workload holding the cluster file can open a :tls database | Regression | `TestEveryHolderOfTheClusterFileCanReachATLSDatabase` | +| U-171 | Every workload holding the cluster file can open a TLS database | Regression | `TestEveryHolderOfTheClusterFileCanReachATLSDatabase` | ### Deletion (design §4.4) diff --git a/operator/internal/controllers/controlplane/managementapi.go b/operator/internal/controllers/controlplane/managementapi.go index 2c8037869..09c4c9212 100644 --- a/operator/internal/controllers/controlplane/managementapi.go +++ b/operator/internal/controllers/controlplane/managementapi.go @@ -503,7 +503,7 @@ func fdbExporterDeployment(cp *simplyblockv1alpha2.ControlPlane) *appsv1.Deploym const tmpVolume = "tmp" labels := map[string]string{appLabel: ComponentFDBExporter} // The exporter opens the database like any other client, so peer TLS is its - // decision too: the cluster file it is given names :tls coordinators, and a + // decision too: the cluster file it is given names TLS coordinators, and a // client with no certificate hangs in the handshake rather than failing. peerTLS := fdbPeerTLS(cp) From 3feaf61c6e29d1af0c28d9342a6eff19300d0c94 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Mon, 21 Sep 2026 11:00:11 +0200 Subject: [PATCH 111/206] fix(operator): a control plane that answers does not hold an expansion MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The gate compared the phase to Available and nothing else, so every phase that was not that word waited, Degraded included. design-controlplane.md §4.3 says the opposite, and says it directly: Degraded is a control plane that answers, so nothing holds on it. It is the phase a management API pod restarting behind a Service produces, and a FoundationDB pod recycled without losing quorum, and a metrics exporter that is not running at all. The phase exists to keep a survivable event survivable, and the gate turned each of them into a stop. On the deployment this was found on, an approved config sat at ControlPlane simplyblock is Degraded rather than Available with every other component ready, because the FoundationDB metrics exporter was crash-looping. That component is non-essential by the table components.go derives its verdicts from -- its absence loses no work and stops nothing -- and it held six workers out of a cluster. What the gate asks now is whether the control plane answers. Unavailable is the phase worth waiting for, because it is the one that means the probe failed or an essential component is at zero, and every call the expansion is about to make would fail against it. Installing and an unset phase hold too: there is nothing to call yet, which is a different thing from something impaired. The controlPlaneAvailable constant is gone rather than joined by a second one. Which phases let an expansion through is a decision about answering, and a pair of constants to compare against invites the next reader to add a third. Test plan: U-178 through U-181. Co-Authored-By: Claude Opus 5 (1M context) --- .../test-plan-clusterdeploymentconfig.md | 22 +++-- .../clusterdeploymentconfig_controller.go | 47 +++++++-- .../deployment/controlplane_gate_test.go | 97 +++++++++++++++++++ 3 files changed, 147 insertions(+), 19 deletions(-) create mode 100644 operator/internal/controllers/deployment/controlplane_gate_test.go diff --git a/operator/docs/tests/test-plan-clusterdeploymentconfig.md b/operator/docs/tests/test-plan-clusterdeploymentconfig.md index c1dc9a7a1..2d048a145 100644 --- a/operator/docs/tests/test-plan-clusterdeploymentconfig.md +++ b/operator/docs/tests/test-plan-clusterdeploymentconfig.md @@ -84,15 +84,19 @@ File: `operator/internal/controllers/deployment/clusterdeploymentconfig_expand_t ### Approval (design §5) -| # | Scenario | Type | Test | -|------|---------------------------------------------------------------------------------|----------|------| -| U-35 | `spec.approved` false: expansion is never entered | Negative | — | -| U-36 | `spec.approved` set true: expansion begins on the next reconcile | Positive | — | -| U-37 | The `ready-to-deploy` label is written on an approved config | Positive | — | -| U-38 | The `ready-to-deploy` label is read from nowhere: setting it alone does nothing | Negative | — | -| U-39 | Expansion is held while the `ControlPlane` is not `Ready` | Negative | — | -| U-40 | The `ControlPlane` becomes `Ready`: the held expansion proceeds unattended | Positive | — | -| U-41 | No `ControlPlane` at all: held with a clear reason, not failed | Negative | — | +| # | Scenario | Type | Test | +|-------|---------------------------------------------------------------------------------|------------|-----------------------------------------------------| +| U-35 | `spec.approved` false: expansion is never entered | Negative | — | +| U-36 | `spec.approved` set true: expansion begins on the next reconcile | Positive | — | +| U-37 | The `ready-to-deploy` label is written on an approved config | Positive | — | +| U-38 | The `ready-to-deploy` label is read from nowhere: setting it alone does nothing | Negative | — | +| U-39 | Expansion is held while the `ControlPlane` is not `Ready` | Negative | — | +| U-40 | The `ControlPlane` becomes `Ready`: the held expansion proceeds unattended | Positive | — | +| U-41 | No `ControlPlane` at all: held with a clear reason, not failed | Negative | — | +| U-178 | A `Degraded` control plane answers, so the expansion proceeds | Regression | `TestADegradedControlPlaneDoesNotHoldTheDeployment` | +| U-179 | An `Available` control plane proceeds | Positive | `TestAnAvailableControlPlaneProceeds` | +| U-180 | An `Unavailable` control plane holds, with a reason | Negative | `TestAnUnavailableControlPlaneHolds` | +| U-181 | A control plane still being installed, or reporting no phase, holds | Boundary | `TestAControlPlaneStillBeingBuiltHolds` | ### Deletion (design §4.3) diff --git a/operator/internal/controllers/deployment/clusterdeploymentconfig_controller.go b/operator/internal/controllers/deployment/clusterdeploymentconfig_controller.go index 71ec2179c..67d714cf7 100644 --- a/operator/internal/controllers/deployment/clusterdeploymentconfig_controller.go +++ b/operator/internal/controllers/deployment/clusterdeploymentconfig_controller.go @@ -314,8 +314,19 @@ func (r *ClusterDeploymentConfigReconciler) markReadyToDeploy( return r.Patch(ctx, config, patch) } -// controlPlaneReady reports whether the singleton is available, which is the +// controlPlaneReady reports whether the singleton answers, which is the // precondition for creating a cluster at all. +// +// Answering is the test rather than being unimpaired. Degraded is a control +// plane that answers -- a management API pod restarting behind a Service, a +// FoundationDB pod recycled without losing quorum, a metrics exporter that is +// not running at all -- and design-controlplane.md §4.3 says nothing holds on +// it. Waiting for Available instead let one non-essential component, whose +// absence loses no work and stops nothing, hold an entire expansion. +// +// Unavailable is the phase worth waiting for, because it is the one that means +// the probe failed or an essential component is at zero, and every call this +// expansion is about to make would fail. func (r *ClusterDeploymentConfigReconciler) controlPlaneReady( ctx context.Context, ) (bool, string) { @@ -325,11 +336,28 @@ func (r *ClusterDeploymentConfigReconciler) controlPlaneReady( return false, fmt.Sprintf("ControlPlane %s cannot be read: %v", singletonControlPlane, err) } - if controlPlane.Status.Phase != controlPlaneAvailable { - return false, fmt.Sprintf("ControlPlane %s is %s rather than %s", - singletonControlPlane, controlPlane.Status.Phase, controlPlaneAvailable) + + switch controlPlane.Status.Phase { + case simplyblockv1alpha2.ControlPlanePhaseAvailable: + return true, "" + case simplyblockv1alpha2.ControlPlanePhaseDegraded: + // Impaired and answering. The message it carries is the administrator's + // business and not this expansion's. + return true, "" + default: + return false, fmt.Sprintf("ControlPlane %s is %s, so it is not answering yet", + singletonControlPlane, phaseOrUnset(controlPlane.Status.Phase)) } - return true, "" +} + +// phaseOrUnset names the phase an object carries, or says it has none. A +// ControlPlane the operator has not reconciled yet has an empty phase, and a +// message reporting it as "" reads as a bug in the message. +func phaseOrUnset(phase simplyblockv1alpha2.ControlPlanePhase) string { + if phase == "" { + return "not reporting a phase yet" + } + return string(phase) } // everyUnfinishedDocument maps a Node event onto every document that has not @@ -553,8 +581,7 @@ func refusef(reason, format string, args ...any) error { return &refusedError{reason: reason, message: fmt.Sprintf(format, args...)} } -// The ControlPlane the expansion waits on, and the phase it waits for. -const ( - singletonControlPlane = "simplyblock" - controlPlaneAvailable = "Available" -) +// The ControlPlane the expansion waits on. Which phases let it through is +// controlPlaneReady's, because it is a decision about answering rather than a +// single value to compare against. +const singletonControlPlane = "simplyblock" diff --git a/operator/internal/controllers/deployment/controlplane_gate_test.go b/operator/internal/controllers/deployment/controlplane_gate_test.go new file mode 100644 index 000000000..4dfee85f7 --- /dev/null +++ b/operator/internal/controllers/deployment/controlplane_gate_test.go @@ -0,0 +1,97 @@ +// Which control-plane phases let a deployment proceed. +// +// The phase separates a control plane that answers from one that does not, and +// only the second is a reason to wait. design-controlplane.md §4.3 states it +// directly: a Degraded control plane is one that answers, so nothing holds on +// it. It is the phase a management API pod restarting behind a Service produces, +// and a FoundationDB pod recycled without losing quorum, and a metrics exporter +// that is not running at all. + +package deployment + +import ( + "context" + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/testsupport" +) + +// aControlPlaneIn builds the reconciler over a singleton in the phase given. +func aControlPlaneIn( + t *testing.T, phase simplyblockv1alpha2.ControlPlanePhase, +) *ClusterDeploymentConfigReconciler { + t.Helper() + scheme := testsupport.NewScheme(t, corev1.AddToScheme) + + controlPlane := &simplyblockv1alpha2.ControlPlane{ + ObjectMeta: metav1.ObjectMeta{Name: singletonControlPlane, Namespace: "simplyblock"}, + } + controlPlane.Status.Phase = phase + + return &ClusterDeploymentConfigReconciler{ + Client: fake.NewClientBuilder().WithScheme(scheme).WithObjects(controlPlane).Build(), + Scheme: scheme, + Namespace: "simplyblock", + } +} + +// TestADegradedControlPlaneDoesNotHoldTheDeployment is the gate this corrects. +// +// Regression: 2026-09-21-degraded-held-the-expansion — the gate compared the +// phase to Available and nothing else, so every phase that was not that word +// waited, Degraded included. On the deployment this was found on the whole +// expansion stopped at "ControlPlane simplyblock is Degraded rather than +// Available" because the FoundationDB metrics exporter was not running: a +// component that is non-essential by its own table, whose absence loses no work +// and stops nothing, held six workers out of a cluster. +func TestADegradedControlPlaneDoesNotHoldTheDeployment(t *testing.T) { + r := aControlPlaneIn(t, simplyblockv1alpha2.ControlPlanePhaseDegraded) + + ready, message := r.controlPlaneReady(context.Background()) + if !ready { + t.Errorf("a control plane that answers held the deployment: %s", message) + } +} + +// Available proceeds, which is the case that always worked. +func TestAnAvailableControlPlaneProceeds(t *testing.T) { + r := aControlPlaneIn(t, simplyblockv1alpha2.ControlPlanePhaseAvailable) + + if ready, message := r.controlPlaneReady(context.Background()); !ready { + t.Errorf("an available control plane held the deployment: %s", message) + } +} + +// Unavailable is the phase that means the control plane does not answer, and it +// is the one worth waiting for: every call the expansion is about to make would +// fail. +func TestAnUnavailableControlPlaneHolds(t *testing.T) { + r := aControlPlaneIn(t, simplyblockv1alpha2.ControlPlanePhaseUnavailable) + + ready, message := r.controlPlaneReady(context.Background()) + if ready { + t.Error("a control plane that does not answer let the deployment proceed") + } + if message == "" { + t.Error("the hold says nothing about why") + } +} + +// The phases before the control plane exists hold too. There is nothing to call +// yet, rather than something impaired. +func TestAControlPlaneStillBeingBuiltHolds(t *testing.T) { + for _, phase := range []simplyblockv1alpha2.ControlPlanePhase{ + simplyblockv1alpha2.ControlPlanePhaseInstalling, + "", + } { + r := aControlPlaneIn(t, phase) + if ready, _ := r.controlPlaneReady(context.Background()); ready { + t.Errorf("a control plane in %q let the deployment proceed", phase) + } + } +} From 8eb989d043b332a6644f73466a7071d9a02c5e1d Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Mon, 21 Sep 2026 10:10:04 +0100 Subject: [PATCH 112/206] updated doc design-csi-addons-replication.md --- .../designs/design-csi-addons-replication.md | 36 ++++++++++--------- 1 file changed, 19 insertions(+), 17 deletions(-) diff --git a/operator/docs/designs/design-csi-addons-replication.md b/operator/docs/designs/design-csi-addons-replication.md index c822c5789..74741973b 100644 --- a/operator/docs/designs/design-csi-addons-replication.md +++ b/operator/docs/designs/design-csi-addons-replication.md @@ -1,8 +1,8 @@ # Design Document: csi-addons Volume Replication -**Status:** Phase 2 Implemented +**Status:** Phase 3 Partially Implemented (§11) **Author:** Israel Geoffrey (geoffrey1330) -**Date:** 2026-09-16 (last updated 2026-09-17) +**Date:** 2026-09-16 (last updated 2026-09-21) **Test Plan:** [`tests/test-plan-csi-addons-replication.md`](../tests/test-plan-csi-addons-replication.md) --- @@ -13,7 +13,7 @@ |-------------|-------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|--------------| | **Phase 1** | Implemented | The csi-addons machinery and the steady-state contract: CRDs, controller-manager, sidecar, the Replication and csi-addons Identity gRPC services with `EnableVolumeReplication`, `DisableVolumeReplication`, and `GetVolumeReplicationInfo`, backed by a typed backend status endpoint | §4, §5.1, §6 | | **Phase 2** | Implemented | The lifecycle verbs: `PromoteVolume` (planned and forced), `DemoteVolume`, and `ResyncVolume`. Validation end to end against a Ramen `VolumeReplicationGroup` in async mode is still outstanding (§12, E-06/E-07) | §5.2, §9 | -| **Phase 3** | Planned | peerClasses convention and preflight, and the replication observability surface (lag, backlog, RPO compliance) | §7, §11 | +| **Phase 3** | Partially implemented | §11 (the Prometheus metrics) is implemented. §7.2 (peerClasses convention and preflight) is not started: it needs a design pass on cross-cluster Kubernetes API access first, since no mechanism for the operator to reach a peer Kubernetes cluster exists today | §7, §11 | | **Phase 4** | Planned | Test failover: the latest-replicated-snapshot read, the `drtest-*` conventions, and the two drill modes (bubble and test cluster) composed from clone, replication, and the real failover | §14 | Phase 1 is independently useful: a `VolumeReplication` object per PVC whose status truthfully reports the relationship, which no surface provides today. Phase 2 makes the object drivable, which is what Ramen actually needs. Phase 3 makes the whole thing operable at fleet scale. Phase 4 turns the same primitives into a rehearsal: a failover that can be drilled, in a bubble or against a test cluster, without touching production replication. @@ -26,10 +26,10 @@ The phase numbers above are this document's own, not the DR storage foundation g | # | Prerequisite | Kind | Blocks | Status | |------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------|---------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| P0-1 | A typed, steady-state per-volume replication status read: `GET .../volumes/{id}/replication/status` serving what `lvol_controller.get_replication_info` computes today (state, lag, outstanding bytes, failure counters), available for the volume's whole replicated life | Control plane (`sbcli`) | Phase 1 | Not shipped | -| P0-2 | Idempotent attach and detach: attaching a volume to the policy it already follows returns success, and detaching a non-attached volume returns success | Control plane (`sbcli`) | Phase 1 | Not shipped | -| P0-3 | A standalone demote verb: `POST .../volumes/{id}/replication/demote` that converges the peer while still serving (repeated snapshot-and-ship until the remaining delta is small), then quiesces, ships the final delta, confirms it landed on the peer, and fences the data path | Control plane (`sbcli`) | Phase 2 | Not shipped | -| P0-4 | An `rpo_target_seconds` field on `ReplicationPolicy`, so RPO compliance is computable against a declared target rather than the derived lag budget | Control plane (`sbcli`) | Phase 3 | Not shipped | +| P0-1 | A typed, steady-state per-volume replication status read: `GET .../volumes/{id}/replication/status` serving what `lvol_controller.get_replication_info` computes today (state, lag, outstanding bytes, failure counters), available for the volume's whole replicated life | Control plane (`sbcli`) | Phase 1 | Shipped: `GET .../volumes/{v}/replication/status` → `ReplicationStatusDTO` (`simplyblock_web/api/v2/cluster/storage_pool/volume/replication.py:58-69`) | +| P0-2 | Idempotent attach and detach: attaching a volume to the policy it already follows returns success, and detaching a non-attached volume returns success | Control plane (`sbcli`) | Phase 1 | Shipped: `replication_policy_controller.attach_policy`/`detach_policy` (`simplyblock_core/controllers/replication_policy_controller.py:208-262`) | +| P0-3 | A standalone demote verb: `POST .../volumes/{id}/replication/demote` that converges the peer while still serving (repeated snapshot-and-ship until the remaining delta is small), then quiesces, ships the final delta, confirms it landed on the peer, and fences the data path | Control plane (`sbcli`) | Phase 2 | Shipped: `POST .../volumes/{v}/replication/demote` → `lvol_controller.demote_lvol` | +| P0-4 | An `rpo_target_seconds` field on `ReplicationPolicy`, so RPO compliance is computable against a declared target rather than the derived lag budget | Control plane (`sbcli`) | Phase 3 | Shipped: `ReplicationPolicy.rpo_target_seconds` (`simplyblock_core/models/replication.py:89`), wired through the API (`PolicyParams.rpo_target_seconds`) and CLI (`--rpo-target-sec`) | | P0-5 | csi-addons upstream: the `VolumeReplication` and `VolumeReplicationClass` CRDs (`replication.storage.openshift.io/v1alpha1`), the kubernetes-csi-addons controller-manager image, and the csi-addons sidecar image | Ecosystem | Phase 1 | Vendored in the chart at v0.15.0 behind `csiaddons.create` (all twelve upstream CRDs, since the stock manager starts a controller per kind); sidecar wiring is Phase 1 | | P0-6 | A latest-replicated-snapshot read: per volume, and per consistency group as one complete generation, the newest fully replicated snapshot on the secondary addressed as a cloneable object | Control plane (`sbcli`) | Phase 4 | Shipped: `GET .../replication/relationships/{lvol}/latest-snapshot` and `GET .../replication/policies/{policy}/latest-generation` | @@ -356,21 +356,23 @@ The kubernetes-csi-addons controller-manager owns events on `VolumeReplication` | `PeerClassesVerified` | Normal | The preflight confirmed same-named classes and mutually pointing policies on both clusters | | `PeerClassesMismatch` | Warning | A replication-enabled class has no same-named peer, or the named policies do not pair; the message names the class and side | -### Prometheus Metrics +### Prometheus Metrics (Implemented) -Exported by the control plane, labeled `volume`, `policy`, and `peer_cluster`: +Exported by the control plane's existing v2 `Collector`-pattern exporter (`simplyblock_web/api/v2/metrics.py`), rebuilt from FDB on every scrape like every other series in that file. Labeled `lvol`/`lvol_name`/`pvc_name`/`pool`/`pool_name` (not the bare `volume` this section originally specified: the exporter's existing lvol-scoped metrics already use `lvol`/`lvol_name`, and joining the new series against them needs a shared label name) plus `policy`/`policy_name`/`peer_cluster`: -| Metric | Labels | Description | -|---------------------------------------------|------------------------------|---------------------------------------------------------------------------------| -| `simplyblock_replication_lag_seconds` | volume, policy, peer_cluster | Now minus the newest fully replicated snapshot's creation time. | -| `simplyblock_replication_backlog_bytes` | volume, policy, peer_cluster | `outstanding_bytes`: the queued-but-unshipped snapshot sizes. | -| `simplyblock_replication_last_sync_seconds` | volume, policy, peer_cluster | Duration of the last shipping cycle. | -| `simplyblock_replication_last_sync_bytes` | volume, policy, peer_cluster | Size of the last shipped snapshot. | -| `simplyblock_replication_rpo_violation` | volume, policy, peer_cluster | 1 while `lag_seconds` exceeds the policy's `rpo_target_seconds` (P0-4), else 0. | -| `simplyblock_replication_degraded` | volume, policy, peer_cluster | 1 while the status read's state is `degraded` or `error`. | +| Metric | Description | +|----------------------------------------------|---------------------------------------------------------------------------------| +| `simplyblock_replication_lag_seconds` | Now minus the newest fully replicated snapshot's creation time. | +| `simplyblock_replication_backlog_bytes` | `outstanding_bytes`: the queued-but-unshipped snapshot sizes. | +| `simplyblock_replication_last_sync_seconds` | Duration of the last shipping cycle. | +| `simplyblock_replication_last_sync_bytes` | Size of the last shipped snapshot. | +| `simplyblock_replication_rpo_violation` | 1 while `lag_seconds` exceeds the policy's `rpo_target_seconds` (P0-4), else 0. | +| `simplyblock_replication_degraded` | 1 while the status read's state is `degraded` or `error`. | `simplyblock_replication_rpo_violation` is the alert: it is the declared objective against the measured lag, which no timestamp alone can express. `simplyblock_replication_backlog_bytes` is the second load-bearing figure, because it is the input to any honest RTO estimate and the number that distinguishes "slow cycle" from "falling behind." One caveat is recorded rather than hidden: `outstanding_bytes` measures queued snapshot sizes, not dirty bytes written since the last snapshot, so intra-interval writes are invisible to it. A true dirty-delta figure needs storage-plane support and is future work. +Values are computed by `lvol_controller.get_replication_info_bulk`, a bulk-friendly sibling of `get_replication_info` (P0-1) that reads a whole cluster's job tasks, policies, and targets once rather than once per replicating volume -- `get_replication_info` itself makes several effectively global FDB scans internally (repeated per call), which is fine for a single-volume status read but not for a per-scrape loop across a fleet. `simplyblock_replication_last_sync_seconds`/`_bytes` are omitted per volume when no cycle has completed yet, and `rpo_violation` is omitted entirely for a volume whose policy declares no `rpo_target_seconds`, matching the exporter's existing "omit rather than fabricate" convention (`_health_family`) -- a 0 there would misread as "in compliance" absent a declared target. + --- ## 12. Testing Strategy From 5896a2ecc7bcfc1b42f02045098603bceaf7c3b2 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Mon, 21 Sep 2026 11:33:37 +0200 Subject: [PATCH 113/206] fix(operator): a node-add slot is held for the add, not for the UUID The control plane writes the node object at the start of add_node, with status in_creation, and finishes minutes later. The slot was released when the UUID appeared, so every add read as finished the moment it began. On a six-worker cluster with maxParallelNodeAdds 1 that produced five node_add tasks running at once, five SPDK pods coming up together, and four StorageNodes sitting at in_creation with a UUID each. The cap exists because an add reboots its host and two hosts rebooting at once cost the control plane its own fault tolerance, so the one value it must not be measured in is the one that is true before the reboot has happened. A slot is now held until the node leaves in_creation. Every other status is a node the control plane has finished with, whatever it then thinks of it: one that came up and went unreachable has still had its add, and holding the cluster's only slot for it would stop the fleet for a node nothing is working on. Both halves move together. The reap keeps an entry whose holder is still being created, and the release goes from the end of the provisioning machine -- which is the UUID, and therefore the wrong moment -- to the steady-state pass that first reads a finished creation back. The two ends that never reach steady state, the deadline and the delete, release where they did. An add that hangs now holds the cap rather than admitting the fleet behind it. That is the cap doing its job, and the event says so with the age the holder has been holding for. Test plan: U-402, U-403. Co-Authored-By: Claude Opus 5 (1M context) --- operator/docs/tests/test-plan-storagenode.md | 2 + .../controllers/node/provisioningslots.go | 14 +++++- .../node/provisioningslots_test.go | 50 +++++++++++++++++++ .../node/storagenode_controller.go | 20 ++++++-- 4 files changed, 81 insertions(+), 5 deletions(-) diff --git a/operator/docs/tests/test-plan-storagenode.md b/operator/docs/tests/test-plan-storagenode.md index 775d6fc4c..4552288c6 100644 --- a/operator/docs/tests/test-plan-storagenode.md +++ b/operator/docs/tests/test-plan-storagenode.md @@ -104,6 +104,8 @@ Files: `operator/internal/controllers/node/provisioning_test.go`, | U-399 | A worker's second socket takes no second slot and resolves against the first | Regression | `TestASecondSocketDoesNotTakeASecondSlot` | | U-400 | A backend node appearing while a node queues for a slot is adopted, not re-added | Regression | `TestANodeWaitingForASlotAdoptsTheNodeThatAppeared` | | U-401 | No backend node for the worker: the queue takes its slot as before | Negative | `TestANodeWithNoBackendNodeStillTakesItsSlot` | +| U-402 | A slot is held while the control plane still reports the node in_creation | Regression | `TestASlotIsHeldUntilTheAddIsFinished` | +| U-403 | The slot goes back once the node leaves in_creation | Positive | `TestTheSlotGoesBackWhenTheNodeLeavesCreation` | ### Entity: The Provisioning Claim (design §4.2) diff --git a/operator/internal/controllers/node/provisioningslots.go b/operator/internal/controllers/node/provisioningslots.go index 5403070c5..efe4d7c73 100644 --- a/operator/internal/controllers/node/provisioningslots.go +++ b/operator/internal/controllers/node/provisioningslots.go @@ -124,7 +124,7 @@ func (r *StorageNodeReconciler) heldSlots( } return nil, err } - if holder.Status.UUID != "" || + if addFinished(&holder) || holder.Status.Phase == simplyblockv1alpha2.StorageNodePhaseFailed || !holder.DeletionTimestamp.IsZero() { continue @@ -180,3 +180,15 @@ func describeSlots(slots []simplyblockv1alpha2.ProvisioningSlot) string { slices.Sort(parts) return strings.Join(parts, ", ") } + +// addFinished reports whether the add a slot was taken for is over. +// +// A node with no UUID has not been created at all, and one the control plane +// still reports as in_creation is being created now. Every other status is a +// node the control plane has finished with, whatever it then thinks of it: a +// node that came up and went unreachable has still had its add, and holding the +// cluster's only slot for it would stop the fleet for a node nothing is working +// on. +func addFinished(node *simplyblockv1alpha2.StorageNode) bool { + return node.Status.UUID != "" && node.Status.Status != nodeStatusInCreation +} diff --git a/operator/internal/controllers/node/provisioningslots_test.go b/operator/internal/controllers/node/provisioningslots_test.go index 50ea4e7cc..dfad98dae 100644 --- a/operator/internal/controllers/node/provisioningslots_test.go +++ b/operator/internal/controllers/node/provisioningslots_test.go @@ -318,3 +318,53 @@ func TestASecondSocketDoesNotTakeASecondSlot(t *testing.T) { t.Errorf("one worker holds %d slots", len(slots)) } } + +// TestASlotIsHeldUntilTheAddIsFinished is what the cap is measured in. +// +// Regression: 2026-09-21-the-slot-was-released-when-the-uuid-appeared — the +// control plane writes the node object at the start of add_node, with +// status=in_creation, so its UUID exists seconds into an add that runs for +// minutes. Releasing on the UUID made every add look finished the moment it +// began: on a six-worker cluster with maxParallelNodeAdds 1, five node_add tasks +// ran at once and five SPDK pods came up together. The cap exists because an add +// reboots its host. +func TestASlotIsHeldUntilTheAddIsFinished(t *testing.T) { + adding, waiting := waitingNode("w1"), waitingNode("w2") + adding.Status.UUID = "10fe8da7-55d6-4162-a5a2-9575421ee31f" + adding.Status.Status = nodeStatusInCreation + + cluster := slotCluster(1) + cluster.Status.ProvisioningSlots = []simplyblockv1alpha2.ProvisioningSlot{ + {Worker: "w1", Node: adding.Name, TakenAt: metav1.Now()}, + } + reconcilers := blindReconcilers(t, cluster, adding, waiting) + + next, _, err := reconcilers[1].awaitSlot(context.Background(), waiting, cluster.DeepCopy()) + if err == nil && next == stepPosting { + t.Error("a second add was posted while the first worker was still being created") + } + if slots := storedCluster(t, reconcilers[1]).Status.ProvisioningSlots; len(slots) != 1 { + t.Errorf("the cluster records %d slots while one add is running", len(slots)) + } +} + +// Once the add is over the slot goes back, which is what keeps the queue moving. +func TestTheSlotGoesBackWhenTheNodeLeavesCreation(t *testing.T) { + added, waiting := waitingNode("w1"), waitingNode("w2") + added.Status.UUID = "10fe8da7-55d6-4162-a5a2-9575421ee31f" + added.Status.Status = nodeStatusOnline + + cluster := slotCluster(1) + cluster.Status.ProvisioningSlots = []simplyblockv1alpha2.ProvisioningSlot{ + {Worker: "w1", Node: added.Name, TakenAt: metav1.Now()}, + } + reconcilers := blindReconcilers(t, cluster, added, waiting) + + next, _, err := reconcilers[1].awaitSlot(context.Background(), waiting, cluster.DeepCopy()) + if err != nil { + t.Fatalf("awaitSlot: %v", err) + } + if next != stepPosting { + t.Errorf("the next worker was sent to %s after the previous add finished", next) + } +} diff --git a/operator/internal/controllers/node/storagenode_controller.go b/operator/internal/controllers/node/storagenode_controller.go index edea8af61..4d240174d 100644 --- a/operator/internal/controllers/node/storagenode_controller.go +++ b/operator/internal/controllers/node/storagenode_controller.go @@ -383,10 +383,12 @@ func (r *StorageNodeReconciler) provision( // the step that holds while the worker is away, and a finished Resolving // is still the end of the path. // - // The add is over, so the slot it was taken for goes back. This is the - // successful one of the three ends, and the other two are the deadline in - // fail and the object going away in teardown. - return ctrl.Result{RequeueAfter: nodeAdvance}, r.releaseSlot(ctx, node) + // The slot is not given back here. Reaching this step means the node has + // a UUID, and a UUID is the start of the add rather than the end of it: + // the control plane writes the node object when add_node begins, with + // status in_creation, and finishes minutes later. Steady state releases + // it once the control plane says the creation is done. + return ctrl.Result{RequeueAfter: nodeAdvance}, nil } if err := machine.TransitionTo(ctx, next); err != nil { @@ -938,6 +940,16 @@ func (r *StorageNodeReconciler) syncStatus( fmt.Sprintf("Node %s is online and carrying its share", node.Status.UUID)) } r.observePhase(node) + + // The node-add slot goes back here rather than when the UUID arrived. This + // is the first pass on which the control plane has said the creation is + // done, which is what the cluster's cap is counting: an add reboots its + // host, and the reboot is not over when the object appears. + if addFinished(node) { + if err := r.releaseSlot(ctx, node); err != nil { + return ctrl.Result{RequeueAfter: nodeRetry}, err + } + } return ctrl.Result{RequeueAfter: nodeRetry}, nil } From 0879b0af2d7ca0e8c8b5791b372cf0691e9fc5c9 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Mon, 21 Sep 2026 12:09:34 +0200 Subject: [PATCH 114/206] fix(operator): a node with no configuration says so Two faults on the path from the cluster's per-node ConfigMap to the configure, both of the same shape: a missing input answered with a value rather than with a failure, and reported later by the thing that could not use it. The writer answered a missing entry with `touch`, so an empty environment file was indistinguishable from a configured one. The configure then ran with nothing set and reported that max-lvol 0 was invalid, which names neither the node, nor the file, nor the ConfigMap it came from. That is permanent, not transient. The writer runs once per pod. A pod that started before its ConfigMap volume was populated keeps the empty file for the rest of its life, because restarting a later init container never re-runs an earlier one: the configure crash-loops against an empty environment while the ConfigMap beside it has been correct for minutes. One worker of six was in that state, and the other five were configured from the same ConfigMap. The writer now exits non-zero, which puts the retry on the container that can fix it. The zero was the second fault and stands on its own. --max-lvol=${MAX_SUBSYS_COUNT:-0} turned an absent value into one the control plane refuses, so a cluster that states no maxSubsystemCount is refused the same way a pod with no entry is. Absent arguments are left out now, and no opinion means the control plane's own default rather than a number this side invented. Both scripts move into nodeenv.go to be testable at all. They were locals inside the DaemonSet builder, which is why neither fault had a test: there was nothing to assert against short of rendering a workload and reading shell out of it. The configure's first argument is conditional now where it used to be unconditional, under the same set -e. A failing && list does not exit either shell, verified in both. Test plan: the two cases in nodeenv_test.go, each red against the previous scripts. Co-Authored-By: Claude Opus 5 (1M context) --- operator/internal/utils/nodeenv.go | 82 +++++++++++++++++++ operator/internal/utils/nodeenv_test.go | 66 +++++++++++++++ .../internal/utils/storage_node_workload.go | 27 +----- 3 files changed, 150 insertions(+), 25 deletions(-) create mode 100644 operator/internal/utils/nodeenv.go create mode 100644 operator/internal/utils/nodeenv_test.go diff --git a/operator/internal/utils/nodeenv.go b/operator/internal/utils/nodeenv.go new file mode 100644 index 000000000..66d65677a --- /dev/null +++ b/operator/internal/utils/nodeenv.go @@ -0,0 +1,82 @@ +// The two init containers a storage-node pod runs before its API starts: one +// that selects this node's entry out of the cluster's per-node ConfigMap, and +// one that runs the configure with it. +// +// They are separate from the DaemonSet that carries them because what they get +// wrong, they get wrong quietly. A shell script that cannot find its input +// answers with an empty file unless it is told not to, and the container after +// it then fails describing a value rather than the input that was missing. + +package utils + +// Where the per-node entries are mounted, and where the selected one is written +// for the configure to source. The second is an emptyDir shared by the two init +// containers and the pod that follows them. +const ( + perNodeConfigDir = "/etc/per-node-config" + nodeEnvFile = "/etc/node-env/env.sh" +) + +// configureInvocation is the tail every configure script ends with, after the +// per-node and fleet arguments have been assembled into ARGS. +const configureInvocation = `" +eval sudo -E python3 simplyblock_web/node_configure.py ${ARGS} +` + +// nodeEnvScripts returns the configure script and the writer that feeds it. +// +// The configure script is returned without its closing quote: the caller appends +// the fleet-level arguments and configureInvocation, which is what keeps the +// per-node and fleet halves in one string without either being able to inject +// into the other. +func nodeEnvScripts() (configure, writer string) { + return nodeConfigureScript(), nodeEnvWriterScript() +} + +// nodeEnvWriterScript selects this node's entry. +// +// A missing entry is a failure rather than an empty file. The writer runs once +// per pod, and a pod that started before its ConfigMap volume was populated +// would otherwise keep an empty environment for the rest of its life: the +// configure after it crash-loops, and restarting a later init container never +// re-runs an earlier one. Exiting non-zero is what puts the retry on this +// container, where the next attempt reads a volume that has caught up. +// +// HOSTNAME is the node's name, injected from spec.nodeName, and the ConfigMap is +// keyed by it. +func nodeEnvWriterScript() string { + return `set -e +mkdir -p /etc/node-env +if [ ! -f ` + perNodeConfigDir + `/${HOSTNAME} ]; then + echo "no per-node configuration for ${HOSTNAME} in ` + perNodeConfigDir + `" >&2 + echo "the cluster's per-node ConfigMap has no entry for this node yet" >&2 + exit 1 +fi +cp ` + perNodeConfigDir + `/${HOSTNAME} ` + nodeEnvFile + ` +` +} + +// nodeConfigureScript assembles the configure's arguments from that entry. +// +// An unset value is left out rather than sent as a default this side invented. +// --max-lvol=0 was such a default, and zero is refused because max-lvol must be +// a positive integer. A cluster that states no maxSubsystemCount has no opinion +// about it, and no opinion is the control plane's own default, not zero. +func nodeConfigureScript() string { + return `set -e +[ -f ` + nodeEnvFile + ` ] && . ` + nodeEnvFile + ` +ARGS="" +[ -n "${MAX_SUBSYS_COUNT}" ] && ARGS="${ARGS} --max-lvol=${MAX_SUBSYS_COUNT}" +[ -n "${PCI_ALLOWED}" ] && ARGS="${ARGS} --pci-allowed=\"${PCI_ALLOWED}\"" +[ -n "${PCI_BLOCKED}" ] && ARGS="${ARGS} --pci-blocked=\"${PCI_BLOCKED}\"" +[ -n "${NVME_DEVICES}" ] && ARGS="${ARGS} --nvme-devices=\"${NVME_DEVICES}\"" +[ -n "${DEVICE_MODEL}" ] && ARGS="${ARGS} --device-model=\"${DEVICE_MODEL}\"" +[ -n "${SIZE_RANGE}" ] && ARGS="${ARGS} --size-range=\"${SIZE_RANGE}\"" +[ "${LBLK}" = "true" ] && ARGS="${ARGS} --lblk" +[ -n "${BLK_NAMES}" ] && ARGS="${ARGS} --blk-names=\"${BLK_NAMES}\"" +[ -n "${BLK_NAMES_EXCLUDE}" ] && ARGS="${ARGS} --blk-names-exclude=\"${BLK_NAMES_EXCLUDE}\"" +[ -n "${BLK_SERIALS}" ] && ARGS="${ARGS} --blk-serials=\"${BLK_SERIALS}\"" +[ -n "${LBLK_JM_PERCENT}" ] && ARGS="${ARGS} --jm-percent=\"${LBLK_JM_PERCENT}\"" +[ "${LBLK_FORCE_FORMAT}" = "true" ] && ARGS="${ARGS} --force-format" +ARGS="${ARGS}` +} diff --git a/operator/internal/utils/nodeenv_test.go b/operator/internal/utils/nodeenv_test.go new file mode 100644 index 000000000..c0da0758e --- /dev/null +++ b/operator/internal/utils/nodeenv_test.go @@ -0,0 +1,66 @@ +// The two init containers that turn a per-node ConfigMap into a configure run. +// +// They are shell, so what they do wrong they do quietly. The assertions here are +// on the scripts as text, which is the weakest kind of test and the only one +// available short of a cluster: what they are guarding is not a value but a +// failure that must not be swallowed. + +package utils + +import ( + "strings" + "testing" +) + +// scripts returns the two init-container scripts of a storage-node workload. +func scripts(t *testing.T) (writer, generator string) { + t.Helper() + g, w := nodeEnvScripts() + if w == "" || g == "" { + t.Fatal("the workload renders no init scripts") + } + return w, g +} + +// TestAMissingPerNodeEntryFailsTheWriter is the defect this guards. +// +// Regression: 2026-09-21-an-absent-node-entry-was-written-as-an-empty-one — the +// writer answered a missing entry with `touch`, so an empty env file was +// indistinguishable from a configured one. The generator then ran with nothing +// set and reported that max-lvol 0 was invalid, which names neither the +// node nor the file nor the ConfigMap. +// +// It matters because the writer runs once. A pod whose ConfigMap volume was not +// yet populated when it started keeps the empty file for its whole life: the +// generator crash-loops, and restarting a later init container never re-runs an +// earlier one. Failing instead is what makes the retry pick the entry up. +func TestAMissingPerNodeEntryFailsTheWriter(t *testing.T) { + writer, _ := scripts(t) + + if strings.Contains(writer, "touch /etc/node-env/env.sh") { + t.Error("a missing per-node entry still produces an empty env file, " + + "which the generator cannot tell from a configured one") + } + if !strings.Contains(writer, "exit 1") { + t.Error("the writer does not fail on a missing entry, so nothing retries it") + } +} + +// TestAnAbsentSubsystemCountIsNotSentAsZero covers the other half. +// +// Regression: 2026-09-21-max-lvol-zero — the generator passed +// --max-lvol=${MAX_SUBSYS_COUNT:-0}, and zero is a value the control plane +// refuses, saying that max-lvol must be a positive integer. A cluster that +// states no +// maxSubsystemCount is a cluster with no opinion about it, which is the control +// plane's default rather than zero. +func TestAnAbsentSubsystemCountIsNotSentAsZero(t *testing.T) { + _, generator := scripts(t) + + if strings.Contains(generator, "${MAX_SUBSYS_COUNT:-0}") { + t.Error("an unset subsystem count is sent as --max-lvol=0, which is refused") + } + if !strings.Contains(generator, "MAX_SUBSYS_COUNT") { + t.Error("the subsystem count is not passed at all") + } +} diff --git a/operator/internal/utils/storage_node_workload.go b/operator/internal/utils/storage_node_workload.go index cdb3652e0..cab562d1c 100644 --- a/operator/internal/utils/storage_node_workload.go +++ b/operator/internal/utils/storage_node_workload.go @@ -72,33 +72,10 @@ func BuildStorageNodeDaemonSet( // The init container sources the per-node env file (written by node-env-writer) // so that node_configure.py receives per-node values for each pod. - initScript := `set -e -[ -f /etc/node-env/env.sh ] && . /etc/node-env/env.sh -ARGS="--max-lvol=${MAX_SUBSYS_COUNT:-0}" -[ -n "${PCI_ALLOWED}" ] && ARGS="${ARGS} --pci-allowed=\"${PCI_ALLOWED}\"" -[ -n "${PCI_BLOCKED}" ] && ARGS="${ARGS} --pci-blocked=\"${PCI_BLOCKED}\"" -[ -n "${NVME_DEVICES}" ] && ARGS="${ARGS} --nvme-devices=\"${NVME_DEVICES}\"" -[ -n "${DEVICE_MODEL}" ] && ARGS="${ARGS} --device-model=\"${DEVICE_MODEL}\"" -[ -n "${SIZE_RANGE}" ] && ARGS="${ARGS} --size-range=\"${SIZE_RANGE}\"" -[ "${LBLK}" = "true" ] && ARGS="${ARGS} --lblk" -[ -n "${BLK_NAMES}" ] && ARGS="${ARGS} --blk-names=\"${BLK_NAMES}\"" -[ -n "${BLK_NAMES_EXCLUDE}" ] && ARGS="${ARGS} --blk-names-exclude=\"${BLK_NAMES_EXCLUDE}\"" -[ -n "${BLK_SERIALS}" ] && ARGS="${ARGS} --blk-serials=\"${BLK_SERIALS}\"" -[ -n "${LBLK_JM_PERCENT}" ] && ARGS="${ARGS} --jm-percent=\"${LBLK_JM_PERCENT}\"" -[ "${LBLK_FORCE_FORMAT}" = "true" ] && ARGS="${ARGS} --force-format" -ARGS="${ARGS}` + fleetArgs + `" -eval sudo -E python3 simplyblock_web/node_configure.py ${ARGS} -` + initScript, nodeEnvWriterScript := nodeEnvScripts() + initScript += fleetArgs + configureInvocation initCmd := []string{"sh", "-c", initScript} - // nodeEnvWriterScript copies the per-node env file from the mounted ConfigMap - // (keyed by hostname) into the shared node-env emptyDir volume. - nodeEnvWriterScript := `mkdir -p /etc/node-env -if [ -f /etc/per-node-config/${HOSTNAME} ]; then - cp /etc/per-node-config/${HOSTNAME} /etc/node-env/env.sh -else - touch /etc/node-env/env.sh -fi` nodeEnvWriterCmd := []string{"sh", "-c", nodeEnvWriterScript} imagePullPolicy := wl.ImagePullPolicy From 818a6de027bc9fe94789b4d9d1c84899a724b53f Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Mon, 21 Sep 2026 12:25:37 +0200 Subject: [PATCH 115/206] fix(operator): a node with no subsystem count stops rather than asking for zero 0879b0af got this half backwards. It left --max-lvol out when MAX_SUBSYS_COUNT was empty, on the reasoning that a cluster stating no maxSubsystemCount has no opinion and no opinion means the control plane's default. The field says otherwise, and says it in its own documentation: maxSubsystemCount is Required, bounded at 10, and "a node that receives no value fails config generation outright rather than falling back to a default". Leaving the argument out is precisely the fallback the field forbids, and it would have configured a node from a number the cluster never stated -- silently, which is worse than the crash loop it replaced. Being Required also settles where the empty value came from. A valid cluster always has one, so an entry without it did not come from a valid cluster: it is an empty file, or one written before the value was known, which is the same absent configuration the writer now fails on. So the configure stops, and says which file and which field. What was wrong with the original was never that it stopped -- it was that it stopped by sending zero, so the control plane refused a value nobody had set and the message named neither the node nor the configuration that never arrived. Co-Authored-By: Claude Opus 5 (1M context) --- operator/internal/utils/nodeenv.go | 24 +++++++++++++++----- operator/internal/utils/nodeenv_test.go | 29 ++++++++++++++++--------- 2 files changed, 37 insertions(+), 16 deletions(-) diff --git a/operator/internal/utils/nodeenv.go b/operator/internal/utils/nodeenv.go index 66d65677a..102915f7a 100644 --- a/operator/internal/utils/nodeenv.go +++ b/operator/internal/utils/nodeenv.go @@ -58,15 +58,27 @@ cp ` + perNodeConfigDir + `/${HOSTNAME} ` + nodeEnvFile + ` // nodeConfigureScript assembles the configure's arguments from that entry. // -// An unset value is left out rather than sent as a default this side invented. -// --max-lvol=0 was such a default, and zero is refused because max-lvol must be -// a positive integer. A cluster that states no maxSubsystemCount has no opinion -// about it, and no opinion is the control plane's own default, not zero. +// A missing subsystem count stops the configure rather than standing in for it. +// StorageCluster.spec.maxSubsystemCount is Required and bounded at 10, so a +// valid cluster always has one, and its absence here means the entry this pod +// read did not come from a valid cluster -- an empty file, or one written before +// the value was known. The field's own contract is that a node receiving no +// value fails config generation outright rather than falling back to a default, +// which is what this preserves. +// +// What it does not do is express that by sending zero. Zero was refused by the +// control plane for being zero, so the message named a value nobody set instead +// of the configuration that never arrived. func nodeConfigureScript() string { return `set -e [ -f ` + nodeEnvFile + ` ] && . ` + nodeEnvFile + ` -ARGS="" -[ -n "${MAX_SUBSYS_COUNT}" ] && ARGS="${ARGS} --max-lvol=${MAX_SUBSYS_COUNT}" +if [ -z "${MAX_SUBSYS_COUNT}" ]; then + echo "MAX_SUBSYS_COUNT is not set in ` + nodeEnvFile + `" >&2 + echo "it carries the cluster's maxSubsystemCount, which is required, so this node" >&2 + echo "has no configuration to generate from" >&2 + exit 1 +fi +ARGS="--max-lvol=${MAX_SUBSYS_COUNT}" [ -n "${PCI_ALLOWED}" ] && ARGS="${ARGS} --pci-allowed=\"${PCI_ALLOWED}\"" [ -n "${PCI_BLOCKED}" ] && ARGS="${ARGS} --pci-blocked=\"${PCI_BLOCKED}\"" [ -n "${NVME_DEVICES}" ] && ARGS="${ARGS} --nvme-devices=\"${NVME_DEVICES}\"" diff --git a/operator/internal/utils/nodeenv_test.go b/operator/internal/utils/nodeenv_test.go index c0da0758e..9d7bbd0a4 100644 --- a/operator/internal/utils/nodeenv_test.go +++ b/operator/internal/utils/nodeenv_test.go @@ -46,21 +46,30 @@ func TestAMissingPerNodeEntryFailsTheWriter(t *testing.T) { } } -// TestAnAbsentSubsystemCountIsNotSentAsZero covers the other half. +// TestAnAbsentSubsystemCountStopsTheConfigure covers the other half. // // Regression: 2026-09-21-max-lvol-zero — the generator passed -// --max-lvol=${MAX_SUBSYS_COUNT:-0}, and zero is a value the control plane -// refuses, saying that max-lvol must be a positive integer. A cluster that -// states no -// maxSubsystemCount is a cluster with no opinion about it, which is the control -// plane's default rather than zero. -func TestAnAbsentSubsystemCountIsNotSentAsZero(t *testing.T) { +// --max-lvol=${MAX_SUBSYS_COUNT:-0}, so a node that had received no +// configuration asked the control plane for zero subsystems and was refused for +// asking for zero. The message named a value nobody had set rather than the +// configuration that never arrived. +// +// Stopping is the contract rather than a choice: spec.maxSubsystemCount is +// Required and bounded at 10, and its documentation says a node receiving no +// value fails config generation outright rather than falling back to a default. +// So the fix is not to leave the argument out -- that would be the fallback the +// field forbids -- but to fail where the cause can still be named. +func TestAnAbsentSubsystemCountStopsTheConfigure(t *testing.T) { _, generator := scripts(t) if strings.Contains(generator, "${MAX_SUBSYS_COUNT:-0}") { - t.Error("an unset subsystem count is sent as --max-lvol=0, which is refused") + t.Error("an unset subsystem count is sent as --max-lvol=0, and refused for being zero") + } + if !strings.Contains(generator, `exit 1`) { + t.Error("an unset subsystem count does not stop the configure, so the node is " + + "configured from a default the cluster never stated") } - if !strings.Contains(generator, "MAX_SUBSYS_COUNT") { - t.Error("the subsystem count is not passed at all") + if !strings.Contains(generator, "--max-lvol=${MAX_SUBSYS_COUNT}") { + t.Error("the subsystem count the cluster did state is not passed") } } From fe82d5b100732038f163fe22539b3e4de551568a Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Mon, 21 Sep 2026 12:59:39 +0200 Subject: [PATCH 116/206] fix(operator): the workers enrolled and the workers configured are one set The pass writes the per-node ConfigMap and then enrolls the workers, and a worker is only schedulable once enrollment has labeled it. That ordering is what stops a storage-node pod starting against an entry that is not there, and it held only while both halves were looking at the same node set. They were not: each did its own List, so a StorageNode that appeared between the two had its worker enrolled by a pass that never wrote its entry. The pod then started, found no entry, and reached the node configuration script with nothing set. What it reported was that max-lvol 0 is not a positive integer, which names neither the node, nor the file, nor the ConfigMap the file was supposed to come from. ReconcileConfig's own comment predicted this failure and named ordering as the answer to it. Both halves now take one reading. clusterNodes does the filtering once, so "this cluster's nodes" means one thing rather than two lists that agree by convention: a node on its way out is not a reason to keep its worker enrolled, and it is not a reason to write its entry either. This closes the window inside the operator. It does not close the one below it: a pod's view of a ConfigMap is kubelet's cached copy, which can lag the object the API server holds, and no ordering on this side reaches that. What covers the rest is the writer failing instead of writing an empty file (0879b0af), so the init container retries until the volume has caught up. Test plan: U-404. Co-Authored-By: Claude Opus 5 (1M context) --- operator/docs/tests/test-plan-storagenode.md | 1 + .../controllers/node/nodesnapshot_test.go | 136 ++++++++++++++++++ .../controllers/node/pernodeconfig.go | 34 ++--- .../controllers/node/workload_controller.go | 73 +++++++--- 4 files changed, 205 insertions(+), 39 deletions(-) create mode 100644 operator/internal/controllers/node/nodesnapshot_test.go diff --git a/operator/docs/tests/test-plan-storagenode.md b/operator/docs/tests/test-plan-storagenode.md index 4552288c6..5aa02cbd1 100644 --- a/operator/docs/tests/test-plan-storagenode.md +++ b/operator/docs/tests/test-plan-storagenode.md @@ -106,6 +106,7 @@ Files: `operator/internal/controllers/node/provisioning_test.go`, | U-401 | No backend node for the worker: the queue takes its slot as before | Negative | `TestANodeWithNoBackendNodeStillTakesItsSlot` | | U-402 | A slot is held while the control plane still reports the node in_creation | Regression | `TestASlotIsHeldUntilTheAddIsFinished` | | U-403 | The slot goes back once the node leaves in_creation | Positive | `TestTheSlotGoesBackWhenTheNodeLeavesCreation` | +| U-404 | Every enrolled worker has a per-node entry, across a node set that grows mid-pass | Regression | `TestEveryEnrolledWorkerHasAnEntry` | ### Entity: The Provisioning Claim (design §4.2) diff --git a/operator/internal/controllers/node/nodesnapshot_test.go b/operator/internal/controllers/node/nodesnapshot_test.go new file mode 100644 index 000000000..904f7721d --- /dev/null +++ b/operator/internal/controllers/node/nodesnapshot_test.go @@ -0,0 +1,136 @@ +// The workers enrolled and the workers configured come from one reading. +// +// The pass writes the per-node ConfigMap and then labels the workers, and a +// worker is only schedulable once it carries the label. That ordering is what +// stops a storage-node pod starting against an entry that is not there yet -- and +// it only holds while both halves are looking at the same set of nodes. + +package node + +import ( + "context" + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + "sigs.k8s.io/controller-runtime/pkg/client/interceptor" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/testsupport" +) + +// aClusterNode is one storage node of the cluster, on the worker given. +func aClusterNode(name, worker string) *simplyblockv1alpha2.StorageNode { + return &simplyblockv1alpha2.StorageNode{ + ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: "simplyblock"}, + Spec: simplyblockv1alpha2.StorageNodeSpec{ + ClusterRef: "c", + WorkerNode: worker, + Config: simplyblockv1alpha2.StorageNodeConfig{ + Sizing: simplyblockv1alpha2.StorageNodeSizing{VCPUCount: ptrTo32(4)}, + }, + }, + } +} + +func ptrTo32(v int32) *int32 { return &v } + +// aSnapshotCluster carries the sizing ReconcileConfig refuses to render without. +func aSnapshotCluster() *simplyblockv1alpha2.StorageCluster { + return &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{Name: "c", Namespace: "simplyblock"}, + Spec: simplyblockv1alpha2.StorageClusterSpec{ + MaxSubsystemCount: ptrTo32(30), + VCPUCount: ptrTo32(4), + }, + } +} + +// TestEveryEnrolledWorkerHasAnEntry is the window this closes. +// +// Regression: 2026-09-21-two-readings-of-one-node-set — the ConfigMap was built +// from one List and the workers were labeled from another, so a StorageNode that +// appeared between the two had its worker enrolled in a pass that never wrote its +// entry. The pod then started, found no entry, and reached the configure with +// nothing set: on the deployment this was found on, one worker of six sat in +// CrashLoopBackOff reporting an invalid max-lvol of 0 while the ConfigMap beside +// it had been correct for minutes. +// +// The second List is what the interceptor below adds a node to, which is the +// cache advancing mid-pass. Both halves reading one snapshot is what makes the +// two sets the same set rather than two sets that usually agree. +func TestEveryEnrolledWorkerHasAnEntry(t *testing.T) { + scheme := testsupport.NewScheme(t, corev1.AddToScheme) + cluster := aSnapshotCluster() + early := aClusterNode("n-w1", "w1") + late := aClusterNode("n-w2", "w2") + + workers := []client.Object{ + &corev1.Node{ObjectMeta: metav1.ObjectMeta{Name: "w1"}}, + &corev1.Node{ObjectMeta: metav1.ObjectMeta{Name: "w2"}}, + } + base := fake.NewClientBuilder().WithScheme(scheme). + WithObjects(append([]client.Object{cluster, early, late}, workers...)...). + Build() + + // The first reading of the node set sees one node, and every reading after it + // sees both: the shape of a cache catching up in the middle of a pass. + reads := 0 + racing := interceptor.NewClient(base, interceptor.Funcs{ + List: func( + ctx context.Context, c client.WithWatch, list client.ObjectList, opts ...client.ListOption, + ) error { + if err := c.List(ctx, list, opts...); err != nil { + return err + } + nodes, ok := list.(*simplyblockv1alpha2.StorageNodeList) + if !ok { + return nil + } + reads++ + if reads == 1 { + nodes.Items = []simplyblockv1alpha2.StorageNode{*early} + } + return nil + }, + }) + + r := &StorageNodeWorkloadReconciler{ + Client: racing, + Scheme: scheme, + Namespace: "simplyblock", + Workload: &Workload{Client: racing}, + } + + snapshot, err := r.clusterNodes(context.Background(), cluster) + if err != nil { + t.Fatalf("read the node set: %v", err) + } + if err := r.Workload.ReconcileConfig(context.Background(), cluster, snapshot); err != nil { + t.Fatalf("write the per-node config: %v", err) + } + if err := r.enrollWorkers(context.Background(), cluster, snapshot); err != nil { + t.Fatalf("enroll the workers: %v", err) + } + + var config corev1.ConfigMap + key := client.ObjectKey{Namespace: "simplyblock", Name: PerNodeConfigMapName("c")} + if err := base.Get(context.Background(), key, &config); err != nil { + t.Fatalf("read the per-node ConfigMap: %v", err) + } + + for _, worker := range []string{"w1", "w2"} { + var object corev1.Node + if err := base.Get(context.Background(), client.ObjectKey{Name: worker}, &object); err != nil { + t.Fatalf("read worker %s: %v", worker, err) + } + enrolled := len(object.Labels) > 0 + _, configured := config.Data[worker] + if enrolled && !configured { + t.Errorf("worker %s is enrolled with no entry, so a pod can start on it "+ + "with nothing to configure from", worker) + } + } +} diff --git a/operator/internal/controllers/node/pernodeconfig.go b/operator/internal/controllers/node/pernodeconfig.go index db67058eb..a0807575f 100644 --- a/operator/internal/controllers/node/pernodeconfig.go +++ b/operator/internal/controllers/node/pernodeconfig.go @@ -45,15 +45,24 @@ func PerNodeConfigMapName(cluster string) string { return cluster + "-per-node-config" } -// ReconcileConfig writes the ConfigMap from the cluster's nodes. +// ReconcileConfig writes the ConfigMap from the node set it is given. // -// It is written before the DaemonSet on every pass. A pod that starts against a -// missing or empty entry reaches the node configuration script with -// --max-subsys-count=0 and fails there, which is a long way from the cause. For -// the same reason, a cluster missing its required sizing is refused with an error +// It is written before the workers are enrolled on every pass, and a worker is +// only schedulable once it carries the label enrollment puts on it, so the +// ordering is what stops a pod starting against an entry that is not there. A +// pod that starts against a missing entry reaches the node configuration script +// with no sizing and fails there, which is a long way from the cause. For the +// same reason, a cluster missing its required sizing is refused with an error // naming the fields rather than written out as blanks. +// +// The node set is a parameter rather than a list of its own, because the +// ordering only holds while both halves are looking at the same one: a node that +// appeared between two readings had its worker enrolled by a pass that never +// wrote its entry. func (w *Workload) ReconcileConfig( - ctx context.Context, cluster *simplyblockv1alpha2.StorageCluster, + ctx context.Context, + cluster *simplyblockv1alpha2.StorageCluster, + nodes []simplyblockv1alpha2.StorageNode, ) error { if cluster.Spec.MaxSubsystemCount == nil || cluster.Spec.VCPUCount == nil { return fmt.Errorf( @@ -61,18 +70,9 @@ func (w *Workload) ReconcileConfig( "set spec.maxSubsystemCount and spec.vcpuCount", cluster.Name) } - var nodes simplyblockv1alpha2.StorageNodeList - if err := w.List(ctx, &nodes, client.InNamespace(cluster.Namespace)); err != nil { - return fmt.Errorf("list cluster %s's nodes: %w", cluster.Name, err) - } - data := map[string]string{} - for i := range nodes.Items { - node := &nodes.Items[i] - if node.Spec.ClusterRef != cluster.Name { - continue - } - data[node.Spec.WorkerNode] = renderNodeConfig(cluster, node) + for i := range nodes { + data[nodes[i].Spec.WorkerNode] = renderNodeConfig(cluster, &nodes[i]) } return w.applyConfigMap(ctx, cluster, data) diff --git a/operator/internal/controllers/node/workload_controller.go b/operator/internal/controllers/node/workload_controller.go index 1a41fe363..5161424e9 100644 --- a/operator/internal/controllers/node/workload_controller.go +++ b/operator/internal/controllers/node/workload_controller.go @@ -121,11 +121,19 @@ func (r *StorageNodeWorkloadReconciler) Reconcile( return ctrl.Result{}, nil } - // The ConfigMap is written before the DaemonSet on every pass. A pod that - // starts against a missing or empty entry reaches the node configuration - // script with --max-subsys-count=0 and fails there, which is a long way from - // the cause (§5.3). - if err := r.Workload.ReconcileConfig(ctx, &cluster); err != nil { + // One reading of the node set feeds both halves that depend on it. The + // ConfigMap is written before the workers are enrolled, and a worker is only + // schedulable once enrollment has labeled it, so a pod cannot start against + // an entry that is not there -- while the two halves are reading the same + // set. Taken separately, a node that appeared between the readings had its + // worker enrolled by a pass that never wrote its entry, and the pod then + // reached the node configuration script with no sizing and failed there, + // which is a long way from the cause (§5.3). + nodes, err := r.clusterNodes(ctx, &cluster) + if err != nil { + return ctrl.Result{}, err + } + if err := r.Workload.ReconcileConfig(ctx, &cluster, nodes); err != nil { return ctrl.Result{}, err } @@ -137,7 +145,11 @@ func (r *StorageNodeWorkloadReconciler) Reconcile( {"the serving certificates", r.reconcileCertificates}, {"the headless service", r.reconcileService}, {"the endpoint slice", r.reconcileEndpointSlice}, - {"the worker enrollment", r.enrollWorkers}, + {"the worker enrollment", func( + ctx context.Context, cluster *simplyblockv1alpha2.StorageCluster, + ) error { + return r.enrollWorkers(ctx, cluster, nodes) + }}, {"the daemon set", r.reconcileDaemonSet}, } { if err := step.run(ctx, &cluster); err != nil { @@ -147,6 +159,34 @@ func (r *StorageNodeWorkloadReconciler) Reconcile( return ctrl.Result{}, nil } +// clusterNodes is the one reading of this cluster's storage nodes that the pass +// uses wherever the answer has to be the same answer. +// +// The filtering is here rather than in each caller so that "this cluster's +// nodes" means one thing: a node on its way out is not a reason to keep its +// worker enrolled, and it is not a reason to write its entry either. +func (r *StorageNodeWorkloadReconciler) clusterNodes( + ctx context.Context, cluster *simplyblockv1alpha2.StorageCluster, +) ([]simplyblockv1alpha2.StorageNode, error) { + var nodes simplyblockv1alpha2.StorageNodeList + if err := r.List(ctx, &nodes, client.InNamespace(cluster.Namespace)); err != nil { + return nil, fmt.Errorf("list the storage nodes: %w", err) + } + + kept := make([]simplyblockv1alpha2.StorageNode, 0, len(nodes.Items)) + for i := range nodes.Items { + node := &nodes.Items[i] + if node.Spec.ClusterRef != cluster.Name || node.Spec.WorkerNode == "" { + continue + } + if !node.DeletionTimestamp.IsZero() { + continue + } + kept = append(kept, *node) + } + return kept, nil +} + // enrollWorkers puts every worker this cluster has a node on into its storage // plane, which is what gives the DaemonSet somewhere to schedule. // @@ -165,24 +205,13 @@ func (r *StorageNodeWorkloadReconciler) Reconcile( // set from the nodes that want it, so calling it once per worker is both // sufficient and what keeps the set consistent. func (r *StorageNodeWorkloadReconciler) enrollWorkers( - ctx context.Context, cluster *simplyblockv1alpha2.StorageCluster, + ctx context.Context, + cluster *simplyblockv1alpha2.StorageCluster, + nodes []simplyblockv1alpha2.StorageNode, ) error { - var nodes simplyblockv1alpha2.StorageNodeList - if err := r.List(ctx, &nodes, client.InNamespace(cluster.Namespace)); err != nil { - return fmt.Errorf("list the storage nodes: %w", err) - } - seen := map[string]struct{}{} - for i := range nodes.Items { - node := &nodes.Items[i] - if node.Spec.ClusterRef != cluster.Name || node.Spec.WorkerNode == "" { - continue - } - // A node on its way out is not a reason to keep its worker enrolled, and - // releasing it is ReleaseWorker's job on the node's own path. - if !node.DeletionTimestamp.IsZero() { - continue - } + for i := range nodes { + node := &nodes[i] if _, already := seen[node.Spec.WorkerNode]; already { continue } From 0e614de301ec63315fea1e587737d0f5a56fdacd Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Mon, 21 Sep 2026 13:09:32 +0100 Subject: [PATCH 117/206] fixed manifest error --- operator/dist/install.yaml | 32 +++++++++++++++++++ .../internal/controllers/driver/workloads.go | 2 +- 2 files changed, 33 insertions(+), 1 deletion(-) diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index 2d521f443..d5623f0ca 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -8822,6 +8822,36 @@ rules: - patch - update - watch +- apiGroups: + - coordination.k8s.io + resources: + - leases + verbs: + - create + - delete + - get + - list + - update + - watch +- apiGroups: + - csiaddons.openshift.io + resources: + - csiaddonsnodes + verbs: + - create + - delete + - get + - list + - update + - watch +- apiGroups: + - csiaddons.openshift.io + resources: + - csiaddonsnodes/status + verbs: + - get + - patch + - update - apiGroups: - discovery.k8s.io resources: @@ -8868,6 +8898,8 @@ rules: resources: - clusterrolebindings - clusterroles + - rolebindings + - roles verbs: - bind - create diff --git a/operator/internal/controllers/driver/workloads.go b/operator/internal/controllers/driver/workloads.go index 9a531d284..023ede6bb 100644 --- a/operator/internal/controllers/driver/workloads.go +++ b/operator/internal/controllers/driver/workloads.go @@ -13,8 +13,8 @@ package driver import ( - "strconv" "slices" + "strconv" appsv1 "k8s.io/api/apps/v1" corev1 "k8s.io/api/core/v1" From d9ca691e3d554b51368cef2f3d79249567e8488f Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Mon, 21 Sep 2026 14:17:40 +0200 Subject: [PATCH 118/206] fix(operator): the journal-manager count is the control plane's to decide Three was sent whenever a node stated no ha_jm_count. That is the control plane's answer for a cluster that can lose one chunk with failure domains off, and it was sent for every cluster. A 2+2 deployment then refused every node add -- ha_jm_count=3 is too low for max_fault_tolerance=2, minimum required is 4 -- and the refusal happens inside the add rather than at it, so the task retried with the same parameters indefinitely. Since a node-add slot is now held for the whole add, the other five workers queued behind a worker that could never finish, and a cluster that had been laid out correctly looked wedged. The rule is get_required_ha_jm_count in storage_node_ops.py and it has two triggers and a cap: four where max_fault_tolerance is two or more, four where failure domains are enabled at any level, three otherwise. A first attempt here reimplemented it as parity + 2, read off the error message rather than the source. That is right for the case that failed and wrong for two others: it gives two for a 1+0 cluster, and it misses the failure-domain trigger entirely, which would have failed a 2+1 deployment with failure domains on in exactly the same way and for a reason the fix had just walked past. So nothing is reimplemented. resolve_ha_jm_count computes the requirement when the request carries no ha_jm_count, and the operator only ever failed because it always sent one. It now sends what the deployment states, and nothing when the deployment states nothing, which is what lets the rule stay in the one place that owns it and grow a third trigger without this side hearing about it. TestAnUnstatedJournalIsTheDefaultRatherThanNothing asserted the old contract in its name. It is now about the journal share, which is still defaulted here, and says why the count beside it no longer is. Test plan: U-405 through U-407. Co-Authored-By: Claude Opus 5 (1M context) --- operator/docs/tests/test-plan-storagenode.md | 3 + .../controllers/node/journalcount_test.go | 66 +++++++++++++++++++ .../controllers/node/provisioning_test.go | 21 ++++-- .../node/storagenode_controller.go | 19 +++++- 4 files changed, 101 insertions(+), 8 deletions(-) create mode 100644 operator/internal/controllers/node/journalcount_test.go diff --git a/operator/docs/tests/test-plan-storagenode.md b/operator/docs/tests/test-plan-storagenode.md index 5aa02cbd1..36345ab1f 100644 --- a/operator/docs/tests/test-plan-storagenode.md +++ b/operator/docs/tests/test-plan-storagenode.md @@ -107,6 +107,9 @@ Files: `operator/internal/controllers/node/provisioning_test.go`, | U-402 | A slot is held while the control plane still reports the node in_creation | Regression | `TestASlotIsHeldUntilTheAddIsFinished` | | U-403 | The slot goes back once the node leaves in_creation | Positive | `TestTheSlotGoesBackWhenTheNodeLeavesCreation` | | U-404 | Every enrolled worker has a per-node entry, across a node set that grows mid-pass | Regression | `TestEveryEnrolledWorkerHasAnEntry` | +| U-405 | A node stating no journal count leaves ha_jm_count to the control plane | Regression | `TestAnUnstatedJournalCountIsLeftToTheControlPlane` | +| U-406 | An unstated count is absent from the request rather than sent as zero | Boundary | `TestAnUnstatedJournalCountIsNotOnTheWire` | +| U-407 | A stated journal count is sent as it stands | Positive | `TestAStatedJournalCountIsSent` | ### Entity: The Provisioning Claim (design §4.2) diff --git a/operator/internal/controllers/node/journalcount_test.go b/operator/internal/controllers/node/journalcount_test.go new file mode 100644 index 000000000..a3b955374 --- /dev/null +++ b/operator/internal/controllers/node/journalcount_test.go @@ -0,0 +1,66 @@ +// How many journal managers a node is added with. +// +// The requirement is the control plane's and it has more than one trigger: four +// where the cluster can lose two chunks, four where failure domains are enabled +// at any parity level, three otherwise. It also computes that for itself when the +// request carries no ha_jm_count, which is what makes not sending one the whole +// of the operator's part. + +package node + +import ( + "encoding/json" + "strings" + "testing" + + "github.com/simplyblock/atlas/ptr" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/utils" +) + +// TestAnUnstatedJournalCountIsLeftToTheControlPlane is the defect. +// +// Regression: 2026-09-21-ha-jm-count-3-on-a-two-parity-cluster — three was sent +// whenever a node stated no count. That is the control plane's answer for a +// cluster that can lose one chunk with failure domains off, and it was sent for +// every cluster: a 2+2 deployment failed every node add, saying that an +// ha_jm_count of 3 was too low for a max_fault_tolerance of 2 and that the +// minimum required was 4. It retried on a loop, +// and held the cluster's only node-add slot while it did, so the other five +// workers queued behind a worker that could never finish. +func TestAnUnstatedJournalCountIsLeftToTheControlPlane(t *testing.T) { + if got := journalCount(nil); got != 0 { + t.Errorf("a node stating no journal count is added with ha_jm_count=%d, "+ + "which is this side answering a question the control plane answers", got) + } + + empty := &simplyblockv1alpha2.JournalManagerSpec{} + if got := journalCount(empty); got != 0 { + t.Errorf("an empty journal block is added with ha_jm_count=%d", got) + } +} + +// Zero has to leave the request rather than travel as zero, because zero is not +// a count the control plane would accept either. +func TestAnUnstatedJournalCountIsNotOnTheWire(t *testing.T) { + body, err := json.Marshal(utils.StorageNodeSetAddParams{ + HaJMCount: journalCount(nil), + }) + if err != nil { + t.Fatalf("marshal the add: %v", err) + } + if strings.Contains(string(body), "ha_jm_count") { + t.Errorf("the add carries ha_jm_count with nothing to say: %s", body) + } +} + +// A stated count is the deployment's and is sent as it stands. The control plane +// refuses one below its floor, which is the right place for that to be decided. +func TestAStatedJournalCountIsSent(t *testing.T) { + spec := &simplyblockv1alpha2.JournalManagerSpec{Count: ptr.To(int32(6))} + + if got := journalCount(spec); got != 6 { + t.Errorf("a stated count of 6 was sent as %d", got) + } +} diff --git a/operator/internal/controllers/node/provisioning_test.go b/operator/internal/controllers/node/provisioning_test.go index 8fb4ff24d..f36429127 100644 --- a/operator/internal/controllers/node/provisioning_test.go +++ b/operator/internal/controllers/node/provisioning_test.go @@ -170,16 +170,25 @@ func TestTheAddCarriesWhatTheNodeSaysAboutItself(t *testing.T) { } } -// A node that declares no journal settings is added with the defaults the -// control plane was always given, rather than with zeros. -func TestAnUnstatedJournalIsTheDefaultRatherThanNothing(t *testing.T) { +// A node that declares no journal share is added with the default the control +// plane was always given. The count beside it is not defaulted here any more. +// +// It used to be, as the three this case asserted, and three is the control +// plane's answer for one parity chunk with failure domains off rather than an +// answer for every cluster. Sending it made a 2+2 deployment refuse every node +// add. The count is now left out of the request, which is what makes the control +// plane compute the one its own rule requires. +func TestAnUnstatedJournalShareIsTheDefaultAndTheCountIsNot(t *testing.T) { r, _ := aSteadyNode(t, aControlPlane()) params := r.addParams(anUnprovisionedNode(stepPosting), anOpsCluster()) - if params.JMPercent != 3 || params.HaJMCount != 3 { - t.Errorf("journal = %d%% over %d, want the defaults 3 and 3", - params.JMPercent, params.HaJMCount) + if params.JMPercent != 3 { + t.Errorf("the journal share is %d%%, want the default 3", params.JMPercent) + } + if params.HaJMCount != 0 { + t.Errorf("the add states ha_jm_count=%d for a node that states none, "+ + "which answers a question the control plane answers", params.HaJMCount) } } diff --git a/operator/internal/controllers/node/storagenode_controller.go b/operator/internal/controllers/node/storagenode_controller.go index 4d240174d..679f90d04 100644 --- a/operator/internal/controllers/node/storagenode_controller.go +++ b/operator/internal/controllers/node/storagenode_controller.go @@ -1473,11 +1473,26 @@ func journalPercent(spec *simplyblockv1alpha2.JournalManagerSpec) int { return ptr.IntFrom(spec.PercentPerDevice, 3) } +// journalCount is how many journal managers a node is added with, and zero when +// the deployment has not said. +// +// Zero is omitted from the request (ha_jm_count is omitempty), and an absent +// ha_jm_count is what the control plane computes for itself: four where the +// cluster can lose two chunks, four where failure domains are enabled at any +// level, and three otherwise. The rule has two triggers and a cap, it is the +// control plane's, and it is the kind of rule that grows a third trigger. +// +// Three used to be sent here whenever a node stated no count. It was right for +// every deployment until one was laid out 2+2, and then wrong for every node of +// it: the add was refused inside the task rather than at it, retried with the +// same parameters forever, and held the cluster's only node-add slot while it +// did. Copying the formula over would have fixed that one deployment and left +// the failure-domain trigger to be discovered the same way. func journalCount(spec *simplyblockv1alpha2.JournalManagerSpec) int { if spec == nil { - return 3 + return 0 } - return ptr.IntFrom(spec.Count, 3) + return ptr.IntFrom(spec.Count, 0) } // recordStep persists the step the machine is about to be in, with the instant it From ae6809b744bd886dca20ef6a05377edc47d04bd3 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Mon, 21 Sep 2026 13:23:52 +0100 Subject: [PATCH 119/206] updated doc design-csi-addons-replication.md --- .../designs/design-csi-addons-replication.md | 93 ++++--------------- 1 file changed, 18 insertions(+), 75 deletions(-) diff --git a/operator/docs/designs/design-csi-addons-replication.md b/operator/docs/designs/design-csi-addons-replication.md index 74741973b..0f052ff65 100644 --- a/operator/docs/designs/design-csi-addons-replication.md +++ b/operator/docs/designs/design-csi-addons-replication.md @@ -1,6 +1,6 @@ # Design Document: csi-addons Volume Replication -**Status:** Phase 3 Partially Implemented (§11) +**Status:** Phase 3 Implemented **Author:** Israel Geoffrey (geoffrey1330) **Date:** 2026-09-16 (last updated 2026-09-21) **Test Plan:** [`tests/test-plan-csi-addons-replication.md`](../tests/test-plan-csi-addons-replication.md) @@ -13,10 +13,9 @@ |-------------|-------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|--------------| | **Phase 1** | Implemented | The csi-addons machinery and the steady-state contract: CRDs, controller-manager, sidecar, the Replication and csi-addons Identity gRPC services with `EnableVolumeReplication`, `DisableVolumeReplication`, and `GetVolumeReplicationInfo`, backed by a typed backend status endpoint | §4, §5.1, §6 | | **Phase 2** | Implemented | The lifecycle verbs: `PromoteVolume` (planned and forced), `DemoteVolume`, and `ResyncVolume`. Validation end to end against a Ramen `VolumeReplicationGroup` in async mode is still outstanding (§12, E-06/E-07) | §5.2, §9 | -| **Phase 3** | Partially implemented | §11 (the Prometheus metrics) is implemented. §7.2 (peerClasses convention and preflight) is not started: it needs a design pass on cross-cluster Kubernetes API access first, since no mechanism for the operator to reach a peer Kubernetes cluster exists today | §7, §11 | -| **Phase 4** | Planned | Test failover: the latest-replicated-snapshot read, the `drtest-*` conventions, and the two drill modes (bubble and test cluster) composed from clone, replication, and the real failover | §14 | +| **Phase 3** | Implemented | §11 (the Prometheus metrics). §7.2's peerClasses preflight is out of this design's scope entirely (it's Ramen's own `DRPolicy` mechanism) and is deferred to a future Ramen-integration design | §7.1, §11 | -Phase 1 is independently useful: a `VolumeReplication` object per PVC whose status truthfully reports the relationship, which no surface provides today. Phase 2 makes the object drivable, which is what Ramen actually needs. Phase 3 makes the whole thing operable at fleet scale. Phase 4 turns the same primitives into a rehearsal: a failover that can be drilled, in a bubble or against a test cluster, without touching production replication. +Phase 1 is independently useful: a `VolumeReplication` object per PVC whose status truthfully reports the relationship, which no surface provides today. Phase 2 makes the object drivable, which is what Ramen actually needs. Phase 3 makes the whole thing operable at fleet scale. The phase numbers above are this document's own, not the DR storage foundation gap analysis's (§1): its Phase 0 (shipping the csi-addons contract itself) is this design's Phase 1, and its Phase 1 (promote, demote, and resync end to end through Ramen) is this design's Phase 2. @@ -31,7 +30,6 @@ The phase numbers above are this document's own, not the DR storage foundation g | P0-3 | A standalone demote verb: `POST .../volumes/{id}/replication/demote` that converges the peer while still serving (repeated snapshot-and-ship until the remaining delta is small), then quiesces, ships the final delta, confirms it landed on the peer, and fences the data path | Control plane (`sbcli`) | Phase 2 | Shipped: `POST .../volumes/{v}/replication/demote` → `lvol_controller.demote_lvol` | | P0-4 | An `rpo_target_seconds` field on `ReplicationPolicy`, so RPO compliance is computable against a declared target rather than the derived lag budget | Control plane (`sbcli`) | Phase 3 | Shipped: `ReplicationPolicy.rpo_target_seconds` (`simplyblock_core/models/replication.py:89`), wired through the API (`PolicyParams.rpo_target_seconds`) and CLI (`--rpo-target-sec`) | | P0-5 | csi-addons upstream: the `VolumeReplication` and `VolumeReplicationClass` CRDs (`replication.storage.openshift.io/v1alpha1`), the kubernetes-csi-addons controller-manager image, and the csi-addons sidecar image | Ecosystem | Phase 1 | Vendored in the chart at v0.15.0 behind `csiaddons.create` (all twelve upstream CRDs, since the stock manager starts a controller per kind); sidecar wiring is Phase 1 | -| P0-6 | A latest-replicated-snapshot read: per volume, and per consistency group as one complete generation, the newest fully replicated snapshot on the secondary addressed as a cloneable object | Control plane (`sbcli`) | Phase 4 | Shipped: `GET .../replication/relationships/{lvol}/latest-snapshot` and `GET .../replication/policies/{policy}/latest-generation` | Everything else the adapter needs already exists: the attach and detach calls, failover, the failback and commit pair, the relationship read, and the backlog arithmetic inside `get_replication_info`. The adapter is thin precisely because the engine is complete. What is missing is the shape Ramen can drive. @@ -52,8 +50,7 @@ Everything else the adapter needs already exists: the attach and detach calls, f 11. [Observability](#11-observability) 12. [Testing Strategy](#12-testing-strategy) 13. [Migration Strategy](#13-migration-strategy) -14. [Test Failover](#14-test-failover) -15. [Open Questions](#15-open-questions) +14. [Open Questions](#14-open-questions) --- @@ -100,7 +97,7 @@ A reader who stops here has the model: the engine is unchanged, the csi-addons s - `status.lastSyncTime` is truthful for the volume's whole replicated life, sourced from a typed backend status read rather than the cutover-time relationship record. - Every verb is idempotent, because Ramen re-drives every reconcile. - The adapter reuses one shared control-plane client (`atlas-lib/controlplane`), ending the pattern where each consumer hand-rolls the same replication HTTP calls. -- A `StorageClass` and `VolumeSnapshotClass` naming convention across paired clusters that Ramen's peerClasses can express, with a preflight that verifies it. +- A `StorageClass` and `VolumeSnapshotClass` naming convention across paired clusters that Ramen's peerClasses can express (§7.2 documents what Ramen's own contract requires; verifying it is a future Ramen-integration design's concern, not this one's). - The RPO and backlog figures Ramen cannot carry (`bytesBehind`, throughput, RPO compliance) are exported as Prometheus metrics from the control plane. ### Non-Goals @@ -137,8 +134,7 @@ A reader who stops here has the model: the engine is unchanged, the csi-addons s │ │ csi-addons Identity and Replication (this design) │ │ │ └──────────────────────────────┬───────────────────────────────────┘ │ │ │ -│ operator: peerClasses preflight, events on the ReplicationPair (§7.2); │ -│ PVCAnnotationWatcher skips csi-addons-managed volumes (§8); │ +│ operator: PVCAnnotationWatcher skips csi-addons-managed volumes (§8); │ │ ReplicationPair/Policy author the backend target and policy │ └─────────────────────────────────┬────────────────────────────────────────────┘ │ HTTP, resolved per volume handle @@ -153,7 +149,6 @@ A reader who stops here has the model: the engine is unchanged, the csi-addons s │ quiesce, flush, fence) │ │ POST .../volumes/{v}/replication/failback (resync) │ │ GET .../replication/relationships/{lvol} (cutover records) │ -│ GET .../relationships/{lvol}/latest-snapshot (P0-6, test failover) │ │ exports simplyblock_replication_* metrics: lag, backlog, RPO (§11) │ └──────────────────────────────────────────────────────────────────────────────┘ ``` @@ -162,9 +157,7 @@ A reader who stops here has the model: the engine is unchanged, the csi-addons s **The adapter holds no state.** csi-addons RPCs are stateless and idempotent by contract. Every answer the driver gives is derived on the spot from the backend status read and the relationship record. There is no driver-side cache, no persisted step, and no state machine. `PromoteVolumeResponse` and `DemoteVolumeResponse` carry no fields at all in `csi-addons/spec` v0.2.0 -- there is no response field to report partial progress in. A verb whose backend work outlives the call (a demote converging a busy peer) instead returns a retryable `ABORTED` error, and the vendored controller-manager's own reconcile requeue is the retry loop; only once the backend reports the target state reached does the call return success, which the controller then reflects as the CR's `Completed` condition. -**The operator's part is small and off the data path.** The kinds that author the backend state (`ReplicationPair` for the target, `ReplicationPolicy` for cadence and retention) keep working unchanged, and the classes name what they author. On top of them the operator runs the peerClasses preflight (§7.2), surfacing convention drift as events on the pair, and teaches the `PVCAnnotationWatcher` the one-owner rule (§8) so the legacy annotation path and a `VolumeReplication` never fight over one volume. Everything imperative it used to own (`ReplicationOps`, the commit cutover, the cutover-proceed handshake) is off this contract and confined to the legacy path. - -**The drill rides the same surface.** Test failover (§14) adds no machinery to this picture: the bubble mode clones the latest replicated snapshot (the P0-6 read plus the ordinary CSI clone path), and the test-cluster mode composes clone, a `drtest-` policy toward the test cluster, and the real failover. The invariant audit that proves a drill disturbed nothing reads the same P0-1 status endpoint the conditions come from. +**The operator's part is small and off the data path.** The kinds that author the backend state (`ReplicationPair` for the target, `ReplicationPolicy` for cadence and retention) keep working unchanged, and the classes name what they author. On top of them the operator teaches the `PVCAnnotationWatcher` the one-owner rule (§8) so the legacy annotation path and a `VolumeReplication` never fight over one volume. Everything imperative it used to own (`ReplicationOps`, the commit cutover, the cutover-proceed handshake) is off this contract and confined to the legacy path. --- @@ -273,26 +266,26 @@ spec: schedulingInterval: 5m ``` -`replicationPolicy` names the backend `ReplicationPolicy` (resolved per cluster by name), which owns cadence, retention, mode, and the replication target. `schedulingInterval` restates the policy's interval for Ramen's `DRPolicy` matching, and the preflight (§7.2) checks the two agree. No secrets parameter is needed: the driver's credentials come from `secret.json`, as for every other RPC. The chart ships no default class. Classes are the user's to author, matching the `VolumeGroupSnapshotClass` decision. +`replicationPolicy` names the backend `ReplicationPolicy` (resolved per cluster by name), which owns cadence, retention, mode, and the replication target. `schedulingInterval` restates the policy's interval for Ramen's `DRPolicy` matching. No secrets parameter is needed: the driver's credentials come from `secret.json`, as for every other RPC. The chart ships no default class. Classes are the user's to author, matching the `VolumeGroupSnapshotClass` decision. One class per (policy, cadence) is the authoring model: a `VolumeReplicationClass` names exactly one policy, and a volume needing a different cadence follows a different policy under a different class. Per-volume interval overrides are not provided. -### 7.2 peerClasses convention and preflight +### 7.2 peerClasses: Ramen's own mechanism, out of this design's scope + +`peerClasses` is Ramen's `DRPolicy` computation, not this design's: Ramen's hub-side controller pairs each managed cluster's `StorageClass`/`VolumeReplicationClass` objects itself, through the OCM hub-spoke visibility it already has, and a `DRPolicy` that cannot find a valid pairing already reports that failure on its own. Building any verification of that pairing into this operator -- a preflight, an admission check, or otherwise -- belongs to the future design that actually wires Ramen/OCM into this operator, not here: this design's own scope stops at §7.1, authoring one `VolumeReplicationClass` per policy, which is sufficient for the Phase 1/2 adapter to work whether or not Ramen, OCM, or peerClasses are ever in the picture. -Two of the requirements below are Ramen's contract, and the rest is this design's convention; the split matters because only the contract can fail a DRPolicy. +What Ramen's contract requires of a pairing, for whoever writes that future design: **Contract (Ramen requires this):** - **The `StorageClass` name exists on both clusters.** A failover restores the protected PVC with its original `spec.storageClassName`, a by-name reference, so an equivalent class of that exact name must exist on the peer. The peerClasses computation also joins classes across the two clusters by `StorageClass` name. - **The Ramen identity labels.** Each cluster's `StorageClass` carries `ramendr.openshift.io/storageid` (differing per cluster, since the backends differ), and the two `VolumeReplicationClass` objects representing one relationship carry an equal `ramendr.openshift.io/replicationid`. The `VolumeReplicationClass` is selected per cluster by `spec.replicationClassSelector` labels plus `provisioner` and a `schedulingInterval` equal to the `DRPolicy`'s, never by name. -**Convention (this design chooses it for operability):** +**Convention (this design's own authoring choice, independent of Ramen):** - **Same `StorageClass` parameters on both clusters** apart from `cluster_id` (necessarily) and pool when pools differ. The replication target's pool mapping already handles the pool difference at shipping time. - **Same `VolumeSnapshotClass` and `VolumeReplicationClass` names on both clusters**, each side's `replicationPolicy` naming that cluster's policy toward its peer. Ramen does not require the names to match, but one name per relationship is what keeps a fleet legible. -The preflight is a check, not a controller: a validation that runs on demand (and on `ReplicationPair` reconciliation) confirming that for each replication-enabled `StorageClass` the peer cluster has a same-named class, and that the named backend policies exist and point at each other's clusters. Its findings surface as events on the `ReplicationPair`, which is the object that already models the cluster pairing. - --- ## 8. Coexistence with the Legacy Replication Kinds @@ -320,8 +313,6 @@ Every endpoint is scoped as today: volume-scoped under `/api/v2/clusters/{c}/sto | `POST` | `.../volumes/{v}/replication/failback` | Existing. Resync (direction reversal, delta-seeded). | | `POST` | `.../volumes/{v}/replication/commit` | Existing, unchanged, and NOT part of this contract: it stays behind the legacy `ReplicationOps` migration path only (§13). | | `GET` | `.../replication/relationships/{lvol}` | Existing, unchanged. Cutover records only; the node redirect depends on its survive-deletion semantics. | -| `GET` | `.../replication/relationships/{lvol}/latest-snapshot` | **Shipped (P0-6).** The newest fully replicated snapshot for the volume, as a cloneable snapshot handle. Exposes what the failover path already computes internally. | -| `GET` | `.../replication/policies/{policy}/latest-generation` | **Shipped (P0-6), group form.** One complete, fully replicated consistency-group generation, every member as a cloneable snapshot handle. Reuses the group fail-over's mixed-generation refusal instead of its side effect. | | `POST` | `.../volumes/{v}/replication/cutover-proceed` | Existing, unchanged, legacy path only: the adapter never reaches it, because the commit cutover is off this contract. | The unused backend verbs the operator never calls (`start`, `stop`, `trigger`, `tasks`) are unaffected, and `start` and `stop` remain the policy-less legacy path. @@ -349,12 +340,7 @@ The unused backend verbs the operator never calls (`start`, `stop`, `trigger`, ` ### Kubernetes Events -The kubernetes-csi-addons controller-manager owns events on `VolumeReplication` (promote, demote, and resync outcomes), and this design adds none there. The operator emits preflight findings on the `ReplicationPair`: - -| Event | Type | Emitted when | -|-----------------------|---------|-----------------------------------------------------------------------------------------------------------------------------| -| `PeerClassesVerified` | Normal | The preflight confirmed same-named classes and mutually pointing policies on both clusters | -| `PeerClassesMismatch` | Warning | A replication-enabled class has no same-named peer, or the named policies do not pair; the message names the class and side | +The kubernetes-csi-addons controller-manager owns events on `VolumeReplication` (promote, demote, and resync outcomes), and this design adds none there. This design defines no events of its own on `ReplicationPair` either: the peerClasses preflight that would have emitted them is Ramen's own concern, out of scope here (§7.2). ### Prometheus Metrics (Implemented) @@ -380,7 +366,7 @@ Values are computed by `lvol_controller.get_replication_info_bulk`, a bulk-frien Full scenario matrix and coverage status: [`tests/test-plan-csi-addons-replication.md`](../tests/test-plan-csi-addons-replication.md) - **Unit (driver):** each verb against a mock control plane: the idempotency table (repeat enable, repeat disable, repeat promote), the refusal paths (different-policy enable, lagging planned promote, disable during cutover), the condition derivation from every status-read state, and handle parsing failures. -- **Unit (operator):** the peerClasses preflight against fake clients for both clusters, and the `PVCAnnotationWatcher` skip when a `VolumeReplication` exists. +- **Unit (operator):** the `PVCAnnotationWatcher` skip when a `VolumeReplication` exists. - **Integration:** the csi-addons sidecar and controller-manager against the driver with a mock backend under envtest or kind: a `VolumeReplication` flipped `primary` to `secondary` and back walks the verbs in order and lands the conditions. - **E2E (two live clusters):** the Ramen-shaped lifecycle without Ramen: enable on the source, write data, and verify `lastSyncTime` advances; forced promote on the DR side, verifying the clone serves with the source fenced; and resync back with a planned swap (demote then promote), verifying zero loss with a hashed writer. Then the same driven by an actual Ramen VRG in async mode, which is Phase 2's acceptance gate. @@ -399,53 +385,10 @@ Three replication control surfaces exist today: the operator's kinds, the stale --- -## 14. Test Failover - -A DR drill proves that failover works without disturbing production replication. Ramen cannot drive one: its two actions, `Failover` and `Relocate`, move the workload for real. The drill is therefore driven by a simplyblock-native kind (owned by the SiteMap and testing layer, and specified there, not here), and this section defines the storage contract that kind consumes. There are two modes over one substrate: a bubble on the secondary cluster, and a separate test cluster reached like the real failover. - -### 14.1 The substrate: replicated snapshots are cloneable test points - -The replicated snapshots on the secondary are first-class snapshot records on the secondary's own control plane, chained and complete, and for a consistency group they carry the `group_id` and `group_seq` provenance of their generation. The failover path already resolves the newest fully replicated snapshot per volume, and the group-wide resolution picks one complete generation across members. P0-6 exposes that resolution as a read (§9), so a drill can address its test point without reimplementing the selection logic. - -Every object a drill creates, on either side, carries a `drtest-` name prefix and a test-id label. Leftovers are then enumerable, and a teardown can prove completeness instead of assuming it. - -### 14.2 Bubble mode: same cluster, different namespace - -The drill namespace lives on the secondary Kubernetes cluster, whose driver already talks to the backend holding the replicated snapshots, so no data moves at all: - -1. Resolve the test point through P0-6: per volume the newest replicated snapshot, or for a group one complete `group_seq`, so the bubble starts from a single crash-consistent cut. -2. Surface each snapshot as a pre-provisioned `VolumeSnapshotContent` (the handle is the secondary-side snapshot), bind a `VolumeSnapshot` in the drill namespace, and clone it into a PVC through the ordinary `dataSource` path. -3. Deploy the application against the clones. The clones are thin, independent volumes, and writes to them never touch the replication stream. -4. Tear down by deleting the namespace, then enumerate by the test-id label to prove nothing leaked. - -### 14.3 Test-cluster mode: the real failover, aimed at expendable volumes - -The second mode reaches a separate storage cluster, and it is deliberately composed from primitives this design already relies on rather than a new shipping capability: - -1. Clone the resolved test point into `drtest-` volumes on the secondary. For a group, clone one generation into a new `drtest-` consistency group, so the group failover path is exercised too. -2. Attach the clones to a `drtest-` replication policy whose `ReplicationTarget` is the test cluster. The ordinary engine ships them (a full copy, since the test backend shares no ancestry). -3. On the test cluster, run the real failover against the shipped volumes to materialize writable clones, and deploy the application there. - -The property this buys is fidelity: the drill exercises the actual failover machinery, target resolution, clone-from-replicated, identity preservation, and the driver redirect, against a third cluster, while the production relationship is never touched, because the `drtest-` policy is a separate policy with its own target. - -### 14.4 The non-disruption proof - -A drill that silently perturbed replication would be worse than no drill. Before the first clone and after the teardown, the driving kind captures and compares: every production `VolumeReplication`'s conditions, the `lastSyncTime` cadence, the lag and backlog from the typed status read (P0-1), and the count of production `VolumeReplication` objects. Any drift fails the drill as an invariant violation. The status endpoint built for Ramen's conditions is the same instrument this audit reads, which is why the drill contract belongs in this design. - -### 14.5 Costs and bounds - -- **A live clone pins its base snapshot.** Retention defers pruning a snapshot with a dependent clone, which is what keeps the drill safe, and also why a drill must carry a maximum lifetime: a long-lived bubble holds the secondary's replicated chain back. -- **Test-cluster mode consumes real resources:** cross-cluster bandwidth for the full copy, and capacity on both the secondary (the `drtest-` clones) and the test cluster. The `drtest-` clones on the secondary exist only as replication sources and are never served, so they are created as internal volumes, the same treatment the shipping engine's own landing volumes already get: invisible to normal listings and exempt from the per-node subsystem cap. Storage capacity is unaffected either way; a drill's cleanup and audit go through the `drtest-` test-id enumeration (§14.1), not the ordinary volume listing. -- **The group rules apply unchanged:** a `drtest-` consistency group observes the member cap and the placement pin like any other, so a drill of a large group is a capacity event on the secondary. - ---- - -## 15. Open Questions +## 14. Open Questions | # | Question | Owner | |---|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------| | 1 | **Demote semantics for the application.** The P0-3 demote fences the volume (ANA inaccessible) after the final flush, and with convergence folded into the verb it is now the only place a planned swap can stall. This is also the one verb the planned promote's lossless guarantee entirely depends on (§5.2): a planned promote is refused unless a completed demote already fenced the source and confirmed the final delta landed, so an unresolved failure mode here is an unresolved gap in the whole "zero loss" claim. Ramen relocation unmounts the workload first, so the fence is ordinarily unopposed, but that is Ramen's choreography, not a guarantee the driver can rely on: a stuck termination, a stale mount that never released, or a demote invoked outside Ramen's normal flow can all leave writes still arriving when quiesce fires. Confirm the verb's behavior when writes are still in flight at quiesce (block versus fail), whether the converge phase has its own budget separate from the quiesced flush, and whether a timeout in either phase must abort back to serving primary or leave the volume fenced with no automatic recovery. | Backend team | -| 2 | **Where the preflight lives.** §7.2 attaches peerClasses validation to the `ReplicationPair` reconciler. If the redesign retires the pair kind, the preflight needs a new home (the `SimplyblockDriver`, or a standalone check job). | Operator team | -| 3 | **Per-volume policy granularity.** A `VolumeReplicationClass` names one policy, and today one policy implies one target and cadence for all its volumes. Confirm one class per (policy, cadence) is an acceptable authoring model for Ramen's `replicationClassSelector`, or whether per-volume interval overrides are needed. | Operator / Backend team | -| 4 | **Visibility of `drtest-` clones.** The test-cluster mode's clones on the secondary are replication sources only, never served. Decide whether the backend creates them as internal volumes (hidden from listings, exempt from the per-node subsystem cap, like the shipping path's landing volumes) or as ordinary volumes under a naming convention. | Backend team | -| 5 | ~~**Avoiding the clone on day-one protection.**~~ **Resolved:** `POST .../replication/failover?planned=true`'s no-demote branch now checks `lvol_controller.replication_source_online` (the source's own storage-node status) before falling through to `FAILED_PRECONDITION` -- an online source is a no-op (§5.2), so a healthy volume's first-ever `PromoteVolume` no longer materializes a clone. The remaining residual: a source that dies within the last health-check interval still briefly reads online, so one reconcile can treat a genuine disaster as a no-op before the node's status catches up and the controller retries -- bounded by the health-check detection window, not open-ended. | Backend team | +| 2 | **Per-volume policy granularity.** A `VolumeReplicationClass` names one policy, and today one policy implies one target and cadence for all its volumes. Confirm one class per (policy, cadence) is an acceptable authoring model for Ramen's `replicationClassSelector`, or whether per-volume interval overrides are needed. | Operator / Backend team | +| 3 | ~~**Avoiding the clone on day-one protection.**~~ **Resolved:** `POST .../replication/failover?planned=true`'s no-demote branch now checks `lvol_controller.replication_source_online` (the source's own storage-node status) before falling through to `FAILED_PRECONDITION` -- an online source is a no-op (§5.2), so a healthy volume's first-ever `PromoteVolume` no longer materializes a clone. The remaining residual: a source that dies within the last health-check interval still briefly reads online, so one reconcile can treat a genuine disaster as a no-op before the node's status catches up and the controller retries -- bounded by the health-check detection window, not open-ended. | Backend team | From 0b7f62930e7334eba134df4441d350e2827c6ae0 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Mon, 21 Sep 2026 14:45:25 +0200 Subject: [PATCH 120/206] fix(operator): the spdk-proxy names are published again The control plane drives a storage node's SPDK process over RPC at .simplyblock-spdk-proxy..svc.cluster.local:. That name is a headless Service plus one EndpointSlice per port. The Service was applied and the slices were not: BuildSpdkProxyEndpointSlice outlived the reconcile that called it when StorageNodeSet retired (d3380539), kept its own unit tests, kept passing them, and published nothing. Nothing used the names, so nothing noticed. simplyblock_core/models/storage_node.py dials mgmt_ip while tls_connect is "disabled" and the per-pod name only when it is not, and TLS was off by default until 5b1cf4cc turned it on. Every node add after that stalled in the control plane's resolver retry loop against a pod that was running, on a Service that existed. The reconcile is restored with the distinction it was written around. A pod that is gone and a pod that is merely not ready this instant are different: readiness flips on one missed probe tick, and deleting a slice for that takes a node's address away mid-operation, which is what an earlier incident was. So the delete pass tests against every port that has a pod object at all, and a not-ready pod keeps its last known address until the pass that finds it ready refreshes it. The step list moves out of Reconcile into workloadSteps. A case that calls a reconcile directly cannot catch a reconcile nobody calls, which is the whole of what went wrong here, and the list was a literal inside a function with no way to read it from a test. Removing the step now fails a test rather than a cluster. Test plan: U-408 through U-412. Co-Authored-By: Claude Opus 5 (1M context) --- operator/docs/tests/test-plan-storagenode.md | 5 + .../internal/controllers/node/spdkproxydns.go | 219 +++++++++++++++++ .../controllers/node/spdkproxydns_test.go | 227 ++++++++++++++++++ .../controllers/node/workload_controller.go | 35 ++- 4 files changed, 477 insertions(+), 9 deletions(-) create mode 100644 operator/internal/controllers/node/spdkproxydns.go create mode 100644 operator/internal/controllers/node/spdkproxydns_test.go diff --git a/operator/docs/tests/test-plan-storagenode.md b/operator/docs/tests/test-plan-storagenode.md index 36345ab1f..33c26f5b3 100644 --- a/operator/docs/tests/test-plan-storagenode.md +++ b/operator/docs/tests/test-plan-storagenode.md @@ -110,6 +110,11 @@ Files: `operator/internal/controllers/node/provisioning_test.go`, | U-405 | A node stating no journal count leaves ha_jm_count to the control plane | Regression | `TestAnUnstatedJournalCountIsLeftToTheControlPlane` | | U-406 | An unstated count is absent from the request rather than sent as zero | Boundary | `TestAnUnstatedJournalCountIsNotOnTheWire` | | U-407 | A stated journal count is sent as it stands | Positive | `TestAStatedJournalCountIsSent` | +| U-408 | The pass publishes the spdk-proxy endpoints, so the builder has a caller | Regression | `TestThePassPublishesTheProxyEndpoints` | +| U-409 | A worker per-pod name resolves to its address on the right port | Regression | `TestTheSPDKProxyNamesArePublished` | +| U-410 | One slice per RPC port | Positive | `TestEachRPCPortGetsItsOwnSlice` | +| U-411 | A pod with no address is left unpublished rather than published with none | Negative | `TestAPodWithNoAddressIsNotPublished` | +| U-412 | Two workers sharing a first DNS label are refused, not merged | Negative | `TestACollidingWorkerNameIsRefused` | ### Entity: The Provisioning Claim (design §4.2) diff --git a/operator/internal/controllers/node/spdkproxydns.go b/operator/internal/controllers/node/spdkproxydns.go new file mode 100644 index 000000000..a80c4616b --- /dev/null +++ b/operator/internal/controllers/node/spdkproxydns.go @@ -0,0 +1,219 @@ +// The per-pod DNS names the control plane drives an SPDK process by. +// +// A storage node's SPDK process answers RPC on its own port, and the control +// plane reaches it at +// .simplyblock-spdk-proxy..svc.cluster.local:. +// That name is a headless Service plus one EndpointSlice per port, each endpoint +// carrying the worker's name as its hostname -- the Service on its own resolves +// nothing. +// +// It is only that name under TLS. simplyblock_core/models/storage_node.py dials +// mgmt_ip directly while tls_connect is "disabled", so a plaintext deployment +// never asks DNS for any of this and a deployment that serves TLS cannot add a +// single node without it. The reconcile that published these slices belonged to +// StorageNodeSet and went out with that kind; nothing used the names until TLS +// became the default, and then every node add stalled in the control plane's +// resolver retry loop. +// +// What the two passes below are careful about is the difference between a pod +// that is gone and a pod that is merely not ready this instant. Deleting a +// slice for the second costs the node its address mid-operation, which is the +// shape of an earlier outage: readiness flips on a single missed probe tick, +// well short of anything that restarts a container. + +package node + +import ( + "context" + "fmt" + "strconv" + "strings" + + corev1 "k8s.io/api/core/v1" + discoveryv1 "k8s.io/api/discovery/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil" + logf "sigs.k8s.io/controller-runtime/pkg/log" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/utils" +) + +// spdkProxyServiceName is the headless Service these slices attach to, and it is +// a literal because the control plane's own address template is. +const spdkProxyServiceName = "simplyblock-spdk-proxy" + +// reconcileSpdkProxyEndpoints publishes one EndpointSlice per RPC port. +func (r *StorageNodeWorkloadReconciler) reconcileSpdkProxyEndpoints( + ctx context.Context, cluster *simplyblockv1alpha2.StorageCluster, +) error { + log := logf.FromContext(ctx) + + var pods corev1.PodList + if err := r.List(ctx, &pods, + client.InNamespace(cluster.Namespace), + client.MatchingLabels{"role": utils.LabelSpdkProxyRole}, + ); err != nil { + return fmt.Errorf("list the spdk-proxy pods: %w", err) + } + + // Two sets, and the difference between them is the whole of the delete pass + // below. byPort holds the pods that can serve an address now; anyPod holds + // every port that has a pod object at all, ready or not. The RPC port is + // readable from the pod spec the moment it is scheduled, long before it is + // ready, so the second is safe to compute from the full list. + byPort := map[int32][]utils.SpdkProxyEndpoint{} + anyPod := map[int32]bool{} + for i := range pods.Items { + pod := &pods.Items[i] + rpcPort, ok := spdkProxyRPCPort(pod) + if !ok { + log.Info("an spdk-proxy pod names no RPC port and is not published", + "pod", pod.Name) + continue + } + anyPod[rpcPort] = true + if !spdkProxyPodServesAnAddress(pod) { + continue + } + byPort[rpcPort] = append(byPort[rpcPort], utils.SpdkProxyEndpoint{ + NodeName: pod.Spec.NodeName, + PodIP: pod.Status.PodIP, + RpcPort: rpcPort, + }) + } + + for rpcPort, endpoints := range byPort { + if err := r.applyProxySlice(ctx, cluster, rpcPort, endpoints); err != nil { + return err + } + } + return r.pruneProxySlices(ctx, cluster, anyPod) +} + +// applyProxySlice writes one port's slice. +func (r *StorageNodeWorkloadReconciler) applyProxySlice( + ctx context.Context, + cluster *simplyblockv1alpha2.StorageCluster, + rpcPort int32, + endpoints []utils.SpdkProxyEndpoint, +) error { + desired, err := utils.BuildSpdkProxyEndpointSlice(cluster, rpcPort, endpoints) + if err != nil { + return err + } + if err := controllerutil.SetControllerReference(cluster, desired, r.Scheme); err != nil { + return fmt.Errorf("own the spdk-proxy EndpointSlice for port %d: %w", rpcPort, err) + } + + var existing discoveryv1.EndpointSlice + err = r.Get(ctx, client.ObjectKeyFromObject(desired), &existing) + if apierrors.IsNotFound(err) { + return r.Create(ctx, desired) + } + if err != nil { + return err + } + desired.ResourceVersion = existing.ResourceVersion + return r.Update(ctx, desired) +} + +// pruneProxySlices removes the slices of ports that have no pod at all. +// +// The test is against every port with a pod object rather than every port with a +// ready one. A pod that is not ready for one probe tick keeps its last known +// address until the pass that finds it ready refreshes it; taking the name away +// instead is what left the control plane resolving nothing mid-operation. +func (r *StorageNodeWorkloadReconciler) pruneProxySlices( + ctx context.Context, + cluster *simplyblockv1alpha2.StorageCluster, + anyPod map[int32]bool, +) error { + var slices discoveryv1.EndpointSliceList + if err := r.List(ctx, &slices, + client.InNamespace(cluster.Namespace), + client.MatchingLabels{"kubernetes.io/service-name": spdkProxyServiceName}, + ); err != nil { + return fmt.Errorf("list the published spdk-proxy EndpointSlices: %w", err) + } + + for i := range slices.Items { + slice := &slices.Items[i] + if !metav1.IsControlledBy(slice, cluster) { + continue + } + serving := false + for _, port := range slice.Ports { + if port.Port != nil && anyPod[*port.Port] { + serving = true + break + } + } + if serving { + continue + } + if err := r.Delete(ctx, slice); err != nil && !apierrors.IsNotFound(err) { + return fmt.Errorf("delete the stale spdk-proxy EndpointSlice %s: %w", slice.Name, err) + } + } + return nil +} + +// spdkProxyPodServesAnAddress reports whether a pod can be published. +// +// It is not the same question as whether the pod is healthy. What an endpoint +// needs is a worker and an address, and every container reporting ready is what +// says the proxy is listening on it. +func spdkProxyPodServesAnAddress(pod *corev1.Pod) bool { + if pod.Status.Phase != corev1.PodRunning { + return false + } + if pod.Spec.NodeName == "" || pod.Status.PodIP == "" { + return false + } + for _, status := range pod.Status.ContainerStatuses { + if !status.Ready { + return false + } + } + return len(pod.Status.ContainerStatuses) > 0 +} + +// spdkProxyRPCPort is the port a pod's proxy answers on. +// +// RPC_PORT on the proxy container is the statement of it. The pod name is a +// fallback because it carries the same number, and a pod the control plane +// created with an older template is still a pod this has to publish. +func spdkProxyRPCPort(pod *corev1.Pod) (int32, bool) { + for _, container := range pod.Spec.Containers { + if container.Name != "spdk-proxy-container" { + continue + } + for _, env := range container.Env { + if env.Name != "RPC_PORT" || env.Value == "" { + continue + } + port, err := strconv.ParseInt(env.Value, 10, 32) + if err != nil { + return 0, false + } + return int32(port), true + } + } + + rest, ok := strings.CutPrefix(pod.Name, "snode-spdk-pod-") + if !ok { + return 0, false + } + dash := strings.Index(rest, "-") + if dash <= 0 { + return 0, false + } + port, err := strconv.ParseInt(rest[:dash], 10, 32) + if err != nil { + return 0, false + } + return int32(port), true +} diff --git a/operator/internal/controllers/node/spdkproxydns_test.go b/operator/internal/controllers/node/spdkproxydns_test.go new file mode 100644 index 000000000..c7711eb5c --- /dev/null +++ b/operator/internal/controllers/node/spdkproxydns_test.go @@ -0,0 +1,227 @@ +// The per-pod DNS names the control plane reaches an SPDK process by. +// +// A node's SPDK process is created by the control plane and then driven over RPC +// at worker-N.simplyblock-spdk-proxy..svc.cluster.local:. That name +// comes from a headless Service plus an EndpointSlice carrying one endpoint per +// worker, with the worker's name as the endpoint hostname. The Service alone +// resolves nothing. + +package node + +import ( + "context" + "fmt" + "slices" + "strings" + "testing" + + corev1 "k8s.io/api/core/v1" + discoveryv1 "k8s.io/api/discovery/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/testsupport" + "github.com/simplyblock/simplyblock-operator/internal/utils" +) + +// anSPDKPod is one SPDK process as the control plane creates it: host-networked, +// so its pod address is the worker's, and carrying its RPC port in the app label. +func anSPDKPod(name, worker, address string, rpcPort int) *corev1.Pod { + return &corev1.Pod{ + ObjectMeta: metav1.ObjectMeta{ + Name: name, + Namespace: "simplyblock", + Labels: map[string]string{ + "app": fmt.Sprintf("spdk-app-%d", rpcPort), + "role": "simplyblock-storage-node", + }, + }, + Spec: corev1.PodSpec{ + HostNetwork: true, + NodeName: worker, + Containers: []corev1.Container{{ + Name: "spdk-proxy-container", + Env: []corev1.EnvVar{{Name: "RPC_PORT", Value: fmt.Sprintf("%d", rpcPort)}}, + }}, + }, + Status: corev1.PodStatus{ + Phase: corev1.PodRunning, + PodIP: address, + ContainerStatuses: []corev1.ContainerStatus{{Name: "spdk-proxy-container", Ready: true}}, + }, + } +} + +// aProxyReconciler is the workload reconciler over the pods given. +func aProxyReconciler( + t *testing.T, pods ...*corev1.Pod, +) (*StorageNodeWorkloadReconciler, *simplyblockv1alpha2.StorageCluster) { + t.Helper() + scheme := testsupport.NewScheme(t, corev1.AddToScheme, discoveryv1.AddToScheme) + + cluster := &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{Name: "c", Namespace: "simplyblock"}, + } + objects := make([]client.Object, 0, 1+len(pods)) + objects = append(objects, cluster) + for _, p := range pods { + objects = append(objects, p) + } + + return &StorageNodeWorkloadReconciler{ + Client: fake.NewClientBuilder().WithScheme(scheme).WithObjects(objects...).Build(), + Scheme: scheme, + Namespace: "simplyblock", + }, cluster +} + +// TestTheSPDKProxyNamesArePublished is the defect. +// +// Regression: 2026-09-21-the-spdk-proxy-service-had-no-endpoints — the headless +// Service was applied and its EndpointSlice never was. BuildSpdkProxyEndpointSlice +// existed and nothing called it, so every per-pod name resolved to nothing and +// every node add stalled in the control plane's RPC retry loop: "Failed to +// resolve the worker's per-pod name. The +// pod was up, the Service was there, and the add never finished. +func TestTheSPDKProxyNamesArePublished(t *testing.T) { + r, cluster := aProxyReconciler(t, + anSPDKPod("snode-spdk-pod-4420-abc", "worker-0.ocp.simplyblock.ai", "10.0.0.10", 4420), + ) + + if err := r.reconcileSpdkProxyEndpoints(context.Background(), cluster); err != nil { + t.Fatalf("publish the proxy endpoints: %v", err) + } + + var slice discoveryv1.EndpointSlice + key := client.ObjectKey{Namespace: "simplyblock", Name: "spdk-proxy-endpoints-4420"} + if err := r.Get(context.Background(), key, &slice); err != nil { + t.Fatalf("the proxy EndpointSlice was not published: %v", err) + } + + if len(slice.Endpoints) != 1 { + t.Fatalf("the slice carries %d endpoints", len(slice.Endpoints)) + } + endpoint := slice.Endpoints[0] + if endpoint.Hostname == nil || *endpoint.Hostname != "worker-0" { + t.Errorf("the endpoint hostname is %v, and the control plane resolves worker-0", + endpoint.Hostname) + } + if len(endpoint.Addresses) != 1 || endpoint.Addresses[0] != "10.0.0.10" { + t.Errorf("the endpoint addresses are %v", endpoint.Addresses) + } + if slice.Labels["kubernetes.io/service-name"] != "simplyblock-spdk-proxy" { + t.Errorf("the slice is not attached to the headless Service: %v", slice.Labels) + } +} + +// One slice per RPC port, because each socket of each worker answers on its own +// and a slice carries one port. +func TestEachRPCPortGetsItsOwnSlice(t *testing.T) { + r, cluster := aProxyReconciler(t, + anSPDKPod("snode-spdk-pod-4420-abc", "worker-0.ocp.simplyblock.ai", "10.0.0.10", 4420), + anSPDKPod("snode-spdk-pod-4422-abc", "worker-1.ocp.simplyblock.ai", "10.0.0.11", 4422), + ) + + if err := r.reconcileSpdkProxyEndpoints(context.Background(), cluster); err != nil { + t.Fatalf("publish the proxy endpoints: %v", err) + } + + for port, worker := range map[int]string{4420: "worker-0", 4422: "worker-1"} { + var slice discoveryv1.EndpointSlice + key := client.ObjectKey{ + Namespace: "simplyblock", + Name: fmt.Sprintf("spdk-proxy-endpoints-%d", port), + } + if err := r.Get(context.Background(), key, &slice); err != nil { + t.Fatalf("port %d was not published: %v", port, err) + } + if len(slice.Endpoints) != 1 || *slice.Endpoints[0].Hostname != worker { + t.Errorf("port %d published %d endpoints", port, len(slice.Endpoints)) + } + if slice.Ports[0].Port == nil || int(*slice.Ports[0].Port) != port { + t.Errorf("the slice for port %d carries port %v", port, slice.Ports[0].Port) + } + } +} + +// A pod with no address yet is left out rather than published with none: an +// endpoint with no address resolves to nothing and is worse than an absent name, +// because a caller gets a resolution rather than a retry. +func TestAPodWithNoAddressIsNotPublished(t *testing.T) { + r, cluster := aProxyReconciler(t, + anSPDKPod("snode-spdk-pod-4420-abc", "worker-0.ocp.simplyblock.ai", "", 4420), + ) + + if err := r.reconcileSpdkProxyEndpoints(context.Background(), cluster); err != nil { + t.Fatalf("publish the proxy endpoints: %v", err) + } + + var slice discoveryv1.EndpointSlice + key := client.ObjectKey{Namespace: "simplyblock", Name: "spdk-proxy-endpoints-4420"} + err := r.Get(context.Background(), key, &slice) + if err == nil && len(slice.Endpoints) != 0 { + t.Errorf("a pod with no address was published as %v", slice.Endpoints[0].Addresses) + } +} + +// Two workers whose names share a first DNS segment cannot both be published, +// because the endpoint hostname is that segment. It is an error rather than a +// silent overwrite of one worker's address with the other's. +func TestACollidingWorkerNameIsRefused(t *testing.T) { + r, cluster := aProxyReconciler(t, + anSPDKPod("snode-spdk-pod-4420-abc", "worker-0.site-a.example", "10.0.0.10", 4420), + anSPDKPod("snode-spdk-pod-4420-def", "worker-0.site-b.example", "10.0.0.20", 4420), + ) + + err := r.reconcileSpdkProxyEndpoints(context.Background(), cluster) + if err == nil { + t.Fatal("two workers sharing a DNS label were published without complaint") + } + if !strings.Contains(err.Error(), "collision") { + t.Errorf("the error does not name the collision: %v", err) + } +} + +// The builder is the one this uses, so the name it produces is the name the +// control plane was configured to resolve. +func TestTheSliceNameMatchesTheBuilder(t *testing.T) { + slice, err := utils.BuildSpdkProxyEndpointSlice( + &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{Namespace: "simplyblock"}, + }, + 4420, + []utils.SpdkProxyEndpoint{{NodeName: "worker-0.x", PodIP: "10.0.0.10", RpcPort: 4420}}, + ) + if err != nil { + t.Fatalf("build: %v", err) + } + if slice.Name != "spdk-proxy-endpoints-4420" { + t.Errorf("the builder names the slice %q", slice.Name) + } +} + +// TestThePassPublishesTheProxyEndpoints is the gap the builder fell through. +// +// Regression: 2026-09-21-the-spdk-proxy-slice-had-no-caller — +// BuildSpdkProxyEndpointSlice survived the retirement of StorageNodeSet and the +// reconcile that called it did not. The builder kept its unit tests and went on +// passing them, so nothing anywhere failed while every per-pod name resolved to +// nothing. +// +// A case that calls the reconcile directly cannot catch that: it is the wiring +// that was missing, not the behavior. This asserts the step is in the pass. +func TestThePassPublishesTheProxyEndpoints(t *testing.T) { + r, _ := aProxyReconciler(t) + + steps := r.workloadSteps(nil) + names := make([]string, 0, len(steps)) + for _, step := range steps { + names = append(names, step.what) + } + if !slices.Contains(names, "the spdk-proxy endpoints") { + t.Error("no step of the workload pass publishes the spdk-proxy endpoints, " + + "so the control plane resolves nothing under TLS") + } +} diff --git a/operator/internal/controllers/node/workload_controller.go b/operator/internal/controllers/node/workload_controller.go index 5161424e9..42e652f6d 100644 --- a/operator/internal/controllers/node/workload_controller.go +++ b/operator/internal/controllers/node/workload_controller.go @@ -137,26 +137,43 @@ func (r *StorageNodeWorkloadReconciler) Reconcile( return ctrl.Result{}, err } - for _, step := range []struct { - what string - run func(context.Context, *simplyblockv1alpha2.StorageCluster) error - }{ + for _, step := range r.workloadSteps(nodes) { + if err := step.run(ctx, &cluster); err != nil { + return ctrl.Result{}, fmt.Errorf("reconcile %s: %w", step.what, err) + } + } + return ctrl.Result{}, nil +} + +// workloadStep is one thing the pass applies, with the name its failure is +// reported under. +type workloadStep struct { + what string + run func(context.Context, *simplyblockv1alpha2.StorageCluster) error +} + +// workloadSteps is everything the pass applies, in order. +// +// It is a method rather than a literal inside Reconcile so that what the pass +// does is readable from a test. The spdk-proxy endpoints are in this list +// because they were once in another one: the builder outlived the reconcile that +// called it, kept passing its own unit tests, and published nothing. +func (r *StorageNodeWorkloadReconciler) workloadSteps( + nodes []simplyblockv1alpha2.StorageNode, +) []workloadStep { + return []workloadStep{ {"the service account and its role", r.reconcileRBAC}, {"the serving certificates", r.reconcileCertificates}, {"the headless service", r.reconcileService}, {"the endpoint slice", r.reconcileEndpointSlice}, + {"the spdk-proxy endpoints", r.reconcileSpdkProxyEndpoints}, {"the worker enrollment", func( ctx context.Context, cluster *simplyblockv1alpha2.StorageCluster, ) error { return r.enrollWorkers(ctx, cluster, nodes) }}, {"the daemon set", r.reconcileDaemonSet}, - } { - if err := step.run(ctx, &cluster); err != nil { - return ctrl.Result{}, fmt.Errorf("reconcile %s: %w", step.what, err) - } } - return ctrl.Result{}, nil } // clusterNodes is the one reading of this cluster's storage nodes that the pass From 4afa3ca587d8b4f4eb639f7baf89dd0645f6cd3b Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Mon, 21 Sep 2026 14:27:18 +0100 Subject: [PATCH 121/206] added design document design-ramen-integration.md --- csi-driver/internal/csi/common/server.go | 8 +- .../internal/csi/controller/replication.go | 3 +- .../controller/replication_lifecycle_test.go | 6 +- .../csi/controller/replication_test.go | 10 +- .../docs/designs/design-ramen-integration.md | 247 ++++++++++++++++++ .../tests/test-plan-csi-addons-replication.md | 11 +- .../docs/tests/test-plan-ramen-integration.md | 134 ++++++++++ 7 files changed, 404 insertions(+), 15 deletions(-) create mode 100644 operator/docs/designs/design-ramen-integration.md create mode 100644 operator/docs/tests/test-plan-ramen-integration.md diff --git a/csi-driver/internal/csi/common/server.go b/csi-driver/internal/csi/common/server.go index 411cc969c..a83acf61b 100644 --- a/csi-driver/internal/csi/common/server.go +++ b/csi-driver/internal/csi/common/server.go @@ -17,7 +17,13 @@ type NonBlockingGRPCServer interface { // Replication) with the same *grpc.Server the CSI services register on, // so csicommon never has to import csi-addons: the caller builds the // closures and this package only invokes them. - Start(endpoint string, ids csi.IdentityServer, cs csi.ControllerServer, ns csi.NodeServer, register ...func(*grpc.Server)) + Start( + endpoint string, + ids csi.IdentityServer, + cs csi.ControllerServer, + ns csi.NodeServer, + register ...func(*grpc.Server), + ) Wait() Stop() ForceStop() diff --git a/csi-driver/internal/csi/controller/replication.go b/csi-driver/internal/csi/controller/replication.go index f472ba2d6..d8f74930a 100644 --- a/csi-driver/internal/csi/controller/replication.go +++ b/csi-driver/internal/csi/controller/replication.go @@ -70,7 +70,8 @@ func (cs *Server) EnableVolumeReplication( ) (*replication.EnableVolumeReplicationResponse, error) { policyID := req.GetParameters()[replicationPolicyParam] if policyID == "" { - return nil, status.Errorf(codes.InvalidArgument, "VolumeReplicationClass parameter %q is required", replicationPolicyParam) + return nil, status.Errorf(codes.InvalidArgument, + "VolumeReplicationClass parameter %q is required", replicationPolicyParam) } h, err := csicommon.ParseVolumeHandle(volumeIDFrom(req)) if err != nil { diff --git a/csi-driver/internal/csi/controller/replication_lifecycle_test.go b/csi-driver/internal/csi/controller/replication_lifecycle_test.go index 6818eaff7..f52b01504 100644 --- a/csi-driver/internal/csi/controller/replication_lifecycle_test.go +++ b/csi-driver/internal/csi/controller/replication_lifecycle_test.go @@ -92,7 +92,7 @@ func TestPromoteVolumeUsesReplicationSourceWhenVolumeIdIsEmpty(t *testing.T) { cs := newReplicationTestServer(t, mock) _, err := cs.PromoteVolume(context.Background(), &replication.PromoteVolumeRequest{ - ReplicationSource: replicationSourceFor(testReplVolID), Force: true, + ReplicationSource: replicationSourceFor(), Force: true, }) if err != nil { t.Fatal(err) @@ -151,7 +151,7 @@ func TestDemoteVolumeUsesReplicationSourceWhenVolumeIdIsEmpty(t *testing.T) { cs := newReplicationTestServer(t, mock) _, err := cs.DemoteVolume(context.Background(), &replication.DemoteVolumeRequest{ - ReplicationSource: replicationSourceFor(testReplVolID), + ReplicationSource: replicationSourceFor(), }) if err != nil { t.Fatal(err) @@ -177,7 +177,7 @@ func TestResyncVolumeUsesReplicationSourceWhenVolumeIdIsEmpty(t *testing.T) { cs := newReplicationTestServer(t, mock) _, err := cs.ResyncVolume(context.Background(), &replication.ResyncVolumeRequest{ - ReplicationSource: replicationSourceFor(testReplVolID), + ReplicationSource: replicationSourceFor(), }) if err != nil { t.Fatal(err) diff --git a/csi-driver/internal/csi/controller/replication_test.go b/csi-driver/internal/csi/controller/replication_test.go index 578356891..4bf110d32 100644 --- a/csi-driver/internal/csi/controller/replication_test.go +++ b/csi-driver/internal/csi/controller/replication_test.go @@ -25,10 +25,10 @@ func newReplicationTestServer(t *testing.T, mock *mockSBCLI) *Server { // kubernetes-csi-addons v0.15.0 sidecar sends on every Replication RPC // instead of the legacy flat VolumeId field (internal/sidecar/service's // ReplicationServer proxy never sets it). -func replicationSourceFor(volumeID string) *replication.ReplicationSource { +func replicationSourceFor() *replication.ReplicationSource { return &replication.ReplicationSource{ Type: &replication.ReplicationSource_Volume{ - Volume: &replication.ReplicationSource_VolumeSource{VolumeId: volumeID}, + Volume: &replication.ReplicationSource_VolumeSource{VolumeId: testReplVolID}, }, } } @@ -80,7 +80,7 @@ func TestEnableVolumeReplicationUsesReplicationSourceWhenVolumeIdIsEmpty(t *test cs := newReplicationTestServer(t, mock) _, err := cs.EnableVolumeReplication(context.Background(), &replication.EnableVolumeReplicationRequest{ - ReplicationSource: replicationSourceFor(testReplVolID), + ReplicationSource: replicationSourceFor(), Parameters: map[string]string{replicationPolicyParam: testReplPolicyID}, }) if err != nil { @@ -196,7 +196,7 @@ func TestDisableVolumeReplicationUsesReplicationSourceWhenVolumeIdIsEmpty(t *tes mock.volumes[testReplVolumeID].ReplicationPolicyID = testReplPolicyID _, err := cs.DisableVolumeReplication(context.Background(), &replication.DisableVolumeReplicationRequest{ - ReplicationSource: replicationSourceFor(testReplVolID), + ReplicationSource: replicationSourceFor(), }) if err != nil { t.Fatal(err) @@ -275,7 +275,7 @@ func TestGetVolumeReplicationInfoUsesReplicationSourceWhenVolumeIdIsEmpty(t *tes } resp, err := cs.GetVolumeReplicationInfo(context.Background(), &replication.GetVolumeReplicationInfoRequest{ - ReplicationSource: replicationSourceFor(testReplVolID), + ReplicationSource: replicationSourceFor(), }) if err != nil { t.Fatal(err) diff --git a/operator/docs/designs/design-ramen-integration.md b/operator/docs/designs/design-ramen-integration.md new file mode 100644 index 000000000..eb23a235d --- /dev/null +++ b/operator/docs/designs/design-ramen-integration.md @@ -0,0 +1,247 @@ +# Design Document: Ramen Integration + +**Status:** Draft (contract confirmed, peerClasses preflight specified, implementation and E2E validation pending) +**Author:** Israel Geoffrey (geoffrey1330) +**Date:** 2026-09-21 +**Test Plan:** [`tests/test-plan-ramen-integration.md`](../tests/test-plan-ramen-integration.md) + +--- + +## Phase 0: External Prerequisites + +| # | Prerequisite | Kind | Blocks | Status | +|------|--------------------------------------------------------------------------------------------------------------------|-----------|----------------------------------------------|-----------------------------------------------------------------------------------------------------------| +| P0-1 | A live OCM hub with at least two registered managed clusters | Ecosystem | The E2E validation (§6) | Unknown | +| P0-2 | Ramen installed on the hub and on each managed cluster, with a `DRPolicy` naming both clusters | Ecosystem | The E2E validation (§6) | Unknown | +| P0-3 | SiteMap, or a hand-authored `DRPlacementControl` standing in for it, driving the `DRPolicy` | Ecosystem | The E2E validation (§6) | Not shipped. SiteMap is an external document and system, and storage is explicitly outside its own scope. | +| P0-4 | `csi-addons/spec` at a version whose `GetVolumeReplicationInfoResponse` carries `lastSyncBytes`/`lastSyncDuration` | Ecosystem | Full Appendix A.3 `GetVolumeReplicationInfo` | Not shipped: pinned at v0.2.0 today, which has neither field. | + +Without P0-1 through P0-3 nothing in §6 can run, because the validation is E2E-only, live-cluster work that no mock or `envtest` substitutes for. §3's preflight and §4's `VolumeGroupReplication` reconciler are both unaffected: neither needs an OCM hub, a Ramen installation, or a `DRPolicy`, only this cluster's own objects. P0-4's absence is narrower: it leaves `GetVolumeReplicationInfo` reporting only `lastSyncTime`, never cycle size or duration, but Ramen's own `PeerReady` gate (§5.1) does not read either field, so P0-4 does not block the validation itself. + +--- + +## Table of Contents + +1. [Background](#1-background) +2. [Goals and Non-Goals](#2-goals-and-non-goals) +3. [peerClasses Preflight](#3-peerclasses-preflight) +4. [VolumeGroupReplication](#4-volumegroupreplication) +5. [The Contract, Confirmed](#5-the-contract-confirmed) +6. [E2E Validation Plan](#6-e2e-validation-plan) +7. [Testing Strategy](#7-testing-strategy) +8. [Open Questions](#8-open-questions) + +--- + +## Overview + +`design-csi-addons-replication.md` builds the storage-level adapter Ramen's per-volume DR contract requires, and validates it end to end "without Ramen" (its own §12): real backend, real csi-addons machinery, but a hand-driven `VolumeReplication` object rather than a real Ramen reconcile loop. That document deferred two pieces as "the next design": a `peerClasses` preflight (its own §7.2, once it became clear `peerClasses` itself is Ramen's mechanism, not this operator's) and `VolumeGroupReplication` (its own §2 Non-Goals, strictly per volume itself). `design-consistency-groups.md` deferred `VolumeGroupReplication` too, as future work independent of any replication policy. Neither document claims it. This document is where both land: §3 specifies the preflight, §4 specifies group replication on top of the consistency-group primitive, and §5 through §6 confirm the rest of the per-volume contract against what already shipped and specify the E2E validation that closes `design-csi-addons-replication.md` §12's outstanding acceptance gate (E-06, E-07). + +--- + +## 1. Background + +The gap analysis's own headline finding (§2) was that simplyblock's DR machinery was "a complete, parallel, out-of-band system that implements none of the interfaces Ramen drives." Appendix A answered that with a concise spec: the six csi-addons `Replication` gRPC verbs, the one info query, and the `VolumeReplication.status` conditions Ramen actually reads (`Completed`, `Degraded`, `Resyncing`), aggregated by Ramen's own VRG into `DataProtected` and by the DRPC into `PeerReady`, the boolean that gates `Relocate` and failback. Appendix B specified the RPO/RTO figures Ramen cannot carry natively, as metrics. + +`design-csi-addons-replication.md`'s three phases implement that spec. This document does not repeat what it built. §5 below cites, section by section, where each Appendix A and Appendix B item now lives in the shipped code. + +The other half of Appendix A's own premise is that Ramen's hub, not this operator, computes `peerClasses` and drives the VRG, through OCM's hub-spoke visibility into every managed cluster. That hub, and the OCM/SiteMap layer above it, is external to this repository (confirmed this session: no `ManagedCluster`, `DRPolicy`, or `VolumeReplicationGroup` reference exists anywhere in this codebase outside design-doc prose, and no mechanism for this operator to reach a peer cluster's Kubernetes API exists or is needed. `design-management-hub.md`, this repo's own hub design, is a separate fleet-config-distribution concern that cites Ramen's hub/spoke split only as precedent, not as something it builds). Ramen's hub already reports when it cannot find a valid `StorageClass`/`VolumeReplicationClass` pairing across two clusters, once it looks. What it has no way to see is a pairing that is wrong on one cluster alone, before any hub-side comparison happens: a `VolumeReplicationClass` missing the `ramendr.openshift.io/replicationid` label, or carrying no `schedulingInterval`, sits invisibly broken until a `DRPolicy` is authored against it and the hub's own reconcile surfaces a cryptic cross-cluster mismatch instead of the local, fixable cause. §3 closes that local gap. + +The gap analysis's own §6 named a second piece of genuinely new work, in its Phase 2: `VolumeGroupReplication`, "on top of the CG primitive," so that the VRG async group path can protect and fail over a multi-volume app at one point rather than only snapshot it. `design-consistency-groups.md` built the CG primitive that gap analysis cites (the `storage.simplyblock.io/consistency-group` label, `VolumeGroupSnapshot`, the `GroupController`) but explicitly left group replication for later, independent of any policy, exactly so this document could attach it without reshaping the group. §4 is that attachment. Together, §3 and §4 are the only new production code this document proposes: everything else Appendix A and Appendix B specify is confirmed, in §5, against code `design-csi-addons-replication.md` already shipped. + +--- + +## 2. Goals and Non-Goals + +### Goals + +- Specify and implement a same-cluster `peerClasses` preflight: catch a `StorageClass`/`VolumeReplicationClass` pairing that Ramen's contract requires but this cluster's objects do not satisfy, before a `DRPolicy` is ever authored against it (§3). +- Specify and implement `VolumeGroupReplication` on top of the consistency-group primitive: fan a group's `primary`/`secondary`/`resync` intent out to its members' existing per-volume adapter, and fan their status back in, satisfying Appendix A.4's group-readiness status query and the gap analysis's own Phase 2 ask (§4). +- Confirm, against the actual shipped code, that every condition, verb, and status query Appendix A specifies is satisfied, or state precisely which is not and why (§5). +- Confirm, against the actual shipped code, which Appendix B metrics are delivered, which are derivable from what already exists, and which remain future work (§5.4). +- Specify a live-cluster validation that closes `design-csi-addons-replication.md` §12's outstanding acceptance gate: a real Ramen VRG, on a real OCM-registered pair of clusters, driving the adapter through protect, planned relocate, and unplanned failover (§6). + +### Non-Goals + +- **Building any hub, OCM, or cross-cluster Kubernetes access mechanism.** That is Ramen's and OCM's job, external to this operator, confirmed in §1. Neither §3's preflight nor §4's group reconciler reaches past this cluster's own objects, and nothing in §6's validation plan asks this operator to reach a peer cluster's API server: every step drives objects on the cluster where the workload currently runs, exactly as `design-csi-addons-replication.md`'s own architecture already assumes. +- **Resolving a `peerClasses` mismatch across two clusters.** That comparison needs the hub's own visibility into both clusters and is Ramen's job once a `DRPolicy` exists. §3 catches what is locally wrong before that comparison ever runs, and it does not repeat the comparison itself. +- **Global VGR.** RamenDR's newer multi-VRG consensus feature for a replication group spanning several applications is out of scope. §4 covers the base case: one VRG, one storage vendor's PVCs, one `VolumeGroupReplication` (§4.1, §8 Open Question 5). +- **A new gRPC verb for group replication.** §4.2 is explicit: group promote, demote, and resync fan the same three verbs `design-csi-addons-replication.md` §5 already ships out to every member. Nothing changes on the driver. +- **SiteMap.** A separate external document and system. Where §6's topology needs a `DRPlacementControl` and SiteMap is not available to author one, a hand-authored stand-in is explicitly permitted (P0-3). +- **`bytesBehind`'s remaining Appendix B siblings** (throughput, RTO estimation, backup RPO/RTO). §5.4 accounts for each, and none blocks the validation this document specifies. +- **Widening `csi-addons/spec` past v0.2.0.** P0-4 is recorded as a prerequisite, not solved here: it is an upstream dependency version, not something this repository's own code can add a field to. + +--- + +## 3. peerClasses Preflight + +Ramen pairs a `StorageClass` and a `VolumeReplicationClass` across two managed clusters into a `peerClasses` entry on the hub's `DRPolicy`, using labels this operator's own convention already defines: `ramendr.openshift.io/storageid` on the `StorageClass`, `ramendr.openshift.io/replicationid` on the `VolumeReplicationClass`, both cited in `design-csi-addons-replication.md` §7.1. When the pairing is wrong on one cluster alone, the earliest and most legible place to catch it is that cluster, before a `DRPolicy` ever compares it against a peer. + +### 3.1 What it checks + +The `ReplicationPairReconciler` (`operator/internal/controller/replicationpair_controller.go`), on its existing 60-second reconcile cadence (`replPairSyncInterval`), adds a preflight step reading only this cluster's own objects: + +1. **List `StorageClass` objects** (typed `storagev1.StorageClass`, the same type `pvcreplication_controller.go` already reads), filtered to `Provisioner == "csi.simplyblock.io"`. +2. **Enrollment signal.** If none of this cluster's simplyblock `StorageClass` objects carries the `ramendr.openshift.io/storageid` label, there is nothing to preflight, and the check is a no-op: an unlabeled `StorageClass` is not a Ramen-managed one, and flagging it would be noise, not a finding. +3. **List `VolumeReplicationClass` objects.** This is an external CRD (`replication.storage.openshift.io/v1alpha1`, from `csi-addons`'s own `volume-replication-operator`) this repository does not vendor a Go type for, so the list reads `unstructured.UnstructuredList` against `schema.GroupVersionKind{Group: "replication.storage.openshift.io", Version: "v1alpha1", Kind: "VolumeReplicationClassList"}`, the same no-vendored-type pattern `internal/upgrade/discover/kinds.go` already uses for `cert-manager`'s `Certificate`. +4. **Per-class check**, for every listed `VolumeReplicationClass` whose `spec.provisioner` is `csi.simplyblock.io`: + - `metadata.labels["ramendr.openshift.io/replicationid"]` is present and non-empty. + - `spec.parameters.schedulingInterval` is present and non-empty. + + Neither check calls the simplyblock backend: whether the named `schedulingInterval` corresponds to a real `ReplicationPolicy` is `design-csi-addons-replication.md` §5's own province (`replicationpolicy_controller.go` already rejects an unresolvable policy at that layer), and re-checking it here would be the sbcli-backed preflight rejected earlier in this design's own history. This step verifies only that Ramen's own two labels and one parameter are present, which is everything a same-cluster read can verify. +5. **Verdict.** If step 2 found an enrolled `StorageClass` and step 4 found no `VolumeReplicationClass` with a matching provisioner, or found one with a missing label or parameter, the preflight fails. Otherwise, it passes. + +### 3.2 Where the result goes + +One Kubernetes event per reconcile, on the `ReplicationPair` object driving the check, not a new CR and not a status field: the preflight is a diagnostic, and the `ReplicationPair`'s own status already carries the fields Ramen and the backend care about. + +| Event reason | Type | When | +|-----------------------|---------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| `PeerClassesVerified` | Normal | An enrolled `StorageClass` exists and every matching `VolumeReplicationClass` carries both the label and the parameter. | +| `PeerClassesMismatch` | Warning | An enrolled `StorageClass` exists but no matching `VolumeReplicationClass` does, or one is missing the label or the parameter. The message names the missing field and the object. | + +`ReplicationPairReconciler` gains a `Recorder events.EventRecorder` field, following the exact pattern `backuppolicy_controller.go` already uses (`r.Recorder.Eventf(object, nil, eventType, reason, reason, format, args...)`), wired in `operator/cmd/main.go` as `mgr.GetEventRecorder("replicationpair-controller")`. + +### 3.3 RBAC + +Two markers, both read-only, added to the reconciler's existing block: + +```go +// +kubebuilder:rbac:groups=storage.k8s.io,resources=storageclasses,verbs=get;list;watch +// +kubebuilder:rbac:groups=replication.storage.openshift.io,resources=volumereplicationclasses,verbs=get;list;watch +``` + +The second is this repository's second grant on an external CRD group it does not own, alongside `simplyblockstoragenodeset_controller.go`'s existing `cert-manager.io/certificates` marker, the precedent `rbac-hardening` cites for exactly this shape of grant. + +### 3.4 What this is not + +This is not the cross-cluster preflight the first pass at this design attempted and the second pass's user correction rejected: it makes no call to the simplyblock backend, resolves no `ReplicationPolicy` by name, and reaches no peer cluster's API server. It is the narrowest thing that closes the gap Ramen's own hub cannot see: a `StorageClass`/`VolumeReplicationClass` pairing broken on one cluster, before any `DRPolicy` compares it to a peer. + +--- + +## 4. VolumeGroupReplication + +The gap analysis's own Phase 2 named this the piece consistency groups still owed: "Add `VolumeGroupReplication` (csi-addons) on top of the CG primitive so the VRG async group path can protect and fail over multi-volume apps at one point, not just snapshot them." `design-csi-addons-replication.md` §2 deferred it as "the next design," strictly per-volume itself. `design-consistency-groups.md` §2 deferred it too, as "future work," independent of any replication policy so that work could attach later. Neither claims it. This section is that attachment. + +### 4.1 What already exists, upstream + +`VolumeGroupReplication`, `VolumeGroupReplicationClass`, and `VolumeGroupReplicationContent` are shipped CRDs in the same `replication.storage.openshift.io` group as `VolumeReplication` and `VolumeReplicationClass`, from `kubernetes-csi-addons`. `VolumeGroupReplication.spec` carries the same three-state `replicationState` (`primary`/`secondary`/`resync`) the per-volume kind carries, a `source.selector` naming the member PVCs by label, and an `external` boolean: `false` routes reconciliation through the generic kubernetes-csi-addons controller-manager, and `true` hands it to "an external controller managed by the storage vendor." Ramen's own VRG (`RamenDR/ramen`'s `vrg_volgrouprep.go`) creates and drives `VolumeGroupReplication` objects directly, once a VRG's PVCs share a `VolumeGroupReplicationClass` carrying a `ramendr.openshift.io/groupreplicationid` label, the group-level sibling of the per-volume `replicationid` label §3.1 already checks. Ramen's own documentation names this case "offloaded" replication: a storage backend that replicates at the LUN or logical-volume-store level, outside Kubernetes, through a vendor controller, exactly the shape simplyblock's backend already has. + +**Global VGR, RamenDR's newer multi-VRG consensus feature for a replication group spanning several applications' VRGs, is out of scope here.** This section covers the base case Ramen has supported longer: one VRG, one storage vendor's PVCs, one `VolumeGroupReplication`. Whether the base case is sufficient for simplyblock's own use, or Global VGR's cross-VRG consensus is eventually needed too, is §8 Open Question 5. + +### 4.2 No new gRPC contract + +Promoting, demoting, or resyncing a group is fanning the same three verbs `design-csi-addons-replication.md` §5 already ships (`PromoteVolume`, `DemoteVolume`, `ResyncVolume`) out to every member, not a fourth verb on the driver. The `replicationState` values line up one for one with the per-volume kind's, by design: `kubernetes-csi-addons`'s own generic controller reconciles a *non*-external `VolumeGroupReplication` this same way, fanning it out to member `VolumeReplication` objects it creates itself. Setting `external: true` does not change what the fan-out does. It changes who performs it, so that the vendor controller can skip the generic manager's own per-member `VolumeReplication` bookkeeping and go straight to whatever shape fits the backend. simplyblock's shape is already built: reuse the per-volume adapter through the same `VolumeReplication` objects Ramen already drives for a single volume, one per group member. + +### 4.3 The reconciler + +A new `VolumeGroupReplicationReconciler`, alongside `ReplicationPairReconciler` and the rest, owns every `VolumeGroupReplication` whose `spec.external` is `true` and whose `spec.volumeGroupReplicationClassName` names a class with `provisioner: csi.simplyblock.io`: + +1. **Resolve membership.** Read `spec.source.selector` against this cluster's PVCs. Reuse the exact invariant `design-consistency-groups.md` §9.2 already established for `VolumeGroupSnapshot`: the selected set must equal a `storage.simplyblock.io/consistency-group` value's current membership exactly, not merely a subset or superset of it. A selector that does not resolve to one whole group is a configuration error, not a partial group to serve. +2. **Fan out.** For each member PVC, ensure a per-volume `VolumeReplication` object exists, owned by the `VolumeGroupReplication`, named deterministically from the group and the member, with `spec.replicationState` mirroring the group's. The already-shipped `kubernetes-csi-addons` controller-manager reconciles each of these exactly as it does any Ramen-created per-volume `VolumeReplication` (`design-csi-addons-replication.md` §5): this reconciler creates and updates the member objects, and never calls the driver's Replication gRPC itself. +3. **Fan in.** Aggregate every member's `VolumeReplication.status.conditions` into the group's own status: `Completed` is the conjunction across all members, `Degraded` and `Resyncing` are the disjunction (one degraded or resyncing member makes the group so). `status.lastGroupSyncTime`, the field Appendix A.4's group-readiness status query names, is the oldest of the members' `lastSyncTime`: a group's recovery point is only as fresh as its slowest member. + +### 4.4 Admission webhook, extended + +`design-consistency-groups.md` §9.4's validating webhook on `VolumeGroupSnapshot` create already enforces "the selector must equal the group's current membership" at `kubectl apply`, with the same fail-closed label check and fail-open backend check. A sibling webhook on `VolumeGroupReplication` create makes the identical two checks against the identical label, so a `VolumeGroupReplication` that could never resolve to one whole group is rejected before the reconciler in §4.3 ever sees it, rather than sitting unreconciled. + +### 4.5 RBAC and events + +`VolumeGroupReplicationReconciler` needs read access to `persistentvolumeclaims` (already granted elsewhere in this operator) and read-write access to the external `volumegroupreplications.replication.storage.openshift.io` and the per-member `volumereplications.replication.storage.openshift.io` it creates: + +```go +// +kubebuilder:rbac:groups=replication.storage.openshift.io,resources=volumegroupreplications,verbs=get;list;watch;update;patch +// +kubebuilder:rbac:groups=replication.storage.openshift.io,resources=volumegroupreplications/status,verbs=get;update;patch +// +kubebuilder:rbac:groups=replication.storage.openshift.io,resources=volumereplications,verbs=get;list;watch;create;update;patch;delete +``` + +Events follow §3.2's pattern, on the `VolumeGroupReplication` object: `GroupReplicationVerified`/`GroupMembershipMismatch` at the webhook's own admission-time checks, mirrored as reconcile-time events for the case the webhook admitted open, and a `GroupReplicationDegraded` (Warning) when §4.3's fan-in first observes a member `Degraded` after previously reporting none. + +--- + +## 5. The Contract, Confirmed + +### 5.1 Conditions (Appendix A.2) + +Appendix A.2 specified `Completed`, `Degraded`, `Resyncing`, with steady-state healthy async as all three settled (`Completed=True, Degraded=False, Resyncing=False`) and lag alone tripping none of them. `design-csi-addons-replication.md` §6.2 implements exactly this mapping: `Completed` from the last requested state change finishing, `Degraded` from the status read's `degraded`/`error` state or `lag_seconds` exceeding `lag_budget_seconds`, `Resyncing` from the status read's own `resyncing` flag. This exceeds Appendix A.2's own "interim" fallback (a staleness heuristic on `lastReplicatedAt`): `Degraded` is sourced from the backend's real failing-task signal, not a timestamp guess. + +Appendix A.3's `PeerReady` (`Completed=True && Degraded=False && Resyncing=False` on the peer) needs no code of this repository's own. It is Ramen's own DRPC-level aggregation of the three conditions above, computed client-side from what §6.2 already reports. + +### 5.2 gRPC verbs (Appendix A.3) + +All six verbs Appendix A.3 specifies are implemented and idempotent, per `design-csi-addons-replication.md` §5: + +| csi-addons verb | Appendix A.3 requirement | Status | +|----------------------------|--------------------------------------------------------------------------------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| `EnableVolumeReplication` | Explicit, idempotent attach. `Completed` once the relationship is established. | Implemented (§5.1, P0-2) | +| `DisableVolumeReplication` | Explicit stop/teardown verb (implicit before this design) | Implemented (§5.1, P0-2) | +| `PromoteVolume` (`force`) | Planned (peer alive) vs. unplanned (peer gone), wired through `force` | Implemented (§5.2), including the source-health no-op that keeps a healthy day-one promote from cloning unnecessarily | +| `DemoteVolume` | A new standalone verb: quiesce, final flush, confirm on peer | Implemented (§5.2, P0-3), exactly the "new standalone demote" Appendix A called for | +| `ResyncVolume` | A direction-reversing resync with `Resyncing`→`Completed` progress | Implemented (§5.2), maps onto `failback` | +| `GetVolumeReplicationInfo` | `lastSyncTime`, plus `lastSyncBytes`/`lastSyncDuration` | `lastSyncTime` implemented (§5.1, P0-1). `lastSyncBytes`/`lastSyncDuration` blocked on P0-4: the response type in the pinned `csi-addons/spec` v0.2.0 has no field for either, so the backend's own cycle-size and cycle-duration reads (`get_replication_info`'s `last_cycle_bytes`/`last_cycle_seconds`) have nowhere to land in the RPC response. | + +### 5.3 Status queries (Appendix A.4) + +1. **Per-slot health/role, never 404ing:** implemented (P0-1, `design-csi-addons-replication.md` §6.1). `state: not_replicating, role: none` is the valid answer for an unreplicated volume, closing exactly the 404 gap Appendix A.4.1 named. +2. **The `PeerReady` boolean:** covered by §5.1 above, Ramen's own aggregation, needing no new code. +3. **Group readiness:** out of scope here (§2 Non-Goals), `design-consistency-groups.md`'s future work. + +### 5.4 Observability (Appendix B) + +| Appendix B concept | Metric | Status | +|--------------------------|------------------------------------------------------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| RPO, time gap | `simplyblock_replication_lag_seconds` | Implemented (`design-csi-addons-replication.md` §11) | +| RPO, data gap | `simplyblock_replication_backlog_bytes` | Implemented. This is Appendix B's "one genuinely new backend measurement," `bytesBehind`. `get_replication_info`'s existing `outstanding_bytes` (queued-but-unshipped snapshot sizes) already reports it. | +| Last cycle size/duration | `simplyblock_replication_last_sync_bytes`/`_seconds` | Implemented | +| RPO compliance | `simplyblock_replication_rpo_violation` | Implemented, keyed on `ReplicationPolicy.rpo_target_seconds` | +| Degraded state | `simplyblock_replication_degraded` | Implemented | +| Throughput (smoothed) | (none) | Not implemented as its own series. Derivable via PromQL (`last_sync_bytes / last_sync_seconds`) from what already exists. Non-blocking. | +| RTO (estimated) | (none) | Not implemented. Appendix B itself frames this as "always an estimate," lowest priority of the set. Non-blocking. | +| Backup RPO/RTO (S3) | (none) | Out of scope. `design-csi-addons-replication.md` §2's own Non-Goals exclude the VolSync/S3 backup path entirely, a separate phase of the gap analysis. | + +Nothing in this row set blocks §6's validation: every metric `PeerReady` or a planned/unplanned promote depends on is already implemented. + +--- + +## 6. E2E Validation Plan + +### 6.1 Topology + +Two Kubernetes clusters, each running a `SimplyblockDriver` against its own simplyblock storage cluster, paired exactly as `regression_test/21/test_csi_addons_replication.sh` already sets up (one shared simplyblock control plane, `ReplicationPair`/`ReplicationPolicy` authoring the backend relationship). On top of that pair: an OCM hub (a third cluster, or the hub role colocated on one of the two), both storage clusters registered as `ManagedCluster`s, Ramen's dr-cluster operator installed on each, Ramen's hub operator installed on the hub, and one `DRPolicy` naming both clusters via their `StorageClass`/`VolumeReplicationClass` labels (`design-csi-addons-replication.md` §7.1's `ramendr.openshift.io/storageid`/`replicationid` labels, already implemented). A `DRPlacementControl` drives the workload, authored by SiteMap where available and hand-authored otherwise (P0-3). + +### 6.2 Test flow + +0. **Preflight.** Before the `DRPolicy` is authored, confirm §3's preflight passes against both clusters' own `StorageClass`/`VolumeReplicationClass` objects: a `PeerClassesVerified` event on each cluster's `ReplicationPair`, not a `PeerClassesMismatch`. +1. **Protect.** Deploy a workload with a PVC on cluster A, under a `StorageClass`/`VolumeReplicationClass` pair carrying the Ramen labels. Create the `DRPlacementControl`. Confirm Ramen's VRG creates one `VolumeReplication` for the PVC, `EnableVolumeReplication` fires, and `status.lastSyncTime` advances on the policy's ordinary cadence. +2. **Planned relocate.** Trigger Ramen's `Relocate` action. Confirm the sequence design-csi-addons-replication.md §5.2 documents drives correctly through Ramen rather than by hand: the source-side VRG demotes (fence, final flush, `Completed` settles once confirmed), then the target-side VRG promotes with `force=false`, gated on `PeerReady` from step 1's demote. Confirm the workload comes up on cluster B with the data a hashed writer wrote before relocation, and that Ramen's own `PeerReady`/`DataProtected` aggregation reports correctly throughout, not just this design's own conditions in isolation. +3. **Unplanned failover.** Simulate cluster B's loss (or reachability loss, matching this session's `replication_source_online` distinction) and trigger Ramen's `Failover` action against cluster A. Confirm the force-escalation path Ramen's own controller drives (§5.2's "no wait-and-retry grace period" behavior) still lands correctly when Ramen, not a test script, is the one issuing the calls. +4. **Resync.** Recover the failed cluster and confirm Ramen drives `ResyncVolume` to reconcile the diverged copy, without merging or re-triggering a full cutover. + +### 6.3 Pass/fail criteria + +Every step's pass criterion is an observable Ramen already reports on its own objects (`VolumeReplication.status.conditions`, the VRG's `DataProtected`, the DRPC's `PeerReady`), not a simplyblock-specific read: the point of this validation is that Ramen's own surface reflects reality, not that a side-channel confirms it. Data correctness is checked with a hashed writer across every promote, matching `design-csi-addons-replication.md` §12's existing E2E discipline. + +--- + +## 7. Testing Strategy + +Three different classes of coverage, for the three different things this document specifies: + +- **§3's preflight is unit-tested**, with the fake client and event recorder `replicationpair_controller_unit_test.go` already uses for the rest of `ReplicationPairReconciler`. An enrolled `StorageClass` with a matching, complete `VolumeReplicationClass` yields `PeerClassesVerified`. One with a missing label, a missing parameter, or no matching class at all yields `PeerClassesMismatch`, and no enrolled `StorageClass` yields no event at all. No `envtest` is needed, since the fake client registers the unstructured `VolumeReplicationClass` GVK the same way it does any typed kind. +- **§4's `VolumeGroupReplicationReconciler` is unit-tested against a fake client the same way**, standing in a `ConsistencyGroup`-labeled set of PVCs and asserting the member `VolumeReplication` fan-out and the status fan-in: a group whose members all report `Completed` yields a `Completed` group, any one member `Degraded` yields a `Degraded` group, and `status.lastGroupSyncTime` is the oldest member `lastSyncTime`, not the newest. The admission webhook (§4.4) is unit-tested the same way `design-consistency-groups.md` §9.4's tests already exercise its `VolumeGroupSnapshot` sibling: a selector matching a whole group admitted, one matching a subset or spanning two groups rejected. +- **§5 confirms existing code** and adds nothing to test on its own. **§6 is exclusively E2E, live-cluster validation** with no smaller harness to substitute, including its step 0 confirmation that §3's preflight passed before the `DRPolicy` was authored, and a group-protect scenario confirming §4's reconciler under a real VRG's `VolumeGroupReplication`. + +Full scenario detail: [`tests/test-plan-ramen-integration.md`](../tests/test-plan-ramen-integration.md). Its M-01 and M-02 close `design-csi-addons-replication.md` test plan's E-06 and E-07, which have carried no implementing test since they were written. + +--- + +## 8. Open Questions + +| # | Question | Owner | +|-----|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-----------------------| +| 1 | **Is a real OCM hub with Ramen already available for this validation**, or does P0-1/P0-2 need to be stood up from scratch? The answer decides whether §6 is schedulable now or needs its own infrastructure work first. | Operator team / Infra | +| 2 | **Does SiteMap exist in a runnable form yet**, or does every run of §6 use a hand-authored `DRPlacementControl` stand-in (P0-3)? If SiteMap is not yet runnable, note that explicitly rather than blocking on it indefinitely. | Operator team | +| 3 | **`csi-addons/spec` version floor (P0-4).** Confirm whether a newer pinned version already carries `lastSyncBytes`/`lastSyncDuration` before treating this as a real upstream gap to track. | Operator team | +| 4 | **Event volume on a large fleet.** `PeerClassesVerified` fires every 60-second reconcile once the preflight passes, on every `ReplicationPair`. Confirm this matches the existing event-rate expectations `rbac-hardening`'s workload review sets, or whether the verified case should log rather than emit an event, firing only once per state change. | Operator team | +| 5 | **Is Global VGR needed.** §4.1 scopes `VolumeGroupReplication` to the base, single-VRG case. Confirm whether any planned simplyblock deployment spans a replication group across more than one application's VRG before treating RamenDR's multi-VRG consensus feature as work this document should also specify. | Operator team | +| 6 | **`VolumeGroupReplicationContent`'s exact contract.** §4.3 specifies the reconciler against `VolumeGroupReplication.spec`/`.status` only. Confirm what, if anything, this reconciler must also write to `VolumeGroupReplicationContent` before implementation starts, against the CRD's real schema rather than this document's reading of it. | Operator team | diff --git a/operator/docs/tests/test-plan-csi-addons-replication.md b/operator/docs/tests/test-plan-csi-addons-replication.md index 1db5bdc8b..b7f13bc56 100644 --- a/operator/docs/tests/test-plan-csi-addons-replication.md +++ b/operator/docs/tests/test-plan-csi-addons-replication.md @@ -139,10 +139,10 @@ Two live simplyblock clusters with the chart-deployed csi-addons machinery. The ### Ramen-Driven (design §12, Phase 2 gate) -| # | Scenario | Type | Test | -|------|----------------------------------------------------------------------------------------------------------------------------------------------|----------|------| -| E-06 | A Ramen `VolumeReplicationGroup` in async mode selects the class, creates one `VolumeReplication` per PVC, and `lastGroupSyncTime` populates | Positive | — | -| E-07 | Ramen failover (`force`) and relocate (demote plus planned promote) both complete against a live workload | Positive | — | +| # | Scenario | Type | Test | +|------|----------------------------------------------------------------------------------------------------------------------------------------------|----------|---------------------------------------------------------------------------------| +| E-06 | A Ramen `VolumeReplicationGroup` in async mode selects the class, creates one `VolumeReplication` per PVC, and `lastGroupSyncTime` populates | Positive | → [`test-plan-ramen-integration.md`](test-plan-ramen-integration.md) M-01 | +| E-07 | Ramen failover (`force`) and relocate (demote plus planned promote) both complete against a live workload | Positive | → [`test-plan-ramen-integration.md`](test-plan-ramen-integration.md) M-02, M-03 | --- @@ -207,6 +207,7 @@ Phase 1 landed the driver's Replication and Identity services, the error classif | U-18 … U-22 | Condition derivation (`Completed`/`Degraded`/`Resyncing`) from the status read | Not implemented: `csi-addons/spec` v0.2.0's `GetVolumeReplicationInfoResponse` carries only `lastSyncTime`, with no per-condition field at all; deriving these needs either a newer spec version or belongs in the controller-manager's own reconcile, neither examined yet | | U-23 … U-27 | Preflight (`peerClasses` verification) and coexistence (`PVCReplicationController`, the one-owner rule) | Out of Phase 1 and Phase 2's scope: the auto-adapter and preflight webhook are a separate, unbuilt subsystem | | I-01 … I-07 | The sidecar and controller-manager loop | The driver's Replication and Identity services and the sidecar container now exist (Phase 1); no envtest/kind suite exercises them against the real kubernetes-csi-addons controller-manager yet | -| E-01 … E-07 | The live lifecycle and the Ramen gate | Blocked on Phase 2 landing, plus a two-cluster test bed with Ramen dr-cluster installed for E-06 and E-07 | +| E-01 … E-05 | The live lifecycle (non-Ramen half) | Needs a two-cluster live test bed. `regression_test/21/` exercises the same lifecycle by hand but is not wired as an automated E2E suite. | +| E-06, E-07 | The Ramen-driven gate | Detailed scenario ownership moved to [`test-plan-ramen-integration.md`](test-plan-ramen-integration.md) M-01 … M-03, blocked there on a live OCM hub with Ramen installed (see that document's Phase 0) | | — | Repeated resync, class drift after verification, annotated-volume migration onto the adapter, cascaded topologies | Beyond the first coverage pass, recorded so the gaps are explicit rather than assumed covered | | M-01, M-02 | Demote under writes; concurrent ownership race | Need failure injection and precise timing a live two-cluster run does not automate yet | diff --git a/operator/docs/tests/test-plan-ramen-integration.md b/operator/docs/tests/test-plan-ramen-integration.md new file mode 100644 index 000000000..8740c0e72 --- /dev/null +++ b/operator/docs/tests/test-plan-ramen-integration.md @@ -0,0 +1,134 @@ +# Test Plan: Ramen Integration + +Related design: [`designs/design-ramen-integration.md`](../designs/design-ramen-integration.md) +Harness: two classes. The peerClasses preflight and `VolumeGroupReplicationReconciler` (design §3, §4) are unit-tested against fake-client harnesses (`replicationpair_controller_unit_test.go` and its `VolumeGroupReplicationReconciler` sibling). Everything else needs a live two-cluster simplyblock deployment (`regression_test/21/`), plus an OCM hub with Ramen installed on the hub and both managed clusters (design §6.1). + +Scope: three things this document specifies. First, whether the peerClasses preflight (design §3) correctly distinguishes a complete `StorageClass`/`VolumeReplicationClass` pairing from one missing a label or a parameter, as unit scenarios (`U-01` … `U-05`). Second, whether `VolumeGroupReplicationReconciler` (design §4) correctly fans a group's replication state out to its members and their status back in, and correctly validates group membership at admission, as unit scenarios (`U-06` … `U-10`). Third, whether a real Ramen `VolumeReplicationGroup`, driven through a real OCM hub, correctly drives the csi-addons adapter `design-csi-addons-replication.md` implements, for both a single volume and a consistency-group of them, as manual E2E scenarios (`M-`), since no smaller harness substitutes for a real Ramen reconcile loop against a real OCM-registered cluster pair. The manual scenarios close two rows already carried in [`test-plan-csi-addons-replication.md`](test-plan-csi-addons-replication.md): E-06 and E-07, both `—` in that plan's `Test` column since they were written. + +--- + +## 1. Unit Scenarios + +### ReplicationPairReconciler peerClasses preflight (design §3) + +Not yet implemented. The `Test` column is `—` throughout until `replicationpair_controller.go` carries the preflight and `replicationpair_controller_unit_test.go` carries these cases. + +| # | Scenario | Type | Test | +|------|------------------------------------------------------------------------------------------------------------|----------|------| +| U-01 | Enrolled `StorageClass`, matching `VolumeReplicationClass` carries both the label and `schedulingInterval` | Positive | — | +| U-02 | Enrolled `StorageClass`, matching `VolumeReplicationClass` missing `ramendr.openshift.io/replicationid` | Negative | — | +| U-03 | Enrolled `StorageClass`, matching `VolumeReplicationClass` missing `spec.parameters.schedulingInterval` | Negative | — | +| U-04 | Enrolled `StorageClass`, no `VolumeReplicationClass` with a matching `spec.provisioner` exists | Negative | — | +| U-05 | No `StorageClass` carries `ramendr.openshift.io/storageid` (nothing enrolled) | Boundary | — | + +### VolumeGroupReplicationReconciler and its admission webhook (design §4) + +Not yet implemented. The `Test` column is `—` throughout until `VolumeGroupReplicationReconciler` and its webhook exist. + +| # | Scenario | Type | Test | +|------|-------------------------------------------------------------------------------------------------------------------|----------|------| +| U-06 | Every member's `VolumeReplication` reports `Completed=True, Degraded=False`, and the group reports the same | Positive | — | +| U-07 | One member's `VolumeReplication` reports `Degraded=True`, and the group reports `Degraded=True` | Negative | — | +| U-08 | Members report differing `lastSyncTime`, and `status.lastGroupSyncTime` is the oldest, not the newest | Boundary | — | +| U-09 | `spec.source.selector` resolves to exactly one `ConsistencyGroup`'s current membership, and is admitted | Positive | — | +| U-10 | `spec.source.selector` resolves to a subset of a group, or spans two groups, and is rejected, naming the mismatch | Negative | — | + +--- + +## 2. Manual Scenarios and Test Concepts + +### M-01: Ramen protects a workload, and `VolumeReplication` reflects it + +**Design reference:** design §6.2 step 1. Closes `test-plan-csi-addons-replication.md` E-06. + +**What to verify:** A Ramen `VolumeReplicationGroup` in async mode, given a `DRPlacementControl` protecting a namespace, selects the `VolumeReplicationClass` by its `replicationClassSelector` and `provisioner`, creates exactly one `VolumeReplication` per protected PVC, and that object's `status.lastSyncTime` advances on the policy's ordinary cadence, not just when driven by hand (already proven, `design-csi-addons-replication.md` §12), but when Ramen itself is the caller. + +**Open question:** whether SiteMap authors the `DRPlacementControl` or a hand-authored stand-in does (design §8, Open Question 2). Record which was used. + +**Test concept:** +1. Stand up the topology in design §6.1: two managed clusters, one shared simplyblock control plane, an OCM hub with Ramen installed, a `DRPolicy` naming both clusters. +2. Deploy a workload with a PVC on cluster A under a labeled `StorageClass`/`VolumeReplicationClass` pair (`ramendr.openshift.io/storageid`/`replicationid`, `design-csi-addons-replication.md` §7.1). +3. Create the `DRPlacementControl` protecting the workload's namespace. +4. Assert: exactly one `VolumeReplication` exists, named and owned per Ramen's own convention. Its `status.state` reaches `Primary`. `status.lastSyncTime` is non-nil and advances across at least two policy intervals. The VRG's `DataProtected` condition is `True`. + +### M-02: Ramen planned relocate, demote then promote, driven by the VRG + +**Design reference:** design §6.2 step 2. Closes `test-plan-csi-addons-replication.md` E-07 (relocate half). + +**What to verify:** Ramen's `Relocate` action demotes the source VRG and promotes the target VRG with `force=false`, gated on the source's `PeerReady`, and the workload comes up on the target cluster with the data a hashed writer wrote before relocation: zero loss, the same guarantee `design-csi-addons-replication.md` §5.2 already proves by hand, now proven through Ramen's own reconcile loop. + +**Current behavior:** the demote → planned-promote sequence, the `Aborted`/`FailedPrecondition` split that keeps a converging demote from being force-escalated, and the source-health no-op on an already-settled promote are all implemented and unit/E2E-tested against a hand-driven `VolumeReplication` (`design-csi-addons-replication.md` §5.2, §12). Never yet exercised by Ramen's own controller issuing the calls. + +**Test concept:** +1. From M-01's protected state, write and hash a known data set to the workload's volume. +2. Trigger Ramen's `Relocate` action toward cluster B. +3. Assert, in order: cluster A's `VolumeReplication` reaches `Secondary` with `Completed=True` (demote confirmed). Cluster B's `VolumeReplication` reaches `Primary` with `Completed=True` and no force-promotion event fired (the planned path, not an escalation). The workload is schedulable and serving on cluster B. The hash taken before relocation matches the data read after. + +### M-03: Ramen unplanned failover, force promote, driven by the VRG + +**Design reference:** design §6.2 step 3. Closes `test-plan-csi-addons-replication.md` E-07 (failover half). + +**What to verify:** with cluster A unreachable (not merely demoted), Ramen's `Failover` action promotes cluster B with `force=true`, and the force-escalation behavior `design-csi-addons-replication.md` §5.2 documents (the vendored controller's own "no wait-and-retry grace period") still lands correctly when Ramen, not a test script, drives it. + +**Test concept:** +1. From a protected, healthy state (M-01), make cluster A's storage genuinely unreachable (network partition or node shutdown, not merely a demote). +2. Trigger Ramen's `Failover` action toward cluster B. +3. Assert: cluster B's `VolumeReplication` reaches `Primary`. The promote succeeded via the forced path (a force-promotion event is expected here, unlike M-02). The workload serves on cluster B. + +### M-04: Ramen-driven resync after recovery + +**Design reference:** design §6.2 step 4 (new scope this document adds, with no corresponding row in `test-plan-csi-addons-replication.md`). + +**What to verify:** once cluster A recovers after M-03's failover, Ramen reconciles the diverged copy via `ResyncVolume` without merging or re-triggering a full cutover, matching `design-csi-addons-replication.md` §5.2's "it never merges" guarantee. + +**Test concept:** +1. Restore cluster A's connectivity/node after M-03. +2. Confirm Ramen (not a manual `kubectl patch`) drives the recovered volume's `VolumeReplication` toward `Resyncing=True`, then `Completed=True` once caught up. +3. Assert the resync reconciled by delta, not a full copy (compare shipped bytes against the volume's total size), and that cluster B remains primary throughout. + +### M-05: Ramen protects and relocates a multi-volume app through `VolumeGroupReplication` + +**Design reference:** design §4, exercised through the topology and test flow §6 defines for the per-volume case. New scope this document adds, with no corresponding row in `test-plan-csi-addons-replication.md`. + +**What to verify:** a VRG whose PVCs share a `storage.simplyblock.io/consistency-group` label and a `VolumeGroupReplicationClass` carrying `ramendr.openshift.io/groupreplicationid` creates one `VolumeGroupReplication`, `VolumeGroupReplicationReconciler` (design §4.3) fans it out to one `VolumeReplication` per member and fans member status back into the group, and a planned relocate of the whole app moves every member together with none diverging. + +**Test concept:** +1. Extend M-01's topology: a workload with three PVCs sharing one `storage.simplyblock.io/consistency-group` value, protected by one VRG under a `VolumeGroupReplicationClass` naming that group's `ReplicationPolicy`. +2. Confirm exactly one `VolumeGroupReplication` exists and exactly three member `VolumeReplication` objects exist, each owned by it. +3. Trigger Ramen's `Relocate` action, as in M-02. +4. Assert: all three members reach `Secondary` on cluster A and `Primary` on cluster B together, not staggered. The group's own `status.lastGroupSyncTime` reflects the oldest member's `lastSyncTime` throughout (U-08). No member is left behind mid-relocate. + +--- + +## 3. Axis Coverage + +| Axis | Values covered | IDs | Not covered | +|-------------------------------|---------------------------------------------------------------------------------|------------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| peerClasses object state | complete, missing label, missing parameter, no matching class, nothing enrolled | U-01 … U-05 | a `VolumeReplicationClass` whose provisioner does not match (not a gap: design §3.1 step 4 skips it by design, so there is nothing to assert) | +| Group replication aggregation | all members healthy, one member degraded, differing `lastSyncTime` | U-06 … U-08 | more than one degraded member simultaneously (not a gap: the aggregation is a plain disjunction, one member already exercises it) | +| Group membership validation | selector equals membership, selector is a subset or spans two groups | U-09, U-10 | membership changing between admission and reconcile (the fail-open backend-unreachable case `design-consistency-groups.md` §9.4's sibling check already covers for the snapshot path) | +| Orchestrator | Ramen VRG async, hub-driven | M-01 … M-05 | direct `kubectl` lifecycle (already covered in `test-plan-csi-addons-replication.md`) | +| Cluster topology | two managed clusters, one relationship, hub-mediated | M-01 … M-05 | three-cluster (cascaded) topologies. SiteMap-authored `DRPlacementControl` specifically (§8 Open Question 2 may leave this a hand-authored stand-in) | +| Failure mode | planned relocate, unplanned failover, post-recovery resync, group relocate | M-02, M-03, M-04, M-05 | a demote that stalls mid-convergence while Ramen-driven (covered by hand in `test-plan-csi-addons-replication.md` M-01/M-02, not yet by Ramen) | + +--- + +## 4. Coverage Summary + +| Class | Scenarios | Covered | Not covered | +|--------------|-----------|---------|-------------| +| Unit | 10 | 0 | U-01 … U-10 | +| Manual (E2E) | 5 | 0 | M-01 … M-05 | + +--- + +## 5. What Is Not Yet Covered + +| # | Gap | Reason | +|-------------|-------------------------------------------------------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| U-01 … U-05 | The entire peerClasses preflight | Not yet implemented: design §3 specifies the behavior, `replicationpair_controller.go` does not yet carry it | +| U-06 … U-10 | `VolumeGroupReplicationReconciler` and its admission webhook | Not yet implemented: design §4 specifies the behavior, neither the reconciler nor the webhook exists yet | +| M-01 … M-05 | The entire Ramen-driven validation | Blocked on Phase 0 (design §Phase 0): a live OCM hub with Ramen installed across a registered two-cluster pair, not yet confirmed available (§8 Open Question 1). M-05 additionally needs the `VolumeGroupReplication` CRDs installed on both clusters. | +| — | Three-cluster / cascaded topologies | Out of scope for this document, and not part of the gap analysis's Appendix A either | +| — | SiteMap-authored (rather than hand-authored) `DRPlacementControl` | Depends on SiteMap's own availability (§8 Open Question 2) | +| — | Global VGR (multi-VRG consensus) | Out of scope for design §4.1, tracked as design §8 Open Question 5 | From 1fff099b2a5ace3ddb4a3c213f4e8b4c3e3db39f Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Mon, 21 Sep 2026 16:05:49 +0200 Subject: [PATCH 122/206] docs(operator): recovering a namespace that will not finish deleting Three different things blocked a namespace teardown on one cluster in one afternoon, and the recovery for each is a cluster-scoped object outside the namespace. None of them is guessable from the symptom, and the order matters: patching finalizers before deleting the webhook configurations half-works, which is the most confusing way for it to fail. The three are an aggregated APIService whose Service died with the namespace and took cluster-wide discovery down with it, finalizers with no operator left to release them behind webhooks whose failurePolicy is Fail, and pods holding volumes whose CSI node plugin was in the namespace being deleted. The admission policies are left as they are. The delete-guards among them assume a running operator, so failing closed when it is gone protects nothing, but changing that weakens admission for every deployment to make one recovery easier, and the recovery is three commands once they are written down. Also written down: which cluster-scoped objects survive because the finalizers that would have deleted them never ran, why that breaks the next install's ownership metadata, and that the cert-manager chain has issued certificates to things outside this release -- OpenBao's server certificate among them -- so deleting the issuer leaves those unable to renew. Co-Authored-By: Claude Opus 5 (1M context) --- operator/docs/runbooks/namespace-teardown.md | 149 +++++++++++++++++++ 1 file changed, 149 insertions(+) create mode 100644 operator/docs/runbooks/namespace-teardown.md diff --git a/operator/docs/runbooks/namespace-teardown.md b/operator/docs/runbooks/namespace-teardown.md new file mode 100644 index 000000000..89da8fbb8 --- /dev/null +++ b/operator/docs/runbooks/namespace-teardown.md @@ -0,0 +1,149 @@ +# Recovering a namespace that will not finish deleting + +What to do when `kubectl delete namespace simplyblock` sits in `Terminating`, +and why each blocker blocks. + +This is here because the recovery has an order, and the order is not guessable +from the symptom: the thing that unsticks the namespace is usually a +cluster-scoped object outside it. Every blocker below was hit on one cluster in +one afternoon. + +## The shape of the problem + +Deleting the namespace deletes the operator first, or close enough to first that +it makes no difference. Everything that then remains is something that needed the +operator: finalizers it would have released, admission webhooks it would have +answered, cluster-scoped objects its finalizers would have deleted. + +`helm uninstall` before deleting the namespace avoids all of it. The rest of this +page is for when that did not happen. + +## Read the conditions first + +```bash +kubectl get ns simplyblock -o jsonpath='{range .status.conditions[*]}{.type}={.status}: {.message}{"\n"}{end}' +``` + +The condition that is `True` names the blocker, and there may be more than one -- +they are cleared one at a time, so expect to come back here after each step. + +## Blocker 1: discovery fails + +``` +NamespaceDeletionDiscoveryFailure=True: ... metrics.simplyblock.io/v1alpha2: +stale GroupVersion discovery +``` + +An aggregated APIService whose backing Service died with the namespace. The +namespace controller cannot enumerate what to delete, so it does nothing at all. + +This one is not confined to the namespace: aggregated discovery is cluster-wide, +so `kubectl api-resources` is degraded for everyone until it is gone. + +```bash +kubectl delete apiservice v1alpha2.metrics.simplyblock.io +``` + +Nothing in the operator can prevent this. The APIService is cluster-scoped, so it +cannot carry an owner reference to anything in the namespace, and the finalizer +that would delete it needs the operator to still be running. + +## Blocker 2: finalizers with nobody to release them + +``` +NamespaceFinalizersRemaining=True: ... storage.simplyblock.io/storagenode-finalizer +in 6 resource instances +``` + +The operator is gone, so nothing will release them. It cannot be brought back +either: a Terminating namespace accepts no new objects. + +**Delete the webhook configurations first.** They are cluster-scoped, they +outlive the namespace, and the ones that intercept `update` or `delete` carry +`failurePolicy: Fail` -- so with the webhook's Service gone they refuse the very +patch that removes a finalizer: + +``` +Internal error occurred: failed calling webhook "vstoragenode.simplyblock.io": +... service "simplyblock-operator-webhook-service" not found +``` + +```bash +kubectl delete validatingwebhookconfiguration simplyblock-operator-validating-webhook-configuration +kubectl delete mutatingwebhookconfiguration simplyblock-operator-mutating-webhook-configuration + +for k in controlplanes operatorops simplyblockdrivers storageclusters storagenodes storagepools; do + kubectl get $k.storage.simplyblock.io -n simplyblock -o name | + xargs -r -I{} kubectl patch -n simplyblock {} --type=merge -p '{"metadata":{"finalizers":[]}}' +done +``` + +Patching before deleting the webhooks half-works, which is the confusing part: +the kinds whose validator is `create`-only are patched, and the rest are refused. + +What this skips is the operator's own teardown. Storage nodes are not removed +from the control plane -- but on a namespace delete the control plane is going +with them, so there is nothing left to deregister from. + +## Blocker 3: pods that cannot be unmounted + +``` +NamespaceDeletionContentFailure=True: ... unexpected items still remain in +namespace: simplyblock for gvr: /v1, Resource=pods +``` + +Pods with a deletion timestamp, no finalizers, and no progress. Check what they +mount: + +```bash +kubectl get pods -n simplyblock -o custom-columns='NAME:.metadata.name,STATUS:.status.phase,DELETED:.metadata.deletionTimestamp' +kubectl get pods -n simplyblock | grep csi-node +``` + +A pod holding a simplyblock volume cannot be unmounted once the CSI node plugin +is gone, and the plugin is in the namespace being deleted. The kubelet waits for +a `NodeUnstage` that nothing will answer. + +```bash +kubectl delete pod -n simplyblock --grace-period=0 --force +``` + +The mount is then left behind on the worker. Where the volume was a simplyblock +one whose cluster is also being deleted there is nothing to corrupt; where it was +not, unmount it on the node before reusing it. + +## Afterward: what survived + +Cluster-scoped objects the operator's finalizers would have deleted are still +there, because the finalizers never ran: + +```bash +for r in crd clusterrole clusterrolebinding apiservice csidriver sc \ + validatingwebhookconfiguration mutatingwebhookconfiguration; do + echo "$r:"; kubectl get $r 2>/dev/null | grep -i simplyblock | awk '{print " "$1}' +done +``` + +They matter for the next install. Helm's release metadata lives in Secrets inside +the namespace, so it died with it, and an object with no release to own it makes +the next `helm install` fail on ownership metadata. Either delete them or accept +adopting them by hand. + +**Leave the cert-manager chain alone until it is clear who else uses it.** +`simplyblock-certificate-authority-issuer` has issued certificates to things +outside this release before now -- OpenBao's server certificate among them -- and +deleting the issuer leaves those unable to renew: + +```bash +kubectl get certificate -A -o custom-columns='NS:.metadata.namespace,NAME:.metadata.name,ISSUER:.spec.issuerRef.name' +``` + +## Doing it in the right order next time + +```bash +helm uninstall simplyblock -n simplyblock +kubectl delete namespace simplyblock +``` + +The uninstall runs while the operator is alive, so finalizers are released, the +webhooks are answered, and the cluster-scoped objects go with the release. From d9893742105c0c86ee4f62bc814bee73aace3791 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Mon, 21 Sep 2026 16:03:33 +0100 Subject: [PATCH 123/206] added support for volumegroupreplication --- .../templates/roles/manager_role.yaml | 38 ++ .../simplyblock-operator-webhook.yaml | 19 + operator/cmd/main.go | 15 + operator/config/rbac/role.yaml | 38 ++ operator/config/webhook/manifests.yaml | 19 + .../docs/designs/design-ramen-integration.md | 108 ++-- .../docs/tests/test-plan-ramen-integration.md | 61 +-- .../volumegroupreplication_controller.go | 487 ++++++++++++++++++ ...megroupreplication_controller_unit_test.go | 294 +++++++++++ .../volumegroupreplication_validator.go | 201 ++++++++ .../volumegroupreplication_validator_test.go | 193 +++++++ 11 files changed, 1363 insertions(+), 110 deletions(-) create mode 100644 operator/internal/controller/volumegroupreplication_controller.go create mode 100644 operator/internal/controller/volumegroupreplication_controller_unit_test.go create mode 100644 operator/internal/webhook/volumegroupreplication_validator.go create mode 100644 operator/internal/webhook/volumegroupreplication_validator_test.go diff --git a/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml b/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml index 029e6a7e5..3174d4d58 100644 --- a/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml @@ -249,6 +249,44 @@ rules: - patch - update - watch +- apiGroups: + - replication.storage.openshift.io + resources: + - volumegroupreplicationclasses + verbs: + - get + - list + - watch +- apiGroups: + - replication.storage.openshift.io + resources: + - volumegroupreplications + verbs: + - get + - list + - patch + - update + - watch +- apiGroups: + - replication.storage.openshift.io + resources: + - volumegroupreplications/status + verbs: + - get + - patch + - update +- apiGroups: + - replication.storage.openshift.io + resources: + - volumereplications + verbs: + - create + - delete + - get + - list + - patch + - update + - watch - apiGroups: - snapshot.storage.k8s.io resources: diff --git a/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml b/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml index 3da26521a..65efe165e 100644 --- a/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml @@ -242,6 +242,25 @@ webhooks: resources: - storagepools sideEffects: None +- admissionReviewVersions: + - v1 + clientConfig: + service: + name: simplyblock-operator-webhook-service + namespace: {{ .Release.Namespace }} + path: /validate-replication-storage-openshift-io-v1alpha1-volumegroupreplication + failurePolicy: Fail + name: vvolumegroupreplication.simplyblock.io + rules: + - apiGroups: + - replication.storage.openshift.io + apiVersions: + - v1alpha1 + operations: + - CREATE + resources: + - volumegroupreplications + sideEffects: None - admissionReviewVersions: - v1 clientConfig: diff --git a/operator/cmd/main.go b/operator/cmd/main.go index 6c3b2aecf..5b4a1937e 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -742,6 +742,14 @@ func main() { setupLog.Error(err, "unable to create controller", "controller", "ReplicationPair") os.Exit(1) } + if err := (&controller.VolumeGroupReplicationReconciler{ + Client: mgr.GetClient(), + Scheme: mgr.GetScheme(), + Recorder: mgr.GetEventRecorder("volumegroupreplication-controller"), + }).SetupWithManager(mgr); err != nil { + setupLog.Error(err, "unable to create controller", "controller", "VolumeGroupReplication") + os.Exit(1) + } if err := (&controller.ReplicationSlotReconciler{ Client: mgr.GetClient(), Scheme: mgr.GetScheme(), @@ -851,6 +859,13 @@ func main() { }}) setupLog.Info("registered volumegroupsnapshot validating webhook") + mgr.GetWebhookServer().Register("/validate-replication-storage-openshift-io-v1alpha1-volumegroupreplication", + &webhook.Admission{Handler: &internalwebhook.VolumeGroupReplicationValidator{ + Client: mgr.GetClient(), + APIClient: webapi.NewClient(), + }}) + setupLog.Info("registered volumegroupreplication validating webhook") + mgr.GetWebhookServer().Register("/validate-storage-simplyblock-io-v1alpha2-volumegroupsnapshotops", &webhook.Admission{Handler: &internalwebhook.VolumeGroupSnapshotOpsValidator{Client: mgr.GetClient()}}) setupLog.Info("registered volumegroupsnapshotops validating webhook") diff --git a/operator/config/rbac/role.yaml b/operator/config/rbac/role.yaml index bdd175326..fd0597911 100644 --- a/operator/config/rbac/role.yaml +++ b/operator/config/rbac/role.yaml @@ -249,6 +249,44 @@ rules: - patch - update - watch +- apiGroups: + - replication.storage.openshift.io + resources: + - volumegroupreplicationclasses + verbs: + - get + - list + - watch +- apiGroups: + - replication.storage.openshift.io + resources: + - volumegroupreplications + verbs: + - get + - list + - patch + - update + - watch +- apiGroups: + - replication.storage.openshift.io + resources: + - volumegroupreplications/status + verbs: + - get + - patch + - update +- apiGroups: + - replication.storage.openshift.io + resources: + - volumereplications + verbs: + - create + - delete + - get + - list + - patch + - update + - watch - apiGroups: - snapshot.storage.k8s.io resources: diff --git a/operator/config/webhook/manifests.yaml b/operator/config/webhook/manifests.yaml index b403d0f35..52c62ed53 100644 --- a/operator/config/webhook/manifests.yaml +++ b/operator/config/webhook/manifests.yaml @@ -223,6 +223,25 @@ webhooks: resources: - storagepools sideEffects: None +- admissionReviewVersions: + - v1 + clientConfig: + service: + name: webhook-service + namespace: system + path: /validate-replication-storage-openshift-io-v1alpha1-volumegroupreplication + failurePolicy: Fail + name: vvolumegroupreplication.simplyblock.io + rules: + - apiGroups: + - replication.storage.openshift.io + apiVersions: + - v1alpha1 + operations: + - CREATE + resources: + - volumegroupreplications + sideEffects: None - admissionReviewVersions: - v1 clientConfig: diff --git a/operator/docs/designs/design-ramen-integration.md b/operator/docs/designs/design-ramen-integration.md index eb23a235d..5fdb7b235 100644 --- a/operator/docs/designs/design-ramen-integration.md +++ b/operator/docs/designs/design-ramen-integration.md @@ -1,6 +1,6 @@ # Design Document: Ramen Integration -**Status:** Draft (contract confirmed, peerClasses preflight specified, implementation and E2E validation pending) +**Status:** Draft (§3 confirms peerClasses needs no operator code, §4 VolumeGroupReplication implemented, E2E validation pending) **Author:** Israel Geoffrey (geoffrey1330) **Date:** 2026-09-21 **Test Plan:** [`tests/test-plan-ramen-integration.md`](../tests/test-plan-ramen-integration.md) @@ -16,7 +16,7 @@ | P0-3 | SiteMap, or a hand-authored `DRPlacementControl` standing in for it, driving the `DRPolicy` | Ecosystem | The E2E validation (§6) | Not shipped. SiteMap is an external document and system, and storage is explicitly outside its own scope. | | P0-4 | `csi-addons/spec` at a version whose `GetVolumeReplicationInfoResponse` carries `lastSyncBytes`/`lastSyncDuration` | Ecosystem | Full Appendix A.3 `GetVolumeReplicationInfo` | Not shipped: pinned at v0.2.0 today, which has neither field. | -Without P0-1 through P0-3 nothing in §6 can run, because the validation is E2E-only, live-cluster work that no mock or `envtest` substitutes for. §3's preflight and §4's `VolumeGroupReplication` reconciler are both unaffected: neither needs an OCM hub, a Ramen installation, or a `DRPolicy`, only this cluster's own objects. P0-4's absence is narrower: it leaves `GetVolumeReplicationInfo` reporting only `lastSyncTime`, never cycle size or duration, but Ramen's own `PeerReady` gate (§5.1) does not read either field, so P0-4 does not block the validation itself. +Without P0-1 through P0-3 nothing in §6 can run, because the validation is E2E-only, live-cluster work that no mock or `envtest` substitutes for. §4's `VolumeGroupReplication` reconciler is unaffected: it needs no OCM hub, no Ramen installation, and no `DRPolicy`, only this cluster's own objects. P0-4's absence is narrower: it leaves `GetVolumeReplicationInfo` reporting only `lastSyncTime`, never cycle size or duration, but Ramen's own `PeerReady` gate (§5.1) does not read either field, so P0-4 does not block the validation itself. --- @@ -24,8 +24,8 @@ Without P0-1 through P0-3 nothing in §6 can run, because the validation is E2E- 1. [Background](#1-background) 2. [Goals and Non-Goals](#2-goals-and-non-goals) -3. [peerClasses Preflight](#3-peerclasses-preflight) -4. [VolumeGroupReplication](#4-volumegroupreplication) +3. [peerClasses: Ramen's Own Mechanism](#3-peerclasses-ramens-own-mechanism) +4. [VolumeGroupReplication (Implemented)](#4-volumegroupreplication-implemented) 5. [The Contract, Confirmed](#5-the-contract-confirmed) 6. [E2E Validation Plan](#6-e2e-validation-plan) 7. [Testing Strategy](#7-testing-strategy) @@ -35,7 +35,7 @@ Without P0-1 through P0-3 nothing in §6 can run, because the validation is E2E- ## Overview -`design-csi-addons-replication.md` builds the storage-level adapter Ramen's per-volume DR contract requires, and validates it end to end "without Ramen" (its own §12): real backend, real csi-addons machinery, but a hand-driven `VolumeReplication` object rather than a real Ramen reconcile loop. That document deferred two pieces as "the next design": a `peerClasses` preflight (its own §7.2, once it became clear `peerClasses` itself is Ramen's mechanism, not this operator's) and `VolumeGroupReplication` (its own §2 Non-Goals, strictly per volume itself). `design-consistency-groups.md` deferred `VolumeGroupReplication` too, as future work independent of any replication policy. Neither document claims it. This document is where both land: §3 specifies the preflight, §4 specifies group replication on top of the consistency-group primitive, and §5 through §6 confirm the rest of the per-volume contract against what already shipped and specify the E2E validation that closes `design-csi-addons-replication.md` §12's outstanding acceptance gate (E-06, E-07). +`design-csi-addons-replication.md` builds the storage-level adapter Ramen's per-volume DR contract requires, and validates it end to end "without Ramen" (its own §12): real backend, real csi-addons machinery, but a hand-driven `VolumeReplication` object rather than a real Ramen reconcile loop. That document's own §7.2 named `peerClasses` verification Ramen's own hub-side mechanism, out of its scope, and deferred `VolumeGroupReplication` as "the next design," strictly per volume itself. `design-consistency-groups.md` deferred `VolumeGroupReplication` too, as future work independent of any replication policy. This document specifies the one piece that actually needed a new design: §4, `VolumeGroupReplication` on top of the consistency-group primitive. §3 confirms, rather than reopens, `design-csi-addons-replication.md` §7.2's original position on peerClasses, after this document's own history of first building a same-cluster preflight for it and then removing that preflight once it became clear a same-cluster check cannot verify a cross-cluster pairing. §5 through §6 confirm the rest of the per-volume contract against what already shipped and specify the E2E validation that closes `design-csi-addons-replication.md` §12's outstanding acceptance gate (E-06, E-07). --- @@ -45,9 +45,9 @@ The gap analysis's own headline finding (§2) was that simplyblock's DR machiner `design-csi-addons-replication.md`'s three phases implement that spec. This document does not repeat what it built. §5 below cites, section by section, where each Appendix A and Appendix B item now lives in the shipped code. -The other half of Appendix A's own premise is that Ramen's hub, not this operator, computes `peerClasses` and drives the VRG, through OCM's hub-spoke visibility into every managed cluster. That hub, and the OCM/SiteMap layer above it, is external to this repository (confirmed this session: no `ManagedCluster`, `DRPolicy`, or `VolumeReplicationGroup` reference exists anywhere in this codebase outside design-doc prose, and no mechanism for this operator to reach a peer cluster's Kubernetes API exists or is needed. `design-management-hub.md`, this repo's own hub design, is a separate fleet-config-distribution concern that cites Ramen's hub/spoke split only as precedent, not as something it builds). Ramen's hub already reports when it cannot find a valid `StorageClass`/`VolumeReplicationClass` pairing across two clusters, once it looks. What it has no way to see is a pairing that is wrong on one cluster alone, before any hub-side comparison happens: a `VolumeReplicationClass` missing the `ramendr.openshift.io/replicationid` label, or carrying no `schedulingInterval`, sits invisibly broken until a `DRPolicy` is authored against it and the hub's own reconcile surfaces a cryptic cross-cluster mismatch instead of the local, fixable cause. §3 closes that local gap. +The other half of Appendix A's own premise is that Ramen's hub, not this operator, computes `peerClasses` and drives the VRG, through OCM's hub-spoke visibility into every managed cluster. That hub, and the OCM/SiteMap layer above it, is external to this repository (confirmed this session: no `ManagedCluster`, `DRPolicy`, or `VolumeReplicationGroup` reference exists anywhere in this codebase outside design-doc prose, and no mechanism for this operator to reach a peer cluster's Kubernetes API exists or is needed. `design-management-hub.md`, this repo's own hub design, specifies a real hub component of its own (the `fleet-manager`, reading member clusters through OCM's `ManagedClusterView`), but it is a separate fleet-config-distribution concern, and nothing in it builds or is intended to build Ramen's own pairing). Ramen's hub already reports when it cannot find a valid `StorageClass`/`VolumeReplicationClass` pairing across two clusters, once it looks, and only the hub's own cross-cluster visibility can make that comparison at all: a same-cluster read sees whether this cluster's own label is present, never whether its value agrees with the peer's, which is the only question a pairing check actually needs answered. §3 explains why this operator's own history of trying to close that gap locally settled on not closing it. -The gap analysis's own §6 named a second piece of genuinely new work, in its Phase 2: `VolumeGroupReplication`, "on top of the CG primitive," so that the VRG async group path can protect and fail over a multi-volume app at one point rather than only snapshot it. `design-consistency-groups.md` built the CG primitive that gap analysis cites (the `storage.simplyblock.io/consistency-group` label, `VolumeGroupSnapshot`, the `GroupController`) but explicitly left group replication for later, independent of any policy, exactly so this document could attach it without reshaping the group. §4 is that attachment. Together, §3 and §4 are the only new production code this document proposes: everything else Appendix A and Appendix B specify is confirmed, in §5, against code `design-csi-addons-replication.md` already shipped. +The gap analysis's own §6 named a second piece of genuinely new work, in its Phase 2: `VolumeGroupReplication`, "on top of the CG primitive," so that the VRG async group path can protect and fail over a multi-volume app at one point rather than only snapshot it. `design-consistency-groups.md` built the CG primitive that gap analysis cites (the `storage.simplyblock.io/consistency-group` label, `VolumeGroupSnapshot`, the `GroupController`) but explicitly left group replication for later, independent of any policy, exactly so this document could attach it without reshaping the group. §4 is that attachment, and the only new production code this document proposes: everything else Appendix A and Appendix B specify is confirmed, in §5, against code `design-csi-addons-replication.md` already shipped. --- @@ -55,7 +55,7 @@ The gap analysis's own §6 named a second piece of genuinely new work, in its Ph ### Goals -- Specify and implement a same-cluster `peerClasses` preflight: catch a `StorageClass`/`VolumeReplicationClass` pairing that Ramen's contract requires but this cluster's objects do not satisfy, before a `DRPolicy` is ever authored against it (§3). +- Confirm that Ramen's `peerClasses` pairing needs no operator-side code, settling the question this document's own earlier attempt at a same-cluster preflight left open (§3). - Specify and implement `VolumeGroupReplication` on top of the consistency-group primitive: fan a group's `primary`/`secondary`/`resync` intent out to its members' existing per-volume adapter, and fan their status back in, satisfying Appendix A.4's group-readiness status query and the gap analysis's own Phase 2 ask (§4). - Confirm, against the actual shipped code, that every condition, verb, and status query Appendix A specifies is satisfied, or state precisely which is not and why (§5). - Confirm, against the actual shipped code, which Appendix B metrics are delivered, which are derivable from what already exists, and which remain future work (§5.4). @@ -63,9 +63,9 @@ The gap analysis's own §6 named a second piece of genuinely new work, in its Ph ### Non-Goals -- **Building any hub, OCM, or cross-cluster Kubernetes access mechanism.** That is Ramen's and OCM's job, external to this operator, confirmed in §1. Neither §3's preflight nor §4's group reconciler reaches past this cluster's own objects, and nothing in §6's validation plan asks this operator to reach a peer cluster's API server: every step drives objects on the cluster where the workload currently runs, exactly as `design-csi-addons-replication.md`'s own architecture already assumes. -- **Resolving a `peerClasses` mismatch across two clusters.** That comparison needs the hub's own visibility into both clusters and is Ramen's job once a `DRPolicy` exists. §3 catches what is locally wrong before that comparison ever runs, and it does not repeat the comparison itself. -- **Global VGR.** RamenDR's newer multi-VRG consensus feature for a replication group spanning several applications is out of scope. §4 covers the base case: one VRG, one storage vendor's PVCs, one `VolumeGroupReplication` (§4.1, §8 Open Question 5). +- **Building any hub, OCM, or cross-cluster Kubernetes access mechanism.** That is Ramen's and OCM's job, external to this operator, confirmed in §1. §4's group reconciler does not reach past this cluster's own objects, and nothing in §6's validation plan asks this operator to reach a peer cluster's API server: every step drives objects on the cluster where the workload currently runs, exactly as `design-csi-addons-replication.md`'s own architecture already assumes. +- **A same-cluster `peerClasses` preflight, in any form.** Tried once, in this document's own history, and removed (§3): a same-cluster read can confirm a label is present, never that its value agrees with the peer's, and a pairing check that cannot verify agreement is not a pairing check. That comparison needs the hub's own visibility into both clusters and is Ramen's job alone, once a `DRPolicy` exists. +- **Global VGR.** RamenDR's newer multi-VRG consensus feature for a replication group spanning several applications is out of scope. §4 covers the base case: one VRG, one storage vendor's PVCs, one `VolumeGroupReplication` (§4.1, §8 Open Question 4). - **A new gRPC verb for group replication.** §4.2 is explicit: group promote, demote, and resync fan the same three verbs `design-csi-addons-replication.md` §5 already ships out to every member. Nothing changes on the driver. - **SiteMap.** A separate external document and system. Where §6's topology needs a `DRPlacementControl` and SiteMap is not available to author one, a hand-authored stand-in is explicitly permitted (P0-3). - **`bytesBehind`'s remaining Appendix B siblings** (throughput, RTO estimation, backup RPO/RTO). §5.4 accounts for each, and none blocks the validation this document specifies. @@ -73,61 +73,25 @@ The gap analysis's own §6 named a second piece of genuinely new work, in its Ph --- -## 3. peerClasses Preflight +## 3. peerClasses: Ramen's Own Mechanism -Ramen pairs a `StorageClass` and a `VolumeReplicationClass` across two managed clusters into a `peerClasses` entry on the hub's `DRPolicy`, using labels this operator's own convention already defines: `ramendr.openshift.io/storageid` on the `StorageClass`, `ramendr.openshift.io/replicationid` on the `VolumeReplicationClass`, both cited in `design-csi-addons-replication.md` §7.1. When the pairing is wrong on one cluster alone, the earliest and most legible place to catch it is that cluster, before a `DRPolicy` ever compares it against a peer. +Ramen's hub pairs a `StorageClass` and a `VolumeReplicationClass` across two managed clusters into a `peerClasses` entry on its `DRPolicy`, through OCM's hub-spoke visibility into both clusters at once. That visibility is what makes the pairing check meaningful, and it is exactly what a reconciler running inside one managed cluster does not have and cannot substitute for: the only question worth asking about a pairing is whether the two clusters' objects *agree*, and agreement is a cross-cluster comparison, not a same-cluster one. -### 3.1 What it checks +This document tried the same-cluster version anyway, once. A `ReplicationPairReconciler` preflight was specified and implemented, confirming this cluster's own `VolumeReplicationClass` carried the `ramendr.openshift.io/replicationid` label and a `schedulingInterval`, and emitting `PeerClassesVerified`/`PeerClassesMismatch` accordingly. That check has a real, narrow use (it catches an operator who forgot the label entirely), but it is not peerClasses verification: a cluster whose label carries a value that does not match its peer's passes identically to one whose value is correct, because presence, not agreement, is everything a same-cluster read can check. Reporting `PeerClassesVerified` under that name risked being read as a stronger guarantee than it delivered, so it has been removed. Nothing in this repository performs this check today, which is the position `design-csi-addons-replication.md` §7.2 already took before this document first tried to revisit it. -The `ReplicationPairReconciler` (`operator/internal/controller/replicationpair_controller.go`), on its existing 60-second reconcile cadence (`replPairSyncInterval`), adds a preflight step reading only this cluster's own objects: - -1. **List `StorageClass` objects** (typed `storagev1.StorageClass`, the same type `pvcreplication_controller.go` already reads), filtered to `Provisioner == "csi.simplyblock.io"`. -2. **Enrollment signal.** If none of this cluster's simplyblock `StorageClass` objects carries the `ramendr.openshift.io/storageid` label, there is nothing to preflight, and the check is a no-op: an unlabeled `StorageClass` is not a Ramen-managed one, and flagging it would be noise, not a finding. -3. **List `VolumeReplicationClass` objects.** This is an external CRD (`replication.storage.openshift.io/v1alpha1`, from `csi-addons`'s own `volume-replication-operator`) this repository does not vendor a Go type for, so the list reads `unstructured.UnstructuredList` against `schema.GroupVersionKind{Group: "replication.storage.openshift.io", Version: "v1alpha1", Kind: "VolumeReplicationClassList"}`, the same no-vendored-type pattern `internal/upgrade/discover/kinds.go` already uses for `cert-manager`'s `Certificate`. -4. **Per-class check**, for every listed `VolumeReplicationClass` whose `spec.provisioner` is `csi.simplyblock.io`: - - `metadata.labels["ramendr.openshift.io/replicationid"]` is present and non-empty. - - `spec.parameters.schedulingInterval` is present and non-empty. - - Neither check calls the simplyblock backend: whether the named `schedulingInterval` corresponds to a real `ReplicationPolicy` is `design-csi-addons-replication.md` §5's own province (`replicationpolicy_controller.go` already rejects an unresolvable policy at that layer), and re-checking it here would be the sbcli-backed preflight rejected earlier in this design's own history. This step verifies only that Ramen's own two labels and one parameter are present, which is everything a same-cluster read can verify. -5. **Verdict.** If step 2 found an enrolled `StorageClass` and step 4 found no `VolumeReplicationClass` with a matching provisioner, or found one with a missing label or parameter, the preflight fails. Otherwise, it passes. - -### 3.2 Where the result goes - -One Kubernetes event per reconcile, on the `ReplicationPair` object driving the check, not a new CR and not a status field: the preflight is a diagnostic, and the `ReplicationPair`'s own status already carries the fields Ramen and the backend care about. - -| Event reason | Type | When | -|-----------------------|---------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| `PeerClassesVerified` | Normal | An enrolled `StorageClass` exists and every matching `VolumeReplicationClass` carries both the label and the parameter. | -| `PeerClassesMismatch` | Warning | An enrolled `StorageClass` exists but no matching `VolumeReplicationClass` does, or one is missing the label or the parameter. The message names the missing field and the object. | - -`ReplicationPairReconciler` gains a `Recorder events.EventRecorder` field, following the exact pattern `backuppolicy_controller.go` already uses (`r.Recorder.Eventf(object, nil, eventType, reason, reason, format, args...)`), wired in `operator/cmd/main.go` as `mgr.GetEventRecorder("replicationpair-controller")`. - -### 3.3 RBAC - -Two markers, both read-only, added to the reconciler's existing block: - -```go -// +kubebuilder:rbac:groups=storage.k8s.io,resources=storageclasses,verbs=get;list;watch -// +kubebuilder:rbac:groups=replication.storage.openshift.io,resources=volumereplicationclasses,verbs=get;list;watch -``` - -The second is this repository's second grant on an external CRD group it does not own, alongside `simplyblockstoragenodeset_controller.go`'s existing `cert-manager.io/certificates` marker, the precedent `rbac-hardening` cites for exactly this shape of grant. - -### 3.4 What this is not - -This is not the cross-cluster preflight the first pass at this design attempted and the second pass's user correction rejected: it makes no call to the simplyblock backend, resolves no `ReplicationPolicy` by name, and reaches no peer cluster's API server. It is the narrowest thing that closes the gap Ramen's own hub cannot see: a `StorageClass`/`VolumeReplicationClass` pairing broken on one cluster, before any `DRPolicy` compares it to a peer. +The exact contract Ramen's pairing requires of the two clusters' objects (the identity labels, the `StorageClass` name-matching, and the authoring convention this repository follows beyond what Ramen strictly requires) is recorded in `design-csi-addons-replication.md` §7.2, unchanged by this document. §4's `VolumeGroupReplication` work is unaffected by any of this: it needs no cross-cluster comparison of its own, only this cluster's own consistency-group membership (§4.2). --- -## 4. VolumeGroupReplication +## 4. VolumeGroupReplication (Implemented) The gap analysis's own Phase 2 named this the piece consistency groups still owed: "Add `VolumeGroupReplication` (csi-addons) on top of the CG primitive so the VRG async group path can protect and fail over multi-volume apps at one point, not just snapshot them." `design-csi-addons-replication.md` §2 deferred it as "the next design," strictly per-volume itself. `design-consistency-groups.md` §2 deferred it too, as "future work," independent of any replication policy so that work could attach later. Neither claims it. This section is that attachment. ### 4.1 What already exists, upstream -`VolumeGroupReplication`, `VolumeGroupReplicationClass`, and `VolumeGroupReplicationContent` are shipped CRDs in the same `replication.storage.openshift.io` group as `VolumeReplication` and `VolumeReplicationClass`, from `kubernetes-csi-addons`. `VolumeGroupReplication.spec` carries the same three-state `replicationState` (`primary`/`secondary`/`resync`) the per-volume kind carries, a `source.selector` naming the member PVCs by label, and an `external` boolean: `false` routes reconciliation through the generic kubernetes-csi-addons controller-manager, and `true` hands it to "an external controller managed by the storage vendor." Ramen's own VRG (`RamenDR/ramen`'s `vrg_volgrouprep.go`) creates and drives `VolumeGroupReplication` objects directly, once a VRG's PVCs share a `VolumeGroupReplicationClass` carrying a `ramendr.openshift.io/groupreplicationid` label, the group-level sibling of the per-volume `replicationid` label §3.1 already checks. Ramen's own documentation names this case "offloaded" replication: a storage backend that replicates at the LUN or logical-volume-store level, outside Kubernetes, through a vendor controller, exactly the shape simplyblock's backend already has. +`VolumeGroupReplication`, `VolumeGroupReplicationClass`, and `VolumeGroupReplicationContent` are shipped CRDs in the same `replication.storage.openshift.io` group as `VolumeReplication` and `VolumeReplicationClass`, from `kubernetes-csi-addons`. `VolumeGroupReplication.spec` carries the same three-state `replicationState` (`primary`/`secondary`/`resync`) the per-volume kind carries, a `source.selector` naming the member PVCs by label, and an `external` boolean: `false` routes reconciliation through the generic kubernetes-csi-addons controller-manager, and `true` hands it to "an external controller managed by the storage vendor." Ramen's own VRG (`RamenDR/ramen`'s `vrg_volgrouprep.go`) creates and drives `VolumeGroupReplication` objects directly, once a VRG's PVCs share a `VolumeGroupReplicationClass` carrying a `ramendr.openshift.io/groupreplicationid` label, the group-level sibling of the per-volume `replicationid` label `design-csi-addons-replication.md` §7.1 defines. Ramen's own documentation names this case "offloaded" replication: a storage backend that replicates at the LUN or logical-volume-store level, outside Kubernetes, through a vendor controller, exactly the shape simplyblock's backend already has. -**Global VGR, RamenDR's newer multi-VRG consensus feature for a replication group spanning several applications' VRGs, is out of scope here.** This section covers the base case Ramen has supported longer: one VRG, one storage vendor's PVCs, one `VolumeGroupReplication`. Whether the base case is sufficient for simplyblock's own use, or Global VGR's cross-VRG consensus is eventually needed too, is §8 Open Question 5. +**Global VGR, RamenDR's newer multi-VRG consensus feature for a replication group spanning several applications' VRGs, is out of scope here.** This section covers the base case Ramen has supported longer: one VRG, one storage vendor's PVCs, one `VolumeGroupReplication`. Whether the base case is sufficient for simplyblock's own use, or Global VGR's cross-VRG consensus is eventually needed too, is §8 Open Question 4. ### 4.2 No new gRPC contract @@ -135,27 +99,30 @@ Promoting, demoting, or resyncing a group is fanning the same three verbs `desig ### 4.3 The reconciler -A new `VolumeGroupReplicationReconciler`, alongside `ReplicationPairReconciler` and the rest, owns every `VolumeGroupReplication` whose `spec.external` is `true` and whose `spec.volumeGroupReplicationClassName` names a class with `provisioner: csi.simplyblock.io`: +`VolumeGroupReplicationReconciler` (`operator/internal/controller/volumegroupreplication_controller.go`), alongside `ReplicationPairReconciler` and the rest, owns every `VolumeGroupReplication` whose `spec.external` is `true` and whose `spec.volumeGroupReplicationClassName` names a class with `provisioner: csi.simplyblock.io`: 1. **Resolve membership.** Read `spec.source.selector` against this cluster's PVCs. Reuse the exact invariant `design-consistency-groups.md` §9.2 already established for `VolumeGroupSnapshot`: the selected set must equal a `storage.simplyblock.io/consistency-group` value's current membership exactly, not merely a subset or superset of it. A selector that does not resolve to one whole group is a configuration error, not a partial group to serve. 2. **Fan out.** For each member PVC, ensure a per-volume `VolumeReplication` object exists, owned by the `VolumeGroupReplication`, named deterministically from the group and the member, with `spec.replicationState` mirroring the group's. The already-shipped `kubernetes-csi-addons` controller-manager reconciles each of these exactly as it does any Ramen-created per-volume `VolumeReplication` (`design-csi-addons-replication.md` §5): this reconciler creates and updates the member objects, and never calls the driver's Replication gRPC itself. -3. **Fan in.** Aggregate every member's `VolumeReplication.status.conditions` into the group's own status: `Completed` is the conjunction across all members, `Degraded` and `Resyncing` are the disjunction (one degraded or resyncing member makes the group so). `status.lastGroupSyncTime`, the field Appendix A.4's group-readiness status query names, is the oldest of the members' `lastSyncTime`: a group's recovery point is only as fresh as its slowest member. +3. **Fan in.** Aggregate every member's `VolumeReplication.status.conditions` into the group's own status: `Completed` is the conjunction across all members, `Degraded` and `Resyncing` are the disjunction (one degraded or resyncing member makes the group so). `status.lastSyncTime` (the real upstream field on `VolumeGroupReplication.status`, mirroring the per-volume kind's own field name exactly rather than a distinct "group" field, confirmed against the shipped CRD schema) is the oldest of the members' `lastSyncTime`: a group's recovery point is only as fresh as its slowest member. ### 4.4 Admission webhook, extended -`design-consistency-groups.md` §9.4's validating webhook on `VolumeGroupSnapshot` create already enforces "the selector must equal the group's current membership" at `kubectl apply`, with the same fail-closed label check and fail-open backend check. A sibling webhook on `VolumeGroupReplication` create makes the identical two checks against the identical label, so a `VolumeGroupReplication` that could never resolve to one whole group is rejected before the reconciler in §4.3 ever sees it, rather than sitting unreconciled. +`design-consistency-groups.md` §9.4's validating webhook on `VolumeGroupSnapshot` create already enforces "the selector must equal the group's current membership" at `kubectl apply`, with the same fail-closed label check and fail-open backend check. `VolumeGroupReplicationValidator` (`operator/internal/webhook/volumegroupreplication_validator.go`) is the sibling webhook this section specified: the identical two checks against the identical label, plus an ownership gate matching `VolumeGroupSnapshotValidator`'s own (`spec.external: true` and a class attributed to `csi.simplyblock.io`, admitting everything else untouched since `replication.storage.openshift.io` is a shared upstream group), so a `VolumeGroupReplication` that could never resolve to one whole group is rejected before the reconciler in §4.3 ever sees it, rather than sitting unreconciled. ### 4.5 RBAC and events -`VolumeGroupReplicationReconciler` needs read access to `persistentvolumeclaims` (already granted elsewhere in this operator) and read-write access to the external `volumegroupreplications.replication.storage.openshift.io` and the per-member `volumereplications.replication.storage.openshift.io` it creates: +`VolumeGroupReplicationReconciler` needs read access to `persistentvolumeclaims`/`persistentvolumes` (declared again here for self-documentation, though already granted elsewhere in this operator), read access to `volumegroupreplicationclasses` to decide ownership (§4.3's own check, not listed in this section's first draft and added here to match the shipped code), and read-write access to the external `volumegroupreplications.replication.storage.openshift.io` and the per-member `volumereplications.replication.storage.openshift.io` it creates: ```go // +kubebuilder:rbac:groups=replication.storage.openshift.io,resources=volumegroupreplications,verbs=get;list;watch;update;patch // +kubebuilder:rbac:groups=replication.storage.openshift.io,resources=volumegroupreplications/status,verbs=get;update;patch +// +kubebuilder:rbac:groups=replication.storage.openshift.io,resources=volumegroupreplicationclasses,verbs=get;list;watch // +kubebuilder:rbac:groups=replication.storage.openshift.io,resources=volumereplications,verbs=get;list;watch;create;update;patch;delete +// +kubebuilder:rbac:groups="",resources=persistentvolumeclaims,verbs=get;list;watch +// +kubebuilder:rbac:groups="",resources=persistentvolumes,verbs=get;list;watch ``` -Events follow §3.2's pattern, on the `VolumeGroupReplication` object: `GroupReplicationVerified`/`GroupMembershipMismatch` at the webhook's own admission-time checks, mirrored as reconcile-time events for the case the webhook admitted open, and a `GroupReplicationDegraded` (Warning) when §4.3's fan-in first observes a member `Degraded` after previously reporting none. +`VolumeGroupReplicationReconciler` carries a `Recorder events.EventRecorder` field, following the exact pattern `backuppolicy_controller.go` already uses (`r.Recorder.Eventf(object, nil, eventType, reason, reason, format, args...)`), wired in `operator/cmd/main.go` as `mgr.GetEventRecorder("volumegroupreplication-controller")`. Events land on the `VolumeGroupReplication` object: `GroupReplicationVerified`/`GroupMembershipMismatch` on every reconcile's own membership check (§4.3 step 1, the same check the webhook makes at admission, mirrored here for the case the webhook admitted open), and a `GroupReplicationDegraded` (Warning) the first time fan-in observes `Degraded=True` after the group's own previous status reported it false, read back from the group object before the status update that reports the new value. --- @@ -211,7 +178,6 @@ Two Kubernetes clusters, each running a `SimplyblockDriver` against its own simp ### 6.2 Test flow -0. **Preflight.** Before the `DRPolicy` is authored, confirm §3's preflight passes against both clusters' own `StorageClass`/`VolumeReplicationClass` objects: a `PeerClassesVerified` event on each cluster's `ReplicationPair`, not a `PeerClassesMismatch`. 1. **Protect.** Deploy a workload with a PVC on cluster A, under a `StorageClass`/`VolumeReplicationClass` pair carrying the Ramen labels. Create the `DRPlacementControl`. Confirm Ramen's VRG creates one `VolumeReplication` for the PVC, `EnableVolumeReplication` fires, and `status.lastSyncTime` advances on the policy's ordinary cadence. 2. **Planned relocate.** Trigger Ramen's `Relocate` action. Confirm the sequence design-csi-addons-replication.md §5.2 documents drives correctly through Ramen rather than by hand: the source-side VRG demotes (fence, final flush, `Completed` settles once confirmed), then the target-side VRG promotes with `force=false`, gated on `PeerReady` from step 1's demote. Confirm the workload comes up on cluster B with the data a hashed writer wrote before relocation, and that Ramen's own `PeerReady`/`DataProtected` aggregation reports correctly throughout, not just this design's own conditions in isolation. 3. **Unplanned failover.** Simulate cluster B's loss (or reachability loss, matching this session's `replication_source_online` distinction) and trigger Ramen's `Failover` action against cluster A. Confirm the force-escalation path Ramen's own controller drives (§5.2's "no wait-and-retry grace period" behavior) still lands correctly when Ramen, not a test script, is the one issuing the calls. @@ -225,11 +191,10 @@ Every step's pass criterion is an observable Ramen already reports on its own ob ## 7. Testing Strategy -Three different classes of coverage, for the three different things this document specifies: +Two different classes of coverage, for the two different things this document specifies. §3 adds nothing to test: it confirms that no operator-side code exists for peerClasses, and there is no code left to exercise once the same-cluster preflight was removed. -- **§3's preflight is unit-tested**, with the fake client and event recorder `replicationpair_controller_unit_test.go` already uses for the rest of `ReplicationPairReconciler`. An enrolled `StorageClass` with a matching, complete `VolumeReplicationClass` yields `PeerClassesVerified`. One with a missing label, a missing parameter, or no matching class at all yields `PeerClassesMismatch`, and no enrolled `StorageClass` yields no event at all. No `envtest` is needed, since the fake client registers the unstructured `VolumeReplicationClass` GVK the same way it does any typed kind. -- **§4's `VolumeGroupReplicationReconciler` is unit-tested against a fake client the same way**, standing in a `ConsistencyGroup`-labeled set of PVCs and asserting the member `VolumeReplication` fan-out and the status fan-in: a group whose members all report `Completed` yields a `Completed` group, any one member `Degraded` yields a `Degraded` group, and `status.lastGroupSyncTime` is the oldest member `lastSyncTime`, not the newest. The admission webhook (§4.4) is unit-tested the same way `design-consistency-groups.md` §9.4's tests already exercise its `VolumeGroupSnapshot` sibling: a selector matching a whole group admitted, one matching a subset or spanning two groups rejected. -- **§5 confirms existing code** and adds nothing to test on its own. **§6 is exclusively E2E, live-cluster validation** with no smaller harness to substitute, including its step 0 confirmation that §3's preflight passed before the `DRPolicy` was authored, and a group-protect scenario confirming §4's reconciler under a real VRG's `VolumeGroupReplication`. +- **§4's `VolumeGroupReplicationReconciler` is unit-tested against a fake client** (`volumegroupreplication_controller_unit_test.go`), the same harness shape `replicationpair_controller_unit_test.go` already uses elsewhere in this package, pre-seeding member `VolumeReplication` objects with the status a prior fan-out would have created and asserting the fan-in aggregation: a group whose members all report `Completed` yields a `Completed` group, any one member `Degraded` yields a `Degraded` group, and `status.lastSyncTime` is the oldest member's, not the newest. The admission webhook (§4.4) is unit-tested the same way (`volumegroupreplication_validator_test.go`), mirroring how `design-consistency-groups.md` §9.4's own tests exercise its `VolumeGroupSnapshot` sibling: a selector matching a whole group admitted, one matching a subset or spanning two groups rejected, a backend outage admitted (fail-open). +- **§5 confirms existing code** and adds nothing to test on its own. **§6 is exclusively E2E, live-cluster validation** with no smaller harness to substitute, including a group-protect scenario confirming §4's reconciler under a real VRG's `VolumeGroupReplication`. Full scenario detail: [`tests/test-plan-ramen-integration.md`](../tests/test-plan-ramen-integration.md). Its M-01 and M-02 close `design-csi-addons-replication.md` test plan's E-06 and E-07, which have carried no implementing test since they were written. @@ -237,11 +202,10 @@ Full scenario detail: [`tests/test-plan-ramen-integration.md`](../tests/test-pla ## 8. Open Questions -| # | Question | Owner | -|-----|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-----------------------| -| 1 | **Is a real OCM hub with Ramen already available for this validation**, or does P0-1/P0-2 need to be stood up from scratch? The answer decides whether §6 is schedulable now or needs its own infrastructure work first. | Operator team / Infra | -| 2 | **Does SiteMap exist in a runnable form yet**, or does every run of §6 use a hand-authored `DRPlacementControl` stand-in (P0-3)? If SiteMap is not yet runnable, note that explicitly rather than blocking on it indefinitely. | Operator team | -| 3 | **`csi-addons/spec` version floor (P0-4).** Confirm whether a newer pinned version already carries `lastSyncBytes`/`lastSyncDuration` before treating this as a real upstream gap to track. | Operator team | -| 4 | **Event volume on a large fleet.** `PeerClassesVerified` fires every 60-second reconcile once the preflight passes, on every `ReplicationPair`. Confirm this matches the existing event-rate expectations `rbac-hardening`'s workload review sets, or whether the verified case should log rather than emit an event, firing only once per state change. | Operator team | -| 5 | **Is Global VGR needed.** §4.1 scopes `VolumeGroupReplication` to the base, single-VRG case. Confirm whether any planned simplyblock deployment spans a replication group across more than one application's VRG before treating RamenDR's multi-VRG consensus feature as work this document should also specify. | Operator team | -| 6 | **`VolumeGroupReplicationContent`'s exact contract.** §4.3 specifies the reconciler against `VolumeGroupReplication.spec`/`.status` only. Confirm what, if anything, this reconciler must also write to `VolumeGroupReplicationContent` before implementation starts, against the CRD's real schema rather than this document's reading of it. | Operator team | +| # | Question | Owner | +|-----|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-----------------------| +| 1 | **Is a real OCM hub with Ramen already available for this validation**, or does P0-1/P0-2 need to be stood up from scratch? The answer decides whether §6 is schedulable now or needs its own infrastructure work first. | Operator team / Infra | +| 2 | **Does SiteMap exist in a runnable form yet**, or does every run of §6 use a hand-authored `DRPlacementControl` stand-in (P0-3)? If SiteMap is not yet runnable, note that explicitly rather than blocking on it indefinitely. | Operator team | +| 3 | **`csi-addons/spec` version floor (P0-4).** Confirm whether a newer pinned version already carries `lastSyncBytes`/`lastSyncDuration` before treating this as a real upstream gap to track. | Operator team | +| 4 | **Is Global VGR needed.** §4.1 scopes `VolumeGroupReplication` to the base, single-VRG case. Confirm whether any planned simplyblock deployment spans a replication group across more than one application's VRG before treating RamenDR's multi-VRG consensus feature as work this document should also specify. | Operator team | +| 5 | **`VolumeGroupReplicationContent`'s exact contract, provisionally resolved.** The shipped reconciler writes nothing to it: `Content` objects bind a group the generic (non-external) controller-manager provisioned through a real CSI `GroupReplication` call, and this operator's `external: true` path never makes one, fanning out to per-volume `VolumeReplication` objects instead (§4.2). If a future Ramen version or tooling expects a `Content` object to exist even under `external: true`, revisit this. | Operator team | diff --git a/operator/docs/tests/test-plan-ramen-integration.md b/operator/docs/tests/test-plan-ramen-integration.md index 8740c0e72..4e96c5503 100644 --- a/operator/docs/tests/test-plan-ramen-integration.md +++ b/operator/docs/tests/test-plan-ramen-integration.md @@ -1,37 +1,25 @@ # Test Plan: Ramen Integration Related design: [`designs/design-ramen-integration.md`](../designs/design-ramen-integration.md) -Harness: two classes. The peerClasses preflight and `VolumeGroupReplicationReconciler` (design §3, §4) are unit-tested against fake-client harnesses (`replicationpair_controller_unit_test.go` and its `VolumeGroupReplicationReconciler` sibling). Everything else needs a live two-cluster simplyblock deployment (`regression_test/21/`), plus an OCM hub with Ramen installed on the hub and both managed clusters (design §6.1). +Harness: two classes. `VolumeGroupReplicationReconciler` (design §4) is unit-tested against a fake-client harness, the same shape `replicationpair_controller_unit_test.go` already uses elsewhere in this package. Everything else needs a live two-cluster simplyblock deployment (`regression_test/21/`), plus an OCM hub with Ramen installed on the hub and both managed clusters (design §6.1). -Scope: three things this document specifies. First, whether the peerClasses preflight (design §3) correctly distinguishes a complete `StorageClass`/`VolumeReplicationClass` pairing from one missing a label or a parameter, as unit scenarios (`U-01` … `U-05`). Second, whether `VolumeGroupReplicationReconciler` (design §4) correctly fans a group's replication state out to its members and their status back in, and correctly validates group membership at admission, as unit scenarios (`U-06` … `U-10`). Third, whether a real Ramen `VolumeReplicationGroup`, driven through a real OCM hub, correctly drives the csi-addons adapter `design-csi-addons-replication.md` implements, for both a single volume and a consistency-group of them, as manual E2E scenarios (`M-`), since no smaller harness substitutes for a real Ramen reconcile loop against a real OCM-registered cluster pair. The manual scenarios close two rows already carried in [`test-plan-csi-addons-replication.md`](test-plan-csi-addons-replication.md): E-06 and E-07, both `—` in that plan's `Test` column since they were written. +Scope: two things this document specifies. First, whether `VolumeGroupReplicationReconciler` (design §4) correctly fans a group's replication state out to its members and their status back in, and correctly validates group membership at admission, as unit scenarios (`U-01` … `U-05`). Second, whether a real Ramen `VolumeReplicationGroup`, driven through a real OCM hub, correctly drives the csi-addons adapter `design-csi-addons-replication.md` implements, for both a single volume and a consistency-group of them, as manual E2E scenarios (`M-`), since no smaller harness substitutes for a real Ramen reconcile loop against a real OCM-registered cluster pair. The manual scenarios close two rows already carried in [`test-plan-csi-addons-replication.md`](test-plan-csi-addons-replication.md): E-06 and E-07, both `—` in that plan's `Test` column since they were written. Design §3 (peerClasses) adds no scenarios of its own: it confirms that no operator-side code exists to test. --- ## 1. Unit Scenarios -### ReplicationPairReconciler peerClasses preflight (design §3) - -Not yet implemented. The `Test` column is `—` throughout until `replicationpair_controller.go` carries the preflight and `replicationpair_controller_unit_test.go` carries these cases. - -| # | Scenario | Type | Test | -|------|------------------------------------------------------------------------------------------------------------|----------|------| -| U-01 | Enrolled `StorageClass`, matching `VolumeReplicationClass` carries both the label and `schedulingInterval` | Positive | — | -| U-02 | Enrolled `StorageClass`, matching `VolumeReplicationClass` missing `ramendr.openshift.io/replicationid` | Negative | — | -| U-03 | Enrolled `StorageClass`, matching `VolumeReplicationClass` missing `spec.parameters.schedulingInterval` | Negative | — | -| U-04 | Enrolled `StorageClass`, no `VolumeReplicationClass` with a matching `spec.provisioner` exists | Negative | — | -| U-05 | No `StorageClass` carries `ramendr.openshift.io/storageid` (nothing enrolled) | Boundary | — | - ### VolumeGroupReplicationReconciler and its admission webhook (design §4) -Not yet implemented. The `Test` column is `—` throughout until `VolumeGroupReplicationReconciler` and its webhook exist. +Implemented in `volumegroupreplication_controller.go` and `volumegroupreplication_validator.go`, covered by `volumegroupreplication_controller_unit_test.go` and `volumegroupreplication_validator_test.go`. -| # | Scenario | Type | Test | -|------|-------------------------------------------------------------------------------------------------------------------|----------|------| -| U-06 | Every member's `VolumeReplication` reports `Completed=True, Degraded=False`, and the group reports the same | Positive | — | -| U-07 | One member's `VolumeReplication` reports `Degraded=True`, and the group reports `Degraded=True` | Negative | — | -| U-08 | Members report differing `lastSyncTime`, and `status.lastGroupSyncTime` is the oldest, not the newest | Boundary | — | -| U-09 | `spec.source.selector` resolves to exactly one `ConsistencyGroup`'s current membership, and is admitted | Positive | — | -| U-10 | `spec.source.selector` resolves to a subset of a group, or spans two groups, and is rejected, naming the mismatch | Negative | — | +| # | Scenario | Type | Test | +|------|-------------------------------------------------------------------------------------------------------------------|----------|-------------------------------------------------------------------| +| U-01 | Every member's `VolumeReplication` reports `Completed=True, Degraded=False`, and the group reports the same | Positive | `TestVolumeGroupReplication_AllMembersHealthyYieldsGroupHealthy` | +| U-02 | One member's `VolumeReplication` reports `Degraded=True`, and the group reports `Degraded=True` | Negative | `TestVolumeGroupReplication_OneMemberDegradedYieldsGroupDegraded` | +| U-03 | Members report differing `lastSyncTime`, and `status.lastSyncTime` is the oldest, not the newest | Boundary | `TestVolumeGroupReplication_LastSyncTimeIsTheOldestMember` | +| U-04 | `spec.source.selector` resolves to exactly one `ConsistencyGroup`'s current membership, and is admitted | Positive | `TestVolumeGroupReplicationValidator` | +| U-05 | `spec.source.selector` resolves to a subset of a group, or spans two groups, and is rejected, naming the mismatch | Negative | `TestVolumeGroupReplicationValidator` | --- @@ -96,29 +84,28 @@ Not yet implemented. The `Test` column is `—` throughout until `VolumeGroupRep 1. Extend M-01's topology: a workload with three PVCs sharing one `storage.simplyblock.io/consistency-group` value, protected by one VRG under a `VolumeGroupReplicationClass` naming that group's `ReplicationPolicy`. 2. Confirm exactly one `VolumeGroupReplication` exists and exactly three member `VolumeReplication` objects exist, each owned by it. 3. Trigger Ramen's `Relocate` action, as in M-02. -4. Assert: all three members reach `Secondary` on cluster A and `Primary` on cluster B together, not staggered. The group's own `status.lastGroupSyncTime` reflects the oldest member's `lastSyncTime` throughout (U-08). No member is left behind mid-relocate. +4. Assert: all three members reach `Secondary` on cluster A and `Primary` on cluster B together, not staggered. The group's own `status.lastSyncTime` reflects the oldest member's throughout (U-03). No member is left behind mid-relocate. --- ## 3. Axis Coverage -| Axis | Values covered | IDs | Not covered | -|-------------------------------|---------------------------------------------------------------------------------|------------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| peerClasses object state | complete, missing label, missing parameter, no matching class, nothing enrolled | U-01 … U-05 | a `VolumeReplicationClass` whose provisioner does not match (not a gap: design §3.1 step 4 skips it by design, so there is nothing to assert) | -| Group replication aggregation | all members healthy, one member degraded, differing `lastSyncTime` | U-06 … U-08 | more than one degraded member simultaneously (not a gap: the aggregation is a plain disjunction, one member already exercises it) | -| Group membership validation | selector equals membership, selector is a subset or spans two groups | U-09, U-10 | membership changing between admission and reconcile (the fail-open backend-unreachable case `design-consistency-groups.md` §9.4's sibling check already covers for the snapshot path) | -| Orchestrator | Ramen VRG async, hub-driven | M-01 … M-05 | direct `kubectl` lifecycle (already covered in `test-plan-csi-addons-replication.md`) | -| Cluster topology | two managed clusters, one relationship, hub-mediated | M-01 … M-05 | three-cluster (cascaded) topologies. SiteMap-authored `DRPlacementControl` specifically (§8 Open Question 2 may leave this a hand-authored stand-in) | -| Failure mode | planned relocate, unplanned failover, post-recovery resync, group relocate | M-02, M-03, M-04, M-05 | a demote that stalls mid-convergence while Ramen-driven (covered by hand in `test-plan-csi-addons-replication.md` M-01/M-02, not yet by Ramen) | +| Axis | Values covered | IDs | Not covered | +|-------------------------------|----------------------------------------------------------------------------|------------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| Group replication aggregation | all members healthy, one member degraded, differing `lastSyncTime` | U-01 … U-03 | more than one degraded member simultaneously (not a gap: the aggregation is a plain disjunction, one member already exercises it) | +| Group membership validation | selector equals membership, selector is a subset or spans two groups | U-04, U-05 | membership changing between admission and reconcile (the fail-open backend-unreachable case `design-consistency-groups.md` §9.4's sibling check already covers for the snapshot path) | +| Orchestrator | Ramen VRG async, hub-driven | M-01 … M-05 | direct `kubectl` lifecycle (already covered in `test-plan-csi-addons-replication.md`) | +| Cluster topology | two managed clusters, one relationship, hub-mediated | M-01 … M-05 | three-cluster (cascaded) topologies. SiteMap-authored `DRPlacementControl` specifically (§8 Open Question 2 may leave this a hand-authored stand-in) | +| Failure mode | planned relocate, unplanned failover, post-recovery resync, group relocate | M-02, M-03, M-04, M-05 | a demote that stalls mid-convergence while Ramen-driven (covered by hand in `test-plan-csi-addons-replication.md` M-01/M-02, not yet by Ramen) | --- ## 4. Coverage Summary -| Class | Scenarios | Covered | Not covered | -|--------------|-----------|---------|-------------| -| Unit | 10 | 0 | U-01 … U-10 | -| Manual (E2E) | 5 | 0 | M-01 … M-05 | +| Class | Scenarios | Covered | Not covered | +|--------------|-----------|-----------------|-------------| +| Unit | 5 | 5 (U-01 … U-05) | — | +| Manual (E2E) | 5 | 0 | M-01 … M-05 | --- @@ -126,9 +113,7 @@ Not yet implemented. The `Test` column is `—` throughout until `VolumeGroupRep | # | Gap | Reason | |-------------|-------------------------------------------------------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| U-01 … U-05 | The entire peerClasses preflight | Not yet implemented: design §3 specifies the behavior, `replicationpair_controller.go` does not yet carry it | -| U-06 … U-10 | `VolumeGroupReplicationReconciler` and its admission webhook | Not yet implemented: design §4 specifies the behavior, neither the reconciler nor the webhook exists yet | | M-01 … M-05 | The entire Ramen-driven validation | Blocked on Phase 0 (design §Phase 0): a live OCM hub with Ramen installed across a registered two-cluster pair, not yet confirmed available (§8 Open Question 1). M-05 additionally needs the `VolumeGroupReplication` CRDs installed on both clusters. | | — | Three-cluster / cascaded topologies | Out of scope for this document, and not part of the gap analysis's Appendix A either | | — | SiteMap-authored (rather than hand-authored) `DRPlacementControl` | Depends on SiteMap's own availability (§8 Open Question 2) | -| — | Global VGR (multi-VRG consensus) | Out of scope for design §4.1, tracked as design §8 Open Question 5 | +| — | Global VGR (multi-VRG consensus) | Out of scope for design §4.1, tracked as design §8 Open Question 4 | diff --git a/operator/internal/controller/volumegroupreplication_controller.go b/operator/internal/controller/volumegroupreplication_controller.go new file mode 100644 index 000000000..84cac75e2 --- /dev/null +++ b/operator/internal/controller/volumegroupreplication_controller.go @@ -0,0 +1,487 @@ +/* +Copyright 2025. + +Licensed under the Apache License, Version 2.0 (the "License"); +you may not use this file except in compliance with the License. +You may obtain a copy of the License at + + http://www.apache.org/licenses/LICENSE-2.0 + +Unless required by applicable law or agreed to in writing, software +distributed under the License is distributed on an "AS IS" BASIS, +WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +See the License for the specific language governing permissions and +limitations under the License. +*/ + +package controller + +import ( + "context" + "fmt" + "time" + + corev1 "k8s.io/api/core/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" + "k8s.io/apimachinery/pkg/runtime" + "k8s.io/apimachinery/pkg/runtime/schema" + "k8s.io/apimachinery/pkg/types" + "k8s.io/client-go/tools/events" + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + logf "sigs.k8s.io/controller-runtime/pkg/log" + + "github.com/simplyblock/atlas/kube" + atlaslvol "github.com/simplyblock/atlas/lvol" + + "github.com/simplyblock/simplyblock-operator/internal/utils" + "github.com/simplyblock/simplyblock-operator/internal/webapi" +) + +const ( + groupReplSyncInterval = 60 * time.Second + groupReplRequeueError = 30 * time.Second + + // consistencyGroupLabel names a PVC's consistency group (design-consistency-groups.md + // §4.1). Duplicated, not shared: the webhook package and + // internal/controllers/consistencygroup each declare their own copy of this same + // unexported constant, and this package follows that established precedent + // rather than introducing a shared import across packages for one string. + consistencyGroupLabel = "storage.simplyblock.io/consistency-group" + + reasonGroupReplicationVerified = "GroupReplicationVerified" + reasonGroupMembershipMismatch = "GroupMembershipMismatch" + reasonGroupReplicationDegraded = "GroupReplicationDegraded" +) + +var ( + volumeGroupReplicationGVK = schema.GroupVersionKind{ + Group: "replication.storage.openshift.io", Version: "v1alpha1", Kind: "VolumeGroupReplication", + } + volumeGroupReplicationClassGVK = schema.GroupVersionKind{ + Group: "replication.storage.openshift.io", Version: "v1alpha1", Kind: "VolumeGroupReplicationClass", + } + volumeReplicationGVK = schema.GroupVersionKind{ + Group: "replication.storage.openshift.io", Version: "v1alpha1", Kind: "VolumeReplication", + } +) + +// VolumeGroupReplicationReconciler fans a VolumeGroupReplication's group-level +// replication intent out to one per-volume VolumeReplication per consistency-group +// member, and fans the members' status back into the group's own +// (design-ramen-integration.md §4.3). It owns no gRPC call of its own: promoting, +// demoting, or resyncing a member is the already-shipped per-volume adapter's job, +// driven by the kubernetes-csi-addons controller-manager reconciling the member +// VolumeReplication objects this reconciler creates (§4.2). +type VolumeGroupReplicationReconciler struct { + client.Client + Scheme *runtime.Scheme + Recorder events.EventRecorder +} + +// +kubebuilder:rbac:groups=replication.storage.openshift.io,resources=volumegroupreplications,verbs=get;list;watch;update;patch +// +kubebuilder:rbac:groups=replication.storage.openshift.io,resources=volumegroupreplications/status,verbs=get;update;patch +// +kubebuilder:rbac:groups=replication.storage.openshift.io,resources=volumegroupreplicationclasses,verbs=get;list;watch +// +kubebuilder:rbac:groups=replication.storage.openshift.io,resources=volumereplications,verbs=get;list;watch;create;update;patch;delete +// +kubebuilder:rbac:groups="",resources=persistentvolumeclaims,verbs=get;list;watch +// +kubebuilder:rbac:groups="",resources=persistentvolumes,verbs=get;list;watch + +func (r *VolumeGroupReplicationReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) { + log := logf.FromContext(ctx) + + vgr := &unstructured.Unstructured{} + vgr.SetGroupVersionKind(volumeGroupReplicationGVK) + if err := r.Get(ctx, req.NamespacedName, vgr); err != nil { + return ctrl.Result{}, client.IgnoreNotFound(err) + } + + // A non-external VolumeGroupReplication is the generic kubernetes-csi-addons + // controller-manager's to reconcile, not this operator's (§4.2). + external, _, _ := unstructured.NestedBool(vgr.Object, "spec", "external") + if !external { + return ctrl.Result{}, nil + } + className, _, _ := unstructured.NestedString(vgr.Object, "spec", "volumeGroupReplicationClassName") + owned, err := r.ownsClass(ctx, className) + if err != nil { + log.Error(err, "failed to resolve VolumeGroupReplicationClass", "class", className) + return ctrl.Result{RequeueAfter: groupReplRequeueError}, nil + } + if !owned { + return ctrl.Result{}, nil + } + + members, determinable, mismatch, err := r.resolveMembers(ctx, vgr) + if err != nil { + log.Error(err, "failed to resolve group membership") + return ctrl.Result{RequeueAfter: groupReplRequeueError}, nil + } + if !determinable { + // A selected PVC is not yet bound, or its volume is not yet resolvable. + // Transient: requeue quietly, matching the webhook's fail-open disposition + // for the identical case (design §4.4). + return ctrl.Result{RequeueAfter: groupReplSyncInterval}, nil + } + if mismatch != "" { + r.Recorder.Eventf(vgr, nil, corev1.EventTypeWarning, reasonGroupMembershipMismatch, reasonGroupMembershipMismatch, mismatch) + return ctrl.Result{RequeueAfter: groupReplSyncInterval}, nil + } + r.Recorder.Eventf(vgr, nil, corev1.EventTypeNormal, reasonGroupReplicationVerified, reasonGroupReplicationVerified, + "selector resolves to the consistency group's current membership (%d member(s))", len(members)) + + replicationState, _, _ := unstructured.NestedString(vgr.Object, "spec", "replicationState") + volumeReplicationClass, _, _ := unstructured.NestedString(vgr.Object, "spec", "volumeReplicationClassName") + + memberVRs, err := r.fanOut(ctx, vgr, members, replicationState, volumeReplicationClass) + if err != nil { + log.Error(err, "failed to fan out member VolumeReplications") + return ctrl.Result{RequeueAfter: groupReplRequeueError}, nil + } + + if err := r.fanIn(ctx, vgr, members, memberVRs); err != nil { + log.Error(err, "failed to update VolumeGroupReplication status") + return ctrl.Result{RequeueAfter: groupReplRequeueError}, nil + } + + return ctrl.Result{RequeueAfter: groupReplSyncInterval}, nil +} + +// ownsClass reports whether className names a VolumeGroupReplicationClass for +// this driver. A class this operator cannot read is treated as not-owned: +// the generic controller-manager or another vendor's controller may still be +// the right owner, and refusing to reconcile is safer than guessing. +func (r *VolumeGroupReplicationReconciler) ownsClass(ctx context.Context, className string) (bool, error) { + if className == "" { + return false, nil + } + class := &unstructured.Unstructured{} + class.SetGroupVersionKind(volumeGroupReplicationClassGVK) + if err := r.Get(ctx, types.NamespacedName{Name: className}, class); err != nil { + if apierrors.IsNotFound(err) { + return false, nil + } + return false, err + } + provisioner, _, _ := unstructured.NestedString(class.Object, "spec", "provisioner") + return provisioner == utils.CSIProvisioner, nil +} + +// resolveMembers reads spec.source.selector, resolves it to this cluster's own +// PVCs, and verifies the selected set equals a consistency group's current +// membership exactly (design §4.3 step 1, the same invariant +// design-consistency-groups.md §9.2 established for VolumeGroupSnapshot). +// +// determinable is false when a selected PVC is not yet bound, matching the +// webhook's fail-open case for the identical situation. mismatch is non-empty +// when membership is determinable but does not match. +func (r *VolumeGroupReplicationReconciler) resolveMembers( + ctx context.Context, vgr *unstructured.Unstructured, +) (members []corev1.PersistentVolumeClaim, determinable bool, mismatch string, err error) { + selMap, found, err := unstructured.NestedMap(vgr.Object, "spec", "source", "selector") + if err != nil { + return nil, true, "", fmt.Errorf("read spec.source.selector: %w", err) + } + if !found { + return nil, true, "VolumeGroupReplication has no spec.source.selector", nil + } + var labelSelector metav1.LabelSelector + if err := runtime.DefaultUnstructuredConverter.FromUnstructured(selMap, &labelSelector); err != nil { + return nil, true, "", fmt.Errorf("convert spec.source.selector: %w", err) + } + sel, err := metav1.LabelSelectorAsSelector(&labelSelector) + if err != nil { + return nil, true, fmt.Sprintf("invalid label selector: %v", err), nil + } + + var pvcList corev1.PersistentVolumeClaimList + if err := r.List(ctx, &pvcList, + client.InNamespace(vgr.GetNamespace()), client.MatchingLabelsSelector{Selector: sel}); err != nil { + return nil, true, "", fmt.Errorf("list PVCs: %w", err) + } + if len(pvcList.Items) == 0 { + return nil, true, "selector matches no PersistentVolumeClaim", nil + } + + groupName := "" + for i := range pvcList.Items { + val := pvcList.Items[i].Labels[consistencyGroupLabel] + if val == "" { + return nil, true, fmt.Sprintf("PVC %q is not labeled %s", pvcList.Items[i].Name, consistencyGroupLabel), nil + } + if groupName == "" { + groupName = val + } else if val != groupName { + return nil, true, fmt.Sprintf("selector spans two consistency groups (%q and %q)", groupName, val), nil + } + } + + clusterUUID, selectedLvols, ok, err := r.selectedLvols(ctx, pvcList.Items) + if err != nil { + return nil, true, "", err + } + if !ok { + return nil, false, "", nil + } + + apiClient := webapi.NewClient() + group, err := apiClient.GetConsistencyGroupByName(ctx, clusterUUID, groupName) + if err != nil { + return nil, true, "", err + } + if group == nil { + return nil, true, fmt.Sprintf("consistency group %q not found", groupName), nil + } + backendMembers, err := apiClient.GetConsistencyGroupMembers(ctx, clusterUUID, group.UUID) + if err != nil { + return nil, true, "", err + } + if !sameLvolSet(selectedLvols, backendMembers) { + return nil, true, fmt.Sprintf( + "selector resolves to %d volume(s) but consistency group %q has %d member(s); "+ + "the selector must equal the group's current membership", + len(selectedLvols), groupName, len(backendMembers)), nil + } + return pvcList.Items, true, "", nil +} + +// selectedLvols maps each selected PVC to its backing lvol UUID and returns the +// shared cluster UUID. ok is false when a PVC is not yet bound, which makes +// membership undeterminable rather than mismatched. +func (r *VolumeGroupReplicationReconciler) selectedLvols( + ctx context.Context, pvcs []corev1.PersistentVolumeClaim, +) (clusterUUID string, lvols []string, ok bool, err error) { + for i := range pvcs { + pvName := pvcs[i].Spec.VolumeName + if pvName == "" { + return "", nil, false, nil + } + pv := &corev1.PersistentVolume{} + if err := r.Get(ctx, types.NamespacedName{Name: pvName}, pv); err != nil { + return "", nil, false, err + } + raw, err := kube.VolumeHandleFromPV(pv) + if err != nil { + return "", nil, false, nil + } + h, parsed := atlaslvol.ParseHandle(raw) + if !parsed { + return "", nil, false, nil + } + clusterUUID = h.ClusterID + lvols = append(lvols, h.VolumeID) + } + return clusterUUID, lvols, true, nil +} + +// sameLvolSet reports whether a and b hold the same set of lvol ids. +func sameLvolSet(a, b []string) bool { + if len(a) != len(b) { + return false + } + seen := make(map[string]bool, len(a)) + for _, s := range a { + seen[s] = true + } + for _, s := range b { + if !seen[s] { + return false + } + } + return true +} + +// fanOut ensures one per-volume VolumeReplication exists for each member PVC, +// owned by the group, with spec.replicationState mirroring the group's +// (design §4.3 step 2). It never calls the driver's Replication gRPC: the +// already-shipped kubernetes-csi-addons controller-manager reconciles each +// member exactly as it does any Ramen-created per-volume VolumeReplication. +func (r *VolumeGroupReplicationReconciler) fanOut( + ctx context.Context, + vgr *unstructured.Unstructured, + members []corev1.PersistentVolumeClaim, + replicationState, volumeReplicationClass string, +) ([]unstructured.Unstructured, error) { + owner := metav1.OwnerReference{ + APIVersion: volumeGroupReplicationGVK.GroupVersion().String(), + Kind: volumeGroupReplicationGVK.Kind, + Name: vgr.GetName(), + UID: vgr.GetUID(), + } + + result := make([]unstructured.Unstructured, 0, len(members)) + for i := range members { + pvc := &members[i] + name := memberVolumeReplicationName(vgr.GetName(), pvc.Name) + + vr := &unstructured.Unstructured{} + vr.SetGroupVersionKind(volumeReplicationGVK) + err := r.Get(ctx, types.NamespacedName{Namespace: vgr.GetNamespace(), Name: name}, vr) + switch { + case apierrors.IsNotFound(err): + vr.SetName(name) + vr.SetNamespace(vgr.GetNamespace()) + vr.SetOwnerReferences([]metav1.OwnerReference{owner}) + spec := map[string]interface{}{ + "autoResync": false, + "replicationState": replicationState, + "volumeReplicationClass": volumeReplicationClass, + "dataSource": map[string]interface{}{ + "kind": "PersistentVolumeClaim", + "name": pvc.Name, + }, + } + if err := unstructured.SetNestedMap(vr.Object, spec, "spec"); err != nil { + return nil, fmt.Errorf("build VolumeReplication %q spec: %w", name, err) + } + if err := r.Create(ctx, vr); err != nil { + return nil, fmt.Errorf("create VolumeReplication %q: %w", name, err) + } + case err != nil: + return nil, fmt.Errorf("get VolumeReplication %q: %w", name, err) + default: + if current, _, _ := unstructured.NestedString(vr.Object, "spec", "replicationState"); current != replicationState { + if err := unstructured.SetNestedField(vr.Object, replicationState, "spec", "replicationState"); err != nil { + return nil, fmt.Errorf("set VolumeReplication %q replicationState: %w", name, err) + } + if err := r.Update(ctx, vr); err != nil { + return nil, fmt.Errorf("update VolumeReplication %q: %w", name, err) + } + } + } + result = append(result, *vr) + } + return result, nil +} + +// memberVolumeReplicationName deterministically names a group member's +// per-volume VolumeReplication from the group and the member PVC. +func memberVolumeReplicationName(groupName, pvcName string) string { + return groupName + "-" + pvcName +} + +// fanIn aggregates every member's VolumeReplication.status.conditions into the +// group's own status (design §4.3 step 3): Completed is the conjunction across +// members, Degraded and Resyncing are the disjunction, and status.lastSyncTime +// is the oldest of the members' lastSyncTime, since a group's recovery point is +// only as fresh as its slowest member. It also emits GroupReplicationDegraded +// the first time it observes Degraded after the group previously reported it +// false, and writes status.persistentVolumeClaimsRefList to the resolved +// membership. +func (r *VolumeGroupReplicationReconciler) fanIn( + ctx context.Context, + vgr *unstructured.Unstructured, + members []corev1.PersistentVolumeClaim, + memberVRs []unstructured.Unstructured, +) error { + wasDegraded := conditionStatus(vgr, "Degraded") == "True" + + completed := len(memberVRs) > 0 + degraded := false + resyncing := false + var oldestSync *time.Time + for i := range memberVRs { + vr := &memberVRs[i] + if conditionStatus(vr, "Completed") != "True" { + completed = false + } + if conditionStatus(vr, "Degraded") == "True" { + degraded = true + } + if conditionStatus(vr, "Resyncing") == "True" { + resyncing = true + } + if ts, found, _ := unstructured.NestedString(vr.Object, "status", "lastSyncTime"); found && ts != "" { + if parsed, err := time.Parse(time.RFC3339, ts); err == nil { + if oldestSync == nil || parsed.Before(*oldestSync) { + oldestSync = &parsed + } + } + } + } + + now := metav1.Now() + conditions := []interface{}{ + groupCondition("Completed", completed, now), + groupCondition("Degraded", degraded, now), + groupCondition("Resyncing", resyncing, now), + } + if err := unstructured.SetNestedSlice(vgr.Object, conditions, "status", "conditions"); err != nil { + return fmt.Errorf("set status.conditions: %w", err) + } + if oldestSync != nil { + if err := unstructured.SetNestedField(vgr.Object, oldestSync.UTC().Format(time.RFC3339), "status", "lastSyncTime"); err != nil { + return fmt.Errorf("set status.lastSyncTime: %w", err) + } + } + if completed { + replicationState, _, _ := unstructured.NestedString(vgr.Object, "spec", "replicationState") + if err := unstructured.SetNestedField(vgr.Object, replicationState, "status", "state"); err != nil { + return fmt.Errorf("set status.state: %w", err) + } + } + refs := make([]interface{}, 0, len(members)) + for i := range members { + refs = append(refs, map[string]interface{}{"name": members[i].Name}) + } + if err := unstructured.SetNestedSlice(vgr.Object, refs, "status", "persistentVolumeClaimsRefList"); err != nil { + return fmt.Errorf("set status.persistentVolumeClaimsRefList: %w", err) + } + + if err := r.Status().Update(ctx, vgr); err != nil { + return err + } + if degraded && !wasDegraded { + r.Recorder.Eventf(vgr, nil, corev1.EventTypeWarning, reasonGroupReplicationDegraded, reasonGroupReplicationDegraded, + "at least one group member reports Degraded") + } + return nil +} + +// conditionStatus returns a condition's status ("True"/"False"/"Unknown") on +// an unstructured VolumeReplication or VolumeGroupReplication, or "" if the +// object carries no condition of that type. +func conditionStatus(obj *unstructured.Unstructured, condType string) string { + raw, found, _ := unstructured.NestedSlice(obj.Object, "status", "conditions") + if !found { + return "" + } + for _, c := range raw { + cm, ok := c.(map[string]interface{}) + if !ok { + continue + } + if cm["type"] == condType { + if s, ok := cm["status"].(string); ok { + return s + } + } + } + return "" +} + +// groupCondition builds one status.conditions entry. +func groupCondition(condType string, status bool, now metav1.Time) map[string]interface{} { + s := "False" + if status { + s = "True" + } + return map[string]interface{}{ + "type": condType, + "status": s, + "reason": "GroupMemberAggregation", + "message": "", + "lastTransitionTime": now.UTC().Format(time.RFC3339), + } +} + +func (r *VolumeGroupReplicationReconciler) SetupWithManager(mgr ctrl.Manager) error { + target := &unstructured.Unstructured{} + target.SetGroupVersionKind(volumeGroupReplicationGVK) + + return ctrl.NewControllerManagedBy(mgr). + For(target). + Named("volumegroupreplication"). + Complete(r) +} diff --git a/operator/internal/controller/volumegroupreplication_controller_unit_test.go b/operator/internal/controller/volumegroupreplication_controller_unit_test.go new file mode 100644 index 000000000..99b875760 --- /dev/null +++ b/operator/internal/controller/volumegroupreplication_controller_unit_test.go @@ -0,0 +1,294 @@ +/* +Copyright 2025. + +Licensed under the Apache License, Version 2.0 (the "License"); +you may not use this file except in compliance with the License. +You may obtain a copy of the License at + + http://www.apache.org/licenses/LICENSE-2.0 + +Unless required by applicable law or agreed to in writing, software +distributed under the License is distributed on an "AS IS" BASIS, +WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +See the License for the specific language governing permissions and +limitations under the License. +*/ + +package controller + +import ( + "context" + "crypto/sha256" + "encoding/hex" + "encoding/json" + "net/http" + "strings" + "testing" + "time" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" + "k8s.io/apimachinery/pkg/types" + "k8s.io/client-go/tools/events" + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" +) + +// The scenario matrix this file implements is +// docs/tests/test-plan-ramen-integration.md §1's U-01…U-03, verifying +// design-ramen-integration.md §4.3's fan-in aggregation. + +const vgrCluster = "55555555-5555-5555-5555-555555555555" + +// vgrUUID maps a short fixture id to a deterministic canonical UUID, the same +// way replication_test.go's sibling in csi-driver and the webhook package's +// vgsUUID both do, since atlas/lvol.ParseHandle accepts only canonical UUIDs. +func vgrUUID(id string) string { + sum := sha256.Sum256([]byte(id)) + h := hex.EncodeToString(sum[:16]) + return h[0:8] + "-" + h[8:12] + "-" + h[12:16] + "-" + h[16:20] + "-" + h[20:32] +} + +// vgrBackend serves the consistency-group resolve + membership reads. +type vgrBackend struct { + members []string // short fixture ids, converted through vgrUUID +} + +func (b vgrBackend) start(t *testing.T) { + t.Helper() + mux := http.NewServeMux() + mux.HandleFunc("/api/v2/clusters/"+vgrCluster+"/consistency-groups/", func(w http.ResponseWriter, r *http.Request) { + if strings.HasSuffix(r.URL.Path, "/members") { + rows := make([]map[string]any, 0, len(b.members)) + for _, m := range b.members { + rows = append(rows, map[string]any{"lvol_id": vgrUUID(m)}) + } + _ = json.NewEncoder(w).Encode(rows) + return + } + _ = json.NewEncoder(w).Encode([]map[string]any{ + {"id": "grp", "name": "grp1", "member_count": len(b.members)}, + }) + }) + srv := newAPIServer(t, mux.ServeHTTP) + t.Setenv("SIMPLYBLOCK_WEBAPI_BASE_URL", srv.URL) +} + +// vgrMember describes one group member: a bound PVC plus the per-volume +// VolumeReplication a prior fan-out already created for it, pre-seeded with +// the status this test wants fan-in to read. +type vgrMember struct { + pvcName string + volumeID string // short fixture id, converted through vgrUUID + completed bool + degraded bool + lastSync time.Time +} + +func newVGRReconciler(t *testing.T, objects ...client.Object) (*VolumeGroupReplicationReconciler, client.Client) { + t.Helper() + scheme := newTestScheme(t, corev1.AddToScheme) + statusSubresource := &unstructured.Unstructured{} + statusSubresource.SetGroupVersionKind(volumeGroupReplicationGVK) + cl := fake.NewClientBuilder(). + WithScheme(scheme). + WithStatusSubresource(statusSubresource). + WithObjects(objects...). + Build() + return &VolumeGroupReplicationReconciler{ + Client: cl, + Scheme: scheme, + Recorder: events.NewFakeRecorder(32), + }, cl +} + +func vgrClass(name string) *unstructured.Unstructured { + class := &unstructured.Unstructured{} + class.SetGroupVersionKind(volumeGroupReplicationClassGVK) + class.SetName(name) + _ = unstructured.SetNestedField(class.Object, "csi.simplyblock.io", "spec", "provisioner") + return class +} + +// vgrGroup builds the VolumeGroupReplication under test, selecting every +// member by the shared "app: grp1" label. +func vgrGroup(name, className, vrClassName, replicationState string) *unstructured.Unstructured { + vgr := &unstructured.Unstructured{} + vgr.SetGroupVersionKind(volumeGroupReplicationGVK) + vgr.SetName(name) + vgr.SetNamespace("default") + vgr.SetUID(types.UID("vgr-uid-" + name)) + spec := map[string]interface{}{ + "external": true, + "autoResync": false, + "replicationState": replicationState, + "volumeGroupReplicationClassName": className, + "volumeReplicationClassName": vrClassName, + "source": map[string]interface{}{ + "selector": map[string]interface{}{ + "matchLabels": map[string]interface{}{"app": "grp1"}, + }, + }, + } + _ = unstructured.SetNestedMap(vgr.Object, spec, "spec") + return vgr +} + +// vgrPVCAndPV builds a bound PVC/PV pair: labeled for both the group +// selector and the consistency-group membership check, backed by a +// simplyblock CSI volume handle. +func vgrPVCAndPV(pvcName, volumeID string) (*corev1.PersistentVolumeClaim, *corev1.PersistentVolume) { + pvName := "pv-" + pvcName + pvc := &corev1.PersistentVolumeClaim{ + ObjectMeta: metav1.ObjectMeta{ + Name: pvcName, + Namespace: "default", + Labels: map[string]string{ + "app": "grp1", + consistencyGroupLabel: "grp1", + }, + }, + Spec: corev1.PersistentVolumeClaimSpec{VolumeName: pvName}, + } + pv := &corev1.PersistentVolume{ + ObjectMeta: metav1.ObjectMeta{Name: pvName}, + Spec: corev1.PersistentVolumeSpec{PersistentVolumeSource: corev1.PersistentVolumeSource{ + CSI: &corev1.CSIPersistentVolumeSource{ + Driver: "csi.simplyblock.io", + VolumeHandle: vgrCluster + ":pool:" + vgrUUID(volumeID), + }, + }}, + } + return pvc, pv +} + +// vgrMemberVR builds the per-volume VolumeReplication a prior fan-out +// already created for one group member, with the status this test wants +// fan-in to aggregate. +func vgrMemberVR(name, groupName, pvcName string, m vgrMember) *unstructured.Unstructured { + vr := &unstructured.Unstructured{} + vr.SetGroupVersionKind(volumeReplicationGVK) + vr.SetName(name) + vr.SetNamespace("default") + vr.SetOwnerReferences([]metav1.OwnerReference{{ + APIVersion: volumeGroupReplicationGVK.GroupVersion().String(), + Kind: volumeGroupReplicationGVK.Kind, + Name: groupName, + UID: types.UID("vgr-uid-" + groupName), + }}) + _ = unstructured.SetNestedField(vr.Object, pvcName, "spec", "dataSource", "name") + _ = unstructured.SetNestedField(vr.Object, "PersistentVolumeClaim", "spec", "dataSource", "kind") + _ = unstructured.SetNestedField(vr.Object, "primary", "spec", "replicationState") + + cond := func(condType string, status bool) map[string]interface{} { + s := "False" + if status { + s = "True" + } + return map[string]interface{}{"type": condType, "status": s, "reason": "Test", "message": ""} + } + conditions := []interface{}{ + cond("Completed", m.completed), + cond("Degraded", m.degraded), + cond("Resyncing", false), + } + _ = unstructured.SetNestedSlice(vr.Object, conditions, "status", "conditions") + _ = unstructured.SetNestedField(vr.Object, m.lastSync.UTC().Format(time.RFC3339), "status", "lastSyncTime") + return vr +} + +// runVGRReconcile builds the group, its class, the member PVC/PV pairs, and +// their pre-seeded VolumeReplications, reconciles once, and returns the +// group's own status.conditions and status.lastSyncTime for assertion. +func runVGRReconcile(t *testing.T, members []vgrMember) (conditions map[string]string, lastSyncTime string) { + t.Helper() + vgrBackend{members: memberIDs(members)}.start(t) + + group := vgrGroup("vgr1", "sb-group-class", "sb-vr-class", "primary") + objs := []client.Object{group, vgrClass("sb-group-class")} + var memberIDsOnly []string + for _, m := range members { + pvc, pv := vgrPVCAndPV(m.pvcName, m.volumeID) + objs = append(objs, pvc, pv) + memberIDsOnly = append(memberIDsOnly, m.volumeID) + } + r, cl := newVGRReconciler(t, objs...) + for _, m := range members { + vr := vgrMemberVR("vgr1-"+m.pvcName, "vgr1", m.pvcName, m) + if err := cl.Create(context.Background(), vr); err != nil { + t.Fatalf("create member VolumeReplication: %v", err) + } + } + + req := ctrl.Request{NamespacedName: types.NamespacedName{Namespace: "default", Name: "vgr1"}} + if _, err := r.Reconcile(context.Background(), req); err != nil { + t.Fatalf("unexpected error: %v", err) + } + + got := &unstructured.Unstructured{} + got.SetGroupVersionKind(volumeGroupReplicationGVK) + if err := cl.Get(context.Background(), types.NamespacedName{Namespace: "default", Name: "vgr1"}, got); err != nil { + t.Fatalf("get VolumeGroupReplication: %v", err) + } + + conditions = map[string]string{} + rawConds, _, _ := unstructured.NestedSlice(got.Object, "status", "conditions") + for _, c := range rawConds { + cm, ok := c.(map[string]interface{}) + if !ok { + continue + } + conditions[cm["type"].(string)] = cm["status"].(string) + } + lastSyncTime, _, _ = unstructured.NestedString(got.Object, "status", "lastSyncTime") + return conditions, lastSyncTime +} + +func memberIDs(members []vgrMember) []string { + var ids []string + for _, m := range members { + ids = append(ids, m.volumeID) + } + return ids +} + +// U-01: every member Completed/not-Degraded yields a group reporting the same. +func TestVolumeGroupReplication_AllMembersHealthyYieldsGroupHealthy(t *testing.T) { + conditions, _ := runVGRReconcile(t, []vgrMember{ + {pvcName: "pvc-a", volumeID: "v1", completed: true, degraded: false, lastSync: time.Now()}, + {pvcName: "pvc-b", volumeID: "v2", completed: true, degraded: false, lastSync: time.Now()}, + }) + if conditions["Completed"] != "True" { + t.Errorf("Completed = %q, want True", conditions["Completed"]) + } + if conditions["Degraded"] != "False" { + t.Errorf("Degraded = %q, want False", conditions["Degraded"]) + } +} + +// U-02: one member Degraded yields a group reporting Degraded. +func TestVolumeGroupReplication_OneMemberDegradedYieldsGroupDegraded(t *testing.T) { + conditions, _ := runVGRReconcile(t, []vgrMember{ + {pvcName: "pvc-a", volumeID: "v1", completed: true, degraded: false, lastSync: time.Now()}, + {pvcName: "pvc-b", volumeID: "v2", completed: true, degraded: true, lastSync: time.Now()}, + }) + if conditions["Degraded"] != "True" { + t.Errorf("Degraded = %q, want True", conditions["Degraded"]) + } +} + +// U-03: members with differing lastSyncTime yield the oldest, not the newest. +func TestVolumeGroupReplication_LastSyncTimeIsTheOldestMember(t *testing.T) { + older := time.Now().Add(-1 * time.Hour).Truncate(time.Second) + newer := time.Now().Truncate(time.Second) + _, lastSyncTime := runVGRReconcile(t, []vgrMember{ + {pvcName: "pvc-a", volumeID: "v1", completed: true, degraded: false, lastSync: older}, + {pvcName: "pvc-b", volumeID: "v2", completed: true, degraded: false, lastSync: newer}, + }) + want := older.UTC().Format(time.RFC3339) + if lastSyncTime != want { + t.Errorf("status.lastSyncTime = %q, want the oldest member's %q", lastSyncTime, want) + } +} diff --git a/operator/internal/webhook/volumegroupreplication_validator.go b/operator/internal/webhook/volumegroupreplication_validator.go new file mode 100644 index 000000000..fb3a8a99c --- /dev/null +++ b/operator/internal/webhook/volumegroupreplication_validator.go @@ -0,0 +1,201 @@ +package webhook + +import ( + "context" + "encoding/json" + "fmt" + "net/http" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" + "k8s.io/apimachinery/pkg/runtime" + "k8s.io/apimachinery/pkg/runtime/schema" + "sigs.k8s.io/controller-runtime/pkg/client" + logf "sigs.k8s.io/controller-runtime/pkg/log" + "sigs.k8s.io/controller-runtime/pkg/webhook/admission" + + "github.com/simplyblock/simplyblock-operator/internal/webapi" +) + +// volumeGroupReplicationGVK and volumeGroupReplicationClassGVK name the +// csi-addons kinds this repository vendors no Go type for (design +// §4.1). Duplicated, not shared, with internal/controller's identical pair: +// each package that needs an unstructured kind's GVK declares its own, +// matching the layering the design's own §4.4/§4.3 keep as two independent +// enforcement points (design-consistency-groups.md §9's "last line of +// defense, not the first"). +var ( + volumeGroupReplicationGVK = schema.GroupVersionKind{ + Group: "replication.storage.openshift.io", Version: "v1alpha1", Kind: "VolumeGroupReplication", + } + volumeGroupReplicationClassGVK = schema.GroupVersionKind{ + Group: "replication.storage.openshift.io", Version: "v1alpha1", Kind: "VolumeGroupReplicationClass", + } +) + +// +kubebuilder:webhook:path=/validate-replication-storage-openshift-io-v1alpha1-volumegroupreplication,mutating=false,failurePolicy=fail,sideEffects=None,groups=replication.storage.openshift.io,resources=volumegroupreplications,verbs=create,versions=v1alpha1,name=vvolumegroupreplication.simplyblock.io,admissionReviewVersions=v1 + +// VolumeGroupReplicationValidator rejects, at kubectl apply, a +// VolumeGroupReplication whose selector does not resolve to exactly one +// consistency group's current membership, so VolumeGroupReplicationReconciler +// (design §4.3) never has to reconcile an object that could never fan out +// correctly. It makes the same two checks, with the same two dispositions, +// design-consistency-groups.md §9.4 already established for +// VolumeGroupSnapshot: +// +// - The label check is static and fail-closed: every selected PVC must carry +// the same non-empty consistency-group label. A selector that spans groups, +// matches an unlabeled PVC, or matches nothing is rejected outright. +// - The membership check is backend and fail-open: it resolves the group by +// its label value and rejects the object unless the selected set equals the +// group's current membership. When membership cannot be determined (the +// backend is unreachable, or a selected PVC is not yet bound), it admits and +// the reconciler backstops it on its own next reconcile (design §4.3). +type VolumeGroupReplicationValidator struct { + Client client.Client + APIClient *webapi.Client +} + +// ownsVolumeGroupReplication reports whether the object is external and +// attributed to this driver's class, the same ownership signal the +// reconciler applies (design §4.3). Every undecidable case admits, because +// with failurePolicy=fail a webhook error would block every +// VolumeGroupReplication in the cluster, foreign drivers included. +func (v *VolumeGroupReplicationValidator) ownsVolumeGroupReplication( + ctx context.Context, vgr *unstructured.Unstructured, +) (bool, string) { + external, _, _ := unstructured.NestedBool(vgr.Object, "spec", "external") + if !external { + return false, "spec.external is not true; the generic controller-manager reconciles this one" + } + className, _, _ := unstructured.NestedString(vgr.Object, "spec", "volumeGroupReplicationClassName") + if className == "" { + return false, "no volumeGroupReplicationClassName; not attributable to this driver" + } + class := &unstructured.Unstructured{} + class.SetGroupVersionKind(volumeGroupReplicationClassGVK) + if err := v.Client.Get(ctx, client.ObjectKey{Name: className}, class); err != nil { + return false, fmt.Sprintf("volume group replication class %q not readable; not validating", className) + } + provisioner, _, _ := unstructured.NestedString(class.Object, "spec", "provisioner") + if provisioner != "csi.simplyblock.io" { + return false, fmt.Sprintf("class %q belongs to provisioner %q; not validating", className, provisioner) + } + return true, "" +} + +func (v *VolumeGroupReplicationValidator) Handle(ctx context.Context, req admission.Request) admission.Response { + log := logf.FromContext(ctx).WithValues("volumegroupreplication", req.Name, "namespace", req.Namespace) + + vgr := &unstructured.Unstructured{} + if err := json.Unmarshal(req.Object.Raw, vgr); err != nil { + return admission.Errored(http.StatusBadRequest, err) + } + + if ours, reason := v.ownsVolumeGroupReplication(ctx, vgr); !ours { + return admission.Allowed(reason) + } + + pvcs, groupName, denied := v.labelCheck(ctx, req.Namespace, vgr) + if denied != nil { + return *denied + } + + clusterUUID, selected, determinable, err := v.selectedLvols(ctx, pvcs) + if err != nil { + log.Error(err, "cannot resolve selected volumes; admitting (reconciler backstops)") + return admission.Allowed("group membership undeterminable; deferring to the reconciler") + } + if !determinable { + return admission.Allowed("a selected PVC is not yet bound; deferring to the reconciler") + } + group, err := v.APIClient.GetConsistencyGroupByName(ctx, clusterUUID, groupName) + if err != nil || group == nil { + if err != nil { + log.Error(err, "cannot resolve consistency group; admitting (reconciler backstops)") + } + return admission.Allowed("consistency group not resolvable; deferring to the reconciler") + } + members, err := v.APIClient.GetConsistencyGroupMembers(ctx, clusterUUID, group.UUID) + if err != nil { + log.Error(err, "cannot read group membership; admitting (reconciler backstops)") + return admission.Allowed("group membership unreadable; deferring to the reconciler") + } + if !sameStringSet(selected, members) { + return admission.Denied(fmt.Sprintf( + "selector resolves to %d volume(s) but consistency group %q has %d member(s); "+ + "the selector must equal the group's current membership", + len(selected), groupName, len(members))) + } + return admission.Allowed("selector equals the consistency group's membership") +} + +// labelCheck resolves spec.source.selector to PVCs and enforces the +// fail-closed label rule, the same as VolumeGroupSnapshotValidator.labelCheck. +func (v *VolumeGroupReplicationValidator) labelCheck( + ctx context.Context, namespace string, vgr *unstructured.Unstructured, +) ([]corev1.PersistentVolumeClaim, string, *admission.Response) { + deny := func(msg string) *admission.Response { r := admission.Denied(msg); return &r } + + selMap, found, err := unstructured.NestedMap(vgr.Object, "spec", "source", "selector") + if err != nil || !found { + return nil, "", deny("VolumeGroupReplication has no spec.source.selector") + } + var labelSelector metav1.LabelSelector + if err := runtime.DefaultUnstructuredConverter.FromUnstructured(selMap, &labelSelector); err != nil { + return nil, "", deny(fmt.Sprintf("invalid label selector: %v", err)) + } + sel, err := metav1.LabelSelectorAsSelector(&labelSelector) + if err != nil { + return nil, "", deny(fmt.Sprintf("invalid label selector: %v", err)) + } + var pvcList corev1.PersistentVolumeClaimList + if err := v.Client.List(ctx, &pvcList, + client.InNamespace(namespace), client.MatchingLabelsSelector{Selector: sel}); err != nil { + return nil, "", deny(fmt.Sprintf("cannot list PVCs for the selector: %v", err)) + } + if len(pvcList.Items) == 0 { + return nil, "", deny("selector matches no PersistentVolumeClaim") + } + groupName := "" + for i := range pvcList.Items { + val := pvcList.Items[i].Labels[consistencyGroupLabel] + if val == "" { + return nil, "", deny(fmt.Sprintf( + "PVC %q is not labeled %s; a group replication's selector must match only group members", + pvcList.Items[i].Name, consistencyGroupLabel)) + } + if groupName == "" { + groupName = val + } else if val != groupName { + return nil, "", deny(fmt.Sprintf( + "selector spans two consistency groups (%q and %q); it must resolve to exactly one", + groupName, val)) + } + } + return pvcList.Items, groupName, nil +} + +// selectedLvols maps each selected PVC to its backing lvol UUID, the same as +// VolumeGroupSnapshotValidator.selectedLvols. +func (v *VolumeGroupReplicationValidator) selectedLvols( + ctx context.Context, pvcs []corev1.PersistentVolumeClaim, +) (clusterUUID string, lvols []string, determinable bool, err error) { + for i := range pvcs { + pvName := pvcs[i].Spec.VolumeName + if pvName == "" { + return "", nil, false, nil + } + cluster, _, volume, ok, err := pvVolumeHandle(ctx, v.Client, pvName) + if err != nil { + return "", nil, false, err + } + if !ok { + return "", nil, false, nil + } + clusterUUID = cluster + lvols = append(lvols, volume) + } + return clusterUUID, lvols, true, nil +} diff --git a/operator/internal/webhook/volumegroupreplication_validator_test.go b/operator/internal/webhook/volumegroupreplication_validator_test.go new file mode 100644 index 000000000..0ebef8e0c --- /dev/null +++ b/operator/internal/webhook/volumegroupreplication_validator_test.go @@ -0,0 +1,193 @@ +package webhook + +import ( + "context" + "crypto/sha256" + "encoding/hex" + "encoding/json" + "net/http" + "net/http/httptest" + "strings" + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" + "k8s.io/apimachinery/pkg/runtime" + crclient "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + "sigs.k8s.io/controller-runtime/pkg/webhook/admission" + + "github.com/simplyblock/simplyblock-operator/internal/webapi" +) + +// The scenario matrix this file implements is +// docs/tests/test-plan-ramen-integration.md §1's U-04…U-05, verifying +// design-ramen-integration.md §4.4's admission webhook. + +const vgrValCluster = "66666666-6666-6666-6666-666666666666" + +func vgrValUUID(id string) string { + sum := sha256.Sum256([]byte(id)) + h := hex.EncodeToString(sum[:16]) + return h[0:8] + "-" + h[8:12] + "-" + h[12:16] + "-" + h[16:20] + "-" + h[20:32] +} + +type vgrValPVC struct { + name string + cgLabel string + volumeID string + bound bool +} + +type vgrValBackend struct { + members []string + fail bool +} + +func (b vgrValBackend) server(t *testing.T) string { + t.Helper() + mux := http.NewServeMux() + mux.HandleFunc("/api/v2/clusters/"+vgrValCluster+"/consistency-groups/", func(w http.ResponseWriter, r *http.Request) { + if b.fail { + w.WriteHeader(http.StatusInternalServerError) + return + } + if strings.HasSuffix(r.URL.Path, "/members") { + rows := make([]map[string]any, 0, len(b.members)) + for _, m := range b.members { + rows = append(rows, map[string]any{"lvol_id": vgrValUUID(m)}) + } + _ = json.NewEncoder(w).Encode(rows) + return + } + _ = json.NewEncoder(w).Encode([]map[string]any{ + {"id": "grp", "name": "grp1", "member_count": len(b.members)}, + }) + }) + srv := httptest.NewServer(mux) + t.Cleanup(srv.Close) + return srv.URL +} + +func newVGRValidator(t *testing.T, pvcs []vgrValPVC, apiURL string) *VolumeGroupReplicationValidator { + t.Helper() + scheme := runtime.NewScheme() + if err := corev1.AddToScheme(scheme); err != nil { + t.Fatalf("add corev1: %v", err) + } + var objs []crclient.Object + for _, p := range pvcs { + labels := map[string]string{"app": "grp1"} + if p.cgLabel != "" { + labels[consistencyGroupLabel] = p.cgLabel + } + pvc := &corev1.PersistentVolumeClaim{ + ObjectMeta: metav1.ObjectMeta{Name: p.name, Namespace: "sb", Labels: labels}, + } + if p.bound { + pvName := "pv-" + p.name + pvc.Spec.VolumeName = pvName + objs = append(objs, &corev1.PersistentVolume{ + ObjectMeta: metav1.ObjectMeta{Name: pvName}, + Spec: corev1.PersistentVolumeSpec{PersistentVolumeSource: corev1.PersistentVolumeSource{ + CSI: &corev1.CSIPersistentVolumeSource{ + Driver: "csi.simplyblock.io", + VolumeHandle: vgrValCluster + ":pool:" + vgrValUUID(p.volumeID), + }, + }}, + }) + } + objs = append(objs, pvc) + } + class := &unstructured.Unstructured{} + class.SetGroupVersionKind(volumeGroupReplicationClassGVK) + class.SetName("sb-group-class") + _ = unstructured.SetNestedField(class.Object, "csi.simplyblock.io", "spec", "provisioner") + objs = append(objs, class) + cl := fake.NewClientBuilder().WithScheme(scheme).WithObjects(objs...).Build() + return &VolumeGroupReplicationValidator{Client: cl, APIClient: webapi.NewClient(apiURL)} +} + +func vgrValRequest(t *testing.T, selectorValue string) admission.Request { + t.Helper() + vgr := &unstructured.Unstructured{} + vgr.SetGroupVersionKind(volumeGroupReplicationGVK) + vgr.SetName("gen") + vgr.SetNamespace("sb") + _ = unstructured.SetNestedMap(vgr.Object, map[string]interface{}{ + "external": true, + "autoResync": false, + "replicationState": "primary", + "volumeGroupReplicationClassName": "sb-group-class", + "source": map[string]interface{}{ + "selector": map[string]interface{}{ + "matchLabels": map[string]interface{}{"app": selectorValue}, + }, + }, + }, "spec") + raw, err := json.Marshal(vgr.Object) + if err != nil { + t.Fatalf("marshal VolumeGroupReplication: %v", err) + } + req := admission.Request{} + req.Object = runtime.RawExtension{Raw: raw} + return req +} + +func TestVolumeGroupReplicationValidator(t *testing.T) { + tests := []struct { + name string + pvcs []vgrValPVC + selector string + backend vgrValBackend + allowed bool + }{ + { + name: "selector equals membership is admitted", + pvcs: []vgrValPVC{ + {name: "a", cgLabel: "grp1", volumeID: "v1", bound: true}, + {name: "b", cgLabel: "grp1", volumeID: "v2", bound: true}, + }, + selector: "grp1", + backend: vgrValBackend{members: []string{"v1", "v2"}}, + allowed: true, + }, + { + name: "selector resolving to a subset of the group is rejected", + pvcs: []vgrValPVC{ + {name: "a", cgLabel: "grp1", volumeID: "v1", bound: true}, + }, + selector: "grp1", + backend: vgrValBackend{members: []string{"v1", "v2"}}, + allowed: false, + }, + { + name: "selector spanning two groups is rejected", + pvcs: []vgrValPVC{ + {name: "a", cgLabel: "grp1", volumeID: "v1", bound: true}, + {name: "b", cgLabel: "grp2", volumeID: "v2", bound: true}, + }, + selector: "grp1", + allowed: false, + }, + { + name: "backend unreachable admits (fail-open)", + pvcs: []vgrValPVC{ + {name: "a", cgLabel: "grp1", volumeID: "v1", bound: true}, + }, + selector: "grp1", + backend: vgrValBackend{fail: true}, + allowed: true, + }, + } + for _, tc := range tests { + t.Run(tc.name, func(t *testing.T) { + v := newVGRValidator(t, tc.pvcs, tc.backend.server(t)) + resp := v.Handle(context.Background(), vgrValRequest(t, tc.selector)) + if resp.Allowed != tc.allowed { + t.Fatalf("Allowed = %v, want %v (msg: %s)", resp.Allowed, tc.allowed, resp.Result.Message) + } + }) + } +} From c40f4189ba44d88fb5e17d1db00231af125dc411 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Mon, 21 Sep 2026 17:10:14 +0100 Subject: [PATCH 124/206] modified operator/dist/install.yaml --- operator/dist/install.yaml | 57 ++++++++++++++++++++++++++++++++++++++ 1 file changed, 57 insertions(+) diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index d5623f0ca..fa5521ff1 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -8910,6 +8910,44 @@ rules: - patch - update - watch +- apiGroups: + - replication.storage.openshift.io + resources: + - volumegroupreplicationclasses + verbs: + - get + - list + - watch +- apiGroups: + - replication.storage.openshift.io + resources: + - volumegroupreplications + verbs: + - get + - list + - patch + - update + - watch +- apiGroups: + - replication.storage.openshift.io + resources: + - volumegroupreplications/status + verbs: + - get + - patch + - update +- apiGroups: + - replication.storage.openshift.io + resources: + - volumereplications + verbs: + - create + - delete + - get + - list + - patch + - update + - watch - apiGroups: - snapshot.storage.k8s.io resources: @@ -9819,6 +9857,25 @@ webhooks: resources: - storagepools sideEffects: None +- admissionReviewVersions: + - v1 + clientConfig: + service: + name: simplyblock-operator-webhook-service + namespace: simplyblock-operator-system + path: /validate-replication-storage-openshift-io-v1alpha1-volumegroupreplication + failurePolicy: Fail + name: vvolumegroupreplication.simplyblock.io + rules: + - apiGroups: + - replication.storage.openshift.io + apiVersions: + - v1alpha1 + operations: + - CREATE + resources: + - volumegroupreplications + sideEffects: None - admissionReviewVersions: - v1 clientConfig: From f49a3263753fbe4f60e2caa6ce614c3980a78913 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Mon, 21 Sep 2026 19:02:02 +0200 Subject: [PATCH 125/206] fix(operator): a pushed node change wakes the node controller again The storage-node subscription was still feeding the controller and no longer waking it. Steady state reads a node's status out of the stream's cache rather than asking the control plane, which is half of what the informer is for, and the other half -- reconciling when the control plane says something changed -- went out with the move to v1alpha2. The source was there before it (storagenode_controller.go:1201 at c98351b7^) and not after. Nothing failed. NodeSubscription went on publishing triggers into a channel with no reader, the cache went on being correct, and every node status change waited out nodeRetry to be noticed: thirty seconds, on a path whose whole purpose was not to wait. Triggers is on the NodeCache interface now rather than left to a type assertion at the call site, so a cache this package accepts is one that can wake it, and a fake that cannot is a compile error rather than a controller that never fires. The test asserts the source is built, not that an event arrives. A channel source only delivers under a running manager and this package has no envtest suite to start one in, so the case would not catch somebody deleting the WatchesRawSource call while leaving the helper. What it catches is the wiring being dropped, which is how it was lost, and the test says so rather than implying more. The backup-policy subscription's triggers are also unconsumed. That one was never wired and its controller reads policies through ListBackupPolicies rather than the cache, so it is a design question about which source that controller should take, not a regression. Left alone. Test plan: U-413, U-414. Co-Authored-By: Claude Opus 5 (1M context) --- operator/docs/tests/test-plan-storagenode.md | 2 + .../controllers/node/peertargets_test.go | 6 ++ .../controllers/node/pushednodes_test.go | 56 +++++++++++++++++++ .../node/storagenode_controller.go | 31 +++++++++- .../node/storagenodeops_controller.go | 7 +++ 5 files changed, 99 insertions(+), 3 deletions(-) create mode 100644 operator/internal/controllers/node/pushednodes_test.go diff --git a/operator/docs/tests/test-plan-storagenode.md b/operator/docs/tests/test-plan-storagenode.md index 33c26f5b3..f084bd292 100644 --- a/operator/docs/tests/test-plan-storagenode.md +++ b/operator/docs/tests/test-plan-storagenode.md @@ -115,6 +115,8 @@ Files: `operator/internal/controllers/node/provisioning_test.go`, | U-410 | One slice per RPC port | Positive | `TestEachRPCPortGetsItsOwnSlice` | | U-411 | A pod with no address is left unpublished rather than published with none | Negative | `TestAPodWithNoAddressIsNotPublished` | | U-412 | Two workers sharing a first DNS label are refused, not merged | Negative | `TestACollidingWorkerNameIsRefused` | +| U-413 | A pushed control-plane change wakes the node controller | Regression | `TestAPushedNodeChangeWakesTheController` | +| U-414 | A deployment with no informer is given no stream source | Negative | `TestNoInformerIsNoSource` | ### Entity: The Provisioning Claim (design §4.2) diff --git a/operator/internal/controllers/node/peertargets_test.go b/operator/internal/controllers/node/peertargets_test.go index 548f9dd81..6699f3548 100644 --- a/operator/internal/controllers/node/peertargets_test.go +++ b/operator/internal/controllers/node/peertargets_test.go @@ -20,6 +20,8 @@ import ( "errors" "testing" + "sigs.k8s.io/controller-runtime/pkg/event" + "github.com/simplyblock/simplyblock-operator/internal/cpinformer" "github.com/simplyblock/simplyblock-operator/internal/cpinformer/subscriptions" ) @@ -212,3 +214,7 @@ func (d *deliveredNodes) Lookup(nodeID string) (cpinformer.Scope, subscriptions. func (d *deliveredNodes) List(cpinformer.Scope) []subscriptions.NodeDTO { return d.nodes } func (d *deliveredNodes) Synced(cpinformer.Scope) bool { return d.synced } + +// Triggers is nil here: these cases drive the reconcile themselves rather than +// waiting to be woken by the stream. +func (d *deliveredNodes) Triggers() <-chan event.GenericEvent { return nil } diff --git a/operator/internal/controllers/node/pushednodes_test.go b/operator/internal/controllers/node/pushednodes_test.go new file mode 100644 index 000000000..5aaa26c04 --- /dev/null +++ b/operator/internal/controllers/node/pushednodes_test.go @@ -0,0 +1,56 @@ +// Whether a pushed control-plane change wakes the node controller. +// +// These assert the source is built, not that an event arrives: a channel source +// only delivers under a running manager, and this package has no envtest suite +// to start one in. So the case below would not catch somebody deleting the +// WatchesRawSource call while leaving this helper -- what it catches is the +// wiring being dropped, which is how it was lost the first time. + +package node + +import ( + "testing" + + "sigs.k8s.io/controller-runtime/pkg/event" + + "github.com/simplyblock/simplyblock-operator/internal/cpinformer" + "github.com/simplyblock/simplyblock-operator/internal/cpinformer/subscriptions" +) + +// aStreamingCache is a NodeCache with a trigger channel and nothing in it. +type aStreamingCache struct{ ch chan event.GenericEvent } + +func (c *aStreamingCache) Triggers() <-chan event.GenericEvent { return c.ch } + +func (c *aStreamingCache) Lookup(string) (cpinformer.Scope, subscriptions.NodeDTO, bool) { + return cpinformer.Scope{}, subscriptions.NodeDTO{}, false +} +func (c *aStreamingCache) List(cpinformer.Scope) []subscriptions.NodeDTO { return nil } +func (c *aStreamingCache) Synced(cpinformer.Scope) bool { return true } + +// TestAPushedNodeChangeWakesTheController is the wiring that was lost. +// +// Regression: 2026-09-21-the-node-stream-woke-nobody — the storage-node +// subscription kept publishing triggers and the controller stopped watching +// them when the kind moved to v1alpha2. Steady state went on reading node status +// out of the same subscription, so the informer looked connected and every +// status change waited out nodeRetry anyway. The stream was a cache, and it was +// built to be a push. +func TestAPushedNodeChangeWakesTheController(t *testing.T) { + r := &StorageNodeReconciler{Nodes: &aStreamingCache{ch: make(chan event.GenericEvent, 1)}} + + if r.pushedNodeChanges() == nil { + t.Error("the controller is not woken by the storage-node stream, so a status " + + "change is noticed only when the requeue interval next comes round") + } +} + +// A deployment with no informer runs on its requeue interval and must not be +// given a source over a nil cache. +func TestNoInformerIsNoSource(t *testing.T) { + r := &StorageNodeReconciler{} + + if r.pushedNodeChanges() != nil { + t.Error("a deployment with no control-plane informer was given a stream source") + } +} diff --git a/operator/internal/controllers/node/storagenode_controller.go b/operator/internal/controllers/node/storagenode_controller.go index 679f90d04..176b83d75 100644 --- a/operator/internal/controllers/node/storagenode_controller.go +++ b/operator/internal/controllers/node/storagenode_controller.go @@ -43,6 +43,7 @@ import ( "sigs.k8s.io/controller-runtime/pkg/handler" logf "sigs.k8s.io/controller-runtime/pkg/log" "sigs.k8s.io/controller-runtime/pkg/reconcile" + "sigs.k8s.io/controller-runtime/pkg/source" "github.com/simplyblock/atlas/prometheus" "github.com/simplyblock/atlas/ptr" @@ -151,14 +152,38 @@ func (r *StorageNodeReconciler) SetupWithManager(mgr ctrl.Manager) error { return fmt.Errorf("index nodes by their cluster: %w", err) } - return ctrl.NewControllerManagedBy(mgr). + builder := ctrl.NewControllerManagedBy(mgr). For(&simplyblockv1alpha2.StorageNode{}). Named("storagenode"). Watches(&simplyblockv1alpha2.StorageCluster{}, handler.EnqueueRequestsFromMapFunc(r.nodesOf)). Watches(&corev1.Node{}, - handler.EnqueueRequestsFromMapFunc(r.nodesOn)). - Complete(r) + handler.EnqueueRequestsFromMapFunc(r.nodesOn)) + + if pushed := r.pushedNodeChanges(); pushed != nil { + builder = builder.WatchesRawSource(pushed) + } + + return builder.Complete(r) +} + +// pushedNodeChanges is the control-plane stream this controller is woken by, and +// nil for a deployment running without the informer. +// +// A pushed change reconciles the node it named; the events already carry the +// object's own name, so they need no map function. +// +// Without it the subscription is a cache rather than a push. Steady state still +// reads a node's status out of it rather than asking the control plane, so +// nothing looks broken -- and every change still waits out nodeRetry to be +// noticed, which is the whole of what the informer was built to remove. It was +// wired once and went out with the move to v1alpha2 (d3380539), and nothing +// failed to say so. +func (r *StorageNodeReconciler) pushedNodeChanges() source.Source { + if r.Nodes == nil { + return nil + } + return source.Channel(r.Nodes.Triggers(), &handler.EnqueueRequestForObject{}) } // nodesOf enqueues every node of a cluster. diff --git a/operator/internal/controllers/node/storagenodeops_controller.go b/operator/internal/controllers/node/storagenodeops_controller.go index 9f18fcb56..2b9ab2b15 100644 --- a/operator/internal/controllers/node/storagenodeops_controller.go +++ b/operator/internal/controllers/node/storagenodeops_controller.go @@ -42,6 +42,7 @@ import ( ctrl "sigs.k8s.io/controller-runtime" "sigs.k8s.io/controller-runtime/pkg/client" "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil" + "sigs.k8s.io/controller-runtime/pkg/event" "sigs.k8s.io/controller-runtime/pkg/handler" logf "sigs.k8s.io/controller-runtime/pkg/log" "sigs.k8s.io/controller-runtime/pkg/reconcile" @@ -115,6 +116,12 @@ type StorageNodeOpsReconciler struct { // NodeCache is the part of the storage-node subscription this package reads. type NodeCache interface { + // Triggers is the reconcile-trigger stream; each event names a StorageNode. + // Reading the cache without it makes the subscription a cache rather than a + // push: a node's status is read from the stream instead of the control plane + // and still waits out the requeue interval to be noticed. + Triggers() <-chan event.GenericEvent + // Lookup returns one node by its backend id, which is what every completion // condition in this package is a predicate over. The scope it comes back with // is the cluster the node was streamed under, which this package already knows From 15563dc77fd6359f0f93ae8cfd64948a2cc8a2a4 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Mon, 21 Sep 2026 18:36:07 +0100 Subject: [PATCH 126/206] fix Repo: Lint error --- .../volumegroupreplication_controller.go | 24 +++++++++---------- ...megroupreplication_controller_unit_test.go | 19 +++++++-------- 2 files changed, 21 insertions(+), 22 deletions(-) diff --git a/operator/internal/controller/volumegroupreplication_controller.go b/operator/internal/controller/volumegroupreplication_controller.go index 84cac75e2..0b87e490e 100644 --- a/operator/internal/controller/volumegroupreplication_controller.go +++ b/operator/internal/controller/volumegroupreplication_controller.go @@ -375,7 +375,7 @@ func (r *VolumeGroupReplicationReconciler) fanIn( members []corev1.PersistentVolumeClaim, memberVRs []unstructured.Unstructured, ) error { - wasDegraded := conditionStatus(vgr, "Degraded") == "True" + wasDegraded := conditionStatus(vgr, "Degraded") == metav1.ConditionTrue completed := len(memberVRs) > 0 degraded := false @@ -383,13 +383,13 @@ func (r *VolumeGroupReplicationReconciler) fanIn( var oldestSync *time.Time for i := range memberVRs { vr := &memberVRs[i] - if conditionStatus(vr, "Completed") != "True" { + if conditionStatus(vr, "Completed") != metav1.ConditionTrue { completed = false } - if conditionStatus(vr, "Degraded") == "True" { + if conditionStatus(vr, "Degraded") == metav1.ConditionTrue { degraded = true } - if conditionStatus(vr, "Resyncing") == "True" { + if conditionStatus(vr, "Resyncing") == metav1.ConditionTrue { resyncing = true } if ts, found, _ := unstructured.NestedString(vr.Object, "status", "lastSyncTime"); found && ts != "" { @@ -439,10 +439,10 @@ func (r *VolumeGroupReplicationReconciler) fanIn( return nil } -// conditionStatus returns a condition's status ("True"/"False"/"Unknown") on -// an unstructured VolumeReplication or VolumeGroupReplication, or "" if the -// object carries no condition of that type. -func conditionStatus(obj *unstructured.Unstructured, condType string) string { +// conditionStatus returns a condition's status (metav1.ConditionTrue/False/ +// Unknown) on an unstructured VolumeReplication or VolumeGroupReplication, or +// "" if the object carries no condition of that type. +func conditionStatus(obj *unstructured.Unstructured, condType string) metav1.ConditionStatus { raw, found, _ := unstructured.NestedSlice(obj.Object, "status", "conditions") if !found { return "" @@ -454,7 +454,7 @@ func conditionStatus(obj *unstructured.Unstructured, condType string) string { } if cm["type"] == condType { if s, ok := cm["status"].(string); ok { - return s + return metav1.ConditionStatus(s) } } } @@ -463,13 +463,13 @@ func conditionStatus(obj *unstructured.Unstructured, condType string) string { // groupCondition builds one status.conditions entry. func groupCondition(condType string, status bool, now metav1.Time) map[string]interface{} { - s := "False" + s := metav1.ConditionFalse if status { - s = "True" + s = metav1.ConditionTrue } return map[string]interface{}{ "type": condType, - "status": s, + "status": string(s), "reason": "GroupMemberAggregation", "message": "", "lastTransitionTime": now.UTC().Format(time.RFC3339), diff --git a/operator/internal/controller/volumegroupreplication_controller_unit_test.go b/operator/internal/controller/volumegroupreplication_controller_unit_test.go index 99b875760..1ae2bee3f 100644 --- a/operator/internal/controller/volumegroupreplication_controller_unit_test.go +++ b/operator/internal/controller/volumegroupreplication_controller_unit_test.go @@ -183,11 +183,11 @@ func vgrMemberVR(name, groupName, pvcName string, m vgrMember) *unstructured.Uns _ = unstructured.SetNestedField(vr.Object, "primary", "spec", "replicationState") cond := func(condType string, status bool) map[string]interface{} { - s := "False" + s := metav1.ConditionFalse if status { - s = "True" + s = metav1.ConditionTrue } - return map[string]interface{}{"type": condType, "status": s, "reason": "Test", "message": ""} + return map[string]interface{}{"type": condType, "status": string(s), "reason": "Test", "message": ""} } conditions := []interface{}{ cond("Completed", m.completed), @@ -207,12 +207,11 @@ func runVGRReconcile(t *testing.T, members []vgrMember) (conditions map[string]s vgrBackend{members: memberIDs(members)}.start(t) group := vgrGroup("vgr1", "sb-group-class", "sb-vr-class", "primary") - objs := []client.Object{group, vgrClass("sb-group-class")} - var memberIDsOnly []string + objs := make([]client.Object, 0, 2+2*len(members)) + objs = append(objs, group, vgrClass("sb-group-class")) for _, m := range members { pvc, pv := vgrPVCAndPV(m.pvcName, m.volumeID) objs = append(objs, pvc, pv) - memberIDsOnly = append(memberIDsOnly, m.volumeID) } r, cl := newVGRReconciler(t, objs...) for _, m := range members { @@ -247,7 +246,7 @@ func runVGRReconcile(t *testing.T, members []vgrMember) (conditions map[string]s } func memberIDs(members []vgrMember) []string { - var ids []string + ids := make([]string, 0, len(members)) for _, m := range members { ids = append(ids, m.volumeID) } @@ -260,10 +259,10 @@ func TestVolumeGroupReplication_AllMembersHealthyYieldsGroupHealthy(t *testing.T {pvcName: "pvc-a", volumeID: "v1", completed: true, degraded: false, lastSync: time.Now()}, {pvcName: "pvc-b", volumeID: "v2", completed: true, degraded: false, lastSync: time.Now()}, }) - if conditions["Completed"] != "True" { + if conditions["Completed"] != string(metav1.ConditionTrue) { t.Errorf("Completed = %q, want True", conditions["Completed"]) } - if conditions["Degraded"] != "False" { + if conditions["Degraded"] != string(metav1.ConditionFalse) { t.Errorf("Degraded = %q, want False", conditions["Degraded"]) } } @@ -274,7 +273,7 @@ func TestVolumeGroupReplication_OneMemberDegradedYieldsGroupDegraded(t *testing. {pvcName: "pvc-a", volumeID: "v1", completed: true, degraded: false, lastSync: time.Now()}, {pvcName: "pvc-b", volumeID: "v2", completed: true, degraded: true, lastSync: time.Now()}, }) - if conditions["Degraded"] != "True" { + if conditions["Degraded"] != string(metav1.ConditionTrue) { t.Errorf("Degraded = %q, want True", conditions["Degraded"]) } } From 01b7e0dda8f774417f1ae7794c9fcb05bafa11ee Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Mon, 21 Sep 2026 20:31:43 +0200 Subject: [PATCH 127/206] feat(operator): a discovery run can name the workers it inspects spec.discover.nodeSelector was the only way to choose them, and a selector cannot say "these two machines": its entries are ANDed, so two hostnames in one selector match nothing at all. Naming a set meant labeling the nodes first and selecting on the label, which is a cluster-wide edit to express a one-off intent. spec.discover.workers is the list. It is exclusive with nodeSelector rather than intersected, refused by a CEL rule on the block: the intersection of a name list and a label selector is not a question anybody asks deliberately, and reading one as narrowing the other makes a run inspect fewer machines than either field says. A named worker that does not exist is announced. The run was told to consider that machine, and the alternative is a draft short of the machines somebody listed with nothing anywhere saying which -- the same reason the selector path already emits an event for a worker it declined rather than dropping it. A name that resolves is still only a statement about which machines to consider: it is declined for the usual reasons like any other. The filter runs over the node list rather than fetching each name, so everything after it sees one list however the run chose its machines. Test plan: U-182 through U-184. Co-Authored-By: Claude Opus 5 (1M context) --- operator/api/v1alpha2/operatorops_types.go | 24 ++++- .../api/v1alpha2/zz_generated.deepcopy.go | 5 + .../storage.simplyblock.io_operatorops.yaml | 24 ++++- operator/dist/install.yaml | 24 ++++- .../test-plan-clusterdeploymentconfig.md | 3 + .../deployment/namedworkers_test.go | 95 +++++++++++++++++++ .../deployment/operatorops_controller.go | 45 +++++++++ .../storage.simplyblock.io_operatorops.yaml | 24 ++++- test-discovery.yaml | 25 +++++ 9 files changed, 265 insertions(+), 4 deletions(-) create mode 100644 operator/internal/controllers/deployment/namedworkers_test.go create mode 100644 test-discovery.yaml diff --git a/operator/api/v1alpha2/operatorops_types.go b/operator/api/v1alpha2/operatorops_types.go index fccd4aa94..c0502f440 100644 --- a/operator/api/v1alpha2/operatorops_types.go +++ b/operator/api/v1alpha2/operatorops_types.go @@ -137,6 +137,13 @@ type DeviceFilter struct { } // DiscoverSpec parameterizes the Discover action. +// +// Which workers a run inspects is stated one of two ways and never both: by name +// in Workers, or by label in NodeSelector. They are exclusive rather than +// intersected because the intersection of a name list and a label selector is a +// question nobody asks deliberately, and reading one as narrowing the other +// would make a run inspect fewer machines than either field says. +// +kubebuilder:validation:XValidation:rule="!(has(self.workers) && size(self.workers) > 0 && has(self.nodeSelector) && size(self.nodeSelector) > 0)",message="spec.discover names workers and also carries a nodeSelector; state one or the other" type DiscoverSpec struct { // ConfigName is the ClusterDeploymentConfig to write. Absent generates one // from the run's timestamp, so that a second discovery never overwrites the @@ -144,8 +151,23 @@ type DiscoverSpec struct { // +optional ConfigName string `json:"configName,omitempty"` + // Workers are the workers to inspect, by node name. + // + // It is the answer to inspecting two named machines, which a label selector + // can only express by labeling them first: a selector's entries are ANDed, + // so two hostnames in one selector match nothing at all. + // + // A named worker that does not exist, or that is declined for one of the + // reasons any worker is declined, is reported by the same event the selector + // path reports it by. Naming a worker is a statement about which machines to + // consider, not a claim that each of them will be used. + // +kubebuilder:validation:MaxItems=128 + // +kubebuilder:validation:items:MaxLength=253 + // +optional + Workers []string `json:"workers,omitempty"` + // NodeSelector restricts which workers are inspected. Empty inspects every - // schedulable worker. + // schedulable worker, and it is exclusive with Workers. // +optional NodeSelector map[string]string `json:"nodeSelector,omitempty"` diff --git a/operator/api/v1alpha2/zz_generated.deepcopy.go b/operator/api/v1alpha2/zz_generated.deepcopy.go index 2a0d1c58a..ce1ff202b 100644 --- a/operator/api/v1alpha2/zz_generated.deepcopy.go +++ b/operator/api/v1alpha2/zz_generated.deepcopy.go @@ -766,6 +766,11 @@ func (in *DeviceSelection) DeepCopy() *DeviceSelection { // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *DiscoverSpec) DeepCopyInto(out *DiscoverSpec) { *out = *in + if in.Workers != nil { + in, out := &in.Workers, &out.Workers + *out = make([]string, len(*in)) + copy(*out, *in) + } if in.NodeSelector != nil { in, out := &in.NodeSelector, &out.NodeSelector *out = make(map[string]string, len(*in)) diff --git a/operator/config/crd/bases/storage.simplyblock.io_operatorops.yaml b/operator/config/crd/bases/storage.simplyblock.io_operatorops.yaml index e4c509e08..ea81f93a0 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_operatorops.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_operatorops.yaml @@ -210,9 +210,31 @@ spec: type: string description: |- NodeSelector restricts which workers are inspected. Empty inspects every - schedulable worker. + schedulable worker, and it is exclusive with Workers. type: object + workers: + description: |- + Workers are the workers to inspect, by node name. + + It is the answer to inspecting two named machines, which a label selector + can only express by labeling them first: a selector's entries are ANDed, + so two hostnames in one selector match nothing at all. + + A named worker that does not exist, or that is declined for one of the + reasons any worker is declined, is reported by the same event the selector + path reports it by. Naming a worker is a statement about which machines to + consider, not a claim that each of them will be used. + items: + maxLength: 253 + type: string + maxItems: 128 + type: array type: object + x-kubernetes-validations: + - message: spec.discover names workers and also carries a nodeSelector; + state one or the other + rule: '!(has(self.workers) && size(self.workers) > 0 && has(self.nodeSelector) + && size(self.nodeSelector) > 0)' required: - action type: object diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index a9cf29751..32a3938dd 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -1963,9 +1963,31 @@ spec: type: string description: |- NodeSelector restricts which workers are inspected. Empty inspects every - schedulable worker. + schedulable worker, and it is exclusive with Workers. type: object + workers: + description: |- + Workers are the workers to inspect, by node name. + + It is the answer to inspecting two named machines, which a label selector + can only express by labeling them first: a selector's entries are ANDed, + so two hostnames in one selector match nothing at all. + + A named worker that does not exist, or that is declined for one of the + reasons any worker is declined, is reported by the same event the selector + path reports it by. Naming a worker is a statement about which machines to + consider, not a claim that each of them will be used. + items: + maxLength: 253 + type: string + maxItems: 128 + type: array type: object + x-kubernetes-validations: + - message: spec.discover names workers and also carries a nodeSelector; + state one or the other + rule: '!(has(self.workers) && size(self.workers) > 0 && has(self.nodeSelector) + && size(self.nodeSelector) > 0)' required: - action type: object diff --git a/operator/docs/tests/test-plan-clusterdeploymentconfig.md b/operator/docs/tests/test-plan-clusterdeploymentconfig.md index 2d048a145..e5de834df 100644 --- a/operator/docs/tests/test-plan-clusterdeploymentconfig.md +++ b/operator/docs/tests/test-plan-clusterdeploymentconfig.md @@ -97,6 +97,9 @@ File: `operator/internal/controllers/deployment/clusterdeploymentconfig_expand_t | U-179 | An `Available` control plane proceeds | Positive | `TestAnAvailableControlPlaneProceeds` | | U-180 | An `Unavailable` control plane holds, with a reason | Negative | `TestAnUnavailableControlPlaneHolds` | | U-181 | A control plane still being installed, or reporting no phase, holds | Boundary | `TestAControlPlaneStillBeingBuiltHolds` | +| U-182 | A run naming workers inspects those and no others | Positive | `TestOnlyTheNamedWorkersAreInspected` | +| U-183 | A named worker that does not exist is announced, not silently dropped | Regression | `TestANamedWorkerThatDoesNotExistIsAnnounced` | +| U-184 | Naming no worker keeps nothing, and the caller skips the filter | Boundary | `TestNamingNoWorkerKeepsNothing` | ### Deletion (design §4.3) diff --git a/operator/internal/controllers/deployment/namedworkers_test.go b/operator/internal/controllers/deployment/namedworkers_test.go new file mode 100644 index 000000000..174b6d917 --- /dev/null +++ b/operator/internal/controllers/deployment/namedworkers_test.go @@ -0,0 +1,95 @@ +// Naming the workers a discovery run inspects. +// +// A label selector cannot express "these two machines": its entries are ANDed, +// so two hostnames in one selector match nothing. spec.discover.workers is the +// list, and what these cover is that a name matching nothing is announced rather +// than quietly leaving the draft short. + +package deployment + +import ( + "strings" + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/client-go/tools/events" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// aWorker is a node by name. +func aWorker(name string) corev1.Node { + return corev1.Node{ObjectMeta: metav1.ObjectMeta{Name: name}} +} + +// aDiscoveryRun is the reconciler with a recorder its events can be read from. +func aDiscoveryRun(t *testing.T) (*OperatorOpsReconciler, *simplyblockv1alpha2.OperatorOps, *events.FakeRecorder) { + t.Helper() + recorder := events.NewFakeRecorder(16) + ops := &simplyblockv1alpha2.OperatorOps{ + ObjectMeta: metav1.ObjectMeta{Name: "rediscover", Namespace: "simplyblock"}, + } + return &OperatorOpsReconciler{Recorder: recorder}, ops, recorder +} + +// The named workers are the ones kept, and nothing else is. +func TestOnlyTheNamedWorkersAreInspected(t *testing.T) { + r, ops, _ := aDiscoveryRun(t) + + kept := r.namedWorkers(ops, + []string{"worker-4.ocp.simplyblock.ai", "worker-5.ocp.simplyblock.ai"}, + []corev1.Node{ + aWorker("worker-3.ocp.simplyblock.ai"), + aWorker("worker-4.ocp.simplyblock.ai"), + aWorker("worker-5.ocp.simplyblock.ai"), + }) + + if len(kept) != 2 { + t.Fatalf("the run inspects %d workers, want the two it named", len(kept)) + } + for _, node := range kept { + if node.Name == "worker-3.ocp.simplyblock.ai" { + t.Error("a worker the run did not name is inspected") + } + } +} + +// TestANamedWorkerThatDoesNotExistIsAnnounced is why this is a function rather +// than a filter written inline. +// +// Being asked to inspect a machine that is not there is a typo or a node that +// has not joined. Both are worth saying: the alternative is a draft short of the +// machines somebody listed, with nothing anywhere saying which. +func TestANamedWorkerThatDoesNotExistIsAnnounced(t *testing.T) { + r, ops, recorder := aDiscoveryRun(t) + + kept := r.namedWorkers(ops, + []string{"worker-4.ocp.simplyblock.ai", "worker-9.typo"}, + []corev1.Node{aWorker("worker-4.ocp.simplyblock.ai")}) + + if len(kept) != 1 { + t.Fatalf("the run inspects %d workers", len(kept)) + } + + var announced string + select { + case announced = <-recorder.Events: + default: + t.Fatal("a named worker that does not exist was dropped without an event") + } + if !strings.Contains(announced, "worker-9.typo") { + t.Errorf("the event does not name the worker: %s", announced) + } +} + +// Naming nothing is not the same as naming everything, and the caller is what +// decides not to filter. This asserts the helper is never asked to. +func TestNamingNoWorkerKeepsNothing(t *testing.T) { + r, ops, _ := aDiscoveryRun(t) + + if kept := r.namedWorkers(ops, nil, []corev1.Node{aWorker("worker-4")}); len(kept) != 0 { + t.Errorf("an empty name list kept %d workers, and the caller skips the filter "+ + "entirely for that case", len(kept)) + } +} diff --git a/operator/internal/controllers/deployment/operatorops_controller.go b/operator/internal/controllers/deployment/operatorops_controller.go index f6eb0eecd..f3137604f 100644 --- a/operator/internal/controllers/deployment/operatorops_controller.go +++ b/operator/internal/controllers/deployment/operatorops_controller.go @@ -343,6 +343,15 @@ func (r *OperatorOpsReconciler) inspect( return false, err } + // A named set is filtered here rather than fetched one node at a time, so + // that everything after this point sees one node list however the run chose + // its machines. A name matching nothing is announced: the run was told to + // consider that machine, and a draft quietly missing it is the case the + // declined events below exist for. + if len(spec.Workers) > 0 { + nodes.Items = r.namedWorkers(ops, spec.Workers, nodes.Items) + } + taken, err := r.workersAlreadyTaken(ctx, ops.Namespace) if err != nil { return false, err @@ -985,3 +994,39 @@ func boundEventMessage(message string) string { const ellipsis = " […]" return message[:maxEventMessage-len(ellipsis)] + ellipsis } + +// namedWorkers keeps the nodes a run named, in the order the API server +// returned them, and announces each name that matched nothing. +// +// The announcement is the point. Being asked to inspect a machine that is not +// there is either a typo or a node that has not joined, and both are worth one +// event: the alternative is a draft short of the machines somebody listed, with +// nothing anywhere saying which or why. +func (r *OperatorOpsReconciler) namedWorkers( + ops *simplyblockv1alpha2.OperatorOps, + named []string, + nodes []corev1.Node, +) []corev1.Node { + wanted := make(map[string]bool, len(named)) + for _, name := range named { + wanted[name] = false + } + + kept := make([]corev1.Node, 0, len(nodes)) + for i := range nodes { + if _, ok := wanted[nodes[i].Name]; !ok { + continue + } + wanted[nodes[i].Name] = true + kept = append(kept, nodes[i]) + } + + for _, name := range named { + if !wanted[name] { + r.event(ops, corev1.EventTypeWarning, WorkerDeclined, + fmt.Sprintf("%s is named in spec.discover.workers and there is no node "+ + "by that name", name)) + } + } + return kept +} diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_operatorops.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_operatorops.yaml index e4c509e08..ea81f93a0 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_operatorops.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_operatorops.yaml @@ -210,9 +210,31 @@ spec: type: string description: |- NodeSelector restricts which workers are inspected. Empty inspects every - schedulable worker. + schedulable worker, and it is exclusive with Workers. type: object + workers: + description: |- + Workers are the workers to inspect, by node name. + + It is the answer to inspecting two named machines, which a label selector + can only express by labeling them first: a selector's entries are ANDed, + so two hostnames in one selector match nothing at all. + + A named worker that does not exist, or that is declined for one of the + reasons any worker is declined, is reported by the same event the selector + path reports it by. Naming a worker is a statement about which machines to + consider, not a claim that each of them will be used. + items: + maxLength: 253 + type: string + maxItems: 128 + type: array type: object + x-kubernetes-validations: + - message: spec.discover names workers and also carries a nodeSelector; + state one or the other + rule: '!(has(self.workers) && size(self.workers) > 0 && has(self.nodeSelector) + && size(self.nodeSelector) > 0)' required: - action type: object diff --git a/test-discovery.yaml b/test-discovery.yaml new file mode 100644 index 000000000..bb6beeffc --- /dev/null +++ b/test-discovery.yaml @@ -0,0 +1,25 @@ +--- +apiVersion: storage.simplyblock.io/v1alpha2 +kind: OperatorOps +metadata: + name: rediscover-rack-2 + namespace: simplyblock +spec: + action: Discover + discover: + configName: discovered-rack-2 + + clusterRef: discovered-initial-discovery-cluster + + workers: + - worker-4.ocp.simplyblock.ai + - worker-5.ocp.simplyblock.ai + + # nodeSelector: + # simplyblock.io/rack: "2" + + enableControlPlaneNodes: false + + deviceFilter: + enableLogicalBlockDevices: false + enablePartitionedDevices: true From ba8211861c73ab44e4bf841acab4adb4b663947e Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Tue, 22 Sep 2026 12:30:13 +0100 Subject: [PATCH 128/206] Authenticate cluster-scoped control-plane calls with the cluster's own secret, not the operator's SA token --- .../controllers/cluster/controlplane.go | 10 ++ .../controllers/cluster/helpers_test.go | 18 ++- .../cluster/storagecluster_controller.go | 38 ++++++- .../cluster/storagecluster_controller_test.go | 105 ++++++++++++++++++ .../cluster/storageclusterops_controller.go | 43 +++++++ .../storageclusterops_controller_test.go | 40 +++++++ operator/internal/webapi/context.go | 23 ++++ operator/internal/webapi/request.go | 10 +- operator/internal/webapi/request_test.go | 31 ++++++ 9 files changed, 309 insertions(+), 9 deletions(-) create mode 100644 operator/internal/webapi/context.go diff --git a/operator/internal/controllers/cluster/controlplane.go b/operator/internal/controllers/cluster/controlplane.go index 8cef938f3..8c63eb1cc 100644 --- a/operator/internal/controllers/cluster/controlplane.go +++ b/operator/internal/controllers/cluster/controlplane.go @@ -79,6 +79,14 @@ type ControlPlane interface { // outcome being asked for reached by another route, so the implementation // reads a 404 as success. CancelTask(ctx context.Context, clusterID, taskID string) error + + // Endpoint is the base URL this control plane was actually reached at, + // resolved from SIMPLYBLOCK_WEBAPI_BASE_URL when set. upsertCSICredentials + // carries it into the CSI driver's aggregate Secret, since the CSI driver + // dials it directly rather than through this reconciler: a hardcoded + // in-cluster address is wrong the moment the control plane this cluster was + // adopted from lives on a different Kubernetes cluster (a shared hub). + Endpoint() string } // httpControlPlane is the ControlPlane the operator runs with: the shared @@ -88,6 +96,8 @@ type httpControlPlane struct{ client *webapi.Client } // NewControlPlane returns the HTTP-backed control-plane surface. func NewControlPlane() ControlPlane { return &httpControlPlane{client: webapi.NewClient()} } +func (c *httpControlPlane) Endpoint() string { return c.client.BaseURL } + func (c *httpControlPlane) Ready(ctx context.Context) error { _, err := c.call(ctx, http.MethodGet, "/api/v2/_meta/ready", nil) return err diff --git a/operator/internal/controllers/cluster/helpers_test.go b/operator/internal/controllers/cluster/helpers_test.go index 62a99c510..a8839fd47 100644 --- a/operator/internal/controllers/cluster/helpers_test.go +++ b/operator/internal/controllers/cluster/helpers_test.go @@ -168,6 +168,11 @@ var _ events.EventRecorder = (*recorder)(nil) type fakeControlPlane struct { t *testing.T + // endpoint is what Endpoint() returns verbatim -- a pure config read, not + // an action the test scripts, so a zero value is a legitimate "unset" + // rather than an unexpected call. + endpoint string + ready func() error create func(utils.ClusterAddParams) (webapi.ClusterResponse, error) cluster func(string) (webapi.ClusterResponse, error) @@ -183,6 +188,12 @@ type fakeControlPlane struct { tasks func(string) ([]subscriptions.TaskDTO, error) cancelTask func(string, string) error + // clusterCtx, when set, is handed the context Cluster() was called with -- + // how a test observes the bearer-token override a caller attached + // (webapi.BearerTokenFromContext), since the closures above only see the + // clusterID. + clusterCtx func(context.Context) + // The counters the write-ahead tests read. Each is the number of times the // control plane was actually asked to do something, which is the only way // to tell a step that skipped its call from one that made it twice. @@ -196,6 +207,8 @@ type fakeControlPlane struct { cancelTaskCalls int } +func (f *fakeControlPlane) Endpoint() string { return f.endpoint } + func (f *fakeControlPlane) Ready(context.Context) error { if f.ready == nil { return nil @@ -214,8 +227,11 @@ func (f *fakeControlPlane) CreateCluster( } func (f *fakeControlPlane) Cluster( - _ context.Context, clusterID string, + ctx context.Context, clusterID string, ) (webapi.ClusterResponse, error) { + if f.clusterCtx != nil { + f.clusterCtx(ctx) + } if f.cluster == nil { f.t.Fatal("the control plane was asked for a cluster and the test did not script it") } diff --git a/operator/internal/controllers/cluster/storagecluster_controller.go b/operator/internal/controllers/cluster/storagecluster_controller.go index 131c9951d..555eec066 100644 --- a/operator/internal/controllers/cluster/storagecluster_controller.go +++ b/operator/internal/controllers/cluster/storagecluster_controller.go @@ -48,6 +48,7 @@ import ( "github.com/simplyblock/simplyblock-operator/internal/cpinformer" "github.com/simplyblock/simplyblock-operator/internal/cpinformer/subscriptions" "github.com/simplyblock/simplyblock-operator/internal/utils" + "github.com/simplyblock/simplyblock-operator/internal/webapi" ) const ( @@ -528,7 +529,13 @@ func (r *StorageClusterReconciler) upgradeClaim( return adoption{}, false, nil } - found, err := r.API.Cluster(ctx, uuid) + // This read authenticates as the cluster itself, using the secret the + // upgrade Secret names, rather than as this operator's own Kubernetes + // identity: the control plane this cluster belongs to may run on a + // different Kubernetes cluster (a shared hub), where a Kubernetes + // TokenReview of this operator's own service-account token can never + // succeed. + found, err := r.API.Cluster(webapi.WithBearerToken(ctx, clusterSecret), uuid) if err != nil { // The Secret names a cluster the control plane does not have. That is // worth retrying rather than failing: the control plane may be @@ -661,7 +668,8 @@ func (r *StorageClusterReconciler) sync( ) (ctrl.Result, error) { log := logf.FromContext(ctx) - if secret, err := r.clusterSecret(ctx, cluster); err == nil && secret != "" { + secret, err := r.clusterSecret(ctx, cluster) + if err == nil && secret != "" { if err := r.upsertCSICredentials(ctx, cluster.Status.UUID, secret); err != nil { log.Error(err, "the CSI credentials entry could not be restored", "cluster", cluster.Name) @@ -669,13 +677,23 @@ func (r *StorageClusterReconciler) sync( } } - reading, err := r.reading(ctx, cluster.Status.UUID) + // Every read below authenticates as this cluster, using its own recorded + // secret, rather than as this operator's own Kubernetes identity: the + // control plane this cluster belongs to may run on a different Kubernetes + // cluster (a shared hub), where a Kubernetes TokenReview of this + // operator's own service-account token can never succeed. + readCtx := ctx + if secret != "" { + readCtx = webapi.WithBearerToken(ctx, secret) + } + + reading, err := r.reading(readCtx, cluster.Status.UUID) if err != nil { log.Error(err, "the cluster could not be read", "cluster", cluster.Name) return ctrl.Result{RequeueAfter: clusterResync}, nil } - tasks := r.readTasks(ctx, cluster) + tasks := r.readTasks(readCtx, cluster) ftt := int32(reading.MaxFaultTolerance) //nolint:gosec // a fault tolerance is a small count err = r.writeStatus(ctx, cluster, func(status *simplyblockv1alpha2.StorageClusterStatus) { @@ -816,7 +834,15 @@ func (r *StorageClusterReconciler) teardown( if cluster.Status.UUID != "" { r.closeStreams(cluster.Status.UUID) - if err := r.API.DeleteCluster(ctx, cluster.Status.UUID); err != nil { + // Authenticates as this cluster, using its own recorded secret, for + // the same reason sync() does: the control plane this cluster + // belongs to may run on a different Kubernetes cluster than this + // operator does. + deleteCtx := ctx + if secret, err := r.clusterSecret(ctx, cluster); err == nil && secret != "" { + deleteCtx = webapi.WithBearerToken(ctx, secret) + } + if err := r.API.DeleteCluster(deleteCtx, cluster.Status.UUID); err != nil { log.Error(err, "the cluster could not be deleted; retrying", "cluster", cluster.Name, "uuid", cluster.Status.UUID) return ctrl.Result{RequeueAfter: clusterRetry}, nil @@ -1039,7 +1065,7 @@ func (r *StorageClusterReconciler) upsertCSICredentials( return r.editCSICredentials(ctx, func(creds *CSICredentials) { entry := CSIClusterEntry{ ClusterID: clusterID, - ClusterEndpoint: utils.ENDPOINT, + ClusterEndpoint: r.API.Endpoint(), ClusterSecret: clusterSecret, } for i := range creds.Clusters { diff --git a/operator/internal/controllers/cluster/storagecluster_controller_test.go b/operator/internal/controllers/cluster/storagecluster_controller_test.go index 0a0d0f934..27b6c4121 100644 --- a/operator/internal/controllers/cluster/storagecluster_controller_test.go +++ b/operator/internal/controllers/cluster/storagecluster_controller_test.go @@ -15,6 +15,7 @@ package cluster import ( "context" + "encoding/json" "errors" "fmt" "net/http" @@ -130,6 +131,42 @@ func TestACreatedClusterReachesSteadyState(t *testing.T) { } } +// The CSI driver dials ClusterEndpoint directly rather than through this +// reconciler, so a control plane shared across more than one Kubernetes +// cluster -- the hub in a Ramen/DR topology -- needs that entry to carry the +// endpoint this reconciler was actually configured to reach, not a hardcoded +// address that only resolves inside its own cluster. +func TestCSICredentialsCarryTheControlPlanesConfiguredEndpoint(t *testing.T) { + const hubEndpoint = "http://simplyblock-webappapi.hub.example:31500" + api := &fakeControlPlane{ + endpoint: hubEndpoint, + create: func(utils.ClusterAddParams) (webapi.ClusterResponse, error) { + reading := activeCluster() + reading.Secret = testClusterSecret + return reading, nil + }, + cluster: func(string) (webapi.ClusterResponse, error) { return activeCluster(), nil }, + } + r := newClusterReconciler(t, api, &recorder{}, newUncreatedCluster()) + reconcileCluster(t, r, 6) + + var secret corev1.Secret + key := types.NamespacedName{Namespace: testNamespace, Name: csiCredentialsSecret} + if err := r.Get(context.Background(), key, &secret); err != nil { + t.Fatalf("read the CSI credentials secret: %v", err) + } + var creds CSICredentials + if err := json.Unmarshal(secret.Data["secret.json"], &creds); err != nil { + t.Fatalf("unmarshal secret.json: %v", err) + } + if len(creds.Clusters) != 1 { + t.Fatalf("clusters = %d entries, want 1", len(creds.Clusters)) + } + if got := creds.Clusters[0].ClusterEndpoint; got != hubEndpoint { + t.Errorf("clusterEndpoint = %q, want the configured endpoint %q", got, hubEndpoint) + } +} + // spec.deviceClass is the CRD's spelling of what sbcli's cluster-create wire // format calls `device_mode`, so the two must map onto each other rather than // the field simply passing through unmapped. @@ -263,6 +300,74 @@ func TestAnUpgradeSecretAdoptsRatherThanCreating(t *testing.T) { } } +// An adopted cluster's control plane may run on a different Kubernetes +// cluster than this operator does (a shared hub), where a Kubernetes +// TokenReview of this operator's own service-account token can never +// succeed. The read that confirms the adoption must authenticate as the +// cluster itself instead, using the secret the upgrade Secret names. +func TestAnUpgradeSecretAuthenticatesTheAdoptionReadAsTheAdoptedCluster(t *testing.T) { + var gotToken string + var gotOK bool + api := &fakeControlPlane{ + cluster: func(string) (webapi.ClusterResponse, error) { + return activeCluster(), nil + }, + clusterCtx: func(ctx context.Context) { + gotToken, gotOK = webapi.BearerTokenFromContext(ctx) + }, + } + upgrade := &corev1.Secret{ + ObjectMeta: objectMeta("simplyblock-" + testClusterName + "-upgrade"), + Data: map[string][]byte{ + "uuid": []byte(testClusterUUID), + "secret": []byte(testClusterSecret), + }, + } + r := newClusterReconciler(t, api, &recorder{}, newUncreatedCluster(), upgrade) + reconcileCluster(t, r, 6) + + if !gotOK { + t.Fatal("the adoption read carried no bearer-token override") + } + if gotToken != testClusterSecret { + t.Errorf("bearer token = %q, want the upgrade secret's own cluster secret %q", + gotToken, testClusterSecret) + } +} + +// Once a cluster is created and its secret is on record, every later +// steady-state read of it must keep authenticating as that cluster -- the +// same reasoning as the adoption read above, just for the read that runs on +// every reconcile after. +func TestStorageClusterSyncAuthenticatesAsTheClusterOnceItsSecretIsKnown(t *testing.T) { + var gotToken string + var gotOK bool + api := &fakeControlPlane{ + create: func(utils.ClusterAddParams) (webapi.ClusterResponse, error) { + reading := activeCluster() + reading.Secret = testClusterSecret + return reading, nil + }, + cluster: func(string) (webapi.ClusterResponse, error) { return activeCluster(), nil }, + clusterCtx: func(ctx context.Context) { + gotToken, gotOK = webapi.BearerTokenFromContext(ctx) + }, + } + r := newClusterReconciler(t, api, &recorder{}, newUncreatedCluster()) + // The creation machine takes 6 passes to reach steady state + // (TestACreatedClusterReachesSteadyState); one more pass is the first + // steady-state sync, which is the read this test is about. + reconcileCluster(t, r, 7) + + if !gotOK { + t.Fatal("the steady-state read carried no bearer-token override") + } + if gotToken != testClusterSecret { + t.Errorf("bearer token = %q, want the cluster's own recorded secret %q", + gotToken, testClusterSecret) + } +} + // The second route: a POST that failed against a cluster which already exists. // That covers two reconciles that both passed the claim on different // resourceVersions, and a response lost after the backend committed. diff --git a/operator/internal/controllers/cluster/storageclusterops_controller.go b/operator/internal/controllers/cluster/storageclusterops_controller.go index 85ecc9c44..cebace976 100644 --- a/operator/internal/controllers/cluster/storageclusterops_controller.go +++ b/operator/internal/controllers/cluster/storageclusterops_controller.go @@ -51,6 +51,7 @@ import ( "github.com/simplyblock/simplyblock-operator/internal/cpinformer" "github.com/simplyblock/simplyblock-operator/internal/cpinformer/subscriptions" "github.com/simplyblock/simplyblock-operator/internal/utils" + "github.com/simplyblock/simplyblock-operator/internal/webapi" ) const ( @@ -222,6 +223,14 @@ func (r *StorageClusterOpsReconciler) Reconcile( func (r *StorageClusterOpsReconciler) advance( ctx context.Context, ops *simplyblockv1alpha2.StorageClusterOps, ) (ctrl.Result, error) { + // Every control-plane call this step and everything downstream of it makes + // (perform, advanceWalk, and everything under them) authenticates as this + // operation's own cluster when its secret is known, rather than as this + // operator's Kubernetes identity -- the only way to reach a control plane a + // different Kubernetes cluster runs (a shared hub), since a Kubernetes + // TokenReview can never cross a cluster boundary. + ctx = r.authenticatedContext(ctx, ops) + graph := action(ops.Spec.Action) machine, err := graphs().FromSnapshot(ctx, graph, statemachine.FromKube[step](ops.Status.Step)) @@ -868,6 +877,40 @@ func (r *StorageClusterOpsReconciler) clusterReading( }, nil } +// clusterSecret reads the secret StorageClusterReconciler.persist wrote for +// this operation's cluster, keyed by the StorageCluster's Kubernetes name (not +// its backend UUID, which this reconciler is not always given yet at the point +// it needs the credential). It reports the empty string when there is none. +func (r *StorageClusterOpsReconciler) clusterSecret( + ctx context.Context, ops *simplyblockv1alpha2.StorageClusterOps, +) (string, error) { + var secret corev1.Secret + key := types.NamespacedName{ + Name: fmt.Sprintf("simplyblock-cluster-%s", ops.Spec.ClusterRef), + Namespace: ops.Namespace, + } + if err := r.Get(ctx, key, &secret); err != nil { + return "", err + } + return string(secret.Data["secret"]), nil +} + +// authenticatedContext attaches this operation's cluster's own credential to +// ctx when one is known, so every control-plane call the operation makes +// authenticates as that cluster instead of as this operator's Kubernetes +// identity -- the only way to reach a control plane a different Kubernetes +// cluster runs (a shared hub), since a Kubernetes TokenReview can never cross a +// cluster boundary. See StorageClusterReconciler's identically-named method. +func (r *StorageClusterOpsReconciler) authenticatedContext( + ctx context.Context, ops *simplyblockv1alpha2.StorageClusterOps, +) context.Context { + secret, err := r.clusterSecret(ctx, ops) + if err != nil || secret == "" { + return ctx + } + return webapi.WithBearerToken(ctx, secret) +} + // effectiveConcurrentRestarts is min(specVal, FTT), defaulting to 1 when the // spec says nothing. Both inputs may be absent, because the fault tolerance // comes from the control plane. diff --git a/operator/internal/controllers/cluster/storageclusterops_controller_test.go b/operator/internal/controllers/cluster/storageclusterops_controller_test.go index 8194c402d..e6db399ef 100644 --- a/operator/internal/controllers/cluster/storageclusterops_controller_test.go +++ b/operator/internal/controllers/cluster/storageclusterops_controller_test.go @@ -21,6 +21,7 @@ import ( "testing" "time" + corev1 "k8s.io/api/core/v1" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" "k8s.io/apimachinery/pkg/types" ctrl "sigs.k8s.io/controller-runtime" @@ -265,6 +266,45 @@ func TestAShutdownIssuesOneCallAndWaitsForTheCluster(t *testing.T) { } } +// An operation against an already-adopted cluster may reach a control plane on +// a different Kubernetes cluster (a shared hub), the same as +// StorageClusterReconciler: it authenticates as the cluster's own recorded +// secret instead of as this operator's Kubernetes identity, since a Kubernetes +// TokenReview can never cross a cluster boundary. +func TestAnOperationAuthenticatesAsItsClusterOnceItsSecretIsKnown(t *testing.T) { + var gotToken string + var gotOK bool + active := true + api := &fakeControlPlane{ + cluster: func(string) (webapi.ClusterResponse, error) { + reading := activeCluster() + if !active { + reading.Status = "suspended" + } + return reading, nil + }, + clusterCtx: func(ctx context.Context) { + gotToken, gotOK = webapi.BearerTokenFromContext(ctx) + }, + shutdown: func(string) error { active = false; return nil }, + } + secret := &corev1.Secret{ + ObjectMeta: objectMeta("simplyblock-cluster-" + testClusterName), + Data: map[string][]byte{"secret": []byte(testClusterSecret)}, + } + r := newOpsReconciler(t, api, &recorder{}, + newTestCluster(), newTestOps(simplyblockv1alpha2.StorageClusterOpsActionShutdown), secret) + + reconcileOps(t, r, 6) + + if !gotOK { + t.Fatal("the operation's control-plane read carried no bearer-token override") + } + if gotToken != testClusterSecret { + t.Errorf("bearer token = %q, want the cluster's own recorded secret %q", gotToken, testClusterSecret) + } +} + // Restart is the one action with two side effects, because the control plane // has no restart endpoint of its own. func TestARestartShutsDownThenStarts(t *testing.T) { diff --git a/operator/internal/webapi/context.go b/operator/internal/webapi/context.go new file mode 100644 index 000000000..7a0126cd3 --- /dev/null +++ b/operator/internal/webapi/context.go @@ -0,0 +1,23 @@ +package webapi + +import "context" + +type bearerTokenKey struct{} + +// WithBearerToken attaches a bearer credential to ctx that Do and +// DoWithHeaders send instead of the client's own service-account token. It is +// how a call scoped to one cluster authenticates as that cluster rather than +// as this process's own Kubernetes identity -- the only way to reach a +// control plane a different Kubernetes cluster runs (a shared hub in a +// multi-cluster deployment), since a Kubernetes TokenReview can never cross a +// cluster boundary. +func WithBearerToken(ctx context.Context, token string) context.Context { + return context.WithValue(ctx, bearerTokenKey{}, token) +} + +// BearerTokenFromContext returns the token WithBearerToken attached, and +// whether one was. +func BearerTokenFromContext(ctx context.Context) (string, bool) { + token, ok := ctx.Value(bearerTokenKey{}).(string) + return token, ok +} diff --git a/operator/internal/webapi/request.go b/operator/internal/webapi/request.go index 7734cebf2..59f2e3729 100644 --- a/operator/internal/webapi/request.go +++ b/operator/internal/webapi/request.go @@ -51,8 +51,14 @@ func (c *Client) DoWithHeaders( return nil, nil, 0, fmt.Errorf("create request: %w", err) } - // Attach auth header - req.Header.Set("Authorization", fmt.Sprintf("Bearer %s", c.saToken)) + // Attach auth header. A token WithBearerToken attached to ctx wins over + // this process's own service-account token, for a call scoped to a + // cluster whose control plane lives on a different Kubernetes cluster. + token := c.saToken + if override, ok := BearerTokenFromContext(ctx); ok { + token = override + } + req.Header.Set("Authorization", fmt.Sprintf("Bearer %s", token)) req.Header.Set("Content-Type", "application/json") // Execute the request diff --git a/operator/internal/webapi/request_test.go b/operator/internal/webapi/request_test.go index 01406ea25..5a3da827c 100644 --- a/operator/internal/webapi/request_test.go +++ b/operator/internal/webapi/request_test.go @@ -73,6 +73,37 @@ func TestDoAgainstSpecMockSendsHeadersBodyAndReturnsResponse(t *testing.T) { } } +// A call scoped to one cluster authenticates as that cluster when +// WithBearerToken names one, instead of as this process's own service +// account -- the only way to reach a control plane a different Kubernetes +// cluster runs, since a TokenReview can never cross a cluster boundary. +func TestDoSendsTheBearerTokenAttachedToContextInsteadOfTheServiceAccountToken(t *testing.T) { + mock := webapimock.NewSpecServerFromFile(t, "../../../shared/openapi.json", false) + defer mock.Close() + + mock.Register( + http.MethodGet, + "/api/v2/clusters/cluster-uuid/", + webapimock.RouteResponse{Status: http.StatusOK, Body: `{}`}, + ) + + c := NewClient(mock.URL()) + c.saToken = "operators-own-service-account-token" + + ctx := WithBearerToken(context.Background(), "cluster-uuids-own-secret") + if _, _, err := c.Do(ctx, http.MethodGet, "/api/v2/clusters/cluster-uuid/", nil); err != nil { + t.Fatalf("Do returned error: %v", err) + } + + reqs := mock.Requests() + if len(reqs) != 1 { + t.Fatalf("expected one request, got %d", len(reqs)) + } + if got := reqs[0].Headers["Authorization"]; got != "Bearer cluster-uuids-own-secret" { + t.Fatalf("authorization header = %q, want the context's bearer token", got) + } +} + func TestDoAgainstStrictSpecMockReturns400ForUnknownPath(t *testing.T) { mock := webapimock.NewSpecServerFromFile(t, "../../../shared/openapi.json", false) defer mock.Close() From 7032d8f337402004c0b67e78ee3fb9cd4700508c Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Tue, 22 Sep 2026 12:42:25 +0100 Subject: [PATCH 129/206] Revert "Authenticate cluster-scoped control-plane calls with the cluster's own secret, not the operator's SA token" This reverts commit ba8211861c73ab44e4bf841acab4adb4b663947e. --- .../controllers/cluster/controlplane.go | 10 -- .../controllers/cluster/helpers_test.go | 18 +-- .../cluster/storagecluster_controller.go | 38 +------ .../cluster/storagecluster_controller_test.go | 105 ------------------ .../cluster/storageclusterops_controller.go | 43 ------- .../storageclusterops_controller_test.go | 40 ------- operator/internal/webapi/context.go | 23 ---- operator/internal/webapi/request.go | 10 +- operator/internal/webapi/request_test.go | 31 ------ 9 files changed, 9 insertions(+), 309 deletions(-) delete mode 100644 operator/internal/webapi/context.go diff --git a/operator/internal/controllers/cluster/controlplane.go b/operator/internal/controllers/cluster/controlplane.go index 8c63eb1cc..8cef938f3 100644 --- a/operator/internal/controllers/cluster/controlplane.go +++ b/operator/internal/controllers/cluster/controlplane.go @@ -79,14 +79,6 @@ type ControlPlane interface { // outcome being asked for reached by another route, so the implementation // reads a 404 as success. CancelTask(ctx context.Context, clusterID, taskID string) error - - // Endpoint is the base URL this control plane was actually reached at, - // resolved from SIMPLYBLOCK_WEBAPI_BASE_URL when set. upsertCSICredentials - // carries it into the CSI driver's aggregate Secret, since the CSI driver - // dials it directly rather than through this reconciler: a hardcoded - // in-cluster address is wrong the moment the control plane this cluster was - // adopted from lives on a different Kubernetes cluster (a shared hub). - Endpoint() string } // httpControlPlane is the ControlPlane the operator runs with: the shared @@ -96,8 +88,6 @@ type httpControlPlane struct{ client *webapi.Client } // NewControlPlane returns the HTTP-backed control-plane surface. func NewControlPlane() ControlPlane { return &httpControlPlane{client: webapi.NewClient()} } -func (c *httpControlPlane) Endpoint() string { return c.client.BaseURL } - func (c *httpControlPlane) Ready(ctx context.Context) error { _, err := c.call(ctx, http.MethodGet, "/api/v2/_meta/ready", nil) return err diff --git a/operator/internal/controllers/cluster/helpers_test.go b/operator/internal/controllers/cluster/helpers_test.go index a8839fd47..62a99c510 100644 --- a/operator/internal/controllers/cluster/helpers_test.go +++ b/operator/internal/controllers/cluster/helpers_test.go @@ -168,11 +168,6 @@ var _ events.EventRecorder = (*recorder)(nil) type fakeControlPlane struct { t *testing.T - // endpoint is what Endpoint() returns verbatim -- a pure config read, not - // an action the test scripts, so a zero value is a legitimate "unset" - // rather than an unexpected call. - endpoint string - ready func() error create func(utils.ClusterAddParams) (webapi.ClusterResponse, error) cluster func(string) (webapi.ClusterResponse, error) @@ -188,12 +183,6 @@ type fakeControlPlane struct { tasks func(string) ([]subscriptions.TaskDTO, error) cancelTask func(string, string) error - // clusterCtx, when set, is handed the context Cluster() was called with -- - // how a test observes the bearer-token override a caller attached - // (webapi.BearerTokenFromContext), since the closures above only see the - // clusterID. - clusterCtx func(context.Context) - // The counters the write-ahead tests read. Each is the number of times the // control plane was actually asked to do something, which is the only way // to tell a step that skipped its call from one that made it twice. @@ -207,8 +196,6 @@ type fakeControlPlane struct { cancelTaskCalls int } -func (f *fakeControlPlane) Endpoint() string { return f.endpoint } - func (f *fakeControlPlane) Ready(context.Context) error { if f.ready == nil { return nil @@ -227,11 +214,8 @@ func (f *fakeControlPlane) CreateCluster( } func (f *fakeControlPlane) Cluster( - ctx context.Context, clusterID string, + _ context.Context, clusterID string, ) (webapi.ClusterResponse, error) { - if f.clusterCtx != nil { - f.clusterCtx(ctx) - } if f.cluster == nil { f.t.Fatal("the control plane was asked for a cluster and the test did not script it") } diff --git a/operator/internal/controllers/cluster/storagecluster_controller.go b/operator/internal/controllers/cluster/storagecluster_controller.go index 555eec066..131c9951d 100644 --- a/operator/internal/controllers/cluster/storagecluster_controller.go +++ b/operator/internal/controllers/cluster/storagecluster_controller.go @@ -48,7 +48,6 @@ import ( "github.com/simplyblock/simplyblock-operator/internal/cpinformer" "github.com/simplyblock/simplyblock-operator/internal/cpinformer/subscriptions" "github.com/simplyblock/simplyblock-operator/internal/utils" - "github.com/simplyblock/simplyblock-operator/internal/webapi" ) const ( @@ -529,13 +528,7 @@ func (r *StorageClusterReconciler) upgradeClaim( return adoption{}, false, nil } - // This read authenticates as the cluster itself, using the secret the - // upgrade Secret names, rather than as this operator's own Kubernetes - // identity: the control plane this cluster belongs to may run on a - // different Kubernetes cluster (a shared hub), where a Kubernetes - // TokenReview of this operator's own service-account token can never - // succeed. - found, err := r.API.Cluster(webapi.WithBearerToken(ctx, clusterSecret), uuid) + found, err := r.API.Cluster(ctx, uuid) if err != nil { // The Secret names a cluster the control plane does not have. That is // worth retrying rather than failing: the control plane may be @@ -668,8 +661,7 @@ func (r *StorageClusterReconciler) sync( ) (ctrl.Result, error) { log := logf.FromContext(ctx) - secret, err := r.clusterSecret(ctx, cluster) - if err == nil && secret != "" { + if secret, err := r.clusterSecret(ctx, cluster); err == nil && secret != "" { if err := r.upsertCSICredentials(ctx, cluster.Status.UUID, secret); err != nil { log.Error(err, "the CSI credentials entry could not be restored", "cluster", cluster.Name) @@ -677,23 +669,13 @@ func (r *StorageClusterReconciler) sync( } } - // Every read below authenticates as this cluster, using its own recorded - // secret, rather than as this operator's own Kubernetes identity: the - // control plane this cluster belongs to may run on a different Kubernetes - // cluster (a shared hub), where a Kubernetes TokenReview of this - // operator's own service-account token can never succeed. - readCtx := ctx - if secret != "" { - readCtx = webapi.WithBearerToken(ctx, secret) - } - - reading, err := r.reading(readCtx, cluster.Status.UUID) + reading, err := r.reading(ctx, cluster.Status.UUID) if err != nil { log.Error(err, "the cluster could not be read", "cluster", cluster.Name) return ctrl.Result{RequeueAfter: clusterResync}, nil } - tasks := r.readTasks(readCtx, cluster) + tasks := r.readTasks(ctx, cluster) ftt := int32(reading.MaxFaultTolerance) //nolint:gosec // a fault tolerance is a small count err = r.writeStatus(ctx, cluster, func(status *simplyblockv1alpha2.StorageClusterStatus) { @@ -834,15 +816,7 @@ func (r *StorageClusterReconciler) teardown( if cluster.Status.UUID != "" { r.closeStreams(cluster.Status.UUID) - // Authenticates as this cluster, using its own recorded secret, for - // the same reason sync() does: the control plane this cluster - // belongs to may run on a different Kubernetes cluster than this - // operator does. - deleteCtx := ctx - if secret, err := r.clusterSecret(ctx, cluster); err == nil && secret != "" { - deleteCtx = webapi.WithBearerToken(ctx, secret) - } - if err := r.API.DeleteCluster(deleteCtx, cluster.Status.UUID); err != nil { + if err := r.API.DeleteCluster(ctx, cluster.Status.UUID); err != nil { log.Error(err, "the cluster could not be deleted; retrying", "cluster", cluster.Name, "uuid", cluster.Status.UUID) return ctrl.Result{RequeueAfter: clusterRetry}, nil @@ -1065,7 +1039,7 @@ func (r *StorageClusterReconciler) upsertCSICredentials( return r.editCSICredentials(ctx, func(creds *CSICredentials) { entry := CSIClusterEntry{ ClusterID: clusterID, - ClusterEndpoint: r.API.Endpoint(), + ClusterEndpoint: utils.ENDPOINT, ClusterSecret: clusterSecret, } for i := range creds.Clusters { diff --git a/operator/internal/controllers/cluster/storagecluster_controller_test.go b/operator/internal/controllers/cluster/storagecluster_controller_test.go index 27b6c4121..0a0d0f934 100644 --- a/operator/internal/controllers/cluster/storagecluster_controller_test.go +++ b/operator/internal/controllers/cluster/storagecluster_controller_test.go @@ -15,7 +15,6 @@ package cluster import ( "context" - "encoding/json" "errors" "fmt" "net/http" @@ -131,42 +130,6 @@ func TestACreatedClusterReachesSteadyState(t *testing.T) { } } -// The CSI driver dials ClusterEndpoint directly rather than through this -// reconciler, so a control plane shared across more than one Kubernetes -// cluster -- the hub in a Ramen/DR topology -- needs that entry to carry the -// endpoint this reconciler was actually configured to reach, not a hardcoded -// address that only resolves inside its own cluster. -func TestCSICredentialsCarryTheControlPlanesConfiguredEndpoint(t *testing.T) { - const hubEndpoint = "http://simplyblock-webappapi.hub.example:31500" - api := &fakeControlPlane{ - endpoint: hubEndpoint, - create: func(utils.ClusterAddParams) (webapi.ClusterResponse, error) { - reading := activeCluster() - reading.Secret = testClusterSecret - return reading, nil - }, - cluster: func(string) (webapi.ClusterResponse, error) { return activeCluster(), nil }, - } - r := newClusterReconciler(t, api, &recorder{}, newUncreatedCluster()) - reconcileCluster(t, r, 6) - - var secret corev1.Secret - key := types.NamespacedName{Namespace: testNamespace, Name: csiCredentialsSecret} - if err := r.Get(context.Background(), key, &secret); err != nil { - t.Fatalf("read the CSI credentials secret: %v", err) - } - var creds CSICredentials - if err := json.Unmarshal(secret.Data["secret.json"], &creds); err != nil { - t.Fatalf("unmarshal secret.json: %v", err) - } - if len(creds.Clusters) != 1 { - t.Fatalf("clusters = %d entries, want 1", len(creds.Clusters)) - } - if got := creds.Clusters[0].ClusterEndpoint; got != hubEndpoint { - t.Errorf("clusterEndpoint = %q, want the configured endpoint %q", got, hubEndpoint) - } -} - // spec.deviceClass is the CRD's spelling of what sbcli's cluster-create wire // format calls `device_mode`, so the two must map onto each other rather than // the field simply passing through unmapped. @@ -300,74 +263,6 @@ func TestAnUpgradeSecretAdoptsRatherThanCreating(t *testing.T) { } } -// An adopted cluster's control plane may run on a different Kubernetes -// cluster than this operator does (a shared hub), where a Kubernetes -// TokenReview of this operator's own service-account token can never -// succeed. The read that confirms the adoption must authenticate as the -// cluster itself instead, using the secret the upgrade Secret names. -func TestAnUpgradeSecretAuthenticatesTheAdoptionReadAsTheAdoptedCluster(t *testing.T) { - var gotToken string - var gotOK bool - api := &fakeControlPlane{ - cluster: func(string) (webapi.ClusterResponse, error) { - return activeCluster(), nil - }, - clusterCtx: func(ctx context.Context) { - gotToken, gotOK = webapi.BearerTokenFromContext(ctx) - }, - } - upgrade := &corev1.Secret{ - ObjectMeta: objectMeta("simplyblock-" + testClusterName + "-upgrade"), - Data: map[string][]byte{ - "uuid": []byte(testClusterUUID), - "secret": []byte(testClusterSecret), - }, - } - r := newClusterReconciler(t, api, &recorder{}, newUncreatedCluster(), upgrade) - reconcileCluster(t, r, 6) - - if !gotOK { - t.Fatal("the adoption read carried no bearer-token override") - } - if gotToken != testClusterSecret { - t.Errorf("bearer token = %q, want the upgrade secret's own cluster secret %q", - gotToken, testClusterSecret) - } -} - -// Once a cluster is created and its secret is on record, every later -// steady-state read of it must keep authenticating as that cluster -- the -// same reasoning as the adoption read above, just for the read that runs on -// every reconcile after. -func TestStorageClusterSyncAuthenticatesAsTheClusterOnceItsSecretIsKnown(t *testing.T) { - var gotToken string - var gotOK bool - api := &fakeControlPlane{ - create: func(utils.ClusterAddParams) (webapi.ClusterResponse, error) { - reading := activeCluster() - reading.Secret = testClusterSecret - return reading, nil - }, - cluster: func(string) (webapi.ClusterResponse, error) { return activeCluster(), nil }, - clusterCtx: func(ctx context.Context) { - gotToken, gotOK = webapi.BearerTokenFromContext(ctx) - }, - } - r := newClusterReconciler(t, api, &recorder{}, newUncreatedCluster()) - // The creation machine takes 6 passes to reach steady state - // (TestACreatedClusterReachesSteadyState); one more pass is the first - // steady-state sync, which is the read this test is about. - reconcileCluster(t, r, 7) - - if !gotOK { - t.Fatal("the steady-state read carried no bearer-token override") - } - if gotToken != testClusterSecret { - t.Errorf("bearer token = %q, want the cluster's own recorded secret %q", - gotToken, testClusterSecret) - } -} - // The second route: a POST that failed against a cluster which already exists. // That covers two reconciles that both passed the claim on different // resourceVersions, and a response lost after the backend committed. diff --git a/operator/internal/controllers/cluster/storageclusterops_controller.go b/operator/internal/controllers/cluster/storageclusterops_controller.go index cebace976..85ecc9c44 100644 --- a/operator/internal/controllers/cluster/storageclusterops_controller.go +++ b/operator/internal/controllers/cluster/storageclusterops_controller.go @@ -51,7 +51,6 @@ import ( "github.com/simplyblock/simplyblock-operator/internal/cpinformer" "github.com/simplyblock/simplyblock-operator/internal/cpinformer/subscriptions" "github.com/simplyblock/simplyblock-operator/internal/utils" - "github.com/simplyblock/simplyblock-operator/internal/webapi" ) const ( @@ -223,14 +222,6 @@ func (r *StorageClusterOpsReconciler) Reconcile( func (r *StorageClusterOpsReconciler) advance( ctx context.Context, ops *simplyblockv1alpha2.StorageClusterOps, ) (ctrl.Result, error) { - // Every control-plane call this step and everything downstream of it makes - // (perform, advanceWalk, and everything under them) authenticates as this - // operation's own cluster when its secret is known, rather than as this - // operator's Kubernetes identity -- the only way to reach a control plane a - // different Kubernetes cluster runs (a shared hub), since a Kubernetes - // TokenReview can never cross a cluster boundary. - ctx = r.authenticatedContext(ctx, ops) - graph := action(ops.Spec.Action) machine, err := graphs().FromSnapshot(ctx, graph, statemachine.FromKube[step](ops.Status.Step)) @@ -877,40 +868,6 @@ func (r *StorageClusterOpsReconciler) clusterReading( }, nil } -// clusterSecret reads the secret StorageClusterReconciler.persist wrote for -// this operation's cluster, keyed by the StorageCluster's Kubernetes name (not -// its backend UUID, which this reconciler is not always given yet at the point -// it needs the credential). It reports the empty string when there is none. -func (r *StorageClusterOpsReconciler) clusterSecret( - ctx context.Context, ops *simplyblockv1alpha2.StorageClusterOps, -) (string, error) { - var secret corev1.Secret - key := types.NamespacedName{ - Name: fmt.Sprintf("simplyblock-cluster-%s", ops.Spec.ClusterRef), - Namespace: ops.Namespace, - } - if err := r.Get(ctx, key, &secret); err != nil { - return "", err - } - return string(secret.Data["secret"]), nil -} - -// authenticatedContext attaches this operation's cluster's own credential to -// ctx when one is known, so every control-plane call the operation makes -// authenticates as that cluster instead of as this operator's Kubernetes -// identity -- the only way to reach a control plane a different Kubernetes -// cluster runs (a shared hub), since a Kubernetes TokenReview can never cross a -// cluster boundary. See StorageClusterReconciler's identically-named method. -func (r *StorageClusterOpsReconciler) authenticatedContext( - ctx context.Context, ops *simplyblockv1alpha2.StorageClusterOps, -) context.Context { - secret, err := r.clusterSecret(ctx, ops) - if err != nil || secret == "" { - return ctx - } - return webapi.WithBearerToken(ctx, secret) -} - // effectiveConcurrentRestarts is min(specVal, FTT), defaulting to 1 when the // spec says nothing. Both inputs may be absent, because the fault tolerance // comes from the control plane. diff --git a/operator/internal/controllers/cluster/storageclusterops_controller_test.go b/operator/internal/controllers/cluster/storageclusterops_controller_test.go index e6db399ef..8194c402d 100644 --- a/operator/internal/controllers/cluster/storageclusterops_controller_test.go +++ b/operator/internal/controllers/cluster/storageclusterops_controller_test.go @@ -21,7 +21,6 @@ import ( "testing" "time" - corev1 "k8s.io/api/core/v1" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" "k8s.io/apimachinery/pkg/types" ctrl "sigs.k8s.io/controller-runtime" @@ -266,45 +265,6 @@ func TestAShutdownIssuesOneCallAndWaitsForTheCluster(t *testing.T) { } } -// An operation against an already-adopted cluster may reach a control plane on -// a different Kubernetes cluster (a shared hub), the same as -// StorageClusterReconciler: it authenticates as the cluster's own recorded -// secret instead of as this operator's Kubernetes identity, since a Kubernetes -// TokenReview can never cross a cluster boundary. -func TestAnOperationAuthenticatesAsItsClusterOnceItsSecretIsKnown(t *testing.T) { - var gotToken string - var gotOK bool - active := true - api := &fakeControlPlane{ - cluster: func(string) (webapi.ClusterResponse, error) { - reading := activeCluster() - if !active { - reading.Status = "suspended" - } - return reading, nil - }, - clusterCtx: func(ctx context.Context) { - gotToken, gotOK = webapi.BearerTokenFromContext(ctx) - }, - shutdown: func(string) error { active = false; return nil }, - } - secret := &corev1.Secret{ - ObjectMeta: objectMeta("simplyblock-cluster-" + testClusterName), - Data: map[string][]byte{"secret": []byte(testClusterSecret)}, - } - r := newOpsReconciler(t, api, &recorder{}, - newTestCluster(), newTestOps(simplyblockv1alpha2.StorageClusterOpsActionShutdown), secret) - - reconcileOps(t, r, 6) - - if !gotOK { - t.Fatal("the operation's control-plane read carried no bearer-token override") - } - if gotToken != testClusterSecret { - t.Errorf("bearer token = %q, want the cluster's own recorded secret %q", gotToken, testClusterSecret) - } -} - // Restart is the one action with two side effects, because the control plane // has no restart endpoint of its own. func TestARestartShutsDownThenStarts(t *testing.T) { diff --git a/operator/internal/webapi/context.go b/operator/internal/webapi/context.go deleted file mode 100644 index 7a0126cd3..000000000 --- a/operator/internal/webapi/context.go +++ /dev/null @@ -1,23 +0,0 @@ -package webapi - -import "context" - -type bearerTokenKey struct{} - -// WithBearerToken attaches a bearer credential to ctx that Do and -// DoWithHeaders send instead of the client's own service-account token. It is -// how a call scoped to one cluster authenticates as that cluster rather than -// as this process's own Kubernetes identity -- the only way to reach a -// control plane a different Kubernetes cluster runs (a shared hub in a -// multi-cluster deployment), since a Kubernetes TokenReview can never cross a -// cluster boundary. -func WithBearerToken(ctx context.Context, token string) context.Context { - return context.WithValue(ctx, bearerTokenKey{}, token) -} - -// BearerTokenFromContext returns the token WithBearerToken attached, and -// whether one was. -func BearerTokenFromContext(ctx context.Context) (string, bool) { - token, ok := ctx.Value(bearerTokenKey{}).(string) - return token, ok -} diff --git a/operator/internal/webapi/request.go b/operator/internal/webapi/request.go index 59f2e3729..7734cebf2 100644 --- a/operator/internal/webapi/request.go +++ b/operator/internal/webapi/request.go @@ -51,14 +51,8 @@ func (c *Client) DoWithHeaders( return nil, nil, 0, fmt.Errorf("create request: %w", err) } - // Attach auth header. A token WithBearerToken attached to ctx wins over - // this process's own service-account token, for a call scoped to a - // cluster whose control plane lives on a different Kubernetes cluster. - token := c.saToken - if override, ok := BearerTokenFromContext(ctx); ok { - token = override - } - req.Header.Set("Authorization", fmt.Sprintf("Bearer %s", token)) + // Attach auth header + req.Header.Set("Authorization", fmt.Sprintf("Bearer %s", c.saToken)) req.Header.Set("Content-Type", "application/json") // Execute the request diff --git a/operator/internal/webapi/request_test.go b/operator/internal/webapi/request_test.go index 5a3da827c..01406ea25 100644 --- a/operator/internal/webapi/request_test.go +++ b/operator/internal/webapi/request_test.go @@ -73,37 +73,6 @@ func TestDoAgainstSpecMockSendsHeadersBodyAndReturnsResponse(t *testing.T) { } } -// A call scoped to one cluster authenticates as that cluster when -// WithBearerToken names one, instead of as this process's own service -// account -- the only way to reach a control plane a different Kubernetes -// cluster runs, since a TokenReview can never cross a cluster boundary. -func TestDoSendsTheBearerTokenAttachedToContextInsteadOfTheServiceAccountToken(t *testing.T) { - mock := webapimock.NewSpecServerFromFile(t, "../../../shared/openapi.json", false) - defer mock.Close() - - mock.Register( - http.MethodGet, - "/api/v2/clusters/cluster-uuid/", - webapimock.RouteResponse{Status: http.StatusOK, Body: `{}`}, - ) - - c := NewClient(mock.URL()) - c.saToken = "operators-own-service-account-token" - - ctx := WithBearerToken(context.Background(), "cluster-uuids-own-secret") - if _, _, err := c.Do(ctx, http.MethodGet, "/api/v2/clusters/cluster-uuid/", nil); err != nil { - t.Fatalf("Do returned error: %v", err) - } - - reqs := mock.Requests() - if len(reqs) != 1 { - t.Fatalf("expected one request, got %d", len(reqs)) - } - if got := reqs[0].Headers["Authorization"]; got != "Bearer cluster-uuids-own-secret" { - t.Fatalf("authorization header = %q, want the context's bearer token", got) - } -} - func TestDoAgainstStrictSpecMockReturns400ForUnknownPath(t *testing.T) { mock := webapimock.NewSpecServerFromFile(t, "../../../shared/openapi.json", false) defer mock.Close() From 38100681beb2e405033470030f3f4dd9f5b7ea7d Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Tue, 22 Sep 2026 12:56:10 +0100 Subject: [PATCH 130/206] fixed boolean flag csiLinkEnabled --- .../storage.simplyblock.io_operatorops.yaml | 24 ++++++++++- operator/cmd/main.go | 43 ++++++++++--------- 2 files changed, 46 insertions(+), 21 deletions(-) diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_operatorops.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_operatorops.yaml index e4c509e08..ea81f93a0 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_operatorops.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_operatorops.yaml @@ -210,9 +210,31 @@ spec: type: string description: |- NodeSelector restricts which workers are inspected. Empty inspects every - schedulable worker. + schedulable worker, and it is exclusive with Workers. type: object + workers: + description: |- + Workers are the workers to inspect, by node name. + + It is the answer to inspecting two named machines, which a label selector + can only express by labeling them first: a selector's entries are ANDed, + so two hostnames in one selector match nothing at all. + + A named worker that does not exist, or that is declined for one of the + reasons any worker is declined, is reported by the same event the selector + path reports it by. Naming a worker is a statement about which machines to + consider, not a claim that each of them will be used. + items: + maxLength: 253 + type: string + maxItems: 128 + type: array type: object + x-kubernetes-validations: + - message: spec.discover names workers and also carries a nodeSelector; + state one or the other + rule: '!(has(self.workers) && size(self.workers) > 0 && has(self.nodeSelector) + && size(self.nodeSelector) > 0)' required: - action type: object diff --git a/operator/cmd/main.go b/operator/cmd/main.go index 371127c6f..65d050f1c 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -171,6 +171,7 @@ func main() { "be carried across.") var csiLinkEnabled bool var csiLinkAddr, csiLinkCertPath, csiLinkCertName, csiLinkCertKey, csiLinkAudience string + flag.BoolVar(&csiLinkEnabled, "csi-link", false, "Serve the CSI link.") flag.StringVar(&csiLinkAddr, "csi-link-bind-address", ":9500", "The address the CSI link endpoint binds to.") flag.StringVar(&csiLinkCertPath, "csi-link-cert-path", "", @@ -333,27 +334,29 @@ func main() { // plugins; a reconciler reaching a node goes through it, and treats // link.ErrNoSession as a requeue rather than a failure. // - // Always served, because both plugins always dial it. TLS when a - // certificate is configured, plaintext when none is. - var certFile, keyFile string - if csiLinkCertPath != "" { - certFile = filepath.Join(csiLinkCertPath, csiLinkCertName) - keyFile = filepath.Join(csiLinkCertPath, csiLinkCertKey) - } - csiPeers, err := csilink.Setup(mgr, csilink.Config{ - BindAddress: csiLinkAddr, - CertFile: certFile, - KeyFile: keyFile, - Namespace: operatorNamespace, - Audiences: []string{csiLinkAudience}, - NodeServiceAccount: "simplyblock-csi-node-sa", - ControllerServiceAccount: "simplyblock-csi-controller-sa", - }) - if err != nil { - setupLog.Error(err, "unable to set up the CSI link") - os.Exit(1) + // Off by default, on with --csi-link. TLS when a certificate is + // configured, plaintext when none is. + if csiLinkEnabled { + var certFile, keyFile string + if csiLinkCertPath != "" { + certFile = filepath.Join(csiLinkCertPath, csiLinkCertName) + keyFile = filepath.Join(csiLinkCertPath, csiLinkCertKey) + } + csiPeers, err := csilink.Setup(mgr, csilink.Config{ + BindAddress: csiLinkAddr, + CertFile: certFile, + KeyFile: keyFile, + Namespace: operatorNamespace, + Audiences: []string{csiLinkAudience}, + NodeServiceAccount: "simplyblock-csi-node-sa", + ControllerServiceAccount: "simplyblock-csi-controller-sa", + }) + if err != nil { + setupLog.Error(err, "unable to set up the CSI link") + os.Exit(1) + } + _ = csiPeers // handed to reconcilers as they start using it } - _ = csiPeers // handed to reconcilers as they start using it // Control-plane SSE push subscriptions: one leader-only manager, streams // driven by scopes that reconcilers register (the StorageNode controller adds From b97957720a8892aae130136fb2f1a43e325fd6da Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Tue, 22 Sep 2026 14:01:15 +0100 Subject: [PATCH 131/206] Extend cross-cluster auth to StorageNode and StorageNodeOps control-plane calls --- .../controllers/cluster/controlplane.go | 10 ++ .../controllers/cluster/helpers_test.go | 18 ++- .../cluster/storagecluster_controller.go | 39 ++++- .../cluster/storagecluster_controller_test.go | 133 ++++++++++++++++++ .../cluster/storageclusterops_controller.go | 45 ++++++ .../storageclusterops_controller_test.go | 55 ++++++++ operator/internal/controllers/node/auth.go | 65 +++++++++ .../internal/controllers/node/auth_test.go | 80 +++++++++++ .../controllers/node/fixtures_test.go | 65 +++++---- .../node/storagenode_controller.go | 12 ++ .../node/storagenodeops_controller.go | 9 ++ operator/internal/webapi/context.go | 22 +++ operator/internal/webapi/request.go | 10 +- operator/internal/webapi/request_test.go | 31 ++++ 14 files changed, 560 insertions(+), 34 deletions(-) create mode 100644 operator/internal/controllers/node/auth.go create mode 100644 operator/internal/controllers/node/auth_test.go create mode 100644 operator/internal/webapi/context.go diff --git a/operator/internal/controllers/cluster/controlplane.go b/operator/internal/controllers/cluster/controlplane.go index 0c0b79179..e84d07373 100644 --- a/operator/internal/controllers/cluster/controlplane.go +++ b/operator/internal/controllers/cluster/controlplane.go @@ -81,6 +81,14 @@ type ControlPlane interface { // outcome being asked for reached by another route, so the implementation // reads a 404 as success. CancelTask(ctx context.Context, clusterID, taskID string) error + + // Endpoint is the base URL this call would currently reach the control + // plane at, resolved the same way every other call resolves it (§3.3). + // upsertCSICredentials carries it into the CSI driver's aggregate Secret, + // since the CSI driver dials it directly rather than through this + // reconciler, and a hardcoded in-cluster address is wrong the moment the + // control plane is a ControlPlane.spec.source.managed one. + Endpoint(ctx context.Context) string } // httpControlPlane is the ControlPlane the operator runs with: one method per @@ -107,6 +115,8 @@ func NewControlPlane(resolve controlplane.EndpointResolver) ControlPlane { return &httpControlPlane{client: webapi.NewClient(), resolve: resolve} } +func (c *httpControlPlane) Endpoint(ctx context.Context) string { return c.clientFor(ctx).BaseURL } + func (c *httpControlPlane) Ready(ctx context.Context) error { _, err := c.call(ctx, http.MethodGet, "/api/v2/_meta/ready", nil) return err diff --git a/operator/internal/controllers/cluster/helpers_test.go b/operator/internal/controllers/cluster/helpers_test.go index ad1cfb7d2..402cff295 100644 --- a/operator/internal/controllers/cluster/helpers_test.go +++ b/operator/internal/controllers/cluster/helpers_test.go @@ -198,6 +198,17 @@ var _ events.EventRecorder = (*recorder)(nil) type fakeControlPlane struct { t *testing.T + // endpoint is what Endpoint() returns verbatim -- a pure config read, not + // an action the test scripts, so a zero value is a legitimate "unset" + // rather than an unexpected call. + endpoint string + + // clusterCtx, when set, is handed the context Cluster() was called with -- + // how a test observes a bearer-token override a caller attached + // (webapi.BearerTokenFromContext), since the closures below only see the + // clusterID. + clusterCtx func(context.Context) + ready func() error create func(utils.ClusterAddParams) (webapi.ClusterResponse, error) cluster func(string) (webapi.ClusterResponse, error) @@ -226,6 +237,8 @@ type fakeControlPlane struct { cancelTaskCalls int } +func (f *fakeControlPlane) Endpoint(context.Context) string { return f.endpoint } + func (f *fakeControlPlane) Ready(context.Context) error { if f.ready == nil { return nil @@ -244,8 +257,11 @@ func (f *fakeControlPlane) CreateCluster( } func (f *fakeControlPlane) Cluster( - _ context.Context, clusterID string, + ctx context.Context, clusterID string, ) (webapi.ClusterResponse, error) { + if f.clusterCtx != nil { + f.clusterCtx(ctx) + } if f.cluster == nil { f.t.Fatal("the control plane was asked for a cluster and the test did not script it") } diff --git a/operator/internal/controllers/cluster/storagecluster_controller.go b/operator/internal/controllers/cluster/storagecluster_controller.go index 0120f4b0c..6ddec0fae 100644 --- a/operator/internal/controllers/cluster/storagecluster_controller.go +++ b/operator/internal/controllers/cluster/storagecluster_controller.go @@ -46,6 +46,7 @@ import ( "github.com/simplyblock/simplyblock-operator/internal/cpinformer" "github.com/simplyblock/simplyblock-operator/internal/cpinformer/subscriptions" "github.com/simplyblock/simplyblock-operator/internal/utils" + "github.com/simplyblock/simplyblock-operator/internal/webapi" ) const ( @@ -530,7 +531,13 @@ func (r *StorageClusterReconciler) upgradeClaim( return adoption{}, false, nil } - found, err := r.API.Cluster(ctx, uuid) + // This read authenticates as the cluster itself, using the secret the + // upgrade Secret names, rather than as this operator's own Kubernetes + // identity: the control plane this cluster belongs to may be a + // ControlPlane.spec.source.managed one, on a different Kubernetes + // cluster, where a Kubernetes TokenReview of this operator's own + // service-account token can never succeed. + found, err := r.API.Cluster(webapi.WithBearerToken(ctx, clusterSecret), uuid) if err != nil { // The Secret names a cluster the control plane does not have. That is // worth retrying rather than failing: the control plane may be @@ -663,7 +670,8 @@ func (r *StorageClusterReconciler) sync( ) (ctrl.Result, error) { log := logf.FromContext(ctx) - if secret, err := r.clusterSecret(ctx, cluster); err == nil && secret != "" { + secret, err := r.clusterSecret(ctx, cluster) + if err == nil && secret != "" { if err := r.upsertCSICredentials(ctx, cluster.Status.UUID, secret); err != nil { log.Error(err, "the CSI credentials entry could not be restored", "cluster", cluster.Name) @@ -671,13 +679,24 @@ func (r *StorageClusterReconciler) sync( } } - reading, err := r.reading(ctx, cluster.Status.UUID) + // Every read below authenticates as this cluster, using its own recorded + // secret, rather than as this operator's own Kubernetes identity: the + // control plane this cluster belongs to may be a + // ControlPlane.spec.source.managed one, on a different Kubernetes + // cluster, where a Kubernetes TokenReview of this operator's own + // service-account token can never succeed. + readCtx := ctx + if secret != "" { + readCtx = webapi.WithBearerToken(ctx, secret) + } + + reading, err := r.reading(readCtx, cluster.Status.UUID) if err != nil { log.Error(err, "the cluster could not be read", "cluster", cluster.Name) return ctrl.Result{RequeueAfter: clusterResync}, nil } - tasks := r.readTasks(ctx, cluster) + tasks := r.readTasks(readCtx, cluster) ftt := int32(reading.MaxFaultTolerance) //nolint:gosec // a fault tolerance is a small count err = r.writeStatus(ctx, cluster, func(status *simplyblockv1alpha2.StorageClusterStatus) { @@ -871,7 +890,15 @@ func (r *StorageClusterReconciler) teardown( if cluster.Status.UUID != "" { r.closeStreams(cluster.Status.UUID) - if err := r.API.DeleteCluster(ctx, cluster.Status.UUID); err != nil { + // Authenticates as this cluster, using its own recorded secret, for + // the same reason sync() does: the control plane this cluster + // belongs to may be a ControlPlane.spec.source.managed one, on a + // different Kubernetes cluster than this operator. + deleteCtx := ctx + if secret, err := r.clusterSecret(ctx, cluster); err == nil && secret != "" { + deleteCtx = webapi.WithBearerToken(ctx, secret) + } + if err := r.API.DeleteCluster(deleteCtx, cluster.Status.UUID); err != nil { log.Error(err, "the cluster could not be deleted; retrying", "cluster", cluster.Name, "uuid", cluster.Status.UUID) return ctrl.Result{RequeueAfter: clusterRetry}, nil @@ -1094,7 +1121,7 @@ func (r *StorageClusterReconciler) upsertCSICredentials( return r.editCSICredentials(ctx, func(creds *CSICredentials) { entry := CSIClusterEntry{ ClusterID: clusterID, - ClusterEndpoint: utils.ENDPOINT, + ClusterEndpoint: r.API.Endpoint(ctx), ClusterSecret: clusterSecret, } for i := range creds.Clusters { diff --git a/operator/internal/controllers/cluster/storagecluster_controller_test.go b/operator/internal/controllers/cluster/storagecluster_controller_test.go index 96c5f85a4..299a019c7 100644 --- a/operator/internal/controllers/cluster/storagecluster_controller_test.go +++ b/operator/internal/controllers/cluster/storagecluster_controller_test.go @@ -15,6 +15,7 @@ package cluster import ( "context" + "encoding/json" "errors" "fmt" "net/http" @@ -130,6 +131,42 @@ func TestACreatedClusterReachesSteadyState(t *testing.T) { } } +// The CSI driver dials ClusterEndpoint directly rather than through this +// reconciler, so a control plane a ControlPlane.spec.source.managed points at +// elsewhere needs that entry to carry the endpoint this reconciler actually +// resolved, not a hardcoded address that only resolves inside its own +// cluster. +func TestCSICredentialsCarryTheControlPlanesResolvedEndpoint(t *testing.T) { + const hubEndpoint = "http://simplyblock-webappapi.hub.example:31500" + api := &fakeControlPlane{ + endpoint: hubEndpoint, + create: func(utils.ClusterAddParams) (webapi.ClusterResponse, error) { + reading := activeCluster() + reading.Secret = testClusterSecret + return reading, nil + }, + cluster: func(string) (webapi.ClusterResponse, error) { return activeCluster(), nil }, + } + r := newClusterReconciler(t, api, &recorder{}, newUncreatedCluster()) + reconcileCluster(t, r, 6) + + var secret corev1.Secret + key := types.NamespacedName{Namespace: testNamespace, Name: csiCredentialsSecret} + if err := r.Get(context.Background(), key, &secret); err != nil { + t.Fatalf("read the CSI credentials secret: %v", err) + } + var creds CSICredentials + if err := json.Unmarshal(secret.Data["secret.json"], &creds); err != nil { + t.Fatalf("unmarshal secret.json: %v", err) + } + if len(creds.Clusters) != 1 { + t.Fatalf("clusters = %d entries, want 1", len(creds.Clusters)) + } + if got := creds.Clusters[0].ClusterEndpoint; got != hubEndpoint { + t.Errorf("clusterEndpoint = %q, want the resolved endpoint %q", got, hubEndpoint) + } +} + // spec.deviceClass is the CRD's spelling of what sbcli's cluster-create wire // format calls `device_mode`, so the two must map onto each other rather than // the field simply passing through unmapped. @@ -263,6 +300,102 @@ func TestAnUpgradeSecretAdoptsRatherThanCreating(t *testing.T) { } } +// An adopted cluster's control plane may be a ControlPlane.spec.source.managed +// one, on a different Kubernetes cluster, where a Kubernetes TokenReview of +// this operator's own service-account token can never succeed. The read that +// confirms the adoption must authenticate as the cluster itself instead, +// using the secret the upgrade Secret names. +func TestAnUpgradeSecretAuthenticatesTheAdoptionReadAsTheAdoptedCluster(t *testing.T) { + // Every Cluster() call across the run is checked, not just the last: a + // later, correctly-authenticated steady-state read would otherwise + // overwrite a single captured value and hide a regression in the + // adoption read specifically. + var calls []struct { + token string + ok bool + } + api := &fakeControlPlane{ + cluster: func(string) (webapi.ClusterResponse, error) { + return activeCluster(), nil + }, + clusterCtx: func(ctx context.Context) { + token, ok := webapi.BearerTokenFromContext(ctx) + calls = append(calls, struct { + token string + ok bool + }{token, ok}) + }, + } + upgrade := &corev1.Secret{ + ObjectMeta: objectMeta("simplyblock-" + testClusterName + "-upgrade"), + Data: map[string][]byte{ + "uuid": []byte(testClusterUUID), + "secret": []byte(testClusterSecret), + }, + } + r := newClusterReconciler(t, api, &recorder{}, newUncreatedCluster(), upgrade) + reconcileCluster(t, r, 6) + + if len(calls) == 0 { + t.Fatal("the control plane was never asked for the cluster") + } + for i, call := range calls { + if !call.ok { + t.Errorf("call %d: carried no bearer-token override", i) + continue + } + if call.token != testClusterSecret { + t.Errorf("call %d: bearer token = %q, want the cluster's own secret %q", + i, call.token, testClusterSecret) + } + } +} + +// Once a cluster is created and its secret is on record, every later +// steady-state read of it must keep authenticating as that cluster -- the +// same reasoning as the adoption read above, just for the read that runs on +// every reconcile after. +func TestStorageClusterSyncAuthenticatesAsTheClusterOnceItsSecretIsKnown(t *testing.T) { + var calls []struct { + token string + ok bool + } + api := &fakeControlPlane{ + create: func(utils.ClusterAddParams) (webapi.ClusterResponse, error) { + reading := activeCluster() + reading.Secret = testClusterSecret + return reading, nil + }, + cluster: func(string) (webapi.ClusterResponse, error) { return activeCluster(), nil }, + clusterCtx: func(ctx context.Context) { + token, ok := webapi.BearerTokenFromContext(ctx) + calls = append(calls, struct { + token string + ok bool + }{token, ok}) + }, + } + r := newClusterReconciler(t, api, &recorder{}, newUncreatedCluster()) + // The creation machine takes 6 passes to reach steady state + // (TestACreatedClusterReachesSteadyState); one more pass is the first + // steady-state sync, which is the read this test is about. + reconcileCluster(t, r, 7) + + if len(calls) == 0 { + t.Fatal("the control plane was never asked for the cluster") + } + for i, call := range calls { + if !call.ok { + t.Errorf("call %d: carried no bearer-token override", i) + continue + } + if call.token != testClusterSecret { + t.Errorf("call %d: bearer token = %q, want the cluster's own recorded secret %q", + i, call.token, testClusterSecret) + } + } +} + // The second route: a POST that failed against a cluster which already exists. // That covers two reconciles that both passed the claim on different // resourceVersions, and a response lost after the backend committed. diff --git a/operator/internal/controllers/cluster/storageclusterops_controller.go b/operator/internal/controllers/cluster/storageclusterops_controller.go index ed12cf8f1..09283c7f6 100644 --- a/operator/internal/controllers/cluster/storageclusterops_controller.go +++ b/operator/internal/controllers/cluster/storageclusterops_controller.go @@ -51,6 +51,7 @@ import ( "github.com/simplyblock/simplyblock-operator/internal/cpinformer" "github.com/simplyblock/simplyblock-operator/internal/cpinformer/subscriptions" "github.com/simplyblock/simplyblock-operator/internal/utils" + "github.com/simplyblock/simplyblock-operator/internal/webapi" ) const ( @@ -237,6 +238,15 @@ func (r *StorageClusterOpsReconciler) Reconcile( func (r *StorageClusterOpsReconciler) advance( ctx context.Context, ops *simplyblockv1alpha2.StorageClusterOps, ) (ctrl.Result, error) { + // Every control-plane call this step and everything downstream of it + // makes (perform, advanceWalk, and everything under them) authenticates + // as this operation's own cluster when its secret is known, rather than + // as this operator's Kubernetes identity -- the only way to reach a + // control plane a different Kubernetes cluster runs (a + // ControlPlane.spec.source.managed one), since a Kubernetes TokenReview + // can never cross a cluster boundary. + ctx = r.authenticatedContext(ctx, ops) + graph := action(ops.Spec.Action) machine, err := graphs().FromSnapshot(ctx, graph, statemachine.FromKube[step](ops.Status.Step)) @@ -920,6 +930,41 @@ func (r *StorageClusterOpsReconciler) clusterReading( }, nil } +// clusterSecret reads the secret StorageClusterReconciler.persist wrote for +// this operation's cluster, keyed by the StorageCluster's Kubernetes name (not +// its backend UUID, which this reconciler is not always given yet at the point +// it needs the credential). It reports the empty string when there is none. +func (r *StorageClusterOpsReconciler) clusterSecret( + ctx context.Context, ops *simplyblockv1alpha2.StorageClusterOps, +) (string, error) { + var secret corev1.Secret + key := types.NamespacedName{ + Name: fmt.Sprintf("simplyblock-cluster-%s", ops.Spec.ClusterRef), + Namespace: ops.Namespace, + } + if err := r.Get(ctx, key, &secret); err != nil { + return "", err + } + return string(secret.Data["secret"]), nil +} + +// authenticatedContext attaches this operation's cluster's own credential to +// ctx when one is known, so every control-plane call the operation makes +// authenticates as that cluster instead of as this operator's Kubernetes +// identity -- the only way to reach a control plane a different Kubernetes +// cluster runs (a ControlPlane.spec.source.managed one), since a Kubernetes +// TokenReview can never cross a cluster boundary. See StorageClusterReconciler's +// identically-named method. +func (r *StorageClusterOpsReconciler) authenticatedContext( + ctx context.Context, ops *simplyblockv1alpha2.StorageClusterOps, +) context.Context { + secret, err := r.clusterSecret(ctx, ops) + if err != nil || secret == "" { + return ctx + } + return webapi.WithBearerToken(ctx, secret) +} + // effectiveConcurrentRestarts is min(specVal, FTT), defaulting to 1 when the // spec says nothing. Both inputs may be absent, because the fault tolerance // comes from the control plane. diff --git a/operator/internal/controllers/cluster/storageclusterops_controller_test.go b/operator/internal/controllers/cluster/storageclusterops_controller_test.go index 8194c402d..7ee8714e5 100644 --- a/operator/internal/controllers/cluster/storageclusterops_controller_test.go +++ b/operator/internal/controllers/cluster/storageclusterops_controller_test.go @@ -21,6 +21,7 @@ import ( "testing" "time" + corev1 "k8s.io/api/core/v1" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" "k8s.io/apimachinery/pkg/types" ctrl "sigs.k8s.io/controller-runtime" @@ -265,6 +266,60 @@ func TestAShutdownIssuesOneCallAndWaitsForTheCluster(t *testing.T) { } } +// An operation against an already-adopted cluster may reach a control plane +// on a different Kubernetes cluster (a ControlPlane.spec.source.managed one), +// the same as StorageClusterReconciler: it authenticates as the cluster's own +// recorded secret instead of as this operator's Kubernetes identity, since a +// Kubernetes TokenReview can never cross a cluster boundary. Every Cluster() +// call across the run is checked, not just the last, so a later correctly- +// authenticated read can't hide a regression in an earlier one. +func TestAnOperationAuthenticatesAsItsClusterOnceItsSecretIsKnown(t *testing.T) { + var calls []struct { + token string + ok bool + } + active := true + api := &fakeControlPlane{ + cluster: func(string) (webapi.ClusterResponse, error) { + reading := activeCluster() + if !active { + reading.Status = "suspended" + } + return reading, nil + }, + clusterCtx: func(ctx context.Context) { + token, ok := webapi.BearerTokenFromContext(ctx) + calls = append(calls, struct { + token string + ok bool + }{token, ok}) + }, + shutdown: func(string) error { active = false; return nil }, + } + secret := &corev1.Secret{ + ObjectMeta: objectMeta("simplyblock-cluster-" + testClusterName), + Data: map[string][]byte{"secret": []byte(testClusterSecret)}, + } + r := newOpsReconciler(t, api, &recorder{}, + newTestCluster(), newTestOps(simplyblockv1alpha2.StorageClusterOpsActionShutdown), secret) + + reconcileOps(t, r, 6) + + if len(calls) == 0 { + t.Fatal("the control plane was never asked for the cluster") + } + for i, call := range calls { + if !call.ok { + t.Errorf("call %d: carried no bearer-token override", i) + continue + } + if call.token != testClusterSecret { + t.Errorf("call %d: bearer token = %q, want the cluster's own recorded secret %q", + i, call.token, testClusterSecret) + } + } +} + // Restart is the one action with two side effects, because the control plane // has no restart endpoint of its own. func TestARestartShutsDownThenStarts(t *testing.T) { diff --git a/operator/internal/controllers/node/auth.go b/operator/internal/controllers/node/auth.go new file mode 100644 index 000000000..87dbb58d8 --- /dev/null +++ b/operator/internal/controllers/node/auth.go @@ -0,0 +1,65 @@ +// Cross-cluster authentication for this package's two reconcilers. +// +// A node's or node operation's control plane may be a +// ControlPlane.spec.source.managed one, on a different Kubernetes cluster +// than this operator, where a Kubernetes TokenReview of this operator's own +// service-account token can never succeed. What does cross that boundary is a +// cluster's own backend secret -- a plain credential the control plane's +// database matches by value, not a Kubernetes identity -- so every call +// scoped to an already-adopted cluster authenticates with that instead, once +// it is known. See internal/controllers/cluster's identically-motivated fix. + +package node + +import ( + "context" + "fmt" + + corev1 "k8s.io/api/core/v1" + "k8s.io/apimachinery/pkg/types" + "sigs.k8s.io/controller-runtime/pkg/client" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/webapi" +) + +// clusterSecretByName reads the secret StorageClusterReconciler.persist wrote +// for the named StorageCluster, and reports the empty string when there is +// none. +func clusterSecretByName( + ctx context.Context, c client.Client, namespace, clusterName string, +) (string, error) { + var secret corev1.Secret + key := types.NamespacedName{ + Name: fmt.Sprintf("simplyblock-cluster-%s", clusterName), + Namespace: namespace, + } + if err := c.Get(ctx, key, &secret); err != nil { + return "", err + } + return string(secret.Data["secret"]), nil +} + +// clusterSecretForNode reads the same secret for the cluster a named +// StorageNode belongs to, for a caller that only has the node's name (a +// StorageNodeOps names its target node, not the node's cluster directly). +func clusterSecretForNode( + ctx context.Context, c client.Client, namespace, nodeRef string, +) (string, error) { + var node simplyblockv1alpha2.StorageNode + key := types.NamespacedName{Name: nodeRef, Namespace: namespace} + if err := c.Get(ctx, key, &node); err != nil { + return "", err + } + return clusterSecretByName(ctx, c, namespace, node.Spec.ClusterRef) +} + +// authenticatedContext attaches a cluster's own credential to ctx when one is +// known, so a call scoped to it authenticates as that cluster instead of as +// this operator's Kubernetes identity. +func authenticatedContext(ctx context.Context, secret string, err error) context.Context { + if err != nil || secret == "" { + return ctx + } + return webapi.WithBearerToken(ctx, secret) +} diff --git a/operator/internal/controllers/node/auth_test.go b/operator/internal/controllers/node/auth_test.go new file mode 100644 index 000000000..72c37a606 --- /dev/null +++ b/operator/internal/controllers/node/auth_test.go @@ -0,0 +1,80 @@ +// A node's or node operation's control plane may be a +// ControlPlane.spec.source.managed one, on a different Kubernetes cluster +// than this operator, where a Kubernetes TokenReview of this operator's own +// service-account token can never succeed. Every call scoped to an +// already-adopted cluster must authenticate as that cluster's own recorded +// secret instead. See internal/controllers/cluster's identically-motivated +// tests. + +package node + +import ( + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// authTestClusterSecret is the cluster's own recorded credential, the one +// every assertion below wants to see instead of an operator SA token. +const authTestClusterSecret = "the-clusters-own-secret" + +// aClusterSecret is the Secret StorageClusterReconciler.persist writes for +// anOpsCluster(), keyed by its Kubernetes name. +func aClusterSecret() *corev1.Secret { + return &corev1.Secret{ + ObjectMeta: metav1.ObjectMeta{ + Name: "simplyblock-cluster-" + opsCluster, + Namespace: opsNamespace, + }, + Data: map[string][]byte{"secret": []byte(authTestClusterSecret)}, + } +} + +// checkBearerTokens fails the test on any call the control plane recorded +// that did not carry authTestClusterSecret as its bearer-token override. +func checkBearerTokens(t *testing.T, api *scriptedControlPlane) { + t.Helper() + if len(api.calls) == 0 { + t.Fatal("the control plane was never asked anything") + } + for i, obs := range api.bearerTokens { + if !obs.ok { + t.Errorf("call %d (%s): carried no bearer-token override", i, api.calls[i]) + continue + } + if obs.token != authTestClusterSecret { + t.Errorf("call %d (%s): bearer token = %q, want the cluster's own secret %q", + i, api.calls[i], obs.token, authTestClusterSecret) + } + } +} + +// The entity reconciler's steady-state read (StorageNode()) authenticates as +// the node's cluster once that cluster's secret is on record. +func TestANodeReadAuthenticatesAsItsClusterOnceItsSecretIsKnown(t *testing.T) { + api := aControlPlane().reporting(nodeStatusInCreation) + r, _ := aSteadyNode(t, api, aClusterSecret()) + // Forces the direct StorageNode() read rather than the stream's cache. + r.Nodes = &deliveredNodes{synced: false} + + settle(t, r) + + checkBearerTokens(t, api) +} + +// A node operation's calls -- the ones actions.go, classify.go, +// hostmaintenance.go, migrate.go, and remove.go make -- authenticate the same +// way, resolved from the operation's own target node's cluster. +func TestANodeOperationAuthenticatesAsItsClusterOnceItsSecretIsKnown(t *testing.T) { + api := aControlPlane().reporting(nodeStatusOnline) + ops := anAdvancingOperation( + "a-shutdown", simplyblockv1alpha2.StorageNodeOpsActionShutdown, stepRequesting) + r, _ := anOpsWorld(t, api, ops, aClusterSecret()) + + pass(t, r, "a-shutdown") + + checkBearerTokens(t, api) +} diff --git a/operator/internal/controllers/node/fixtures_test.go b/operator/internal/controllers/node/fixtures_test.go index d38a3ec7d..f266e9f73 100644 --- a/operator/internal/controllers/node/fixtures_test.go +++ b/operator/internal/controllers/node/fixtures_test.go @@ -88,11 +88,24 @@ type scriptedControlPlane struct { // once reads. calls []string + // bearerTokens parallels calls: bearerTokens[i] is what + // webapi.BearerTokenFromContext reported for calls[i]'s own context, so a + // test can check that a cluster-scoped call authenticated as the cluster + // it was actually asked about rather than as this operator's own + // Kubernetes identity. + bearerTokens []bearerObservation + // restarts carries the parameters of each restart, because three actions // issue one and they differ precisely in what they fill in. restarts []RestartParams } +// bearerObservation is one entry of scriptedControlPlane.bearerTokens. +type bearerObservation struct { + token string + ok bool +} + // aControlPlane reports one online node and nothing else. func aControlPlane() *scriptedControlPlane { return &scriptedControlPlane{ @@ -147,15 +160,17 @@ func (c *scriptedControlPlane) asked(method string) int { return count } -func (c *scriptedControlPlane) record(method, argument string) error { +func (c *scriptedControlPlane) record(ctx context.Context, method, argument string) error { + token, ok := webapi.BearerTokenFromContext(ctx) c.calls = append(c.calls, method+":"+argument) + c.bearerTokens = append(c.bearerTokens, bearerObservation{token, ok}) return c.refuse[method] } func (c *scriptedControlPlane) StorageNode( - _ context.Context, _, nodeID string, + ctx context.Context, _, nodeID string, ) (NodeReading, bool, error) { - if err := c.record("StorageNode", nodeID); err != nil { + if err := c.record(ctx, "StorageNode", nodeID); err != nil { return NodeReading{}, false, err } reading, found := c.nodes[nodeID] @@ -163,9 +178,9 @@ func (c *scriptedControlPlane) StorageNode( } func (c *scriptedControlPlane) StorageNodes( - _ context.Context, clusterID string, + ctx context.Context, clusterID string, ) ([]NodeReading, error) { - if err := c.record("StorageNodes", clusterID); err != nil { + if err := c.record(ctx, "StorageNodes", clusterID); err != nil { return nil, err } readings := make([]NodeReading, 0, len(c.nodes)) @@ -176,58 +191,58 @@ func (c *scriptedControlPlane) StorageNodes( } func (c *scriptedControlPlane) AddNode( - _ context.Context, clusterID string, _ utils.StorageNodeSetAddParams, + ctx context.Context, clusterID string, _ utils.StorageNodeSetAddParams, ) error { - return c.record("AddNode", clusterID) + return c.record(ctx, "AddNode", clusterID) } -func (c *scriptedControlPlane) Suspend(_ context.Context, _, nodeID string) error { - return c.record("Suspend", nodeID) +func (c *scriptedControlPlane) Suspend(ctx context.Context, _, nodeID string) error { + return c.record(ctx, "Suspend", nodeID) } -func (c *scriptedControlPlane) Resume(_ context.Context, _, nodeID string) error { - return c.record("Resume", nodeID) +func (c *scriptedControlPlane) Resume(ctx context.Context, _, nodeID string) error { + return c.record(ctx, "Resume", nodeID) } -func (c *scriptedControlPlane) ShutdownNode(_ context.Context, _, nodeID string) error { - return c.record("ShutdownNode", nodeID) +func (c *scriptedControlPlane) ShutdownNode(ctx context.Context, _, nodeID string) error { + return c.record(ctx, "ShutdownNode", nodeID) } func (c *scriptedControlPlane) RestartNode( - _ context.Context, _, nodeID string, params RestartParams, + ctx context.Context, _, nodeID string, params RestartParams, ) error { c.restarts = append(c.restarts, params) - return c.record("RestartNode", nodeID) + return c.record(ctx, "RestartNode", nodeID) } -func (c *scriptedControlPlane) Promote(_ context.Context, _, nodeID string) error { - return c.record("Promote", nodeID) +func (c *scriptedControlPlane) Promote(ctx context.Context, _, nodeID string) error { + return c.record(ctx, "Promote", nodeID) } -func (c *scriptedControlPlane) RemoveNode(_ context.Context, _, nodeID string) error { - return c.record("RemoveNode", nodeID) +func (c *scriptedControlPlane) RemoveNode(ctx context.Context, _, nodeID string) error { + return c.record(ctx, "RemoveNode", nodeID) } func (c *scriptedControlPlane) StoragePools( - _ context.Context, clusterID string, + ctx context.Context, clusterID string, ) ([]webapi.StoragePoolInfo, error) { - if err := c.record("StoragePools", clusterID); err != nil { + if err := c.record(ctx, "StoragePools", clusterID); err != nil { return nil, err } return c.pools, nil } func (c *scriptedControlPlane) PoolVolumes( - _ context.Context, _, poolID string, + ctx context.Context, _, poolID string, ) ([]webapi.VolumeInfo, error) { - if err := c.record("PoolVolumes", poolID); err != nil { + if err := c.record(ctx, "PoolVolumes", poolID); err != nil { return nil, err } return c.volumes[poolID], nil } -func (c *scriptedControlPlane) DeleteVolume(_ context.Context, _, _, volumeID string) error { - return c.record("DeleteVolume", volumeID) +func (c *scriptedControlPlane) DeleteVolume(ctx context.Context, _, _, volumeID string) error { + return c.record(ctx, "DeleteVolume", volumeID) } // anOpsNode is the node every operation in these suites targets: provisioned, diff --git a/operator/internal/controllers/node/storagenode_controller.go b/operator/internal/controllers/node/storagenode_controller.go index 176b83d75..b864087cc 100644 --- a/operator/internal/controllers/node/storagenode_controller.go +++ b/operator/internal/controllers/node/storagenode_controller.go @@ -239,6 +239,18 @@ func (r *StorageNodeReconciler) Reconcile( if err != nil { return ctrl.Result{}, err } + // Every control-plane call below authenticates as this node's cluster, + // using its own recorded secret, rather than as this operator's own + // Kubernetes identity -- the only way to reach a control plane a + // different Kubernetes cluster runs (a ControlPlane.spec.source.managed + // one), since a Kubernetes TokenReview can never cross a cluster + // boundary. cluster is nil for one the object outlived (§3.4), which + // leaves ctx unauthenticated the same as before this fix: nothing below + // reaches the control plane for a node whose cluster is gone. + if cluster != nil { + secret, err := clusterSecretByName(ctx, r.Client, cluster.Namespace, cluster.Name) + ctx = authenticatedContext(ctx, secret, err) + } if !node.DeletionTimestamp.IsZero() { return r.teardown(ctx, &node) diff --git a/operator/internal/controllers/node/storagenodeops_controller.go b/operator/internal/controllers/node/storagenodeops_controller.go index 2b9ab2b15..a0c1a7381 100644 --- a/operator/internal/controllers/node/storagenodeops_controller.go +++ b/operator/internal/controllers/node/storagenodeops_controller.go @@ -256,6 +256,15 @@ func (r *StorageNodeOpsReconciler) Reconcile( func (r *StorageNodeOpsReconciler) advance( ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, ) (ctrl.Result, error) { + // Every control-plane call this step and everything downstream of it + // makes authenticates as this operation's target node's cluster when its + // secret is known, rather than as this operator's own Kubernetes + // identity -- the only way to reach a control plane a different + // Kubernetes cluster runs (a ControlPlane.spec.source.managed one), + // since a Kubernetes TokenReview can never cross a cluster boundary. + secret, secretErr := clusterSecretForNode(ctx, r.Client, ops.Namespace, ops.Spec.NodeRef) + ctx = authenticatedContext(ctx, secret, secretErr) + machine, err := graphs().FromSnapshot(ctx, action(ops.Spec.Action), statemachine.FromKube[step](ops.Status.Step)) if err != nil { diff --git a/operator/internal/webapi/context.go b/operator/internal/webapi/context.go new file mode 100644 index 000000000..1c9f0082e --- /dev/null +++ b/operator/internal/webapi/context.go @@ -0,0 +1,22 @@ +package webapi + +import "context" + +type bearerTokenKey struct{} + +// WithBearerToken attaches a bearer credential to ctx that Do and +// DoWithHeaders send instead of the client's own service-account token. It is +// how a call scoped to one cluster authenticates as that cluster rather than +// as this process's own Kubernetes identity -- the only way to reach a +// control plane a different Kubernetes cluster runs (ControlPlane.spec.source.managed), +// since a Kubernetes TokenReview can never cross a cluster boundary. +func WithBearerToken(ctx context.Context, token string) context.Context { + return context.WithValue(ctx, bearerTokenKey{}, token) +} + +// BearerTokenFromContext returns the token WithBearerToken attached, and +// whether one was. +func BearerTokenFromContext(ctx context.Context) (string, bool) { + token, ok := ctx.Value(bearerTokenKey{}).(string) + return token, ok +} diff --git a/operator/internal/webapi/request.go b/operator/internal/webapi/request.go index 7734cebf2..59f2e3729 100644 --- a/operator/internal/webapi/request.go +++ b/operator/internal/webapi/request.go @@ -51,8 +51,14 @@ func (c *Client) DoWithHeaders( return nil, nil, 0, fmt.Errorf("create request: %w", err) } - // Attach auth header - req.Header.Set("Authorization", fmt.Sprintf("Bearer %s", c.saToken)) + // Attach auth header. A token WithBearerToken attached to ctx wins over + // this process's own service-account token, for a call scoped to a + // cluster whose control plane lives on a different Kubernetes cluster. + token := c.saToken + if override, ok := BearerTokenFromContext(ctx); ok { + token = override + } + req.Header.Set("Authorization", fmt.Sprintf("Bearer %s", token)) req.Header.Set("Content-Type", "application/json") // Execute the request diff --git a/operator/internal/webapi/request_test.go b/operator/internal/webapi/request_test.go index 01406ea25..5a3da827c 100644 --- a/operator/internal/webapi/request_test.go +++ b/operator/internal/webapi/request_test.go @@ -73,6 +73,37 @@ func TestDoAgainstSpecMockSendsHeadersBodyAndReturnsResponse(t *testing.T) { } } +// A call scoped to one cluster authenticates as that cluster when +// WithBearerToken names one, instead of as this process's own service +// account -- the only way to reach a control plane a different Kubernetes +// cluster runs, since a TokenReview can never cross a cluster boundary. +func TestDoSendsTheBearerTokenAttachedToContextInsteadOfTheServiceAccountToken(t *testing.T) { + mock := webapimock.NewSpecServerFromFile(t, "../../../shared/openapi.json", false) + defer mock.Close() + + mock.Register( + http.MethodGet, + "/api/v2/clusters/cluster-uuid/", + webapimock.RouteResponse{Status: http.StatusOK, Body: `{}`}, + ) + + c := NewClient(mock.URL()) + c.saToken = "operators-own-service-account-token" + + ctx := WithBearerToken(context.Background(), "cluster-uuids-own-secret") + if _, _, err := c.Do(ctx, http.MethodGet, "/api/v2/clusters/cluster-uuid/", nil); err != nil { + t.Fatalf("Do returned error: %v", err) + } + + reqs := mock.Requests() + if len(reqs) != 1 { + t.Fatalf("expected one request, got %d", len(reqs)) + } + if got := reqs[0].Headers["Authorization"]; got != "Bearer cluster-uuids-own-secret" { + t.Fatalf("authorization header = %q, want the context's bearer token", got) + } +} + func TestDoAgainstStrictSpecMockReturns400ForUnknownPath(t *testing.T) { mock := webapimock.NewSpecServerFromFile(t, "../../../shared/openapi.json", false) defer mock.Close() From 2aac5713f867df49457acd74b191d49a2e45c26a Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Tue, 22 Sep 2026 14:55:06 +0100 Subject: [PATCH 132/206] moved csiAddons into sidecarImages --- .../charts/simplyblock-operator/values.schema.json | 4 ++++ helm-charts/charts/simplyblock-operator/values.yaml | 8 ++++++-- .../controllers/driver/registration_test.go | 13 +++++++++++++ operator/internal/controllers/driver/sidecars.go | 13 +++++++------ 4 files changed, 30 insertions(+), 8 deletions(-) diff --git a/helm-charts/charts/simplyblock-operator/values.schema.json b/helm-charts/charts/simplyblock-operator/values.schema.json index 13e8d13c6..45cc781b5 100644 --- a/helm-charts/charts/simplyblock-operator/values.schema.json +++ b/helm-charts/charts/simplyblock-operator/values.schema.json @@ -433,6 +433,10 @@ "nodeDriverRegistrar": { "type": "string", "pattern": "^($|(quay\\.io/simplyblock-io|docker\\.io/simplyblock|public\\.ecr\\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$" + }, + "csiAddons": { + "type": "string", + "pattern": "^($|(quay\\.io/simplyblock-io|docker\\.io/simplyblock|public\\.ecr\\.aws/simply-block|quay\\.io/csiaddons)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$" } } } diff --git a/helm-charts/charts/simplyblock-operator/values.yaml b/helm-charts/charts/simplyblock-operator/values.yaml index 32f713ebd..8d3cf551c 100644 --- a/helm-charts/charts/simplyblock-operator/values.yaml +++ b/helm-charts/charts/simplyblock-operator/values.yaml @@ -232,9 +232,12 @@ driver: # accounts nothing presents. enableServiceAccountAuth: false - # Overrides for the six CSI sidecars, as full `repository:tag` references. + # Overrides for the seven CSI sidecars, as full `repository:tag` references. # Empty takes the version this operator release ships, which is the - # combination it was tested against. + # combination it was tested against. csiAddons's own default is already the + # real upstream image (quay.io/csiaddons/k8s-sidecar) rather than a + # quay.io/simplyblock-io mirror, which does not exist yet -- see + # SimplyblockDriverSpec.SidecarImages.CSIAddons. sidecarImages: provisioner: "" attacher: "" @@ -242,6 +245,7 @@ driver: snapshotter: "" healthMonitor: "" nodeDriverRegistrar: "" + csiAddons: "" controlplane: # Where the control plane is, for deployment.profile: managed. diff --git a/operator/internal/controllers/driver/registration_test.go b/operator/internal/controllers/driver/registration_test.go index 1cbb1a2bd..4456ec910 100644 --- a/operator/internal/controllers/driver/registration_test.go +++ b/operator/internal/controllers/driver/registration_test.go @@ -142,6 +142,19 @@ func TestSidecarsDefaultToTheOperatorsRelease(t *testing.T) { } } +// The csi-addons sidecar's default names the real upstream image +// (quay.io/csiaddons/k8s-sidecar) rather than a quay.io/simplyblock-io mirror +// that does not exist yet -- unlike the other six sidecars, which do have one. +// TestSidecarsDefaultToTheOperatorsRelease checks defaultCSIAddonsImage +// against itself, so a wrong constant would still pass it; this pins the +// literal a pod actually pulls. +func TestTheCSIAddonsSidecarDefaultsToTheRealUpstreamImage(t *testing.T) { + const wantImage = "quay.io/csiaddons/k8s-sidecar:v0.15.0" + if got := sidecars(testDriver("simplyblock")).csiAddons; got != wantImage { + t.Errorf("csiAddons default = %q, want the real upstream image %q", got, wantImage) + } +} + // U-86: one override reaches its own sidecar and no other, which is the property // that makes a pin survivable without freezing the rest of the deployment. func TestOneSidecarOverrideReachesOnlyItsOwn(t *testing.T) { diff --git a/operator/internal/controllers/driver/sidecars.go b/operator/internal/controllers/driver/sidecars.go index 3e4491abb..582bb50c2 100644 --- a/operator/internal/controllers/driver/sidecars.go +++ b/operator/internal/controllers/driver/sidecars.go @@ -25,12 +25,13 @@ const ( defaultSnapshotterImage = "quay.io/simplyblock-io/csi-snapshotter:v8.2.0" defaultHealthMonitorImage = "quay.io/simplyblock-io/csi-external-health-monitor-controller:v0.14.0" defaultNodeDriverRegistrarImage = "quay.io/simplyblock-io/csi-node-driver-registrar:v2.12.0" - // defaultCSIAddonsImage is the kubernetes-csi-addons sidecar (upstream - // quay.io/csiaddons/k8s-sidecar), pinned at the same v0.15.0 the chart's - // controller-manager runs (design P0-5). Named for the eventual - // quay.io/simplyblock-io mirror this field's validation pattern requires, - // which does not exist yet; mirroring it is a release task. - defaultCSIAddonsImage = "quay.io/simplyblock-io/csi-addons-sidecar:v0.15.0" + // defaultCSIAddonsImage is the kubernetes-csi-addons sidecar, pinned at the + // same v0.15.0 the chart's controller-manager runs (design P0-5). The real + // upstream image (quay.io/csiaddons/k8s-sidecar) rather than a + // quay.io/simplyblock-io mirror: that mirror does not exist yet, and + // mirroring it is a release task, not something to assume has already + // happened. + defaultCSIAddonsImage = "quay.io/csiaddons/k8s-sidecar:v0.15.0" ) // resolvedSidecars is the image each sidecar runs, after the overrides. From 742fbf25e8a5d471ec64a9a714fa027dab416313 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Tue, 22 Sep 2026 16:04:52 +0100 Subject: [PATCH 133/206] Thread a managed control plane's admin credential into CreateCluster/ClusterByName so a remote cluster can create its own backend identity --- .../storage.simplyblock.io_controlplanes.yaml | 26 ++++ .../templates/controlplane_cr.yaml | 4 + .../charts/simplyblock-operator/values.yaml | 11 ++ operator/api/v1alpha2/controlplane_types.go | 15 +++ .../api/v1alpha2/zz_generated.deepcopy.go | 5 + operator/cmd/main.go | 12 +- .../storage.simplyblock.io_controlplanes.yaml | 26 ++++ .../controllers/cluster/controlplane.go | 45 ++++++- .../cluster/controlplane_admin_test.go | 104 +++++++++++++++ .../controllers/controlplane/credential.go | 80 ++++++++++++ .../controlplane/credential_test.go | 121 ++++++++++++++++++ .../controllers/controlplane/helpers_test.go | 14 ++ .../controllers/controlplane/managementapi.go | 15 +++ .../controlplane/workloads_test.go | 36 ++++++ .../storage.simplyblock.io_controlplanes.yaml | 26 ++++ 15 files changed, 532 insertions(+), 8 deletions(-) create mode 100644 operator/internal/controllers/cluster/controlplane_admin_test.go create mode 100644 operator/internal/controllers/controlplane/credential.go create mode 100644 operator/internal/controllers/controlplane/credential_test.go diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_controlplanes.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_controlplanes.yaml index 95afa9396..077d5ea80 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_controlplanes.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_controlplanes.yaml @@ -163,6 +163,32 @@ spec: local: description: Local is a control plane the operator installs. properties: + adminTokenSecretRef: + description: |- + AdminTokenSecretRef names a Secret in this namespace holding a static + admin bearer token this control plane accepts, under the `token` key, in + addition to this deployment's own Kubernetes identity + (SB_K8S_ADMIN_SERVICE_ACCOUNTS). It is what lets a cluster this control + plane manages remotely (spec.source.managed there, + ManagedControlPlane.CredentialsSecretRef naming the same value) + authenticate a CreateCluster call, since a Kubernetes TokenReview can + never cross a cluster boundary. + + The Secret is projected into the management API container's environment + with secretKeyRef, so this operator never itself reads the plaintext. + Absent grants no credential beyond the operator's own service account. + properties: + name: + default: "" + description: |- + Name of the referent. + This field is effectively required, but due to backwards compatibility is + allowed to be empty. Instances of this type with an empty value here are + almost certainly wrong. + More info: https://kubernetes.io/docs/concepts/overview/working-with-objects/names/#names + type: string + type: object + x-kubernetes-map-type: atomic foundationDB: description: FoundationDB sizes the FoundationDB the management API stores its state in. properties: diff --git a/helm-charts/charts/simplyblock-operator/templates/controlplane_cr.yaml b/helm-charts/charts/simplyblock-operator/templates/controlplane_cr.yaml index ea8aa8c6a..78ad25aa1 100644 --- a/helm-charts/charts/simplyblock-operator/templates/controlplane_cr.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/controlplane_cr.yaml @@ -75,6 +75,10 @@ spec: {{- end }} {{- end }} {{- end }} + {{- if .Values.controlplane.local.adminTokenSecretRef }} + adminTokenSecretRef: + name: {{ .Values.controlplane.local.adminTokenSecretRef | quote }} + {{- end }} {{- else }} # A control plane elsewhere manages this cluster's storage. The operator # installs nothing here: it resolves the endpoint, probes it, and reports. diff --git a/helm-charts/charts/simplyblock-operator/values.yaml b/helm-charts/charts/simplyblock-operator/values.yaml index 8d3cf551c..b8325f055 100644 --- a/helm-charts/charts/simplyblock-operator/values.yaml +++ b/helm-charts/charts/simplyblock-operator/values.yaml @@ -248,6 +248,17 @@ driver: csiAddons: "" controlplane: + # deployment.profile: standalone options -- this cluster's own control plane. + local: + # Secret in this namespace holding a static admin bearer token this + # control plane accepts, under `token`, in addition to this operator's + # own Kubernetes identity. It is what lets a cluster this control plane + # manages remotely (deployment.profile: managed there, + # controlplane.managed.credentialsSecretRef naming the same value) create + # its backend cluster identity, since a Kubernetes TokenReview can never + # cross a cluster boundary. Empty grants no additional admin credential. + adminTokenSecretRef: "" + # Where the control plane is, for deployment.profile: managed. managed: # Base URL of the management API. Required for that profile. A loopback or diff --git a/operator/api/v1alpha2/controlplane_types.go b/operator/api/v1alpha2/controlplane_types.go index 5818dd96d..10b068de0 100644 --- a/operator/api/v1alpha2/controlplane_types.go +++ b/operator/api/v1alpha2/controlplane_types.go @@ -222,6 +222,21 @@ type LocalControlPlane struct { // +kubebuilder:default={} // +optional TLS ControlPlaneTLS `json:"tls,omitempty"` + + // AdminTokenSecretRef names a Secret in this namespace holding a static + // admin bearer token this control plane accepts, under the `token` key, in + // addition to this deployment's own Kubernetes identity + // (SB_K8S_ADMIN_SERVICE_ACCOUNTS). It is what lets a cluster this control + // plane manages remotely (spec.source.managed there, + // ManagedControlPlane.CredentialsSecretRef naming the same value) + // authenticate a CreateCluster call, since a Kubernetes TokenReview can + // never cross a cluster boundary. + // + // The Secret is projected into the management API container's environment + // with secretKeyRef, so this operator never itself reads the plaintext. + // Absent grants no credential beyond the operator's own service account. + // +optional + AdminTokenSecretRef *corev1.LocalObjectReference `json:"adminTokenSecretRef,omitempty"` } // ManagedControlPlane is a control plane somewhere else, which this cluster's diff --git a/operator/api/v1alpha2/zz_generated.deepcopy.go b/operator/api/v1alpha2/zz_generated.deepcopy.go index ce1ff202b..2075f626b 100644 --- a/operator/api/v1alpha2/zz_generated.deepcopy.go +++ b/operator/api/v1alpha2/zz_generated.deepcopy.go @@ -935,6 +935,11 @@ func (in *LocalControlPlane) DeepCopyInto(out *LocalControlPlane) { } } in.TLS.DeepCopyInto(&out.TLS) + if in.AdminTokenSecretRef != nil { + in, out := &in.AdminTokenSecretRef, &out.AdminTokenSecretRef + *out = new(v1.LocalObjectReference) + **out = **in + } } // DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new LocalControlPlane. diff --git a/operator/cmd/main.go b/operator/cmd/main.go index 65d050f1c..6ec496b5a 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -522,6 +522,14 @@ func main() { controlPlaneEndpoint := controlplanecontroller.NewEndpointResolver( mgr.GetClient(), operatorNamespace) + // What a managed control plane's CreateCluster and ClusterByName calls + // authenticate with, read from the same ControlPlane object + // (spec.source.managed.credentialsSecretRef) for the same reason: a + // StorageCluster CR applied against a remote control plane has no other + // credential to create its backend identity with. + controlPlaneCredential := controlplanecontroller.NewCredentialResolver( + mgr.GetClient(), operatorNamespace) + if err := (&controlplanecontroller.ControlPlaneReconciler{ Client: mgr.GetClient(), Scheme: mgr.GetScheme(), @@ -542,7 +550,7 @@ func main() { Client: mgr.GetClient(), Scheme: mgr.GetScheme(), Recorder: mgr.GetEventRecorder("storagecluster-controller"), - API: clustercontroller.NewControlPlane(controlPlaneEndpoint), + API: clustercontroller.NewControlPlane(controlPlaneEndpoint, controlPlaneCredential), Namespace: operatorNamespace, Clusters: clusterSubscription, Tasks: taskSubscription, @@ -808,7 +816,7 @@ func main() { Client: mgr.GetClient(), Scheme: mgr.GetScheme(), Recorder: mgr.GetEventRecorder("storageclusterops-controller"), - API: clustercontroller.NewControlPlane(controlPlaneEndpoint), + API: clustercontroller.NewControlPlane(controlPlaneEndpoint, controlPlaneCredential), Clusters: clusterSubscription, Nodes: nodeSubscription, Tasks: taskSubscription, diff --git a/operator/config/crd/bases/storage.simplyblock.io_controlplanes.yaml b/operator/config/crd/bases/storage.simplyblock.io_controlplanes.yaml index 95afa9396..077d5ea80 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_controlplanes.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_controlplanes.yaml @@ -163,6 +163,32 @@ spec: local: description: Local is a control plane the operator installs. properties: + adminTokenSecretRef: + description: |- + AdminTokenSecretRef names a Secret in this namespace holding a static + admin bearer token this control plane accepts, under the `token` key, in + addition to this deployment's own Kubernetes identity + (SB_K8S_ADMIN_SERVICE_ACCOUNTS). It is what lets a cluster this control + plane manages remotely (spec.source.managed there, + ManagedControlPlane.CredentialsSecretRef naming the same value) + authenticate a CreateCluster call, since a Kubernetes TokenReview can + never cross a cluster boundary. + + The Secret is projected into the management API container's environment + with secretKeyRef, so this operator never itself reads the plaintext. + Absent grants no credential beyond the operator's own service account. + properties: + name: + default: "" + description: |- + Name of the referent. + This field is effectively required, but due to backwards compatibility is + allowed to be empty. Instances of this type with an empty value here are + almost certainly wrong. + More info: https://kubernetes.io/docs/concepts/overview/working-with-objects/names/#names + type: string + type: object + x-kubernetes-map-type: atomic foundationDB: description: FoundationDB sizes the FoundationDB the management API stores its state in. properties: diff --git a/operator/internal/controllers/cluster/controlplane.go b/operator/internal/controllers/cluster/controlplane.go index e84d07373..1b1ab5df9 100644 --- a/operator/internal/controllers/cluster/controlplane.go +++ b/operator/internal/controllers/cluster/controlplane.go @@ -102,6 +102,12 @@ type httpControlPlane struct { // startup client is the only one. resolve controlplane.EndpointResolver + // credential answers what a managed control plane's CreateCluster and + // ClusterByName calls authenticate with (see adminContext). Nil, or a + // resolver that answers false, means neither call adds a bearer token + // beyond whatever webapi.Client already carries. + credential controlplane.CredentialResolver + // mu guards resolved, which is rebuilt when the published endpoint changes. mu sync.Mutex resolved *webapi.Client @@ -109,10 +115,13 @@ type httpControlPlane struct { // NewControlPlane returns the HTTP-backed control-plane surface. // -// The resolver may be nil, which is what a test passes: calls then go to the -// startup client and nothing reads a ControlPlane object. -func NewControlPlane(resolve controlplane.EndpointResolver) ControlPlane { - return &httpControlPlane{client: webapi.NewClient(), resolve: resolve} +// Either resolver may be nil, which is what a test passes: calls then go to +// the startup client, unauthenticated beyond its own saToken, and nothing +// reads a ControlPlane object. +func NewControlPlane( + resolve controlplane.EndpointResolver, credential controlplane.CredentialResolver, +) ControlPlane { + return &httpControlPlane{client: webapi.NewClient(), resolve: resolve, credential: credential} } func (c *httpControlPlane) Endpoint(ctx context.Context) string { return c.clientFor(ctx).BaseURL } @@ -125,7 +134,7 @@ func (c *httpControlPlane) Ready(ctx context.Context) error { func (c *httpControlPlane) CreateCluster( ctx context.Context, params utils.ClusterAddParams, ) (webapi.ClusterResponse, error) { - body, err := c.call(ctx, http.MethodPost, "/api/v2/clusters/", params) + body, err := c.call(c.adminContext(ctx), http.MethodPost, "/api/v2/clusters/", params) if err != nil { return webapi.ClusterResponse{}, err } @@ -145,7 +154,7 @@ func (c *httpControlPlane) Cluster( func (c *httpControlPlane) ClusterByName( ctx context.Context, name string, ) (utils.ClusterListEntry, bool, error) { - body, err := c.call(ctx, http.MethodGet, "/api/v2/clusters/", nil) + body, err := c.call(c.adminContext(ctx), http.MethodGet, "/api/v2/clusters/", nil) if err != nil { return utils.ClusterListEntry{}, false, err } @@ -306,3 +315,27 @@ func (c *httpControlPlane) clientFor(ctx context.Context) *webapi.Client { } return c.resolved } + +// adminContext is what CreateCluster and ClusterByName call with, instead of +// ctx directly. +// +// Both run before any cluster secret exists -- there is no cluster yet to +// have one, and no adoption has happened either -- so a managed control +// plane's admin credential is the only thing that can authenticate them, and +// nothing else in this package has a stronger claim to the context's bearer +// token slot. A ctx that already carries one (a caller more specific than +// this method wins, though none exists on this interface today) is left +// alone. +func (c *httpControlPlane) adminContext(ctx context.Context) context.Context { + if c.credential == nil { + return ctx + } + if _, ok := webapi.BearerTokenFromContext(ctx); ok { + return ctx + } + token, ok := c.credential(ctx) + if !ok { + return ctx + } + return webapi.WithBearerToken(ctx, token) +} diff --git a/operator/internal/controllers/cluster/controlplane_admin_test.go b/operator/internal/controllers/cluster/controlplane_admin_test.go new file mode 100644 index 000000000..ea668241f --- /dev/null +++ b/operator/internal/controllers/cluster/controlplane_admin_test.go @@ -0,0 +1,104 @@ +// Whether CreateCluster and ClusterByName carry a managed control plane's +// admin credential. +// +// These two are the only calls made before any cluster secret exists to +// authenticate with -- there is no cluster yet to have one, and no adoption +// has happened yet either. Without the credential reaching them, a +// StorageCluster CR applied against a managed control plane could never +// create its backend identity there at all, which is the gap this closes. +// Every other call on this interface keeps authenticating however it already +// did: as the cluster it was adopted or created as, via the context wrap the +// reconciler applies at the call site (storagecluster_controller.go). + +package cluster + +import ( + "context" + "net/http" + "testing" + + "github.com/simplyblock/simplyblock-operator/internal/utils" + webapimock "github.com/simplyblock/simplyblock-operator/internal/webapi/mock" +) + +const clusterAdminSpecPath = "../../../../shared/openapi.json" + +func TestCreateClusterAuthenticatesWithTheManagedAdminCredentialWhenOneResolves(t *testing.T) { + mock := webapimock.NewSpecServerFromFile(t, clusterAdminSpecPath, false) + defer mock.Close() + + mock.Register(http.MethodPost, "/api/v2/clusters/", webapimock.RouteResponse{ + Status: http.StatusCreated, + Body: `{"id":"cluster-uuid","secret":"cluster-secret"}`, + }) + + resolveEndpoint := func(context.Context) string { return mock.URL() } + resolveCredential := func(context.Context) (string, bool) { return "admin-token", true } + + api := NewControlPlane(resolveEndpoint, resolveCredential) + if _, err := api.CreateCluster(context.Background(), utils.ClusterAddParams{Name: "b"}); err != nil { + t.Fatalf("CreateCluster: %v", err) + } + + reqs := mock.Requests() + if len(reqs) != 1 { + t.Fatalf("expected one request, got %d", len(reqs)) + } + if got := reqs[0].Headers["Authorization"]; got != "Bearer admin-token" { + t.Errorf("authorization header = %q, want the managed admin credential", got) + } +} + +func TestClusterByNameAuthenticatesWithTheManagedAdminCredentialWhenOneResolves(t *testing.T) { + mock := webapimock.NewSpecServerFromFile(t, clusterAdminSpecPath, false) + defer mock.Close() + + mock.Register(http.MethodGet, "/api/v2/clusters/", webapimock.RouteResponse{ + Status: http.StatusOK, Body: `[]`, + }) + + resolveEndpoint := func(context.Context) string { return mock.URL() } + resolveCredential := func(context.Context) (string, bool) { return "admin-token", true } + + api := NewControlPlane(resolveEndpoint, resolveCredential) + if _, _, err := api.ClusterByName(context.Background(), "b"); err != nil { + t.Fatalf("ClusterByName: %v", err) + } + + reqs := mock.Requests() + if len(reqs) != 1 { + t.Fatalf("expected one request, got %d", len(reqs)) + } + if got := reqs[0].Headers["Authorization"]; got != "Bearer admin-token" { + t.Errorf("authorization header = %q, want the managed admin credential", got) + } +} + +// A local control plane's resolver names no credential, so CreateCluster +// authenticates exactly as it always has: the client's own (empty) saToken, +// not the admin credential. +func TestCreateClusterCarriesNoAdminCredentialWhenNoneResolves(t *testing.T) { + mock := webapimock.NewSpecServerFromFile(t, clusterAdminSpecPath, false) + defer mock.Close() + + mock.Register(http.MethodPost, "/api/v2/clusters/", webapimock.RouteResponse{ + Status: http.StatusCreated, + Body: `{"id":"cluster-uuid","secret":"cluster-secret"}`, + }) + + resolveEndpoint := func(context.Context) string { return mock.URL() } + resolveCredential := func(context.Context) (string, bool) { return "", false } + + api := NewControlPlane(resolveEndpoint, resolveCredential) + if _, err := api.CreateCluster(context.Background(), utils.ClusterAddParams{Name: "b"}); err != nil { + t.Fatalf("CreateCluster: %v", err) + } + + reqs := mock.Requests() + if len(reqs) != 1 { + t.Fatalf("expected one request, got %d", len(reqs)) + } + if got := reqs[0].Headers["Authorization"]; got == "Bearer admin-token" { + t.Errorf("authorization header = %q, want the client's own (empty) token, not the admin credential", got) + } +} diff --git a/operator/internal/controllers/controlplane/credential.go b/operator/internal/controllers/controlplane/credential.go new file mode 100644 index 000000000..3afe53bf3 --- /dev/null +++ b/operator/internal/controllers/controlplane/credential.go @@ -0,0 +1,80 @@ +// What the rest of the operator authenticates a call to a managed control +// plane with. +// +// This mirrors NewEndpointResolver (endpoint.go, resolver.go): the same +// singleton, the same cache interval, and the same "unreadable is a transient +// miss, not an error" answer, because a caller with no credential falls back +// to whatever it already carries -- its own cluster-secret authentication, or +// none at all for a local control plane -- and that fallback is exactly what an +// admitting deployment looked like before this existed. + +package controlplane + +import ( + "context" + "strings" + "sync" + "time" + + corev1 "k8s.io/api/core/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// CredentialResolver answers what a call to the control plane authenticates +// with, per call. The second return is false when the ControlPlane names no +// managed credential -- the caller then keeps whatever it already had. +type CredentialResolver func(ctx context.Context) (token string, ok bool) + +// NewCredentialResolver returns a resolver reading the ControlPlane +// singleton's spec.source.managed.credentialsSecretRef, in the operator's +// namespace. +func NewCredentialResolver(reader client.Reader, namespace string) CredentialResolver { + var ( + mu sync.Mutex + cached string + cachedOK bool + cachedAt time.Time + ) + + return func(ctx context.Context) (string, bool) { + mu.Lock() + defer mu.Unlock() + + if !cachedAt.IsZero() && time.Since(cachedAt) < endpointCacheTTL { + return cached, cachedOK + } + + var cp simplyblockv1alpha2.ControlPlane + key := client.ObjectKey{Namespace: namespace, Name: SingletonName} + if err := reader.Get(ctx, key, &cp); err != nil { + // Not cached, for the same reason NewEndpointResolver does not cache + // this case: an unreadable singleton is transient, and caching the + // negative answer would hold every caller on no credential for the + // rest of the interval. + return "", false + } + + managed := cp.Spec.Source.Managed + if managed == nil || managed.CredentialsSecretRef == nil || managed.CredentialsSecretRef.Name == "" { + cached, cachedOK, cachedAt = "", false, time.Now() + return cached, cachedOK + } + + var secret corev1.Secret + secretKey := client.ObjectKey{Namespace: namespace, Name: managed.CredentialsSecretRef.Name} + if err := reader.Get(ctx, secretKey, &secret); err != nil { + return "", false + } + + for _, k := range credentialKeys { + if value := strings.TrimSpace(string(secret.Data[k])); value != "" { + cached, cachedOK, cachedAt = value, true, time.Now() + return cached, cachedOK + } + } + cached, cachedOK, cachedAt = "", false, time.Now() + return cached, cachedOK + } +} diff --git a/operator/internal/controllers/controlplane/credential_test.go b/operator/internal/controllers/controlplane/credential_test.go new file mode 100644 index 000000000..8aa0f06e9 --- /dev/null +++ b/operator/internal/controllers/controlplane/credential_test.go @@ -0,0 +1,121 @@ +// The credential resolver: what the rest of the operator authenticates a call +// to a managed control plane with. +// +// It follows NewEndpointResolver's shape deliberately: a ControlPlane or +// Secret this reads cannot answer right now is the same, to every caller, as +// one that names no credential, and both mean "fall back to whatever this +// caller already had." Nothing here is required to reach a control plane the +// deployment itself installed, which is what makes it additive. + +package controlplane + +import ( + "context" + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "sigs.k8s.io/controller-runtime/pkg/client" +) + +func credentialSecret(token string) *corev1.Secret { + return &corev1.Secret{ + ObjectMeta: metav1.ObjectMeta{Name: "cp-token", Namespace: testNamespace}, + Data: map[string][]byte{"token": []byte(token)}, + } +} + +// A managed control plane naming a credential resolves to it, which is what +// lets a call carry the token the remote control plane's admin auth expects. +func TestTheCredentialResolverAnswersWithTheManagedToken(t *testing.T) { + cp := managedControlPlane("https://sb-control.example.com:5000") + secret := credentialSecret("a-bearer-token") + + resolve := NewCredentialResolver(newClient(t, cp, secret), testNamespace) + + token, ok := resolve(context.Background()) + if !ok || token != "a-bearer-token" { + t.Errorf("resolved (%q, %v), want (%q, true)", token, ok, "a-bearer-token") + } +} + +// A local control plane names no managed credential at all, so the resolver +// says so rather than inventing one. +func TestALocalControlPlaneResolvesToNoCredential(t *testing.T) { + cp := localControlPlane() + + resolve := NewCredentialResolver(newClient(t, cp), testNamespace) + + if token, ok := resolve(context.Background()); ok { + t.Errorf("resolved (%q, true) from a local control plane, want no credential", token) + } +} + +// A managed control plane that names no credentialsSecretRef is the in-cluster +// case the chart writes when it installs the control plane itself -- no token +// to carry. +func TestAManagedControlPlaneWithNoCredentialsRefResolvesToNoCredential(t *testing.T) { + cp := managedControlPlane("https://sb-control.example.com:5000") + cp.Spec.Source.Managed.CredentialsSecretRef = nil + + resolve := NewCredentialResolver(newClient(t, cp), testNamespace) + + if token, ok := resolve(context.Background()); ok { + t.Errorf("resolved (%q, true) with no ref set, want no credential", token) + } +} + +// A singleton that cannot be read resolves to no credential rather than to an +// error, for the same reason the endpoint resolver does: an unreadable object +// is a transient miss, not a failed call. +func TestAnAbsentControlPlaneResolvesToNoCredential(t *testing.T) { + resolve := NewCredentialResolver(newClient(t), testNamespace) + + if token, ok := resolve(context.Background()); ok { + t.Errorf("resolved (%q, true) against a cluster with no ControlPlane", token) + } +} + +// A Secret the ref names but that does not exist is the same as no credential +// to the caller -- there is nothing to authenticate with either way. +func TestAMissingCredentialsSecretResolvesToNoCredential(t *testing.T) { + cp := managedControlPlane("https://sb-control.example.com:5000") + + resolve := NewCredentialResolver(newClient(t, cp), testNamespace) + + if token, ok := resolve(context.Background()); ok { + t.Errorf("resolved (%q, true) with the named Secret absent", token) + } +} + +// A change to the credential reaches the next caller, which is the property +// that makes this a resolver rather than a value captured once at startup. +func TestAChangedCredentialReachesTheNextCaller(t *testing.T) { + ctx := context.Background() + + cp := managedControlPlane("https://sb-control.example.com:5000") + secret := credentialSecret("first-token") + c := newClient(t, cp, secret) + + resolve := NewCredentialResolver(c, testNamespace) + if token, ok := resolve(ctx); !ok || token != "first-token" { + t.Fatalf("resolved (%q, %v) before the change", token, ok) + } + + var current corev1.Secret + key := client.ObjectKey{Name: "cp-token", Namespace: testNamespace} + if err := c.Get(ctx, key, ¤t); err != nil { + t.Fatalf("read the secret: %v", err) + } + current.Data["token"] = []byte("second-token") + if err := c.Update(ctx, ¤t); err != nil { + t.Fatalf("publish the new token: %v", err) + } + + // The resolver caches for a short interval, so a fresh one stands in for the + // interval passing, exactly as TestAChangedEndpointReachesTheNextCaller does. + resolve = NewCredentialResolver(c, testNamespace) + if token, ok := resolve(ctx); !ok || token != "second-token" { + t.Errorf("resolved (%q, %v), want the token the Secret now carries", token, ok) + } +} diff --git a/operator/internal/controllers/controlplane/helpers_test.go b/operator/internal/controllers/controlplane/helpers_test.go index 08e6f2441..dad4510d8 100644 --- a/operator/internal/controllers/controlplane/helpers_test.go +++ b/operator/internal/controllers/controlplane/helpers_test.go @@ -172,6 +172,20 @@ func findDeployment(t *testing.T, objects []client.Object, name string) *appsv1. return d } +// findEnvVar locates a container env entry by name in a Deployment's first +// container, failing rather than returning a zero value so a missing entry +// reads as the assertion it is instead of a nil-field panic later. +func findEnvVar(t *testing.T, d *appsv1.Deployment, name string) corev1.EnvVar { + t.Helper() + for _, e := range d.Spec.Template.Spec.Containers[0].Env { + if e.Name == name { + return e + } + } + t.Fatalf("no %q env var on %s", name, d.Name) + return corev1.EnvVar{} +} + func findClusterRole(t *testing.T, objects []client.Object, name string) *rbacv1.ClusterRole { t.Helper() obj := findObject(objects, name) diff --git a/operator/internal/controllers/controlplane/managementapi.go b/operator/internal/controllers/controlplane/managementapi.go index 09c4c9212..6a5f3c565 100644 --- a/operator/internal/controllers/controlplane/managementapi.go +++ b/operator/internal/controllers/controlplane/managementapi.go @@ -232,6 +232,21 @@ func webAPIDeployment(cp *simplyblockv1alpha2.ControlPlane) *appsv1.Deployment { {Name: "SB_K8S_METRICS_SERVICE_ACCOUNTS", Value: "system:serviceaccount:" + cp.Namespace + ":simplyblock-prometheus"}, } + if ref := managed.AdminTokenSecretRef; ref != nil && ref.Name != "" { + // Sourced from the Secret directly rather than read and copied in here, + // so this operator never itself holds the plaintext -- the same reason + // resolveManaged reads ManagedControlPlane.CredentialsSecretRef on the + // other side only to attach it to a request, never to log or store it. + env = append(env, corev1.EnvVar{ + Name: "SB_ADMIN_TOKENS", + ValueFrom: &corev1.EnvVarSource{ + SecretKeyRef: &corev1.SecretKeySelector{ + LocalObjectReference: *ref, + Key: "token", + }, + }, + }) + } env = append(env, prometheusEnv()...) env = append(env, tlsEnv(managed)...) diff --git a/operator/internal/controllers/controlplane/workloads_test.go b/operator/internal/controllers/controlplane/workloads_test.go index 340b5dd36..3851987ee 100644 --- a/operator/internal/controllers/controlplane/workloads_test.go +++ b/operator/internal/controllers/controlplane/workloads_test.go @@ -87,6 +87,42 @@ func TestASingleManagementAPIInstanceStaysExpressible(t *testing.T) { } } +// AdminTokenSecretRef reaches the management API as SB_ADMIN_TOKENS, sourced +// via secretKeyRef rather than a literal value, so this operator never itself +// reads the plaintext. It is what lets a cluster this control plane manages +// remotely authenticate a CreateCluster call. +func TestAnAdminTokenSecretRefReachesTheManagementAPIsEnvironment(t *testing.T) { + cp := localControlPlane() + cp.Spec.Source.Local.AdminTokenSecretRef = &corev1.LocalObjectReference{Name: "hub-admin-token"} + + api := findDeployment(t, managementAPIObjects(cp), ComponentWebAPI) + env := findEnvVar(t, api, "SB_ADMIN_TOKENS") + + if env.ValueFrom == nil || env.ValueFrom.SecretKeyRef == nil { + t.Fatalf("SB_ADMIN_TOKENS is not sourced from a Secret: %#v", env) + } + if env.ValueFrom.SecretKeyRef.Name != "hub-admin-token" { + t.Errorf("secretKeyRef.name = %q, want %q", env.ValueFrom.SecretKeyRef.Name, "hub-admin-token") + } + if env.ValueFrom.SecretKeyRef.Key != "token" { + t.Errorf("secretKeyRef.key = %q, want %q", env.ValueFrom.SecretKeyRef.Key, "token") + } +} + +// Absent names no additional credential: the management API runs exactly as +// it always has, authenticating only this operator's own service account. +func TestNoAdminTokenSecretRefMeansNoExtraEnvVar(t *testing.T) { + cp := localControlPlane() + + api := findDeployment(t, managementAPIObjects(cp), ComponentWebAPI) + + for _, e := range api.Spec.Template.Spec.Containers[0].Env { + if e.Name == "SB_ADMIN_TOKENS" { + t.Fatalf("SB_ADMIN_TOKENS set with no adminTokenSecretRef: %#v", e) + } + } +} + // Every workload built from the control plane's own image runs it, so an upgrade // that writes one image onto the entity moves all of them. func TestEveryWorkloadOfTheControlPlaneRunsTheSpecsImage(t *testing.T) { diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_controlplanes.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_controlplanes.yaml index 95afa9396..077d5ea80 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_controlplanes.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_controlplanes.yaml @@ -163,6 +163,32 @@ spec: local: description: Local is a control plane the operator installs. properties: + adminTokenSecretRef: + description: |- + AdminTokenSecretRef names a Secret in this namespace holding a static + admin bearer token this control plane accepts, under the `token` key, in + addition to this deployment's own Kubernetes identity + (SB_K8S_ADMIN_SERVICE_ACCOUNTS). It is what lets a cluster this control + plane manages remotely (spec.source.managed there, + ManagedControlPlane.CredentialsSecretRef naming the same value) + authenticate a CreateCluster call, since a Kubernetes TokenReview can + never cross a cluster boundary. + + The Secret is projected into the management API container's environment + with secretKeyRef, so this operator never itself reads the plaintext. + Absent grants no credential beyond the operator's own service account. + properties: + name: + default: "" + description: |- + Name of the referent. + This field is effectively required, but due to backwards compatibility is + allowed to be empty. Instances of this type with an empty value here are + almost certainly wrong. + More info: https://kubernetes.io/docs/concepts/overview/working-with-objects/names/#names + type: string + type: object + x-kubernetes-map-type: atomic foundationDB: description: FoundationDB sizes the FoundationDB the management API stores its state in. properties: From 9dd1143123d0617bc326f716ce81a235044ecedc Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Tue, 22 Sep 2026 16:09:54 +0100 Subject: [PATCH 134/206] ran make operator-build-installer --- operator/dist/install.yaml | 26 ++++++++++++++++++++++++++ 1 file changed, 26 insertions(+) diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index 3e7f345e7..181bca865 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -1323,6 +1323,32 @@ spec: local: description: Local is a control plane the operator installs. properties: + adminTokenSecretRef: + description: |- + AdminTokenSecretRef names a Secret in this namespace holding a static + admin bearer token this control plane accepts, under the `token` key, in + addition to this deployment's own Kubernetes identity + (SB_K8S_ADMIN_SERVICE_ACCOUNTS). It is what lets a cluster this control + plane manages remotely (spec.source.managed there, + ManagedControlPlane.CredentialsSecretRef naming the same value) + authenticate a CreateCluster call, since a Kubernetes TokenReview can + never cross a cluster boundary. + + The Secret is projected into the management API container's environment + with secretKeyRef, so this operator never itself reads the plaintext. + Absent grants no credential beyond the operator's own service account. + properties: + name: + default: "" + description: |- + Name of the referent. + This field is effectively required, but due to backwards compatibility is + allowed to be empty. Instances of this type with an empty value here are + almost certainly wrong. + More info: https://kubernetes.io/docs/concepts/overview/working-with-objects/names/#names + type: string + type: object + x-kubernetes-map-type: atomic foundationDB: description: FoundationDB sizes the FoundationDB the management API stores its state in. From 8a7c58540958bb53e23b0936280a8a694edffaf5 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Tue, 22 Sep 2026 17:34:36 +0100 Subject: [PATCH 135/206] pool reach hub cluster endpoint --- operator/cmd/main.go | 9 +- .../pool/controlplane_endpoint_test.go | 92 +++++++++++++++++++ .../internal/controllers/pool/helpers_test.go | 5 + .../pool/storagepool_controller.go | 75 ++++++++++++++- 4 files changed, 173 insertions(+), 8 deletions(-) create mode 100644 operator/internal/controllers/pool/controlplane_endpoint_test.go diff --git a/operator/cmd/main.go b/operator/cmd/main.go index 6ec496b5a..67efb922a 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -584,10 +584,11 @@ func main() { os.Exit(1) } if err := (&pool.StoragePoolReconciler{ - Client: mgr.GetClient(), - Scheme: mgr.GetScheme(), - Recorder: mgr.GetEventRecorder("storagepool-controller"), - VolumeScopes: volumeScopes, + Client: mgr.GetClient(), + Scheme: mgr.GetScheme(), + Recorder: mgr.GetEventRecorder("storagepool-controller"), + VolumeScopes: volumeScopes, + EndpointResolver: controlPlaneEndpoint, }).SetupWithManager(mgr); err != nil { setupLog.Error(err, "unable to create controller", "controller", "StoragePool") os.Exit(1) diff --git a/operator/internal/controllers/pool/controlplane_endpoint_test.go b/operator/internal/controllers/pool/controlplane_endpoint_test.go new file mode 100644 index 000000000..99a877817 --- /dev/null +++ b/operator/internal/controllers/pool/controlplane_endpoint_test.go @@ -0,0 +1,92 @@ +// Whether StoragePoolReconciler reaches the control plane the ControlPlane +// object resolves to, and authenticates as its cluster once that cluster's +// secret is known -- rather than always the hardcoded in-cluster Service +// webapi.NewClient() defaults to. Before EndpointResolver existed on this +// reconciler, a pool on a ControlPlane.spec.source.managed deployment could +// never reach its control plane at all. See internal/controllers/cluster's +// identically-motivated fix (storagecluster_controller.go's clusterSecret and +// controlplane.go's clientFor). + +package pool + +import ( + "context" + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" +) + +// A pool's create reaches whatever endpoint the resolver publishes, not the +// package's built-in default. Without this, the request would try to dial +// the hardcoded in-cluster Service and never reach the stub at all -- the +// stub's own posts count is the proof, since the reconciler has no other way +// to answer testPoolUUID. +func TestAPoolReachesTheEndpointTheControlPlaneResolverPublishes(t *testing.T) { + cp, rec := newControlPlane(t), &recorder{} + c := newClient(t, newCluster(testClusterUUID), newPool("tenant-a")) + r := &StoragePoolReconciler{ + Client: c, + Scheme: testScheme(t), + Recorder: rec, + EndpointResolver: func(context.Context) string { return cp.server.URL }, + } + + p, _ := reconcileSettled(t, r, "tenant-a") + + if p.Status.UUID != testPoolUUID { + t.Fatalf("status.uuid = %q, want %q -- the create never reached the resolved endpoint", + p.Status.UUID, testPoolUUID) + } + if cp.posts != 1 { + t.Errorf("the control plane was asked to create the pool %d times, want 1", cp.posts) + } +} + +// Once the pool's cluster has its own recorded secret, the create +// authenticates as that cluster instead of as this operator's own Kubernetes +// identity -- the only way to reach a control plane a different Kubernetes +// cluster runs, since a TokenReview can never cross that boundary. +func TestAPoolAuthenticatesAsItsClusterOnceTheSecretIsKnown(t *testing.T) { + cp, rec := newControlPlane(t), &recorder{} + secret := &corev1.Secret{ + ObjectMeta: metav1.ObjectMeta{ + Name: "simplyblock-cluster-" + testCluster, + Namespace: testNamespace, + }, + Data: map[string][]byte{"secret": []byte("cluster-own-secret")}, + } + c := newClient(t, newCluster(testClusterUUID), newPool("tenant-a"), secret) + r := &StoragePoolReconciler{ + Client: c, + Scheme: testScheme(t), + Recorder: rec, + EndpointResolver: func(context.Context) string { return cp.server.URL }, + } + + reconcileSettled(t, r, "tenant-a") + + if cp.lastAuth != "Bearer cluster-own-secret" { + t.Errorf("authorization = %q, want the cluster's own secret", cp.lastAuth) + } +} + +// With no secret recorded yet (a cluster still being created), the call still +// goes out rather than being held -- the operator's own identity is what it +// falls back to, the same as before this fix. +func TestAPoolCallsWithNoClusterSecretDoesNotBlock(t *testing.T) { + cp, rec := newControlPlane(t), &recorder{} + c := newClient(t, newCluster(testClusterUUID), newPool("tenant-a")) + r := &StoragePoolReconciler{ + Client: c, + Scheme: testScheme(t), + Recorder: rec, + EndpointResolver: func(context.Context) string { return cp.server.URL }, + } + + p, _ := reconcileSettled(t, r, "tenant-a") + + if p.Status.UUID != testPoolUUID { + t.Errorf("status.uuid = %q, want %q", p.Status.UUID, testPoolUUID) + } +} diff --git a/operator/internal/controllers/pool/helpers_test.go b/operator/internal/controllers/pool/helpers_test.go index 67bb29dea..314f1cdfd 100644 --- a/operator/internal/controllers/pool/helpers_test.go +++ b/operator/internal/controllers/pool/helpers_test.go @@ -130,6 +130,10 @@ type controlPlane struct { deletes int hosts []string + // lastAuth is the Authorization header of the most recently handled + // request, which is what a test checks a call authenticated with. + lastAuth string + server *httptest.Server } @@ -150,6 +154,7 @@ func (cp *controlPlane) client() func() *webapi.Client { } func (cp *controlPlane) handle(w http.ResponseWriter, r *http.Request) { + cp.lastAuth = r.Header.Get("Authorization") switch { case r.Method == http.MethodPost && hasSuffix(r.URL.Path, "/host"): cp.hosts = append(cp.hosts, "added") diff --git a/operator/internal/controllers/pool/storagepool_controller.go b/operator/internal/controllers/pool/storagepool_controller.go index 3e5f59a0a..74fe8c1f8 100644 --- a/operator/internal/controllers/pool/storagepool_controller.go +++ b/operator/internal/controllers/pool/storagepool_controller.go @@ -51,6 +51,7 @@ import ( "github.com/simplyblock/atlas/ptr" simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/controlplane" "github.com/simplyblock/simplyblock-operator/internal/cpinformer" "github.com/simplyblock/simplyblock-operator/internal/utils" "github.com/simplyblock/simplyblock-operator/internal/webapi" @@ -83,15 +84,33 @@ type StoragePoolReconciler struct { VolumeScopes *cpinformer.ScopeSet // NewAPIClient builds the control-plane client. It is a field so a test can - // point the reconciler at a mock server; nil selects the real one. + // point the reconciler at a mock server; nil selects the real one, resolved + // through EndpointResolver. NewAPIClient func() *webapi.Client + // EndpointResolver answers where the control plane currently is, resolved + // per reconcile the same way internal/controllers/cluster and + // internal/controllers/node do (controlplane.NewEndpointResolver). Nil, or a + // resolver that answers nothing, means the startup client -- this cluster's + // own in-cluster address, or SIMPLYBLOCK_WEBAPI_BASE_URL -- is the only one, + // which is what a standalone deployment and every pre-existing test still + // get. Without this, a pool on a ControlPlane.spec.source.managed deployment + // could never reach its control plane at all: webapi.NewClient() defaults to + // a Service this Kubernetes cluster never runs. + EndpointResolver controlplane.EndpointResolver + // reportedMissingNodes remembers which unresolved spec.allowedNodes entries // have already been announced, so a name left behind by a removed node is // one event rather than one per reconcile forever. The authored list is // deliberately not pruned, so without this the event would repeat for the // life of the pool. reportedMissingNodes sync.Map + + // mu guards startupClient and resolvedClient, which clientFor rebuilds when + // the resolved endpoint changes. + mu sync.Mutex + startupClient *webapi.Client + resolvedClient *webapi.Client } // poolDTO is the control plane's storage-pool response. @@ -212,7 +231,14 @@ func (r *StoragePoolReconciler) Reconcile(ctx context.Context, req ctrl.Request) return ctrl.Result{}, err } - api := r.apiClient() + // Authenticates as this cluster, using its own recorded secret, for the + // same reason internal/controllers/cluster's sync() does: the control + // plane this cluster belongs to may be a ControlPlane.spec.source.managed + // one, on a different Kubernetes cluster than this operator. + if secret, err := r.clusterSecret(ctx, cluster); err == nil && secret != "" { + ctx = webapi.WithBearerToken(ctx, secret) + } + api := r.apiClient(ctx) if !p.DeletionTimestamp.IsZero() { return r.reconcileDeletion(ctx, p, api, clusterUUID) @@ -806,11 +832,52 @@ func (r *StoragePoolReconciler) event( r.Recorder.Eventf(object, nil, eventType, reason, reason, format, args...) } -func (r *StoragePoolReconciler) apiClient() *webapi.Client { +// apiClient is the client one reconcile call uses, resolved through +// EndpointResolver the same way internal/controllers/cluster's +// httpControlPlane.clientFor is. +func (r *StoragePoolReconciler) apiClient(ctx context.Context) *webapi.Client { if r.NewAPIClient != nil { return r.NewAPIClient() } - return webapi.NewClient() + + r.mu.Lock() + defer r.mu.Unlock() + + if r.startupClient == nil { + r.startupClient = webapi.NewClient() + } + if r.EndpointResolver == nil { + return r.startupClient + } + endpoint := r.EndpointResolver(ctx) + if endpoint == "" || endpoint == r.startupClient.BaseURL { + return r.startupClient + } + if r.resolvedClient == nil || r.resolvedClient.BaseURL != endpoint { + r.resolvedClient = webapi.NewClient(endpoint) + } + return r.resolvedClient +} + +// clusterSecret reads the credential StorageClusterReconciler.persist wrote +// for the pool's cluster, so a call scoped to it authenticates as that +// cluster instead of as this operator's own Kubernetes identity -- the only +// way to reach a control plane a different Kubernetes cluster runs, since a +// TokenReview can never cross that boundary. Mirrors +// internal/controllers/cluster's identically named method and +// internal/controllers/node's clusterSecretByName. +func (r *StoragePoolReconciler) clusterSecret( + ctx context.Context, cluster *simplyblockv1alpha2.StorageCluster, +) (string, error) { + var secret corev1.Secret + key := client.ObjectKey{ + Name: fmt.Sprintf("simplyblock-cluster-%s", cluster.Name), + Namespace: cluster.Namespace, + } + if err := r.Get(ctx, key, &secret); err != nil { + return "", err + } + return string(secret.Data["secret"]), nil } // SetupWithManager registers the reconciler and the one watch that is not on the From 72a5832246e18811aa157c5f5772edede3ff575b Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Tue, 22 Sep 2026 18:54:13 +0100 Subject: [PATCH 136/206] fix cluster naming --- .../deployment/cluster_disambiguation_test.go | 110 ++++++++++++++++++ .../deployment/operatorops_controller.go | 69 ++++++++++- 2 files changed, 177 insertions(+), 2 deletions(-) create mode 100644 operator/internal/controllers/deployment/cluster_disambiguation_test.go diff --git a/operator/internal/controllers/deployment/cluster_disambiguation_test.go b/operator/internal/controllers/deployment/cluster_disambiguation_test.go new file mode 100644 index 000000000..d7b032346 --- /dev/null +++ b/operator/internal/controllers/deployment/cluster_disambiguation_test.go @@ -0,0 +1,110 @@ +// Whether the cluster name a discovery run proposes carries something that +// distinguishes this Kubernetes cluster from another one pointed at the same +// control plane. +// +// InitialDiscoveryName is deliberately identical on every install, and that is +// fine for the OperatorOps and the ClusterDeploymentConfig discovery writes -- +// both live in this Kubernetes cluster's own API server, where the name +// collides with nothing else. The StorageCluster the document proposes does +// not stay local: it is registered on the control plane by name +// (StorageClusterReconciler.creationParams), and two separate Kubernetes +// clusters pointed at the same control plane (ControlPlane.spec.source.managed) +// otherwise propose the identical one. StorageClusterReconciler.postCluster's +// create-conflict fallback exists to resume a retried create of the SAME +// cluster and cannot tell that apart from a name that belongs to an entirely +// different Kubernetes cluster's own -- so it silently adopts the other one. + +package deployment + +import ( + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/types" +) + +// kubeSystemNamespace is the fixture that gives a fake cluster its own +// identity, the same as the real kube-system Namespace every Kubernetes +// cluster carries from the moment its API server first comes up. +func kubeSystemNamespace(uid types.UID) *corev1.Namespace { + return &corev1.Namespace{ + ObjectMeta: metav1.ObjectMeta{Name: metav1.NamespaceSystem, UID: uid}, + } +} + +// writtenClusterName drives a discovery run to completion and returns the +// cluster name its document proposes. +func writtenClusterName(t *testing.T, r *runner) string { + t.Helper() + r.step() // start + r.step() // inspect + r.step() // probing: creates the Jobs + + if err := r.client.Create(t.Context(), + reportConfigMap(t, "worker-1", "0000:5e:00.0")); err != nil { + t.Fatalf("write a report: %v", err) + } + + r.step() // probing: sees the report, moves to Writing + r.step() // writing + + configs := r.configs() + if len(configs) != 1 { + t.Fatalf("wrote %d documents, want 1", len(configs)) + } + if configs[0].Spec.Cluster == nil { + t.Fatalf("the document names no cluster template") + } + return configs[0].Spec.Cluster.Name +} + +// Two Kubernetes clusters running the identical, unmodified bootstrap +// discovery -- same OperatorOps name, same fleet shape -- propose different +// cluster names once each has its own kube-system identity. This is the +// property that keeps a second cluster's discovery from ever being mistaken, +// on the shared control plane, for the first cluster's own. +func TestTwoKubernetesClustersProposeDifferentClusterNames(t *testing.T) { + clusterA := newRunner(t, + discoverRun(nil), worker("worker-1"), kubeSystemNamespace("11111111-1111-1111-1111-111111111111")) + clusterB := newRunner(t, + discoverRun(nil), worker("worker-1"), kubeSystemNamespace("22222222-2222-2222-2222-222222222222")) + + nameA := writtenClusterName(t, clusterA) + nameB := writtenClusterName(t, clusterB) + + if nameA == nameB { + t.Fatalf("both clusters proposed %q, which is exactly the collision this fixes", nameA) + } +} + +// The same kube-system identity proposes the same cluster name on a second +// run: the suffix is derived, not random, which is what keeps the "create is +// idempotent by name" property bootstrap.go documents. +func TestTheSameKubernetesClusterProposesTheSameNameAcrossRuns(t *testing.T) { + const uid = types.UID("33333333-3333-3333-3333-333333333333") + + first := newRunner(t, discoverRun(nil), worker("worker-1"), kubeSystemNamespace(uid)) + second := newRunner(t, discoverRun(nil), worker("worker-1"), kubeSystemNamespace(uid)) + + nameFirst := writtenClusterName(t, first) + nameSecond := writtenClusterName(t, second) + + if nameFirst != nameSecond { + t.Errorf("proposed %q then %q for the same kube-system identity", nameFirst, nameSecond) + } +} + +// Without a readable kube-system Namespace -- every fixture in this package +// before this file, and any real cluster whose RBAC has not yet caught up -- +// the proposed name is exactly what it always was. This is what keeps the fix +// additive. +func TestWithNoKubeSystemNamespaceTheNameIsUnchanged(t *testing.T) { + r := newRunner(t, discoverRun(nil), worker("worker-1")) + + got := writtenClusterName(t, r) + + if got != configNamePrefix+opsName+clusterNameSuffix { + t.Errorf("proposed %q, want %q", got, configNamePrefix+opsName+clusterNameSuffix) + } +} diff --git a/operator/internal/controllers/deployment/operatorops_controller.go b/operator/internal/controllers/deployment/operatorops_controller.go index f3137604f..3e4e63ee4 100644 --- a/operator/internal/controllers/deployment/operatorops_controller.go +++ b/operator/internal/controllers/deployment/operatorops_controller.go @@ -29,6 +29,7 @@ import ( "errors" "fmt" "slices" + "strings" "time" batchv1 "k8s.io/api/batch/v1" @@ -70,6 +71,19 @@ const ( // when the caller named nothing. configNamePrefix = "discovered-" clusterNameSuffix = "-cluster" + + // clusterNameLimit is the longest name a StorageCluster (and so the + // backend cluster it registers under) can carry. Restated here because + // draftFor enforces it directly rather than relying on the apiserver to + // refuse an over-length name later, which is what CreatingCluster would do + // with no way to say why. + clusterNameLimit = 63 + + // disambiguatorLength is how much of the kube-system Namespace's UID + // clusterDisambiguator keeps. Long enough that two Kubernetes clusters + // collide by chance only astronomically rarely, short enough that it + // barely touches clusterNameLimit. + disambiguatorLength = 8 ) // The reasons a discovery run emits. They are constants rather than literals at @@ -132,6 +146,7 @@ type OperatorOpsReconciler struct { // +kubebuilder:rbac:groups=batch,resources=jobs,verbs=get;list;watch;create // +kubebuilder:rbac:groups="",resources=nodes,verbs=get;list;watch // +kubebuilder:rbac:groups="",resources=configmaps,verbs=get;list;watch +// +kubebuilder:rbac:groups="",resources=namespaces,verbs=get // +kubebuilder:rbac:groups="",resources=events,verbs=create;patch // Reconcile advances one operator operation by one step. @@ -582,7 +597,7 @@ func (r *OperatorOpsReconciler) write( plan.ExplainWithin(maxEventMessage-len(refusalPreamble))) } - config, notes := r.draftFor(ops, spec, plan) + config, notes := r.draftFor(ctx, ops, spec, plan) if err := r.Create(ctx, config); err != nil { if !apierrors.IsAlreadyExists(err) { return false, err @@ -638,9 +653,45 @@ func refuseUnreadableFilter(spec *simplyblockv1alpha2.DiscoverSpec) error { return nil } +// clusterDisambiguator is a short, stable-per-Kubernetes-cluster suffix for the +// cluster name a discovery run proposes. +// +// InitialDiscoveryName is deliberately identical on every install, which is +// fine for the OperatorOps and the ClusterDeploymentConfig it writes -- both +// live in this Kubernetes cluster's own API server, where the name collides +// with nothing else. The StorageCluster the document proposes does not stay +// local, though: it is registered on the control plane by name +// (StorageClusterReconciler.creationParams), and a second Kubernetes cluster +// pointed at the same one (ControlPlane.spec.source.managed) would otherwise +// propose the identical name. StorageClusterReconciler.postCluster's +// create-conflict fallback exists to resume a retried create of the SAME +// cluster and cannot tell that apart from a name that belongs to an entirely +// different Kubernetes cluster's own -- so it silently adopts the other one. +// +// The kube-system Namespace's UID is the closest thing a Kubernetes cluster has +// to its own fixed identity: present from the moment its API server first +// comes up, and never reissued afterward. An unreadable namespace answers the +// empty string rather than an error -- a cluster this cannot be read from is +// not going to succeed at registering on the control plane a moment later +// either, and the caller falls back to the name exactly as it was before this +// existed. +func (r *OperatorOpsReconciler) clusterDisambiguator(ctx context.Context) string { + var ns corev1.Namespace + key := client.ObjectKey{Name: metav1.NamespaceSystem} + if err := r.Get(ctx, key, &ns); err != nil || ns.UID == "" { + return "" + } + id := strings.ReplaceAll(string(ns.UID), "-", "") + if len(id) > disambiguatorLength { + id = id[:disambiguatorLength] + } + return id +} + // draftFor builds the document, and the notes explaining the numbers in it that // were not read off the hardware. func (r *OperatorOpsReconciler) draftFor( + ctx context.Context, ops *simplyblockv1alpha2.OperatorOps, spec *simplyblockv1alpha2.DiscoverSpec, plan discoverypkg.Plan, @@ -679,7 +730,21 @@ func (r *OperatorOpsReconciler) draftFor( notes = append(notes, fmt.Sprintf( "the draft grows the existing cluster %s, so it proposes no cluster layout", spec.ClusterRef)) } else { - template := discoverypkg.ClusterTemplateFor(name+clusterNameSuffix, plan) + clusterName := name + clusterNameSuffix + if disambiguator := r.clusterDisambiguator(ctx); disambiguator != "" { + base := name + // Truncated so the disambiguated name still fits, the same way a + // name too long for the limit already does not (§ gap G-11): this + // does not newly break a base name that already did not fit, only + // keeps this suffix from being the reason a borderline one no + // longer does. + room := clusterNameLimit - len(clusterNameSuffix) - len(disambiguator) - 1 + if room > 0 && len(base) > room { + base = base[:room] + } + clusterName = base + "-" + disambiguator + clusterNameSuffix + } + template := discoverypkg.ClusterTemplateFor(clusterName, plan) config.Spec.Cluster = template.Template notes = append(notes, template.Notes...) } From b3bb56c97c16052884153b47cfa9c3c8e2874df9 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Tue, 22 Sep 2026 19:52:19 +0100 Subject: [PATCH 137/206] Fix cross-cluster managed control plane support: admin auth, endpoint resolution, node addressing, and image defaults --- .../storage.simplyblock.io_controlplanes.yaml | 14 ++ .../templates/controlplane_cr.yaml | 3 + .../charts/simplyblock-operator/values.yaml | 7 + operator/api/v1alpha2/controlplane_types.go | 14 ++ .../storage.simplyblock.io_controlplanes.yaml | 14 ++ .../node/managed_node_address_test.go | 123 ++++++++++++++++++ operator/internal/controllers/node/migrate.go | 2 +- .../controllers/node/provisioning_test.go | 10 +- .../node/storagenode_controller.go | 5 +- .../internal/controllers/node/workload.go | 48 ++++++- .../controllers/node/workload_controller.go | 11 +- .../node/workload_controller_test.go | 103 +++++++++++++++ .../storage.simplyblock.io_controlplanes.yaml | 14 ++ 13 files changed, 352 insertions(+), 16 deletions(-) create mode 100644 operator/internal/controllers/node/managed_node_address_test.go create mode 100644 operator/internal/controllers/node/workload_controller_test.go diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_controlplanes.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_controlplanes.yaml index 077d5ea80..52a1707c4 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_controlplanes.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_controlplanes.yaml @@ -482,6 +482,20 @@ spec: so a loopback or link-local address is rejected. pattern: ^https?://[a-zA-Z0-9.-]+(:[0-9]{1,5})?(/.*)?$ type: string + storageNodeImage: + description: |- + StorageNodeImage is the storage-node image a StorageCluster on this + Kubernetes cluster defaults to when its own spec.storageNodes.image is + unset. A local control plane's own spec.source.local.image doubles as + this default (StorageNodeWorkloadReconciler.image), because a + self-hosted deployment's control plane and its storage nodes are one + release. A managed one is a different Kubernetes cluster's install and + says nothing about what this cluster's storage nodes should run, so + there is no equivalent to fall back to without this field -- every + StorageCluster on a managed deployment must get an image from here or + from its own spec. + pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ + type: string required: - endpoint type: object diff --git a/helm-charts/charts/simplyblock-operator/templates/controlplane_cr.yaml b/helm-charts/charts/simplyblock-operator/templates/controlplane_cr.yaml index 78ad25aa1..7c31aa4d8 100644 --- a/helm-charts/charts/simplyblock-operator/templates/controlplane_cr.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/controlplane_cr.yaml @@ -92,5 +92,8 @@ spec: caBundleSecretRef: name: {{ .Values.controlplane.managed.caBundleSecretRef | quote }} {{- end }} + {{- if .Values.controlplane.managed.storageNodeImage }} + storageNodeImage: {{ .Values.controlplane.managed.storageNodeImage | quote }} + {{- end }} {{- end }} diff --git a/helm-charts/charts/simplyblock-operator/values.yaml b/helm-charts/charts/simplyblock-operator/values.yaml index b8325f055..651b3ad91 100644 --- a/helm-charts/charts/simplyblock-operator/values.yaml +++ b/helm-charts/charts/simplyblock-operator/values.yaml @@ -270,6 +270,13 @@ controlplane: # Secret holding the CA the endpoint is verified against, under `ca.crt` or # `tls.crt`. Empty uses the system trust store. Named and unusable fails. caBundleSecretRef: "" + # Default storage-node image for a StorageCluster on this Kubernetes + # cluster whose own spec.storageNodes.image is unset. The managed control + # plane is a different Kubernetes cluster's install and says nothing about + # what this cluster's storage nodes should run, so there is no equivalent + # to the local profile's spec.image to fall back to without this -- + # required in practice for any StorageCluster that does not set its own. + storageNodeImage: "" # Accept a pod's service-account token from the CSI driver's two accounts # instead of requiring the static cluster secret. This is the control diff --git a/operator/api/v1alpha2/controlplane_types.go b/operator/api/v1alpha2/controlplane_types.go index 10b068de0..f3e34d8a5 100644 --- a/operator/api/v1alpha2/controlplane_types.go +++ b/operator/api/v1alpha2/controlplane_types.go @@ -270,6 +270,20 @@ type ManagedControlPlane struct { // is verified against. Absent means the system trust store. // +optional CABundleSecretRef *corev1.LocalObjectReference `json:"caBundleSecretRef,omitempty"` + + // StorageNodeImage is the storage-node image a StorageCluster on this + // Kubernetes cluster defaults to when its own spec.storageNodes.image is + // unset. A local control plane's own spec.source.local.image doubles as + // this default (StorageNodeWorkloadReconciler.image), because a + // self-hosted deployment's control plane and its storage nodes are one + // release. A managed one is a different Kubernetes cluster's install and + // says nothing about what this cluster's storage nodes should run, so + // there is no equivalent to fall back to without this field -- every + // StorageCluster on a managed deployment must get an image from here or + // from its own spec. + // +kubebuilder:validation:Pattern=`^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$` + // +optional + StorageNodeImage string `json:"storageNodeImage,omitempty"` } // ControlPlaneSource selects whether this cluster hosts its control plane or is diff --git a/operator/config/crd/bases/storage.simplyblock.io_controlplanes.yaml b/operator/config/crd/bases/storage.simplyblock.io_controlplanes.yaml index 077d5ea80..52a1707c4 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_controlplanes.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_controlplanes.yaml @@ -482,6 +482,20 @@ spec: so a loopback or link-local address is rejected. pattern: ^https?://[a-zA-Z0-9.-]+(:[0-9]{1,5})?(/.*)?$ type: string + storageNodeImage: + description: |- + StorageNodeImage is the storage-node image a StorageCluster on this + Kubernetes cluster defaults to when its own spec.storageNodes.image is + unset. A local control plane's own spec.source.local.image doubles as + this default (StorageNodeWorkloadReconciler.image), because a + self-hosted deployment's control plane and its storage nodes are one + release. A managed one is a different Kubernetes cluster's install and + says nothing about what this cluster's storage nodes should run, so + there is no equivalent to fall back to without this field -- every + StorageCluster on a managed deployment must get an image from here or + from its own spec. + pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ + type: string required: - endpoint type: object diff --git a/operator/internal/controllers/node/managed_node_address_test.go b/operator/internal/controllers/node/managed_node_address_test.go new file mode 100644 index 000000000..3fd97d3af --- /dev/null +++ b/operator/internal/controllers/node/managed_node_address_test.go @@ -0,0 +1,123 @@ +// Whether NodeAddress gives the control plane something it can actually +// reach: the per-pod Service DNS name a control plane on this Kubernetes +// cluster resolves itself, or the worker's own real address when the control +// plane runs on a different one entirely and could never resolve that DNS. + +package node + +import ( + "context" + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/testsupport" + "github.com/simplyblock/simplyblock-operator/internal/utils" +) + +func managedControlPlaneSingleton() *simplyblockv1alpha2.ControlPlane { + return &simplyblockv1alpha2.ControlPlane{ + ObjectMeta: metav1.ObjectMeta{Name: SingletonControlPlaneName, Namespace: opsNamespace}, + Spec: simplyblockv1alpha2.ControlPlaneSpec{ + Source: simplyblockv1alpha2.ControlPlaneSource{ + Managed: &simplyblockv1alpha2.ManagedControlPlane{ + Endpoint: "https://hub.example.com:5000", + }, + }, + }, + } +} + +func localControlPlaneSingleton() *simplyblockv1alpha2.ControlPlane { + return &simplyblockv1alpha2.ControlPlane{ + ObjectMeta: metav1.ObjectMeta{Name: SingletonControlPlaneName, Namespace: opsNamespace}, + Spec: simplyblockv1alpha2.ControlPlaneSpec{ + Source: simplyblockv1alpha2.ControlPlaneSource{ + Local: &simplyblockv1alpha2.LocalControlPlane{Image: "docker.io/simplyblock/simplyblock:1.0"}, + }, + }, + } +} + +func workerNode(name, internalIP string) *corev1.Node { + return &corev1.Node{ + ObjectMeta: metav1.ObjectMeta{Name: name}, + Status: corev1.NodeStatus{ + Addresses: []corev1.NodeAddress{ + {Type: corev1.NodeInternalIP, Address: internalIP}, + }, + }, + } +} + +// A managed control plane runs on a Kubernetes cluster that can never resolve +// this cluster's own Service DNS, so it is given the worker's real address -- +// one the storage-node-api pod already answers on directly, since it runs +// with hostNetwork. +func TestNodeAddressUsesTheWorkersRealIPWhenTheControlPlaneIsManaged(t *testing.T) { + scheme := testsupport.NewScheme(t, corev1.AddToScheme) + c := fake.NewClientBuilder().WithScheme(scheme). + WithObjects(managedControlPlaneSingleton(), workerNode(opsWorker, "192.168.10.112")). + Build() + w := &Workload{Client: c} + + got := w.NodeAddress(context.Background(), opsWorker, opsNamespace) + + if want := "192.168.10.112:5000"; got != want { + t.Errorf("NodeAddress = %q, want %q", got, want) + } +} + +// A local control plane resolves the per-pod DNS name itself, and this is the +// existing, well-tested precondition for that: nothing about a same-cluster +// deployment changes. +func TestNodeAddressKeepsTheDNSNameForALocalControlPlane(t *testing.T) { + scheme := testsupport.NewScheme(t, corev1.AddToScheme) + c := fake.NewClientBuilder().WithScheme(scheme). + WithObjects(localControlPlaneSingleton(), workerNode(opsWorker, "192.168.10.112")). + Build() + w := &Workload{Client: c} + + got := w.NodeAddress(context.Background(), opsWorker, opsNamespace) + + if want := utils.StorageNodeSetAPIAddress(opsWorker, opsNamespace); got != want { + t.Errorf("NodeAddress = %q, want the per-pod DNS name %q", got, want) + } +} + +// No ControlPlane singleton at all -- every fixture in this package before +// this file -- falls back to exactly the address it always resolved to. This +// is what keeps the fix additive. +func TestNodeAddressFallsBackToDNSWhenTheControlPlaneCannotBeRead(t *testing.T) { + scheme := testsupport.NewScheme(t, corev1.AddToScheme) + c := fake.NewClientBuilder().WithScheme(scheme). + WithObjects(workerNode(opsWorker, "192.168.10.112")). + Build() + w := &Workload{Client: c} + + got := w.NodeAddress(context.Background(), opsWorker, opsNamespace) + + if want := utils.StorageNodeSetAPIAddress(opsWorker, opsNamespace); got != want { + t.Errorf("NodeAddress = %q, want the per-pod DNS name %q", got, want) + } +} + +// A managed control plane but a worker Node this reader cannot find (a stale +// cache, a name that does not match) falls back the same way, rather than +// handing the control plane an empty or malformed address. +func TestNodeAddressFallsBackToDNSWhenTheWorkerNodeCannotBeRead(t *testing.T) { + scheme := testsupport.NewScheme(t, corev1.AddToScheme) + c := fake.NewClientBuilder().WithScheme(scheme). + WithObjects(managedControlPlaneSingleton()). + Build() + w := &Workload{Client: c} + + got := w.NodeAddress(context.Background(), opsWorker, opsNamespace) + + if want := utils.StorageNodeSetAPIAddress(opsWorker, opsNamespace); got != want { + t.Errorf("NodeAddress = %q, want the per-pod DNS name %q", got, want) + } +} diff --git a/operator/internal/controllers/node/migrate.go b/operator/internal/controllers/node/migrate.go index c1cb8c835..b1a28a565 100644 --- a/operator/internal/controllers/node/migrate.go +++ b/operator/internal/controllers/node/migrate.go @@ -160,7 +160,7 @@ func (r *StorageNodeOpsReconciler) migrateRelocate( force = *ops.Spec.Force } params := RestartParams{ - NodeAddress: r.Workload.NodeAddress(target, node.Namespace), + NodeAddress: r.Workload.NodeAddress(ctx, target, node.Namespace), Force: force, ReattachVolume: boolValue(ops.Spec.ReattachVolume), NewSsdPcie: ops.Spec.MigrateParams().NewSsdPcie, diff --git a/operator/internal/controllers/node/provisioning_test.go b/operator/internal/controllers/node/provisioning_test.go index f36429127..135a271af 100644 --- a/operator/internal/controllers/node/provisioning_test.go +++ b/operator/internal/controllers/node/provisioning_test.go @@ -144,7 +144,7 @@ func TestTheAddCarriesWhatTheNodeSaysAboutItself(t *testing.T) { } r, _ := aSteadyNode(t, aControlPlane()) - params := r.addParams(node, cluster) + params := r.addParams(context.Background(), node, cluster) if params.SPDKImage != "example.test/spdk:v1" || params.SPDKProxyImage != "example.test/spdk-proxy:v1" || @@ -181,7 +181,7 @@ func TestTheAddCarriesWhatTheNodeSaysAboutItself(t *testing.T) { func TestAnUnstatedJournalShareIsTheDefaultAndTheCountIsNot(t *testing.T) { r, _ := aSteadyNode(t, aControlPlane()) - params := r.addParams(anUnprovisionedNode(stepPosting), anOpsCluster()) + params := r.addParams(context.Background(), anUnprovisionedNode(stepPosting), anOpsCluster()) if params.JMPercent != 3 { t.Errorf("the journal share is %d%%, want the default 3", params.JMPercent) @@ -201,18 +201,18 @@ func TestOnlyAFaultGroupThatIsANumberIsSentToTheControlPlane(t *testing.T) { numbered := anUnprovisionedNode(stepPosting) numbered.Spec.Config.FailureDomain = "2" - if index := r.addParams(numbered, anOpsCluster()).FailureDomain; index == nil || *index != 2 { + if index := r.addParams(context.Background(), numbered, anOpsCluster()).FailureDomain; index == nil || *index != 2 { t.Errorf("failureDomain = %v, want the index the label spells", index) } named := anUnprovisionedNode(stepPosting) named.Spec.Config.FailureDomain = "rack-1" - if index := r.addParams(named, anOpsCluster()).FailureDomain; index != nil { + if index := r.addParams(context.Background(), named, anOpsCluster()).FailureDomain; index != nil { t.Errorf("failureDomain = %v, want none: the control plane has no field for a name", index) } - if index := r.addParams(anUnprovisionedNode(stepPosting), anOpsCluster()).FailureDomain; index != nil { + if index := r.addParams(context.Background(), anUnprovisionedNode(stepPosting), anOpsCluster()).FailureDomain; index != nil { t.Errorf("failureDomain = %v, want none for a node that declares no group", index) } } diff --git a/operator/internal/controllers/node/storagenode_controller.go b/operator/internal/controllers/node/storagenode_controller.go index b864087cc..b41cb0ad3 100644 --- a/operator/internal/controllers/node/storagenode_controller.go +++ b/operator/internal/controllers/node/storagenode_controller.go @@ -774,7 +774,7 @@ func (r *StorageNodeReconciler) postNode( node *simplyblockv1alpha2.StorageNode, cluster *simplyblockv1alpha2.StorageCluster, ) error { - params := r.addParams(node, cluster) + params := r.addParams(ctx, node, cluster) if err := r.API.AddNode(ctx, cluster.Status.UUID, params); err != nil { return fmt.Errorf("add node %s on worker %s: %w", node.Name, node.Spec.WorkerNode, err) @@ -1426,6 +1426,7 @@ func (r *StorageNodeReconciler) upgradeAdoption( // addParams is what the node-add call carries. The node describes itself, so every // value but the subsystem cap comes from its own spec.config (§3.1). func (r *StorageNodeReconciler) addParams( + ctx context.Context, node *simplyblockv1alpha2.StorageNode, cluster *simplyblockv1alpha2.StorageCluster, ) utils.StorageNodeSetAddParams { @@ -1436,7 +1437,7 @@ func (r *StorageNodeReconciler) addParams( } params := utils.StorageNodeSetAddParams{ - NodeAddress: r.Workload.NodeAddress(node.Spec.WorkerNode, node.Namespace), + NodeAddress: r.Workload.NodeAddress(ctx, node.Spec.WorkerNode, node.Namespace), InterfaceName: workload.MgmtInterface, SPDKImage: config.SpdkImage, SPDKProxyImage: config.SpdkProxyImage, diff --git a/operator/internal/controllers/node/workload.go b/operator/internal/controllers/node/workload.go index 957e02a1d..bdab7dbf1 100644 --- a/operator/internal/controllers/node/workload.go +++ b/operator/internal/controllers/node/workload.go @@ -89,16 +89,52 @@ type Workload struct { ManagerNode string } -// NodeAddress is the per-pod DNS name the control plane is given as node_address -// when a node is added or restarted. +// NodeAddress is what the control plane is given as node_address when a node +// is added or restarted. // -// It is the precondition for both: a restart issued against a name that does not -// yet resolve fails name resolution inside the control plane, and the control -// plane's response to that is to reset the node to offline (§5.4). -func (w *Workload) NodeAddress(worker, namespace string) string { +// A local control plane resolves the per-pod DNS name itself, which is the +// existing precondition for both: a restart issued against a name that does +// not yet resolve fails name resolution inside the control plane, and the +// control plane's response to that is to reset the node to offline (§5.4). +// +// A managed control plane runs on a different Kubernetes cluster and can +// never resolve this cluster's own Service DNS, so it is given the worker's +// real, routable address instead -- one the storage-node-api pod already +// answers on directly, since it runs with hostNetwork (BuildStorageNodeDaemonSet). +// Reading the worker Node's own reported address is what keeps this additive: +// a deployment with no managed control plane takes exactly the path it always +// did. +func (w *Workload) NodeAddress(ctx context.Context, worker, namespace string) string { + if address, ok := w.managedNodeAddress(ctx, worker, namespace); ok { + return address + } return utils.StorageNodeSetAPIAddress(worker, namespace) } +// managedNodeAddress answers the worker's real address when the singleton +// ControlPlane names a managed control plane, and false otherwise -- including +// when the singleton or the worker Node cannot be read, since an operator that +// cannot tell falls back to the address that has always worked for a control +// plane this cluster hosts. +func (w *Workload) managedNodeAddress(ctx context.Context, worker, namespace string) (string, bool) { + var cp simplyblockv1alpha2.ControlPlane + key := client.ObjectKey{Namespace: namespace, Name: SingletonControlPlaneName} + if err := w.Get(ctx, key, &cp); err != nil || cp.Spec.Source.Managed == nil { + return "", false + } + + var node corev1.Node + if err := w.Get(ctx, client.ObjectKey{Name: worker}, &node); err != nil { + return "", false + } + for _, addr := range node.Status.Addresses { + if addr.Type == corev1.NodeInternalIP && addr.Address != "" { + return fmt.Sprintf("%s:5000", addr.Address), true + } + } + return "", false +} + // LabelWorker puts one worker into a cluster's storage plane and rewrites the // per-slot labels of every node on it. // diff --git a/operator/internal/controllers/node/workload_controller.go b/operator/internal/controllers/node/workload_controller.go index 42e652f6d..747c91f53 100644 --- a/operator/internal/controllers/node/workload_controller.go +++ b/operator/internal/controllers/node/workload_controller.go @@ -311,8 +311,15 @@ func (r *StorageNodeWorkloadReconciler) image( "spec.storageNodes.image is unset and ControlPlane %s cannot be read: %w", SingletonControlPlaneName, err) } - if managed := controlPlane.Spec.Source.Local; managed != nil && managed.Image != "" { - return managed.Image, nil + if local := controlPlane.Spec.Source.Local; local != nil && local.Image != "" { + return local.Image, nil + } + // A managed control plane is a different Kubernetes cluster's install and + // carries no image of its own here to fall back to -- spec.source.managed + // says where the control plane is, not what this cluster's storage nodes + // should run. StorageNodeImage is the only source of a default left. + if managed := controlPlane.Spec.Source.Managed; managed != nil && managed.StorageNodeImage != "" { + return managed.StorageNodeImage, nil } return "", fmt.Errorf( "spec.storageNodes.image is unset and ControlPlane %s states no managed image", diff --git a/operator/internal/controllers/node/workload_controller_test.go b/operator/internal/controllers/node/workload_controller_test.go new file mode 100644 index 000000000..44bfc26b0 --- /dev/null +++ b/operator/internal/controllers/node/workload_controller_test.go @@ -0,0 +1,103 @@ +// What a StorageCluster's storage-node DaemonSet defaults its image to when +// spec.storageNodes.image is unset. +// +// A local control plane's own image doubles as the default, since a +// self-hosted deployment's control plane and its storage nodes are one +// release. A managed control plane is a different Kubernetes cluster's +// install and says nothing about this one's storage nodes, so +// ManagedControlPlane.StorageNodeImage is the only source of a default there +// -- without it, the workload can never be built at all. + +package node + +import ( + "context" + "testing" + + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/testsupport" +) + +func aStorageCluster() *simplyblockv1alpha2.StorageCluster { + return &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{Name: opsCluster, Namespace: opsNamespace}, + } +} + +// An explicit spec.storageNodes.image always wins, whatever the ControlPlane +// says. +func TestImagePrefersTheClustersOwnSpec(t *testing.T) { + scheme := testsupport.NewScheme(t) + c := fake.NewClientBuilder().WithScheme(scheme). + WithObjects(localControlPlaneSingleton()). + Build() + r := &StorageNodeWorkloadReconciler{Client: c, Namespace: opsNamespace} + + cluster := aStorageCluster() + cluster.Spec.StorageNodes = &simplyblockv1alpha2.StorageNodesSpec{ + Image: "docker.io/simplyblock/simplyblock:from-the-spec", + } + + got, err := r.image(context.Background(), cluster) + if err != nil { + t.Fatalf("image: %v", err) + } + if want := "docker.io/simplyblock/simplyblock:from-the-spec"; got != want { + t.Errorf("image = %q, want %q", got, want) + } +} + +// A local control plane's own image is the default, exactly as it always was. +func TestImageDefaultsToTheLocalControlPlanesImage(t *testing.T) { + scheme := testsupport.NewScheme(t) + c := fake.NewClientBuilder().WithScheme(scheme). + WithObjects(localControlPlaneSingleton()). + Build() + r := &StorageNodeWorkloadReconciler{Client: c, Namespace: opsNamespace} + + got, err := r.image(context.Background(), aStorageCluster()) + if err != nil { + t.Fatalf("image: %v", err) + } + if want := "docker.io/simplyblock/simplyblock:1.0"; got != want { + t.Errorf("image = %q, want the local control plane's own %q", got, want) + } +} + +// A managed control plane names a storage-node image of its own, since its +// own spec.source.managed carries no image at all -- it is a different +// Kubernetes cluster's install. +func TestImageDefaultsToTheManagedControlPlanesStorageNodeImage(t *testing.T) { + cp := managedControlPlaneSingleton() + cp.Spec.Source.Managed.StorageNodeImage = "docker.io/simplyblock/simplyblock:from-managed" + + scheme := testsupport.NewScheme(t) + c := fake.NewClientBuilder().WithScheme(scheme).WithObjects(cp).Build() + r := &StorageNodeWorkloadReconciler{Client: c, Namespace: opsNamespace} + + got, err := r.image(context.Background(), aStorageCluster()) + if err != nil { + t.Fatalf("image: %v", err) + } + if want := "docker.io/simplyblock/simplyblock:from-managed"; got != want { + t.Errorf("image = %q, want the managed control plane's storage-node image %q", got, want) + } +} + +// A managed control plane that names no storage-node image either is a +// deployment nothing can build a DaemonSet for, and that has to fail loudly +// rather than build one with an empty image. +func TestImageErrorsWhenTheManagedControlPlaneNamesNoStorageNodeImage(t *testing.T) { + scheme := testsupport.NewScheme(t) + c := fake.NewClientBuilder().WithScheme(scheme). + WithObjects(managedControlPlaneSingleton()). + Build() + r := &StorageNodeWorkloadReconciler{Client: c, Namespace: opsNamespace} + + if _, err := r.image(context.Background(), aStorageCluster()); err == nil { + t.Error("image returned no error for a managed control plane naming no storage-node image") + } +} diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_controlplanes.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_controlplanes.yaml index 077d5ea80..52a1707c4 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_controlplanes.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_controlplanes.yaml @@ -482,6 +482,20 @@ spec: so a loopback or link-local address is rejected. pattern: ^https?://[a-zA-Z0-9.-]+(:[0-9]{1,5})?(/.*)?$ type: string + storageNodeImage: + description: |- + StorageNodeImage is the storage-node image a StorageCluster on this + Kubernetes cluster defaults to when its own spec.storageNodes.image is + unset. A local control plane's own spec.source.local.image doubles as + this default (StorageNodeWorkloadReconciler.image), because a + self-hosted deployment's control plane and its storage nodes are one + release. A managed one is a different Kubernetes cluster's install and + says nothing about what this cluster's storage nodes should run, so + there is no equivalent to fall back to without this field -- every + StorageCluster on a managed deployment must get an image from here or + from its own spec. + pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ + type: string required: - endpoint type: object From 6adbc0a42b1f2f9b013f095534c82077c1cfceda Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Tue, 22 Sep 2026 20:10:18 +0100 Subject: [PATCH 138/206] fixed linter issue --- operator/dist/install.yaml | 14 ++++++++++++++ .../cluster/controlplane_admin_test.go | 16 ++++++++++------ .../cluster/storageclusterops_controller_test.go | 6 +++--- 3 files changed, 27 insertions(+), 9 deletions(-) diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index 181bca865..a8c778da4 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -1647,6 +1647,20 @@ spec: so a loopback or link-local address is rejected. pattern: ^https?://[a-zA-Z0-9.-]+(:[0-9]{1,5})?(/.*)?$ type: string + storageNodeImage: + description: |- + StorageNodeImage is the storage-node image a StorageCluster on this + Kubernetes cluster defaults to when its own spec.storageNodes.image is + unset. A local control plane's own spec.source.local.image doubles as + this default (StorageNodeWorkloadReconciler.image), because a + self-hosted deployment's control plane and its storage nodes are one + release. A managed one is a different Kubernetes cluster's install and + says nothing about what this cluster's storage nodes should run, so + there is no equivalent to fall back to without this field -- every + StorageCluster on a managed deployment must get an image from here or + from its own spec. + pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ + type: string required: - endpoint type: object diff --git a/operator/internal/controllers/cluster/controlplane_admin_test.go b/operator/internal/controllers/cluster/controlplane_admin_test.go index ea668241f..a4f19e35e 100644 --- a/operator/internal/controllers/cluster/controlplane_admin_test.go +++ b/operator/internal/controllers/cluster/controlplane_admin_test.go @@ -21,7 +21,11 @@ import ( webapimock "github.com/simplyblock/simplyblock-operator/internal/webapi/mock" ) -const clusterAdminSpecPath = "../../../../shared/openapi.json" +const ( + clusterAdminSpecPath = "../../../../shared/openapi.json" + testAdminToken = "admin-token" + testAdminAuthHeader = "Bearer " + testAdminToken +) func TestCreateClusterAuthenticatesWithTheManagedAdminCredentialWhenOneResolves(t *testing.T) { mock := webapimock.NewSpecServerFromFile(t, clusterAdminSpecPath, false) @@ -33,7 +37,7 @@ func TestCreateClusterAuthenticatesWithTheManagedAdminCredentialWhenOneResolves( }) resolveEndpoint := func(context.Context) string { return mock.URL() } - resolveCredential := func(context.Context) (string, bool) { return "admin-token", true } + resolveCredential := func(context.Context) (string, bool) { return testAdminToken, true } api := NewControlPlane(resolveEndpoint, resolveCredential) if _, err := api.CreateCluster(context.Background(), utils.ClusterAddParams{Name: "b"}); err != nil { @@ -44,7 +48,7 @@ func TestCreateClusterAuthenticatesWithTheManagedAdminCredentialWhenOneResolves( if len(reqs) != 1 { t.Fatalf("expected one request, got %d", len(reqs)) } - if got := reqs[0].Headers["Authorization"]; got != "Bearer admin-token" { + if got := reqs[0].Headers["Authorization"]; got != testAdminAuthHeader { t.Errorf("authorization header = %q, want the managed admin credential", got) } } @@ -58,7 +62,7 @@ func TestClusterByNameAuthenticatesWithTheManagedAdminCredentialWhenOneResolves( }) resolveEndpoint := func(context.Context) string { return mock.URL() } - resolveCredential := func(context.Context) (string, bool) { return "admin-token", true } + resolveCredential := func(context.Context) (string, bool) { return testAdminToken, true } api := NewControlPlane(resolveEndpoint, resolveCredential) if _, _, err := api.ClusterByName(context.Background(), "b"); err != nil { @@ -69,7 +73,7 @@ func TestClusterByNameAuthenticatesWithTheManagedAdminCredentialWhenOneResolves( if len(reqs) != 1 { t.Fatalf("expected one request, got %d", len(reqs)) } - if got := reqs[0].Headers["Authorization"]; got != "Bearer admin-token" { + if got := reqs[0].Headers["Authorization"]; got != testAdminAuthHeader { t.Errorf("authorization header = %q, want the managed admin credential", got) } } @@ -98,7 +102,7 @@ func TestCreateClusterCarriesNoAdminCredentialWhenNoneResolves(t *testing.T) { if len(reqs) != 1 { t.Fatalf("expected one request, got %d", len(reqs)) } - if got := reqs[0].Headers["Authorization"]; got == "Bearer admin-token" { + if got := reqs[0].Headers["Authorization"]; got == testAdminAuthHeader { t.Errorf("authorization header = %q, want the client's own (empty) token, not the admin credential", got) } } diff --git a/operator/internal/controllers/cluster/storageclusterops_controller_test.go b/operator/internal/controllers/cluster/storageclusterops_controller_test.go index 7ee8714e5..6b1e10548 100644 --- a/operator/internal/controllers/cluster/storageclusterops_controller_test.go +++ b/operator/internal/controllers/cluster/storageclusterops_controller_test.go @@ -248,7 +248,7 @@ func TestAShutdownIssuesOneCallAndWaitsForTheCluster(t *testing.T) { cluster: func(string) (webapi.ClusterResponse, error) { reading := activeCluster() if !active { - reading.Status = "suspended" + reading.Status = statusSuspended } return reading, nil }, @@ -283,7 +283,7 @@ func TestAnOperationAuthenticatesAsItsClusterOnceItsSecretIsKnown(t *testing.T) cluster: func(string) (webapi.ClusterResponse, error) { reading := activeCluster() if !active { - reading.Status = "suspended" + reading.Status = statusSuspended } return reading, nil }, @@ -330,7 +330,7 @@ func TestARestartShutsDownThenStarts(t *testing.T) { reading.Status = status return reading, nil }, - shutdown: func(string) error { status = "suspended"; return nil }, + shutdown: func(string) error { status = statusSuspended; return nil }, start: func(string) error { status = utils.ClusterStatusActive; return nil }, } r := newOpsReconciler(t, api, &recorder{}, From a7e0aceb8721e7b9d151e40a6d857929ca533b84 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Wed, 23 Sep 2026 00:02:22 +0100 Subject: [PATCH 139/206] Pass raw cluster UUIDs through ResolveClusterUUID for cross-cluster ReplicationPair targets --- operator/internal/utils/objects.go | 10 ++++++++++ operator/internal/utils/objects_test.go | 15 +++++++++++++++ 2 files changed, 25 insertions(+) diff --git a/operator/internal/utils/objects.go b/operator/internal/utils/objects.go index ffc327121..4273bee03 100644 --- a/operator/internal/utils/objects.go +++ b/operator/internal/utils/objects.go @@ -99,6 +99,16 @@ func ResolveClusterUUID( clusterName string, ) (string, error) { + // A cluster on a different physical Kubernetes cluster has no local + // StorageCluster object to match by name -- the object only exists on + // that other cluster's own API server. The backend control plane is + // shared across clusters and addresses cluster pairs by this same UUID, + // so a raw UUID is passed through unresolved rather than requiring a + // local name match that can never succeed for a genuinely remote target. + if IsUUID(clusterName) { + return clusterName, nil + } + var clusters simplyblockv1alpha2.StorageClusterList if err := c.List(ctx, &clusters, client.InNamespace(namespace)); err != nil { return "", err diff --git a/operator/internal/utils/objects_test.go b/operator/internal/utils/objects_test.go index 4ea275f6c..348481501 100644 --- a/operator/internal/utils/objects_test.go +++ b/operator/internal/utils/objects_test.go @@ -138,6 +138,21 @@ func TestResolveClusterAndPoolUUID(t *testing.T) { t.Fatalf("ResolveClusterCRByUUID should fail for an unknown UUID") } + // A ReplicationPair naming a cluster that lives on a different physical + // Kubernetes cluster carries only that cluster's backend UUID -- there is + // no local StorageCluster object to match by name, because the object + // itself only exists on the other cluster's own API server. Passing a raw + // UUID through unresolved (rather than requiring a local name match) is + // what makes cross-cluster ReplicationPair authoring possible at all. + const remoteUUID = "e7afccef-1d5a-4b77-88aa-bfbe90d8b3a3" + remoteResolved, err := ResolveClusterUUID(ctx, c, "ns1", remoteUUID) + if err != nil { + t.Fatalf("ResolveClusterUUID should pass a raw UUID through even with no local match: %v", err) + } + if remoteResolved != remoteUUID { + t.Fatalf("ResolveClusterUUID got %q want pass-through of %q", remoteResolved, remoteUUID) + } + if _, err := ResolveClusterCRByUUID(ctx, c, "ns2", "uuid-a"); err == nil { t.Fatalf("ResolveClusterCRByUUID should not find a cluster from a different namespace") } From 5902d8b1fb7e2d913f8269bf151c9782be2e936a Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Wed, 23 Sep 2026 12:46:15 +0100 Subject: [PATCH 140/206] Resolve the control-plane endpoint for ReplicationPair/ReplicationPolicy on managed clusters --- operator/cmd/main.go | 10 ++-- .../controller/replicationpair_controller.go | 28 ++++++++++- .../replicationpair_controller_unit_test.go | 50 +++++++++++++++++++ .../replicationpolicy_controller.go | 24 ++++++++- .../replicationpolicy_controller_unit_test.go | 44 ++++++++++++++++ 5 files changed, 150 insertions(+), 6 deletions(-) diff --git a/operator/cmd/main.go b/operator/cmd/main.go index 67efb922a..584f9f0d2 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -853,8 +853,9 @@ func main() { os.Exit(1) } if err := (&controller.ReplicationPolicyReconciler{ - Client: mgr.GetClient(), - Scheme: mgr.GetScheme(), + Client: mgr.GetClient(), + Scheme: mgr.GetScheme(), + EndpointResolver: controlPlaneEndpoint, }).SetupWithManager(mgr); err != nil { setupLog.Error(err, "unable to create controller", "controller", "ReplicationPolicy") os.Exit(1) @@ -867,8 +868,9 @@ func main() { os.Exit(1) } if err := (&controller.ReplicationPairReconciler{ - Client: mgr.GetClient(), - Scheme: mgr.GetScheme(), + Client: mgr.GetClient(), + Scheme: mgr.GetScheme(), + EndpointResolver: controlPlaneEndpoint, }).SetupWithManager(mgr); err != nil { setupLog.Error(err, "unable to create controller", "controller", "ReplicationPair") os.Exit(1) diff --git a/operator/internal/controller/replicationpair_controller.go b/operator/internal/controller/replicationpair_controller.go index 20b8909ce..51febe229 100644 --- a/operator/internal/controller/replicationpair_controller.go +++ b/operator/internal/controller/replicationpair_controller.go @@ -30,6 +30,7 @@ import ( logf "sigs.k8s.io/controller-runtime/pkg/log" simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" + "github.com/simplyblock/simplyblock-operator/internal/controllers/controlplane" "github.com/simplyblock/simplyblock-operator/internal/utils" "github.com/simplyblock/simplyblock-operator/internal/webapi" ) @@ -46,6 +47,31 @@ const ( type ReplicationPairReconciler struct { client.Client Scheme *runtime.Scheme + + // EndpointResolver answers where the control plane currently is, resolved + // per reconcile the same way internal/controllers/pool's StoragePoolReconciler + // does. Nil means the default/SIMPLYBLOCK_WEBAPI_BASE_URL client is the only + // one, which is what a standalone deployment and every pre-existing test + // still get. Without this, a ReplicationPair reconciled on a + // ControlPlane.spec.source.managed member cluster can never reach its + // control plane: webapi.NewClient() defaults to a Service that cluster + // never runs (confirmed live: cross-cluster ReplicationPair authoring + // failed with "dial tcp: lookup simplyblock-webappapi ... no such host"). + EndpointResolver controlplane.EndpointResolver +} + +// apiClient resolves the control-plane client for one reconcile call, through +// EndpointResolver when set, exactly as StoragePoolReconciler.apiClient does. +func (r *ReplicationPairReconciler) apiClient(ctx context.Context) *webapi.Client { + if r.EndpointResolver == nil { + return webapi.NewClient() + } + + if endpoint := r.EndpointResolver(ctx); endpoint != "" { + return webapi.NewClient(endpoint) + } + + return webapi.NewClient() } // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=replicationpairs,verbs=get;list;watch;create;update;patch;delete @@ -62,7 +88,7 @@ func (r *ReplicationPairReconciler) Reconcile(ctx context.Context, req ctrl.Requ return ctrl.Result{}, client.IgnoreNotFound(err) } - apiClient := webapi.NewClient() + apiClient := r.apiClient(ctx) if !pair.DeletionTimestamp.IsZero() { return r.reconcileDelete(ctx, &pair, apiClient) diff --git a/operator/internal/controller/replicationpair_controller_unit_test.go b/operator/internal/controller/replicationpair_controller_unit_test.go index 2ec8d1fde..e8ecf1e42 100644 --- a/operator/internal/controller/replicationpair_controller_unit_test.go +++ b/operator/internal/controller/replicationpair_controller_unit_test.go @@ -159,6 +159,56 @@ func TestSitePair_CreatesBackendTarget(t *testing.T) { } } +// ---------- EndpointResolver overrides the default/env-var endpoint ---------- + +// TestSitePair_UsesEndpointResolver proves the reconciler reaches the control +// plane through EndpointResolver, the same mechanism StoragePoolReconciler +// already uses (controlplane.NewEndpointResolver, wired from +// ControlPlane.status.endpoint) -- required for a managed cluster whose own +// operator has no reachable simplyblock-webappapi Service, only the hub's +// externally-published one. SIMPLYBLOCK_WEBAPI_BASE_URL is deliberately left +// unset here so a pass proves the resolver path, not the pre-existing +// env-var fallback TestSitePair_CreatesBackendTarget already covers. +func TestSitePair_UsesEndpointResolver(t *testing.T) { + cluster1 := testCluster("default", "cluster1", "src-uuid") + cluster2 := testCluster("default", "cluster2", "tgt-uuid") + pair := newSitePair() + + srv := newAPIServer(t, func(w http.ResponseWriter, req *http.Request) { + if req.Method == http.MethodGet { + w.WriteHeader(http.StatusOK) + _, _ = w.Write([]byte(`[]`)) + return + } + if req.Method == http.MethodPost { + w.WriteHeader(http.StatusCreated) + resp, _ := json.Marshal(map[string]string{"id": "tgt-backend-uuid"}) + _, _ = w.Write(resp) + return + } + w.WriteHeader(http.StatusMethodNotAllowed) + }) + + r, cl := newSitePairReconciler(t, cluster1, cluster2, pair) + r.EndpointResolver = func(context.Context) string { return srv.URL } + + res, err := r.Reconcile(context.Background(), sitePairRequest("pair1")) + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + if res.RequeueAfter != replPairSyncInterval { + t.Errorf("RequeueAfter = %v, want %v", res.RequeueAfter, replPairSyncInterval) + } + + got := getSitePair(t, cl) + if !got.Status.Ready { + t.Errorf("pair.Status.Ready = false, want true (EndpointResolver should have been used)") + } + if got.Status.BackendTargetID != "tgt-backend-uuid" { + t.Errorf("BackendTargetID = %q, want tgt-backend-uuid", got.Status.BackendTargetID) + } +} + // ---------- backend target already exists → reuses it ---------- func TestSitePair_ReuseExistingTarget(t *testing.T) { diff --git a/operator/internal/controller/replicationpolicy_controller.go b/operator/internal/controller/replicationpolicy_controller.go index d355b8f56..7faf3d70d 100644 --- a/operator/internal/controller/replicationpolicy_controller.go +++ b/operator/internal/controller/replicationpolicy_controller.go @@ -35,6 +35,7 @@ import ( "sigs.k8s.io/controller-runtime/pkg/reconcile" simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" + "github.com/simplyblock/simplyblock-operator/internal/controllers/controlplane" "github.com/simplyblock/simplyblock-operator/internal/utils" "github.com/simplyblock/simplyblock-operator/internal/webapi" ) @@ -63,6 +64,27 @@ type idResponse struct { type ReplicationPolicyReconciler struct { client.Client Scheme *runtime.Scheme + + // EndpointResolver answers where the control plane currently is -- see + // ReplicationPairReconciler's identical field (replicationpair_controller.go) + // for why this exists: without it, a ReplicationPolicy reconciled on a + // ControlPlane.spec.source.managed member cluster can never reach its + // control plane. + EndpointResolver controlplane.EndpointResolver +} + +// apiClient resolves the control-plane client for one reconcile call, through +// EndpointResolver when set. Mirrors ReplicationPairReconciler.apiClient. +func (r *ReplicationPolicyReconciler) apiClient(ctx context.Context) *webapi.Client { + if r.EndpointResolver == nil { + return webapi.NewClient() + } + + if endpoint := r.EndpointResolver(ctx); endpoint != "" { + return webapi.NewClient(endpoint) + } + + return webapi.NewClient() } // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=replicationpolicies,verbs=get;list;watch;create;update;patch;delete @@ -101,7 +123,7 @@ func (r *ReplicationPolicyReconciler) Reconcile(ctx context.Context, req ctrl.Re return ctrl.Result{RequeueAfter: 10 * time.Second}, nil } - apiClient := webapi.NewClient() + apiClient := r.apiClient(ctx) if !policy.DeletionTimestamp.IsZero() { return r.reconcileDelete(ctx, &policy, apiClient, clusterUUID) diff --git a/operator/internal/controller/replicationpolicy_controller_unit_test.go b/operator/internal/controller/replicationpolicy_controller_unit_test.go index 41e07dc52..324a17eec 100644 --- a/operator/internal/controller/replicationpolicy_controller_unit_test.go +++ b/operator/internal/controller/replicationpolicy_controller_unit_test.go @@ -183,6 +183,50 @@ func TestPolicy_CreatesBackendPolicy_WhenAbsent(t *testing.T) { } } +// ---------- EndpointResolver overrides the default/env-var endpoint ---------- + +// TestPolicy_UsesEndpointResolver mirrors +// TestSitePair_UsesEndpointResolver (replicationpair_controller_unit_test.go): +// same bug, same fix, same reason -- a ReplicationPolicy reconciled on a +// ControlPlane.spec.source.managed member cluster needs the hub's externally +// published endpoint, not the in-cluster Service this reconciler otherwise +// defaults to. SIMPLYBLOCK_WEBAPI_BASE_URL is deliberately left unset so a +// pass proves the resolver path, not the env-var fallback +// TestPolicy_CreatesBackendPolicy_WhenAbsent already covers. +func TestPolicy_UsesEndpointResolver(t *testing.T) { + pair := readyPairForPolicy() + policy := &simplyblockv1alpha1.ReplicationPolicy{ + ObjectMeta: metav1.ObjectMeta{ + Name: "pol", Namespace: "default", + Finalizers: []string{utils.FinalizerReplicationPolicy}, + }, + Spec: simplyblockv1alpha1.ReplicationPolicySpec{PairRef: "pair1"}, + } + r, cl := newPolicyReconciler(t, pair, policy) + + srv := newAPIServer(t, func(w http.ResponseWriter, req *http.Request) { + switch { + case req.Method == http.MethodGet && req.URL.Path == apiPathReplicationPolicies: + writeJSON(w, []interface{}{}) + case req.Method == http.MethodPost && req.URL.Path == apiPathReplicationPolicies: + writeJSON(w, map[string]string{"id": "pol-backend-uuid"}) + default: + w.WriteHeader(http.StatusOK) + } + }) + r.EndpointResolver = func(context.Context) string { return srv.URL } + + _, err := r.Reconcile(context.Background(), policyRequest("pol")) + if err != nil { + t.Fatalf("unexpected error: %v", err) + } + + got := getPolicy(t, cl) + if got.Status.BackendPolicyID != "pol-backend-uuid" { + t.Errorf("BackendPolicyID = %q, want pol-backend-uuid (EndpointResolver should have been used)", got.Status.BackendPolicyID) + } +} + // ---------- reuse existing backend policy ---------- func TestPolicy_ReusesExistingBackendPolicy(t *testing.T) { From 60385759bb41d77fdf371c3ee6a03f06d53d671b Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Wed, 23 Sep 2026 13:10:12 +0100 Subject: [PATCH 141/206] Authenticate ReplicationPair/ReplicationPolicy as their own cluster, not this operator's identity --- .../controller/replicationpair_controller.go | 26 +++++++++++ .../replicationpair_controller_unit_test.go | 44 ++++++++++++++++++ .../replicationpolicy_controller.go | 22 +++++++++ .../replicationpolicy_controller_unit_test.go | 45 +++++++++++++++++++ 4 files changed, 137 insertions(+) diff --git a/operator/internal/controller/replicationpair_controller.go b/operator/internal/controller/replicationpair_controller.go index 51febe229..92e2a8d25 100644 --- a/operator/internal/controller/replicationpair_controller.go +++ b/operator/internal/controller/replicationpair_controller.go @@ -23,6 +23,7 @@ import ( "net/http" "time" + corev1 "k8s.io/api/core/v1" "k8s.io/apimachinery/pkg/runtime" ctrl "sigs.k8s.io/controller-runtime" "sigs.k8s.io/controller-runtime/pkg/client" @@ -74,6 +75,27 @@ func (r *ReplicationPairReconciler) apiClient(ctx context.Context) *webapi.Clien return webapi.NewClient() } +// clusterSecret reads the credential StorageClusterReconciler.persist wrote +// for the named local StorageCluster, so a call authenticates as that +// cluster instead of as this operator's own Kubernetes identity -- the only +// way to reach a control plane a different Kubernetes cluster runs, since a +// TokenReview can never cross that boundary. Mirrors +// internal/controllers/pool's identically named method (and +// internal/controllers/cluster's, internal/controllers/node's). +func (r *ReplicationPairReconciler) clusterSecret( + ctx context.Context, namespace, clusterName string, +) (string, error) { + var secret corev1.Secret + key := client.ObjectKey{ + Name: fmt.Sprintf("simplyblock-cluster-%s", clusterName), + Namespace: namespace, + } + if err := r.Get(ctx, key, &secret); err != nil { + return "", err + } + return string(secret.Data["secret"]), nil +} + // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=replicationpairs,verbs=get;list;watch;create;update;patch;delete // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=replicationpairs/status,verbs=get;update;patch // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=replicationpairs/finalizers,verbs=update @@ -88,6 +110,10 @@ func (r *ReplicationPairReconciler) Reconcile(ctx context.Context, req ctrl.Requ return ctrl.Result{}, client.IgnoreNotFound(err) } + if secret, err := r.clusterSecret(ctx, pair.Namespace, pair.Spec.SourceCluster); err == nil && secret != "" { + ctx = webapi.WithBearerToken(ctx, secret) + } + apiClient := r.apiClient(ctx) if !pair.DeletionTimestamp.IsZero() { diff --git a/operator/internal/controller/replicationpair_controller_unit_test.go b/operator/internal/controller/replicationpair_controller_unit_test.go index e8ecf1e42..68fdf7cd2 100644 --- a/operator/internal/controller/replicationpair_controller_unit_test.go +++ b/operator/internal/controller/replicationpair_controller_unit_test.go @@ -209,6 +209,50 @@ func TestSitePair_UsesEndpointResolver(t *testing.T) { } } +// TestSitePair_AuthenticatesAsSourceClusterOnceSecretIsKnown mirrors +// internal/controllers/pool's TestAPoolAuthenticatesAsItsClusterOnceTheSecretIsKnown +// -- same fix, same reason: cluster B's operator authenticating as its own +// Kubernetes identity can never pass a TokenReview on the hub's cluster (the +// whole reason per-cluster secrets exist at all), confirmed live this session +// as the very next error once EndpointResolver alone let the request reach +// the right endpoint ("status 401: Invalid token"). +func TestSitePair_AuthenticatesAsSourceClusterOnceSecretIsKnown(t *testing.T) { + cluster1 := testCluster("default", "cluster1", "src-uuid") + cluster2 := testCluster("default", "cluster2", "tgt-uuid") + pair := newSitePair() + secret := &corev1.Secret{ + ObjectMeta: metav1.ObjectMeta{ + Name: "simplyblock-cluster-cluster1", + Namespace: "default", + }, + Data: map[string][]byte{"secret": []byte("cluster1-own-secret")}, + } + + var gotAuth string + srv := newAPIServer(t, func(w http.ResponseWriter, req *http.Request) { + gotAuth = req.Header.Get("Authorization") + if req.Method == http.MethodGet { + w.WriteHeader(http.StatusOK) + _, _ = w.Write([]byte(`[]`)) + return + } + w.WriteHeader(http.StatusCreated) + resp, _ := json.Marshal(map[string]string{"id": "tgt-backend-uuid"}) + _, _ = w.Write(resp) + }) + + r, _ := newSitePairReconciler(t, cluster1, cluster2, pair, secret) + r.EndpointResolver = func(context.Context) string { return srv.URL } + + if _, err := r.Reconcile(context.Background(), sitePairRequest("pair1")); err != nil { + t.Fatalf("unexpected error: %v", err) + } + + if gotAuth != "Bearer cluster1-own-secret" { + t.Errorf("authorization = %q, want the source cluster's own secret", gotAuth) + } +} + // ---------- backend target already exists → reuses it ---------- func TestSitePair_ReuseExistingTarget(t *testing.T) { diff --git a/operator/internal/controller/replicationpolicy_controller.go b/operator/internal/controller/replicationpolicy_controller.go index 7faf3d70d..e5ae43fac 100644 --- a/operator/internal/controller/replicationpolicy_controller.go +++ b/operator/internal/controller/replicationpolicy_controller.go @@ -24,6 +24,7 @@ import ( "net/http" "time" + corev1 "k8s.io/api/core/v1" apierrors "k8s.io/apimachinery/pkg/api/errors" "k8s.io/apimachinery/pkg/runtime" "k8s.io/apimachinery/pkg/types" @@ -87,6 +88,23 @@ func (r *ReplicationPolicyReconciler) apiClient(ctx context.Context) *webapi.Cli return webapi.NewClient() } +// clusterSecret reads the credential StorageClusterReconciler.persist wrote +// for the named local StorageCluster. Mirrors +// ReplicationPairReconciler.clusterSecret. +func (r *ReplicationPolicyReconciler) clusterSecret( + ctx context.Context, namespace, clusterName string, +) (string, error) { + var secret corev1.Secret + key := client.ObjectKey{ + Name: fmt.Sprintf("simplyblock-cluster-%s", clusterName), + Namespace: namespace, + } + if err := r.Get(ctx, key, &secret); err != nil { + return "", err + } + return string(secret.Data["secret"]), nil +} + // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=replicationpolicies,verbs=get;list;watch;create;update;patch;delete // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=replicationpolicies/status,verbs=get;update;patch // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=replicationpolicies/finalizers,verbs=update @@ -123,6 +141,10 @@ func (r *ReplicationPolicyReconciler) Reconcile(ctx context.Context, req ctrl.Re return ctrl.Result{RequeueAfter: 10 * time.Second}, nil } + if secret, err := r.clusterSecret(ctx, policy.Namespace, pair.Spec.SourceCluster); err == nil && secret != "" { + ctx = webapi.WithBearerToken(ctx, secret) + } + apiClient := r.apiClient(ctx) if !policy.DeletionTimestamp.IsZero() { diff --git a/operator/internal/controller/replicationpolicy_controller_unit_test.go b/operator/internal/controller/replicationpolicy_controller_unit_test.go index 324a17eec..0ddd9a8b2 100644 --- a/operator/internal/controller/replicationpolicy_controller_unit_test.go +++ b/operator/internal/controller/replicationpolicy_controller_unit_test.go @@ -227,6 +227,51 @@ func TestPolicy_UsesEndpointResolver(t *testing.T) { } } +// TestPolicy_AuthenticatesAsSourceClusterOnceSecretIsKnown mirrors +// replicationpair_controller_unit_test.go's identically named test -- same +// bug, same fix, same reason: authenticating as this operator's own +// Kubernetes identity can never pass a TokenReview on the hub's cluster. +func TestPolicy_AuthenticatesAsSourceClusterOnceSecretIsKnown(t *testing.T) { + pair := readyPairForPolicy() + policy := &simplyblockv1alpha1.ReplicationPolicy{ + ObjectMeta: metav1.ObjectMeta{ + Name: "pol", Namespace: "default", + Finalizers: []string{utils.FinalizerReplicationPolicy}, + }, + Spec: simplyblockv1alpha1.ReplicationPolicySpec{PairRef: "pair1"}, + } + secret := &corev1.Secret{ + ObjectMeta: metav1.ObjectMeta{ + Name: "simplyblock-cluster-" + testClusterName, + Namespace: "default", + }, + Data: map[string][]byte{"secret": []byte("cluster-own-secret")}, + } + r, _ := newPolicyReconciler(t, pair, policy, secret) + + var gotAuth string + srv := newAPIServer(t, func(w http.ResponseWriter, req *http.Request) { + gotAuth = req.Header.Get("Authorization") + switch { + case req.Method == http.MethodGet && req.URL.Path == apiPathReplicationPolicies: + writeJSON(w, []interface{}{}) + case req.Method == http.MethodPost && req.URL.Path == apiPathReplicationPolicies: + writeJSON(w, map[string]string{"id": "pol-backend-uuid"}) + default: + w.WriteHeader(http.StatusOK) + } + }) + r.EndpointResolver = func(context.Context) string { return srv.URL } + + if _, err := r.Reconcile(context.Background(), policyRequest("pol")); err != nil { + t.Fatalf("unexpected error: %v", err) + } + + if gotAuth != "Bearer cluster-own-secret" { + t.Errorf("authorization = %q, want the source cluster's own secret", gotAuth) + } +} + // ---------- reuse existing backend policy ---------- func TestPolicy_ReusesExistingBackendPolicy(t *testing.T) { From 7aa02ab95ef61d83428702f2a0a13a499c5870c7 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Wed, 23 Sep 2026 17:09:12 +0100 Subject: [PATCH 142/206] Resolve foreign source-side volume handles to the local replica before replication RPCs --- atlas-lib/controlplane/replication.go | 58 ++++++++++++ atlas-lib/controlplane/replication_test.go | 86 ++++++++++++++++++ .../csi/controller/mock_controlplane_test.go | 37 +++++++- .../internal/csi/controller/replication.go | 48 ++++++++++ .../controller/replication_lifecycle_test.go | 88 +++++++++++++++++++ .../csi/controller/replication_test.go | 43 ++++++++- .../node/storagenode_controller.go | 14 --- .../node/storagenodeops_controller.go | 3 - .../internal/controllers/node/workload.go | 13 --- .../controllers/node/workload_controller.go | 5 -- 10 files changed, 353 insertions(+), 42 deletions(-) diff --git a/atlas-lib/controlplane/replication.go b/atlas-lib/controlplane/replication.go index 4b88df585..aa55fd16d 100644 --- a/atlas-lib/controlplane/replication.go +++ b/atlas-lib/controlplane/replication.go @@ -7,10 +7,19 @@ import ( "strings" "time" + openapi_types "github.com/oapi-codegen/runtime/types" + "github.com/simplyblock/atlas/internal/cpapi" "github.com/simplyblock/atlas/lvol" ) +func uuidPtrString(u *openapi_types.UUID) string { + if u == nil { + return "" + } + return u.String() +} + // ReplicationStatus is the typed steady-state replication status of one // volume, for the volume's whole replicated life -- unlike a cutover-record // relationship read, this is never a 404 for a volume that exists. @@ -223,3 +232,52 @@ func (c *Client) ResyncVolume(ctx context.Context, h lvol.VolumeHandle, sourceCl } return nil } + +// Relationship is the replication pairing h belongs to: which volume +// replicates to which, and which side h itself names (IsSource). TargetClusterID/ +// TargetPoolID/TargetLvolID always name the same, fixed target (replica) side +// of the pairing regardless of whether h names the source or the target -- +// querying by either volume's own id returns the identical target_* answer. +type Relationship struct { + IsSource bool + + SourceClusterID string + SourceLvolID string + + TargetClusterID string + TargetPoolID string + TargetLvolID string +} + +// GetVolumeReplicationRelationship resolves h to its replication pairing. +// This is what a caller handed a volume identity inherited from the OTHER +// side of a pairing (e.g. a destination PVC whose PV was restored carrying +// the source's own volumeHandle) uses to find the volume it should actually +// operate on locally: TargetClusterID/TargetPoolID/TargetLvolID name that +// volume regardless of which side h itself named. Returns an error +// unwrapping to errs.ErrNotFound when h has no replication relationship at +// all yet (e.g. a volume never enabled for replication) -- callers treat that +// as "use h unchanged," not a failure. +func (c *Client) GetVolumeReplicationRelationship(ctx context.Context, h lvol.VolumeHandle) (Relationship, error) { + cluster, pool, volume, err := h.Split() + if err != nil { + return Relationship{}, err + } + resp, err := c.api.ClustersStoragePoolsVolumesReplicationDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationGetWithResponse( + ctx, cluster, pool, volume) + if err != nil { + return Relationship{}, fmt.Errorf("replication relationship of volume %s: %w", h, err) + } + d, err := payload("replication relationship of volume "+string(h), resp.JSON200, resp.StatusCode(), resp.Body) + if err != nil { + return Relationship{}, err + } + return Relationship{ + IsSource: d.IsSource, + SourceClusterID: uuidPtrString(d.SourceClusterId), + SourceLvolID: uuidPtrString(d.SourceLvolId), + TargetClusterID: uuidPtrString(d.TargetClusterId), + TargetPoolID: uuidPtrString(d.TargetPoolId), + TargetLvolID: uuidPtrString(d.TargetLvolId), + }, nil +} diff --git a/atlas-lib/controlplane/replication_test.go b/atlas-lib/controlplane/replication_test.go index 9567f0fee..44d500b37 100644 --- a/atlas-lib/controlplane/replication_test.go +++ b/atlas-lib/controlplane/replication_test.go @@ -300,3 +300,89 @@ func TestClientResyncVolumeWithoutSourceCluster(t *testing.T) { t.Errorf("request body = %q, want no source_cluster_id when none is given", gotBody) } } + +const ( + testTargetCluster = "44444444-4444-4444-4444-444444444444" + testTargetPool = "55555555-5555-5555-5555-555555555555" + testTargetVolume = "66666666-6666-6666-6666-666666666667" +) + +func TestClientGetVolumeReplicationRelationship(t *testing.T) { + c := newTestClient(t, func(w http.ResponseWriter, r *http.Request) { + if !strings.HasSuffix(r.URL.Path, "/replication/") && !strings.HasSuffix(r.URL.Path, "/replication") { + t.Errorf("unexpected path %q", r.URL.Path) + } + w.Header().Set("Content-Type", "application/json") + _, _ = w.Write([]byte(`{ + "replication_id": "` + testCluster + `", + "direction": "to_target", + "mode": "migration", + "state": "replicating", + "is_source": true, + "source_cluster_id": "` + testCluster + `", + "source_lvol_id": "` + testVolume + `", + "target_cluster_id": "` + testTargetCluster + `", + "target_pool_id": "` + testTargetPool + `", + "target_lvol_id": "` + testTargetVolume + `", + "target_nqn": "nqn.test", + "target_ns_id": 1 + }`)) + }) + + rel, err := c.GetVolumeReplicationRelationship(context.Background(), testHandle) + if err != nil { + t.Fatal(err) + } + if !rel.IsSource { + t.Error("IsSource = false, want true (queried by the source volume)") + } + if rel.TargetClusterID != testTargetCluster || rel.TargetPoolID != testTargetPool || rel.TargetLvolID != testTargetVolume { + t.Errorf("target = %s/%s/%s, want %s/%s/%s", + rel.TargetClusterID, rel.TargetPoolID, rel.TargetLvolID, + testTargetCluster, testTargetPool, testTargetVolume) + } +} + +// Queried by the TARGET volume's own id, the relationship still reports the +// SAME fixed target_* fields -- this is what lets a caller always resolve to +// target_* regardless of which side of the pairing it was handed, without +// having to branch on IsSource first. +func TestClientGetVolumeReplicationRelationshipQueriedByTarget(t *testing.T) { + c := newTestClient(t, func(w http.ResponseWriter, r *http.Request) { + w.Header().Set("Content-Type", "application/json") + _, _ = w.Write([]byte(`{ + "replication_id": "` + testCluster + `", + "direction": "to_target", + "mode": "migration", + "state": "replicating", + "is_source": false, + "source_cluster_id": "` + testCluster + `", + "source_lvol_id": "` + testVolume + `", + "target_cluster_id": "` + testTargetCluster + `", + "target_pool_id": "` + testTargetPool + `", + "target_lvol_id": "` + testTargetVolume + `", + "target_nqn": "nqn.test", + "target_ns_id": 1 + }`)) + }) + + rel, err := c.GetVolumeReplicationRelationship(context.Background(), testHandle) + if err != nil { + t.Fatal(err) + } + if rel.IsSource { + t.Error("IsSource = true, want false (queried by the target volume)") + } + if rel.TargetLvolID != testTargetVolume { + t.Errorf("TargetLvolID = %s, want %s (unchanged regardless of which side was queried)", rel.TargetLvolID, testTargetVolume) + } +} + +func TestClientGetVolumeReplicationRelationshipNotFound(t *testing.T) { + c := newTestClient(t, func(w http.ResponseWriter, r *http.Request) { + w.WriteHeader(http.StatusNotFound) + }) + if _, err := c.GetVolumeReplicationRelationship(context.Background(), testHandle); !errors.Is(err, errs.ErrNotFound) { + t.Errorf("err = %v, want ErrNotFound for a volume with no replication relationship yet", err) + } +} diff --git a/csi-driver/internal/csi/controller/mock_controlplane_test.go b/csi-driver/internal/csi/controller/mock_controlplane_test.go index 75942a6b5..9dfcc2adc 100644 --- a/csi-driver/internal/csi/controller/mock_controlplane_test.go +++ b/csi-driver/internal/csi/controller/mock_controlplane_test.go @@ -117,6 +117,15 @@ type mockSBCLI struct { // only has to prove the driver maps whatever shape the endpoint returns. replicationStatus map[string]map[string]any + // replicationRelationship, keyed by volume id (either side of the + // pairing), is the raw JSON body GET .../replication/ serves for that + // volume. Absent means no relationship exists yet (404, matching + // sbcli's get_relationship returning None for a volume never enabled for + // replication) -- this look-up deliberately does not require the id to + // exist in m.volumes, matching the real backend, whose relationship + // records outlive a deleted source volume. + replicationRelationship map[string]map[string]any + // replicationPUTStatus, when set, makes every PUT carrying // replication_policy_id respond with this HTTP status instead of the // normal idempotent update, modeling a backend refusal (e.g. a policy @@ -131,6 +140,10 @@ type mockSBCLI struct { // call, so a test can assert the driver actually sent planned=true/false // rather than only checking the resulting gRPC code. lastFailoverQuery string + // lastFailoverVolumeID captures which volume's path the last failover + // call landed on, so a test can assert a relationship-resolved call + // reached the TARGET volume rather than the one it was originally given. + lastFailoverVolumeID string // demoteStatus, when set, is the HTTP status POST .../demote answers with // instead of its default success (204). 202 models "still converging." @@ -145,10 +158,11 @@ type mockSBCLI struct { func newMockSBCLI() *mockSBCLI { m := &mockSBCLI{ - volumes: make(map[string]*mockVolume), - snapshots: make(map[string]*mockSnapshot), - groups: make(map[string]*mockGroup), - replicationStatus: make(map[string]map[string]any), + volumes: make(map[string]*mockVolume), + snapshots: make(map[string]*mockSnapshot), + groups: make(map[string]*mockGroup), + replicationStatus: make(map[string]map[string]any), + replicationRelationship: make(map[string]map[string]any), } mux := http.NewServeMux() @@ -171,6 +185,10 @@ func newMockSBCLI() *mockSBCLI { "GET /api/v2/clusters/{clusterID}/storage-pools/{poolID}/volumes/{volumeID}/replication/status", m.locked(m.handleReplicationStatus), ) + mux.HandleFunc( + "GET /api/v2/clusters/{clusterID}/storage-pools/{poolID}/volumes/{volumeID}/replication/", + m.locked(m.handleReplicationRelationship), + ) mux.HandleFunc( "POST /api/v2/clusters/{clusterID}/storage-pools/{poolID}/volumes/{volumeID}/replication/failover", m.locked(m.handleFailover), @@ -368,12 +386,23 @@ func (m *mockSBCLI) handleReplicationStatus(w http.ResponseWriter, r *http.Reque writeJSON(w, http.StatusOK, body) } +func (m *mockSBCLI) handleReplicationRelationship(w http.ResponseWriter, r *http.Request) { + volumeID := r.PathValue("volumeID") + body, ok := m.replicationRelationship[volumeID] + if !ok { + writeJSON(w, http.StatusNotFound, map[string]string{"detail": "Volume has no replication relationship"}) + return + } + writeJSON(w, http.StatusOK, body) +} + func (m *mockSBCLI) handleFailover(w http.ResponseWriter, r *http.Request) { volumeID := r.PathValue("volumeID") if m.lookupVolume(w, volumeID) == nil { return } m.lastFailoverQuery = r.URL.RawQuery + m.lastFailoverVolumeID = volumeID if m.failoverStatus != 0 { writeJSON(w, m.failoverStatus, map[string]string{"detail": "injected status"}) return diff --git a/csi-driver/internal/csi/controller/replication.go b/csi-driver/internal/csi/controller/replication.go index d8f74930a..8fa1a0f8d 100644 --- a/csi-driver/internal/csi/controller/replication.go +++ b/csi-driver/internal/csi/controller/replication.go @@ -8,12 +8,16 @@ package controller import ( "context" + "errors" "github.com/csi-addons/spec/lib/go/replication" "google.golang.org/grpc/codes" "google.golang.org/grpc/status" "google.golang.org/protobuf/types/known/timestamppb" + atlascp "github.com/simplyblock/atlas/controlplane" + "github.com/simplyblock/atlas/errs" + "github.com/simplyblock/atlas/lvol" "github.com/simplyblock/csi-driver/internal/clusters" csicommon "github.com/simplyblock/csi-driver/internal/csi/common" ) @@ -51,6 +55,42 @@ func volumeIDFrom(req volumeIDCarrier) string { return req.GetVolumeId() } +// resolveToLocalReplica rewrites h to the volume that actually replicates +// data on this side of an existing pairing, when h names the OTHER (foreign) +// side of it instead. Ramen's S3-restore recreates a destination PV carrying +// the ORIGINAL source's own volumeHandle verbatim (confirmed live +// 2026-09-23, relocate M-02), and every Replication RPC parses its target +// straight from the handle it's given -- without this resolution, "promote" +// or "enable replication" would operate on the foreign, original volume +// instead of the local replica that has actually been receiving replicated +// data. The returned handle and client change together, since the target +// side may live on a different cluster with its own secret.json entry. +// +// Returns h and client unchanged when h has no replication relationship yet +// (errs.ErrNotFound -- the ordinary case for a volume never enabled for +// replication, e.g. M-01's first-ever protect) or when h already names the +// target side. +func resolveToLocalReplica( + ctx context.Context, h *lvol.Handle, client *atlascp.Client, +) (*lvol.Handle, *atlascp.Client, error) { + rel, err := client.GetVolumeReplicationRelationship(ctx, h.Handle()) + if err != nil { + if errors.Is(err, errs.ErrNotFound) { + return h, client, nil + } + return nil, nil, err + } + if !rel.IsSource { + return h, client, nil + } + target := &lvol.Handle{ClusterID: rel.TargetClusterID, PoolRef: rel.TargetPoolID, VolumeID: rel.TargetLvolID} + targetClient, err := clusters.ReplicationClient(ctx, target.ClusterID) + if err != nil { + return nil, nil, err + } + return target, targetClient, nil +} + // EnableVolumeReplication attaches the volume to the policy named by the // VolumeReplicationClass. Attaching to the policy the volume already follows // is success (the backend's own idempotency, P0-2). @@ -81,6 +121,10 @@ func (cs *Server) EnableVolumeReplication( if err != nil { return nil, status.Error(codes.Unavailable, err.Error()) } + h, client, err = resolveToLocalReplica(ctx, h, client) + if err != nil { + return nil, status.Error(codes.Unavailable, err.Error()) + } if err := client.EnableVolumeReplication(ctx, h.Handle(), policyID); err != nil { return nil, classifyEnableVolumeReplicationError(err) } @@ -159,6 +203,10 @@ func (cs *Server) PromoteVolume( if err != nil { return nil, status.Error(codes.Unavailable, err.Error()) } + h, client, err = resolveToLocalReplica(ctx, h, client) + if err != nil { + return nil, status.Error(codes.Unavailable, err.Error()) + } if err := client.PromoteVolume(ctx, h.Handle(), req.GetForce()); err != nil { return nil, classifyPromoteVolumeError(err) } diff --git a/csi-driver/internal/csi/controller/replication_lifecycle_test.go b/csi-driver/internal/csi/controller/replication_lifecycle_test.go index f52b01504..dc764f3dd 100644 --- a/csi-driver/internal/csi/controller/replication_lifecycle_test.go +++ b/csi-driver/internal/csi/controller/replication_lifecycle_test.go @@ -27,6 +27,94 @@ func TestPromoteVolumeForced(t *testing.T) { } } +// A Ramen-restored destination PV inherits the ORIGINAL source's own +// volumeHandle verbatim (confirmed live 2026-09-23, relocate M-02): Ramen's +// S3-restore recreates the exact PV/PVC object it archived at protect time, +// including the source cluster+lvol identity, on a cluster that never +// provisioned that volume at all. Every Replication RPC parses its target +// straight from the given handle, so without resolving through the backend's +// own source->target relationship first, "promote" would be asking the +// ORIGINAL, foreign volume to fail over -- not the local replica that has +// actually been receiving replicated data. PromoteVolume must resolve a +// handle whose relationship says IsSource and redirect to TargetLvolId +// (on TargetClusterId/TargetPoolId) before calling failover. +func TestPromoteVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + mock.volumes[testReplTargetVolumeID] = &mockVolume{UUID: testReplTargetVolumeID, Name: "repl-vol-target", Size: 1 << 30} + mock.replicationRelationship[testReplVolumeID] = map[string]any{ + "replication_id": "aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee", + "direction": "to_target", + "mode": "migration", + "state": "replicating", + "is_source": true, + "source_cluster_id": sanityClusterID, "source_lvol_id": testReplVolumeID, + "target_cluster_id": sanityClusterID, "target_pool_id": sanityPoolUUID, "target_lvol_id": testReplTargetVolumeID, + "target_nqn": "nqn.test", "target_ns_id": 1, + } + + _, err := cs.PromoteVolume(context.Background(), &replication.PromoteVolumeRequest{ + VolumeId: testReplVolID, Force: true, + }) + if err != nil { + t.Fatal(err) + } + if mock.lastFailoverVolumeID != testReplTargetVolumeID { + t.Errorf("failover landed on volume %q, want the resolved target %q", mock.lastFailoverVolumeID, testReplTargetVolumeID) + } +} + +// A volume already naming the target side of its own relationship (IsSource +// false) is promoted directly, unchanged -- resolving again would be a +// harmless no-op, but this proves it takes that path rather than one that +// happens to work only by coincidence. +func TestPromoteVolumeAlreadyNamingTheTargetIsUnchanged(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + mock.replicationRelationship[testReplVolumeID] = map[string]any{ + "replication_id": "aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee", + "direction": "to_target", + "mode": "migration", + "state": "replicating", + "is_source": false, + "source_cluster_id": sanityClusterID, "source_lvol_id": "99999999-9999-9999-9999-999999999998", + "target_cluster_id": sanityClusterID, "target_pool_id": sanityPoolUUID, "target_lvol_id": testReplVolumeID, + "target_nqn": "nqn.test", "target_ns_id": 1, + } + + _, err := cs.PromoteVolume(context.Background(), &replication.PromoteVolumeRequest{ + VolumeId: testReplVolID, Force: true, + }) + if err != nil { + t.Fatal(err) + } + if mock.lastFailoverVolumeID != testReplVolumeID { + t.Errorf("failover landed on volume %q, want %q unchanged", mock.lastFailoverVolumeID, testReplVolumeID) + } +} + +// The ordinary case, and by far the most common: a volume never enabled for +// replication (M-01's own first-ever protect) has no relationship at all yet. +// This must promote the given handle directly rather than fail the whole +// call over a 404 that just means "nothing to resolve." +func TestPromoteVolumeWithNoRelationshipYetUsesTheGivenHandle(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + + _, err := cs.PromoteVolume(context.Background(), &replication.PromoteVolumeRequest{ + VolumeId: testReplVolID, Force: true, + }) + if err != nil { + t.Fatal(err) + } + if mock.lastFailoverVolumeID != testReplVolumeID { + t.Errorf("failover landed on volume %q, want %q", mock.lastFailoverVolumeID, testReplVolumeID) + } +} + func TestPromoteVolumePlannedSendsThePlannedFlag(t *testing.T) { mock := newMockSBCLI() defer mock.Close() diff --git a/csi-driver/internal/csi/controller/replication_test.go b/csi-driver/internal/csi/controller/replication_test.go index 4bf110d32..89f2163ee 100644 --- a/csi-driver/internal/csi/controller/replication_test.go +++ b/csi-driver/internal/csi/controller/replication_test.go @@ -10,9 +10,10 @@ import ( ) const ( - testReplVolumeID = "88888888-8888-8888-8888-888888888888" - testReplVolID = sanityClusterID + ":" + sanityPoolUUID + ":" + testReplVolumeID - testReplPolicyID = "77777777-7777-7777-7777-777777777777" + testReplVolumeID = "88888888-8888-8888-8888-888888888888" + testReplVolID = sanityClusterID + ":" + sanityPoolUUID + ":" + testReplVolumeID + testReplPolicyID = "77777777-7777-7777-7777-777777777777" + testReplTargetVolumeID = "88888888-8888-8888-8888-888888888889" ) func newReplicationTestServer(t *testing.T, mock *mockSBCLI) *Server { @@ -91,6 +92,42 @@ func TestEnableVolumeReplicationUsesReplicationSourceWhenVolumeIdIsEmpty(t *test } } +// Same relationship-resolution requirement as PromoteVolume (see its own +// TestPromoteVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship): +// csi-addons calls EnableVolumeReplication as a promote precondition too, so +// a handle inherited from the foreign source side must resolve to the local +// target volume before the policy attach is attempted against it. +func TestEnableVolumeReplicationResolvesToTargetWhenGivenTheSourceSideOfARelationship(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + mock.volumes[testReplTargetVolumeID] = &mockVolume{UUID: testReplTargetVolumeID, Name: "repl-vol-target", Size: 1 << 30} + mock.replicationRelationship[testReplVolumeID] = map[string]any{ + "replication_id": "aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee", + "direction": "to_target", + "mode": "migration", + "state": "replicating", + "is_source": true, + "source_cluster_id": sanityClusterID, "source_lvol_id": testReplVolumeID, + "target_cluster_id": sanityClusterID, "target_pool_id": sanityPoolUUID, "target_lvol_id": testReplTargetVolumeID, + "target_nqn": "nqn.test", "target_ns_id": 1, + } + + _, err := cs.EnableVolumeReplication(context.Background(), &replication.EnableVolumeReplicationRequest{ + VolumeId: testReplVolID, + Parameters: map[string]string{replicationPolicyParam: testReplPolicyID}, + }) + if err != nil { + t.Fatal(err) + } + if got := mock.volumes[testReplTargetVolumeID].ReplicationPolicyID; got != testReplPolicyID { + t.Errorf("target volume's ReplicationPolicyID = %q, want %q", got, testReplPolicyID) + } + if got := mock.volumes[testReplVolumeID].ReplicationPolicyID; got != "" { + t.Errorf("source volume's ReplicationPolicyID = %q, want untouched", got) + } +} + func TestEnableVolumeReplicationMissingPolicyParam(t *testing.T) { mock := newMockSBCLI() defer mock.Close() diff --git a/operator/internal/controllers/node/storagenode_controller.go b/operator/internal/controllers/node/storagenode_controller.go index 21d2fe2df..b41cb0ad3 100644 --- a/operator/internal/controllers/node/storagenode_controller.go +++ b/operator/internal/controllers/node/storagenode_controller.go @@ -239,7 +239,6 @@ func (r *StorageNodeReconciler) Reconcile( if err != nil { return ctrl.Result{}, err } -<<<<<<< HEAD // Every control-plane call below authenticates as this node's cluster, // using its own recorded secret, rather than as this operator's own // Kubernetes identity -- the only way to reach a control plane a @@ -252,8 +251,6 @@ func (r *StorageNodeReconciler) Reconcile( secret, err := clusterSecretByName(ctx, r.Client, cluster.Namespace, cluster.Name) ctx = authenticatedContext(ctx, secret, err) } -======= ->>>>>>> main if !node.DeletionTimestamp.IsZero() { return r.teardown(ctx, &node) @@ -777,11 +774,7 @@ func (r *StorageNodeReconciler) postNode( node *simplyblockv1alpha2.StorageNode, cluster *simplyblockv1alpha2.StorageCluster, ) error { -<<<<<<< HEAD params := r.addParams(ctx, node, cluster) -======= - params := r.addParams(node, cluster) ->>>>>>> main if err := r.API.AddNode(ctx, cluster.Status.UUID, params); err != nil { return fmt.Errorf("add node %s on worker %s: %w", node.Name, node.Spec.WorkerNode, err) @@ -1433,10 +1426,7 @@ func (r *StorageNodeReconciler) upgradeAdoption( // addParams is what the node-add call carries. The node describes itself, so every // value but the subsystem cap comes from its own spec.config (§3.1). func (r *StorageNodeReconciler) addParams( -<<<<<<< HEAD ctx context.Context, -======= ->>>>>>> main node *simplyblockv1alpha2.StorageNode, cluster *simplyblockv1alpha2.StorageCluster, ) utils.StorageNodeSetAddParams { @@ -1447,11 +1437,7 @@ func (r *StorageNodeReconciler) addParams( } params := utils.StorageNodeSetAddParams{ -<<<<<<< HEAD NodeAddress: r.Workload.NodeAddress(ctx, node.Spec.WorkerNode, node.Namespace), -======= - NodeAddress: r.Workload.NodeAddress(node.Spec.WorkerNode, node.Namespace), ->>>>>>> main InterfaceName: workload.MgmtInterface, SPDKImage: config.SpdkImage, SPDKProxyImage: config.SpdkProxyImage, diff --git a/operator/internal/controllers/node/storagenodeops_controller.go b/operator/internal/controllers/node/storagenodeops_controller.go index a24afa7ca..a0c1a7381 100644 --- a/operator/internal/controllers/node/storagenodeops_controller.go +++ b/operator/internal/controllers/node/storagenodeops_controller.go @@ -256,7 +256,6 @@ func (r *StorageNodeOpsReconciler) Reconcile( func (r *StorageNodeOpsReconciler) advance( ctx context.Context, ops *simplyblockv1alpha2.StorageNodeOps, ) (ctrl.Result, error) { -<<<<<<< HEAD // Every control-plane call this step and everything downstream of it // makes authenticates as this operation's target node's cluster when its // secret is known, rather than as this operator's own Kubernetes @@ -266,8 +265,6 @@ func (r *StorageNodeOpsReconciler) advance( secret, secretErr := clusterSecretForNode(ctx, r.Client, ops.Namespace, ops.Spec.NodeRef) ctx = authenticatedContext(ctx, secret, secretErr) -======= ->>>>>>> main machine, err := graphs().FromSnapshot(ctx, action(ops.Spec.Action), statemachine.FromKube[step](ops.Status.Step)) if err != nil { diff --git a/operator/internal/controllers/node/workload.go b/operator/internal/controllers/node/workload.go index a4ac7daa0..bdab7dbf1 100644 --- a/operator/internal/controllers/node/workload.go +++ b/operator/internal/controllers/node/workload.go @@ -89,7 +89,6 @@ type Workload struct { ManagerNode string } -<<<<<<< HEAD // NodeAddress is what the control plane is given as node_address when a node // is added or restarted. // @@ -136,18 +135,6 @@ func (w *Workload) managedNodeAddress(ctx context.Context, worker, namespace str return "", false } -======= -// NodeAddress is the per-pod DNS name the control plane is given as node_address -// when a node is added or restarted. -// -// It is the precondition for both: a restart issued against a name that does not -// yet resolve fails name resolution inside the control plane, and the control -// plane's response to that is to reset the node to offline (§5.4). -func (w *Workload) NodeAddress(worker, namespace string) string { - return utils.StorageNodeSetAPIAddress(worker, namespace) -} - ->>>>>>> main // LabelWorker puts one worker into a cluster's storage plane and rewrites the // per-slot labels of every node on it. // diff --git a/operator/internal/controllers/node/workload_controller.go b/operator/internal/controllers/node/workload_controller.go index 22c232706..747c91f53 100644 --- a/operator/internal/controllers/node/workload_controller.go +++ b/operator/internal/controllers/node/workload_controller.go @@ -311,7 +311,6 @@ func (r *StorageNodeWorkloadReconciler) image( "spec.storageNodes.image is unset and ControlPlane %s cannot be read: %w", SingletonControlPlaneName, err) } -<<<<<<< HEAD if local := controlPlane.Spec.Source.Local; local != nil && local.Image != "" { return local.Image, nil } @@ -321,10 +320,6 @@ func (r *StorageNodeWorkloadReconciler) image( // should run. StorageNodeImage is the only source of a default left. if managed := controlPlane.Spec.Source.Managed; managed != nil && managed.StorageNodeImage != "" { return managed.StorageNodeImage, nil -======= - if managed := controlPlane.Spec.Source.Local; managed != nil && managed.Image != "" { - return managed.Image, nil ->>>>>>> main } return "", fmt.Errorf( "spec.storageNodes.image is unset and ControlPlane %s states no managed image", From a6bddd2aa4890588a573afc9bfd510a152575af1 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Wed, 23 Sep 2026 17:25:14 +0100 Subject: [PATCH 143/206] fixed linter issue --- .../controller/replication_lifecycle_test.go | 7 +++++-- .../csi/controller/replication_test.go | 4 +++- .../templates/numa-resource-plugin.yaml | 9 +++------ .../simplyblock-operator-webhook.yaml | 3 --- operator/dist/install.yaml | 19 +++++++++++++++++++ operator/test/e2e/rbac_test.go | 16 ++++++++-------- 6 files changed, 38 insertions(+), 20 deletions(-) diff --git a/csi-driver/internal/csi/controller/replication_lifecycle_test.go b/csi-driver/internal/csi/controller/replication_lifecycle_test.go index dc764f3dd..06c62cf81 100644 --- a/csi-driver/internal/csi/controller/replication_lifecycle_test.go +++ b/csi-driver/internal/csi/controller/replication_lifecycle_test.go @@ -42,7 +42,9 @@ func TestPromoteVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship(t *t mock := newMockSBCLI() defer mock.Close() cs := newReplicationTestServer(t, mock) - mock.volumes[testReplTargetVolumeID] = &mockVolume{UUID: testReplTargetVolumeID, Name: "repl-vol-target", Size: 1 << 30} + mock.volumes[testReplTargetVolumeID] = &mockVolume{ + UUID: testReplTargetVolumeID, Name: "repl-vol-target", Size: 1 << 30, + } mock.replicationRelationship[testReplVolumeID] = map[string]any{ "replication_id": "aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee", "direction": "to_target", @@ -61,7 +63,8 @@ func TestPromoteVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship(t *t t.Fatal(err) } if mock.lastFailoverVolumeID != testReplTargetVolumeID { - t.Errorf("failover landed on volume %q, want the resolved target %q", mock.lastFailoverVolumeID, testReplTargetVolumeID) + t.Errorf("failover landed on volume %q, want the resolved target %q", + mock.lastFailoverVolumeID, testReplTargetVolumeID) } } diff --git a/csi-driver/internal/csi/controller/replication_test.go b/csi-driver/internal/csi/controller/replication_test.go index 89f2163ee..0a061ec6b 100644 --- a/csi-driver/internal/csi/controller/replication_test.go +++ b/csi-driver/internal/csi/controller/replication_test.go @@ -101,7 +101,9 @@ func TestEnableVolumeReplicationResolvesToTargetWhenGivenTheSourceSideOfARelatio mock := newMockSBCLI() defer mock.Close() cs := newReplicationTestServer(t, mock) - mock.volumes[testReplTargetVolumeID] = &mockVolume{UUID: testReplTargetVolumeID, Name: "repl-vol-target", Size: 1 << 30} + mock.volumes[testReplTargetVolumeID] = &mockVolume{ + UUID: testReplTargetVolumeID, Name: "repl-vol-target", Size: 1 << 30, + } mock.replicationRelationship[testReplVolumeID] = map[string]any{ "replication_id": "aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee", "direction": "to_target", diff --git a/helm-charts/charts/simplyblock-operator/templates/numa-resource-plugin.yaml b/helm-charts/charts/simplyblock-operator/templates/numa-resource-plugin.yaml index fd5f463bd..99e7d0856 100644 --- a/helm-charts/charts/simplyblock-operator/templates/numa-resource-plugin.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/numa-resource-plugin.yaml @@ -1,8 +1,5 @@ -<<<<<<< HEAD {{- $storagenodeEnabled := and .Values.storagenode.create .Values.storagenode.enableCpuTopology .Values.storagenode.enableDevicePlugin -}} {{- if or $storagenodeEnabled true -}} -======= ->>>>>>> main --- apiVersion: v1 kind: ServiceAccount @@ -40,7 +37,6 @@ spec: spec: serviceAccountName: simplyblock-numa-resource-plugin priorityClassName: system-node-critical -<<<<<<< HEAD {{- $dsList := .Values.storagenode.daemonsets | default (list) -}} {{- if and $storagenodeEnabled (gt (len $dsList) 0) }} affinity: @@ -57,8 +53,6 @@ spec: {{- end }} {{- end }} {{- else if true }} -======= ->>>>>>> main affinity: nodeAffinity: requiredDuringSchedulingIgnoredDuringExecution: @@ -70,6 +64,7 @@ spec: - matchExpressions: - key: io.simplyblock.storagenodeset operator: Exists + {{- end }} tolerations: # Run on all nodes including control plane - operator: Exists @@ -128,3 +123,5 @@ spec: type: RollingUpdate rollingUpdate: maxUnavailable: 1 + +{{- end -}} diff --git a/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml b/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml index 352304989..9c738e997 100644 --- a/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml @@ -132,8 +132,6 @@ webhooks: service: name: simplyblock-operator-webhook-service namespace: {{ .Release.Namespace }} -<<<<<<< HEAD -======= path: /validate-v1-pvc-vdo-size failurePolicy: Ignore name: vdo-size-floor-validator.simplyblock.io @@ -153,7 +151,6 @@ webhooks: service: name: simplyblock-operator-webhook-service namespace: {{ .Release.Namespace }} ->>>>>>> main path: /validate-storage-simplyblock-io-v1alpha2-operatorops failurePolicy: Fail name: voperatorops.simplyblock.io diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index 29dbc026c..f4e2c29db 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -11684,6 +11684,25 @@ webhooks: resources: - controlplaneops sideEffects: None +- admissionReviewVersions: + - v1 + clientConfig: + service: + name: simplyblock-operator-webhook-service + namespace: simplyblock-operator-system + path: /validate-v1-pvc-vdo-size + failurePolicy: Ignore + name: vdo-size-floor-validator.simplyblock.io + rules: + - apiGroups: + - "" + apiVersions: + - v1 + operations: + - CREATE + resources: + - persistentvolumeclaims + sideEffects: None - admissionReviewVersions: - v1 clientConfig: diff --git a/operator/test/e2e/rbac_test.go b/operator/test/e2e/rbac_test.go index d27c4ad60..35c1ac80b 100644 --- a/operator/test/e2e/rbac_test.go +++ b/operator/test/e2e/rbac_test.go @@ -50,14 +50,14 @@ import ( // after the Manager Describe regardless of Ginkgo's randomized container order. const ( - rbacFooNS = "rbac-cluster-foo" - rbacBarNS = "rbac-cluster-bar" - rbacViewerSA = "viewer-sa" - rbacEditorSA = "editor-sa" - rbacOutsiderSA = "outsider-sa" - rbacScopedSA = "scoped-sa" - rbacScopedRoleName = "rbac-foo-admin" - rbacScopedClusterAllowed = "rbac-allowed" + rbacFooNS = "rbac-cluster-foo" + rbacBarNS = "rbac-cluster-bar" + rbacViewerSA = "viewer-sa" + rbacEditorSA = "editor-sa" + rbacOutsiderSA = "outsider-sa" + rbacScopedSA = "scoped-sa" + rbacScopedRoleName = "rbac-foo-admin" + rbacScopedClusterAllowed = "rbac-allowed" rbacScopedClusterForbidden = "rbac-forbidden" ) From 345dd83fc6179fef239f64d606e1ee90b87b1533 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Wed, 23 Sep 2026 23:09:22 +0100 Subject: [PATCH 144/206] Enable-volume-replication no-ops when the destination volume doesn't exist yet --- .../internal/csi/controller/replication.go | 16 +++++++++++ .../csi/controller/replication_test.go | 27 +++++++++++++++++++ 2 files changed, 43 insertions(+) diff --git a/csi-driver/internal/csi/controller/replication.go b/csi-driver/internal/csi/controller/replication.go index 8fa1a0f8d..1a8b3b181 100644 --- a/csi-driver/internal/csi/controller/replication.go +++ b/csi-driver/internal/csi/controller/replication.go @@ -126,6 +126,22 @@ func (cs *Server) EnableVolumeReplication( return nil, status.Error(codes.Unavailable, err.Error()) } if err := client.EnableVolumeReplication(ctx, h.Handle(), policyID); err != nil { + if errors.Is(err, errs.ErrNotFound) { + // This backend's replication is one-way: the destination never + // carries a persistent, independently-provisioned LVol of its + // own -- the writable clone only comes into existence when + // PromoteVolume clones the last replicated snapshot. csi-addons + // always calls Enable before Promote, unconditionally, for + // whichever side is becoming Primary, so on a first-ever + // relocate (nothing for resolveToLocalReplica to redirect + // through either, since no relationship exists until a promote + // has actually happened) Enable is legitimately handed a handle + // that names nothing yet. There is nothing to attach a policy + // to, and nothing wrong either -- PromoteVolume is what actually + // creates and validates the volume, and is what surfaces a real + // error if there truly is nothing to clone from. + return &replication.EnableVolumeReplicationResponse{}, nil + } return nil, classifyEnableVolumeReplicationError(err) } return &replication.EnableVolumeReplicationResponse{}, nil diff --git a/csi-driver/internal/csi/controller/replication_test.go b/csi-driver/internal/csi/controller/replication_test.go index 0a061ec6b..78a1dd7a0 100644 --- a/csi-driver/internal/csi/controller/replication_test.go +++ b/csi-driver/internal/csi/controller/replication_test.go @@ -130,6 +130,33 @@ func TestEnableVolumeReplicationResolvesToTargetWhenGivenTheSourceSideOfARelatio } } +// simplyblock's replication is one-way and the destination never carries a +// persistent, independently-provisioned LVol of its own (confirmed live +// 2026-09-23, relocate M-02): the writable clone only comes into existence +// when PromoteVolume clones the last replicated snapshot. csi-addons always +// calls EnableVolumeReplication before PromoteVolume, unconditionally, for +// the side that's becoming Primary -- so on a first-ever relocate (no prior +// relationship for resolveToLocalReplica to redirect through either), Enable +// is handed a handle that legitimately names nothing yet. That must no-op +// rather than fail: PromoteVolume is what actually creates and validates the +// volume, and is what surfaces a real error if there's genuinely nothing to +// clone from. +func TestEnableVolumeReplicationNoOpsWhenVolumeDoesNotExistYet(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + notYetClonedVolumeID := "99999999-8888-8888-8888-888888888888" + notYetClonedVolID := sanityClusterID + ":" + sanityPoolUUID + ":" + notYetClonedVolumeID + + _, err := cs.EnableVolumeReplication(context.Background(), &replication.EnableVolumeReplicationRequest{ + VolumeId: notYetClonedVolID, + Parameters: map[string]string{replicationPolicyParam: testReplPolicyID}, + }) + if err != nil { + t.Errorf("EnableVolumeReplication = %v, want nil (no-op: nothing to attach a policy to yet)", err) + } +} + func TestEnableVolumeReplicationMissingPolicyParam(t *testing.T) { mock := newMockSBCLI() defer mock.Close() From eab20fadcd8040d377da4261a449990963a93028 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Thu, 24 Sep 2026 12:16:11 +0100 Subject: [PATCH 145/206] fix(csi-driver): resolve foreign volumeHandle to local replica in DemoteVolume and DisableVolumeReplication --- .../csi/controller/mock_controlplane_test.go | 5 +++ .../internal/csi/controller/replication.go | 8 ++++ .../controller/replication_lifecycle_test.go | 41 +++++++++++++++++++ .../csi/controller/replication_test.go | 34 +++++++++++++++ 4 files changed, 88 insertions(+) diff --git a/csi-driver/internal/csi/controller/mock_controlplane_test.go b/csi-driver/internal/csi/controller/mock_controlplane_test.go index 9dfcc2adc..ea8e6ba9e 100644 --- a/csi-driver/internal/csi/controller/mock_controlplane_test.go +++ b/csi-driver/internal/csi/controller/mock_controlplane_test.go @@ -148,6 +148,10 @@ type mockSBCLI struct { // demoteStatus, when set, is the HTTP status POST .../demote answers with // instead of its default success (204). 202 models "still converging." demoteStatus int + // lastDemoteVolumeID captures which volume's path the last demote call + // landed on, so a test can assert a relationship-resolved call reached + // the TARGET volume rather than the one it was originally given. + lastDemoteVolumeID string // failbackStatus, when set, is the HTTP status POST .../failback answers // with instead of its default success (204). @@ -415,6 +419,7 @@ func (m *mockSBCLI) handleDemote(w http.ResponseWriter, r *http.Request) { if m.lookupVolume(w, volumeID) == nil { return } + m.lastDemoteVolumeID = volumeID if m.demoteStatus != 0 { writeJSON(w, m.demoteStatus, map[string]bool{"demoted": m.demoteStatus == http.StatusNoContent}) return diff --git a/csi-driver/internal/csi/controller/replication.go b/csi-driver/internal/csi/controller/replication.go index 1a8b3b181..38537b1e2 100644 --- a/csi-driver/internal/csi/controller/replication.go +++ b/csi-driver/internal/csi/controller/replication.go @@ -162,6 +162,10 @@ func (cs *Server) DisableVolumeReplication( if err != nil { return nil, status.Error(codes.Unavailable, err.Error()) } + h, client, err = resolveToLocalReplica(ctx, h, client) + if err != nil { + return nil, status.Error(codes.Unavailable, err.Error()) + } if err := client.DisableVolumeReplication(ctx, h.Handle()); err != nil { return nil, classifyDisableVolumeReplicationError(err) } @@ -247,6 +251,10 @@ func (cs *Server) DemoteVolume( if err != nil { return nil, status.Error(codes.Unavailable, err.Error()) } + h, client, err = resolveToLocalReplica(ctx, h, client) + if err != nil { + return nil, status.Error(codes.Unavailable, err.Error()) + } done, err := client.DemoteVolume(ctx, h.Handle()) if err != nil { return nil, classifyDemoteVolumeError(err) diff --git a/csi-driver/internal/csi/controller/replication_lifecycle_test.go b/csi-driver/internal/csi/controller/replication_lifecycle_test.go index 06c62cf81..d6b4ea530 100644 --- a/csi-driver/internal/csi/controller/replication_lifecycle_test.go +++ b/csi-driver/internal/csi/controller/replication_lifecycle_test.go @@ -190,6 +190,47 @@ func TestPromoteVolumeUsesReplicationSourceWhenVolumeIdIsEmpty(t *testing.T) { } } +// Same relationship-resolution requirement as PromoteVolume/EnableVolumeReplication +// (see TestPromoteVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship): +// a second (or later) relocate demotes a volume that itself came into +// existence via an earlier promote, so its PV/PVC still carries the +// ORIGINAL, foreign source's volumeHandle. Confirmed live 2026-09-24 (relocate +// M-02's round trip, B -> A): DemoteVolume issued the RPC against that +// foreign, pre-promote lvol id and got a 404 from the control plane, well +// before ever reaching the actual local replica that had been serving as +// primary. DemoteVolume must resolve a handle whose relationship says +// IsSource and redirect to TargetLvolId first, exactly like Promote and +// Enable already do. +func TestDemoteVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + mock.volumes[testReplTargetVolumeID] = &mockVolume{ + UUID: testReplTargetVolumeID, Name: "repl-vol-target", Size: 1 << 30, + } + mock.replicationRelationship[testReplVolumeID] = map[string]any{ + "replication_id": "aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee", + "direction": "to_target", + "mode": "failover", + "state": "failed_over", + "is_source": true, + "source_cluster_id": sanityClusterID, "source_lvol_id": testReplVolumeID, + "target_cluster_id": sanityClusterID, "target_pool_id": sanityPoolUUID, "target_lvol_id": testReplTargetVolumeID, + "target_nqn": "nqn.test", "target_ns_id": 1, + } + + _, err := cs.DemoteVolume(context.Background(), &replication.DemoteVolumeRequest{ + VolumeId: testReplVolID, + }) + if err != nil { + t.Fatal(err) + } + if mock.lastDemoteVolumeID != testReplTargetVolumeID { + t.Errorf("demote landed on volume %q, want the resolved target %q", + mock.lastDemoteVolumeID, testReplTargetVolumeID) + } +} + func TestDemoteVolumeDone(t *testing.T) { mock := newMockSBCLI() defer mock.Close() diff --git a/csi-driver/internal/csi/controller/replication_test.go b/csi-driver/internal/csi/controller/replication_test.go index 78a1dd7a0..34ae4d881 100644 --- a/csi-driver/internal/csi/controller/replication_test.go +++ b/csi-driver/internal/csi/controller/replication_test.go @@ -255,6 +255,40 @@ func TestDisableVolumeReplicationNotAttachedIsSuccess(t *testing.T) { } } +// Same relationship-resolution requirement as DemoteVolume/PromoteVolume/ +// EnableVolumeReplication: Ramen's own teardown sequence calls Disable right +// after a successful demote, on the SAME volume -- so it inherits the SAME +// foreign, pre-promote handle and must resolve to the local replica before +// the policy detach is attempted against it. +func TestDisableVolumeReplicationResolvesToTargetWhenGivenTheSourceSideOfARelationship(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + mock.volumes[testReplTargetVolumeID] = &mockVolume{ + UUID: testReplTargetVolumeID, Name: "repl-vol-target", Size: 1 << 30, ReplicationPolicyID: testReplPolicyID, + } + mock.replicationRelationship[testReplVolumeID] = map[string]any{ + "replication_id": "aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee", + "direction": "to_target", + "mode": "failover", + "state": "failed_over", + "is_source": true, + "source_cluster_id": sanityClusterID, "source_lvol_id": testReplVolumeID, + "target_cluster_id": sanityClusterID, "target_pool_id": sanityPoolUUID, "target_lvol_id": testReplTargetVolumeID, + "target_nqn": "nqn.test", "target_ns_id": 1, + } + + _, err := cs.DisableVolumeReplication(context.Background(), &replication.DisableVolumeReplicationRequest{ + VolumeId: testReplVolID, + }) + if err != nil { + t.Fatal(err) + } + if got := mock.volumes[testReplTargetVolumeID].ReplicationPolicyID; got != "" { + t.Errorf("target volume's ReplicationPolicyID = %q, want cleared", got) + } +} + func TestDisableVolumeReplicationUsesReplicationSourceWhenVolumeIdIsEmpty(t *testing.T) { mock := newMockSBCLI() defer mock.Close() From b529b860ab5429f06118440a274464cd5deabbbf Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Thu, 24 Sep 2026 13:17:17 +0100 Subject: [PATCH 146/206] fix(atlas-lib): resolve replication relationships via the cluster-scoped endpoint, not the pool-scoped one that requires the volume to still exist --- atlas-lib/controlplane/replication.go | 15 ++++++++++++--- atlas-lib/controlplane/replication_test.go | 18 ++++++++++++++++-- .../csi/controller/mock_controlplane_test.go | 7 ++++++- 3 files changed, 34 insertions(+), 6 deletions(-) diff --git a/atlas-lib/controlplane/replication.go b/atlas-lib/controlplane/replication.go index aa55fd16d..0c610e209 100644 --- a/atlas-lib/controlplane/replication.go +++ b/atlas-lib/controlplane/replication.go @@ -258,13 +258,22 @@ type Relationship struct { // unwrapping to errs.ErrNotFound when h has no replication relationship at // all yet (e.g. a volume never enabled for replication) -- callers treat that // as "use h unchanged," not a failure. +// GetVolumeReplicationRelationship reads the relationship through the +// cluster-scoped endpoint, not the pool-scoped one: the pool-scoped route +// requires the queried volume to still exist (sbcli's FastAPI Volume +// dependency 404s before the handler body runs), but a relationship must stay +// resolvable by SOURCE id after the source volume itself is deleted -- e.g. a +// demoted volume whose fail-over already completed and was reaped by +// lvol_monitor's deferred-removal hold, confirmed live 2026-09-24 (relocate +// M-02's round trip: resolveToLocalReplica needs exactly this to redirect +// DemoteVolume/DisableVolumeReplication on the SECOND hop of a relocate). func (c *Client) GetVolumeReplicationRelationship(ctx context.Context, h lvol.VolumeHandle) (Relationship, error) { - cluster, pool, volume, err := h.Split() + cluster, _, volume, err := h.Split() if err != nil { return Relationship{}, err } - resp, err := c.api.ClustersStoragePoolsVolumesReplicationDetailApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationGetWithResponse( - ctx, cluster, pool, volume) + resp, err := c.api.ClustersReplicationRelationshipsDetailApiV2ClustersClusterIdReplicationRelationshipsLvolIdGetWithResponse( + ctx, cluster, volume) if err != nil { return Relationship{}, fmt.Errorf("replication relationship of volume %s: %w", h, err) } diff --git a/atlas-lib/controlplane/replication_test.go b/atlas-lib/controlplane/replication_test.go index 44d500b37..740371e41 100644 --- a/atlas-lib/controlplane/replication_test.go +++ b/atlas-lib/controlplane/replication_test.go @@ -307,10 +307,24 @@ const ( testTargetVolume = "66666666-6666-6666-6666-666666666667" ) +// GetVolumeReplicationRelationship must call the CLUSTER-scoped endpoint +// (GET /clusters/{cluster_id}/replication/relationships/{lvol_id}), never the +// pool-scoped one (.../storage-pools/{pool}/volumes/{volume}/replication/): +// the pool-scoped route requires the queried volume to still exist (sbcli's +// FastAPI Volume dependency 404s before the handler body even runs), but a +// relationship must stay resolvable by SOURCE id after the source volume +// itself is gone -- e.g. a demoted volume whose fail-over already completed +// and was reaped by lvol_monitor's deferred-removal hold (confirmed live +// 2026-09-24, relocate M-02's round trip: DemoteVolume/DisableVolumeReplication +// on the SECOND hop 404'd resolving through the pool-scoped endpoint against +// exactly this). sbcli's cluster-scoped endpoint is built for this case -- +// its own docstring: "resolvable even when the source volume has been +// deleted... The CSI driver uses this to redirect". func TestClientGetVolumeReplicationRelationship(t *testing.T) { c := newTestClient(t, func(w http.ResponseWriter, r *http.Request) { - if !strings.HasSuffix(r.URL.Path, "/replication/") && !strings.HasSuffix(r.URL.Path, "/replication") { - t.Errorf("unexpected path %q", r.URL.Path) + wantPath := "/api/v2/clusters/" + testCluster + "/replication/relationships/" + testVolume + if r.URL.Path != wantPath { + t.Errorf("path = %q, want %q", r.URL.Path, wantPath) } w.Header().Set("Content-Type", "application/json") _, _ = w.Write([]byte(`{ diff --git a/csi-driver/internal/csi/controller/mock_controlplane_test.go b/csi-driver/internal/csi/controller/mock_controlplane_test.go index ea8e6ba9e..d72e7d1fd 100644 --- a/csi-driver/internal/csi/controller/mock_controlplane_test.go +++ b/csi-driver/internal/csi/controller/mock_controlplane_test.go @@ -190,7 +190,12 @@ func newMockSBCLI() *mockSBCLI { m.locked(m.handleReplicationStatus), ) mux.HandleFunc( - "GET /api/v2/clusters/{clusterID}/storage-pools/{poolID}/volumes/{volumeID}/replication/", + // Cluster-scoped, not pool-scoped: GetVolumeReplicationRelationship + // must stay resolvable by source id after the source volume itself is + // deleted, which the pool-scoped route (requiring the volume to still + // exist) cannot do -- see TestClientGetVolumeReplicationRelationship + // in atlas-lib/controlplane for the live-confirmed reason. + "GET /api/v2/clusters/{clusterID}/replication/relationships/{volumeID}", m.locked(m.handleReplicationRelationship), ) mux.HandleFunc( From 5a01bb1745ae0c9f8a7b51b44d955adf0cb49b18 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Thu, 24 Sep 2026 15:08:50 +0100 Subject: [PATCH 147/206] csi-driver: resolve resync and info reads to the local replica --- .../csi/controller/mock_controlplane_test.go | 7 +- .../internal/csi/controller/replication.go | 24 ++++- .../controller/replication_lifecycle_test.go | 98 ++++++++++++++++++- .../tests/test-plan-csi-addons-replication.md | 41 ++++---- 4 files changed, 146 insertions(+), 24 deletions(-) diff --git a/csi-driver/internal/csi/controller/mock_controlplane_test.go b/csi-driver/internal/csi/controller/mock_controlplane_test.go index d72e7d1fd..ace3b66ea 100644 --- a/csi-driver/internal/csi/controller/mock_controlplane_test.go +++ b/csi-driver/internal/csi/controller/mock_controlplane_test.go @@ -128,7 +128,7 @@ type mockSBCLI struct { // replicationPUTStatus, when set, makes every PUT carrying // replication_policy_id respond with this HTTP status instead of the - // normal idempotent update, modeling a backend refusal (e.g. a policy + // normal idempotent update, modeling a backend refusal (e.g., a policy // that is not active) or a transient failure. replicationPUTStatus int @@ -158,6 +158,10 @@ type mockSBCLI struct { failbackStatus int // lastFailbackBody captures the raw JSON body of the last failback call. lastFailbackBody []byte + // lastFailbackVolumeID captures which volume's path the last failback + // call landed on, so a test can assert the relationship resolution + // redirected it -- the same capture handleFailover keeps for promote. + lastFailbackVolumeID string } func newMockSBCLI() *mockSBCLI { @@ -438,6 +442,7 @@ func (m *mockSBCLI) handleFailback(w http.ResponseWriter, r *http.Request) { return } m.lastFailbackBody, _ = io.ReadAll(r.Body) + m.lastFailbackVolumeID = volumeID if m.failbackStatus != 0 { writeJSON(w, m.failbackStatus, map[string]string{"detail": "injected status"}) return diff --git a/csi-driver/internal/csi/controller/replication.go b/csi-driver/internal/csi/controller/replication.go index 38537b1e2..2ac73aa18 100644 --- a/csi-driver/internal/csi/controller/replication.go +++ b/csi-driver/internal/csi/controller/replication.go @@ -68,7 +68,7 @@ func volumeIDFrom(req volumeIDCarrier) string { // // Returns h and client unchanged when h has no replication relationship yet // (errs.ErrNotFound -- the ordinary case for a volume never enabled for -// replication, e.g. M-01's first-ever protect) or when h already names the +// replication, e.g., M-01's first-ever protect) or when h already names the // target side. func resolveToLocalReplica( ctx context.Context, h *lvol.Handle, client *atlascp.Client, @@ -128,7 +128,7 @@ func (cs *Server) EnableVolumeReplication( if err := client.EnableVolumeReplication(ctx, h.Handle(), policyID); err != nil { if errors.Is(err, errs.ErrNotFound) { // This backend's replication is one-way: the destination never - // carries a persistent, independently-provisioned LVol of its + // carries a persistent, independently provisioned LVol of its // own -- the writable clone only comes into existence when // PromoteVolume clones the last replicated snapshot. csi-addons // always calls Enable before Promote, unconditionally, for @@ -172,7 +172,10 @@ func (cs *Server) DisableVolumeReplication( return &replication.DisableVolumeReplicationResponse{}, nil } -// GetVolumeReplicationInfo returns the volume's replicated-life status. +// GetVolumeReplicationInfo returns the volume's replicated-life status, +// resolved to the local replica first for the same reason as every other +// verb: Ramen polls this with the S3-restored, foreign volumeHandle, whose +// own lvol record may already be reaped (see ResyncVolume). // // The spec's response carries only LastSyncTime in this version // (github.com/csi-addons/spec v0.2.0); lastSyncBytes/lastSyncDuration are not @@ -191,6 +194,10 @@ func (cs *Server) GetVolumeReplicationInfo( if err != nil { return nil, status.Error(codes.Unavailable, err.Error()) } + h, client, err = resolveToLocalReplica(ctx, h, client) + if err != nil { + return nil, status.Error(codes.Unavailable, err.Error()) + } info, err := client.GetVolumeReplicationInfo(ctx, h.Handle()) if err != nil { return nil, classifyGetVolumeReplicationInfoError(err) @@ -270,6 +277,13 @@ func (cs *Server) DemoteVolume( // off the ordinary lag read -- it never cuts over, matching the design's own // "it never merges" (§5.2): cutover is PromoteVolume's job, on a separate, // later call. +// +// The resolveToLocalReplica step is what keeps a relocate's round trip alive: +// the Secondary side's VR carries the ORIGINAL source's volumeHandle +// (S3-restored verbatim), and by the second hop that source lvol record has +// been reaped by lvol_monitor's post-failover hold -- confirmed live +// 2026-09-24, when resync (then the only verb without the resolution) 404ed +// against the dead handle on every reconcile and stalled the relocate back. func (cs *Server) ResyncVolume( ctx context.Context, req *replication.ResyncVolumeRequest, @@ -282,6 +296,10 @@ func (cs *Server) ResyncVolume( if err != nil { return nil, status.Error(codes.Unavailable, err.Error()) } + h, client, err = resolveToLocalReplica(ctx, h, client) + if err != nil { + return nil, status.Error(codes.Unavailable, err.Error()) + } sourceClusterID := req.GetParameters()[sourceClusterIDParam] if err := client.ResyncVolume(ctx, h.Handle(), sourceClusterID); err != nil { return nil, classifyResyncVolumeError(err) diff --git a/csi-driver/internal/csi/controller/replication_lifecycle_test.go b/csi-driver/internal/csi/controller/replication_lifecycle_test.go index d6b4ea530..b1b05b365 100644 --- a/csi-driver/internal/csi/controller/replication_lifecycle_test.go +++ b/csi-driver/internal/csi/controller/replication_lifecycle_test.go @@ -316,7 +316,7 @@ func TestResyncVolumeUsesReplicationSourceWhenVolumeIdIsEmpty(t *testing.T) { } } -func TestResyncVolumeForwardsTheSourceClusterParameter(t *testing.T) { +func TestResyncVolumeSendsTheSourceClusterParameter(t *testing.T) { mock := newMockSBCLI() defer mock.Close() cs := newReplicationTestServer(t, mock) @@ -377,6 +377,102 @@ func TestResyncVolumeNotReadyWhileLagExceedsBudget(t *testing.T) { } } +// Regression: 2026-09-24-resync-foreign-handle-404 — on relocate M-02's round +// trip (B -> A), cluster B's Secondary-role VolumeReplication still carries the +// ORIGINAL cluster-A volumeHandle (Ramen's S3-restore preserves it verbatim), +// and by then the A-side lvol record has been reaped by lvol_monitor's +// LVOL_DEMOTE_FAILOVER_HOLD_SEC deferred removal. ResyncVolume was the only +// Replication verb that skipped resolveToLocalReplica, so it fired the +// failback call at that dead, foreign lvol and got a permanent 404 ("LVol +// 00660ccf... not found") on every reconcile -- the VR never finished becoming +// Secondary and the whole relocate-back stalled. The demote in the very same +// reconcile succeeded, because DemoteVolume resolves. Resync must redirect a +// handle whose relationship says IsSource to the live local replica +// (TargetLvolId), and the relationship must carry it there even though the +// source volume itself no longer exists (the cluster-scoped relationship +// endpoint stays resolvable by a deleted source id, confirmed live +// 2026-09-24). +func TestResyncVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + mock.volumes[testReplTargetVolumeID] = &mockVolume{ + UUID: testReplTargetVolumeID, Name: "repl-vol-target", Size: 1 << 30, + } + mock.replicationRelationship[testReplVolumeID] = map[string]any{ + "replication_id": "aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee", + "direction": "to_target", + "mode": "failover", + "state": "failed_over", + "is_source": true, + "source_cluster_id": sanityClusterID, "source_lvol_id": testReplVolumeID, + "target_cluster_id": sanityClusterID, "target_pool_id": sanityPoolUUID, "target_lvol_id": testReplTargetVolumeID, + "target_nqn": "nqn.test", "target_ns_id": 1, + } + // The source lvol record is gone -- reaped after the failover hold -- so + // any call landing on it 404s, exactly as the live control plane did. + delete(mock.volumes, testReplVolumeID) + + resp, err := cs.ResyncVolume(context.Background(), &replication.ResyncVolumeRequest{ + VolumeId: testReplVolID, + }) + if err != nil { + t.Fatal(err) + } + if mock.lastFailbackVolumeID != testReplTargetVolumeID { + t.Errorf("failback landed on volume %q, want the resolved target %q", + mock.lastFailbackVolumeID, testReplTargetVolumeID) + } + if !resp.Ready { + t.Error("Ready = false, want true: the resolved target reports no lag at all") + } +} + +// Regression: 2026-09-24-resync-foreign-handle-404 — the same missing +// resolution as TestResyncVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship, +// on the standalone info read: Ramen polls GetVolumeReplicationInfo for +// lastSyncTime against the same S3-restored, foreign volumeHandle, so once the +// source record is reaped the read 404s instead of reporting the local +// replica's status. +func TestGetVolumeReplicationInfoResolvesToTargetWhenGivenTheSourceSideOfARelationship(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + mock.volumes[testReplTargetVolumeID] = &mockVolume{ + UUID: testReplTargetVolumeID, Name: "repl-vol-target", Size: 1 << 30, + } + mock.replicationRelationship[testReplVolumeID] = map[string]any{ + "replication_id": "aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee", + "direction": "to_target", + "mode": "failover", + "state": "failed_over", + "is_source": true, + "source_cluster_id": sanityClusterID, "source_lvol_id": testReplVolumeID, + "target_cluster_id": sanityClusterID, "target_pool_id": sanityPoolUUID, "target_lvol_id": testReplTargetVolumeID, + "target_nqn": "nqn.test", "target_ns_id": 1, + } + mock.replicationStatus[testReplTargetVolumeID] = map[string]any{ + "role": "target", "state": "in_sync", + "last_replicated_at": "2026-09-24T13:00:00Z", + "outstanding_count": 0, "outstanding_bytes": 0, + "failing_count": 0, "max_retry_reached": false, "resyncing": false, + } + delete(mock.volumes, testReplVolumeID) + + resp, err := cs.GetVolumeReplicationInfo(context.Background(), &replication.GetVolumeReplicationInfoRequest{ + VolumeId: testReplVolID, + }) + if err != nil { + t.Fatal(err) + } + if resp.LastSyncTime == nil { + t.Fatal("LastSyncTime = nil, want the resolved target's last_replicated_at") + } + if got := resp.LastSyncTime.AsTime().UTC().Format("2006-01-02T15:04:05Z"); got != "2026-09-24T13:00:00Z" { + t.Errorf("LastSyncTime = %s, want the resolved target's 2026-09-24T13:00:00Z", got) + } +} + func TestResyncVolumeBackendFailureIsUnavailable(t *testing.T) { mock := newMockSBCLI() defer mock.Close() diff --git a/operator/docs/tests/test-plan-csi-addons-replication.md b/operator/docs/tests/test-plan-csi-addons-replication.md index b7f13bc56..249c578b1 100644 --- a/operator/docs/tests/test-plan-csi-addons-replication.md +++ b/operator/docs/tests/test-plan-csi-addons-replication.md @@ -39,24 +39,27 @@ File: `csi-driver/internal/csi/controller/replication_test.go` (planned) File: `csi-driver/internal/csi/controller/replication_lifecycle_test.go` -| # | Scenario | Type | Test | -|------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| U-10 | Forced promote: the failover endpoint is called without `planned=true`, ignoring demote state entirely | Positive | `TestPromoteVolumeForced` | -| U-11 | Forced promote repeated after completion: success without a second failover (backend idempotency honored) | Boundary | — | -| U-12 | Planned promote sends `planned=true` on the failover call | Positive | `TestPromoteVolumePlannedSendsThePlannedFlag` | -| U-13 | Planned promote with no demote ever requested: `FAILED_PRECONDITION`, the one case meant to let the vendored controller's own force-escalation take over | Negative | `TestPromoteVolumePlannedWithNoDemoteIsFailedPrecondition` | -| U-14 | Demote: the demote endpoint is called, and the RPC succeeds only once the backend confirms the final snapshot landed | Positive | `TestDemoteVolumeDone` | -| U-15 | Demote repeated on a demoted volume: success (idempotency) | Boundary | `sbcli: test_demote_is_idempotent_once_done` (backend tier; the driver's `TestDemoteVolumeDone` exercises the same success path) | -| U-16 | Resync: the failback endpoint is called, forwarding the `sourceClusterID` class parameter when given | Positive | `TestResyncVolume`, `TestResyncVolumeForwardsTheSourceClusterParameter` | -| U-17 | Resync's `ready` field reflects lag against budget: false while lag exceeds it, true once caught up | Boundary | `TestResyncVolumeNotReadyWhileLagExceedsBudget`, `TestResyncVolumeReadyReflectsLag` | -| U-40 | Planned promote while demote is still converging: `ABORTED` (retryable) — never `FAILED_PRECONDITION`, which the vendored controller auto-escalates to a forced, lossy promote inline with no wait-and-retry grace period of its own | Negative | `TestPromoteVolumePlannedWhileDemoteConvergingIsAborted` | -| U-41 | Demote still converging: `ABORTED` (retryable), non-blocking — the RPC never waits out the backend's own convergence loop | Boundary | `TestDemoteVolumeNotYetDoneIsAborted` | -| U-42 | `demote_lvol` fences the source strictly before triggering the final snapshot, never after (a write landing in the gap would be silently lost) | Positive | `sbcli: test_demote_fences_before_triggering_the_final_snapshot` | -| U-43 | `demote_lvol` re-invoked while pending: checks the marker only, never re-fences or re-triggers | Boundary | `sbcli: test_demote_does_not_refence_or_retrigger_once_pending` | -| U-44 | `demote_lvol` completes once the triggered snapshot carries the replicated marker | Positive | `sbcli: test_demote_completes_once_the_snapshot_carries_the_replicated_marker` | -| U-45 | `demote_lvol` surfaces a snapshot-creation failure without recording pending state | Negative | `sbcli: test_demote_surfaces_a_snapshot_creation_failure` | -| U-46 | The `failover` route's three-way planned-gate branch: demoted proceeds, converging is 409, no relationship is 412 | Positive/Negative | `sbcli: test_planned_failover_proceeds_once_demoted`, `test_planned_failover_while_demote_is_converging_is_409`, `test_planned_failover_without_any_demote_is_412` | -| U-47 | Unplanned failover ignores demote state, unchanged from before P0-3 | Regression | `sbcli: test_unplanned_failover_ignores_demote_state` | +| # | Scenario | Type | Test | +|------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| U-10 | Forced promote: the failover endpoint is called without `planned=true`, ignoring demote state entirely | Positive | `TestPromoteVolumeForced` | +| U-11 | Forced promote repeated after completion: success without a second failover (backend idempotency honored) | Boundary | — | +| U-12 | Planned promote sends `planned=true` on the failover call | Positive | `TestPromoteVolumePlannedSendsThePlannedFlag` | +| U-13 | Planned promote with no demote ever requested: `FAILED_PRECONDITION`, the one case meant to let the vendored controller's own force-escalation take over | Negative | `TestPromoteVolumePlannedWithNoDemoteIsFailedPrecondition` | +| U-14 | Demote: the demote endpoint is called, and the RPC succeeds only once the backend confirms the final snapshot landed | Positive | `TestDemoteVolumeDone` | +| U-15 | Demote repeated on a demoted volume: success (idempotency) | Boundary | `sbcli: test_demote_is_idempotent_once_done` (backend tier; the driver's `TestDemoteVolumeDone` exercises the same success path) | +| U-16 | Resync: the failback endpoint is called, forwarding the `sourceClusterID` class parameter when given | Positive | `TestResyncVolume`, `TestResyncVolumeSendsTheSourceClusterParameter` | +| U-17 | Resync's `ready` field reflects lag against budget: false while lag exceeds it, true once caught up | Boundary | `TestResyncVolumeNotReadyWhileLagExceedsBudget`, `TestResyncVolumeReadyReflectsLag` | +| U-40 | Planned promote while demote is still converging: `ABORTED` (retryable) — never `FAILED_PRECONDITION`, which the vendored controller auto-escalates to a forced, lossy promote inline with no wait-and-retry grace period of its own | Negative | `TestPromoteVolumePlannedWhileDemoteConvergingIsAborted` | +| U-41 | Demote still converging: `ABORTED` (retryable), non-blocking — the RPC never waits out the backend's own convergence loop | Boundary | `TestDemoteVolumeNotYetDoneIsAborted` | +| U-42 | `demote_lvol` fences the source strictly before triggering the final snapshot, never after (a write landing in the gap would be silently lost) | Positive | `sbcli: test_demote_fences_before_triggering_the_final_snapshot` | +| U-43 | `demote_lvol` re-invoked while pending: checks the marker only, never re-fences or re-triggers | Boundary | `sbcli: test_demote_does_not_refence_or_retrigger_once_pending` | +| U-44 | `demote_lvol` completes once the triggered snapshot carries the replicated marker | Positive | `sbcli: test_demote_completes_once_the_snapshot_carries_the_replicated_marker` | +| U-45 | `demote_lvol` surfaces a snapshot-creation failure without recording pending state | Negative | `sbcli: test_demote_surfaces_a_snapshot_creation_failure` | +| U-46 | The `failover` route's three-way planned-gate branch: demoted proceeds, converging is 409, no relationship is 412 | Positive/Negative | `sbcli: test_planned_failover_proceeds_once_demoted`, `test_planned_failover_while_demote_is_converging_is_409`, `test_planned_failover_without_any_demote_is_412` | +| U-47 | Unplanned failover ignores demote state, unchanged from before P0-3 | Regression | `sbcli: test_unplanned_failover_ignores_demote_state` | +| U-48 | Every verb given the SOURCE side of an existing relationship resolves to the local replica (`TargetLvolId`) before acting — Ramen's S3-restore hands the destination cluster the original source's volumeHandle verbatim | Positive | `TestEnableVolumeReplicationResolvesToTargetWhenGivenTheSourceSideOfARelationship`, `TestDisableVolumeReplicationResolvesToTargetWhenGivenTheSourceSideOfARelationship`, `TestPromoteVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship`, `TestDemoteVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship` | +| U-49 | Resync resolves the foreign source handle even after the source lvol record itself was reaped (2026-09-24-resync-foreign-handle-404: the only verb without the resolution 404ed on every reconcile and stalled the relocate back) | Regression | `TestResyncVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship` | +| U-50 | The standalone info read resolves the foreign source handle likewise, so lastSyncTime keeps flowing after the source record is reaped (2026-09-24-resync-foreign-handle-404) | Regression | `TestGetVolumeReplicationInfoResolvesToTargetWhenGivenTheSourceSideOfARelationship` | ### Condition Derivation (design §6.2) @@ -175,7 +178,7 @@ Two live simplyblock clusters with the chart-deployed csi-addons machinery. The | Axis | Values covered | IDs | Not covered | |------------------|------------------------------------------------------------------------|-----------------------------------------------------------------------------|-----------------------------------------------------------| -| Verb lifecycle | enable, disable, info, forced promote, planned promote, demote, resync | U-01, U-02, U-04 … U-10, U-12 … U-17, U-40 … U-47, I-02 … I-04, E-01 … E-04 | — | +| Verb lifecycle | enable, disable, info, forced promote, planned promote, demote, resync | U-01, U-02, U-04 … U-10, U-12 … U-17, U-40 … U-50, I-02 … I-04, E-01 … E-04 | — | | Idempotency | repeat enable, disable, demote; re-drive after restart | U-02, U-05, U-15, I-07 | repeated promote (U-11), repeated resync | | Conditions | healthy, degraded, error, staleness, resyncing, disabled | U-18 … U-22, E-05 | condition behavior across backend upgrade | | Coexistence | slot skip, one-owner refusal, concurrent claim | U-26, U-27, M-02 | migration of an annotated volume onto a VolumeReplication | From 097695777df07f36189ecea17638c1e1659cbc8a Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Thu, 24 Sep 2026 16:27:25 +0100 Subject: [PATCH 148/206] csi-driver: clean up retired replica chain on DeleteVolume, never the active copy; parse ActiveLvolID in atlas-lib --- atlas-lib/controlplane/replication.go | 16 +++- atlas-lib/controlplane/replication_test.go | 42 +++++++++- csi-driver/internal/csi/controller/volume.go | 55 +++++++++++++ .../internal/csi/controller/volume_test.go | 81 +++++++++++++++++++ .../tests/test-plan-csi-addons-replication.md | 47 ++++++----- 5 files changed, 212 insertions(+), 29 deletions(-) diff --git a/atlas-lib/controlplane/replication.go b/atlas-lib/controlplane/replication.go index 0c610e209..ca85cd07d 100644 --- a/atlas-lib/controlplane/replication.go +++ b/atlas-lib/controlplane/replication.go @@ -247,22 +247,31 @@ type Relationship struct { TargetClusterID string TargetPoolID string TargetLvolID string + + // ActiveLvolID names the volume currently serving the pairing's data, + // resolved transitively by the backend across chained fail-overs (a + // relocate round trip leaves the FIRST pairing's target pointing at a + // clone that a SECOND pairing has since superseded). A caller cleaning up + // retired copies keys off this: any member that is not the active volume + // is garbage, and the active one must never be touched through a stale + // handle. Empty when the backend predates the field. + ActiveLvolID string } // GetVolumeReplicationRelationship resolves h to its replication pairing. // This is what a caller handed a volume identity inherited from the OTHER -// side of a pairing (e.g. a destination PVC whose PV was restored carrying +// side of a pairing (e.g., a destination PVC whose PV was restored carrying // the source's own volumeHandle) uses to find the volume it should actually // operate on locally: TargetClusterID/TargetPoolID/TargetLvolID name that // volume regardless of which side h itself named. Returns an error // unwrapping to errs.ErrNotFound when h has no replication relationship at -// all yet (e.g. a volume never enabled for replication) -- callers treat that +// all yet (e.g., a volume never enabled for replication) -- callers treat that // as "use h unchanged," not a failure. // GetVolumeReplicationRelationship reads the relationship through the // cluster-scoped endpoint, not the pool-scoped one: the pool-scoped route // requires the queried volume to still exist (sbcli's FastAPI Volume // dependency 404s before the handler body runs), but a relationship must stay -// resolvable by SOURCE id after the source volume itself is deleted -- e.g. a +// resolvable by SOURCE id after the source volume itself is deleted -- e.g., a // demoted volume whose fail-over already completed and was reaped by // lvol_monitor's deferred-removal hold, confirmed live 2026-09-24 (relocate // M-02's round trip: resolveToLocalReplica needs exactly this to redirect @@ -288,5 +297,6 @@ func (c *Client) GetVolumeReplicationRelationship(ctx context.Context, h lvol.Vo TargetClusterID: uuidPtrString(d.TargetClusterId), TargetPoolID: uuidPtrString(d.TargetPoolId), TargetLvolID: uuidPtrString(d.TargetLvolId), + ActiveLvolID: uuidPtrString(d.ActiveLvolId), }, nil } diff --git a/atlas-lib/controlplane/replication_test.go b/atlas-lib/controlplane/replication_test.go index 740371e41..e5a1ffa35 100644 --- a/atlas-lib/controlplane/replication_test.go +++ b/atlas-lib/controlplane/replication_test.go @@ -138,8 +138,8 @@ func TestClientGetVolumeReplicationInfo(t *testing.T) { } } -// A volume that never replicated is a valid, non-error answer: role "none", -// state "not_replicating", and every timing/lag field null. +// A volume that never replicated is a valid, non-error answer: role "none," +// state "not_replicating," and every timing/lag field null. func TestClientGetVolumeReplicationInfoNeverReplicated(t *testing.T) { c := newTestClient(t, func(w http.ResponseWriter, r *http.Request) { w.Header().Set("Content-Type", "application/json") @@ -313,13 +313,13 @@ const ( // the pool-scoped route requires the queried volume to still exist (sbcli's // FastAPI Volume dependency 404s before the handler body even runs), but a // relationship must stay resolvable by SOURCE id after the source volume -// itself is gone -- e.g. a demoted volume whose fail-over already completed +// itself is gone -- e.g., a demoted volume whose fail-over already completed // and was reaped by lvol_monitor's deferred-removal hold (confirmed live // 2026-09-24, relocate M-02's round trip: DemoteVolume/DisableVolumeReplication // on the SECOND hop 404'd resolving through the pool-scoped endpoint against // exactly this). sbcli's cluster-scoped endpoint is built for this case -- // its own docstring: "resolvable even when the source volume has been -// deleted... The CSI driver uses this to redirect". +// deleted... The CSI driver uses this to redirect." func TestClientGetVolumeReplicationRelationship(t *testing.T) { c := newTestClient(t, func(w http.ResponseWriter, r *http.Request) { wantPath := "/api/v2/clusters/" + testCluster + "/replication/relationships/" + testVolume @@ -357,6 +357,40 @@ func TestClientGetVolumeReplicationRelationship(t *testing.T) { } } +// Regression: 2026-09-24-delete-foreign-handle-leak — the backend resolves +// active_lvol_id transitively across chained fail-overs, and the driver's +// retired-copy cleanup keys off it: without it parsed, the cleanup cannot +// tell a stale clone from the volume the workload currently runs on. +func TestClientGetVolumeReplicationRelationshipParsesTheActiveVolume(t *testing.T) { + c := newTestClient(t, func(w http.ResponseWriter, r *http.Request) { + w.Header().Set("Content-Type", "application/json") + _, _ = w.Write([]byte(`{ + "replication_id": "` + testCluster + `", + "direction": "to_target", + "mode": "failover", + "state": "failed_over", + "is_source": true, + "source_cluster_id": "` + testCluster + `", + "source_lvol_id": "` + testVolume + `", + "target_cluster_id": "` + testTargetCluster + `", + "target_pool_id": "` + testTargetPool + `", + "target_lvol_id": "` + testTargetVolume + `", + "target_nqn": "nqn.test", + "target_ns_id": 1, + "active": "target", + "active_lvol_id": "` + testTargetVolume + `" + }`)) + }) + + rel, err := c.GetVolumeReplicationRelationship(context.Background(), testHandle) + if err != nil { + t.Fatal(err) + } + if rel.ActiveLvolID != testTargetVolume { + t.Errorf("ActiveLvolID = %q, want %q", rel.ActiveLvolID, testTargetVolume) + } +} + // Queried by the TARGET volume's own id, the relationship still reports the // SAME fixed target_* fields -- this is what lets a caller always resolve to // target_* regardless of which side of the pairing it was handed, without diff --git a/csi-driver/internal/csi/controller/volume.go b/csi-driver/internal/csi/controller/volume.go index da802d331..ef69fe4db 100644 --- a/csi-driver/internal/csi/controller/volume.go +++ b/csi-driver/internal/csi/controller/volume.go @@ -11,7 +11,9 @@ import ( "strings" "github.com/container-storage-interface/spec/lib/go/csi" + "github.com/simplyblock/atlas/errs" "github.com/simplyblock/atlas/kube" + "github.com/simplyblock/atlas/lvol" "google.golang.org/grpc/codes" "google.golang.org/grpc/status" "k8s.io/klog" @@ -199,9 +201,62 @@ func (cs *Server) DeleteVolume( return nil, classifyDeleteVolumeError(err) } + if err := cs.deleteRetiredReplicaChain(ctx, volumeID); err != nil { + klog.Errorf("failed to delete retired replication copies of volume %s: %v", volumeID, err) + return nil, classifyDeleteVolumeError(err) + } + return &csi.DeleteVolumeResponse{}, nil } +// deleteRetiredReplicaChain removes the RETIRED members of the volume's +// replication pairing chain. Ramen's S3-restore keeps every PV on the +// ORIGINAL volumeHandle across fail-overs, so the DeleteVolume that cleans up +// a retired side arrives carrying an identity whose own lvol record is +// already reaped, while the actual local copy -- the pairing's superseded +// target clone -- lives on untouched (confirmed live 2026-09-24, relocate +// M-02's round trip: cluster B kept clone 6102a48e, policy still attached, +// after its PV was deleted). Walking source->target and deleting every member +// that is NOT the pairing's active volume removes exactly those leftovers. +// +// The active volume is never deleted through a stale handle: the workload is +// running on it, and its own deletion arrives through this same path once no +// pairing supersedes it. A missing ActiveLvolID (backend predating the field) +// deletes nothing, erring toward leaking a clone over destroying live data. +func (cs *Server) deleteRetiredReplicaChain(ctx context.Context, volumeID string) error { + h, err := csicommon.ParseVolumeHandle(volumeID) + if err != nil { + return nil // unparseable handles were already tolerated as deleted above + } + cur := h + for range 8 { // one hop per past fail-over; capped far above any real chain + client, err := clusters.ReplicationClient(ctx, cur.ClusterID) + if err != nil { + return err + } + rel, err := client.GetVolumeReplicationRelationship(ctx, cur.Handle()) + if err != nil { + if errors.Is(err, errs.ErrNotFound) { + return nil // no pairing: an ordinary volume, nothing retired to clean + } + return err + } + if rel.ActiveLvolID == "" || rel.TargetLvolID == "" || rel.TargetLvolID == rel.ActiveLvolID { + return nil + } + target := &lvol.Handle{ClusterID: rel.TargetClusterID, PoolRef: rel.TargetPoolID, VolumeID: rel.TargetLvolID} + targetClient, err := clusters.ReplicationClient(ctx, target.ClusterID) + if err != nil { + return err + } + if err := targetClient.DeleteVolume(ctx, target.Handle()); err != nil { + return err + } + cur = target + } + return nil +} + func (cs *Server) prepareCreateVolumeReq( ctx context.Context, req *csi.CreateVolumeRequest, diff --git a/csi-driver/internal/csi/controller/volume_test.go b/csi-driver/internal/csi/controller/volume_test.go index c35d0f698..67e16f60e 100644 --- a/csi-driver/internal/csi/controller/volume_test.go +++ b/csi-driver/internal/csi/controller/volume_test.go @@ -100,6 +100,87 @@ func TestDeleteVolume_ControlPlaneErrorMapping(t *testing.T) { }) } +// testReplActiveVolumeID stands in for the volume a relocate round trip +// leaves actually serving the workload -- the SECOND hop's clone, which the +// FIRST pairing's records know only as active_lvol_id. +const testReplActiveVolumeID = "88888888-8888-8888-8888-888888888890" + +// Regression: 2026-09-24-delete-foreign-handle-leak — Ramen keeps every PV on +// the ORIGINAL volumeHandle across fail-overs, so the DeleteVolume that +// cleans up a retired side arrives carrying an identity whose own lvol record +// is already reaped, while the actual local copy -- the pairing's superseded +// target clone -- lives on untouched (confirmed live 2026-09-24, relocate +// M-02 round trip: cluster B kept clone 6102a48e, with the replication policy +// still attached to it, after its PV was deleted). DeleteVolume must follow +// the relationship and remove the retired, non-active members. +func TestDeleteVolumeRemovesTheRetiredReplicaBehindAForeignHandle(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + mock.volumes[testReplTargetVolumeID] = &mockVolume{ + UUID: testReplTargetVolumeID, Name: "repl-vol-retired-clone", Size: 1 << 30, + } + mock.volumes[testReplActiveVolumeID] = &mockVolume{ + UUID: testReplActiveVolumeID, Name: "repl-vol-active", Size: 1 << 30, + } + mock.replicationRelationship[testReplVolumeID] = map[string]any{ + "replication_id": "aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee", + "direction": "to_target", + "mode": "failover", + "state": "failed_over", + "is_source": true, + "source_cluster_id": sanityClusterID, "source_lvol_id": testReplVolumeID, + "target_cluster_id": sanityClusterID, "target_pool_id": sanityPoolUUID, "target_lvol_id": testReplTargetVolumeID, + "target_nqn": "nqn.test", "target_ns_id": 1, + "active": "target", "active_lvol_id": testReplActiveVolumeID, + } + delete(mock.volumes, testReplVolumeID) // the original's record, reaped after fail-over + + _, err := cs.DeleteVolume(context.Background(), &csi.DeleteVolumeRequest{VolumeId: testReplVolID}) + if err != nil { + t.Fatal(err) + } + if _, ok := mock.volumes[testReplTargetVolumeID]; ok { + t.Error("the retired clone still exists: DeleteVolume never followed the relationship to it") + } + if _, ok := mock.volumes[testReplActiveVolumeID]; !ok { + t.Error("the ACTIVE volume was deleted: the workload was running on it") + } +} + +// The safety half of the same contract, pinned so the cleanup above can never +// be "fixed" into deleting the live side: when the pairing's target IS the +// active volume (a fail-over whose destination still serves the workload), +// a DeleteVolume carrying the dead source handle must delete nothing. +func TestDeleteVolumeNeverDeletesTheActiveReplicaThroughADeadHandle(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + mock.volumes[testReplTargetVolumeID] = &mockVolume{ + UUID: testReplTargetVolumeID, Name: "repl-vol-live", Size: 1 << 30, + } + mock.replicationRelationship[testReplVolumeID] = map[string]any{ + "replication_id": "aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee", + "direction": "to_target", + "mode": "failover", + "state": "failed_over", + "is_source": true, + "source_cluster_id": sanityClusterID, "source_lvol_id": testReplVolumeID, + "target_cluster_id": sanityClusterID, "target_pool_id": sanityPoolUUID, "target_lvol_id": testReplTargetVolumeID, + "target_nqn": "nqn.test", "target_ns_id": 1, + "active": "target", "active_lvol_id": testReplTargetVolumeID, + } + delete(mock.volumes, testReplVolumeID) + + _, err := cs.DeleteVolume(context.Background(), &csi.DeleteVolumeRequest{VolumeId: testReplVolID}) + if err != nil { + t.Fatal(err) + } + if _, ok := mock.volumes[testReplTargetVolumeID]; !ok { + t.Error("the ACTIVE volume was deleted through the dead source handle") + } +} + // TestControllerExpandVolume_ControlPlaneErrorMapping drives ControllerExpandVolume // through every control-plane response. func TestControllerExpandVolume_ControlPlaneErrorMapping(t *testing.T) { diff --git a/operator/docs/tests/test-plan-csi-addons-replication.md b/operator/docs/tests/test-plan-csi-addons-replication.md index 249c578b1..752074200 100644 --- a/operator/docs/tests/test-plan-csi-addons-replication.md +++ b/operator/docs/tests/test-plan-csi-addons-replication.md @@ -39,27 +39,30 @@ File: `csi-driver/internal/csi/controller/replication_test.go` (planned) File: `csi-driver/internal/csi/controller/replication_lifecycle_test.go` -| # | Scenario | Type | Test | -|------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| U-10 | Forced promote: the failover endpoint is called without `planned=true`, ignoring demote state entirely | Positive | `TestPromoteVolumeForced` | -| U-11 | Forced promote repeated after completion: success without a second failover (backend idempotency honored) | Boundary | — | -| U-12 | Planned promote sends `planned=true` on the failover call | Positive | `TestPromoteVolumePlannedSendsThePlannedFlag` | -| U-13 | Planned promote with no demote ever requested: `FAILED_PRECONDITION`, the one case meant to let the vendored controller's own force-escalation take over | Negative | `TestPromoteVolumePlannedWithNoDemoteIsFailedPrecondition` | -| U-14 | Demote: the demote endpoint is called, and the RPC succeeds only once the backend confirms the final snapshot landed | Positive | `TestDemoteVolumeDone` | -| U-15 | Demote repeated on a demoted volume: success (idempotency) | Boundary | `sbcli: test_demote_is_idempotent_once_done` (backend tier; the driver's `TestDemoteVolumeDone` exercises the same success path) | -| U-16 | Resync: the failback endpoint is called, forwarding the `sourceClusterID` class parameter when given | Positive | `TestResyncVolume`, `TestResyncVolumeSendsTheSourceClusterParameter` | -| U-17 | Resync's `ready` field reflects lag against budget: false while lag exceeds it, true once caught up | Boundary | `TestResyncVolumeNotReadyWhileLagExceedsBudget`, `TestResyncVolumeReadyReflectsLag` | -| U-40 | Planned promote while demote is still converging: `ABORTED` (retryable) — never `FAILED_PRECONDITION`, which the vendored controller auto-escalates to a forced, lossy promote inline with no wait-and-retry grace period of its own | Negative | `TestPromoteVolumePlannedWhileDemoteConvergingIsAborted` | -| U-41 | Demote still converging: `ABORTED` (retryable), non-blocking — the RPC never waits out the backend's own convergence loop | Boundary | `TestDemoteVolumeNotYetDoneIsAborted` | -| U-42 | `demote_lvol` fences the source strictly before triggering the final snapshot, never after (a write landing in the gap would be silently lost) | Positive | `sbcli: test_demote_fences_before_triggering_the_final_snapshot` | -| U-43 | `demote_lvol` re-invoked while pending: checks the marker only, never re-fences or re-triggers | Boundary | `sbcli: test_demote_does_not_refence_or_retrigger_once_pending` | -| U-44 | `demote_lvol` completes once the triggered snapshot carries the replicated marker | Positive | `sbcli: test_demote_completes_once_the_snapshot_carries_the_replicated_marker` | -| U-45 | `demote_lvol` surfaces a snapshot-creation failure without recording pending state | Negative | `sbcli: test_demote_surfaces_a_snapshot_creation_failure` | -| U-46 | The `failover` route's three-way planned-gate branch: demoted proceeds, converging is 409, no relationship is 412 | Positive/Negative | `sbcli: test_planned_failover_proceeds_once_demoted`, `test_planned_failover_while_demote_is_converging_is_409`, `test_planned_failover_without_any_demote_is_412` | -| U-47 | Unplanned failover ignores demote state, unchanged from before P0-3 | Regression | `sbcli: test_unplanned_failover_ignores_demote_state` | -| U-48 | Every verb given the SOURCE side of an existing relationship resolves to the local replica (`TargetLvolId`) before acting — Ramen's S3-restore hands the destination cluster the original source's volumeHandle verbatim | Positive | `TestEnableVolumeReplicationResolvesToTargetWhenGivenTheSourceSideOfARelationship`, `TestDisableVolumeReplicationResolvesToTargetWhenGivenTheSourceSideOfARelationship`, `TestPromoteVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship`, `TestDemoteVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship` | -| U-49 | Resync resolves the foreign source handle even after the source lvol record itself was reaped (2026-09-24-resync-foreign-handle-404: the only verb without the resolution 404ed on every reconcile and stalled the relocate back) | Regression | `TestResyncVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship` | -| U-50 | The standalone info read resolves the foreign source handle likewise, so lastSyncTime keeps flowing after the source record is reaped (2026-09-24-resync-foreign-handle-404) | Regression | `TestGetVolumeReplicationInfoResolvesToTargetWhenGivenTheSourceSideOfARelationship` | +| # | Scenario | Type | Test | +|------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| U-10 | Forced promote: the failover endpoint is called without `planned=true`, ignoring demote state entirely | Positive | `TestPromoteVolumeForced` | +| U-11 | Forced promote repeated after completion: success without a second failover (backend idempotency honored) | Boundary | — | +| U-12 | Planned promote sends `planned=true` on the failover call | Positive | `TestPromoteVolumePlannedSendsThePlannedFlag` | +| U-13 | Planned promote with no demote ever requested: `FAILED_PRECONDITION`, the one case meant to let the vendored controller's own force-escalation take over | Negative | `TestPromoteVolumePlannedWithNoDemoteIsFailedPrecondition` | +| U-14 | Demote: the demote endpoint is called, and the RPC succeeds only once the backend confirms the final snapshot landed | Positive | `TestDemoteVolumeDone` | +| U-15 | Demote repeated on a demoted volume: success (idempotency) | Boundary | `sbcli: test_demote_is_idempotent_once_done` (backend tier; the driver's `TestDemoteVolumeDone` exercises the same success path) | +| U-16 | Resync: the failback endpoint is called, forwarding the `sourceClusterID` class parameter when given | Positive | `TestResyncVolume`, `TestResyncVolumeSendsTheSourceClusterParameter` | +| U-17 | Resync's `ready` field reflects lag against budget: false while lag exceeds it, true once caught up | Boundary | `TestResyncVolumeNotReadyWhileLagExceedsBudget`, `TestResyncVolumeReadyReflectsLag` | +| U-40 | Planned promote while demote is still converging: `ABORTED` (retryable) — never `FAILED_PRECONDITION`, which the vendored controller auto-escalates to a forced, lossy promote inline with no wait-and-retry grace period of its own | Negative | `TestPromoteVolumePlannedWhileDemoteConvergingIsAborted` | +| U-41 | Demote still converging: `ABORTED` (retryable), non-blocking — the RPC never waits out the backend's own convergence loop | Boundary | `TestDemoteVolumeNotYetDoneIsAborted` | +| U-42 | `demote_lvol` fences the source strictly before triggering the final snapshot, never after (a write landing in the gap would be silently lost) | Positive | `sbcli: test_demote_fences_before_triggering_the_final_snapshot` | +| U-43 | `demote_lvol` re-invoked while pending: checks the marker only, never re-fences or re-triggers | Boundary | `sbcli: test_demote_does_not_refence_or_retrigger_once_pending` | +| U-44 | `demote_lvol` completes once the triggered snapshot carries the replicated marker | Positive | `sbcli: test_demote_completes_once_the_snapshot_carries_the_replicated_marker` | +| U-45 | `demote_lvol` surfaces a snapshot-creation failure without recording pending state | Negative | `sbcli: test_demote_surfaces_a_snapshot_creation_failure` | +| U-46 | The `failover` route's three-way planned-gate branch: demoted proceeds, converging is 409, no relationship is 412 | Positive/Negative | `sbcli: test_planned_failover_proceeds_once_demoted`, `test_planned_failover_while_demote_is_converging_is_409`, `test_planned_failover_without_any_demote_is_412` | +| U-47 | Unplanned failover ignores demote state, unchanged from before P0-3 | Regression | `sbcli: test_unplanned_failover_ignores_demote_state` | +| U-48 | Every verb given the SOURCE side of an existing relationship resolves to the local replica (`TargetLvolId`) before acting — Ramen's S3-restore hands the destination cluster the original source's volumeHandle verbatim | Positive | `TestEnableVolumeReplicationResolvesToTargetWhenGivenTheSourceSideOfARelationship`, `TestDisableVolumeReplicationResolvesToTargetWhenGivenTheSourceSideOfARelationship`, `TestPromoteVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship`, `TestDemoteVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship` | +| U-49 | Resync resolves the foreign source handle even after the source lvol record itself was reaped (2026-09-24-resync-foreign-handle-404: the only verb without the resolution 404ed on every reconcile and stalled the relocate back) | Regression | `TestResyncVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship` | +| U-50 | The standalone info read resolves the foreign source handle likewise, so lastSyncTime keeps flowing after the source record is reaped (2026-09-24-resync-foreign-handle-404) | Regression | `TestGetVolumeReplicationInfoResolvesToTargetWhenGivenTheSourceSideOfARelationship` | +| U-51 | Promote while the fail-back's first reverse-replicated snapshot is still in flight: control-plane 409 (converging, retryable), never a stack-traced 500; the same backend message with no replication configured stays 500 | Regression | `sbcli: test_failover_while_the_first_replicated_snapshot_is_in_flight_is_409`, `test_failover_with_no_snapshot_and_no_replication_configured_stays_500` | +| U-52 | DeleteVolume carrying a reaped, foreign handle follows the relationship and removes the retired, non-active members (2026-09-24-delete-foreign-handle-leak); the `Relationship.ActiveLvolID` parse in atlas-lib is part of the same fix | Regression | `TestDeleteVolumeRemovesTheRetiredReplicaBehindAForeignHandle`, `atlas-lib: TestClientGetVolumeReplicationRelationshipParsesTheActiveVolume` | +| U-53 | DeleteVolume never deletes the pairing's ACTIVE volume through a dead source handle — the workload is running on it | Negative | `TestDeleteVolumeNeverDeletesTheActiveReplicaThroughADeadHandle` | ### Condition Derivation (design §6.2) @@ -178,7 +181,7 @@ Two live simplyblock clusters with the chart-deployed csi-addons machinery. The | Axis | Values covered | IDs | Not covered | |------------------|------------------------------------------------------------------------|-----------------------------------------------------------------------------|-----------------------------------------------------------| -| Verb lifecycle | enable, disable, info, forced promote, planned promote, demote, resync | U-01, U-02, U-04 … U-10, U-12 … U-17, U-40 … U-50, I-02 … I-04, E-01 … E-04 | — | +| Verb lifecycle | enable, disable, info, forced promote, planned promote, demote, resync | U-01, U-02, U-04 … U-10, U-12 … U-17, U-40 … U-53, I-02 … I-04, E-01 … E-04 | — | | Idempotency | repeat enable, disable, demote; re-drive after restart | U-02, U-05, U-15, I-07 | repeated promote (U-11), repeated resync | | Conditions | healthy, degraded, error, staleness, resyncing, disabled | U-18 … U-22, E-05 | condition behavior across backend upgrade | | Coexistence | slot skip, one-owner refusal, concurrent claim | U-26, U-27, M-02 | migration of an annotated volume onto a VolumeReplication | From dba955aefd119345569abf040d3980fcd418725c Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Thu, 24 Sep 2026 16:41:16 +0100 Subject: [PATCH 149/206] fixed unparsable --- csi-driver/internal/csi/controller/volume.go | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/csi-driver/internal/csi/controller/volume.go b/csi-driver/internal/csi/controller/volume.go index ef69fe4db..872b73525 100644 --- a/csi-driver/internal/csi/controller/volume.go +++ b/csi-driver/internal/csi/controller/volume.go @@ -226,7 +226,7 @@ func (cs *Server) DeleteVolume( func (cs *Server) deleteRetiredReplicaChain(ctx context.Context, volumeID string) error { h, err := csicommon.ParseVolumeHandle(volumeID) if err != nil { - return nil // unparseable handles were already tolerated as deleted above + return nil // unparsable handles were already tolerated as deleted above } cur := h for range 8 { // one hop per past fail-over; capped far above any real chain From 7438eb4a488ff99205a23feb9778d59913724dcc Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Thu, 24 Sep 2026 17:20:25 +0100 Subject: [PATCH 150/206] csi-driver: walk chained replication pairings to the active volume in node redirect and controller resolution --- .../internal/csi/controller/replication.go | 44 +++-- .../controller/replication_lifecycle_test.go | 57 +++++++ csi-driver/internal/csi/node/stage.go | 2 +- csi-driver/internal/csi/node/stats.go | 102 +++++++---- csi-driver/internal/csi/node/stats_test.go | 159 ++++++++++++++++++ .../tests/test-plan-csi-addons-replication.md | 4 +- 6 files changed, 316 insertions(+), 52 deletions(-) create mode 100644 csi-driver/internal/csi/node/stats_test.go diff --git a/csi-driver/internal/csi/controller/replication.go b/csi-driver/internal/csi/controller/replication.go index 2ac73aa18..90ffd2e93 100644 --- a/csi-driver/internal/csi/controller/replication.go +++ b/csi-driver/internal/csi/controller/replication.go @@ -70,25 +70,43 @@ func volumeIDFrom(req volumeIDCarrier) string { // (errs.ErrNotFound -- the ordinary case for a volume never enabled for // replication, e.g., M-01's first-ever protect) or when h already names the // target side. +// +// The resolution WALKS chained pairings rather than taking one step: a +// relocate round trip leaves original -> hop-1 clone -> hop-2 clone, and the +// volume actually serving the workload is the LAST hop (each record names it +// as active_lvol_id, resolved transitively by the backend). Stopping at the +// first pairing's target landed every post-round-trip verb -- including the +// policy attach Enable performs -- on the retired middle clone, on the wrong +// cluster (confirmed live 2026-09-24). Each hop uses that record's own +// target triple, whose cluster/pool/lvol are consistent with each other; +// combining active_lvol_id with ANOTHER record's cluster is exactly the bug +// this walk exists to avoid. An empty ActiveLvolID (backend predating the +// field) stops after the first hop, the old single-step behavior. func resolveToLocalReplica( ctx context.Context, h *lvol.Handle, client *atlascp.Client, ) (*lvol.Handle, *atlascp.Client, error) { - rel, err := client.GetVolumeReplicationRelationship(ctx, h.Handle()) - if err != nil { - if errors.Is(err, errs.ErrNotFound) { + for range 8 { // one hop per past fail-over; capped far above any real chain + rel, err := client.GetVolumeReplicationRelationship(ctx, h.Handle()) + if err != nil { + if errors.Is(err, errs.ErrNotFound) { + return h, client, nil + } + return nil, nil, err + } + if !rel.IsSource { + return h, client, nil + } + target := &lvol.Handle{ClusterID: rel.TargetClusterID, PoolRef: rel.TargetPoolID, VolumeID: rel.TargetLvolID} + targetClient, err := clusters.ReplicationClient(ctx, target.ClusterID) + if err != nil { + return nil, nil, err + } + h, client = target, targetClient + if rel.ActiveLvolID == "" || rel.ActiveLvolID == rel.TargetLvolID { return h, client, nil } - return nil, nil, err - } - if !rel.IsSource { - return h, client, nil - } - target := &lvol.Handle{ClusterID: rel.TargetClusterID, PoolRef: rel.TargetPoolID, VolumeID: rel.TargetLvolID} - targetClient, err := clusters.ReplicationClient(ctx, target.ClusterID) - if err != nil { - return nil, nil, err } - return target, targetClient, nil + return h, client, nil } // EnableVolumeReplication attaches the volume to the policy named by the diff --git a/csi-driver/internal/csi/controller/replication_lifecycle_test.go b/csi-driver/internal/csi/controller/replication_lifecycle_test.go index b1b05b365..106f248fa 100644 --- a/csi-driver/internal/csi/controller/replication_lifecycle_test.go +++ b/csi-driver/internal/csi/controller/replication_lifecycle_test.go @@ -473,6 +473,63 @@ func TestGetVolumeReplicationInfoResolvesToTargetWhenGivenTheSourceSideOfARelati } } +// Regression: 2026-09-24-chained-relationship-resolves-one-hop-short — a +// relocate ROUND TRIP leaves two chained pairings: original -> hop-1 clone +// (cluster B), and hop-1 clone -> hop-2 clone (cluster A, the volume actually +// serving the workload, named by active_lvol_id on every record in the +// chain). Single-step resolution stopped at the FIRST pairing's target -- the +// retired hop-1 clone -- so post-round-trip Replication verbs (and the policy +// attach that Enable performs) landed on a superseded volume on the wrong +// cluster (confirmed live 2026-09-24: the policy stuck to B's clone while A's +// new primary ran unprotected). Resolution must walk hop by hop, using each +// record's own consistent target triple, until the hop whose target IS the +// active volume. +func TestPromoteVolumeResolvesAcrossAChainedRelationshipToTheActiveVolume(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + mock.volumes[testReplTargetVolumeID] = &mockVolume{ + UUID: testReplTargetVolumeID, Name: "repl-vol-hop1-clone", Size: 1 << 30, + } + mock.volumes[testReplActiveVolumeID] = &mockVolume{ + UUID: testReplActiveVolumeID, Name: "repl-vol-active", Size: 1 << 30, + } + mock.replicationRelationship[testReplVolumeID] = map[string]any{ + "replication_id": "aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeee1", + "direction": "to_target", + "mode": "failover", + "state": "failed_over", + "is_source": true, + "source_cluster_id": sanityClusterID, "source_lvol_id": testReplVolumeID, + "target_cluster_id": sanityClusterID, "target_pool_id": sanityPoolUUID, "target_lvol_id": testReplTargetVolumeID, + "target_nqn": "nqn.test", "target_ns_id": 1, + "active": "target", "active_lvol_id": testReplActiveVolumeID, + } + mock.replicationRelationship[testReplTargetVolumeID] = map[string]any{ + "replication_id": "aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeee2", + "direction": "to_target", + "mode": "failover", + "state": "failed_over", + "is_source": true, + "source_cluster_id": sanityClusterID, "source_lvol_id": testReplTargetVolumeID, + "target_cluster_id": sanityClusterID, "target_pool_id": sanityPoolUUID, "target_lvol_id": testReplActiveVolumeID, + "target_nqn": "nqn.test", "target_ns_id": 1, + "active": "target", "active_lvol_id": testReplActiveVolumeID, + } + delete(mock.volumes, testReplVolumeID) + + _, err := cs.PromoteVolume(context.Background(), &replication.PromoteVolumeRequest{ + VolumeId: testReplVolID, Force: true, + }) + if err != nil { + t.Fatal(err) + } + if mock.lastFailoverVolumeID != testReplActiveVolumeID { + t.Errorf("failover landed on volume %q, want the chain's ACTIVE volume %q, not the retired middle hop", + mock.lastFailoverVolumeID, testReplActiveVolumeID) + } +} + func TestResyncVolumeBackendFailureIsUnavailable(t *testing.T) { mock := newMockSBCLI() defer mock.Close() diff --git a/csi-driver/internal/csi/node/stage.go b/csi-driver/internal/csi/node/stage.go index 7375ea401..0b52f0d20 100644 --- a/csi-driver/internal/csi/node/stage.go +++ b/csi-driver/internal/csi/node/stage.go @@ -545,7 +545,7 @@ func (ns *Server) refreshVolumeContext(ctx context.Context, volumeID string, vc // The source volume was deleted by a migration with --delete-source. // The replication relationship survives it and names the active // volume on the target cluster, which is what this redirects to. - connInfo = ns.redirectToActiveVolume(ctx, sbcClient, spdkVol.VolumeID, volumeID, vc) + connInfo = redirectToActiveVolume(ctx, sbcClient, spdkVol.VolumeID, volumeID, vc) } if connInfo == nil { klog.Warningf("failed to fetch volume connection info for %s: %v", volumeID, infoErr) diff --git a/csi-driver/internal/csi/node/stats.go b/csi-driver/internal/csi/node/stats.go index 7f7d6c80f..6185e9257 100644 --- a/csi-driver/internal/csi/node/stats.go +++ b/csi-driver/internal/csi/node/stats.go @@ -99,48 +99,76 @@ func (ns *Server) NodeGetVolumeStats( }, nil } +// clusterClientFor resolves a cluster and pool to a control-plane client. A +// package variable, so redirectToActiveVolume's chain walk is testable with +// fake clients instead of a live secret file. +var clusterClientFor = func(ctx context.Context, clusterID, poolID string) (controlplane.ClusterAPI, error) { + return clusters.Client(ctx, clusterID, poolID) +} + // redirectToActiveVolume is called when VolumeInfo returns ErrVolumeNotFound for -// the source volume, typically after a migration with --delete-source removed it. -// It queries the replication relationship on the source cluster (which survives -// volume deletion) to find the active volume on the target cluster, then fetches -// connection info from the target. Returns nil if redirection is not possible. -func (ns *Server) redirectToActiveVolume( +// the source volume, typically after a migration with --delete-source removed it +// or after a fail-over retired it. It follows the replication relationship +// (which survives volume deletion) to the volume actually serving the data and +// fetches connection info from there. Returns nil if redirection is not possible. +// +// The walk follows CHAINED pairings hop by hop: a relocate round trip leaves +// original -> hop-1 clone -> hop-2 clone, where only the last hop is live +// (each record names it as active_lvol_id, resolved transitively by the +// backend). Each hop uses that record's own target triple -- its cluster, +// pool, and lvol are consistent with EACH OTHER, while active_lvol_id may +// live on an entirely different cluster than the record's target fields +// describe. Pairing the first record's active_lvol_id with its target +// cluster asked cluster B for a volume living on cluster A, fell back to a +// stale stashed context, and timed the mount out (confirmed live 2026-09-24, +// relocate M-02's round trip). +func redirectToActiveVolume( ctx context.Context, srcClient controlplane.ClusterAPI, srcLvolID, volumeID string, vc map[string]string, ) map[string]string { - rel, err := srcClient.GetRelationship(ctx, srcLvolID) - if err != nil || rel == nil { - klog.Warningf("replication relationship lookup failed for deleted volume %s: %v", volumeID, err) - return nil - } - activeLvolID := rel.ActiveLvolID - targetClusterID := rel.TargetClusterID - targetPoolID := rel.TargetPoolID - if activeLvolID == "" || targetClusterID == "" || targetPoolID == "" { - klog.Warningf("relationship for %s has incomplete target info (cluster=%s pool=%s active=%s)", - volumeID, targetClusterID, targetPoolID, activeLvolID) - return nil - } - tgtClient, err := clusters.Client(ctx, targetClusterID, targetPoolID) - if err != nil { - klog.Warningf("target cluster %s not in secret file for deleted volume %s: %v", - targetClusterID, volumeID, err) - return nil - } - connInfo, err := tgtClient.VolumeInfo(ctx, activeLvolID, vc["hostNQN"]) - if err != nil { - klog.Warningf("failed to fetch connection info from target cluster %s for volume %s: %v", - targetClusterID, activeLvolID, err) - return nil + client, lvolID := srcClient, srcLvolID + for range 8 { // one hop per past fail-over; capped far above any real chain + rel, err := client.GetRelationship(ctx, lvolID) + if err != nil || rel == nil { + klog.Warningf("replication relationship lookup failed for deleted volume %s (at hop %s): %v", + volumeID, lvolID, err) + return nil + } + if rel.TargetLvolID == "" || rel.TargetClusterID == "" || rel.TargetPoolID == "" { + klog.Warningf("relationship for %s has incomplete target info (cluster=%s pool=%s lvol=%s)", + volumeID, rel.TargetClusterID, rel.TargetPoolID, rel.TargetLvolID) + return nil + } + tgtClient, err := clusterClientFor(ctx, rel.TargetClusterID, rel.TargetPoolID) + if err != nil { + klog.Warningf("target cluster %s not in secret file for deleted volume %s: %v", + rel.TargetClusterID, volumeID, err) + return nil + } + if rel.ActiveLvolID != "" && rel.ActiveLvolID != rel.TargetLvolID { + // This pairing's target was itself superseded by a later + // fail-over; keep walking from it toward the active volume. + client, lvolID = tgtClient, rel.TargetLvolID + continue + } + connInfo, err := tgtClient.VolumeInfo(ctx, rel.TargetLvolID, vc["hostNQN"]) + if err != nil { + klog.Warningf("failed to fetch connection info from target cluster %s for volume %s: %v", + rel.TargetClusterID, rel.TargetLvolID, err) + return nil + } + klog.Infof("redirected deleted volume %s → active volume %s on cluster %s", + volumeID, rel.TargetLvolID, rel.TargetClusterID) + // Override cluster_id and poolID so the initiator uses the active + // volume's cluster for any subsequent API calls. Without this the + // initiator inherits the source cluster_id from vc and fails looking + // up the volume there. + connInfo[csicommon.ParamClusterID] = rel.TargetClusterID + connInfo["poolID"] = rel.TargetPoolID + return connInfo } - klog.Infof("redirected deleted volume %s → active volume %s on cluster %s", - volumeID, activeLvolID, targetClusterID) - // Override cluster_id and poolID so the initiator uses the target cluster - // for any subsequent API calls. Without this the initiator inherits the - // source cluster_id from vc and fails looking up the target volume there. - connInfo[csicommon.ParamClusterID] = targetClusterID - connInfo["poolID"] = targetPoolID - return connInfo + klog.Warningf("replication chain for deleted volume %s did not converge within 8 hops", volumeID) + return nil } diff --git a/csi-driver/internal/csi/node/stats_test.go b/csi-driver/internal/csi/node/stats_test.go new file mode 100644 index 000000000..c1572a709 --- /dev/null +++ b/csi-driver/internal/csi/node/stats_test.go @@ -0,0 +1,159 @@ +// The redirect-to-active-volume walk: how a NodeStage/NodeGetVolumeStats call +// carrying a volume handle whose own lvol record is gone finds the volume that +// actually serves the data, across one fail-over or a whole relocate round +// trip's chain of them. +package node + +import ( + "context" + "errors" + "testing" + + "github.com/simplyblock/csi-driver/internal/controlplane" + csicommon "github.com/simplyblock/csi-driver/internal/csi/common" +) + +// fakeRelationshipAPI stubs exactly the two ClusterAPI calls the redirect +// walk makes; every other method panics via the embedded nil interface, +// which is the point -- the walk must touch nothing else. +type fakeRelationshipAPI struct { + controlplane.ClusterAPI + rels map[string]*controlplane.ReplicationRelationship + conn map[string]map[string]string +} + +func (f *fakeRelationshipAPI) GetRelationship(_ context.Context, lvolID string) (*controlplane.ReplicationRelationship, error) { + rel, ok := f.rels[lvolID] + if !ok { + return nil, errors.New("no replication relationship") + } + return rel, nil +} + +func (f *fakeRelationshipAPI) VolumeInfo(_ context.Context, lvolID, _ string) (map[string]string, error) { + c, ok := f.conn[lvolID] + if !ok { + return nil, controlplane.ErrVolumeNotFound + } + return c, nil +} + +// Regression: 2026-09-24-chained-relationship-resolves-one-hop-short — after +// a relocate ROUND TRIP the chain is original(A) -> hop-1 clone(B) -> hop-2 +// clone(A, the live volume, named by active_lvol_id on every record). The +// redirect used the FIRST record's active_lvol_id but paired it with that +// same record's target cluster/pool -- fields describing a DIFFERENT hop -- +// and asked cluster B for a volume that lives on cluster A ("volume not +// found," confirmed live 2026-09-24), then fell back to the stale stashed +// context and the mount timed out. The walk must follow each record's own +// consistent target triple, hop by hop, until the hop whose target IS the +// active volume. +func TestRedirectToActiveVolumeWalksAChainedRelationship(t *testing.T) { + const ( + clusterA = "aaaaaaaa-0000-0000-0000-000000000001" + clusterB = "bbbbbbbb-0000-0000-0000-000000000001" + poolA = "aaaaaaaa-0000-0000-0000-00000000000a" + poolB = "bbbbbbbb-0000-0000-0000-00000000000b" + original = "11111111-1111-1111-1111-111111111111" + hop1 = "22222222-2222-2222-2222-222222222222" + active = "33333333-3333-3333-3333-333333333333" + ) + + srcClient := &fakeRelationshipAPI{ + rels: map[string]*controlplane.ReplicationRelationship{ + original: { + SourceLvolID: original, TargetLvolID: hop1, + SourceClusterID: clusterA, TargetClusterID: clusterB, TargetPoolID: poolB, + ActiveLvolID: active, + }, + }, + } + clusterBClient := &fakeRelationshipAPI{ + rels: map[string]*controlplane.ReplicationRelationship{ + hop1: { + SourceLvolID: hop1, TargetLvolID: active, + SourceClusterID: clusterB, TargetClusterID: clusterA, TargetPoolID: poolA, + ActiveLvolID: active, + }, + }, + } + clusterAClient := &fakeRelationshipAPI{ + conn: map[string]map[string]string{ + active: {"nqn": "nqn.test:" + active, "ip": "10.0.0.1", "port": "4420"}, + }, + } + + orig := clusterClientFor + defer func() { clusterClientFor = orig }() + clusterClientFor = func(_ context.Context, clusterID, poolID string) (controlplane.ClusterAPI, error) { + switch clusterID + "/" + poolID { + case clusterB + "/" + poolB: + return clusterBClient, nil + case clusterA + "/" + poolA: + return clusterAClient, nil + } + return nil, errors.New("unexpected cluster " + clusterID + "/" + poolID) + } + + connInfo := redirectToActiveVolume(context.Background(), srcClient, original, + clusterA+":"+poolA+":"+original, map[string]string{"hostNQN": "nqn.host"}) + if connInfo == nil { + t.Fatal("redirect returned nil: the walk never reached the active volume") + } + if got := connInfo["nqn"]; got != "nqn.test:"+active { + t.Errorf("connection nqn = %q, want the ACTIVE volume's %q", got, "nqn.test:"+active) + } + if got := connInfo[csicommon.ParamClusterID]; got != clusterA { + t.Errorf("cluster_id = %q, want the active volume's cluster %q, not the first hop's", got, clusterA) + } + if got := connInfo["poolID"]; got != poolA { + t.Errorf("poolID = %q, want the active volume's pool %q", got, poolA) + } +} + +// The single-pairing case the redirect was originally written for (a +// migration with --delete-source, or one fail-over): the first record's +// target IS the active volume, and the walk must behave exactly as the +// one-step redirect always did. +func TestRedirectToActiveVolumeSinglePairingIsUnchanged(t *testing.T) { + const ( + clusterA = "aaaaaaaa-0000-0000-0000-000000000001" + clusterB = "bbbbbbbb-0000-0000-0000-000000000001" + poolB = "bbbbbbbb-0000-0000-0000-00000000000b" + original = "11111111-1111-1111-1111-111111111111" + active = "22222222-2222-2222-2222-222222222222" + ) + + srcClient := &fakeRelationshipAPI{ + rels: map[string]*controlplane.ReplicationRelationship{ + original: { + SourceLvolID: original, TargetLvolID: active, + SourceClusterID: clusterA, TargetClusterID: clusterB, TargetPoolID: poolB, + ActiveLvolID: active, + }, + }, + } + clusterBClient := &fakeRelationshipAPI{ + conn: map[string]map[string]string{ + active: {"nqn": "nqn.test:" + active}, + }, + } + + orig := clusterClientFor + defer func() { clusterClientFor = orig }() + clusterClientFor = func(_ context.Context, clusterID, poolID string) (controlplane.ClusterAPI, error) { + if clusterID == clusterB && poolID == poolB { + return clusterBClient, nil + } + return nil, errors.New("unexpected cluster " + clusterID) + } + + connInfo := redirectToActiveVolume(context.Background(), srcClient, original, + clusterA+":pool:"+original, map[string]string{}) + if connInfo == nil { + t.Fatal("redirect returned nil for the plain single-pairing case") + } + if got := connInfo[csicommon.ParamClusterID]; got != clusterB { + t.Errorf("cluster_id = %q, want %q", got, clusterB) + } +} diff --git a/operator/docs/tests/test-plan-csi-addons-replication.md b/operator/docs/tests/test-plan-csi-addons-replication.md index 752074200..c3368330a 100644 --- a/operator/docs/tests/test-plan-csi-addons-replication.md +++ b/operator/docs/tests/test-plan-csi-addons-replication.md @@ -63,6 +63,8 @@ File: `csi-driver/internal/csi/controller/replication_lifecycle_test.go` | U-51 | Promote while the fail-back's first reverse-replicated snapshot is still in flight: control-plane 409 (converging, retryable), never a stack-traced 500; the same backend message with no replication configured stays 500 | Regression | `sbcli: test_failover_while_the_first_replicated_snapshot_is_in_flight_is_409`, `test_failover_with_no_snapshot_and_no_replication_configured_stays_500` | | U-52 | DeleteVolume carrying a reaped, foreign handle follows the relationship and removes the retired, non-active members (2026-09-24-delete-foreign-handle-leak); the `Relationship.ActiveLvolID` parse in atlas-lib is part of the same fix | Regression | `TestDeleteVolumeRemovesTheRetiredReplicaBehindAForeignHandle`, `atlas-lib: TestClientGetVolumeReplicationRelationshipParsesTheActiveVolume` | | U-53 | DeleteVolume never deletes the pairing's ACTIVE volume through a dead source handle — the workload is running on it | Negative | `TestDeleteVolumeNeverDeletesTheActiveReplicaThroughADeadHandle` | +| U-54 | Resolution walks CHAINED pairings to the active volume: after a relocate round trip, verbs on the original handle land on the last hop's clone, never the retired middle hop (2026-09-24-chained-relationship-resolves-one-hop-short) | Regression | `TestPromoteVolumeResolvesAcrossAChainedRelationshipToTheActiveVolume` | +| U-55 | The node's deleted-volume redirect walks the same chain using each record's own target triple, never pairing active_lvol_id with another hop's cluster; the single-pairing migration case is unchanged | Regression | `node: TestRedirectToActiveVolumeWalksAChainedRelationship`, `TestRedirectToActiveVolumeSinglePairingIsUnchanged` | ### Condition Derivation (design §6.2) @@ -181,7 +183,7 @@ Two live simplyblock clusters with the chart-deployed csi-addons machinery. The | Axis | Values covered | IDs | Not covered | |------------------|------------------------------------------------------------------------|-----------------------------------------------------------------------------|-----------------------------------------------------------| -| Verb lifecycle | enable, disable, info, forced promote, planned promote, demote, resync | U-01, U-02, U-04 … U-10, U-12 … U-17, U-40 … U-53, I-02 … I-04, E-01 … E-04 | — | +| Verb lifecycle | enable, disable, info, forced promote, planned promote, demote, resync | U-01, U-02, U-04 … U-10, U-12 … U-17, U-40 … U-55, I-02 … I-04, E-01 … E-04 | — | | Idempotency | repeat enable, disable, demote; re-drive after restart | U-02, U-05, U-15, I-07 | repeated promote (U-11), repeated resync | | Conditions | healthy, degraded, error, staleness, resyncing, disabled | U-18 … U-22, E-05 | condition behavior across backend upgrade | | Coexistence | slot skip, one-owner refusal, concurrent claim | U-26, U-27, M-02 | migration of an annotated volume onto a VolumeReplication | From ad47cab162224e12bda81c1bc4cbbcffa13d6753 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Thu, 24 Sep 2026 17:27:09 +0100 Subject: [PATCH 151/206] fixed helm linter issue --- .../charts/simplyblock-operator/values.yaml | 86 ++----------------- 1 file changed, 6 insertions(+), 80 deletions(-) diff --git a/helm-charts/charts/simplyblock-operator/values.yaml b/helm-charts/charts/simplyblock-operator/values.yaml index 651b3ad91..a3f6656e4 100644 --- a/helm-charts/charts/simplyblock-operator/values.yaml +++ b/helm-charts/charts/simplyblock-operator/values.yaml @@ -74,93 +74,19 @@ csiSecret: simplybk: secret: -logicalVolume: - pool_name: testing1 - qos_rw_iops: "0" - qos_rw_mbytes: "0" - qos_r_mbytes: "0" - qos_w_mbytes: "0" - max_size: "0" - encryption: "False" - numDataChunks: "1" - numParityChunks: "1" - max_namespace_per_subsys: "1" - tune2fs_reserved_blocks: "0" - fabric: tcp podAnnotations: {} # SimplyBlock Daemonset, Deployment, Statefulset annotations simplyBlockAnnotations: {} -benchmarks: 0 - -# FIXME: this will not work if there are group of nodes with different AMI types like: AL2, AL2023 -# AL2_x86_64: eth0 -# AL2023_x86_64_STANDARD: ens5 - - -storagenode: - create: false - ifname: eth0 - numDataChunks: "1" - numParityChunks: "1" - spdkImage: - spdkProxyImage: - maxLogicalVolumes: 10 - maxSnapshots: 10 - maxSize: - numPartitions: 1 - isolateCores: false - journalManager: - dataNic: - pciAllowed: - pciBlocked: - deviceNames: - format4k: false - socketsToUse: - nodesPerSocket: - deviceModel: - sizeRange: - coresPercentage: - ubuntuHost: false - enableCpuTopology: false - enableDevicePlugin: true - openShiftCluster: false - reservedSystemCpu: - multiCluster: - enable: false - clusters: - - cluster_id: - secret: - workers: - daemonsets: - - name: simplyblock-storage-node-ds - appLabel: storage-node - nodeSelector: - key: io.simplyblock.node-type - value: simplyblock-storage-plane - tolerations: - create: false - list: - - operator: Exists - effect: - key: - value: - - name: simplyblock-storage-node-ds-restart - appLabel: storage-node-restart - nodeSelector: - key: io.simplyblock.node-type - value: simplyblock-storage-plane-restart - tolerations: - create: false - list: - - operator: Exists - effect: - key: - value: +multiCluster: + enable: false + clusters: + - cluster_id: + secret: + workers: - deployment: # One of: standalone, managed. # From 71d75f283f38a9f6f9cda253c417764fecb9851f Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Thu, 24 Sep 2026 17:29:45 +0100 Subject: [PATCH 152/206] removed storage-node manifest --- .../templates/storage-node-controller.yaml | 85 -------- .../templates/storage-node.yaml | 204 ------------------ 2 files changed, 289 deletions(-) delete mode 100644 helm-charts/charts/simplyblock-operator/templates/storage-node-controller.yaml delete mode 100644 helm-charts/charts/simplyblock-operator/templates/storage-node.yaml diff --git a/helm-charts/charts/simplyblock-operator/templates/storage-node-controller.yaml b/helm-charts/charts/simplyblock-operator/templates/storage-node-controller.yaml deleted file mode 100644 index 019d80a07..000000000 --- a/helm-charts/charts/simplyblock-operator/templates/storage-node-controller.yaml +++ /dev/null @@ -1,85 +0,0 @@ -{{- if .Values.storagenode.create -}} ---- -apiVersion: v1 -kind: ServiceAccount -metadata: - name: simplyblock-storage-node-service-account - ---- -apiVersion: apps/v1 -kind: Deployment -metadata: - name: simplyblock-storage-node-controller - labels: - app: storage-node-controller -spec: - replicas: 1 - selector: - matchLabels: - app: storage-node-controller - template: - metadata: - labels: - app: storage-node-controller - spec: - serviceAccountName: simplyblock-storage-node-service-account - containers: - - name: storage-node-controller - image: "{{ .Values.image.storageNode.repository }}:{{ .Values.image.storageNode.tag }}" - imagePullPolicy: "{{ .Values.image.storageNode.pullPolicy }}" - env: - - name: SPDKCSI_SECRET - valueFrom: - secretKeyRef: - name: simplyblock-csi-secret - key: secret.json - - name: CLUSTER_CONFIG - valueFrom: - configMapKeyRef: - name: simplyblock-csi-cm - key: config.json - - name: IFNAME - value: "{{- if kindIs "string" .Values.storagenode.ifname -}} - {{ .Values.storagenode.ifname }} - {{- else -}} - {{ join "," .Values.storagenode.ifname }} - {{- end -}}" - - name: MAXSNAP - value: "{{ .Values.storagenode.maxSnapshots }}" - - name: JMPERCENT - value: "3" - - name: NUMPARTITIONS - value: "{{ .Values.storagenode.numPartitions }}" - - name: DISABLEHAJM - value: "false" - - name: ENABLETESTDEVICE - value: "false" - - name: DATANICS - value: "{{ .Values.storagenode.dataNic }}" - - name: SPDKIMAGE - value: "{{ .Values.storagenode.spdkImage }}" - - name: SPDKPROXYIMAGE - value: "{{ .Values.storagenode.spdkProxyImage }}" - - name: FORMAT4K - value: "{{ .Values.storagenode.format4k }}" - - name: NAMESPACE - valueFrom: - fieldRef: - fieldPath: metadata.namespace - volumeMounts: - - name: cluster-config - mountPath: /etc/simplyblock/config - readOnly: true - - name: cluster-secret - mountPath: /etc/simplyblock/secret - readOnly: true - volumes: - - name: cluster-config - configMap: - name: simplyblock-clusters - - name: cluster-secret - secret: - secretName: simplyblock-csi-secret-v2 - restartPolicy: Always - -{{- end -}} diff --git a/helm-charts/charts/simplyblock-operator/templates/storage-node.yaml b/helm-charts/charts/simplyblock-operator/templates/storage-node.yaml deleted file mode 100644 index 0fd141cb0..000000000 --- a/helm-charts/charts/simplyblock-operator/templates/storage-node.yaml +++ /dev/null @@ -1,204 +0,0 @@ -{{- if .Values.storagenode.create -}} ---- -apiVersion: v1 -kind: ServiceAccount -metadata: - name: simplyblock-storage-node-sa - ---- -kind: ClusterRole -apiVersion: rbac.authorization.k8s.io/v1 -metadata: - name: simplyblock-storage-node-role -rules: -- apiGroups: [""] - resources: ["pods", "namespaces", "pods/exec"] - verbs: ["list", "get", "create", "delete", "watch"] -- apiGroups: ["apps"] - resources: ["deployments"] - verbs: ["create", "delete"] -- apiGroups: ["batch"] - resources: ["jobs"] - verbs: ["create", "delete", "get", "list", "watch"] -- apiGroups: [""] - resources: ["nodes"] - verbs: ["get", "list", "watch"] - -{{- if .Values.storagenode.openShiftCluster }} -- apiGroups: ["machineconfiguration.openshift.io"] - resources: ["machineconfigs", "machineconfigpools", "kubeletconfigs"] - verbs: ["list", "get", "create", "update", "patch", "watch"] -- apiGroups: [""] - resources: ["nodes"] - verbs: ["list", "get", "update", "patch", "watch"] -{{- end }} - ---- -kind: ClusterRoleBinding -apiVersion: rbac.authorization.k8s.io/v1 -metadata: - name: simplyblock-pods-list-sn -subjects: -- kind: ServiceAccount - name: simplyblock-storage-node-sa - namespace: {{ .Release.Namespace }} -roleRef: - kind: ClusterRole - name: simplyblock-storage-node-role - apiGroup: rbac.authorization.k8s.io - -{{- range .Values.storagenode.daemonsets }} ---- -apiVersion: apps/v1 -kind: DaemonSet -metadata: - name: {{ .name }} - annotations: - {{- with $.Values.simplyBlockAnnotations }} - {{- toYaml . | nindent 4 }} - {{- end }} -spec: - selector: - matchLabels: - app: {{ .appLabel }} - updateStrategy: - type: RollingUpdate - rollingUpdate: - maxUnavailable: 1 - template: - metadata: - labels: - app: {{ .appLabel }} - annotations: - {{- range $key, $value := $.Values.podAnnotations }} - {{ $key }}: {{ $value | quote }} - {{- end }} - spec: - serviceAccountName: simplyblock-storage-node-sa - nodeSelector: - {{ .nodeSelector.key }}: {{ .nodeSelector.value }} - volumes: - - name: dev-vol - hostPath: - path: /dev - - name: etc-simplyblock - hostPath: - path: /var/simplyblock - - name: host-sys - hostPath: - path: /sys - - name: host-mnt - hostPath: - path: /mnt - - name: host-modules - hostPath: - path: /lib/modules - {{- include "simplyblock.tlsVolume" (dict "ctx" $ "secret" "simplyblock-storage-node-api-tls") | nindent 8 }} - hostNetwork: true - {{- if .tolerations.create }} - tolerations: - {{- range .tolerations.list }} - - operator: {{ .operator | quote }} - {{- if .effect }} - effect: {{ .effect | quote }} - {{- end }} - {{- if .key }} - key: {{ .key | quote }} - {{- end }} - {{- if .value }} - value: {{ .value | quote }} - {{- end }} - {{- end }} - {{- end }} - initContainers: - - name: s-node-api-config-generator - image: "{{ $.Values.image.simplyblock.repository }}:{{ $.Values.image.simplyblock.tag }}" - imagePullPolicy: "{{ $.Values.image.simplyblock.pullPolicy }}" - env: - {{- include "simplyblock.tlsEnv" $ | nindent 8 }} - - name: HOSTNAME - valueFrom: - fieldRef: - fieldPath: spec.nodeName - command: - - "python" - - "simplyblock_web/node_configure.py" - - "--max-lvol={{ $.Values.storagenode.maxLogicalVolumes }}" - - "--max-size={{ $.Values.storagenode.maxSize }}" - {{- with (default list $.Values.storagenode.pciAllowed) }} - - "--pci-allowed={{ join "," . }}" - {{- end }} - {{- with (default list $.Values.storagenode.pciBlocked) }} - - "--pci-blocked={{ join "," . }}" - {{- end }} - {{- with (default list $.Values.storagenode.deviceNames) }} - - "--nvme-devices={{ join "," . }}" - {{- end }} - {{- if $.Values.storagenode.socketsToUse }} - - "--sockets-to-use={{ $.Values.storagenode.socketsToUse }}" - {{- end }} - {{- if $.Values.storagenode.nodesPerSocket }} - - "--nodes-per-socket={{ $.Values.storagenode.nodesPerSocket }}" - {{- end }} - {{- if $.Values.storagenode.deviceModel }} - - "--device-model={{ $.Values.storagenode.deviceModel }}" - {{- end }} - {{- if $.Values.storagenode.sizeRange }} - - "--size-range={{ $.Values.storagenode.sizeRange }}" - {{- end }} - {{- if $.Values.storagenode.coresPercentage }} - - "--cores-percentage={{ $.Values.storagenode.coresPercentage }}" - {{- end }} - volumeMounts: - - name: etc-simplyblock - mountPath: /etc/simplyblock - - name: host-modules - mountPath: /lib/modules - readOnly: true - - name: host-mnt - mountPath: /mnt - {{- include "simplyblock.tlsVolumeMount" $ | nindent 10 }} - securityContext: - privileged: true - containers: - - name: s-node-api-container - image: "{{ $.Values.image.simplyblock.repository }}:{{ $.Values.image.simplyblock.tag }}" - imagePullPolicy: "{{ $.Values.image.simplyblock.pullPolicy }}" - command: ["python", "simplyblock_web/node_webapp.py", "storage_node_k8s"] - env: - {{- include "simplyblock.tlsEnv" $ | nindent 8 }} - - name: CORE_ISOLATION - value: "{{ $.Values.storagenode.isolateCores }}" - - name: UBUNTU_HOST - value: "{{ $.Values.storagenode.ubuntuHost }}" - - name: OPENSHIFT_CLUSTER - value: "{{ $.Values.storagenode.openShiftCluster }}" - - name: CPU_TOPOLOGY_ENABLED - value: "{{ $.Values.storagenode.enableCpuTopology }}" - {{- if $.Values.storagenode.reservedSystemCpu }} - - name: RESERVED_SYSTEM_CPUS - value: "{{ $.Values.storagenode.reservedSystemCpu }}" - {{- end }} - - name: HOSTNAME - valueFrom: - fieldRef: - fieldPath: spec.nodeName - securityContext: - privileged: true - readinessProbe: - httpGet: - path: /snode/check - port: 5000 - initialDelaySeconds: 10 - periodSeconds: 5 - volumeMounts: - - name: dev-vol - mountPath: /dev - - name: etc-simplyblock - mountPath: /etc/simplyblock - - name: host-sys - mountPath: /sys - -{{- end }} - -{{- end -}} From 43523f9daa844713f59fc1da3807f504a9a38f1b Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Thu, 24 Sep 2026 17:31:33 +0100 Subject: [PATCH 153/206] fixed manifest numa-resource-plugin.yaml --- .../templates/numa-resource-plugin.yaml | 21 ------------------- 1 file changed, 21 deletions(-) diff --git a/helm-charts/charts/simplyblock-operator/templates/numa-resource-plugin.yaml b/helm-charts/charts/simplyblock-operator/templates/numa-resource-plugin.yaml index 99e7d0856..9e162b9c6 100644 --- a/helm-charts/charts/simplyblock-operator/templates/numa-resource-plugin.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/numa-resource-plugin.yaml @@ -1,5 +1,3 @@ -{{- $storagenodeEnabled := and .Values.storagenode.create .Values.storagenode.enableCpuTopology .Values.storagenode.enableDevicePlugin -}} -{{- if or $storagenodeEnabled true -}} --- apiVersion: v1 kind: ServiceAccount @@ -37,22 +35,6 @@ spec: spec: serviceAccountName: simplyblock-numa-resource-plugin priorityClassName: system-node-critical - {{- $dsList := .Values.storagenode.daemonsets | default (list) -}} - {{- if and $storagenodeEnabled (gt (len $dsList) 0) }} - affinity: - nodeAffinity: - requiredDuringSchedulingIgnoredDuringExecution: - nodeSelectorTerms: - - matchExpressions: - - key: {{ (index $dsList 0).nodeSelector.key | quote }} - operator: In - values: - {{- range $i, $ds := $dsList }} - {{- if and $ds.nodeSelector $ds.nodeSelector.value }} - - {{ $ds.nodeSelector.value | quote }} - {{- end }} - {{- end }} - {{- else if true }} affinity: nodeAffinity: requiredDuringSchedulingIgnoredDuringExecution: @@ -64,7 +46,6 @@ spec: - matchExpressions: - key: io.simplyblock.storagenodeset operator: Exists - {{- end }} tolerations: # Run on all nodes including control plane - operator: Exists @@ -123,5 +104,3 @@ spec: type: RollingUpdate rollingUpdate: maxUnavailable: 1 - -{{- end -}} From 414f0b3ddf6cfd3e55099c72563ada372b1bbeff Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Thu, 24 Sep 2026 17:59:27 +0100 Subject: [PATCH 154/206] csi-driver: attach the class policy to the newly created primary after promote --- .../internal/csi/controller/replication.go | 43 +++++++++ .../controller/replication_lifecycle_test.go | 90 +++++++++++++++++++ csi-driver/internal/csi/node/stats_test.go | 4 +- .../tests/test-plan-csi-addons-replication.md | 55 ++++++------ 4 files changed, 164 insertions(+), 28 deletions(-) diff --git a/csi-driver/internal/csi/controller/replication.go b/csi-driver/internal/csi/controller/replication.go index 90ffd2e93..7a75b0fea 100644 --- a/csi-driver/internal/csi/controller/replication.go +++ b/csi-driver/internal/csi/controller/replication.go @@ -255,9 +255,52 @@ func (cs *Server) PromoteVolume( if err := client.PromoteVolume(ctx, h.Handle(), req.GetForce()); err != nil { return nil, classifyPromoteVolumeError(err) } + if err := attachPolicyToPromotedVolume(ctx, req); err != nil { + return nil, err + } return &replication.PromoteVolumeResponse{}, nil } +// attachPolicyToPromotedVolume re-resolves the request's handle AFTER a +// successful promote and attaches the VolumeReplicationClass's policy to the +// volume the walk now reaches -- the clone the promote just created. The +// vendored csi-addons controller calls EnableVolumeReplication only at VR +// creation, which on a relocate is BEFORE that clone exists: Enable lands on +// the pre-promote side and no later reconcile re-issues it (confirmed live +// 2026-09-24, twice -- the new primary ran unprotected, policy=NONE, while +// the retired clone kept the policy). Promote is the one call that knows the +// new primary exists, and the class parameters ride on every RPC, so it +// finishes the job itself. No policy parameter means nothing to attach: +// promote outside a policy stays exactly what it was. +func attachPolicyToPromotedVolume(ctx context.Context, req *replication.PromoteVolumeRequest) error { + policyID := req.GetParameters()[replicationPolicyParam] + if policyID == "" { + return nil + } + h, err := csicommon.ParseVolumeHandle(volumeIDFrom(req)) + if err != nil { + return status.Error(codes.InvalidArgument, err.Error()) + } + client, err := clusters.ReplicationClient(ctx, h.ClusterID) + if err != nil { + return status.Error(codes.Unavailable, err.Error()) + } + h, client, err = resolveToLocalReplica(ctx, h, client) + if err != nil { + return status.Error(codes.Unavailable, err.Error()) + } + if err := client.EnableVolumeReplication(ctx, h.Handle(), policyID); err != nil { + if errors.Is(err, errs.ErrNotFound) { + // The backend has not registered the promoted clone yet; the + // controller re-drives promote (idempotent) and the attach lands + // on a later pass. + return nil + } + return classifyEnableVolumeReplicationError(err) + } + return nil +} + // DemoteVolume fences the source and confirms the last write replicated // (P0-3) -- the lossless half of a planned swap. Synchronous and // non-blocking: it never waits out the backend's own convergence loop. diff --git a/csi-driver/internal/csi/controller/replication_lifecycle_test.go b/csi-driver/internal/csi/controller/replication_lifecycle_test.go index 106f248fa..6fc90800e 100644 --- a/csi-driver/internal/csi/controller/replication_lifecycle_test.go +++ b/csi-driver/internal/csi/controller/replication_lifecycle_test.go @@ -530,6 +530,96 @@ func TestPromoteVolumeResolvesAcrossAChainedRelationshipToTheActiveVolume(t *tes } } +// Regression: 2026-09-24-promote-leaves-new-primary-unprotected — the +// vendored csi-addons controller calls EnableVolumeReplication only at VR +// creation, which on a relocate happens BEFORE the promote creates the new +// primary clone: Enable lands on the pre-promote side and nothing ever +// attaches the policy to the volume the workload actually ends up on +// (confirmed live 2026-09-24, twice: the policy stayed on the retired clone +// while the new primary ran with policy=NONE, do_replicate=False, and even a +// forced reconcile-all after a manager restart issued no further Enable). +// The VolumeReplicationClass parameters ride on EVERY replication RPC, +// promote included, so PromoteVolume itself must re-resolve after a +// successful promote -- the walk now reaches the newly created clone -- and +// attach the class's policy there. +func TestPromoteVolumeAttachesThePolicyToTheNewlyActiveVolume(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + mock.volumes[testReplTargetVolumeID] = &mockVolume{ + UUID: testReplTargetVolumeID, Name: "repl-vol-hop1-clone", Size: 1 << 30, + } + mock.volumes[testReplActiveVolumeID] = &mockVolume{ + UUID: testReplActiveVolumeID, Name: "repl-vol-new-primary", Size: 1 << 30, + } + mock.replicationRelationship[testReplVolumeID] = map[string]any{ + "replication_id": "aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeee1", + "direction": "to_target", + "mode": "failover", + "state": "failed_over", + "is_source": true, + "source_cluster_id": sanityClusterID, "source_lvol_id": testReplVolumeID, + "target_cluster_id": sanityClusterID, "target_pool_id": sanityPoolUUID, "target_lvol_id": testReplTargetVolumeID, + "target_nqn": "nqn.test", "target_ns_id": 1, + "active": "target", "active_lvol_id": testReplActiveVolumeID, + } + mock.replicationRelationship[testReplTargetVolumeID] = map[string]any{ + "replication_id": "aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeee2", + "direction": "to_target", + "mode": "failover", + "state": "failed_over", + "is_source": true, + "source_cluster_id": sanityClusterID, "source_lvol_id": testReplTargetVolumeID, + "target_cluster_id": sanityClusterID, "target_pool_id": sanityPoolUUID, "target_lvol_id": testReplActiveVolumeID, + "target_nqn": "nqn.test", "target_ns_id": 1, + "active": "target", "active_lvol_id": testReplActiveVolumeID, + } + delete(mock.volumes, testReplVolumeID) + + _, err := cs.PromoteVolume(context.Background(), &replication.PromoteVolumeRequest{ + VolumeId: testReplVolID, + Force: true, + Parameters: map[string]string{replicationPolicyParam: testReplPolicyID}, + }) + if err != nil { + t.Fatal(err) + } + if got := mock.volumes[testReplActiveVolumeID].ReplicationPolicyID; got != testReplPolicyID { + t.Errorf("new primary's policy = %q, want %q attached by the promote itself", got, testReplPolicyID) + } +} + +// Without the class parameter there is nothing to attach, and promote must +// stay exactly what it was -- some callers drive promote outside any policy. +func TestPromoteVolumeWithoutThePolicyParameterAttachesNothing(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + mock.volumes[testReplTargetVolumeID] = &mockVolume{ + UUID: testReplTargetVolumeID, Name: "repl-vol-target", Size: 1 << 30, + } + mock.replicationRelationship[testReplVolumeID] = map[string]any{ + "replication_id": "aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee", + "direction": "to_target", + "mode": "failover", + "state": "failed_over", + "is_source": true, + "source_cluster_id": sanityClusterID, "source_lvol_id": testReplVolumeID, + "target_cluster_id": sanityClusterID, "target_pool_id": sanityPoolUUID, "target_lvol_id": testReplTargetVolumeID, + "target_nqn": "nqn.test", "target_ns_id": 1, + } + + _, err := cs.PromoteVolume(context.Background(), &replication.PromoteVolumeRequest{ + VolumeId: testReplVolID, Force: true, + }) + if err != nil { + t.Fatal(err) + } + if got := mock.volumes[testReplTargetVolumeID].ReplicationPolicyID; got != "" { + t.Errorf("policy = %q attached with no class parameter, want none", got) + } +} + func TestResyncVolumeBackendFailureIsUnavailable(t *testing.T) { mock := newMockSBCLI() defer mock.Close() diff --git a/csi-driver/internal/csi/node/stats_test.go b/csi-driver/internal/csi/node/stats_test.go index c1572a709..0a78fa2b7 100644 --- a/csi-driver/internal/csi/node/stats_test.go +++ b/csi-driver/internal/csi/node/stats_test.go @@ -22,7 +22,9 @@ type fakeRelationshipAPI struct { conn map[string]map[string]string } -func (f *fakeRelationshipAPI) GetRelationship(_ context.Context, lvolID string) (*controlplane.ReplicationRelationship, error) { +func (f *fakeRelationshipAPI) GetRelationship( + _ context.Context, lvolID string, +) (*controlplane.ReplicationRelationship, error) { rel, ok := f.rels[lvolID] if !ok { return nil, errors.New("no replication relationship") diff --git a/operator/docs/tests/test-plan-csi-addons-replication.md b/operator/docs/tests/test-plan-csi-addons-replication.md index c3368330a..59a29f15b 100644 --- a/operator/docs/tests/test-plan-csi-addons-replication.md +++ b/operator/docs/tests/test-plan-csi-addons-replication.md @@ -39,32 +39,33 @@ File: `csi-driver/internal/csi/controller/replication_test.go` (planned) File: `csi-driver/internal/csi/controller/replication_lifecycle_test.go` -| # | Scenario | Type | Test | -|------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| U-10 | Forced promote: the failover endpoint is called without `planned=true`, ignoring demote state entirely | Positive | `TestPromoteVolumeForced` | -| U-11 | Forced promote repeated after completion: success without a second failover (backend idempotency honored) | Boundary | — | -| U-12 | Planned promote sends `planned=true` on the failover call | Positive | `TestPromoteVolumePlannedSendsThePlannedFlag` | -| U-13 | Planned promote with no demote ever requested: `FAILED_PRECONDITION`, the one case meant to let the vendored controller's own force-escalation take over | Negative | `TestPromoteVolumePlannedWithNoDemoteIsFailedPrecondition` | -| U-14 | Demote: the demote endpoint is called, and the RPC succeeds only once the backend confirms the final snapshot landed | Positive | `TestDemoteVolumeDone` | -| U-15 | Demote repeated on a demoted volume: success (idempotency) | Boundary | `sbcli: test_demote_is_idempotent_once_done` (backend tier; the driver's `TestDemoteVolumeDone` exercises the same success path) | -| U-16 | Resync: the failback endpoint is called, forwarding the `sourceClusterID` class parameter when given | Positive | `TestResyncVolume`, `TestResyncVolumeSendsTheSourceClusterParameter` | -| U-17 | Resync's `ready` field reflects lag against budget: false while lag exceeds it, true once caught up | Boundary | `TestResyncVolumeNotReadyWhileLagExceedsBudget`, `TestResyncVolumeReadyReflectsLag` | -| U-40 | Planned promote while demote is still converging: `ABORTED` (retryable) — never `FAILED_PRECONDITION`, which the vendored controller auto-escalates to a forced, lossy promote inline with no wait-and-retry grace period of its own | Negative | `TestPromoteVolumePlannedWhileDemoteConvergingIsAborted` | -| U-41 | Demote still converging: `ABORTED` (retryable), non-blocking — the RPC never waits out the backend's own convergence loop | Boundary | `TestDemoteVolumeNotYetDoneIsAborted` | -| U-42 | `demote_lvol` fences the source strictly before triggering the final snapshot, never after (a write landing in the gap would be silently lost) | Positive | `sbcli: test_demote_fences_before_triggering_the_final_snapshot` | -| U-43 | `demote_lvol` re-invoked while pending: checks the marker only, never re-fences or re-triggers | Boundary | `sbcli: test_demote_does_not_refence_or_retrigger_once_pending` | -| U-44 | `demote_lvol` completes once the triggered snapshot carries the replicated marker | Positive | `sbcli: test_demote_completes_once_the_snapshot_carries_the_replicated_marker` | -| U-45 | `demote_lvol` surfaces a snapshot-creation failure without recording pending state | Negative | `sbcli: test_demote_surfaces_a_snapshot_creation_failure` | -| U-46 | The `failover` route's three-way planned-gate branch: demoted proceeds, converging is 409, no relationship is 412 | Positive/Negative | `sbcli: test_planned_failover_proceeds_once_demoted`, `test_planned_failover_while_demote_is_converging_is_409`, `test_planned_failover_without_any_demote_is_412` | -| U-47 | Unplanned failover ignores demote state, unchanged from before P0-3 | Regression | `sbcli: test_unplanned_failover_ignores_demote_state` | -| U-48 | Every verb given the SOURCE side of an existing relationship resolves to the local replica (`TargetLvolId`) before acting — Ramen's S3-restore hands the destination cluster the original source's volumeHandle verbatim | Positive | `TestEnableVolumeReplicationResolvesToTargetWhenGivenTheSourceSideOfARelationship`, `TestDisableVolumeReplicationResolvesToTargetWhenGivenTheSourceSideOfARelationship`, `TestPromoteVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship`, `TestDemoteVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship` | -| U-49 | Resync resolves the foreign source handle even after the source lvol record itself was reaped (2026-09-24-resync-foreign-handle-404: the only verb without the resolution 404ed on every reconcile and stalled the relocate back) | Regression | `TestResyncVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship` | -| U-50 | The standalone info read resolves the foreign source handle likewise, so lastSyncTime keeps flowing after the source record is reaped (2026-09-24-resync-foreign-handle-404) | Regression | `TestGetVolumeReplicationInfoResolvesToTargetWhenGivenTheSourceSideOfARelationship` | -| U-51 | Promote while the fail-back's first reverse-replicated snapshot is still in flight: control-plane 409 (converging, retryable), never a stack-traced 500; the same backend message with no replication configured stays 500 | Regression | `sbcli: test_failover_while_the_first_replicated_snapshot_is_in_flight_is_409`, `test_failover_with_no_snapshot_and_no_replication_configured_stays_500` | -| U-52 | DeleteVolume carrying a reaped, foreign handle follows the relationship and removes the retired, non-active members (2026-09-24-delete-foreign-handle-leak); the `Relationship.ActiveLvolID` parse in atlas-lib is part of the same fix | Regression | `TestDeleteVolumeRemovesTheRetiredReplicaBehindAForeignHandle`, `atlas-lib: TestClientGetVolumeReplicationRelationshipParsesTheActiveVolume` | -| U-53 | DeleteVolume never deletes the pairing's ACTIVE volume through a dead source handle — the workload is running on it | Negative | `TestDeleteVolumeNeverDeletesTheActiveReplicaThroughADeadHandle` | -| U-54 | Resolution walks CHAINED pairings to the active volume: after a relocate round trip, verbs on the original handle land on the last hop's clone, never the retired middle hop (2026-09-24-chained-relationship-resolves-one-hop-short) | Regression | `TestPromoteVolumeResolvesAcrossAChainedRelationshipToTheActiveVolume` | -| U-55 | The node's deleted-volume redirect walks the same chain using each record's own target triple, never pairing active_lvol_id with another hop's cluster; the single-pairing migration case is unchanged | Regression | `node: TestRedirectToActiveVolumeWalksAChainedRelationship`, `TestRedirectToActiveVolumeSinglePairingIsUnchanged` | +| # | Scenario | Type | Test | +|------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| U-10 | Forced promote: the failover endpoint is called without `planned=true`, ignoring demote state entirely | Positive | `TestPromoteVolumeForced` | +| U-11 | Forced promote repeated after completion: success without a second failover (backend idempotency honored) | Boundary | — | +| U-12 | Planned promote sends `planned=true` on the failover call | Positive | `TestPromoteVolumePlannedSendsThePlannedFlag` | +| U-13 | Planned promote with no demote ever requested: `FAILED_PRECONDITION`, the one case meant to let the vendored controller's own force-escalation take over | Negative | `TestPromoteVolumePlannedWithNoDemoteIsFailedPrecondition` | +| U-14 | Demote: the demote endpoint is called, and the RPC succeeds only once the backend confirms the final snapshot landed | Positive | `TestDemoteVolumeDone` | +| U-15 | Demote repeated on a demoted volume: success (idempotency) | Boundary | `sbcli: test_demote_is_idempotent_once_done` (backend tier; the driver's `TestDemoteVolumeDone` exercises the same success path) | +| U-16 | Resync: the failback endpoint is called, forwarding the `sourceClusterID` class parameter when given | Positive | `TestResyncVolume`, `TestResyncVolumeSendsTheSourceClusterParameter` | +| U-17 | Resync's `ready` field reflects lag against budget: false while lag exceeds it, true once caught up | Boundary | `TestResyncVolumeNotReadyWhileLagExceedsBudget`, `TestResyncVolumeReadyReflectsLag` | +| U-40 | Planned promote while demote is still converging: `ABORTED` (retryable) — never `FAILED_PRECONDITION`, which the vendored controller auto-escalates to a forced, lossy promote inline with no wait-and-retry grace period of its own | Negative | `TestPromoteVolumePlannedWhileDemoteConvergingIsAborted` | +| U-41 | Demote still converging: `ABORTED` (retryable), non-blocking — the RPC never waits out the backend's own convergence loop | Boundary | `TestDemoteVolumeNotYetDoneIsAborted` | +| U-42 | `demote_lvol` fences the source strictly before triggering the final snapshot, never after (a write landing in the gap would be silently lost) | Positive | `sbcli: test_demote_fences_before_triggering_the_final_snapshot` | +| U-43 | `demote_lvol` re-invoked while pending: checks the marker only, never re-fences or re-triggers | Boundary | `sbcli: test_demote_does_not_refence_or_retrigger_once_pending` | +| U-44 | `demote_lvol` completes once the triggered snapshot carries the replicated marker | Positive | `sbcli: test_demote_completes_once_the_snapshot_carries_the_replicated_marker` | +| U-45 | `demote_lvol` surfaces a snapshot-creation failure without recording pending state | Negative | `sbcli: test_demote_surfaces_a_snapshot_creation_failure` | +| U-46 | The `failover` route's three-way planned-gate branch: demoted proceeds, converging is 409, no relationship is 412 | Positive/Negative | `sbcli: test_planned_failover_proceeds_once_demoted`, `test_planned_failover_while_demote_is_converging_is_409`, `test_planned_failover_without_any_demote_is_412` | +| U-47 | Unplanned failover ignores demote state, unchanged from before P0-3 | Regression | `sbcli: test_unplanned_failover_ignores_demote_state` | +| U-48 | Every verb given the SOURCE side of an existing relationship resolves to the local replica (`TargetLvolId`) before acting — Ramen's S3-restore hands the destination cluster the original source's volumeHandle verbatim | Positive | `TestEnableVolumeReplicationResolvesToTargetWhenGivenTheSourceSideOfARelationship`, `TestDisableVolumeReplicationResolvesToTargetWhenGivenTheSourceSideOfARelationship`, `TestPromoteVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship`, `TestDemoteVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship` | +| U-49 | Resync resolves the foreign source handle even after the source lvol record itself was reaped (2026-09-24-resync-foreign-handle-404: the only verb without the resolution 404ed on every reconcile and stalled the relocate back) | Regression | `TestResyncVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship` | +| U-50 | The standalone info read resolves the foreign source handle likewise, so lastSyncTime keeps flowing after the source record is reaped (2026-09-24-resync-foreign-handle-404) | Regression | `TestGetVolumeReplicationInfoResolvesToTargetWhenGivenTheSourceSideOfARelationship` | +| U-51 | Promote while the fail-back's first reverse-replicated snapshot is still in flight: control-plane 409 (converging, retryable), never a stack-traced 500; the same backend message with no replication configured stays 500 | Regression | `sbcli: test_failover_while_the_first_replicated_snapshot_is_in_flight_is_409`, `test_failover_with_no_snapshot_and_no_replication_configured_stays_500` | +| U-52 | DeleteVolume carrying a reaped, foreign handle follows the relationship and removes the retired, non-active members (2026-09-24-delete-foreign-handle-leak); the `Relationship.ActiveLvolID` parse in atlas-lib is part of the same fix | Regression | `TestDeleteVolumeRemovesTheRetiredReplicaBehindAForeignHandle`, `atlas-lib: TestClientGetVolumeReplicationRelationshipParsesTheActiveVolume` | +| U-53 | DeleteVolume never deletes the pairing's ACTIVE volume through a dead source handle — the workload is running on it | Negative | `TestDeleteVolumeNeverDeletesTheActiveReplicaThroughADeadHandle` | +| U-54 | Resolution walks CHAINED pairings to the active volume: after a relocate round trip, verbs on the original handle land on the last hop's clone, never the retired middle hop (2026-09-24-chained-relationship-resolves-one-hop-short) | Regression | `TestPromoteVolumeResolvesAcrossAChainedRelationshipToTheActiveVolume` | +| U-55 | The node's deleted-volume redirect walks the same chain using each record's own target triple, never pairing active_lvol_id with another hop's cluster; the single-pairing migration case is unchanged | Regression | `node: TestRedirectToActiveVolumeWalksAChainedRelationship`, `TestRedirectToActiveVolumeSinglePairingIsUnchanged` | +| U-56 | Promote re-resolves after success and attaches the class's policy to the newly created primary (2026-09-24-promote-leaves-new-primary-unprotected: Enable only runs at VR creation, before the clone exists); no class parameter attaches nothing | Regression | `TestPromoteVolumeAttachesThePolicyToTheNewlyActiveVolume`, `TestPromoteVolumeWithoutThePolicyParameterAttachesNothing` | ### Condition Derivation (design §6.2) @@ -183,7 +184,7 @@ Two live simplyblock clusters with the chart-deployed csi-addons machinery. The | Axis | Values covered | IDs | Not covered | |------------------|------------------------------------------------------------------------|-----------------------------------------------------------------------------|-----------------------------------------------------------| -| Verb lifecycle | enable, disable, info, forced promote, planned promote, demote, resync | U-01, U-02, U-04 … U-10, U-12 … U-17, U-40 … U-55, I-02 … I-04, E-01 … E-04 | — | +| Verb lifecycle | enable, disable, info, forced promote, planned promote, demote, resync | U-01, U-02, U-04 … U-10, U-12 … U-17, U-40 … U-56, I-02 … I-04, E-01 … E-04 | — | | Idempotency | repeat enable, disable, demote; re-drive after restart | U-02, U-05, U-15, I-07 | repeated promote (U-11), repeated resync | | Conditions | healthy, degraded, error, staleness, resyncing, disabled | U-18 … U-22, E-05 | condition behavior across backend upgrade | | Coexistence | slot skip, one-owner refusal, concurrent claim | U-26, U-27, M-02 | migration of an annotated volume onto a VolumeReplication | From c9bce7621ed1e24b90b6a00e47d5189889061816 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Thu, 24 Sep 2026 18:57:56 +0100 Subject: [PATCH 155/206] Revert post-promote policy attach --- .../internal/csi/controller/replication.go | 55 +++---------- .../controller/replication_lifecycle_test.go | 77 ++++--------------- .../tests/test-plan-csi-addons-replication.md | 54 ++++++------- 3 files changed, 54 insertions(+), 132 deletions(-) diff --git a/csi-driver/internal/csi/controller/replication.go b/csi-driver/internal/csi/controller/replication.go index 7a75b0fea..1477a68d1 100644 --- a/csi-driver/internal/csi/controller/replication.go +++ b/csi-driver/internal/csi/controller/replication.go @@ -255,52 +255,21 @@ func (cs *Server) PromoteVolume( if err := client.PromoteVolume(ctx, h.Handle(), req.GetForce()); err != nil { return nil, classifyPromoteVolumeError(err) } - if err := attachPolicyToPromotedVolume(ctx, req); err != nil { - return nil, err - } + // Deliberately NOT attaching the class's policy to the promoted volume + // here, even though the new primary comes up unprotected (policy=NONE) + // until the next VR creation re-runs Enable. Attaching at promote time + // was tried and confirmed harmful live (2026-09-24, relocate M-02): the + // attach starts policy-driven replication of the fresh clone immediately, + // and when the relocate's fail-back then re-points the SAME volume's + // replication at the original source node, the two chains collide -- the + // fail-over clone build picked a snapshot from the policy's chain and + // died on "Failed to create BDev" on the wrong LVS, wedging the whole + // relocate. Re-protecting the new primary belongs to the planned-cutover + // (replication_commit) flow, where it can be sequenced strictly after the + // cutover completes instead of racing the fail-back. return &replication.PromoteVolumeResponse{}, nil } -// attachPolicyToPromotedVolume re-resolves the request's handle AFTER a -// successful promote and attaches the VolumeReplicationClass's policy to the -// volume the walk now reaches -- the clone the promote just created. The -// vendored csi-addons controller calls EnableVolumeReplication only at VR -// creation, which on a relocate is BEFORE that clone exists: Enable lands on -// the pre-promote side and no later reconcile re-issues it (confirmed live -// 2026-09-24, twice -- the new primary ran unprotected, policy=NONE, while -// the retired clone kept the policy). Promote is the one call that knows the -// new primary exists, and the class parameters ride on every RPC, so it -// finishes the job itself. No policy parameter means nothing to attach: -// promote outside a policy stays exactly what it was. -func attachPolicyToPromotedVolume(ctx context.Context, req *replication.PromoteVolumeRequest) error { - policyID := req.GetParameters()[replicationPolicyParam] - if policyID == "" { - return nil - } - h, err := csicommon.ParseVolumeHandle(volumeIDFrom(req)) - if err != nil { - return status.Error(codes.InvalidArgument, err.Error()) - } - client, err := clusters.ReplicationClient(ctx, h.ClusterID) - if err != nil { - return status.Error(codes.Unavailable, err.Error()) - } - h, client, err = resolveToLocalReplica(ctx, h, client) - if err != nil { - return status.Error(codes.Unavailable, err.Error()) - } - if err := client.EnableVolumeReplication(ctx, h.Handle(), policyID); err != nil { - if errors.Is(err, errs.ErrNotFound) { - // The backend has not registered the promoted clone yet; the - // controller re-drives promote (idempotent) and the attach lands - // on a later pass. - return nil - } - return classifyEnableVolumeReplicationError(err) - } - return nil -} - // DemoteVolume fences the source and confirms the last write replicated // (P0-3) -- the lossless half of a planned swap. Synchronous and // non-blocking: it never waits out the backend's own convergence loop. diff --git a/csi-driver/internal/csi/controller/replication_lifecycle_test.go b/csi-driver/internal/csi/controller/replication_lifecycle_test.go index 6fc90800e..4ebf8a687 100644 --- a/csi-driver/internal/csi/controller/replication_lifecycle_test.go +++ b/csi-driver/internal/csi/controller/replication_lifecycle_test.go @@ -530,30 +530,25 @@ func TestPromoteVolumeResolvesAcrossAChainedRelationshipToTheActiveVolume(t *tes } } -// Regression: 2026-09-24-promote-leaves-new-primary-unprotected — the -// vendored csi-addons controller calls EnableVolumeReplication only at VR -// creation, which on a relocate happens BEFORE the promote creates the new -// primary clone: Enable lands on the pre-promote side and nothing ever -// attaches the policy to the volume the workload actually ends up on -// (confirmed live 2026-09-24, twice: the policy stayed on the retired clone -// while the new primary ran with policy=NONE, do_replicate=False, and even a -// forced reconcile-all after a manager restart issued no further Enable). -// The VolumeReplicationClass parameters ride on EVERY replication RPC, -// promote included, so PromoteVolume itself must re-resolve after a -// successful promote -- the walk now reaches the newly created clone -- and -// attach the class's policy there. -func TestPromoteVolumeAttachesThePolicyToTheNewlyActiveVolume(t *testing.T) { +// Regression: 2026-09-24-promote-must-not-attach-the-policy — attaching the +// class's policy to the promoted volume inside PromoteVolume was tried (to +// close the "new primary comes up with policy=NONE" gap) and confirmed +// harmful live the same day: the attach starts policy-driven replication of +// the fresh clone immediately, the relocate's fail-back then re-points the +// SAME volume's replication at the original source node, and the two chains +// collide -- the fail-over clone build died on "Failed to create BDev" and +// the whole relocate wedged at WaitForReadiness. Promote must promote and +// nothing else, even when the class parameters (which ride on every RPC) +// name a policy; re-protection is the planned-cutover flow's job. +func TestPromoteVolumeDoesNotAttachThePolicyItWasHanded(t *testing.T) { mock := newMockSBCLI() defer mock.Close() cs := newReplicationTestServer(t, mock) mock.volumes[testReplTargetVolumeID] = &mockVolume{ - UUID: testReplTargetVolumeID, Name: "repl-vol-hop1-clone", Size: 1 << 30, - } - mock.volumes[testReplActiveVolumeID] = &mockVolume{ - UUID: testReplActiveVolumeID, Name: "repl-vol-new-primary", Size: 1 << 30, + UUID: testReplTargetVolumeID, Name: "repl-vol-target", Size: 1 << 30, } mock.replicationRelationship[testReplVolumeID] = map[string]any{ - "replication_id": "aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeee1", + "replication_id": "aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee", "direction": "to_target", "mode": "failover", "state": "failed_over", @@ -561,18 +556,7 @@ func TestPromoteVolumeAttachesThePolicyToTheNewlyActiveVolume(t *testing.T) { "source_cluster_id": sanityClusterID, "source_lvol_id": testReplVolumeID, "target_cluster_id": sanityClusterID, "target_pool_id": sanityPoolUUID, "target_lvol_id": testReplTargetVolumeID, "target_nqn": "nqn.test", "target_ns_id": 1, - "active": "target", "active_lvol_id": testReplActiveVolumeID, - } - mock.replicationRelationship[testReplTargetVolumeID] = map[string]any{ - "replication_id": "aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeee2", - "direction": "to_target", - "mode": "failover", - "state": "failed_over", - "is_source": true, - "source_cluster_id": sanityClusterID, "source_lvol_id": testReplTargetVolumeID, - "target_cluster_id": sanityClusterID, "target_pool_id": sanityPoolUUID, "target_lvol_id": testReplActiveVolumeID, - "target_nqn": "nqn.test", "target_ns_id": 1, - "active": "target", "active_lvol_id": testReplActiveVolumeID, + "active": "target", "active_lvol_id": testReplTargetVolumeID, } delete(mock.volumes, testReplVolumeID) @@ -584,39 +568,8 @@ func TestPromoteVolumeAttachesThePolicyToTheNewlyActiveVolume(t *testing.T) { if err != nil { t.Fatal(err) } - if got := mock.volumes[testReplActiveVolumeID].ReplicationPolicyID; got != testReplPolicyID { - t.Errorf("new primary's policy = %q, want %q attached by the promote itself", got, testReplPolicyID) - } -} - -// Without the class parameter there is nothing to attach, and promote must -// stay exactly what it was -- some callers drive promote outside any policy. -func TestPromoteVolumeWithoutThePolicyParameterAttachesNothing(t *testing.T) { - mock := newMockSBCLI() - defer mock.Close() - cs := newReplicationTestServer(t, mock) - mock.volumes[testReplTargetVolumeID] = &mockVolume{ - UUID: testReplTargetVolumeID, Name: "repl-vol-target", Size: 1 << 30, - } - mock.replicationRelationship[testReplVolumeID] = map[string]any{ - "replication_id": "aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee", - "direction": "to_target", - "mode": "failover", - "state": "failed_over", - "is_source": true, - "source_cluster_id": sanityClusterID, "source_lvol_id": testReplVolumeID, - "target_cluster_id": sanityClusterID, "target_pool_id": sanityPoolUUID, "target_lvol_id": testReplTargetVolumeID, - "target_nqn": "nqn.test", "target_ns_id": 1, - } - - _, err := cs.PromoteVolume(context.Background(), &replication.PromoteVolumeRequest{ - VolumeId: testReplVolID, Force: true, - }) - if err != nil { - t.Fatal(err) - } if got := mock.volumes[testReplTargetVolumeID].ReplicationPolicyID; got != "" { - t.Errorf("policy = %q attached with no class parameter, want none", got) + t.Errorf("promote attached policy %q to the promoted volume; it must attach nothing", got) } } diff --git a/operator/docs/tests/test-plan-csi-addons-replication.md b/operator/docs/tests/test-plan-csi-addons-replication.md index 59a29f15b..7787f1e09 100644 --- a/operator/docs/tests/test-plan-csi-addons-replication.md +++ b/operator/docs/tests/test-plan-csi-addons-replication.md @@ -39,33 +39,33 @@ File: `csi-driver/internal/csi/controller/replication_test.go` (planned) File: `csi-driver/internal/csi/controller/replication_lifecycle_test.go` -| # | Scenario | Type | Test | -|------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| U-10 | Forced promote: the failover endpoint is called without `planned=true`, ignoring demote state entirely | Positive | `TestPromoteVolumeForced` | -| U-11 | Forced promote repeated after completion: success without a second failover (backend idempotency honored) | Boundary | — | -| U-12 | Planned promote sends `planned=true` on the failover call | Positive | `TestPromoteVolumePlannedSendsThePlannedFlag` | -| U-13 | Planned promote with no demote ever requested: `FAILED_PRECONDITION`, the one case meant to let the vendored controller's own force-escalation take over | Negative | `TestPromoteVolumePlannedWithNoDemoteIsFailedPrecondition` | -| U-14 | Demote: the demote endpoint is called, and the RPC succeeds only once the backend confirms the final snapshot landed | Positive | `TestDemoteVolumeDone` | -| U-15 | Demote repeated on a demoted volume: success (idempotency) | Boundary | `sbcli: test_demote_is_idempotent_once_done` (backend tier; the driver's `TestDemoteVolumeDone` exercises the same success path) | -| U-16 | Resync: the failback endpoint is called, forwarding the `sourceClusterID` class parameter when given | Positive | `TestResyncVolume`, `TestResyncVolumeSendsTheSourceClusterParameter` | -| U-17 | Resync's `ready` field reflects lag against budget: false while lag exceeds it, true once caught up | Boundary | `TestResyncVolumeNotReadyWhileLagExceedsBudget`, `TestResyncVolumeReadyReflectsLag` | -| U-40 | Planned promote while demote is still converging: `ABORTED` (retryable) — never `FAILED_PRECONDITION`, which the vendored controller auto-escalates to a forced, lossy promote inline with no wait-and-retry grace period of its own | Negative | `TestPromoteVolumePlannedWhileDemoteConvergingIsAborted` | -| U-41 | Demote still converging: `ABORTED` (retryable), non-blocking — the RPC never waits out the backend's own convergence loop | Boundary | `TestDemoteVolumeNotYetDoneIsAborted` | -| U-42 | `demote_lvol` fences the source strictly before triggering the final snapshot, never after (a write landing in the gap would be silently lost) | Positive | `sbcli: test_demote_fences_before_triggering_the_final_snapshot` | -| U-43 | `demote_lvol` re-invoked while pending: checks the marker only, never re-fences or re-triggers | Boundary | `sbcli: test_demote_does_not_refence_or_retrigger_once_pending` | -| U-44 | `demote_lvol` completes once the triggered snapshot carries the replicated marker | Positive | `sbcli: test_demote_completes_once_the_snapshot_carries_the_replicated_marker` | -| U-45 | `demote_lvol` surfaces a snapshot-creation failure without recording pending state | Negative | `sbcli: test_demote_surfaces_a_snapshot_creation_failure` | -| U-46 | The `failover` route's three-way planned-gate branch: demoted proceeds, converging is 409, no relationship is 412 | Positive/Negative | `sbcli: test_planned_failover_proceeds_once_demoted`, `test_planned_failover_while_demote_is_converging_is_409`, `test_planned_failover_without_any_demote_is_412` | -| U-47 | Unplanned failover ignores demote state, unchanged from before P0-3 | Regression | `sbcli: test_unplanned_failover_ignores_demote_state` | -| U-48 | Every verb given the SOURCE side of an existing relationship resolves to the local replica (`TargetLvolId`) before acting — Ramen's S3-restore hands the destination cluster the original source's volumeHandle verbatim | Positive | `TestEnableVolumeReplicationResolvesToTargetWhenGivenTheSourceSideOfARelationship`, `TestDisableVolumeReplicationResolvesToTargetWhenGivenTheSourceSideOfARelationship`, `TestPromoteVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship`, `TestDemoteVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship` | -| U-49 | Resync resolves the foreign source handle even after the source lvol record itself was reaped (2026-09-24-resync-foreign-handle-404: the only verb without the resolution 404ed on every reconcile and stalled the relocate back) | Regression | `TestResyncVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship` | -| U-50 | The standalone info read resolves the foreign source handle likewise, so lastSyncTime keeps flowing after the source record is reaped (2026-09-24-resync-foreign-handle-404) | Regression | `TestGetVolumeReplicationInfoResolvesToTargetWhenGivenTheSourceSideOfARelationship` | -| U-51 | Promote while the fail-back's first reverse-replicated snapshot is still in flight: control-plane 409 (converging, retryable), never a stack-traced 500; the same backend message with no replication configured stays 500 | Regression | `sbcli: test_failover_while_the_first_replicated_snapshot_is_in_flight_is_409`, `test_failover_with_no_snapshot_and_no_replication_configured_stays_500` | -| U-52 | DeleteVolume carrying a reaped, foreign handle follows the relationship and removes the retired, non-active members (2026-09-24-delete-foreign-handle-leak); the `Relationship.ActiveLvolID` parse in atlas-lib is part of the same fix | Regression | `TestDeleteVolumeRemovesTheRetiredReplicaBehindAForeignHandle`, `atlas-lib: TestClientGetVolumeReplicationRelationshipParsesTheActiveVolume` | -| U-53 | DeleteVolume never deletes the pairing's ACTIVE volume through a dead source handle — the workload is running on it | Negative | `TestDeleteVolumeNeverDeletesTheActiveReplicaThroughADeadHandle` | -| U-54 | Resolution walks CHAINED pairings to the active volume: after a relocate round trip, verbs on the original handle land on the last hop's clone, never the retired middle hop (2026-09-24-chained-relationship-resolves-one-hop-short) | Regression | `TestPromoteVolumeResolvesAcrossAChainedRelationshipToTheActiveVolume` | -| U-55 | The node's deleted-volume redirect walks the same chain using each record's own target triple, never pairing active_lvol_id with another hop's cluster; the single-pairing migration case is unchanged | Regression | `node: TestRedirectToActiveVolumeWalksAChainedRelationship`, `TestRedirectToActiveVolumeSinglePairingIsUnchanged` | -| U-56 | Promote re-resolves after success and attaches the class's policy to the newly created primary (2026-09-24-promote-leaves-new-primary-unprotected: Enable only runs at VR creation, before the clone exists); no class parameter attaches nothing | Regression | `TestPromoteVolumeAttachesThePolicyToTheNewlyActiveVolume`, `TestPromoteVolumeWithoutThePolicyParameterAttachesNothing` | +| # | Scenario | Type | Test | +|------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| U-10 | Forced promote: the failover endpoint is called without `planned=true`, ignoring demote state entirely | Positive | `TestPromoteVolumeForced` | +| U-11 | Forced promote repeated after completion: success without a second failover (backend idempotency honored) | Boundary | — | +| U-12 | Planned promote sends `planned=true` on the failover call | Positive | `TestPromoteVolumePlannedSendsThePlannedFlag` | +| U-13 | Planned promote with no demote ever requested: `FAILED_PRECONDITION`, the one case meant to let the vendored controller's own force-escalation take over | Negative | `TestPromoteVolumePlannedWithNoDemoteIsFailedPrecondition` | +| U-14 | Demote: the demote endpoint is called, and the RPC succeeds only once the backend confirms the final snapshot landed | Positive | `TestDemoteVolumeDone` | +| U-15 | Demote repeated on a demoted volume: success (idempotency) | Boundary | `sbcli: test_demote_is_idempotent_once_done` (backend tier; the driver's `TestDemoteVolumeDone` exercises the same success path) | +| U-16 | Resync: the failback endpoint is called, forwarding the `sourceClusterID` class parameter when given | Positive | `TestResyncVolume`, `TestResyncVolumeSendsTheSourceClusterParameter` | +| U-17 | Resync's `ready` field reflects lag against budget: false while lag exceeds it, true once caught up | Boundary | `TestResyncVolumeNotReadyWhileLagExceedsBudget`, `TestResyncVolumeReadyReflectsLag` | +| U-40 | Planned promote while demote is still converging: `ABORTED` (retryable) — never `FAILED_PRECONDITION`, which the vendored controller auto-escalates to a forced, lossy promote inline with no wait-and-retry grace period of its own | Negative | `TestPromoteVolumePlannedWhileDemoteConvergingIsAborted` | +| U-41 | Demote still converging: `ABORTED` (retryable), non-blocking — the RPC never waits out the backend's own convergence loop | Boundary | `TestDemoteVolumeNotYetDoneIsAborted` | +| U-42 | `demote_lvol` fences the source strictly before triggering the final snapshot, never after (a write landing in the gap would be silently lost) | Positive | `sbcli: test_demote_fences_before_triggering_the_final_snapshot` | +| U-43 | `demote_lvol` re-invoked while pending: checks the marker only, never re-fences or re-triggers | Boundary | `sbcli: test_demote_does_not_refence_or_retrigger_once_pending` | +| U-44 | `demote_lvol` completes once the triggered snapshot carries the replicated marker | Positive | `sbcli: test_demote_completes_once_the_snapshot_carries_the_replicated_marker` | +| U-45 | `demote_lvol` surfaces a snapshot-creation failure without recording pending state | Negative | `sbcli: test_demote_surfaces_a_snapshot_creation_failure` | +| U-46 | The `failover` route's three-way planned-gate branch: demoted proceeds, converging is 409, no relationship is 412 | Positive/Negative | `sbcli: test_planned_failover_proceeds_once_demoted`, `test_planned_failover_while_demote_is_converging_is_409`, `test_planned_failover_without_any_demote_is_412` | +| U-47 | Unplanned failover ignores demote state, unchanged from before P0-3 | Regression | `sbcli: test_unplanned_failover_ignores_demote_state` | +| U-48 | Every verb given the SOURCE side of an existing relationship resolves to the local replica (`TargetLvolId`) before acting — Ramen's S3-restore hands the destination cluster the original source's volumeHandle verbatim | Positive | `TestEnableVolumeReplicationResolvesToTargetWhenGivenTheSourceSideOfARelationship`, `TestDisableVolumeReplicationResolvesToTargetWhenGivenTheSourceSideOfARelationship`, `TestPromoteVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship`, `TestDemoteVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship` | +| U-49 | Resync resolves the foreign source handle even after the source lvol record itself was reaped (2026-09-24-resync-foreign-handle-404: the only verb without the resolution 404ed on every reconcile and stalled the relocate back) | Regression | `TestResyncVolumeResolvesToTargetWhenGivenTheSourceSideOfARelationship` | +| U-50 | The standalone info read resolves the foreign source handle likewise, so lastSyncTime keeps flowing after the source record is reaped (2026-09-24-resync-foreign-handle-404) | Regression | `TestGetVolumeReplicationInfoResolvesToTargetWhenGivenTheSourceSideOfARelationship` | +| U-51 | Promote while the fail-back's first reverse-replicated snapshot is still in flight: control-plane 409 (converging, retryable), never a stack-traced 500; the same backend message with no replication configured stays 500 | Regression | `sbcli: test_failover_while_the_first_replicated_snapshot_is_in_flight_is_409`, `test_failover_with_no_snapshot_and_no_replication_configured_stays_500` | +| U-52 | DeleteVolume carrying a reaped, foreign handle follows the relationship and removes the retired, non-active members (2026-09-24-delete-foreign-handle-leak); the `Relationship.ActiveLvolID` parse in atlas-lib is part of the same fix | Regression | `TestDeleteVolumeRemovesTheRetiredReplicaBehindAForeignHandle`, `atlas-lib: TestClientGetVolumeReplicationRelationshipParsesTheActiveVolume` | +| U-53 | DeleteVolume never deletes the pairing's ACTIVE volume through a dead source handle — the workload is running on it | Negative | `TestDeleteVolumeNeverDeletesTheActiveReplicaThroughADeadHandle` | +| U-54 | Resolution walks CHAINED pairings to the active volume: after a relocate round trip, verbs on the original handle land on the last hop's clone, never the retired middle hop (2026-09-24-chained-relationship-resolves-one-hop-short) | Regression | `TestPromoteVolumeResolvesAcrossAChainedRelationshipToTheActiveVolume` | +| U-55 | The node's deleted-volume redirect walks the same chain using each record's own target triple, never pairing active_lvol_id with another hop's cluster; the single-pairing migration case is unchanged | Regression | `node: TestRedirectToActiveVolumeWalksAChainedRelationship`, `TestRedirectToActiveVolumeSinglePairingIsUnchanged` | +| U-56 | Promote does NOT attach the class's policy to the promoted volume, even when handed one (2026-09-24-promote-must-not-attach-the-policy: an attach at promote time races the fail-back's re-pointing of the same volume's replication and wedged the relocate on a wrong-LVS clone build); re-protection belongs to the planned-cutover flow | Regression | `TestPromoteVolumeDoesNotAttachThePolicyItWasHanded` | ### Condition Derivation (design §6.2) From fab913d5304c5ae81d49ff94def44026cf9805d5 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Fri, 25 Sep 2026 20:56:06 +0100 Subject: [PATCH 156/206] docs(csi-addons): document VolumeGroupReplication in the csi-addons design and test plan --- .../designs/design-csi-addons-replication.md | 164 ++++++++++++++---- .../docs/designs/design-ramen-integration.md | 4 +- .../tests/test-plan-csi-addons-replication.md | 70 +++++--- .../docs/tests/test-plan-ramen-integration.md | 2 +- 4 files changed, 179 insertions(+), 61 deletions(-) diff --git a/operator/docs/designs/design-csi-addons-replication.md b/operator/docs/designs/design-csi-addons-replication.md index 0f052ff65..d304b860d 100644 --- a/operator/docs/designs/design-csi-addons-replication.md +++ b/operator/docs/designs/design-csi-addons-replication.md @@ -1,8 +1,8 @@ # Design Document: csi-addons Volume Replication -**Status:** Phase 3 Implemented +**Status:** Phase 4 Implemented **Author:** Israel Geoffrey (geoffrey1330) -**Date:** 2026-09-16 (last updated 2026-09-21) +**Date:** 2026-09-16 (last updated 2026-09-25) **Test Plan:** [`tests/test-plan-csi-addons-replication.md`](../tests/test-plan-csi-addons-replication.md) --- @@ -13,9 +13,10 @@ |-------------|-------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|--------------| | **Phase 1** | Implemented | The csi-addons machinery and the steady-state contract: CRDs, controller-manager, sidecar, the Replication and csi-addons Identity gRPC services with `EnableVolumeReplication`, `DisableVolumeReplication`, and `GetVolumeReplicationInfo`, backed by a typed backend status endpoint | §4, §5.1, §6 | | **Phase 2** | Implemented | The lifecycle verbs: `PromoteVolume` (planned and forced), `DemoteVolume`, and `ResyncVolume`. Validation end to end against a Ramen `VolumeReplicationGroup` in async mode is still outstanding (§12, E-06/E-07) | §5.2, §9 | -| **Phase 3** | Implemented | §11 (the Prometheus metrics). §7.2's peerClasses preflight is out of this design's scope entirely (it's Ramen's own `DRPolicy` mechanism) and is deferred to a future Ramen-integration design | §7.1, §11 | +| **Phase 3** | Implemented | §11 (the Prometheus metrics). §7.2's peerClasses preflight is out of this design's scope entirely (it's Ramen's own `DRPolicy` mechanism) and is deferred to a future Ramen-integration design | §7.1, §11 | +| **Phase 4** | Implemented | §14 (`VolumeGroupReplication`): the group-level csi-addons surface on top of the consistency-group primitive, fanning the §5 verbs out to a group's members. This is the gap analysis's own Phase 2 group-replication item | §14 | -Phase 1 is independently useful: a `VolumeReplication` object per PVC whose status truthfully reports the relationship, which no surface provides today. Phase 2 makes the object drivable, which is what Ramen actually needs. Phase 3 makes the whole thing operable at fleet scale. +Phase 1 is independently useful: a `VolumeReplication` object per PVC whose status truthfully reports the relationship, which no surface provides today. Phase 2 makes the object drivable, which is what Ramen actually needs. Phase 3 makes the whole thing operable at fleet scale. Phase 4 lifts the same surface from one volume to a consistency group. The phase numbers above are this document's own, not the DR storage foundation gap analysis's (§1): its Phase 0 (shipping the csi-addons contract itself) is this design's Phase 1, and its Phase 1 (promote, demote, and resync end to end through Ramen) is this design's Phase 2. @@ -23,13 +24,13 @@ The phase numbers above are this document's own, not the DR storage foundation g ## Phase 0 — External Prerequisites -| # | Prerequisite | Kind | Blocks | Status | -|------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------|---------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| P0-1 | A typed, steady-state per-volume replication status read: `GET .../volumes/{id}/replication/status` serving what `lvol_controller.get_replication_info` computes today (state, lag, outstanding bytes, failure counters), available for the volume's whole replicated life | Control plane (`sbcli`) | Phase 1 | Shipped: `GET .../volumes/{v}/replication/status` → `ReplicationStatusDTO` (`simplyblock_web/api/v2/cluster/storage_pool/volume/replication.py:58-69`) | -| P0-2 | Idempotent attach and detach: attaching a volume to the policy it already follows returns success, and detaching a non-attached volume returns success | Control plane (`sbcli`) | Phase 1 | Shipped: `replication_policy_controller.attach_policy`/`detach_policy` (`simplyblock_core/controllers/replication_policy_controller.py:208-262`) | -| P0-3 | A standalone demote verb: `POST .../volumes/{id}/replication/demote` that converges the peer while still serving (repeated snapshot-and-ship until the remaining delta is small), then quiesces, ships the final delta, confirms it landed on the peer, and fences the data path | Control plane (`sbcli`) | Phase 2 | Shipped: `POST .../volumes/{v}/replication/demote` → `lvol_controller.demote_lvol` | +| # | Prerequisite | Kind | Blocks | Status | +|------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------|---------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| P0-1 | A typed, steady-state per-volume replication status read: `GET .../volumes/{id}/replication/status` serving what `lvol_controller.get_replication_info` computes today (state, lag, outstanding bytes, failure counters), available for the volume's whole replicated life | Control plane (`sbcli`) | Phase 1 | Shipped: `GET .../volumes/{v}/replication/status` → `ReplicationStatusDTO` (`simplyblock_web/api/v2/cluster/storage_pool/volume/replication.py:58-69`) | +| P0-2 | Idempotent attach and detach: attaching a volume to the policy it already follows returns success, and detaching a non-attached volume returns success | Control plane (`sbcli`) | Phase 1 | Shipped: `replication_policy_controller.attach_policy`/`detach_policy` (`simplyblock_core/controllers/replication_policy_controller.py:208-262`) | +| P0-3 | A standalone demote verb: `POST .../volumes/{id}/replication/demote` that converges the peer while still serving (repeated snapshot-and-ship until the remaining delta is small), then quiesces, ships the final delta, confirms it landed on the peer, and fences the data path | Control plane (`sbcli`) | Phase 2 | Shipped: `POST .../volumes/{v}/replication/demote` → `lvol_controller.demote_lvol` | | P0-4 | An `rpo_target_seconds` field on `ReplicationPolicy`, so RPO compliance is computable against a declared target rather than the derived lag budget | Control plane (`sbcli`) | Phase 3 | Shipped: `ReplicationPolicy.rpo_target_seconds` (`simplyblock_core/models/replication.py:89`), wired through the API (`PolicyParams.rpo_target_seconds`) and CLI (`--rpo-target-sec`) | -| P0-5 | csi-addons upstream: the `VolumeReplication` and `VolumeReplicationClass` CRDs (`replication.storage.openshift.io/v1alpha1`), the kubernetes-csi-addons controller-manager image, and the csi-addons sidecar image | Ecosystem | Phase 1 | Vendored in the chart at v0.15.0 behind `csiaddons.create` (all twelve upstream CRDs, since the stock manager starts a controller per kind); sidecar wiring is Phase 1 | +| P0-5 | csi-addons upstream: the `VolumeReplication` and `VolumeReplicationClass` CRDs (`replication.storage.openshift.io/v1alpha1`), the kubernetes-csi-addons controller-manager image, and the csi-addons sidecar image | Ecosystem | Phase 1 | Vendored in the chart at v0.15.0 behind `csiaddons.create` (all twelve upstream CRDs, since the stock manager starts a controller per kind); sidecar wiring is Phase 1 | Everything else the adapter needs already exists: the attach and detach calls, failover, the failback and commit pair, the relationship read, and the backlog arithmetic inside `get_replication_info`. The adapter is thin precisely because the engine is complete. What is missing is the shape Ramen can drive. @@ -50,7 +51,8 @@ Everything else the adapter needs already exists: the attach and detach calls, f 11. [Observability](#11-observability) 12. [Testing Strategy](#12-testing-strategy) 13. [Migration Strategy](#13-migration-strategy) -14. [Open Questions](#14-open-questions) +14. [VolumeGroupReplication](#14-volumegroupreplication) +15. [Open Questions](#15-open-questions) --- @@ -78,7 +80,7 @@ A reader who stops here has the model: the engine is unchanged, the csi-addons s **The contract this design targets.** Ramen's dr-cluster operator reconciles one `VolumeReplication` per protected PVC. It flips `spec.replicationState` between `primary` and `secondary` and waits for the driver's conditions (`Completed`, `Degraded`, `Resyncing`) to report the operation done and the relationship healthy. It reads `status.lastSyncTime` for RPO. It never calls a vendor API. The interfaces are the csi-addons specification's Replication gRPC, served by the driver, and the kubernetes-csi-addons controller-manager, which turns `VolumeReplication` objects into those RPCs through a per-driver sidecar. -**The direction is already committed.** The CRD redesign excludes the four replication kinds from its model because "that subsystem is being redesigned against the CSI Addons specification, whose `VolumeReplication` and `VolumeGroupReplication` kinds already carry the per-volume and per-group replication contract that a backup tool or a DR orchestrator understands" (`crd-redesign/design-crd-model.md`, Non-Goals). This document is that redesign's first, per-volume half. The DR storage foundation gap analysis (Phase 0 and Appendix A of that document) is its requirements source. +**The direction is already committed.** The CRD redesign excludes the four replication kinds from its model because "that subsystem is being redesigned against the CSI Addons specification, whose `VolumeReplication` and `VolumeGroupReplication` kinds already carry the per-volume and per-group replication contract that a backup tool or a DR orchestrator understands" (`crd-redesign/design-crd-model.md`, Non-Goals). This document is that redesign's replication chapter: the per-volume half in §5, and the per-group half (`VolumeGroupReplication`) in §14. The DR storage foundation gap analysis (Phase 0 and Appendix A of that document) is its requirements source. **Three facts about today's surface shape the design.** @@ -102,7 +104,7 @@ A reader who stops here has the model: the engine is unchanged, the csi-addons s ### Non-Goals -- **`VolumeGroupReplication`.** Continuous group promote and demote on top of consistency groups is the next design. This one is strictly per volume. The group snapshot surface (`design-consistency-groups.md`) is untouched. +- **Global `VolumeGroupReplication`.** The base group surface (one VRG, one vendor's PVCs, one `VolumeGroupReplication` on top of a consistency group) is specified in §14. RamenDR's newer multi-VRG "Global VGR" consensus, spanning a replication group across several applications, is out of scope (§14.8, Open Question 4). The group snapshot surface (`design-consistency-groups.md`) is untouched. - **VolSync and the S3 backup path.** The recurrent-immutable-snapshots DR type is a separate phase of the gap analysis and does not pass through this adapter. - **Synchronous replication.** The engine is asynchronous snapshot shipping, and nothing here changes that. - **Ramen hub components.** DRPolicy, DRPC, and hub orchestration are consumers of this contract, not part of it. @@ -304,16 +306,16 @@ The consolidation direction (§13) is that the annotation path becomes a compati Every endpoint is scoped as today: volume-scoped under `/api/v2/clusters/{c}/storage-pools/{p}/volumes/{v}`, cluster-scoped under `/api/v2/clusters/{c}/replication`. -| Method | Endpoint | Notes | -|--------|---------------------------------------------------------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| `PUT` | `.../volumes/{v}` (`replication_policy_id`) | Existing attach and detach. P0-2 makes both idempotent: same-policy attach and non-attached detach return success. | -| `GET` | `.../volumes/{v}/replication/status` | **New (P0-1).** The typed steady-state status of §6.1. Never 404s for a volume that exists; `state: not_replicating, role: none` is a valid answer. | -| `POST` | `.../volumes/{v}/replication/failover` (+ planned gate) | Existing; the one promote, both forms. Gains a planned form that is refused unless the peer holds every acknowledged write (a completed demote). Idempotent by NQN probe. | -| `POST` | `.../volumes/{v}/replication/demote` | **New (P0-3).** Quiesce, final ship, confirm on peer, fence. Idempotent: demoting a demoted volume returns success. | -| `POST` | `.../volumes/{v}/replication/failback` | Existing. Resync (direction reversal, delta-seeded). | -| `POST` | `.../volumes/{v}/replication/commit` | Existing, unchanged, and NOT part of this contract: it stays behind the legacy `ReplicationOps` migration path only (§13). | -| `GET` | `.../replication/relationships/{lvol}` | Existing, unchanged. Cutover records only; the node redirect depends on its survive-deletion semantics. | -| `POST` | `.../volumes/{v}/replication/cutover-proceed` | Existing, unchanged, legacy path only: the adapter never reaches it, because the commit cutover is off this contract. | +| Method | Endpoint | Notes | +|--------|---------------------------------------------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| `PUT` | `.../volumes/{v}` (`replication_policy_id`) | Existing attach and detach. P0-2 makes both idempotent: same-policy attach and non-attached detach return success. | +| `GET` | `.../volumes/{v}/replication/status` | **New (P0-1).** The typed steady-state status of §6.1. Never 404s for a volume that exists; `state: not_replicating, role: none` is a valid answer. | +| `POST` | `.../volumes/{v}/replication/failover` (+ planned gate) | Existing; the one promote, both forms. Gains a planned form that is refused unless the peer holds every acknowledged write (a completed demote). Idempotent by NQN probe. | +| `POST` | `.../volumes/{v}/replication/demote` | **New (P0-3).** Quiesce, final ship, confirm on peer, fence. Idempotent: demoting a demoted volume returns success. | +| `POST` | `.../volumes/{v}/replication/failback` | Existing. Resync (direction reversal, delta-seeded). | +| `POST` | `.../volumes/{v}/replication/commit` | Existing, unchanged, and NOT part of this contract: it stays behind the legacy `ReplicationOps` migration path only (§13). | +| `GET` | `.../replication/relationships/{lvol}` | Existing, unchanged. Cutover records only; the node redirect depends on its survive-deletion semantics. | +| `POST` | `.../volumes/{v}/replication/cutover-proceed` | Existing, unchanged, legacy path only: the adapter never reaches it, because the commit cutover is off this contract. | The unused backend verbs the operator never calls (`start`, `stop`, `trigger`, `tasks`) are unaffected, and `start` and `stop` remain the policy-less legacy path. @@ -347,7 +349,7 @@ The kubernetes-csi-addons controller-manager owns events on `VolumeReplication` Exported by the control plane's existing v2 `Collector`-pattern exporter (`simplyblock_web/api/v2/metrics.py`), rebuilt from FDB on every scrape like every other series in that file. Labeled `lvol`/`lvol_name`/`pvc_name`/`pool`/`pool_name` (not the bare `volume` this section originally specified: the exporter's existing lvol-scoped metrics already use `lvol`/`lvol_name`, and joining the new series against them needs a shared label name) plus `policy`/`policy_name`/`peer_cluster`: | Metric | Description | -|----------------------------------------------|---------------------------------------------------------------------------------| +|---------------------------------------------|---------------------------------------------------------------------------------| | `simplyblock_replication_lag_seconds` | Now minus the newest fully replicated snapshot's creation time. | | `simplyblock_replication_backlog_bytes` | `outstanding_bytes`: the queued-but-unshipped snapshot sizes. | | `simplyblock_replication_last_sync_seconds` | Duration of the last shipping cycle. | @@ -366,7 +368,7 @@ Values are computed by `lvol_controller.get_replication_info_bulk`, a bulk-frien Full scenario matrix and coverage status: [`tests/test-plan-csi-addons-replication.md`](../tests/test-plan-csi-addons-replication.md) - **Unit (driver):** each verb against a mock control plane: the idempotency table (repeat enable, repeat disable, repeat promote), the refusal paths (different-policy enable, lagging planned promote, disable during cutover), the condition derivation from every status-read state, and handle parsing failures. -- **Unit (operator):** the `PVCAnnotationWatcher` skip when a `VolumeReplication` exists. +- **Unit (operator):** the `PVCAnnotationWatcher` skip when a `VolumeReplication` exists; and, for §14, `VolumeGroupReplicationReconciler`'s fan-in aggregation (group `Completed`/`Degraded`/`Resyncing` from the members, oldest-member `lastSyncTime`) and the admission webhook's membership check, against a fake client. - **Integration:** the csi-addons sidecar and controller-manager against the driver with a mock backend under envtest or kind: a `VolumeReplication` flipped `primary` to `secondary` and back walks the verbs in order and lands the conditions. - **E2E (two live clusters):** the Ramen-shaped lifecycle without Ramen: enable on the source, write data, and verify `lastSyncTime` advances; forced promote on the DR side, verifying the clone serves with the source fenced; and resync back with a planned swap (demote then promote), verifying zero loss with a hashed writer. Then the same driven by an actual Ramen VRG in async mode, which is Phase 2's acceptance gate. @@ -385,10 +387,110 @@ Three replication control surfaces exist today: the operator's kinds, the stale --- -## 14. Open Questions +## 14. VolumeGroupReplication -| # | Question | Owner | -|---|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------| -| 1 | **Demote semantics for the application.** The P0-3 demote fences the volume (ANA inaccessible) after the final flush, and with convergence folded into the verb it is now the only place a planned swap can stall. This is also the one verb the planned promote's lossless guarantee entirely depends on (§5.2): a planned promote is refused unless a completed demote already fenced the source and confirmed the final delta landed, so an unresolved failure mode here is an unresolved gap in the whole "zero loss" claim. Ramen relocation unmounts the workload first, so the fence is ordinarily unopposed, but that is Ramen's choreography, not a guarantee the driver can rely on: a stuck termination, a stale mount that never released, or a demote invoked outside Ramen's normal flow can all leave writes still arriving when quiesce fires. Confirm the verb's behavior when writes are still in flight at quiesce (block versus fail), whether the converge phase has its own budget separate from the quiesced flush, and whether a timeout in either phase must abort back to serving primary or leave the volume fenced with no automatic recovery. | Backend team | -| 2 | **Per-volume policy granularity.** A `VolumeReplicationClass` names one policy, and today one policy implies one target and cadence for all its volumes. Confirm one class per (policy, cadence) is an acceptable authoring model for Ramen's `replicationClassSelector`, or whether per-volume interval overrides are needed. | Operator / Backend team | -| 3 | ~~**Avoiding the clone on day-one protection.**~~ **Resolved:** `POST .../replication/failover?planned=true`'s no-demote branch now checks `lvol_controller.replication_source_online` (the source's own storage-node status) before falling through to `FAILED_PRECONDITION` -- an online source is a no-op (§5.2), so a healthy volume's first-ever `PromoteVolume` no longer materializes a clone. The remaining residual: a source that dies within the last health-check interval still briefly reads online, so one reconcile can treat a genuine disaster as a no-op before the node's status catches up and the controller retries -- bounded by the health-check detection window, not open-ended. | Backend team | +`VolumeGroupReplication` extends this adapter from one volume to a consistency group. It is the group-level sibling of the `VolumeReplication` surface (§5): the same csi-addons `replication.storage.openshift.io` API group, the same `replicationState` intent, and the same gRPC verbs (§5.2), fanned out to every member of a consistency group rather than driven per volume. This is the async **group** path the gap analysis's Phase 2 named ("`VolumeGroupReplication` on top of the CG primitive so the VRG async group path can protect and fail over multi-volume apps at one point"), and it is implemented. + +Ramen's `VolumeReplicationGroup` creates one `VolumeGroupReplication` when a protected application's PVCs share a `VolumeGroupReplicationClass` carrying a `ramendr.openshift.io/groupreplicationid` label, the group-level sibling of the per-volume `replicationid` label of §7.1. The reconciler and validator specified here own that object. The driver's Replication gRPC (§5) is unchanged, and no new verb is added. The end-to-end Ramen validation of this path is `design-ramen-integration.md` §6, whose test plan carries its E2E scenario as M-05. + +### 14.1 Upstream kinds and ownership + +`VolumeGroupReplication`, `VolumeGroupReplicationClass`, and `VolumeGroupReplicationContent` are upstream kubernetes-csi-addons CRDs in the same `replication.storage.openshift.io` group this design already vendors `VolumeReplication` and `VolumeReplicationClass` from (§4.1). This repository declares no Go type for them, and the reconciler and validator read and write them as `unstructured`, matching how the per-volume kinds are handled. + +`VolumeGroupReplication.spec` carries `replicationState` (`primary`/`secondary`/`resync`), a `source.selector` naming the member PVCs by label, a `volumeGroupReplicationClassName`, a `volumeReplicationClassName` (the per-member class the fan-out stamps on each member `VolumeReplication`), and an `external` boolean. `external: false` routes the object through the generic kubernetes-csi-addons controller-manager, which fans it out to member `VolumeReplication` objects itself. `external: true` hands it to a vendor controller. simplyblock is the `external: true`, "offloaded" case: the backend replicates below Kubernetes, so this operator performs the fan-out. + +`VolumeGroupReplicationReconciler` (`operator/internal/controller/volumegroupreplication_controller.go`) owns a `VolumeGroupReplication` only when both `spec.external` is `true` and `spec.volumeGroupReplicationClassName` names a `VolumeGroupReplicationClass` whose `spec.provisioner` is `csi.simplyblock.io`. Every other object (non-external, or a foreign provisioner's class) is skipped, because `replication.storage.openshift.io` is a shared group and the generic controller-manager or another vendor may be the right owner. + +### 14.2 No new gRPC contract + +Group promote, demote, and resync are the same three verbs of §5.2 (`PromoteVolume`, `DemoteVolume`, `ResyncVolume`) fanned out to every member, not a fourth verb on the driver. Each member is an ordinary `VolumeReplication` object, reconciled by the already-shipped kubernetes-csi-addons controller-manager exactly as it reconciles any Ramen-created per-volume object (§5). `external: true` changes who performs the fan-out (this operator, rather than the generic manager), not what the fan-out does. + +### 14.3 The reconciler + +`VolumeGroupReplicationReconciler` reconciles each owned `VolumeGroupReplication` on a 60-second resync (30 seconds on a transient error): + +1. **Resolve membership.** Read `spec.source.selector` and list the matching PVCs in the object's namespace. Every matched PVC must carry the same non-empty `storage.simplyblock.io/consistency-group` label; a selector that matches an unlabeled PVC, spans two group values, or matches nothing is a membership mismatch. Map each PVC through its PV handle (`{clusterID}:{poolID}:{volumeID}`) to a backend lvol, read the consistency group named by the label, and require the selected lvol set to equal the group's current backend membership exactly, the same invariant `design-consistency-groups.md` §9.2 established for `VolumeGroupSnapshot`. A selected PVC that is not yet bound makes membership undeterminable, which requeues quietly rather than failing. +2. **Fan out.** For each member PVC, ensure a per-volume `VolumeReplication` exists, owned by the group, named `-`, with `spec.replicationState` mirroring the group's, `spec.volumeReplicationClass` set to the group's `volumeReplicationClassName`, `spec.autoResync: false`, and a `dataSource` naming the member PVC. An existing member whose `replicationState` has drifted from the group's is updated. The reconciler never calls the driver's Replication gRPC; the controller-manager drives each member. +3. **Fan in.** Aggregate the members' `VolumeReplication.status.conditions` into the group's own. `Completed` is the conjunction (true only when at least one member exists and every member is `Completed=True`), and `Degraded` and `Resyncing` are the disjunction (one degraded or resyncing member makes the group so). `status.lastSyncTime` is the oldest of the members' `lastSyncTime`, because a group's recovery point is only as fresh as its slowest member. `status.state` mirrors `spec.replicationState` once `Completed`. `status.persistentVolumeClaimsRefList` records the resolved membership. + +### 14.4 Admission webhook + +`VolumeGroupReplicationValidator` (`operator/internal/webhook/volumegroupreplication_validator.go`) rejects, at create, a `VolumeGroupReplication` whose selector cannot resolve to one whole consistency group, so the reconciler never has to reconcile an object that could never fan out correctly. It is the sibling of `VolumeGroupSnapshotValidator` (`design-consistency-groups.md` §9.4) and makes the same two checks with the same dispositions: + +- **Ownership gate.** The object is validated only when `spec.external` is `true` and its class is attributed to `csi.simplyblock.io`. Every other case is admitted untouched, because `failurePolicy: fail` on a shared group means a webhook error would block foreign drivers' objects too. +- **Label check (fail-closed).** Every selected PVC must carry the same non-empty consistency-group label. A selector that spans groups, matches an unlabeled PVC, or matches nothing is denied. +- **Membership check (fail-open).** The selected lvol set must equal the backend group's membership. When membership cannot be determined (the backend is unreachable, or a selected PVC is not yet bound), the object is admitted and the reconciler backstops it (§14.3). + +### 14.5 The class and CR + +A `VolumeGroupReplicationClass` binds a group to this driver, and Ramen selects it by its `ramendr.openshift.io/groupreplicationid` label rather than by name: + +```yaml +apiVersion: replication.storage.openshift.io/v1alpha1 +kind: VolumeGroupReplicationClass +metadata: + name: simplyblock-group-async-5m + labels: + ramendr.openshift.io/groupreplicationid: simplyblock-async-5m +spec: + provisioner: csi.simplyblock.io + parameters: {} +``` + +```yaml +apiVersion: replication.storage.openshift.io/v1alpha1 +kind: VolumeGroupReplication +metadata: + name: app-group + namespace: team-a +spec: + external: true + replicationState: primary + volumeGroupReplicationClassName: simplyblock-group-async-5m + volumeReplicationClassName: simplyblock-async-5m + source: + selector: + matchLabels: + storage.simplyblock.io/consistency-group: app-group-cg +``` + +The per-member backend policy is not named on the group class. It comes from the group's `volumeReplicationClassName`, which the fan-out stamps on each member `VolumeReplication` and which resolves to a `ReplicationPolicy` exactly as a standalone `VolumeReplication` does (§7.1). + +### 14.6 Backend API + +Group replication adds no replication endpoint. The reconciler reads the consistency group to verify membership, and every promote, demote, and resync reaches the backend through the member `VolumeReplication` objects, which use the per-volume endpoints of §9. The two reads it makes: + +| Method | Endpoint | Notes | +|--------|--------------------------------------|--------------------------------------------------------------------| +| `GET` | `.../consistency-group` (by name) | Resolve the group named by the members' shared label. | +| `GET` | `.../consistency-group/{id}/members` | The group's current membership, compared against the selected set. | + +Both are consistency-group reads owned by `design-consistency-groups.md`; this design consumes them and adds none. + +### 14.7 Observability + +`VolumeGroupReplicationReconciler` carries an `events.EventRecorder` (wired in `operator/cmd/main.go` as `volumegroupreplication-controller`) and emits on the `VolumeGroupReplication` object: + +| Event | Type | Reason | +|---------------------------------------------------------------------|---------|----------------------------| +| The selector resolves to the consistency group's current membership | Normal | `GroupReplicationVerified` | +| The selector does not resolve to exactly one whole group | Warning | `GroupMembershipMismatch` | +| At least one member reports `Degraded`, on the transition into it | Warning | `GroupReplicationDegraded` | + +The per-member replication metrics of §11 already cover each group member, so the group adds no metric of its own: its aggregate state (`Completed`/`Degraded`/`Resyncing` and the oldest-member `lastSyncTime`) is derivable from the members' series. + +### 14.8 Scope + +The base case is one VRG, one storage vendor's PVCs, one `VolumeGroupReplication`. RamenDR's newer multi-VRG "Global VGR" consensus, spanning a replication group across several applications' VRGs, is out of scope (Open Question 4). `VolumeGroupReplicationContent` is unused: `Content` objects bind a group the generic controller-manager provisioned through a real CSI `GroupReplication` call, and the `external: true` fan-out never makes one (Open Question 5). + +--- + +## 15. Open Questions + +| # | Question | Owner | +|-----|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------| +| 1 | **Demote semantics for the application.** The P0-3 demote fences the volume (ANA inaccessible) after the final flush, and with convergence folded into the verb it is now the only place a planned swap can stall. This is also the one verb the planned promote's lossless guarantee entirely depends on (§5.2): a planned promote is refused unless a completed demote already fenced the source and confirmed the final delta landed, so an unresolved failure mode here is an unresolved gap in the whole "zero loss" claim. Ramen relocation unmounts the workload first, so the fence is ordinarily unopposed, but that is Ramen's choreography, not a guarantee the driver can rely on: a stuck termination, a stale mount that never released, or a demote invoked outside Ramen's normal flow can all leave writes still arriving when quiesce fires. Confirm the verb's behavior when writes are still in flight at quiesce (block versus fail), whether the converge phase has its own budget separate from the quiesced flush, and whether a timeout in either phase must abort back to serving primary or leave the volume fenced with no automatic recovery. | Backend team | +| 2 | **Per-volume policy granularity.** A `VolumeReplicationClass` names one policy, and today one policy implies one target and cadence for all its volumes. Confirm one class per (policy, cadence) is an acceptable authoring model for Ramen's `replicationClassSelector`, or whether per-volume interval overrides are needed. | Operator / Backend team | +| 3 | ~~**Avoiding the clone on day-one protection.**~~ **Resolved:** `POST .../replication/failover?planned=true`'s no-demote branch now checks `lvol_controller.replication_source_online` (the source's own storage-node status) before falling through to `FAILED_PRECONDITION` -- an online source is a no-op (§5.2), so a healthy volume's first-ever `PromoteVolume` no longer materializes a clone. The remaining residual: a source that dies within the last health-check interval still briefly reads online, so one reconcile can treat a genuine disaster as a no-op before the node's status catches up and the controller retries -- bounded by the health-check detection window, not open-ended. | Backend team | +| 4 | **Is Global VGR needed (§14.8).** The `VolumeGroupReplication` reconciler covers the base case: one VRG, one vendor's PVCs, one group. Confirm whether any planned simplyblock deployment spans a replication group across more than one application's VRG before treating RamenDR's multi-VRG "Global VGR" consensus as work this design should also specify. | Operator team | +| 5 | **`VolumeGroupReplicationContent`'s exact contract (§14.8).** The shipped reconciler writes nothing to it: `Content` objects bind a group the generic (non-external) controller-manager provisioned through a real CSI `GroupReplication` call, and the `external: true` path never makes one, fanning out to per-member `VolumeReplication` objects instead (§14.2). If a future Ramen version or tooling expects a `Content` object to exist even under `external: true`, revisit this. | Operator team | diff --git a/operator/docs/designs/design-ramen-integration.md b/operator/docs/designs/design-ramen-integration.md index 5fdb7b235..06fbdb737 100644 --- a/operator/docs/designs/design-ramen-integration.md +++ b/operator/docs/designs/design-ramen-integration.md @@ -85,7 +85,9 @@ The exact contract Ramen's pairing requires of the two clusters' objects (the id ## 4. VolumeGroupReplication (Implemented) -The gap analysis's own Phase 2 named this the piece consistency groups still owed: "Add `VolumeGroupReplication` (csi-addons) on top of the CG primitive so the VRG async group path can protect and fail over multi-volume apps at one point, not just snapshot them." `design-csi-addons-replication.md` §2 deferred it as "the next design," strictly per-volume itself. `design-consistency-groups.md` §2 deferred it too, as "future work," independent of any replication policy so that work could attach later. Neither claims it. This section is that attachment. +The gap analysis's own Phase 2 named this the piece consistency groups still owed: "Add `VolumeGroupReplication` (csi-addons) on top of the CG primitive so the VRG async group path can protect and fail over multi-volume apps at one point, not just snapshot them." `design-consistency-groups.md` §2 deferred it as "future work," independent of any replication policy so that work could attach later. This section is that attachment. + +**Authoritative specification.** `design-csi-addons-replication.md` §14 now carries the csi-addons group surface as the per-group half of that document's replication chapter, beside the per-volume `VolumeReplication` surface, and §14.1 … §14.8 are the reference for the reconciler, the admission webhook, the class, the backend reads, and the events. This section keeps the Ramen-integration framing (the offloaded `external` case in §4.1, and why no new gRPC verb is needed in §4.2) and the E2E validation (§6); where the two documents describe the same reconciler, §14 is authoritative. ### 4.1 What already exists, upstream diff --git a/operator/docs/tests/test-plan-csi-addons-replication.md b/operator/docs/tests/test-plan-csi-addons-replication.md index 7787f1e09..f4f0db7b2 100644 --- a/operator/docs/tests/test-plan-csi-addons-replication.md +++ b/operator/docs/tests/test-plan-csi-addons-replication.md @@ -6,7 +6,7 @@ Scope is the CSI driver's Replication service, the operator's preflight and coex Scenario IDs are permanent and are never reused or renumbered. `U-` is unit (no cluster: mock control plane, fake `client.Client`), `I-` is integration (the sidecar and controller-manager against the driver with a mock backend), `E-` is end-to-end (two live simplyblock clusters), and `M-` is manual. Types are `Positive`, `Negative`, `Boundary`, and `Regression`. A `—` in the `Test` column means nothing implements the scenario yet, and every such row reappears in §7 with its reason. -Phase 1 (the csi-addons machinery, §4, §5.1's three verbs, and §6's steady-state contract) and Phase 2 (§5.2's lifecycle verbs and P0-3) have both landed; their unit rows below are filled in. The operator's preflight and coexistence controllers (peerClasses, `PVCReplicationController`) remain a separate, unbuilt subsystem, and the integration and E2E tiers wait on a test bed neither phase has built yet. +Phase 1 (the csi-addons machinery, §4, §5.1's three verbs, and §6's steady-state contract) and Phase 2 (§5.2's lifecycle verbs and P0-3) have both landed; their unit rows below are filled in. Phase 4 (§14, `VolumeGroupReplication`) has landed too: its reconciler and admission webhook are unit-tested (U-57 … U-61), and its live group relocate is the E2E row E-08. The operator's preflight and coexistence controllers (peerClasses, `PVCReplicationController`) remain a separate, unbuilt subsystem, and the integration and E2E tiers wait on a test bed neither phase has built yet. --- @@ -112,6 +112,18 @@ File: `operator/internal/controllers/driver/workloads_test.go`, `operator/intern | U-38 | The sidecar's grant is a namespaced Role bound to the controller plugin's account, not a ClusterRole, since CSIAddonsNode is namespaced | Positive | `TestCSIAddonsRoleIsNamespacedAndBoundToTheControllerAccount` | | U-39 | The namespaced Role's rules are scoped to the sidecar's own job: its CSIAddonsNode and its own leader-election Lease | Positive | `TestCSIAddonsRoleRulesAreScopedToItsOwnJob` | +### VolumeGroupReplication: fan-in and membership (design §14) + +Files: `operator/internal/controller/volumegroupreplication_controller_unit_test.go`, `operator/internal/webhook/volumegroupreplication_validator_test.go`. These are the same scenarios [`test-plan-ramen-integration.md`](test-plan-ramen-integration.md) §1 tracks as its U-01 … U-05; the design they verify now lives in §14 of this plan's design doc, so the coverage is recorded here too. `VolumeGroupReplicationReconciler` fans a group's `replicationState` out to its members and their status back in (§14.3), and its admission webhook (§14.4) enforces the one-whole-group invariant. The reconciler calls no driver gRPC, so the fan-out itself is exercised only at the E2E tier (E-08). + +| # | Scenario | Type | Test | +|------|-------------------------------------------------------------------------------------------------------------------|----------|-------------------------------------------------------------------| +| U-57 | Every member's `VolumeReplication` reports `Completed=True, Degraded=False`, and the group reports the same | Positive | `TestVolumeGroupReplication_AllMembersHealthyYieldsGroupHealthy` | +| U-58 | One member reports `Degraded=True`, and the group reports `Degraded=True` (disjunction) | Negative | `TestVolumeGroupReplication_OneMemberDegradedYieldsGroupDegraded` | +| U-59 | Members report differing `lastSyncTime`, and `status.lastSyncTime` is the oldest, not the newest | Boundary | `TestVolumeGroupReplication_LastSyncTimeIsTheOldestMember` | +| U-60 | `spec.source.selector` resolves to exactly one consistency group's current membership, and is admitted | Positive | `TestVolumeGroupReplicationValidator` | +| U-61 | `spec.source.selector` resolves to a subset of a group, or spans two groups, and is rejected, naming the mismatch | Negative | `TestVolumeGroupReplicationValidator` | + --- ## 2. Integration Tests @@ -148,10 +160,11 @@ Two live simplyblock clusters with the chart-deployed csi-addons machinery. The ### Ramen-Driven (design §12, Phase 2 gate) -| # | Scenario | Type | Test | -|------|----------------------------------------------------------------------------------------------------------------------------------------------|----------|---------------------------------------------------------------------------------| -| E-06 | A Ramen `VolumeReplicationGroup` in async mode selects the class, creates one `VolumeReplication` per PVC, and `lastGroupSyncTime` populates | Positive | → [`test-plan-ramen-integration.md`](test-plan-ramen-integration.md) M-01 | -| E-07 | Ramen failover (`force`) and relocate (demote plus planned promote) both complete against a live workload | Positive | → [`test-plan-ramen-integration.md`](test-plan-ramen-integration.md) M-02, M-03 | +| # | Scenario | Type | Test | +|------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|----------|---------------------------------------------------------------------------------| +| E-06 | A Ramen `VolumeReplicationGroup` in async mode selects the class, creates one `VolumeReplication` per PVC, and `lastGroupSyncTime` populates | Positive | → [`test-plan-ramen-integration.md`](test-plan-ramen-integration.md) M-01 | +| E-07 | Ramen failover (`force`) and relocate (demote plus planned promote) both complete against a live workload | Positive | → [`test-plan-ramen-integration.md`](test-plan-ramen-integration.md) M-02, M-03 | +| E-08 | A Ramen VRG protects and relocates a multi-volume app through one `VolumeGroupReplication`: all members move together, none diverging, and the group's `status.lastSyncTime` tracks the oldest member (design §14) | Positive | → [`test-plan-ramen-integration.md`](test-plan-ramen-integration.md) M-05 | --- @@ -182,15 +195,16 @@ Two live simplyblock clusters with the chart-deployed csi-addons machinery. The ## 5. Axis Coverage -| Axis | Values covered | IDs | Not covered | -|------------------|------------------------------------------------------------------------|-----------------------------------------------------------------------------|-----------------------------------------------------------| -| Verb lifecycle | enable, disable, info, forced promote, planned promote, demote, resync | U-01, U-02, U-04 … U-10, U-12 … U-17, U-40 … U-56, I-02 … I-04, E-01 … E-04 | — | -| Idempotency | repeat enable, disable, demote; re-drive after restart | U-02, U-05, U-15, I-07 | repeated promote (U-11), repeated resync | -| Conditions | healthy, degraded, error, staleness, resyncing, disabled | U-18 … U-22, E-05 | condition behavior across backend upgrade | -| Coexistence | slot skip, one-owner refusal, concurrent claim | U-26, U-27, M-02 | migration of an annotated volume onto a VolumeReplication | -| peerClasses | verified, missing class, mispaired policies | U-23 … U-25 | drift after verification | -| Orchestrator | direct kubectl lifecycle, Ramen VRG async | I-02 … I-06, E-06, E-07 | Ramen hub failover of multiple apps | -| Cluster topology | two clusters, one relationship addressed from both sides | E-01 … E-07 | three-cluster (cascaded) topologies | +| Axis | Values covered | IDs | Not covered | +|-------------------|------------------------------------------------------------------------|-----------------------------------------------------------------------------|------------------------------------------------------------------------| +| Verb lifecycle | enable, disable, info, forced promote, planned promote, demote, resync | U-01, U-02, U-04 … U-10, U-12 … U-17, U-40 … U-56, I-02 … I-04, E-01 … E-04 | — | +| Idempotency | repeat enable, disable, demote; re-drive after restart | U-02, U-05, U-15, I-07 | repeated promote (U-11), repeated resync | +| Conditions | healthy, degraded, error, staleness, resyncing, disabled | U-18 … U-22, E-05 | condition behavior across backend upgrade | +| Coexistence | slot skip, one-owner refusal, concurrent claim | U-26, U-27, M-02 | migration of an annotated volume onto a VolumeReplication | +| peerClasses | verified, missing class, mispaired policies | U-23 … U-25 | drift after verification | +| Orchestrator | direct kubectl lifecycle, Ramen VRG async (per volume and per group) | I-02 … I-06, E-06, E-07, E-08 | Ramen hub failover of multiple apps | +| Group replication | fan-in aggregation, membership validation, live group relocate | U-57 … U-61, E-08 | Global (multi-VRG) VGR; the fan-out at the integration tier (no I-row) | +| Cluster topology | two clusters, one relationship addressed from both sides | E-01 … E-08 | three-cluster (cascaded) topologies | --- @@ -198,25 +212,25 @@ Two live simplyblock clusters with the chart-deployed csi-addons machinery. The | Class | Scenarios | Covered | Not covered | |-------------|-----------|---------|-------------------------| -| Unit | 46 | 34 | U-03, U-11, U-18 … U-27 | +| Unit | 51 | 39 | U-03, U-11, U-18 … U-27 | | Integration | 7 | 0 | I-01 … I-07 | -| E2E | 7 | 0 | E-01 … E-07 | +| E2E | 8 | 0 | E-01 … E-08 | | Manual | 2 | 0 | M-01, M-02 | -Phase 1 landed the driver's Replication and Identity services, the error classifier, and the operator's sidecar and RBAC wiring, covering every Phase 1 unit scenario except U-03 (§7). Phase 2 landed P0-3 (the demote endpoint and the planned gate on `failover`, in sbcli) and the driver's `PromoteVolume`/`DemoteVolume`/`ResyncVolume`, covering every Phase 2 unit scenario except U-11 (§7). The operator's preflight and coexistence controllers (peerClasses, `PVCReplicationController`) remain a separate, unbuilt subsystem, and neither a sidecar-and-controller-manager integration suite nor a live two-cluster E2E bed exists yet, so those tiers remain fully uncovered. +Phase 1 landed the driver's Replication and Identity services, the error classifier, and the operator's sidecar and RBAC wiring, covering every Phase 1 unit scenario except U-03 (§7). Phase 2 landed P0-3 (the demote endpoint and the planned gate on `failover`, in sbcli) and the driver's `PromoteVolume`/`DemoteVolume`/`ResyncVolume`, covering every Phase 2 unit scenario except U-11 (§7). Phase 4 landed the `VolumeGroupReplication` reconciler and its admission webhook, covering all of U-57 … U-61; its live group relocate (E-08) waits on the same Ramen bed as E-06/E-07. The operator's preflight and coexistence controllers (peerClasses, `PVCReplicationController`) remain a separate, unbuilt subsystem, and neither a sidecar-and-controller-manager integration suite nor a live two-cluster E2E bed exists yet, so those tiers remain fully uncovered. --- ## 7. What Is Not Yet Covered -| # | Gap | Reason | -|-------------|-------------------------------------------------------------------------------------------------------------------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| U-03 | Different-policy enable refused with `FAILED_PRECONDITION` | Not implemented: sbcli's `attach_policy` silently re-attaches onto the new policy rather than refusing, and no endpoint exposes the policy id a volume is currently attached to, for the driver to compare against before attaching (design §5.1 assumed this refusal exists; it does not against the current backend) | -| U-11 | Forced promote repeated after completion: success without a second failover | Not implemented: the backend's own idempotency (`_failover_volumes`' `already failed_over` skip) is real and exercised transitively, but no driver-level test asserts it directly for `PromoteVolume` | -| U-18 … U-22 | Condition derivation (`Completed`/`Degraded`/`Resyncing`) from the status read | Not implemented: `csi-addons/spec` v0.2.0's `GetVolumeReplicationInfoResponse` carries only `lastSyncTime`, with no per-condition field at all; deriving these needs either a newer spec version or belongs in the controller-manager's own reconcile, neither examined yet | -| U-23 … U-27 | Preflight (`peerClasses` verification) and coexistence (`PVCReplicationController`, the one-owner rule) | Out of Phase 1 and Phase 2's scope: the auto-adapter and preflight webhook are a separate, unbuilt subsystem | -| I-01 … I-07 | The sidecar and controller-manager loop | The driver's Replication and Identity services and the sidecar container now exist (Phase 1); no envtest/kind suite exercises them against the real kubernetes-csi-addons controller-manager yet | -| E-01 … E-05 | The live lifecycle (non-Ramen half) | Needs a two-cluster live test bed. `regression_test/21/` exercises the same lifecycle by hand but is not wired as an automated E2E suite. | -| E-06, E-07 | The Ramen-driven gate | Detailed scenario ownership moved to [`test-plan-ramen-integration.md`](test-plan-ramen-integration.md) M-01 … M-03, blocked there on a live OCM hub with Ramen installed (see that document's Phase 0) | -| — | Repeated resync, class drift after verification, annotated-volume migration onto the adapter, cascaded topologies | Beyond the first coverage pass, recorded so the gaps are explicit rather than assumed covered | -| M-01, M-02 | Demote under writes; concurrent ownership race | Need failure injection and precise timing a live two-cluster run does not automate yet | +| # | Gap | Reason | +|------------------|-------------------------------------------------------------------------------------------------------------------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| U-03 | Different-policy enable refused with `FAILED_PRECONDITION` | Not implemented: sbcli's `attach_policy` silently re-attaches onto the new policy rather than refusing, and no endpoint exposes the policy id a volume is currently attached to, for the driver to compare against before attaching (design §5.1 assumed this refusal exists; it does not against the current backend) | +| U-11 | Forced promote repeated after completion: success without a second failover | Not implemented: the backend's own idempotency (`_failover_volumes`' `already failed_over` skip) is real and exercised transitively, but no driver-level test asserts it directly for `PromoteVolume` | +| U-18 … U-22 | Condition derivation (`Completed`/`Degraded`/`Resyncing`) from the status read | Not implemented: `csi-addons/spec` v0.2.0's `GetVolumeReplicationInfoResponse` carries only `lastSyncTime`, with no per-condition field at all; deriving these needs either a newer spec version or belongs in the controller-manager's own reconcile, neither examined yet | +| U-23 … U-27 | Preflight (`peerClasses` verification) and coexistence (`PVCReplicationController`, the one-owner rule) | Out of Phase 1 and Phase 2's scope: the auto-adapter and preflight webhook are a separate, unbuilt subsystem | +| I-01 … I-07 | The sidecar and controller-manager loop | The driver's Replication and Identity services and the sidecar container now exist (Phase 1); no envtest/kind suite exercises them against the real kubernetes-csi-addons controller-manager yet | +| E-01 … E-05 | The live lifecycle (non-Ramen half) | Needs a two-cluster live test bed. `regression_test/21/` exercises the same lifecycle by hand but is not wired as an automated E2E suite. | +| E-06, E-07, E-08 | The Ramen-driven gate (per-volume and per-group) | Detailed scenario ownership moved to [`test-plan-ramen-integration.md`](test-plan-ramen-integration.md) M-01 … M-03 (per volume) and M-05 (the `VolumeGroupReplication` group relocate), blocked there on a live OCM hub with Ramen installed (see that document's Phase 0) | +| — | Repeated resync, class drift after verification, annotated-volume migration onto the adapter, cascaded topologies | Beyond the first coverage pass, recorded so the gaps are explicit rather than assumed covered | +| M-01, M-02 | Demote under writes; concurrent ownership race | Need failure injection and precise timing a live two-cluster run does not automate yet | diff --git a/operator/docs/tests/test-plan-ramen-integration.md b/operator/docs/tests/test-plan-ramen-integration.md index 4e96c5503..92059df6a 100644 --- a/operator/docs/tests/test-plan-ramen-integration.md +++ b/operator/docs/tests/test-plan-ramen-integration.md @@ -11,7 +11,7 @@ Scope: two things this document specifies. First, whether `VolumeGroupReplicatio ### VolumeGroupReplicationReconciler and its admission webhook (design §4) -Implemented in `volumegroupreplication_controller.go` and `volumegroupreplication_validator.go`, covered by `volumegroupreplication_controller_unit_test.go` and `volumegroupreplication_validator_test.go`. +Implemented in `volumegroupreplication_controller.go` and `volumegroupreplication_validator.go`, covered by `volumegroupreplication_controller_unit_test.go` and `volumegroupreplication_validator_test.go`. Since the reconciler's design now lives in `design-csi-addons-replication.md` §14 (the csi-addons group surface), these same scenarios are mirrored in [`test-plan-csi-addons-replication.md`](test-plan-csi-addons-replication.md) as its U-57 … U-61, and the group's live relocate as that plan's E-08. | # | Scenario | Type | Test | |------|-------------------------------------------------------------------------------------------------------------------|----------|-------------------------------------------------------------------| From 0a59a974695f7d684fd0e9d9037952fbe3723817 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Fri, 25 Sep 2026 23:39:56 +0100 Subject: [PATCH 157/206] refactor(operator): remove offloaded VolumeGroupReplication reconciler and validator --- .../templates/roles/manager_role.yaml | 38 -- .../simplyblock-operator-webhook.yaml | 19 - operator/cmd/main.go | 15 - operator/config/rbac/role.yaml | 38 -- operator/config/webhook/manifests.yaml | 19 - .../designs/design-csi-addons-replication.md | 175 ++++--- .../docs/designs/design-ramen-integration.md | 81 +-- .../tests/test-plan-csi-addons-replication.md | 65 ++- .../docs/tests/test-plan-ramen-integration.md | 67 +-- .../volumegroupreplication_controller.go | 487 ------------------ ...megroupreplication_controller_unit_test.go | 293 ----------- .../volumegroupreplication_validator.go | 201 -------- .../volumegroupreplication_validator_test.go | 193 ------- 13 files changed, 188 insertions(+), 1503 deletions(-) delete mode 100644 operator/internal/controller/volumegroupreplication_controller.go delete mode 100644 operator/internal/controller/volumegroupreplication_controller_unit_test.go delete mode 100644 operator/internal/webhook/volumegroupreplication_validator.go delete mode 100644 operator/internal/webhook/volumegroupreplication_validator_test.go diff --git a/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml b/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml index 52e8928d8..13c23aeab 100644 --- a/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml @@ -254,44 +254,6 @@ rules: - patch - update - watch -- apiGroups: - - replication.storage.openshift.io - resources: - - volumegroupreplicationclasses - verbs: - - get - - list - - watch -- apiGroups: - - replication.storage.openshift.io - resources: - - volumegroupreplications - verbs: - - get - - list - - patch - - update - - watch -- apiGroups: - - replication.storage.openshift.io - resources: - - volumegroupreplications/status - verbs: - - get - - patch - - update -- apiGroups: - - replication.storage.openshift.io - resources: - - volumereplications - verbs: - - create - - delete - - get - - list - - patch - - update - - watch - apiGroups: - snapshot.storage.k8s.io resources: diff --git a/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml b/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml index 9c738e997..af8c2c450 100644 --- a/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator-webhook.yaml @@ -380,25 +380,6 @@ webhooks: resources: - storagepools sideEffects: None -- admissionReviewVersions: - - v1 - clientConfig: - service: - name: simplyblock-operator-webhook-service - namespace: {{ .Release.Namespace }} - path: /validate-replication-storage-openshift-io-v1alpha1-volumegroupreplication - failurePolicy: Fail - name: vvolumegroupreplication.simplyblock.io - rules: - - apiGroups: - - replication.storage.openshift.io - apiVersions: - - v1alpha1 - operations: - - CREATE - resources: - - volumegroupreplications - sideEffects: None - admissionReviewVersions: - v1 clientConfig: diff --git a/operator/cmd/main.go b/operator/cmd/main.go index b8eddf49e..313b386f6 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -875,14 +875,6 @@ func main() { setupLog.Error(err, "unable to create controller", "controller", "ReplicationPair") os.Exit(1) } - if err := (&controller.VolumeGroupReplicationReconciler{ - Client: mgr.GetClient(), - Scheme: mgr.GetScheme(), - Recorder: mgr.GetEventRecorder("volumegroupreplication-controller"), - }).SetupWithManager(mgr); err != nil { - setupLog.Error(err, "unable to create controller", "controller", "VolumeGroupReplication") - os.Exit(1) - } if err := (&controller.ReplicationSlotReconciler{ Client: mgr.GetClient(), Scheme: mgr.GetScheme(), @@ -1028,13 +1020,6 @@ func main() { }}) setupLog.Info("registered volumegroupsnapshot validating webhook") - mgr.GetWebhookServer().Register("/validate-replication-storage-openshift-io-v1alpha1-volumegroupreplication", - &webhook.Admission{Handler: &internalwebhook.VolumeGroupReplicationValidator{ - Client: mgr.GetClient(), - APIClient: webapi.NewClient(), - }}) - setupLog.Info("registered volumegroupreplication validating webhook") - mgr.GetWebhookServer().Register("/validate-storage-simplyblock-io-v1alpha2-volumegroupsnapshotops", &webhook.Admission{Handler: &internalwebhook.VolumeGroupSnapshotOpsValidator{Client: mgr.GetClient()}}) setupLog.Info("registered volumegroupsnapshotops validating webhook") diff --git a/operator/config/rbac/role.yaml b/operator/config/rbac/role.yaml index 7bc3aabe3..50dc7ce64 100644 --- a/operator/config/rbac/role.yaml +++ b/operator/config/rbac/role.yaml @@ -254,44 +254,6 @@ rules: - patch - update - watch -- apiGroups: - - replication.storage.openshift.io - resources: - - volumegroupreplicationclasses - verbs: - - get - - list - - watch -- apiGroups: - - replication.storage.openshift.io - resources: - - volumegroupreplications - verbs: - - get - - list - - patch - - update - - watch -- apiGroups: - - replication.storage.openshift.io - resources: - - volumegroupreplications/status - verbs: - - get - - patch - - update -- apiGroups: - - replication.storage.openshift.io - resources: - - volumereplications - verbs: - - create - - delete - - get - - list - - patch - - update - - watch - apiGroups: - snapshot.storage.k8s.io resources: diff --git a/operator/config/webhook/manifests.yaml b/operator/config/webhook/manifests.yaml index 8ba121344..d39ccba07 100644 --- a/operator/config/webhook/manifests.yaml +++ b/operator/config/webhook/manifests.yaml @@ -362,25 +362,6 @@ webhooks: resources: - storagepools sideEffects: None -- admissionReviewVersions: - - v1 - clientConfig: - service: - name: webhook-service - namespace: system - path: /validate-replication-storage-openshift-io-v1alpha1-volumegroupreplication - failurePolicy: Fail - name: vvolumegroupreplication.simplyblock.io - rules: - - apiGroups: - - replication.storage.openshift.io - apiVersions: - - v1alpha1 - operations: - - CREATE - resources: - - volumegroupreplications - sideEffects: None - admissionReviewVersions: - v1 clientConfig: diff --git a/operator/docs/designs/design-csi-addons-replication.md b/operator/docs/designs/design-csi-addons-replication.md index d304b860d..8b773548b 100644 --- a/operator/docs/designs/design-csi-addons-replication.md +++ b/operator/docs/designs/design-csi-addons-replication.md @@ -1,6 +1,6 @@ # Design Document: csi-addons Volume Replication -**Status:** Phase 4 Implemented +**Status:** Phase 3 Implemented (Phase 4, group replication: Planned) **Author:** Israel Geoffrey (geoffrey1330) **Date:** 2026-09-16 (last updated 2026-09-25) **Test Plan:** [`tests/test-plan-csi-addons-replication.md`](../tests/test-plan-csi-addons-replication.md) @@ -9,14 +9,14 @@ ## Phasing Overview -| Phase | Status | Scope | Sections | -|-------------|-------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|--------------| -| **Phase 1** | Implemented | The csi-addons machinery and the steady-state contract: CRDs, controller-manager, sidecar, the Replication and csi-addons Identity gRPC services with `EnableVolumeReplication`, `DisableVolumeReplication`, and `GetVolumeReplicationInfo`, backed by a typed backend status endpoint | §4, §5.1, §6 | -| **Phase 2** | Implemented | The lifecycle verbs: `PromoteVolume` (planned and forced), `DemoteVolume`, and `ResyncVolume`. Validation end to end against a Ramen `VolumeReplicationGroup` in async mode is still outstanding (§12, E-06/E-07) | §5.2, §9 | -| **Phase 3** | Implemented | §11 (the Prometheus metrics). §7.2's peerClasses preflight is out of this design's scope entirely (it's Ramen's own `DRPolicy` mechanism) and is deferred to a future Ramen-integration design | §7.1, §11 | -| **Phase 4** | Implemented | §14 (`VolumeGroupReplication`): the group-level csi-addons surface on top of the consistency-group primitive, fanning the §5 verbs out to a group's members. This is the gap analysis's own Phase 2 group-replication item | §14 | +| Phase | Status | Scope | Sections | +|-------------|-------------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|--------------| +| **Phase 1** | Implemented | The csi-addons machinery and the steady-state contract: CRDs, controller-manager, sidecar, the Replication and csi-addons Identity gRPC services with `EnableVolumeReplication`, `DisableVolumeReplication`, and `GetVolumeReplicationInfo`, backed by a typed backend status endpoint | §4, §5.1, §6 | +| **Phase 2** | Implemented | The lifecycle verbs: `PromoteVolume` (planned and forced), `DemoteVolume`, and `ResyncVolume`. Validation end to end against a Ramen `VolumeReplicationGroup` in async mode is still outstanding (§12, E-06/E-07) | §5.2, §9 | +| **Phase 3** | Implemented | §11 (the Prometheus metrics). §7.2's peerClasses preflight is out of this design's scope entirely (it's Ramen's own `DRPolicy` mechanism) and is deferred to a future Ramen-integration design | §7.1, §11 | +| **Phase 4** | Planned | §14 (`VolumeGroupReplication`): the group-level csi-addons surface, driven by the stock controller-manager through a new driver VolumeGroup service and a group handle the existing Replication verbs act on, backed by new group-replication endpoints (P0-6/P0-7). This is the gap analysis's own Phase 2 group-replication item | §14 | -Phase 1 is independently useful: a `VolumeReplication` object per PVC whose status truthfully reports the relationship, which no surface provides today. Phase 2 makes the object drivable, which is what Ramen actually needs. Phase 3 makes the whole thing operable at fleet scale. Phase 4 lifts the same surface from one volume to a consistency group. +Phase 1 is independently useful: a `VolumeReplication` object per PVC whose status truthfully reports the relationship, which no surface provides today. Phase 2 makes the object drivable, which is what Ramen actually needs. Phase 3 makes the whole thing operable at fleet scale. Phase 4 lifts the same surface from one volume to a consistency group, replicated as one unit. The phase numbers above are this document's own, not the DR storage foundation gap analysis's (§1): its Phase 0 (shipping the csi-addons contract itself) is this design's Phase 1, and its Phase 1 (promote, demote, and resync end to end through Ramen) is this design's Phase 2. @@ -24,15 +24,17 @@ The phase numbers above are this document's own, not the DR storage foundation g ## Phase 0 — External Prerequisites -| # | Prerequisite | Kind | Blocks | Status | -|------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------|---------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| P0-1 | A typed, steady-state per-volume replication status read: `GET .../volumes/{id}/replication/status` serving what `lvol_controller.get_replication_info` computes today (state, lag, outstanding bytes, failure counters), available for the volume's whole replicated life | Control plane (`sbcli`) | Phase 1 | Shipped: `GET .../volumes/{v}/replication/status` → `ReplicationStatusDTO` (`simplyblock_web/api/v2/cluster/storage_pool/volume/replication.py:58-69`) | -| P0-2 | Idempotent attach and detach: attaching a volume to the policy it already follows returns success, and detaching a non-attached volume returns success | Control plane (`sbcli`) | Phase 1 | Shipped: `replication_policy_controller.attach_policy`/`detach_policy` (`simplyblock_core/controllers/replication_policy_controller.py:208-262`) | -| P0-3 | A standalone demote verb: `POST .../volumes/{id}/replication/demote` that converges the peer while still serving (repeated snapshot-and-ship until the remaining delta is small), then quiesces, ships the final delta, confirms it landed on the peer, and fences the data path | Control plane (`sbcli`) | Phase 2 | Shipped: `POST .../volumes/{v}/replication/demote` → `lvol_controller.demote_lvol` | -| P0-4 | An `rpo_target_seconds` field on `ReplicationPolicy`, so RPO compliance is computable against a declared target rather than the derived lag budget | Control plane (`sbcli`) | Phase 3 | Shipped: `ReplicationPolicy.rpo_target_seconds` (`simplyblock_core/models/replication.py:89`), wired through the API (`PolicyParams.rpo_target_seconds`) and CLI (`--rpo-target-sec`) | -| P0-5 | csi-addons upstream: the `VolumeReplication` and `VolumeReplicationClass` CRDs (`replication.storage.openshift.io/v1alpha1`), the kubernetes-csi-addons controller-manager image, and the csi-addons sidecar image | Ecosystem | Phase 1 | Vendored in the chart at v0.15.0 behind `csiaddons.create` (all twelve upstream CRDs, since the stock manager starts a controller per kind); sidecar wiring is Phase 1 | +| # | Prerequisite | Kind | Blocks | Status | +|------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------|---------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| P0-1 | A typed, steady-state per-volume replication status read: `GET .../volumes/{id}/replication/status` serving what `lvol_controller.get_replication_info` computes today (state, lag, outstanding bytes, failure counters), available for the volume's whole replicated life | Control plane (`sbcli`) | Phase 1 | Shipped: `GET .../volumes/{v}/replication/status` → `ReplicationStatusDTO` (`simplyblock_web/api/v2/cluster/storage_pool/volume/replication.py:58-69`) | +| P0-2 | Idempotent attach and detach: attaching a volume to the policy it already follows returns success, and detaching a non-attached volume returns success | Control plane (`sbcli`) | Phase 1 | Shipped: `replication_policy_controller.attach_policy`/`detach_policy` (`simplyblock_core/controllers/replication_policy_controller.py:208-262`) | +| P0-3 | A standalone demote verb: `POST .../volumes/{id}/replication/demote` that converges the peer while still serving (repeated snapshot-and-ship until the remaining delta is small), then quiesces, ships the final delta, confirms it landed on the peer, and fences the data path | Control plane (`sbcli`) | Phase 2 | Shipped: `POST .../volumes/{v}/replication/demote` → `lvol_controller.demote_lvol` | +| P0-4 | An `rpo_target_seconds` field on `ReplicationPolicy`, so RPO compliance is computable against a declared target rather than the derived lag budget | Control plane (`sbcli`) | Phase 3 | Shipped: `ReplicationPolicy.rpo_target_seconds` (`simplyblock_core/models/replication.py:89`), wired through the API (`PolicyParams.rpo_target_seconds`) and CLI (`--rpo-target-sec`) | +| P0-5 | csi-addons upstream: the `VolumeReplication` and `VolumeReplicationClass` CRDs (`replication.storage.openshift.io/v1alpha1`), the kubernetes-csi-addons controller-manager image, and the csi-addons sidecar image | Ecosystem | Phase 1 | Vendored in the chart at v0.15.0 behind `csiaddons.create` (all twelve upstream CRDs, since the stock manager starts a controller per kind); sidecar wiring is Phase 1 | +| P0-6 | Group-level replication in the control plane: a consistency group replicated as one unit, with attach/detach to a group replication policy, group failover, group demote, group failback, and a typed group status, using group snapshots (`bdev_lvol_snapshot_group`) as the recovery generations (§14.4, §14.5) | Control plane (`sbcli`) | Phase 4 | Not shipped. New group-replication engine and the `/consistency-groups/{id}/replication/*` endpoints | +| P0-7 | A group replication policy: cadence and target for a whole consistency group, the group twin of `ReplicationPolicy` (Open Question 7) | Control plane (`sbcli`) | Phase 4 | Not shipped | -Everything else the adapter needs already exists: the attach and detach calls, failover, the failback and commit pair, the relationship read, and the backlog arithmetic inside `get_replication_info`. The adapter is thin precisely because the engine is complete. What is missing is the shape Ramen can drive. +Everything the per-volume adapter (Phases 1-3) needs already exists: the attach and detach calls, failover, the failback and commit pair, the relationship read, and the backlog arithmetic inside `get_replication_info`. That adapter is thin precisely because the engine is complete. Group replication (Phase 4) is the exception: P0-6/P0-7 are genuinely new backend work, because the engine replicates per volume today and a consistency group has no group-level replication of its own (§14.4). --- @@ -176,7 +178,8 @@ A reader who stops here has the model: the engine is unchanged, the csi-addons s The plugin registers two additional gRPC services on the existing socket, beside the CSI services, following the GroupController precedent in `csicommon.NonBlockingGRPCServer`: - **csi-addons Identity:** `GetIdentity`, `GetCapabilities` (advertising `VOLUME_REPLICATION`), and `Probe`. This is a distinct service from CSI Identity, so it is a new small server type, not an extension of the existing one. -- **Replication:** the six verbs of §5, implemented on the controller `Server` through an embedded `replication.UnimplementedControllerServer` from `github.com/csi-addons/spec` (pinned at v0.2.0), mirroring how the GroupController embeds its unimplemented base. +- **Replication:** the six verbs of §5, implemented on the controller `Server` through an embedded `replication.UnimplementedControllerServer` from `github.com/csi-addons/spec` (pinned at v0.2.0), mirroring how the GroupController embeds its unimplemented base. The same verbs accept a group handle for group replication (§14.4). +- **csi-addons VolumeGroup (GroupController)** (Phase 4, Planned): `CreateVolumeGroup`, `ModifyVolumeGroupMembership`, and `DeleteVolumeGroup`, mapping a set of volume handles to the backend consistency group and back (§14.3). csi-addons Identity advertises the capability so the stock controller-manager dials it. `NonBlockingGRPCServer.Start` today takes exactly the three CSI servers. It gains a registration hook so the driver package can register additional services without `csicommon` importing csi-addons. @@ -368,7 +371,9 @@ Values are computed by `lvol_controller.get_replication_info_bulk`, a bulk-frien Full scenario matrix and coverage status: [`tests/test-plan-csi-addons-replication.md`](../tests/test-plan-csi-addons-replication.md) - **Unit (driver):** each verb against a mock control plane: the idempotency table (repeat enable, repeat disable, repeat promote), the refusal paths (different-policy enable, lagging planned promote, disable during cutover), the condition derivation from every status-read state, and handle parsing failures. -- **Unit (operator):** the `PVCAnnotationWatcher` skip when a `VolumeReplication` exists; and, for §14, `VolumeGroupReplicationReconciler`'s fan-in aggregation (group `Completed`/`Degraded`/`Resyncing` from the members, oldest-member `lastSyncTime`) and the admission webhook's membership check, against a fake client. +- **Unit (operator):** the `PVCAnnotationWatcher` skip when a `VolumeReplication` exists. +- **Unit (driver, §14):** the VolumeGroup service verbs against a mock control plane (`CreateVolumeGroup` resolves the label-formed group idempotently; `ModifyVolumeGroupMembership` refused off-placement; `DeleteVolumeGroup` leaves members), and the group-handle branch of each Replication verb routing to the group endpoints. +- **Unit (backend, §14.5):** the group-replication engine in `sbcli` (group failover clones the last group generation for every member atomically; group demote quiesces all then ships one final group snapshot; group failback reverses direction), tests-first. - **Integration:** the csi-addons sidecar and controller-manager against the driver with a mock backend under envtest or kind: a `VolumeReplication` flipped `primary` to `secondary` and back walks the verbs in order and lands the conditions. - **E2E (two live clusters):** the Ramen-shaped lifecycle without Ramen: enable on the source, write data, and verify `lastSyncTime` advances; forced promote on the DR side, verifying the clone serves with the source fenced; and resync back with a planned swap (demote then promote), verifying zero loss with a hashed writer. Then the same driven by an actual Ramen VRG in async mode, which is Phase 2's acceptance gate. @@ -384,104 +389,103 @@ Three replication control surfaces exist today: the operator's kinds, the stale 2. **Opportunistic:** the operator's replication reconcilers move their inline HTTP calls onto the shared client, and the slot controller sources `lastReplicatedAt` from the typed status read. No behavior change, one client. 3. **After Ramen validation:** the annotation path is declared a compatibility layer, and new volumes are protected through `VolumeReplication`. The stale CSI-shipped CRDs are deleted (they were never installable), and `SnapshotReplication` is retired from the charts. 4. **The redesign's replication chapter:** whether `ReplicationPair` and `ReplicationPolicy` survive as the backend-policy authoring surface (classes need policies to name) or are re-cut is decided there, not here. This design only requires that a policy exists per cluster pair, however it is authored. +5. **Removing the offloaded group reconciler.** An earlier revision implemented `VolumeGroupReplication` as an operator-owned, `external: true` path: `VolumeGroupReplicationReconciler` (`operator/internal/controller/volumegroupreplication_controller.go`) and `VolumeGroupReplicationValidator` (`operator/internal/webhook/volumegroupreplication_validator.go`), which fanned a group out to per-member `VolumeReplication` objects. §14 supersedes that with the stock, `external: false` driver path, so both files, their unit tests, the `main.go` wiring, and the webhook registration are removed. Group protection then rides the driver's VolumeGroup service and the group-replication endpoints (P0-6/P0-7), with no operator code on the path. --- ## 14. VolumeGroupReplication -`VolumeGroupReplication` extends this adapter from one volume to a consistency group. It is the group-level sibling of the `VolumeReplication` surface (§5): the same csi-addons `replication.storage.openshift.io` API group, the same `replicationState` intent, and the same gRPC verbs (§5.2), fanned out to every member of a consistency group rather than driven per volume. This is the async **group** path the gap analysis's Phase 2 named ("`VolumeGroupReplication` on top of the CG primitive so the VRG async group path can protect and fail over multi-volume apps at one point"), and it is implemented. +`VolumeGroupReplication` protects a consistency group as one replicated unit. It is the group-level sibling of the `VolumeReplication` surface (§5), in the same csi-addons `replication.storage.openshift.io` API group, and it is driven entirely by the stock kubernetes-csi-addons machinery: the generic controller-manager forms a backend volume group through the driver's csi-addons **VolumeGroup** service, then replicates that group as a single unit through the Replication service (§5), addressing it by one group handle. No operator reconciler and no vendor-specific controller take part. This is the gap analysis's Phase 2 group-replication item, and it is **Planned** (§13 records the earlier attempt this supersedes). -Ramen's `VolumeReplicationGroup` creates one `VolumeGroupReplication` when a protected application's PVCs share a `VolumeGroupReplicationClass` carrying a `ramendr.openshift.io/groupreplicationid` label, the group-level sibling of the per-volume `replicationid` label of §7.1. The reconciler and validator specified here own that object. The driver's Replication gRPC (§5) is unchanged, and no new verb is added. The end-to-end Ramen validation of this path is `design-ramen-integration.md` §6, whose test plan carries its E2E scenario as M-05. +Ramen's `VolumeReplicationGroup` creates one `VolumeGroupReplication` when a protected application's PVCs share a `StorageClass` carrying `ramendr.openshift.io/groupreplicationid` (the group-level sibling of the per-volume `replicationid` of §7.1). The end-to-end Ramen validation is `design-ramen-integration.md` §6, whose test plan carries its E2E scenario as M-05. -### 14.1 Upstream kinds and ownership +### 14.1 Why the driver, not the operator -`VolumeGroupReplication`, `VolumeGroupReplicationClass`, and `VolumeGroupReplicationContent` are upstream kubernetes-csi-addons CRDs in the same `replication.storage.openshift.io` group this design already vendors `VolumeReplication` and `VolumeReplicationClass` from (§4.1). This repository declares no Go type for them, and the reconciler and validator read and write them as `unstructured`, matching how the per-volume kinds are handled. +csi-addons's `VolumeGroupReplication` controller (v0.15.0) branches on `spec.external`: -`VolumeGroupReplication.spec` carries `replicationState` (`primary`/`secondary`/`resync`), a `source.selector` naming the member PVCs by label, a `volumeGroupReplicationClassName`, a `volumeReplicationClassName` (the per-member class the fan-out stamps on each member `VolumeReplication`), and an `external` boolean. `external: false` routes the object through the generic kubernetes-csi-addons controller-manager, which fans it out to member `VolumeReplication` objects itself. `external: true` hands it to a vendor controller. simplyblock is the `external: true`, "offloaded" case: the backend replicates below Kubernetes, so this operator performs the fan-out. +- **`external: false` (this design's path).** The stock controller creates a `VolumeGroupReplicationContent`, calls the driver's VolumeGroup service to group the member volumes and obtain a group handle, then creates **one** `VolumeReplication` addressed by that group handle, which the ordinary `VolumeReplication` controller drives through the driver's Replication service (§5). The whole group is one replicated object, exactly as a single volume is. +- **`external: true`.** The stock controller skips the object entirely, handing it to a vendor controller. -`VolumeGroupReplicationReconciler` (`operator/internal/controller/volumegroupreplication_controller.go`) owns a `VolumeGroupReplication` only when both `spec.external` is `true` and `spec.volumeGroupReplicationClassName` names a `VolumeGroupReplicationClass` whose `spec.provisioner` is `csi.simplyblock.io`. Every other object (non-external, or a foreign provisioner's class) is skipped, because `replication.storage.openshift.io` is a shared group and the generic controller-manager or another vendor may be the right owner. +This design takes the `external: false` path: the group is a first-class backend object the driver creates, and the standard machinery replicates it. The alternative, an operator `VolumeGroupReplicationReconciler` owning `external: true` objects and fanning them out to per-member `VolumeReplication` objects, is not used. It re-implements in the operator what the csi-addons controller already does, and it drives the group through per-member calls rather than one group operation, losing the crash-consistent, all-at-one-point failover a consistency group exists to give. An earlier revision built that reconciler and its admission webhook; §13 removes them. -### 14.2 No new gRPC contract +### 14.2 Ramen configuration -Group promote, demote, and resync are the same three verbs of §5.2 (`PromoteVolume`, `DemoteVolume`, `ResyncVolume`) fanned out to every member, not a fourth verb on the driver. Each member is an ordinary `VolumeReplication` object, reconciled by the already-shipped kubernetes-csi-addons controller-manager exactly as it reconciles any Ramen-created per-volume object (§5). `external: true` changes who performs the fan-out (this operator, rather than the generic manager), not what the fan-out does. - -### 14.3 The reconciler - -`VolumeGroupReplicationReconciler` reconciles each owned `VolumeGroupReplication` on a 60-second resync (30 seconds on a transient error): - -1. **Resolve membership.** Read `spec.source.selector` and list the matching PVCs in the object's namespace. Every matched PVC must carry the same non-empty `storage.simplyblock.io/consistency-group` label; a selector that matches an unlabeled PVC, spans two group values, or matches nothing is a membership mismatch. Map each PVC through its PV handle (`{clusterID}:{poolID}:{volumeID}`) to a backend lvol, read the consistency group named by the label, and require the selected lvol set to equal the group's current backend membership exactly, the same invariant `design-consistency-groups.md` §9.2 established for `VolumeGroupSnapshot`. A selected PVC that is not yet bound makes membership undeterminable, which requeues quietly rather than failing. -2. **Fan out.** For each member PVC, ensure a per-volume `VolumeReplication` exists, owned by the group, named `-`, with `spec.replicationState` mirroring the group's, `spec.volumeReplicationClass` set to the group's `volumeReplicationClassName`, `spec.autoResync: false`, and a `dataSource` naming the member PVC. An existing member whose `replicationState` has drifted from the group's is updated. The reconciler never calls the driver's Replication gRPC; the controller-manager drives each member. -3. **Fan in.** Aggregate the members' `VolumeReplication.status.conditions` into the group's own. `Completed` is the conjunction (true only when at least one member exists and every member is `Completed=True`), and `Degraded` and `Resyncing` are the disjunction (one degraded or resyncing member makes the group so). `status.lastSyncTime` is the oldest of the members' `lastSyncTime`, because a group's recovery point is only as fresh as its slowest member. `status.state` mirrors `spec.replicationState` once `Completed`. `status.persistentVolumeClaimsRefList` records the resolved membership. - -### 14.4 Admission webhook - -`VolumeGroupReplicationValidator` (`operator/internal/webhook/volumegroupreplication_validator.go`) rejects, at create, a `VolumeGroupReplication` whose selector cannot resolve to one whole consistency group, so the reconciler never has to reconcile an object that could never fan out correctly. It is the sibling of `VolumeGroupSnapshotValidator` (`design-consistency-groups.md` §9.4) and makes the same two checks with the same dispositions: - -- **Ownership gate.** The object is validated only when `spec.external` is `true` and its class is attributed to `csi.simplyblock.io`. Every other case is admitted untouched, because `failurePolicy: fail` on a shared group means a webhook error would block foreign drivers' objects too. -- **Label check (fail-closed).** Every selected PVC must carry the same non-empty consistency-group label. A selector that spans groups, matches an unlabeled PVC, or matches nothing is denied. -- **Membership check (fail-open).** The selected lvol set must equal the backend group's membership. When membership cannot be determined (the backend is unreachable, or a selected PVC is not yet bound), the object is admitted and the reconciler backstops it (§14.3). - -### 14.5 The class and CR - -A `VolumeGroupReplicationClass` binds a group to this driver, and Ramen selects it by its `ramendr.openshift.io/groupreplicationid` label rather than by name: +Ramen sets `spec.external` from whether the member PVCs' `StorageClass` carries `ramendr.openshift.io/offloaded`. For this driver-driven design the `StorageClass` carries `groupreplicationid` but **not** `offloaded`, so Ramen creates the object with `external: false` and a real `volumeReplicationClassName`, and the stock controller reconciles it. ```yaml +apiVersion: storage.k8s.io/v1 +kind: StorageClass +metadata: + name: simplyblock-group-sc + labels: + ramendr.openshift.io/storageid: # differs per cluster + ramendr.openshift.io/replicationid: # names the per-volume relationship + ramendr.openshift.io/groupreplicationid: # triggers group replication; no `offloaded` label +provisioner: csi.simplyblock.io +parameters: { cluster_id: , pool_name: } # plus the usual StorageClass parameters +--- apiVersion: replication.storage.openshift.io/v1alpha1 kind: VolumeGroupReplicationClass metadata: name: simplyblock-group-async-5m labels: - ramendr.openshift.io/groupreplicationid: simplyblock-async-5m + ramendr.openshift.io/groupreplicationid: # must equal the StorageClass's + ramendr.openshift.io/storageid: # must equal the StorageClass's spec: provisioner: csi.simplyblock.io - parameters: {} + parameters: + schedulingInterval: "5m" # must equal the DRPolicy's ``` -```yaml -apiVersion: replication.storage.openshift.io/v1alpha1 -kind: VolumeGroupReplication -metadata: - name: app-group - namespace: team-a -spec: - external: true - replicationState: primary - volumeGroupReplicationClassName: simplyblock-group-async-5m - volumeReplicationClassName: simplyblock-async-5m - source: - selector: - matchLabels: - storage.simplyblock.io/consistency-group: app-group-cg -``` +Member PVCs carry `storage.simplyblock.io/consistency-group` (backend grouping and group snapshots, `design-consistency-groups.md`) and the `pvcSelector` label the VRG matches. Ramen additionally reads its own `ramendr.openshift.io/consistency-group` label to build the group's selector, so a Ramen-protected member carries both consistency-group labels with the same value. + +### 14.3 The VolumeGroup service (driver) + +The plugin serves the csi-addons **VolumeGroup** (GroupController) service beside csi-addons Identity and Replication (§4.2), advertising the capability through csi-addons Identity so the stock controller-manager dials it. Its verbs map onto the backend consistency group that already exists as a first-class object (`design-consistency-groups.md`): + +| Verb | Backend mapping | Semantics | +|-------------------------------|------------------------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| `CreateVolumeGroup` | Resolve (or ensure) the consistency group over the given volume handles; return its id as the group handle | Idempotent over the group the `storage.simplyblock.io/consistency-group` label already formed at provisioning: the members are co-placed, so the call returns the existing group's id. | +| `ModifyVolumeGroupMembership` | Add or remove members | Refused when placement forbids it (a volume not co-placed on the group's node/LVS); dynamic membership is `design-consistency-groups.md`'s own Phase 4. | +| `DeleteVolumeGroup` | Dissolve the group | The member volumes survive; only the grouping is removed. Idempotent. | + +The label remains the single source of truth for membership (`design-consistency-groups.md`): `CreateVolumeGroup` does not introduce a competing grouping, it addresses the same backend group the label formed. One backend consistency group is therefore reachable three ways: by label at provisioning, by label for `VolumeGroupSnapshot`, and by handle for group replication. -The per-member backend policy is not named on the group class. It comes from the group's `volumeReplicationClassName`, which the fan-out stamps on each member `VolumeReplication` and which resolves to a `ReplicationPolicy` exactly as a standalone `VolumeReplication` does (§7.1). +### 14.4 Group-wide replication (driver and backend) -### 14.6 Backend API +Once grouped, csi-addons addresses the whole group by its handle through the **same** Replication verbs of §5. Each verb operates on the group as one unit and maps to a new backend group-replication endpoint (§9, and P0-6/P0-7): -Group replication adds no replication endpoint. The reconciler reads the consistency group to verify membership, and every promote, demote, and resync reaches the backend through the member `VolumeReplication` objects, which use the per-volume endpoints of §9. The two reads it makes: +| Verb (on the group handle) | Backend group operation | +|---------------------------------|---------------------------------------------------------------------------------------------------------------------------| +| `EnableVolumeReplication` | Attach the consistency group to a group replication policy (cadence and target for the whole group). | +| `PromoteVolume` (force/planned) | Group failover: clone the last fully replicated group-snapshot generation for **every** member on the target, atomically. | +| `DemoteVolume` | Group demote: quiesce every member, take and ship one final group snapshot, confirm it landed, then fence every member. | +| `ResyncVolume` | Group failback: reverse the shipping direction for the whole group. | +| `GetVolumeReplicationInfo` | Group `lastSyncTime`: the creation time of the newest fully replicated group-snapshot generation. | +| `DisableVolumeReplication` | Detach the group from its group replication policy. | -| Method | Endpoint | Notes | -|--------|--------------------------------------|--------------------------------------------------------------------| -| `GET` | `.../consistency-group` (by name) | Resolve the group named by the members' shared label. | -| `GET` | `.../consistency-group/{id}/members` | The group's current membership, compared against the selected set. | +The recovery generation is a **group snapshot** (`bdev_lvol_snapshot_group`, `design-consistency-groups.md` P0-1): one frozen snapshot of every member at a single point, shipped atomically, so a group failover always lands every member at one crash-consistent point. Group promote, demote, and failback are all-or-nothing across members (Open Question 6). The driver detects a group handle (versus a per-volume handle) and routes to these endpoints; a per-volume handle still takes the §5 path unchanged. -Both are consistency-group reads owned by `design-consistency-groups.md`; this design consumes them and adds none. +### 14.5 Backend API -### 14.7 Observability +The group-replication endpoints, cluster-scoped like the per-volume ones under `/api/v2/clusters/{c}`, are new backend work (P0-6, P0-7): -`VolumeGroupReplicationReconciler` carries an `events.EventRecorder` (wired in `operator/cmd/main.go` as `volumegroupreplication-controller`) and emits on the `VolumeGroupReplication` object: +| Method | Endpoint | Notes | +|--------|---------------------------------------------------------|-------------------------------------------------------------------------------------------------------| +| `PUT` | `.../consistency-groups/{id}` (`replication_policy_id`) | Attach and detach the group to a group replication policy. Idempotent (P0-2's group twin). | +| `POST` | `.../consistency-groups/{id}/replication/failover` | Group promote, planned and forced. Clones the last group generation for every member atomically. | +| `POST` | `.../consistency-groups/{id}/replication/demote` | Group demote: quiesce all, ship one final group snapshot, confirm, fence all. | +| `POST` | `.../consistency-groups/{id}/replication/failback` | Group resync (direction reversal). | +| `GET` | `.../consistency-groups/{id}/replication/status` | Typed group status: role, group `lastSyncTime`, lag, per-member health rollup. | +| `GET` | `.../consistency-groups/` (by name), `/{id}/members` | Existing reads (`design-consistency-groups.md`); the VolumeGroup service resolves the handle by them. | -| Event | Type | Reason | -|---------------------------------------------------------------------|---------|----------------------------| -| The selector resolves to the consistency group's current membership | Normal | `GroupReplicationVerified` | -| The selector does not resolve to exactly one whole group | Warning | `GroupMembershipMismatch` | -| At least one member reports `Degraded`, on the transition into it | Warning | `GroupReplicationDegraded` | +### 14.6 Observability -The per-member replication metrics of §11 already cover each group member, so the group adds no metric of its own: its aggregate state (`Completed`/`Degraded`/`Resyncing` and the oldest-member `lastSyncTime`) is derivable from the members' series. +The stock kubernetes-csi-addons controller-manager owns Kubernetes events on `VolumeGroupReplication` (the group promote, demote, and resync outcomes), the same way it owns them for `VolumeReplication`; this design adds none there. The per-member replication metrics of §11 cover each group member; a group rollup (`simplyblock_replication_group_lag_seconds`, keyed by `consistency_group`) is derivable from the group status read of §14.5 and is exported by the same v2 exporter, its worst-member lag being the figure a group RPO alert watches. -### 14.8 Scope +### 14.7 Scope -The base case is one VRG, one storage vendor's PVCs, one `VolumeGroupReplication`. RamenDR's newer multi-VRG "Global VGR" consensus, spanning a replication group across several applications' VRGs, is out of scope (Open Question 4). `VolumeGroupReplicationContent` is unused: `Content` objects bind a group the generic controller-manager provisioned through a real CSI `GroupReplication` call, and the `external: true` fan-out never makes one (Open Question 5). +The base case is one VRG, one storage vendor's PVCs, one `VolumeGroupReplication`. RamenDR's newer multi-VRG "Global VGR" consensus, spanning a replication group across several applications' VRGs, is out of scope (Open Question 4). `VolumeGroupReplicationContent` is created and owned by the stock controller-manager on the `external: false` path (it carries the group handle `CreateVolumeGroup` returns); this design writes nothing to it directly. --- @@ -492,5 +496,8 @@ The base case is one VRG, one storage vendor's PVCs, one `VolumeGroupReplication | 1 | **Demote semantics for the application.** The P0-3 demote fences the volume (ANA inaccessible) after the final flush, and with convergence folded into the verb it is now the only place a planned swap can stall. This is also the one verb the planned promote's lossless guarantee entirely depends on (§5.2): a planned promote is refused unless a completed demote already fenced the source and confirmed the final delta landed, so an unresolved failure mode here is an unresolved gap in the whole "zero loss" claim. Ramen relocation unmounts the workload first, so the fence is ordinarily unopposed, but that is Ramen's choreography, not a guarantee the driver can rely on: a stuck termination, a stale mount that never released, or a demote invoked outside Ramen's normal flow can all leave writes still arriving when quiesce fires. Confirm the verb's behavior when writes are still in flight at quiesce (block versus fail), whether the converge phase has its own budget separate from the quiesced flush, and whether a timeout in either phase must abort back to serving primary or leave the volume fenced with no automatic recovery. | Backend team | | 2 | **Per-volume policy granularity.** A `VolumeReplicationClass` names one policy, and today one policy implies one target and cadence for all its volumes. Confirm one class per (policy, cadence) is an acceptable authoring model for Ramen's `replicationClassSelector`, or whether per-volume interval overrides are needed. | Operator / Backend team | | 3 | ~~**Avoiding the clone on day-one protection.**~~ **Resolved:** `POST .../replication/failover?planned=true`'s no-demote branch now checks `lvol_controller.replication_source_online` (the source's own storage-node status) before falling through to `FAILED_PRECONDITION` -- an online source is a no-op (§5.2), so a healthy volume's first-ever `PromoteVolume` no longer materializes a clone. The remaining residual: a source that dies within the last health-check interval still briefly reads online, so one reconcile can treat a genuine disaster as a no-op before the node's status catches up and the controller retries -- bounded by the health-check detection window, not open-ended. | Backend team | -| 4 | **Is Global VGR needed (§14.8).** The `VolumeGroupReplication` reconciler covers the base case: one VRG, one vendor's PVCs, one group. Confirm whether any planned simplyblock deployment spans a replication group across more than one application's VRG before treating RamenDR's multi-VRG "Global VGR" consensus as work this design should also specify. | Operator team | -| 5 | **`VolumeGroupReplicationContent`'s exact contract (§14.8).** The shipped reconciler writes nothing to it: `Content` objects bind a group the generic (non-external) controller-manager provisioned through a real CSI `GroupReplication` call, and the `external: true` path never makes one, fanning out to per-member `VolumeReplication` objects instead (§14.2). If a future Ramen version or tooling expects a `Content` object to exist even under `external: true`, revisit this. | Operator team | +| 4 | **Is Global VGR needed (§14.7).** §14 covers the base case: one VRG, one vendor's PVCs, one group. Confirm whether any planned simplyblock deployment spans a replication group across more than one application's VRG before treating RamenDR's multi-VRG "Global VGR" consensus as work this design should also specify. | Operator team | +| 5 | **`VolumeGroupReplicationContent` on the `external: false` path (§14.7).** The stock controller-manager creates and owns the `Content`, populating it with the group handle `CreateVolumeGroup` returns. Confirm the driver need only return a stable handle and never reads or writes the `Content` itself, across the csi-addons versions in scope. | Operator team | +| 6 | **Atomicity of group failover, demote, and failback (§14.4).** The design states all-or-nothing across members. Confirm the backend can guarantee it, and define the behavior when one member cannot complete: abort the whole group, or serve a partial group and report it. | Backend team | +| 7 | **The group replication policy (§14.4, P0-7).** `design-consistency-groups.md` removed the replication policy from the consistency group; group replication needs a cadence and target back. Confirm a group replication policy attached to the CG (the group twin of `ReplicationPolicy`) is the model, versus per-member policies coordinated at the group level. | Backend team | +| 8 | **Grouping already-provisioned, non-co-placed volumes (§14.3).** `CreateVolumeGroup` is idempotent over the label-formed, co-placed group. Confirm whether `ModifyVolumeGroupMembership` must support adding a volume that is not already co-placed (a data move), or whether that stays refused as `design-consistency-groups.md`'s Phase 4 dynamic-membership work. | Backend team | diff --git a/operator/docs/designs/design-ramen-integration.md b/operator/docs/designs/design-ramen-integration.md index 06fbdb737..851815218 100644 --- a/operator/docs/designs/design-ramen-integration.md +++ b/operator/docs/designs/design-ramen-integration.md @@ -1,6 +1,6 @@ # Design Document: Ramen Integration -**Status:** Draft (§3 confirms peerClasses needs no operator code, §4 VolumeGroupReplication implemented, E2E validation pending) +**Status:** Draft (§3 confirms peerClasses needs no operator code; §4 defers `VolumeGroupReplication` to `design-csi-addons-replication.md` §14, driver-and-backend, Planned; E2E validation pending) **Author:** Israel Geoffrey (geoffrey1330) **Date:** 2026-09-21 **Test Plan:** [`tests/test-plan-ramen-integration.md`](../tests/test-plan-ramen-integration.md) @@ -16,7 +16,7 @@ | P0-3 | SiteMap, or a hand-authored `DRPlacementControl` standing in for it, driving the `DRPolicy` | Ecosystem | The E2E validation (§6) | Not shipped. SiteMap is an external document and system, and storage is explicitly outside its own scope. | | P0-4 | `csi-addons/spec` at a version whose `GetVolumeReplicationInfoResponse` carries `lastSyncBytes`/`lastSyncDuration` | Ecosystem | Full Appendix A.3 `GetVolumeReplicationInfo` | Not shipped: pinned at v0.2.0 today, which has neither field. | -Without P0-1 through P0-3 nothing in §6 can run, because the validation is E2E-only, live-cluster work that no mock or `envtest` substitutes for. §4's `VolumeGroupReplication` reconciler is unaffected: it needs no OCM hub, no Ramen installation, and no `DRPolicy`, only this cluster's own objects. P0-4's absence is narrower: it leaves `GetVolumeReplicationInfo` reporting only `lastSyncTime`, never cycle size or duration, but Ramen's own `PeerReady` gate (§5.1) does not read either field, so P0-4 does not block the validation itself. +Without P0-1 through P0-3 nothing in §6 can run, because the validation is E2E-only, live-cluster work that no mock or `envtest` substitutes for. §4's `VolumeGroupReplication` work is a driver-and-backend feature specified in `design-csi-addons-replication.md` §14 and does not depend on the OCM hub for its own unit and backend tests, only for the E2E group scenario (M-05). P0-4's absence is narrower: it leaves `GetVolumeReplicationInfo` reporting only `lastSyncTime`, never cycle size or duration, but Ramen's own `PeerReady` gate (§5.1) does not read either field, so P0-4 does not block the validation itself. --- @@ -25,7 +25,7 @@ Without P0-1 through P0-3 nothing in §6 can run, because the validation is E2E- 1. [Background](#1-background) 2. [Goals and Non-Goals](#2-goals-and-non-goals) 3. [peerClasses: Ramen's Own Mechanism](#3-peerclasses-ramens-own-mechanism) -4. [VolumeGroupReplication (Implemented)](#4-volumegroupreplication-implemented) +4. [VolumeGroupReplication](#4-volumegroupreplication) 5. [The Contract, Confirmed](#5-the-contract-confirmed) 6. [E2E Validation Plan](#6-e2e-validation-plan) 7. [Testing Strategy](#7-testing-strategy) @@ -35,7 +35,7 @@ Without P0-1 through P0-3 nothing in §6 can run, because the validation is E2E- ## Overview -`design-csi-addons-replication.md` builds the storage-level adapter Ramen's per-volume DR contract requires, and validates it end to end "without Ramen" (its own §12): real backend, real csi-addons machinery, but a hand-driven `VolumeReplication` object rather than a real Ramen reconcile loop. That document's own §7.2 named `peerClasses` verification Ramen's own hub-side mechanism, out of its scope, and deferred `VolumeGroupReplication` as "the next design," strictly per volume itself. `design-consistency-groups.md` deferred `VolumeGroupReplication` too, as future work independent of any replication policy. This document specifies the one piece that actually needed a new design: §4, `VolumeGroupReplication` on top of the consistency-group primitive. §3 confirms, rather than reopens, `design-csi-addons-replication.md` §7.2's original position on peerClasses, after this document's own history of first building a same-cluster preflight for it and then removing that preflight once it became clear a same-cluster check cannot verify a cross-cluster pairing. §5 through §6 confirm the rest of the per-volume contract against what already shipped and specify the E2E validation that closes `design-csi-addons-replication.md` §12's outstanding acceptance gate (E-06, E-07). +`design-csi-addons-replication.md` builds the storage-level adapter Ramen's per-volume DR contract requires, and validates it end to end "without Ramen" (its own §12): real backend, real csi-addons machinery, but a hand-driven `VolumeReplication` object rather than a real Ramen reconcile loop. That document's own §7.2 named `peerClasses` verification Ramen's own hub-side mechanism, out of its scope, and deferred `VolumeGroupReplication` as "the next design," strictly per volume itself. `design-consistency-groups.md` deferred `VolumeGroupReplication` too, as future work independent of any replication policy. This document records Ramen's part in the two pieces that needed new design: peerClasses (§3) and `VolumeGroupReplication` (§4), the latter now specified in full in `design-csi-addons-replication.md` §14. §3 confirms, rather than reopens, `design-csi-addons-replication.md` §7.2's original position on peerClasses, after this document's own history of first building a same-cluster preflight for it and then removing that preflight once it became clear a same-cluster check cannot verify a cross-cluster pairing. §5 through §6 confirm the rest of the per-volume contract against what already shipped and specify the E2E validation that closes `design-csi-addons-replication.md` §12's outstanding acceptance gate (E-06, E-07). --- @@ -47,7 +47,7 @@ The gap analysis's own headline finding (§2) was that simplyblock's DR machiner The other half of Appendix A's own premise is that Ramen's hub, not this operator, computes `peerClasses` and drives the VRG, through OCM's hub-spoke visibility into every managed cluster. That hub, and the OCM/SiteMap layer above it, is external to this repository (confirmed this session: no `ManagedCluster`, `DRPolicy`, or `VolumeReplicationGroup` reference exists anywhere in this codebase outside design-doc prose, and no mechanism for this operator to reach a peer cluster's Kubernetes API exists or is needed. `design-management-hub.md`, this repo's own hub design, specifies a real hub component of its own (the `fleet-manager`, reading member clusters through OCM's `ManagedClusterView`), but it is a separate fleet-config-distribution concern, and nothing in it builds or is intended to build Ramen's own pairing). Ramen's hub already reports when it cannot find a valid `StorageClass`/`VolumeReplicationClass` pairing across two clusters, once it looks, and only the hub's own cross-cluster visibility can make that comparison at all: a same-cluster read sees whether this cluster's own label is present, never whether its value agrees with the peer's, which is the only question a pairing check actually needs answered. §3 explains why this operator's own history of trying to close that gap locally settled on not closing it. -The gap analysis's own §6 named a second piece of genuinely new work, in its Phase 2: `VolumeGroupReplication`, "on top of the CG primitive," so that the VRG async group path can protect and fail over a multi-volume app at one point rather than only snapshot it. `design-consistency-groups.md` built the CG primitive that gap analysis cites (the `storage.simplyblock.io/consistency-group` label, `VolumeGroupSnapshot`, the `GroupController`) but explicitly left group replication for later, independent of any policy, exactly so this document could attach it without reshaping the group. §4 is that attachment, and the only new production code this document proposes: everything else Appendix A and Appendix B specify is confirmed, in §5, against code `design-csi-addons-replication.md` already shipped. +The gap analysis's own §6 named a second piece of genuinely new work, in its Phase 2: `VolumeGroupReplication`, "on top of the CG primitive," so that the VRG async group path can protect and fail over a multi-volume app at one point rather than only snapshot it. `design-consistency-groups.md` built the CG primitive that gap analysis cites (the `storage.simplyblock.io/consistency-group` label, `VolumeGroupSnapshot`, the `GroupController`) but explicitly left group replication for later, independent of any policy, exactly so this document could attach it without reshaping the group. §4 records Ramen's part in that attachment; the group surface itself is specified as driver-and-backend work in `design-csi-addons-replication.md` §14 (Planned). Everything else Appendix A and Appendix B specify is confirmed, in §5, against code `design-csi-addons-replication.md` already shipped. --- @@ -56,17 +56,17 @@ The gap analysis's own §6 named a second piece of genuinely new work, in its Ph ### Goals - Confirm that Ramen's `peerClasses` pairing needs no operator-side code, settling the question this document's own earlier attempt at a same-cluster preflight left open (§3). -- Specify and implement `VolumeGroupReplication` on top of the consistency-group primitive: fan a group's `primary`/`secondary`/`resync` intent out to its members' existing per-volume adapter, and fan their status back in, satisfying Appendix A.4's group-readiness status query and the gap analysis's own Phase 2 ask (§4). +- Confirm Ramen drives `VolumeGroupReplication` through the stock csi-addons machinery (the `external: false` path) against the driver-and-backend group surface `design-csi-addons-replication.md` §14 specifies, satisfying Appendix A.4's group-readiness query and the gap analysis's own Phase 2 ask (§4). - Confirm, against the actual shipped code, that every condition, verb, and status query Appendix A specifies is satisfied, or state precisely which is not and why (§5). - Confirm, against the actual shipped code, which Appendix B metrics are delivered, which are derivable from what already exists, and which remain future work (§5.4). - Specify a live-cluster validation that closes `design-csi-addons-replication.md` §12's outstanding acceptance gate: a real Ramen VRG, on a real OCM-registered pair of clusters, driving the adapter through protect, planned relocate, and unplanned failover (§6). ### Non-Goals -- **Building any hub, OCM, or cross-cluster Kubernetes access mechanism.** That is Ramen's and OCM's job, external to this operator, confirmed in §1. §4's group reconciler does not reach past this cluster's own objects, and nothing in §6's validation plan asks this operator to reach a peer cluster's API server: every step drives objects on the cluster where the workload currently runs, exactly as `design-csi-addons-replication.md`'s own architecture already assumes. +- **Building any hub, OCM, or cross-cluster Kubernetes access mechanism.** That is Ramen's and OCM's job, external to this operator, confirmed in §1. Nothing in §6's validation plan asks this operator to reach a peer cluster's API server: every step drives objects on the cluster where the workload currently runs, exactly as `design-csi-addons-replication.md`'s own architecture already assumes. - **A same-cluster `peerClasses` preflight, in any form.** Tried once, in this document's own history, and removed (§3): a same-cluster read can confirm a label is present, never that its value agrees with the peer's, and a pairing check that cannot verify agreement is not a pairing check. That comparison needs the hub's own visibility into both clusters and is Ramen's job alone, once a `DRPolicy` exists. -- **Global VGR.** RamenDR's newer multi-VRG consensus feature for a replication group spanning several applications is out of scope. §4 covers the base case: one VRG, one storage vendor's PVCs, one `VolumeGroupReplication` (§4.1, §8 Open Question 4). -- **A new gRPC verb for group replication.** §4.2 is explicit: group promote, demote, and resync fan the same three verbs `design-csi-addons-replication.md` §5 already ships out to every member. Nothing changes on the driver. +- **Global VGR.** RamenDR's newer multi-VRG consensus feature for a replication group spanning several applications is out of scope. §4 covers the base case: one VRG, one storage vendor's PVCs, one `VolumeGroupReplication` (§8 Open Question 4). +- **A new group gRPC verb on the driver's Replication service.** The same `PromoteVolume`/`DemoteVolume`/`ResyncVolume` verbs act on a group handle rather than a single volume (`design-csi-addons-replication.md` §14.4). The new driver surface is the csi-addons VolumeGroup service (`CreateVolumeGroup` and siblings), not a fourth Replication verb. - **SiteMap.** A separate external document and system. Where §6's topology needs a `DRPlacementControl` and SiteMap is not available to author one, a hand-authored stand-in is explicitly permitted (P0-3). - **`bytesBehind`'s remaining Appendix B siblings** (throughput, RTO estimation, backup RPO/RTO). §5.4 accounts for each, and none blocks the validation this document specifies. - **Widening `csi-addons/spec` past v0.2.0.** P0-4 is recorded as a prerequisite, not solved here: it is an upstream dependency version, not something this repository's own code can add a field to. @@ -79,52 +79,21 @@ Ramen's hub pairs a `StorageClass` and a `VolumeReplicationClass` across two man This document tried the same-cluster version anyway, once. A `ReplicationPairReconciler` preflight was specified and implemented, confirming this cluster's own `VolumeReplicationClass` carried the `ramendr.openshift.io/replicationid` label and a `schedulingInterval`, and emitting `PeerClassesVerified`/`PeerClassesMismatch` accordingly. That check has a real, narrow use (it catches an operator who forgot the label entirely), but it is not peerClasses verification: a cluster whose label carries a value that does not match its peer's passes identically to one whose value is correct, because presence, not agreement, is everything a same-cluster read can check. Reporting `PeerClassesVerified` under that name risked being read as a stronger guarantee than it delivered, so it has been removed. Nothing in this repository performs this check today, which is the position `design-csi-addons-replication.md` §7.2 already took before this document first tried to revisit it. -The exact contract Ramen's pairing requires of the two clusters' objects (the identity labels, the `StorageClass` name-matching, and the authoring convention this repository follows beyond what Ramen strictly requires) is recorded in `design-csi-addons-replication.md` §7.2, unchanged by this document. §4's `VolumeGroupReplication` work is unaffected by any of this: it needs no cross-cluster comparison of its own, only this cluster's own consistency-group membership (§4.2). +The exact contract Ramen's pairing requires of the two clusters' objects (the identity labels, the `StorageClass` name-matching, and the authoring convention this repository follows beyond what Ramen strictly requires) is recorded in `design-csi-addons-replication.md` §7.2, unchanged by this document. `VolumeGroupReplication` (§4) is unaffected by any of this: it needs no cross-cluster comparison of its own, only this cluster's own consistency-group membership. --- -## 4. VolumeGroupReplication (Implemented) +## 4. VolumeGroupReplication -The gap analysis's own Phase 2 named this the piece consistency groups still owed: "Add `VolumeGroupReplication` (csi-addons) on top of the CG primitive so the VRG async group path can protect and fail over multi-volume apps at one point, not just snapshot them." `design-consistency-groups.md` §2 deferred it as "future work," independent of any replication policy so that work could attach later. This section is that attachment. +The gap analysis's Phase 2 named this the piece consistency groups still owed: "Add `VolumeGroupReplication` (csi-addons) on top of the CG primitive so the VRG async group path can protect and fail over multi-volume apps at one point, not just snapshot them." It is now specified in full in `design-csi-addons-replication.md` §14, as a **driver-and-backend** feature driven by the stock kubernetes-csi-addons machinery (Phase 4, Planned there). This section records only what Ramen contributes to that path and where this document validates it. -**Authoritative specification.** `design-csi-addons-replication.md` §14 now carries the csi-addons group surface as the per-group half of that document's replication chapter, beside the per-volume `VolumeReplication` surface, and §14.1 … §14.8 are the reference for the reconciler, the admission webhook, the class, the backend reads, and the events. This section keeps the Ramen-integration framing (the offloaded `external` case in §4.1, and why no new gRPC verb is needed in §4.2) and the E2E validation (§6); where the two documents describe the same reconciler, §14 is authoritative. +**Ramen's part.** Ramen's VRG (`RamenDR/ramen`'s `vrg_volgrouprep.go`, confirmed present at `v0.1.0-rc1`) creates one `VolumeGroupReplication` when the member PVCs' `StorageClass` carries `ramendr.openshift.io/groupreplicationid` (the group-level sibling of the per-volume `replicationid` of `design-csi-addons-replication.md` §7.1), and sets `spec.external` from whether that same `StorageClass` carries `ramendr.openshift.io/offloaded`. simplyblock takes the `external: false` path (`design-csi-addons-replication.md` §14.1): the `StorageClass` carries `groupreplicationid` but **not** `offloaded`, so the stock controller-manager owns the object, groups the members through the driver's csi-addons VolumeGroup service, and replicates the whole group as one unit through the driver's Replication service, addressed by a group handle. No operator code is on the path. -### 4.1 What already exists, upstream +**No new group verb.** Group promote, demote, and resync are the same three verbs `design-csi-addons-replication.md` §5 already ships, applied to the group handle rather than a single volume (§14.4). The recovery generation is a group snapshot, so a group failover lands every member at one crash-consistent point. -`VolumeGroupReplication`, `VolumeGroupReplicationClass`, and `VolumeGroupReplicationContent` are shipped CRDs in the same `replication.storage.openshift.io` group as `VolumeReplication` and `VolumeReplicationClass`, from `kubernetes-csi-addons`. `VolumeGroupReplication.spec` carries the same three-state `replicationState` (`primary`/`secondary`/`resync`) the per-volume kind carries, a `source.selector` naming the member PVCs by label, and an `external` boolean: `false` routes reconciliation through the generic kubernetes-csi-addons controller-manager, and `true` hands it to "an external controller managed by the storage vendor." Ramen's own VRG (`RamenDR/ramen`'s `vrg_volgrouprep.go`) creates and drives `VolumeGroupReplication` objects directly, once a VRG's PVCs share a `VolumeGroupReplicationClass` carrying a `ramendr.openshift.io/groupreplicationid` label, the group-level sibling of the per-volume `replicationid` label `design-csi-addons-replication.md` §7.1 defines. Ramen's own documentation names this case "offloaded" replication: a storage backend that replicates at the LUN or logical-volume-store level, outside Kubernetes, through a vendor controller, exactly the shape simplyblock's backend already has. +**Global VGR out of scope.** RamenDR's newer multi-VRG consensus, spanning a replication group across several applications' VRGs, is not covered here; §8 Open Question 4. -**Global VGR, RamenDR's newer multi-VRG consensus feature for a replication group spanning several applications' VRGs, is out of scope here.** This section covers the base case Ramen has supported longer: one VRG, one storage vendor's PVCs, one `VolumeGroupReplication`. Whether the base case is sufficient for simplyblock's own use, or Global VGR's cross-VRG consensus is eventually needed too, is §8 Open Question 4. - -### 4.2 No new gRPC contract - -Promoting, demoting, or resyncing a group is fanning the same three verbs `design-csi-addons-replication.md` §5 already ships (`PromoteVolume`, `DemoteVolume`, `ResyncVolume`) out to every member, not a fourth verb on the driver. The `replicationState` values line up one for one with the per-volume kind's, by design: `kubernetes-csi-addons`'s own generic controller reconciles a *non*-external `VolumeGroupReplication` this same way, fanning it out to member `VolumeReplication` objects it creates itself. Setting `external: true` does not change what the fan-out does. It changes who performs it, so that the vendor controller can skip the generic manager's own per-member `VolumeReplication` bookkeeping and go straight to whatever shape fits the backend. simplyblock's shape is already built: reuse the per-volume adapter through the same `VolumeReplication` objects Ramen already drives for a single volume, one per group member. - -### 4.3 The reconciler - -`VolumeGroupReplicationReconciler` (`operator/internal/controller/volumegroupreplication_controller.go`), alongside `ReplicationPairReconciler` and the rest, owns every `VolumeGroupReplication` whose `spec.external` is `true` and whose `spec.volumeGroupReplicationClassName` names a class with `provisioner: csi.simplyblock.io`: - -1. **Resolve membership.** Read `spec.source.selector` against this cluster's PVCs. Reuse the exact invariant `design-consistency-groups.md` §9.2 already established for `VolumeGroupSnapshot`: the selected set must equal a `storage.simplyblock.io/consistency-group` value's current membership exactly, not merely a subset or superset of it. A selector that does not resolve to one whole group is a configuration error, not a partial group to serve. -2. **Fan out.** For each member PVC, ensure a per-volume `VolumeReplication` object exists, owned by the `VolumeGroupReplication`, named deterministically from the group and the member, with `spec.replicationState` mirroring the group's. The already-shipped `kubernetes-csi-addons` controller-manager reconciles each of these exactly as it does any Ramen-created per-volume `VolumeReplication` (`design-csi-addons-replication.md` §5): this reconciler creates and updates the member objects, and never calls the driver's Replication gRPC itself. -3. **Fan in.** Aggregate every member's `VolumeReplication.status.conditions` into the group's own status: `Completed` is the conjunction across all members, `Degraded` and `Resyncing` are the disjunction (one degraded or resyncing member makes the group so). `status.lastSyncTime` (the real upstream field on `VolumeGroupReplication.status`, mirroring the per-volume kind's own field name exactly rather than a distinct "group" field, confirmed against the shipped CRD schema) is the oldest of the members' `lastSyncTime`: a group's recovery point is only as fresh as its slowest member. - -### 4.4 Admission webhook, extended - -`design-consistency-groups.md` §9.4's validating webhook on `VolumeGroupSnapshot` create already enforces "the selector must equal the group's current membership" at `kubectl apply`, with the same fail-closed label check and fail-open backend check. `VolumeGroupReplicationValidator` (`operator/internal/webhook/volumegroupreplication_validator.go`) is the sibling webhook this section specified: the identical two checks against the identical label, plus an ownership gate matching `VolumeGroupSnapshotValidator`'s own (`spec.external: true` and a class attributed to `csi.simplyblock.io`, admitting everything else untouched since `replication.storage.openshift.io` is a shared upstream group), so a `VolumeGroupReplication` that could never resolve to one whole group is rejected before the reconciler in §4.3 ever sees it, rather than sitting unreconciled. - -### 4.5 RBAC and events - -`VolumeGroupReplicationReconciler` needs read access to `persistentvolumeclaims`/`persistentvolumes` (declared again here for self-documentation, though already granted elsewhere in this operator), read access to `volumegroupreplicationclasses` to decide ownership (§4.3's own check, not listed in this section's first draft and added here to match the shipped code), and read-write access to the external `volumegroupreplications.replication.storage.openshift.io` and the per-member `volumereplications.replication.storage.openshift.io` it creates: - -```go -// +kubebuilder:rbac:groups=replication.storage.openshift.io,resources=volumegroupreplications,verbs=get;list;watch;update;patch -// +kubebuilder:rbac:groups=replication.storage.openshift.io,resources=volumegroupreplications/status,verbs=get;update;patch -// +kubebuilder:rbac:groups=replication.storage.openshift.io,resources=volumegroupreplicationclasses,verbs=get;list;watch -// +kubebuilder:rbac:groups=replication.storage.openshift.io,resources=volumereplications,verbs=get;list;watch;create;update;patch;delete -// +kubebuilder:rbac:groups="",resources=persistentvolumeclaims,verbs=get;list;watch -// +kubebuilder:rbac:groups="",resources=persistentvolumes,verbs=get;list;watch -``` - -`VolumeGroupReplicationReconciler` carries a `Recorder events.EventRecorder` field, following the exact pattern `backuppolicy_controller.go` already uses (`r.Recorder.Eventf(object, nil, eventType, reason, reason, format, args...)`), wired in `operator/cmd/main.go` as `mgr.GetEventRecorder("volumegroupreplication-controller")`. Events land on the `VolumeGroupReplication` object: `GroupReplicationVerified`/`GroupMembershipMismatch` on every reconcile's own membership check (§4.3 step 1, the same check the webhook makes at admission, mirrored here for the case the webhook admitted open), and a `GroupReplicationDegraded` (Warning) the first time fan-in observes `Degraded=True` after the group's own previous status reported it false, read back from the group object before the status update that reports the new value. +**An earlier operator-owned revision is removed.** A previous version of this section specified a `VolumeGroupReplicationReconciler` and a `VolumeGroupReplicationValidator` in this operator, owning `external: true` objects and fanning them out to per-member `VolumeReplication` objects. `design-csi-addons-replication.md` §13 removes them in favor of the driver-driven path above: it re-implemented in the operator what the stock controller already does, and drove the group through per-member calls rather than one group operation. Nothing in this operator handles `VolumeGroupReplication` now. --- @@ -195,8 +164,8 @@ Every step's pass criterion is an observable Ramen already reports on its own ob Two different classes of coverage, for the two different things this document specifies. §3 adds nothing to test: it confirms that no operator-side code exists for peerClasses, and there is no code left to exercise once the same-cluster preflight was removed. -- **§4's `VolumeGroupReplicationReconciler` is unit-tested against a fake client** (`volumegroupreplication_controller_unit_test.go`), the same harness shape `replicationpair_controller_unit_test.go` already uses elsewhere in this package, pre-seeding member `VolumeReplication` objects with the status a prior fan-out would have created and asserting the fan-in aggregation: a group whose members all report `Completed` yields a `Completed` group, any one member `Degraded` yields a `Degraded` group, and `status.lastSyncTime` is the oldest member's, not the newest. The admission webhook (§4.4) is unit-tested the same way (`volumegroupreplication_validator_test.go`), mirroring how `design-consistency-groups.md` §9.4's own tests exercise its `VolumeGroupSnapshot` sibling: a selector matching a whole group admitted, one matching a subset or spanning two groups rejected, a backend outage admitted (fail-open). -- **§5 confirms existing code** and adds nothing to test on its own. **§6 is exclusively E2E, live-cluster validation** with no smaller harness to substitute, including a group-protect scenario confirming §4's reconciler under a real VRG's `VolumeGroupReplication`. +- **§4's group surface is tested where it lives**, in `design-csi-addons-replication.md` §14: the driver's VolumeGroup service and group-handle Replication routing, and the backend group-replication engine, as unit and backend scenarios in that document's test plan (U-62 … U-69). This document adds no unit test of its own for it, since no operator code implements it any longer. +- **§5 confirms existing code** and adds nothing to test on its own. **§6 is exclusively E2E, live-cluster validation** with no smaller harness to substitute, including a group-protect scenario (M-05) confirming a real VRG's `VolumeGroupReplication` driving the driver's group surface end to end. Full scenario detail: [`tests/test-plan-ramen-integration.md`](../tests/test-plan-ramen-integration.md). Its M-01 and M-02 close `design-csi-addons-replication.md` test plan's E-06 and E-07, which have carried no implementing test since they were written. @@ -204,10 +173,10 @@ Full scenario detail: [`tests/test-plan-ramen-integration.md`](../tests/test-pla ## 8. Open Questions -| # | Question | Owner | -|-----|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-----------------------| -| 1 | **Is a real OCM hub with Ramen already available for this validation**, or does P0-1/P0-2 need to be stood up from scratch? The answer decides whether §6 is schedulable now or needs its own infrastructure work first. | Operator team / Infra | -| 2 | **Does SiteMap exist in a runnable form yet**, or does every run of §6 use a hand-authored `DRPlacementControl` stand-in (P0-3)? If SiteMap is not yet runnable, note that explicitly rather than blocking on it indefinitely. | Operator team | -| 3 | **`csi-addons/spec` version floor (P0-4).** Confirm whether a newer pinned version already carries `lastSyncBytes`/`lastSyncDuration` before treating this as a real upstream gap to track. | Operator team | -| 4 | **Is Global VGR needed.** §4.1 scopes `VolumeGroupReplication` to the base, single-VRG case. Confirm whether any planned simplyblock deployment spans a replication group across more than one application's VRG before treating RamenDR's multi-VRG consensus feature as work this document should also specify. | Operator team | -| 5 | **`VolumeGroupReplicationContent`'s exact contract, provisionally resolved.** The shipped reconciler writes nothing to it: `Content` objects bind a group the generic (non-external) controller-manager provisioned through a real CSI `GroupReplication` call, and this operator's `external: true` path never makes one, fanning out to per-volume `VolumeReplication` objects instead (§4.2). If a future Ramen version or tooling expects a `Content` object to exist even under `external: true`, revisit this. | Operator team | +| # | Question | Owner | +|-----|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-----------------------| +| 1 | **Is a real OCM hub with Ramen already available for this validation**, or does P0-1/P0-2 need to be stood up from scratch? The answer decides whether §6 is schedulable now or needs its own infrastructure work first. | Operator team / Infra | +| 2 | **Does SiteMap exist in a runnable form yet**, or does every run of §6 use a hand-authored `DRPlacementControl` stand-in (P0-3)? If SiteMap is not yet runnable, note that explicitly rather than blocking on it indefinitely. | Operator team | +| 3 | **`csi-addons/spec` version floor (P0-4).** Confirm whether a newer pinned version already carries `lastSyncBytes`/`lastSyncDuration` before treating this as a real upstream gap to track. | Operator team | +| 4 | **Is Global VGR needed.** §4 scopes `VolumeGroupReplication` to the base, single-VRG case. Confirm whether any planned simplyblock deployment spans a replication group across more than one application's VRG before treating RamenDR's multi-VRG consensus feature as work this document should also specify. Tracked as `design-csi-addons-replication.md` Open Question 4. | Operator team | +| 5 | ~~**`VolumeGroupReplicationContent`'s exact contract.**~~ **Moved.** With group replication now the driver-and-backend, `external: false` path (§4), the `Content` object is created and owned by the stock controller-manager; its contract is tracked in `design-csi-addons-replication.md` Open Question 5, not here. | Operator team | diff --git a/operator/docs/tests/test-plan-csi-addons-replication.md b/operator/docs/tests/test-plan-csi-addons-replication.md index f4f0db7b2..9280fb443 100644 --- a/operator/docs/tests/test-plan-csi-addons-replication.md +++ b/operator/docs/tests/test-plan-csi-addons-replication.md @@ -6,7 +6,7 @@ Scope is the CSI driver's Replication service, the operator's preflight and coex Scenario IDs are permanent and are never reused or renumbered. `U-` is unit (no cluster: mock control plane, fake `client.Client`), `I-` is integration (the sidecar and controller-manager against the driver with a mock backend), `E-` is end-to-end (two live simplyblock clusters), and `M-` is manual. Types are `Positive`, `Negative`, `Boundary`, and `Regression`. A `—` in the `Test` column means nothing implements the scenario yet, and every such row reappears in §7 with its reason. -Phase 1 (the csi-addons machinery, §4, §5.1's three verbs, and §6's steady-state contract) and Phase 2 (§5.2's lifecycle verbs and P0-3) have both landed; their unit rows below are filled in. Phase 4 (§14, `VolumeGroupReplication`) has landed too: its reconciler and admission webhook are unit-tested (U-57 … U-61), and its live group relocate is the E2E row E-08. The operator's preflight and coexistence controllers (peerClasses, `PVCReplicationController`) remain a separate, unbuilt subsystem, and the integration and E2E tiers wait on a test bed neither phase has built yet. +Phase 1 (the csi-addons machinery, §4, §5.1's three verbs, and §6's steady-state contract) and Phase 2 (§5.2's lifecycle verbs and P0-3) have both landed; their unit rows below are filled in. Phase 4 (§14, `VolumeGroupReplication`) is **Planned**: the driver VolumeGroup service, the group-handle Replication routing, and the backend group-replication engine (U-62 … U-69) are not built yet, and its live group relocate is E-08. The earlier operator-owned reconciler and validator (U-57 … U-61) are retired (design §13). The operator's preflight and coexistence controllers (peerClasses, `PVCReplicationController`) remain a separate, unbuilt subsystem, and the integration and E2E tiers wait on a test bed neither phase has built yet. --- @@ -112,17 +112,27 @@ File: `operator/internal/controllers/driver/workloads_test.go`, `operator/intern | U-38 | The sidecar's grant is a namespaced Role bound to the controller plugin's account, not a ClusterRole, since CSIAddonsNode is namespaced | Positive | `TestCSIAddonsRoleIsNamespacedAndBoundToTheControllerAccount` | | U-39 | The namespaced Role's rules are scoped to the sidecar's own job: its CSIAddonsNode and its own leader-election Lease | Positive | `TestCSIAddonsRoleRulesAreScopedToItsOwnJob` | -### VolumeGroupReplication: fan-in and membership (design §14) +### VolumeGroupReplication: VolumeGroup service and group-wide replication (design §14, Phase 4 — Planned) -Files: `operator/internal/controller/volumegroupreplication_controller_unit_test.go`, `operator/internal/webhook/volumegroupreplication_validator_test.go`. These are the same scenarios [`test-plan-ramen-integration.md`](test-plan-ramen-integration.md) §1 tracks as its U-01 … U-05; the design they verify now lives in §14 of this plan's design doc, so the coverage is recorded here too. `VolumeGroupReplicationReconciler` fans a group's `replicationState` out to its members and their status back in (§14.3), and its admission webhook (§14.4) enforces the one-whole-group invariant. The reconciler calls no driver gRPC, so the fan-out itself is exercised only at the E2E tier (E-08). +The driver-driven, `external: false` model (design §14): the plugin's csi-addons VolumeGroup service maps a set of volume handles to the backend consistency group (§14.3), and the existing Replication verbs route a group handle to the new backend group-replication endpoints (§14.4). None of this is built yet, so every row's `Test` is `—` and reappears in §7. Files (planned): `csi-driver/internal/csi/controller/volumegroup_test.go`, `csi-driver/internal/csi/controller/replication_group_test.go`, and the backend tier in `sbcli`. -| # | Scenario | Type | Test | -|------|-------------------------------------------------------------------------------------------------------------------|----------|-------------------------------------------------------------------| -| U-57 | Every member's `VolumeReplication` reports `Completed=True, Degraded=False`, and the group reports the same | Positive | `TestVolumeGroupReplication_AllMembersHealthyYieldsGroupHealthy` | -| U-58 | One member reports `Degraded=True`, and the group reports `Degraded=True` (disjunction) | Negative | `TestVolumeGroupReplication_OneMemberDegradedYieldsGroupDegraded` | -| U-59 | Members report differing `lastSyncTime`, and `status.lastSyncTime` is the oldest, not the newest | Boundary | `TestVolumeGroupReplication_LastSyncTimeIsTheOldestMember` | -| U-60 | `spec.source.selector` resolves to exactly one consistency group's current membership, and is admitted | Positive | `TestVolumeGroupReplicationValidator` | -| U-61 | `spec.source.selector` resolves to a subset of a group, or spans two groups, and is rejected, naming the mismatch | Negative | `TestVolumeGroupReplicationValidator` | +U-57 … U-61 are **retired**: they covered the earlier operator-owned `VolumeGroupReplicationReconciler`/`VolumeGroupReplicationValidator`, which design §13 removes in favor of this path. Their Go tests are deleted with the reconciler. + +| # | Scenario | Type | Test | +|----------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------|----------|------| +| ~~U-57~~ | **Retired** (operator fan-in aggregation; reconciler removed, design §13) | — | — | +| ~~U-58~~ | **Retired** (operator fan-in disjunction; reconciler removed) | — | — | +| ~~U-59~~ | **Retired** (operator oldest-`lastSyncTime`; reconciler removed) | — | — | +| ~~U-60~~ | **Retired** (operator webhook admits a whole group; validator removed) | — | — | +| ~~U-61~~ | **Retired** (operator webhook rejects a subset/cross-group selector; validator removed) | — | — | +| U-62 | `CreateVolumeGroup` over the label-formed, co-placed group resolves and returns the existing group's id (idempotent), making no second group | Positive | — | +| U-63 | `CreateVolumeGroup` with a handle that resolves to no consistency group: `FAILED_PRECONDITION` (or the group's own not-found), no group created | Negative | — | +| U-64 | `ModifyVolumeGroupMembership` adding a volume not co-placed on the group's node/LVS: refused (design §14.3, dynamic membership deferred) | Negative | — | +| U-65 | `DeleteVolumeGroup` dissolves the grouping and leaves every member volume intact; repeating it succeeds (idempotent) | Boundary | — | +| U-66 | A group handle on `PromoteVolume`/`DemoteVolume`/`ResyncVolume` routes to the group endpoints; a per-volume handle still takes the §5 path unchanged | Positive | — | +| U-67 | `GetVolumeReplicationInfo` on a group handle returns the group `lastSyncTime` (newest fully replicated group generation), never `NOT_FOUND` for a real group | Positive | — | +| U-68 | Backend group failover clones the last group-snapshot generation for every member atomically; if one member cannot clone, the whole group operation aborts (§14.4) | Boundary | — | +| U-69 | Backend group demote quiesces every member, ships one final group snapshot, confirms it landed, then fences every member | Positive | — | --- @@ -195,29 +205,29 @@ Two live simplyblock clusters with the chart-deployed csi-addons machinery. The ## 5. Axis Coverage -| Axis | Values covered | IDs | Not covered | -|-------------------|------------------------------------------------------------------------|-----------------------------------------------------------------------------|------------------------------------------------------------------------| -| Verb lifecycle | enable, disable, info, forced promote, planned promote, demote, resync | U-01, U-02, U-04 … U-10, U-12 … U-17, U-40 … U-56, I-02 … I-04, E-01 … E-04 | — | -| Idempotency | repeat enable, disable, demote; re-drive after restart | U-02, U-05, U-15, I-07 | repeated promote (U-11), repeated resync | -| Conditions | healthy, degraded, error, staleness, resyncing, disabled | U-18 … U-22, E-05 | condition behavior across backend upgrade | -| Coexistence | slot skip, one-owner refusal, concurrent claim | U-26, U-27, M-02 | migration of an annotated volume onto a VolumeReplication | -| peerClasses | verified, missing class, mispaired policies | U-23 … U-25 | drift after verification | -| Orchestrator | direct kubectl lifecycle, Ramen VRG async (per volume and per group) | I-02 … I-06, E-06, E-07, E-08 | Ramen hub failover of multiple apps | -| Group replication | fan-in aggregation, membership validation, live group relocate | U-57 … U-61, E-08 | Global (multi-VRG) VGR; the fan-out at the integration tier (no I-row) | -| Cluster topology | two clusters, one relationship addressed from both sides | E-01 … E-08 | three-cluster (cascaded) topologies | +| Axis | Values covered | IDs | Not covered | +|-------------------|-----------------------------------------------------------------------------|-----------------------------------------------------------------------------|----------------------------------------------------------------------------------------------| +| Verb lifecycle | enable, disable, info, forced promote, planned promote, demote, resync | U-01, U-02, U-04 … U-10, U-12 … U-17, U-40 … U-56, I-02 … I-04, E-01 … E-04 | — | +| Idempotency | repeat enable, disable, demote; re-drive after restart | U-02, U-05, U-15, I-07 | repeated promote (U-11), repeated resync | +| Conditions | healthy, degraded, error, staleness, resyncing, disabled | U-18 … U-22, E-05 | condition behavior across backend upgrade | +| Coexistence | slot skip, one-owner refusal, concurrent claim | U-26, U-27, M-02 | migration of an annotated volume onto a VolumeReplication | +| peerClasses | verified, missing class, mispaired policies | U-23 … U-25 | drift after verification | +| Orchestrator | direct kubectl lifecycle, Ramen VRG async (per volume and per group) | I-02 … I-06, E-06, E-07, E-08 | Ramen hub failover of multiple apps | +| Group replication | VolumeGroup service, group-handle routing, backend group ops, live relocate | U-62 … U-69, E-08 | Global (multi-VRG) VGR; the controller-manager group loop at the integration tier (no I-row) | +| Cluster topology | two clusters, one relationship addressed from both sides | E-01 … E-08 | three-cluster (cascaded) topologies | --- ## 6. Coverage Summary -| Class | Scenarios | Covered | Not covered | -|-------------|-----------|---------|-------------------------| -| Unit | 51 | 39 | U-03, U-11, U-18 … U-27 | -| Integration | 7 | 0 | I-01 … I-07 | -| E2E | 8 | 0 | E-01 … E-08 | -| Manual | 2 | 0 | M-01, M-02 | +| Class | Scenarios | Covered | Not covered | +|-------------|-----------|---------|--------------------------------------| +| Unit | 54 | 34 | U-03, U-11, U-18 … U-27, U-62 … U-69 | +| Integration | 7 | 0 | I-01 … I-07 | +| E2E | 8 | 0 | E-01 … E-08 | +| Manual | 2 | 0 | M-01, M-02 | -Phase 1 landed the driver's Replication and Identity services, the error classifier, and the operator's sidecar and RBAC wiring, covering every Phase 1 unit scenario except U-03 (§7). Phase 2 landed P0-3 (the demote endpoint and the planned gate on `failover`, in sbcli) and the driver's `PromoteVolume`/`DemoteVolume`/`ResyncVolume`, covering every Phase 2 unit scenario except U-11 (§7). Phase 4 landed the `VolumeGroupReplication` reconciler and its admission webhook, covering all of U-57 … U-61; its live group relocate (E-08) waits on the same Ramen bed as E-06/E-07. The operator's preflight and coexistence controllers (peerClasses, `PVCReplicationController`) remain a separate, unbuilt subsystem, and neither a sidecar-and-controller-manager integration suite nor a live two-cluster E2E bed exists yet, so those tiers remain fully uncovered. +Phase 1 landed the driver's Replication and Identity services, the error classifier, and the operator's sidecar and RBAC wiring, covering every Phase 1 unit scenario except U-03 (§7). Phase 2 landed P0-3 (the demote endpoint and the planned gate on `failover`, in sbcli) and the driver's `PromoteVolume`/`DemoteVolume`/`ResyncVolume`, covering every Phase 2 unit scenario except U-11 (§7). Phase 4 (group replication, §14) is Planned: U-62 … U-69 (the driver VolumeGroup service, group-handle routing, and the backend group engine) and E-08 are uncovered, and the earlier operator reconciler's rows (U-57 … U-61) are retired with it. The operator's preflight and coexistence controllers (peerClasses, `PVCReplicationController`) remain a separate, unbuilt subsystem, and neither a sidecar-and-controller-manager integration suite nor a live two-cluster E2E bed exists yet, so those tiers remain fully uncovered. --- @@ -232,5 +242,6 @@ Phase 1 landed the driver's Replication and Identity services, the error classif | I-01 … I-07 | The sidecar and controller-manager loop | The driver's Replication and Identity services and the sidecar container now exist (Phase 1); no envtest/kind suite exercises them against the real kubernetes-csi-addons controller-manager yet | | E-01 … E-05 | The live lifecycle (non-Ramen half) | Needs a two-cluster live test bed. `regression_test/21/` exercises the same lifecycle by hand but is not wired as an automated E2E suite. | | E-06, E-07, E-08 | The Ramen-driven gate (per-volume and per-group) | Detailed scenario ownership moved to [`test-plan-ramen-integration.md`](test-plan-ramen-integration.md) M-01 … M-03 (per volume) and M-05 (the `VolumeGroupReplication` group relocate), blocked there on a live OCM hub with Ramen installed (see that document's Phase 0) | +| U-62 … U-69 | Group replication (§14): the driver VolumeGroup service, group-handle routing, and the backend group engine | Phase 4 is Planned, not built. P0-6/P0-7 (the `sbcli` group-replication engine and group policy) are the blocking backend work; the driver's VolumeGroup service and group-handle routing depend on them | | — | Repeated resync, class drift after verification, annotated-volume migration onto the adapter, cascaded topologies | Beyond the first coverage pass, recorded so the gaps are explicit rather than assumed covered | | M-01, M-02 | Demote under writes; concurrent ownership race | Need failure injection and precise timing a live two-cluster run does not automate yet | diff --git a/operator/docs/tests/test-plan-ramen-integration.md b/operator/docs/tests/test-plan-ramen-integration.md index 92059df6a..210376608 100644 --- a/operator/docs/tests/test-plan-ramen-integration.md +++ b/operator/docs/tests/test-plan-ramen-integration.md @@ -1,25 +1,27 @@ # Test Plan: Ramen Integration Related design: [`designs/design-ramen-integration.md`](../designs/design-ramen-integration.md) -Harness: two classes. `VolumeGroupReplicationReconciler` (design §4) is unit-tested against a fake-client harness, the same shape `replicationpair_controller_unit_test.go` already uses elsewhere in this package. Everything else needs a live two-cluster simplyblock deployment (`regression_test/21/`), plus an OCM hub with Ramen installed on the hub and both managed clusters (design §6.1). +Harness: one class. Everything this document specifies needs a live two-cluster simplyblock deployment (`regression_test/21/`), plus an OCM hub with Ramen installed on the hub and both managed clusters (design §6.1). `VolumeGroupReplication` (design §4) is now a driver-and-backend feature specified in `design-csi-addons-replication.md` §14; its unit and backend scenarios live in that document's test plan (U-62 … U-69), not here. -Scope: two things this document specifies. First, whether `VolumeGroupReplicationReconciler` (design §4) correctly fans a group's replication state out to its members and their status back in, and correctly validates group membership at admission, as unit scenarios (`U-01` … `U-05`). Second, whether a real Ramen `VolumeReplicationGroup`, driven through a real OCM hub, correctly drives the csi-addons adapter `design-csi-addons-replication.md` implements, for both a single volume and a consistency-group of them, as manual E2E scenarios (`M-`), since no smaller harness substitutes for a real Ramen reconcile loop against a real OCM-registered cluster pair. The manual scenarios close two rows already carried in [`test-plan-csi-addons-replication.md`](test-plan-csi-addons-replication.md): E-06 and E-07, both `—` in that plan's `Test` column since they were written. Design §3 (peerClasses) adds no scenarios of its own: it confirms that no operator-side code exists to test. +Scope: this document specifies whether a real Ramen `VolumeReplicationGroup`, driven through a real OCM hub, correctly drives the csi-addons surface `design-csi-addons-replication.md` implements, for both a single volume and a consistency group of them, as manual E2E scenarios (`M-`), since no smaller harness substitutes for a real Ramen reconcile loop against a real OCM-registered cluster pair. The manual scenarios close two rows already carried in [`test-plan-csi-addons-replication.md`](test-plan-csi-addons-replication.md): E-06 and E-07, both `—` in that plan's `Test` column since they were written. Design §3 (peerClasses) adds no scenarios of its own: it confirms that no operator-side code exists to test. --- ## 1. Unit Scenarios -### VolumeGroupReplicationReconciler and its admission webhook (design §4) +### VolumeGroupReplication (design §4) -Implemented in `volumegroupreplication_controller.go` and `volumegroupreplication_validator.go`, covered by `volumegroupreplication_controller_unit_test.go` and `volumegroupreplication_validator_test.go`. Since the reconciler's design now lives in `design-csi-addons-replication.md` §14 (the csi-addons group surface), these same scenarios are mirrored in [`test-plan-csi-addons-replication.md`](test-plan-csi-addons-replication.md) as its U-57 … U-61, and the group's live relocate as that plan's E-08. +`VolumeGroupReplication` is no longer an operator reconciler; it is the driver-and-backend, `external: false` path of `design-csi-addons-replication.md` §14. Its unit and backend scenarios are that document's test plan U-62 … U-69, and its live group relocate is that plan's E-08 (which points back to M-05 here). This plan owns no unit scenario for it. -| # | Scenario | Type | Test | -|------|-------------------------------------------------------------------------------------------------------------------|----------|-------------------------------------------------------------------| -| U-01 | Every member's `VolumeReplication` reports `Completed=True, Degraded=False`, and the group reports the same | Positive | `TestVolumeGroupReplication_AllMembersHealthyYieldsGroupHealthy` | -| U-02 | One member's `VolumeReplication` reports `Degraded=True`, and the group reports `Degraded=True` | Negative | `TestVolumeGroupReplication_OneMemberDegradedYieldsGroupDegraded` | -| U-03 | Members report differing `lastSyncTime`, and `status.lastSyncTime` is the oldest, not the newest | Boundary | `TestVolumeGroupReplication_LastSyncTimeIsTheOldestMember` | -| U-04 | `spec.source.selector` resolves to exactly one `ConsistencyGroup`'s current membership, and is admitted | Positive | `TestVolumeGroupReplicationValidator` | -| U-05 | `spec.source.selector` resolves to a subset of a group, or spans two groups, and is rejected, naming the mismatch | Negative | `TestVolumeGroupReplicationValidator` | +U-01 … U-05 are **retired**: they covered the operator-owned `VolumeGroupReplicationReconciler`/`VolumeGroupReplicationValidator`, which `design-csi-addons-replication.md` §13 removes. + +| # | Scenario | Type | Test | +|----------|-----------------------------------------------------------------------------------------|------|------| +| ~~U-01~~ | **Retired** (operator fan-in, all members healthy; reconciler removed) | — | — | +| ~~U-02~~ | **Retired** (operator fan-in, one member degraded; reconciler removed) | — | — | +| ~~U-03~~ | **Retired** (operator oldest-`lastSyncTime`; reconciler removed) | — | — | +| ~~U-04~~ | **Retired** (operator webhook admits a whole group; validator removed) | — | — | +| ~~U-05~~ | **Retired** (operator webhook rejects a subset/cross-group selector; validator removed) | — | — | --- @@ -76,44 +78,43 @@ Implemented in `volumegroupreplication_controller.go` and `volumegroupreplicatio ### M-05: Ramen protects and relocates a multi-volume app through `VolumeGroupReplication` -**Design reference:** design §4, exercised through the topology and test flow §6 defines for the per-volume case. New scope this document adds, with no corresponding row in `test-plan-csi-addons-replication.md`. +**Design reference:** design §4 and `design-csi-addons-replication.md` §14 (the driver-and-backend group surface), exercised through the topology and test flow §6 defines for the per-volume case. New scope this document adds; it closes `test-plan-csi-addons-replication.md` E-08. It cannot run until Phase 4 (that plan's U-62 … U-69) is built. -**What to verify:** a VRG whose PVCs share a `storage.simplyblock.io/consistency-group` label and a `VolumeGroupReplicationClass` carrying `ramendr.openshift.io/groupreplicationid` creates one `VolumeGroupReplication`, `VolumeGroupReplicationReconciler` (design §4.3) fans it out to one `VolumeReplication` per member and fans member status back into the group, and a planned relocate of the whole app moves every member together with none diverging. +**What to verify:** a VRG whose PVCs share the consistency-group labels and a `StorageClass` carrying `ramendr.openshift.io/groupreplicationid` (and **not** `offloaded`) creates one `VolumeGroupReplication` with `spec.external: false`; the stock kubernetes-csi-addons controller-manager forms the backend group through the driver's `CreateVolumeGroup`, replicates it as one unit addressed by the group handle, and a planned relocate of the whole app moves every member together at one crash-consistent point, none diverging. **Test concept:** -1. Extend M-01's topology: a workload with three PVCs sharing one `storage.simplyblock.io/consistency-group` value, protected by one VRG under a `VolumeGroupReplicationClass` naming that group's `ReplicationPolicy`. -2. Confirm exactly one `VolumeGroupReplication` exists and exactly three member `VolumeReplication` objects exist, each owned by it. +1. Extend M-01's topology: a workload with three PVCs sharing one `storage.simplyblock.io/consistency-group` value (and Ramen's own `ramendr.openshift.io/consistency-group`), under a group `StorageClass` carrying `groupreplicationid` and a `VolumeGroupReplicationClass` matching it. +2. Confirm exactly one `VolumeGroupReplication` (`external: false`) and its `VolumeGroupReplicationContent` exist, that the backend consistency group holds all three members, and that the group is replicating (the group `lastSyncTime` advances). 3. Trigger Ramen's `Relocate` action, as in M-02. -4. Assert: all three members reach `Secondary` on cluster A and `Primary` on cluster B together, not staggered. The group's own `status.lastSyncTime` reflects the oldest member's throughout (U-03). No member is left behind mid-relocate. +4. Assert: the whole group promotes on cluster B and demotes on cluster A as one unit, each member's PVC restored and serving on cluster B, every member's pre-relocate data intact, and the group's `lastSyncTime` tracking one group generation throughout. No member is left behind mid-relocate. --- ## 3. Axis Coverage -| Axis | Values covered | IDs | Not covered | -|-------------------------------|----------------------------------------------------------------------------|------------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| Group replication aggregation | all members healthy, one member degraded, differing `lastSyncTime` | U-01 … U-03 | more than one degraded member simultaneously (not a gap: the aggregation is a plain disjunction, one member already exercises it) | -| Group membership validation | selector equals membership, selector is a subset or spans two groups | U-04, U-05 | membership changing between admission and reconcile (the fail-open backend-unreachable case `design-consistency-groups.md` §9.4's sibling check already covers for the snapshot path) | -| Orchestrator | Ramen VRG async, hub-driven | M-01 … M-05 | direct `kubectl` lifecycle (already covered in `test-plan-csi-addons-replication.md`) | -| Cluster topology | two managed clusters, one relationship, hub-mediated | M-01 … M-05 | three-cluster (cascaded) topologies. SiteMap-authored `DRPlacementControl` specifically (§8 Open Question 2 may leave this a hand-authored stand-in) | -| Failure mode | planned relocate, unplanned failover, post-recovery resync, group relocate | M-02, M-03, M-04, M-05 | a demote that stalls mid-convergence while Ramen-driven (covered by hand in `test-plan-csi-addons-replication.md` M-01/M-02, not yet by Ramen) | +| Axis | Values covered | IDs | Not covered | +|----------------------------------|----------------------------------------------------------------------------|--------------------------------------------------------|------------------------------------------------------------------------------------------------------------------------------------------------------| +| Group replication (unit/backend) | the driver VolumeGroup service and the backend group engine | (in `test-plan-csi-addons-replication.md` U-62 … U-69) | owned by the csi-addons plan now that group replication is driver-and-backend (design §4), not this plan | +| Orchestrator | Ramen VRG async, hub-driven | M-01 … M-05 | direct `kubectl` lifecycle (already covered in `test-plan-csi-addons-replication.md`) | +| Cluster topology | two managed clusters, one relationship, hub-mediated | M-01 … M-05 | three-cluster (cascaded) topologies. SiteMap-authored `DRPlacementControl` specifically (§8 Open Question 2 may leave this a hand-authored stand-in) | +| Failure mode | planned relocate, unplanned failover, post-recovery resync, group relocate | M-02, M-03, M-04, M-05 | a demote that stalls mid-convergence while Ramen-driven (covered by hand in `test-plan-csi-addons-replication.md` M-01/M-02, not yet by Ramen) | --- ## 4. Coverage Summary -| Class | Scenarios | Covered | Not covered | -|--------------|-----------|-----------------|-------------| -| Unit | 5 | 5 (U-01 … U-05) | — | -| Manual (E2E) | 5 | 0 | M-01 … M-05 | +| Class | Scenarios | Covered | Not covered | +|--------------|-----------|---------|-------------------------------------------------------------------------------------------------------------| +| Unit | 0 | 0 | — (U-01 … U-05 retired; group unit/backend scenarios are `test-plan-csi-addons-replication.md` U-62 … U-69) | +| Manual (E2E) | 5 | 0 | M-01 … M-05 | --- ## 5. What Is Not Yet Covered -| # | Gap | Reason | -|-------------|-------------------------------------------------------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| -| M-01 … M-05 | The entire Ramen-driven validation | Blocked on Phase 0 (design §Phase 0): a live OCM hub with Ramen installed across a registered two-cluster pair, not yet confirmed available (§8 Open Question 1). M-05 additionally needs the `VolumeGroupReplication` CRDs installed on both clusters. | -| — | Three-cluster / cascaded topologies | Out of scope for this document, and not part of the gap analysis's Appendix A either | -| — | SiteMap-authored (rather than hand-authored) `DRPlacementControl` | Depends on SiteMap's own availability (§8 Open Question 2) | -| — | Global VGR (multi-VRG consensus) | Out of scope for design §4.1, tracked as design §8 Open Question 4 | +| # | Gap | Reason | +|-------------|-------------------------------------------------------------------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| M-01 … M-05 | The entire Ramen-driven validation | Blocked on Phase 0 (design §Phase 0): a live OCM hub with Ramen installed across a registered two-cluster pair, not yet confirmed available (§8 Open Question 1). M-05 additionally needs Phase 4 built (`test-plan-csi-addons-replication.md` U-62 … U-69: the driver VolumeGroup service and the backend group engine) and the `VolumeGroupReplication` CRDs installed on both clusters. | +| — | Three-cluster / cascaded topologies | Out of scope for this document, and not part of the gap analysis's Appendix A either | +| — | SiteMap-authored (rather than hand-authored) `DRPlacementControl` | Depends on SiteMap's own availability (§8 Open Question 2) | +| — | Global VGR (multi-VRG consensus) | Out of scope for design §4.1, tracked as design §8 Open Question 4 | diff --git a/operator/internal/controller/volumegroupreplication_controller.go b/operator/internal/controller/volumegroupreplication_controller.go deleted file mode 100644 index 0b87e490e..000000000 --- a/operator/internal/controller/volumegroupreplication_controller.go +++ /dev/null @@ -1,487 +0,0 @@ -/* -Copyright 2025. - -Licensed under the Apache License, Version 2.0 (the "License"); -you may not use this file except in compliance with the License. -You may obtain a copy of the License at - - http://www.apache.org/licenses/LICENSE-2.0 - -Unless required by applicable law or agreed to in writing, software -distributed under the License is distributed on an "AS IS" BASIS, -WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. -See the License for the specific language governing permissions and -limitations under the License. -*/ - -package controller - -import ( - "context" - "fmt" - "time" - - corev1 "k8s.io/api/core/v1" - apierrors "k8s.io/apimachinery/pkg/api/errors" - metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" - "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" - "k8s.io/apimachinery/pkg/runtime" - "k8s.io/apimachinery/pkg/runtime/schema" - "k8s.io/apimachinery/pkg/types" - "k8s.io/client-go/tools/events" - ctrl "sigs.k8s.io/controller-runtime" - "sigs.k8s.io/controller-runtime/pkg/client" - logf "sigs.k8s.io/controller-runtime/pkg/log" - - "github.com/simplyblock/atlas/kube" - atlaslvol "github.com/simplyblock/atlas/lvol" - - "github.com/simplyblock/simplyblock-operator/internal/utils" - "github.com/simplyblock/simplyblock-operator/internal/webapi" -) - -const ( - groupReplSyncInterval = 60 * time.Second - groupReplRequeueError = 30 * time.Second - - // consistencyGroupLabel names a PVC's consistency group (design-consistency-groups.md - // §4.1). Duplicated, not shared: the webhook package and - // internal/controllers/consistencygroup each declare their own copy of this same - // unexported constant, and this package follows that established precedent - // rather than introducing a shared import across packages for one string. - consistencyGroupLabel = "storage.simplyblock.io/consistency-group" - - reasonGroupReplicationVerified = "GroupReplicationVerified" - reasonGroupMembershipMismatch = "GroupMembershipMismatch" - reasonGroupReplicationDegraded = "GroupReplicationDegraded" -) - -var ( - volumeGroupReplicationGVK = schema.GroupVersionKind{ - Group: "replication.storage.openshift.io", Version: "v1alpha1", Kind: "VolumeGroupReplication", - } - volumeGroupReplicationClassGVK = schema.GroupVersionKind{ - Group: "replication.storage.openshift.io", Version: "v1alpha1", Kind: "VolumeGroupReplicationClass", - } - volumeReplicationGVK = schema.GroupVersionKind{ - Group: "replication.storage.openshift.io", Version: "v1alpha1", Kind: "VolumeReplication", - } -) - -// VolumeGroupReplicationReconciler fans a VolumeGroupReplication's group-level -// replication intent out to one per-volume VolumeReplication per consistency-group -// member, and fans the members' status back into the group's own -// (design-ramen-integration.md §4.3). It owns no gRPC call of its own: promoting, -// demoting, or resyncing a member is the already-shipped per-volume adapter's job, -// driven by the kubernetes-csi-addons controller-manager reconciling the member -// VolumeReplication objects this reconciler creates (§4.2). -type VolumeGroupReplicationReconciler struct { - client.Client - Scheme *runtime.Scheme - Recorder events.EventRecorder -} - -// +kubebuilder:rbac:groups=replication.storage.openshift.io,resources=volumegroupreplications,verbs=get;list;watch;update;patch -// +kubebuilder:rbac:groups=replication.storage.openshift.io,resources=volumegroupreplications/status,verbs=get;update;patch -// +kubebuilder:rbac:groups=replication.storage.openshift.io,resources=volumegroupreplicationclasses,verbs=get;list;watch -// +kubebuilder:rbac:groups=replication.storage.openshift.io,resources=volumereplications,verbs=get;list;watch;create;update;patch;delete -// +kubebuilder:rbac:groups="",resources=persistentvolumeclaims,verbs=get;list;watch -// +kubebuilder:rbac:groups="",resources=persistentvolumes,verbs=get;list;watch - -func (r *VolumeGroupReplicationReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) { - log := logf.FromContext(ctx) - - vgr := &unstructured.Unstructured{} - vgr.SetGroupVersionKind(volumeGroupReplicationGVK) - if err := r.Get(ctx, req.NamespacedName, vgr); err != nil { - return ctrl.Result{}, client.IgnoreNotFound(err) - } - - // A non-external VolumeGroupReplication is the generic kubernetes-csi-addons - // controller-manager's to reconcile, not this operator's (§4.2). - external, _, _ := unstructured.NestedBool(vgr.Object, "spec", "external") - if !external { - return ctrl.Result{}, nil - } - className, _, _ := unstructured.NestedString(vgr.Object, "spec", "volumeGroupReplicationClassName") - owned, err := r.ownsClass(ctx, className) - if err != nil { - log.Error(err, "failed to resolve VolumeGroupReplicationClass", "class", className) - return ctrl.Result{RequeueAfter: groupReplRequeueError}, nil - } - if !owned { - return ctrl.Result{}, nil - } - - members, determinable, mismatch, err := r.resolveMembers(ctx, vgr) - if err != nil { - log.Error(err, "failed to resolve group membership") - return ctrl.Result{RequeueAfter: groupReplRequeueError}, nil - } - if !determinable { - // A selected PVC is not yet bound, or its volume is not yet resolvable. - // Transient: requeue quietly, matching the webhook's fail-open disposition - // for the identical case (design §4.4). - return ctrl.Result{RequeueAfter: groupReplSyncInterval}, nil - } - if mismatch != "" { - r.Recorder.Eventf(vgr, nil, corev1.EventTypeWarning, reasonGroupMembershipMismatch, reasonGroupMembershipMismatch, mismatch) - return ctrl.Result{RequeueAfter: groupReplSyncInterval}, nil - } - r.Recorder.Eventf(vgr, nil, corev1.EventTypeNormal, reasonGroupReplicationVerified, reasonGroupReplicationVerified, - "selector resolves to the consistency group's current membership (%d member(s))", len(members)) - - replicationState, _, _ := unstructured.NestedString(vgr.Object, "spec", "replicationState") - volumeReplicationClass, _, _ := unstructured.NestedString(vgr.Object, "spec", "volumeReplicationClassName") - - memberVRs, err := r.fanOut(ctx, vgr, members, replicationState, volumeReplicationClass) - if err != nil { - log.Error(err, "failed to fan out member VolumeReplications") - return ctrl.Result{RequeueAfter: groupReplRequeueError}, nil - } - - if err := r.fanIn(ctx, vgr, members, memberVRs); err != nil { - log.Error(err, "failed to update VolumeGroupReplication status") - return ctrl.Result{RequeueAfter: groupReplRequeueError}, nil - } - - return ctrl.Result{RequeueAfter: groupReplSyncInterval}, nil -} - -// ownsClass reports whether className names a VolumeGroupReplicationClass for -// this driver. A class this operator cannot read is treated as not-owned: -// the generic controller-manager or another vendor's controller may still be -// the right owner, and refusing to reconcile is safer than guessing. -func (r *VolumeGroupReplicationReconciler) ownsClass(ctx context.Context, className string) (bool, error) { - if className == "" { - return false, nil - } - class := &unstructured.Unstructured{} - class.SetGroupVersionKind(volumeGroupReplicationClassGVK) - if err := r.Get(ctx, types.NamespacedName{Name: className}, class); err != nil { - if apierrors.IsNotFound(err) { - return false, nil - } - return false, err - } - provisioner, _, _ := unstructured.NestedString(class.Object, "spec", "provisioner") - return provisioner == utils.CSIProvisioner, nil -} - -// resolveMembers reads spec.source.selector, resolves it to this cluster's own -// PVCs, and verifies the selected set equals a consistency group's current -// membership exactly (design §4.3 step 1, the same invariant -// design-consistency-groups.md §9.2 established for VolumeGroupSnapshot). -// -// determinable is false when a selected PVC is not yet bound, matching the -// webhook's fail-open case for the identical situation. mismatch is non-empty -// when membership is determinable but does not match. -func (r *VolumeGroupReplicationReconciler) resolveMembers( - ctx context.Context, vgr *unstructured.Unstructured, -) (members []corev1.PersistentVolumeClaim, determinable bool, mismatch string, err error) { - selMap, found, err := unstructured.NestedMap(vgr.Object, "spec", "source", "selector") - if err != nil { - return nil, true, "", fmt.Errorf("read spec.source.selector: %w", err) - } - if !found { - return nil, true, "VolumeGroupReplication has no spec.source.selector", nil - } - var labelSelector metav1.LabelSelector - if err := runtime.DefaultUnstructuredConverter.FromUnstructured(selMap, &labelSelector); err != nil { - return nil, true, "", fmt.Errorf("convert spec.source.selector: %w", err) - } - sel, err := metav1.LabelSelectorAsSelector(&labelSelector) - if err != nil { - return nil, true, fmt.Sprintf("invalid label selector: %v", err), nil - } - - var pvcList corev1.PersistentVolumeClaimList - if err := r.List(ctx, &pvcList, - client.InNamespace(vgr.GetNamespace()), client.MatchingLabelsSelector{Selector: sel}); err != nil { - return nil, true, "", fmt.Errorf("list PVCs: %w", err) - } - if len(pvcList.Items) == 0 { - return nil, true, "selector matches no PersistentVolumeClaim", nil - } - - groupName := "" - for i := range pvcList.Items { - val := pvcList.Items[i].Labels[consistencyGroupLabel] - if val == "" { - return nil, true, fmt.Sprintf("PVC %q is not labeled %s", pvcList.Items[i].Name, consistencyGroupLabel), nil - } - if groupName == "" { - groupName = val - } else if val != groupName { - return nil, true, fmt.Sprintf("selector spans two consistency groups (%q and %q)", groupName, val), nil - } - } - - clusterUUID, selectedLvols, ok, err := r.selectedLvols(ctx, pvcList.Items) - if err != nil { - return nil, true, "", err - } - if !ok { - return nil, false, "", nil - } - - apiClient := webapi.NewClient() - group, err := apiClient.GetConsistencyGroupByName(ctx, clusterUUID, groupName) - if err != nil { - return nil, true, "", err - } - if group == nil { - return nil, true, fmt.Sprintf("consistency group %q not found", groupName), nil - } - backendMembers, err := apiClient.GetConsistencyGroupMembers(ctx, clusterUUID, group.UUID) - if err != nil { - return nil, true, "", err - } - if !sameLvolSet(selectedLvols, backendMembers) { - return nil, true, fmt.Sprintf( - "selector resolves to %d volume(s) but consistency group %q has %d member(s); "+ - "the selector must equal the group's current membership", - len(selectedLvols), groupName, len(backendMembers)), nil - } - return pvcList.Items, true, "", nil -} - -// selectedLvols maps each selected PVC to its backing lvol UUID and returns the -// shared cluster UUID. ok is false when a PVC is not yet bound, which makes -// membership undeterminable rather than mismatched. -func (r *VolumeGroupReplicationReconciler) selectedLvols( - ctx context.Context, pvcs []corev1.PersistentVolumeClaim, -) (clusterUUID string, lvols []string, ok bool, err error) { - for i := range pvcs { - pvName := pvcs[i].Spec.VolumeName - if pvName == "" { - return "", nil, false, nil - } - pv := &corev1.PersistentVolume{} - if err := r.Get(ctx, types.NamespacedName{Name: pvName}, pv); err != nil { - return "", nil, false, err - } - raw, err := kube.VolumeHandleFromPV(pv) - if err != nil { - return "", nil, false, nil - } - h, parsed := atlaslvol.ParseHandle(raw) - if !parsed { - return "", nil, false, nil - } - clusterUUID = h.ClusterID - lvols = append(lvols, h.VolumeID) - } - return clusterUUID, lvols, true, nil -} - -// sameLvolSet reports whether a and b hold the same set of lvol ids. -func sameLvolSet(a, b []string) bool { - if len(a) != len(b) { - return false - } - seen := make(map[string]bool, len(a)) - for _, s := range a { - seen[s] = true - } - for _, s := range b { - if !seen[s] { - return false - } - } - return true -} - -// fanOut ensures one per-volume VolumeReplication exists for each member PVC, -// owned by the group, with spec.replicationState mirroring the group's -// (design §4.3 step 2). It never calls the driver's Replication gRPC: the -// already-shipped kubernetes-csi-addons controller-manager reconciles each -// member exactly as it does any Ramen-created per-volume VolumeReplication. -func (r *VolumeGroupReplicationReconciler) fanOut( - ctx context.Context, - vgr *unstructured.Unstructured, - members []corev1.PersistentVolumeClaim, - replicationState, volumeReplicationClass string, -) ([]unstructured.Unstructured, error) { - owner := metav1.OwnerReference{ - APIVersion: volumeGroupReplicationGVK.GroupVersion().String(), - Kind: volumeGroupReplicationGVK.Kind, - Name: vgr.GetName(), - UID: vgr.GetUID(), - } - - result := make([]unstructured.Unstructured, 0, len(members)) - for i := range members { - pvc := &members[i] - name := memberVolumeReplicationName(vgr.GetName(), pvc.Name) - - vr := &unstructured.Unstructured{} - vr.SetGroupVersionKind(volumeReplicationGVK) - err := r.Get(ctx, types.NamespacedName{Namespace: vgr.GetNamespace(), Name: name}, vr) - switch { - case apierrors.IsNotFound(err): - vr.SetName(name) - vr.SetNamespace(vgr.GetNamespace()) - vr.SetOwnerReferences([]metav1.OwnerReference{owner}) - spec := map[string]interface{}{ - "autoResync": false, - "replicationState": replicationState, - "volumeReplicationClass": volumeReplicationClass, - "dataSource": map[string]interface{}{ - "kind": "PersistentVolumeClaim", - "name": pvc.Name, - }, - } - if err := unstructured.SetNestedMap(vr.Object, spec, "spec"); err != nil { - return nil, fmt.Errorf("build VolumeReplication %q spec: %w", name, err) - } - if err := r.Create(ctx, vr); err != nil { - return nil, fmt.Errorf("create VolumeReplication %q: %w", name, err) - } - case err != nil: - return nil, fmt.Errorf("get VolumeReplication %q: %w", name, err) - default: - if current, _, _ := unstructured.NestedString(vr.Object, "spec", "replicationState"); current != replicationState { - if err := unstructured.SetNestedField(vr.Object, replicationState, "spec", "replicationState"); err != nil { - return nil, fmt.Errorf("set VolumeReplication %q replicationState: %w", name, err) - } - if err := r.Update(ctx, vr); err != nil { - return nil, fmt.Errorf("update VolumeReplication %q: %w", name, err) - } - } - } - result = append(result, *vr) - } - return result, nil -} - -// memberVolumeReplicationName deterministically names a group member's -// per-volume VolumeReplication from the group and the member PVC. -func memberVolumeReplicationName(groupName, pvcName string) string { - return groupName + "-" + pvcName -} - -// fanIn aggregates every member's VolumeReplication.status.conditions into the -// group's own status (design §4.3 step 3): Completed is the conjunction across -// members, Degraded and Resyncing are the disjunction, and status.lastSyncTime -// is the oldest of the members' lastSyncTime, since a group's recovery point is -// only as fresh as its slowest member. It also emits GroupReplicationDegraded -// the first time it observes Degraded after the group previously reported it -// false, and writes status.persistentVolumeClaimsRefList to the resolved -// membership. -func (r *VolumeGroupReplicationReconciler) fanIn( - ctx context.Context, - vgr *unstructured.Unstructured, - members []corev1.PersistentVolumeClaim, - memberVRs []unstructured.Unstructured, -) error { - wasDegraded := conditionStatus(vgr, "Degraded") == metav1.ConditionTrue - - completed := len(memberVRs) > 0 - degraded := false - resyncing := false - var oldestSync *time.Time - for i := range memberVRs { - vr := &memberVRs[i] - if conditionStatus(vr, "Completed") != metav1.ConditionTrue { - completed = false - } - if conditionStatus(vr, "Degraded") == metav1.ConditionTrue { - degraded = true - } - if conditionStatus(vr, "Resyncing") == metav1.ConditionTrue { - resyncing = true - } - if ts, found, _ := unstructured.NestedString(vr.Object, "status", "lastSyncTime"); found && ts != "" { - if parsed, err := time.Parse(time.RFC3339, ts); err == nil { - if oldestSync == nil || parsed.Before(*oldestSync) { - oldestSync = &parsed - } - } - } - } - - now := metav1.Now() - conditions := []interface{}{ - groupCondition("Completed", completed, now), - groupCondition("Degraded", degraded, now), - groupCondition("Resyncing", resyncing, now), - } - if err := unstructured.SetNestedSlice(vgr.Object, conditions, "status", "conditions"); err != nil { - return fmt.Errorf("set status.conditions: %w", err) - } - if oldestSync != nil { - if err := unstructured.SetNestedField(vgr.Object, oldestSync.UTC().Format(time.RFC3339), "status", "lastSyncTime"); err != nil { - return fmt.Errorf("set status.lastSyncTime: %w", err) - } - } - if completed { - replicationState, _, _ := unstructured.NestedString(vgr.Object, "spec", "replicationState") - if err := unstructured.SetNestedField(vgr.Object, replicationState, "status", "state"); err != nil { - return fmt.Errorf("set status.state: %w", err) - } - } - refs := make([]interface{}, 0, len(members)) - for i := range members { - refs = append(refs, map[string]interface{}{"name": members[i].Name}) - } - if err := unstructured.SetNestedSlice(vgr.Object, refs, "status", "persistentVolumeClaimsRefList"); err != nil { - return fmt.Errorf("set status.persistentVolumeClaimsRefList: %w", err) - } - - if err := r.Status().Update(ctx, vgr); err != nil { - return err - } - if degraded && !wasDegraded { - r.Recorder.Eventf(vgr, nil, corev1.EventTypeWarning, reasonGroupReplicationDegraded, reasonGroupReplicationDegraded, - "at least one group member reports Degraded") - } - return nil -} - -// conditionStatus returns a condition's status (metav1.ConditionTrue/False/ -// Unknown) on an unstructured VolumeReplication or VolumeGroupReplication, or -// "" if the object carries no condition of that type. -func conditionStatus(obj *unstructured.Unstructured, condType string) metav1.ConditionStatus { - raw, found, _ := unstructured.NestedSlice(obj.Object, "status", "conditions") - if !found { - return "" - } - for _, c := range raw { - cm, ok := c.(map[string]interface{}) - if !ok { - continue - } - if cm["type"] == condType { - if s, ok := cm["status"].(string); ok { - return metav1.ConditionStatus(s) - } - } - } - return "" -} - -// groupCondition builds one status.conditions entry. -func groupCondition(condType string, status bool, now metav1.Time) map[string]interface{} { - s := metav1.ConditionFalse - if status { - s = metav1.ConditionTrue - } - return map[string]interface{}{ - "type": condType, - "status": string(s), - "reason": "GroupMemberAggregation", - "message": "", - "lastTransitionTime": now.UTC().Format(time.RFC3339), - } -} - -func (r *VolumeGroupReplicationReconciler) SetupWithManager(mgr ctrl.Manager) error { - target := &unstructured.Unstructured{} - target.SetGroupVersionKind(volumeGroupReplicationGVK) - - return ctrl.NewControllerManagedBy(mgr). - For(target). - Named("volumegroupreplication"). - Complete(r) -} diff --git a/operator/internal/controller/volumegroupreplication_controller_unit_test.go b/operator/internal/controller/volumegroupreplication_controller_unit_test.go deleted file mode 100644 index 1ae2bee3f..000000000 --- a/operator/internal/controller/volumegroupreplication_controller_unit_test.go +++ /dev/null @@ -1,293 +0,0 @@ -/* -Copyright 2025. - -Licensed under the Apache License, Version 2.0 (the "License"); -you may not use this file except in compliance with the License. -You may obtain a copy of the License at - - http://www.apache.org/licenses/LICENSE-2.0 - -Unless required by applicable law or agreed to in writing, software -distributed under the License is distributed on an "AS IS" BASIS, -WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. -See the License for the specific language governing permissions and -limitations under the License. -*/ - -package controller - -import ( - "context" - "crypto/sha256" - "encoding/hex" - "encoding/json" - "net/http" - "strings" - "testing" - "time" - - corev1 "k8s.io/api/core/v1" - metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" - "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" - "k8s.io/apimachinery/pkg/types" - "k8s.io/client-go/tools/events" - ctrl "sigs.k8s.io/controller-runtime" - "sigs.k8s.io/controller-runtime/pkg/client" - "sigs.k8s.io/controller-runtime/pkg/client/fake" -) - -// The scenario matrix this file implements is -// docs/tests/test-plan-ramen-integration.md §1's U-01…U-03, verifying -// design-ramen-integration.md §4.3's fan-in aggregation. - -const vgrCluster = "55555555-5555-5555-5555-555555555555" - -// vgrUUID maps a short fixture id to a deterministic canonical UUID, the same -// way replication_test.go's sibling in csi-driver and the webhook package's -// vgsUUID both do, since atlas/lvol.ParseHandle accepts only canonical UUIDs. -func vgrUUID(id string) string { - sum := sha256.Sum256([]byte(id)) - h := hex.EncodeToString(sum[:16]) - return h[0:8] + "-" + h[8:12] + "-" + h[12:16] + "-" + h[16:20] + "-" + h[20:32] -} - -// vgrBackend serves the consistency-group resolve + membership reads. -type vgrBackend struct { - members []string // short fixture ids, converted through vgrUUID -} - -func (b vgrBackend) start(t *testing.T) { - t.Helper() - mux := http.NewServeMux() - mux.HandleFunc("/api/v2/clusters/"+vgrCluster+"/consistency-groups/", func(w http.ResponseWriter, r *http.Request) { - if strings.HasSuffix(r.URL.Path, "/members") { - rows := make([]map[string]any, 0, len(b.members)) - for _, m := range b.members { - rows = append(rows, map[string]any{"lvol_id": vgrUUID(m)}) - } - _ = json.NewEncoder(w).Encode(rows) - return - } - _ = json.NewEncoder(w).Encode([]map[string]any{ - {"id": "grp", "name": "grp1", "member_count": len(b.members)}, - }) - }) - srv := newAPIServer(t, mux.ServeHTTP) - t.Setenv("SIMPLYBLOCK_WEBAPI_BASE_URL", srv.URL) -} - -// vgrMember describes one group member: a bound PVC plus the per-volume -// VolumeReplication a prior fan-out already created for it, pre-seeded with -// the status this test wants fan-in to read. -type vgrMember struct { - pvcName string - volumeID string // short fixture id, converted through vgrUUID - completed bool - degraded bool - lastSync time.Time -} - -func newVGRReconciler(t *testing.T, objects ...client.Object) (*VolumeGroupReplicationReconciler, client.Client) { - t.Helper() - scheme := newTestScheme(t, corev1.AddToScheme) - statusSubresource := &unstructured.Unstructured{} - statusSubresource.SetGroupVersionKind(volumeGroupReplicationGVK) - cl := fake.NewClientBuilder(). - WithScheme(scheme). - WithStatusSubresource(statusSubresource). - WithObjects(objects...). - Build() - return &VolumeGroupReplicationReconciler{ - Client: cl, - Scheme: scheme, - Recorder: events.NewFakeRecorder(32), - }, cl -} - -func vgrClass(name string) *unstructured.Unstructured { - class := &unstructured.Unstructured{} - class.SetGroupVersionKind(volumeGroupReplicationClassGVK) - class.SetName(name) - _ = unstructured.SetNestedField(class.Object, "csi.simplyblock.io", "spec", "provisioner") - return class -} - -// vgrGroup builds the VolumeGroupReplication under test, selecting every -// member by the shared "app: grp1" label. -func vgrGroup(name, className, vrClassName, replicationState string) *unstructured.Unstructured { - vgr := &unstructured.Unstructured{} - vgr.SetGroupVersionKind(volumeGroupReplicationGVK) - vgr.SetName(name) - vgr.SetNamespace("default") - vgr.SetUID(types.UID("vgr-uid-" + name)) - spec := map[string]interface{}{ - "external": true, - "autoResync": false, - "replicationState": replicationState, - "volumeGroupReplicationClassName": className, - "volumeReplicationClassName": vrClassName, - "source": map[string]interface{}{ - "selector": map[string]interface{}{ - "matchLabels": map[string]interface{}{"app": "grp1"}, - }, - }, - } - _ = unstructured.SetNestedMap(vgr.Object, spec, "spec") - return vgr -} - -// vgrPVCAndPV builds a bound PVC/PV pair: labeled for both the group -// selector and the consistency-group membership check, backed by a -// simplyblock CSI volume handle. -func vgrPVCAndPV(pvcName, volumeID string) (*corev1.PersistentVolumeClaim, *corev1.PersistentVolume) { - pvName := "pv-" + pvcName - pvc := &corev1.PersistentVolumeClaim{ - ObjectMeta: metav1.ObjectMeta{ - Name: pvcName, - Namespace: "default", - Labels: map[string]string{ - "app": "grp1", - consistencyGroupLabel: "grp1", - }, - }, - Spec: corev1.PersistentVolumeClaimSpec{VolumeName: pvName}, - } - pv := &corev1.PersistentVolume{ - ObjectMeta: metav1.ObjectMeta{Name: pvName}, - Spec: corev1.PersistentVolumeSpec{PersistentVolumeSource: corev1.PersistentVolumeSource{ - CSI: &corev1.CSIPersistentVolumeSource{ - Driver: "csi.simplyblock.io", - VolumeHandle: vgrCluster + ":pool:" + vgrUUID(volumeID), - }, - }}, - } - return pvc, pv -} - -// vgrMemberVR builds the per-volume VolumeReplication a prior fan-out -// already created for one group member, with the status this test wants -// fan-in to aggregate. -func vgrMemberVR(name, groupName, pvcName string, m vgrMember) *unstructured.Unstructured { - vr := &unstructured.Unstructured{} - vr.SetGroupVersionKind(volumeReplicationGVK) - vr.SetName(name) - vr.SetNamespace("default") - vr.SetOwnerReferences([]metav1.OwnerReference{{ - APIVersion: volumeGroupReplicationGVK.GroupVersion().String(), - Kind: volumeGroupReplicationGVK.Kind, - Name: groupName, - UID: types.UID("vgr-uid-" + groupName), - }}) - _ = unstructured.SetNestedField(vr.Object, pvcName, "spec", "dataSource", "name") - _ = unstructured.SetNestedField(vr.Object, "PersistentVolumeClaim", "spec", "dataSource", "kind") - _ = unstructured.SetNestedField(vr.Object, "primary", "spec", "replicationState") - - cond := func(condType string, status bool) map[string]interface{} { - s := metav1.ConditionFalse - if status { - s = metav1.ConditionTrue - } - return map[string]interface{}{"type": condType, "status": string(s), "reason": "Test", "message": ""} - } - conditions := []interface{}{ - cond("Completed", m.completed), - cond("Degraded", m.degraded), - cond("Resyncing", false), - } - _ = unstructured.SetNestedSlice(vr.Object, conditions, "status", "conditions") - _ = unstructured.SetNestedField(vr.Object, m.lastSync.UTC().Format(time.RFC3339), "status", "lastSyncTime") - return vr -} - -// runVGRReconcile builds the group, its class, the member PVC/PV pairs, and -// their pre-seeded VolumeReplications, reconciles once, and returns the -// group's own status.conditions and status.lastSyncTime for assertion. -func runVGRReconcile(t *testing.T, members []vgrMember) (conditions map[string]string, lastSyncTime string) { - t.Helper() - vgrBackend{members: memberIDs(members)}.start(t) - - group := vgrGroup("vgr1", "sb-group-class", "sb-vr-class", "primary") - objs := make([]client.Object, 0, 2+2*len(members)) - objs = append(objs, group, vgrClass("sb-group-class")) - for _, m := range members { - pvc, pv := vgrPVCAndPV(m.pvcName, m.volumeID) - objs = append(objs, pvc, pv) - } - r, cl := newVGRReconciler(t, objs...) - for _, m := range members { - vr := vgrMemberVR("vgr1-"+m.pvcName, "vgr1", m.pvcName, m) - if err := cl.Create(context.Background(), vr); err != nil { - t.Fatalf("create member VolumeReplication: %v", err) - } - } - - req := ctrl.Request{NamespacedName: types.NamespacedName{Namespace: "default", Name: "vgr1"}} - if _, err := r.Reconcile(context.Background(), req); err != nil { - t.Fatalf("unexpected error: %v", err) - } - - got := &unstructured.Unstructured{} - got.SetGroupVersionKind(volumeGroupReplicationGVK) - if err := cl.Get(context.Background(), types.NamespacedName{Namespace: "default", Name: "vgr1"}, got); err != nil { - t.Fatalf("get VolumeGroupReplication: %v", err) - } - - conditions = map[string]string{} - rawConds, _, _ := unstructured.NestedSlice(got.Object, "status", "conditions") - for _, c := range rawConds { - cm, ok := c.(map[string]interface{}) - if !ok { - continue - } - conditions[cm["type"].(string)] = cm["status"].(string) - } - lastSyncTime, _, _ = unstructured.NestedString(got.Object, "status", "lastSyncTime") - return conditions, lastSyncTime -} - -func memberIDs(members []vgrMember) []string { - ids := make([]string, 0, len(members)) - for _, m := range members { - ids = append(ids, m.volumeID) - } - return ids -} - -// U-01: every member Completed/not-Degraded yields a group reporting the same. -func TestVolumeGroupReplication_AllMembersHealthyYieldsGroupHealthy(t *testing.T) { - conditions, _ := runVGRReconcile(t, []vgrMember{ - {pvcName: "pvc-a", volumeID: "v1", completed: true, degraded: false, lastSync: time.Now()}, - {pvcName: "pvc-b", volumeID: "v2", completed: true, degraded: false, lastSync: time.Now()}, - }) - if conditions["Completed"] != string(metav1.ConditionTrue) { - t.Errorf("Completed = %q, want True", conditions["Completed"]) - } - if conditions["Degraded"] != string(metav1.ConditionFalse) { - t.Errorf("Degraded = %q, want False", conditions["Degraded"]) - } -} - -// U-02: one member Degraded yields a group reporting Degraded. -func TestVolumeGroupReplication_OneMemberDegradedYieldsGroupDegraded(t *testing.T) { - conditions, _ := runVGRReconcile(t, []vgrMember{ - {pvcName: "pvc-a", volumeID: "v1", completed: true, degraded: false, lastSync: time.Now()}, - {pvcName: "pvc-b", volumeID: "v2", completed: true, degraded: true, lastSync: time.Now()}, - }) - if conditions["Degraded"] != string(metav1.ConditionTrue) { - t.Errorf("Degraded = %q, want True", conditions["Degraded"]) - } -} - -// U-03: members with differing lastSyncTime yield the oldest, not the newest. -func TestVolumeGroupReplication_LastSyncTimeIsTheOldestMember(t *testing.T) { - older := time.Now().Add(-1 * time.Hour).Truncate(time.Second) - newer := time.Now().Truncate(time.Second) - _, lastSyncTime := runVGRReconcile(t, []vgrMember{ - {pvcName: "pvc-a", volumeID: "v1", completed: true, degraded: false, lastSync: older}, - {pvcName: "pvc-b", volumeID: "v2", completed: true, degraded: false, lastSync: newer}, - }) - want := older.UTC().Format(time.RFC3339) - if lastSyncTime != want { - t.Errorf("status.lastSyncTime = %q, want the oldest member's %q", lastSyncTime, want) - } -} diff --git a/operator/internal/webhook/volumegroupreplication_validator.go b/operator/internal/webhook/volumegroupreplication_validator.go deleted file mode 100644 index fb3a8a99c..000000000 --- a/operator/internal/webhook/volumegroupreplication_validator.go +++ /dev/null @@ -1,201 +0,0 @@ -package webhook - -import ( - "context" - "encoding/json" - "fmt" - "net/http" - - corev1 "k8s.io/api/core/v1" - metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" - "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" - "k8s.io/apimachinery/pkg/runtime" - "k8s.io/apimachinery/pkg/runtime/schema" - "sigs.k8s.io/controller-runtime/pkg/client" - logf "sigs.k8s.io/controller-runtime/pkg/log" - "sigs.k8s.io/controller-runtime/pkg/webhook/admission" - - "github.com/simplyblock/simplyblock-operator/internal/webapi" -) - -// volumeGroupReplicationGVK and volumeGroupReplicationClassGVK name the -// csi-addons kinds this repository vendors no Go type for (design -// §4.1). Duplicated, not shared, with internal/controller's identical pair: -// each package that needs an unstructured kind's GVK declares its own, -// matching the layering the design's own §4.4/§4.3 keep as two independent -// enforcement points (design-consistency-groups.md §9's "last line of -// defense, not the first"). -var ( - volumeGroupReplicationGVK = schema.GroupVersionKind{ - Group: "replication.storage.openshift.io", Version: "v1alpha1", Kind: "VolumeGroupReplication", - } - volumeGroupReplicationClassGVK = schema.GroupVersionKind{ - Group: "replication.storage.openshift.io", Version: "v1alpha1", Kind: "VolumeGroupReplicationClass", - } -) - -// +kubebuilder:webhook:path=/validate-replication-storage-openshift-io-v1alpha1-volumegroupreplication,mutating=false,failurePolicy=fail,sideEffects=None,groups=replication.storage.openshift.io,resources=volumegroupreplications,verbs=create,versions=v1alpha1,name=vvolumegroupreplication.simplyblock.io,admissionReviewVersions=v1 - -// VolumeGroupReplicationValidator rejects, at kubectl apply, a -// VolumeGroupReplication whose selector does not resolve to exactly one -// consistency group's current membership, so VolumeGroupReplicationReconciler -// (design §4.3) never has to reconcile an object that could never fan out -// correctly. It makes the same two checks, with the same two dispositions, -// design-consistency-groups.md §9.4 already established for -// VolumeGroupSnapshot: -// -// - The label check is static and fail-closed: every selected PVC must carry -// the same non-empty consistency-group label. A selector that spans groups, -// matches an unlabeled PVC, or matches nothing is rejected outright. -// - The membership check is backend and fail-open: it resolves the group by -// its label value and rejects the object unless the selected set equals the -// group's current membership. When membership cannot be determined (the -// backend is unreachable, or a selected PVC is not yet bound), it admits and -// the reconciler backstops it on its own next reconcile (design §4.3). -type VolumeGroupReplicationValidator struct { - Client client.Client - APIClient *webapi.Client -} - -// ownsVolumeGroupReplication reports whether the object is external and -// attributed to this driver's class, the same ownership signal the -// reconciler applies (design §4.3). Every undecidable case admits, because -// with failurePolicy=fail a webhook error would block every -// VolumeGroupReplication in the cluster, foreign drivers included. -func (v *VolumeGroupReplicationValidator) ownsVolumeGroupReplication( - ctx context.Context, vgr *unstructured.Unstructured, -) (bool, string) { - external, _, _ := unstructured.NestedBool(vgr.Object, "spec", "external") - if !external { - return false, "spec.external is not true; the generic controller-manager reconciles this one" - } - className, _, _ := unstructured.NestedString(vgr.Object, "spec", "volumeGroupReplicationClassName") - if className == "" { - return false, "no volumeGroupReplicationClassName; not attributable to this driver" - } - class := &unstructured.Unstructured{} - class.SetGroupVersionKind(volumeGroupReplicationClassGVK) - if err := v.Client.Get(ctx, client.ObjectKey{Name: className}, class); err != nil { - return false, fmt.Sprintf("volume group replication class %q not readable; not validating", className) - } - provisioner, _, _ := unstructured.NestedString(class.Object, "spec", "provisioner") - if provisioner != "csi.simplyblock.io" { - return false, fmt.Sprintf("class %q belongs to provisioner %q; not validating", className, provisioner) - } - return true, "" -} - -func (v *VolumeGroupReplicationValidator) Handle(ctx context.Context, req admission.Request) admission.Response { - log := logf.FromContext(ctx).WithValues("volumegroupreplication", req.Name, "namespace", req.Namespace) - - vgr := &unstructured.Unstructured{} - if err := json.Unmarshal(req.Object.Raw, vgr); err != nil { - return admission.Errored(http.StatusBadRequest, err) - } - - if ours, reason := v.ownsVolumeGroupReplication(ctx, vgr); !ours { - return admission.Allowed(reason) - } - - pvcs, groupName, denied := v.labelCheck(ctx, req.Namespace, vgr) - if denied != nil { - return *denied - } - - clusterUUID, selected, determinable, err := v.selectedLvols(ctx, pvcs) - if err != nil { - log.Error(err, "cannot resolve selected volumes; admitting (reconciler backstops)") - return admission.Allowed("group membership undeterminable; deferring to the reconciler") - } - if !determinable { - return admission.Allowed("a selected PVC is not yet bound; deferring to the reconciler") - } - group, err := v.APIClient.GetConsistencyGroupByName(ctx, clusterUUID, groupName) - if err != nil || group == nil { - if err != nil { - log.Error(err, "cannot resolve consistency group; admitting (reconciler backstops)") - } - return admission.Allowed("consistency group not resolvable; deferring to the reconciler") - } - members, err := v.APIClient.GetConsistencyGroupMembers(ctx, clusterUUID, group.UUID) - if err != nil { - log.Error(err, "cannot read group membership; admitting (reconciler backstops)") - return admission.Allowed("group membership unreadable; deferring to the reconciler") - } - if !sameStringSet(selected, members) { - return admission.Denied(fmt.Sprintf( - "selector resolves to %d volume(s) but consistency group %q has %d member(s); "+ - "the selector must equal the group's current membership", - len(selected), groupName, len(members))) - } - return admission.Allowed("selector equals the consistency group's membership") -} - -// labelCheck resolves spec.source.selector to PVCs and enforces the -// fail-closed label rule, the same as VolumeGroupSnapshotValidator.labelCheck. -func (v *VolumeGroupReplicationValidator) labelCheck( - ctx context.Context, namespace string, vgr *unstructured.Unstructured, -) ([]corev1.PersistentVolumeClaim, string, *admission.Response) { - deny := func(msg string) *admission.Response { r := admission.Denied(msg); return &r } - - selMap, found, err := unstructured.NestedMap(vgr.Object, "spec", "source", "selector") - if err != nil || !found { - return nil, "", deny("VolumeGroupReplication has no spec.source.selector") - } - var labelSelector metav1.LabelSelector - if err := runtime.DefaultUnstructuredConverter.FromUnstructured(selMap, &labelSelector); err != nil { - return nil, "", deny(fmt.Sprintf("invalid label selector: %v", err)) - } - sel, err := metav1.LabelSelectorAsSelector(&labelSelector) - if err != nil { - return nil, "", deny(fmt.Sprintf("invalid label selector: %v", err)) - } - var pvcList corev1.PersistentVolumeClaimList - if err := v.Client.List(ctx, &pvcList, - client.InNamespace(namespace), client.MatchingLabelsSelector{Selector: sel}); err != nil { - return nil, "", deny(fmt.Sprintf("cannot list PVCs for the selector: %v", err)) - } - if len(pvcList.Items) == 0 { - return nil, "", deny("selector matches no PersistentVolumeClaim") - } - groupName := "" - for i := range pvcList.Items { - val := pvcList.Items[i].Labels[consistencyGroupLabel] - if val == "" { - return nil, "", deny(fmt.Sprintf( - "PVC %q is not labeled %s; a group replication's selector must match only group members", - pvcList.Items[i].Name, consistencyGroupLabel)) - } - if groupName == "" { - groupName = val - } else if val != groupName { - return nil, "", deny(fmt.Sprintf( - "selector spans two consistency groups (%q and %q); it must resolve to exactly one", - groupName, val)) - } - } - return pvcList.Items, groupName, nil -} - -// selectedLvols maps each selected PVC to its backing lvol UUID, the same as -// VolumeGroupSnapshotValidator.selectedLvols. -func (v *VolumeGroupReplicationValidator) selectedLvols( - ctx context.Context, pvcs []corev1.PersistentVolumeClaim, -) (clusterUUID string, lvols []string, determinable bool, err error) { - for i := range pvcs { - pvName := pvcs[i].Spec.VolumeName - if pvName == "" { - return "", nil, false, nil - } - cluster, _, volume, ok, err := pvVolumeHandle(ctx, v.Client, pvName) - if err != nil { - return "", nil, false, err - } - if !ok { - return "", nil, false, nil - } - clusterUUID = cluster - lvols = append(lvols, volume) - } - return clusterUUID, lvols, true, nil -} diff --git a/operator/internal/webhook/volumegroupreplication_validator_test.go b/operator/internal/webhook/volumegroupreplication_validator_test.go deleted file mode 100644 index 0ebef8e0c..000000000 --- a/operator/internal/webhook/volumegroupreplication_validator_test.go +++ /dev/null @@ -1,193 +0,0 @@ -package webhook - -import ( - "context" - "crypto/sha256" - "encoding/hex" - "encoding/json" - "net/http" - "net/http/httptest" - "strings" - "testing" - - corev1 "k8s.io/api/core/v1" - metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" - "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" - "k8s.io/apimachinery/pkg/runtime" - crclient "sigs.k8s.io/controller-runtime/pkg/client" - "sigs.k8s.io/controller-runtime/pkg/client/fake" - "sigs.k8s.io/controller-runtime/pkg/webhook/admission" - - "github.com/simplyblock/simplyblock-operator/internal/webapi" -) - -// The scenario matrix this file implements is -// docs/tests/test-plan-ramen-integration.md §1's U-04…U-05, verifying -// design-ramen-integration.md §4.4's admission webhook. - -const vgrValCluster = "66666666-6666-6666-6666-666666666666" - -func vgrValUUID(id string) string { - sum := sha256.Sum256([]byte(id)) - h := hex.EncodeToString(sum[:16]) - return h[0:8] + "-" + h[8:12] + "-" + h[12:16] + "-" + h[16:20] + "-" + h[20:32] -} - -type vgrValPVC struct { - name string - cgLabel string - volumeID string - bound bool -} - -type vgrValBackend struct { - members []string - fail bool -} - -func (b vgrValBackend) server(t *testing.T) string { - t.Helper() - mux := http.NewServeMux() - mux.HandleFunc("/api/v2/clusters/"+vgrValCluster+"/consistency-groups/", func(w http.ResponseWriter, r *http.Request) { - if b.fail { - w.WriteHeader(http.StatusInternalServerError) - return - } - if strings.HasSuffix(r.URL.Path, "/members") { - rows := make([]map[string]any, 0, len(b.members)) - for _, m := range b.members { - rows = append(rows, map[string]any{"lvol_id": vgrValUUID(m)}) - } - _ = json.NewEncoder(w).Encode(rows) - return - } - _ = json.NewEncoder(w).Encode([]map[string]any{ - {"id": "grp", "name": "grp1", "member_count": len(b.members)}, - }) - }) - srv := httptest.NewServer(mux) - t.Cleanup(srv.Close) - return srv.URL -} - -func newVGRValidator(t *testing.T, pvcs []vgrValPVC, apiURL string) *VolumeGroupReplicationValidator { - t.Helper() - scheme := runtime.NewScheme() - if err := corev1.AddToScheme(scheme); err != nil { - t.Fatalf("add corev1: %v", err) - } - var objs []crclient.Object - for _, p := range pvcs { - labels := map[string]string{"app": "grp1"} - if p.cgLabel != "" { - labels[consistencyGroupLabel] = p.cgLabel - } - pvc := &corev1.PersistentVolumeClaim{ - ObjectMeta: metav1.ObjectMeta{Name: p.name, Namespace: "sb", Labels: labels}, - } - if p.bound { - pvName := "pv-" + p.name - pvc.Spec.VolumeName = pvName - objs = append(objs, &corev1.PersistentVolume{ - ObjectMeta: metav1.ObjectMeta{Name: pvName}, - Spec: corev1.PersistentVolumeSpec{PersistentVolumeSource: corev1.PersistentVolumeSource{ - CSI: &corev1.CSIPersistentVolumeSource{ - Driver: "csi.simplyblock.io", - VolumeHandle: vgrValCluster + ":pool:" + vgrValUUID(p.volumeID), - }, - }}, - }) - } - objs = append(objs, pvc) - } - class := &unstructured.Unstructured{} - class.SetGroupVersionKind(volumeGroupReplicationClassGVK) - class.SetName("sb-group-class") - _ = unstructured.SetNestedField(class.Object, "csi.simplyblock.io", "spec", "provisioner") - objs = append(objs, class) - cl := fake.NewClientBuilder().WithScheme(scheme).WithObjects(objs...).Build() - return &VolumeGroupReplicationValidator{Client: cl, APIClient: webapi.NewClient(apiURL)} -} - -func vgrValRequest(t *testing.T, selectorValue string) admission.Request { - t.Helper() - vgr := &unstructured.Unstructured{} - vgr.SetGroupVersionKind(volumeGroupReplicationGVK) - vgr.SetName("gen") - vgr.SetNamespace("sb") - _ = unstructured.SetNestedMap(vgr.Object, map[string]interface{}{ - "external": true, - "autoResync": false, - "replicationState": "primary", - "volumeGroupReplicationClassName": "sb-group-class", - "source": map[string]interface{}{ - "selector": map[string]interface{}{ - "matchLabels": map[string]interface{}{"app": selectorValue}, - }, - }, - }, "spec") - raw, err := json.Marshal(vgr.Object) - if err != nil { - t.Fatalf("marshal VolumeGroupReplication: %v", err) - } - req := admission.Request{} - req.Object = runtime.RawExtension{Raw: raw} - return req -} - -func TestVolumeGroupReplicationValidator(t *testing.T) { - tests := []struct { - name string - pvcs []vgrValPVC - selector string - backend vgrValBackend - allowed bool - }{ - { - name: "selector equals membership is admitted", - pvcs: []vgrValPVC{ - {name: "a", cgLabel: "grp1", volumeID: "v1", bound: true}, - {name: "b", cgLabel: "grp1", volumeID: "v2", bound: true}, - }, - selector: "grp1", - backend: vgrValBackend{members: []string{"v1", "v2"}}, - allowed: true, - }, - { - name: "selector resolving to a subset of the group is rejected", - pvcs: []vgrValPVC{ - {name: "a", cgLabel: "grp1", volumeID: "v1", bound: true}, - }, - selector: "grp1", - backend: vgrValBackend{members: []string{"v1", "v2"}}, - allowed: false, - }, - { - name: "selector spanning two groups is rejected", - pvcs: []vgrValPVC{ - {name: "a", cgLabel: "grp1", volumeID: "v1", bound: true}, - {name: "b", cgLabel: "grp2", volumeID: "v2", bound: true}, - }, - selector: "grp1", - allowed: false, - }, - { - name: "backend unreachable admits (fail-open)", - pvcs: []vgrValPVC{ - {name: "a", cgLabel: "grp1", volumeID: "v1", bound: true}, - }, - selector: "grp1", - backend: vgrValBackend{fail: true}, - allowed: true, - }, - } - for _, tc := range tests { - t.Run(tc.name, func(t *testing.T) { - v := newVGRValidator(t, tc.pvcs, tc.backend.server(t)) - resp := v.Handle(context.Background(), vgrValRequest(t, tc.selector)) - if resp.Allowed != tc.allowed { - t.Fatalf("Allowed = %v, want %v (msg: %s)", resp.Allowed, tc.allowed, resp.Result.Message) - } - }) - } -} From a1bc4872d2a5a8313a022cefd9018ba0c16446b5 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Sat, 26 Sep 2026 01:31:09 +0100 Subject: [PATCH 158/206] feat(csi-addons): drive VolumeGroupReplication through the stock controller --- atlas-lib/controlplane/consistencygroups.go | 95 ++ atlas-lib/controlplane/groupreplication.go | 174 +++ atlas-lib/internal/cpapi/cpapi.gen.go | 1353 +++++++++++++++-- atlas-lib/lvol/grouphandle.go | 63 + atlas-lib/lvol/grouphandle_test.go | 80 + csi-driver/go.mod | 3 +- csi-driver/go.sum | 4 +- .../csi/controller/mock_controlplane_test.go | 118 ++ .../internal/csi/controller/replication.go | 81 + .../csi/controller/replication_group_test.go | 148 ++ csi-driver/internal/csi/controller/server.go | 5 + .../internal/csi/controller/volumegroup.go | 116 ++ .../csi/controller/volumegroup_test.go | 129 ++ .../csi/csiaddons/identity/identity.go | 27 + .../csi/csiaddons/identity/identity_test.go | 29 + csi-driver/internal/driver/driver.go | 7 + operator/dist/install.yaml | 57 - shared/openapi.json | 421 ++++- 18 files changed, 2750 insertions(+), 160 deletions(-) create mode 100644 atlas-lib/controlplane/consistencygroups.go create mode 100644 atlas-lib/controlplane/groupreplication.go create mode 100644 atlas-lib/lvol/grouphandle.go create mode 100644 atlas-lib/lvol/grouphandle_test.go create mode 100644 csi-driver/internal/csi/controller/replication_group_test.go create mode 100644 csi-driver/internal/csi/controller/volumegroup.go create mode 100644 csi-driver/internal/csi/controller/volumegroup_test.go diff --git a/atlas-lib/controlplane/consistencygroups.go b/atlas-lib/controlplane/consistencygroups.go new file mode 100644 index 000000000..d45464253 --- /dev/null +++ b/atlas-lib/controlplane/consistencygroups.go @@ -0,0 +1,95 @@ +// Consistency-group reads the csi-addons VolumeGroup service needs: resolving the +// backend consistency group a set of member volumes already belongs to (design +// design-csi-addons-replication.md §14.3). A group is formed at provisioning by +// the storage.simplyblock.io/consistency-group label; this only reads it back. +package controlplane + +import ( + "context" + "fmt" + + "github.com/google/uuid" + + "github.com/simplyblock/atlas/internal/cpapi" +) + +// ConsistencyGroupForLvols returns the id of the backend consistency group in +// clusterID whose current (open-epoch) membership is exactly lvolIDs. +// +// The members provisioned under one storage.simplyblock.io/consistency-group +// label share one backend group, so the VolumeGroup service resolves the group +// by its members rather than by the name it is handed (which is the caller's own +// generated name, not the group's). Returns an error when no group's membership +// matches exactly, so a partial or mixed selection is never silently grouped. +func (c *Client) ConsistencyGroupForLvols(ctx context.Context, clusterID string, lvolIDs []string) (string, error) { + cluster, err := parseUUID("cluster id", clusterID) + if err != nil { + return "", err + } + if len(lvolIDs) == 0 { + return "", fmt.Errorf("no volumes given to resolve a consistency group") + } + want := make(map[string]bool, len(lvolIDs)) + for _, id := range lvolIDs { + want[id] = true + } + + listResp, err := c.api.ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetWithResponse( + ctx, cluster, &cpapi.ClustersConsistencyGroupsListApiV2ClustersClusterIdConsistencyGroupsGetParams{}) + if err != nil { + return "", fmt.Errorf("list consistency groups in cluster %s: %w", clusterID, err) + } + groups, err := payload("consistency groups in cluster "+clusterID, + listResp.JSON200, listResp.StatusCode(), listResp.Body) + if err != nil { + return "", err + } + + for _, g := range *groups { + members, err := c.consistencyGroupMembers(ctx, cluster, g.Id) + if err != nil { + return "", err + } + if sameStringSet(members, want) { + return g.Id.String(), nil + } + } + return "", fmt.Errorf("no consistency group in cluster %s has exactly the %d requested member(s)", + clusterID, len(lvolIDs)) +} + +// consistencyGroupMembers returns a group's current open-epoch member lvol ids. +func (c *Client) consistencyGroupMembers(ctx context.Context, cluster, group uuid.UUID) ([]string, error) { + resp, err := c.api.ClustersConsistencyGroupsMembersApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersGetWithResponse( + ctx, cluster, group) + if err != nil { + return nil, fmt.Errorf("members of consistency group %s: %w", group, err) + } + dtos, err := payload("members of consistency group "+group.String(), + resp.JSON200, resp.StatusCode(), resp.Body) + if err != nil { + return nil, err + } + open := make([]string, 0, len(*dtos)) + for _, m := range *dtos { + if m.RemovedSeq == 0 { + open = append(open, m.LvolId) + } + } + return open, nil +} + +// sameStringSet reports whether ids is exactly the set want (same length, same +// elements), so a resolved group is neither a superset nor a subset of the +// requested members. +func sameStringSet(ids []string, want map[string]bool) bool { + if len(ids) != len(want) { + return false + } + for _, id := range ids { + if !want[id] { + return false + } + } + return true +} diff --git a/atlas-lib/controlplane/groupreplication.go b/atlas-lib/controlplane/groupreplication.go new file mode 100644 index 000000000..4b806d867 --- /dev/null +++ b/atlas-lib/controlplane/groupreplication.go @@ -0,0 +1,174 @@ +// Group replication verbs on a consistency-group handle (design +// design-csi-addons-replication.md §14.4): the whole group replicates as one +// unit through the cluster-scoped /consistency-groups/{id}/replication/* +// endpoints, the group twins of the per-volume verbs in replication.go. The +// driver's Replication verbs route a cg: handle here (§14.4). +package controlplane + +import ( + "context" + "fmt" + "net/http" + + "github.com/google/uuid" + + "github.com/simplyblock/atlas/internal/cpapi" + "github.com/simplyblock/atlas/lvol" +) + +// groupIDs parses a group handle into the cluster and group UUIDs the v2 path +// parameters require. +func groupIDs(gh lvol.GroupHandle) (cluster, group uuid.UUID, err error) { + cluster, err = parseUUID("cluster id", gh.ClusterID) + if err != nil { + return + } + group, err = parseUUID("consistency group id", gh.GroupID) + return +} + +// EnableGroupReplication attaches the consistency group to a group replication +// policy, so the whole group replicates as one unit. +func (c *Client) EnableGroupReplication(ctx context.Context, gh lvol.GroupHandle, policyID string) error { + cluster, group, err := groupIDs(gh) + if err != nil { + return err + } + policy, err := parseUUID("replication policy id", policyID) + if err != nil { + return err + } + resp, err := c.api.ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutWithResponse( + ctx, cluster, group, cpapi.ConsistencyGroupReplicationIntentDTO{ReplicationPolicyId: &policy}) + if err != nil { + return fmt.Errorf("enable group replication %s: %w", gh.Handle(), err) + } + if code := resp.StatusCode(); code != http.StatusOK && code != http.StatusNoContent { + return respError("enable group replication "+string(gh.Handle()), code, resp.Body) + } + return nil +} + +// DisableGroupReplication detaches the consistency group from its policy, +// stopping replication without dissolving the group. +func (c *Client) DisableGroupReplication(ctx context.Context, gh lvol.GroupHandle) error { + cluster, group, err := groupIDs(gh) + if err != nil { + return err + } + // ReplicationPolicyId is not omitempty on the intent DTO, so a nil pointer + // marshals as an explicit null. The backend reads that null as a detach, + // distinct from an absent key. + resp, err := c.api.ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutWithResponse( + ctx, cluster, group, cpapi.ConsistencyGroupReplicationIntentDTO{ReplicationPolicyId: nil}) + if err != nil { + return fmt.Errorf("disable group replication %s: %w", gh.Handle(), err) + } + if code := resp.StatusCode(); code != http.StatusOK && code != http.StatusNoContent { + return respError("disable group replication "+string(gh.Handle()), code, resp.Body) + } + return nil +} + +// PromoteGroup fails the whole group over as one unit: every member is cloned +// from the same group-snapshot generation on the target, atomically. +func (c *Client) PromoteGroup(ctx context.Context, gh lvol.GroupHandle) error { + cluster, group, err := groupIDs(gh) + if err != nil { + return err + } + resp, err := c.api.ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPostWithResponse( + ctx, cluster, group) + if err != nil { + return fmt.Errorf("promote group %s: %w", gh.Handle(), err) + } + if code := resp.StatusCode(); code != http.StatusOK && code != http.StatusNoContent { + return respError("promote group "+string(gh.Handle()), code, resp.Body) + } + return nil +} + +// DemoteGroup demotes the whole group: quiesce every member, ship one final +// group snapshot, confirm, fence all. Re-drivable, not queued: done=false while +// any member is still converging, so the caller calls again. +func (c *Client) DemoteGroup(ctx context.Context, gh lvol.GroupHandle) (bool, error) { + cluster, group, err := groupIDs(gh) + if err != nil { + return false, err + } + resp, err := c.api.ClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePostWithResponse( + ctx, cluster, group) + if err != nil { + return false, fmt.Errorf("demote group %s: %w", gh.Handle(), err) + } + switch code := resp.StatusCode(); code { + case http.StatusNoContent: + return true, nil + case http.StatusAccepted: + return false, nil + default: + return false, respError("demote group "+string(gh.Handle()), code, resp.Body) + } +} + +// ResyncGroup reverses the shipping direction for the whole group. It never cuts +// over, matching the per-volume ResyncVolume. sourceClusterID selects the source +// explicitly; "" leaves it unset. +func (c *Client) ResyncGroup(ctx context.Context, gh lvol.GroupHandle, sourceClusterID string) error { + cluster, group, err := groupIDs(gh) + if err != nil { + return err + } + var body cpapi.GroupFailbackParams + if sourceClusterID != "" { + id, err := parseUUID("source cluster id", sourceClusterID) + if err != nil { + return err + } + body.SourceClusterId = &id + } + resp, err := c.api.ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostWithResponse( + ctx, cluster, group, body) + if err != nil { + return fmt.Errorf("resync group %s: %w", gh.Handle(), err) + } + if code := resp.StatusCode(); code != http.StatusOK && code != http.StatusNoContent { + return respError("resync group "+string(gh.Handle()), code, resp.Body) + } + return nil +} + +// GetGroupReplicationInfo returns the group's aggregate replication status +// (oldest recovery point, worst lag, and health; design §14.6), in the same +// shape as the per-volume status so the driver's Info verb maps it uniformly. +func (c *Client) GetGroupReplicationInfo(ctx context.Context, gh lvol.GroupHandle) (ReplicationStatus, error) { + cluster, group, err := groupIDs(gh) + if err != nil { + return ReplicationStatus{}, err + } + resp, err := c.api.ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetWithResponse( + ctx, cluster, group) + if err != nil { + return ReplicationStatus{}, fmt.Errorf("group replication status %s: %w", gh.Handle(), err) + } + d, err := payload("group replication status "+string(gh.Handle()), resp.JSON200, resp.StatusCode(), resp.Body) + if err != nil { + return ReplicationStatus{}, err + } + return ReplicationStatus{ + Role: string(d.Role), + State: string(d.State), + LastReplicatedAt: d.LastReplicatedAt, + LagSeconds: d.LagSeconds, + OutstandingCount: derefInt(d.OutstandingCount), + OutstandingBytes: derefInt(d.OutstandingBytes), + Resyncing: d.Resyncing != nil && *d.Resyncing, + }, nil +} + +func derefInt(p *int) int { + if p == nil { + return 0 + } + return *p +} diff --git a/atlas-lib/internal/cpapi/cpapi.gen.go b/atlas-lib/internal/cpapi/cpapi.gen.go index 7bbeb5c91..7aec772a2 100644 --- a/atlas-lib/internal/cpapi/cpapi.gen.go +++ b/atlas-lib/internal/cpapi/cpapi.gen.go @@ -144,6 +144,60 @@ func (e ClusterParamsHaType) Valid() bool { } } +// Defines values for ConsistencyGroupReplicationStatusDTORole. +const ( + ConsistencyGroupReplicationStatusDTORoleFailedOver ConsistencyGroupReplicationStatusDTORole = "failed_over" + ConsistencyGroupReplicationStatusDTORoleNone ConsistencyGroupReplicationStatusDTORole = "none" + ConsistencyGroupReplicationStatusDTORoleSecondary ConsistencyGroupReplicationStatusDTORole = "secondary" + ConsistencyGroupReplicationStatusDTORoleSource ConsistencyGroupReplicationStatusDTORole = "source" +) + +// Valid indicates whether the value is a known member of the ConsistencyGroupReplicationStatusDTORole enum. +func (e ConsistencyGroupReplicationStatusDTORole) Valid() bool { + switch e { + case ConsistencyGroupReplicationStatusDTORoleFailedOver: + return true + case ConsistencyGroupReplicationStatusDTORoleNone: + return true + case ConsistencyGroupReplicationStatusDTORoleSecondary: + return true + case ConsistencyGroupReplicationStatusDTORoleSource: + return true + default: + return false + } +} + +// Defines values for ConsistencyGroupReplicationStatusDTOState. +const ( + ConsistencyGroupReplicationStatusDTOStateDegraded ConsistencyGroupReplicationStatusDTOState = "degraded" + ConsistencyGroupReplicationStatusDTOStateError ConsistencyGroupReplicationStatusDTOState = "error" + ConsistencyGroupReplicationStatusDTOStateInSync ConsistencyGroupReplicationStatusDTOState = "in_sync" + ConsistencyGroupReplicationStatusDTOStateLagging ConsistencyGroupReplicationStatusDTOState = "lagging" + ConsistencyGroupReplicationStatusDTOStateNotReplicating ConsistencyGroupReplicationStatusDTOState = "not_replicating" + ConsistencyGroupReplicationStatusDTOStateReplicating ConsistencyGroupReplicationStatusDTOState = "replicating" +) + +// Valid indicates whether the value is a known member of the ConsistencyGroupReplicationStatusDTOState enum. +func (e ConsistencyGroupReplicationStatusDTOState) Valid() bool { + switch e { + case ConsistencyGroupReplicationStatusDTOStateDegraded: + return true + case ConsistencyGroupReplicationStatusDTOStateError: + return true + case ConsistencyGroupReplicationStatusDTOStateInSync: + return true + case ConsistencyGroupReplicationStatusDTOStateLagging: + return true + case ConsistencyGroupReplicationStatusDTOStateNotReplicating: + return true + case ConsistencyGroupReplicationStatusDTOStateReplicating: + return true + default: + return false + } +} + // Defines values for FailoverResultDTOStatus. const ( FailoverResultDTOStatusFailed FailoverResultDTOStatus = "failed" @@ -998,6 +1052,40 @@ type ConsistencyGroupMemberJoinDTO struct { LvolId string `json:"lvol_id"` } +// ConsistencyGroupReplicationIntentDTO Request body to enable or disable group replication (design §14.4). +// +// A UUID attaches the whole consistency group to that group replication policy; +// an explicit “null“ detaches it (the group and its members stay grouped by +// label, only replication stops). The field is required, so omitting it is a +// 422 rather than an ambiguous no-op. +type ConsistencyGroupReplicationIntentDTO struct { + ReplicationPolicyId *openapi_types.UUID `json:"replication_policy_id"` +} + +// ConsistencyGroupReplicationStatusDTO The replication status of a consistency group as one unit. +// +// A group's recovery point is its OLDEST member's, its lag and health its +// WORST member's, and its backlog the sum, because a group is only as +// protected as its slowest, sickest member. Never a 404, matching the +// per-volume “ReplicationStatusDTO“ (design-csi-addons-replication.md +// §14.4/§14.6). +type ConsistencyGroupReplicationStatusDTO struct { + LagSeconds *int `json:"lag_seconds,omitempty"` + LastReplicatedAt *time.Time `json:"last_replicated_at,omitempty"` + MemberCount *int `json:"member_count,omitempty"` + OutstandingBytes *int `json:"outstanding_bytes,omitempty"` + OutstandingCount *int `json:"outstanding_count,omitempty"` + Resyncing *bool `json:"resyncing,omitempty"` + Role ConsistencyGroupReplicationStatusDTORole `json:"role"` + State ConsistencyGroupReplicationStatusDTOState `json:"state"` +} + +// ConsistencyGroupReplicationStatusDTORole defines model for ConsistencyGroupReplicationStatusDTO.Role. +type ConsistencyGroupReplicationStatusDTORole string + +// ConsistencyGroupReplicationStatusDTOState defines model for ConsistencyGroupReplicationStatusDTO.State. +type ConsistencyGroupReplicationStatusDTOState string + // DeviceDTO defines model for DeviceDTO. type DeviceDTO struct { BdevType *string `json:"bdev_type,omitempty"` @@ -1065,6 +1153,11 @@ type FailoverResultDTO struct { // FailoverResultDTOStatus defines model for FailoverResultDTO.Status. type FailoverResultDTOStatus string +// GroupFailbackParams defines model for GroupFailbackParams. +type GroupFailbackParams struct { + SourceClusterId *openapi_types.UUID `json:"source_cluster_id,omitempty"` +} + // HTTPValidationError defines model for HTTPValidationError. type HTTPValidationError struct { Detail *[]ValidationError `json:"detail,omitempty"` @@ -1981,6 +2074,12 @@ type ClustersBackupsSourceSwitchApiV2ClustersClusterIdBackupsSourceSwitchPostJSO // ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostJSONRequestBody defines body for ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPost for application/json ContentType. type ClustersConsistencyGroupsMembersJoinApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersPostJSONRequestBody = ConsistencyGroupMemberJoinDTO +// ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutJSONRequestBody defines body for ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPut for application/json ContentType. +type ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutJSONRequestBody = ConsistencyGroupReplicationIntentDTO + +// ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostJSONRequestBody defines body for ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPost for application/json ContentType. +type ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostJSONRequestBody = GroupFailbackParams + // ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostJSONRequestBody defines body for ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPost for application/json ContentType. type ClustersReplicationPoliciesCreateApiV2ClustersClusterIdReplicationPoliciesPostJSONRequestBody = PolicyParams @@ -2656,6 +2755,84 @@ type ClientInterface interface { // Corresponds with DELETE /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/members/{lvol_id} (the `ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDelete` operationId). ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDelete(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, lvolId string, reqEditors ...RequestEditorFn) (*http.Response, error) + // ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutWithBody Clusters:Consistency-Groups:Replication:Configure + // + // Enable or disable group replication (design-csi-addons-replication.md + // §14.4): a policy id attaches the whole group to that group replication + // policy; ``null`` detaches it (the group and its members stay grouped by + // label). A refused attach (not a consistency-group policy, missing policy) is + // a 409. + // + // Takes any type of body and a specified content type. + // + // Corresponds with PUT /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication (the `ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPut` operationId). + ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutWithBody(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*http.Response, error) + + // ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPut Clusters:Consistency-Groups:Replication:Configure + // + // Enable or disable group replication (design-csi-addons-replication.md + // §14.4): a policy id attaches the whole group to that group replication + // policy; ``null`` detaches it (the group and its members stay grouped by + // label). A refused attach (not a consistency-group policy, missing policy) is + // a 409. + // + // Takes a body of the `application/json` content type. + // + // Corresponds with PUT /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication (the `ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPut` operationId). + ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPut(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, body ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutJSONRequestBody, reqEditors ...RequestEditorFn) (*http.Response, error) + + // ClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePost Clusters:Consistency-Groups:Replication:Demote + // + // Demote the whole group: fence every member and confirm each one's last + // write replicated (design-csi-addons-replication.md §14.4). Re-drivable, not + // queued: 204 once every member is demoted, 202 (with per-member detail) while + // any is still converging, 500 on a hard failure. + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/demote (the `ClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePost` operationId). + ClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePost(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) + + // ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostWithBody Clusters:Consistency-Groups:Replication:Failback + // + // Fail the whole group back: point every member's replication back at the + // source cluster (design-csi-addons-replication.md §14.4). The cutover itself is + // each member's own commit. + // + // Takes any type of body and a specified content type. + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/failback (the `ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPost` operationId). + ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostWithBody(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*http.Response, error) + + // ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPost Clusters:Consistency-Groups:Replication:Failback + // + // Fail the whole group back: point every member's replication back at the + // source cluster (design-csi-addons-replication.md §14.4). The cutover itself is + // each member's own commit. + // + // Takes a body of the `application/json` content type. + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/failback (the `ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPost` operationId). + ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPost(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, body ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostJSONRequestBody, reqEditors ...RequestEditorFn) (*http.Response, error) + + // ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPost Clusters:Consistency-Groups:Replication:Failover + // + // Fail the whole group over as ONE unit through its replication policy + // (design-csi-addons-replication.md §14.4): every member is pinned to the same + // group generation, all-or-nothing. Refuses (412) a group not attached to a + // policy. + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/failover (the `ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPost` operationId). + ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPost(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) + + // ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGet Clusters:Consistency-Groups:Replication:Status + // + // The group's replication status as one unit: oldest recovery point, worst + // member lag and health, summed backlog (design-csi-addons-replication.md + // §14.4/§14.6). Never 404s -- a group with no replicating member reports + // ``state: not_replicating``. + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/status (the `ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGet` operationId). + ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGet(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) + // ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGet Clusters:Consistency-Groups:Snapshots:List // // Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/snapshots (the `ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGet` operationId). @@ -3264,16 +3441,20 @@ type ClientInterface interface { // faithfully replicated. // // ``planned=True`` gates on a completed demote (P0-3) so a planned swap - // loses nothing: 412 when no demote was ever requested for this volume (the - // caller's premise that the source is reachable to demote was wrong, and a - // 412 is what lets the csi-addons controller's own force-escalation take - // over), 409 while demote is still converging (retryable -- 409 must never - // become a code the controller reads as permission to force, since that - // controller escalates on ANY FAILED_PRECONDITION from a force=false - // promote with no wait-and-retry grace period of its own). Unplanned - // failover (the default) ignores demote state entirely, unchanged from - // today: its whole premise is that the source may never have been - // reachable to demote. + // loses nothing: 409 while demote is still converging (retryable -- 409 + // must never become a code the controller reads as permission to force, + // since that controller escalates on ANY FAILED_PRECONDITION from a + // force=false promote with no wait-and-retry grace period of its own). + // When no demote was ever requested, the source's own health decides: a + // genuinely healthy, still-serving source means there is nothing to fail + // over -- this is the vendored csi-addons controller's OWN first-ever + // reconcile of a `VolumeReplication` that already lives here, not a + // disaster, and this call succeeds as the no-op it is. A source that is + // NOT healthy gets 412, the caller's premise that it was reachable to + // demote was wrong, and 412 is what lets the controller's own + // force-escalation take over. Unplanned failover (the default) ignores + // demote state entirely, unchanged from today: its whole premise is that + // the source may never have been reachable to demote. // // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/failover (the `ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPost` operationId). ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPost(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, params *ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPostParams, reqEditors ...RequestEditorFn) (*http.Response, error) @@ -4147,6 +4328,154 @@ func (c *Client) ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdCon return c.Client.Do(req) } +// ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutWithBody Clusters:Consistency-Groups:Replication:Configure +// +// Enable or disable group replication (design-csi-addons-replication.md +// §14.4): a policy id attaches the whole group to that group replication +// policy; “null“ detaches it (the group and its members stay grouped by +// label). A refused attach (not a consistency-group policy, missing policy) is +// a 409. +// +// Takes any type of body and a specified content type. +// +// Corresponds with PUT /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication (the `ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPut` operationId). +func (c *Client) ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutWithBody(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutRequestWithBody(c.Server, clusterId, groupId, contentType, body) + if err != nil { + return nil, err + } + req = req.WithContext(ctx) + if err := c.applyEditors(ctx, req, reqEditors); err != nil { + return nil, err + } + return c.Client.Do(req) +} + +// ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPut Clusters:Consistency-Groups:Replication:Configure +// +// Enable or disable group replication (design-csi-addons-replication.md +// §14.4): a policy id attaches the whole group to that group replication +// policy; “null“ detaches it (the group and its members stay grouped by +// label). A refused attach (not a consistency-group policy, missing policy) is +// a 409. +// +// Takes a body of the `application/json` content type. +// +// Corresponds with PUT /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication (the `ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPut` operationId). +func (c *Client) ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPut(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, body ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutJSONRequestBody, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutRequest(c.Server, clusterId, groupId, body) + if err != nil { + return nil, err + } + req = req.WithContext(ctx) + if err := c.applyEditors(ctx, req, reqEditors); err != nil { + return nil, err + } + return c.Client.Do(req) +} + +// ClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePost Clusters:Consistency-Groups:Replication:Demote +// +// Demote the whole group: fence every member and confirm each one's last +// write replicated (design-csi-addons-replication.md §14.4). Re-drivable, not +// queued: 204 once every member is demoted, 202 (with per-member detail) while +// any is still converging, 500 on a hard failure. +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/demote (the `ClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePost` operationId). +func (c *Client) ClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePost(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePostRequest(c.Server, clusterId, groupId) + if err != nil { + return nil, err + } + req = req.WithContext(ctx) + if err := c.applyEditors(ctx, req, reqEditors); err != nil { + return nil, err + } + return c.Client.Do(req) +} + +// ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostWithBody Clusters:Consistency-Groups:Replication:Failback +// +// Fail the whole group back: point every member's replication back at the +// source cluster (design-csi-addons-replication.md §14.4). The cutover itself is +// each member's own commit. +// +// Takes any type of body and a specified content type. +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/failback (the `ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPost` operationId). +func (c *Client) ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostWithBody(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostRequestWithBody(c.Server, clusterId, groupId, contentType, body) + if err != nil { + return nil, err + } + req = req.WithContext(ctx) + if err := c.applyEditors(ctx, req, reqEditors); err != nil { + return nil, err + } + return c.Client.Do(req) +} + +// ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPost Clusters:Consistency-Groups:Replication:Failback +// +// Fail the whole group back: point every member's replication back at the +// source cluster (design-csi-addons-replication.md §14.4). The cutover itself is +// each member's own commit. +// +// Takes a body of the `application/json` content type. +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/failback (the `ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPost` operationId). +func (c *Client) ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPost(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, body ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostJSONRequestBody, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostRequest(c.Server, clusterId, groupId, body) + if err != nil { + return nil, err + } + req = req.WithContext(ctx) + if err := c.applyEditors(ctx, req, reqEditors); err != nil { + return nil, err + } + return c.Client.Do(req) +} + +// ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPost Clusters:Consistency-Groups:Replication:Failover +// +// Fail the whole group over as ONE unit through its replication policy +// (design-csi-addons-replication.md §14.4): every member is pinned to the same +// group generation, all-or-nothing. Refuses (412) a group not attached to a +// policy. +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/failover (the `ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPost` operationId). +func (c *Client) ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPost(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPostRequest(c.Server, clusterId, groupId) + if err != nil { + return nil, err + } + req = req.WithContext(ctx) + if err := c.applyEditors(ctx, req, reqEditors); err != nil { + return nil, err + } + return c.Client.Do(req) +} + +// ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGet Clusters:Consistency-Groups:Replication:Status +// +// The group's replication status as one unit: oldest recovery point, worst +// member lag and health, summed backlog (design-csi-addons-replication.md +// §14.4/§14.6). Never 404s -- a group with no replicating member reports +// “state: not_replicating“. +// +// Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/status (the `ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGet` operationId). +func (c *Client) ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGet(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetRequest(c.Server, clusterId, groupId) + if err != nil { + return nil, err + } + req = req.WithContext(ctx) + if err := c.applyEditors(ctx, req, reqEditors); err != nil { + return nil, err + } + return c.Client.Do(req) +} + // ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGet Clusters:Consistency-Groups:Snapshots:List // // Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/snapshots (the `ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGet` operationId). @@ -5675,16 +6004,20 @@ func (c *Client) ClustersStoragePoolsVolumesReplicationFailbackApiV2ClustersClus // faithfully replicated. // // “planned=True“ gates on a completed demote (P0-3) so a planned swap -// loses nothing: 412 when no demote was ever requested for this volume (the -// caller's premise that the source is reachable to demote was wrong, and a -// 412 is what lets the csi-addons controller's own force-escalation take -// over), 409 while demote is still converging (retryable -- 409 must never -// become a code the controller reads as permission to force, since that -// controller escalates on ANY FAILED_PRECONDITION from a force=false -// promote with no wait-and-retry grace period of its own). Unplanned -// failover (the default) ignores demote state entirely, unchanged from -// today: its whole premise is that the source may never have been -// reachable to demote. +// loses nothing: 409 while demote is still converging (retryable -- 409 +// must never become a code the controller reads as permission to force, +// since that controller escalates on ANY FAILED_PRECONDITION from a +// force=false promote with no wait-and-retry grace period of its own). +// When no demote was ever requested, the source's own health decides: a +// genuinely healthy, still-serving source means there is nothing to fail +// over -- this is the vendored csi-addons controller's OWN first-ever +// reconcile of a `VolumeReplication` that already lives here, not a +// disaster, and this call succeeds as the no-op it is. A source that is +// NOT healthy gets 412, the caller's premise that it was reachable to +// demote was wrong, and 412 is what lets the controller's own +// force-escalation take over. Unplanned failover (the default) ignores +// demote state entirely, unchanged from today: its whole premise is that +// the source may never have been reachable to demote. // // Corresponds with POST /api/v2/clusters/{cluster_id}/storage-pools/{pool_id}/volumes/{volume_id}/replication/failover (the `ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPost` operationId). func (c *Client) ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPost(ctx context.Context, clusterId openapi_types.UUID, poolId openapi_types.UUID, volumeId openapi_types.UUID, params *ClustersStoragePoolsVolumesReplicationFailoverApiV2ClustersClusterIdStoragePoolsPoolIdVolumesVolumeIdReplicationFailoverPostParams, reqEditors ...RequestEditorFn) (*http.Response, error) { @@ -7588,8 +7921,19 @@ func NewClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyG return req, nil } -// NewClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetRequest constructs an http.Request for the ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGet method -func NewClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetRequest(server string, clusterId openapi_types.UUID, groupId openapi_types.UUID) (*http.Request, error) { +// NewClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutRequest calls the generic ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPut builder with application/json body +func NewClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutRequest(server string, clusterId openapi_types.UUID, groupId openapi_types.UUID, body ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutJSONRequestBody) (*http.Request, error) { + var bodyReader io.Reader + buf, err := json.Marshal(body) + if err != nil { + return nil, err + } + bodyReader = bytes.NewReader(buf) + return NewClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutRequestWithBody(server, clusterId, groupId, "application/json", bodyReader) +} + +// NewClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutRequestWithBody constructs an http.Request for the ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPut method, with any body, and a specified content type +func NewClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutRequestWithBody(server string, clusterId openapi_types.UUID, groupId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { var err error var pathParam0 string @@ -7611,7 +7955,7 @@ func NewClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyG return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/consistency-groups/%s/snapshots", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/consistency-groups/%s/replication", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -7621,16 +7965,18 @@ func NewClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyG return nil, err } - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodPut, queryURL.String(), body) if err != nil { return nil, err } + req.Header.Add("Content-Type", contentType) + return req, nil } -// NewClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPostRequest constructs an http.Request for the ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPost method -func NewClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPostRequest(server string, clusterId openapi_types.UUID, groupId openapi_types.UUID) (*http.Request, error) { +// NewClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePostRequest constructs an http.Request for the ClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePost method +func NewClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePostRequest(server string, clusterId openapi_types.UUID, groupId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -7652,7 +7998,7 @@ func NewClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyG return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/consistency-groups/%s/snapshots", pathParam0, pathParam1) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/consistency-groups/%s/replication/demote", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -7670,8 +8016,19 @@ func NewClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyG return req, nil } -// NewClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDeleteRequest constructs an http.Request for the ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDelete method -func NewClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDeleteRequest(server string, clusterId openapi_types.UUID, groupId openapi_types.UUID, seq int) (*http.Request, error) { +// NewClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostRequest calls the generic ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPost builder with application/json body +func NewClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostRequest(server string, clusterId openapi_types.UUID, groupId openapi_types.UUID, body ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostJSONRequestBody) (*http.Request, error) { + var bodyReader io.Reader + buf, err := json.Marshal(body) + if err != nil { + return nil, err + } + bodyReader = bytes.NewReader(buf) + return NewClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostRequestWithBody(server, clusterId, groupId, "application/json", bodyReader) +} + +// NewClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostRequestWithBody constructs an http.Request for the ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPost method, with any body, and a specified content type +func NewClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostRequestWithBody(server string, clusterId openapi_types.UUID, groupId openapi_types.UUID, contentType string, body io.Reader) (*http.Request, error) { var err error var pathParam0 string @@ -7688,19 +8045,12 @@ func NewClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistenc return nil, err } - var pathParam2 string - - pathParam2, err = runtime.StyleParamWithOptions("simple", false, "seq", seq, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "integer", Format: ""}) - if err != nil { - return nil, err - } - serverURL, err := url.Parse(server) if err != nil { return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/consistency-groups/%s/snapshots/%s", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/consistency-groups/%s/replication/failback", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -7710,16 +8060,18 @@ func NewClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistenc return nil, err } - req, err := http.NewRequest(http.MethodDelete, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodPost, queryURL.String(), body) if err != nil { return nil, err } + req.Header.Add("Content-Type", contentType) + return req, nil } -// NewClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGetRequest constructs an http.Request for the ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGet method -func NewClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGetRequest(server string, clusterId openapi_types.UUID, groupId openapi_types.UUID, seq int) (*http.Request, error) { +// NewClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPostRequest constructs an http.Request for the ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPost method +func NewClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPostRequest(server string, clusterId openapi_types.UUID, groupId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -7736,19 +8088,12 @@ func NewClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistenc return nil, err } - var pathParam2 string - - pathParam2, err = runtime.StyleParamWithOptions("simple", false, "seq", seq, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "integer", Format: ""}) - if err != nil { - return nil, err - } - serverURL, err := url.Parse(server) if err != nil { return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/consistency-groups/%s/snapshots/%s", pathParam0, pathParam1, pathParam2) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/consistency-groups/%s/replication/failover", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -7758,7 +8103,7 @@ func NewClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistenc return nil, err } - req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) if err != nil { return nil, err } @@ -7766,8 +8111,8 @@ func NewClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistenc return req, nil } -// NewClustersExpandApiV2ClustersClusterIdExpandPostRequest constructs an http.Request for the ClustersExpandApiV2ClustersClusterIdExpandPost method -func NewClustersExpandApiV2ClustersClusterIdExpandPostRequest(server string, clusterId openapi_types.UUID) (*http.Request, error) { +// NewClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetRequest constructs an http.Request for the ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGet method +func NewClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetRequest(server string, clusterId openapi_types.UUID, groupId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -7777,12 +8122,19 @@ func NewClustersExpandApiV2ClustersClusterIdExpandPostRequest(server string, clu return nil, err } + var pathParam1 string + + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "group_id", groupId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + serverURL, err := url.Parse(server) if err != nil { return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/expand", pathParam0) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/consistency-groups/%s/replication/status", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -7792,7 +8144,7 @@ func NewClustersExpandApiV2ClustersClusterIdExpandPostRequest(server string, clu return nil, err } - req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) if err != nil { return nil, err } @@ -7800,8 +8152,8 @@ func NewClustersExpandApiV2ClustersClusterIdExpandPostRequest(server string, clu return req, nil } -// NewClustersIostatsApiV2ClustersClusterIdIostatsGetRequest constructs an http.Request for the ClustersIostatsApiV2ClustersClusterIdIostatsGet method -func NewClustersIostatsApiV2ClustersClusterIdIostatsGetRequest(server string, clusterId openapi_types.UUID, params *ClustersIostatsApiV2ClustersClusterIdIostatsGetParams) (*http.Request, error) { +// NewClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetRequest constructs an http.Request for the ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGet method +func NewClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetRequest(server string, clusterId openapi_types.UUID, groupId openapi_types.UUID) (*http.Request, error) { var err error var pathParam0 string @@ -7811,12 +8163,19 @@ func NewClustersIostatsApiV2ClustersClusterIdIostatsGetRequest(server string, cl return nil, err } + var pathParam1 string + + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "group_id", groupId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + serverURL, err := url.Parse(server) if err != nil { return nil, err } - operationPath := fmt.Sprintf("/api/v2/clusters/%s/iostats", pathParam0) + operationPath := fmt.Sprintf("/api/v2/clusters/%s/consistency-groups/%s/snapshots", pathParam0, pathParam1) if operationPath[0] == '/' { operationPath = "." + operationPath } @@ -7826,19 +8185,224 @@ func NewClustersIostatsApiV2ClustersClusterIdIostatsGetRequest(server string, cl return nil, err } - if params != nil { - // queryValues collects non-styled parameters (passthrough, JSON) - // that are safe to round-trip through url.Values.Encode(). - queryValues := queryURL.Query() - // rawQueryFragments collects pre-encoded query fragments from - // styled parameters, preserving literal commas as delimiters - // per the OpenAPI spec (e.g. "color=blue,black,brown"). - var rawQueryFragments []string + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + if err != nil { + return nil, err + } - if params.History != nil { + return req, nil +} - if queryFrag, err := runtime.StyleParamWithOptions("form", true, "history", *params.History, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "", Format: ""}); err != nil { - return nil, err +// NewClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPostRequest constructs an http.Request for the ClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPost method +func NewClustersConsistencyGroupsSnapshotsTakeApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsPostRequest(server string, clusterId openapi_types.UUID, groupId openapi_types.UUID) (*http.Request, error) { + var err error + + var pathParam0 string + + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + var pathParam1 string + + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "group_id", groupId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + serverURL, err := url.Parse(server) + if err != nil { + return nil, err + } + + operationPath := fmt.Sprintf("/api/v2/clusters/%s/consistency-groups/%s/snapshots", pathParam0, pathParam1) + if operationPath[0] == '/' { + operationPath = "." + operationPath + } + + queryURL, err := serverURL.Parse(operationPath) + if err != nil { + return nil, err + } + + req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) + if err != nil { + return nil, err + } + + return req, nil +} + +// NewClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDeleteRequest constructs an http.Request for the ClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDelete method +func NewClustersConsistencyGroupsSnapshotsDeleteApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqDeleteRequest(server string, clusterId openapi_types.UUID, groupId openapi_types.UUID, seq int) (*http.Request, error) { + var err error + + var pathParam0 string + + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + var pathParam1 string + + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "group_id", groupId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + var pathParam2 string + + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "seq", seq, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "integer", Format: ""}) + if err != nil { + return nil, err + } + + serverURL, err := url.Parse(server) + if err != nil { + return nil, err + } + + operationPath := fmt.Sprintf("/api/v2/clusters/%s/consistency-groups/%s/snapshots/%s", pathParam0, pathParam1, pathParam2) + if operationPath[0] == '/' { + operationPath = "." + operationPath + } + + queryURL, err := serverURL.Parse(operationPath) + if err != nil { + return nil, err + } + + req, err := http.NewRequest(http.MethodDelete, queryURL.String(), nil) + if err != nil { + return nil, err + } + + return req, nil +} + +// NewClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGetRequest constructs an http.Request for the ClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGet method +func NewClustersConsistencyGroupsSnapshotsDetailApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsSeqGetRequest(server string, clusterId openapi_types.UUID, groupId openapi_types.UUID, seq int) (*http.Request, error) { + var err error + + var pathParam0 string + + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + var pathParam1 string + + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "group_id", groupId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + var pathParam2 string + + pathParam2, err = runtime.StyleParamWithOptions("simple", false, "seq", seq, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "integer", Format: ""}) + if err != nil { + return nil, err + } + + serverURL, err := url.Parse(server) + if err != nil { + return nil, err + } + + operationPath := fmt.Sprintf("/api/v2/clusters/%s/consistency-groups/%s/snapshots/%s", pathParam0, pathParam1, pathParam2) + if operationPath[0] == '/' { + operationPath = "." + operationPath + } + + queryURL, err := serverURL.Parse(operationPath) + if err != nil { + return nil, err + } + + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + if err != nil { + return nil, err + } + + return req, nil +} + +// NewClustersExpandApiV2ClustersClusterIdExpandPostRequest constructs an http.Request for the ClustersExpandApiV2ClustersClusterIdExpandPost method +func NewClustersExpandApiV2ClustersClusterIdExpandPostRequest(server string, clusterId openapi_types.UUID) (*http.Request, error) { + var err error + + var pathParam0 string + + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + serverURL, err := url.Parse(server) + if err != nil { + return nil, err + } + + operationPath := fmt.Sprintf("/api/v2/clusters/%s/expand", pathParam0) + if operationPath[0] == '/' { + operationPath = "." + operationPath + } + + queryURL, err := serverURL.Parse(operationPath) + if err != nil { + return nil, err + } + + req, err := http.NewRequest(http.MethodPost, queryURL.String(), nil) + if err != nil { + return nil, err + } + + return req, nil +} + +// NewClustersIostatsApiV2ClustersClusterIdIostatsGetRequest constructs an http.Request for the ClustersIostatsApiV2ClustersClusterIdIostatsGet method +func NewClustersIostatsApiV2ClustersClusterIdIostatsGetRequest(server string, clusterId openapi_types.UUID, params *ClustersIostatsApiV2ClustersClusterIdIostatsGetParams) (*http.Request, error) { + var err error + + var pathParam0 string + + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + serverURL, err := url.Parse(server) + if err != nil { + return nil, err + } + + operationPath := fmt.Sprintf("/api/v2/clusters/%s/iostats", pathParam0) + if operationPath[0] == '/' { + operationPath = "." + operationPath + } + + queryURL, err := serverURL.Parse(operationPath) + if err != nil { + return nil, err + } + + if params != nil { + // queryValues collects non-styled parameters (passthrough, JSON) + // that are safe to round-trip through url.Values.Encode(). + queryValues := queryURL.Query() + // rawQueryFragments collects pre-encoded query fragments from + // styled parameters, preserving literal commas as delimiters + // per the OpenAPI spec (e.g. "color=blue,black,brown"). + var rawQueryFragments []string + + if params.History != nil { + + if queryFrag, err := runtime.StyleParamWithOptions("form", true, "history", *params.History, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationQuery, Type: "", Format: ""}); err != nil { + return nil, err } else { for _, qp := range strings.Split(queryFrag, "&") { rawQueryFragments = append(rawQueryFragments, qp) @@ -13405,6 +13969,90 @@ type ClientWithResponsesInterface interface { // Corresponds with DELETE /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/members/{lvol_id} (the `ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDelete` operationId). ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDeleteWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, lvolId string, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDeleteResponse, error) + // ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutWithBodyWithResponse Clusters:Consistency-Groups:Replication:Configure + // + // Enable or disable group replication (design-csi-addons-replication.md + // §14.4): a policy id attaches the whole group to that group replication + // policy; ``null`` detaches it (the group and its members stay grouped by + // label). A refused attach (not a consistency-group policy, missing policy) is + // a 409. + // + // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with PUT /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication (the `ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPut` operationId). + ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutResponse, error) + + // ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutWithResponse Clusters:Consistency-Groups:Replication:Configure + // + // Enable or disable group replication (design-csi-addons-replication.md + // §14.4): a policy id attaches the whole group to that group replication + // policy; ``null`` detaches it (the group and its members stay grouped by + // label). A refused attach (not a consistency-group policy, missing policy) is + // a 409. + // + // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with PUT /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication (the `ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPut` operationId). + ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, body ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutResponse, error) + + // ClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePostWithResponse Clusters:Consistency-Groups:Replication:Demote + // + // Demote the whole group: fence every member and confirm each one's last + // write replicated (design-csi-addons-replication.md §14.4). Re-drivable, not + // queued: 204 once every member is demoted, 202 (with per-member detail) while + // any is still converging, 500 on a hard failure. + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/demote (the `ClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePost` operationId). + ClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePostResponse, error) + + // ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostWithBodyWithResponse Clusters:Consistency-Groups:Replication:Failback + // + // Fail the whole group back: point every member's replication back at the + // source cluster (design-csi-addons-replication.md §14.4). The cutover itself is + // each member's own commit. + // + // Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/failback (the `ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPost` operationId). + ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostResponse, error) + + // ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostWithResponse Clusters:Consistency-Groups:Replication:Failback + // + // Fail the whole group back: point every member's replication back at the + // source cluster (design-csi-addons-replication.md §14.4). The cutover itself is + // each member's own commit. + // + // Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/failback (the `ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPost` operationId). + ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, body ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostResponse, error) + + // ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPostWithResponse Clusters:Consistency-Groups:Replication:Failover + // + // Fail the whole group over as ONE unit through its replication policy + // (design-csi-addons-replication.md §14.4): every member is pinned to the same + // group generation, all-or-nothing. Refuses (412) a group not attached to a + // policy. + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/failover (the `ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPost` operationId). + ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPostResponse, error) + + // ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetWithResponse Clusters:Consistency-Groups:Replication:Status + // + // The group's replication status as one unit: oldest recovery point, worst + // member lag and health, summed backlog (design-csi-addons-replication.md + // §14.4/§14.6). Never 404s -- a group with no replicating member reports + // ``state: not_replicating``. + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/status (the `ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGet` operationId). + ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetResponse, error) + // ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetWithResponse Clusters:Consistency-Groups:Snapshots:List // // Returns a wrapper object for the known response body format(s). @@ -14137,16 +14785,20 @@ type ClientWithResponsesInterface interface { // faithfully replicated. // // ``planned=True`` gates on a completed demote (P0-3) so a planned swap - // loses nothing: 412 when no demote was ever requested for this volume (the - // caller's premise that the source is reachable to demote was wrong, and a - // 412 is what lets the csi-addons controller's own force-escalation take - // over), 409 while demote is still converging (retryable -- 409 must never - // become a code the controller reads as permission to force, since that - // controller escalates on ANY FAILED_PRECONDITION from a force=false - // promote with no wait-and-retry grace period of its own). Unplanned - // failover (the default) ignores demote state entirely, unchanged from - // today: its whole premise is that the source may never have been - // reachable to demote. + // loses nothing: 409 while demote is still converging (retryable -- 409 + // must never become a code the controller reads as permission to force, + // since that controller escalates on ANY FAILED_PRECONDITION from a + // force=false promote with no wait-and-retry grace period of its own). + // When no demote was ever requested, the source's own health decides: a + // genuinely healthy, still-serving source means there is nothing to fail + // over -- this is the vendored csi-addons controller's OWN first-ever + // reconcile of a `VolumeReplication` that already lives here, not a + // disaster, and this call succeeds as the no-op it is. A source that is + // NOT healthy gets 412, the caller's premise that it was reachable to + // demote was wrong, and 412 is what lets the controller's own + // force-escalation take over. Unplanned failover (the default) ignores + // demote state entirely, unchanged from today: its whole premise is that + // the source may never have been reachable to demote. // // Returns a wrapper object for the known response body format(s). // @@ -15688,32 +16340,251 @@ func (r ClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyG return "" } -type ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetResponse struct { +type ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutResponse struct { Body []byte HTTPResponse *http.Response - // JSON200 the response for an HTTP 200 `application/json` response - JSON200 *[]ConsistencyGroupGenerationDTO // JSON422 the response for an HTTP 422 `application/json` response JSON422 *HTTPValidationError } -// GetJSON200 returns the response for an HTTP 200 `application/json` response -func (r ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetResponse) GetJSON200() *[]ConsistencyGroupGenerationDTO { - return r.JSON200 -} - // GetJSON422 returns the response for an HTTP 422 `application/json` response -func (r ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetResponse) GetJSON422() *HTTPValidationError { +func (r ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutResponse) GetJSON422() *HTTPValidationError { return r.JSON422 } // GetBody returns the raw response body bytes -func (r ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetResponse) GetBody() []byte { +func (r ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutResponse) GetBody() []byte { return r.Body } // Status returns HTTPResponse.Status -func (r ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetResponse) Status() string { +func (r ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutResponse) Status() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Status + } + return http.StatusText(0) +} + +// StatusCode returns HTTPResponse.StatusCode +func (r ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutResponse) StatusCode() int { + if r.HTTPResponse != nil { + return r.HTTPResponse.StatusCode + } + return 0 +} + +// ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers +func (r ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutResponse) ContentType() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Header.Get("Content-Type") + } + return "" +} + +type ClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePostResponse struct { + Body []byte + HTTPResponse *http.Response + // JSON422 the response for an HTTP 422 `application/json` response + JSON422 *HTTPValidationError +} + +// GetJSON422 returns the response for an HTTP 422 `application/json` response +func (r ClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePostResponse) GetJSON422() *HTTPValidationError { + return r.JSON422 +} + +// GetBody returns the raw response body bytes +func (r ClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePostResponse) GetBody() []byte { + return r.Body +} + +// Status returns HTTPResponse.Status +func (r ClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePostResponse) Status() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Status + } + return http.StatusText(0) +} + +// StatusCode returns HTTPResponse.StatusCode +func (r ClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePostResponse) StatusCode() int { + if r.HTTPResponse != nil { + return r.HTTPResponse.StatusCode + } + return 0 +} + +// ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers +func (r ClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePostResponse) ContentType() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Header.Get("Content-Type") + } + return "" +} + +type ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostResponse struct { + Body []byte + HTTPResponse *http.Response + // JSON422 the response for an HTTP 422 `application/json` response + JSON422 *HTTPValidationError +} + +// GetJSON422 returns the response for an HTTP 422 `application/json` response +func (r ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostResponse) GetJSON422() *HTTPValidationError { + return r.JSON422 +} + +// GetBody returns the raw response body bytes +func (r ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostResponse) GetBody() []byte { + return r.Body +} + +// Status returns HTTPResponse.Status +func (r ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostResponse) Status() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Status + } + return http.StatusText(0) +} + +// StatusCode returns HTTPResponse.StatusCode +func (r ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostResponse) StatusCode() int { + if r.HTTPResponse != nil { + return r.HTTPResponse.StatusCode + } + return 0 +} + +// ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers +func (r ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostResponse) ContentType() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Header.Get("Content-Type") + } + return "" +} + +type ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPostResponse struct { + Body []byte + HTTPResponse *http.Response + // JSON200 the response for an HTTP 200 `application/json` response + JSON200 *map[string]interface{} + // JSON422 the response for an HTTP 422 `application/json` response + JSON422 *HTTPValidationError +} + +// GetJSON200 returns the response for an HTTP 200 `application/json` response +func (r ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPostResponse) GetJSON200() *map[string]interface{} { + return r.JSON200 +} + +// GetJSON422 returns the response for an HTTP 422 `application/json` response +func (r ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPostResponse) GetJSON422() *HTTPValidationError { + return r.JSON422 +} + +// GetBody returns the raw response body bytes +func (r ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPostResponse) GetBody() []byte { + return r.Body +} + +// Status returns HTTPResponse.Status +func (r ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPostResponse) Status() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Status + } + return http.StatusText(0) +} + +// StatusCode returns HTTPResponse.StatusCode +func (r ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPostResponse) StatusCode() int { + if r.HTTPResponse != nil { + return r.HTTPResponse.StatusCode + } + return 0 +} + +// ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers +func (r ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPostResponse) ContentType() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Header.Get("Content-Type") + } + return "" +} + +type ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetResponse struct { + Body []byte + HTTPResponse *http.Response + // JSON200 the response for an HTTP 200 `application/json` response + JSON200 *ConsistencyGroupReplicationStatusDTO + // JSON422 the response for an HTTP 422 `application/json` response + JSON422 *HTTPValidationError +} + +// GetJSON200 returns the response for an HTTP 200 `application/json` response +func (r ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetResponse) GetJSON200() *ConsistencyGroupReplicationStatusDTO { + return r.JSON200 +} + +// GetJSON422 returns the response for an HTTP 422 `application/json` response +func (r ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetResponse) GetJSON422() *HTTPValidationError { + return r.JSON422 +} + +// GetBody returns the raw response body bytes +func (r ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetResponse) GetBody() []byte { + return r.Body +} + +// Status returns HTTPResponse.Status +func (r ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetResponse) Status() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Status + } + return http.StatusText(0) +} + +// StatusCode returns HTTPResponse.StatusCode +func (r ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetResponse) StatusCode() int { + if r.HTTPResponse != nil { + return r.HTTPResponse.StatusCode + } + return 0 +} + +// ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers +func (r ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetResponse) ContentType() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Header.Get("Content-Type") + } + return "" +} + +type ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetResponse struct { + Body []byte + HTTPResponse *http.Response + // JSON200 the response for an HTTP 200 `application/json` response + JSON200 *[]ConsistencyGroupGenerationDTO + // JSON422 the response for an HTTP 422 `application/json` response + JSON422 *HTTPValidationError +} + +// GetJSON200 returns the response for an HTTP 200 `application/json` response +func (r ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetResponse) GetJSON200() *[]ConsistencyGroupGenerationDTO { + return r.JSON200 +} + +// GetJSON422 returns the response for an HTTP 422 `application/json` response +func (r ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetResponse) GetJSON422() *HTTPValidationError { + return r.JSON422 +} + +// GetBody returns the raw response body bytes +func (r ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetResponse) GetBody() []byte { + return r.Body +} + +// Status returns HTTPResponse.Status +func (r ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetResponse) Status() string { if r.HTTPResponse != nil { return r.HTTPResponse.Status } @@ -20565,6 +21436,132 @@ func (c *ClientWithResponses) ClustersConsistencyGroupsMembersDetachApiV2Cluster return ParseClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistencyGroupsGroupIdMembersLvolIdDeleteResponse(rsp) } +// ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutWithBodyWithResponse Clusters:Consistency-Groups:Replication:Configure +// +// Enable or disable group replication (design-csi-addons-replication.md +// §14.4): a policy id attaches the whole group to that group replication +// policy; “null“ detaches it (the group and its members stay grouped by +// label). A refused attach (not a consistency-group policy, missing policy) is +// a 409. +// +// Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). +// +// Corresponds with PUT /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication (the `ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPut` operationId). +func (c *ClientWithResponses) ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutResponse, error) { + rsp, err := c.ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutWithBody(ctx, clusterId, groupId, contentType, body, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutResponse(rsp) +} + +// ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutWithResponse Clusters:Consistency-Groups:Replication:Configure +// +// Enable or disable group replication (design-csi-addons-replication.md +// §14.4): a policy id attaches the whole group to that group replication +// policy; “null“ detaches it (the group and its members stay grouped by +// label). A refused attach (not a consistency-group policy, missing policy) is +// a 409. +// +// Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). +// +// Corresponds with PUT /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication (the `ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPut` operationId). +func (c *ClientWithResponses) ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, body ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutResponse, error) { + rsp, err := c.ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPut(ctx, clusterId, groupId, body, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutResponse(rsp) +} + +// ClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePostWithResponse Clusters:Consistency-Groups:Replication:Demote +// +// Demote the whole group: fence every member and confirm each one's last +// write replicated (design-csi-addons-replication.md §14.4). Re-drivable, not +// queued: 204 once every member is demoted, 202 (with per-member detail) while +// any is still converging, 500 on a hard failure. +// +// Returns a wrapper object for the known response body format(s). +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/demote (the `ClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePost` operationId). +func (c *ClientWithResponses) ClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePostWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePostResponse, error) { + rsp, err := c.ClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePost(ctx, clusterId, groupId, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePostResponse(rsp) +} + +// ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostWithBodyWithResponse Clusters:Consistency-Groups:Replication:Failback +// +// Fail the whole group back: point every member's replication back at the +// source cluster (design-csi-addons-replication.md §14.4). The cutover itself is +// each member's own commit. +// +// Takes any type of body and a specified content type, and returns a wrapper object for the known response body format(s). +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/failback (the `ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPost` operationId). +func (c *ClientWithResponses) ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostWithBodyWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, contentType string, body io.Reader, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostResponse, error) { + rsp, err := c.ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostWithBody(ctx, clusterId, groupId, contentType, body, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostResponse(rsp) +} + +// ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostWithResponse Clusters:Consistency-Groups:Replication:Failback +// +// Fail the whole group back: point every member's replication back at the +// source cluster (design-csi-addons-replication.md §14.4). The cutover itself is +// each member's own commit. +// +// Takes a body of the `application/json` content type, and returns a wrapper object for the known response body format(s). +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/failback (the `ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPost` operationId). +func (c *ClientWithResponses) ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, body ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostJSONRequestBody, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostResponse, error) { + rsp, err := c.ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPost(ctx, clusterId, groupId, body, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostResponse(rsp) +} + +// ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPostWithResponse Clusters:Consistency-Groups:Replication:Failover +// +// Fail the whole group over as ONE unit through its replication policy +// (design-csi-addons-replication.md §14.4): every member is pinned to the same +// group generation, all-or-nothing. Refuses (412) a group not attached to a +// policy. +// +// Returns a wrapper object for the known response body format(s). +// +// Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/failover (the `ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPost` operationId). +func (c *ClientWithResponses) ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPostResponse, error) { + rsp, err := c.ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPost(ctx, clusterId, groupId, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPostResponse(rsp) +} + +// ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetWithResponse Clusters:Consistency-Groups:Replication:Status +// +// The group's replication status as one unit: oldest recovery point, worst +// member lag and health, summed backlog (design-csi-addons-replication.md +// §14.4/§14.6). Never 404s -- a group with no replicating member reports +// “state: not_replicating“. +// +// Returns a wrapper object for the known response body format(s). +// +// Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/status (the `ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGet` operationId). +func (c *ClientWithResponses) ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetResponse, error) { + rsp, err := c.ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGet(ctx, clusterId, groupId, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetResponse(rsp) +} + // ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetWithResponse Clusters:Consistency-Groups:Snapshots:List // // Returns a wrapper object for the known response body format(s). @@ -21849,16 +22846,20 @@ func (c *ClientWithResponses) ClustersStoragePoolsVolumesReplicationFailbackApiV // faithfully replicated. // // “planned=True“ gates on a completed demote (P0-3) so a planned swap -// loses nothing: 412 when no demote was ever requested for this volume (the -// caller's premise that the source is reachable to demote was wrong, and a -// 412 is what lets the csi-addons controller's own force-escalation take -// over), 409 while demote is still converging (retryable -- 409 must never -// become a code the controller reads as permission to force, since that -// controller escalates on ANY FAILED_PRECONDITION from a force=false -// promote with no wait-and-retry grace period of its own). Unplanned -// failover (the default) ignores demote state entirely, unchanged from -// today: its whole premise is that the source may never have been -// reachable to demote. +// loses nothing: 409 while demote is still converging (retryable -- 409 +// must never become a code the controller reads as permission to force, +// since that controller escalates on ANY FAILED_PRECONDITION from a +// force=false promote with no wait-and-retry grace period of its own). +// When no demote was ever requested, the source's own health decides: a +// genuinely healthy, still-serving source means there is nothing to fail +// over -- this is the vendored csi-addons controller's OWN first-ever +// reconcile of a `VolumeReplication` that already lives here, not a +// disaster, and this call succeeds as the no-op it is. A source that is +// NOT healthy gets 412, the caller's premise that it was reachable to +// demote was wrong, and 412 is what lets the controller's own +// force-escalation take over. Unplanned failover (the default) ignores +// demote state entirely, unchanged from today: its whole premise is that +// the source may never have been reachable to demote. // // Returns a wrapper object for the known response body format(s). // @@ -23123,6 +24124,162 @@ func ParseClustersConsistencyGroupsMembersDetachApiV2ClustersClusterIdConsistenc return response, nil } +// ParseClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutResponse parses an HTTP response from a ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutWithResponse call +func ParseClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutResponse(rsp *http.Response) (*ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutResponse, error) { + bodyBytes, err := io.ReadAll(rsp.Body) + defer func() { _ = rsp.Body.Close() }() + if err != nil { + return nil, err + } + + response := &ClustersConsistencyGroupsReplicationConfigureApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationPutResponse{ + Body: bodyBytes, + HTTPResponse: rsp, + } + + switch { + case rsp.StatusCode == 204: + break // No content-type + + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: + var dest HTTPValidationError + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON422 = &dest + + } + + return response, nil +} + +// ParseClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePostResponse parses an HTTP response from a ClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePostWithResponse call +func ParseClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePostResponse(rsp *http.Response) (*ClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePostResponse, error) { + bodyBytes, err := io.ReadAll(rsp.Body) + defer func() { _ = rsp.Body.Close() }() + if err != nil { + return nil, err + } + + response := &ClustersConsistencyGroupsReplicationDemoteApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationDemotePostResponse{ + Body: bodyBytes, + HTTPResponse: rsp, + } + + switch { + case rsp.StatusCode == 202: + break // No content-type + + case rsp.StatusCode == 204: + break // No content-type + + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: + var dest HTTPValidationError + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON422 = &dest + + } + + return response, nil +} + +// ParseClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostResponse parses an HTTP response from a ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostWithResponse call +func ParseClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostResponse(rsp *http.Response) (*ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostResponse, error) { + bodyBytes, err := io.ReadAll(rsp.Body) + defer func() { _ = rsp.Body.Close() }() + if err != nil { + return nil, err + } + + response := &ClustersConsistencyGroupsReplicationFailbackApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailbackPostResponse{ + Body: bodyBytes, + HTTPResponse: rsp, + } + + switch { + case rsp.StatusCode == 204: + break // No content-type + + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: + var dest HTTPValidationError + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON422 = &dest + + } + + return response, nil +} + +// ParseClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPostResponse parses an HTTP response from a ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPostWithResponse call +func ParseClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPostResponse(rsp *http.Response) (*ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPostResponse, error) { + bodyBytes, err := io.ReadAll(rsp.Body) + defer func() { _ = rsp.Body.Close() }() + if err != nil { + return nil, err + } + + response := &ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPostResponse{ + Body: bodyBytes, + HTTPResponse: rsp, + } + + switch { + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: + var dest map[string]interface{} + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON200 = &dest + + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: + var dest HTTPValidationError + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON422 = &dest + + } + + return response, nil +} + +// ParseClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetResponse parses an HTTP response from a ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetWithResponse call +func ParseClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetResponse(rsp *http.Response) (*ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetResponse, error) { + bodyBytes, err := io.ReadAll(rsp.Body) + defer func() { _ = rsp.Body.Close() }() + if err != nil { + return nil, err + } + + response := &ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetResponse{ + Body: bodyBytes, + HTTPResponse: rsp, + } + + switch { + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: + var dest ConsistencyGroupReplicationStatusDTO + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON200 = &dest + + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: + var dest HTTPValidationError + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON422 = &dest + + } + + return response, nil +} + // ParseClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetResponse parses an HTTP response from a ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetWithResponse call func ParseClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetResponse(rsp *http.Response) (*ClustersConsistencyGroupsSnapshotsListApiV2ClustersClusterIdConsistencyGroupsGroupIdSnapshotsGetResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) diff --git a/atlas-lib/lvol/grouphandle.go b/atlas-lib/lvol/grouphandle.go new file mode 100644 index 000000000..ced8902e5 --- /dev/null +++ b/atlas-lib/lvol/grouphandle.go @@ -0,0 +1,63 @@ +// Consistency-group volume handles: the identity a csi-addons VolumeGroup and +// its group-level replication verbs address, kept distinct from a per-volume +// handle so the driver's Replication verbs can route a whole group to the +// cluster-scoped group-replication endpoints instead of the per-volume ones +// (design-csi-addons-replication.md §14.4). It lives beside handle.go because it +// is the same handle grammar with a group sentinel. +package lvol + +import "strings" + +// groupHandlePrefix marks a handle as naming a consistency group rather than a +// single volume. A per-volume handle's first segment is a cluster UUID, so a +// non-UUID sentinel here keeps the two grammars unambiguous: ParseHandle rejects +// a group handle (its first segment is not a UUID), and ParseGroupHandle rejects +// a per-volume one (it lacks the sentinel). +const groupHandlePrefix = "cg" + +// GroupHandle is a consistency-group handle taken apart: the cluster and the +// group it names. The csi-addons VolumeGroup service returns one of these as a +// group's replication handle. The cluster-scoped group-replication endpoints +// need only the cluster and the group id, so a handle carries no pool segment. +type GroupHandle struct { + // ClusterID and GroupID are canonical UUIDs, spelled exactly as the handle + // spells them (not normalized), for the same string-comparison reason + // Handle keeps its segments verbatim. + ClusterID string + GroupID string +} + +// ParseGroupHandle splits a consistency-group handle (cg:{clusterID}:{groupID}) +// into the cluster and group it names, reporting whether it was well formed. +// Both ids must be canonical UUIDs, and the leading "cg" sentinel is what +// distinguishes it from a per-volume handle. Surrounding whitespace is trimmed, +// as ParseHandle trims it, because a handle is read back out of a YAML object. +func ParseGroupHandle(h VolumeHandle) (GroupHandle, bool) { + parts := strings.Split(strings.TrimSpace(string(h)), handleSeparator) + if len(parts) != 3 || parts[0] != groupHandlePrefix { + return GroupHandle{}, false + } + clusterID, groupID := parts[1], parts[2] + if !IsCanonicalUUID(clusterID) || !IsCanonicalUUID(groupID) { + return GroupHandle{}, false + } + return GroupHandle{ClusterID: clusterID, GroupID: groupID}, true +} + +// String renders the group handle back into the form ParseGroupHandle reads. +func (h GroupHandle) String() string { + return strings.Join([]string{groupHandlePrefix, h.ClusterID, h.GroupID}, handleSeparator) +} + +// Handle renders the parts into a VolumeHandle. +func (h GroupHandle) Handle() VolumeHandle { + return VolumeHandle(h.String()) +} + +// IsGroupHandle reports whether a handle names a consistency group rather than a +// single volume. The driver's Replication verbs branch on this to route a group +// handle to the group-replication endpoints (design §14.4). +func IsGroupHandle(h VolumeHandle) bool { + _, ok := ParseGroupHandle(h) + return ok +} diff --git a/atlas-lib/lvol/grouphandle_test.go b/atlas-lib/lvol/grouphandle_test.go new file mode 100644 index 000000000..0562ec629 --- /dev/null +++ b/atlas-lib/lvol/grouphandle_test.go @@ -0,0 +1,80 @@ +package lvol + +import "testing" + +func TestParseGroupHandle(t *testing.T) { + const ( + cluster = "8ffac363-0c46-4714-a71b-f9c0b58a1269" + group = "a1111111-1111-4111-8111-111111111111" + pool = "df34f16c-1a2b-3c4d-5e6f-7a8b9c0d1e2f" + volume = "b2222222-2222-4222-8222-222222222222" + ) + + t.Run("parses a well-formed group handle", func(t *testing.T) { + got, ok := ParseGroupHandle(VolumeHandle("cg:" + cluster + ":" + group)) + if !ok { + t.Fatalf("ParseGroupHandle rejected a well-formed handle") + } + if got.ClusterID != cluster || got.GroupID != group { + t.Fatalf("ParseGroupHandle = %+v, want cluster=%s group=%s", got, cluster, group) + } + }) + + t.Run("trims surrounding whitespace", func(t *testing.T) { + if _, ok := ParseGroupHandle(VolumeHandle(" cg:" + cluster + ":" + group + "\n")); !ok { + t.Fatalf("ParseGroupHandle did not trim whitespace") + } + }) + + t.Run("round-trips through Handle", func(t *testing.T) { + h := GroupHandle{ClusterID: cluster, GroupID: group} + got, ok := ParseGroupHandle(h.Handle()) + if !ok || got != h { + t.Fatalf("round-trip: got %+v ok=%v, want %+v", got, ok, h) + } + }) + + reject := []struct { + name, handle string + }{ + {"a per-volume handle", cluster + ":" + pool + ":" + volume}, + {"the wrong sentinel", "vg:" + cluster + ":" + group}, + {"a missing sentinel", cluster + ":" + group}, + {"a non-UUID cluster", "cg:not-a-uuid:" + group}, + {"a non-UUID group", "cg:" + cluster + ":not-a-uuid"}, + {"too few segments", "cg:" + cluster}, + {"too many segments", "cg:" + cluster + ":" + group + ":extra"}, + {"empty", ""}, + } + for _, tc := range reject { + t.Run("rejects "+tc.name, func(t *testing.T) { + if _, ok := ParseGroupHandle(VolumeHandle(tc.handle)); ok { + t.Fatalf("ParseGroupHandle accepted %q", tc.handle) + } + }) + } +} + +// The two grammars must not overlap: a per-volume handle is never a group +// handle, and a group handle is never a per-volume handle, so the routing branch +// can tell them apart. +func TestGroupAndVolumeHandlesAreDisjoint(t *testing.T) { + const ( + cluster = "8ffac363-0c46-4714-a71b-f9c0b58a1269" + group = "a1111111-1111-4111-8111-111111111111" + pool = "df34f16c-1a2b-3c4d-5e6f-7a8b9c0d1e2f" + volume = "b2222222-2222-4222-8222-222222222222" + ) + groupHandle := VolumeHandle("cg:" + cluster + ":" + group) + volumeHandle := VolumeHandle(cluster + ":" + pool + ":" + volume) + + if _, ok := ParseHandle(groupHandle); ok { + t.Errorf("ParseHandle accepted a group handle %q", groupHandle) + } + if !IsGroupHandle(groupHandle) { + t.Errorf("IsGroupHandle(%q) = false, want true", groupHandle) + } + if IsGroupHandle(volumeHandle) { + t.Errorf("IsGroupHandle(%q) = true, want false", volumeHandle) + } +} diff --git a/csi-driver/go.mod b/csi-driver/go.mod index 3544e2a6d..b574f9b3b 100644 --- a/csi-driver/go.mod +++ b/csi-driver/go.mod @@ -4,6 +4,7 @@ go 1.26.2 require ( github.com/container-storage-interface/spec v1.12.0 + github.com/csi-addons/spec v0.2.1-0.20260515055340-d4a373713b9a github.com/kubernetes-csi/csi-lib-utils v0.24.0 github.com/kubernetes-csi/csi-test/v5 v5.5.0 github.com/onsi/gomega v1.42.1 @@ -25,7 +26,6 @@ require ( github.com/antlr4-go/antlr/v4 v4.13.0 // indirect github.com/apapsch/go-jsonmerge/v2 v2.0.0 // indirect github.com/cenkalti/backoff/v5 v5.0.3 // indirect - github.com/csi-addons/spec v0.2.0 // indirect github.com/distribution/reference v0.6.0 // indirect github.com/dprotaso/go-yit v0.0.0-20220510233725-9ba8df137936 // indirect github.com/fsnotify/fsnotify v1.9.0 // indirect @@ -37,7 +37,6 @@ require ( github.com/go-playground/universal-translator v0.18.1 // indirect github.com/go-playground/validator/v10 v10.30.3 // indirect github.com/go-task/slim-sprig/v3 v3.0.0 // indirect - github.com/golang/protobuf v1.5.4 // indirect github.com/google/cel-go v0.26.0 // indirect github.com/google/gnostic-models v0.7.0 // indirect github.com/google/pprof v0.0.0-20260402051712-545e8a4df936 // indirect diff --git a/csi-driver/go.sum b/csi-driver/go.sum index fa9883c5c..40a2cc779 100644 --- a/csi-driver/go.sum +++ b/csi-driver/go.sum @@ -24,8 +24,8 @@ github.com/chzyer/test v0.0.0-20180213035817-a1ea475d72b1/go.mod h1:Q3SI9o4m/ZMn github.com/container-storage-interface/spec v1.12.0 h1:zrFOEqpR5AghNaaDG4qyedwPBqU2fU0dWjLQMP/azK0= github.com/container-storage-interface/spec v1.12.0/go.mod h1:txsm+MA2B2WDa5kW69jNbqPnvTtfvZma7T/zsAZ9qX8= github.com/cpuguy83/go-md2man/v2 v2.0.6/go.mod h1:oOW0eioCTA6cOiMLiUPZOpcVxMig6NIQQ7OS05n1F4g= -github.com/csi-addons/spec v0.2.0 h1:Ews7bxpN9P6nFxl1XvMg87cR1wLROdH1FzSfLfb4VfI= -github.com/csi-addons/spec v0.2.0/go.mod h1:Mwq4iLiUV4s+K1bszcWU6aMsR5KPsbIYzzszJ6+56vI= +github.com/csi-addons/spec v0.2.1-0.20260515055340-d4a373713b9a h1:AxvSsTN8TxwpeP6X/ScNK5i+6eXu3APRfuIXSEMENVQ= +github.com/csi-addons/spec v0.2.1-0.20260515055340-d4a373713b9a/go.mod h1:Mwq4iLiUV4s+K1bszcWU6aMsR5KPsbIYzzszJ6+56vI= github.com/davecgh/go-spew v1.1.0/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38= github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38= github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc h1:U9qPSI2PIWSS1VwoXQT9A3Wy9MM3WgvqSxFWenqJduM= diff --git a/csi-driver/internal/csi/controller/mock_controlplane_test.go b/csi-driver/internal/csi/controller/mock_controlplane_test.go index ace3b66ea..1d29d950f 100644 --- a/csi-driver/internal/csi/controller/mock_controlplane_test.go +++ b/csi-driver/internal/csi/controller/mock_controlplane_test.go @@ -234,10 +234,34 @@ func newMockSBCLI() *mockSBCLI { "DELETE /api/v2/clusters/{clusterID}/storage-pools/{poolID}/snapshots/{snapshotID}/", m.locked(m.handleDeleteSnapshot), ) + mux.HandleFunc( + "GET /api/v2/clusters/{clusterID}/consistency-groups/{$}", + m.locked(m.handleListGroups), + ) mux.HandleFunc( "GET /api/v2/clusters/{clusterID}/consistency-groups/{groupID}/members", m.locked(m.handleGroupMembers), ) + mux.HandleFunc( + "PUT /api/v2/clusters/{clusterID}/consistency-groups/{groupID}/replication", + m.locked(m.handleConfigureGroupReplication), + ) + mux.HandleFunc( + "POST /api/v2/clusters/{clusterID}/consistency-groups/{groupID}/replication/failover", + m.locked(m.handleGroupFailover), + ) + mux.HandleFunc( + "POST /api/v2/clusters/{clusterID}/consistency-groups/{groupID}/replication/demote", + m.locked(m.handleGroupDemote), + ) + mux.HandleFunc( + "POST /api/v2/clusters/{clusterID}/consistency-groups/{groupID}/replication/failback", + m.locked(m.handleGroupFailback), + ) + mux.HandleFunc( + "GET /api/v2/clusters/{clusterID}/consistency-groups/{groupID}/replication/status", + m.locked(m.handleGroupReplicationStatus), + ) mux.HandleFunc( "POST /api/v2/clusters/{clusterID}/consistency-groups/{groupID}/snapshots", m.locked(m.handleTakeGroupSnapshot), @@ -639,6 +663,13 @@ type mockGroup struct { Members []string // lvol UUIDs with an open epoch LastSeq int Gens map[int][]mockGenMember + // Group-replication state the driver's group-handle routing drives. + PolicyID string + Promoted bool + Demoted bool + DemoteConverging bool // when set, /demote answers 202 (still converging) + FailbackSource string + LastReplicatedAt int64 // unix seconds surfaced by /replication/status } // seedGroup registers a group with the given member lvol UUIDs and stamps each @@ -654,6 +685,93 @@ func (m *mockSBCLI) seedGroup(groupID string, memberUUIDs ...string) { } } +func (m *mockSBCLI) handleListGroups(w http.ResponseWriter, r *http.Request) { + name := r.URL.Query().Get("name") + rows := make([]map[string]any, 0, len(m.groups)) + for id, g := range m.groups { + gname := "cg-" + id + if name != "" && name != gname { + continue + } + rows = append(rows, map[string]any{ + "id": id, "cluster_id": r.PathValue("clusterID"), "name": gname, + "member_count": len(g.Members), "last_group_seq": g.LastSeq, + }) + } + writeJSON(w, http.StatusOK, rows) +} + +func (m *mockSBCLI) handleConfigureGroupReplication(w http.ResponseWriter, r *http.Request) { + g := m.groups[r.PathValue("groupID")] + if g == nil { + writeJSON(w, http.StatusNotFound, map[string]string{"detail": "group not found"}) + return + } + var body struct { + ReplicationPolicyID *string `json:"replication_policy_id"` + } + _ = json.NewDecoder(r.Body).Decode(&body) + if body.ReplicationPolicyID == nil { + g.PolicyID = "" + } else { + g.PolicyID = *body.ReplicationPolicyID + } + w.WriteHeader(http.StatusNoContent) +} + +func (m *mockSBCLI) handleGroupFailover(w http.ResponseWriter, r *http.Request) { + g := m.groups[r.PathValue("groupID")] + if g == nil { + writeJSON(w, http.StatusNotFound, map[string]string{"detail": "group not found"}) + return + } + g.Promoted = true + writeJSON(w, http.StatusOK, map[string]any{"members": []any{}}) +} + +func (m *mockSBCLI) handleGroupDemote(w http.ResponseWriter, r *http.Request) { + g := m.groups[r.PathValue("groupID")] + if g == nil { + writeJSON(w, http.StatusNotFound, map[string]string{"detail": "group not found"}) + return + } + if g.DemoteConverging { + writeJSON(w, http.StatusAccepted, map[string]any{"demoted": false}) + return + } + g.Demoted = true + w.WriteHeader(http.StatusNoContent) +} + +func (m *mockSBCLI) handleGroupFailback(w http.ResponseWriter, r *http.Request) { + g := m.groups[r.PathValue("groupID")] + if g == nil { + writeJSON(w, http.StatusNotFound, map[string]string{"detail": "group not found"}) + return + } + var body struct { + SourceClusterID *string `json:"source_cluster_id"` + } + _ = json.NewDecoder(r.Body).Decode(&body) + if body.SourceClusterID != nil { + g.FailbackSource = *body.SourceClusterID + } + w.WriteHeader(http.StatusNoContent) +} + +func (m *mockSBCLI) handleGroupReplicationStatus(w http.ResponseWriter, r *http.Request) { + g := m.groups[r.PathValue("groupID")] + if g == nil { + writeJSON(w, http.StatusNotFound, map[string]string{"detail": "group not found"}) + return + } + out := map[string]any{"role": "source", "state": "in_sync", "member_count": len(g.Members)} + if g.LastReplicatedAt != 0 { + out["last_replicated_at"] = time.Unix(g.LastReplicatedAt, 0).UTC().Format(time.RFC3339) + } + writeJSON(w, http.StatusOK, out) +} + func (m *mockSBCLI) handleGroupMembers(w http.ResponseWriter, r *http.Request) { g := m.groups[r.PathValue("groupID")] if g == nil { diff --git a/csi-driver/internal/csi/controller/replication.go b/csi-driver/internal/csi/controller/replication.go index 1477a68d1..88ddea84d 100644 --- a/csi-driver/internal/csi/controller/replication.go +++ b/csi-driver/internal/csi/controller/replication.go @@ -131,6 +131,19 @@ func (cs *Server) EnableVolumeReplication( return nil, status.Errorf(codes.InvalidArgument, "VolumeReplicationClass parameter %q is required", replicationPolicyParam) } + // A group handle drives the whole consistency group as one unit through the + // group-replication endpoints (design §14.4); a per-volume handle takes the + // §5 path below unchanged. + if gh, ok := lvol.ParseGroupHandle(lvol.VolumeHandle(volumeIDFrom(req))); ok { + client, err := clusters.ReplicationClient(ctx, gh.ClusterID) + if err != nil { + return nil, status.Error(codes.Unavailable, err.Error()) + } + if err := client.EnableGroupReplication(ctx, gh, policyID); err != nil { + return nil, classifyEnableVolumeReplicationError(err) + } + return &replication.EnableVolumeReplicationResponse{}, nil + } h, err := csicommon.ParseVolumeHandle(volumeIDFrom(req)) if err != nil { return nil, status.Error(codes.InvalidArgument, err.Error()) @@ -172,6 +185,16 @@ func (cs *Server) DisableVolumeReplication( ctx context.Context, req *replication.DisableVolumeReplicationRequest, ) (*replication.DisableVolumeReplicationResponse, error) { + if gh, ok := lvol.ParseGroupHandle(lvol.VolumeHandle(volumeIDFrom(req))); ok { + client, err := clusters.ReplicationClient(ctx, gh.ClusterID) + if err != nil { + return nil, status.Error(codes.Unavailable, err.Error()) + } + if err := client.DisableGroupReplication(ctx, gh); err != nil { + return nil, classifyDisableVolumeReplicationError(err) + } + return &replication.DisableVolumeReplicationResponse{}, nil + } h, err := csicommon.ParseVolumeHandle(volumeIDFrom(req)) if err != nil { return nil, status.Error(codes.InvalidArgument, err.Error()) @@ -204,6 +227,21 @@ func (cs *Server) GetVolumeReplicationInfo( ctx context.Context, req *replication.GetVolumeReplicationInfoRequest, ) (*replication.GetVolumeReplicationInfoResponse, error) { + if gh, ok := lvol.ParseGroupHandle(lvol.VolumeHandle(volumeIDFrom(req))); ok { + client, err := clusters.ReplicationClient(ctx, gh.ClusterID) + if err != nil { + return nil, status.Error(codes.Unavailable, err.Error()) + } + info, err := client.GetGroupReplicationInfo(ctx, gh) + if err != nil { + return nil, classifyGetVolumeReplicationInfoError(err) + } + resp := &replication.GetVolumeReplicationInfoResponse{} + if info.LastReplicatedAt != nil { + resp.LastSyncTime = timestamppb.New(*info.LastReplicatedAt) + } + return resp, nil + } h, err := csicommon.ParseVolumeHandle(volumeIDFrom(req)) if err != nil { return nil, status.Error(codes.InvalidArgument, err.Error()) @@ -240,6 +278,20 @@ func (cs *Server) PromoteVolume( ctx context.Context, req *replication.PromoteVolumeRequest, ) (*replication.PromoteVolumeResponse, error) { + // A group handle promotes the whole consistency group atomically (design + // §14.4): every member is cloned from the same group generation. The + // planned/forced split is the backend group failover's own concern, so + // force is not forwarded here. + if gh, ok := lvol.ParseGroupHandle(lvol.VolumeHandle(volumeIDFrom(req))); ok { + client, err := clusters.ReplicationClient(ctx, gh.ClusterID) + if err != nil { + return nil, status.Error(codes.Unavailable, err.Error()) + } + if err := client.PromoteGroup(ctx, gh); err != nil { + return nil, classifyPromoteVolumeError(err) + } + return &replication.PromoteVolumeResponse{}, nil + } h, err := csicommon.ParseVolumeHandle(volumeIDFrom(req)) if err != nil { return nil, status.Error(codes.InvalidArgument, err.Error()) @@ -280,6 +332,20 @@ func (cs *Server) DemoteVolume( ctx context.Context, req *replication.DemoteVolumeRequest, ) (*replication.DemoteVolumeResponse, error) { + if gh, ok := lvol.ParseGroupHandle(lvol.VolumeHandle(volumeIDFrom(req))); ok { + client, err := clusters.ReplicationClient(ctx, gh.ClusterID) + if err != nil { + return nil, status.Error(codes.Unavailable, err.Error()) + } + done, err := client.DemoteGroup(ctx, gh) + if err != nil { + return nil, classifyDemoteVolumeError(err) + } + if !done { + return nil, status.Error(codes.Aborted, "group demote is still converging") + } + return &replication.DemoteVolumeResponse{}, nil + } h, err := csicommon.ParseVolumeHandle(volumeIDFrom(req)) if err != nil { return nil, status.Error(codes.InvalidArgument, err.Error()) @@ -318,6 +384,21 @@ func (cs *Server) ResyncVolume( ctx context.Context, req *replication.ResyncVolumeRequest, ) (*replication.ResyncVolumeResponse, error) { + if gh, ok := lvol.ParseGroupHandle(lvol.VolumeHandle(volumeIDFrom(req))); ok { + client, err := clusters.ReplicationClient(ctx, gh.ClusterID) + if err != nil { + return nil, status.Error(codes.Unavailable, err.Error()) + } + if err := client.ResyncGroup(ctx, gh, req.GetParameters()[sourceClusterIDParam]); err != nil { + return nil, classifyResyncVolumeError(err) + } + info, err := client.GetGroupReplicationInfo(ctx, gh) + if err != nil { + return nil, classifyGetVolumeReplicationInfoError(err) + } + ready := info.LagSeconds == nil || info.LagBudgetSeconds == nil || *info.LagSeconds <= *info.LagBudgetSeconds + return &replication.ResyncVolumeResponse{Ready: ready}, nil + } h, err := csicommon.ParseVolumeHandle(volumeIDFrom(req)) if err != nil { return nil, status.Error(codes.InvalidArgument, err.Error()) diff --git a/csi-driver/internal/csi/controller/replication_group_test.go b/csi-driver/internal/csi/controller/replication_group_test.go new file mode 100644 index 000000000..94e6cd950 --- /dev/null +++ b/csi-driver/internal/csi/controller/replication_group_test.go @@ -0,0 +1,148 @@ +package controller + +import ( + "context" + "testing" + + "github.com/csi-addons/spec/lib/go/replication" + "google.golang.org/grpc/codes" + "google.golang.org/grpc/status" +) + +// The Replication verbs route a cg: group handle to the group-replication +// endpoints, driving the whole consistency group as one unit (design §14.4). A +// per-volume handle still takes the §5 path (covered by replication_test.go). + +const ( + vgGroupHandle = "cg:" + sanityClusterID + ":" + vgGroupID + vgPolicyID = "dddddddd-dddd-4ddd-8ddd-dddddddddddd" +) + +func groupSource(handle string) *replication.ReplicationSource { + return &replication.ReplicationSource{ + Type: &replication.ReplicationSource_Volume{ + Volume: &replication.ReplicationSource_VolumeSource{VolumeId: handle}, + }, + } +} + +func newGroupReplTestServer(t *testing.T, mock *mockSBCLI) *Server { + t.Helper() + mock.seedGroup(vgGroupID, vgMember1, vgMember2) + return newTestControllerServer(t, mock) +} + +func TestEnableVolumeReplicationRoutesAGroupHandle(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newGroupReplTestServer(t, mock) + + _, err := cs.EnableVolumeReplication(context.Background(), &replication.EnableVolumeReplicationRequest{ + ReplicationSource: groupSource(vgGroupHandle), + Parameters: map[string]string{replicationPolicyParam: vgPolicyID}, + }) + if err != nil { + t.Fatalf("EnableVolumeReplication: %v", err) + } + if got := mock.groups[vgGroupID].PolicyID; got != vgPolicyID { + t.Fatalf("group policy = %q, want %q", got, vgPolicyID) + } +} + +func TestDisableVolumeReplicationRoutesAGroupHandle(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newGroupReplTestServer(t, mock) + mock.groups[vgGroupID].PolicyID = vgPolicyID + + _, err := cs.DisableVolumeReplication(context.Background(), &replication.DisableVolumeReplicationRequest{ + ReplicationSource: groupSource(vgGroupHandle), + }) + if err != nil { + t.Fatalf("DisableVolumeReplication: %v", err) + } + if got := mock.groups[vgGroupID].PolicyID; got != "" { + t.Fatalf("group policy = %q after disable, want empty", got) + } +} + +func TestPromoteVolumeRoutesAGroupHandle(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newGroupReplTestServer(t, mock) + + if _, err := cs.PromoteVolume(context.Background(), &replication.PromoteVolumeRequest{ + ReplicationSource: groupSource(vgGroupHandle), Force: true, + }); err != nil { + t.Fatalf("PromoteVolume: %v", err) + } + if !mock.groups[vgGroupID].Promoted { + t.Fatal("group was not promoted") + } +} + +func TestDemoteVolumeRoutesAGroupHandle(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newGroupReplTestServer(t, mock) + + if _, err := cs.DemoteVolume(context.Background(), &replication.DemoteVolumeRequest{ + ReplicationSource: groupSource(vgGroupHandle), + }); err != nil { + t.Fatalf("DemoteVolume: %v", err) + } + if !mock.groups[vgGroupID].Demoted { + t.Fatal("group was not demoted") + } +} + +func TestDemoteVolumeGroupStillConvergingIsAborted(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newGroupReplTestServer(t, mock) + mock.groups[vgGroupID].DemoteConverging = true + + _, err := cs.DemoteVolume(context.Background(), &replication.DemoteVolumeRequest{ + ReplicationSource: groupSource(vgGroupHandle), + }) + if status.Code(err) != codes.Aborted { + t.Fatalf("err = %v, want Aborted", err) + } +} + +func TestResyncVolumeRoutesAGroupHandle(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newGroupReplTestServer(t, mock) + + resp, err := cs.ResyncVolume(context.Background(), &replication.ResyncVolumeRequest{ + ReplicationSource: groupSource(vgGroupHandle), + Parameters: map[string]string{sourceClusterIDParam: sanityClusterID}, + }) + if err != nil { + t.Fatalf("ResyncVolume: %v", err) + } + if !resp.GetReady() { + t.Error("expected Ready (group status has no lag budget, so ready)") + } + if got := mock.groups[vgGroupID].FailbackSource; got != sanityClusterID { + t.Fatalf("failback source = %q, want %q", got, sanityClusterID) + } +} + +func TestGetVolumeReplicationInfoRoutesAGroupHandle(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newGroupReplTestServer(t, mock) + mock.groups[vgGroupID].LastReplicatedAt = 1_700_000_000 + + resp, err := cs.GetVolumeReplicationInfo(context.Background(), &replication.GetVolumeReplicationInfoRequest{ + ReplicationSource: groupSource(vgGroupHandle), + }) + if err != nil { + t.Fatalf("GetVolumeReplicationInfo: %v", err) + } + if resp.GetLastSyncTime() == nil { + t.Fatal("expected a LastSyncTime for the group") + } +} diff --git a/csi-driver/internal/csi/controller/server.go b/csi-driver/internal/csi/controller/server.go index 46e3f801b..ffe66d755 100644 --- a/csi-driver/internal/csi/controller/server.go +++ b/csi-driver/internal/csi/controller/server.go @@ -21,6 +21,11 @@ type Server struct { // (PromoteVolume, DemoteVolume, ResyncVolume) fall through to this // embedded default until Phase 2. replication.UnimplementedControllerServer + // The csi-addons VolumeGroup (GroupController) service (design §14.3) is + // implemented in volumegroup.go. Its unimplemented base is embedded through + // a named wrapper (volumeGroupUnimplemented) because the replication base + // above already occupies the UnimplementedControllerServer embed name. + volumeGroupUnimplemented volumeLocks *csicommon.VolumeLocks // kubeClient reads/patches PVC annotations (host_id resolution, placement-hint // cleanup). Built once at construction and reused, and nil when no in-cluster diff --git a/csi-driver/internal/csi/controller/volumegroup.go b/csi-driver/internal/csi/controller/volumegroup.go new file mode 100644 index 000000000..9123cea8e --- /dev/null +++ b/csi-driver/internal/csi/controller/volumegroup.go @@ -0,0 +1,116 @@ +// The csi-addons VolumeGroup (GroupController) service (design +// design-csi-addons-replication.md §14.3): the stock kubernetes-csi-addons +// controller-manager dials it to form a backend consistency group before +// replicating it as one unit. CreateVolumeGroup resolves the group the member +// volumes already belong to (they joined at provisioning by the +// storage.simplyblock.io/consistency-group label) and hands back a group handle +// the Replication verbs route on. Membership and lifecycle are owned by that +// label, not by this service, so ModifyVolumeGroupMembership and +// DeleteVolumeGroup never reshape or delete the backend group. +package controller + +import ( + "context" + + "github.com/csi-addons/spec/lib/go/volumegroup" + "google.golang.org/grpc/codes" + "google.golang.org/grpc/status" + + "github.com/simplyblock/atlas/lvol" + + "github.com/simplyblock/csi-driver/internal/clusters" +) + +// volumeGroupUnimplemented embeds the csi-addons VolumeGroup unimplemented server +// so Server can satisfy the service's forward-compat guard under a name that does +// not collide with the replication base's own UnimplementedControllerServer. +type volumeGroupUnimplemented struct { + volumegroup.UnimplementedControllerServer +} + +// CreateVolumeGroup resolves the backend consistency group the given member +// volumes belong to and returns its group handle (design §14.3). It creates no +// group of its own: the group is label-formed at provisioning, so this is +// idempotent and returns the existing group's handle. +func (cs *Server) CreateVolumeGroup( + ctx context.Context, + req *volumegroup.CreateVolumeGroupRequest, +) (*volumegroup.CreateVolumeGroupResponse, error) { + clusterID, lvolIDs, err := parseGroupMembers(req.GetVolumeIds()) + if err != nil { + return nil, err + } + client, err := clusters.ReplicationClient(ctx, clusterID) + if err != nil { + return nil, status.Error(codes.Unavailable, err.Error()) + } + groupID, err := client.ConsistencyGroupForLvols(ctx, clusterID, lvolIDs) + if err != nil { + // No backend group matches the selection exactly: the members are not + // one whole consistency group (design §14.3, the admission webhook's + // invariant), so refuse rather than group a partial set. + return nil, status.Error(codes.FailedPrecondition, err.Error()) + } + gh := lvol.GroupHandle{ClusterID: clusterID, GroupID: groupID} + return &volumegroup.CreateVolumeGroupResponse{ + VolumeGroup: &volumegroup.VolumeGroup{VolumeGroupId: string(gh.Handle())}, + }, nil +} + +// ModifyVolumeGroupMembership is a success no-op: a group's membership is owned +// by the storage.simplyblock.io/consistency-group label at provisioning, and +// dynamic membership through this RPC is deferred (design §14.3, +// design-consistency-groups.md Phase 4). Returning success also lets the +// controller-manager's teardown, which empties a group before deleting it, +// proceed without error. +func (cs *Server) ModifyVolumeGroupMembership( + _ context.Context, + req *volumegroup.ModifyVolumeGroupMembershipRequest, +) (*volumegroup.ModifyVolumeGroupMembershipResponse, error) { + if _, ok := lvol.ParseGroupHandle(lvol.VolumeHandle(req.GetVolumeGroupId())); !ok { + return nil, status.Errorf(codes.InvalidArgument, + "not a consistency-group handle: %q", req.GetVolumeGroupId()) + } + return &volumegroup.ModifyVolumeGroupMembershipResponse{ + VolumeGroup: &volumegroup.VolumeGroup{VolumeGroupId: req.GetVolumeGroupId()}, + }, nil +} + +// DeleteVolumeGroup is a success no-op: the backend consistency group is +// label-formed and lives as long as a labeled member exists, so deleting the +// csi-addons grouping never deletes the backend group or its volumes (design +// §14.3). Idempotent. +func (cs *Server) DeleteVolumeGroup( + _ context.Context, + req *volumegroup.DeleteVolumeGroupRequest, +) (*volumegroup.DeleteVolumeGroupResponse, error) { + if _, ok := lvol.ParseGroupHandle(lvol.VolumeHandle(req.GetVolumeGroupId())); !ok { + return nil, status.Errorf(codes.InvalidArgument, + "not a consistency-group handle: %q", req.GetVolumeGroupId()) + } + return &volumegroup.DeleteVolumeGroupResponse{}, nil +} + +// parseGroupMembers parses the member volume handles of a CreateVolumeGroup +// request into a shared cluster id and the member lvol ids, rejecting a +// malformed handle or members that span clusters. +func parseGroupMembers(volumeIDs []string) (clusterID string, lvolIDs []string, err error) { + if len(volumeIDs) == 0 { + return "", nil, status.Error(codes.InvalidArgument, + "CreateVolumeGroup requires at least one volume") + } + lvolIDs = make([]string, 0, len(volumeIDs)) + for _, vid := range volumeIDs { + h, ok := lvol.ParseHandle(lvol.VolumeHandle(vid)) + if !ok { + return "", nil, status.Errorf(codes.InvalidArgument, "malformed volume handle %q", vid) + } + if clusterID != "" && h.ClusterID != clusterID { + return "", nil, status.Error(codes.InvalidArgument, + "a volume group's members must all live in one cluster") + } + clusterID = h.ClusterID + lvolIDs = append(lvolIDs, h.VolumeID) + } + return clusterID, lvolIDs, nil +} diff --git a/csi-driver/internal/csi/controller/volumegroup_test.go b/csi-driver/internal/csi/controller/volumegroup_test.go new file mode 100644 index 000000000..788d7a6b4 --- /dev/null +++ b/csi-driver/internal/csi/controller/volumegroup_test.go @@ -0,0 +1,129 @@ +package controller + +import ( + "context" + "testing" + + "github.com/csi-addons/spec/lib/go/volumegroup" + "google.golang.org/grpc/codes" + "google.golang.org/grpc/status" +) + +const ( + vgGroupID = "c9c9c9c9-c9c9-4c9c-8c9c-c9c9c9c9c9c9" + vgMember1 = "a1111111-1111-4111-8111-111111111111" + vgMember2 = "b2222222-2222-4222-8222-222222222222" +) + +func vgHandle(volumeID string) string { + return sanityClusterID + ":" + sanityPoolUUID + ":" + volumeID +} + +// CreateVolumeGroup resolves the backend consistency group its member volumes +// already belong to, and returns that group's handle (design §14.3). +func TestCreateVolumeGroupResolvesTheBackendGroup(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + mock.seedGroup(vgGroupID, vgMember1, vgMember2) + cs := newTestControllerServer(t, mock) + + resp, err := cs.CreateVolumeGroup(context.Background(), &volumegroup.CreateVolumeGroupRequest{ + Name: "vgrcontent-generated-name", + VolumeIds: []string{vgHandle(vgMember1), vgHandle(vgMember2)}, + }) + if err != nil { + t.Fatalf("CreateVolumeGroup: %v", err) + } + want := "cg:" + sanityClusterID + ":" + vgGroupID + if got := resp.GetVolumeGroup().GetVolumeGroupId(); got != want { + t.Fatalf("group handle = %q, want %q", got, want) + } +} + +// A selection that is not exactly one backend group's membership must not be +// silently grouped. +func TestCreateVolumeGroupRefusesAPartialSelection(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + mock.seedGroup(vgGroupID, vgMember1, vgMember2) + cs := newTestControllerServer(t, mock) + + _, err := cs.CreateVolumeGroup(context.Background(), &volumegroup.CreateVolumeGroupRequest{ + VolumeIds: []string{vgHandle(vgMember1)}, // only one of the two members + }) + if status.Code(err) != codes.FailedPrecondition { + t.Fatalf("err = %v, want FailedPrecondition", err) + } +} + +func TestCreateVolumeGroupRejectsAMalformedHandle(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newTestControllerServer(t, mock) + + _, err := cs.CreateVolumeGroup(context.Background(), &volumegroup.CreateVolumeGroupRequest{ + VolumeIds: []string{"not-a-volume-handle"}, + }) + if status.Code(err) != codes.InvalidArgument { + t.Fatalf("err = %v, want InvalidArgument", err) + } +} + +func TestCreateVolumeGroupRequiresAVolume(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newTestControllerServer(t, mock) + + _, err := cs.CreateVolumeGroup(context.Background(), &volumegroup.CreateVolumeGroupRequest{}) + if status.Code(err) != codes.InvalidArgument { + t.Fatalf("err = %v, want InvalidArgument", err) + } +} + +// ModifyVolumeGroupMembership is a success no-op (membership is label-driven, +// design §14.3), so the controller-manager's teardown does not error. +func TestModifyVolumeGroupMembershipIsANoOpSuccess(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newTestControllerServer(t, mock) + + gh := "cg:" + sanityClusterID + ":" + vgGroupID + resp, err := cs.ModifyVolumeGroupMembership(context.Background(), + &volumegroup.ModifyVolumeGroupMembershipRequest{VolumeGroupId: gh}) + if err != nil { + t.Fatalf("ModifyVolumeGroupMembership: %v", err) + } + if got := resp.GetVolumeGroup().GetVolumeGroupId(); got != gh { + t.Fatalf("group handle = %q, want %q", got, gh) + } +} + +// DeleteVolumeGroup is a success no-op: the backend group is label-formed and +// survives the csi-addons grouping (design §14.3). +func TestDeleteVolumeGroupIsANoOpSuccess(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newTestControllerServer(t, mock) + + gh := "cg:" + sanityClusterID + ":" + vgGroupID + if _, err := cs.DeleteVolumeGroup(context.Background(), + &volumegroup.DeleteVolumeGroupRequest{VolumeGroupId: gh}); err != nil { + t.Fatalf("DeleteVolumeGroup: %v", err) + } +} + +func TestVolumeGroupVerbsRejectANonGroupHandle(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newTestControllerServer(t, mock) + perVolume := vgHandle(vgMember1) // a per-volume handle, not a group handle + + if _, err := cs.ModifyVolumeGroupMembership(context.Background(), + &volumegroup.ModifyVolumeGroupMembershipRequest{VolumeGroupId: perVolume}); status.Code(err) != codes.InvalidArgument { + t.Errorf("Modify: err = %v, want InvalidArgument", err) + } + if _, err := cs.DeleteVolumeGroup(context.Background(), + &volumegroup.DeleteVolumeGroupRequest{VolumeGroupId: perVolume}); status.Code(err) != codes.InvalidArgument { + t.Errorf("Delete: err = %v, want InvalidArgument", err) + } +} diff --git a/csi-driver/internal/csi/csiaddons/identity/identity.go b/csi-driver/internal/csi/csiaddons/identity/identity.go index 8def40c96..7e659d6e4 100644 --- a/csi-driver/internal/csi/csiaddons/identity/identity.go +++ b/csi-driver/internal/csi/csiaddons/identity/identity.go @@ -49,6 +49,33 @@ func (s *Server) GetCapabilities( }, }, }, + // The VolumeGroup service (design §14.3), which the stock + // kubernetes-csi-addons controller-manager dials to form a backend + // consistency group before replicating it as one unit. + { + Type: &identity.Capability_VolumeGroup_{ + VolumeGroup: &identity.Capability_VolumeGroup{ + Type: identity.Capability_VolumeGroup_VOLUME_GROUP, + }, + }, + }, + { + Type: &identity.Capability_VolumeGroup_{ + VolumeGroup: &identity.Capability_VolumeGroup{ + Type: identity.Capability_VolumeGroup_MODIFY_VOLUME_GROUP, + }, + }, + }, + // DeleteVolumeGroup dissolves the group but keeps its member volumes + // (design §14.3): a group is a label-formed set, never an owner of + // the volumes' lifecycle. + { + Type: &identity.Capability_VolumeGroup_{ + VolumeGroup: &identity.Capability_VolumeGroup{ + Type: identity.Capability_VolumeGroup_DO_NOT_ALLOW_VG_TO_DELETE_VOLUMES, + }, + }, + }, }, }, nil } diff --git a/csi-driver/internal/csi/csiaddons/identity/identity_test.go b/csi-driver/internal/csi/csiaddons/identity/identity_test.go index ad47630d1..d9e18a4dc 100644 --- a/csi-driver/internal/csi/csiaddons/identity/identity_test.go +++ b/csi-driver/internal/csi/csiaddons/identity/identity_test.go @@ -41,6 +41,35 @@ func TestGetCapabilitiesAdvertisesVolumeReplication(t *testing.T) { } } +func TestGetCapabilitiesAdvertisesVolumeGroup(t *testing.T) { + s := New("test.csi.simplyblock.io", "v1.2.3") + resp, err := s.GetCapabilities(context.Background(), &identity.GetCapabilitiesRequest{}) + if err != nil { + t.Fatal(err) + } + var sawVolumeGroup, sawDoNotDeleteVolumes bool + for _, c := range resp.Capabilities { + vg := c.GetVolumeGroup() + if vg == nil { + continue + } + switch vg.Type { + case identity.Capability_VolumeGroup_VOLUME_GROUP: + sawVolumeGroup = true + case identity.Capability_VolumeGroup_DO_NOT_ALLOW_VG_TO_DELETE_VOLUMES: + sawDoNotDeleteVolumes = true + } + } + if !sawVolumeGroup { + t.Error("capabilities do not advertise VOLUME_GROUP") + } + // DeleteVolumeGroup dissolves the group but keeps its member volumes + // (design §14.3), which is exactly what this capability promises. + if !sawDoNotDeleteVolumes { + t.Error("capabilities do not advertise DO_NOT_ALLOW_VG_TO_DELETE_VOLUMES") + } +} + func TestProbeReportsReady(t *testing.T) { s := New("test.csi.simplyblock.io", "v1.2.3") resp, err := s.Probe(context.Background(), &identity.ProbeRequest{}) diff --git a/csi-driver/internal/driver/driver.go b/csi-driver/internal/driver/driver.go index aee013d73..fa0ab78d4 100644 --- a/csi-driver/internal/driver/driver.go +++ b/csi-driver/internal/driver/driver.go @@ -30,6 +30,7 @@ import ( "github.com/container-storage-interface/spec/lib/go/csi" csiaddonsidentity "github.com/csi-addons/spec/lib/go/identity" csiaddonsreplication "github.com/csi-addons/spec/lib/go/replication" + csiaddonsvolumegroup "github.com/csi-addons/spec/lib/go/volumegroup" "google.golang.org/grpc" "k8s.io/client-go/kubernetes" "k8s.io/client-go/rest" @@ -148,6 +149,12 @@ func Run(conf *config.Config) { register = append(register, func(gs *grpc.Server) { csiaddonsreplication.RegisterControllerServer(gs, cs) }) + // The csi-addons VolumeGroup service (design §14.3): the stock + // controller-manager dials it to form a backend consistency group before + // replicating the group as one unit. + register = append(register, func(gs *grpc.Server) { + csiaddonsvolumegroup.RegisterControllerServer(gs, cs) + }) } s := csicommon.NewNonBlockingGRPCServer() diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index f4e2c29db..0272bab84 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -10841,44 +10841,6 @@ rules: - patch - update - watch -- apiGroups: - - replication.storage.openshift.io - resources: - - volumegroupreplicationclasses - verbs: - - get - - list - - watch -- apiGroups: - - replication.storage.openshift.io - resources: - - volumegroupreplications - verbs: - - get - - list - - patch - - update - - watch -- apiGroups: - - replication.storage.openshift.io - resources: - - volumegroupreplications/status - verbs: - - get - - patch - - update -- apiGroups: - - replication.storage.openshift.io - resources: - - volumereplications - verbs: - - create - - delete - - get - - list - - patch - - update - - watch - apiGroups: - snapshot.storage.k8s.io resources: @@ -11938,25 +11900,6 @@ webhooks: resources: - storagepools sideEffects: None -- admissionReviewVersions: - - v1 - clientConfig: - service: - name: simplyblock-operator-webhook-service - namespace: simplyblock-operator-system - path: /validate-replication-storage-openshift-io-v1alpha1-volumegroupreplication - failurePolicy: Fail - name: vvolumegroupreplication.simplyblock.io - rules: - - apiGroups: - - replication.storage.openshift.io - apiVersions: - - v1alpha1 - operations: - - CREATE - resources: - - volumegroupreplications - sideEffects: None - admissionReviewVersions: - v1 clientConfig: diff --git a/shared/openapi.json b/shared/openapi.json index 12eff74ef..7a8da607c 100644 --- a/shared/openapi.json +++ b/shared/openapi.json @@ -4579,7 +4579,7 @@ "replication" ], "summary": "Clusters:Storage-Pools:Volumes:Replication:Failover", - "description": "Bring the volume up on the target cluster.\n\nThe counterpart's id is read back from this volume's replication\nrelationship, its connection paths from the target volume's `connect`.\n\n``generation`` selects WHICH retained point-in-time to come up on: 0 (the\ndefault) is the newest, 1 the one before it, and so on through the\nhistory a retention schedule keeps. Failing over to an older generation\nis the recovery path for a logical corruption, which the newest copy has\nfaithfully replicated.\n\n``planned=True`` gates on a completed demote (P0-3) so a planned swap\nloses nothing: 412 when no demote was ever requested for this volume (the\ncaller's premise that the source is reachable to demote was wrong, and a\n412 is what lets the csi-addons controller's own force-escalation take\nover), 409 while demote is still converging (retryable -- 409 must never\nbecome a code the controller reads as permission to force, since that\ncontroller escalates on ANY FAILED_PRECONDITION from a force=false\npromote with no wait-and-retry grace period of its own). Unplanned\nfailover (the default) ignores demote state entirely, unchanged from\ntoday: its whole premise is that the source may never have been\nreachable to demote.", + "description": "Bring the volume up on the target cluster.\n\nThe counterpart's id is read back from this volume's replication\nrelationship, its connection paths from the target volume's `connect`.\n\n``generation`` selects WHICH retained point-in-time to come up on: 0 (the\ndefault) is the newest, 1 the one before it, and so on through the\nhistory a retention schedule keeps. Failing over to an older generation\nis the recovery path for a logical corruption, which the newest copy has\nfaithfully replicated.\n\n``planned=True`` gates on a completed demote (P0-3) so a planned swap\nloses nothing: 409 while demote is still converging (retryable -- 409\nmust never become a code the controller reads as permission to force,\nsince that controller escalates on ANY FAILED_PRECONDITION from a\nforce=false promote with no wait-and-retry grace period of its own).\nWhen no demote was ever requested, the source's own health decides: a\ngenuinely healthy, still-serving source means there is nothing to fail\nover -- this is the vendored csi-addons controller's OWN first-ever\nreconcile of a `VolumeReplication` that already lives here, not a\ndisaster, and this call succeeds as the no-op it is. A source that is\nNOT healthy gets 412, the caller's premise that it was reachable to\ndemote was wrong, and 412 is what lets the controller's own\nforce-escalation take over. Unplanned failover (the default) ignores\ndemote state entirely, unchanged from today: its whole premise is that\nthe source may never have been reachable to demote.", "operationId": "clusters_storage_pools_volumes_replication_failover_api_v2_clusters__cluster_id__storage_pools__pool_id__volumes__volume_id__replication_failover_post", "security": [ { @@ -7707,6 +7707,305 @@ } } }, + "/api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication": { + "put": { + "tags": [ + "consistency-groups" + ], + "summary": "Clusters:Consistency-Groups:Replication:Configure", + "description": "Enable or disable group replication (design-csi-addons-replication.md\n\u00a714.4): a policy id attaches the whole group to that group replication\npolicy; ``null`` detaches it (the group and its members stay grouped by\nlabel). A refused attach (not a consistency-group policy, missing policy) is\na 409.", + "operationId": "clusters_consistency_groups_replication_configure_api_v2_clusters__cluster_id__consistency_groups__group_id__replication_put", + "security": [ + { + "HTTPBearer": [] + } + ], + "parameters": [ + { + "name": "cluster_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Cluster Id" + } + }, + { + "name": "group_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Group Id" + } + } + ], + "requestBody": { + "required": true, + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/ConsistencyGroupReplicationIntentDTO" + } + } + } + }, + "responses": { + "204": { + "description": "Successful Response" + }, + "422": { + "description": "Validation Error", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/HTTPValidationError" + } + } + } + } + } + } + }, + "/api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/status": { + "get": { + "tags": [ + "consistency-groups" + ], + "summary": "Clusters:Consistency-Groups:Replication:Status", + "description": "The group's replication status as one unit: oldest recovery point, worst\nmember lag and health, summed backlog (design-csi-addons-replication.md\n\u00a714.4/\u00a714.6). Never 404s -- a group with no replicating member reports\n``state: not_replicating``.", + "operationId": "clusters_consistency_groups_replication_status_api_v2_clusters__cluster_id__consistency_groups__group_id__replication_status_get", + "security": [ + { + "HTTPBearer": [] + } + ], + "parameters": [ + { + "name": "cluster_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Cluster Id" + } + }, + { + "name": "group_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Group Id" + } + } + ], + "responses": { + "200": { + "description": "Successful Response", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/ConsistencyGroupReplicationStatusDTO" + } + } + } + }, + "422": { + "description": "Validation Error", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/HTTPValidationError" + } + } + } + } + } + } + }, + "/api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/failover": { + "post": { + "tags": [ + "consistency-groups" + ], + "summary": "Clusters:Consistency-Groups:Replication:Failover", + "description": "Fail the whole group over as ONE unit through its replication policy\n(design-csi-addons-replication.md \u00a714.4): every member is pinned to the same\ngroup generation, all-or-nothing. Refuses (412) a group not attached to a\npolicy.", + "operationId": "clusters_consistency_groups_replication_failover_api_v2_clusters__cluster_id__consistency_groups__group_id__replication_failover_post", + "security": [ + { + "HTTPBearer": [] + } + ], + "parameters": [ + { + "name": "cluster_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Cluster Id" + } + }, + { + "name": "group_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Group Id" + } + } + ], + "responses": { + "200": { + "description": "Successful Response", + "content": { + "application/json": { + "schema": { + "type": "object", + "additionalProperties": true, + "title": "Response Clusters Consistency Groups Replication Failover Api V2 Clusters Cluster Id Consistency Groups Group Id Replication Failover Post" + } + } + } + }, + "422": { + "description": "Validation Error", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/HTTPValidationError" + } + } + } + } + } + } + }, + "/api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/demote": { + "post": { + "tags": [ + "consistency-groups" + ], + "summary": "Clusters:Consistency-Groups:Replication:Demote", + "description": "Demote the whole group: fence every member and confirm each one's last\nwrite replicated (design-csi-addons-replication.md \u00a714.4). Re-drivable, not\nqueued: 204 once every member is demoted, 202 (with per-member detail) while\nany is still converging, 500 on a hard failure.", + "operationId": "clusters_consistency_groups_replication_demote_api_v2_clusters__cluster_id__consistency_groups__group_id__replication_demote_post", + "security": [ + { + "HTTPBearer": [] + } + ], + "parameters": [ + { + "name": "cluster_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Cluster Id" + } + }, + { + "name": "group_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Group Id" + } + } + ], + "responses": { + "204": { + "description": "Successful Response" + }, + "202": { + "description": "Accepted" + }, + "422": { + "description": "Validation Error", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/HTTPValidationError" + } + } + } + } + } + } + }, + "/api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/failback": { + "post": { + "tags": [ + "consistency-groups" + ], + "summary": "Clusters:Consistency-Groups:Replication:Failback", + "description": "Fail the whole group back: point every member's replication back at the\nsource cluster (design-csi-addons-replication.md \u00a714.4). The cutover itself is\neach member's own commit.", + "operationId": "clusters_consistency_groups_replication_failback_api_v2_clusters__cluster_id__consistency_groups__group_id__replication_failback_post", + "security": [ + { + "HTTPBearer": [] + } + ], + "parameters": [ + { + "name": "cluster_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Cluster Id" + } + }, + { + "name": "group_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Group Id" + } + } + ], + "requestBody": { + "required": true, + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/GroupFailbackParams" + } + } + } + }, + "responses": { + "204": { + "description": "Successful Response" + }, + "422": { + "description": "Validation Error", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/HTTPValidationError" + } + } + } + } + } + } + }, "/api/v2/management-nodes/": { "get": { "summary": "Management Nodes:List", @@ -8913,6 +9212,108 @@ "title": "ConsistencyGroupMemberJoinDTO", "description": "Request body for the late join of an existing volume (design \u00a74.5)." }, + "ConsistencyGroupReplicationIntentDTO": { + "properties": { + "replication_policy_id": { + "anyOf": [ + { + "type": "string", + "format": "uuid" + }, + { + "type": "null" + } + ], + "title": "Replication Policy Id" + } + }, + "type": "object", + "required": [ + "replication_policy_id" + ], + "title": "ConsistencyGroupReplicationIntentDTO", + "description": "Request body to enable or disable group replication (design \u00a714.4).\n\nA UUID attaches the whole consistency group to that group replication policy;\nan explicit ``null`` detaches it (the group and its members stay grouped by\nlabel, only replication stops). The field is required, so omitting it is a\n422 rather than an ambiguous no-op." + }, + "ConsistencyGroupReplicationStatusDTO": { + "properties": { + "role": { + "type": "string", + "enum": [ + "source", + "secondary", + "failed_over", + "none" + ], + "title": "Role" + }, + "state": { + "type": "string", + "enum": [ + "in_sync", + "replicating", + "lagging", + "degraded", + "error", + "not_replicating" + ], + "title": "State" + }, + "member_count": { + "type": "integer", + "minimum": 0.0, + "title": "Member Count", + "default": 0 + }, + "last_replicated_at": { + "anyOf": [ + { + "type": "string", + "format": "date-time" + }, + { + "type": "null" + } + ], + "title": "Last Replicated At" + }, + "lag_seconds": { + "anyOf": [ + { + "type": "integer", + "minimum": 0.0 + }, + { + "type": "null" + } + ], + "title": "Lag Seconds" + }, + "outstanding_count": { + "type": "integer", + "minimum": 0.0, + "title": "Outstanding Count", + "default": 0 + }, + "outstanding_bytes": { + "type": "integer", + "minimum": 0.0, + "title": "Outstanding Bytes", + "default": 0 + }, + "resyncing": { + "type": "boolean", + "title": "Resyncing", + "default": false + } + }, + "type": "object", + "required": [ + "role", + "state" + ], + "title": "ConsistencyGroupReplicationStatusDTO", + "description": "The replication status of a consistency group as one unit.\n\nA group's recovery point is its OLDEST member's, its lag and health its\nWORST member's, and its backlog the sum, because a group is only as\nprotected as its slowest, sickest member. Never a 404, matching the\nper-volume ``ReplicationStatusDTO`` (design-csi-addons-replication.md\n\u00a714.4/\u00a714.6)." + }, "DeviceDTO": { "properties": { "id": { @@ -9227,6 +9628,24 @@ ], "title": "FailoverResultDTO" }, + "GroupFailbackParams": { + "properties": { + "source_cluster_id": { + "anyOf": [ + { + "type": "string", + "format": "uuid" + }, + { + "type": "null" + } + ], + "title": "Source Cluster Id" + } + }, + "type": "object", + "title": "GroupFailbackParams" + }, "HTTPValidationError": { "properties": { "detail": { From 31ddf0e125e0f456e89bfd3ccc8a83052bb1ebfc Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Sat, 26 Sep 2026 23:41:59 +0100 Subject: [PATCH 159/206] csi-driver: route group Replication verbs from the volumegroup source oneof --- csi-driver/internal/csi/controller/replication.go | 11 +++++++++++ .../internal/csi/controller/replication_group_test.go | 10 ++++++++-- 2 files changed, 19 insertions(+), 2 deletions(-) diff --git a/csi-driver/internal/csi/controller/replication.go b/csi-driver/internal/csi/controller/replication.go index 88ddea84d..766ea7359 100644 --- a/csi-driver/internal/csi/controller/replication.go +++ b/csi-driver/internal/csi/controller/replication.go @@ -48,10 +48,21 @@ type volumeIDCarrier interface { // proxies every Replication RPC through ReplicationSource and never sets the // legacy flat VolumeId, so that field is checked first; the flat field is // kept as a fallback for any caller that still sends it. +// +// A VolumeGroupReplication drives the SAME Replication verbs, but the sidecar +// sets the group source oneof (ReplicationSource.volumegroup.volume_group_id) +// rather than the per-volume one, carrying the "cg:{cluster}:{group}" handle. +// Reading only the volume oneof left the handle empty and every group verb +// failed with `invalid volume handle ""` (VGR promote, confirmed live +// 2026-09-26). The group handle is returned as-is so ParseGroupHandle routes it +// to the group-replication path. func volumeIDFrom(req volumeIDCarrier) string { if v := req.GetReplicationSource().GetVolume().GetVolumeId(); v != "" { return v } + if g := req.GetReplicationSource().GetVolumegroup().GetVolumeGroupId(); g != "" { + return g + } return req.GetVolumeId() } diff --git a/csi-driver/internal/csi/controller/replication_group_test.go b/csi-driver/internal/csi/controller/replication_group_test.go index 94e6cd950..f4bb30610 100644 --- a/csi-driver/internal/csi/controller/replication_group_test.go +++ b/csi-driver/internal/csi/controller/replication_group_test.go @@ -18,10 +18,16 @@ const ( vgPolicyID = "dddddddd-dddd-4ddd-8ddd-dddddddddddd" ) +// groupSource builds the ReplicationSource the csi-addons sidecar actually sends +// when driving a VolumeGroupReplication: the group handle rides the volumegroup +// oneof, NOT the per-volume one. Using the volume oneof here (as this helper +// once did) exercised the same wrong field the code read, so every group-routing +// test passed while live VGR promote failed with `invalid volume handle ""` +// (2026-09-26). Regression: 2026-09-26-vgr-source-oneof. func groupSource(handle string) *replication.ReplicationSource { return &replication.ReplicationSource{ - Type: &replication.ReplicationSource_Volume{ - Volume: &replication.ReplicationSource_VolumeSource{VolumeId: handle}, + Type: &replication.ReplicationSource_Volumegroup{ + Volumegroup: &replication.ReplicationSource_VolumeGroupSource{VolumeGroupId: handle}, }, } } From 6652adb4df5f475972b2394d02b340a54314f61e Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Sun, 27 Sep 2026 04:03:20 +0100 Subject: [PATCH 160/206] csi-driver: treat an empty replicationPolicyID as a no-op enable --- .../internal/csi/controller/replication.go | 14 +++++++++++-- .../csi/controller/replication_group_test.go | 20 +++++++++++++++++++ .../csi/controller/replication_test.go | 20 ++++++++++++++----- 3 files changed, 47 insertions(+), 7 deletions(-) diff --git a/csi-driver/internal/csi/controller/replication.go b/csi-driver/internal/csi/controller/replication.go index 766ea7359..242763056 100644 --- a/csi-driver/internal/csi/controller/replication.go +++ b/csi-driver/internal/csi/controller/replication.go @@ -138,9 +138,19 @@ func (cs *Server) EnableVolumeReplication( req *replication.EnableVolumeReplicationRequest, ) (*replication.EnableVolumeReplicationResponse, error) { policyID := req.GetParameters()[replicationPolicyParam] + // An empty replicationPolicyID is the FAIL-OVER TARGET: the side becoming + // primary carries no reverse-direction policy yet, because the reverse + // direction is a fail-back-time concern (design-ramen-integration.md §6.2) -- + // nothing on that side replicates until it becomes primary. csi-addons always + // calls Enable before Promote for whichever side is becoming Primary, so this + // Enable is a legitimate no-op there: there is nothing to attach, and + // Promote/PromoteGroup is what clones the replicated snapshot (and, for a + // group, reconstitutes it) and does the real work. Rejecting it as a hard + // "required" error blocked every fail-over whose target class had no policy, + // both per-volume and group. Mirrors the per-volume ErrNotFound tolerance + // below (Enable handed a handle that names nothing yet). if policyID == "" { - return nil, status.Errorf(codes.InvalidArgument, - "VolumeReplicationClass parameter %q is required", replicationPolicyParam) + return &replication.EnableVolumeReplicationResponse{}, nil } // A group handle drives the whole consistency group as one unit through the // group-replication endpoints (design §14.4); a per-volume handle takes the diff --git a/csi-driver/internal/csi/controller/replication_group_test.go b/csi-driver/internal/csi/controller/replication_group_test.go index f4bb30610..69805a009 100644 --- a/csi-driver/internal/csi/controller/replication_group_test.go +++ b/csi-driver/internal/csi/controller/replication_group_test.go @@ -55,6 +55,26 @@ func TestEnableVolumeReplicationRoutesAGroupHandle(t *testing.T) { } } +// The group fail-over target has no reverse policy on its VolumeGroupReplication +// class either, so Enable on the cg: handle must be a no-op there; PromoteGroup +// clones and reconstitutes the group. Regression: 2026-09-27-failover-empty-policy. +func TestEnableVolumeReplicationEmptyPolicyIsNoOpForGroup(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newGroupReplTestServer(t, mock) + + _, err := cs.EnableVolumeReplication(context.Background(), &replication.EnableVolumeReplicationRequest{ + ReplicationSource: groupSource(vgGroupHandle), + Parameters: map[string]string{}, + }) + if err != nil { + t.Fatalf("empty policy should be a no-op on the group fail-over target, got: %v", err) + } + if got := mock.groups[vgGroupID].PolicyID; got != "" { + t.Fatalf("group policy = %q, want empty (nothing attached when no policy)", got) + } +} + func TestDisableVolumeReplicationRoutesAGroupHandle(t *testing.T) { mock := newMockSBCLI() defer mock.Close() diff --git a/csi-driver/internal/csi/controller/replication_test.go b/csi-driver/internal/csi/controller/replication_test.go index 34ae4d881..433ceca5a 100644 --- a/csi-driver/internal/csi/controller/replication_test.go +++ b/csi-driver/internal/csi/controller/replication_test.go @@ -131,7 +131,7 @@ func TestEnableVolumeReplicationResolvesToTargetWhenGivenTheSourceSideOfARelatio } // simplyblock's replication is one-way and the destination never carries a -// persistent, independently-provisioned LVol of its own (confirmed live +// persistent, independently provisioned LVol of its own (confirmed live // 2026-09-23, relocate M-02): the writable clone only comes into existence // when PromoteVolume clones the last replicated snapshot. csi-addons always // calls EnableVolumeReplication before PromoteVolume, unconditionally, for @@ -157,7 +157,15 @@ func TestEnableVolumeReplicationNoOpsWhenVolumeDoesNotExistYet(t *testing.T) { } } -func TestEnableVolumeReplicationMissingPolicyParam(t *testing.T) { +// An empty replicationPolicyID is the FAIL-OVER TARGET: the side becoming +// primary carries no reverse-direction policy yet (the reverse direction is a +// fail-back-time concern, design-ramen-integration.md §6.2). csi-addons always +// calls Enable before Promote for the side becoming Primary, so Enable must be a +// no-op here -- there is nothing to attach, and PromoteVolume clones from the +// replicated snapshot and does the real work. Rejecting it as "required" blocked +// every fail-over whose target VolumeReplicationClass had no policy. +// Regression: 2026-09-27-failover-empty-policy. +func TestEnableVolumeReplicationEmptyPolicyIsNoOp(t *testing.T) { mock := newMockSBCLI() defer mock.Close() cs := newReplicationTestServer(t, mock) @@ -165,9 +173,11 @@ func TestEnableVolumeReplicationMissingPolicyParam(t *testing.T) { _, err := cs.EnableVolumeReplication(context.Background(), &replication.EnableVolumeReplicationRequest{ VolumeId: testReplVolID, }) - st, _ := status.FromError(err) - if st.Code() != codes.InvalidArgument { - t.Errorf("code = %v, want InvalidArgument", st.Code()) + if err != nil { + t.Fatalf("empty policy should be a no-op on the fail-over target, got: %v", err) + } + if got := mock.volumes[testReplVolumeID].ReplicationPolicyID; got != "" { + t.Errorf("ReplicationPolicyID = %q, want empty (nothing attached when no policy)", got) } } From c6fc14664d33523e4058dcb0ffc2f2645c31636f Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Tue, 29 Sep 2026 08:55:08 +0100 Subject: [PATCH 161/206] csi-driver: fix golangci-lint lll and unparam in group replication tests --- .../csi/controller/replication_group_test.go | 20 +++++++++---------- .../csi/controller/volumegroup_test.go | 5 +++-- 2 files changed, 13 insertions(+), 12 deletions(-) diff --git a/csi-driver/internal/csi/controller/replication_group_test.go b/csi-driver/internal/csi/controller/replication_group_test.go index 69805a009..3414b6d47 100644 --- a/csi-driver/internal/csi/controller/replication_group_test.go +++ b/csi-driver/internal/csi/controller/replication_group_test.go @@ -24,10 +24,10 @@ const ( // once did) exercised the same wrong field the code read, so every group-routing // test passed while live VGR promote failed with `invalid volume handle ""` // (2026-09-26). Regression: 2026-09-26-vgr-source-oneof. -func groupSource(handle string) *replication.ReplicationSource { +func groupSource() *replication.ReplicationSource { return &replication.ReplicationSource{ Type: &replication.ReplicationSource_Volumegroup{ - Volumegroup: &replication.ReplicationSource_VolumeGroupSource{VolumeGroupId: handle}, + Volumegroup: &replication.ReplicationSource_VolumeGroupSource{VolumeGroupId: vgGroupHandle}, }, } } @@ -44,7 +44,7 @@ func TestEnableVolumeReplicationRoutesAGroupHandle(t *testing.T) { cs := newGroupReplTestServer(t, mock) _, err := cs.EnableVolumeReplication(context.Background(), &replication.EnableVolumeReplicationRequest{ - ReplicationSource: groupSource(vgGroupHandle), + ReplicationSource: groupSource(), Parameters: map[string]string{replicationPolicyParam: vgPolicyID}, }) if err != nil { @@ -64,7 +64,7 @@ func TestEnableVolumeReplicationEmptyPolicyIsNoOpForGroup(t *testing.T) { cs := newGroupReplTestServer(t, mock) _, err := cs.EnableVolumeReplication(context.Background(), &replication.EnableVolumeReplicationRequest{ - ReplicationSource: groupSource(vgGroupHandle), + ReplicationSource: groupSource(), Parameters: map[string]string{}, }) if err != nil { @@ -82,7 +82,7 @@ func TestDisableVolumeReplicationRoutesAGroupHandle(t *testing.T) { mock.groups[vgGroupID].PolicyID = vgPolicyID _, err := cs.DisableVolumeReplication(context.Background(), &replication.DisableVolumeReplicationRequest{ - ReplicationSource: groupSource(vgGroupHandle), + ReplicationSource: groupSource(), }) if err != nil { t.Fatalf("DisableVolumeReplication: %v", err) @@ -98,7 +98,7 @@ func TestPromoteVolumeRoutesAGroupHandle(t *testing.T) { cs := newGroupReplTestServer(t, mock) if _, err := cs.PromoteVolume(context.Background(), &replication.PromoteVolumeRequest{ - ReplicationSource: groupSource(vgGroupHandle), Force: true, + ReplicationSource: groupSource(), Force: true, }); err != nil { t.Fatalf("PromoteVolume: %v", err) } @@ -113,7 +113,7 @@ func TestDemoteVolumeRoutesAGroupHandle(t *testing.T) { cs := newGroupReplTestServer(t, mock) if _, err := cs.DemoteVolume(context.Background(), &replication.DemoteVolumeRequest{ - ReplicationSource: groupSource(vgGroupHandle), + ReplicationSource: groupSource(), }); err != nil { t.Fatalf("DemoteVolume: %v", err) } @@ -129,7 +129,7 @@ func TestDemoteVolumeGroupStillConvergingIsAborted(t *testing.T) { mock.groups[vgGroupID].DemoteConverging = true _, err := cs.DemoteVolume(context.Background(), &replication.DemoteVolumeRequest{ - ReplicationSource: groupSource(vgGroupHandle), + ReplicationSource: groupSource(), }) if status.Code(err) != codes.Aborted { t.Fatalf("err = %v, want Aborted", err) @@ -142,7 +142,7 @@ func TestResyncVolumeRoutesAGroupHandle(t *testing.T) { cs := newGroupReplTestServer(t, mock) resp, err := cs.ResyncVolume(context.Background(), &replication.ResyncVolumeRequest{ - ReplicationSource: groupSource(vgGroupHandle), + ReplicationSource: groupSource(), Parameters: map[string]string{sourceClusterIDParam: sanityClusterID}, }) if err != nil { @@ -163,7 +163,7 @@ func TestGetVolumeReplicationInfoRoutesAGroupHandle(t *testing.T) { mock.groups[vgGroupID].LastReplicatedAt = 1_700_000_000 resp, err := cs.GetVolumeReplicationInfo(context.Background(), &replication.GetVolumeReplicationInfoRequest{ - ReplicationSource: groupSource(vgGroupHandle), + ReplicationSource: groupSource(), }) if err != nil { t.Fatalf("GetVolumeReplicationInfo: %v", err) diff --git a/csi-driver/internal/csi/controller/volumegroup_test.go b/csi-driver/internal/csi/controller/volumegroup_test.go index 788d7a6b4..095aa2206 100644 --- a/csi-driver/internal/csi/controller/volumegroup_test.go +++ b/csi-driver/internal/csi/controller/volumegroup_test.go @@ -118,8 +118,9 @@ func TestVolumeGroupVerbsRejectANonGroupHandle(t *testing.T) { cs := newTestControllerServer(t, mock) perVolume := vgHandle(vgMember1) // a per-volume handle, not a group handle - if _, err := cs.ModifyVolumeGroupMembership(context.Background(), - &volumegroup.ModifyVolumeGroupMembershipRequest{VolumeGroupId: perVolume}); status.Code(err) != codes.InvalidArgument { + _, err := cs.ModifyVolumeGroupMembership(context.Background(), + &volumegroup.ModifyVolumeGroupMembershipRequest{VolumeGroupId: perVolume}) + if status.Code(err) != codes.InvalidArgument { t.Errorf("Modify: err = %v, want InvalidArgument", err) } if _, err := cs.DeleteVolumeGroup(context.Background(), From ac972619c190e1db2a34b2b753555d5141fa431d Mon Sep 17 00:00:00 2001 From: Geoffrey Israel Date: Thu, 1 Oct 2026 09:25:16 +0100 Subject: [PATCH 162/206] docs(operator): design non-disruptive test failover (TestFailover CRD) (#583) * docs(operator): design non-disruptive test failover (TestFailover CRD) * docs(operator): updated design non-disruptive test failover (TestFailover CRD) * feat(operator): TestFailover CRD and controller for non-disruptive test failover * ran make operator-build-installer * fixed linter issue * fix(testfailover): stage the bubble clone with a real VolumeContext * ran make operator-build-installer * fix(testfailover): stage the bubble clone with a complete PV * feat(testfailover): recover a consistency group (scope=Group) * fix linter issue * fix(testfailover): resolve group source UUID and member PVCs correctly * fix(testfailover): read group policy from the group, not a placement heuristic * fix(testfailover): recover in-place CG drills from a fresh source generation, not the DR target * reverted test failover within the same cluster * fixed failing operator manifest CI --- csi-driver/internal/csi/node/stack_test.go | 33 + csi-driver/internal/csi/node/stage.go | 7 + .../storage.simplyblock.io_testfailovers.yaml | 305 +++++ .../templates/roles/manager_role.yaml | 27 + operator/api/v1alpha2/testfailover_types.go | 291 ++++ .../api/v1alpha2/zz_generated.deepcopy.go | 160 +++ operator/cmd/main.go | 12 + .../storage.simplyblock.io_testfailovers.yaml | 305 +++++ operator/config/crd/kustomization.yaml | 1 + ...yblock-operator.clusterserviceversion.yaml | 8 + operator/config/rbac/role.yaml | 27 + operator/dist/install.yaml | 332 +++++ operator/docs/designs/design-test-failover.md | 612 +++++++++ .../docs/tests/test-plan-test-failover.md | 280 ++++ operator/go.mod | 1 + operator/go.sum | 5 +- .../controller/testfailover_controller.go | 1217 +++++++++++++++++ .../testfailover_controller_unit_test.go | 1195 ++++++++++++++++ .../storage.simplyblock.io_testfailovers.yaml | 305 +++++ operator/internal/webapi/consistency_group.go | 9 +- operator/internal/webapi/group_failover.go | 171 +++ .../internal/webapi/group_failover_test.go | 99 ++ 22 files changed, 5400 insertions(+), 2 deletions(-) create mode 100644 helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_testfailovers.yaml create mode 100644 operator/api/v1alpha2/testfailover_types.go create mode 100644 operator/config/crd/bases/storage.simplyblock.io_testfailovers.yaml create mode 100644 operator/docs/designs/design-test-failover.md create mode 100644 operator/docs/tests/test-plan-test-failover.md create mode 100644 operator/internal/controller/testfailover_controller.go create mode 100644 operator/internal/controller/testfailover_controller_unit_test.go create mode 100644 operator/internal/upgrade/crds/manifests/storage.simplyblock.io_testfailovers.yaml create mode 100644 operator/internal/webapi/group_failover.go create mode 100644 operator/internal/webapi/group_failover_test.go diff --git a/csi-driver/internal/csi/node/stack_test.go b/csi-driver/internal/csi/node/stack_test.go index 914ed8321..f234306ee 100644 --- a/csi-driver/internal/csi/node/stack_test.go +++ b/csi-driver/internal/csi/node/stack_test.go @@ -18,6 +18,8 @@ import ( "testing" "github.com/container-storage-interface/spec/lib/go/csi" + "google.golang.org/grpc/codes" + "google.golang.org/grpc/status" corev1 "k8s.io/api/core/v1" k8smount "k8s.io/mount-utils" @@ -285,6 +287,37 @@ func TestStageBringsTheStackUp(t *testing.T) { } } +// Regression: 2026-09-29-testfailover-nil-volumecontext — a statically +// provisioned PV (the TestFailover bubble PV) carries no csi.volumeAttributes, so +// NodeStageVolume received a nil VolumeContext and panicked with "assignment to +// entry in nil map" at the first vc[...] write. The node plugin crash-looped, its +// socket refused connections, and no pod could mount the clone. Staging must +// tolerate a nil VolumeContext: the volume's identity is re-resolved from its +// handle regardless. +func TestStageToleratesNilVolumeContext(t *testing.T) { + runner := newRecordingRunner() + ns, _ := newStackedServer(t, runner) + + req := &csi.NodeStageVolumeRequest{ + VolumeId: pvcTestHandle, + StagingTargetPath: t.TempDir(), + VolumeCapability: mountCapability(), + VolumeContext: nil, // a static PV with no volumeAttributes + } + + // Before the fix this panicked on the nil map at the first vc[...] write. It + // must return instead. With no control plane in this unit harness to resolve + // the volume's identity from its handle, staging fails cleanly (fail-safe) + // rather than crashing the node plugin or attaching the wrong target. + _, err := ns.NodeStageVolume(context.Background(), req) + if status.Code(err) != codes.Internal { + t.Fatalf("want a clean Internal error for an unresolvable nil-context stage, got %v", err) + } + if runner.called("up") { + t.Fatalf("stage attached a target with no resolved subsystem: %v", runner.calls) + } +} + // An unstage releases and never destroys. It fires whenever no pod on this node // needs the volume mounted, which includes an ordinary pod restart, and // conflating the two verbs is what once ran vgremove over a volume holding diff --git a/csi-driver/internal/csi/node/stage.go b/csi-driver/internal/csi/node/stage.go index 0b52f0d20..807ed1250 100644 --- a/csi-driver/internal/csi/node/stage.go +++ b/csi-driver/internal/csi/node/stage.go @@ -80,6 +80,13 @@ func (ns *Server) NodeStageVolume( } vc := req.GetVolumeContext() + if vc == nil { + // A statically provisioned PV can carry no csi.volumeAttributes, which + // arrives here as a nil map. The volume's identity is re-resolved from its + // handle by refreshVolumeContext regardless, so an empty context is enough + // to stage; a nil one would panic on the first write below. + vc = map[string]string{} + } vc["stagingParentPath"] = stagingParentPath ns.refreshVolumeContext(ctx, volumeID, vc) diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_testfailovers.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_testfailovers.yaml new file mode 100644 index 000000000..d68cf6ad7 --- /dev/null +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_testfailovers.yaml @@ -0,0 +1,305 @@ +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition +metadata: + annotations: + controller-gen.kubebuilder.io/version: v0.21.0 + name: testfailovers.storage.simplyblock.io +spec: + group: storage.simplyblock.io + names: + kind: TestFailover + listKind: TestFailoverList + plural: testfailovers + shortNames: + - tfo + singular: testfailover + scope: Namespaced + versions: + - additionalPrinterColumns: + - jsonPath: .spec.scope + name: Scope + type: string + - jsonPath: .spec.sourceRef + name: Source + type: string + - jsonPath: .spec.sourceCluster + name: "On" + type: string + - jsonPath: .spec.bubbleCluster + name: Bubble + type: string + - jsonPath: .status.phase + name: Phase + type: string + - jsonPath: .status.step.state + name: Step + type: string + - jsonPath: .status.message + name: Message + priority: 1 + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha2 + schema: + openAPIV3Schema: + description: |- + TestFailover is a one-way, non-disruptive test-failover drill. It recovers a + source volume, or a consistency group, from a snapshot into an isolated + namespace on a chosen cluster as bound PVCs, without touching the source. The + hub reads the source on its cluster and places the bubble on the recovery + cluster through OCM. It runs to a terminal phase, or holds Ready until it is + deleted, and deletion reclaims the clones and any snapshots the drill took. + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: |- + TestFailoverSpec is the request for one non-disruptive test-failover drill. + + The source is named by where it runs and what it is, so the hub can find it + without anyone extracting a backend handle by hand. SourceNamespace is + required for a Volume drill, where the source is a PVC, and unused for a Group + drill, where SourceRef names a consistency group. + properties: + bubbleCluster: + description: |- + BubbleCluster is the OCM ManagedCluster to recover onto: a DR target holding + the replicated point, or another cluster. It must differ from SourceCluster; + test-failover recovers onto a different cluster, never in place. Immutable. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + bubbleNamespace: + default: bubble + description: |- + BubbleNamespace is the namespace on the bubble cluster where the recovered + PVCs are created. Immutable. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + scope: + description: Scope selects what the drill recovers. Immutable. + enum: + - Volume + - Group + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + sourceCluster: + description: |- + SourceCluster is the OCM ManagedCluster the source runs on. The hub reads + the source there through a ManagedClusterView. Immutable. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + sourceNamespace: + description: |- + SourceNamespace is the namespace of the source PVC on SourceCluster. + Required for scope=Volume. Immutable. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + sourceRef: + description: |- + SourceRef names the source on SourceCluster: a PersistentVolumeClaim in + SourceNamespace (scope=Volume), or a consistency group (scope=Group). + Immutable. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + ttlSeconds: + description: |- + TTLSeconds is an optional maximum lifetime: the drill is torn down after it + even without a delete, so a forgotten drill cannot hold a clone forever. + format: int64 + minimum: 0 + type: integer + required: + - bubbleCluster + - scope + - sourceCluster + - sourceRef + type: object + x-kubernetes-validations: + - message: field sourceNamespace is immutable once set + rule: '!has(oldSelf.sourceNamespace) || has(self.sourceNamespace)' + - message: field bubbleNamespace is immutable once set + rule: '!has(oldSelf.bubbleNamespace) || has(self.bubbleNamespace)' + - message: sourceNamespace is required for scope=Volume + rule: self.scope != 'Volume' || has(self.sourceNamespace) + status: + description: TestFailoverStatus is the observed state of one drill. + properties: + clones: + description: Clones is one entry per recovered volume. + items: + description: |- + TestFailoverClone is one recovered volume: the source it came from, the + snapshot and clone the drill built, and the PVC placed on the bubble cluster. + properties: + cloneID: + description: CloneID is the backend id of the writable clone. + type: string + pvcName: + description: PVCName is the bound PVC in the bubble namespace + on the bubble cluster. + type: string + sizeBytes: + description: SizeBytes is the recovered volume's size. + format: int64 + type: integer + snapshotID: + description: |- + SnapshotID is the recovery-point snapshot: the replicated snapshot already on + the bubble cluster's backend that the clone is built from. + type: string + sourceFSType: + description: |- + SourceFSType is the source PV's CSI fsType, carried onto the bubble PV so the + node plugin stages the clone with the filesystem it actually carries. The + clone is a block copy of the source, so its filesystem is the source's; an + empty fsType makes the node plugin default to ext4 and refuse to mount an XFS + volume. + type: string + sourceHandle: + description: SourceHandle is the source volume's backend handle, + read from its PV. + type: string + sourceRef: + description: |- + SourceRef is the source volume, or group member, the recovered volume maps + to. + type: string + sourceVolumeContext: + additionalProperties: + type: string + description: |- + SourceVolumeContext is the source PV's CSI volumeAttributes, minus the + identity and provisioner keys, carried onto the bubble PV so the node plugin + receives a non-nil VolumeContext when it stages the clone. The clone's own + identity (NQN, connections, nsId, and so on) is re-resolved from the clone + handle at stage time, so only the class-level parameters are carried; the + identity keys are dropped so a failed clone lookup can never point the mount + back at the source. + type: object + required: + - sourceRef + type: object + type: array + x-kubernetes-list-map-keys: + - sourceRef + x-kubernetes-list-type: map + completedAt: + description: CompletedAt is when the drill reached a terminal phase. + format: date-time + type: string + message: + description: |- + Message is the reason the phase is what it is: one sentence, replaced as the + drill moves, and never a log. + type: string + observedGeneration: + description: |- + ObservedGeneration is the generation the rest of this status was computed + from, so a stale status can be told from a current one. + format: int64 + type: integer + phase: + description: Phase is the drill's own progress. + enum: + - Pending + - Provisioning + - Ready + - Failed + - TearingDown + type: string + readyAt: + description: ReadyAt is when every recovered PVC became bound. + format: date-time + type: string + report: + description: Report is the drill's evidence, populated as it reaches + Ready. + properties: + bubbleCluster: + description: BubbleCluster is the cluster the drill recovered + onto. + type: string + invariantsHeld: + description: |- + InvariantsHeld is true only when the source fingerprint taken before the + drill matches the one taken at Ready. A Ready drill with this false is a + defect. + type: boolean + recoveryPoint: + description: RecoveryPoint is the snapshot or group generation + the drill recovered. + type: string + recoveryPointAgeSeconds: + description: RecoveryPointAgeSeconds is the drill time minus the + recovery-point time. + format: int64 + type: integer + recoveryPointTime: + description: RecoveryPointTime is when that point was taken. + format: date-time + type: string + type: object + startedAt: + description: StartedAt is when the drill started. + format: date-time + type: string + step: + description: Step is the position of the running drill's state machine. + properties: + deadline: + description: |- + Deadline is when that state expires, absent when it has none. It is an + absolute instant, so a state whose deadline passed while the controller + was down restores as already expired. + format: date-time + type: string + state: + description: |- + State is the state the machine was in. Empty means the resource has not + been reconciled yet, and restores to the graph's initial state. + type: string + type: object + x-kubernetes-validations: + - message: unknown step + rule: '!has(self.state) || self.state in [''ResolvingSource'',''ResolvingPoint'',''Shipping'',''Cloning'',''Placing'',''Releasing'']' + triggered: + description: |- + Triggered records that the current step's side effect was issued, so a + restart does not repeat it. + type: boolean + type: object + type: object + served: true + storage: true + subresources: + status: {} diff --git a/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml b/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml index 2b1b2b541..5c72b25d6 100644 --- a/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml @@ -325,6 +325,7 @@ rules: - storagepoolops - storagepools - tasks + - testfailovers - volumemigrations verbs: - create @@ -360,6 +361,7 @@ rules: - storagepoolops/finalizers - storagepools/finalizers - tasks/finalizers + - testfailovers/finalizers - volumemigrations/finalizers verbs: - update @@ -392,6 +394,7 @@ rules: - storagepoolops/status - storagepools/status - tasks/status + - testfailovers/status - volumegroupsnapshotops/status - volumemigrations/status verbs: @@ -418,3 +421,27 @@ rules: - get - list - watch +- apiGroups: + - view.open-cluster-management.io + resources: + - managedclusterviews + verbs: + - create + - delete + - get + - list + - patch + - update + - watch +- apiGroups: + - work.open-cluster-management.io + resources: + - manifestworks + verbs: + - create + - delete + - get + - list + - patch + - update + - watch diff --git a/operator/api/v1alpha2/testfailover_types.go b/operator/api/v1alpha2/testfailover_types.go new file mode 100644 index 000000000..de033a936 --- /dev/null +++ b/operator/api/v1alpha2/testfailover_types.go @@ -0,0 +1,291 @@ +// One non-disruptive test-failover drill. +// +// A test failover proves an application can be recovered from a point-in-time +// copy, in isolation, without disturbing the running production. The recovery +// point is always a snapshot and the result is always a clone, so the source is +// never touched. The hub coordinates the drill: it reads the source on its +// cluster, resolves a recovery point on the recovery cluster's backend, clones +// it there, and places the clone as a bound PVC in an isolated namespace on the +// recovery cluster. Deleting the object reclaims the clone and any snapshot the +// drill took. +// +// Specified by operator/docs/designs/design-test-failover.md, whose Appendix A +// is this file. + +package v1alpha2 + +import ( + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + + "github.com/simplyblock/atlas/statemachine" +) + +// TestFailoverScope selects what a drill recovers. +// +kubebuilder:validation:Enum=Volume;Group +type TestFailoverScope string + +const ( + // TestFailoverScopeVolume recovers a single source volume, named by a PVC. + TestFailoverScopeVolume TestFailoverScope = "Volume" + + // TestFailoverScopeGroup recovers a consistency group from one + // group-consistent point. + TestFailoverScopeGroup TestFailoverScope = "Group" +) + +// TestFailoverPhase is the drill's own progress. +// +kubebuilder:validation:Enum=Pending;Provisioning;Ready;Failed;TearingDown +type TestFailoverPhase string + +const ( + TestFailoverPhasePending TestFailoverPhase = "Pending" + TestFailoverPhaseProvisioning TestFailoverPhase = "Provisioning" + TestFailoverPhaseReady TestFailoverPhase = "Ready" + TestFailoverPhaseFailed TestFailoverPhase = "Failed" + TestFailoverPhaseTearingDown TestFailoverPhase = "TearingDown" +) + +// TestFailoverStep is one step of a running drill. Which steps belong to which +// phase of the flow is declared by the drill's graph rather than by this type, +// which is why the enum stays flat. +// +kubebuilder:validation:Enum=ResolvingSource;ResolvingPoint;Shipping;Cloning;Placing;Releasing +type TestFailoverStep string + +const ( + // TestFailoverStepResolvingSource reads the source PVC and PV on the source + // cluster through a ManagedClusterView to learn the source volume's handle. + TestFailoverStepResolvingSource TestFailoverStep = "ResolvingSource" + + // TestFailoverStepResolvingPoint resolves the recovery point on the recovery + // cluster's backend: a fresh source snapshot, or the latest replicated one. + TestFailoverStepResolvingPoint TestFailoverStep = "ResolvingPoint" + + // TestFailoverStepShipping ships the recovery point to a backend that holds no + // copy of it (the cross-cluster, separate-backend case). + TestFailoverStepShipping TestFailoverStep = "Shipping" + + // TestFailoverStepCloning clones the recovery point into a writable volume on + // the recovery cluster's backend. + TestFailoverStepCloning TestFailoverStep = "Cloning" + + // TestFailoverStepPlacing delivers the bubble PV and PVC to the recovery + // cluster and waits for the PVC to bind. + TestFailoverStepPlacing TestFailoverStep = "Placing" + + // TestFailoverStepReleasing tears the drill down: removes the placed objects, + // reclaims the clone, and deletes any snapshot the drill took. + TestFailoverStepReleasing TestFailoverStep = "Releasing" +) + +// TestFailoverSpec is the request for one non-disruptive test-failover drill. +// +// The source is named by where it runs and what it is, so the hub can find it +// without anyone extracting a backend handle by hand. SourceNamespace is +// required for a Volume drill, where the source is a PVC, and unused for a Group +// drill, where SourceRef names a consistency group. +// +kubebuilder:validation:XValidation:rule="self.scope != 'Volume' || has(self.sourceNamespace)",message="sourceNamespace is required for scope=Volume" +type TestFailoverSpec struct { + // Scope selects what the drill recovers. Immutable. + // +kubebuilder:validation:Required + // +k8s:immutable + Scope TestFailoverScope `json:"scope"` + + // SourceCluster is the OCM ManagedCluster the source runs on. The hub reads + // the source there through a ManagedClusterView. Immutable. + // +kubebuilder:validation:Required + // +k8s:immutable + SourceCluster string `json:"sourceCluster"` + + // SourceNamespace is the namespace of the source PVC on SourceCluster. + // Required for scope=Volume. Immutable. + // +optional + // +k8s:immutable + SourceNamespace string `json:"sourceNamespace,omitempty"` + + // SourceRef names the source on SourceCluster: a PersistentVolumeClaim in + // SourceNamespace (scope=Volume), or a consistency group (scope=Group). + // Immutable. + // +kubebuilder:validation:Required + // +k8s:immutable + SourceRef string `json:"sourceRef"` + + // BubbleCluster is the OCM ManagedCluster to recover onto: a DR target holding + // the replicated point, or another cluster. It must differ from SourceCluster; + // test-failover recovers onto a different cluster, never in place. Immutable. + // +kubebuilder:validation:Required + // +k8s:immutable + BubbleCluster string `json:"bubbleCluster"` + + // BubbleNamespace is the namespace on the bubble cluster where the recovered + // PVCs are created. Immutable. + // +kubebuilder:default=bubble + // +optional + // +k8s:immutable + BubbleNamespace string `json:"bubbleNamespace,omitempty"` + + // TTLSeconds is an optional maximum lifetime: the drill is torn down after it + // even without a delete, so a forgotten drill cannot hold a clone forever. + // +kubebuilder:validation:Minimum=0 + // +optional + TTLSeconds *int64 `json:"ttlSeconds,omitempty"` +} + +// TestFailoverClone is one recovered volume: the source it came from, the +// snapshot and clone the drill built, and the PVC placed on the bubble cluster. +type TestFailoverClone struct { + // SourceRef is the source volume, or group member, the recovered volume maps + // to. + SourceRef string `json:"sourceRef"` + + // SourceHandle is the source volume's backend handle, read from its PV. + // +optional + SourceHandle string `json:"sourceHandle,omitempty"` + + // SnapshotID is the recovery-point snapshot: the replicated snapshot already on + // the bubble cluster's backend that the clone is built from. + // +optional + SnapshotID string `json:"snapshotID,omitempty"` + + // CloneID is the backend id of the writable clone. + // +optional + CloneID string `json:"cloneID,omitempty"` + + // PVCName is the bound PVC in the bubble namespace on the bubble cluster. + // +optional + PVCName string `json:"pvcName,omitempty"` + + // SizeBytes is the recovered volume's size. + // +optional + SizeBytes int64 `json:"sizeBytes,omitempty"` + + // SourceVolumeContext is the source PV's CSI volumeAttributes, minus the + // identity and provisioner keys, carried onto the bubble PV so the node plugin + // receives a non-nil VolumeContext when it stages the clone. The clone's own + // identity (NQN, connections, nsId, and so on) is re-resolved from the clone + // handle at stage time, so only the class-level parameters are carried; the + // identity keys are dropped so a failed clone lookup can never point the mount + // back at the source. + // +optional + SourceVolumeContext map[string]string `json:"sourceVolumeContext,omitempty"` + + // SourceFSType is the source PV's CSI fsType, carried onto the bubble PV so the + // node plugin stages the clone with the filesystem it actually carries. The + // clone is a block copy of the source, so its filesystem is the source's; an + // empty fsType makes the node plugin default to ext4 and refuse to mount an XFS + // volume. + // +optional + SourceFSType string `json:"sourceFSType,omitempty"` +} + +// TestFailoverReport is the evidence a drill produces. +type TestFailoverReport struct { + // BubbleCluster is the cluster the drill recovered onto. + // +optional + BubbleCluster string `json:"bubbleCluster,omitempty"` + + // RecoveryPoint is the snapshot or group generation the drill recovered. + // +optional + RecoveryPoint string `json:"recoveryPoint,omitempty"` + + // RecoveryPointTime is when that point was taken. + // +optional + RecoveryPointTime *metav1.Time `json:"recoveryPointTime,omitempty"` + + // RecoveryPointAgeSeconds is the drill time minus the recovery-point time. + // +optional + RecoveryPointAgeSeconds int64 `json:"recoveryPointAgeSeconds,omitempty"` + + // InvariantsHeld is true only when the source fingerprint taken before the + // drill matches the one taken at Ready. A Ready drill with this false is a + // defect. + // +optional + InvariantsHeld bool `json:"invariantsHeld,omitempty"` +} + +// TestFailoverStatus is the observed state of one drill. +type TestFailoverStatus struct { + // Phase is the drill's own progress. + // +optional + Phase TestFailoverPhase `json:"phase,omitempty"` + + // Step is the position of the running drill's state machine. + // +kubebuilder:validation:XValidation:rule="!has(self.state) || self.state in ['ResolvingSource','ResolvingPoint','Shipping','Cloning','Placing','Releasing']",message="unknown step" + // +optional + Step statemachine.KubeSnapshot `json:"step,omitempty"` + + // Message is the reason the phase is what it is: one sentence, replaced as the + // drill moves, and never a log. + // +optional + Message string `json:"message,omitempty"` + + // Triggered records that the current step's side effect was issued, so a + // restart does not repeat it. + // +optional + Triggered bool `json:"triggered,omitempty"` + + // ObservedGeneration is the generation the rest of this status was computed + // from, so a stale status can be told from a current one. + // +optional + ObservedGeneration int64 `json:"observedGeneration,omitempty"` + + // Clones is one entry per recovered volume. + // +optional + // +listType=map + // +listMapKey=sourceRef + Clones []TestFailoverClone `json:"clones,omitempty"` + + // Report is the drill's evidence, populated as it reaches Ready. + // +optional + Report *TestFailoverReport `json:"report,omitempty"` + + // StartedAt is when the drill started. + // +optional + StartedAt *metav1.Time `json:"startedAt,omitempty"` + + // ReadyAt is when every recovered PVC became bound. + // +optional + ReadyAt *metav1.Time `json:"readyAt,omitempty"` + + // CompletedAt is when the drill reached a terminal phase. + // +optional + CompletedAt *metav1.Time `json:"completedAt,omitempty"` +} + +// +kubebuilder:object:root=true +// +kubebuilder:subresource:status +// +kubebuilder:resource:scope=Namespaced,shortName=tfo +// +kubebuilder:printcolumn:name="Scope",type=string,JSONPath=".spec.scope" +// +kubebuilder:printcolumn:name="Source",type=string,JSONPath=".spec.sourceRef" +// +kubebuilder:printcolumn:name="On",type=string,JSONPath=".spec.sourceCluster" +// +kubebuilder:printcolumn:name="Bubble",type=string,JSONPath=".spec.bubbleCluster" +// +kubebuilder:printcolumn:name="Phase",type=string,JSONPath=".status.phase" +// +kubebuilder:printcolumn:name="Step",type=string,JSONPath=".status.step.state" +// +kubebuilder:printcolumn:name="Message",type=string,JSONPath=".status.message",priority=1 +// +kubebuilder:printcolumn:name="Age",type=date,JSONPath=".metadata.creationTimestamp" + +// TestFailover is a one-way, non-disruptive test-failover drill. It recovers a +// source volume, or a consistency group, from a snapshot into an isolated +// namespace on a chosen cluster as bound PVCs, without touching the source. The +// hub reads the source on its cluster and places the bubble on the recovery +// cluster through OCM. It runs to a terminal phase, or holds Ready until it is +// deleted, and deletion reclaims the clones and any snapshots the drill took. +type TestFailover struct { + metav1.TypeMeta `json:",inline"` + metav1.ObjectMeta `json:"metadata,omitempty"` + + Spec TestFailoverSpec `json:"spec,omitempty"` + Status TestFailoverStatus `json:"status,omitempty"` +} + +// +kubebuilder:object:root=true + +// TestFailoverList contains a list of TestFailover. +type TestFailoverList struct { + metav1.TypeMeta `json:",inline"` + metav1.ListMeta `json:"metadata,omitempty"` + Items []TestFailover `json:"items"` +} + +func init() { + SchemeBuilder.Register(&TestFailover{}, &TestFailoverList{}) +} diff --git a/operator/api/v1alpha2/zz_generated.deepcopy.go b/operator/api/v1alpha2/zz_generated.deepcopy.go index cd68ea0b4..c4659eb15 100644 --- a/operator/api/v1alpha2/zz_generated.deepcopy.go +++ b/operator/api/v1alpha2/zz_generated.deepcopy.go @@ -3579,6 +3579,166 @@ func (in *StripeSpec) DeepCopy() *StripeSpec { return out } +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *TestFailover) DeepCopyInto(out *TestFailover) { + *out = *in + out.TypeMeta = in.TypeMeta + in.ObjectMeta.DeepCopyInto(&out.ObjectMeta) + in.Spec.DeepCopyInto(&out.Spec) + in.Status.DeepCopyInto(&out.Status) +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TestFailover. +func (in *TestFailover) DeepCopy() *TestFailover { + if in == nil { + return nil + } + out := new(TestFailover) + in.DeepCopyInto(out) + return out +} + +// DeepCopyObject is an autogenerated deepcopy function, copying the receiver, creating a new runtime.Object. +func (in *TestFailover) DeepCopyObject() runtime.Object { + if c := in.DeepCopy(); c != nil { + return c + } + return nil +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *TestFailoverClone) DeepCopyInto(out *TestFailoverClone) { + *out = *in + if in.SourceVolumeContext != nil { + in, out := &in.SourceVolumeContext, &out.SourceVolumeContext + *out = make(map[string]string, len(*in)) + for key, val := range *in { + (*out)[key] = val + } + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TestFailoverClone. +func (in *TestFailoverClone) DeepCopy() *TestFailoverClone { + if in == nil { + return nil + } + out := new(TestFailoverClone) + in.DeepCopyInto(out) + return out +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *TestFailoverList) DeepCopyInto(out *TestFailoverList) { + *out = *in + out.TypeMeta = in.TypeMeta + in.ListMeta.DeepCopyInto(&out.ListMeta) + if in.Items != nil { + in, out := &in.Items, &out.Items + *out = make([]TestFailover, len(*in)) + for i := range *in { + (*in)[i].DeepCopyInto(&(*out)[i]) + } + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TestFailoverList. +func (in *TestFailoverList) DeepCopy() *TestFailoverList { + if in == nil { + return nil + } + out := new(TestFailoverList) + in.DeepCopyInto(out) + return out +} + +// DeepCopyObject is an autogenerated deepcopy function, copying the receiver, creating a new runtime.Object. +func (in *TestFailoverList) DeepCopyObject() runtime.Object { + if c := in.DeepCopy(); c != nil { + return c + } + return nil +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *TestFailoverReport) DeepCopyInto(out *TestFailoverReport) { + *out = *in + if in.RecoveryPointTime != nil { + in, out := &in.RecoveryPointTime, &out.RecoveryPointTime + *out = (*in).DeepCopy() + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TestFailoverReport. +func (in *TestFailoverReport) DeepCopy() *TestFailoverReport { + if in == nil { + return nil + } + out := new(TestFailoverReport) + in.DeepCopyInto(out) + return out +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *TestFailoverSpec) DeepCopyInto(out *TestFailoverSpec) { + *out = *in + if in.TTLSeconds != nil { + in, out := &in.TTLSeconds, &out.TTLSeconds + *out = new(int64) + **out = **in + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TestFailoverSpec. +func (in *TestFailoverSpec) DeepCopy() *TestFailoverSpec { + if in == nil { + return nil + } + out := new(TestFailoverSpec) + in.DeepCopyInto(out) + return out +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *TestFailoverStatus) DeepCopyInto(out *TestFailoverStatus) { + *out = *in + in.Step.DeepCopyInto(&out.Step) + if in.Clones != nil { + in, out := &in.Clones, &out.Clones + *out = make([]TestFailoverClone, len(*in)) + for i := range *in { + (*in)[i].DeepCopyInto(&(*out)[i]) + } + } + if in.Report != nil { + in, out := &in.Report, &out.Report + *out = new(TestFailoverReport) + (*in).DeepCopyInto(*out) + } + if in.StartedAt != nil { + in, out := &in.StartedAt, &out.StartedAt + *out = (*in).DeepCopy() + } + if in.ReadyAt != nil { + in, out := &in.ReadyAt, &out.ReadyAt + *out = (*in).DeepCopy() + } + if in.CompletedAt != nil { + in, out := &in.CompletedAt, &out.CompletedAt + *out = (*in).DeepCopy() + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new TestFailoverStatus. +func (in *TestFailoverStatus) DeepCopy() *TestFailoverStatus { + if in == nil { + return nil + } + out := new(TestFailoverStatus) + in.DeepCopyInto(out) + return out +} + // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *ThroughputLimits) DeepCopyInto(out *ThroughputLimits) { *out = *in diff --git a/operator/cmd/main.go b/operator/cmd/main.go index 3611bdfde..2bb8b3a6a 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -53,6 +53,7 @@ import ( volumegroupsnapshotv1beta1 "github.com/kubernetes-csi/external-snapshotter/client/v8/apis/volumegroupsnapshot/v1beta1" snapshotv1 "github.com/kubernetes-csi/external-snapshotter/client/v8/apis/volumesnapshot/v1" + workv1 "open-cluster-management.io/api/work/v1" simplyblockv1alpha1 "github.com/simplyblock/simplyblock-operator/api/v1alpha1" simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" @@ -106,6 +107,9 @@ func init() { // external-snapshotter VolumeSnapshot: the VolumeGroupSnapshotOps restore // enumerates a group snapshot's member snapshots (design §7.4). utilruntime.Must(snapshotv1.AddToScheme(scheme)) + // OCM ManifestWork: the TestFailover controller places the bubble PV/PVC on a + // recovery cluster through it (design-test-failover.md §7.6). + utilruntime.Must(workv1.Install(scheme)) // +kubebuilder:scaffold:scheme } @@ -899,6 +903,14 @@ func main() { setupLog.Error(err, "unable to create controller", "controller", "VolumeGroupSnapshotOps") os.Exit(1) } + if err := (&controller.TestFailoverReconciler{ + Client: mgr.GetClient(), + Scheme: mgr.GetScheme(), + Recorder: mgr.GetEventRecorder("testfailover-controller"), + }).SetupWithManager(mgr); err != nil { + setupLog.Error(err, "unable to create controller", "controller", "TestFailover") + os.Exit(1) + } // +kubebuilder:scaffold:builder // Provision the admission webhooks' serving certificate at runtime (self-signed diff --git a/operator/config/crd/bases/storage.simplyblock.io_testfailovers.yaml b/operator/config/crd/bases/storage.simplyblock.io_testfailovers.yaml new file mode 100644 index 000000000..d68cf6ad7 --- /dev/null +++ b/operator/config/crd/bases/storage.simplyblock.io_testfailovers.yaml @@ -0,0 +1,305 @@ +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition +metadata: + annotations: + controller-gen.kubebuilder.io/version: v0.21.0 + name: testfailovers.storage.simplyblock.io +spec: + group: storage.simplyblock.io + names: + kind: TestFailover + listKind: TestFailoverList + plural: testfailovers + shortNames: + - tfo + singular: testfailover + scope: Namespaced + versions: + - additionalPrinterColumns: + - jsonPath: .spec.scope + name: Scope + type: string + - jsonPath: .spec.sourceRef + name: Source + type: string + - jsonPath: .spec.sourceCluster + name: "On" + type: string + - jsonPath: .spec.bubbleCluster + name: Bubble + type: string + - jsonPath: .status.phase + name: Phase + type: string + - jsonPath: .status.step.state + name: Step + type: string + - jsonPath: .status.message + name: Message + priority: 1 + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha2 + schema: + openAPIV3Schema: + description: |- + TestFailover is a one-way, non-disruptive test-failover drill. It recovers a + source volume, or a consistency group, from a snapshot into an isolated + namespace on a chosen cluster as bound PVCs, without touching the source. The + hub reads the source on its cluster and places the bubble on the recovery + cluster through OCM. It runs to a terminal phase, or holds Ready until it is + deleted, and deletion reclaims the clones and any snapshots the drill took. + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: |- + TestFailoverSpec is the request for one non-disruptive test-failover drill. + + The source is named by where it runs and what it is, so the hub can find it + without anyone extracting a backend handle by hand. SourceNamespace is + required for a Volume drill, where the source is a PVC, and unused for a Group + drill, where SourceRef names a consistency group. + properties: + bubbleCluster: + description: |- + BubbleCluster is the OCM ManagedCluster to recover onto: a DR target holding + the replicated point, or another cluster. It must differ from SourceCluster; + test-failover recovers onto a different cluster, never in place. Immutable. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + bubbleNamespace: + default: bubble + description: |- + BubbleNamespace is the namespace on the bubble cluster where the recovered + PVCs are created. Immutable. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + scope: + description: Scope selects what the drill recovers. Immutable. + enum: + - Volume + - Group + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + sourceCluster: + description: |- + SourceCluster is the OCM ManagedCluster the source runs on. The hub reads + the source there through a ManagedClusterView. Immutable. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + sourceNamespace: + description: |- + SourceNamespace is the namespace of the source PVC on SourceCluster. + Required for scope=Volume. Immutable. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + sourceRef: + description: |- + SourceRef names the source on SourceCluster: a PersistentVolumeClaim in + SourceNamespace (scope=Volume), or a consistency group (scope=Group). + Immutable. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + ttlSeconds: + description: |- + TTLSeconds is an optional maximum lifetime: the drill is torn down after it + even without a delete, so a forgotten drill cannot hold a clone forever. + format: int64 + minimum: 0 + type: integer + required: + - bubbleCluster + - scope + - sourceCluster + - sourceRef + type: object + x-kubernetes-validations: + - message: field sourceNamespace is immutable once set + rule: '!has(oldSelf.sourceNamespace) || has(self.sourceNamespace)' + - message: field bubbleNamespace is immutable once set + rule: '!has(oldSelf.bubbleNamespace) || has(self.bubbleNamespace)' + - message: sourceNamespace is required for scope=Volume + rule: self.scope != 'Volume' || has(self.sourceNamespace) + status: + description: TestFailoverStatus is the observed state of one drill. + properties: + clones: + description: Clones is one entry per recovered volume. + items: + description: |- + TestFailoverClone is one recovered volume: the source it came from, the + snapshot and clone the drill built, and the PVC placed on the bubble cluster. + properties: + cloneID: + description: CloneID is the backend id of the writable clone. + type: string + pvcName: + description: PVCName is the bound PVC in the bubble namespace + on the bubble cluster. + type: string + sizeBytes: + description: SizeBytes is the recovered volume's size. + format: int64 + type: integer + snapshotID: + description: |- + SnapshotID is the recovery-point snapshot: the replicated snapshot already on + the bubble cluster's backend that the clone is built from. + type: string + sourceFSType: + description: |- + SourceFSType is the source PV's CSI fsType, carried onto the bubble PV so the + node plugin stages the clone with the filesystem it actually carries. The + clone is a block copy of the source, so its filesystem is the source's; an + empty fsType makes the node plugin default to ext4 and refuse to mount an XFS + volume. + type: string + sourceHandle: + description: SourceHandle is the source volume's backend handle, + read from its PV. + type: string + sourceRef: + description: |- + SourceRef is the source volume, or group member, the recovered volume maps + to. + type: string + sourceVolumeContext: + additionalProperties: + type: string + description: |- + SourceVolumeContext is the source PV's CSI volumeAttributes, minus the + identity and provisioner keys, carried onto the bubble PV so the node plugin + receives a non-nil VolumeContext when it stages the clone. The clone's own + identity (NQN, connections, nsId, and so on) is re-resolved from the clone + handle at stage time, so only the class-level parameters are carried; the + identity keys are dropped so a failed clone lookup can never point the mount + back at the source. + type: object + required: + - sourceRef + type: object + type: array + x-kubernetes-list-map-keys: + - sourceRef + x-kubernetes-list-type: map + completedAt: + description: CompletedAt is when the drill reached a terminal phase. + format: date-time + type: string + message: + description: |- + Message is the reason the phase is what it is: one sentence, replaced as the + drill moves, and never a log. + type: string + observedGeneration: + description: |- + ObservedGeneration is the generation the rest of this status was computed + from, so a stale status can be told from a current one. + format: int64 + type: integer + phase: + description: Phase is the drill's own progress. + enum: + - Pending + - Provisioning + - Ready + - Failed + - TearingDown + type: string + readyAt: + description: ReadyAt is when every recovered PVC became bound. + format: date-time + type: string + report: + description: Report is the drill's evidence, populated as it reaches + Ready. + properties: + bubbleCluster: + description: BubbleCluster is the cluster the drill recovered + onto. + type: string + invariantsHeld: + description: |- + InvariantsHeld is true only when the source fingerprint taken before the + drill matches the one taken at Ready. A Ready drill with this false is a + defect. + type: boolean + recoveryPoint: + description: RecoveryPoint is the snapshot or group generation + the drill recovered. + type: string + recoveryPointAgeSeconds: + description: RecoveryPointAgeSeconds is the drill time minus the + recovery-point time. + format: int64 + type: integer + recoveryPointTime: + description: RecoveryPointTime is when that point was taken. + format: date-time + type: string + type: object + startedAt: + description: StartedAt is when the drill started. + format: date-time + type: string + step: + description: Step is the position of the running drill's state machine. + properties: + deadline: + description: |- + Deadline is when that state expires, absent when it has none. It is an + absolute instant, so a state whose deadline passed while the controller + was down restores as already expired. + format: date-time + type: string + state: + description: |- + State is the state the machine was in. Empty means the resource has not + been reconciled yet, and restores to the graph's initial state. + type: string + type: object + x-kubernetes-validations: + - message: unknown step + rule: '!has(self.state) || self.state in [''ResolvingSource'',''ResolvingPoint'',''Shipping'',''Cloning'',''Placing'',''Releasing'']' + triggered: + description: |- + Triggered records that the current step's side effect was issued, so a + restart does not repeat it. + type: boolean + type: object + type: object + served: true + storage: true + subresources: + status: {} diff --git a/operator/config/crd/kustomization.yaml b/operator/config/crd/kustomization.yaml index 66e4a63c0..760b4ffb6 100644 --- a/operator/config/crd/kustomization.yaml +++ b/operator/config/crd/kustomization.yaml @@ -30,6 +30,7 @@ resources: - bases/storage.simplyblock.io_persistentvolumeops.yaml - bases/storage.simplyblock.io_controlplaneops.yaml - bases/storage.simplyblock.io_storagedeviceops.yaml +- bases/storage.simplyblock.io_testfailovers.yaml # +kubebuilder:scaffold:crdkustomizeresource patches: [] diff --git a/operator/config/manifests/bases/simplyblock-operator.clusterserviceversion.yaml b/operator/config/manifests/bases/simplyblock-operator.clusterserviceversion.yaml index 802b65438..fc70292e9 100644 --- a/operator/config/manifests/bases/simplyblock-operator.clusterserviceversion.yaml +++ b/operator/config/manifests/bases/simplyblock-operator.clusterserviceversion.yaml @@ -593,6 +593,14 @@ spec: displayName: Tasks path: tasks version: v1alpha1 + - description: TestFailover is a one-way, non-disruptive test-failover drill. + It recovers a source volume, or a consistency group, from a snapshot into + an isolated namespace on a chosen cluster as bound PVCs, without touching + the source. + displayName: Test Failover + kind: TestFailover + name: testfailovers.storage.simplyblock.io + version: v1alpha2 description: The Simplyblock Operator helps with installation, operation, and management of Simplyblock Control Planes, Storage Planes, and the CSI Driver. displayName: Simplyblock Operator diff --git a/operator/config/rbac/role.yaml b/operator/config/rbac/role.yaml index c5dc860b5..7fbbf83d5 100644 --- a/operator/config/rbac/role.yaml +++ b/operator/config/rbac/role.yaml @@ -325,6 +325,7 @@ rules: - storagepoolops - storagepools - tasks + - testfailovers - volumemigrations verbs: - create @@ -360,6 +361,7 @@ rules: - storagepoolops/finalizers - storagepools/finalizers - tasks/finalizers + - testfailovers/finalizers - volumemigrations/finalizers verbs: - update @@ -392,6 +394,7 @@ rules: - storagepoolops/status - storagepools/status - tasks/status + - testfailovers/status - volumegroupsnapshotops/status - volumemigrations/status verbs: @@ -418,3 +421,27 @@ rules: - get - list - watch +- apiGroups: + - view.open-cluster-management.io + resources: + - managedclusterviews + verbs: + - create + - delete + - get + - list + - patch + - update + - watch +- apiGroups: + - work.open-cluster-management.io + resources: + - manifestworks + verbs: + - create + - delete + - get + - list + - patch + - update + - watch diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index 9a39c6e50..6ce448b2d 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -11205,6 +11205,311 @@ spec: --- apiVersion: apiextensions.k8s.io/v1 kind: CustomResourceDefinition +metadata: + annotations: + controller-gen.kubebuilder.io/version: v0.21.0 + name: testfailovers.storage.simplyblock.io +spec: + group: storage.simplyblock.io + names: + kind: TestFailover + listKind: TestFailoverList + plural: testfailovers + shortNames: + - tfo + singular: testfailover + scope: Namespaced + versions: + - additionalPrinterColumns: + - jsonPath: .spec.scope + name: Scope + type: string + - jsonPath: .spec.sourceRef + name: Source + type: string + - jsonPath: .spec.sourceCluster + name: "On" + type: string + - jsonPath: .spec.bubbleCluster + name: Bubble + type: string + - jsonPath: .status.phase + name: Phase + type: string + - jsonPath: .status.step.state + name: Step + type: string + - jsonPath: .status.message + name: Message + priority: 1 + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha2 + schema: + openAPIV3Schema: + description: |- + TestFailover is a one-way, non-disruptive test-failover drill. It recovers a + source volume, or a consistency group, from a snapshot into an isolated + namespace on a chosen cluster as bound PVCs, without touching the source. The + hub reads the source on its cluster and places the bubble on the recovery + cluster through OCM. It runs to a terminal phase, or holds Ready until it is + deleted, and deletion reclaims the clones and any snapshots the drill took. + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: |- + TestFailoverSpec is the request for one non-disruptive test-failover drill. + + The source is named by where it runs and what it is, so the hub can find it + without anyone extracting a backend handle by hand. SourceNamespace is + required for a Volume drill, where the source is a PVC, and unused for a Group + drill, where SourceRef names a consistency group. + properties: + bubbleCluster: + description: |- + BubbleCluster is the OCM ManagedCluster to recover onto: a DR target holding + the replicated point, or another cluster. It must differ from SourceCluster; + test-failover recovers onto a different cluster, never in place. Immutable. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + bubbleNamespace: + default: bubble + description: |- + BubbleNamespace is the namespace on the bubble cluster where the recovered + PVCs are created. Immutable. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + scope: + description: Scope selects what the drill recovers. Immutable. + enum: + - Volume + - Group + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + sourceCluster: + description: |- + SourceCluster is the OCM ManagedCluster the source runs on. The hub reads + the source there through a ManagedClusterView. Immutable. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + sourceNamespace: + description: |- + SourceNamespace is the namespace of the source PVC on SourceCluster. + Required for scope=Volume. Immutable. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + sourceRef: + description: |- + SourceRef names the source on SourceCluster: a PersistentVolumeClaim in + SourceNamespace (scope=Volume), or a consistency group (scope=Group). + Immutable. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + ttlSeconds: + description: |- + TTLSeconds is an optional maximum lifetime: the drill is torn down after it + even without a delete, so a forgotten drill cannot hold a clone forever. + format: int64 + minimum: 0 + type: integer + required: + - bubbleCluster + - scope + - sourceCluster + - sourceRef + type: object + x-kubernetes-validations: + - message: field sourceNamespace is immutable once set + rule: '!has(oldSelf.sourceNamespace) || has(self.sourceNamespace)' + - message: field bubbleNamespace is immutable once set + rule: '!has(oldSelf.bubbleNamespace) || has(self.bubbleNamespace)' + - message: sourceNamespace is required for scope=Volume + rule: self.scope != 'Volume' || has(self.sourceNamespace) + status: + description: TestFailoverStatus is the observed state of one drill. + properties: + clones: + description: Clones is one entry per recovered volume. + items: + description: |- + TestFailoverClone is one recovered volume: the source it came from, the + snapshot and clone the drill built, and the PVC placed on the bubble cluster. + properties: + cloneID: + description: CloneID is the backend id of the writable clone. + type: string + pvcName: + description: PVCName is the bound PVC in the bubble namespace + on the bubble cluster. + type: string + sizeBytes: + description: SizeBytes is the recovered volume's size. + format: int64 + type: integer + snapshotID: + description: |- + SnapshotID is the recovery-point snapshot: the replicated snapshot already on + the bubble cluster's backend that the clone is built from. + type: string + sourceFSType: + description: |- + SourceFSType is the source PV's CSI fsType, carried onto the bubble PV so the + node plugin stages the clone with the filesystem it actually carries. The + clone is a block copy of the source, so its filesystem is the source's; an + empty fsType makes the node plugin default to ext4 and refuse to mount an XFS + volume. + type: string + sourceHandle: + description: SourceHandle is the source volume's backend handle, + read from its PV. + type: string + sourceRef: + description: |- + SourceRef is the source volume, or group member, the recovered volume maps + to. + type: string + sourceVolumeContext: + additionalProperties: + type: string + description: |- + SourceVolumeContext is the source PV's CSI volumeAttributes, minus the + identity and provisioner keys, carried onto the bubble PV so the node plugin + receives a non-nil VolumeContext when it stages the clone. The clone's own + identity (NQN, connections, nsId, and so on) is re-resolved from the clone + handle at stage time, so only the class-level parameters are carried; the + identity keys are dropped so a failed clone lookup can never point the mount + back at the source. + type: object + required: + - sourceRef + type: object + type: array + x-kubernetes-list-map-keys: + - sourceRef + x-kubernetes-list-type: map + completedAt: + description: CompletedAt is when the drill reached a terminal phase. + format: date-time + type: string + message: + description: |- + Message is the reason the phase is what it is: one sentence, replaced as the + drill moves, and never a log. + type: string + observedGeneration: + description: |- + ObservedGeneration is the generation the rest of this status was computed + from, so a stale status can be told from a current one. + format: int64 + type: integer + phase: + description: Phase is the drill's own progress. + enum: + - Pending + - Provisioning + - Ready + - Failed + - TearingDown + type: string + readyAt: + description: ReadyAt is when every recovered PVC became bound. + format: date-time + type: string + report: + description: Report is the drill's evidence, populated as it reaches + Ready. + properties: + bubbleCluster: + description: BubbleCluster is the cluster the drill recovered + onto. + type: string + invariantsHeld: + description: |- + InvariantsHeld is true only when the source fingerprint taken before the + drill matches the one taken at Ready. A Ready drill with this false is a + defect. + type: boolean + recoveryPoint: + description: RecoveryPoint is the snapshot or group generation + the drill recovered. + type: string + recoveryPointAgeSeconds: + description: RecoveryPointAgeSeconds is the drill time minus the + recovery-point time. + format: int64 + type: integer + recoveryPointTime: + description: RecoveryPointTime is when that point was taken. + format: date-time + type: string + type: object + startedAt: + description: StartedAt is when the drill started. + format: date-time + type: string + step: + description: Step is the position of the running drill's state machine. + properties: + deadline: + description: |- + Deadline is when that state expires, absent when it has none. It is an + absolute instant, so a state whose deadline passed while the controller + was down restores as already expired. + format: date-time + type: string + state: + description: |- + State is the state the machine was in. Empty means the resource has not + been reconciled yet, and restores to the graph's initial state. + type: string + type: object + x-kubernetes-validations: + - message: unknown step + rule: '!has(self.state) || self.state in [''ResolvingSource'',''ResolvingPoint'',''Shipping'',''Cloning'',''Placing'',''Releasing'']' + triggered: + description: |- + Triggered records that the current step's side effect was issued, so a + restart does not repeat it. + type: boolean + type: object + type: object + served: true + storage: true + subresources: + status: {} +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition metadata: annotations: controller-gen.kubebuilder.io/version: v0.21.0 @@ -12054,6 +12359,7 @@ rules: - storagepoolops - storagepools - tasks + - testfailovers - volumemigrations verbs: - create @@ -12089,6 +12395,7 @@ rules: - storagepoolops/finalizers - storagepools/finalizers - tasks/finalizers + - testfailovers/finalizers - volumemigrations/finalizers verbs: - update @@ -12121,6 +12428,7 @@ rules: - storagepoolops/status - storagepools/status - tasks/status + - testfailovers/status - volumegroupsnapshotops/status - volumemigrations/status verbs: @@ -12147,6 +12455,30 @@ rules: - get - list - watch +- apiGroups: + - view.open-cluster-management.io + resources: + - managedclusterviews + verbs: + - create + - delete + - get + - list + - patch + - update + - watch +- apiGroups: + - work.open-cluster-management.io + resources: + - manifestworks + verbs: + - create + - delete + - get + - list + - patch + - update + - watch --- apiVersion: rbac.authorization.k8s.io/v1 kind: ClusterRole diff --git a/operator/docs/designs/design-test-failover.md b/operator/docs/designs/design-test-failover.md new file mode 100644 index 000000000..930da44a2 --- /dev/null +++ b/operator/docs/designs/design-test-failover.md @@ -0,0 +1,612 @@ +# Design Document: Non-Disruptive Test Failover + +**Status:** Draft +**Author:** Israel Geoffrey (geoffrey1330) +**Date:** 2026-09-29 +**Test Plan:** [`tests/test-plan-test-failover.md`](../tests/test-plan-test-failover.md) + +--- + +## Phasing Overview + +| Phase | Status | Where the bubble runs | Recovery point | New capability | Sections | +|-------------|---------|---------------------------------|-----------------------------------------------|------------------------------------------------------------------|-----------------------| +| **Phase 1** | Planned | The DR target cluster | The replicated snapshot already on the target | Cross-cluster read and placement from the hub (OCM) | §4, §5.1–§5.5, §6, §7 | +| **Phase 2** | Planned | A cluster that holds no replica | A point shipped there on demand | On-demand shipping of a recovery point to a named backend (P0-6) | §5.6 | + +Test-failover recovers onto a cluster other than the source's, never in place: the point is to rehearse the site a real failover would move to, and recovering into the source's own backend would clone into the subsystem serving the live volume. The two phases differ on one axis: whether the bubble cluster already holds a replicated copy of the source. In Phase 1 the bubble runs on the DR target, where replication has already landed a snapshot, so no data moves. In Phase 2 the bubble runs on a cluster with no copy, the only case that needs data shipped on demand and the design's long pole. + +The hub coordinates both phases. It never has to run the source or the bubble itself, and both may be any managed cluster, but the bubble must differ from the source. What Phase 1 needs, and Phase 2 inherits, is the ability to read the source object on its cluster and place the bubble object on the recovery cluster, both from the hub, which OCM provides. + +--- + +## Phase 0 — External Prerequisites + +| # | Prerequisite | Kind | Blocks | Status | +|------|----------------------------------------------------------------------------------------------------------------------------------------------|-------------------------|---------|----------------------------------------------------------------------------------------------------------------------------------------------------------| +| P0-2 | Clone a snapshot into a writable volume in a chosen pool, and return the clone's volume handle | Control plane (`sbcli`) | All | Shipped: `snapshot_controller.clone`, and the CSI clone-from-snapshot path | +| P0-3 | Delete a volume, idempotent | Control plane (`sbcli`) | All | Shipped | +| P0-4 | Resolve the latest replicated snapshot on a DR-target backend for a source relationship | Control plane (`sbcli`) | Phase 1 | Shipped: `lvol_controller.latest_replicated_snapshot` and `replication_policy_controller.latest_replicated_generation`, with v2 endpoints | +| P0-5 | From the hub, read an object on a managed cluster (`ManagedClusterView`) and place objects on it (`ManifestWork`), each with status feedback | Ecosystem (OCM) | Phase 1 | Available: both are OCM primitives, once the cluster is a registered `ManagedCluster` with a working view controller (04-bootstrap-ocm.sh) | +| P0-6 | On-demand shipping of a specific snapshot or group generation to a named target backend that holds no copy, followed by a clone there | Control plane (`sbcli`) | Phase 2 | Not shipped. The long pole of Phase 2. Today's cross-cluster reach is the continuous replication engine or the S3 backup path, neither an on-demand push | + +The recovery point numbering keeps its original P0- ids so cross-references hold; P0-1 (taking a fresh source snapshot) is gone with the in-place case. Phase 1 has no unmet storage prerequisite: the recovery point is the replicated snapshot already on the target backend (P0-4), and it clones and deletes with calls that ship. Its one non-storage need is OCM (P0-5), which the DR setup already establishes, and it is needed because the hub, as coordinator, reaches the source and the bubble on their own clusters through it. Phase 2 is the exception: P0-6 is genuinely new, because moving one recovery point to a backend that holds no copy of it, on demand, is a capability the engine does not have. + +--- + +## Table of Contents + +1. [Background](#1-background) +2. [Goals and Non-Goals](#2-goals-and-non-goals) +3. [Architecture Overview](#3-architecture-overview) +4. [API Design — New CRD](#4-api-design--new-crd) +5. [Core Mechanism](#5-core-mechanism) +6. [State Machine](#6-state-machine) +7. [Controller Design](#7-controller-design) +8. [Backend API Requirements](#8-backend-api-requirements) +9. [Configuration](#9-configuration) +10. [Failure Modes and Fallback](#10-failure-modes-and-fallback) +11. [Observability](#11-observability) +12. [Testing Strategy](#12-testing-strategy) +13. [Open Questions](#13-open-questions) +14. [Appendix A: `testfailover_types.go`](#appendix-a-testfailover_typesgo) + +--- + +## Overview + +A test failover proves an application can be recovered from a point-in-time copy, in isolation, without disturbing the running production. The recovery point is always a snapshot and the result is always a clone, so the source is never touched. What varies is where the bubble runs, which is what a real DR test cares about: recovering onto the site a real failover would move to. + +The feature is a new CRD, `TestFailover`, and its controller, both on the hub, which coordinates the drill. The object names the source, by the cluster it runs on and its PVC, and a place to recover it, the `bubbleCluster`, which must be a different cluster. The controller reads the source PVC on its cluster to learn its volume, resolves the recovery point on the bubble's backend, clones it there, and places the clone on the bubble cluster as a bound PVC in an isolated namespace, `bubble` by default. The operator boots the application there, confirms the data, and deletes the drill, which reclaims the clone. The source serves throughout. + +`bubbleCluster` selects the topology, and always names a cluster other than the source's. Naming the DR target recovers from the replicated snapshot already sitting on its backend (Phase 1), a genuine "fail over to the target site" test with no data moved. Naming a cluster that holds no copy is the case that needs the point shipped there first (Phase 2). Both are one object, one controller, and one state machine, differing only in whether the point is already on the bubble's backend or must be shipped there. + +--- + +## 1. Background + +simplyblock has a real failover, driven either by Ramen (`DRPC.spec.action: Failover`) or imperatively by the simplyblock-native `ReplicationOps` CR. Both promote a replicated copy on the DR target and land the recovered workload there. That copy is, mechanically, a clone of the last replicated snapshot on the target, so a real failover is a clone-and-promote of a recovery point that already lives on the target's backend. + +Three facts shape a test failover. First, replication is between two clusters: `add_target` refuses `target_cluster_id == cluster_id` with "A cluster cannot replicate to itself" (`simplyblock_core/controllers/replication_policy_controller.py`). The DR target holds the replicated copy, which `lvol_controller.latest_replicated_snapshot` resolves without triggering anything, so the drill recovers there and never in the source's own cluster. Second, the primitives to recover from that point already exist: `snapshot_controller.clone` clones the replicated snapshot into a writable volume, and a volume can be deleted. A clone is a first-class volume with its own handle, which CSI static provisioning adopts as a `PersistentVolume`. Third, the hub already reaches its managed clusters both ways in this deployment: it reads an object on one through an OCM `ManagedClusterView`, the way Ramen reads a spoke's status, and writes one through a `ManifestWork`, the way Ramen places a workload. + +What is missing is the orchestration: an object that finds the source on its cluster, resolves the right recovery point, clones it on the right backend, places the bubble PVC on the right cluster, proves the copy is recoverable, and tears it down, all without touching the source. Ramen orchestrates none of it, because Ramen only fails over and relocates for real. This design is that object. + +--- + +## 2. Goals and Non-Goals + +### Goals + +- A `TestFailover` CRD and controller that, from one object on the hub, produce a bound PVC per source volume in an isolated namespace, on the cluster the drill recovers onto. +- Locate the source from the hub. The object names the source by its cluster and PVC, and the controller reads that PVC through OCM to learn its volume, so nobody has to hand-extract a backend handle. +- Recover onto the DR target site, from the replicated snapshot already there, with the bubble PVC placed on that cluster. This is the case that makes a test failover a real rehearsal of the DR target. +- Non-disruptive by construction. The recovery point is a snapshot and the result is a clone, so the source's data and I/O are never touched, and the controller records a before-and-after fingerprint of the source and its replication relationship so a regression is caught rather than assumed. +- Volume-scoped and consistency-group-scoped drills. A group drill recovers one PVC per member from one group-consistent point. +- DR target now, an arbitrary cluster later (§5), behind one object and one state machine. +- A finalizer-driven teardown that reclaims the clone, removes the placed PVC from the bubble cluster, and proves nothing test-labeled remains. +- Restart safety. The controller records which side effect each step issued, so a restart mid-drill resumes rather than repeats. + +### Non-Goals + +- **Bringing up the application.** The object produces bound PVCs and stops. The workload that consumes them is the operator's to deploy. An application lifecycle and its workload spec are a separate concern, out of scope here. +- **A test failback.** The drill is one-way. Tearing it down reclaims the clone. There is no promote-back. +- **Recovering in place, on the source's own cluster.** `bubbleCluster` must name a different cluster, and a same-cluster drill is rejected up front. Recovering into the source's backend would clone into the subsystem serving the live volume, and it rehearses no failover site. +- **Recovering onto a cluster with no copy in Phase 1.** A bubble cluster whose backend holds no replica needs the point shipped there, which is Phase 2 (§5.6), gated on a backend primitive that does not exist (P0-6). +- **Snapshot scheduling and evidence export.** A recurring schedule and a signed test report are a layer above this object and are out of scope here. +- **Replacing Ramen's own test paths.** Ramen has no non-disruptive test. This design does not add one to Ramen. It is a simplyblock-native object that reuses OCM only as the transport for reading the source and placing the bubble. + +--- + +## 3. Architecture Overview + +``` + TestFailover CR ──▶┌──────────────────────────────────────────────────────┐ + (on the hub) │ hub operator: TestFailoverReconciler │ + spec: scope, │ 1. read source PVC on sourceCluster (ManagedClusterView) → handle │ + sourceCluster, │ 2. resolve recovery point on the bubble's backend │ + sourceNamespace, │ replicated snapshot (P0-4) │ + sourceRef, │ 3. [Phase 2] ship point to that backend (P0-6) │ + bubbleCluster, │ 4. clone the point on that backend (P0-2) │ + bubbleNamespace │ 5. place PV + PVC on bubbleCluster (ManifestWork) │ + │ 6. Ready; hold until deleted │ + │ 7. finalizer: remove PVC, reclaim clone │ + └──┬──────────────┬───────────────────────┬────────────┘ + ManagedClusterView│ │ REST (webapi.Client) │ ManifestWork + (read source) ▼ ▼ ▼ (place bubble) + ┌────────────────────────┐ ┌──────────────────┐ ┌────────────────────────┐ + │ source cluster │ │ control plane │ │ bubble cluster │ + │ PVC + PV (volumeHandle)│ │ clone P0-2 │ │ work-agent applies: │ + │ projected to the hub │ │ delete P0-3 │ │ PersistentVolume │ + └────────────────────────┘ │ latest-repl P0-4│ │ PersistentVolumeClaim│ + │ ship P0-6 (P2) │ │ (bound to the clone │ + └──────────────────┘ │ on its backend) │ + └────────────────────────┘ +``` + +The controller runs on the hub, which coordinates the drill and runs neither the source nor the bubble. It learns the source's volume by reading the source PVC and its PV on `sourceCluster` through an OCM `ManagedClusterView`, which projects their current state back to the hub. It reads the source only to fingerprint it, never to change it. The clone is built on the bubble cluster's own backend, and the bubble cluster's CSI driver adopts it through a static `PersistentVolume` naming the clone's handle, with a `PersistentVolumeClaim` bound to it in the bubble namespace. + +Placement is uniform: the controller delivers the bubble PV and PVC to `bubbleCluster` as an OCM `ManifestWork`, whose work-agent applies them and reports the bind result back through the `ManifestWork` status. The hub addresses both the source and the bubble this way, including its own cluster when it is self-managed. The trust boundary for storage is the control-plane REST API, reached with the admin bearer token the operator already uses. The trust boundary for cross-cluster read and write is OCM, which the DR setup already establishes (04-bootstrap-ocm.sh). + +--- + +## 4. API Design — New CRD + +`TestFailover` is a namespaced object in the `storage.simplyblock.io` group, created on the hub. One object drives one drill. Its spec is immutable, because the object is a request and a drill whose target moved under the controller mid-flight has no coherent meaning. The full type is [Appendix A](#appendix-a-testfailover_typesgo), and the body shows only the fields an argument turns on. + +### 4.1 `TestFailover` Spec + +The spec names the source by where it runs and what it is, and names where to recover it. `sourceCluster` and `sourceRef` are what let the hub find the source without anyone extracting a backend handle by hand. + +```go +// SourceCluster is the OCM ManagedCluster the source runs on. The hub reads the +// source there through a ManagedClusterView. Immutable. +// +kubebuilder:validation:Required +// +k8s:immutable +SourceCluster string `json:"sourceCluster"` + +// SourceRef names the source on SourceCluster: a PersistentVolumeClaim in +// SourceNamespace (scope=Volume), or a consistency group (scope=Group). +// Immutable. +// +kubebuilder:validation:Required +// +k8s:immutable +SourceRef string `json:"sourceRef"` + +// BubbleCluster is the OCM ManagedCluster to recover onto: a DR target holding +// the replicated point, or another cluster. It must differ from SourceCluster; +// test-failover recovers onto a different cluster, never in place. Immutable. +// +kubebuilder:validation:Required +// +k8s:immutable +BubbleCluster string `json:"bubbleCluster"` +``` + +`sourceCluster` answers "where is the PVC to test," and the controller reads it there rather than requiring a handle. `bubbleCluster` selects the topology (§5) and carries the "recover onto the target site" intent; it must name a cluster other than the source's, and a same-cluster drill is rejected up front. `bubbleNamespace` defaults to `bubble` and is the namespace on the bubble cluster where every recovered PVC lands, isolated so it cannot collide with the source workload's PVCs, which carry the same names. + +### 4.2 `TestFailover` Status + +The status carries the state-machine position, the resolved source and recovery point, one entry per recovered volume, and a report. `phase` is the coarse lifecycle and `step` is the durable machine position with its deadline (§6). `clones` is the list the finalizer reclaims from, and `report` is the evidence a reader takes away. + +```go +// Clones is one entry per recovered volume: the source it came from, the +// snapshot and clone the drill built, and the PVC placed on the bubble cluster. +// +optional +// +listType=map +// +listMapKey=sourceRef +Clones []TestFailoverClone `json:"clones,omitempty"` + +// Report is the drill's evidence: the source and point recovered, its age, the +// cluster it ran on, and whether the source was untouched. Populated as the +// drill reaches Ready. +// +optional +Report *TestFailoverReport `json:"report,omitempty"` +``` + +An invariant the controller enforces and the status records: the drill is non-disruptive. `status.report.invariantsHeld` is set only when the fingerprint of the source and its replication relationship taken before the drill matches the one taken after (§7.4). A drill that reached `Ready` with `invariantsHeld: false` is a defect, not a passing test. + +The object owns a finalizer, `storage.simplyblock.io/testfailover-teardown`. Deletion runs the teardown state (§6) before the finalizer is removed, so a clone or a placed PVC is never orphaned by a delete that races the controller. + +--- + +## 5. Core Mechanism + +### 5.1 Locating the source + +The hub does not run the source, so it reads it. The controller creates an OCM `ManagedClusterView` on `sourceCluster` for the source PVC named by `sourceRef` in `sourceNamespace`, and for the PV it is bound to, which projects their current state back to the hub. From the PV's `spec.csi.volumeHandle` it learns the source volume's backend handle, the identity every later step keys on. For a group drill, `sourceRef` names a consistency group, whose member volumes the control plane resolves from the group id on `sourceCluster`'s backend, so the drill recovers the whole set. + +Reading the source is also where the non-disruptiveness fingerprint begins: the handle, the PVC's binding, and the replication relationship's state are captured here and compared again at the end (§7.4). + +### 5.2 Resolving the recovery point + +The recovery point is a snapshot already on the bubble cluster's backend. For a volume drill it is the latest replicated snapshot for the source relationship, which `latest_replicated_snapshot` resolves from the source handle (P0-4). For a group drill it is one group-consistent snapshot, resolved as the latest replicated generation on the target. Either way the resolution triggers nothing, because replication already produced the point. + +Resolving the replicated point touches nothing: replication already produced it, and the drill neither promotes nor commits anything, which is what keeps the source and the live replication relationship untouched. The drill takes no snapshot of its own, so it has none to delete on teardown. + +### 5.3 Cloning the point on the bubble's backend + +The controller clones the recovery-point snapshot into a writable volume on the bubble cluster's own backend (P0-2), and the backend returns the clone's volume handle. The replicated snapshot already sits on that backend, so the clone is local to the point and no data crosses a cluster boundary. The clone is a first-class volume, tagged with the drill's `test-id` for reclaim. + +### 5.4 Placing the bubble PVC + +The clone is adopted as a static `PersistentVolume` whose `spec.csi.volumeHandle` is the clone's handle, with a `PersistentVolumeClaim` bound to it in the bubble namespace, sized from the source's request. The controller delivers the PV and PVC to `bubbleCluster` as an OCM `ManifestWork` (P0-5), and that cluster's work-agent applies them and reports the PVC bound through the `ManifestWork` status. The PV carries `persistentVolumeReclaimPolicy: Retain`, so deleting the PVC does not delete the clone, which the controller reclaims at the backend on teardown. + +The result is one bound PVC per source volume in the bubble namespace on the bubble cluster. For a group drill, one PVC per member from the one point, which is what makes the recovered set crash-consistent. + +### 5.5 Recovering onto the DR target (Phase 1) + +The DR-target case is why the source is named by cluster and PVC rather than assumed local. The source runs on one cluster and the bubble on the DR target, so the controller, from the hub, reads the source PVC on its cluster (§5.1), resolves its handle to the replicated snapshot on the target's backend (P0-4), clones it there (§5.3), and places the bubble PVC on the target through `ManifestWork` (§5.4). Nothing is shipped, because replication already put the point on the target. This is a true rehearsal of the site that would take over in a real failover, and it leaves the running replication relationship exactly as it was. + +### 5.6 Shipping to a cluster with no copy (Phase 2) + +A bubble cluster whose backend holds no replica of the source needs the point moved there before it can be cloned. The controller calls the shipping verb (P0-6), which replicates the specific recovery point to that backend as a cloneable object, and then the clone and placement steps run there as in §5.3 and §5.4. + +This is the design's long pole, because the shipping primitive does not exist. The continuous replication engine is pair-scoped and aimed at the DR target, and the S3 backup path is not a cluster-to-cluster push. Neither is an on-demand "ship this one point to cluster X now." P0-6 is that new capability, and Phase 2 does not ship until it does. + +### 5.7 Teardown + +Deleting the `TestFailover` runs the teardown state before the finalizer clears. The controller removes the PVC and its static PV from the bubble cluster by deleting their `ManifestWork`, removes the `ManagedClusterView` it created on the source, reclaims each clone at the backend, then enumerates by the drill's `test-id` label to prove nothing remains. The recovery point is a replicated snapshot the drill only resolved, never created, so it is left alone. Only then is the finalizer removed. A teardown that cannot confirm a reclaim holds the object in `TearingDown` with the reason on `status.message`, rather than removing the finalizer and orphaning backend storage. + +--- + +## 6. State Machine + +``` +Pending + │ spec admitted, finalizer added + ▼ +Provisioning ──(step: ResolvingSource)──▶ read source PVC/PV on sourceCluster + │ via ManagedClusterView → handle + │ (step: ResolvingPoint) latest replicated snapshot (P0-4) + │ ← status.report.recoveryPoint set + │ (step: Shipping) [Phase 2 only] ship point to the bubble's backend (P0-6) + │ (step: Cloning) clone on the bubble's backend (P0-2) + │ ← status.clones[].cloneID set + │ (step: Placing) deliver PV + PVC to bubbleCluster via ManifestWork; + │ wait Bound + ▼ +Ready ────────────────────────────────── every PVC Bound; report populated + │ (holds here until the object is deleted) + │ .metadata.deletionTimestamp set + ▼ +TearingDown ──(step: Releasing)──▶ delete ManifestWork + ManagedClusterView, + │ reclaim clones (the replicated point is left alone) + ▼ +(finalizer removed, object gone) + +Failed ◀── any step's deadline expires, or a backend, read, or placement call + fails terminally (object stays; teardown still runs on delete) +``` + +The machine position lives in `status.step` as a snapshot carrying the state and the deadline that state expires at, so a restored controller times a stalled step out rather than waiting forever. `status.phase` is the coarse view for `kubectl get`. Each step records `status.step.triggered` once its side effect is issued, so a restart between cloning and recording the handle does not clone twice, and a restart between placing and confirming does not place twice. + +| Condition | Step | Result | +|---------------------------------|-----------------|----------------------------------------------------------------------------------------------------| +| User deletes mid-drill | any | `TearingDown`: reclaim whatever `status.clones` records, then clear the finalizer | +| Operator restart | any | resume from `status.step`, and `triggered` prevents re-issuing the current step's side effect | +| Source PVC not found on cluster | ResolvingSource | `Failed`: the view returns nothing for `sourceRef` on `sourceCluster` | +| Source cluster not managed | ResolvingSource | `Failed`: `sourceCluster` is not a registered `ManagedCluster` | +| No replicated point on target | ResolvingPoint | `Failed` for a Phase 1 drill: replication has landed nothing on the bubble's backend yet | +| Bubble equals source cluster | ResolvingSource | `Failed`: `bubbleCluster` must differ from `sourceCluster`; a same-cluster drill is rejected | +| Backend clone error | Cloning | `Failed`. The replicated point is left alone; the clone is reclaimed on delete | +| Bubble cluster not managed | Placing | `Failed`: `bubbleCluster` is not a registered `ManagedCluster` | +| PVC never binds | Placing | `Failed` at the deadline. The clone is recorded and reclaimed on delete | +| Clone reclaim fails | Releasing | hold in `TearingDown` with the reason. The finalizer is not removed until the reclaim is confirmed | + +--- + +## 7. Controller Design + +### 7.1 Location + +`internal/controller/testfailover_controller.go`, `TestFailoverReconciler`, on the hub. It reuses the `webapi.Client` the replication controllers use for control-plane calls, an OCM `ManagedClusterView` to read the source, and an OCM `ManifestWork` to place the bubble. + +### 7.2 Reconciliation Trigger + +Watches `TestFailover` and owns the `ManagedClusterView` and `ManifestWork` it creates for each drill, reconciling on their status feedback, which carries the source's projected state and the placed PVC's bind state back to the hub. It requeues on its own step deadline so a stalled step is detected without an external event. + +### 7.3 Concurrency and Mutual Exclusion + +A drill does not mutate the source, so two drills against one source cannot corrupt it. What they can do is duplicate snapshots and clones, so the controller keys one active drill per resolved `(scope, sourceCluster, sourceRef, bubbleCluster)` and refuses a second with an admission rule where CEL can express it and a `Failed` phase otherwise. This is a lighter lock than the `ReplicationOps` entity lock (`ActiveOpsRef`), because there is no production object whose single-writer invariant has to be defended. + +### 7.4 Interaction with Existing Controllers + +The drill is invisible to production by construction. It only reads the source, resolves or takes a snapshot, and clones the snapshot, none of which changes the source volume or the live replication. To make "invisible" checkable rather than asserted, the controller fingerprints the source at `ResolvingSource` and again at `Ready`: the source PVC is still bound to the same volume, and the relationship's `ReplicationSlot`, and any `VolumeReplication` or `VolumeGroupReplication` object, are unchanged in state and lag. A mismatch sets `status.report.invariantsHeld: false` and moves the object to `Failed`, because a drill that changed production has failed at its one core promise. + +### 7.5 RBAC + +New rules: `testfailovers` and `testfailovers/status` and `testfailovers/finalizers` (full), and `create`/`delete`/`get`/`list`/`watch` on `managedclusterviews` and `manifestworks` in a cluster's namespace on the hub. The source's projection and the bubble's PV and PVC are handled by the managed clusters' own agents, so the hub operator needs no direct `persistentvolume`, `persistentvolumeclaim`, or snapshot-API permission. No new permission on any production replication CR either: the controller reads them, and read is a permission the operator already holds. + +### 7.6 Cross-Cluster Read and Placement + +The hub reaches its managed clusters through OCM, the transport already established for DR. To read the source, the controller creates a `ManagedClusterView` in `sourceCluster`'s namespace naming the PVC and PV, and the view controller on that cluster projects them back into the view's status. To place the bubble, the controller creates a `ManifestWork` in `bubbleCluster`'s namespace carrying the PV and PVC, and the work-agent on that cluster applies them and reports their status back, which is how the hub learns the bubble is bound without a direct connection to either cluster's API server. Both clusters must be registered `ManagedCluster`s (04-bootstrap-ocm.sh). Teardown deletes both objects, and OCM garbage-collects the applied PV and PVC on the bubble cluster. + +--- + +## 8. Backend API Requirements + +| Method | Endpoint | Notes | +|--------|---------------------------------------------------------------------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------| +| GET | `.../clusters/{c}/replication/relationships/{lvol}/latest-snapshot` | **Exists** (P0-4). Resolves the latest replicated snapshot on the DR-target backend. Pure read. Group form: `latest-generation` | +| POST | `.../clusters/{c}/snapshots/{id}/clone` | **Exists** (P0-2). Clones the snapshot into a chosen pool and returns the clone's volume handle. Backed by `snapshot_controller.clone` | +| DELETE | `.../clusters/{c}/storage-pools/{p}/volumes/{id}` | **Exists** (P0-3). Reclaims a clone. Idempotent | +| POST | `.../clusters/{c}/replication/ship-snapshot` | **New** (P0-6, Phase 2). Ships a named recovery point to a backend that holds no copy. Idempotent per `(recovery-point, target)`. Long-running: returns a handle to poll | + +Phase 1 uses only endpoints that exist. The one mutating call the controller may retry after a restart, the clone, is made idempotent by keying on the drill's `test-id`, so a retry that finds a matching clone reuses it. The exact v2 route spellings mirror the existing clone route and are confirmed against the API at implementation time. P0-6's ship is a long-running call and returns a handle the controller polls, with a deadline that moves the step to `Failed` on expiry. The source read and bubble placement are OCM, not backend calls (§7.6). + +--- + +## 9. Configuration + +| Field | Type | Default | Description | +|------------------------|--------|-----------------------|--------------------------------------------------------------------------------------------------------------------------------------------------------| +| `spec.sourceCluster` | string | (required) | The OCM `ManagedCluster` the source runs on. Immutable | +| `spec.sourceNamespace` | string | (required for Volume) | The namespace of the source PVC. Immutable | +| `spec.sourceRef` | string | (required) | The source PVC (Volume) or consistency group (Group) on `sourceCluster`. Immutable | +| `spec.bubbleCluster` | string | (required) | The OCM `ManagedCluster` to recover onto. Must differ from `sourceCluster`. Immutable | +| `spec.bubbleNamespace` | string | `bubble` | Namespace on the bubble cluster the recovered PVCs are created in. Immutable | +| `spec.ttlSeconds` | int | unset | Optional maximum lifetime. When set, the drill is torn down after the deadline even without a delete, so a forgotten drill cannot hold a clone forever | + +`ttlSeconds` is the only runtime-relevant knob and it is advisory: the controller reads it once at `Ready` and schedules a teardown. Changing the rest of the spec after creation is refused by immutability, because a drill whose target moved mid-flight has no coherent meaning. + +--- + +## 10. Failure Modes and Fallback + +| Failure | Detection | Behavior | +|------------------------------------|------------------------------|----------------------------------------------------------------------------------------------------------------------| +| Control plane unreachable | REST call error | Requeue with backoff, and the step deadline eventually moves the object to `Failed`. No partial state is committed | +| Source cluster not managed | ManagedClusterView error | `Failed` at `ResolvingSource`: `sourceCluster` must be registered on the hub | +| Source PVC not found | view projects nothing | `Failed` at `ResolvingSource` with the `sourceRef` in `status.message` | +| No replicated point on the target | latest-snapshot 404 | `Failed` at `ResolvingPoint` for a Phase 1 drill. Replication has landed nothing on the bubble's backend yet | +| Bubble cluster not managed | ManifestWork placement error | `Failed` at `Placing`: the target must be registered on the hub | +| PVC never binds on the bubble | ManifestWork status timeout | `Failed` at `Placing`. The clone is recorded and reclaimed on delete | +| Fingerprint drift (source changed) | Compare at `Ready` | `Failed`, `invariantsHeld: false`. This is the guard, not an expected path | +| Reclaim cannot be confirmed | delete call non-success | Hold in `TearingDown`, finalizer retained, reason on `status.message`. Never orphan a clone, snapshot, or placed PVC | + +Every path degrades to a named state. The one path that must never degrade silently is the fingerprint guard: a drill that cannot prove it left the source untouched fails, rather than passing on the assumption that it did. + +--- + +## 11. Observability + +The operator has no metrics or events for a non-disruptive test today, because the capability does not exist. Everything below is new, on the new kind. + +### Kubernetes Events + +Events land on the `TestFailover` object, which lives on the hub and outlives each step it reports. + +| Event | Type | Reason | +|---------------------------------------------------------------|---------|-----------------------| +| The source resolved to volume X on cluster Y | Normal | SourceResolved | +| The recovery point resolved to snapshot X at time T | Normal | RecoveryPointResolved | +| The clone was built on the bubble's backend, source untouched | Normal | CloneBuilt | +| The bubble PVC is bound on cluster X and the drill is ready | Normal | BubbleReady | +| The drill failed because the source changed during the drill | Warning | InvariantViolated | +| A clone, snapshot, or placed PVC could not be reclaimed | Warning | ReclaimPending | + +`InvariantViolated` and `ReclaimPending` are the two that matter most. The first says the drill stopped being non-disruptive, and the second is a teardown correctly refusing to orphan storage or a peer object, which is a hold that would otherwise look like a hang. + +### Prometheus Metrics + +| Metric | Labels | Description | +|-------------------------------------------------------|----------------------------|----------------------------------------------------------------------------------| +| `simplyblock_testfailover_drills_total` | `scope`, `phase`, `result` | Counter of completed drills by outcome (`ready`, `failed`, `invariant_violated`) | +| `simplyblock_testfailover_duration_seconds` | `scope`, `phase`, `step` | Histogram of time spent per step, for the recover-time estimate | +| `simplyblock_testfailover_active` | `phase` | Gauge of drills currently holding a clone, for capacity watch | +| `simplyblock_testfailover_recovery_point_age_seconds` | `scope` | Gauge of the recovery point's age at drill time (now minus the snapshot time) | + +The two load-bearing metrics are `simplyblock_testfailover_drills_total` with `result="invariant_violated"`, which is the alert that a test failover stopped being non-disruptive, and `simplyblock_testfailover_active`, which is the alert that clones are accumulating because teardowns are not completing. + +--- + +## 12. Testing Strategy + +Full scenario matrix, coverage status, and hand-off test concepts: +[`tests/test-plan-test-failover.md`](../tests/test-plan-test-failover.md) + +- **Unit:** the state-machine transitions and their deadlines, the fingerprint comparison (drift detected and no-drift accepted), the source resolution from a `ManagedClusterView` projection, the DR-target recovery-point resolution, the same-cluster rejection, the idempotency keying, and the scope-to-source resolution. +- **Integration:** the reconcile loop against `envtest`, a mock control plane, and a mock OCM (`ManagedClusterView` and `ManifestWork` with status feedback). The DR-target drill, and the teardown that reclaims and proves no leftovers. The restart case per step. The non-disruptiveness guard, asserting a source change fails the drill. +- **E2E:** a live two-cluster DR setup where the replicated point on the target is cloned there, the bubble PVC is placed on the target, a pod boots on it, and the recovered marker matches, with the source and its replication lag asserted untouched. +- **Load / long-running:** none in Phase 1 or 2. + +The Phase 2 scenarios (§5.6) become testable only when P0-6 exists. The risk concentrates in the non-disruptiveness guard (§7.4), the cross-cluster read and placement and their status feedback (§7.6), and the teardown reclaim (§5.7): the first is the feature's core promise, and the last is where a bug leaks backend storage or a peer object. + +--- + +## 13. Open Questions + +| # | Question | Owner | +|-----|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|---------------------| +| 1 | **Group point on the replica.** `latest_replicated_generation` resolves a group generation on the DR target, but a group-consistent point on a replica was flagged as unconfirmed. Is the generation crash-consistent across members on the target, or only on the source? | SPDK / Backend team | +| 2 | **Clone accounting.** A drill's clone consumes the bubble backend's lvstore object budget for the drill's life. Does the backend expose a per-clone reservation the controller can pre-check, or does a drill risk failing at `Cloning` on a full lvstore with no admission-time warning? | SPDK / Backend team | +| 3 | **Leftover proof under a disabled `LIST_VOLUMES`.** The CSI driver does not advertise `LIST_VOLUMES`, so teardown cannot cross-check backend volumes against Kubernetes objects through CSI. Is the label enumeration sufficient, or is a backend enumeration needed to guarantee no leaked clone or snapshot? | SPDK / Backend team | +| 4 | **P0-6 shape (Phase 2).** Is on-demand shipping a new engine mode (a one-shot pair-and-transfer) or a distinct primitive? Its API shape (§8) is provisional until this is decided. | SPDK / Backend team | + +--- + +## Appendix A: `testfailover_types.go` + +```go +package v1alpha2 + +import ( + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + + "github.com/simplyblock/atlas/statemachine" +) + +// TestFailoverScope selects what a drill recovers. +// +kubebuilder:validation:Enum=Volume;Group +type TestFailoverScope string + +const ( + // TestFailoverScopeVolume recovers a single source volume. + TestFailoverScopeVolume TestFailoverScope = "Volume" + // TestFailoverScopeGroup recovers a consistency group from one point. + TestFailoverScopeGroup TestFailoverScope = "Group" +) + +// TestFailoverPhase is the coarse lifecycle phase of a drill. +// +kubebuilder:validation:Enum=Pending;Provisioning;Ready;Failed;TearingDown +type TestFailoverPhase string + +const ( + TestFailoverPhasePending TestFailoverPhase = "Pending" + TestFailoverPhaseProvisioning TestFailoverPhase = "Provisioning" + TestFailoverPhaseReady TestFailoverPhase = "Ready" + TestFailoverPhaseFailed TestFailoverPhase = "Failed" + TestFailoverPhaseTearingDown TestFailoverPhase = "TearingDown" +) + +// TestFailoverStep is one step of a running drill. The enum is the union of every +// phase's steps; which steps belong to which phase is declared by the graph rather +// than by this type. +// +kubebuilder:validation:Enum=ResolvingSource;ResolvingPoint;Shipping;Cloning;Placing;Releasing +type TestFailoverStep string + +const ( + TestFailoverStepResolvingSource TestFailoverStep = "ResolvingSource" + TestFailoverStepResolvingPoint TestFailoverStep = "ResolvingPoint" + TestFailoverStepShipping TestFailoverStep = "Shipping" + TestFailoverStepCloning TestFailoverStep = "Cloning" + TestFailoverStepPlacing TestFailoverStep = "Placing" + TestFailoverStepReleasing TestFailoverStep = "Releasing" +) + +// TestFailoverSpec is the request for one non-disruptive test-failover drill. +type TestFailoverSpec struct { + // Scope selects what the drill recovers. Immutable. + // +kubebuilder:validation:Required + // +k8s:immutable + Scope TestFailoverScope `json:"scope"` + + // SourceCluster is the OCM ManagedCluster the source runs on. The hub reads + // the source there through a ManagedClusterView. Immutable. + // +kubebuilder:validation:Required + // +k8s:immutable + SourceCluster string `json:"sourceCluster"` + + // SourceNamespace is the namespace of the source PVC on SourceCluster. + // Required for scope=Volume. Immutable. + // +optional + // +k8s:immutable + SourceNamespace string `json:"sourceNamespace,omitempty"` + + // SourceRef names the source on SourceCluster: a PersistentVolumeClaim in + // SourceNamespace (scope=Volume), or a consistency group (scope=Group). + // Immutable. + // +kubebuilder:validation:Required + // +k8s:immutable + SourceRef string `json:"sourceRef"` + + // BubbleCluster is the OCM ManagedCluster to recover onto: a DR target holding + // the replicated point, or another cluster. It must differ from SourceCluster; + // test-failover recovers onto a different cluster, never in place. Immutable. + // +kubebuilder:validation:Required + // +k8s:immutable + BubbleCluster string `json:"bubbleCluster"` + + // BubbleNamespace is the namespace on the bubble cluster where the recovered + // PVCs are created. Immutable. + // +kubebuilder:default=bubble + // +optional + // +k8s:immutable + BubbleNamespace string `json:"bubbleNamespace,omitempty"` + + // TTLSeconds is an optional maximum lifetime: the drill is torn down after it + // even without a delete, so a forgotten drill cannot hold a clone forever. + // +optional + // +kubebuilder:validation:Minimum=0 + TTLSeconds *int64 `json:"ttlSeconds,omitempty"` +} + +// TestFailoverClone is one recovered volume: the source it came from, the +// snapshot and clone the drill built, and the PVC placed on the bubble cluster. +type TestFailoverClone struct { + // SourceRef is the source volume (or group member) the recovered volume maps to. + SourceRef string `json:"sourceRef"` + // SourceHandle is the source volume's backend handle, read from its PV. + // +optional + SourceHandle string `json:"sourceHandle,omitempty"` + // SnapshotID is the recovery-point snapshot: the replicated snapshot already on + // the bubble cluster's backend that the clone is built from. + // +optional + SnapshotID string `json:"snapshotID,omitempty"` + // CloneID is the backend id of the writable clone. + // +optional + CloneID string `json:"cloneID,omitempty"` + // PVCName is the bound PVC in the bubble namespace on the bubble cluster. + // +optional + PVCName string `json:"pvcName,omitempty"` + // SizeBytes is the recovered volume's size. + // +optional + SizeBytes int64 `json:"sizeBytes,omitempty"` +} + +// TestFailoverReport is the evidence a drill produces. +type TestFailoverReport struct { + // BubbleCluster is the cluster the drill recovered onto. + // +optional + BubbleCluster string `json:"bubbleCluster,omitempty"` + // RecoveryPoint is the snapshot or group generation the drill recovered. + // +optional + RecoveryPoint string `json:"recoveryPoint,omitempty"` + // RecoveryPointTime is when that point was taken. + // +optional + RecoveryPointTime *metav1.Time `json:"recoveryPointTime,omitempty"` + // RecoveryPointAgeSeconds is the drill time minus the recovery-point time. + // +optional + RecoveryPointAgeSeconds int64 `json:"recoveryPointAgeSeconds,omitempty"` + // InvariantsHeld is true only when the source fingerprint taken before the + // drill matches the one taken at Ready. A Ready drill with this false is a + // defect. + // +optional + InvariantsHeld bool `json:"invariantsHeld,omitempty"` +} + +// TestFailoverStatus is the observed state of a drill. +type TestFailoverStatus struct { + // +optional + Phase TestFailoverPhase `json:"phase,omitempty"` + // +kubebuilder:validation:XValidation:rule="!has(self.state) || self.state in ['ResolvingSource','ResolvingPoint','Shipping','Cloning','Placing','Releasing']",message="unknown step" + // +optional + Step statemachine.KubeSnapshot `json:"step,omitempty"` + // +optional + Message string `json:"message,omitempty"` + // Triggered records that the current step's side effect was issued, so a + // restart does not repeat it. + // +optional + Triggered bool `json:"triggered,omitempty"` + // +optional + ObservedGeneration int64 `json:"observedGeneration,omitempty"` + // +optional + // +listType=map + // +listMapKey=sourceRef + Clones []TestFailoverClone `json:"clones,omitempty"` + // +optional + Report *TestFailoverReport `json:"report,omitempty"` + // +optional + StartedAt *metav1.Time `json:"startedAt,omitempty"` + // +optional + ReadyAt *metav1.Time `json:"readyAt,omitempty"` + // +optional + CompletedAt *metav1.Time `json:"completedAt,omitempty"` +} + +// +kubebuilder:object:root=true +// +kubebuilder:subresource:status +// +kubebuilder:resource:scope=Namespaced,shortName=tfo +// +kubebuilder:printcolumn:name="Scope",type=string,JSONPath=`.spec.scope` +// +kubebuilder:printcolumn:name="Source",type=string,JSONPath=`.spec.sourceRef` +// +kubebuilder:printcolumn:name="On",type=string,JSONPath=`.spec.sourceCluster` +// +kubebuilder:printcolumn:name="Bubble",type=string,JSONPath=`.spec.bubbleCluster` +// +kubebuilder:printcolumn:name="Phase",type=string,JSONPath=`.status.phase` +// +kubebuilder:printcolumn:name="Step",type=string,JSONPath=`.status.step.state` +// +kubebuilder:printcolumn:name="Age",type=date,JSONPath=`.metadata.creationTimestamp` + +// TestFailover is a one-way, non-disruptive test-failover drill. It recovers a +// source volume, or a consistency group, from a snapshot into an isolated +// namespace on a chosen cluster as bound PVCs, without touching the source: the +// recovery point is a snapshot and the result is a clone. The hub reads the +// source on its cluster and places the bubble on the recovery cluster through +// OCM, which must differ from the source cluster. Deleting the object reclaims +// the clones; the replicated recovery point is left alone. +type TestFailover struct { + metav1.TypeMeta `json:",inline"` + metav1.ObjectMeta `json:"metadata,omitempty"` + + Spec TestFailoverSpec `json:"spec,omitempty"` + Status TestFailoverStatus `json:"status,omitempty"` +} + +// +kubebuilder:object:root=true + +// TestFailoverList contains a list of TestFailover. +type TestFailoverList struct { + metav1.TypeMeta `json:",inline"` + metav1.ListMeta `json:"metadata,omitempty"` + Items []TestFailover `json:"items"` +} + +func init() { + SchemeBuilder.Register(&TestFailover{}, &TestFailoverList{}) +} +``` diff --git a/operator/docs/tests/test-plan-test-failover.md b/operator/docs/tests/test-plan-test-failover.md new file mode 100644 index 000000000..365467674 --- /dev/null +++ b/operator/docs/tests/test-plan-test-failover.md @@ -0,0 +1,280 @@ +# Test Plan: Non-Disruptive Test Failover + +Related design: [`designs/design-test-failover.md`](../designs/design-test-failover.md) +Harness: [`operator/internal/controller`](../../internal/controller) + +Scope: the operator, the CSI driver, and the Kubernetes surface of this +repository. Control-plane (`sbcli`), SPDK, and OCM behavior is a dependency, +faked at the boundary. See the `test-scenarios` skill. + +Scenario IDs are permanent: `U-` unit (no cluster, pure functions, fake +`client.Client`, mock HTTP), `I-` integration (full reconcile loop against +`envtest`, a mock backend, and a mock OCM), `E-` end-to-end (live clusters, real +data path), `M-` manual (needs failure injection or orchestration not yet +automated). Types are `Positive`, `Negative`, `Boundary`, `Regression`. The +`Test` column names the implementing function, or `—` when the scenario is not +yet covered. Every `—` also appears in §8. + +This design is `Draft` and nothing is implemented, so every `Test` cell is `—`. +The plan is the specification the implementation is written against, and §8 +carries the whole matrix as the gap list until the work lands. + +--- + +## 1. Unit Tests + +Pure helpers and controller methods in `testfailover_controller.go`, covered with +a fake `client.Client` and a mock control-plane HTTP server. Numbering runs +continuously across the groups. + +### Source resolution (design §4.1, §5.1) + +File: `internal/controller/testfailover_unit_test.go` + +| # | Scenario | Type | Test | +|------|------------------------------------------------------------------------------------------------------|----------|------| +| U-01 | `scope: Volume`: a `ManagedClusterView` projection of the source PVC and PV yields the volume handle | Positive | — | +| U-02 | `scope: Group`: `sourceRef` resolves to the group's member volumes, one recovered volume per member | Positive | — | +| U-03 | The view projects nothing for `sourceRef` → clean `Failed`, no panic | Negative | — | +| U-04 | `scope: Group` where the group has zero live members → `Failed`, nothing to recover | Boundary | — | + +### Recovery-point resolution (design §5.2) + +File: `internal/controller/testfailover_unit_test.go` + +| # | Scenario | Type | Test | +|------|------------------------------------------------------------------------------------------------------------------|----------|------| +| U-05 | `bubbleCluster` equals `sourceCluster` → `Failed`, rejected up front (recovery is onto a DIFFERENT cluster only) | Negative | — | +| U-06 | `bubbleCluster` is a DR target → latest replicated snapshot on that backend | Positive | — | +| U-07 | Group drill on a DR target → latest replicated generation, one snapshot per member | Positive | — | +| U-08 | DR target with no replicated point yet → `Failed` at `ResolvingPoint` | Negative | — | + +### Non-disruptiveness fingerprint (design §7.4) + +File: `internal/controller/testfailover_fingerprint_unit_test.go` + +| # | Scenario | Type | Test | +|------|------------------------------------------------------------------------------------------------------------|----------|------| +| U-09 | Source volume and its replication lag unchanged from `ResolvingSource` to `Ready` → `invariantsHeld: true` | Positive | — | +| U-10 | Source PVC rebound to a different volume between captures → drift detected | Negative | — | +| U-11 | A `ReplicationSlot` or VGR state change between captures → drift detected | Negative | — | + +### State machine and idempotency (design §6, §8) + +File: `internal/controller/testfailover_statemachine_unit_test.go` + +| # | Scenario | Type | Test | +|------|--------------------------------------------------------------------------------------|----------|------| +| U-12 | Each step advances to the next on its success condition | Positive | — | +| U-13 | A step whose deadline has passed moves the object to `Failed` | Boundary | — | +| U-14 | Deadline exactly at now is not yet expired, and now plus ε is (strict `>`) | Boundary | — | +| U-15 | Terminal `Ready` re-reconcile is a no-op | Positive | — | +| U-16 | Terminal `Failed` re-reconcile is a no-op | Positive | — | +| U-17 | The snapshot and clone idempotency key is `(test-id, source)`, stable across a retry | Positive | — | + +### Defaults and immutability (design §4.1, §9) + +File: `internal/controller/testfailover_unit_test.go` + +| # | Scenario | Type | Test | +|------|---------------------------------------------------------------------|----------|------| +| U-18 | `bubbleNamespace` defaults to `bubble` when unset | Boundary | — | +| U-19 | `ttlSeconds` unset → no auto-teardown scheduled | Boundary | — | +| U-20 | `ttlSeconds` set → teardown scheduled at creation time plus the TTL | Positive | — | + +--- + +## 2. Integration Tests + +Run the full controller reconcile loop against a real Kubernetes API via +`envtest`, a mock control-plane HTTP server that records call counts, and a mock +OCM (`ManagedClusterView` and `ManifestWork` with status feedback). + +### Source read and DR-target drill (design §5.1–§5.5, §6) + +| # | Scenario | Type | Test | +|------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|----------|------| +| I-01 | Create → `ResolvingSource` creates a `ManagedClusterView` on `sourceCluster`, reads the PVC handle, resolves the target's replicated snapshot (P0-4), clones it (P0-2), places PV and PVC, reaches `Ready` with a `BubbleReady` event | Positive | — | +| I-02 | `bubbleCluster` equals `sourceCluster` → `Failed` at `ResolvingSource`, message explains a same-cluster drill is unsupported | Negative | — | +| I-03 | `sourceCluster` is not a registered `ManagedCluster` → `Failed` at `ResolvingSource` | Negative | — | +| I-04 | The `ManagedClusterView` projects nothing for `sourceRef` → `Failed` at `ResolvingSource`, ref in `status.message` | Negative | — | +| I-05 | Clone returns 5xx → retried, no state advance, mock shows repeated calls | Negative | — | +| I-06 | Group drill: one group-consistent replicated generation → one PVC per member, all from that point | Positive | — | +| I-07 | Control plane unreachable (connection refused) → requeue with backoff, no partial state committed | Negative | — | + +### DR-target drill via OCM (design §5.5, §7.6) + +| # | Scenario | Type | Test | +|------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|----------|------| +| I-08 | `bubbleCluster` set → the recovery point resolves to the target's latest replicated snapshot (P0-4), the clone is built on the target backend, and the PV and PVC are delivered as a `ManifestWork` to that cluster | Positive | — | +| I-09 | `ManifestWork` status feedback reports the PVC `Bound` → the drill reaches `Ready` | Positive | — | +| I-10 | `bubbleCluster` is not a registered `ManagedCluster` → `Failed` at `Placing` | Negative | — | +| I-11 | Phase 1 drill where replication has landed nothing on the target yet (P0-4 404) → `Failed` at `ResolvingPoint` | Negative | — | +| I-12 | `ManifestWork` never reports Bound before its deadline → `Failed` at `Placing`, clone recorded for reclaim | Negative | — | + +### Restart safety (design §6, §7.4) + +| # | Scenario | Type | Test | +|------|--------------------------------------------------------------------------------------------------------------|----------|------| +| I-13 | Restart at `ResolvingSource` → resumes, does not create a second `ManagedClusterView` | Negative | — | +| I-14 | Restart at `Cloning` after the clone was issued → resumes, clone call count is 1 | Negative | — | +| I-15 | Restart at `Placing` after the `ManifestWork` was created → resumes, does not create a second `ManifestWork` | Negative | — | +| I-16 | Restart at `Releasing` during teardown → resumes the reclaim, does not double-reclaim | Negative | — | + +### Non-disruptiveness guard (design §7.4) + +| # | Scenario | Type | Test | +|------|---------------------------------------------------------------------------------------------------------------------------------|----------|------| +| I-17 | The source and its replication lag are unchanged before and after a `Ready` drill, and `invariantsHeld: true` | Positive | — | +| I-18 | Injected source or relationship change between start and `Ready` → `Failed`, `InvariantViolated` event, `invariantsHeld: false` | Negative | — | +| I-19 | The whole drill issues zero mutating calls against the source volume or its relationship | Positive | — | + +### Teardown and finalizer (design §5.7, §6) + +| # | Scenario | Type | Test | +|------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|----------|------| +| I-20 | Delete a `Ready` drill → `TearingDown` deletes the `ManifestWork` and `ManagedClusterView`, reclaims the clone (P0-3), leaves the replicated recovery point alone, removes the finalizer | Positive | — | +| I-21 | A DR-target drill that only resolved a replicated snapshot → teardown reclaims the clone but does NOT delete that snapshot | Positive | — | +| I-22 | Delete a `Failed` drill → teardown still runs, finalizer removed on the failure path | Negative | — | +| I-23 | Reclaim returns non-success → holds `TearingDown`, finalizer retained, `ReclaimPending` event | Negative | — | +| I-24 | Reclaim of an already-gone clone or snapshot returns success (404-as-success), teardown completes | Negative | — | +| I-25 | After teardown, no object carrying the drill's `test-id` label remains | Positive | — | + +### Admission and concurrency (design §4.1, §7.3) + +| # | Scenario | Type | Test | +|------|------------------------------------------------------------------------------------------------------------------------------------------------------------------|----------|------| +| I-26 | Patch any immutable spec field (`scope`, `sourceCluster`, `sourceNamespace`, `sourceRef`, `bubbleCluster`, `bubbleNamespace`) after creation → admission rejects | Negative | — | +| I-27 | A second drill on the same `(scope, sourceCluster, sourceRef, bubbleCluster)` while the first is active → refused | Negative | — | +| I-28 | Two drills on different sources run independently, neither blocks the other | Positive | — | +| I-29 | RBAC sufficiency: the controller creates a `ManagedClusterView` and a `ManifestWork` without a forbidden verb, and needs no PV/PVC or snapshot-API permission | Positive | — | + +--- + +## 3. E2E Tests + +Run against a live two-cluster DR setup (a source cluster and a DR target). +Data-path rows assert recovered-data correctness (a marker written to the +source), not merely that a PVC bound. + +### Same-cluster rejection (design §2 Non-Goals, §5) + +| # | Scenario | Type | Test | +|------|---------------------------------------------------------------------------------------------------------------------|----------|------| +| E-01 | `bubbleCluster` = the source's own cluster → the drill fails fast at `ResolvingSource`, nothing is cloned or placed | Negative | — | + +### DR-target drill (design §5.5) + +| # | Scenario | Type | Test | +|------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|----------|------| +| E-03 | `bubbleCluster` = the DR target: the replicated snapshot there is cloned, the bubble PVC is placed on the target via `ManifestWork`, a pod boots on the target, and the recovered marker matches the source | Positive | — | +| E-04 | Across a DR-target drill, the source and the running source-to-target replication lag are unchanged | Positive | — | +| E-05 | Group drill onto the target: every member PVC recovers from one replicated generation and each carries the marker (crash-consistent set) | Positive | — | +| E-06 | Teardown of a DR-target drill deletes the OCM objects and reclaims the clone, and the target shows no leaked volume, and the replicated snapshot it resolved is left intact | Positive | — | +| E-07 | Operator restart mid-drill → a single view, snapshot, clone, and `ManifestWork`, the drill still reaches `Ready` | Negative | — | + +--- + +## 4. Ship-to-Non-Target — Phase 2 (Planned) + +Testable only once P0-6 (on-demand shipping to a backend with no copy) exists +(design §5.6). Type and Test are decided when the phase is scoped. + +| # | Scenario | +|---------|---------------------------------------------------------------------------------------------------------------------------------------| +| U-P2-01 | Recovery-point resolution routes through the `Shipping` step only when the bubble backend holds no copy | +| I-P2-01 | `bubbleCluster` has no replica → `Shipping` calls P0-6, polls the returned handle, then clones on that backend and places the PVC | +| I-P2-02 | Ship handle poll exceeds its deadline → `Failed` at `Shipping` | +| E-P2-01 | Drill onto a third cluster with no replica: the point is shipped, cloned, and a pod boots on the recovered PVC with the marker intact | + +--- + +## 5. Manual Scenarios and Test Concepts + +### M-01 — Non-disruptiveness under sustained production I/O + +**Design reference:** design §7.4, §2 (Goals) + +**What to verify:** a drill run while the source application is actively writing, +and while source-to-target replication is running, disturbs neither. No write is +lost and the replication lag does not regress. + +**Test concept:** +1. Run fio in verify mode against the source workload, with replication to the target active. +2. While it runs, create a `TestFailover` with `bubbleCluster` = the target and let it reach `Ready`. +3. Assert `status.report.invariantsHeld` is true, the source is bound to the same volume, and the replication lag is within normal variance. +4. Tear the drill down and confirm fio still verifies with no errors. + +### M-02 — Replicated recovery point pruned before the clone + +**Design reference:** design §5.2, §10 (Failure Modes) + +**What to verify:** a drill degrades cleanly if the replicated snapshot it +resolved is pruned by retention before the clone runs. + +**Test concept:** +1. Create a DR-target `TestFailover` and let it resolve the replicated recovery point. +2. Prune that snapshot on the target before the clone call. +3. Assert the drill reports `Failed` at `Cloning` with a not-found reason, and that no clone was created. + +### M-03 — Leftover proof with `LIST_VOLUMES` disabled + +**Design reference:** design §5.7, Open Question 3 + +**What to verify:** teardown leaves no leaked backend clone on the bubble +backend, even though the CSI driver does not advertise `LIST_VOLUMES`, so the +Kubernetes-side label enumeration cannot be cross-checked through CSI. + +**Open question:** whether the label enumeration on Kubernetes objects is +sufficient, or a backend enumeration is required (design Open Question 3). + +**Test concept:** +1. Run and tear down a DR-target drill. +2. Enumerate backend volumes and snapshots directly on the target's storage cluster (out of band) and assert none carry the drill's `test-id`. + +--- + +## 6. Axis Coverage + +| Axis | Values covered | IDs | Not covered | +|------------------------------|--------------------------------------------------------------------------------------|------------------------------------|--------------------------------------------------------| +| Where the bubble runs | DR target; same-cluster rejected | I-01, I-08, E-03; U-05, I-02, E-01 | non-target cluster (Phase 2: I-P2-01, E-P2-01) | +| Source location | read on a managed cluster via ManagedClusterView | U-01, I-01 | source on the hub's own self-managed cluster | +| Scope | Volume, Group | U-01, U-02, I-06, E-05 | — | +| Recovery point | replicated (volume and group) | U-06, U-07 | shipped (Phase 2) | +| Cross-cluster transport | ManagedClusterView read, ManifestWork write | I-01, I-08 | — | +| Lifecycle / restart | mid-step restart (each step), delete mid-drill, TTL teardown | I-13, I-15, I-20, U-20 | control-plane restart mid-call | +| Control-plane / OCM response | 404, 5xx, connection refused, ManifestWork timeout, idempotent retry, 404-as-success | I-04, I-05, I-07, I-12, I-14, I-24 | partial-write then crash | +| Concurrency | same source, different sources | I-27, I-28 | spec mutated mid-drill (blocked by immutability, I-26) | +| Data correctness | recovered marker, source untouched, crash-consistent group | E-03, E-04, E-05 | recovered-data checksum under load (M-01) | + +--- + +## 7. Coverage Summary + +| Class | Scenarios | Covered | Not covered | +|-------------------|-----------|---------|------------------------------------| +| Unit | 20 | 0 | U-01 … U-20 | +| Integration | 29 | 0 | I-01 … I-29 | +| E2E | 6 | 0 | E-01, E-03 … E-07 | +| Manual | 3 | 0 | M-01 … M-03 | +| Phase 2 (planned) | 4 | 0 | U-P2-01, I-P2-01, I-P2-02, E-P2-01 | + +Nothing is covered: the design is `Draft` and the CRD and controller are unbuilt. +Phase 1 has no unbuilt backend dependency, so its scenarios become implementable +as soon as the controller exists; Phase 2 waits on P0-6. + +--- + +## 8. What Is Not Yet Covered + +| # | Gap | Reason | +|------------------------------------|----------------------------------------------|-----------------------------------------------------------------------------------------------------------------------------------------------| +| U-01 … U-20 | All unit scenarios | The `TestFailover` type and controller do not exist yet | +| I-01 … I-29 | All integration scenarios | The controller, its `envtest` suite, and the mock OCM are unwritten | +| E-01 … E-07 | All E2E scenarios | Needs a live two-cluster DR setup and the shipped feature | +| M-01 … M-03 | All manual scenarios | Need the shipped feature plus failure injection (snapshot delete, backend enumeration) | +| U-P2-01, I-P2-01, I-P2-02, E-P2-01 | Ship-to-non-target | Blocked on P0-6 (on-demand shipping to a backend with no copy), which does not exist | +| — | Source on the hub's own self-managed cluster | An edge of the topology axis. The primary path reads the source on a managed cluster, and the self-managed case is not exercised separately | +| — | Partial-write then crash on the backend | The backend verbs are the idempotency boundary, and asserting a mid-write crash needs backend fault injection this repository's harness lacks | +| — | Recovered-data checksum under sustained load | Covered as a manual concept (M-01), and not automatable without a live-cluster fio harness | diff --git a/operator/go.mod b/operator/go.mod index df4e93a08..2deeab4bc 100644 --- a/operator/go.mod +++ b/operator/go.mod @@ -30,6 +30,7 @@ require ( k8s.io/component-base v0.36.2 k8s.io/kube-openapi v0.0.0-20260317180543-43fb72c5454a k8s.io/utils v0.0.0-20260210185600-b8788abfbbc2 + open-cluster-management.io/api v0.15.0 sigs.k8s.io/controller-runtime v0.24.1 sigs.k8s.io/yaml v1.6.0 ) diff --git a/operator/go.sum b/operator/go.sum index 1a94375fa..5efdd335f 100644 --- a/operator/go.sum +++ b/operator/go.sum @@ -349,8 +349,9 @@ github.com/oasdiff/yaml3 v0.0.14/go.mod h1:csto2xfDjYccdUn/yw/bPjj/cYTdp6HtFA0J4 github.com/onsi/ginkgo v1.6.0/go.mod h1:lLunBs/Ym6LB5Z9jYTR76FiuTmxDTDusOGeTQH+WWjE= github.com/onsi/ginkgo v1.10.2/go.mod h1:lLunBs/Ym6LB5Z9jYTR76FiuTmxDTDusOGeTQH+WWjE= github.com/onsi/ginkgo v1.12.1/go.mod h1:zj2OWP4+oCPe1qIXoGWkgMRwljMUYCdkwsT2108oapk= -github.com/onsi/ginkgo v1.16.4 h1:29JGrr5oVBm5ulCWet69zQkzWipVXIol6ygQUe/EzNc= github.com/onsi/ginkgo v1.16.4/go.mod h1:dX+/inL/fNMqNlz0e9LfyB9TswhZpCVdJM/Z6Vvnwo0= +github.com/onsi/ginkgo v1.16.5 h1:8xi0RTUf59SOSfEtZMvwTvXYMzG4gV23XVHOZiXNtnE= +github.com/onsi/ginkgo v1.16.5/go.mod h1:+E8gABHa3K6zRBolWtd+ROzc/U5bkGt0FwiG042wbpU= github.com/onsi/ginkgo/v2 v2.1.3/go.mod h1:vw5CSIxN1JObi/U8gcbwft7ZxR2dgaR70JSE3/PpL4c= github.com/onsi/ginkgo/v2 v2.28.1 h1:S4hj+HbZp40fNKuLUQOYLDgZLwNUVn19N3Atb98NCyI= github.com/onsi/ginkgo/v2 v2.28.1/go.mod h1:CLtbVInNckU3/+gC8LzkGUb9oF+e8W8TdUsxPwvdOgE= @@ -681,6 +682,8 @@ k8s.io/streaming v0.36.2 h1:NSKthPPg9UFSKsRauVJUVGH2Dvn8fhKmY4qrMkw/p98= k8s.io/streaming v0.36.2/go.mod h1:z6fV3D+NVkoeqRMtWwlUZK6U17SY/LqNzOxWL6GyR/s= k8s.io/utils v0.0.0-20260210185600-b8788abfbbc2 h1:AZYQSJemyQB5eRxqcPky+/7EdBj0xi3g0ZcxxJ7vbWU= k8s.io/utils v0.0.0-20260210185600-b8788abfbbc2/go.mod h1:xDxuJ0whA3d0I4mf/C4ppKHxXynQ+fxnkmQH0vTHnuk= +open-cluster-management.io/api v0.15.0 h1:lRee1KOlGHZb2scTA7ff9E9Fxt2hJc7jpkHnaCbvkOU= +open-cluster-management.io/api v0.15.0/go.mod h1:9erZEWEn4bEqh0nIX2wA7f/s3KCuFycQdBrPrRzi0QM= oras.land/oras-go/v2 v2.6.1 h1:bonOEkjLfp8tt6qXWRRWP6p1F+9octchOf2EqnWB4Zs= oras.land/oras-go/v2 v2.6.1/go.mod h1:dhtFrFOuZuDtAVeZ9FUnaa5zfzplG3ZnFX9/uH1J/Yk= sigs.k8s.io/apiserver-network-proxy/konnectivity-client v0.34.0 h1:hSfpvjjTQXQY2Fol2CS0QHMNs/WI1MOSGzCm1KhM5ec= diff --git a/operator/internal/controller/testfailover_controller.go b/operator/internal/controller/testfailover_controller.go new file mode 100644 index 000000000..53cab8521 --- /dev/null +++ b/operator/internal/controller/testfailover_controller.go @@ -0,0 +1,1217 @@ +// The TestFailover controller drives one non-disruptive test-failover drill to a +// terminal phase and holds it there until the object is deleted. +// +// It coordinates from the hub: it reads the source on its cluster, resolves a +// recovery point on the recovery cluster's backend, clones it there, and places +// the clone as a bound PVC in an isolated namespace on the recovery cluster, +// without ever touching the source. Deleting the object reclaims what the drill +// created. See operator/docs/designs/design-test-failover.md. +// +// This file is built in slices: the state graph and the drill's lifecycle +// scaffolding land first, and each step's side effects (source read, snapshot, +// clone, placement, teardown) fill in behind the graph the reconcile already +// walks. + +package controller + +import ( + "context" + "crypto/sha256" + "encoding/json" + "fmt" + "net/http" + "strings" + "time" + + corev1 "k8s.io/api/core/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" + "k8s.io/apimachinery/pkg/api/resource" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" + "k8s.io/apimachinery/pkg/runtime" + "k8s.io/apimachinery/pkg/runtime/schema" + "k8s.io/client-go/tools/events" + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil" + logf "sigs.k8s.io/controller-runtime/pkg/log" + + "github.com/simplyblock/atlas/statemachine" + workv1 "open-cluster-management.io/api/work/v1" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/utils" + "github.com/simplyblock/simplyblock-operator/internal/webapi" +) + +// csiDriverName is the CSI driver the bubble's PersistentVolume adopts the clone +// through, on the recovery cluster. +const csiDriverName = "csi.simplyblock.io" + +// testFailoverIDLabel tags every object a drill creates, so teardown can +// enumerate and prove nothing was left behind. +const testFailoverIDLabel = "storage.simplyblock.io/test-id" + +// managedClusterViewGVK is the OCM read primitive the hub uses to project a +// managed cluster's object back to itself. It is driven unstructured to avoid a +// dependency on the multicloud-operators-foundation module that defines it. +var managedClusterViewGVK = schema.GroupVersionKind{ + Group: "view.open-cluster-management.io", + Version: "v1beta1", + Kind: "ManagedClusterView", +} + +// finalizerTestFailover holds the object until its drill is torn down, so a +// clone, a drill-taken snapshot, or a placed PVC is never orphaned by a delete +// that races the controller. +const finalizerTestFailover = "storage.simplyblock.io/testfailover-teardown" + +// testFailoverStepRequeue is how long the reconcile waits before re-entering a +// step that is still in progress. Named so the controller's cadence is tunable +// in one place (reconciler-patterns §7). +const testFailoverStepRequeue = 10 * time.Second + +// testFailoverStepTimeout bounds how long any one step may take. A step that +// blows it fails the drill rather than holding forever, and because the deadline +// lives in status it survives an operator restart (reconciler-patterns §3). +const testFailoverStepTimeout = 15 * time.Minute + +// testFailoverGraph is the drill's provisioning state graph: the ordered steps +// from reading the source to a placed, bound PVC. Every state arms a per-step +// deadline on entry. Teardown (Releasing) is not in this graph; it runs on the +// deletion path, off the finalizer, not as a forward transition. +func testFailoverGraph() statemachine.Config[simplyblockv1alpha2.TestFailoverStep] { + armDeadline := func(context.Context, simplyblockv1alpha2.TestFailoverStep, simplyblockv1alpha2.TestFailoverStep) (time.Duration, error) { + return testFailoverStepTimeout, nil + } + return statemachine.Config[simplyblockv1alpha2.TestFailoverStep]{ + Initial: simplyblockv1alpha2.TestFailoverStepResolvingSource, + States: map[simplyblockv1alpha2.TestFailoverStep]statemachine.StateDef[simplyblockv1alpha2.TestFailoverStep]{ + simplyblockv1alpha2.TestFailoverStepResolvingSource: { + To: []simplyblockv1alpha2.TestFailoverStep{simplyblockv1alpha2.TestFailoverStepResolvingPoint}, + OnEnter: armDeadline, + }, + simplyblockv1alpha2.TestFailoverStepResolvingPoint: { + To: []simplyblockv1alpha2.TestFailoverStep{ + simplyblockv1alpha2.TestFailoverStepShipping, + simplyblockv1alpha2.TestFailoverStepCloning, + }, + OnEnter: armDeadline, + }, + simplyblockv1alpha2.TestFailoverStepShipping: { + To: []simplyblockv1alpha2.TestFailoverStep{simplyblockv1alpha2.TestFailoverStepCloning}, + OnEnter: armDeadline, + }, + simplyblockv1alpha2.TestFailoverStepCloning: { + To: []simplyblockv1alpha2.TestFailoverStep{simplyblockv1alpha2.TestFailoverStepPlacing}, + OnEnter: armDeadline, + }, + // Placing is terminal in the provisioning graph: once the bubble PVC is + // bound, the drill's phase is Ready and it holds until deleted. + simplyblockv1alpha2.TestFailoverStepPlacing: {OnEnter: armDeadline}, + }, + } +} + +// TestFailoverReconciler reconciles a TestFailover object. +type TestFailoverReconciler struct { + client.Client + Scheme *runtime.Scheme + Recorder events.EventRecorder +} + +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=testfailovers,verbs=get;list;watch;create;update;patch;delete +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=testfailovers/status,verbs=get;update;patch +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=testfailovers/finalizers,verbs=update +// +kubebuilder:rbac:groups=work.open-cluster-management.io,resources=manifestworks,verbs=get;list;watch;create;update;patch;delete +// +kubebuilder:rbac:groups=view.open-cluster-management.io,resources=managedclusterviews,verbs=get;list;watch;create;update;patch;delete +// +kubebuilder:rbac:groups=events.k8s.io,resources=events,verbs=create;patch + +// Reconcile drives one drill: it ensures the finalizer, walks the provisioning +// graph to Ready, holds there, and tears the drill down on deletion. +func (r *TestFailoverReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) { + log := logf.FromContext(ctx) + + var tf simplyblockv1alpha2.TestFailover + if err := r.Get(ctx, req.NamespacedName, &tf); err != nil { + return ctrl.Result{}, client.IgnoreNotFound(err) + } + + if !tf.DeletionTimestamp.IsZero() { + return r.reconcileDeletion(ctx, &tf) + } + + // A drill that never got its finalizer gets it before any side effect, so + // teardown is guaranteed a chance to run. + if controllerutil.AddFinalizer(&tf, finalizerTestFailover) { + if err := r.Update(ctx, &tf); err != nil { + return ctrl.Result{}, err + } + return ctrl.Result{Requeue: true}, nil + } + + // A terminal phase re-reconciles to nothing. Ready holds until the object is + // deleted, and Failed stays as the record. The finalizer is retained in both, + // so teardown still runs on delete. + if tf.Status.Phase == simplyblockv1alpha2.TestFailoverPhaseReady || + tf.Status.Phase == simplyblockv1alpha2.TestFailoverPhaseFailed { + return ctrl.Result{}, nil + } + + log.V(1).Info("reconciling test-failover drill", "phase", tf.Status.Phase, "step", tf.Status.Step.State) + return r.advanceDrill(ctx, &tf) +} + +// advanceDrill walks the provisioning graph one step per reconcile. It builds +// the machine from status.step, so the position survives a restart, and never +// runs a step's side effect from construction. +func (r *TestFailoverReconciler) advanceDrill(ctx context.Context, tf *simplyblockv1alpha2.TestFailover) (ctrl.Result, error) { + machine, err := statemachine.NewFromSnapshot(ctx, testFailoverGraph(), + statemachine.FromKube[simplyblockv1alpha2.TestFailoverStep](tf.Status.Step)) + if err != nil { + return r.fail(ctx, tf, "invalid drill step: "+err.Error()) + } + defer machine.Close() + + // Nobody has reconciled this yet: enter the initial step. + if tf.Status.Step.State == "" { + return r.begin(ctx, tf, machine.CurrentState()) + } + + // A step that blew its deadline fails the drill rather than holding forever. + if machine.TimeoutReached() { + return r.fail(ctx, tf, "step "+string(machine.CurrentState())+" exceeded its deadline") + } + + switch machine.CurrentState() { + case simplyblockv1alpha2.TestFailoverStepResolvingSource: + return r.resolveSource(ctx, tf) + case simplyblockv1alpha2.TestFailoverStepResolvingPoint: + return r.resolvePoint(ctx, tf) + case simplyblockv1alpha2.TestFailoverStepCloning: + return r.cloneRecoveryPoint(ctx, tf) + case simplyblockv1alpha2.TestFailoverStepPlacing: + return r.placeBubble(ctx, tf) + default: + // Later slices implement the remaining steps; until then a step in + // progress holds rather than blocks, and re-enters on the requeue. + return r.hold(ctx, tf, "step "+string(machine.CurrentState())+" not yet implemented") + } +} + +// resolveSource reads the source on its cluster through ManagedClusterViews to +// learn the source volume's backend handle, then advances to ResolvingPoint. It +// is non-blocking: each view is created once and its result awaited across +// reconciles, so a restart re-enters rather than re-creates. +func (r *TestFailoverReconciler) resolveSource(ctx context.Context, tf *simplyblockv1alpha2.TestFailover) (ctrl.Result, error) { + // Test-failover recovers onto a DIFFERENT cluster (a DR target or another + // cluster), never in place: recovering a bubble into the source's own + // subsystem cannot be mounted alongside the live source, and the point is to + // rehearse the site that would take over. Reject a same-cluster drill up front. + if tf.Spec.BubbleCluster == tf.Spec.SourceCluster { + return r.fail(ctx, tf, "test-failover within the same cluster is not supported: bubbleCluster must differ from sourceCluster") + } + if tf.Spec.Scope == simplyblockv1alpha2.TestFailoverScopeGroup { + return r.resolveSourceGroup(ctx, tf) + } + + pvc, ready, err := r.projectedSource(ctx, tf, "src-pvc", "persistentvolumeclaims", tf.Spec.SourceRef, tf.Spec.SourceNamespace) + if err != nil { + return ctrl.Result{}, err + } + if !ready { + return r.hold(ctx, tf, "waiting for the source PVC projection from cluster "+tf.Spec.SourceCluster) + } + pvName, _, _ := unstructured.NestedString(pvc, "spec", "volumeName") + if pvName == "" { + return r.fail(ctx, tf, "source PVC "+tf.Spec.SourceRef+" is not bound to a volume") + } + + pv, ready, err := r.projectedSource(ctx, tf, "src-pv", "persistentvolumes", pvName, "") + if err != nil { + return ctrl.Result{}, err + } + if !ready { + return r.hold(ctx, tf, "waiting for the source PV projection from cluster "+tf.Spec.SourceCluster) + } + handle, _, _ := unstructured.NestedString(pv, "spec", "csi", "volumeHandle") + if handle == "" { + return r.fail(ctx, tf, "source PV "+pvName+" has no CSI volume handle") + } + srcAttrs, _, _ := unstructured.NestedStringMap(pv, "spec", "csi", "volumeAttributes") + bubbleVC := bubbleVolumeContext(srcAttrs) + fsType, _, _ := unstructured.NestedString(pv, "spec", "csi", "fsType") + + if err := r.transitionTo(ctx, tf, simplyblockv1alpha2.TestFailoverStepResolvingPoint, func(s *simplyblockv1alpha2.TestFailoverStatus) { + s.Clones = []simplyblockv1alpha2.TestFailoverClone{{ + SourceRef: tf.Spec.SourceRef, + SourceHandle: handle, + SourceVolumeContext: bubbleVC, + SourceFSType: fsType, + }} + s.Message = "resolved the source volume; resolving the recovery point" + }); err != nil { + return ctrl.Result{}, err + } + r.Recorder.Eventf(tf, nil, corev1.EventTypeNormal, "SourceResolved", "SourceResolved", + "resolved source volume %s on cluster %s", handle, tf.Spec.SourceCluster) + return ctrl.Result{Requeue: true}, nil +} + +// resolveSourceGroup resolves a consistency group's members into one clone slot +// each, then advances to ResolvingPoint. The members and their K8s identity come +// from the source cluster's backend (a group member carries only an lvol id); +// the shared class metadata (fsType and volumeAttributes, identical across +// members of one StorageClass) is read once from a representative member's PV +// through a ManagedClusterView. It is non-blocking: the representative views are +// created once and awaited across reconciles. +func (r *TestFailoverReconciler) resolveSourceGroup(ctx context.Context, tf *simplyblockv1alpha2.TestFailover) (ctrl.Result, error) { + srcUUID, err := r.resolveGroupSourceUUID(ctx, tf) + if err != nil { + return r.hold(ctx, tf, "resolving the source cluster's backend UUID: "+err.Error()) + } + + api := webapi.NewClient() + if secret, secErr := r.clusterSecret(ctx, tf.Namespace, tf.Spec.SourceCluster); secErr == nil && secret != "" { + ctx = webapi.WithBearerToken(ctx, secret) + } + + group, err := api.GetConsistencyGroupByName(ctx, srcUUID, tf.Spec.SourceRef) + if err != nil { + return ctrl.Result{}, err + } + if group == nil { + return r.fail(ctx, tf, "consistency group "+tf.Spec.SourceRef+" not found on cluster "+tf.Spec.SourceCluster) + } + memberIDs, err := api.GetConsistencyGroupMembers(ctx, srcUUID, group.UUID) + if err != nil { + return ctrl.Result{}, err + } + if len(memberIDs) == 0 { + return r.fail(ctx, tf, "consistency group "+tf.Spec.SourceRef+" has no members") + } + memberVols, err := api.ResolveMemberVolumes(ctx, srcUUID, memberIDs) + if err != nil { + return ctrl.Result{}, err + } + if len(memberVols) != len(memberIDs) { + return r.hold(ctx, tf, fmt.Sprintf("resolved %d of %d group members' source volumes; retrying", len(memberVols), len(memberIDs))) + } + + // One member's PV carries the class metadata every member shares, so a single + // projection serves the whole group. + rep := memberVols[memberIDs[0]] + pvc, ready, err := r.projectedSource(ctx, tf, "src-pvc", "persistentvolumeclaims", rep.PVCName, rep.PVCNamespace) + if err != nil { + return ctrl.Result{}, err + } + if !ready { + return r.hold(ctx, tf, "waiting for the source PVC projection for group member "+rep.PVCName) + } + pvName, _, _ := unstructured.NestedString(pvc, "spec", "volumeName") + if pvName == "" { + return r.fail(ctx, tf, "group member PVC "+rep.PVCName+" is not bound to a volume") + } + pv, ready, err := r.projectedSource(ctx, tf, "src-pv", "persistentvolumes", pvName, "") + if err != nil { + return ctrl.Result{}, err + } + if !ready { + return r.hold(ctx, tf, "waiting for the source PV projection for group member "+rep.PVCName) + } + srcAttrs, _, _ := unstructured.NestedStringMap(pv, "spec", "csi", "volumeAttributes") + bubbleVC := bubbleVolumeContext(srcAttrs) + fsType, _, _ := unstructured.NestedString(pv, "spec", "csi", "fsType") + + clones := make([]simplyblockv1alpha2.TestFailoverClone, 0, len(memberIDs)) + for _, id := range memberIDs { + v := memberVols[id] + clones = append(clones, simplyblockv1alpha2.TestFailoverClone{ + SourceRef: v.PVCName, + SourceHandle: srcUUID + ":" + v.PoolID + ":" + v.LvolID, + SourceFSType: fsType, + SourceVolumeContext: bubbleVC, + SizeBytes: v.Size, + }) + } + + if err := r.transitionTo(ctx, tf, simplyblockv1alpha2.TestFailoverStepResolvingPoint, func(s *simplyblockv1alpha2.TestFailoverStatus) { + s.Clones = clones + s.Message = fmt.Sprintf("resolved the consistency group's %d members; resolving the recovery point", len(clones)) + }); err != nil { + return ctrl.Result{}, err + } + r.Recorder.Eventf(tf, nil, corev1.EventTypeNormal, "SourceResolved", "SourceResolved", + "resolved consistency group %s (%d members) on cluster %s", tf.Spec.SourceRef, len(clones), tf.Spec.SourceCluster) + return ctrl.Result{Requeue: true}, nil +} + +// resolveGroupSourceUUID returns the backend UUID of the source cluster. It +// prefers the canonical resolution (a raw UUID, or a local StorageCluster named +// like the cluster), and falls back to the sole local StorageCluster when the +// hub is colocated on the source cluster, where the OCM cluster name does not +// match the StorageCluster name. +func (r *TestFailoverReconciler) resolveGroupSourceUUID(ctx context.Context, tf *simplyblockv1alpha2.TestFailover) (string, error) { + if utils.IsUUID(tf.Spec.SourceCluster) { + return tf.Spec.SourceCluster, nil + } + // StorageClusters live in the operator's namespace, not the drill's, so list + // cluster-wide rather than in the CR's namespace. + var clusters simplyblockv1alpha2.StorageClusterList + if err := r.List(ctx, &clusters); err != nil { + return "", err + } + // Prefer a StorageCluster named like the source cluster. + for i := range clusters.Items { + if clusters.Items[i].Name == tf.Spec.SourceCluster && clusters.Items[i].Status.UUID != "" { + return clusters.Items[i].Status.UUID, nil + } + } + // Fall back to the sole StorageCluster with a UUID: on a colocated hub the OCM + // cluster name does not match the StorageCluster name, but there is one. + uuid, ready := "", 0 + for i := range clusters.Items { + if clusters.Items[i].Status.UUID != "" { + uuid = clusters.Items[i].Status.UUID + ready++ + } + } + if ready == 1 { + return uuid, nil + } + return "", fmt.Errorf("no unique backend UUID for source cluster %q (%d Storage Clusters with a UUID)", tf.Spec.SourceCluster, ready) +} + +// projectedSource ensures a ManagedClusterView for one source object exists on +// the source cluster and returns the projected object once the view controller +// has fetched it. ready is false while the projection is still pending, which is +// the reconcile's cue to hold. +func (r *TestFailoverReconciler) projectedSource(ctx context.Context, tf *simplyblockv1alpha2.TestFailover, suffix, resourceKind, name, namespace string) (obj map[string]interface{}, ready bool, err error) { + viewName := testFailoverViewName(tf, suffix) + view := &unstructured.Unstructured{} + view.SetGroupVersionKind(managedClusterViewGVK) + getErr := r.Get(ctx, client.ObjectKey{Namespace: tf.Spec.SourceCluster, Name: viewName}, view) + if apierrors.IsNotFound(getErr) { + if createErr := r.Create(ctx, newManagedClusterView(tf, viewName, resourceKind, name, namespace)); createErr != nil { + return nil, false, createErr + } + return nil, false, nil + } + if getErr != nil { + return nil, false, getErr + } + result, found, nestedErr := unstructured.NestedMap(view.Object, "status", "result") + if nestedErr != nil || !found || len(result) == 0 { + return nil, false, nil + } + return result, true, nil +} + +// newManagedClusterView builds a view that asks the source cluster to project one +// object back to the hub. It lives in the source cluster's namespace on the hub +// and is labeled with the drill's test-id for teardown enumeration. +func newManagedClusterView(tf *simplyblockv1alpha2.TestFailover, name, resourceKind, targetName, targetNamespace string) *unstructured.Unstructured { + scope := map[string]interface{}{"resource": resourceKind, "name": targetName} + if targetNamespace != "" { + scope["namespace"] = targetNamespace + } + view := &unstructured.Unstructured{} + view.SetGroupVersionKind(managedClusterViewGVK) + view.SetNamespace(tf.Spec.SourceCluster) + view.SetName(name) + view.SetLabels(map[string]string{testFailoverIDLabel: string(tf.UID)}) + _ = unstructured.SetNestedMap(view.Object, scope, "spec", "scope") + return view +} + +// testFailoverViewName is a deterministic, bounded name for one of a drill's +// views, so a restart finds the existing view instead of creating a second. +func testFailoverViewName(tf *simplyblockv1alpha2.TestFailover, suffix string) string { + h := sha256.Sum256([]byte(tf.Namespace + "/" + tf.Name)) + return fmt.Sprintf("tfo-%x-%s", h[:6], suffix) +} + +// transitionTo validates the edge against the graph and persists the new step +// together with any status mutation. +func (r *TestFailoverReconciler) transitionTo(ctx context.Context, tf *simplyblockv1alpha2.TestFailover, to simplyblockv1alpha2.TestFailoverStep, mutate func(*simplyblockv1alpha2.TestFailoverStatus)) error { + machine, err := statemachine.NewFromSnapshot(ctx, testFailoverGraph(), + statemachine.FromKube[simplyblockv1alpha2.TestFailoverStep](tf.Status.Step)) + if err != nil { + return err + } + defer machine.Close() + if err := machine.TransitionTo(ctx, to); err != nil { + return err + } + snap := statemachine.ToKube(machine.Snapshot()) + return r.patchStatus(ctx, tf, func(s *simplyblockv1alpha2.TestFailoverStatus) { + s.Step = snap + if mutate != nil { + mutate(s) + } + }) +} + +// begin records that the drill has started and enters the initial step. It +// refuses to start a second active drill against the same source and bubble, so +// two drills cannot duplicate each other's snapshots and clones (design §7.3). +func (r *TestFailoverReconciler) begin(ctx context.Context, tf *simplyblockv1alpha2.TestFailover, initial simplyblockv1alpha2.TestFailoverStep) (ctrl.Result, error) { + if other, err := r.conflictingActiveDrill(ctx, tf); err != nil { + return ctrl.Result{}, err + } else if other != "" { + return r.fail(ctx, tf, "another active drill "+other+" is running for the same source and bubble cluster") + } + + now := metav1.Now() + deadline := metav1.NewTime(now.Add(testFailoverStepTimeout)) + if err := r.patchStatus(ctx, tf, func(s *simplyblockv1alpha2.TestFailoverStatus) { + s.Phase = simplyblockv1alpha2.TestFailoverPhaseProvisioning + s.Step = statemachine.KubeSnapshot{State: string(initial), Deadline: &deadline} + s.Message = "resolving the source" + if s.StartedAt == nil { + s.StartedAt = &now + } + }); err != nil { + return ctrl.Result{}, err + } + r.Recorder.Eventf(tf, nil, corev1.EventTypeNormal, "DrillStarted", "DrillStarted", "the test-failover drill started") + return ctrl.Result{Requeue: true}, nil +} + +// conflictingActiveDrill returns the name of another non-terminal drill against +// the same source and bubble, or empty when there is none. A drill that has +// Failed or is tearing down no longer holds resources and does not conflict. +func (r *TestFailoverReconciler) conflictingActiveDrill(ctx context.Context, tf *simplyblockv1alpha2.TestFailover) (string, error) { + var drills simplyblockv1alpha2.TestFailoverList + if err := r.List(ctx, &drills, client.InNamespace(tf.Namespace)); err != nil { + return "", err + } + for i := range drills.Items { + other := &drills.Items[i] + if other.UID == tf.UID || !other.DeletionTimestamp.IsZero() { + continue + } + if other.Status.Phase == simplyblockv1alpha2.TestFailoverPhaseFailed || + other.Status.Phase == simplyblockv1alpha2.TestFailoverPhaseTearingDown { + continue + } + if other.Spec.Scope == tf.Spec.Scope && + other.Spec.SourceRef == tf.Spec.SourceRef && + other.Spec.BubbleCluster == tf.Spec.BubbleCluster { + return other.Name, nil + } + } + return "", nil +} + +// hold keeps the drill on its current step and re-enters after the requeue. +func (r *TestFailoverReconciler) hold(ctx context.Context, tf *simplyblockv1alpha2.TestFailover, message string) (ctrl.Result, error) { + if err := r.patchStatus(ctx, tf, func(s *simplyblockv1alpha2.TestFailoverStatus) { + s.Message = message + }); err != nil { + return ctrl.Result{}, err + } + return ctrl.Result{RequeueAfter: testFailoverStepRequeue}, nil +} + +// fail records a terminal failure. The finalizer is retained, so teardown still +// runs on delete. +func (r *TestFailoverReconciler) fail(ctx context.Context, tf *simplyblockv1alpha2.TestFailover, message string) (ctrl.Result, error) { + now := metav1.Now() + if err := r.patchStatus(ctx, tf, func(s *simplyblockv1alpha2.TestFailoverStatus) { + s.Phase = simplyblockv1alpha2.TestFailoverPhaseFailed + s.Message = message + if s.CompletedAt == nil { + s.CompletedAt = &now + } + }); err != nil { + return ctrl.Result{}, err + } + r.Recorder.Eventf(tf, nil, corev1.EventTypeWarning, "DrillFailed", "DrillFailed", "%s", message) + return ctrl.Result{}, nil +} + +// reconcileDeletion tears the drill down and removes the finalizer only once +// every drill resource is confirmed gone. It reclaims in dependency order, and +// each step tolerates a not-found (a re-run after a partial teardown is safe). A +// reclaim that cannot be confirmed holds the object in TearingDown rather than +// clearing the finalizer and orphaning backend storage. +func (r *TestFailoverReconciler) reconcileDeletion(ctx context.Context, tf *simplyblockv1alpha2.TestFailover) (ctrl.Result, error) { + if !controllerutil.ContainsFinalizer(tf, finalizerTestFailover) { + return ctrl.Result{}, nil + } + if tf.Status.Phase != simplyblockv1alpha2.TestFailoverPhaseTearingDown { + _ = r.patchStatus(ctx, tf, func(s *simplyblockv1alpha2.TestFailoverStatus) { + s.Phase = simplyblockv1alpha2.TestFailoverPhaseTearingDown + s.Step = statemachine.KubeSnapshot{State: string(simplyblockv1alpha2.TestFailoverStepReleasing)} + s.Message = "tearing down the bubble" + }) + } + + // The ManifestWork's removal garbage-collects the PV and PVC on the recovery + // cluster, so it goes first. + if err := r.deleteManifestWork(ctx, tf); err != nil { + return r.reclaimPending(ctx, tf, "remove the bubble placement", err) + } + + // Reclaim every clone slot (one for a volume drill, one per member for a + // group drill). Each reclaim tolerates a not-found, so a re-run after a + // partial teardown is safe. + for i := range tf.Status.Clones { + clone := tf.Status.Clones[i] + if clone.CloneID != "" { + if err := r.reclaimClone(ctx, tf, clone.CloneID); err != nil { + return r.reclaimPending(ctx, tf, "reclaim the clone for "+clone.SourceRef, err) + } + } + // The recovery point is the replicated snapshot already on the target, + // resolved rather than created by the drill, so it is left alone. + } + + // The read-side views cost nothing to leave, but teardown proves no test-id + // object remains, so they go too. + if err := r.deleteView(ctx, tf, "src-pvc"); err != nil { + return r.reclaimPending(ctx, tf, "remove the source view", err) + } + if err := r.deleteView(ctx, tf, "src-pv"); err != nil { + return r.reclaimPending(ctx, tf, "remove the source view", err) + } + + controllerutil.RemoveFinalizer(tf, finalizerTestFailover) + if err := r.Update(ctx, tf); err != nil { + return ctrl.Result{}, err + } + return ctrl.Result{}, nil +} + +// reclaimPending records that teardown is holding on a reclaim it could not +// confirm, keeping the finalizer so nothing is orphaned. +func (r *TestFailoverReconciler) reclaimPending(ctx context.Context, tf *simplyblockv1alpha2.TestFailover, what string, cause error) (ctrl.Result, error) { + _ = r.patchStatus(ctx, tf, func(s *simplyblockv1alpha2.TestFailoverStatus) { + s.Message = "teardown is holding: could not " + what + ": " + cause.Error() + }) + r.Recorder.Eventf(tf, nil, corev1.EventTypeWarning, "ReclaimPending", "ReclaimPending", + "teardown could not %s: %v", what, cause) + return ctrl.Result{}, cause +} + +// deleteManifestWork removes the bubble's ManifestWork, tolerating an already-gone +// or already-deleting one. +func (r *TestFailoverReconciler) deleteManifestWork(ctx context.Context, tf *simplyblockv1alpha2.TestFailover) error { + var mw workv1.ManifestWork + err := r.Get(ctx, client.ObjectKey{Namespace: tf.Spec.BubbleCluster, Name: testFailoverManifestWorkName(tf)}, &mw) + if apierrors.IsNotFound(err) { + return nil + } + if err != nil { + return err + } + if !mw.DeletionTimestamp.IsZero() { + return nil + } + return client.IgnoreNotFound(r.Delete(ctx, &mw)) +} + +// deleteView removes one of the drill's ManagedClusterViews, tolerating a +// not-found. +func (r *TestFailoverReconciler) deleteView(ctx context.Context, tf *simplyblockv1alpha2.TestFailover, suffix string) error { + view := &unstructured.Unstructured{} + view.SetGroupVersionKind(managedClusterViewGVK) + view.SetNamespace(tf.Spec.SourceCluster) + view.SetName(testFailoverViewName(tf, suffix)) + return client.IgnoreNotFound(r.Delete(ctx, view)) +} + +// reclaimClone deletes the drill's clone from the recovery cluster's backend. +func (r *TestFailoverReconciler) reclaimClone(ctx context.Context, tf *simplyblockv1alpha2.TestFailover, handle string) error { + cluster, pool, vol, ok := splitHandle(handle) + if !ok { + return nil + } + apiClient := webapi.NewClient() + if secret, err := r.clusterSecret(ctx, tf.Namespace, tf.Spec.BubbleCluster); err == nil && secret != "" { + ctx = webapi.WithBearerToken(ctx, secret) + } + return backendDelete(ctx, apiClient, "reclaim clone", + fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes/%s", cluster, pool, vol)) +} + +// backendDelete issues an idempotent DELETE, treating a not-found as success. +func backendDelete(ctx context.Context, api *webapi.Client, what, endpoint string) error { + body, status, err := api.Do(ctx, http.MethodDelete, endpoint, nil) + if err != nil { + return fmt.Errorf("%s: %w", what, err) + } + if status == http.StatusNotFound || status < 300 { + return nil + } + return fmt.Errorf("%s: status %d: %s", what, status, string(body)) +} + +// sourceUnchanged re-reads the source PV projection and reports whether the +// source is still the same backend volume the drill recovered from. It is +// best-effort: when the projection cannot be re-read it does not fail the drill, +// but a projection that shows a different handle is a genuine invariant breach. +func (r *TestFailoverReconciler) sourceUnchanged(ctx context.Context, tf *simplyblockv1alpha2.TestFailover) bool { + if len(tf.Status.Clones) == 0 || tf.Status.Clones[0].SourceHandle == "" { + return true + } + pv, ok := r.readViewResult(ctx, tf, "src-pv") + if !ok { + return true + } + handle, _, _ := unstructured.NestedString(pv, "spec", "csi", "volumeHandle") + if handle == "" { + return true + } + return handle == tf.Status.Clones[0].SourceHandle +} + +// readViewResult reads an existing view's projected object without creating one. +func (r *TestFailoverReconciler) readViewResult(ctx context.Context, tf *simplyblockv1alpha2.TestFailover, suffix string) (map[string]interface{}, bool) { + view := &unstructured.Unstructured{} + view.SetGroupVersionKind(managedClusterViewGVK) + if err := r.Get(ctx, client.ObjectKey{Namespace: tf.Spec.SourceCluster, Name: testFailoverViewName(tf, suffix)}, view); err != nil { + return nil, false + } + result, found, err := unstructured.NestedMap(view.Object, "status", "result") + if err != nil || !found || len(result) == 0 { + return nil, false + } + return result, true +} + +// patchStatus applies mutate to the status and writes it, recording +// observedGeneration so a stale status can be told from a current one. +func (r *TestFailoverReconciler) patchStatus(ctx context.Context, tf *simplyblockv1alpha2.TestFailover, mutate func(*simplyblockv1alpha2.TestFailoverStatus)) error { + base := client.MergeFrom(tf.DeepCopy()) + mutate(&tf.Status) + tf.Status.ObservedGeneration = tf.Generation + return r.Status().Patch(ctx, tf, base) +} + +// replicatedSnapshotResult is the control plane's ReplicatedSnapshotDTO for the +// latest replicated snapshot on a DR-target backend. +type replicatedSnapshotResult struct { + SnapshotID string `json:"snapshot_id"` + ClusterID string `json:"cluster_id"` + PoolID string `json:"pool_id"` + CreatedAt time.Time `json:"created_at"` +} + +// resolvePoint resolves the recovery point on the bubble cluster's backend and +// advances to Cloning. The bubble is a DR target holding the latest replicated +// snapshot already there, so no data moves and nothing is triggered. It records +// the point as a CSI snapshot handle so Cloning is self-contained. +func (r *TestFailoverReconciler) resolvePoint(ctx context.Context, tf *simplyblockv1alpha2.TestFailover) (ctrl.Result, error) { + if tf.Spec.Scope == simplyblockv1alpha2.TestFailoverScopeGroup { + return r.resolvePointGroup(ctx, tf) + } + if len(tf.Status.Clones) == 0 || tf.Status.Clones[0].SourceHandle == "" { + return r.fail(ctx, tf, "internal: the source was not resolved before ResolvingPoint") + } + srcCluster, srcPool, srcLvol, ok := splitHandle(tf.Status.Clones[0].SourceHandle) + if !ok { + return r.fail(ctx, tf, "source handle is malformed: "+tf.Status.Clones[0].SourceHandle) + } + + apiClient := webapi.NewClient() + if secret, err := r.clusterSecret(ctx, tf.Namespace, tf.Spec.SourceCluster); err == nil && secret != "" { + ctx = webapi.WithBearerToken(ctx, secret) + } + + // The replicated snapshot is already on the DR target's backend. + dto, found, err := r.latestReplicatedSnapshot(ctx, apiClient, srcCluster, srcLvol) + if err != nil { + return ctrl.Result{}, err + } + if !found { + return r.fail(ctx, tf, "no replicated snapshot on the target for volume "+srcLvol+" yet") + } + snapCluster, snapPool, snapUUID := dto.ClusterID, dto.PoolID, dto.SnapshotID + var pointTime *metav1.Time + if !dto.CreatedAt.IsZero() { + pt := metav1.NewTime(dto.CreatedAt) + pointTime = &pt + } + if snapPool == "" { + snapPool = srcPool + } + pointHandle := snapCluster + ":" + snapPool + ":" + snapUUID + + if err := r.transitionTo(ctx, tf, simplyblockv1alpha2.TestFailoverStepCloning, func(s *simplyblockv1alpha2.TestFailoverStatus) { + s.Clones[0].SnapshotID = pointHandle + if s.Report == nil { + s.Report = &simplyblockv1alpha2.TestFailoverReport{} + } + s.Report.BubbleCluster = tf.Spec.BubbleCluster + s.Report.RecoveryPoint = snapUUID + s.Report.RecoveryPointTime = pointTime + s.Message = "resolved the recovery point; cloning it into the bubble" + }); err != nil { + return ctrl.Result{}, err + } + r.Recorder.Eventf(tf, nil, corev1.EventTypeNormal, "RecoveryPointResolved", "RecoveryPointResolved", + "recovery point %s on cluster %s", snapUUID, snapCluster) + return ctrl.Result{Requeue: true}, nil +} + +// resolvePointGroup resolves the group's one group-consistent recovery point on +// the target and records one snapshot handle per clone slot, then advances to +// Cloning. The point comes from the group's replication policy's latest +// generation, which the control plane returns only when every member has a +// snapshot at the same generation, so the recovered set is crash-consistent. +func (r *TestFailoverReconciler) resolvePointGroup(ctx context.Context, tf *simplyblockv1alpha2.TestFailover) (ctrl.Result, error) { + if len(tf.Status.Clones) == 0 { + return r.fail(ctx, tf, "internal: the group source was not resolved before ResolvingPoint") + } + srcUUID, err := r.resolveGroupSourceUUID(ctx, tf) + if err != nil { + return r.hold(ctx, tf, "resolving the source cluster's backend UUID: "+err.Error()) + } + + api := webapi.NewClient() + if secret, secErr := r.clusterSecret(ctx, tf.Namespace, tf.Spec.SourceCluster); secErr == nil && secret != "" { + ctx = webapi.WithBearerToken(ctx, secret) + } + + group, err := api.GetConsistencyGroupByName(ctx, srcUUID, tf.Spec.SourceRef) + if err != nil { + return ctrl.Result{}, err + } + if group == nil { + return r.fail(ctx, tf, "consistency group "+tf.Spec.SourceRef+" not found on cluster "+tf.Spec.SourceCluster) + } + + // The policy is read off the group: attach_group_policy stores it on the group + // record, and a group-first attach leaves the policy's own placement empty, so + // the recovery point is keyed on group.policy_id, not the policy list. + if group.PolicyID == "" { + return r.fail(ctx, tf, "no replication policy attached to group "+tf.Spec.SourceRef) + } + + groupSeq, members, found, err := api.LatestReplicatedGeneration(ctx, srcUUID, group.PolicyID) + if err != nil { + return ctrl.Result{}, err + } + if !found { + return r.fail(ctx, tf, "no replicated generation on the target for group "+tf.Spec.SourceRef+" yet") + } + if len(members) != len(tf.Status.Clones) { + return r.fail(ctx, tf, fmt.Sprintf("the group generation has %d members but %d were resolved; group membership changed mid-drill", len(members), len(tf.Status.Clones))) + } + + if err := r.transitionTo(ctx, tf, simplyblockv1alpha2.TestFailoverStepCloning, func(s *simplyblockv1alpha2.TestFailoverStatus) { + // Every member is at one generation, so any one-to-one assignment of the + // generation's snapshots to the clone slots yields a crash-consistent set; + // a precise source-to-target mapping is a later refinement. + for i := range members { + s.Clones[i].SnapshotID = members[i].ClusterID + ":" + members[i].PoolID + ":" + members[i].SnapshotID + } + if s.Report == nil { + s.Report = &simplyblockv1alpha2.TestFailoverReport{} + } + s.Report.BubbleCluster = tf.Spec.BubbleCluster + s.Report.RecoveryPoint = fmt.Sprintf("generation %d", groupSeq) + s.Message = fmt.Sprintf("resolved the group-consistent point (generation %d); cloning %d members", groupSeq, len(members)) + }); err != nil { + return ctrl.Result{}, err + } + r.Recorder.Eventf(tf, nil, corev1.EventTypeNormal, "RecoveryPointResolved", "RecoveryPointResolved", + "group-consistent generation %d on cluster %s (%d members)", groupSeq, tf.Spec.BubbleCluster, len(members)) + return ctrl.Result{Requeue: true}, nil +} + +// latestReplicatedSnapshot reads the latest replicated snapshot for a source +// volume on its DR target. found is false when replication has landed nothing. +func (r *TestFailoverReconciler) latestReplicatedSnapshot(ctx context.Context, api *webapi.Client, sourceClusterUUID, sourceLvolUUID string) (replicatedSnapshotResult, bool, error) { + endpoint := fmt.Sprintf("/api/v2/clusters/%s/replication/relationships/%s/latest-snapshot", sourceClusterUUID, sourceLvolUUID) + body, status, err := api.Do(ctx, http.MethodGet, endpoint, nil) + if status == http.StatusNotFound { + return replicatedSnapshotResult{}, false, nil + } + if err != nil || status >= 300 { + return replicatedSnapshotResult{}, false, requestError("resolve latest replicated snapshot", body, status, err) + } + var dto replicatedSnapshotResult + if err := json.Unmarshal(body, &dto); err != nil { + return replicatedSnapshotResult{}, false, fmt.Errorf("decode latest-snapshot: %w", err) + } + return dto, dto.SnapshotID != "", nil +} + +// cloneRecoveryPoint clones the resolved recovery point into a writable volume on +// the recovery cluster's backend and advances to Placing. The clone is built on +// the same backend the point lives on, so no data crosses a cluster boundary. +func (r *TestFailoverReconciler) cloneRecoveryPoint(ctx context.Context, tf *simplyblockv1alpha2.TestFailover) (ctrl.Result, error) { + if len(tf.Status.Clones) == 0 { + return r.fail(ctx, tf, "internal: no recovery point was resolved before Cloning") + } + + apiClient := webapi.NewClient() + if secret, err := r.clusterSecret(ctx, tf.Namespace, tf.Spec.BubbleCluster); err == nil && secret != "" { + ctx = webapi.WithBearerToken(ctx, secret) + } + + // One clone per slot, on the backend the point lives on, so no data crosses a + // cluster boundary. ensureClone is idempotent, so re-entry reuses any clone + // already built. A volume drill has one slot; a group drill has one per member. + handles := make([]string, len(tf.Status.Clones)) + sizes := make([]int64, len(tf.Status.Clones)) + for i := range tf.Status.Clones { + c := tf.Status.Clones[i] + if c.SnapshotID == "" { + return r.fail(ctx, tf, "internal: the recovery point was not resolved for member "+c.SourceRef) + } + snapCluster, snapPool, snapUUID, ok := splitHandle(c.SnapshotID) + if !ok { + return r.fail(ctx, tf, "recovery point handle is malformed: "+c.SnapshotID) + } + name := testFailoverCloneName(tf) + if tf.Spec.Scope == simplyblockv1alpha2.TestFailoverScopeGroup { + name = testFailoverMemberName(tf, c.SourceRef, "clone") + } + cloneUUID, sizeBytes, err := r.ensureClone(ctx, apiClient, snapCluster, snapPool, snapUUID, name) + if err != nil { + return ctrl.Result{}, err + } + handles[i] = snapCluster + ":" + snapPool + ":" + cloneUUID + sizes[i] = sizeBytes + } + + if err := r.transitionTo(ctx, tf, simplyblockv1alpha2.TestFailoverStepPlacing, func(s *simplyblockv1alpha2.TestFailoverStatus) { + for i := range s.Clones { + s.Clones[i].CloneID = handles[i] + if sizes[i] > 0 { + s.Clones[i].SizeBytes = sizes[i] + } + } + s.Message = fmt.Sprintf("cloned %d recovery point(s); placing on %s", len(s.Clones), tf.Spec.BubbleCluster) + }); err != nil { + return ctrl.Result{}, err + } + r.Recorder.Eventf(tf, nil, corev1.EventTypeNormal, "CloneBuilt", "CloneBuilt", + "cloned %d recovery point(s) on cluster %s, source untouched", len(tf.Status.Clones), tf.Spec.BubbleCluster) + return ctrl.Result{Requeue: true}, nil +} + +// ensureClone returns the id of the clone named name, cloning the snapshot if it +// does not exist yet. Ask-then-act: it lists the pool's volumes first, so a retry +// after a crash between clone and status-write reuses the clone rather than +// building a second (reconciler-patterns §4). +func (r *TestFailoverReconciler) ensureClone(ctx context.Context, api *webapi.Client, clusterUUID, poolID, snapUUID, name string) (string, int64, error) { + listEndpoint := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes", clusterUUID, poolID) + body, status, err := api.Do(ctx, http.MethodGet, listEndpoint, nil) + if err != nil || status >= 300 { + return "", 0, requestError("list volumes", body, status, err) + } + var existing []struct { + ID string `json:"id"` + Name string `json:"name"` + Size int64 `json:"size"` + } + _ = json.Unmarshal(body, &existing) + for _, v := range existing { + if v.Name == name { + return v.ID, v.Size, nil + } + } + + createEndpoint := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes?response_format=full", clusterUUID, poolID) + body, status, err = api.Do(ctx, http.MethodPost, createEndpoint, + map[string]interface{}{"name": name, "snapshot_id": snapUUID}) + if err != nil || status >= 300 { + return "", 0, requestError("clone snapshot", body, status, err) + } + var created struct { + ID string `json:"id"` + Size int64 `json:"size"` + } + if err := json.Unmarshal(body, &created); err != nil || created.ID == "" { + return "", 0, fmt.Errorf("clone snapshot: no volume id in the response") + } + return created.ID, created.Size, nil +} + +// placeBubble delivers the bubble PV and PVC to the recovery cluster through an +// OCM ManifestWork and marks the drill Ready once the PVC binds. It is +// non-blocking: the ManifestWork is created once and its bind status awaited +// across reconciles via its status feedback. +func (r *TestFailoverReconciler) placeBubble(ctx context.Context, tf *simplyblockv1alpha2.TestFailover) (ctrl.Result, error) { + if len(tf.Status.Clones) == 0 || tf.Status.Clones[0].CloneID == "" { + return r.fail(ctx, tf, "internal: the clone was not built before Placing") + } + + var mw workv1.ManifestWork + err := r.Get(ctx, client.ObjectKey{Namespace: tf.Spec.BubbleCluster, Name: testFailoverManifestWorkName(tf)}, &mw) + if apierrors.IsNotFound(err) { + desired, buildErr := r.bubbleManifestWork(tf) + if buildErr != nil { + return ctrl.Result{}, buildErr + } + if createErr := r.Create(ctx, desired); createErr != nil { + return ctrl.Result{}, createErr + } + return r.hold(ctx, tf, "placing the bubble PVC on cluster "+tf.Spec.BubbleCluster) + } + if err != nil { + return ctrl.Result{}, err + } + + if bound := boundBubblePVCs(&mw); bound < len(tf.Status.Clones) { + return r.hold(ctx, tf, fmt.Sprintf("waiting for the bubble PVCs to bind on cluster %s (%d of %d bound)", tf.Spec.BubbleCluster, bound, len(tf.Status.Clones))) + } + + // The non-disruptiveness guard: the drill only ever read the source, so it + // must still be the same volume it started from. A drill that cannot prove + // this is a defect, not a pass. + held := r.sourceUnchanged(ctx, tf) + now := metav1.Now() + if err := r.patchStatus(ctx, tf, func(s *simplyblockv1alpha2.TestFailoverStatus) { + if s.Report == nil { + s.Report = &simplyblockv1alpha2.TestFailoverReport{} + } + s.Report.InvariantsHeld = held + if held { + s.Phase = simplyblockv1alpha2.TestFailoverPhaseReady + s.Message = "bubble ready: the recovered PVC is bound on cluster " + tf.Spec.BubbleCluster + if s.ReadyAt == nil { + s.ReadyAt = &now + } + } else { + s.Phase = simplyblockv1alpha2.TestFailoverPhaseFailed + s.Message = "the source changed during the drill; the test was not non-disruptive" + if s.CompletedAt == nil { + s.CompletedAt = &now + } + } + }); err != nil { + return ctrl.Result{}, err + } + if held { + r.Recorder.Eventf(tf, nil, corev1.EventTypeNormal, "BubbleReady", "BubbleReady", + "the bubble PVC is bound on cluster %s", tf.Spec.BubbleCluster) + } else { + r.Recorder.Eventf(tf, nil, corev1.EventTypeWarning, "InvariantViolated", "InvariantViolated", + "the source changed during the drill") + } + return ctrl.Result{}, nil +} + +// bubbleVolumeContextStripKeys are the source PV volumeAttributes that must NOT +// be carried onto the bubble PV: they identify the SOURCE volume and its NVMe-oF +// target. The node plugin re-resolves the clone's own identity from the clone +// handle at stage time, so these are redundant on success; on a failed clone +// lookup, a stale source NQN/connections here would silently point the mount back +// at the source (reachable across clusters on a flat network), so they are +// dropped and staging fails safe instead. +var bubbleVolumeContextStripKeys = map[string]struct{}{ + "cluster_id": {}, "pool_name": {}, "nqn": {}, "connections": {}, + "model": {}, "name": {}, "uuid": {}, "nsId": {}, "targetLvolID": {}, +} + +// bubbleVolumeContext copies the source PV's volumeAttributes minus the identity +// keys above and the provisioner-injected keys (csi.storage.k8s.io/*, +// storage.kubernetes.io/*), leaving the class-level parameters the node plugin +// needs. It returns nil when nothing survives, which the driver tolerates. +func bubbleVolumeContext(src map[string]string) map[string]string { + if len(src) == 0 { + return nil + } + out := make(map[string]string, len(src)) + for k, v := range src { + if _, strip := bubbleVolumeContextStripKeys[k]; strip { + continue + } + if strings.HasPrefix(k, "csi.storage.k8s.io/") || strings.HasPrefix(k, "storage.kubernetes.io/") { + continue + } + out[k] = v + } + if len(out) == 0 { + return nil + } + return out +} + +// bubbleManifestWork wraps the bubble namespace, a static PersistentVolume bound +// to the clone, and its PersistentVolumeClaim, in a ManifestWork addressed to the +// recovery cluster, with a feedback rule that reports the PVC's bind phase back to +// the hub. +func (r *TestFailoverReconciler) bubbleManifestWork(tf *simplyblockv1alpha2.TestFailover) (*workv1.ManifestWork, error) { + ns := tf.Spec.BubbleNamespace + labels := map[string]string{testFailoverIDLabel: string(tf.UID)} + scName := "" + group := tf.Spec.Scope == simplyblockv1alpha2.TestFailoverScopeGroup + + namespace := &corev1.Namespace{ + TypeMeta: metav1.TypeMeta{Kind: "Namespace", APIVersion: "v1"}, + ObjectMeta: metav1.ObjectMeta{Name: ns, Labels: labels}, + } + // One namespace manifest plus a PV+PVC pair per clone slot. + manifests := make([]workv1.Manifest, 0, 1+2*len(tf.Status.Clones)) + raw, err := json.Marshal(namespace) + if err != nil { + return nil, fmt.Errorf("marshal bubble namespace: %w", err) + } + manifests = append(manifests, workv1.Manifest{RawExtension: runtime.RawExtension{Raw: raw}}) + + // One PV+PVC pair per clone slot, each reporting its own bind phase back to the + // hub. A volume drill has one; a group drill has one per member, all in the one + // bubble namespace so the recovered set is crash-consistent. + configs := make([]workv1.ManifestConfigOption, 0, len(tf.Status.Clones)) + for i := range tf.Status.Clones { + clone := tf.Status.Clones[i] + pvName := testFailoverPVName(tf) + pvcName := tf.Spec.SourceRef + if group { + pvName = testFailoverMemberName(tf, clone.SourceRef, "pv") + pvcName = clone.SourceRef + } + capacity := *resource.NewQuantity(clone.SizeBytes, resource.BinarySI) + + pv := &corev1.PersistentVolume{ + TypeMeta: metav1.TypeMeta{Kind: "PersistentVolume", APIVersion: "v1"}, + ObjectMeta: metav1.ObjectMeta{Name: pvName, Labels: labels}, + Spec: corev1.PersistentVolumeSpec{ + Capacity: corev1.ResourceList{corev1.ResourceStorage: capacity}, + AccessModes: []corev1.PersistentVolumeAccessMode{corev1.ReadWriteOnce}, + PersistentVolumeReclaimPolicy: corev1.PersistentVolumeReclaimRetain, + StorageClassName: scName, + ClaimRef: &corev1.ObjectReference{ + Kind: "PersistentVolumeClaim", APIVersion: "v1", Namespace: ns, Name: pvcName, + }, + PersistentVolumeSource: corev1.PersistentVolumeSource{ + CSI: &corev1.CSIPersistentVolumeSource{ + Driver: csiDriverName, + VolumeHandle: clone.CloneID, + FSType: clone.SourceFSType, + VolumeAttributes: clone.SourceVolumeContext, + }, + }, + }, + } + pvc := &corev1.PersistentVolumeClaim{ + TypeMeta: metav1.TypeMeta{Kind: "PersistentVolumeClaim", APIVersion: "v1"}, + ObjectMeta: metav1.ObjectMeta{Name: pvcName, Namespace: ns, Labels: labels}, + Spec: corev1.PersistentVolumeClaimSpec{ + AccessModes: []corev1.PersistentVolumeAccessMode{corev1.ReadWriteOnce}, + Resources: corev1.VolumeResourceRequirements{Requests: corev1.ResourceList{corev1.ResourceStorage: capacity}}, + StorageClassName: &scName, + VolumeName: pvName, + }, + } + for _, obj := range []client.Object{pv, pvc} { + raw, err := json.Marshal(obj) + if err != nil { + return nil, fmt.Errorf("marshal bubble manifest: %w", err) + } + manifests = append(manifests, workv1.Manifest{RawExtension: runtime.RawExtension{Raw: raw}}) + } + configs = append(configs, workv1.ManifestConfigOption{ + ResourceIdentifier: workv1.ResourceIdentifier{ + Group: "", Resource: "persistentvolumeclaims", Namespace: ns, Name: pvcName, + }, + FeedbackRules: []workv1.FeedbackRule{{ + Type: workv1.JSONPathsType, + JsonPaths: []workv1.JsonPath{{Name: "phase", Path: ".status.phase"}}, + }}, + }) + } + + return &workv1.ManifestWork{ + ObjectMeta: metav1.ObjectMeta{ + Name: testFailoverManifestWorkName(tf), + Namespace: tf.Spec.BubbleCluster, + Labels: labels, + }, + Spec: workv1.ManifestWorkSpec{ + Workload: workv1.ManifestsTemplate{Manifests: manifests}, + ManifestConfigs: configs, + }, + }, nil +} + +// boundBubblePVCs counts the bubble PVCs the ManifestWork's status feedback +// reports Bound. Placing is complete only when every clone's PVC is bound. +func boundBubblePVCs(mw *workv1.ManifestWork) int { + bound := 0 + for _, m := range mw.Status.ResourceStatus.Manifests { + if m.ResourceMeta.Resource != "persistentvolumeclaims" { + continue + } + for _, v := range m.StatusFeedbacks.Values { + if v.Name == "phase" && v.Value.String != nil && *v.Value.String == string(corev1.ClaimBound) { + bound++ + break + } + } + } + return bound +} + +// testFailoverPVName is the deterministic name of the drill's static +// PersistentVolume on the recovery cluster. +func testFailoverPVName(tf *simplyblockv1alpha2.TestFailover) string { + h := sha256.Sum256([]byte(tf.Namespace + "/" + tf.Name)) + return fmt.Sprintf("tfo-%x-pv", h[:6]) +} + +// testFailoverManifestWorkName is the deterministic name of the drill's +// ManifestWork in the recovery cluster's namespace on the hub. +func testFailoverManifestWorkName(tf *simplyblockv1alpha2.TestFailover) string { + h := sha256.Sum256([]byte(tf.Namespace + "/" + tf.Name)) + return fmt.Sprintf("tfo-%x-bubble", h[:6]) +} + +// clusterSecret returns the cluster secret the control plane authenticates a +// per-cluster call with, from the hub Secret simplyblock-cluster-. +func (r *TestFailoverReconciler) clusterSecret(ctx context.Context, namespace, clusterName string) (string, error) { + var secret corev1.Secret + if err := r.Get(ctx, client.ObjectKey{Namespace: namespace, Name: "simplyblock-cluster-" + clusterName}, &secret); err != nil { + return "", err + } + return string(secret.Data["secret"]), nil +} + +// testFailoverCloneName is the deterministic name of the clone a drill builds, so +// ask-then-act can find it on a retry. +func testFailoverCloneName(tf *simplyblockv1alpha2.TestFailover) string { + h := sha256.Sum256([]byte(tf.Namespace + "/" + tf.Name)) + return fmt.Sprintf("tfo-%x-clone", h[:6]) +} + +// testFailoverMemberName is the deterministic name of one group member's drill +// object (clone or PV), keyed by the drill and the member's source ref, so a +// group drill's members do not collide and each is re-findable on a retry. +func testFailoverMemberName(tf *simplyblockv1alpha2.TestFailover, memberRef, kind string) string { + h := sha256.Sum256([]byte(tf.Namespace + "/" + tf.Name)) + m := sha256.Sum256([]byte(memberRef)) + return fmt.Sprintf("tfo-%x-%x-%s", h[:6], m[:4], kind) +} + +// splitHandle splits a CSI handle "cluster:pool:uuid" into its three parts. +func splitHandle(handle string) (cluster, pool, uuid string, ok bool) { + parts := strings.Split(handle, ":") + if len(parts) != 3 || parts[0] == "" || parts[1] == "" || parts[2] == "" { + return "", "", "", false + } + return parts[0], parts[1], parts[2], true +} + +// requestError builds an error for a failed control-plane call. +func requestError(what string, body []byte, status int, err error) error { + if err != nil { + return fmt.Errorf("%s: %w", what, err) + } + return fmt.Errorf("%s: status %d: %s", what, status, string(body)) +} + +// SetupWithManager registers the controller. +func (r *TestFailoverReconciler) SetupWithManager(mgr ctrl.Manager) error { + return ctrl.NewControllerManagedBy(mgr). + For(&simplyblockv1alpha2.TestFailover{}). + Owns(&workv1.ManifestWork{}). + Named("testfailover"). + Complete(r) +} diff --git a/operator/internal/controller/testfailover_controller_unit_test.go b/operator/internal/controller/testfailover_controller_unit_test.go new file mode 100644 index 000000000..2df684ed3 --- /dev/null +++ b/operator/internal/controller/testfailover_controller_unit_test.go @@ -0,0 +1,1195 @@ +package controller + +import ( + "context" + "encoding/json" + "net/http" + "strings" + "testing" + "time" + + corev1 "k8s.io/api/core/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" + "k8s.io/apimachinery/pkg/types" + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil" + + "github.com/simplyblock/atlas/statemachine" + workv1 "open-cluster-management.io/api/work/v1" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +// The handles the drill fixtures resolve to at each step. The source lives on +// clusterA and the recovery bubble on clusterB, so bubble != source in the sample. +const ( + testSourceHandle = "clusterA:poolA:lvolX" + testSnapshotHandle = "clusterB:poolB:snapY" + testCloneHandle = "clusterB:poolB:cloneVol" + testFSTypeXFS = "xfs" + testFabricTCP = "tcp" + testSourceCluster = "ramen-cluster-a" +) + +// atResolvingPoint returns a drill seeded at ResolvingPoint with its source +// already resolved to testSourceHandle. +func atResolvingPoint() *simplyblockv1alpha2.TestFailover { + tf := atResolvingSource() + tf.Status.Step = statemachine.KubeSnapshot{State: string(simplyblockv1alpha2.TestFailoverStepResolvingPoint)} + tf.Status.Clones = []simplyblockv1alpha2.TestFailoverClone{{ + SourceRef: tf.Spec.SourceRef, + SourceHandle: testSourceHandle, + }} + return tf +} + +func newTestFailoverReconciler(t *testing.T, objects ...client.Object) (*TestFailoverReconciler, client.Client) { + t.Helper() + scheme := newTestScheme(t) + // The ManagedClusterView is driven unstructured, so the fake client needs its + // GVK (and list GVK) registered to create and read it. + scheme.AddKnownTypeWithName(managedClusterViewGVK, &unstructured.Unstructured{}) + listGVK := managedClusterViewGVK + listGVK.Kind += "List" + scheme.AddKnownTypeWithName(listGVK, &unstructured.UnstructuredList{}) + if err := workv1.Install(scheme); err != nil { + t.Fatalf("register work/v1 scheme: %v", err) + } + cl := newTestClient(t, scheme, + []client.Object{&simplyblockv1alpha2.TestFailover{}, &workv1.ManifestWork{}}, + objects...) + return &TestFailoverReconciler{Client: cl, Scheme: scheme, Recorder: &fakeRecorder{}}, cl +} + +// getView reads a ManagedClusterView the controller created. +func getView(t *testing.T, cl client.Client, namespace, name string) *unstructured.Unstructured { + t.Helper() + v := &unstructured.Unstructured{} + v.SetGroupVersionKind(managedClusterViewGVK) + if err := cl.Get(context.Background(), client.ObjectKey{Namespace: namespace, Name: name}, v); err != nil { + t.Fatalf("get ManagedClusterView %s/%s: %v", namespace, name, err) + } + return v +} + +// setViewResult simulates the OCM view controller having projected an object, +// by writing status.result onto the view. +func setViewResult(t *testing.T, cl client.Client, v *unstructured.Unstructured, result map[string]interface{}) { + t.Helper() + if err := unstructured.SetNestedMap(v.Object, result, "status", "result"); err != nil { + t.Fatalf("set status.result: %v", err) + } + if err := cl.Update(context.Background(), v); err != nil { + t.Fatalf("update view with result: %v", err) + } +} + +// atResolvingSource returns a drill seeded at the ResolvingSource step, the state +// the entry transition leaves it in. +func atResolvingSource() *simplyblockv1alpha2.TestFailover { + tf := sampleTestFailover() + tf.Finalizers = []string{finalizerTestFailover} + tf.Status.Phase = simplyblockv1alpha2.TestFailoverPhaseProvisioning + tf.Status.Step = statemachine.KubeSnapshot{State: string(simplyblockv1alpha2.TestFailoverStepResolvingSource)} + return tf +} + +func sampleTestFailover() *simplyblockv1alpha2.TestFailover { + return &simplyblockv1alpha2.TestFailover{ + ObjectMeta: metav1.ObjectMeta{Name: "drill-1", Namespace: "simplyblock"}, + Spec: simplyblockv1alpha2.TestFailoverSpec{ + Scope: simplyblockv1alpha2.TestFailoverScopeVolume, + SourceCluster: testSourceCluster, + SourceNamespace: "prod-app", + SourceRef: "postgres-data", + BubbleCluster: "ramen-cluster-b", + }, + } +} + +func testFailoverRequest(tf *simplyblockv1alpha2.TestFailover) ctrl.Request { + return ctrl.Request{NamespacedName: types.NamespacedName{Name: tf.Name, Namespace: tf.Namespace}} +} + +// TestFailoverAddsFinalizerThenEntersResolvingSource covers the entry into the +// drill: the first reconcile adds the teardown finalizer, and the next begins +// the drill at the initial step with the phase and bookkeeping set. +func TestFailoverAddsFinalizerThenEntersResolvingSource(t *testing.T) { + tf := sampleTestFailover() + r, cl := newTestFailoverReconciler(t, tf) + ctx := context.Background() + key := testFailoverRequest(tf).NamespacedName + + // First pass: the finalizer is added and nothing else has happened yet. + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("first reconcile: %v", err) + } + var got simplyblockv1alpha2.TestFailover + if err := cl.Get(ctx, key, &got); err != nil { + t.Fatalf("get after first reconcile: %v", err) + } + if !controllerutil.ContainsFinalizer(&got, finalizerTestFailover) { + t.Fatalf("finalizer %q was not added", finalizerTestFailover) + } + if got.Status.Phase != "" { + t.Errorf("phase = %q, want empty before the drill begins", got.Status.Phase) + } + + // Second pass: the drill begins at ResolvingSource. + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("second reconcile: %v", err) + } + if err := cl.Get(ctx, key, &got); err != nil { + t.Fatalf("get after second reconcile: %v", err) + } + if got.Status.Phase != simplyblockv1alpha2.TestFailoverPhaseProvisioning { + t.Errorf("phase = %q, want %q", got.Status.Phase, simplyblockv1alpha2.TestFailoverPhaseProvisioning) + } + if got.Status.Step.State != string(simplyblockv1alpha2.TestFailoverStepResolvingSource) { + t.Errorf("step = %q, want %q", got.Status.Step.State, simplyblockv1alpha2.TestFailoverStepResolvingSource) + } + if got.Status.ObservedGeneration != got.Generation { + t.Errorf("observedGeneration = %d, want %d", got.Status.ObservedGeneration, got.Generation) + } + if got.Status.StartedAt == nil { + t.Errorf("startedAt was not set") + } +} + +// TestFailoverDeletionClearsTheFinalizer covers that a deleted drill is torn +// down and does not hang on its finalizer. +func TestFailoverDeletionClearsTheFinalizer(t *testing.T) { + tf := sampleTestFailover() + tf.Finalizers = []string{finalizerTestFailover} + now := metav1.Now() + tf.DeletionTimestamp = &now + r, cl := newTestFailoverReconciler(t, tf) + ctx := context.Background() + + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("reconcile of a deleting drill: %v", err) + } + var got simplyblockv1alpha2.TestFailover + err := cl.Get(ctx, testFailoverRequest(tf).NamespacedName, &got) + if err == nil && controllerutil.ContainsFinalizer(&got, finalizerTestFailover) { + t.Fatalf("finalizer still present after deletion reconcile") + } +} + +// TestFailoverResolvingSourceResolvesHandleThenAdvances covers the full source +// read: the controller creates a ManagedClusterView for the PVC, then for its +// PV, and once both are projected it records the volume handle and advances to +// ResolvingPoint. Each projection arrives across reconciles, never blocking. +func TestFailoverResolvingSourceResolvesHandleThenAdvances(t *testing.T) { + tf := atResolvingSource() + r, cl := newTestFailoverReconciler(t, tf) + ctx := context.Background() + key := testFailoverRequest(tf).NamespacedName + + // Pass 1: the PVC view is created and the drill holds on its projection. + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("pass 1: %v", err) + } + pvcView := getView(t, cl, tf.Spec.SourceCluster, testFailoverViewName(tf, "src-pvc")) + var got simplyblockv1alpha2.TestFailover + if err := cl.Get(ctx, key, &got); err != nil { + t.Fatal(err) + } + if got.Status.Step.State != string(simplyblockv1alpha2.TestFailoverStepResolvingSource) { + t.Fatalf("step advanced before the source was projected: %q", got.Status.Step.State) + } + setViewResult(t, cl, pvcView, map[string]interface{}{ + "spec": map[string]interface{}{"volumeName": "pv-1"}, + }) + + // Pass 2: the PVC result is read, and the PV view is created. + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("pass 2: %v", err) + } + pvView := getView(t, cl, tf.Spec.SourceCluster, testFailoverViewName(tf, "src-pv")) + setViewResult(t, cl, pvView, map[string]interface{}{ + "spec": map[string]interface{}{ + "csi": map[string]interface{}{"volumeHandle": "clusterA:pool:lvolX"}, + }, + }) + + // Pass 3: the handle is read and the drill advances to ResolvingPoint. + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("pass 3: %v", err) + } + if err := cl.Get(ctx, key, &got); err != nil { + t.Fatal(err) + } + if got.Status.Step.State != string(simplyblockv1alpha2.TestFailoverStepResolvingPoint) { + t.Errorf("step = %q, want %q", got.Status.Step.State, simplyblockv1alpha2.TestFailoverStepResolvingPoint) + } + if len(got.Status.Clones) != 1 || got.Status.Clones[0].SourceHandle != "clusterA:pool:lvolX" { + t.Errorf("clones = %+v, want one with sourceHandle clusterA:pool:lvolX", got.Status.Clones) + } +} + +// Regression: 2026-09-29-testfailover-nil-volumecontext — the bubble PV must +// carry a VolumeContext or the node plugin panics staging it. The source PV's +// volumeAttributes are captured here, minus the identity and provisioner keys: +// the class params are needed to stage, but the identity keys would point a +// failed clone lookup back at the source, so they are dropped. +func TestFailoverResolvingSourceCapturesStrippedVolumeContext(t *testing.T) { + tf := atResolvingSource() + r, cl := newTestFailoverReconciler(t, tf) + ctx := context.Background() + key := testFailoverRequest(tf).NamespacedName + + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("pass 1: %v", err) + } + setViewResult(t, cl, getView(t, cl, tf.Spec.SourceCluster, testFailoverViewName(tf, "src-pvc")), + map[string]interface{}{"spec": map[string]interface{}{"volumeName": "pv-1"}}) + + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("pass 2: %v", err) + } + setViewResult(t, cl, getView(t, cl, tf.Spec.SourceCluster, testFailoverViewName(tf, "src-pv")), + map[string]interface{}{"spec": map[string]interface{}{"csi": map[string]interface{}{ + "volumeHandle": "clusterA:pool:lvolX", + "fsType": testFSTypeXFS, + "volumeAttributes": map[string]interface{}{ + // class params — kept + "tune2fs_reserved_blocks": "", + "fabric": testFabricTCP, + "qos_rw_iops": "0", + // identity — stripped (would mis-point a failed clone lookup) + "cluster_id": "clusterA", + "pool_name": "poolA", + "nqn": "nqn.source", + "connections": "[{\"ip\":\"10.0.0.1\",\"port\":4420}]", + "uuid": "lvolX", + "nsId": "1", + "model": "lvolX", + // provisioner-injected — stripped (stale source metadata) + "csi.storage.k8s.io/pv/name": "pv-1", + "storage.kubernetes.io/csiProvisionerIdentity": "x", + }, + }}}) + + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("pass 3: %v", err) + } + var got simplyblockv1alpha2.TestFailover + if err := cl.Get(ctx, key, &got); err != nil { + t.Fatal(err) + } + if len(got.Status.Clones) != 1 { + t.Fatalf("clones = %+v, want one", got.Status.Clones) + } + if got.Status.Clones[0].SourceFSType != testFSTypeXFS { + t.Errorf("sourceFSType = %q, want the source PV's xfs", got.Status.Clones[0].SourceFSType) + } + vc := got.Status.Clones[0].SourceVolumeContext + for _, k := range []string{"tune2fs_reserved_blocks", "fabric", "qos_rw_iops"} { + if _, ok := vc[k]; !ok { + t.Errorf("class param %q was dropped from the bubble VolumeContext: %+v", k, vc) + } + } + for _, k := range []string{ + "cluster_id", "pool_name", "nqn", "connections", "uuid", "nsId", "model", + "csi.storage.k8s.io/pv/name", "storage.kubernetes.io/csiProvisionerIdentity", + } { + if _, ok := vc[k]; ok { + t.Errorf("identity/provisioner key %q leaked into the bubble VolumeContext: %+v", k, vc) + } + } +} + +// TestFailoverResolvingSourceGroupResolvesMembers covers the group source path: +// the members come from the source cluster's backend (each carries only an lvol +// id, resolved to its PVC), one representative PV supplies the shared class +// metadata, and the drill advances to ResolvingPoint with one clone slot per +// member. +func TestFailoverResolvingSourceGroupResolvesMembers(t *testing.T) { + tf := sampleTestFailover() + tf.Finalizers = []string{finalizerTestFailover} + tf.Spec.Scope = simplyblockv1alpha2.TestFailoverScopeGroup + tf.Spec.SourceRef = "cg" + tf.Spec.SourceCluster = testSourceCluster // not a UUID: exercises the sole-StorageCluster fallback + tf.Status.Phase = simplyblockv1alpha2.TestFailoverPhaseProvisioning + tf.Status.Step = statemachine.KubeSnapshot{State: string(simplyblockv1alpha2.TestFailoverStepResolvingSource)} + + sc := &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{Name: "local-sc", Namespace: tf.Namespace}, + Status: simplyblockv1alpha2.StorageClusterStatus{UUID: "C"}, + } + r, cl := newTestFailoverReconciler(t, tf, sc) + ctx := context.Background() + key := testFailoverRequest(tf).NamespacedName + + srv := newAPIServer(t, func(w http.ResponseWriter, req *http.Request) { + w.Header().Set("Content-Type", "application/json") + p := req.URL.Path + switch { + case strings.HasSuffix(p, "/consistency-groups/") && req.URL.Query().Get("name") == "cg": + _, _ = w.Write([]byte(`[{"id":"g1","name":"cg","lvs_name":"lvs-a","node_id":"node-a"}]`)) + case strings.HasSuffix(p, "/consistency-groups/g1/members"): + _, _ = w.Write([]byte(`[{"lvol_id":"lvol-a"},{"lvol_id":"lvol-b"}]`)) + case strings.HasSuffix(p, "/storage-pools/"): + _, _ = w.Write([]byte(`[{"id":"pool-1"}]`)) + case strings.HasSuffix(p, "/storage-pools/pool-1/volumes"): + _, _ = w.Write([]byte(`[{"id":"lvol-a","pvc_name":"app/data-1","namespace":"nvme-ns","pool_id":null,"size":1073741824},` + + `{"id":"lvol-b","pvc_name":"app/data-2","namespace":"nvme-ns","pool_id":null,"size":1073741824}]`)) + default: + t.Errorf("unexpected request %s %s", req.Method, p) + w.WriteHeader(http.StatusInternalServerError) + } + }) + t.Setenv("SIMPLYBLOCK_WEBAPI_BASE_URL", srv.URL) + + // Pass 1: backend resolution done, the representative PVC view is created. + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("pass 1: %v", err) + } + setViewResult(t, cl, getView(t, cl, tf.Spec.SourceCluster, testFailoverViewName(tf, "src-pvc")), + map[string]interface{}{"spec": map[string]interface{}{"volumeName": "pv-1"}}) + + // Pass 2: the representative PV view is created. + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("pass 2: %v", err) + } + setViewResult(t, cl, getView(t, cl, tf.Spec.SourceCluster, testFailoverViewName(tf, "src-pv")), + map[string]interface{}{"spec": map[string]interface{}{"csi": map[string]interface{}{ + "volumeHandle": "C:pool-1:lvol-a", + "fsType": testFSTypeXFS, + "volumeAttributes": map[string]interface{}{ + "fabric": testFabricTCP, + "nqn": "nqn.source", // identity: must be stripped + }, + }}}) + + // Pass 3: the clones are built and the drill advances. + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("pass 3: %v", err) + } + var got simplyblockv1alpha2.TestFailover + if err := cl.Get(ctx, key, &got); err != nil { + t.Fatal(err) + } + if got.Status.Step.State != string(simplyblockv1alpha2.TestFailoverStepResolvingPoint) { + t.Fatalf("step = %q, want ResolvingPoint", got.Status.Step.State) + } + if len(got.Status.Clones) != 2 { + t.Fatalf("clones = %+v, want one per member (2)", got.Status.Clones) + } + want := map[string]string{"data-1": "C:pool-1:lvol-a", "data-2": "C:pool-1:lvol-b"} + for _, c := range got.Status.Clones { + if want[c.SourceRef] != c.SourceHandle { + t.Errorf("clone %q handle = %q, want %q", c.SourceRef, c.SourceHandle, want[c.SourceRef]) + } + if c.SourceFSType != testFSTypeXFS { + t.Errorf("clone %q fsType = %q, want xfs", c.SourceRef, c.SourceFSType) + } + if c.SourceVolumeContext["fabric"] != testFabricTCP { + t.Errorf("clone %q did not carry the shared class attrs: %+v", c.SourceRef, c.SourceVolumeContext) + } + if _, leaked := c.SourceVolumeContext["nqn"]; leaked { + t.Errorf("clone %q leaked the identity key nqn", c.SourceRef) + } + if c.SizeBytes != 1073741824 { + t.Errorf("clone %q size = %d, want 1Gi", c.SourceRef, c.SizeBytes) + } + } +} + +// TestFailoverResolvingSourceHoldsWithoutAProjection covers that a pending view +// holds the drill on its step rather than advancing or failing. +func TestFailoverResolvingSourceHoldsWithoutAProjection(t *testing.T) { + tf := atResolvingSource() + r, cl := newTestFailoverReconciler(t, tf) + ctx := context.Background() + + res, err := r.Reconcile(ctx, testFailoverRequest(tf)) + if err != nil { + t.Fatalf("reconcile: %v", err) + } + if res.RequeueAfter == 0 { + t.Errorf("expected a requeue while waiting for the projection") + } + var got simplyblockv1alpha2.TestFailover + if err := cl.Get(ctx, testFailoverRequest(tf).NamespacedName, &got); err != nil { + t.Fatal(err) + } + if got.Status.Phase == simplyblockv1alpha2.TestFailoverPhaseFailed { + t.Errorf("drill failed while merely waiting for a projection") + } + if got.Status.Step.State != string(simplyblockv1alpha2.TestFailoverStepResolvingSource) { + t.Errorf("step moved off ResolvingSource while waiting: %q", got.Status.Step.State) + } +} + +// TestFailoverResolvingSourceReuseViewOnRestart covers restart safety: a second +// reconcile before the projection arrives finds the existing view rather than +// creating a duplicate. +func TestFailoverResolvingSourceReuseViewOnRestart(t *testing.T) { + tf := atResolvingSource() + r, cl := newTestFailoverReconciler(t, tf) + ctx := context.Background() + + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("pass 1: %v", err) + } + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("pass 2 (restart): %v", err) + } + + list := &unstructured.UnstructuredList{} + gvk := managedClusterViewGVK + gvk.Kind += "List" + list.SetGroupVersionKind(gvk) + if err := cl.List(ctx, list, client.InNamespace(tf.Spec.SourceCluster)); err != nil { + t.Fatalf("list views: %v", err) + } + if len(list.Items) != 1 { + t.Errorf("got %d ManagedClusterViews, want exactly 1 (no duplicate on restart)", len(list.Items)) + } +} + +// TestFailoverResolvingSourceUnboundPVCFails covers that a source PVC bound to no +// volume is a terminal failure, not an endless hold. +func TestFailoverResolvingSourceUnboundPVCFails(t *testing.T) { + tf := atResolvingSource() + r, cl := newTestFailoverReconciler(t, tf) + ctx := context.Background() + + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("pass 1: %v", err) + } + pvcView := getView(t, cl, tf.Spec.SourceCluster, testFailoverViewName(tf, "src-pvc")) + setViewResult(t, cl, pvcView, map[string]interface{}{ + "spec": map[string]interface{}{}, // no volumeName: unbound + }) + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("pass 2: %v", err) + } + var got simplyblockv1alpha2.TestFailover + if err := cl.Get(ctx, testFailoverRequest(tf).NamespacedName, &got); err != nil { + t.Fatal(err) + } + if got.Status.Phase != simplyblockv1alpha2.TestFailoverPhaseFailed { + t.Errorf("phase = %q, want Failed for an unbound source PVC", got.Status.Phase) + } +} + +// TestFailoverResolvingPointDRTargetUsesReplicatedSnapshot covers the DR-target +// path: the recovery point is the latest replicated snapshot already on the +// target backend, read (never taken), and the drill advances to Cloning. +func TestFailoverResolvingPointDRTargetUsesReplicatedSnapshot(t *testing.T) { + tf := atResolvingPoint() + r, cl := newTestFailoverReconciler(t, tf) + ctx := context.Background() + + srv := newAPIServer(t, func(w http.ResponseWriter, req *http.Request) { + if req.Method == http.MethodGet && strings.HasSuffix(req.URL.Path, "/relationships/lvolX/latest-snapshot") { + w.Header().Set("Content-Type", "application/json") + _ = json.NewEncoder(w).Encode(map[string]interface{}{ + "snapshot_id": "snapY", "cluster_id": "clusterB", "pool_id": "poolB", + "created_at": "2026-09-29T00:00:00Z", + }) + return + } + t.Errorf("unexpected request %s %s", req.Method, req.URL.Path) + w.WriteHeader(http.StatusInternalServerError) + }) + t.Setenv("SIMPLYBLOCK_WEBAPI_BASE_URL", srv.URL) + + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("reconcile: %v", err) + } + var got simplyblockv1alpha2.TestFailover + if err := cl.Get(ctx, testFailoverRequest(tf).NamespacedName, &got); err != nil { + t.Fatal(err) + } + if got.Status.Step.State != string(simplyblockv1alpha2.TestFailoverStepCloning) { + t.Errorf("step = %q, want Cloning", got.Status.Step.State) + } + if got.Status.Clones[0].SnapshotID != "clusterB:poolB:snapY" { + t.Errorf("snapshot handle = %q, want clusterB:poolB:snapY", got.Status.Clones[0].SnapshotID) + } + if got.Status.Report == nil || got.Status.Report.RecoveryPoint != "snapY" { + t.Errorf("report.recoveryPoint not set to snapY: %+v", got.Status.Report) + } +} + +// TestFailoverResolvingPointGroupResolvesGeneration covers the group recovery +// point: the drill resolves the group's replication policy, reads its latest +// group-consistent generation, and records one target snapshot handle per clone +// slot before advancing to Cloning. +// +// Regression (2026-09-30): the drill inferred the policy from the policies list +// by matching the group's placement (group_lvs_name/group_node_id) and a +// consistency_group flag. A group attached with attach_group_policy sets +// group.policy_id and leaves both empty, so the heuristic matched nothing and the +// group drill failed at ResolvingPoint with "no consistency-group replication +// policy found." The policy id is read off the group DTO, which now carries it. +func TestFailoverResolvingPointGroupResolvesGeneration(t *testing.T) { + tf := sampleTestFailover() + tf.Finalizers = []string{finalizerTestFailover} + tf.Spec.Scope = simplyblockv1alpha2.TestFailoverScopeGroup + tf.Spec.SourceRef = "cg" + tf.Spec.SourceCluster = testSourceCluster + tf.Status.Phase = simplyblockv1alpha2.TestFailoverPhaseProvisioning + tf.Status.Step = statemachine.KubeSnapshot{State: string(simplyblockv1alpha2.TestFailoverStepResolvingPoint)} + tf.Status.Clones = []simplyblockv1alpha2.TestFailoverClone{ + {SourceRef: "data-1", SourceHandle: "C:pool-1:lvol-a", SourceFSType: testFSTypeXFS, SizeBytes: 1073741824}, + {SourceRef: "data-2", SourceHandle: "C:pool-1:lvol-b", SourceFSType: testFSTypeXFS, SizeBytes: 1073741824}, + } + + sc := &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{Name: "local-sc", Namespace: tf.Namespace}, + Status: simplyblockv1alpha2.StorageClusterStatus{UUID: "C"}, + } + r, cl := newTestFailoverReconciler(t, tf, sc) + ctx := context.Background() + + srv := newAPIServer(t, func(w http.ResponseWriter, req *http.Request) { + w.Header().Set("Content-Type", "application/json") + p := req.URL.Path + switch { + // The live group-first shape: the group carries its policy_id, and the + // policy itself exposes no placement and no consistency_group flag, so the + // drill must read the policy off the group rather than the policies list. + case strings.HasSuffix(p, "/consistency-groups/") && req.URL.Query().Get("name") == "cg": + _, _ = w.Write([]byte(`[{"id":"g1","name":"cg","lvs_name":"lvs-a","node_id":"node-a","policy_id":"p1"}]`)) + case strings.HasSuffix(p, "/replication/policies/p1/latest-generation"): + _, _ = w.Write([]byte(`{"group_seq":7,"members":[` + + `{"snapshot_id":"s1","cluster_id":"B","pool_id":"pb","lvol_id":"t1","size":1073741824,"group_seq":7},` + + `{"snapshot_id":"s2","cluster_id":"B","pool_id":"pb","lvol_id":"t2","size":1073741824,"group_seq":7}]}`)) + default: + t.Errorf("unexpected request %s %s", req.Method, p) + w.WriteHeader(http.StatusInternalServerError) + } + }) + t.Setenv("SIMPLYBLOCK_WEBAPI_BASE_URL", srv.URL) + + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("reconcile: %v", err) + } + var got simplyblockv1alpha2.TestFailover + if err := cl.Get(ctx, testFailoverRequest(tf).NamespacedName, &got); err != nil { + t.Fatal(err) + } + if got.Status.Step.State != string(simplyblockv1alpha2.TestFailoverStepCloning) { + t.Fatalf("step = %q, want Cloning", got.Status.Step.State) + } + if len(got.Status.Clones) != 2 || got.Status.Clones[0].SnapshotID != "B:pb:s1" || got.Status.Clones[1].SnapshotID != "B:pb:s2" { + t.Errorf("clone snapshot handles = %+v, want B:pb:s1 and B:pb:s2", got.Status.Clones) + } + if got.Status.Report == nil || got.Status.Report.RecoveryPoint != "generation 7" { + t.Errorf("report.recoveryPoint = %+v, want 'generation 7'", got.Status.Report) + } +} + +// TestFailoverResolvingPointDRTargetNoReplicaFails covers that a target with no +// replicated point yet is a terminal failure, not an endless hold. +func TestFailoverResolvingPointDRTargetNoReplicaFails(t *testing.T) { + tf := atResolvingPoint() + r, cl := newTestFailoverReconciler(t, tf) + ctx := context.Background() + + srv := newAPIServer(t, func(w http.ResponseWriter, _ *http.Request) { + w.WriteHeader(http.StatusNotFound) + }) + t.Setenv("SIMPLYBLOCK_WEBAPI_BASE_URL", srv.URL) + + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("reconcile: %v", err) + } + var got simplyblockv1alpha2.TestFailover + if err := cl.Get(ctx, testFailoverRequest(tf).NamespacedName, &got); err != nil { + t.Fatal(err) + } + if got.Status.Phase != simplyblockv1alpha2.TestFailoverPhaseFailed { + t.Errorf("phase = %q, want Failed when no replicated point exists", got.Status.Phase) + } +} + +// TestFailoverSameClusterIsRejected covers that a drill whose bubble is the +// source's own cluster fails immediately: test-failover recovers onto a DIFFERENT +// cluster, never in place. +func TestFailoverSameClusterIsRejected(t *testing.T) { + tf := sampleTestFailover() + tf.Finalizers = []string{finalizerTestFailover} + tf.Spec.BubbleCluster = tf.Spec.SourceCluster + tf.Status.Phase = simplyblockv1alpha2.TestFailoverPhaseProvisioning + tf.Status.Step = statemachine.KubeSnapshot{State: string(simplyblockv1alpha2.TestFailoverStepResolvingSource)} + r, cl := newTestFailoverReconciler(t, tf) + ctx := context.Background() + + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("reconcile: %v", err) + } + var got simplyblockv1alpha2.TestFailover + if err := cl.Get(ctx, testFailoverRequest(tf).NamespacedName, &got); err != nil { + t.Fatal(err) + } + if got.Status.Phase != simplyblockv1alpha2.TestFailoverPhaseFailed { + t.Fatalf("phase = %q, want Failed for a same-cluster drill; message=%q", got.Status.Phase, got.Status.Message) + } + if !strings.Contains(got.Status.Message, "same cluster") { + t.Errorf("message = %q, want it to explain same-cluster is unsupported", got.Status.Message) + } +} + +// atCloning returns a drill seeded at Cloning with its recovery point resolved. +func atCloning() *simplyblockv1alpha2.TestFailover { + tf := atResolvingPoint() + tf.Status.Step = statemachine.KubeSnapshot{State: string(simplyblockv1alpha2.TestFailoverStepCloning)} + tf.Status.Clones[0].SnapshotID = testSnapshotHandle + return tf +} + +// TestFailoverCloningClonesThenAdvances covers cloning the recovery point into a +// writable volume and advancing to Placing with the clone handle recorded. +func TestFailoverCloningClonesThenAdvances(t *testing.T) { + tf := atCloning() + r, cl := newTestFailoverReconciler(t, tf) + ctx := context.Background() + + srv := newAPIServer(t, func(w http.ResponseWriter, req *http.Request) { + switch { + case req.Method == http.MethodGet && strings.HasSuffix(req.URL.Path, "/storage-pools/poolB/volumes"): + w.Header().Set("Content-Type", "application/json") + _, _ = w.Write([]byte("[]")) + case req.Method == http.MethodPost && strings.HasSuffix(req.URL.Path, "/storage-pools/poolB/volumes"): + w.Header().Set("Content-Type", "application/json") + _ = json.NewEncoder(w).Encode(map[string]interface{}{"id": "cloneVol", "size": 1073741824}) + default: + t.Errorf("unexpected request %s %s", req.Method, req.URL.Path) + w.WriteHeader(http.StatusInternalServerError) + } + }) + t.Setenv("SIMPLYBLOCK_WEBAPI_BASE_URL", srv.URL) + + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("reconcile: %v", err) + } + var got simplyblockv1alpha2.TestFailover + if err := cl.Get(ctx, testFailoverRequest(tf).NamespacedName, &got); err != nil { + t.Fatal(err) + } + if got.Status.Step.State != string(simplyblockv1alpha2.TestFailoverStepPlacing) { + t.Errorf("step = %q, want Placing", got.Status.Step.State) + } + if got.Status.Clones[0].CloneID != "clusterB:poolB:cloneVol" { + t.Errorf("clone handle = %q, want clusterB:poolB:cloneVol", got.Status.Clones[0].CloneID) + } + if got.Status.Clones[0].SizeBytes != 1073741824 { + t.Errorf("clone size = %d, want 1073741824", got.Status.Clones[0].SizeBytes) + } +} + +// TestFailoverCloningReusesExistingClone covers ask-then-act idempotency: an +// existing clone with the drill's name is reused, and no second clone is built. +func TestFailoverCloningReusesExistingClone(t *testing.T) { + tf := atCloning() + r, cl := newTestFailoverReconciler(t, tf) + ctx := context.Background() + wantName := testFailoverCloneName(tf) + + srv := newAPIServer(t, func(w http.ResponseWriter, req *http.Request) { + if req.Method == http.MethodGet && strings.HasSuffix(req.URL.Path, "/storage-pools/poolB/volumes") { + w.Header().Set("Content-Type", "application/json") + _ = json.NewEncoder(w).Encode([]map[string]interface{}{ + {"id": "cloneExisting", "name": wantName, "size": 2048}, + }) + return + } + if req.Method == http.MethodPost { + t.Errorf("a second clone was built though one already existed: %s", req.URL.Path) + } + w.WriteHeader(http.StatusInternalServerError) + }) + t.Setenv("SIMPLYBLOCK_WEBAPI_BASE_URL", srv.URL) + + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("reconcile: %v", err) + } + var got simplyblockv1alpha2.TestFailover + if err := cl.Get(ctx, testFailoverRequest(tf).NamespacedName, &got); err != nil { + t.Fatal(err) + } + if got.Status.Clones[0].CloneID != "clusterB:poolB:cloneExisting" { + t.Errorf("clone handle = %q, want clusterB:poolB:cloneExisting", got.Status.Clones[0].CloneID) + } +} + +// TestFailoverCloningRetriesOnServerError covers that a transient control-plane +// error is retried (error returned, no state advance), not swallowed. +func TestFailoverCloningRetriesOnServerError(t *testing.T) { + tf := atCloning() + r, cl := newTestFailoverReconciler(t, tf) + ctx := context.Background() + + srv := newAPIServer(t, func(w http.ResponseWriter, _ *http.Request) { + w.WriteHeader(http.StatusInternalServerError) + }) + t.Setenv("SIMPLYBLOCK_WEBAPI_BASE_URL", srv.URL) + + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err == nil { + t.Fatalf("expected an error to trigger a retry on a 5xx") + } + var got simplyblockv1alpha2.TestFailover + if err := cl.Get(ctx, testFailoverRequest(tf).NamespacedName, &got); err != nil { + t.Fatal(err) + } + if got.Status.Step.State != string(simplyblockv1alpha2.TestFailoverStepCloning) { + t.Errorf("step = %q, want it to stay Cloning after a transient error", got.Status.Step.State) + } +} + +// atPlacing returns a drill seeded at Placing with its clone built. +func atPlacing() *simplyblockv1alpha2.TestFailover { + tf := atCloning() + tf.Status.Step = statemachine.KubeSnapshot{State: string(simplyblockv1alpha2.TestFailoverStepPlacing)} + tf.Status.Clones[0].CloneID = testCloneHandle + tf.Status.Clones[0].SizeBytes = 1073741824 + return tf +} + +func getManifestWork(t *testing.T, cl client.Client, namespace, name string) *workv1.ManifestWork { + t.Helper() + var mw workv1.ManifestWork + if err := cl.Get(context.Background(), client.ObjectKey{Namespace: namespace, Name: name}, &mw); err != nil { + t.Fatalf("get ManifestWork %s/%s: %v", namespace, name, err) + } + return &mw +} + +// markManifestWorkPVCBound simulates the recovery cluster's work-agent reporting +// the bubble PVC bound through the ManifestWork status feedback. +func markManifestWorkPVCBound(t *testing.T, cl client.Client, mw *workv1.ManifestWork) { + t.Helper() + bound := string(corev1.ClaimBound) + mw.Status.ResourceStatus.Manifests = []workv1.ManifestCondition{{ + ResourceMeta: workv1.ManifestResourceMeta{Resource: "persistentvolumeclaims"}, + StatusFeedbacks: workv1.StatusFeedbackResult{Values: []workv1.FeedbackValue{{ + Name: "phase", + Value: workv1.FieldValue{Type: workv1.String, String: &bound}, + }}}, + }} + if err := cl.Status().Update(context.Background(), mw); err != nil { + t.Fatalf("update ManifestWork status: %v", err) + } +} + +// TestFailoverPlacingDeliversManifestWorkThenReady covers the placement step: a +// ManifestWork carrying the bubble namespace, PV, and PVC is delivered to the +// recovery cluster, and the drill reaches Ready once the PVC binds. +func TestFailoverPlacingDeliversManifestWorkThenReady(t *testing.T) { + tf := atPlacing() + r, cl := newTestFailoverReconciler(t, tf) + ctx := context.Background() + key := testFailoverRequest(tf).NamespacedName + + // Pass 1: the ManifestWork is created and the drill holds on the bind. + res, err := r.Reconcile(ctx, testFailoverRequest(tf)) + if err != nil { + t.Fatalf("pass 1: %v", err) + } + if res.RequeueAfter == 0 { + t.Errorf("expected a requeue while waiting for the bubble PVC to bind") + } + mw := getManifestWork(t, cl, tf.Spec.BubbleCluster, testFailoverManifestWorkName(tf)) + if len(mw.Spec.Workload.Manifests) != 3 { + t.Errorf("ManifestWork carries %d manifests, want 3 (namespace, PV, PVC)", len(mw.Spec.Workload.Manifests)) + } + if len(mw.Spec.ManifestConfigs) != 1 || mw.Spec.ManifestConfigs[0].ResourceIdentifier.Name != tf.Spec.SourceRef { + t.Errorf("feedback rule not set on the bubble PVC %q: %+v", tf.Spec.SourceRef, mw.Spec.ManifestConfigs) + } + var got simplyblockv1alpha2.TestFailover + if err := cl.Get(ctx, key, &got); err != nil { + t.Fatal(err) + } + if got.Status.Phase == simplyblockv1alpha2.TestFailoverPhaseReady { + t.Errorf("drill reached Ready before the PVC was reported bound") + } + + // The work-agent reports the PVC bound. + markManifestWorkPVCBound(t, cl, mw) + + // Pass 2: the drill reaches Ready. + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("pass 2: %v", err) + } + if err := cl.Get(ctx, key, &got); err != nil { + t.Fatal(err) + } + if got.Status.Phase != simplyblockv1alpha2.TestFailoverPhaseReady { + t.Errorf("phase = %q, want Ready", got.Status.Phase) + } + if got.Status.ReadyAt == nil { + t.Errorf("readyAt was not set") + } +} + +// Regression: 2026-09-29-testfailover-nil-volumecontext — the bubble PV must +// carry the source's class-level VolumeContext (so the node plugin has a non-nil +// context to stage) while still pointing at the clone by handle. +func TestFailoverPlacingPVCarriesSourceVolumeContext(t *testing.T) { + tf := atPlacing() + tf.Status.Clones[0].SourceVolumeContext = map[string]string{ + "tune2fs_reserved_blocks": "", + "fabric": testFabricTCP, + } + tf.Status.Clones[0].SourceFSType = testFSTypeXFS + r, cl := newTestFailoverReconciler(t, tf) + ctx := context.Background() + + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("reconcile: %v", err) + } + mw := getManifestWork(t, cl, tf.Spec.BubbleCluster, testFailoverManifestWorkName(tf)) + // manifests are [namespace, PV, PVC]; decode the PV. + if len(mw.Spec.Workload.Manifests) != 3 { + t.Fatalf("ManifestWork carries %d manifests, want 3", len(mw.Spec.Workload.Manifests)) + } + var pv corev1.PersistentVolume + if err := json.Unmarshal(mw.Spec.Workload.Manifests[1].Raw, &pv); err != nil { + t.Fatalf("decode bubble PV manifest: %v", err) + } + if pv.Spec.CSI == nil { + t.Fatal("bubble PV has no CSI source") + } + if pv.Spec.CSI.VolumeHandle != tf.Status.Clones[0].CloneID { + t.Errorf("bubble PV points at %q, want the clone handle %q", pv.Spec.CSI.VolumeHandle, tf.Status.Clones[0].CloneID) + } + if pv.Spec.CSI.VolumeAttributes["fabric"] != testFabricTCP { + t.Errorf("bubble PV VolumeAttributes did not carry the source class params: %+v", pv.Spec.CSI.VolumeAttributes) + } + // The clone carries the source's filesystem; without this the node plugin + // defaults to ext4 and refuses to mount the XFS volume. + if pv.Spec.CSI.FSType != testFSTypeXFS { + t.Errorf("bubble PV fsType = %q, want the source's xfs", pv.Spec.CSI.FSType) + } +} + +// TestFailoverPlacingGroupDeliversAllMembersThenReady covers the group placement: +// one ManifestWork carries the namespace and a PV+PVC pair per member, and the +// drill reaches Ready only once every member's PVC binds. +func TestFailoverPlacingGroupDeliversAllMembersThenReady(t *testing.T) { + tf := sampleTestFailover() + tf.Finalizers = []string{finalizerTestFailover} + tf.Spec.Scope = simplyblockv1alpha2.TestFailoverScopeGroup + tf.Spec.SourceRef = "cg" + tf.Status.Phase = simplyblockv1alpha2.TestFailoverPhaseProvisioning + tf.Status.Step = statemachine.KubeSnapshot{State: string(simplyblockv1alpha2.TestFailoverStepPlacing)} + tf.Status.Clones = []simplyblockv1alpha2.TestFailoverClone{ + {SourceRef: "data-1", SourceHandle: "C:pool-1:lvol-a", SnapshotID: "B:pb:s1", CloneID: "B:pb:c1", SourceFSType: testFSTypeXFS, SizeBytes: 1073741824}, + {SourceRef: "data-2", SourceHandle: "C:pool-1:lvol-b", SnapshotID: "B:pb:s2", CloneID: "B:pb:c2", SourceFSType: testFSTypeXFS, SizeBytes: 1073741824}, + } + r, cl := newTestFailoverReconciler(t, tf) + ctx := context.Background() + key := testFailoverRequest(tf).NamespacedName + + // Pass 1: the ManifestWork is created carrying ns + 2*(PV,PVC) = 5 manifests + // and one feedback config per member PVC; the drill holds until both bind. + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("pass 1: %v", err) + } + mw := getManifestWork(t, cl, tf.Spec.BubbleCluster, testFailoverManifestWorkName(tf)) + if len(mw.Spec.Workload.Manifests) != 5 { + t.Errorf("ManifestWork carries %d manifests, want 5 (namespace + 2*(PV,PVC))", len(mw.Spec.Workload.Manifests)) + } + if len(mw.Spec.ManifestConfigs) != 2 { + t.Errorf("ManifestWork has %d feedback configs, want one per member (2)", len(mw.Spec.ManifestConfigs)) + } + var got simplyblockv1alpha2.TestFailover + if err := cl.Get(ctx, key, &got); err != nil { + t.Fatal(err) + } + if got.Status.Phase == simplyblockv1alpha2.TestFailoverPhaseReady { + t.Errorf("drill reached Ready before any PVC bound") + } + + // Only one member bound: still not Ready. + bound := string(corev1.ClaimBound) + pvcBound := func(name string) workv1.ManifestCondition { + return workv1.ManifestCondition{ + ResourceMeta: workv1.ManifestResourceMeta{Resource: "persistentvolumeclaims", Name: name}, + StatusFeedbacks: workv1.StatusFeedbackResult{Values: []workv1.FeedbackValue{{Name: "phase", Value: workv1.FieldValue{Type: workv1.String, String: &bound}}}}, + } + } + mw.Status.ResourceStatus.Manifests = []workv1.ManifestCondition{pvcBound("data-1")} + if err := cl.Status().Update(ctx, mw); err != nil { + t.Fatalf("update MW status (one bound): %v", err) + } + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("pass 2: %v", err) + } + if err := cl.Get(ctx, key, &got); err != nil { + t.Fatal(err) + } + if got.Status.Phase == simplyblockv1alpha2.TestFailoverPhaseReady { + t.Errorf("drill reached Ready with only one of two member PVCs bound") + } + + // Both bound: Ready. + mw.Status.ResourceStatus.Manifests = []workv1.ManifestCondition{pvcBound("data-1"), pvcBound("data-2")} + if err := cl.Status().Update(ctx, mw); err != nil { + t.Fatalf("update MW status (both bound): %v", err) + } + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("pass 3: %v", err) + } + if err := cl.Get(ctx, key, &got); err != nil { + t.Fatal(err) + } + if got.Status.Phase != simplyblockv1alpha2.TestFailoverPhaseReady { + t.Errorf("phase = %q, want Ready once both member PVCs are bound", got.Status.Phase) + } +} + +// TestFailoverPlacingReuseManifestWorkOnRestart covers restart safety: a second +// reconcile before the PVC binds finds the existing ManifestWork, not a duplicate. +func TestFailoverPlacingReuseManifestWorkOnRestart(t *testing.T) { + tf := atPlacing() + r, cl := newTestFailoverReconciler(t, tf) + ctx := context.Background() + + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("pass 1: %v", err) + } + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("pass 2 (restart): %v", err) + } + var list workv1.ManifestWorkList + if err := cl.List(ctx, &list, client.InNamespace(tf.Spec.BubbleCluster)); err != nil { + t.Fatalf("list ManifestWorks: %v", err) + } + if len(list.Items) != 1 { + t.Errorf("got %d ManifestWorks, want exactly 1 (no duplicate on restart)", len(list.Items)) + } +} + +// deletingReadyDrill returns a Ready drill with a clone recorded, being deleted. +func deletingReadyDrill() *simplyblockv1alpha2.TestFailover { + tf := atPlacing() + tf.Status.Phase = simplyblockv1alpha2.TestFailoverPhaseReady + tf.Status.Clones[0].SnapshotID = "clusterA:poolA:snapS" + now := metav1.Now() + tf.DeletionTimestamp = &now + return tf +} + +func srcPVView(tf *simplyblockv1alpha2.TestFailover, handle string) *unstructured.Unstructured { + v := &unstructured.Unstructured{} + v.SetGroupVersionKind(managedClusterViewGVK) + v.SetNamespace(tf.Spec.SourceCluster) + v.SetName(testFailoverViewName(tf, "src-pv")) + _ = unstructured.SetNestedMap(v.Object, map[string]interface{}{ + "spec": map[string]interface{}{"csi": map[string]interface{}{"volumeHandle": handle}}, + }, "status", "result") + return v +} + +// TestFailoverTeardownReclaimsThenClearsFinalizer covers that deleting a drill +// reclaims the clone, removes the ManifestWork, and only then clears the +// finalizer. The recovery point is a replicated snapshot the drill only resolved, +// so it is left alone. +func TestFailoverTeardownReclaimsThenClearsFinalizer(t *testing.T) { + tf := deletingReadyDrill() + mw := &workv1.ManifestWork{ObjectMeta: metav1.ObjectMeta{ + Name: testFailoverManifestWorkName(tf), Namespace: tf.Spec.BubbleCluster, + }} + r, cl := newTestFailoverReconciler(t, tf, mw) + ctx := context.Background() + + var reclaimedClone bool + srv := newAPIServer(t, func(w http.ResponseWriter, req *http.Request) { + switch { + case req.Method == http.MethodDelete && strings.Contains(req.URL.Path, "/volumes/cloneVol"): + reclaimedClone = true + w.WriteHeader(http.StatusNoContent) + case req.Method == http.MethodDelete && strings.Contains(req.URL.Path, "/snapshots/"): + t.Errorf("teardown deleted a snapshot it only resolved: %s", req.URL.Path) + w.WriteHeader(http.StatusInternalServerError) + default: + t.Errorf("unexpected request %s %s", req.Method, req.URL.Path) + w.WriteHeader(http.StatusInternalServerError) + } + }) + t.Setenv("SIMPLYBLOCK_WEBAPI_BASE_URL", srv.URL) + + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("teardown reconcile: %v", err) + } + if !reclaimedClone { + t.Errorf("the clone was not reclaimed") + } + err := cl.Get(ctx, testFailoverRequest(tf).NamespacedName, &simplyblockv1alpha2.TestFailover{}) + if err == nil || !apierrors.IsNotFound(err) { + t.Errorf("drill still present after teardown (finalizer not cleared): %v", err) + } + if err := cl.Get(ctx, client.ObjectKey{Namespace: tf.Spec.BubbleCluster, Name: mw.Name}, &workv1.ManifestWork{}); !apierrors.IsNotFound(err) { + t.Errorf("ManifestWork still present after teardown: %v", err) + } +} + +// TestFailoverTeardownTolersatesAlreadyGoneResources covers idempotency: a 404 +// from every reclaim is treated as success, so a re-run after a partial teardown +// still completes. +func TestFailoverTeardownToleratesAlreadyGone(t *testing.T) { + tf := deletingReadyDrill() + r, cl := newTestFailoverReconciler(t, tf) // no ManifestWork seeded + ctx := context.Background() + + srv := newAPIServer(t, func(w http.ResponseWriter, _ *http.Request) { + w.WriteHeader(http.StatusNotFound) + }) + t.Setenv("SIMPLYBLOCK_WEBAPI_BASE_URL", srv.URL) + + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("teardown reconcile: %v", err) + } + err := cl.Get(ctx, testFailoverRequest(tf).NamespacedName, &simplyblockv1alpha2.TestFailover{}) + if err == nil || !apierrors.IsNotFound(err) { + t.Errorf("drill still present after teardown of already-gone resources: %v", err) + } +} + +// TestFailoverTeardownHoldsWhenReclaimFails covers that a reclaim that cannot be +// confirmed holds the object with its finalizer rather than orphaning storage. +func TestFailoverTeardownHoldsWhenReclaimFails(t *testing.T) { + tf := deletingReadyDrill() + r, cl := newTestFailoverReconciler(t, tf) + ctx := context.Background() + + srv := newAPIServer(t, func(w http.ResponseWriter, _ *http.Request) { + w.WriteHeader(http.StatusInternalServerError) + }) + t.Setenv("SIMPLYBLOCK_WEBAPI_BASE_URL", srv.URL) + + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err == nil { + t.Fatalf("expected an error to hold teardown when a reclaim fails") + } + var got simplyblockv1alpha2.TestFailover + if err := cl.Get(ctx, testFailoverRequest(tf).NamespacedName, &got); err != nil { + t.Fatalf("drill was removed despite a failed reclaim: %v", err) + } + if !controllerutil.ContainsFinalizer(&got, finalizerTestFailover) { + t.Errorf("finalizer was cleared despite a failed reclaim") + } + if got.Status.Phase != simplyblockv1alpha2.TestFailoverPhaseTearingDown { + t.Errorf("phase = %q, want TearingDown while holding", got.Status.Phase) + } +} + +// TestFailoverPlacingFailsWhenSourceChanged covers the non-disruptiveness guard: +// if the source's projected volume handle differs at Ready, the drill fails +// rather than reporting a passing, non-disruptive test. +func TestFailoverPlacingFailsWhenSourceChanged(t *testing.T) { + tf := atPlacing() // SourceHandle is clusterA:poolA:lvolX + view := srcPVView(tf, "clusterA:poolA:DIFFERENT") + mw := &workv1.ManifestWork{ObjectMeta: metav1.ObjectMeta{ + Name: testFailoverManifestWorkName(tf), Namespace: tf.Spec.BubbleCluster, + }} + r, cl := newTestFailoverReconciler(t, tf, view, mw) + ctx := context.Background() + + markManifestWorkPVCBound(t, cl, getManifestWork(t, cl, tf.Spec.BubbleCluster, mw.Name)) + + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("reconcile: %v", err) + } + var got simplyblockv1alpha2.TestFailover + if err := cl.Get(ctx, testFailoverRequest(tf).NamespacedName, &got); err != nil { + t.Fatal(err) + } + if got.Status.Phase != simplyblockv1alpha2.TestFailoverPhaseFailed { + t.Errorf("phase = %q, want Failed when the source changed", got.Status.Phase) + } + if got.Status.Report == nil || got.Status.Report.InvariantsHeld { + t.Errorf("invariantsHeld = true, want false when the source changed") + } +} + +// TestFailoverPlacingConfirmsInvariantHeld covers the passing guard: an unchanged +// source projection yields Ready with invariantsHeld true. +func TestFailoverPlacingConfirmsInvariantHeld(t *testing.T) { + tf := atPlacing() + view := srcPVView(tf, "clusterA:poolA:lvolX") // same as SourceHandle + mw := &workv1.ManifestWork{ObjectMeta: metav1.ObjectMeta{ + Name: testFailoverManifestWorkName(tf), Namespace: tf.Spec.BubbleCluster, + }} + r, cl := newTestFailoverReconciler(t, tf, view, mw) + ctx := context.Background() + + markManifestWorkPVCBound(t, cl, getManifestWork(t, cl, tf.Spec.BubbleCluster, mw.Name)) + + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("reconcile: %v", err) + } + var got simplyblockv1alpha2.TestFailover + if err := cl.Get(ctx, testFailoverRequest(tf).NamespacedName, &got); err != nil { + t.Fatal(err) + } + if got.Status.Phase != simplyblockv1alpha2.TestFailoverPhaseReady { + t.Errorf("phase = %q, want Ready", got.Status.Phase) + } + if got.Status.Report == nil || !got.Status.Report.InvariantsHeld { + t.Errorf("invariantsHeld = false, want true for an unchanged source") + } +} + +// TestFailoverRefusesSecondDrillForSameSource covers the concurrency guard: a +// second drill against the same source and bubble as an active one is refused. +func TestFailoverRefusesSecondDrillForSameSource(t *testing.T) { + existing := atResolvingSource() // active (Provisioning) on the sample source/bubble + existing.Name = "drill-existing" + existing.UID = "uid-existing" + + second := sampleTestFailover() // same source/bubble, not yet started + second.Name = "drill-second" + second.UID = "uid-second" + second.Finalizers = []string{finalizerTestFailover} + + r, cl := newTestFailoverReconciler(t, existing, second) + ctx := context.Background() + + if _, err := r.Reconcile(ctx, testFailoverRequest(second)); err != nil { + t.Fatalf("reconcile: %v", err) + } + var got simplyblockv1alpha2.TestFailover + if err := cl.Get(ctx, testFailoverRequest(second).NamespacedName, &got); err != nil { + t.Fatal(err) + } + if got.Status.Phase != simplyblockv1alpha2.TestFailoverPhaseFailed { + t.Errorf("phase = %q, want Failed for a conflicting second drill", got.Status.Phase) + } +} + +// TestFailoverStepDeadlineFailsTheDrill covers that a step that blew its deadline +// fails the drill rather than holding forever. +func TestFailoverStepDeadlineFailsTheDrill(t *testing.T) { + tf := atResolvingSource() + past := metav1.NewTime(time.Now().Add(-time.Hour)) + tf.Status.Step = statemachine.KubeSnapshot{ + State: string(simplyblockv1alpha2.TestFailoverStepResolvingSource), + Deadline: &past, + } + r, cl := newTestFailoverReconciler(t, tf) + ctx := context.Background() + + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("reconcile: %v", err) + } + var got simplyblockv1alpha2.TestFailover + if err := cl.Get(ctx, testFailoverRequest(tf).NamespacedName, &got); err != nil { + t.Fatal(err) + } + if got.Status.Phase != simplyblockv1alpha2.TestFailoverPhaseFailed { + t.Errorf("phase = %q, want Failed after the step deadline passed", got.Status.Phase) + } + if !strings.Contains(got.Status.Message, "deadline") { + t.Errorf("message = %q, want it to mention the deadline", got.Status.Message) + } +} diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_testfailovers.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_testfailovers.yaml new file mode 100644 index 000000000..d68cf6ad7 --- /dev/null +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_testfailovers.yaml @@ -0,0 +1,305 @@ +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition +metadata: + annotations: + controller-gen.kubebuilder.io/version: v0.21.0 + name: testfailovers.storage.simplyblock.io +spec: + group: storage.simplyblock.io + names: + kind: TestFailover + listKind: TestFailoverList + plural: testfailovers + shortNames: + - tfo + singular: testfailover + scope: Namespaced + versions: + - additionalPrinterColumns: + - jsonPath: .spec.scope + name: Scope + type: string + - jsonPath: .spec.sourceRef + name: Source + type: string + - jsonPath: .spec.sourceCluster + name: "On" + type: string + - jsonPath: .spec.bubbleCluster + name: Bubble + type: string + - jsonPath: .status.phase + name: Phase + type: string + - jsonPath: .status.step.state + name: Step + type: string + - jsonPath: .status.message + name: Message + priority: 1 + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha2 + schema: + openAPIV3Schema: + description: |- + TestFailover is a one-way, non-disruptive test-failover drill. It recovers a + source volume, or a consistency group, from a snapshot into an isolated + namespace on a chosen cluster as bound PVCs, without touching the source. The + hub reads the source on its cluster and places the bubble on the recovery + cluster through OCM. It runs to a terminal phase, or holds Ready until it is + deleted, and deletion reclaims the clones and any snapshots the drill took. + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: |- + TestFailoverSpec is the request for one non-disruptive test-failover drill. + + The source is named by where it runs and what it is, so the hub can find it + without anyone extracting a backend handle by hand. SourceNamespace is + required for a Volume drill, where the source is a PVC, and unused for a Group + drill, where SourceRef names a consistency group. + properties: + bubbleCluster: + description: |- + BubbleCluster is the OCM ManagedCluster to recover onto: a DR target holding + the replicated point, or another cluster. It must differ from SourceCluster; + test-failover recovers onto a different cluster, never in place. Immutable. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + bubbleNamespace: + default: bubble + description: |- + BubbleNamespace is the namespace on the bubble cluster where the recovered + PVCs are created. Immutable. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + scope: + description: Scope selects what the drill recovers. Immutable. + enum: + - Volume + - Group + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + sourceCluster: + description: |- + SourceCluster is the OCM ManagedCluster the source runs on. The hub reads + the source there through a ManagedClusterView. Immutable. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + sourceNamespace: + description: |- + SourceNamespace is the namespace of the source PVC on SourceCluster. + Required for scope=Volume. Immutable. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + sourceRef: + description: |- + SourceRef names the source on SourceCluster: a PersistentVolumeClaim in + SourceNamespace (scope=Volume), or a consistency group (scope=Group). + Immutable. + type: string + x-kubernetes-validations: + - message: field is immutable + rule: self == oldSelf + ttlSeconds: + description: |- + TTLSeconds is an optional maximum lifetime: the drill is torn down after it + even without a delete, so a forgotten drill cannot hold a clone forever. + format: int64 + minimum: 0 + type: integer + required: + - bubbleCluster + - scope + - sourceCluster + - sourceRef + type: object + x-kubernetes-validations: + - message: field sourceNamespace is immutable once set + rule: '!has(oldSelf.sourceNamespace) || has(self.sourceNamespace)' + - message: field bubbleNamespace is immutable once set + rule: '!has(oldSelf.bubbleNamespace) || has(self.bubbleNamespace)' + - message: sourceNamespace is required for scope=Volume + rule: self.scope != 'Volume' || has(self.sourceNamespace) + status: + description: TestFailoverStatus is the observed state of one drill. + properties: + clones: + description: Clones is one entry per recovered volume. + items: + description: |- + TestFailoverClone is one recovered volume: the source it came from, the + snapshot and clone the drill built, and the PVC placed on the bubble cluster. + properties: + cloneID: + description: CloneID is the backend id of the writable clone. + type: string + pvcName: + description: PVCName is the bound PVC in the bubble namespace + on the bubble cluster. + type: string + sizeBytes: + description: SizeBytes is the recovered volume's size. + format: int64 + type: integer + snapshotID: + description: |- + SnapshotID is the recovery-point snapshot: the replicated snapshot already on + the bubble cluster's backend that the clone is built from. + type: string + sourceFSType: + description: |- + SourceFSType is the source PV's CSI fsType, carried onto the bubble PV so the + node plugin stages the clone with the filesystem it actually carries. The + clone is a block copy of the source, so its filesystem is the source's; an + empty fsType makes the node plugin default to ext4 and refuse to mount an XFS + volume. + type: string + sourceHandle: + description: SourceHandle is the source volume's backend handle, + read from its PV. + type: string + sourceRef: + description: |- + SourceRef is the source volume, or group member, the recovered volume maps + to. + type: string + sourceVolumeContext: + additionalProperties: + type: string + description: |- + SourceVolumeContext is the source PV's CSI volumeAttributes, minus the + identity and provisioner keys, carried onto the bubble PV so the node plugin + receives a non-nil VolumeContext when it stages the clone. The clone's own + identity (NQN, connections, nsId, and so on) is re-resolved from the clone + handle at stage time, so only the class-level parameters are carried; the + identity keys are dropped so a failed clone lookup can never point the mount + back at the source. + type: object + required: + - sourceRef + type: object + type: array + x-kubernetes-list-map-keys: + - sourceRef + x-kubernetes-list-type: map + completedAt: + description: CompletedAt is when the drill reached a terminal phase. + format: date-time + type: string + message: + description: |- + Message is the reason the phase is what it is: one sentence, replaced as the + drill moves, and never a log. + type: string + observedGeneration: + description: |- + ObservedGeneration is the generation the rest of this status was computed + from, so a stale status can be told from a current one. + format: int64 + type: integer + phase: + description: Phase is the drill's own progress. + enum: + - Pending + - Provisioning + - Ready + - Failed + - TearingDown + type: string + readyAt: + description: ReadyAt is when every recovered PVC became bound. + format: date-time + type: string + report: + description: Report is the drill's evidence, populated as it reaches + Ready. + properties: + bubbleCluster: + description: BubbleCluster is the cluster the drill recovered + onto. + type: string + invariantsHeld: + description: |- + InvariantsHeld is true only when the source fingerprint taken before the + drill matches the one taken at Ready. A Ready drill with this false is a + defect. + type: boolean + recoveryPoint: + description: RecoveryPoint is the snapshot or group generation + the drill recovered. + type: string + recoveryPointAgeSeconds: + description: RecoveryPointAgeSeconds is the drill time minus the + recovery-point time. + format: int64 + type: integer + recoveryPointTime: + description: RecoveryPointTime is when that point was taken. + format: date-time + type: string + type: object + startedAt: + description: StartedAt is when the drill started. + format: date-time + type: string + step: + description: Step is the position of the running drill's state machine. + properties: + deadline: + description: |- + Deadline is when that state expires, absent when it has none. It is an + absolute instant, so a state whose deadline passed while the controller + was down restores as already expired. + format: date-time + type: string + state: + description: |- + State is the state the machine was in. Empty means the resource has not + been reconciled yet, and restores to the graph's initial state. + type: string + type: object + x-kubernetes-validations: + - message: unknown step + rule: '!has(self.state) || self.state in [''ResolvingSource'',''ResolvingPoint'',''Shipping'',''Cloning'',''Placing'',''Releasing'']' + triggered: + description: |- + Triggered records that the current step's side effect was issued, so a + restart does not repeat it. + type: boolean + type: object + type: object + served: true + storage: true + subresources: + status: {} diff --git a/operator/internal/webapi/consistency_group.go b/operator/internal/webapi/consistency_group.go index cfbd574a8..31cdba847 100644 --- a/operator/internal/webapi/consistency_group.go +++ b/operator/internal/webapi/consistency_group.go @@ -9,11 +9,18 @@ import ( "net/http" ) -// ConsistencyGroupInfo is a group summary (design §10). +// ConsistencyGroupInfo is a group summary (design §10). PolicyID is the group's +// replication policy: a group attached with attach_group_policy stores its policy +// on the group record, so the group drill reads the policy off the group rather +// than inferring it from the policy list's placement (which a group-first attach +// leaves empty). LvsName and NodeID carry the group's pinned placement. type ConsistencyGroupInfo struct { UUID string `json:"id"` Name string `json:"name"` MemberCount int `json:"member_count"` + LvsName string `json:"lvs_name"` + NodeID string `json:"node_id"` + PolicyID string `json:"policy_id"` } // ConsistencyGroupMember is one current member of a group (design §10 /members). diff --git a/operator/internal/webapi/group_failover.go b/operator/internal/webapi/group_failover.go new file mode 100644 index 000000000..9592f1194 --- /dev/null +++ b/operator/internal/webapi/group_failover.go @@ -0,0 +1,171 @@ +// Group test-failover reads: the control-plane calls the TestFailover controller +// makes to recover a whole consistency group. A group drill needs three things +// the per-volume path does not: the group's replication policy (the group form +// of the recovery point is keyed on the policy, not the group), the one +// group-consistent generation of replicated snapshots on the target, and each +// member volume's K8s identity (PVC name and namespace) so the recovered PVCs +// can be named and their source PVs read for staging metadata. +package webapi + +import ( + "context" + "encoding/json" + "fmt" + "net/http" + "strings" +) + +// storagePoolsListPathFmt is the cluster-scoped storage-pools list endpoint. +const storagePoolsListPathFmt = "/api/v2/clusters/%s/storage-pools/" + +// ReplicatedGroupSnapshot is one member's replicated snapshot on the target +// cluster, at one group-consistent generation. It is the cloneable point for +// that member: cluster and pool address the target backend, snapshot is the +// snapshot to clone, and size sizes the recovered PVC. +type ReplicatedGroupSnapshot struct { + SnapshotID string `json:"snapshot_id"` + ClusterID string `json:"cluster_id"` + PoolID string `json:"pool_id"` + LvolID string `json:"lvol_id"` + Size int64 `json:"size"` + GroupSeq int `json:"group_seq"` +} + +// latestGenerationResponse is the group form of latest-snapshot: one generation +// number and one replicated snapshot per current member, all on the target. +type latestGenerationResponse struct { + GroupSeq int `json:"group_seq"` + Members []ReplicatedGroupSnapshot `json:"members"` +} + +// LatestReplicatedGeneration resolves the latest group-consistent generation on +// the target for a policy, returning the generation and one replicated snapshot +// per member. The control plane refuses (400) when no generation is complete for +// every member or when members straddle generations, which is what makes the +// recovered set crash-consistent; found is false only when there is no +// generation yet (nothing has replicated). +func (c *Client) LatestReplicatedGeneration( + ctx context.Context, + clusterUUID, policyID string, +) (groupSeq int, members []ReplicatedGroupSnapshot, found bool, err error) { + endpoint := fmt.Sprintf("/api/v2/clusters/%s/replication/policies/%s/latest-generation", clusterUUID, policyID) + body, statusCode, doErr := c.Do(ctx, http.MethodGet, endpoint, nil) + if statusCode == http.StatusNotFound { + return 0, nil, false, nil + } + if doErr != nil { + return 0, nil, false, fmt.Errorf("resolve latest replicated generation: %w", doErr) + } + if statusCode >= 300 { + return 0, nil, false, fmt.Errorf("resolve latest replicated generation: status %d: %s", statusCode, string(body)) + } + var dto latestGenerationResponse + if err := json.Unmarshal(body, &dto); err != nil { + return 0, nil, false, fmt.Errorf("decode latest-generation: %w", err) + } + return dto.GroupSeq, dto.Members, len(dto.Members) > 0, nil +} + +// MemberVolume is a group member's source volume, resolved to the K8s identity +// the drill needs: the PVC name and namespace it was provisioned for, so the +// recovered PVC can be named and the source PV read for staging metadata. +// +// The control plane stores the PVC identity in one field as "namespace/name" and +// uses the separate "namespace" field for the NVMe namespace, not the K8s one, so +// the K8s namespace and name are split out of PVCRef rather than read from the +// volume's namespace field. +type MemberVolume struct { + LvolID string + PVCName string + PVCNamespace string + PoolID string + Size int64 +} + +// memberVolumeDTO is the wire shape ResolveMemberVolumes decodes before splitting +// the namespaced PVC reference into a namespace and a name. +type memberVolumeDTO struct { + LvolID string `json:"id"` + PVCRef string `json:"pvc_name"` + PoolID string `json:"pool_id"` + Size int64 `json:"size"` +} + +// ResolveMemberVolumes maps each of the given member lvol ids to its source +// volume's K8s identity on the source cluster. A group member carries only an +// lvol id, so the pool is not known up front; this enumerates the cluster's +// pools and their volumes once and matches. Members it cannot find are omitted +// from the result, so the caller can tell an incomplete resolution from a +// complete one by the map size. +func (c *Client) ResolveMemberVolumes( + ctx context.Context, + clusterUUID string, + lvolIDs []string, +) (map[string]MemberVolume, error) { + want := make(map[string]struct{}, len(lvolIDs)) + for _, id := range lvolIDs { + want[id] = struct{}{} + } + + poolsEndpoint := fmt.Sprintf(storagePoolsListPathFmt, clusterUUID) + body, statusCode, err := c.Do(ctx, http.MethodGet, poolsEndpoint, nil) + if err != nil { + return nil, fmt.Errorf("list storage pools: %w", err) + } + if statusCode >= 300 { + return nil, fmt.Errorf("list storage pools: status %d: %s", statusCode, string(body)) + } + var pools []struct { + ID string `json:"id"` + } + if err := json.Unmarshal(body, &pools); err != nil { + return nil, fmt.Errorf("unmarshal storage pools: %w", err) + } + + found := make(map[string]MemberVolume, len(lvolIDs)) + for _, pool := range pools { + if len(found) == len(want) { + break + } + volsEndpoint := fmt.Sprintf("/api/v2/clusters/%s/storage-pools/%s/volumes", clusterUUID, pool.ID) + vbody, vstatus, verr := c.Do(ctx, http.MethodGet, volsEndpoint, nil) + if verr != nil { + return nil, fmt.Errorf("list volumes in pool %s: %w", pool.ID, verr) + } + if vstatus >= 300 { + return nil, fmt.Errorf("list volumes in pool %s: status %d: %s", pool.ID, vstatus, string(vbody)) + } + var vols []memberVolumeDTO + if err := json.Unmarshal(vbody, &vols); err != nil { + return nil, fmt.Errorf("unmarshal volumes in pool %s: %w", pool.ID, err) + } + for i := range vols { + v := vols[i] + if _, ok := want[v.LvolID]; !ok { + continue + } + ns, name := splitPVCRef(v.PVCRef) + poolID := v.PoolID + if poolID == "" { + poolID = pool.ID + } + found[v.LvolID] = MemberVolume{ + LvolID: v.LvolID, + PVCName: name, + PVCNamespace: ns, + PoolID: poolID, + Size: v.Size, + } + } + } + return found, nil +} + +// splitPVCRef splits a "namespace/name" PVC reference into its namespace and +// name. A reference with no slash is taken as a bare name in no namespace. +func splitPVCRef(ref string) (namespace, name string) { + if i := strings.IndexByte(ref, '/'); i >= 0 { + return ref[:i], ref[i+1:] + } + return "", ref +} diff --git a/operator/internal/webapi/group_failover_test.go b/operator/internal/webapi/group_failover_test.go new file mode 100644 index 000000000..df6aee72c --- /dev/null +++ b/operator/internal/webapi/group_failover_test.go @@ -0,0 +1,99 @@ +package webapi + +import ( + "context" + "net/http" + "net/http/httptest" + "testing" +) + +// routingServer serves fixed JSON bodies keyed by request path, so a test can +// stand in for the control plane without the openapi spec mock. +func routingServer(t *testing.T, routes map[string]string) *Client { + t.Helper() + srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + body, ok := routes[r.URL.Path] + if !ok { + http.NotFound(w, r) + return + } + w.Header().Set("Content-Type", "application/json") + _, _ = w.Write([]byte(body)) + })) + t.Cleanup(srv.Close) + return NewClient(srv.URL) +} + +func TestLatestReplicatedGenerationReturnsPerMemberSnapshots(t *testing.T) { + c := routingServer(t, map[string]string{ + "/api/v2/clusters/C/replication/policies/P/latest-generation": `{ + "group_seq": 7, + "members": [ + {"snapshot_id":"s1","cluster_id":"B","pool_id":"pb","lvol_id":"t1","size":1073741824,"group_seq":7}, + {"snapshot_id":"s2","cluster_id":"B","pool_id":"pb","lvol_id":"t2","size":1073741824,"group_seq":7} + ] + }`, + }) + seq, members, found, err := c.LatestReplicatedGeneration(context.Background(), "C", "P") + if err != nil { + t.Fatalf("LatestReplicatedGeneration: %v", err) + } + if !found { + t.Fatalf("found = false, want true when a generation exists") + } + if seq != 7 { + t.Errorf("group_seq = %d, want 7", seq) + } + if len(members) != 2 || members[0].SnapshotID != "s1" || members[1].SnapshotID != "s2" { + t.Errorf("members = %+v, want two with s1,s2", members) + } + if members[0].ClusterID != "B" || members[0].PoolID != "pb" || members[0].Size != 1073741824 { + t.Errorf("member[0] = %+v, want the target cluster/pool/size", members[0]) + } +} + +func TestLatestReplicatedGenerationNotFoundWhenNothingReplicated(t *testing.T) { + // No route registered -> 404 -> found=false, no error (nothing has replicated). + c := routingServer(t, map[string]string{}) + _, _, found, err := c.LatestReplicatedGeneration(context.Background(), "C", "P") + if err != nil { + t.Fatalf("LatestReplicatedGeneration on 404: %v", err) + } + if found { + t.Errorf("found = true, want false on 404") + } +} + +func TestResolveMemberVolumesMapsLvolIDsToPVCNames(t *testing.T) { + // Two pools; the two members live in different pools. Resolution must find + // both and carry their pvc name, namespace, pool, and size. + // pvc_name is the namespaced "namespace/name" form, and pool_id is null in the + // list (the iterating pool supplies it); "namespace" is the NVMe namespace and + // must be ignored. + c := routingServer(t, map[string]string{ + "/api/v2/clusters/C/storage-pools/": `[{"id":"pool-1"},{"id":"pool-2"}]`, + "/api/v2/clusters/C/storage-pools/pool-1/volumes": `[ + {"id":"lvol-a","pvc_name":"app/data-1","namespace":"nvme-ns-uuid","pool_id":null,"size":1073741824}, + {"id":"lvol-x","pvc_name":"app/other","namespace":"nvme-ns-uuid","pool_id":null,"size":1073741824} + ]`, + "/api/v2/clusters/C/storage-pools/pool-2/volumes": `[ + {"id":"lvol-b","pvc_name":"app/data-2","namespace":"nvme-ns-uuid","pool_id":null,"size":1073741824} + ]`, + }) + got, err := c.ResolveMemberVolumes(context.Background(), "C", []string{"lvol-a", "lvol-b"}) + if err != nil { + t.Fatalf("ResolveMemberVolumes: %v", err) + } + if len(got) != 2 { + t.Fatalf("resolved %d members, want 2: %+v", len(got), got) + } + if got["lvol-a"].PVCName != "data-1" || got["lvol-a"].PVCNamespace != "app" || got["lvol-a"].PoolID != "pool-1" { + t.Errorf("lvol-a = %+v, want name=data-1 ns=app pool=pool-1", got["lvol-a"]) + } + if got["lvol-b"].PVCName != "data-2" || got["lvol-b"].PVCNamespace != "app" || got["lvol-b"].PoolID != "pool-2" { + t.Errorf("lvol-b = %+v, want name=data-2 ns=app pool=pool-2", got["lvol-b"]) + } + if _, ok := got["lvol-x"]; ok { + t.Errorf("resolved an unrequested member lvol-x") + } +} From 663764f754f5b128266f00ff86388000208d066c Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Thu, 1 Oct 2026 09:48:02 +0100 Subject: [PATCH 163/206] fixed merge linter issue --- .../controllers/deployment/draftseed_test.go | 11 ++++++----- .../deployment/operatorops_controller.go | 19 ++++++++++++++++++- .../deployment/operatorops_unit_test.go | 8 ++++---- 3 files changed, 28 insertions(+), 10 deletions(-) diff --git a/operator/internal/controllers/deployment/draftseed_test.go b/operator/internal/controllers/deployment/draftseed_test.go index c6ad3c6d6..265b9866e 100644 --- a/operator/internal/controllers/deployment/draftseed_test.go +++ b/operator/internal/controllers/deployment/draftseed_test.go @@ -15,6 +15,7 @@ package deployment import ( + "context" "strings" "testing" @@ -73,7 +74,7 @@ func noteMentioning(notes []string, want string) bool { func TestTheStatedLayoutSeedsTheInitialRunsDraft(t *testing.T) { r := &OperatorOpsReconciler{} - draft, notes, err := r.draftFor(initialRun(), &simplyblockv1alpha2.DiscoverSpec{}, + draft, notes, err := r.draftFor(context.Background(), initialRun(), &simplyblockv1alpha2.DiscoverSpec{}, discoverypkg.Plan{}, statedLayout()) if err != nil { t.Fatalf("the fleet was refused: %v", err) @@ -120,7 +121,7 @@ func TestTheStatedLayoutSeedsTheInitialRunsDraft(t *testing.T) { func TestTheStatedDraftFieldsReachTheDocument(t *testing.T) { r := &OperatorOpsReconciler{} - draft, _, err := r.draftFor(initialRun(), &simplyblockv1alpha2.DiscoverSpec{}, + draft, _, err := r.draftFor(context.Background(), initialRun(), &simplyblockv1alpha2.DiscoverSpec{}, discoverypkg.Plan{}, statedLayout()) if err != nil { t.Fatalf("the fleet was refused: %v", err) @@ -157,7 +158,7 @@ func TestARunNobodyLabeledGetsNoSeed(t *testing.T) { ObjectMeta: metav1.ObjectMeta{Name: "discover-again"}, } - draft, _, err := r.draftFor(theirs, &simplyblockv1alpha2.DiscoverSpec{}, + draft, _, err := r.draftFor(context.Background(), theirs, &simplyblockv1alpha2.DiscoverSpec{}, discoverypkg.Plan{}, statedLayout()) if err != nil { t.Fatalf("the fleet was refused: %v", err) @@ -198,7 +199,7 @@ func TestAnUnstatedFieldStaysDerived(t *testing.T) { Cluster: bootstrap.ClusterConfig{EnableChecksumValidation: ptr.To(true)}, }} - draft, _, err := r.draftFor(initialRun(), &simplyblockv1alpha2.DiscoverSpec{}, + draft, _, err := r.draftFor(context.Background(), initialRun(), &simplyblockv1alpha2.DiscoverSpec{}, discoverypkg.Plan{}, partial) if err != nil { t.Fatalf("the fleet was refused: %v", err) @@ -230,7 +231,7 @@ func TestAGrowthDraftIsNotSeeded(t *testing.T) { r := &OperatorOpsReconciler{} spec := &simplyblockv1alpha2.DiscoverSpec{ClusterRef: "simplyblock-cluster"} - draft, _, err := r.draftFor(initialRun(), spec, discoverypkg.Plan{}, statedLayout()) + draft, _, err := r.draftFor(context.Background(), initialRun(), spec, discoverypkg.Plan{}, statedLayout()) if err != nil { t.Fatalf("the fleet was refused: %v", err) } diff --git a/operator/internal/controllers/deployment/operatorops_controller.go b/operator/internal/controllers/deployment/operatorops_controller.go index 0f1d274c9..5f82f4c59 100644 --- a/operator/internal/controllers/deployment/operatorops_controller.go +++ b/operator/internal/controllers/deployment/operatorops_controller.go @@ -645,7 +645,7 @@ func (r *OperatorOpsReconciler) write( "could not be parsed; the draft states what this run found") } - config, notes, err := r.draftFor(ops, spec, plan, installation) + config, notes, err := r.draftFor(ctx, ops, spec, plan, installation) if err != nil { // A fleet the run read and cannot draft a document for. The reason names // the worker and the shape of its disks, and the run's own message is one @@ -733,6 +733,11 @@ func refuseUnreadableFilter(spec *simplyblockv1alpha2.DiscoverSpec) error { // either, and the caller falls back to the name exactly as it was before this // existed. func (r *OperatorOpsReconciler) clusterDisambiguator(ctx context.Context) string { + // No client to read kube-system from is the same as an unreadable one: no + // disambiguator, and the name is left as it was. + if r.Client == nil { + return "" + } var ns corev1.Namespace key := client.ObjectKey{Name: metav1.NamespaceSystem} if err := r.Get(ctx, key, &ns); err != nil || ns.UID == "" { @@ -817,6 +822,18 @@ func (r *OperatorOpsReconciler) draftFor( // and cannot be changed, so a stated one is taken over the one derived // from the draft's own name. clusterName := name + clusterNameSuffix + if disambiguator := r.clusterDisambiguator(ctx); disambiguator != "" { + base := name + // Truncated so the disambiguated name still fits, the same way a + // name too long for the limit already does not: this does not newly + // break a base name that already did not fit, only keeps this suffix + // from being the reason a borderline one no longer does. + room := clusterNameLimit - len(clusterNameSuffix) - len(disambiguator) - 1 + if room > 0 && len(base) > room { + base = base[:room] + } + clusterName = base + "-" + disambiguator + clusterNameSuffix + } if seed != nil && seed.Name != "" { clusterName = seed.Name } diff --git a/operator/internal/controllers/deployment/operatorops_unit_test.go b/operator/internal/controllers/deployment/operatorops_unit_test.go index ff39aa6ad..9295f4489 100644 --- a/operator/internal/controllers/deployment/operatorops_unit_test.go +++ b/operator/internal/controllers/deployment/operatorops_unit_test.go @@ -827,7 +827,7 @@ func TestTheDraftStatesTheHostOSTheProbesRead(t *testing.T) { }}, }}} - config, notes, err := r.draftFor(ops, &simplyblockv1alpha2.DiscoverSpec{}, plan, nil) + config, notes, err := r.draftFor(context.Background(), ops, &simplyblockv1alpha2.DiscoverSpec{}, plan, nil) if err != nil { t.Fatalf("the fleet was refused: %v", err) } @@ -864,7 +864,7 @@ func TestTheDraftStatesNoHostOSForAFleetThatDisagrees(t *testing.T) { worker("worker-02", "rocky"), }} - config, notes, err := r.draftFor(ops, &simplyblockv1alpha2.DiscoverSpec{}, plan, nil) + config, notes, err := r.draftFor(context.Background(), ops, &simplyblockv1alpha2.DiscoverSpec{}, plan, nil) if err != nil { t.Fatalf("the fleet was refused: %v", err) } @@ -916,7 +916,7 @@ func TestTheDraftCarriesTheTolerationsTheRunProbedWith(t *testing.T) { ops := &simplyblockv1alpha2.OperatorOps{ObjectMeta: metav1.ObjectMeta{Name: "discover-1"}} spec := &simplyblockv1alpha2.DiscoverSpec{Tolerations: storagePlaneTaint} - config, notes, err := r.draftFor(ops, spec, discoverypkg.Plan{}, nil) + config, notes, err := r.draftFor(context.Background(), ops, spec, discoverypkg.Plan{}, nil) if err != nil { t.Fatalf("the fleet was refused: %v", err) } @@ -942,7 +942,7 @@ func TestAGrowthDraftCarriesNoTolerations(t *testing.T) { Tolerations: storagePlaneTaint, } - config, _, err := r.draftFor(ops, spec, discoverypkg.Plan{}, nil) + config, _, err := r.draftFor(context.Background(), ops, spec, discoverypkg.Plan{}, nil) if err != nil { t.Fatalf("the fleet was refused: %v", err) } From df26de2d4b9feafda350c016fcbeb332f5c7bb78 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Thu, 1 Oct 2026 17:22:33 +0100 Subject: [PATCH 164/206] fix(operator): skip hub-only TestFailover controller where OCM work API is absent The operator crashed at startup on every cluster outside the hub. The TestFailover controller unconditionally watched OCM ManifestWork (.Owns(&workv1.ManifestWork{})), but ManifestWork is served only on the hub -- managed clusters read it from the hub and never serve it locally. Without the CRD the controller's cache never syncs and the manager exits ("failed to wait for testfailover caches to sync ... *v1.ManifestWork"), taking every other controller (replication, storagecluster, ...) down with it. That blocked DR failback, since the managed cluster's operator reconciles the reverse ReplicationPair. Gate the registration on the work.open-cluster-management.io API group actually being served (serverHasAPIGroup, via the existing discovery client pattern). On the hub it registers as before; off the hub it logs and skips. Installing the ManifestWork CRD on managed clusters is the wrong workaround -- it belongs only on the hub. Co-Authored-By: Claude Opus 4.8 --- operator/cmd/main.go | 52 ++++++++++++++++++++++++++++++++++----- operator/cmd/main_test.go | 31 +++++++++++++++++++++++ 2 files changed, 77 insertions(+), 6 deletions(-) diff --git a/operator/cmd/main.go b/operator/cmd/main.go index d539b9f21..c783275b6 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -83,12 +83,33 @@ var ( const ( openShiftConfigAPIGroup = "config.openshift.io" certManagerAPIGroup = "cert-manager.io" + // ocmWorkAPIGroup is OCM's ManifestWork API group. It is served on the hub + // and absent on managed clusters, which read ManifestWork from the hub rather + // than serving it locally. The TestFailover controller watches ManifestWork, + // so it is registered only where this group is served. + ocmWorkAPIGroup = "work.open-cluster-management.io" ) type serverGroupsGetter interface { ServerGroups() (*metav1.APIGroupList, error) } +// serverHasAPIGroup reports whether the API server serves the named group. Used +// to skip controllers that watch a kind the cluster does not define, since a +// watch whose cache can never sync takes the whole manager down at startup. +func serverHasAPIGroup(discoveryClient serverGroupsGetter, group string) (bool, error) { + groupList, err := discoveryClient.ServerGroups() + if err != nil { + return false, fmt.Errorf("discover API groups: %w", err) + } + for _, g := range groupList.Groups { + if g.Name == group { + return true, nil + } + } + return false, nil +} + func init() { utilruntime.Must(clientgoscheme.AddToScheme(scheme)) // The cert-manager webhook provisioner injects its CA bundle into the @@ -904,14 +925,33 @@ func main() { setupLog.Error(err, "unable to create controller", "controller", "VolumeGroupSnapshotOps") os.Exit(1) } - if err := (&controller.TestFailoverReconciler{ - Client: mgr.GetClient(), - Scheme: mgr.GetScheme(), - Recorder: mgr.GetEventRecorder("testfailover-controller"), - }).SetupWithManager(mgr); err != nil { - setupLog.Error(err, "unable to create controller", "controller", "TestFailover") + // The TestFailover controller watches OCM ManifestWork, which only the hub + // serves; registering it where the work API is absent leaves a watch whose + // cache never syncs and the manager exits at startup, taking every other + // controller with it. It is a hub-only drill, so skip it off the hub. + workDiscovery, err := discovery.NewDiscoveryClientForConfig(cfg) + if err != nil { + setupLog.Error(err, "unable to build a discovery client to check for the OCM work API") os.Exit(1) } + hasWorkAPI, err := serverHasAPIGroup(workDiscovery, ocmWorkAPIGroup) + if err != nil { + setupLog.Error(err, "unable to determine whether the OCM work API is served") + os.Exit(1) + } + if hasWorkAPI { + if err := (&controller.TestFailoverReconciler{ + Client: mgr.GetClient(), + Scheme: mgr.GetScheme(), + Recorder: mgr.GetEventRecorder("testfailover-controller"), + }).SetupWithManager(mgr); err != nil { + setupLog.Error(err, "unable to create controller", "controller", "TestFailover") + os.Exit(1) + } + } else { + setupLog.Info("OCM ManifestWork API not served; skipping TestFailover controller (hub-only)", + "apiGroup", ocmWorkAPIGroup) + } // +kubebuilder:scaffold:builder // Provision the admission webhooks' serving certificate at runtime (self-signed diff --git a/operator/cmd/main_test.go b/operator/cmd/main_test.go index 938893810..9b892ac8e 100644 --- a/operator/cmd/main_test.go +++ b/operator/cmd/main_test.go @@ -21,6 +21,37 @@ func (f fakeServerGroupsGetter) ServerGroups() (*metav1.APIGroupList, error) { return list, nil } +// Regression: 2026-10-01 — the operator crashed at startup on every cluster +// outside the hub. The TestFailover controller unconditionally watched OCM's +// ManifestWork (.Owns(&workv1.ManifestWork{})), but ManifestWork is served only +// on the hub — managed clusters read it from the hub and never serve it locally. +// Without the CRD the controller's cache never syncs and the manager exits +// ("failed to wait for testfailover caches to sync ... *v1.ManifestWork"), +// taking every other controller (replication, storagecluster, …) down with it. +// Registration must be gated on the work API actually being served. +func TestServerHasAPIGroup(t *testing.T) { + tests := []struct { + name string + groups []string + want bool + }{ + {name: "work api served (hub)", groups: []string{ocmWorkAPIGroup, "storage.simplyblock.io"}, want: true}, + {name: "work api absent (managed cluster)", groups: []string{"storage.simplyblock.io"}, want: false}, + {name: "no groups at all", groups: nil, want: false}, + } + for _, tc := range tests { + t.Run(tc.name, func(t *testing.T) { + got, err := serverHasAPIGroup(fakeServerGroupsGetter{groups: tc.groups}, ocmWorkAPIGroup) + if err != nil { + t.Fatalf("serverHasAPIGroup returned error: %v", err) + } + if got != tc.want { + t.Fatalf("serverHasAPIGroup = %v, want %v", got, tc.want) + } + }) + } +} + func TestValidateTLSConfiguration(t *testing.T) { tests := []struct { name string From a18f13ac160050d494494c44539522c4ad3b1de7 Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Thu, 1 Oct 2026 18:30:17 +0100 Subject: [PATCH 165/206] fix(operator): gate TestFailover on the ManifestWork resource, not its API group The previous gate checked the work.open-cluster-management.io API *group*, but a managed cluster serves that same group/version for AppliedManifestWork (the work-agent's local record) while never serving ManifestWork. So the group-level check saw the group as served, registered the hub-only TestFailover controller anyway, and the manager still crashed at startup when the ManifestWork watch could not sync (confirmed live 2026-10-01: operator on the managed cluster Error at 2m32s). Check the ManifestWork resource specifically via ServerResourcesForGroupVersion. The regression test now covers the AppliedManifestWork-only managed cluster, which is the case the group-level check got wrong. Co-Authored-By: Claude Opus 4.8 --- operator/cmd/main.go | 50 +++++++++++++++++++------------ operator/cmd/main_test.go | 62 +++++++++++++++++++++++++++++++-------- 2 files changed, 81 insertions(+), 31 deletions(-) diff --git a/operator/cmd/main.go b/operator/cmd/main.go index c783275b6..b3258ecd9 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -38,6 +38,7 @@ import ( _ "k8s.io/client-go/plugin/pkg/client/auth" apiextensionsv1 "k8s.io/apiextensions-apiserver/pkg/apis/apiextensions/v1" + apierrors "k8s.io/apimachinery/pkg/api/errors" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" "k8s.io/apimachinery/pkg/runtime" utilruntime "k8s.io/apimachinery/pkg/util/runtime" @@ -83,27 +84,40 @@ var ( const ( openShiftConfigAPIGroup = "config.openshift.io" certManagerAPIGroup = "cert-manager.io" - // ocmWorkAPIGroup is OCM's ManifestWork API group. It is served on the hub - // and absent on managed clusters, which read ManifestWork from the hub rather - // than serving it locally. The TestFailover controller watches ManifestWork, - // so it is registered only where this group is served. - ocmWorkAPIGroup = "work.open-cluster-management.io" + // ocmWorkGroupVersion is OCM's work API group/version, and ocmManifestWorkResource + // the ManifestWork resource within it. ManifestWork is served only on the hub; + // a managed cluster serves the SAME group/version for AppliedManifestWork (the + // work-agent's local record) but NOT ManifestWork, so the presence check must + // be resource-level, not group-level — a group-level check sees the group as + // served everywhere AppliedManifestWork exists. The TestFailover controller + // watches ManifestWork, so it is registered only where ManifestWork is served. + ocmWorkGroupVersion = "work.open-cluster-management.io/v1" + ocmManifestWorkResource = "manifestworks" ) type serverGroupsGetter interface { ServerGroups() (*metav1.APIGroupList, error) } -// serverHasAPIGroup reports whether the API server serves the named group. Used -// to skip controllers that watch a kind the cluster does not define, since a -// watch whose cache can never sync takes the whole manager down at startup. -func serverHasAPIGroup(discoveryClient serverGroupsGetter, group string) (bool, error) { - groupList, err := discoveryClient.ServerGroups() +type serverResourcesGetter interface { + ServerResourcesForGroupVersion(groupVersion string) (*metav1.APIResourceList, error) +} + +// serverHasResource reports whether the API server serves the named resource in +// the given group/version. Used to skip a controller that watches a kind the +// cluster does not serve, since a watch whose cache can never sync takes the +// whole manager down at startup. A group/version the server does not serve at +// all is reported as absent rather than an error. +func serverHasResource(discoveryClient serverResourcesGetter, groupVersion, resource string) (bool, error) { + list, err := discoveryClient.ServerResourcesForGroupVersion(groupVersion) if err != nil { - return false, fmt.Errorf("discover API groups: %w", err) + if apierrors.IsNotFound(err) { + return false, nil + } + return false, fmt.Errorf("discover resources for %s: %w", groupVersion, err) } - for _, g := range groupList.Groups { - if g.Name == group { + for _, r := range list.APIResources { + if r.Name == resource { return true, nil } } @@ -934,12 +948,12 @@ func main() { setupLog.Error(err, "unable to build a discovery client to check for the OCM work API") os.Exit(1) } - hasWorkAPI, err := serverHasAPIGroup(workDiscovery, ocmWorkAPIGroup) + hasManifestWork, err := serverHasResource(workDiscovery, ocmWorkGroupVersion, ocmManifestWorkResource) if err != nil { - setupLog.Error(err, "unable to determine whether the OCM work API is served") + setupLog.Error(err, "unable to determine whether the OCM ManifestWork resource is served") os.Exit(1) } - if hasWorkAPI { + if hasManifestWork { if err := (&controller.TestFailoverReconciler{ Client: mgr.GetClient(), Scheme: mgr.GetScheme(), @@ -949,8 +963,8 @@ func main() { os.Exit(1) } } else { - setupLog.Info("OCM ManifestWork API not served; skipping TestFailover controller (hub-only)", - "apiGroup", ocmWorkAPIGroup) + setupLog.Info("OCM ManifestWork resource not served; skipping TestFailover controller (hub-only)", + "groupVersion", ocmWorkGroupVersion, "resource", ocmManifestWorkResource) } // +kubebuilder:scaffold:builder diff --git a/operator/cmd/main_test.go b/operator/cmd/main_test.go index 9b892ac8e..d6c615cb1 100644 --- a/operator/cmd/main_test.go +++ b/operator/cmd/main_test.go @@ -4,7 +4,9 @@ import ( "strings" "testing" + apierrors "k8s.io/apimachinery/pkg/api/errors" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/runtime/schema" "github.com/simplyblock/simplyblock-operator/internal/utils" ) @@ -21,32 +23,66 @@ func (f fakeServerGroupsGetter) ServerGroups() (*metav1.APIGroupList, error) { return list, nil } +type fakeServerResourcesGetter struct { + // groupVersion -> resource names served. A missing key means the server does + // not serve that group/version at all (discovery returns NotFound). + resources map[string][]string +} + +func (f fakeServerResourcesGetter) ServerResourcesForGroupVersion(gv string) (*metav1.APIResourceList, error) { + names, ok := f.resources[gv] + if !ok { + return nil, apierrors.NewNotFound(schema.GroupResource{Resource: gv}, "") + } + list := &metav1.APIResourceList{GroupVersion: gv} + for _, n := range names { + list.APIResources = append(list.APIResources, metav1.APIResource{Name: n}) + } + return list, nil +} + // Regression: 2026-10-01 — the operator crashed at startup on every cluster // outside the hub. The TestFailover controller unconditionally watched OCM's // ManifestWork (.Owns(&workv1.ManifestWork{})), but ManifestWork is served only -// on the hub — managed clusters read it from the hub and never serve it locally. -// Without the CRD the controller's cache never syncs and the manager exits +// on the hub — managed clusters serve the same group/version for +// AppliedManifestWork (the work-agent's local record) yet never serve +// ManifestWork. The controller's cache never syncs and the manager exits // ("failed to wait for testfailover caches to sync ... *v1.ManifestWork"), // taking every other controller (replication, storagecluster, …) down with it. -// Registration must be gated on the work API actually being served. -func TestServerHasAPIGroup(t *testing.T) { +// Registration must be gated on the ManifestWork RESOURCE being served, not the +// work API group — a group-level check sees the group as served wherever +// AppliedManifestWork exists and so does not skip the controller (live 2026-10-01). +func TestServerHasManifestWork(t *testing.T) { tests := []struct { - name string - groups []string - want bool + name string + resources map[string][]string + want bool }{ - {name: "work api served (hub)", groups: []string{ocmWorkAPIGroup, "storage.simplyblock.io"}, want: true}, - {name: "work api absent (managed cluster)", groups: []string{"storage.simplyblock.io"}, want: false}, - {name: "no groups at all", groups: nil, want: false}, + { + name: "hub serves ManifestWork", + resources: map[string][]string{ocmWorkGroupVersion: {"manifestworks", "appliedmanifestworks"}}, + want: true, + }, + { + name: "managed cluster serves only AppliedManifestWork", + resources: map[string][]string{ocmWorkGroupVersion: {"appliedmanifestworks"}}, + want: false, + }, + { + name: "cluster without OCM at all", + resources: map[string][]string{}, + want: false, + }, } for _, tc := range tests { t.Run(tc.name, func(t *testing.T) { - got, err := serverHasAPIGroup(fakeServerGroupsGetter{groups: tc.groups}, ocmWorkAPIGroup) + got, err := serverHasResource( + fakeServerResourcesGetter{resources: tc.resources}, ocmWorkGroupVersion, ocmManifestWorkResource) if err != nil { - t.Fatalf("serverHasAPIGroup returned error: %v", err) + t.Fatalf("serverHasResource returned error: %v", err) } if got != tc.want { - t.Fatalf("serverHasAPIGroup = %v, want %v", got, tc.want) + t.Fatalf("serverHasResource = %v, want %v", got, tc.want) } }) } From 5bb7084e8608d0a5ff7987013e1ba631427844dd Mon Sep 17 00:00:00 2001 From: geoffrey1330 Date: Thu, 1 Oct 2026 23:31:01 +0100 Subject: [PATCH 166/206] fix(operator): point tasks-runner-backup-merge at its real module The tasks-runner-backup-merge container ran simplyblock_core/services/backup_merge_service.py, which does not exist in the control-plane image -- the service is tasks_runner_backup_merge.py, following the same tasks_runner_* convention as every sibling runner. The container crash-looped ("python3: can't open file ... backup_merge_service.py") and pinned the whole tasks pod in CrashLoopBackOff (120+ restarts observed live 2026-10-01). The existing TestTheServicePoolsRunWhatTheyDeclare only checks the command is a .py module, so the wrong name passed it. Add a regression test pinning the real module name. Co-Authored-By: Claude Opus 4.8 --- .../controllers/controlplane/managementapi.go | 2 +- .../controlplane/workloads_test.go | 26 +++++++++++++++++++ 2 files changed, 27 insertions(+), 1 deletion(-) diff --git a/operator/internal/controllers/controlplane/managementapi.go b/operator/internal/controllers/controlplane/managementapi.go index d1191c7ae..ee1905e7e 100644 --- a/operator/internal/controllers/controlplane/managementapi.go +++ b/operator/internal/controllers/controlplane/managementapi.go @@ -422,7 +422,7 @@ func taskServices() []service { {name: "tasks-runner-node-removal", module: "simplyblock_core/services/tasks_runner_node_removal.py"}, {name: "tasks-runner-snapshot-replication", module: "simplyblock_core/services/snapshot_replication.py"}, {name: "tasks-runner-backup", module: "simplyblock_core/services/tasks_runner_backup.py"}, - {name: "tasks-runner-backup-merge", module: "simplyblock_core/services/backup_merge_service.py"}, + {name: "tasks-runner-backup-merge", module: "simplyblock_core/services/tasks_runner_backup_merge.py"}, {name: "tasks-runner-replication-final", module: "simplyblock_core/services/tasks_runner_replication_final.py"}, } } diff --git a/operator/internal/controllers/controlplane/workloads_test.go b/operator/internal/controllers/controlplane/workloads_test.go index 8a506810d..3b158153f 100644 --- a/operator/internal/controllers/controlplane/workloads_test.go +++ b/operator/internal/controllers/controlplane/workloads_test.go @@ -450,6 +450,32 @@ func TestTheServicePoolsRunWhatTheyDeclare(t *testing.T) { } } +// Regression: 2026-10-01 — the tasks-runner-backup-merge container named a module +// that does not exist in the control-plane image (backup_merge_service.py); the +// real service is tasks_runner_backup_merge.py, like every other tasks-runner-*. +// The container crash-looped ("python3: can't open file +// '/app/simplyblock_core/services/backup_merge_service.py'"), which pinned the +// whole tasks pod in CrashLoopBackOff. A .py-suffix check does not catch it, so +// pin the real module name. +func TestTheBackupMergeRunnerNamesItsRealModule(t *testing.T) { + cp := localControlPlane() + d := findDeployment(t, managementAPIObjects(cp), ComponentTasks) + const want = "simplyblock_core/services/tasks_runner_backup_merge.py" + found := false + for _, container := range d.Spec.Template.Spec.Containers { + if container.Name != "tasks-runner-backup-merge" { + continue + } + found = true + if len(container.Command) != 2 || container.Command[1] != want { + t.Errorf("tasks-runner-backup-merge runs %v, want python3 %q", container.Command, want) + } + } + if !found { + t.Fatal("no tasks-runner-backup-merge container in the tasks deployment") + } +} + // The control plane's account is granted exec on pods, which is the strongest // thing in its role and the one an audit has to be able to find. Losing it would // stop the control plane driving the storage nodes' processes, which is not a From 280662f42e5990666bdad923483f68f891b0e5b5 Mon Sep 17 00:00:00 2001 From: michael Date: Fri, 2 Oct 2026 22:29:38 +0300 Subject: [PATCH 167/206] csi-driver: replication RPCs act on the chain's last LOCAL member, not its end A volume's PV keeps the handle it was created with across fail-overs, and the chain of replication relationships behind that handle alternates between the sites: every fail-over adds a hop to the other side. resolveToLocalReplica walked the chain to its active end for every csi-addons Replication RPC. After an unplanned fail-over A->B, Ramen makes the old primary on A secondary: DemoteVolume (and DisableVolumeReplication on VR deletion) on site A resolved to the chain's end -- the NEW primary on site B -- and demoted it (live 2026-10-02, realbed WordPress: the demote fenced the live primary's paths and took demote snapshots of it; the VRG on A never became secondary, so the fail-back never got PeerReady). The driver needs to know which clusters are its own. The operator marks the clusters it manages `local` in the CSI secret's entries (an entry another site registered for cross-cluster handle resolution is not), and the resolver returns the chain's last member on a local cluster. A secret that marks no cluster local -- an operator predating the flag -- keeps the previous behaviour. Co-Authored-By: Claude Fable 5.1 --- csi-driver/internal/clusters/clusters.go | 22 ++++++++ .../internal/csi/controller/replication.go | 51 +++++++++++++++++-- .../csi/controller/replication_local_test.go | 50 ++++++++++++++++++ .../cluster/storagecluster_controller.go | 9 ++++ 4 files changed, 128 insertions(+), 4 deletions(-) create mode 100644 csi-driver/internal/csi/controller/replication_local_test.go diff --git a/csi-driver/internal/clusters/clusters.go b/csi-driver/internal/clusters/clusters.go index a71616866..daadbacf6 100644 --- a/csi-driver/internal/clusters/clusters.go +++ b/csi-driver/internal/clusters/clusters.go @@ -39,6 +39,10 @@ type Config struct { ClusterID string `json:"cluster_id"` ClusterEndpoint string `json:"cluster_endpoint"` ClusterSecret string `json:"cluster_secret"` + // Local marks a cluster of the site this driver runs on (written by the + // site's operator); an entry without it belongs to another site, kept so + // that a failed-over volume's handle still resolves. + Local bool `json:"local,omitempty"` } // Info is the secret file as a whole. @@ -86,6 +90,24 @@ func Load() (Info, error) { return clusters, nil } +// Local returns the ids of the clusters the secret marks local, and whether +// the secret marks any: a secret written by an operator that predates the +// flag marks none, and callers then fall back to treating every cluster as +// local. +func Local() (map[string]bool, bool, error) { + clusters, err := Load() + if err != nil { + return nil, false, err + } + local := map[string]bool{} + for _, cluster := range clusters.Clusters { + if cluster.Local { + local[cluster.ClusterID] = true + } + } + return local, len(local) > 0, nil +} + // List returns the ID of every cluster in the secret. func List() ([]string, error) { clusters, err := Load() diff --git a/csi-driver/internal/csi/controller/replication.go b/csi-driver/internal/csi/controller/replication.go index 242763056..f81a55ebc 100644 --- a/csi-driver/internal/csi/controller/replication.go +++ b/csi-driver/internal/csi/controller/replication.go @@ -93,19 +93,36 @@ func volumeIDFrom(req volumeIDCarrier) string { // combining active_lvol_id with ANOTHER record's cluster is exactly the bug // this walk exists to avoid. An empty ActiveLvolID (backend predating the // field) stops after the first hop, the old single-step behavior. +// +// The chain behind a handle alternates between the sites: every fail-over +// adds a hop to the other side. The volume this driver must act on is the +// chain's last member on a LOCAL cluster (the secret marks the site's own +// clusters, clusters.Local), not the chain's end: after an unplanned +// fail-over A->B, Ramen makes the old primary on A secondary, and the +// chain's end is the NEW primary on B. Resolving to the end demoted -- and +// on VR deletion detached -- the live production volume on the other site +// (2026-10-02, WordPress: the demote fenced the live primary's paths and +// took demote snapshots of it; the fail-back never got PeerReady). A secret +// that marks no cluster local (an operator predating the flag) keeps the +// previous behaviour, the chain's active end. func resolveToLocalReplica( ctx context.Context, h *lvol.Handle, client *atlascp.Client, ) (*lvol.Handle, *atlascp.Client, error) { + local, flagged, err := clusters.Local() + if err != nil { + return nil, nil, err + } + hops := []chainHop{{h: h, client: client}} for range 8 { // one hop per past fail-over; capped far above any real chain rel, err := client.GetVolumeReplicationRelationship(ctx, h.Handle()) if err != nil { if errors.Is(err, errs.ErrNotFound) { - return h, client, nil + break } return nil, nil, err } if !rel.IsSource { - return h, client, nil + break } target := &lvol.Handle{ClusterID: rel.TargetClusterID, PoolRef: rel.TargetPoolID, VolumeID: rel.TargetLvolID} targetClient, err := clusters.ReplicationClient(ctx, target.ClusterID) @@ -113,11 +130,37 @@ func resolveToLocalReplica( return nil, nil, err } h, client = target, targetClient + hops = append(hops, chainHop{h: h, client: client}) if rel.ActiveLvolID == "" || rel.ActiveLvolID == rel.TargetLvolID { - return h, client, nil + break + } + } + pick := chooseReplica(hops, local, flagged) + return pick.h, pick.client, nil +} + +// chainHop is one member of a replication chain, with the client of its +// cluster. +type chainHop struct { + h *lvol.Handle + client *atlascp.Client +} + +// chooseReplica picks the chain member a Replication RPC acts on: the last +// member on a local cluster when the secret marks local clusters (and the +// chain's end when none of the members is local, e.g. a volume that only +// ever lived elsewhere), else the chain's end. +func chooseReplica(hops []chainHop, local map[string]bool, flagged bool) chainHop { + end := hops[len(hops)-1] + if !flagged { + return end + } + for i := len(hops) - 1; i >= 0; i-- { + if local[hops[i].h.ClusterID] { + return hops[i] } } - return h, client, nil + return end } // EnableVolumeReplication attaches the volume to the policy named by the diff --git a/csi-driver/internal/csi/controller/replication_local_test.go b/csi-driver/internal/csi/controller/replication_local_test.go new file mode 100644 index 000000000..2dab4de1e --- /dev/null +++ b/csi-driver/internal/csi/controller/replication_local_test.go @@ -0,0 +1,50 @@ +package controller + +import ( + "testing" + + "github.com/simplyblock/atlas/lvol" +) + +// The chain of 2026-10-02 (realbed, WordPress): created on A (0aea, gone), +// failed over to B (6e83), relocated back to A (80e3), failed over to B +// (e3d4, the live primary). Ramen then made the old primary on A secondary. +func liveChain() []chainHop { + mk := func(cluster, id string) chainHop { return chainHop{h: &lvol.Handle{ClusterID: cluster, PoolRef: "p", VolumeID: id}} } + return []chainHop{mk("A", "0aea"), mk("B", "6e83"), mk("A", "80e3"), mk("B", "e3d4")} +} + +func TestChooseReplicaOnTheOldPrimarysSiteIsTheOldPrimaryNotTheLivePrimary(t *testing.T) { + got := chooseReplica(liveChain(), map[string]bool{"A": true}, true) + if got.h.VolumeID != "80e3" { + t.Fatalf("site A acts on %s, want 80e3 (its own, superseded primary); e3d4 is the live primary on B", got.h.VolumeID) + } +} + +func TestChooseReplicaOnTheNewPrimarysSiteIsTheLivePrimary(t *testing.T) { + got := chooseReplica(liveChain(), map[string]bool{"B": true}, true) + if got.h.VolumeID != "e3d4" { + t.Fatalf("site B acts on %s, want e3d4", got.h.VolumeID) + } +} + +func TestChooseReplicaWithoutLocalFlagsKeepsTheChainsEnd(t *testing.T) { + got := chooseReplica(liveChain(), nil, false) + if got.h.VolumeID != "e3d4" { + t.Fatalf("unflagged secret: %s, want the chain's end e3d4", got.h.VolumeID) + } +} + +func TestChooseReplicaWithNoLocalMemberKeepsTheChainsEnd(t *testing.T) { + got := chooseReplica(liveChain(), map[string]bool{"C": true}, true) + if got.h.VolumeID != "e3d4" { + t.Fatalf("no member on C: %s, want the chain's end e3d4", got.h.VolumeID) + } +} + +func TestChooseReplicaOfAVolumeWithoutARelationshipIsTheVolume(t *testing.T) { + one := liveChain()[:1] + if got := chooseReplica(one, map[string]bool{"B": true}, true); got.h.VolumeID != "0aea" { + t.Fatalf("got %s", got.h.VolumeID) + } +} diff --git a/operator/internal/controllers/cluster/storagecluster_controller.go b/operator/internal/controllers/cluster/storagecluster_controller.go index 6f30bef45..14a963ce3 100644 --- a/operator/internal/controllers/cluster/storagecluster_controller.go +++ b/operator/internal/controllers/cluster/storagecluster_controller.go @@ -201,6 +201,14 @@ type CSIClusterEntry struct { ClusterID string `json:"cluster_id"` ClusterEndpoint string `json:"cluster_endpoint"` ClusterSecret string `json:"cluster_secret"` + // Local marks a cluster this operator manages, i.e. the storage of the + // Kubernetes cluster the driver runs on, as opposed to an entry another + // site registered so that a failed-over volume's handle still resolves. + // The driver's csi-addons Replication RPCs act on the LOCAL member of a + // replication chain: a volume's PV keeps the handle it was created with + // across fail-overs, and the chain of relationships behind it alternates + // between the sites. + Local bool `json:"local,omitempty"` } // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storageclusters,verbs=get;list;watch;create;update;patch;delete @@ -1126,6 +1134,7 @@ func (r *StorageClusterReconciler) upsertCSICredentials( ClusterID: clusterID, ClusterEndpoint: r.API.Endpoint(ctx), ClusterSecret: clusterSecret, + Local: true, } for i := range creds.Clusters { if creds.Clusters[i].ClusterID == clusterID { From 5106e180ce770c1a9be38d3e68fd96d43d141d49 Mon Sep 17 00:00:00 2001 From: michael Date: Fri, 2 Oct 2026 22:41:53 +0300 Subject: [PATCH 168/206] csi-driver: lint (gofmt, goconst) for the local-replica resolver Co-Authored-By: Claude Fable 5.1 --- .../csi/controller/replication_local_test.go | 26 ++++++++++++------- 1 file changed, 17 insertions(+), 9 deletions(-) diff --git a/csi-driver/internal/csi/controller/replication_local_test.go b/csi-driver/internal/csi/controller/replication_local_test.go index 2dab4de1e..9dda81232 100644 --- a/csi-driver/internal/csi/controller/replication_local_test.go +++ b/csi-driver/internal/csi/controller/replication_local_test.go @@ -9,42 +9,50 @@ import ( // The chain of 2026-10-02 (realbed, WordPress): created on A (0aea, gone), // failed over to B (6e83), relocated back to A (80e3), failed over to B // (e3d4, the live primary). Ramen then made the old primary on A secondary. +const ( + siteA, siteB = "A", "B" + oldPrimaryOnA = "80e3" + livePrimaryOnB = "e3d4" +) + func liveChain() []chainHop { - mk := func(cluster, id string) chainHop { return chainHop{h: &lvol.Handle{ClusterID: cluster, PoolRef: "p", VolumeID: id}} } - return []chainHop{mk("A", "0aea"), mk("B", "6e83"), mk("A", "80e3"), mk("B", "e3d4")} + mk := func(cluster, id string) chainHop { + return chainHop{h: &lvol.Handle{ClusterID: cluster, PoolRef: "p", VolumeID: id}} + } + return []chainHop{mk(siteA, "0aea"), mk(siteB, "6e83"), mk(siteA, oldPrimaryOnA), mk(siteB, livePrimaryOnB)} } func TestChooseReplicaOnTheOldPrimarysSiteIsTheOldPrimaryNotTheLivePrimary(t *testing.T) { - got := chooseReplica(liveChain(), map[string]bool{"A": true}, true) - if got.h.VolumeID != "80e3" { + got := chooseReplica(liveChain(), map[string]bool{siteA: true}, true) + if got.h.VolumeID != oldPrimaryOnA { t.Fatalf("site A acts on %s, want 80e3 (its own, superseded primary); e3d4 is the live primary on B", got.h.VolumeID) } } func TestChooseReplicaOnTheNewPrimarysSiteIsTheLivePrimary(t *testing.T) { - got := chooseReplica(liveChain(), map[string]bool{"B": true}, true) - if got.h.VolumeID != "e3d4" { + got := chooseReplica(liveChain(), map[string]bool{siteB: true}, true) + if got.h.VolumeID != livePrimaryOnB { t.Fatalf("site B acts on %s, want e3d4", got.h.VolumeID) } } func TestChooseReplicaWithoutLocalFlagsKeepsTheChainsEnd(t *testing.T) { got := chooseReplica(liveChain(), nil, false) - if got.h.VolumeID != "e3d4" { + if got.h.VolumeID != livePrimaryOnB { t.Fatalf("unflagged secret: %s, want the chain's end e3d4", got.h.VolumeID) } } func TestChooseReplicaWithNoLocalMemberKeepsTheChainsEnd(t *testing.T) { got := chooseReplica(liveChain(), map[string]bool{"C": true}, true) - if got.h.VolumeID != "e3d4" { + if got.h.VolumeID != livePrimaryOnB { t.Fatalf("no member on C: %s, want the chain's end e3d4", got.h.VolumeID) } } func TestChooseReplicaOfAVolumeWithoutARelationshipIsTheVolume(t *testing.T) { one := liveChain()[:1] - if got := chooseReplica(one, map[string]bool{"B": true}, true); got.h.VolumeID != "0aea" { + if got := chooseReplica(one, map[string]bool{siteB: true}, true); got.h.VolumeID != "0aea" { t.Fatalf("got %s", got.h.VolumeID) } } From 18c0824117bd224d1352fe576adaa604dd99fe70 Mon Sep 17 00:00:00 2001 From: michael Date: Fri, 2 Oct 2026 23:03:35 +0300 Subject: [PATCH 169/206] csi-driver: demoting or detaching a reaped chain member succeeds The old primary of an unplanned fail-over is reaped by the control plane once its fail-over completed; when the site returns Ramen still demotes it and deletes its VolumeReplication. Resolved to that member, the demote and the detach got a 404 and the VR stayed Degraded (live 2026-10-02, site A, 80e3e748). A 404 on a member of a chain the backend records is nothing left to demote or detach: success. A 404 on a handle without any relationship stays NotFound. Co-Authored-By: Claude Fable 5.1 --- .../internal/csi/controller/replication.go | 45 +++++++++++++++---- .../csi/controller/replication_local_test.go | 30 +++++++++++++ 2 files changed, 67 insertions(+), 8 deletions(-) diff --git a/csi-driver/internal/csi/controller/replication.go b/csi-driver/internal/csi/controller/replication.go index f81a55ebc..5a18a2710 100644 --- a/csi-driver/internal/csi/controller/replication.go +++ b/csi-driver/internal/csi/controller/replication.go @@ -108,10 +108,23 @@ func volumeIDFrom(req volumeIDCarrier) string { func resolveToLocalReplica( ctx context.Context, h *lvol.Handle, client *atlascp.Client, ) (*lvol.Handle, *atlascp.Client, error) { + h, client, _, err := resolveReplica(ctx, h, client) + return h, client, err +} + +// resolveReplica is resolveToLocalReplica reporting also whether h has a +// replication relationship at all (known): the member it resolves to is then +// one of a chain the backend records, and a volume of that chain that no +// longer exists is a superseded, reaped old primary -- nothing left to demote +// or detach -- rather than an unknown handle. +func resolveReplica( + ctx context.Context, h *lvol.Handle, client *atlascp.Client, +) (*lvol.Handle, *atlascp.Client, bool, error) { local, flagged, err := clusters.Local() if err != nil { - return nil, nil, err + return nil, nil, false, err } + known := false hops := []chainHop{{h: h, client: client}} for range 8 { // one hop per past fail-over; capped far above any real chain rel, err := client.GetVolumeReplicationRelationship(ctx, h.Handle()) @@ -119,15 +132,16 @@ func resolveToLocalReplica( if errors.Is(err, errs.ErrNotFound) { break } - return nil, nil, err + return nil, nil, false, err } + known = true if !rel.IsSource { break } target := &lvol.Handle{ClusterID: rel.TargetClusterID, PoolRef: rel.TargetPoolID, VolumeID: rel.TargetLvolID} targetClient, err := clusters.ReplicationClient(ctx, target.ClusterID) if err != nil { - return nil, nil, err + return nil, nil, false, err } h, client = target, targetClient hops = append(hops, chainHop{h: h, client: client}) @@ -136,7 +150,16 @@ func resolveToLocalReplica( } } pick := chooseReplica(hops, local, flagged) - return pick.h, pick.client, nil + return pick.h, pick.client, known, nil +} + +// reapedChainMember is whether a Replication verb on a resolved chain member +// found the volume gone (404): a superseded old primary the control plane +// has reaped after its fail-over completed (deferred removal; live +// 2026-10-02 on site A). Demoting or detaching it is a no-op that succeeds; +// a 404 on a handle with no relationship stays NotFound. +func reapedChainMember(known bool, ce classifiedError) bool { + return known && status.Code(ce) == codes.NotFound } // chainHop is one member of a replication chain, with the client of its @@ -267,12 +290,14 @@ func (cs *Server) DisableVolumeReplication( if err != nil { return nil, status.Error(codes.Unavailable, err.Error()) } - h, client, err = resolveToLocalReplica(ctx, h, client) + h, client, known, err := resolveReplica(ctx, h, client) if err != nil { return nil, status.Error(codes.Unavailable, err.Error()) } if err := client.DisableVolumeReplication(ctx, h.Handle()); err != nil { - return nil, classifyDisableVolumeReplicationError(err) + if ce := classifyDisableVolumeReplicationError(err); !reapedChainMember(known, ce) { + return nil, ce + } } return &replication.DisableVolumeReplicationResponse{}, nil } @@ -418,13 +443,17 @@ func (cs *Server) DemoteVolume( if err != nil { return nil, status.Error(codes.Unavailable, err.Error()) } - h, client, err = resolveToLocalReplica(ctx, h, client) + h, client, known, err := resolveReplica(ctx, h, client) if err != nil { return nil, status.Error(codes.Unavailable, err.Error()) } done, err := client.DemoteVolume(ctx, h.Handle()) if err != nil { - return nil, classifyDemoteVolumeError(err) + ce := classifyDemoteVolumeError(err) + if !reapedChainMember(known, ce) { + return nil, ce + } + done = true } if !done { return nil, status.Error(codes.Aborted, "demote is still converging") diff --git a/csi-driver/internal/csi/controller/replication_local_test.go b/csi-driver/internal/csi/controller/replication_local_test.go index 9dda81232..96fbe9e60 100644 --- a/csi-driver/internal/csi/controller/replication_local_test.go +++ b/csi-driver/internal/csi/controller/replication_local_test.go @@ -1,8 +1,10 @@ package controller import ( + "context" "testing" + "github.com/csi-addons/spec/lib/go/replication" "github.com/simplyblock/atlas/lvol" ) @@ -56,3 +58,31 @@ func TestChooseReplicaOfAVolumeWithoutARelationshipIsTheVolume(t *testing.T) { t.Fatalf("got %s", got.h.VolumeID) } } + +// The old primary of an unplanned fail-over is reaped by the control plane +// once its fail-over completed; Ramen still demotes it (and deletes its VR) +// when the site returns. Nothing is left to demote: success, not NotFound +// (live 2026-10-02: the VR on site A stayed Degraded on a 404). +func TestDemoteAndDisableOfAReapedChainMemberSucceed(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newReplicationTestServer(t, mock) + gone := "99999999-aaaa-bbbb-cccc-dddddddddddd" + mock.replicationRelationship[testReplVolumeID] = map[string]any{ + "replication_id": "aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee", + "direction": "to_target", + "mode": "failover", + "state": "failed_over", + "is_source": true, + "source_cluster_id": sanityClusterID, "source_lvol_id": testReplVolumeID, + "target_cluster_id": sanityClusterID, "target_pool_id": sanityPoolUUID, "target_lvol_id": gone, + "active_lvol_id": gone, + "target_nqn": "nqn.test", "target_ns_id": 1, + } + if _, err := cs.DemoteVolume(context.Background(), &replication.DemoteVolumeRequest{VolumeId: testReplVolID}); err != nil { + t.Fatalf("demote of a reaped chain member: %v", err) + } + if _, err := cs.DisableVolumeReplication(context.Background(), &replication.DisableVolumeReplicationRequest{VolumeId: testReplVolID}); err != nil { + t.Fatalf("disable of a reaped chain member: %v", err) + } +} From 9ee0e79de87bef6739fe448d3a5303684dafd1d4 Mon Sep 17 00:00:00 2001 From: michael Date: Fri, 2 Oct 2026 23:26:10 +0300 Subject: [PATCH 170/206] csi-driver: Resync and the status read of a reaped local member go to the chain's active end After an unplanned fail-over the recovered site's old primary is reaped once its fail-over completed. Ramen still resyncs that site's secondary VR (and reads its status). Addressed to the reaped member both got a 404 and the VR stayed Degraded (live 2026-10-02, site A). sbcli's replication_failback is addressed to the failed-over clone and, given the original source cluster, re-aims the clone's replication at that site's node so only the delta ships: Resync falls back to the chain's active end with the local cluster as the fail-back source, and the status read follows. Co-Authored-By: Claude Fable 5.1 --- .../internal/csi/controller/replication.go | 88 +++++++++++++++++-- .../csi/controller/replication_local_test.go | 14 ++- 2 files changed, 92 insertions(+), 10 deletions(-) diff --git a/csi-driver/internal/csi/controller/replication.go b/csi-driver/internal/csi/controller/replication.go index 5a18a2710..335774a6c 100644 --- a/csi-driver/internal/csi/controller/replication.go +++ b/csi-driver/internal/csi/controller/replication.go @@ -9,6 +9,7 @@ package controller import ( "context" "errors" + "sort" "github.com/csi-addons/spec/lib/go/replication" "google.golang.org/grpc/codes" @@ -120,10 +121,22 @@ func resolveToLocalReplica( func resolveReplica( ctx context.Context, h *lvol.Handle, client *atlascp.Client, ) (*lvol.Handle, *atlascp.Client, bool, error) { + hops, known, err := resolveChain(ctx, h, client) + if err != nil { + return nil, nil, false, err + } local, flagged, err := clusters.Local() if err != nil { return nil, nil, false, err } + pick := chooseReplica(hops, local, flagged) + return pick.h, pick.client, known, nil +} + +// resolveChain walks the replication chain behind h: h itself, then every +// IsSource->target hop to the active end. known is whether h has a +// relationship at all. +func resolveChain(ctx context.Context, h *lvol.Handle, client *atlascp.Client) ([]chainHop, bool, error) { known := false hops := []chainHop{{h: h, client: client}} for range 8 { // one hop per past fail-over; capped far above any real chain @@ -132,7 +145,7 @@ func resolveReplica( if errors.Is(err, errs.ErrNotFound) { break } - return nil, nil, false, err + return nil, false, err } known = true if !rel.IsSource { @@ -141,7 +154,7 @@ func resolveReplica( target := &lvol.Handle{ClusterID: rel.TargetClusterID, PoolRef: rel.TargetPoolID, VolumeID: rel.TargetLvolID} targetClient, err := clusters.ReplicationClient(ctx, target.ClusterID) if err != nil { - return nil, nil, false, err + return nil, false, err } h, client = target, targetClient hops = append(hops, chainHop{h: h, client: client}) @@ -149,8 +162,36 @@ func resolveReplica( break } } - pick := chooseReplica(hops, local, flagged) - return pick.h, pick.client, known, nil + return hops, known, nil +} + +// activeEndFallback is where a Resync or a status read goes when the local +// member of the chain is reaped: the chain's active end -- the live primary +// on the other site -- and, as the cluster to fail back to, the local one. +// sbcli's replication_failback is addressed to the failed-over clone and +// re-aims its replication at the original site's node (the recovered-source +// case: only the delta ships), which is exactly the fail-back of a site that +// lost its primary (live 2026-10-02, site A after the unplanned fail-over of +// WordPress: the old primary 80e3e748 was reaped, the clone e3d439ca on B +// holds the data). +func activeEndFallback(hops []chainHop, sourceClusterID string) (chainHop, string) { + end := hops[len(hops)-1] + if local, flagged, err := clusters.Local(); err == nil && flagged { + ids := make([]string, 0, len(local)) + for id := range local { + ids = append(ids, id) + } + sort.Strings(ids) + for _, hop := range hops { + if local[hop.h.ClusterID] { + return end, hop.h.ClusterID + } + } + if len(ids) > 0 { + return end, ids[0] + } + } + return end, sourceClusterID } // reapedChainMember is whether a Replication verb on a resolved chain member @@ -339,13 +380,29 @@ func (cs *Server) GetVolumeReplicationInfo( if err != nil { return nil, status.Error(codes.Unavailable, err.Error()) } - h, client, err = resolveToLocalReplica(ctx, h, client) + hops, known, err := resolveChain(ctx, h, client) + if err != nil { + return nil, status.Error(codes.Unavailable, err.Error()) + } + local, flagged, err := clusters.Local() if err != nil { return nil, status.Error(codes.Unavailable, err.Error()) } + pick := chooseReplica(hops, local, flagged) + h, client = pick.h, pick.client info, err := client.GetVolumeReplicationInfo(ctx, h.Handle()) if err != nil { - return nil, classifyGetVolumeReplicationInfoError(err) + ce := classifyGetVolumeReplicationInfoError(err) + if !reapedChainMember(known, ce) { + return nil, ce + } + // The local member is reaped: the status that matters is the active + // end's, replicating back to this site after a Resync. + end, _ := activeEndFallback(hops, "") + h, client = end.h, end.client + if info, err = client.GetVolumeReplicationInfo(ctx, h.Handle()); err != nil { + return nil, classifyGetVolumeReplicationInfoError(err) + } } resp := &replication.GetVolumeReplicationInfoResponse{} if info.LastReplicatedAt != nil { @@ -500,13 +557,28 @@ func (cs *Server) ResyncVolume( if err != nil { return nil, status.Error(codes.Unavailable, err.Error()) } - h, client, err = resolveToLocalReplica(ctx, h, client) + hops, known, err := resolveChain(ctx, h, client) if err != nil { return nil, status.Error(codes.Unavailable, err.Error()) } + local, flagged, err := clusters.Local() + if err != nil { + return nil, status.Error(codes.Unavailable, err.Error()) + } + pick := chooseReplica(hops, local, flagged) + h, client = pick.h, pick.client sourceClusterID := req.GetParameters()[sourceClusterIDParam] if err := client.ResyncVolume(ctx, h.Handle(), sourceClusterID); err != nil { - return nil, classifyResyncVolumeError(err) + ce := classifyResyncVolumeError(err) + if !reapedChainMember(known, ce) { + return nil, ce + } + var end chainHop + end, sourceClusterID = activeEndFallback(hops, sourceClusterID) + h, client = end.h, end.client + if err := client.ResyncVolume(ctx, h.Handle(), sourceClusterID); err != nil { + return nil, classifyResyncVolumeError(err) + } } info, err := client.GetVolumeReplicationInfo(ctx, h.Handle()) if err != nil { diff --git a/csi-driver/internal/csi/controller/replication_local_test.go b/csi-driver/internal/csi/controller/replication_local_test.go index 96fbe9e60..83f93c855 100644 --- a/csi-driver/internal/csi/controller/replication_local_test.go +++ b/csi-driver/internal/csi/controller/replication_local_test.go @@ -79,10 +79,20 @@ func TestDemoteAndDisableOfAReapedChainMemberSucceed(t *testing.T) { "active_lvol_id": gone, "target_nqn": "nqn.test", "target_ns_id": 1, } - if _, err := cs.DemoteVolume(context.Background(), &replication.DemoteVolumeRequest{VolumeId: testReplVolID}); err != nil { + ctx := context.Background() + if _, err := cs.DemoteVolume(ctx, &replication.DemoteVolumeRequest{VolumeId: testReplVolID}); err != nil { t.Fatalf("demote of a reaped chain member: %v", err) } - if _, err := cs.DisableVolumeReplication(context.Background(), &replication.DisableVolumeReplicationRequest{VolumeId: testReplVolID}); err != nil { + disable := &replication.DisableVolumeReplicationRequest{VolumeId: testReplVolID} + if _, err := cs.DisableVolumeReplication(ctx, disable); err != nil { t.Fatalf("disable of a reaped chain member: %v", err) } } + +func TestActiveEndFallbackAimsAtTheChainsEndFromTheLocalSite(t *testing.T) { + // Not flagged in the test secret: the class parameter stays. + end, src := activeEndFallback(liveChain(), "param") + if end.h.VolumeID != livePrimaryOnB || src != "param" { + t.Fatalf("end %s source %s", end.h.VolumeID, src) + } +} From 65279a1efd3165195f057a55329bd76feb8123ff Mon Sep 17 00:00:00 2001 From: michael Date: Fri, 2 Oct 2026 23:39:35 +0300 Subject: [PATCH 171/206] operator: register every cluster of the control plane with the CSI driver A volume replicated to another site is promoted there under a PV that keeps the handle it was created with, which names the cluster the volume came from: that site's driver must be able to reach the control plane for that cluster too. Until now each operator wrote only its own cluster into simplyblock-csi-secret-v2, and the cross-registration was a manual merge of the sites' secrets (realbed deploy.sh csi). Each operator now writes, next to its own cluster (marked local, the flag the driver's replication-chain resolver keys on), one entry per other cluster the control plane lists, with the secret the list carries -- empty when the control plane withholds it, and the driver then authenticates with its API token. Entries another operator on the same Kubernetes cluster marked local are left alone; a foreign entry whose cluster the control plane no longer lists is dropped. The own entry never waits on the list: when the list cannot be read, peers are registered on the next sync. Co-Authored-By: Claude Fable 5.1 --- .../controllers/cluster/controlplane.go | 21 ++++- .../cluster/csicredentials_test.go | 64 +++++++++++++ .../controllers/cluster/helpers_test.go | 8 ++ .../cluster/storagecluster_controller.go | 91 +++++++++++++++++-- 4 files changed, 171 insertions(+), 13 deletions(-) create mode 100644 operator/internal/controllers/cluster/csicredentials_test.go diff --git a/operator/internal/controllers/cluster/controlplane.go b/operator/internal/controllers/cluster/controlplane.go index 1b1ab5df9..67470e1f7 100644 --- a/operator/internal/controllers/cluster/controlplane.go +++ b/operator/internal/controllers/cluster/controlplane.go @@ -53,6 +53,9 @@ type ControlPlane interface { // reports false rather than an error when there is none, because "no such // cluster" is the ordinary answer on the creation path. ClusterByName(ctx context.Context, name string) (utils.ClusterListEntry, bool, error) + // Clusters lists every cluster of the control plane, as ClusterByName + // reads them. + Clusters(ctx context.Context) ([]utils.ClusterListEntry, error) // DeleteCluster is retried until it succeeds, and the finalizer is not // removed before it does. @@ -151,16 +154,24 @@ func (c *httpControlPlane) Cluster( return webapi.ParseClusterResponse(body) } -func (c *httpControlPlane) ClusterByName( - ctx context.Context, name string, -) (utils.ClusterListEntry, bool, error) { +func (c *httpControlPlane) Clusters(ctx context.Context) ([]utils.ClusterListEntry, error) { body, err := c.call(c.adminContext(ctx), http.MethodGet, "/api/v2/clusters/", nil) if err != nil { - return utils.ClusterListEntry{}, false, err + return nil, err } var entries []utils.ClusterListEntry if err := json.Unmarshal(body, &entries); err != nil { - return utils.ClusterListEntry{}, false, fmt.Errorf("read the cluster list: %w", err) + return nil, fmt.Errorf("read the cluster list: %w", err) + } + return entries, nil +} + +func (c *httpControlPlane) ClusterByName( + ctx context.Context, name string, +) (utils.ClusterListEntry, bool, error) { + entries, err := c.Clusters(ctx) + if err != nil { + return utils.ClusterListEntry{}, false, err } for _, entry := range entries { if entry.Name == name { diff --git a/operator/internal/controllers/cluster/csicredentials_test.go b/operator/internal/controllers/cluster/csicredentials_test.go new file mode 100644 index 000000000..8e30bd80b --- /dev/null +++ b/operator/internal/controllers/cluster/csicredentials_test.go @@ -0,0 +1,64 @@ +package cluster + +import ( + "testing" + + "github.com/simplyblock/simplyblock-operator/internal/utils" +) + +func entryIDs(c CSICredentials) map[string]CSIClusterEntry { + out := map[string]CSIClusterEntry{} + for _, e := range c.Clusters { + out[e.ClusterID] = e + } + return out +} + +func TestMergeCSICredentialsRegistersEveryClusterOfTheControlPlane(t *testing.T) { + // Site A's operator manages cluster A; the control plane also runs + // cluster B (site B). A volume failed over from B to A arrives under a + // PV whose handle names B: the driver on A must reach B's cluster too. + creds := CSICredentials{} + own := CSIClusterEntry{ClusterID: "A", ClusterEndpoint: "http://cp:5000", ClusterSecret: "sa", Local: true} + peers := []utils.ClusterListEntry{{UUID: "A", Secret: "sa"}, {UUID: "B", Secret: "sb"}} + mergeCSICredentials(&creds, own, peers, true) + got := entryIDs(creds) + if len(got) != 2 || !got["A"].Local || got["B"].Local || got["B"].ClusterSecret != "sb" || got["B"].ClusterEndpoint != "http://cp:5000" { + t.Fatalf("entries %+v", creds.Clusters) + } +} + +func TestMergeCSICredentialsKeepsAnotherLocalEntryAndPrunesStaleForeignOnes(t *testing.T) { + creds := CSICredentials{Clusters: []CSIClusterEntry{ + {ClusterID: "A2", ClusterEndpoint: "http://cp:5000", ClusterSecret: "sa2", Local: true}, // another operator here + {ClusterID: "OLD", ClusterEndpoint: "http://cp:5000", ClusterSecret: "x"}, // a cluster since removed + {ClusterID: "B", ClusterEndpoint: "http://cp:5000", ClusterSecret: "kept"}, + }} + own := CSIClusterEntry{ClusterID: "A", ClusterEndpoint: "http://cp:5000", ClusterSecret: "sa", Local: true} + // The list withholds B's secret: the recorded one stays. + peers := []utils.ClusterListEntry{{UUID: "A"}, {UUID: "A2"}, {UUID: "B"}} + mergeCSICredentials(&creds, own, peers, true) + got := entryIDs(creds) + if _, stale := got["OLD"]; stale { + t.Fatal("a cluster the control plane no longer lists stays registered") + } + if !got["A2"].Local || got["A2"].ClusterSecret != "sa2" { + t.Fatalf("the other operator's local entry changed: %+v", got["A2"]) + } + if got["B"].Local || got["B"].ClusterSecret != "kept" { + t.Fatalf("B: %+v", got["B"]) + } + if !got["A"].Local { + t.Fatalf("own: %+v", got["A"]) + } +} + +func TestMergeCSICredentialsWithoutTheListOnlyWritesTheOwnEntry(t *testing.T) { + creds := CSICredentials{Clusters: []CSIClusterEntry{{ClusterID: "B", ClusterSecret: "sb"}}} + own := CSIClusterEntry{ClusterID: "A", ClusterSecret: "sa", Local: true} + mergeCSICredentials(&creds, own, nil, false) + got := entryIDs(creds) + if len(got) != 2 || got["B"].ClusterSecret != "sb" || !got["A"].Local { + t.Fatalf("entries %+v", creds.Clusters) + } +} diff --git a/operator/internal/controllers/cluster/helpers_test.go b/operator/internal/controllers/cluster/helpers_test.go index 402cff295..bfb8989be 100644 --- a/operator/internal/controllers/cluster/helpers_test.go +++ b/operator/internal/controllers/cluster/helpers_test.go @@ -213,6 +213,7 @@ type fakeControlPlane struct { create func(utils.ClusterAddParams) (webapi.ClusterResponse, error) cluster func(string) (webapi.ClusterResponse, error) byName func(string) (utils.ClusterListEntry, bool, error) + clusters func() ([]utils.ClusterListEntry, error) deleteCall func(string) error activate func(string) error expand func(string) error @@ -277,6 +278,13 @@ func (f *fakeControlPlane) ClusterByName( return f.byName(name) } +func (f *fakeControlPlane) Clusters(context.Context) ([]utils.ClusterListEntry, error) { + if f.clusters == nil { + return nil, nil + } + return f.clusters() +} + func (f *fakeControlPlane) DeleteCluster(_ context.Context, clusterID string) error { f.deleteCalls++ if f.deleteCall == nil { diff --git a/operator/internal/controllers/cluster/storagecluster_controller.go b/operator/internal/controllers/cluster/storagecluster_controller.go index 6f30bef45..ac4fb7c00 100644 --- a/operator/internal/controllers/cluster/storagecluster_controller.go +++ b/operator/internal/controllers/cluster/storagecluster_controller.go @@ -201,6 +201,14 @@ type CSIClusterEntry struct { ClusterID string `json:"cluster_id"` ClusterEndpoint string `json:"cluster_endpoint"` ClusterSecret string `json:"cluster_secret"` + // Local marks a cluster this operator manages, i.e. the storage of the + // Kubernetes cluster the driver runs on, as opposed to an entry another + // site registered so that a failed-over volume's handle still resolves. + // The driver's csi-addons Replication RPCs act on the LOCAL member of a + // replication chain: a volume's PV keeps the handle it was created with + // across fail-overs, and the chain of relationships behind it alternates + // between the sites. + Local bool `json:"local,omitempty"` } // +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storageclusters,verbs=get;list;watch;create;update;patch;delete @@ -1117,24 +1125,91 @@ func (r *StorageClusterReconciler) clusterSecret( } // upsertCSICredentials adds or replaces this cluster's entry in the aggregate -// Secret the CSI driver reads. +// Secret the CSI driver reads, and registers every other cluster of the +// same control plane beside it. +// +// A volume replicated to another site is promoted there under a PV that +// keeps the handle it was created with -- the handle names the cluster the +// volume came from -- so that site's driver must be able to reach the +// control plane for that cluster too. Each operator therefore writes, next +// to its own cluster (Local), one entry per other cluster the control plane +// lists, with the secret the list carries (empty when the control plane +// withholds it: the driver then authenticates with its API token). Entries +// another operator on this Kubernetes cluster marked Local are left alone; +// a foreign entry whose cluster the control plane no longer lists is +// dropped. Until this, the cross-registration was a manual merge of the +// sites' secrets (realbed deploy.sh csi). func (r *StorageClusterReconciler) upsertCSICredentials( ctx context.Context, clusterID, clusterSecret string, ) error { + endpoint := r.API.Endpoint(ctx) + peers, err := r.API.Clusters(ctx) + if err != nil { + // The own entry must never wait on the peer list: the driver reaches + // this cluster through it. Peers are registered on the next sync. + logf.FromContext(ctx).Info("the control plane's cluster list could not be read; "+ + "other clusters are registered with the CSI driver on the next sync", "error", err.Error()) + peers = nil + } return r.editCSICredentials(ctx, func(creds *CSICredentials) { - entry := CSIClusterEntry{ + mergeCSICredentials(creds, CSIClusterEntry{ ClusterID: clusterID, - ClusterEndpoint: r.API.Endpoint(ctx), + ClusterEndpoint: endpoint, ClusterSecret: clusterSecret, + Local: true, + }, peers, err == nil) + }) +} + +// mergeCSICredentials writes own into creds and, when the control plane's +// cluster list was read, one non-local entry per other cluster of that +// list, pruning non-local entries the list no longer has. +func mergeCSICredentials(creds *CSICredentials, own CSIClusterEntry, peers []utils.ClusterListEntry, listed bool) { + replaced := false + for i := range creds.Clusters { + if creds.Clusters[i].ClusterID == own.ClusterID { + creds.Clusters[i] = own + replaced = true } + } + if !replaced { + creds.Clusters = append(creds.Clusters, own) + } + if !listed { + return + } + known := map[string]bool{} + for _, p := range peers { + known[p.UUID] = true + if p.UUID == own.ClusterID { + continue + } + entry := CSIClusterEntry{ClusterID: p.UUID, ClusterEndpoint: own.ClusterEndpoint, ClusterSecret: p.Secret} + found := false for i := range creds.Clusters { - if creds.Clusters[i].ClusterID == clusterID { - creds.Clusters[i] = entry - return + if creds.Clusters[i].ClusterID != p.UUID { + continue + } + found = true + if creds.Clusters[i].Local { + break // another operator here manages it; its entry stands } + if entry.ClusterSecret == "" { + entry.ClusterSecret = creds.Clusters[i].ClusterSecret + } + creds.Clusters[i] = entry } - creds.Clusters = append(creds.Clusters, entry) - }) + if !found { + creds.Clusters = append(creds.Clusters, entry) + } + } + kept := creds.Clusters[:0] + for _, e := range creds.Clusters { + if e.Local || known[e.ClusterID] { + kept = append(kept, e) + } + } + creds.Clusters = kept } // removeCSICredentials drops this cluster's entry from it. From bede3ddee6f5f688d8eb10ac321ce1f8cad6e5a3 Mon Sep 17 00:00:00 2001 From: michael Date: Sat, 3 Oct 2026 01:05:14 +0300 Subject: [PATCH 172/206] csi-driver: PromoteVolume addresses the chain's active end, not the local member The control plane's failover endpoint takes the volume that holds the data (the pairing's source) and creates the clone on the replication target. The local member on the promoting site is the volume being replaced -- demoted, or already reaped: on 2026-10-02 the fail-back promote on site A hit the reaped 80e3e748 and 404ed while the live primary on B held the data. Co-Authored-By: Claude Fable 5.1 --- csi-driver/internal/csi/controller/replication.go | 12 +++++++++++- 1 file changed, 11 insertions(+), 1 deletion(-) diff --git a/csi-driver/internal/csi/controller/replication.go b/csi-driver/internal/csi/controller/replication.go index 335774a6c..ec12e7d71 100644 --- a/csi-driver/internal/csi/controller/replication.go +++ b/csi-driver/internal/csi/controller/replication.go @@ -446,10 +446,20 @@ func (cs *Server) PromoteVolume( if err != nil { return nil, status.Error(codes.Unavailable, err.Error()) } - h, client, err = resolveToLocalReplica(ctx, h, client) + // Promote is addressed to the chain's ACTIVE END, never to the local + // member: the control plane's failover endpoint takes the volume that + // currently holds the data (the source of the pairing) and creates the + // clone on the replication target -- this site. The local member here + // is the volume being replaced: a demoted old primary, or one the + // control plane already reaped (2026-10-02: promote on site A hit the + // reaped 80e3e748 and 404ed while the live primary e3d439ca on B held + // the data). + hops, _, err := resolveChain(ctx, h, client) if err != nil { return nil, status.Error(codes.Unavailable, err.Error()) } + end := hops[len(hops)-1] + h, client = end.h, end.client if err := client.PromoteVolume(ctx, h.Handle(), req.GetForce()); err != nil { return nil, classifyPromoteVolumeError(err) } From 2b6e4284488c872b18929c305ff1035381e44bd0 Mon Sep 17 00:00:00 2001 From: michael Date: Sat, 3 Oct 2026 01:58:58 +0300 Subject: [PATCH 173/206] csi-driver: Resync and the status read address the chain's active end sbcli's replication_failback takes the volume that holds the data (the failed-over clone) and, given the recovered site as source cluster, re-aims its replication at that site's node. Addressed to the local member -- the demoted or reaped old primary -- it configured nothing for the live clone: after the unplanned fail-over of Gitea the clones on A had no replication and lastGroupSyncTime stayed empty (2026-10-02). The pairing's status is the active end's for the same reason. Demote and Disable keep the local member; Promote already uses the active end. Co-Authored-By: Claude Fable 5.1 --- .../internal/csi/controller/replication.go | 55 +++++++------------ 1 file changed, 19 insertions(+), 36 deletions(-) diff --git a/csi-driver/internal/csi/controller/replication.go b/csi-driver/internal/csi/controller/replication.go index ec12e7d71..e7b483ecf 100644 --- a/csi-driver/internal/csi/controller/replication.go +++ b/csi-driver/internal/csi/controller/replication.go @@ -380,29 +380,19 @@ func (cs *Server) GetVolumeReplicationInfo( if err != nil { return nil, status.Error(codes.Unavailable, err.Error()) } - hops, known, err := resolveChain(ctx, h, client) - if err != nil { - return nil, status.Error(codes.Unavailable, err.Error()) - } - local, flagged, err := clusters.Local() + // The pairing's status is the ACTIVE END's: the volume that holds the + // data and replicates. On the primary site that is the local volume; on + // the secondary site the local member is the demoted or reaped old + // primary, whose status says nothing about the pipe back to this site. + hops, _, err := resolveChain(ctx, h, client) if err != nil { return nil, status.Error(codes.Unavailable, err.Error()) } - pick := chooseReplica(hops, local, flagged) - h, client = pick.h, pick.client + end := hops[len(hops)-1] + h, client = end.h, end.client info, err := client.GetVolumeReplicationInfo(ctx, h.Handle()) if err != nil { - ce := classifyGetVolumeReplicationInfoError(err) - if !reapedChainMember(known, ce) { - return nil, ce - } - // The local member is reaped: the status that matters is the active - // end's, replicating back to this site after a Resync. - end, _ := activeEndFallback(hops, "") - h, client = end.h, end.client - if info, err = client.GetVolumeReplicationInfo(ctx, h.Handle()); err != nil { - return nil, classifyGetVolumeReplicationInfoError(err) - } + return nil, classifyGetVolumeReplicationInfoError(err) } resp := &replication.GetVolumeReplicationInfoResponse{} if info.LastReplicatedAt != nil { @@ -567,28 +557,21 @@ func (cs *Server) ResyncVolume( if err != nil { return nil, status.Error(codes.Unavailable, err.Error()) } - hops, known, err := resolveChain(ctx, h, client) - if err != nil { - return nil, status.Error(codes.Unavailable, err.Error()) - } - local, flagged, err := clusters.Local() + // Resync is addressed to the chain's ACTIVE END with this site as the + // cluster to fail back to: sbcli's replication_failback takes the volume + // that holds the data (the failed-over clone) and re-aims its replication + // at the recovered site's node, shipping only the delta. The local member + // is the demoted or reaped old primary; re-aiming IT configured nothing + // for the live clone (2026-10-02, Gitea after the unplanned fail-over: + // the clones on A had no replication, lastGroupSyncTime stayed empty). + hops, _, err := resolveChain(ctx, h, client) if err != nil { return nil, status.Error(codes.Unavailable, err.Error()) } - pick := chooseReplica(hops, local, flagged) - h, client = pick.h, pick.client - sourceClusterID := req.GetParameters()[sourceClusterIDParam] + end, sourceClusterID := activeEndFallback(hops, req.GetParameters()[sourceClusterIDParam]) + h, client = end.h, end.client if err := client.ResyncVolume(ctx, h.Handle(), sourceClusterID); err != nil { - ce := classifyResyncVolumeError(err) - if !reapedChainMember(known, ce) { - return nil, ce - } - var end chainHop - end, sourceClusterID = activeEndFallback(hops, sourceClusterID) - h, client = end.h, end.client - if err := client.ResyncVolume(ctx, h.Handle(), sourceClusterID); err != nil { - return nil, classifyResyncVolumeError(err) - } + return nil, classifyResyncVolumeError(err) } info, err := client.GetVolumeReplicationInfo(ctx, h.Handle()) if err != nil { From 8988179c6238edc5a965b5c9abc32608d73f388f Mon Sep 17 00:00:00 2001 From: michael Date: Sat, 3 Oct 2026 04:27:13 +0300 Subject: [PATCH 174/206] volstack/csi-driver: a plan that names no filesystem mounts what the volume carries A PersistentVolume without fsType -- a static PV adopting an existing volume (the storage operator's test fail-over clone), or one Ramen restored without the field -- gave the node plugin an empty fsType, which it turned into "ext4" before the device was looked at; the filesystem layer then refused every such XFS volume ("the volume carries xfs and this plan asks for ext4"), and the bubble VM never started (realbed 2026-10-03). An empty FsType now expresses no opinion: the layer mounts the filesystem the device carries, formats a blank device as DefaultFsType (ext4), records what it found as its params, and the node plugin remembers that type for the volume. A named filesystem keeps refusing a device carrying another. Co-Authored-By: Claude Fable 5.1 --- atlas-lib/volstack/layers/filesystem.go | 64 ++++++++++++++++++-- atlas-lib/volstack/layers/filesystem_test.go | 38 ++++++++++++ atlas-lib/volstack/plans/node.go | 1 + atlas-lib/volstack/plans/plans.go | 7 ++- csi-driver/internal/csi/node/plan.go | 7 ++- csi-driver/internal/csi/node/stage.go | 34 +++++++++-- 6 files changed, 139 insertions(+), 12 deletions(-) diff --git a/atlas-lib/volstack/layers/filesystem.go b/atlas-lib/volstack/layers/filesystem.go index 7d799c903..10d8a472f 100644 --- a/atlas-lib/volstack/layers/filesystem.go +++ b/atlas-lib/volstack/layers/filesystem.go @@ -63,8 +63,19 @@ type FilesystemConfig struct { // formatted as, and it is also the only filesystem the layer will mount: a // device carrying another is refused, because neither reformatting it nor // serving what is on it is safe. + // + // Empty, the plan expresses no opinion: a device carrying a filesystem is + // mounted as what it carries, and a blank one is formatted as DefaultFsType. + // That is what a PersistentVolume without fsType means -- a static PV that + // adopts an existing volume (a test fail-over's clone, 2026-10-03), or one + // Ramen restored without the field -- and turning it into "ext4" before the + // device was looked at made the layer refuse every such XFS volume. FsType string + // DefaultFsType is what a blank device is formatted as when FsType names + // nothing. Empty means ext4. + DefaultFsType string + // StagingPath is where the filesystem is mounted. StagingPath string @@ -107,6 +118,43 @@ type FilesystemConfig struct { // the volume is, and refuses every other device. type Filesystem struct { cfg FilesystemConfig + // detected is the filesystem found on the device when the plan named none. + detected string +} + +// effective is the filesystem this layer acts with: the one the plan named, +// else the one the device carries, else the default for a blank device. +func (f *Filesystem) effective(reading blockdev.Reading) string { + if f.cfg.FsType != "" { + return f.cfg.FsType + } + if reading.Content == blockdev.ContentFilesystem && reading.Type != "" { + f.detected = reading.Type + return reading.Type + } + if f.detected != "" { + return f.detected + } + return f.defaultFsType() +} + +// known is the filesystem this layer stands for once it has acted: named, +// detected, or the default it formats with. +func (f *Filesystem) known() string { + if f.cfg.FsType != "" { + return f.cfg.FsType + } + if f.detected != "" { + return f.detected + } + return f.defaultFsType() +} + +func (f *Filesystem) defaultFsType() string { + if f.cfg.DefaultFsType != "" { + return f.cfg.DefaultFsType + } + return "ext4" } // NewFilesystem returns the filesystem layer for one volume. @@ -206,11 +254,12 @@ func (f *Filesystem) Ensure(ctx context.Context, below volstack.Artifact) (volst // decide the state and again to act on it. The reading itself is not needed // past that, because the filesystem to act on is the one the plan named and // observe has already refused every device carrying another. - state, _, own, err := f.observe(ctx, below) + state, reading, own, err := f.observe(ctx, below) if err != nil { return volstack.Artifact{}, err } if state == volstack.StateReady { + f.effective(reading) return own, nil } @@ -219,7 +268,7 @@ func (f *Filesystem) Ensure(ctx context.Context, below volstack.Artifact) (volst // disagreement, which is the point: the only two ways to reconcile one are to // reformat, which destroys the volume, and to serve the other filesystem, // which hides the misconfiguration until something else acts on it. - fsType := f.cfg.FsType + fsType := f.effective(reading) if state == volstack.StateAbsent { if err := f.cfg.Ops.Format(ctx, dev.Path, fsType, f.formatOptions(below)); err != nil { return volstack.Artifact{}, fmt.Errorf("filesystem: format %s as %s: %w", dev.Path, fsType, err) @@ -321,7 +370,7 @@ func (f *Filesystem) Heal(ctx context.Context, below, _ volstack.Artifact) error return err } - if err := f.cfg.Ops.Mount(ctx, dev.Path, f.cfg.StagingPath, f.cfg.FsType, f.mountFlags()); err != nil { + if err := f.cfg.Ops.Mount(ctx, dev.Path, f.cfg.StagingPath, f.effective(reading), f.mountFlags()); err != nil { return fmt.Errorf("filesystem: remount %s at %s: %w", dev.Path, f.cfg.StagingPath, err) } return nil @@ -379,7 +428,7 @@ type FilesystemParams struct { // recorded is the one the volume asked for, and a teardown needs no more than // that: what is actually on the device is read from the device. func (f *Filesystem) Params() any { - return FilesystemParams{FsType: f.cfg.FsType} + return FilesystemParams{FsType: f.known()} } // agrees reports whether the filesystem on the device is the one the plan asked @@ -436,13 +485,16 @@ func (f *Filesystem) blank( switch { case prior == "": return volstack.StateAbsent, reading, volstack.Artifact{}, nil - case prior != f.cfg.FsType: + case f.cfg.FsType != "" && prior != f.cfg.FsType: return volstack.StateAbsent, reading, volstack.Artifact{}, fmt.Errorf( "filesystem: refusing to stage %s, which is recorded as carrying %s where the plan "+ "asks for %s: reformatting would destroy the volume, and mounting it as %s would "+ "serve a filesystem the plan does not declare", deviceOf(below), prior, f.cfg.FsType, prior) default: + if f.cfg.FsType == "" { + f.detected = prior + } // Recorded as formatted while nothing was found on it: the reading is a // failed probe rather than an empty device, so the filesystem is treated as // present and unmounted. Mounting it is the honest next step, and a mount @@ -486,7 +538,7 @@ func (f *Filesystem) mountFlags() []string { // asked for. That is also the only one the layer acts on, since a device // carrying another is refused rather than reconciled. func (f *Filesystem) strategy() FilesystemLayerStrategy { - return FilesystemStrategyFor(f.cfg.FsType) + return FilesystemStrategyFor(f.known()) } // deviceOf names the device below for an error message, without asserting there diff --git a/atlas-lib/volstack/layers/filesystem_test.go b/atlas-lib/volstack/layers/filesystem_test.go index aa107980b..742791b16 100644 --- a/atlas-lib/volstack/layers/filesystem_test.go +++ b/atlas-lib/volstack/layers/filesystem_test.go @@ -592,3 +592,41 @@ func TestReleaseForcesWhenAPlainUnmountRefuses(t *testing.T) { t.Fatal("a plain unmount refused and the release did not fall back to its force path") } } + +// A plan that names no filesystem -- a static PV without fsType, such as a +// test fail-over's clone or a PV Ramen restored without the field -- mounts +// what the device carries instead of refusing it for not being ext4 +// (2026-10-03), and formats a blank device as the default. +func TestAPlanNamingNoFilesystemMountsWhatTheDeviceCarries(t *testing.T) { + fs := newFakeFS() + l := newFSAsking(t, fs, "", blockdev.Reading{Content: blockdev.ContentFilesystem, Type: "xfs"}, nil) + if _, err := l.Ensure(context.Background(), belowArtifact()); err != nil { + t.Fatal(err) + } + if len(fs.formatted) != 0 { + t.Fatalf("formatted a device that carries a filesystem: %+v", fs.formatted) + } + if len(fs.mounted) != 1 || fs.mounted[0].fsType != "xfs" { + t.Fatalf("mounted %+v, want once as xfs", fs.mounted) + } + if p, _ := l.Params().(FilesystemParams); p.FsType != "xfs" { + t.Fatalf("params %+v, want the detected xfs recorded", p) + } +} + +func TestAPlanNamingNoFilesystemFormatsABlankDeviceAsTheDefault(t *testing.T) { + fs := newFakeFS() + l := NewFilesystem(FilesystemConfig{ + FsType: "", DefaultFsType: "ext4", StagingPath: stagingPath, Ops: fs, + Content: fakeReader{reading: blockdev.Reading{Content: blockdev.ContentBlank}}, + }) + if _, err := l.Ensure(context.Background(), belowArtifact()); err != nil { + t.Fatal(err) + } + if len(fs.formatted) != 1 || fs.formatted[0].fsType != "ext4" { + t.Fatalf("formatted %+v, want once as ext4", fs.formatted) + } + if len(fs.mounted) != 1 || fs.mounted[0].fsType != "ext4" { + t.Fatalf("mounted %+v, want once as ext4", fs.mounted) + } +} diff --git a/atlas-lib/volstack/plans/node.go b/atlas-lib/volstack/plans/node.go index 5aa5c02d8..f31045d40 100644 --- a/atlas-lib/volstack/plans/node.go +++ b/atlas-lib/volstack/plans/node.go @@ -109,6 +109,7 @@ func (n *Node) fabric(connection lvol.Connection) volstack.Layer { func (n *Node) filesystem(volume Volume) volstack.Layer { return layers.NewFilesystem(layers.FilesystemConfig{ FsType: volume.FsType, + DefaultFsType: volume.DefaultFsType, StagingPath: volume.StagingPath, MountFlags: volume.MountFlags, FormatOptions: volume.FormatOptions, diff --git a/atlas-lib/volstack/plans/plans.go b/atlas-lib/volstack/plans/plans.go index c514dae4a..8395a160f 100644 --- a/atlas-lib/volstack/plans/plans.go +++ b/atlas-lib/volstack/plans/plans.go @@ -52,9 +52,14 @@ type Volume struct { // FsType is the filesystem this volume is. It decides what a blank device is // formatted as, and it is also the only filesystem that will be mounted: a - // device carrying another is refused. + // device carrying another is refused. Empty: whatever the device carries, + // and DefaultFsType for a blank one. FsType string + // DefaultFsType is what a blank device is formatted as when FsType names + // nothing. + DefaultFsType string + // MountFlags are the flags the volume asked for, ahead of the ones the // filesystem layer derives from the filesystem itself. MountFlags []string diff --git a/csi-driver/internal/csi/node/plan.go b/csi-driver/internal/csi/node/plan.go index e0dcaa747..aa86c704d 100644 --- a/csi-driver/internal/csi/node/plan.go +++ b/csi-driver/internal/csi/node/plan.go @@ -159,6 +159,10 @@ func stackVolume( vc map[string]string, volCap *csi.VolumeCapability, ) plans.Volume { + // Named by the record or the capability; empty otherwise, so the layer + // mounts what the device carries (a static PV without fsType: a test + // fail-over's clone, a PV Ramen restored without the field) and formats + // a blank device as ext4. fsType := stagedFsType(vc, volCap) return plans.Volume{ UUID: deviceLvolID(vc), @@ -167,8 +171,9 @@ func stackVolume( PVCName: vc[csicommon.CSIStorageNameKey], StagingPath: stagingPath, FsType: fsType, + DefaultFsType: "ext4", MountFlags: volumeMountFlags(volCap), - FormatOptions: mount.FormatOptions(fsType, vc, wantsVDO(vc)), + FormatOptions: mount.FormatOptions(fsTypeOrDefault(volCap), vc, wantsVDO(vc)), ReservedBlocksPercent: vc["tune2fs_reserved_blocks"], Encrypted: boolFromContext(vc[csicommon.ParamEncryption]), } diff --git a/csi-driver/internal/csi/node/stage.go b/csi-driver/internal/csi/node/stage.go index 2e52d0cd1..004fdb243 100644 --- a/csi-driver/internal/csi/node/stage.go +++ b/csi-driver/internal/csi/node/stage.go @@ -102,7 +102,7 @@ func (ns *Server) NodeStageVolume( return nil, status.Error(codes.Internal, err.Error()) } - ns.rememberStagedVolume(ctx, volumeID, vc, artifact, req.GetVolumeCapability()) + ns.rememberStagedVolume(ctx, volumeID, vc, artifact, plan, req.GetVolumeCapability()) // The CSI spec passes VolumeContext to this RPC and to nothing after it, so // what the later RPCs need is written beside the staging path. @@ -619,6 +619,7 @@ func (ns *Server) rememberStagedVolume( volumeID string, vc map[string]string, artifact volstack.Artifact, + plan volstack.Plan, volCap *csi.VolumeCapability, ) { if device, ok := artifact.Device(); ok { @@ -636,10 +637,33 @@ func (ns *Server) rememberStagedVolume( // The device carries this filesystem, because the layer either put it there // or refused to stage a device carrying another. fsType := stagedFsType(vc, volCap) + if fsType == "" { + // Nobody named it: the layer mounted what the device carries, or + // formatted a blank device as its default, and knows which. + fsType = planFsType(plan) + } + if fsType == "" { + return + } vc[stagedFsTypeKey] = fsType ns.recordOnDiskFilesystem(ctx, volumeID, vc, fsType) } +// planFsType is the filesystem the plan's filesystem layer stands for after +// it acted (layers.FilesystemParams), "" when the plan has none. +func planFsType(plan volstack.Plan) string { + for _, l := range plan { + recorded, ok := l.(interface{ Params() any }) + if !ok { + continue + } + if p, ok := recorded.Params().(layers.FilesystemParams); ok { + return p.FsType + } + } + return "" +} + // priorFormat is what the volume is recorded as carrying, for the layer that // has to decide whether a device reading blank is empty or merely unreadable. // @@ -704,13 +728,15 @@ func (ns *Server) volumeIsBeingDeleted(ctx context.Context, volumeID string) boo } // stagedFsType returns the filesystem a volume was staged with: the one -// recorded at stage time when it is there, and otherwise the one the volume -// capability asks for, which is all a volume staged by an older driver has. +// recorded at stage time when it is there, else the one the volume capability +// asks for, else "" -- no opinion, which lets the filesystem layer mount what +// the device carries instead of refusing an XFS volume for not being the ext4 +// nobody asked for (a static PV without fsType, 2026-10-03). func stagedFsType(volumeContext map[string]string, volCap *csi.VolumeCapability) string { if fsType := strings.TrimSpace(volumeContext[stagedFsTypeKey]); fsType != "" { return fsType } - return fsTypeOrDefault(volCap) + return volCap.GetMount().GetFsType() } // fsTypeOrDefault returns the requested filesystem type, defaulting to ext4. From 599a7415e357895a02b0701b6fc35bcb933c0228 Mon Sep 17 00:00:00 2001 From: michael Date: Sat, 3 Oct 2026 05:21:04 +0300 Subject: [PATCH 175/206] operator: a test fail-over's bubble PV and PVC carry the source's volumeMode A VM's disk is a Block claim. The drill's bubble PV and PVC omitted the mode, defaulted to Filesystem, and the kubelet asked the node plugin to mount the raw guest disk (mount: exit status 32); the bubble VM never started (realbed 2026-10-03). The source PV's volumeMode is recorded on the clone and set on both objects. Co-Authored-By: Claude Fable 5.1 --- .../storage.simplyblock.io_testfailovers.yaml | 7 +++++++ operator/api/v1alpha2/testfailover_types.go | 6 ++++++ .../storage.simplyblock.io_testfailovers.yaml | 7 +++++++ .../controller/testfailover_controller.go | 17 +++++++++++++++++ .../storage.simplyblock.io_testfailovers.yaml | 7 +++++++ 5 files changed, 44 insertions(+) diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_testfailovers.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_testfailovers.yaml index d68cf6ad7..1cca0647d 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_testfailovers.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_testfailovers.yaml @@ -185,6 +185,13 @@ spec: empty fsType makes the node plugin default to ext4 and refuse to mount an XFS volume. type: string + sourceVolumeMode: + description: |- + SourceVolumeMode is the source PV's volumeMode (Filesystem or Block), + carried onto the bubble PV and PVC. A VM's disk is a Block claim; a bubble + claim that omitted the mode defaulted to Filesystem and the kubelet asked + the node plugin to mount a raw guest disk (2026-10-03). + type: string sourceHandle: description: SourceHandle is the source volume's backend handle, read from its PV. diff --git a/operator/api/v1alpha2/testfailover_types.go b/operator/api/v1alpha2/testfailover_types.go index de033a936..cf2e1e852 100644 --- a/operator/api/v1alpha2/testfailover_types.go +++ b/operator/api/v1alpha2/testfailover_types.go @@ -175,6 +175,12 @@ type TestFailoverClone struct { // volume. // +optional SourceFSType string `json:"sourceFSType,omitempty"` + // SourceVolumeMode is the source PV's volumeMode (Filesystem or Block), + // carried onto the bubble PV and PVC. A VM's disk is a Block claim; a bubble + // claim that omitted the mode defaulted to Filesystem and the kubelet asked + // the node plugin to mount a raw guest disk (2026-10-03). + // +optional + SourceVolumeMode string `json:"sourceVolumeMode,omitempty"` } // TestFailoverReport is the evidence a drill produces. diff --git a/operator/config/crd/bases/storage.simplyblock.io_testfailovers.yaml b/operator/config/crd/bases/storage.simplyblock.io_testfailovers.yaml index d68cf6ad7..1cca0647d 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_testfailovers.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_testfailovers.yaml @@ -185,6 +185,13 @@ spec: empty fsType makes the node plugin default to ext4 and refuse to mount an XFS volume. type: string + sourceVolumeMode: + description: |- + SourceVolumeMode is the source PV's volumeMode (Filesystem or Block), + carried onto the bubble PV and PVC. A VM's disk is a Block claim; a bubble + claim that omitted the mode defaulted to Filesystem and the kubelet asked + the node plugin to mount a raw guest disk (2026-10-03). + type: string sourceHandle: description: SourceHandle is the source volume's backend handle, read from its PV. diff --git a/operator/internal/controller/testfailover_controller.go b/operator/internal/controller/testfailover_controller.go index 53cab8521..3f13340be 100644 --- a/operator/internal/controller/testfailover_controller.go +++ b/operator/internal/controller/testfailover_controller.go @@ -241,6 +241,7 @@ func (r *TestFailoverReconciler) resolveSource(ctx context.Context, tf *simplybl srcAttrs, _, _ := unstructured.NestedStringMap(pv, "spec", "csi", "volumeAttributes") bubbleVC := bubbleVolumeContext(srcAttrs) fsType, _, _ := unstructured.NestedString(pv, "spec", "csi", "fsType") + volumeMode, _, _ := unstructured.NestedString(pv, "spec", "volumeMode") if err := r.transitionTo(ctx, tf, simplyblockv1alpha2.TestFailoverStepResolvingPoint, func(s *simplyblockv1alpha2.TestFailoverStatus) { s.Clones = []simplyblockv1alpha2.TestFailoverClone{{ @@ -248,6 +249,7 @@ func (r *TestFailoverReconciler) resolveSource(ctx context.Context, tf *simplybl SourceHandle: handle, SourceVolumeContext: bubbleVC, SourceFSType: fsType, + SourceVolumeMode: volumeMode, }} s.Message = "resolved the source volume; resolving the recovery point" }); err != nil { @@ -322,6 +324,7 @@ func (r *TestFailoverReconciler) resolveSourceGroup(ctx context.Context, tf *sim srcAttrs, _, _ := unstructured.NestedStringMap(pv, "spec", "csi", "volumeAttributes") bubbleVC := bubbleVolumeContext(srcAttrs) fsType, _, _ := unstructured.NestedString(pv, "spec", "csi", "fsType") + volumeMode, _, _ := unstructured.NestedString(pv, "spec", "volumeMode") clones := make([]simplyblockv1alpha2.TestFailoverClone, 0, len(memberIDs)) for _, id := range memberIDs { @@ -330,6 +333,7 @@ func (r *TestFailoverReconciler) resolveSourceGroup(ctx context.Context, tf *sim SourceRef: v.PVCName, SourceHandle: srcUUID + ":" + v.PoolID + ":" + v.LvolID, SourceFSType: fsType, + SourceVolumeMode: volumeMode, SourceVolumeContext: bubbleVC, SizeBytes: v.Size, }) @@ -998,6 +1002,17 @@ func (r *TestFailoverReconciler) placeBubble(ctx context.Context, tf *simplybloc return ctrl.Result{}, nil } +// volumeModeOf is the bubble PV's and PVC's volumeMode: the source's, so a +// VM's Block disk stays a block device instead of being mounted as a +// filesystem; nil (the default, Filesystem) when the source did not say. +func volumeModeOf(clone simplyblockv1alpha2.TestFailoverClone) *corev1.PersistentVolumeMode { + if clone.SourceVolumeMode == "" { + return nil + } + mode := corev1.PersistentVolumeMode(clone.SourceVolumeMode) + return &mode +} + // bubbleVolumeContextStripKeys are the source PV volumeAttributes that must NOT // be carried onto the bubble PV: they identify the SOURCE volume and its NVMe-oF // target. The node plugin re-resolves the clone's own identity from the clone @@ -1078,6 +1093,7 @@ func (r *TestFailoverReconciler) bubbleManifestWork(tf *simplyblockv1alpha2.Test AccessModes: []corev1.PersistentVolumeAccessMode{corev1.ReadWriteOnce}, PersistentVolumeReclaimPolicy: corev1.PersistentVolumeReclaimRetain, StorageClassName: scName, + VolumeMode: volumeModeOf(clone), ClaimRef: &corev1.ObjectReference{ Kind: "PersistentVolumeClaim", APIVersion: "v1", Namespace: ns, Name: pvcName, }, @@ -1099,6 +1115,7 @@ func (r *TestFailoverReconciler) bubbleManifestWork(tf *simplyblockv1alpha2.Test Resources: corev1.VolumeResourceRequirements{Requests: corev1.ResourceList{corev1.ResourceStorage: capacity}}, StorageClassName: &scName, VolumeName: pvName, + VolumeMode: volumeModeOf(clone), }, } for _, obj := range []client.Object{pv, pvc} { diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_testfailovers.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_testfailovers.yaml index d68cf6ad7..1cca0647d 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_testfailovers.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_testfailovers.yaml @@ -185,6 +185,13 @@ spec: empty fsType makes the node plugin default to ext4 and refuse to mount an XFS volume. type: string + sourceVolumeMode: + description: |- + SourceVolumeMode is the source PV's volumeMode (Filesystem or Block), + carried onto the bubble PV and PVC. A VM's disk is a Block claim; a bubble + claim that omitted the mode defaulted to Filesystem and the kubelet asked + the node plugin to mount a raw guest disk (2026-10-03). + type: string sourceHandle: description: SourceHandle is the source volume's backend handle, read from its PV. From 7d5f38bc45f393bb596f90c25463c6ac54403a40 Mon Sep 17 00:00:00 2001 From: michael Date: Sat, 3 Oct 2026 12:27:06 +0300 Subject: [PATCH 176/206] chart: the Control Center on the integrate_csi_addons line The console templates and their values block of feature/control-center-ui (PR #512), copied as they are: self-contained (sbcc.* helpers only), so the chart deploys the UI with the control plane on this line too. One addition: controlCenter.service.nodePort pins the node port of a NodePort Service. Co-Authored-By: Claude Fable 5.1 --- .../templates/_control_center_helpers.tpl | 61 ++++++ .../templates/control-center-mock.yaml | 96 +++++++++ .../control-center-networkpolicy.yaml | 76 +++++++ .../templates/control-center-rbac.yaml | 176 ++++++++++++++++ .../templates/control-center.yaml | 196 ++++++++++++++++++ .../charts/simplyblock-operator/values.yaml | 152 ++++++++++++++ 6 files changed, 757 insertions(+) create mode 100644 helm-charts/charts/simplyblock-operator/templates/_control_center_helpers.tpl create mode 100644 helm-charts/charts/simplyblock-operator/templates/control-center-mock.yaml create mode 100644 helm-charts/charts/simplyblock-operator/templates/control-center-networkpolicy.yaml create mode 100644 helm-charts/charts/simplyblock-operator/templates/control-center-rbac.yaml create mode 100644 helm-charts/charts/simplyblock-operator/templates/control-center.yaml diff --git a/helm-charts/charts/simplyblock-operator/templates/_control_center_helpers.tpl b/helm-charts/charts/simplyblock-operator/templates/_control_center_helpers.tpl new file mode 100644 index 000000000..97a31717e --- /dev/null +++ b/helm-charts/charts/simplyblock-operator/templates/_control_center_helpers.tpl @@ -0,0 +1,61 @@ +{{/* + Control Center helpers — self-contained on purpose. + + The console ships as a fragment (see control-center/ in the monorepo) and + deliberately does NOT call this chart's other helpers: everything is + namespaced under `sbcc.` so it cannot collide, and the fragment renders + unchanged if it is ever lifted out of this chart again. +*/}} + +{{- define "sbcc.name" -}} +{{- default "control-center" .Values.controlCenter.nameOverride | trunc 63 | trimSuffix "-" -}} +{{- end -}} + +{{- define "sbcc.fullname" -}} +{{- if .Values.controlCenter.fullnameOverride -}} +{{- .Values.controlCenter.fullnameOverride | trunc 63 | trimSuffix "-" -}} +{{- else -}} +{{- printf "%s-%s" .Release.Name (include "sbcc.name" .) | trunc 63 | trimSuffix "-" -}} +{{- end -}} +{{- end -}} + +{{- define "sbcc.selectorLabels" -}} +app.kubernetes.io/name: {{ include "sbcc.name" . }} +app.kubernetes.io/instance: {{ .Release.Name }} +{{- end -}} + +{{- define "sbcc.labels" -}} +{{ include "sbcc.selectorLabels" . }} +app.kubernetes.io/component: ui +app.kubernetes.io/part-of: simplyblock +app.kubernetes.io/managed-by: {{ .Release.Service }} +{{- with .Chart }} +helm.sh/chart: {{ printf "%s-%s" .Name .Version | replace "+" "_" | trunc 63 | trimSuffix "-" }} +{{- end }} +{{- end -}} + +{{/* Image reference. The registry falls back to the public one and the tag to + the chart's appVersion — pin controlCenter.image.tag for production, this + chart's appVersion is a floating tag. */}} +{{- define "sbcc.image" -}} +{{- $i := .Values.controlCenter.image -}} +{{- $reg := $i.registry | default "quay.io" -}} +{{- $tag := $i.tag | default (.Chart.AppVersion | default "latest") -}} +{{- printf "%s/%s:%s" $reg $i.repository $tag -}} +{{- end -}} + +{{/* Default upstream URLs. This chart names its objects statically + (simplyblock-operator, simplyblock-prometheus, …) rather than deriving + them from the release, so the fallbacks here are static too. */}} +{{- define "sbcc.operatorUrl" -}} +{{- .Values.controlCenter.operatorUrl | default "http://simplyblock-operator:8080" -}} +{{- end -}} + +{{- define "sbcc.prometheusUrl" -}} +{{- if .Values.controlCenter.prometheusUrl -}} +{{- .Values.controlCenter.prometheusUrl -}} +{{- else -}} +{{- $p := ((.Values.prometheus).simplyblock) | default dict -}} +{{- printf "http://%s:%v" ($p.prometheusURL | default "simplyblock-prometheus") ($p.prometheusPORT | default 9090) -}} +{{- end -}} +{{- end -}} diff --git a/helm-charts/charts/simplyblock-operator/templates/control-center-mock.yaml b/helm-charts/charts/simplyblock-operator/templates/control-center-mock.yaml new file mode 100644 index 000000000..8cc406a13 --- /dev/null +++ b/helm-charts/charts/simplyblock-operator/templates/control-center-mock.yaml @@ -0,0 +1,96 @@ +{{/* + sb-mock — the Control Center's test backend (control-center/mock). + + Deployed only for testing and demos: it impersonates the Kubernetes API, + the operator API and Prometheus with generated data, and the console's + Deployment points its proxy here when controlCenter.mock.enabled is true. + Writes persist into the mock's memory and perform no real change; nothing + in this pod touches the actual cluster, and it needs no RBAC at all. +*/}} +{{- if and .Values.controlCenter.enabled .Values.controlCenter.mock.enabled }} +{{- $cc := .Values.controlCenter }} +{{- $mock := $cc.mock }} +{{- $name := printf "%s-mock" (include "sbcc.fullname" .) }} +{{- $reg := $mock.image.registry | default "quay.io" }} +{{- $tag := $mock.image.tag | default (.Chart.AppVersion | default "latest") }} +apiVersion: apps/v1 +kind: Deployment +metadata: + name: {{ $name }} + namespace: {{ .Release.Namespace }} + labels: + {{- include "sbcc.labels" . | nindent 4 }} + app.kubernetes.io/component: ui-mock +spec: + # one replica on purpose: the world lives in this pod's memory + replicas: 1 + selector: + matchLabels: + app.kubernetes.io/name: {{ include "sbcc.name" . }}-mock + app.kubernetes.io/instance: {{ .Release.Name }} + template: + metadata: + labels: + app.kubernetes.io/name: {{ include "sbcc.name" . }}-mock + app.kubernetes.io/instance: {{ .Release.Name }} + spec: + automountServiceAccountToken: false + securityContext: + runAsNonRoot: true + runAsUser: 65532 + runAsGroup: 65532 + seccompProfile: + type: RuntimeDefault + {{- with $cc.nodeSelector }} + nodeSelector: {{- toYaml . | nindent 8 }} + {{- end }} + {{- with $cc.tolerations }} + tolerations: {{- toYaml . | nindent 8 }} + {{- end }} + containers: + - name: mock + image: "{{ $reg }}/{{ $mock.image.repository }}:{{ $tag }}" + imagePullPolicy: {{ $mock.image.pullPolicy }} + args: + - --listen=:8080 + - --dataset={{ $mock.dataset }} + {{- if $mock.seed }} + - --seed={{ $mock.seed }} + {{- end }} + - --sim-interval={{ $mock.simInterval }} + - --fail-rate={{ $mock.failRate }} + - --namespace={{ .Release.Namespace }} + ports: + - name: http + containerPort: 8080 + resources: {{- toYaml $mock.resources | nindent 12 }} + securityContext: + allowPrivilegeEscalation: false + readOnlyRootFilesystem: true + capabilities: + drop: ["ALL"] + readinessProbe: + httpGet: {path: /healthz, port: http} + periodSeconds: 5 + livenessProbe: + httpGet: {path: /healthz, port: http} + periodSeconds: 20 +--- +apiVersion: v1 +kind: Service +metadata: + name: {{ $name }} + namespace: {{ .Release.Namespace }} + labels: + {{- include "sbcc.labels" . | nindent 4 }} + app.kubernetes.io/component: ui-mock +spec: + type: ClusterIP + ports: + - name: http + port: 8080 + targetPort: http + selector: + app.kubernetes.io/name: {{ include "sbcc.name" . }}-mock + app.kubernetes.io/instance: {{ .Release.Name }} +{{- end }} diff --git a/helm-charts/charts/simplyblock-operator/templates/control-center-networkpolicy.yaml b/helm-charts/charts/simplyblock-operator/templates/control-center-networkpolicy.yaml new file mode 100644 index 000000000..969fabc3e --- /dev/null +++ b/helm-charts/charts/simplyblock-operator/templates/control-center-networkpolicy.yaml @@ -0,0 +1,76 @@ +{{/* + NetworkPolicy for the Control Center. + + The console has exactly three upstreams — the Kubernetes API, the operator + API and Prometheus — and no route to the internet: React, Babel, the + typefaces and the brand mark are vendored into the image. This caps egress + to those upstreams plus DNS. + + Off by default because the API-server rule is cluster-specific: it ships as + the RFC1918 ranges, and you should narrow controlCenter.networkPolicy + .apiServerCidrs to your API server endpoint (`kubectl get endpoints + kubernetes -n default`) or your service CIDR. A rule with ports but no `to:` + would allow 443 to the whole internet, which is why one is never rendered. +*/}} +{{- if and .Values.controlCenter.enabled .Values.controlCenter.networkPolicy.enabled }} +{{- $cc := .Values.controlCenter }} +{{- $name := include "sbcc.fullname" . }} +apiVersion: networking.k8s.io/v1 +kind: NetworkPolicy +metadata: + name: {{ $name }} + namespace: {{ .Release.Namespace }} + labels: + {{- include "sbcc.labels" . | nindent 4 }} +spec: + podSelector: + matchLabels: + {{- include "sbcc.selectorLabels" . | nindent 6 }} + policyTypes: ["Ingress", "Egress"] + ingress: + # Who may reach the console. Defaults to the release namespace (which + # covers kubectl port-forward); add your ingress controller's namespace + # when the Ingress is enabled. + - from: + {{- range $ns := $cc.networkPolicy.ingressFromNamespaces }} + - namespaceSelector: + matchLabels: + kubernetes.io/metadata.name: {{ $ns }} + {{- end }} + - podSelector: {} + ports: + - protocol: TCP + port: 8080 + egress: + # The operator API and Prometheus, inside the release namespace. + - to: + - namespaceSelector: + matchLabels: + kubernetes.io/metadata.name: {{ .Release.Namespace }} + ports: + - protocol: TCP + port: 8080 + - protocol: TCP + port: 9090 + # The Kubernetes API server. Narrow these CIDRs — see the header comment. + - to: + {{- range $cidr := $cc.networkPolicy.apiServerCidrs }} + - ipBlock: + cidr: {{ $cidr }} + {{- end }} + ports: + - protocol: TCP + port: 443 + - protocol: TCP + port: 6443 + # DNS + - to: + - namespaceSelector: + matchLabels: + kubernetes.io/metadata.name: kube-system + ports: + - protocol: UDP + port: 53 + - protocol: TCP + port: 53 +{{- end }} diff --git a/helm-charts/charts/simplyblock-operator/templates/control-center-rbac.yaml b/helm-charts/charts/simplyblock-operator/templates/control-center-rbac.yaml new file mode 100644 index 000000000..c1bd77089 --- /dev/null +++ b/helm-charts/charts/simplyblock-operator/templates/control-center-rbac.yaml @@ -0,0 +1,176 @@ +{{/* + RBAC for the Control Center. + + In serviceaccount mode this role IS the console's authority, for anyone who + can reach the Service. In passthrough mode the pod needs nothing at all — + Kubernetes enforces each user's own RBAC — so set controlCenter.rbac.create + to false there. + + Kept in step with control-center/deploy/k8s/rbac.yaml; that file is the + plain-manifest equivalent of this one. +*/}} +{{- if and .Values.controlCenter.enabled .Values.controlCenter.rbac.create }} +{{- $name := include "sbcc.fullname" . }} +apiVersion: rbac.authorization.k8s.io/v1 +kind: ClusterRole +metadata: + name: {{ $name }} + labels: + {{- include "sbcc.labels" . | nindent 4 }} +rules: + # ---- simplyblock entities: read ---- + - apiGroups: ["storage.simplyblock.io"] + resources: + - storageclusters + - storagenodes + - storagenodesets + - storagedevices + - storagepools + - storagebackups + - backuppolicies + - backuprestores + - backupimports + - controlplanes + - replicationpairs + - replicationpolicies + - replicationslots + - volumemigrations + - tasks + - simplyblockdrivers + - clusterdeploymentconfigs + verbs: ["get", "list", "watch"] + - apiGroups: ["storage.simplyblock.io"] + resources: + - storageclusters/status + - storagenodes/status + - storagepools/status + - replicationpolicies/status + - replicationslots/status + verbs: ["get"] + + # ---- day-2 operations ---- + # An action is not a verb against an entity: it is an Ops object the operator + # reconciles. The console creates one, reads its phase, patches spec.abort to + # unwind it, and deletes it only to abort. + - apiGroups: ["storage.simplyblock.io"] + resources: + - storageclusterops + - storagenodeops + - storagedeviceops + - storagepoolops + - storagebackupops + - controlplaneops + - persistentvolumeops + - operatorops + - replicationops + verbs: ["get", "list", "watch", "create", "patch", "delete"] + + # ---- resources the console authors ---- + - apiGroups: ["storage.simplyblock.io"] + resources: ["replicationpairs", "replicationpolicies"] + verbs: ["create", "update", "patch", "delete"] + - apiGroups: ["storage.simplyblock.io"] + resources: ["storagebackups", "backuppolicies", "backuprestores", "backupimports", "volumemigrations"] + verbs: ["create", "update", "patch", "delete"] + - apiGroups: ["storage.simplyblock.io"] + resources: ["storageclusters", "storagepools", "storagenodesets"] + verbs: ["update", "patch"] + + # ---- core Kubernetes ---- + - apiGroups: [""] + resources: ["nodes", "persistentvolumes", "persistentvolumeclaims", "pods", "events", "namespaces"] + verbs: ["get", "list", "watch"] + - apiGroups: [""] + resources: ["pods/log"] + verbs: ["get"] + # A PVC annotation is the whole membership model for replication and for + # backup policies, so patching one is how a volume is attached. + - apiGroups: [""] + resources: ["persistentvolumeclaims"] + verbs: ["patch", "update"] + - apiGroups: ["storage.k8s.io"] + resources: ["storageclasses"] + verbs: ["get", "list", "watch", "patch"] + - apiGroups: ["snapshot.storage.k8s.io"] + resources: ["volumesnapshots", "volumesnapshotcontents", "volumesnapshotclasses"] + verbs: ["get", "list", "watch"] + + # ---- workloads, read only, for the Ramen recipe editor ---- + - apiGroups: ["apps"] + resources: ["deployments", "statefulsets", "daemonsets", "replicasets"] + verbs: ["get", "list"] + - apiGroups: [""] + resources: ["services", "configmaps"] + verbs: ["get", "list"] + - apiGroups: ["networking.k8s.io"] + resources: ["ingresses"] + verbs: ["get", "list"] + - apiGroups: ["kubevirt.io"] + resources: ["virtualmachines", "virtualmachineinstances"] + verbs: ["get", "list"] + + # ---- Ramen: instances only, never the CRDs ---- + - apiGroups: ["ramendr.openshift.io"] + resources: ["drpolicies", "drclusters", "drclusterconfigs", "volumereplicationgroups"] + verbs: ["get", "list", "watch"] + - apiGroups: ["ramendr.openshift.io"] + resources: ["drplacementcontrols"] + verbs: ["get", "list", "watch", "create", "update", "patch", "delete"] + - apiGroups: ["ramendr.openshift.io"] + resources: ["recipes"] + verbs: ["get", "list", "watch", "create", "update", "patch", "delete"] + - apiGroups: ["cluster.open-cluster-management.io"] + resources: ["managedclusters"] + verbs: ["get", "list"] + + # ---- DR hub (dr-simplyblock): every DR screen is a CR ---- + # In serviceaccount mode this is the console's authority for DR: reads on + # everything, the dr-operator writes, the dr-admin writes on plans/paths/ + # restores, and the override verb the hub's webhook checks before admitting a + # readiness override. Absent CRDs cost nothing — the DR section then reports + # that the hub is not installed. Trim to your policy, or use passthrough. + - apiGroups: ["dr.simplyblock.io"] + resources: ["protectionplans", "drpaths", "protectedapplications", "recoveryplans", "recoveryactions", "testbubbles", "testschedules", "restoreactions", "drconfigs"] + verbs: ["get", "list", "watch"] + - apiGroups: ["dr.simplyblock.io"] + resources: ["protectionplans", "drpaths", "protectedapplications", "recoveryplans", "recoveryactions", "testbubbles", "testschedules", "restoreactions"] + verbs: ["create", "update", "patch", "delete"] + - apiGroups: ["dr.simplyblock.io"] + resources: ["recoveryactions"] + verbs: ["override"] + - apiGroups: ["sitemap.simplyblock.io"] + resources: ["siteprofiles", "dhcpservers"] + verbs: ["get", "list", "watch"] + - apiGroups: ["sitemap.simplyblock.io"] + resources: ["dhcpservers"] + verbs: ["create", "update", "patch", "delete"] + # the console asks the API server what its identity may do, to disable controls + - apiGroups: ["authorization.k8s.io"] + resources: ["selfsubjectaccessreviews", "selfsubjectrulesreviews"] + verbs: ["create"] + - apiGroups: ["authentication.k8s.io"] + resources: ["selfsubjectreviews"] + verbs: ["create"] + + # ---- Helm release view ---- + # Helm stores releases as Secrets. This is the only reason the console reads a + # Secret, and a backup credentialsSecretRef is only ever shown by name. + - apiGroups: [""] + resources: ["secrets"] + verbs: ["get", "list"] +--- +apiVersion: rbac.authorization.k8s.io/v1 +kind: ClusterRoleBinding +metadata: + name: {{ $name }} + labels: + {{- include "sbcc.labels" . | nindent 4 }} +roleRef: + apiGroup: rbac.authorization.k8s.io + kind: ClusterRole + name: {{ $name }} +subjects: + - kind: ServiceAccount + name: {{ $name }} + namespace: {{ .Release.Namespace }} +{{- end }} diff --git a/helm-charts/charts/simplyblock-operator/templates/control-center.yaml b/helm-charts/charts/simplyblock-operator/templates/control-center.yaml new file mode 100644 index 000000000..3a81f5e3b --- /dev/null +++ b/helm-charts/charts/simplyblock-operator/templates/control-center.yaml @@ -0,0 +1,196 @@ +{{/* + Control Center — the simplyblock web UI. + + Installs with the operator when controlCenter.enabled is true, so the console + arrives as part of the deployment rather than as a second thing to install. + Self-contained on purpose: it calls only its own `sbcc.*` helpers (in + _control_center_helpers.tpl), never this chart's. Sources and image build + live in control-center/ at the repository root. +*/}} +{{- if .Values.controlCenter.enabled }} +{{- $cc := .Values.controlCenter }} +{{- $name := include "sbcc.fullname" . }} +apiVersion: v1 +kind: ServiceAccount +metadata: + name: {{ $name }} + namespace: {{ .Release.Namespace }} + labels: + {{- include "sbcc.labels" . | nindent 4 }} +{{- with $cc.imagePullSecrets }} +imagePullSecrets: + {{- toYaml . | nindent 2 }} +{{- end }} +--- +apiVersion: apps/v1 +kind: Deployment +metadata: + name: {{ $name }} + namespace: {{ .Release.Namespace }} + labels: + {{- include "sbcc.labels" . | nindent 4 }} +spec: + replicas: {{ $cc.replicas }} + revisionHistoryLimit: 3 + selector: + matchLabels: + {{- include "sbcc.selectorLabels" . | nindent 6 }} + strategy: + type: RollingUpdate + rollingUpdate: + maxUnavailable: 0 + maxSurge: 1 + template: + metadata: + labels: + {{- include "sbcc.selectorLabels" . | nindent 8 }} + {{- with $cc.podAnnotations }} + annotations: + {{- toYaml . | nindent 8 }} + {{- end }} + spec: + serviceAccountName: {{ $name }} + automountServiceAccountToken: true + securityContext: + runAsNonRoot: true + runAsUser: 101 + runAsGroup: 101 + fsGroup: 101 + seccompProfile: + type: RuntimeDefault + {{- with $cc.nodeSelector }} + nodeSelector: {{- toYaml . | nindent 8 }} + {{- end }} + {{- with $cc.tolerations }} + tolerations: {{- toYaml . | nindent 8 }} + {{- end }} + topologySpreadConstraints: + - maxSkew: 1 + topologyKey: kubernetes.io/hostname + whenUnsatisfiable: ScheduleAnyway + labelSelector: + matchLabels: + {{- include "sbcc.selectorLabels" . | nindent 14 }} + containers: + - name: ui + image: {{ include "sbcc.image" . | quote }} + imagePullPolicy: {{ $cc.image.pullPolicy }} + ports: + - name: http + containerPort: 8080 + env: + - name: SB_NAMESPACE + valueFrom: + fieldRef: + fieldPath: metadata.namespace + - name: SB_MODE + value: {{ $cc.mode | default "full" | quote }} + - name: SB_DR_NAMESPACE + value: {{ $cc.drNamespace | default "ramen-ops" | quote }} + - name: SB_AUTH_MODE + value: {{ $cc.authMode | quote }} + {{- /* With the mock enabled every upstream is the sb-mock service: + the console runs unmodified against generated data. */}} + {{- $mockURL := printf "http://%s-mock:8080" $name }} + - name: SB_K8S_API + value: {{ ternary $mockURL $cc.kubernetesApi $cc.mock.enabled | quote }} + # SNI and Host for the API server. Must match a name on the + # apiserver's certificate — see controlCenter.kubernetesApiHost. + # (Ignored for a plain-http upstream, i.e. in mock mode.) + - name: SB_K8S_HOST + value: {{ $cc.kubernetesApiHost | quote }} + - name: SB_K8S_CA_FILE + value: {{ $cc.kubernetesApiCaFile | quote }} + - name: SB_OPERATOR_URL + value: {{ ternary $mockURL (include "sbcc.operatorUrl" .) $cc.mock.enabled | quote }} + - name: SB_HELM_URL + value: {{ ternary $mockURL ($cc.helmUrl | default (include "sbcc.operatorUrl" .)) $cc.mock.enabled | quote }} + - name: SB_PROMETHEUS_URL + value: {{ ternary $mockURL (include "sbcc.prometheusUrl" .) $cc.mock.enabled | quote }} + - name: SB_MOCK + value: "false" + - name: SB_LISTEN_PORT + value: "8080" + - name: SB_TOKEN_REFRESH_SECONDS + value: {{ $cc.tokenRefreshSeconds | quote }} + resources: {{- toYaml $cc.resources | nindent 12 }} + securityContext: + allowPrivilegeEscalation: false + readOnlyRootFilesystem: true + capabilities: + drop: ["ALL"] + startupProbe: + httpGet: {path: /healthz, port: http} + periodSeconds: 2 + failureThreshold: 15 + readinessProbe: + httpGet: {path: /healthz, port: http} + periodSeconds: 10 + livenessProbe: + httpGet: {path: /healthz, port: http} + periodSeconds: 20 + volumeMounts: + # readOnlyRootFilesystem: the scratch mount holds the generated + # config.js, the proxied token and nginx's temp files; conf.d is + # writable because the image renders its server block at startup. + - name: nginx-tmp + mountPath: /tmp/nginx + - name: nginx-confd + mountPath: /etc/nginx/conf.d + volumes: + - name: nginx-tmp + emptyDir: {medium: Memory, sizeLimit: 16Mi} + - name: nginx-confd + emptyDir: {medium: Memory, sizeLimit: 1Mi} +--- +apiVersion: v1 +kind: Service +metadata: + name: {{ $name }} + namespace: {{ .Release.Namespace }} + labels: + {{- include "sbcc.labels" . | nindent 4 }} +spec: + type: {{ $cc.service.type }} + ports: + - name: http + port: {{ $cc.service.port }} + targetPort: http + {{- with $cc.service.nodePort }} + nodePort: {{ . }} + {{- end }} + selector: + {{- include "sbcc.selectorLabels" . | nindent 4 }} +{{- if $cc.ingress.enabled }} +--- +apiVersion: networking.k8s.io/v1 +kind: Ingress +metadata: + name: {{ $name }} + namespace: {{ .Release.Namespace }} + labels: + {{- include "sbcc.labels" . | nindent 4 }} + {{- with $cc.ingress.annotations }} + annotations: + {{- toYaml . | nindent 4 }} + {{- end }} +spec: + {{- with $cc.ingress.className }} + ingressClassName: {{ . }} + {{- end }} + {{- with $cc.ingress.tls }} + tls: {{- toYaml . | nindent 4 }} + {{- end }} + rules: + - host: {{ required "controlCenter.ingress.host is required when the ingress is enabled" $cc.ingress.host }} + http: + paths: + - path: / + pathType: Prefix + backend: + service: + name: {{ $name }} + port: + name: http +{{- end }} +{{- end }} diff --git a/helm-charts/charts/simplyblock-operator/values.yaml b/helm-charts/charts/simplyblock-operator/values.yaml index 420bb7951..276d44491 100644 --- a/helm-charts/charts/simplyblock-operator/values.yaml +++ b/helm-charts/charts/simplyblock-operator/values.yaml @@ -974,3 +974,155 @@ tls: # and FoundationDB's peers, so there is nothing left for a deployment to # provision before it can be required. mutual_enabled: true + +# The Control Center — the simplyblock web console (control-center/ in the +# monorepo). Off by default: it proxies the Kubernetes API with the pod's +# ServiceAccount token, so enabling it is a deliberate decision about who may +# reach the Service. See control-center/README.md, "Who can do what". +controlCenter: + enabled: false + + # Naming. The console templates are self-contained and do not borrow the + # chart's helpers, so these control its object names. Default object name: + # -control-center. + nameOverride: "" + fullnameOverride: "" + + image: + # defaults to quay.io when unset + registry: "" + repository: simplyblock-io/control-center + # inherits .Chart.AppVersion when unset. This chart's appVersion is a + # floating tag, so pin a released console tag here for production. + tag: "" + pullPolicy: IfNotPresent + imagePullSecrets: [] + + # Stateless — location lives in the browser — so two replicas cost little and + # keep the console up through a node loss, which is when it matters most. + replicas: 2 + + # full | dr + # + # full: the storage console — clusters, Kubernetes, control plane — with a + # Disaster recovery section that reads the DR hub's dr.simplyblock.io CRDs + # when dr-simplyblock is installed on the same cluster. + # dr: the DR-only console. Nothing but the DR section; no storage CRDs, no + # operator API, no Prometheus. This is what the dr-simplyblock-hub chart + # deploys (console.enabled) on a hub without a simplyblock control plane; + # set it here only to run the same stripped-down console from this chart. + mode: full + # Ramen's ops namespace on the DR hub: where discovered ProtectedApplications + # live and where the console asks the API server what it may do. + drNamespace: ramen-ops + + # serviceaccount | passthrough + # + # serviceaccount: this pod attaches its own token to proxied requests. The + # browser holds no credential. The console's authority is the ClusterRole, + # shared by everyone who can reach it, so put authentication in front. + # + # passthrough: the browser supplies the bearer token and Kubernetes enforces + # that user's own RBAC. Per-user authority and a real audit trail, at the + # cost of needing an auth proxy that injects the token. If you use this, + # set rbac.create=false — the pod needs no permissions of its own. + authMode: serviceaccount + tokenRefreshSeconds: 600 + + kubernetesApi: https://kubernetes.default.svc + # SNI and Host presented to the API server. Must be a name on its + # certificate — change this only together with kubernetesApi. + kubernetesApiHost: kubernetes.default.svc + kubernetesApiCaFile: /var/run/secrets/kubernetes.io/serviceaccount/ca.crt + # Empty defaults resolve to this chart's own services: the operator API on + # http://simplyblock-operator:8080 and Prometheus per + # prometheus.simplyblock.prometheusURL/prometheusPORT. + operatorUrl: "" + helmUrl: "" + prometheusUrl: "" + + rbac: + # Required in serviceaccount mode: the proxied token needs these rules or + # every request returns 403 and the console renders empty. Set false only + # with authMode: passthrough, where each user's own RBAC applies instead. + create: true + + service: + type: ClusterIP + port: 80 + # With type NodePort: the node port to pin, else Kubernetes picks one. + nodePort: null + + ingress: + enabled: false + className: nginx + host: "" + annotations: + # Authentication is not optional in serviceaccount mode. Replace this with + # your OIDC/OAuth2 proxy annotations, or keep basic auth as a stop-gap. + nginx.ingress.kubernetes.io/auth-type: basic + nginx.ingress.kubernetes.io/auth-secret: simplyblock-control-center-auth + nginx.ingress.kubernetes.io/auth-realm: simplyblock Control Center + nginx.ingress.kubernetes.io/proxy-body-size: 2m + nginx.ingress.kubernetes.io/proxy-read-timeout: "3600" + tls: [] + + networkPolicy: + # Caps the console's egress to its three upstreams plus DNS. Off by default + # because the API-server CIDRs below are cluster-specific: narrow them to + # your API server endpoint (`kubectl get endpoints kubernetes -n default`) + # or your service CIDR before enabling. + enabled: false + apiServerCidrs: + - 10.0.0.0/8 + - 172.16.0.0/12 + - 192.168.0.0/16 + # Namespaces allowed to reach the console, in addition to the release + # namespace itself — add your ingress controller's namespace when the + # Ingress is enabled, e.g. [ingress-nginx]. + ingressFromNamespaces: [] + + # Deploys sb-mock (control-center/mock) next to the console and points the + # console's proxy at it instead of the real cluster: the Kubernetes API, + # operator API and Prometheus are all impersonated, reads come from a + # generated dataset, and writes persist without performing any change. For + # testing the UI and demo installs only — never enable in production. + mock: + enabled: false + image: + # defaults to quay.io when unset + registry: "" + repository: simplyblock-io/control-center-mock + # inherits .Chart.AppVersion when unset + tag: "" + pullPolicy: IfNotPresent + # small-healthy | medium-degraded | large-scale | dr-failover | chaos, + # or auto for a seeded random pick + dataset: auto + # 0 picks a random seed; a fixed value makes the world reproducible + seed: 0 + # simulator tick (Ops phase progression); "0" disables the simulator so + # e2e suites can drive it deterministically via POST /mockctl/advance + simInterval: 4s + # probability a simulated operation ends Failed + failRate: 0.1 + resources: + requests: + cpu: 10m + memory: 32Mi + limits: + cpu: 200m + memory: 256Mi + + resources: + requests: + cpu: 20m + memory: 48Mi + limits: + cpu: 500m + memory: 192Mi + + nodeSelector: {} + tolerations: [] + podAnnotations: {} + From e356b56c25b4493b0fa224bc648d7b3242eac981 Mon Sep 17 00:00:00 2001 From: michael Date: Sat, 3 Oct 2026 14:29:00 +0300 Subject: [PATCH 177/206] csi-driver: a replication-chain walk is bounded by a cycle guard, not by 8 hops The PV keeps the original handle while every relocate and fail-over appends a clone, so the chain behind a volume grows by one per move and is never compacted. Both walks (the node plugin's redirect to the active volume and the controller's resolveChain) stopped after 8 hops; the ninth move of a volume -- an unplanned fail-over on 2026-10-03 -- left its clone one hop out of reach, the node fell back to the stashed context and attached the original on the partitioned site, and the VM never started. The walks now track the members visited and stop on a loop, with a bound of 256 as a guard far above any real chain. Tests: a 12-move chain alternating clusters, and a looping pair. Co-Authored-By: Claude Fable 5.1 --- .../internal/csi/controller/replication.go | 18 +++- csi-driver/internal/csi/node/stats.go | 21 ++++- csi-driver/internal/csi/node/stats_test.go | 82 +++++++++++++++++++ 3 files changed, 118 insertions(+), 3 deletions(-) diff --git a/csi-driver/internal/csi/controller/replication.go b/csi-driver/internal/csi/controller/replication.go index e7b483ecf..85f5b21c4 100644 --- a/csi-driver/internal/csi/controller/replication.go +++ b/csi-driver/internal/csi/controller/replication.go @@ -9,6 +9,7 @@ package controller import ( "context" "errors" + "fmt" "sort" "github.com/csi-addons/spec/lib/go/replication" @@ -139,7 +140,15 @@ func resolveReplica( func resolveChain(ctx context.Context, h *lvol.Handle, client *atlascp.Client) ([]chainHop, bool, error) { known := false hops := []chainHop{{h: h, client: client}} - for range 8 { // one hop per past fail-over; capped far above any real chain + // One hop per past fail-over, never compacted (the PV keeps the original + // handle): a cap of 8 ended the walk one hop short of a ninth move's + // clone (2026-10-03). The bound is a cycle guard, not a length estimate. + visited := map[lvol.VolumeHandle]bool{} + for range maxChainHops { + if visited[h.Handle()] { + return nil, false, fmt.Errorf("replication chain of %s loops at %s", hops[0].h.Handle(), h.Handle()) + } + visited[h.Handle()] = true rel, err := client.GetVolumeReplicationRelationship(ctx, h.Handle()) if err != nil { if errors.Is(err, errs.ErrNotFound) { @@ -162,9 +171,16 @@ func resolveChain(ctx context.Context, h *lvol.Handle, client *atlascp.Client) ( break } } + if len(hops) > maxChainHops { + return nil, false, fmt.Errorf("replication chain of %s did not converge within %d hops", hops[0].h.Handle(), maxChainHops) + } return hops, known, nil } +// maxChainHops bounds a replication-chain walk: a guard against a looping +// record, far above any chain a volume accumulates in its lifetime. +const maxChainHops = 256 + // activeEndFallback is where a Resync or a status read goes when the local // member of the chain is reaped: the chain's active end -- the live primary // on the other site -- and, as the cluster to fail back to, the local one. diff --git a/csi-driver/internal/csi/node/stats.go b/csi-driver/internal/csi/node/stats.go index 6185e9257..0e24a3085 100644 --- a/csi-driver/internal/csi/node/stats.go +++ b/csi-driver/internal/csi/node/stats.go @@ -129,7 +129,19 @@ func redirectToActiveVolume( vc map[string]string, ) map[string]string { client, lvolID := srcClient, srcLvolID - for range 8 { // one hop per past fail-over; capped far above any real chain + // One hop per past fail-over, and the chain never shrinks: the PV keeps + // the original handle while every relocate and fail-over appends a clone, + // so a volume moved nine times is nine hops out. A cap of 8 stranded a + // fail-over's clone behind the ninth hop and the node attached the + // partitioned original instead (2026-10-03, WordPress's fifth move of + // the day). The bound is a cycle guard now, not a length estimate. + visited := map[string]bool{} + for range maxChainHops { + if visited[lvolID] { + klog.Warningf("replication chain for deleted volume %s loops at %s", volumeID, lvolID) + return nil + } + visited[lvolID] = true rel, err := client.GetRelationship(ctx, lvolID) if err != nil || rel == nil { klog.Warningf("replication relationship lookup failed for deleted volume %s (at hop %s): %v", @@ -169,6 +181,11 @@ func redirectToActiveVolume( connInfo["poolID"] = rel.TargetPoolID return connInfo } - klog.Warningf("replication chain for deleted volume %s did not converge within 8 hops", volumeID) + klog.Warningf("replication chain for deleted volume %s did not converge within %d hops", volumeID, maxChainHops) return nil } + +// maxChainHops bounds a replication-chain walk. A chain grows by one member +// per move and is never compacted, so this is a guard against a looping +// record, far above any chain a volume accumulates in its lifetime. +const maxChainHops = 256 diff --git a/csi-driver/internal/csi/node/stats_test.go b/csi-driver/internal/csi/node/stats_test.go index 0a78fa2b7..bb01510f7 100644 --- a/csi-driver/internal/csi/node/stats_test.go +++ b/csi-driver/internal/csi/node/stats_test.go @@ -7,6 +7,7 @@ package node import ( "context" "errors" + "fmt" "testing" "github.com/simplyblock/csi-driver/internal/controlplane" @@ -159,3 +160,84 @@ func TestRedirectToActiveVolumeSinglePairingIsUnchanged(t *testing.T) { t.Errorf("cluster_id = %q, want %q", got, clusterB) } } + +// A volume moved many times: the PV keeps the original handle while every +// relocate and fail-over appends a clone, alternating between the two +// clusters, so the live copy sits one hop further out after each move. The +// walk must reach it however long the chain has grown; a cap of 8 stranded +// the ninth move's clone and the node attached the original on the +// partitioned site instead (2026-10-03). +func TestRedirectToActiveVolumeFollowsALongChain(t *testing.T) { + const ( + clusterA = "aaaaaaaa-0000-0000-0000-000000000001" + clusterB = "bbbbbbbb-0000-0000-0000-000000000001" + poolA = "aaaaaaaa-0000-0000-0000-00000000000a" + poolB = "bbbbbbbb-0000-0000-0000-00000000000b" + moves = 12 + ) + member := func(i int) string { return fmt.Sprintf("%08d-0000-0000-0000-000000000000", i) } + cluster := func(i int) (string, string) { + if i%2 == 0 { + return clusterA, poolA + } + return clusterB, poolB + } + active := member(moves) + clients := map[string]*fakeRelationshipAPI{ + clusterA + "/" + poolA: {rels: map[string]*controlplane.ReplicationRelationship{}, conn: map[string]map[string]string{}}, + clusterB + "/" + poolB: {rels: map[string]*controlplane.ReplicationRelationship{}, conn: map[string]map[string]string{}}, + } + for i := 0; i < moves; i++ { + srcC, srcP := cluster(i) + tgtC, tgtP := cluster(i + 1) + clients[srcC+"/"+srcP].rels[member(i)] = &controlplane.ReplicationRelationship{ + SourceLvolID: member(i), TargetLvolID: member(i + 1), + SourceClusterID: srcC, TargetClusterID: tgtC, TargetPoolID: tgtP, + ActiveLvolID: active, + } + } + activeC, activeP := cluster(moves) + clients[activeC+"/"+activeP].conn[active] = map[string]string{"nqn": "nqn.test:" + active} + + orig := clusterClientFor + defer func() { clusterClientFor = orig }() + clusterClientFor = func(_ context.Context, clusterID, poolID string) (controlplane.ClusterAPI, error) { + if c, ok := clients[clusterID+"/"+poolID]; ok { + return c, nil + } + return nil, errors.New("unexpected cluster " + clusterID + "/" + poolID) + } + + connInfo := redirectToActiveVolume(context.Background(), clients[clusterA+"/"+poolA], member(0), + clusterA+":"+poolA+":"+member(0), map[string]string{"hostNQN": "nqn.host"}) + if connInfo == nil { + t.Fatalf("redirect returned nil: the walk gave up before the %d-hop chain's active volume", moves) + } + if got := connInfo["nqn"]; got != "nqn.test:"+active { + t.Errorf("connection nqn = %q, want the active volume's %q", got, "nqn.test:"+active) + } + if got := connInfo[csicommon.ParamClusterID]; got != activeC { + t.Errorf("cluster_id = %q, want the active volume's cluster %q", got, activeC) + } +} + +// A relationship that points back at a member already walked must end the +// walk instead of spinning to the bound. +func TestRedirectToActiveVolumeStopsOnALoop(t *testing.T) { + const ( + clusterA = "aaaaaaaa-0000-0000-0000-000000000001" + poolA = "aaaaaaaa-0000-0000-0000-00000000000a" + x = "11111111-1111-1111-1111-111111111111" + y = "22222222-2222-2222-2222-222222222222" + ) + client := &fakeRelationshipAPI{rels: map[string]*controlplane.ReplicationRelationship{ + x: {SourceLvolID: x, TargetLvolID: y, SourceClusterID: clusterA, TargetClusterID: clusterA, TargetPoolID: poolA, ActiveLvolID: "zz"}, + y: {SourceLvolID: y, TargetLvolID: x, SourceClusterID: clusterA, TargetClusterID: clusterA, TargetPoolID: poolA, ActiveLvolID: "zz"}, + }} + orig := clusterClientFor + defer func() { clusterClientFor = orig }() + clusterClientFor = func(context.Context, string, string) (controlplane.ClusterAPI, error) { return client, nil } + if got := redirectToActiveVolume(context.Background(), client, x, clusterA+":"+poolA+":"+x, map[string]string{}); got != nil { + t.Fatalf("a looping chain returned %v, want nil", got) + } +} From f8760ac5af3af96cd21b5bcc146621160da412d8 Mon Sep 17 00:00:00 2001 From: michael Date: Sat, 3 Oct 2026 14:31:26 +0300 Subject: [PATCH 178/206] operator: TestFailover CRD as controller-gen renders it The hand-written sourceVolumeMode entry differed from the generated one in wrapping; the Manifests check keeps the three copies identical to the generator's output. Co-Authored-By: Claude Fable 5.1 --- .../crds/storage.simplyblock.io_testfailovers.yaml | 14 +++++++------- .../storage.simplyblock.io_testfailovers.yaml | 14 +++++++------- .../storage.simplyblock.io_testfailovers.yaml | 14 +++++++------- 3 files changed, 21 insertions(+), 21 deletions(-) diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_testfailovers.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_testfailovers.yaml index 1cca0647d..d8b008df6 100644 --- a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_testfailovers.yaml +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_testfailovers.yaml @@ -185,13 +185,6 @@ spec: empty fsType makes the node plugin default to ext4 and refuse to mount an XFS volume. type: string - sourceVolumeMode: - description: |- - SourceVolumeMode is the source PV's volumeMode (Filesystem or Block), - carried onto the bubble PV and PVC. A VM's disk is a Block claim; a bubble - claim that omitted the mode defaulted to Filesystem and the kubelet asked - the node plugin to mount a raw guest disk (2026-10-03). - type: string sourceHandle: description: SourceHandle is the source volume's backend handle, read from its PV. @@ -213,6 +206,13 @@ spec: identity keys are dropped so a failed clone lookup can never point the mount back at the source. type: object + sourceVolumeMode: + description: |- + SourceVolumeMode is the source PV's volumeMode (Filesystem or Block), + carried onto the bubble PV and PVC. A VM's disk is a Block claim; a bubble + claim that omitted the mode defaulted to Filesystem and the kubelet asked + the node plugin to mount a raw guest disk (2026-10-03). + type: string required: - sourceRef type: object diff --git a/operator/config/crd/bases/storage.simplyblock.io_testfailovers.yaml b/operator/config/crd/bases/storage.simplyblock.io_testfailovers.yaml index 1cca0647d..d8b008df6 100644 --- a/operator/config/crd/bases/storage.simplyblock.io_testfailovers.yaml +++ b/operator/config/crd/bases/storage.simplyblock.io_testfailovers.yaml @@ -185,13 +185,6 @@ spec: empty fsType makes the node plugin default to ext4 and refuse to mount an XFS volume. type: string - sourceVolumeMode: - description: |- - SourceVolumeMode is the source PV's volumeMode (Filesystem or Block), - carried onto the bubble PV and PVC. A VM's disk is a Block claim; a bubble - claim that omitted the mode defaulted to Filesystem and the kubelet asked - the node plugin to mount a raw guest disk (2026-10-03). - type: string sourceHandle: description: SourceHandle is the source volume's backend handle, read from its PV. @@ -213,6 +206,13 @@ spec: identity keys are dropped so a failed clone lookup can never point the mount back at the source. type: object + sourceVolumeMode: + description: |- + SourceVolumeMode is the source PV's volumeMode (Filesystem or Block), + carried onto the bubble PV and PVC. A VM's disk is a Block claim; a bubble + claim that omitted the mode defaulted to Filesystem and the kubelet asked + the node plugin to mount a raw guest disk (2026-10-03). + type: string required: - sourceRef type: object diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_testfailovers.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_testfailovers.yaml index 1cca0647d..d8b008df6 100644 --- a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_testfailovers.yaml +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_testfailovers.yaml @@ -185,13 +185,6 @@ spec: empty fsType makes the node plugin default to ext4 and refuse to mount an XFS volume. type: string - sourceVolumeMode: - description: |- - SourceVolumeMode is the source PV's volumeMode (Filesystem or Block), - carried onto the bubble PV and PVC. A VM's disk is a Block claim; a bubble - claim that omitted the mode defaulted to Filesystem and the kubelet asked - the node plugin to mount a raw guest disk (2026-10-03). - type: string sourceHandle: description: SourceHandle is the source volume's backend handle, read from its PV. @@ -213,6 +206,13 @@ spec: identity keys are dropped so a failed clone lookup can never point the mount back at the source. type: object + sourceVolumeMode: + description: |- + SourceVolumeMode is the source PV's volumeMode (Filesystem or Block), + carried onto the bubble PV and PVC. A VM's disk is a Block claim; a bubble + claim that omitted the mode defaulted to Filesystem and the kubelet asked + the node plugin to mount a raw guest disk (2026-10-03). + type: string required: - sourceRef type: object From 8233d4bba45563abf7a45e21c83213df9706a160 Mon Sep 17 00:00:00 2001 From: michael Date: Sat, 3 Oct 2026 14:59:47 +0300 Subject: [PATCH 179/206] operator: StorageSiteDeployment, a managed site's storage deployed from the hub A site's storage cluster is built from objects on the site's API server (the OperatorOps discovery, the ClusterDeploymentConfig draft it writes, the StorageCluster the approved draft expands into), which a hub managing the site through OCM does not reach. StorageSiteDeployment is the hub-side request: the site, the discovery filters, the sizing, and the one-way approval. Its controller (hub-only, registered with TestFailover when the ManifestWork API is served) carries the request through one ManifestWork in the site's hub namespace -- the discovery first, then a server-side apply of the sizing and the approval onto the draft once it names nodes -- and projects the draft, the StorageCluster and its nodes back through ManagedClusterViews. Phases Pending, Discovering, Drafted, Deploying, Online, Failed; conditions Delivered, Discovered, Approved, Ready; the cluster's uuid and pool are in the status, which is what a StorageClass names. Deleting the request orphans the work's resources: the storage cluster is never torn down by withdrawing the request. The console's ClusterRole may manage the kind and read ManagedClusters. The TestFailover CRD copies are the generator's rendering (as on #618). Design: simplyblock-dr docs/design/control-center-managed-discovery.md. Co-Authored-By: Claude Opus 5.5 --- ...simplyblock.io_storagesitedeployments.yaml | 1057 +++++++++++++++++ .../templates/control-center-rbac.yaml | 8 + .../templates/roles/manager_role.yaml | 11 + .../v1alpha2/storagesitedeployment_types.go | 298 +++++ .../api/v1alpha2/zz_generated.deepcopy.go | 251 ++++ operator/cmd/main.go | 12 +- ...simplyblock.io_storagesitedeployments.yaml | 1057 +++++++++++++++++ operator/config/crd/kustomization.yaml | 1 + operator/config/rbac/role.yaml | 11 + .../storagesitedeployment_controller.go | 719 +++++++++++ ...ragesitedeployment_controller_unit_test.go | 335 ++++++ ...simplyblock.io_storagesitedeployments.yaml | 1057 +++++++++++++++++ 12 files changed, 4816 insertions(+), 1 deletion(-) create mode 100644 helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagesitedeployments.yaml create mode 100644 operator/api/v1alpha2/storagesitedeployment_types.go create mode 100644 operator/config/crd/bases/storage.simplyblock.io_storagesitedeployments.yaml create mode 100644 operator/internal/controller/storagesitedeployment_controller.go create mode 100644 operator/internal/controller/storagesitedeployment_controller_unit_test.go create mode 100644 operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagesitedeployments.yaml diff --git a/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagesitedeployments.yaml b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagesitedeployments.yaml new file mode 100644 index 000000000..9c2718a8f --- /dev/null +++ b/helm-charts/charts/simplyblock-operator/crds/storage.simplyblock.io_storagesitedeployments.yaml @@ -0,0 +1,1057 @@ +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition +metadata: + annotations: + controller-gen.kubebuilder.io/version: v0.21.0 + name: storagesitedeployments.storage.simplyblock.io +spec: + group: storage.simplyblock.io + names: + kind: StorageSiteDeployment + listKind: StorageSiteDeploymentList + plural: storagesitedeployments + shortNames: + - sbsd + singular: storagesitedeployment + scope: Namespaced + versions: + - additionalPrinterColumns: + - jsonPath: .spec.cluster + name: Cluster + type: string + - jsonPath: .spec.approved + name: Approved + type: boolean + - jsonPath: .status.phase + name: Phase + type: string + - jsonPath: .status.draft.phase + name: Draft + type: string + - jsonPath: .status.storageCluster.phase + name: Storage + type: string + - jsonPath: .status.message + name: Message + priority: 1 + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha2 + schema: + openAPIV3Schema: + description: |- + StorageSiteDeployment requests a managed site's storage cluster from the hub: + a discovery on the site, the sizing of the draft it writes, and the approval + that expands the draft into a StorageCluster. The hub carries the request + through OCM and projects the site's draft and cluster into the status. + Deleting the request leaves the storage cluster alone. + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: StorageSiteDeploymentSpec is the request for one site's storage + cluster. + properties: + approved: + default: false + description: |- + Approved is the review gate, delivered to the draft on the site. One-way, + as the draft's own gate is. + type: boolean + cluster: + description: |- + Cluster is the OCM ManagedCluster the storage is deployed on. The request's + ManifestWork and views live in its namespace on the hub. Immutable. + maxLength: 63 + minLength: 1 + type: string + x-kubernetes-validations: + - message: cluster is immutable + rule: self == oldSelf + discover: + description: |- + Discover is the discovery the site runs first. Changing it runs another + discovery, which rewrites the draft. + properties: + enableControlPlaneNodes: + description: |- + EnableControlPlaneNodes lets the discovery consider the nodes that run the + API server. Every server of a small distribution is one, so a three-node + site has no storage without it. + type: boolean + nodeSelector: + additionalProperties: + type: string + description: NodeSelector limits the discovery to the nodes carrying + these labels. + type: object + workers: + description: Workers limits the discovery to these nodes. Empty + is every worker. + items: + type: string + type: array + x-kubernetes-list-type: set + type: object + draftName: + default: site-draft + description: |- + DraftName is the ClusterDeploymentConfig the discovery writes on the site + and the request sizes and approves. Immutable. + maxLength: 63 + type: string + x-kubernetes-validations: + - message: draftName is immutable + rule: self == oldSelf + siteNamespace: + default: simplyblock + description: |- + SiteNamespace is the simplyblock operator's namespace on the site, where + the discovery and the draft live. + maxLength: 63 + type: string + sizing: + description: |- + Sizing is written onto the draft's cluster template once the draft exists, + so the reviewer sees the sized draft before approving it. + properties: + enableDriveFormat: + description: EnableDriveFormat lets the deployment format the + devices it takes. + type: boolean + enableJournalDevice: + description: EnableJournalDevice dedicates one device per node + to the journal. + type: boolean + maxSubsystemCount: + description: MaxSubsystemCount is the number of NVMe-oF subsystems + each node serves. + format: int32 + minimum: 1 + type: integer + minHugePagesSize: + description: |- + MinHugePagesSize is the hugepage memory each storage node takes, as a + quantity ("8G"). + type: string + name: + description: Name is the StorageCluster's name on the site. + maxLength: 63 + type: string + stripe: + description: Stripe is the erasure-coding layout. + properties: + dataChunks: + description: DataChunks is the number of data chunks per stripe + (ndcs). + format: int32 + minimum: 1 + type: integer + parityChunks: + description: |- + ParityChunks is the number of parity chunks per stripe (npcs), and + therefore how many chunk losses a stripe survives. + format: int32 + minimum: 0 + type: integer + type: object + x-kubernetes-validations: + - message: the erasure-coding scheme must be one of 1+0, 1+1, + 2+1, 4+1, 1+2, 2+2, or 4+2, written as dataChunks+parityChunks, + and an unstated half is 1 + rule: '[has(self.dataChunks) ? self.dataChunks : 1, has(self.parityChunks) + ? self.parityChunks : 1] in [[1, 0], [1, 1], [2, 1], [4, 1], + [1, 2], [2, 2], [4, 2]]' + vcpuCount: + description: VCPUCount is the number of vCPUs each storage node + takes. + format: int32 + minimum: 1 + type: integer + type: object + required: + - cluster + type: object + x-kubernetes-validations: + - message: 'approval is one-way: an approved deployment cannot be un-approved' + rule: '!has(oldSelf.approved) || !oldSelf.approved || self.approved' + status: + description: StorageSiteDeploymentStatus is what the site reports back, + projected. + properties: + conditions: + description: |- + Conditions: Delivered (the work is applied on the site), Discovered (the + draft names nodes), Approved (the site's draft is approved), Ready (the + StorageCluster is Online). + items: + description: Condition contains details for one aspect of the current + state of this API Resource. + properties: + lastTransitionTime: + description: |- + lastTransitionTime is the last time the condition transitioned from one status to another. + This should be when the underlying condition changed. If that is not known, then using the time when the API field changed is acceptable. + format: date-time + type: string + message: + description: |- + message is a human readable message indicating details about the transition. + This may be an empty string. + maxLength: 32768 + type: string + observedGeneration: + description: |- + observedGeneration represents the .metadata.generation that the condition was set based upon. + For instance, if .metadata.generation is currently 12, but the .status.conditions[x].observedGeneration is 9, the condition is out of date + with respect to the current state of the instance. + format: int64 + minimum: 0 + type: integer + reason: + description: |- + reason contains a programmatic identifier indicating the reason for the condition's last transition. + Producers of specific condition types may define expected values and meanings for this field, + and whether the values are considered a guaranteed API. + The value should be a CamelCase string. + This field may not be empty. + maxLength: 1024 + minLength: 1 + pattern: ^[A-Za-z]([A-Za-z0-9_,:]*[A-Za-z0-9_])?$ + type: string + status: + description: status of the condition, one of True, False, Unknown. + enum: + - "True" + - "False" + - Unknown + type: string + type: + description: type of condition in CamelCase or in foo.example.com/CamelCase. + maxLength: 316 + pattern: ^([a-z0-9]([-a-z0-9]*[a-z0-9])?(\.[a-z0-9]([-a-z0-9]*[a-z0-9])?)*/)?(([A-Za-z0-9][-A-Za-z0-9_.]*)?[A-Za-z0-9])$ + type: string + required: + - lastTransitionTime + - message + - reason + - status + - type + type: object + type: array + x-kubernetes-list-map-keys: + - type + x-kubernetes-list-type: map + draft: + description: Draft is the draft as the site reports it. + properties: + approved: + description: Approved is whether the draft is approved on the + site. + type: boolean + cluster: + description: Cluster is the draft's cluster template, with the + sizing applied. + properties: + backup: + description: |- + Backup is where this cluster's backups live, and it expands into + StorageCluster.spec.backup unchanged. + + It is here for the reason KMS is: a store stated on the document is + present when the cluster is created rather than patched in afterward by + whoever remembers. Unlike most of what this template carries, the field it + fills is mutable, so a document that states none costs nothing permanent. + A cluster can be given a store whenever there is one to give. + + The Secret it names is not resolved at admission. It is a core object a + deployment legitimately creates alongside the document or after it, and + the cluster's own creation is where its absence is reported. + properties: + bucket: + description: Bucket is the bucket backups are written + to and read from. + type: string + credentialsSecretRef: + description: |- + CredentialsSecretRef names the Secret holding the access key and the + secret key. It is a reference rather than the values, because a spec is + readable by anybody who can read the object. + properties: + name: + default: "" + description: |- + Name of the referent. + This field is effectively required, but due to backwards compatibility is + allowed to be empty. Instances of this type with an empty value here are + almost certainly wrong. + More info: https://kubernetes.io/docs/concepts/overview/working-with-objects/names/#names + type: string + type: object + x-kubernetes-map-type: atomic + endpoint: + description: Endpoint is the S3 endpoint, for example, + https://s3.example.com. + pattern: ^https?://[a-zA-Z0-9.-]+(:[0-9]{1,5})?(/.*)?$ + type: string + prefix: + description: |- + Prefix narrows the store to one key prefix, so that several clusters can + share a bucket without each walking the others' backups. + type: string + region: + description: Region is the bucket's region, for endpoints + that do not imply one. + type: string + required: + - bucket + - credentialsSecretRef + - endpoint + type: object + containerResources: + description: |- + ContainerResources sizes the storage-node container, and expands into the + cluster's own spec.storageNodes.containerResources. + + The container it sizes is the node's management API rather than SPDK, + which runs in a pod of its own: what outgrows the default is a node + answering for many subsystems, not a node moving more data. It is on the + document because a deployment is where a fleet's sizing is decided, and + a cluster written from a document that could not say so had to be edited + afterward on a field the document owns everywhere else. + + Stating either half replaces both. The defaults apply to a cluster that + states neither requests nor limits, so a document stating requests alone + produces a container with no limits rather than one with the default + limits, and a memory limit is what has the kubelet evict a leaking agent + rather than losing the worker. + + It is a pointer because a resource block is a struct, and a struct with + omitempty is serialized whether or not anything is in it: as a value, + every document a discovery run writes would carry an empty + containerResources that says nothing and that a reviewer has to decide + about. + properties: + claims: + description: |- + Claims lists the names of resources, defined in spec.resourceClaims, + that are used by this container. + + This field depends on the + DynamicResourceAllocation feature gate. + + This field is immutable. It can only be set for containers. + items: + description: ResourceClaim references one entry in PodSpec.ResourceClaims. + properties: + name: + description: |- + Name must match the name of one entry in pod.spec.resourceClaims of + the Pod where this field is used. It makes that resource available + inside a container. + type: string + request: + description: |- + Request is the name chosen for a request in the referenced claim. + If empty, everything from the claim is made available, otherwise + only the result of this request. + type: string + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + limits: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Limits describes the maximum amount of compute resources allowed. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + requests: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Requests describes the minimum amount of compute resources required. + If Requests is omitted for a container, it defaults to Limits if that is explicitly specified, + otherwise to an implementation-defined value. Requests cannot exceed Limits. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + type: object + enableAtomicity4K: + description: |- + EnableAtomicity4K enforces 4K write atomicity on every device this + deployment names, which is what lets checksum validation run on devices + whose logical block size is under the data plane's 4K minimum. + + It is the route to checked I/O on a device that cannot be reformatted: a + logical block device's block size is fixed by the drive, and some NVMe + devices offer no 4K format either. Where a device can be reformatted, + EnableDriveFormat is the other route and this is unnecessary. + + It is an enforcement because the question is often unanswerable. A SATA + drive presenting 512-byte logical blocks over a 4K physical sector reports + 512 and nothing more, and a kernel older than 6.11 publishes no atomic + write attributes at all. Where a device does answer, the storage node's + report carries it, and a reviewer approves this against that rather than + against a vendor's datasheet -- because enforcing a guarantee the hardware + does not keep is how a torn write becomes a checksum that silently + disagrees with it. + + It means nothing unless EnableChecksumValidation is set, which is the + cluster's own rule and is left to the cluster to enforce. + type: boolean + enableChecksumValidation: + description: |- + EnableChecksumValidation turns on inline CRC validation of every I/O, for + silent-data-error protection. + + It is on the document because it is immutable on the cluster it lands on: + the backend bakes the checksum method into each device when the cluster is + created and never re-applies it, so a cluster created without this is one + nobody can turn it on for. A deployment that wants its data checked has to + say so here or not at all. + type: boolean + enableDriveFormat: + description: |- + EnableDriveFormat formats every device the document names before a storage + node takes it, which is how a drive carrying anything already is made + usable. + + It says what is wanted rather than how, because the how differs by device + class: an NVMe device is formatted to a 4K block size, and a logical block + device has its signatures wiped. One field covers both, so a document does + not have to know which class the expansion will resolve it to. + + It is on the document rather than defaulted further down because it is + destructive and the document is what somebody approves. A reviewer reading + a draft has to see that the drives it lists will be formatted, and be able + to strike it before approving; the cluster's own field is immutable once + the cluster exists, so a default nobody saw could not be undone either. + type: boolean + enableFailureDomains: + description: |- + EnableFailureDomains opts the cluster into failure-domain mode, in which + every group must label the fault group its workers belong to. + type: boolean + enableJournalDevice: + description: |- + EnableJournalDevice dedicates the smallest NVMe device on each of this + deployment's workers to the journal manager, instead of carving a journal + partition out of every device. + + It is here rather than on a node set because it is immutable on the cluster + it lands on, for the reason SocketsToUse is: the on-disk layout a fleet was + built with is not one a later document can vary. It also costs a drive of + capacity per node, which is a trade a reviewer approves rather than one a + default makes for them. + type: boolean + enableNodeAffinity: + description: |- + EnableNodeAffinity has the data plane serve an erasure-coded volume's I/O + from the local node's own devices where it can, before crossing the + network. + + It is not Kubernetes affinity, and the name is the one place this API + invites that reading: nothing about it schedules a pod, labels a worker, + or places a volume's primary node. The control plane carries it into the + cluster map it pushes to each node, where it sets the local node's index, + and what changes is which copy of a chunk is read. + Co-locating a workload with the primary node of its volume is a separate + mechanism and is not configured here. + + It is on the document because it is immutable on the cluster: the control + plane takes it at cluster create and never re-applies it, so this is the + only moment it can be set at all. + type: boolean + fabricType: + description: FabricType is the storage fabric. + maxLength: 32 + type: string + initContainerResources: + description: |- + InitContainerResources sizes both of the storage node's init containers, + and expands into the cluster's own spec.storageNodes.initContainerResources. + + They are sized apart from the container because they do a different job + and are gone before it starts: one writes the node's env file and the + other runs node_configure.py once, so what they need is a short burst + rather than the footprint of a process that runs for the node's life. + + Stating either half replaces both, as with containerResources, and it is + a pointer for the same reason. + properties: + claims: + description: |- + Claims lists the names of resources, defined in spec.resourceClaims, + that are used by this container. + + This field depends on the + DynamicResourceAllocation feature gate. + + This field is immutable. It can only be set for containers. + items: + description: ResourceClaim references one entry in PodSpec.ResourceClaims. + properties: + name: + description: |- + Name must match the name of one entry in pod.spec.resourceClaims of + the Pod where this field is used. It makes that resource available + inside a container. + type: string + request: + description: |- + Request is the name chosen for a request in the referenced claim. + If empty, everything from the claim is made available, otherwise + only the result of this request. + type: string + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + limits: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Limits describes the maximum amount of compute resources allowed. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + requests: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Requests describes the minimum amount of compute resources required. + If Requests is omitted for a container, it defaults to Limits if that is explicitly specified, + otherwise to an implementation-defined value. Requests cannot exceed Limits. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + type: object + kms: + description: |- + KMS selects where the cluster stores volume encryption keys. Stating it on + the document is what makes it present when the cluster is created, where + setting it on the StorageCluster afterward races with that creation. + properties: + vault: + description: Vault stores keys in HashiCorp Vault. + properties: + endpoint: + description: |- + Endpoint is the Vault endpoint, for example, https://vault.example.com:8200. + Rejected unless it resolves to an external address. + pattern: ^https?://[a-zA-Z0-9.-]+(:[0-9]{1,5})?(/.*)?$ + type: string + required: + - endpoint + type: object + type: object + maxSubsystemCount: + description: |- + MaxSubsystemCount is the maximum number of NVMe-oF subsystems each storage + node of this cluster serves. Required, because the StorageCluster's own + field is, and no StorageNode carries a copy of it. + format: int32 + maximum: 75 + minimum: 10 + type: integer + minHugePagesSize: + description: |- + MinHugePagesSize is the smallest huge-page allocation each storage node of + this cluster makes: 100G or 1T, where a bare number is gigabytes. Like + VCPUCount it is the cluster's and is copied onto every node the expansion + writes. Omitted, each node uses the computed minimum. + maxLength: 32 + type: string + name: + description: |- + Name is the StorageCluster's name, and is therefore held to what such a + name may be rather than to what an object name may be. A longer value is a + document the API server accepts and a CreatingCluster step that can never + succeed, since the cluster it would write is one the API server refuses. + maxLength: 63 + type: string + nodeProvisioningBudget: + description: |- + NodeProvisioningBudget is how many workers the expansion may have in the + node-add process at once. It expands into the cluster's own + spec.storageNodes.nodeProvisioningBudget, whose meaning it shares: the cap + is counted by distinct worker, so a two-socket host spends one of the + budget, and a worker hosting a FoundationDB pod is sequential whatever the + budget says. + + It is on the document because a document is what states the size of a + deployment, and a deployment of thirty workers added one at a time is the + difference between an afternoon and a week. Omitted, the cluster's default + of one applies, which is the serial behavior. + format: int32 + minimum: 1 + type: integer + nodesPerSocket: + description: |- + NodesPerSocket is how many storage nodes run per NUMA socket. See + SocketsToUse, which it multiplies. + format: int32 + maximum: 8 + minimum: 1 + type: integer + openshift: + description: |- + OpenShift is what this deployment states because it runs on OpenShift. It + expands into StorageCluster.spec.storageNodes.openshift, whose shape it + shares, and it is read only for a document whose environment is + OpenShift: the environment is what says which distribution this is, and + the block is what that distribution needs said beyond it. + properties: + machineConfigPool: + default: worker + description: |- + MachineConfigPool names a machine-config role the storage nodes' own pool + inherits from, beyond the worker role it always inherits. + + It is not the pool the nodes end up in, which the description it carried + before said and which cost a reader the reboot they were trying to avoid. + Adding a node creates a pool of its own, storage-, and moves the + node into it; a node belongs to exactly one custom pool, so whatever + machine configuration its previous pool carried is lost unless that + pool's role is named here for the new one to select as well. The default + is the role every pool already selects, which is what makes it a no-op + for a fleet whose workers are ordinary workers. + maxLength: 253 + pattern: ^[a-z0-9]([-a-z0-9]*[a-z0-9])?$ + type: string + type: object + ports: + description: |- + Ports are where this cluster's storage nodes listen. Unstated, and for + each member left unstated, the cluster's own defaults decide. + properties: + nodeAgent: + default: 50001 + description: |- + NodeAgent is the port each node's agent API listens on. It expands into + StorageCluster.spec.snodeApiPort, and it is named for the component + rather than for that field: the agent is what spec.images.nodeAgent pins + and what the storage-node DaemonSet runs. + format: int32 + maximum: 65535 + minimum: 1024 + type: integer + nvmf: + default: 4420 + description: |- + NVMf is the base of the NVMe-oF port range every node binds. It expands + into StorageCluster.spec.nvmfBasePort. + format: int32 + maximum: 65535 + minimum: 1024 + type: integer + rpc: + default: 8080 + description: |- + Rpc is the base of the RPC port range every node binds. It expands into + StorageCluster.spec.rpcBasePort. + format: int32 + maximum: 65535 + minimum: 1024 + type: integer + type: object + socketsToUse: + description: |- + SocketsToUse restricts the deployment to selected NUMA sockets, and empty + means socket 0 alone. With NodesPerSocket it decides how many storage nodes + each worker runs, so a group of two workers on a two-socket layout expands + to four nodes. + + It is here rather than on a node set because it is immutable on the cluster + it lands on: the layout a fleet was built with is not one a later document + can vary, and a reviewer should see it before the cluster exists. + items: + maxLength: 16 + type: string + maxItems: 16 + type: array + x-kubernetes-list-type: set + stripe: + description: Stripe is the erasure-coding layout. + properties: + dataChunks: + description: DataChunks is the number of data chunks per + stripe (ndcs). + format: int32 + minimum: 1 + type: integer + parityChunks: + description: |- + ParityChunks is the number of parity chunks per stripe (npcs), and + therefore how many chunk losses a stripe survives. + format: int32 + minimum: 0 + type: integer + type: object + x-kubernetes-validations: + - message: the erasure-coding scheme must be one of 1+0, 1+1, + 2+1, 4+1, 1+2, 2+2, or 4+2, written as dataChunks+parityChunks, + and an unstated half is 1 + rule: '[has(self.dataChunks) ? self.dataChunks : 1, has(self.parityChunks) + ? self.parityChunks : 1] in [[1, 0], [1, 1], [2, 1], [4, + 1], [1, 2], [2, 2], [4, 2]]' + tolerations: + description: |- + Tolerations are what the storage-node pods tolerate, and they expand into + the cluster's own spec.storageNodes.tolerations. + + A fleet that dedicates machines to storage taints them, which is what + keeps everything else off. The DaemonSet that lands on those machines has + to tolerate the taint or it schedules nowhere, and a document that could + not say so described a deployment that does not start: the correction was + an edit to the cluster the document had just created, on a field the + document owns everywhere else. + + A growth document states none. It names a cluster rather than describing + one, and that cluster already carries what its storage nodes tolerate. + items: + description: |- + The pod this Toleration is attached to tolerates any taint that matches + the triple using the matching operator . + properties: + effect: + description: |- + Effect indicates the taint effect to match. Empty means match all taint effects. + When specified, allowed values are NoSchedule, PreferNoSchedule and NoExecute. + type: string + key: + description: |- + Key is the taint key that the toleration applies to. Empty means match all taint keys. + If the key is empty, operator must be Exists; this combination means to match all values and all keys. + type: string + operator: + description: |- + Operator represents a key's relationship to the value. + Valid operators are Exists, Equal, Lt, and Gt. Defaults to Equal. + Exists is equivalent to wildcard for value, so that a pod can + tolerate all taints of a particular category. + Lt and Gt perform numeric comparisons (requires feature gate TaintTolerationComparisonOperators). + type: string + tolerationSeconds: + description: |- + TolerationSeconds represents the period of time the toleration (which must be + of effect NoExecute, otherwise this field is ignored) tolerates the taint. By default, + it is not set, which means tolerate the taint forever (do not evict). Zero and + negative values will be treated as 0 (evict immediately) by the system. + format: int64 + type: integer + value: + description: |- + Value is the taint value the toleration matches to. + If the operator is Exists, the value should be empty, otherwise just a regular string. + type: string + type: object + maxItems: 32 + type: array + vcpuCount: + description: |- + VCPUCount is the number of vCPUs allocated to SPDK on each storage node of + this cluster. It is stated here and nowhere below, because the control + plane assumes it uniform across a cluster's nodes; CreatingNodes copies it + into every StorageNode.spec.config.sizing it writes. Required, because the + StorageCluster's own field is. + The floor is 4 rather than a hardware limit: a node must carry one core + beyond this budget for the system, and the control plane's core layout + assigns no NVMe-oF poller core at all for a 2-vCPU budget. + format: int32 + minimum: 4 + type: integer + required: + - maxSubsystemCount + - name + - vcpuCount + type: object + message: + description: |- + Message is what the site says about the draft: validation findings while + it is a draft, the expansion's step afterwards. + type: string + name: + description: Name is the ClusterDeploymentConfig on the site. + type: string + nodeRefs: + description: NodeRefs are the StorageNode objects the expansion + created. + items: + type: string + type: array + x-kubernetes-list-type: set + nodeSets: + description: NodeSets are the nodes and devices the discovery + found, for review. + items: + description: |- + NodeSet is the organizational grouping of a deployment, usually a rack: the + workers a document adds or grows together. It carries no sizing, because sizing + is uniform across a cluster and is stated once in ClusterTemplate. + properties: + groups: + description: Groups are the sets of workers sharing one + configuration. + items: + description: |- + NodeGroup is a set of workers that share one configuration, which is what + makes ten identical machines one entry rather than ten. + properties: + dataInterfaces: + description: DataInterfaces are the data-plane network + interfaces. + items: + maxLength: 63 + type: string + maxItems: 32 + type: array + devices: + description: Devices selects the storage devices every + worker in the group uses. + properties: + block: + description: |- + Block names logical block devices by path ("/dev/sdb"). It expands into the + same config.deviceNames as NVMe, which takes a PCI address and a device + path in one list. It is the alternative to NVMe rather than a companion of + it: the two classes are not mixed within a cluster. + items: + maxLength: 255 + pattern: ^/dev/[a-zA-Z0-9._/-]+$ + type: string + maxItems: 128 + type: array + x-kubernetes-list-type: set + nvme: + description: NVMe names NVMe devices by PCI address + ("0000:5e:00.0"). + items: + maxLength: 32 + pattern: ^[0-9a-fA-F]{4}:[0-9a-fA-F]{2}:[0-9a-fA-F]{2}\.[0-9a-fA-F]$ + type: string + maxItems: 128 + type: array + x-kubernetes-list-type: set + type: object + x-kubernetes-validations: + - message: a device selection names NVMe addresses + or block devices, not both + rule: has(self.nvme) != has(self.block) + failureDomain: + description: |- + FailureDomain is the label of the fault group every worker in this group + belongs to ("rack-b"), which is usually the name of the rack, zone, or + power feed they share. Discovery seeds it from topology.kubernetes.io/zone + and leaves it unset where the Kubernetes API carries no topology, which + holds provisioning with a clear reason rather than guessing. It expands + into StorageNode.spec.config.failureDomain, whose shape it shares. + maxLength: 63 + pattern: ^[a-zA-Z0-9]([-_.a-zA-Z0-9]*[a-zA-Z0-9])?$ + type: string + journalManager: + description: JournalManager tunes the journal managers + on these nodes. + properties: + count: + description: Count is the number of journal managers + to configure. + format: int32 + minimum: 1 + type: integer + percentPerDevice: + description: PercentPerDevice is the share of + each device given to the journal. + format: int32 + maximum: 100 + minimum: 1 + type: integer + type: object + mgmtInterface: + description: MgmtInterface is the management network + interface the storage nodes bind. + maxLength: 63 + type: string + name: + description: |- + Name identifies the group within its node set, for a reader and for the + events a validation failure emits. + maxLength: 253 + type: string + reservedSystemCPU: + description: |- + ReservedSystemCPU is the CPU set held back from SPDK for the system on + these nodes, as a core list such as 0,1 or 0-3. + + It is a group's rather than the cluster's because it names core ids, and a + group is what a document calls the workers that share their hardware: 0,1 + on a sixteen-core worker and 0,1 on a ninety-six-core worker are different + fractions of the machine. It expands into + StorageNode.spec.config.reservedSystemCPU, whose shape it shares, and a + group that states none leaves the cluster's fleet-wide value to decide. + + On OpenShift it reaches the kubelet through a KubeletConfig for the + machine config pool, which is the cluster's, so groups that disagree there + are writing over one another's pool configuration. + maxLength: 63 + pattern: ^[0-9]+(-[0-9]+)?(,[0-9]+(-[0-9]+)?)*$ + type: string + spdkSystemMemory: + description: |- + SpdkSystemMemory is the memory the control plane starts SPDK with on these + nodes. + maxLength: 32 + pattern: ^[0-9]+(G|GI|GB|GiB|M|MI|MB|MiB|g|gi|gb|gib|m|mi|mb|mib)?$ + type: string + workers: + description: Workers are the Kubernetes worker hostnames + in this group. + items: + maxLength: 253 + type: string + maxItems: 200 + minItems: 1 + type: array + x-kubernetes-list-type: set + required: + - name + - workers + type: object + maxItems: 64 + minItems: 1 + type: array + name: + description: |- + Name is the node set's name. It is copied to StorageNode.spec.nodeSet, so + that a node can be traced back to the part of the document that produced + it. + maxLength: 253 + type: string + required: + - groups + - name + type: object + type: array + phase: + description: |- + Phase is the draft's own phase on the site (Draft, Expanding, Expanded, + Failed). + type: string + required: + - name + type: object + message: + description: |- + Message is the reason the phase is what it is: one sentence, replaced as + the request moves, and never a log. + type: string + observedGeneration: + description: |- + ObservedGeneration is the generation the rest of this status was computed + from. + format: int64 + type: integer + phase: + description: Phase is the request's own progress. + enum: + - Pending + - Discovering + - Drafted + - Deploying + - Online + - Failed + type: string + storageCluster: + description: StorageCluster is the cluster the approved draft produced. + properties: + name: + description: Name is the StorageCluster object on the site. + type: string + nodes: + description: Nodes are the cluster's storage nodes. + items: + description: |- + StorageSiteNode is one storage node of the deployed cluster, as the site + reports it. + properties: + hostname: + description: Hostname is the Kubernetes node it runs on. + type: string + name: + description: Name is the StorageNode object on the site. + type: string + phase: + description: Phase is the node's phase on the site. + type: string + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + phase: + description: Phase is the StorageCluster's phase on the site. + type: string + pool: + description: |- + Pool is the pool the cluster was created with, which a StorageClass names + in pool_name. + type: string + uuid: + description: |- + UUID is the storage cluster's id in the control plane, which a + StorageClass names in cluster_id. + type: string + required: + - name + type: object + workName: + description: WorkName is the ManifestWork carrying the request to + the site. + type: string + type: object + type: object + served: true + storage: true + subresources: + status: {} diff --git a/helm-charts/charts/simplyblock-operator/templates/control-center-rbac.yaml b/helm-charts/charts/simplyblock-operator/templates/control-center-rbac.yaml index c1bd77089..4a638e58c 100644 --- a/helm-charts/charts/simplyblock-operator/templates/control-center-rbac.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/control-center-rbac.yaml @@ -144,6 +144,14 @@ rules: - apiGroups: ["sitemap.simplyblock.io"] resources: ["dhcpservers"] verbs: ["create", "update", "patch", "delete"] + # a managed site's storage deployment is requested, sized and approved + # from the hub console (StorageSiteDeployment, carried to the site by OCM) + - apiGroups: ["storage.simplyblock.io"] + resources: ["storagesitedeployments"] + verbs: ["get", "list", "watch", "create", "update", "patch", "delete"] + - apiGroups: ["cluster.open-cluster-management.io"] + resources: ["managedclusters"] + verbs: ["get", "list", "watch"] # the console asks the API server what its identity may do, to disable controls - apiGroups: ["authorization.k8s.io"] resources: ["selfsubjectaccessreviews", "selfsubjectrulesreviews"] diff --git a/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml b/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml index c6801eb7f..3f8c28526 100644 --- a/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/roles/manager_role.yaml @@ -176,6 +176,14 @@ rules: - patch - update - watch +- apiGroups: + - cluster.open-cluster-management.io + resources: + - managedclusters + verbs: + - get + - list + - watch - apiGroups: - coordination.k8s.io resources: @@ -325,6 +333,7 @@ rules: - storagenodes - storagepoolops - storagepools + - storagesitedeployments - tasks - testfailovers - volumemigrations @@ -361,6 +370,7 @@ rules: - storagenodes/finalizers - storagepoolops/finalizers - storagepools/finalizers + - storagesitedeployments/finalizers - tasks/finalizers - testfailovers/finalizers - volumemigrations/finalizers @@ -394,6 +404,7 @@ rules: - storagenodesets/status - storagepoolops/status - storagepools/status + - storagesitedeployments/status - tasks/status - testfailovers/status - volumegroupsnapshotops/status diff --git a/operator/api/v1alpha2/storagesitedeployment_types.go b/operator/api/v1alpha2/storagesitedeployment_types.go new file mode 100644 index 000000000..93a18006c --- /dev/null +++ b/operator/api/v1alpha2/storagesitedeployment_types.go @@ -0,0 +1,298 @@ +// The storage deployment of a managed site, requested from the hub. +// +// A site's storage cluster is built from objects that live on the site's API +// server: an OperatorOps discovery, the ClusterDeploymentConfig draft it +// writes, and the StorageCluster the approved draft expands into. A hub that +// manages the site through Open Cluster Management does not reach that API +// server, so this kind is the hub-side request: it names the site and the +// sizing, and a controller carries the request to the site through a +// ManifestWork and projects the site's answer back through ManagedClusterViews. +// Approval is the same one-way gate the draft has on the site; it is flipped +// here and delivered there. +// +// Deleting the object withdraws nothing on the site: the storage cluster it +// requested stays, as a storage cluster is never torn down by deleting a +// request. Specified by docs/design/control-center-managed-discovery.md of the +// simplyblock-dr repository. + +package v1alpha2 + +import ( + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" +) + +// StorageSiteDeploymentPhase is the request's own progress. +// +kubebuilder:validation:Enum=Pending;Discovering;Drafted;Deploying;Online;Failed +type StorageSiteDeploymentPhase string + +const ( + // StorageSiteDeploymentPhasePending is the request before the hub delivered + // anything to the site. + StorageSiteDeploymentPhasePending StorageSiteDeploymentPhase = "Pending" + + // StorageSiteDeploymentPhaseDiscovering is the discovery running on the site: + // the draft is not written yet, or names no node yet. + StorageSiteDeploymentPhaseDiscovering StorageSiteDeploymentPhase = "Discovering" + + // StorageSiteDeploymentPhaseDrafted is a draft with nodes on the site, sized + // as the request says, awaiting approval. + StorageSiteDeploymentPhaseDrafted StorageSiteDeploymentPhase = "Drafted" + + // StorageSiteDeploymentPhaseDeploying is an approved draft expanding into a + // StorageCluster that is not Online yet. + StorageSiteDeploymentPhaseDeploying StorageSiteDeploymentPhase = "Deploying" + + // StorageSiteDeploymentPhaseOnline is the StorageCluster Online on the site. + StorageSiteDeploymentPhaseOnline StorageSiteDeploymentPhase = "Online" + + // StorageSiteDeploymentPhaseFailed is the site's own failure: the draft or the + // StorageCluster failed, or the hub could not deliver the request. + StorageSiteDeploymentPhaseFailed StorageSiteDeploymentPhase = "Failed" +) + +// StorageSiteDiscovery is the discovery the site runs: which nodes are +// inspected. It is the hub-side form of OperatorOps.spec.discover. +type StorageSiteDiscovery struct { + // EnableControlPlaneNodes lets the discovery consider the nodes that run the + // API server. Every server of a small distribution is one, so a three-node + // site has no storage without it. + // +optional + EnableControlPlaneNodes *bool `json:"enableControlPlaneNodes,omitempty"` + + // Workers limits the discovery to these nodes. Empty is every worker. + // +optional + // +listType=set + Workers []string `json:"workers,omitempty"` + + // NodeSelector limits the discovery to the nodes carrying these labels. + // +optional + NodeSelector map[string]string `json:"nodeSelector,omitempty"` +} + +// StorageSiteSizing is the cluster template written onto the draft before it +// is approved: the fields of ClusterDeploymentConfig.spec.cluster a reviewer +// decides. Absent fields keep what the discovery wrote. +type StorageSiteSizing struct { + // Name is the StorageCluster's name on the site. + // +kubebuilder:validation:MaxLength=63 + // +optional + Name string `json:"name,omitempty"` + + // VCPUCount is the number of vCPUs each storage node takes. + // +kubebuilder:validation:Minimum=1 + // +optional + VCPUCount *int32 `json:"vcpuCount,omitempty"` + + // MinHugePagesSize is the hugepage memory each storage node takes, as a + // quantity ("8G"). + // +optional + MinHugePagesSize string `json:"minHugePagesSize,omitempty"` + + // MaxSubsystemCount is the number of NVMe-oF subsystems each node serves. + // +kubebuilder:validation:Minimum=1 + // +optional + MaxSubsystemCount *int32 `json:"maxSubsystemCount,omitempty"` + + // EnableDriveFormat lets the deployment format the devices it takes. + // +optional + EnableDriveFormat *bool `json:"enableDriveFormat,omitempty"` + + // EnableJournalDevice dedicates one device per node to the journal. + // +optional + EnableJournalDevice *bool `json:"enableJournalDevice,omitempty"` + + // Stripe is the erasure-coding layout. + // +optional + Stripe *StripeSpec `json:"stripe,omitempty"` +} + +// StorageSiteDeploymentSpec is the request for one site's storage cluster. +// +kubebuilder:validation:XValidation:rule="!has(oldSelf.approved) || !oldSelf.approved || self.approved",message="approval is one-way: an approved deployment cannot be un-approved" +type StorageSiteDeploymentSpec struct { + // Cluster is the OCM ManagedCluster the storage is deployed on. The request's + // ManifestWork and views live in its namespace on the hub. Immutable. + // +kubebuilder:validation:Required + // +kubebuilder:validation:MinLength=1 + // +kubebuilder:validation:MaxLength=63 + // +kubebuilder:validation:XValidation:rule="self == oldSelf",message="cluster is immutable" + Cluster string `json:"cluster"` + + // SiteNamespace is the simplyblock operator's namespace on the site, where + // the discovery and the draft live. + // +kubebuilder:default=simplyblock + // +kubebuilder:validation:MaxLength=63 + // +optional + SiteNamespace string `json:"siteNamespace,omitempty"` + + // DraftName is the ClusterDeploymentConfig the discovery writes on the site + // and the request sizes and approves. Immutable. + // +kubebuilder:default=site-draft + // +kubebuilder:validation:MaxLength=63 + // +kubebuilder:validation:XValidation:rule="self == oldSelf",message="draftName is immutable" + // +optional + DraftName string `json:"draftName,omitempty"` + + // Discover is the discovery the site runs first. Changing it runs another + // discovery, which rewrites the draft. + // +optional + Discover StorageSiteDiscovery `json:"discover,omitempty"` + + // Sizing is written onto the draft's cluster template once the draft exists, + // so the reviewer sees the sized draft before approving it. + // +optional + Sizing *StorageSiteSizing `json:"sizing,omitempty"` + + // Approved is the review gate, delivered to the draft on the site. One-way, + // as the draft's own gate is. + // +kubebuilder:default=false + // +optional + Approved bool `json:"approved"` +} + +// StorageSiteDraft is the draft as the site reports it. +type StorageSiteDraft struct { + // Name is the ClusterDeploymentConfig on the site. + Name string `json:"name"` + + // Phase is the draft's own phase on the site (Draft, Expanding, Expanded, + // Failed). + // +optional + Phase string `json:"phase,omitempty"` + + // Message is what the site says about the draft: validation findings while + // it is a draft, the expansion's step afterwards. + // +optional + Message string `json:"message,omitempty"` + + // Approved is whether the draft is approved on the site. + // +optional + Approved bool `json:"approved,omitempty"` + + // Cluster is the draft's cluster template, with the sizing applied. + // +optional + Cluster *ClusterTemplate `json:"cluster,omitempty"` + + // NodeSets are the nodes and devices the discovery found, for review. + // +optional + NodeSets []NodeSet `json:"nodeSets,omitempty"` + + // NodeRefs are the StorageNode objects the expansion created. + // +optional + // +listType=set + NodeRefs []string `json:"nodeRefs,omitempty"` +} + +// StorageSiteNode is one storage node of the deployed cluster, as the site +// reports it. +type StorageSiteNode struct { + // Name is the StorageNode object on the site. + Name string `json:"name"` + + // Phase is the node's phase on the site. + // +optional + Phase string `json:"phase,omitempty"` + + // Hostname is the Kubernetes node it runs on. + // +optional + Hostname string `json:"hostname,omitempty"` +} + +// StorageSiteCluster is the StorageCluster the approved draft produced. +type StorageSiteCluster struct { + // Name is the StorageCluster object on the site. + Name string `json:"name"` + + // UUID is the storage cluster's id in the control plane, which a + // StorageClass names in cluster_id. + // +optional + UUID string `json:"uuid,omitempty"` + + // Phase is the StorageCluster's phase on the site. + // +optional + Phase string `json:"phase,omitempty"` + + // Pool is the pool the cluster was created with, which a StorageClass names + // in pool_name. + // +optional + Pool string `json:"pool,omitempty"` + + // Nodes are the cluster's storage nodes. + // +optional + // +listType=map + // +listMapKey=name + Nodes []StorageSiteNode `json:"nodes,omitempty"` +} + +// StorageSiteDeploymentStatus is what the site reports back, projected. +type StorageSiteDeploymentStatus struct { + // Phase is the request's own progress. + // +optional + Phase StorageSiteDeploymentPhase `json:"phase,omitempty"` + + // Message is the reason the phase is what it is: one sentence, replaced as + // the request moves, and never a log. + // +optional + Message string `json:"message,omitempty"` + + // ObservedGeneration is the generation the rest of this status was computed + // from. + // +optional + ObservedGeneration int64 `json:"observedGeneration,omitempty"` + + // WorkName is the ManifestWork carrying the request to the site. + // +optional + WorkName string `json:"workName,omitempty"` + + // Draft is the draft as the site reports it. + // +optional + Draft *StorageSiteDraft `json:"draft,omitempty"` + + // StorageCluster is the cluster the approved draft produced. + // +optional + StorageCluster *StorageSiteCluster `json:"storageCluster,omitempty"` + + // Conditions: Delivered (the work is applied on the site), Discovered (the + // draft names nodes), Approved (the site's draft is approved), Ready (the + // StorageCluster is Online). + // +optional + // +listType=map + // +listMapKey=type + Conditions []metav1.Condition `json:"conditions,omitempty"` +} + +// +kubebuilder:object:root=true +// +kubebuilder:subresource:status +// +kubebuilder:resource:scope=Namespaced,shortName=sbsd +// +kubebuilder:printcolumn:name="Cluster",type=string,JSONPath=".spec.cluster" +// +kubebuilder:printcolumn:name="Approved",type=boolean,JSONPath=".spec.approved" +// +kubebuilder:printcolumn:name="Phase",type=string,JSONPath=".status.phase" +// +kubebuilder:printcolumn:name="Draft",type=string,JSONPath=".status.draft.phase" +// +kubebuilder:printcolumn:name="Storage",type=string,JSONPath=".status.storageCluster.phase" +// +kubebuilder:printcolumn:name="Message",type=string,JSONPath=".status.message",priority=1 +// +kubebuilder:printcolumn:name="Age",type=date,JSONPath=".metadata.creationTimestamp" + +// StorageSiteDeployment requests a managed site's storage cluster from the hub: +// a discovery on the site, the sizing of the draft it writes, and the approval +// that expands the draft into a StorageCluster. The hub carries the request +// through OCM and projects the site's draft and cluster into the status. +// Deleting the request leaves the storage cluster alone. +type StorageSiteDeployment struct { + metav1.TypeMeta `json:",inline"` + metav1.ObjectMeta `json:"metadata,omitempty"` + + Spec StorageSiteDeploymentSpec `json:"spec,omitempty"` + Status StorageSiteDeploymentStatus `json:"status,omitempty"` +} + +// +kubebuilder:object:root=true + +// StorageSiteDeploymentList contains a list of StorageSiteDeployment. +type StorageSiteDeploymentList struct { + metav1.TypeMeta `json:",inline"` + metav1.ListMeta `json:"metadata,omitempty"` + Items []StorageSiteDeployment `json:"items"` +} + +func init() { + SchemeBuilder.Register(&StorageSiteDeployment{}, &StorageSiteDeploymentList{}) +} diff --git a/operator/api/v1alpha2/zz_generated.deepcopy.go b/operator/api/v1alpha2/zz_generated.deepcopy.go index 4913a6e6f..9dc34a411 100644 --- a/operator/api/v1alpha2/zz_generated.deepcopy.go +++ b/operator/api/v1alpha2/zz_generated.deepcopy.go @@ -3566,6 +3566,257 @@ func (in *StoragePoolStatus) DeepCopy() *StoragePoolStatus { return out } +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *StorageSiteCluster) DeepCopyInto(out *StorageSiteCluster) { + *out = *in + if in.Nodes != nil { + in, out := &in.Nodes, &out.Nodes + *out = make([]StorageSiteNode, len(*in)) + copy(*out, *in) + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new StorageSiteCluster. +func (in *StorageSiteCluster) DeepCopy() *StorageSiteCluster { + if in == nil { + return nil + } + out := new(StorageSiteCluster) + in.DeepCopyInto(out) + return out +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *StorageSiteDeployment) DeepCopyInto(out *StorageSiteDeployment) { + *out = *in + out.TypeMeta = in.TypeMeta + in.ObjectMeta.DeepCopyInto(&out.ObjectMeta) + in.Spec.DeepCopyInto(&out.Spec) + in.Status.DeepCopyInto(&out.Status) +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new StorageSiteDeployment. +func (in *StorageSiteDeployment) DeepCopy() *StorageSiteDeployment { + if in == nil { + return nil + } + out := new(StorageSiteDeployment) + in.DeepCopyInto(out) + return out +} + +// DeepCopyObject is an autogenerated deepcopy function, copying the receiver, creating a new runtime.Object. +func (in *StorageSiteDeployment) DeepCopyObject() runtime.Object { + if c := in.DeepCopy(); c != nil { + return c + } + return nil +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *StorageSiteDeploymentList) DeepCopyInto(out *StorageSiteDeploymentList) { + *out = *in + out.TypeMeta = in.TypeMeta + in.ListMeta.DeepCopyInto(&out.ListMeta) + if in.Items != nil { + in, out := &in.Items, &out.Items + *out = make([]StorageSiteDeployment, len(*in)) + for i := range *in { + (*in)[i].DeepCopyInto(&(*out)[i]) + } + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new StorageSiteDeploymentList. +func (in *StorageSiteDeploymentList) DeepCopy() *StorageSiteDeploymentList { + if in == nil { + return nil + } + out := new(StorageSiteDeploymentList) + in.DeepCopyInto(out) + return out +} + +// DeepCopyObject is an autogenerated deepcopy function, copying the receiver, creating a new runtime.Object. +func (in *StorageSiteDeploymentList) DeepCopyObject() runtime.Object { + if c := in.DeepCopy(); c != nil { + return c + } + return nil +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *StorageSiteDeploymentSpec) DeepCopyInto(out *StorageSiteDeploymentSpec) { + *out = *in + in.Discover.DeepCopyInto(&out.Discover) + if in.Sizing != nil { + in, out := &in.Sizing, &out.Sizing + *out = new(StorageSiteSizing) + (*in).DeepCopyInto(*out) + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new StorageSiteDeploymentSpec. +func (in *StorageSiteDeploymentSpec) DeepCopy() *StorageSiteDeploymentSpec { + if in == nil { + return nil + } + out := new(StorageSiteDeploymentSpec) + in.DeepCopyInto(out) + return out +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *StorageSiteDeploymentStatus) DeepCopyInto(out *StorageSiteDeploymentStatus) { + *out = *in + if in.Draft != nil { + in, out := &in.Draft, &out.Draft + *out = new(StorageSiteDraft) + (*in).DeepCopyInto(*out) + } + if in.StorageCluster != nil { + in, out := &in.StorageCluster, &out.StorageCluster + *out = new(StorageSiteCluster) + (*in).DeepCopyInto(*out) + } + if in.Conditions != nil { + in, out := &in.Conditions, &out.Conditions + *out = make([]metav1.Condition, len(*in)) + for i := range *in { + (*in)[i].DeepCopyInto(&(*out)[i]) + } + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new StorageSiteDeploymentStatus. +func (in *StorageSiteDeploymentStatus) DeepCopy() *StorageSiteDeploymentStatus { + if in == nil { + return nil + } + out := new(StorageSiteDeploymentStatus) + in.DeepCopyInto(out) + return out +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *StorageSiteDiscovery) DeepCopyInto(out *StorageSiteDiscovery) { + *out = *in + if in.EnableControlPlaneNodes != nil { + in, out := &in.EnableControlPlaneNodes, &out.EnableControlPlaneNodes + *out = new(bool) + **out = **in + } + if in.Workers != nil { + in, out := &in.Workers, &out.Workers + *out = make([]string, len(*in)) + copy(*out, *in) + } + if in.NodeSelector != nil { + in, out := &in.NodeSelector, &out.NodeSelector + *out = make(map[string]string, len(*in)) + for key, val := range *in { + (*out)[key] = val + } + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new StorageSiteDiscovery. +func (in *StorageSiteDiscovery) DeepCopy() *StorageSiteDiscovery { + if in == nil { + return nil + } + out := new(StorageSiteDiscovery) + in.DeepCopyInto(out) + return out +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *StorageSiteDraft) DeepCopyInto(out *StorageSiteDraft) { + *out = *in + if in.Cluster != nil { + in, out := &in.Cluster, &out.Cluster + *out = new(ClusterTemplate) + (*in).DeepCopyInto(*out) + } + if in.NodeSets != nil { + in, out := &in.NodeSets, &out.NodeSets + *out = make([]NodeSet, len(*in)) + for i := range *in { + (*in)[i].DeepCopyInto(&(*out)[i]) + } + } + if in.NodeRefs != nil { + in, out := &in.NodeRefs, &out.NodeRefs + *out = make([]string, len(*in)) + copy(*out, *in) + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new StorageSiteDraft. +func (in *StorageSiteDraft) DeepCopy() *StorageSiteDraft { + if in == nil { + return nil + } + out := new(StorageSiteDraft) + in.DeepCopyInto(out) + return out +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *StorageSiteNode) DeepCopyInto(out *StorageSiteNode) { + *out = *in +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new StorageSiteNode. +func (in *StorageSiteNode) DeepCopy() *StorageSiteNode { + if in == nil { + return nil + } + out := new(StorageSiteNode) + in.DeepCopyInto(out) + return out +} + +// DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. +func (in *StorageSiteSizing) DeepCopyInto(out *StorageSiteSizing) { + *out = *in + if in.VCPUCount != nil { + in, out := &in.VCPUCount, &out.VCPUCount + *out = new(int32) + **out = **in + } + if in.MaxSubsystemCount != nil { + in, out := &in.MaxSubsystemCount, &out.MaxSubsystemCount + *out = new(int32) + **out = **in + } + if in.EnableDriveFormat != nil { + in, out := &in.EnableDriveFormat, &out.EnableDriveFormat + *out = new(bool) + **out = **in + } + if in.EnableJournalDevice != nil { + in, out := &in.EnableJournalDevice, &out.EnableJournalDevice + *out = new(bool) + **out = **in + } + if in.Stripe != nil { + in, out := &in.Stripe, &out.Stripe + *out = new(StripeSpec) + (*in).DeepCopyInto(*out) + } +} + +// DeepCopy is an autogenerated deepcopy function, copying the receiver, creating a new StorageSiteSizing. +func (in *StorageSiteSizing) DeepCopy() *StorageSiteSizing { + if in == nil { + return nil + } + out := new(StorageSiteSizing) + in.DeepCopyInto(out) + return out +} + // DeepCopyInto is an autogenerated deepcopy function, copying the receiver, writing into out. in must be non-nil. func (in *StripeSpec) DeepCopyInto(out *StripeSpec) { *out = *in diff --git a/operator/cmd/main.go b/operator/cmd/main.go index b3258ecd9..42b37093a 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -962,8 +962,18 @@ func main() { setupLog.Error(err, "unable to create controller", "controller", "TestFailover") os.Exit(1) } + // A managed site's storage deployment is requested from the hub through + // the same work API, so the controller is hub-only too. + if err := (&controller.StorageSiteDeploymentReconciler{ + Client: mgr.GetClient(), + Scheme: mgr.GetScheme(), + Recorder: mgr.GetEventRecorder("storagesitedeployment-controller"), + }).SetupWithManager(mgr); err != nil { + setupLog.Error(err, "unable to create controller", "controller", "StorageSiteDeployment") + os.Exit(1) + } } else { - setupLog.Info("OCM ManifestWork resource not served; skipping TestFailover controller (hub-only)", + setupLog.Info("OCM ManifestWork resource not served; skipping the TestFailover and StorageSiteDeployment controllers (hub-only)", "groupVersion", ocmWorkGroupVersion, "resource", ocmManifestWorkResource) } // +kubebuilder:scaffold:builder diff --git a/operator/config/crd/bases/storage.simplyblock.io_storagesitedeployments.yaml b/operator/config/crd/bases/storage.simplyblock.io_storagesitedeployments.yaml new file mode 100644 index 000000000..9c2718a8f --- /dev/null +++ b/operator/config/crd/bases/storage.simplyblock.io_storagesitedeployments.yaml @@ -0,0 +1,1057 @@ +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition +metadata: + annotations: + controller-gen.kubebuilder.io/version: v0.21.0 + name: storagesitedeployments.storage.simplyblock.io +spec: + group: storage.simplyblock.io + names: + kind: StorageSiteDeployment + listKind: StorageSiteDeploymentList + plural: storagesitedeployments + shortNames: + - sbsd + singular: storagesitedeployment + scope: Namespaced + versions: + - additionalPrinterColumns: + - jsonPath: .spec.cluster + name: Cluster + type: string + - jsonPath: .spec.approved + name: Approved + type: boolean + - jsonPath: .status.phase + name: Phase + type: string + - jsonPath: .status.draft.phase + name: Draft + type: string + - jsonPath: .status.storageCluster.phase + name: Storage + type: string + - jsonPath: .status.message + name: Message + priority: 1 + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha2 + schema: + openAPIV3Schema: + description: |- + StorageSiteDeployment requests a managed site's storage cluster from the hub: + a discovery on the site, the sizing of the draft it writes, and the approval + that expands the draft into a StorageCluster. The hub carries the request + through OCM and projects the site's draft and cluster into the status. + Deleting the request leaves the storage cluster alone. + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: StorageSiteDeploymentSpec is the request for one site's storage + cluster. + properties: + approved: + default: false + description: |- + Approved is the review gate, delivered to the draft on the site. One-way, + as the draft's own gate is. + type: boolean + cluster: + description: |- + Cluster is the OCM ManagedCluster the storage is deployed on. The request's + ManifestWork and views live in its namespace on the hub. Immutable. + maxLength: 63 + minLength: 1 + type: string + x-kubernetes-validations: + - message: cluster is immutable + rule: self == oldSelf + discover: + description: |- + Discover is the discovery the site runs first. Changing it runs another + discovery, which rewrites the draft. + properties: + enableControlPlaneNodes: + description: |- + EnableControlPlaneNodes lets the discovery consider the nodes that run the + API server. Every server of a small distribution is one, so a three-node + site has no storage without it. + type: boolean + nodeSelector: + additionalProperties: + type: string + description: NodeSelector limits the discovery to the nodes carrying + these labels. + type: object + workers: + description: Workers limits the discovery to these nodes. Empty + is every worker. + items: + type: string + type: array + x-kubernetes-list-type: set + type: object + draftName: + default: site-draft + description: |- + DraftName is the ClusterDeploymentConfig the discovery writes on the site + and the request sizes and approves. Immutable. + maxLength: 63 + type: string + x-kubernetes-validations: + - message: draftName is immutable + rule: self == oldSelf + siteNamespace: + default: simplyblock + description: |- + SiteNamespace is the simplyblock operator's namespace on the site, where + the discovery and the draft live. + maxLength: 63 + type: string + sizing: + description: |- + Sizing is written onto the draft's cluster template once the draft exists, + so the reviewer sees the sized draft before approving it. + properties: + enableDriveFormat: + description: EnableDriveFormat lets the deployment format the + devices it takes. + type: boolean + enableJournalDevice: + description: EnableJournalDevice dedicates one device per node + to the journal. + type: boolean + maxSubsystemCount: + description: MaxSubsystemCount is the number of NVMe-oF subsystems + each node serves. + format: int32 + minimum: 1 + type: integer + minHugePagesSize: + description: |- + MinHugePagesSize is the hugepage memory each storage node takes, as a + quantity ("8G"). + type: string + name: + description: Name is the StorageCluster's name on the site. + maxLength: 63 + type: string + stripe: + description: Stripe is the erasure-coding layout. + properties: + dataChunks: + description: DataChunks is the number of data chunks per stripe + (ndcs). + format: int32 + minimum: 1 + type: integer + parityChunks: + description: |- + ParityChunks is the number of parity chunks per stripe (npcs), and + therefore how many chunk losses a stripe survives. + format: int32 + minimum: 0 + type: integer + type: object + x-kubernetes-validations: + - message: the erasure-coding scheme must be one of 1+0, 1+1, + 2+1, 4+1, 1+2, 2+2, or 4+2, written as dataChunks+parityChunks, + and an unstated half is 1 + rule: '[has(self.dataChunks) ? self.dataChunks : 1, has(self.parityChunks) + ? self.parityChunks : 1] in [[1, 0], [1, 1], [2, 1], [4, 1], + [1, 2], [2, 2], [4, 2]]' + vcpuCount: + description: VCPUCount is the number of vCPUs each storage node + takes. + format: int32 + minimum: 1 + type: integer + type: object + required: + - cluster + type: object + x-kubernetes-validations: + - message: 'approval is one-way: an approved deployment cannot be un-approved' + rule: '!has(oldSelf.approved) || !oldSelf.approved || self.approved' + status: + description: StorageSiteDeploymentStatus is what the site reports back, + projected. + properties: + conditions: + description: |- + Conditions: Delivered (the work is applied on the site), Discovered (the + draft names nodes), Approved (the site's draft is approved), Ready (the + StorageCluster is Online). + items: + description: Condition contains details for one aspect of the current + state of this API Resource. + properties: + lastTransitionTime: + description: |- + lastTransitionTime is the last time the condition transitioned from one status to another. + This should be when the underlying condition changed. If that is not known, then using the time when the API field changed is acceptable. + format: date-time + type: string + message: + description: |- + message is a human readable message indicating details about the transition. + This may be an empty string. + maxLength: 32768 + type: string + observedGeneration: + description: |- + observedGeneration represents the .metadata.generation that the condition was set based upon. + For instance, if .metadata.generation is currently 12, but the .status.conditions[x].observedGeneration is 9, the condition is out of date + with respect to the current state of the instance. + format: int64 + minimum: 0 + type: integer + reason: + description: |- + reason contains a programmatic identifier indicating the reason for the condition's last transition. + Producers of specific condition types may define expected values and meanings for this field, + and whether the values are considered a guaranteed API. + The value should be a CamelCase string. + This field may not be empty. + maxLength: 1024 + minLength: 1 + pattern: ^[A-Za-z]([A-Za-z0-9_,:]*[A-Za-z0-9_])?$ + type: string + status: + description: status of the condition, one of True, False, Unknown. + enum: + - "True" + - "False" + - Unknown + type: string + type: + description: type of condition in CamelCase or in foo.example.com/CamelCase. + maxLength: 316 + pattern: ^([a-z0-9]([-a-z0-9]*[a-z0-9])?(\.[a-z0-9]([-a-z0-9]*[a-z0-9])?)*/)?(([A-Za-z0-9][-A-Za-z0-9_.]*)?[A-Za-z0-9])$ + type: string + required: + - lastTransitionTime + - message + - reason + - status + - type + type: object + type: array + x-kubernetes-list-map-keys: + - type + x-kubernetes-list-type: map + draft: + description: Draft is the draft as the site reports it. + properties: + approved: + description: Approved is whether the draft is approved on the + site. + type: boolean + cluster: + description: Cluster is the draft's cluster template, with the + sizing applied. + properties: + backup: + description: |- + Backup is where this cluster's backups live, and it expands into + StorageCluster.spec.backup unchanged. + + It is here for the reason KMS is: a store stated on the document is + present when the cluster is created rather than patched in afterward by + whoever remembers. Unlike most of what this template carries, the field it + fills is mutable, so a document that states none costs nothing permanent. + A cluster can be given a store whenever there is one to give. + + The Secret it names is not resolved at admission. It is a core object a + deployment legitimately creates alongside the document or after it, and + the cluster's own creation is where its absence is reported. + properties: + bucket: + description: Bucket is the bucket backups are written + to and read from. + type: string + credentialsSecretRef: + description: |- + CredentialsSecretRef names the Secret holding the access key and the + secret key. It is a reference rather than the values, because a spec is + readable by anybody who can read the object. + properties: + name: + default: "" + description: |- + Name of the referent. + This field is effectively required, but due to backwards compatibility is + allowed to be empty. Instances of this type with an empty value here are + almost certainly wrong. + More info: https://kubernetes.io/docs/concepts/overview/working-with-objects/names/#names + type: string + type: object + x-kubernetes-map-type: atomic + endpoint: + description: Endpoint is the S3 endpoint, for example, + https://s3.example.com. + pattern: ^https?://[a-zA-Z0-9.-]+(:[0-9]{1,5})?(/.*)?$ + type: string + prefix: + description: |- + Prefix narrows the store to one key prefix, so that several clusters can + share a bucket without each walking the others' backups. + type: string + region: + description: Region is the bucket's region, for endpoints + that do not imply one. + type: string + required: + - bucket + - credentialsSecretRef + - endpoint + type: object + containerResources: + description: |- + ContainerResources sizes the storage-node container, and expands into the + cluster's own spec.storageNodes.containerResources. + + The container it sizes is the node's management API rather than SPDK, + which runs in a pod of its own: what outgrows the default is a node + answering for many subsystems, not a node moving more data. It is on the + document because a deployment is where a fleet's sizing is decided, and + a cluster written from a document that could not say so had to be edited + afterward on a field the document owns everywhere else. + + Stating either half replaces both. The defaults apply to a cluster that + states neither requests nor limits, so a document stating requests alone + produces a container with no limits rather than one with the default + limits, and a memory limit is what has the kubelet evict a leaking agent + rather than losing the worker. + + It is a pointer because a resource block is a struct, and a struct with + omitempty is serialized whether or not anything is in it: as a value, + every document a discovery run writes would carry an empty + containerResources that says nothing and that a reviewer has to decide + about. + properties: + claims: + description: |- + Claims lists the names of resources, defined in spec.resourceClaims, + that are used by this container. + + This field depends on the + DynamicResourceAllocation feature gate. + + This field is immutable. It can only be set for containers. + items: + description: ResourceClaim references one entry in PodSpec.ResourceClaims. + properties: + name: + description: |- + Name must match the name of one entry in pod.spec.resourceClaims of + the Pod where this field is used. It makes that resource available + inside a container. + type: string + request: + description: |- + Request is the name chosen for a request in the referenced claim. + If empty, everything from the claim is made available, otherwise + only the result of this request. + type: string + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + limits: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Limits describes the maximum amount of compute resources allowed. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + requests: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Requests describes the minimum amount of compute resources required. + If Requests is omitted for a container, it defaults to Limits if that is explicitly specified, + otherwise to an implementation-defined value. Requests cannot exceed Limits. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + type: object + enableAtomicity4K: + description: |- + EnableAtomicity4K enforces 4K write atomicity on every device this + deployment names, which is what lets checksum validation run on devices + whose logical block size is under the data plane's 4K minimum. + + It is the route to checked I/O on a device that cannot be reformatted: a + logical block device's block size is fixed by the drive, and some NVMe + devices offer no 4K format either. Where a device can be reformatted, + EnableDriveFormat is the other route and this is unnecessary. + + It is an enforcement because the question is often unanswerable. A SATA + drive presenting 512-byte logical blocks over a 4K physical sector reports + 512 and nothing more, and a kernel older than 6.11 publishes no atomic + write attributes at all. Where a device does answer, the storage node's + report carries it, and a reviewer approves this against that rather than + against a vendor's datasheet -- because enforcing a guarantee the hardware + does not keep is how a torn write becomes a checksum that silently + disagrees with it. + + It means nothing unless EnableChecksumValidation is set, which is the + cluster's own rule and is left to the cluster to enforce. + type: boolean + enableChecksumValidation: + description: |- + EnableChecksumValidation turns on inline CRC validation of every I/O, for + silent-data-error protection. + + It is on the document because it is immutable on the cluster it lands on: + the backend bakes the checksum method into each device when the cluster is + created and never re-applies it, so a cluster created without this is one + nobody can turn it on for. A deployment that wants its data checked has to + say so here or not at all. + type: boolean + enableDriveFormat: + description: |- + EnableDriveFormat formats every device the document names before a storage + node takes it, which is how a drive carrying anything already is made + usable. + + It says what is wanted rather than how, because the how differs by device + class: an NVMe device is formatted to a 4K block size, and a logical block + device has its signatures wiped. One field covers both, so a document does + not have to know which class the expansion will resolve it to. + + It is on the document rather than defaulted further down because it is + destructive and the document is what somebody approves. A reviewer reading + a draft has to see that the drives it lists will be formatted, and be able + to strike it before approving; the cluster's own field is immutable once + the cluster exists, so a default nobody saw could not be undone either. + type: boolean + enableFailureDomains: + description: |- + EnableFailureDomains opts the cluster into failure-domain mode, in which + every group must label the fault group its workers belong to. + type: boolean + enableJournalDevice: + description: |- + EnableJournalDevice dedicates the smallest NVMe device on each of this + deployment's workers to the journal manager, instead of carving a journal + partition out of every device. + + It is here rather than on a node set because it is immutable on the cluster + it lands on, for the reason SocketsToUse is: the on-disk layout a fleet was + built with is not one a later document can vary. It also costs a drive of + capacity per node, which is a trade a reviewer approves rather than one a + default makes for them. + type: boolean + enableNodeAffinity: + description: |- + EnableNodeAffinity has the data plane serve an erasure-coded volume's I/O + from the local node's own devices where it can, before crossing the + network. + + It is not Kubernetes affinity, and the name is the one place this API + invites that reading: nothing about it schedules a pod, labels a worker, + or places a volume's primary node. The control plane carries it into the + cluster map it pushes to each node, where it sets the local node's index, + and what changes is which copy of a chunk is read. + Co-locating a workload with the primary node of its volume is a separate + mechanism and is not configured here. + + It is on the document because it is immutable on the cluster: the control + plane takes it at cluster create and never re-applies it, so this is the + only moment it can be set at all. + type: boolean + fabricType: + description: FabricType is the storage fabric. + maxLength: 32 + type: string + initContainerResources: + description: |- + InitContainerResources sizes both of the storage node's init containers, + and expands into the cluster's own spec.storageNodes.initContainerResources. + + They are sized apart from the container because they do a different job + and are gone before it starts: one writes the node's env file and the + other runs node_configure.py once, so what they need is a short burst + rather than the footprint of a process that runs for the node's life. + + Stating either half replaces both, as with containerResources, and it is + a pointer for the same reason. + properties: + claims: + description: |- + Claims lists the names of resources, defined in spec.resourceClaims, + that are used by this container. + + This field depends on the + DynamicResourceAllocation feature gate. + + This field is immutable. It can only be set for containers. + items: + description: ResourceClaim references one entry in PodSpec.ResourceClaims. + properties: + name: + description: |- + Name must match the name of one entry in pod.spec.resourceClaims of + the Pod where this field is used. It makes that resource available + inside a container. + type: string + request: + description: |- + Request is the name chosen for a request in the referenced claim. + If empty, everything from the claim is made available, otherwise + only the result of this request. + type: string + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + limits: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Limits describes the maximum amount of compute resources allowed. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + requests: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Requests describes the minimum amount of compute resources required. + If Requests is omitted for a container, it defaults to Limits if that is explicitly specified, + otherwise to an implementation-defined value. Requests cannot exceed Limits. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + type: object + kms: + description: |- + KMS selects where the cluster stores volume encryption keys. Stating it on + the document is what makes it present when the cluster is created, where + setting it on the StorageCluster afterward races with that creation. + properties: + vault: + description: Vault stores keys in HashiCorp Vault. + properties: + endpoint: + description: |- + Endpoint is the Vault endpoint, for example, https://vault.example.com:8200. + Rejected unless it resolves to an external address. + pattern: ^https?://[a-zA-Z0-9.-]+(:[0-9]{1,5})?(/.*)?$ + type: string + required: + - endpoint + type: object + type: object + maxSubsystemCount: + description: |- + MaxSubsystemCount is the maximum number of NVMe-oF subsystems each storage + node of this cluster serves. Required, because the StorageCluster's own + field is, and no StorageNode carries a copy of it. + format: int32 + maximum: 75 + minimum: 10 + type: integer + minHugePagesSize: + description: |- + MinHugePagesSize is the smallest huge-page allocation each storage node of + this cluster makes: 100G or 1T, where a bare number is gigabytes. Like + VCPUCount it is the cluster's and is copied onto every node the expansion + writes. Omitted, each node uses the computed minimum. + maxLength: 32 + type: string + name: + description: |- + Name is the StorageCluster's name, and is therefore held to what such a + name may be rather than to what an object name may be. A longer value is a + document the API server accepts and a CreatingCluster step that can never + succeed, since the cluster it would write is one the API server refuses. + maxLength: 63 + type: string + nodeProvisioningBudget: + description: |- + NodeProvisioningBudget is how many workers the expansion may have in the + node-add process at once. It expands into the cluster's own + spec.storageNodes.nodeProvisioningBudget, whose meaning it shares: the cap + is counted by distinct worker, so a two-socket host spends one of the + budget, and a worker hosting a FoundationDB pod is sequential whatever the + budget says. + + It is on the document because a document is what states the size of a + deployment, and a deployment of thirty workers added one at a time is the + difference between an afternoon and a week. Omitted, the cluster's default + of one applies, which is the serial behavior. + format: int32 + minimum: 1 + type: integer + nodesPerSocket: + description: |- + NodesPerSocket is how many storage nodes run per NUMA socket. See + SocketsToUse, which it multiplies. + format: int32 + maximum: 8 + minimum: 1 + type: integer + openshift: + description: |- + OpenShift is what this deployment states because it runs on OpenShift. It + expands into StorageCluster.spec.storageNodes.openshift, whose shape it + shares, and it is read only for a document whose environment is + OpenShift: the environment is what says which distribution this is, and + the block is what that distribution needs said beyond it. + properties: + machineConfigPool: + default: worker + description: |- + MachineConfigPool names a machine-config role the storage nodes' own pool + inherits from, beyond the worker role it always inherits. + + It is not the pool the nodes end up in, which the description it carried + before said and which cost a reader the reboot they were trying to avoid. + Adding a node creates a pool of its own, storage-, and moves the + node into it; a node belongs to exactly one custom pool, so whatever + machine configuration its previous pool carried is lost unless that + pool's role is named here for the new one to select as well. The default + is the role every pool already selects, which is what makes it a no-op + for a fleet whose workers are ordinary workers. + maxLength: 253 + pattern: ^[a-z0-9]([-a-z0-9]*[a-z0-9])?$ + type: string + type: object + ports: + description: |- + Ports are where this cluster's storage nodes listen. Unstated, and for + each member left unstated, the cluster's own defaults decide. + properties: + nodeAgent: + default: 50001 + description: |- + NodeAgent is the port each node's agent API listens on. It expands into + StorageCluster.spec.snodeApiPort, and it is named for the component + rather than for that field: the agent is what spec.images.nodeAgent pins + and what the storage-node DaemonSet runs. + format: int32 + maximum: 65535 + minimum: 1024 + type: integer + nvmf: + default: 4420 + description: |- + NVMf is the base of the NVMe-oF port range every node binds. It expands + into StorageCluster.spec.nvmfBasePort. + format: int32 + maximum: 65535 + minimum: 1024 + type: integer + rpc: + default: 8080 + description: |- + Rpc is the base of the RPC port range every node binds. It expands into + StorageCluster.spec.rpcBasePort. + format: int32 + maximum: 65535 + minimum: 1024 + type: integer + type: object + socketsToUse: + description: |- + SocketsToUse restricts the deployment to selected NUMA sockets, and empty + means socket 0 alone. With NodesPerSocket it decides how many storage nodes + each worker runs, so a group of two workers on a two-socket layout expands + to four nodes. + + It is here rather than on a node set because it is immutable on the cluster + it lands on: the layout a fleet was built with is not one a later document + can vary, and a reviewer should see it before the cluster exists. + items: + maxLength: 16 + type: string + maxItems: 16 + type: array + x-kubernetes-list-type: set + stripe: + description: Stripe is the erasure-coding layout. + properties: + dataChunks: + description: DataChunks is the number of data chunks per + stripe (ndcs). + format: int32 + minimum: 1 + type: integer + parityChunks: + description: |- + ParityChunks is the number of parity chunks per stripe (npcs), and + therefore how many chunk losses a stripe survives. + format: int32 + minimum: 0 + type: integer + type: object + x-kubernetes-validations: + - message: the erasure-coding scheme must be one of 1+0, 1+1, + 2+1, 4+1, 1+2, 2+2, or 4+2, written as dataChunks+parityChunks, + and an unstated half is 1 + rule: '[has(self.dataChunks) ? self.dataChunks : 1, has(self.parityChunks) + ? self.parityChunks : 1] in [[1, 0], [1, 1], [2, 1], [4, + 1], [1, 2], [2, 2], [4, 2]]' + tolerations: + description: |- + Tolerations are what the storage-node pods tolerate, and they expand into + the cluster's own spec.storageNodes.tolerations. + + A fleet that dedicates machines to storage taints them, which is what + keeps everything else off. The DaemonSet that lands on those machines has + to tolerate the taint or it schedules nowhere, and a document that could + not say so described a deployment that does not start: the correction was + an edit to the cluster the document had just created, on a field the + document owns everywhere else. + + A growth document states none. It names a cluster rather than describing + one, and that cluster already carries what its storage nodes tolerate. + items: + description: |- + The pod this Toleration is attached to tolerates any taint that matches + the triple using the matching operator . + properties: + effect: + description: |- + Effect indicates the taint effect to match. Empty means match all taint effects. + When specified, allowed values are NoSchedule, PreferNoSchedule and NoExecute. + type: string + key: + description: |- + Key is the taint key that the toleration applies to. Empty means match all taint keys. + If the key is empty, operator must be Exists; this combination means to match all values and all keys. + type: string + operator: + description: |- + Operator represents a key's relationship to the value. + Valid operators are Exists, Equal, Lt, and Gt. Defaults to Equal. + Exists is equivalent to wildcard for value, so that a pod can + tolerate all taints of a particular category. + Lt and Gt perform numeric comparisons (requires feature gate TaintTolerationComparisonOperators). + type: string + tolerationSeconds: + description: |- + TolerationSeconds represents the period of time the toleration (which must be + of effect NoExecute, otherwise this field is ignored) tolerates the taint. By default, + it is not set, which means tolerate the taint forever (do not evict). Zero and + negative values will be treated as 0 (evict immediately) by the system. + format: int64 + type: integer + value: + description: |- + Value is the taint value the toleration matches to. + If the operator is Exists, the value should be empty, otherwise just a regular string. + type: string + type: object + maxItems: 32 + type: array + vcpuCount: + description: |- + VCPUCount is the number of vCPUs allocated to SPDK on each storage node of + this cluster. It is stated here and nowhere below, because the control + plane assumes it uniform across a cluster's nodes; CreatingNodes copies it + into every StorageNode.spec.config.sizing it writes. Required, because the + StorageCluster's own field is. + The floor is 4 rather than a hardware limit: a node must carry one core + beyond this budget for the system, and the control plane's core layout + assigns no NVMe-oF poller core at all for a 2-vCPU budget. + format: int32 + minimum: 4 + type: integer + required: + - maxSubsystemCount + - name + - vcpuCount + type: object + message: + description: |- + Message is what the site says about the draft: validation findings while + it is a draft, the expansion's step afterwards. + type: string + name: + description: Name is the ClusterDeploymentConfig on the site. + type: string + nodeRefs: + description: NodeRefs are the StorageNode objects the expansion + created. + items: + type: string + type: array + x-kubernetes-list-type: set + nodeSets: + description: NodeSets are the nodes and devices the discovery + found, for review. + items: + description: |- + NodeSet is the organizational grouping of a deployment, usually a rack: the + workers a document adds or grows together. It carries no sizing, because sizing + is uniform across a cluster and is stated once in ClusterTemplate. + properties: + groups: + description: Groups are the sets of workers sharing one + configuration. + items: + description: |- + NodeGroup is a set of workers that share one configuration, which is what + makes ten identical machines one entry rather than ten. + properties: + dataInterfaces: + description: DataInterfaces are the data-plane network + interfaces. + items: + maxLength: 63 + type: string + maxItems: 32 + type: array + devices: + description: Devices selects the storage devices every + worker in the group uses. + properties: + block: + description: |- + Block names logical block devices by path ("/dev/sdb"). It expands into the + same config.deviceNames as NVMe, which takes a PCI address and a device + path in one list. It is the alternative to NVMe rather than a companion of + it: the two classes are not mixed within a cluster. + items: + maxLength: 255 + pattern: ^/dev/[a-zA-Z0-9._/-]+$ + type: string + maxItems: 128 + type: array + x-kubernetes-list-type: set + nvme: + description: NVMe names NVMe devices by PCI address + ("0000:5e:00.0"). + items: + maxLength: 32 + pattern: ^[0-9a-fA-F]{4}:[0-9a-fA-F]{2}:[0-9a-fA-F]{2}\.[0-9a-fA-F]$ + type: string + maxItems: 128 + type: array + x-kubernetes-list-type: set + type: object + x-kubernetes-validations: + - message: a device selection names NVMe addresses + or block devices, not both + rule: has(self.nvme) != has(self.block) + failureDomain: + description: |- + FailureDomain is the label of the fault group every worker in this group + belongs to ("rack-b"), which is usually the name of the rack, zone, or + power feed they share. Discovery seeds it from topology.kubernetes.io/zone + and leaves it unset where the Kubernetes API carries no topology, which + holds provisioning with a clear reason rather than guessing. It expands + into StorageNode.spec.config.failureDomain, whose shape it shares. + maxLength: 63 + pattern: ^[a-zA-Z0-9]([-_.a-zA-Z0-9]*[a-zA-Z0-9])?$ + type: string + journalManager: + description: JournalManager tunes the journal managers + on these nodes. + properties: + count: + description: Count is the number of journal managers + to configure. + format: int32 + minimum: 1 + type: integer + percentPerDevice: + description: PercentPerDevice is the share of + each device given to the journal. + format: int32 + maximum: 100 + minimum: 1 + type: integer + type: object + mgmtInterface: + description: MgmtInterface is the management network + interface the storage nodes bind. + maxLength: 63 + type: string + name: + description: |- + Name identifies the group within its node set, for a reader and for the + events a validation failure emits. + maxLength: 253 + type: string + reservedSystemCPU: + description: |- + ReservedSystemCPU is the CPU set held back from SPDK for the system on + these nodes, as a core list such as 0,1 or 0-3. + + It is a group's rather than the cluster's because it names core ids, and a + group is what a document calls the workers that share their hardware: 0,1 + on a sixteen-core worker and 0,1 on a ninety-six-core worker are different + fractions of the machine. It expands into + StorageNode.spec.config.reservedSystemCPU, whose shape it shares, and a + group that states none leaves the cluster's fleet-wide value to decide. + + On OpenShift it reaches the kubelet through a KubeletConfig for the + machine config pool, which is the cluster's, so groups that disagree there + are writing over one another's pool configuration. + maxLength: 63 + pattern: ^[0-9]+(-[0-9]+)?(,[0-9]+(-[0-9]+)?)*$ + type: string + spdkSystemMemory: + description: |- + SpdkSystemMemory is the memory the control plane starts SPDK with on these + nodes. + maxLength: 32 + pattern: ^[0-9]+(G|GI|GB|GiB|M|MI|MB|MiB|g|gi|gb|gib|m|mi|mb|mib)?$ + type: string + workers: + description: Workers are the Kubernetes worker hostnames + in this group. + items: + maxLength: 253 + type: string + maxItems: 200 + minItems: 1 + type: array + x-kubernetes-list-type: set + required: + - name + - workers + type: object + maxItems: 64 + minItems: 1 + type: array + name: + description: |- + Name is the node set's name. It is copied to StorageNode.spec.nodeSet, so + that a node can be traced back to the part of the document that produced + it. + maxLength: 253 + type: string + required: + - groups + - name + type: object + type: array + phase: + description: |- + Phase is the draft's own phase on the site (Draft, Expanding, Expanded, + Failed). + type: string + required: + - name + type: object + message: + description: |- + Message is the reason the phase is what it is: one sentence, replaced as + the request moves, and never a log. + type: string + observedGeneration: + description: |- + ObservedGeneration is the generation the rest of this status was computed + from. + format: int64 + type: integer + phase: + description: Phase is the request's own progress. + enum: + - Pending + - Discovering + - Drafted + - Deploying + - Online + - Failed + type: string + storageCluster: + description: StorageCluster is the cluster the approved draft produced. + properties: + name: + description: Name is the StorageCluster object on the site. + type: string + nodes: + description: Nodes are the cluster's storage nodes. + items: + description: |- + StorageSiteNode is one storage node of the deployed cluster, as the site + reports it. + properties: + hostname: + description: Hostname is the Kubernetes node it runs on. + type: string + name: + description: Name is the StorageNode object on the site. + type: string + phase: + description: Phase is the node's phase on the site. + type: string + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + phase: + description: Phase is the StorageCluster's phase on the site. + type: string + pool: + description: |- + Pool is the pool the cluster was created with, which a StorageClass names + in pool_name. + type: string + uuid: + description: |- + UUID is the storage cluster's id in the control plane, which a + StorageClass names in cluster_id. + type: string + required: + - name + type: object + workName: + description: WorkName is the ManifestWork carrying the request to + the site. + type: string + type: object + type: object + served: true + storage: true + subresources: + status: {} diff --git a/operator/config/crd/kustomization.yaml b/operator/config/crd/kustomization.yaml index 760b4ffb6..520d8f4be 100644 --- a/operator/config/crd/kustomization.yaml +++ b/operator/config/crd/kustomization.yaml @@ -31,6 +31,7 @@ resources: - bases/storage.simplyblock.io_controlplaneops.yaml - bases/storage.simplyblock.io_storagedeviceops.yaml - bases/storage.simplyblock.io_testfailovers.yaml +- bases/storage.simplyblock.io_storagesitedeployments.yaml # +kubebuilder:scaffold:crdkustomizeresource patches: [] diff --git a/operator/config/rbac/role.yaml b/operator/config/rbac/role.yaml index 4d6684b03..1335dea16 100644 --- a/operator/config/rbac/role.yaml +++ b/operator/config/rbac/role.yaml @@ -176,6 +176,14 @@ rules: - patch - update - watch +- apiGroups: + - cluster.open-cluster-management.io + resources: + - managedclusters + verbs: + - get + - list + - watch - apiGroups: - coordination.k8s.io resources: @@ -325,6 +333,7 @@ rules: - storagenodes - storagepoolops - storagepools + - storagesitedeployments - tasks - testfailovers - volumemigrations @@ -361,6 +370,7 @@ rules: - storagenodes/finalizers - storagepoolops/finalizers - storagepools/finalizers + - storagesitedeployments/finalizers - tasks/finalizers - testfailovers/finalizers - volumemigrations/finalizers @@ -394,6 +404,7 @@ rules: - storagenodesets/status - storagepoolops/status - storagepools/status + - storagesitedeployments/status - tasks/status - testfailovers/status - volumegroupsnapshotops/status diff --git a/operator/internal/controller/storagesitedeployment_controller.go b/operator/internal/controller/storagesitedeployment_controller.go new file mode 100644 index 000000000..c4a24c091 --- /dev/null +++ b/operator/internal/controller/storagesitedeployment_controller.go @@ -0,0 +1,719 @@ +// The StorageSiteDeployment controller carries a managed site's storage +// deployment request from the hub to the site and projects the site's answer +// back. +// +// It never holds a site kubeconfig: every write to the site is a ManifestWork +// in the site's hub namespace, every read a ManagedClusterView there, the same +// two primitives the TestFailover controller uses. The work carries the +// OperatorOps discovery first; once the site has written a draft with nodes, it +// carries a server-side apply of the draft's sizing, and when the request is +// approved, the draft's approval. The views project the draft, the +// StorageCluster the approved draft expands into, and that cluster's nodes. +// +// The request withdraws nothing on deletion: the work is released with its +// resources orphaned, so a storage cluster is never torn down by deleting the +// request that asked for it. See docs/design/control-center-managed-discovery.md +// in the simplyblock-dr repository. + +package controller + +import ( + "context" + "crypto/sha256" + "encoding/json" + "fmt" + "sort" + "time" + + corev1 "k8s.io/api/core/v1" + apiequality "k8s.io/apimachinery/pkg/api/equality" + apierrors "k8s.io/apimachinery/pkg/api/errors" + "k8s.io/apimachinery/pkg/api/meta" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" + "k8s.io/apimachinery/pkg/runtime" + "k8s.io/client-go/tools/events" + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/controller/controllerutil" + logf "sigs.k8s.io/controller-runtime/pkg/log" + + workv1 "open-cluster-management.io/api/work/v1" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" + "github.com/simplyblock/simplyblock-operator/internal/controllers/pool" +) + +// storageSiteDeploymentIDLabel tags the work and the views of one request, so +// its release can enumerate them. +const storageSiteDeploymentIDLabel = "storage.simplyblock.io/site-deployment" + +// finalizerStorageSiteDeployment holds the request until its work and views +// are released. The work is released with its resources orphaned: the +// discovery, the draft and the storage cluster stay on the site. +const finalizerStorageSiteDeployment = "storage.simplyblock.io/storagesitedeployment-release" + +// hubDeployFieldManager is the field manager the work agent applies the +// draft's sizing and approval with, so the discovery's own fields on the draft +// are left to their owner. +const hubDeployFieldManager = "hub-deploy" + +// storageSiteDeploymentRequeue is how long the reconcile waits before reading +// the site's views again while the request is in progress. +const storageSiteDeploymentRequeue = 15 * time.Second + +// storageSiteDeploymentOnlineRequeue keeps an Online request's projection of +// the storage cluster fresh without polling the site hard. +const storageSiteDeploymentOnlineRequeue = 2 * time.Minute + +// maxStorageSiteNodeViews bounds the per-node views a request keeps: list +// views are not supported by ManagedClusterView, so there is one per node +// named in the draft's nodeRefs. +const maxStorageSiteNodeViews = 64 + +// annotationStorageClusterDefaultPool is the annotation the StorageCluster +// controller records the cluster's first pool under (controllers/cluster). +const annotationStorageClusterDefaultPool = "storage.simplyblock.io/default-pool" + +// Conditions of a request. +const ( + ConditionStorageSiteDelivered = "Delivered" + ConditionStorageSiteDiscovered = "Discovered" + ConditionStorageSiteApproved = "Approved" + ConditionStorageSiteReady = "Ready" +) + +// StorageSiteDeploymentReconciler reconciles a StorageSiteDeployment object. +type StorageSiteDeploymentReconciler struct { + client.Client + Scheme *runtime.Scheme + Recorder events.EventRecorder +} + +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagesitedeployments,verbs=get;list;watch;create;update;patch;delete +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagesitedeployments/status,verbs=get;update;patch +// +kubebuilder:rbac:groups=storage.simplyblock.io,resources=storagesitedeployments/finalizers,verbs=update +// +kubebuilder:rbac:groups=work.open-cluster-management.io,resources=manifestworks,verbs=get;list;watch;create;update;patch;delete +// +kubebuilder:rbac:groups=view.open-cluster-management.io,resources=managedclusterviews,verbs=get;list;watch;create;update;patch;delete +// +kubebuilder:rbac:groups=cluster.open-cluster-management.io,resources=managedclusters,verbs=get;list;watch +// +kubebuilder:rbac:groups=events.k8s.io,resources=events,verbs=create;patch + +// Reconcile carries the request to the site and projects the site's answer: +// it ensures the finalizer and the work, reads the draft, the storage cluster +// and its nodes through views, and derives the phase from what they report. +func (r *StorageSiteDeploymentReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) { + log := logf.FromContext(ctx) + + var sd simplyblockv1alpha2.StorageSiteDeployment + if err := r.Get(ctx, req.NamespacedName, &sd); err != nil { + return ctrl.Result{}, client.IgnoreNotFound(err) + } + + if !sd.DeletionTimestamp.IsZero() { + return r.reconcileDeletion(ctx, &sd) + } + if !controllerutil.ContainsFinalizer(&sd, finalizerStorageSiteDeployment) { + controllerutil.AddFinalizer(&sd, finalizerStorageSiteDeployment) + if err := r.Update(ctx, &sd); err != nil { + return ctrl.Result{}, err + } + return ctrl.Result{Requeue: true}, nil + } + + work, err := r.ensureWork(ctx, &sd) + if err != nil { + log.Error(err, "ensure ManifestWork") + return ctrl.Result{}, err + } + return r.project(ctx, &sd, work) +} + +// ensureWork creates or updates the request's ManifestWork: the discovery +// always, and the draft's sizing and approval once the site has a draft with +// nodes to apply them to. +func (r *StorageSiteDeploymentReconciler) ensureWork(ctx context.Context, sd *simplyblockv1alpha2.StorageSiteDeployment) (*workv1.ManifestWork, error) { + want, err := r.manifestWork(sd) + if err != nil { + return nil, err + } + var have workv1.ManifestWork + err = r.Get(ctx, client.ObjectKeyFromObject(want), &have) + if apierrors.IsNotFound(err) { + if err := r.Create(ctx, want); err != nil { + return nil, err + } + return want, nil + } + if err != nil { + return nil, err + } + if !apiequality.Semantic.DeepEqual(have.Spec, want.Spec) || !apiequality.Semantic.DeepEqual(have.Labels, want.Labels) { + have.Spec = want.Spec + have.Labels = want.Labels + if err := r.Update(ctx, &have); err != nil { + return nil, err + } + } + return &have, nil +} + +// manifestWork builds the request's work as it should be now. Its resources +// are orphaned on delete: the discovery, the draft and what it expanded into +// are the site's, and deleting the request must not take them away. +func (r *StorageSiteDeploymentReconciler) manifestWork(sd *simplyblockv1alpha2.StorageSiteDeployment) (*workv1.ManifestWork, error) { + ns := siteNamespace(sd) + draft := draftName(sd) + labels := map[string]string{storageSiteDeploymentIDLabel: string(sd.UID)} + + ops := map[string]any{ + "apiVersion": simplyblockv1alpha2.GroupVersion.String(), + "kind": "OperatorOps", + "metadata": map[string]any{"name": discoveryName(sd), "namespace": ns, "labels": labels}, + "spec": map[string]any{ + "action": string(simplyblockv1alpha2.OperatorOpsActionDiscover), + "discover": discoverSpec(sd), + }, + } + manifests := []workv1.Manifest{} + var configs []workv1.ManifestConfigOption + raw, err := json.Marshal(ops) + if err != nil { + return nil, fmt.Errorf("marshal discovery: %w", err) + } + manifests = append(manifests, workv1.Manifest{RawExtension: runtime.RawExtension{Raw: raw}}) + + // The sizing and the approval are applied onto the draft the discovery + // wrote, never before it exists: an apply that created the draft would + // make a document with no nodes, which the site refuses. + if draftHasNodes(sd.Status.Draft) && (sd.Spec.Sizing != nil || sd.Spec.Approved) { + spec := map[string]any{"approved": sd.Spec.Approved} + if tpl := sizingTemplate(sd.Spec.Sizing); len(tpl) > 0 { + spec["cluster"] = tpl + } + cdc := map[string]any{ + "apiVersion": simplyblockv1alpha2.GroupVersion.String(), + "kind": "ClusterDeploymentConfig", + "metadata": map[string]any{"name": draft, "namespace": ns}, + "spec": spec, + } + raw, err := json.Marshal(cdc) + if err != nil { + return nil, fmt.Errorf("marshal draft apply: %w", err) + } + manifests = append(manifests, workv1.Manifest{RawExtension: runtime.RawExtension{Raw: raw}}) + configs = append(configs, workv1.ManifestConfigOption{ + ResourceIdentifier: workv1.ResourceIdentifier{ + Group: simplyblockv1alpha2.GroupVersion.Group, Resource: "clusterdeploymentconfigs", Namespace: ns, Name: draft, + }, + UpdateStrategy: &workv1.UpdateStrategy{ + Type: workv1.UpdateStrategyTypeServerSideApply, + ServerSideApply: &workv1.ServerSideApplyConfig{ + Force: true, + FieldManager: hubDeployFieldManager, + }, + }, + FeedbackRules: []workv1.FeedbackRule{{ + Type: workv1.JSONPathsType, + JsonPaths: []workv1.JsonPath{ + {Name: "phase", Path: ".status.phase"}, + {Name: "approved", Path: ".spec.approved"}, + }, + }}, + }) + } + + return &workv1.ManifestWork{ + ObjectMeta: metav1.ObjectMeta{ + Name: workName(sd), + Namespace: sd.Spec.Cluster, + Labels: labels, + }, + Spec: workv1.ManifestWorkSpec{ + Workload: workv1.ManifestsTemplate{Manifests: manifests}, + ManifestConfigs: configs, + DeleteOption: &workv1.DeleteOption{PropagationPolicy: workv1.DeletePropagationPolicyTypeOrphan}, + }, + }, nil +} + +// discoverSpec is the OperatorOps discover block of the request. +func discoverSpec(sd *simplyblockv1alpha2.StorageSiteDeployment) map[string]any { + d := map[string]any{"configName": draftName(sd)} + if sd.Spec.Discover.EnableControlPlaneNodes != nil { + d["enableControlPlaneNodes"] = *sd.Spec.Discover.EnableControlPlaneNodes + } + if len(sd.Spec.Discover.Workers) > 0 { + d["workers"] = sd.Spec.Discover.Workers + } + if len(sd.Spec.Discover.NodeSelector) > 0 { + d["nodeSelector"] = sd.Spec.Discover.NodeSelector + } + return d +} + +// sizingTemplate is the draft's cluster template fields the request sets. +// Only the stated fields are applied, so what the discovery wrote stays. +func sizingTemplate(s *simplyblockv1alpha2.StorageSiteSizing) map[string]any { + tpl := map[string]any{} + if s == nil { + return tpl + } + if s.Name != "" { + tpl["name"] = s.Name + } + if s.VCPUCount != nil { + tpl["vcpuCount"] = *s.VCPUCount + } + if s.MinHugePagesSize != "" { + tpl["minHugePagesSize"] = s.MinHugePagesSize + } + if s.MaxSubsystemCount != nil { + tpl["maxSubsystemCount"] = *s.MaxSubsystemCount + } + if s.EnableDriveFormat != nil { + tpl["enableDriveFormat"] = *s.EnableDriveFormat + } + if s.EnableJournalDevice != nil { + tpl["enableJournalDevice"] = *s.EnableJournalDevice + } + if s.Stripe != nil { + stripe := map[string]any{} + if s.Stripe.DataChunks != nil { + stripe["dataChunks"] = *s.Stripe.DataChunks + } + if s.Stripe.ParityChunks != nil { + stripe["parityChunks"] = *s.Stripe.ParityChunks + } + tpl["stripe"] = stripe + } + return tpl +} + +// project reads the site's views and derives the request's phase. +func (r *StorageSiteDeploymentReconciler) project(ctx context.Context, sd *simplyblockv1alpha2.StorageSiteDeployment, work *workv1.ManifestWork) (ctrl.Result, error) { + delivered, deliveryMessage := workDelivery(work) + + draftObj, haveDraft, err := r.projected(ctx, sd, "draft", "clusterdeploymentconfigs", draftName(sd), siteNamespace(sd)) + if err != nil { + return ctrl.Result{}, err + } + if !haveDraft { + if deliveryMessage != "" { + return r.setPhase(ctx, sd, simplyblockv1alpha2.StorageSiteDeploymentPhaseFailed, deliveryMessage, func(s *simplyblockv1alpha2.StorageSiteDeploymentStatus) { + setCondition(s, ConditionStorageSiteDelivered, false, "NotApplied", deliveryMessage) + }) + } + return r.setPhase(ctx, sd, simplyblockv1alpha2.StorageSiteDeploymentPhaseDiscovering, + fmt.Sprintf("waiting for site %s to write draft %s/%s", sd.Spec.Cluster, siteNamespace(sd), draftName(sd)), + func(s *simplyblockv1alpha2.StorageSiteDeploymentStatus) { + s.WorkName = work.Name + setCondition(s, ConditionStorageSiteDelivered, delivered, deliveryReason(delivered), deliveryNote(delivered, deliveryMessage)) + }) + } + + var cdc simplyblockv1alpha2.ClusterDeploymentConfig + if err := runtime.DefaultUnstructuredConverter.FromUnstructured(draftObj, &cdc); err != nil { + return r.setPhase(ctx, sd, simplyblockv1alpha2.StorageSiteDeploymentPhaseFailed, + fmt.Sprintf("the site's draft could not be read: %v", err), nil) + } + draft := projectDraft(&cdc) + base := func(s *simplyblockv1alpha2.StorageSiteDeploymentStatus) { + s.WorkName = work.Name + s.Draft = draft + setCondition(s, ConditionStorageSiteDelivered, delivered, deliveryReason(delivered), deliveryNote(delivered, deliveryMessage)) + setCondition(s, ConditionStorageSiteDiscovered, draftHasNodes(draft), "Nodes", fmt.Sprintf("%d node(s) in the draft", draftNodeCount(draft))) + setCondition(s, ConditionStorageSiteApproved, draft.Approved, "SiteDraft", fmt.Sprintf("the site's draft approved=%t", draft.Approved)) + } + + if !draftHasNodes(draft) { + return r.setPhase(ctx, sd, simplyblockv1alpha2.StorageSiteDeploymentPhaseDiscovering, + "the site's draft names no node yet: discovery is running", base) + } + if deliveryMessage != "" { + // The draft exists, so the message is about the sizing or the approval + // the work could not apply: the site refused it. + return r.setPhase(ctx, sd, simplyblockv1alpha2.StorageSiteDeploymentPhaseFailed, deliveryMessage, base) + } + switch { + case cdc.Status.Phase == simplyblockv1alpha2.ClusterDeploymentConfigPhaseFailed: + return r.setPhase(ctx, sd, simplyblockv1alpha2.StorageSiteDeploymentPhaseFailed, + "the site's draft failed: "+orDefault(cdc.Status.Message, "no message"), base) + case !cdc.Spec.Approved: + msg := "the draft awaits approval" + if sd.Spec.Approved { + msg = "approval requested; waiting for the site's draft to take it" + } else if sd.Spec.Sizing != nil && !sizingApplied(sd.Spec.Sizing, cdc.Spec.Cluster) { + msg = "the draft awaits approval; the sizing is being applied" + } + return r.setPhase(ctx, sd, simplyblockv1alpha2.StorageSiteDeploymentPhaseDrafted, msg, base) + } + + // Approved on the site: follow the StorageCluster it expands into. + clusterName := cdc.Status.ClusterRef + if clusterName == "" && cdc.Spec.Cluster != nil { + clusterName = cdc.Spec.Cluster.Name + } + if clusterName == "" { + return r.setPhase(ctx, sd, simplyblockv1alpha2.StorageSiteDeploymentPhaseDeploying, + "the draft is approved; waiting for the site to name its StorageCluster", base) + } + scObj, haveSC, err := r.projected(ctx, sd, "cluster", "storageclusters", clusterName, siteNamespace(sd)) + if err != nil { + return ctrl.Result{}, err + } + sc := &simplyblockv1alpha2.StorageSiteCluster{Name: clusterName} + if haveSC { + var cluster simplyblockv1alpha2.StorageCluster + if err := runtime.DefaultUnstructuredConverter.FromUnstructured(scObj, &cluster); err == nil { + sc.UUID = cluster.Status.UUID + sc.Phase = string(cluster.Status.Phase) + sc.Pool = cluster.Annotations[annotationStorageClusterDefaultPool] + } + } + if sc.Pool == "" { + sc.Pool = pool.DefaultPoolName(clusterName) + } + nodes, err := r.projectNodes(ctx, sd, draft.NodeRefs) + if err != nil { + return ctrl.Result{}, err + } + sc.Nodes = nodes + withCluster := func(s *simplyblockv1alpha2.StorageSiteDeploymentStatus) { + base(s) + s.StorageCluster = sc + } + + switch { + case sc.Phase == string(simplyblockv1alpha2.StorageClusterPhaseOnline) && sc.UUID != "": + return r.setPhase(ctx, sd, simplyblockv1alpha2.StorageSiteDeploymentPhaseOnline, + fmt.Sprintf("StorageCluster %s is Online (%d node(s))", clusterName, len(nodes)), func(s *simplyblockv1alpha2.StorageSiteDeploymentStatus) { + withCluster(s) + setCondition(s, ConditionStorageSiteReady, true, "Online", "the StorageCluster is Online") + }) + case cdc.Status.Phase == simplyblockv1alpha2.ClusterDeploymentConfigPhaseExpanded && sc.Phase == string(simplyblockv1alpha2.StorageClusterPhaseUnavailable): + return r.setPhase(ctx, sd, simplyblockv1alpha2.StorageSiteDeploymentPhaseFailed, + fmt.Sprintf("StorageCluster %s is Unavailable after the expansion", clusterName), withCluster) + } + msg := fmt.Sprintf("draft %s, StorageCluster %s %s", orDefault(string(cdc.Status.Phase), "Expanding"), clusterName, orDefault(sc.Phase, "not reported yet")) + if step := cdc.Status.Step.State; step != "" && cdc.Status.Phase == simplyblockv1alpha2.ClusterDeploymentConfigPhaseExpanding { + msg = fmt.Sprintf("draft Expanding (%s), StorageCluster %s %s", step, clusterName, orDefault(sc.Phase, "not reported yet")) + } + return r.setPhase(ctx, sd, simplyblockv1alpha2.StorageSiteDeploymentPhaseDeploying, msg, func(s *simplyblockv1alpha2.StorageSiteDeploymentStatus) { + withCluster(s) + setCondition(s, ConditionStorageSiteReady, false, "Deploying", msg) + }) +} + +// projectNodes projects the draft's StorageNodes, one view each. +func (r *StorageSiteDeploymentReconciler) projectNodes(ctx context.Context, sd *simplyblockv1alpha2.StorageSiteDeployment, refs []string) ([]simplyblockv1alpha2.StorageSiteNode, error) { + sorted := append([]string(nil), refs...) + sort.Strings(sorted) + if len(sorted) > maxStorageSiteNodeViews { + sorted = sorted[:maxStorageSiteNodeViews] + } + nodes := make([]simplyblockv1alpha2.StorageSiteNode, 0, len(sorted)) + for i, name := range sorted { + obj, ok, err := r.projected(ctx, sd, fmt.Sprintf("node-%d", i), "storagenodes", name, siteNamespace(sd)) + if err != nil { + return nil, err + } + n := simplyblockv1alpha2.StorageSiteNode{Name: name} + if ok { + var sn simplyblockv1alpha2.StorageNode + if err := runtime.DefaultUnstructuredConverter.FromUnstructured(obj, &sn); err == nil { + n.Phase = string(sn.Status.Phase) + n.Hostname = sn.Status.Hostname + } + } + nodes = append(nodes, n) + } + return nodes, nil +} + +// projected reads one of the request's views, creating it when it is missing, +// and reports whether the site has projected the object yet. +func (r *StorageSiteDeploymentReconciler) projected(ctx context.Context, sd *simplyblockv1alpha2.StorageSiteDeployment, suffix, resource, name, namespace string) (map[string]interface{}, bool, error) { + viewName := storageSiteViewName(sd, suffix) + view := &unstructured.Unstructured{} + view.SetGroupVersionKind(managedClusterViewGVK) + getErr := r.Get(ctx, client.ObjectKey{Namespace: sd.Spec.Cluster, Name: viewName}, view) + if apierrors.IsNotFound(getErr) { + want := newStorageSiteView(sd, viewName, resource, name, namespace) + if createErr := r.Create(ctx, want); createErr != nil { + return nil, false, createErr + } + return nil, false, nil + } + if getErr != nil { + return nil, false, getErr + } + // A view that names another object (the draft's cluster changed) is + // pointed at the right one. + scope, _, _ := unstructured.NestedMap(view.Object, "spec", "scope") + if scope["name"] != name || scope["resource"] != resource { + want := newStorageSiteView(sd, viewName, resource, name, namespace) + view.Object["spec"] = want.Object["spec"] + if err := r.Update(ctx, view); err != nil { + return nil, false, err + } + return nil, false, nil + } + result, found, nestedErr := unstructured.NestedMap(view.Object, "status", "result") + if nestedErr != nil || !found || len(result) == 0 { + return nil, false, nil + } + return result, true, nil +} + +// newStorageSiteView asks the site to project one object back to the hub. +func newStorageSiteView(sd *simplyblockv1alpha2.StorageSiteDeployment, name, resource, targetName, targetNamespace string) *unstructured.Unstructured { + scope := map[string]interface{}{"resource": resource, "name": targetName} + if targetNamespace != "" { + scope["namespace"] = targetNamespace + } + view := &unstructured.Unstructured{} + view.SetGroupVersionKind(managedClusterViewGVK) + view.SetNamespace(sd.Spec.Cluster) + view.SetName(name) + view.SetLabels(map[string]string{storageSiteDeploymentIDLabel: string(sd.UID)}) + _ = unstructured.SetNestedMap(view.Object, scope, "spec", "scope") + return view +} + +// setPhase writes the phase, the message and the mutation, and records a +// phase change as an event. +func (r *StorageSiteDeploymentReconciler) setPhase(ctx context.Context, sd *simplyblockv1alpha2.StorageSiteDeployment, phase simplyblockv1alpha2.StorageSiteDeploymentPhase, message string, mutate func(*simplyblockv1alpha2.StorageSiteDeploymentStatus)) (ctrl.Result, error) { + previous := sd.Status.Phase + if err := r.patchStatus(ctx, sd, func(s *simplyblockv1alpha2.StorageSiteDeploymentStatus) { + if mutate != nil { + mutate(s) + } + s.Phase = phase + s.Message = message + }); err != nil { + return ctrl.Result{}, err + } + if previous != phase && r.Recorder != nil { + kind := corev1.EventTypeNormal + if phase == simplyblockv1alpha2.StorageSiteDeploymentPhaseFailed { + kind = corev1.EventTypeWarning + } + r.Recorder.Eventf(sd, nil, kind, string(phase), string(phase), "%s", message) + } + switch phase { + case simplyblockv1alpha2.StorageSiteDeploymentPhaseOnline: + return ctrl.Result{RequeueAfter: storageSiteDeploymentOnlineRequeue}, nil + case simplyblockv1alpha2.StorageSiteDeploymentPhaseFailed: + // A failure on the site may clear (a node comes back, a draft is + // corrected on the site): keep reading at the slow cadence. + return ctrl.Result{RequeueAfter: storageSiteDeploymentOnlineRequeue}, nil + } + return ctrl.Result{RequeueAfter: storageSiteDeploymentRequeue}, nil +} + +// patchStatus applies mutate to the status and writes it with the generation +// it was computed from. +func (r *StorageSiteDeploymentReconciler) patchStatus(ctx context.Context, sd *simplyblockv1alpha2.StorageSiteDeployment, mutate func(*simplyblockv1alpha2.StorageSiteDeploymentStatus)) error { + base := client.MergeFrom(sd.DeepCopy()) + mutate(&sd.Status) + sd.Status.ObservedGeneration = sd.Generation + return r.Status().Patch(ctx, sd, base) +} + +// reconcileDeletion releases the request's work and views and removes the +// finalizer. The work orphans its resources, so nothing on the site goes. +func (r *StorageSiteDeploymentReconciler) reconcileDeletion(ctx context.Context, sd *simplyblockv1alpha2.StorageSiteDeployment) (ctrl.Result, error) { + if !controllerutil.ContainsFinalizer(sd, finalizerStorageSiteDeployment) { + return ctrl.Result{}, nil + } + var work workv1.ManifestWork + err := r.Get(ctx, client.ObjectKey{Namespace: sd.Spec.Cluster, Name: workName(sd)}, &work) + switch { + case apierrors.IsNotFound(err): + case err != nil: + return ctrl.Result{}, err + case work.DeletionTimestamp.IsZero(): + if err := client.IgnoreNotFound(r.Delete(ctx, &work)); err != nil { + return ctrl.Result{}, err + } + return ctrl.Result{RequeueAfter: 5 * time.Second}, nil + default: + // Deleting: wait for the work agent to release it. + return ctrl.Result{RequeueAfter: 5 * time.Second}, nil + } + views := &unstructured.UnstructuredList{} + listGVK := managedClusterViewGVK + listGVK.Kind += "List" + views.SetGroupVersionKind(listGVK) + if err := r.List(ctx, views, client.InNamespace(sd.Spec.Cluster), client.MatchingLabels{storageSiteDeploymentIDLabel: string(sd.UID)}); err != nil && !meta.IsNoMatchError(err) { + return ctrl.Result{}, err + } + for i := range views.Items { + if err := client.IgnoreNotFound(r.Delete(ctx, &views.Items[i])); err != nil { + return ctrl.Result{}, err + } + } + controllerutil.RemoveFinalizer(sd, finalizerStorageSiteDeployment) + return ctrl.Result{}, r.Update(ctx, sd) +} + +// workDelivery reads the work's status: whether every manifest is applied, +// and the first manifest's refusal when one is not. +func workDelivery(work *workv1.ManifestWork) (applied bool, message string) { + if work == nil { + return false, "" + } + for _, m := range work.Status.ResourceStatus.Manifests { + for _, c := range m.Conditions { + if c.Type == workv1.ManifestApplied && c.Status == metav1.ConditionFalse { + return false, fmt.Sprintf("the site did not apply %s %s: %s", m.ResourceMeta.Kind, m.ResourceMeta.Name, c.Message) + } + } + } + for _, c := range work.Status.Conditions { + if c.Type == workv1.WorkApplied { + return c.Status == metav1.ConditionTrue, "" + } + } + return false, "" +} + +func deliveryReason(delivered bool) string { + if delivered { + return "Applied" + } + return "Pending" +} + +func deliveryNote(delivered bool, message string) string { + switch { + case message != "": + return message + case delivered: + return "the work is applied on the site" + } + return "the work is not applied on the site yet" +} + +// projectDraft is the draft as the status carries it. +func projectDraft(cdc *simplyblockv1alpha2.ClusterDeploymentConfig) *simplyblockv1alpha2.StorageSiteDraft { + d := &simplyblockv1alpha2.StorageSiteDraft{ + Name: cdc.Name, + Phase: string(cdc.Status.Phase), + Message: cdc.Status.Message, + Approved: cdc.Spec.Approved, + NodeSets: cdc.Spec.NodeSets, + NodeRefs: cdc.Status.NodeRefs, + } + if cdc.Spec.Cluster != nil { + d.Cluster = cdc.Spec.Cluster.DeepCopy() + } + if d.Phase == "" { + d.Phase = string(simplyblockv1alpha2.ClusterDeploymentConfigPhaseDraft) + } + return d +} + +// draftHasNodes is whether the draft names at least one worker. +func draftHasNodes(d *simplyblockv1alpha2.StorageSiteDraft) bool { + return draftNodeCount(d) > 0 +} + +func draftNodeCount(d *simplyblockv1alpha2.StorageSiteDraft) int { + if d == nil { + return 0 + } + n := 0 + for _, set := range d.NodeSets { + for _, g := range set.Groups { + n += len(g.Workers) + } + } + return n +} + +// sizingApplied is whether the draft's template carries the request's sizing. +func sizingApplied(s *simplyblockv1alpha2.StorageSiteSizing, tpl *simplyblockv1alpha2.ClusterTemplate) bool { + if s == nil { + return true + } + if tpl == nil { + return false + } + eq32 := func(want, have *int32) bool { return want == nil || (have != nil && *have == *want) } + eqBool := func(want, have *bool) bool { return want == nil || (have != nil && *have == *want) } + if s.Name != "" && tpl.Name != s.Name { + return false + } + if s.MinHugePagesSize != "" && tpl.MinHugePagesSize != s.MinHugePagesSize { + return false + } + if !eq32(s.VCPUCount, tpl.VCPUCount) || !eq32(s.MaxSubsystemCount, tpl.MaxSubsystemCount) || + !eqBool(s.EnableDriveFormat, tpl.EnableDriveFormat) || !eqBool(s.EnableJournalDevice, tpl.EnableJournalDevice) { + return false + } + if s.Stripe != nil { + if tpl.Stripe == nil || !eq32(s.Stripe.DataChunks, tpl.Stripe.DataChunks) || !eq32(s.Stripe.ParityChunks, tpl.Stripe.ParityChunks) { + return false + } + } + return true +} + +func setCondition(s *simplyblockv1alpha2.StorageSiteDeploymentStatus, kind string, ok bool, reason, message string) { + status := metav1.ConditionFalse + if ok { + status = metav1.ConditionTrue + } + meta.SetStatusCondition(&s.Conditions, metav1.Condition{Type: kind, Status: status, Reason: reason, Message: message, ObservedGeneration: s.ObservedGeneration}) +} + +func orDefault(s, d string) string { + if s == "" { + return d + } + return s +} + +func siteNamespace(sd *simplyblockv1alpha2.StorageSiteDeployment) string { + return orDefault(sd.Spec.SiteNamespace, "simplyblock") +} + +func draftName(sd *simplyblockv1alpha2.StorageSiteDeployment) string { + return orDefault(sd.Spec.DraftName, "site-draft") +} + +// storageSiteHash is a short, deterministic id of the request for the names +// of its work and views, which must be bounded and found again on a restart. +func storageSiteHash(sd *simplyblockv1alpha2.StorageSiteDeployment) string { + h := sha256.Sum256([]byte(sd.Namespace + "/" + sd.Name)) + return fmt.Sprintf("%x", h[:6]) +} + +func workName(sd *simplyblockv1alpha2.StorageSiteDeployment) string { + return "sbsd-" + storageSiteHash(sd) +} + +func storageSiteViewName(sd *simplyblockv1alpha2.StorageSiteDeployment, suffix string) string { + return "sbsd-" + storageSiteHash(sd) + "-" + suffix +} + +// discoveryName is the OperatorOps on the site. It carries a hash of the +// discovery's parameters, so a changed discovery is a new run rather than an +// edit of a finished one. +func discoveryName(sd *simplyblockv1alpha2.StorageSiteDeployment) string { + raw, _ := json.Marshal(sd.Spec.Discover) + h := sha256.Sum256(raw) + return fmt.Sprintf("hub-discover-%s-%x", draftName(sd), h[:3]) +} + +// SetupWithManager registers the controller. The work and the views live in +// the site's namespace on the hub, where an owner reference to the request +// cannot point, so the site's answers are read on the requeue cadence of the +// phase rather than through a watch. +func (r *StorageSiteDeploymentReconciler) SetupWithManager(mgr ctrl.Manager) error { + return ctrl.NewControllerManagedBy(mgr). + For(&simplyblockv1alpha2.StorageSiteDeployment{}). + Named("storagesitedeployment"). + Complete(r) +} diff --git a/operator/internal/controller/storagesitedeployment_controller_unit_test.go b/operator/internal/controller/storagesitedeployment_controller_unit_test.go new file mode 100644 index 000000000..63f9bbf76 --- /dev/null +++ b/operator/internal/controller/storagesitedeployment_controller_unit_test.go @@ -0,0 +1,335 @@ +package controller + +import ( + "context" + "encoding/json" + "testing" + + "k8s.io/apimachinery/pkg/api/meta" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" + "k8s.io/apimachinery/pkg/runtime" + "k8s.io/apimachinery/pkg/types" + ctrl "sigs.k8s.io/controller-runtime" + "sigs.k8s.io/controller-runtime/pkg/client" + + workv1 "open-cluster-management.io/api/work/v1" + + simplyblockv1alpha2 "github.com/simplyblock/simplyblock-operator/api/v1alpha2" +) + +const ( + testSDNamespace = "simplyblock" + testSDName = "site-a" + testSDCluster = "site-a" +) + +func newStorageSiteDeploymentReconciler(t *testing.T, objects ...client.Object) (*StorageSiteDeploymentReconciler, client.Client) { + t.Helper() + scheme := newTestScheme(t) + scheme.AddKnownTypeWithName(managedClusterViewGVK, &unstructured.Unstructured{}) + listGVK := managedClusterViewGVK + listGVK.Kind += "List" + scheme.AddKnownTypeWithName(listGVK, &unstructured.UnstructuredList{}) + if err := workv1.Install(scheme); err != nil { + t.Fatalf("register work/v1 scheme: %v", err) + } + cl := newTestClient(t, scheme, + []client.Object{&simplyblockv1alpha2.StorageSiteDeployment{}, &workv1.ManifestWork{}}, + objects...) + return &StorageSiteDeploymentReconciler{Client: cl, Scheme: scheme, Recorder: &fakeRecorder{}}, cl +} + +func newSiteDeployment(mutate func(*simplyblockv1alpha2.StorageSiteDeployment)) *simplyblockv1alpha2.StorageSiteDeployment { + enable := true + sd := &simplyblockv1alpha2.StorageSiteDeployment{ + ObjectMeta: metav1.ObjectMeta{Name: testSDName, Namespace: testSDNamespace, UID: types.UID("uid-site-a")}, + Spec: simplyblockv1alpha2.StorageSiteDeploymentSpec{ + Cluster: testSDCluster, + Discover: simplyblockv1alpha2.StorageSiteDiscovery{EnableControlPlaneNodes: &enable}, + }, + } + if mutate != nil { + mutate(sd) + } + return sd +} + +// reconcileSD runs the reconcile n times (the first adds the finalizer) and +// returns the request as stored. +func reconcileSD(t *testing.T, r *StorageSiteDeploymentReconciler, cl client.Client, n int) *simplyblockv1alpha2.StorageSiteDeployment { + t.Helper() + req := ctrl.Request{NamespacedName: types.NamespacedName{Namespace: testSDNamespace, Name: testSDName}} + for i := 0; i < n; i++ { + if _, err := r.Reconcile(context.Background(), req); err != nil { + t.Fatalf("reconcile %d: %v", i, err) + } + } + var sd simplyblockv1alpha2.StorageSiteDeployment + if err := cl.Get(context.Background(), req.NamespacedName, &sd); err != nil { + t.Fatalf("get request: %v", err) + } + return &sd +} + +func getSiteWork(t *testing.T, cl client.Client, sd *simplyblockv1alpha2.StorageSiteDeployment) *workv1.ManifestWork { + t.Helper() + var w workv1.ManifestWork + if err := cl.Get(context.Background(), client.ObjectKey{Namespace: sd.Spec.Cluster, Name: workName(sd)}, &w); err != nil { + t.Fatalf("get ManifestWork: %v", err) + } + return &w +} + +// workManifests decodes the work's manifests by kind. +func workManifests(t *testing.T, w *workv1.ManifestWork) map[string]map[string]any { + t.Helper() + out := map[string]map[string]any{} + for _, m := range w.Spec.Workload.Manifests { + obj := map[string]any{} + if err := json.Unmarshal(m.Raw, &obj); err != nil { + t.Fatalf("decode manifest: %v", err) + } + out[obj["kind"].(string)] = obj + } + return out +} + +// projectSiteObject writes what the site's view controller would project. +func projectSiteObject(t *testing.T, cl client.Client, sd *simplyblockv1alpha2.StorageSiteDeployment, suffix string, obj any) { + t.Helper() + v := getView(t, cl, sd.Spec.Cluster, storageSiteViewName(sd, suffix)) + result, err := runtime.DefaultUnstructuredConverter.ToUnstructured(obj) + if err != nil { + t.Fatalf("convert projection: %v", err) + } + setViewResult(t, cl, v, result) +} + +func siteDraft(approved bool, phase simplyblockv1alpha2.ClusterDeploymentConfigPhase, mutate func(*simplyblockv1alpha2.ClusterDeploymentConfig)) *simplyblockv1alpha2.ClusterDeploymentConfig { + cdc := &simplyblockv1alpha2.ClusterDeploymentConfig{ + TypeMeta: metav1.TypeMeta{APIVersion: simplyblockv1alpha2.GroupVersion.String(), Kind: "ClusterDeploymentConfig"}, + ObjectMeta: metav1.ObjectMeta{Name: "site-draft", Namespace: "simplyblock"}, + Spec: simplyblockv1alpha2.ClusterDeploymentConfigSpec{ + Approved: approved, + NodeSets: []simplyblockv1alpha2.NodeSet{{ + Name: "default", + Groups: []simplyblockv1alpha2.NodeGroup{{Workers: []string{"n1", "n2", "n3"}}}, + }}, + }, + } + cdc.Status.Phase = phase + if mutate != nil { + mutate(cdc) + } + return cdc +} + +func TestStorageSiteDeploymentDeliversTheDiscoveryAndWaitsForTheDraft(t *testing.T) { + r, cl := newStorageSiteDeploymentReconciler(t, newSiteDeployment(nil)) + sd := reconcileSD(t, r, cl, 2) + + if sd.Status.Phase != simplyblockv1alpha2.StorageSiteDeploymentPhaseDiscovering { + t.Fatalf("phase = %q, want Discovering", sd.Status.Phase) + } + w := getSiteWork(t, cl, sd) + if w.Spec.DeleteOption == nil || w.Spec.DeleteOption.PropagationPolicy != workv1.DeletePropagationPolicyTypeOrphan { + t.Errorf("work delete option = %+v, want Orphan: a request must never take the site's storage away", w.Spec.DeleteOption) + } + ms := workManifests(t, w) + ops, ok := ms["OperatorOps"] + if !ok || len(ms) != 1 { + t.Fatalf("manifests = %v, want the discovery alone before the draft exists", keys(ms)) + } + discover := ops["spec"].(map[string]any)["discover"].(map[string]any) + if discover["configName"] != "site-draft" || discover["enableControlPlaneNodes"] != true { + t.Errorf("discover = %v, want configName site-draft and control-plane nodes enabled", discover) + } + // The draft view exists so the site can project the draft. + getView(t, cl, testSDCluster, storageSiteViewName(sd, "draft")) +} + +func TestStorageSiteDeploymentSizesTheDraftOnceItNamesNodes(t *testing.T) { + vcpu := int32(8) + data, parity := int32(1), int32(1) + r, cl := newStorageSiteDeploymentReconciler(t, newSiteDeployment(func(sd *simplyblockv1alpha2.StorageSiteDeployment) { + sd.Spec.Sizing = &simplyblockv1alpha2.StorageSiteSizing{ + Name: "sb-site-a", VCPUCount: &vcpu, MinHugePagesSize: "8G", + Stripe: &simplyblockv1alpha2.StripeSpec{DataChunks: &data, ParityChunks: &parity}, + } + })) + sd := reconcileSD(t, r, cl, 2) + projectSiteObject(t, cl, sd, "draft", siteDraft(false, simplyblockv1alpha2.ClusterDeploymentConfigPhaseDraft, nil)) + sd = reconcileSD(t, r, cl, 2) + + if sd.Status.Phase != simplyblockv1alpha2.StorageSiteDeploymentPhaseDrafted { + t.Fatalf("phase = %q (%s), want Drafted", sd.Status.Phase, sd.Status.Message) + } + if got := draftNodeCount(sd.Status.Draft); got != 3 { + t.Errorf("projected draft nodes = %d, want 3", got) + } + ms := workManifests(t, getSiteWork(t, cl, sd)) + cdc, ok := ms["ClusterDeploymentConfig"] + if !ok { + t.Fatalf("manifests = %v, want the draft's sizing applied", keys(ms)) + } + spec := cdc["spec"].(map[string]any) + if spec["approved"] != false { + t.Errorf("approved = %v, want false before the request is approved", spec["approved"]) + } + cluster := spec["cluster"].(map[string]any) + if cluster["name"] != "sb-site-a" || cluster["vcpuCount"] != float64(8) || cluster["minHugePagesSize"] != "8G" { + t.Errorf("cluster template = %v, want the sizing", cluster) + } + if _, ok := spec["nodeSets"]; ok { + t.Error("the sizing apply must not carry nodeSets: they are the discovery's") + } + if !meta.IsStatusConditionTrue(sd.Status.Conditions, ConditionStorageSiteDiscovered) { + t.Error("Discovered condition is not True for a draft with nodes") + } +} + +func TestStorageSiteDeploymentApprovalFollowsTheClusterToOnline(t *testing.T) { + r, cl := newStorageSiteDeploymentReconciler(t, newSiteDeployment(func(sd *simplyblockv1alpha2.StorageSiteDeployment) { + sd.Spec.Approved = true + })) + sd := reconcileSD(t, r, cl, 2) + projectSiteObject(t, cl, sd, "draft", siteDraft(false, simplyblockv1alpha2.ClusterDeploymentConfigPhaseDraft, nil)) + // One pass projects the draft, the next delivers the approval onto it. + sd = reconcileSD(t, r, cl, 2) + + ms := workManifests(t, getSiteWork(t, cl, sd)) + if ms["ClusterDeploymentConfig"]["spec"].(map[string]any)["approved"] != true { + t.Fatalf("the work does not carry the approval: %v", ms["ClusterDeploymentConfig"]) + } + if sd.Status.Phase != simplyblockv1alpha2.StorageSiteDeploymentPhaseDrafted { + t.Fatalf("phase = %q, want Drafted until the site's draft takes the approval", sd.Status.Phase) + } + + // The site's draft takes it and starts expanding. + projectSiteObject(t, cl, sd, "draft", siteDraft(true, simplyblockv1alpha2.ClusterDeploymentConfigPhaseExpanding, func(c *simplyblockv1alpha2.ClusterDeploymentConfig) { + c.Status.ClusterRef = "sb-site-a" + c.Status.NodeRefs = []string{"sn-1", "sn-2", "sn-3"} + })) + sd = reconcileSD(t, r, cl, 1) + if sd.Status.Phase != simplyblockv1alpha2.StorageSiteDeploymentPhaseDeploying { + t.Fatalf("phase = %q (%s), want Deploying", sd.Status.Phase, sd.Status.Message) + } + + // The cluster comes Online. + sc := &simplyblockv1alpha2.StorageCluster{ + TypeMeta: metav1.TypeMeta{APIVersion: simplyblockv1alpha2.GroupVersion.String(), Kind: "StorageCluster"}, + ObjectMeta: metav1.ObjectMeta{Name: "sb-site-a", Namespace: "simplyblock"}, + } + sc.Status.UUID = "8f8dd277-1544-4177-9a74-e0f66eb2672c" + sc.Status.Phase = simplyblockv1alpha2.StorageClusterPhaseOnline + projectSiteObject(t, cl, sd, "cluster", sc) + sd = reconcileSD(t, r, cl, 1) + + if sd.Status.Phase != simplyblockv1alpha2.StorageSiteDeploymentPhaseOnline { + t.Fatalf("phase = %q (%s), want Online", sd.Status.Phase, sd.Status.Message) + } + if sd.Status.StorageCluster == nil || sd.Status.StorageCluster.UUID != sc.Status.UUID { + t.Fatalf("storageCluster = %+v, want the site's uuid", sd.Status.StorageCluster) + } + if sd.Status.StorageCluster.Pool == "" { + t.Error("storageCluster.pool is empty: a StorageClass needs it") + } + if n := len(sd.Status.StorageCluster.Nodes); n != 3 { + t.Errorf("projected nodes = %d, want 3", n) + } + if !meta.IsStatusConditionTrue(sd.Status.Conditions, ConditionStorageSiteReady) { + t.Error("Ready condition is not True for an Online cluster") + } +} + +func TestStorageSiteDeploymentReportsAFailedDraft(t *testing.T) { + r, cl := newStorageSiteDeploymentReconciler(t, newSiteDeployment(func(sd *simplyblockv1alpha2.StorageSiteDeployment) { + sd.Spec.Approved = true + })) + sd := reconcileSD(t, r, cl, 2) + projectSiteObject(t, cl, sd, "draft", siteDraft(true, simplyblockv1alpha2.ClusterDeploymentConfigPhaseFailed, func(c *simplyblockv1alpha2.ClusterDeploymentConfig) { + c.Status.Message = "node n2 has no free device" + })) + sd = reconcileSD(t, r, cl, 1) + if sd.Status.Phase != simplyblockv1alpha2.StorageSiteDeploymentPhaseFailed { + t.Fatalf("phase = %q, want Failed", sd.Status.Phase) + } + if sd.Status.Message != "the site's draft failed: node n2 has no free device" { + t.Errorf("message = %q, want the site's own reason", sd.Status.Message) + } +} + +func TestStorageSiteDeploymentADraftWithoutNodesIsStillDiscovering(t *testing.T) { + r, cl := newStorageSiteDeploymentReconciler(t, newSiteDeployment(func(sd *simplyblockv1alpha2.StorageSiteDeployment) { + sd.Spec.Approved = true + })) + sd := reconcileSD(t, r, cl, 2) + projectSiteObject(t, cl, sd, "draft", siteDraft(false, simplyblockv1alpha2.ClusterDeploymentConfigPhaseDraft, func(c *simplyblockv1alpha2.ClusterDeploymentConfig) { + c.Spec.NodeSets = nil + })) + sd = reconcileSD(t, r, cl, 1) + if sd.Status.Phase != simplyblockv1alpha2.StorageSiteDeploymentPhaseDiscovering { + t.Fatalf("phase = %q, want Discovering", sd.Status.Phase) + } + if _, ok := workManifests(t, getSiteWork(t, cl, sd))["ClusterDeploymentConfig"]; ok { + t.Error("the approval was delivered onto a draft without nodes") + } +} + +func TestStorageSiteDeploymentAChangedDiscoveryIsANewRun(t *testing.T) { + a := newSiteDeployment(nil) + b := newSiteDeployment(func(sd *simplyblockv1alpha2.StorageSiteDeployment) { + sd.Spec.Discover.Workers = []string{"n1"} + }) + if discoveryName(a) == discoveryName(b) { + t.Fatalf("discovery name %q did not change with the discovery's parameters", discoveryName(a)) + } + if discoveryName(a) != discoveryName(newSiteDeployment(nil)) { + t.Fatal("the discovery name is not stable for the same parameters") + } +} + +func TestStorageSiteDeploymentDeletionOrphansTheSitesStorage(t *testing.T) { + r, cl := newStorageSiteDeploymentReconciler(t, newSiteDeployment(nil)) + sd := reconcileSD(t, r, cl, 2) + if err := cl.Delete(context.Background(), sd); err != nil { + t.Fatalf("delete request: %v", err) + } + // First pass deletes the work; the fake client removes it at once. + reconcileOnce := func() { + req := ctrl.Request{NamespacedName: types.NamespacedName{Namespace: testSDNamespace, Name: testSDName}} + if _, err := r.Reconcile(context.Background(), req); err != nil { + t.Fatalf("reconcile: %v", err) + } + } + reconcileOnce() + reconcileOnce() + + var w workv1.ManifestWork + if err := cl.Get(context.Background(), client.ObjectKey{Namespace: testSDCluster, Name: workName(sd)}, &w); err == nil { + t.Error("the work is still there after the request was deleted") + } + views := &unstructured.UnstructuredList{} + listGVK := managedClusterViewGVK + listGVK.Kind += "List" + views.SetGroupVersionKind(listGVK) + if err := cl.List(context.Background(), views, client.InNamespace(testSDCluster)); err != nil { + t.Fatalf("list views: %v", err) + } + if len(views.Items) != 0 { + t.Errorf("%d view(s) left after the request was deleted", len(views.Items)) + } + var gone simplyblockv1alpha2.StorageSiteDeployment + if err := cl.Get(context.Background(), client.ObjectKeyFromObject(sd), &gone); err == nil { + t.Error("the request is still there: the finalizer was not released") + } +} + +func keys(m map[string]map[string]any) []string { + out := make([]string, 0, len(m)) + for k := range m { + out = append(out, k) + } + return out +} diff --git a/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagesitedeployments.yaml b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagesitedeployments.yaml new file mode 100644 index 000000000..9c2718a8f --- /dev/null +++ b/operator/internal/upgrade/crds/manifests/storage.simplyblock.io_storagesitedeployments.yaml @@ -0,0 +1,1057 @@ +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition +metadata: + annotations: + controller-gen.kubebuilder.io/version: v0.21.0 + name: storagesitedeployments.storage.simplyblock.io +spec: + group: storage.simplyblock.io + names: + kind: StorageSiteDeployment + listKind: StorageSiteDeploymentList + plural: storagesitedeployments + shortNames: + - sbsd + singular: storagesitedeployment + scope: Namespaced + versions: + - additionalPrinterColumns: + - jsonPath: .spec.cluster + name: Cluster + type: string + - jsonPath: .spec.approved + name: Approved + type: boolean + - jsonPath: .status.phase + name: Phase + type: string + - jsonPath: .status.draft.phase + name: Draft + type: string + - jsonPath: .status.storageCluster.phase + name: Storage + type: string + - jsonPath: .status.message + name: Message + priority: 1 + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha2 + schema: + openAPIV3Schema: + description: |- + StorageSiteDeployment requests a managed site's storage cluster from the hub: + a discovery on the site, the sizing of the draft it writes, and the approval + that expands the draft into a StorageCluster. The hub carries the request + through OCM and projects the site's draft and cluster into the status. + Deleting the request leaves the storage cluster alone. + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: StorageSiteDeploymentSpec is the request for one site's storage + cluster. + properties: + approved: + default: false + description: |- + Approved is the review gate, delivered to the draft on the site. One-way, + as the draft's own gate is. + type: boolean + cluster: + description: |- + Cluster is the OCM ManagedCluster the storage is deployed on. The request's + ManifestWork and views live in its namespace on the hub. Immutable. + maxLength: 63 + minLength: 1 + type: string + x-kubernetes-validations: + - message: cluster is immutable + rule: self == oldSelf + discover: + description: |- + Discover is the discovery the site runs first. Changing it runs another + discovery, which rewrites the draft. + properties: + enableControlPlaneNodes: + description: |- + EnableControlPlaneNodes lets the discovery consider the nodes that run the + API server. Every server of a small distribution is one, so a three-node + site has no storage without it. + type: boolean + nodeSelector: + additionalProperties: + type: string + description: NodeSelector limits the discovery to the nodes carrying + these labels. + type: object + workers: + description: Workers limits the discovery to these nodes. Empty + is every worker. + items: + type: string + type: array + x-kubernetes-list-type: set + type: object + draftName: + default: site-draft + description: |- + DraftName is the ClusterDeploymentConfig the discovery writes on the site + and the request sizes and approves. Immutable. + maxLength: 63 + type: string + x-kubernetes-validations: + - message: draftName is immutable + rule: self == oldSelf + siteNamespace: + default: simplyblock + description: |- + SiteNamespace is the simplyblock operator's namespace on the site, where + the discovery and the draft live. + maxLength: 63 + type: string + sizing: + description: |- + Sizing is written onto the draft's cluster template once the draft exists, + so the reviewer sees the sized draft before approving it. + properties: + enableDriveFormat: + description: EnableDriveFormat lets the deployment format the + devices it takes. + type: boolean + enableJournalDevice: + description: EnableJournalDevice dedicates one device per node + to the journal. + type: boolean + maxSubsystemCount: + description: MaxSubsystemCount is the number of NVMe-oF subsystems + each node serves. + format: int32 + minimum: 1 + type: integer + minHugePagesSize: + description: |- + MinHugePagesSize is the hugepage memory each storage node takes, as a + quantity ("8G"). + type: string + name: + description: Name is the StorageCluster's name on the site. + maxLength: 63 + type: string + stripe: + description: Stripe is the erasure-coding layout. + properties: + dataChunks: + description: DataChunks is the number of data chunks per stripe + (ndcs). + format: int32 + minimum: 1 + type: integer + parityChunks: + description: |- + ParityChunks is the number of parity chunks per stripe (npcs), and + therefore how many chunk losses a stripe survives. + format: int32 + minimum: 0 + type: integer + type: object + x-kubernetes-validations: + - message: the erasure-coding scheme must be one of 1+0, 1+1, + 2+1, 4+1, 1+2, 2+2, or 4+2, written as dataChunks+parityChunks, + and an unstated half is 1 + rule: '[has(self.dataChunks) ? self.dataChunks : 1, has(self.parityChunks) + ? self.parityChunks : 1] in [[1, 0], [1, 1], [2, 1], [4, 1], + [1, 2], [2, 2], [4, 2]]' + vcpuCount: + description: VCPUCount is the number of vCPUs each storage node + takes. + format: int32 + minimum: 1 + type: integer + type: object + required: + - cluster + type: object + x-kubernetes-validations: + - message: 'approval is one-way: an approved deployment cannot be un-approved' + rule: '!has(oldSelf.approved) || !oldSelf.approved || self.approved' + status: + description: StorageSiteDeploymentStatus is what the site reports back, + projected. + properties: + conditions: + description: |- + Conditions: Delivered (the work is applied on the site), Discovered (the + draft names nodes), Approved (the site's draft is approved), Ready (the + StorageCluster is Online). + items: + description: Condition contains details for one aspect of the current + state of this API Resource. + properties: + lastTransitionTime: + description: |- + lastTransitionTime is the last time the condition transitioned from one status to another. + This should be when the underlying condition changed. If that is not known, then using the time when the API field changed is acceptable. + format: date-time + type: string + message: + description: |- + message is a human readable message indicating details about the transition. + This may be an empty string. + maxLength: 32768 + type: string + observedGeneration: + description: |- + observedGeneration represents the .metadata.generation that the condition was set based upon. + For instance, if .metadata.generation is currently 12, but the .status.conditions[x].observedGeneration is 9, the condition is out of date + with respect to the current state of the instance. + format: int64 + minimum: 0 + type: integer + reason: + description: |- + reason contains a programmatic identifier indicating the reason for the condition's last transition. + Producers of specific condition types may define expected values and meanings for this field, + and whether the values are considered a guaranteed API. + The value should be a CamelCase string. + This field may not be empty. + maxLength: 1024 + minLength: 1 + pattern: ^[A-Za-z]([A-Za-z0-9_,:]*[A-Za-z0-9_])?$ + type: string + status: + description: status of the condition, one of True, False, Unknown. + enum: + - "True" + - "False" + - Unknown + type: string + type: + description: type of condition in CamelCase or in foo.example.com/CamelCase. + maxLength: 316 + pattern: ^([a-z0-9]([-a-z0-9]*[a-z0-9])?(\.[a-z0-9]([-a-z0-9]*[a-z0-9])?)*/)?(([A-Za-z0-9][-A-Za-z0-9_.]*)?[A-Za-z0-9])$ + type: string + required: + - lastTransitionTime + - message + - reason + - status + - type + type: object + type: array + x-kubernetes-list-map-keys: + - type + x-kubernetes-list-type: map + draft: + description: Draft is the draft as the site reports it. + properties: + approved: + description: Approved is whether the draft is approved on the + site. + type: boolean + cluster: + description: Cluster is the draft's cluster template, with the + sizing applied. + properties: + backup: + description: |- + Backup is where this cluster's backups live, and it expands into + StorageCluster.spec.backup unchanged. + + It is here for the reason KMS is: a store stated on the document is + present when the cluster is created rather than patched in afterward by + whoever remembers. Unlike most of what this template carries, the field it + fills is mutable, so a document that states none costs nothing permanent. + A cluster can be given a store whenever there is one to give. + + The Secret it names is not resolved at admission. It is a core object a + deployment legitimately creates alongside the document or after it, and + the cluster's own creation is where its absence is reported. + properties: + bucket: + description: Bucket is the bucket backups are written + to and read from. + type: string + credentialsSecretRef: + description: |- + CredentialsSecretRef names the Secret holding the access key and the + secret key. It is a reference rather than the values, because a spec is + readable by anybody who can read the object. + properties: + name: + default: "" + description: |- + Name of the referent. + This field is effectively required, but due to backwards compatibility is + allowed to be empty. Instances of this type with an empty value here are + almost certainly wrong. + More info: https://kubernetes.io/docs/concepts/overview/working-with-objects/names/#names + type: string + type: object + x-kubernetes-map-type: atomic + endpoint: + description: Endpoint is the S3 endpoint, for example, + https://s3.example.com. + pattern: ^https?://[a-zA-Z0-9.-]+(:[0-9]{1,5})?(/.*)?$ + type: string + prefix: + description: |- + Prefix narrows the store to one key prefix, so that several clusters can + share a bucket without each walking the others' backups. + type: string + region: + description: Region is the bucket's region, for endpoints + that do not imply one. + type: string + required: + - bucket + - credentialsSecretRef + - endpoint + type: object + containerResources: + description: |- + ContainerResources sizes the storage-node container, and expands into the + cluster's own spec.storageNodes.containerResources. + + The container it sizes is the node's management API rather than SPDK, + which runs in a pod of its own: what outgrows the default is a node + answering for many subsystems, not a node moving more data. It is on the + document because a deployment is where a fleet's sizing is decided, and + a cluster written from a document that could not say so had to be edited + afterward on a field the document owns everywhere else. + + Stating either half replaces both. The defaults apply to a cluster that + states neither requests nor limits, so a document stating requests alone + produces a container with no limits rather than one with the default + limits, and a memory limit is what has the kubelet evict a leaking agent + rather than losing the worker. + + It is a pointer because a resource block is a struct, and a struct with + omitempty is serialized whether or not anything is in it: as a value, + every document a discovery run writes would carry an empty + containerResources that says nothing and that a reviewer has to decide + about. + properties: + claims: + description: |- + Claims lists the names of resources, defined in spec.resourceClaims, + that are used by this container. + + This field depends on the + DynamicResourceAllocation feature gate. + + This field is immutable. It can only be set for containers. + items: + description: ResourceClaim references one entry in PodSpec.ResourceClaims. + properties: + name: + description: |- + Name must match the name of one entry in pod.spec.resourceClaims of + the Pod where this field is used. It makes that resource available + inside a container. + type: string + request: + description: |- + Request is the name chosen for a request in the referenced claim. + If empty, everything from the claim is made available, otherwise + only the result of this request. + type: string + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + limits: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Limits describes the maximum amount of compute resources allowed. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + requests: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Requests describes the minimum amount of compute resources required. + If Requests is omitted for a container, it defaults to Limits if that is explicitly specified, + otherwise to an implementation-defined value. Requests cannot exceed Limits. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + type: object + enableAtomicity4K: + description: |- + EnableAtomicity4K enforces 4K write atomicity on every device this + deployment names, which is what lets checksum validation run on devices + whose logical block size is under the data plane's 4K minimum. + + It is the route to checked I/O on a device that cannot be reformatted: a + logical block device's block size is fixed by the drive, and some NVMe + devices offer no 4K format either. Where a device can be reformatted, + EnableDriveFormat is the other route and this is unnecessary. + + It is an enforcement because the question is often unanswerable. A SATA + drive presenting 512-byte logical blocks over a 4K physical sector reports + 512 and nothing more, and a kernel older than 6.11 publishes no atomic + write attributes at all. Where a device does answer, the storage node's + report carries it, and a reviewer approves this against that rather than + against a vendor's datasheet -- because enforcing a guarantee the hardware + does not keep is how a torn write becomes a checksum that silently + disagrees with it. + + It means nothing unless EnableChecksumValidation is set, which is the + cluster's own rule and is left to the cluster to enforce. + type: boolean + enableChecksumValidation: + description: |- + EnableChecksumValidation turns on inline CRC validation of every I/O, for + silent-data-error protection. + + It is on the document because it is immutable on the cluster it lands on: + the backend bakes the checksum method into each device when the cluster is + created and never re-applies it, so a cluster created without this is one + nobody can turn it on for. A deployment that wants its data checked has to + say so here or not at all. + type: boolean + enableDriveFormat: + description: |- + EnableDriveFormat formats every device the document names before a storage + node takes it, which is how a drive carrying anything already is made + usable. + + It says what is wanted rather than how, because the how differs by device + class: an NVMe device is formatted to a 4K block size, and a logical block + device has its signatures wiped. One field covers both, so a document does + not have to know which class the expansion will resolve it to. + + It is on the document rather than defaulted further down because it is + destructive and the document is what somebody approves. A reviewer reading + a draft has to see that the drives it lists will be formatted, and be able + to strike it before approving; the cluster's own field is immutable once + the cluster exists, so a default nobody saw could not be undone either. + type: boolean + enableFailureDomains: + description: |- + EnableFailureDomains opts the cluster into failure-domain mode, in which + every group must label the fault group its workers belong to. + type: boolean + enableJournalDevice: + description: |- + EnableJournalDevice dedicates the smallest NVMe device on each of this + deployment's workers to the journal manager, instead of carving a journal + partition out of every device. + + It is here rather than on a node set because it is immutable on the cluster + it lands on, for the reason SocketsToUse is: the on-disk layout a fleet was + built with is not one a later document can vary. It also costs a drive of + capacity per node, which is a trade a reviewer approves rather than one a + default makes for them. + type: boolean + enableNodeAffinity: + description: |- + EnableNodeAffinity has the data plane serve an erasure-coded volume's I/O + from the local node's own devices where it can, before crossing the + network. + + It is not Kubernetes affinity, and the name is the one place this API + invites that reading: nothing about it schedules a pod, labels a worker, + or places a volume's primary node. The control plane carries it into the + cluster map it pushes to each node, where it sets the local node's index, + and what changes is which copy of a chunk is read. + Co-locating a workload with the primary node of its volume is a separate + mechanism and is not configured here. + + It is on the document because it is immutable on the cluster: the control + plane takes it at cluster create and never re-applies it, so this is the + only moment it can be set at all. + type: boolean + fabricType: + description: FabricType is the storage fabric. + maxLength: 32 + type: string + initContainerResources: + description: |- + InitContainerResources sizes both of the storage node's init containers, + and expands into the cluster's own spec.storageNodes.initContainerResources. + + They are sized apart from the container because they do a different job + and are gone before it starts: one writes the node's env file and the + other runs node_configure.py once, so what they need is a short burst + rather than the footprint of a process that runs for the node's life. + + Stating either half replaces both, as with containerResources, and it is + a pointer for the same reason. + properties: + claims: + description: |- + Claims lists the names of resources, defined in spec.resourceClaims, + that are used by this container. + + This field depends on the + DynamicResourceAllocation feature gate. + + This field is immutable. It can only be set for containers. + items: + description: ResourceClaim references one entry in PodSpec.ResourceClaims. + properties: + name: + description: |- + Name must match the name of one entry in pod.spec.resourceClaims of + the Pod where this field is used. It makes that resource available + inside a container. + type: string + request: + description: |- + Request is the name chosen for a request in the referenced claim. + If empty, everything from the claim is made available, otherwise + only the result of this request. + type: string + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + limits: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Limits describes the maximum amount of compute resources allowed. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + requests: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Requests describes the minimum amount of compute resources required. + If Requests is omitted for a container, it defaults to Limits if that is explicitly specified, + otherwise to an implementation-defined value. Requests cannot exceed Limits. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + type: object + kms: + description: |- + KMS selects where the cluster stores volume encryption keys. Stating it on + the document is what makes it present when the cluster is created, where + setting it on the StorageCluster afterward races with that creation. + properties: + vault: + description: Vault stores keys in HashiCorp Vault. + properties: + endpoint: + description: |- + Endpoint is the Vault endpoint, for example, https://vault.example.com:8200. + Rejected unless it resolves to an external address. + pattern: ^https?://[a-zA-Z0-9.-]+(:[0-9]{1,5})?(/.*)?$ + type: string + required: + - endpoint + type: object + type: object + maxSubsystemCount: + description: |- + MaxSubsystemCount is the maximum number of NVMe-oF subsystems each storage + node of this cluster serves. Required, because the StorageCluster's own + field is, and no StorageNode carries a copy of it. + format: int32 + maximum: 75 + minimum: 10 + type: integer + minHugePagesSize: + description: |- + MinHugePagesSize is the smallest huge-page allocation each storage node of + this cluster makes: 100G or 1T, where a bare number is gigabytes. Like + VCPUCount it is the cluster's and is copied onto every node the expansion + writes. Omitted, each node uses the computed minimum. + maxLength: 32 + type: string + name: + description: |- + Name is the StorageCluster's name, and is therefore held to what such a + name may be rather than to what an object name may be. A longer value is a + document the API server accepts and a CreatingCluster step that can never + succeed, since the cluster it would write is one the API server refuses. + maxLength: 63 + type: string + nodeProvisioningBudget: + description: |- + NodeProvisioningBudget is how many workers the expansion may have in the + node-add process at once. It expands into the cluster's own + spec.storageNodes.nodeProvisioningBudget, whose meaning it shares: the cap + is counted by distinct worker, so a two-socket host spends one of the + budget, and a worker hosting a FoundationDB pod is sequential whatever the + budget says. + + It is on the document because a document is what states the size of a + deployment, and a deployment of thirty workers added one at a time is the + difference between an afternoon and a week. Omitted, the cluster's default + of one applies, which is the serial behavior. + format: int32 + minimum: 1 + type: integer + nodesPerSocket: + description: |- + NodesPerSocket is how many storage nodes run per NUMA socket. See + SocketsToUse, which it multiplies. + format: int32 + maximum: 8 + minimum: 1 + type: integer + openshift: + description: |- + OpenShift is what this deployment states because it runs on OpenShift. It + expands into StorageCluster.spec.storageNodes.openshift, whose shape it + shares, and it is read only for a document whose environment is + OpenShift: the environment is what says which distribution this is, and + the block is what that distribution needs said beyond it. + properties: + machineConfigPool: + default: worker + description: |- + MachineConfigPool names a machine-config role the storage nodes' own pool + inherits from, beyond the worker role it always inherits. + + It is not the pool the nodes end up in, which the description it carried + before said and which cost a reader the reboot they were trying to avoid. + Adding a node creates a pool of its own, storage-, and moves the + node into it; a node belongs to exactly one custom pool, so whatever + machine configuration its previous pool carried is lost unless that + pool's role is named here for the new one to select as well. The default + is the role every pool already selects, which is what makes it a no-op + for a fleet whose workers are ordinary workers. + maxLength: 253 + pattern: ^[a-z0-9]([-a-z0-9]*[a-z0-9])?$ + type: string + type: object + ports: + description: |- + Ports are where this cluster's storage nodes listen. Unstated, and for + each member left unstated, the cluster's own defaults decide. + properties: + nodeAgent: + default: 50001 + description: |- + NodeAgent is the port each node's agent API listens on. It expands into + StorageCluster.spec.snodeApiPort, and it is named for the component + rather than for that field: the agent is what spec.images.nodeAgent pins + and what the storage-node DaemonSet runs. + format: int32 + maximum: 65535 + minimum: 1024 + type: integer + nvmf: + default: 4420 + description: |- + NVMf is the base of the NVMe-oF port range every node binds. It expands + into StorageCluster.spec.nvmfBasePort. + format: int32 + maximum: 65535 + minimum: 1024 + type: integer + rpc: + default: 8080 + description: |- + Rpc is the base of the RPC port range every node binds. It expands into + StorageCluster.spec.rpcBasePort. + format: int32 + maximum: 65535 + minimum: 1024 + type: integer + type: object + socketsToUse: + description: |- + SocketsToUse restricts the deployment to selected NUMA sockets, and empty + means socket 0 alone. With NodesPerSocket it decides how many storage nodes + each worker runs, so a group of two workers on a two-socket layout expands + to four nodes. + + It is here rather than on a node set because it is immutable on the cluster + it lands on: the layout a fleet was built with is not one a later document + can vary, and a reviewer should see it before the cluster exists. + items: + maxLength: 16 + type: string + maxItems: 16 + type: array + x-kubernetes-list-type: set + stripe: + description: Stripe is the erasure-coding layout. + properties: + dataChunks: + description: DataChunks is the number of data chunks per + stripe (ndcs). + format: int32 + minimum: 1 + type: integer + parityChunks: + description: |- + ParityChunks is the number of parity chunks per stripe (npcs), and + therefore how many chunk losses a stripe survives. + format: int32 + minimum: 0 + type: integer + type: object + x-kubernetes-validations: + - message: the erasure-coding scheme must be one of 1+0, 1+1, + 2+1, 4+1, 1+2, 2+2, or 4+2, written as dataChunks+parityChunks, + and an unstated half is 1 + rule: '[has(self.dataChunks) ? self.dataChunks : 1, has(self.parityChunks) + ? self.parityChunks : 1] in [[1, 0], [1, 1], [2, 1], [4, + 1], [1, 2], [2, 2], [4, 2]]' + tolerations: + description: |- + Tolerations are what the storage-node pods tolerate, and they expand into + the cluster's own spec.storageNodes.tolerations. + + A fleet that dedicates machines to storage taints them, which is what + keeps everything else off. The DaemonSet that lands on those machines has + to tolerate the taint or it schedules nowhere, and a document that could + not say so described a deployment that does not start: the correction was + an edit to the cluster the document had just created, on a field the + document owns everywhere else. + + A growth document states none. It names a cluster rather than describing + one, and that cluster already carries what its storage nodes tolerate. + items: + description: |- + The pod this Toleration is attached to tolerates any taint that matches + the triple using the matching operator . + properties: + effect: + description: |- + Effect indicates the taint effect to match. Empty means match all taint effects. + When specified, allowed values are NoSchedule, PreferNoSchedule and NoExecute. + type: string + key: + description: |- + Key is the taint key that the toleration applies to. Empty means match all taint keys. + If the key is empty, operator must be Exists; this combination means to match all values and all keys. + type: string + operator: + description: |- + Operator represents a key's relationship to the value. + Valid operators are Exists, Equal, Lt, and Gt. Defaults to Equal. + Exists is equivalent to wildcard for value, so that a pod can + tolerate all taints of a particular category. + Lt and Gt perform numeric comparisons (requires feature gate TaintTolerationComparisonOperators). + type: string + tolerationSeconds: + description: |- + TolerationSeconds represents the period of time the toleration (which must be + of effect NoExecute, otherwise this field is ignored) tolerates the taint. By default, + it is not set, which means tolerate the taint forever (do not evict). Zero and + negative values will be treated as 0 (evict immediately) by the system. + format: int64 + type: integer + value: + description: |- + Value is the taint value the toleration matches to. + If the operator is Exists, the value should be empty, otherwise just a regular string. + type: string + type: object + maxItems: 32 + type: array + vcpuCount: + description: |- + VCPUCount is the number of vCPUs allocated to SPDK on each storage node of + this cluster. It is stated here and nowhere below, because the control + plane assumes it uniform across a cluster's nodes; CreatingNodes copies it + into every StorageNode.spec.config.sizing it writes. Required, because the + StorageCluster's own field is. + The floor is 4 rather than a hardware limit: a node must carry one core + beyond this budget for the system, and the control plane's core layout + assigns no NVMe-oF poller core at all for a 2-vCPU budget. + format: int32 + minimum: 4 + type: integer + required: + - maxSubsystemCount + - name + - vcpuCount + type: object + message: + description: |- + Message is what the site says about the draft: validation findings while + it is a draft, the expansion's step afterwards. + type: string + name: + description: Name is the ClusterDeploymentConfig on the site. + type: string + nodeRefs: + description: NodeRefs are the StorageNode objects the expansion + created. + items: + type: string + type: array + x-kubernetes-list-type: set + nodeSets: + description: NodeSets are the nodes and devices the discovery + found, for review. + items: + description: |- + NodeSet is the organizational grouping of a deployment, usually a rack: the + workers a document adds or grows together. It carries no sizing, because sizing + is uniform across a cluster and is stated once in ClusterTemplate. + properties: + groups: + description: Groups are the sets of workers sharing one + configuration. + items: + description: |- + NodeGroup is a set of workers that share one configuration, which is what + makes ten identical machines one entry rather than ten. + properties: + dataInterfaces: + description: DataInterfaces are the data-plane network + interfaces. + items: + maxLength: 63 + type: string + maxItems: 32 + type: array + devices: + description: Devices selects the storage devices every + worker in the group uses. + properties: + block: + description: |- + Block names logical block devices by path ("/dev/sdb"). It expands into the + same config.deviceNames as NVMe, which takes a PCI address and a device + path in one list. It is the alternative to NVMe rather than a companion of + it: the two classes are not mixed within a cluster. + items: + maxLength: 255 + pattern: ^/dev/[a-zA-Z0-9._/-]+$ + type: string + maxItems: 128 + type: array + x-kubernetes-list-type: set + nvme: + description: NVMe names NVMe devices by PCI address + ("0000:5e:00.0"). + items: + maxLength: 32 + pattern: ^[0-9a-fA-F]{4}:[0-9a-fA-F]{2}:[0-9a-fA-F]{2}\.[0-9a-fA-F]$ + type: string + maxItems: 128 + type: array + x-kubernetes-list-type: set + type: object + x-kubernetes-validations: + - message: a device selection names NVMe addresses + or block devices, not both + rule: has(self.nvme) != has(self.block) + failureDomain: + description: |- + FailureDomain is the label of the fault group every worker in this group + belongs to ("rack-b"), which is usually the name of the rack, zone, or + power feed they share. Discovery seeds it from topology.kubernetes.io/zone + and leaves it unset where the Kubernetes API carries no topology, which + holds provisioning with a clear reason rather than guessing. It expands + into StorageNode.spec.config.failureDomain, whose shape it shares. + maxLength: 63 + pattern: ^[a-zA-Z0-9]([-_.a-zA-Z0-9]*[a-zA-Z0-9])?$ + type: string + journalManager: + description: JournalManager tunes the journal managers + on these nodes. + properties: + count: + description: Count is the number of journal managers + to configure. + format: int32 + minimum: 1 + type: integer + percentPerDevice: + description: PercentPerDevice is the share of + each device given to the journal. + format: int32 + maximum: 100 + minimum: 1 + type: integer + type: object + mgmtInterface: + description: MgmtInterface is the management network + interface the storage nodes bind. + maxLength: 63 + type: string + name: + description: |- + Name identifies the group within its node set, for a reader and for the + events a validation failure emits. + maxLength: 253 + type: string + reservedSystemCPU: + description: |- + ReservedSystemCPU is the CPU set held back from SPDK for the system on + these nodes, as a core list such as 0,1 or 0-3. + + It is a group's rather than the cluster's because it names core ids, and a + group is what a document calls the workers that share their hardware: 0,1 + on a sixteen-core worker and 0,1 on a ninety-six-core worker are different + fractions of the machine. It expands into + StorageNode.spec.config.reservedSystemCPU, whose shape it shares, and a + group that states none leaves the cluster's fleet-wide value to decide. + + On OpenShift it reaches the kubelet through a KubeletConfig for the + machine config pool, which is the cluster's, so groups that disagree there + are writing over one another's pool configuration. + maxLength: 63 + pattern: ^[0-9]+(-[0-9]+)?(,[0-9]+(-[0-9]+)?)*$ + type: string + spdkSystemMemory: + description: |- + SpdkSystemMemory is the memory the control plane starts SPDK with on these + nodes. + maxLength: 32 + pattern: ^[0-9]+(G|GI|GB|GiB|M|MI|MB|MiB|g|gi|gb|gib|m|mi|mb|mib)?$ + type: string + workers: + description: Workers are the Kubernetes worker hostnames + in this group. + items: + maxLength: 253 + type: string + maxItems: 200 + minItems: 1 + type: array + x-kubernetes-list-type: set + required: + - name + - workers + type: object + maxItems: 64 + minItems: 1 + type: array + name: + description: |- + Name is the node set's name. It is copied to StorageNode.spec.nodeSet, so + that a node can be traced back to the part of the document that produced + it. + maxLength: 253 + type: string + required: + - groups + - name + type: object + type: array + phase: + description: |- + Phase is the draft's own phase on the site (Draft, Expanding, Expanded, + Failed). + type: string + required: + - name + type: object + message: + description: |- + Message is the reason the phase is what it is: one sentence, replaced as + the request moves, and never a log. + type: string + observedGeneration: + description: |- + ObservedGeneration is the generation the rest of this status was computed + from. + format: int64 + type: integer + phase: + description: Phase is the request's own progress. + enum: + - Pending + - Discovering + - Drafted + - Deploying + - Online + - Failed + type: string + storageCluster: + description: StorageCluster is the cluster the approved draft produced. + properties: + name: + description: Name is the StorageCluster object on the site. + type: string + nodes: + description: Nodes are the cluster's storage nodes. + items: + description: |- + StorageSiteNode is one storage node of the deployed cluster, as the site + reports it. + properties: + hostname: + description: Hostname is the Kubernetes node it runs on. + type: string + name: + description: Name is the StorageNode object on the site. + type: string + phase: + description: Phase is the node's phase on the site. + type: string + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + phase: + description: Phase is the StorageCluster's phase on the site. + type: string + pool: + description: |- + Pool is the pool the cluster was created with, which a StorageClass names + in pool_name. + type: string + uuid: + description: |- + UUID is the storage cluster's id in the control plane, which a + StorageClass names in cluster_id. + type: string + required: + - name + type: object + workName: + description: WorkName is the ManifestWork carrying the request to + the site. + type: string + type: object + type: object + served: true + storage: true + subresources: + status: {} From 2ab649a73213cbaeca1c934c9316dd70ab61f5ef Mon Sep 17 00:00:00 2001 From: michael Date: Sat, 3 Oct 2026 17:14:37 +0300 Subject: [PATCH 180/206] csi-driver: wrap lines over 120 characters; operator: regenerate dist/install.yaml Lint (lll) on the chain-walk change and its tests; the installer bundle carries the TestFailover CRD's sourceVolumeMode. Co-Authored-By: Claude Opus 5.5 --- .../internal/csi/controller/replication.go | 3 ++- csi-driver/internal/csi/node/stats_test.go | 23 +++++++++++++++---- operator/dist/install.yaml | 7 ++++++ 3 files changed, 27 insertions(+), 6 deletions(-) diff --git a/csi-driver/internal/csi/controller/replication.go b/csi-driver/internal/csi/controller/replication.go index 85f5b21c4..bdc9045c5 100644 --- a/csi-driver/internal/csi/controller/replication.go +++ b/csi-driver/internal/csi/controller/replication.go @@ -172,7 +172,8 @@ func resolveChain(ctx context.Context, h *lvol.Handle, client *atlascp.Client) ( } } if len(hops) > maxChainHops { - return nil, false, fmt.Errorf("replication chain of %s did not converge within %d hops", hops[0].h.Handle(), maxChainHops) + return nil, false, fmt.Errorf("replication chain of %s did not converge within %d hops", + hops[0].h.Handle(), maxChainHops) } return hops, known, nil } diff --git a/csi-driver/internal/csi/node/stats_test.go b/csi-driver/internal/csi/node/stats_test.go index bb01510f7..2ddc16854 100644 --- a/csi-driver/internal/csi/node/stats_test.go +++ b/csi-driver/internal/csi/node/stats_test.go @@ -184,8 +184,14 @@ func TestRedirectToActiveVolumeFollowsALongChain(t *testing.T) { } active := member(moves) clients := map[string]*fakeRelationshipAPI{ - clusterA + "/" + poolA: {rels: map[string]*controlplane.ReplicationRelationship{}, conn: map[string]map[string]string{}}, - clusterB + "/" + poolB: {rels: map[string]*controlplane.ReplicationRelationship{}, conn: map[string]map[string]string{}}, + clusterA + "/" + poolA: { + rels: map[string]*controlplane.ReplicationRelationship{}, + conn: map[string]map[string]string{}, + }, + clusterB + "/" + poolB: { + rels: map[string]*controlplane.ReplicationRelationship{}, + conn: map[string]map[string]string{}, + }, } for i := 0; i < moves; i++ { srcC, srcP := cluster(i) @@ -231,13 +237,20 @@ func TestRedirectToActiveVolumeStopsOnALoop(t *testing.T) { y = "22222222-2222-2222-2222-222222222222" ) client := &fakeRelationshipAPI{rels: map[string]*controlplane.ReplicationRelationship{ - x: {SourceLvolID: x, TargetLvolID: y, SourceClusterID: clusterA, TargetClusterID: clusterA, TargetPoolID: poolA, ActiveLvolID: "zz"}, - y: {SourceLvolID: y, TargetLvolID: x, SourceClusterID: clusterA, TargetClusterID: clusterA, TargetPoolID: poolA, ActiveLvolID: "zz"}, + x: { + SourceLvolID: x, TargetLvolID: y, SourceClusterID: clusterA, + TargetClusterID: clusterA, TargetPoolID: poolA, ActiveLvolID: "zz", + }, + y: { + SourceLvolID: y, TargetLvolID: x, SourceClusterID: clusterA, + TargetClusterID: clusterA, TargetPoolID: poolA, ActiveLvolID: "zz", + }, }} orig := clusterClientFor defer func() { clusterClientFor = orig }() clusterClientFor = func(context.Context, string, string) (controlplane.ClusterAPI, error) { return client, nil } - if got := redirectToActiveVolume(context.Background(), client, x, clusterA+":"+poolA+":"+x, map[string]string{}); got != nil { + got := redirectToActiveVolume(context.Background(), client, x, clusterA+":"+poolA+":"+x, map[string]string{}) + if got != nil { t.Fatalf("a looping chain returned %v, want nil", got) } } diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index 97cc7fead..c7c996599 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -11463,6 +11463,13 @@ spec: identity keys are dropped so a failed clone lookup can never point the mount back at the source. type: object + sourceVolumeMode: + description: |- + SourceVolumeMode is the source PV's volumeMode (Filesystem or Block), + carried onto the bubble PV and PVC. A VM's disk is a Block claim; a bubble + claim that omitted the mode defaulted to Filesystem and the kubelet asked + the node plugin to mount a raw guest disk (2026-10-03). + type: string required: - sourceRef type: object From cce7357fb80c335695409e0463b844e34d253e91 Mon Sep 17 00:00:00 2001 From: michael Date: Sat, 3 Oct 2026 17:32:22 +0300 Subject: [PATCH 181/206] =?UTF-8?q?operator:=20lint=20=E2=80=94=20a=20list?= =?UTF-8?q?-kind=20suffix=20constant;=20wrap=20the=20hub-only=20log=20line?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 --- operator/cmd/main.go | 3 ++- .../internal/controller/storagesitedeployment_controller.go | 6 +++++- .../storagesitedeployment_controller_unit_test.go | 4 ++-- .../controller/testfailover_controller_unit_test.go | 4 ++-- 4 files changed, 11 insertions(+), 6 deletions(-) diff --git a/operator/cmd/main.go b/operator/cmd/main.go index 42b37093a..004db856e 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -973,7 +973,8 @@ func main() { os.Exit(1) } } else { - setupLog.Info("OCM ManifestWork resource not served; skipping the TestFailover and StorageSiteDeployment controllers (hub-only)", + setupLog.Info("OCM ManifestWork resource not served; skipping the TestFailover and "+ + "StorageSiteDeployment controllers (hub-only)", "groupVersion", ocmWorkGroupVersion, "resource", ocmManifestWorkResource) } // +kubebuilder:scaffold:builder diff --git a/operator/internal/controller/storagesitedeployment_controller.go b/operator/internal/controller/storagesitedeployment_controller.go index c4a24c091..87108d4fc 100644 --- a/operator/internal/controller/storagesitedeployment_controller.go +++ b/operator/internal/controller/storagesitedeployment_controller.go @@ -543,7 +543,7 @@ func (r *StorageSiteDeploymentReconciler) reconcileDeletion(ctx context.Context, } views := &unstructured.UnstructuredList{} listGVK := managedClusterViewGVK - listGVK.Kind += "List" + listGVK.Kind += listKindSuffix views.SetGroupVersionKind(listGVK) if err := r.List(ctx, views, client.InNamespace(sd.Spec.Cluster), client.MatchingLabels{storageSiteDeploymentIDLabel: string(sd.UID)}); err != nil && !meta.IsNoMatchError(err) { return ctrl.Result{}, err @@ -717,3 +717,7 @@ func (r *StorageSiteDeploymentReconciler) SetupWithManager(mgr ctrl.Manager) err Named("storagesitedeployment"). Complete(r) } + +// listKindSuffix turns a kind into its list kind (StorageCluster -> +// StorageClusterList) for an unstructured list read. +const listKindSuffix = "List" diff --git a/operator/internal/controller/storagesitedeployment_controller_unit_test.go b/operator/internal/controller/storagesitedeployment_controller_unit_test.go index 63f9bbf76..4bb445d6f 100644 --- a/operator/internal/controller/storagesitedeployment_controller_unit_test.go +++ b/operator/internal/controller/storagesitedeployment_controller_unit_test.go @@ -29,7 +29,7 @@ func newStorageSiteDeploymentReconciler(t *testing.T, objects ...client.Object) scheme := newTestScheme(t) scheme.AddKnownTypeWithName(managedClusterViewGVK, &unstructured.Unstructured{}) listGVK := managedClusterViewGVK - listGVK.Kind += "List" + listGVK.Kind += listKindSuffix scheme.AddKnownTypeWithName(listGVK, &unstructured.UnstructuredList{}) if err := workv1.Install(scheme); err != nil { t.Fatalf("register work/v1 scheme: %v", err) @@ -312,7 +312,7 @@ func TestStorageSiteDeploymentDeletionOrphansTheSitesStorage(t *testing.T) { } views := &unstructured.UnstructuredList{} listGVK := managedClusterViewGVK - listGVK.Kind += "List" + listGVK.Kind += listKindSuffix views.SetGroupVersionKind(listGVK) if err := cl.List(context.Background(), views, client.InNamespace(testSDCluster)); err != nil { t.Fatalf("list views: %v", err) diff --git a/operator/internal/controller/testfailover_controller_unit_test.go b/operator/internal/controller/testfailover_controller_unit_test.go index 2df684ed3..7c8effd6d 100644 --- a/operator/internal/controller/testfailover_controller_unit_test.go +++ b/operator/internal/controller/testfailover_controller_unit_test.go @@ -53,7 +53,7 @@ func newTestFailoverReconciler(t *testing.T, objects ...client.Object) (*TestFai // GVK (and list GVK) registered to create and read it. scheme.AddKnownTypeWithName(managedClusterViewGVK, &unstructured.Unstructured{}) listGVK := managedClusterViewGVK - listGVK.Kind += "List" + listGVK.Kind += listKindSuffix scheme.AddKnownTypeWithName(listGVK, &unstructured.UnstructuredList{}) if err := workv1.Install(scheme); err != nil { t.Fatalf("register work/v1 scheme: %v", err) @@ -443,7 +443,7 @@ func TestFailoverResolvingSourceReuseViewOnRestart(t *testing.T) { list := &unstructured.UnstructuredList{} gvk := managedClusterViewGVK - gvk.Kind += "List" + gvk.Kind += listKindSuffix list.SetGroupVersionKind(gvk) if err := cl.List(ctx, list, client.InNamespace(tf.Spec.SourceCluster)); err != nil { t.Fatalf("list views: %v", err) From 16f9713754db642cb0e6e2d1738683cb165fc2b4 Mon Sep 17 00:00:00 2001 From: michael Date: Sat, 3 Oct 2026 17:46:29 +0300 Subject: [PATCH 182/206] operator: regenerate dist/install.yaml with the StorageSiteDeployment CRD Co-Authored-By: Claude Opus 5.5 --- operator/dist/install.yaml | 1068 ++++++++++++++++++++++++++++++++++++ 1 file changed, 1068 insertions(+) diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index c7c996599..c91db2421 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -11152,6 +11152,1063 @@ spec: --- apiVersion: apiextensions.k8s.io/v1 kind: CustomResourceDefinition +metadata: + annotations: + controller-gen.kubebuilder.io/version: v0.21.0 + name: storagesitedeployments.storage.simplyblock.io +spec: + group: storage.simplyblock.io + names: + kind: StorageSiteDeployment + listKind: StorageSiteDeploymentList + plural: storagesitedeployments + shortNames: + - sbsd + singular: storagesitedeployment + scope: Namespaced + versions: + - additionalPrinterColumns: + - jsonPath: .spec.cluster + name: Cluster + type: string + - jsonPath: .spec.approved + name: Approved + type: boolean + - jsonPath: .status.phase + name: Phase + type: string + - jsonPath: .status.draft.phase + name: Draft + type: string + - jsonPath: .status.storageCluster.phase + name: Storage + type: string + - jsonPath: .status.message + name: Message + priority: 1 + type: string + - jsonPath: .metadata.creationTimestamp + name: Age + type: date + name: v1alpha2 + schema: + openAPIV3Schema: + description: |- + StorageSiteDeployment requests a managed site's storage cluster from the hub: + a discovery on the site, the sizing of the draft it writes, and the approval + that expands the draft into a StorageCluster. The hub carries the request + through OCM and projects the site's draft and cluster into the status. + Deleting the request leaves the storage cluster alone. + properties: + apiVersion: + description: |- + APIVersion defines the versioned schema of this representation of an object. + Servers should convert recognized schemas to the latest internal value, and + may reject unrecognized values. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources + type: string + kind: + description: |- + Kind is a string value representing the REST resource this object represents. + Servers may infer this from the endpoint the client submits requests to. + Cannot be updated. + In CamelCase. + More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds + type: string + metadata: + type: object + spec: + description: StorageSiteDeploymentSpec is the request for one site's storage + cluster. + properties: + approved: + default: false + description: |- + Approved is the review gate, delivered to the draft on the site. One-way, + as the draft's own gate is. + type: boolean + cluster: + description: |- + Cluster is the OCM ManagedCluster the storage is deployed on. The request's + ManifestWork and views live in its namespace on the hub. Immutable. + maxLength: 63 + minLength: 1 + type: string + x-kubernetes-validations: + - message: cluster is immutable + rule: self == oldSelf + discover: + description: |- + Discover is the discovery the site runs first. Changing it runs another + discovery, which rewrites the draft. + properties: + enableControlPlaneNodes: + description: |- + EnableControlPlaneNodes lets the discovery consider the nodes that run the + API server. Every server of a small distribution is one, so a three-node + site has no storage without it. + type: boolean + nodeSelector: + additionalProperties: + type: string + description: NodeSelector limits the discovery to the nodes carrying + these labels. + type: object + workers: + description: Workers limits the discovery to these nodes. Empty + is every worker. + items: + type: string + type: array + x-kubernetes-list-type: set + type: object + draftName: + default: site-draft + description: |- + DraftName is the ClusterDeploymentConfig the discovery writes on the site + and the request sizes and approves. Immutable. + maxLength: 63 + type: string + x-kubernetes-validations: + - message: draftName is immutable + rule: self == oldSelf + siteNamespace: + default: simplyblock + description: |- + SiteNamespace is the simplyblock operator's namespace on the site, where + the discovery and the draft live. + maxLength: 63 + type: string + sizing: + description: |- + Sizing is written onto the draft's cluster template once the draft exists, + so the reviewer sees the sized draft before approving it. + properties: + enableDriveFormat: + description: EnableDriveFormat lets the deployment format the + devices it takes. + type: boolean + enableJournalDevice: + description: EnableJournalDevice dedicates one device per node + to the journal. + type: boolean + maxSubsystemCount: + description: MaxSubsystemCount is the number of NVMe-oF subsystems + each node serves. + format: int32 + minimum: 1 + type: integer + minHugePagesSize: + description: |- + MinHugePagesSize is the hugepage memory each storage node takes, as a + quantity ("8G"). + type: string + name: + description: Name is the StorageCluster's name on the site. + maxLength: 63 + type: string + stripe: + description: Stripe is the erasure-coding layout. + properties: + dataChunks: + description: DataChunks is the number of data chunks per stripe + (ndcs). + format: int32 + minimum: 1 + type: integer + parityChunks: + description: |- + ParityChunks is the number of parity chunks per stripe (npcs), and + therefore how many chunk losses a stripe survives. + format: int32 + minimum: 0 + type: integer + type: object + x-kubernetes-validations: + - message: the erasure-coding scheme must be one of 1+0, 1+1, + 2+1, 4+1, 1+2, 2+2, or 4+2, written as dataChunks+parityChunks, + and an unstated half is 1 + rule: '[has(self.dataChunks) ? self.dataChunks : 1, has(self.parityChunks) + ? self.parityChunks : 1] in [[1, 0], [1, 1], [2, 1], [4, 1], + [1, 2], [2, 2], [4, 2]]' + vcpuCount: + description: VCPUCount is the number of vCPUs each storage node + takes. + format: int32 + minimum: 1 + type: integer + type: object + required: + - cluster + type: object + x-kubernetes-validations: + - message: 'approval is one-way: an approved deployment cannot be un-approved' + rule: '!has(oldSelf.approved) || !oldSelf.approved || self.approved' + status: + description: StorageSiteDeploymentStatus is what the site reports back, + projected. + properties: + conditions: + description: |- + Conditions: Delivered (the work is applied on the site), Discovered (the + draft names nodes), Approved (the site's draft is approved), Ready (the + StorageCluster is Online). + items: + description: Condition contains details for one aspect of the current + state of this API Resource. + properties: + lastTransitionTime: + description: |- + lastTransitionTime is the last time the condition transitioned from one status to another. + This should be when the underlying condition changed. If that is not known, then using the time when the API field changed is acceptable. + format: date-time + type: string + message: + description: |- + message is a human readable message indicating details about the transition. + This may be an empty string. + maxLength: 32768 + type: string + observedGeneration: + description: |- + observedGeneration represents the .metadata.generation that the condition was set based upon. + For instance, if .metadata.generation is currently 12, but the .status.conditions[x].observedGeneration is 9, the condition is out of date + with respect to the current state of the instance. + format: int64 + minimum: 0 + type: integer + reason: + description: |- + reason contains a programmatic identifier indicating the reason for the condition's last transition. + Producers of specific condition types may define expected values and meanings for this field, + and whether the values are considered a guaranteed API. + The value should be a CamelCase string. + This field may not be empty. + maxLength: 1024 + minLength: 1 + pattern: ^[A-Za-z]([A-Za-z0-9_,:]*[A-Za-z0-9_])?$ + type: string + status: + description: status of the condition, one of True, False, Unknown. + enum: + - "True" + - "False" + - Unknown + type: string + type: + description: type of condition in CamelCase or in foo.example.com/CamelCase. + maxLength: 316 + pattern: ^([a-z0-9]([-a-z0-9]*[a-z0-9])?(\.[a-z0-9]([-a-z0-9]*[a-z0-9])?)*/)?(([A-Za-z0-9][-A-Za-z0-9_.]*)?[A-Za-z0-9])$ + type: string + required: + - lastTransitionTime + - message + - reason + - status + - type + type: object + type: array + x-kubernetes-list-map-keys: + - type + x-kubernetes-list-type: map + draft: + description: Draft is the draft as the site reports it. + properties: + approved: + description: Approved is whether the draft is approved on the + site. + type: boolean + cluster: + description: Cluster is the draft's cluster template, with the + sizing applied. + properties: + backup: + description: |- + Backup is where this cluster's backups live, and it expands into + StorageCluster.spec.backup unchanged. + + It is here for the reason KMS is: a store stated on the document is + present when the cluster is created rather than patched in afterward by + whoever remembers. Unlike most of what this template carries, the field it + fills is mutable, so a document that states none costs nothing permanent. + A cluster can be given a store whenever there is one to give. + + The Secret it names is not resolved at admission. It is a core object a + deployment legitimately creates alongside the document or after it, and + the cluster's own creation is where its absence is reported. + properties: + bucket: + description: Bucket is the bucket backups are written + to and read from. + type: string + credentialsSecretRef: + description: |- + CredentialsSecretRef names the Secret holding the access key and the + secret key. It is a reference rather than the values, because a spec is + readable by anybody who can read the object. + properties: + name: + default: "" + description: |- + Name of the referent. + This field is effectively required, but due to backwards compatibility is + allowed to be empty. Instances of this type with an empty value here are + almost certainly wrong. + More info: https://kubernetes.io/docs/concepts/overview/working-with-objects/names/#names + type: string + type: object + x-kubernetes-map-type: atomic + endpoint: + description: Endpoint is the S3 endpoint, for example, + https://s3.example.com. + pattern: ^https?://[a-zA-Z0-9.-]+(:[0-9]{1,5})?(/.*)?$ + type: string + prefix: + description: |- + Prefix narrows the store to one key prefix, so that several clusters can + share a bucket without each walking the others' backups. + type: string + region: + description: Region is the bucket's region, for endpoints + that do not imply one. + type: string + required: + - bucket + - credentialsSecretRef + - endpoint + type: object + containerResources: + description: |- + ContainerResources sizes the storage-node container, and expands into the + cluster's own spec.storageNodes.containerResources. + + The container it sizes is the node's management API rather than SPDK, + which runs in a pod of its own: what outgrows the default is a node + answering for many subsystems, not a node moving more data. It is on the + document because a deployment is where a fleet's sizing is decided, and + a cluster written from a document that could not say so had to be edited + afterward on a field the document owns everywhere else. + + Stating either half replaces both. The defaults apply to a cluster that + states neither requests nor limits, so a document stating requests alone + produces a container with no limits rather than one with the default + limits, and a memory limit is what has the kubelet evict a leaking agent + rather than losing the worker. + + It is a pointer because a resource block is a struct, and a struct with + omitempty is serialized whether or not anything is in it: as a value, + every document a discovery run writes would carry an empty + containerResources that says nothing and that a reviewer has to decide + about. + properties: + claims: + description: |- + Claims lists the names of resources, defined in spec.resourceClaims, + that are used by this container. + + This field depends on the + DynamicResourceAllocation feature gate. + + This field is immutable. It can only be set for containers. + items: + description: ResourceClaim references one entry in PodSpec.ResourceClaims. + properties: + name: + description: |- + Name must match the name of one entry in pod.spec.resourceClaims of + the Pod where this field is used. It makes that resource available + inside a container. + type: string + request: + description: |- + Request is the name chosen for a request in the referenced claim. + If empty, everything from the claim is made available, otherwise + only the result of this request. + type: string + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + limits: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Limits describes the maximum amount of compute resources allowed. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + requests: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Requests describes the minimum amount of compute resources required. + If Requests is omitted for a container, it defaults to Limits if that is explicitly specified, + otherwise to an implementation-defined value. Requests cannot exceed Limits. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + type: object + enableAtomicity4K: + description: |- + EnableAtomicity4K enforces 4K write atomicity on every device this + deployment names, which is what lets checksum validation run on devices + whose logical block size is under the data plane's 4K minimum. + + It is the route to checked I/O on a device that cannot be reformatted: a + logical block device's block size is fixed by the drive, and some NVMe + devices offer no 4K format either. Where a device can be reformatted, + EnableDriveFormat is the other route and this is unnecessary. + + It is an enforcement because the question is often unanswerable. A SATA + drive presenting 512-byte logical blocks over a 4K physical sector reports + 512 and nothing more, and a kernel older than 6.11 publishes no atomic + write attributes at all. Where a device does answer, the storage node's + report carries it, and a reviewer approves this against that rather than + against a vendor's datasheet -- because enforcing a guarantee the hardware + does not keep is how a torn write becomes a checksum that silently + disagrees with it. + + It means nothing unless EnableChecksumValidation is set, which is the + cluster's own rule and is left to the cluster to enforce. + type: boolean + enableChecksumValidation: + description: |- + EnableChecksumValidation turns on inline CRC validation of every I/O, for + silent-data-error protection. + + It is on the document because it is immutable on the cluster it lands on: + the backend bakes the checksum method into each device when the cluster is + created and never re-applies it, so a cluster created without this is one + nobody can turn it on for. A deployment that wants its data checked has to + say so here or not at all. + type: boolean + enableDriveFormat: + description: |- + EnableDriveFormat formats every device the document names before a storage + node takes it, which is how a drive carrying anything already is made + usable. + + It says what is wanted rather than how, because the how differs by device + class: an NVMe device is formatted to a 4K block size, and a logical block + device has its signatures wiped. One field covers both, so a document does + not have to know which class the expansion will resolve it to. + + It is on the document rather than defaulted further down because it is + destructive and the document is what somebody approves. A reviewer reading + a draft has to see that the drives it lists will be formatted, and be able + to strike it before approving; the cluster's own field is immutable once + the cluster exists, so a default nobody saw could not be undone either. + type: boolean + enableFailureDomains: + description: |- + EnableFailureDomains opts the cluster into failure-domain mode, in which + every group must label the fault group its workers belong to. + type: boolean + enableJournalDevice: + description: |- + EnableJournalDevice dedicates the smallest NVMe device on each of this + deployment's workers to the journal manager, instead of carving a journal + partition out of every device. + + It is here rather than on a node set because it is immutable on the cluster + it lands on, for the reason SocketsToUse is: the on-disk layout a fleet was + built with is not one a later document can vary. It also costs a drive of + capacity per node, which is a trade a reviewer approves rather than one a + default makes for them. + type: boolean + enableNodeAffinity: + description: |- + EnableNodeAffinity has the data plane serve an erasure-coded volume's I/O + from the local node's own devices where it can, before crossing the + network. + + It is not Kubernetes affinity, and the name is the one place this API + invites that reading: nothing about it schedules a pod, labels a worker, + or places a volume's primary node. The control plane carries it into the + cluster map it pushes to each node, where it sets the local node's index, + and what changes is which copy of a chunk is read. + Co-locating a workload with the primary node of its volume is a separate + mechanism and is not configured here. + + It is on the document because it is immutable on the cluster: the control + plane takes it at cluster create and never re-applies it, so this is the + only moment it can be set at all. + type: boolean + fabricType: + description: FabricType is the storage fabric. + maxLength: 32 + type: string + initContainerResources: + description: |- + InitContainerResources sizes both of the storage node's init containers, + and expands into the cluster's own spec.storageNodes.initContainerResources. + + They are sized apart from the container because they do a different job + and are gone before it starts: one writes the node's env file and the + other runs node_configure.py once, so what they need is a short burst + rather than the footprint of a process that runs for the node's life. + + Stating either half replaces both, as with containerResources, and it is + a pointer for the same reason. + properties: + claims: + description: |- + Claims lists the names of resources, defined in spec.resourceClaims, + that are used by this container. + + This field depends on the + DynamicResourceAllocation feature gate. + + This field is immutable. It can only be set for containers. + items: + description: ResourceClaim references one entry in PodSpec.ResourceClaims. + properties: + name: + description: |- + Name must match the name of one entry in pod.spec.resourceClaims of + the Pod where this field is used. It makes that resource available + inside a container. + type: string + request: + description: |- + Request is the name chosen for a request in the referenced claim. + If empty, everything from the claim is made available, otherwise + only the result of this request. + type: string + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + limits: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Limits describes the maximum amount of compute resources allowed. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + requests: + additionalProperties: + anyOf: + - type: integer + - type: string + pattern: ^(\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))(([KMGTPE]i)|[numkMGTPE]|([eE](\+|-)?(([0-9]+(\.[0-9]*)?)|(\.[0-9]+))))?$ + x-kubernetes-int-or-string: true + description: |- + Requests describes the minimum amount of compute resources required. + If Requests is omitted for a container, it defaults to Limits if that is explicitly specified, + otherwise to an implementation-defined value. Requests cannot exceed Limits. + More info: https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/ + type: object + type: object + kms: + description: |- + KMS selects where the cluster stores volume encryption keys. Stating it on + the document is what makes it present when the cluster is created, where + setting it on the StorageCluster afterward races with that creation. + properties: + vault: + description: Vault stores keys in HashiCorp Vault. + properties: + endpoint: + description: |- + Endpoint is the Vault endpoint, for example, https://vault.example.com:8200. + Rejected unless it resolves to an external address. + pattern: ^https?://[a-zA-Z0-9.-]+(:[0-9]{1,5})?(/.*)?$ + type: string + required: + - endpoint + type: object + type: object + maxSubsystemCount: + description: |- + MaxSubsystemCount is the maximum number of NVMe-oF subsystems each storage + node of this cluster serves. Required, because the StorageCluster's own + field is, and no StorageNode carries a copy of it. + format: int32 + maximum: 75 + minimum: 10 + type: integer + minHugePagesSize: + description: |- + MinHugePagesSize is the smallest huge-page allocation each storage node of + this cluster makes: 100G or 1T, where a bare number is gigabytes. Like + VCPUCount it is the cluster's and is copied onto every node the expansion + writes. Omitted, each node uses the computed minimum. + maxLength: 32 + type: string + name: + description: |- + Name is the StorageCluster's name, and is therefore held to what such a + name may be rather than to what an object name may be. A longer value is a + document the API server accepts and a CreatingCluster step that can never + succeed, since the cluster it would write is one the API server refuses. + maxLength: 63 + type: string + nodeProvisioningBudget: + description: |- + NodeProvisioningBudget is how many workers the expansion may have in the + node-add process at once. It expands into the cluster's own + spec.storageNodes.nodeProvisioningBudget, whose meaning it shares: the cap + is counted by distinct worker, so a two-socket host spends one of the + budget, and a worker hosting a FoundationDB pod is sequential whatever the + budget says. + + It is on the document because a document is what states the size of a + deployment, and a deployment of thirty workers added one at a time is the + difference between an afternoon and a week. Omitted, the cluster's default + of one applies, which is the serial behavior. + format: int32 + minimum: 1 + type: integer + nodesPerSocket: + description: |- + NodesPerSocket is how many storage nodes run per NUMA socket. See + SocketsToUse, which it multiplies. + format: int32 + maximum: 8 + minimum: 1 + type: integer + openshift: + description: |- + OpenShift is what this deployment states because it runs on OpenShift. It + expands into StorageCluster.spec.storageNodes.openshift, whose shape it + shares, and it is read only for a document whose environment is + OpenShift: the environment is what says which distribution this is, and + the block is what that distribution needs said beyond it. + properties: + machineConfigPool: + default: worker + description: |- + MachineConfigPool names a machine-config role the storage nodes' own pool + inherits from, beyond the worker role it always inherits. + + It is not the pool the nodes end up in, which the description it carried + before said and which cost a reader the reboot they were trying to avoid. + Adding a node creates a pool of its own, storage-, and moves the + node into it; a node belongs to exactly one custom pool, so whatever + machine configuration its previous pool carried is lost unless that + pool's role is named here for the new one to select as well. The default + is the role every pool already selects, which is what makes it a no-op + for a fleet whose workers are ordinary workers. + maxLength: 253 + pattern: ^[a-z0-9]([-a-z0-9]*[a-z0-9])?$ + type: string + type: object + ports: + description: |- + Ports are where this cluster's storage nodes listen. Unstated, and for + each member left unstated, the cluster's own defaults decide. + properties: + nodeAgent: + default: 50001 + description: |- + NodeAgent is the port each node's agent API listens on. It expands into + StorageCluster.spec.snodeApiPort, and it is named for the component + rather than for that field: the agent is what spec.images.nodeAgent pins + and what the storage-node DaemonSet runs. + format: int32 + maximum: 65535 + minimum: 1024 + type: integer + nvmf: + default: 4420 + description: |- + NVMf is the base of the NVMe-oF port range every node binds. It expands + into StorageCluster.spec.nvmfBasePort. + format: int32 + maximum: 65535 + minimum: 1024 + type: integer + rpc: + default: 8080 + description: |- + Rpc is the base of the RPC port range every node binds. It expands into + StorageCluster.spec.rpcBasePort. + format: int32 + maximum: 65535 + minimum: 1024 + type: integer + type: object + socketsToUse: + description: |- + SocketsToUse restricts the deployment to selected NUMA sockets, and empty + means socket 0 alone. With NodesPerSocket it decides how many storage nodes + each worker runs, so a group of two workers on a two-socket layout expands + to four nodes. + + It is here rather than on a node set because it is immutable on the cluster + it lands on: the layout a fleet was built with is not one a later document + can vary, and a reviewer should see it before the cluster exists. + items: + maxLength: 16 + type: string + maxItems: 16 + type: array + x-kubernetes-list-type: set + stripe: + description: Stripe is the erasure-coding layout. + properties: + dataChunks: + description: DataChunks is the number of data chunks per + stripe (ndcs). + format: int32 + minimum: 1 + type: integer + parityChunks: + description: |- + ParityChunks is the number of parity chunks per stripe (npcs), and + therefore how many chunk losses a stripe survives. + format: int32 + minimum: 0 + type: integer + type: object + x-kubernetes-validations: + - message: the erasure-coding scheme must be one of 1+0, 1+1, + 2+1, 4+1, 1+2, 2+2, or 4+2, written as dataChunks+parityChunks, + and an unstated half is 1 + rule: '[has(self.dataChunks) ? self.dataChunks : 1, has(self.parityChunks) + ? self.parityChunks : 1] in [[1, 0], [1, 1], [2, 1], [4, + 1], [1, 2], [2, 2], [4, 2]]' + tolerations: + description: |- + Tolerations are what the storage-node pods tolerate, and they expand into + the cluster's own spec.storageNodes.tolerations. + + A fleet that dedicates machines to storage taints them, which is what + keeps everything else off. The DaemonSet that lands on those machines has + to tolerate the taint or it schedules nowhere, and a document that could + not say so described a deployment that does not start: the correction was + an edit to the cluster the document had just created, on a field the + document owns everywhere else. + + A growth document states none. It names a cluster rather than describing + one, and that cluster already carries what its storage nodes tolerate. + items: + description: |- + The pod this Toleration is attached to tolerates any taint that matches + the triple using the matching operator . + properties: + effect: + description: |- + Effect indicates the taint effect to match. Empty means match all taint effects. + When specified, allowed values are NoSchedule, PreferNoSchedule and NoExecute. + type: string + key: + description: |- + Key is the taint key that the toleration applies to. Empty means match all taint keys. + If the key is empty, operator must be Exists; this combination means to match all values and all keys. + type: string + operator: + description: |- + Operator represents a key's relationship to the value. + Valid operators are Exists, Equal, Lt, and Gt. Defaults to Equal. + Exists is equivalent to wildcard for value, so that a pod can + tolerate all taints of a particular category. + Lt and Gt perform numeric comparisons (requires feature gate TaintTolerationComparisonOperators). + type: string + tolerationSeconds: + description: |- + TolerationSeconds represents the period of time the toleration (which must be + of effect NoExecute, otherwise this field is ignored) tolerates the taint. By default, + it is not set, which means tolerate the taint forever (do not evict). Zero and + negative values will be treated as 0 (evict immediately) by the system. + format: int64 + type: integer + value: + description: |- + Value is the taint value the toleration matches to. + If the operator is Exists, the value should be empty, otherwise just a regular string. + type: string + type: object + maxItems: 32 + type: array + vcpuCount: + description: |- + VCPUCount is the number of vCPUs allocated to SPDK on each storage node of + this cluster. It is stated here and nowhere below, because the control + plane assumes it uniform across a cluster's nodes; CreatingNodes copies it + into every StorageNode.spec.config.sizing it writes. Required, because the + StorageCluster's own field is. + The floor is 4 rather than a hardware limit: a node must carry one core + beyond this budget for the system, and the control plane's core layout + assigns no NVMe-oF poller core at all for a 2-vCPU budget. + format: int32 + minimum: 4 + type: integer + required: + - maxSubsystemCount + - name + - vcpuCount + type: object + message: + description: |- + Message is what the site says about the draft: validation findings while + it is a draft, the expansion's step afterwards. + type: string + name: + description: Name is the ClusterDeploymentConfig on the site. + type: string + nodeRefs: + description: NodeRefs are the StorageNode objects the expansion + created. + items: + type: string + type: array + x-kubernetes-list-type: set + nodeSets: + description: NodeSets are the nodes and devices the discovery + found, for review. + items: + description: |- + NodeSet is the organizational grouping of a deployment, usually a rack: the + workers a document adds or grows together. It carries no sizing, because sizing + is uniform across a cluster and is stated once in ClusterTemplate. + properties: + groups: + description: Groups are the sets of workers sharing one + configuration. + items: + description: |- + NodeGroup is a set of workers that share one configuration, which is what + makes ten identical machines one entry rather than ten. + properties: + dataInterfaces: + description: DataInterfaces are the data-plane network + interfaces. + items: + maxLength: 63 + type: string + maxItems: 32 + type: array + devices: + description: Devices selects the storage devices every + worker in the group uses. + properties: + block: + description: |- + Block names logical block devices by path ("/dev/sdb"). It expands into the + same config.deviceNames as NVMe, which takes a PCI address and a device + path in one list. It is the alternative to NVMe rather than a companion of + it: the two classes are not mixed within a cluster. + items: + maxLength: 255 + pattern: ^/dev/[a-zA-Z0-9._/-]+$ + type: string + maxItems: 128 + type: array + x-kubernetes-list-type: set + nvme: + description: NVMe names NVMe devices by PCI address + ("0000:5e:00.0"). + items: + maxLength: 32 + pattern: ^[0-9a-fA-F]{4}:[0-9a-fA-F]{2}:[0-9a-fA-F]{2}\.[0-9a-fA-F]$ + type: string + maxItems: 128 + type: array + x-kubernetes-list-type: set + type: object + x-kubernetes-validations: + - message: a device selection names NVMe addresses + or block devices, not both + rule: has(self.nvme) != has(self.block) + failureDomain: + description: |- + FailureDomain is the label of the fault group every worker in this group + belongs to ("rack-b"), which is usually the name of the rack, zone, or + power feed they share. Discovery seeds it from topology.kubernetes.io/zone + and leaves it unset where the Kubernetes API carries no topology, which + holds provisioning with a clear reason rather than guessing. It expands + into StorageNode.spec.config.failureDomain, whose shape it shares. + maxLength: 63 + pattern: ^[a-zA-Z0-9]([-_.a-zA-Z0-9]*[a-zA-Z0-9])?$ + type: string + journalManager: + description: JournalManager tunes the journal managers + on these nodes. + properties: + count: + description: Count is the number of journal managers + to configure. + format: int32 + minimum: 1 + type: integer + percentPerDevice: + description: PercentPerDevice is the share of + each device given to the journal. + format: int32 + maximum: 100 + minimum: 1 + type: integer + type: object + mgmtInterface: + description: MgmtInterface is the management network + interface the storage nodes bind. + maxLength: 63 + type: string + name: + description: |- + Name identifies the group within its node set, for a reader and for the + events a validation failure emits. + maxLength: 253 + type: string + reservedSystemCPU: + description: |- + ReservedSystemCPU is the CPU set held back from SPDK for the system on + these nodes, as a core list such as 0,1 or 0-3. + + It is a group's rather than the cluster's because it names core ids, and a + group is what a document calls the workers that share their hardware: 0,1 + on a sixteen-core worker and 0,1 on a ninety-six-core worker are different + fractions of the machine. It expands into + StorageNode.spec.config.reservedSystemCPU, whose shape it shares, and a + group that states none leaves the cluster's fleet-wide value to decide. + + On OpenShift it reaches the kubelet through a KubeletConfig for the + machine config pool, which is the cluster's, so groups that disagree there + are writing over one another's pool configuration. + maxLength: 63 + pattern: ^[0-9]+(-[0-9]+)?(,[0-9]+(-[0-9]+)?)*$ + type: string + spdkSystemMemory: + description: |- + SpdkSystemMemory is the memory the control plane starts SPDK with on these + nodes. + maxLength: 32 + pattern: ^[0-9]+(G|GI|GB|GiB|M|MI|MB|MiB|g|gi|gb|gib|m|mi|mb|mib)?$ + type: string + workers: + description: Workers are the Kubernetes worker hostnames + in this group. + items: + maxLength: 253 + type: string + maxItems: 200 + minItems: 1 + type: array + x-kubernetes-list-type: set + required: + - name + - workers + type: object + maxItems: 64 + minItems: 1 + type: array + name: + description: |- + Name is the node set's name. It is copied to StorageNode.spec.nodeSet, so + that a node can be traced back to the part of the document that produced + it. + maxLength: 253 + type: string + required: + - groups + - name + type: object + type: array + phase: + description: |- + Phase is the draft's own phase on the site (Draft, Expanding, Expanded, + Failed). + type: string + required: + - name + type: object + message: + description: |- + Message is the reason the phase is what it is: one sentence, replaced as + the request moves, and never a log. + type: string + observedGeneration: + description: |- + ObservedGeneration is the generation the rest of this status was computed + from. + format: int64 + type: integer + phase: + description: Phase is the request's own progress. + enum: + - Pending + - Discovering + - Drafted + - Deploying + - Online + - Failed + type: string + storageCluster: + description: StorageCluster is the cluster the approved draft produced. + properties: + name: + description: Name is the StorageCluster object on the site. + type: string + nodes: + description: Nodes are the cluster's storage nodes. + items: + description: |- + StorageSiteNode is one storage node of the deployed cluster, as the site + reports it. + properties: + hostname: + description: Hostname is the Kubernetes node it runs on. + type: string + name: + description: Name is the StorageNode object on the site. + type: string + phase: + description: Phase is the node's phase on the site. + type: string + required: + - name + type: object + type: array + x-kubernetes-list-map-keys: + - name + x-kubernetes-list-type: map + phase: + description: Phase is the StorageCluster's phase on the site. + type: string + pool: + description: |- + Pool is the pool the cluster was created with, which a StorageClass names + in pool_name. + type: string + uuid: + description: |- + UUID is the storage cluster's id in the control plane, which a + StorageClass names in cluster_id. + type: string + required: + - name + type: object + workName: + description: WorkName is the ManifestWork carrying the request to + the site. + type: string + type: object + type: object + served: true + storage: true + subresources: + status: {} +--- +apiVersion: apiextensions.k8s.io/v1 +kind: CustomResourceDefinition metadata: annotations: controller-gen.kubebuilder.io/version: v0.21.0 @@ -12270,6 +13327,14 @@ rules: - patch - update - watch +- apiGroups: + - cluster.open-cluster-management.io + resources: + - managedclusters + verbs: + - get + - list + - watch - apiGroups: - coordination.k8s.io resources: @@ -12419,6 +13484,7 @@ rules: - storagenodes - storagepoolops - storagepools + - storagesitedeployments - tasks - testfailovers - volumemigrations @@ -12455,6 +13521,7 @@ rules: - storagenodes/finalizers - storagepoolops/finalizers - storagepools/finalizers + - storagesitedeployments/finalizers - tasks/finalizers - testfailovers/finalizers - volumemigrations/finalizers @@ -12488,6 +13555,7 @@ rules: - storagenodesets/status - storagepoolops/status - storagepools/status + - storagesitedeployments/status - tasks/status - testfailovers/status - volumegroupsnapshotops/status From 81cdbf5f081d2c54792d8b1fee1bf0d66780a456 Mon Sep 17 00:00:00 2001 From: michael Date: Sat, 3 Oct 2026 17:49:20 +0300 Subject: [PATCH 183/206] operator: the CSV owns StorageSiteDeployment Co-Authored-By: Claude Opus 5.5 --- .../bases/simplyblock-operator.clusterserviceversion.yaml | 8 ++++++++ 1 file changed, 8 insertions(+) diff --git a/operator/config/manifests/bases/simplyblock-operator.clusterserviceversion.yaml b/operator/config/manifests/bases/simplyblock-operator.clusterserviceversion.yaml index fc70292e9..ced5c9b1d 100644 --- a/operator/config/manifests/bases/simplyblock-operator.clusterserviceversion.yaml +++ b/operator/config/manifests/bases/simplyblock-operator.clusterserviceversion.yaml @@ -601,6 +601,14 @@ spec: kind: TestFailover name: testfailovers.storage.simplyblock.io version: v1alpha2 + - description: Deploys a managed site's storage from the hub, through OCM. + Discovers the site's nodes, applies the requested sizing to the draft and, + once approved, deploys the storage cluster; status projects the draft and + the resulting storage cluster. + displayName: Storage Site Deployment + kind: StorageSiteDeployment + name: storagesitedeployments.storage.simplyblock.io + version: v1alpha2 description: The Simplyblock Operator helps with installation, operation, and management of Simplyblock Control Planes, Storage Planes, and the CSI Driver. displayName: Simplyblock Operator From c6579ab8abc6a63ab53d3332e730915034e34b71 Mon Sep 17 00:00:00 2001 From: michael Date: Sat, 3 Oct 2026 21:10:44 +0300 Subject: [PATCH 184/206] chart: the console's operator-API upstream is empty when the release has no operator API Unset, controlCenter.operatorUrl defaulted to http://simplyblock-operator:8080 whether or not that Service exists. On a release without it nginx exits with "host not found in upstream" and the console crash-loops (2026-10-03, the DR test bed on integrate_csi_addons). The default is now that Service only when it exists in the release namespace, and empty otherwise; the console then answers 503 on the operator-backed screens and serves everything else. Co-Authored-By: Claude Opus 5.5 --- .../templates/_control_center_helpers.tpl | 12 +++++++++++- 1 file changed, 11 insertions(+), 1 deletion(-) diff --git a/helm-charts/charts/simplyblock-operator/templates/_control_center_helpers.tpl b/helm-charts/charts/simplyblock-operator/templates/_control_center_helpers.tpl index 97a31717e..e7ed59f80 100644 --- a/helm-charts/charts/simplyblock-operator/templates/_control_center_helpers.tpl +++ b/helm-charts/charts/simplyblock-operator/templates/_control_center_helpers.tpl @@ -47,8 +47,18 @@ helm.sh/chart: {{ printf "%s-%s" .Name .Version | replace "+" "_" | trunc 63 | t {{/* Default upstream URLs. This chart names its objects statically (simplyblock-operator, simplyblock-prometheus, …) rather than deriving them from the release, so the fallbacks here are static too. */}} +{{/* The operator's HTTP API, when this release serves one. Unset, it is the + simplyblock-operator Service if that exists in the release namespace, and + empty otherwise: nginx refuses to start on an upstream host it cannot + resolve ("host not found in upstream"), and an empty upstream makes the + console answer 503 "not part of this deployment" on the operator-backed + screens instead (2026-10-03, a release without the operator API). */}} {{- define "sbcc.operatorUrl" -}} -{{- .Values.controlCenter.operatorUrl | default "http://simplyblock-operator:8080" -}} +{{- if .Values.controlCenter.operatorUrl -}} +{{- .Values.controlCenter.operatorUrl -}} +{{- else if (lookup "v1" "Service" .Release.Namespace "simplyblock-operator") -}} +http://simplyblock-operator:8080 +{{- end -}} {{- end -}} {{- define "sbcc.prometheusUrl" -}} From b71f0fa1a11d0dd8b7f860b4fda70d27d1eac5f6 Mon Sep 17 00:00:00 2001 From: michael Date: Sat, 3 Oct 2026 23:18:39 +0300 Subject: [PATCH 185/206] csi-driver: a late consistency-group join live-migrates the volume to the group's node first A PVC labeled into a consistency group while its volume lives off the group's pinned node/LVS was refused for good: a join never moves a volume. The membership watcher now asks the backend for the join plan (POST .../members/plan) and, when placement is the only obstacle, requests the move as a VolumeMigration (cg-join-, spec.pvName + targetNodeUUID): the operator attaches the target paths on the consumer host and runs the live migration, then the next pass joins. Failed or aborted migrations, and joins that can never succeed, are Warning events. After a join the watcher can move the member into its group's subsystem (POST .../members/{id}/colocate); off by default (SPDKCSI_CG_COLOCATE), and a refusal is a Normal event. SPDKCSI_CG_PREJOIN_MIGRATION (default on) and SPDKCSI_CG_CLIENT_SWAP_READY control the rest. Co-Authored-By: Claude Opus 5.5 --- .../controlplane/consistency_group.go | 76 ++++++++ .../csi/controller/cg_membership_watcher.go | 133 +++++++++++++- .../controller/cg_membership_watcher_test.go | 163 ++++++++++++++++++ .../csi/controller/cg_prejoin_migration.go | 104 +++++++++++ .../controller/cg_prejoin_migration_test.go | 72 ++++++++ csi-driver/internal/driver/driver.go | 11 +- 6 files changed, 552 insertions(+), 7 deletions(-) create mode 100644 csi-driver/internal/csi/controller/cg_prejoin_migration.go create mode 100644 csi-driver/internal/csi/controller/cg_prejoin_migration_test.go diff --git a/csi-driver/internal/controlplane/consistency_group.go b/csi-driver/internal/controlplane/consistency_group.go index 72e7b14f0..eebedeaea 100644 --- a/csi-driver/internal/controlplane/consistency_group.go +++ b/csi-driver/internal/controlplane/consistency_group.go @@ -283,3 +283,79 @@ func (c *ClusterClient) JoinConsistencyGroupMember(ctx context.Context, groupID, func (c *ClusterClient) DetachConsistencyGroupMember(ctx context.Context, groupID, lvolID string) error { return c.API.detachConsistencyGroupMember(ctx, groupUUID(groupID), lvolID) } + +// Join plan step names (sbcli cg_colocation, docs/consistency-group-colocation.md). +const ( + JoinStepMigrate = "migrate" + JoinStepJoin = "join" + JoinStepColocate = "colocate" +) + +// ErrColocationRefused wraps a backend 409 on a co-location step: namespace +// moves are disabled, or a host is connected and the client cannot swap +// paths. Like a membership refusal it is a standing state, not a fault. +var ErrColocationRefused = errors.New("consistency-group co-location refused") + +// JoinPlan is what joining an existing volume to a group takes: a live +// migration of MigrateLvolIDs to TargetNodeID first (the volume is off the +// group's pinned node/LVS), the join, and a namespace move into TargetNQN. +type JoinPlan struct { + Steps []string `json:"steps"` + TargetNodeID string `json:"target_node_id"` + MigrateLvolIDs []string `json:"migrate_lvol_ids"` + TargetNQN string `json:"target_nqn"` +} + +// Has reports whether the plan contains step. +func (p *JoinPlan) Has(step string) bool { + for _, s := range p.Steps { + if s == step { + return true + } + } + return false +} + +func (client APIClient) planConsistencyGroupJoin(ctx context.Context, gid, lvolID string) (*JoinPlan, error) { + body := map[string]string{"lvol_id": lvolID} + raw, err := client.do(ctx, http.MethodPost, client.v2consistencyGroupMembers(gid)+"/plan", body) + if err != nil { + if isHTTPStatus(err, http.StatusConflict) { + return nil, fmt.Errorf("%w: %s", ErrMembershipRefused, err.Error()) + } + return nil, err + } + var plan JoinPlan + if err := json.Unmarshal(raw, &plan); err != nil { + return nil, fmt.Errorf("unexpected response for join plan: %w", err) + } + return &plan, nil +} + +func (client APIClient) colocateConsistencyGroupMember( + ctx context.Context, gid, lvolID string, clientSwapReady bool, +) error { + body := map[string]bool{"client_swap_ready": clientSwapReady} + _, err := client.do(ctx, http.MethodPost, client.v2consistencyGroupMember(gid, lvolID)+"/colocate", body) + if err != nil && isHTTPStatus(err, http.StatusConflict) { + return fmt.Errorf("%w: %s", ErrColocationRefused, err.Error()) + } + return err +} + +// PlanConsistencyGroupJoin asks the backend which steps joining lvolID to the +// group takes, without taking any. A join that can never succeed (another +// group, another pool, another group's subsystem siblings) is +// ErrMembershipRefused. +func (c *ClusterClient) PlanConsistencyGroupJoin(ctx context.Context, groupID, lvolID string) (*JoinPlan, error) { + return c.API.planConsistencyGroupJoin(ctx, groupUUID(groupID), lvolID) +} + +// ColocateConsistencyGroupMember moves a member's namespace into its group's +// subsystem. clientSwapReady asserts the client stages the volume behind the +// device-mapper indirection and swaps paths itself. +func (c *ClusterClient) ColocateConsistencyGroupMember( + ctx context.Context, groupID, lvolID string, clientSwapReady bool, +) error { + return c.API.colocateConsistencyGroupMember(ctx, groupUUID(groupID), lvolID, clientSwapReady) +} diff --git a/csi-driver/internal/csi/controller/cg_membership_watcher.go b/csi-driver/internal/csi/controller/cg_membership_watcher.go index 96b8aaba2..28d8d50f5 100644 --- a/csi-driver/internal/csi/controller/cg_membership_watcher.go +++ b/csi-driver/internal/csi/controller/cg_membership_watcher.go @@ -16,6 +16,7 @@ import ( corev1 "k8s.io/api/core/v1" metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/client-go/dynamic" "k8s.io/client-go/informers" "k8s.io/client-go/kubernetes" "k8s.io/client-go/kubernetes/scheme" @@ -42,6 +43,17 @@ type membershipClient interface { GetConsistencyGroup(ctx context.Context, groupID string) (*controlplane.ConsistencyGroupSummary, error) JoinConsistencyGroupMember(ctx context.Context, groupID, lvolID string) error DetachConsistencyGroupMember(ctx context.Context, groupID, lvolID string) error + PlanConsistencyGroupJoin(ctx context.Context, groupID, lvolID string) (*controlplane.JoinPlan, error) + ColocateConsistencyGroupMember(ctx context.Context, groupID, lvolID string, clientSwapReady bool) error +} + +// migrationRequester asks the operator to live-migrate a PersistentVolume's +// backing volume to a storage node, through a VolumeMigration: the operator +// attaches the target paths on the consumer host before the backend moves the +// data, which only it can do. Ensure creates the request once (by name) and +// reports its phase; "" while the operator has not picked it up. +type migrationRequester interface { + Ensure(ctx context.Context, name, pvName, targetNodeUUID string) (string, error) } // cgMembershipWatcher reconciles PVC label state to backend group membership. @@ -52,12 +64,29 @@ type cgMembershipWatcher struct { // pool; production wires clusters.Client, tests substitute a stub. clientFor func(ctx context.Context, clusterID, poolRef string) (membershipClient, error) recorder record.EventRecorder + // migrations runs the pre-join live migration of a volume off its group's + // pinned node (co-location design §5); nil disables it, and an off-pin + // join stays refused as before. + migrations migrationRequester + // colocate moves a member into its group's subsystem after the join; off + // by default, because the backend refuses the move for an attached volume + // until the node plugin swaps paths (design §6). + colocate bool + // clientSwapReady asserts every node stages volumes behind the + // device-mapper indirection and swaps paths itself (design §6). + clientSwapReady bool } // StartConsistencyGroupLabelWatcher runs the membership watcher until ctx is // canceled. It is a no-op (with a log line) when kube is nil, mirroring how // the annotation helpers degrade without an in-cluster config. -func StartConsistencyGroupLabelWatcher(ctx context.Context, kube kubernetes.Interface, driverName string) { +// +// dyn, when non-nil, lets the watcher request the pre-join live migration of a +// volume off its group's pinned node (a VolumeMigration); without it an +// off-pin join stays refused. +func StartConsistencyGroupLabelWatcher( + ctx context.Context, kube kubernetes.Interface, dyn dynamic.Interface, driverName string, +) { if kube == nil { klog.Warning("consistency-group label watcher disabled: no Kubernetes client") return @@ -72,6 +101,12 @@ func StartConsistencyGroupLabelWatcher(ctx context.Context, kube kubernetes.Inte }, recorder: broadcaster.NewRecorder(scheme.Scheme, corev1.EventSource{Component: "spdkcsi-cg-membership"}), } + preJoin, colocate, swapReady := watcherOptions() + if preJoin && dyn != nil { + watcher.migrations = newVolumeMigrations(dyn) + } + watcher.colocate = colocate + watcher.clientSwapReady = swapReady factory := informers.NewSharedInformerFactory(kube, membershipResync) informer := factory.Core().V1().PersistentVolumeClaims().Informer() @@ -136,7 +171,7 @@ func (w *cgMembershipWatcher) reconcile(ctx context.Context, pvc *corev1.Persist switch { case label != "" && groupID == "": - return w.join(ctx, client, pvc, label, handle.VolumeID) + return w.join(ctx, client, pvc, pv.Name, label, handle.VolumeID) case label == "" && groupID != "": return w.detach(ctx, client, pvc, groupID, handle.VolumeID) case label != "" && groupID != "": @@ -152,14 +187,20 @@ func (w *cgMembershipWatcher) reconcile(ctx context.Context, pvc *corev1.Persist "volume is a member of consistency group %q but the PVC is labeled %q; "+ "membership is one-way, so relabeling cannot move a volume between groups", group.Name, label) + return nil } + return w.colocateMember(ctx, client, pvc, groupID, label, handle.VolumeID) } return nil } +// preJoinMigrationName is the VolumeMigration a late join requests for a +// volume: one per volume, so a resync finds the request it already made. +func preJoinMigrationName(lvolID string) string { return "cg-join-" + lvolID } + func (w *cgMembershipWatcher) join( ctx context.Context, client membershipClient, - pvc *corev1.PersistentVolumeClaim, label, lvolID string, + pvc *corev1.PersistentVolumeClaim, pvName, label, lvolID string, ) error { group, err := client.ResolveConsistencyGroupByName(ctx, label) if err != nil { @@ -174,17 +215,97 @@ func (w *cgMembershipWatcher) join( } if err := client.JoinConsistencyGroupMember(ctx, group.ID, lvolID); err != nil { if errors.Is(err, controlplane.ErrMembershipRefused) { - w.recorder.Eventf(pvc, corev1.EventTypeWarning, "ConsistencyGroupJoinRefused", - "volume cannot join consistency group %q: %v", label, err) - return nil + return w.joinRefused(ctx, client, pvc, pvName, label, group.ID, lvolID, err) } return fmt.Errorf("join consistency group %q: %w", label, err) } w.recorder.Eventf(pvc, corev1.EventTypeNormal, "ConsistencyGroupJoined", "volume joined consistency group %q; it is included from the next generation", label) + return w.colocateMember(ctx, client, pvc, group.ID, label, lvolID) +} + +// joinRefused turns a refused join into the pre-join migration when the only +// obstacle is placement: the backend's plan says the volume (with its +// subsystem siblings) must first move to the group's pinned node. Anything +// else stays a refusal on the PVC. +func (w *cgMembershipWatcher) joinRefused( + ctx context.Context, client membershipClient, pvc *corev1.PersistentVolumeClaim, + pvName, label, groupID, lvolID string, refusal error, +) error { + refuse := func(err error) error { + w.recorder.Eventf(pvc, corev1.EventTypeWarning, "ConsistencyGroupJoinRefused", + "volume cannot join consistency group %q: %v", label, err) + return nil + } + if w.migrations == nil { + return refuse(refusal) + } + plan, err := client.PlanConsistencyGroupJoin(ctx, groupID, lvolID) + if err != nil { + if errors.Is(err, controlplane.ErrMembershipRefused) { + return refuse(err) + } + return fmt.Errorf("plan the join to consistency group %q: %w", label, err) + } + if !plan.Has(controlplane.JoinStepMigrate) || plan.TargetNodeID == "" { + return refuse(refusal) + } + name := preJoinMigrationName(lvolID) + phase, err := w.migrations.Ensure(ctx, name, pvName, plan.TargetNodeID) + if err != nil { + return fmt.Errorf("request the pre-join migration of %s: %w", pvName, err) + } + switch phase { + case migrationCompleted: + // Moved: the placement precondition holds now. The resync or the next + // PVC event joins; joining here would race the backend's record switch + // of the last subsystem sibling. + w.recorder.Eventf(pvc, corev1.EventTypeNormal, "ConsistencyGroupMigrated", + "volume moved to the group's node %s; joining consistency group %q on the next pass", + plan.TargetNodeID, label) + case migrationFailed, migrationAborted: + w.recorder.Eventf(pvc, corev1.EventTypeWarning, "ConsistencyGroupMigrationFailed", + "the migration to consistency group %q's node %s ended %s; delete VolumeMigration %s to retry", + label, plan.TargetNodeID, phase, name) + default: + w.recorder.Eventf(pvc, corev1.EventTypeNormal, "ConsistencyGroupMigrating", + "volume is off consistency group %q's node: migrating it (with %d volume(s) of its "+ + "subsystem) to node %s through VolumeMigration %s before the join", label, + len(plan.MigrateLvolIDs), plan.TargetNodeID, name) + } return nil } +// The VolumeMigration phases the watcher acts on (operator api/v1alpha1). +const ( + migrationCompleted = "Completed" + migrationFailed = "Failed" + migrationAborted = "Aborted" +) + +// colocateMember moves a current member into its group's subsystem when the +// watcher is configured to; a refusal (moves disabled, or a connected host and +// no client swap) is a Normal event, since the member is valid where it is. +func (w *cgMembershipWatcher) colocateMember( + ctx context.Context, client membershipClient, pvc *corev1.PersistentVolumeClaim, + groupID, label, lvolID string, +) error { + if !w.colocate { + return nil + } + err := client.ColocateConsistencyGroupMember(ctx, groupID, lvolID, w.clientSwapReady) + switch { + case err == nil: + return nil + case errors.Is(err, controlplane.ErrColocationRefused): + w.recorder.Eventf(pvc, corev1.EventTypeNormal, "ConsistencyGroupColocationDeferred", + "volume stays in its own subsystem for now (consistency group %q): %v", label, err) + return nil + default: + return fmt.Errorf("co-locate with consistency group %q: %w", label, err) + } +} + func (w *cgMembershipWatcher) detach( ctx context.Context, client membershipClient, pvc *corev1.PersistentVolumeClaim, groupID, lvolID string, diff --git a/csi-driver/internal/csi/controller/cg_membership_watcher_test.go b/csi-driver/internal/csi/controller/cg_membership_watcher_test.go index ba9f958eb..850daa973 100644 --- a/csi-driver/internal/csi/controller/cg_membership_watcher_test.go +++ b/csi-driver/internal/csi/controller/cg_membership_watcher_test.go @@ -32,6 +32,45 @@ type fakeMembership struct { joined [][2]string // {groupID, lvolID} detached [][2]string + + plan *controlplane.JoinPlan + planErr error + colocateErr error + colocated [][2]string + swapReady []bool +} + +func (f *fakeMembership) PlanConsistencyGroupJoin( + _ context.Context, _, _ string, +) (*controlplane.JoinPlan, error) { + if f.planErr != nil { + return nil, f.planErr + } + if f.plan == nil { + return &controlplane.JoinPlan{Steps: []string{controlplane.JoinStepJoin}}, nil + } + return f.plan, nil +} + +func (f *fakeMembership) ColocateConsistencyGroupMember( + _ context.Context, groupID, lvolID string, clientSwapReady bool, +) error { + f.colocated = append(f.colocated, [2]string{groupID, lvolID}) + f.swapReady = append(f.swapReady, clientSwapReady) + return f.colocateErr +} + +// fakeMigrations records the VolumeMigrations the watcher requests and +// reports a scripted phase for them. +type fakeMigrations struct { + phase string + err error + requests [][3]string // {name, pvName, target} +} + +func (f *fakeMigrations) Ensure(_ context.Context, name, pvName, target string) (string, error) { + f.requests = append(f.requests, [3]string{name, pvName, target}) + return f.phase, f.err } func (f *fakeMembership) GetVolumeGroupID(_ context.Context, lvolID string) (string, error) { @@ -253,3 +292,127 @@ func TestConflictingLabelIsSurfacedNotActedOn(t *testing.T) { } requireEvent(t, recorder, "ConsistencyGroupConflict") } + +func offPinFixture(t *testing.T, phase string) ( + *cgMembershipWatcher, *corev1.PersistentVolumeClaim, *record.FakeRecorder, + *fakeMembership, *fakeMigrations, +) { + t.Helper() + membership := &fakeMembership{ + groupIDByLvol: map[string]string{}, + groupsByName: map[string]*controlplane.ConsistencyGroupSummary{ + "db-group": {ID: "gid-1", Name: "db-group"}, + }, + joinErr: controlplane.ErrMembershipRefused, + plan: &controlplane.JoinPlan{ + Steps: []string{controlplane.JoinStepMigrate, controlplane.JoinStepJoin}, + TargetNodeID: "node-pin", + MigrateLvolIDs: []string{testLvol, "sibling"}, + }, + } + migrations := &fakeMigrations{phase: phase} + watcher, pvc, recorder := watcherFixture(t, + map[string]string{consistencyGroupLabel: "db-group"}, testDriver, membership) + watcher.migrations = migrations + return watcher, pvc, recorder, membership, migrations +} + +func TestOffPinJoinRequestsThePreJoinMigration(t *testing.T) { + watcher, pvc, recorder, membership, migrations := offPinFixture(t, "") + if err := watcher.reconcile(context.Background(), pvc); err != nil { + t.Fatalf("reconcile: %v", err) + } + if len(migrations.requests) != 1 || + migrations.requests[0] != [3]string{"cg-join-" + testLvol, "pv-data", "node-pin"} { + t.Fatalf("expected one VolumeMigration of pv-data to node-pin, got %v", migrations.requests) + } + if len(membership.joined) != 0 { + t.Fatalf("no join may happen before the volume is on the pin, got %v", membership.joined) + } + requireEvent(t, recorder, "ConsistencyGroupMigrating") +} + +func TestACompletedPreJoinMigrationIsReportedAndTheJoinFollowsOnTheNextPass(t *testing.T) { + watcher, pvc, recorder, _, migrations := offPinFixture(t, migrationCompleted) + if err := watcher.reconcile(context.Background(), pvc); err != nil { + t.Fatalf("reconcile: %v", err) + } + requireEvent(t, recorder, "ConsistencyGroupMigrated") + if len(migrations.requests) != 1 { + t.Fatalf("the existing request is read, not duplicated: %v", migrations.requests) + } +} + +func TestAFailedPreJoinMigrationIsAWarning(t *testing.T) { + watcher, pvc, recorder, _, _ := offPinFixture(t, migrationFailed) + if err := watcher.reconcile(context.Background(), pvc); err != nil { + t.Fatalf("reconcile: %v", err) + } + requireEvent(t, recorder, "ConsistencyGroupMigrationFailed") +} + +func TestAJoinThatCanNeverSucceedIsNotMigrated(t *testing.T) { + watcher, pvc, recorder, membership, migrations := offPinFixture(t, "") + membership.planErr = controlplane.ErrMembershipRefused + if err := watcher.reconcile(context.Background(), pvc); err != nil { + t.Fatalf("reconcile: %v", err) + } + if len(migrations.requests) != 0 { + t.Fatalf("a refused plan must not request a migration: %v", migrations.requests) + } + requireEvent(t, recorder, "ConsistencyGroupJoinRefused") +} + +func TestWithoutTheMigrationRequesterAnOffPinJoinStaysRefused(t *testing.T) { + watcher, pvc, recorder, _, _ := offPinFixture(t, "") + watcher.migrations = nil + if err := watcher.reconcile(context.Background(), pvc); err != nil { + t.Fatalf("reconcile: %v", err) + } + requireEvent(t, recorder, "ConsistencyGroupJoinRefused") +} + +func TestColocationRunsAfterAJoinOnlyWhenEnabled(t *testing.T) { + membership := &fakeMembership{ + groupIDByLvol: map[string]string{}, + groupsByName: map[string]*controlplane.ConsistencyGroupSummary{ + "db-group": {ID: "gid-1", Name: "db-group"}, + }, + } + watcher, pvc, recorder := watcherFixture(t, + map[string]string{consistencyGroupLabel: "db-group"}, testDriver, membership) + if err := watcher.reconcile(context.Background(), pvc); err != nil { + t.Fatalf("reconcile: %v", err) + } + if len(membership.colocated) != 0 { + t.Fatalf("co-location is off by default, got %v", membership.colocated) + } + requireEvent(t, recorder, "ConsistencyGroupJoined") + + membership.joined = nil + watcher.colocate, watcher.clientSwapReady = true, true + if err := watcher.reconcile(context.Background(), pvc); err != nil { + t.Fatalf("reconcile: %v", err) + } + if len(membership.colocated) != 1 || !membership.swapReady[0] { + t.Fatalf("expected one co-location with the client swap asserted, got %v %v", + membership.colocated, membership.swapReady) + } +} + +func TestARefusedColocationIsANormalEvent(t *testing.T) { + membership := &fakeMembership{ + groupIDByLvol: map[string]string{testLvol: "gid-1"}, + groupsByID: map[string]*controlplane.ConsistencyGroupSummary{ + "gid-1": {ID: "gid-1", Name: "db-group"}, + }, + colocateErr: controlplane.ErrColocationRefused, + } + watcher, pvc, recorder := watcherFixture(t, + map[string]string{consistencyGroupLabel: "db-group"}, testDriver, membership) + watcher.colocate = true + if err := watcher.reconcile(context.Background(), pvc); err != nil { + t.Fatalf("a refused co-location is not an error: %v", err) + } + requireEvent(t, recorder, "ConsistencyGroupColocationDeferred") +} diff --git a/csi-driver/internal/csi/controller/cg_prejoin_migration.go b/csi-driver/internal/csi/controller/cg_prejoin_migration.go new file mode 100644 index 000000000..156cb687d --- /dev/null +++ b/csi-driver/internal/csi/controller/cg_prejoin_migration.go @@ -0,0 +1,104 @@ +// The pre-join live migration (docs/consistency-group-colocation.md §5, in +// sbcli): a volume whose PVC gets the consistency-group label while it lives +// off the group's pinned node/LVS is moved there first, then joined. +// +// The move is requested as a VolumeMigration, not made against the control +// plane directly: a live migration needs the consumer host attached to the +// target's paths before the backend moves the data, and attaching them is the +// operator's VolumeMigration controller's job. The CSI driver only asks for +// the move and reads how it went, through a dynamic client, so it does not +// depend on the operator's API types. +package controller + +import ( + "context" + "fmt" + "os" + "strings" + + apierrors "k8s.io/apimachinery/pkg/api/errors" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" + "k8s.io/apimachinery/pkg/runtime/schema" + "k8s.io/client-go/dynamic" +) + +var volumeMigrationGVR = schema.GroupVersionResource{ + Group: "storage.simplyblock.io", Version: "v1alpha1", Resource: "volumemigrations", +} + +// preJoinPurposeLabel marks the VolumeMigrations the watcher creates, so they +// are told apart from the rebalancer's and an operator's own. +const preJoinPurposeLabel = "storage.simplyblock.io/purpose" + +// volumeMigrations creates VolumeMigrations in the driver's namespace. +type volumeMigrations struct { + client dynamic.Interface + namespace string +} + +func newVolumeMigrations(client dynamic.Interface) *volumeMigrations { + ns := strings.TrimSpace(os.Getenv("POD_NAMESPACE")) + if ns == "" { + ns = "simplyblock" + } + return &volumeMigrations{client: client, namespace: ns} +} + +// Ensure creates the named VolumeMigration when it does not exist and returns +// its status.phase. An existing one is never rewritten: a request for another +// target (the group re-pinned meanwhile) waits until the old one is deleted. +func (m *volumeMigrations) Ensure(ctx context.Context, name, pvName, targetNodeUUID string) (string, error) { + res := m.client.Resource(volumeMigrationGVR).Namespace(m.namespace) + obj, err := res.Get(ctx, name, metav1.GetOptions{}) + if apierrors.IsNotFound(err) { + obj = &unstructured.Unstructured{Object: map[string]any{ + "apiVersion": "storage.simplyblock.io/v1alpha1", + "kind": "VolumeMigration", + "metadata": map[string]any{ + "name": name, + "namespace": m.namespace, + "labels": map[string]any{ + preJoinPurposeLabel: "consistency-group-join", + "app.kubernetes.io/managed-by": "spdkcsi", + }, + }, + "spec": map[string]any{ + "pvName": pvName, + "targetNodeUUID": targetNodeUUID, + }, + }} + created, cerr := res.Create(ctx, obj, metav1.CreateOptions{}) + if cerr != nil && !apierrors.IsAlreadyExists(cerr) { + return "", fmt.Errorf("create VolumeMigration %s/%s: %w", m.namespace, name, cerr) + } + if cerr == nil { + obj = created + } else if obj, err = res.Get(ctx, name, metav1.GetOptions{}); err != nil { + return "", fmt.Errorf("read VolumeMigration %s/%s: %w", m.namespace, name, err) + } + } else if err != nil { + return "", fmt.Errorf("read VolumeMigration %s/%s: %w", m.namespace, name, err) + } + phase, _, _ := unstructured.NestedString(obj.Object, "status", "phase") + return phase, nil +} + +// watcherOptions reads the co-location switches from the environment: +// SPDKCSI_CG_PREJOIN_MIGRATION (default true), SPDKCSI_CG_COLOCATE (default +// false) and SPDKCSI_CG_CLIENT_SWAP_READY (default false). +func watcherOptions() (preJoin, colocate, swapReady bool) { + flag := func(name string, def bool) bool { + switch strings.ToLower(strings.TrimSpace(os.Getenv(name))) { + case "1", "true", "yes", "on": + return true + case "0", "false", "no", "off": + return false + default: + return def + } + } + return flag("SPDKCSI_CG_PREJOIN_MIGRATION", true), + flag("SPDKCSI_CG_COLOCATE", false), + flag("SPDKCSI_CG_CLIENT_SWAP_READY", false) +} diff --git a/csi-driver/internal/csi/controller/cg_prejoin_migration_test.go b/csi-driver/internal/csi/controller/cg_prejoin_migration_test.go new file mode 100644 index 000000000..ce1a98e61 --- /dev/null +++ b/csi-driver/internal/csi/controller/cg_prejoin_migration_test.go @@ -0,0 +1,72 @@ +package controller + +import ( + "context" + "testing" + + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" + "k8s.io/apimachinery/pkg/runtime" + "k8s.io/apimachinery/pkg/runtime/schema" + dynamicfake "k8s.io/client-go/dynamic/fake" +) + +func newFakeVolumeMigrations(objects ...runtime.Object) (*volumeMigrations, *dynamicfake.FakeDynamicClient) { + client := dynamicfake.NewSimpleDynamicClientWithCustomListKinds(runtime.NewScheme(), + map[schema.GroupVersionResource]string{volumeMigrationGVR: "VolumeMigrationList"}, objects...) + return &volumeMigrations{client: client, namespace: "simplyblock"}, client +} + +func TestEnsureCreatesTheRequestOnceAndReportsItsPhase(t *testing.T) { + m, client := newFakeVolumeMigrations() + phase, err := m.Ensure(context.Background(), "cg-join-v1", "pv-1", "node-pin") + if err != nil || phase != "" { + t.Fatalf("first Ensure: phase %q err %v", phase, err) + } + got, err := client.Resource(volumeMigrationGVR).Namespace("simplyblock"). + Get(context.Background(), "cg-join-v1", metav1.GetOptions{}) + if err != nil { + t.Fatalf("the VolumeMigration was not created: %v", err) + } + pv, _, _ := unstructured.NestedString(got.Object, "spec", "pvName") + target, _, _ := unstructured.NestedString(got.Object, "spec", "targetNodeUUID") + if pv != "pv-1" || target != "node-pin" { + t.Fatalf("spec = %s -> %s, want pv-1 -> node-pin", pv, target) + } + if got.GetLabels()[preJoinPurposeLabel] != "consistency-group-join" { + t.Fatalf("missing the purpose label: %v", got.GetLabels()) + } + + if err := unstructured.SetNestedField(got.Object, "Completed", "status", "phase"); err != nil { + t.Fatal(err) + } + if _, err := client.Resource(volumeMigrationGVR).Namespace("simplyblock"). + Update(context.Background(), got, metav1.UpdateOptions{}); err != nil { + t.Fatal(err) + } + phase, err = m.Ensure(context.Background(), "cg-join-v1", "pv-1", "other-node") + if err != nil || phase != "Completed" { + t.Fatalf("second Ensure: phase %q err %v", phase, err) + } + again, _ := client.Resource(volumeMigrationGVR).Namespace("simplyblock"). + Get(context.Background(), "cg-join-v1", metav1.GetOptions{}) + if target, _, _ := unstructured.NestedString(again.Object, "spec", "targetNodeUUID"); target != "node-pin" { + t.Fatalf("an existing request must never be rewritten, target is now %s", target) + } +} + +func TestWatcherOptionsDefaults(t *testing.T) { + t.Setenv("SPDKCSI_CG_PREJOIN_MIGRATION", "") + t.Setenv("SPDKCSI_CG_COLOCATE", "") + t.Setenv("SPDKCSI_CG_CLIENT_SWAP_READY", "") + preJoin, colocate, swap := watcherOptions() + if !preJoin || colocate || swap { + t.Fatalf("defaults = %v %v %v, want true false false", preJoin, colocate, swap) + } + t.Setenv("SPDKCSI_CG_PREJOIN_MIGRATION", "off") + t.Setenv("SPDKCSI_CG_COLOCATE", "true") + preJoin, colocate, _ = watcherOptions() + if preJoin || !colocate { + t.Fatalf("overrides = %v %v, want false true", preJoin, colocate) + } +} diff --git a/csi-driver/internal/driver/driver.go b/csi-driver/internal/driver/driver.go index fa0ab78d4..1589f62e5 100644 --- a/csi-driver/internal/driver/driver.go +++ b/csi-driver/internal/driver/driver.go @@ -32,6 +32,7 @@ import ( csiaddonsreplication "github.com/csi-addons/spec/lib/go/replication" csiaddonsvolumegroup "github.com/csi-addons/spec/lib/go/volumegroup" "google.golang.org/grpc" + "k8s.io/client-go/dynamic" "k8s.io/client-go/kubernetes" "k8s.io/client-go/rest" "k8s.io/klog" @@ -94,12 +95,20 @@ func Run(conf *config.Config) { // its own in-cluster config + clientset. A missing in-cluster config is // non-fatal, and the features that need it degrade to no-ops. var kubeClient kubernetes.Interface + var dynClient dynamic.Interface if k8sConfig, err := rest.InClusterConfig(); err != nil { klog.Warningf("no in-cluster config; Kubernetes API features disabled: %v", err) } else if clientset, err := kubernetes.NewForConfig(k8sConfig); err != nil { klog.Warningf("failed to create kubernetes client; Kubernetes API features disabled: %v", err) } else { kubeClient = clientset + // The consistency-group watcher requests VolumeMigrations (the + // pre-join live migration) without importing the operator's types. + if d, err := dynamic.NewForConfig(k8sConfig); err != nil { + klog.Warningf("failed to create dynamic client; pre-join migrations disabled: %v", err) + } else { + dynClient = d + } } if conf.IsNodeServer { @@ -122,7 +131,7 @@ func Run(conf *config.Config) { // without a Kubernetes client, like the other kube-backed features. watcherCtx, watcherCancel := context.WithCancel(context.Background()) defer watcherCancel() - controller.StartConsistencyGroupLabelWatcher(watcherCtx, kubeClient, conf.DriverName) + controller.StartConsistencyGroupLabelWatcher(watcherCtx, kubeClient, dynClient, conf.DriverName) } // The link to the operator, when enabled. It is independent of the CSI From 161dad789a2fd413dc9fb6e86983c46f0222c660 Mon Sep 17 00:00:00 2001 From: michael Date: Sat, 3 Oct 2026 23:18:39 +0300 Subject: [PATCH 186/206] operator: the rebalancer skips consistency-group members; the CSI controller may request VolumeMigrations A group's members live on one node/LVS; the rebalancer moving one alone splits the group (the control plane now refuses it). BuildPinnedSet treats a PVC with the consistency-group label like a pinned one. The VolumeMigration webhook's refusal points at the control plane's group migration. The CSI provisioner role gains get/create on volumemigrations for the pre-join migration. Co-Authored-By: Claude Opus 5.5 --- .../autoplacement/logical_volume_selector.go | 11 +++- .../internal/autoplacement/pinned_set_test.go | 55 +++++++++++++++++++ operator/internal/controllers/driver/rbac.go | 7 +++ .../internal/controllers/driver/rbac_test.go | 1 + .../webhook/volumemigration_validator.go | 9 ++- 5 files changed, 79 insertions(+), 4 deletions(-) create mode 100644 operator/internal/autoplacement/pinned_set_test.go diff --git a/operator/internal/autoplacement/logical_volume_selector.go b/operator/internal/autoplacement/logical_volume_selector.go index 2a07a174e..2ab3804f8 100644 --- a/operator/internal/autoplacement/logical_volume_selector.go +++ b/operator/internal/autoplacement/logical_volume_selector.go @@ -235,8 +235,15 @@ func (lvs *LogicalVolumeSelector) selectMigrationSet(ranked []RankedCandidate, m return out } +// consistencyGroupLabel marks a PVC whose volume is a consistency-group member. +const consistencyGroupLabel = "storage.simplyblock.io/consistency-group" + // BuildPinnedSet returns the set of volume UUIDs whose bound PVC carries a pin -// annotation (see kube.IsPinnedVolume). It scans all PersistentVolumes and +// annotation (see kube.IsPinnedVolume) or the consistency-group label. A +// group's members live on one node/LVS, so the rebalancer must not move one +// of them alone; it skips them (the control plane refuses such a migration +// too), and a group moves only as a whole through the group migration +// (co-location design §3). It scans all PersistentVolumes and // resolves the volume UUID from the CSI volume handle // ("::"). Pass an empty clusterUUID to include // volumes from all clusters. @@ -270,7 +277,7 @@ func (lvs *LogicalVolumeSelector) BuildPinnedSet(ctx context.Context, clusterUUI }, pvc); err != nil { continue } - if atlaskube.IsPinnedVolume(pvc.Annotations) { + if atlaskube.IsPinnedVolume(pvc.Annotations) || pvc.Labels[consistencyGroupLabel] != "" { pinned[lvolID] = true } } diff --git a/operator/internal/autoplacement/pinned_set_test.go b/operator/internal/autoplacement/pinned_set_test.go new file mode 100644 index 000000000..46be610ff --- /dev/null +++ b/operator/internal/autoplacement/pinned_set_test.go @@ -0,0 +1,55 @@ +package autoplacement + +import ( + "context" + "testing" + + corev1 "k8s.io/api/core/v1" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "sigs.k8s.io/controller-runtime/pkg/client" + "sigs.k8s.io/controller-runtime/pkg/client/fake" + + atlaskube "github.com/simplyblock/atlas/kube" +) + +func boundPV(name, handle, claim string) *corev1.PersistentVolume { + pv := csiPV(name, handle, "sc") + pv.Spec.ClaimRef = &corev1.ObjectReference{Name: claim, Namespace: "app"} + return pv +} + +func claim(name string, labels, annotations map[string]string) *corev1.PersistentVolumeClaim { + return &corev1.PersistentVolumeClaim{ObjectMeta: metav1.ObjectMeta{ + Name: name, Namespace: "app", Labels: labels, Annotations: annotations, + }} +} + +// A consistency group's members live on one node/LVS: the rebalancer moving +// one of them alone would split the group, so a member is treated like a +// pinned volume and skipped. +func TestBuildPinnedSetSkipsConsistencyGroupMembers(t *testing.T) { + objects := []client.Object{ + boundPV("pv-member", "cluster-a:pool:vol-member", "member"), + boundPV("pv-plain", "cluster-a:pool:vol-plain", "plain"), + boundPV("pv-pinned", "cluster-a:pool:vol-pinned", "pinned"), + claim("member", map[string]string{consistencyGroupLabel: "db"}, nil), + claim("plain", nil, nil), + claim("pinned", nil, map[string]string{atlaskube.AnnoSelectedStorageNode: "node-1"}), + } + cl := fake.NewClientBuilder().WithScheme(namespacedTestScheme(t)).WithObjects(objects...).Build() + lvs := NewLogicalVolumeSelector(nil, cl, nil) + + got, err := lvs.BuildPinnedSet(context.Background(), "cluster-a") + if err != nil { + t.Fatalf("BuildPinnedSet: %v", err) + } + if !got["vol-member"] { + t.Error("a consistency-group member must be skipped by the rebalancer") + } + if !got["vol-pinned"] { + t.Error("a pinned volume must stay skipped") + } + if got["vol-plain"] { + t.Error("an ordinary volume must stay a rebalancing candidate") + } +} diff --git a/operator/internal/controllers/driver/rbac.go b/operator/internal/controllers/driver/rbac.go index ecc283c83..a4a940f2f 100644 --- a/operator/internal/controllers/driver/rbac.go +++ b/operator/internal/controllers/driver/rbac.go @@ -33,6 +33,7 @@ var ( snapshot = []string{"snapshot.storage.k8s.io"} groupsnapshot = []string{"groupsnapshot.storage.k8s.io"} csiaddons = []string{"csiaddons.openshift.io"} + simplyblock = []string{"storage.simplyblock.io"} coordination = []string{"coordination.k8s.io"} ) @@ -72,6 +73,12 @@ var clusterRoleRules = map[string][]rbacv1.PolicyRule{ rule(storage, []string{"csinodes"}, "get", "list", "watch"), rule(core, []string{"nodes"}, "get", "list", "watch"), rule(storage, []string{"volumeattachments"}, "get", "list", "watch"), + // The consistency-group watcher's pre-join live migration: a volume + // labeled into a group while it lives off the group's pinned node is + // moved there first through a VolumeMigration, which the operator + // runs (co-location design §5). It creates and reads its own + // requests; it never updates or deletes one. + rule(simplyblock, []string{"volumemigrations"}, "get", "create"), }, "attacher": { rule(core, []string{"persistentvolumes"}, "get", "list", "watch", "update", "patch"), diff --git a/operator/internal/controllers/driver/rbac_test.go b/operator/internal/controllers/driver/rbac_test.go index d03f6a870..413a3220e 100644 --- a/operator/internal/controllers/driver/rbac_test.go +++ b/operator/internal/controllers/driver/rbac_test.go @@ -105,6 +105,7 @@ func TestRulesMatchTheChart(t *testing.T) { {nodeComponent, "", "events", []string{"create", "patch"}}, {"provisioner", "snapshot.storage.k8s.io", "volumesnapshotcontents/status", []string{"get", "update", "patch"}}, {"provisioner", "", "persistentvolumes", []string{"get", "list", "watch", "create", "delete", "patch"}}, + {"provisioner", "storage.simplyblock.io", "volumemigrations", []string{"get", "create"}}, {"attacher", "storage.k8s.io", "volumeattachments/status", []string{"patch"}}, {"resizer", "", "persistentvolumeclaims/status", []string{"patch"}}, {"health-monitor", "", "events", []string{"get", "list", "watch", "create", "patch"}}, diff --git a/operator/internal/webhook/volumemigration_validator.go b/operator/internal/webhook/volumemigration_validator.go index e34f9d7b0..681ddb1b6 100644 --- a/operator/internal/webhook/volumemigration_validator.go +++ b/operator/internal/webhook/volumemigration_validator.go @@ -61,9 +61,14 @@ func (v *VolumeMigrationValidator) Handle(ctx context.Context, req admission.Req return admission.Allowed("consistency-group membership undeterminable; deferring to the backend") } if member { + // A group moves only as a whole: every member, and every subsystem + // holding one, in one group migration (sbcli co-location design §3). + // A VolumeMigration moves one subsystem, so it would split the group. return admission.Denied(fmt.Sprintf( - "volume %s is a member of consistency group %s and cannot be migrated; "+ - "a group's members are pinned to one logical volume store (§8.4)", volumeUUID, groupID)) + "volume %s is a member of consistency group %s and cannot be migrated alone: "+ + "a group's members live on one logical volume store, so the group moves as a whole "+ + "through the control plane's group migration "+ + "(POST /api/v2/clusters//consistency-groups//migration)", volumeUUID, groupID)) } return admission.Allowed("target volume is not a consistency-group member") } From bed4b73da2385313ed0ca0af0de0f5ff5f328e82 Mon Sep 17 00:00:00 2001 From: michael Date: Sat, 3 Oct 2026 23:18:39 +0300 Subject: [PATCH 187/206] volstack: a dm-linear indirection under raw block and plain volumes (flag-gated) Moving a volume's namespace to another NVMe subsystem (consistency-group co-location) changes the block device the namespace appears as. A new dmLinear layer between the fabric and the consumer exposes /dev/mapper/sb- over the namespace device; when a heal brings the fabric up on the new namespace, the layer's Heal re-points the mapping (dmsetup suspend / reload / resume, resume on the old table if the reload fails) and the pod's device or mounted filesystem stays. - atlas-lib/devmapper: one-segment linear mappings (create, table, swap, remove, size). - plans: IndirectRawBlock (fabric -> dmLinear), IndirectPlain (fabric -> dmLinear -> filesystem). - node plugin: SPDKCSI_DM_INDIRECTION (default off) picks the indirect rows for a fresh stage only; a staged volume follows its stack record, so the layer is never inserted under or pulled from under a consumer. Teardown recognises the new records. What does not exist yet: a node-side trigger that heals a staged volume when its namespace moved (today a heal runs at publish), and the release of the old namespace's path. Until then the control plane keeps namespace moves of attached volumes off. Co-Authored-By: Claude Opus 5.5 --- atlas-lib/devmapper/devmapper.go | 150 ++++++++++++++ atlas-lib/devmapper/devmapper_test.go | 106 ++++++++++ atlas-lib/volstack/layers/contract_test.go | 7 + atlas-lib/volstack/layers/dmlinear.go | 184 ++++++++++++++++++ atlas-lib/volstack/layers/dmlinear_test.go | 171 ++++++++++++++++ atlas-lib/volstack/plans/node.go | 20 ++ atlas-lib/volstack/plans/plans.go | 14 ++ .../internal/csi/node/indirection_test.go | 77 ++++++++ csi-driver/internal/csi/node/plan.go | 41 ++++ csi-driver/internal/csi/node/stage.go | 28 ++- 10 files changed, 797 insertions(+), 1 deletion(-) create mode 100644 atlas-lib/devmapper/devmapper.go create mode 100644 atlas-lib/devmapper/devmapper_test.go create mode 100644 atlas-lib/volstack/layers/dmlinear.go create mode 100644 atlas-lib/volstack/layers/dmlinear_test.go create mode 100644 csi-driver/internal/csi/node/indirection_test.go diff --git a/atlas-lib/devmapper/devmapper.go b/atlas-lib/devmapper/devmapper.go new file mode 100644 index 000000000..0baab0e25 --- /dev/null +++ b/atlas-lib/devmapper/devmapper.go @@ -0,0 +1,150 @@ +// Package devmapper manages single-segment dm-linear devices: the indirection +// between an NVMe-oF namespace and what a pod or a filesystem uses. +// +// A dm-linear device maps its whole range onto one underlying block device and +// writes nothing to it, so it can be put in front of a volume that already +// carries data. What it buys is a device whose identity does not change when +// the namespace behind it does: when a volume's namespace moves to another NVMe +// subsystem (consistency-group co-location), the new namespace shows up as a +// different block device, and Swap re-points the mapping at it -- suspend, load +// the new table, resume -- while the consumer keeps the device it opened. I/O +// issued during the suspend is queued by device-mapper, not failed. +package devmapper + +import ( + "context" + "errors" + "fmt" + "os" + "os/exec" + "strconv" + "strings" +) + +// Runner execs a command and returns its combined output. Tests substitute a +// recorder; production execs the host's dmsetup and blockdev. +type Runner func(ctx context.Context, args ...string) (string, error) + +// ErrNotFound reports that the named mapping does not exist. +var ErrNotFound = errors.New("device-mapper mapping not found") + +// Mapper creates, reads, re-points and removes dm-linear mappings. +type Mapper struct { + run Runner +} + +// New returns a Mapper over run; nil runs the host's commands. +func New(run Runner) *Mapper { + if run == nil { + run = runCommand + } + return &Mapper{run: run} +} + +func runCommand(ctx context.Context, args ...string) (string, error) { + //nolint:gosec // fixed set of binaries (dmsetup, blockdev), structured args + cmd := exec.CommandContext(ctx, args[0], args[1:]...) + // No udev in the node plugin's container to complete the handshake. + cmd.Env = append(os.Environ(), "DM_DISABLE_UDEV=1") + out, err := cmd.CombinedOutput() + if err != nil { + return string(out), fmt.Errorf("%v: %w: %s", args, err, strings.TrimSpace(string(out))) + } + return string(out), nil +} + +// Path is the device node of mapping name. +func Path(name string) string { return "/dev/mapper/" + name } + +// Target is what a mapping currently points at. +type Target struct { + // Sectors is the mapped length in 512-byte sectors. + Sectors uint64 + // Device is the underlying device as "major:minor". + Device string +} + +// table renders the one-segment linear table for device. +func table(sectors uint64, device string) string { + return fmt.Sprintf("0 %d linear %s 0", sectors, device) +} + +// Sectors is device's size in 512-byte sectors. +func (m *Mapper) Sectors(ctx context.Context, device string) (uint64, error) { + out, err := m.run(ctx, "blockdev", "--getsz", device) + if err != nil { + return 0, fmt.Errorf("size of %s: %w", device, err) + } + n, err := strconv.ParseUint(strings.TrimSpace(out), 10, 64) + if err != nil { + return 0, fmt.Errorf("size of %s: unexpected %q", device, strings.TrimSpace(out)) + } + return n, nil +} + +// Table reads mapping name's target, or ErrNotFound. +func (m *Mapper) Table(ctx context.Context, name string) (Target, error) { + out, err := m.run(ctx, "dmsetup", "table", name) + if err != nil { + if strings.Contains(out, "No such device") || strings.Contains(err.Error(), "No such device") { + return Target{}, ErrNotFound + } + return Target{}, fmt.Errorf("dmsetup table %s: %w", name, err) + } + return parseTable(name, out) +} + +func parseTable(name, out string) (Target, error) { + lines := strings.Split(strings.TrimSpace(out), "\n") + if len(lines) != 1 { + return Target{}, fmt.Errorf("mapping %s has %d segments; a linear indirection has one", name, len(lines)) + } + f := strings.Fields(lines[0]) + if len(f) != 5 || f[2] != "linear" || f[0] != "0" || f[4] != "0" { + return Target{}, fmt.Errorf("mapping %s is not a whole-device linear mapping: %q", name, lines[0]) + } + sectors, err := strconv.ParseUint(f[1], 10, 64) + if err != nil { + return Target{}, fmt.Errorf("mapping %s: length %q: %w", name, f[1], err) + } + return Target{Sectors: sectors, Device: f[3]}, nil +} + +// Create maps name onto device over sectors. +func (m *Mapper) Create(ctx context.Context, name, device string, sectors uint64) error { + if _, err := m.run(ctx, "dmsetup", "create", name, "--table", table(sectors, device)); err != nil { + return fmt.Errorf("dmsetup create %s: %w", name, err) + } + return nil +} + +// Swap re-points name at device: suspend (in-flight I/O drains, new I/O is +// queued), load the new table, resume. A failed load resumes on the old table, +// so a swap never leaves the mapping suspended. +func (m *Mapper) Swap(ctx context.Context, name, device string, sectors uint64) error { + if _, err := m.run(ctx, "dmsetup", "suspend", name); err != nil { + return fmt.Errorf("dmsetup suspend %s: %w", name, err) + } + if _, err := m.run(ctx, "dmsetup", "reload", name, "--table", table(sectors, device)); err != nil { + if _, rerr := m.run(ctx, "dmsetup", "resume", name); rerr != nil { + return fmt.Errorf("dmsetup reload %s: %w; resume on the old table also failed: %v", name, err, rerr) + } + return fmt.Errorf("dmsetup reload %s: %w (resumed on the old table)", name, err) + } + if _, err := m.run(ctx, "dmsetup", "resume", name); err != nil { + return fmt.Errorf("dmsetup resume %s: %w", name, err) + } + return nil +} + +// Remove deletes mapping name; a missing mapping is success. +func (m *Mapper) Remove(ctx context.Context, name string) error { + out, err := m.run(ctx, "dmsetup", "remove", "--retry", name) + if err != nil { + if strings.Contains(out, "No such device") || strings.Contains(err.Error(), "No such device") { + return nil + } + return fmt.Errorf("dmsetup remove %s: %w", name, err) + } + return nil +} diff --git a/atlas-lib/devmapper/devmapper_test.go b/atlas-lib/devmapper/devmapper_test.go new file mode 100644 index 000000000..2bf468a7a --- /dev/null +++ b/atlas-lib/devmapper/devmapper_test.go @@ -0,0 +1,106 @@ +package devmapper + +import ( + "context" + "errors" + "strings" + "testing" +) + +type recorder struct { + calls []string + out map[string]string + err map[string]error +} + +func (r *recorder) run(_ context.Context, args ...string) (string, error) { + key := strings.Join(args, " ") + r.calls = append(r.calls, key) + for prefix, err := range r.err { + if strings.HasPrefix(key, prefix) { + return r.out[prefix], err + } + } + for prefix, out := range r.out { + if strings.HasPrefix(key, prefix) { + return out, nil + } + } + return "", nil +} + +func TestTableParsesAWholeDeviceLinearMapping(t *testing.T) { + r := &recorder{out: map[string]string{"dmsetup table sb-v": "0 2097152 linear 259:3 0\n"}} + got, err := New(r.run).Table(context.Background(), "sb-v") + if err != nil || got != (Target{Sectors: 2097152, Device: "259:3"}) { + t.Fatalf("Table = %+v, %v", got, err) + } +} + +func TestTableOfAMissingMappingIsErrNotFound(t *testing.T) { + r := &recorder{ + out: map[string]string{"dmsetup table": "device-mapper: table ioctl on sb-v failed: No such device or address"}, + err: map[string]error{"dmsetup table": errors.New("exit status 1")}, + } + if _, err := New(r.run).Table(context.Background(), "sb-v"); !errors.Is(err, ErrNotFound) { + t.Fatalf("err = %v, want ErrNotFound", err) + } +} + +func TestTableRefusesAMappingThatIsNotOneLinearSegment(t *testing.T) { + for _, out := range []string{ + "0 100 striped 2 128 259:3 0 259:4 0", + "0 100 linear 259:3 0\n100 100 linear 259:4 0", + "0 100 linear 259:3 8", + } { + if _, err := parseTable("sb-v", out); err == nil { + t.Errorf("accepted %q", out) + } + } +} + +func TestSwapSuspendsReloadsAndResumes(t *testing.T) { + r := &recorder{} + if err := New(r.run).Swap(context.Background(), "sb-v", "259:11", 2048); err != nil { + t.Fatal(err) + } + want := []string{ + "dmsetup suspend sb-v", + "dmsetup reload sb-v --table 0 2048 linear 259:11 0", + "dmsetup resume sb-v", + } + if strings.Join(r.calls, "|") != strings.Join(want, "|") { + t.Fatalf("calls = %v, want %v", r.calls, want) + } +} + +// A refused reload must not leave the mapping suspended: I/O queued during the +// suspend would hang until someone resumed it by hand. +func TestAFailedReloadResumesOnTheOldTable(t *testing.T) { + r := &recorder{err: map[string]error{"dmsetup reload": errors.New("invalid table")}} + err := New(r.run).Swap(context.Background(), "sb-v", "259:11", 2048) + if err == nil || !strings.Contains(err.Error(), "resumed on the old table") { + t.Fatalf("err = %v", err) + } + if r.calls[len(r.calls)-1] != "dmsetup resume sb-v" { + t.Fatalf("last call %q, want a resume", r.calls[len(r.calls)-1]) + } +} + +func TestRemovingAMissingMappingSucceeds(t *testing.T) { + r := &recorder{ + out: map[string]string{"dmsetup remove": "No such device or address"}, + err: map[string]error{"dmsetup remove": errors.New("exit status 1")}, + } + if err := New(r.run).Remove(context.Background(), "sb-v"); err != nil { + t.Fatal(err) + } +} + +func TestSectorsReadsBlockdev(t *testing.T) { + r := &recorder{out: map[string]string{"blockdev --getsz /dev/nvme1n1": "2097152\n"}} + n, err := New(r.run).Sectors(context.Background(), "/dev/nvme1n1") + if err != nil || n != 2097152 { + t.Fatalf("Sectors = %d, %v", n, err) + } +} diff --git a/atlas-lib/volstack/layers/contract_test.go b/atlas-lib/volstack/layers/contract_test.go index 5d23b055f..6cdd86853 100644 --- a/atlas-lib/volstack/layers/contract_test.go +++ b/atlas-lib/volstack/layers/contract_test.go @@ -68,6 +68,13 @@ func shippedLayers() map[string]volstack.Layer { Ops: newFakeFS(), Content: fakeReader{reading: blockdev.Reading{Content: blockdev.ContentBlank}}, }), + "dmLinear": NewDMLinear(DMLinearConfig{ + Name: DMLinearName("vol-x"), + Mapper: newFakeMapper(), + Resolve: func(path string) (blockdev.Device, error) { + return blockdev.Device{Path: path, Major: 253, Minor: 1}, nil + }, + }), } } diff --git a/atlas-lib/volstack/layers/dmlinear.go b/atlas-lib/volstack/layers/dmlinear.go new file mode 100644 index 000000000..6e3c40b77 --- /dev/null +++ b/atlas-lib/volstack/layers/dmlinear.go @@ -0,0 +1,184 @@ +// The dmLinear layer: a dm-linear device between the fabric and whatever uses +// the volume, so the device the consumer holds survives a change of the +// namespace behind it. +// +// A volume's namespace can move to another NVMe subsystem (consistency-group +// co-location, sbcli docs/consistency-group-colocation.md §6). The new +// namespace is a different block device: without an indirection the pod's +// raw block device or the mounted filesystem would have to be torn down and +// set up again. With it, the fabric layer below brings up the new namespace, +// and this layer's Heal re-points its mapping there (suspend, load, resume) +// while the device above stays the same node with the same data. +package layers + +import ( + "context" + "errors" + "fmt" + + "github.com/simplyblock/atlas/blockdev" + "github.com/simplyblock/atlas/devmapper" + "github.com/simplyblock/atlas/volstack" +) + +// DMMapper is the device-mapper surface the layer uses; *devmapper.Mapper +// satisfies it, and tests substitute a recorder. +type DMMapper interface { + Table(ctx context.Context, name string) (devmapper.Target, error) + Sectors(ctx context.Context, device string) (uint64, error) + Create(ctx context.Context, name, device string, sectors uint64) error + Swap(ctx context.Context, name, device string, sectors uint64) error + Remove(ctx context.Context, name string) error +} + +// DMLinearConfig is what a dmLinear layer is built with. +type DMLinearConfig struct { + // Name is the mapping's name, derived from the volume's identity so a + // restage finds the mapping it made (sb-). + Name string + // Mapper runs the device-mapper commands. + Mapper DMMapper + // Resolve describes the mapping's device node upward; defaults to + // blockdev.ResolveDevice. + Resolve DeviceResolver +} + +// DMLinear is the indirection layer. +type DMLinear struct { + cfg DMLinearConfig + // stale is what the last Observe found: the mapping points somewhere else + // than the device below now is. The heal loop asks Healthy without the + // layer below, so the comparison is made where both are known. + stale bool +} + +// NewDMLinear returns the layer. +func NewDMLinear(cfg DMLinearConfig) *DMLinear { + if cfg.Resolve == nil { + cfg.Resolve = blockdev.ResolveDevice + } + return &DMLinear{cfg: cfg} +} + +// Name is what the record calls this layer. +func (d *DMLinear) Name() string { return "dmLinear" } + +// DMLinearName is the mapping name of the volume with UUID uuid. +func DMLinearName(uuid string) string { return "sb-" + uuid } + +func deviceNumber(dev blockdev.Device) string { return fmt.Sprintf("%d:%d", dev.Major, dev.Minor) } + +func belowDevice(below volstack.Artifact) (blockdev.Device, bool) { + if len(below.Devices) != 1 { + return blockdev.Device{}, false + } + return below.Devices[0], true +} + +// Observe reports the mapping: absent, ready when it points at the device +// below, and partial when it points elsewhere (the namespace moved and the +// fabric below brought up the new one) or there is nothing below to point at. +func (d *DMLinear) Observe(ctx context.Context, below volstack.Artifact) (volstack.State, volstack.Artifact, error) { + target, err := d.cfg.Mapper.Table(ctx, d.cfg.Name) + if errors.Is(err, devmapper.ErrNotFound) { + d.stale = false + return volstack.StateAbsent, volstack.Artifact{}, nil + } + if err != nil { + return volstack.StateAbsent, volstack.Artifact{}, fmt.Errorf("dmLinear: read %s: %w", d.cfg.Name, err) + } + own, err := d.own() + if err != nil { + return volstack.StateAbsent, volstack.Artifact{}, err + } + dev, ok := belowDevice(below) + d.stale = !ok || target.Device != deviceNumber(dev) + if d.stale { + return volstack.StatePartial, own, nil + } + return volstack.StateReady, own, nil +} + +func (d *DMLinear) own() (volstack.Artifact, error) { + dev, err := d.cfg.Resolve(devmapper.Path(d.cfg.Name)) + if err != nil { + return volstack.Artifact{}, fmt.Errorf("dmLinear: resolve %s: %w", devmapper.Path(d.cfg.Name), err) + } + return volstack.Artifact{Devices: []blockdev.Device{dev}}, nil +} + +// Ensure creates the mapping over the device below, or re-points it there. +func (d *DMLinear) Ensure(ctx context.Context, below volstack.Artifact) (volstack.Artifact, error) { + state, own, err := d.Observe(ctx, below) + if err != nil { + return volstack.Artifact{}, err + } + switch state { + case volstack.StateReady: + return own, nil + case volstack.StatePartial: + if err := d.swap(ctx, below); err != nil { + return volstack.Artifact{}, err + } + return d.own() + } + dev, ok := belowDevice(below) + if !ok { + return volstack.Artifact{}, fmt.Errorf("dmLinear: %s needs exactly one device below, got %d", + d.cfg.Name, len(below.Devices)) + } + sectors, err := d.cfg.Mapper.Sectors(ctx, dev.Path) + if err != nil { + return volstack.Artifact{}, fmt.Errorf("dmLinear: %w", err) + } + if err := d.cfg.Mapper.Create(ctx, d.cfg.Name, deviceNumber(dev), sectors); err != nil { + return volstack.Artifact{}, fmt.Errorf("dmLinear: %w", err) + } + return d.own() +} + +func (d *DMLinear) swap(ctx context.Context, below volstack.Artifact) error { + dev, ok := belowDevice(below) + if !ok { + return fmt.Errorf("dmLinear: cannot re-point %s: no device below", d.cfg.Name) + } + sectors, err := d.cfg.Mapper.Sectors(ctx, dev.Path) + if err != nil { + return fmt.Errorf("dmLinear: %w", err) + } + if err := d.cfg.Mapper.Swap(ctx, d.cfg.Name, deviceNumber(dev), sectors); err != nil { + return fmt.Errorf("dmLinear: %w", err) + } + d.stale = false + return nil +} + +// Release removes the mapping and keeps the data, which lives below it. +func (d *DMLinear) Release(ctx context.Context, _ volstack.Artifact) error { + return d.cfg.Mapper.Remove(ctx, d.cfg.Name) +} + +// Destroy has nothing durable to remove: the mapping writes no metadata. +func (d *DMLinear) Destroy(context.Context, volstack.Artifact) error { return nil } + +// Healthy reports whether the mapping points at the device below, as the last +// Observe found it. +func (d *DMLinear) Healthy(context.Context, volstack.Artifact) (bool, error) { return !d.stale, nil } + +// Heal re-points the mapping at the device below, which the fabric layer may +// just have brought up on another subsystem. +func (d *DMLinear) Heal(ctx context.Context, below, _ volstack.Artifact) error { + if _, err := d.cfg.Mapper.Table(ctx, d.cfg.Name); errors.Is(err, devmapper.ErrNotFound) { + _, err := d.Ensure(ctx, below) + return err + } + return d.swap(ctx, below) +} + +// DMLinearParams is what the record keeps of this layer. +type DMLinearParams struct { + Name string `json:"name"` +} + +// Params is the record's view of the layer. +func (d *DMLinear) Params() any { return DMLinearParams{Name: d.cfg.Name} } diff --git a/atlas-lib/volstack/layers/dmlinear_test.go b/atlas-lib/volstack/layers/dmlinear_test.go new file mode 100644 index 000000000..452a17706 --- /dev/null +++ b/atlas-lib/volstack/layers/dmlinear_test.go @@ -0,0 +1,171 @@ +package layers + +import ( + "context" + "errors" + "testing" + + "github.com/simplyblock/atlas/blockdev" + "github.com/simplyblock/atlas/devmapper" + "github.com/simplyblock/atlas/volstack" +) + +// fakeMapper keeps mappings in memory and records the verbs run against them. +type fakeMapper struct { + tables map[string]devmapper.Target + sectors map[string]uint64 + calls []string + swapErr error +} + +func newFakeMapper() *fakeMapper { + return &fakeMapper{tables: map[string]devmapper.Target{}, sectors: map[string]uint64{}} +} + +func (f *fakeMapper) Table(_ context.Context, name string) (devmapper.Target, error) { + t, ok := f.tables[name] + if !ok { + return devmapper.Target{}, devmapper.ErrNotFound + } + return t, nil +} + +func (f *fakeMapper) Sectors(_ context.Context, device string) (uint64, error) { + if n, ok := f.sectors[device]; ok { + return n, nil + } + return 2048, nil +} + +func (f *fakeMapper) Create(_ context.Context, name, device string, sectors uint64) error { + f.calls = append(f.calls, "create "+name+" "+device) + f.tables[name] = devmapper.Target{Sectors: sectors, Device: device} + return nil +} + +func (f *fakeMapper) Swap(_ context.Context, name, device string, sectors uint64) error { + f.calls = append(f.calls, "swap "+name+" "+device) + if f.swapErr != nil { + return f.swapErr + } + f.tables[name] = devmapper.Target{Sectors: sectors, Device: device} + return nil +} + +func (f *fakeMapper) Remove(_ context.Context, name string) error { + f.calls = append(f.calls, "remove "+name) + delete(f.tables, name) + return nil +} + +func nvmeBelow(path string, major, minor uint32) volstack.Artifact { + return volstack.Artifact{Devices: []blockdev.Device{{Path: path, Major: major, Minor: minor}}} +} + +func dmLayer(m *fakeMapper) *DMLinear { + return NewDMLinear(DMLinearConfig{ + Name: DMLinearName("vol-1"), + Mapper: m, + Resolve: func(path string) (blockdev.Device, error) { + return blockdev.Device{Path: path, Name: "dm-7", Major: 253, Minor: 7}, nil + }, + }) +} + +func TestDMLinearEnsureMapsTheDeviceBelowAndExposesTheMapping(t *testing.T) { + m := newFakeMapper() + layer := dmLayer(m) + own, err := layer.Ensure(context.Background(), nvmeBelow("/dev/nvme1n1", 259, 3)) + if err != nil { + t.Fatalf("ensure: %v", err) + } + if len(m.calls) != 1 || m.calls[0] != "create sb-vol-1 259:3" { + t.Fatalf("calls = %v, want one create over 259:3", m.calls) + } + if own.Devices[0].Path != "/dev/mapper/sb-vol-1" { + t.Fatalf("exposes %s, want the mapping", own.Devices[0].Path) + } + state, _, _ := layer.Observe(context.Background(), nvmeBelow("/dev/nvme1n1", 259, 3)) + if state != volstack.StateReady { + t.Fatalf("state after ensure = %v, want Ready", state) + } +} + +// The namespace moved to another subsystem: the fabric below now exposes a +// different device. The mapping is reported stale and the heal re-points it +// at the new device; the device above keeps its identity. +func TestDMLinearHealRepointsAMappingWhoseNamespaceMoved(t *testing.T) { + m := newFakeMapper() + layer := dmLayer(m) + if _, err := layer.Ensure(context.Background(), nvmeBelow("/dev/nvme1n1", 259, 3)); err != nil { + t.Fatal(err) + } + moved := nvmeBelow("/dev/nvme4n2", 259, 11) + state, own, err := layer.Observe(context.Background(), moved) + if err != nil || state != volstack.StatePartial { + t.Fatalf("observe after the move = %v %v, want Partial", state, err) + } + if healthy, _ := layer.Healthy(context.Background(), own); healthy { + t.Fatal("a mapping pointing at the old namespace is not healthy") + } + if err := layer.Heal(context.Background(), moved, own); err != nil { + t.Fatalf("heal: %v", err) + } + if got := m.tables["sb-vol-1"].Device; got != "259:11" { + t.Fatalf("mapping points at %s, want the new namespace 259:11", got) + } + if healthy, _ := layer.Healthy(context.Background(), own); !healthy { + t.Fatal("healthy after the swap") + } + if own.Devices[0].Path != "/dev/mapper/sb-vol-1" { + t.Fatalf("the device above changed: %s", own.Devices[0].Path) + } +} + +func TestDMLinearHealWithNothingBelowFailsAndLeavesTheMapping(t *testing.T) { + m := newFakeMapper() + layer := dmLayer(m) + if _, err := layer.Ensure(context.Background(), nvmeBelow("/dev/nvme1n1", 259, 3)); err != nil { + t.Fatal(err) + } + if err := layer.Heal(context.Background(), volstack.Artifact{}, volstack.Artifact{}); err == nil { + t.Fatal("re-pointing at nothing must fail") + } + if got := m.tables["sb-vol-1"].Device; got != "259:3" { + t.Fatalf("a failed heal must leave the mapping, now %s", got) + } +} + +func TestDMLinearAFailedSwapKeepsTheLayerUnhealthy(t *testing.T) { + m := newFakeMapper() + layer := dmLayer(m) + if _, err := layer.Ensure(context.Background(), nvmeBelow("/dev/nvme1n1", 259, 3)); err != nil { + t.Fatal(err) + } + moved := nvmeBelow("/dev/nvme4n2", 259, 11) + _, own, _ := layer.Observe(context.Background(), moved) + m.swapErr = errors.New("reload refused") + if err := layer.Heal(context.Background(), moved, own); err == nil { + t.Fatal("expected the swap error") + } + if healthy, _ := layer.Healthy(context.Background(), own); healthy { + t.Fatal("still stale after a failed swap") + } +} + +func TestDMLinearReleaseRemovesTheMappingAndDestroyKeepsNothing(t *testing.T) { + m := newFakeMapper() + layer := dmLayer(m) + if _, err := layer.Ensure(context.Background(), nvmeBelow("/dev/nvme1n1", 259, 3)); err != nil { + t.Fatal(err) + } + if err := layer.Release(context.Background(), volstack.Artifact{}); err != nil { + t.Fatal(err) + } + if _, ok := m.tables["sb-vol-1"]; ok { + t.Fatal("release must remove the mapping") + } + if err := layer.Destroy(context.Background(), volstack.Artifact{}); err != nil { + t.Fatal(err) + } +} diff --git a/atlas-lib/volstack/plans/node.go b/atlas-lib/volstack/plans/node.go index f31045d40..6e1ea8156 100644 --- a/atlas-lib/volstack/plans/node.go +++ b/atlas-lib/volstack/plans/node.go @@ -13,6 +13,7 @@ import ( "context" "github.com/simplyblock/atlas/blockdev" + "github.com/simplyblock/atlas/devmapper" "github.com/simplyblock/atlas/lvm" "github.com/simplyblock/atlas/lvol" "github.com/simplyblock/atlas/nvme" @@ -55,6 +56,10 @@ type NodeConfig struct { // it nil, and the reading decides alone. PriorFormat func(ctx context.Context, volume Volume) (string, error) + // Mapper runs the device-mapper commands of the dmLinear indirection, and + // defaults to the host's dmsetup. + Mapper layers.DMMapper + // Resolve answers what the kernel says about a device path, and defaults to // blockdev.ResolveDevice. It is a seam only because the logical-volume layer // creates a device-mapper node and has to describe it upward, which a test @@ -95,6 +100,21 @@ func NewNode(cfg NodeConfig) *Node { } // fabric is the bottom layer of every plan: one namespace, attached. +// dmLinear is the indirection between the fabric and what uses the volume +// (docs/consistency-group-colocation.md §6 in sbcli): its device survives a +// move of the volume's namespace to another subsystem. +func (n *Node) dmLinear(volume Volume) volstack.Layer { + mapper := n.cfg.Mapper + if mapper == nil { + mapper = devmapper.New(nil) + } + return layers.NewDMLinear(layers.DMLinearConfig{ + Name: layers.DMLinearName(volume.UUID), + Mapper: mapper, + Resolve: n.cfg.Resolve, + }) +} + func (n *Node) fabric(connection lvol.Connection) volstack.Layer { return layers.NewFabric(layers.FabricConfig{ Connection: connection, diff --git a/atlas-lib/volstack/plans/plans.go b/atlas-lib/volstack/plans/plans.go index 8395a160f..481236621 100644 --- a/atlas-lib/volstack/plans/plans.go +++ b/atlas-lib/volstack/plans/plans.go @@ -183,6 +183,20 @@ func (n *Node) Plain(connection lvol.Connection, volume Volume) volstack.Plan { return volstack.Plan{n.fabric(connection), n.filesystem(volume)} } +// IndirectRawBlock is `fabric` → `dmLinear`: a raw block volume behind the +// device-mapper indirection, so the device the pod holds survives a move of +// the volume's namespace to another subsystem (the indirection's Heal re-points +// it at the namespace the fabric brought up). +func (n *Node) IndirectRawBlock(connection lvol.Connection, volume Volume) volstack.Plan { + return volstack.Plan{n.fabric(connection), n.dmLinear(volume)} +} + +// IndirectPlain is `fabric` → `dmLinear` → `filesystem`: Plain with the +// indirection under the filesystem, so a namespace move does not unmount it. +func (n *Node) IndirectPlain(connection lvol.Connection, volume Volume) volstack.Plan { + return volstack.Plan{n.fabric(connection), n.dmLinear(volume), n.filesystem(volume)} +} + // LVM is `fabric` → `lvmPhysicalVolume` → `lvmVolumeGroup` → `lvmLogicalVolume` // → `filesystem`, the shape a volume with client-side deduplication or // compression takes. What the logical volume is to be lives in options, so this diff --git a/csi-driver/internal/csi/node/indirection_test.go b/csi-driver/internal/csi/node/indirection_test.go new file mode 100644 index 000000000..f2f78a8bb --- /dev/null +++ b/csi-driver/internal/csi/node/indirection_test.go @@ -0,0 +1,77 @@ +package node + +import ( + "context" + "strings" + "testing" +) + +// A volume staged behind the dm indirection is torn down through it: the +// record names the layer, and the teardown walks it. +func TestTeardownPlanOfAnIndirectVolumeWalksTheMapping(t *testing.T) { + for _, layers := range [][]string{ + {"fabric", "dmLinear"}, + {"fabric", "dmLinear", "filesystem"}, + } { + ns, _ := newStackedServer(t, newRecordingRunner()) + writeRecord(t, ns.stack, pvcTestHandle, layers) + plan, err := ns.teardownPlan(context.Background(), pvcTestHandle, "/staging", stagedContext()) + if err != nil { + t.Fatalf("teardownPlan(%v): %v", layers, err) + } + if got, want := strings.Join(plan.Names(), " → "), strings.Join(layers, " → "); got != want { + t.Errorf("the teardown walks %s, want %s", got, want) + } + } +} + +// A fresh stage follows SPDKCSI_DM_INDIRECTION; a staged volume follows its +// record, so the layer is never inserted under (or pulled from under) a live +// consumer by a later heal. +func TestAttachShapeFollowsTheFlagOnlyForAFreshStage(t *testing.T) { + ns, _ := newStackedServer(t, newRecordingRunner()) + + t.Setenv("SPDKCSI_DM_INDIRECTION", "") + if got := ns.attachShape(pvcTestHandle, shapePlain); got != shapePlain { + t.Errorf("flag off, no record: shape %v, want plain", got) + } + t.Setenv("SPDKCSI_DM_INDIRECTION", "true") + if got := ns.attachShape(pvcTestHandle, shapePlain); got != shapeIndirectPlain { + t.Errorf("flag on, no record: shape %v, want indirect plain", got) + } + if got := ns.attachShape(pvcTestHandle, shapeRawBlock); got != shapeIndirectRawBlock { + t.Errorf("flag on, no record: shape %v, want indirect raw block", got) + } + if got := ns.attachShape(pvcTestHandle, shapeLVM); got != shapeLVM { + t.Errorf("an LVM stack has no indirect variant, got %v", got) + } + + writeRecord(t, ns.stack, pvcTestHandle, []string{"fabric", "filesystem"}) + if got := ns.attachShape(pvcTestHandle, shapePlain); got != shapePlain { + t.Errorf("flag on, staged without the layer: shape %v, want plain", got) + } + + t.Setenv("SPDKCSI_DM_INDIRECTION", "") + writeRecord(t, ns.stack, pvcTestHandle, []string{"fabric", "dmLinear", "filesystem"}) + if got := ns.attachShape(pvcTestHandle, shapePlain); got != shapeIndirectPlain { + t.Errorf("flag off, staged with the layer: shape %v, want indirect plain", got) + } +} + +func TestPlanForBuildsTheIndirectRows(t *testing.T) { + s, _ := newTestStack(t, newRecordingRunner()) + node := s.node("", nil) + vc := stagedContext() + volume := stackVolume("/staging", vc, mountCapability()) + build := func(shape stackShape) string { + return strings.Join(planFor(node, connectionFromContext(vc), volume, vdoOptions(vc), shape).Names(), " → ") + } + got := build(shapeIndirectPlain) + if got != "fabric → dmLinear → filesystem" { + t.Errorf("indirect plain = %s", got) + } + got = build(shapeIndirectRawBlock) + if got != "fabric → dmLinear" { + t.Errorf("indirect raw block = %s", got) + } +} diff --git a/csi-driver/internal/csi/node/plan.go b/csi-driver/internal/csi/node/plan.go index aa86c704d..fbe2b86cc 100644 --- a/csi-driver/internal/csi/node/plan.go +++ b/csi-driver/internal/csi/node/plan.go @@ -14,6 +14,7 @@ package node import ( "encoding/json" "fmt" + "os" "slices" "strconv" "strings" @@ -43,6 +44,7 @@ const ( layerLVMPhysicalVolume = "lvmPhysicalVolume" layerLVMVolumeGroup = "lvmVolumeGroup" layerLVMLogicalVolume = "lvmLogicalVolume" + layerDMLinear = "dmLinear" ) // vdoPoolName is the pool `lvcreate --type vdo` creates alongside the logical @@ -76,6 +78,15 @@ const ( // asked for client-side compression or deduplication and is opened as a // block device. shapeLVMRawBlock + + // shapeIndirectRawBlock and shapeIndirectPlain are shapeRawBlock and + // shapePlain with the dmLinear indirection above the fabric, so a volume's + // namespace can move to another subsystem under a live consumer + // (consistency-group co-location). Chosen for a fresh stage when the node + // runs with SPDKCSI_DM_INDIRECTION, and afterwards by the stack record: + // a volume never gains or loses the layer under a staged consumer. + shapeIndirectRawBlock + shapeIndirectPlain ) // planFor is the layer list one of the shapes means, built with the seams the @@ -94,6 +105,10 @@ func planFor( switch shape { case shapeRawBlock: return node.RawBlock(connection) + case shapeIndirectRawBlock: + return node.IndirectRawBlock(connection, volume) + case shapeIndirectPlain: + return node.IndirectPlain(connection, volume) case shapeLVM: return node.LVM(connection, volume, options) case shapeLVMRawBlock: @@ -107,6 +122,29 @@ func planFor( // capability decides whether there is a filesystem, and the class parameters // decide whether the LVM layers that provide client-side compression and // deduplication sit between it and the fabric. +// indirect turns a raw block or plain shape into its dmLinear variant. Only +// those two have one: an LVM stack already re-points through its own device +// mapper nodes, and is left as it is. +func indirect(shape stackShape) stackShape { + switch shape { + case shapeRawBlock: + return shapeIndirectRawBlock + case shapePlain: + return shapeIndirectPlain + } + return shape +} + +// dmIndirectionEnabled is SPDKCSI_DM_INDIRECTION: stage new raw block and +// plain volumes behind the dmLinear indirection. Off by default. +func dmIndirectionEnabled() bool { + switch strings.ToLower(strings.TrimSpace(os.Getenv("SPDKCSI_DM_INDIRECTION"))) { + case "1", "true", "yes", "on": + return true + } + return false +} + func shapeFor(vc map[string]string, volCap *csi.VolumeCapability) stackShape { block := volCap.GetBlock() != nil switch { @@ -402,6 +440,8 @@ var recordedShapes = []struct { }{ {[]string{layerFabric}, shapeRawBlock}, {[]string{layerFabric, layerFilesystem}, shapePlain}, + {[]string{layerFabric, layerDMLinear}, shapeIndirectRawBlock}, + {[]string{layerFabric, layerDMLinear, layerFilesystem}, shapeIndirectPlain}, { []string{layerFabric, layerLVMPhysicalVolume, layerLVMVolumeGroup, layerLVMLogicalVolume}, shapeLVMRawBlock, @@ -421,6 +461,7 @@ var knownLayers = map[string]bool{ layerLVMPhysicalVolume: true, layerLVMVolumeGroup: true, layerLVMLogicalVolume: true, + layerDMLinear: true, } // shapeFromRecord is the plan shape a recorded layer list describes. diff --git a/csi-driver/internal/csi/node/stage.go b/csi-driver/internal/csi/node/stage.go index 004fdb243..c5d59b43d 100644 --- a/csi-driver/internal/csi/node/stage.go +++ b/csi-driver/internal/csi/node/stage.go @@ -18,6 +18,7 @@ import ( "encoding/json" "errors" "fmt" + "slices" "strings" "time" @@ -287,7 +288,32 @@ func (ns *Server) attachPlan( } node := ns.stack.node(hostNQN, ns.priorFormat(volumeID, vc)) volume := stackVolume(stagingTargetPath, vc, volCap) - return planFor(node, connection, volume, vdoOptions(vc), shapeFor(vc, volCap)), nil + shape := ns.attachShape(volumeID, shapeFor(vc, volCap)) + return planFor(node, connection, volume, vdoOptions(vc), shape), nil +} + +// attachShape decides whether a stage or heal builds the dmLinear indirection. +// A volume's stack never gains or loses the layer under a staged consumer: +// with a stack record the record decides (the layer is there or it is not), +// and only a fresh stage, which has none, follows SPDKCSI_DM_INDIRECTION. A +// record that cannot be read keeps the shape without the layer, which is what +// every volume staged before the indirection existed is. +func (ns *Server) attachShape(volumeID string, shape stackShape) stackShape { + record, err := ns.stack.store.Load(volumeID) + switch { + case errors.Is(err, volstack.ErrNoRecord): + if dmIndirectionEnabled() { + return indirect(shape) + } + return shape + case err != nil: + klog.Warningf("volume %s: stack record unreadable (%v); staging without the dm indirection", volumeID, err) + return shape + } + if slices.Contains(recordedLayers(record), layerDMLinear) { + return indirect(shape) + } + return shape } // teardownPlan is the plan an unstage walks, which is the shape that was built From 7545be9539f932dcf719a4df7c2563d80b35600d Mon Sep 17 00:00:00 2001 From: michael Date: Sat, 3 Oct 2026 23:20:22 +0300 Subject: [PATCH 188/206] docs(design): consistency groups point at the co-location design Co-Authored-By: Claude Opus 5.5 --- .../docs/designs/design-consistency-groups.md | 22 +++++++++++++++++++ 1 file changed, 22 insertions(+) diff --git a/operator/docs/designs/design-consistency-groups.md b/operator/docs/designs/design-consistency-groups.md index aa17b552c..d8c605819 100644 --- a/operator/docs/designs/design-consistency-groups.md +++ b/operator/docs/designs/design-consistency-groups.md @@ -208,6 +208,17 @@ A group lives exactly as long as its members. Removing the last member deletes t ### 4.5 Membership after provisioning (Phase 4 — Planned) +> **Update 2026-10-03 (co-location).** Group-wide migration, the pre-join +> live migration of a late joiner, create-time subsystem forcing and namespace +> moves between subsystems are specified in sbcli +> `docs/consistency-group-colocation.md`. In short: the control plane refuses a +> migration that would split a group and migrates a group's whole scope (every +> subsystem holding a member) through `.../consistency-groups/{g}/migration`; +> the CSI watcher turns a join refused for placement into a VolumeMigration to +> the pin and joins afterwards; the rebalancer skips group members; namespace +> moves and the node's dm-linear indirection are behind flags. + + Phase 4 makes the membership label live for the volume's whole life rather than read once at creation. Adding `storage.simplyblock.io/consistency-group` to an existing PVC joins its volume to the named group, and removing the label detaches the volume in §8.2's sense: the epoch closes and every generation that contains the member stays restorable. **What made this safe.** The group's subsystem-scoped operations reach every volume that shares a member's NVMe subsystem: a migration moves the whole subsystem (§9.5), and the frozen cut stalls a shared subsystem's I/O path as one unit. While a subsystem could hold volumes of more than one storage pool, only the create path could guarantee that a member's subsystem held nothing those operations must not touch, which is why membership was fixed at creation. The control plane now enforces subsystem and pool alignment as an invariant (P0-5): a shared subsystem belongs to exactly one pool, checked at lvol placement and again inside the transactional namespace-slot claim, so any volume in the group's pool sits in a subsystem wholly owned by that pool. A late join can therefore no longer entangle another pool's volumes in the group's freeze or migration scope, and the join reduces to the epoch bookkeeping the backend already has. @@ -358,6 +369,17 @@ If a member's snapshot in some generation is gone (pruned, or its volume hard-de ### 8.4 Members are excluded from migration +> **Update 2026-10-03: superseded for whole groups.** Group-wide migration, the pre-join +> live migration of a late joiner, create-time subsystem forcing and namespace +> moves between subsystems are specified in sbcli +> `docs/consistency-group-colocation.md`. In short: the control plane refuses a +> migration that would split a group and migrates a group's whole scope (every +> subsystem holding a member) through `.../consistency-groups/{g}/migration`; +> the CSI watcher turns a join refused for placement into a VolumeMigration to +> the pin and joins afterwards; the rebalancer skips group members; namespace +> moves and the node's dm-linear indirection are behind flags. + + A group's members are pinned to one logical volume store (§4.2), so a member that moved off the store would break the frozen group snapshot. For now, the design excludes consistency-group members from volume migration entirely, and it does so at two layers. The operator's validating webhook on `VolumeMigration` (§9.5) declines the request at `kubectl apply` when the target PV's backing volume is a group member, so the operator never even starts the migration. The backend refusing to migrate a group member is the last line of defense behind it, catching any migration reached by a path the webhook does not cover. Between the two, a group's placement stays fixed for its life and the frozen-snapshot invariant holds by construction rather than being checked after the fact. This is the conservative first cut, and it keeps the feature simple while the group model settles. Migrating a whole group as a unit, moving the shared placement pin together so every member stays colocated, is a later option and is Open Question 2. The §12 failure path, a member found off the pinned store at snapshot time, stays as defense in depth against a member moved by some path other than migration (a node-failure recovery, say), because a group snapshot must fail loudly rather than freeze an inconsistent subset. From b18cab7b64943d67207ee349ea77e1d2686c599b Mon Sep 17 00:00:00 2001 From: michael Date: Sat, 3 Oct 2026 23:34:22 +0300 Subject: [PATCH 189/206] test(node): the record-governed shape case uses a second volume handle Co-Authored-By: Claude Opus 5.5 --- csi-driver/internal/csi/node/indirection_test.go | 12 ++++++++---- 1 file changed, 8 insertions(+), 4 deletions(-) diff --git a/csi-driver/internal/csi/node/indirection_test.go b/csi-driver/internal/csi/node/indirection_test.go index f2f78a8bb..9985886a9 100644 --- a/csi-driver/internal/csi/node/indirection_test.go +++ b/csi-driver/internal/csi/node/indirection_test.go @@ -28,6 +28,10 @@ func TestTeardownPlanOfAnIndirectVolumeWalksTheMapping(t *testing.T) { // A fresh stage follows SPDKCSI_DM_INDIRECTION; a staged volume follows its // record, so the layer is never inserted under (or pulled from under) a live // consumer by a later heal. +// indirectHandle is a second volume, so the record of one never answers for +// another. +const indirectHandle = "0c5b4c4e-6f1d-4b8e-9d2a-3a1f2b7c9e10:pool-1:7d1e0f5a-2b3c-4d5e-8f90-a1b2c3d4e5f6" + func TestAttachShapeFollowsTheFlagOnlyForAFreshStage(t *testing.T) { ns, _ := newStackedServer(t, newRecordingRunner()) @@ -46,14 +50,14 @@ func TestAttachShapeFollowsTheFlagOnlyForAFreshStage(t *testing.T) { t.Errorf("an LVM stack has no indirect variant, got %v", got) } - writeRecord(t, ns.stack, pvcTestHandle, []string{"fabric", "filesystem"}) - if got := ns.attachShape(pvcTestHandle, shapePlain); got != shapePlain { + writeRecord(t, ns.stack, indirectHandle, []string{"fabric", "filesystem"}) + if got := ns.attachShape(indirectHandle, shapePlain); got != shapePlain { t.Errorf("flag on, staged without the layer: shape %v, want plain", got) } t.Setenv("SPDKCSI_DM_INDIRECTION", "") - writeRecord(t, ns.stack, pvcTestHandle, []string{"fabric", "dmLinear", "filesystem"}) - if got := ns.attachShape(pvcTestHandle, shapePlain); got != shapeIndirectPlain { + writeRecord(t, ns.stack, indirectHandle, []string{"fabric", "dmLinear", "filesystem"}) + if got := ns.attachShape(indirectHandle, shapePlain); got != shapeIndirectPlain { t.Errorf("flag off, staged with the layer: shape %v, want indirect plain", got) } } From 1d7008c72d86f7bca4a7bc02522aa395622c2f5e Mon Sep 17 00:00:00 2001 From: michael Date: Sun, 4 Oct 2026 00:27:46 +0300 Subject: [PATCH 190/206] chart: the Control Center may update site profiles Edit bindings in the console patches a SiteProfile; the console's ClusterRole only read them, so the patch was refused on the DR test bed (2026-10-04). Co-Authored-By: Claude Opus 5.5 --- .../simplyblock-operator/templates/control-center-rbac.yaml | 5 +++++ 1 file changed, 5 insertions(+) diff --git a/helm-charts/charts/simplyblock-operator/templates/control-center-rbac.yaml b/helm-charts/charts/simplyblock-operator/templates/control-center-rbac.yaml index 4a638e58c..29ac17f9b 100644 --- a/helm-charts/charts/simplyblock-operator/templates/control-center-rbac.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/control-center-rbac.yaml @@ -144,6 +144,11 @@ rules: - apiGroups: ["sitemap.simplyblock.io"] resources: ["dhcpservers"] verbs: ["create", "update", "patch", "delete"] + # Site profile bindings (logical and guest networks, the DHCP server) are + # edited in the console; dr-hub creates and deletes the profiles themselves. + - apiGroups: ["sitemap.simplyblock.io"] + resources: ["siteprofiles"] + verbs: ["update", "patch"] # a managed site's storage deployment is requested, sized and approved # from the hub console (StorageSiteDeployment, carried to the site by OCM) - apiGroups: ["storage.simplyblock.io"] From 94ca3d90bc72ec0448572bc6a3cdc4d9eb756e90 Mon Sep 17 00:00:00 2001 From: michael Date: Sun, 4 Oct 2026 02:43:19 +0300 Subject: [PATCH 191/206] csi-driver: GetReplicationDestinationInfo, so Ramen can restore a consistency group kubernetes-csi-addons v0.15 lists every member of a VolumeGroupReplicationContent in status.persistentVolumeMappingList and fills each destinationVolumeHandle and status.destinationVolumeGroupID only from GetReplicationDestinationInfo. The driver did not offer it, so the destinations stayed empty and Ramen's restore on the target refused the group: "destination volume ID is empty for VGRC ..." -- the first consistency-group relocate on the DR test bed stopped with the application demoted on the source and not started on the target (2026-10-03). The driver advertises GET_REPLICATION_DESTINATION_INFO and answers with the source's own handles: a simplyblock PV keeps its original handle across moves and every Replication verb resolves it through the replication relationship, the behaviour every relocate and fail-over case was proven with. - a volume: its own handle, with no lookup -- a failure would set DestinationInfoAvailable=False and make Ramen refuse the VRG - a group: its own cg: handle and a complete identity map over its current members, keyed by their volume handles; UNAVAILABLE when the group cannot be listed, never a partial map atlas-lib gains ConsistencyGroupMemberHandles: the member listing carries bare lvol ids, and the shared pool (a group never spans pools) is resolved from the first member. Co-Authored-By: Claude Opus 5.5 --- atlas-lib/controlplane/consistencygroups.go | 57 +++++++++ .../consistencygroups_members_test.go | 72 ++++++++++++ .../csi/controller/mock_controlplane_test.go | 11 +- .../csi/controller/replication_destination.go | 105 +++++++++++++++++ .../replication_destination_test.go | 109 ++++++++++++++++++ .../csi/csiaddons/identity/identity.go | 11 ++ .../csi/csiaddons/identity/identity_test.go | 19 +++ 7 files changed, 380 insertions(+), 4 deletions(-) create mode 100644 atlas-lib/controlplane/consistencygroups_members_test.go create mode 100644 csi-driver/internal/csi/controller/replication_destination.go create mode 100644 csi-driver/internal/csi/controller/replication_destination_test.go diff --git a/atlas-lib/controlplane/consistencygroups.go b/atlas-lib/controlplane/consistencygroups.go index d45464253..35fa05edb 100644 --- a/atlas-lib/controlplane/consistencygroups.go +++ b/atlas-lib/controlplane/consistencygroups.go @@ -6,11 +6,14 @@ package controlplane import ( "context" + "errors" "fmt" "github.com/google/uuid" + "github.com/simplyblock/atlas/errs" "github.com/simplyblock/atlas/internal/cpapi" + "github.com/simplyblock/atlas/lvol" ) // ConsistencyGroupForLvols returns the id of the backend consistency group in @@ -93,3 +96,57 @@ func sameStringSet(ids []string, want map[string]bool) bool { } return true } + +// ConsistencyGroupMemberHandles returns the volume handles +// ("::") of a consistency group's current (open-epoch) +// members, in the order the control plane lists them. +// +// The member listing carries bare lvol ids. Every member of one group lives in +// one storage pool (the backend refuses a cross-pool join), so the pool is +// resolved once, from the first member, by probing the cluster's pools. An +// empty group is an error: a caller mapping members (csi-addons destination +// info) must never return an empty map as if it were complete. +func (c *Client) ConsistencyGroupMemberHandles( + ctx context.Context, gh lvol.GroupHandle, +) ([]lvol.VolumeHandle, error) { + cluster, err := parseUUID("cluster id", gh.ClusterID) + if err != nil { + return nil, err + } + group, err := parseUUID("consistency group id", gh.GroupID) + if err != nil { + return nil, err + } + members, err := c.consistencyGroupMembers(ctx, cluster, group) + if err != nil { + return nil, err + } + if len(members) == 0 { + return nil, fmt.Errorf("consistency group %s has no current members", gh) + } + pools, err := c.ListStoragePools(ctx, gh.ClusterID) + if err != nil { + return nil, err + } + poolID := "" + for _, p := range pools { + h := lvol.Handle{ClusterID: gh.ClusterID, PoolRef: p.ID, VolumeID: members[0]} + _, err := c.Volume(ctx, h.Handle()) + if err == nil { + poolID = p.ID + break + } + if !errors.Is(err, errs.ErrNotFound) { + return nil, err + } + } + if poolID == "" { + return nil, fmt.Errorf("member %s of consistency group %s is in none of the %d pools of cluster %s: %w", + members[0], gh, len(pools), gh.ClusterID, errs.ErrNotFound) + } + out := make([]lvol.VolumeHandle, 0, len(members)) + for _, m := range members { + out = append(out, lvol.Handle{ClusterID: gh.ClusterID, PoolRef: poolID, VolumeID: m}.Handle()) + } + return out, nil +} diff --git a/atlas-lib/controlplane/consistencygroups_members_test.go b/atlas-lib/controlplane/consistencygroups_members_test.go new file mode 100644 index 000000000..f57b42a42 --- /dev/null +++ b/atlas-lib/controlplane/consistencygroups_members_test.go @@ -0,0 +1,72 @@ +package controlplane + +import ( + "context" + "net/http" + "strings" + "testing" + + "github.com/simplyblock/atlas/lvol" +) + +// The member listing carries bare lvol ids; the handles need the pool, which +// the group's first member is probed for across the cluster's pools. Here it +// lives in the SECOND pool, and a removed member is left out. +func TestConsistencyGroupMemberHandlesResolvesThePool(t *testing.T) { + const ( + group = "c9c9c9c9-c9c9-4c9c-8c9c-c9c9c9c9c9c9" + m1 = "a1111111-1111-4111-8111-111111111111" + m2 = "b2222222-2222-4222-8222-222222222222" + gone = "d4444444-4444-4444-8444-444444444444" + otherPool = "55555555-5555-5555-5555-555555555555" + ) + pool := func(id, name string) string { + return `{"id":"` + id + `","cluster_id":"` + testCluster + `","name":"` + name + `",` + + `"max_size":1000,"capacity":{},"max_r_mbytes":0,"max_rw_iops":0,"max_rw_mbytes":0,` + + `"max_w_mbytes":0,"volume_max_size":0,"status":"active"}` + } + member := func(id string, removed int) string { + return `{"lvol_id":"` + id + `","joined_seq":1,"removed_seq":` + string(rune('0'+removed)) + + `,"node_id":"n","lvs_name":"l","online":true}` + } + c := newTestClient(t, func(w http.ResponseWriter, r *http.Request) { + w.Header().Set("Content-Type", "application/json") + switch { + case strings.HasSuffix(r.URL.Path, "/consistency-groups/"+group+"/members"): + _, _ = w.Write([]byte("[" + member(m1, 0) + "," + member(gone, 3) + "," + member(m2, 0) + "]")) + case strings.HasSuffix(r.URL.Path, "/storage-pools/"): + _, _ = w.Write([]byte("[" + pool(otherPool, "other") + "," + pool(testPool, "pool1") + "]")) + case strings.Contains(r.URL.Path, "/storage-pools/"+testPool+"/volumes/"+m1): + _, _ = w.Write([]byte(`{"id":"` + m1 + `","name":"pvc-1","pool_name":"pool1",` + + `"size":20971520,"ns_id":1,"nqn":"nqn.2023-02.io.simplyblock:c:lvol:v"}`)) + default: + w.WriteHeader(http.StatusNotFound) + _, _ = w.Write([]byte(`{"detail":"not found"}`)) + } + }) + + got, err := c.ConsistencyGroupMemberHandles(context.Background(), + lvol.GroupHandle{ClusterID: testCluster, GroupID: group}) + if err != nil { + t.Fatal(err) + } + want := []lvol.VolumeHandle{ + lvol.VolumeHandle(testCluster + ":" + testPool + ":" + m1), + lvol.VolumeHandle(testCluster + ":" + testPool + ":" + m2), + } + if len(got) != len(want) || got[0] != want[0] || got[1] != want[1] { + t.Fatalf("handles = %v, want %v", got, want) + } +} + +func TestConsistencyGroupMemberHandlesOfAnEmptyGroupIsAnError(t *testing.T) { + c := newTestClient(t, func(w http.ResponseWriter, r *http.Request) { + w.Header().Set("Content-Type", "application/json") + _, _ = w.Write([]byte(`[]`)) + }) + _, err := c.ConsistencyGroupMemberHandles(context.Background(), + lvol.GroupHandle{ClusterID: testCluster, GroupID: "c9c9c9c9-c9c9-4c9c-8c9c-c9c9c9c9c9c9"}) + if err == nil { + t.Fatal("an empty group answered without an error") + } +} diff --git a/csi-driver/internal/csi/controller/mock_controlplane_test.go b/csi-driver/internal/csi/controller/mock_controlplane_test.go index 1d29d950f..49f75dfb0 100644 --- a/csi-driver/internal/csi/controller/mock_controlplane_test.go +++ b/csi-driver/internal/csi/controller/mock_controlplane_test.go @@ -328,9 +328,11 @@ func (m *mockSBCLI) lookupSnapshot(w http.ResponseWriter, snapshotID string) *mo } func (m *mockSBCLI) handleListPools(w http.ResponseWriter, _ *http.Request) { - writeJSON(w, http.StatusOK, []map[string]string{ - {"name": sanityPoolName, "id": sanityPoolUUID}, - }) + writeJSON(w, http.StatusOK, []map[string]any{{ + "name": sanityPoolName, "id": sanityPoolUUID, "cluster_id": sanityClusterID, + "max_size": 0, "capacity": map[string]any{}, "max_r_mbytes": 0, "max_rw_iops": 0, + "max_rw_mbytes": 0, "max_w_mbytes": 0, "volume_max_size": 0, "status": "active", + }}) } func (m *mockSBCLI) handleListVolumes(w http.ResponseWriter, _ *http.Request) { @@ -354,7 +356,8 @@ func (m *mockSBCLI) handleGetVolume(w http.ResponseWriter, r *http.Request) { } writeJSON(w, http.StatusOK, map[string]any{ "id": volume.UUID, "name": volume.Name, "size": volume.Size, "status": volume.status(), - "group_id": volume.GroupID, + "group_id": volume.GroupID, "pool_name": sanityPoolName, "ns_id": 1, + "nqn": "nqn.2023-02.io.simplyblock:" + sanityClusterID + ":lvol:" + volume.UUID, }) } diff --git a/csi-driver/internal/csi/controller/replication_destination.go b/csi-driver/internal/csi/controller/replication_destination.go new file mode 100644 index 000000000..32c882e05 --- /dev/null +++ b/csi-driver/internal/csi/controller/replication_destination.go @@ -0,0 +1,105 @@ +package controller + +import ( + "context" + + "github.com/csi-addons/spec/lib/go/replication" + "google.golang.org/grpc/codes" + "google.golang.org/grpc/status" + + "github.com/simplyblock/atlas/lvol" + "github.com/simplyblock/csi-driver/internal/clusters" + csicommon "github.com/simplyblock/csi-driver/internal/csi/common" +) + +// GetReplicationDestinationInfo answers csi-addons' destination-info call +// (capability GET_REPLICATION_DESTINATION_INFO, csi-addons/spec replication.proto, +// kubernetes-csi-addons >= v0.15.0). +// +// The csi-addons v0.15 controller lists every member of a +// VolumeGroupReplicationContent in status.persistentVolumeMappingList and fills +// each destinationVolumeHandle (and status.destinationVolumeGroupID) only from +// this call. Without it the destinations stay empty, and Ramen's restore on the +// target refuses the whole group: "destination volume ID is empty for VGRC …" +// (vrg_volgrouprep.go updateVGRCVolumeHandlesForRestore). That stopped the +// first consistency-group relocate on the DR test bed with the application +// demoted on the source and not started on the target (2026-10-03, WordPress +// site-a -> site-b). A single volume did not need it: Ramen skips a +// VolumeReplication without the destination-info condition. +// +// The destination of a simplyblock volume or group is addressed by its own +// handle. One control plane manages both clusters of a replication pair; every +// Replication verb resolves a handle through the replication relationship to +// the member it must act on (resolveChain, the group endpoints), and the PV +// keeps the original handle across every move -- the behaviour all relocate and +// fail-over cases were proven with. Answering with the replica's raw volume +// (the landing copy on the peer) would make Ramen rewrite the restored PV to a +// volume that a promote replaces with a clone, bypassing that resolution. So +// the answer is the source handle itself. +// +// A single volume's answer needs no lookup and never fails for a valid handle: +// once the capability is advertised csi-addons asks for every +// VolumeReplication too, and a failure sets DestinationInfoAvailable=False, +// which makes Ramen refuse the VRG instead of skipping it (vrg_volrep.go +// destinationInfoAvailableOrSkip) -- an error here would block protecting a +// volume that has not replicated yet. A group needs its current members for the +// complete map; a group the control plane cannot list is UNAVAILABLE +// (retryable). +func (cs *Server) GetReplicationDestinationInfo( + ctx context.Context, + req *replication.GetReplicationDestinationInfoRequest, +) (*replication.GetReplicationDestinationInfoResponse, error) { + if g := req.GetReplicationSource().GetVolumegroup().GetVolumeGroupId(); g != "" { + return groupDestinationInfo(ctx, g) + } + volumeID := req.GetReplicationSource().GetVolume().GetVolumeId() + if volumeID == "" { + return nil, status.Error(codes.InvalidArgument, "replication source names no volume or volume group") + } + if _, err := csicommon.ParseVolumeHandle(volumeID); err != nil { + return nil, status.Error(codes.InvalidArgument, err.Error()) + } + return &replication.GetReplicationDestinationInfoResponse{ + ReplicationDestination: &replication.ReplicationDestination{ + Type: &replication.ReplicationDestination_Volume{ + Volume: &replication.ReplicationDestination_VolumeDestination{VolumeId: volumeID}, + }, + }, + }, nil +} + +// groupDestinationInfo is the group branch: the group's own handle, and a +// complete source -> destination map over the group's current members (the spec +// forbids a partial map). The keys are the members' volume handles exactly as +// their PersistentVolumes carry them, which is what csi-addons matches them +// against. +func groupDestinationInfo( + ctx context.Context, groupID string, +) (*replication.GetReplicationDestinationInfoResponse, error) { + gh, ok := lvol.ParseGroupHandle(lvol.VolumeHandle(groupID)) + if !ok { + return nil, status.Errorf(codes.InvalidArgument, "invalid volume group handle %q", groupID) + } + client, err := clusters.ReplicationClient(ctx, gh.ClusterID) + if err != nil { + return nil, status.Error(codes.Unavailable, err.Error()) + } + members, err := client.ConsistencyGroupMemberHandles(ctx, gh) + if err != nil { + return nil, status.Errorf(codes.Unavailable, "members of %s: %v", groupID, err) + } + ids := make(map[string]string, len(members)) + for _, m := range members { + ids[string(m)] = string(m) + } + return &replication.GetReplicationDestinationInfoResponse{ + ReplicationDestination: &replication.ReplicationDestination{ + Type: &replication.ReplicationDestination_Volumegroup{ + Volumegroup: &replication.ReplicationDestination_VolumeGroupDestination{ + VolumeGroupId: groupID, + VolumeIds: ids, + }, + }, + }, + }, nil +} diff --git a/csi-driver/internal/csi/controller/replication_destination_test.go b/csi-driver/internal/csi/controller/replication_destination_test.go new file mode 100644 index 000000000..e8db93fb5 --- /dev/null +++ b/csi-driver/internal/csi/controller/replication_destination_test.go @@ -0,0 +1,109 @@ +package controller + +import ( + "context" + "testing" + + "github.com/csi-addons/spec/lib/go/replication" + "google.golang.org/grpc/codes" + "google.golang.org/grpc/status" +) + +func volumeSource(id string) *replication.ReplicationSource { + return &replication.ReplicationSource{ + Type: &replication.ReplicationSource_Volume{ + Volume: &replication.ReplicationSource_VolumeSource{VolumeId: id}, + }, + } +} + +// A single volume's destination is its own handle: the PV keeps it across +// moves and every Replication verb resolves it to the replica. The answer needs +// no relationship yet -- an error would make Ramen refuse the VRG. +func TestReplicationDestinationOfAVolumeIsItsOwnHandle(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newTestControllerServer(t, mock) + + resp, err := cs.GetReplicationDestinationInfo(context.Background(), + &replication.GetReplicationDestinationInfoRequest{ReplicationSource: volumeSource(testReplVolID)}) + if err != nil { + t.Fatalf("GetReplicationDestinationInfo: %v", err) + } + if got := resp.GetReplicationDestination().GetVolume().GetVolumeId(); got != testReplVolID { + t.Fatalf("destination volume = %q, want the source handle %q", got, testReplVolID) + } + if resp.GetReplicationDestination().GetVolumegroup() != nil { + t.Fatal("a volume source answered with a group destination") + } +} + +// A group's answer is its own handle and a COMPLETE map over its current +// members, keyed by the members' volume handles exactly as their PVs carry +// them: csi-addons matches persistentVolumeMappingList[].volumeHandle against +// the keys and leaves the destination empty for any miss, which is what made +// Ramen refuse the restore (2026-10-03). +func TestReplicationDestinationOfAGroupMapsEveryMember(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newGroupReplTestServer(t, mock) + + resp, err := cs.GetReplicationDestinationInfo(context.Background(), + &replication.GetReplicationDestinationInfoRequest{ReplicationSource: groupSource()}) + if err != nil { + t.Fatalf("GetReplicationDestinationInfo: %v", err) + } + vg := resp.GetReplicationDestination().GetVolumegroup() + if vg == nil { + t.Fatal("a group source answered without a group destination") + } + if vg.GetVolumeGroupId() != vgGroupHandle { + t.Fatalf("destination group = %q, want %q", vg.GetVolumeGroupId(), vgGroupHandle) + } + want := map[string]string{ + sanityClusterID + ":" + sanityPoolUUID + ":" + vgMember1: sanityClusterID + ":" + sanityPoolUUID + ":" + vgMember1, + sanityClusterID + ":" + sanityPoolUUID + ":" + vgMember2: sanityClusterID + ":" + sanityPoolUUID + ":" + vgMember2, + } + got := vg.GetVolumeIds() + if len(got) != len(want) { + t.Fatalf("volume_ids = %v, want %v", got, want) + } + for k, v := range want { + if got[k] != v { + t.Fatalf("volume_ids[%s] = %q, want %q (all: %v)", k, got[k], v, got) + } + } +} + +func TestReplicationDestinationRejectsAnEmptyOrInvalidSource(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newTestControllerServer(t, mock) + + for name, src := range map[string]*replication.ReplicationSource{ + "empty": nil, + "not-handle": volumeSource("not-a-handle"), + "bad-group": {Type: &replication.ReplicationSource_Volumegroup{ + Volumegroup: &replication.ReplicationSource_VolumeGroupSource{VolumeGroupId: "cg:x:y"}, + }}, + } { + _, err := cs.GetReplicationDestinationInfo(context.Background(), + &replication.GetReplicationDestinationInfoRequest{ReplicationSource: src}) + if status.Code(err) != codes.InvalidArgument { + t.Errorf("%s: code %v (%v), want InvalidArgument", name, status.Code(err), err) + } + } +} + +// A group the control plane does not know is retryable, never an empty map. +func TestReplicationDestinationOfAnUnknownGroupIsUnavailable(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newTestControllerServer(t, mock) + + _, err := cs.GetReplicationDestinationInfo(context.Background(), + &replication.GetReplicationDestinationInfoRequest{ReplicationSource: groupSource()}) + if status.Code(err) != codes.Unavailable { + t.Fatalf("code %v (%v), want Unavailable", status.Code(err), err) + } +} diff --git a/csi-driver/internal/csi/csiaddons/identity/identity.go b/csi-driver/internal/csi/csiaddons/identity/identity.go index 7e659d6e4..4f4bf3bff 100644 --- a/csi-driver/internal/csi/csiaddons/identity/identity.go +++ b/csi-driver/internal/csi/csiaddons/identity/identity.go @@ -49,6 +49,17 @@ func (s *Server) GetCapabilities( }, }, }, + // GetReplicationDestinationInfo (kubernetes-csi-addons >= v0.15.0): + // without it the VolumeGroupReplicationContent's per-volume + // destinations stay empty and Ramen cannot restore a consistency + // group on the target (2026-10-03). + { + Type: &identity.Capability_VolumeReplication_{ + VolumeReplication: &identity.Capability_VolumeReplication{ + Type: identity.Capability_VolumeReplication_GET_REPLICATION_DESTINATION_INFO, + }, + }, + }, // The VolumeGroup service (design §14.3), which the stock // kubernetes-csi-addons controller-manager dials to form a backend // consistency group before replicating it as one unit. diff --git a/csi-driver/internal/csi/csiaddons/identity/identity_test.go b/csi-driver/internal/csi/csiaddons/identity/identity_test.go index d9e18a4dc..ffbdd0956 100644 --- a/csi-driver/internal/csi/csiaddons/identity/identity_test.go +++ b/csi-driver/internal/csi/csiaddons/identity/identity_test.go @@ -82,3 +82,22 @@ func TestProbeReportsReady(t *testing.T) { t.Errorf("Ready = %v, want nil (assume ready)", resp.Ready) } } + +// Without GET_REPLICATION_DESTINATION_INFO the csi-addons v0.15 controller never +// asks for destinations, a group's VolumeGroupReplicationContent keeps empty +// destination handles, and Ramen cannot restore the group on the target +// (2026-10-03). +func TestGetCapabilitiesAdvertisesReplicationDestinationInfo(t *testing.T) { + s := New("test.csi.simplyblock.io", "v1.2.3") + resp, err := s.GetCapabilities(context.Background(), &identity.GetCapabilitiesRequest{}) + if err != nil { + t.Fatal(err) + } + for _, c := range resp.Capabilities { + if vr := c.GetVolumeReplication(); vr != nil && + vr.Type == identity.Capability_VolumeReplication_GET_REPLICATION_DESTINATION_INFO { + return + } + } + t.Error("capabilities do not advertise GET_REPLICATION_DESTINATION_INFO") +} From 1f94d8e075d637db50f92911247c9934d92f4453 Mon Sep 17 00:00:00 2001 From: michael Date: Sun, 4 Oct 2026 14:02:41 +0300 Subject: [PATCH 192/206] csi-driver: a group handle resolves to where the group's data lives now A VolumeGroupReplication keeps its original group handle across a relocate, while the group it names is emptied by design -- its demoted members are deleted so a relocate back stays possible -- and the data lives in the peer group of the same name on the other site. Every group verb reached the empty source group; GetReplicationDestinationInfo listed it and Ramen's VRG on the target waited for destination info for ever (2026-10-04, WordPress A -> B). The control plane resolves a group handle (sbcli GET /consistency-groups/{id}/replication/resolution, added to shared/openapi.json, client regenerated; atlas ResolveGroup). The verbs act on it like the per-volume chain: Info, Resync and Enable on the live group; Demote and Disable on the group of a cluster flagged local; Promote stays on the named group, whose control-plane fail-over resolves the peer itself. GetReplicationDestinationInfo maps every PV handle the group protected (the lineages' origins) to itself. CreateVolumeGroup of PVs whose volumes were replaced groups the live volumes at the ends of their chains. A control plane without the endpoint keeps the previous behaviour. Co-Authored-By: Claude Opus 5.5 --- atlas-lib/controlplane/groupreplication.go | 69 ++++++ atlas-lib/internal/cpapi/cpapi.gen.go | 219 ++++++++++++++++++ .../csi/controller/mock_controlplane_test.go | 20 ++ .../internal/csi/controller/replication.go | 32 ++- .../csi/controller/replication_destination.go | 31 ++- .../controller/replication_groupresolve.go | 98 ++++++++ .../replication_groupresolve_test.go | 214 +++++++++++++++++ .../internal/csi/controller/volumegroup.go | 48 ++++ shared/openapi.json | 115 +++++++++ 9 files changed, 835 insertions(+), 11 deletions(-) create mode 100644 csi-driver/internal/csi/controller/replication_groupresolve.go create mode 100644 csi-driver/internal/csi/controller/replication_groupresolve_test.go diff --git a/atlas-lib/controlplane/groupreplication.go b/atlas-lib/controlplane/groupreplication.go index 4b806d867..987358209 100644 --- a/atlas-lib/controlplane/groupreplication.go +++ b/atlas-lib/controlplane/groupreplication.go @@ -6,6 +6,7 @@ package controlplane import ( + "bytes" "context" "fmt" "net/http" @@ -172,3 +173,71 @@ func derefInt(p *int) int { } return *p } + +// GroupResolution is where a consistency group's data lives now, keyed by the +// handles its PersistentVolumes keep (sbcli GET +// /consistency-groups/{id}/replication/resolution). +type GroupResolution struct { + // Active is the group holding live members: the group itself while it has + // any, else its peer group of the same name on another cluster. Nil when no + // group holds a live member. + Active *lvol.GroupHandle + // Members has one entry per protected volume with a live volume at the end + // of its lineage: Origin is the handle its PV carries, Active the volume + // serving the data now. + Members []GroupMemberResolution + // Legacy reports a control plane without the resolution endpoint (404): the + // caller then treats the group as live where it is, as before. + Legacy bool +} + +// GroupMemberResolution maps one PV handle to the volume serving its data. +type GroupMemberResolution struct { + Origin lvol.VolumeHandle + Active lvol.VolumeHandle +} + +// ResolveGroup resolves a group handle to where the group's data lives now. +// +// A VolumeGroupReplication keeps its original group handle across a relocate, +// while the group it names is emptied by design (its demoted members are +// deleted so a relocate back stays possible) and the data moves to the peer +// group as clones. The group verbs resolve the handle here, the group analogue +// of the per-volume relationship chain (2026-10-04: WordPress's VRG waited for +// destination info for ever against the emptied source group). +func (c *Client) ResolveGroup(ctx context.Context, gh lvol.GroupHandle) (GroupResolution, error) { + cluster, group, err := groupIDs(gh) + if err != nil { + return GroupResolution{}, err + } + resp, err := c.api.ClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGetWithResponse( + ctx, cluster, group) + if err != nil { + return GroupResolution{}, fmt.Errorf("resolve group %s: %w", gh.Handle(), err) + } + if resp.StatusCode() == http.StatusNotFound && resp.JSON200 == nil && !groupNotFound(resp.Body) { + return GroupResolution{Active: &gh, Legacy: true}, nil + } + d, err := payload("resolve group "+string(gh.Handle()), resp.JSON200, resp.StatusCode(), resp.Body) + if err != nil { + return GroupResolution{}, err + } + out := GroupResolution{} + if d.ActiveGroupId != nil && *d.ActiveGroupId != "" && d.ActiveClusterId != nil && *d.ActiveClusterId != "" { + out.Active = &lvol.GroupHandle{ClusterID: *d.ActiveClusterId, GroupID: *d.ActiveGroupId} + } + if d.Members != nil { + for _, m := range *d.Members { + out.Members = append(out.Members, GroupMemberResolution{ + Origin: lvol.VolumeHandle(m.OriginHandle), Active: lvol.VolumeHandle(m.ActiveHandle)}) + } + } + return out, nil +} + +// groupNotFound tells the group-level 404 ("ConsistencyGroup not found", +// the resource dependency's answer) from a route-level 404 (FastAPI's +// {"detail":"Not Found"} on a control plane that predates the endpoint). +func groupNotFound(body []byte) bool { + return bytes.Contains(bytes.ToLower(body), []byte("consistencygroup")) +} diff --git a/atlas-lib/internal/cpapi/cpapi.gen.go b/atlas-lib/internal/cpapi/cpapi.gen.go index 7aec772a2..6e63c361a 100644 --- a/atlas-lib/internal/cpapi/cpapi.gen.go +++ b/atlas-lib/internal/cpapi/cpapi.gen.go @@ -1037,6 +1037,13 @@ type ConsistencyGroupGenerationMemberDTO struct { SnapshotId string `json:"snapshot_id"` } +// ConsistencyGroupLineageMemberDTO One protected volume of a consistency group: the handle its +// PersistentVolume carries and the volume serving its data now. +type ConsistencyGroupLineageMemberDTO struct { + ActiveHandle string `json:"active_handle"` + OriginHandle string `json:"origin_handle"` +} + // ConsistencyGroupMemberDTO One current member of a consistency group (design §10 /members). type ConsistencyGroupMemberDTO struct { JoinedSeq int `json:"joined_seq"` @@ -1086,6 +1093,16 @@ type ConsistencyGroupReplicationStatusDTORole string // ConsistencyGroupReplicationStatusDTOState defines model for ConsistencyGroupReplicationStatusDTO.State. type ConsistencyGroupReplicationStatusDTOState string +// ConsistencyGroupResolutionDTO Where a consistency group's data lives now (replication_policy_controller. +// resolve_group). “active_*“ are empty when no group holds a live member. +type ConsistencyGroupResolutionDTO struct { + ActiveClusterId *string `json:"active_cluster_id,omitempty"` + ActiveGroupId *string `json:"active_group_id,omitempty"` + ClusterId string `json:"cluster_id"` + GroupId string `json:"group_id"` + Members *[]ConsistencyGroupLineageMemberDTO `json:"members,omitempty"` +} + // DeviceDTO defines model for DeviceDTO. type DeviceDTO struct { BdevType *string `json:"bdev_type,omitempty"` @@ -2823,6 +2840,21 @@ type ClientInterface interface { // Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/failover (the `ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPost` operationId). ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPost(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) + // ClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGet Clusters:Consistency-Groups:Replication:Resolution + // + // Where the group's data lives now, keyed by the handles its PVs keep. + // + // After a relocate the group a VGR names is empty -- its demoted members were + // deleted so the way back stays open -- and its data lives in the peer group of + // the same name. The CSI driver resolves the VGR's original group handle here: + // the group holding live members, and each protected volume's original handle + // with the volume serving it now (2026-10-04: WordPress's VRG waited for + // destination info for ever against the emptied source group). Never a 404 + // for an existing group: ``active_group_id`` is empty when nothing serves it. + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/resolution (the `ClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGet` operationId). + ClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGet(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) + // ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGet Clusters:Consistency-Groups:Replication:Status // // The group's replication status as one unit: oldest recovery point, worst @@ -4456,6 +4488,31 @@ func (c *Client) ClustersConsistencyGroupsReplicationFailoverApiV2ClustersCluste return c.Client.Do(req) } +// ClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGet Clusters:Consistency-Groups:Replication:Resolution +// +// Where the group's data lives now, keyed by the handles its PVs keep. +// +// After a relocate the group a VGR names is empty -- its demoted members were +// deleted so the way back stays open -- and its data lives in the peer group of +// the same name. The CSI driver resolves the VGR's original group handle here: +// the group holding live members, and each protected volume's original handle +// with the volume serving it now (2026-10-04: WordPress's VRG waited for +// destination info for ever against the emptied source group). Never a 404 +// for an existing group: “active_group_id“ is empty when nothing serves it. +// +// Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/resolution (the `ClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGet` operationId). +func (c *Client) ClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGet(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*http.Response, error) { + req, err := NewClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGetRequest(c.Server, clusterId, groupId) + if err != nil { + return nil, err + } + req = req.WithContext(ctx) + if err := c.applyEditors(ctx, req, reqEditors); err != nil { + return nil, err + } + return c.Client.Do(req) +} + // ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGet Clusters:Consistency-Groups:Replication:Status // // The group's replication status as one unit: oldest recovery point, worst @@ -8111,6 +8168,47 @@ func NewClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsis return req, nil } +// NewClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGetRequest constructs an http.Request for the ClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGet method +func NewClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGetRequest(server string, clusterId openapi_types.UUID, groupId openapi_types.UUID) (*http.Request, error) { + var err error + + var pathParam0 string + + pathParam0, err = runtime.StyleParamWithOptions("simple", false, "cluster_id", clusterId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + var pathParam1 string + + pathParam1, err = runtime.StyleParamWithOptions("simple", false, "group_id", groupId, runtime.StyleParamOptions{ParamLocation: runtime.ParamLocationPath, Type: "string", Format: "uuid"}) + if err != nil { + return nil, err + } + + serverURL, err := url.Parse(server) + if err != nil { + return nil, err + } + + operationPath := fmt.Sprintf("/api/v2/clusters/%s/consistency-groups/%s/replication/resolution", pathParam0, pathParam1) + if operationPath[0] == '/' { + operationPath = "." + operationPath + } + + queryURL, err := serverURL.Parse(operationPath) + if err != nil { + return nil, err + } + + req, err := http.NewRequest(http.MethodGet, queryURL.String(), nil) + if err != nil { + return nil, err + } + + return req, nil +} + // NewClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetRequest constructs an http.Request for the ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGet method func NewClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetRequest(server string, clusterId openapi_types.UUID, groupId openapi_types.UUID) (*http.Request, error) { var err error @@ -14041,6 +14139,23 @@ type ClientWithResponsesInterface interface { // Corresponds with POST /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/failover (the `ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPost` operationId). ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPostWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPostResponse, error) + // ClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGetWithResponse Clusters:Consistency-Groups:Replication:Resolution + // + // Where the group's data lives now, keyed by the handles its PVs keep. + // + // After a relocate the group a VGR names is empty -- its demoted members were + // deleted so the way back stays open -- and its data lives in the peer group of + // the same name. The CSI driver resolves the VGR's original group handle here: + // the group holding live members, and each protected volume's original handle + // with the volume serving it now (2026-10-04: WordPress's VRG waited for + // destination info for ever against the emptied source group). Never a 404 + // for an existing group: ``active_group_id`` is empty when nothing serves it. + // + // Returns a wrapper object for the known response body format(s). + // + // Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/resolution (the `ClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGet` operationId). + ClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGetResponse, error) + // ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetWithResponse Clusters:Consistency-Groups:Replication:Status // // The group's replication status as one unit: oldest recovery point, worst @@ -16511,6 +16626,54 @@ func (r ClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsis return "" } +type ClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGetResponse struct { + Body []byte + HTTPResponse *http.Response + // JSON200 the response for an HTTP 200 `application/json` response + JSON200 *ConsistencyGroupResolutionDTO + // JSON422 the response for an HTTP 422 `application/json` response + JSON422 *HTTPValidationError +} + +// GetJSON200 returns the response for an HTTP 200 `application/json` response +func (r ClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGetResponse) GetJSON200() *ConsistencyGroupResolutionDTO { + return r.JSON200 +} + +// GetJSON422 returns the response for an HTTP 422 `application/json` response +func (r ClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGetResponse) GetJSON422() *HTTPValidationError { + return r.JSON422 +} + +// GetBody returns the raw response body bytes +func (r ClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGetResponse) GetBody() []byte { + return r.Body +} + +// Status returns HTTPResponse.Status +func (r ClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGetResponse) Status() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Status + } + return http.StatusText(0) +} + +// StatusCode returns HTTPResponse.StatusCode +func (r ClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGetResponse) StatusCode() int { + if r.HTTPResponse != nil { + return r.HTTPResponse.StatusCode + } + return 0 +} + +// ContentType is a convenience method to retrieve the Content-Type value from the HTTP response headers +func (r ClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGetResponse) ContentType() string { + if r.HTTPResponse != nil { + return r.HTTPResponse.Header.Get("Content-Type") + } + return "" +} + type ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetResponse struct { Body []byte HTTPResponse *http.Response @@ -21544,6 +21707,29 @@ func (c *ClientWithResponses) ClustersConsistencyGroupsReplicationFailoverApiV2C return ParseClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationFailoverPostResponse(rsp) } +// ClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGetWithResponse Clusters:Consistency-Groups:Replication:Resolution +// +// Where the group's data lives now, keyed by the handles its PVs keep. +// +// After a relocate the group a VGR names is empty -- its demoted members were +// deleted so the way back stays open -- and its data lives in the peer group of +// the same name. The CSI driver resolves the VGR's original group handle here: +// the group holding live members, and each protected volume's original handle +// with the volume serving it now (2026-10-04: WordPress's VRG waited for +// destination info for ever against the emptied source group). Never a 404 +// for an existing group: “active_group_id“ is empty when nothing serves it. +// +// Returns a wrapper object for the known response body format(s). +// +// Corresponds with GET /api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/resolution (the `ClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGet` operationId). +func (c *ClientWithResponses) ClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGetWithResponse(ctx context.Context, clusterId openapi_types.UUID, groupId openapi_types.UUID, reqEditors ...RequestEditorFn) (*ClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGetResponse, error) { + rsp, err := c.ClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGet(ctx, clusterId, groupId, reqEditors...) + if err != nil { + return nil, err + } + return ParseClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGetResponse(rsp) +} + // ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetWithResponse Clusters:Consistency-Groups:Replication:Status // // The group's replication status as one unit: oldest recovery point, worst @@ -24247,6 +24433,39 @@ func ParseClustersConsistencyGroupsReplicationFailoverApiV2ClustersClusterIdCons return response, nil } +// ParseClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGetResponse parses an HTTP response from a ClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGetWithResponse call +func ParseClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGetResponse(rsp *http.Response) (*ClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGetResponse, error) { + bodyBytes, err := io.ReadAll(rsp.Body) + defer func() { _ = rsp.Body.Close() }() + if err != nil { + return nil, err + } + + response := &ClustersConsistencyGroupsReplicationResolutionApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationResolutionGetResponse{ + Body: bodyBytes, + HTTPResponse: rsp, + } + + switch { + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 200: + var dest ConsistencyGroupResolutionDTO + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON200 = &dest + + case strings.Contains(rsp.Header.Get("Content-Type"), "json") && rsp.StatusCode == 422: + var dest HTTPValidationError + if err := json.Unmarshal(bodyBytes, &dest); err != nil { + return nil, err + } + response.JSON422 = &dest + + } + + return response, nil +} + // ParseClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetResponse parses an HTTP response from a ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetWithResponse call func ParseClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetResponse(rsp *http.Response) (*ClustersConsistencyGroupsReplicationStatusApiV2ClustersClusterIdConsistencyGroupsGroupIdReplicationStatusGetResponse, error) { bodyBytes, err := io.ReadAll(rsp.Body) diff --git a/csi-driver/internal/csi/controller/mock_controlplane_test.go b/csi-driver/internal/csi/controller/mock_controlplane_test.go index 49f75dfb0..99fe69220 100644 --- a/csi-driver/internal/csi/controller/mock_controlplane_test.go +++ b/csi-driver/internal/csi/controller/mock_controlplane_test.go @@ -162,6 +162,12 @@ type mockSBCLI struct { // call landed on, so a test can assert the relationship resolution // redirected it -- the same capture handleFailover keeps for promote. lastFailbackVolumeID string + + // groupResolution, keyed by group id, is the body GET + // .../consistency-groups/{id}/replication/resolution serves. Absent means + // a control plane that predates the endpoint (a route-level 404), so the + // driver keeps its pre-resolution behaviour. + groupResolution map[string]map[string]any } func newMockSBCLI() *mockSBCLI { @@ -171,6 +177,7 @@ func newMockSBCLI() *mockSBCLI { groups: make(map[string]*mockGroup), replicationStatus: make(map[string]map[string]any), replicationRelationship: make(map[string]map[string]any), + groupResolution: make(map[string]map[string]any), } mux := http.NewServeMux() @@ -262,6 +269,10 @@ func newMockSBCLI() *mockSBCLI { "GET /api/v2/clusters/{clusterID}/consistency-groups/{groupID}/replication/status", m.locked(m.handleGroupReplicationStatus), ) + mux.HandleFunc( + "GET /api/v2/clusters/{clusterID}/consistency-groups/{groupID}/replication/resolution", + m.locked(m.handleGroupResolution), + ) mux.HandleFunc( "POST /api/v2/clusters/{clusterID}/consistency-groups/{groupID}/snapshots", m.locked(m.handleTakeGroupSnapshot), @@ -775,6 +786,15 @@ func (m *mockSBCLI) handleGroupReplicationStatus(w http.ResponseWriter, r *http. writeJSON(w, http.StatusOK, out) } +func (m *mockSBCLI) handleGroupResolution(w http.ResponseWriter, r *http.Request) { + body, ok := m.groupResolution[r.PathValue("groupID")] + if !ok { + writeJSON(w, http.StatusNotFound, map[string]string{"detail": "Not Found"}) + return + } + writeJSON(w, http.StatusOK, body) +} + func (m *mockSBCLI) handleGroupMembers(w http.ResponseWriter, r *http.Request) { g := m.groups[r.PathValue("groupID")] if g == nil { diff --git a/csi-driver/internal/csi/controller/replication.go b/csi-driver/internal/csi/controller/replication.go index bdc9045c5..5bc76275f 100644 --- a/csi-driver/internal/csi/controller/replication.go +++ b/csi-driver/internal/csi/controller/replication.go @@ -280,9 +280,11 @@ func (cs *Server) EnableVolumeReplication( // group-replication endpoints (design §14.4); a per-volume handle takes the // §5 path below unchanged. if gh, ok := lvol.ParseGroupHandle(lvol.VolumeHandle(volumeIDFrom(req))); ok { - client, err := clusters.ReplicationClient(ctx, gh.ClusterID) + // The live group: after a relocate the named group is empty and the + // re-protection attaches the group serving the data. + gh, client, err := resolveGroupTarget(ctx, gh, groupActiveEnd) if err != nil { - return nil, status.Error(codes.Unavailable, err.Error()) + return nil, err } if err := client.EnableGroupReplication(ctx, gh, policyID); err != nil { return nil, classifyEnableVolumeReplicationError(err) @@ -331,9 +333,11 @@ func (cs *Server) DisableVolumeReplication( req *replication.DisableVolumeReplicationRequest, ) (*replication.DisableVolumeReplicationResponse, error) { if gh, ok := lvol.ParseGroupHandle(lvol.VolumeHandle(volumeIDFrom(req))); ok { - client, err := clusters.ReplicationClient(ctx, gh.ClusterID) + // The local group: this site detaches what it holds, never the live + // group on the other site. + gh, client, err := resolveGroupTarget(ctx, gh, groupLocalSite) if err != nil { - return nil, status.Error(codes.Unavailable, err.Error()) + return nil, err } if err := client.DisableGroupReplication(ctx, gh); err != nil { return nil, classifyDisableVolumeReplicationError(err) @@ -375,9 +379,9 @@ func (cs *Server) GetVolumeReplicationInfo( req *replication.GetVolumeReplicationInfoRequest, ) (*replication.GetVolumeReplicationInfoResponse, error) { if gh, ok := lvol.ParseGroupHandle(lvol.VolumeHandle(volumeIDFrom(req))); ok { - client, err := clusters.ReplicationClient(ctx, gh.ClusterID) + gh, client, err := resolveGroupTarget(ctx, gh, groupActiveEnd) if err != nil { - return nil, status.Error(codes.Unavailable, err.Error()) + return nil, err } info, err := client.GetGroupReplicationInfo(ctx, gh) if err != nil { @@ -435,6 +439,12 @@ func (cs *Server) PromoteVolume( // §14.4): every member is cloned from the same group generation. The // planned/forced split is the backend group failover's own concern, so // force is not forwarded here. + // + // The named group, unresolved: the control plane's group fail-over resolves + // the peer itself and tells apart "already promoted there" (a no-op), a + // fail-back (clone the peer's members home) and a fail-over from the newest + // replicated generation (sbcli failover_group). Promoting the live group + // instead would turn a fail-back into a no-op on the group being left. if gh, ok := lvol.ParseGroupHandle(lvol.VolumeHandle(volumeIDFrom(req))); ok { client, err := clusters.ReplicationClient(ctx, gh.ClusterID) if err != nil { @@ -496,9 +506,11 @@ func (cs *Server) DemoteVolume( req *replication.DemoteVolumeRequest, ) (*replication.DemoteVolumeResponse, error) { if gh, ok := lvol.ParseGroupHandle(lvol.VolumeHandle(volumeIDFrom(req))); ok { - client, err := clusters.ReplicationClient(ctx, gh.ClusterID) + // The local group: a relocate back demotes the group this site serves, + // which after the first move is the peer of the group the VGR names. + gh, client, err := resolveGroupTarget(ctx, gh, groupLocalSite) if err != nil { - return nil, status.Error(codes.Unavailable, err.Error()) + return nil, err } done, err := client.DemoteGroup(ctx, gh) if err != nil { @@ -552,9 +564,9 @@ func (cs *Server) ResyncVolume( req *replication.ResyncVolumeRequest, ) (*replication.ResyncVolumeResponse, error) { if gh, ok := lvol.ParseGroupHandle(lvol.VolumeHandle(volumeIDFrom(req))); ok { - client, err := clusters.ReplicationClient(ctx, gh.ClusterID) + gh, client, err := resolveGroupTarget(ctx, gh, groupActiveEnd) if err != nil { - return nil, status.Error(codes.Unavailable, err.Error()) + return nil, err } if err := client.ResyncGroup(ctx, gh, req.GetParameters()[sourceClusterIDParam]); err != nil { return nil, classifyResyncVolumeError(err) diff --git a/csi-driver/internal/csi/controller/replication_destination.go b/csi-driver/internal/csi/controller/replication_destination.go index 32c882e05..190297954 100644 --- a/csi-driver/internal/csi/controller/replication_destination.go +++ b/csi-driver/internal/csi/controller/replication_destination.go @@ -2,11 +2,13 @@ package controller import ( "context" + "fmt" "github.com/csi-addons/spec/lib/go/replication" "google.golang.org/grpc/codes" "google.golang.org/grpc/status" + atlascp "github.com/simplyblock/atlas/controlplane" "github.com/simplyblock/atlas/lvol" "github.com/simplyblock/csi-driver/internal/clusters" csicommon "github.com/simplyblock/csi-driver/internal/csi/common" @@ -68,6 +70,33 @@ func (cs *Server) GetReplicationDestinationInfo( }, nil } +// groupPVHandles is every handle a PersistentVolume of the group carries. It +// is resolved by the control plane (ResolveGroup): after a relocate the group a +// VGR names is empty -- its demoted members were deleted so a relocate back +// stays possible -- while its PVs keep their original handles, and listing the +// empty group left Ramen's VRG waiting for destination info for ever +// (2026-10-04, WordPress A -> B). The handles are the lineages' origins, the +// keys csi-addons matches; each maps to itself, the contract above. A control +// plane without the resolution endpoint is asked for the current members as +// before. +func groupPVHandles(ctx context.Context, client *atlascp.Client, gh lvol.GroupHandle) ([]lvol.VolumeHandle, error) { + res, err := client.ResolveGroup(ctx, gh) + if err != nil { + return nil, err + } + if res.Legacy { + return client.ConsistencyGroupMemberHandles(ctx, gh) + } + if len(res.Members) == 0 { + return nil, fmt.Errorf("consistency group %s has no live member", gh.Handle()) + } + handles := make([]lvol.VolumeHandle, 0, len(res.Members)) + for _, m := range res.Members { + handles = append(handles, m.Origin) + } + return handles, nil +} + // groupDestinationInfo is the group branch: the group's own handle, and a // complete source -> destination map over the group's current members (the spec // forbids a partial map). The keys are the members' volume handles exactly as @@ -84,7 +113,7 @@ func groupDestinationInfo( if err != nil { return nil, status.Error(codes.Unavailable, err.Error()) } - members, err := client.ConsistencyGroupMemberHandles(ctx, gh) + members, err := groupPVHandles(ctx, client, gh) if err != nil { return nil, status.Errorf(codes.Unavailable, "members of %s: %v", groupID, err) } diff --git a/csi-driver/internal/csi/controller/replication_groupresolve.go b/csi-driver/internal/csi/controller/replication_groupresolve.go new file mode 100644 index 000000000..48851404c --- /dev/null +++ b/csi-driver/internal/csi/controller/replication_groupresolve.go @@ -0,0 +1,98 @@ +package controller + +import ( + "context" + + "google.golang.org/grpc/codes" + "google.golang.org/grpc/status" + "k8s.io/klog" + + atlascp "github.com/simplyblock/atlas/controlplane" + "github.com/simplyblock/atlas/lvol" + "github.com/simplyblock/csi-driver/internal/clusters" +) + +// groupSide says which group of a moved consistency group a verb acts on, the +// group analogue of chooseReplica's local member and the chain's active end. +type groupSide int + +const ( + // groupActiveEnd is the group serving the data now: Info, Resync and + // Enable (re-protection attaches the live group). + groupActiveEnd groupSide = iota + // groupLocalSite is the group on this driver's own cluster: Demote and + // Disable act on what this site holds, never on the live primary elsewhere + // (the per-volume rule of 2026-10-02, PR #618). + groupLocalSite +) + +// resolveGroupTarget resolves a VolumeGroupReplication's group handle to the +// group a verb must act on, and a client for that group's cluster. +// +// A VGR keeps its original group handle across a relocate, while the group it +// names is emptied by design -- its demoted members are deleted so a relocate +// back stays possible -- and the data lives in the peer group of the same name +// as clones (2026-10-04, WordPress A -> B: every group verb on the original +// handle reached the empty source group). The control plane resolves the +// handle (ResolveGroup); the two candidates are the named group and the group +// holding live members. groupActiveEnd picks the latter; groupLocalSite picks +// the one on a cluster flagged local in the driver's secret, preferring the +// live one, and falls back to the live one when no cluster is flagged, as +// chooseReplica does. +// +// A control plane without the resolution endpoint, or a resolution that fails, +// leaves the handle as it is: the verbs then behave as before this resolution. +func resolveGroupTarget( + ctx context.Context, gh lvol.GroupHandle, side groupSide, +) (lvol.GroupHandle, *atlascp.Client, error) { + client, err := clusters.ReplicationClient(ctx, gh.ClusterID) + if err != nil { + return gh, nil, status.Error(codes.Unavailable, err.Error()) + } + res, err := client.ResolveGroup(ctx, gh) + if err != nil { + klog.Warningf("resolve group %s: %v; acting on the handle as named", gh.Handle(), err) + return gh, client, nil + } + target := chooseGroup(gh, res, side, localClusters()) + if target == gh { + return gh, client, nil + } + targetClient, err := clusters.ReplicationClient(ctx, target.ClusterID) + if err != nil { + return gh, nil, status.Error(codes.Unavailable, err.Error()) + } + klog.Infof("group %s resolves to %s", gh.Handle(), target.Handle()) + return target, targetClient, nil +} + +// chooseGroup is resolveGroupTarget's pure choice: the named group or the live +// one, by side and the local clusters (nil when none is flagged). +func chooseGroup( + gh lvol.GroupHandle, res atlascp.GroupResolution, side groupSide, local map[string]bool, +) lvol.GroupHandle { + if res.Active == nil || *res.Active == gh { + return gh + } + live := *res.Active + if side == groupActiveEnd || local == nil { + return live + } + if local[live.ClusterID] { + return live + } + if local[gh.ClusterID] { + return gh + } + return live +} + +// localClusters is the set of clusters the driver's secret flags local, nil +// when none is flagged (an older operator) or the secret cannot be read. +func localClusters() map[string]bool { + local, flagged, err := clusters.Local() + if err != nil || !flagged { + return nil + } + return local +} diff --git a/csi-driver/internal/csi/controller/replication_groupresolve_test.go b/csi-driver/internal/csi/controller/replication_groupresolve_test.go new file mode 100644 index 000000000..7f8c6e771 --- /dev/null +++ b/csi-driver/internal/csi/controller/replication_groupresolve_test.go @@ -0,0 +1,214 @@ +package controller + +import ( + "context" + "encoding/json" + "os" + "path/filepath" + "testing" + + "github.com/csi-addons/spec/lib/go/replication" + "google.golang.org/grpc/codes" + "google.golang.org/grpc/status" + + atlascp "github.com/simplyblock/atlas/controlplane" + "github.com/simplyblock/atlas/lvol" + "github.com/simplyblock/csi-driver/internal/clusters" +) + +// A VolumeGroupReplication keeps its original group handle across a relocate, +// while the group it names is emptied by design (its demoted members are +// deleted so a relocate back stays possible) and the data lives in the peer +// group of the same name on the other site. The group verbs resolve the handle +// through the control plane (2026-10-04, WordPress A -> B: Ramen's VRG waited +// for destination info for ever against the emptied source group). + +const ( + peerClusterID = "bbbbbbbb-bbbb-4bbb-8bbb-bbbbbbbbbbbb" + peerGroupID = "e5e5e5e5-e5e5-4e5e-8e5e-e5e5e5e5e5e5" + peerClone = "f6666666-6666-4666-8666-666666666666" + origPV = sanityClusterID + ":p-a:" + vgMember1 +) + +// writeTwoSiteSecret registers the named cluster (site A) and its peer (site +// B), both answered by the one mock, with site B flagged local: the driver runs +// on site B, where the relocate landed. +func writeTwoSiteSecret(t *testing.T, mock *mockSBCLI, localCluster string) { + t.Helper() + data, _ := json.Marshal(clusters.Info{Clusters: []clusters.Config{ + {ClusterID: sanityClusterID, ClusterEndpoint: mock.URL(), ClusterSecret: sanitySecret, + Local: localCluster == sanityClusterID}, + {ClusterID: peerClusterID, ClusterEndpoint: mock.URL(), ClusterSecret: sanitySecret, + Local: localCluster == peerClusterID}, + }}) + f := filepath.Join(t.TempDir(), "secret.json") + if err := os.WriteFile(f, data, 0o600); err != nil { + t.Fatal(err) + } + t.Setenv("SPDKCSI_SECRET", f) +} + +// movedGroup is the incident's state: the named group on A is empty, the data +// lives in the peer group on B as a clone of the original volume. +func movedGroup(t *testing.T, localCluster string) (*Server, *mockSBCLI) { + t.Helper() + mock := newMockSBCLI() + t.Cleanup(mock.Close) + mock.seedGroup(vgGroupID) + mock.seedGroup(peerGroupID, peerClone) + mock.groupResolution[vgGroupID] = map[string]any{ + "cluster_id": sanityClusterID, "group_id": vgGroupID, + "active_cluster_id": peerClusterID, "active_group_id": peerGroupID, + "members": []map[string]string{{ + "origin_handle": origPV, "active_handle": peerClusterID + ":p-b:" + peerClone}}, + } + cs := newTestControllerServer(t, mock) + writeTwoSiteSecret(t, mock, localCluster) + return cs, mock +} + +func TestDestinationInfoOfAGroupEmptiedByARelocateMapsItsPVs(t *testing.T) { + cs, _ := movedGroup(t, peerClusterID) + + resp, err := cs.GetReplicationDestinationInfo(context.Background(), + &replication.GetReplicationDestinationInfoRequest{ReplicationSource: groupSource()}) + if err != nil { + t.Fatalf("GetReplicationDestinationInfo: %v", err) + } + vg := resp.GetReplicationDestination().GetVolumegroup() + if vg.GetVolumeGroupId() != vgGroupHandle { + t.Fatalf("destination group = %q, want the VGR's own handle %q", vg.GetVolumeGroupId(), vgGroupHandle) + } + if got := vg.GetVolumeIds(); len(got) != 1 || got[origPV] != origPV { + t.Fatalf("map = %v, want the PV's original handle mapped to itself", got) + } +} + +func TestDestinationInfoOfAGroupWithNoLiveMemberIsUnavailable(t *testing.T) { + cs, mock := movedGroup(t, peerClusterID) + mock.groupResolution[vgGroupID] = map[string]any{"cluster_id": sanityClusterID, "group_id": vgGroupID} + + _, err := cs.GetReplicationDestinationInfo(context.Background(), + &replication.GetReplicationDestinationInfoRequest{ReplicationSource: groupSource()}) + if status.Code(err) != codes.Unavailable { + t.Fatalf("code = %v, want Unavailable (retryable, never a partial map)", status.Code(err)) + } +} + +// The relocate back B -> A: Ramen demotes the VGR on B, which names the group +// on A. Site B must demote the group it serves -- the peer group -- not the +// empty one on A. +func TestDemoteOfAMovedGroupActsOnTheLocalGroup(t *testing.T) { + cs, mock := movedGroup(t, peerClusterID) + + if _, err := cs.DemoteVolume(context.Background(), + &replication.DemoteVolumeRequest{ReplicationSource: groupSource()}); err != nil { + t.Fatalf("DemoteVolume: %v", err) + } + if !mock.groups[peerGroupID].Demoted || mock.groups[vgGroupID].Demoted { + t.Fatalf("demoted: peer %v, named %v; want the peer group on site B only", + mock.groups[peerGroupID].Demoted, mock.groups[vgGroupID].Demoted) + } +} + +// Site A, the old primary, demoting its side must stay on its own group. +func TestDemoteOnTheOldPrimarysSiteStaysOnItsOwnGroup(t *testing.T) { + cs, mock := movedGroup(t, sanityClusterID) + + if _, err := cs.DemoteVolume(context.Background(), + &replication.DemoteVolumeRequest{ReplicationSource: groupSource()}); err != nil { + t.Fatalf("DemoteVolume: %v", err) + } + if mock.groups[peerGroupID].Demoted || !mock.groups[vgGroupID].Demoted { + t.Fatal("site A demoted the live group on site B") + } +} + +// Re-protection B -> A: Enable on the VGR's handle attaches the live group. +func TestEnableOfAMovedGroupAttachesTheLiveGroup(t *testing.T) { + cs, mock := movedGroup(t, peerClusterID) + + _, err := cs.EnableVolumeReplication(context.Background(), &replication.EnableVolumeReplicationRequest{ + ReplicationSource: groupSource(), + Parameters: map[string]string{replicationPolicyParam: vgPolicyID}, + }) + if err != nil { + t.Fatalf("EnableVolumeReplication: %v", err) + } + if mock.groups[peerGroupID].PolicyID != vgPolicyID || mock.groups[vgGroupID].PolicyID != "" { + t.Fatalf("policy: peer %q, named %q; want it on the live group only", + mock.groups[peerGroupID].PolicyID, mock.groups[vgGroupID].PolicyID) + } +} + +func TestInfoOfAMovedGroupReadsTheLiveGroup(t *testing.T) { + cs, mock := movedGroup(t, peerClusterID) + mock.groups[peerGroupID].LastReplicatedAt = 1791105000 + + resp, err := cs.GetVolumeReplicationInfo(context.Background(), + &replication.GetVolumeReplicationInfoRequest{ReplicationSource: groupSource()}) + if err != nil { + t.Fatalf("GetVolumeReplicationInfo: %v", err) + } + if resp.GetLastSyncTime().GetSeconds() != 1791105000 { + t.Fatalf("last sync = %v, want the live group's", resp.GetLastSyncTime()) + } +} + +// Promote stays on the named group: the control plane's group fail-over +// resolves the peer itself, and redirecting it would turn a fail-back into a +// no-op on the group being left. +func TestPromoteOfAMovedGroupStaysOnTheNamedGroup(t *testing.T) { + cs, mock := movedGroup(t, peerClusterID) + + if _, err := cs.PromoteVolume(context.Background(), + &replication.PromoteVolumeRequest{ReplicationSource: groupSource()}); err != nil { + t.Fatalf("PromoteVolume: %v", err) + } + if !mock.groups[vgGroupID].Promoted || mock.groups[peerGroupID].Promoted { + t.Fatal("promote did not reach the named group") + } +} + +// A group still live where it was created resolves to itself. +func TestAGroupLiveAtItsSourceIsActedOnAsNamed(t *testing.T) { + mock := newMockSBCLI() + defer mock.Close() + cs := newGroupReplTestServer(t, mock) + mock.groupResolution[vgGroupID] = map[string]any{ + "cluster_id": sanityClusterID, "group_id": vgGroupID, + "active_cluster_id": sanityClusterID, "active_group_id": vgGroupID, + "members": []map[string]string{{"origin_handle": origPV, "active_handle": origPV}}, + } + if _, err := cs.DemoteVolume(context.Background(), + &replication.DemoteVolumeRequest{ReplicationSource: groupSource()}); err != nil { + t.Fatalf("DemoteVolume: %v", err) + } + if !mock.groups[vgGroupID].Demoted { + t.Fatal("the named, live group was not demoted") + } +} + +func TestChooseGroup(t *testing.T) { + named := lvol.GroupHandle{ClusterID: sanityClusterID, GroupID: vgGroupID} + live := lvol.GroupHandle{ClusterID: peerClusterID, GroupID: peerGroupID} + moved := atlascp.GroupResolution{Active: &live} + for _, tc := range []struct { + name string + res atlascp.GroupResolution + side groupSide + local map[string]bool + want lvol.GroupHandle + }{ + {"not moved", atlascp.GroupResolution{Active: &named}, groupLocalSite, map[string]bool{peerClusterID: true}, named}, + {"nothing live", atlascp.GroupResolution{}, groupActiveEnd, nil, named}, + {"active end", moved, groupActiveEnd, map[string]bool{sanityClusterID: true}, live}, + {"local on the live site", moved, groupLocalSite, map[string]bool{peerClusterID: true}, live}, + {"local on the named site", moved, groupLocalSite, map[string]bool{sanityClusterID: true}, named}, + {"no local flags", moved, groupLocalSite, nil, live}, + } { + if got := chooseGroup(named, tc.res, tc.side, tc.local); got != tc.want { + t.Errorf("%s: got %s, want %s", tc.name, got.Handle(), tc.want.Handle()) + } + } +} diff --git a/csi-driver/internal/csi/controller/volumegroup.go b/csi-driver/internal/csi/controller/volumegroup.go index 9123cea8e..e3ee4b0ed 100644 --- a/csi-driver/internal/csi/controller/volumegroup.go +++ b/csi-driver/internal/csi/controller/volumegroup.go @@ -46,6 +46,14 @@ func (cs *Server) CreateVolumeGroup( } groupID, err := client.ConsistencyGroupForLvols(ctx, clusterID, lvolIDs) if err != nil { + // The PVs may keep the handles of volumes a relocate replaced: their + // data lives at the end of each relationship chain, in the group the + // move formed there. Group that live set instead. + if gh, ok := liveGroupForMembers(ctx, req.GetVolumeIds()); ok { + return &volumegroup.CreateVolumeGroupResponse{ + VolumeGroup: &volumegroup.VolumeGroup{VolumeGroupId: string(gh.Handle())}, + }, nil + } // No backend group matches the selection exactly: the members are not // one whole consistency group (design §14.3, the admission webhook's // invariant), so refuse rather than group a partial set. @@ -94,6 +102,46 @@ func (cs *Server) DeleteVolumeGroup( // parseGroupMembers parses the member volume handles of a CreateVolumeGroup // request into a shared cluster id and the member lvol ids, rejecting a // malformed handle or members that span clusters. +// liveGroupForMembers resolves each PV handle through its relationship chain to +// the volume serving it now and returns the consistency group that holds +// exactly those volumes, when they all live in one cluster and form one group. +func liveGroupForMembers(ctx context.Context, volumeIDs []string) (lvol.GroupHandle, bool) { + cluster := "" + ids := make([]string, 0, len(volumeIDs)) + for _, vid := range volumeIDs { + h, ok := lvol.ParseHandle(lvol.VolumeHandle(vid)) + if !ok { + return lvol.GroupHandle{}, false + } + client, err := clusters.ReplicationClient(ctx, h.ClusterID) + if err != nil { + return lvol.GroupHandle{}, false + } + hops, _, err := resolveChain(ctx, &h, client) + if err != nil || len(hops) == 0 { + return lvol.GroupHandle{}, false + } + end := hops[len(hops)-1].h + if cluster != "" && end.ClusterID != cluster { + return lvol.GroupHandle{}, false + } + cluster = end.ClusterID + ids = append(ids, end.VolumeID) + } + if cluster == "" { + return lvol.GroupHandle{}, false + } + client, err := clusters.ReplicationClient(ctx, cluster) + if err != nil { + return lvol.GroupHandle{}, false + } + groupID, err := client.ConsistencyGroupForLvols(ctx, cluster, ids) + if err != nil { + return lvol.GroupHandle{}, false + } + return lvol.GroupHandle{ClusterID: cluster, GroupID: groupID}, true +} + func parseGroupMembers(volumeIDs []string) (clusterID string, lvolIDs []string, err error) { if len(volumeIDs) == 0 { return "", nil, status.Error(codes.InvalidArgument, diff --git a/shared/openapi.json b/shared/openapi.json index 7a8da607c..d94dbb0d1 100644 --- a/shared/openapi.json +++ b/shared/openapi.json @@ -8133,6 +8133,65 @@ } } } + }, + "/api/v2/clusters/{cluster_id}/consistency-groups/{group_id}/replication/resolution": { + "get": { + "tags": [ + "consistency-groups" + ], + "summary": "Clusters:Consistency-Groups:Replication:Resolution", + "description": "Where the group's data lives now, keyed by the handles its PVs keep.\n\nAfter a relocate the group a VGR names is empty -- its demoted members were\ndeleted so the way back stays open -- and its data lives in the peer group of\nthe same name. The CSI driver resolves the VGR's original group handle here:\nthe group holding live members, and each protected volume's original handle\nwith the volume serving it now (2026-10-04: WordPress's VRG waited for\ndestination info for ever against the emptied source group). Never a 404\nfor an existing group: ``active_group_id`` is empty when nothing serves it.", + "operationId": "clusters_consistency_groups_replication_resolution_api_v2_clusters__cluster_id__consistency_groups__group_id__replication_resolution_get", + "security": [ + { + "HTTPBearer": [] + } + ], + "parameters": [ + { + "name": "cluster_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Cluster Id" + } + }, + { + "name": "group_id", + "in": "path", + "required": true, + "schema": { + "type": "string", + "format": "uuid", + "title": "Group Id" + } + } + ], + "responses": { + "200": { + "description": "Successful Response", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/ConsistencyGroupResolutionDTO" + } + } + } + }, + "422": { + "description": "Validation Error", + "content": { + "application/json": { + "schema": { + "$ref": "#/components/schemas/HTTPValidationError" + } + } + } + } + } + } } }, "components": { @@ -12422,6 +12481,62 @@ "spdk_image" ], "title": "_UpdateParams" + }, + "ConsistencyGroupResolutionDTO": { + "properties": { + "cluster_id": { + "type": "string", + "title": "Cluster Id" + }, + "group_id": { + "type": "string", + "title": "Group Id" + }, + "active_cluster_id": { + "type": "string", + "title": "Active Cluster Id", + "default": "" + }, + "active_group_id": { + "type": "string", + "title": "Active Group Id", + "default": "" + }, + "members": { + "items": { + "$ref": "#/components/schemas/ConsistencyGroupLineageMemberDTO" + }, + "type": "array", + "title": "Members", + "default": [] + } + }, + "type": "object", + "required": [ + "cluster_id", + "group_id" + ], + "title": "ConsistencyGroupResolutionDTO", + "description": "Where a consistency group's data lives now (replication_policy_controller.\nresolve_group). ``active_*`` are empty when no group holds a live member." + }, + "ConsistencyGroupLineageMemberDTO": { + "properties": { + "origin_handle": { + "type": "string", + "title": "Origin Handle" + }, + "active_handle": { + "type": "string", + "title": "Active Handle" + } + }, + "type": "object", + "required": [ + "origin_handle", + "active_handle" + ], + "title": "ConsistencyGroupLineageMemberDTO", + "description": "One protected volume of a consistency group: the handle its\nPersistentVolume carries and the volume serving its data now." } }, "securitySchemes": { From 981ea4b1d51998e3a9f09d42a48c7b18193de5b6 Mon Sep 17 00:00:00 2001 From: michael Date: Sun, 4 Oct 2026 14:21:49 +0300 Subject: [PATCH 193/206] operator: start the hub-only OCM controllers when ManifestWork appears after startup The TestFailover and StorageSiteDeployment controllers were registered only if OCM's ManifestWork resource was served at process start, and never again. The operator is commonly installed before OCM (the DR stack brings it), and on the DR test bed every TestFailover stayed without a reconciler until the operator restarted (2026-10-04). When the resource is absent at start, a runnable now polls discovery (60 s, doubling to 10 min; errors polled through) and registers both controllers with the running manager once it is served. controller-runtime starts runnables added after Start right away, and neither controller registers field indexes, so no process restart is needed. Co-Authored-By: Claude Opus 5.5 --- operator/cmd/main.go | 23 +++++-- operator/cmd/ocm_late.go | 70 ++++++++++++++++++++ operator/cmd/ocm_late_test.go | 117 ++++++++++++++++++++++++++++++++++ 3 files changed, 204 insertions(+), 6 deletions(-) create mode 100644 operator/cmd/ocm_late.go create mode 100644 operator/cmd/ocm_late_test.go diff --git a/operator/cmd/main.go b/operator/cmd/main.go index 004db856e..6796e75bf 100644 --- a/operator/cmd/main.go +++ b/operator/cmd/main.go @@ -953,14 +953,13 @@ func main() { setupLog.Error(err, "unable to determine whether the OCM ManifestWork resource is served") os.Exit(1) } - if hasManifestWork { + registerOCMControllers := func() error { if err := (&controller.TestFailoverReconciler{ Client: mgr.GetClient(), Scheme: mgr.GetScheme(), Recorder: mgr.GetEventRecorder("testfailover-controller"), }).SetupWithManager(mgr); err != nil { - setupLog.Error(err, "unable to create controller", "controller", "TestFailover") - os.Exit(1) + return fmt.Errorf("controller TestFailover: %w", err) } // A managed site's storage deployment is requested from the hub through // the same work API, so the controller is hub-only too. @@ -969,13 +968,25 @@ func main() { Scheme: mgr.GetScheme(), Recorder: mgr.GetEventRecorder("storagesitedeployment-controller"), }).SetupWithManager(mgr); err != nil { - setupLog.Error(err, "unable to create controller", "controller", "StorageSiteDeployment") + return fmt.Errorf("controller StorageSiteDeployment: %w", err) + } + return nil + } + if hasManifestWork { + if err := registerOCMControllers(); err != nil { + setupLog.Error(err, "unable to create controller") os.Exit(1) } } else { - setupLog.Info("OCM ManifestWork resource not served; skipping the TestFailover and "+ - "StorageSiteDeployment controllers (hub-only)", + // OCM may be installed after the operator (the DR stack brings it): keep + // looking, and start the controllers once ManifestWork is served. + setupLog.Info("OCM ManifestWork resource not served yet; the TestFailover and "+ + "StorageSiteDeployment controllers start once it is (hub-only)", "groupVersion", ocmWorkGroupVersion, "resource", ocmManifestWorkResource) + if err := mgr.Add(&ocmLateStart{log: setupLog, disc: workDiscovery, register: registerOCMControllers}); err != nil { + setupLog.Error(err, "unable to add the OCM ManifestWork watcher") + os.Exit(1) + } } // +kubebuilder:scaffold:builder diff --git a/operator/cmd/ocm_late.go b/operator/cmd/ocm_late.go new file mode 100644 index 000000000..1a38c14af --- /dev/null +++ b/operator/cmd/ocm_late.go @@ -0,0 +1,70 @@ +package main + +import ( + "context" + "time" + + "github.com/go-logr/logr" +) + +const ( + // ocmPollInterval is how often the operator looks again for the OCM + // ManifestWork resource when it was not served at start; it doubles after + // each miss up to ocmPollMaxInterval. + ocmPollInterval = 60 * time.Second + ocmPollMaxInterval = 10 * time.Minute +) + +// waitForResource polls discovery until the API server serves resource in +// groupVersion, and reports true; false when ctx ends first. Discovery errors +// are logged and polled through: they are as likely to be a passing API server +// hiccup as a real absence. after is time.After outside of tests. +func waitForResource(ctx context.Context, log logr.Logger, disc serverResourcesGetter, groupVersion, resource string, + interval, maxInterval time.Duration, after func(time.Duration) <-chan time.Time) bool { + for { + select { + case <-ctx.Done(): + return false + case <-after(interval): + } + served, err := serverHasResource(disc, groupVersion, resource) + if err != nil { + log.Info("could not check for the OCM ManifestWork resource; trying again", + "groupVersion", groupVersion, "resource", resource, "error", err.Error()) + } else if served { + return true + } + if interval *= 2; interval > maxInterval { + interval = maxInterval + } + } +} + +// ocmLateStart registers the hub-only controllers once OCM's ManifestWork is +// served. The operator is commonly installed before OCM (the DR stack brings +// it); checking only at process start left every TestFailover without a +// reconciler until the operator happened to restart (2026-10-04). Controllers +// added to a running manager are started by controller-runtime right away +// (manager runnableGroup.Add), and neither controller registers field indexes, +// so no restart of the process is needed. +type ocmLateStart struct { + log logr.Logger + disc serverResourcesGetter + register func() error + after func(time.Duration) <-chan time.Time +} + +// Start implements manager.Runnable. +func (o *ocmLateStart) Start(ctx context.Context) error { + after := o.after + if after == nil { + after = time.After + } + if !waitForResource(ctx, o.log, o.disc, ocmWorkGroupVersion, ocmManifestWorkResource, + ocmPollInterval, ocmPollMaxInterval, after) { + return nil + } + o.log.Info("OCM ManifestWork resource is now served; starting the TestFailover and "+ + "StorageSiteDeployment controllers", "groupVersion", ocmWorkGroupVersion, "resource", ocmManifestWorkResource) + return o.register() +} diff --git a/operator/cmd/ocm_late_test.go b/operator/cmd/ocm_late_test.go new file mode 100644 index 000000000..9f2494189 --- /dev/null +++ b/operator/cmd/ocm_late_test.go @@ -0,0 +1,117 @@ +package main + +import ( + "context" + "errors" + "testing" + "time" + + "github.com/go-logr/logr" + apierrors "k8s.io/apimachinery/pkg/api/errors" + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/runtime/schema" +) + +// scriptedDiscovery answers ServerResourcesForGroupVersion from a script, one +// answer per call; the last answer repeats. +type scriptedDiscovery struct { + answers []func() (*metav1.APIResourceList, error) + calls int +} + +func (d *scriptedDiscovery) ServerResourcesForGroupVersion(string) (*metav1.APIResourceList, error) { + i := d.calls + if i >= len(d.answers) { + i = len(d.answers) - 1 + } + d.calls++ + return d.answers[i]() +} + +func notServed() (*metav1.APIResourceList, error) { + return nil, apierrors.NewNotFound(schema.GroupResource{Group: "work.open-cluster-management.io"}, "v1") +} + +func servedWithout() (*metav1.APIResourceList, error) { + // a managed cluster: the group serves AppliedManifestWork, not ManifestWork + return &metav1.APIResourceList{APIResources: []metav1.APIResource{{Name: "appliedmanifestworks"}}}, nil +} + +func served() (*metav1.APIResourceList, error) { + return &metav1.APIResourceList{APIResources: []metav1.APIResource{{Name: ocmManifestWorkResource}}}, nil +} + +func discoveryError() (*metav1.APIResourceList, error) { + return nil, errors.New("the server is currently unable to handle the request") +} + +// instant fires every wait at once and records the intervals asked for. +func instant(waits *[]time.Duration) func(time.Duration) <-chan time.Time { + return func(d time.Duration) <-chan time.Time { + *waits = append(*waits, d) + ch := make(chan time.Time, 1) + ch <- time.Time{} + return ch + } +} + +func TestWaitForResourceFindsManifestWorkInstalledLater(t *testing.T) { + disc := &scriptedDiscovery{answers: []func() (*metav1.APIResourceList, error){ + notServed, servedWithout, discoveryError, served}} + var waits []time.Duration + ok := waitForResource(context.Background(), logr.Discard(), disc, ocmWorkGroupVersion, ocmManifestWorkResource, + time.Minute, 3*time.Minute, instant(&waits)) + if !ok { + t.Fatal("waitForResource gave up although ManifestWork became served") + } + if disc.calls != 4 { + t.Errorf("discovery called %d times, want 4 (absent, other kind only, error, served)", disc.calls) + } + want := []time.Duration{time.Minute, 2 * time.Minute, 3 * time.Minute, 3 * time.Minute} + if len(waits) != len(want) { + t.Fatalf("waits %v, want %v", waits, want) + } + for i := range want { + if waits[i] != want[i] { + t.Errorf("wait %d = %v, want %v (doubling, capped)", i, waits[i], want[i]) + } + } +} + +func TestWaitForResourceStopsWithItsContext(t *testing.T) { + disc := &scriptedDiscovery{answers: []func() (*metav1.APIResourceList, error){notServed}} + ctx, cancel := context.WithCancel(context.Background()) + cancel() + never := func(time.Duration) <-chan time.Time { return make(chan time.Time) } + if waitForResource(ctx, logr.Discard(), disc, ocmWorkGroupVersion, ocmManifestWorkResource, + time.Minute, time.Minute, never) { + t.Fatal("waitForResource reported the resource served after its context ended") + } + if disc.calls != 0 { + t.Errorf("discovery called %d times after cancellation, want 0", disc.calls) + } +} + +func TestOCMLateStartRegistersTheControllersOnce(t *testing.T) { + disc := &scriptedDiscovery{answers: []func() (*metav1.APIResourceList, error){notServed, served}} + var waits []time.Duration + registered := 0 + o := &ocmLateStart{log: logr.Discard(), disc: disc, after: instant(&waits), + register: func() error { registered++; return nil }} + if err := o.Start(context.Background()); err != nil { + t.Fatalf("Start: %v", err) + } + if registered != 1 { + t.Errorf("controllers registered %d times, want 1", registered) + } +} + +func TestOCMLateStartReportsARegistrationFailure(t *testing.T) { + disc := &scriptedDiscovery{answers: []func() (*metav1.APIResourceList, error){served}} + var waits []time.Duration + o := &ocmLateStart{log: logr.Discard(), disc: disc, after: instant(&waits), + register: func() error { return errors.New("boom") }} + if err := o.Start(context.Background()); err == nil { + t.Fatal("Start hid the registration failure") + } +} From 582cdde245afe5442ac9329161fad1582952b831 Mon Sep 17 00:00:00 2001 From: michael Date: Sun, 4 Oct 2026 14:30:09 +0300 Subject: [PATCH 194/206] operator: a test fail-over places its PV and PVC CreateOnly; the bubble namespace has one work The work agent re-applied the whole bubble PV under the default update strategy and wiped spec.claimRef.uid, which the PV controller sets on binding: the PV fell back to Available and the PVC became Lost ("ClaimMisbound: Two claims are bound to the same volume", DR test bed 2026-10-04, Gitea bubble). Every object a drill places is CreateOnly now; ServerSideApply would not help, claimRef is an atomic struct. Each per-PVC drill of one test also carried the same bubble Namespace in its own work, with its own test-id label. The namespace now lives in one shared work per bubble (tfo-ns-), created by whichever drill places first and removed with the last drill torn down. Co-Authored-By: Claude Opus 5.5 --- .../controller/testfailover_controller.go | 135 ++++++++++++- .../testfailover_controller_unit_test.go | 189 ++++++++++++++++-- 2 files changed, 297 insertions(+), 27 deletions(-) diff --git a/operator/internal/controller/testfailover_controller.go b/operator/internal/controller/testfailover_controller.go index 3f13340be..475996ffe 100644 --- a/operator/internal/controller/testfailover_controller.go +++ b/operator/internal/controller/testfailover_controller.go @@ -557,6 +557,10 @@ func (r *TestFailoverReconciler) reconcileDeletion(ctx context.Context, tf *simp if err := r.deleteManifestWork(ctx, tf); err != nil { return r.reclaimPending(ctx, tf, "remove the bubble placement", err) } + // The namespace goes with the last drill placing into it. + if err := r.deleteBubbleNamespaceWork(ctx, tf); err != nil { + return r.reclaimPending(ctx, tf, "remove the bubble namespace", err) + } // Reclaim every clone slot (one for a volume drill, one per member for a // group drill). Each reclaim tolerates a not-found, so a re-run after a @@ -946,6 +950,9 @@ func (r *TestFailoverReconciler) placeBubble(ctx context.Context, tf *simplybloc return r.fail(ctx, tf, "internal: the clone was not built before Placing") } + if err := r.ensureBubbleNamespaceWork(ctx, tf); err != nil { + return ctrl.Result{}, err + } var mw workv1.ManifestWork err := r.Get(ctx, client.ObjectKey{Namespace: tf.Spec.BubbleCluster, Name: testFailoverManifestWorkName(tf)}, &mw) if apierrors.IsNotFound(err) { @@ -1059,17 +1066,10 @@ func (r *TestFailoverReconciler) bubbleManifestWork(tf *simplyblockv1alpha2.Test scName := "" group := tf.Spec.Scope == simplyblockv1alpha2.TestFailoverScopeGroup - namespace := &corev1.Namespace{ - TypeMeta: metav1.TypeMeta{Kind: "Namespace", APIVersion: "v1"}, - ObjectMeta: metav1.ObjectMeta{Name: ns, Labels: labels}, - } - // One namespace manifest plus a PV+PVC pair per clone slot. - manifests := make([]workv1.Manifest, 0, 1+2*len(tf.Status.Clones)) - raw, err := json.Marshal(namespace) - if err != nil { - return nil, fmt.Errorf("marshal bubble namespace: %w", err) - } - manifests = append(manifests, workv1.Manifest{RawExtension: runtime.RawExtension{Raw: raw}}) + // A PV+PVC pair per clone slot. The bubble namespace is not in this work: + // every drill of one test shares it, so it has one owner, the bubble's + // namespace work (bubbleNamespaceWork). + manifests := make([]workv1.Manifest, 0, 2*len(tf.Status.Clones)) // One PV+PVC pair per clone slot, each reporting its own bind phase back to the // hub. A volume drill has one; a group drill has one per member, all in the one @@ -1133,6 +1133,10 @@ func (r *TestFailoverReconciler) bubbleManifestWork(tf *simplyblockv1alpha2.Test Type: workv1.JSONPathsType, JsonPaths: []workv1.JsonPath{{Name: "phase", Path: ".status.phase"}}, }}, + UpdateStrategy: createOnly(), + }, workv1.ManifestConfigOption{ + ResourceIdentifier: workv1.ResourceIdentifier{Group: "", Resource: "persistentvolumes", Name: pvName}, + UpdateStrategy: createOnly(), }) } @@ -1181,6 +1185,115 @@ func testFailoverManifestWorkName(tf *simplyblockv1alpha2.TestFailover) string { return fmt.Sprintf("tfo-%x-bubble", h[:6]) } +// createOnly is the update strategy of every object a drill places: created once, +// never re-applied. The default (Update) re-applied the whole PersistentVolume and +// wiped spec.claimRef.uid, which the PV controller sets on binding; the PV fell +// back to Available and the PVC became Lost ("ClaimMisbound: Two claims are bound +// to the same volume", 2026-10-04). ServerSideApply does not help: claimRef is an +// atomic struct, so the work agent would still own all of it. Nothing a drill +// places changes after creation. +func createOnly() *workv1.UpdateStrategy { + return &workv1.UpdateStrategy{Type: workv1.UpdateStrategyTypeCreateOnly} +} + +// bubbleNamespaceWorkName is the name of the one ManifestWork that owns a bubble +// namespace on its recovery cluster, shared by every drill placing into it. +func bubbleNamespaceWorkName(cluster, namespace string) string { + h := sha256.Sum256([]byte(cluster + "/" + namespace)) + return fmt.Sprintf("tfo-ns-%x", h[:6]) +} + +// bubbleNamespaceWork is the ManifestWork holding only the bubble namespace. One +// test fail-over creates a drill per protected PVC, all in the same bubble +// namespace; with the namespace in each drill's own work, the works fought over +// its labels and the first drill torn down deleted the namespace under the +// others' PVCs. +func bubbleNamespaceWork(tf *simplyblockv1alpha2.TestFailover) (*workv1.ManifestWork, error) { + labels := map[string]string{testFailoverIDLabel: string(tf.UID)} + namespace := &corev1.Namespace{ + TypeMeta: metav1.TypeMeta{Kind: "Namespace", APIVersion: "v1"}, + ObjectMeta: metav1.ObjectMeta{Name: tf.Spec.BubbleNamespace, Labels: labels}, + } + raw, err := json.Marshal(namespace) + if err != nil { + return nil, fmt.Errorf("marshal bubble namespace: %w", err) + } + return &workv1.ManifestWork{ + ObjectMeta: metav1.ObjectMeta{ + Name: bubbleNamespaceWorkName(tf.Spec.BubbleCluster, tf.Spec.BubbleNamespace), + Namespace: tf.Spec.BubbleCluster, + Labels: labels, + }, + Spec: workv1.ManifestWorkSpec{ + Workload: workv1.ManifestsTemplate{ + Manifests: []workv1.Manifest{{RawExtension: runtime.RawExtension{Raw: raw}}}, + }, + ManifestConfigs: []workv1.ManifestConfigOption{{ + ResourceIdentifier: workv1.ResourceIdentifier{ + Group: "", Resource: "namespaces", Name: tf.Spec.BubbleNamespace, + }, + UpdateStrategy: createOnly(), + }}, + }, + }, nil +} + +// ensureBubbleNamespaceWork creates the bubble's namespace work unless another +// drill of the same bubble already did. +func (r *TestFailoverReconciler) ensureBubbleNamespaceWork( + ctx context.Context, tf *simplyblockv1alpha2.TestFailover, +) error { + desired, err := bubbleNamespaceWork(tf) + if err != nil { + return err + } + if err := r.Create(ctx, desired); err != nil && !apierrors.IsAlreadyExists(err) { + return err + } + return nil +} + +// liveDrillOnBubble reports whether a drill other than tf, not itself being +// deleted, places into the same bubble namespace on the same cluster. +func liveDrillOnBubble(items []simplyblockv1alpha2.TestFailover, tf *simplyblockv1alpha2.TestFailover) bool { + for i := range items { + o := &items[i] + if o.UID == tf.UID || !o.DeletionTimestamp.IsZero() { + continue + } + if o.Spec.BubbleCluster == tf.Spec.BubbleCluster && o.Spec.BubbleNamespace == tf.Spec.BubbleNamespace { + return true + } + } + return false +} + +// deleteBubbleNamespaceWork removes the bubble's namespace work once no other live +// drill places into it; the last drill torn down takes the namespace with it. +func (r *TestFailoverReconciler) deleteBubbleNamespaceWork( + ctx context.Context, tf *simplyblockv1alpha2.TestFailover, +) error { + var list simplyblockv1alpha2.TestFailoverList + if err := r.List(ctx, &list); err != nil { + return err + } + if liveDrillOnBubble(list.Items, tf) { + return nil + } + var mw workv1.ManifestWork + key := client.ObjectKey{ + Namespace: tf.Spec.BubbleCluster, + Name: bubbleNamespaceWorkName(tf.Spec.BubbleCluster, tf.Spec.BubbleNamespace), + } + if err := r.Get(ctx, key, &mw); err != nil { + return client.IgnoreNotFound(err) + } + if !mw.DeletionTimestamp.IsZero() { + return nil + } + return client.IgnoreNotFound(r.Delete(ctx, &mw)) +} + // clusterSecret returns the cluster secret the control plane authenticates a // per-cluster call with, from the hub Secret simplyblock-cluster-. func (r *TestFailoverReconciler) clusterSecret(ctx context.Context, namespace, clusterName string) (string, error) { diff --git a/operator/internal/controller/testfailover_controller_unit_test.go b/operator/internal/controller/testfailover_controller_unit_test.go index 7c8effd6d..5789f226c 100644 --- a/operator/internal/controller/testfailover_controller_unit_test.go +++ b/operator/internal/controller/testfailover_controller_unit_test.go @@ -782,7 +782,7 @@ func markManifestWorkPVCBound(t *testing.T, cl client.Client, mw *workv1.Manifes } // TestFailoverPlacingDeliversManifestWorkThenReady covers the placement step: a -// ManifestWork carrying the bubble namespace, PV, and PVC is delivered to the +// ManifestWork carrying the bubble PV and PVC (the namespace has its own work) is delivered to the // recovery cluster, and the drill reaches Ready once the PVC binds. func TestFailoverPlacingDeliversManifestWorkThenReady(t *testing.T) { tf := atPlacing() @@ -799,10 +799,10 @@ func TestFailoverPlacingDeliversManifestWorkThenReady(t *testing.T) { t.Errorf("expected a requeue while waiting for the bubble PVC to bind") } mw := getManifestWork(t, cl, tf.Spec.BubbleCluster, testFailoverManifestWorkName(tf)) - if len(mw.Spec.Workload.Manifests) != 3 { - t.Errorf("ManifestWork carries %d manifests, want 3 (namespace, PV, PVC)", len(mw.Spec.Workload.Manifests)) + if len(mw.Spec.Workload.Manifests) != 2 { + t.Errorf("ManifestWork carries %d manifests, want 2 (PV, PVC)", len(mw.Spec.Workload.Manifests)) } - if len(mw.Spec.ManifestConfigs) != 1 || mw.Spec.ManifestConfigs[0].ResourceIdentifier.Name != tf.Spec.SourceRef { + if len(mw.Spec.ManifestConfigs) != 2 || mw.Spec.ManifestConfigs[0].ResourceIdentifier.Name != tf.Spec.SourceRef { t.Errorf("feedback rule not set on the bubble PVC %q: %+v", tf.Spec.SourceRef, mw.Spec.ManifestConfigs) } var got simplyblockv1alpha2.TestFailover @@ -848,12 +848,12 @@ func TestFailoverPlacingPVCarriesSourceVolumeContext(t *testing.T) { t.Fatalf("reconcile: %v", err) } mw := getManifestWork(t, cl, tf.Spec.BubbleCluster, testFailoverManifestWorkName(tf)) - // manifests are [namespace, PV, PVC]; decode the PV. - if len(mw.Spec.Workload.Manifests) != 3 { - t.Fatalf("ManifestWork carries %d manifests, want 3", len(mw.Spec.Workload.Manifests)) + // manifests are [PV, PVC]; decode the PV. + if len(mw.Spec.Workload.Manifests) != 2 { + t.Fatalf("ManifestWork carries %d manifests, want 2", len(mw.Spec.Workload.Manifests)) } var pv corev1.PersistentVolume - if err := json.Unmarshal(mw.Spec.Workload.Manifests[1].Raw, &pv); err != nil { + if err := json.Unmarshal(mw.Spec.Workload.Manifests[0].Raw, &pv); err != nil { t.Fatalf("decode bubble PV manifest: %v", err) } if pv.Spec.CSI == nil { @@ -890,17 +890,17 @@ func TestFailoverPlacingGroupDeliversAllMembersThenReady(t *testing.T) { ctx := context.Background() key := testFailoverRequest(tf).NamespacedName - // Pass 1: the ManifestWork is created carrying ns + 2*(PV,PVC) = 5 manifests - // and one feedback config per member PVC; the drill holds until both bind. + // Pass 1: the ManifestWork is created carrying 2*(PV,PVC) = 4 manifests and + // two configs per member (PVC feedback, PV strategy); the drill holds until both bind. if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { t.Fatalf("pass 1: %v", err) } mw := getManifestWork(t, cl, tf.Spec.BubbleCluster, testFailoverManifestWorkName(tf)) - if len(mw.Spec.Workload.Manifests) != 5 { - t.Errorf("ManifestWork carries %d manifests, want 5 (namespace + 2*(PV,PVC))", len(mw.Spec.Workload.Manifests)) + if len(mw.Spec.Workload.Manifests) != 4 { + t.Errorf("ManifestWork carries %d manifests, want 4 (2*(PV,PVC))", len(mw.Spec.Workload.Manifests)) } - if len(mw.Spec.ManifestConfigs) != 2 { - t.Errorf("ManifestWork has %d feedback configs, want one per member (2)", len(mw.Spec.ManifestConfigs)) + if len(mw.Spec.ManifestConfigs) != 4 { + t.Errorf("ManifestWork has %d configs, want two per member (4)", len(mw.Spec.ManifestConfigs)) } var got simplyblockv1alpha2.TestFailover if err := cl.Get(ctx, key, &got); err != nil { @@ -965,8 +965,9 @@ func TestFailoverPlacingReuseManifestWorkOnRestart(t *testing.T) { if err := cl.List(ctx, &list, client.InNamespace(tf.Spec.BubbleCluster)); err != nil { t.Fatalf("list ManifestWorks: %v", err) } - if len(list.Items) != 1 { - t.Errorf("got %d ManifestWorks, want exactly 1 (no duplicate on restart)", len(list.Items)) + // The drill's own work plus the bubble's namespace work; neither duplicated. + if len(list.Items) != 2 { + t.Errorf("got %d ManifestWorks, want exactly 2 (drill + namespace, no duplicate on restart)", len(list.Items)) } } @@ -1193,3 +1194,159 @@ func TestFailoverStepDeadlineFailsTheDrill(t *testing.T) { t.Errorf("message = %q, want it to mention the deadline", got.Status.Message) } } + +// Regression: 2026-10-04 bubble PVC Lost. The work agent re-applied the whole PV +// under the default update strategy and wiped spec.claimRef.uid; every object a +// drill places is CreateOnly now, and the namespace is in its own work. +func TestFailoverPlacingPlacesEverythingCreateOnly(t *testing.T) { + tf := atPlacing() + tf.Spec.BubbleNamespace = "app-drtest-1" + r, cl := newTestFailoverReconciler(t, tf) + if _, err := r.Reconcile(context.Background(), testFailoverRequest(tf)); err != nil { + t.Fatalf("reconcile: %v", err) + } + mw := getManifestWork(t, cl, tf.Spec.BubbleCluster, testFailoverManifestWorkName(tf)) + seen := map[string]bool{} + for _, c := range mw.Spec.ManifestConfigs { + if c.UpdateStrategy == nil || c.UpdateStrategy.Type != workv1.UpdateStrategyTypeCreateOnly { + t.Errorf("%s/%s: update strategy %+v, want CreateOnly", + c.ResourceIdentifier.Resource, c.ResourceIdentifier.Name, c.UpdateStrategy) + } + seen[c.ResourceIdentifier.Resource] = true + } + if !seen["persistentvolumes"] || !seen["persistentvolumeclaims"] { + t.Errorf("configs cover %v, want both the PV and the PVC", seen) + } + for _, m := range mw.Spec.Workload.Manifests { + var obj metav1.TypeMeta + if err := json.Unmarshal(m.Raw, &obj); err != nil { + t.Fatal(err) + } + if obj.Kind == "Namespace" { + t.Errorf("the drill's work carries the bubble namespace; it belongs to the namespace work") + } + } + nsw := getManifestWork(t, cl, tf.Spec.BubbleCluster, + bubbleNamespaceWorkName(tf.Spec.BubbleCluster, tf.Spec.BubbleNamespace)) + if len(nsw.Spec.Workload.Manifests) != 1 || len(nsw.Spec.ManifestConfigs) != 1 || + nsw.Spec.ManifestConfigs[0].UpdateStrategy.Type != workv1.UpdateStrategyTypeCreateOnly { + t.Errorf("namespace work = %+v, want the one namespace, CreateOnly", nsw.Spec) + } +} + +// Two drills of one test share the bubble namespace: the second finds the +// namespace work and does not fail; one work owns the namespace. +func TestFailoverPlacingSharesOneNamespaceWork(t *testing.T) { + a := atPlacing() + a.Spec.BubbleNamespace = "app-drtest-1" + b := atPlacing() + b.Name, b.UID = "drill-2", "uid-2" + b.Spec.BubbleNamespace = "app-drtest-1" + b.Spec.SourceRef = "other-data" + r, cl := newTestFailoverReconciler(t, a, b) + for _, tf := range []*simplyblockv1alpha2.TestFailover{a, b} { + if _, err := r.Reconcile(context.Background(), testFailoverRequest(tf)); err != nil { + t.Fatalf("reconcile %s: %v", tf.Name, err) + } + } + var works workv1.ManifestWorkList + if err := cl.List(context.Background(), &works, client.InNamespace(a.Spec.BubbleCluster)); err != nil { + t.Fatal(err) + } + owners := 0 + for i := range works.Items { + for _, m := range works.Items[i].Spec.Workload.Manifests { + var obj metav1.TypeMeta + if err := json.Unmarshal(m.Raw, &obj); err != nil { + t.Fatal(err) + } + if obj.Kind == "Namespace" { + owners++ + } + } + } + if owners != 1 { + t.Errorf("%d works carry the bubble namespace, want exactly one", owners) + } +} + +func TestLiveDrillOnBubble(t *testing.T) { + self := sampleTestFailover() + self.UID = "self" + self.Spec.BubbleNamespace = "ns-1" + other := func(uid, cluster, ns string, deleting bool) simplyblockv1alpha2.TestFailover { + o := *sampleTestFailover() + o.UID = types.UID(uid) + o.Spec.BubbleCluster, o.Spec.BubbleNamespace = cluster, ns + if deleting { + now := metav1.Now() + o.DeletionTimestamp = &now + } + return o + } + cases := []struct { + name string + items []simplyblockv1alpha2.TestFailover + want bool + }{ + {"only itself", []simplyblockv1alpha2.TestFailover{*self}, false}, + {"another live drill on the bubble", []simplyblockv1alpha2.TestFailover{*self, + other("o1", self.Spec.BubbleCluster, "ns-1", false)}, true}, + {"the other is being deleted", []simplyblockv1alpha2.TestFailover{*self, + other("o1", self.Spec.BubbleCluster, "ns-1", true)}, false}, + {"another bubble namespace", []simplyblockv1alpha2.TestFailover{*self, + other("o1", self.Spec.BubbleCluster, "ns-2", false)}, false}, + {"another cluster", []simplyblockv1alpha2.TestFailover{*self, + other("o1", "elsewhere", "ns-1", false)}, false}, + } + for _, c := range cases { + if got := liveDrillOnBubble(c.items, self); got != c.want { + t.Errorf("%s: got %v, want %v", c.name, got, c.want) + } + } +} + +// Teardown keeps the namespace work while another live drill places into it, and +// removes it with the last one. +func TestFailoverDeletionRemovesTheNamespaceWorkWithTheLastDrill(t *testing.T) { + mk := func(name, uid string, deleting bool) *simplyblockv1alpha2.TestFailover { + tf := sampleTestFailover() + tf.Name, tf.UID = name, types.UID(uid) + tf.Spec.BubbleNamespace = "app-drtest-1" + tf.Finalizers = []string{finalizerTestFailover} + if deleting { + now := metav1.Now() + tf.DeletionTimestamp = &now + } + return tf + } + first, second := mk("drill-1", "u1", true), mk("drill-2", "u2", false) + nsw, err := bubbleNamespaceWork(first) + if err != nil { + t.Fatal(err) + } + r, cl := newTestFailoverReconciler(t, first, second, nsw) + ctx := context.Background() + key := client.ObjectKey{Namespace: nsw.Namespace, Name: nsw.Name} + + if _, err := r.Reconcile(ctx, testFailoverRequest(first)); err != nil { + t.Fatalf("teardown of drill-1: %v", err) + } + if err := cl.Get(ctx, key, &workv1.ManifestWork{}); err != nil { + t.Fatalf("namespace work removed while drill-2 still places into it: %v", err) + } + + var live simplyblockv1alpha2.TestFailover + if err := cl.Get(ctx, testFailoverRequest(second).NamespacedName, &live); err != nil { + t.Fatal(err) + } + if err := cl.Delete(ctx, &live); err != nil { + t.Fatal(err) + } + if _, err := r.Reconcile(ctx, testFailoverRequest(second)); err != nil { + t.Fatalf("teardown of drill-2: %v", err) + } + if err := cl.Get(ctx, key, &workv1.ManifestWork{}); !apierrors.IsNotFound(err) { + t.Errorf("namespace work still present after the last drill: %v", err) + } +} From 7e0f931b7d8ce89a90b98681b082f4f6197a8658 Mon Sep 17 00:00:00 2001 From: michael Date: Sun, 4 Oct 2026 16:49:05 +0300 Subject: [PATCH 195/206] operator: tests name the bubble namespace once (goconst) Co-Authored-By: Claude Opus 5.5 --- .../controller/testfailover_controller_unit_test.go | 11 +++++++---- 1 file changed, 7 insertions(+), 4 deletions(-) diff --git a/operator/internal/controller/testfailover_controller_unit_test.go b/operator/internal/controller/testfailover_controller_unit_test.go index 5789f226c..c47f8e1f3 100644 --- a/operator/internal/controller/testfailover_controller_unit_test.go +++ b/operator/internal/controller/testfailover_controller_unit_test.go @@ -1200,7 +1200,7 @@ func TestFailoverStepDeadlineFailsTheDrill(t *testing.T) { // drill places is CreateOnly now, and the namespace is in its own work. func TestFailoverPlacingPlacesEverythingCreateOnly(t *testing.T) { tf := atPlacing() - tf.Spec.BubbleNamespace = "app-drtest-1" + tf.Spec.BubbleNamespace = bubbleNS r, cl := newTestFailoverReconciler(t, tf) if _, err := r.Reconcile(context.Background(), testFailoverRequest(tf)); err != nil { t.Fatalf("reconcile: %v", err) @@ -1238,10 +1238,10 @@ func TestFailoverPlacingPlacesEverythingCreateOnly(t *testing.T) { // namespace work and does not fail; one work owns the namespace. func TestFailoverPlacingSharesOneNamespaceWork(t *testing.T) { a := atPlacing() - a.Spec.BubbleNamespace = "app-drtest-1" + a.Spec.BubbleNamespace = bubbleNS b := atPlacing() b.Name, b.UID = "drill-2", "uid-2" - b.Spec.BubbleNamespace = "app-drtest-1" + b.Spec.BubbleNamespace = bubbleNS b.Spec.SourceRef = "other-data" r, cl := newTestFailoverReconciler(t, a, b) for _, tf := range []*simplyblockv1alpha2.TestFailover{a, b} { @@ -1312,7 +1312,7 @@ func TestFailoverDeletionRemovesTheNamespaceWorkWithTheLastDrill(t *testing.T) { mk := func(name, uid string, deleting bool) *simplyblockv1alpha2.TestFailover { tf := sampleTestFailover() tf.Name, tf.UID = name, types.UID(uid) - tf.Spec.BubbleNamespace = "app-drtest-1" + tf.Spec.BubbleNamespace = bubbleNS tf.Finalizers = []string{finalizerTestFailover} if deleting { now := metav1.Now() @@ -1350,3 +1350,6 @@ func TestFailoverDeletionRemovesTheNamespaceWorkWithTheLastDrill(t *testing.T) { t.Errorf("namespace work still present after the last drill: %v", err) } } + +// bubbleNS is the bubble namespace the binding tests place their clones in. +const bubbleNS = "app-drtest-1" From 22efd7e58cc2ff58bb8b8d287a8e45ccfbcaf576 Mon Sep 17 00:00:00 2001 From: michael Date: Sun, 4 Oct 2026 17:56:06 +0300 Subject: [PATCH 196/206] operator: build the manager from its package, not from main.go alone `go build cmd/main.go` compiles that one file; a second file of package main (cmd/ocm_late.go) was left out (undefined: ocmLateStart) in the image build and make build/run. Co-Authored-By: Claude Opus 5.5 --- operator/Dockerfile | 2 +- operator/Makefile | 4 ++-- 2 files changed, 3 insertions(+), 3 deletions(-) diff --git a/operator/Dockerfile b/operator/Dockerfile index 332bdd507..c51792d67 100644 --- a/operator/Dockerfile +++ b/operator/Dockerfile @@ -36,7 +36,7 @@ COPY operator/ . RUN CC=$([ "${TARGETARCH}" = "arm64" ] && echo "aarch64-linux-gnu-gcc" || echo "gcc") \ CGO_ENABLED=1 GOEXPERIMENT=boringcrypto \ GOOS=${TARGETOS:-linux} GOARCH=${TARGETARCH} \ - go build -a -o manager cmd/main.go + go build -a -o manager ./cmd # The worker probe ships in this image rather than in one of its own, so that # the Job a discovery run creates cannot be a version out of step with the diff --git a/operator/Makefile b/operator/Makefile index b03a001a7..5acd38bf0 100644 --- a/operator/Makefile +++ b/operator/Makefile @@ -177,11 +177,11 @@ lint-config: golangci-lint ## Verify golangci-lint linter configuration .PHONY: build build: manifests generate fmt vet ## Build manager binary. mkdir -p build - go build -o build/manager cmd/main.go + go build -o build/manager ./cmd .PHONY: run run: manifests generate fmt vet ## Run a controller from your host. - go run ./cmd/main.go + go run ./cmd # The API upgrade tool, built on its own and not in the operator image: it is a # prerequisite for running the new operator, so shipping it inside that From cbeaf344e1417a4bf5bbe1913767264e9273eccc Mon Sep 17 00:00:00 2001 From: michael Date: Sun, 4 Oct 2026 22:16:05 +0300 Subject: [PATCH 197/206] fix(testfailover): pair group generation snapshots with PVCs by volume, not position resolvePointGroup handed the group generation's snapshots to the clone slots in list order ("a precise source-to-target mapping is a later refinement"). The control plane lists the members in its own order, so a group drill could clone one member's snapshot onto another member's PVC (the database disk onto the web volume): one generation, but the wrong disks. Each slot now gets the snapshot of its own member: - by source_lvol_id when every generation member reports it; - otherwise by the replica volume on the target (the member's lvol_id), read per slot from the source volume's latest-snapshot (its lvol_id is the same replica volume). A slot without a member, a member no slot claims, a member naming no volume or two members of one volume fail the drill with the PVC and volume named. There is no positional fallback. Co-Authored-By: Claude Opus 5.5 --- .../controller/testfailover_controller.go | 62 ++++++++- .../testfailover_controller_unit_test.go | 118 ++++++++++++++++- .../controller/testfailover_group_mapping.go | 99 ++++++++++++++ .../testfailover_group_mapping_test.go | 123 ++++++++++++++++++ operator/internal/webapi/group_failover.go | 18 ++- 5 files changed, 406 insertions(+), 14 deletions(-) create mode 100644 operator/internal/controller/testfailover_group_mapping.go create mode 100644 operator/internal/controller/testfailover_group_mapping_test.go diff --git a/operator/internal/controller/testfailover_controller.go b/operator/internal/controller/testfailover_controller.go index 53cab8521..dbc5e8465 100644 --- a/operator/internal/controller/testfailover_controller.go +++ b/operator/internal/controller/testfailover_controller.go @@ -692,10 +692,12 @@ func (r *TestFailoverReconciler) patchStatus(ctx context.Context, tf *simplybloc // replicatedSnapshotResult is the control plane's ReplicatedSnapshotDTO for the // latest replicated snapshot on a DR-target backend. +// LvolID is the replica volume on the target the snapshot belongs to. type replicatedSnapshotResult struct { SnapshotID string `json:"snapshot_id"` ClusterID string `json:"cluster_id"` PoolID string `json:"pool_id"` + LvolID string `json:"lvol_id"` CreatedAt time.Time `json:"created_at"` } @@ -801,12 +803,29 @@ func (r *TestFailoverReconciler) resolvePointGroup(ctx context.Context, tf *simp return r.fail(ctx, tf, fmt.Sprintf("the group generation has %d members but %d were resolved; group membership changed mid-drill", len(members), len(tf.Status.Clones))) } + // Each slot gets the snapshot of ITS member, matched by volume identity: + // one generation makes the set crash-consistent, but only the identity + // keeps one member's disk off another's PVC. + slots := make([]groupMemberSource, len(tf.Status.Clones)) + for i, c := range tf.Status.Clones { + _, _, lvol, ok := splitHandle(c.SourceHandle) + if !ok { + return r.fail(ctx, tf, "group member "+c.SourceRef+" has a malformed source handle: "+c.SourceHandle) + } + slots[i] = groupMemberSource{PVC: c.SourceRef, LvolID: lvol} + } + targetOf, err := r.replicaVolumes(ctx, api, srcUUID, slots, members) + if err != nil { + return ctrl.Result{}, err + } + assigned, err := assignGenerationMembers(slots, members, targetOf) + if err != nil { + return r.fail(ctx, tf, fmt.Sprintf("cannot pair generation %d with the group's PVCs: %v", groupSeq, err)) + } + if err := r.transitionTo(ctx, tf, simplyblockv1alpha2.TestFailoverStepCloning, func(s *simplyblockv1alpha2.TestFailoverStatus) { - // Every member is at one generation, so any one-to-one assignment of the - // generation's snapshots to the clone slots yields a crash-consistent set; - // a precise source-to-target mapping is a later refinement. - for i := range members { - s.Clones[i].SnapshotID = members[i].ClusterID + ":" + members[i].PoolID + ":" + members[i].SnapshotID + for i, m := range assigned { + s.Clones[i].SnapshotID = members[m].ClusterID + ":" + members[m].PoolID + ":" + members[m].SnapshotID } if s.Report == nil { s.Report = &simplyblockv1alpha2.TestFailoverReport{} @@ -822,6 +841,39 @@ func (r *TestFailoverReconciler) resolvePointGroup(ctx context.Context, tf *simp return ctrl.Result{Requeue: true}, nil } +// replicaVolumes maps each slot's source lvol to its replica volume on the +// target, for a control plane whose generation members do not name their source +// volume. It returns nil, and reads nothing, when every member names it. +func (r *TestFailoverReconciler) replicaVolumes( + ctx context.Context, + api *webapi.Client, + sourceClusterUUID string, + slots []groupMemberSource, + members []webapi.ReplicatedGroupSnapshot, +) (map[string]string, error) { + complete := true + for i := range members { + if members[i].SourceLvolID == "" { + complete = false + break + } + } + if complete { + return nil, nil + } + targetOf := make(map[string]string, len(slots)) + for _, slot := range slots { + dto, found, err := r.latestReplicatedSnapshot(ctx, api, sourceClusterUUID, slot.LvolID) + if err != nil { + return nil, err + } + if found { + targetOf[slot.LvolID] = dto.LvolID + } + } + return targetOf, nil +} + // latestReplicatedSnapshot reads the latest replicated snapshot for a source // volume on its DR target. found is false when replication has landed nothing. func (r *TestFailoverReconciler) latestReplicatedSnapshot(ctx context.Context, api *webapi.Client, sourceClusterUUID, sourceLvolUUID string) (replicatedSnapshotResult, bool, error) { diff --git a/operator/internal/controller/testfailover_controller_unit_test.go b/operator/internal/controller/testfailover_controller_unit_test.go index 2df684ed3..701e4c3e5 100644 --- a/operator/internal/controller/testfailover_controller_unit_test.go +++ b/operator/internal/controller/testfailover_controller_unit_test.go @@ -560,9 +560,13 @@ func TestFailoverResolvingPointGroupResolvesGeneration(t *testing.T) { case strings.HasSuffix(p, "/consistency-groups/") && req.URL.Query().Get("name") == "cg": _, _ = w.Write([]byte(`[{"id":"g1","name":"cg","lvs_name":"lvs-a","node_id":"node-a","policy_id":"p1"}]`)) case strings.HasSuffix(p, "/replication/policies/p1/latest-generation"): + // Listed in the reverse order of the drill's PVCs: the pairing must be + // by source volume, not by position. _, _ = w.Write([]byte(`{"group_seq":7,"members":[` + - `{"snapshot_id":"s1","cluster_id":"B","pool_id":"pb","lvol_id":"t1","size":1073741824,"group_seq":7},` + - `{"snapshot_id":"s2","cluster_id":"B","pool_id":"pb","lvol_id":"t2","size":1073741824,"group_seq":7}]}`)) + `{"snapshot_id":"s2","cluster_id":"B","pool_id":"pb","lvol_id":"t2","source_lvol_id":"lvol-b",` + + `"size":1073741824,"group_seq":7},` + + `{"snapshot_id":"s1","cluster_id":"B","pool_id":"pb","lvol_id":"t1","source_lvol_id":"lvol-a",` + + `"size":1073741824,"group_seq":7}]}`)) default: t.Errorf("unexpected request %s %s", req.Method, p) w.WriteHeader(http.StatusInternalServerError) @@ -588,6 +592,116 @@ func TestFailoverResolvingPointGroupResolvesGeneration(t *testing.T) { } } +// groupAtResolvingPoint returns a two-member group drill seeded at +// ResolvingPoint, and the StorageCluster that resolves its source UUID. +func groupAtResolvingPoint() (*simplyblockv1alpha2.TestFailover, *simplyblockv1alpha2.StorageCluster) { + tf := sampleTestFailover() + tf.Finalizers = []string{finalizerTestFailover} + tf.Spec.Scope = simplyblockv1alpha2.TestFailoverScopeGroup + tf.Spec.SourceRef = "cg" + tf.Spec.SourceCluster = testSourceCluster + tf.Status.Phase = simplyblockv1alpha2.TestFailoverPhaseProvisioning + tf.Status.Step = statemachine.KubeSnapshot{State: string(simplyblockv1alpha2.TestFailoverStepResolvingPoint)} + tf.Status.Clones = []simplyblockv1alpha2.TestFailoverClone{ + {SourceRef: "data-1", SourceHandle: "C:pool-1:lvol-a", SourceFSType: testFSTypeXFS, SizeBytes: 1073741824}, + {SourceRef: "data-2", SourceHandle: "C:pool-1:lvol-b", SourceFSType: testFSTypeXFS, SizeBytes: 1073741824}, + } + sc := &simplyblockv1alpha2.StorageCluster{ + ObjectMeta: metav1.ObjectMeta{Name: "local-sc", Namespace: tf.Namespace}, + Status: simplyblockv1alpha2.StorageClusterStatus{UUID: "C"}, + } + return tf, sc +} + +// TestFailoverResolvingPointGroupPairsByReplicaVolume covers a control plane +// whose generation names only the target replica volume: the drill reads each +// member's replica volume from its per-volume latest snapshot and pairs on it. +func TestFailoverResolvingPointGroupPairsByReplicaVolume(t *testing.T) { + tf, sc := groupAtResolvingPoint() + r, cl := newTestFailoverReconciler(t, tf, sc) + ctx := context.Background() + + srv := newAPIServer(t, func(w http.ResponseWriter, req *http.Request) { + w.Header().Set("Content-Type", "application/json") + p := req.URL.Path + switch { + case strings.HasSuffix(p, "/consistency-groups/") && req.URL.Query().Get("name") == "cg": + _, _ = w.Write([]byte(`[{"id":"g1","name":"cg","policy_id":"p1"}]`)) + case strings.HasSuffix(p, "/replication/policies/p1/latest-generation"): + _, _ = w.Write([]byte(`{"group_seq":7,"members":[` + + `{"snapshot_id":"s2","cluster_id":"B","pool_id":"pb","lvol_id":"t2","group_seq":7},` + + `{"snapshot_id":"s1","cluster_id":"B","pool_id":"pb","lvol_id":"t1","group_seq":7}]}`)) + case strings.HasSuffix(p, "/relationships/lvol-a/latest-snapshot"): + _, _ = w.Write([]byte(`{"snapshot_id":"s1","cluster_id":"B","pool_id":"pb","lvol_id":"t1"}`)) + case strings.HasSuffix(p, "/relationships/lvol-b/latest-snapshot"): + _, _ = w.Write([]byte(`{"snapshot_id":"s2","cluster_id":"B","pool_id":"pb","lvol_id":"t2"}`)) + default: + t.Errorf("unexpected request %s %s", req.Method, p) + w.WriteHeader(http.StatusInternalServerError) + } + }) + t.Setenv("SIMPLYBLOCK_WEBAPI_BASE_URL", srv.URL) + + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("reconcile: %v", err) + } + var got simplyblockv1alpha2.TestFailover + if err := cl.Get(ctx, testFailoverRequest(tf).NamespacedName, &got); err != nil { + t.Fatal(err) + } + if got.Status.Step.State != string(simplyblockv1alpha2.TestFailoverStepCloning) { + t.Fatalf("step = %q, want Cloning; message=%q", got.Status.Step.State, got.Status.Message) + } + if got.Status.Clones[0].SnapshotID != "B:pb:s1" || got.Status.Clones[1].SnapshotID != "B:pb:s2" { + t.Errorf("clone snapshot handles = %+v, want data-1=B:pb:s1, data-2=B:pb:s2", got.Status.Clones) + } +} + +// TestFailoverResolvingPointGroupMissingMemberFails covers a generation that +// lacks one PVC's member: the drill fails naming the PVC, never assigning +// another member's snapshot by position. +func TestFailoverResolvingPointGroupMissingMemberFails(t *testing.T) { + tf, sc := groupAtResolvingPoint() + r, cl := newTestFailoverReconciler(t, tf, sc) + ctx := context.Background() + + srv := newAPIServer(t, func(w http.ResponseWriter, req *http.Request) { + w.Header().Set("Content-Type", "application/json") + p := req.URL.Path + switch { + case strings.HasSuffix(p, "/consistency-groups/") && req.URL.Query().Get("name") == "cg": + _, _ = w.Write([]byte(`[{"id":"g1","name":"cg","policy_id":"p1"}]`)) + case strings.HasSuffix(p, "/replication/policies/p1/latest-generation"): + _, _ = w.Write([]byte(`{"group_seq":7,"members":[` + + `{"snapshot_id":"s1","cluster_id":"B","pool_id":"pb","lvol_id":"t1","source_lvol_id":"lvol-a"},` + + `{"snapshot_id":"s9","cluster_id":"B","pool_id":"pb","lvol_id":"t9","source_lvol_id":"lvol-z"}]}`)) + default: + t.Errorf("unexpected request %s %s", req.Method, p) + w.WriteHeader(http.StatusInternalServerError) + } + }) + t.Setenv("SIMPLYBLOCK_WEBAPI_BASE_URL", srv.URL) + + if _, err := r.Reconcile(ctx, testFailoverRequest(tf)); err != nil { + t.Fatalf("reconcile: %v", err) + } + var got simplyblockv1alpha2.TestFailover + if err := cl.Get(ctx, testFailoverRequest(tf).NamespacedName, &got); err != nil { + t.Fatal(err) + } + if got.Status.Phase != simplyblockv1alpha2.TestFailoverPhaseFailed { + t.Fatalf("phase = %q, want Failed; message=%q", got.Status.Phase, got.Status.Message) + } + if !strings.Contains(got.Status.Message, "no snapshot for data-2") { + t.Errorf("message = %q, want it to name data-2", got.Status.Message) + } + for _, c := range got.Status.Clones { + if c.SnapshotID != "" { + t.Errorf("clone %s got snapshot %q despite the failed pairing", c.SourceRef, c.SnapshotID) + } + } +} + // TestFailoverResolvingPointDRTargetNoReplicaFails covers that a target with no // replicated point yet is a terminal failure, not an endless hold. func TestFailoverResolvingPointDRTargetNoReplicaFails(t *testing.T) { diff --git a/operator/internal/controller/testfailover_group_mapping.go b/operator/internal/controller/testfailover_group_mapping.go new file mode 100644 index 000000000..723f1d986 --- /dev/null +++ b/operator/internal/controller/testfailover_group_mapping.go @@ -0,0 +1,99 @@ +package controller + +import ( + "fmt" + "strings" + + "github.com/simplyblock/simplyblock-operator/internal/webapi" +) + +// groupMemberSource is one clone slot of a group drill as the mapping sees it: +// the PVC it recovers and the source lvol it was resolved from. +type groupMemberSource struct { + PVC string + LvolID string +} + +// assignGenerationMembers pairs every clone slot with the generation member that +// holds THAT slot's data, returning the member index per slot. The pairing is by +// identity, never by position: the control plane lists a generation's members in +// no particular order, so a positional pairing can restore one member's disk onto +// another's PVC (the database disk onto the web volume). +// +// The key is the member's source lvol (source_lvol_id) when every member carries +// it. A control plane that does not report it yet returns only the replica volume +// on the target (lvol_id); then targetOf maps each slot's source lvol to its +// replica volume, read from the per-volume latest-snapshot, and the pairing is on +// the replica volume. Anything short of a one-to-one pairing is an error: a slot +// without a member, a member no slot claims, an empty or a duplicate key. +func assignGenerationMembers( + slots []groupMemberSource, + members []webapi.ReplicatedGroupSnapshot, + targetOf map[string]string, +) ([]int, error) { + bySource := true + for i := range members { + if members[i].SourceLvolID == "" { + bySource = false + break + } + } + + memberKey := func(m webapi.ReplicatedGroupSnapshot) string { + if bySource { + return m.SourceLvolID + } + return m.LvolID + } + index := make(map[string]int, len(members)) + for i := range members { + k := memberKey(members[i]) + if k == "" { + return nil, fmt.Errorf("generation member snapshot %s names no volume; cannot tell whose data it holds", + members[i].SnapshotID) + } + if j, dup := index[k]; dup { + return nil, fmt.Errorf("generation members %s and %s both belong to volume %s", + members[j].SnapshotID, members[i].SnapshotID, k) + } + index[k] = i + } + + assigned := make([]int, len(slots)) + used := make(map[int]bool, len(members)) + var missing []string + for s, slot := range slots { + k := slot.LvolID + if !bySource { + k = targetOf[slot.LvolID] + if k == "" { + missing = append(missing, fmt.Sprintf("%s (source volume %s has no replica volume on the target)", + slot.PVC, slot.LvolID)) + continue + } + } + i, ok := index[k] + if !ok { + missing = append(missing, fmt.Sprintf("%s (volume %s)", slot.PVC, k)) + continue + } + if used[i] { + return nil, fmt.Errorf("generation member snapshot %s matches more than one PVC", members[i].SnapshotID) + } + used[i] = true + assigned[s] = i + } + if len(missing) > 0 { + return nil, fmt.Errorf("the generation holds no snapshot for %s", strings.Join(missing, ", ")) + } + var stray []string + for i := range members { + if !used[i] { + stray = append(stray, fmt.Sprintf("%s (volume %s)", members[i].SnapshotID, memberKey(members[i]))) + } + } + if len(stray) > 0 { + return nil, fmt.Errorf("generation snapshots %s match no PVC of the drill", strings.Join(stray, ", ")) + } + return assigned, nil +} diff --git a/operator/internal/controller/testfailover_group_mapping_test.go b/operator/internal/controller/testfailover_group_mapping_test.go new file mode 100644 index 000000000..35887c129 --- /dev/null +++ b/operator/internal/controller/testfailover_group_mapping_test.go @@ -0,0 +1,123 @@ +package controller + +import ( + "strings" + "testing" + + "github.com/simplyblock/simplyblock-operator/internal/webapi" +) + +const ( + mapSnapWeb = "snap-web" + mapSnapDB = "snap-db" + mapTgtWeb = "tgt-web" + mapTgtDB = "tgt-db" + mapSrcWeb = "src-web" + mapSrcDB = "src-db" +) + +// The two members of a web+db group, as the drill resolved them. +var mappingSlots = []groupMemberSource{ + {PVC: "web-disk", LvolID: mapSrcWeb}, + {PVC: "db-disk", LvolID: mapSrcDB}, +} + +// TestAssignGenerationMembersBySourceShuffled covers the regression: the control +// plane lists the generation's members in its own order (db first here), and each +// PVC must still get its own member's snapshot, not the one at its position. +func TestAssignGenerationMembersBySourceShuffled(t *testing.T) { + members := []webapi.ReplicatedGroupSnapshot{ + {SnapshotID: mapSnapDB, LvolID: mapTgtDB, SourceLvolID: mapSrcDB}, + {SnapshotID: mapSnapWeb, LvolID: mapTgtWeb, SourceLvolID: mapSrcWeb}, + } + got, err := assignGenerationMembers(mappingSlots, members, nil) + if err != nil { + t.Fatal(err) + } + if members[got[0]].SnapshotID != mapSnapWeb || members[got[1]].SnapshotID != mapSnapDB { + t.Errorf("web-disk got %s, db-disk got %s; want snap-web, snap-db", + members[got[0]].SnapshotID, members[got[1]].SnapshotID) + } +} + +// TestAssignGenerationMembersByReplicaVolume covers a control plane that names +// only the replica volume on the target: the pairing goes through the source to +// replica map and is still order-independent. +func TestAssignGenerationMembersByReplicaVolume(t *testing.T) { + members := []webapi.ReplicatedGroupSnapshot{ + {SnapshotID: mapSnapDB, LvolID: mapTgtDB}, + {SnapshotID: mapSnapWeb, LvolID: mapTgtWeb}, + } + targetOf := map[string]string{mapSrcWeb: mapTgtWeb, mapSrcDB: mapTgtDB} + got, err := assignGenerationMembers(mappingSlots, members, targetOf) + if err != nil { + t.Fatal(err) + } + if members[got[0]].SnapshotID != mapSnapWeb || members[got[1]].SnapshotID != mapSnapDB { + t.Errorf("web-disk got %s, db-disk got %s; want snap-web, snap-db", + members[got[0]].SnapshotID, members[got[1]].SnapshotID) + } +} + +// TestAssignGenerationMembersRefusesMismatch covers every way the pairing can be +// incomplete: none of them falls back to position. +func TestAssignGenerationMembersRefusesMismatch(t *testing.T) { + cases := []struct { + name string + members []webapi.ReplicatedGroupSnapshot + targetOf map[string]string + want string + }{ + { + name: "missing member", + members: []webapi.ReplicatedGroupSnapshot{ + {SnapshotID: mapSnapWeb, LvolID: mapTgtWeb, SourceLvolID: mapSrcWeb}, + {SnapshotID: "snap-other", LvolID: "tgt-other", SourceLvolID: "src-other"}, + }, + want: "no snapshot for db-disk", + }, + { + name: "stray member", + members: []webapi.ReplicatedGroupSnapshot{ + {SnapshotID: mapSnapWeb, LvolID: mapTgtWeb, SourceLvolID: mapSrcWeb}, + {SnapshotID: mapSnapDB, LvolID: mapTgtDB, SourceLvolID: mapSrcDB}, + {SnapshotID: "snap-x", LvolID: "tgt-x", SourceLvolID: "src-x"}, + }, + want: "snap-x (volume src-x) match no PVC", + }, + { + name: "duplicate member", + members: []webapi.ReplicatedGroupSnapshot{ + {SnapshotID: "snap-a", LvolID: mapTgtWeb, SourceLvolID: mapSrcWeb}, + {SnapshotID: "snap-b", LvolID: mapTgtWeb, SourceLvolID: mapSrcWeb}, + }, + want: "both belong to volume src-web", + }, + { + name: "no replica volume known", + members: []webapi.ReplicatedGroupSnapshot{ + {SnapshotID: mapSnapWeb, LvolID: mapTgtWeb}, + {SnapshotID: mapSnapDB, LvolID: mapTgtDB}, + }, + targetOf: map[string]string{mapSrcWeb: mapTgtWeb}, + want: "db-disk (source volume src-db has no replica volume", + }, + { + name: "member without a volume", + members: []webapi.ReplicatedGroupSnapshot{ + {SnapshotID: mapSnapWeb, LvolID: mapTgtWeb}, + {SnapshotID: mapSnapDB}, + }, + targetOf: map[string]string{mapSrcWeb: mapTgtWeb, mapSrcDB: mapTgtDB}, + want: "snap-db names no volume", + }, + } + for _, tc := range cases { + t.Run(tc.name, func(t *testing.T) { + _, err := assignGenerationMembers(mappingSlots, tc.members, tc.targetOf) + if err == nil || !strings.Contains(err.Error(), tc.want) { + t.Errorf("err = %v, want it to contain %q", err, tc.want) + } + }) + } +} diff --git a/operator/internal/webapi/group_failover.go b/operator/internal/webapi/group_failover.go index 9592f1194..029b7fa6d 100644 --- a/operator/internal/webapi/group_failover.go +++ b/operator/internal/webapi/group_failover.go @@ -21,14 +21,18 @@ const storagePoolsListPathFmt = "/api/v2/clusters/%s/storage-pools/" // ReplicatedGroupSnapshot is one member's replicated snapshot on the target // cluster, at one group-consistent generation. It is the cloneable point for // that member: cluster and pool address the target backend, snapshot is the -// snapshot to clone, and size sizes the recovered PVC. +// snapshot to clone, and size sizes the recovered PVC. LvolID is the replica +// volume on the TARGET the snapshot belongs to, not the source member; +// SourceLvolID is the source member whose data it holds, empty on a control +// plane that does not report it yet. type ReplicatedGroupSnapshot struct { - SnapshotID string `json:"snapshot_id"` - ClusterID string `json:"cluster_id"` - PoolID string `json:"pool_id"` - LvolID string `json:"lvol_id"` - Size int64 `json:"size"` - GroupSeq int `json:"group_seq"` + SnapshotID string `json:"snapshot_id"` + ClusterID string `json:"cluster_id"` + PoolID string `json:"pool_id"` + LvolID string `json:"lvol_id"` + SourceLvolID string `json:"source_lvol_id"` + Size int64 `json:"size"` + GroupSeq int `json:"group_seq"` } // latestGenerationResponse is the group form of latest-snapshot: one generation From 7d31da0774dfea198ca32bb32969929846120401 Mon Sep 17 00:00:00 2001 From: michael Date: Sun, 4 Oct 2026 22:31:19 +0300 Subject: [PATCH 198/206] chart: the Control Center may create the console's S3 and health probe requests The plan form's S3 Test button and the probe Test buttons (console #632) create dr.simplyblock.io S3ProbeRequest / HealthProbeRequest objects that dr-hub answers (simplyblock-dr #5). Co-Authored-By: Claude Opus 5.5 --- .../simplyblock-operator/templates/control-center-rbac.yaml | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/helm-charts/charts/simplyblock-operator/templates/control-center-rbac.yaml b/helm-charts/charts/simplyblock-operator/templates/control-center-rbac.yaml index 29ac17f9b..91967a6f1 100644 --- a/helm-charts/charts/simplyblock-operator/templates/control-center-rbac.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/control-center-rbac.yaml @@ -138,6 +138,10 @@ rules: - apiGroups: ["dr.simplyblock.io"] resources: ["recoveryactions"] verbs: ["override"] + # the forms' Test buttons: dr-hub probes an S3 store, dr-agent health probes, on demand + - apiGroups: ["dr.simplyblock.io"] + resources: ["s3proberequests", "healthproberequests"] + verbs: ["get", "create", "delete"] - apiGroups: ["sitemap.simplyblock.io"] resources: ["siteprofiles", "dhcpservers"] verbs: ["get", "list", "watch"] From e8318531607842f54e17901f228ad85cf2d2e76b Mon Sep 17 00:00:00 2001 From: michael Date: Mon, 5 Oct 2026 20:23:21 +0300 Subject: [PATCH 199/206] chart: the Control Center reads the dr-agent views, creates DHCP probes and label requests; opt-in hub labelling The console's forms offer what each site's dr-agent reports (through the dr-agent-status ManagedClusterViews, by name), ask a network's DHCP servers (dhcpproberequests) and label a site's objects through dr-agent (labelrequests). Patching the hub's own objects is behind controlCenter.rbac.hubLabelling (default off): RBAC cannot narrow a patch to label keys. Same rules as console PR #651. Co-Authored-By: Claude Opus 5.5 --- .../templates/control-center-rbac.yaml | 37 +++++++++++++++++++ .../charts/simplyblock-operator/values.yaml | 5 +++ 2 files changed, 42 insertions(+) diff --git a/helm-charts/charts/simplyblock-operator/templates/control-center-rbac.yaml b/helm-charts/charts/simplyblock-operator/templates/control-center-rbac.yaml index 91967a6f1..af75477c5 100644 --- a/helm-charts/charts/simplyblock-operator/templates/control-center-rbac.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/control-center-rbac.yaml @@ -142,6 +142,23 @@ rules: - apiGroups: ["dr.simplyblock.io"] resources: ["s3proberequests", "healthproberequests"] verbs: ["get", "create", "delete"] + # "Ask" on a guest network: dr-agent asks the NAD's DHCP servers (a + # short-lived pod on the network; dr-hub needs dr-admin for it too) + - apiGroups: ["dr.simplyblock.io"] + resources: ["dhcpproberequests"] + verbs: ["get", "create", "delete"] + # Label for DR on a managed site: dr-agent applies the allow-listed labels; + # the request stays on the hub as the record (no delete) + - apiGroups: ["dr.simplyblock.io"] + resources: ["labelrequests"] + verbs: ["get", "list", "create"] + # what each site's dr-agent reports, through the view dr-hub keeps of it: + # the forms offer discovered values (clusters' zones, Velero, classes, + # namespaces, labels, PVCs, NADs, DHCP servers). Only that view, by name. + - apiGroups: ["view.open-cluster-management.io"] + resources: ["managedclusterviews"] + resourceNames: ["dr-agent-status"] + verbs: ["get"] - apiGroups: ["sitemap.simplyblock.io"] resources: ["siteprofiles", "dhcpservers"] verbs: ["get", "list", "watch"] @@ -161,6 +178,26 @@ rules: - apiGroups: ["cluster.open-cluster-management.io"] resources: ["managedclusters"] verbs: ["get", "list", "watch"] + {{- if .Values.controlCenter.rbac.hubLabelling }} + # Label for DR on the hub itself: the console patches the hub's objects + # directly. RBAC cannot narrow a patch to label keys (the console offers + # only DR's), so this is opt-in (controlCenter.rbac.hubLabelling). + - apiGroups: ["storage.k8s.io"] + resources: ["storageclasses"] + verbs: ["patch"] + - apiGroups: [""] + resources: ["nodes", "persistentvolumeclaims"] + verbs: ["patch"] + - apiGroups: ["apps"] + resources: ["deployments", "statefulsets"] + verbs: ["list", "patch"] + - apiGroups: ["kubevirt.io"] + resources: ["virtualmachines"] + verbs: ["list", "patch"] + - apiGroups: [""] + resources: ["events"] + verbs: ["create"] + {{- end }} # the console asks the API server what its identity may do, to disable controls - apiGroups: ["authorization.k8s.io"] resources: ["selfsubjectaccessreviews", "selfsubjectrulesreviews"] diff --git a/helm-charts/charts/simplyblock-operator/values.yaml b/helm-charts/charts/simplyblock-operator/values.yaml index 276d44491..92f3b2c28 100644 --- a/helm-charts/charts/simplyblock-operator/values.yaml +++ b/helm-charts/charts/simplyblock-operator/values.yaml @@ -1046,6 +1046,11 @@ controlCenter: # every request returns 403 and the console renders empty. Set false only # with authMode: passthrough, where each user's own RBAC applies instead. create: true + # Let "Label for DR" patch the hub's own StorageClasses, nodes, PVCs and + # workloads. RBAC cannot narrow a patch to label keys, so this grants + # patch on those kinds; managed sites are labelled through dr-agent + # (LabelRequest) and need no grant here. + hubLabelling: false service: type: ClusterIP From a10f266241742792055d62f0d27120c5f7a8dfec Mon Sep 17 00:00:00 2001 From: michael Date: Mon, 5 Oct 2026 20:35:36 +0300 Subject: [PATCH 200/206] chart+operator: the Control Center reads storage through the control plane API On a hub whose control plane manages storage clusters on other sites the clusters are not CRDs on the hub, so the console's Clusters and Control plane screens were empty. The console now proxies the management API (read-only, credentials scrubbed in the pod) and this wires it: - controlCenter.controlPlane (enabled, url, trustServiceAccount, tokenSecret): SB_CONTROLPLANE_URL defaults to this release's management API on the standalone profile (https when tls.enabled), off on managed sites. - The operator appends SB_EXTRA_ADMIN_SERVICE_ACCOUNTS (set by the chart to the console's service account) to the management API's admin accounts; entries that are not service account usernames are ignored. - Under mutual TLS the console gets its own client certificate (simplyblock-control-center-client) and the CA; a static admin token Secret can be used instead. - The console's NetworkPolicy allows the management API port. Co-Authored-By: Claude Opus 5.5 --- .../templates/_control_center_helpers.tpl | 21 +++++++++++ .../control-center-networkpolicy.yaml | 3 ++ .../templates/control-center.yaml | 36 +++++++++++++++++++ .../templates/controlplane_certificates.yaml | 20 +++++++++++ .../templates/simplyblock-operator.yaml | 7 ++++ .../charts/simplyblock-operator/values.yaml | 17 +++++++++ .../controllers/controlplane/managementapi.go | 28 +++++++++++++-- .../controlplane/workloads_test.go | 31 ++++++++++++++++ 8 files changed, 161 insertions(+), 2 deletions(-) diff --git a/helm-charts/charts/simplyblock-operator/templates/_control_center_helpers.tpl b/helm-charts/charts/simplyblock-operator/templates/_control_center_helpers.tpl index e7ed59f80..71089375d 100644 --- a/helm-charts/charts/simplyblock-operator/templates/_control_center_helpers.tpl +++ b/helm-charts/charts/simplyblock-operator/templates/_control_center_helpers.tpl @@ -61,6 +61,27 @@ http://simplyblock-operator:8080 {{- end -}} {{- end -}} +{{/* The control plane API the console reads storage from. Fully qualified: + nginx resolves it per request without the pod's search domains. */}} +{{- define "sbcc.controlPlaneUrl" -}} +{{- $cp := .Values.controlCenter.controlPlane | default dict -}} +{{- if and $cp.enabled (not .Values.controlCenter.mock.enabled) -}} +{{- if $cp.url -}} +{{- $cp.url -}} +{{- else if eq .Values.deployment.profile "standalone" -}} +{{- printf "%s://simplyblock-webappapi.%s.svc.cluster.local:5000" (ternary "https" "http" .Values.tls.enabled) .Release.Namespace -}} +{{- end -}} +{{- end -}} +{{- end -}} + +{{/* The service account the operator adds to the management API's admins. */}} +{{- define "sbcc.trustedAccount" -}} +{{- $cp := .Values.controlCenter.controlPlane | default dict -}} +{{- if and .Values.controlCenter.enabled $cp.trustServiceAccount (include "sbcc.controlPlaneUrl" .) (eq .Values.controlCenter.authMode "serviceaccount") (not $cp.tokenSecret) -}} +{{- printf "system:serviceaccount:%s:%s" .Release.Namespace (include "sbcc.fullname" .) -}} +{{- end -}} +{{- end -}} + {{- define "sbcc.prometheusUrl" -}} {{- if .Values.controlCenter.prometheusUrl -}} {{- .Values.controlCenter.prometheusUrl -}} diff --git a/helm-charts/charts/simplyblock-operator/templates/control-center-networkpolicy.yaml b/helm-charts/charts/simplyblock-operator/templates/control-center-networkpolicy.yaml index 969fabc3e..7efbffe29 100644 --- a/helm-charts/charts/simplyblock-operator/templates/control-center-networkpolicy.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/control-center-networkpolicy.yaml @@ -52,6 +52,9 @@ spec: port: 8080 - protocol: TCP port: 9090 + # the control plane API (management API), read-only via the proxy + - protocol: TCP + port: 5000 # The Kubernetes API server. Narrow these CIDRs — see the header comment. - to: {{- range $cidr := $cc.networkPolicy.apiServerCidrs }} diff --git a/helm-charts/charts/simplyblock-operator/templates/control-center.yaml b/helm-charts/charts/simplyblock-operator/templates/control-center.yaml index 3a81f5e3b..84ff20fbc 100644 --- a/helm-charts/charts/simplyblock-operator/templates/control-center.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/control-center.yaml @@ -107,6 +107,23 @@ spec: value: {{ ternary $mockURL ($cc.helmUrl | default (include "sbcc.operatorUrl" .)) $cc.mock.enabled | quote }} - name: SB_PROMETHEUS_URL value: {{ ternary $mockURL (include "sbcc.prometheusUrl" .) $cc.mock.enabled | quote }} + {{- $cpURL := include "sbcc.controlPlaneUrl" . }} + - name: SB_CONTROLPLANE_URL + value: {{ $cpURL | quote }} + {{- if and $cpURL $cc.controlPlane.tokenSecret }} + - name: SB_CONTROLPLANE_TOKEN_FILE + value: /etc/simplyblock/controlplane-token/token + {{- end }} + {{- if and $cpURL .Values.tls.enabled }} + - name: SB_CONTROLPLANE_CA_FILE + value: /etc/simplyblock/tls/ca.crt + {{- if .Values.tls.mutual_enabled }} + - name: SB_CONTROLPLANE_CLIENT_CERT + value: /etc/simplyblock/tls/tls.crt + - name: SB_CONTROLPLANE_CLIENT_KEY + value: /etc/simplyblock/tls/tls.key + {{- end }} + {{- end }} - name: SB_MOCK value: "false" - name: SB_LISTEN_PORT @@ -137,11 +154,30 @@ spec: mountPath: /tmp/nginx - name: nginx-confd mountPath: /etc/nginx/conf.d + {{- if and $cpURL $cc.controlPlane.tokenSecret }} + - name: controlplane-token + mountPath: /etc/simplyblock/controlplane-token + readOnly: true + {{- end }} + {{- if $cpURL }} + {{- include "simplyblock.tlsVolumeMount" . | nindent 12 }} + {{- end }} volumes: - name: nginx-tmp emptyDir: {medium: Memory, sizeLimit: 16Mi} - name: nginx-confd emptyDir: {medium: Memory, sizeLimit: 1Mi} + {{- if and $cpURL $cc.controlPlane.tokenSecret }} + - name: controlplane-token + secret: + secretName: {{ $cc.controlPlane.tokenSecret }} + items: + - key: token + path: token + {{- end }} + {{- if $cpURL }} + {{- include "simplyblock.clientTlsVolume" (dict "ctx" . "clientSecret" "simplyblock-control-center-client-tls") | nindent 8 }} + {{- end }} --- apiVersion: v1 kind: Service diff --git a/helm-charts/charts/simplyblock-operator/templates/controlplane_certificates.yaml b/helm-charts/charts/simplyblock-operator/templates/controlplane_certificates.yaml index 78205a832..059afadc1 100644 --- a/helm-charts/charts/simplyblock-operator/templates/controlplane_certificates.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/controlplane_certificates.yaml @@ -124,5 +124,25 @@ spec: - digital signature - key encipherment - client auth +{{- if and .Values.controlCenter.enabled (include "sbcc.controlPlaneUrl" .) }} +--- +# The Control Center's proxy presents this to the management API when it reads +# storage through it (controlCenter.controlPlane). +apiVersion: cert-manager.io/v1 +kind: Certificate +metadata: + name: simplyblock-control-center-client + namespace: {{ .Release.Namespace }} +spec: + commonName: simplyblock-control-center + secretName: simplyblock-control-center-client-tls + issuerRef: + kind: ClusterIssuer + name: simplyblock-certificate-authority-issuer + usages: + - digital signature + - key encipherment + - client auth +{{- end }} {{- end }} {{- end }} diff --git a/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator.yaml b/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator.yaml index e1d3746f2..cab838eff 100644 --- a/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/simplyblock-operator.yaml @@ -96,6 +96,13 @@ spec: # pod for the reason above. - name: SB_OPERATOR_IMAGE value: "{{ .Values.image.operator.repository }}:{{ .Values.image.operator.tag }}" + {{- with include "sbcc.trustedAccount" . }} + # The Control Center reads storage through the management API with + # its own service account; the management API trusts it next to + # this operator (controlCenter.controlPlane.trustServiceAccount). + - name: SB_EXTRA_ADMIN_SERVICE_ACCOUNTS + value: {{ . | quote }} + {{- end }} {{- include "simplyblock.tlsEnv" . | nindent 12 }} resources: limits: diff --git a/helm-charts/charts/simplyblock-operator/values.yaml b/helm-charts/charts/simplyblock-operator/values.yaml index 276d44491..25e8bbd1c 100644 --- a/helm-charts/charts/simplyblock-operator/values.yaml +++ b/helm-charts/charts/simplyblock-operator/values.yaml @@ -1041,6 +1041,23 @@ controlCenter: helmUrl: "" prometheusUrl: "" + # The control plane API, read-only, behind the console's own proxy. On a hub + # whose control plane manages storage clusters on other sites the clusters + # are not CRDs here, and this is where the Clusters and Control plane screens + # read them from. Every response is scrubbed of credentials before it leaves + # the console's pod, and the proxy refuses anything but GET. + controlPlane: + enabled: true + # Empty: this release's management API when deployment.profile is + # standalone (https when tls.enabled), and off otherwise. + url: "" + # Have the operator add the console's service account to the management + # API's admin accounts (serviceaccount mode). The console proxies only GETs. + trustServiceAccount: true + # Or a Secret (key: token) holding one of the control plane's static admin + # tokens (controlplane.local.adminTokenSecretRef); it wins when set. + tokenSecret: "" + rbac: # Required in serviceaccount mode: the proxied token needs these rules or # every request returns 403 and the console renders empty. Set false only diff --git a/operator/internal/controllers/controlplane/managementapi.go b/operator/internal/controllers/controlplane/managementapi.go index ee1905e7e..a26fadcea 100644 --- a/operator/internal/controllers/controlplane/managementapi.go +++ b/operator/internal/controllers/controlplane/managementapi.go @@ -17,6 +17,10 @@ package controlplane import ( + "os" + "slices" + "strings" + appsv1 "k8s.io/api/apps/v1" corev1 "k8s.io/api/core/v1" rbacv1 "k8s.io/api/rbac/v1" @@ -269,8 +273,7 @@ func webAPIDeployment(cp *simplyblockv1alpha2.ControlPlane) *appsv1.Deployment { // The operator's own account is the one caller the control plane has to // trust before anything else works: every controller in this process // authenticates with its token. - {Name: "SB_K8S_ADMIN_SERVICE_ACCOUNTS", - Value: "system:serviceaccount:" + cp.Namespace + ":simplyblock-operator"}, + {Name: "SB_K8S_ADMIN_SERVICE_ACCOUNTS", Value: adminServiceAccounts(cp.Namespace)}, {Name: "SB_K8S_METRICS_SERVICE_ACCOUNTS", Value: "system:serviceaccount:" + cp.Namespace + ":simplyblock-prometheus"}, } @@ -666,3 +669,24 @@ func fdbExporterService(namespace string) *corev1.Service { func intstrFromInt(port int32) intstr.IntOrString { return intstr.FromInt32(port) } + +// extraAdminAccountsEnv names the operator's own environment variable listing +// further service accounts the management API must trust as administrators: +// the Control Center's, which reads every storage cluster through the control +// plane API on a hub whose clusters are not CRDs here. The chart sets it; an +// operator without it trusts only itself, as before. +const extraAdminAccountsEnv = "SB_EXTRA_ADMIN_SERVICE_ACCOUNTS" + +// adminServiceAccounts is the operator's own account followed by the extra +// ones. An entry that is not a service account username is ignored rather than +// passed on, so a typo can only narrow who is trusted, never widen it. +func adminServiceAccounts(namespace string) string { + accounts := []string{"system:serviceaccount:" + namespace + ":simplyblock-operator"} + for _, a := range strings.Split(os.Getenv(extraAdminAccountsEnv), ",") { + a = strings.TrimSpace(a) + if strings.Count(a, ":") == 3 && strings.HasPrefix(a, "system:serviceaccount:") && !slices.Contains(accounts, a) { + accounts = append(accounts, a) + } + } + return strings.Join(accounts, ",") +} diff --git a/operator/internal/controllers/controlplane/workloads_test.go b/operator/internal/controllers/controlplane/workloads_test.go index 3b158153f..d2d6fcb47 100644 --- a/operator/internal/controllers/controlplane/workloads_test.go +++ b/operator/internal/controllers/controlplane/workloads_test.go @@ -599,3 +599,34 @@ func podSpecOf(obj client.Object) *corev1.PodSpec { return nil } } + +// The Control Center reads every storage cluster through the management API +// with its own service account; the chart names it in the operator's +// environment and the management API trusts it next to the operator. +func TestExtraAdminServiceAccountsReachTheManagementAPI(t *testing.T) { + t.Setenv(extraAdminAccountsEnv, + " system:serviceaccount:simplyblock:console , not-an-account,system:serviceaccount:a:b:c,"+ + "system:serviceaccount:simplyblock:console") + cp := localControlPlane() + + api := findDeployment(t, managementAPIObjects(cp), ComponentWebAPI) + env := findEnvVar(t, api, "SB_K8S_ADMIN_SERVICE_ACCOUNTS") + + want := "system:serviceaccount:" + cp.Namespace + ":simplyblock-operator,system:serviceaccount:simplyblock:console" + if env.Value != want { + t.Errorf("SB_K8S_ADMIN_SERVICE_ACCOUNTS = %q, want %q", env.Value, want) + } +} + +// Without the operator's variable the management API trusts the operator alone. +func TestNoExtraAdminServiceAccountsMeansTheOperatorAlone(t *testing.T) { + t.Setenv(extraAdminAccountsEnv, "") + cp := localControlPlane() + + api := findDeployment(t, managementAPIObjects(cp), ComponentWebAPI) + env := findEnvVar(t, api, "SB_K8S_ADMIN_SERVICE_ACCOUNTS") + + if want := "system:serviceaccount:" + cp.Namespace + ":simplyblock-operator"; env.Value != want { + t.Errorf("SB_K8S_ADMIN_SERVICE_ACCOUNTS = %q, want %q", env.Value, want) + } +} From 29411104c9e52d195e28aa9cc195e2a4b2c5587d Mon Sep 17 00:00:00 2001 From: michael Date: Mon, 5 Oct 2026 20:39:39 +0300 Subject: [PATCH 201/206] test(node): main's provisioner fixture carries a Workload addParams reads the node address through Workload on this branch (managed control planes); main's awaitworker fixture predates that and panicked on a nil Workload in TestThePostedAddsTaskIsRecordedOnTheNode. Co-Authored-By: Claude Opus 5.5 --- operator/internal/controllers/node/awaitworker_test.go | 1 + 1 file changed, 1 insertion(+) diff --git a/operator/internal/controllers/node/awaitworker_test.go b/operator/internal/controllers/node/awaitworker_test.go index ffb3ac496..5f6d45cb9 100644 --- a/operator/internal/controllers/node/awaitworker_test.go +++ b/operator/internal/controllers/node/awaitworker_test.go @@ -72,6 +72,7 @@ func aProvisioner(t *testing.T, worker *corev1.Node, node *simplyblockv1alpha2.S Scheme: scheme, Recorder: events.NewFakeRecorder(64), API: countingBackend{adds: &adds}, + Workload: &Workload{Client: apiClient}, }, cluster, apiClient, &adds } From d47e2a08b48f588a7b85e5e30674c0137ca55371 Mon Sep 17 00:00:00 2001 From: michael Date: Mon, 5 Oct 2026 20:53:26 +0300 Subject: [PATCH 202/206] operator: regenerate dist/install.yaml after the main merge Carries the regenerated StorageSiteDeployment and TestFailover CRDs (make build-installer, kustomize v5.7.1). Co-Authored-By: Claude Opus 5.5 --- operator/dist/install.yaml | 36 ++++++++++++++++++++++++++++-------- 1 file changed, 28 insertions(+), 8 deletions(-) diff --git a/operator/dist/install.yaml b/operator/dist/install.yaml index f225f07f8..7669c13b5 100644 --- a/operator/dist/install.yaml +++ b/operator/dist/install.yaml @@ -11854,11 +11854,6 @@ spec: https://s3.example.com. pattern: ^https?://[a-zA-Z0-9.-]+(:[0-9]{1,5})?(/.*)?$ type: string - prefix: - description: |- - Prefix narrows the store to one key prefix, so that several clusters can - share a bucket without each walking the others' backups. - type: string region: description: Region is the bucket's region, for endpoints that do not imply one. @@ -12435,10 +12430,11 @@ spec: on these nodes. properties: count: - description: Count is the number of journal managers - to configure. + description: |- + Count is the number of journal managers to configure. The control plane + requires at least 3. format: int32 - minimum: 1 + minimum: 3 type: integer percentPerDevice: description: PercentPerDevice is the share of @@ -12988,6 +12984,30 @@ spec: step: description: Step is the position of the running drill's state machine. properties: + claim: + description: |- + Claim records that this state's side effect was started. Absent means + no pass has started it since the state was entered. + properties: + attempt: + description: Attempt counts the claims taken on this state, + starting at 1. + format: int32 + type: integer + leaseUntil: + description: |- + LeaseUntil is when the claim expires and the side effect may be fired + again. + format: date-time + type: string + state: + description: State is the state the claim was taken in. + type: string + required: + - attempt + - leaseUntil + - state + type: object deadline: description: |- Deadline is when that state expires, absent when it has none. It is an From a841327bcd0c0c51cf61482c3996c1632272df6a Mon Sep 17 00:00:00 2001 From: michael Date: Mon, 5 Oct 2026 21:37:32 +0200 Subject: [PATCH 203/206] fix(operator): tasks-runner-backup-merge runs sbcli main's backup_merge_service.py sbcli main's task-runner rework (#1226) ships the backup-merge service as backup_merge_service.py. 5bb7084e pointed the container at the older integrate_csi_addons_p0 name (tasks_runner_backup_merge.py); p0 went into sbcli main on 2026-10-05, and a control plane on sbcli main crash-looped the container (python3 exit 2: can't open file), leaving the ControlPlane Degraded. Co-Authored-By: Claude Opus 5.5 --- .../controllers/controlplane/managementapi.go | 2 +- .../controllers/controlplane/workloads_test.go | 16 ++++++++-------- 2 files changed, 9 insertions(+), 9 deletions(-) diff --git a/operator/internal/controllers/controlplane/managementapi.go b/operator/internal/controllers/controlplane/managementapi.go index a26fadcea..1cd8b7b6e 100644 --- a/operator/internal/controllers/controlplane/managementapi.go +++ b/operator/internal/controllers/controlplane/managementapi.go @@ -425,7 +425,7 @@ func taskServices() []service { {name: "tasks-runner-node-removal", module: "simplyblock_core/services/tasks_runner_node_removal.py"}, {name: "tasks-runner-snapshot-replication", module: "simplyblock_core/services/snapshot_replication.py"}, {name: "tasks-runner-backup", module: "simplyblock_core/services/tasks_runner_backup.py"}, - {name: "tasks-runner-backup-merge", module: "simplyblock_core/services/tasks_runner_backup_merge.py"}, + {name: "tasks-runner-backup-merge", module: "simplyblock_core/services/backup_merge_service.py"}, {name: "tasks-runner-replication-final", module: "simplyblock_core/services/tasks_runner_replication_final.py"}, } } diff --git a/operator/internal/controllers/controlplane/workloads_test.go b/operator/internal/controllers/controlplane/workloads_test.go index d2d6fcb47..aa24ea213 100644 --- a/operator/internal/controllers/controlplane/workloads_test.go +++ b/operator/internal/controllers/controlplane/workloads_test.go @@ -450,17 +450,17 @@ func TestTheServicePoolsRunWhatTheyDeclare(t *testing.T) { } } -// Regression: 2026-10-01 — the tasks-runner-backup-merge container named a module -// that does not exist in the control-plane image (backup_merge_service.py); the -// real service is tasks_runner_backup_merge.py, like every other tasks-runner-*. -// The container crash-looped ("python3: can't open file -// '/app/simplyblock_core/services/backup_merge_service.py'"), which pinned the -// whole tasks pod in CrashLoopBackOff. A .py-suffix check does not catch it, so -// pin the real module name. +// Regression: the tasks-runner-backup-merge container must name the module the +// control-plane image ships, or it crash-loops ("python3: can't open file") and +// pins the whole tasks pod. sbcli main's task-runner rework (#1226) ships it as +// backup_merge_service.py; the older integrate_csi_addons_p0 image named it +// tasks_runner_backup_merge.py (2026-10-01). The control plane now runs sbcli +// main (integrate_csi_addons_p0 merged into it on 2026-10-05). A .py-suffix +// check does not catch a wrong name, so pin it. func TestTheBackupMergeRunnerNamesItsRealModule(t *testing.T) { cp := localControlPlane() d := findDeployment(t, managementAPIObjects(cp), ComponentTasks) - const want = "simplyblock_core/services/tasks_runner_backup_merge.py" + const want = "simplyblock_core/services/backup_merge_service.py" found := false for _, container := range d.Spec.Template.Spec.Containers { if container.Name != "tasks-runner-backup-merge" { From da540a49d1b8bf9afa43d10204f9f290a5daefea Mon Sep 17 00:00:00 2001 From: michael Date: Mon, 5 Oct 2026 22:12:46 +0200 Subject: [PATCH 204/206] chart: the Control Center's log store (Graylog search, read-only) controlCenter.logStore: SB_GRAYLOG_URL defaults to the release's Graylog when controlplane.observability.enabled, with the observability stack's secret (simplyblock-grafana-secrets / MONITORING_SECRET, Graylog's root password) mounted for the console's proxy; a URL, password Secret/key or access-token Secret override it. Egress to 9000 in the console's NetworkPolicy. Same templates as the console branch (simplyblock-operator #657). Co-Authored-By: Claude Opus 5.5 --- .../templates/_control_center_helpers.tpl | 26 ++++++++++++++ .../control-center-networkpolicy.yaml | 3 ++ .../templates/control-center.yaml | 35 +++++++++++++++++++ .../charts/simplyblock-operator/values.yaml | 16 +++++++++ 4 files changed, 80 insertions(+) diff --git a/helm-charts/charts/simplyblock-operator/templates/_control_center_helpers.tpl b/helm-charts/charts/simplyblock-operator/templates/_control_center_helpers.tpl index 71089375d..d2ec36174 100644 --- a/helm-charts/charts/simplyblock-operator/templates/_control_center_helpers.tpl +++ b/helm-charts/charts/simplyblock-operator/templates/_control_center_helpers.tpl @@ -74,6 +74,32 @@ http://simplyblock-operator:8080 {{- end -}} {{- end -}} +{{/* The log store the console searches: Graylog of this release's + observability stack, unless an explicit URL is given. Fully qualified for + the same reason as the control plane URL. */}} +{{- define "sbcc.graylogUrl" -}} +{{- $ls := .Values.controlCenter.logStore | default dict -}} +{{- if and $ls.enabled (not .Values.controlCenter.mock.enabled) -}} +{{- if $ls.url -}} +{{- $ls.url -}} +{{- else if ((.Values.controlplane).observability).enabled -}} +{{- printf "http://simplyblock-graylog.%s.svc.cluster.local:9000" .Release.Namespace -}} +{{- end -}} +{{- end -}} +{{- end -}} + +{{/* The Secret and key holding the log store's password: the given one, or + the observability stack's own secret (its root password is + controlplane.observability.secret, kept in simplyblock-grafana-secrets). */}} +{{- define "sbcc.graylogPasswordSecret" -}} +{{- $ls := .Values.controlCenter.logStore | default dict -}} +{{- $ls.passwordSecret | default "simplyblock-grafana-secrets" -}} +{{- end -}} +{{- define "sbcc.graylogPasswordKey" -}} +{{- $ls := .Values.controlCenter.logStore | default dict -}} +{{- $ls.passwordKey | default "MONITORING_SECRET" -}} +{{- end -}} + {{/* The service account the operator adds to the management API's admins. */}} {{- define "sbcc.trustedAccount" -}} {{- $cp := .Values.controlCenter.controlPlane | default dict -}} diff --git a/helm-charts/charts/simplyblock-operator/templates/control-center-networkpolicy.yaml b/helm-charts/charts/simplyblock-operator/templates/control-center-networkpolicy.yaml index 7efbffe29..110e41077 100644 --- a/helm-charts/charts/simplyblock-operator/templates/control-center-networkpolicy.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/control-center-networkpolicy.yaml @@ -55,6 +55,9 @@ spec: # the control plane API (management API), read-only via the proxy - protocol: TCP port: 5000 + # the log store (Graylog search API), read-only via the proxy + - protocol: TCP + port: 9000 # The Kubernetes API server. Narrow these CIDRs — see the header comment. - to: {{- range $cidr := $cc.networkPolicy.apiServerCidrs }} diff --git a/helm-charts/charts/simplyblock-operator/templates/control-center.yaml b/helm-charts/charts/simplyblock-operator/templates/control-center.yaml index 84ff20fbc..f32c90ef0 100644 --- a/helm-charts/charts/simplyblock-operator/templates/control-center.yaml +++ b/helm-charts/charts/simplyblock-operator/templates/control-center.yaml @@ -124,6 +124,21 @@ spec: value: /etc/simplyblock/tls/tls.key {{- end }} {{- end }} + {{- $glURL := include "sbcc.graylogUrl" . }} + # the log store (Graylog search API), read-only via the proxy + - name: SB_GRAYLOG_URL + value: {{ $glURL | quote }} + {{- if $glURL }} + - name: SB_GRAYLOG_USER + value: {{ $cc.logStore.user | default "admin" | quote }} + {{- if $cc.logStore.tokenSecret }} + - name: SB_GRAYLOG_TOKEN_FILE + value: /etc/simplyblock/graylog/token + {{- else }} + - name: SB_GRAYLOG_PASSWORD_FILE + value: /etc/simplyblock/graylog/password + {{- end }} + {{- end }} - name: SB_MOCK value: "false" - name: SB_LISTEN_PORT @@ -162,6 +177,11 @@ spec: {{- if $cpURL }} {{- include "simplyblock.tlsVolumeMount" . | nindent 12 }} {{- end }} + {{- if $glURL }} + - name: graylog-credentials + mountPath: /etc/simplyblock/graylog + readOnly: true + {{- end }} volumes: - name: nginx-tmp emptyDir: {medium: Memory, sizeLimit: 16Mi} @@ -178,6 +198,21 @@ spec: {{- if $cpURL }} {{- include "simplyblock.clientTlsVolume" (dict "ctx" . "clientSecret" "simplyblock-control-center-client-tls") | nindent 8 }} {{- end }} + {{- if $glURL }} + - name: graylog-credentials + secret: + {{- if $cc.logStore.tokenSecret }} + secretName: {{ $cc.logStore.tokenSecret }} + items: + - key: token + path: token + {{- else }} + secretName: {{ include "sbcc.graylogPasswordSecret" . }} + items: + - key: {{ include "sbcc.graylogPasswordKey" . }} + path: password + {{- end }} + {{- end }} --- apiVersion: v1 kind: Service diff --git a/helm-charts/charts/simplyblock-operator/values.yaml b/helm-charts/charts/simplyblock-operator/values.yaml index 6d6016ee0..a84c36806 100644 --- a/helm-charts/charts/simplyblock-operator/values.yaml +++ b/helm-charts/charts/simplyblock-operator/values.yaml @@ -1074,6 +1074,22 @@ controlCenter: # tokens (controlplane.local.adminTokenSecretRef); it wins when set. tokenSecret: "" + # The log store the Logs view searches: Graylog's search API (the shipped + # logs of every pod fluent-bit collects), read-only through the console's + # proxy, credentials attached server-side, responses scrubbed of credentials. + # Without it the Logs view falls back to a pod's live tail. + logStore: + enabled: true + # Empty: this release's Graylog when controlplane.observability.enabled. + url: "" + user: admin + # The Secret and key holding the user's password. Empty: the observability + # stack's own (simplyblock-grafana-secrets / MONITORING_SECRET). + passwordSecret: "" + passwordKey: "" + # Or a Secret (key: token) holding a Graylog access token; it wins when set. + tokenSecret: "" + rbac: # Required in serviceaccount mode: the proxied token needs these rules or # every request returns 403 and the console renders empty. Set false only From 6e2582207fd81da477d6cddf0ca80429a67395a3 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Tue, 6 Oct 2026 11:55:38 +0200 Subject: [PATCH 205/206] fix(operator): the admin service account tests are declared once The merge of main brought adminaccounts_test.go from #512, and this branch already carried the same two tests in workloads_test.go. The package redeclared TestNoExtraAdminServiceAccountsMeansTheOperatorAlone and its test binary did not compile. The workloads_test.go pair stays: it asserts the same inputs through the rendered management API Deployment rather than the helper alone. Co-Authored-By: Claude Opus 5.5 --- .../controlplane/adminaccounts_test.go | 24 ------------------- 1 file changed, 24 deletions(-) delete mode 100644 operator/internal/controllers/controlplane/adminaccounts_test.go diff --git a/operator/internal/controllers/controlplane/adminaccounts_test.go b/operator/internal/controllers/controlplane/adminaccounts_test.go deleted file mode 100644 index 129914ed7..000000000 --- a/operator/internal/controllers/controlplane/adminaccounts_test.go +++ /dev/null @@ -1,24 +0,0 @@ -package controlplane - -import "testing" - -// The Control Center reads every storage cluster through the management API -// with its own service account; the chart names it in the operator's -// environment and the management API trusts it next to the operator. -func TestExtraAdminServiceAccountsAreAppendedOnceAndValidated(t *testing.T) { - t.Setenv(extraAdminAccountsEnv, - " system:serviceaccount:simplyblock:console , not-an-account,system:serviceaccount:a:b:c,"+ - "system:serviceaccount:simplyblock:console") - want := "system:serviceaccount:sb:simplyblock-operator,system:serviceaccount:simplyblock:console" - if got := adminServiceAccounts("sb"); got != want { - t.Errorf("adminServiceAccounts = %q, want %q", got, want) - } -} - -// Without the operator's variable the management API trusts the operator alone. -func TestNoExtraAdminServiceAccountsMeansTheOperatorAlone(t *testing.T) { - t.Setenv(extraAdminAccountsEnv, "") - if got, want := adminServiceAccounts("sb"), "system:serviceaccount:sb:simplyblock-operator"; got != want { - t.Errorf("adminServiceAccounts = %q, want %q", got, want) - } -} From 39af1c22875161518d50a4a25bc5e708e28a8ab9 Mon Sep 17 00:00:00 2001 From: "Christoph Engelbert (noctarius)" Date: Tue, 6 Oct 2026 11:54:38 +0200 Subject: [PATCH 206/206] chore(fleet): regenerate the DriverDeployment CRD for the csi-addons sidecar DriverDeployment embeds the SimplyblockDriver spec, and the csiAddons sidecar image field with its quay.io/csiaddons pattern was added there without regenerating the fleet CRD. make build produces this diff on the branch alone. The chart-sync check does not cover fleet/, so CI did not report it. Co-Authored-By: Claude Opus 5.5 --- ...leet.simplyblock.io_driverdeployments.yaml | 20 +++++++++++++++++-- 1 file changed, 18 insertions(+), 2 deletions(-) diff --git a/fleet/config/crd/bases/fleet.simplyblock.io_driverdeployments.yaml b/fleet/config/crd/bases/fleet.simplyblock.io_driverdeployments.yaml index a2331a6a9..41c615170 100644 --- a/fleet/config/crd/bases/fleet.simplyblock.io_driverdeployments.yaml +++ b/fleet/config/crd/bases/fleet.simplyblock.io_driverdeployments.yaml @@ -298,13 +298,29 @@ spec: type: object sidecarImages: description: |- - SidecarImages overrides the six CSI sidecars, one field each. Unset takes - the version this operator release ships. + SidecarImages overrides the seven CSI sidecars, one field each. Unset + takes the version this operator release ships. properties: attacher: description: Attacher is csi-attacher, on the controller plugin. pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ type: string + csiAddons: + description: |- + CSIAddons is the kubernetes-csi-addons sidecar, on the controller + plugin. It connects to the plugin's socket, probes the csi-addons + Identity service for capabilities, and publishes a CSIAddonsNode so the + kubernetes-csi-addons controller-manager (design + design-csi-addons-replication.md §4.1) can reach the Replication + service this driver serves. + + Unlike the other sidecars above, this one's upstream home is the + csi-addons project's own registry, not simplyblock's: the allowlist + carries quay.io/csiaddons alongside the simplyblock registries so a + deployment can run the stock kubernetes-csi-addons sidecar image + directly, ahead of (or instead of) a quay.io/simplyblock-io mirror. + pattern: ^($|(quay\.io/simplyblock-io|docker\.io/simplyblock|public\.ecr\.aws/simply-block|quay\.io/csiaddons)/[a-z0-9][a-z0-9._-]*:[a-zA-Z0-9][a-zA-Z0-9._-]*(@sha256:[a-f0-9]{64})?)$ + type: string healthMonitor: description: |- HealthMonitor is csi-external-health-monitor-controller, on the